{"case_id":"12860774","cluster":"DISTRACTOR-CASSANDRA-10233","comments":[{"body":"We haven't seen this in our upgrade suite. Can you provide any further information as to your environment, what version did you upgrade from? Any non-default cassandra yaml settings? Can you procide your (sanitized) schema, and any other relevant logs?\n\n/cc [~krummas] if he has any hint as to the problem.","created":"2015-09-02T16:10:01.043+0000"},{"body":"Upgraded from 2.1.8. I upgraded today to 2.2.1 and still see the issue.\n\nI am not sure which table is causing this issue and we have quite a few tables so pasting the whole schema would be impossible. If you have any pointers on how to narrow to a single table (or even a single keyspace) I'd be happy to provide it's schema.\n\nThe only changes to cassandra.yaml are increasing the write timeout and increasing the size of out of heap memtable to 6GB.","created":"2015-09-02T22:19:28.148+0000"},{"body":"In fact I don't believe that this problem is related exclusively to versions 2.2.X.\nI'm facing the very same problem here. We have upgrade from version 2.1.5 to 2.1.8, after migration we noticed the exception above in the logs. \n\nLooks like something related to hinded hand off, according to the staktrace:\nHintedHandOffManager.java:515\n\nHere is the stack for the 2.1.8 version:\n{code}\nERROR [HintedHandoff:2] 2015-09-30 16:48:38,729 CassandraDaemon.java:223 - Exception in thread Thread[HintedHandoff:2,1,main]\njava.lang.IndexOutOfBoundsException: null\n at java.nio.Buffer.checkIndex(Buffer.java:546) ~[na:1.8.0_45]\n at java.nio.HeapByteBuffer.getLong(HeapByteBuffer.java:416) ~[na:1.8.0_45]\n at org.apache.cassandra.utils.UUIDGen.getUUID(UUIDGen.java:106) ~[apache-cassandra-2.1.8.jar:2.1.8]\n at org.apache.cassandra.db.HintedHandOffManager.scheduleAllDeliveries(HintedHandOffManager.java:514) ~[apache-cassandra-2.1.8.jar:2.1.8]\n at org.apache.cassandra.db.HintedHandOffManager.access$000(HintedHandOffManager.java:88) ~[apache-cassandra-2.1.8.jar:2.1.8]\n at org.apache.cassandra.db.HintedHandOffManager$1.run(HintedHandOffManager.java:168) ~[apache-cassandra-2.1.8.jar:2.1.8]\n at org.apache.cassandra.concurrent.DebuggableScheduledThreadPoolExecutor$UncomplainingRunnable.run(DebuggableScheduledThreadPoolExecutor.java:118) ~[apache-cassandra-2.1.8.jar:2.1.8]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) [na:1.8.0_45]\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) [na:1.8.0_45]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180) [na:1.8.0_45]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294) [na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_45]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_45]\n{code}\n\nObservations\n- the nodes are in UN state\n- after a node decommission, looks like the error 'moved' from one node to another\n\nDo have any ideas what are the consequences of this error? Is hinted hand off still working? ","created":"2015-09-30T17:00:47.037+0000"},{"body":"it will prevent null pointer exception when hints primary_key is null. \nIt is not allowing the hintedhandoff to be scheduled and performed. \nThe intent of this patch is to fix it. ","created":"2015-09-30T17:56:13.548+0000"},{"body":"patch attached","created":"2015-09-30T17:56:41.942+0000"},{"body":"Guys I code a little patch for the problem. Looks like it is affecting the hinted hand off scheduler.\n\n{code}\ncqlsh:system> SELECT target_id,hint_id,message_version FROM hints LIMIT 5;\n\n target_id | hint_id | message_version\n--------------------------------------+--------------------------------------+-----------------\n | 2f5e0320-62d3-11e5-877e-77558ae77cc8 | 8\n 72888e32-dae5-41cd-a033-3c5871a3e045 | fad152f0-662a-11e5-89ed-77558ae77cc8 | 8\n 72888e32-dae5-41cd-a033-3c5871a3e045 | fad152f1-662a-11e5-89ed-77558ae77cc8 | 8\n 72888e32-dae5-41cd-a033-3c5871a3e045 | fb69e970-662a-11e5-89ed-77558ae77cc8 | 8\n 72888e32-dae5-41cd-a033-3c5871a3e045 | fb69e971-662a-11e5-89ed-77558ae77cc8 | 8\n\n(5 rows)\n{code}\n\nWe have an empty target_id in our hints table. That is causing the problem. \nCan you check my patch?\n","created":"2015-09-30T17:59:08.270+0000"},{"body":"new patch available","created":"2015-10-01T14:29:24.149+0000"},{"body":"I'll give the patch a look. I'm also curious on how the broken hint got in there in the first place. I would like to try and reproduce can you give more details of your upgrade procedure? \n- how many nodes?\n- assuming rolling upgrade\n- jdk change?\n- roughly how long was each node unavailable\n- gc_grace value of table with broken hint\n- values of max_hint_window_in_ms, max_hints_delivery_threads, hinted_handoff_enabled, hinted_handoff_throttle_in_kb in cassandra.yaml\n- what type of mutation was the hint without a target_id?","created":"2015-10-01T14:39:26.840+0000"},{"body":"[~nutbunnies] please see this file:\nhttps://issues.apache.org/jira/secure/attachment/12764620/cassandra-2.1.8-10233-v2.txt\n\nThis is a new attachment without code formatting.\nIt will keep showing the exception on logs, but will not prevent hinted handoff scheduler to run.\nIt is very weird, looks like something corrupted, it is not possible to have an empty primary key. May be another issue. In fact the main objective here is to keep hinted hand off operations healthy and working.\n\nAccording to [~pauloricardomg] we can just truncate the system.hints table as a palliative solution. \n\n[~fhsgoncalves] will provide the information you requested.\nThanks.","created":"2015-10-01T14:51:07.176+0000"},{"body":"Hi [~nutbunnies], I work together with Eiti Kimura at Movile, and this issue is happening in one of our cluster of cassandra.\n\nI'll try answer your questions:\n\n- how many nodes?\nCurrently we are running with 15 nodes, in 2 racks, in the same datacenter. One rack has 7 nodes and the other has 8 nodes.\n\n- assuming rolling upgrade\nI did't understand if this is a question, but what I can say is that we already upgraded to version 2.1.9 yesterday, and the problem started when we added 7 new nodes to the cluster a week ago. We add one node a time, waiting for each node join the cluster before start the joining of the next node.\n\n- jdk change?\nWe are using the same version for a long time, Java Hotspot 1.8.0_45-b14.\n\n- roughly how long was each node unavailable\nI'm sending the uptime of each node. The nodes were not unavailable, only very slow to respond requests when we added 7 more nodes last week.\npompeia1 14:52:37 up 126 days\npompeia2 14:52:37 up 126 days\npompeia3 14:52:37 up 126 days\npompeia4 14:52:37 up 126 days\npompeia5 14:52:37 up 126 days\npompeia6 14:52:37 up 126 days\npompeia7 14:52:37 up 82 days\npompeia8 14:52:37 up 82 days\npompeia9 14:52:37 up 7 days\npompeia10 14:52:37 up 7 days\npompeia11 14:52:37 up 7 days\npompeia12 14:52:37 up 7 days\npompeia13 14:52:37 up 7 days\npompeia14 14:52:37 up 7 days\npompeia15 14:52:37 up 7 days\n\n- gc_grace value of table with broken hint\nvalues of max_hint_window_in_ms, max_hints_delivery_threads, hinted_handoff_enabled, hinted_handoff_throttle_in_kb in cassandra.yaml\nWe are not sure about the table that is problematic, but we think that is the most large (considering the records count and number of columns) and most used table that we have, and I'm going to inform the its values:\n-- gc_grace_seconds = 864000\nThe values in the application.yml\n-- max_hint_window_in_ms: 10800000\n-- max_hints_delivery_threads: 2\n-- hinted_handoff_enabled: true\n-- hinted_handoff_throttle_in_kb: 1024\n\n- what type of mutation was the hint without a target_id?\nI don't know how to get the type of mutation, only the mutation value, that is a blob in the table. Can you help me here?\n\nIf you need any other information, I can send to you!\nThank you!","created":"2015-10-01T15:09:59.424+0000"},{"body":"[~eitikimura] as [~pauloricardomg] mentioned truncating the hints table will get things rolling again, just make sure to repair in near future to prevent possible dataloss or zombie records related to the mutations truncated out of the hints table.\n\n[~fhsgoncalves] thanks, that helps a ton. Couple more questions to help reproduce it:\n- Just want to make sure order of operation is correct, the error showed up when you added additional nodes then persisted after upgrades? \n- Are all the new nodes in the second rack? \n- Are you running with vnodes?\n- Are your keyspaces set with a replication of NetworkTopologyStrategy, SimpleStrategy or ?\n- In the cassandra.yaml what is the value for: endpoint_snitch","created":"2015-10-01T22:29:44.724+0000"},{"body":"[~nutbunnies]:\n- Just want to make sure order of operation is correct, the error showed up when you added additional nodes then persisted after upgrades?\nWe are not sure about when the error started, but we think that was after the addition of nodes, and yes, it persisted after the upgrade to version 2.1.9.\n\n- Are all the new nodes in the second rack?\nNo, the new nodes are added to all racks using a round robin policy, each rack contains both old and new nodes always.\n\n- Are you running with vnodes?\nYes.\n\n- Are your keyspaces set with a replication of NetworkTopologyStrategy, SimpleStrategy or ?\nNetworkTopologyStrategy\n\n- In the cassandra.yaml what is the value for: endpoint_snitch\nendpoint_snitch: GossipingPropertyFileSnitch\n\n---\n\nWe did the truncate on table system.hints, and the errors for process the table disappeared and the table now are empty every time I checked :)\n\nWe are now running a repair on each node in the cluster, sequentially. But we still are afraid that this error happen again.\n\nThanks for the attention until now!","created":"2015-10-02T00:14:34.451+0000"},{"body":"Good set of info [~fhsgoncalves], thanks. \n[~nutbunnies] it is working now. My main concern is related to consistence, if it happens again (another issue) or with other people affected by problem may cause cluster problems if hints were not replied correctly.\n\nThe change in the patch will continue show the exception in system.log but will not stop the flow, we'll skip the inconsistent line and move to the next, continue to schedule the hints and hand off process without affecting the entire cluster consistency. \nI think having a bunch of nodes in the cluster with problems sending hints a big issue. The patch is not to fix the hints but will prevent things get any worse.\n\nDo you consider the possibility to apply this patch in the near future? \n\n","created":"2015-10-02T12:46:58.244+0000"},{"body":"[~eitikimura] [~fhsgoncalves] are JVM assertions enabled or disabled in your cluster (-ea JVM_OPTS on cassandra-env.sh) ?\n\nBecause in case assertions are disabled, there is a possibility hints are inserted with {{target_id}} null on {{HintedHandoffManager.writeHintForMutation}}, when the TokenMetadata is not fully populated for some reason:\n\n{code:java}\n public static void writeHintForMutation(Mutation mutation, long now, int ttl, InetAddress target)\n {\n assert ttl > 0;\n UUID hostId = StorageService.instance.getTokenMetadata().getHostId(target);\n assert hostId != null : \"Missing host ID for \" + target.getHostAddress();\n HintedHandOffManager.instance.hintFor(mutation, now, ttl, hostId).apply();\n StorageMetrics.totalHints.inc();\n }\n{code}","created":"2015-10-03T00:42:00.776+0000"},{"body":"assigning for patch review","created":"2015-10-05T16:11:49.633+0000"},{"body":"[~pauloricardomg] I think this method writeHintForMutation is from class StorageProxy and not for the HintedHandoffManager, right? ","created":"2015-10-05T17:02:25.260+0000"},{"body":"[~pauloricardomg] please, see the file: cassandra-2.1.8-10233-v3.txt\nI did the following improvement:\n\n{code:java}\n public static void writeHintForMutation(Mutation mutation, long now, int ttl, InetAddress target)\n {\n Preconditions.checkArgument(ttl > 0, \"the ttl must be > 0\");\n UUID hostId = StorageService.instance.getTokenMetadata().getHostId(target);\n Preconditions.checkNotNull(hostId, \"Missing host ID for \" + target.getHostAddress());\n HintedHandOffManager.instance.hintFor(mutation, now, ttl, hostId).apply();\n StorageMetrics.totalHints.inc();\n }\n{code}","created":"2015-10-05T17:15:35.124+0000"},{"body":"[~nutbunnies]: the JVM assertions are disabled!\nWe just follow the following comment in order to improve the performance: (we considered that it was a good practice). \nhttps://github.com/apache/cassandra/blob/cassandra-2.1/conf/cassandra-env.sh#L172","created":"2015-10-05T18:51:13.515+0000"},{"body":"I think that jvm assertions should be disabled in production, if we need to validate something in production it should be done like [~eitikimura] said.","created":"2015-10-05T18:53:32.127+0000"},{"body":"[~eitikimura] I'd prefer not add the try-catch block on {{HintedHandOffManager.scheduleAllDeliveries()}} as in the general case stored hints shouldn't be corrupted, and it could make hints be silently dropped which could lead to more serious issues. Since we already know the issue was caused on {{StorageProxy.writeHintForMutation}} I think it suffices to perform the check there. And if someone hit this bug before it's fixed, the workaround should be truncate hints + repair.\n\nI think what you did on {{StorageProxy.writeHintForMutation}} looks awesome, but the thrown exception might be ignored silently by the hints executor, so it's better to perform an explicit check, log a warn and throw an AssertionError if {{hostId \\!= null}}, so we'll be able to track if it happens again in the logs. Could you please make these changes and re-submit the patch? Please check if your patch apply to cassandra-2.2 branch, and if it doesn't please also submit a patch for 2.2. It should not be necessary to create a patch for 3.0, as the hints engine was rewritten from scratch.\n\nThanks for that [~eitikimura]!\n\n[~fhsgoncalves] yep, afaik assertions should be optional in production, but they should never happen in the first place. probably this is being caused by some other issue I was not able to track in the latest changes, but [~eitikimura]'s patch should help us troubleshoot if it happens in the future.\n\n[~nutbunnies] [~mambocab] maybe it would be interesting to have a dtest job with assertions disabled, since we rely a lot on assertions for pre-condition checking, and many people disable them in production.","created":"2015-10-05T21:32:00.412+0000"},{"body":"[~nutbunnies] Thanks!","created":"2015-10-05T23:44:32.058+0000"},{"body":"Hello [~pauloricardomg] you can se the new patches attached the 2.1-v4 and the 2.2.1 patch for 2.2.X versions.\nI'm expliciting check the values, logging a warn message and throing an AssertionError as you asked. I hope it helps. \nThanks ","created":"2015-10-06T00:45:05.776+0000"},{"body":"Folks,\n\nI am seeing the same exception in 2.0.14 release. Any plans to apply the patch in that release. What is the work-around?","created":"2015-10-07T13:59:04.965+0000"},{"body":"Hello [~kf200467] the workaround is just truncat the system.hints table. \n\nThe exception occurs due to inconsistent data in hints table, as described in my comments above, I found records where the hints primary key were null. Probably it was caused by assertions being disabled when running on production setup. \n\nThe patch I code and we discussed here is to prevent inconsistent data being persisted in to the hints table. \n I'll code a patch to 2.0.X version as soo as possible. ","created":"2015-10-07T14:27:21.585+0000"},{"body":"[~kenfailbus] can you please paste the output of the following command on cqlsh on the node throwing the exceptions:\n{noformat}\nSELECT target_id,hint_id,message_version FROM system.hints LIMIT 5;\n{noformat}\n\n","created":"2015-10-07T14:27:28.654+0000"},{"body":"I was wondering if it's possible to remove just the faulty hint without truncating the whole table, but I'm not sure if deletion of null keys is supported. If you confirm there is a hint with a null/empty target_id, could you try executing the following command to see if it works [~kenfailbus]? (did you try that, [~eitikimura]?)\n\n{noformat}\nDELETE FROM system.hints WHERE target_id = null;\nDELETE FROM system.hints WHERE target_id = '';\n{noformat}","created":"2015-10-07T14:32:37.618+0000"},{"body":"This is what I saw when running the above command.\n{code}\ncqlsh> SELECT target_id,hint_id,message_version FROM system.hints LIMIT 5;\n\n target_id | hint_id | message_version\n-----------+--------------------------------------+-----------------\n | af72b680-6622-11e5-bb6e-5b476c07be60 | 7\n | b27f3330-6622-11e5-bb6e-5b476c07be60 | 7\n | b5f0da50-6622-11e5-bb6e-5b476c07be60 | 7\n | b6a1b3c0-6622-11e5-bb6e-5b476c07be60 | 7\n | b6bcb5d0-6622-11e5-bb6e-5b476c07be60 | 7\n{code}","created":"2015-10-07T14:57:12.116+0000"},{"body":"Since the output above in select query looks very that there is no target_id present, would it be better off truncating the whole hints table and restarting cassandra. Also, we have reduced our hint window to 5min. beyond which we just bootstrap the node back or run repairs on it.","created":"2015-10-07T15:10:43.218+0000"},{"body":"[~eitikimura] thanks for your patch, I think we can keep the ttl > 0 assertion instead of the explicit check, since we won't run into problems if ttl <= 0 (even though that's not the case), so the assertion should suffice. Also, in the 2.1 patch there are some extra lines in the imports, could you please remove them? Sorry for not pointing that out before.\n\nAfter that it should be ready to commit!! :)","created":"2015-10-07T19:31:18.897+0000"},{"body":"[~pauloricardomg] nice! \nI changed the files as you suggested and I just added one more patch for 2.0 version as well. The new files attached are:\ncassandra-2.1-10233-v5.txt\ncassandra-2.0-10233.txt\ncassandra-2.2.1-10233-v2.txt","created":"2015-10-08T02:22:45.388+0000"},{"body":"LGTM, marking as \"Ready to Commit\". Thanks [~eitikimura].","created":"2015-10-08T14:32:02.984+0000"},{"body":"This fix only prevents hints with null target_id from being written in nodes with disabled assertions. I have opened CASSANDRA-10485 to address the root cause of the issue.","created":"2015-10-08T18:44:44.838+0000"},{"body":"Committed to 2.1 as [bc1058f8ea1e50d57db6e06fb027845871d9927c|https://github.com/apache/cassandra/commit/bc1058f8ea1e50d57db6e06fb027845871d9927c] and merged with 2.2, leaving 3.0 and trunk as is. Thank you.","created":"2015-10-14T15:59:12.551+0000"},{"body":"[~fhsgoncalves] [~eitikimura]\n\nbq. the problem started when we added 7 new nodes to the cluster a week ago\n\nDo you remember if there was any failed bootstrap when you added these nodes? For example, you started bootstrapping, streams hanged, you wiped the node and restarted the process again? Or did all bootstraps succeed in the first time? (I'm investigating the root cause of the issue on CASSANDRA-10485)","created":"2015-10-24T00:24:43.027+0000"}],"conversations":[{"body":"After upgrading our cluster to 2.2.0, the following error started showing exectly every 10 minutes on every server in the cluster:\n\n{noformat}\nINFO [CompactionExecutor:1381] 2015-08-31 18:31:55,506 CompactionTask.java:142 - Compacting (8e7e1520-500e-11e5-b1e3-e95897ba4d20) [/cassandra/data/system/hints-2666e20573ef38b390fefecf96e8f0c7/la-540-big-Data.db:level=0, ]\nINFO [CompactionExecutor:1381] 2015-08-31 18:31:55,599 CompactionTask.java:224 - Compacted (8e7e1520-500e-11e5-b1e3-e95897ba4d20) 1 sstables to [/cassandra/data/system/hints-2666e20573ef38b390fefecf96e8f0c7/la-541-big,] to level=0. 1,544,495 bytes to 1,544,495 (~100% of original) in 93ms = 15.838121MB/s. 0 total partitions merged to 4. Partition merge counts were {1:4, }\nERROR [HintedHandoff:1] 2015-08-31 18:31:55,600 CassandraDaemon.java:182 - Exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.IndexOutOfBoundsException: null\n\tat java.nio.Buffer.checkIndex(Buffer.java:538) ~[na:1.7.0_79]\n\tat java.nio.HeapByteBuffer.getLong(HeapByteBuffer.java:410) ~[na:1.7.0_79]\n\tat org.apache.cassandra.utils.UUIDGen.getUUID(UUIDGen.java:106) ~[apache-cassandra-2.2.0.jar:2.2.0]\n\tat org.apache.cassandra.db.HintedHandOffManager.scheduleAllDeliveries(HintedHandOffManager.java:515) ~[apache-cassandra-2.2.0.jar:2.2.0]\n\tat org.apache.cassandra.db.HintedHandOffManager.access$000(HintedHandOffManager.java:88) ~[apache-cassandra-2.2.0.jar:2.2.0]\n\tat org.apache.cassandra.db.HintedHandOffManager$1.run(HintedHandOffManager.java:168) ~[apache-cassandra-2.2.0.jar:2.2.0]\n\tat org.apache.cassandra.concurrent.DebuggableScheduledThreadPoolExecutor$UncomplainingRunnable.run(DebuggableScheduledThreadPoolExecutor.java:118) ~[apache-cassandra-2.2.0.jar:2.2.0]\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_79]\n\tat java.util.concurrent.FutureTask.runAndReset(FutureTask.java:304) [na:1.7.0_79]\n\tat java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:178) [na:1.7.0_79]\n\tat java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:293) [na:1.7.0_79]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) [na:1.7.0_79]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_79]\n\tat java.lang.Thread.run(Thread.java:745) [na:1.7.0_79]\n{noformat}","from":"reporter","subject":"IndexOutOfBoundsException in HintedHandOffManager"},{"body":"We haven't seen this in our upgrade suite. Can you provide any further information as to your environment, what version did you upgrade from? Any non-default cassandra yaml settings? Can you procide your (sanitized) schema, and any other relevant logs?\n\n/cc [~krummas] if he has any hint as to the problem.","from":"developer"},{"body":"Upgraded from 2.1.8. I upgraded today to 2.2.1 and still see the issue.\n\nI am not sure which table is causing this issue and we have quite a few tables so pasting the whole schema would be impossible. If you have any pointers on how to narrow to a single table (or even a single keyspace) I'd be happy to provide it's schema.\n\nThe only changes to cassandra.yaml are increasing the write timeout and increasing the size of out of heap memtable to 6GB.","from":"developer"},{"body":"In fact I don't believe that this problem is related exclusively to versions 2.2.X.\nI'm facing the very same problem here. We have upgrade from version 2.1.5 to 2.1.8, after migration we noticed the exception above in the logs. \n\nLooks like something related to hinded hand off, according to the staktrace:\nHintedHandOffManager.java:515\n\nHere is the stack for the 2.1.8 version:\n{code}\nERROR [HintedHandoff:2] 2015-09-30 16:48:38,729 CassandraDaemon.java:223 - Exception in thread Thread[HintedHandoff:2,1,main]\njava.lang.IndexOutOfBoundsException: null\n at java.nio.Buffer.checkIndex(Buffer.java:546) ~[na:1.8.0_45]\n at java.nio.HeapByteBuffer.getLong(HeapByteBuffer.java:416) ~[na:1.8.0_45]\n at org.apache.cassandra.utils.UUIDGen.getUUID(UUIDGen.java:106) ~[apache-cassandra-2.1.8.jar:2.1.8]\n at org.apache.cassandra.db.HintedHandOffManager.scheduleAllDeliveries(HintedHandOffManager.java:514) ~[apache-cassandra-2.1.8.jar:2.1.8]\n at org.apache.cassandra.db.HintedHandOffManager.access$000(HintedHandOffManager.java:88) ~[apache-cassandra-2.1.8.jar:2.1.8]\n at org.apache.cassandra.db.HintedHandOffManager$1.run(HintedHandOffManager.java:168) ~[apache-cassandra-2.1.8.jar:2.1.8]\n at org.apache.cassandra.concurrent.DebuggableScheduledThreadPoolExecutor$UncomplainingRunnable.run(DebuggableScheduledThreadPoolExecutor.java:118) ~[apache-cassandra-2.1.8.jar:2.1.8]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) [na:1.8.0_45]\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) [na:1.8.0_45]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180) [na:1.8.0_45]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294) [na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_45]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_45]\n{code}\n\nObservations\n- the nodes are in UN state\n- after a node decommission, looks like the error 'moved' from one node to another\n\nDo have any ideas what are the consequences of this error? Is hinted hand off still working? ","from":"developer"},{"body":"it will prevent null pointer exception when hints primary_key is null. \nIt is not allowing the hintedhandoff to be scheduled and performed. \nThe intent of this patch is to fix it. ","from":"developer"},{"body":"patch attached","from":"developer"},{"body":"Guys I code a little patch for the problem. Looks like it is affecting the hinted hand off scheduler.\n\n{code}\ncqlsh:system> SELECT target_id,hint_id,message_version FROM hints LIMIT 5;\n\n target_id | hint_id | message_version\n--------------------------------------+--------------------------------------+-----------------\n | 2f5e0320-62d3-11e5-877e-77558ae77cc8 | 8\n 72888e32-dae5-41cd-a033-3c5871a3e045 | fad152f0-662a-11e5-89ed-77558ae77cc8 | 8\n 72888e32-dae5-41cd-a033-3c5871a3e045 | fad152f1-662a-11e5-89ed-77558ae77cc8 | 8\n 72888e32-dae5-41cd-a033-3c5871a3e045 | fb69e970-662a-11e5-89ed-77558ae77cc8 | 8\n 72888e32-dae5-41cd-a033-3c5871a3e045 | fb69e971-662a-11e5-89ed-77558ae77cc8 | 8\n\n(5 rows)\n{code}\n\nWe have an empty target_id in our hints table. That is causing the problem. \nCan you check my patch?\n","from":"developer"},{"body":"new patch available","from":"developer"},{"body":"I'll give the patch a look. I'm also curious on how the broken hint got in there in the first place. I would like to try and reproduce can you give more details of your upgrade procedure? \n- how many nodes?\n- assuming rolling upgrade\n- jdk change?\n- roughly how long was each node unavailable\n- gc_grace value of table with broken hint\n- values of max_hint_window_in_ms, max_hints_delivery_threads, hinted_handoff_enabled, hinted_handoff_throttle_in_kb in cassandra.yaml\n- what type of mutation was the hint without a target_id?","from":"developer"},{"body":"[~nutbunnies] please see this file:\nhttps://issues.apache.org/jira/secure/attachment/12764620/cassandra-2.1.8-10233-v2.txt\n\nThis is a new attachment without code formatting.\nIt will keep showing the exception on logs, but will not prevent hinted handoff scheduler to run.\nIt is very weird, looks like something corrupted, it is not possible to have an empty primary key. May be another issue. In fact the main objective here is to keep hinted hand off operations healthy and working.\n\nAccording to [~pauloricardomg] we can just truncate the system.hints table as a palliative solution. \n\n[~fhsgoncalves] will provide the information you requested.\nThanks.","from":"developer"},{"body":"Hi [~nutbunnies], I work together with Eiti Kimura at Movile, and this issue is happening in one of our cluster of cassandra.\n\nI'll try answer your questions:\n\n- how many nodes?\nCurrently we are running with 15 nodes, in 2 racks, in the same datacenter. One rack has 7 nodes and the other has 8 nodes.\n\n- assuming rolling upgrade\nI did't understand if this is a question, but what I can say is that we already upgraded to version 2.1.9 yesterday, and the problem started when we added 7 new nodes to the cluster a week ago. We add one node a time, waiting for each node join the cluster before start the joining of the next node.\n\n- jdk change?\nWe are using the same version for a long time, Java Hotspot 1.8.0_45-b14.\n\n- roughly how long was each node unavailable\nI'm sending the uptime of each node. The nodes were not unavailable, only very slow to respond requests when we added 7 more nodes last week.\npompeia1 14:52:37 up 126 days\npompeia2 14:52:37 up 126 days\npompeia3 14:52:37 up 126 days\npompeia4 14:52:37 up 126 days\npompeia5 14:52:37 up 126 days\npompeia6 14:52:37 up 126 days\npompeia7 14:52:37 up 82 days\npompeia8 14:52:37 up 82 days\npompeia9 14:52:37 up 7 days\npompeia10 14:52:37 up 7 days\npompeia11 14:52:37 up 7 days\npompeia12 14:52:37 up 7 days\npompeia13 14:52:37 up 7 days\npompeia14 14:52:37 up 7 days\npompeia15 14:52:37 up 7 days\n\n- gc_grace value of table with broken hint\nvalues of max_hint_window_in_ms, max_hints_delivery_threads, hinted_handoff_enabled, hinted_handoff_throttle_in_kb in cassandra.yaml\nWe are not sure about the table that is problematic, but we think that is the most large (considering the records count and number of columns) and most used table that we have, and I'm going to inform the its values:\n-- gc_grace_seconds = 864000\nThe values in the application.yml\n-- max_hint_window_in_ms: 10800000\n-- max_hints_delivery_threads: 2\n-- hinted_handoff_enabled: true\n-- hinted_handoff_throttle_in_kb: 1024\n\n- what type of mutation was the hint without a target_id?\nI don't know how to get the type of mutation, only the mutation value, that is a blob in the table. Can you help me here?\n\nIf you need any other information, I can send to you!\nThank you!","from":"developer"},{"body":"[~eitikimura] as [~pauloricardomg] mentioned truncating the hints table will get things rolling again, just make sure to repair in near future to prevent possible dataloss or zombie records related to the mutations truncated out of the hints table.\n\n[~fhsgoncalves] thanks, that helps a ton. Couple more questions to help reproduce it:\n- Just want to make sure order of operation is correct, the error showed up when you added additional nodes then persisted after upgrades? \n- Are all the new nodes in the second rack? \n- Are you running with vnodes?\n- Are your keyspaces set with a replication of NetworkTopologyStrategy, SimpleStrategy or ?\n- In the cassandra.yaml what is the value for: endpoint_snitch","from":"developer"},{"body":"[~nutbunnies]:\n- Just want to make sure order of operation is correct, the error showed up when you added additional nodes then persisted after upgrades?\nWe are not sure about when the error started, but we think that was after the addition of nodes, and yes, it persisted after the upgrade to version 2.1.9.\n\n- Are all the new nodes in the second rack?\nNo, the new nodes are added to all racks using a round robin policy, each rack contains both old and new nodes always.\n\n- Are you running with vnodes?\nYes.\n\n- Are your keyspaces set with a replication of NetworkTopologyStrategy, SimpleStrategy or ?\nNetworkTopologyStrategy\n\n- In the cassandra.yaml what is the value for: endpoint_snitch\nendpoint_snitch: GossipingPropertyFileSnitch\n\n---\n\nWe did the truncate on table system.hints, and the errors for process the table disappeared and the table now are empty every time I checked :)\n\nWe are now running a repair on each node in the cluster, sequentially. But we still are afraid that this error happen again.\n\nThanks for the attention until now!","from":"developer"},{"body":"Good set of info [~fhsgoncalves], thanks. \n[~nutbunnies] it is working now. My main concern is related to consistence, if it happens again (another issue) or with other people affected by problem may cause cluster problems if hints were not replied correctly.\n\nThe change in the patch will continue show the exception in system.log but will not stop the flow, we'll skip the inconsistent line and move to the next, continue to schedule the hints and hand off process without affecting the entire cluster consistency. \nI think having a bunch of nodes in the cluster with problems sending hints a big issue. The patch is not to fix the hints but will prevent things get any worse.\n\nDo you consider the possibility to apply this patch in the near future? \n\n","from":"developer"},{"body":"[~eitikimura] [~fhsgoncalves] are JVM assertions enabled or disabled in your cluster (-ea JVM_OPTS on cassandra-env.sh) ?\n\nBecause in case assertions are disabled, there is a possibility hints are inserted with {{target_id}} null on {{HintedHandoffManager.writeHintForMutation}}, when the TokenMetadata is not fully populated for some reason:\n\n{code:java}\n public static void writeHintForMutation(Mutation mutation, long now, int ttl, InetAddress target)\n {\n assert ttl > 0;\n UUID hostId = StorageService.instance.getTokenMetadata().getHostId(target);\n assert hostId != null : \"Missing host ID for \" + target.getHostAddress();\n HintedHandOffManager.instance.hintFor(mutation, now, ttl, hostId).apply();\n StorageMetrics.totalHints.inc();\n }\n{code}","from":"developer"},{"body":"assigning for patch review","from":"developer"},{"body":"[~pauloricardomg] I think this method writeHintForMutation is from class StorageProxy and not for the HintedHandoffManager, right? ","from":"developer"},{"body":"[~pauloricardomg] please, see the file: cassandra-2.1.8-10233-v3.txt\nI did the following improvement:\n\n{code:java}\n public static void writeHintForMutation(Mutation mutation, long now, int ttl, InetAddress target)\n {\n Preconditions.checkArgument(ttl > 0, \"the ttl must be > 0\");\n UUID hostId = StorageService.instance.getTokenMetadata().getHostId(target);\n Preconditions.checkNotNull(hostId, \"Missing host ID for \" + target.getHostAddress());\n HintedHandOffManager.instance.hintFor(mutation, now, ttl, hostId).apply();\n StorageMetrics.totalHints.inc();\n }\n{code}","from":"developer"},{"body":"[~nutbunnies]: the JVM assertions are disabled!\nWe just follow the following comment in order to improve the performance: (we considered that it was a good practice). \nhttps://github.com/apache/cassandra/blob/cassandra-2.1/conf/cassandra-env.sh#L172","from":"developer"},{"body":"I think that jvm assertions should be disabled in production, if we need to validate something in production it should be done like [~eitikimura] said.","from":"developer"},{"body":"[~eitikimura] I'd prefer not add the try-catch block on {{HintedHandOffManager.scheduleAllDeliveries()}} as in the general case stored hints shouldn't be corrupted, and it could make hints be silently dropped which could lead to more serious issues. Since we already know the issue was caused on {{StorageProxy.writeHintForMutation}} I think it suffices to perform the check there. And if someone hit this bug before it's fixed, the workaround should be truncate hints + repair.\n\nI think what you did on {{StorageProxy.writeHintForMutation}} looks awesome, but the thrown exception might be ignored silently by the hints executor, so it's better to perform an explicit check, log a warn and throw an AssertionError if {{hostId \\!= null}}, so we'll be able to track if it happens again in the logs. Could you please make these changes and re-submit the patch? Please check if your patch apply to cassandra-2.2 branch, and if it doesn't please also submit a patch for 2.2. It should not be necessary to create a patch for 3.0, as the hints engine was rewritten from scratch.\n\nThanks for that [~eitikimura]!\n\n[~fhsgoncalves] yep, afaik assertions should be optional in production, but they should never happen in the first place. probably this is being caused by some other issue I was not able to track in the latest changes, but [~eitikimura]'s patch should help us troubleshoot if it happens in the future.\n\n[~nutbunnies] [~mambocab] maybe it would be interesting to have a dtest job with assertions disabled, since we rely a lot on assertions for pre-condition checking, and many people disable them in production.","from":"developer"},{"body":"[~nutbunnies] Thanks!","from":"developer"},{"body":"Hello [~pauloricardomg] you can se the new patches attached the 2.1-v4 and the 2.2.1 patch for 2.2.X versions.\nI'm expliciting check the values, logging a warn message and throing an AssertionError as you asked. I hope it helps. \nThanks ","from":"developer"},{"body":"Folks,\n\nI am seeing the same exception in 2.0.14 release. Any plans to apply the patch in that release. What is the work-around?","from":"developer"},{"body":"Hello [~kf200467] the workaround is just truncat the system.hints table. \n\nThe exception occurs due to inconsistent data in hints table, as described in my comments above, I found records where the hints primary key were null. Probably it was caused by assertions being disabled when running on production setup. \n\nThe patch I code and we discussed here is to prevent inconsistent data being persisted in to the hints table. \n I'll code a patch to 2.0.X version as soo as possible. ","from":"developer"},{"body":"[~kenfailbus] can you please paste the output of the following command on cqlsh on the node throwing the exceptions:\n{noformat}\nSELECT target_id,hint_id,message_version FROM system.hints LIMIT 5;\n{noformat}\n\n","from":"developer"},{"body":"I was wondering if it's possible to remove just the faulty hint without truncating the whole table, but I'm not sure if deletion of null keys is supported. If you confirm there is a hint with a null/empty target_id, could you try executing the following command to see if it works [~kenfailbus]? (did you try that, [~eitikimura]?)\n\n{noformat}\nDELETE FROM system.hints WHERE target_id = null;\nDELETE FROM system.hints WHERE target_id = '';\n{noformat}","from":"developer"},{"body":"This is what I saw when running the above command.\n{code}\ncqlsh> SELECT target_id,hint_id,message_version FROM system.hints LIMIT 5;\n\n target_id | hint_id | message_version\n-----------+--------------------------------------+-----------------\n | af72b680-6622-11e5-bb6e-5b476c07be60 | 7\n | b27f3330-6622-11e5-bb6e-5b476c07be60 | 7\n | b5f0da50-6622-11e5-bb6e-5b476c07be60 | 7\n | b6a1b3c0-6622-11e5-bb6e-5b476c07be60 | 7\n | b6bcb5d0-6622-11e5-bb6e-5b476c07be60 | 7\n{code}","from":"developer"},{"body":"Since the output above in select query looks very that there is no target_id present, would it be better off truncating the whole hints table and restarting cassandra. Also, we have reduced our hint window to 5min. beyond which we just bootstrap the node back or run repairs on it.","from":"developer"},{"body":"[~eitikimura] thanks for your patch, I think we can keep the ttl > 0 assertion instead of the explicit check, since we won't run into problems if ttl <= 0 (even though that's not the case), so the assertion should suffice. Also, in the 2.1 patch there are some extra lines in the imports, could you please remove them? Sorry for not pointing that out before.\n\nAfter that it should be ready to commit!! :)","from":"developer"},{"body":"[~pauloricardomg] nice! \nI changed the files as you suggested and I just added one more patch for 2.0 version as well. The new files attached are:\ncassandra-2.1-10233-v5.txt\ncassandra-2.0-10233.txt\ncassandra-2.2.1-10233-v2.txt","from":"developer"},{"body":"LGTM, marking as \"Ready to Commit\". Thanks [~eitikimura].","from":"developer"},{"body":"This fix only prevents hints with null target_id from being written in nodes with disabled assertions. I have opened CASSANDRA-10485 to address the root cause of the issue.","from":"developer"},{"body":"Committed to 2.1 as [bc1058f8ea1e50d57db6e06fb027845871d9927c|https://github.com/apache/cassandra/commit/bc1058f8ea1e50d57db6e06fb027845871d9927c] and merged with 2.2, leaving 3.0 and trunk as is. Thank you.","from":"developer"},{"body":"[~fhsgoncalves] [~eitikimura]\n\nbq. the problem started when we added 7 new nodes to the cluster a week ago\n\nDo you remember if there was any failed bootstrap when you added these nodes? For example, you started bootstrapping, streams hanged, you wiped the node and restarted the process again? Or did all bootstraps succeed in the first time? (I'm investigating the root cause of the issue on CASSANDRA-10485)","from":"developer"}],"created":"2015-08-31T18:44:26.000+0000","description":"After upgrading our cluster to 2.2.0, the following error started showing exectly every 10 minutes on every server in the cluster:\n\n{noformat}\nINFO [CompactionExecutor:1381] 2015-08-31 18:31:55,506 CompactionTask.java:142 - Compacting (8e7e1520-500e-11e5-b1e3-e95897ba4d20) [/cassandra/data/system/hints-2666e20573ef38b390fefecf96e8f0c7/la-540-big-Data.db:level=0, ]\nINFO [CompactionExecutor:1381] 2015-08-31 18:31:55,599 CompactionTask.java:224 - Compacted (8e7e1520-500e-11e5-b1e3-e95897ba4d20) 1 sstables to [/cassandra/data/system/hints-2666e20573ef38b390fefecf96e8f0c7/la-541-big,] to level=0. 1,544,495 bytes to 1,544,495 (~100% of original) in 93ms = 15.838121MB/s. 0 total partitions merged to 4. Partition merge counts were {1:4, }\nERROR [HintedHandoff:1] 2015-08-31 18:31:55,600 CassandraDaemon.java:182 - Exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.IndexOutOfBoundsException: null\n\tat java.nio.Buffer.checkIndex(Buffer.java:538) ~[na:1.7.0_79]\n\tat java.nio.HeapByteBuffer.getLong(HeapByteBuffer.java:410) ~[na:1.7.0_79]\n\tat org.apache.cassandra.utils.UUIDGen.getUUID(UUIDGen.java:106) ~[apache-cassandra-2.2.0.jar:2.2.0]\n\tat org.apache.cassandra.db.HintedHandOffManager.scheduleAllDeliveries(HintedHandOffManager.java:515) ~[apache-cassandra-2.2.0.jar:2.2.0]\n\tat org.apache.cassandra.db.HintedHandOffManager.access$000(HintedHandOffManager.java:88) ~[apache-cassandra-2.2.0.jar:2.2.0]\n\tat org.apache.cassandra.db.HintedHandOffManager$1.run(HintedHandOffManager.java:168) ~[apache-cassandra-2.2.0.jar:2.2.0]\n\tat org.apache.cassandra.concurrent.DebuggableScheduledThreadPoolExecutor$UncomplainingRunnable.run(DebuggableScheduledThreadPoolExecutor.java:118) ~[apache-cassandra-2.2.0.jar:2.2.0]\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_79]\n\tat java.util.concurrent.FutureTask.runAndReset(FutureTask.java:304) [na:1.7.0_79]\n\tat java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:178) [na:1.7.0_79]\n\tat java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:293) [na:1.7.0_79]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) [na:1.7.0_79]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_79]\n\tat java.lang.Thread.run(Thread.java:745) [na:1.7.0_79]\n{noformat}","issue_id":"12860774","key":"CASSANDRA-10233","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-10-14T15:59:46.000+0000","role":"fixed_distractor","summary":"IndexOutOfBoundsException in HintedHandOffManager"} {"case_id":"12861596","cluster":"DISTRACTOR-CASSANDRA-10258","comments":[{"body":"We have the same behavior on counter's table in cassandra 2.1.8. In the case of inserting data only through sstable the data was successfully inserts and reads. In the case of mixed inserts (java-driver CQL and jmx-sstableloader) in same column-family we had read timeout and warning in cassandra.log (below):\n\njava.lang.AssertionError: Wrong class type: class org.apache.cassandra.db.BufferCounterUpdateCell\n at org.apache.cassandra.db.AbstractCell.reconcileCounter(AbstractCell.java:211) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.BufferCounterCell.reconcile(BufferCounterCell.java:118) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.filter.QueryFilter$1.reduce(QueryFilter.java:122) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.filter.QueryFilter$1.reduce(QueryFilter.java:116) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:114) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:100) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143) ~[guava-16.0.1.jar:na]\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138) ~[guava-16.0.1.jar:na]\n at org.apache.cassandra.db.filter.SliceQueryFilter.collectReducedColumns(SliceQueryFilter.java:264) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.filter.QueryFilter.collateColumns(QueryFilter.java:108) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:82) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.RowIteratorFactory$2.getReduced(RowIteratorFactory.java:99) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.RowIteratorFactory$2.getReduced(RowIteratorFactory.java:71) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:117) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:100) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143) ~[guava-16.0.1.jar:na]\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138) ~[guava-16.0.1.jar:na]\n at org.apache.cassandra.db.ColumnFamilyStore$8.computeNext(ColumnFamilyStore.java:2048) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.ColumnFamilyStore$8.computeNext(ColumnFamilyStore.java:2044) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143) ~[guava-16.0.1.jar:na]\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138) ~[guava-16.0.1.jar:na]\n","created":"2015-10-19T12:32:28.601+0000"},{"body":"I was able to reproduce the issue above with a simple [test case|https://github.com/apache/cassandra/commit/24ab53e4b3b6512488b9ba119f777a84ebd6fa0b] on Cassandra 2.1.\n\nThe problem is that {{CQLSSTableWriter}} serializes {{CounterUpdateCell}} objects without being applied by a counter leader replica, so they cannot be reconciled with ordinary {{CounterCell}} objects during reads/compactions. Another limitation with {{CQLSStableWriter}} and counters on 2.1, is that multiple counter updates to the same partition are not coalesced by {{CQLSSTableWriter}}, so only the last counter update is applied.\n\nSurprisingly enough, on 3.0, it's not possible to reproduce this issue, because CASSANDRA-8099 abolished the {{CounterUpdateCells}} altogether, so {{CQLSSTableWriter}} counter cells are treated as remote shards during reconcile. Furthermore, the {{PartitionUpdate}} abstraction, also introduced by CASSANDRA-8099, merges all updates to the same partition before serialization on {{CQLSSTableWriter}}, so multiple updates to the same partition are indeed coalesced on 3.0.\n\nWe could fix the 2.1 behavior by special-casing the counter cell creation on {{UpdateParameters.addCounter()}} when {{Config.isClientMode()}} (an indication that the {{CQLSStableWriter}} is running) and creating a remote {{CounterCell}} instead, as done [in this commented block|https://github.com/apache/cassandra/blob/5449ad7b7e1b1930c07e0ea80ac3c3e2882633b5/src/java/org/apache/cassandra/cql3/UpdateParameters.java#L79]. This would mimic the 3.0 behavior and treat {{CQLSStableWriter}} counters as a remote shard.\n\nA hidden nit in this solution (and also on 3.0) is that a remote counter context is created for each {{CQLSSTableWriter}} session that updates counters, and afaik we currently have no way of discarding old counter contexts. So, new counter contexts created by {{CQLSStableWriter}} can potentially grow unbounded, what could affect performance significantly. Maybe after CASSANDRA-6506 we will be able to discard \"inactive\" counter contexts, and be able to support counters on {{CQLSStableWriter}} more efficiently.\n\nWhile we don't have a solution for cleaning inactive/temporary counter contexts, I propose we disable counters usage by default on {{CQLSSTableWriter}}, or at least make it harder to use (with a flag/warning, for instance), since it might still be useful for testing scenarios. What do you think [~slebresne], [~iamaleksey]? Do you see a better alternative?","created":"2015-10-21T02:36:39.695+0000"},{"body":"I think we should just forbid the use of counters in CQLSSTableWriter (in all version to be clear). By design, counters require for their efficiency that shards are only created on replicas, so that the set of \"node id\" contained in any counter stay relatively low. CQLSSTableWriter just doesn't fulfill that assumption and I don't think there is anything we can do about it. Generating a fake \"node id\" for each CQLSSTableWriter (which is what ends up happening) has a very high potential to blow the context sizes and it's a bad idea: you are much better off just using normal CQL updates if you want to load counters (sure that may be slower, but slower is better than broken).\n\nOn top of that problem, there is the fact that if a given instance of CQLSSTableWriter increments the same counter twice, we *cannot* guarantee that both increment will be added together because we cannot guarantee both increment will make it into the same sstable. If they don't (make it into the same sstable), then whatever node the sstables ends up to will consider those 2 separate shards as \"remote\" and won't add these together (we'll probably end up with a broken shard warning in fact).\n\nSo really, the simplest and imo best solution is to add \"you cannot write sstable yourself with CQLSSTableWriter\" to the list of counters limitations.","created":"2015-10-21T13:24:35.294+0000"},{"body":"Thanks for the quick feedback [~slebresne].\n\nAttached trivial patch and simple test forbidding counter updates on {{CQLSStableWriter}}. Mind reviewing/committing?\n\nThanks!","created":"2015-10-22T18:58:29.735+0000"},{"body":"bq. So really, the simplest and imo best solution is to add \"you cannot write sstable yourself with CQLSSTableWriter\" to the list of counters limitations.\n\nBasically, this.","created":"2015-10-23T16:25:16.204+0000"},{"body":"Committed as [3674ad9dab8f29173d7d4ee82488a8e9ea586240|https://github.com/apache/cassandra/commit/3674ad9dab8f29173d7d4ee82488a8e9ea586240] to 2.1, merged with 2.2, 3.0, 3.1, and trunk. Thanks.","created":"2015-11-10T13:48:53.423+0000"}],"conversations":[{"body":"We use CQLSStableWriter to produce testing datasets.\n\nHere are the steps to reproduce this issue :\n\n1) definition of a table with counter\n{code}\nCREATE TABLE my_counter (\n my_id text,\n my_counter counter,\n PRIMARY KEY (my_id)\n)\n{code}\n\n2) with CQLSSTableWriter initialize this table (about 2millions entries) with this insert order (one insert / key only)\n{{UPDATE myks.my_counter SET my_counter = my_counter + ? WHERE my_id = ?}}\n\n3) load the files written by CQLSSTableWriter with sstableloader in your cassandra cluster (tested on a single node and a 3 nodes cluster)\n\n4) start a process that updates the counters (we used 3millions entries distributed on the key my_id)\n\n5) after a while try to query a key in the my_counter table\n{{cqlsh:myks> select * from my_counter where my_id='0000001';}}\nRequest did not complete within rpc_timeout.\n\nIn the logs of cassandra (2.0.12) :\n{code}\nERROR [CompactionExecutor:3] 2015-05-28 15:53:39,491 CassandraDaemon.java (line 258) Exception in thread Thread[CompactionExecutor:3,1,main]\njava.lang.AssertionError: Wrong class type.\n at org.apache.cassandra.db.CounterUpdateColumn.reconcile(CounterUpdateColumn.java:70)\n at org.apache.cassandra.db.ArrayBackedSortedColumns.resolveAgainst(ArrayBackedSortedColumns.java:147)\n at org.apache.cassandra.db.ArrayBackedSortedColumns.addColumn(ArrayBackedSortedColumns.java:126)\n at org.apache.cassandra.db.ColumnFamily.addColumn(ColumnFamily.java:121)\n at org.apache.cassandra.db.compaction.PrecompactedRow$1.reduce(PrecompactedRow.java:120)\n at org.apache.cassandra.db.compaction.PrecompactedRow$1.reduce(PrecompactedRow.java:115)\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:112)\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:98)\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n at org.apache.cassandra.db.filter.SliceQueryFilter.collectReducedColumns(SliceQueryFilter.java:191)\n at org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:144)\n at org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:103)\n at org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:85)\n at org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:196)\n at org.apache.cassandra.db.compaction.CompactionIterable$Reducer.getReduced(CompactionIterable.java:74)\n at org.apache.cassandra.db.compaction.CompactionIterable$Reducer.getReduced(CompactionIterable.java:55)\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:115)\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:98)\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:164)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:60)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:59)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:198)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\nIn the logs of cassandra v2.1.5 :\n{code}\nWARN [SharedPool-Worker-38] 2015-06-11 16:39:06,008 AbstractTracingAwareExecutorService.java:169 - Uncaught exception on thread Thread[SharedPool-Worker-38,5,\nmain]: {}\njava.lang.AssertionError: Wrong class type: class org.apache.cassandra.db.BufferCounterUpdateCell\n at org.apache.cassandra.db.AbstractCell.reconcileCounter(AbstractCell.java:211) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.BufferCounterCell.reconcile(BufferCounterCell.java:118) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.filter.QueryFilter$1.reduce(QueryFilter.java:122) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.filter.QueryFilter$1.reduce(QueryFilter.java:116) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:114) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:100) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143) ~[guava-16.0.1.jar:na]\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138) ~[guava-16.0.1.jar:na]\n at org.apache.cassandra.db.filter.NamesQueryFilter.collectReducedColumns(NamesQueryFilter.java:100) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.filter.QueryFilter.collateColumns(QueryFilter.java:108) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:82) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:69) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CollationController.collectAllData(CollationController.java:314) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CollationController.getTopLevelColumns(CollationController.java:62) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.ColumnFamilyStore.getTopLevelColumns(ColumnFamilyStore.java:1900) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.ColumnFamilyStore.getColumnFamily(ColumnFamilyStore.java:1758) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.Keyspace.getRow(Keyspace.java:346) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.SliceByNamesReadCommand.getRow(SliceByNamesReadCommand.java:53) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CounterMutation.getCurrentValuesFromCFS(CounterMutation.java:262) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CounterMutation.getCurrentValues(CounterMutation.java:229) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CounterMutation.processModifications(CounterMutation.java:197) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CounterMutation.apply(CounterMutation.java:124) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.service.StorageProxy$8.runMayThrow(StorageProxy.java:1155) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2191) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_45]\n at org.apache.cassandra.concurrent.AbstractTracingAwareExecutorService$FutureTask.run(AbstractTracingAwareExecutorService.java:164) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [cassandra-all-2.1.5.469.jar:2.1.5.469]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_45]\n{code}\n\nsame exception as issue https://issues.apache.org/jira/browse/CASSANDRA-7188","from":"reporter","subject":"Reject counter writes in CQLSSTableWriter"},{"body":"We have the same behavior on counter's table in cassandra 2.1.8. In the case of inserting data only through sstable the data was successfully inserts and reads. In the case of mixed inserts (java-driver CQL and jmx-sstableloader) in same column-family we had read timeout and warning in cassandra.log (below):\n\njava.lang.AssertionError: Wrong class type: class org.apache.cassandra.db.BufferCounterUpdateCell\n at org.apache.cassandra.db.AbstractCell.reconcileCounter(AbstractCell.java:211) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.BufferCounterCell.reconcile(BufferCounterCell.java:118) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.filter.QueryFilter$1.reduce(QueryFilter.java:122) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.filter.QueryFilter$1.reduce(QueryFilter.java:116) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:114) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:100) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143) ~[guava-16.0.1.jar:na]\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138) ~[guava-16.0.1.jar:na]\n at org.apache.cassandra.db.filter.SliceQueryFilter.collectReducedColumns(SliceQueryFilter.java:264) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.filter.QueryFilter.collateColumns(QueryFilter.java:108) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:82) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.RowIteratorFactory$2.getReduced(RowIteratorFactory.java:99) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.RowIteratorFactory$2.getReduced(RowIteratorFactory.java:71) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:117) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:100) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143) ~[guava-16.0.1.jar:na]\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138) ~[guava-16.0.1.jar:na]\n at org.apache.cassandra.db.ColumnFamilyStore$8.computeNext(ColumnFamilyStore.java:2048) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at org.apache.cassandra.db.ColumnFamilyStore$8.computeNext(ColumnFamilyStore.java:2044) ~[cassandra-all-2.1.8.621.jar:2.1.8.621]\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143) ~[guava-16.0.1.jar:na]\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138) ~[guava-16.0.1.jar:na]\n","from":"developer"},{"body":"I was able to reproduce the issue above with a simple [test case|https://github.com/apache/cassandra/commit/24ab53e4b3b6512488b9ba119f777a84ebd6fa0b] on Cassandra 2.1.\n\nThe problem is that {{CQLSSTableWriter}} serializes {{CounterUpdateCell}} objects without being applied by a counter leader replica, so they cannot be reconciled with ordinary {{CounterCell}} objects during reads/compactions. Another limitation with {{CQLSStableWriter}} and counters on 2.1, is that multiple counter updates to the same partition are not coalesced by {{CQLSSTableWriter}}, so only the last counter update is applied.\n\nSurprisingly enough, on 3.0, it's not possible to reproduce this issue, because CASSANDRA-8099 abolished the {{CounterUpdateCells}} altogether, so {{CQLSSTableWriter}} counter cells are treated as remote shards during reconcile. Furthermore, the {{PartitionUpdate}} abstraction, also introduced by CASSANDRA-8099, merges all updates to the same partition before serialization on {{CQLSSTableWriter}}, so multiple updates to the same partition are indeed coalesced on 3.0.\n\nWe could fix the 2.1 behavior by special-casing the counter cell creation on {{UpdateParameters.addCounter()}} when {{Config.isClientMode()}} (an indication that the {{CQLSStableWriter}} is running) and creating a remote {{CounterCell}} instead, as done [in this commented block|https://github.com/apache/cassandra/blob/5449ad7b7e1b1930c07e0ea80ac3c3e2882633b5/src/java/org/apache/cassandra/cql3/UpdateParameters.java#L79]. This would mimic the 3.0 behavior and treat {{CQLSStableWriter}} counters as a remote shard.\n\nA hidden nit in this solution (and also on 3.0) is that a remote counter context is created for each {{CQLSSTableWriter}} session that updates counters, and afaik we currently have no way of discarding old counter contexts. So, new counter contexts created by {{CQLSStableWriter}} can potentially grow unbounded, what could affect performance significantly. Maybe after CASSANDRA-6506 we will be able to discard \"inactive\" counter contexts, and be able to support counters on {{CQLSStableWriter}} more efficiently.\n\nWhile we don't have a solution for cleaning inactive/temporary counter contexts, I propose we disable counters usage by default on {{CQLSSTableWriter}}, or at least make it harder to use (with a flag/warning, for instance), since it might still be useful for testing scenarios. What do you think [~slebresne], [~iamaleksey]? Do you see a better alternative?","from":"developer"},{"body":"I think we should just forbid the use of counters in CQLSSTableWriter (in all version to be clear). By design, counters require for their efficiency that shards are only created on replicas, so that the set of \"node id\" contained in any counter stay relatively low. CQLSSTableWriter just doesn't fulfill that assumption and I don't think there is anything we can do about it. Generating a fake \"node id\" for each CQLSSTableWriter (which is what ends up happening) has a very high potential to blow the context sizes and it's a bad idea: you are much better off just using normal CQL updates if you want to load counters (sure that may be slower, but slower is better than broken).\n\nOn top of that problem, there is the fact that if a given instance of CQLSSTableWriter increments the same counter twice, we *cannot* guarantee that both increment will be added together because we cannot guarantee both increment will make it into the same sstable. If they don't (make it into the same sstable), then whatever node the sstables ends up to will consider those 2 separate shards as \"remote\" and won't add these together (we'll probably end up with a broken shard warning in fact).\n\nSo really, the simplest and imo best solution is to add \"you cannot write sstable yourself with CQLSSTableWriter\" to the list of counters limitations.","from":"developer"},{"body":"Thanks for the quick feedback [~slebresne].\n\nAttached trivial patch and simple test forbidding counter updates on {{CQLSStableWriter}}. Mind reviewing/committing?\n\nThanks!","from":"developer"},{"body":"bq. So really, the simplest and imo best solution is to add \"you cannot write sstable yourself with CQLSSTableWriter\" to the list of counters limitations.\n\nBasically, this.","from":"developer"},{"body":"Committed as [3674ad9dab8f29173d7d4ee82488a8e9ea586240|https://github.com/apache/cassandra/commit/3674ad9dab8f29173d7d4ee82488a8e9ea586240] to 2.1, merged with 2.2, 3.0, 3.1, and trunk. Thanks.","from":"developer"}],"created":"2015-09-03T14:47:21.000+0000","description":"We use CQLSStableWriter to produce testing datasets.\n\nHere are the steps to reproduce this issue :\n\n1) definition of a table with counter\n{code}\nCREATE TABLE my_counter (\n my_id text,\n my_counter counter,\n PRIMARY KEY (my_id)\n)\n{code}\n\n2) with CQLSSTableWriter initialize this table (about 2millions entries) with this insert order (one insert / key only)\n{{UPDATE myks.my_counter SET my_counter = my_counter + ? WHERE my_id = ?}}\n\n3) load the files written by CQLSSTableWriter with sstableloader in your cassandra cluster (tested on a single node and a 3 nodes cluster)\n\n4) start a process that updates the counters (we used 3millions entries distributed on the key my_id)\n\n5) after a while try to query a key in the my_counter table\n{{cqlsh:myks> select * from my_counter where my_id='0000001';}}\nRequest did not complete within rpc_timeout.\n\nIn the logs of cassandra (2.0.12) :\n{code}\nERROR [CompactionExecutor:3] 2015-05-28 15:53:39,491 CassandraDaemon.java (line 258) Exception in thread Thread[CompactionExecutor:3,1,main]\njava.lang.AssertionError: Wrong class type.\n at org.apache.cassandra.db.CounterUpdateColumn.reconcile(CounterUpdateColumn.java:70)\n at org.apache.cassandra.db.ArrayBackedSortedColumns.resolveAgainst(ArrayBackedSortedColumns.java:147)\n at org.apache.cassandra.db.ArrayBackedSortedColumns.addColumn(ArrayBackedSortedColumns.java:126)\n at org.apache.cassandra.db.ColumnFamily.addColumn(ColumnFamily.java:121)\n at org.apache.cassandra.db.compaction.PrecompactedRow$1.reduce(PrecompactedRow.java:120)\n at org.apache.cassandra.db.compaction.PrecompactedRow$1.reduce(PrecompactedRow.java:115)\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:112)\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:98)\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n at org.apache.cassandra.db.filter.SliceQueryFilter.collectReducedColumns(SliceQueryFilter.java:191)\n at org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:144)\n at org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:103)\n at org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:85)\n at org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:196)\n at org.apache.cassandra.db.compaction.CompactionIterable$Reducer.getReduced(CompactionIterable.java:74)\n at org.apache.cassandra.db.compaction.CompactionIterable$Reducer.getReduced(CompactionIterable.java:55)\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:115)\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:98)\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:164)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:60)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:59)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:198)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\nIn the logs of cassandra v2.1.5 :\n{code}\nWARN [SharedPool-Worker-38] 2015-06-11 16:39:06,008 AbstractTracingAwareExecutorService.java:169 - Uncaught exception on thread Thread[SharedPool-Worker-38,5,\nmain]: {}\njava.lang.AssertionError: Wrong class type: class org.apache.cassandra.db.BufferCounterUpdateCell\n at org.apache.cassandra.db.AbstractCell.reconcileCounter(AbstractCell.java:211) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.BufferCounterCell.reconcile(BufferCounterCell.java:118) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.filter.QueryFilter$1.reduce(QueryFilter.java:122) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.filter.QueryFilter$1.reduce(QueryFilter.java:116) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:114) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:100) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143) ~[guava-16.0.1.jar:na]\n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138) ~[guava-16.0.1.jar:na]\n at org.apache.cassandra.db.filter.NamesQueryFilter.collectReducedColumns(NamesQueryFilter.java:100) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.filter.QueryFilter.collateColumns(QueryFilter.java:108) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:82) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:69) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CollationController.collectAllData(CollationController.java:314) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CollationController.getTopLevelColumns(CollationController.java:62) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.ColumnFamilyStore.getTopLevelColumns(ColumnFamilyStore.java:1900) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.ColumnFamilyStore.getColumnFamily(ColumnFamilyStore.java:1758) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.Keyspace.getRow(Keyspace.java:346) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.SliceByNamesReadCommand.getRow(SliceByNamesReadCommand.java:53) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CounterMutation.getCurrentValuesFromCFS(CounterMutation.java:262) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CounterMutation.getCurrentValues(CounterMutation.java:229) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CounterMutation.processModifications(CounterMutation.java:197) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.db.CounterMutation.apply(CounterMutation.java:124) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.service.StorageProxy$8.runMayThrow(StorageProxy.java:1155) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2191) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_45]\n at org.apache.cassandra.concurrent.AbstractTracingAwareExecutorService$FutureTask.run(AbstractTracingAwareExecutorService.java:164) ~[cassandra-all-2.1.5.469.jar:2.1.5.469]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [cassandra-all-2.1.5.469.jar:2.1.5.469]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_45]\n{code}\n\nsame exception as issue https://issues.apache.org/jira/browse/CASSANDRA-7188","issue_id":"12861596","key":"CASSANDRA-10258","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-11-10T13:49:44.000+0000","role":"fixed_distractor","summary":"Reject counter writes in CQLSSTableWriter"} {"case_id":"12903277","cluster":"DISTRACTOR-CASSANDRA-10477","comments":[{"body":"[~iamaleksey], who should be assigned to this?\n\n[~leonhardt], can you attach the system.log file from one of the affected nodes?","created":"2015-10-09T17:32:09.540+0000"},{"body":"[~philipthompson] I can't make the logfile publicly accessible but I can share it by mail. Let me know to whom I should send it.","created":"2015-10-12T13:27:08.688+0000"},{"body":"Once I find someone to work on this, they'll share their contact info with you. Thanks.","created":"2015-10-12T16:16:56.258+0000"},{"body":"[~leonhardt] can you email me the log file to the address in my profile?","created":"2015-11-10T21:20:35.951+0000"},{"body":"This assertion makes it look like the node either think it's broadcast address is that of another node, or alternatively it is connecting with itself which is causing it to submit hints to itself which is what the assertion is checking for. Neither condition should occur.\n\nIf you could also get me the output of \"netstat -tlnp\" along with node tool status, ring, and netstats when the problem occurs that would be helpful. You could do it now just to see if something shows up, but definitely when the problem occurs.","created":"2015-11-10T21:29:25.459+0000"},{"body":"Hello, we just observed this on a cluster running 2.1.11, Oracle Java 1.8.0_66.\n\nA single machine experienced this issue, causing our entire cluster to grind to a halt on any quorum operations.\n\nOur logs feature an extremely large number of:\n\n{code}\nERROR [EXPIRING-MAP-REAPER:1] 2015-11-11 05:10:22,894 CassandraDaemon.java:227 - Exception in threa\nd Thread[EXPIRING-MAP-REAPER:1,5,main]\njava.lang.AssertionError: /172.31.3.33\n at org.apache.cassandra.service.StorageProxy.submitHint(StorageProxy.java:949) ~[apache-cas\nsandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:383) ~[apache-ca\nssandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:363) ~[apache-ca\nssandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.utils.ExpiringMap$1.run(ExpiringMap.java:98) ~[apache-cassandra-2.1\n.11.jar:2.1.11]\n at org.apache.cassandra.concurrent.DebuggableScheduledThreadPoolExecutor$UncomplainingRunna\nble.run(DebuggableScheduledThreadPoolExecutor.java:118) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) [na:1.8.0_66]\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) [na:1.8.0_66]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(Schedule\ndThreadPoolExecutor.java:180) [na:1.8.0_66]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThread\nPoolExecutor.java:294) [na:1.8.0_66]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [na:1.8.\n0_66]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.\n0_66]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_66]\n{code}\n\nAdditionally, this is interspersed with every nearly every other neighbor node being marked down:\n\n{code}\nINFO [GossipStage:1] 2015-11-11 04:14:25,369 Gossiper.java:1020 - Node /172.31.55.172 has restarted, now UP\nINFO [GossipStage:1] 2015-11-11 04:14:25,369 TokenMetadata.java:414 - Updating topology for /172.31.55.172\nINFO [GossipStage:1] 2015-11-11 04:14:25,369 TokenMetadata.java:414 - Updating topology for /172.31.55.172\nINFO [GossipStage:1] 2015-11-11 04:14:25,370 StorageService.java:1698 - Node /172.31.55.172 state jump to normal\nINFO [GossipStage:1] 2015-11-11 04:14:25,372 TokenMetadata.java:414 - Updating topology for /172.31.55.172\nINFO [GossipStage:1] 2015-11-11 04:14:25,372 TokenMetadata.java:414 - Updating topology for /172.31.55.172\nINFO [SharedPool-Worker-3] 2015-11-11 04:14:25,531 Gossiper.java:987 - InetAddress /172.31.55.172 is now UP\nINFO [SharedPool-Worker-5] 2015-11-11 04:14:25,536 Gossiper.java:987 - InetAddress /172.31.55.172 is now UP\nINFO [SharedPool-Worker-3] 2015-11-11 04:14:25,536 Gossiper.java:987 - InetAddress /172.31.55.172 is now UP\nINFO [SharedPool-Worker-1] 2015-11-11 04:14:25,536 Gossiper.java:987 - InetAddress /172.31.55.172 is now UP\nINFO [SharedPool-Worker-4] 2015-11-11 04:14:25,536 Gossiper.java:987 - InetAddress /172.31.55.172 is now UP\nINFO [HANDSHAKE-/172.31.55.172] 2015-11-11 04:14:25,537 OutboundTcpConnection.java:485 - Handshaking version with /172.31.55.172\n[snipped]\nWARN [GossipTasks:1] 2015-11-11 04:18:26,379 Gossiper.java:747 - Gossip stage has 15 pending tasks\n; skipping status check (no nodes will be marked down)\nWARN [GossipTasks:1] 2015-11-11 04:18:27,480 Gossiper.java:747 - Gossip stage has 17 pending tasks\n; skipping status check (no nodes will be marked down)\nWARN [GossipTasks:1] 2015-11-11 04:18:28,580 Gossiper.java:747 - Gossip stage has 19 pending tasks\n; skipping status check (no nodes will be marked down)\nWARN [GossipTasks:1] 2015-11-11 04:18:29,681 Gossiper.java:747 - Gossip stage has 21 pending tasks\n; skipping status check (no nodes will be marked down)\nWARN [GossipTasks:1] 2015-11-11 04:18:30,781 Gossiper.java:747 - Gossip stage has 25 pending tasks\n; skipping status check (no nodes will be marked down)\n...\n{code}\n\nNo other nodes were restarted in this time frame.\n\nPlease let us know if there is any additional information we can provide.\n","created":"2015-11-11T05:37:34.449+0000"},{"body":"A few additional details:\n\nUnfortunately, I didn't get any data while the issue was happening. Afterwards, netstat, nodetool status, etc. are all nominal.\n\nDuring the period of time when this node was experiencing difficulty, no other nodes reported any unhealthy hosts. However, we do have our phi convict threshold tuned up from 8 to 10, due to running on AWS.\n\nThis event was localized to one node out of 12. Keyspace RF ranges from 3-5. Queries at LOCAL_QUORUM were timing out with insufficient responses.","created":"2015-11-11T06:36:35.751+0000"},{"body":"Just observed this issue again. Node was undergoing anticompaction when it occurred- once again brought the ring to a halt.\n\nCouldn't get all the required information due to the urgency of the situation, but did confirm that nodetool status reported the node as up with no issue (on another node).\n\nI have fresh logs to offer out-of-band to anyone who is investigating this issue- feel free to email or ping here.","created":"2015-11-17T03:09:09.708+0000"},{"body":"Can you attach or send me the yaml's your are using at each node?","created":"2015-11-17T19:33:37.144+0000"},{"body":"Theory time. [There is a path by which tasks that are supposed to go through the local hint process for inserts need to use.|https://github.com/apache/cassandra/blob/cassandra-2.1.9/src/java/org/apache/cassandra/service/StorageProxy.java#L1027] Since we have a case where an insert does not go down this path it kind of implies that one of the other call sites for inserts is incorrect and is going through the remote message service path.\n\nIt only happens when the node is overloaded and local inserts start timing out. The reason you don't normally see it is that local inserts probably don't time out most of the time. One thing you could do is increase the mutation timeouts to see if you can get past the low performance period without timing out and hitting this.\n\nHowever I think that the assertion is a symptom of a different problem and not the cause for the performance/availability issues. It's the canary in the coal mine letting you know this broken path is being taken due timeouts of local mutations.\n\nI think the thing to do is search the call hierarchy of {{[StorageProxy.submitHint|https://github.com/apache/cassandra/blob/cassandra-2.1.9/src/java/org/apache/cassandra/service/StorageProxy.java#L944]}} to find a path where it can be reached when timing out a local write. We know it's coming through MessageService in this instance which makes it a little trickier because the type of the callback isn't known. It looks like PAXOS might in some cases go down this path incorrectly.\n\nI am going to try running a few things locally with some assertions to see if I can get it to send a message with hint delivery to itself.","created":"2015-11-17T22:31:40.226+0000"},{"body":"Good news is that I am at least partially correct and PAXOS is heading down the road to submitting hints for the local node.\n\n[New failing utests from this assertion|https://github.com/apache/cassandra/compare/cassandra-2.1...aweisberg:CASSANDRA-10477-test?expand=1#diff-5e7d892105f1fa0706dbedf919b5dd99L46]\nhttp://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/1/testReport/junit/org.apache.cassandra.triggers/TriggersTest/executeTriggerOnCqlInsertWithConditions/\nhttp://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/1/testReport/junit/org.apache.cassandra.triggers/TriggersTest/executeTriggerOnCqlBatchWithConditions/\nhttp://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/1/testReport/junit/org.apache.cassandra.triggers/TriggersTest/executeTriggerOnThriftCASOperation/\n\nAlso several [failing dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-dtest/1/#showFailuresLink]\n\nI'll try getting the PAXOS code to do something similar to the insertLocal where it doesn't submit a real hint.\n","created":"2015-11-18T16:43:48.117+0000"},{"body":"[~bdeggleston] [~slebresne] can you chime in on whether I am on the right track here?\n\nShould {{[StorageProxy.commitPaxos|https://github.com/apache/cassandra/blob/cassandra-2.1.11/src/java/org/apache/cassandra/service/StorageProxy.java#L494]}} not be sending messages to the local node that are eligible for hinting on timeout?\n","created":"2015-11-18T17:04:16.612+0000"},{"body":"It seems like adding a paxos commit equivalent of StorageProxy.insertLocal, and submitting local commits that way would be the safest thing to do here. In theory, you should be able to add a check against the local address to StorageProxy.shouldHint and just drop the commit message if the node is overloaded, it should get back up to speed on the next paxos round. However there may be subtleties and edge cases that I'm not thinking of, so I don't want to recommend that without giving this more thought.","created":"2015-11-18T18:56:01.949+0000"},{"body":"Proposed fix\n\n|[2.1 code|https://github.com/apache/cassandra/compare/cassandra-2.1...aweisberg:CASSANDRA-10477-test]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-dtest/]|\n|[3.0 code|https://github.com/apache/cassandra/compare/cassandra-3.0...aweisberg:CASSANDRA-10477-3.0]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-dtest/]|","created":"2015-11-18T23:25:43.522+0000"},{"body":"* The failure detector will never return false for the local host, so the changes in the 2nd branch of commitPaxos are unnecessary.\n* We're kind of dodging the hint \"overload\" protection on the paxos path as we don't use {{sendToHintedEndpoints}} (which in particular makes the comment on {{commitPaxosLocal}} misleading as it suggests otherwise). I think the simplest solution is to move the overload test from {{sendToHintedEndpoints}} to some {{checkOverloaded()}} method and call that in {{commitPaxos}} too.\n* Instead of adding the {{droppable()}} method to {{LocalMutationRunnable}}, we should probably use {{MessagingService.DROPPABLE_VERBS.contains(verb)}}.\n* In theory, we could still run into the problem of that ticket if {{OPTIMIZE_LOCAL_REQUESTS}} is {{false}}. And in fact, I believe this option is unsafe since at least CASSANDRA-4753 as we somewhat strongly assume writes to the localhost do *not* go through {{MessagingService}}. So I would suggest ditching that option. Not only is it unsafe, but it's not used anywhere by the code and it's hardcoded so you have to change the code and recompile to even use it (which means I doubt anyone has even tried it in a long long time). And if we end up needing it in the future, we'll have to figure out how to make it safe.\n* Why isn't the added assertion in {{WriteCallbackInfo}} on 3.0 not using {{!shouldHint}} lie in the 2.1 patch?\n","created":"2015-12-04T10:52:23.442+0000"},{"body":"bq. We're kind of dodging the hint \"overload\" protection on the paxos path as we don't use sendToHintedEndpoints (which in particular makes the comment on commitPaxosLocal misleading as it suggests otherwise). I think the simplest solution is to move the overload test from sendToHintedEndpoints to some checkOverloaded() method and call that in commitPaxos too.\nWhich aspect of hint \"overload\" protection is missing? [I see it increments a counter which I thought was the signal upstream.|https://github.com/apache/cassandra/blob/cassandra-2.1/src/java/org/apache/cassandra/service/StorageProxy.java#L976]\n\nLooking at it further is it because it doesn't throw {{OverloadedException}}? So a better behavior would be to have the check and exception in a helper method and use that in commitPaxos() so that it can now throw {{OverloadedException}}?\n\nI do wonder what the unforeseen consequences of having {{CAS}} capable of throwing {{OE}} is going to do that we haven't seen or tested before. Where this gets interesting is that the read path now throws {{OE}} where it didn't before because apparently serial consistency reads can end up calling {{beginAndRepairPaxos}}. I need to take a close look at how we test this path to make sure it's going to behave well once exercised.\n\nbq. In theory, we could still run into the problem of that ticket if OPTIMIZE_LOCAL_REQUESTS is false. And in fact, I believe this option is unsafe since at least CASSANDRA-4753 as we somewhat strongly assume writes to the localhost do not go through MessagingService. So I would suggest ditching that option. Not only is it unsafe, but it's not used anywhere by the code and it's hardcoded so you have to change the code and recompile to even use it (which means I doubt anyone has even tried it in a long long time). And if we end up needing it in the future, we'll have to figure out how to make it safe.\nIt's already removed from 2.2. Yeah I don't think anyone uses it.\n\nbq. Why isn't the added assertion in WriteCallbackInfo on 3.0 not using !shouldHint lie in the 2.1 patch?\nIt's an oversight from merging.","created":"2015-12-04T17:13:17.892+0000"},{"body":"bq. Which aspect of hint \"overload\" protection is missing? I see it increments a counter which I thought was the signal upstream.\n\nThis is about whom is looking at said counter (to do something about it if it's too high). The normal write path is, and so incrementing the counter in CAS will potentially apply back-pressure on normal write, but not on CAS request themselves.\n\nbq. Looking at it further is it because it doesn't throw OverloadedException? So a better behavior would be to have the check and exception in a helper method and use that in commitPaxos() so that it can now throw OverloadedException?\n\nExactly.\n\nbq. I do wonder what the unforeseen consequences of having CAS capable of throwing OE is going to do that we haven't seen or tested before.\n\nIt's a good question, and to be honest I'm not sure we have any test that cover {{OverloadException}} at all (but I could be wrong). But in general, the commit part of Paxos is not very \"sensible\": worst case, if not enough replica get the commit, the next serial operation (including a read) on the partition will re-commit. So the main question is whether potentially throwing {{OverloadedException}} would surprise people. I would argue it shouldn't because normal writes can do so and we never specified it was any different for CAS. That said, if we're uncomfortable with it, I'm totally fine committing that part of the change only in 3.2 (aka trunk currently).\n\nbq. the read path now throws OE where it didn't before\n\nRight. That's probably more justification for keeping that part in 3.2 only.\n\n","created":"2015-12-04T17:57:31.819+0000"},{"body":"bq. Why isn't the added assertion in WriteCallbackInfo on 3.0 not using !shouldHint lie in the 2.1 patch?\nThis turns out to be because shouldHint() has additional stuff that the assertion doesn't want. The assertion doesn't care if hints are disabled along with several of the other things that are added.\n\nI think I managed to shuffle everything correctly. Going to let the tests run and prognosticate on how I want to test OE. I grepped the dtests and unit tests for OverloadedException and didn't get a single hit!\n\n|[2.1 code|https://github.com/apache/cassandra/compare/cassandra-2.1...aweisberg:CASSANDRA-10477-test]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-dtest/]|\n|[3.0 code|https://github.com/apache/cassandra/compare/cassandra-3.0...aweisberg:CASSANDRA-10477-3.0]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-dtest/]|\n|[Trunk code|https://github.com/apache/cassandra/compare/trunk...aweisberg:CASSANDRA-10477-trunk]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-dtest/]|","created":"2015-12-04T20:39:25.048+0000"},{"body":"bq. The assertion doesn't care if hints are disabled along with several of the other things that are added.\n\nFirst, I still don't understand why it's not consistent between 2.1 and 3.0. As far as I can tell, the {{WriteCallbackInfo.shouldHint()}} mostly method calls {{StorageProxy.shouldHint()}} which does pretty much the same thing in both versions. Second, I'd argue the assertion _must_ use {{!shouldHint()}} because what we're trying to assert is that {{submitHint}} is never called for localhost on the expiration of a callback, and that depends on the result of {{shouldHint()}}. That said, I think it would almost be better to have the assertion just be {{!target.equals(FBUtilities.getBroadcastAddress())}} as we're basically saying a local write should always use the specific local path, not {{MessagingService}}. In any case, I think the assertion is worth a quick comment to explain why we're asserting that here.\n\nThe rest of the changes lgtm, but the unit tests on 3.0 don't seem to have run due to some problem with an {{@Override}}.\n\nbq. and prognosticate on how I want to test OE\n\nThe lack of coverage of OE is certainly something we should fix (it's not trivial though), but I would suggest not blocking that fix for that since it's not directly related (meaning, we should probably open a separate ticket for it).\n","created":"2015-12-07T09:00:30.063+0000"},{"body":"I'm having the same problem on 2.2.3:\n{code}\n10:55:54.203 [ERROR] CassandraDaemon - Exception in thread Thread[EXPIRING-MAP-REAPER:1,5,UCS-Threads] java.lang.AssertionError: /\n at org.apache.cassandra.service.StorageProxy.submitHint(StorageProxy.java:978)\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:399)\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:379)\n at org.apache.cassandra.utils.ExpiringMap$1.run(ExpiringMap.java:98)\n at org.apache.cassandra.concurrent.DebuggableScheduledThreadPoolExecutor$UncomplainingRunnable.run(DebuggableScheduledThreadPoolExecutor.java:118)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\nWhat is the status on this issue? Can I expect this to be fixed in 2.2.x branch?","created":"2016-01-11T11:17:46.037+0000"},{"body":"[~aweisberg] the ball is in your court I believe.","created":"2016-01-11T11:33:53.535+0000"},{"body":"The tests are passing, but enough time has passed that I should rebase and test again. Will do that today.","created":"2016-01-11T15:52:02.554+0000"},{"body":"Can you also answer my comments on the assertion?","created":"2016-01-11T16:17:51.724+0000"},{"body":"I agree the assertion should just be on the address. Already made the change back in December just need to get the tests done.","created":"2016-01-11T16:28:13.374+0000"},{"body":"Rebased, updated commit message, updated test matrix, started tests.\n|[2.1 code|https://github.com/apache/cassandra/compare/cassandra-2.1...aweisberg:CASSANDRA-10477-test]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-dtest/]|\n|[3.0 code|https://github.com/apache/cassandra/compare/cassandra-3.0...aweisberg:CASSANDRA-10477-3.0]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-dtest/]|\n|[3.3 code|https://github.com/apache/cassandra/compare/cassandra-3.3...aweisberg:CASSANDRA-10477-3.3]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.3-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.3-dtest/]|\n|[Trunk code|https://github.com/apache/cassandra/compare/trunk...aweisberg:CASSANDRA-10477-trunk]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-dtest/]|","created":"2016-01-11T17:33:00.226+0000"},{"body":"Will it be fixed in 2.2.x too?","created":"2016-01-12T08:53:44.531+0000"},{"body":"[~aweisberg] I think you have a bad merge on 3.0 (though strangely the 3.3 and trunk branches seem fine), the test run failed at compilation time.\n\nbq. Will it be fixed in 2.2.x too?\n\nIt will.","created":"2016-01-12T13:36:43.801+0000"},{"body":"Fixed 3.0 compilation issue. Added 2.2 branch and updated test matrix.\n|[2.1 code|https://github.com/apache/cassandra/compare/cassandra-2.1...aweisberg:CASSANDRA-10477-test]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-dtest/]|\n|[2.2 code|https://github.com/apache/cassandra/compare/cassandra-2.2...aweisberg:CASSANDRA-10477-2.2]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-2.2-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-2.2-dtest/]|\n|[3.0 code|https://github.com/apache/cassandra/compare/cassandra-3.0...aweisberg:CASSANDRA-10477-3.0]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-dtest/]|\n|[3.3 code|https://github.com/apache/cassandra/compare/cassandra-3.3...aweisberg:CASSANDRA-10477-3.3]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.3-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.3-dtest/]|\n|[Trunk code|https://github.com/apache/cassandra/compare/trunk...aweisberg:CASSANDRA-10477-trunk]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-dtest/]|","created":"2016-01-12T16:05:40.137+0000"},{"body":"Great, thank you.","created":"2016-01-12T16:13:00.743+0000"},{"body":"Committed, thanks.","created":"2016-01-13T10:38:44.524+0000"}],"conversations":[{"body":"A few days after updating from 2.0.15 to 2.1.9 we have the following log entry on 2 of 5 machines:\n\n{noformat}\nERROR [EXPIRING-MAP-REAPER:1] 2015-10-07 17:01:08,041 CassandraDaemon.java:223 - Exception in thread Thread[EXPIRING-MAP-REAPER:1,5,main]\njava.lang.AssertionError: /192.168.11.88\n at org.apache.cassandra.service.StorageProxy.submitHint(StorageProxy.java:949) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:383) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:363) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.utils.ExpiringMap$1.run(ExpiringMap.java:98) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.concurrent.DebuggableScheduledThreadPoolExecutor$UncomplainingRunnable.run(DebuggableScheduledThreadPoolExecutor.java:118) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) [na:1.8.0_45]\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) [na:1.8.0_45]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180) [na:1.8.0_45]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294) [na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_45]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_45]\n{noformat}\n\n192.168.11.88 is the broadcast address of the local machine.\n\nWhen this is logged the read request latency of the whole cluster becomes very bad, from 6 ms/op to more than 100 ms/op according to OpsCenter. Clients get a lot of timeouts. We need to restart the affected Cassandra node to get back normal read latencies. It seems write latency is not affected.\n\nDisabling hinted handoff using {{nodetool disablehandoff}} only prevents the assert from being logged. At some point the read latency becomes bad again. Restarting the node where hinted handoff was disabled results in the read latency being better again.","from":"reporter","subject":"java.lang.AssertionError in StorageProxy.submitHint"},{"body":"[~iamaleksey], who should be assigned to this?\n\n[~leonhardt], can you attach the system.log file from one of the affected nodes?","from":"developer"},{"body":"[~philipthompson] I can't make the logfile publicly accessible but I can share it by mail. Let me know to whom I should send it.","from":"developer"},{"body":"Once I find someone to work on this, they'll share their contact info with you. Thanks.","from":"developer"},{"body":"[~leonhardt] can you email me the log file to the address in my profile?","from":"developer"},{"body":"This assertion makes it look like the node either think it's broadcast address is that of another node, or alternatively it is connecting with itself which is causing it to submit hints to itself which is what the assertion is checking for. Neither condition should occur.\n\nIf you could also get me the output of \"netstat -tlnp\" along with node tool status, ring, and netstats when the problem occurs that would be helpful. You could do it now just to see if something shows up, but definitely when the problem occurs.","from":"developer"},{"body":"Hello, we just observed this on a cluster running 2.1.11, Oracle Java 1.8.0_66.\n\nA single machine experienced this issue, causing our entire cluster to grind to a halt on any quorum operations.\n\nOur logs feature an extremely large number of:\n\n{code}\nERROR [EXPIRING-MAP-REAPER:1] 2015-11-11 05:10:22,894 CassandraDaemon.java:227 - Exception in threa\nd Thread[EXPIRING-MAP-REAPER:1,5,main]\njava.lang.AssertionError: /172.31.3.33\n at org.apache.cassandra.service.StorageProxy.submitHint(StorageProxy.java:949) ~[apache-cas\nsandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:383) ~[apache-ca\nssandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:363) ~[apache-ca\nssandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.utils.ExpiringMap$1.run(ExpiringMap.java:98) ~[apache-cassandra-2.1\n.11.jar:2.1.11]\n at org.apache.cassandra.concurrent.DebuggableScheduledThreadPoolExecutor$UncomplainingRunna\nble.run(DebuggableScheduledThreadPoolExecutor.java:118) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) [na:1.8.0_66]\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) [na:1.8.0_66]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(Schedule\ndThreadPoolExecutor.java:180) [na:1.8.0_66]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThread\nPoolExecutor.java:294) [na:1.8.0_66]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [na:1.8.\n0_66]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.\n0_66]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_66]\n{code}\n\nAdditionally, this is interspersed with every nearly every other neighbor node being marked down:\n\n{code}\nINFO [GossipStage:1] 2015-11-11 04:14:25,369 Gossiper.java:1020 - Node /172.31.55.172 has restarted, now UP\nINFO [GossipStage:1] 2015-11-11 04:14:25,369 TokenMetadata.java:414 - Updating topology for /172.31.55.172\nINFO [GossipStage:1] 2015-11-11 04:14:25,369 TokenMetadata.java:414 - Updating topology for /172.31.55.172\nINFO [GossipStage:1] 2015-11-11 04:14:25,370 StorageService.java:1698 - Node /172.31.55.172 state jump to normal\nINFO [GossipStage:1] 2015-11-11 04:14:25,372 TokenMetadata.java:414 - Updating topology for /172.31.55.172\nINFO [GossipStage:1] 2015-11-11 04:14:25,372 TokenMetadata.java:414 - Updating topology for /172.31.55.172\nINFO [SharedPool-Worker-3] 2015-11-11 04:14:25,531 Gossiper.java:987 - InetAddress /172.31.55.172 is now UP\nINFO [SharedPool-Worker-5] 2015-11-11 04:14:25,536 Gossiper.java:987 - InetAddress /172.31.55.172 is now UP\nINFO [SharedPool-Worker-3] 2015-11-11 04:14:25,536 Gossiper.java:987 - InetAddress /172.31.55.172 is now UP\nINFO [SharedPool-Worker-1] 2015-11-11 04:14:25,536 Gossiper.java:987 - InetAddress /172.31.55.172 is now UP\nINFO [SharedPool-Worker-4] 2015-11-11 04:14:25,536 Gossiper.java:987 - InetAddress /172.31.55.172 is now UP\nINFO [HANDSHAKE-/172.31.55.172] 2015-11-11 04:14:25,537 OutboundTcpConnection.java:485 - Handshaking version with /172.31.55.172\n[snipped]\nWARN [GossipTasks:1] 2015-11-11 04:18:26,379 Gossiper.java:747 - Gossip stage has 15 pending tasks\n; skipping status check (no nodes will be marked down)\nWARN [GossipTasks:1] 2015-11-11 04:18:27,480 Gossiper.java:747 - Gossip stage has 17 pending tasks\n; skipping status check (no nodes will be marked down)\nWARN [GossipTasks:1] 2015-11-11 04:18:28,580 Gossiper.java:747 - Gossip stage has 19 pending tasks\n; skipping status check (no nodes will be marked down)\nWARN [GossipTasks:1] 2015-11-11 04:18:29,681 Gossiper.java:747 - Gossip stage has 21 pending tasks\n; skipping status check (no nodes will be marked down)\nWARN [GossipTasks:1] 2015-11-11 04:18:30,781 Gossiper.java:747 - Gossip stage has 25 pending tasks\n; skipping status check (no nodes will be marked down)\n...\n{code}\n\nNo other nodes were restarted in this time frame.\n\nPlease let us know if there is any additional information we can provide.\n","from":"developer"},{"body":"A few additional details:\n\nUnfortunately, I didn't get any data while the issue was happening. Afterwards, netstat, nodetool status, etc. are all nominal.\n\nDuring the period of time when this node was experiencing difficulty, no other nodes reported any unhealthy hosts. However, we do have our phi convict threshold tuned up from 8 to 10, due to running on AWS.\n\nThis event was localized to one node out of 12. Keyspace RF ranges from 3-5. Queries at LOCAL_QUORUM were timing out with insufficient responses.","from":"developer"},{"body":"Just observed this issue again. Node was undergoing anticompaction when it occurred- once again brought the ring to a halt.\n\nCouldn't get all the required information due to the urgency of the situation, but did confirm that nodetool status reported the node as up with no issue (on another node).\n\nI have fresh logs to offer out-of-band to anyone who is investigating this issue- feel free to email or ping here.","from":"developer"},{"body":"Can you attach or send me the yaml's your are using at each node?","from":"developer"},{"body":"Theory time. [There is a path by which tasks that are supposed to go through the local hint process for inserts need to use.|https://github.com/apache/cassandra/blob/cassandra-2.1.9/src/java/org/apache/cassandra/service/StorageProxy.java#L1027] Since we have a case where an insert does not go down this path it kind of implies that one of the other call sites for inserts is incorrect and is going through the remote message service path.\n\nIt only happens when the node is overloaded and local inserts start timing out. The reason you don't normally see it is that local inserts probably don't time out most of the time. One thing you could do is increase the mutation timeouts to see if you can get past the low performance period without timing out and hitting this.\n\nHowever I think that the assertion is a symptom of a different problem and not the cause for the performance/availability issues. It's the canary in the coal mine letting you know this broken path is being taken due timeouts of local mutations.\n\nI think the thing to do is search the call hierarchy of {{[StorageProxy.submitHint|https://github.com/apache/cassandra/blob/cassandra-2.1.9/src/java/org/apache/cassandra/service/StorageProxy.java#L944]}} to find a path where it can be reached when timing out a local write. We know it's coming through MessageService in this instance which makes it a little trickier because the type of the callback isn't known. It looks like PAXOS might in some cases go down this path incorrectly.\n\nI am going to try running a few things locally with some assertions to see if I can get it to send a message with hint delivery to itself.","from":"developer"},{"body":"Good news is that I am at least partially correct and PAXOS is heading down the road to submitting hints for the local node.\n\n[New failing utests from this assertion|https://github.com/apache/cassandra/compare/cassandra-2.1...aweisberg:CASSANDRA-10477-test?expand=1#diff-5e7d892105f1fa0706dbedf919b5dd99L46]\nhttp://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/1/testReport/junit/org.apache.cassandra.triggers/TriggersTest/executeTriggerOnCqlInsertWithConditions/\nhttp://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/1/testReport/junit/org.apache.cassandra.triggers/TriggersTest/executeTriggerOnCqlBatchWithConditions/\nhttp://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/1/testReport/junit/org.apache.cassandra.triggers/TriggersTest/executeTriggerOnThriftCASOperation/\n\nAlso several [failing dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-dtest/1/#showFailuresLink]\n\nI'll try getting the PAXOS code to do something similar to the insertLocal where it doesn't submit a real hint.\n","from":"developer"},{"body":"[~bdeggleston] [~slebresne] can you chime in on whether I am on the right track here?\n\nShould {{[StorageProxy.commitPaxos|https://github.com/apache/cassandra/blob/cassandra-2.1.11/src/java/org/apache/cassandra/service/StorageProxy.java#L494]}} not be sending messages to the local node that are eligible for hinting on timeout?\n","from":"developer"},{"body":"It seems like adding a paxos commit equivalent of StorageProxy.insertLocal, and submitting local commits that way would be the safest thing to do here. In theory, you should be able to add a check against the local address to StorageProxy.shouldHint and just drop the commit message if the node is overloaded, it should get back up to speed on the next paxos round. However there may be subtleties and edge cases that I'm not thinking of, so I don't want to recommend that without giving this more thought.","from":"developer"},{"body":"Proposed fix\n\n|[2.1 code|https://github.com/apache/cassandra/compare/cassandra-2.1...aweisberg:CASSANDRA-10477-test]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-dtest/]|\n|[3.0 code|https://github.com/apache/cassandra/compare/cassandra-3.0...aweisberg:CASSANDRA-10477-3.0]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-dtest/]|","from":"developer"},{"body":"* The failure detector will never return false for the local host, so the changes in the 2nd branch of commitPaxos are unnecessary.\n* We're kind of dodging the hint \"overload\" protection on the paxos path as we don't use {{sendToHintedEndpoints}} (which in particular makes the comment on {{commitPaxosLocal}} misleading as it suggests otherwise). I think the simplest solution is to move the overload test from {{sendToHintedEndpoints}} to some {{checkOverloaded()}} method and call that in {{commitPaxos}} too.\n* Instead of adding the {{droppable()}} method to {{LocalMutationRunnable}}, we should probably use {{MessagingService.DROPPABLE_VERBS.contains(verb)}}.\n* In theory, we could still run into the problem of that ticket if {{OPTIMIZE_LOCAL_REQUESTS}} is {{false}}. And in fact, I believe this option is unsafe since at least CASSANDRA-4753 as we somewhat strongly assume writes to the localhost do *not* go through {{MessagingService}}. So I would suggest ditching that option. Not only is it unsafe, but it's not used anywhere by the code and it's hardcoded so you have to change the code and recompile to even use it (which means I doubt anyone has even tried it in a long long time). And if we end up needing it in the future, we'll have to figure out how to make it safe.\n* Why isn't the added assertion in {{WriteCallbackInfo}} on 3.0 not using {{!shouldHint}} lie in the 2.1 patch?\n","from":"developer"},{"body":"bq. We're kind of dodging the hint \"overload\" protection on the paxos path as we don't use sendToHintedEndpoints (which in particular makes the comment on commitPaxosLocal misleading as it suggests otherwise). I think the simplest solution is to move the overload test from sendToHintedEndpoints to some checkOverloaded() method and call that in commitPaxos too.\nWhich aspect of hint \"overload\" protection is missing? [I see it increments a counter which I thought was the signal upstream.|https://github.com/apache/cassandra/blob/cassandra-2.1/src/java/org/apache/cassandra/service/StorageProxy.java#L976]\n\nLooking at it further is it because it doesn't throw {{OverloadedException}}? So a better behavior would be to have the check and exception in a helper method and use that in commitPaxos() so that it can now throw {{OverloadedException}}?\n\nI do wonder what the unforeseen consequences of having {{CAS}} capable of throwing {{OE}} is going to do that we haven't seen or tested before. Where this gets interesting is that the read path now throws {{OE}} where it didn't before because apparently serial consistency reads can end up calling {{beginAndRepairPaxos}}. I need to take a close look at how we test this path to make sure it's going to behave well once exercised.\n\nbq. In theory, we could still run into the problem of that ticket if OPTIMIZE_LOCAL_REQUESTS is false. And in fact, I believe this option is unsafe since at least CASSANDRA-4753 as we somewhat strongly assume writes to the localhost do not go through MessagingService. So I would suggest ditching that option. Not only is it unsafe, but it's not used anywhere by the code and it's hardcoded so you have to change the code and recompile to even use it (which means I doubt anyone has even tried it in a long long time). And if we end up needing it in the future, we'll have to figure out how to make it safe.\nIt's already removed from 2.2. Yeah I don't think anyone uses it.\n\nbq. Why isn't the added assertion in WriteCallbackInfo on 3.0 not using !shouldHint lie in the 2.1 patch?\nIt's an oversight from merging.","from":"developer"},{"body":"bq. Which aspect of hint \"overload\" protection is missing? I see it increments a counter which I thought was the signal upstream.\n\nThis is about whom is looking at said counter (to do something about it if it's too high). The normal write path is, and so incrementing the counter in CAS will potentially apply back-pressure on normal write, but not on CAS request themselves.\n\nbq. Looking at it further is it because it doesn't throw OverloadedException? So a better behavior would be to have the check and exception in a helper method and use that in commitPaxos() so that it can now throw OverloadedException?\n\nExactly.\n\nbq. I do wonder what the unforeseen consequences of having CAS capable of throwing OE is going to do that we haven't seen or tested before.\n\nIt's a good question, and to be honest I'm not sure we have any test that cover {{OverloadException}} at all (but I could be wrong). But in general, the commit part of Paxos is not very \"sensible\": worst case, if not enough replica get the commit, the next serial operation (including a read) on the partition will re-commit. So the main question is whether potentially throwing {{OverloadedException}} would surprise people. I would argue it shouldn't because normal writes can do so and we never specified it was any different for CAS. That said, if we're uncomfortable with it, I'm totally fine committing that part of the change only in 3.2 (aka trunk currently).\n\nbq. the read path now throws OE where it didn't before\n\nRight. That's probably more justification for keeping that part in 3.2 only.\n\n","from":"developer"},{"body":"bq. Why isn't the added assertion in WriteCallbackInfo on 3.0 not using !shouldHint lie in the 2.1 patch?\nThis turns out to be because shouldHint() has additional stuff that the assertion doesn't want. The assertion doesn't care if hints are disabled along with several of the other things that are added.\n\nI think I managed to shuffle everything correctly. Going to let the tests run and prognosticate on how I want to test OE. I grepped the dtests and unit tests for OverloadedException and didn't get a single hit!\n\n|[2.1 code|https://github.com/apache/cassandra/compare/cassandra-2.1...aweisberg:CASSANDRA-10477-test]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-dtest/]|\n|[3.0 code|https://github.com/apache/cassandra/compare/cassandra-3.0...aweisberg:CASSANDRA-10477-3.0]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-dtest/]|\n|[Trunk code|https://github.com/apache/cassandra/compare/trunk...aweisberg:CASSANDRA-10477-trunk]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-dtest/]|","from":"developer"},{"body":"bq. The assertion doesn't care if hints are disabled along with several of the other things that are added.\n\nFirst, I still don't understand why it's not consistent between 2.1 and 3.0. As far as I can tell, the {{WriteCallbackInfo.shouldHint()}} mostly method calls {{StorageProxy.shouldHint()}} which does pretty much the same thing in both versions. Second, I'd argue the assertion _must_ use {{!shouldHint()}} because what we're trying to assert is that {{submitHint}} is never called for localhost on the expiration of a callback, and that depends on the result of {{shouldHint()}}. That said, I think it would almost be better to have the assertion just be {{!target.equals(FBUtilities.getBroadcastAddress())}} as we're basically saying a local write should always use the specific local path, not {{MessagingService}}. In any case, I think the assertion is worth a quick comment to explain why we're asserting that here.\n\nThe rest of the changes lgtm, but the unit tests on 3.0 don't seem to have run due to some problem with an {{@Override}}.\n\nbq. and prognosticate on how I want to test OE\n\nThe lack of coverage of OE is certainly something we should fix (it's not trivial though), but I would suggest not blocking that fix for that since it's not directly related (meaning, we should probably open a separate ticket for it).\n","from":"developer"},{"body":"I'm having the same problem on 2.2.3:\n{code}\n10:55:54.203 [ERROR] CassandraDaemon - Exception in thread Thread[EXPIRING-MAP-REAPER:1,5,UCS-Threads] java.lang.AssertionError: /\n at org.apache.cassandra.service.StorageProxy.submitHint(StorageProxy.java:978)\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:399)\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:379)\n at org.apache.cassandra.utils.ExpiringMap$1.run(ExpiringMap.java:98)\n at org.apache.cassandra.concurrent.DebuggableScheduledThreadPoolExecutor$UncomplainingRunnable.run(DebuggableScheduledThreadPoolExecutor.java:118)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\nWhat is the status on this issue? Can I expect this to be fixed in 2.2.x branch?","from":"developer"},{"body":"[~aweisberg] the ball is in your court I believe.","from":"developer"},{"body":"The tests are passing, but enough time has passed that I should rebase and test again. Will do that today.","from":"developer"},{"body":"Can you also answer my comments on the assertion?","from":"developer"},{"body":"I agree the assertion should just be on the address. Already made the change back in December just need to get the tests done.","from":"developer"},{"body":"Rebased, updated commit message, updated test matrix, started tests.\n|[2.1 code|https://github.com/apache/cassandra/compare/cassandra-2.1...aweisberg:CASSANDRA-10477-test]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-dtest/]|\n|[3.0 code|https://github.com/apache/cassandra/compare/cassandra-3.0...aweisberg:CASSANDRA-10477-3.0]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-dtest/]|\n|[3.3 code|https://github.com/apache/cassandra/compare/cassandra-3.3...aweisberg:CASSANDRA-10477-3.3]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.3-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.3-dtest/]|\n|[Trunk code|https://github.com/apache/cassandra/compare/trunk...aweisberg:CASSANDRA-10477-trunk]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-dtest/]|","from":"developer"},{"body":"Will it be fixed in 2.2.x too?","from":"developer"},{"body":"[~aweisberg] I think you have a bad merge on 3.0 (though strangely the 3.3 and trunk branches seem fine), the test run failed at compilation time.\n\nbq. Will it be fixed in 2.2.x too?\n\nIt will.","from":"developer"},{"body":"Fixed 3.0 compilation issue. Added 2.2 branch and updated test matrix.\n|[2.1 code|https://github.com/apache/cassandra/compare/cassandra-2.1...aweisberg:CASSANDRA-10477-test]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-test-dtest/]|\n|[2.2 code|https://github.com/apache/cassandra/compare/cassandra-2.2...aweisberg:CASSANDRA-10477-2.2]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-2.2-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-2.2-dtest/]|\n|[3.0 code|https://github.com/apache/cassandra/compare/cassandra-3.0...aweisberg:CASSANDRA-10477-3.0]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.0-dtest/]|\n|[3.3 code|https://github.com/apache/cassandra/compare/cassandra-3.3...aweisberg:CASSANDRA-10477-3.3]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.3-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-3.3-dtest/]|\n|[Trunk code|https://github.com/apache/cassandra/compare/trunk...aweisberg:CASSANDRA-10477-trunk]|[utests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/aweisberg/job/aweisberg-CASSANDRA-10477-trunk-dtest/]|","from":"developer"},{"body":"Great, thank you.","from":"developer"},{"body":"Committed, thanks.","from":"developer"}],"created":"2015-10-08T07:57:05.000+0000","description":"A few days after updating from 2.0.15 to 2.1.9 we have the following log entry on 2 of 5 machines:\n\n{noformat}\nERROR [EXPIRING-MAP-REAPER:1] 2015-10-07 17:01:08,041 CassandraDaemon.java:223 - Exception in thread Thread[EXPIRING-MAP-REAPER:1,5,main]\njava.lang.AssertionError: /192.168.11.88\n at org.apache.cassandra.service.StorageProxy.submitHint(StorageProxy.java:949) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:383) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.net.MessagingService$5.apply(MessagingService.java:363) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.utils.ExpiringMap$1.run(ExpiringMap.java:98) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.concurrent.DebuggableScheduledThreadPoolExecutor$UncomplainingRunnable.run(DebuggableScheduledThreadPoolExecutor.java:118) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) [na:1.8.0_45]\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) [na:1.8.0_45]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180) [na:1.8.0_45]\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294) [na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_45]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_45]\n{noformat}\n\n192.168.11.88 is the broadcast address of the local machine.\n\nWhen this is logged the read request latency of the whole cluster becomes very bad, from 6 ms/op to more than 100 ms/op according to OpsCenter. Clients get a lot of timeouts. We need to restart the affected Cassandra node to get back normal read latencies. It seems write latency is not affected.\n\nDisabling hinted handoff using {{nodetool disablehandoff}} only prevents the assert from being logged. At some point the read latency becomes bad again. Restarting the node where hinted handoff was disabled results in the read latency being better again.","issue_id":"12903277","key":"CASSANDRA-10477","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-01-13T10:38:44.000+0000","role":"fixed_distractor","summary":"java.lang.AssertionError in StorageProxy.submitHint"} {"case_id":"12905207","cluster":"DISTRACTOR-CASSANDRA-10534","comments":[{"body":"[~benedict], [~sharvanath] analysis is correct: since CASSANDRA-6916 we no longer fsync compression metadata after writing it. I've attached a small [patch|https://github.com/stef1927/cassandra/commits/10534-2.1] that should fix this, can you take a look?\n\nIf the patch is fine I will run CI on the 2.1+ branches. ","created":"2015-10-16T07:06:17.140+0000"},{"body":"Hi [~sharvanath]: thanks for taking the time to strace this and find our (my) mistake. \n\n[~stefania]: thanks for providing a patch. -LGTM. I'll commit once we have clean CI results.- Why did you put the {{sync()}} call in the finally block before the close, instead of in the try block? At the very least we should ensure the close is called after, but it seems easiest to place it in the try block.","created":"2015-10-16T07:09:12.134+0000"},{"body":"[~benedict] [~Stefania] thanks for taking quick action on it.","created":"2015-10-16T07:29:47.314+0000"},{"body":"[~Benedict]: I merely wanted to fsync in case of partial write but if it looks too unusual we can have it in the try block. I amended the commit and force pushed. The 2.2 patch is a rewrite because the code is too divergent. I believe the FOS close is idempotent so we are OK if we close it twice but please double check. The 2.2 patch then merges without conflicts into 3.0.\n\nCI should eventually appear here:\n\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-2.1-dtest\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-2.1-testall\n\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-2.2-dtest\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-2.2-testall\n\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-3.0-dtest\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-3.0-testall\n","created":"2015-10-16T07:39:37.516+0000"},{"body":"In all cases we will need to call flush before calling sync, since we have a buffered writer. {{close}} is idempotent, so that should not be a problem.","created":"2015-10-16T07:47:56.328+0000"},{"body":"I've added the call to {{flush}} in a separate commit. \n\nI tried to abort the CI jobs and restart new ones but it seems the abort only removed the jobs in the queue so it's the old jobs (without flush) that are running at the moment. If they are aborted later on I will restart them or if I am offline you can restart yourself with cassci {{!build stef1927-10534-3.0-dtest}} etc.","created":"2015-10-16T08:24:07.385+0000"},{"body":"I've rebased and restarted all 3 jobs.","created":"2015-10-19T00:48:49.957+0000"},{"body":"This bug also happened in 2.1.10.","created":"2015-11-05T09:17:41.762+0000"},{"body":"[~benedict], I've rebased and re-run CI on 2.1 and 2.2. It looks OK to me (a few failing tests but inline with the main branches). \n\nLet me know if there is anything else to be done before committing.","created":"2015-11-10T05:40:00.053+0000"},{"body":"I don't see why we would ever not sync the other files? Are they really not necessary for the sstable to be readable? If they are required then they need to be synced as well otherwise we are going to take actions based on the sstable being durable/readable when it isn't really readable.","created":"2015-11-11T17:54:40.170+0000"},{"body":"_TOC.txt_ is only used by standalone tools, we do a listing of the folder in the CFS constructor. I could not find where the digest file is read, at least not on the 2.1 code. Other components I have not checked yet. For sure index and data files are sync-ed.\n\nWe should probably sync all sstable components that we write but should we fix this regression and open a new ticket for better visibility?","created":"2015-11-12T07:14:06.363+0000"},{"body":"Ok. +1 LGTM.\n\nIt looks like the other components all use SequentialWriter which if you use the transaction stuff does the right thing without any extra work. Except we don't use the transaction proxy for the other metadata components instead we call finish manually which goes through the transaction proxy.\n\nI can't quite tell since it looks like it can throw and cause other code not to execute.","created":"2015-11-12T20:38:34.249+0000"},{"body":"Thank you for the review and for checking the remaining components. I agree with your analysis. \n\nThese are the components defined in {{SSTableWriter.components()}}:\n\n|| Component || Notes |\n| Component.DATA | Sync-ed, SequentialWriter | \n| Component.PRIMARY_INDEX | Sync-ed, SequentialWriter |\n| Component.STATS | Sync-ed, SequentialWriter |\n| Component.SUMMARY | {{SSTableReader.saveSummary()}}, called in finish, not sync-ed but we write a magic number at the end and we regenerate the summary when loading it if we don't find this magic number, |\n| Component.TOC | Not sync-ed but it's read only by standalone tools | \n| Component.DIGEST | Written by {{DataIntegrityMetadata.ChecksumWriter}}, not sync-ed and not used but intended for users so they can validate uncompressed data files via sha1sum. In 2.2 this becomes the adler32 checksum that can be verified with nodetool verify or the standalone verifier.|\n| Component.FILTER | Written and sync-ed manually in {{IndexWrit er.close()}} |\n| Component.COMPRESSION_INFO | To be sync-ed by this patch in {{ompressionMetadata.Writer.close()}} |\n| Component.CRC | Sync-ed, SequentialWriter |\n\nbq. I can't quite tell since it looks like it can throw and cause other code not to execute.\n\nYes but this was improved in 2.2 with {{LifecycleTransaction}} and in 3.0 even further with the removal of unfinished left overs via {{LogTransaction}}.\n\nSo, IMO, we could have problems with standalone tools and we should probably sync TOC and DIGEST at some point but it is not critical and probaly best addressed in another ticket. Shall I open one?","created":"2015-11-13T06:36:00.722+0000"},{"body":"Yes please open new low priority ticket. It can't be a ton of work to sync the TOC and digest and it would reduce the window where they are inconsistent.\n\nA bigger issue brought up by this is that we don't do power failure testing in a loop so we can't actually verify the correctness of Cassandra under those conditions. Can you create a major priority ticket for that?\n\nI am thinking maybe we just want to have a mock filesystem that defers writes so we can simulate the kernel not flushing after the process exits.","created":"2015-11-13T14:42:17.681+0000"},{"body":"[2.1 commit|https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=7e056fa27047a868660ff796734dcbd485e1b29a]\n[2.2 commit|https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=73a730f926d25a7d4f693507937b8565b701259c]\n\nMerged up.","created":"2015-11-13T14:59:59.604+0000"},{"body":"Created CASSANDRA-10709 and CASSANDRA-10710.","created":"2015-11-16T09:57:21.743+0000"}],"conversations":[{"body":"I was seeing SSTable corruption due to a CompressionInfo.db file of size 0, this happened multiple times in our testing with hard node reboots. After some investigation it seems like these file is not being fsynced, and that can potentially lead to data corruption. I am working with version 2.1.9.\n\nI checked for fsync calls using strace, and found them happening for all but the following components: CompressionInfo, TOC.txt and digest.sha1. All of these but the CompressionInfo seem tolerable. Also a quick look through the code did not reveal any fsync calls. Moreover, I suspect the commit 4e95953f29d89a441dfe06d3f0393ed7dd8586df (https://github.com/apache/cassandra/commit/4e95953f29d89a441dfe06d3f0393ed7dd8586df#diff-b7e48a1398e39a936c11d0397d5d1966R344) has caused the regression, which removed the line\n{noformat}\n getChannel().force(true);\n{noformat}\nfrom CompressionMetadata.Writer.close.\n\nFollowing is the trace I saw in system.log:\n{noformat}\nINFO [SSTableBatchOpen:1] 2015-09-29 19:24:39,170 SSTableReader.java:478 - Opening /var/lib/cassandra/data/system/compactions_in_progress-55080ab05d9c388690a4acb25fe1f77b/system-compactions_in_progress-ka-13368 (79 bytes)\nERROR [SSTableBatchOpen:1] 2015-09-29 19:24:39,177 FileUtils.java:447 - Exiting forcefully due to file system exception on startup, disk failure policy \"stop\"\norg.apache.cassandra.io.sstable.CorruptSSTableException: java.io.EOFException\n at org.apache.cassandra.io.compress.CompressionMetadata.(CompressionMetadata.java:131) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.compress.CompressionMetadata.create(CompressionMetadata.java:85) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.util.CompressedSegmentedFile$Builder.metadata(CompressedSegmentedFile.java:79) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.util.CompressedPoolingSegmentedFile$Builder.complete(CompressedPoolingSegmentedFile.java:72) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.util.SegmentedFile$Builder.complete(SegmentedFile.java:168) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.sstable.SSTableReader.load(SSTableReader.java:752) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.sstable.SSTableReader.load(SSTableReader.java:703) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.sstable.SSTableReader.open(SSTableReader.java:491) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.sstable.SSTableReader.open(SSTableReader.java:387) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.sstable.SSTableReader$4.run(SSTableReader.java:534) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_80]\n at java.util.concurrent.FutureTask.run(FutureTask.java:262) [na:1.7.0_80]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) [na:1.7.0_80]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_80]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_80]\nCaused by: java.io.EOFException: null\n at java.io.DataInputStream.readUnsignedShort(DataInputStream.java:340) ~[na:1.7.0_80]\n at java.io.DataInputStream.readUTF(DataInputStream.java:589) ~[na:1.7.0_80]\n at java.io.DataInputStream.readUTF(DataInputStream.java:564) ~[na:1.7.0_80]\n at org.apache.cassandra.io.compress.CompressionMetadata.(CompressionMetadata.java:106) ~[apache-cassandra-2.1.9.jar:2.1.9]\n ... 14 common frames omitted\n{noformat}\n\nFollowing is the result of ls on the data directory of a corrupted SSTable after the hard reboot:\n{noformat}\n$ ls -l /var/lib/cassandra/data/system/sstable_activity-5a1ff267ace03f128563cfae6103c65e/\ntotal 60\n-rw-r--r-- 1 cassandra cassandra 0 Oct 15 09:31 system-sstable_activity-ka-1-CompressionInfo.db\n-rw-r--r-- 1 cassandra cassandra 9740 Oct 15 09:31 system-sstable_activity-ka-1-Data.db\n-rw-r--r-- 1 cassandra cassandra 0 Oct 15 09:31 system-sstable_activity-ka-1-Digest.sha1\n-rw-r--r-- 1 cassandra cassandra 880 Oct 15 09:31 system-sstable_activity-ka-1-Filter.db\n-rw-r--r-- 1 cassandra cassandra 34000 Oct 15 09:31 system-sstable_activity-ka-1-Index.db\n-rw-r--r-- 1 cassandra cassandra 7338 Oct 15 09:31 system-sstable_activity-ka-1-Statistics.db\n-rw-r--r-- 1 cassandra cassandra 0 Oct 15 09:31 system-sstable_activity-ka-1-TOC.txt\n{noformat}","from":"reporter","subject":"CompressionInfo not being fsynced on close"},{"body":"[~benedict], [~sharvanath] analysis is correct: since CASSANDRA-6916 we no longer fsync compression metadata after writing it. I've attached a small [patch|https://github.com/stef1927/cassandra/commits/10534-2.1] that should fix this, can you take a look?\n\nIf the patch is fine I will run CI on the 2.1+ branches. ","from":"developer"},{"body":"Hi [~sharvanath]: thanks for taking the time to strace this and find our (my) mistake. \n\n[~stefania]: thanks for providing a patch. -LGTM. I'll commit once we have clean CI results.- Why did you put the {{sync()}} call in the finally block before the close, instead of in the try block? At the very least we should ensure the close is called after, but it seems easiest to place it in the try block.","from":"developer"},{"body":"[~benedict] [~Stefania] thanks for taking quick action on it.","from":"developer"},{"body":"[~Benedict]: I merely wanted to fsync in case of partial write but if it looks too unusual we can have it in the try block. I amended the commit and force pushed. The 2.2 patch is a rewrite because the code is too divergent. I believe the FOS close is idempotent so we are OK if we close it twice but please double check. The 2.2 patch then merges without conflicts into 3.0.\n\nCI should eventually appear here:\n\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-2.1-dtest\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-2.1-testall\n\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-2.2-dtest\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-2.2-testall\n\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-3.0-dtest\nhttp://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-10534-3.0-testall\n","from":"developer"},{"body":"In all cases we will need to call flush before calling sync, since we have a buffered writer. {{close}} is idempotent, so that should not be a problem.","from":"developer"},{"body":"I've added the call to {{flush}} in a separate commit. \n\nI tried to abort the CI jobs and restart new ones but it seems the abort only removed the jobs in the queue so it's the old jobs (without flush) that are running at the moment. If they are aborted later on I will restart them or if I am offline you can restart yourself with cassci {{!build stef1927-10534-3.0-dtest}} etc.","from":"developer"},{"body":"I've rebased and restarted all 3 jobs.","from":"developer"},{"body":"This bug also happened in 2.1.10.","from":"developer"},{"body":"[~benedict], I've rebased and re-run CI on 2.1 and 2.2. It looks OK to me (a few failing tests but inline with the main branches). \n\nLet me know if there is anything else to be done before committing.","from":"developer"},{"body":"I don't see why we would ever not sync the other files? Are they really not necessary for the sstable to be readable? If they are required then they need to be synced as well otherwise we are going to take actions based on the sstable being durable/readable when it isn't really readable.","from":"developer"},{"body":"_TOC.txt_ is only used by standalone tools, we do a listing of the folder in the CFS constructor. I could not find where the digest file is read, at least not on the 2.1 code. Other components I have not checked yet. For sure index and data files are sync-ed.\n\nWe should probably sync all sstable components that we write but should we fix this regression and open a new ticket for better visibility?","from":"developer"},{"body":"Ok. +1 LGTM.\n\nIt looks like the other components all use SequentialWriter which if you use the transaction stuff does the right thing without any extra work. Except we don't use the transaction proxy for the other metadata components instead we call finish manually which goes through the transaction proxy.\n\nI can't quite tell since it looks like it can throw and cause other code not to execute.","from":"developer"},{"body":"Thank you for the review and for checking the remaining components. I agree with your analysis. \n\nThese are the components defined in {{SSTableWriter.components()}}:\n\n|| Component || Notes |\n| Component.DATA | Sync-ed, SequentialWriter | \n| Component.PRIMARY_INDEX | Sync-ed, SequentialWriter |\n| Component.STATS | Sync-ed, SequentialWriter |\n| Component.SUMMARY | {{SSTableReader.saveSummary()}}, called in finish, not sync-ed but we write a magic number at the end and we regenerate the summary when loading it if we don't find this magic number, |\n| Component.TOC | Not sync-ed but it's read only by standalone tools | \n| Component.DIGEST | Written by {{DataIntegrityMetadata.ChecksumWriter}}, not sync-ed and not used but intended for users so they can validate uncompressed data files via sha1sum. In 2.2 this becomes the adler32 checksum that can be verified with nodetool verify or the standalone verifier.|\n| Component.FILTER | Written and sync-ed manually in {{IndexWrit er.close()}} |\n| Component.COMPRESSION_INFO | To be sync-ed by this patch in {{ompressionMetadata.Writer.close()}} |\n| Component.CRC | Sync-ed, SequentialWriter |\n\nbq. I can't quite tell since it looks like it can throw and cause other code not to execute.\n\nYes but this was improved in 2.2 with {{LifecycleTransaction}} and in 3.0 even further with the removal of unfinished left overs via {{LogTransaction}}.\n\nSo, IMO, we could have problems with standalone tools and we should probably sync TOC and DIGEST at some point but it is not critical and probaly best addressed in another ticket. Shall I open one?","from":"developer"},{"body":"Yes please open new low priority ticket. It can't be a ton of work to sync the TOC and digest and it would reduce the window where they are inconsistent.\n\nA bigger issue brought up by this is that we don't do power failure testing in a loop so we can't actually verify the correctness of Cassandra under those conditions. Can you create a major priority ticket for that?\n\nI am thinking maybe we just want to have a mock filesystem that defers writes so we can simulate the kernel not flushing after the process exits.","from":"developer"},{"body":"[2.1 commit|https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=7e056fa27047a868660ff796734dcbd485e1b29a]\n[2.2 commit|https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=73a730f926d25a7d4f693507937b8565b701259c]\n\nMerged up.","from":"developer"},{"body":"Created CASSANDRA-10709 and CASSANDRA-10710.","from":"developer"}],"created":"2015-10-15T14:28:07.000+0000","description":"I was seeing SSTable corruption due to a CompressionInfo.db file of size 0, this happened multiple times in our testing with hard node reboots. After some investigation it seems like these file is not being fsynced, and that can potentially lead to data corruption. I am working with version 2.1.9.\n\nI checked for fsync calls using strace, and found them happening for all but the following components: CompressionInfo, TOC.txt and digest.sha1. All of these but the CompressionInfo seem tolerable. Also a quick look through the code did not reveal any fsync calls. Moreover, I suspect the commit 4e95953f29d89a441dfe06d3f0393ed7dd8586df (https://github.com/apache/cassandra/commit/4e95953f29d89a441dfe06d3f0393ed7dd8586df#diff-b7e48a1398e39a936c11d0397d5d1966R344) has caused the regression, which removed the line\n{noformat}\n getChannel().force(true);\n{noformat}\nfrom CompressionMetadata.Writer.close.\n\nFollowing is the trace I saw in system.log:\n{noformat}\nINFO [SSTableBatchOpen:1] 2015-09-29 19:24:39,170 SSTableReader.java:478 - Opening /var/lib/cassandra/data/system/compactions_in_progress-55080ab05d9c388690a4acb25fe1f77b/system-compactions_in_progress-ka-13368 (79 bytes)\nERROR [SSTableBatchOpen:1] 2015-09-29 19:24:39,177 FileUtils.java:447 - Exiting forcefully due to file system exception on startup, disk failure policy \"stop\"\norg.apache.cassandra.io.sstable.CorruptSSTableException: java.io.EOFException\n at org.apache.cassandra.io.compress.CompressionMetadata.(CompressionMetadata.java:131) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.compress.CompressionMetadata.create(CompressionMetadata.java:85) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.util.CompressedSegmentedFile$Builder.metadata(CompressedSegmentedFile.java:79) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.util.CompressedPoolingSegmentedFile$Builder.complete(CompressedPoolingSegmentedFile.java:72) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.util.SegmentedFile$Builder.complete(SegmentedFile.java:168) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.sstable.SSTableReader.load(SSTableReader.java:752) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.sstable.SSTableReader.load(SSTableReader.java:703) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.sstable.SSTableReader.open(SSTableReader.java:491) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.sstable.SSTableReader.open(SSTableReader.java:387) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at org.apache.cassandra.io.sstable.SSTableReader$4.run(SSTableReader.java:534) ~[apache-cassandra-2.1.9.jar:2.1.9]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_80]\n at java.util.concurrent.FutureTask.run(FutureTask.java:262) [na:1.7.0_80]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) [na:1.7.0_80]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_80]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_80]\nCaused by: java.io.EOFException: null\n at java.io.DataInputStream.readUnsignedShort(DataInputStream.java:340) ~[na:1.7.0_80]\n at java.io.DataInputStream.readUTF(DataInputStream.java:589) ~[na:1.7.0_80]\n at java.io.DataInputStream.readUTF(DataInputStream.java:564) ~[na:1.7.0_80]\n at org.apache.cassandra.io.compress.CompressionMetadata.(CompressionMetadata.java:106) ~[apache-cassandra-2.1.9.jar:2.1.9]\n ... 14 common frames omitted\n{noformat}\n\nFollowing is the result of ls on the data directory of a corrupted SSTable after the hard reboot:\n{noformat}\n$ ls -l /var/lib/cassandra/data/system/sstable_activity-5a1ff267ace03f128563cfae6103c65e/\ntotal 60\n-rw-r--r-- 1 cassandra cassandra 0 Oct 15 09:31 system-sstable_activity-ka-1-CompressionInfo.db\n-rw-r--r-- 1 cassandra cassandra 9740 Oct 15 09:31 system-sstable_activity-ka-1-Data.db\n-rw-r--r-- 1 cassandra cassandra 0 Oct 15 09:31 system-sstable_activity-ka-1-Digest.sha1\n-rw-r--r-- 1 cassandra cassandra 880 Oct 15 09:31 system-sstable_activity-ka-1-Filter.db\n-rw-r--r-- 1 cassandra cassandra 34000 Oct 15 09:31 system-sstable_activity-ka-1-Index.db\n-rw-r--r-- 1 cassandra cassandra 7338 Oct 15 09:31 system-sstable_activity-ka-1-Statistics.db\n-rw-r--r-- 1 cassandra cassandra 0 Oct 15 09:31 system-sstable_activity-ka-1-TOC.txt\n{noformat}","issue_id":"12905207","key":"CASSANDRA-10534","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-11-13T14:59:59.000+0000","role":"fixed_distractor","summary":"CompressionInfo not being fsynced on close"} {"case_id":"12907489","cluster":"DISTRACTOR-CASSANDRA-10583","comments":[{"body":"This seems to be related to bulk loading. \n\nTo reproduce:\n\n1. Clone https://github.com/depend/issues/tree/master/CASSANDRA-10583. Build and run it, this application will generate an sstable with 10 rows.\n2. Load it into C* with sstableloader.\n3. \n{noformat} \ncqlsh:timeseries_test> select * from double_daily;\n\n tag | group | timestamp | value\n------+-------+--------------------------+-------\n TEST | 1 | 2002-05-01 04:00:00+0000 | 0\n TEST | 1 | 2002-05-02 04:00:00+0000 | 1\n TEST | 1 | 2002-05-03 04:00:00+0000 | 2\n TEST | 1 | 2002-05-04 04:00:00+0000 | 3\n TEST | 1 | 2002-05-05 04:00:00+0000 | 4\n TEST | 1 | 2002-05-06 04:00:00+0000 | 5\n TEST | 1 | 2002-05-07 04:00:00+0000 | 6\n TEST | 1 | 2002-05-08 04:00:00+0000 | 7\n TEST | 1 | 2002-05-09 04:00:00+0000 | 8\n TEST | 1 | 2002-05-10 04:00:00+0000 | 9\n\n(10 rows)\n{noformat} \n\n4. \n{noformat} \ncqlsh:timeseries_test> select * from double_daily where tag='TEST' and group = 1 and timestamp > '2002-05-01 00:00:00-0400';\n\n tag | group | timestamp | value\n------+-------+--------------------------+-------\n TEST | 1 | 2002-05-01 04:00:00+0000 | 0\n TEST | 1 | 2002-05-02 04:00:00+0000 | 1\n TEST | 1 | 2002-05-03 04:00:00+0000 | 2\n TEST | 1 | 2002-05-04 04:00:00+0000 | 3\n TEST | 1 | 2002-05-05 04:00:00+0000 | 4\n TEST | 1 | 2002-05-06 04:00:00+0000 | 5\n TEST | 1 | 2002-05-07 04:00:00+0000 | 6\n TEST | 1 | 2002-05-08 04:00:00+0000 | 7\n TEST | 1 | 2002-05-09 04:00:00+0000 | 8\n TEST | 1 | 2002-05-10 04:00:00+0000 | 9\n\n(10 rows)\n{noformat} \n\n5. \n{noformat} \ncqlsh:timeseries_test> select * from double_daily where tag='TEST' and group = 1 and timestamp > '2002-05-02 00:00:00-0400';\n\n tag | group | timestamp | value\n-----+-------+-----------+-------\n\n(0 rows)\n{noformat} \n\nI wasn't able to find that \"equal\" condition which returns everything. But query #5 still shows nothing is later than 2002/5/2 which is not true.","created":"2015-10-23T18:46:13.742+0000"},{"body":"After reporting this error, I converted all timestamps in our application to bigint and thought the workaround worked. Then I upgraded C* to 2.2.4. But once again we are seeing the same error with bigint - query returns data that clearly should not be returned.\n\nIt could be a bug in old CQLSSTableWriter. I will use CQLSSTableWriter in 2.2.4 and regenerate sstables then bulk load to see if this still happens.","created":"2015-12-31T22:41:59.571+0000"},{"body":"[~depend] Thanks for the report.\n\nI used your repository and generated SSTable, create 1 node cassandra cluster for v2.1.9 and the latest cassandra-2.1, and loaded generated SSTable using sstableloader.\n\nUnfortunately, I could not reproduce your issue in both clusters.\n\n{code}\ncqlsh:timeseries_test> select * from double_daily where tag='TEST' and group = 1 and timestamp > '2002-05-05 00:00:00-0400';\n tag | group | timestamp | value\n------+-------+--------------------------+-------\n TEST | 1 | 2002-05-05 05:00:00+0000 | 4\n TEST | 1 | 2002-05-06 05:00:00+0000 | 5\n TEST | 1 | 2002-05-07 05:00:00+0000 | 6\n TEST | 1 | 2002-05-08 05:00:00+0000 | 7\n TEST | 1 | 2002-05-09 05:00:00+0000 | 8\n TEST | 1 | 2002-05-10 05:00:00+0000 | 9\n(6 rows)\n{code}\n\nThe result of {{sstable2json}} does not seem problematic.\n{code}\n[\n{\"key\": \"TEST\",\n \"cells\": [[\"1:2002-05-01 00\\\\:00-0500:\",\"\",1455036610854000],\n [\"1:2002-05-01 00\\\\:00-0500:value\",\"0.0\",1455036610854000],\n [\"1:2002-05-02 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-02 00\\\\:00-0500:value\",\"1.0\",1455036610861000],\n [\"1:2002-05-03 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-03 00\\\\:00-0500:value\",\"2.0\",1455036610861000],\n [\"1:2002-05-04 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-04 00\\\\:00-0500:value\",\"3.0\",1455036610861000],\n [\"1:2002-05-05 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-05 00\\\\:00-0500:value\",\"4.0\",1455036610861000],\n [\"1:2002-05-06 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-06 00\\\\:00-0500:value\",\"5.0\",1455036610861000],\n [\"1:2002-05-07 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-07 00\\\\:00-0500:value\",\"6.0\",1455036610861000],\n [\"1:2002-05-08 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-08 00\\\\:00-0500:value\",\"7.0\",1455036610861000],\n [\"1:2002-05-09 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-09 00\\\\:00-0500:value\",\"8.0\",1455036610861000],\n [\"1:2002-05-10 00\\\\:00-0500:\",\"\",1455036610862000],\n [\"1:2002-05-10 00\\\\:00-0500:value\",\"9.0\",1455036610862000]]}\n]\n{code}\n\nIs there anything special to your environment?","created":"2016-02-09T17:40:26.550+0000"},{"body":"Yuki,\n\nThanks for looking into this. I can still reproduce the problem. My sstable2json seems fine.\n\n{noformat}\n[\n{\"key\": \"TEST\",\n \"cells\": [[\"1:2002-05-01 00\\\\:00-0400:\",\"\",1455044687135000],\n [\"1:2002-05-01 00\\\\:00-0400:value\",\"0.0\",1455044687135000],\n [\"1:2002-05-02 00\\\\:00-0400:\",\"\",1455044687143000],\n [\"1:2002-05-02 00\\\\:00-0400:value\",\"1.0\",1455044687143000],\n [\"1:2002-05-03 00\\\\:00-0400:\",\"\",1455044687143000],\n [\"1:2002-05-03 00\\\\:00-0400:value\",\"2.0\",1455044687143000],\n [\"1:2002-05-04 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-04 00\\\\:00-0400:value\",\"3.0\",1455044687144000],\n [\"1:2002-05-05 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-05 00\\\\:00-0400:value\",\"4.0\",1455044687144000],\n [\"1:2002-05-06 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-06 00\\\\:00-0400:value\",\"5.0\",1455044687144000],\n [\"1:2002-05-07 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-07 00\\\\:00-0400:value\",\"6.0\",1455044687144000],\n [\"1:2002-05-08 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-08 00\\\\:00-0400:value\",\"7.0\",1455044687144000],\n [\"1:2002-05-09 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-09 00\\\\:00-0400:value\",\"8.0\",1455044687144000],\n [\"1:2002-05-10 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-10 00\\\\:00-0400:value\",\"9.0\",1455044687144000]]}\n]\n{noformat}","created":"2016-02-09T19:26:50.105+0000"},{"body":"[~yukim]\n\nI uploaded sstables and screenshots.\n\nI tried to create sstables on both Windows 10 and CentOS 7.\n\nI tested 2.2.4 on Windows and Linux. Both have the same problem.","created":"2016-02-09T19:32:53.993+0000"},{"body":"Can you post the resut with {{TRACING ON}}?","created":"2016-02-09T19:40:39.545+0000"},{"body":"uploaded","created":"2016-02-09T19:46:25.856+0000"},{"body":"[~yukim]\n\nIs there any update on this ticket. I could reproduce it using the SSTable attached to the ticket.","created":"2016-05-25T03:38:59.370+0000"},{"body":"I think the problem here is in your CQLSSTableWriter program you didn't specify {{CLUSTERING ORDER BY}}, so the data are inserted in ascending order, but in Cassandra the schema treated it as descending since you created that way, query result is messed up.\n\nCan you update your generation program to use schema that matches in Cassandra?","created":"2016-05-25T15:43:29.650+0000"},{"body":"That's it! After I add {{WITH CLUSTERING ORDER BY (group ASC, timestamp DESC)}} it started to work. I don't know why I missed it. It makes sense to include cluster options when generating sstables. ","created":"2016-05-25T16:21:19.347+0000"},{"body":"Adding {{WITH CLUSTERING ORDER BY ...}} from client side fixed the issue.","created":"2016-05-25T16:23:40.186+0000"}],"conversations":[{"body":"I have this table:\n\n{noformat}\nCREATE TABLE test (\n tag text,\n group int,\n timestamp timestamp,\n value double,\n PRIMARY KEY (tag, group, timestamp)\n) WITH CLUSTERING ORDER BY (group ASC, timestamp DESC)\n{noformat}\n\nFirst I used CQLSSTableWriter to bulk load a bunch of sstables. Then I ran this query:\n\n{noformat}\ncqlsh> select * from test where tag = 'MSFT' and group = 1 and timestamp ='2004-12-15 16:00:00-0500';\n\n tag | group | timestamp | value\n------+-------+--------------------------+-------\n MSFT | 1 | 2004-12-15 21:00:00+0000 | 27.11\n MSFT | 1 | 2004-12-16 21:00:00+0000 | 27.16\n MSFT | 1 | 2004-12-17 21:00:00+0000 | 26.96\n MSFT | 1 | 2004-12-20 21:00:00+0000 | 26.95\n MSFT | 1 | 2004-12-21 21:00:00+0000 | 27.07\n MSFT | 1 | 2004-12-22 21:00:00+0000 | 26.98\n MSFT | 1 | 2004-12-23 21:00:00+0000 | 27.01\n MSFT | 1 | 2004-12-27 21:00:00+0000 | 26.85\n MSFT | 1 | 2004-12-28 21:00:00+0000 | 26.95\n MSFT | 1 | 2004-12-29 21:00:00+0000 | 26.9\n MSFT | 1 | 2004-12-30 21:00:00+0000 | 26.76\n(11 rows)\n{noformat}\n\nThe result is obviously wrong.\n\nIf I run this query:\n\n{noformat}\ncqlsh> select * from test where tag = 'MSFT' and group = 1 and timestamp ='2004-12-16 16:00:00-0500';\n\n tag | group | timestamp | value\n-----+-------+-----------+-------\n\n(0 rows)\n{noformat}\n\nIn DevCenter I tried to create a similar table and insert a few rows but couldn't reproduce this. This may have something to do with the bulk loading process. But still, the fact cqlsh returns data that doesn't match the query is concerning.","from":"reporter","subject":"After bulk loading CQL query on timestamp column returns wrong result"},{"body":"This seems to be related to bulk loading. \n\nTo reproduce:\n\n1. Clone https://github.com/depend/issues/tree/master/CASSANDRA-10583. Build and run it, this application will generate an sstable with 10 rows.\n2. Load it into C* with sstableloader.\n3. \n{noformat} \ncqlsh:timeseries_test> select * from double_daily;\n\n tag | group | timestamp | value\n------+-------+--------------------------+-------\n TEST | 1 | 2002-05-01 04:00:00+0000 | 0\n TEST | 1 | 2002-05-02 04:00:00+0000 | 1\n TEST | 1 | 2002-05-03 04:00:00+0000 | 2\n TEST | 1 | 2002-05-04 04:00:00+0000 | 3\n TEST | 1 | 2002-05-05 04:00:00+0000 | 4\n TEST | 1 | 2002-05-06 04:00:00+0000 | 5\n TEST | 1 | 2002-05-07 04:00:00+0000 | 6\n TEST | 1 | 2002-05-08 04:00:00+0000 | 7\n TEST | 1 | 2002-05-09 04:00:00+0000 | 8\n TEST | 1 | 2002-05-10 04:00:00+0000 | 9\n\n(10 rows)\n{noformat} \n\n4. \n{noformat} \ncqlsh:timeseries_test> select * from double_daily where tag='TEST' and group = 1 and timestamp > '2002-05-01 00:00:00-0400';\n\n tag | group | timestamp | value\n------+-------+--------------------------+-------\n TEST | 1 | 2002-05-01 04:00:00+0000 | 0\n TEST | 1 | 2002-05-02 04:00:00+0000 | 1\n TEST | 1 | 2002-05-03 04:00:00+0000 | 2\n TEST | 1 | 2002-05-04 04:00:00+0000 | 3\n TEST | 1 | 2002-05-05 04:00:00+0000 | 4\n TEST | 1 | 2002-05-06 04:00:00+0000 | 5\n TEST | 1 | 2002-05-07 04:00:00+0000 | 6\n TEST | 1 | 2002-05-08 04:00:00+0000 | 7\n TEST | 1 | 2002-05-09 04:00:00+0000 | 8\n TEST | 1 | 2002-05-10 04:00:00+0000 | 9\n\n(10 rows)\n{noformat} \n\n5. \n{noformat} \ncqlsh:timeseries_test> select * from double_daily where tag='TEST' and group = 1 and timestamp > '2002-05-02 00:00:00-0400';\n\n tag | group | timestamp | value\n-----+-------+-----------+-------\n\n(0 rows)\n{noformat} \n\nI wasn't able to find that \"equal\" condition which returns everything. But query #5 still shows nothing is later than 2002/5/2 which is not true.","from":"developer"},{"body":"After reporting this error, I converted all timestamps in our application to bigint and thought the workaround worked. Then I upgraded C* to 2.2.4. But once again we are seeing the same error with bigint - query returns data that clearly should not be returned.\n\nIt could be a bug in old CQLSSTableWriter. I will use CQLSSTableWriter in 2.2.4 and regenerate sstables then bulk load to see if this still happens.","from":"developer"},{"body":"[~depend] Thanks for the report.\n\nI used your repository and generated SSTable, create 1 node cassandra cluster for v2.1.9 and the latest cassandra-2.1, and loaded generated SSTable using sstableloader.\n\nUnfortunately, I could not reproduce your issue in both clusters.\n\n{code}\ncqlsh:timeseries_test> select * from double_daily where tag='TEST' and group = 1 and timestamp > '2002-05-05 00:00:00-0400';\n tag | group | timestamp | value\n------+-------+--------------------------+-------\n TEST | 1 | 2002-05-05 05:00:00+0000 | 4\n TEST | 1 | 2002-05-06 05:00:00+0000 | 5\n TEST | 1 | 2002-05-07 05:00:00+0000 | 6\n TEST | 1 | 2002-05-08 05:00:00+0000 | 7\n TEST | 1 | 2002-05-09 05:00:00+0000 | 8\n TEST | 1 | 2002-05-10 05:00:00+0000 | 9\n(6 rows)\n{code}\n\nThe result of {{sstable2json}} does not seem problematic.\n{code}\n[\n{\"key\": \"TEST\",\n \"cells\": [[\"1:2002-05-01 00\\\\:00-0500:\",\"\",1455036610854000],\n [\"1:2002-05-01 00\\\\:00-0500:value\",\"0.0\",1455036610854000],\n [\"1:2002-05-02 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-02 00\\\\:00-0500:value\",\"1.0\",1455036610861000],\n [\"1:2002-05-03 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-03 00\\\\:00-0500:value\",\"2.0\",1455036610861000],\n [\"1:2002-05-04 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-04 00\\\\:00-0500:value\",\"3.0\",1455036610861000],\n [\"1:2002-05-05 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-05 00\\\\:00-0500:value\",\"4.0\",1455036610861000],\n [\"1:2002-05-06 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-06 00\\\\:00-0500:value\",\"5.0\",1455036610861000],\n [\"1:2002-05-07 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-07 00\\\\:00-0500:value\",\"6.0\",1455036610861000],\n [\"1:2002-05-08 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-08 00\\\\:00-0500:value\",\"7.0\",1455036610861000],\n [\"1:2002-05-09 00\\\\:00-0500:\",\"\",1455036610861000],\n [\"1:2002-05-09 00\\\\:00-0500:value\",\"8.0\",1455036610861000],\n [\"1:2002-05-10 00\\\\:00-0500:\",\"\",1455036610862000],\n [\"1:2002-05-10 00\\\\:00-0500:value\",\"9.0\",1455036610862000]]}\n]\n{code}\n\nIs there anything special to your environment?","from":"developer"},{"body":"Yuki,\n\nThanks for looking into this. I can still reproduce the problem. My sstable2json seems fine.\n\n{noformat}\n[\n{\"key\": \"TEST\",\n \"cells\": [[\"1:2002-05-01 00\\\\:00-0400:\",\"\",1455044687135000],\n [\"1:2002-05-01 00\\\\:00-0400:value\",\"0.0\",1455044687135000],\n [\"1:2002-05-02 00\\\\:00-0400:\",\"\",1455044687143000],\n [\"1:2002-05-02 00\\\\:00-0400:value\",\"1.0\",1455044687143000],\n [\"1:2002-05-03 00\\\\:00-0400:\",\"\",1455044687143000],\n [\"1:2002-05-03 00\\\\:00-0400:value\",\"2.0\",1455044687143000],\n [\"1:2002-05-04 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-04 00\\\\:00-0400:value\",\"3.0\",1455044687144000],\n [\"1:2002-05-05 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-05 00\\\\:00-0400:value\",\"4.0\",1455044687144000],\n [\"1:2002-05-06 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-06 00\\\\:00-0400:value\",\"5.0\",1455044687144000],\n [\"1:2002-05-07 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-07 00\\\\:00-0400:value\",\"6.0\",1455044687144000],\n [\"1:2002-05-08 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-08 00\\\\:00-0400:value\",\"7.0\",1455044687144000],\n [\"1:2002-05-09 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-09 00\\\\:00-0400:value\",\"8.0\",1455044687144000],\n [\"1:2002-05-10 00\\\\:00-0400:\",\"\",1455044687144000],\n [\"1:2002-05-10 00\\\\:00-0400:value\",\"9.0\",1455044687144000]]}\n]\n{noformat}","from":"developer"},{"body":"[~yukim]\n\nI uploaded sstables and screenshots.\n\nI tried to create sstables on both Windows 10 and CentOS 7.\n\nI tested 2.2.4 on Windows and Linux. Both have the same problem.","from":"developer"},{"body":"Can you post the resut with {{TRACING ON}}?","from":"developer"},{"body":"uploaded","from":"developer"},{"body":"[~yukim]\n\nIs there any update on this ticket. I could reproduce it using the SSTable attached to the ticket.","from":"developer"},{"body":"I think the problem here is in your CQLSSTableWriter program you didn't specify {{CLUSTERING ORDER BY}}, so the data are inserted in ascending order, but in Cassandra the schema treated it as descending since you created that way, query result is messed up.\n\nCan you update your generation program to use schema that matches in Cassandra?","from":"developer"},{"body":"That's it! After I add {{WITH CLUSTERING ORDER BY (group ASC, timestamp DESC)}} it started to work. I don't know why I missed it. It makes sense to include cluster options when generating sstables. ","from":"developer"},{"body":"Adding {{WITH CLUSTERING ORDER BY ...}} from client side fixed the issue.","from":"developer"}],"created":"2015-10-23T17:47:55.000+0000","description":"I have this table:\n\n{noformat}\nCREATE TABLE test (\n tag text,\n group int,\n timestamp timestamp,\n value double,\n PRIMARY KEY (tag, group, timestamp)\n) WITH CLUSTERING ORDER BY (group ASC, timestamp DESC)\n{noformat}\n\nFirst I used CQLSSTableWriter to bulk load a bunch of sstables. Then I ran this query:\n\n{noformat}\ncqlsh> select * from test where tag = 'MSFT' and group = 1 and timestamp ='2004-12-15 16:00:00-0500';\n\n tag | group | timestamp | value\n------+-------+--------------------------+-------\n MSFT | 1 | 2004-12-15 21:00:00+0000 | 27.11\n MSFT | 1 | 2004-12-16 21:00:00+0000 | 27.16\n MSFT | 1 | 2004-12-17 21:00:00+0000 | 26.96\n MSFT | 1 | 2004-12-20 21:00:00+0000 | 26.95\n MSFT | 1 | 2004-12-21 21:00:00+0000 | 27.07\n MSFT | 1 | 2004-12-22 21:00:00+0000 | 26.98\n MSFT | 1 | 2004-12-23 21:00:00+0000 | 27.01\n MSFT | 1 | 2004-12-27 21:00:00+0000 | 26.85\n MSFT | 1 | 2004-12-28 21:00:00+0000 | 26.95\n MSFT | 1 | 2004-12-29 21:00:00+0000 | 26.9\n MSFT | 1 | 2004-12-30 21:00:00+0000 | 26.76\n(11 rows)\n{noformat}\n\nThe result is obviously wrong.\n\nIf I run this query:\n\n{noformat}\ncqlsh> select * from test where tag = 'MSFT' and group = 1 and timestamp ='2004-12-16 16:00:00-0500';\n\n tag | group | timestamp | value\n-----+-------+-----------+-------\n\n(0 rows)\n{noformat}\n\nIn DevCenter I tried to create a similar table and insert a few rows but couldn't reproduce this. This may have something to do with the bulk loading process. But still, the fact cqlsh returns data that doesn't match the query is concerning.","issue_id":"12907489","key":"CASSANDRA-10583","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-05-25T16:23:40.000+0000","role":"fixed_distractor","summary":"After bulk loading CQL query on timestamp column returns wrong result"} {"case_id":"12928185","cluster":"DISTRACTOR-CASSANDRA-10979","comments":[{"body":"{code:title=src/java/org/apache/cassandra/db/compaction/LeveledManifest.java}\ndiff --git a/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java b/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java\nindex 1e67a9e..ed45a63 100644\n--- a/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java\n+++ b/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java\n@@ -602,6 +602,8 @@ public class LeveledManifest\n candidates = Sets.union(candidates, l1overlapping);\n }\n if (candidates.size() < 2)\n+ if (!DatabaseDescriptor.getDisableSTCSInL0() && remaining.size() > MAX_COMPACTING_L0)\n+ return getSSTablesForSTCS(remaining)\n return Collections.emptyList();\n else\n return candidates;\n\n{code}","created":"2016-01-07T19:24:38.002+0000"},{"body":"Debug logging using the current (unpatched) code: https://gist.github.com/autocracy/346afa253175af475770","created":"2016-01-07T21:24:26.556+0000"},{"body":"[~krummas] I mentioned this briefly this morning over chat but now we have more information. Specifically we observe a large L0 to L1 compaction that blocks additional L0 STCS compactions, probably because the remaining L0 sstables overlap compactingL0. \n\nYou can see in the log above that {quote}L0 is too far behind, performing size-tiering there first{quote} does not appear. What do you think?","created":"2016-01-07T22:32:13.289+0000"},{"body":"Because we will drop out of the method early due to the overlapping L1 sstables at [L598|https://github.com/apache/cassandra/blob/cassandra-2.1/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java#L598], this won't return the proper set of sstables. We also need to make sure that we are returning a {{CompactionCandidate}} which will compact to L0, otherwise we will create overlapping sstables in L1.\n\nI've attached a new patch which factors out the selection of the STCS compaction, and if we don't select anything in L0 to compact, we will check to see if we should perform an STCS compaction.","created":"2016-01-08T16:02:00.195+0000"},{"body":"This LGTM, could you push a branch so we get the test runs in?\n\nAnd, should we really target 2.1 for this?","created":"2016-01-12T12:31:45.802+0000"},{"body":"I pushed them; forgot to link the tests here.\n\n||trunk||\n|[branch|https://github.com/carlyeks/cassandra/tree/ticket/10979/trunk]|\n|[utest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-trunk-testall]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-trunk-dtest]|\n\nGood point on not targeting 2.1; since this is an improvement and not a bugfix, we should only target trunk for this fix.","created":"2016-01-12T16:57:45.415+0000"},{"body":"The problem I've got and why I considered it a bug is that the behavior as it is right now doesn't appear to be as it was intended, and particularly negative for me is that it destroys performance of my cluster by an order of magnitude with LCS tables during a repair as potentially thousands of tables build up in L0.","created":"2016-01-12T17:31:16.717+0000"},{"body":"Agree with Jeff here. This seems like a bug in the intended behavior that we should at least fix in 2.2+, since it's not a critical bug fix to go in 2.1.","created":"2016-01-12T22:33:57.720+0000"},{"body":"[~autocracy] have you confirmed that the patch fixes your issue?\n\nI would assume that any slowdown after incremental repair is due to https://issues.apache.org/jira/browse/CASSANDRA-10831","created":"2016-01-13T09:40:10.409+0000"},{"body":"I'm fine with this going in 2.2 as it's not an intrusive fix, but I also think this is more of an improvement than a bug fix.\n\nWhile this will reduce the number of sstables on a read when L0 is backed up, the problem remains that L0 cannot keep up with the amount of incoming data. Once we finish the L0->L1 compaction, we will be running the L0 STCS since we will have a level that is oversized which will also cause us to check L0 again.","created":"2016-01-13T16:28:59.084+0000"},{"body":"[~sebastian.estevez@datastax.com]: Can you cut a build of DSE 4.8.4 with this patch added? I'll try my best to test within a week.","created":"2016-01-14T15:48:36.394+0000"},{"body":"Applied the patch to a node that had just come up from streaming and was 700+ tables in L0. After restart, observed that as L0 was shifted into L1, L0 continued to compact new tables to prevent the extreme growth of tables.\n\nThank you. Again still requesting a merge into 2.1 since streaming and repair with LCS are practically broken for us without this patch.","created":"2016-01-20T19:01:25.619+0000"},{"body":"I've rebased and pushed up new 2.2-trunk branches.\n\n||2.2||3.0||3.3||trunk||\n|[branch|https://github.com/carlyeks/cassandra/tree/ticket/10979/2.2]|[branch|https://github.com/carlyeks/cassandra/tree/ticket/10979/3.0]|[branch|https://github.com/carlyeks/cassandra/tree/ticket/10979/3.3]|[branch|https://github.com/carlyeks/cassandra/tree/ticket/10979/trunk]|\n|[utest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-2.2-testall/]|[utest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-3.0-testall/]|[utest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-3.3-testall/]|[utest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-trunk-testall/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-2.2-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-3.3-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-trunk-dtest/]|","created":"2016-01-20T22:50:10.132+0000"},{"body":"committed to 2.2+ - [~autocracy], the 2.1 patch still applies if you can't upgrade to 2.2","created":"2016-01-22T06:36:14.533+0000"}],"conversations":[{"body":"Reading code from https://github.com/apache/cassandra/blob/cassandra-2.1/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java and comparing with behavior shown in https://gist.github.com/autocracy/c95aca6b00e42215daaf, the following happens:\n\nScore for L1,L2,and L3 is all < 1 (paste shows 20/10 and 200/100, due to incremental repair).\n\nRelevant code from here is\n\n if (Sets.intersection(l1overlapping, compacting).size() > 0)\n return Collections.emptyList();\n\nSince there will be overlap between what is compacting and L1 (in my case, pushing over 1,000 tables in to L1 from L0 SCTS), I get a pile up of 1,000 smaller tables in L0 while awaiting the transition from L0 to L1 and destroy my performance.\n\nRequested outcome is to continue to perform SCTS on non-compacting L0 tables.","from":"reporter","subject":"LCS doesn't do L0 STC on new tables while an L0->L1 compaction is in progress"},{"body":"{code:title=src/java/org/apache/cassandra/db/compaction/LeveledManifest.java}\ndiff --git a/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java b/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java\nindex 1e67a9e..ed45a63 100644\n--- a/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java\n+++ b/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java\n@@ -602,6 +602,8 @@ public class LeveledManifest\n candidates = Sets.union(candidates, l1overlapping);\n }\n if (candidates.size() < 2)\n+ if (!DatabaseDescriptor.getDisableSTCSInL0() && remaining.size() > MAX_COMPACTING_L0)\n+ return getSSTablesForSTCS(remaining)\n return Collections.emptyList();\n else\n return candidates;\n\n{code}","from":"developer"},{"body":"Debug logging using the current (unpatched) code: https://gist.github.com/autocracy/346afa253175af475770","from":"developer"},{"body":"[~krummas] I mentioned this briefly this morning over chat but now we have more information. Specifically we observe a large L0 to L1 compaction that blocks additional L0 STCS compactions, probably because the remaining L0 sstables overlap compactingL0. \n\nYou can see in the log above that {quote}L0 is too far behind, performing size-tiering there first{quote} does not appear. What do you think?","from":"developer"},{"body":"Because we will drop out of the method early due to the overlapping L1 sstables at [L598|https://github.com/apache/cassandra/blob/cassandra-2.1/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java#L598], this won't return the proper set of sstables. We also need to make sure that we are returning a {{CompactionCandidate}} which will compact to L0, otherwise we will create overlapping sstables in L1.\n\nI've attached a new patch which factors out the selection of the STCS compaction, and if we don't select anything in L0 to compact, we will check to see if we should perform an STCS compaction.","from":"developer"},{"body":"This LGTM, could you push a branch so we get the test runs in?\n\nAnd, should we really target 2.1 for this?","from":"developer"},{"body":"I pushed them; forgot to link the tests here.\n\n||trunk||\n|[branch|https://github.com/carlyeks/cassandra/tree/ticket/10979/trunk]|\n|[utest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-trunk-testall]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-trunk-dtest]|\n\nGood point on not targeting 2.1; since this is an improvement and not a bugfix, we should only target trunk for this fix.","from":"developer"},{"body":"The problem I've got and why I considered it a bug is that the behavior as it is right now doesn't appear to be as it was intended, and particularly negative for me is that it destroys performance of my cluster by an order of magnitude with LCS tables during a repair as potentially thousands of tables build up in L0.","from":"developer"},{"body":"Agree with Jeff here. This seems like a bug in the intended behavior that we should at least fix in 2.2+, since it's not a critical bug fix to go in 2.1.","from":"developer"},{"body":"[~autocracy] have you confirmed that the patch fixes your issue?\n\nI would assume that any slowdown after incremental repair is due to https://issues.apache.org/jira/browse/CASSANDRA-10831","from":"developer"},{"body":"I'm fine with this going in 2.2 as it's not an intrusive fix, but I also think this is more of an improvement than a bug fix.\n\nWhile this will reduce the number of sstables on a read when L0 is backed up, the problem remains that L0 cannot keep up with the amount of incoming data. Once we finish the L0->L1 compaction, we will be running the L0 STCS since we will have a level that is oversized which will also cause us to check L0 again.","from":"developer"},{"body":"[~sebastian.estevez@datastax.com]: Can you cut a build of DSE 4.8.4 with this patch added? I'll try my best to test within a week.","from":"developer"},{"body":"Applied the patch to a node that had just come up from streaming and was 700+ tables in L0. After restart, observed that as L0 was shifted into L1, L0 continued to compact new tables to prevent the extreme growth of tables.\n\nThank you. Again still requesting a merge into 2.1 since streaming and repair with LCS are practically broken for us without this patch.","from":"developer"},{"body":"I've rebased and pushed up new 2.2-trunk branches.\n\n||2.2||3.0||3.3||trunk||\n|[branch|https://github.com/carlyeks/cassandra/tree/ticket/10979/2.2]|[branch|https://github.com/carlyeks/cassandra/tree/ticket/10979/3.0]|[branch|https://github.com/carlyeks/cassandra/tree/ticket/10979/3.3]|[branch|https://github.com/carlyeks/cassandra/tree/ticket/10979/trunk]|\n|[utest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-2.2-testall/]|[utest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-3.0-testall/]|[utest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-3.3-testall/]|[utest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-trunk-testall/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-2.2-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-3.3-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/carlyeks/job/carlyeks-ticket-10979-trunk-dtest/]|","from":"developer"},{"body":"committed to 2.2+ - [~autocracy], the 2.1 patch still applies if you can't upgrade to 2.2","from":"developer"}],"created":"2016-01-07T00:29:49.000+0000","description":"Reading code from https://github.com/apache/cassandra/blob/cassandra-2.1/src/java/org/apache/cassandra/db/compaction/LeveledManifest.java and comparing with behavior shown in https://gist.github.com/autocracy/c95aca6b00e42215daaf, the following happens:\n\nScore for L1,L2,and L3 is all < 1 (paste shows 20/10 and 200/100, due to incremental repair).\n\nRelevant code from here is\n\n if (Sets.intersection(l1overlapping, compacting).size() > 0)\n return Collections.emptyList();\n\nSince there will be overlap between what is compacting and L1 (in my case, pushing over 1,000 tables in to L1 from L0 SCTS), I get a pile up of 1,000 smaller tables in L0 while awaiting the transition from L0 to L1 and destroy my performance.\n\nRequested outcome is to continue to perform SCTS on non-compacting L0 tables.","issue_id":"12928185","key":"CASSANDRA-10979","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-01-22T06:36:14.000+0000","role":"fixed_distractor","summary":"LCS doesn't do L0 STC on new tables while an L0->L1 compaction is in progress"} {"case_id":"12928641","cluster":"DISTRACTOR-CASSANDRA-10988","comments":[{"body":"We managed to reproduce the bug using Cassandra 2.2.5 with the following table schema (unset parameters use default values):\n{code:sql}\nCREATE TABLE mytable (\n u text,\n t timeuuid,\n PRIMARY KEY (u, t)\n) WITH COMPACT STORAGE\n AND CLUSTERING ORDER BY (t DESC)\n AND compaction = {'min_threshold': '2', 'class': 'org.apache.cassandra.db.compaction.DateTieredCompactionStrategy', 'base_time_seconds': '1'}\n AND compression = {'sstable_compression': 'org.apache.cassandra.io.compress.LZ4Compressor'}\n AND dclocal_read_repair_chance = 0.0\n AND default_time_to_live = 10\n AND gc_grace_seconds = 0\n AND read_repair_chance = 0.0;\n{code}\nAnd the following query:\n{code:sql}\nSELECT COUNT(1) as cnt\nFROM mytable\nWHERE u = :user AND t > :timestamp\nORDER BY t DESC\nLIMIT 110;\n{code}\nRemoving {{COMPACT STORAGE}} definitely helps, so it is somehow connected.","created":"2016-04-07T17:22:35.786+0000"},{"body":"I've tested it with {{2.2}} branch and was able to reproduce it with unit tests (updated summary). {{3.x}} is unaffected.","created":"2016-04-13T16:12:11.747+0000"},{"body":"Created a patch that works, although I'm not entirely satisfied with it. \n\nCurrent problem is that in {{SelectStatement::makeExclusiveSliceBound}}, call to {{areRequestedBoundsInclusive}} (which calls {{isInclusive}} internally), which doesn't reverse {{Bound}} even when clustering order is reversed. On the other hand, next call, {{SelectStatement::getClusteringColumnsBounds}}, reverses the bound internally, which results into the empty {{ByteBuffer}} for the {{Bound}} that's present and non-empty one for the {{Bound}} that's unset.\n\nThe patch is making {{PrimaryKeyRestrictionSet::isInclusive}} to reverse {{Bound}} when needed. Problem is a) that would work only for {{SingleColumnRestriction}}, as there's just one column and b) all \"underlying\" {{isInclusive}} methods are not flipping bounds. It's possible to introduce a different method that'd be more transparent (although that one is only used there iirc) in terms of whether or not it flips the bound. \n\nTests are, however, passing\n\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/10988-2.2]|[testall|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-2.2-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-2.2-dtest/]|","created":"2016-04-15T06:20:24.930+0000"},{"body":"I think that it will be clearer to reverse the {{Bound}} in {{SelectStatement::makeExclusiveSliceBound}} and add a comment to explain it.\n{{SelectStatement::makeExclusiveSliceBound}} will be only called for non-composite slices ({{SingleColumnRestriction.Slice}}).\n\nI noticed that the {{SelectMultiColumnRelationTest::testMixedOrderColumnsX}} methods do not tests {{COMPACT}} tables. It is not directly related to the patch but this problem proves that our testing for {{COMPACT}} tables is not good enough. By consequence, extending it might be a good idea.\n\nIt will be nice if you can also create a patch for {{3.0}} with the unit tests changes. The problem does not affect that version but adding tests is always a good idea to prevent regressions.\n\n","created":"2016-04-20T10:02:32.368+0000"},{"body":"Made the suggested changes (move the logic to {{SelectStatement}}). Also, made the mixed order columns tests also test with {{COMPACT}} tables and added all tests to both {{3.0}} and {{trunk}} (all are passing}}. \n\n|code|utest|dtest|\n|[2.2|https://github.com/ifesdjeen/cassandra/tree/10988-2.2]|[testall|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-2.2-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-2.2-dtest/]|\n|[3.0|https://github.com/ifesdjeen/cassandra/tree/10988-3.0]|[testall|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-3.0-testall/]|\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/10988-trunk]|[testall|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-trunk-testall/]|\n\n(I skipped dtests for {{3.0}} and {{trunk}} since there were no code changes, tests only)\nRan all failing unit tests locally, they're all passing. Dtest shows no failures.","created":"2016-04-21T07:24:26.922+0000"},{"body":"Thanks. +1","created":"2016-04-28T08:56:41.349+0000"},{"body":"Committed into 2.2 at c8d955533b6968368907e5b090a309ac57bf419f and merged into 3.0 and trunk","created":"2016-04-28T09:26:06.562+0000"},{"body":"Thank you!","created":"2016-04-28T09:57:07.103+0000"},{"body":"If this is happening in 2.1.11, why is this only fixed in 2.2? cc [~blerer]","created":"2016-07-01T01:50:15.891+0000"},{"body":"[~kohlisankalp] It was decided since some time to only push fixes for critical issues to 2.1. The problem becoming do decide what is a critical issue and what is not.\nIn this case, the description is suggesting that the problem is a new one but the location of the problem made us believe that the issue had always been there. By consequence, it was arguable that the problem was a critical one. I concluded, like Alex probably did, that the issue might not be really critical.\nLooking at it now, I agree that we should probably have provided a patch for 2.1. \nAre you also facing the same issue?","created":"2016-07-01T07:07:26.822+0000"},{"body":"The issue was not reproducible on {{2.1}}. The tests are passing. Problem was related to bounds being taken in different order on different stages of query, because of {{EOCs}}, which seems not be the case on 2.1. ","created":"2016-07-01T10:29:45.515+0000"}],"conversations":[{"body":"After we've upgraded our cluster to version 2.1.11, we started getting the bellow exceptions for some of our queries. Issue seems to be very similar to CASSANDRA-7284.\n\nCode to reproduce:\n\n{code:java}\n createTable(\"CREATE TABLE %s (\" +\n \" a text,\" +\n \" b int,\" +\n \" PRIMARY KEY (a, b)\" +\n \") WITH COMPACT STORAGE\" +\n \" AND CLUSTERING ORDER BY (b DESC)\");\n\n execute(\"insert into %s (a, b) values ('a', 2)\");\n execute(\"SELECT * FROM %s WHERE a = 'a' AND b > 0\");\n{code}\n\n{code:java}\njava.lang.ClassCastException: org.apache.cassandra.db.composites.Composites$EmptyComposite cannot be cast to org.apache.cassandra.db.composites.CellName\n at org.apache.cassandra.db.composites.AbstractCellNameType.cellFromByteBuffer(AbstractCellNameType.java:188) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.db.composites.AbstractSimpleCellNameType.makeCellName(AbstractSimpleCellNameType.java:125) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.db.composites.AbstractCellNameType.makeCellName(AbstractCellNameType.java:254) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.makeExclusiveSliceBound(SelectStatement.java:1197) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.applySliceRestriction(SelectStatement.java:1205) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.processColumnFamily(SelectStatement.java:1283) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.process(SelectStatement.java:1250) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.processResults(SelectStatement.java:299) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:276) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:224) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:67) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.QueryProcessor.processStatement(QueryProcessor.java:238) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.QueryProcessor.processPrepared(QueryProcessor.java:493) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.transport.messages.ExecuteMessage.execute(ExecuteMessage.java:138) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:439) [apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:335) [apache-cassandra-2.1.11.jar:2.1.11]\n at io.netty.channel.SimpleChannelInboundHandler.channelRead(SimpleChannelInboundHandler.java:105) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:333) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext.access$700(AbstractChannelHandlerContext.java:32) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext$8.run(AbstractChannelHandlerContext.java:324) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) [na:1.8.0_66]\n at org.apache.cassandra.concurrent.AbstractTracingAwareExecutorService$FutureTask.run(AbstractTracingAwareExecutorService.java:164) [apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [apache-cassandra-2.1.11.jar:2.1.11]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_66]\n{code}","from":"reporter","subject":"isInclusive and boundsAsComposites in Restriction take bounds in different order"},{"body":"We managed to reproduce the bug using Cassandra 2.2.5 with the following table schema (unset parameters use default values):\n{code:sql}\nCREATE TABLE mytable (\n u text,\n t timeuuid,\n PRIMARY KEY (u, t)\n) WITH COMPACT STORAGE\n AND CLUSTERING ORDER BY (t DESC)\n AND compaction = {'min_threshold': '2', 'class': 'org.apache.cassandra.db.compaction.DateTieredCompactionStrategy', 'base_time_seconds': '1'}\n AND compression = {'sstable_compression': 'org.apache.cassandra.io.compress.LZ4Compressor'}\n AND dclocal_read_repair_chance = 0.0\n AND default_time_to_live = 10\n AND gc_grace_seconds = 0\n AND read_repair_chance = 0.0;\n{code}\nAnd the following query:\n{code:sql}\nSELECT COUNT(1) as cnt\nFROM mytable\nWHERE u = :user AND t > :timestamp\nORDER BY t DESC\nLIMIT 110;\n{code}\nRemoving {{COMPACT STORAGE}} definitely helps, so it is somehow connected.","from":"developer"},{"body":"I've tested it with {{2.2}} branch and was able to reproduce it with unit tests (updated summary). {{3.x}} is unaffected.","from":"developer"},{"body":"Created a patch that works, although I'm not entirely satisfied with it. \n\nCurrent problem is that in {{SelectStatement::makeExclusiveSliceBound}}, call to {{areRequestedBoundsInclusive}} (which calls {{isInclusive}} internally), which doesn't reverse {{Bound}} even when clustering order is reversed. On the other hand, next call, {{SelectStatement::getClusteringColumnsBounds}}, reverses the bound internally, which results into the empty {{ByteBuffer}} for the {{Bound}} that's present and non-empty one for the {{Bound}} that's unset.\n\nThe patch is making {{PrimaryKeyRestrictionSet::isInclusive}} to reverse {{Bound}} when needed. Problem is a) that would work only for {{SingleColumnRestriction}}, as there's just one column and b) all \"underlying\" {{isInclusive}} methods are not flipping bounds. It's possible to introduce a different method that'd be more transparent (although that one is only used there iirc) in terms of whether or not it flips the bound. \n\nTests are, however, passing\n\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/10988-2.2]|[testall|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-2.2-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-2.2-dtest/]|","from":"developer"},{"body":"I think that it will be clearer to reverse the {{Bound}} in {{SelectStatement::makeExclusiveSliceBound}} and add a comment to explain it.\n{{SelectStatement::makeExclusiveSliceBound}} will be only called for non-composite slices ({{SingleColumnRestriction.Slice}}).\n\nI noticed that the {{SelectMultiColumnRelationTest::testMixedOrderColumnsX}} methods do not tests {{COMPACT}} tables. It is not directly related to the patch but this problem proves that our testing for {{COMPACT}} tables is not good enough. By consequence, extending it might be a good idea.\n\nIt will be nice if you can also create a patch for {{3.0}} with the unit tests changes. The problem does not affect that version but adding tests is always a good idea to prevent regressions.\n\n","from":"developer"},{"body":"Made the suggested changes (move the logic to {{SelectStatement}}). Also, made the mixed order columns tests also test with {{COMPACT}} tables and added all tests to both {{3.0}} and {{trunk}} (all are passing}}. \n\n|code|utest|dtest|\n|[2.2|https://github.com/ifesdjeen/cassandra/tree/10988-2.2]|[testall|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-2.2-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-2.2-dtest/]|\n|[3.0|https://github.com/ifesdjeen/cassandra/tree/10988-3.0]|[testall|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-3.0-testall/]|\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/10988-trunk]|[testall|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-10988-trunk-testall/]|\n\n(I skipped dtests for {{3.0}} and {{trunk}} since there were no code changes, tests only)\nRan all failing unit tests locally, they're all passing. Dtest shows no failures.","from":"developer"},{"body":"Thanks. +1","from":"developer"},{"body":"Committed into 2.2 at c8d955533b6968368907e5b090a309ac57bf419f and merged into 3.0 and trunk","from":"developer"},{"body":"Thank you!","from":"developer"},{"body":"If this is happening in 2.1.11, why is this only fixed in 2.2? cc [~blerer]","from":"developer"},{"body":"[~kohlisankalp] It was decided since some time to only push fixes for critical issues to 2.1. The problem becoming do decide what is a critical issue and what is not.\nIn this case, the description is suggesting that the problem is a new one but the location of the problem made us believe that the issue had always been there. By consequence, it was arguable that the problem was a critical one. I concluded, like Alex probably did, that the issue might not be really critical.\nLooking at it now, I agree that we should probably have provided a patch for 2.1. \nAre you also facing the same issue?","from":"developer"},{"body":"The issue was not reproducible on {{2.1}}. The tests are passing. Problem was related to bounds being taken in different order on different stages of query, because of {{EOCs}}, which seems not be the case on 2.1. ","from":"developer"}],"created":"2016-01-08T13:34:58.000+0000","description":"After we've upgraded our cluster to version 2.1.11, we started getting the bellow exceptions for some of our queries. Issue seems to be very similar to CASSANDRA-7284.\n\nCode to reproduce:\n\n{code:java}\n createTable(\"CREATE TABLE %s (\" +\n \" a text,\" +\n \" b int,\" +\n \" PRIMARY KEY (a, b)\" +\n \") WITH COMPACT STORAGE\" +\n \" AND CLUSTERING ORDER BY (b DESC)\");\n\n execute(\"insert into %s (a, b) values ('a', 2)\");\n execute(\"SELECT * FROM %s WHERE a = 'a' AND b > 0\");\n{code}\n\n{code:java}\njava.lang.ClassCastException: org.apache.cassandra.db.composites.Composites$EmptyComposite cannot be cast to org.apache.cassandra.db.composites.CellName\n at org.apache.cassandra.db.composites.AbstractCellNameType.cellFromByteBuffer(AbstractCellNameType.java:188) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.db.composites.AbstractSimpleCellNameType.makeCellName(AbstractSimpleCellNameType.java:125) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.db.composites.AbstractCellNameType.makeCellName(AbstractCellNameType.java:254) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.makeExclusiveSliceBound(SelectStatement.java:1197) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.applySliceRestriction(SelectStatement.java:1205) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.processColumnFamily(SelectStatement.java:1283) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.process(SelectStatement.java:1250) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.processResults(SelectStatement.java:299) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:276) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:224) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:67) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.QueryProcessor.processStatement(QueryProcessor.java:238) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.cql3.QueryProcessor.processPrepared(QueryProcessor.java:493) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.transport.messages.ExecuteMessage.execute(ExecuteMessage.java:138) ~[apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:439) [apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:335) [apache-cassandra-2.1.11.jar:2.1.11]\n at io.netty.channel.SimpleChannelInboundHandler.channelRead(SimpleChannelInboundHandler.java:105) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:333) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext.access$700(AbstractChannelHandlerContext.java:32) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext$8.run(AbstractChannelHandlerContext.java:324) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) [na:1.8.0_66]\n at org.apache.cassandra.concurrent.AbstractTracingAwareExecutorService$FutureTask.run(AbstractTracingAwareExecutorService.java:164) [apache-cassandra-2.1.11.jar:2.1.11]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [apache-cassandra-2.1.11.jar:2.1.11]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_66]\n{code}","issue_id":"12928641","key":"CASSANDRA-10988","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-07-04T07:27:43.000+0000","role":"fixed_distractor","summary":"isInclusive and boundsAsComposites in Restriction take bounds in different order"} {"case_id":"12932691","cluster":"DISTRACTOR-CASSANDRA-11043","comments":[{"body":"Any idea about this, [~beobal]? Where do you think is the best place to call to an hypothetical {{validate(Expression)}} method?","created":"2016-01-28T10:16:39.892+0000"},{"body":"[~adelapena], sorry about the slow response - you're right though, the intention is to validate those expressions earlier and reject before getting to the point of execution. Unfortunately, the unit test covering that is not properly representative of running on a real server so it didn't catch the problem you described. I think we also need to add back the early validation of non-custom expressions, as although the custom expression syntax is a better fit for indexes which don't map to a specific column, for backwards compatibility it's still possible to do things in the old style, by creating a fake column & adding an index on it. The fix should be reasonably straightforward, but I'd like to have a think about how to test this better (using custom indexes etc in dtests is not so straightforward). I'll aim to have a patch ready sometime this week.","created":"2016-02-02T18:38:08.330+0000"},{"body":"Yes, it would be good to preserve the query validation also for the old fake column approach. Maybe a {{validate(ReadCommand)}} method, called from CQL layer before executing the command itself, could work for both kinds of indexes. It's not our use case, but I think that other implementations could be interested in validating the extra information contained in the {{ReadCommand}}, such as the query limits or the allow filtering clauses. It's just an idea.\n\nI'm not sure about how to test this feature. I guess that using dtests would involve to add stub index implementations to the tested server's class path, which sounds complicated...\n\nPlease let me know if I can help you in some way.","created":"2016-02-04T10:19:47.794+0000"},{"body":"Turns out modifying the testing was more straightforward than I'd thought. I've modified all of the tests in {{CustomIndexTest}} which are verifying query time checks to execute queries at the protocol level using the java driver, rather than by calling {{QueryProcessor.execute}} directly. To fix the reported issue I've added a {{validate(ReadCommand command)}} method to {{Index}}, which is called when the read command is first constructed (in {{SelectStatement}} and {{CassandraServer}}). There's a default no-op implementation provided, so it's fully backwards compatible with existing implementations. \n\n||branch||testall||dtest||\n|[11043-3.0|https://github.com/beobal/cassandra/tree/11043-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-11043-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-11043-3.0-dtest]|\n|[11043-trunk|https://github.com/beobal/cassandra/tree/11043-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-11043-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-11043-trunk-dtest]|\n","created":"2016-02-09T09:51:58.984+0000"},{"body":"It looks awesome! The new {{validate(ReadCommand command)}} method is perfect, and not requiring dtests to test validation is also great.","created":"2016-02-10T11:52:47.881+0000"},{"body":"It works like a charm, thanks!","created":"2016-02-12T14:42:19.130+0000"},{"body":"thanks, committed to 3.0 in {{9cfbc31bc29685bd60355a823e0cf261a89858f0}} and merged to trunk.\n\n[~adelapena] just for future reference, the \"Resolved\" status is for when the fix has actually been committed. Where the reviewer isn't a committer, the \"Ready to Commit\" status is what you want. \n","created":"2016-02-15T13:18:15.965+0000"},{"body":"Ok, I'm sorry, I didn't realize. I´ll keep it in mind in the future. Thanks for your help.","created":"2016-02-15T13:39:03.655+0000"},{"body":"The validation method works great with CQL queries producing a {{PartitionRangeReadCommand}}, but it is not called with CQL queries producing one or more {{SinglePartitionReadCommand}}.\n\nFor example, the following query properly validates the expression:\n{code}\nSELECT * FROM test WHERE expr(test_idx, 'error');\n{code}\nHowever the following queries skip validation:\n{code}\nSELECT * FROM test WHERE expr(test_idx, 'error') AND id=1;\nSELECT * FROM test WHERE expr(test_idx, 'error') AND id IN (1,2);\n{code}","created":"2016-04-19T16:09:19.938+0000"},{"body":"The problem here is that queries involving indexes should always be executed as a {{PartitionRangeReadCommand}}. When the statement restrictions are processed, we don't set the {{usesSecondaryIndexing}} flag if custom expressions are present, which we should do. If the flag were set, then we'd never get into the situation described here. We encountered a similar problem on CASSANDRA-11310 recently, and as I mentioned in [the comment there|https://issues.apache.org/jira/browse/CASSANDRA-11310?focusedCommentId=15245922&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15245922], although we could (and most likely will) move to enable 2i usage on single partition reads, we need to do that in controlled way to ensure we don't violate an assumptions being made about the read command. In the meantime, I've opened CASSANDRA-11617 to fix the bug with custom expressions.\n","created":"2016-04-20T12:12:39.973+0000"}],"conversations":[{"body":"It seems that [CASSANDRA-7575|https://issues.apache.org/jira/browse/CASSANDRA-7575] is broken in Cassandra 3.x. As stated in the secondary indexes' API documentation, custom index implementations should perform any validation of query expressions at {{Index#searcherFor(ReadCommand)}}, throwing an {{InvalidRequestException}} if the expressions are not valid. I assume these validation errors should produce an {{InvalidRequest}} error on cqlsh, or raise an {{InvalidQueryException}} on Java driver. However, when {{Index#searcherFor(ReadCommand)}} throws its {{InvalidRequestException}}, I get this cqlsh output:\n{noformat}\nTraceback (most recent call last):\n File \"bin/cqlsh.py\", line 1246, in perform_simple_statement\n result = future.result()\n File \"/Users/adelapena/stratio/platform/src/cassandra-3.2.1/bin/../lib/cassandra-driver-internal-only-3.0.0-6af642d.zip/cassandra-driver-3.0.0-6af642d/cassandra/cluster.py\", line 3122, in result\n raise self._final_exception\nReadFailure: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 1 failures\" info={'failures': 1, 'received_responses': 0, 'required_responses': 1, 'consistency': 'ONE'}\n{noformat}\n\nI attach a dummy index implementation to reproduce the error:\n{noformat}\nCREATE KEYSPACE test with replication = {'class' : 'SimpleStrategy', 'replication_factor' : '1' }; \nCREATE TABLE test.test (id int PRIMARY KEY, value varchar); \nCREATE CUSTOM INDEX test_index ON test.test() USING 'com.stratio.TestIndex'; \nSELECT * FROM test.test WHERE expr(test_index,'ok');\nSELECT * FROM test.test WHERE expr(test_index,'error');\n{noformat}\nThis is specially problematic when using Cassandra Java Driver, because one of these server exceptions can produce subsequent queries fail (even if they are valid) with a no host available exception.\n\nMaybe the validation method added with [CASSANDRA-7575|https://issues.apache.org/jira/browse/CASSANDRA-7575] should be restored, unless there is a way to properly manage the exception.","from":"reporter","subject":"Secondary indexes doesn't properly validate custom expressions"},{"body":"Any idea about this, [~beobal]? Where do you think is the best place to call to an hypothetical {{validate(Expression)}} method?","from":"developer"},{"body":"[~adelapena], sorry about the slow response - you're right though, the intention is to validate those expressions earlier and reject before getting to the point of execution. Unfortunately, the unit test covering that is not properly representative of running on a real server so it didn't catch the problem you described. I think we also need to add back the early validation of non-custom expressions, as although the custom expression syntax is a better fit for indexes which don't map to a specific column, for backwards compatibility it's still possible to do things in the old style, by creating a fake column & adding an index on it. The fix should be reasonably straightforward, but I'd like to have a think about how to test this better (using custom indexes etc in dtests is not so straightforward). I'll aim to have a patch ready sometime this week.","from":"developer"},{"body":"Yes, it would be good to preserve the query validation also for the old fake column approach. Maybe a {{validate(ReadCommand)}} method, called from CQL layer before executing the command itself, could work for both kinds of indexes. It's not our use case, but I think that other implementations could be interested in validating the extra information contained in the {{ReadCommand}}, such as the query limits or the allow filtering clauses. It's just an idea.\n\nI'm not sure about how to test this feature. I guess that using dtests would involve to add stub index implementations to the tested server's class path, which sounds complicated...\n\nPlease let me know if I can help you in some way.","from":"developer"},{"body":"Turns out modifying the testing was more straightforward than I'd thought. I've modified all of the tests in {{CustomIndexTest}} which are verifying query time checks to execute queries at the protocol level using the java driver, rather than by calling {{QueryProcessor.execute}} directly. To fix the reported issue I've added a {{validate(ReadCommand command)}} method to {{Index}}, which is called when the read command is first constructed (in {{SelectStatement}} and {{CassandraServer}}). There's a default no-op implementation provided, so it's fully backwards compatible with existing implementations. \n\n||branch||testall||dtest||\n|[11043-3.0|https://github.com/beobal/cassandra/tree/11043-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-11043-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-11043-3.0-dtest]|\n|[11043-trunk|https://github.com/beobal/cassandra/tree/11043-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-11043-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-11043-trunk-dtest]|\n","from":"developer"},{"body":"It looks awesome! The new {{validate(ReadCommand command)}} method is perfect, and not requiring dtests to test validation is also great.","from":"developer"},{"body":"It works like a charm, thanks!","from":"developer"},{"body":"thanks, committed to 3.0 in {{9cfbc31bc29685bd60355a823e0cf261a89858f0}} and merged to trunk.\n\n[~adelapena] just for future reference, the \"Resolved\" status is for when the fix has actually been committed. Where the reviewer isn't a committer, the \"Ready to Commit\" status is what you want. \n","from":"developer"},{"body":"Ok, I'm sorry, I didn't realize. I´ll keep it in mind in the future. Thanks for your help.","from":"developer"},{"body":"The validation method works great with CQL queries producing a {{PartitionRangeReadCommand}}, but it is not called with CQL queries producing one or more {{SinglePartitionReadCommand}}.\n\nFor example, the following query properly validates the expression:\n{code}\nSELECT * FROM test WHERE expr(test_idx, 'error');\n{code}\nHowever the following queries skip validation:\n{code}\nSELECT * FROM test WHERE expr(test_idx, 'error') AND id=1;\nSELECT * FROM test WHERE expr(test_idx, 'error') AND id IN (1,2);\n{code}","from":"developer"},{"body":"The problem here is that queries involving indexes should always be executed as a {{PartitionRangeReadCommand}}. When the statement restrictions are processed, we don't set the {{usesSecondaryIndexing}} flag if custom expressions are present, which we should do. If the flag were set, then we'd never get into the situation described here. We encountered a similar problem on CASSANDRA-11310 recently, and as I mentioned in [the comment there|https://issues.apache.org/jira/browse/CASSANDRA-11310?focusedCommentId=15245922&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15245922], although we could (and most likely will) move to enable 2i usage on single partition reads, we need to do that in controlled way to ensure we don't violate an assumptions being made about the read command. In the meantime, I've opened CASSANDRA-11617 to fix the bug with custom expressions.\n","from":"developer"}],"created":"2016-01-20T12:24:23.000+0000","description":"It seems that [CASSANDRA-7575|https://issues.apache.org/jira/browse/CASSANDRA-7575] is broken in Cassandra 3.x. As stated in the secondary indexes' API documentation, custom index implementations should perform any validation of query expressions at {{Index#searcherFor(ReadCommand)}}, throwing an {{InvalidRequestException}} if the expressions are not valid. I assume these validation errors should produce an {{InvalidRequest}} error on cqlsh, or raise an {{InvalidQueryException}} on Java driver. However, when {{Index#searcherFor(ReadCommand)}} throws its {{InvalidRequestException}}, I get this cqlsh output:\n{noformat}\nTraceback (most recent call last):\n File \"bin/cqlsh.py\", line 1246, in perform_simple_statement\n result = future.result()\n File \"/Users/adelapena/stratio/platform/src/cassandra-3.2.1/bin/../lib/cassandra-driver-internal-only-3.0.0-6af642d.zip/cassandra-driver-3.0.0-6af642d/cassandra/cluster.py\", line 3122, in result\n raise self._final_exception\nReadFailure: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 1 failures\" info={'failures': 1, 'received_responses': 0, 'required_responses': 1, 'consistency': 'ONE'}\n{noformat}\n\nI attach a dummy index implementation to reproduce the error:\n{noformat}\nCREATE KEYSPACE test with replication = {'class' : 'SimpleStrategy', 'replication_factor' : '1' }; \nCREATE TABLE test.test (id int PRIMARY KEY, value varchar); \nCREATE CUSTOM INDEX test_index ON test.test() USING 'com.stratio.TestIndex'; \nSELECT * FROM test.test WHERE expr(test_index,'ok');\nSELECT * FROM test.test WHERE expr(test_index,'error');\n{noformat}\nThis is specially problematic when using Cassandra Java Driver, because one of these server exceptions can produce subsequent queries fail (even if they are valid) with a no host available exception.\n\nMaybe the validation method added with [CASSANDRA-7575|https://issues.apache.org/jira/browse/CASSANDRA-7575] should be restored, unless there is a way to properly manage the exception.","issue_id":"12932691","key":"CASSANDRA-11043","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-04-20T12:12:39.000+0000","role":"fixed_distractor","summary":"Secondary indexes doesn't properly validate custom expressions"} {"case_id":"12938653","cluster":"DISTRACTOR-CASSANDRA-11163","comments":[{"body":"Also, the bloom filter is rebuilt on load(), but the rebuild (corrected) filter isn't persisted to disk (could make the argument that rebuilding the BF is right, but if you're going to take the time to do that, should also persist it to disk in the same way index summary is persisted).\n\n","created":"2016-02-12T00:58:08.827+0000"},{"body":"FWIW - this is causing startup on nodes in our cluster to take in excess of 15mins. We do have roughly 1.5TB on disk. We're on a patched version of 3.9.\n\nWhat I don't quite understand is why this behavior persists past the first restart of the BF FP?","created":"2017-01-20T21:26:29.030+0000"},{"body":"It persists until all sstables are rewritten - the bloom filter is recalculated on startup but not persisted to disk until compaction runs. You can force the bloom filter to be rewritten without waiting dircompaction or this bug to be fixed by running \"nodetool upgradesstables -a\" to rewrite all of the sstables (note that this can be IO intensive and impact your latencies)","created":"2017-01-21T02:36:06.618+0000"},{"body":"[~KurtG] when you're patch available on this, ping me, I'll review and commit it for you.\r\n\r\nAlso marking as 3.0+ because it's a legit bug and not a feature; if you don't plan on doing a 3.0 patch, let me know and I'll adjust.\r\n\r\n \r\n\r\n ","created":"2018-01-25T01:55:52.777+0000"},{"body":"I'll have a patch up soon (next few days). Just need to write some tests. Will likely also solve CASSANDRA-14166","created":"2018-01-29T00:28:26.243+0000"},{"body":"Patches for each branch below. So far I've only got a green test run for trunk. Failures for 3.0 and 3.11 seem flaky/unrelated so will keep trying.\r\n|[trunk|https://github.com/apache/cassandra/compare/trunk...kgreav:14166-trunk]|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...kgreav:14166-3.11]|[3.0|https://github.com/apache/cassandra/compare/cassandra-3.0...kgreav:14166-3.0]|\r\n|[utests|https://circleci.com/gh/kgreav/cassandra/66]|\r\n\r\nThis patch also solves CASSANDRA-14166. Basically I've completely stopped regeneration of Summaries on startup (if BFFP is changed), and also stopped the behaviour implemented in CASSANDRA-5015 which would also regenerate the bloomfilters on startup.\r\nThere's definitely no reason to regenerate Summaries in this case, and as previously mentioned it's not great regenerating the bloomfilter unless you're going to persist it. I have added persistence for the bloomfilter (when it is regenerated), however I think it's a bad idea to do this on startup as it will likely be more time consuming than regenerating the summaries.\r\n\r\nIf an operator chooses to these can be regenerated through {{upgradesstables -a}}. Albeit that's not super efficient if you're just updating the bloomfilter, however I think it's good enough for the moment and a potential follow up ticket would be to add a nodetool command to regenerate bloomfilter/summaries/index/etc.\r\n\r\nThe new behaviour would be to:\r\n# Never recreate Summary or bloomfilter when using an offline tool. Note that if the summary doesn't exist the tool will still create it in memory, but it won't persist it to disk (solving CASSANDRA-14166)\r\n# Only regenerate the summary when it can't be loaded, OR when we've said to recreate the bloomfilter. However we only save the summary if we've been explicitly told to (which should be always EXCEPT for offline tools).\r\n# Only regenerate and persist the bloomfilter when it's missing - not when it has changed. This means we rely on compactions/upgradesstables to update the bloomfilter. \r\n\r\nI've updated {{org.apache.cassandra.io.sstable.SSTableReaderTest#testOpeningSSTable}} to hopefully test all of these cases.","created":"2018-02-01T23:35:28.512+0000"},{"body":"* In {{load(ValidationMetadata validation, boolean isOffline)}} everywhere your calling {{load( bool , true )}} you can instead call \\{{ load( bool, !isOffline) }} since you never want to save the summary in those other situations either. This will break your test but IMHO thats checking that the wrong case occurs. If the summary file is not there, it should not create it. Tools and such may be running with a different user, if someone runs this on a data directory and this occurs it will create a file that C* would be unable to delete, causing compaction threads to die and backup etc. I think, in offline mode the tools should _never_ delete, touch or create unnecessary files, especially the summary/bf files since they are mostly there to speed up startup and not necessary for the reader to work anyway. You can also make the \"recreateBloomFilter\" always false in offline mode (whenever its true, instead put !isOffline) since it will then just use whats there. With one exception of where the FILTER component is missing, where you can just put AlwaysPresent bf and skip so that code that uses it doesn't NPE.\r\n\r\n * In unit tests, is the 1000ms sleep necessary? the lastModified is in ms so I thought it may be ok to set lower\r\n\r\n * Just checking it out and running it over and over, the unit tests fails occasionally (rarely) (line 407 check {{assertNotEquals(bloomModified, bloomFile.lastModified());}} is the same)\r\n\r\n * NP: I think you can reuse the last option (track hotness) since its only false currently in situations where we dont want or need to recreate currently. If rename it to like \"allowChanges\". That way we are not adding additional booleans to end of that load function.","created":"2018-02-21T08:24:38.038+0000"},{"body":"Thanks for the review Chris. Made changes according to your comments and updated the branches.\r\n# That works. It also didn't break tests because there was that initial catchall isOfflline at the top which covered all those cases. Note that ATM the only way the summary is recreated is by online operations (using open, rather than openNoValidation).\r\n# No it wasn't, assumed in Java we only had access to second precision which is not true in Java 8, as long as you use {{Files}}. I've updated the test to use 10ms sleeps with ms precision and it seems to be much more reliable. 99% sure the rare failures before were just because the second precision of {{lastModified()}}\r\n# I've updated to just use {{isOffline}}. It seemed to make more sense than {{allowChanges}} to me, in that \"we don't track hotness or touch files if we're offline\". I'll note that this changed the behaviour of {{org.apache.cassandra.db.ColumnFamilyStore#getSnapshotSSTableReader}} to be \"offline\", however I think that makes sense in this case, as we shouldn't be regenerating summaries/BF's for snapshots anyway.\r\n\r\n|[3.0|https://github.com/apache/cassandra/compare/cassandra-3.0...kgreav:14166-3.0]|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...kgreav:14166-3.11]|[trunk|https://github.com/apache/cassandra/compare/trunk...kgreav:14166-trunk]|\r\n","created":"2018-03-02T02:04:46.073+0000"},{"body":"Doesnt this now miss case where under normal (not offline) run and they change bf ratio it no longer recreates them? {{(validation.bloomFilterFPChance != metadata().params.bloomFilterFpChance)}}. With this it will just load the old bf so theres no way to change it.","created":"2018-03-05T16:51:43.918+0000"},{"body":"Correct. As I noted previously \r\n\r\nbq. Only regenerate and persist the bloomfilter when it's missing - not when it has changed. This means we rely on compactions/upgradesstables to update the bloomfilter.\r\nbq. There's definitely no reason to regenerate Summaries in this case, and as previously mentioned it's not great regenerating the bloomfilter unless you're going to persist it. I have added persistence for the bloomfilter (when it is regenerated), however I think it's a bad idea to do this on startup as it will likely be more time consuming than regenerating the summaries.\r\n\r\nSo the previous behaviour was to regenerate the BF in this case but *not* persist it on the next startup (this meant it would happen on every startup until compactions/upgrades had occured). The summaries would be regenerated and persisted on the next startup (pointlessly). Both of these things would slow startup time pretty significantly depending on how much data you had.\r\n\r\nThe new behaviour would be to avoid regenerating BF/Summaries at all on startup and instead rely on upgradesstables/compactions to update them. Summaries would only be recreated when necessary (when not loaded/corrupt/missing).\r\n\r\nIn trunk it might make sense to also add a nodetool command that will allow us to regenerate the bloomfilters/summaries/etc without re-writing the whole data file.","created":"2018-03-06T00:00:58.568+0000"},{"body":"gotcha +1 on patch from me\r\n\r\nI like idea of a {{nodetool recreatecomponent}} or something but that would be a different jira.","created":"2018-03-06T00:05:25.731+0000"},{"body":"Thanks for the review [~cnlwsu]. Created CASSANDRA-14291 as a follow up.","created":"2018-03-06T00:47:58.587+0000"},{"body":"||branch||testall||dtest||\r\n|[14166-3.0|https://github.com/kgreav/cassandra/tree/14166-3.0]|[testall|https://circleci.com/gh/kgreav/cassandra/tree/14166-3.0]|[dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/506/]|\r\n|[14166-3.11|https://github.com/kgreav/cassandra/tree/14166-3.11]|[testall|https://circleci.com/gh/kgreav/cassandra/tree/14166-3.11]|[dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/507]|\r\n|[14166-trunk|https://github.com/kgreav/cassandra/tree/14166-trunk]|[testall|https://circleci.com/gh/kgreav/cassandra/tree/14166-trunk]|[dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/508]|\r\n\r\nEDIT: i had [troubles|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/510/] getting the trunk patch dtests to run. As upstream trunk dtests appear to now be [working|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-trunk-dtest/462/], I, off a [thelastpickle|https://github.com/thelastpickle/cassandra/tree/14166-trunk] fork, rebased that patch off trunk and am running the dtests [again|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/511/]…","created":"2018-03-09T07:25:02.008+0000"},{"body":"Rebased trunk patch run the dtests. \r\n\r\nI don't believe the failed test has anything to do with this ticket: [repair_tests.repair_test.TestRepair.test_dc_parallel_repair|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/511/testReport/junit/repair_tests.repair_test/TestRepair/test_dc_parallel_repair/]","created":"2018-03-12T02:44:35.554+0000"},{"body":"Committed.","created":"2018-03-12T08:37:19.302+0000"},{"body":"This commit broke 3.0 unit tests (maybe 3.11 too, haven't checked).\r\n\r\nSee https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-3.0-test-all/lastCompletedBuild/testReport/org.apache.cassandra.io.sstable/SSTableReaderTest/testOpeningSSTable/history/ and https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-3.0-test-all/248/changes. It's also failing locally and I'm pretty sure will be failing [here|https://circleci.com/gh/iamaleksey/cassandra/234] too.\r\n\r\nWe should be more careful about committing things. But in the meantime, [~mck] can you please fix this or revert this until there is a fix? Cheers.","created":"2018-04-30T13:01:35.727+0000"},{"body":"thought this was fixed in https://issues.apache.org/jira/browse/CASSANDRA-14387 ? time resolution being different on causing it to fail on some systems.","created":"2018-04-30T13:12:51.484+0000"},{"body":"Yeah. As you (and Marcus, offline) mentioned, CASSANDRA-14387 was only committed to trunk, so 3.0 tests are still failing elsewhere (incl. Jenkins, but not Circle). We just need to commit 14387 to 3.0 and 3.11.","created":"2018-04-30T13:18:05.301+0000"},{"body":"{quote}Cherry-plicked into 3.0 as [e16f0ed0698c5cb47ab2bb0a0b04966d5bdbcde0|https://github.com/apache/cassandra/commit/e16f0ed0698c5cb47ab2bb0a0b04966d5bdbcde0] and merged upwards\r\n{quote}\r\nThanks [~iamaleksey] for fixing this.\r\n{quote}We should be more careful about committing things.\r\n{quote}\r\nYes, I should have noticed that {{SSTableReaderTest}} was a new failure in  [https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-3.0-test-all/248/#showFailuresLink] \r\n\r\nUp until now I've only been using circleci unit test reports, as illustrated in the \"branch|testall|dtest\" comment above, which reported it all green.\r\n\r\nAs CircleCI hasn't failed on this test, what's the difference here? \r\nAleksey's circleci build #234 [passed SSTableReaderTest|https://circleci.com/gh/iamaleksey/cassandra/234#tests/containers/2] , as did Kurt's [here|https://circleci.com/gh/kgreav/cassandra/116#tests/containers/1]\r\n\r\n \r\n\r\nIf CircleCi and Jenkins can both identify failures separately then \"test-all\" column should be split into two, to link to each.\r\n\r\nOtherwise, does Jenkins have any form of alerting or reporting on tests that go from stable to flakey? Would have been ideal to have automatically have receive the breakage information before it affected you [~iamaleksey]","created":"2018-04-30T23:30:54.052+0000"},{"body":"the granularity of the file timestamp depends on the filesystem","created":"2018-04-30T23:32:20.176+0000"},{"body":"[~cnlwsu], oh! not os but filesystem, and circleci and jenkins differ there. Ouch. \r\nHow should we test against that?","created":"2018-04-30T23:36:44.706+0000"},{"body":"Hm, sorry about that. Originally started with second delays but changed it to speed up the test. Thought ms would be safe in this day and age. 2018 and all that.\r\n\r\n{quote}How should we test against that?\r\n{quote}\r\nWe don't, just have to go with the lowest resolution which will hopefully be seconds. No point putting workarounds in to speed up a test by a few seconds on _some_ platforms. Having said that, what filesystem is jenkins running? Seems odd it'd only have second precision. But then again it's probably an ancient server.","created":"2018-05-01T01:28:37.093+0000"},{"body":"There isn't a clean or standard way to determine if a file system supports sub-second date/time resolution for file modifications.","created":"2018-05-01T06:25:40.355+0000"},{"body":"bq. Yes, I should have noticed that SSTableReaderTest was a new failure in https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-3.0-test-all/248/#showFailuresLink \r\n\r\nRealistically, you shouldn't have. Looking at Circle should be sufficient enough. I just assumed that it would break on Circle just like it broke in our internal CI and Jenkins, and that was wrong. My apologies for the somewhat passive-aggressive righteous tone.\r\n\r\nIf Circle is green, we should be free to commit, and we should be checking with ASF Jenkins from time to time, but it should not be required for every commit.","created":"2018-05-01T10:39:12.774+0000"}],"conversations":[{"body":"This is from trunk, but I also saw this happen on 2.0:\n\nBefore:\n{noformat}\nroot@bw-1:/srv/cassandra# ls -ltr /var/lib/cassandra/data/keyspace1/standard1-071efdc0d11811e590c3413ee28a6c90/\ntotal 221460\ndrwxr-xr-x 2 root root 4096 Feb 11 23:34 backups\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-6-big-TOC.txt\n-rw-r--r-- 1 root root 26518 Feb 11 23:50 ma-6-big-Summary.db\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-6-big-Statistics.db\n-rw-r--r-- 1 root root 2607705 Feb 11 23:50 ma-6-big-Index.db\n-rw-r--r-- 1 root root 192440 Feb 11 23:50 ma-6-big-Filter.db\n-rw-r--r-- 1 root root 10 Feb 11 23:50 ma-6-big-Digest.crc32\n-rw-r--r-- 1 root root 35212125 Feb 11 23:50 ma-6-big-Data.db\n-rw-r--r-- 1 root root 2156 Feb 11 23:50 ma-6-big-CRC.db\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-7-big-TOC.txt\n-rw-r--r-- 1 root root 26518 Feb 11 23:50 ma-7-big-Summary.db\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-7-big-Statistics.db\n-rw-r--r-- 1 root root 2607614 Feb 11 23:50 ma-7-big-Index.db\n-rw-r--r-- 1 root root 192432 Feb 11 23:50 ma-7-big-Filter.db\n-rw-r--r-- 1 root root 9 Feb 11 23:50 ma-7-big-Digest.crc32\n-rw-r--r-- 1 root root 35190400 Feb 11 23:50 ma-7-big-Data.db\n-rw-r--r-- 1 root root 2152 Feb 11 23:50 ma-7-big-CRC.db\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-5-big-TOC.txt\n-rw-r--r-- 1 root root 104178 Feb 11 23:50 ma-5-big-Summary.db\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-5-big-Statistics.db\n-rw-r--r-- 1 root root 10289077 Feb 11 23:50 ma-5-big-Index.db\n-rw-r--r-- 1 root root 757384 Feb 11 23:50 ma-5-big-Filter.db\n-rw-r--r-- 1 root root 9 Feb 11 23:50 ma-5-big-Digest.crc32\n-rw-r--r-- 1 root root 139201355 Feb 11 23:50 ma-5-big-Data.db\n-rw-r--r-- 1 root root 8508 Feb 11 23:50 ma-5-big-CRC.db\nroot@bw-1:/srv/cassandra# md5sum /var/lib/cassandra/data/keyspace1/standard1-071efdc0d11811e590c3413ee28a6c90/ma-5-big-Summary.db\n5fca154fc790f7cfa37e8ad6d1c7552c\n{noformat}\n\nBF ratio changed, node restarted:\n{noformat}\nroot@bw-1:/srv/cassandra# ls -ltr /var/lib/cassandra/data/keyspace1/standard1-071efdc0d11811e590c3413ee28a6c90/\ntotal 242168\ndrwxr-xr-x 2 root root 4096 Feb 11 23:34 backups\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-6-big-TOC.txt\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-6-big-Statistics.db\n-rw-r--r-- 1 root root 2607705 Feb 11 23:50 ma-6-big-Index.db\n-rw-r--r-- 1 root root 192440 Feb 11 23:50 ma-6-big-Filter.db\n-rw-r--r-- 1 root root 10 Feb 11 23:50 ma-6-big-Digest.crc32\n-rw-r--r-- 1 root root 35212125 Feb 11 23:50 ma-6-big-Data.db\n-rw-r--r-- 1 root root 2156 Feb 11 23:50 ma-6-big-CRC.db\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-7-big-TOC.txt\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-7-big-Statistics.db\n-rw-r--r-- 1 root root 2607614 Feb 11 23:50 ma-7-big-Index.db\n-rw-r--r-- 1 root root 192432 Feb 11 23:50 ma-7-big-Filter.db\n-rw-r--r-- 1 root root 9 Feb 11 23:50 ma-7-big-Digest.crc32\n-rw-r--r-- 1 root root 35190400 Feb 11 23:50 ma-7-big-Data.db\n-rw-r--r-- 1 root root 2152 Feb 11 23:50 ma-7-big-CRC.db\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-5-big-TOC.txt\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-5-big-Statistics.db\n-rw-r--r-- 1 root root 10289077 Feb 11 23:50 ma-5-big-Index.db\n-rw-r--r-- 1 root root 757384 Feb 11 23:50 ma-5-big-Filter.db\n-rw-r--r-- 1 root root 9 Feb 11 23:50 ma-5-big-Digest.crc32\n-rw-r--r-- 1 root root 139201355 Feb 11 23:50 ma-5-big-Data.db\n-rw-r--r-- 1 root root 8508 Feb 11 23:50 ma-5-big-CRC.db\n-rw-r--r-- 1 root root 80 Feb 12 00:03 ma-8-big-TOC.txt\n-rw-r--r-- 1 root root 14902 Feb 12 00:03 ma-8-big-Summary.db\n-rw-r--r-- 1 root root 10264 Feb 12 00:03 ma-8-big-Statistics.db\n-rw-r--r-- 1 root root 1458631 Feb 12 00:03 ma-8-big-Index.db\n-rw-r--r-- 1 root root 10808 Feb 12 00:03 ma-8-big-Filter.db\n-rw-r--r-- 1 root root 10 Feb 12 00:03 ma-8-big-Digest.crc32\n-rw-r--r-- 1 root root 19660275 Feb 12 00:03 ma-8-big-Data.db\n-rw-r--r-- 1 root root 1204 Feb 12 00:03 ma-8-big-CRC.db\n-rw-r--r-- 1 root root 26518 Feb 12 00:04 ma-7-big-Summary.db\n-rw-r--r-- 1 root root 26518 Feb 12 00:04 ma-6-big-Summary.db\n-rw-r--r-- 1 root root 104178 Feb 12 00:04 ma-5-big-Summary.db\nroot@bw-1:/srv/cassandra# md5sum /var/lib/cassandra/data/keyspace1/standard1-071efdc0d11811e590c3413ee28a6c90/ma-5-big-Summary.db \n5fca154fc790f7cfa37e8ad6d1c7552c \n{noformat}\n\nThis hurts startup time and appears to do nothing useful whatsoever.","from":"reporter","subject":"Summaries are needlessly rebuilt when the BF FP ratio is changed"},{"body":"Also, the bloom filter is rebuilt on load(), but the rebuild (corrected) filter isn't persisted to disk (could make the argument that rebuilding the BF is right, but if you're going to take the time to do that, should also persist it to disk in the same way index summary is persisted).\n\n","from":"developer"},{"body":"FWIW - this is causing startup on nodes in our cluster to take in excess of 15mins. We do have roughly 1.5TB on disk. We're on a patched version of 3.9.\n\nWhat I don't quite understand is why this behavior persists past the first restart of the BF FP?","from":"developer"},{"body":"It persists until all sstables are rewritten - the bloom filter is recalculated on startup but not persisted to disk until compaction runs. You can force the bloom filter to be rewritten without waiting dircompaction or this bug to be fixed by running \"nodetool upgradesstables -a\" to rewrite all of the sstables (note that this can be IO intensive and impact your latencies)","from":"developer"},{"body":"[~KurtG] when you're patch available on this, ping me, I'll review and commit it for you.\r\n\r\nAlso marking as 3.0+ because it's a legit bug and not a feature; if you don't plan on doing a 3.0 patch, let me know and I'll adjust.\r\n\r\n \r\n\r\n ","from":"developer"},{"body":"I'll have a patch up soon (next few days). Just need to write some tests. Will likely also solve CASSANDRA-14166","from":"developer"},{"body":"Patches for each branch below. So far I've only got a green test run for trunk. Failures for 3.0 and 3.11 seem flaky/unrelated so will keep trying.\r\n|[trunk|https://github.com/apache/cassandra/compare/trunk...kgreav:14166-trunk]|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...kgreav:14166-3.11]|[3.0|https://github.com/apache/cassandra/compare/cassandra-3.0...kgreav:14166-3.0]|\r\n|[utests|https://circleci.com/gh/kgreav/cassandra/66]|\r\n\r\nThis patch also solves CASSANDRA-14166. Basically I've completely stopped regeneration of Summaries on startup (if BFFP is changed), and also stopped the behaviour implemented in CASSANDRA-5015 which would also regenerate the bloomfilters on startup.\r\nThere's definitely no reason to regenerate Summaries in this case, and as previously mentioned it's not great regenerating the bloomfilter unless you're going to persist it. I have added persistence for the bloomfilter (when it is regenerated), however I think it's a bad idea to do this on startup as it will likely be more time consuming than regenerating the summaries.\r\n\r\nIf an operator chooses to these can be regenerated through {{upgradesstables -a}}. Albeit that's not super efficient if you're just updating the bloomfilter, however I think it's good enough for the moment and a potential follow up ticket would be to add a nodetool command to regenerate bloomfilter/summaries/index/etc.\r\n\r\nThe new behaviour would be to:\r\n# Never recreate Summary or bloomfilter when using an offline tool. Note that if the summary doesn't exist the tool will still create it in memory, but it won't persist it to disk (solving CASSANDRA-14166)\r\n# Only regenerate the summary when it can't be loaded, OR when we've said to recreate the bloomfilter. However we only save the summary if we've been explicitly told to (which should be always EXCEPT for offline tools).\r\n# Only regenerate and persist the bloomfilter when it's missing - not when it has changed. This means we rely on compactions/upgradesstables to update the bloomfilter. \r\n\r\nI've updated {{org.apache.cassandra.io.sstable.SSTableReaderTest#testOpeningSSTable}} to hopefully test all of these cases.","from":"developer"},{"body":"* In {{load(ValidationMetadata validation, boolean isOffline)}} everywhere your calling {{load( bool , true )}} you can instead call \\{{ load( bool, !isOffline) }} since you never want to save the summary in those other situations either. This will break your test but IMHO thats checking that the wrong case occurs. If the summary file is not there, it should not create it. Tools and such may be running with a different user, if someone runs this on a data directory and this occurs it will create a file that C* would be unable to delete, causing compaction threads to die and backup etc. I think, in offline mode the tools should _never_ delete, touch or create unnecessary files, especially the summary/bf files since they are mostly there to speed up startup and not necessary for the reader to work anyway. You can also make the \"recreateBloomFilter\" always false in offline mode (whenever its true, instead put !isOffline) since it will then just use whats there. With one exception of where the FILTER component is missing, where you can just put AlwaysPresent bf and skip so that code that uses it doesn't NPE.\r\n\r\n * In unit tests, is the 1000ms sleep necessary? the lastModified is in ms so I thought it may be ok to set lower\r\n\r\n * Just checking it out and running it over and over, the unit tests fails occasionally (rarely) (line 407 check {{assertNotEquals(bloomModified, bloomFile.lastModified());}} is the same)\r\n\r\n * NP: I think you can reuse the last option (track hotness) since its only false currently in situations where we dont want or need to recreate currently. If rename it to like \"allowChanges\". That way we are not adding additional booleans to end of that load function.","from":"developer"},{"body":"Thanks for the review Chris. Made changes according to your comments and updated the branches.\r\n# That works. It also didn't break tests because there was that initial catchall isOfflline at the top which covered all those cases. Note that ATM the only way the summary is recreated is by online operations (using open, rather than openNoValidation).\r\n# No it wasn't, assumed in Java we only had access to second precision which is not true in Java 8, as long as you use {{Files}}. I've updated the test to use 10ms sleeps with ms precision and it seems to be much more reliable. 99% sure the rare failures before were just because the second precision of {{lastModified()}}\r\n# I've updated to just use {{isOffline}}. It seemed to make more sense than {{allowChanges}} to me, in that \"we don't track hotness or touch files if we're offline\". I'll note that this changed the behaviour of {{org.apache.cassandra.db.ColumnFamilyStore#getSnapshotSSTableReader}} to be \"offline\", however I think that makes sense in this case, as we shouldn't be regenerating summaries/BF's for snapshots anyway.\r\n\r\n|[3.0|https://github.com/apache/cassandra/compare/cassandra-3.0...kgreav:14166-3.0]|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...kgreav:14166-3.11]|[trunk|https://github.com/apache/cassandra/compare/trunk...kgreav:14166-trunk]|\r\n","from":"developer"},{"body":"Doesnt this now miss case where under normal (not offline) run and they change bf ratio it no longer recreates them? {{(validation.bloomFilterFPChance != metadata().params.bloomFilterFpChance)}}. With this it will just load the old bf so theres no way to change it.","from":"developer"},{"body":"Correct. As I noted previously \r\n\r\nbq. Only regenerate and persist the bloomfilter when it's missing - not when it has changed. This means we rely on compactions/upgradesstables to update the bloomfilter.\r\nbq. There's definitely no reason to regenerate Summaries in this case, and as previously mentioned it's not great regenerating the bloomfilter unless you're going to persist it. I have added persistence for the bloomfilter (when it is regenerated), however I think it's a bad idea to do this on startup as it will likely be more time consuming than regenerating the summaries.\r\n\r\nSo the previous behaviour was to regenerate the BF in this case but *not* persist it on the next startup (this meant it would happen on every startup until compactions/upgrades had occured). The summaries would be regenerated and persisted on the next startup (pointlessly). Both of these things would slow startup time pretty significantly depending on how much data you had.\r\n\r\nThe new behaviour would be to avoid regenerating BF/Summaries at all on startup and instead rely on upgradesstables/compactions to update them. Summaries would only be recreated when necessary (when not loaded/corrupt/missing).\r\n\r\nIn trunk it might make sense to also add a nodetool command that will allow us to regenerate the bloomfilters/summaries/etc without re-writing the whole data file.","from":"developer"},{"body":"gotcha +1 on patch from me\r\n\r\nI like idea of a {{nodetool recreatecomponent}} or something but that would be a different jira.","from":"developer"},{"body":"Thanks for the review [~cnlwsu]. Created CASSANDRA-14291 as a follow up.","from":"developer"},{"body":"||branch||testall||dtest||\r\n|[14166-3.0|https://github.com/kgreav/cassandra/tree/14166-3.0]|[testall|https://circleci.com/gh/kgreav/cassandra/tree/14166-3.0]|[dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/506/]|\r\n|[14166-3.11|https://github.com/kgreav/cassandra/tree/14166-3.11]|[testall|https://circleci.com/gh/kgreav/cassandra/tree/14166-3.11]|[dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/507]|\r\n|[14166-trunk|https://github.com/kgreav/cassandra/tree/14166-trunk]|[testall|https://circleci.com/gh/kgreav/cassandra/tree/14166-trunk]|[dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/508]|\r\n\r\nEDIT: i had [troubles|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/510/] getting the trunk patch dtests to run. As upstream trunk dtests appear to now be [working|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-trunk-dtest/462/], I, off a [thelastpickle|https://github.com/thelastpickle/cassandra/tree/14166-trunk] fork, rebased that patch off trunk and am running the dtests [again|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/511/]…","from":"developer"},{"body":"Rebased trunk patch run the dtests. \r\n\r\nI don't believe the failed test has anything to do with this ticket: [repair_tests.repair_test.TestRepair.test_dc_parallel_repair|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/511/testReport/junit/repair_tests.repair_test/TestRepair/test_dc_parallel_repair/]","from":"developer"},{"body":"Committed.","from":"developer"},{"body":"This commit broke 3.0 unit tests (maybe 3.11 too, haven't checked).\r\n\r\nSee https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-3.0-test-all/lastCompletedBuild/testReport/org.apache.cassandra.io.sstable/SSTableReaderTest/testOpeningSSTable/history/ and https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-3.0-test-all/248/changes. It's also failing locally and I'm pretty sure will be failing [here|https://circleci.com/gh/iamaleksey/cassandra/234] too.\r\n\r\nWe should be more careful about committing things. But in the meantime, [~mck] can you please fix this or revert this until there is a fix? Cheers.","from":"developer"},{"body":"thought this was fixed in https://issues.apache.org/jira/browse/CASSANDRA-14387 ? time resolution being different on causing it to fail on some systems.","from":"developer"},{"body":"Yeah. As you (and Marcus, offline) mentioned, CASSANDRA-14387 was only committed to trunk, so 3.0 tests are still failing elsewhere (incl. Jenkins, but not Circle). We just need to commit 14387 to 3.0 and 3.11.","from":"developer"},{"body":"{quote}Cherry-plicked into 3.0 as [e16f0ed0698c5cb47ab2bb0a0b04966d5bdbcde0|https://github.com/apache/cassandra/commit/e16f0ed0698c5cb47ab2bb0a0b04966d5bdbcde0] and merged upwards\r\n{quote}\r\nThanks [~iamaleksey] for fixing this.\r\n{quote}We should be more careful about committing things.\r\n{quote}\r\nYes, I should have noticed that {{SSTableReaderTest}} was a new failure in  [https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-3.0-test-all/248/#showFailuresLink] \r\n\r\nUp until now I've only been using circleci unit test reports, as illustrated in the \"branch|testall|dtest\" comment above, which reported it all green.\r\n\r\nAs CircleCI hasn't failed on this test, what's the difference here? \r\nAleksey's circleci build #234 [passed SSTableReaderTest|https://circleci.com/gh/iamaleksey/cassandra/234#tests/containers/2] , as did Kurt's [here|https://circleci.com/gh/kgreav/cassandra/116#tests/containers/1]\r\n\r\n \r\n\r\nIf CircleCi and Jenkins can both identify failures separately then \"test-all\" column should be split into two, to link to each.\r\n\r\nOtherwise, does Jenkins have any form of alerting or reporting on tests that go from stable to flakey? Would have been ideal to have automatically have receive the breakage information before it affected you [~iamaleksey]","from":"developer"},{"body":"the granularity of the file timestamp depends on the filesystem","from":"developer"},{"body":"[~cnlwsu], oh! not os but filesystem, and circleci and jenkins differ there. Ouch. \r\nHow should we test against that?","from":"developer"},{"body":"Hm, sorry about that. Originally started with second delays but changed it to speed up the test. Thought ms would be safe in this day and age. 2018 and all that.\r\n\r\n{quote}How should we test against that?\r\n{quote}\r\nWe don't, just have to go with the lowest resolution which will hopefully be seconds. No point putting workarounds in to speed up a test by a few seconds on _some_ platforms. Having said that, what filesystem is jenkins running? Seems odd it'd only have second precision. But then again it's probably an ancient server.","from":"developer"},{"body":"There isn't a clean or standard way to determine if a file system supports sub-second date/time resolution for file modifications.","from":"developer"},{"body":"bq. Yes, I should have noticed that SSTableReaderTest was a new failure in https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-3.0-test-all/248/#showFailuresLink \r\n\r\nRealistically, you shouldn't have. Looking at Circle should be sufficient enough. I just assumed that it would break on Circle just like it broke in our internal CI and Jenkins, and that was wrong. My apologies for the somewhat passive-aggressive righteous tone.\r\n\r\nIf Circle is green, we should be free to commit, and we should be checking with ASF Jenkins from time to time, but it should not be required for every commit.","from":"developer"}],"created":"2016-02-12T00:08:41.000+0000","description":"This is from trunk, but I also saw this happen on 2.0:\n\nBefore:\n{noformat}\nroot@bw-1:/srv/cassandra# ls -ltr /var/lib/cassandra/data/keyspace1/standard1-071efdc0d11811e590c3413ee28a6c90/\ntotal 221460\ndrwxr-xr-x 2 root root 4096 Feb 11 23:34 backups\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-6-big-TOC.txt\n-rw-r--r-- 1 root root 26518 Feb 11 23:50 ma-6-big-Summary.db\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-6-big-Statistics.db\n-rw-r--r-- 1 root root 2607705 Feb 11 23:50 ma-6-big-Index.db\n-rw-r--r-- 1 root root 192440 Feb 11 23:50 ma-6-big-Filter.db\n-rw-r--r-- 1 root root 10 Feb 11 23:50 ma-6-big-Digest.crc32\n-rw-r--r-- 1 root root 35212125 Feb 11 23:50 ma-6-big-Data.db\n-rw-r--r-- 1 root root 2156 Feb 11 23:50 ma-6-big-CRC.db\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-7-big-TOC.txt\n-rw-r--r-- 1 root root 26518 Feb 11 23:50 ma-7-big-Summary.db\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-7-big-Statistics.db\n-rw-r--r-- 1 root root 2607614 Feb 11 23:50 ma-7-big-Index.db\n-rw-r--r-- 1 root root 192432 Feb 11 23:50 ma-7-big-Filter.db\n-rw-r--r-- 1 root root 9 Feb 11 23:50 ma-7-big-Digest.crc32\n-rw-r--r-- 1 root root 35190400 Feb 11 23:50 ma-7-big-Data.db\n-rw-r--r-- 1 root root 2152 Feb 11 23:50 ma-7-big-CRC.db\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-5-big-TOC.txt\n-rw-r--r-- 1 root root 104178 Feb 11 23:50 ma-5-big-Summary.db\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-5-big-Statistics.db\n-rw-r--r-- 1 root root 10289077 Feb 11 23:50 ma-5-big-Index.db\n-rw-r--r-- 1 root root 757384 Feb 11 23:50 ma-5-big-Filter.db\n-rw-r--r-- 1 root root 9 Feb 11 23:50 ma-5-big-Digest.crc32\n-rw-r--r-- 1 root root 139201355 Feb 11 23:50 ma-5-big-Data.db\n-rw-r--r-- 1 root root 8508 Feb 11 23:50 ma-5-big-CRC.db\nroot@bw-1:/srv/cassandra# md5sum /var/lib/cassandra/data/keyspace1/standard1-071efdc0d11811e590c3413ee28a6c90/ma-5-big-Summary.db\n5fca154fc790f7cfa37e8ad6d1c7552c\n{noformat}\n\nBF ratio changed, node restarted:\n{noformat}\nroot@bw-1:/srv/cassandra# ls -ltr /var/lib/cassandra/data/keyspace1/standard1-071efdc0d11811e590c3413ee28a6c90/\ntotal 242168\ndrwxr-xr-x 2 root root 4096 Feb 11 23:34 backups\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-6-big-TOC.txt\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-6-big-Statistics.db\n-rw-r--r-- 1 root root 2607705 Feb 11 23:50 ma-6-big-Index.db\n-rw-r--r-- 1 root root 192440 Feb 11 23:50 ma-6-big-Filter.db\n-rw-r--r-- 1 root root 10 Feb 11 23:50 ma-6-big-Digest.crc32\n-rw-r--r-- 1 root root 35212125 Feb 11 23:50 ma-6-big-Data.db\n-rw-r--r-- 1 root root 2156 Feb 11 23:50 ma-6-big-CRC.db\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-7-big-TOC.txt\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-7-big-Statistics.db\n-rw-r--r-- 1 root root 2607614 Feb 11 23:50 ma-7-big-Index.db\n-rw-r--r-- 1 root root 192432 Feb 11 23:50 ma-7-big-Filter.db\n-rw-r--r-- 1 root root 9 Feb 11 23:50 ma-7-big-Digest.crc32\n-rw-r--r-- 1 root root 35190400 Feb 11 23:50 ma-7-big-Data.db\n-rw-r--r-- 1 root root 2152 Feb 11 23:50 ma-7-big-CRC.db\n-rw-r--r-- 1 root root 80 Feb 11 23:50 ma-5-big-TOC.txt\n-rw-r--r-- 1 root root 10264 Feb 11 23:50 ma-5-big-Statistics.db\n-rw-r--r-- 1 root root 10289077 Feb 11 23:50 ma-5-big-Index.db\n-rw-r--r-- 1 root root 757384 Feb 11 23:50 ma-5-big-Filter.db\n-rw-r--r-- 1 root root 9 Feb 11 23:50 ma-5-big-Digest.crc32\n-rw-r--r-- 1 root root 139201355 Feb 11 23:50 ma-5-big-Data.db\n-rw-r--r-- 1 root root 8508 Feb 11 23:50 ma-5-big-CRC.db\n-rw-r--r-- 1 root root 80 Feb 12 00:03 ma-8-big-TOC.txt\n-rw-r--r-- 1 root root 14902 Feb 12 00:03 ma-8-big-Summary.db\n-rw-r--r-- 1 root root 10264 Feb 12 00:03 ma-8-big-Statistics.db\n-rw-r--r-- 1 root root 1458631 Feb 12 00:03 ma-8-big-Index.db\n-rw-r--r-- 1 root root 10808 Feb 12 00:03 ma-8-big-Filter.db\n-rw-r--r-- 1 root root 10 Feb 12 00:03 ma-8-big-Digest.crc32\n-rw-r--r-- 1 root root 19660275 Feb 12 00:03 ma-8-big-Data.db\n-rw-r--r-- 1 root root 1204 Feb 12 00:03 ma-8-big-CRC.db\n-rw-r--r-- 1 root root 26518 Feb 12 00:04 ma-7-big-Summary.db\n-rw-r--r-- 1 root root 26518 Feb 12 00:04 ma-6-big-Summary.db\n-rw-r--r-- 1 root root 104178 Feb 12 00:04 ma-5-big-Summary.db\nroot@bw-1:/srv/cassandra# md5sum /var/lib/cassandra/data/keyspace1/standard1-071efdc0d11811e590c3413ee28a6c90/ma-5-big-Summary.db \n5fca154fc790f7cfa37e8ad6d1c7552c \n{noformat}\n\nThis hurts startup time and appears to do nothing useful whatsoever.","issue_id":"12938653","key":"CASSANDRA-11163","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2018-03-12T08:36:41.000+0000","role":"fixed_distractor","summary":"Summaries are needlessly rebuilt when the BF FP ratio is changed"} {"case_id":"12939330","cluster":"DISTRACTOR-CASSANDRA-11172","comments":[{"body":"could you post logs?","created":"2016-02-15T18:42:00.690+0000"},{"body":"It's this: `INFO [CompactionExecutor:12750] 2016-02-16 17:47:48,642 LeveledManifest.java:415 - Adding high-level (L0) SSTableReader(path='/mnt/cassandra/data/youtube/youtube_videos-2d16275b7ff93269bea0aac894e1abaa/youtube-youtube_videos-ka-104968-Data.db') to candidates` repeating endlessly. That one's extra special this time because of the L0. I'll look in our aggregation server and try to get value from the time when it starts.","created":"2016-02-16T17:55:23.877+0000"},{"body":"Possible duplicate of CASSANDRA-10831 ?","created":"2016-02-16T18:00:56.039+0000"},{"body":"do you run incremental repair?","created":"2016-02-16T18:12:22.714+0000"},{"body":"Yes. It's after incremental repair that I'm seeing this. Next time around I'll check that the file listed doesn't exist before restart, but I think this is a duplicate.\n\nAlternately, though, I'm also seeing the gossip thread lockup at times after incremental repair without mentioning higher level sstables, so that might be a new ticket to file next time around.","created":"2016-02-16T18:14:54.142+0000"},{"body":"you should definitely upgrade\n\nbut I have never seen this happen, but LCS with inc repair was very broken before CASSANDRA-10831","created":"2016-02-16T19:50:19.833+0000"},{"body":"Seeing this on C* 3.3 after running a full repair on another node. {{nodetool repair --full -pr -j 4 my_keyspace_name}} I can provide whatever further details you'd like, obviously.","created":"2016-02-17T04:15:17.804+0000"},{"body":"[~_wsh] do you have the logs leading up to the breakage? The debug.log would be helpful\n\nDoes it happen every time for this node?","created":"2016-02-17T06:27:45.146+0000"},{"body":"greping the logs for ma-3663 would be helpful as well","created":"2016-02-17T06:31:09.301+0000"},{"body":"debug.log? I can go hunt through my compressed files but this quickly pushes the other logs out of rotation. I've watched it happen. Thousands upon thousands of lines. And it's happened for a bunch of different nodes. The solution has been to kill them and restart but it's temporary.\n\nI just checked--all 20 zipped system.logs are filled with 40,000+ lines like the ones in my sample file. I was going to upload some to you but they're just identical, including the earliest lines I have:\n{noformat}\nINFO [CompactionExecutor:35] 2016-02-17 01:59:03,111 LeveledManifest.java:438 - Adding high-level (L0) BigTableReader(path='/var/lib/cassandra/data/segmentation/domain_events_by_event_domain_time-e81d74f0cd3a11e5aad8e7b84e29e52f/ma-3663-big-Data.db') to candidates\nINFO [CompactionExecutor:37] 2016-02-17 01:59:03,112 LeveledManifest.java:438 - Adding high-level (L3) BigTableReader(path='/var/lib/cassandra/data/segmentation/times_by_event_domain_user-e6112a30cd3a11e5ba896547d15a24f6/ma-4586-big-Data.db') to candidates\nINFO [CompactionExecutor:35] 2016-02-17 01:59:03,276 LeveledManifest.java:438 - Adding high-level (L0) BigTableReader(path='/var/lib/cassandra/data/segmentation/domain_events_by_event_domain_time-e81d74f0cd3a11e5aad8e7b84e29e52f/ma-3663-big-Data.db') to candidates\nINFO [CompactionExecutor:37] 2016-02-17 01:59:03,276 LeveledManifest.java:438 - Adding high-level (L3) BigTableReader(path='/var/lib/cassandra/data/segmentation/times_by_event_domain_user-e6112a30cd3a11e5ba896547d15a24f6/ma-4586-big-Data.db') to candidates\n{noformat}","created":"2016-02-17T07:35:57.024+0000"},{"body":"I don't know if these are worth anything but, in case they are: {{nodetool tpstats}} before I shut down the node and both mixed (native) and only-Java thread dumps.","created":"2016-02-17T08:07:09.075+0000"},{"body":"The problem is (probably) that we have an sstable in the wrong level in the manifest.\n\nI have not been able to reproduce this locally but it would make sense (ie, the sstable would not get removed from the manifest and we could end up in an infinite loop), but I think we should see an exception before things explode like this, so it would be really helpful if anyone has logs an hour or so *before* this happens","created":"2016-02-17T08:45:09.157+0000"},{"body":"Just got hit with this today on 2.2.5.","created":"2016-02-20T21:46:01.403+0000"},{"body":"Think I found the issue, this happens if you do nodetool repair -full (when replacing repaired sstables there was an assumption that the sstables was unrepaired before, but that is not the case when you do -full repairs)\n\n||branch||testall||dtest||\n|[marcuse/11172|https://github.com/krummas/cassandra/tree/marcuse/11172]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-dtest]|\n|[marcuse/11172-3.0|https://github.com/krummas/cassandra/tree/marcuse/11172-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-3.0-dtest]|\n|[marcuse/11172-trunk|https://github.com/krummas/cassandra/tree/marcuse/11172-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-trunk-dtest]|\n","created":"2016-02-22T10:09:35.380+0000"},{"body":"and the problem is not the \"Adding high-level ...\" - it is the fact that we keep around sstables in the wrong compaction strategy instance..","created":"2016-02-22T10:11:12.202+0000"},{"body":"The code looks good, but it needs a regression test.","created":"2016-02-22T17:01:48.653+0000"},{"body":"I used [this|https://github.com/krummas/cassandra-dtest/commits/marcuse/11172] to reproduce locally, but I was a bit skeptical to committing as it takes quite a long time to run and the failure case is that it hangs. But I guess it is a good test to run anyway.","created":"2016-02-22T21:40:57.611+0000"},{"body":"[~krummas] Fix version should also include 2.1? ","created":"2016-02-22T23:40:55.712+0000"},{"body":"+1 vote for inclusion in 2.1","created":"2016-02-23T01:21:47.372+0000"},{"body":"[~kohlisankalp] we have no way of triggering this in 2.1, it only happens if we repair (and anticompact) an already repaired sstable, and we can't do that in 2.1 - the only way to trigger anticompaction in 2.1 is with \"-inc\" and then we only include non-repaired sstables. But I guess I can commit it to 2.1 for correctness etc","created":"2016-02-23T07:26:45.359+0000"},{"body":"Wouldn't it also hang with a 10x smaller stress size and {{sstable_size_in_mb}}?","created":"2016-02-23T07:48:28.056+0000"},{"body":"I will try","created":"2016-02-23T07:51:52.011+0000"},{"body":"pushed an updated dtest which finishes in 300s with the patch and never without the patch","created":"2016-02-23T09:50:06.318+0000"},{"body":"Thank you, +1 on the patch.","created":"2016-02-23T10:00:19.103+0000"},{"body":"Committed, thanks","created":"2016-02-23T10:18:55.552+0000"},{"body":"We are seeing this issue on multiple nodes and it brings down our cassandra cluster. Is there some way to mitigate the issue until 3.4 is out?","created":"2016-02-26T04:50:36.159+0000"},{"body":"yes, don't do {{-full}} repairs","created":"2016-02-26T06:25:39.554+0000"},{"body":"[~krummas] Ok I guess that is to prevent it from occurring but we've already done this. Is there a way to fix the issue afterwards e.g. would crub'ing the tables bring something?","created":"2016-02-26T09:21:13.051+0000"},{"body":"[~marco.zattoo] just restart the affected nodes and it is fixed","created":"2016-02-26T12:29:57.630+0000"},{"body":"[~krummas] I've restarted nodes many times and the issue just comes up again. Does this mean that it is not related to the fix you've done? The issue happens during normal compactions on just restarted node after some time (I think after it tries to do the first compaction) with the result that the system.log starts get spammed with the infinite amount to the messages.\n{code}\nINFO [CompactionExecutor:5] 2016-02-26 05:45:56,479 LeveledManifest.java:438 - Adding high-level (L3) BigTableReader(path='/var/lib/cassandra/data2/ham/raw_sessions-417de7c0bb4711e4972d05e7bd5b0c2f/ma-45159-big-Data.db') to candidates\n{code}","created":"2016-02-26T12:51:52.437+0000"},{"body":"Unfortunately I can't provide any exception trace from before the issue starts because all the 20 zip logs are already full with the log line above. I also started a {{nodetool scrub ham raw_sessions}} but I'm wondering whether this will do any good and there is lot's data so it might take a while until it hits the sstables causing the issue.","created":"2016-02-26T12:55:07.937+0000"},{"body":"no, it should fix itself on restart, you sure you are not running repairs at all?","created":"2016-02-26T12:55:19.181+0000"},{"body":"No I didn't start any repair after restarting the nodes. Actually on one node I just restarted and will try to catch the exception happening in the thread doing the compaction. I hope that helps to figure out the root cause.","created":"2016-02-26T13:09:49.718+0000"},{"body":"[~krummas] I've added nodetool tpstats output. Would you like to see some other output from nodetool?","created":"2016-02-26T13:14:20.217+0000"},{"body":"I've hit the bug again without doing any repair. Again the logs are so quickly filled with the message that I was unable to get any message prior the LCS message.\n{code}\nroot@cassandra1:/var/log/cassandra# nodetool compactionstats\npending tasks: 41\n- ham.raw_sessions: 41\n{code}","created":"2016-02-26T13:58:44.719+0000"},{"body":"Here is the exception what happens prior the log spamming starts\n{code}\nERROR [CompactionExecutor:5] 2016-02-26 14:05:26,622 CassandraDaemon.java:195 - Exception in thread Thread[CompactionExecutor:5,1,main]\njava.lang.AssertionError: null\n at org.apache.cassandra.db.rows.BufferCell.(BufferCell.java:49) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BufferCell.tombstone(BufferCell.java:88) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BufferCell.tombstone(BufferCell.java:83) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BufferCell.purge(BufferCell.java:175) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.ComplexColumnData.lambda$purge$107(ComplexColumnData.java:165) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.ComplexColumnData$$Lambda$126/379481423.apply(Unknown Source) ~[na:na]\n at org.apache.cassandra.utils.btree.BTree$FiltrationTracker.apply(BTree.java:650) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.utils.btree.BTree.transformAndFilter(BTree.java:693) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.utils.btree.BTree.transformAndFilter(BTree.java:668) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.ComplexColumnData.transformAndFilter(ComplexColumnData.java:170) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.ComplexColumnData.purge(ComplexColumnData.java:165) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.ComplexColumnData.purge(ComplexColumnData.java:43) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BTreeRow.lambda$purge$102(BTreeRow.java:333) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BTreeRow$$Lambda$125/1572342504.apply(Unknown Source) ~[na:na]\n at org.apache.cassandra.utils.btree.BTree$FiltrationTracker.apply(BTree.java:650) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.utils.btree.BTree.transformAndFilter(BTree.java:693) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.utils.btree.BTree.transformAndFilter(BTree.java:668) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BTreeRow.transformAndFilter(BTreeRow.java:338) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BTreeRow.purge(BTreeRow.java:333) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToRow(PurgeFunction.java:88) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.transform.BaseRows.hasNext(BaseRows.java:116) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.ColumnIndex$Builder.build(ColumnIndex.java:120) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.ColumnIndex.writeAndBuildIndex(ColumnIndex.java:57) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.append(BigTableWriter.java:153) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.io.sstable.SSTableRewriter.append(SSTableRewriter.java:118) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.writers.MaxSSTableSizeWriter.realAppend(MaxSSTableSizeWriter.java:74) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.writers.CompactionAwareWriter.append(CompactionAwareWriter.java:132) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:182) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:78) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionCandidate.run(CompactionManager.java:264) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_45]\n at java.util.concurrent.FutureTask.run(FutureTask.java:266) ~[na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) ~[na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_45]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_45]\n{code}","created":"2016-02-26T14:08:27.607+0000"},{"body":"If that is the cause, then it is another bug (I will have a look)\n\nUnless you have already fixed this, could you try an offline scrub before you start?","created":"2016-02-29T12:53:13.805+0000"},{"body":"I had to bring the cluster into a healthy state and deleted the whole keyspace which was affected. Node that to me it seemed that the problem somehow spread from first only occurring on one node and then spread to other nodes. Also I guess rate limiting the amount of how many times this log message gets logged would be nice in order to be able to debug issues. \n\nSince the deletion of the whole keyspace I haven't seen any issues.","created":"2016-03-01T08:25:05.615+0000"},{"body":"[~marco.zattoo] I think your problem will be fixed with CASSANDRA-11373 - the exception you see will abort the compaction and cause the issues mentioned there","created":"2016-03-23T14:22:25.168+0000"}],"conversations":[{"body":"Observed that after a large repair on LCS that sometimes the system will enter an infinite loop with vast amounts of logs lines recording, \"Adding high-level (L${LEVEL}) SSTableReader(path='${TABLE}') to candidates\"\n\nThis results in an outage of the node and eventual crashing. The log spam quickly rotates out possibly useful earlier debugging.","from":"reporter","subject":"Infinite loop bug adding high-level SSTableReader in compaction"},{"body":"could you post logs?","from":"developer"},{"body":"It's this: `INFO [CompactionExecutor:12750] 2016-02-16 17:47:48,642 LeveledManifest.java:415 - Adding high-level (L0) SSTableReader(path='/mnt/cassandra/data/youtube/youtube_videos-2d16275b7ff93269bea0aac894e1abaa/youtube-youtube_videos-ka-104968-Data.db') to candidates` repeating endlessly. That one's extra special this time because of the L0. I'll look in our aggregation server and try to get value from the time when it starts.","from":"developer"},{"body":"Possible duplicate of CASSANDRA-10831 ?","from":"developer"},{"body":"do you run incremental repair?","from":"developer"},{"body":"Yes. It's after incremental repair that I'm seeing this. Next time around I'll check that the file listed doesn't exist before restart, but I think this is a duplicate.\n\nAlternately, though, I'm also seeing the gossip thread lockup at times after incremental repair without mentioning higher level sstables, so that might be a new ticket to file next time around.","from":"developer"},{"body":"you should definitely upgrade\n\nbut I have never seen this happen, but LCS with inc repair was very broken before CASSANDRA-10831","from":"developer"},{"body":"Seeing this on C* 3.3 after running a full repair on another node. {{nodetool repair --full -pr -j 4 my_keyspace_name}} I can provide whatever further details you'd like, obviously.","from":"developer"},{"body":"[~_wsh] do you have the logs leading up to the breakage? The debug.log would be helpful\n\nDoes it happen every time for this node?","from":"developer"},{"body":"greping the logs for ma-3663 would be helpful as well","from":"developer"},{"body":"debug.log? I can go hunt through my compressed files but this quickly pushes the other logs out of rotation. I've watched it happen. Thousands upon thousands of lines. And it's happened for a bunch of different nodes. The solution has been to kill them and restart but it's temporary.\n\nI just checked--all 20 zipped system.logs are filled with 40,000+ lines like the ones in my sample file. I was going to upload some to you but they're just identical, including the earliest lines I have:\n{noformat}\nINFO [CompactionExecutor:35] 2016-02-17 01:59:03,111 LeveledManifest.java:438 - Adding high-level (L0) BigTableReader(path='/var/lib/cassandra/data/segmentation/domain_events_by_event_domain_time-e81d74f0cd3a11e5aad8e7b84e29e52f/ma-3663-big-Data.db') to candidates\nINFO [CompactionExecutor:37] 2016-02-17 01:59:03,112 LeveledManifest.java:438 - Adding high-level (L3) BigTableReader(path='/var/lib/cassandra/data/segmentation/times_by_event_domain_user-e6112a30cd3a11e5ba896547d15a24f6/ma-4586-big-Data.db') to candidates\nINFO [CompactionExecutor:35] 2016-02-17 01:59:03,276 LeveledManifest.java:438 - Adding high-level (L0) BigTableReader(path='/var/lib/cassandra/data/segmentation/domain_events_by_event_domain_time-e81d74f0cd3a11e5aad8e7b84e29e52f/ma-3663-big-Data.db') to candidates\nINFO [CompactionExecutor:37] 2016-02-17 01:59:03,276 LeveledManifest.java:438 - Adding high-level (L3) BigTableReader(path='/var/lib/cassandra/data/segmentation/times_by_event_domain_user-e6112a30cd3a11e5ba896547d15a24f6/ma-4586-big-Data.db') to candidates\n{noformat}","from":"developer"},{"body":"I don't know if these are worth anything but, in case they are: {{nodetool tpstats}} before I shut down the node and both mixed (native) and only-Java thread dumps.","from":"developer"},{"body":"The problem is (probably) that we have an sstable in the wrong level in the manifest.\n\nI have not been able to reproduce this locally but it would make sense (ie, the sstable would not get removed from the manifest and we could end up in an infinite loop), but I think we should see an exception before things explode like this, so it would be really helpful if anyone has logs an hour or so *before* this happens","from":"developer"},{"body":"Just got hit with this today on 2.2.5.","from":"developer"},{"body":"Think I found the issue, this happens if you do nodetool repair -full (when replacing repaired sstables there was an assumption that the sstables was unrepaired before, but that is not the case when you do -full repairs)\n\n||branch||testall||dtest||\n|[marcuse/11172|https://github.com/krummas/cassandra/tree/marcuse/11172]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-dtest]|\n|[marcuse/11172-3.0|https://github.com/krummas/cassandra/tree/marcuse/11172-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-3.0-dtest]|\n|[marcuse/11172-trunk|https://github.com/krummas/cassandra/tree/marcuse/11172-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-marcuse-11172-trunk-dtest]|\n","from":"developer"},{"body":"and the problem is not the \"Adding high-level ...\" - it is the fact that we keep around sstables in the wrong compaction strategy instance..","from":"developer"},{"body":"The code looks good, but it needs a regression test.","from":"developer"},{"body":"I used [this|https://github.com/krummas/cassandra-dtest/commits/marcuse/11172] to reproduce locally, but I was a bit skeptical to committing as it takes quite a long time to run and the failure case is that it hangs. But I guess it is a good test to run anyway.","from":"developer"},{"body":"[~krummas] Fix version should also include 2.1? ","from":"developer"},{"body":"+1 vote for inclusion in 2.1","from":"developer"},{"body":"[~kohlisankalp] we have no way of triggering this in 2.1, it only happens if we repair (and anticompact) an already repaired sstable, and we can't do that in 2.1 - the only way to trigger anticompaction in 2.1 is with \"-inc\" and then we only include non-repaired sstables. But I guess I can commit it to 2.1 for correctness etc","from":"developer"},{"body":"Wouldn't it also hang with a 10x smaller stress size and {{sstable_size_in_mb}}?","from":"developer"},{"body":"I will try","from":"developer"},{"body":"pushed an updated dtest which finishes in 300s with the patch and never without the patch","from":"developer"},{"body":"Thank you, +1 on the patch.","from":"developer"},{"body":"Committed, thanks","from":"developer"},{"body":"We are seeing this issue on multiple nodes and it brings down our cassandra cluster. Is there some way to mitigate the issue until 3.4 is out?","from":"developer"},{"body":"yes, don't do {{-full}} repairs","from":"developer"},{"body":"[~krummas] Ok I guess that is to prevent it from occurring but we've already done this. Is there a way to fix the issue afterwards e.g. would crub'ing the tables bring something?","from":"developer"},{"body":"[~marco.zattoo] just restart the affected nodes and it is fixed","from":"developer"},{"body":"[~krummas] I've restarted nodes many times and the issue just comes up again. Does this mean that it is not related to the fix you've done? The issue happens during normal compactions on just restarted node after some time (I think after it tries to do the first compaction) with the result that the system.log starts get spammed with the infinite amount to the messages.\n{code}\nINFO [CompactionExecutor:5] 2016-02-26 05:45:56,479 LeveledManifest.java:438 - Adding high-level (L3) BigTableReader(path='/var/lib/cassandra/data2/ham/raw_sessions-417de7c0bb4711e4972d05e7bd5b0c2f/ma-45159-big-Data.db') to candidates\n{code}","from":"developer"},{"body":"Unfortunately I can't provide any exception trace from before the issue starts because all the 20 zip logs are already full with the log line above. I also started a {{nodetool scrub ham raw_sessions}} but I'm wondering whether this will do any good and there is lot's data so it might take a while until it hits the sstables causing the issue.","from":"developer"},{"body":"no, it should fix itself on restart, you sure you are not running repairs at all?","from":"developer"},{"body":"No I didn't start any repair after restarting the nodes. Actually on one node I just restarted and will try to catch the exception happening in the thread doing the compaction. I hope that helps to figure out the root cause.","from":"developer"},{"body":"[~krummas] I've added nodetool tpstats output. Would you like to see some other output from nodetool?","from":"developer"},{"body":"I've hit the bug again without doing any repair. Again the logs are so quickly filled with the message that I was unable to get any message prior the LCS message.\n{code}\nroot@cassandra1:/var/log/cassandra# nodetool compactionstats\npending tasks: 41\n- ham.raw_sessions: 41\n{code}","from":"developer"},{"body":"Here is the exception what happens prior the log spamming starts\n{code}\nERROR [CompactionExecutor:5] 2016-02-26 14:05:26,622 CassandraDaemon.java:195 - Exception in thread Thread[CompactionExecutor:5,1,main]\njava.lang.AssertionError: null\n at org.apache.cassandra.db.rows.BufferCell.(BufferCell.java:49) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BufferCell.tombstone(BufferCell.java:88) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BufferCell.tombstone(BufferCell.java:83) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BufferCell.purge(BufferCell.java:175) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.ComplexColumnData.lambda$purge$107(ComplexColumnData.java:165) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.ComplexColumnData$$Lambda$126/379481423.apply(Unknown Source) ~[na:na]\n at org.apache.cassandra.utils.btree.BTree$FiltrationTracker.apply(BTree.java:650) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.utils.btree.BTree.transformAndFilter(BTree.java:693) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.utils.btree.BTree.transformAndFilter(BTree.java:668) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.ComplexColumnData.transformAndFilter(ComplexColumnData.java:170) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.ComplexColumnData.purge(ComplexColumnData.java:165) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.ComplexColumnData.purge(ComplexColumnData.java:43) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BTreeRow.lambda$purge$102(BTreeRow.java:333) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BTreeRow$$Lambda$125/1572342504.apply(Unknown Source) ~[na:na]\n at org.apache.cassandra.utils.btree.BTree$FiltrationTracker.apply(BTree.java:650) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.utils.btree.BTree.transformAndFilter(BTree.java:693) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.utils.btree.BTree.transformAndFilter(BTree.java:668) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BTreeRow.transformAndFilter(BTreeRow.java:338) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.rows.BTreeRow.purge(BTreeRow.java:333) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToRow(PurgeFunction.java:88) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.transform.BaseRows.hasNext(BaseRows.java:116) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.ColumnIndex$Builder.build(ColumnIndex.java:120) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.ColumnIndex.writeAndBuildIndex(ColumnIndex.java:57) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.append(BigTableWriter.java:153) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.io.sstable.SSTableRewriter.append(SSTableRewriter.java:118) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.writers.MaxSSTableSizeWriter.realAppend(MaxSSTableSizeWriter.java:74) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.writers.CompactionAwareWriter.append(CompactionAwareWriter.java:132) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:182) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:78) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionCandidate.run(CompactionManager.java:264) ~[apache-cassandra-3.3.0.jar:3.3.0]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_45]\n at java.util.concurrent.FutureTask.run(FutureTask.java:266) ~[na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) ~[na:1.8.0_45]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_45]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_45]\n{code}","from":"developer"},{"body":"If that is the cause, then it is another bug (I will have a look)\n\nUnless you have already fixed this, could you try an offline scrub before you start?","from":"developer"},{"body":"I had to bring the cluster into a healthy state and deleted the whole keyspace which was affected. Node that to me it seemed that the problem somehow spread from first only occurring on one node and then spread to other nodes. Also I guess rate limiting the amount of how many times this log message gets logged would be nice in order to be able to debug issues. \n\nSince the deletion of the whole keyspace I haven't seen any issues.","from":"developer"},{"body":"[~marco.zattoo] I think your problem will be fixed with CASSANDRA-11373 - the exception you see will abort the compaction and cause the issues mentioned there","from":"developer"}],"created":"2016-02-15T17:43:47.000+0000","description":"Observed that after a large repair on LCS that sometimes the system will enter an infinite loop with vast amounts of logs lines recording, \"Adding high-level (L${LEVEL}) SSTableReader(path='${TABLE}') to candidates\"\n\nThis results in an outage of the node and eventual crashing. The log spam quickly rotates out possibly useful earlier debugging.","issue_id":"12939330","key":"CASSANDRA-11172","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-02-23T10:18:55.000+0000","role":"fixed_distractor","summary":"Infinite loop bug adding high-level SSTableReader in compaction"} {"case_id":"12949310","cluster":"DISTRACTOR-CASSANDRA-11345","comments":[{"body":"[~jfgosselin] do u remember if you were using sequential or parallel repairs when this issue showed up?","created":"2016-04-12T14:12:34.025+0000"},{"body":"During a sequential repair.","created":"2016-04-25T12:49:31.793+0000"},{"body":"I get this error when trying to boostrap a node. I had a [stackoverflow question|http://stackoverflow.com/questions/36719592/cant-add-a-new-cassandra-datacenter-due-to-streaming-errors/] opened (thought it was a DSE issue) but didn't think about looking at the out node logs.\n\nI tried bootstraping and rebuilding the node. Neither ever finish successfully.","created":"2016-05-10T08:28:35.278+0000"},{"body":"Also happening in 2.1.14","created":"2016-05-10T14:12:19.877+0000"},{"body":"[~vineus] What compaction strategy are you using? what is the average sstable size in the sending node?\n\nI suspect the following race is happening:\n- Node A sends SSTable X to bootstrapping node B\n- SSTable X is obsoleted by compaction during streaming\n- Node A finishes sending SSTable X to B and starts [calculating header size|https://github.com/apache/cassandra/blob/21448c50891642f95097a9e5ed0a3802bd90a877/src/java/org/apache/cassandra/streaming/StreamSession.java#L563] to update metrics (this became more expensive for compressed tables after CASSANDRA-10680)\n- Meanwhile, node B finishes receiving SSTable X and sends \"received\" message to node A\n- Node A receives \"received\" message for SSTable X and [releases its reference|https://github.com/apache/cassandra/blob/21448c50891642f95097a9e5ed0a3802bd90a877/src/java/org/apache/cassandra/streaming/messages/OutgoingFileMessage.java#L99], causing it to be garbage collected since it was obsoleted\n- Since {{CompressionMetadata}} was released, [header size calculation|https://github.com/apache/cassandra/blob/21448c50891642f95097a9e5ed0a3802bd90a877/src/java/org/apache/cassandra/streaming/messages/FileMessageHeader.java#L133] fails with {{Memory was freed}}\n\nDo you think this is plausible [~yukim] ? \n\nSimple solution would be to cache header size on first calculation, but we need to confirm first this is really what is happening because it could be hiding some other issue.\n\n[~vineus] Can you [set log level|https://docs.datastax.com/en/cassandra/2.1/cassandra/tools/toolsSetLogLev.html] of package {{org.apache.cassandra.streaming}} on sending on receiving nodes and attach system logs after error happens again?","created":"2016-05-10T21:34:28.667+0000"},{"body":"SSTable X in node B is added (activated) when streaming all SSTables in the same table finishes.\nSo I guess compaction of SSTable X shouldn't be happening during streaming.","created":"2016-05-10T22:13:24.816+0000"},{"body":"[~yukim] the \"memory freed\" problem happens on Node A, so obsoleting compaction happens on node A while it's streaming. Is this possible?","created":"2016-05-10T22:16:41.595+0000"},{"body":"My bad, and yes, it seems possible from the code.","created":"2016-05-10T22:26:49.156+0000"},{"body":"We are using the default {{SizeTieredCompactionStrategy}} for all tables\nThe average SSTable ({{-Data.db}} files) size of the sending nodes varies between a few MB to 40GB\n\nAfter turning the DEBUG loglevel on for {{org.apache.cassandra.streaming}} on a sending node, bootstrap a new node, and wait for the error, this is what I got next to the error:\n\n(I removed the file name from the log, said file still exist)\n\n{noformat}\nDEBUG [STREAM-OUT-/172.31.45.28] 2016-05-11 13:28:49,947 CompressedStreamWriter.java:90 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Finished streaming file /raid0/cassandra/data/SSTABLEFILE-Data.db to /172.31.45.28, bytesTransferred = 5854462750, totalSize = 5854462750\nDEBUG [STREAM-IN-/172.31.45.28] 2016-05-11 13:28:49,947 ConnectionHandler.java:106 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Closing stream connection handler on /172.31.45.28\nINFO [STREAM-IN-/172.31.45.28] 2016-05-11 13:28:49,947 StreamResultFuture.java:180 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Session with /172.31.45.28 is complete\nWARN [STREAM-IN-/172.31.45.28] 2016-05-11 13:28:49,948 StreamResultFuture.java:207 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Stream failed\nERROR [STREAM-OUT-/172.31.45.28] 2016-05-11 13:28:49,949 StreamSession.java:505 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Streaming error occurred\njava.lang.AssertionError: Memory was freed\n at org.apache.cassandra.io.util.SafeMemory.checkBounds(SafeMemory.java:97) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.io.util.Memory.getLong(Memory.java:249) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.io.compress.CompressionMetadata.getTotalSizeForSections(CompressionMetadata.java:247) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.messages.FileMessageHeader.size(FileMessageHeader.java:112) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.StreamSession.fileSent(StreamSession.java:546) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.messages.OutgoingFileMessage$1.serialize(OutgoingFileMessage.java:50) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.messages.OutgoingFileMessage$1.serialize(OutgoingFileMessage.java:41) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.messages.StreamMessage.serialize(StreamMessage.java:45) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.sendMessage(ConnectionHandler.java:358) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.run(ConnectionHandler.java:338) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_80]\n{noformat}","created":"2016-05-11T10:30:05.546+0000"},{"body":"Can you post the trace a few lines before that to see what were the last messages received/processed on [STREAM-IN-/172.31.45.28] ?\nAlso, can you check {{/raid0/cassandra/data/SSTABLEFILE-Data.db}} was logged before as participating in any compaction or other operation? Also can you double check the file still exists?\nThanks","created":"2016-05-11T15:29:34.498+0000"},{"body":"Ok, I'm sorry, I completely missed this, there was a timeout error prior to the \"Memory was freed\" error for this particular stream session:\n\n{noformat}\nERROR [STREAM-IN-/172.31.45.28] 2016-05-11 13:10:43,842 StreamSession.java:505 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Streaming error occurred\njava.net.SocketTimeoutException: null\n at sun.nio.ch.SocketAdaptor$SocketInputStream.read(SocketAdaptor.java:229) ~[na:1.7.0_80]\n at sun.nio.ch.ChannelInputStream.read(ChannelInputStream.java:103) ~[na:1.7.0_80]\n at java.nio.channels.Channels$ReadableByteChannelImpl.read(Channels.java:385) ~[na:1.7.0_80]\n at org.apache.cassandra.streaming.messages.StreamMessage.deserialize(StreamMessage.java:51) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.ConnectionHandler$IncomingMessageHandler.run(ConnectionHandler.java:257) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_80]\n{noformat}\n\nThis particular file is mentioned only once in the log prior to the \"Finished streaming\" message, when is started to be streamed (with apparently the same size):\n\n{noformat}\nDEBUG [STREAM-OUT-/172.31.45.28] 2016-05-11 12:10:23,960 CompressedStreamWriter.java:60 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Start streaming file /raid0/cassandra/data/SSTABLEFILE-Data.db to /172.31.45.28, repairedAt = 0, totalSize = 5854462750\n{noformat}\n\nThere are compactions for sstables of the same keyspace/table but not for this particular file.\n\nIt still exist, weighting 52595900317 bytes and was last modified on Apr 16","created":"2016-05-11T16:00:27.313+0000"},{"body":"[~vineus] It seems you are seeing the effects CASSANDRA-8343. The problem is that it takes more than {{streaming_socket_timeout_in_ms}} (default=1h) to transfer this 50GB file, so the sending node does not receive any message from the receiving node in the incoming socket and times out the stream session. As a temporary workaround you should increase this property to a larger value (10h or 20h) in sending nodes.\n\nIn any case we should probably keep this ticket open to prevent this memory was freed error from happening when the stream session fails for some other reason.","created":"2016-05-11T18:15:38.327+0000"},{"body":"bq. The problem is that it takes more than streaming_socket_timeout_in_ms (default=1h) to transfer this 50GB file, so the sending node does not receive any message from the receiving node in the incoming socket and times out the stream session.\nIs there a reason we don't have a simple heartbeat mechanism on these streaming sessions to prevent things like this?","created":"2016-05-11T19:39:58.735+0000"},{"body":"No reason not to, this is going to be provided by CASSANDRA-8343. We haven't noticed this much before due to CASSANDRA-11286.","created":"2016-05-11T21:59:05.725+0000"},{"body":"Thank! you [~pauloricardomg]. It was indeed the issue, setting {streaming_socket_timeout_in_ms} to a high value (72000000ms or 20h in my case) solved the issue! Will answer my stackoverflow question in the hope it helps others.","created":"2016-05-12T12:04:09.465+0000"},{"body":"While the root cause of the reported issue is CASSANDRA-11840, as a consequence of that the stream session failed while an sstable was being transferred, making its reference be released by {{StreamTransferTask.abort()}} -> {{OutgoingFileMessage.complete()}} before the transfer was complete, potentially causing the subsequent {{Memory was freed}} error if the sstable is garbage collected.\n\nThe ideal solution is CASSANDRA-11956 (to interrupt the stream transfer task right away), but while that is not in place we should release the sstable reference only after the ongoing transfer is finished (even if the session is aborted). So the idea of the patch is to set a {{transferring}} flag on {{OutgoingFileMessage}} while the file is being transferred, and if the session is failed before that the reference is not released, only after the {{transferring}} flag is unset at the end of the file transfer. I added a unit test that verifies that an sstable is not unreferenced if it's being transferred and the stream session fails, only after the transfer is finished.\n\nI also added a ninja fix to cache the {{FileMessageHeader}} size to avoid paying a high cost of recalculation of the size of large compressed files after the file is transferred on [StreamSession.fileSent(FileMessageHeader)|https://github.com/apache/cassandra/blob/9c8ee4c73f4e4e8d5b7693d34a0d6d6397418e90/src/java/org/apache/cassandra/streaming/StreamSession.java#L574]\n\nPatch and tests available below:\n||2.1||2.2||3.0||3.7||trunk||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.1...pauloricardomg:2.1-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.0...pauloricardomg:3.0-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.7...pauloricardomg:3.7-11345]|[branch|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:trunk-11345]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.7-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.7-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-dtest/lastCompletedBuild/testReport/]|\n\nWould you mind reviewing [~yukim]? Thanks in advance!\n\ncommit info: minor conflicts all the way up to 3.7 that merges cleanly to trunk","created":"2016-06-03T22:31:47.805+0000"},{"body":"[StreamTransferTaskTest is failing.|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-testall/lastCompletedBuild/testReport/org.apache.cassandra.streaming/StreamTransferTaskTest/testScheduleTimeout/]\n\nI think this is due to [opposite check here|https://github.com/pauloricardomg/cassandra/blob/d2784f0bab9b629aead1625d28955cc513be2e79/src/java/org/apache/cassandra/streaming/messages/OutgoingFileMessage.java#L124]. Shouldn't it be {{if (transferring)}}?\n","created":"2016-06-07T16:13:10.362+0000"},{"body":"It was just a test cleanup problem (since before there was only one test in the suite), so I added a tearDown step, rebased and resubmitted tests:\n\n||2.1||2.2||3.0||trunk||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.1...pauloricardomg:2.1-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.0...pauloricardomg:3.0-11345]|[branch|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:trunk-11345]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-dtest/lastCompletedBuild/testReport/]|","created":"2016-06-21T16:06:37.581+0000"},{"body":"Thanks for update. One more thing I want to discuss is that I think we can calculate {{size()}} at constructor and cache it to {{final long size}}. {{size()}} can be called at anytime from several thread so it seems safer that way.","created":"2016-07-07T16:23:04.497+0000"},{"body":"Thanks for the review. I updated the patch to cache the transfer size during {{FileMessageHeader}} construction.\n\nSince this is mostly a consequence CASSANDRA-11839, which was mitigated by CASSANDRA-11840 on 2.1, I think we can commit this only to 2.2+. WDYT?\n\nNew patch and CI result available below:\n\n||2.1||2.2||3.0||3.9||trunk||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.1...pauloricardomg:2.1-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.0...pauloricardomg:3.0-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.9...pauloricardomg:3.9-11345]|[branch|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:trunk-11345]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.9-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.9-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-dtest/lastCompletedBuild/testReport/]|\n","created":"2016-07-22T04:12:37.581+0000"},{"body":"3.9/trunk branch do not compile newly added unit test. (see github comment.)\nI wonder why tests were running on CI though.","created":"2016-08-04T01:41:16.799+0000"},{"body":"Fixed nit, CI results look good now. I think tests were from the previous rebase or dtests do not compile unit tests (where the compilation error was).","created":"2016-08-04T13:03:40.284+0000"},{"body":"Thanks, committed to 2.2+ as {{03985212644112d2751cdabc72bd954dda9ff3ba}}.","created":"2016-08-04T14:19:30.262+0000"},{"body":"if a file starts transferring right after a stream session is aborted, it may cause the sstable ref to be released twice (once when the session is aborted, another when the file finishes transfer), causing the following exception (reproduce in this [multiplexer run|https://cassci.datastax.com/view/Parameterized/job/parameterized_dtest_multiplexer/244/] for CASSANDRA-10810:\n{noformat}\nERROR [STREAM-OUT-/127.0.0.3] 2016-08-12 05:36:20,278 Ref.java:196 - BAD RELEASE: attempted to release a reference (org.apache.cassandra.utils.concurrent.Ref$State@62ddd804) that has already been released\n{noformat}\n\nSimple fix is to fail-fast if start transferring an already completed {{OutgoingFileMessage}}. I submitted [another multiplexer run|https://cassci.datastax.com/view/Parameterized/job/parameterized_dtest_multiplexer/247/] with [this fix|https://github.com/pauloricardomg/cassandra/commit/a932d66cf68046c732103d88c5b5f150f8d58553] and it passed.\n\n[~yukim] coul you have a quick look and if it looks good ninja this fix as a complement to the original commit? Thanks!","created":"2016-08-15T11:01:45.094+0000"},{"body":"Thanks, committed follow up as {{88b3cfc3b6b4417a913a512b461897378b48c039}}.","created":"2016-08-15T16:51:57.609+0000"}],"conversations":[{"body":"We encountered the following AssertionError (twice on the same node) during a repair :\n\nOn node /172.16.63.41\n\n{noformat}\nINFO [STREAM-IN-/10.174.216.160] 2016-03-09 02:38:13,900 StreamResultFuture.java:180 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Session with /10.174.216.160 is complete \nWARN [STREAM-IN-/10.174.216.160] 2016-03-09 02:38:13,900 StreamResultFuture.java:207 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Stream failed \nERROR [STREAM-OUT-/10.174.216.160] 2016-03-09 02:38:13,906 StreamSession.java:505 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Streaming error occurred \njava.lang.AssertionError: Memory was freed \n at org.apache.cassandra.io.util.SafeMemory.checkBounds(SafeMemory.java:97) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.io.util.Memory.getLong(Memory.java:249) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.io.compress.CompressionMetadata.getTotalSizeForSections(CompressionMetadata.java:247) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.FileMessageHeader.size(FileMessageHeader.java:112) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.StreamSession.fileSent(StreamSession.java:546) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.OutgoingFileMessage$1.serialize(OutgoingFileMessage.java:50) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.OutgoingFileMessage$1.serialize(OutgoingFileMessage.java:41) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.StreamMessage.serialize(StreamMessage.java:45) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.sendMessage(ConnectionHandler.java:351) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.run(ConnectionHandler.java:331) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_65] \n{noformat} \n\nOn node /10.174.216.160\n \n{noformat} \nERROR [STREAM-OUT-/172.16.63.41] 2016-03-09 02:38:14,140 StreamSession.java:505 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Streaming error occurred \njava.io.IOException: Connection reset by peer \n at sun.nio.ch.FileDispatcherImpl.write0(Native Method) ~[na:1.7.0_65] \n at sun.nio.ch.SocketDispatcher.write(SocketDispatcher.java:47) ~[na:1.7.0_65] \n at sun.nio.ch.IOUtil.writeFromNativeBuffer(IOUtil.java:93) ~[na:1.7.0_65] \n at sun.nio.ch.IOUtil.write(IOUtil.java:65) ~[na:1.7.0_65] \n at sun.nio.ch.SocketChannelImpl.write(SocketChannelImpl.java:487) ~[na:1.7.0_65] \n at org.apache.cassandra.io.util.DataOutputStreamAndChannel.write(DataOutputStreamAndChannel.java:48) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.StreamMessage.serialize(StreamMessage.java:44) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.sendMessage(ConnectionHandler.java:351) [apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.run(ConnectionHandler.java:323) [apache-cassandra-2.1.13.jar:2.1.13] \n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_65] \nINFO [STREAM-IN-/172.16.63.41] 2016-03-09 02:38:14,142 StreamResultFuture.java:180 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Session with /172.16.63.41 is complete\nWARN [STREAM-IN-/172.16.63.41] 2016-03-09 02:38:14,142 StreamResultFuture.java:207 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Stream failed \nERROR [STREAM-OUT-/172.16.63.41] 2016-03-09 02:38:14,143 StreamSession.java:505 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Streaming error occurred \njava.io.IOException: Broken pipe \n at sun.nio.ch.FileDispatcherImpl.write0(Native Method) ~[na:1.7.0_65] \n at sun.nio.ch.SocketDispatcher.write(SocketDispatcher.java:47) ~[na:1.7.0_65] \n at sun.nio.ch.IOUtil.writeFromNativeBuffer(IOUtil.java:93) ~[na:1.7.0_65] \n at sun.nio.ch.IOUtil.write(IOUtil.java:65) ~[na:1.7.0_65] \n at sun.nio.ch.SocketChannelImpl.write(SocketChannelImpl.java:487) ~[na:1.7.0_65] \n at org.apache.cassandra.io.util.DataOutputStreamAndChannel.write(DataOutputStreamAndChannel.java:48) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.StreamMessage.serialize(StreamMessage.java:44) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.sendMessage(ConnectionHandler.java:351) [apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.run(ConnectionHandler.java:331) [apache-cassandra-2.1.13.jar:2.1.13] \n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_65] \n{noformat}","from":"reporter","subject":"Assertion Errors \"Memory was freed\" during streaming"},{"body":"[~jfgosselin] do u remember if you were using sequential or parallel repairs when this issue showed up?","from":"developer"},{"body":"During a sequential repair.","from":"developer"},{"body":"I get this error when trying to boostrap a node. I had a [stackoverflow question|http://stackoverflow.com/questions/36719592/cant-add-a-new-cassandra-datacenter-due-to-streaming-errors/] opened (thought it was a DSE issue) but didn't think about looking at the out node logs.\n\nI tried bootstraping and rebuilding the node. Neither ever finish successfully.","from":"developer"},{"body":"Also happening in 2.1.14","from":"developer"},{"body":"[~vineus] What compaction strategy are you using? what is the average sstable size in the sending node?\n\nI suspect the following race is happening:\n- Node A sends SSTable X to bootstrapping node B\n- SSTable X is obsoleted by compaction during streaming\n- Node A finishes sending SSTable X to B and starts [calculating header size|https://github.com/apache/cassandra/blob/21448c50891642f95097a9e5ed0a3802bd90a877/src/java/org/apache/cassandra/streaming/StreamSession.java#L563] to update metrics (this became more expensive for compressed tables after CASSANDRA-10680)\n- Meanwhile, node B finishes receiving SSTable X and sends \"received\" message to node A\n- Node A receives \"received\" message for SSTable X and [releases its reference|https://github.com/apache/cassandra/blob/21448c50891642f95097a9e5ed0a3802bd90a877/src/java/org/apache/cassandra/streaming/messages/OutgoingFileMessage.java#L99], causing it to be garbage collected since it was obsoleted\n- Since {{CompressionMetadata}} was released, [header size calculation|https://github.com/apache/cassandra/blob/21448c50891642f95097a9e5ed0a3802bd90a877/src/java/org/apache/cassandra/streaming/messages/FileMessageHeader.java#L133] fails with {{Memory was freed}}\n\nDo you think this is plausible [~yukim] ? \n\nSimple solution would be to cache header size on first calculation, but we need to confirm first this is really what is happening because it could be hiding some other issue.\n\n[~vineus] Can you [set log level|https://docs.datastax.com/en/cassandra/2.1/cassandra/tools/toolsSetLogLev.html] of package {{org.apache.cassandra.streaming}} on sending on receiving nodes and attach system logs after error happens again?","from":"developer"},{"body":"SSTable X in node B is added (activated) when streaming all SSTables in the same table finishes.\nSo I guess compaction of SSTable X shouldn't be happening during streaming.","from":"developer"},{"body":"[~yukim] the \"memory freed\" problem happens on Node A, so obsoleting compaction happens on node A while it's streaming. Is this possible?","from":"developer"},{"body":"My bad, and yes, it seems possible from the code.","from":"developer"},{"body":"We are using the default {{SizeTieredCompactionStrategy}} for all tables\nThe average SSTable ({{-Data.db}} files) size of the sending nodes varies between a few MB to 40GB\n\nAfter turning the DEBUG loglevel on for {{org.apache.cassandra.streaming}} on a sending node, bootstrap a new node, and wait for the error, this is what I got next to the error:\n\n(I removed the file name from the log, said file still exist)\n\n{noformat}\nDEBUG [STREAM-OUT-/172.31.45.28] 2016-05-11 13:28:49,947 CompressedStreamWriter.java:90 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Finished streaming file /raid0/cassandra/data/SSTABLEFILE-Data.db to /172.31.45.28, bytesTransferred = 5854462750, totalSize = 5854462750\nDEBUG [STREAM-IN-/172.31.45.28] 2016-05-11 13:28:49,947 ConnectionHandler.java:106 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Closing stream connection handler on /172.31.45.28\nINFO [STREAM-IN-/172.31.45.28] 2016-05-11 13:28:49,947 StreamResultFuture.java:180 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Session with /172.31.45.28 is complete\nWARN [STREAM-IN-/172.31.45.28] 2016-05-11 13:28:49,948 StreamResultFuture.java:207 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Stream failed\nERROR [STREAM-OUT-/172.31.45.28] 2016-05-11 13:28:49,949 StreamSession.java:505 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Streaming error occurred\njava.lang.AssertionError: Memory was freed\n at org.apache.cassandra.io.util.SafeMemory.checkBounds(SafeMemory.java:97) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.io.util.Memory.getLong(Memory.java:249) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.io.compress.CompressionMetadata.getTotalSizeForSections(CompressionMetadata.java:247) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.messages.FileMessageHeader.size(FileMessageHeader.java:112) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.StreamSession.fileSent(StreamSession.java:546) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.messages.OutgoingFileMessage$1.serialize(OutgoingFileMessage.java:50) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.messages.OutgoingFileMessage$1.serialize(OutgoingFileMessage.java:41) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.messages.StreamMessage.serialize(StreamMessage.java:45) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.sendMessage(ConnectionHandler.java:358) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.run(ConnectionHandler.java:338) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_80]\n{noformat}","from":"developer"},{"body":"Can you post the trace a few lines before that to see what were the last messages received/processed on [STREAM-IN-/172.31.45.28] ?\nAlso, can you check {{/raid0/cassandra/data/SSTABLEFILE-Data.db}} was logged before as participating in any compaction or other operation? Also can you double check the file still exists?\nThanks","from":"developer"},{"body":"Ok, I'm sorry, I completely missed this, there was a timeout error prior to the \"Memory was freed\" error for this particular stream session:\n\n{noformat}\nERROR [STREAM-IN-/172.31.45.28] 2016-05-11 13:10:43,842 StreamSession.java:505 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Streaming error occurred\njava.net.SocketTimeoutException: null\n at sun.nio.ch.SocketAdaptor$SocketInputStream.read(SocketAdaptor.java:229) ~[na:1.7.0_80]\n at sun.nio.ch.ChannelInputStream.read(ChannelInputStream.java:103) ~[na:1.7.0_80]\n at java.nio.channels.Channels$ReadableByteChannelImpl.read(Channels.java:385) ~[na:1.7.0_80]\n at org.apache.cassandra.streaming.messages.StreamMessage.deserialize(StreamMessage.java:51) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at org.apache.cassandra.streaming.ConnectionHandler$IncomingMessageHandler.run(ConnectionHandler.java:257) ~[cassandra-all-2.1.14.1272.jar:2.1.14.1272]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_80]\n{noformat}\n\nThis particular file is mentioned only once in the log prior to the \"Finished streaming\" message, when is started to be streamed (with apparently the same size):\n\n{noformat}\nDEBUG [STREAM-OUT-/172.31.45.28] 2016-05-11 12:10:23,960 CompressedStreamWriter.java:60 - [Stream #ecfe0390-1763-11e6-b6c8-c1820b05e9ae] Start streaming file /raid0/cassandra/data/SSTABLEFILE-Data.db to /172.31.45.28, repairedAt = 0, totalSize = 5854462750\n{noformat}\n\nThere are compactions for sstables of the same keyspace/table but not for this particular file.\n\nIt still exist, weighting 52595900317 bytes and was last modified on Apr 16","from":"developer"},{"body":"[~vineus] It seems you are seeing the effects CASSANDRA-8343. The problem is that it takes more than {{streaming_socket_timeout_in_ms}} (default=1h) to transfer this 50GB file, so the sending node does not receive any message from the receiving node in the incoming socket and times out the stream session. As a temporary workaround you should increase this property to a larger value (10h or 20h) in sending nodes.\n\nIn any case we should probably keep this ticket open to prevent this memory was freed error from happening when the stream session fails for some other reason.","from":"developer"},{"body":"bq. The problem is that it takes more than streaming_socket_timeout_in_ms (default=1h) to transfer this 50GB file, so the sending node does not receive any message from the receiving node in the incoming socket and times out the stream session.\nIs there a reason we don't have a simple heartbeat mechanism on these streaming sessions to prevent things like this?","from":"developer"},{"body":"No reason not to, this is going to be provided by CASSANDRA-8343. We haven't noticed this much before due to CASSANDRA-11286.","from":"developer"},{"body":"Thank! you [~pauloricardomg]. It was indeed the issue, setting {streaming_socket_timeout_in_ms} to a high value (72000000ms or 20h in my case) solved the issue! Will answer my stackoverflow question in the hope it helps others.","from":"developer"},{"body":"While the root cause of the reported issue is CASSANDRA-11840, as a consequence of that the stream session failed while an sstable was being transferred, making its reference be released by {{StreamTransferTask.abort()}} -> {{OutgoingFileMessage.complete()}} before the transfer was complete, potentially causing the subsequent {{Memory was freed}} error if the sstable is garbage collected.\n\nThe ideal solution is CASSANDRA-11956 (to interrupt the stream transfer task right away), but while that is not in place we should release the sstable reference only after the ongoing transfer is finished (even if the session is aborted). So the idea of the patch is to set a {{transferring}} flag on {{OutgoingFileMessage}} while the file is being transferred, and if the session is failed before that the reference is not released, only after the {{transferring}} flag is unset at the end of the file transfer. I added a unit test that verifies that an sstable is not unreferenced if it's being transferred and the stream session fails, only after the transfer is finished.\n\nI also added a ninja fix to cache the {{FileMessageHeader}} size to avoid paying a high cost of recalculation of the size of large compressed files after the file is transferred on [StreamSession.fileSent(FileMessageHeader)|https://github.com/apache/cassandra/blob/9c8ee4c73f4e4e8d5b7693d34a0d6d6397418e90/src/java/org/apache/cassandra/streaming/StreamSession.java#L574]\n\nPatch and tests available below:\n||2.1||2.2||3.0||3.7||trunk||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.1...pauloricardomg:2.1-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.0...pauloricardomg:3.0-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.7...pauloricardomg:3.7-11345]|[branch|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:trunk-11345]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.7-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.7-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-dtest/lastCompletedBuild/testReport/]|\n\nWould you mind reviewing [~yukim]? Thanks in advance!\n\ncommit info: minor conflicts all the way up to 3.7 that merges cleanly to trunk","from":"developer"},{"body":"[StreamTransferTaskTest is failing.|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-testall/lastCompletedBuild/testReport/org.apache.cassandra.streaming/StreamTransferTaskTest/testScheduleTimeout/]\n\nI think this is due to [opposite check here|https://github.com/pauloricardomg/cassandra/blob/d2784f0bab9b629aead1625d28955cc513be2e79/src/java/org/apache/cassandra/streaming/messages/OutgoingFileMessage.java#L124]. Shouldn't it be {{if (transferring)}}?\n","from":"developer"},{"body":"It was just a test cleanup problem (since before there was only one test in the suite), so I added a tearDown step, rebased and resubmitted tests:\n\n||2.1||2.2||3.0||trunk||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.1...pauloricardomg:2.1-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.0...pauloricardomg:3.0-11345]|[branch|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:trunk-11345]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-dtest/lastCompletedBuild/testReport/]|","from":"developer"},{"body":"Thanks for update. One more thing I want to discuss is that I think we can calculate {{size()}} at constructor and cache it to {{final long size}}. {{size()}} can be called at anytime from several thread so it seems safer that way.","from":"developer"},{"body":"Thanks for the review. I updated the patch to cache the transfer size during {{FileMessageHeader}} construction.\n\nSince this is mostly a consequence CASSANDRA-11839, which was mitigated by CASSANDRA-11840 on 2.1, I think we can commit this only to 2.2+. WDYT?\n\nNew patch and CI result available below:\n\n||2.1||2.2||3.0||3.9||trunk||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.1...pauloricardomg:2.1-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.0...pauloricardomg:3.0-11345]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.9...pauloricardomg:3.9-11345]|[branch|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:trunk-11345]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.9-11345-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.1-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.9-11345-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-11345-dtest/lastCompletedBuild/testReport/]|\n","from":"developer"},{"body":"3.9/trunk branch do not compile newly added unit test. (see github comment.)\nI wonder why tests were running on CI though.","from":"developer"},{"body":"Fixed nit, CI results look good now. I think tests were from the previous rebase or dtests do not compile unit tests (where the compilation error was).","from":"developer"},{"body":"Thanks, committed to 2.2+ as {{03985212644112d2751cdabc72bd954dda9ff3ba}}.","from":"developer"},{"body":"if a file starts transferring right after a stream session is aborted, it may cause the sstable ref to be released twice (once when the session is aborted, another when the file finishes transfer), causing the following exception (reproduce in this [multiplexer run|https://cassci.datastax.com/view/Parameterized/job/parameterized_dtest_multiplexer/244/] for CASSANDRA-10810:\n{noformat}\nERROR [STREAM-OUT-/127.0.0.3] 2016-08-12 05:36:20,278 Ref.java:196 - BAD RELEASE: attempted to release a reference (org.apache.cassandra.utils.concurrent.Ref$State@62ddd804) that has already been released\n{noformat}\n\nSimple fix is to fail-fast if start transferring an already completed {{OutgoingFileMessage}}. I submitted [another multiplexer run|https://cassci.datastax.com/view/Parameterized/job/parameterized_dtest_multiplexer/247/] with [this fix|https://github.com/pauloricardomg/cassandra/commit/a932d66cf68046c732103d88c5b5f150f8d58553] and it passed.\n\n[~yukim] coul you have a quick look and if it looks good ninja this fix as a complement to the original commit? Thanks!","from":"developer"},{"body":"Thanks, committed follow up as {{88b3cfc3b6b4417a913a512b461897378b48c039}}.","from":"developer"}],"created":"2016-03-11T21:16:45.000+0000","description":"We encountered the following AssertionError (twice on the same node) during a repair :\n\nOn node /172.16.63.41\n\n{noformat}\nINFO [STREAM-IN-/10.174.216.160] 2016-03-09 02:38:13,900 StreamResultFuture.java:180 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Session with /10.174.216.160 is complete \nWARN [STREAM-IN-/10.174.216.160] 2016-03-09 02:38:13,900 StreamResultFuture.java:207 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Stream failed \nERROR [STREAM-OUT-/10.174.216.160] 2016-03-09 02:38:13,906 StreamSession.java:505 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Streaming error occurred \njava.lang.AssertionError: Memory was freed \n at org.apache.cassandra.io.util.SafeMemory.checkBounds(SafeMemory.java:97) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.io.util.Memory.getLong(Memory.java:249) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.io.compress.CompressionMetadata.getTotalSizeForSections(CompressionMetadata.java:247) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.FileMessageHeader.size(FileMessageHeader.java:112) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.StreamSession.fileSent(StreamSession.java:546) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.OutgoingFileMessage$1.serialize(OutgoingFileMessage.java:50) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.OutgoingFileMessage$1.serialize(OutgoingFileMessage.java:41) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.StreamMessage.serialize(StreamMessage.java:45) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.sendMessage(ConnectionHandler.java:351) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.run(ConnectionHandler.java:331) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_65] \n{noformat} \n\nOn node /10.174.216.160\n \n{noformat} \nERROR [STREAM-OUT-/172.16.63.41] 2016-03-09 02:38:14,140 StreamSession.java:505 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Streaming error occurred \njava.io.IOException: Connection reset by peer \n at sun.nio.ch.FileDispatcherImpl.write0(Native Method) ~[na:1.7.0_65] \n at sun.nio.ch.SocketDispatcher.write(SocketDispatcher.java:47) ~[na:1.7.0_65] \n at sun.nio.ch.IOUtil.writeFromNativeBuffer(IOUtil.java:93) ~[na:1.7.0_65] \n at sun.nio.ch.IOUtil.write(IOUtil.java:65) ~[na:1.7.0_65] \n at sun.nio.ch.SocketChannelImpl.write(SocketChannelImpl.java:487) ~[na:1.7.0_65] \n at org.apache.cassandra.io.util.DataOutputStreamAndChannel.write(DataOutputStreamAndChannel.java:48) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.StreamMessage.serialize(StreamMessage.java:44) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.sendMessage(ConnectionHandler.java:351) [apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.run(ConnectionHandler.java:323) [apache-cassandra-2.1.13.jar:2.1.13] \n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_65] \nINFO [STREAM-IN-/172.16.63.41] 2016-03-09 02:38:14,142 StreamResultFuture.java:180 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Session with /172.16.63.41 is complete\nWARN [STREAM-IN-/172.16.63.41] 2016-03-09 02:38:14,142 StreamResultFuture.java:207 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Stream failed \nERROR [STREAM-OUT-/172.16.63.41] 2016-03-09 02:38:14,143 StreamSession.java:505 - [Stream #f6980580-e55f-11e5-8f08-ef9e099ce99e] Streaming error occurred \njava.io.IOException: Broken pipe \n at sun.nio.ch.FileDispatcherImpl.write0(Native Method) ~[na:1.7.0_65] \n at sun.nio.ch.SocketDispatcher.write(SocketDispatcher.java:47) ~[na:1.7.0_65] \n at sun.nio.ch.IOUtil.writeFromNativeBuffer(IOUtil.java:93) ~[na:1.7.0_65] \n at sun.nio.ch.IOUtil.write(IOUtil.java:65) ~[na:1.7.0_65] \n at sun.nio.ch.SocketChannelImpl.write(SocketChannelImpl.java:487) ~[na:1.7.0_65] \n at org.apache.cassandra.io.util.DataOutputStreamAndChannel.write(DataOutputStreamAndChannel.java:48) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.messages.StreamMessage.serialize(StreamMessage.java:44) ~[apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.sendMessage(ConnectionHandler.java:351) [apache-cassandra-2.1.13.jar:2.1.13] \n at org.apache.cassandra.streaming.ConnectionHandler$OutgoingMessageHandler.run(ConnectionHandler.java:331) [apache-cassandra-2.1.13.jar:2.1.13] \n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_65] \n{noformat}","issue_id":"12949310","key":"CASSANDRA-11345","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-08-15T16:51:57.000+0000","role":"fixed_distractor","summary":"Assertion Errors \"Memory was freed\" during streaming"} {"case_id":"12949698","cluster":"DISTRACTOR-CASSANDRA-11349","comments":[{"body":"Looks like the {{MergeIterator.ManyToOne}} logic gets in the way of {{LazilyCompactedRow.Reducer}} doing it's job. The iterator will stop adding atoms to the reducer and continue to advance, once two range tombstones with different deletion times are about to be merged.","created":"2016-03-21T19:51:34.923+0000"},{"body":"\n\nI gave the patch some more thoughts and I'm now confident that the proposed change is the best way to address the issue. \n\nBasically what happens during validation compaction is that a scanner is created for each sstable. The {{CompactionIterable.Reducer}} will then create a {{LazilyCompactedRow}} with an iterable of {{OnDiskAtom}} for the same key in each sstable. The purpose of {{LazilyCompactedRow}} during validation compaction is to create a digest of the compacted version of all atoms that would represent a single row. This is done cell by cell, where each collection of atoms for a single cell name is consumed by {{LazilyCompactedRow.Reducer}}.\n \n The decision on whether {{LazilyCompactedRow.Reducer}} should finish to merge a cell and move to the next one is currently being done by {{AbstractCellNameType.onDiskAtomComparator}}, as evaluated by {{MergeIterator.ManyToOne}}. However, the comparator does not only compare by name, but also by {{DeletionTime}} in case of {{RangeTombstone}}. As a consequence, {{MergeIterator.ManyToOne}} will advance in case two {{RangeTombstone}} with different deletion times are read, which breaks the \"_will be called one or more times with cells that share the same column name_\" contract in {{LazilyCompactedRow.Reducer}}.\n\nThe submitted patch will introduce a new {{Comparator}} that will basically work like {{onDiskAtomComparator}}, but does not compare deletion time. As simple as that.\n\n\n||2.1||2.2||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.1]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.2]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-testall/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-dtest/]|\n\n\nThe only other places where {{LazilyCompactedRow}} is being used except validation compaction are the cleanup and scrub functions, which shouldn't be affected, as those are working on individual sstables and I assume that there's no case where an sstable can have multiple identical range tombstones with different timestamps.\n","created":"2016-03-23T12:37:21.567+0000"},{"body":"Nice patch.\n\nI will be able to test it on a dev environment either this week or at the beginning of next week.\n\nThere is still one case not covered (though not sure it can happen).\nSuppose that in SSTable 1, there is a range tombstone covering the columns \"a\" through \"g\" at time t1, and in SSTable 2, there is a range tombstone covering the columns \"c\" through \"d\" at time t2.\nIf those two SSTable are merged (for example on another replica), it will be split in 3 range tombstones in one SSTable (one range tombstone \"a\" -> \"b\" at t1, \"c\" -> \"d\" at t2, \"e\" -> \"f\" at t1).\nComputing the merkle tree for those two hosts will still be different (not the same range tombstones).\n\nAs said above, not sure if it can happen, and anyway, this patch is a good improvement and probably fits 99% cases.\n\n","created":"2016-03-23T17:45:11.048+0000"},{"body":"I tested the patch on a dev environment containing production data and there were still some differences.\n\nThe test procedure was:\n - use ccm & the branch linked to this ticket (I verified that the classpath is ok)\n - copy a fresh backup of production data\n - did a first full repair -> it had some differences on all CFs but this can be explained if there is a small delay when snapshotting all hosts for the backup (this cluster receive a few thousands writes per second)\n - did a second full repair -> only one CF had no differences (the one without range tombstone) and all others had differences (while they should not because there are no reads & writes on this dev environment)\n\nI will continue to investigate and try to isolate those differences...","created":"2016-03-30T21:50:04.844+0000"},{"body":"Thanks, the patch is OK.\n\nIn fact, the differences were produced by another bug that I will create separately.","created":"2016-04-01T15:39:34.738+0000"},{"body":"I suspect this doesn't affect 3.x: has someone checked, and if not, can someone do so we know if some 3.x version is needed or not for this?","created":"2016-04-01T15:56:40.126+0000"},{"body":"I tested against the 3.0.4 and it is not affected (not tested the 3.X, but assumed that it's not affected).\nThere is another similar ticket (the 3.0.4 is not affected): https://issues.apache.org/jira/browse/CASSANDRA-11477 ","created":"2016-04-01T16:22:32.232+0000"},{"body":"[~spodxx@gmail.com] [~slebresne] [~frousseau] I'm confused here. Why should repair be special cased over normal compaction in this case? If the times are different then you *do* still need to resolve it as you need to take the greater time.\n\nIt seems to me the crux of the current patch is to \"fix\" this by special casing the comparator to just compare just the max value of the interval during repair validation:\n\n{code}\n // only compare interval, but not deletion time\n+ return AbstractCellNameType.this.compare(((RangeTombstone)c1).max, ((RangeTombstone)c2).max);\n{code}\n\nI just did my best to merge and compare the code between 2.0 and 2.1 and I'm still trying to parse how this code is different in 2.0 vs. 2.1... We've been unable to reproduce this in 2.0 so far, but the bits of the code being touched here don't seem to be different so I'm trying to understand why 2.1 would hit this and not 2.0.\n\nCould you please explain a bit more why we we can ignore the timestamp?\n","created":"2016-04-01T20:26:16.647+0000"},{"body":"And just for my sanity and for discussion in the Jira, here is the current handling in the comparator\n\n{code}\nif (c1 instanceof RangeTombstone)\n{\n if (c2 instanceof RangeTombstone)\n {\n RangeTombstone t1 = (RangeTombstone)c1;\n RangeTombstone t2 = (RangeTombstone)c2;\n int comp2 = AbstractCellNameType.this.compare(t1.max, t2.max);\n return comp2 == 0 ? t1.data.compareTo(t2.data) : comp2;\n }\n else\n {\n return -1;\n }\n }\n else\n {\n return c2 instanceof RangeTombstone ? 1 : 0;\n }\n}\n{code}","created":"2016-04-01T20:30:51.270+0000"},{"body":"I'm also not sure how this is meant to fix it. Special casing validation compaction may fix repairs but you'd still get the digest mismatches on reads.","created":"2016-04-01T20:56:07.623+0000"},{"body":"I think there are a couple of things wrong with the current patch.\n\nFirst, the comparator needs to continue to compare the full cell name first, and only break ties on range tombstones with {{compare(t1.max, t2.max)}}.\n\nSecond, I believe we should remove the timestamp tie-breaking behavior from the comparator in general, and not just for validation compactions. In other words, I think we're doing the comparison incorrectly for all compactions right now.\n\nWe want the comparison to return 0 whenever range tombstones have equal names and ranges, even if they have different timestamps. This will result in {{LazilyCompactedRow.Reducer.reduce()}} being called in one round with each of the tombstones that only differ in timestamp. The logic in {{LCR.Reducer.reduce()}} already handles the case of multiple range tombstones with different timestamps by picking the one with the highest timestamp, so these will correctly be reduced to a single RT. It looks like the current codebase will keep both range tombstones during a compaction, which isn't necessarily harmful, but is suboptimal. For repair purposes, though, this is incorrect as it produces a different digest.\n\nTo summarize: I think all we need to do is remove the timestamp tie-breaking logic from the existing comparator.\n\n[~slebresne] should double-check my logic, though.","created":"2016-04-01T21:28:05.397+0000"},{"body":"It makes sense just to modify {{onDiskAtomComparator}}. Given the generic name I assumed the comparator is used in other places as well, but since its only used in {{LazyCompactedRow}} we can just change the patch as suggested and simply remove the timestamp tie-break behaviour in {{onDiskAtomComparator}}. \n\nAs for regular compactions, I agree with Tyler that this should not effect compactions in a way that it does with validation compaction. Before the patch, {{LazyCompactedRow}} would not reduce both RTs but instead have {{ColumnIndex.buildForCompaction()}} iterate over both RTs and have them added to the {{RangeTombstone.Tracker}}. The tracker would merge them in a way {{LCR.Reducer.getReduced}} would after the patch. However, I’m not fully sure if there could be some other for more complex cases where this still would cause problems.\n\nAlthough the patch should fix the described issue, the way we deal with RTs during validation compaction is still not ideal. The problem is that LCR lacks some handling of relationships between RTs compared to {{RangeTombstone.Tracker}}. If we create digests column by column, we get wrong results for shadowing tombstones not sharing the same intervals.\n\n{noformat}\nCREATE KEYSPACE test_rt WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 2};\nUSE test_rt;\nCREATE TABLE IF NOT EXISTS table1 (\n c1 text,\n c2 text,\n c3 text,\n c4 float,\n PRIMARY KEY (c1, c2, c3)\n) WITH compaction = {'class': 'SizeTieredCompactionStrategy', 'enabled': 'false'};\nDELETE FROM table1 WHERE c1 = 'a' AND c2 = 'b' AND c3 = 'c';\n\nccm node1 flush\n\nDELETE FROM table1 WHERE c1 = 'a' AND c2 = 'b';\n\nccm node1 repair test_rt table1\n{noformat}\n\n\nIn this case the (c1, c2, c3) RT will always be repaired after it has been compacted with (c1, c2) on any node. \nSo I’m wondering if we shouldn’t take a more bold approach here than the patch does. \n","created":"2016-04-02T14:14:43.172+0000"},{"body":"Using the RangeTombstone.Tracker can help in the situation described just above.\n\nIn fact, the RT should always update the tracker (see CASSANDRA-11477).\nThe trick here is to always considered it as \"expired\" in the tracker (even if not) so the tombstones are not accumulated during compaction (if expired the tracker keeps only the list of opened RTs and if not, it keeps all unwritten RTs, ie all RTs because it's a validation compaction...).\n\nHaving a look at the update method of the Tracker, it already check if the tombstone is superseded by another one (and don't add it as \"opened\" if superseded).\n\nThus, the v2 patch:\n - includes the previous patch\n - always update the tracker with the RT (considering it as expired even if not, just to not retain too many of them in memory, and because it's for validation, it's a read only and won't affect anything)\n - test if the RT was added in the openedTombstones list, and if that's not the case, skip it for digest.\n\nI know that the patch may be a bit rough (at least on the \"isLastOpened\" method) but it is more to validate the approach first and did not want the patch to be too invasive (by modifying the returned value of the update method).\n\nWDYT ?\n\nNote: I have not yet tested it against our production data\nNote2: Regarding the read-repair, this seems to be a different story and can't see anything for now that could explain those differences (will dig later on this as this is less urgent)","created":"2016-04-05T21:28:42.087+0000"},{"body":"[~frousseau], what makes things more complicated here is that changes to LCR will effect regular compactions as well. Adding all tombstones as expired in your {{11349-2.1-v2.patch}} will have unwanted side effects for regular compactions, e.g. try {{RangeTombstoneMergeTest}} with it.\n\nI've now spend some time trying to make use of the RT.Tracker there but without much success. Adding non-expired range tombstones to the tracker from within LCR would cause corrupted sstables. Even creating an edge case just for validation compaction would not handle all potential TS shadowing scenarios and will probably cause more harm than good (and potential digest mismatch storms). I'm not even sure it's possible given the current iterative MergeIterator > LazilyCompactedRow > RT.Tracker interaction. \n\nI'm now at a point where I'd suggest to just stick with {{11349-2.1.patch}} unless someone else has a better idea how to solve this. I've updated the [dtest PR|https://github.com/riptano/cassandra-dtest/pull/881] with two of the described shadowing scenarios that will only work with 3.0+ even after the patch, if someone wants to give it a try.\n\nCassci results for {{11349-2.1.patch}}:\n\n||2.1||2.2||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.1]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.2]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-testall/]|\n\n","created":"2016-04-07T11:52:27.136+0000"},{"body":"Can we keep the conversation going to get this patch into the next 2.x release?","created":"2016-04-22T08:10:32.984+0000"},{"body":"Sorry for not being reactive lately, I'm rather busy atm...\n\nI'd be more than happy[1] to see this patch in the next release.\nI haven't tested it yet and probably can find some time next week to test it on a dev cluster if it can help.\nNevertheless, I won't be able to tell if it really worked because there will still have some mismatches (due to CASSANDRA-11477).\n\nI have started working on a patch which should be able to handle both CASSANDRA-11477 and the last edge case.\n\nWhat it basically does:\n - Tracker is now an interface\n - there are two implementations: one called RegularCompactionTracker and another one ValidationCompactionTracker\n - the ColumnIndexer.Builder has one more optional parameter : a boolean to know if it is built for validation\n - the RegularCompactionTracker is identical to the existing Tracker + one empty method\n - the ValidationCompactionTracker is similar to the existing Tracker but retain only opened tombstones (most methods are thus empty)\n - the Reducer slightly changed but its behaviour is the same regarding the regular compactions\n\nI can share it if you're interested (code compiles but I still haven't tested it at all and plan to do it soon and share it after).\n\n[1] Just to share more information: those issues are important to us, because a few of our clusters are impacted and a few days after filing the bug, we decided to temporarily stop repairing some tables (knowing that we could live with inconsistencies on those tables) which were heavily impacted by those bugs (each repair increased disk occupancy by a few percent), and did a major compaction. This resulted in two to three times less disk occupancy (One table shrinked from 243GB to 79GB. Note that this was not due to tombstones reclaiming old data because, it's been nearly a month now, and the big SSTable resulting from the major compaction is still there but disk usage has not grown that much). ","created":"2016-04-22T20:58:12.504+0000"},{"body":"Ok, I uploaded a new version of the patch (11349-2.1-v2.patch)\n\nAs said above, there are two Tracker implementations now: one for regular compaction and another one for validation compaction.\nIt solves both cases described here (the one in the ticket + the one in the comment of Stefan) and CASSANDRA-11477.","created":"2016-04-28T13:57:23.137+0000"},{"body":"I'm not sure introducing a new tracker interface is the best way to handle this. It took me a while to actually figure out the differences between the {{update}} implementations in both trackers, since for most parts it's sharing the same copied code. It would probably be better to have ValidationTracker subclass RegularCompactionTracker, add {{remove/addUnwrittenTombstone}} implemented empty for validation.\n\nThe {{addRangeTombstone}} semantics also look like a case of leaky abstractions to me. It's adding nothing at all for regular compaction, but serves as early exit path for validation. \n\nGood news is that the dtests and unit tests seem to pass with the patch. :)","created":"2016-05-02T12:02:52.073+0000"},{"body":"Great.\n\nI just created a new patch (11349-2.1-v3.patch) where the 'update' method is empty in the validation tracker (in fact, it was a left-over from previous attempts and should have been empty ).\n\nThe main difference for example between the \"update\" method from the regular compaction, and the addRangeTombstone from the validation compaction is the returned value. In the latter case, it's wether the RT is superseded/shadowed by another previously met tombstone.\nTo be honest, I did not managed to factorize both of them without compromising readability even if they share some similarities.\n\nI'm a bit skeptical with ValidationCompactionTracker extending RegularCompactionTracker because RegularCompaction has more fields (unwrittenTombstones, atomCount) which would not be used by the ValidationCompactionTracker (and it feels odd to have unused fields).\nDoing it the other side, ie RegularCompactionTracker extending ValidationCompactionTracker, seemed a better fit (RegularCompaction reuses the comparator and openedTombstones), adds more fields, but there is not much to win: only the isDeleted method is in common...\nThus the interface did not seem a bad choice: implementations are less coupled (and could diverge more in the future if needed).\nBut this can be changed if needed (I just wanted to explain design choices and am not opposed to inheritance)\n\nI agree that this way of doing is a \"leaky abstraction\". Nevertheless, the main idea is to have a patch doing minimal architectural changes to the current code base (did not want to refactor anything) to avoid introducing bugs. Moreover, because the 3.X and 3.0.X are not affected, this will stay in the 2.1.X and 2.2.X branches (and won't be technical debt).\nAnyway, it's more a pragmatic solution than an elegant one (and evidently, I am open to a more elegant solution).","created":"2016-05-02T14:28:43.317+0000"},{"body":"Sounds reasonable and I agree that the code changes should be less invasive as possible. We're talking about 2.x so we should avoid heavy refactoring. Mentioned class design could possibly still be improved, but that depends where to go from here..\n\nTo wrap up available patch options:\n\n1) {{1349-2.1.patch}} with 2 changed lines would address the issue initially described\n2) {{11349-2.1-v3.patch}} introduces a bit more changes but will also create correct digests for shadowed range tombstones and cells (see [dtest|https://github.com/spodkowinski/cassandra-dtest/blob/CASSANDRA-11349/repair_tests/repair_test.py#L425])\n\nAny opinions on this except from Fabian's and mine? It would be good to get some feedback from someone how would be actually willing to commit something like this.","created":"2016-05-03T12:35:41.010+0000"},{"body":"As I see it neither solution will be sufficient. A lot of the visible effects of the problem come as a side effect of CASSANDRA-7953, but there are some underlying issues that are only really solved in 3.0 by the new tombstone handling from CASSANDRA-8099.\n\nWhether we change {{onDiskAtomComparator}} or not, we will still get disordered or multiple equal range tombstones from a single source as that's how they are written in the sstables. {{MergeIterator}} will not combine equal entries from the same source, even if it did and everything was written using the {{onDiskAtomComparator}} (which I don't believe to be the case), it is still in the wrong order for resolving which tombstones can be deleted without delaying their processing.\n\nIn other words the problem cannot be solved by changing the reducer; we can, however, do it if we change {{update}} to follow closely or, better still, _call_ {{IndexBuilder.buildForCompaction}} and make the builder accept a prepared atom serializer (or some subinterface) instead of an output file, and update the digest in the calls to that serializer.","created":"2016-05-04T12:54:15.029+0000"},{"body":"I have something like [this|https://github.com/apache/cassandra/compare/trunk...blambov:11349] in mind.","created":"2016-05-05T14:07:02.704+0000"},{"body":"To quickly sum up the current behavior.. {{ColumnIndex.Builder}} is created for each {{LazilyCompactedRow.update()}} call. The builder will iterate through all atoms produced by the {{MergeIterator}} and uses a {{RangeTombstone.Tracker}} instance for tombstone normalization. Tombstones will be added to the tracker from {{Builder.add()}} and by {{LCR.Reducer.getReduced()}}, which in turn will be called once for all atoms for the same column as considered by {{onDiskAtomComparator}}. \n\n[~blambov], so what you're saying is that we can't be sure that the {{MergeIterator}} will always be able to provide deterministically ordered values, as write order may be different and we therefor cannot simply iterate through the reducer to create a correct digest. \n\nWhat I'm a bit concerned about while trying to understand Branimir's approach is that at some point {{getReduced()}} will add the RT to the tracker while in another scenario the RT will be added later and will cause the serializer be called differently as well. Or to put this in other words, if we can't be sure about the reducer returning deterministically ordered values, won't this effect the tracker and digest calculation in the builder as well?","created":"2016-05-10T11:43:11.034+0000"},{"body":"Not precisely: One part of the problem is that we cannot ensure that e.g. the same range tombstone will not come twice from the same sstable (in which case {{MergeIterator}} would issue two separate {{getReduced}} calls) or from two different sstables (in which case {{MergeIterator}} would call {{getReduced}} once). Another is that while compaction uses the tracker to identify when a tombstone is redundant and can be omitted, {{getReduced}} does not have that information at the time it processes that tombstone because the covering tombstone has not arrived yet.\n\nThe tracker can properly resolve these situations, but it can't do it without delaying which causes the necessity for abusing the serializer.\n\nThe reducer only adds RTs to the tracker if it would not return them in the output for some reason (e.g. expiration), the point being to always pass on the full stream of RTs; the order should not be affected by it choosing to do that.","created":"2016-05-10T12:42:08.108+0000"},{"body":"Thanks for the clarification. It's really helpful to understand the intention of how those parts are suppose to work together. \n\nThe serializer approach seems to be a good idea how to handle this, but there are still [cases|https://github.com/spodkowinski/cassandra-dtest/blob/b110685bceddbcb63ebc744ba54a25cb268f2478/repair_tests/repair_test.py#L438:L451] \\[1\\] not handled correctly. I'm going to take a closer look to understand why. I'd also like to do some more testing for potential digest mismatch storms during rolling upgrades, but wouldn't expect any blockers so far. \n\n\\[1\\] nosetests repair_tests/repair_test.py:TestRepair.shadowed_range_tombstone_digest_parallel_repair_test\n","created":"2016-05-11T13:13:33.227+0000"},{"body":"Ok, this seems a better approach (with this approach RT are added to the tracker through the ColumnIndex.add method).\nI had some time to test it on a dev environment and repair did not find any difference (which is a good thing).\n\nRegarding the case that is not working correctly, I think the solution is to use a RangeTombstoneList before writing RangeTombstones.\n\nThe current implementation Tracker.writeUnwrittenTombstones(...) is:\n{noformat}\n for (RangeTombstone rt : unwrittenTombstones)\n {\n size += writeTombstone(rt, out, atomSerializer);\n }\n{noformat}\nAnd should be replaced by:\n {noformat}\n RangeTombstoneList rtl = new RangeTombstoneList(comparator, unwrittenTombstones.size());\n for (RangeTombstone rt : unwrittenTombstones)\n {\n rtl.add(rt);\n }\n for (RangeTombstone rt : rtl)\n {\n size += writeTombstone(rt, out, atomSerializer);\n }\n{noformat}\nI haven't tested this but it should work.\nThe explanation for this is the following:\n - on node1, due to the flushes, each RT is written in its own SSTable\n - on node2, because all RTs are kept in memory, they're kept in a RangeTombstoneList. This RangeTombstoneList will keep non overlapping RTs.\n\nDuring repair, on node1, RTs are merged but are kept as is (ie some RTs can be overlapped) while on node2, they can't.\nBy using the RangeTombstoneList before serializing the unwritten RT, no RT can overlap another RT.\n\nNote: doing the change above will also change the way RT are serialized during normal compactions...","created":"2016-05-17T21:53:44.363+0000"},{"body":"There will be cases where this {{RangeTombstoneList}} solution is not sufficient (e.g. inserting {{c1 = 'a' AND c2 = 'b' AND c3 = 'a'}} at the end of the test above).\n\nIs it imperative that we fix all scenarios here if 3.0 has the proper solution?","created":"2016-05-19T07:55:51.737+0000"},{"body":"I've been debuging the latest mentioned error case using the following cql/ccm statements and a local 2 node cluster.\n\n{code}\ncreate keyspace ks WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 2};\nuse ks;\nCREATE TABLE IF NOT EXISTS table1 ( c1 text, c2 text, c3 text, c4 float,\n PRIMARY KEY (c1, c2, c3)\n) WITH compaction = {'class': 'SizeTieredCompactionStrategy', 'enabled': 'false'};\nDELETE FROM table1 USING TIMESTAMP 1463656272791 WHERE c1 = 'a' AND c2 = 'b' AND c3 = 'c';\nccm node1 flush\nDELETE FROM table1 USING TIMESTAMP 1463656272792 WHERE c1 = 'a' AND c2 = 'b';\nccm node1 flush\nDELETE FROM table1 USING TIMESTAMP 1463656272793 WHERE c1 = 'a' AND c2 = 'b' AND c3 = 'd';\nccm node1 flush\n{code}\n\nTimestamps have been added for easier tracking of the specific tombstone in the debugger.\n\nColmnIndex.Builder.buildForCompaction() will add tombstones in the following order to the tracker:\n\n*Node1*\n\n{{1463656272792: c1 = 'a' AND c2 = 'b'}}\nFirst RT, added to unwritten + opened tombstones\n\n{{1463656272791: c1 = 'a' AND c2 = 'b' AND c3 = 'c'}}\nOvershadowed by RT added before while being older at the same time. Will not be added and simply ignored.\n\n{{1463656272793: c1 = 'a' AND c2 = 'b' AND c3 = 'd'}}\nOvershaded by first and only RT added to opened so far, but newer and will thus be added to unwritten+opened\n\nWe end up with 2 unwritten tombstones (..92+..93) passed to the serializer for message digest.\n\n\n*Node2*\n\n{{1463656272792: c1 = 'a' AND c2 = 'b'}} (EOC.START)\nFirst RT, added to unwritten + opened tombstones\n\n{{1463656272793: c1 = 'a' AND c2 = 'b' AND c3 = 'd'}} (EOC.END)\ncomparision of EOC flag (Tracker:251) of previously added RT will cause having it removed from the opened list (Tracker:258). Afterwards the current RT will be added to unwritten + opened.\n\n{{1463656272792: c1 = 'a' AND c2 = 'b'}} ({color:red}again!{color})\nGets compared with prev. added RT, which supersedes the current one and thus stays in the list. Will again be added to unwritten + opened list.\n\nWe end up with 3 unwritten RTs, including 1463656272792 twice.\n\n-I still haven't been able to exactly pinpoint why the reducer will be called twice with the same TS, but since [~blambov] explicitly mentioned that possibility, I guess it's intended behavior (but why? :)).-\n\nRunning sstable2json makes it more obvious how node2 flushes the RTs:\n\n{noformat}\n[\n{\"key\": \"a\",\n \"cells\": [[\"b:_\",\"b:d:_\",1463656272792,\"t\",1463731877],\n [\"b:d:_\",\"b:d:!\",1463656272793,\"t\",1463731886],\n [\"b:d:!\",\"b:!\",1463656272792,\"t\",1463731877]]}\n]\n{noformat}","created":"2016-05-19T14:49:20.304+0000"},{"body":"[~blambov] We have 4 clusters impacted by this bug, and for 3 out of 4, what you have in mind works.\nI still need to verify for the 4th one. I'll try to verify this today.\nRegarding the 3.0, migrating 60 nodes is not something done easily.\n\n[~spodxx@gmail.com] Yes, there are 3 RT on node2, because, in memory, RT are stored in a RangeTombstoneList (then serialized). The RangeTombstoneList automatically split the tombstones which are overlapping.","created":"2016-05-20T09:10:39.801+0000"},{"body":"Ok, it appears that the initial idea by [~blambov] is sufficient (after having done some basic testing for our 4th cluster).\n\nNevertheless, I'm surprised that we seems to be the only one affected by this issue. Maybe it's because it took us some time to realize it and investigate it, and there was no clear sign apart from big streams during repairs + data set size increasing too fast. So this may explain why not many people reported it, but there may be others affected out in the wild. That's why it's probably best to try to fix most of it (if it's not possible to fix it entirely), but I also understand that the less changes there are, the less risky it is...\n\nSo I'm good with it either partially fixed or mostly fixed. ","created":"2016-05-20T16:57:27.261+0000"},{"body":"I've now created a new patch version [here|https://github.com/spodkowinski/cassandra/commit/c8601f8cd3921e754bcbe8c9362cf3d2e7072e1e] that basically combines both of your ideas of doing the digest updates in the serializer and using {{RangeTombstonesList}} to normalize RT intervals. Tests look good, feel free to add your own. [~blambov], can you think of any further cases that would not be covered by this approach? ","created":"2016-05-20T16:58:58.732+0000"},{"body":"Thanks Stefan.\n\nSo if I understand well, your latest branch does not changes how SSTables are serialized on disk (by using 2 specialized serializers: one for compaction and one for validation) but still solves all cases (or at least all known cases).\n\nAny chance that this patch can be included in 2.1.15 ?","created":"2016-05-31T16:31:36.745+0000"},{"body":"I've now attached a patch for the last mentioned implementation as {{11349-2.1-v4.patch}} and {{11349-2.2-v4.patch}} to the ticket.\n\nTest results are as follows (reported failures cannot be reproduced locally and seem to be unrelated to me):\n\n||2.1||2.2||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.1]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.2]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-testall/]|\n\nAnyone willing to take another look and actually commit a patch for this issue? I've pushed my WIP branch [here|https://github.com/spodkowinski/cassandra/commits/WIP2-11349] with individual commits that might help during the review.\n","created":"2016-06-02T10:15:07.154+0000"},{"body":"Does this really solve the problem with the test you mentioned? Putting the tombstones through {{RangeTombstoneList}} will normalize them, but they may not be issued in the right position, i.e. the RTL solution only works if the data contains only tombstones. For example, the {{\\[\"b:d:\\!\",\"b:\\!\",1463656272792,\"t\",1463731877\\]}} part from the test above gets issued before a potential token that may come before {{b:d:!}}.\n\nThe test needs to be extended to include live tokens, for example by adding each of\n{code}\nINSERT INTO table1 (c1, c2, c3, c4) VALUES ('b', 'b', 'a', 1)\n{code}\nor\n{code}\nINSERT INTO table1 (c1, c2, c3, c4) VALUES ('b', 'd', 'a', 1)\n{code}\nor\n{code}\nINSERT INTO table1 (c1, c2, c3, c4) VALUES ('b', 'e', 'a', 1)\n{code}\nafter the deletions.\n\nThe RTL solution will break (in different ways) for at least two of the above. It also has performance implications that I am not really happy to take. A proper solution is to either fully replicate what RTL does in the tombstone tracker (which may be not be worth it so late in the lifespan of 2.1 and 2.2), or make the tombstone tracker wrap around an RTL (which may be inefficient and is still somewhat tricky).\n\nIf (as Fabien's testing seems to imply) doing the digest update as serialization solves the majority of the differences and repair pain, I would prefer to stop there.","created":"2016-06-08T21:08:07.694+0000"},{"body":"You're correct by pointing out that live columns can prevent fully normalizing all RTs using the RTL approach in patch v4. It will still be more accurate than without RTL consolidation, but the question is if the additional complexity is worth it. If you'd be more comfortable going with the patch initially suggested by yourself, I'm confident that this will still be a big improvement. \n","created":"2016-06-10T13:47:41.890+0000"},{"body":"Just to let you know that we packaged the patch done by Branimir (as it is the one that have more chances to be included mainstream).\n\nWe restored one cluster (3 nodes, 100GB of data per node, affected table is 25GB) from a snapshot on new hardware, and did a full repair. So far, so good, not much differences are found for the affected table but this was expected because repairs are not run for a few months (around a hundred VS a few hundred of thousands before).\n\nWe will continue testing by recreating all of our clusters, and then, deploy it on our production (and I'll let you know once this is done).","created":"2016-06-16T09:38:28.449+0000"},{"body":"Had a look here, and I'm more comfortable with sticking to [~blambov] approach. For 2.1 and 2.2, we're now in \"only critical bug fixes\" and running things through RTL definitively changes things too much for my comfort. That imply I'm fine not fixing every possible problems if that gets us too far (especially since it's properly fixed in 3.0 and not that many people seems to have reported this). And Branimir's approach seems to be making a good enough impact in practice.\n\nSo [~blambov], could you rebase your patch for 2.1 and 2.2 and run CI. After which, if tests are good, I'm +1 committing. ","created":"2016-06-28T10:18:28.005+0000"},{"body":"Rebased patch here:\n|[2.1|https://github.com/blambov/cassandra/tree/11349]|[utests|http://cassci.datastax.com/view/Dev/view/blambov/job/blambov-11349-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blambov/job/blambov-11349-dtest/]|\n|[2.2|https://github.com/blambov/cassandra/tree/11349-2.2]|[utests|http://cassci.datastax.com/view/Dev/view/blambov/job/blambov-11349-2.2-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blambov/job/blambov-11349-2.2-dtest/]|\n","created":"2016-07-04T08:33:54.950+0000"},{"body":"Tests look ok, all failures are either failing on base or failed very recently.","created":"2016-07-04T11:08:47.559+0000"},{"body":"Perfect, thanks, committed.","created":"2016-07-05T09:27:30.943+0000"},{"body":"Here is some quick feedback: patch is deployed on production for more than a week and it diminished a lot the streaming during repairs.\n\nA patched version of C* 2.1.14 (containing the patch) has been deployed on all of our production clusters more than a week ago and it works well.\nThere are still a few differences during repairs, but a lot less than before, and this is \"manageable\".\n\nThanks to all of you for your help.","created":"2016-07-19T10:14:16.466+0000"},{"body":"I've looked at some metrics today for one of our clusters that has been updated to 2.1.16 a couple of weeks ago. We used to see tens of thousands of sstables getting streamed each night during repairs with many GBs.\n\nWith 2.1.16 the number of streamed sstables went down to almost none. Thanks for fixing this to everyone involved! :)","created":"2017-01-02T12:19:26.437+0000"}],"conversations":[{"body":"We observed that repair, for some of our clusters, streamed a lot of data and many partitions were \"out of sync\".\nMoreover, the read repair mismatch ratio is around 3% on those clusters, which is really high.\n\nAfter investigation, it appears that, if two range tombstones exists for a partition for the same range/interval, they're both included in the merkle tree computation.\nBut, if for some reason, on another node, the two range tombstones were already compacted into a single range tombstone, this will result in a merkle tree difference.\nCurrently, this is clearly bad because MerkleTree differences are dependent on compactions (and if a partition is deleted and created multiple times, the only way to ensure that repair \"works correctly\"/\"don't overstream data\" is to major compact before each repair... which is not really feasible).\n\nBelow is a list of steps allowing to easily reproduce this case:\n{noformat}\nccm create test -v 2.1.13 -n 2 -s\nccm node1 cqlsh\nCREATE KEYSPACE test_rt WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 2};\nUSE test_rt;\nCREATE TABLE IF NOT EXISTS table1 (\n c1 text,\n c2 text,\n c3 float,\n c4 float,\n PRIMARY KEY ((c1), c2)\n);\nINSERT INTO table1 (c1, c2, c3, c4) VALUES ( 'a', 'b', 1, 2);\nDELETE FROM table1 WHERE c1 = 'a' AND c2 = 'b';\nctrl ^d\n# now flush only one of the two nodes\nccm node1 flush \nccm node1 cqlsh\nUSE test_rt;\nINSERT INTO table1 (c1, c2, c3, c4) VALUES ( 'a', 'b', 1, 3);\nDELETE FROM table1 WHERE c1 = 'a' AND c2 = 'b';\nctrl ^d\nccm node1 repair\n# now grep the log and observe that there was some inconstencies detected between nodes (while it shouldn't have detected any)\nccm node1 showlog | grep \"out of sync\"\n{noformat}\nConsequences of this are a costly repair, accumulating many small SSTables (up to thousands for a rather short period of time when using VNodes, the time for compaction to absorb those small files), but also an increased size on disk.\n","from":"reporter","subject":"MerkleTree mismatch when multiple range tombstones exists for the same partition and interval"},{"body":"Looks like the {{MergeIterator.ManyToOne}} logic gets in the way of {{LazilyCompactedRow.Reducer}} doing it's job. The iterator will stop adding atoms to the reducer and continue to advance, once two range tombstones with different deletion times are about to be merged.","from":"developer"},{"body":"\n\nI gave the patch some more thoughts and I'm now confident that the proposed change is the best way to address the issue. \n\nBasically what happens during validation compaction is that a scanner is created for each sstable. The {{CompactionIterable.Reducer}} will then create a {{LazilyCompactedRow}} with an iterable of {{OnDiskAtom}} for the same key in each sstable. The purpose of {{LazilyCompactedRow}} during validation compaction is to create a digest of the compacted version of all atoms that would represent a single row. This is done cell by cell, where each collection of atoms for a single cell name is consumed by {{LazilyCompactedRow.Reducer}}.\n \n The decision on whether {{LazilyCompactedRow.Reducer}} should finish to merge a cell and move to the next one is currently being done by {{AbstractCellNameType.onDiskAtomComparator}}, as evaluated by {{MergeIterator.ManyToOne}}. However, the comparator does not only compare by name, but also by {{DeletionTime}} in case of {{RangeTombstone}}. As a consequence, {{MergeIterator.ManyToOne}} will advance in case two {{RangeTombstone}} with different deletion times are read, which breaks the \"_will be called one or more times with cells that share the same column name_\" contract in {{LazilyCompactedRow.Reducer}}.\n\nThe submitted patch will introduce a new {{Comparator}} that will basically work like {{onDiskAtomComparator}}, but does not compare deletion time. As simple as that.\n\n\n||2.1||2.2||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.1]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.2]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-testall/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-dtest/]|\n\n\nThe only other places where {{LazilyCompactedRow}} is being used except validation compaction are the cleanup and scrub functions, which shouldn't be affected, as those are working on individual sstables and I assume that there's no case where an sstable can have multiple identical range tombstones with different timestamps.\n","from":"developer"},{"body":"Nice patch.\n\nI will be able to test it on a dev environment either this week or at the beginning of next week.\n\nThere is still one case not covered (though not sure it can happen).\nSuppose that in SSTable 1, there is a range tombstone covering the columns \"a\" through \"g\" at time t1, and in SSTable 2, there is a range tombstone covering the columns \"c\" through \"d\" at time t2.\nIf those two SSTable are merged (for example on another replica), it will be split in 3 range tombstones in one SSTable (one range tombstone \"a\" -> \"b\" at t1, \"c\" -> \"d\" at t2, \"e\" -> \"f\" at t1).\nComputing the merkle tree for those two hosts will still be different (not the same range tombstones).\n\nAs said above, not sure if it can happen, and anyway, this patch is a good improvement and probably fits 99% cases.\n\n","from":"developer"},{"body":"I tested the patch on a dev environment containing production data and there were still some differences.\n\nThe test procedure was:\n - use ccm & the branch linked to this ticket (I verified that the classpath is ok)\n - copy a fresh backup of production data\n - did a first full repair -> it had some differences on all CFs but this can be explained if there is a small delay when snapshotting all hosts for the backup (this cluster receive a few thousands writes per second)\n - did a second full repair -> only one CF had no differences (the one without range tombstone) and all others had differences (while they should not because there are no reads & writes on this dev environment)\n\nI will continue to investigate and try to isolate those differences...","from":"developer"},{"body":"Thanks, the patch is OK.\n\nIn fact, the differences were produced by another bug that I will create separately.","from":"developer"},{"body":"I suspect this doesn't affect 3.x: has someone checked, and if not, can someone do so we know if some 3.x version is needed or not for this?","from":"developer"},{"body":"I tested against the 3.0.4 and it is not affected (not tested the 3.X, but assumed that it's not affected).\nThere is another similar ticket (the 3.0.4 is not affected): https://issues.apache.org/jira/browse/CASSANDRA-11477 ","from":"developer"},{"body":"[~spodxx@gmail.com] [~slebresne] [~frousseau] I'm confused here. Why should repair be special cased over normal compaction in this case? If the times are different then you *do* still need to resolve it as you need to take the greater time.\n\nIt seems to me the crux of the current patch is to \"fix\" this by special casing the comparator to just compare just the max value of the interval during repair validation:\n\n{code}\n // only compare interval, but not deletion time\n+ return AbstractCellNameType.this.compare(((RangeTombstone)c1).max, ((RangeTombstone)c2).max);\n{code}\n\nI just did my best to merge and compare the code between 2.0 and 2.1 and I'm still trying to parse how this code is different in 2.0 vs. 2.1... We've been unable to reproduce this in 2.0 so far, but the bits of the code being touched here don't seem to be different so I'm trying to understand why 2.1 would hit this and not 2.0.\n\nCould you please explain a bit more why we we can ignore the timestamp?\n","from":"developer"},{"body":"And just for my sanity and for discussion in the Jira, here is the current handling in the comparator\n\n{code}\nif (c1 instanceof RangeTombstone)\n{\n if (c2 instanceof RangeTombstone)\n {\n RangeTombstone t1 = (RangeTombstone)c1;\n RangeTombstone t2 = (RangeTombstone)c2;\n int comp2 = AbstractCellNameType.this.compare(t1.max, t2.max);\n return comp2 == 0 ? t1.data.compareTo(t2.data) : comp2;\n }\n else\n {\n return -1;\n }\n }\n else\n {\n return c2 instanceof RangeTombstone ? 1 : 0;\n }\n}\n{code}","from":"developer"},{"body":"I'm also not sure how this is meant to fix it. Special casing validation compaction may fix repairs but you'd still get the digest mismatches on reads.","from":"developer"},{"body":"I think there are a couple of things wrong with the current patch.\n\nFirst, the comparator needs to continue to compare the full cell name first, and only break ties on range tombstones with {{compare(t1.max, t2.max)}}.\n\nSecond, I believe we should remove the timestamp tie-breaking behavior from the comparator in general, and not just for validation compactions. In other words, I think we're doing the comparison incorrectly for all compactions right now.\n\nWe want the comparison to return 0 whenever range tombstones have equal names and ranges, even if they have different timestamps. This will result in {{LazilyCompactedRow.Reducer.reduce()}} being called in one round with each of the tombstones that only differ in timestamp. The logic in {{LCR.Reducer.reduce()}} already handles the case of multiple range tombstones with different timestamps by picking the one with the highest timestamp, so these will correctly be reduced to a single RT. It looks like the current codebase will keep both range tombstones during a compaction, which isn't necessarily harmful, but is suboptimal. For repair purposes, though, this is incorrect as it produces a different digest.\n\nTo summarize: I think all we need to do is remove the timestamp tie-breaking logic from the existing comparator.\n\n[~slebresne] should double-check my logic, though.","from":"developer"},{"body":"It makes sense just to modify {{onDiskAtomComparator}}. Given the generic name I assumed the comparator is used in other places as well, but since its only used in {{LazyCompactedRow}} we can just change the patch as suggested and simply remove the timestamp tie-break behaviour in {{onDiskAtomComparator}}. \n\nAs for regular compactions, I agree with Tyler that this should not effect compactions in a way that it does with validation compaction. Before the patch, {{LazyCompactedRow}} would not reduce both RTs but instead have {{ColumnIndex.buildForCompaction()}} iterate over both RTs and have them added to the {{RangeTombstone.Tracker}}. The tracker would merge them in a way {{LCR.Reducer.getReduced}} would after the patch. However, I’m not fully sure if there could be some other for more complex cases where this still would cause problems.\n\nAlthough the patch should fix the described issue, the way we deal with RTs during validation compaction is still not ideal. The problem is that LCR lacks some handling of relationships between RTs compared to {{RangeTombstone.Tracker}}. If we create digests column by column, we get wrong results for shadowing tombstones not sharing the same intervals.\n\n{noformat}\nCREATE KEYSPACE test_rt WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 2};\nUSE test_rt;\nCREATE TABLE IF NOT EXISTS table1 (\n c1 text,\n c2 text,\n c3 text,\n c4 float,\n PRIMARY KEY (c1, c2, c3)\n) WITH compaction = {'class': 'SizeTieredCompactionStrategy', 'enabled': 'false'};\nDELETE FROM table1 WHERE c1 = 'a' AND c2 = 'b' AND c3 = 'c';\n\nccm node1 flush\n\nDELETE FROM table1 WHERE c1 = 'a' AND c2 = 'b';\n\nccm node1 repair test_rt table1\n{noformat}\n\n\nIn this case the (c1, c2, c3) RT will always be repaired after it has been compacted with (c1, c2) on any node. \nSo I’m wondering if we shouldn’t take a more bold approach here than the patch does. \n","from":"developer"},{"body":"Using the RangeTombstone.Tracker can help in the situation described just above.\n\nIn fact, the RT should always update the tracker (see CASSANDRA-11477).\nThe trick here is to always considered it as \"expired\" in the tracker (even if not) so the tombstones are not accumulated during compaction (if expired the tracker keeps only the list of opened RTs and if not, it keeps all unwritten RTs, ie all RTs because it's a validation compaction...).\n\nHaving a look at the update method of the Tracker, it already check if the tombstone is superseded by another one (and don't add it as \"opened\" if superseded).\n\nThus, the v2 patch:\n - includes the previous patch\n - always update the tracker with the RT (considering it as expired even if not, just to not retain too many of them in memory, and because it's for validation, it's a read only and won't affect anything)\n - test if the RT was added in the openedTombstones list, and if that's not the case, skip it for digest.\n\nI know that the patch may be a bit rough (at least on the \"isLastOpened\" method) but it is more to validate the approach first and did not want the patch to be too invasive (by modifying the returned value of the update method).\n\nWDYT ?\n\nNote: I have not yet tested it against our production data\nNote2: Regarding the read-repair, this seems to be a different story and can't see anything for now that could explain those differences (will dig later on this as this is less urgent)","from":"developer"},{"body":"[~frousseau], what makes things more complicated here is that changes to LCR will effect regular compactions as well. Adding all tombstones as expired in your {{11349-2.1-v2.patch}} will have unwanted side effects for regular compactions, e.g. try {{RangeTombstoneMergeTest}} with it.\n\nI've now spend some time trying to make use of the RT.Tracker there but without much success. Adding non-expired range tombstones to the tracker from within LCR would cause corrupted sstables. Even creating an edge case just for validation compaction would not handle all potential TS shadowing scenarios and will probably cause more harm than good (and potential digest mismatch storms). I'm not even sure it's possible given the current iterative MergeIterator > LazilyCompactedRow > RT.Tracker interaction. \n\nI'm now at a point where I'd suggest to just stick with {{11349-2.1.patch}} unless someone else has a better idea how to solve this. I've updated the [dtest PR|https://github.com/riptano/cassandra-dtest/pull/881] with two of the described shadowing scenarios that will only work with 3.0+ even after the patch, if someone wants to give it a try.\n\nCassci results for {{11349-2.1.patch}}:\n\n||2.1||2.2||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.1]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.2]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-testall/]|\n\n","from":"developer"},{"body":"Can we keep the conversation going to get this patch into the next 2.x release?","from":"developer"},{"body":"Sorry for not being reactive lately, I'm rather busy atm...\n\nI'd be more than happy[1] to see this patch in the next release.\nI haven't tested it yet and probably can find some time next week to test it on a dev cluster if it can help.\nNevertheless, I won't be able to tell if it really worked because there will still have some mismatches (due to CASSANDRA-11477).\n\nI have started working on a patch which should be able to handle both CASSANDRA-11477 and the last edge case.\n\nWhat it basically does:\n - Tracker is now an interface\n - there are two implementations: one called RegularCompactionTracker and another one ValidationCompactionTracker\n - the ColumnIndexer.Builder has one more optional parameter : a boolean to know if it is built for validation\n - the RegularCompactionTracker is identical to the existing Tracker + one empty method\n - the ValidationCompactionTracker is similar to the existing Tracker but retain only opened tombstones (most methods are thus empty)\n - the Reducer slightly changed but its behaviour is the same regarding the regular compactions\n\nI can share it if you're interested (code compiles but I still haven't tested it at all and plan to do it soon and share it after).\n\n[1] Just to share more information: those issues are important to us, because a few of our clusters are impacted and a few days after filing the bug, we decided to temporarily stop repairing some tables (knowing that we could live with inconsistencies on those tables) which were heavily impacted by those bugs (each repair increased disk occupancy by a few percent), and did a major compaction. This resulted in two to three times less disk occupancy (One table shrinked from 243GB to 79GB. Note that this was not due to tombstones reclaiming old data because, it's been nearly a month now, and the big SSTable resulting from the major compaction is still there but disk usage has not grown that much). ","from":"developer"},{"body":"Ok, I uploaded a new version of the patch (11349-2.1-v2.patch)\n\nAs said above, there are two Tracker implementations now: one for regular compaction and another one for validation compaction.\nIt solves both cases described here (the one in the ticket + the one in the comment of Stefan) and CASSANDRA-11477.","from":"developer"},{"body":"I'm not sure introducing a new tracker interface is the best way to handle this. It took me a while to actually figure out the differences between the {{update}} implementations in both trackers, since for most parts it's sharing the same copied code. It would probably be better to have ValidationTracker subclass RegularCompactionTracker, add {{remove/addUnwrittenTombstone}} implemented empty for validation.\n\nThe {{addRangeTombstone}} semantics also look like a case of leaky abstractions to me. It's adding nothing at all for regular compaction, but serves as early exit path for validation. \n\nGood news is that the dtests and unit tests seem to pass with the patch. :)","from":"developer"},{"body":"Great.\n\nI just created a new patch (11349-2.1-v3.patch) where the 'update' method is empty in the validation tracker (in fact, it was a left-over from previous attempts and should have been empty ).\n\nThe main difference for example between the \"update\" method from the regular compaction, and the addRangeTombstone from the validation compaction is the returned value. In the latter case, it's wether the RT is superseded/shadowed by another previously met tombstone.\nTo be honest, I did not managed to factorize both of them without compromising readability even if they share some similarities.\n\nI'm a bit skeptical with ValidationCompactionTracker extending RegularCompactionTracker because RegularCompaction has more fields (unwrittenTombstones, atomCount) which would not be used by the ValidationCompactionTracker (and it feels odd to have unused fields).\nDoing it the other side, ie RegularCompactionTracker extending ValidationCompactionTracker, seemed a better fit (RegularCompaction reuses the comparator and openedTombstones), adds more fields, but there is not much to win: only the isDeleted method is in common...\nThus the interface did not seem a bad choice: implementations are less coupled (and could diverge more in the future if needed).\nBut this can be changed if needed (I just wanted to explain design choices and am not opposed to inheritance)\n\nI agree that this way of doing is a \"leaky abstraction\". Nevertheless, the main idea is to have a patch doing minimal architectural changes to the current code base (did not want to refactor anything) to avoid introducing bugs. Moreover, because the 3.X and 3.0.X are not affected, this will stay in the 2.1.X and 2.2.X branches (and won't be technical debt).\nAnyway, it's more a pragmatic solution than an elegant one (and evidently, I am open to a more elegant solution).","from":"developer"},{"body":"Sounds reasonable and I agree that the code changes should be less invasive as possible. We're talking about 2.x so we should avoid heavy refactoring. Mentioned class design could possibly still be improved, but that depends where to go from here..\n\nTo wrap up available patch options:\n\n1) {{1349-2.1.patch}} with 2 changed lines would address the issue initially described\n2) {{11349-2.1-v3.patch}} introduces a bit more changes but will also create correct digests for shadowed range tombstones and cells (see [dtest|https://github.com/spodkowinski/cassandra-dtest/blob/CASSANDRA-11349/repair_tests/repair_test.py#L425])\n\nAny opinions on this except from Fabian's and mine? It would be good to get some feedback from someone how would be actually willing to commit something like this.","from":"developer"},{"body":"As I see it neither solution will be sufficient. A lot of the visible effects of the problem come as a side effect of CASSANDRA-7953, but there are some underlying issues that are only really solved in 3.0 by the new tombstone handling from CASSANDRA-8099.\n\nWhether we change {{onDiskAtomComparator}} or not, we will still get disordered or multiple equal range tombstones from a single source as that's how they are written in the sstables. {{MergeIterator}} will not combine equal entries from the same source, even if it did and everything was written using the {{onDiskAtomComparator}} (which I don't believe to be the case), it is still in the wrong order for resolving which tombstones can be deleted without delaying their processing.\n\nIn other words the problem cannot be solved by changing the reducer; we can, however, do it if we change {{update}} to follow closely or, better still, _call_ {{IndexBuilder.buildForCompaction}} and make the builder accept a prepared atom serializer (or some subinterface) instead of an output file, and update the digest in the calls to that serializer.","from":"developer"},{"body":"I have something like [this|https://github.com/apache/cassandra/compare/trunk...blambov:11349] in mind.","from":"developer"},{"body":"To quickly sum up the current behavior.. {{ColumnIndex.Builder}} is created for each {{LazilyCompactedRow.update()}} call. The builder will iterate through all atoms produced by the {{MergeIterator}} and uses a {{RangeTombstone.Tracker}} instance for tombstone normalization. Tombstones will be added to the tracker from {{Builder.add()}} and by {{LCR.Reducer.getReduced()}}, which in turn will be called once for all atoms for the same column as considered by {{onDiskAtomComparator}}. \n\n[~blambov], so what you're saying is that we can't be sure that the {{MergeIterator}} will always be able to provide deterministically ordered values, as write order may be different and we therefor cannot simply iterate through the reducer to create a correct digest. \n\nWhat I'm a bit concerned about while trying to understand Branimir's approach is that at some point {{getReduced()}} will add the RT to the tracker while in another scenario the RT will be added later and will cause the serializer be called differently as well. Or to put this in other words, if we can't be sure about the reducer returning deterministically ordered values, won't this effect the tracker and digest calculation in the builder as well?","from":"developer"},{"body":"Not precisely: One part of the problem is that we cannot ensure that e.g. the same range tombstone will not come twice from the same sstable (in which case {{MergeIterator}} would issue two separate {{getReduced}} calls) or from two different sstables (in which case {{MergeIterator}} would call {{getReduced}} once). Another is that while compaction uses the tracker to identify when a tombstone is redundant and can be omitted, {{getReduced}} does not have that information at the time it processes that tombstone because the covering tombstone has not arrived yet.\n\nThe tracker can properly resolve these situations, but it can't do it without delaying which causes the necessity for abusing the serializer.\n\nThe reducer only adds RTs to the tracker if it would not return them in the output for some reason (e.g. expiration), the point being to always pass on the full stream of RTs; the order should not be affected by it choosing to do that.","from":"developer"},{"body":"Thanks for the clarification. It's really helpful to understand the intention of how those parts are suppose to work together. \n\nThe serializer approach seems to be a good idea how to handle this, but there are still [cases|https://github.com/spodkowinski/cassandra-dtest/blob/b110685bceddbcb63ebc744ba54a25cb268f2478/repair_tests/repair_test.py#L438:L451] \\[1\\] not handled correctly. I'm going to take a closer look to understand why. I'd also like to do some more testing for potential digest mismatch storms during rolling upgrades, but wouldn't expect any blockers so far. \n\n\\[1\\] nosetests repair_tests/repair_test.py:TestRepair.shadowed_range_tombstone_digest_parallel_repair_test\n","from":"developer"},{"body":"Ok, this seems a better approach (with this approach RT are added to the tracker through the ColumnIndex.add method).\nI had some time to test it on a dev environment and repair did not find any difference (which is a good thing).\n\nRegarding the case that is not working correctly, I think the solution is to use a RangeTombstoneList before writing RangeTombstones.\n\nThe current implementation Tracker.writeUnwrittenTombstones(...) is:\n{noformat}\n for (RangeTombstone rt : unwrittenTombstones)\n {\n size += writeTombstone(rt, out, atomSerializer);\n }\n{noformat}\nAnd should be replaced by:\n {noformat}\n RangeTombstoneList rtl = new RangeTombstoneList(comparator, unwrittenTombstones.size());\n for (RangeTombstone rt : unwrittenTombstones)\n {\n rtl.add(rt);\n }\n for (RangeTombstone rt : rtl)\n {\n size += writeTombstone(rt, out, atomSerializer);\n }\n{noformat}\nI haven't tested this but it should work.\nThe explanation for this is the following:\n - on node1, due to the flushes, each RT is written in its own SSTable\n - on node2, because all RTs are kept in memory, they're kept in a RangeTombstoneList. This RangeTombstoneList will keep non overlapping RTs.\n\nDuring repair, on node1, RTs are merged but are kept as is (ie some RTs can be overlapped) while on node2, they can't.\nBy using the RangeTombstoneList before serializing the unwritten RT, no RT can overlap another RT.\n\nNote: doing the change above will also change the way RT are serialized during normal compactions...","from":"developer"},{"body":"There will be cases where this {{RangeTombstoneList}} solution is not sufficient (e.g. inserting {{c1 = 'a' AND c2 = 'b' AND c3 = 'a'}} at the end of the test above).\n\nIs it imperative that we fix all scenarios here if 3.0 has the proper solution?","from":"developer"},{"body":"I've been debuging the latest mentioned error case using the following cql/ccm statements and a local 2 node cluster.\n\n{code}\ncreate keyspace ks WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 2};\nuse ks;\nCREATE TABLE IF NOT EXISTS table1 ( c1 text, c2 text, c3 text, c4 float,\n PRIMARY KEY (c1, c2, c3)\n) WITH compaction = {'class': 'SizeTieredCompactionStrategy', 'enabled': 'false'};\nDELETE FROM table1 USING TIMESTAMP 1463656272791 WHERE c1 = 'a' AND c2 = 'b' AND c3 = 'c';\nccm node1 flush\nDELETE FROM table1 USING TIMESTAMP 1463656272792 WHERE c1 = 'a' AND c2 = 'b';\nccm node1 flush\nDELETE FROM table1 USING TIMESTAMP 1463656272793 WHERE c1 = 'a' AND c2 = 'b' AND c3 = 'd';\nccm node1 flush\n{code}\n\nTimestamps have been added for easier tracking of the specific tombstone in the debugger.\n\nColmnIndex.Builder.buildForCompaction() will add tombstones in the following order to the tracker:\n\n*Node1*\n\n{{1463656272792: c1 = 'a' AND c2 = 'b'}}\nFirst RT, added to unwritten + opened tombstones\n\n{{1463656272791: c1 = 'a' AND c2 = 'b' AND c3 = 'c'}}\nOvershadowed by RT added before while being older at the same time. Will not be added and simply ignored.\n\n{{1463656272793: c1 = 'a' AND c2 = 'b' AND c3 = 'd'}}\nOvershaded by first and only RT added to opened so far, but newer and will thus be added to unwritten+opened\n\nWe end up with 2 unwritten tombstones (..92+..93) passed to the serializer for message digest.\n\n\n*Node2*\n\n{{1463656272792: c1 = 'a' AND c2 = 'b'}} (EOC.START)\nFirst RT, added to unwritten + opened tombstones\n\n{{1463656272793: c1 = 'a' AND c2 = 'b' AND c3 = 'd'}} (EOC.END)\ncomparision of EOC flag (Tracker:251) of previously added RT will cause having it removed from the opened list (Tracker:258). Afterwards the current RT will be added to unwritten + opened.\n\n{{1463656272792: c1 = 'a' AND c2 = 'b'}} ({color:red}again!{color})\nGets compared with prev. added RT, which supersedes the current one and thus stays in the list. Will again be added to unwritten + opened list.\n\nWe end up with 3 unwritten RTs, including 1463656272792 twice.\n\n-I still haven't been able to exactly pinpoint why the reducer will be called twice with the same TS, but since [~blambov] explicitly mentioned that possibility, I guess it's intended behavior (but why? :)).-\n\nRunning sstable2json makes it more obvious how node2 flushes the RTs:\n\n{noformat}\n[\n{\"key\": \"a\",\n \"cells\": [[\"b:_\",\"b:d:_\",1463656272792,\"t\",1463731877],\n [\"b:d:_\",\"b:d:!\",1463656272793,\"t\",1463731886],\n [\"b:d:!\",\"b:!\",1463656272792,\"t\",1463731877]]}\n]\n{noformat}","from":"developer"},{"body":"[~blambov] We have 4 clusters impacted by this bug, and for 3 out of 4, what you have in mind works.\nI still need to verify for the 4th one. I'll try to verify this today.\nRegarding the 3.0, migrating 60 nodes is not something done easily.\n\n[~spodxx@gmail.com] Yes, there are 3 RT on node2, because, in memory, RT are stored in a RangeTombstoneList (then serialized). The RangeTombstoneList automatically split the tombstones which are overlapping.","from":"developer"},{"body":"Ok, it appears that the initial idea by [~blambov] is sufficient (after having done some basic testing for our 4th cluster).\n\nNevertheless, I'm surprised that we seems to be the only one affected by this issue. Maybe it's because it took us some time to realize it and investigate it, and there was no clear sign apart from big streams during repairs + data set size increasing too fast. So this may explain why not many people reported it, but there may be others affected out in the wild. That's why it's probably best to try to fix most of it (if it's not possible to fix it entirely), but I also understand that the less changes there are, the less risky it is...\n\nSo I'm good with it either partially fixed or mostly fixed. ","from":"developer"},{"body":"I've now created a new patch version [here|https://github.com/spodkowinski/cassandra/commit/c8601f8cd3921e754bcbe8c9362cf3d2e7072e1e] that basically combines both of your ideas of doing the digest updates in the serializer and using {{RangeTombstonesList}} to normalize RT intervals. Tests look good, feel free to add your own. [~blambov], can you think of any further cases that would not be covered by this approach? ","from":"developer"},{"body":"Thanks Stefan.\n\nSo if I understand well, your latest branch does not changes how SSTables are serialized on disk (by using 2 specialized serializers: one for compaction and one for validation) but still solves all cases (or at least all known cases).\n\nAny chance that this patch can be included in 2.1.15 ?","from":"developer"},{"body":"I've now attached a patch for the last mentioned implementation as {{11349-2.1-v4.patch}} and {{11349-2.2-v4.patch}} to the ticket.\n\nTest results are as follows (reported failures cannot be reproduced locally and seem to be unrelated to me):\n\n||2.1||2.2||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.1]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-11349-2.2]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.1-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-11349-2.2-testall/]|\n\nAnyone willing to take another look and actually commit a patch for this issue? I've pushed my WIP branch [here|https://github.com/spodkowinski/cassandra/commits/WIP2-11349] with individual commits that might help during the review.\n","from":"developer"},{"body":"Does this really solve the problem with the test you mentioned? Putting the tombstones through {{RangeTombstoneList}} will normalize them, but they may not be issued in the right position, i.e. the RTL solution only works if the data contains only tombstones. For example, the {{\\[\"b:d:\\!\",\"b:\\!\",1463656272792,\"t\",1463731877\\]}} part from the test above gets issued before a potential token that may come before {{b:d:!}}.\n\nThe test needs to be extended to include live tokens, for example by adding each of\n{code}\nINSERT INTO table1 (c1, c2, c3, c4) VALUES ('b', 'b', 'a', 1)\n{code}\nor\n{code}\nINSERT INTO table1 (c1, c2, c3, c4) VALUES ('b', 'd', 'a', 1)\n{code}\nor\n{code}\nINSERT INTO table1 (c1, c2, c3, c4) VALUES ('b', 'e', 'a', 1)\n{code}\nafter the deletions.\n\nThe RTL solution will break (in different ways) for at least two of the above. It also has performance implications that I am not really happy to take. A proper solution is to either fully replicate what RTL does in the tombstone tracker (which may be not be worth it so late in the lifespan of 2.1 and 2.2), or make the tombstone tracker wrap around an RTL (which may be inefficient and is still somewhat tricky).\n\nIf (as Fabien's testing seems to imply) doing the digest update as serialization solves the majority of the differences and repair pain, I would prefer to stop there.","from":"developer"},{"body":"You're correct by pointing out that live columns can prevent fully normalizing all RTs using the RTL approach in patch v4. It will still be more accurate than without RTL consolidation, but the question is if the additional complexity is worth it. If you'd be more comfortable going with the patch initially suggested by yourself, I'm confident that this will still be a big improvement. \n","from":"developer"},{"body":"Just to let you know that we packaged the patch done by Branimir (as it is the one that have more chances to be included mainstream).\n\nWe restored one cluster (3 nodes, 100GB of data per node, affected table is 25GB) from a snapshot on new hardware, and did a full repair. So far, so good, not much differences are found for the affected table but this was expected because repairs are not run for a few months (around a hundred VS a few hundred of thousands before).\n\nWe will continue testing by recreating all of our clusters, and then, deploy it on our production (and I'll let you know once this is done).","from":"developer"},{"body":"Had a look here, and I'm more comfortable with sticking to [~blambov] approach. For 2.1 and 2.2, we're now in \"only critical bug fixes\" and running things through RTL definitively changes things too much for my comfort. That imply I'm fine not fixing every possible problems if that gets us too far (especially since it's properly fixed in 3.0 and not that many people seems to have reported this). And Branimir's approach seems to be making a good enough impact in practice.\n\nSo [~blambov], could you rebase your patch for 2.1 and 2.2 and run CI. After which, if tests are good, I'm +1 committing. ","from":"developer"},{"body":"Rebased patch here:\n|[2.1|https://github.com/blambov/cassandra/tree/11349]|[utests|http://cassci.datastax.com/view/Dev/view/blambov/job/blambov-11349-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blambov/job/blambov-11349-dtest/]|\n|[2.2|https://github.com/blambov/cassandra/tree/11349-2.2]|[utests|http://cassci.datastax.com/view/Dev/view/blambov/job/blambov-11349-2.2-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blambov/job/blambov-11349-2.2-dtest/]|\n","from":"developer"},{"body":"Tests look ok, all failures are either failing on base or failed very recently.","from":"developer"},{"body":"Perfect, thanks, committed.","from":"developer"},{"body":"Here is some quick feedback: patch is deployed on production for more than a week and it diminished a lot the streaming during repairs.\n\nA patched version of C* 2.1.14 (containing the patch) has been deployed on all of our production clusters more than a week ago and it works well.\nThere are still a few differences during repairs, but a lot less than before, and this is \"manageable\".\n\nThanks to all of you for your help.","from":"developer"},{"body":"I've looked at some metrics today for one of our clusters that has been updated to 2.1.16 a couple of weeks ago. We used to see tens of thousands of sstables getting streamed each night during repairs with many GBs.\n\nWith 2.1.16 the number of streamed sstables went down to almost none. Thanks for fixing this to everyone involved! :)","from":"developer"}],"created":"2016-03-13T11:48:45.000+0000","description":"We observed that repair, for some of our clusters, streamed a lot of data and many partitions were \"out of sync\".\nMoreover, the read repair mismatch ratio is around 3% on those clusters, which is really high.\n\nAfter investigation, it appears that, if two range tombstones exists for a partition for the same range/interval, they're both included in the merkle tree computation.\nBut, if for some reason, on another node, the two range tombstones were already compacted into a single range tombstone, this will result in a merkle tree difference.\nCurrently, this is clearly bad because MerkleTree differences are dependent on compactions (and if a partition is deleted and created multiple times, the only way to ensure that repair \"works correctly\"/\"don't overstream data\" is to major compact before each repair... which is not really feasible).\n\nBelow is a list of steps allowing to easily reproduce this case:\n{noformat}\nccm create test -v 2.1.13 -n 2 -s\nccm node1 cqlsh\nCREATE KEYSPACE test_rt WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 2};\nUSE test_rt;\nCREATE TABLE IF NOT EXISTS table1 (\n c1 text,\n c2 text,\n c3 float,\n c4 float,\n PRIMARY KEY ((c1), c2)\n);\nINSERT INTO table1 (c1, c2, c3, c4) VALUES ( 'a', 'b', 1, 2);\nDELETE FROM table1 WHERE c1 = 'a' AND c2 = 'b';\nctrl ^d\n# now flush only one of the two nodes\nccm node1 flush \nccm node1 cqlsh\nUSE test_rt;\nINSERT INTO table1 (c1, c2, c3, c4) VALUES ( 'a', 'b', 1, 3);\nDELETE FROM table1 WHERE c1 = 'a' AND c2 = 'b';\nctrl ^d\nccm node1 repair\n# now grep the log and observe that there was some inconstencies detected between nodes (while it shouldn't have detected any)\nccm node1 showlog | grep \"out of sync\"\n{noformat}\nConsequences of this are a costly repair, accumulating many small SSTables (up to thousands for a rather short period of time when using VNodes, the time for compaction to absorb those small files), but also an increased size on disk.\n","issue_id":"12949698","key":"CASSANDRA-11349","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-07-05T09:27:30.000+0000","role":"fixed_distractor","summary":"MerkleTree mismatch when multiple range tombstones exists for the same partition and interval"} {"case_id":"12950933","cluster":"DISTRACTOR-CASSANDRA-11363","comments":[{"body":"Also reproducible in 3.0.3.\n\nI have 2 clusters showing exactly the same problem, one running 2.1.13 (5 nodes, 4 cores, 32GB Ram, CMS Java 7) and another 3 nodes, 4cores, 16GB Ram running 3.0.3 (CMS Java 8). \n\nBoth get the problems described: \"The issue seems to coincide with the number of connections OR the number of total requests being processed at a given time (as the latter increases with the former in our system)\"\n\nIt is normal for OS Load to shoot to 33+ with around 40connected clients.","created":"2016-03-16T18:45:55.107+0000"},{"body":"For troubleshooting we set up a coordinator-only node and pointed one app server at it. This resulted in roughly 90 connections to the node. We witnessed many timeouts of requests from the app server's perspective. We downgraded the coordinator-only node to 2.1.9 and upgraded point-release by point-release (DSE point releases) until we saw the same behavior in DSE 4.8.4 (Cassandra 2.1.12).\n\nI'm not certain this has to do with connection count anymore. \n\nWe have several different types of workloads going on, but we found that only the workloads that use batches were timing out. Additionally this is only happening when the node is used as a coordinator. We are not seeing this issue when we disable the binary protocol, effectively making the node no-longer a coordinator.","created":"2016-03-17T17:00:23.800+0000"},{"body":"I can confirm that in both my clusters batches are in use.","created":"2016-03-17T17:09:40.398+0000"},{"body":"I also observed in C*2.1.12 that a certain percentage of the Native-Transport-Requests are blocked, yet no major CPU or resources issue on my side though, it might then not be related.\n\nFor what it is worth here is something I observed about Native-Transport-Requests: increasing the 'native_transport_max_threads' value help mitigating this, as expected, but Native-Transport-Requests number is still a non zero value.\n\n{noformat}\n[alain~]$ knife ssh \"role:cassandra\" \"nodetool tpstats | grep Native-Transport-Requests\" | grep -e server1 -e server2 -e server3 -e server4 | sort | awk 'BEGIN { printf \"%50s %10s\",\"Server |\",\" Blocked ratio:\\n\" } { printf \"%50s %10f%\\n\", $1, (($7/$5)*100) }'\n Server | Blocked ratio:\n server3 0.044902%\n server4 0.030127%\n server2 0.045759%\n server1 0.082763%\n{noformat}\n\nI waited long enough between the change and the result capture, many days, probably a few weeks. As all the nodes are in the same datacenter, under a (fairly) balanced load, this is probably relevant.\n\nHere are the result for those nodes, in our use case.\n\n||Server||native_transport_max_threads||Percentage of blocked threads||\n|server1|128|0.082763%|\n|server2|384|0.044902%|\n|server3|512|0.045759%|\n|server4|1024|0.030127%|\n\nAlso from the mailing list outputs, it looks like it is quite common to have some Native-Transport-Requests blocked, probably unavoidable depending on the network and use cases (spiky workloads?). ","created":"2016-04-07T14:35:39.879+0000"},{"body":"Raised this to critical. \n\nYou have four long-time users with multiple large deployments who are seeing a quantifiable percentage of client errors on non-resource constrained clusters across multiple versions. ","created":"2016-04-07T16:51:03.997+0000"},{"body":"fixed by CASSANDRA-10200 ?","created":"2016-04-07T17:35:19.298+0000"},{"body":"I went through 2.1.12 changes and didn't find anything suspicious. On 2.1.13 and 3.0.3 though, we changed the {{ServerConnection}} query state map from a {{NonBlockingHashMap}} to a {{ConcurrentHashMap}} on CASSANDRA-10938, which might be misbehaving for some reason.\n\nIs anyone willing to try the revert patch below on 2.1.13 or 3.0.3 and check if that changes anything?\n\n{noformat}\ndiff --git a/src/java/org/apache/cassandra/transport/ServerConnection.java b/src/java/org/apache/cassandra/transport/ServerConnection.java\nindex ce4d164..5991b33 100644\n--- a/src/java/org/apache/cassandra/transport/ServerConnection.java\n+++ b/src/java/org/apache/cassandra/transport/ServerConnection.java\n@@ -17,7 +17,6 @@\n */\n package org.apache.cassandra.transport;\n \n-import java.util.concurrent.ConcurrentHashMap;\n import java.util.concurrent.ConcurrentMap;\n \n import io.netty.channel.Channel;\n@@ -29,6 +28,8 @@ import org.apache.cassandra.config.DatabaseDescriptor;\n import org.apache.cassandra.service.ClientState;\n import org.apache.cassandra.service.QueryState;\n \n+import org.cliffc.high_scale_lib.NonBlockingHashMap;\n+\n public class ServerConnection extends Connection\n {\n private enum State { UNINITIALIZED, AUTHENTICATION, READY }\n@@ -37,7 +38,7 @@ public class ServerConnection extends Connection\n private final ClientState clientState;\n private volatile State state;\n \n- private final ConcurrentMap queryStates = new ConcurrentHashMap<>();\n+ private final ConcurrentMap queryStates = new NonBlockingHashMap<>();\n \n public ServerConnection(Channel channel, int version, Connection.Tracker tracker)\n {\n{noformat}","created":"2016-04-07T19:50:26.870+0000"},{"body":"Unfortunately we are on DSE, so I can't run the revert","created":"2016-04-07T19:52:51.167+0000"},{"body":"[~pauloricardomg] unfortunately, this may be been latent for some time. CASSANDRA-10044 re-introduced the tpstats counters as they had been missing (which is mostly likely why this has not been noticed until recently). ","created":"2016-04-07T19:58:51.711+0000"},{"body":"that's right, that was just a wild guess. I focused my quick investigation on 2.1.12 and 2.1.13 commits, but some deeper investigation is probably needed.\n\nprobably it's a good idea to backport CASSANDRA-10044 to versions < 2.1.12 and setup a coordinator-only node and check where it started happening to narrow down the scope.","created":"2016-04-07T20:44:01.264+0000"},{"body":"we may have two separate issues here, in mine, the issue is 100% CPU utilization and ultra high load when using batches. According to the jfr all of hot threads are spinning on \n{code}\norg.apache.cassandra.locator.NetworkTopologyStrategy.hasSufficientReplicas(String, Map, Multimap)\n -> org.apache.cassandra.locator.NetworkTopologyStrategy.hasSufficientReplicas(Map, Multimap)\n -> org.apache.cassandra.locator.NetworkTopologyStrategy.calculateNaturalEndpoints(Token, TokenMetadata)\n{code}","created":"2016-04-07T20:55:59.753+0000"},{"body":"Yeah, that is quite different that what [~arodrime] and I have seen recently. Nodes in our case were otherwise well within utilization thresholds.\n\n[~devdazed] Your issue looks like it would be addressed by [~cnlwsu]'s reference to CASSANDRA-10200 (lean on DSE folks for a patch). ","created":"2016-04-07T21:06:08.934+0000"},{"body":"well, it may be possible that this is the same issue, just very much exacerbated by the batches","created":"2016-04-07T21:07:24.394+0000"},{"body":"[~devdazed] are you using unlogged batches by any chance?","created":"2016-04-07T21:19:36.621+0000"},{"body":"[~pauloricardomg] yes, we are using unlogged batches that cross partitions","created":"2016-04-07T21:31:38.909+0000"},{"body":"[~devdazed] Ok, then it's very likely that you're hit by CASSANDRA-11529. I created another ticket in case this one is a different issue.\n\n[~zznate] [~CRolo] can you double check you're not hitting CASSANDRA-11529?","created":"2016-04-07T21:35:23.812+0000"},{"body":"any chance the number of vnodes in a cluster affects how bad this issue is?","created":"2016-04-07T21:37:00.039+0000"},{"body":"yes, cluster with more vnodes will be more affected by that.","created":"2016-04-07T21:40:50.576+0000"},{"body":"[~pauloricardomg] CASSANDRA-11529 makes sense because CASSANDRA-9303 was backported to 2.1.12 in DSE 4.8.4. Hence why we see it in that version vs only in 2.1.13.","created":"2016-04-07T21:42:31.005+0000"},{"body":"CASSANDRA-11529 makes sense for [~devdazed], but the numbers from [~arodrime] are on a *non-vnode* cluster. \n\nKeep in mind: \n- we actually get this metric to go down a bit by increasing native transport threads\n- we are not CPU bound or spiking (noticeably at least)\n\nIf this was an inefficiency in a hot code path per CASSANDRA-11529, I feel like increasing parallelism would exacerbate the issue. \n\nGood thoughts though - thanks for digging in!","created":"2016-04-07T21:59:19.476+0000"},{"body":"On my description above we where using physical nodes (initial token set, num_token commented, no vnodes enable), fwiw.","created":"2016-04-08T07:47:10.419+0000"},{"body":"On my systems: 3.0.3 -> Logged Batches, \n 2.1.13 -> Unlogged Batches.\n\nBut on 2.1.13 batches where taken out at some point to see if it would improve, but not much improvement was seen. So it might not be related to batches. The test was also small (With batches disabled), so would not make that count as solid test. ","created":"2016-04-08T09:27:30.065+0000"},{"body":"[~CRolo] [~arodrime] if you could record and attach JFR files of servers with high numbers of blocked NTR threads that would be of great help in investigating this issue. (I will also have a look on the already attached files, but they could be affected by CASSANDRA-11529). Thanks in advance!","created":"2016-04-08T20:08:23.719+0000"},{"body":"I will try to get it today, I might able to get those for the 3.0.3 cluster, the 2.1.13 might prove more difficult. I will update this ticket later.","created":"2016-04-11T10:56:23.977+0000"},{"body":"I wasn't able to reproduce this condition so far in a 2.1 [cstar_perf|http://cstar.datastax.com/] cluster with the following spec: 1 stress, 3 Cassandra, each node 2x Intel(R) Xeon(R) CPU E5-2620 v2 @ 2.10GHz (12 cores total), 64G, 3 Samsung SSD 845DC EVO 240GB, mdadm RAID 0.\n\nThe test consistent in the following sequence of stress steps followed by {{nodetool tpstats}}:\n* {{user profile=https://raw.githubusercontent.com/mesosphere/cassandra-mesos/master/driver-extensions/cluster-loadtest/cqlstress-example.yaml ops\\(insert=1\\) n=1M -rate threads=300}}\n* {{user profile=https://raw.githubusercontent.com/mesosphere/cassandra-mesos/master/driver-extensions/cluster-loadtest/cqlstress-example.yaml ops\\(simple1=1\\) n=1M -rate threads=300}}\n* {{user profile=https://raw.githubusercontent.com/mesosphere/cassandra-mesos/master/driver-extensions/cluster-loadtest/cqlstress-example.yaml ops\\(range1=1\\) n=1M -rate threads=300}}\n\nAt the end of 5 runs, the total number of blocked NTR threads was negligible (0 for all runs, except one with 0.004% blocked). I will try running on a larger mixed workload, ramping up the number of stress threads and also try it on 3.0.\n\nMeanwhile, some JFR files, reproduction steps or at least more detailed description on the environment/workload to reproduce this would be greatly appreciated.","created":"2016-04-27T18:12:00.353+0000"},{"body":"[~pauloricardomg] Thanks for continuing to dig into this. However, it looks like you have RF=1 in that stress script:\n{noformat}\nkeyspace_definition: |\n CREATE KEYSPACE stresscql WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 1};\n{noformat}\n\n\n\n\n","created":"2016-04-28T16:03:58.271+0000"},{"body":"Tried reproducing this in a variety of workloads without success in the following environment:\n* 3 nodes m3.xlarge (m3.xlarge, 4 vcpu, 15GB RAM, 2 x 40)\n* 8GB Heap, CMS\n* 100M keys (24GB per node)\n* C* 3.0.3 with default settings and a variation with native_transport_max_threads=10\n* no-vnodes\n\nThe workloads were based a variation of https://raw.githubusercontent.com/mesosphere/cassandra-mesos/master/driver-extensions/cluster-loadtest/cqlstress-example.yaml with 100M keys and RF=3 executed from 2 m3.xlarge stress nodes for 1, 2 and 6 hours with 10, 20 and 30 threads without exhausting the cluster CPU/IO capacity. I tried several combinations of the following workloads:\n* read-only\n* write-only\n* range-only\n* triggering repairs during the execution\n* unthrottling compaction\n\nI recorded and analyzed flight recordings during the tests but didn't find anything suspicious. No blocked native transport threads were verified during tests with above scenarios, so this might indicate that this condition is not a widespread bug like CASSANDRA-11529 but probably some edgy combination of workload, environment and bad scheduling that happens in production but is harder to reproduce with synthetic workloads.\n \nA thread dump of when this condition happens would probably help us detect where is the bottleneck or contention, so I created CASSANDRA-11713 to add ability of logging a thread dump when the thread pool queue is full. If someone could install that patch and enable it in production to capture a thread dump when the blockage happens that would probably help us elucidate what's going on here.","created":"2016-05-04T22:04:37.126+0000"},{"body":"Closing as Cannot Reproduce (and not for lack of effort). Feel free to reopen with concrete reproducible steps, if you are able to come up with them.","created":"2016-05-31T20:19:24.757+0000"},{"body":"The Native Transport Request pool is the only thread pool that has a bounded limit (128)\n\nThe NTR uses the SEPExecutor which effectively blocks till the queue has room.\n\nHowever if I'm reading it correctly the SEPWorker goes into a spin loop for some scenarios when there us no work so perhaps we are hitting some edge case when the tasks are blocked. [~pauloricardomg] perhaps try setting this queue to something small like 4 to force blocking?","created":"2016-06-16T20:17:58.984+0000"},{"body":"I also see this behavior after upgrading to 2.1.14 (from 2.0.17). The JMX counter was at 0 before the upgrade. \nI increased the native max threads to 256 but the NTR blocked were still at 0.6%.\n\nAs for now, the better change was to move memtable offheap (less GC time => less load => less all time blocked NTR). But I still see up to 0.3% of all time blocked...\n\nNote: I see this behavior on two different clusters (hosted on AWS). The load after the upgrade in 2.1 is higher than in 2.0 on the cluster using LCS. This cluster has the worst blocked NTR ratio. \n","created":"2016-07-13T19:46:06.912+0000"},{"body":"[~rha] would you be able to try the attached patch {{thread-queue-2.1.txt}} and see if that helps?","created":"2016-07-18T20:15:33.336+0000"},{"body":"I see a lower NTR blocked percentage with 1024 max queued requests.\n \nI attached {{max_queued_ntr_property.txt}} to set this value in {{cassandra-env.sh}} and it turns out that 1536 was a good value in my case. I don't see any blocked NTR so far. That said it's just a workaround because the root cause might be elsewhere. Anyway I think it's better to have a property to set this value instead of a hard coded number. WDYT?\n\nUPDATE: I had to increase up to {{-Dcassandra.max_queued_native_transport_requests=3072}} on the other DC (same cluster) in order to see 0 blocked NTR.","created":"2016-08-05T17:10:08.915+0000"},{"body":"[~tjake] We are preparing this patch for a deploy to two different cluster: 2.2.6 and 2.1.16. Both clusters can exhibit high-burst workloads (which is why I think you can't reproduce this with the stress tool, and I agree with your suspicions on the cause). We'll let you know how it goes. ","created":"2016-08-09T19:53:04.416+0000"},{"body":"Sorry, but I am re-opening this. Given we have a potential explanation and workaround available and that I still see blocked NTR (to a varying degree) on *every* 2.x+ cluster I have come in contact with recently. ","created":"2016-08-09T19:54:37.334+0000"},{"body":"Great, looks like this is the issue. I guess the question is what is a reasonable default for this? Should we set it high? what's too high? \n\nAny suggestions [~benedict]?","created":"2016-08-09T19:55:02.028+0000"},{"body":"This blocking behaviour and default queue limit was carried forward from the prior code, so I'm afraid I don't have any insights. It may be that the increased baseline performance of 2.1 permits worse outlier states to accumulate if the user exploits it. \n\nThe old code was using the jboss MemoryAwareExecutorService, but estimated the size of each request as 1. A value of 128 does seem very small for users performing very small operations, but conversely a few large reads could destroy the box, so we will have complaints whatever we pick. Perhaps configuring this parameter should be explicitly called out in whatever best practices docs we have. \n\nIdeally, this limit would be removed entirely and better dynamic constraints applied - I think we have some tickets already for keeping the number of requests at a coordinator constrained. If that were dealt with (for all request types), this limit could be removed entirely.","created":"2016-08-10T08:36:05.137+0000"},{"body":"Ok thanks, I think we are best off using the System.property with a larger default like 4096. I'll wait for [~zznate] to confirm this fixes his issue as well...","created":"2016-08-10T19:44:25.789+0000"},{"body":"Is the following setting in the {{cassandra.yaml}} not the same as {{-Dcassandra.max_queued_native_transport_requests}}\n\n{{native_transport_max_threads: 128}}","created":"2016-08-11T14:28:59.049+0000"},{"body":"These are two different things.\n\nThe latter is the number of threads. The former is the max length of the queue'd requests which are allowed to build up before being rejected for those threads.","created":"2016-08-11T14:38:03.476+0000"},{"body":"4096 seems a safe default value. Over time I see some blocked NTR on one DC even with 3072.","created":"2016-08-11T18:00:57.996+0000"},{"body":"[~tjake] I am probably going to give this a try as well early next week, I will keep you posted about results over here using 4096 as it seems people tend to converge to this value.","created":"2016-08-12T15:58:08.314+0000"},{"body":"So after giving this a try it looks like 4096 was not high enough here, but it is already a way better:\n\n{noformat}\n[alain@ip-172-17-xx-xx ~]$ nodetool tpstats | grep -e Native-Transport-Requests -e Blocked\nPool Name Active Pending Completed Blocked All time blocked\nNative-Transport-Requests 1 2 199191 0 27\n[alain@ip-172-17-xx-xx ~]$ nodetool tpstats | grep -e Native-Transport-Requests -e Blocked\nPool Name Active Pending Completed Blocked All time blocked\nNative-Transport-Requests 1 0 507360 0 36\n{noformat}\n\nI would say:\n\n1 - This definitely solve the issue on my side as well\n2 - 4096 is not enough on my specific case but is probably still a good default value. We probably need to smooth operations on our side as well, this workload is very spiky.\n\nI know [~zznate] is giving it a try as well currently, and it looks great so far.\n\nThanks for the patch [~tjake] !","created":"2016-08-17T09:09:16.018+0000"},{"body":"[~arodrime] do you see an improvement in client observed latencies with the value increased? I just want to make sure we have an observable improvement changing this, not just \"the metric that sounds bad went down\". As less things blocked up front could actually make latency worse if it bogs down other things in the system or causes more GC churn than your box can handle.","created":"2016-08-17T14:36:14.401+0000"},{"body":"I will see what I can find. This environment I am working on is not very well monitored. I should be able to give you a latency trend though. Let me check that.","created":"2016-08-17T14:43:11.791+0000"},{"body":"[~tjake] \nThe patch was deployed running with the following setting:\n{noformat}\n-Dcassandra.max_queued_native_transport_requests=1024\n{noformat}\n(EDIT: changed 4096 to 1024 as I pasted the wrong value initially).\n\nAfter an hour of soaking in, it looks good so far:\n{noformat}\nNative-Transport-Requests 5 0 17738639 0 0\n{noformat}\n\nOn this particular cluster, the NTR all-time would start ticking up immediately (but sporadically) after a restart so the above is an excellent sign. The workload contains a large number of very small rows mixed read and write. \n\nThere is no discernable impact on client latencies, nor with GC churn. If anything, both are marginally smaller. ","created":"2016-08-18T03:48:07.154+0000"},{"body":"Committed to: 2.1.16, 2.2.8, 3.0.9, 3.10, trunk.","created":"2016-09-20T02:52:02.933+0000"},{"body":"For future reference this went into 3.0.10 not 3.0.9","created":"2016-09-27T23:59:46.056+0000"},{"body":"I'm using Apache Cassandra™ 3.0.12.1656 which packages with DSE 5.0.8. I'm still seeing high numbers native request blocked. Should i add the *-Dcassandra.max_queued_native_transport_requests=1024* or its already in the patch because ,im seeing this issue.","created":"2017-08-30T14:34:44.561+0000"},{"body":"[~sadagopan88] When using Open Source Apache Cassandra you have to specify it in {{cassandra-env.sh}}:\n{code}\nJVM_OPTS=\"$JVM_OPTS -Dcassandra.max_queued_native_transport_requests=1024\"\n{code}\nI don't know if DSE set this setting to something different than default value (128). You can ask to DataStax.","created":"2017-08-30T15:46:16.708+0000"},{"body":"Thanks it worked!. ","created":"2017-09-06T16:23:09.153+0000"},{"body":"I have same issue in open source cassandra 3.8, I have high blocked Native-transport-requests in nodetool tpstats. I also see higher client latencies on co-ordinator when the local latencies are much low and no network issues. I am not sure if the client latencies are related to this. I am currently using default of 128 threads.\r\n\r\n!tablestats.png!!tpstats.png!!tablehistograms.png!\r\n\r\n!proxyhistograms.png!","created":"2018-06-06T19:14:12.824+0000"}],"conversations":[{"body":"When upgrading from 2.1.9 to 2.1.13, we are witnessing an issue where the machine load increases to very high levels (> 120 on an 8 core machine) and native transport requests get blocked in tpstats.\n\nI was able to reproduce this in both CMS and G1GC as well as on JVM 7 and 8.\n\nThe issue does not seem to affect the nodes running 2.1.9.\n\nThe issue seems to coincide with the number of connections OR the number of total requests being processed at a given time (as the latter increases with the former in our system)\n\nCurrently there is between 600 and 800 client connections on each machine and each machine is handling roughly 2000-3000 client requests per second.\n\nDisabling the binary protocol fixes the issue for this node but isn't a viable option cluster-wide.\n\nHere is the output from tpstats:\n\n{code}\nPool Name Active Pending Completed Blocked All time blocked\nMutationStage 0 8 8387821 0 0\nReadStage 0 0 355860 0 0\nRequestResponseStage 0 7 2532457 0 0\nReadRepairStage 0 0 150 0 0\nCounterMutationStage 32 104 897560 0 0\nMiscStage 0 0 0 0 0\nHintedHandoff 0 0 65 0 0\nGossipStage 0 0 2338 0 0\nCacheCleanupExecutor 0 0 0 0 0\nInternalResponseStage 0 0 0 0 0\nCommitLogArchiver 0 0 0 0 0\nCompactionExecutor 2 190 474 0 0\nValidationExecutor 0 0 0 0 0\nMigrationStage 0 0 10 0 0\nAntiEntropyStage 0 0 0 0 0\nPendingRangeCalculator 0 0 310 0 0\nSampler 0 0 0 0 0\nMemtableFlushWriter 1 10 94 0 0\nMemtablePostFlush 1 34 257 0 0\nMemtableReclaimMemory 0 0 94 0 0\nNative-Transport-Requests 128 156 387957 16 278451\n\nMessage type Dropped\nREAD 0\nRANGE_SLICE 0\n_TRACE 0\nMUTATION 0\nCOUNTER_MUTATION 0\nBINARY 0\nREQUEST_RESPONSE 0\nPAGED_RANGE 0\nREAD_REPAIR 0\n{code}\n\nAttached is the jstack output for both CMS and G1GC.\n\nFlight recordings are here:\nhttps://s3.amazonaws.com/simple-logs/cassandra-102-cms.jfr\nhttps://s3.amazonaws.com/simple-logs/cassandra-102-g1gc.jfr\n\nIt is interesting to note that while the flight recording was taking place, the load on the machine went back to healthy, and when the flight recording finished the load went back to > 100.","from":"reporter","subject":"High Blocked NTR When Connecting"},{"body":"Also reproducible in 3.0.3.\n\nI have 2 clusters showing exactly the same problem, one running 2.1.13 (5 nodes, 4 cores, 32GB Ram, CMS Java 7) and another 3 nodes, 4cores, 16GB Ram running 3.0.3 (CMS Java 8). \n\nBoth get the problems described: \"The issue seems to coincide with the number of connections OR the number of total requests being processed at a given time (as the latter increases with the former in our system)\"\n\nIt is normal for OS Load to shoot to 33+ with around 40connected clients.","from":"developer"},{"body":"For troubleshooting we set up a coordinator-only node and pointed one app server at it. This resulted in roughly 90 connections to the node. We witnessed many timeouts of requests from the app server's perspective. We downgraded the coordinator-only node to 2.1.9 and upgraded point-release by point-release (DSE point releases) until we saw the same behavior in DSE 4.8.4 (Cassandra 2.1.12).\n\nI'm not certain this has to do with connection count anymore. \n\nWe have several different types of workloads going on, but we found that only the workloads that use batches were timing out. Additionally this is only happening when the node is used as a coordinator. We are not seeing this issue when we disable the binary protocol, effectively making the node no-longer a coordinator.","from":"developer"},{"body":"I can confirm that in both my clusters batches are in use.","from":"developer"},{"body":"I also observed in C*2.1.12 that a certain percentage of the Native-Transport-Requests are blocked, yet no major CPU or resources issue on my side though, it might then not be related.\n\nFor what it is worth here is something I observed about Native-Transport-Requests: increasing the 'native_transport_max_threads' value help mitigating this, as expected, but Native-Transport-Requests number is still a non zero value.\n\n{noformat}\n[alain~]$ knife ssh \"role:cassandra\" \"nodetool tpstats | grep Native-Transport-Requests\" | grep -e server1 -e server2 -e server3 -e server4 | sort | awk 'BEGIN { printf \"%50s %10s\",\"Server |\",\" Blocked ratio:\\n\" } { printf \"%50s %10f%\\n\", $1, (($7/$5)*100) }'\n Server | Blocked ratio:\n server3 0.044902%\n server4 0.030127%\n server2 0.045759%\n server1 0.082763%\n{noformat}\n\nI waited long enough between the change and the result capture, many days, probably a few weeks. As all the nodes are in the same datacenter, under a (fairly) balanced load, this is probably relevant.\n\nHere are the result for those nodes, in our use case.\n\n||Server||native_transport_max_threads||Percentage of blocked threads||\n|server1|128|0.082763%|\n|server2|384|0.044902%|\n|server3|512|0.045759%|\n|server4|1024|0.030127%|\n\nAlso from the mailing list outputs, it looks like it is quite common to have some Native-Transport-Requests blocked, probably unavoidable depending on the network and use cases (spiky workloads?). ","from":"developer"},{"body":"Raised this to critical. \n\nYou have four long-time users with multiple large deployments who are seeing a quantifiable percentage of client errors on non-resource constrained clusters across multiple versions. ","from":"developer"},{"body":"fixed by CASSANDRA-10200 ?","from":"developer"},{"body":"I went through 2.1.12 changes and didn't find anything suspicious. On 2.1.13 and 3.0.3 though, we changed the {{ServerConnection}} query state map from a {{NonBlockingHashMap}} to a {{ConcurrentHashMap}} on CASSANDRA-10938, which might be misbehaving for some reason.\n\nIs anyone willing to try the revert patch below on 2.1.13 or 3.0.3 and check if that changes anything?\n\n{noformat}\ndiff --git a/src/java/org/apache/cassandra/transport/ServerConnection.java b/src/java/org/apache/cassandra/transport/ServerConnection.java\nindex ce4d164..5991b33 100644\n--- a/src/java/org/apache/cassandra/transport/ServerConnection.java\n+++ b/src/java/org/apache/cassandra/transport/ServerConnection.java\n@@ -17,7 +17,6 @@\n */\n package org.apache.cassandra.transport;\n \n-import java.util.concurrent.ConcurrentHashMap;\n import java.util.concurrent.ConcurrentMap;\n \n import io.netty.channel.Channel;\n@@ -29,6 +28,8 @@ import org.apache.cassandra.config.DatabaseDescriptor;\n import org.apache.cassandra.service.ClientState;\n import org.apache.cassandra.service.QueryState;\n \n+import org.cliffc.high_scale_lib.NonBlockingHashMap;\n+\n public class ServerConnection extends Connection\n {\n private enum State { UNINITIALIZED, AUTHENTICATION, READY }\n@@ -37,7 +38,7 @@ public class ServerConnection extends Connection\n private final ClientState clientState;\n private volatile State state;\n \n- private final ConcurrentMap queryStates = new ConcurrentHashMap<>();\n+ private final ConcurrentMap queryStates = new NonBlockingHashMap<>();\n \n public ServerConnection(Channel channel, int version, Connection.Tracker tracker)\n {\n{noformat}","from":"developer"},{"body":"Unfortunately we are on DSE, so I can't run the revert","from":"developer"},{"body":"[~pauloricardomg] unfortunately, this may be been latent for some time. CASSANDRA-10044 re-introduced the tpstats counters as they had been missing (which is mostly likely why this has not been noticed until recently). ","from":"developer"},{"body":"that's right, that was just a wild guess. I focused my quick investigation on 2.1.12 and 2.1.13 commits, but some deeper investigation is probably needed.\n\nprobably it's a good idea to backport CASSANDRA-10044 to versions < 2.1.12 and setup a coordinator-only node and check where it started happening to narrow down the scope.","from":"developer"},{"body":"we may have two separate issues here, in mine, the issue is 100% CPU utilization and ultra high load when using batches. According to the jfr all of hot threads are spinning on \n{code}\norg.apache.cassandra.locator.NetworkTopologyStrategy.hasSufficientReplicas(String, Map, Multimap)\n -> org.apache.cassandra.locator.NetworkTopologyStrategy.hasSufficientReplicas(Map, Multimap)\n -> org.apache.cassandra.locator.NetworkTopologyStrategy.calculateNaturalEndpoints(Token, TokenMetadata)\n{code}","from":"developer"},{"body":"Yeah, that is quite different that what [~arodrime] and I have seen recently. Nodes in our case were otherwise well within utilization thresholds.\n\n[~devdazed] Your issue looks like it would be addressed by [~cnlwsu]'s reference to CASSANDRA-10200 (lean on DSE folks for a patch). ","from":"developer"},{"body":"well, it may be possible that this is the same issue, just very much exacerbated by the batches","from":"developer"},{"body":"[~devdazed] are you using unlogged batches by any chance?","from":"developer"},{"body":"[~pauloricardomg] yes, we are using unlogged batches that cross partitions","from":"developer"},{"body":"[~devdazed] Ok, then it's very likely that you're hit by CASSANDRA-11529. I created another ticket in case this one is a different issue.\n\n[~zznate] [~CRolo] can you double check you're not hitting CASSANDRA-11529?","from":"developer"},{"body":"any chance the number of vnodes in a cluster affects how bad this issue is?","from":"developer"},{"body":"yes, cluster with more vnodes will be more affected by that.","from":"developer"},{"body":"[~pauloricardomg] CASSANDRA-11529 makes sense because CASSANDRA-9303 was backported to 2.1.12 in DSE 4.8.4. Hence why we see it in that version vs only in 2.1.13.","from":"developer"},{"body":"CASSANDRA-11529 makes sense for [~devdazed], but the numbers from [~arodrime] are on a *non-vnode* cluster. \n\nKeep in mind: \n- we actually get this metric to go down a bit by increasing native transport threads\n- we are not CPU bound or spiking (noticeably at least)\n\nIf this was an inefficiency in a hot code path per CASSANDRA-11529, I feel like increasing parallelism would exacerbate the issue. \n\nGood thoughts though - thanks for digging in!","from":"developer"},{"body":"On my description above we where using physical nodes (initial token set, num_token commented, no vnodes enable), fwiw.","from":"developer"},{"body":"On my systems: 3.0.3 -> Logged Batches, \n 2.1.13 -> Unlogged Batches.\n\nBut on 2.1.13 batches where taken out at some point to see if it would improve, but not much improvement was seen. So it might not be related to batches. The test was also small (With batches disabled), so would not make that count as solid test. ","from":"developer"},{"body":"[~CRolo] [~arodrime] if you could record and attach JFR files of servers with high numbers of blocked NTR threads that would be of great help in investigating this issue. (I will also have a look on the already attached files, but they could be affected by CASSANDRA-11529). Thanks in advance!","from":"developer"},{"body":"I will try to get it today, I might able to get those for the 3.0.3 cluster, the 2.1.13 might prove more difficult. I will update this ticket later.","from":"developer"},{"body":"I wasn't able to reproduce this condition so far in a 2.1 [cstar_perf|http://cstar.datastax.com/] cluster with the following spec: 1 stress, 3 Cassandra, each node 2x Intel(R) Xeon(R) CPU E5-2620 v2 @ 2.10GHz (12 cores total), 64G, 3 Samsung SSD 845DC EVO 240GB, mdadm RAID 0.\n\nThe test consistent in the following sequence of stress steps followed by {{nodetool tpstats}}:\n* {{user profile=https://raw.githubusercontent.com/mesosphere/cassandra-mesos/master/driver-extensions/cluster-loadtest/cqlstress-example.yaml ops\\(insert=1\\) n=1M -rate threads=300}}\n* {{user profile=https://raw.githubusercontent.com/mesosphere/cassandra-mesos/master/driver-extensions/cluster-loadtest/cqlstress-example.yaml ops\\(simple1=1\\) n=1M -rate threads=300}}\n* {{user profile=https://raw.githubusercontent.com/mesosphere/cassandra-mesos/master/driver-extensions/cluster-loadtest/cqlstress-example.yaml ops\\(range1=1\\) n=1M -rate threads=300}}\n\nAt the end of 5 runs, the total number of blocked NTR threads was negligible (0 for all runs, except one with 0.004% blocked). I will try running on a larger mixed workload, ramping up the number of stress threads and also try it on 3.0.\n\nMeanwhile, some JFR files, reproduction steps or at least more detailed description on the environment/workload to reproduce this would be greatly appreciated.","from":"developer"},{"body":"[~pauloricardomg] Thanks for continuing to dig into this. However, it looks like you have RF=1 in that stress script:\n{noformat}\nkeyspace_definition: |\n CREATE KEYSPACE stresscql WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 1};\n{noformat}\n\n\n\n\n","from":"developer"},{"body":"Tried reproducing this in a variety of workloads without success in the following environment:\n* 3 nodes m3.xlarge (m3.xlarge, 4 vcpu, 15GB RAM, 2 x 40)\n* 8GB Heap, CMS\n* 100M keys (24GB per node)\n* C* 3.0.3 with default settings and a variation with native_transport_max_threads=10\n* no-vnodes\n\nThe workloads were based a variation of https://raw.githubusercontent.com/mesosphere/cassandra-mesos/master/driver-extensions/cluster-loadtest/cqlstress-example.yaml with 100M keys and RF=3 executed from 2 m3.xlarge stress nodes for 1, 2 and 6 hours with 10, 20 and 30 threads without exhausting the cluster CPU/IO capacity. I tried several combinations of the following workloads:\n* read-only\n* write-only\n* range-only\n* triggering repairs during the execution\n* unthrottling compaction\n\nI recorded and analyzed flight recordings during the tests but didn't find anything suspicious. No blocked native transport threads were verified during tests with above scenarios, so this might indicate that this condition is not a widespread bug like CASSANDRA-11529 but probably some edgy combination of workload, environment and bad scheduling that happens in production but is harder to reproduce with synthetic workloads.\n \nA thread dump of when this condition happens would probably help us detect where is the bottleneck or contention, so I created CASSANDRA-11713 to add ability of logging a thread dump when the thread pool queue is full. If someone could install that patch and enable it in production to capture a thread dump when the blockage happens that would probably help us elucidate what's going on here.","from":"developer"},{"body":"Closing as Cannot Reproduce (and not for lack of effort). Feel free to reopen with concrete reproducible steps, if you are able to come up with them.","from":"developer"},{"body":"The Native Transport Request pool is the only thread pool that has a bounded limit (128)\n\nThe NTR uses the SEPExecutor which effectively blocks till the queue has room.\n\nHowever if I'm reading it correctly the SEPWorker goes into a spin loop for some scenarios when there us no work so perhaps we are hitting some edge case when the tasks are blocked. [~pauloricardomg] perhaps try setting this queue to something small like 4 to force blocking?","from":"developer"},{"body":"I also see this behavior after upgrading to 2.1.14 (from 2.0.17). The JMX counter was at 0 before the upgrade. \nI increased the native max threads to 256 but the NTR blocked were still at 0.6%.\n\nAs for now, the better change was to move memtable offheap (less GC time => less load => less all time blocked NTR). But I still see up to 0.3% of all time blocked...\n\nNote: I see this behavior on two different clusters (hosted on AWS). The load after the upgrade in 2.1 is higher than in 2.0 on the cluster using LCS. This cluster has the worst blocked NTR ratio. \n","from":"developer"},{"body":"[~rha] would you be able to try the attached patch {{thread-queue-2.1.txt}} and see if that helps?","from":"developer"},{"body":"I see a lower NTR blocked percentage with 1024 max queued requests.\n \nI attached {{max_queued_ntr_property.txt}} to set this value in {{cassandra-env.sh}} and it turns out that 1536 was a good value in my case. I don't see any blocked NTR so far. That said it's just a workaround because the root cause might be elsewhere. Anyway I think it's better to have a property to set this value instead of a hard coded number. WDYT?\n\nUPDATE: I had to increase up to {{-Dcassandra.max_queued_native_transport_requests=3072}} on the other DC (same cluster) in order to see 0 blocked NTR.","from":"developer"},{"body":"[~tjake] We are preparing this patch for a deploy to two different cluster: 2.2.6 and 2.1.16. Both clusters can exhibit high-burst workloads (which is why I think you can't reproduce this with the stress tool, and I agree with your suspicions on the cause). We'll let you know how it goes. ","from":"developer"},{"body":"Sorry, but I am re-opening this. Given we have a potential explanation and workaround available and that I still see blocked NTR (to a varying degree) on *every* 2.x+ cluster I have come in contact with recently. ","from":"developer"},{"body":"Great, looks like this is the issue. I guess the question is what is a reasonable default for this? Should we set it high? what's too high? \n\nAny suggestions [~benedict]?","from":"developer"},{"body":"This blocking behaviour and default queue limit was carried forward from the prior code, so I'm afraid I don't have any insights. It may be that the increased baseline performance of 2.1 permits worse outlier states to accumulate if the user exploits it. \n\nThe old code was using the jboss MemoryAwareExecutorService, but estimated the size of each request as 1. A value of 128 does seem very small for users performing very small operations, but conversely a few large reads could destroy the box, so we will have complaints whatever we pick. Perhaps configuring this parameter should be explicitly called out in whatever best practices docs we have. \n\nIdeally, this limit would be removed entirely and better dynamic constraints applied - I think we have some tickets already for keeping the number of requests at a coordinator constrained. If that were dealt with (for all request types), this limit could be removed entirely.","from":"developer"},{"body":"Ok thanks, I think we are best off using the System.property with a larger default like 4096. I'll wait for [~zznate] to confirm this fixes his issue as well...","from":"developer"},{"body":"Is the following setting in the {{cassandra.yaml}} not the same as {{-Dcassandra.max_queued_native_transport_requests}}\n\n{{native_transport_max_threads: 128}}","from":"developer"},{"body":"These are two different things.\n\nThe latter is the number of threads. The former is the max length of the queue'd requests which are allowed to build up before being rejected for those threads.","from":"developer"},{"body":"4096 seems a safe default value. Over time I see some blocked NTR on one DC even with 3072.","from":"developer"},{"body":"[~tjake] I am probably going to give this a try as well early next week, I will keep you posted about results over here using 4096 as it seems people tend to converge to this value.","from":"developer"},{"body":"So after giving this a try it looks like 4096 was not high enough here, but it is already a way better:\n\n{noformat}\n[alain@ip-172-17-xx-xx ~]$ nodetool tpstats | grep -e Native-Transport-Requests -e Blocked\nPool Name Active Pending Completed Blocked All time blocked\nNative-Transport-Requests 1 2 199191 0 27\n[alain@ip-172-17-xx-xx ~]$ nodetool tpstats | grep -e Native-Transport-Requests -e Blocked\nPool Name Active Pending Completed Blocked All time blocked\nNative-Transport-Requests 1 0 507360 0 36\n{noformat}\n\nI would say:\n\n1 - This definitely solve the issue on my side as well\n2 - 4096 is not enough on my specific case but is probably still a good default value. We probably need to smooth operations on our side as well, this workload is very spiky.\n\nI know [~zznate] is giving it a try as well currently, and it looks great so far.\n\nThanks for the patch [~tjake] !","from":"developer"},{"body":"[~arodrime] do you see an improvement in client observed latencies with the value increased? I just want to make sure we have an observable improvement changing this, not just \"the metric that sounds bad went down\". As less things blocked up front could actually make latency worse if it bogs down other things in the system or causes more GC churn than your box can handle.","from":"developer"},{"body":"I will see what I can find. This environment I am working on is not very well monitored. I should be able to give you a latency trend though. Let me check that.","from":"developer"},{"body":"[~tjake] \nThe patch was deployed running with the following setting:\n{noformat}\n-Dcassandra.max_queued_native_transport_requests=1024\n{noformat}\n(EDIT: changed 4096 to 1024 as I pasted the wrong value initially).\n\nAfter an hour of soaking in, it looks good so far:\n{noformat}\nNative-Transport-Requests 5 0 17738639 0 0\n{noformat}\n\nOn this particular cluster, the NTR all-time would start ticking up immediately (but sporadically) after a restart so the above is an excellent sign. The workload contains a large number of very small rows mixed read and write. \n\nThere is no discernable impact on client latencies, nor with GC churn. If anything, both are marginally smaller. ","from":"developer"},{"body":"Committed to: 2.1.16, 2.2.8, 3.0.9, 3.10, trunk.","from":"developer"},{"body":"For future reference this went into 3.0.10 not 3.0.9","from":"developer"},{"body":"I'm using Apache Cassandra™ 3.0.12.1656 which packages with DSE 5.0.8. I'm still seeing high numbers native request blocked. Should i add the *-Dcassandra.max_queued_native_transport_requests=1024* or its already in the patch because ,im seeing this issue.","from":"developer"},{"body":"[~sadagopan88] When using Open Source Apache Cassandra you have to specify it in {{cassandra-env.sh}}:\n{code}\nJVM_OPTS=\"$JVM_OPTS -Dcassandra.max_queued_native_transport_requests=1024\"\n{code}\nI don't know if DSE set this setting to something different than default value (128). You can ask to DataStax.","from":"developer"},{"body":"Thanks it worked!. ","from":"developer"},{"body":"I have same issue in open source cassandra 3.8, I have high blocked Native-transport-requests in nodetool tpstats. I also see higher client latencies on co-ordinator when the local latencies are much low and no network issues. I am not sure if the client latencies are related to this. I am currently using default of 128 threads.\r\n\r\n!tablestats.png!!tpstats.png!!tablehistograms.png!\r\n\r\n!proxyhistograms.png!","from":"developer"}],"created":"2016-03-16T18:11:49.000+0000","description":"When upgrading from 2.1.9 to 2.1.13, we are witnessing an issue where the machine load increases to very high levels (> 120 on an 8 core machine) and native transport requests get blocked in tpstats.\n\nI was able to reproduce this in both CMS and G1GC as well as on JVM 7 and 8.\n\nThe issue does not seem to affect the nodes running 2.1.9.\n\nThe issue seems to coincide with the number of connections OR the number of total requests being processed at a given time (as the latter increases with the former in our system)\n\nCurrently there is between 600 and 800 client connections on each machine and each machine is handling roughly 2000-3000 client requests per second.\n\nDisabling the binary protocol fixes the issue for this node but isn't a viable option cluster-wide.\n\nHere is the output from tpstats:\n\n{code}\nPool Name Active Pending Completed Blocked All time blocked\nMutationStage 0 8 8387821 0 0\nReadStage 0 0 355860 0 0\nRequestResponseStage 0 7 2532457 0 0\nReadRepairStage 0 0 150 0 0\nCounterMutationStage 32 104 897560 0 0\nMiscStage 0 0 0 0 0\nHintedHandoff 0 0 65 0 0\nGossipStage 0 0 2338 0 0\nCacheCleanupExecutor 0 0 0 0 0\nInternalResponseStage 0 0 0 0 0\nCommitLogArchiver 0 0 0 0 0\nCompactionExecutor 2 190 474 0 0\nValidationExecutor 0 0 0 0 0\nMigrationStage 0 0 10 0 0\nAntiEntropyStage 0 0 0 0 0\nPendingRangeCalculator 0 0 310 0 0\nSampler 0 0 0 0 0\nMemtableFlushWriter 1 10 94 0 0\nMemtablePostFlush 1 34 257 0 0\nMemtableReclaimMemory 0 0 94 0 0\nNative-Transport-Requests 128 156 387957 16 278451\n\nMessage type Dropped\nREAD 0\nRANGE_SLICE 0\n_TRACE 0\nMUTATION 0\nCOUNTER_MUTATION 0\nBINARY 0\nREQUEST_RESPONSE 0\nPAGED_RANGE 0\nREAD_REPAIR 0\n{code}\n\nAttached is the jstack output for both CMS and G1GC.\n\nFlight recordings are here:\nhttps://s3.amazonaws.com/simple-logs/cassandra-102-cms.jfr\nhttps://s3.amazonaws.com/simple-logs/cassandra-102-g1gc.jfr\n\nIt is interesting to note that while the flight recording was taking place, the load on the machine went back to healthy, and when the flight recording finished the load went back to > 100.","issue_id":"12950933","key":"CASSANDRA-11363","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-09-20T02:52:02.000+0000","role":"fixed_distractor","summary":"High Blocked NTR When Connecting"} {"case_id":"12951429","cluster":"DISTRACTOR-CASSANDRA-11381","comments":[{"body":"Is this ticket Patch Available?","created":"2016-03-18T15:46:22.081+0000"},{"body":"[~JoshuaMcKenzie], I've only run unit tests against trunk. And I'm trying to figure out how to write a dtest for this (since unit tests and system property don't mix (afaik)).\n\nIf you're happy to proceed without the dtest, say and I'll 'submit patch'. And I do want to double-check that it's not possible to create auth tables without even the system tables having been created (eg bootstrap & join_ring=false).","created":"2016-03-18T20:09:36.759+0000"},{"body":"bq. trying to figure out how to write a dtest for this\nI say we hold off until you've worked out a dtest for it. Just wanted to make sure there wasn't a patch waiting for it and it slip through the cracks.","created":"2016-03-21T17:17:38.350+0000"},{"body":"[~JoshuaMcKenzie], the dtest is attached.","created":"2016-03-23T01:57:46.884+0000"},{"body":"Small update to the dtest. I had been writing dtests in the trunk branch instead of master. More info [here|https://github.com/pcmanus/ccm/pull/479] and [here|https://github.com/riptano/cassandra-dtest/pull/892].","created":"2016-03-27T09:58:34.627+0000"},{"body":"[~JoshuaMcKenzie], [~jkni], what's up?","created":"2016-04-20T05:54:19.235+0000"},{"body":"No fault of Josh here, just a bit of a review backlog on my end. Promise I'll get to this today, sorry for the delay [~mck].","created":"2016-04-20T14:22:11.320+0000"},{"body":"Thanks for the patches [~mck] and apologies for the delay again. This is at the top of my queue now, so I can work on getting this in quickly.\n\nI can confirm that this is a genuine issue - a few thoughts on the patches:\n* Is there a reason that we only setup auth in {{StorageService.initServer}} if the saved tokens aren't empty? This doesn't handle the case when a node starts with join_ring false on its first boot (a coordinator-only node). This can be reproduced with a small variant of your dtest in which we do not start node3 in the initial preparation. Some experimentation on my end suggests this would be fine.\n* On a similar note, is there a reason we can't just move auth setup slightly earlier in {{StorageService.initServer()}}? If we moved it before the check of the join_ring property, we could do without the idempotency parts of the patches. Again, this looks workable in my experimentation.\n* Was there any method by which you arrived at the 300 second timeouts for the cql connection on the dtests? This really drags the test on in failing cases, and I was able to reduce it to 30 seconds without any false negatives.\n* Tests handling the coordinator-only case and using JMX to join the node to the ring would be great. I think a minimal parameterized test could handle these all these permutations.","created":"2016-04-21T06:37:55.371+0000"},{"body":"{quote} - Is there a reason that we only setup auth in StorageService.initServer if the saved tokens aren't empty? This doesn't handle the case when a node starts with join_ring false on its first boot (a coordinator-only node). This can be reproduced with a small variant of your dtest in which we do not start node3 in the initial preparation. Some experimentation on my end suggests this would be fine.{quote}\n\nGood point. Will fix. There's a similar fault in the trunk patch in regards to {{isSurveyMode}}.\n\n{quote} - On a similar note, is there a reason we can't just move auth setup slightly earlier in StorageService.initServer()? If we moved it before the check of the join_ring property, we could do without the idempotency parts of the patches. Again, this looks workable in my experimentation.{quote}\n\nThat did seem like the cleaner solution. But my understanding on the auth model and setup isn't solid, so my intention here was to offer the less-risk patch. I still suspect this (less-risk approach) is the smart approach (for a patch from me) against the non-trunk patches.\n\nI'll take a look at moving the auth setup to an earlier place in just the trunk patch, ok? It seems to make sense to move the call in {{prepareToJoin()}}.\n\n{quote} - Was there any method by which you arrived at the 300 second timeouts for the cql connection on the dtests? This really drags the test on in failing cases, and I was able to reduce it to 30 seconds without any false negatives.{quote}\n\nNo. Just playing it safe for my dual-core thinkpad. (I did have to bump it up once).\n\n{quote} - Tests handling the coordinator-only case and using JMX to join the node to the ring would be great. I think a minimal parameterized test could handle these all these permutations.{quote}\n\nI'll take a look.","created":"2016-04-26T11:11:00.327+0000"},{"body":"[~mck], I definitely understand the inclination toward a lower-risk approach. [~beobal], as resident person-who-I-know-knows-how-auth-works, do you have any time to weigh in on whether moving auth setup earlier {{StorageService.initServer()}} is a risky proposition for any reason? I've done some research and couldn't find any cause for concern. In particular, this ticket would be simpler to fix if auth was moved before join_ring on 2.1/2.2/3.0.x and before start_gossip on trunk.","created":"2016-04-26T14:23:04.941+0000"},{"body":"Auth setup has to wait until the node has joined the ring in order to be able to use the auth tables, as {{system_auth}} is distributed. The same applies to the other pseudo-system keyspaces {{system_traces}} & {{system_distributed}} which is why those are also created in {{joinTokenRing}}. ","created":"2016-04-26T15:13:44.471+0000"},{"body":"On top of that, custom authenticator/authorizer/rolemanager impls may well use their own tables (& possibly other resources), which would obviously be user-defined from a system perspective and so also not accessible before joining.","created":"2016-04-26T15:16:26.267+0000"},{"body":"Giving it some further thought, I think that the moving {{doAuthSetup}} to before joining is only going to be problematic for the first node in a fresh cluster, and then only if it doesn't join (at least in 3.0+). Starting a new coordinator-only node in an existing cluster should be fine, and re-starting an existing node without joining should also be ok, as long as the replication of {{system_auth}} can cope with it. \n","created":"2016-04-26T17:18:34.427+0000"},{"body":"I've tested some scenarios and can confirm that the scenario Sam described above is problematic.\n\nAnother concern I have is potentially performing auth setup before a node is added to its own TokenMetadata in the join_ring false case. It seems to me that this might cause problems with our own implementations and could feasibly cause problems for 3rd party implementations, so I think the original approach of handling auth setup down both branches may be worth pursuing here in the interest of safety.\n\nOn trunk, we have another property \"start_gossip\" that stops {{StorageService.initServer}} even earlier - I don't think we can support auth at this point, so I don't care about supporting it in this ticket (my earlier comment hinted at the fact that I was considering whether it was in scope).","created":"2016-04-27T04:12:12.293+0000"},{"body":"Coming back to this issue now.\n\n{quote}I've tested some scenarios and can confirm that the scenario Sam described above is problematic.{quote}\n[~jkni], which scenario Sam describes?\n\n{quote}...so I think the original approach of handling auth setup down both branches may be worth pursuing here in the interest of safety.{quote}\n\nSo I've uploaded new patches for 2.1 and dtest that \n - set up auth also when tokens are empty\n - add a dtest to test ^^\n\nWorking on the other patches.","created":"2016-08-03T10:35:35.982+0000"},{"body":"…and reduced the timeouts in the dtest to 30 seconds.\n\nAll patches are updated.","created":"2016-08-03T12:12:25.987+0000"},{"body":"[~mck] - The scenario I was referring to was moving {{doAuthSetup}} to before joining for the first node in a fresh cluster.\n\nThanks for keeping this moving. This is near the top of my review queue and I should be able to get to it soon.","created":"2016-08-03T19:24:55.555+0000"},{"body":"ping. is there anything outstanding here?","created":"2016-10-07T01:15:59.975+0000"},{"body":"[~jkni] ping…? ","created":"2016-11-07T04:02:32.113+0000"},{"body":"Thanks for pinging me on this - as you suspected, it slipped through the cracks. \n\nOn reviewing the final version of this patch, I found one problem with the 2.2+ patches. The proposed patch technically breaks the documented {{IRoleManager}}, {{IAuthenticator}}, and {{IAuthorizer}} interfaces. With the implementation given, {{doAuthSetup}} will be called twice for a node started with {{join_ring=False}}, so {{setup()}} will be called twice for the role manager, authenticator, and authorizer. In the documentation for these three public interfaces, we state that {{setup()}} will only be called once after starting a node. I think we should preserve this documented behavior. While slightly less elegant, I think we should instead track whether we've run {{doAuthSetup}} and not repeat this call for a node started with {{join_ring=False}} that is asked to join. This means the parts of the patch implementing idempotency for the MigrationManager listener registration become unnecessary.","created":"2016-11-07T20:10:08.764+0000"},{"body":"Makes sense, will update patches…","created":"2016-11-16T20:22:32.463+0000"},{"body":"[~jkni], all patches (but 2.1) have been updated.","created":"2016-11-17T19:59:34.279+0000"},{"body":"Thanks - the patches look good and I put them through CI.\n\n||branch||testall||dtest||\n|[CASSANDRA-11381-2.2|https://github.com/jkni/cassandra/tree/CASSANDRA-11381-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-2.2-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-2.2-dtest]|\n|[CASSANDRA-11381-3.0|https://github.com/jkni/cassandra/tree/CASSANDRA-11381-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-3.0-dtest]|\n|[CASSANDRA-11381-3.X|https://github.com/jkni/cassandra/tree/CASSANDRA-11381-3.X]|[testall|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-3.X-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-3.X-dtest]|\n|[CASSANDRA-11381-trunk|https://github.com/jkni/cassandra/tree/CASSANDRA-11381-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-trunk-dtest]|\n\nCI looks good for the most part, and I checked that your added dtest passes on all branches. CI revealed one small problem - when a fresh node is started with join_ring=False that has no tokens for other nodes discovered through gossip and no saved tokens, it hits an AssertionError in {{CassandraRoleManager}} setup that is not handled and gets logged as an error by a top level error handler, as seen [here|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-2.2-dtest/1/testReport/junit/topology_test/TestTopology/do_not_join_ring_test/]. In this specific test case, this behavior is hit because a single node cluster is started with join_ring=False. Since this prevents setup from being retried within the {{CassandraRoleManager}}, it seems to me that it is probably worth checking for an absence of tokens in {{CassandraRoleManager.setupDefaultRole}} and throwing a catchable exception/printing a warning so that setup can be retried. What do you think? There may be another alternative I haven't considered.","created":"2016-11-29T04:35:31.335+0000"},{"body":"[~jkni], i have finally come back to this…\n\n{quote}CI revealed one small problem - when a fresh node is started with join_ring=False that has no tokens for other nodes discovered through gossip and no saved tokens, it hits an AssertionError in CassandraRoleManager setup that is not handled{quote}\n\n{{CassandraRoleManager.setupDefaultRole()}} now throws an {{IllegalStateException}} if there are no known tokens in the ring. This permits {{CassandraRoleManager.scheduleSetupTask(..)}} to retry.\n\nThe {{topology_test.py:TestTopology.do_not_join_ring_test}} dtest looks ok now, locally.\n\n\\\\\n\n|| branch || testall || dtest ||\n| [cassandra-2.2_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-2.2_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-2.2_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/40/] |\n| [cassandra-3.0_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-3.0_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-3.0_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/41/] |\n| [cassandra-3.11_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-3.11_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-3.11_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/42/] |\n| [trunk_11381|https://github.com/michaelsembwever/cassandra/tree/mck/trunk_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Ftrunk_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/43/] |\n\n(the asf jenkins dtests take a while, and won't include the new [dtest|https://github.com/michaelsembwever/cassandra-dtest/tree/mck/master_11381])","created":"2017-05-03T08:31:25.780+0000"},{"body":"[~jkni], dtests look good. are we ready to commit this?","created":"2017-05-10T03:43:00.702+0000"},{"body":"Thanks - this looks just about good. A few comments:\n* The 3.0 branch looks like it is an older version of the patch than the 2.2, 3.11, and trunk patches - it's missing the atomic guard ensuring we only run the set up one. Is this just an oversight?\n* The new exception looks good, but the condition is too restrictive. We'll only hit an error when there are no tokens in the ring; the attached patch fails whenever the local node has no tokens, even if there are other tokens in the ring. Changing the condition to something like {{StorageService.instance.getTokenMetadata().sortedTokens().isEmpty()}} should suffice.\n\nIt would be great it we could cut down the time the new tests take to run. I have a few suggestions that I'll post on the dtest PR once this is ready to go.","created":"2017-05-26T20:57:27.446+0000"},{"body":"{quote}The 3.0 branch looks like it is an older version of the patch than the 2.2, 3.11, and trunk patches - it's missing the atomic guard ensuring we only run the set up one. Is this just an oversight?{quote}\n\nYes, thanks for catching that. Has been corrected.\n\n\n{quote}The new exception looks good, but the condition is too restrictive. {quote}\n\nThe condition has been changed to use {{StorageService.instance.getTokenMetadata().sortedTokens().isEmpty()}}.\n\n--\nAll four patches updated (and rebased):\n\n|| branch || testall || dtest ||\n| [cassandra-2.2_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-2.2_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-2.2_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/85] |\n| [cassandra-3.0_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-3.0_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-3.0_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/84] |\n| [cassandra-3.11_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-3.11_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-3.11_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/83] |\n| [trunk_11381|https://github.com/michaelsembwever/cassandra/tree/mck/trunk_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Ftrunk_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/82] |\n\nAll dtests are waiting on -[INFRA-14153|https://issues.apache.org/jira/browse/INFRA-14153]-.\nEDIT: dtests are running now.","created":"2017-06-05T03:48:39.378+0000"},{"body":"Great! I also ran dtests on some other machines this weekend and they looked good on all branches (relative to the current state of the branches).\n\n+1 - one minor nit that can be fixed on commit: new CHANGES.txt entries should go at the top of the list.","created":"2017-06-12T14:12:27.804+0000"},{"body":"{quote}one minor nit that can be fixed on commit: new CHANGES.txt entries should go at the top of the list.{quote}\n\nI had no idea! :-) One of those tribal habits that one either notices or doesn't… thanks for pointing it out.\n\nWill correct CHANGES.txt, and commit…","created":"2017-06-13T03:02:04.120+0000"},{"body":"[~jkni], work on the dtest in https://github.com/riptano/cassandra-dtest/pull/1479\nYou said you had some ideas about making it faster/more-efficient.","created":"2017-06-13T05:15:32.649+0000"}],"conversations":[{"body":"Starting up a node with {{-Dcassandra.join_ring=false}} in a cluster that has authentication configured, eg PasswordAuthenticator, won't be able to serve requests. This is because {{Auth.setup()}} never gets called during the startup.\n\n\nWithout {{Auth.setup()}} having been called in {{StorageService}} clients connecting to the node fail with the node throwing\n\n{noformat}\njava.lang.NullPointerException\n at org.apache.cassandra.auth.PasswordAuthenticator.authenticate(PasswordAuthenticator.java:119)\n at org.apache.cassandra.thrift.CassandraServer.login(CassandraServer.java:1471)\n at org.apache.cassandra.thrift.Cassandra$Processor$login.getResult(Cassandra.java:3505)\n at org.apache.cassandra.thrift.Cassandra$Processor$login.getResult(Cassandra.java:3489)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at com.thinkaurelius.thrift.Message.invoke(Message.java:314)\n at com.thinkaurelius.thrift.Message$Invocation.execute(Message.java:90)\n at com.thinkaurelius.thrift.TDisruptorServer$InvocationHandler.onEvent(TDisruptorServer.java:695)\n at com.thinkaurelius.thrift.TDisruptorServer$InvocationHandler.onEvent(TDisruptorServer.java:689)\n at com.lmax.disruptor.WorkProcessor.run(WorkProcessor.java:112)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{noformat}\n\nThe exception thrown from the [code|https://github.com/apache/cassandra/blob/cassandra-2.0.16/src/java/org/apache/cassandra/auth/PasswordAuthenticator.java#L119]\n{code}\nResultMessage.Rows rows = authenticateStatement.execute(QueryState.forInternalCalls(), new QueryOptions(consistencyForUser(username),\n Lists.newArrayList(ByteBufferUtil.bytes(username))));\n{code}","from":"reporter","subject":"Node running with join_ring=false and authentication can not serve requests"},{"body":"Is this ticket Patch Available?","from":"developer"},{"body":"[~JoshuaMcKenzie], I've only run unit tests against trunk. And I'm trying to figure out how to write a dtest for this (since unit tests and system property don't mix (afaik)).\n\nIf you're happy to proceed without the dtest, say and I'll 'submit patch'. And I do want to double-check that it's not possible to create auth tables without even the system tables having been created (eg bootstrap & join_ring=false).","from":"developer"},{"body":"bq. trying to figure out how to write a dtest for this\nI say we hold off until you've worked out a dtest for it. Just wanted to make sure there wasn't a patch waiting for it and it slip through the cracks.","from":"developer"},{"body":"[~JoshuaMcKenzie], the dtest is attached.","from":"developer"},{"body":"Small update to the dtest. I had been writing dtests in the trunk branch instead of master. More info [here|https://github.com/pcmanus/ccm/pull/479] and [here|https://github.com/riptano/cassandra-dtest/pull/892].","from":"developer"},{"body":"[~JoshuaMcKenzie], [~jkni], what's up?","from":"developer"},{"body":"No fault of Josh here, just a bit of a review backlog on my end. Promise I'll get to this today, sorry for the delay [~mck].","from":"developer"},{"body":"Thanks for the patches [~mck] and apologies for the delay again. This is at the top of my queue now, so I can work on getting this in quickly.\n\nI can confirm that this is a genuine issue - a few thoughts on the patches:\n* Is there a reason that we only setup auth in {{StorageService.initServer}} if the saved tokens aren't empty? This doesn't handle the case when a node starts with join_ring false on its first boot (a coordinator-only node). This can be reproduced with a small variant of your dtest in which we do not start node3 in the initial preparation. Some experimentation on my end suggests this would be fine.\n* On a similar note, is there a reason we can't just move auth setup slightly earlier in {{StorageService.initServer()}}? If we moved it before the check of the join_ring property, we could do without the idempotency parts of the patches. Again, this looks workable in my experimentation.\n* Was there any method by which you arrived at the 300 second timeouts for the cql connection on the dtests? This really drags the test on in failing cases, and I was able to reduce it to 30 seconds without any false negatives.\n* Tests handling the coordinator-only case and using JMX to join the node to the ring would be great. I think a minimal parameterized test could handle these all these permutations.","from":"developer"},{"body":"{quote} - Is there a reason that we only setup auth in StorageService.initServer if the saved tokens aren't empty? This doesn't handle the case when a node starts with join_ring false on its first boot (a coordinator-only node). This can be reproduced with a small variant of your dtest in which we do not start node3 in the initial preparation. Some experimentation on my end suggests this would be fine.{quote}\n\nGood point. Will fix. There's a similar fault in the trunk patch in regards to {{isSurveyMode}}.\n\n{quote} - On a similar note, is there a reason we can't just move auth setup slightly earlier in StorageService.initServer()? If we moved it before the check of the join_ring property, we could do without the idempotency parts of the patches. Again, this looks workable in my experimentation.{quote}\n\nThat did seem like the cleaner solution. But my understanding on the auth model and setup isn't solid, so my intention here was to offer the less-risk patch. I still suspect this (less-risk approach) is the smart approach (for a patch from me) against the non-trunk patches.\n\nI'll take a look at moving the auth setup to an earlier place in just the trunk patch, ok? It seems to make sense to move the call in {{prepareToJoin()}}.\n\n{quote} - Was there any method by which you arrived at the 300 second timeouts for the cql connection on the dtests? This really drags the test on in failing cases, and I was able to reduce it to 30 seconds without any false negatives.{quote}\n\nNo. Just playing it safe for my dual-core thinkpad. (I did have to bump it up once).\n\n{quote} - Tests handling the coordinator-only case and using JMX to join the node to the ring would be great. I think a minimal parameterized test could handle these all these permutations.{quote}\n\nI'll take a look.","from":"developer"},{"body":"[~mck], I definitely understand the inclination toward a lower-risk approach. [~beobal], as resident person-who-I-know-knows-how-auth-works, do you have any time to weigh in on whether moving auth setup earlier {{StorageService.initServer()}} is a risky proposition for any reason? I've done some research and couldn't find any cause for concern. In particular, this ticket would be simpler to fix if auth was moved before join_ring on 2.1/2.2/3.0.x and before start_gossip on trunk.","from":"developer"},{"body":"Auth setup has to wait until the node has joined the ring in order to be able to use the auth tables, as {{system_auth}} is distributed. The same applies to the other pseudo-system keyspaces {{system_traces}} & {{system_distributed}} which is why those are also created in {{joinTokenRing}}. ","from":"developer"},{"body":"On top of that, custom authenticator/authorizer/rolemanager impls may well use their own tables (& possibly other resources), which would obviously be user-defined from a system perspective and so also not accessible before joining.","from":"developer"},{"body":"Giving it some further thought, I think that the moving {{doAuthSetup}} to before joining is only going to be problematic for the first node in a fresh cluster, and then only if it doesn't join (at least in 3.0+). Starting a new coordinator-only node in an existing cluster should be fine, and re-starting an existing node without joining should also be ok, as long as the replication of {{system_auth}} can cope with it. \n","from":"developer"},{"body":"I've tested some scenarios and can confirm that the scenario Sam described above is problematic.\n\nAnother concern I have is potentially performing auth setup before a node is added to its own TokenMetadata in the join_ring false case. It seems to me that this might cause problems with our own implementations and could feasibly cause problems for 3rd party implementations, so I think the original approach of handling auth setup down both branches may be worth pursuing here in the interest of safety.\n\nOn trunk, we have another property \"start_gossip\" that stops {{StorageService.initServer}} even earlier - I don't think we can support auth at this point, so I don't care about supporting it in this ticket (my earlier comment hinted at the fact that I was considering whether it was in scope).","from":"developer"},{"body":"Coming back to this issue now.\n\n{quote}I've tested some scenarios and can confirm that the scenario Sam described above is problematic.{quote}\n[~jkni], which scenario Sam describes?\n\n{quote}...so I think the original approach of handling auth setup down both branches may be worth pursuing here in the interest of safety.{quote}\n\nSo I've uploaded new patches for 2.1 and dtest that \n - set up auth also when tokens are empty\n - add a dtest to test ^^\n\nWorking on the other patches.","from":"developer"},{"body":"…and reduced the timeouts in the dtest to 30 seconds.\n\nAll patches are updated.","from":"developer"},{"body":"[~mck] - The scenario I was referring to was moving {{doAuthSetup}} to before joining for the first node in a fresh cluster.\n\nThanks for keeping this moving. This is near the top of my review queue and I should be able to get to it soon.","from":"developer"},{"body":"ping. is there anything outstanding here?","from":"developer"},{"body":"[~jkni] ping…? ","from":"developer"},{"body":"Thanks for pinging me on this - as you suspected, it slipped through the cracks. \n\nOn reviewing the final version of this patch, I found one problem with the 2.2+ patches. The proposed patch technically breaks the documented {{IRoleManager}}, {{IAuthenticator}}, and {{IAuthorizer}} interfaces. With the implementation given, {{doAuthSetup}} will be called twice for a node started with {{join_ring=False}}, so {{setup()}} will be called twice for the role manager, authenticator, and authorizer. In the documentation for these three public interfaces, we state that {{setup()}} will only be called once after starting a node. I think we should preserve this documented behavior. While slightly less elegant, I think we should instead track whether we've run {{doAuthSetup}} and not repeat this call for a node started with {{join_ring=False}} that is asked to join. This means the parts of the patch implementing idempotency for the MigrationManager listener registration become unnecessary.","from":"developer"},{"body":"Makes sense, will update patches…","from":"developer"},{"body":"[~jkni], all patches (but 2.1) have been updated.","from":"developer"},{"body":"Thanks - the patches look good and I put them through CI.\n\n||branch||testall||dtest||\n|[CASSANDRA-11381-2.2|https://github.com/jkni/cassandra/tree/CASSANDRA-11381-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-2.2-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-2.2-dtest]|\n|[CASSANDRA-11381-3.0|https://github.com/jkni/cassandra/tree/CASSANDRA-11381-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-3.0-dtest]|\n|[CASSANDRA-11381-3.X|https://github.com/jkni/cassandra/tree/CASSANDRA-11381-3.X]|[testall|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-3.X-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-3.X-dtest]|\n|[CASSANDRA-11381-trunk|https://github.com/jkni/cassandra/tree/CASSANDRA-11381-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-trunk-dtest]|\n\nCI looks good for the most part, and I checked that your added dtest passes on all branches. CI revealed one small problem - when a fresh node is started with join_ring=False that has no tokens for other nodes discovered through gossip and no saved tokens, it hits an AssertionError in {{CassandraRoleManager}} setup that is not handled and gets logged as an error by a top level error handler, as seen [here|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-CASSANDRA-11381-2.2-dtest/1/testReport/junit/topology_test/TestTopology/do_not_join_ring_test/]. In this specific test case, this behavior is hit because a single node cluster is started with join_ring=False. Since this prevents setup from being retried within the {{CassandraRoleManager}}, it seems to me that it is probably worth checking for an absence of tokens in {{CassandraRoleManager.setupDefaultRole}} and throwing a catchable exception/printing a warning so that setup can be retried. What do you think? There may be another alternative I haven't considered.","from":"developer"},{"body":"[~jkni], i have finally come back to this…\n\n{quote}CI revealed one small problem - when a fresh node is started with join_ring=False that has no tokens for other nodes discovered through gossip and no saved tokens, it hits an AssertionError in CassandraRoleManager setup that is not handled{quote}\n\n{{CassandraRoleManager.setupDefaultRole()}} now throws an {{IllegalStateException}} if there are no known tokens in the ring. This permits {{CassandraRoleManager.scheduleSetupTask(..)}} to retry.\n\nThe {{topology_test.py:TestTopology.do_not_join_ring_test}} dtest looks ok now, locally.\n\n\\\\\n\n|| branch || testall || dtest ||\n| [cassandra-2.2_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-2.2_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-2.2_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/40/] |\n| [cassandra-3.0_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-3.0_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-3.0_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/41/] |\n| [cassandra-3.11_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-3.11_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-3.11_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/42/] |\n| [trunk_11381|https://github.com/michaelsembwever/cassandra/tree/mck/trunk_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Ftrunk_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/43/] |\n\n(the asf jenkins dtests take a while, and won't include the new [dtest|https://github.com/michaelsembwever/cassandra-dtest/tree/mck/master_11381])","from":"developer"},{"body":"[~jkni], dtests look good. are we ready to commit this?","from":"developer"},{"body":"Thanks - this looks just about good. A few comments:\n* The 3.0 branch looks like it is an older version of the patch than the 2.2, 3.11, and trunk patches - it's missing the atomic guard ensuring we only run the set up one. Is this just an oversight?\n* The new exception looks good, but the condition is too restrictive. We'll only hit an error when there are no tokens in the ring; the attached patch fails whenever the local node has no tokens, even if there are other tokens in the ring. Changing the condition to something like {{StorageService.instance.getTokenMetadata().sortedTokens().isEmpty()}} should suffice.\n\nIt would be great it we could cut down the time the new tests take to run. I have a few suggestions that I'll post on the dtest PR once this is ready to go.","from":"developer"},{"body":"{quote}The 3.0 branch looks like it is an older version of the patch than the 2.2, 3.11, and trunk patches - it's missing the atomic guard ensuring we only run the set up one. Is this just an oversight?{quote}\n\nYes, thanks for catching that. Has been corrected.\n\n\n{quote}The new exception looks good, but the condition is too restrictive. {quote}\n\nThe condition has been changed to use {{StorageService.instance.getTokenMetadata().sortedTokens().isEmpty()}}.\n\n--\nAll four patches updated (and rebased):\n\n|| branch || testall || dtest ||\n| [cassandra-2.2_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-2.2_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-2.2_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/85] |\n| [cassandra-3.0_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-3.0_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-3.0_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/84] |\n| [cassandra-3.11_11381|https://github.com/michaelsembwever/cassandra/tree/mck/cassandra-3.11_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Fcassandra-3.11_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/83] |\n| [trunk_11381|https://github.com/michaelsembwever/cassandra/tree/mck/trunk_11381]\t| [testall|https://circleci.com/gh/michaelsembwever/cassandra/tree/mck%2Ftrunk_11381]\t| [dtest|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/82] |\n\nAll dtests are waiting on -[INFRA-14153|https://issues.apache.org/jira/browse/INFRA-14153]-.\nEDIT: dtests are running now.","from":"developer"},{"body":"Great! I also ran dtests on some other machines this weekend and they looked good on all branches (relative to the current state of the branches).\n\n+1 - one minor nit that can be fixed on commit: new CHANGES.txt entries should go at the top of the list.","from":"developer"},{"body":"{quote}one minor nit that can be fixed on commit: new CHANGES.txt entries should go at the top of the list.{quote}\n\nI had no idea! :-) One of those tribal habits that one either notices or doesn't… thanks for pointing it out.\n\nWill correct CHANGES.txt, and commit…","from":"developer"},{"body":"[~jkni], work on the dtest in https://github.com/riptano/cassandra-dtest/pull/1479\nYou said you had some ideas about making it faster/more-efficient.","from":"developer"}],"created":"2016-03-18T03:19:59.000+0000","description":"Starting up a node with {{-Dcassandra.join_ring=false}} in a cluster that has authentication configured, eg PasswordAuthenticator, won't be able to serve requests. This is because {{Auth.setup()}} never gets called during the startup.\n\n\nWithout {{Auth.setup()}} having been called in {{StorageService}} clients connecting to the node fail with the node throwing\n\n{noformat}\njava.lang.NullPointerException\n at org.apache.cassandra.auth.PasswordAuthenticator.authenticate(PasswordAuthenticator.java:119)\n at org.apache.cassandra.thrift.CassandraServer.login(CassandraServer.java:1471)\n at org.apache.cassandra.thrift.Cassandra$Processor$login.getResult(Cassandra.java:3505)\n at org.apache.cassandra.thrift.Cassandra$Processor$login.getResult(Cassandra.java:3489)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at com.thinkaurelius.thrift.Message.invoke(Message.java:314)\n at com.thinkaurelius.thrift.Message$Invocation.execute(Message.java:90)\n at com.thinkaurelius.thrift.TDisruptorServer$InvocationHandler.onEvent(TDisruptorServer.java:695)\n at com.thinkaurelius.thrift.TDisruptorServer$InvocationHandler.onEvent(TDisruptorServer.java:689)\n at com.lmax.disruptor.WorkProcessor.run(WorkProcessor.java:112)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{noformat}\n\nThe exception thrown from the [code|https://github.com/apache/cassandra/blob/cassandra-2.0.16/src/java/org/apache/cassandra/auth/PasswordAuthenticator.java#L119]\n{code}\nResultMessage.Rows rows = authenticateStatement.execute(QueryState.forInternalCalls(), new QueryOptions(consistencyForUser(username),\n Lists.newArrayList(ByteBufferUtil.bytes(username))));\n{code}","issue_id":"12951429","key":"CASSANDRA-11381","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-06-13T04:42:31.000+0000","role":"fixed_distractor","summary":"Node running with join_ring=false and authentication can not serve requests"} {"case_id":"12960014","cluster":"DISTRACTOR-CASSANDRA-11604","comments":[{"body":"as a workaround:\nnodetool scrub followed by a cassandra restart on every node makes the table accessible again","created":"2016-04-20T14:27:11.384+0000"},{"body":"Whoever is gonna deal with the ticket, please have a look into why scrub is able to fix the issue (it shouldn't be - likely there is a bug in scrub).","created":"2016-04-29T15:26:35.050+0000"},{"body":"This assertion was removed in [trunk|https://github.com/apache/cassandra/commit/677230df694752c7ecf6d5459eee60ad7cf45ecf#diff-bc19f192ef82fbca9abd27526054bb0fL254] (appl in 3.5 has similar effect).\n\nFrom what I can say, scrub doesn't fix that issue. Node restart alone has the same effect, or the flush:\nDuring the node restart, commit log will replay mutation with the same schema as the table itself. During the flush and consequent reads, all Cells will get the correct Column Definition. Although this assert doesn't change the behaviour, since ALTER statements only allow \"backward-compatible\" changes (after the schema change, it'll be possible to work with the old version, too).\n\nI've added the test for this particular edge case (updating UDT within inserted non-frozen map) for {{trunk}} and removed assert in {{3.0.x}} (along with adding the test):\n\n||[3.0|https://github.com/ifesdjeen/cassandra/tree/11604-3.0]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-11604-3.0-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-11604-3.0-dtest/]|\n||[trunk|https://github.com/ifesdjeen/cassandra/tree/11604-trunk]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-11604-trunk-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-11604-trunk-dtest/]|\n\n{{LegacySSTableTest}} is failing locally on trunk, too, although it happened before this commit (also, there were no code changes for trunk). The rest of tests are passing locally, too.","created":"2016-05-09T11:51:11.384+0000"},{"body":"+1 - the assert doesn't seem valid to me, and looking over Tyler's changes in trunk, I can't see a reason the assertion would be valid here but not on trunk.\n\nTests look good.","created":"2016-06-01T17:13:19.767+0000"},{"body":"[Committed test|https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=0d5984b9dbd54a42fbbe6a71a045b13a612208b6] and merged up.","created":"2016-06-14T14:21:44.404+0000"}],"conversations":[{"body":"in cassandra 3.5 i get the following exception when i run this cqls:\n{code}\n--DROP KEYSPACE bugtest ;\nCREATE KEYSPACE bugtest\n WITH REPLICATION = { 'class' : 'SimpleStrategy', 'replication_factor' : 1 };\nuse bugtest;\nCREATE TYPE tt (\n\ta boolean\n);\ncreate table t1 (\n\tk text,\n\tv map>,\n\tPRIMARY KEY(k)\n);\ninsert into t1 (k,v) values ('k2',{'mk':{a:false}});\nALTER TYPE tt ADD b boolean;\nUPDATE t1 SET v['mk'] = { b:true } WHERE k = 'k2';\nselect * from t1; \n{code}\nthe last select fails.\n{code}\nWARN [SharedPool-Worker-5] 2016-04-19 14:18:49,885 AbstractLocalAwareExecutorService.java:169 - Uncaught exception on thread Thread[SharedPool-Worker-5,5,main]: {}\njava.lang.AssertionError: null\n at org.apache.cassandra.db.rows.ComplexColumnData$Builder.addCell(ComplexColumnData.java:254) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.Row$Merger$ColumnDataReducer.getReduced(Row.java:623) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.Row$Merger$ColumnDataReducer.getReduced(Row.java:549) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:217) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:156) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.Row$Merger.merge(Row.java:526) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator$MergeReducer.getReduced(UnfilteredRowIterators.java:473) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator$MergeReducer.getReduced(UnfilteredRowIterators.java:437) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:217) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:156) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator.computeNext(UnfilteredRowIterators.java:419) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator.computeNext(UnfilteredRowIterators.java:279) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:100) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:32) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.transform.BaseRows.hasNext(BaseRows.java:112) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.transform.UnfilteredRows.isEmpty(UnfilteredRows.java:38) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:64) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:24) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.transform.BasePartitions.hasNext(BasePartitions.java:76) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:289) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:134) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:127) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:123) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:65) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:292) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1799) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2466) ~[apache-cassandra-3.5.jar:3.5]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_72-internal]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [apache-cassandra-3.5.jar:3.5]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_72-internal]\n{code}","from":"reporter","subject":"select on table fails after changing user defined type in map"},{"body":"as a workaround:\nnodetool scrub followed by a cassandra restart on every node makes the table accessible again","from":"developer"},{"body":"Whoever is gonna deal with the ticket, please have a look into why scrub is able to fix the issue (it shouldn't be - likely there is a bug in scrub).","from":"developer"},{"body":"This assertion was removed in [trunk|https://github.com/apache/cassandra/commit/677230df694752c7ecf6d5459eee60ad7cf45ecf#diff-bc19f192ef82fbca9abd27526054bb0fL254] (appl in 3.5 has similar effect).\n\nFrom what I can say, scrub doesn't fix that issue. Node restart alone has the same effect, or the flush:\nDuring the node restart, commit log will replay mutation with the same schema as the table itself. During the flush and consequent reads, all Cells will get the correct Column Definition. Although this assert doesn't change the behaviour, since ALTER statements only allow \"backward-compatible\" changes (after the schema change, it'll be possible to work with the old version, too).\n\nI've added the test for this particular edge case (updating UDT within inserted non-frozen map) for {{trunk}} and removed assert in {{3.0.x}} (along with adding the test):\n\n||[3.0|https://github.com/ifesdjeen/cassandra/tree/11604-3.0]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-11604-3.0-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-11604-3.0-dtest/]|\n||[trunk|https://github.com/ifesdjeen/cassandra/tree/11604-trunk]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-11604-trunk-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-11604-trunk-dtest/]|\n\n{{LegacySSTableTest}} is failing locally on trunk, too, although it happened before this commit (also, there were no code changes for trunk). The rest of tests are passing locally, too.","from":"developer"},{"body":"+1 - the assert doesn't seem valid to me, and looking over Tyler's changes in trunk, I can't see a reason the assertion would be valid here but not on trunk.\n\nTests look good.","from":"developer"},{"body":"[Committed test|https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=0d5984b9dbd54a42fbbe6a71a045b13a612208b6] and merged up.","from":"developer"}],"created":"2016-04-19T14:31:35.000+0000","description":"in cassandra 3.5 i get the following exception when i run this cqls:\n{code}\n--DROP KEYSPACE bugtest ;\nCREATE KEYSPACE bugtest\n WITH REPLICATION = { 'class' : 'SimpleStrategy', 'replication_factor' : 1 };\nuse bugtest;\nCREATE TYPE tt (\n\ta boolean\n);\ncreate table t1 (\n\tk text,\n\tv map>,\n\tPRIMARY KEY(k)\n);\ninsert into t1 (k,v) values ('k2',{'mk':{a:false}});\nALTER TYPE tt ADD b boolean;\nUPDATE t1 SET v['mk'] = { b:true } WHERE k = 'k2';\nselect * from t1; \n{code}\nthe last select fails.\n{code}\nWARN [SharedPool-Worker-5] 2016-04-19 14:18:49,885 AbstractLocalAwareExecutorService.java:169 - Uncaught exception on thread Thread[SharedPool-Worker-5,5,main]: {}\njava.lang.AssertionError: null\n at org.apache.cassandra.db.rows.ComplexColumnData$Builder.addCell(ComplexColumnData.java:254) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.Row$Merger$ColumnDataReducer.getReduced(Row.java:623) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.Row$Merger$ColumnDataReducer.getReduced(Row.java:549) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:217) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:156) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.Row$Merger.merge(Row.java:526) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator$MergeReducer.getReduced(UnfilteredRowIterators.java:473) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator$MergeReducer.getReduced(UnfilteredRowIterators.java:437) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:217) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:156) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator.computeNext(UnfilteredRowIterators.java:419) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator.computeNext(UnfilteredRowIterators.java:279) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:100) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:32) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.transform.BaseRows.hasNext(BaseRows.java:112) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.transform.UnfilteredRows.isEmpty(UnfilteredRows.java:38) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:64) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:24) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.transform.BasePartitions.hasNext(BasePartitions.java:76) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:289) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:134) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:127) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:123) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:65) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:292) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1799) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2466) ~[apache-cassandra-3.5.jar:3.5]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_72-internal]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [apache-cassandra-3.5.jar:3.5]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [apache-cassandra-3.5.jar:3.5]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_72-internal]\n{code}","issue_id":"12960014","key":"CASSANDRA-11604","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-06-14T14:21:44.000+0000","role":"fixed_distractor","summary":"select on table fails after changing user defined type in map"} {"case_id":"12967361","cluster":"DISTRACTOR-CASSANDRA-11750","comments":[{"body":"I think on trunk this is not the case anymore after CASSANDRA-11578 which decoupled disk failure policy handling and only applied it when online.\nBut still there can be the case where CorruptSSTableException is handled in offline scrub that prevent it to continue.\nLet me check.","created":"2016-05-11T13:10:36.496+0000"},{"body":"So, to be clear, the issue happens when one of {{system}} tables is corrupted.\nIn the description above, OP tried to scrub {{system.compactions_in_progress}} table, but the actual exception happened during loading schema (this opens all system tables) not during scrubbing SSTables.\n\nIf {{system}} tables are fine, then scrubbing continues to work in 2.1/2.2.\nIn 3.0 and above, schema moved to its own keyspace, so in those version if schema SSTables are ok then you can scrub system keyspace.\n\nProbably backporting CASSANDRA-11578 to 2.1 and 2.2 (and even 3.0) should do the job.\n","created":"2016-05-18T20:26:51.262+0000"},{"body":"||branch||testall||dtest||\n|[11750-2.1|https://github.com/yukim/cassandra/tree/11750-2.1]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-2.1-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-2.1-dtest/lastCompletedBuild/testReport/]|\n|[11750-2.2|https://github.com/yukim/cassandra/tree/11750-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-2.2-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-2.2-dtest/lastCompletedBuild/testReport/]|\n\nBackported CASSANDRA-11578 to 2.1/2.2. This ables tools to skip disk failure policy check.","created":"2016-05-19T13:23:43.579+0000"},{"body":"Awesome - thanks Yuki!","created":"2016-05-19T13:36:02.447+0000"},{"body":"[~yukim] is the a reason for not putting this in 3.0 as well? Seems strange to not merge the change all the way forward and only have it in 2.1/2.2/3.8?","created":"2016-05-19T21:52:58.742+0000"},{"body":"you are right. Here is 3.0 version.\n\n||branch||testall||dtest||\n|[11750-3.0|https://github.com/yukim/cassandra/tree/11750-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-3.0-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-3.0-dtest/lastCompletedBuild/testReport/]|\n","created":"2016-05-20T03:54:51.200+0000"},{"body":"+1, reproduced locally with ccm and is fixed by patch. Backport looks good.","created":"2016-05-24T23:36:21.758+0000"},{"body":"Thanks, committed as {{b851792c4e3ae32b8d863d9079cca6d135f1cf23}}.","created":"2016-05-26T15:02:42.067+0000"}],"conversations":[{"body":"Hit a failure on startup due to corruption of some sstables in system keyspace. Deleted the listed file and restarted - came down again with another file.\n\nFigured that I may as well run scrub to clean up all the files. Got following error:\n\n{noformat}\nsstablescrub system compaction_history \nERROR 17:21:34 Exiting forcefully due to file system exception on startup, disk failure policy \"stop\" \norg.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /cassandra/data/system/compaction_history-b4dbb7b4dc493fb5b3bfce6e434832ca/system-compaction_history-ka-1936-CompressionInfo.db \nat org.apache.cassandra.io.compress.CompressionMetadata.(CompressionMetadata.java:131) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.compress.CompressionMetadata.create(CompressionMetadata.java:85) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.util.CompressedSegmentedFile$Builder.metadata(CompressedSegmentedFile.java:79) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.util.CompressedPoolingSegmentedFile$Builder.complete(CompressedPoolingSegmentedFile.java:72) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.util.SegmentedFile$Builder.complete(SegmentedFile.java:169) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.sstable.SSTableReader.load(SSTableReader.java:741) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.sstable.SSTableReader.load(SSTableReader.java:692) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.sstable.SSTableReader.open(SSTableReader.java:480) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046]\nat org.apache.cassandra.io.sstable.SSTableReader.open(SSTableReader.java:376) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046]\nat org.apache.cassandra.io.sstable.SSTableReader$4.run(SSTableReader.java:523) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_79] \nat java.util.concurrent.FutureTask.run(FutureTask.java:262) [na:1.7.0_79] \nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) [na:1.7.0_79] \nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_79] \nat java.lang.Thread.run(Thread.java:745) [na:1.7.0_79] \nCaused by: java.io.EOFException: null \nat java.io.DataInputStream.readUnsignedShort(DataInputStream.java:340) ~[na:1.7.0_79] \nat java.io.DataInputStream.readUTF(DataInputStream.java:589) ~[na:1.7.0_79] \nat java.io.DataInputStream.readUTF(DataInputStream.java:564) ~[na:1.7.0_79] \nat org.apache.cassandra.io.compress.CompressionMetadata.(CompressionMetadata.java:106) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \n... 14 common frames omitted \n{noformat}\n\nI guess it might be by design - but I'd argue that I should at least have the option to continue and let it do it's thing. I'd prefer that sstablescrub ignored the disk failure policy. ","from":"reporter","subject":"Offline scrub should not abort when it hits corruption"},{"body":"I think on trunk this is not the case anymore after CASSANDRA-11578 which decoupled disk failure policy handling and only applied it when online.\nBut still there can be the case where CorruptSSTableException is handled in offline scrub that prevent it to continue.\nLet me check.","from":"developer"},{"body":"So, to be clear, the issue happens when one of {{system}} tables is corrupted.\nIn the description above, OP tried to scrub {{system.compactions_in_progress}} table, but the actual exception happened during loading schema (this opens all system tables) not during scrubbing SSTables.\n\nIf {{system}} tables are fine, then scrubbing continues to work in 2.1/2.2.\nIn 3.0 and above, schema moved to its own keyspace, so in those version if schema SSTables are ok then you can scrub system keyspace.\n\nProbably backporting CASSANDRA-11578 to 2.1 and 2.2 (and even 3.0) should do the job.\n","from":"developer"},{"body":"||branch||testall||dtest||\n|[11750-2.1|https://github.com/yukim/cassandra/tree/11750-2.1]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-2.1-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-2.1-dtest/lastCompletedBuild/testReport/]|\n|[11750-2.2|https://github.com/yukim/cassandra/tree/11750-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-2.2-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-2.2-dtest/lastCompletedBuild/testReport/]|\n\nBackported CASSANDRA-11578 to 2.1/2.2. This ables tools to skip disk failure policy check.","from":"developer"},{"body":"Awesome - thanks Yuki!","from":"developer"},{"body":"[~yukim] is the a reason for not putting this in 3.0 as well? Seems strange to not merge the change all the way forward and only have it in 2.1/2.2/3.8?","from":"developer"},{"body":"you are right. Here is 3.0 version.\n\n||branch||testall||dtest||\n|[11750-3.0|https://github.com/yukim/cassandra/tree/11750-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-3.0-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-11750-3.0-dtest/lastCompletedBuild/testReport/]|\n","from":"developer"},{"body":"+1, reproduced locally with ccm and is fixed by patch. Backport looks good.","from":"developer"},{"body":"Thanks, committed as {{b851792c4e3ae32b8d863d9079cca6d135f1cf23}}.","from":"developer"}],"created":"2016-05-11T10:25:07.000+0000","description":"Hit a failure on startup due to corruption of some sstables in system keyspace. Deleted the listed file and restarted - came down again with another file.\n\nFigured that I may as well run scrub to clean up all the files. Got following error:\n\n{noformat}\nsstablescrub system compaction_history \nERROR 17:21:34 Exiting forcefully due to file system exception on startup, disk failure policy \"stop\" \norg.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /cassandra/data/system/compaction_history-b4dbb7b4dc493fb5b3bfce6e434832ca/system-compaction_history-ka-1936-CompressionInfo.db \nat org.apache.cassandra.io.compress.CompressionMetadata.(CompressionMetadata.java:131) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.compress.CompressionMetadata.create(CompressionMetadata.java:85) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.util.CompressedSegmentedFile$Builder.metadata(CompressedSegmentedFile.java:79) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.util.CompressedPoolingSegmentedFile$Builder.complete(CompressedPoolingSegmentedFile.java:72) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.util.SegmentedFile$Builder.complete(SegmentedFile.java:169) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.sstable.SSTableReader.load(SSTableReader.java:741) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.sstable.SSTableReader.load(SSTableReader.java:692) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat org.apache.cassandra.io.sstable.SSTableReader.open(SSTableReader.java:480) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046]\nat org.apache.cassandra.io.sstable.SSTableReader.open(SSTableReader.java:376) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046]\nat org.apache.cassandra.io.sstable.SSTableReader$4.run(SSTableReader.java:523) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \nat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_79] \nat java.util.concurrent.FutureTask.run(FutureTask.java:262) [na:1.7.0_79] \nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) [na:1.7.0_79] \nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_79] \nat java.lang.Thread.run(Thread.java:745) [na:1.7.0_79] \nCaused by: java.io.EOFException: null \nat java.io.DataInputStream.readUnsignedShort(DataInputStream.java:340) ~[na:1.7.0_79] \nat java.io.DataInputStream.readUTF(DataInputStream.java:589) ~[na:1.7.0_79] \nat java.io.DataInputStream.readUTF(DataInputStream.java:564) ~[na:1.7.0_79] \nat org.apache.cassandra.io.compress.CompressionMetadata.(CompressionMetadata.java:106) ~[cassandra-all-2.1.12.1046.jar:2.1.12.1046] \n... 14 common frames omitted \n{noformat}\n\nI guess it might be by design - but I'd argue that I should at least have the option to continue and let it do it's thing. I'd prefer that sstablescrub ignored the disk failure policy. ","issue_id":"12967361","key":"CASSANDRA-11750","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-05-26T15:02:42.000+0000","role":"fixed_distractor","summary":"Offline scrub should not abort when it hits corruption"} {"case_id":"12977666","cluster":"DISTRACTOR-CASSANDRA-11991","comments":[{"body":"I can review the patch. We should try to get it in for 2.1.15 if possible ","created":"2016-06-15T16:57:55.019+0000"},{"body":"For context, the problem is basically the one I described in [my comment|https://issues.apache.org/jira/browse/CASSANDRA-9649?focusedCommentId=14601016&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-14601016] on CASSANDRA-9649 and for which I suggested reverting CASSANDRA-7801.\n\nNow, I was kind of wrong about reverting CASSANDRA-7801 since since CASSANDRA-9649 we were relying on {{ClientState.getTimestamp()}} to give use timestamp that were unique for the running VM, which meant we can't blindly revert CASSANDRA-7801.\n\nWhat I think is the simplest solution however is to stop relying on that property (of {{ClientState.getTimestamp()}}) for the uniqueness of our ballots, but instead randomize the non-timestamp parts of the ballot for every new ballot. With that, we don't have to revert CASSANDRA-7801, we just have to ensure that if we use the last known proposal timestamp (i.e. if whomever clock generated that timestamp is \"in the future\"), we don't persist it in the local clock (this in turn means the timestamp might not be unique in the VM for 2 concurrent paxos operation and hence the need to randomize the rest of the UUID).\n\nI've pushed a patch for this for 2.1. I'll attach branches for 2.2+ with tests tomorrow (but was waiting on the 2.1 results before doing that) but I don't think the modified code has changed since 2.1 so marking ready for review in the meantime.\n\n| [2.1|https://github.com/pcmanus/cassandra/commits/11991-2.1] | [utests|http://cassci.datastax.com/job/pcmanus-11991-2.1-testall/] | [dtests|http://cassci.datastax.com/job/pcmanus-11991-2.1-dtest/] |\n","created":"2016-06-16T16:01:54.176+0000"},{"body":"lgtm. The only minor nit I have is CASSANDRA-9649 made {{ClientState#lastTimestampMicros}} a static field. I believe this change helped accelerate the cluster get into a bad state wrt the propagation of bad timestamps. wdyt about switching it back to being an instance field (not static)?","created":"2016-06-21T22:53:54.415+0000"},{"body":"bq. The only minor nit I have is CASSANDRA-9649 made ClientState#lastTimestampMicros a static field. believe this change helped accelerate the cluster get into a bad state wrt the propagation of bad timestamps. wdyt about switching it back to being an instance field (not static)?\n\nThere is really 2 reasons I made it static:\n# CASSANDRA-7801: it's not a huge thing, but it helps user being less confused.\n# because I feel that having it not static was a mistake in the first place. That is, even if we completely forget about Paxos, I think having our CQL clock strictly monotonic per-node (rather than per-connection) is cleaner and less surprising to people. So I'm not a fan of getting back to the older behavior, and doing so could be considered a breaking change (easier to give new guarantees than give some away).\n\nI don't disagree that fact made the consequences of this bug worst, but I don't think removing the {{static}} is minor at all. And that code feel easy enough to convince oneself that we're not modifying {{lastTimestampMicros}} in bad ways anymore, so hopefully we won't make that mistake anymore.\n\nBesides, if you're smart, you should be switching to client generated timestamps anyway :)","created":"2016-06-22T07:49:37.567+0000"},{"body":"Also, I've merged the patch up (no conflict whatsoever) and started CI on all branches:\n\n|| version || utests || dtests||\n| [2.1|https://github.com/pcmanus/cassandra/commits/11991-2.1] | [utests|http://cassci.datastax.com/job/pcmanus-11991-2.1-testall/] | [dtests|http://cassci.datastax.com/job/pcmanus-11991-2.1-dtest/] |\n| [2.2|https://github.com/pcmanus/cassandra/commits/11991-2.2] | [utests|http://cassci.datastax.com/job/pcmanus-11991-2.2-testall/] | [dtests|http://cassci.datastax.com/job/pcmanus-11991-2.2-dtest/]|\n| [3.0|https://github.com/pcmanus/cassandra/commits/11991-3.0] | [utests|http://cassci.datastax.com/job/pcmanus-11991-3.0-testall/] | [dtests|http://cassci.datastax.com/job/pcmanus-11991-3.0-dtest/]|\n| [trunk|https://github.com/pcmanus/cassandra/commits/11991-trunk] | [utests|http://cassci.datastax.com/job/pcmanus-11991-trunk-testall/] | [dtests|http://cassci.datastax.com/job/pcmanus-11991-trunk-dtest/]|\n\nI also updated the comment on top of {{ClientState#lastTimestampMicros}} since it wasn't complete on the real motivation.\n","created":"2016-06-22T08:16:14.540+0000"},{"body":"bq. I think having our CQL clock strictly monotonic per-node (rather than per-connection) is cleaner and less surprising to people.\n\nok, I'll buy that.\n\nOtherwise, +1 on all branches. I checked out the test results, and the ones that failed were either 1) unrelated to this CAS change or 2) test timeouts (and still unrelated to CAS).","created":"2016-06-22T12:28:58.693+0000"},{"body":"+1 [~slebresne] Please commit it. ","created":"2016-06-23T03:46:03.445+0000"},{"body":"Committed, thanks.","created":"2016-06-23T07:59:01.111+0000"}],"conversations":[{"body":"W made a mistake in CASSANDRA-9649 so that a temporal clock skew on one node can \"corrupt\" other node clocks through Paxos. That wasn't intended and we should fix that. I'll attach a patch later.","from":"reporter","subject":"On clock skew, paxos may \"corrupt\" the node clock"},{"body":"I can review the patch. We should try to get it in for 2.1.15 if possible ","from":"developer"},{"body":"For context, the problem is basically the one I described in [my comment|https://issues.apache.org/jira/browse/CASSANDRA-9649?focusedCommentId=14601016&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-14601016] on CASSANDRA-9649 and for which I suggested reverting CASSANDRA-7801.\n\nNow, I was kind of wrong about reverting CASSANDRA-7801 since since CASSANDRA-9649 we were relying on {{ClientState.getTimestamp()}} to give use timestamp that were unique for the running VM, which meant we can't blindly revert CASSANDRA-7801.\n\nWhat I think is the simplest solution however is to stop relying on that property (of {{ClientState.getTimestamp()}}) for the uniqueness of our ballots, but instead randomize the non-timestamp parts of the ballot for every new ballot. With that, we don't have to revert CASSANDRA-7801, we just have to ensure that if we use the last known proposal timestamp (i.e. if whomever clock generated that timestamp is \"in the future\"), we don't persist it in the local clock (this in turn means the timestamp might not be unique in the VM for 2 concurrent paxos operation and hence the need to randomize the rest of the UUID).\n\nI've pushed a patch for this for 2.1. I'll attach branches for 2.2+ with tests tomorrow (but was waiting on the 2.1 results before doing that) but I don't think the modified code has changed since 2.1 so marking ready for review in the meantime.\n\n| [2.1|https://github.com/pcmanus/cassandra/commits/11991-2.1] | [utests|http://cassci.datastax.com/job/pcmanus-11991-2.1-testall/] | [dtests|http://cassci.datastax.com/job/pcmanus-11991-2.1-dtest/] |\n","from":"developer"},{"body":"lgtm. The only minor nit I have is CASSANDRA-9649 made {{ClientState#lastTimestampMicros}} a static field. I believe this change helped accelerate the cluster get into a bad state wrt the propagation of bad timestamps. wdyt about switching it back to being an instance field (not static)?","from":"developer"},{"body":"bq. The only minor nit I have is CASSANDRA-9649 made ClientState#lastTimestampMicros a static field. believe this change helped accelerate the cluster get into a bad state wrt the propagation of bad timestamps. wdyt about switching it back to being an instance field (not static)?\n\nThere is really 2 reasons I made it static:\n# CASSANDRA-7801: it's not a huge thing, but it helps user being less confused.\n# because I feel that having it not static was a mistake in the first place. That is, even if we completely forget about Paxos, I think having our CQL clock strictly monotonic per-node (rather than per-connection) is cleaner and less surprising to people. So I'm not a fan of getting back to the older behavior, and doing so could be considered a breaking change (easier to give new guarantees than give some away).\n\nI don't disagree that fact made the consequences of this bug worst, but I don't think removing the {{static}} is minor at all. And that code feel easy enough to convince oneself that we're not modifying {{lastTimestampMicros}} in bad ways anymore, so hopefully we won't make that mistake anymore.\n\nBesides, if you're smart, you should be switching to client generated timestamps anyway :)","from":"developer"},{"body":"Also, I've merged the patch up (no conflict whatsoever) and started CI on all branches:\n\n|| version || utests || dtests||\n| [2.1|https://github.com/pcmanus/cassandra/commits/11991-2.1] | [utests|http://cassci.datastax.com/job/pcmanus-11991-2.1-testall/] | [dtests|http://cassci.datastax.com/job/pcmanus-11991-2.1-dtest/] |\n| [2.2|https://github.com/pcmanus/cassandra/commits/11991-2.2] | [utests|http://cassci.datastax.com/job/pcmanus-11991-2.2-testall/] | [dtests|http://cassci.datastax.com/job/pcmanus-11991-2.2-dtest/]|\n| [3.0|https://github.com/pcmanus/cassandra/commits/11991-3.0] | [utests|http://cassci.datastax.com/job/pcmanus-11991-3.0-testall/] | [dtests|http://cassci.datastax.com/job/pcmanus-11991-3.0-dtest/]|\n| [trunk|https://github.com/pcmanus/cassandra/commits/11991-trunk] | [utests|http://cassci.datastax.com/job/pcmanus-11991-trunk-testall/] | [dtests|http://cassci.datastax.com/job/pcmanus-11991-trunk-dtest/]|\n\nI also updated the comment on top of {{ClientState#lastTimestampMicros}} since it wasn't complete on the real motivation.\n","from":"developer"},{"body":"bq. I think having our CQL clock strictly monotonic per-node (rather than per-connection) is cleaner and less surprising to people.\n\nok, I'll buy that.\n\nOtherwise, +1 on all branches. I checked out the test results, and the ones that failed were either 1) unrelated to this CAS change or 2) test timeouts (and still unrelated to CAS).","from":"developer"},{"body":"+1 [~slebresne] Please commit it. ","from":"developer"},{"body":"Committed, thanks.","from":"developer"}],"created":"2016-06-10T16:17:15.000+0000","description":"W made a mistake in CASSANDRA-9649 so that a temporal clock skew on one node can \"corrupt\" other node clocks through Paxos. That wasn't intended and we should fix that. I'll attach a patch later.","issue_id":"12977666","key":"CASSANDRA-11991","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-06-23T07:59:00.000+0000","role":"fixed_distractor","summary":"On clock skew, paxos may \"corrupt\" the node clock"} {"case_id":"12978100","cluster":"DISTRACTOR-CASSANDRA-11995","comments":[{"body":"Got the same error, using DSE Enterprise, but reported Cassandra version is 3.0.7.1159\n StorageService.java:625 - Cassandra version: 3.0.7.1159\n\n{noformat}\nERROR [main] 2016-07-27 17:06:50,956 JVMStabilityInspector.java:82 - Exiting due to error while processing commit log during initialization.\norg.apache.cassandra.db.commitlog.CommitLogReplayer$CommitLogReplayException: Could not read commit log descriptor in file /ssd/cassandra/commitlog/CommitLog-6-1469633515550.log\n at org.apache.cassandra.db.commitlog.CommitLogReplayer.handleReplayError(CommitLogReplayer.java:650) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at org.apache.cassandra.db.commitlog.CommitLogReplayer.recover(CommitLogReplayer.java:327) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at org.apache.cassandra.db.commitlog.CommitLogReplayer.recover(CommitLogReplayer.java:148) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at org.apache.cassandra.db.commitlog.CommitLog.recover(CommitLog.java:181) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at org.apache.cassandra.db.commitlog.CommitLog.recover(CommitLog.java:161) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:289) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at com.datastax.bdp.server.DseDaemon.setup(DseDaemon.java:440) [dse-core-5.0.1.jar:5.0.1]\n at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:557) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at com.datastax.bdp.DseModule.main(DseModule.java:49) [dse-core-5.0.1.jar:5.0.1]\nINFO [Daemon shutdown] 2016-07-27 17:06:50,958 DseDaemon.java:549 - DSE shutting down...\n{noformat}\n","created":"2016-07-27T17:22:18.788+0000"},{"body":"Just happened to me again.\n[~jogarcia], are you also running on Windows?","created":"2016-09-29T15:24:48.273+0000"},{"body":"We observed this phenomenon running Apache Cassandra 3.7 on Windows Server 2012 VM. The precise history and state of the VM in question was unclear, but apparently the administrators took a live volume snapshot, and after reverting to the snapshot, Cassandra would no longer start due to the error above. In this case, the commit log file was 48MB of NUL bytes. The Cassandra system log showed an ungraceful termination and a handful of warnings related to slow commit log syncs, but no errors prior to the CommitLogReplayException on subsequent startups.\n\nCassandra was configured for batch mode commit log syncing with a window size of 50ms. It's not clear whether the system was writing to Cassandra at the time of the snapshot.","created":"2016-11-07T15:28:58.167+0000"},{"body":"I'm not too familiar with the Cassandra code base, but it looks like the CommitLogSegment.createSegment() method returns an instance that has had the log header written, but not synced. So until a sync is triggered (periodic, batch, or graceful shutdown), the segment file is left in a corrupt state because it has no header on disk. If I'm right about this, I'd suggest making CommitLogSegment.createSegment() call sync() on the new segment before returning it.","created":"2016-11-07T16:14:38.096+0000"},{"body":"Mine is running on Ubuntu. What I found was that the specific node had a bad memory stick. The machine in question was crashing constantly, every few hours. I think when the crash happens at the wrong moment, there is a log being created and initialized with NULL characters, or probably just allocated to be sure it gets enough space. When the crash happens at that time, the log stays allocated with NULL characters and that prevents the cassandra node from starting, as its format is incorrect. I think what is needed after reboot is to check for that condition and realize the failire and roll it back. and restart.\n\nIn our specific case, we replaced the memory stick, the machine stopped crashing and I've never seen the error again. Probably the condition is rare and only happens during bad timing of a crash, but loosing power on a large cluster would make such condition appear with some frequency. ","created":"2016-11-07T16:27:03.763+0000"},{"body":"I've been seeing this issue quite frequently on my dev machine.\n\nFor various reasons, that machine is currently a bit unstable and tends to lock up quite often, meaning I'm forced to hit the reset button on the front, which results in ungraceful termination of all the running programs. Since I use it for other things, in some boots of the system the Cassandra server starts but I never actually connect to it or use it in any way. In those cases, if I perform a forced reset of the machine while Cassandra is in a \"running but hasn't been accessed yet\" state, I almost always encounter this issue.\n\nI was able to work around the problem by just deleting the commit log that was filled with NULs. I checked the tables and data and nothing seems to be missing in the database, so this definitely seems like a minor glitch. It looks like the Cassandra server is creating an empty commit log file, but the header for it isn't being synced to disk, possibly because there haven't been any changes to data happening which would warrant bothering with one. Then if there is a forced reset while the server is in this state, it isn't able to start up again on reboot because there is a commit log filled with NUL bytes that lacks a header.","created":"2017-02-10T12:01:47.841+0000"},{"body":"|| branch || utests || dtests ||\n| [3.0|https://github.com/jeffjirsa/cassandra/tree/cassandra-3.0-11995] | [testall|http://cassci.datastax.com/job/jeffjirsa-cassandra-3.0-11995-testall/] | [dtest|http://cassci.datastax.com/job/jeffjirsa-cassandra-3.0-11995-dtest/] |\n| [3.11|https://github.com/jeffjirsa/cassandra/tree/cassandra-3.11-11995] | [testall|http://cassci.datastax.com/job/jeffjirsa-cassandra-3.11-11995-testall/] | [dtest|http://cassci.datastax.com/job/jeffjirsa-cassandra-3.11-11995-dtest/] |\n| [trunk|https://github.com/jeffjirsa/cassandra/tree/cassandra-11995] | [testall|http://cassci.datastax.com/job/jeffjirsa-cassandra-11995-testall/] | [dtest|http://cassci.datastax.com/job/jeffjirsa-cassandra-11995-dtest/] |\n\nNote to reviewer: [~aweisberg] and I talked about this offline a bit, and one of the things worth questioning is \"how do we even get in this position\". It seems like there may be a window after [CommitLogDescriptor.writeHeader|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/commitlog/CommitLogSegment.java#L168-L171] is called where we don't actually sync, but short of a system reboot, we should still have the data in memory and the kernel should keep it consistent - however, if we're crashing for some other reason, we could certainly have an all-0 file, which will fail to replay. We may want to open up a subsequent JIRA to talk address that particular problem, but we see it as distinct from the replay problem. \n\nThis patch, then, is only dealing with the problem of replaying the final all-0 file, which we consider to be a change in behavior from 2.x. Continuing to replay a \"corrupt\" all-null file is the 2.x behavior, and presumably should only allowed if we're the last segment, which we already explicitly tolerate in the rest of that segment via {{tolerateTruncation}} flag - this patch just makes {{tolerateTruncation}} also tolerate truncation of the header without interrupting replay and startup.\n","created":"2017-03-17T21:26:01.553+0000"},{"body":"[~blambov] or [~aweisberg] - either of you interested in reviewing? I can rebase as needed.\n","created":"2017-07-22T01:01:54.209+0000"},{"body":"The patch looks good.\n\nDid this happen during normal shutdown? The log is supposed to clean-up such unused segments, but probably isn't able to delete them on Windows because they are still memory-mapped. Still, they should have a header if that's the case, shouldn't they?\n\nThere may also be a further complication related to this. If the empty segment did not sync on a crash, the one before it may be truncated too. We may have to treat that situation as tolerable too.","created":"2017-07-24T14:18:57.180+0000"},{"body":"[~blambov] - I do think there's a few areas for improvement, though this is the low hanging fruit of it. The one case where I was able to see this myself was in a test environment where failures were expected (and indeed, injected), but based on the descriptions above, I believe happens on any unclean shutdown triggered during startup.\n\nI'd personally like to commit this as an intermediate step, and not try to fix ALL of the potential problems at once (I think this removes the most obvious cause, and does it with very little risk). Are you ok with that, and we can open some follow-up tickets as they arise? \n\n\n","created":"2017-07-31T22:55:44.955+0000"},{"body":"Thanks for the 'ready-to-commit' change. Committed as {{0493545dd08d29c34d757bb2f1d90052d03d24c6}} to 3.0 and merged up through 3.11 and trunk.\n\n","created":"2017-09-05T18:51:00.336+0000"}],"conversations":[{"body":"I noticed this morning that Cassandra was failing to start, after being shut down on Friday.\n{code}\nERROR 09:13:37 Exiting due to error while processing commit log during initialization.\norg.apache.cassandra.db.commitlog.CommitLogReplayer$CommitLogReplayException: Could not read commit log descriptor in file C:\\Program Files\\DataStax Community\\data\\commitlog\\CommitLog-5-1465571056722.log\n\tat org.apache.cassandra.db.commitlog.CommitLogReplayer.handleReplayError(CommitLogReplayer.java:622) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.db.commitlog.CommitLogReplayer.recover(CommitLogReplayer.java:302) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.db.commitlog.CommitLogReplayer.recover(CommitLogReplayer.java:147) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.db.commitlog.CommitLog.recover(CommitLog.java:189) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.db.commitlog.CommitLog.recover(CommitLog.java:169) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:273) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:513) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:622) [apache-cassandra-2.2.3.jar:2.2.3]\n{code}\nChecking the referenced file reveals it comprises 33,554,432 (32 * 1024 * 1024) NUL bytes.\nNo logs (stdout, stderr, prunsrv) from the shutdown show any other issues and appear exactly as normal.\nIs installed as a service via DataStax's distribution.","from":"reporter","subject":"Commitlog replaced with all NULs"},{"body":"Got the same error, using DSE Enterprise, but reported Cassandra version is 3.0.7.1159\n StorageService.java:625 - Cassandra version: 3.0.7.1159\n\n{noformat}\nERROR [main] 2016-07-27 17:06:50,956 JVMStabilityInspector.java:82 - Exiting due to error while processing commit log during initialization.\norg.apache.cassandra.db.commitlog.CommitLogReplayer$CommitLogReplayException: Could not read commit log descriptor in file /ssd/cassandra/commitlog/CommitLog-6-1469633515550.log\n at org.apache.cassandra.db.commitlog.CommitLogReplayer.handleReplayError(CommitLogReplayer.java:650) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at org.apache.cassandra.db.commitlog.CommitLogReplayer.recover(CommitLogReplayer.java:327) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at org.apache.cassandra.db.commitlog.CommitLogReplayer.recover(CommitLogReplayer.java:148) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at org.apache.cassandra.db.commitlog.CommitLog.recover(CommitLog.java:181) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at org.apache.cassandra.db.commitlog.CommitLog.recover(CommitLog.java:161) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:289) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at com.datastax.bdp.server.DseDaemon.setup(DseDaemon.java:440) [dse-core-5.0.1.jar:5.0.1]\n at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:557) [cassandra-all-3.0.7.1159.jar:3.0.7.1159]\n at com.datastax.bdp.DseModule.main(DseModule.java:49) [dse-core-5.0.1.jar:5.0.1]\nINFO [Daemon shutdown] 2016-07-27 17:06:50,958 DseDaemon.java:549 - DSE shutting down...\n{noformat}\n","from":"developer"},{"body":"Just happened to me again.\n[~jogarcia], are you also running on Windows?","from":"developer"},{"body":"We observed this phenomenon running Apache Cassandra 3.7 on Windows Server 2012 VM. The precise history and state of the VM in question was unclear, but apparently the administrators took a live volume snapshot, and after reverting to the snapshot, Cassandra would no longer start due to the error above. In this case, the commit log file was 48MB of NUL bytes. The Cassandra system log showed an ungraceful termination and a handful of warnings related to slow commit log syncs, but no errors prior to the CommitLogReplayException on subsequent startups.\n\nCassandra was configured for batch mode commit log syncing with a window size of 50ms. It's not clear whether the system was writing to Cassandra at the time of the snapshot.","from":"developer"},{"body":"I'm not too familiar with the Cassandra code base, but it looks like the CommitLogSegment.createSegment() method returns an instance that has had the log header written, but not synced. So until a sync is triggered (periodic, batch, or graceful shutdown), the segment file is left in a corrupt state because it has no header on disk. If I'm right about this, I'd suggest making CommitLogSegment.createSegment() call sync() on the new segment before returning it.","from":"developer"},{"body":"Mine is running on Ubuntu. What I found was that the specific node had a bad memory stick. The machine in question was crashing constantly, every few hours. I think when the crash happens at the wrong moment, there is a log being created and initialized with NULL characters, or probably just allocated to be sure it gets enough space. When the crash happens at that time, the log stays allocated with NULL characters and that prevents the cassandra node from starting, as its format is incorrect. I think what is needed after reboot is to check for that condition and realize the failire and roll it back. and restart.\n\nIn our specific case, we replaced the memory stick, the machine stopped crashing and I've never seen the error again. Probably the condition is rare and only happens during bad timing of a crash, but loosing power on a large cluster would make such condition appear with some frequency. ","from":"developer"},{"body":"I've been seeing this issue quite frequently on my dev machine.\n\nFor various reasons, that machine is currently a bit unstable and tends to lock up quite often, meaning I'm forced to hit the reset button on the front, which results in ungraceful termination of all the running programs. Since I use it for other things, in some boots of the system the Cassandra server starts but I never actually connect to it or use it in any way. In those cases, if I perform a forced reset of the machine while Cassandra is in a \"running but hasn't been accessed yet\" state, I almost always encounter this issue.\n\nI was able to work around the problem by just deleting the commit log that was filled with NULs. I checked the tables and data and nothing seems to be missing in the database, so this definitely seems like a minor glitch. It looks like the Cassandra server is creating an empty commit log file, but the header for it isn't being synced to disk, possibly because there haven't been any changes to data happening which would warrant bothering with one. Then if there is a forced reset while the server is in this state, it isn't able to start up again on reboot because there is a commit log filled with NUL bytes that lacks a header.","from":"developer"},{"body":"|| branch || utests || dtests ||\n| [3.0|https://github.com/jeffjirsa/cassandra/tree/cassandra-3.0-11995] | [testall|http://cassci.datastax.com/job/jeffjirsa-cassandra-3.0-11995-testall/] | [dtest|http://cassci.datastax.com/job/jeffjirsa-cassandra-3.0-11995-dtest/] |\n| [3.11|https://github.com/jeffjirsa/cassandra/tree/cassandra-3.11-11995] | [testall|http://cassci.datastax.com/job/jeffjirsa-cassandra-3.11-11995-testall/] | [dtest|http://cassci.datastax.com/job/jeffjirsa-cassandra-3.11-11995-dtest/] |\n| [trunk|https://github.com/jeffjirsa/cassandra/tree/cassandra-11995] | [testall|http://cassci.datastax.com/job/jeffjirsa-cassandra-11995-testall/] | [dtest|http://cassci.datastax.com/job/jeffjirsa-cassandra-11995-dtest/] |\n\nNote to reviewer: [~aweisberg] and I talked about this offline a bit, and one of the things worth questioning is \"how do we even get in this position\". It seems like there may be a window after [CommitLogDescriptor.writeHeader|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/commitlog/CommitLogSegment.java#L168-L171] is called where we don't actually sync, but short of a system reboot, we should still have the data in memory and the kernel should keep it consistent - however, if we're crashing for some other reason, we could certainly have an all-0 file, which will fail to replay. We may want to open up a subsequent JIRA to talk address that particular problem, but we see it as distinct from the replay problem. \n\nThis patch, then, is only dealing with the problem of replaying the final all-0 file, which we consider to be a change in behavior from 2.x. Continuing to replay a \"corrupt\" all-null file is the 2.x behavior, and presumably should only allowed if we're the last segment, which we already explicitly tolerate in the rest of that segment via {{tolerateTruncation}} flag - this patch just makes {{tolerateTruncation}} also tolerate truncation of the header without interrupting replay and startup.\n","from":"developer"},{"body":"[~blambov] or [~aweisberg] - either of you interested in reviewing? I can rebase as needed.\n","from":"developer"},{"body":"The patch looks good.\n\nDid this happen during normal shutdown? The log is supposed to clean-up such unused segments, but probably isn't able to delete them on Windows because they are still memory-mapped. Still, they should have a header if that's the case, shouldn't they?\n\nThere may also be a further complication related to this. If the empty segment did not sync on a crash, the one before it may be truncated too. We may have to treat that situation as tolerable too.","from":"developer"},{"body":"[~blambov] - I do think there's a few areas for improvement, though this is the low hanging fruit of it. The one case where I was able to see this myself was in a test environment where failures were expected (and indeed, injected), but based on the descriptions above, I believe happens on any unclean shutdown triggered during startup.\n\nI'd personally like to commit this as an intermediate step, and not try to fix ALL of the potential problems at once (I think this removes the most obvious cause, and does it with very little risk). Are you ok with that, and we can open some follow-up tickets as they arise? \n\n\n","from":"developer"},{"body":"Thanks for the 'ready-to-commit' change. Committed as {{0493545dd08d29c34d757bb2f1d90052d03d24c6}} to 3.0 and merged up through 3.11 and trunk.\n\n","from":"developer"}],"created":"2016-06-13T10:05:11.000+0000","description":"I noticed this morning that Cassandra was failing to start, after being shut down on Friday.\n{code}\nERROR 09:13:37 Exiting due to error while processing commit log during initialization.\norg.apache.cassandra.db.commitlog.CommitLogReplayer$CommitLogReplayException: Could not read commit log descriptor in file C:\\Program Files\\DataStax Community\\data\\commitlog\\CommitLog-5-1465571056722.log\n\tat org.apache.cassandra.db.commitlog.CommitLogReplayer.handleReplayError(CommitLogReplayer.java:622) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.db.commitlog.CommitLogReplayer.recover(CommitLogReplayer.java:302) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.db.commitlog.CommitLogReplayer.recover(CommitLogReplayer.java:147) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.db.commitlog.CommitLog.recover(CommitLog.java:189) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.db.commitlog.CommitLog.recover(CommitLog.java:169) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:273) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:513) [apache-cassandra-2.2.3.jar:2.2.3]\n\tat org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:622) [apache-cassandra-2.2.3.jar:2.2.3]\n{code}\nChecking the referenced file reveals it comprises 33,554,432 (32 * 1024 * 1024) NUL bytes.\nNo logs (stdout, stderr, prunsrv) from the shutdown show any other issues and appear exactly as normal.\nIs installed as a service via DataStax's distribution.","issue_id":"12978100","key":"CASSANDRA-11995","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-09-05T18:51:00.000+0000","role":"fixed_distractor","summary":"Commitlog replaced with all NULs"} {"case_id":"12986267","cluster":"DISTRACTOR-CASSANDRA-12126","comments":[{"body":"cc [~slebresne] and [~jbellis]","created":"2016-07-01T18:05:11.346+0000"},{"body":"I \"think\" you are right that this can happen, and that committing an empty commit on SERIAL reads \"should\" fix it. Paxos is however subtle enough that I would feel more confident with this if we had a reproduction test first, if only so we can validate whatever patch we come up with. [~jkni] I believe you've spend some time on jespen-like tests for paxos, do you think this is something we could use to try to reproduce something like that relatively consistently?","created":"2016-07-05T09:33:51.612+0000"},{"body":"Sure - I have Jepsen tests for LWT and a few other similar tests I've been building. I haven't seen an issue like this reproduced through them so far, but I'll try and see if I can reproduce this reliably. It certainly seems feasible to hit. I definitely think it is a good idea to run any proposed patch through LWT tests for a while.","created":"2016-07-05T15:02:09.391+0000"},{"body":"I was able to repro it in a not so good way :) Here are the steps\n1. Create a 3 node C* cluster(A,B and C)\n2. Create a keyspace with RF=3 and a simple table\n3. Insert System.exit PaxosState.propose method to simulate a failure in A and B. \n4. Do a CAS Write. Now with this, only C will be able to store the propose and not A and B. \n4. Now bring down machine C and remove the System.exit from A and B and bring them up again. \n5. Do a CAS Read and this will involve only A and B and will not return the data. \n6. Bring down A and bring up C. \n7. Do a CAS Read and you will be able to read the data. \n\nIf we can use some test framework to simulate such failures then that will be better. ","created":"2016-07-05T18:04:32.687+0000"},{"body":"Another take on how to test coordination aspects for this ticket would be to make use of the MessagingService mocking classes implemented in CASSANDRA-12016. I've created a couple of tests [here|https://github.com/spodkowinski/cassandra/tree/WIP-12126/test/unit/org/apache/cassandra/service/paxos] to get a better idea how this would look like. Although limited to observing the behavior of a single node/state machine, it's probably more lightweight and easier to implement than doing the same using dtests or Jepsen.\n\nAs of the described edge case, I'd agree with [~kohlisankalp]'s suggestion (if I understood correctly) to do an additional proposal round. However, it would be nice to optimize this a bit so we don't trigger new proposals for each and every SERIAL read. I did a first implementation for this [here|https://github.com/spodkowinski/cassandra/commit/96ec151992f49c773e5af5d85ce69ec87d8b7bc5] (with [CASReadTriggerEmptyProposal|https://github.com/spodkowinski/cassandra/blob/WIP-12126/test/unit/org/apache/cassandra/service/paxos/CASReadTriggerEmptyProposal.java] as corresponding test) for the sake of discussion. ","created":"2016-07-12T10:41:15.307+0000"},{"body":"I've now revisited this issue again and took a closer look at the work I've done back then months ago (after rebasing on trunk). The patch follows the suggested solution by [~kohlisankalp] by using an empty commit as additional propose step. It also implements an optimization to avoid this step in case no paxos rounds for writing new values have been conducted between serial reads.\n\n* [CASSANDRA-12126-trunk|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12126-trunk]\n* [CircleCI results|https://circleci.com/gh/spodkowinski/cassandra/5]\n\nExcept the actual fix, there's also a lot of tests included, which I'd hate to throw away, as we're clearly lacking tests for CAS operations.\n","created":"2017-04-07T10:25:26.723+0000"},{"body":"I see you outlining two series of steps:\n\n1 -> 2 -> 3. The value V from 1 is not seen in 2, but once it is seen in 3 it is always seen.\n\n1 -> 2 -> 4. V is never seen.\n\nIt seems to me that both of these maintain linearizability. What am I missing?","created":"2017-04-18T14:42:08.949+0000"},{"body":"bq. 1 -> 2 -> 3. The value V from 1 is not seen in 2, but once it is seen in 3 it is always seen.\n\nAs 1 \"completes\" before 2, it's result should be visible by 2 (or not ever) for linearizability (taken in the sense discussed here for instance: http://www.bailis.org/blog/linearizability-versus-serializability/). More pragmatically, outside of any theoretical definition, if serial read can't guarantee they will see any previous operation (even failed one, as long as they returned to the client), then they are not very useful in the first place.","created":"2017-04-18T15:01:57.909+0000"},{"body":"But we stipulated that 1 times out and did not complete. (If it did complete it would be guaranteed to be visible by any majority of course.)","created":"2017-04-18T15:08:23.979+0000"},{"body":"Bailis:\n\n# once a write completes, all later reads (where “later” is defined by wall-clock start time) should return the value of that write or the value of a later write. \n# Once a read returns a particular value, all later reads should return that value or the value of a later write.\n\nI think we all agree that our current behavior satisfies (2). I am saying that we actually also satisfy (1) because the write is not complete until Sankalp's step 3.\n","created":"2017-04-18T15:12:50.449+0000"},{"body":"bq. But we stipulated that 1 times out and did not complete.\n\nIt completed in the sense that the operation is finished and returned to the client (albeit with a timeout). Don't get me wrong, theory is blind, so if you want to define than an operation completes only if it \"finished and did not timeout\" and define linearizability only in term of completing operation (with that definition of completion), then sure, I agree this particular definition of linearizability is not broken by the description of this ticket.\n\nBut how useful is a definition of linearizability that says nothing about operations that timeouts (especially keeping in mind that our implementation is particularly prone to timeouts)?","created":"2017-04-18T15:16:38.810+0000"},{"body":"bq. I am saying that we actually also satisfy (1) because the write is not complete until Sankalp's step 3.\n\nYour definition of \"completion\" doesn't work logically. You're suggesting (1) only complete when it is visible (in (3)), but linearizabilty is about when operations that complete are visible, so if you define completion as \"when it's visible\", you're in for trouble.","created":"2017-04-18T15:21:03.723+0000"},{"body":"I'm confused, because it sounds like you're saying \"all operations should be visible once finished.\" Of course that's not actually what you mean that would require participation from all replicas to finish in-flight requests and not just a majority. What is the distinction you are proposing?","created":"2017-04-18T17:31:11.245+0000"},{"body":"bq. What is the distinction you are proposing?\n\nNot sure, I think we don't put the same definitions on operation visibility. What I'm saying is that \"if an operation has a visible outcome, then that outcome should be visible (by serial operations) by any subsequent operation (so as soon as the operation returns to the client if you will)\". In particular, if a serial read follows a serial write (meaning that it's started after the write returned, even with a timeout), then if the write has any effect, the read should see it.\n\nNote that when you get a timeout on the initial write, you don't know if the write has been applied or not, but the whole point of a serial read is to be able to unequivocally decide what was that outcome. If we can't guarantee that, if there is no way to observe if a timed-out write has been applied or not, then I'm not sure how one would use LWT in the first place.\n","created":"2017-04-19T08:58:18.956+0000"},{"body":"I see. So you are saying that\n\n1: Write\n2: Read -> Nothing\n3: Read -> Something\n\nIs broken because to go from Nothing to Something [in a linearized system] there needs to be a write in between.","created":"2017-04-19T14:56:12.074+0000"},{"body":"Exactly.","created":"2017-04-19T15:03:26.677+0000"},{"body":"Hi all,\r\n\r\nOur team from UCARE University of Chicago, have been able to reproduce similar manifestation to this bug consistently with our model checker. (Our scenario is different with what [~kohlisankalp] proposed)\r\nHere are the workload and scenario of the bug:\r\n\r\nWorkload: 3 nodes-cluster, 3 client requests (but no crash event)\r\n\r\nScenario:\r\n # Start 3-nodes cluster and inject all of 3 client requests to 3 different nodes (node X, Y, Z)\r\n # Node X sends its prepare messages (ballot number=1) to all nodes and all nodes accept it\r\n # Node X sends its propose message to itself, causing its inProgress value to be \"X\".\r\n # Node Y sends its prepare messages (ballot number=2) to all nodes.\r\nThis also causes the rest of node X propose messages to be invalid because its ballot number is smaller than node Y prepare messages.\r\n # In our scenario, the prepare response messages from node Y and Z comes first before prepare response message from node X, causing the node Y to unrecognize the state of node X which already accepted value \"X\" (step 3).\r\n # But since our query of client request 2 has an IF, that said IF value_1='X', therefore node Y will not continue on sending propose messages to all nodes.\r\nUp to this point, it means none of the queries are committed to the server.\r\n # Node Z now sends its prepare messages to all nodes and all nodes accept it.\r\n # In our scenario, now the node X returns its response first where it also let node Z knows about its inProgress Value \"X\".\r\nFrom here, node Z will propose and commit client request-1 (with value \"X\") instead of client-request-3.\r\n # Therefore, we ended up having client request 1 stored to the server, although client request-3 was the one that is said successful.\r\n\r\nWe are ready to assist, if any further information is needed.\r\n\r\n \r\n\r\n ","created":"2018-09-26T21:55:59.397+0000"},{"body":"[~jeffreyflukman] thanks for this report. Suspect that most of the folks who are interested in this are already cc'd and received an email notification of your response, but explicitly tagging [~benedict] [~iamaleksey] and [~bdeggleston] as people who aren't yet watching it but may be interested.\r\n\r\nAlso, very much interested in the model you mentioned - is that available publicly at this point? ","created":"2018-09-26T23:46:28.504+0000"},{"body":"To complete our scenario, here is the setup for our Cassandra:\r\n We run the scenario with Cassandra-v2.0.15.\r\n Here is the scheme that we use:\r\n * CREATE KEYSPACE test WITH REPLICATION = \\{'class': 'SimpleStrategy', 'replication_factor': 3};\r\n * CREATE TABLE tests ( name text PRIMARY KEY, owner text, value_1 text, value_2 text, value_3 text);\r\n\r\nHere are the queries that we submit:\r\n * client request to node X (1st): UPDATE test.tests SET value_1 = 'A' WHERE name = 'testing' IF owner = 'user_1';\r\n * client request to node Y (2nd): UPDATE test.tests SET value_2 = 'B' WHERE name = 'testing' IF value_1 = 'A';\r\n * client request to node Z (3rd): UPDATE test.tests SET value_3 = 'C' WHERE name = 'testing' IF value_1 = 'A';\r\n\r\nTo confirm, when the bug is manifested, the end result will be: value_1='A', value_2=null, value_3=null\r\n\r\n[~jjirsa], regarding our tool, at this point, it is not open for public. ","created":"2018-09-27T02:50:42.506+0000"},{"body":"Why is end result not correct? second and third operation did not succeed because first 1 did not finish? Can you combine the example with the earlier comment please","created":"2018-09-27T02:58:40.566+0000"},{"body":"[~jeffreyflukman] it would help if you could explicitly state the client responses returned for each of your operations. The options are: time out, rejected (condition not met), success (condition met, and mutation applied)\r\n\r\nFor completeness, as with CASSANDRA-12438, the read queries you are performing, to which nodes, at what point and with what consistency levels would be helpful to know. Are you verifying the state with a SERIAL read after the last query, most specifically?  Also, can we assume that the state of the table began with \\\\{name:'testing', owner:'user_1', value1:null, value2:null, value3:null}\\?","created":"2018-09-27T08:44:34.699+0000"},{"body":"During our testing with our model checker, we limit the round of Paxos for each query,\r\nbecause if not, it is possible that we get stuck in a very long sequence of message transactions among the nodes without progressing anywhere. So, what we do is we only execute one round of Paxos for each query.\r\n\r\nTo enlight our test and combine our whole story, here is what happened in detail:\r\n * We first prepared the 3 node-cluster with the test.tests table as initial table structure and yes, the initial table began with:\r\n{name:'testing', owner:'user_1', value1:null, value2:null, value3:null}\r\n * Next, we run the model checker that will start the 3 node-cluster.\r\n * Inject the 3 client requests in order: query 1, then query 2, then query 3.\r\nThis cause query 1 to have ballot number < query 2 ballot number < query 3 ballot number.\r\n * Now this means, in the beginning, the model checker already see there will be 9 prepare messages in its queue that will be reordered in some way.\r\n * When the bug is manifested, we ended up having:\r\n ** Node X's prepare messages proceed and all nodes response with true back to node X.\r\n ** Node X sends its propose message with value_1='A' to itself first and get a response true as well.\r\n ** At this moment, Node X inProgress value is updated to the proposed value, value_1='A'\r\n ** But then node Y prepare messages proceed and all nodes response with true back to node Y,\r\nbecause prepare messages of node Y have a higher ballot number.\r\n ** But when node Y about to proceed the propose messages it realized that the current data does not fulfill the IF condition, so it does not proceed to propose messages. --> Client request 2 to node Y is therefore rejected\r\n ** Continuing node X propose messages to node Y and Z, both requests are returned with false to node X\r\n ** Now at this point node X should be able to retry the Paxos with a higher ballot number, but since we limit the round of Paxos for each query to one, therefore client request 1 to node X is timed out.\r\n ** Lastly, node Z sends its prepare messages to all nodes, and get response true messages from all nodes,\r\nbecause the ballot number is higher as well.\r\n ** At this point, if the node X response message is returned first to node X, what will happen is node Z will realize that node X still has an inProgress value in the process (value_1='A'). This cause node Z to send propose messages and commit messages but for client request 1 using the current highest ballot number.\r\nHere we have our first data update saved: value_1='A', value_2=null, value_3=null.\r\n ** Back to our constraint of one round Paxos for each query, we ended up not retrying client request-3 because we reached timeout.\r\n * To sum up:\r\n ** client request-1: Timed out\r\n ** client request-2: Rejected\r\n ** client request-3: Timed out\r\n\r\nThere we get an inconsistency between the client side and the server side, where all requests actually failed, but when we read the end result again from all nodes, we get value_1='A', value_2=null, value_3=null.\r\n\r\n \r\n\r\nI made a wrong statement at the end of my first comment:\r\n{quote}9. Therefore, we ended up having client request 1 stored to the server, although client request-3 was the one that is said successful.\r\n{quote}\r\nIt should be failed due to timeout.","created":"2018-09-27T16:20:08.110+0000"},{"body":"bq. client request-1: Timed out, client request-2: Rejected, client request-3: Timed out\r\n\r\nGiven those responses to the queries. The client side does not know the state of the system without issuing a READ at SERIAL (or doing another INSERT that gets a success which the state can be inferred from).\r\n\r\nbq. There we get an inconsistency between the client side and the server side, where all requests actually failed, but when we read the end result again from all nodes, we get value_1='A', value_2=null, value_3=null.\r\n\r\nGiven the responses you got, there is no inconsistency. The client received \"timed out\" exceptions. A timed out exception means \"your query may or may not have been applied, the server doesn't know, you should retry it if you want to ensure it goes through\". In this case request-1 was successful, and request-3 failed. So {{value_1='A', value_2=null, value_3=null}} is a valid state and not inconsistent.","created":"2018-09-27T16:30:44.203+0000"},{"body":"I agree with [~jjordan] that this is a correct response. \r\n\r\nAlso in the future please open a new Jira if it is a different issue. ","created":"2018-09-27T17:28:37.567+0000"},{"body":"Thank you for your responses, [~jjordan] and [~kohlisankalp].\r\nI think you have cleared up some misunderstandings for me (and our team) where timeout is a \"gray area\" for the client \r\nto determine whether a request has been successfully processed.\r\n\r\nOne thing that we would like to point out maybe, based on the early discussion in this bug description, quote\r\n{quote}However we need to fix step 2, since it caused reads to not be linearizable with respect to writes and other reads. In this case, we know that majority of acceptors have no inflight commit which means we have majority that nothing was accepted by majority. I think we should run a propose step here with empty commit and that will cause write written in step 1 to not be visible ever after.\r\n{quote}\r\nWhat we tried to mimic with our model checker in the beginning actually was this scenario where node Y saw that the majority of nodes do not have inProgress value, but then suddenly node Z saw that there is an inProgress value from node X and tried to repair and commit it.\r\nSo, we confirm that we can also see this behavior:\r\n{quote}2: Read -> Nothing\r\n3: Read -> Something\r\n{quote}\r\nWe read nothing in node Y, yet node Z read something in the next request.\r\n\r\n\r\n\r\nTo sum up, at least, our scenario explains this behavior: Node Y does not try to repair the Paxos because node X's prepare response comes last, therefore node Y ignores the node X's prepare response and based its decision to not repair the Paxos.\r\nBut in node Z's client request, node Z decides to repair the Paxos based on node X's existing inProgress value_1=\"A\" because node X's prepare response comes early (1st or 2nd). Which cause an inconsistent reaction in some way between node Y and node Z (although this is correct based on the original Paxos algorithm).\r\n\r\n\r\nA solution to avoid this inconsistent reactions from these two nodes maybe is for each node to decide whether to repair a Paxos or not based on the complete view of the alive nodes, therefore if the response X's comes last with an inProgress value, node Y will still repair the Paxos.","created":"2018-09-27T18:15:30.568+0000"},{"body":"{quote}We read nothing in node Y, yet node Z read something in the next request.\r\n{quote}\r\nI think the problem here is that, at the API level, there isn't enough information to say that X didn't simply 'occur' *after* both Y and Z. That is, unless the rejection of Y occurs after X's timeout. In this case, it would seem to be an API-visible error, as at the point of timeout the indeterminacy should be fixed. Timeouts should not ‘live forever’ as the bogeyman, ready to mess with history.\r\n\r\nI think, though, that the suggested mechanism could result in this.\r\n\r\nTake three nodes (RF=3) A, B and C; and any three CAS operations X, Y and Z such that:\r\n * X and Y can always succeed\r\n * Z can only succeed if X has succeeded\r\n\r\nSetup:\r\n # Prepare _and_ Propose X with ballot 1; proposal accepted only by A\r\n ** this will be the last and only node’s proposal acceptance\r\n # Prepare Y with ballot 2; reach B and C before ballot 1, so they do not accept\r\n # Now, lock X and Y in battle, always failing to proceed to the propose step before the other reaches the prepare step again\r\n # X and Y both timeout having failed to cleanly apply\r\n\r\nPart 2:\r\n # Z is now attempted; it prepares to only B and C, seeing no in-progress proposal\r\n # As a result, it does not see X; it is rejected, so there is no new proposal/commit\r\n # Z is attempted again; this time, A is consulted\r\n # Suddenly, a wild X appears. From nowhere.  Z succeeds, despite no intervening operations.\r\n\r\nIt does seem, in essence, to be an incidence of the bug (or a very similar one) described in the ticket.","created":"2018-09-28T00:56:45.055+0000"},{"body":"Having reviewed the code I think what Benedict says is correct. The criteria we use for identifying if there is an progress paxos round that needs resolution is incorrect because it assumes we have visibility to all accepted ballots when we only have visibility to a majority.\r\n\r\nI think this optimization can be done correctly, but it's a bit of surgery. Right now reads do a prepare and modify the promised ballot at each acceptor. If instead we only read the promised ballot from each acceptor then we could check the promised ballot matches the most recent committed ballot. If those are the same we know nothing is in progress because a higher ballot than the most recent accepted/committed ballot has not been promised by a majority which means there can be no lingering accepted ballot since a majority of promises must be collected first.\r\n\r\nIf that isn't the case then we can go ahead and do a prepare + propose to make them match and subsequent reads won't have to do a propose.\r\n\r\nThis may also impact our choice of how many replicas to contact in each phase since we want them to have consistent paxos state so reads can be one roundtrip. I am not sure if we contact them all (like with mutations) or just a majority.","created":"2018-12-21T17:38:48.275+0000"},{"body":"It looks like I never noted that IMO the real problem is that we do not serialize the evaluation of a condition if the condition is not met, so other commands are also not serialized wrt the evaluation of such conditions either.  The evaluation of a condition should always be a serialized wrt other events, so we should be agreeing on the conditional operation, and then performing it, not evaluating the condition and then serializing the choice of a new value.","created":"2019-10-31T16:44:35.375+0000"},{"body":"It definitively doesn't look good that this messages comes so late, but I feel this is a serious issue of the {{SERIAL}}/{{LOCAL_SERIAL}} consistency levels since this breaks the basic guarantee they exist to provide, and as such should be fixed all the way down 3.0, and the sooner, the better.\r\n\r\nIn an attempt to sum this up quickly, the problem we have here affects both serial reads _and_ LWT updates that do not apply (whose condition evaluates to {{false}}). In both case, while the current code replays \"effectively committed\" proposals (those whose proposal has been accepted by majority of replica) with {{beginAndRepairPaxos}}, neither make proposals of their own, so nothing will prevent a proposal accepted by a minority of replica (say just one) to be later replayed (and thus committed).\r\n\r\nI've pushed [2 in-jvm dtests|https://github.com/pcmanus/cassandra/commit/3442277905362b38e0d6a2b8170916fcfd18d469] that demonstrate the issue for both cases (again, serial reads and non-applying updates). They use \"filters\" to selectively drop messages to make failure consistent but aren't otherwise very involved.\r\n\r\nAs [~kohlisankalp] mentioned initially, the \"simplest\"\\[1\\] way to fix this that I see is to commit an empty update in both cases. Actually committing, which sets the {{mostRecentCommit}} value in the Paxos state, ensures that no prior proposal can ever be replayed. I've pushed a patch to do so on 3.0/3.11 below (will merge up on 4.0, but wanted to make sure we're ok on the approach first):\r\n\r\n||version||\r\n| [3.0|https://github.com/pcmanus/cassandra/commits/C-12126-3.0] |\r\n| [3.11|https://github.com/pcmanus/cassandra/commits/C-12126-3.11] |\r\n\r\nThe big downside of this patch however is the performance impact. Currently, a {{SERIAL}} read (that finds nothing in progress it needs to replay) is 2 round-trips (a prepare phase, followed by the actual read). With this patch, it is 3 round-trips as we have to propose our empty commit and get acceptance (we don't really have to wait for responses on the commit though), which will be noticeable for performance sensitive use-cases. Similarly, the performance of LWT that don't apply will be impacted.\r\n\r\nThat said, I don't seen another approach to fixing this that would be as acceptable for 3.0/3.11 in terms of risks, and imo 'slower but correct' beats 'faster but broken' any day, so I'm in favor of moving forward with this fix.\r\n\r\nOpinions?\r\n\r\n\r\n\r\n\\[1\\]: I mean by that both the simplicity of the change, but also of validating that this fix the problem at hand without creating new correctness problems.\r\n","created":"2020-05-26T09:56:07.482+0000"},{"body":"Agreed that's the most straightforward way to address both issues (although I've only skimmed your patch).\r\n\r\nIn 3.x though, and at least for the serial read fix, I think we should include a flag to disable the fix, in case a) there's a problem with the fix or b) operators would rather trade the performance impact for linearizability for whatever reason.\r\n\r\nThere's also a variant of the non-applying update issue where it's exposed by a read, not another insert. It would be good to have a test for that as well.","created":"2020-05-26T22:08:47.857+0000"},{"body":"So, a thought has occurred to me: what do we actually claim our consistency properties are for SERIAL? \r\n\r\nMy understanding was that we claimed only serializability, in which case I don't think that strictly speaking this is a bug. I think it's only a bug if we claim strict serializability. However the only docs I can find claiming either are DataStax's which mixes linearizable up with serializable.\r\n\r\nFWIW, I consider this to be a bug, as we should at least support the more intuitively correct semantics. But perhaps we should instead introduce a new STRICT_SERIAL consistency level to solve it, and clarify what SERIAL means in our docs?\r\n\r\nI would also be OK with simply claiming strict serializability for SERIAL. But perhaps this technicality/ambiguity buys us some time and cover to solve the problem without introducing major performance penalties?\r\n\r\nI also have some relevant test cases I will share tomorrow, along with test cases for other correctness failures of LWTs.","created":"2020-05-26T23:31:40.658+0000"},{"body":"I've pushed various test cases [here|https://github.com/belliottsmith/cassandra/tree/12126-tests-3.0] - most of them marked {{@Ignore}} because they are known to fail, and won't be resolved immediately.","created":"2020-05-27T09:59:10.711+0000"},{"body":"{quote}I think we should include a flag to disable the fix\r\n{quote}\r\nThe option of having a flag occurred to me, but I rejected it initially because I continue to believe the current behavior is wrong (a moral judgment, I guess) and in principle, having a \"please, make my database broken\" flag does not feel like a good idea.\r\n\r\nBut I reckon that it _may_ exists advanced users that did noticed the lack of linearizability for reads and effectively built around it knowingly, for which the performance impact may be considered a regression with no upside (but if you sense skepticism on my part when reading that sentence, you're radar is not completely off).\r\n\r\nAnd as we're talking minor upgrade here, I'm amenable to such flag, though I'd prefer making it clear somehow that it is unsafe/risky and something we may remove in the future with no particular warning.\r\n{quote}It would be good to have a test for that as well.\r\n{quote}\r\nCertainly, good point, I can add the 2 missing interleaving.\r\n{quote}do we actually claim our consistency properties are for SERIAL?\r\n{quote}\r\nWhile our official doc on the matter is certainly lacking (not spelling much guarantee at all afaict, and I'm happy to piggy-back on this ticket to correct that), we've always implied linearizability. I have, at least, and I'm sure I can dig up other doing it as well on the mailing list if necessary. We did this both by throwing the linearizable word out from time to time, but also by repeatedly recommending that when a write times out, one needs to issue a SERIAL read to 'observe' if that write went through or not (and as an aside, if you can't rely on either reads or non-applying CAS for that, I'm not even sure how to use LWTs, except maybe for excessively specific cases).\r\n{quote}perhaps we should instead introduce a new STRICT_SERIAL consistency level\r\n{quote}\r\nI'm rather cold on that because, tbh. I think non-strict serializability is a theoretical notion that is useless in practice and that it is something we should not offer. And I'd rather avoid one more \"feature\" for which we spend our time saying \"don't use it\".\r\n{quote}I've pushed various test cases\r\n{quote}\r\nAwesome, thanks. I'll look at integrating those in the branch if you don't mind.","created":"2020-05-27T10:21:30.072+0000"},{"body":"bq. I'm amenable to such flag\r\n\r\nActually, let me rephrase that a bit. I'd *really* prefer not adding such flag. If someone is ok with serializability without linearizability, then they can use QUORUM reads, and given how things are implemented, it provides (non-strict) serializability. Granted, for someone that uses SERIAL today, is ok with the lack of linearizability and can't afford the performance penalty, it'll require a client side change, which this flag would avoid, so there is not zero value to such flag. But I suspect user fitting that category (knowingly ok with lack of linearizability) is really really small, and we always have to make trade-offs. So in that case I feel adding one more flag, one I consider dangerous, is not worth it. So to clarify, if a consensus appears for such flag, so be it, I'll add it, but I'm personally not neutral either.","created":"2020-05-27T10:37:55.052+0000"},{"body":"bq. I'm rather cold on that because, tbh. I think non-strict serializability is a theoretical notion that is useless in practice and that it is something we should not offer. And I'd rather avoid one more \"feature\" for which we spend our time saying \"don't use it\".\r\n\r\nYeah, I'm very sympathetic to this view, and have always assumed linearizability with partitions as the object. I'm just really trying to morally justify providing some time to fix this without any negative repercussions. \r\n\r\nEither way, we should definitely clarify what we mean by SERIAL in some official project documentation somewhere though. We probably need to do so in terms of strict serializability as opposed to linearizability, so that it can be consistent with a future in which we support multi-partition transactions (which as a project we really need to deliver in the not-too-distant future).\r\n\r\nbq. non-applying CAS for that\r\n\r\nFWIW, I think this particular case is a no-brainer; there's no real cost to strengthening the semantics of non-applying CAS IMO, since users should anticipate their CAS operations will ordinarily take this long. Whatever the conclusion of our discussion, I think we should apply a fix at least for the non-applying case immediately, and I do not believe any flag to disable this part of the fix is necessary.\r\n\r\nReads are trickier, because the user will see a significant performance penalty on patch version upgrade. I'm sympathetic to the view we should just fix the read part immediately, performance regressions be damned. But we do have other serious consistency violations that should also be fixed. I think it is worth _considering_ if we should instead aggressively try to remedy all of the known issues, have a strong verification push, and then roll out all of the changes at-once - including a fix for this that does not regress performance. It might seem a lot for a patch version, but I'm not sure risk is a concern when we know there are several serious problems today, and have been for years.\r\n\r\nI'm not going to advocate super strongly for either approach, as I don't think there's a clear answer, I just want to raise the alternative as an option to expressly consider.\r\n\r\nbq. Awesome, thanks. I'll look at integrating those in the branch if you don't mind.\r\n\r\nAbsolutely, that was my intention.","created":"2020-05-27T11:03:56.731+0000"},{"body":"bq. But we do have other serious consistency violations that should also be fixed.\r\n\r\nCould you expand on that?\r\n","created":"2020-05-27T15:01:26.496+0000"},{"body":"The test cases I provided demonstrate several consistency violations during range movements. I've just thought of another one, and am writing a test case for it. Perhaps we could claim that range movements are always (potentially) consistency violations, but they are particularly keenly felt when you claim a linearisable history.\r\n\r\nThere are also (more debatably) issues with TTL on {{system.paxos}}, particularly when mixed with non-global commit; perhaps we could claim this is the user's problem, but it's not clear why we support global consensus that can be lost through local commit, and I don't think we communicate clearly the consistency implications to not call this a bug.\r\n\r\nAlso, mixing LOCAL_SERIAL and SERIAL is entirely unsafe, and even supporting them both is arguably a consistency violation without mechanisms to safely transition from one level to another.","created":"2020-05-27T15:35:17.558+0000"},{"body":"bq. The test cases I provided demonstrate several consistency violations during range movements.\r\n\r\nYes, sorry I hadn't read them before commenting. And I certainly agree those are problematic (I was about to open a ticket so it's tracked, but I'd say CASSANDRA-15745 kind of cover those).\r\n\r\nbq. There are also (more debatably) issues with TTL on system.paxos\r\n\r\nAgreed this has always been a weak point. It does feel somewhat separated of other consistency points though, and maybe short term we can just offer a way to override the TTL (with documentation on the tradeoffs involved)?\r\n\r\nbq. Also, mixing LOCAL_SERIAL and SERIAL is entirely unsafe\r\n\r\nYeah. I'm not sure how to fix that one without a breaking API change though (namely, limiting their unrestricted use together). It's not \"that\" different from the fact we allow unrestricted mixing of serial and non-serial operations. Which is something I don't like and I'm happy to discuss moving forward, but imo post-3.X material in the best of cases.\r\n\r\nbq. I think it is worth considering if we should instead aggressively try to remedy all of the known issues, have a strong verification push, and then roll out all of the changes at-once - including a fix for this that does not regress performance.\r\n\r\nIt is certainly an option worth bringing, and thank you for that. I'm not sure how to really know what is the best option though, so I can only offer my current opinion.\r\n\r\nWhich is that I feel this issue is a very serious issue. And I don't mean that in a way that diminishes the seriousness of the other problems you mentioned, I mean that in absolute terms (the range movement issues are also fairly bad imo for instance). But leaving less of our known serious unaddressed feels better than not, so I'd personally prefer fixing that issue ASAP. Basically, I'm worried that waiting for a more all-encompassing fix might take us quite some time, with no absolute guarantee that we'll be collectively at ease with pushing that to 3.X.\r\n\r\nAnyway, I'd like to move this forward personally. How do we decide if we do?\r\n","created":"2020-06-05T16:43:34.394+0000"},{"body":"> I'd like to move this forward personally. \r\n\r\nSure, go for it.","created":"2020-06-10T11:05:32.771+0000"},{"body":"Ok, I've rebased the patch against 4.0 and started CI on it all:\r\n||branch||CI||\r\n|[3.0|https://github.com/pcmanus/cassandra/tree/C-12126-3.0]|[Run #146|https://ci-cassandra.apache.org/job/Cassandra-devbranch/146/]|\r\n|[3.11|https://github.com/pcmanus/cassandra/tree/C-12126-3.11]|[Run #147|https://ci-cassandra.apache.org/job/Cassandra-devbranch/147/]|\r\n|[4.0|https://github.com/pcmanus/cassandra/tree/C-12126-4.0]|[Run #148|https://ci-cassandra.apache.org/job/Cassandra-devbranch/148/]|\r\n\r\nI included a commit to add the flag that disables the new empty commit for SERIAL reads as suggested by [~bdeggleston] earlier. Still slightly on the fence on the need for such flag, but I call it \"unsafe\" ({{-Dcassandra.unsafe.disable-serial-reads-linearizability}} to be specific) and log a warning when used, so I'm at peace with that.\r\n\r\nI'll note for future reviewers that while the 3.11 branch is almost a straight away merge up of 3.0, there is a minor differences on the 4.0 branch, namely:\r\n * the added in-jvm dtests needed a few changes to reflect 4.0 changes. To make that easier, I squashed 2 of the commits from the 3.0/3.11 branches, which is why that branch has one less commit.\r\n * There is a few changes related to the translation of {{WriteTimeoutException}} into {{CasWriteTimeoutException}} (I pushed it down in some cases). I believe this fixes a minor \"bug\" where the \"contentions\" number we returned with {{CasWriteTimeoutException}} was potentially inaccurate (namely, if we timed out in {{beginRepairAndPaxos}}, contention leading to that exception would be ignored)\r\n\r\nI'll wait on getting usable CI results to officially mark it 'ready to review', but it is in spirit if anyone is burning to look at this.","created":"2020-06-12T13:55:31.279+0000"},{"body":"I'm only semi-sure how to parse Jenkins CI results these days but from what I can tell, all failures are unrelated so marking ready for review.","created":"2020-06-15T13:03:33.198+0000"},{"body":"I noticed that the previous version of the patches wasn't working in all cases due to an existing quirk of the CAS implementation.\r\n\r\nNamely, accepted updates that were empty were not replayed by {{beginAndRepairPaxos}}. Which is a problem for the new empty commits made during serial reads/non-applying CAS. I added tests to show that if the commit messages for those empty commits were lost/delayed, we could still have linearizability violations.\r\n\r\nNow, the logic of not replaying empty updates looks wrong to me. There shouldn't be anything special about an empty update, and if one is explicitely accepted by a quorum of nodes, we shouldn't ignore it, or that's a break of the Paxos algorithm (as kind of can be demonstrated by the tests I added).\r\n\r\nTo be clear, that logic was added *by me* in CASSANDRA-6012 and that was the sole purpose of that ticket. Except that I can't make sense of my reasoning back then, and since I didn't included a test to demonstrate the problem I was solving back then (which was wrong, mea culpa), I have to assume that I was just confused (maybe I mixed in my head promised ballots and accepted ones?). Anyway, I think the fix here is simply to remove that bad logic, which fixes the issue, and I included an additional commit for that.\r\n","created":"2020-06-17T11:55:46.100+0000"},{"body":"So, I'm reasonably sure it cannot be necessary for us to commit an empty proposal, because we do not ever need to witness it. Either the proposal was agreed by a quorum (and the proposer can report this) but it has no visible effect on future proposals, and does not need to be witnessed by anybody else, or it was not agreed and it does not need to be either proposed again, committed or witnessed by anybody else.\r\n\r\nHowever we have to be consistent about it: we either need to _never_ commit them, or _always_ commit them.","created":"2020-06-17T13:39:33.271+0000"},{"body":"bq. I'm reasonably sure it cannot be necessary for us to commit an empty proposal, because we do not ever need to witness it.\r\n\r\nWe may have to be precise. We do not need to \"apply\" an empty commit, since it's a no-op, and the patch actually ensures we don't bother. But \"committed\" do something else, it update the \"mrc\" value, and _that_ needs to be done. Otherwise, if we _accept_ an empty proposal, yet does not update the \"mrc\" value, we will not do progress anymore (well, without additional modification to the algorithm that is).\r\n\r\nBut I could be misunderstanding what you are suggesting here. I'll note though, just in case that help, that the logic I'm calling faulty is not the _commit_ of empty updates (though, as said above, I think it's necessary for the sake of the mrc value), it's the fact the don't replay the _proposal_ of empty updates. ","created":"2020-06-17T13:53:59.403+0000"},{"body":"The problem here stems only from the overload of {{mostRecentInProgressCommitWithUpdate}}, which (seems to) assume that an empty update is for a higher promise (since the meaning is overloaded in the response message) rather than an \"incomplete\" proposal. If the empty proposal were to be correctly merged with {{mostRecentInProgressCommitWithUpdate}}, it would override the early non-empty incomplete proposal.\r\n\r\nWhich is a long-winded way of saying that I am fairly confident there's no need to update the paxos state table with the \"committed\" status of this empty proposal so long as it remains in the table _as an accepted proposal_, and so long as this accepted proposal continues to override earlier in progress proposals.\r\n","created":"2020-06-17T14:10:03.440+0000"},{"body":"I'll have to apologize, but I don't understand what you are suggesting.","created":"2020-06-17T14:16:49.182+0000"},{"body":"I'm not proposing we do anything different for your patch, just clarifying that this isn't strictly necessary - it is quite possible to modify the algorithm to never commit empty proposals. The problem today is that we:\r\n\r\n# \"Refresh\" a quorum with the MRC if not witnessed by all promisers \r\n# Filter out empty proposals when deciding if we have an in progress proposal ({{mostRecentInProgressCommit}} vs {{mostRecentInProgressCommitWithUpdate}})\r\n\r\nIf instead we did not refresh empty commits, and we did not filter out empty proposals when _updating_ {{mostRecentInProgressCommitWithUpdate}} but did not _complete_ any empty proposals we found then everything would be fine.\r\n\r\n{{mostRecentInProgressCommitWithUpdate}} confuses matters because it is poorly named, and is updated by its naming rather than intent - I think it is _meant_ to be {{mostRecentInProgressProposal}} whereas {{mostRecentInProgressCommit}} should be e.g. {{mostRecentInProgressPromiseOrProposal}}, and {{mostRecentInProgressProposal}} would gain empty proposals as well as non-empty ones, and correctly discount the older in progress proposal that was invalidated by the newer read that did not witness it.\r\n\r\nTo be clear, I'm mostly participating in this discussion for my own benefit and for the benefit of future work, not trying to solicit changes to your work.","created":"2020-06-17T14:40:33.144+0000"},{"body":"To say it another way: the only purpose of an empty proposal is to poison earlier proposals, and this can be done just as well without moving the proposal to the MRC column in the table. If we treat it as any other \"in progress\" proposal for invalidating earlier proposals, then once we reach a quorum we must in future be witnessed alongside any earlier proposals and invalidate them. If we didn't reach a quorum, then it doesn't matter if we are witnessed or not, or if any earlier proposals are invalidated.\r\n\r\n","created":"2020-06-17T15:02:54.859+0000"},{"body":"Ok, I understand what you are suggesting now and I agree this should work as well. And it does is more optimal.\r\n\r\nI like to think of our algorithm as \"pure Paxos instances\" separated by the MRC to tell us when we can forget the previous instance and start a new one. Committing empty updates as any other updates still fits that mental model, while your suggestion adds a bit of a special case in that it bends the Paxos rules slightly, allowing to sometime ignore a previously accepted value in a promise (when it's empty). Which is not a criticism, just thinking out loud. It's more performant and this is likely worth the slight special casing since it's not too hard to reason about its correctness.\r\n\r\nI'll sleep on it and modify to your suggestion tomorrow (which is trivial, just need to massage an appropriate comment to explain it).\r\n","created":"2020-06-17T15:38:21.060+0000"},{"body":"Alright, my \"tomorrow\" is off by 1, but pushed an additional commit to implement the optimization suggested by Benedict above. Restarted CI for good measure.\r\n\r\n||branch||CI||\r\n|[3.0|https://github.com/pcmanus/cassandra/tree/C-12126-3.0]|[Run #155|https://ci-cassandra.apache.org/job/Cassandra-devbranch/155/]|\r\n|[3.11|https://github.com/pcmanus/cassandra/tree/C-12126-3.11]|[Run #156|https://ci-cassandra.apache.org/job/Cassandra-devbranch/156/]|\r\n|[4.0|https://github.com/pcmanus/cassandra/tree/C-12126-4.0]|[Run #157|https://ci-cassandra.apache.org/job/Cassandra-devbranch/157/]|\r\n","created":"2020-06-19T14:02:27.361+0000"},{"body":"Sorry, for the delay. The patches looks good.","created":"2020-10-26T14:02:55.132+0000"},{"body":"Thanks for the review. I've rebased the branches, but since the last runs were a while ago, I restarted CI runs. I'll commit if those look clean.\r\n\r\n||branch||CI||\r\n|[3.0|https://github.com/pcmanus/cassandra/tree/C-12126-3.0]|[Run #171|https://ci-cassandra.apache.org/job/Cassandra-devbranch/171/]|\r\n|[3.11|https://github.com/pcmanus/cassandra/tree/C-12126-3.11]|[Run #172|https://ci-cassandra.apache.org/job/Cassandra-devbranch/172/]|\r\n|[4.0|https://github.com/pcmanus/cassandra/tree/C-12126-4.0]|[Run #173|https://ci-cassandra.apache.org/job/Cassandra-devbranch/173/]|\r\n","created":"2020-11-05T10:58:56.541+0000"},{"body":"So, before we commit this I wanted to share that some experimentation found that this can lead to a significant increase in timeouts, particularly for read-heavy workloads, that previously would not have competed with each other. I think committing this to a patch release is honestly problematic, as it could surprise users with a service outage. At the very least, there should be HUGE warnings in {{NEWS.txt}}, but honestly I would prefer to have users opt-in for patch releases.\r\n\r\nAs much as I agree that it is problematic to provide the wrong semantics, I think it is also problematic to force a decision between stability and correctness onto our users without their informed and positive consent.\r\n\r\nI hope that I will be able to provide the community with an alternative solution in the near future, without these (and many other existing) pitfalls. However I'm not sure how that should affect this decision.","created":"2020-11-05T11:55:47.709+0000"},{"body":"{quote}I hope that I will be able to provide the community with an alternative solution in the near future, without these (and many other existing) pitfalls.{quote}\r\n\r\n[~benedict] Few questions regarding your comment:\r\n* What timeframe do you have in mind? \r\n* Is it a solution only for 4.0 or for all the branches?\r\n* Can we help you with that?","created":"2020-11-05T13:17:40.310+0000"},{"body":"To some extent that is all up for debate.\r\n\r\n\r\n My plan so far has been to avoid interfering with 4.0 release, so I have been working towards targeting 4.x. This would also permit time to produce documentation and reach out to the list to begin the slow handshake to see if the project wants the work, and in what manner. However, the main body of work is essentially complete, so it is possible that this could be brought forwards if there were appetite.\r\n As to target version, it would be possible to target 3.0+, at least for a portion of the work that would encompass this issue, without a great deal of work. The project's appetite would be the main decider, as it's a significant body of work.\r\n\r\n\r\n The main contribution would be a parallel implementation of the same underlying Paxos algorithm, that is able to run concurrently alongside it (supporting live migration), but with several latency improvements, as well as several fixes to correctness. Alongside this is related work to guarantee linearizability across range movements in the form of modifications to repair, bootstrap, replace etc.\r\n\r\n\r\n Related to this work are several patches to wider Cassandra to support automated verification of its correctness, by permitting deterministic simulation of Cassandra clusters with adversarial ordering of events. We have so far simulated billions of transactions to verify its linearizability. I anticipate that this work will be useful for the project's overall goal of improving quality, but they are themselves quite significant and will require their own discussions around timeline and scope.","created":"2020-11-06T11:39:42.469+0000"},{"body":"It seem to me that there are several options here:\r\n# Try to use your proposal for 4.0 if the community has the appetite for it. The main issue there is some potential extra delay for 4.0\r\n# Do nothing for 4.0. Meaning do not commit the patch. We have lived a long time with that issue and we can probably wait a bit more for a proper solution.\r\n# Commit the patch as such, fixing the correctness but introducting potentially some performance issue until we release a better solution.\r\n# Changing the patch to default to the current behavior but allowing people to enable the new one if the correctness is a problem for them.\r\n\r\nMay be we should trigger a discussion on the mailing list and see what is other people opinion.\r\n\r\nI can take care of that next week if you think it is a good idea.","created":"2020-11-06T12:42:48.743+0000"},{"body":"Yes, that sounds like a great idea, and I really appreciate you offering to take that to the list. I'll chime in with any necessary details to help inform the decision, but will try not to influence it otherwise. I don't have a strong opinion about which of those four options we select, except that my experiments do suggest (3) is perhaps dangerous for some of our users. It's probably a trade-off that should be made with careful business consideration and experimentation by each end user.\r\n\r\nAs far as delaying 4.0 is concerned, that's probably also a matter of community decision-making. We could quite quickly have a patch, that has been reviewed by multiple committers, posted in fairly short order - perhaps before we exit beta. This work will have had much greater validation than the current implementation, but publishing all of this validation work will take longer - likely also achievable before GA, but we might have to invert our process a little. Perhaps this is acceptable, given the balance of correctness and regression we're considering as an alternative, but given my proximity to the work (and that I also don't have a strong position either way), I would prefer to let others make that call.","created":"2020-11-06T15:24:13.480+0000"},{"body":"Committed following the dev mailing list discussion. Thanks.","created":"2020-11-27T16:18:29.271+0000"},{"body":"> 3) Issue another CAS Read and it goes to A and B. Now we will discover that there is something inflight from A and will propose and commit it with the current ballot. Now we can read the value written in step 1 as part of this CAS read.\r\n\r\nSorry, I'm not fully sure about the current implementation and how realistic my proposal is but,\r\n\r\ncan we read all the replicas to do the read recovery in step 3 to solve the issue?\r\n\r\nIt only reads A and B but if we read C as well, we know that the proposal is not accepted by the majority.\r\n\r\n ","created":"2020-12-01T02:50:18.995+0000"},{"body":"[~feeblefakie] C can be unreachable for different reasons.","created":"2020-12-01T08:40:50.868+0000"},{"body":"Just a note that the bug that this fixes usually pops up as the following timeout for people looking for reasons why SERIAL or LOCAL_SERIAL are seeing read timeouts >3.11.10.  Setting the flag to the opt-out option will `fix` it but probably shouldn't be reading at this level if you run into this.\r\n{code:java}\r\n! com.datastax.driver.core.exceptions.ReadTimeoutException: Cassandra timeout during read query at consistency LOCAL_SERIAL (2 responses were required but only 0 replica responded)\r\n{code}","created":"2021-04-06T23:51:46.158+0000"}],"conversations":[{"body":"While looking at the CAS code in Cassandra, I found a potential issue with CAS Reads. Here is how it can happen with RF=3\n\n1) You issue a CAS Write and it fails in the propose phase. A machine replies true to a propose and saves the commit in accepted filed. The other two machines B and C does not get to the accept phase. \n\nCurrent state is that machine A has this commit in paxos table as accepted but not committed and B and C does not. \n\n2) Issue a CAS Read and it goes to only B and C. You wont be able to read the value written in step 1. This step is as if nothing is inflight. \n\n3) Issue another CAS Read and it goes to A and B. Now we will discover that there is something inflight from A and will propose and commit it with the current ballot. Now we can read the value written in step 1 as part of this CAS read.\n\nIf we skip step 3 and instead run step 4, we will never learn about value written in step 1. \n\n4. Issue a CAS Write and it involves only B and C. This will succeed and commit a different value than step 1. Step 1 value will never be seen again and was never seen before. \n\n\n\nIf you read the Lamport “paxos made simple” paper and read section 2.3. It talks about this issue which is how learners can find out if majority of the acceptors have accepted the proposal. \n\nIn step 3, it is correct that we propose the value again since we dont know if it was accepted by majority of acceptors. When we ask majority of acceptors, and more than one acceptors but not majority has something in flight, we have no way of knowing if it is accepted by majority of acceptors. So this behavior is correct. \n\nHowever we need to fix step 2, since it caused reads to not be linearizable with respect to writes and other reads. In this case, we know that majority of acceptors have no inflight commit which means we have majority that nothing was accepted by majority. I think we should run a propose step here with empty commit and that will cause write written in step 1 to not be visible ever after. \n\nWith this fix, we will either see data written in step 1 on next serial read or will never see it which is what we want. \n","from":"reporter","subject":"CAS Reads Inconsistencies "},{"body":"cc [~slebresne] and [~jbellis]","from":"developer"},{"body":"I \"think\" you are right that this can happen, and that committing an empty commit on SERIAL reads \"should\" fix it. Paxos is however subtle enough that I would feel more confident with this if we had a reproduction test first, if only so we can validate whatever patch we come up with. [~jkni] I believe you've spend some time on jespen-like tests for paxos, do you think this is something we could use to try to reproduce something like that relatively consistently?","from":"developer"},{"body":"Sure - I have Jepsen tests for LWT and a few other similar tests I've been building. I haven't seen an issue like this reproduced through them so far, but I'll try and see if I can reproduce this reliably. It certainly seems feasible to hit. I definitely think it is a good idea to run any proposed patch through LWT tests for a while.","from":"developer"},{"body":"I was able to repro it in a not so good way :) Here are the steps\n1. Create a 3 node C* cluster(A,B and C)\n2. Create a keyspace with RF=3 and a simple table\n3. Insert System.exit PaxosState.propose method to simulate a failure in A and B. \n4. Do a CAS Write. Now with this, only C will be able to store the propose and not A and B. \n4. Now bring down machine C and remove the System.exit from A and B and bring them up again. \n5. Do a CAS Read and this will involve only A and B and will not return the data. \n6. Bring down A and bring up C. \n7. Do a CAS Read and you will be able to read the data. \n\nIf we can use some test framework to simulate such failures then that will be better. ","from":"developer"},{"body":"Another take on how to test coordination aspects for this ticket would be to make use of the MessagingService mocking classes implemented in CASSANDRA-12016. I've created a couple of tests [here|https://github.com/spodkowinski/cassandra/tree/WIP-12126/test/unit/org/apache/cassandra/service/paxos] to get a better idea how this would look like. Although limited to observing the behavior of a single node/state machine, it's probably more lightweight and easier to implement than doing the same using dtests or Jepsen.\n\nAs of the described edge case, I'd agree with [~kohlisankalp]'s suggestion (if I understood correctly) to do an additional proposal round. However, it would be nice to optimize this a bit so we don't trigger new proposals for each and every SERIAL read. I did a first implementation for this [here|https://github.com/spodkowinski/cassandra/commit/96ec151992f49c773e5af5d85ce69ec87d8b7bc5] (with [CASReadTriggerEmptyProposal|https://github.com/spodkowinski/cassandra/blob/WIP-12126/test/unit/org/apache/cassandra/service/paxos/CASReadTriggerEmptyProposal.java] as corresponding test) for the sake of discussion. ","from":"developer"},{"body":"I've now revisited this issue again and took a closer look at the work I've done back then months ago (after rebasing on trunk). The patch follows the suggested solution by [~kohlisankalp] by using an empty commit as additional propose step. It also implements an optimization to avoid this step in case no paxos rounds for writing new values have been conducted between serial reads.\n\n* [CASSANDRA-12126-trunk|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12126-trunk]\n* [CircleCI results|https://circleci.com/gh/spodkowinski/cassandra/5]\n\nExcept the actual fix, there's also a lot of tests included, which I'd hate to throw away, as we're clearly lacking tests for CAS operations.\n","from":"developer"},{"body":"I see you outlining two series of steps:\n\n1 -> 2 -> 3. The value V from 1 is not seen in 2, but once it is seen in 3 it is always seen.\n\n1 -> 2 -> 4. V is never seen.\n\nIt seems to me that both of these maintain linearizability. What am I missing?","from":"developer"},{"body":"bq. 1 -> 2 -> 3. The value V from 1 is not seen in 2, but once it is seen in 3 it is always seen.\n\nAs 1 \"completes\" before 2, it's result should be visible by 2 (or not ever) for linearizability (taken in the sense discussed here for instance: http://www.bailis.org/blog/linearizability-versus-serializability/). More pragmatically, outside of any theoretical definition, if serial read can't guarantee they will see any previous operation (even failed one, as long as they returned to the client), then they are not very useful in the first place.","from":"developer"},{"body":"But we stipulated that 1 times out and did not complete. (If it did complete it would be guaranteed to be visible by any majority of course.)","from":"developer"},{"body":"Bailis:\n\n# once a write completes, all later reads (where “later” is defined by wall-clock start time) should return the value of that write or the value of a later write. \n# Once a read returns a particular value, all later reads should return that value or the value of a later write.\n\nI think we all agree that our current behavior satisfies (2). I am saying that we actually also satisfy (1) because the write is not complete until Sankalp's step 3.\n","from":"developer"},{"body":"bq. But we stipulated that 1 times out and did not complete.\n\nIt completed in the sense that the operation is finished and returned to the client (albeit with a timeout). Don't get me wrong, theory is blind, so if you want to define than an operation completes only if it \"finished and did not timeout\" and define linearizability only in term of completing operation (with that definition of completion), then sure, I agree this particular definition of linearizability is not broken by the description of this ticket.\n\nBut how useful is a definition of linearizability that says nothing about operations that timeouts (especially keeping in mind that our implementation is particularly prone to timeouts)?","from":"developer"},{"body":"bq. I am saying that we actually also satisfy (1) because the write is not complete until Sankalp's step 3.\n\nYour definition of \"completion\" doesn't work logically. You're suggesting (1) only complete when it is visible (in (3)), but linearizabilty is about when operations that complete are visible, so if you define completion as \"when it's visible\", you're in for trouble.","from":"developer"},{"body":"I'm confused, because it sounds like you're saying \"all operations should be visible once finished.\" Of course that's not actually what you mean that would require participation from all replicas to finish in-flight requests and not just a majority. What is the distinction you are proposing?","from":"developer"},{"body":"bq. What is the distinction you are proposing?\n\nNot sure, I think we don't put the same definitions on operation visibility. What I'm saying is that \"if an operation has a visible outcome, then that outcome should be visible (by serial operations) by any subsequent operation (so as soon as the operation returns to the client if you will)\". In particular, if a serial read follows a serial write (meaning that it's started after the write returned, even with a timeout), then if the write has any effect, the read should see it.\n\nNote that when you get a timeout on the initial write, you don't know if the write has been applied or not, but the whole point of a serial read is to be able to unequivocally decide what was that outcome. If we can't guarantee that, if there is no way to observe if a timed-out write has been applied or not, then I'm not sure how one would use LWT in the first place.\n","from":"developer"},{"body":"I see. So you are saying that\n\n1: Write\n2: Read -> Nothing\n3: Read -> Something\n\nIs broken because to go from Nothing to Something [in a linearized system] there needs to be a write in between.","from":"developer"},{"body":"Exactly.","from":"developer"},{"body":"Hi all,\r\n\r\nOur team from UCARE University of Chicago, have been able to reproduce similar manifestation to this bug consistently with our model checker. (Our scenario is different with what [~kohlisankalp] proposed)\r\nHere are the workload and scenario of the bug:\r\n\r\nWorkload: 3 nodes-cluster, 3 client requests (but no crash event)\r\n\r\nScenario:\r\n # Start 3-nodes cluster and inject all of 3 client requests to 3 different nodes (node X, Y, Z)\r\n # Node X sends its prepare messages (ballot number=1) to all nodes and all nodes accept it\r\n # Node X sends its propose message to itself, causing its inProgress value to be \"X\".\r\n # Node Y sends its prepare messages (ballot number=2) to all nodes.\r\nThis also causes the rest of node X propose messages to be invalid because its ballot number is smaller than node Y prepare messages.\r\n # In our scenario, the prepare response messages from node Y and Z comes first before prepare response message from node X, causing the node Y to unrecognize the state of node X which already accepted value \"X\" (step 3).\r\n # But since our query of client request 2 has an IF, that said IF value_1='X', therefore node Y will not continue on sending propose messages to all nodes.\r\nUp to this point, it means none of the queries are committed to the server.\r\n # Node Z now sends its prepare messages to all nodes and all nodes accept it.\r\n # In our scenario, now the node X returns its response first where it also let node Z knows about its inProgress Value \"X\".\r\nFrom here, node Z will propose and commit client request-1 (with value \"X\") instead of client-request-3.\r\n # Therefore, we ended up having client request 1 stored to the server, although client request-3 was the one that is said successful.\r\n\r\nWe are ready to assist, if any further information is needed.\r\n\r\n \r\n\r\n ","from":"developer"},{"body":"[~jeffreyflukman] thanks for this report. Suspect that most of the folks who are interested in this are already cc'd and received an email notification of your response, but explicitly tagging [~benedict] [~iamaleksey] and [~bdeggleston] as people who aren't yet watching it but may be interested.\r\n\r\nAlso, very much interested in the model you mentioned - is that available publicly at this point? ","from":"developer"},{"body":"To complete our scenario, here is the setup for our Cassandra:\r\n We run the scenario with Cassandra-v2.0.15.\r\n Here is the scheme that we use:\r\n * CREATE KEYSPACE test WITH REPLICATION = \\{'class': 'SimpleStrategy', 'replication_factor': 3};\r\n * CREATE TABLE tests ( name text PRIMARY KEY, owner text, value_1 text, value_2 text, value_3 text);\r\n\r\nHere are the queries that we submit:\r\n * client request to node X (1st): UPDATE test.tests SET value_1 = 'A' WHERE name = 'testing' IF owner = 'user_1';\r\n * client request to node Y (2nd): UPDATE test.tests SET value_2 = 'B' WHERE name = 'testing' IF value_1 = 'A';\r\n * client request to node Z (3rd): UPDATE test.tests SET value_3 = 'C' WHERE name = 'testing' IF value_1 = 'A';\r\n\r\nTo confirm, when the bug is manifested, the end result will be: value_1='A', value_2=null, value_3=null\r\n\r\n[~jjirsa], regarding our tool, at this point, it is not open for public. ","from":"developer"},{"body":"Why is end result not correct? second and third operation did not succeed because first 1 did not finish? Can you combine the example with the earlier comment please","from":"developer"},{"body":"[~jeffreyflukman] it would help if you could explicitly state the client responses returned for each of your operations. The options are: time out, rejected (condition not met), success (condition met, and mutation applied)\r\n\r\nFor completeness, as with CASSANDRA-12438, the read queries you are performing, to which nodes, at what point and with what consistency levels would be helpful to know. Are you verifying the state with a SERIAL read after the last query, most specifically?  Also, can we assume that the state of the table began with \\\\{name:'testing', owner:'user_1', value1:null, value2:null, value3:null}\\?","from":"developer"},{"body":"During our testing with our model checker, we limit the round of Paxos for each query,\r\nbecause if not, it is possible that we get stuck in a very long sequence of message transactions among the nodes without progressing anywhere. So, what we do is we only execute one round of Paxos for each query.\r\n\r\nTo enlight our test and combine our whole story, here is what happened in detail:\r\n * We first prepared the 3 node-cluster with the test.tests table as initial table structure and yes, the initial table began with:\r\n{name:'testing', owner:'user_1', value1:null, value2:null, value3:null}\r\n * Next, we run the model checker that will start the 3 node-cluster.\r\n * Inject the 3 client requests in order: query 1, then query 2, then query 3.\r\nThis cause query 1 to have ballot number < query 2 ballot number < query 3 ballot number.\r\n * Now this means, in the beginning, the model checker already see there will be 9 prepare messages in its queue that will be reordered in some way.\r\n * When the bug is manifested, we ended up having:\r\n ** Node X's prepare messages proceed and all nodes response with true back to node X.\r\n ** Node X sends its propose message with value_1='A' to itself first and get a response true as well.\r\n ** At this moment, Node X inProgress value is updated to the proposed value, value_1='A'\r\n ** But then node Y prepare messages proceed and all nodes response with true back to node Y,\r\nbecause prepare messages of node Y have a higher ballot number.\r\n ** But when node Y about to proceed the propose messages it realized that the current data does not fulfill the IF condition, so it does not proceed to propose messages. --> Client request 2 to node Y is therefore rejected\r\n ** Continuing node X propose messages to node Y and Z, both requests are returned with false to node X\r\n ** Now at this point node X should be able to retry the Paxos with a higher ballot number, but since we limit the round of Paxos for each query to one, therefore client request 1 to node X is timed out.\r\n ** Lastly, node Z sends its prepare messages to all nodes, and get response true messages from all nodes,\r\nbecause the ballot number is higher as well.\r\n ** At this point, if the node X response message is returned first to node X, what will happen is node Z will realize that node X still has an inProgress value in the process (value_1='A'). This cause node Z to send propose messages and commit messages but for client request 1 using the current highest ballot number.\r\nHere we have our first data update saved: value_1='A', value_2=null, value_3=null.\r\n ** Back to our constraint of one round Paxos for each query, we ended up not retrying client request-3 because we reached timeout.\r\n * To sum up:\r\n ** client request-1: Timed out\r\n ** client request-2: Rejected\r\n ** client request-3: Timed out\r\n\r\nThere we get an inconsistency between the client side and the server side, where all requests actually failed, but when we read the end result again from all nodes, we get value_1='A', value_2=null, value_3=null.\r\n\r\n \r\n\r\nI made a wrong statement at the end of my first comment:\r\n{quote}9. Therefore, we ended up having client request 1 stored to the server, although client request-3 was the one that is said successful.\r\n{quote}\r\nIt should be failed due to timeout.","from":"developer"},{"body":"bq. client request-1: Timed out, client request-2: Rejected, client request-3: Timed out\r\n\r\nGiven those responses to the queries. The client side does not know the state of the system without issuing a READ at SERIAL (or doing another INSERT that gets a success which the state can be inferred from).\r\n\r\nbq. There we get an inconsistency between the client side and the server side, where all requests actually failed, but when we read the end result again from all nodes, we get value_1='A', value_2=null, value_3=null.\r\n\r\nGiven the responses you got, there is no inconsistency. The client received \"timed out\" exceptions. A timed out exception means \"your query may or may not have been applied, the server doesn't know, you should retry it if you want to ensure it goes through\". In this case request-1 was successful, and request-3 failed. So {{value_1='A', value_2=null, value_3=null}} is a valid state and not inconsistent.","from":"developer"},{"body":"I agree with [~jjordan] that this is a correct response. \r\n\r\nAlso in the future please open a new Jira if it is a different issue. ","from":"developer"},{"body":"Thank you for your responses, [~jjordan] and [~kohlisankalp].\r\nI think you have cleared up some misunderstandings for me (and our team) where timeout is a \"gray area\" for the client \r\nto determine whether a request has been successfully processed.\r\n\r\nOne thing that we would like to point out maybe, based on the early discussion in this bug description, quote\r\n{quote}However we need to fix step 2, since it caused reads to not be linearizable with respect to writes and other reads. In this case, we know that majority of acceptors have no inflight commit which means we have majority that nothing was accepted by majority. I think we should run a propose step here with empty commit and that will cause write written in step 1 to not be visible ever after.\r\n{quote}\r\nWhat we tried to mimic with our model checker in the beginning actually was this scenario where node Y saw that the majority of nodes do not have inProgress value, but then suddenly node Z saw that there is an inProgress value from node X and tried to repair and commit it.\r\nSo, we confirm that we can also see this behavior:\r\n{quote}2: Read -> Nothing\r\n3: Read -> Something\r\n{quote}\r\nWe read nothing in node Y, yet node Z read something in the next request.\r\n\r\n\r\n\r\nTo sum up, at least, our scenario explains this behavior: Node Y does not try to repair the Paxos because node X's prepare response comes last, therefore node Y ignores the node X's prepare response and based its decision to not repair the Paxos.\r\nBut in node Z's client request, node Z decides to repair the Paxos based on node X's existing inProgress value_1=\"A\" because node X's prepare response comes early (1st or 2nd). Which cause an inconsistent reaction in some way between node Y and node Z (although this is correct based on the original Paxos algorithm).\r\n\r\n\r\nA solution to avoid this inconsistent reactions from these two nodes maybe is for each node to decide whether to repair a Paxos or not based on the complete view of the alive nodes, therefore if the response X's comes last with an inProgress value, node Y will still repair the Paxos.","from":"developer"},{"body":"{quote}We read nothing in node Y, yet node Z read something in the next request.\r\n{quote}\r\nI think the problem here is that, at the API level, there isn't enough information to say that X didn't simply 'occur' *after* both Y and Z. That is, unless the rejection of Y occurs after X's timeout. In this case, it would seem to be an API-visible error, as at the point of timeout the indeterminacy should be fixed. Timeouts should not ‘live forever’ as the bogeyman, ready to mess with history.\r\n\r\nI think, though, that the suggested mechanism could result in this.\r\n\r\nTake three nodes (RF=3) A, B and C; and any three CAS operations X, Y and Z such that:\r\n * X and Y can always succeed\r\n * Z can only succeed if X has succeeded\r\n\r\nSetup:\r\n # Prepare _and_ Propose X with ballot 1; proposal accepted only by A\r\n ** this will be the last and only node’s proposal acceptance\r\n # Prepare Y with ballot 2; reach B and C before ballot 1, so they do not accept\r\n # Now, lock X and Y in battle, always failing to proceed to the propose step before the other reaches the prepare step again\r\n # X and Y both timeout having failed to cleanly apply\r\n\r\nPart 2:\r\n # Z is now attempted; it prepares to only B and C, seeing no in-progress proposal\r\n # As a result, it does not see X; it is rejected, so there is no new proposal/commit\r\n # Z is attempted again; this time, A is consulted\r\n # Suddenly, a wild X appears. From nowhere.  Z succeeds, despite no intervening operations.\r\n\r\nIt does seem, in essence, to be an incidence of the bug (or a very similar one) described in the ticket.","from":"developer"},{"body":"Having reviewed the code I think what Benedict says is correct. The criteria we use for identifying if there is an progress paxos round that needs resolution is incorrect because it assumes we have visibility to all accepted ballots when we only have visibility to a majority.\r\n\r\nI think this optimization can be done correctly, but it's a bit of surgery. Right now reads do a prepare and modify the promised ballot at each acceptor. If instead we only read the promised ballot from each acceptor then we could check the promised ballot matches the most recent committed ballot. If those are the same we know nothing is in progress because a higher ballot than the most recent accepted/committed ballot has not been promised by a majority which means there can be no lingering accepted ballot since a majority of promises must be collected first.\r\n\r\nIf that isn't the case then we can go ahead and do a prepare + propose to make them match and subsequent reads won't have to do a propose.\r\n\r\nThis may also impact our choice of how many replicas to contact in each phase since we want them to have consistent paxos state so reads can be one roundtrip. I am not sure if we contact them all (like with mutations) or just a majority.","from":"developer"},{"body":"It looks like I never noted that IMO the real problem is that we do not serialize the evaluation of a condition if the condition is not met, so other commands are also not serialized wrt the evaluation of such conditions either.  The evaluation of a condition should always be a serialized wrt other events, so we should be agreeing on the conditional operation, and then performing it, not evaluating the condition and then serializing the choice of a new value.","from":"developer"},{"body":"It definitively doesn't look good that this messages comes so late, but I feel this is a serious issue of the {{SERIAL}}/{{LOCAL_SERIAL}} consistency levels since this breaks the basic guarantee they exist to provide, and as such should be fixed all the way down 3.0, and the sooner, the better.\r\n\r\nIn an attempt to sum this up quickly, the problem we have here affects both serial reads _and_ LWT updates that do not apply (whose condition evaluates to {{false}}). In both case, while the current code replays \"effectively committed\" proposals (those whose proposal has been accepted by majority of replica) with {{beginAndRepairPaxos}}, neither make proposals of their own, so nothing will prevent a proposal accepted by a minority of replica (say just one) to be later replayed (and thus committed).\r\n\r\nI've pushed [2 in-jvm dtests|https://github.com/pcmanus/cassandra/commit/3442277905362b38e0d6a2b8170916fcfd18d469] that demonstrate the issue for both cases (again, serial reads and non-applying updates). They use \"filters\" to selectively drop messages to make failure consistent but aren't otherwise very involved.\r\n\r\nAs [~kohlisankalp] mentioned initially, the \"simplest\"\\[1\\] way to fix this that I see is to commit an empty update in both cases. Actually committing, which sets the {{mostRecentCommit}} value in the Paxos state, ensures that no prior proposal can ever be replayed. I've pushed a patch to do so on 3.0/3.11 below (will merge up on 4.0, but wanted to make sure we're ok on the approach first):\r\n\r\n||version||\r\n| [3.0|https://github.com/pcmanus/cassandra/commits/C-12126-3.0] |\r\n| [3.11|https://github.com/pcmanus/cassandra/commits/C-12126-3.11] |\r\n\r\nThe big downside of this patch however is the performance impact. Currently, a {{SERIAL}} read (that finds nothing in progress it needs to replay) is 2 round-trips (a prepare phase, followed by the actual read). With this patch, it is 3 round-trips as we have to propose our empty commit and get acceptance (we don't really have to wait for responses on the commit though), which will be noticeable for performance sensitive use-cases. Similarly, the performance of LWT that don't apply will be impacted.\r\n\r\nThat said, I don't seen another approach to fixing this that would be as acceptable for 3.0/3.11 in terms of risks, and imo 'slower but correct' beats 'faster but broken' any day, so I'm in favor of moving forward with this fix.\r\n\r\nOpinions?\r\n\r\n\r\n\r\n\\[1\\]: I mean by that both the simplicity of the change, but also of validating that this fix the problem at hand without creating new correctness problems.\r\n","from":"developer"},{"body":"Agreed that's the most straightforward way to address both issues (although I've only skimmed your patch).\r\n\r\nIn 3.x though, and at least for the serial read fix, I think we should include a flag to disable the fix, in case a) there's a problem with the fix or b) operators would rather trade the performance impact for linearizability for whatever reason.\r\n\r\nThere's also a variant of the non-applying update issue where it's exposed by a read, not another insert. It would be good to have a test for that as well.","from":"developer"},{"body":"So, a thought has occurred to me: what do we actually claim our consistency properties are for SERIAL? \r\n\r\nMy understanding was that we claimed only serializability, in which case I don't think that strictly speaking this is a bug. I think it's only a bug if we claim strict serializability. However the only docs I can find claiming either are DataStax's which mixes linearizable up with serializable.\r\n\r\nFWIW, I consider this to be a bug, as we should at least support the more intuitively correct semantics. But perhaps we should instead introduce a new STRICT_SERIAL consistency level to solve it, and clarify what SERIAL means in our docs?\r\n\r\nI would also be OK with simply claiming strict serializability for SERIAL. But perhaps this technicality/ambiguity buys us some time and cover to solve the problem without introducing major performance penalties?\r\n\r\nI also have some relevant test cases I will share tomorrow, along with test cases for other correctness failures of LWTs.","from":"developer"},{"body":"I've pushed various test cases [here|https://github.com/belliottsmith/cassandra/tree/12126-tests-3.0] - most of them marked {{@Ignore}} because they are known to fail, and won't be resolved immediately.","from":"developer"},{"body":"{quote}I think we should include a flag to disable the fix\r\n{quote}\r\nThe option of having a flag occurred to me, but I rejected it initially because I continue to believe the current behavior is wrong (a moral judgment, I guess) and in principle, having a \"please, make my database broken\" flag does not feel like a good idea.\r\n\r\nBut I reckon that it _may_ exists advanced users that did noticed the lack of linearizability for reads and effectively built around it knowingly, for which the performance impact may be considered a regression with no upside (but if you sense skepticism on my part when reading that sentence, you're radar is not completely off).\r\n\r\nAnd as we're talking minor upgrade here, I'm amenable to such flag, though I'd prefer making it clear somehow that it is unsafe/risky and something we may remove in the future with no particular warning.\r\n{quote}It would be good to have a test for that as well.\r\n{quote}\r\nCertainly, good point, I can add the 2 missing interleaving.\r\n{quote}do we actually claim our consistency properties are for SERIAL?\r\n{quote}\r\nWhile our official doc on the matter is certainly lacking (not spelling much guarantee at all afaict, and I'm happy to piggy-back on this ticket to correct that), we've always implied linearizability. I have, at least, and I'm sure I can dig up other doing it as well on the mailing list if necessary. We did this both by throwing the linearizable word out from time to time, but also by repeatedly recommending that when a write times out, one needs to issue a SERIAL read to 'observe' if that write went through or not (and as an aside, if you can't rely on either reads or non-applying CAS for that, I'm not even sure how to use LWTs, except maybe for excessively specific cases).\r\n{quote}perhaps we should instead introduce a new STRICT_SERIAL consistency level\r\n{quote}\r\nI'm rather cold on that because, tbh. I think non-strict serializability is a theoretical notion that is useless in practice and that it is something we should not offer. And I'd rather avoid one more \"feature\" for which we spend our time saying \"don't use it\".\r\n{quote}I've pushed various test cases\r\n{quote}\r\nAwesome, thanks. I'll look at integrating those in the branch if you don't mind.","from":"developer"},{"body":"bq. I'm amenable to such flag\r\n\r\nActually, let me rephrase that a bit. I'd *really* prefer not adding such flag. If someone is ok with serializability without linearizability, then they can use QUORUM reads, and given how things are implemented, it provides (non-strict) serializability. Granted, for someone that uses SERIAL today, is ok with the lack of linearizability and can't afford the performance penalty, it'll require a client side change, which this flag would avoid, so there is not zero value to such flag. But I suspect user fitting that category (knowingly ok with lack of linearizability) is really really small, and we always have to make trade-offs. So in that case I feel adding one more flag, one I consider dangerous, is not worth it. So to clarify, if a consensus appears for such flag, so be it, I'll add it, but I'm personally not neutral either.","from":"developer"},{"body":"bq. I'm rather cold on that because, tbh. I think non-strict serializability is a theoretical notion that is useless in practice and that it is something we should not offer. And I'd rather avoid one more \"feature\" for which we spend our time saying \"don't use it\".\r\n\r\nYeah, I'm very sympathetic to this view, and have always assumed linearizability with partitions as the object. I'm just really trying to morally justify providing some time to fix this without any negative repercussions. \r\n\r\nEither way, we should definitely clarify what we mean by SERIAL in some official project documentation somewhere though. We probably need to do so in terms of strict serializability as opposed to linearizability, so that it can be consistent with a future in which we support multi-partition transactions (which as a project we really need to deliver in the not-too-distant future).\r\n\r\nbq. non-applying CAS for that\r\n\r\nFWIW, I think this particular case is a no-brainer; there's no real cost to strengthening the semantics of non-applying CAS IMO, since users should anticipate their CAS operations will ordinarily take this long. Whatever the conclusion of our discussion, I think we should apply a fix at least for the non-applying case immediately, and I do not believe any flag to disable this part of the fix is necessary.\r\n\r\nReads are trickier, because the user will see a significant performance penalty on patch version upgrade. I'm sympathetic to the view we should just fix the read part immediately, performance regressions be damned. But we do have other serious consistency violations that should also be fixed. I think it is worth _considering_ if we should instead aggressively try to remedy all of the known issues, have a strong verification push, and then roll out all of the changes at-once - including a fix for this that does not regress performance. It might seem a lot for a patch version, but I'm not sure risk is a concern when we know there are several serious problems today, and have been for years.\r\n\r\nI'm not going to advocate super strongly for either approach, as I don't think there's a clear answer, I just want to raise the alternative as an option to expressly consider.\r\n\r\nbq. Awesome, thanks. I'll look at integrating those in the branch if you don't mind.\r\n\r\nAbsolutely, that was my intention.","from":"developer"},{"body":"bq. But we do have other serious consistency violations that should also be fixed.\r\n\r\nCould you expand on that?\r\n","from":"developer"},{"body":"The test cases I provided demonstrate several consistency violations during range movements. I've just thought of another one, and am writing a test case for it. Perhaps we could claim that range movements are always (potentially) consistency violations, but they are particularly keenly felt when you claim a linearisable history.\r\n\r\nThere are also (more debatably) issues with TTL on {{system.paxos}}, particularly when mixed with non-global commit; perhaps we could claim this is the user's problem, but it's not clear why we support global consensus that can be lost through local commit, and I don't think we communicate clearly the consistency implications to not call this a bug.\r\n\r\nAlso, mixing LOCAL_SERIAL and SERIAL is entirely unsafe, and even supporting them both is arguably a consistency violation without mechanisms to safely transition from one level to another.","from":"developer"},{"body":"bq. The test cases I provided demonstrate several consistency violations during range movements.\r\n\r\nYes, sorry I hadn't read them before commenting. And I certainly agree those are problematic (I was about to open a ticket so it's tracked, but I'd say CASSANDRA-15745 kind of cover those).\r\n\r\nbq. There are also (more debatably) issues with TTL on system.paxos\r\n\r\nAgreed this has always been a weak point. It does feel somewhat separated of other consistency points though, and maybe short term we can just offer a way to override the TTL (with documentation on the tradeoffs involved)?\r\n\r\nbq. Also, mixing LOCAL_SERIAL and SERIAL is entirely unsafe\r\n\r\nYeah. I'm not sure how to fix that one without a breaking API change though (namely, limiting their unrestricted use together). It's not \"that\" different from the fact we allow unrestricted mixing of serial and non-serial operations. Which is something I don't like and I'm happy to discuss moving forward, but imo post-3.X material in the best of cases.\r\n\r\nbq. I think it is worth considering if we should instead aggressively try to remedy all of the known issues, have a strong verification push, and then roll out all of the changes at-once - including a fix for this that does not regress performance.\r\n\r\nIt is certainly an option worth bringing, and thank you for that. I'm not sure how to really know what is the best option though, so I can only offer my current opinion.\r\n\r\nWhich is that I feel this issue is a very serious issue. And I don't mean that in a way that diminishes the seriousness of the other problems you mentioned, I mean that in absolute terms (the range movement issues are also fairly bad imo for instance). But leaving less of our known serious unaddressed feels better than not, so I'd personally prefer fixing that issue ASAP. Basically, I'm worried that waiting for a more all-encompassing fix might take us quite some time, with no absolute guarantee that we'll be collectively at ease with pushing that to 3.X.\r\n\r\nAnyway, I'd like to move this forward personally. How do we decide if we do?\r\n","from":"developer"},{"body":"> I'd like to move this forward personally. \r\n\r\nSure, go for it.","from":"developer"},{"body":"Ok, I've rebased the patch against 4.0 and started CI on it all:\r\n||branch||CI||\r\n|[3.0|https://github.com/pcmanus/cassandra/tree/C-12126-3.0]|[Run #146|https://ci-cassandra.apache.org/job/Cassandra-devbranch/146/]|\r\n|[3.11|https://github.com/pcmanus/cassandra/tree/C-12126-3.11]|[Run #147|https://ci-cassandra.apache.org/job/Cassandra-devbranch/147/]|\r\n|[4.0|https://github.com/pcmanus/cassandra/tree/C-12126-4.0]|[Run #148|https://ci-cassandra.apache.org/job/Cassandra-devbranch/148/]|\r\n\r\nI included a commit to add the flag that disables the new empty commit for SERIAL reads as suggested by [~bdeggleston] earlier. Still slightly on the fence on the need for such flag, but I call it \"unsafe\" ({{-Dcassandra.unsafe.disable-serial-reads-linearizability}} to be specific) and log a warning when used, so I'm at peace with that.\r\n\r\nI'll note for future reviewers that while the 3.11 branch is almost a straight away merge up of 3.0, there is a minor differences on the 4.0 branch, namely:\r\n * the added in-jvm dtests needed a few changes to reflect 4.0 changes. To make that easier, I squashed 2 of the commits from the 3.0/3.11 branches, which is why that branch has one less commit.\r\n * There is a few changes related to the translation of {{WriteTimeoutException}} into {{CasWriteTimeoutException}} (I pushed it down in some cases). I believe this fixes a minor \"bug\" where the \"contentions\" number we returned with {{CasWriteTimeoutException}} was potentially inaccurate (namely, if we timed out in {{beginRepairAndPaxos}}, contention leading to that exception would be ignored)\r\n\r\nI'll wait on getting usable CI results to officially mark it 'ready to review', but it is in spirit if anyone is burning to look at this.","from":"developer"},{"body":"I'm only semi-sure how to parse Jenkins CI results these days but from what I can tell, all failures are unrelated so marking ready for review.","from":"developer"},{"body":"I noticed that the previous version of the patches wasn't working in all cases due to an existing quirk of the CAS implementation.\r\n\r\nNamely, accepted updates that were empty were not replayed by {{beginAndRepairPaxos}}. Which is a problem for the new empty commits made during serial reads/non-applying CAS. I added tests to show that if the commit messages for those empty commits were lost/delayed, we could still have linearizability violations.\r\n\r\nNow, the logic of not replaying empty updates looks wrong to me. There shouldn't be anything special about an empty update, and if one is explicitely accepted by a quorum of nodes, we shouldn't ignore it, or that's a break of the Paxos algorithm (as kind of can be demonstrated by the tests I added).\r\n\r\nTo be clear, that logic was added *by me* in CASSANDRA-6012 and that was the sole purpose of that ticket. Except that I can't make sense of my reasoning back then, and since I didn't included a test to demonstrate the problem I was solving back then (which was wrong, mea culpa), I have to assume that I was just confused (maybe I mixed in my head promised ballots and accepted ones?). Anyway, I think the fix here is simply to remove that bad logic, which fixes the issue, and I included an additional commit for that.\r\n","from":"developer"},{"body":"So, I'm reasonably sure it cannot be necessary for us to commit an empty proposal, because we do not ever need to witness it. Either the proposal was agreed by a quorum (and the proposer can report this) but it has no visible effect on future proposals, and does not need to be witnessed by anybody else, or it was not agreed and it does not need to be either proposed again, committed or witnessed by anybody else.\r\n\r\nHowever we have to be consistent about it: we either need to _never_ commit them, or _always_ commit them.","from":"developer"},{"body":"bq. I'm reasonably sure it cannot be necessary for us to commit an empty proposal, because we do not ever need to witness it.\r\n\r\nWe may have to be precise. We do not need to \"apply\" an empty commit, since it's a no-op, and the patch actually ensures we don't bother. But \"committed\" do something else, it update the \"mrc\" value, and _that_ needs to be done. Otherwise, if we _accept_ an empty proposal, yet does not update the \"mrc\" value, we will not do progress anymore (well, without additional modification to the algorithm that is).\r\n\r\nBut I could be misunderstanding what you are suggesting here. I'll note though, just in case that help, that the logic I'm calling faulty is not the _commit_ of empty updates (though, as said above, I think it's necessary for the sake of the mrc value), it's the fact the don't replay the _proposal_ of empty updates. ","from":"developer"},{"body":"The problem here stems only from the overload of {{mostRecentInProgressCommitWithUpdate}}, which (seems to) assume that an empty update is for a higher promise (since the meaning is overloaded in the response message) rather than an \"incomplete\" proposal. If the empty proposal were to be correctly merged with {{mostRecentInProgressCommitWithUpdate}}, it would override the early non-empty incomplete proposal.\r\n\r\nWhich is a long-winded way of saying that I am fairly confident there's no need to update the paxos state table with the \"committed\" status of this empty proposal so long as it remains in the table _as an accepted proposal_, and so long as this accepted proposal continues to override earlier in progress proposals.\r\n","from":"developer"},{"body":"I'll have to apologize, but I don't understand what you are suggesting.","from":"developer"},{"body":"I'm not proposing we do anything different for your patch, just clarifying that this isn't strictly necessary - it is quite possible to modify the algorithm to never commit empty proposals. The problem today is that we:\r\n\r\n# \"Refresh\" a quorum with the MRC if not witnessed by all promisers \r\n# Filter out empty proposals when deciding if we have an in progress proposal ({{mostRecentInProgressCommit}} vs {{mostRecentInProgressCommitWithUpdate}})\r\n\r\nIf instead we did not refresh empty commits, and we did not filter out empty proposals when _updating_ {{mostRecentInProgressCommitWithUpdate}} but did not _complete_ any empty proposals we found then everything would be fine.\r\n\r\n{{mostRecentInProgressCommitWithUpdate}} confuses matters because it is poorly named, and is updated by its naming rather than intent - I think it is _meant_ to be {{mostRecentInProgressProposal}} whereas {{mostRecentInProgressCommit}} should be e.g. {{mostRecentInProgressPromiseOrProposal}}, and {{mostRecentInProgressProposal}} would gain empty proposals as well as non-empty ones, and correctly discount the older in progress proposal that was invalidated by the newer read that did not witness it.\r\n\r\nTo be clear, I'm mostly participating in this discussion for my own benefit and for the benefit of future work, not trying to solicit changes to your work.","from":"developer"},{"body":"To say it another way: the only purpose of an empty proposal is to poison earlier proposals, and this can be done just as well without moving the proposal to the MRC column in the table. If we treat it as any other \"in progress\" proposal for invalidating earlier proposals, then once we reach a quorum we must in future be witnessed alongside any earlier proposals and invalidate them. If we didn't reach a quorum, then it doesn't matter if we are witnessed or not, or if any earlier proposals are invalidated.\r\n\r\n","from":"developer"},{"body":"Ok, I understand what you are suggesting now and I agree this should work as well. And it does is more optimal.\r\n\r\nI like to think of our algorithm as \"pure Paxos instances\" separated by the MRC to tell us when we can forget the previous instance and start a new one. Committing empty updates as any other updates still fits that mental model, while your suggestion adds a bit of a special case in that it bends the Paxos rules slightly, allowing to sometime ignore a previously accepted value in a promise (when it's empty). Which is not a criticism, just thinking out loud. It's more performant and this is likely worth the slight special casing since it's not too hard to reason about its correctness.\r\n\r\nI'll sleep on it and modify to your suggestion tomorrow (which is trivial, just need to massage an appropriate comment to explain it).\r\n","from":"developer"},{"body":"Alright, my \"tomorrow\" is off by 1, but pushed an additional commit to implement the optimization suggested by Benedict above. Restarted CI for good measure.\r\n\r\n||branch||CI||\r\n|[3.0|https://github.com/pcmanus/cassandra/tree/C-12126-3.0]|[Run #155|https://ci-cassandra.apache.org/job/Cassandra-devbranch/155/]|\r\n|[3.11|https://github.com/pcmanus/cassandra/tree/C-12126-3.11]|[Run #156|https://ci-cassandra.apache.org/job/Cassandra-devbranch/156/]|\r\n|[4.0|https://github.com/pcmanus/cassandra/tree/C-12126-4.0]|[Run #157|https://ci-cassandra.apache.org/job/Cassandra-devbranch/157/]|\r\n","from":"developer"},{"body":"Sorry, for the delay. The patches looks good.","from":"developer"},{"body":"Thanks for the review. I've rebased the branches, but since the last runs were a while ago, I restarted CI runs. I'll commit if those look clean.\r\n\r\n||branch||CI||\r\n|[3.0|https://github.com/pcmanus/cassandra/tree/C-12126-3.0]|[Run #171|https://ci-cassandra.apache.org/job/Cassandra-devbranch/171/]|\r\n|[3.11|https://github.com/pcmanus/cassandra/tree/C-12126-3.11]|[Run #172|https://ci-cassandra.apache.org/job/Cassandra-devbranch/172/]|\r\n|[4.0|https://github.com/pcmanus/cassandra/tree/C-12126-4.0]|[Run #173|https://ci-cassandra.apache.org/job/Cassandra-devbranch/173/]|\r\n","from":"developer"},{"body":"So, before we commit this I wanted to share that some experimentation found that this can lead to a significant increase in timeouts, particularly for read-heavy workloads, that previously would not have competed with each other. I think committing this to a patch release is honestly problematic, as it could surprise users with a service outage. At the very least, there should be HUGE warnings in {{NEWS.txt}}, but honestly I would prefer to have users opt-in for patch releases.\r\n\r\nAs much as I agree that it is problematic to provide the wrong semantics, I think it is also problematic to force a decision between stability and correctness onto our users without their informed and positive consent.\r\n\r\nI hope that I will be able to provide the community with an alternative solution in the near future, without these (and many other existing) pitfalls. However I'm not sure how that should affect this decision.","from":"developer"},{"body":"{quote}I hope that I will be able to provide the community with an alternative solution in the near future, without these (and many other existing) pitfalls.{quote}\r\n\r\n[~benedict] Few questions regarding your comment:\r\n* What timeframe do you have in mind? \r\n* Is it a solution only for 4.0 or for all the branches?\r\n* Can we help you with that?","from":"developer"},{"body":"To some extent that is all up for debate.\r\n\r\n\r\n My plan so far has been to avoid interfering with 4.0 release, so I have been working towards targeting 4.x. This would also permit time to produce documentation and reach out to the list to begin the slow handshake to see if the project wants the work, and in what manner. However, the main body of work is essentially complete, so it is possible that this could be brought forwards if there were appetite.\r\n As to target version, it would be possible to target 3.0+, at least for a portion of the work that would encompass this issue, without a great deal of work. The project's appetite would be the main decider, as it's a significant body of work.\r\n\r\n\r\n The main contribution would be a parallel implementation of the same underlying Paxos algorithm, that is able to run concurrently alongside it (supporting live migration), but with several latency improvements, as well as several fixes to correctness. Alongside this is related work to guarantee linearizability across range movements in the form of modifications to repair, bootstrap, replace etc.\r\n\r\n\r\n Related to this work are several patches to wider Cassandra to support automated verification of its correctness, by permitting deterministic simulation of Cassandra clusters with adversarial ordering of events. We have so far simulated billions of transactions to verify its linearizability. I anticipate that this work will be useful for the project's overall goal of improving quality, but they are themselves quite significant and will require their own discussions around timeline and scope.","from":"developer"},{"body":"It seem to me that there are several options here:\r\n# Try to use your proposal for 4.0 if the community has the appetite for it. The main issue there is some potential extra delay for 4.0\r\n# Do nothing for 4.0. Meaning do not commit the patch. We have lived a long time with that issue and we can probably wait a bit more for a proper solution.\r\n# Commit the patch as such, fixing the correctness but introducting potentially some performance issue until we release a better solution.\r\n# Changing the patch to default to the current behavior but allowing people to enable the new one if the correctness is a problem for them.\r\n\r\nMay be we should trigger a discussion on the mailing list and see what is other people opinion.\r\n\r\nI can take care of that next week if you think it is a good idea.","from":"developer"},{"body":"Yes, that sounds like a great idea, and I really appreciate you offering to take that to the list. I'll chime in with any necessary details to help inform the decision, but will try not to influence it otherwise. I don't have a strong opinion about which of those four options we select, except that my experiments do suggest (3) is perhaps dangerous for some of our users. It's probably a trade-off that should be made with careful business consideration and experimentation by each end user.\r\n\r\nAs far as delaying 4.0 is concerned, that's probably also a matter of community decision-making. We could quite quickly have a patch, that has been reviewed by multiple committers, posted in fairly short order - perhaps before we exit beta. This work will have had much greater validation than the current implementation, but publishing all of this validation work will take longer - likely also achievable before GA, but we might have to invert our process a little. Perhaps this is acceptable, given the balance of correctness and regression we're considering as an alternative, but given my proximity to the work (and that I also don't have a strong position either way), I would prefer to let others make that call.","from":"developer"},{"body":"Committed following the dev mailing list discussion. Thanks.","from":"developer"},{"body":"> 3) Issue another CAS Read and it goes to A and B. Now we will discover that there is something inflight from A and will propose and commit it with the current ballot. Now we can read the value written in step 1 as part of this CAS read.\r\n\r\nSorry, I'm not fully sure about the current implementation and how realistic my proposal is but,\r\n\r\ncan we read all the replicas to do the read recovery in step 3 to solve the issue?\r\n\r\nIt only reads A and B but if we read C as well, we know that the proposal is not accepted by the majority.\r\n\r\n ","from":"developer"},{"body":"[~feeblefakie] C can be unreachable for different reasons.","from":"developer"},{"body":"Just a note that the bug that this fixes usually pops up as the following timeout for people looking for reasons why SERIAL or LOCAL_SERIAL are seeing read timeouts >3.11.10.  Setting the flag to the opt-out option will `fix` it but probably shouldn't be reading at this level if you run into this.\r\n{code:java}\r\n! com.datastax.driver.core.exceptions.ReadTimeoutException: Cassandra timeout during read query at consistency LOCAL_SERIAL (2 responses were required but only 0 replica responded)\r\n{code}","from":"developer"}],"created":"2016-07-01T17:46:09.000+0000","description":"While looking at the CAS code in Cassandra, I found a potential issue with CAS Reads. Here is how it can happen with RF=3\n\n1) You issue a CAS Write and it fails in the propose phase. A machine replies true to a propose and saves the commit in accepted filed. The other two machines B and C does not get to the accept phase. \n\nCurrent state is that machine A has this commit in paxos table as accepted but not committed and B and C does not. \n\n2) Issue a CAS Read and it goes to only B and C. You wont be able to read the value written in step 1. This step is as if nothing is inflight. \n\n3) Issue another CAS Read and it goes to A and B. Now we will discover that there is something inflight from A and will propose and commit it with the current ballot. Now we can read the value written in step 1 as part of this CAS read.\n\nIf we skip step 3 and instead run step 4, we will never learn about value written in step 1. \n\n4. Issue a CAS Write and it involves only B and C. This will succeed and commit a different value than step 1. Step 1 value will never be seen again and was never seen before. \n\n\n\nIf you read the Lamport “paxos made simple” paper and read section 2.3. It talks about this issue which is how learners can find out if majority of the acceptors have accepted the proposal. \n\nIn step 3, it is correct that we propose the value again since we dont know if it was accepted by majority of acceptors. When we ask majority of acceptors, and more than one acceptors but not majority has something in flight, we have no way of knowing if it is accepted by majority of acceptors. So this behavior is correct. \n\nHowever we need to fix step 2, since it caused reads to not be linearizable with respect to writes and other reads. In this case, we know that majority of acceptors have no inflight commit which means we have majority that nothing was accepted by majority. I think we should run a propose step here with empty commit and that will cause write written in step 1 to not be visible ever after. \n\nWith this fix, we will either see data written in step 1 on next serial read or will never see it which is what we want. \n","issue_id":"12986267","key":"CASSANDRA-12126","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2020-11-27T16:18:29.000+0000","role":"fixed_distractor","summary":"CAS Reads Inconsistencies "} {"case_id":"12467688","cluster":"DISTRACTOR-CASSANDRA-1221","comments":[{"body":"Hi Gary,\n\nI was able to reproduce this using today's nightly build. This time i used a smaller data set (500000 keys) and I got the following:\n\n[agoudarzi@cas-test3 scripts]$ nodetool --host 10.50.26.132 ring \nAddress Status State Load Token \n 160348796167900510561059505917619274541 \n10.50.26.134 Up Normal 116.98 MB 32717880524093094169411234083126184860 \n10.50.26.132 Up Leaving 58.58 MB 75101027859180840627831025901565139619 \n10.50.26.133 Up Normal 117.09 MB 160348796167900510561059505917619274541 \n\n[agoudarzi@cas-test3 scripts]$ nodetool --host 10.50.26.132 streams\nMode: Leaving: streaming data to other nodes\nStreaming to: /10.50.26.133\n /var/lib/cassandra/data/Keyspace1/Standard1-d-17-Data.db/[(0,54080834)]\nNot receiving any streams.\n[agoudarzi@cas-test3 scripts]$ nodetool --host 10.50.26.133 streams\nMode: Normal\nNot sending any streams.\nNot receiving any streams.\n\nFrom the logs of 10.50.26.132 it seams that it tried to tell 10.50.26.133 to claim its stream:\n\nINFO [STREAM-STAGE:1] 2010-07-13 16:50:35,994 StreamOut.java (line 135) Sending a stream initiate message to /10.50.26.133 ...\nINFO [STREAM-STAGE:1] 2010-07-13 16:50:35,994 StreamOut.java (line 140) Waiting for transfer to /10.50.26.133 to complete\n\nBut nothing in 133's log acknowledges the receipt of the request from 132 and as you see above it shows that it is getting no streams and this has been going for the past hour or so.\n\n-Arya\n\n","created":"2010-07-14T00:05:00.782+0000"},{"body":"Arya, can you supply the nodetool commands you are using that constitute \"cleanup\"? I've tried a few times now and can't get the failure you describe. In your latest test was .132 the second or third node booted?","created":"2010-07-14T21:50:34.388+0000"},{"body":"132 is node1\n133 is node2\n134 is node3\n\nGive me some time and I'll regenerate all the commands for you in details. ","created":"2010-07-14T22:55:52.800+0000"},{"body":"Gary, I missed one thing (Step 9-13) where I took a node down and insert. I tested this without step 9-13 and loadbalance worked. Not sure why step 9-13 changes everything. Here are the full production steps (I am also attaching the logs from all 3 nodes if it helps):\n\nThis is the ring topology I discuss here:\n\nNode1 10.50.26.132 (Hostname: cas-test1)\nNode2 10.50.26.133 (Hostname: cas-test2)\nNode3 10.50.26.134 (Hostname: cas-test3)\n\nThis run is using today's nightly built from a clean setup.\n\nStep 1: Startup Node1\n\n[agoudarzi@cas-test1 ~]$ sudo /etc/init.d/cassandra start\n\nStep 2: LoadSchemadFromYAML\n\nI go to JConsole and call the function from o.a.c.service StorageService MBeans\n\nStep 3: I insert 500000 keys into Standard1 CF using py_stress\n\n$ python stress.py --num-keys 500000 --threads 8 --nodes 10.50.26.132 --keep-going --operation insert\nKeyspace already exists.\ntotal,interval_op_rate,interval_key_rate,avg_latency,elapsed_time\n62455,6245,6245,0.00128478354559,10\n121893,5943,5943,0.00134767375398,21\n184298,6240,6240,0.00128336335573,31\n248124,6382,6382,0.00124898537112,42\n297205,4908,4908,0.00163303957852,52\n340338,4313,4313,0.00189026848124,63\n380233,3989,3989,0.00203818801591,73\n444452,6421,6421,0.00124198496903,84\n500000,5554,5554,0.00114441244599,93\n\nStep 3: Let's Take a Look at Ring\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n10.50.26.132 Up Normal 206.05 MB 139380634429053457983268837561452509806 \n\nStep 4: Bootstrap Node 2 into cluster\n\n[agoudarzi@cas-test2 ~]$ sudo /etc/init.d/cassandra start\n\nStep 5: Check the ring\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Joining 5.84 KB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 206.05 MB 139380634429053457983268837561452509806 \n\nStep 6: Check the streams on Node 1\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 streams\nMode: Normal\nStreaming to: /10.50.26.133\n /var/lib/cassandra/data/Keyspace1/Standard1-d-5-Data.db/[(0,34658183), (89260032,109057810)]\n /var/lib/cassandra/data/Keyspace1/Standard1-d-6-Data.db/[(0,8746823), (22272929,27264363)]\n /var/lib/cassandra/data/Keyspace1/Standard1-d-8-Data.db/[(0,8749389), (22336617,27264253)]\n /var/lib/cassandra/data/Keyspace1/Standard1-d-9-Data.db/[(0,8235190), (21174782,25822054)]\n /var/lib/cassandra/data/Keyspace1/Standard1-d-7-Data.db/[(0,8642472), (22333239,27264347)]\nNot receiving any streams.\n\nStep 7: Check the ring from Node1 and Node2 and make sure they agree\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 233.93 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 233.93 MB 139380634429053457983268837561452509806\n\nStep 8: Cleanup Node1\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 cleanup\n\nStep 9: Check the ring agreement again\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 117 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 117 MB 139380634429053457983268837561452509806 \n\nStep 9: Let's kill Node 2\n\n[agoudarzi@cas-test2 ~]$ sudo /etc/init.d/cassandra stop \n\nStep 10: Check the ring on Node 1\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Down Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 117 MB 139380634429053457983268837561452509806\n\nStep 11: Let's try to insert 500000 more keys expecting lots of unavailable exceptions ad the only replica for some keys is dead and py_stress does not use CLevel.ANY or ZERO\n\n$ python stress.py --num-keys 500000 --threads 8 --nodes 10.50.26.132 --keep-going --operation insert\n\nKeyspace already exists.\nUnavailableException()\nUnavailableException()\nUnavailableException()\n....\n...\n..\n.\ntotal,interval_op_rate,interval_key_rate,avg_latency,elapsed_time\n500000,2922,2922,0.000816067281446,67\n\nStep 12: Check the ring\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Down Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 205.35 MB 139380634429053457983268837561452509806 \n\nNode 1 got more data as expected\n\nStep 13: Bring up Node 2 again\n\n[agoudarzi@cas-test2 ~]$ sudo /etc/init.d/cassandra start\n\nStep 14: Check the Ring\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 205.35 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 205.35 MB 139380634429053457983268837561452509806\n\nStep 15: Bootstrap Node 3\n\n[agoudarzi@cas-test3 ~]$ sudo /etc/init.d/cassandra start\n\nStep 11: Check Ring \n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 234 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 234 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.134 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 234 MB 139380634429053457983268837561452509806 \n\nStep 12: Cleanup Node 1 (132)\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 cleanup\n\nStep 13: Check Ring\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 58.89 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 58.89 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.134 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 58.89 MB 139380634429053457983268837561452509806 \n\nLooks as expected. Node 1 (132) has the least load. Let's loadbalance it.\n\nStep 14: Loadbalance Node 1\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 loadbalance &\n[1] 27457\n\nStep 15: Check Ring\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Leaving 58.89 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Leaving 58.89 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.134 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Leaving 58.89 MB 139380634429053457983268837561452509806 \n\nStep 16: Check the Streams\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 streams\nMode: Leaving: streaming data to other nodes\nStreaming to: /10.50.26.133\n /var/lib/cassandra/data/Keyspace1/Standard1-d-17-Data.db/[(0,54278564)]\nNot receiving any streams.\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 streams\nMode: Normal\nNot sending any streams.\nNot receiving any streams.\n\nPROBLEM:\nNotice 132 says I am streaming to 133 but 133 says \"Not receiving any streams!\" \n\n\n","created":"2010-07-15T00:20:55.567+0000"},{"body":"Cassandra System Logs for node 1-3","created":"2010-07-15T00:24:27.997+0000"},{"body":"Thanks Arya. I can reproduce this. Now just to fix it.","created":"2010-07-15T20:30:59.392+0000"},{"body":"Several problems on this ticket.\n1. MessaingService implemented IFailureDetector and was in charge of shutting down TCP connections during a partition. However, it was never added to the listeners in FD. This meant that MS.convict() was never getting called.\n2. Under the right conditions (I still don't understand this fully), java sockets still give every indication they are connected even though the host on the other end is down. A single write will succeed, even though no bytes are sent. Seriously. In our case the single write was a StreamInitiateMessage that we then wait forever to be acked the the [dead] remote host.\n\nMy solution was to use Gossiper.convict, which calls SS.onDead, to call MS.convict(). I don't think it makes sense to have MS implement IFailureDetector since we treat Gossiper as authoritative with respect to node alive-ness.\n\nI'll spend some time checking this morning, but I suspect we have the same situation in 0.6.","created":"2010-07-20T11:48:34.378+0000"},{"body":"patch for 0.6. I couldn't get stress.py to work in my branch, but the same problem should be present. All tests pass with this patch.","created":"2010-07-20T15:25:48.804+0000"},{"body":"+1\n\n(but fix brace placement in onDead please)","created":"2010-07-20T16:17:32.837+0000"}],"conversations":[{"body":"Arya Goudarzi reports:\n\nPlease confirm if this is an issue and should be reported or I am doing something wrong. I could not find anything relevant on JIRA:\n\nPlaying with 0.7 nightly (today's build), I setup a 3 node cluster this way:\n\n - Added one node;\n - Loaded default schema with RF 1 from YAML using JMX;\n - Loaded 2M keys using py_stress;\n - Bootstrapped a second node;\n - Cleaned up the first node;\n - Bootstrapped a third node;\n - Cleaned up the second node;\n\nI got the following ring:\n\nAddress Status Load Range Ring\n 154293670372423273273390365393543806425\n10.50.26.132 Up 518.63 MB 69164917636305877859094619660693892452 |<--|\n10.50.26.134 Up 234.8 MB 111685517405103688771527967027648896391 | |\n10.50.26.133 Up 235.26 MB 154293670372423273273390365393543806425 |-->|\n\nNow I ran:\n\nnodetool --host 10.50.26.132 loadbalance\n\nIt's been going for a while. I checked the streams\n\nnodetool --host 10.50.26.134 streams\nMode: Normal\nNot sending any streams.\nStreaming from: /10.50.26.132\n Keyspace1: /var/lib/cassandra/data/Keyspace1/Standard1-tmp-d-3-Data.db/[(0,22206096), (22206096,27271682)]\n Keyspace1: /var/lib/cassandra/data/Keyspace1/Standard1-tmp-d-4-Data.db/[(0,15180462), (15180462,18656982)]\n Keyspace1: /var/lib/cassandra/data/Keyspace1/Standard1-tmp-d-5-Data.db/[(0,353139829), (353139829,433883659)]\n Keyspace1: /var/lib/cassandra/data/Keyspace1/Standard1-tmp-d-6-Data.db/[(0,366336059), (366336059,450095320)]\n\nnodetool --host 10.50.26.132 streams\nMode: Leaving: streaming data to other nodes\nStreaming to: /10.50.26.134\n /var/lib/cassandra/data/Keyspace1/Standard1-d-48-Data.db/[(0,366336059), (366336059,450095320)]\nNot receiving any streams.\n\nThese have been going for the past 2 hours.\n\nI see in the logs of the node with 134 IP address and I saw this:\n\nINFO [GOSSIP_STAGE:1] 2010-06-22 16:30:54,679 StorageService.java (line 603) Will not change my token ownership to /10.50.26.132\n\nSo, to my understanding from wikis loadbalance supposed to decommission and re-bootstrap again by sending its tokens to other nodes and then bootstrap again. It's been stuck in streaming for the past 2 hours and the size of ring has not changed. The log in the first node says it has started streaming for the past hours:\n\nINFO [STREAM-STAGE:1] 2010-06-22 16:35:56,255 StreamOut.java (line 72) Beginning transfer process to /10.50.26.134 for ranges (154293670372423273273390365393543806425,69164917636305877859094619660693892452]\n INFO [STREAM-STAGE:1] 2010-06-22 16:35:56,255 StreamOut.java (line 82) Flushing memtables for Keyspace1...\n INFO [STREAM-STAGE:1] 2010-06-22 16:35:56,266 StreamOut.java (line 128) Stream context metadata [/var/lib/cassandra/data/Keyspace1/Standard1-d-48-Data.db/[(0,366336059), (366336059,450095320)]] 1 sstables.\n INFO [STREAM-STAGE:1] 2010-06-22 16:35:56,267 StreamOut.java (line 135) Sending a stream initiate message to /10.50.26.134 ...\n INFO [STREAM-STAGE:1] 2010-06-22 16:35:56,267 StreamOut.java (line 140) Waiting for transfer to /10.50.26.134 to complete\n INFO [FLUSH-TIMER] 2010-06-22 17:36:53,370 ColumnFamilyStore.java (line 359) LocationInfo has reached its threshold; switching in a fresh Memtable at CommitLogContext(file='/var/lib/cassandra/commitlog/CommitLog-1277249454413.log', position=720)\n INFO [FLUSH-TIMER] 2010-06-22 17:36:53,370 ColumnFamilyStore.java (line 622) Enqueuing flush of Memtable(LocationInfo)@1637794189\n INFO [FLUSH-WRITER-POOL:1] 2010-06-22 17:36:53,370 Memtable.java (line 149) Writing Memtable(LocationInfo)@1637794189\n INFO [FLUSH-WRITER-POOL:1] 2010-06-22 17:36:53,528 Memtable.java (line 163) Completed flushing /var/lib/cassandra/data/system/LocationInfo-d-9-Data.db\n INFO [MEMTABLE-POST-FLUSHER:1] 2010-06-22 17:36:53,529 ColumnFamilyStore.java (line 374) Discarding 1000\n\n\nNothing more after this line.\n\nAm I doing something wrong?","from":"reporter","subject":"loadbalance operation never completes on a 3 node cluster"},{"body":"Hi Gary,\n\nI was able to reproduce this using today's nightly build. This time i used a smaller data set (500000 keys) and I got the following:\n\n[agoudarzi@cas-test3 scripts]$ nodetool --host 10.50.26.132 ring \nAddress Status State Load Token \n 160348796167900510561059505917619274541 \n10.50.26.134 Up Normal 116.98 MB 32717880524093094169411234083126184860 \n10.50.26.132 Up Leaving 58.58 MB 75101027859180840627831025901565139619 \n10.50.26.133 Up Normal 117.09 MB 160348796167900510561059505917619274541 \n\n[agoudarzi@cas-test3 scripts]$ nodetool --host 10.50.26.132 streams\nMode: Leaving: streaming data to other nodes\nStreaming to: /10.50.26.133\n /var/lib/cassandra/data/Keyspace1/Standard1-d-17-Data.db/[(0,54080834)]\nNot receiving any streams.\n[agoudarzi@cas-test3 scripts]$ nodetool --host 10.50.26.133 streams\nMode: Normal\nNot sending any streams.\nNot receiving any streams.\n\nFrom the logs of 10.50.26.132 it seams that it tried to tell 10.50.26.133 to claim its stream:\n\nINFO [STREAM-STAGE:1] 2010-07-13 16:50:35,994 StreamOut.java (line 135) Sending a stream initiate message to /10.50.26.133 ...\nINFO [STREAM-STAGE:1] 2010-07-13 16:50:35,994 StreamOut.java (line 140) Waiting for transfer to /10.50.26.133 to complete\n\nBut nothing in 133's log acknowledges the receipt of the request from 132 and as you see above it shows that it is getting no streams and this has been going for the past hour or so.\n\n-Arya\n\n","from":"developer"},{"body":"Arya, can you supply the nodetool commands you are using that constitute \"cleanup\"? I've tried a few times now and can't get the failure you describe. In your latest test was .132 the second or third node booted?","from":"developer"},{"body":"132 is node1\n133 is node2\n134 is node3\n\nGive me some time and I'll regenerate all the commands for you in details. ","from":"developer"},{"body":"Gary, I missed one thing (Step 9-13) where I took a node down and insert. I tested this without step 9-13 and loadbalance worked. Not sure why step 9-13 changes everything. Here are the full production steps (I am also attaching the logs from all 3 nodes if it helps):\n\nThis is the ring topology I discuss here:\n\nNode1 10.50.26.132 (Hostname: cas-test1)\nNode2 10.50.26.133 (Hostname: cas-test2)\nNode3 10.50.26.134 (Hostname: cas-test3)\n\nThis run is using today's nightly built from a clean setup.\n\nStep 1: Startup Node1\n\n[agoudarzi@cas-test1 ~]$ sudo /etc/init.d/cassandra start\n\nStep 2: LoadSchemadFromYAML\n\nI go to JConsole and call the function from o.a.c.service StorageService MBeans\n\nStep 3: I insert 500000 keys into Standard1 CF using py_stress\n\n$ python stress.py --num-keys 500000 --threads 8 --nodes 10.50.26.132 --keep-going --operation insert\nKeyspace already exists.\ntotal,interval_op_rate,interval_key_rate,avg_latency,elapsed_time\n62455,6245,6245,0.00128478354559,10\n121893,5943,5943,0.00134767375398,21\n184298,6240,6240,0.00128336335573,31\n248124,6382,6382,0.00124898537112,42\n297205,4908,4908,0.00163303957852,52\n340338,4313,4313,0.00189026848124,63\n380233,3989,3989,0.00203818801591,73\n444452,6421,6421,0.00124198496903,84\n500000,5554,5554,0.00114441244599,93\n\nStep 3: Let's Take a Look at Ring\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n10.50.26.132 Up Normal 206.05 MB 139380634429053457983268837561452509806 \n\nStep 4: Bootstrap Node 2 into cluster\n\n[agoudarzi@cas-test2 ~]$ sudo /etc/init.d/cassandra start\n\nStep 5: Check the ring\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Joining 5.84 KB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 206.05 MB 139380634429053457983268837561452509806 \n\nStep 6: Check the streams on Node 1\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 streams\nMode: Normal\nStreaming to: /10.50.26.133\n /var/lib/cassandra/data/Keyspace1/Standard1-d-5-Data.db/[(0,34658183), (89260032,109057810)]\n /var/lib/cassandra/data/Keyspace1/Standard1-d-6-Data.db/[(0,8746823), (22272929,27264363)]\n /var/lib/cassandra/data/Keyspace1/Standard1-d-8-Data.db/[(0,8749389), (22336617,27264253)]\n /var/lib/cassandra/data/Keyspace1/Standard1-d-9-Data.db/[(0,8235190), (21174782,25822054)]\n /var/lib/cassandra/data/Keyspace1/Standard1-d-7-Data.db/[(0,8642472), (22333239,27264347)]\nNot receiving any streams.\n\nStep 7: Check the ring from Node1 and Node2 and make sure they agree\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 233.93 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 233.93 MB 139380634429053457983268837561452509806\n\nStep 8: Cleanup Node1\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 cleanup\n\nStep 9: Check the ring agreement again\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 117 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 117 MB 139380634429053457983268837561452509806 \n\nStep 9: Let's kill Node 2\n\n[agoudarzi@cas-test2 ~]$ sudo /etc/init.d/cassandra stop \n\nStep 10: Check the ring on Node 1\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Down Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 117 MB 139380634429053457983268837561452509806\n\nStep 11: Let's try to insert 500000 more keys expecting lots of unavailable exceptions ad the only replica for some keys is dead and py_stress does not use CLevel.ANY or ZERO\n\n$ python stress.py --num-keys 500000 --threads 8 --nodes 10.50.26.132 --keep-going --operation insert\n\nKeyspace already exists.\nUnavailableException()\nUnavailableException()\nUnavailableException()\n....\n...\n..\n.\ntotal,interval_op_rate,interval_key_rate,avg_latency,elapsed_time\n500000,2922,2922,0.000816067281446,67\n\nStep 12: Check the ring\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Down Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 205.35 MB 139380634429053457983268837561452509806 \n\nNode 1 got more data as expected\n\nStep 13: Bring up Node 2 again\n\n[agoudarzi@cas-test2 ~]$ sudo /etc/init.d/cassandra start\n\nStep 14: Check the Ring\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 205.35 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.132 Up Normal 205.35 MB 139380634429053457983268837561452509806\n\nStep 15: Bootstrap Node 3\n\n[agoudarzi@cas-test3 ~]$ sudo /etc/init.d/cassandra start\n\nStep 11: Check Ring \n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 234 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 234 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.134 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 234 MB 139380634429053457983268837561452509806 \n\nStep 12: Cleanup Node 1 (132)\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 cleanup\n\nStep 13: Check Ring\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 58.89 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 58.89 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.134 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Normal 58.89 MB 139380634429053457983268837561452509806 \n\nLooks as expected. Node 1 (132) has the least load. Let's loadbalance it.\n\nStep 14: Loadbalance Node 1\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 loadbalance &\n[1] 27457\n\nStep 15: Check Ring\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Leaving 58.89 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Leaving 58.89 MB 139380634429053457983268837561452509806 \n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.134 ring\nAddress Status State Load Token \n 139380634429053457983268837561452509806 \n10.50.26.133 Up Normal 116.94 MB 54081521187303805945240848606999860232 \n10.50.26.134 Up Normal 116.68 MB 96565427321648203609592911083606603165 \n10.50.26.132 Up Leaving 58.89 MB 139380634429053457983268837561452509806 \n\nStep 16: Check the Streams\n\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.132 streams\nMode: Leaving: streaming data to other nodes\nStreaming to: /10.50.26.133\n /var/lib/cassandra/data/Keyspace1/Standard1-d-17-Data.db/[(0,54278564)]\nNot receiving any streams.\n[agoudarzi@cas-test1 ~]$ nodetool --host=10.50.26.133 streams\nMode: Normal\nNot sending any streams.\nNot receiving any streams.\n\nPROBLEM:\nNotice 132 says I am streaming to 133 but 133 says \"Not receiving any streams!\" \n\n\n","from":"developer"},{"body":"Cassandra System Logs for node 1-3","from":"developer"},{"body":"Thanks Arya. I can reproduce this. Now just to fix it.","from":"developer"},{"body":"Several problems on this ticket.\n1. MessaingService implemented IFailureDetector and was in charge of shutting down TCP connections during a partition. However, it was never added to the listeners in FD. This meant that MS.convict() was never getting called.\n2. Under the right conditions (I still don't understand this fully), java sockets still give every indication they are connected even though the host on the other end is down. A single write will succeed, even though no bytes are sent. Seriously. In our case the single write was a StreamInitiateMessage that we then wait forever to be acked the the [dead] remote host.\n\nMy solution was to use Gossiper.convict, which calls SS.onDead, to call MS.convict(). I don't think it makes sense to have MS implement IFailureDetector since we treat Gossiper as authoritative with respect to node alive-ness.\n\nI'll spend some time checking this morning, but I suspect we have the same situation in 0.6.","from":"developer"},{"body":"patch for 0.6. I couldn't get stress.py to work in my branch, but the same problem should be present. All tests pass with this patch.","from":"developer"},{"body":"+1\n\n(but fix brace placement in onDead please)","from":"developer"}],"created":"2010-06-23T12:39:42.000+0000","description":"Arya Goudarzi reports:\n\nPlease confirm if this is an issue and should be reported or I am doing something wrong. I could not find anything relevant on JIRA:\n\nPlaying with 0.7 nightly (today's build), I setup a 3 node cluster this way:\n\n - Added one node;\n - Loaded default schema with RF 1 from YAML using JMX;\n - Loaded 2M keys using py_stress;\n - Bootstrapped a second node;\n - Cleaned up the first node;\n - Bootstrapped a third node;\n - Cleaned up the second node;\n\nI got the following ring:\n\nAddress Status Load Range Ring\n 154293670372423273273390365393543806425\n10.50.26.132 Up 518.63 MB 69164917636305877859094619660693892452 |<--|\n10.50.26.134 Up 234.8 MB 111685517405103688771527967027648896391 | |\n10.50.26.133 Up 235.26 MB 154293670372423273273390365393543806425 |-->|\n\nNow I ran:\n\nnodetool --host 10.50.26.132 loadbalance\n\nIt's been going for a while. I checked the streams\n\nnodetool --host 10.50.26.134 streams\nMode: Normal\nNot sending any streams.\nStreaming from: /10.50.26.132\n Keyspace1: /var/lib/cassandra/data/Keyspace1/Standard1-tmp-d-3-Data.db/[(0,22206096), (22206096,27271682)]\n Keyspace1: /var/lib/cassandra/data/Keyspace1/Standard1-tmp-d-4-Data.db/[(0,15180462), (15180462,18656982)]\n Keyspace1: /var/lib/cassandra/data/Keyspace1/Standard1-tmp-d-5-Data.db/[(0,353139829), (353139829,433883659)]\n Keyspace1: /var/lib/cassandra/data/Keyspace1/Standard1-tmp-d-6-Data.db/[(0,366336059), (366336059,450095320)]\n\nnodetool --host 10.50.26.132 streams\nMode: Leaving: streaming data to other nodes\nStreaming to: /10.50.26.134\n /var/lib/cassandra/data/Keyspace1/Standard1-d-48-Data.db/[(0,366336059), (366336059,450095320)]\nNot receiving any streams.\n\nThese have been going for the past 2 hours.\n\nI see in the logs of the node with 134 IP address and I saw this:\n\nINFO [GOSSIP_STAGE:1] 2010-06-22 16:30:54,679 StorageService.java (line 603) Will not change my token ownership to /10.50.26.132\n\nSo, to my understanding from wikis loadbalance supposed to decommission and re-bootstrap again by sending its tokens to other nodes and then bootstrap again. It's been stuck in streaming for the past 2 hours and the size of ring has not changed. The log in the first node says it has started streaming for the past hours:\n\nINFO [STREAM-STAGE:1] 2010-06-22 16:35:56,255 StreamOut.java (line 72) Beginning transfer process to /10.50.26.134 for ranges (154293670372423273273390365393543806425,69164917636305877859094619660693892452]\n INFO [STREAM-STAGE:1] 2010-06-22 16:35:56,255 StreamOut.java (line 82) Flushing memtables for Keyspace1...\n INFO [STREAM-STAGE:1] 2010-06-22 16:35:56,266 StreamOut.java (line 128) Stream context metadata [/var/lib/cassandra/data/Keyspace1/Standard1-d-48-Data.db/[(0,366336059), (366336059,450095320)]] 1 sstables.\n INFO [STREAM-STAGE:1] 2010-06-22 16:35:56,267 StreamOut.java (line 135) Sending a stream initiate message to /10.50.26.134 ...\n INFO [STREAM-STAGE:1] 2010-06-22 16:35:56,267 StreamOut.java (line 140) Waiting for transfer to /10.50.26.134 to complete\n INFO [FLUSH-TIMER] 2010-06-22 17:36:53,370 ColumnFamilyStore.java (line 359) LocationInfo has reached its threshold; switching in a fresh Memtable at CommitLogContext(file='/var/lib/cassandra/commitlog/CommitLog-1277249454413.log', position=720)\n INFO [FLUSH-TIMER] 2010-06-22 17:36:53,370 ColumnFamilyStore.java (line 622) Enqueuing flush of Memtable(LocationInfo)@1637794189\n INFO [FLUSH-WRITER-POOL:1] 2010-06-22 17:36:53,370 Memtable.java (line 149) Writing Memtable(LocationInfo)@1637794189\n INFO [FLUSH-WRITER-POOL:1] 2010-06-22 17:36:53,528 Memtable.java (line 163) Completed flushing /var/lib/cassandra/data/system/LocationInfo-d-9-Data.db\n INFO [MEMTABLE-POST-FLUSHER:1] 2010-06-22 17:36:53,529 ColumnFamilyStore.java (line 374) Discarding 1000\n\n\nNothing more after this line.\n\nAm I doing something wrong?","issue_id":"12467688","key":"CASSANDRA-1221","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2010-07-20T17:12:36.000+0000","role":"fixed_distractor","summary":"loadbalance operation never completes on a 3 node cluster"} {"case_id":"12999303","cluster":"DISTRACTOR-CASSANDRA-12525","comments":[{"body":"[~asood]: sorry I'm a bit late asking this, but do you have the logs collected during this period (in particular, the logs from the 5 new nodes would be useful here).","created":"2016-08-31T08:17:07.330+0000"},{"body":"Sam, unfortunately don't have these handy ATM, but let me try to get these by next week. We were able to consistently replicate this and have the cluster handy in dev, so let me get you the logs as soon as I free up from my day to day work.\n\nAppreciate you looking into this.","created":"2016-09-01T02:22:30.966+0000"},{"body":"We were able to reliably reproduce this issue on C* 4.0.5\r\n\r\nOur steps:\r\n # Create DC 1 with say 3 nodes\r\n # Set a password for the cassandra user\r\n # Create other users, etc.\r\n # Add another DC DC2\r\n ## We observe in I think 3.11 cassandra logs (4.0 is similar)  for the first node brought up on the new DC:\r\n\r\n||source_file    || source_line ||message||\r\n|CassandraRoleManager.java |374  | Created default superuser role 'cassandra'|\r\n # Then we run a full repair on system_auth so other users propagate from DC1 to DC2\r\n # Result: cassandra user has original casssandra password\r\n\r\nWe also have seen this sporadically in C* 3.11.13\r\n\r\nI believe there is some race conditions where that first node believes it's the only node in the cluster and ignores the other DC Our expectation is that this code does not run for the second data center and if it needs to run, a simple fix might be to set the write_time to epoch 0 or something so a subsequent repair overwrites it.\r\n\r\n \r\n\r\nA workaround is deleting the system_auth on the node in question but this still exposes the node to being accessed with the default password for a short period of time\r\n\r\n \r\n\r\n ","created":"2022-11-15T21:30:31.659+0000"},{"body":"Linking German's PR that [~mck] pointed me to, thanks!\r\n\r\nI think this approach makes sense, effectively making the automatic creation idempotent. WDYT, [~samt]? (I think we should still check the 4.0.5 bug described though.)","created":"2022-11-29T21:39:59.740+0000"},{"body":"Seems like a sensible solution to me.","created":"2022-12-01T10:24:31.462+0000"},{"body":" I put some comments in PR.\r\n\r\n[~xgerman42] could you please move this to \"patch available\" so I can review it formally etc? We should follow Jira workflow here.","created":"2022-12-05T09:24:30.515+0000"},{"body":"I am sorry for all the mistakes. This is my first patch so still learning the process...","created":"2022-12-05T20:42:42.553+0000"},{"body":"[~xgerman42] I think we should do a test you described in the description of this ticket to be sure that it works. The test currently in place just checks if these timestamps are 0 / are same. It would be nice to have more \"real world\" test here if you do not mind. The test you did consists of 10 nodes together. I do not think that is necessary. What is the minimal amount of nodes you are able to reproduce this problem on? Do 2 nodes work as well?","created":"2022-12-06T15:18:12.104+0000"},{"body":"Hi again [~xgerman42] ,\r\n\r\nI wanted to replicate your issue locally but I was not able to do that. I used 3 nodes per dc in two dc's (6 nodes in total).\r\n\r\nI was also checking the logic when it comes to role creation (1) and (2). What it does is that it will create local cassandra role only in case it is nowhere to be found. First it checks with consistency level ONE, if it is not there, it will check with QUORUM CL. I put more debugging to that and every time it executed only query with ONE and it saw that cassandra role is there and it just moved on.\r\n\r\nThe only ever case when cassandra role was created was when I started the very first node in dc1. The fact that you see cassandra role creation in other node in the second dc is very interesting. I am not completely sure how you got that behavior.\r\n\r\nThere is also this (3), it will postpone the check of (1) for 10 seconds. So what could in theory happen is that the node did not see any peers in this window of 10 seconds which means that it evaluated that it is alone in the cluster (all conditions in (1) were evaluated as true (and false as they were negated)). This is the most reasonable explanation why you see this from time to time.\r\n\r\nThe interesting consequence of that logic in (3) is that it can not be blocking because that node about to start does not know in advance if it is going to be the only one in the cluster or not. The delay is controlled by system property (4) so you could pro-actively increase this to some higher value, like 30 seconds to minimize the chance that this might happen.\r\n\r\nTo improve this, we might make the default waiting time bigger but that is not solving it entirely, we are just kicking the can down the road here.\r\n\r\nSo if we can not completely prevent this, the next best option is to do what you suggested.\r\n\r\nIs not there any third, better way? What about waiting on _something_ in that scheduled tasked, postponed to 10s by default, to wait for something which is not initialized fully yet? The fact that these peers are not there yet seems like Gossip did not have a chance to see the topology fully yet? \r\n\r\n(1) https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/auth/CassandraRoleManager.java#L351-L356\r\n(2) https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/auth/CassandraRoleManager.java#L376-L384\r\n(3) https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/auth/CassandraRoleManager.java#L386-L405\r\n(4) https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/auth/AuthKeyspace.java#L58","created":"2022-12-08T17:03:28.644+0000"},{"body":"I don't think a third way precludes also adding the protection that creating the user in an idempotent manner grants. We can do both. ","created":"2022-12-08T17:06:38.988+0000"},{"body":"[~smiklosovic]  it's best practice after adding a second DC to run a (full) repair of the system_auth keyspace so with my patch this will effectively mitigate the issue for most people.\r\n\r\nAlso consider this scenario:\r\n1. We bring up DC 1 in network A and change the cassandra user's password\r\n\r\n2. We bring up DC 2 on network B - but due to some error the networks are not connected\r\n\r\n3. We discover our error and connect network A and B so the DCs can see each other\r\n\r\n \r\n\r\nIn this case DC2 would not know that there is another DC (someone might just have misconfigured seed nodes) until the connection is established.To help with this scenario (other than my patch and repair) we would need some command line parameter so an operator can skip the initial role generation...","created":"2022-12-08T17:42:49.549+0000"},{"body":"Ok lets to this then. I would not increase default delay for that system property. If we kept increasing this and other time that property, after a while all the \"waitings\" together would slow-downed the overall startup unbearably.\r\n\r\nFor your scenario, maybe I am completely wrong here, but I would expect that nodes in dc2 would have as a seed a node from dc1. So the startup should not even propagate that far to create the default role. It should not even start if it can not see any seeds.\r\n\r\nTo test this, we could do this then with 2 nodes only. Start the first, change the password, partition the network, start the second and it should create the default role and then start the first node and repair the second. You should be able to connect to the second node with the changed password.\r\n\r\nTo partition the network, there are already some examples how to drop the communication, start to take a look at this (1) and you eventually figure it out.\r\n\r\nPlease, tell me if it is too much for you to do this. We would be glad if you figured it out, if you dont we are here to help to guide you. We are trying to guide new contributors so they will be more comfortable here.\r\n\r\n(1) https://github.com/apache/cassandra/blob/trunk/test/distributed/org/apache/cassandra/distributed/shared/ClusterUtils.java#L124\r\n\r\n","created":"2022-12-08T19:50:03.487+0000"},{"body":"[~smiklosovic]  Sorry for the delayed response. I added the test with the two data centers.\r\n\r\nThe error condition is only hit when the new node is in the seed node list which changes the behavior of the shadow round - otherwise startup will fail there and not get to messing up the system_auth keyspace.\r\n\r\nAdding all potential seed nodes to casandra.yaml is common practice to avoid carefully managing the list for each node...","created":"2022-12-29T23:31:56.597+0000"},{"body":"Hi [~xgerman42] ,\r\n\r\nthanks for being so persistent!\r\n\r\nI have checked your latest changes and the test is not doing what I was mentioning in my last comment. My suggestion was:\r\n * start the first node (done)\r\n * change the password (not done)\r\n * partition the network (done)\r\n * start the second node (done)\r\n * check that it created the default role (not done)\r\n * \"unpartition\" the network (not done)\r\n * repair the second node (not done)\r\n * you should be able to connect to the second node with the changed password (not done)\r\n\r\nDoing CQL against a node can be done through Cassandra Java driver (logging in, changing password ...). Repairing of the node can be done via nodetool. There is \"nodetool\" method on IInstance you get from calling cluster.get like \"cluster.get(2).nodetool(\"repair\")\". You got the idea.\r\n\r\nDo you plan to finish this or do you have any other idea how how to test this differently? I humbly think the approach I outlined is the most comprehensive in order to mimic the real-world usage here.\r\n\r\nThanks","created":"2023-01-09T11:00:07.878+0000"},{"body":"[~xgerman42]\r\n\r\nI was thinking about this a little bit more lately and I will try to jump in to fill the gaps, if you do not mind. I have to admit that it might be little bit off-putting to jump through all these hurdles suddenly at once. Also, maybe we realize that, for some reason, the steps I suggested are not entirely correct (might happen, right!?) and we would just kill more time on this than necessary.\r\n","created":"2023-01-10T07:05:50.537+0000"},{"body":"[~smiklosovic]  I really learned a lot writing those tests (would have never imagined that by adding the node's ip to the seed node list the behavior of bootstrap would change) - so super helpful. Though eventually I like to do more substantial work currently I just don't have the time.  Especially, I feel for implementing the changes you suggest I would need to change the way the tests work - for instance `cluster.coordinator(1).execute` doesn't use an admin context so there is some refactoring required to make it roles aware and in my opinion this is the right way forward. But that likely requires refactoring and design work and is a bigger conversation than CAAS-12525.\r\n\r\nIn any case curious how you implement the changes and I am sure I can learn a lot by looking over your shoulder/studying the PR.","created":"2023-01-10T17:58:46.186+0000"},{"body":"[~xgerman42]\r\n\r\nplease find the patch here: https://github.com/apache/cassandra/pull/2085/files\r\n\r\nI took what you did and finished the test, also all changes are squashed into one commit. Please go through it and verify my logic is correct. We might then ask for the second review.","created":"2023-01-11T16:05:20.092+0000"},{"body":"[~smiklosovic] this looks great - especially the use of the cql client.","created":"2023-01-11T18:53:38.327+0000"},{"body":"LGTM, I think we just need to run CI.","created":"2023-01-11T19:09:28.443+0000"},{"body":"PR: https://github.com/apache/cassandra/pull/2085/files\r\n\r\nj8 pre-commit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1701/workflows/6adcd572-b3dd-4f9f-af86-4f20f78d36d8\r\nj11 pre-commit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1701/workflows/314b5982-03a2-464a-ad07-d697261847cd\r\n\r\nfailing tests in j8 are \r\n\r\nhttps://issues.apache.org/jira/browse/CASSANDRA-17708\r\nhttps://issues.apache.org/jira/browse/CASSANDRA-17819 (there is ongoing discussion about possible re-opening of this ticket as it started to fail for more people lately)\r\nfailing testForcedNormalRepairWithOneNodeDown is mentioned here https://issues.apache.org/jira/browse/CASSANDRA-14752 but it was evaluated as not-related which is probably the same here, it seems it does not have dedicated ticket yet, I may create it.\r\n\r\ntest was run repeatedly too","created":"2023-01-12T10:43:24.763+0000"},{"body":"ah wait a minute, we need to fix this from 3.0 up, right? ","created":"2023-01-12T10:50:59.024+0000"},{"body":"Sorry, I meant 'branch and run CI', if you are keen to fix it in all the branches.","created":"2023-01-12T11:58:02.935+0000"},{"body":"Yeah I'll backport that. It seems to be very easy patch.","created":"2023-01-12T11:59:09.578+0000"},{"body":"3.0 https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/2185/\r\n3.11 https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/2186/\r\n4.0 https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/2190/\r\n\r\n4.1\r\nj8 https://app.circleci.com/pipelines/github/instaclustr/cassandra/1714/workflows/aa3cd518-a827-43b4-b378-2b5ba31759c0\r\nj11 https://app.circleci.com/pipelines/github/instaclustr/cassandra/1714/workflows/bbb07956-0556-4154-bf7d-cde850bd2a1f \r\n\r\ntrunk\r\nj8 https://app.circleci.com/pipelines/github/instaclustr/cassandra/1701/workflows/6adcd572-b3dd-4f9f-af86-4f20f78d36d8\r\nj11 https://app.circleci.com/pipelines/github/instaclustr/cassandra/1701/workflows/314b5982-03a2-464a-ad07-d697261847cd\r\n\r\nhttps://github.com/instaclustr/cassandra/tree/CASSANDRA-12525-3.0\r\nhttps://github.com/instaclustr/cassandra/tree/CASSANDRA-12525-3.11\r\nhttps://github.com/instaclustr/cassandra/tree/CASSANDRA-12525-4.0\r\nhttps://github.com/instaclustr/cassandra/tree/CASSANDRA-12525-4.1\r\nhttps://github.com/instaclustr/cassandra/tree/CASSANDRA-12525","created":"2023-01-14T16:43:16.105+0000"},{"body":"[~brandon.williams] I am waiting patiently for your explicit +1. I think I run multiplexer only for trunk but I consider that to be enough as it is exactly same test.","created":"2023-01-17T08:34:49.322+0000"},{"body":"Thanks for your patience, yesterday was a holiday in the US. +1, I too am fine with a multiplexer run on a single branch.","created":"2023-01-17T13:36:09.872+0000"},{"body":"[~xgerman42] thanks for your contribution. Keep them coming!","created":"2023-01-18T14:30:33.370+0000"}],"conversations":[{"body":"Made the following observation:\n\nWhen adding new nodes to an existing C* cluster with authentication enabled we end up loosing password information about `cassandra` user. \n\nInitial Setup\n- Create a 5 node cluster with system_auth having RF=5 and NetworkTopologyStrategy\n- Enable PasswordAuthenticator on this cluster and update the password for 'cassandra' user to say 'password' via the alter query\n- Make sure you run nodetool repair on all the nodes\n\nTest case\n- Now go ahead and add 5 more nodes to this cluster.\n- Run nodetool repair on all the 10 nodes now\n- Decommission the original 5 nodes such that only the new 5 nodes are in the cluster now\n\n- Run cqlsh and try to connect to this cluster using old user name and password, cassandra/password\n\nI was unable to connect to the nodes with the original credentials and was only able to connect using the default cassandra/cassandra credentials\n\nFrom the conversation over IIRC\n\n`beobal: sood: that definitely shouldn't happen. The new nodes should only create the default superuser role if there are 0 roles currently defined (including that default one)`","from":"reporter","subject":"When adding new nodes to a cluster which has authentication enabled, we end up losing cassandra user's current crendentials and they get reverted back to default cassandra/cassandra credentials"},{"body":"[~asood]: sorry I'm a bit late asking this, but do you have the logs collected during this period (in particular, the logs from the 5 new nodes would be useful here).","from":"developer"},{"body":"Sam, unfortunately don't have these handy ATM, but let me try to get these by next week. We were able to consistently replicate this and have the cluster handy in dev, so let me get you the logs as soon as I free up from my day to day work.\n\nAppreciate you looking into this.","from":"developer"},{"body":"We were able to reliably reproduce this issue on C* 4.0.5\r\n\r\nOur steps:\r\n # Create DC 1 with say 3 nodes\r\n # Set a password for the cassandra user\r\n # Create other users, etc.\r\n # Add another DC DC2\r\n ## We observe in I think 3.11 cassandra logs (4.0 is similar)  for the first node brought up on the new DC:\r\n\r\n||source_file    || source_line ||message||\r\n|CassandraRoleManager.java |374  | Created default superuser role 'cassandra'|\r\n # Then we run a full repair on system_auth so other users propagate from DC1 to DC2\r\n # Result: cassandra user has original casssandra password\r\n\r\nWe also have seen this sporadically in C* 3.11.13\r\n\r\nI believe there is some race conditions where that first node believes it's the only node in the cluster and ignores the other DC Our expectation is that this code does not run for the second data center and if it needs to run, a simple fix might be to set the write_time to epoch 0 or something so a subsequent repair overwrites it.\r\n\r\n \r\n\r\nA workaround is deleting the system_auth on the node in question but this still exposes the node to being accessed with the default password for a short period of time\r\n\r\n \r\n\r\n ","from":"developer"},{"body":"Linking German's PR that [~mck] pointed me to, thanks!\r\n\r\nI think this approach makes sense, effectively making the automatic creation idempotent. WDYT, [~samt]? (I think we should still check the 4.0.5 bug described though.)","from":"developer"},{"body":"Seems like a sensible solution to me.","from":"developer"},{"body":" I put some comments in PR.\r\n\r\n[~xgerman42] could you please move this to \"patch available\" so I can review it formally etc? We should follow Jira workflow here.","from":"developer"},{"body":"I am sorry for all the mistakes. This is my first patch so still learning the process...","from":"developer"},{"body":"[~xgerman42] I think we should do a test you described in the description of this ticket to be sure that it works. The test currently in place just checks if these timestamps are 0 / are same. It would be nice to have more \"real world\" test here if you do not mind. The test you did consists of 10 nodes together. I do not think that is necessary. What is the minimal amount of nodes you are able to reproduce this problem on? Do 2 nodes work as well?","from":"developer"},{"body":"Hi again [~xgerman42] ,\r\n\r\nI wanted to replicate your issue locally but I was not able to do that. I used 3 nodes per dc in two dc's (6 nodes in total).\r\n\r\nI was also checking the logic when it comes to role creation (1) and (2). What it does is that it will create local cassandra role only in case it is nowhere to be found. First it checks with consistency level ONE, if it is not there, it will check with QUORUM CL. I put more debugging to that and every time it executed only query with ONE and it saw that cassandra role is there and it just moved on.\r\n\r\nThe only ever case when cassandra role was created was when I started the very first node in dc1. The fact that you see cassandra role creation in other node in the second dc is very interesting. I am not completely sure how you got that behavior.\r\n\r\nThere is also this (3), it will postpone the check of (1) for 10 seconds. So what could in theory happen is that the node did not see any peers in this window of 10 seconds which means that it evaluated that it is alone in the cluster (all conditions in (1) were evaluated as true (and false as they were negated)). This is the most reasonable explanation why you see this from time to time.\r\n\r\nThe interesting consequence of that logic in (3) is that it can not be blocking because that node about to start does not know in advance if it is going to be the only one in the cluster or not. The delay is controlled by system property (4) so you could pro-actively increase this to some higher value, like 30 seconds to minimize the chance that this might happen.\r\n\r\nTo improve this, we might make the default waiting time bigger but that is not solving it entirely, we are just kicking the can down the road here.\r\n\r\nSo if we can not completely prevent this, the next best option is to do what you suggested.\r\n\r\nIs not there any third, better way? What about waiting on _something_ in that scheduled tasked, postponed to 10s by default, to wait for something which is not initialized fully yet? The fact that these peers are not there yet seems like Gossip did not have a chance to see the topology fully yet? \r\n\r\n(1) https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/auth/CassandraRoleManager.java#L351-L356\r\n(2) https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/auth/CassandraRoleManager.java#L376-L384\r\n(3) https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/auth/CassandraRoleManager.java#L386-L405\r\n(4) https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/auth/AuthKeyspace.java#L58","from":"developer"},{"body":"I don't think a third way precludes also adding the protection that creating the user in an idempotent manner grants. We can do both. ","from":"developer"},{"body":"[~smiklosovic]  it's best practice after adding a second DC to run a (full) repair of the system_auth keyspace so with my patch this will effectively mitigate the issue for most people.\r\n\r\nAlso consider this scenario:\r\n1. We bring up DC 1 in network A and change the cassandra user's password\r\n\r\n2. We bring up DC 2 on network B - but due to some error the networks are not connected\r\n\r\n3. We discover our error and connect network A and B so the DCs can see each other\r\n\r\n \r\n\r\nIn this case DC2 would not know that there is another DC (someone might just have misconfigured seed nodes) until the connection is established.To help with this scenario (other than my patch and repair) we would need some command line parameter so an operator can skip the initial role generation...","from":"developer"},{"body":"Ok lets to this then. I would not increase default delay for that system property. If we kept increasing this and other time that property, after a while all the \"waitings\" together would slow-downed the overall startup unbearably.\r\n\r\nFor your scenario, maybe I am completely wrong here, but I would expect that nodes in dc2 would have as a seed a node from dc1. So the startup should not even propagate that far to create the default role. It should not even start if it can not see any seeds.\r\n\r\nTo test this, we could do this then with 2 nodes only. Start the first, change the password, partition the network, start the second and it should create the default role and then start the first node and repair the second. You should be able to connect to the second node with the changed password.\r\n\r\nTo partition the network, there are already some examples how to drop the communication, start to take a look at this (1) and you eventually figure it out.\r\n\r\nPlease, tell me if it is too much for you to do this. We would be glad if you figured it out, if you dont we are here to help to guide you. We are trying to guide new contributors so they will be more comfortable here.\r\n\r\n(1) https://github.com/apache/cassandra/blob/trunk/test/distributed/org/apache/cassandra/distributed/shared/ClusterUtils.java#L124\r\n\r\n","from":"developer"},{"body":"[~smiklosovic]  Sorry for the delayed response. I added the test with the two data centers.\r\n\r\nThe error condition is only hit when the new node is in the seed node list which changes the behavior of the shadow round - otherwise startup will fail there and not get to messing up the system_auth keyspace.\r\n\r\nAdding all potential seed nodes to casandra.yaml is common practice to avoid carefully managing the list for each node...","from":"developer"},{"body":"Hi [~xgerman42] ,\r\n\r\nthanks for being so persistent!\r\n\r\nI have checked your latest changes and the test is not doing what I was mentioning in my last comment. My suggestion was:\r\n * start the first node (done)\r\n * change the password (not done)\r\n * partition the network (done)\r\n * start the second node (done)\r\n * check that it created the default role (not done)\r\n * \"unpartition\" the network (not done)\r\n * repair the second node (not done)\r\n * you should be able to connect to the second node with the changed password (not done)\r\n\r\nDoing CQL against a node can be done through Cassandra Java driver (logging in, changing password ...). Repairing of the node can be done via nodetool. There is \"nodetool\" method on IInstance you get from calling cluster.get like \"cluster.get(2).nodetool(\"repair\")\". You got the idea.\r\n\r\nDo you plan to finish this or do you have any other idea how how to test this differently? I humbly think the approach I outlined is the most comprehensive in order to mimic the real-world usage here.\r\n\r\nThanks","from":"developer"},{"body":"[~xgerman42]\r\n\r\nI was thinking about this a little bit more lately and I will try to jump in to fill the gaps, if you do not mind. I have to admit that it might be little bit off-putting to jump through all these hurdles suddenly at once. Also, maybe we realize that, for some reason, the steps I suggested are not entirely correct (might happen, right!?) and we would just kill more time on this than necessary.\r\n","from":"developer"},{"body":"[~smiklosovic]  I really learned a lot writing those tests (would have never imagined that by adding the node's ip to the seed node list the behavior of bootstrap would change) - so super helpful. Though eventually I like to do more substantial work currently I just don't have the time.  Especially, I feel for implementing the changes you suggest I would need to change the way the tests work - for instance `cluster.coordinator(1).execute` doesn't use an admin context so there is some refactoring required to make it roles aware and in my opinion this is the right way forward. But that likely requires refactoring and design work and is a bigger conversation than CAAS-12525.\r\n\r\nIn any case curious how you implement the changes and I am sure I can learn a lot by looking over your shoulder/studying the PR.","from":"developer"},{"body":"[~xgerman42]\r\n\r\nplease find the patch here: https://github.com/apache/cassandra/pull/2085/files\r\n\r\nI took what you did and finished the test, also all changes are squashed into one commit. Please go through it and verify my logic is correct. We might then ask for the second review.","from":"developer"},{"body":"[~smiklosovic] this looks great - especially the use of the cql client.","from":"developer"},{"body":"LGTM, I think we just need to run CI.","from":"developer"},{"body":"PR: https://github.com/apache/cassandra/pull/2085/files\r\n\r\nj8 pre-commit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1701/workflows/6adcd572-b3dd-4f9f-af86-4f20f78d36d8\r\nj11 pre-commit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1701/workflows/314b5982-03a2-464a-ad07-d697261847cd\r\n\r\nfailing tests in j8 are \r\n\r\nhttps://issues.apache.org/jira/browse/CASSANDRA-17708\r\nhttps://issues.apache.org/jira/browse/CASSANDRA-17819 (there is ongoing discussion about possible re-opening of this ticket as it started to fail for more people lately)\r\nfailing testForcedNormalRepairWithOneNodeDown is mentioned here https://issues.apache.org/jira/browse/CASSANDRA-14752 but it was evaluated as not-related which is probably the same here, it seems it does not have dedicated ticket yet, I may create it.\r\n\r\ntest was run repeatedly too","from":"developer"},{"body":"ah wait a minute, we need to fix this from 3.0 up, right? ","from":"developer"},{"body":"Sorry, I meant 'branch and run CI', if you are keen to fix it in all the branches.","from":"developer"},{"body":"Yeah I'll backport that. It seems to be very easy patch.","from":"developer"},{"body":"3.0 https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/2185/\r\n3.11 https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/2186/\r\n4.0 https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/2190/\r\n\r\n4.1\r\nj8 https://app.circleci.com/pipelines/github/instaclustr/cassandra/1714/workflows/aa3cd518-a827-43b4-b378-2b5ba31759c0\r\nj11 https://app.circleci.com/pipelines/github/instaclustr/cassandra/1714/workflows/bbb07956-0556-4154-bf7d-cde850bd2a1f \r\n\r\ntrunk\r\nj8 https://app.circleci.com/pipelines/github/instaclustr/cassandra/1701/workflows/6adcd572-b3dd-4f9f-af86-4f20f78d36d8\r\nj11 https://app.circleci.com/pipelines/github/instaclustr/cassandra/1701/workflows/314b5982-03a2-464a-ad07-d697261847cd\r\n\r\nhttps://github.com/instaclustr/cassandra/tree/CASSANDRA-12525-3.0\r\nhttps://github.com/instaclustr/cassandra/tree/CASSANDRA-12525-3.11\r\nhttps://github.com/instaclustr/cassandra/tree/CASSANDRA-12525-4.0\r\nhttps://github.com/instaclustr/cassandra/tree/CASSANDRA-12525-4.1\r\nhttps://github.com/instaclustr/cassandra/tree/CASSANDRA-12525","from":"developer"},{"body":"[~brandon.williams] I am waiting patiently for your explicit +1. I think I run multiplexer only for trunk but I consider that to be enough as it is exactly same test.","from":"developer"},{"body":"Thanks for your patience, yesterday was a holiday in the US. +1, I too am fine with a multiplexer run on a single branch.","from":"developer"},{"body":"[~xgerman42] thanks for your contribution. Keep them coming!","from":"developer"}],"created":"2016-08-23T17:32:10.000+0000","description":"Made the following observation:\n\nWhen adding new nodes to an existing C* cluster with authentication enabled we end up loosing password information about `cassandra` user. \n\nInitial Setup\n- Create a 5 node cluster with system_auth having RF=5 and NetworkTopologyStrategy\n- Enable PasswordAuthenticator on this cluster and update the password for 'cassandra' user to say 'password' via the alter query\n- Make sure you run nodetool repair on all the nodes\n\nTest case\n- Now go ahead and add 5 more nodes to this cluster.\n- Run nodetool repair on all the 10 nodes now\n- Decommission the original 5 nodes such that only the new 5 nodes are in the cluster now\n\n- Run cqlsh and try to connect to this cluster using old user name and password, cassandra/password\n\nI was unable to connect to the nodes with the original credentials and was only able to connect using the default cassandra/cassandra credentials\n\nFrom the conversation over IIRC\n\n`beobal: sood: that definitely shouldn't happen. The new nodes should only create the default superuser role if there are 0 roles currently defined (including that default one)`","issue_id":"12999303","key":"CASSANDRA-12525","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2023-01-18T14:29:53.000+0000","role":"fixed_distractor","summary":"When adding new nodes to a cluster which has authentication enabled, we end up losing cassandra user's current crendentials and they get reverted back to default cassandra/cassandra credentials"} {"case_id":"13001795","cluster":"DISTRACTOR-CASSANDRA-12582","comments":[{"body":"If there is anything else needed to reproduce or understand this issue, please let us know. \n\nWe're eager to help however we can to try to get this fixed for 3.0.9 to complete all the fixes related to issues with static columns in this issue, CASSANDRA-11988 and CASSANDRA-12336.","created":"2016-09-07T21:41:42.454+0000"},{"body":"[~tjake] I don't want to speak out of turn, but it seems to me that this not getting attention in 3.0.9 will leave that build without a complete fix to the known issues dealing with static columns (other related issues linked above). Is there any way this can get attention in 3.0.9?","created":"2016-09-20T13:54:48.544+0000"},{"body":"Just saw that 3.0.9 was released, so i guess ignore that question. Bummer that this didn't make it in. Is there a way to know when versions are planning on being released? Seems like theres a level of organizational conversation going on that I'm not aware of.","created":"2016-09-20T14:00:51.411+0000"},{"body":"[~eprothro] The versions get bumped when I do a release. We periodically do releases and 3.0.9 was long overdue (See changelog). I'm not sure who would be best to look into this issue but I'll ask if someone can help. thanks for the repro script! ","created":"2016-09-20T14:12:18.256+0000"},{"body":" I'll have a go at reproducing this and see if I can understand what is going on. Thanks for providing such detailed information.","created":"2016-09-22T08:09:36.277+0000"},{"body":"Reproduced without problems.","created":"2016-09-22T08:14:42.700+0000"},{"body":"The reason of the corruption is that the sstable iterators try to read a static column as a regular column. This is due to the fact that the serialization header is missing the dropped static column.\n\nWhen the header is deserialized, it relies on {{CFMetadata}} to provide a fake dropped column. The problem is that {{CFMetadata}} always assumes that dropped columns are regular, see [here|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/config/CFMetaData.java#L682]. Then when the iterators read the static column, they rely on the header to make the decision of whether a static column is present or not, [here|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/columniterator/AbstractSSTableIterator.java#L169]. As a consequence, they don't attempt to skip the static column and they try to read it as a regular column later on.","created":"2016-09-22T09:28:56.882+0000"},{"body":"This patch should fix the problem:\n\n||3.0||trunk||\n|[patch|https://github.com/stef1927/cassandra/commits/12582-3.0]|[patch|https://github.com/stef1927/cassandra/commits/12582]|\n|[testall|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-testall/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-dtest/]|\n\nLuckily the serialization header stores static and regular columns separately, so I was able to pass this information to {{CFMetadata}} in order to create a fake static column but I wonder if we may have issues in other parts of the code, where we create a regular column def instead of a static one. Basically all callers of [{{CFMetadata.getDroppedColumnDefinition}}|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/config/CFMetaData.java#L672] are at risk. So I wonder if we should also add {{ColumnDefinition.Kind}} to {{SystemKeyspace.DroppedColumns}}. Currently only regular or static columns can be dropped, so technically we could also only add a boolean.\n\nWDYT [~iamaleksey]?\n","created":"2016-09-23T06:07:26.164+0000"},{"body":"bq. I wonder if we should also add {{ColumnDefinition.Kind}} to {{SystemKeyspace.DroppedColumns}}\n\nWe definitively should imo. Though that imply a schema change and that's kind of problematic at the moment until 4.0, so we might have to do so in a separate ticket and rely on work-around like you implemented until then.\n\nbq. Currently only regular or static columns can be dropped, so technically we could also only add a boolean\n\nI can't really see us ever allowing the dropping of partition key or clustering columns (it sounds way too painful to support in the storage engine to be worth it), but I don't think we have much to lose by saving the full kind so it's probably fine to be overly cautious.","created":"2016-09-23T07:35:07.166+0000"},{"body":"It's worth noting that the exception I reproduced is actually different that the one in the ticket description. It seems the exception in the description is failing to read a static row whilst the exception I reproduced is failing to read a regular row (a clustering to be exact), I hope the workaround in the serialization header covers both, but I cannot reproduce the exact exception in the description. \n\n{code}\njava.lang.RuntimeException: org.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /home/stefi/git/cstar/cassandra/data/data/issue12582/apples_by_tree-3980ba90809c11e6b62e6be1d44ebd9b/mc-1-big-Data.db\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2470) ~[main/:na]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_101]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[main/:na]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [main/:na]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_101]\nCaused by: org.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /home/stefi/git/cstar/cassandra/data/data/issue12582/apples_by_tree-3980ba90809c11e6b62e6be1d44ebd9b/mc-1-big-Data.db\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator$Reader.hasNext(AbstractSSTableIterator.java:353) ~[main/:na]\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator.hasNext(AbstractSSTableIterator.java:219) ~[main/:na]\n at org.apache.cassandra.db.columniterator.SSTableIterator.hasNext(SSTableIterator.java:32) ~[main/:na]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:95) ~[main/:na]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:32) ~[main/:na]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:95) ~[main/:na]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:32) ~[main/:na]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n at org.apache.cassandra.db.transform.BaseRows.hasNext(BaseRows.java:129) ~[main/:na]\n at org.apache.cassandra.db.transform.UnfilteredRows.isEmpty(UnfilteredRows.java:58) ~[main/:na]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:65) ~[main/:na]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:24) ~[main/:na]\n at org.apache.cassandra.db.transform.BasePartitions.hasNext(BasePartitions.java:96) ~[main/:na]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:295) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:145) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:138) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:134) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:76) ~[main/:na]\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:320) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1796) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2466) ~[main/:na]\n ... 5 common frames omitted\nCaused by: java.io.IOException: Corrupt flags value for clustering prefix (isStatic flag set): 160\n at org.apache.cassandra.db.ClusteringPrefix$Deserializer.prepare(ClusteringPrefix.java:422) ~[main/:na]\n at org.apache.cassandra.db.UnfilteredDeserializer$CurrentDeserializer.prepareNext(UnfilteredDeserializer.java:172) ~[main/:na]\n at org.apache.cassandra.db.UnfilteredDeserializer$CurrentDeserializer.hasNext(UnfilteredDeserializer.java:153) ~[main/:na]\n at org.apache.cassandra.db.columniterator.SSTableIterator$ForwardReader.computeNext(SSTableIterator.java:126) ~[main/:na]\n at org.apache.cassandra.db.columniterator.SSTableIterator$ForwardReader.hasNextInternal(SSTableIterator.java:153) ~[main/:na]\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator$Reader.hasNext(AbstractSSTableIterator.java:340) ~[main/:na]\n ... 26 common frames omitted\n{code}","created":"2016-09-23T08:12:43.185+0000"},{"body":"Thank you for your comment [~slebresne]. Do we need to wait until 4.0? If we add a field to {{SystemKeyspace.DroppedColumns}}, and we are careful to read it as an optional field that may be missing, given that the table has local replication, isn't this enough to cover all columns that will be dropped from this point onwards? For columns that have already been dropped, there's nothing we can do, even with a full schema migration we have lost the information of whether the column was static or not. Am I missing something?\n\nNoted regarding storing the full kind, and I agree.","created":"2016-09-23T08:48:51.038+0000"},{"body":"bq. If we add a field to SystemKeyspace.DroppedColumns, and we are careful to read it as an optional field that may be missing, given that the table has local replication\n\nDon't get fooled by the keyspace replication strategy, schema tables *are* replicated, just by their own manual mechanism (see {{MigrationManager}} for the gory details). Adding a column to a schema table exposes us to the same problem than in CASSANDRA-12236 (and it's follow-up, CASSANDRA-12697). And while for the {{cdc}} column we have been able to use the {{cdc_enabled}} to \"solve\" this problem (basically saying, \"keep cdc_enabled to false until fully upgraded\"), we don't have such trick here.\n\nDon't get me wrong, we could probably come up with something along the line of CASSANDRA-12236 if we really wanted to, but there is a fair chance it won't be very user friendly, and it's a tad involved anyway. So let's say that it's a lot easier to make the change in 4.0, and that even if we think it's important enough to try it before that, we might still want to defer that change to a follow-up ticket just so we get your work-around in in the meantime.","created":"2016-09-23T10:06:37.423+0000"},{"body":"[~Stefania] confirmed that your patch fixes the issue for me in both the reproduction script, as well as for our real-world use case. Thank you!","created":"2016-09-23T16:23:44.875+0000"},{"body":"Thanks for the explanation [~slebresne], I created CASSANDRA-12705 for adding a field to SystemKeyspace.DroppedColumns. I totally missed that the schema mutations are pushed to other nodes and would therefore cause exceptions during a rolling upgrade.\n\nI've checked the CI results and they are all clean except for testall on trunk, but the failures there also occur on trunk. Do you think you could be the reviewer since you've already looked at the change anyway?\n\nI also converted the reproduction script into a dtest, pull request [here|https://github.com/riptano/cassandra-dtest/pull/1342].\n\n--\n\n[~eprothro]: it's great to hear that the patch fixes your real-world use case as well, thank you so much for this update and, once again, for the reproduction script.","created":"2016-09-26T02:10:59.056+0000"},{"body":"Rebased on the latest branches, the 3.0 patch applies cleanly upwards:\n\n||3.0||3.X||trunk||\n|[patch|https://github.com/stef1927/cassandra/commits/12582-3.0]|[patch|https://github.com/stef1927/cassandra/commits/12582-3.X]|[patch|https://github.com/stef1927/cassandra/commits/12582]|\n|[testall|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.X-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-testall/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.X-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-dtest/]|\n\nSince there were no conflicts when rebasing, I did not relaunch the CI jobs which were clean.","created":"2016-10-07T04:48:36.003+0000"},{"body":"+1","created":"2016-10-07T15:45:58.418+0000"},{"body":"Thanks for the review, committed to 3.0 as 2f0e365dc495336332e174b30e2d15e56fc344e2 and merged into 3.X and trunk.","created":"2016-10-10T01:56:36.839+0000"}],"conversations":[{"body":"We ran into an issue on production where reads began to fail for certain queries, depending on the range within the relation for those queries. Cassandra system log showed an unhandled {{CorruptSSTableException}} exception.\n\nCQL read failure:\n{code}\nReadFailure: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 1 failures\" info={'failures': 1, 'received_responses': 0, 'required_responses': 1, 'consistency': 'ONE'}\n{code}\n\nCassandra exception:\n{code}\nWARN [SharedPool-Worker-2] 2016-08-31 12:49:27,979 AbstractLocalAwareExecutorService.java:169 - Uncaught exception on thread Thread[SharedPool-Worker-2,5,main]: {}\njava.lang.RuntimeException: org.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /usr/local/apache-cassandra-3.0.8/data/data/issue309/apples_by_tree-006748a06fa311e6a7f8ef8b642e977b/mb-1-big-Data.db\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2453) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_72]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [apache-cassandra-3.0.8.jar:3.0.8]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_72]\nCaused by: org.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /usr/local/apache-cassandra-3.0.8/data/data/issue309/apples_by_tree-006748a06fa311e6a7f8ef8b642e977b/mb-1-big-Data.db\n at org.apache.cassandra.io.sstable.format.big.BigTableScanner$KeyScanningIterator$1.initializeIterator(BigTableScanner.java:343) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.maybeInit(LazilyInitializedUnfilteredRowIterator.java:48) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.isReverseOrder(LazilyInitializedUnfilteredRowIterator.java:65) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.isReverseOrder(LazilyInitializedUnfilteredRowIterator.java:66) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:62) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:24) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.transform.BasePartitions.hasNext(BasePartitions.java:96) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:295) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:134) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:127) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:123) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:65) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:289) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1796) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2449) ~[apache-cassandra-3.0.8.jar:3.0.8]\n ... 5 common frames omitted\nCaused by: org.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /usr/local/apache-cassandra-3.0.8/data/data/issue309/apples_by_tree-006748a06fa311e6a7f8ef8b642e977b/mb-1-big-Data.db\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator.(AbstractSSTableIterator.java:130) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.columniterator.SSTableIterator.(SSTableIterator.java:46) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.io.sstable.format.big.BigTableReader.iterator(BigTableReader.java:69) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.io.sstable.format.big.BigTableScanner$KeyScanningIterator$1.initializeIterator(BigTableScanner.java:338) ~[apache-cassandra-3.0.8.jar:3.0.8]\n ... 19 common frames omitted\nCaused by: java.io.IOException: Corrupt (negative) value length encountered\n at org.apache.cassandra.db.marshal.AbstractType.readValue(AbstractType.java:399) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.BufferCell$Serializer.deserialize(BufferCell.java:302) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.readSimpleColumn(UnfilteredSerializer.java:462) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.deserializeRowBody(UnfilteredSerializer.java:440) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.deserializeStaticRow(UnfilteredSerializer.java:381) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator.readStaticRow(AbstractSSTableIterator.java:179) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator.(AbstractSSTableIterator.java:103) ~[apache-cassandra-3.0.8.jar:3.0.8]\n ... 22 common frames omitted\n{code}\n\nAfter debugging, it appears that a previously dropped static column (weeks prior) was the instigator of the issue. As a workaround we added back the column, restarted all cassandra processes within the cluster, and the read error and corruption exception went away.\n\nAttached is a script to reproduce with a simple schema.\n\nAlso noteworthy (and shown in the script) is that when in this state, compaction silently failed (exit 0) to remove the dropped static columns from the \"corrupted\" sstable.\n","from":"reporter","subject":"Removing static column results in ReadFailure due to CorruptSSTableException"},{"body":"If there is anything else needed to reproduce or understand this issue, please let us know. \n\nWe're eager to help however we can to try to get this fixed for 3.0.9 to complete all the fixes related to issues with static columns in this issue, CASSANDRA-11988 and CASSANDRA-12336.","from":"developer"},{"body":"[~tjake] I don't want to speak out of turn, but it seems to me that this not getting attention in 3.0.9 will leave that build without a complete fix to the known issues dealing with static columns (other related issues linked above). Is there any way this can get attention in 3.0.9?","from":"developer"},{"body":"Just saw that 3.0.9 was released, so i guess ignore that question. Bummer that this didn't make it in. Is there a way to know when versions are planning on being released? Seems like theres a level of organizational conversation going on that I'm not aware of.","from":"developer"},{"body":"[~eprothro] The versions get bumped when I do a release. We periodically do releases and 3.0.9 was long overdue (See changelog). I'm not sure who would be best to look into this issue but I'll ask if someone can help. thanks for the repro script! ","from":"developer"},{"body":" I'll have a go at reproducing this and see if I can understand what is going on. Thanks for providing such detailed information.","from":"developer"},{"body":"Reproduced without problems.","from":"developer"},{"body":"The reason of the corruption is that the sstable iterators try to read a static column as a regular column. This is due to the fact that the serialization header is missing the dropped static column.\n\nWhen the header is deserialized, it relies on {{CFMetadata}} to provide a fake dropped column. The problem is that {{CFMetadata}} always assumes that dropped columns are regular, see [here|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/config/CFMetaData.java#L682]. Then when the iterators read the static column, they rely on the header to make the decision of whether a static column is present or not, [here|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/columniterator/AbstractSSTableIterator.java#L169]. As a consequence, they don't attempt to skip the static column and they try to read it as a regular column later on.","from":"developer"},{"body":"This patch should fix the problem:\n\n||3.0||trunk||\n|[patch|https://github.com/stef1927/cassandra/commits/12582-3.0]|[patch|https://github.com/stef1927/cassandra/commits/12582]|\n|[testall|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-testall/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-dtest/]|\n\nLuckily the serialization header stores static and regular columns separately, so I was able to pass this information to {{CFMetadata}} in order to create a fake static column but I wonder if we may have issues in other parts of the code, where we create a regular column def instead of a static one. Basically all callers of [{{CFMetadata.getDroppedColumnDefinition}}|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/config/CFMetaData.java#L672] are at risk. So I wonder if we should also add {{ColumnDefinition.Kind}} to {{SystemKeyspace.DroppedColumns}}. Currently only regular or static columns can be dropped, so technically we could also only add a boolean.\n\nWDYT [~iamaleksey]?\n","from":"developer"},{"body":"bq. I wonder if we should also add {{ColumnDefinition.Kind}} to {{SystemKeyspace.DroppedColumns}}\n\nWe definitively should imo. Though that imply a schema change and that's kind of problematic at the moment until 4.0, so we might have to do so in a separate ticket and rely on work-around like you implemented until then.\n\nbq. Currently only regular or static columns can be dropped, so technically we could also only add a boolean\n\nI can't really see us ever allowing the dropping of partition key or clustering columns (it sounds way too painful to support in the storage engine to be worth it), but I don't think we have much to lose by saving the full kind so it's probably fine to be overly cautious.","from":"developer"},{"body":"It's worth noting that the exception I reproduced is actually different that the one in the ticket description. It seems the exception in the description is failing to read a static row whilst the exception I reproduced is failing to read a regular row (a clustering to be exact), I hope the workaround in the serialization header covers both, but I cannot reproduce the exact exception in the description. \n\n{code}\njava.lang.RuntimeException: org.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /home/stefi/git/cstar/cassandra/data/data/issue12582/apples_by_tree-3980ba90809c11e6b62e6be1d44ebd9b/mc-1-big-Data.db\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2470) ~[main/:na]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_101]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[main/:na]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [main/:na]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_101]\nCaused by: org.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /home/stefi/git/cstar/cassandra/data/data/issue12582/apples_by_tree-3980ba90809c11e6b62e6be1d44ebd9b/mc-1-big-Data.db\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator$Reader.hasNext(AbstractSSTableIterator.java:353) ~[main/:na]\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator.hasNext(AbstractSSTableIterator.java:219) ~[main/:na]\n at org.apache.cassandra.db.columniterator.SSTableIterator.hasNext(SSTableIterator.java:32) ~[main/:na]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:95) ~[main/:na]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:32) ~[main/:na]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:95) ~[main/:na]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.computeNext(LazilyInitializedUnfilteredRowIterator.java:32) ~[main/:na]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n at org.apache.cassandra.db.transform.BaseRows.hasNext(BaseRows.java:129) ~[main/:na]\n at org.apache.cassandra.db.transform.UnfilteredRows.isEmpty(UnfilteredRows.java:58) ~[main/:na]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:65) ~[main/:na]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:24) ~[main/:na]\n at org.apache.cassandra.db.transform.BasePartitions.hasNext(BasePartitions.java:96) ~[main/:na]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:295) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:145) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:138) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:134) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:76) ~[main/:na]\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:320) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1796) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2466) ~[main/:na]\n ... 5 common frames omitted\nCaused by: java.io.IOException: Corrupt flags value for clustering prefix (isStatic flag set): 160\n at org.apache.cassandra.db.ClusteringPrefix$Deserializer.prepare(ClusteringPrefix.java:422) ~[main/:na]\n at org.apache.cassandra.db.UnfilteredDeserializer$CurrentDeserializer.prepareNext(UnfilteredDeserializer.java:172) ~[main/:na]\n at org.apache.cassandra.db.UnfilteredDeserializer$CurrentDeserializer.hasNext(UnfilteredDeserializer.java:153) ~[main/:na]\n at org.apache.cassandra.db.columniterator.SSTableIterator$ForwardReader.computeNext(SSTableIterator.java:126) ~[main/:na]\n at org.apache.cassandra.db.columniterator.SSTableIterator$ForwardReader.hasNextInternal(SSTableIterator.java:153) ~[main/:na]\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator$Reader.hasNext(AbstractSSTableIterator.java:340) ~[main/:na]\n ... 26 common frames omitted\n{code}","from":"developer"},{"body":"Thank you for your comment [~slebresne]. Do we need to wait until 4.0? If we add a field to {{SystemKeyspace.DroppedColumns}}, and we are careful to read it as an optional field that may be missing, given that the table has local replication, isn't this enough to cover all columns that will be dropped from this point onwards? For columns that have already been dropped, there's nothing we can do, even with a full schema migration we have lost the information of whether the column was static or not. Am I missing something?\n\nNoted regarding storing the full kind, and I agree.","from":"developer"},{"body":"bq. If we add a field to SystemKeyspace.DroppedColumns, and we are careful to read it as an optional field that may be missing, given that the table has local replication\n\nDon't get fooled by the keyspace replication strategy, schema tables *are* replicated, just by their own manual mechanism (see {{MigrationManager}} for the gory details). Adding a column to a schema table exposes us to the same problem than in CASSANDRA-12236 (and it's follow-up, CASSANDRA-12697). And while for the {{cdc}} column we have been able to use the {{cdc_enabled}} to \"solve\" this problem (basically saying, \"keep cdc_enabled to false until fully upgraded\"), we don't have such trick here.\n\nDon't get me wrong, we could probably come up with something along the line of CASSANDRA-12236 if we really wanted to, but there is a fair chance it won't be very user friendly, and it's a tad involved anyway. So let's say that it's a lot easier to make the change in 4.0, and that even if we think it's important enough to try it before that, we might still want to defer that change to a follow-up ticket just so we get your work-around in in the meantime.","from":"developer"},{"body":"[~Stefania] confirmed that your patch fixes the issue for me in both the reproduction script, as well as for our real-world use case. Thank you!","from":"developer"},{"body":"Thanks for the explanation [~slebresne], I created CASSANDRA-12705 for adding a field to SystemKeyspace.DroppedColumns. I totally missed that the schema mutations are pushed to other nodes and would therefore cause exceptions during a rolling upgrade.\n\nI've checked the CI results and they are all clean except for testall on trunk, but the failures there also occur on trunk. Do you think you could be the reviewer since you've already looked at the change anyway?\n\nI also converted the reproduction script into a dtest, pull request [here|https://github.com/riptano/cassandra-dtest/pull/1342].\n\n--\n\n[~eprothro]: it's great to hear that the patch fixes your real-world use case as well, thank you so much for this update and, once again, for the reproduction script.","from":"developer"},{"body":"Rebased on the latest branches, the 3.0 patch applies cleanly upwards:\n\n||3.0||3.X||trunk||\n|[patch|https://github.com/stef1927/cassandra/commits/12582-3.0]|[patch|https://github.com/stef1927/cassandra/commits/12582-3.X]|[patch|https://github.com/stef1927/cassandra/commits/12582]|\n|[testall|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.X-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-testall/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-3.X-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/stef1927/job/stef1927-12582-dtest/]|\n\nSince there were no conflicts when rebasing, I did not relaunch the CI jobs which were clean.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Thanks for the review, committed to 3.0 as 2f0e365dc495336332e174b30e2d15e56fc344e2 and merged into 3.X and trunk.","from":"developer"}],"created":"2016-08-31T19:12:10.000+0000","description":"We ran into an issue on production where reads began to fail for certain queries, depending on the range within the relation for those queries. Cassandra system log showed an unhandled {{CorruptSSTableException}} exception.\n\nCQL read failure:\n{code}\nReadFailure: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 1 failures\" info={'failures': 1, 'received_responses': 0, 'required_responses': 1, 'consistency': 'ONE'}\n{code}\n\nCassandra exception:\n{code}\nWARN [SharedPool-Worker-2] 2016-08-31 12:49:27,979 AbstractLocalAwareExecutorService.java:169 - Uncaught exception on thread Thread[SharedPool-Worker-2,5,main]: {}\njava.lang.RuntimeException: org.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /usr/local/apache-cassandra-3.0.8/data/data/issue309/apples_by_tree-006748a06fa311e6a7f8ef8b642e977b/mb-1-big-Data.db\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2453) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_72]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [apache-cassandra-3.0.8.jar:3.0.8]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_72]\nCaused by: org.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /usr/local/apache-cassandra-3.0.8/data/data/issue309/apples_by_tree-006748a06fa311e6a7f8ef8b642e977b/mb-1-big-Data.db\n at org.apache.cassandra.io.sstable.format.big.BigTableScanner$KeyScanningIterator$1.initializeIterator(BigTableScanner.java:343) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.maybeInit(LazilyInitializedUnfilteredRowIterator.java:48) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.isReverseOrder(LazilyInitializedUnfilteredRowIterator.java:65) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.LazilyInitializedUnfilteredRowIterator.isReverseOrder(LazilyInitializedUnfilteredRowIterator.java:66) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:62) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.partitions.PurgeFunction.applyToPartition(PurgeFunction.java:24) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.transform.BasePartitions.hasNext(BasePartitions.java:96) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:295) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:134) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:127) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:123) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:65) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:289) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1796) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2449) ~[apache-cassandra-3.0.8.jar:3.0.8]\n ... 5 common frames omitted\nCaused by: org.apache.cassandra.io.sstable.CorruptSSTableException: Corrupted: /usr/local/apache-cassandra-3.0.8/data/data/issue309/apples_by_tree-006748a06fa311e6a7f8ef8b642e977b/mb-1-big-Data.db\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator.(AbstractSSTableIterator.java:130) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.columniterator.SSTableIterator.(SSTableIterator.java:46) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.io.sstable.format.big.BigTableReader.iterator(BigTableReader.java:69) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.io.sstable.format.big.BigTableScanner$KeyScanningIterator$1.initializeIterator(BigTableScanner.java:338) ~[apache-cassandra-3.0.8.jar:3.0.8]\n ... 19 common frames omitted\nCaused by: java.io.IOException: Corrupt (negative) value length encountered\n at org.apache.cassandra.db.marshal.AbstractType.readValue(AbstractType.java:399) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.BufferCell$Serializer.deserialize(BufferCell.java:302) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.readSimpleColumn(UnfilteredSerializer.java:462) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.deserializeRowBody(UnfilteredSerializer.java:440) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.deserializeStaticRow(UnfilteredSerializer.java:381) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator.readStaticRow(AbstractSSTableIterator.java:179) ~[apache-cassandra-3.0.8.jar:3.0.8]\n at org.apache.cassandra.db.columniterator.AbstractSSTableIterator.(AbstractSSTableIterator.java:103) ~[apache-cassandra-3.0.8.jar:3.0.8]\n ... 22 common frames omitted\n{code}\n\nAfter debugging, it appears that a previously dropped static column (weeks prior) was the instigator of the issue. As a workaround we added back the column, restarted all cassandra processes within the cluster, and the read error and corruption exception went away.\n\nAttached is a script to reproduce with a simple schema.\n\nAlso noteworthy (and shown in the script) is that when in this state, compaction silently failed (exit 0) to remove the dropped static columns from the \"corrupted\" sstable.\n","issue_id":"13001795","key":"CASSANDRA-12582","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-10-10T01:57:00.000+0000","role":"fixed_distractor","summary":"Removing static column results in ReadFailure due to CorruptSSTableException"} {"case_id":"13003422","cluster":"DISTRACTOR-CASSANDRA-12620","comments":[{"body":"[~blerer] could you have a look?","created":"2016-09-08T02:58:08.787+0000"},{"body":"Here's another table that is exhibiting this behaviour, but does not have the \"DESC+ASC\" clustering as discussed.\n\nThis data was inserted using the Spark connector {{saveToCassandra}} with {{WriteConf(ttl = TTLOption.constant(Duration.standardDays(2)))}} (it also has a default ttl of 2 days on the table definition)\n\nAs in the initial example, this data is never explicitly deleted and only TTLd.\n\n{code}\nCREATE TABLE applicationservices.aggregate_bucket_event_analysis_v3 (\n bucket_type int,\n bucket_id text,\n processed_date timestamp,\n results list>,\n PRIMARY KEY ((bucket_type, bucket_id), processed_date)\n) WITH CLUSTERING ORDER BY (processed_date DESC)\n AND bloom_filter_fp_chance = 0.01\n AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\n AND comment = ''\n AND compaction = {'class': 'org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy', 'max_threshold': '32', 'min_threshold': '4'}\n AND compression = {'chunk_length_in_kb': '64', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}\n AND crc_check_chance = 1.0\n AND dclocal_read_repair_chance = 0.1\n AND default_time_to_live = 172800\n AND gc_grace_seconds = 864000\n AND max_index_interval = 2048\n AND memtable_flush_period_in_ms = 0\n AND min_index_interval = 128\n AND read_repair_chance = 0.0\n AND speculative_retry = '99PERCENTILE';\n{code}\n\n{code}\n {\n \"partition\" : {\n \"key\" : [ \"0\", \"26316\" ],\n \"position\" : 18119\n },\n \"rows\" : [\n\t\t//.....\n\t\t{\n\t\t\t\"type\" : \"row\",\n\t\t\t\"position\" : 25035,\n\t\t\t\"clustering\" : [ \"2016-08-16 14:00Z\" ],\n\t\t\t\"liveness_info\" : { \"tstamp\" : \"2016-08-16T14:00:23.927005Z\" },\n\t\t\t\"cells\" : [ ]\n\t\t},\n\t\t{\n\t\t\t\"type\" : \"row\",\n\t\t\t\"position\" : 25054,\n\t\t\t\"clustering\" : [ \"2016-08-16 12:00Z\" ],\n\t\t\t\"liveness_info\" : { \"tstamp\" : \"2016-08-16T12:00:13.880002Z\" },\n\t\t\t\"cells\" : [ ]\n\t\t},\n\t\t{\n\t\t\t\"type\" : \"row\",\n\t\t\t\"position\" : 25073,\n\t\t\t\"clustering\" : [ \"2016-08-16 10:00Z\" ],\n\t\t\t\"liveness_info\" : { \"tstamp\" : \"2016-08-16T10:00:32.998006Z\" },\n\t\t\t\"cells\" : [ ]\n\t\t},\t\n\t\t//.....\n{code}","created":"2016-09-08T05:15:45.545+0000"},{"body":"If needed, for above, the UDTs are:\n\n{code}\ncqlsh:applicationservices> describe type udt_aggregate_bucket_event_analysis_v3_result;\n\nCREATE TYPE applicationservices.udt_aggregate_bucket_event_analysis_v3_result (\n number_of_uses int,\n entity_type int,\n entity_value text,\n event_details frozen>>\n);\n\ncqlsh:applicationservices> describe type udt_aggregate_bucket_event_analysis_v3_details;\n\nCREATE TYPE applicationservices.udt_aggregate_bucket_event_analysis_v3_details (\n aggregate_id text,\n identity_sid text,\n event_type int,\n event_id text,\n event_date timestamp\n);\n{code}","created":"2016-09-08T05:17:21.608+0000"},{"body":"[~csauve] I did not manage to reproduce the problem. So guess, that I did something different that you.\nHere is what I did:\n# On the latest 2.1: I created the {{aggregate_bucket_event_v3}} table\n# Inserted a couple of dummy rows with TTL.\n# Waited for the TTL to expire.\n# Flushed the memtables to disk with nodetool\n# Stopped the server\n# Removed the commitlog and saved_cache directory\n# Started the 3.0 server\n# upgraded the SSTables with node tools\n\nDid you do something differently? Do you have the same problems for all your rows that have expired?","created":"2016-09-27T08:51:50.307+0000"},{"body":"Thank you [~blerer] \nFist of all we had 3-node DSE cluster. The upgrade procedure was:\n\n1. Disable OpsCenter repair service and Spark job-server and workers \n2. On the node being upgraded\n a. Drain the node using nodetool\n b. Stop DSE\n c. Upgrade to 3.0\n d. Start DSE \n e. upgrade SSTables using nodetool\n3. Repeat for all nodes in the cluster\n\nNote that upgradesstables stack a few times and we had to re-run it. More than one table was affected and many rows in those tables. ","created":"2016-09-27T13:52:25.941+0000"},{"body":"[~mguissine], [~csauve] could you give me the original content of one of the expired rows that was resurrected ?\n\nI start to wonder if the problem is not specific to some data. ","created":"2016-10-21T13:45:00.487+0000"},{"body":"[~blerer] Unfortunately we don't have backups of the data from pre-upgrade, those have been purged a while back. ","created":"2016-10-24T15:25:04.177+0000"},{"body":"Hi,\n\nWe just had the same problem, upgrading from Cassandra 2.1.13 / DSE 4.8.6 to Cassandra 3.0.10 / DSE 5.0.4. The problem has been identified on 3 different environments.\n\nDead / tombstoned rows with an expired TTL show up after the upgrade. The thing is only the primary key is non null, all other fields are null. We only insert non null values using TTL that have various TTL from 60s to 1day.\n\n{code}\nselect * from ttl_entries where key='d7c10084-724c-4117-9927-b927a972203b';\n\nkey | field1 | field2 | field3 | field4 | field5 | field6 | field7 | field8\n--------------------------------------+----------+-----------+-----------+----------+----------+--------------+--------+-----------\nd7c10084-724c-4117-9927-b927a972203b | null | null | null | null | null | null | null | null\n{code}\n\n\nHere's an {{sstabledump}} of one of the impacted row. Note that the {{liveness_info}} is missing TTL informations:\n\n{code:title=sstabledump row d7c10084-724c-4117-9927-b927a972203b}\n{\n \"partition\" : { \n \"key\" : [ \"d7c10084-724c-4117-9927-b927a972203b\" ], \n \"position\" : 6557684 \n }, \n \"rows\" : [ \n { \n \"type\" : \"row\", \n \"position\" : 6557734, \n \"liveness_info\" : { \"tstamp\" : \"2016-11-25T02:24:06.237Z\" }, \n \"cells\" : [ \n { \"name\" : \"field1\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field2\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field3\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field4\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field5\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field6\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field8\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"deletion_info\" : { \"marked_deleted\" : \"2016-11-25T02:24:06.236999Z\", \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_01\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_02\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_03\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_04\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_05\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_06\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_07\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_08\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_09\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_10\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_11\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_12\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_13\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_14\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_15\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_16\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_17\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_18\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_19\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_20\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_21\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_22\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_23\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_24\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_25\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } } \n ] \n } \n ] \n},\n{code}\n\n\nData inserted after the upgrade with TTL have correct TTL metadata:\n\n{code}\n\"liveness_info\" : { \"tstamp\" : \"2016-11-29T09:43:24.953Z\", \"ttl\" : 3600, \"expires_at\" : \"2016-11-29T10:43:24Z\", \"expired\" : true }\n{code}\n\nHere's the describe table.\n\n{code:title=table description after the upgrade}\nCREATE TABLE ttl_entries (\n key text PRIMARY KEY,\n field1 text,\n field2 int,\n field3 text,\n field4 bigint,\n field5 int,\n field6 text,\n field7 set,\n field8 text\n) WITH bloom_filter_fp_chance = 0.01\n AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\n AND compaction = {'class': 'org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy', 'max_threshold': '32', 'min_threshold': '4'}\n AND compression = {'chunk_length_in_kb': '64', 'class': 'org.apache.cassandra.io.compress.SnappyCompressor'}\n AND crc_check_chance = 1.0\n AND dclocal_read_repair_chance = 0.0\n AND default_time_to_live = 0\n AND gc_grace_seconds = 864000\n AND max_index_interval = 2048\n AND memtable_flush_period_in_ms = 0\n AND min_index_interval = 128\n AND read_repair_chance = 0.1\n AND speculative_retry = '99PERCENTILE';\n{code}\n\nOur upgrade procedure :\n\nFor each node before upgrading\n* nodetool repair -pr\n\nThen on each node \n* nodetool drain \n* stop the node \n* change files config \n* link to new binaries \n* start the nodes \n* nodetool upgradesstables \n\nNote: code snippet have been edited to remove sensible data (field names and values)\n\nPS : I think the title should be edited to indicates that it's rows with _expired TTL_ that shows up.","created":"2016-12-05T15:21:25.204+0000"},{"body":"[~bric3] do you have a dump of the original data (before upgrade)?","created":"2016-12-05T15:29:56.637+0000"},{"body":"[~blerer] We have rollbacked an environment. I'm not sure I can upload the dump (maybe in private). But here's another primary key that is *expired* before and after the upgrade :\n\nFor information 1479813209 : Tue, 22 Nov 2016 11:13:29 GMT\n\n{code:title=sstable2json / DSE 4.8.6 / Cassandra 2.1.13}\n{\"key\": \"1c5598b3-70de-4ba5-a9ec-a57a61d3f494\",\n \"cells\": [[\"\",1479813209,1479813209820000,\"d\"],\n [\"field1\",1479813209,1479813209820000,\"d\"],\n [\"field2\",1479813209,1479813209820000,\"d\"],\n [\"field3\",1479813209,1479813209820000,\"d\"],\n [\"field4\",1479813209,1479813209820000,\"d\"],\n [\"field5\",1479813209,1479813209820000,\"d\"],\n [\"field7:_\",\"scopes:!\",1479813209819999,\"t\",1479813209],\n [\"field7:41444d494e5f4150504c49434154494f4e\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f434154414c4f475f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f434154414c4f475f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f435245444954535f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f435245444954535f5452414e53414354494f4e535f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f435245444954535f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f454e41424c494e475f5041434b\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f464f524249445f4d534953444e5f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f464f524249445f4d534953444e5f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f484f54425f434845434b\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f484f54425f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f484f54425f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f4f464645525f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f4f464645525f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f5041434b5f50524943494e475f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f5041434b5f50524943494e475f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f504152544e45525f50524f564953494f4e494e475f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f504152544e45525f50524f564953494f4e494e475f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f50524f4d4f54494f4e5f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f50524f4d4f54494f4e5f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f5245504f52545f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f53434f5045535f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f534d535f53454e44\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f555345525f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f555345525f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f56414c49444154494f4e5f53454e44\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f564f4943455f44455354494e4154494f4e5f42554e444c45535f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f434c49454e545f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f434c49454e545f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f47524f55505f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f47524f55505f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f555345525f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f555345525f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field6\",1479813209,1479813209820000,\"d\"],\n [\"field8\",1479813209,1479813209820000,\"d\"]]},\n{code}\n\n{code:title=sstabledump / DSE 5.0.4 / Cassandra 3.0.10}\n{\n \"partition\" : {\n \"key\" : [ \"1c5598b3-70de-4ba5-a9ec-a57a61d3f494\" ],\n \"position\" : 0\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 50,\n \"liveness_info\" : { \"tstamp\" : \"2016-11-22T11:13:29.820Z\" },\n \"cells\" : [\n { \"name\" : \"field1\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field2\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field3\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field4\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field5\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field6\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field8\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"deletion_info\" : { \"marked_deleted\" : \"2016-11-22T11:13:29.819999Z\", \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_APPLICATION\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_CATALOG_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_CATALOG_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_CREDITS_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_CREDITS_TRANSACTIONS_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_CREDITS_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_ENABLING_PACK\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_FORBID_MSISDN_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_FORBID_MSISDN_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_HOTB_CHECK\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_HOTB_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_HOTB_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_OFFER_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" }\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_OFFER_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PACK_PRICING_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PACK_PRICING_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PARTNER_PROVISIONING_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PARTNER_PROVISIONING_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PROMOTION_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PROMOTION_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_REPORT_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_SCOPES_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_SMS_SEND\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_USER_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_USER_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_VALIDATION_SEND\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_VOICE_DESTINATION_BUNDLES_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_CLIENT_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_CLIENT_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_GROUP_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_GROUP_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_USER_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_USER_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } }\n ]\n }\n ]\n},\n{code}\n\n\n*Edit:* the sstable2json has been done *after* a repair.\n","created":"2016-12-05T15:58:36.770+0000"},{"body":"Not sure but the liveness info looks wrong, it seems the conversion for the the _ka_ format to the new _mc_ format _forgot_ to report the expired / delete row.\n\nIf I read it correctly {{\"cells\": \\[\\[\"\",1479813209,1479813209820000,\"d\"],}} is the row marker, and it shows the row has the status deleted. However the upgraded entry don't show this deleted status ({{\"liveness_info\" : \\{ \"tstamp\" : \"2016-11-22T11:13:29.820Z\" \\},}}).\n\n","created":"2016-12-05T16:45:25.687+0000"},{"body":"Thanks for all the information. I will focus on this problem this week.","created":"2016-12-06T07:57:58.311+0000"},{"body":"[~blerer] If you'll need db files we'll have to discuss privately ;)","created":"2016-12-06T14:21:58.285+0000"},{"body":"Thanks for the offer. I will try with the information you gave me and tell you if I need more. ","created":"2016-12-06T14:43:00.788+0000"},{"body":"[~blerer] Actually investigating a bit more we've found out that the issue may not be related to the upgradesstables.\n\nWe've managed to reproduce the issue with our dataset on 3 nodes (RF = 3). Here's the scenario *without upgrading sstables*:\n\n* all nodes are on DSE 4.8.6 / Cassandra 2.1.13 => select only show valid and live data.\n* on a single node\n** nodetool drain\n** stop the node\n** change files config\n** link to new binaries (5.0.4)\n** start the node\n* now select show dead / expired rows\n\nCould it be related to how cassandra reads or translate the old sstable format (in our case {{ka}})","created":"2016-12-06T15:46:43.691+0000"},{"body":"Sorry, for the delay. It tooks me some time and a certain amount of back and forth between the {{2.1}} and {{3.0}} code base to fully understand the problem.\n[~bric3] Thanks for the SSTables they helped a lot.\n\nIn {{2.1}} when a row is created with a {{TTL}} the row marker cell is created as an {{ExpiringCell}} but when the row is compacted if the {{TTL}} has expired the cell will be converted into a {{DeletedCell}}.\n\nIn {{3.0}} the code used to read the post 3.0 format was only expecting an {{ExpiringCell}} for rows with {{TTL}} and was ignoring the local deletion time of the {{DeletedCell}}. Due to that the row was not marked as deleted anymore.\n\nAs the problem was in the code used to deserialize the post 3.0 format it was not necessary to run {{upgradesstables}} to see the problem.\n\nI originally could not reproduce the problem because I did not think of running a compaction before upgrading (SizeTieredCompaction).\n\nI pushed a patch [here|https://github.com/apache/cassandra/compare/trunk...blerer:12620-3.0?expand=1].\n\n \n\n","created":"2016-12-15T10:47:28.037+0000"},{"body":"Not sure this is the proper fix (as in, it's likely fine for anyone that isn't doing wacky things, but it doesn't stricly respect backward compatibility).\n\nIn 2.x, I believe we never delete a row marker explicitely, so it can only be a tombstone due to TTLs. But a row marker expiring does not absolutely guarantee the whole row is gone. Consider the following example (which is weird, for sure, but possible):\n{noformat}\nCREATE TABLE t (k int PRIMARY KEY, v1 int, v2 int);\nUPDATE t USING TIMESTAMP 0 SET v1 = 1 WHERE k = 0;\nINSERT INTO t(k, v2) VALUES (0, 2) USING TIMESTAMP 1 AND TTL 3;\n{noformat}\nWhen the 2nd insert expires, we'll still have a row with {{v1=1}} (though no row marker internally), and this both in 2.x and 3.x. With this patch however, we might end up generating a row deletion with timestamp 1, which _would_ delete the first UPDATE, which is stricly speaking incorrect (for the CQL semantic).\n\nLong story short, a deleted row marker is not equivalent to a row deletion, it just mean there is no row marker, and as the equivalent in 3.x of a row marker is the primary key liveness info, all we want here is make sure we don't have one. In other words, I believe the proper patch is:\n{noformat}\n assert !cell.value.hasRemaining();\n- builder.addPrimaryKeyLivenessInfo(LivenessInfo.create(cell.timestamp, cell.ttl, cell.localDeletionTime));\n+ if (!cell.isTombstone())\n+ builder.addPrimaryKeyLivenessInfo(LivenessInfo.create(cell.timestamp, cell.ttl, cell.localDeletionTime));\n{noformat}\n","created":"2016-12-15T11:07:13.940+0000"},{"body":"Thanks for the review.\nI updated my [branch|https://github.com/apache/cassandra/compare/trunk...blerer:12620-3.0?expand=1] with your change and tested it. It work perfectly including the usecase that was not working with my original fix.\n\nI will run the upgrade tests and post the results once I got them.","created":"2016-12-15T13:09:46.280+0000"},{"body":"Thanks, +1 assuming the tests are clean (nit: I don't think you need the {{Row.Deletion}} import anymore).","created":"2016-12-15T13:16:44.739+0000"},{"body":"The upgrade tests results are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-12620-3.0-upgrade-upgrade/lastCompletedBuild/testReport/]. I checked the failing tests and they were already failing before the change.","created":"2016-12-16T11:06:34.907+0000"},{"body":"Committed into 3.0 at c612cd8d7dbd24888c216ad53f974686b88dd601 and merged into 3.11, 3.X and trunk","created":"2016-12-16T12:05:36.566+0000"},{"body":"Interesting, we didn't think about running {{nodetool compact}} just {{nodetool repair}}.\n\nFor self information purpose where are the sources of the tests run the jenksin instance ?","created":"2016-12-16T12:37:52.958+0000"},{"body":"The upgrade test are part of the Dtests [here|https://github.com/riptano/cassandra-dtest/tree/master/upgrade_tests]. They are run for all the possible upgrade paths.","created":"2016-12-19T08:58:28.818+0000"}],"conversations":[{"body":"We had the below table on C* 2.x (dse 4.8.4, we assume was 2.1.15.1423 according to documentation), and were entering TTLs at write-time using the DataStax C# Driver (using the POCO mapper).\n\nUpon upgrade to 3.0.8.1293 (DSE 5.0.2), we are seeing a lot of rows that:\n\n* should have been TTL'd\n* have no non-primary-key column data\n\n{code}\nCREATE TABLE applicationservices.aggregate_bucket_event_v3 (\n bucket_type int,\n bucket_id text,\n date timestamp,\n aggregate_id text,\n event_type int,\n event_id text,\n entities list>>,\n identity_sid text,\n PRIMARY KEY ((bucket_type, bucket_id), date, aggregate_id, event_type, event_id)\n) WITH CLUSTERING ORDER BY (date DESC, aggregate_id ASC, event_type ASC, event_id ASC)\n AND bloom_filter_fp_chance = 0.1\n AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\n AND comment = ''\n AND compaction = {'class': 'org.apache.cassandra.db.compaction.LeveledCompactionStrategy'}\n AND compression = {'chunk_length_in_kb': '64', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}\n AND crc_check_chance = 1.0\n AND dclocal_read_repair_chance = 0.1\n AND default_time_to_live = 0\n AND gc_grace_seconds = 864000\n AND max_index_interval = 2048\n AND memtable_flush_period_in_ms = 0\n AND min_index_interval = 128\n AND read_repair_chance = 0.0\n AND speculative_retry = '99PERCENTILE';\n{code}\n\n{code}\n{\n \"partition\" : {\n \"key\" : [ \"0\", \"26492\" ],\n \"position\" : 54397932\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 54397961,\n \"clustering\" : [ \"2016-09-07 23:33Z\", \"3651664\", \"0\", \"773665449947099136\" ],\n \"liveness_info\" : { \"tstamp\" : \"2016-09-07T23:34:09.758Z\", \"ttl\" : 172741, \"expires_at\" : \"2016-09-09T23:33:10Z\", \"expired\" : false },\n \"cells\" : [\n { \"name\" : \"identity_sid\", \"value\" : \"p_tw_zahidana\" },\n { \"name\" : \"entities\", \"deletion_info\" : { \"marked_deleted\" : \"2016-09-07T23:34:09.757999Z\", \"local_delete_time\" : \"2016-09-07T23:34:09Z\" } },\n { \"name\" : \"entities\", \"path\" : [ \"936e17e1-7553-11e6-9b92-29a33b5827c3\" ], \"value\" : \"0:https\\\\://www.youtube.com/watch?v=pwAJAssv6As\" },\n { \"name\" : \"entities\", \"path\" : [ \"936e17e2-7553-11e6-9b92-29a33b5827c3\" ], \"value\" : \"2:youtube\" }\n ]\n },\n {\n \"type\" : \"row\",\n\n },\n {\n \"type\" : \"row\",\n \"position\" : 54397177,\n \"clustering\" : [ \"2016-08-17 10:00Z\", \"6387376\", \"0\", \"765850666296225792\" ],\n \"liveness_info\" : { \"tstamp\" : \"2016-08-17T11:26:15.917001Z\" },\n \"cells\" : [ ]\n },\n {\n \"type\" : \"row\",\n \"position\" : 54397227,\n \"clustering\" : [ \"2016-08-17 07:00Z\", \"6387376\", \"0\", \"765805367347601409\" ],\n \"liveness_info\" : { \"tstamp\" : \"2016-08-17T08:11:17.587Z\" },\n \"cells\" : [ ]\n },\n {\n \"type\" : \"row\",\n \"position\" : 54397276,\n \"clustering\" : [ \"2016-08-17 04:00Z\", \"6387376\", \"0\", \"765760069858365441\" ],\n \"liveness_info\" : { \"tstamp\" : \"2016-08-17T05:58:11.228Z\" },\n \"cells\" : [ ]\n },\n{code}","from":"reporter","subject":"Resurrected rows with expired TTL on update to 3.x"},{"body":"[~blerer] could you have a look?","from":"developer"},{"body":"Here's another table that is exhibiting this behaviour, but does not have the \"DESC+ASC\" clustering as discussed.\n\nThis data was inserted using the Spark connector {{saveToCassandra}} with {{WriteConf(ttl = TTLOption.constant(Duration.standardDays(2)))}} (it also has a default ttl of 2 days on the table definition)\n\nAs in the initial example, this data is never explicitly deleted and only TTLd.\n\n{code}\nCREATE TABLE applicationservices.aggregate_bucket_event_analysis_v3 (\n bucket_type int,\n bucket_id text,\n processed_date timestamp,\n results list>,\n PRIMARY KEY ((bucket_type, bucket_id), processed_date)\n) WITH CLUSTERING ORDER BY (processed_date DESC)\n AND bloom_filter_fp_chance = 0.01\n AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\n AND comment = ''\n AND compaction = {'class': 'org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy', 'max_threshold': '32', 'min_threshold': '4'}\n AND compression = {'chunk_length_in_kb': '64', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}\n AND crc_check_chance = 1.0\n AND dclocal_read_repair_chance = 0.1\n AND default_time_to_live = 172800\n AND gc_grace_seconds = 864000\n AND max_index_interval = 2048\n AND memtable_flush_period_in_ms = 0\n AND min_index_interval = 128\n AND read_repair_chance = 0.0\n AND speculative_retry = '99PERCENTILE';\n{code}\n\n{code}\n {\n \"partition\" : {\n \"key\" : [ \"0\", \"26316\" ],\n \"position\" : 18119\n },\n \"rows\" : [\n\t\t//.....\n\t\t{\n\t\t\t\"type\" : \"row\",\n\t\t\t\"position\" : 25035,\n\t\t\t\"clustering\" : [ \"2016-08-16 14:00Z\" ],\n\t\t\t\"liveness_info\" : { \"tstamp\" : \"2016-08-16T14:00:23.927005Z\" },\n\t\t\t\"cells\" : [ ]\n\t\t},\n\t\t{\n\t\t\t\"type\" : \"row\",\n\t\t\t\"position\" : 25054,\n\t\t\t\"clustering\" : [ \"2016-08-16 12:00Z\" ],\n\t\t\t\"liveness_info\" : { \"tstamp\" : \"2016-08-16T12:00:13.880002Z\" },\n\t\t\t\"cells\" : [ ]\n\t\t},\n\t\t{\n\t\t\t\"type\" : \"row\",\n\t\t\t\"position\" : 25073,\n\t\t\t\"clustering\" : [ \"2016-08-16 10:00Z\" ],\n\t\t\t\"liveness_info\" : { \"tstamp\" : \"2016-08-16T10:00:32.998006Z\" },\n\t\t\t\"cells\" : [ ]\n\t\t},\t\n\t\t//.....\n{code}","from":"developer"},{"body":"If needed, for above, the UDTs are:\n\n{code}\ncqlsh:applicationservices> describe type udt_aggregate_bucket_event_analysis_v3_result;\n\nCREATE TYPE applicationservices.udt_aggregate_bucket_event_analysis_v3_result (\n number_of_uses int,\n entity_type int,\n entity_value text,\n event_details frozen>>\n);\n\ncqlsh:applicationservices> describe type udt_aggregate_bucket_event_analysis_v3_details;\n\nCREATE TYPE applicationservices.udt_aggregate_bucket_event_analysis_v3_details (\n aggregate_id text,\n identity_sid text,\n event_type int,\n event_id text,\n event_date timestamp\n);\n{code}","from":"developer"},{"body":"[~csauve] I did not manage to reproduce the problem. So guess, that I did something different that you.\nHere is what I did:\n# On the latest 2.1: I created the {{aggregate_bucket_event_v3}} table\n# Inserted a couple of dummy rows with TTL.\n# Waited for the TTL to expire.\n# Flushed the memtables to disk with nodetool\n# Stopped the server\n# Removed the commitlog and saved_cache directory\n# Started the 3.0 server\n# upgraded the SSTables with node tools\n\nDid you do something differently? Do you have the same problems for all your rows that have expired?","from":"developer"},{"body":"Thank you [~blerer] \nFist of all we had 3-node DSE cluster. The upgrade procedure was:\n\n1. Disable OpsCenter repair service and Spark job-server and workers \n2. On the node being upgraded\n a. Drain the node using nodetool\n b. Stop DSE\n c. Upgrade to 3.0\n d. Start DSE \n e. upgrade SSTables using nodetool\n3. Repeat for all nodes in the cluster\n\nNote that upgradesstables stack a few times and we had to re-run it. More than one table was affected and many rows in those tables. ","from":"developer"},{"body":"[~mguissine], [~csauve] could you give me the original content of one of the expired rows that was resurrected ?\n\nI start to wonder if the problem is not specific to some data. ","from":"developer"},{"body":"[~blerer] Unfortunately we don't have backups of the data from pre-upgrade, those have been purged a while back. ","from":"developer"},{"body":"Hi,\n\nWe just had the same problem, upgrading from Cassandra 2.1.13 / DSE 4.8.6 to Cassandra 3.0.10 / DSE 5.0.4. The problem has been identified on 3 different environments.\n\nDead / tombstoned rows with an expired TTL show up after the upgrade. The thing is only the primary key is non null, all other fields are null. We only insert non null values using TTL that have various TTL from 60s to 1day.\n\n{code}\nselect * from ttl_entries where key='d7c10084-724c-4117-9927-b927a972203b';\n\nkey | field1 | field2 | field3 | field4 | field5 | field6 | field7 | field8\n--------------------------------------+----------+-----------+-----------+----------+----------+--------------+--------+-----------\nd7c10084-724c-4117-9927-b927a972203b | null | null | null | null | null | null | null | null\n{code}\n\n\nHere's an {{sstabledump}} of one of the impacted row. Note that the {{liveness_info}} is missing TTL informations:\n\n{code:title=sstabledump row d7c10084-724c-4117-9927-b927a972203b}\n{\n \"partition\" : { \n \"key\" : [ \"d7c10084-724c-4117-9927-b927a972203b\" ], \n \"position\" : 6557684 \n }, \n \"rows\" : [ \n { \n \"type\" : \"row\", \n \"position\" : 6557734, \n \"liveness_info\" : { \"tstamp\" : \"2016-11-25T02:24:06.237Z\" }, \n \"cells\" : [ \n { \"name\" : \"field1\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field2\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field3\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field4\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field5\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field6\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field8\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"deletion_info\" : { \"marked_deleted\" : \"2016-11-25T02:24:06.236999Z\", \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_01\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_02\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_03\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_04\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_05\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_06\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_07\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_08\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_09\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_10\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_11\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_12\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_13\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_14\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_15\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_16\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_17\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_18\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_19\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_20\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_21\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_22\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_23\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_24\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } }, \n { \"name\" : \"field7\", \"path\" : [ \"VALUE_25\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-25T02:24:06Z\" } } \n ] \n } \n ] \n},\n{code}\n\n\nData inserted after the upgrade with TTL have correct TTL metadata:\n\n{code}\n\"liveness_info\" : { \"tstamp\" : \"2016-11-29T09:43:24.953Z\", \"ttl\" : 3600, \"expires_at\" : \"2016-11-29T10:43:24Z\", \"expired\" : true }\n{code}\n\nHere's the describe table.\n\n{code:title=table description after the upgrade}\nCREATE TABLE ttl_entries (\n key text PRIMARY KEY,\n field1 text,\n field2 int,\n field3 text,\n field4 bigint,\n field5 int,\n field6 text,\n field7 set,\n field8 text\n) WITH bloom_filter_fp_chance = 0.01\n AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\n AND compaction = {'class': 'org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy', 'max_threshold': '32', 'min_threshold': '4'}\n AND compression = {'chunk_length_in_kb': '64', 'class': 'org.apache.cassandra.io.compress.SnappyCompressor'}\n AND crc_check_chance = 1.0\n AND dclocal_read_repair_chance = 0.0\n AND default_time_to_live = 0\n AND gc_grace_seconds = 864000\n AND max_index_interval = 2048\n AND memtable_flush_period_in_ms = 0\n AND min_index_interval = 128\n AND read_repair_chance = 0.1\n AND speculative_retry = '99PERCENTILE';\n{code}\n\nOur upgrade procedure :\n\nFor each node before upgrading\n* nodetool repair -pr\n\nThen on each node \n* nodetool drain \n* stop the node \n* change files config \n* link to new binaries \n* start the nodes \n* nodetool upgradesstables \n\nNote: code snippet have been edited to remove sensible data (field names and values)\n\nPS : I think the title should be edited to indicates that it's rows with _expired TTL_ that shows up.","from":"developer"},{"body":"[~bric3] do you have a dump of the original data (before upgrade)?","from":"developer"},{"body":"[~blerer] We have rollbacked an environment. I'm not sure I can upload the dump (maybe in private). But here's another primary key that is *expired* before and after the upgrade :\n\nFor information 1479813209 : Tue, 22 Nov 2016 11:13:29 GMT\n\n{code:title=sstable2json / DSE 4.8.6 / Cassandra 2.1.13}\n{\"key\": \"1c5598b3-70de-4ba5-a9ec-a57a61d3f494\",\n \"cells\": [[\"\",1479813209,1479813209820000,\"d\"],\n [\"field1\",1479813209,1479813209820000,\"d\"],\n [\"field2\",1479813209,1479813209820000,\"d\"],\n [\"field3\",1479813209,1479813209820000,\"d\"],\n [\"field4\",1479813209,1479813209820000,\"d\"],\n [\"field5\",1479813209,1479813209820000,\"d\"],\n [\"field7:_\",\"scopes:!\",1479813209819999,\"t\",1479813209],\n [\"field7:41444d494e5f4150504c49434154494f4e\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f434154414c4f475f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f434154414c4f475f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f435245444954535f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f435245444954535f5452414e53414354494f4e535f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f435245444954535f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f454e41424c494e475f5041434b\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f464f524249445f4d534953444e5f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f464f524249445f4d534953444e5f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f484f54425f434845434b\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f484f54425f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f484f54425f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f4f464645525f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f4f464645525f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f5041434b5f50524943494e475f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f5041434b5f50524943494e475f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f504152544e45525f50524f564953494f4e494e475f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f504152544e45525f50524f564953494f4e494e475f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f50524f4d4f54494f4e5f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f50524f4d4f54494f4e5f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f5245504f52545f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f53434f5045535f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f534d535f53454e44\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f555345525f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f555345525f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f56414c49444154494f4e5f53454e44\",1479813209,1479813209820000,\"d\"],\n [\"field7:41444d494e5f564f4943455f44455354494e4154494f4e5f42554e444c45535f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f434c49454e545f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f434c49454e545f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f47524f55505f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f47524f55505f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f555345525f52454144\",1479813209,1479813209820000,\"d\"],\n [\"field7:415554485f41444d494e5f555345525f5752495445\",1479813209,1479813209820000,\"d\"],\n [\"field6\",1479813209,1479813209820000,\"d\"],\n [\"field8\",1479813209,1479813209820000,\"d\"]]},\n{code}\n\n{code:title=sstabledump / DSE 5.0.4 / Cassandra 3.0.10}\n{\n \"partition\" : {\n \"key\" : [ \"1c5598b3-70de-4ba5-a9ec-a57a61d3f494\" ],\n \"position\" : 0\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 50,\n \"liveness_info\" : { \"tstamp\" : \"2016-11-22T11:13:29.820Z\" },\n \"cells\" : [\n { \"name\" : \"field1\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field2\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field3\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field4\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field5\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field6\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field8\", \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"deletion_info\" : { \"marked_deleted\" : \"2016-11-22T11:13:29.819999Z\", \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_APPLICATION\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_CATALOG_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_CATALOG_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_CREDITS_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_CREDITS_TRANSACTIONS_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_CREDITS_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_ENABLING_PACK\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_FORBID_MSISDN_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_FORBID_MSISDN_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_HOTB_CHECK\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_HOTB_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_HOTB_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_OFFER_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" }\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_OFFER_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PACK_PRICING_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PACK_PRICING_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PARTNER_PROVISIONING_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PARTNER_PROVISIONING_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PROMOTION_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_PROMOTION_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_REPORT_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_SCOPES_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_SMS_SEND\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_USER_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_USER_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_VALIDATION_SEND\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"ADMIN_VOICE_DESTINATION_BUNDLES_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_CLIENT_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_CLIENT_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_GROUP_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_GROUP_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_USER_READ\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } },\n { \"name\" : \"field7\", \"path\" : [ \"AUTH_ADMIN_USER_WRITE\" ], \"deletion_info\" : { \"local_delete_time\" : \"2016-11-22T11:13:29Z\" } }\n ]\n }\n ]\n},\n{code}\n\n\n*Edit:* the sstable2json has been done *after* a repair.\n","from":"developer"},{"body":"Not sure but the liveness info looks wrong, it seems the conversion for the the _ka_ format to the new _mc_ format _forgot_ to report the expired / delete row.\n\nIf I read it correctly {{\"cells\": \\[\\[\"\",1479813209,1479813209820000,\"d\"],}} is the row marker, and it shows the row has the status deleted. However the upgraded entry don't show this deleted status ({{\"liveness_info\" : \\{ \"tstamp\" : \"2016-11-22T11:13:29.820Z\" \\},}}).\n\n","from":"developer"},{"body":"Thanks for all the information. I will focus on this problem this week.","from":"developer"},{"body":"[~blerer] If you'll need db files we'll have to discuss privately ;)","from":"developer"},{"body":"Thanks for the offer. I will try with the information you gave me and tell you if I need more. ","from":"developer"},{"body":"[~blerer] Actually investigating a bit more we've found out that the issue may not be related to the upgradesstables.\n\nWe've managed to reproduce the issue with our dataset on 3 nodes (RF = 3). Here's the scenario *without upgrading sstables*:\n\n* all nodes are on DSE 4.8.6 / Cassandra 2.1.13 => select only show valid and live data.\n* on a single node\n** nodetool drain\n** stop the node\n** change files config\n** link to new binaries (5.0.4)\n** start the node\n* now select show dead / expired rows\n\nCould it be related to how cassandra reads or translate the old sstable format (in our case {{ka}})","from":"developer"},{"body":"Sorry, for the delay. It tooks me some time and a certain amount of back and forth between the {{2.1}} and {{3.0}} code base to fully understand the problem.\n[~bric3] Thanks for the SSTables they helped a lot.\n\nIn {{2.1}} when a row is created with a {{TTL}} the row marker cell is created as an {{ExpiringCell}} but when the row is compacted if the {{TTL}} has expired the cell will be converted into a {{DeletedCell}}.\n\nIn {{3.0}} the code used to read the post 3.0 format was only expecting an {{ExpiringCell}} for rows with {{TTL}} and was ignoring the local deletion time of the {{DeletedCell}}. Due to that the row was not marked as deleted anymore.\n\nAs the problem was in the code used to deserialize the post 3.0 format it was not necessary to run {{upgradesstables}} to see the problem.\n\nI originally could not reproduce the problem because I did not think of running a compaction before upgrading (SizeTieredCompaction).\n\nI pushed a patch [here|https://github.com/apache/cassandra/compare/trunk...blerer:12620-3.0?expand=1].\n\n \n\n","from":"developer"},{"body":"Not sure this is the proper fix (as in, it's likely fine for anyone that isn't doing wacky things, but it doesn't stricly respect backward compatibility).\n\nIn 2.x, I believe we never delete a row marker explicitely, so it can only be a tombstone due to TTLs. But a row marker expiring does not absolutely guarantee the whole row is gone. Consider the following example (which is weird, for sure, but possible):\n{noformat}\nCREATE TABLE t (k int PRIMARY KEY, v1 int, v2 int);\nUPDATE t USING TIMESTAMP 0 SET v1 = 1 WHERE k = 0;\nINSERT INTO t(k, v2) VALUES (0, 2) USING TIMESTAMP 1 AND TTL 3;\n{noformat}\nWhen the 2nd insert expires, we'll still have a row with {{v1=1}} (though no row marker internally), and this both in 2.x and 3.x. With this patch however, we might end up generating a row deletion with timestamp 1, which _would_ delete the first UPDATE, which is stricly speaking incorrect (for the CQL semantic).\n\nLong story short, a deleted row marker is not equivalent to a row deletion, it just mean there is no row marker, and as the equivalent in 3.x of a row marker is the primary key liveness info, all we want here is make sure we don't have one. In other words, I believe the proper patch is:\n{noformat}\n assert !cell.value.hasRemaining();\n- builder.addPrimaryKeyLivenessInfo(LivenessInfo.create(cell.timestamp, cell.ttl, cell.localDeletionTime));\n+ if (!cell.isTombstone())\n+ builder.addPrimaryKeyLivenessInfo(LivenessInfo.create(cell.timestamp, cell.ttl, cell.localDeletionTime));\n{noformat}\n","from":"developer"},{"body":"Thanks for the review.\nI updated my [branch|https://github.com/apache/cassandra/compare/trunk...blerer:12620-3.0?expand=1] with your change and tested it. It work perfectly including the usecase that was not working with my original fix.\n\nI will run the upgrade tests and post the results once I got them.","from":"developer"},{"body":"Thanks, +1 assuming the tests are clean (nit: I don't think you need the {{Row.Deletion}} import anymore).","from":"developer"},{"body":"The upgrade tests results are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-12620-3.0-upgrade-upgrade/lastCompletedBuild/testReport/]. I checked the failing tests and they were already failing before the change.","from":"developer"},{"body":"Committed into 3.0 at c612cd8d7dbd24888c216ad53f974686b88dd601 and merged into 3.11, 3.X and trunk","from":"developer"},{"body":"Interesting, we didn't think about running {{nodetool compact}} just {{nodetool repair}}.\n\nFor self information purpose where are the sources of the tests run the jenksin instance ?","from":"developer"},{"body":"The upgrade test are part of the Dtests [here|https://github.com/riptano/cassandra-dtest/tree/master/upgrade_tests]. They are run for all the possible upgrade paths.","from":"developer"}],"created":"2016-09-08T02:56:45.000+0000","description":"We had the below table on C* 2.x (dse 4.8.4, we assume was 2.1.15.1423 according to documentation), and were entering TTLs at write-time using the DataStax C# Driver (using the POCO mapper).\n\nUpon upgrade to 3.0.8.1293 (DSE 5.0.2), we are seeing a lot of rows that:\n\n* should have been TTL'd\n* have no non-primary-key column data\n\n{code}\nCREATE TABLE applicationservices.aggregate_bucket_event_v3 (\n bucket_type int,\n bucket_id text,\n date timestamp,\n aggregate_id text,\n event_type int,\n event_id text,\n entities list>>,\n identity_sid text,\n PRIMARY KEY ((bucket_type, bucket_id), date, aggregate_id, event_type, event_id)\n) WITH CLUSTERING ORDER BY (date DESC, aggregate_id ASC, event_type ASC, event_id ASC)\n AND bloom_filter_fp_chance = 0.1\n AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\n AND comment = ''\n AND compaction = {'class': 'org.apache.cassandra.db.compaction.LeveledCompactionStrategy'}\n AND compression = {'chunk_length_in_kb': '64', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}\n AND crc_check_chance = 1.0\n AND dclocal_read_repair_chance = 0.1\n AND default_time_to_live = 0\n AND gc_grace_seconds = 864000\n AND max_index_interval = 2048\n AND memtable_flush_period_in_ms = 0\n AND min_index_interval = 128\n AND read_repair_chance = 0.0\n AND speculative_retry = '99PERCENTILE';\n{code}\n\n{code}\n{\n \"partition\" : {\n \"key\" : [ \"0\", \"26492\" ],\n \"position\" : 54397932\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 54397961,\n \"clustering\" : [ \"2016-09-07 23:33Z\", \"3651664\", \"0\", \"773665449947099136\" ],\n \"liveness_info\" : { \"tstamp\" : \"2016-09-07T23:34:09.758Z\", \"ttl\" : 172741, \"expires_at\" : \"2016-09-09T23:33:10Z\", \"expired\" : false },\n \"cells\" : [\n { \"name\" : \"identity_sid\", \"value\" : \"p_tw_zahidana\" },\n { \"name\" : \"entities\", \"deletion_info\" : { \"marked_deleted\" : \"2016-09-07T23:34:09.757999Z\", \"local_delete_time\" : \"2016-09-07T23:34:09Z\" } },\n { \"name\" : \"entities\", \"path\" : [ \"936e17e1-7553-11e6-9b92-29a33b5827c3\" ], \"value\" : \"0:https\\\\://www.youtube.com/watch?v=pwAJAssv6As\" },\n { \"name\" : \"entities\", \"path\" : [ \"936e17e2-7553-11e6-9b92-29a33b5827c3\" ], \"value\" : \"2:youtube\" }\n ]\n },\n {\n \"type\" : \"row\",\n\n },\n {\n \"type\" : \"row\",\n \"position\" : 54397177,\n \"clustering\" : [ \"2016-08-17 10:00Z\", \"6387376\", \"0\", \"765850666296225792\" ],\n \"liveness_info\" : { \"tstamp\" : \"2016-08-17T11:26:15.917001Z\" },\n \"cells\" : [ ]\n },\n {\n \"type\" : \"row\",\n \"position\" : 54397227,\n \"clustering\" : [ \"2016-08-17 07:00Z\", \"6387376\", \"0\", \"765805367347601409\" ],\n \"liveness_info\" : { \"tstamp\" : \"2016-08-17T08:11:17.587Z\" },\n \"cells\" : [ ]\n },\n {\n \"type\" : \"row\",\n \"position\" : 54397276,\n \"clustering\" : [ \"2016-08-17 04:00Z\", \"6387376\", \"0\", \"765760069858365441\" ],\n \"liveness_info\" : { \"tstamp\" : \"2016-08-17T05:58:11.228Z\" },\n \"cells\" : [ ]\n },\n{code}","issue_id":"13003422","key":"CASSANDRA-12620","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-12-16T12:05:36.000+0000","role":"fixed_distractor","summary":"Resurrected rows with expired TTL on update to 3.x"} {"case_id":"13005368","cluster":"DISTRACTOR-CASSANDRA-12653","comments":[{"body":"You are definitely correct about the cause of this - it occasionally surfaces in test failures such as CASSANDRA-11689. I agree with your conclusion that this doesn't cause any serious harm, but I think it is worth fixing.","created":"2016-09-16T15:43:09.811+0000"},{"body":"It's been a while since I've looked at this, but I agree that we should try to fix it. I think the easiest would be to throw away additional shadow round replies, if we can flag them as such.","created":"2016-09-19T06:35:21.897+0000"},{"body":"I think we can get around changing gossip message values by making sure that we don't process any ACKs for which (any) SYN hasn't been send before by the {{Gossiper}}. I've added {{Gossiper.lastSynSendAt}} that can be used to drop messages in the verb handler.\n","created":"2016-09-19T12:42:10.011+0000"},{"body":"\nThe attached patch will solve this issue by using a separate data structure for the information gathered during shadow rounds and by using a timestamp for the first syn send. \n\n||trunk||3.0||2.2||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-trunk]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-3.0]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-2.2]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-testall/]|\n\n\nI can also create patches for merge conflicts if needed after agreeing on which versions to target.\n","created":"2016-09-22T14:15:52.244+0000"},{"body":"Are you still planing to review the patch, [~jkni]?\n","created":"2016-11-03T08:52:53.363+0000"},{"body":"[~spodxx@gmail.com] - yes! I sincerely apologize for the delay here. If anyone else is interested in reviewing this, they're welcome to pick it up, but it's near the top of my list and I hope to get to this soon.","created":"2016-11-04T20:16:00.953+0000"},{"body":"What's status here [~jkni]?","created":"2017-01-10T15:43:55.140+0000"},{"body":"Thanks for the ping - I gave the patch another skim and have some questions.\n\nThe approach of returning endpoint states obtained through gossip seems sound. I definitely like this idea because it (mostly) prevents us from needing to reason about how shadow rounds affect the proper gossip process. That said, I'm not sure it fully accomplishes this goal. At the moment, it should be safe if a later response comes back after the gossiper has started, but we need to be careful to preserve this property in the future.\n\nI'm not sure how the timestamp check is supposed to work. It initializes the field using System.nanoTime() the first time we send a gossip, but in the gossip digest ack verb handler, we check timestamps using the endpoint state update timestamp, which is not serialized inter-node and also initialized using System.nanoTime() by the local JVM. It seems to me that this reduces to a boolean check that the gossiper has been properly started at least once, since this check will only fail when firstSynSendAt == 0. Am I missing something here? It also seems to me that we should initialize the field on starting the gossiper rather than checking and possibly initializing it every time we send a gossip message.","created":"2017-01-10T20:33:59.944+0000"},{"body":"The firstSynSendAt value is only set while sending the first syn during regular gossip. All ACKs will be ignored that have been received and queued before that point. Please keep in mind that the first SYN is send by GossipTasks by using delayed execution, so it's not enough to simply check if the Gossiper has already been started.\n\nYou are correct in pointing out that the timestamp of the incoming ACK will be a local nanoTime() value. In-flight ACKs triggered during the shadow round will not be filtered by the timestamp check unless already deserialized. But I'd prefer comparing local timestamps anyways, as they will not require some kind of clock skew tolerance, which would be hard to come up with given the kind of races we like to address. Also this shouldn't be a problem as, like you mentioned, the current handling of empty shadow round ACK replies should be safe. But it's not a very clean solution and I'd rather be able to either have a causal relationship to the corresponding SYN/Gossip round or even a custom ShadowGossipMessage type. But this is probably better taken care of in e.g. CASSANDRA-12345 instead of this bug fix.\n","created":"2017-01-11T15:29:39.131+0000"},{"body":"Thanks for the quick response; I agree with all the points in your message. My gut instinct is to make the patch as small as possible since we agree that establishing a causal relationship or explicitly separating the shadow gossip round is the proper long-term solution, but the patch isn't particularly large either way, so I'll move forward with the the patch as proposed.\n\nI'll give the patches another review for any small fixes.","created":"2017-01-11T16:00:47.466+0000"},{"body":"[~jkni] - this still on your radar?\n","created":"2017-02-11T06:27:43.686+0000"},{"body":"[~jjirsa] - yes. That said, I don't anticipate getting to it within the next couple days, so feel free to give it the final review if it is higher priority for you than that. \n\nI've given several passes of review to the core concepts and they seem good. I think a final code style/details pass is all that remains.","created":"2017-02-13T05:32:26.995+0000"},{"body":"[~spodxx@gmail.com] - is there some deeper purpose of moving the {{FD.instance.isAlive()}} check higher in {{MigrationTask#runMayThrow()}} method beyond \"check to see if it's dead before we bother checking to see if it's worth sending a migration task\"? Is there a reason we don't let {{MM#shouldPullSchemaFrom}} return false if FD says the instance is dead? \n\nGiven that the shadow round is meant to just get ring state without changing anything, should we add an explicit check to {{MigrationManager#scheduleSchemaPull()}} to ensure that {{Gossiper.instance.isInShadowRound()}} is false before scheduling?\n","created":"2017-02-13T21:11:10.517+0000"},{"body":"bq. Stefan Podkowinski - is there some deeper purpose of moving the FD.instance.isAlive() check higher in MigrationTask#runMayThrow() method beyond \"check to see if it's dead before we bother checking to see if it's worth sending a migration task\"? Is there a reason we don't let MM#shouldPullSchemaFrom return false if FD says the instance is dead?\n\nWe could move FS.isAlive into MM.shouldPullSchemaFrom, yes. Not totally against it, but the log message in MigrationTask in case of a false return value would have to be changed and actually the isAlive status should only be relevant at task execution, as there's a 60 second delay after submitting it. So in theory you could submit a task for a node that has been dead but will be alive again at time of execution.\n\nbq. Given that the shadow round is meant to just get ring state without changing anything, should we add an explicit check to MigrationManager#scheduleSchemaPull() to ensure that Gossiper.instance.isInShadowRound() is false before scheduling?\n\nThe MigrationManager should never issue a schema pull during shadow round. If we add such check, I'd prefer to throw an exception instead of failing silently, instead of letting the process run in an undefined state. On the other hand, it's not really the business of the MM to monitor the gossiper life-cycle when it comes to separations of concerns.","created":"2017-02-14T10:02:23.390+0000"},{"body":"Fair points all around. I've deployed this as-is into a decent sized test cluster just to see how it behaves - first impressions are good compared to 3.0.10. Would love to see this land in 3.0.11 - if [~jkni] doesn't get to it in time for cutting 3.0.11, I'll try to get a more formal review pass on it.\n","created":"2017-02-14T15:59:38.534+0000"},{"body":"A few questions/comments on the latest patches:\n* On all versions, a space is missing in the conditional {{if(firstSynSendAt == 0)}}\n* On 2.2, it looks like the patch adds the {{valuesEqual}} method from later versions. Was this intentional? It looks unused.\n* On all versions, using {{firstSynSendAt == 0}} to check if it has been initialized isn't entirely safe. It's entirely legal (although admittedly rare) for {{System.nanoTime}} to return 0. If this happened, all acks would be rejected.\n* Comparisons for two {{System.nanoTime}} values (such as in {{GossipDigestAckVerbHandler}}) should not use t1 < t2. Instead, one should check the difference (t1 - t2 < 0) because numerical overflow could occur in the {{System.nanoTime}} long.\n* In {{maybeFinishShadowRound}}/{{finishShadowRound}}, we should add the states to the {{endpointShadowStateMap}} before setting {{inShadowRound}} to false. It looks like the current behavior admits a race where {{doShadowRound}} could read {{inShadowRound == false}} and exit its loop and copy the endpointShadowStateMap before it is filled by shadow round finish.\n* I believe {{firstSynSendAt}} is accessed from multiple threads and needs to be volatile.","created":"2017-02-20T06:06:34.021+0000"},{"body":"Thanks for your comments, Joel! Patch has been rebased and updated by following your suggestions, except from the issue mentioned below. \n\nI've also remove two @VisibleForTesting annotations in 2.2, since the related test isn't available in this branch.\n\nbq. On all versions, using firstSynSendAt == 0 to check if it has been initialized isn't entirely safe. It's entirely legal (although admittedly rare) for System.nanoTime to return 0. If this happened, all acks would be rejected.\n\nIn this very rare case we'd only discard the acks of 1 or 2 nodes from a single gossip round, as the firstSynSendAt value would afterwards be set by the next periodic sendGossip call. I'd therefor prefer to leave the variable as is, beside making it volatile.\n\n\n||2.2||3.0||3.11||trunk||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-2.2]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-3.0]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-3.11]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-trunk]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.11-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.11-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-testall/]|\n\n","created":"2017-02-20T14:56:21.983+0000"},{"body":"On the whole, I'm pretty good with this patch - nice work, Stefan!\n\nI have two nits:\n- the 3.11 and trunk branches have the tests included, but 2.2 and 3.0 do not. I that just an oversight? Can we add the tests to those branches, as well?\n- {{GossipDigestAckVerbHandler#doVerb}}, you get the [following timestamp|https://github.com/spodkowinski/cassandra/commit/9179b7ee06c51a79881f5be18cd01261ebe62143#diff-787d4963a51f20221468e976df1b121aR64]: {code} long ts = epStateMap.values().iterator().next().getUpdateTimestamp();{code}\nWe [do not serialize|https://github.com/apache/cassandra/blob/cassandra-2.2/src/java/org/apache/cassandra/gms/EndpointState.java#L166] {{EndpointState.updateTimestamp}}, so the value at the receiver ends up being the receiver's {{System.nanoTime()}}, as can be seen from the [{{EndpointState}} constructor|https://github.com/apache/cassandra/blob/cassandra-2.2/src/java/org/apache/cassandra/gms/EndpointState.java#L58]. What is it you are looking to confirm here? I'm kinda leaning toward eliminating that check altogether as {code}Gossiper.instance.firstSynSendAt == 0{code} might be a sufficient check anyways.\n\n","created":"2017-02-22T14:36:58.631+0000"},{"body":"Also, I'm not sure {{Gossiper#doShadowRound}} needs to be {{synchronized}}. Currently, that method is only invoked from {{StorageService#prepareToJoin}} (in all current branches). Is it necessary to mark it as such? If it is, can you please add a comment to the method.","created":"2017-02-22T14:45:52.722+0000"},{"body":"I think I can answer these - feel free to correct me, [~spodxx@gmail.com]\n\nIn order,\n* Presently, the tests depend on the mock MessagingService, which was added in [CASSANDRA-12016] to 3.10+. We'd new tests for 2.2/3.0+, which is desirable, but I have no great ideas how to do it other than fiddly byteman tests. \n* I agree with this. Stefan and I discussed it on the first pass of review, and I wouldn't mind eliminating that check altogether and making it a boolean. OTOH, it's cheap to check deserialization time and excludes the messages that were deserialized prior to the check. OTOH, there's no meaningful distinction in correctness-preserving behaviors between that and arbitrarily delayed gossip messages, and we need to handle the latter correctly anyway. I'm most concerned about this check giving future readers false hope :).\n* It also seems to be me that it doesn't presently need to be synchronized. That said, I assumed it was a defensive choice because the internals are definitely not safe to call on multiple threads, and someone may make that mistake in the future.","created":"2017-02-22T15:21:58.566+0000"},{"body":"The goal of introducing the firstSynSendAt timestamp was to prevent processing ACKs (lagged over from shadow gossip round) before any regular SYN has been send. I still think checking this value against the local message construction time is a good idea, as messages will be queued before getting dispatched and it's not that unlikely that the scheduler will make this happen for older messages just after we send the syn. As [mentioned before|https://issues.apache.org/jira/browse/CASSANDRA-12653?focusedCommentId=15818621&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15818621] the proper way to do this would be to be able to correlate between messages or do the shadow round conversation based on new dedicated message types.\n","created":"2017-02-23T10:00:04.917+0000"},{"body":"Do those answers address your questions well enough, [~jasobrown]? The latest patch addressed my concerns, but I don't want to step on your toes.\n\nI had to restart dtests for 2.2, but the latest patch/CI looks good to me otherwise.","created":"2017-02-24T17:08:39.690+0000"},{"body":"[~spod] I'm on board with {{firstSynSendAt}} being a timestamp, I'm just questioning what you compare it against. Instead of this:\n\n{code}\nlong ts = epStateMap.values().iterator().next().getUpdateTimestamp();\nif ((ts - Gossiper.instance.firstSynSendAt) < 0 || Gossiper.instance.firstSynSendAt == 0)\n{code}\n\nI'm suggesting this:\n{code}\nif ((System.nanoTime() - Gossiper.instance.firstSynSendAt) < 0 || Gossiper.instance.firstSynSendAt == 0)\n{code}\n\nIt amounts to the same thing (a check against {{System.nanoTime()}}), but it's more obvious where the value comes from.\n\n[~jkni] wrt to {{synchronized}}, on one hand agree with you, but then that makes the argument for making many more of the methods on {{Gossiper}} synchronized. If we're going to make this method different from the rest (which is fine with me), can we add in a relevant, detailed comment why it is?","created":"2017-02-25T13:09:57.255+0000"},{"body":"As for the firstSynSendAt timestamp, I was just about to add a code comment why the updateTimestamp comparison would work fine, when I realized that I had to figure out _again_ how this value was initialized and that the getUpdateTimestamp method description was not exactly what I wanted for the use case. So lets just stick with Jasons suggestion for sake of clarity and maintainability.\n\nI've also added some javadoc comments regarding the synchronized use.\n\n\n||2.2||3.0||3.11||trunk||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-2.2]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-3.0]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-3.11]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-trunk]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.11-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.11-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-testall/]|","created":"2017-03-01T08:47:29.960+0000"},{"body":"Nice work, +1. [~jkni] Do you want to give this one more reading, or do you trust [~spodxx@gmail.com] and I not to screw things up too badly? :)","created":"2017-03-01T13:37:27.779+0000"},{"body":"Thanks! The latest changes look good - however, if moving the System.nanoTime() call to the comparison site, it seems that the {{firstSynSendAt}} truly does reduce to a boolean, since the comparison will now always be true if {{firstSynSendAt}} has been set. I don't think the existing patch will cause any problems, but it may be more complicated than it needs to be.","created":"2017-03-02T18:47:38.728+0000"},{"body":"The value might be effectively reduced to a boolean with the latest version, but it doesn't have to stay that way in the future. But I honestly don't really feel that the level of additional complexity we're talking about here is worth further discussion. At this point it's probably more a matter of personal code style preferences. So are you ok keeping this issue as \"ready to commit\", [~jkni]?","created":"2017-03-09T09:47:05.542+0000"},{"body":"Sure - while I'd argue that a need for a change in the future could be introduced in the future patch, I agree that this distinction is very minor and won't cause any problems. Thanks for the patch and your patience!","created":"2017-03-09T15:27:30.507+0000"},{"body":"Which of you two new committers (congrats!) feels like committing this? \n","created":"2017-03-13T05:07:59.526+0000"},{"body":"I planned on leaving that honor to [~spodxx@gmail.com] as patch author, but if he doesn't, I'm happy to do so.","created":"2017-03-13T14:15:55.742+0000"},{"body":"Committed to 2.2 as {{bf0906b92cf65161d828e31bc46436d427bbb4b8}} and merged forward through 3.0, 3.11, and trunk. Added Jason Brown as an additional reviewer in the commit since his feedback was incorporated in the latest round of patches.\n\nThanks everyone!","created":"2017-03-22T21:07:42.213+0000"}],"conversations":[{"body":"Bootstrapping or replacing a node in the cluster requires to gather and check some host IDs or tokens by doing a gossip \"shadow round\" once before joining the cluster. This is done by sending a gossip SYN to all seeds until we receive a response with the cluster state, from where we can move on in the bootstrap process. Receiving a response will call the shadow round done and calls {{Gossiper.resetEndpointStateMap}} for cleaning up the received state again.\n\nThe issue here is that at this point there might be other in-flight requests and it's very likely that shadow round responses from other seeds will be received afterwards, while the current state of the bootstrap process doesn't expect this to happen (e.g. gossiper may or may not be enabled). \n\nOne side effect will be that MigrationTasks are spawned for each shadow round reply except the first. Tasks might or might not execute based on whether at execution time {{Gossiper.resetEndpointStateMap}} had been called, which effects the outcome of {{FailureDetector.instance.isAlive(endpoint))}} at start of the task. You'll see error log messages such as follows when this happend:\n\n{noformat}\nINFO [SharedPool-Worker-1] 2016-09-08 08:36:39,255 Gossiper.java:993 - InetAddress /xx.xx.xx.xx is now UP\nERROR [MigrationStage:1] 2016-09-08 08:36:39,255 FailureDetector.java:223 - unknown endpoint /xx.xx.xx.xx\n{noformat}\n\n\nAlthough is isn't pretty, I currently don't see any serious harm from this, but it would be good to get a second opinion (feel free to close as \"wont fix\").\n\n/cc [~Stefania] [~thobbs]","from":"reporter","subject":"In-flight shadow round requests"},{"body":"You are definitely correct about the cause of this - it occasionally surfaces in test failures such as CASSANDRA-11689. I agree with your conclusion that this doesn't cause any serious harm, but I think it is worth fixing.","from":"developer"},{"body":"It's been a while since I've looked at this, but I agree that we should try to fix it. I think the easiest would be to throw away additional shadow round replies, if we can flag them as such.","from":"developer"},{"body":"I think we can get around changing gossip message values by making sure that we don't process any ACKs for which (any) SYN hasn't been send before by the {{Gossiper}}. I've added {{Gossiper.lastSynSendAt}} that can be used to drop messages in the verb handler.\n","from":"developer"},{"body":"\nThe attached patch will solve this issue by using a separate data structure for the information gathered during shadow rounds and by using a timestamp for the first syn send. \n\n||trunk||3.0||2.2||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-trunk]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-3.0]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-2.2]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-testall/]|\n\n\nI can also create patches for merge conflicts if needed after agreeing on which versions to target.\n","from":"developer"},{"body":"Are you still planing to review the patch, [~jkni]?\n","from":"developer"},{"body":"[~spodxx@gmail.com] - yes! I sincerely apologize for the delay here. If anyone else is interested in reviewing this, they're welcome to pick it up, but it's near the top of my list and I hope to get to this soon.","from":"developer"},{"body":"What's status here [~jkni]?","from":"developer"},{"body":"Thanks for the ping - I gave the patch another skim and have some questions.\n\nThe approach of returning endpoint states obtained through gossip seems sound. I definitely like this idea because it (mostly) prevents us from needing to reason about how shadow rounds affect the proper gossip process. That said, I'm not sure it fully accomplishes this goal. At the moment, it should be safe if a later response comes back after the gossiper has started, but we need to be careful to preserve this property in the future.\n\nI'm not sure how the timestamp check is supposed to work. It initializes the field using System.nanoTime() the first time we send a gossip, but in the gossip digest ack verb handler, we check timestamps using the endpoint state update timestamp, which is not serialized inter-node and also initialized using System.nanoTime() by the local JVM. It seems to me that this reduces to a boolean check that the gossiper has been properly started at least once, since this check will only fail when firstSynSendAt == 0. Am I missing something here? It also seems to me that we should initialize the field on starting the gossiper rather than checking and possibly initializing it every time we send a gossip message.","from":"developer"},{"body":"The firstSynSendAt value is only set while sending the first syn during regular gossip. All ACKs will be ignored that have been received and queued before that point. Please keep in mind that the first SYN is send by GossipTasks by using delayed execution, so it's not enough to simply check if the Gossiper has already been started.\n\nYou are correct in pointing out that the timestamp of the incoming ACK will be a local nanoTime() value. In-flight ACKs triggered during the shadow round will not be filtered by the timestamp check unless already deserialized. But I'd prefer comparing local timestamps anyways, as they will not require some kind of clock skew tolerance, which would be hard to come up with given the kind of races we like to address. Also this shouldn't be a problem as, like you mentioned, the current handling of empty shadow round ACK replies should be safe. But it's not a very clean solution and I'd rather be able to either have a causal relationship to the corresponding SYN/Gossip round or even a custom ShadowGossipMessage type. But this is probably better taken care of in e.g. CASSANDRA-12345 instead of this bug fix.\n","from":"developer"},{"body":"Thanks for the quick response; I agree with all the points in your message. My gut instinct is to make the patch as small as possible since we agree that establishing a causal relationship or explicitly separating the shadow gossip round is the proper long-term solution, but the patch isn't particularly large either way, so I'll move forward with the the patch as proposed.\n\nI'll give the patches another review for any small fixes.","from":"developer"},{"body":"[~jkni] - this still on your radar?\n","from":"developer"},{"body":"[~jjirsa] - yes. That said, I don't anticipate getting to it within the next couple days, so feel free to give it the final review if it is higher priority for you than that. \n\nI've given several passes of review to the core concepts and they seem good. I think a final code style/details pass is all that remains.","from":"developer"},{"body":"[~spodxx@gmail.com] - is there some deeper purpose of moving the {{FD.instance.isAlive()}} check higher in {{MigrationTask#runMayThrow()}} method beyond \"check to see if it's dead before we bother checking to see if it's worth sending a migration task\"? Is there a reason we don't let {{MM#shouldPullSchemaFrom}} return false if FD says the instance is dead? \n\nGiven that the shadow round is meant to just get ring state without changing anything, should we add an explicit check to {{MigrationManager#scheduleSchemaPull()}} to ensure that {{Gossiper.instance.isInShadowRound()}} is false before scheduling?\n","from":"developer"},{"body":"bq. Stefan Podkowinski - is there some deeper purpose of moving the FD.instance.isAlive() check higher in MigrationTask#runMayThrow() method beyond \"check to see if it's dead before we bother checking to see if it's worth sending a migration task\"? Is there a reason we don't let MM#shouldPullSchemaFrom return false if FD says the instance is dead?\n\nWe could move FS.isAlive into MM.shouldPullSchemaFrom, yes. Not totally against it, but the log message in MigrationTask in case of a false return value would have to be changed and actually the isAlive status should only be relevant at task execution, as there's a 60 second delay after submitting it. So in theory you could submit a task for a node that has been dead but will be alive again at time of execution.\n\nbq. Given that the shadow round is meant to just get ring state without changing anything, should we add an explicit check to MigrationManager#scheduleSchemaPull() to ensure that Gossiper.instance.isInShadowRound() is false before scheduling?\n\nThe MigrationManager should never issue a schema pull during shadow round. If we add such check, I'd prefer to throw an exception instead of failing silently, instead of letting the process run in an undefined state. On the other hand, it's not really the business of the MM to monitor the gossiper life-cycle when it comes to separations of concerns.","from":"developer"},{"body":"Fair points all around. I've deployed this as-is into a decent sized test cluster just to see how it behaves - first impressions are good compared to 3.0.10. Would love to see this land in 3.0.11 - if [~jkni] doesn't get to it in time for cutting 3.0.11, I'll try to get a more formal review pass on it.\n","from":"developer"},{"body":"A few questions/comments on the latest patches:\n* On all versions, a space is missing in the conditional {{if(firstSynSendAt == 0)}}\n* On 2.2, it looks like the patch adds the {{valuesEqual}} method from later versions. Was this intentional? It looks unused.\n* On all versions, using {{firstSynSendAt == 0}} to check if it has been initialized isn't entirely safe. It's entirely legal (although admittedly rare) for {{System.nanoTime}} to return 0. If this happened, all acks would be rejected.\n* Comparisons for two {{System.nanoTime}} values (such as in {{GossipDigestAckVerbHandler}}) should not use t1 < t2. Instead, one should check the difference (t1 - t2 < 0) because numerical overflow could occur in the {{System.nanoTime}} long.\n* In {{maybeFinishShadowRound}}/{{finishShadowRound}}, we should add the states to the {{endpointShadowStateMap}} before setting {{inShadowRound}} to false. It looks like the current behavior admits a race where {{doShadowRound}} could read {{inShadowRound == false}} and exit its loop and copy the endpointShadowStateMap before it is filled by shadow round finish.\n* I believe {{firstSynSendAt}} is accessed from multiple threads and needs to be volatile.","from":"developer"},{"body":"Thanks for your comments, Joel! Patch has been rebased and updated by following your suggestions, except from the issue mentioned below. \n\nI've also remove two @VisibleForTesting annotations in 2.2, since the related test isn't available in this branch.\n\nbq. On all versions, using firstSynSendAt == 0 to check if it has been initialized isn't entirely safe. It's entirely legal (although admittedly rare) for System.nanoTime to return 0. If this happened, all acks would be rejected.\n\nIn this very rare case we'd only discard the acks of 1 or 2 nodes from a single gossip round, as the firstSynSendAt value would afterwards be set by the next periodic sendGossip call. I'd therefor prefer to leave the variable as is, beside making it volatile.\n\n\n||2.2||3.0||3.11||trunk||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-2.2]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-3.0]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-3.11]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-trunk]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.11-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.11-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-testall/]|\n\n","from":"developer"},{"body":"On the whole, I'm pretty good with this patch - nice work, Stefan!\n\nI have two nits:\n- the 3.11 and trunk branches have the tests included, but 2.2 and 3.0 do not. I that just an oversight? Can we add the tests to those branches, as well?\n- {{GossipDigestAckVerbHandler#doVerb}}, you get the [following timestamp|https://github.com/spodkowinski/cassandra/commit/9179b7ee06c51a79881f5be18cd01261ebe62143#diff-787d4963a51f20221468e976df1b121aR64]: {code} long ts = epStateMap.values().iterator().next().getUpdateTimestamp();{code}\nWe [do not serialize|https://github.com/apache/cassandra/blob/cassandra-2.2/src/java/org/apache/cassandra/gms/EndpointState.java#L166] {{EndpointState.updateTimestamp}}, so the value at the receiver ends up being the receiver's {{System.nanoTime()}}, as can be seen from the [{{EndpointState}} constructor|https://github.com/apache/cassandra/blob/cassandra-2.2/src/java/org/apache/cassandra/gms/EndpointState.java#L58]. What is it you are looking to confirm here? I'm kinda leaning toward eliminating that check altogether as {code}Gossiper.instance.firstSynSendAt == 0{code} might be a sufficient check anyways.\n\n","from":"developer"},{"body":"Also, I'm not sure {{Gossiper#doShadowRound}} needs to be {{synchronized}}. Currently, that method is only invoked from {{StorageService#prepareToJoin}} (in all current branches). Is it necessary to mark it as such? If it is, can you please add a comment to the method.","from":"developer"},{"body":"I think I can answer these - feel free to correct me, [~spodxx@gmail.com]\n\nIn order,\n* Presently, the tests depend on the mock MessagingService, which was added in [CASSANDRA-12016] to 3.10+. We'd new tests for 2.2/3.0+, which is desirable, but I have no great ideas how to do it other than fiddly byteman tests. \n* I agree with this. Stefan and I discussed it on the first pass of review, and I wouldn't mind eliminating that check altogether and making it a boolean. OTOH, it's cheap to check deserialization time and excludes the messages that were deserialized prior to the check. OTOH, there's no meaningful distinction in correctness-preserving behaviors between that and arbitrarily delayed gossip messages, and we need to handle the latter correctly anyway. I'm most concerned about this check giving future readers false hope :).\n* It also seems to be me that it doesn't presently need to be synchronized. That said, I assumed it was a defensive choice because the internals are definitely not safe to call on multiple threads, and someone may make that mistake in the future.","from":"developer"},{"body":"The goal of introducing the firstSynSendAt timestamp was to prevent processing ACKs (lagged over from shadow gossip round) before any regular SYN has been send. I still think checking this value against the local message construction time is a good idea, as messages will be queued before getting dispatched and it's not that unlikely that the scheduler will make this happen for older messages just after we send the syn. As [mentioned before|https://issues.apache.org/jira/browse/CASSANDRA-12653?focusedCommentId=15818621&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15818621] the proper way to do this would be to be able to correlate between messages or do the shadow round conversation based on new dedicated message types.\n","from":"developer"},{"body":"Do those answers address your questions well enough, [~jasobrown]? The latest patch addressed my concerns, but I don't want to step on your toes.\n\nI had to restart dtests for 2.2, but the latest patch/CI looks good to me otherwise.","from":"developer"},{"body":"[~spod] I'm on board with {{firstSynSendAt}} being a timestamp, I'm just questioning what you compare it against. Instead of this:\n\n{code}\nlong ts = epStateMap.values().iterator().next().getUpdateTimestamp();\nif ((ts - Gossiper.instance.firstSynSendAt) < 0 || Gossiper.instance.firstSynSendAt == 0)\n{code}\n\nI'm suggesting this:\n{code}\nif ((System.nanoTime() - Gossiper.instance.firstSynSendAt) < 0 || Gossiper.instance.firstSynSendAt == 0)\n{code}\n\nIt amounts to the same thing (a check against {{System.nanoTime()}}), but it's more obvious where the value comes from.\n\n[~jkni] wrt to {{synchronized}}, on one hand agree with you, but then that makes the argument for making many more of the methods on {{Gossiper}} synchronized. If we're going to make this method different from the rest (which is fine with me), can we add in a relevant, detailed comment why it is?","from":"developer"},{"body":"As for the firstSynSendAt timestamp, I was just about to add a code comment why the updateTimestamp comparison would work fine, when I realized that I had to figure out _again_ how this value was initialized and that the getUpdateTimestamp method description was not exactly what I wanted for the use case. So lets just stick with Jasons suggestion for sake of clarity and maintainability.\n\nI've also added some javadoc comments regarding the synchronized use.\n\n\n||2.2||3.0||3.11||trunk||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-2.2]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-3.0]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-3.11]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-12653-trunk]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.11-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-2.2-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-3.11-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-12653-trunk-testall/]|","from":"developer"},{"body":"Nice work, +1. [~jkni] Do you want to give this one more reading, or do you trust [~spodxx@gmail.com] and I not to screw things up too badly? :)","from":"developer"},{"body":"Thanks! The latest changes look good - however, if moving the System.nanoTime() call to the comparison site, it seems that the {{firstSynSendAt}} truly does reduce to a boolean, since the comparison will now always be true if {{firstSynSendAt}} has been set. I don't think the existing patch will cause any problems, but it may be more complicated than it needs to be.","from":"developer"},{"body":"The value might be effectively reduced to a boolean with the latest version, but it doesn't have to stay that way in the future. But I honestly don't really feel that the level of additional complexity we're talking about here is worth further discussion. At this point it's probably more a matter of personal code style preferences. So are you ok keeping this issue as \"ready to commit\", [~jkni]?","from":"developer"},{"body":"Sure - while I'd argue that a need for a change in the future could be introduced in the future patch, I agree that this distinction is very minor and won't cause any problems. Thanks for the patch and your patience!","from":"developer"},{"body":"Which of you two new committers (congrats!) feels like committing this? \n","from":"developer"},{"body":"I planned on leaving that honor to [~spodxx@gmail.com] as patch author, but if he doesn't, I'm happy to do so.","from":"developer"},{"body":"Committed to 2.2 as {{bf0906b92cf65161d828e31bc46436d427bbb4b8}} and merged forward through 3.0, 3.11, and trunk. Added Jason Brown as an additional reviewer in the commit since his feedback was incorporated in the latest round of patches.\n\nThanks everyone!","from":"developer"}],"created":"2016-09-16T08:33:09.000+0000","description":"Bootstrapping or replacing a node in the cluster requires to gather and check some host IDs or tokens by doing a gossip \"shadow round\" once before joining the cluster. This is done by sending a gossip SYN to all seeds until we receive a response with the cluster state, from where we can move on in the bootstrap process. Receiving a response will call the shadow round done and calls {{Gossiper.resetEndpointStateMap}} for cleaning up the received state again.\n\nThe issue here is that at this point there might be other in-flight requests and it's very likely that shadow round responses from other seeds will be received afterwards, while the current state of the bootstrap process doesn't expect this to happen (e.g. gossiper may or may not be enabled). \n\nOne side effect will be that MigrationTasks are spawned for each shadow round reply except the first. Tasks might or might not execute based on whether at execution time {{Gossiper.resetEndpointStateMap}} had been called, which effects the outcome of {{FailureDetector.instance.isAlive(endpoint))}} at start of the task. You'll see error log messages such as follows when this happend:\n\n{noformat}\nINFO [SharedPool-Worker-1] 2016-09-08 08:36:39,255 Gossiper.java:993 - InetAddress /xx.xx.xx.xx is now UP\nERROR [MigrationStage:1] 2016-09-08 08:36:39,255 FailureDetector.java:223 - unknown endpoint /xx.xx.xx.xx\n{noformat}\n\n\nAlthough is isn't pretty, I currently don't see any serious harm from this, but it would be good to get a second opinion (feel free to close as \"wont fix\").\n\n/cc [~Stefania] [~thobbs]","issue_id":"13005368","key":"CASSANDRA-12653","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-03-22T21:02:47.000+0000","role":"fixed_distractor","summary":"In-flight shadow round requests"} {"case_id":"13007018","cluster":"DISTRACTOR-CASSANDRA-12694","comments":[{"body":"I think this problem is being caused by this addition:\nhttps://github.com/ifesdjeen/cassandra/commit/ef9225fea660b46ed4905c10b91e7efe2746da5b#diff-c06541855022eca5fd794dd24ff02f89\n\nwhich was added for CASSANDRA-9530.\n\nIn this scenario, I think the CAS check is correctly returning an empty row (the matching row has a null value) which the above change errors out on, before StorageProxy.cas can check if the row applies to the CAS conditions.\n\nI'm not sure what the impact of removing the check is, as the comment indicates that it's also repeated for compactions.","created":"2016-09-28T01:05:23.607+0000"},{"body":"Thanks for the pointer to the commit. Although the problem was not really caused by this addition, it just got surfaced. What we expect there is correct and protects users from reading data from the corrupt SSTable. In the case of an empty row we would have liveness (which is used to distinguish empty from dead row [more details here|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/rows/Row.java#L71-L86](. \n\nWhat happens is that we're following a read path that reads out only the columns that are affected by the LWT (and this column was never set), which results into this check failing, since there's no liveness that would confirm that this row exists but is empty. In fact, we'll run into the same situation if we use \"regular\" (not LWT) update in the first place, like that:\n\n{code}\nupdate test.lwt_corruption_test set last_updated = 555 where test_id = 'test3';\nupdate test.lwt_corruption_test set last_updated = 555 where test_id = 'test3' if message_id = null;\n;; would result into the same error\n{code}\n\nHowever, running the LWT with a column that actually exists would result into successful LWT:\n\n{code}\nupdate test.lwt_corruption_test set last_updated = 555 where test_id = 'test3' if message_id = null and last_updated = 555;\n{code}\n\nNow, {{last_updated}} will be returned and non-empty.\n\nIn order to avoid this problem and make sure we never have a situation where row is empty, we have to fetch enough columns to satisfy the column condition and distinguish between \"all conditions are null\" and \"row does not exist\".\n\nIf we update static row, we won't have any conditions (or updates) on regular rows, so we fetch _all_ static rows.\nIf we update regular rows, we fetch only the static rows that take part in column condition and _all_ regular rows.\n\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/12694-trunk] |[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-trunk-testall/] |[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-trunk-dtest/] |\n","created":"2016-09-30T14:42:58.218+0000"},{"body":"The trunk patch looks OK with some nits, see this [branch|https://github.com/stef1927/cassandra/commits/12694] for a rebased version. I had minor conflicts with CASSANDRA-12060, so you may want to double check your initial commit as well, particularly the comments in {{readCommand()}}: they were removed by this patch but changed by 12060, so I left them there since they seem to make sense.\n\nAs discussed offline, we still need to:\n* check that the logic in {{columnsToRead()}} is aligned with {{readCommand()}}\n* backport to 3.0 and 3.x\n* add a unit test \n\n\n","created":"2016-10-05T08:36:43.836+0000"},{"body":"Actually, while I'm sure the patch fixes the issue, I think it's uselessly inefficient. We shouldn't touch {{CQL3CasRequest.columnsToRead()}}, but we should change {{CQL3CasRequest.readCommand()}}, changing the {{ColumnFilter.selection(columnsToRead())}} to instead use {{ColumnFilter.allColumnsBuilder()}} and add the columns from {{ColumnsToRead()}}. Or alternatively (and probably better), add to {{ColumnFilter}}:\n{noformat}\npublic static ColumnFilter all(CFMetaData metadata, PartitionColumns queried)\n{\n return new ColumnFilter(true, metadata, queried, null);\n}\n{noformat}\nand use that.\n\nI'll refer to the javadoc on {{ColumnFilter}} on why this is different (and better), but for the record, the rule to remember is that {{ColumnFilter.selection()}} and {{ColumnFilter.seelctionBuilder()}} should never be used for CQL code.","created":"2016-10-05T09:08:30.470+0000"},{"body":"Thank you for the comment. \nThe javadoc on {{ColumnFilter}} also says that it's ok to use selection for some internal purposes, and it seems to me that this is exactly the case. \n\nProblem with switching to {{allColumns}} is that we (yet again) are going to break the failure result (which is not a problem on in unit tests, but we'll have even more special-casing on LWTs in dtest code), which was exactly the reason why I've chosen changing {{columnsToRead}} instead. I have tried the approach you describe and have opted out for the current version.\n\nI went through all combinations and cases and having the mentioned conditions satisfies all the requirements and can guarantee that we know whether or not partition exists, both in case with static and partition conditions and does not break the existing behaviour of what we return on CAS failure. We also did not do the same change (query all columns) in [CASSANDRA-12060] and opted out for special-casing static columns with {{LIMIT 1}} in order to make sure that partition exists. ","created":"2016-10-05T11:16:36.925+0000"},{"body":"bq. The javadoc on {{ColumnFilter}} also says that it's ok to use selection for some internal purposes, and it seems to me that this is exactly the case.\n\nWell, as luck as it, I wrote both the class and comment, and this is not a case that the comment meant to cover. The sole purpose of having different \"selection\" and \"all columns\" builders is to distinguish between the cases where you care about the difference between \"the row exists but has null values for every column queried\" and \"the row doesn't exist\" (job for the \"all columns\" builder) and cases where it doesn't matter (job for the \"selection\" builder). Unless I'm misunderstanding this ticket, the problem here is exactly that we want to make the difference between those two cases, so it should use a \"all columns\" builder or you're defeating the whole purpose of the {{queried}} versus {{fetched}} nuance {{ColumnFilter}} makes.\n\nbq. Problem with switching to {{allColumns}} is that we (yet again) are going to break the failure result\n\nIf you could be more specific on what you mean here, that would be helpful. Is the problem specific to static columns maybe?\n\nbq. We also did not do the same change (query all columns) in CASSANDRA-12060\n\nYou're mixing different things. CASSANDRA-12060 was about checking partition existence. Which columns you read impact row existence.","created":"2016-10-05T13:32:12.970+0000"},{"body":"re:12060, I meant that we could achieve similar results by using {{all}} and without special-casing static columns (modulo output, just as we have here now). Might be our tests do not reflect all reality, but with current behaviour I haven't found cases where it'd fail.\n\nbq. Is the problem specific to static columns maybe?\n\nUnfortunately, not. I've pushed what you have suggested [here|https://github.com/ifesdjeen/cassandra/tree/12694-reviewed] (tests reflect some changes, most likely there will be more in dtests, with three-way differences between versions). \nI trust your intuition, and most likely you're right: if there are no guarantees to make the distinctions, we should always use {{all}}. This also will be future proof, in case there's some other edge case hiding somewhere. I just wanted to bring up the behaviour/output change to your attention.","created":"2016-10-05T14:32:04.528+0000"},{"body":"I'll note that my suggestion was a tiny bit different: I didn't meant to remove the {{columsToRead()}} entirely, I mean to use the variation of the {{ColumnFilter.all()}} function of my earlier comment which still takes the columns to read. The difference with using {{ColumnFilter.selection()}} using the terminology in {{ColumnFilter}} is that I want to _fetch_ all columns but only _query_ the one from {{columnsToRead()}}, while {{ColumnFilter.selection()}} only _fetch_ the columns provided (and the existing {{ColumnFilter.all()}} as used in your modification _fetch_ and _query_ all). I believe this is why you have the different output (but if it's not, I'll need to have a closer look).","created":"2016-10-05T15:10:14.331+0000"},{"body":"bq. I want to fetch all columns but only query the one from columnsToRead()\n\nIn both cases ({{all(cfm)}} and {{all(cfm, columns)}}), the output is similar, with several exceptions (for example, when only static columns are used in condition or only regular columns are used: in these cases we will return only them). I've added more tests for such behaviour.\nAlthough after looking at it again I think that new output is better/more correct, as we do have a partition and now the output corresponds to that fact (in case with {{NOT EXISTS}} in tests. \n\nYou're right that it's better to avoid using {{selection}}, and example with {{NOT EXISTS}} kind of proves it. As with {{selection}} the output was as if partition did not exist at all, but it did exist, even though all the rows were deleted.\n\nIf you think this is ok, I'll rebase the other versions, too.\n\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/12694-reviewed]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-reviewed-dtest/]|[testall|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-reviewed-testall/]|","created":"2016-10-06T08:07:36.711+0000"},{"body":"Had a closer look at this. But first, I don't really agree that the new outputs from test are better. The case in question deletes a row if it exists. As the condition is entirely based on the row existence, I don't think making the output change based on partition-related information is necessary meaningful. In fact, we shouldn't *have* to read static content to even decide if a row exists or not.\n\nIn all fairness, we never carefully defined what the CAS result set output contained, so what's better or not is debatable, but I think we should preserver the output for 2 main reasons: 1) to not break backward compatibility (we shouldn't unless we think a behavior was really a bug and I don't think this can be said) and 2) forcing to query static content when the existence is on a row is inefficient and I'd rather not back that inefficient in.\n\nAnyway, the reason we get this output difference is actually the genuine bug/inefficiency from CASSANDRA-12768: we have a condition on a row, so no reason to read any static content.\n","created":"2016-10-10T14:25:50.585+0000"},{"body":"I'm able to reproduce this on 3.9 as well","created":"2016-11-17T12:51:45.174+0000"},{"body":"And 3.0.10, 2.2.8 seems to work though...","created":"2016-11-17T13:05:41.737+0000"},{"body":"As [CASSANDRA-12768], I've updated the patch. It now became much smaller, as we rely on fetch all/query only selected columns distinction in {{ColumnFilter}}. \n\n|[3.0|https://github.com/ifesdjeen/cassandra/tree/12694-3.0]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-3.0-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-3.0-dtest/]|\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/12694-trunk]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-trunk-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-trunk-dtest/]|\n|[dtest patch|https://github.com/riptano/cassandra-dtest/pull/1403]|\n\n","created":"2016-12-08T13:57:49.317+0000"},{"body":"Alright, updated patch lgtm, so committed, thanks.","created":"2016-12-08T16:30:20.059+0000"},{"body":"This commit appears to be a bug fix for 3.10 and has not been merged to the cassandra-3.11 branch, where 3.10 will be released from. Please merge. I've updated fixver to 3.10.\n\nThis commit also has a typo in commit message for the JIRA number, and contained no CHANGES.txt entry, as far as I can tell.\n\n{noformat}\n(cassandra-3.11)mshuler@hana:~/git/cassandra$ git log ..cassandra-3.0\ncommit d9b06e8af41c42244f76058641aeecda53a9bf75 (origin/cassandra-3.0, cassandra-3.0)\nAuthor: Alex Petrov \nDate: Wed Dec 7 16:04:51 2016 +0100\n\n Make distinction between unset row and non-existing partition for LWTs\n \n Patch by Alex Petrov; reviewed by Sylvain Lebresne for CASSANDRA-12964.\n\ncommit 3aefe356545bcc2f5c4453e0d9d704bef339a008 (github/cassandra-3.0)\nMerge: 95d0b67 0b97c5d\nAuthor: Robert Stupp \nDate: Wed Dec 7 13:08:35 2016 +0100\n\n Merge branch 'cassandra-2.2' into cassandra-3.0\n\ncommit 0b97c5d1f717be30a04c59c766465d9c62a4e9ee (origin/cassandra-2.2, github/cassandra-2.2, cassandra-2.2)\nAuthor: Robert Stupp \nDate: Wed Dec 7 13:06:55 2016 +0100\n\n testall failure in org.apache.cassandra.cql3.validation.entities.UFTest.testAllNativeTypes\n \n patch by Robert Stupp; reviewed by Alex Petrov for CASSANDRA-12817\n\ncommit 95d0b671d1af154eaf1c1e81992c7f3f51469eee\nAuthor: Sylvain Lebresne \nDate: Mon Oct 10 13:36:15 2016 +0200\n\n CQL often queries static columns unnecessarily\n \n patch by Sylvain Lebresne; reviewed by Tyler Hobbs for CASSANDRA-12768\n{noformat}","created":"2016-12-08T18:07:51.916+0000"},{"body":"Merged to the misnamed branch.","created":"2016-12-09T10:37:55.962+0000"},{"body":"Was the issue resolved in 3.10, or 3.11? Do I have to do anything special to fix a table after upgrading to 3.10? I'm still hitting this error when trying to update a static column with LWT.","created":"2017-07-19T19:16:32.856+0000"},{"body":"It appears to be [here|https://github.com/apache/cassandra/commit/1e067746e432dc0a450ad111a8ec545011bb5bc7] , which seems to be 3.0.10, 3.10, 3.11.0\n\nIt's still missing the CHANGES, and the JIRA is typo'd in it, but it appears to be in 3.10\n\nIf you're still seeing it in 3.10, can you please paste the full stack trace? \n\n\n","created":"2017-07-19T20:32:08.476+0000"},{"body":"Here's the stacktrace I saw with 3.10:\n\n{code}\nERROR [ReadRepairStage:9] 2017-07-19 18:51:32,607 CassandraDaemon.java:229 - Exception in thread Thread[ReadRepairStage:9,5,main]\njava.io.IOError: java.io.IOException: Corrupt empty row found in unfiltered partition\n\tat org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer$1.computeNext(UnfilteredRowIteratorSerializer.java:227) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer$1.computeNext(UnfilteredRowIteratorSerializer.java:215) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIterators.digest(UnfilteredRowIterators.java:178) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.partitions.UnfilteredPartitionIterators.digest(UnfilteredPartitionIterators.java:270) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.ReadResponse.makeDigest(ReadResponse.java:98) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.ReadResponse$DataResponse.digest(ReadResponse.java:203) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.service.DigestResolver.compareResponses(DigestResolver.java:87) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.service.ReadCallback$AsyncRepairRunner.run(ReadCallback.java:233) ~[apache-cassandra-3.10.jar:3.10]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) ~[na:1.8.0_91]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) ~[na:1.8.0_91]\n\tat org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:79) ~[apache-cassandra-3.10.jar:3.10]\n\tat java.lang.Thread.run(Thread.java:745) ~[na:1.8.0_91]\nCaused by: java.io.IOException: Corrupt empty row found in unfiltered partition\n\tat org.apache.cassandra.db.rows.UnfilteredSerializer.deserialize(UnfilteredSerializer.java:445) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer$1.computeNext(UnfilteredRowIteratorSerializer.java:222) ~[apache-cassandra-3.10.jar:3.10]\n\t... 12 common frames omitted\n{code}\n\nHowever, I upgraded to 3.11, and the issue seems to have gone away.","created":"2017-07-19T20:35:07.156+0000"}],"conversations":[{"body":"{noformat}\ncqlsh> create table test.test (test_id TEXT, last_updated TIMESTAMP, message_id TEXT, PRIMARY KEY(test_id));\nupdate test.test set last_updated = 1474494363669 where test_id = 'test1' if message_id = null;\n{noformat}\n\nThen nodetool flush on the all 3 nodes.\n\n{noformat}\ncqlsh> update test.test set last_updated = 1474494363669 where test_id = 'test1' if message_id = null;\nServerError: \n{noformat}\n\nFrom cassandra log\n{noformat}\nERROR [SharedPool-Worker-1] 2016-09-23 12:09:13,179 Message.java:611 - Unexpected exception during request; channel = [id: 0x7a22599e, L:/127.0.0.1:9042 - R:/127.0.0.1:58297]\njava.io.IOError: java.io.IOException: Corrupt empty row found in unfiltered partition\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer$1.computeNext(UnfilteredRowIteratorSerializer.java:224) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer$1.computeNext(UnfilteredRowIteratorSerializer.java:212) ~[main/:na]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIterators.digest(UnfilteredRowIterators.java:125) ~[main/:na]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators.digest(UnfilteredPartitionIterators.java:249) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse.makeDigest(ReadResponse.java:87) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$DataResponse.digest(ReadResponse.java:192) ~[main/:na]\n at org.apache.cassandra.service.DigestResolver.resolve(DigestResolver.java:80) ~[main/:na]\n at org.apache.cassandra.service.ReadCallback.get(ReadCallback.java:139) ~[main/:na]\n at org.apache.cassandra.service.AbstractReadExecutor.get(AbstractReadExecutor.java:145) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$SinglePartitionReadLifecycle.awaitResultsAndRetryOnDigestMismatch(StorageProxy.java:1714) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.fetchRows(StorageProxy.java:1663) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.readRegular(StorageProxy.java:1604) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.read(StorageProxy.java:1523) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.readOne(StorageProxy.java:1497) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.readOne(StorageProxy.java:1491) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.cas(StorageProxy.java:249) ~[main/:na]\n at org.apache.cassandra.cql3.statements.ModificationStatement.executeWithCondition(ModificationStatement.java:441) ~[main/:na]\n at org.apache.cassandra.cql3.statements.ModificationStatement.execute(ModificationStatement.java:416) ~[main/:na]\n at org.apache.cassandra.cql3.QueryProcessor.processStatement(QueryProcessor.java:208) ~[main/:na]\n at org.apache.cassandra.cql3.QueryProcessor.process(QueryProcessor.java:239) ~[main/:na]\n at org.apache.cassandra.cql3.QueryProcessor.process(QueryProcessor.java:224) ~[main/:na]\n at org.apache.cassandra.transport.messages.QueryMessage.execute(QueryMessage.java:115) ~[main/:na]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:507) [main/:na]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:401) [main/:na]\n{noformat}","from":"reporter","subject":"PAXOS Update Corrupted empty row exception"},{"body":"I think this problem is being caused by this addition:\nhttps://github.com/ifesdjeen/cassandra/commit/ef9225fea660b46ed4905c10b91e7efe2746da5b#diff-c06541855022eca5fd794dd24ff02f89\n\nwhich was added for CASSANDRA-9530.\n\nIn this scenario, I think the CAS check is correctly returning an empty row (the matching row has a null value) which the above change errors out on, before StorageProxy.cas can check if the row applies to the CAS conditions.\n\nI'm not sure what the impact of removing the check is, as the comment indicates that it's also repeated for compactions.","from":"developer"},{"body":"Thanks for the pointer to the commit. Although the problem was not really caused by this addition, it just got surfaced. What we expect there is correct and protects users from reading data from the corrupt SSTable. In the case of an empty row we would have liveness (which is used to distinguish empty from dead row [more details here|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/rows/Row.java#L71-L86](. \n\nWhat happens is that we're following a read path that reads out only the columns that are affected by the LWT (and this column was never set), which results into this check failing, since there's no liveness that would confirm that this row exists but is empty. In fact, we'll run into the same situation if we use \"regular\" (not LWT) update in the first place, like that:\n\n{code}\nupdate test.lwt_corruption_test set last_updated = 555 where test_id = 'test3';\nupdate test.lwt_corruption_test set last_updated = 555 where test_id = 'test3' if message_id = null;\n;; would result into the same error\n{code}\n\nHowever, running the LWT with a column that actually exists would result into successful LWT:\n\n{code}\nupdate test.lwt_corruption_test set last_updated = 555 where test_id = 'test3' if message_id = null and last_updated = 555;\n{code}\n\nNow, {{last_updated}} will be returned and non-empty.\n\nIn order to avoid this problem and make sure we never have a situation where row is empty, we have to fetch enough columns to satisfy the column condition and distinguish between \"all conditions are null\" and \"row does not exist\".\n\nIf we update static row, we won't have any conditions (or updates) on regular rows, so we fetch _all_ static rows.\nIf we update regular rows, we fetch only the static rows that take part in column condition and _all_ regular rows.\n\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/12694-trunk] |[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-trunk-testall/] |[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-trunk-dtest/] |\n","from":"developer"},{"body":"The trunk patch looks OK with some nits, see this [branch|https://github.com/stef1927/cassandra/commits/12694] for a rebased version. I had minor conflicts with CASSANDRA-12060, so you may want to double check your initial commit as well, particularly the comments in {{readCommand()}}: they were removed by this patch but changed by 12060, so I left them there since they seem to make sense.\n\nAs discussed offline, we still need to:\n* check that the logic in {{columnsToRead()}} is aligned with {{readCommand()}}\n* backport to 3.0 and 3.x\n* add a unit test \n\n\n","from":"developer"},{"body":"Actually, while I'm sure the patch fixes the issue, I think it's uselessly inefficient. We shouldn't touch {{CQL3CasRequest.columnsToRead()}}, but we should change {{CQL3CasRequest.readCommand()}}, changing the {{ColumnFilter.selection(columnsToRead())}} to instead use {{ColumnFilter.allColumnsBuilder()}} and add the columns from {{ColumnsToRead()}}. Or alternatively (and probably better), add to {{ColumnFilter}}:\n{noformat}\npublic static ColumnFilter all(CFMetaData metadata, PartitionColumns queried)\n{\n return new ColumnFilter(true, metadata, queried, null);\n}\n{noformat}\nand use that.\n\nI'll refer to the javadoc on {{ColumnFilter}} on why this is different (and better), but for the record, the rule to remember is that {{ColumnFilter.selection()}} and {{ColumnFilter.seelctionBuilder()}} should never be used for CQL code.","from":"developer"},{"body":"Thank you for the comment. \nThe javadoc on {{ColumnFilter}} also says that it's ok to use selection for some internal purposes, and it seems to me that this is exactly the case. \n\nProblem with switching to {{allColumns}} is that we (yet again) are going to break the failure result (which is not a problem on in unit tests, but we'll have even more special-casing on LWTs in dtest code), which was exactly the reason why I've chosen changing {{columnsToRead}} instead. I have tried the approach you describe and have opted out for the current version.\n\nI went through all combinations and cases and having the mentioned conditions satisfies all the requirements and can guarantee that we know whether or not partition exists, both in case with static and partition conditions and does not break the existing behaviour of what we return on CAS failure. We also did not do the same change (query all columns) in [CASSANDRA-12060] and opted out for special-casing static columns with {{LIMIT 1}} in order to make sure that partition exists. ","from":"developer"},{"body":"bq. The javadoc on {{ColumnFilter}} also says that it's ok to use selection for some internal purposes, and it seems to me that this is exactly the case.\n\nWell, as luck as it, I wrote both the class and comment, and this is not a case that the comment meant to cover. The sole purpose of having different \"selection\" and \"all columns\" builders is to distinguish between the cases where you care about the difference between \"the row exists but has null values for every column queried\" and \"the row doesn't exist\" (job for the \"all columns\" builder) and cases where it doesn't matter (job for the \"selection\" builder). Unless I'm misunderstanding this ticket, the problem here is exactly that we want to make the difference between those two cases, so it should use a \"all columns\" builder or you're defeating the whole purpose of the {{queried}} versus {{fetched}} nuance {{ColumnFilter}} makes.\n\nbq. Problem with switching to {{allColumns}} is that we (yet again) are going to break the failure result\n\nIf you could be more specific on what you mean here, that would be helpful. Is the problem specific to static columns maybe?\n\nbq. We also did not do the same change (query all columns) in CASSANDRA-12060\n\nYou're mixing different things. CASSANDRA-12060 was about checking partition existence. Which columns you read impact row existence.","from":"developer"},{"body":"re:12060, I meant that we could achieve similar results by using {{all}} and without special-casing static columns (modulo output, just as we have here now). Might be our tests do not reflect all reality, but with current behaviour I haven't found cases where it'd fail.\n\nbq. Is the problem specific to static columns maybe?\n\nUnfortunately, not. I've pushed what you have suggested [here|https://github.com/ifesdjeen/cassandra/tree/12694-reviewed] (tests reflect some changes, most likely there will be more in dtests, with three-way differences between versions). \nI trust your intuition, and most likely you're right: if there are no guarantees to make the distinctions, we should always use {{all}}. This also will be future proof, in case there's some other edge case hiding somewhere. I just wanted to bring up the behaviour/output change to your attention.","from":"developer"},{"body":"I'll note that my suggestion was a tiny bit different: I didn't meant to remove the {{columsToRead()}} entirely, I mean to use the variation of the {{ColumnFilter.all()}} function of my earlier comment which still takes the columns to read. The difference with using {{ColumnFilter.selection()}} using the terminology in {{ColumnFilter}} is that I want to _fetch_ all columns but only _query_ the one from {{columnsToRead()}}, while {{ColumnFilter.selection()}} only _fetch_ the columns provided (and the existing {{ColumnFilter.all()}} as used in your modification _fetch_ and _query_ all). I believe this is why you have the different output (but if it's not, I'll need to have a closer look).","from":"developer"},{"body":"bq. I want to fetch all columns but only query the one from columnsToRead()\n\nIn both cases ({{all(cfm)}} and {{all(cfm, columns)}}), the output is similar, with several exceptions (for example, when only static columns are used in condition or only regular columns are used: in these cases we will return only them). I've added more tests for such behaviour.\nAlthough after looking at it again I think that new output is better/more correct, as we do have a partition and now the output corresponds to that fact (in case with {{NOT EXISTS}} in tests. \n\nYou're right that it's better to avoid using {{selection}}, and example with {{NOT EXISTS}} kind of proves it. As with {{selection}} the output was as if partition did not exist at all, but it did exist, even though all the rows were deleted.\n\nIf you think this is ok, I'll rebase the other versions, too.\n\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/12694-reviewed]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-reviewed-dtest/]|[testall|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-reviewed-testall/]|","from":"developer"},{"body":"Had a closer look at this. But first, I don't really agree that the new outputs from test are better. The case in question deletes a row if it exists. As the condition is entirely based on the row existence, I don't think making the output change based on partition-related information is necessary meaningful. In fact, we shouldn't *have* to read static content to even decide if a row exists or not.\n\nIn all fairness, we never carefully defined what the CAS result set output contained, so what's better or not is debatable, but I think we should preserver the output for 2 main reasons: 1) to not break backward compatibility (we shouldn't unless we think a behavior was really a bug and I don't think this can be said) and 2) forcing to query static content when the existence is on a row is inefficient and I'd rather not back that inefficient in.\n\nAnyway, the reason we get this output difference is actually the genuine bug/inefficiency from CASSANDRA-12768: we have a condition on a row, so no reason to read any static content.\n","from":"developer"},{"body":"I'm able to reproduce this on 3.9 as well","from":"developer"},{"body":"And 3.0.10, 2.2.8 seems to work though...","from":"developer"},{"body":"As [CASSANDRA-12768], I've updated the patch. It now became much smaller, as we rely on fetch all/query only selected columns distinction in {{ColumnFilter}}. \n\n|[3.0|https://github.com/ifesdjeen/cassandra/tree/12694-3.0]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-3.0-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-3.0-dtest/]|\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/12694-trunk]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-trunk-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12694-trunk-dtest/]|\n|[dtest patch|https://github.com/riptano/cassandra-dtest/pull/1403]|\n\n","from":"developer"},{"body":"Alright, updated patch lgtm, so committed, thanks.","from":"developer"},{"body":"This commit appears to be a bug fix for 3.10 and has not been merged to the cassandra-3.11 branch, where 3.10 will be released from. Please merge. I've updated fixver to 3.10.\n\nThis commit also has a typo in commit message for the JIRA number, and contained no CHANGES.txt entry, as far as I can tell.\n\n{noformat}\n(cassandra-3.11)mshuler@hana:~/git/cassandra$ git log ..cassandra-3.0\ncommit d9b06e8af41c42244f76058641aeecda53a9bf75 (origin/cassandra-3.0, cassandra-3.0)\nAuthor: Alex Petrov \nDate: Wed Dec 7 16:04:51 2016 +0100\n\n Make distinction between unset row and non-existing partition for LWTs\n \n Patch by Alex Petrov; reviewed by Sylvain Lebresne for CASSANDRA-12964.\n\ncommit 3aefe356545bcc2f5c4453e0d9d704bef339a008 (github/cassandra-3.0)\nMerge: 95d0b67 0b97c5d\nAuthor: Robert Stupp \nDate: Wed Dec 7 13:08:35 2016 +0100\n\n Merge branch 'cassandra-2.2' into cassandra-3.0\n\ncommit 0b97c5d1f717be30a04c59c766465d9c62a4e9ee (origin/cassandra-2.2, github/cassandra-2.2, cassandra-2.2)\nAuthor: Robert Stupp \nDate: Wed Dec 7 13:06:55 2016 +0100\n\n testall failure in org.apache.cassandra.cql3.validation.entities.UFTest.testAllNativeTypes\n \n patch by Robert Stupp; reviewed by Alex Petrov for CASSANDRA-12817\n\ncommit 95d0b671d1af154eaf1c1e81992c7f3f51469eee\nAuthor: Sylvain Lebresne \nDate: Mon Oct 10 13:36:15 2016 +0200\n\n CQL often queries static columns unnecessarily\n \n patch by Sylvain Lebresne; reviewed by Tyler Hobbs for CASSANDRA-12768\n{noformat}","from":"developer"},{"body":"Merged to the misnamed branch.","from":"developer"},{"body":"Was the issue resolved in 3.10, or 3.11? Do I have to do anything special to fix a table after upgrading to 3.10? I'm still hitting this error when trying to update a static column with LWT.","from":"developer"},{"body":"It appears to be [here|https://github.com/apache/cassandra/commit/1e067746e432dc0a450ad111a8ec545011bb5bc7] , which seems to be 3.0.10, 3.10, 3.11.0\n\nIt's still missing the CHANGES, and the JIRA is typo'd in it, but it appears to be in 3.10\n\nIf you're still seeing it in 3.10, can you please paste the full stack trace? \n\n\n","from":"developer"},{"body":"Here's the stacktrace I saw with 3.10:\n\n{code}\nERROR [ReadRepairStage:9] 2017-07-19 18:51:32,607 CassandraDaemon.java:229 - Exception in thread Thread[ReadRepairStage:9,5,main]\njava.io.IOError: java.io.IOException: Corrupt empty row found in unfiltered partition\n\tat org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer$1.computeNext(UnfilteredRowIteratorSerializer.java:227) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer$1.computeNext(UnfilteredRowIteratorSerializer.java:215) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIterators.digest(UnfilteredRowIterators.java:178) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.partitions.UnfilteredPartitionIterators.digest(UnfilteredPartitionIterators.java:270) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.ReadResponse.makeDigest(ReadResponse.java:98) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.ReadResponse$DataResponse.digest(ReadResponse.java:203) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.service.DigestResolver.compareResponses(DigestResolver.java:87) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.service.ReadCallback$AsyncRepairRunner.run(ReadCallback.java:233) ~[apache-cassandra-3.10.jar:3.10]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) ~[na:1.8.0_91]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) ~[na:1.8.0_91]\n\tat org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:79) ~[apache-cassandra-3.10.jar:3.10]\n\tat java.lang.Thread.run(Thread.java:745) ~[na:1.8.0_91]\nCaused by: java.io.IOException: Corrupt empty row found in unfiltered partition\n\tat org.apache.cassandra.db.rows.UnfilteredSerializer.deserialize(UnfilteredSerializer.java:445) ~[apache-cassandra-3.10.jar:3.10]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer$1.computeNext(UnfilteredRowIteratorSerializer.java:222) ~[apache-cassandra-3.10.jar:3.10]\n\t... 12 common frames omitted\n{code}\n\nHowever, I upgraded to 3.11, and the issue seems to have gone away.","from":"developer"}],"created":"2016-09-23T02:12:06.000+0000","description":"{noformat}\ncqlsh> create table test.test (test_id TEXT, last_updated TIMESTAMP, message_id TEXT, PRIMARY KEY(test_id));\nupdate test.test set last_updated = 1474494363669 where test_id = 'test1' if message_id = null;\n{noformat}\n\nThen nodetool flush on the all 3 nodes.\n\n{noformat}\ncqlsh> update test.test set last_updated = 1474494363669 where test_id = 'test1' if message_id = null;\nServerError: \n{noformat}\n\nFrom cassandra log\n{noformat}\nERROR [SharedPool-Worker-1] 2016-09-23 12:09:13,179 Message.java:611 - Unexpected exception during request; channel = [id: 0x7a22599e, L:/127.0.0.1:9042 - R:/127.0.0.1:58297]\njava.io.IOError: java.io.IOException: Corrupt empty row found in unfiltered partition\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer$1.computeNext(UnfilteredRowIteratorSerializer.java:224) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer$1.computeNext(UnfilteredRowIteratorSerializer.java:212) ~[main/:na]\n at org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIterators.digest(UnfilteredRowIterators.java:125) ~[main/:na]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators.digest(UnfilteredPartitionIterators.java:249) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse.makeDigest(ReadResponse.java:87) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$DataResponse.digest(ReadResponse.java:192) ~[main/:na]\n at org.apache.cassandra.service.DigestResolver.resolve(DigestResolver.java:80) ~[main/:na]\n at org.apache.cassandra.service.ReadCallback.get(ReadCallback.java:139) ~[main/:na]\n at org.apache.cassandra.service.AbstractReadExecutor.get(AbstractReadExecutor.java:145) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$SinglePartitionReadLifecycle.awaitResultsAndRetryOnDigestMismatch(StorageProxy.java:1714) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.fetchRows(StorageProxy.java:1663) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.readRegular(StorageProxy.java:1604) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.read(StorageProxy.java:1523) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.readOne(StorageProxy.java:1497) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.readOne(StorageProxy.java:1491) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy.cas(StorageProxy.java:249) ~[main/:na]\n at org.apache.cassandra.cql3.statements.ModificationStatement.executeWithCondition(ModificationStatement.java:441) ~[main/:na]\n at org.apache.cassandra.cql3.statements.ModificationStatement.execute(ModificationStatement.java:416) ~[main/:na]\n at org.apache.cassandra.cql3.QueryProcessor.processStatement(QueryProcessor.java:208) ~[main/:na]\n at org.apache.cassandra.cql3.QueryProcessor.process(QueryProcessor.java:239) ~[main/:na]\n at org.apache.cassandra.cql3.QueryProcessor.process(QueryProcessor.java:224) ~[main/:na]\n at org.apache.cassandra.transport.messages.QueryMessage.execute(QueryMessage.java:115) ~[main/:na]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:507) [main/:na]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:401) [main/:na]\n{noformat}","issue_id":"13007018","key":"CASSANDRA-12694","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-12-09T10:37:55.000+0000","role":"fixed_distractor","summary":"PAXOS Update Corrupted empty row exception"} {"case_id":"13008887","cluster":"DISTRACTOR-CASSANDRA-12734","comments":[{"body":"[~ifesdjeen] I think we also need CREATE MV WITH ID for this to work.","created":"2016-10-03T15:06:50.722+0000"},{"body":"Hello,\r\n\r\nAny news on this bug ? Is someone working on it ?\r\n\r\nSincerely,\r\nJean-Sébastien","created":"2019-04-30T08:00:23.187+0000"},{"body":"\r\n{code:java}\r\ncqlsh -e \"DESCRIBE SCHEMA\" > schema.cql\r\n{code}\r\n\r\nIs able to export the creation command for the dematerialized view. One can thus create the right tables prior importing the snapshots.","created":"2019-05-09T02:05:33.221+0000"},{"body":"Zhao has created a patch applicable to 3.0 and 3.11 which I ported. \r\n||[3.0|https://github.com/ekaterinadimitrova2/cassandra/pull/new/12734-3.0]||[3.11|https://github.com/ekaterinadimitrova2/cassandra/pull/new/12734-3.11]||\r\n\r\nHe also added two tests which I ported to [4.0|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-4.0?expand=1] where the issue was already fixed before but the new tests revealed issues not related to the issue we are trying to fix as part of this ticket. \r\n\r\nThe tests showed 4.0 regression and change of behavior, not related to this issue though. \r\n\r\n1) The clustering order is always ASC when we create a materialized view, even when we want it explicitly DESC.\r\n\r\n2) While both 3.0 and 3.11 will produce the same [expectedViewSnapshot|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-3.11?expand=1#diff-41e29794572dd219fb3b5b15b0fccfe2d0b92f6eeb0b84a91b3ee22bacd83d61R740-R744], In 4.0 this will look in the following way:\r\n{code:java}\r\nCREATE MATERIALIZED VIEW IF NOT EXISTS %s.%s AS\r\n SELECT pk2, pk1, ck2, ck1, reg1, reg2\r\n FROM %s.%s\r\n WHERE pk2 IS NOT NULL AND pk1 IS NOT NULL AND ck2 IS NOT NULL AND ck1 IS NOT NULL\r\n PRIMARY KEY ((pk2, pk1), ck2, ck1)\r\n WITH ID = %s\r\n AND CLUSTERING ORDER BY (ck2 ASC, ck1 ASC){code}\r\nProbably worth it to open separate ticket for Cassandra 4, the code was refactored there, instead of fixing everything as part of this one.\r\n\r\nCC [~blerer] as I know he was also looking into this ticket. \r\n\r\n \r\n\r\n ","created":"2021-08-24T21:59:05.467+0000"},{"body":"[~e.dimitrova] I ran some tests on 4.0 and hit another issue related to DESCRIBE. I need to dig a deeper. ","created":"2021-08-25T08:42:06.032+0000"},{"body":"The underlying issue on 4.0 is now fixed (see CASSANDRA-16898). The patch for 3.0 and 3.11 LGTM. We probably just need to make a patch for 4.0 and trunk with only the unit tests. ","created":"2021-09-01T12:00:17.372+0000"},{"body":"Thank you for raising CASSANDRA-16898 and the quick fix.\r\n\r\nTests added to [4.0|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-4.0?expand=1] and [trunk|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-trunk?expand=1].\r\n\r\n I am running in a loop the two new tests [here|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra?branch=12734-4.0].\r\n\r\nI also rebased and pushed full CI Jenkins run for the other two branches:\r\n|| Patch || CI ||\r\n|[3.0|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-3.0?expand=1]|[CI|https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1094/]|\r\n|[3.11|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-3.11?expand=1]|[CI|https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1095/]|\r\n\r\n ","created":"2021-09-02T18:45:38.755+0000"},{"body":"The patches look good to me. Thanks :-) ","created":"2021-09-03T08:02:36.806+0000"},{"body":"To https://github.com/apache/cassandra.git\r\n\r\n   e4b37c3271..67eb22ec9d  cassandra-3.0 -> cassandra-3.0\r\n\r\n   957c6264ef..d6e1c41c48  cassandra-3.11 -> cassandra-3.11\r\n\r\n   6a4a93a808..49e83027e2  cassandra-4.0 -> cassandra-4.0\r\n\r\n   f9aa19e3b1..163a4d7137  trunk -> trunk\r\n\r\n \r\n\r\nCI also looked ok, committed, thanks!","created":"2021-09-03T15:03:41.382+0000"}],"conversations":[{"body":"The materialized view schema file that gets created and stored with the sstables is created as a table instead of a materialized view. \n\nCan the materialized view be created and added to the corresponding table's schema file?\n","from":"reporter","subject":"Materialized View schema file for snapshots created as tables"},{"body":"[~ifesdjeen] I think we also need CREATE MV WITH ID for this to work.","from":"developer"},{"body":"Hello,\r\n\r\nAny news on this bug ? Is someone working on it ?\r\n\r\nSincerely,\r\nJean-Sébastien","from":"developer"},{"body":"\r\n{code:java}\r\ncqlsh -e \"DESCRIBE SCHEMA\" > schema.cql\r\n{code}\r\n\r\nIs able to export the creation command for the dematerialized view. One can thus create the right tables prior importing the snapshots.","from":"developer"},{"body":"Zhao has created a patch applicable to 3.0 and 3.11 which I ported. \r\n||[3.0|https://github.com/ekaterinadimitrova2/cassandra/pull/new/12734-3.0]||[3.11|https://github.com/ekaterinadimitrova2/cassandra/pull/new/12734-3.11]||\r\n\r\nHe also added two tests which I ported to [4.0|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-4.0?expand=1] where the issue was already fixed before but the new tests revealed issues not related to the issue we are trying to fix as part of this ticket. \r\n\r\nThe tests showed 4.0 regression and change of behavior, not related to this issue though. \r\n\r\n1) The clustering order is always ASC when we create a materialized view, even when we want it explicitly DESC.\r\n\r\n2) While both 3.0 and 3.11 will produce the same [expectedViewSnapshot|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-3.11?expand=1#diff-41e29794572dd219fb3b5b15b0fccfe2d0b92f6eeb0b84a91b3ee22bacd83d61R740-R744], In 4.0 this will look in the following way:\r\n{code:java}\r\nCREATE MATERIALIZED VIEW IF NOT EXISTS %s.%s AS\r\n SELECT pk2, pk1, ck2, ck1, reg1, reg2\r\n FROM %s.%s\r\n WHERE pk2 IS NOT NULL AND pk1 IS NOT NULL AND ck2 IS NOT NULL AND ck1 IS NOT NULL\r\n PRIMARY KEY ((pk2, pk1), ck2, ck1)\r\n WITH ID = %s\r\n AND CLUSTERING ORDER BY (ck2 ASC, ck1 ASC){code}\r\nProbably worth it to open separate ticket for Cassandra 4, the code was refactored there, instead of fixing everything as part of this one.\r\n\r\nCC [~blerer] as I know he was also looking into this ticket. \r\n\r\n \r\n\r\n ","from":"developer"},{"body":"[~e.dimitrova] I ran some tests on 4.0 and hit another issue related to DESCRIBE. I need to dig a deeper. ","from":"developer"},{"body":"The underlying issue on 4.0 is now fixed (see CASSANDRA-16898). The patch for 3.0 and 3.11 LGTM. We probably just need to make a patch for 4.0 and trunk with only the unit tests. ","from":"developer"},{"body":"Thank you for raising CASSANDRA-16898 and the quick fix.\r\n\r\nTests added to [4.0|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-4.0?expand=1] and [trunk|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-trunk?expand=1].\r\n\r\n I am running in a loop the two new tests [here|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra?branch=12734-4.0].\r\n\r\nI also rebased and pushed full CI Jenkins run for the other two branches:\r\n|| Patch || CI ||\r\n|[3.0|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-3.0?expand=1]|[CI|https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1094/]|\r\n|[3.11|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:12734-3.11?expand=1]|[CI|https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1095/]|\r\n\r\n ","from":"developer"},{"body":"The patches look good to me. Thanks :-) ","from":"developer"},{"body":"To https://github.com/apache/cassandra.git\r\n\r\n   e4b37c3271..67eb22ec9d  cassandra-3.0 -> cassandra-3.0\r\n\r\n   957c6264ef..d6e1c41c48  cassandra-3.11 -> cassandra-3.11\r\n\r\n   6a4a93a808..49e83027e2  cassandra-4.0 -> cassandra-4.0\r\n\r\n   f9aa19e3b1..163a4d7137  trunk -> trunk\r\n\r\n \r\n\r\nCI also looked ok, committed, thanks!","from":"developer"}],"created":"2016-09-30T14:11:08.000+0000","description":"The materialized view schema file that gets created and stored with the sstables is created as a table instead of a materialized view. \n\nCan the materialized view be created and added to the corresponding table's schema file?\n","issue_id":"13008887","key":"CASSANDRA-12734","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2021-09-03T15:05:21.000+0000","role":"fixed_distractor","summary":"Materialized View schema file for snapshots created as tables"} {"case_id":"13009228","cluster":"DISTRACTOR-CASSANDRA-12743","comments":[{"body":"[~jbleduigou] Is this still happening for you?","created":"2017-09-06T09:10:03.323+0000"},{"body":"[~krummas] I'm on the same team as [~jbleduigou], we just had a 60h perf testing run without such issue, so it looks like it's not happening anymore.","created":"2017-09-06T09:20:03.363+0000"},{"body":"Ok, thanks, closing as cannot reproduce, please reopen if you see it again","created":"2017-09-06T09:26:14.536+0000"},{"body":"We're seeing the same issue in 3.0.14, here is the log:\r\n{noformat}\r\nERROR [CompactionExecutor:73398] 2018-01-30 02:23:32,177 CassandraDaemon.java:207 - Exception in thread Thread[CompactionExecutor:73398,1,main]\r\njava.lang.AssertionError: null\r\n at org.apache.cassandra.io.compress.CompressionMetadata$Chunk.(CompressionMetadata.java:475) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.compress.CompressionMetadata.chunkFor(CompressionMetadata.java:240) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.updateState(MmappedRegions.java:155) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:70) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:58) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.map(MmappedRegions.java:96) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.CompressedSegmentedFile.(CompressedSegmentedFile.java:47) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.CompressedSegmentedFile$Builder.complete(CompressedSegmentedFile.java:132) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.SegmentedFile$Builder.complete(SegmentedFile.java:177) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.SegmentedFile$Builder.buildData(SegmentedFile.java:188) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.openEarly(BigTableWriter.java:245) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.maybeReopenEarly(SSTableRewriter.java:172) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.append(SSTableRewriter.java:124) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.writers.DefaultCompactionWriter.realAppend(DefaultCompactionWriter.java:57) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.writers.CompactionAwareWriter.append(CompactionAwareWriter.java:109) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:195) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:89) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:61) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionCandidate.run(CompactionManager.java:264) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_121]\r\n at java.util.concurrent.FutureTask.run(FutureTask.java:266) ~[na:1.8.0_121]\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) ~[na:1.8.0_121]\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_121]\r\n at org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:79) [apache-cassandra-3.0.14.jar:3.0.14]\r\n at java.lang.Thread.run(Thread.java:745) ~[na:1.8.0_121]\r\n{noformat}\r\n\r\nSometimes the error is\r\n{noformat}\r\nERROR [CompactionExecutor:1324] 2018-01-31 17:00:21,922 CassandraDaemon.java:207 - Exception in thread Thread[CompactionExecutor:1324,1,main]\r\njava.lang.AssertionError: Illegal bounds [12800..12808); size: 12800\r\n at org.apache.cassandra.io.util.Memory.checkBounds(Memory.java:339) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.SafeMemory.checkBounds(SafeMemory.java:104) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.Memory.getLong(Memory.java:260) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.compress.CompressionMetadata.chunkFor(CompressionMetadata.java:238) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.updateState(MmappedRegions.java:155) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:70) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:58) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.map(MmappedRegions.java:96) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.CompressedSegmentedFile.(CompressedSegmentedFile.java:47) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.CompressedSegmentedFile$Builder.complete(CompressedSegmentedFile.java:132) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.SegmentedFile$Builder.complete(SegmentedFile.java:177) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.SegmentedFile$Builder.buildData(SegmentedFile.java:188) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.openEarly(BigTableWriter.java:245) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.maybeReopenEarly(SSTableRewriter.java:172) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.append(SSTableRewriter.java:124) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.writers.MaxSSTableSizeWriter.realAppend(MaxSSTableSizeWriter.java:88) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.writers.CompactionAwareWriter.append(CompactionAwareWriter.java:109) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:195) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:89) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:61) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionCandidate.run(CompactionManager.java:264) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_121]\r\n at java.util.concurrent.FutureTask.run(FutureTask.java:266) ~[na:1.8.0_121]\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) ~[na:1.8.0_121]\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_121]\r\n at org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:79) [apache-cassandra-3.0.14.jar:3.0.14]\r\n at java.lang.Thread.run(Thread.java:745) ~[na:1.8.0_121]\r\n{noformat}\r\n\r\n{{sstable_preemptive_open_interval_in_mb}} is set to 50 by default. Maybe we could try to disable it.\r\n\r\nIt happens for both LCS and STCS.","created":"2018-02-01T03:06:25.987+0000"},{"body":"Found one on 3.11.1 as well:\r\n{code:java}\r\nERROR [CompactionExecutor:1696] 2018-02-07 09:58:10,510 CassandraDaemon.java:228 - Exception in thread Thread[CompactionExecutor:1696,1,main]\r\njava.lang.AssertionError: null\r\n at org.apache.cassandra.io.compress.CompressionMetadata$Chunk.(CompressionMetadata.java:474) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.compress.CompressionMetadata.chunkFor(CompressionMetadata.java:239) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.util.MmappedRegions.updateState(MmappedRegions.java:163) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:73) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:61) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.util.MmappedRegions.map(MmappedRegions.java:104) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.util.FileHandle$Builder.complete(FileHandle.java:362) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.openEarly(BigTableWriter.java:290) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.maybeReopenEarly(SSTableRewriter.java:179) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.append(SSTableRewriter.java:134) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.writers.MaxSSTableSizeWriter.realAppend(MaxSSTableSizeWriter.java:98) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.writers.CompactionAwareWriter.append(CompactionAwareWriter.java:141) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:201) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:85) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:61) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionCandidate.run(CompactionManager.java:268) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source) ~[na:1.8.0_152]\r\n at java.util.concurrent.FutureTask.run(Unknown Source) ~[na:1.8.0_152]\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [na:1.8.0_152]\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [na:1.8.0_152]\r\n at org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:81) [apache-cassandra-3.11.1.jar:3.11.1]\r\n at java.lang.Thread.run(Unknown Source) ~[na:1.8.0_152]\r\n{code}","created":"2018-02-07T10:42:03.475+0000"},{"body":"The problem is because [{{dataSyncPosition}}|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/IndexSummaryBuilder.java#L64] is the compressed file size (set here: [{{BigTableWriter.java:445}}|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/format/big/BigTableWriter.java#L445]), VS. [{{lastReadableByData}}|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/IndexSummaryBuilder.java#L58] is having [uncompressed data size|https://github.com/apache/cassandra/blob/5dc55e715eba6667c388da9f8f1eb7a46489b35c/src/java/org/apache/cassandra/io/sstable/format/big/BigTableWriter.java#L185]: [{{IndexSummaryBuilder.java:222}}|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/IndexSummaryBuilder.java#L222].\r\nSo if the compression ratio is bigger or around {{1.0}} and [the index file is synced faster than the data file|https://github.com/apache/cassandra/blob/5dc55e715eba6667c388da9f8f1eb7a46489b35c/src/java/org/apache/cassandra/io/sstable/IndexSummaryBuilder.java#L174], [{{openEarly()}}|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/format/big/BigTableWriter.java#L287] may open data that haven't been synced.\r\n\r\nHere is the patch, please review:\r\n| Branch | uTest | dTest |\r\n| [12743-2.2|https://github.com/cooldoger/cassandra/tree/12743-2.2] | [!https://circleci.com/gh/cooldoger/cassandra/tree/12743-2.2.svg?style=svg!|https://circleci.com/gh/cooldoger/cassandra/tree/12743-2.2] | [#526|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/526]\r\n| [12743-3.0|https://github.com/cooldoger/cassandra/tree/12743-3.0] | [!https://circleci.com/gh/cooldoger/cassandra/tree/12743-3.0.svg?style=svg!|https://circleci.com/gh/cooldoger/cassandra/tree/12743-3.0] | [#527|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/527]\r\n| [12743-3.11|https://github.com/cooldoger/cassandra/tree/12743-3.11] | [!https://circleci.com/gh/cooldoger/cassandra/tree/12743-3.11.svg?style=svg!|https://circleci.com/gh/cooldoger/cassandra/tree/12743-3.11] | [#528|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/528]\r\n| [12743-trunk|https://github.com/cooldoger/cassandra/tree/12743-trunk] | [!https://circleci.com/gh/cooldoger/cassandra/tree/12743-trunk.svg?style=svg!|https://circleci.com/gh/cooldoger/cassandra/tree/12743-trunk] | [#529|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/529]","created":"2018-04-25T00:59:04.276+0000"},{"body":"nice catch [~jay.zhuang]!\r\n\r\nDo you think it would it be possible to create a regression test? Ie, to create an sstable where this reproduces?\r\n\r\nI think we need the truncate(...) fix in 2.2 and 3.0 as well (including the fix in SequentialWriter which only seems to be in 3.11+ (see [this|https://issues.apache.org/jira/browse/CASSANDRA-11579?focusedCommentId=15241980&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#comment-15241980] comment))?","created":"2018-04-25T12:31:06.139+0000"},{"body":"Thanks [~krummas] for the review. It makes sense to backport the {{truncate()}} part to {{2.2}} and {{3.0}}. The branch is updated, please review.\r\n\r\nIt's really hard to reproduce it locally, as the most of time the index file is synced slower than the data file, so it uses index file synced position (even the data synced position is wrong, it won't cause the problem): [{{IndexSummaryBuilder.java:174}}|https://github.com/apache/cassandra/blob/5dc55e715eba6667c388da9f8f1eb7a46489b35c/src/java/org/apache/cassandra/io/sstable/IndexSummaryBuilder.java#L174].\r\nBut in one of our cluster, it happens every dozens hours per node. I patched the fix and no longer see the issue.","created":"2018-04-25T22:02:33.402+0000"},{"body":"pushed a repro here: https://github.com/krummas/cassandra/commits/marcuse/repro_12743 - it runs a major compaction with a compressor that doubles the data size","created":"2018-04-26T07:05:56.359+0000"},{"body":"Cool, thanks [~krummas]. Updated each branch with the unittest.","created":"2018-04-26T21:43:20.768+0000"},{"body":"restarted the trunk dtests - +1 if they look ok","created":"2018-04-27T08:48:09.500+0000"},{"body":"There're a few dTest failures for each branch. Trying to confirm they're not introduced by this patch.","created":"2018-04-30T18:37:09.496+0000"},{"body":"Thank you Marcus again for the review. Committed as [3a71382|https://github.com/apache/cassandra/commit/3a713827f48399f389ea851a19b8ec8cd2cc5773].","created":"2018-05-01T22:32:17.153+0000"}],"conversations":[{"body":"While running compaction I run into an error sometimes :\n{noformat}\nnodetool compact\nerror: null\n-- StackTrace --\njava.lang.AssertionError\n at org.apache.cassandra.io.compress.CompressionMetadata$Chunk.(CompressionMetadata.java:463)\n at org.apache.cassandra.io.compress.CompressionMetadata.chunkFor(CompressionMetadata.java:228)\n at org.apache.cassandra.io.util.CompressedSegmentedFile.createMappedSegments(CompressedSegmentedFile.java:80)\n at org.apache.cassandra.io.util.CompressedPoolingSegmentedFile.(CompressedPoolingSegmentedFile.java:38)\n at org.apache.cassandra.io.util.CompressedPoolingSegmentedFile$Builder.complete(CompressedPoolingSegmentedFile.java:101)\n at org.apache.cassandra.io.util.SegmentedFile$Builder.complete(SegmentedFile.java:198)\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.openEarly(BigTableWriter.java:315)\n at org.apache.cassandra.io.sstable.SSTableRewriter.maybeReopenEarly(SSTableRewriter.java:171)\n at org.apache.cassandra.io.sstable.SSTableRewriter.append(SSTableRewriter.java:116)\n at org.apache.cassandra.db.compaction.writers.DefaultCompactionWriter.append(DefaultCompactionWriter.java:64)\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:184)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:74)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:59)\n at org.apache.cassandra.db.compaction.CompactionManager$8.runMayThrow(CompactionManager.java:599)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{noformat}\n\nWhy is that happening?\nIs there anyway to provide more details (e.g. which SSTable cannot be compacted)?\n\nWe are using Cassandra 2.2.7","from":"reporter","subject":"Assertion error while running compaction"},{"body":"[~jbleduigou] Is this still happening for you?","from":"developer"},{"body":"[~krummas] I'm on the same team as [~jbleduigou], we just had a 60h perf testing run without such issue, so it looks like it's not happening anymore.","from":"developer"},{"body":"Ok, thanks, closing as cannot reproduce, please reopen if you see it again","from":"developer"},{"body":"We're seeing the same issue in 3.0.14, here is the log:\r\n{noformat}\r\nERROR [CompactionExecutor:73398] 2018-01-30 02:23:32,177 CassandraDaemon.java:207 - Exception in thread Thread[CompactionExecutor:73398,1,main]\r\njava.lang.AssertionError: null\r\n at org.apache.cassandra.io.compress.CompressionMetadata$Chunk.(CompressionMetadata.java:475) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.compress.CompressionMetadata.chunkFor(CompressionMetadata.java:240) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.updateState(MmappedRegions.java:155) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:70) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:58) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.map(MmappedRegions.java:96) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.CompressedSegmentedFile.(CompressedSegmentedFile.java:47) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.CompressedSegmentedFile$Builder.complete(CompressedSegmentedFile.java:132) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.SegmentedFile$Builder.complete(SegmentedFile.java:177) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.SegmentedFile$Builder.buildData(SegmentedFile.java:188) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.openEarly(BigTableWriter.java:245) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.maybeReopenEarly(SSTableRewriter.java:172) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.append(SSTableRewriter.java:124) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.writers.DefaultCompactionWriter.realAppend(DefaultCompactionWriter.java:57) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.writers.CompactionAwareWriter.append(CompactionAwareWriter.java:109) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:195) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:89) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:61) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionCandidate.run(CompactionManager.java:264) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_121]\r\n at java.util.concurrent.FutureTask.run(FutureTask.java:266) ~[na:1.8.0_121]\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) ~[na:1.8.0_121]\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_121]\r\n at org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:79) [apache-cassandra-3.0.14.jar:3.0.14]\r\n at java.lang.Thread.run(Thread.java:745) ~[na:1.8.0_121]\r\n{noformat}\r\n\r\nSometimes the error is\r\n{noformat}\r\nERROR [CompactionExecutor:1324] 2018-01-31 17:00:21,922 CassandraDaemon.java:207 - Exception in thread Thread[CompactionExecutor:1324,1,main]\r\njava.lang.AssertionError: Illegal bounds [12800..12808); size: 12800\r\n at org.apache.cassandra.io.util.Memory.checkBounds(Memory.java:339) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.SafeMemory.checkBounds(SafeMemory.java:104) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.Memory.getLong(Memory.java:260) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.compress.CompressionMetadata.chunkFor(CompressionMetadata.java:238) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.updateState(MmappedRegions.java:155) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:70) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:58) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.MmappedRegions.map(MmappedRegions.java:96) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.CompressedSegmentedFile.(CompressedSegmentedFile.java:47) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.CompressedSegmentedFile$Builder.complete(CompressedSegmentedFile.java:132) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.SegmentedFile$Builder.complete(SegmentedFile.java:177) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.util.SegmentedFile$Builder.buildData(SegmentedFile.java:188) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.openEarly(BigTableWriter.java:245) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.maybeReopenEarly(SSTableRewriter.java:172) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.append(SSTableRewriter.java:124) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.writers.MaxSSTableSizeWriter.realAppend(MaxSSTableSizeWriter.java:88) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.writers.CompactionAwareWriter.append(CompactionAwareWriter.java:109) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:195) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:89) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:61) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionCandidate.run(CompactionManager.java:264) ~[apache-cassandra-3.0.14.jar:3.0.14]\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_121]\r\n at java.util.concurrent.FutureTask.run(FutureTask.java:266) ~[na:1.8.0_121]\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) ~[na:1.8.0_121]\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_121]\r\n at org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:79) [apache-cassandra-3.0.14.jar:3.0.14]\r\n at java.lang.Thread.run(Thread.java:745) ~[na:1.8.0_121]\r\n{noformat}\r\n\r\n{{sstable_preemptive_open_interval_in_mb}} is set to 50 by default. Maybe we could try to disable it.\r\n\r\nIt happens for both LCS and STCS.","from":"developer"},{"body":"Found one on 3.11.1 as well:\r\n{code:java}\r\nERROR [CompactionExecutor:1696] 2018-02-07 09:58:10,510 CassandraDaemon.java:228 - Exception in thread Thread[CompactionExecutor:1696,1,main]\r\njava.lang.AssertionError: null\r\n at org.apache.cassandra.io.compress.CompressionMetadata$Chunk.(CompressionMetadata.java:474) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.compress.CompressionMetadata.chunkFor(CompressionMetadata.java:239) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.util.MmappedRegions.updateState(MmappedRegions.java:163) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:73) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.util.MmappedRegions.(MmappedRegions.java:61) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.util.MmappedRegions.map(MmappedRegions.java:104) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.util.FileHandle$Builder.complete(FileHandle.java:362) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.openEarly(BigTableWriter.java:290) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.maybeReopenEarly(SSTableRewriter.java:179) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.io.sstable.SSTableRewriter.append(SSTableRewriter.java:134) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.writers.MaxSSTableSizeWriter.realAppend(MaxSSTableSizeWriter.java:98) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.writers.CompactionAwareWriter.append(CompactionAwareWriter.java:141) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:201) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:85) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:61) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionCandidate.run(CompactionManager.java:268) ~[apache-cassandra-3.11.1.jar:3.11.1]\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source) ~[na:1.8.0_152]\r\n at java.util.concurrent.FutureTask.run(Unknown Source) ~[na:1.8.0_152]\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [na:1.8.0_152]\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [na:1.8.0_152]\r\n at org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:81) [apache-cassandra-3.11.1.jar:3.11.1]\r\n at java.lang.Thread.run(Unknown Source) ~[na:1.8.0_152]\r\n{code}","from":"developer"},{"body":"The problem is because [{{dataSyncPosition}}|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/IndexSummaryBuilder.java#L64] is the compressed file size (set here: [{{BigTableWriter.java:445}}|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/format/big/BigTableWriter.java#L445]), VS. [{{lastReadableByData}}|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/IndexSummaryBuilder.java#L58] is having [uncompressed data size|https://github.com/apache/cassandra/blob/5dc55e715eba6667c388da9f8f1eb7a46489b35c/src/java/org/apache/cassandra/io/sstable/format/big/BigTableWriter.java#L185]: [{{IndexSummaryBuilder.java:222}}|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/IndexSummaryBuilder.java#L222].\r\nSo if the compression ratio is bigger or around {{1.0}} and [the index file is synced faster than the data file|https://github.com/apache/cassandra/blob/5dc55e715eba6667c388da9f8f1eb7a46489b35c/src/java/org/apache/cassandra/io/sstable/IndexSummaryBuilder.java#L174], [{{openEarly()}}|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/io/sstable/format/big/BigTableWriter.java#L287] may open data that haven't been synced.\r\n\r\nHere is the patch, please review:\r\n| Branch | uTest | dTest |\r\n| [12743-2.2|https://github.com/cooldoger/cassandra/tree/12743-2.2] | [!https://circleci.com/gh/cooldoger/cassandra/tree/12743-2.2.svg?style=svg!|https://circleci.com/gh/cooldoger/cassandra/tree/12743-2.2] | [#526|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/526]\r\n| [12743-3.0|https://github.com/cooldoger/cassandra/tree/12743-3.0] | [!https://circleci.com/gh/cooldoger/cassandra/tree/12743-3.0.svg?style=svg!|https://circleci.com/gh/cooldoger/cassandra/tree/12743-3.0] | [#527|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/527]\r\n| [12743-3.11|https://github.com/cooldoger/cassandra/tree/12743-3.11] | [!https://circleci.com/gh/cooldoger/cassandra/tree/12743-3.11.svg?style=svg!|https://circleci.com/gh/cooldoger/cassandra/tree/12743-3.11] | [#528|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/528]\r\n| [12743-trunk|https://github.com/cooldoger/cassandra/tree/12743-trunk] | [!https://circleci.com/gh/cooldoger/cassandra/tree/12743-trunk.svg?style=svg!|https://circleci.com/gh/cooldoger/cassandra/tree/12743-trunk] | [#529|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/529]","from":"developer"},{"body":"nice catch [~jay.zhuang]!\r\n\r\nDo you think it would it be possible to create a regression test? Ie, to create an sstable where this reproduces?\r\n\r\nI think we need the truncate(...) fix in 2.2 and 3.0 as well (including the fix in SequentialWriter which only seems to be in 3.11+ (see [this|https://issues.apache.org/jira/browse/CASSANDRA-11579?focusedCommentId=15241980&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#comment-15241980] comment))?","from":"developer"},{"body":"Thanks [~krummas] for the review. It makes sense to backport the {{truncate()}} part to {{2.2}} and {{3.0}}. The branch is updated, please review.\r\n\r\nIt's really hard to reproduce it locally, as the most of time the index file is synced slower than the data file, so it uses index file synced position (even the data synced position is wrong, it won't cause the problem): [{{IndexSummaryBuilder.java:174}}|https://github.com/apache/cassandra/blob/5dc55e715eba6667c388da9f8f1eb7a46489b35c/src/java/org/apache/cassandra/io/sstable/IndexSummaryBuilder.java#L174].\r\nBut in one of our cluster, it happens every dozens hours per node. I patched the fix and no longer see the issue.","from":"developer"},{"body":"pushed a repro here: https://github.com/krummas/cassandra/commits/marcuse/repro_12743 - it runs a major compaction with a compressor that doubles the data size","from":"developer"},{"body":"Cool, thanks [~krummas]. Updated each branch with the unittest.","from":"developer"},{"body":"restarted the trunk dtests - +1 if they look ok","from":"developer"},{"body":"There're a few dTest failures for each branch. Trying to confirm they're not introduced by this patch.","from":"developer"},{"body":"Thank you Marcus again for the review. Committed as [3a71382|https://github.com/apache/cassandra/commit/3a713827f48399f389ea851a19b8ec8cd2cc5773].","from":"developer"}],"created":"2016-10-03T12:57:39.000+0000","description":"While running compaction I run into an error sometimes :\n{noformat}\nnodetool compact\nerror: null\n-- StackTrace --\njava.lang.AssertionError\n at org.apache.cassandra.io.compress.CompressionMetadata$Chunk.(CompressionMetadata.java:463)\n at org.apache.cassandra.io.compress.CompressionMetadata.chunkFor(CompressionMetadata.java:228)\n at org.apache.cassandra.io.util.CompressedSegmentedFile.createMappedSegments(CompressedSegmentedFile.java:80)\n at org.apache.cassandra.io.util.CompressedPoolingSegmentedFile.(CompressedPoolingSegmentedFile.java:38)\n at org.apache.cassandra.io.util.CompressedPoolingSegmentedFile$Builder.complete(CompressedPoolingSegmentedFile.java:101)\n at org.apache.cassandra.io.util.SegmentedFile$Builder.complete(SegmentedFile.java:198)\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.openEarly(BigTableWriter.java:315)\n at org.apache.cassandra.io.sstable.SSTableRewriter.maybeReopenEarly(SSTableRewriter.java:171)\n at org.apache.cassandra.io.sstable.SSTableRewriter.append(SSTableRewriter.java:116)\n at org.apache.cassandra.db.compaction.writers.DefaultCompactionWriter.append(DefaultCompactionWriter.java:64)\n at org.apache.cassandra.db.compaction.CompactionTask.runMayThrow(CompactionTask.java:184)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:74)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:59)\n at org.apache.cassandra.db.compaction.CompactionManager$8.runMayThrow(CompactionManager.java:599)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{noformat}\n\nWhy is that happening?\nIs there anyway to provide more details (e.g. which SSTable cannot be compacted)?\n\nWe are using Cassandra 2.2.7","issue_id":"13009228","key":"CASSANDRA-12743","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2018-05-01T22:32:50.000+0000","role":"fixed_distractor","summary":"Assertion error while running compaction"} {"case_id":"13010902","cluster":"DISTRACTOR-CASSANDRA-12765","comments":[{"body":"Traced the issue to\n\n{code:title=CollationController.java|borderStyle=solid}\n private ColumnFamily collectAllData(boolean copyOnHeap)\n {\n // omitted for brevity\n if (!filter.shouldInclude(sstable))\n {\n nonIntersectingSSTables++;\n // sstable contains no tombstone if maxLocalDeletionTime == Integer.MAX_VALUE, so we can safely skip those entirely\n if (sstable.getSSTableMetadata().maxLocalDeletionTime != Integer.MAX_VALUE)\n {\n if (skippedSSTables == null)\n skippedSSTables = new ArrayList<>();\n skippedSSTables.add(sstable);\n }\n continue;\n }\n{code}\n\nThe sstable is excluded by the filter because:\n\n{code:title=SliceQueryFilter.java|borderStyle=solid}\n public boolean shouldInclude(SSTableReader sstable)\n {\n List minColumnNames = sstable.getSSTableMetadata().minColumnNames;\n List maxColumnNames = sstable.getSSTableMetadata().maxColumnNames;\n CellNameType comparator = sstable.metadata.comparator;\n\n if (minColumnNames.isEmpty() || maxColumnNames.isEmpty())\n return true;\n\n for (ColumnSlice slice : slices)\n if (slice.intersects(minColumnNames, maxColumnNames, comparator, reversed))\n return true;\n\n return false;\n }\n{code}\n\nThe other partition (eg. 8772618c9009cf8f5a5e0c19) means minColumnNames and maxColumnNames are not empty, and because the cluster key is different (eg. test2) it also doesn't intersect. So that means if moves inside the if (!filter.shouldInclude(sstable)).\n\nThe comment about if maxLocalDeletionTime == Integer.MAX_VALUE means the sstable contains no tombstones is wrong. As shown in the steps to reproduce the sstable that contains the row level deletion and another partition the metadata has maxLocalDeletionTime == Integer.MAX_VALUE because of the live cell.\n\n{code:title=ColumnFamily.java|borderStyle=solid}\n public ColumnStats getColumnStats()\n {\n // omitted for brevity\n for (Cell cell : this)\n {\n minTimestampTracker.update(cell.timestamp());\n maxTimestampTracker.update(cell.timestamp());\n maxDeletionTimeTracker.update(cell.getLocalDeletionTime());\n{code}\n\nWith the patch the sstable is added to skippedSSTables and therefore gets included due to tombstones.\n\nAs far as I can tell this issue dates back to https://issues.apache.org/jira/browse/CASSANDRA-5514 but I haven't attempted to reproduce in any version earlier then 2.1.15 and its been an issue on a cluster managing which started on 2.1.13, so I have currently tagged this bug as since 2.0 beta 1 since that corresponds to #5514","created":"2016-10-10T05:20:56.709+0000"},{"body":"Also reproducible on 3.x including latest trunk.","created":"2016-10-10T09:19:45.506+0000"},{"body":"[~krummas] to review","created":"2016-10-17T21:57:49.319+0000"},{"body":"Looks like Stefania and Branimir already encountered this in CASSANDRA-8180. Seems they implemented a fix as well but I guess it didn't work? https://issues.apache.org/jira/browse/CASSANDRA-8180?focusedCommentId=15116781&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15116781","created":"2016-10-19T10:40:09.402+0000"},{"body":"[~KurtG] yeah, I'm working on that, there is another bug in 3.x that breaks this","created":"2016-10-19T10:42:04.622+0000"},{"body":"I created a test and pushed the patch and the fix for 3.0+ - running tests:\n\n||branch||testall||dtest||\n|[cameron/12765-2.1|https://github.com/krummas/cassandra/tree/cameron/12765-2.1]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-2.1-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-2.1-dtest]|\n|[cameron/12765-2.2|https://github.com/krummas/cassandra/tree/cameron/12765-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-2.2-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-2.2-dtest]|\n|[cameron/12765-3.0|https://github.com/krummas/cassandra/tree/cameron/12765-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-3.0-dtest]|\n|[cameron/12765-3.X|https://github.com/krummas/cassandra/tree/cameron/12765-3.X]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-3.X-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-3.X-dtest]|\n|[cameron/12765-trunk|https://github.com/krummas/cassandra/tree/cameron/12765-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-trunk-dtest]|\n\nthe 3.0+ fix is [here|https://github.com/apache/cassandra/compare/trunk...krummas:cameron/12765-trunk?expand=1] - we were only including the partition level deletion if it was live\n","created":"2016-10-19T11:52:30.622+0000"},{"body":"if {{sstable.getSSTableMetadata().minLocalDeletionTime != Integer.MAX_VALUE}} is the fix, how come we don't apply that same fix to 2.1/2?","created":"2016-10-20T01:22:56.246+0000"},{"body":"[~KurtG] we didn't start collecting that information until 3.0\n\nhttps://github.com/apache/cassandra/blob/cassandra-2.2/src/java/org/apache/cassandra/io/sstable/metadata/StatsMetadata.java#L50\nvs\nhttps://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/io/sstable/metadata/StatsMetadata.java#L51\n","created":"2016-10-20T05:28:19.970+0000"},{"body":"well, that makes sense :p. thanks for the explanation.","created":"2016-10-20T08:02:18.219+0000"},{"body":"Committed, thanks for the patch","created":"2016-10-21T07:18:09.399+0000"}],"conversations":[{"body":"{noformat}\nCREATE TABLE test.payload(\n bucket_id TEXT,\n name TEXT,\n data TEXT,\n PRIMARY KEY (bucket_id, name)\n);\ninsert into test.payload (bucket_id, name, data) values ('8772618c9009cf8f5a5e0c18', 'test', 'hello');\n{noformat}\n\nFlush nodes (nodetool flush)\n\n{noformat}\ninsert into test.payload (bucket_id, name, data) values ('8772618c9009cf8f5a5e0c19', 'test2', 'hello');\ndelete from test.payload where bucket_id = '8772618c9009cf8f5a5e0c18';\n{noformat}\n\nFlush nodes (nodetool flush)\n\n{noformat}\nselect * from test.payload where bucket_id = '8772618c9009cf8f5a5e0c18' and name = 'test';\n{noformat}\n\nExpected 0 rows but get 1 row back.","from":"reporter","subject":"SSTable ignored incorrectly with partition level tombstone"},{"body":"Traced the issue to\n\n{code:title=CollationController.java|borderStyle=solid}\n private ColumnFamily collectAllData(boolean copyOnHeap)\n {\n // omitted for brevity\n if (!filter.shouldInclude(sstable))\n {\n nonIntersectingSSTables++;\n // sstable contains no tombstone if maxLocalDeletionTime == Integer.MAX_VALUE, so we can safely skip those entirely\n if (sstable.getSSTableMetadata().maxLocalDeletionTime != Integer.MAX_VALUE)\n {\n if (skippedSSTables == null)\n skippedSSTables = new ArrayList<>();\n skippedSSTables.add(sstable);\n }\n continue;\n }\n{code}\n\nThe sstable is excluded by the filter because:\n\n{code:title=SliceQueryFilter.java|borderStyle=solid}\n public boolean shouldInclude(SSTableReader sstable)\n {\n List minColumnNames = sstable.getSSTableMetadata().minColumnNames;\n List maxColumnNames = sstable.getSSTableMetadata().maxColumnNames;\n CellNameType comparator = sstable.metadata.comparator;\n\n if (minColumnNames.isEmpty() || maxColumnNames.isEmpty())\n return true;\n\n for (ColumnSlice slice : slices)\n if (slice.intersects(minColumnNames, maxColumnNames, comparator, reversed))\n return true;\n\n return false;\n }\n{code}\n\nThe other partition (eg. 8772618c9009cf8f5a5e0c19) means minColumnNames and maxColumnNames are not empty, and because the cluster key is different (eg. test2) it also doesn't intersect. So that means if moves inside the if (!filter.shouldInclude(sstable)).\n\nThe comment about if maxLocalDeletionTime == Integer.MAX_VALUE means the sstable contains no tombstones is wrong. As shown in the steps to reproduce the sstable that contains the row level deletion and another partition the metadata has maxLocalDeletionTime == Integer.MAX_VALUE because of the live cell.\n\n{code:title=ColumnFamily.java|borderStyle=solid}\n public ColumnStats getColumnStats()\n {\n // omitted for brevity\n for (Cell cell : this)\n {\n minTimestampTracker.update(cell.timestamp());\n maxTimestampTracker.update(cell.timestamp());\n maxDeletionTimeTracker.update(cell.getLocalDeletionTime());\n{code}\n\nWith the patch the sstable is added to skippedSSTables and therefore gets included due to tombstones.\n\nAs far as I can tell this issue dates back to https://issues.apache.org/jira/browse/CASSANDRA-5514 but I haven't attempted to reproduce in any version earlier then 2.1.15 and its been an issue on a cluster managing which started on 2.1.13, so I have currently tagged this bug as since 2.0 beta 1 since that corresponds to #5514","from":"developer"},{"body":"Also reproducible on 3.x including latest trunk.","from":"developer"},{"body":"[~krummas] to review","from":"developer"},{"body":"Looks like Stefania and Branimir already encountered this in CASSANDRA-8180. Seems they implemented a fix as well but I guess it didn't work? https://issues.apache.org/jira/browse/CASSANDRA-8180?focusedCommentId=15116781&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15116781","from":"developer"},{"body":"[~KurtG] yeah, I'm working on that, there is another bug in 3.x that breaks this","from":"developer"},{"body":"I created a test and pushed the patch and the fix for 3.0+ - running tests:\n\n||branch||testall||dtest||\n|[cameron/12765-2.1|https://github.com/krummas/cassandra/tree/cameron/12765-2.1]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-2.1-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-2.1-dtest]|\n|[cameron/12765-2.2|https://github.com/krummas/cassandra/tree/cameron/12765-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-2.2-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-2.2-dtest]|\n|[cameron/12765-3.0|https://github.com/krummas/cassandra/tree/cameron/12765-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-3.0-dtest]|\n|[cameron/12765-3.X|https://github.com/krummas/cassandra/tree/cameron/12765-3.X]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-3.X-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-3.X-dtest]|\n|[cameron/12765-trunk|https://github.com/krummas/cassandra/tree/cameron/12765-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/krummas/job/krummas-cameron-12765-trunk-dtest]|\n\nthe 3.0+ fix is [here|https://github.com/apache/cassandra/compare/trunk...krummas:cameron/12765-trunk?expand=1] - we were only including the partition level deletion if it was live\n","from":"developer"},{"body":"if {{sstable.getSSTableMetadata().minLocalDeletionTime != Integer.MAX_VALUE}} is the fix, how come we don't apply that same fix to 2.1/2?","from":"developer"},{"body":"[~KurtG] we didn't start collecting that information until 3.0\n\nhttps://github.com/apache/cassandra/blob/cassandra-2.2/src/java/org/apache/cassandra/io/sstable/metadata/StatsMetadata.java#L50\nvs\nhttps://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/io/sstable/metadata/StatsMetadata.java#L51\n","from":"developer"},{"body":"well, that makes sense :p. thanks for the explanation.","from":"developer"},{"body":"Committed, thanks for the patch","from":"developer"}],"created":"2016-10-10T05:04:39.000+0000","description":"{noformat}\nCREATE TABLE test.payload(\n bucket_id TEXT,\n name TEXT,\n data TEXT,\n PRIMARY KEY (bucket_id, name)\n);\ninsert into test.payload (bucket_id, name, data) values ('8772618c9009cf8f5a5e0c18', 'test', 'hello');\n{noformat}\n\nFlush nodes (nodetool flush)\n\n{noformat}\ninsert into test.payload (bucket_id, name, data) values ('8772618c9009cf8f5a5e0c19', 'test2', 'hello');\ndelete from test.payload where bucket_id = '8772618c9009cf8f5a5e0c18';\n{noformat}\n\nFlush nodes (nodetool flush)\n\n{noformat}\nselect * from test.payload where bucket_id = '8772618c9009cf8f5a5e0c18' and name = 'test';\n{noformat}\n\nExpected 0 rows but get 1 row back.","issue_id":"13010902","key":"CASSANDRA-12765","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-10-21T07:18:07.000+0000","role":"fixed_distractor","summary":"SSTable ignored incorrectly with partition level tombstone"} {"case_id":"13012791","cluster":"DISTRACTOR-CASSANDRA-12796","comments":[{"body":"I am able to reproduce this issue in Apache Cassandra 3.0.8 with a wide partition and a secondary index defined over it.\n\nThe code has changed significantly between the version reported here and 3.0.8 however the characteristics of the failure are fairly similar, i.e. when a secondary index is rebuild there is a build up of large number of pending memtable flush runnables and the node gets overwhelmed and crashes due to an OOM.\n\nAdjusting the granule on which the write barrier applies (taking a pass with the suggested patch's logic on the 3.0.8 code) does seem to alleviate the problem and I do not see the memtable flush runnables queue up, however I am not sure if there are other unintended consequences of tweaking this write barrier granule which need to be considered.\n\nI would like to know if this patch can be brought into Cassandra 3.0.x or are there other solutions to deal with large partitions with secondary indexes?","created":"2016-11-04T20:10:23.960+0000"},{"body":"b.q. I would like to know if this patch can be brought into Cassandra 3.0.x or are there other solutions to deal with large partitions with secondary indexes?\n\nThere aren't I'm afraid, so if you could post your patches for 2.2 & 3.0 I'll make sure they get reviewed (I shouldn't think a separate patch for trunk will be necessary).\nThanks.","created":"2016-11-07T16:25:57.208+0000"},{"body":"I'll post the formal patch for 2.2 shortly.","created":"2016-11-08T11:02:11.103+0000"},{"body":"The patch for *2.2* can be found at [https://github.com/mmajercik/cassandra/tree/12796-2.2]\n\nI made some attempt to adjust this patch for branch *3.0* but was overwhelmed by sheer extent of refactoring that virtually left no stone untouched. I suppose it takes a while until I'll be able to issue patch for *3.0*.","created":"2016-11-10T09:06:01.465+0000"},{"body":"[~mmajercik] no worries, I can take care of porting your patch to 3.0. Leave it with me and I'll try to get to it in the next day or two.","created":"2016-11-10T11:11:55.230+0000"},{"body":"Forward ported the original patch to 3.0/3.X/trunk and submitted CI jobs:\n\n||branch||testall||dtest||\n|[12796-2.2|https://github.com/beobal/cassandra/tree/12796-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-2.2-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-2.2-dtest]|\n|[12796-3.0|https://github.com/beobal/cassandra/tree/12796-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.0-dtest]|\n|[12796-3.X|https://github.com/beobal/cassandra/tree/12796-3.X]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.X-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.X-dtest]|\n|[12796-trunk|https://github.com/beobal/cassandra/tree/12796-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-trunk-dtest]|\n","created":"2016-11-11T15:49:40.665+0000"},{"body":"It looks like the proposed solution is only partially adequate for 3.0.x; while it enables memtable _flushes_ to proceed while a partition is being indexed, the read ordering {{OpGroup}} introduced in CASSANDRA-11905 continues to block memtable memory from being _reclaimed_. In our environment, this ultimately blocks new memtables from being allocated while indexing is underway, which in turn blocks new mutations from finishing while they for new memtables to become available. That ultimately also leads to heap exhaustion as blocked mutations accumulate.\n\nSo it looks like the granularity of the read {{OpOrder}} lock also needs to be reduced. Following the logic of CASSANDRA-11905, before 3.0.x that was implicitly accomplished by using a pager to read the partition, which handled the read {{OpOrder}} correctly behind the scenes. Is doing something like that an option now?","created":"2016-11-22T20:29:47.723+0000"},{"body":"I observed the same issue when did brief testing on 3.0 branch","created":"2016-11-23T08:07:43.395+0000"},{"body":"bq.So it looks like the granularity of the read OpOrder lock also needs to be reduced\n\nGood point, and like you said it's possible to use a Pager here, as in pre 3.0 versions. \n\nI've force-pushed new versions for 3.0/3.11/3.X/trunk (the 2.2 version is unchanged of course). There have been some issues with dtests in CI today, so I'll kick those off when the infra is stable again.\n\n||branch||testall||dtest||\n|[12796-2.2|https://github.com/beobal/cassandra/tree/12796-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-2.2-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-2.2-dtest]|\n|[12796-3.0|https://github.com/beobal/cassandra/tree/12796-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.0-dtest]|\n|[12796-3.11|https://github.com/beobal/cassandra/tree/12796-3.11]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.11-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.11-dtest]|\n|[12796-3.X|https://github.com/beobal/cassandra/tree/12796-3.X]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.X-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.X-dtest]|\n|[12796-trunk|https://github.com/beobal/cassandra/tree/12796-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-trunk-dtest]|\n","created":"2016-11-24T18:50:13.954+0000"},{"body":"The CI looks reasonable: 3 dtest failures on the 3.0 branch, which all have corresponding failures upstream, plus a couple of failures on the original 2.2 branch which have since been addressed by other tickets. \n\nThe internal paging will aim to read rows in chunks of ~4mb and I've used the default CQL page size of 10k rows for a floor as it seems like as good a place to start as any. \n\n[~mmajercik], [~anmols] how does this latest 3.0 version look with your testing? ","created":"2016-11-29T12:50:49.567+0000"},{"body":"[~beobal], the patch for branch _3.0_ works fine, however the page size for single partition pager appears to be calculated incorrectly. My table's average partition size is around *7GB* and yet the page size got calculated as *1*.\n\n{code:java}\n private int calculateIndexingPageSize()\n {\n double averageRowSize = baseCfs.getMeanPartitionSize();\n if (averageRowSize <= 0)\n return DEFAULT_PAGE_SIZE;\n\n return (int) Math.max(1, Math.min(DEFAULT_PAGE_SIZE, 4 * 1024 * 1024 / averageRowSize));\n }\n{code}\n\nThis rendered index rebuild extremely slow as registering read/write order group implies significant performance overhead and for this reason the page size should have reasonable value.\n\nI think there is no harm if we set page size to {{DEFAULT_PAGE_SIZE}} as the pager doesn't span across different partitions in case the partition is small ([https://github.com/mmajercik/cassandra/commit/3fc016e73d3032f4d04584a45945141151a49213])\n\n[12796-3.0|https://github.com/mmajercik/cassandra/tree/12796-3.0]\n","created":"2016-12-06T12:37:40.064+0000"},{"body":"I think you want (row per part / (mean part size / 4 MB)) if you are guessing at \"rows per 4 MB\"","created":"2016-12-07T12:01:55.656+0000"},{"body":"Yep, the method was totally wrong. I'm not 100% comfortable with hardcoding the page size as suggested so I'm just running some quick experiments to figure out a reasonable X for \"rows per X MB\".","created":"2016-12-07T12:07:52.961+0000"},{"body":"Pushed another commit with an implementation of {{calculateIndexingPageSize}} which actually works. It isn't perfect of course, but when the average row size is smallish, it will tend to use the default page size of 10000, which is no problem as the each ordering will still only be held open for a matter of milliseconds. When rows are larger though, that default page size could cause problems as the orderings will be held open for much longer, so this will help to keep that window more consistently bounded. \n\nI've also made the page size an argument to {{indexPartition}} so we don't perform the redundant calculation for every partition. \n","created":"2016-12-08T08:58:39.437+0000"},{"body":"Excellent. This is the closest we can get to keep target page size in bytes. Do you want me to test it with our test data? If not, could it be merged into the primary repository?","created":"2016-12-08T09:56:58.175+0000"},{"body":"If you could double check with your test data, that'd be awesome. I'm just waiting for the latest CI runs to finish, so if all looks good with that & your test data I'll commit asap.","created":"2016-12-08T11:26:07.323+0000"},{"body":"I think the new algorithm looks good. It has a much better chance of doing the right thing. Couple nits on just that commit, I would put the table name in the TRACE log message, and you might add the new -D to the jvm.options file in 3.11+.","created":"2016-12-08T13:48:45.156+0000"},{"body":"[~beobal], sorry for delay in response. Just tested your fix on our sample data and it works fine. Thank you for your help.","created":"2016-12-12T09:25:57.198+0000"},{"body":"Thanks, committed to 2.2 in {{fb2940050e27a5642a23f3e9b5aaa7ae65e018b1}} and to 3.0 (with the additional log info) in {{36ce4e02b429b1297d71c5c8a963623c62d9e159}}. Also added the new -D switch to {{jvm.options}} from 3.10 (i.e. the cassandra-3.11 branch) onwards.","created":"2016-12-13T10:41:54.317+0000"}],"conversations":[{"body":"We have a table with rather wide partition and a secondary index defined over it. As soon as we try to rebuild the index we observed exhaustion of Java heap and eventual OOM error. After a lengthy investigation we have managed to find a culprit which appears to be a wrong granule of barrier issuances in method {{org.apache.cassandra.db.Keyspace.indexRow}}:\n\n{code}\n try (OpOrder.Group opGroup = cfs.keyspace.writeOrder.start()){html}\n {\n Set indexes = cfs.indexManager.getIndexesByNames(idxNames);\n\n Iterator pager = QueryPagers.pageRowLocally(cfs, key.getKey(), DEFAULT_PAGE_SIZE);\n while (pager.hasNext())\n {\n ColumnFamily cf = pager.next();\n ColumnFamily cf2 = cf.cloneMeShallow();\n for (Cell cell : cf)\n {\n if (cfs.indexManager.indexes(cell.name(), indexes))\n cf2.addColumn(cell);\n }\n cfs.indexManager.indexRow(key.getKey(), cf2, opGroup);\n }\n }\n{code}\n\nPlease note the operation group granule is a partition of the source table which poses a problem for wide partition tables as flush runnable ({{org.apache.cassandra.db.ColumnFamilyStore.Flush.run()}}) won't proceed with flushing secondary index memtable before completing operations prior recent issue of the barrier. In our situation the flush runnable waits until whole wide partition gets indexed into the secondary index memtable before flushing it. This causes an exhaustion of the heap and eventual OOM error.\n\nAfter we changed granule of barrier issue in method {{org.apache.cassandra.db.Keyspace.indexRow}} to query page as opposed to table partition secondary index (see [https://github.com/mmajercik/cassandra/commit/7e10e5aa97f1de483c2a5faf867315ecbf65f3d6?diff=unified]), rebuild started to work without heap exhaustion. ","from":"reporter","subject":"Heap exhaustion when rebuilding secondary index over a table with wide partitions"},{"body":"I am able to reproduce this issue in Apache Cassandra 3.0.8 with a wide partition and a secondary index defined over it.\n\nThe code has changed significantly between the version reported here and 3.0.8 however the characteristics of the failure are fairly similar, i.e. when a secondary index is rebuild there is a build up of large number of pending memtable flush runnables and the node gets overwhelmed and crashes due to an OOM.\n\nAdjusting the granule on which the write barrier applies (taking a pass with the suggested patch's logic on the 3.0.8 code) does seem to alleviate the problem and I do not see the memtable flush runnables queue up, however I am not sure if there are other unintended consequences of tweaking this write barrier granule which need to be considered.\n\nI would like to know if this patch can be brought into Cassandra 3.0.x or are there other solutions to deal with large partitions with secondary indexes?","from":"developer"},{"body":"b.q. I would like to know if this patch can be brought into Cassandra 3.0.x or are there other solutions to deal with large partitions with secondary indexes?\n\nThere aren't I'm afraid, so if you could post your patches for 2.2 & 3.0 I'll make sure they get reviewed (I shouldn't think a separate patch for trunk will be necessary).\nThanks.","from":"developer"},{"body":"I'll post the formal patch for 2.2 shortly.","from":"developer"},{"body":"The patch for *2.2* can be found at [https://github.com/mmajercik/cassandra/tree/12796-2.2]\n\nI made some attempt to adjust this patch for branch *3.0* but was overwhelmed by sheer extent of refactoring that virtually left no stone untouched. I suppose it takes a while until I'll be able to issue patch for *3.0*.","from":"developer"},{"body":"[~mmajercik] no worries, I can take care of porting your patch to 3.0. Leave it with me and I'll try to get to it in the next day or two.","from":"developer"},{"body":"Forward ported the original patch to 3.0/3.X/trunk and submitted CI jobs:\n\n||branch||testall||dtest||\n|[12796-2.2|https://github.com/beobal/cassandra/tree/12796-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-2.2-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-2.2-dtest]|\n|[12796-3.0|https://github.com/beobal/cassandra/tree/12796-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.0-dtest]|\n|[12796-3.X|https://github.com/beobal/cassandra/tree/12796-3.X]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.X-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.X-dtest]|\n|[12796-trunk|https://github.com/beobal/cassandra/tree/12796-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-trunk-dtest]|\n","from":"developer"},{"body":"It looks like the proposed solution is only partially adequate for 3.0.x; while it enables memtable _flushes_ to proceed while a partition is being indexed, the read ordering {{OpGroup}} introduced in CASSANDRA-11905 continues to block memtable memory from being _reclaimed_. In our environment, this ultimately blocks new memtables from being allocated while indexing is underway, which in turn blocks new mutations from finishing while they for new memtables to become available. That ultimately also leads to heap exhaustion as blocked mutations accumulate.\n\nSo it looks like the granularity of the read {{OpOrder}} lock also needs to be reduced. Following the logic of CASSANDRA-11905, before 3.0.x that was implicitly accomplished by using a pager to read the partition, which handled the read {{OpOrder}} correctly behind the scenes. Is doing something like that an option now?","from":"developer"},{"body":"I observed the same issue when did brief testing on 3.0 branch","from":"developer"},{"body":"bq.So it looks like the granularity of the read OpOrder lock also needs to be reduced\n\nGood point, and like you said it's possible to use a Pager here, as in pre 3.0 versions. \n\nI've force-pushed new versions for 3.0/3.11/3.X/trunk (the 2.2 version is unchanged of course). There have been some issues with dtests in CI today, so I'll kick those off when the infra is stable again.\n\n||branch||testall||dtest||\n|[12796-2.2|https://github.com/beobal/cassandra/tree/12796-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-2.2-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-2.2-dtest]|\n|[12796-3.0|https://github.com/beobal/cassandra/tree/12796-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.0-dtest]|\n|[12796-3.11|https://github.com/beobal/cassandra/tree/12796-3.11]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.11-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.11-dtest]|\n|[12796-3.X|https://github.com/beobal/cassandra/tree/12796-3.X]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.X-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-3.X-dtest]|\n|[12796-trunk|https://github.com/beobal/cassandra/tree/12796-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/beobal/job/beobal-12796-trunk-dtest]|\n","from":"developer"},{"body":"The CI looks reasonable: 3 dtest failures on the 3.0 branch, which all have corresponding failures upstream, plus a couple of failures on the original 2.2 branch which have since been addressed by other tickets. \n\nThe internal paging will aim to read rows in chunks of ~4mb and I've used the default CQL page size of 10k rows for a floor as it seems like as good a place to start as any. \n\n[~mmajercik], [~anmols] how does this latest 3.0 version look with your testing? ","from":"developer"},{"body":"[~beobal], the patch for branch _3.0_ works fine, however the page size for single partition pager appears to be calculated incorrectly. My table's average partition size is around *7GB* and yet the page size got calculated as *1*.\n\n{code:java}\n private int calculateIndexingPageSize()\n {\n double averageRowSize = baseCfs.getMeanPartitionSize();\n if (averageRowSize <= 0)\n return DEFAULT_PAGE_SIZE;\n\n return (int) Math.max(1, Math.min(DEFAULT_PAGE_SIZE, 4 * 1024 * 1024 / averageRowSize));\n }\n{code}\n\nThis rendered index rebuild extremely slow as registering read/write order group implies significant performance overhead and for this reason the page size should have reasonable value.\n\nI think there is no harm if we set page size to {{DEFAULT_PAGE_SIZE}} as the pager doesn't span across different partitions in case the partition is small ([https://github.com/mmajercik/cassandra/commit/3fc016e73d3032f4d04584a45945141151a49213])\n\n[12796-3.0|https://github.com/mmajercik/cassandra/tree/12796-3.0]\n","from":"developer"},{"body":"I think you want (row per part / (mean part size / 4 MB)) if you are guessing at \"rows per 4 MB\"","from":"developer"},{"body":"Yep, the method was totally wrong. I'm not 100% comfortable with hardcoding the page size as suggested so I'm just running some quick experiments to figure out a reasonable X for \"rows per X MB\".","from":"developer"},{"body":"Pushed another commit with an implementation of {{calculateIndexingPageSize}} which actually works. It isn't perfect of course, but when the average row size is smallish, it will tend to use the default page size of 10000, which is no problem as the each ordering will still only be held open for a matter of milliseconds. When rows are larger though, that default page size could cause problems as the orderings will be held open for much longer, so this will help to keep that window more consistently bounded. \n\nI've also made the page size an argument to {{indexPartition}} so we don't perform the redundant calculation for every partition. \n","from":"developer"},{"body":"Excellent. This is the closest we can get to keep target page size in bytes. Do you want me to test it with our test data? If not, could it be merged into the primary repository?","from":"developer"},{"body":"If you could double check with your test data, that'd be awesome. I'm just waiting for the latest CI runs to finish, so if all looks good with that & your test data I'll commit asap.","from":"developer"},{"body":"I think the new algorithm looks good. It has a much better chance of doing the right thing. Couple nits on just that commit, I would put the table name in the TRACE log message, and you might add the new -D to the jvm.options file in 3.11+.","from":"developer"},{"body":"[~beobal], sorry for delay in response. Just tested your fix on our sample data and it works fine. Thank you for your help.","from":"developer"},{"body":"Thanks, committed to 2.2 in {{fb2940050e27a5642a23f3e9b5aaa7ae65e018b1}} and to 3.0 (with the additional log info) in {{36ce4e02b429b1297d71c5c8a963623c62d9e159}}. Also added the new -D switch to {{jvm.options}} from 3.10 (i.e. the cassandra-3.11 branch) onwards.","from":"developer"}],"created":"2016-10-17T09:10:54.000+0000","description":"We have a table with rather wide partition and a secondary index defined over it. As soon as we try to rebuild the index we observed exhaustion of Java heap and eventual OOM error. After a lengthy investigation we have managed to find a culprit which appears to be a wrong granule of barrier issuances in method {{org.apache.cassandra.db.Keyspace.indexRow}}:\n\n{code}\n try (OpOrder.Group opGroup = cfs.keyspace.writeOrder.start()){html}\n {\n Set indexes = cfs.indexManager.getIndexesByNames(idxNames);\n\n Iterator pager = QueryPagers.pageRowLocally(cfs, key.getKey(), DEFAULT_PAGE_SIZE);\n while (pager.hasNext())\n {\n ColumnFamily cf = pager.next();\n ColumnFamily cf2 = cf.cloneMeShallow();\n for (Cell cell : cf)\n {\n if (cfs.indexManager.indexes(cell.name(), indexes))\n cf2.addColumn(cell);\n }\n cfs.indexManager.indexRow(key.getKey(), cf2, opGroup);\n }\n }\n{code}\n\nPlease note the operation group granule is a partition of the source table which poses a problem for wide partition tables as flush runnable ({{org.apache.cassandra.db.ColumnFamilyStore.Flush.run()}}) won't proceed with flushing secondary index memtable before completing operations prior recent issue of the barrier. In our situation the flush runnable waits until whole wide partition gets indexed into the secondary index memtable before flushing it. This causes an exhaustion of the heap and eventual OOM error.\n\nAfter we changed granule of barrier issue in method {{org.apache.cassandra.db.Keyspace.indexRow}} to query page as opposed to table partition secondary index (see [https://github.com/mmajercik/cassandra/commit/7e10e5aa97f1de483c2a5faf867315ecbf65f3d6?diff=unified]), rebuild started to work without heap exhaustion. ","issue_id":"13012791","key":"CASSANDRA-12796","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-12-13T10:41:54.000+0000","role":"fixed_distractor","summary":"Heap exhaustion when rebuilding secondary index over a table with wide partitions"} {"case_id":"13023365","cluster":"DISTRACTOR-CASSANDRA-12956","comments":[{"body":"The problem is only present in Cassandra starting from 3.0. Versions before that will replay commit log despite the exception, possibly generating multiple indentical sstables.","created":"2016-11-29T13:41:09.887+0000"},{"body":"Patch for {{3.0}} is quite different and is much bigger. Main problem is that there's no transactionality on the same level as in {{3.X}}. {{3.0}} memtables are flushed and renamed to non-tmp names, readers are returned. We need a bit better granularity, since after we may have to abort all the flushed sstables if 2i failed. I've changed it a bit in {{3.x}} fashion, although since we flush to just one sstable, I thought that extracting {{txn}} to the top level will not give us anything.\n\nBoth patches introduce the second latch. I'm usually not the biggest fan of two threads that have to wait for one another, but here the ordering is an issue. Problem is that post-flush executor is single-threaded (for ordering), and flush executor is multi-threaded, so we can't return future backed with that multi-threaded executor as it will break order. On the other hand, if we move 2i flush to flush executor, we'll have to sequentially wait for 2i, then all memtables. Current approach allows to keep these actions parallel. \n\nWe only need to synchronise the non-cf 2i flush with memtable holding data for current cf. All the cf-index memtables will be in sync with data one anyways since they're combined in the transaction. \n\n|[3.X|https://github.com/ifesdjeen/cassandra/tree/12956-3.X]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.X-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.X-dtest/]|\n|[3.0|https://github.com/ifesdjeen/cassandra/tree/12956-3.0-v2]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.0-v2-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.0-v2-dtest/]|\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/12956-trunk]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-trunk-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-trunk-dtest/]|\n","created":"2016-11-30T12:28:24.314+0000"},{"body":"[~blambov] to review.","created":"2016-11-30T18:07:13.737+0000"},{"body":"Looking at the 3.X patch, as flush threads waiting for post-flush waiting for flush threads is a recipe for disaster (poor performance and deadlocks in particular), I would much prefer the 2i flush to be done on a different thread. In particular, as the flush runnables doing the actual work proceed on their per-disk executors, the flush thread itself [here|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:12956-3.X#diff-98f5acb96aa6d684781936c141132e2aR1152] looks like a better candidate.\n\nCould the inverse of the current problem also cause issues (i.e. 2i flushing without the data being in sstables)? It appears that the 2i flush also needs to eventually become transactional -- doing the above would make that easier too.","created":"2016-12-01T10:05:10.810+0000"},{"body":"Thanks for the review!\n\nI was also skeptical about two threads waiting for one another. Also, tried the approach you've suggested.\nI hesitated mostly because it'd be blocking the flush thread (although you're right that it's going to be\nwaiting for flushes anyways) and because {{flushMemtable}} is called from loop, so I wasn't sure if it's\na good place.\n\nIn retrospect, I think your suggestion is much better than the previous version. I've re-implemented a patch\nfor 3.0 as well, it got much simpler. Now, we do 2i flush before memtable flush during flush of \"data\" memtable\n(first one). We could bring back changes that expose sstable writer and run 2i flush on the other\nthread and commit sstable only on successful 2i flush, although since all cf memtables are flushed sequentially\nanyways and it might be a bit out of scope of the bugfix I decided to leave it this way. Simply running 2i\nflush in a different thread is not enough, as we need to ensure it's in sync with \"data\" memtable flush.\n\nOrder of 2i/memtable flush does not matter, as for 2i it's only important that data is present either in\nsstable or in memtable. We can have the following situations: flush running (memtable is queried), flush\nsuccessfull (sstable is queried), flush unsuccessful (memtable is queried), flush unsuccessful + node restarted\n(CL will replay the data and it'll be available from memtable again). So 3.0 patch relies on this behaviour.\n\n|[3.0|https://github.com/ifesdjeen/cassandra/tree/12956-3.0-v2]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.0-v2-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.0-v2-dtest/]|\n|[3.X|https://github.com/ifesdjeen/cassandra/tree/12956-3.X]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.X-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.X-dtest/]|\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/12956-trunk]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-trunk-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-trunk-dtest/]|\n","created":"2016-12-01T20:14:31.709+0000"},{"body":"Thanks, much cleaner and safer indeed. One small issue and a nit:\n- The [2i flush|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:12956-3.X#diff-98f5acb96aa6d684781936c141132e2aR1082] should only be done if {{truncate}} is false.\n- The [barrier await|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:12956-3.X#diff-98f5acb96aa6d684781936c141132e2aR1130] is not necessary as the flush does not commence [until that has happened|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:12956-3.X#diff-98f5acb96aa6d684781936c141132e2aR1071].\n","created":"2016-12-02T09:14:55.865+0000"},{"body":"Great, thank you.\nI've removed the duplicate {{await}}, thanks for catching that. {{truncate}} check for {{false}} is done in [flush memtable|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:12956-3.X#diff-98f5acb96aa6d684781936c141132e2aR1097].\n\nI've applied the change to all branches and did CI.","created":"2016-12-05T15:57:50.819+0000"},{"body":"Thanks, committed as 6f90e55e7e23cbe814a3232c8d1ec67f2ff2a537.","created":"2016-12-06T11:22:33.627+0000"},{"body":"Thank you!","created":"2016-12-07T11:57:39.171+0000"},{"body":"What are the fix versions for this?","created":"2016-12-09T22:40:07.797+0000"}],"conversations":[{"body":"If during the node shutdown / drain the custom (non-cf) 2i throws an exception, CommitLog will get correctly preserved (segments won't get discarded because segment tracking is correct). \n\nHowever, when it gets replayed on node startup, we're making a decision whether or not to replay the commit log. CL segment starts getting replayed, since there are non-discarded segments and during this process we're checking whether there every [individual mutation|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/db/commitlog/CommitLogReplayer.java#L215] in commit log is already committed or no. Information about the sstables is taken from [live sstables on disk|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/db/commitlog/CommitLogReplayer.java#L250-L256].\n","from":"reporter","subject":"CL is not replayed on custom 2i exception"},{"body":"The problem is only present in Cassandra starting from 3.0. Versions before that will replay commit log despite the exception, possibly generating multiple indentical sstables.","from":"developer"},{"body":"Patch for {{3.0}} is quite different and is much bigger. Main problem is that there's no transactionality on the same level as in {{3.X}}. {{3.0}} memtables are flushed and renamed to non-tmp names, readers are returned. We need a bit better granularity, since after we may have to abort all the flushed sstables if 2i failed. I've changed it a bit in {{3.x}} fashion, although since we flush to just one sstable, I thought that extracting {{txn}} to the top level will not give us anything.\n\nBoth patches introduce the second latch. I'm usually not the biggest fan of two threads that have to wait for one another, but here the ordering is an issue. Problem is that post-flush executor is single-threaded (for ordering), and flush executor is multi-threaded, so we can't return future backed with that multi-threaded executor as it will break order. On the other hand, if we move 2i flush to flush executor, we'll have to sequentially wait for 2i, then all memtables. Current approach allows to keep these actions parallel. \n\nWe only need to synchronise the non-cf 2i flush with memtable holding data for current cf. All the cf-index memtables will be in sync with data one anyways since they're combined in the transaction. \n\n|[3.X|https://github.com/ifesdjeen/cassandra/tree/12956-3.X]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.X-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.X-dtest/]|\n|[3.0|https://github.com/ifesdjeen/cassandra/tree/12956-3.0-v2]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.0-v2-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.0-v2-dtest/]|\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/12956-trunk]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-trunk-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-trunk-dtest/]|\n","from":"developer"},{"body":"[~blambov] to review.","from":"developer"},{"body":"Looking at the 3.X patch, as flush threads waiting for post-flush waiting for flush threads is a recipe for disaster (poor performance and deadlocks in particular), I would much prefer the 2i flush to be done on a different thread. In particular, as the flush runnables doing the actual work proceed on their per-disk executors, the flush thread itself [here|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:12956-3.X#diff-98f5acb96aa6d684781936c141132e2aR1152] looks like a better candidate.\n\nCould the inverse of the current problem also cause issues (i.e. 2i flushing without the data being in sstables)? It appears that the 2i flush also needs to eventually become transactional -- doing the above would make that easier too.","from":"developer"},{"body":"Thanks for the review!\n\nI was also skeptical about two threads waiting for one another. Also, tried the approach you've suggested.\nI hesitated mostly because it'd be blocking the flush thread (although you're right that it's going to be\nwaiting for flushes anyways) and because {{flushMemtable}} is called from loop, so I wasn't sure if it's\na good place.\n\nIn retrospect, I think your suggestion is much better than the previous version. I've re-implemented a patch\nfor 3.0 as well, it got much simpler. Now, we do 2i flush before memtable flush during flush of \"data\" memtable\n(first one). We could bring back changes that expose sstable writer and run 2i flush on the other\nthread and commit sstable only on successful 2i flush, although since all cf memtables are flushed sequentially\nanyways and it might be a bit out of scope of the bugfix I decided to leave it this way. Simply running 2i\nflush in a different thread is not enough, as we need to ensure it's in sync with \"data\" memtable flush.\n\nOrder of 2i/memtable flush does not matter, as for 2i it's only important that data is present either in\nsstable or in memtable. We can have the following situations: flush running (memtable is queried), flush\nsuccessfull (sstable is queried), flush unsuccessful (memtable is queried), flush unsuccessful + node restarted\n(CL will replay the data and it'll be available from memtable again). So 3.0 patch relies on this behaviour.\n\n|[3.0|https://github.com/ifesdjeen/cassandra/tree/12956-3.0-v2]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.0-v2-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.0-v2-dtest/]|\n|[3.X|https://github.com/ifesdjeen/cassandra/tree/12956-3.X]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.X-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-3.X-dtest/]|\n|[trunk|https://github.com/ifesdjeen/cassandra/tree/12956-trunk]|[utest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-trunk-testall/]|[dtest|https://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-12956-trunk-dtest/]|\n","from":"developer"},{"body":"Thanks, much cleaner and safer indeed. One small issue and a nit:\n- The [2i flush|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:12956-3.X#diff-98f5acb96aa6d684781936c141132e2aR1082] should only be done if {{truncate}} is false.\n- The [barrier await|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:12956-3.X#diff-98f5acb96aa6d684781936c141132e2aR1130] is not necessary as the flush does not commence [until that has happened|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:12956-3.X#diff-98f5acb96aa6d684781936c141132e2aR1071].\n","from":"developer"},{"body":"Great, thank you.\nI've removed the duplicate {{await}}, thanks for catching that. {{truncate}} check for {{false}} is done in [flush memtable|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:12956-3.X#diff-98f5acb96aa6d684781936c141132e2aR1097].\n\nI've applied the change to all branches and did CI.","from":"developer"},{"body":"Thanks, committed as 6f90e55e7e23cbe814a3232c8d1ec67f2ff2a537.","from":"developer"},{"body":"Thank you!","from":"developer"},{"body":"What are the fix versions for this?","from":"developer"}],"created":"2016-11-25T11:12:44.000+0000","description":"If during the node shutdown / drain the custom (non-cf) 2i throws an exception, CommitLog will get correctly preserved (segments won't get discarded because segment tracking is correct). \n\nHowever, when it gets replayed on node startup, we're making a decision whether or not to replay the commit log. CL segment starts getting replayed, since there are non-discarded segments and during this process we're checking whether there every [individual mutation|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/db/commitlog/CommitLogReplayer.java#L215] in commit log is already committed or no. Information about the sstables is taken from [live sstables on disk|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/db/commitlog/CommitLogReplayer.java#L250-L256].\n","issue_id":"13023365","key":"CASSANDRA-12956","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-12-07T11:57:38.000+0000","role":"fixed_distractor","summary":"CL is not replayed on custom 2i exception"} {"case_id":"13028717","cluster":"DISTRACTOR-CASSANDRA-13052","comments":[{"body":"Cristian, please be aware that token ranges are start-exclusive and end-inclusive. The result of a midpoint calculation is probably undefined, in case an invalid range has been specified.\n","created":"2016-12-19T12:19:12.842+0000"},{"body":"Stefan, the above code is to highlight the root cause for the problem not the actual range. If you provide a small range the method provided in the above description will try, recursively, to divide the ranges until left and right token have same value. Next recursive iteration will generate a midpoint token way out of the suggested repair range.\n\nHere is an example (C* 2.0.14):\n\nSee the token range provided for repair: (7792951013348769424,7792951013348769525]. That's 100 tokens.\nBut below you can see the Differencer.java: \"have 119 range(s) out of sync for testCF\". I would say you cannot split 100 tokens in 119 ranges.\n\nINFO [AntiEntropySessions:1] 2016-12-15 14:52:51,951 RepairSession.java (line 246) [repair #27437120-c2d6-11e6-b49f-8b496c707234] new session: will sync /127.0.0.1, /127.0.0.2, /127.0.0.3 on range (7792951013348769424,7792951013348769525] for testKS.[testCF]\nINFO [AntiEntropySessions:1] 2016-12-15 14:52:51,960 RepairJob.java (line 161) [repair #27437120-c2d6-11e6-b49f-8b496c707234] requesting merkle trees for testCF (to [/127.0.0.2, /127.0.0.3, /127.0.0.1])\nINFO [AntiEntropyStage:1] 2016-12-15 14:52:52,054 RepairSession.java (line 166) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Received merkle tree for testCF from /127.0.0.2\nINFO [AntiEntropyStage:1] 2016-12-15 14:52:52,064 RepairSession.java (line 166) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Received merkle tree for testCF from /127.0.0.1\nINFO [AntiEntropyStage:1] 2016-12-15 14:52:52,065 RepairSession.java (line 166) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Received merkle tree for testCF from /127.0.0.3\nINFO [RepairJobTask:1] 2016-12-15 14:52:52,071 Differencer.java (line 67) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Endpoints /127.0.0.2 and /127.0.0.1 are consistent for testCF\nINFO [RepairJobTask:3] 2016-12-15 14:52:52,105 Differencer.java (line 74) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Endpoints /127.0.0.1 and /127.0.0.3 have 119 range(s) out of sync for testCF\nINFO [RepairJobTask:2] 2016-12-15 14:52:52,108 Differencer.java (line 74) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Endpoints /127.0.0.2 and /127.0.0.3 have 119 range(s) out of sync for testCF\nINFO [RepairJobTask:2] 2016-12-15 14:52:52,110 StreamingRepairTask.java (line 77) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Forwarding streaming repair of 119 ranges to /127.0.0.2 (to be streamed with /127.0.0.3)\nINFO [RepairJobTask:3] 2016-12-15 14:52:52,118 StreamingRepairTask.java (line 64) [streaming task #27437120-c2d6-11e6-b49f-8b496c707234] Performing streaming repair of 119 ranges with /127.0.0.3\nINFO [STREAM-IN-/127.0.0.3] 2016-12-15 14:52:53,363 StreamingRepairTask.java (line 92) [repair #27437120-c2d6-11e6-b49f-8b496c707234] streaming task succeed, returning response to /127.0.0.1\nINFO [AntiEntropyStage:1] 2016-12-15 14:52:53,372 RepairSession.java (line 223) [repair #27437120-c2d6-11e6-b49f-8b496c707234] testCF is fully synced\nINFO [AntiEntropySessions:1] 2016-12-15 14:52:53,373 RepairSession.java (line 284) [repair #27437120-c2d6-11e6-b49f-8b496c707234] session completed successfully\nINFO [Thread-13] 2016-12-15 14:52:53,378 StorageService.java (line 2644) Repair session 27437120-c2d6-11e6-b49f-8b496c707234 for range (7792951013348769424,7792951013348769525] finished","created":"2016-12-19T13:19:12.502+0000"},{"body":"I've now reproduced the described situation locally, but I'm not fully sure about the implications. It's certainly not the intended behavior, so thanks a lot for reporting!\n","created":"2016-12-19T15:53:06.274+0000"},{"body":"Please find instructions how to locally reproduce this issue attached in {{ccm_reproduce-13052.txt}}. I've also linked my WIP branch, which includes debug statements for relevant parts of the MerkleTree class that will traverse and hash two trees based on the midpoint calculation. See {{system-dev-debug-13052.log}} for example log output.","created":"2016-12-20T13:53:43.605+0000"},{"body":"My [WIP branch|https://github.com/apache/cassandra/compare/cassandra-2.1...spodkowinski:WIP-13052] has now been updated with an exit condition that should avoid running into the situation to have the midpoint method called with an invalid range.\n\nThe current MT code recursively looks for the node closest to the leafs by continually dividing the range by it's midpoint and returning at a point when any child is not contained anymore in the range we're looking for. Unfortunately this condition will not apply when we have a leaf in the MT that exactly spans the range of a single token we're searching. See [MerkleTree.findHelper|https://github.com/apache/cassandra/blob/4a2464192e9e69457f5a5ecf26c094f9298bf069/src/java/org/apache/cassandra/utils/MerkleTree.java#L420].\n\nI still want to investigate how likely this going to happen for larger ranges, e.g. complete vnodes.","created":"2016-12-21T12:32:16.767+0000"},{"body":"{quote}\nWe soon notice heavy streaming and according to the logs the number of ranges streamed was in thousands.\n{quote}\n\n[~cmposto], I'm currently a bit confused while trying to evaluate the actual effect of this behaviour. I was first a bit concerned about either having ranges streamed that shouldn't be repaired or to send identical ranges thousand of times. But I've come to the conclusion that none of this should happen, as {{StreamSession.addTransferRanges()}} should normalize all redundant ranges into a singe range, before starting to stream any files. Can you further describe how this bug resulted into \"heavy streaming\" on your cluster? What was is that you noticed before further looking into this issue?\n","created":"2017-01-03T15:20:48.711+0000"},{"body":"Hi Stefan,\n\nIn our PROD cluster (version 2.0.14) we saw large amount of data being streamed for a single token range (order of GBs). When we run a repair for a full vnode range we see less data being streamed (order of MBs).\nI would suggest to log the ranges streamed for the above test. Worth doing after the call to normalize.\n\nRegards,\nCristian","created":"2017-01-04T14:32:24.277+0000"},{"body":"I've just checked your attached log file and if I understood correctly I see you tried to run a repair for a range of 10 tokens only: (2321271983248423860,2321271983248423870].\nBut if you see the ranges marked as inconsistent in the logs you will find something like below. The start token in below example (-6902100053606351945) is way out of the suggested range for repair.\n\nDEBUG [RepairJobTask:2] 2016-12-20 14:24:26,756 MerkleTree.java:298 - (126) Left sub-range fully inconsistent #\n\nDEBUG [RepairJobTask:2] 2016-12-20 14:24:26,757 MerkleTree.java:319 - (126) Right sub-range fully inconsistent #\n\nDEBUG [RepairJobTask:2] 2016-12-20 14:24:26,759 MerkleTree.java:326 - (126) Fully inconsistent range [#, #]\n","created":"2017-01-04T14:49:00.239+0000"},{"body":"The issue is not that ranges will be added multiple times, but the fact that the added range has an equal start and endpoint and is thus invalid by definition. Things starts to go awry in the normalize method, which will handle the range as a wrap-around value and turns the specified range into another range of (MINIMUM, MINIMUM). This range is handled special in some occassions and seems to cause {{DataTracker.sstablesInBounds()}} return basically all sstables. Based on those tables, the SSTable section calculation will return all content from each SSTable as well and so it looks to me that this will cause streaming of all data. \n\n","created":"2017-01-05T12:01:05.971+0000"},{"body":"Which repairs are affected? The tree depth will effectively calculated as {{min(log(partitions), 20)}}, ie. the depth will increase logarithmic until capped at 20. In the worst case scenario, the described bug will happen when the repaired token range will be less than 2^20. Using the Murmur3Partitioner, this is still less than individual ranges created by using vnodes in a large cluster {{(2^64 / (vnodes * nodes))}}. As vnodes are created based on random values, there's still a chance that two vnodes are close enough together to get affected by this bug regulary. Although this requires a fairly large amount of partitions and I'm wondering if it would be practically possible to run repairs in this setup anyways. More likely this could be an issue for tools that would work by calculating their own (smaller) ranges instead of running repairs based on (v)nodes.","created":"2017-01-05T12:03:40.979+0000"},{"body":"I've started tests for my attached patch.\n\n||2.1||2.2||3.0||3.x||trunk||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-2.1]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-2.2]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.0]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.x]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-trunk]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-2.1-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-2.2-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-3.x-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-trunk-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-2.1-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-2.2-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-3.x-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-trunk-testall/]|\n\nHappy to contribute a dtest as well if someone would step up reviewing the patch.\n\n","created":"2017-01-05T15:37:16.583+0000"},{"body":"[~spodxx@gmail.com] this LGTM. If you don't mind, I've pushed up 2 changes to my repo [here|https://github.com/bdeggleston/cassandra/tree/13052-3.0]. I've added a unit test, and added a check which prevents using the same start and end tokens for a subrange repair. \n\nAlso, I think we should probably only apply this to 3.0 and up. It's not really critical enough to apply to 2.x imo.","created":"2017-05-09T15:04:55.487+0000"},{"body":"Also vote for 3.0 and newer - it seems straight forward, but it's been like this a very long time. \n","created":"2017-05-09T23:35:29.776+0000"},{"body":"[~spodxx@gmail.com] any thoughts on this?","created":"2017-05-17T16:42:37.455+0000"},{"body":"I've merged your commits, rebased and fired up tests a couple of hours ago. \n\n* trunk [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-trunk] [testall|https://circleci.com/gh/spodkowinski/cassandra/47] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/53/]\n* 3.11 [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.11] [testall|https://circleci.com/gh/spodkowinski/cassandra/45] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/54/]\n* 3.0 [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.0] [testall|https://circleci.com/gh/spodkowinski/cassandra/46] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/51/]\n\nI'd be fine to see this committed if tests don't indicate any other issues.","created":"2017-05-17T17:05:23.690+0000"},{"body":"Cool, +1 then. -What about only applying this to trunk?-\n\nEDIT: never mind the trunk only part, was thinking of another ticket :)","created":"2017-05-17T17:07:48.117+0000"},{"body":"Looks like [RepairOptionTest.java#L154|https://github.com/apache/cassandra/blob/cassandra-3.11/test/unit/org/apache/cassandra/repair/messages/RepairOptionTest.java#L154] uses a range that will now cause an \"java.lang.IllegalArgumentException: Start and end tokens must be different.\" error. If we define a token range to be start-exclusive and end-inclusive, creating a range based on a single token shouldn't be possible and we shouldn't allow \"42:42\" as \"ranges\" argument. The correct range for a single token in this case would be \"41:42\". \n\n/cc [~aweisberg]","created":"2017-05-18T08:24:25.840+0000"},{"body":"That makes sense. IIRC, a range with the same start and end token means 'the entire token range', which is unlikely to be what the person doing the repairing wanted. In the case of this test though, it's only checking that providing a subrange causes isGlobal to return false... so it should be ok to change it to \"41:42\"","created":"2017-05-18T17:27:59.584+0000"},{"body":"Mentioned test has been fixed and re-run, along with the aborted dtest. All tests have finished and reported errors look unrelated to the patch.\n\n* trunk [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-trunk] [testall|https://circleci.com/gh/spodkowinski/cassandra/47] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/53/]\n* 3.11 [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.11] [testall|https://circleci.com/gh/spodkowinski/cassandra/48] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/57/]\n* 3.0 [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.0] [testall|https://circleci.com/gh/spodkowinski/cassandra/49] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/51/]\n","created":"2017-05-19T10:13:28.584+0000"},{"body":"Merged as 38725802456dfcfcc2584cdfd061ac22215c2dfb to 3.0 -> 3.11 -> 4.0","created":"2017-05-24T12:57:01.629+0000"}],"conversations":[{"body":"We tried to do a single token repair by providing 2 consecutive token values for a large column family. We soon notice heavy streaming and according to the logs the number of ranges streamed was in thousands.\nAfter investigation we found a bug in the two partitioner classes we use (RandomPartitioner and Murmur3Partitioner).\nThe midpoint method used by MerkleTree.differenceHelper method to find ranges with differences for streaming returns abnormal values (way out of the initial range requested for repair) if the repair requested range is small (I expect smaller than 2^15).\nHere is the simple code to reproduce the bug for Murmur3Partitioner:\n\nToken left = new Murmur3Partitioner.LongToken(123456789L);\nToken right = new Murmur3Partitioner.LongToken(123456789L);\nIPartitioner partitioner = new Murmur3Partitioner();\nToken midpoint = partitioner.midpoint(left, right);\nSystem.out.println(\"Murmur3: [ \" + left.getToken() + \" : \" + midpoint.getToken() + \" : \" + right.getToken() + \" ]\");\n\nThe output is:\nMurmur3: [ 123456789 : -9223372036731319019 : 123456789 ]\n\nNote that the midpoint token is nowhere near the suggested repair range. This will happen if during the parsing of the tree (in MerkleTree.differenceHelper) in search for differences there isn't enough tokens for the split and the subrange becomes 0 (left.token=right.token) as in the above test.","from":"reporter","subject":"Repair process is violating the start/end token limits for small ranges"},{"body":"Cristian, please be aware that token ranges are start-exclusive and end-inclusive. The result of a midpoint calculation is probably undefined, in case an invalid range has been specified.\n","from":"developer"},{"body":"Stefan, the above code is to highlight the root cause for the problem not the actual range. If you provide a small range the method provided in the above description will try, recursively, to divide the ranges until left and right token have same value. Next recursive iteration will generate a midpoint token way out of the suggested repair range.\n\nHere is an example (C* 2.0.14):\n\nSee the token range provided for repair: (7792951013348769424,7792951013348769525]. That's 100 tokens.\nBut below you can see the Differencer.java: \"have 119 range(s) out of sync for testCF\". I would say you cannot split 100 tokens in 119 ranges.\n\nINFO [AntiEntropySessions:1] 2016-12-15 14:52:51,951 RepairSession.java (line 246) [repair #27437120-c2d6-11e6-b49f-8b496c707234] new session: will sync /127.0.0.1, /127.0.0.2, /127.0.0.3 on range (7792951013348769424,7792951013348769525] for testKS.[testCF]\nINFO [AntiEntropySessions:1] 2016-12-15 14:52:51,960 RepairJob.java (line 161) [repair #27437120-c2d6-11e6-b49f-8b496c707234] requesting merkle trees for testCF (to [/127.0.0.2, /127.0.0.3, /127.0.0.1])\nINFO [AntiEntropyStage:1] 2016-12-15 14:52:52,054 RepairSession.java (line 166) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Received merkle tree for testCF from /127.0.0.2\nINFO [AntiEntropyStage:1] 2016-12-15 14:52:52,064 RepairSession.java (line 166) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Received merkle tree for testCF from /127.0.0.1\nINFO [AntiEntropyStage:1] 2016-12-15 14:52:52,065 RepairSession.java (line 166) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Received merkle tree for testCF from /127.0.0.3\nINFO [RepairJobTask:1] 2016-12-15 14:52:52,071 Differencer.java (line 67) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Endpoints /127.0.0.2 and /127.0.0.1 are consistent for testCF\nINFO [RepairJobTask:3] 2016-12-15 14:52:52,105 Differencer.java (line 74) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Endpoints /127.0.0.1 and /127.0.0.3 have 119 range(s) out of sync for testCF\nINFO [RepairJobTask:2] 2016-12-15 14:52:52,108 Differencer.java (line 74) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Endpoints /127.0.0.2 and /127.0.0.3 have 119 range(s) out of sync for testCF\nINFO [RepairJobTask:2] 2016-12-15 14:52:52,110 StreamingRepairTask.java (line 77) [repair #27437120-c2d6-11e6-b49f-8b496c707234] Forwarding streaming repair of 119 ranges to /127.0.0.2 (to be streamed with /127.0.0.3)\nINFO [RepairJobTask:3] 2016-12-15 14:52:52,118 StreamingRepairTask.java (line 64) [streaming task #27437120-c2d6-11e6-b49f-8b496c707234] Performing streaming repair of 119 ranges with /127.0.0.3\nINFO [STREAM-IN-/127.0.0.3] 2016-12-15 14:52:53,363 StreamingRepairTask.java (line 92) [repair #27437120-c2d6-11e6-b49f-8b496c707234] streaming task succeed, returning response to /127.0.0.1\nINFO [AntiEntropyStage:1] 2016-12-15 14:52:53,372 RepairSession.java (line 223) [repair #27437120-c2d6-11e6-b49f-8b496c707234] testCF is fully synced\nINFO [AntiEntropySessions:1] 2016-12-15 14:52:53,373 RepairSession.java (line 284) [repair #27437120-c2d6-11e6-b49f-8b496c707234] session completed successfully\nINFO [Thread-13] 2016-12-15 14:52:53,378 StorageService.java (line 2644) Repair session 27437120-c2d6-11e6-b49f-8b496c707234 for range (7792951013348769424,7792951013348769525] finished","from":"developer"},{"body":"I've now reproduced the described situation locally, but I'm not fully sure about the implications. It's certainly not the intended behavior, so thanks a lot for reporting!\n","from":"developer"},{"body":"Please find instructions how to locally reproduce this issue attached in {{ccm_reproduce-13052.txt}}. I've also linked my WIP branch, which includes debug statements for relevant parts of the MerkleTree class that will traverse and hash two trees based on the midpoint calculation. See {{system-dev-debug-13052.log}} for example log output.","from":"developer"},{"body":"My [WIP branch|https://github.com/apache/cassandra/compare/cassandra-2.1...spodkowinski:WIP-13052] has now been updated with an exit condition that should avoid running into the situation to have the midpoint method called with an invalid range.\n\nThe current MT code recursively looks for the node closest to the leafs by continually dividing the range by it's midpoint and returning at a point when any child is not contained anymore in the range we're looking for. Unfortunately this condition will not apply when we have a leaf in the MT that exactly spans the range of a single token we're searching. See [MerkleTree.findHelper|https://github.com/apache/cassandra/blob/4a2464192e9e69457f5a5ecf26c094f9298bf069/src/java/org/apache/cassandra/utils/MerkleTree.java#L420].\n\nI still want to investigate how likely this going to happen for larger ranges, e.g. complete vnodes.","from":"developer"},{"body":"{quote}\nWe soon notice heavy streaming and according to the logs the number of ranges streamed was in thousands.\n{quote}\n\n[~cmposto], I'm currently a bit confused while trying to evaluate the actual effect of this behaviour. I was first a bit concerned about either having ranges streamed that shouldn't be repaired or to send identical ranges thousand of times. But I've come to the conclusion that none of this should happen, as {{StreamSession.addTransferRanges()}} should normalize all redundant ranges into a singe range, before starting to stream any files. Can you further describe how this bug resulted into \"heavy streaming\" on your cluster? What was is that you noticed before further looking into this issue?\n","from":"developer"},{"body":"Hi Stefan,\n\nIn our PROD cluster (version 2.0.14) we saw large amount of data being streamed for a single token range (order of GBs). When we run a repair for a full vnode range we see less data being streamed (order of MBs).\nI would suggest to log the ranges streamed for the above test. Worth doing after the call to normalize.\n\nRegards,\nCristian","from":"developer"},{"body":"I've just checked your attached log file and if I understood correctly I see you tried to run a repair for a range of 10 tokens only: (2321271983248423860,2321271983248423870].\nBut if you see the ranges marked as inconsistent in the logs you will find something like below. The start token in below example (-6902100053606351945) is way out of the suggested range for repair.\n\nDEBUG [RepairJobTask:2] 2016-12-20 14:24:26,756 MerkleTree.java:298 - (126) Left sub-range fully inconsistent #\n\nDEBUG [RepairJobTask:2] 2016-12-20 14:24:26,757 MerkleTree.java:319 - (126) Right sub-range fully inconsistent #\n\nDEBUG [RepairJobTask:2] 2016-12-20 14:24:26,759 MerkleTree.java:326 - (126) Fully inconsistent range [#, #]\n","from":"developer"},{"body":"The issue is not that ranges will be added multiple times, but the fact that the added range has an equal start and endpoint and is thus invalid by definition. Things starts to go awry in the normalize method, which will handle the range as a wrap-around value and turns the specified range into another range of (MINIMUM, MINIMUM). This range is handled special in some occassions and seems to cause {{DataTracker.sstablesInBounds()}} return basically all sstables. Based on those tables, the SSTable section calculation will return all content from each SSTable as well and so it looks to me that this will cause streaming of all data. \n\n","from":"developer"},{"body":"Which repairs are affected? The tree depth will effectively calculated as {{min(log(partitions), 20)}}, ie. the depth will increase logarithmic until capped at 20. In the worst case scenario, the described bug will happen when the repaired token range will be less than 2^20. Using the Murmur3Partitioner, this is still less than individual ranges created by using vnodes in a large cluster {{(2^64 / (vnodes * nodes))}}. As vnodes are created based on random values, there's still a chance that two vnodes are close enough together to get affected by this bug regulary. Although this requires a fairly large amount of partitions and I'm wondering if it would be practically possible to run repairs in this setup anyways. More likely this could be an issue for tools that would work by calculating their own (smaller) ranges instead of running repairs based on (v)nodes.","from":"developer"},{"body":"I've started tests for my attached patch.\n\n||2.1||2.2||3.0||3.x||trunk||\n|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-2.1]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-2.2]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.0]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.x]|[branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-trunk]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-2.1-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-2.2-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-3.0-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-3.x-dtest/]|[dtest|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-trunk-dtest/]|\n|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-2.1-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-2.2-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-3.0-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-3.x-testall/]|[testall|http://cassci.datastax.com/view/Dev/view/spodkowinski/job/spodkowinski-CASSANDRA-13052-trunk-testall/]|\n\nHappy to contribute a dtest as well if someone would step up reviewing the patch.\n\n","from":"developer"},{"body":"[~spodxx@gmail.com] this LGTM. If you don't mind, I've pushed up 2 changes to my repo [here|https://github.com/bdeggleston/cassandra/tree/13052-3.0]. I've added a unit test, and added a check which prevents using the same start and end tokens for a subrange repair. \n\nAlso, I think we should probably only apply this to 3.0 and up. It's not really critical enough to apply to 2.x imo.","from":"developer"},{"body":"Also vote for 3.0 and newer - it seems straight forward, but it's been like this a very long time. \n","from":"developer"},{"body":"[~spodxx@gmail.com] any thoughts on this?","from":"developer"},{"body":"I've merged your commits, rebased and fired up tests a couple of hours ago. \n\n* trunk [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-trunk] [testall|https://circleci.com/gh/spodkowinski/cassandra/47] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/53/]\n* 3.11 [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.11] [testall|https://circleci.com/gh/spodkowinski/cassandra/45] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/54/]\n* 3.0 [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.0] [testall|https://circleci.com/gh/spodkowinski/cassandra/46] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/51/]\n\nI'd be fine to see this committed if tests don't indicate any other issues.","from":"developer"},{"body":"Cool, +1 then. -What about only applying this to trunk?-\n\nEDIT: never mind the trunk only part, was thinking of another ticket :)","from":"developer"},{"body":"Looks like [RepairOptionTest.java#L154|https://github.com/apache/cassandra/blob/cassandra-3.11/test/unit/org/apache/cassandra/repair/messages/RepairOptionTest.java#L154] uses a range that will now cause an \"java.lang.IllegalArgumentException: Start and end tokens must be different.\" error. If we define a token range to be start-exclusive and end-inclusive, creating a range based on a single token shouldn't be possible and we shouldn't allow \"42:42\" as \"ranges\" argument. The correct range for a single token in this case would be \"41:42\". \n\n/cc [~aweisberg]","from":"developer"},{"body":"That makes sense. IIRC, a range with the same start and end token means 'the entire token range', which is unlikely to be what the person doing the repairing wanted. In the case of this test though, it's only checking that providing a subrange causes isGlobal to return false... so it should be ok to change it to \"41:42\"","from":"developer"},{"body":"Mentioned test has been fixed and re-run, along with the aborted dtest. All tests have finished and reported errors look unrelated to the patch.\n\n* trunk [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-trunk] [testall|https://circleci.com/gh/spodkowinski/cassandra/47] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/53/]\n* 3.11 [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.11] [testall|https://circleci.com/gh/spodkowinski/cassandra/48] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/57/]\n* 3.0 [branch|https://github.com/spodkowinski/cassandra/tree/CASSANDRA-13052-3.0] [testall|https://circleci.com/gh/spodkowinski/cassandra/49] [dtest|https://builds.apache.org/user/spod/my-views/view/Cassandra%20List%20View/job/Cassandra-devbranch-dtest/51/]\n","from":"developer"},{"body":"Merged as 38725802456dfcfcc2584cdfd061ac22215c2dfb to 3.0 -> 3.11 -> 4.0","from":"developer"}],"created":"2016-12-16T17:00:37.000+0000","description":"We tried to do a single token repair by providing 2 consecutive token values for a large column family. We soon notice heavy streaming and according to the logs the number of ranges streamed was in thousands.\nAfter investigation we found a bug in the two partitioner classes we use (RandomPartitioner and Murmur3Partitioner).\nThe midpoint method used by MerkleTree.differenceHelper method to find ranges with differences for streaming returns abnormal values (way out of the initial range requested for repair) if the repair requested range is small (I expect smaller than 2^15).\nHere is the simple code to reproduce the bug for Murmur3Partitioner:\n\nToken left = new Murmur3Partitioner.LongToken(123456789L);\nToken right = new Murmur3Partitioner.LongToken(123456789L);\nIPartitioner partitioner = new Murmur3Partitioner();\nToken midpoint = partitioner.midpoint(left, right);\nSystem.out.println(\"Murmur3: [ \" + left.getToken() + \" : \" + midpoint.getToken() + \" : \" + right.getToken() + \" ]\");\n\nThe output is:\nMurmur3: [ 123456789 : -9223372036731319019 : 123456789 ]\n\nNote that the midpoint token is nowhere near the suggested repair range. This will happen if during the parsing of the tree (in MerkleTree.differenceHelper) in search for differences there isn't enough tokens for the split and the subrange becomes 0 (left.token=right.token) as in the above test.","issue_id":"13028717","key":"CASSANDRA-13052","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-05-24T12:57:01.000+0000","role":"fixed_distractor","summary":"Repair process is violating the start/end token limits for small ranges"} {"case_id":"13032647","cluster":"DISTRACTOR-CASSANDRA-13109","comments":[{"body":"Attaching the patch.","created":"2017-01-06T19:53:00.467+0000"},{"body":"Thanks for the report and reproduction steps. I _was_ able to reproduce this consistently with 3.0.10 but this actually happens to be fixed on the current 3.0 branch (future 3.0.11).\n\nAnd the reason this is not a problem anymore is CASSANDRA-12694 and that's because it basically made us fetch all columns for CAS (much like we do for any other CQL query really), thus side-stepping the \"the code to read legacy sstable can return stuffs for non fetched columns\".\n\nThat said, that problem in {{LegacyLayout}} is kind of legit: we should include something that is not fetched and while we correctly skip non-fetched column for \"cells\", we did miss doing the same for collection tombsones. But I say \"kind of\" because thruth is, after CASSANDRA-12694, I'm not sure we can really run into this problem anymore: only thrift queries uses the mode where we don't fetch all columns, but thrift don't support collections so actually cannot create a query that would be problematic for this case.\n\nStill, the code is definitively not doing what it should and the fix is trivial enough that it's not worth risking running into problems in the future, or if I happen to miss a genuine case where this could happen, so I've pushed the patch for CI below and will commit if tests are clear. I did modify said patch a bit because if the column for which we have a collection tombstone is not fetched, we also want to skip the few lines that were before the condition your patch added as otherwise we might end up returning an empty row which would break assumptions of the code. Still, mostly the same thing, just moving the condition a bit up.\n| [13109-3.0|https://github.com/pcmanus/cassandra/commits/13109-3.0] | [utests|http://cassci.datastax.com/job/pcmanus-13109-3.0-testall] | [dtests|http://cassci.datastax.com/job/pcmanus-13109-3.0-dtest] |\n| [13109-3.11|https://github.com/pcmanus/cassandra/commits/13109-3.11] | [utests|http://cassci.datastax.com/job/pcmanus-13109-3.11-testall] | [dtests|http://cassci.datastax.com/job/pcmanus-13109-3.11-dtest] |\n","created":"2017-02-08T10:30:24.639+0000"},{"body":"CI was clean so committed, thanks.","created":"2017-02-09T09:29:20.811+0000"}],"conversations":[{"body":"We've observed this upgrading from 2.1.15 to 3.0.8 and from 2.1.16 to 3.0.10: some lightweight transactions executed on upgraded nodes fail with a read failure. The following conditions seem relevant to this occurring:\n\n* The transaction must be conditioned on the current value of at least one column, e.g., {{IF NOT EXISTS}} transactions don't seem to be affected.\n* There should be a collection column (in our case, a map) defined on the table on which the transaction is executed.\n* The transaction should be executed before sstables on the node are upgraded. The failure does not occur after the sstables have been upgraded (whether via {{nodetool upgradesstables}} or effectively via compaction).\n* Upgraded nodes seem to be able to participate in lightweight transactions as long as they're not the coordinator.\n* The values in the row being manipulated by the transaction must have been consistently manipulated by lightweight transactions (perhaps the existence of Paxos state for the partition is somehow relevant?).\n* In 3.0.10, it _seems_ to be necessary to have the partition split across multiple legacy sstables. This was not necessary to reproduce the bug in 3.0.8 or .9.\n\nFor applications affected by this bug, a possible workaround is to prevent nodes being upgraded from coordinating requests until sstables have been upgraded.\n\nWe're able to reproduce this when upgrading from 2.1.16 to 3.0.10 with the following steps on a single-node cluster using a mostly pristine {{cassandra.yaml}} from the source distribution.\n\n# Start Cassandra-2.1.16 on the node.\n# Create a table with a collection column and insert some data into it.\n{code:sql}\nCREATE KEYSPACE test WITH REPLICATION = {'class': 'SimpleStrategy', 'replication_factor': 1};\nCREATE TABLE test.test (key TEXT PRIMARY KEY, cas_target TEXT, some_collection MAP);\nINSERT INTO test.test (key, cas_target, some_collection) VALUES ('key', 'value', {}) IF NOT EXISTS;\n{code}\n# Flush the row to an sstable: {{nodetool flush}}.\n# Update the row:\n{code:sql}\nUPDATE test.test SET cas_target = 'newvalue', some_collection = {} WHERE key = 'key' IF cas_target = 'value';\n{code}\n# Drain the node: {{nodetool drain}}\n# Stop the node, upgrade to 3.0.10, and start the node.\n# Attempt to update the row again:\n{code:sql}\nUPDATE test.test SET cas_target = 'lastvalue' WHERE key = 'key' IF cas_target = 'newvalue';\n{code}\nUsing {{cqlsh}}, if the error is reproduced, the following output will be returned:\n{code:sql}\n$ ./cqlsh <<< \"UPDATE test.test SET cas_target = 'newvalue', some_collection = {} WHERE key = 'key' IF cas_target = 'value';\"\n(start: 2016-12-22 10:14:27 EST)\n:2:ReadFailure: Error from server: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 1 failures\" info={'failures': 1, 'received_responses': 0, 'required_responses': 1, 'consistency': 'QUORUM'}\n{code}\nand the following stack trace will be present in the system log:\n{noformat}\nWARN 15:14:28 Uncaught exception on thread Thread[SharedPool-Worker-10,10,main]: {}\njava.lang.RuntimeException: java.lang.NullPointerException\n\tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2476) ~[main/:na]\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_101]\n\tat org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[main/:na]\n\tat org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [main/:na]\n\tat org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n\tat java.lang.Thread.run(Thread.java:745) [na:1.8.0_101]\nCaused by: java.lang.NullPointerException: null\n\tat org.apache.cassandra.db.rows.Row$Merger$ColumnDataReducer.getReduced(Row.java:617) ~[main/:na]\n\tat org.apache.cassandra.db.rows.Row$Merger$ColumnDataReducer.getReduced(Row.java:569) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:220) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:159) ~[main/:na]\n\tat org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n\tat org.apache.cassandra.db.rows.Row$Merger.merge(Row.java:546) ~[main/:na]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator$MergeReducer.getReduced(UnfilteredRowIterators.java:563) ~[main/:na]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator$MergeReducer.getReduced(UnfilteredRowIterators.java:527) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:220) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:159) ~[main/:na]\n\tat org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator.computeNext(UnfilteredRowIterators.java:509) ~[main/:na]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator.computeNext(UnfilteredRowIterators.java:369) ~[main/:na]\n\tat org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n\tat org.apache.cassandra.db.partitions.AbstractBTreePartition.build(AbstractBTreePartition.java:334) ~[main/:na]\n\tat org.apache.cassandra.db.partitions.ImmutableBTreePartition.create(ImmutableBTreePartition.java:111) ~[main/:na]\n\tat org.apache.cassandra.db.partitions.ImmutableBTreePartition.create(ImmutableBTreePartition.java:94) ~[main/:na]\n\tat org.apache.cassandra.db.SinglePartitionReadCommand.add(SinglePartitionReadCommand.java:810) ~[main/:na]\n\tat org.apache.cassandra.db.SinglePartitionReadCommand.queryMemtableAndSSTablesInTimestampOrder(SinglePartitionReadCommand.java:760) ~[main/:na]\n\tat org.apache.cassandra.db.SinglePartitionReadCommand.queryMemtableAndDiskInternal(SinglePartitionReadCommand.java:519) ~[main/:na]\n\tat org.apache.cassandra.db.SinglePartitionReadCommand.queryMemtableAndDisk(SinglePartitionReadCommand.java:496) ~[main/:na]\n\tat org.apache.cassandra.db.SinglePartitionReadCommand.queryStorage(SinglePartitionReadCommand.java:358) ~[main/:na]\n\tat org.apache.cassandra.db.ReadCommand.executeLocally(ReadCommand.java:394) ~[main/:na]\n\tat org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1794) ~[main/:na]\n\tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2472) ~[main/:na]\n\t... 5 common frames omitted\n{noformat}\n\nUnder both 3.0.8 and .9, the {{nodetool flush}} and additional {{UPDATE}} statement before upgrading to 3.0 are not necessary to reproduce this. In that case (when Cassandra only has to read the data from one sstable?), a different stack trace appears in the log. Here's a sample from 3.0.8:\n\n{noformat}\n WARN [SharedPool-Worker-3] 2016-12-13 15:19:48,863 AbstractLocalAwareExecutorService.java (line 169) Uncaught exception on thread Thread[SharedPool-Worker-3,5,main]: {}\njava.lang.RuntimeException: java.lang.IllegalStateException: [ColumnDefinition{name=REDACTED, type=org.apache.cassandra.db.marshal.UTF8Type, kind=REGULAR, position=-1}, ColumnDefinition{name=REDACTED2, type=org.apache.cassandra.db.marshal.MapType(org.apache.cassandra.db.marshal.UTF8Type,org.apache.cassandra.db.marshal.UTF8Type), kind=REGULAR, position=-1}] is not a subset of [REDACTED]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2453) ~[main/:na]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_101]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[main/:na]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [main/:na]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_101]\nCaused by: java.lang.IllegalStateException: [ColumnDefinition{name=REDACTED, type=org.apache.cassandra.db.marshal.UTF8Type, kind=REGULAR, position=-1}, ColumnDefinition{name=REDACTED2, type=org.apache.cassandra.db.marshal.MapType(org.apache.cassandra.db.marshal.UTF8Type,org.apache.cassandra.db.marshal.UTF8Type), kind=REGULAR, position=-1}] is not a subset of [REDACTED]\n at org.apache.cassandra.db.Columns$Serializer.encodeBitmap(Columns.java:531) ~[main/:na]\n at org.apache.cassandra.db.Columns$Serializer.serializeSubset(Columns.java:465) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:178) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:108) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:96) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:132) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:87) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:77) ~[main/:na]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:300) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:134) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:127) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:123) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:65) ~[main/:na]\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:289) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1796) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2449) ~[main/:na]\n ... 5 common frames omitted\n WARN [SharedPool-Worker-1] 2016-12-13 15:19:48,943 AbstractLocalAwareExecutorService.java (line 169) Uncaught exception on thread Thread[SharedPool-Worker-1,5,main]: {}\njava.lang.IllegalStateException: [ColumnDefinition{name=REDACTED, type=org.apache.cassandra.db.marshal.UTF8Type, kind=REGULAR, position=-1}, ColumnDefinition{name=REDACTED2, type=org.apache.cassandra.db.marshal.MapType(org.apache.cassandra.db.marshal.UTF8Type,org.apache.cassandra.db.marshal.UTF8Type), kind=REGULAR, position=-1}] is not a subset of [REDACTED]\n at org.apache.cassandra.db.Columns$Serializer.encodeBitmap(Columns.java:531) ~[main/:na]\n at org.apache.cassandra.db.Columns$Serializer.serializeSubset(Columns.java:465) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:178) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:108) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:96) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:132) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:87) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:77) ~[main/:na]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:300) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:134) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:127) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:123) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:65) ~[main/:na]\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:289) ~[main/:na]\n at org.apache.cassandra.db.ReadCommandVerbHandler.doVerb(ReadCommandVerbHandler.java:47) ~[main/:na]\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:67) ~[main/:na]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_101]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[main/:na]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [main/:na]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_101]\n{noformat}\n\nIt's not clear to us what changed in 3.0.10 to make this behavior somewhat more difficult to reproduce.\n\nWe spent some time trying to track down the cause in 3.0.8, and we've identified a very small patch (which I will attach to this issue) that _seems_ to fix it. The problem appears to be that the logic that reads data from legacy sstables can pull range tombstones covering collection columns that weren't requested, which then breaks downstream logic that doesn't expect those tombstones to be present in the data. The patch attempts to include those tombstones only if they're explicitly requested. However, there's enough going on in that logic that it's not clear to us whether the change is safe, so it is definitely in need of review from someone knowledgable about what that area of the code is intended to do.","from":"reporter","subject":"Lightweight transactions temporarily fail after upgrade from 2.1 to 3.0"},{"body":"Attaching the patch.","from":"developer"},{"body":"Thanks for the report and reproduction steps. I _was_ able to reproduce this consistently with 3.0.10 but this actually happens to be fixed on the current 3.0 branch (future 3.0.11).\n\nAnd the reason this is not a problem anymore is CASSANDRA-12694 and that's because it basically made us fetch all columns for CAS (much like we do for any other CQL query really), thus side-stepping the \"the code to read legacy sstable can return stuffs for non fetched columns\".\n\nThat said, that problem in {{LegacyLayout}} is kind of legit: we should include something that is not fetched and while we correctly skip non-fetched column for \"cells\", we did miss doing the same for collection tombsones. But I say \"kind of\" because thruth is, after CASSANDRA-12694, I'm not sure we can really run into this problem anymore: only thrift queries uses the mode where we don't fetch all columns, but thrift don't support collections so actually cannot create a query that would be problematic for this case.\n\nStill, the code is definitively not doing what it should and the fix is trivial enough that it's not worth risking running into problems in the future, or if I happen to miss a genuine case where this could happen, so I've pushed the patch for CI below and will commit if tests are clear. I did modify said patch a bit because if the column for which we have a collection tombstone is not fetched, we also want to skip the few lines that were before the condition your patch added as otherwise we might end up returning an empty row which would break assumptions of the code. Still, mostly the same thing, just moving the condition a bit up.\n| [13109-3.0|https://github.com/pcmanus/cassandra/commits/13109-3.0] | [utests|http://cassci.datastax.com/job/pcmanus-13109-3.0-testall] | [dtests|http://cassci.datastax.com/job/pcmanus-13109-3.0-dtest] |\n| [13109-3.11|https://github.com/pcmanus/cassandra/commits/13109-3.11] | [utests|http://cassci.datastax.com/job/pcmanus-13109-3.11-testall] | [dtests|http://cassci.datastax.com/job/pcmanus-13109-3.11-dtest] |\n","from":"developer"},{"body":"CI was clean so committed, thanks.","from":"developer"}],"created":"2017-01-06T19:50:41.000+0000","description":"We've observed this upgrading from 2.1.15 to 3.0.8 and from 2.1.16 to 3.0.10: some lightweight transactions executed on upgraded nodes fail with a read failure. The following conditions seem relevant to this occurring:\n\n* The transaction must be conditioned on the current value of at least one column, e.g., {{IF NOT EXISTS}} transactions don't seem to be affected.\n* There should be a collection column (in our case, a map) defined on the table on which the transaction is executed.\n* The transaction should be executed before sstables on the node are upgraded. The failure does not occur after the sstables have been upgraded (whether via {{nodetool upgradesstables}} or effectively via compaction).\n* Upgraded nodes seem to be able to participate in lightweight transactions as long as they're not the coordinator.\n* The values in the row being manipulated by the transaction must have been consistently manipulated by lightweight transactions (perhaps the existence of Paxos state for the partition is somehow relevant?).\n* In 3.0.10, it _seems_ to be necessary to have the partition split across multiple legacy sstables. This was not necessary to reproduce the bug in 3.0.8 or .9.\n\nFor applications affected by this bug, a possible workaround is to prevent nodes being upgraded from coordinating requests until sstables have been upgraded.\n\nWe're able to reproduce this when upgrading from 2.1.16 to 3.0.10 with the following steps on a single-node cluster using a mostly pristine {{cassandra.yaml}} from the source distribution.\n\n# Start Cassandra-2.1.16 on the node.\n# Create a table with a collection column and insert some data into it.\n{code:sql}\nCREATE KEYSPACE test WITH REPLICATION = {'class': 'SimpleStrategy', 'replication_factor': 1};\nCREATE TABLE test.test (key TEXT PRIMARY KEY, cas_target TEXT, some_collection MAP);\nINSERT INTO test.test (key, cas_target, some_collection) VALUES ('key', 'value', {}) IF NOT EXISTS;\n{code}\n# Flush the row to an sstable: {{nodetool flush}}.\n# Update the row:\n{code:sql}\nUPDATE test.test SET cas_target = 'newvalue', some_collection = {} WHERE key = 'key' IF cas_target = 'value';\n{code}\n# Drain the node: {{nodetool drain}}\n# Stop the node, upgrade to 3.0.10, and start the node.\n# Attempt to update the row again:\n{code:sql}\nUPDATE test.test SET cas_target = 'lastvalue' WHERE key = 'key' IF cas_target = 'newvalue';\n{code}\nUsing {{cqlsh}}, if the error is reproduced, the following output will be returned:\n{code:sql}\n$ ./cqlsh <<< \"UPDATE test.test SET cas_target = 'newvalue', some_collection = {} WHERE key = 'key' IF cas_target = 'value';\"\n(start: 2016-12-22 10:14:27 EST)\n:2:ReadFailure: Error from server: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 1 failures\" info={'failures': 1, 'received_responses': 0, 'required_responses': 1, 'consistency': 'QUORUM'}\n{code}\nand the following stack trace will be present in the system log:\n{noformat}\nWARN 15:14:28 Uncaught exception on thread Thread[SharedPool-Worker-10,10,main]: {}\njava.lang.RuntimeException: java.lang.NullPointerException\n\tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2476) ~[main/:na]\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_101]\n\tat org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[main/:na]\n\tat org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [main/:na]\n\tat org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n\tat java.lang.Thread.run(Thread.java:745) [na:1.8.0_101]\nCaused by: java.lang.NullPointerException: null\n\tat org.apache.cassandra.db.rows.Row$Merger$ColumnDataReducer.getReduced(Row.java:617) ~[main/:na]\n\tat org.apache.cassandra.db.rows.Row$Merger$ColumnDataReducer.getReduced(Row.java:569) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:220) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:159) ~[main/:na]\n\tat org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n\tat org.apache.cassandra.db.rows.Row$Merger.merge(Row.java:546) ~[main/:na]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator$MergeReducer.getReduced(UnfilteredRowIterators.java:563) ~[main/:na]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator$MergeReducer.getReduced(UnfilteredRowIterators.java:527) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:220) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:159) ~[main/:na]\n\tat org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator.computeNext(UnfilteredRowIterators.java:509) ~[main/:na]\n\tat org.apache.cassandra.db.rows.UnfilteredRowIterators$UnfilteredRowMergeIterator.computeNext(UnfilteredRowIterators.java:369) ~[main/:na]\n\tat org.apache.cassandra.utils.AbstractIterator.hasNext(AbstractIterator.java:47) ~[main/:na]\n\tat org.apache.cassandra.db.partitions.AbstractBTreePartition.build(AbstractBTreePartition.java:334) ~[main/:na]\n\tat org.apache.cassandra.db.partitions.ImmutableBTreePartition.create(ImmutableBTreePartition.java:111) ~[main/:na]\n\tat org.apache.cassandra.db.partitions.ImmutableBTreePartition.create(ImmutableBTreePartition.java:94) ~[main/:na]\n\tat org.apache.cassandra.db.SinglePartitionReadCommand.add(SinglePartitionReadCommand.java:810) ~[main/:na]\n\tat org.apache.cassandra.db.SinglePartitionReadCommand.queryMemtableAndSSTablesInTimestampOrder(SinglePartitionReadCommand.java:760) ~[main/:na]\n\tat org.apache.cassandra.db.SinglePartitionReadCommand.queryMemtableAndDiskInternal(SinglePartitionReadCommand.java:519) ~[main/:na]\n\tat org.apache.cassandra.db.SinglePartitionReadCommand.queryMemtableAndDisk(SinglePartitionReadCommand.java:496) ~[main/:na]\n\tat org.apache.cassandra.db.SinglePartitionReadCommand.queryStorage(SinglePartitionReadCommand.java:358) ~[main/:na]\n\tat org.apache.cassandra.db.ReadCommand.executeLocally(ReadCommand.java:394) ~[main/:na]\n\tat org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1794) ~[main/:na]\n\tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2472) ~[main/:na]\n\t... 5 common frames omitted\n{noformat}\n\nUnder both 3.0.8 and .9, the {{nodetool flush}} and additional {{UPDATE}} statement before upgrading to 3.0 are not necessary to reproduce this. In that case (when Cassandra only has to read the data from one sstable?), a different stack trace appears in the log. Here's a sample from 3.0.8:\n\n{noformat}\n WARN [SharedPool-Worker-3] 2016-12-13 15:19:48,863 AbstractLocalAwareExecutorService.java (line 169) Uncaught exception on thread Thread[SharedPool-Worker-3,5,main]: {}\njava.lang.RuntimeException: java.lang.IllegalStateException: [ColumnDefinition{name=REDACTED, type=org.apache.cassandra.db.marshal.UTF8Type, kind=REGULAR, position=-1}, ColumnDefinition{name=REDACTED2, type=org.apache.cassandra.db.marshal.MapType(org.apache.cassandra.db.marshal.UTF8Type,org.apache.cassandra.db.marshal.UTF8Type), kind=REGULAR, position=-1}] is not a subset of [REDACTED]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2453) ~[main/:na]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_101]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[main/:na]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [main/:na]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_101]\nCaused by: java.lang.IllegalStateException: [ColumnDefinition{name=REDACTED, type=org.apache.cassandra.db.marshal.UTF8Type, kind=REGULAR, position=-1}, ColumnDefinition{name=REDACTED2, type=org.apache.cassandra.db.marshal.MapType(org.apache.cassandra.db.marshal.UTF8Type,org.apache.cassandra.db.marshal.UTF8Type), kind=REGULAR, position=-1}] is not a subset of [REDACTED]\n at org.apache.cassandra.db.Columns$Serializer.encodeBitmap(Columns.java:531) ~[main/:na]\n at org.apache.cassandra.db.Columns$Serializer.serializeSubset(Columns.java:465) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:178) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:108) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:96) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:132) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:87) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:77) ~[main/:na]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:300) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:134) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:127) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:123) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:65) ~[main/:na]\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:289) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1796) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2449) ~[main/:na]\n ... 5 common frames omitted\n WARN [SharedPool-Worker-1] 2016-12-13 15:19:48,943 AbstractLocalAwareExecutorService.java (line 169) Uncaught exception on thread Thread[SharedPool-Worker-1,5,main]: {}\njava.lang.IllegalStateException: [ColumnDefinition{name=REDACTED, type=org.apache.cassandra.db.marshal.UTF8Type, kind=REGULAR, position=-1}, ColumnDefinition{name=REDACTED2, type=org.apache.cassandra.db.marshal.MapType(org.apache.cassandra.db.marshal.UTF8Type,org.apache.cassandra.db.marshal.UTF8Type), kind=REGULAR, position=-1}] is not a subset of [REDACTED]\n at org.apache.cassandra.db.Columns$Serializer.encodeBitmap(Columns.java:531) ~[main/:na]\n at org.apache.cassandra.db.Columns$Serializer.serializeSubset(Columns.java:465) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:178) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:108) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredSerializer.serialize(UnfilteredSerializer.java:96) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:132) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:87) ~[main/:na]\n at org.apache.cassandra.db.rows.UnfilteredRowIteratorSerializer.serialize(UnfilteredRowIteratorSerializer.java:77) ~[main/:na]\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:300) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:134) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:127) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:123) ~[main/:na]\n at org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:65) ~[main/:na]\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:289) ~[main/:na]\n at org.apache.cassandra.db.ReadCommandVerbHandler.doVerb(ReadCommandVerbHandler.java:47) ~[main/:na]\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:67) ~[main/:na]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_101]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[main/:na]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [main/:na]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_101]\n{noformat}\n\nIt's not clear to us what changed in 3.0.10 to make this behavior somewhat more difficult to reproduce.\n\nWe spent some time trying to track down the cause in 3.0.8, and we've identified a very small patch (which I will attach to this issue) that _seems_ to fix it. The problem appears to be that the logic that reads data from legacy sstables can pull range tombstones covering collection columns that weren't requested, which then breaks downstream logic that doesn't expect those tombstones to be present in the data. The patch attempts to include those tombstones only if they're explicitly requested. However, there's enough going on in that logic that it's not clear to us whether the change is safe, so it is definitely in need of review from someone knowledgable about what that area of the code is intended to do.","issue_id":"13032647","key":"CASSANDRA-13109","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-02-09T09:29:20.000+0000","role":"fixed_distractor","summary":"Lightweight transactions temporarily fail after upgrade from 2.1 to 3.0"} {"case_id":"13034648","cluster":"DISTRACTOR-CASSANDRA-13120","comments":[{"body":"I wanted to share some more details from digging and testing. First thing that I tried is LCS on clean node, to rule out TWCS as a problem. After some data loaded via *cassandra-stress* and some manually I managed to have same behavior.\n\n{code:title=getsstables|borderStyle=solid}\nvagrant@cassandra:/var/lib/cassandra/data/rts/lcs_test-2be45680e3b111e6b03e5fb4fc6cee63$ nodetool getsstables rts lcs_test 1eea2a12-e3ad-11e6-bf01-fe55135034f3;\n/var/lib/cassandra/data/rts/lcs_test-2be45680e3b111e6b03e5fb4fc6cee63/mb-993-big-Data.db\n{code}\n\nSo UUID is in single SSTable. Than I went to *cqlsh* and tried tracing and select and here is outcome:\n{code:title=tracing|borderStyle=solid}\ncqlsh> SELECT * FROM rts.lcs_test WHERE id = 1eea2a12-e3ad-11e6-bf01-fe55135034f3;\n\n id | tp_id | ex_uuid\n--------------------------------------+-------+----------\n 1eea2a12-e3ad-11e6-bf01-fe55135034f3 | 3 | asddddad\n\n(1 rows)\n\nTracing session: aad45de0-e3bc-11e6-8de0-89a4196d20ae\n\n activity | timestamp | source | source_elapsed\n-----------------------------------------------------------------------------------------------------------+----------------------------+---------------+----------------\n Execute CQL3 query | 2017-01-26 11:43:34.078000 | 192.168.34.20 | 0\n Parsing SELECT * FROM rts.lcs_test WHERE id = 1eea2a12-e3ad-11e6-bf01-fe55135034f3; [SharedPool-Worker-1] | 2017-01-26 11:43:34.078000 | 192.168.34.20 | 199\n Preparing statement [SharedPool-Worker-1] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 361\n Executing single-partition query on lcs_test [SharedPool-Worker-2] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 660\n Acquiring sstable references [SharedPool-Worker-2] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 684\n Bloom filter allows skipping sstable 1406 [SharedPool-Worker-2] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 732\n Key cache hit for sstable 993 [SharedPool-Worker-2] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 755\n Skipped 0/2 non-slice-intersecting sstables, included 0 due to tombstones [SharedPool-Worker-2] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 771\n Merging data from memtables and 2 sstables [SharedPool-Worker-2] | 2017-01-26 11:43:34.079001 | 192.168.34.20 | 785\n Read 1 live and 0 tombstone cells [SharedPool-Worker-2] | 2017-01-26 11:43:34.079001 | 192.168.34.20 | 849\n Request complete | 2017-01-26 11:43:34.082327 | 192.168.34.20 | 4327\n\n{code}\n\nSo again 2 sstables are merged while bloom filter says one can be skipped.\n\nNext think that I checked is access time of files in linux system with *ls -l --time=atime* command:\n{code:title=accesstime|borderStyle=solid}\nvagrant@cassandra:/var/lib/cassandra/data/rts/lcs_test-2be45680e3b111e6b03e5fb4fc6cee63$ ls -l --time=atime mb-993-big-*\n-rw-r--r-- 1 cassandra cassandra 76 Jan 26 11:42 mb-993-big-CRC.db\n-rw-r--r-- 1 cassandra cassandra 1048687 Jan 26 11:42 mb-993-big-Data.db\n-rw-r--r-- 1 cassandra cassandra 9 Jan 26 11:42 mb-993-big-Digest.crc32\n-rw-r--r-- 1 cassandra cassandra 5744 Jan 26 11:42 mb-993-big-Filter.db\n-rw-r--r-- 1 cassandra cassandra 100769 Jan 26 11:42 mb-993-big-Index.db\n-rw-r--r-- 1 cassandra cassandra 17703 Jan 26 11:42 mb-993-big-Statistics.db\n-rw-r--r-- 1 cassandra cassandra 1072 Jan 26 11:42 mb-993-big-Summary.db\n-rw-r--r-- 1 cassandra cassandra 80 Jan 26 11:42 mb-993-big-TOC.txt\n{code}\n\nThis timestamp of 11:42 is not changing when I do select on CQLSH and it is timestamp when Cassandra probably read file first time and placed it in memory. This actually proves that Cassandra is not accessing SSTables each time query is done, it is using memory so probably logging is misleading in this case. I would expect if bloom filter allowed skipping of some sstable to log \"Merging data from memtables and 1 sstables\" instead of 2.\n\nWhy [this line in SinglePartitionReadCommand|https://github.com/apache/cassandra/blob/554d6beb0920cba9bcc0124fa9f33c580012b761/src/java/org/apache/cassandra/db/SinglePartitionReadCommand.java#L509] returns 2 sstables when it applies PK filter instead of 1?","created":"2017-01-26T11:58:03.667+0000"},{"body":"The trace and histogram output are effectively wrong. The problem is caused by the fact that the code within {{SinglePartitionReadCommand}} is not aware of which SSTables are skipped due to the Bloom filters.\nI will try to find a way to correct that problem.","created":"2017-02-07T10:35:13.323+0000"},{"body":"Basically counter is incremented (_sstablesIterated_) on each loop no matter if Bloom Filter says it should be skipped or not. I was waiting for someone to check if that is intended or not.\n\nI have suggestion how we can fix that (I can create a patch as well). Out of all sstables candidates for a file, only sstables which return non null from [BigTableReader| https://github.com/apache/cassandra/blob/af3fe39dcabd9ef77a00309ce6741268423206df/src/java/org/apache/cassandra/io/sstable/format/big/BigTableReader.java] will add up to result. Read path and this code is executed from [this line|https://github.com/apache/cassandra/blob/cassandra-3.0.10/src/java/org/apache/cassandra/db/SinglePartitionReadCommand.java#L584].\n\nWe can introduce counter _sstablesWithData_ if _sstablesIterated_ is used somewhere else, and increment _sstablesWithData_ only when _iter_ has non-null result returned from BigTableReader.\n\nWhat do you think? That would give correct number of sstables based on read path including Bloom Filter.\n\n\n\n","created":"2017-02-07T12:02:03.611+0000"},{"body":"bq. Out of all sstables candidates for a file, only sstables which return non null from BigTableReader will add up to result.\n\nI would prefer to not use that approach as it could break quite easily without anybody noticing. If somebody was replacing the reader by an {{Optional}} for example.\nTo be honest, I am not sure yet how to fix it in the best way which is why I want to think a bit about it.","created":"2017-02-07T12:36:53.671+0000"},{"body":"I pushed a patch to solve the problem [here|https://github.com/blerer/cassandra/tree/13120-3.0]. The problem does not affect the {{3.11}} branch.\n\nThe patch passed CI without problems.\n\n[~nbozicns] I ended up relying on checking the {{RowIndexEntry}} value. I do not like this approach much but as it is properly fixed in {{3.11}} I think it is acceptable.\n\n[~Stefania] Could you review?","created":"2017-05-09T09:41:23.535+0000"},{"body":"The 3.0 patch LGTM, as you said it's not the cleanest, but it is safe and a good choice for 3.0.\n\nI am not quite sure why you think this is _properly fixed in 3.11_? {{UnfilteredRowIteratorWithLowerBound}} relies on cached RIEs only when we cannot use the metadata bounds and by then the iterator is initialized, it could return a lower bound from metadata which is before the key and in this case the iterator would be initialized. Even when the lower bound is null, merge iterator will initialize the iterator by calling {{hasNext}}. Iterator initialized means that sstables iterated is incremented ({{SPRC.withSSTablesIterated}}). Besides, {{SPRC.queryMemtableAndSSTablesInTimestampOrder}} doesn't use this iterator at all. I don't see any other check on the BF in 3.11 so is it just that it cannot be reproduced and or am I missing something?\n","created":"2017-05-10T01:06:41.835+0000"},{"body":"Sorry, my mistake. I did not checked the code correctly. I will make a new patch for 3.11","created":"2017-05-10T15:08:47.382+0000"},{"body":"Right now, what {{CFHistograms}} expose is the number of SSTables on which we do a partition lookup.\nThe partition look up can lead to skipping the SSTable (BF, min max, partition index lookup or index entry not found) or not.\nNevertheless, partition lookups are not cheap. Especially when an index lookup has to be done. Due to that, [~tjake] suggested to me, in an offline discussion, to keep the metric as it is and to add a new one {{mergedSSTable}} to track how many SSTables have been actually merged.\n\nThe number of actually merged SSTables should also be the one used for the {{Trace}} message and for determining if the {{SSTables}} must be compacted.","created":"2017-05-19T10:35:09.031+0000"},{"body":"bq. Right now, what CFHistograms expose is the number of SSTables on which we do a partition lookup.\n\nExcept it has never been documented, as far as I can see neither the [ASF docs|http://cassandra.apache.org/doc/latest/tools/nodetool/tablehistograms.html] nor the [DS docs|http://docs.datastax.com/en/cassandra/3.0/cassandra/tools/toolsTablehisto.html] say what the number of sstables actually is and [people|https://www.smartcat.io/blog/2017/where-is-my-data-debugging-sstables-in-cassandra/] tend to think this is the number of sstables it touches on each read, so it is very misleading. Our own comment says {{/** Histogram of the number of sstable data files accessed per read */}}, which to me would indicate that if the BF excludes a table then it should not be counted. We should at a minimum improve our own comments.\n\nbq. keep the metric as it is and to add a new one mergedSSTable to track how many SSTables have been actually merged.\n\nAre we thinking of a new metrics histogram? I'm not opposed, as long as we document {{nodetool \\[cf|table\\]histograms}} accordingly. My only concern is that adding a new histogram on each read may have a performance impact - but I do understand if we don't want to change the existing behavior, especially in 3.0.\n\n","created":"2017-05-22T00:59:18.121+0000"},{"body":"Just to clarify what the current metric excludes: sstables that are not live, not selected by the view interval tree (which relies on the first and last partition key in the sstable), not selected by the filters (which rely on the clustering min and max values) unless they may have tombstones, not selected because of a max timestamp older than the latest tombstone. Basically anything not selected by {{queryMemtableAndDisk}}. Also, for 3.11+, sstables that were not accessed because the query was already satisfied ({{UnfilteredRowIteratorWithLowerBound}}).\n\nHowever, sstables that are discarded by {{SSTableReader.getPosition()}}, so BF, min and max keys (again) and partition not found in the sstable (BF false positive) end up in the current metric. The partition lookup in the index is the only expensive part, the first two are not.\n\nIMO the current metric is not at all consistent, so OK for two metrics if we must, but let's clarify the boundaries in a way meaningful to the users, not dependent on how we organized the code.","created":"2017-05-22T06:09:05.565+0000"},{"body":"What did the 2.x metric include/exclude? Did we diverge with 8099, or has it always been like this?\n","created":"2017-05-22T06:18:43.836+0000"},{"body":"bq. What did the 2.x metric include/exclude? Did we diverge with 8099, or has it always been like this?\n\nFrom a very quick look, the 2.x code looks similar to 3.x, so I think it has always been like this. The entry point is {{CFS.getTopLevelColumns()}}.","created":"2017-05-22T06:32:12.999+0000"},{"body":"We double checked the {{2.2}} code with [~Stefania] and the behavior was actually changed by CASSANDRA-8099.\n\nAfter some discussion, we agreed that it is probably best to fix the metric and leave the addition of a new metric to a followup ticket (if people feel the need for a new metric).\n\nRegarding, {{SSTableReader:readCount}} as it is used by {{SizedTierd}} compaction to determine the SSTable hotness, it makes sense to also take into account parition range queries. ","created":"2017-05-22T10:01:58.853+0000"},{"body":"To clarify further, the code before 8099 did not increment the number of sstables iterated if the sstable iterator had a [null column family|https://github.com/apache/cassandra/blob/cassandra-2.2/src/java/org/apache/cassandra/db/CollationController.java#L131], which is the equivalent of {{SSTableReader.getPosition()}} returning null. Therefore, sstables excluded by the BF or a failed partition lookup were not considered. This patch will restore this behavior.","created":"2017-05-22T10:09:23.043+0000"},{"body":"I made an initial patch for 3.11 [here|https://github.com/apache/cassandra/compare/trunk...blerer:13120-3.11].\n[~Stefania] could you check the patch and tell me if you are fine with it before I rewrite it for {{3.0}} and {{trunk}}?\n","created":"2017-05-24T08:40:43.623+0000"},{"body":"The approach looks good. \n\nSome nits [here|https://github.com/stef1927/cassandra/commit/4ff0ccf5e290749b60204509b7d4b7d5d469de59]. \n\nI have two suggestions:\n\n* Pass a boolean or a new enum to the {{SSTableReadMetricsCollector}} constructor that indicates the query type (single or range). This way we know for sure if we need to increment {{mergedSSTables}}. At the moment we rely on knowing which methods get called where, which is a bit brittle.\n\n* Make the listener symmetric w.r.t. sstables skipped and selected. At the moment we have {{skippingSSTable}} with a reason to indicate that an sstable was skipped and different methods to indicate that an sstable was selected. I would personally prefer to only have two methods, something like {{onSSTableSkipped}} and {{onSSTableSelected}}, with two parameters: the sstable and the reason. The reason enums should probably be two distinct enums. We loose the RIE parameter but it is not really used at the moment. WDYT?\n\nRegarding the unit tests:\n* The javadoc of {{SSTablesIteratedTest}} needs updating, since it is no longer only limited to CASSANDRA-8180.\n* I'm not sure the new test covers {{queryMemtableAndSSTablesInTimestampOrder()}}.\n* It may be useful to also have a range query with the test checking that the metric is not updated and with a comment explaining why that it the case.\n","created":"2017-05-26T03:01:28.867+0000"},{"body":"I took the feedbacks into account and made some new patches for [3.0|https://github.com/apache/cassandra/compare/cassandra-3.0...blerer:13120-3.0], [3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...blerer:13120-3.11] and [trunk|https://github.com/apache/cassandra/compare/trunk...blerer:13120-trunk]. I ran the patch on our internal CI. There are not failures for the unit tests and the failing dtests are not related to the change. ","created":"2017-05-31T15:54:17.058+0000"},{"body":"+1, latest approach looks pretty good, great job!","created":"2017-06-01T03:36:46.111+0000"},{"body":"Thanks for the review :-)","created":"2017-06-01T08:25:01.681+0000"},{"body":"Committed into 3.0 at e22cb278b63a6ee5f03c7213071d07fd3b198659 and merged into 3.11 and trunk","created":"2017-06-01T08:26:09.324+0000"}],"conversations":[{"body":"If we look at the following output:\n\n{noformat}\n[centos@cassandra-c-3]$ nodetool getsstables -- keyspace table 60ea4399-6b9f-4419-9ccb-ff2e6742de10\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-647146-big-Data.db\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-647147-big-Data.db\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-647145-big-Data.db\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-647152-big-Data.db\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-647157-big-Data.db\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-648137-big-Data.db\n{noformat}\n\nWe can see that this key value appears in just 6 sstables. However, when we run a select against the table and key we get:\n\n{noformat}\nTracing session: a6c81330-d670-11e6-b00b-c1d403fd6e84\n\n activity | timestamp | source | source_elapsed\n-------------------------------------------------------------------------------------------------------------------+----------------------------+----------------+----------------\n Execute CQL3 query | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 0\n Parsing SELECT * FROM keyspace.table WHERE id = 60ea4399-6b9f-4419-9ccb-ff2e6742de10; [SharedPool-Worker-2] | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 104\n Preparing statement [SharedPool-Worker-2] | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 220\n Executing single-partition query on table [SharedPool-Worker-1] | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 450\n Acquiring sstable references [SharedPool-Worker-1] | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 477\n Bloom filter allows skipping sstable 648146 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 496\n Bloom filter allows skipping sstable 648145 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 503\n Key cache hit for sstable 648140 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 513\n Bloom filter allows skipping sstable 648135 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 520\n Bloom filter allows skipping sstable 648130 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 526\n Bloom filter allows skipping sstable 648048 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 530\n Bloom filter allows skipping sstable 647749 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 535\n Bloom filter allows skipping sstable 647404 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 540\n Key cache hit for sstable 647145 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 548\n Key cache hit for sstable 647146 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 556\n Key cache hit for sstable 647147 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 564\n Bloom filter allows skipping sstable 647148 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 570\n Bloom filter allows skipping sstable 647149 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 575\n Bloom filter allows skipping sstable 647150 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 580\n Bloom filter allows skipping sstable 647151 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 585\n Key cache hit for sstable 647152 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 591\n Bloom filter allows skipping sstable 647153 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 597\n Bloom filter allows skipping sstable 647154 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 601\n Bloom filter allows skipping sstable 647155 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 606\n Bloom filter allows skipping sstable 647156 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 611\n Key cache hit for sstable 647157 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 617\n Bloom filter allows skipping sstable 647158 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419003 | 10.200.254.141 | 623\n Skipped 0/22 non-slice-intersecting sstables, included 0 due to tombstones [SharedPool-Worker-1] | 2017-01-09 13:36:40.419003 | 10.200.254.141 | 644\n Merging data from memtables and 22 sstables [SharedPool-Worker-1] | 2017-01-09 13:36:40.419003 | 10.200.254.141 | 654\n Read 9 live and 0 tombstone cells [SharedPool-Worker-1] | 2017-01-09 13:36:40.419003 | 10.200.254.141 | 732\n Request complete | 2017-01-09 13:36:40.419808 | 10.200.254.141 | 808\n{noformat}\n\nYou'll note we claim not to have skipped any files due to bloom filters - even though we know the data is only in 6 files.\n\nCFHistograms also report that we're hitting every sstables:\n\n{noformat}\nPercentile SSTables Write Latency Read Latency Partition Size Cell Count \n(micros) (micros) (bytes) \n50% 24.00 14.24 182.79 103 1 \n75% 24.00 17.08 315.85 149 2 \n95% 24.00 20.50 7007.51 372 7 \n98% 24.00 24.60 10090.81 642 12 \n99% 24.00 29.52 12108.97 770 14 \nMin 21.00 3.31 29.52 43 0 \nMax 29.00 1358.10 62479.63 1597 35\n{noformat}\n\nCode for the read is here:\n\nhttps://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/SinglePartitionReadCommand.java#L561\n\nWe seem to iterate over all the sstables and increment the metric as part of that iteration.\n\nEither the reporting is incorrect - or we should maybe check the bloom filters first and then iterate the tombstones after? \n\nIn this particular case we were using TWCS which makes the problem more apparent. TWCS guarantees that we'll keep more sstables in an un-merged state. With STCS we have to search them all, but most of them should be merged together if Compaction is keeping up. LCS the read path is restricted which will disguise the impact.\n","from":"reporter","subject":"Trace and Histogram output misleading"},{"body":"I wanted to share some more details from digging and testing. First thing that I tried is LCS on clean node, to rule out TWCS as a problem. After some data loaded via *cassandra-stress* and some manually I managed to have same behavior.\n\n{code:title=getsstables|borderStyle=solid}\nvagrant@cassandra:/var/lib/cassandra/data/rts/lcs_test-2be45680e3b111e6b03e5fb4fc6cee63$ nodetool getsstables rts lcs_test 1eea2a12-e3ad-11e6-bf01-fe55135034f3;\n/var/lib/cassandra/data/rts/lcs_test-2be45680e3b111e6b03e5fb4fc6cee63/mb-993-big-Data.db\n{code}\n\nSo UUID is in single SSTable. Than I went to *cqlsh* and tried tracing and select and here is outcome:\n{code:title=tracing|borderStyle=solid}\ncqlsh> SELECT * FROM rts.lcs_test WHERE id = 1eea2a12-e3ad-11e6-bf01-fe55135034f3;\n\n id | tp_id | ex_uuid\n--------------------------------------+-------+----------\n 1eea2a12-e3ad-11e6-bf01-fe55135034f3 | 3 | asddddad\n\n(1 rows)\n\nTracing session: aad45de0-e3bc-11e6-8de0-89a4196d20ae\n\n activity | timestamp | source | source_elapsed\n-----------------------------------------------------------------------------------------------------------+----------------------------+---------------+----------------\n Execute CQL3 query | 2017-01-26 11:43:34.078000 | 192.168.34.20 | 0\n Parsing SELECT * FROM rts.lcs_test WHERE id = 1eea2a12-e3ad-11e6-bf01-fe55135034f3; [SharedPool-Worker-1] | 2017-01-26 11:43:34.078000 | 192.168.34.20 | 199\n Preparing statement [SharedPool-Worker-1] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 361\n Executing single-partition query on lcs_test [SharedPool-Worker-2] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 660\n Acquiring sstable references [SharedPool-Worker-2] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 684\n Bloom filter allows skipping sstable 1406 [SharedPool-Worker-2] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 732\n Key cache hit for sstable 993 [SharedPool-Worker-2] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 755\n Skipped 0/2 non-slice-intersecting sstables, included 0 due to tombstones [SharedPool-Worker-2] | 2017-01-26 11:43:34.079000 | 192.168.34.20 | 771\n Merging data from memtables and 2 sstables [SharedPool-Worker-2] | 2017-01-26 11:43:34.079001 | 192.168.34.20 | 785\n Read 1 live and 0 tombstone cells [SharedPool-Worker-2] | 2017-01-26 11:43:34.079001 | 192.168.34.20 | 849\n Request complete | 2017-01-26 11:43:34.082327 | 192.168.34.20 | 4327\n\n{code}\n\nSo again 2 sstables are merged while bloom filter says one can be skipped.\n\nNext think that I checked is access time of files in linux system with *ls -l --time=atime* command:\n{code:title=accesstime|borderStyle=solid}\nvagrant@cassandra:/var/lib/cassandra/data/rts/lcs_test-2be45680e3b111e6b03e5fb4fc6cee63$ ls -l --time=atime mb-993-big-*\n-rw-r--r-- 1 cassandra cassandra 76 Jan 26 11:42 mb-993-big-CRC.db\n-rw-r--r-- 1 cassandra cassandra 1048687 Jan 26 11:42 mb-993-big-Data.db\n-rw-r--r-- 1 cassandra cassandra 9 Jan 26 11:42 mb-993-big-Digest.crc32\n-rw-r--r-- 1 cassandra cassandra 5744 Jan 26 11:42 mb-993-big-Filter.db\n-rw-r--r-- 1 cassandra cassandra 100769 Jan 26 11:42 mb-993-big-Index.db\n-rw-r--r-- 1 cassandra cassandra 17703 Jan 26 11:42 mb-993-big-Statistics.db\n-rw-r--r-- 1 cassandra cassandra 1072 Jan 26 11:42 mb-993-big-Summary.db\n-rw-r--r-- 1 cassandra cassandra 80 Jan 26 11:42 mb-993-big-TOC.txt\n{code}\n\nThis timestamp of 11:42 is not changing when I do select on CQLSH and it is timestamp when Cassandra probably read file first time and placed it in memory. This actually proves that Cassandra is not accessing SSTables each time query is done, it is using memory so probably logging is misleading in this case. I would expect if bloom filter allowed skipping of some sstable to log \"Merging data from memtables and 1 sstables\" instead of 2.\n\nWhy [this line in SinglePartitionReadCommand|https://github.com/apache/cassandra/blob/554d6beb0920cba9bcc0124fa9f33c580012b761/src/java/org/apache/cassandra/db/SinglePartitionReadCommand.java#L509] returns 2 sstables when it applies PK filter instead of 1?","from":"developer"},{"body":"The trace and histogram output are effectively wrong. The problem is caused by the fact that the code within {{SinglePartitionReadCommand}} is not aware of which SSTables are skipped due to the Bloom filters.\nI will try to find a way to correct that problem.","from":"developer"},{"body":"Basically counter is incremented (_sstablesIterated_) on each loop no matter if Bloom Filter says it should be skipped or not. I was waiting for someone to check if that is intended or not.\n\nI have suggestion how we can fix that (I can create a patch as well). Out of all sstables candidates for a file, only sstables which return non null from [BigTableReader| https://github.com/apache/cassandra/blob/af3fe39dcabd9ef77a00309ce6741268423206df/src/java/org/apache/cassandra/io/sstable/format/big/BigTableReader.java] will add up to result. Read path and this code is executed from [this line|https://github.com/apache/cassandra/blob/cassandra-3.0.10/src/java/org/apache/cassandra/db/SinglePartitionReadCommand.java#L584].\n\nWe can introduce counter _sstablesWithData_ if _sstablesIterated_ is used somewhere else, and increment _sstablesWithData_ only when _iter_ has non-null result returned from BigTableReader.\n\nWhat do you think? That would give correct number of sstables based on read path including Bloom Filter.\n\n\n\n","from":"developer"},{"body":"bq. Out of all sstables candidates for a file, only sstables which return non null from BigTableReader will add up to result.\n\nI would prefer to not use that approach as it could break quite easily without anybody noticing. If somebody was replacing the reader by an {{Optional}} for example.\nTo be honest, I am not sure yet how to fix it in the best way which is why I want to think a bit about it.","from":"developer"},{"body":"I pushed a patch to solve the problem [here|https://github.com/blerer/cassandra/tree/13120-3.0]. The problem does not affect the {{3.11}} branch.\n\nThe patch passed CI without problems.\n\n[~nbozicns] I ended up relying on checking the {{RowIndexEntry}} value. I do not like this approach much but as it is properly fixed in {{3.11}} I think it is acceptable.\n\n[~Stefania] Could you review?","from":"developer"},{"body":"The 3.0 patch LGTM, as you said it's not the cleanest, but it is safe and a good choice for 3.0.\n\nI am not quite sure why you think this is _properly fixed in 3.11_? {{UnfilteredRowIteratorWithLowerBound}} relies on cached RIEs only when we cannot use the metadata bounds and by then the iterator is initialized, it could return a lower bound from metadata which is before the key and in this case the iterator would be initialized. Even when the lower bound is null, merge iterator will initialize the iterator by calling {{hasNext}}. Iterator initialized means that sstables iterated is incremented ({{SPRC.withSSTablesIterated}}). Besides, {{SPRC.queryMemtableAndSSTablesInTimestampOrder}} doesn't use this iterator at all. I don't see any other check on the BF in 3.11 so is it just that it cannot be reproduced and or am I missing something?\n","from":"developer"},{"body":"Sorry, my mistake. I did not checked the code correctly. I will make a new patch for 3.11","from":"developer"},{"body":"Right now, what {{CFHistograms}} expose is the number of SSTables on which we do a partition lookup.\nThe partition look up can lead to skipping the SSTable (BF, min max, partition index lookup or index entry not found) or not.\nNevertheless, partition lookups are not cheap. Especially when an index lookup has to be done. Due to that, [~tjake] suggested to me, in an offline discussion, to keep the metric as it is and to add a new one {{mergedSSTable}} to track how many SSTables have been actually merged.\n\nThe number of actually merged SSTables should also be the one used for the {{Trace}} message and for determining if the {{SSTables}} must be compacted.","from":"developer"},{"body":"bq. Right now, what CFHistograms expose is the number of SSTables on which we do a partition lookup.\n\nExcept it has never been documented, as far as I can see neither the [ASF docs|http://cassandra.apache.org/doc/latest/tools/nodetool/tablehistograms.html] nor the [DS docs|http://docs.datastax.com/en/cassandra/3.0/cassandra/tools/toolsTablehisto.html] say what the number of sstables actually is and [people|https://www.smartcat.io/blog/2017/where-is-my-data-debugging-sstables-in-cassandra/] tend to think this is the number of sstables it touches on each read, so it is very misleading. Our own comment says {{/** Histogram of the number of sstable data files accessed per read */}}, which to me would indicate that if the BF excludes a table then it should not be counted. We should at a minimum improve our own comments.\n\nbq. keep the metric as it is and to add a new one mergedSSTable to track how many SSTables have been actually merged.\n\nAre we thinking of a new metrics histogram? I'm not opposed, as long as we document {{nodetool \\[cf|table\\]histograms}} accordingly. My only concern is that adding a new histogram on each read may have a performance impact - but I do understand if we don't want to change the existing behavior, especially in 3.0.\n\n","from":"developer"},{"body":"Just to clarify what the current metric excludes: sstables that are not live, not selected by the view interval tree (which relies on the first and last partition key in the sstable), not selected by the filters (which rely on the clustering min and max values) unless they may have tombstones, not selected because of a max timestamp older than the latest tombstone. Basically anything not selected by {{queryMemtableAndDisk}}. Also, for 3.11+, sstables that were not accessed because the query was already satisfied ({{UnfilteredRowIteratorWithLowerBound}}).\n\nHowever, sstables that are discarded by {{SSTableReader.getPosition()}}, so BF, min and max keys (again) and partition not found in the sstable (BF false positive) end up in the current metric. The partition lookup in the index is the only expensive part, the first two are not.\n\nIMO the current metric is not at all consistent, so OK for two metrics if we must, but let's clarify the boundaries in a way meaningful to the users, not dependent on how we organized the code.","from":"developer"},{"body":"What did the 2.x metric include/exclude? Did we diverge with 8099, or has it always been like this?\n","from":"developer"},{"body":"bq. What did the 2.x metric include/exclude? Did we diverge with 8099, or has it always been like this?\n\nFrom a very quick look, the 2.x code looks similar to 3.x, so I think it has always been like this. The entry point is {{CFS.getTopLevelColumns()}}.","from":"developer"},{"body":"We double checked the {{2.2}} code with [~Stefania] and the behavior was actually changed by CASSANDRA-8099.\n\nAfter some discussion, we agreed that it is probably best to fix the metric and leave the addition of a new metric to a followup ticket (if people feel the need for a new metric).\n\nRegarding, {{SSTableReader:readCount}} as it is used by {{SizedTierd}} compaction to determine the SSTable hotness, it makes sense to also take into account parition range queries. ","from":"developer"},{"body":"To clarify further, the code before 8099 did not increment the number of sstables iterated if the sstable iterator had a [null column family|https://github.com/apache/cassandra/blob/cassandra-2.2/src/java/org/apache/cassandra/db/CollationController.java#L131], which is the equivalent of {{SSTableReader.getPosition()}} returning null. Therefore, sstables excluded by the BF or a failed partition lookup were not considered. This patch will restore this behavior.","from":"developer"},{"body":"I made an initial patch for 3.11 [here|https://github.com/apache/cassandra/compare/trunk...blerer:13120-3.11].\n[~Stefania] could you check the patch and tell me if you are fine with it before I rewrite it for {{3.0}} and {{trunk}}?\n","from":"developer"},{"body":"The approach looks good. \n\nSome nits [here|https://github.com/stef1927/cassandra/commit/4ff0ccf5e290749b60204509b7d4b7d5d469de59]. \n\nI have two suggestions:\n\n* Pass a boolean or a new enum to the {{SSTableReadMetricsCollector}} constructor that indicates the query type (single or range). This way we know for sure if we need to increment {{mergedSSTables}}. At the moment we rely on knowing which methods get called where, which is a bit brittle.\n\n* Make the listener symmetric w.r.t. sstables skipped and selected. At the moment we have {{skippingSSTable}} with a reason to indicate that an sstable was skipped and different methods to indicate that an sstable was selected. I would personally prefer to only have two methods, something like {{onSSTableSkipped}} and {{onSSTableSelected}}, with two parameters: the sstable and the reason. The reason enums should probably be two distinct enums. We loose the RIE parameter but it is not really used at the moment. WDYT?\n\nRegarding the unit tests:\n* The javadoc of {{SSTablesIteratedTest}} needs updating, since it is no longer only limited to CASSANDRA-8180.\n* I'm not sure the new test covers {{queryMemtableAndSSTablesInTimestampOrder()}}.\n* It may be useful to also have a range query with the test checking that the metric is not updated and with a comment explaining why that it the case.\n","from":"developer"},{"body":"I took the feedbacks into account and made some new patches for [3.0|https://github.com/apache/cassandra/compare/cassandra-3.0...blerer:13120-3.0], [3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...blerer:13120-3.11] and [trunk|https://github.com/apache/cassandra/compare/trunk...blerer:13120-trunk]. I ran the patch on our internal CI. There are not failures for the unit tests and the failing dtests are not related to the change. ","from":"developer"},{"body":"+1, latest approach looks pretty good, great job!","from":"developer"},{"body":"Thanks for the review :-)","from":"developer"},{"body":"Committed into 3.0 at e22cb278b63a6ee5f03c7213071d07fd3b198659 and merged into 3.11 and trunk","from":"developer"}],"created":"2017-01-13T13:41:38.000+0000","description":"If we look at the following output:\n\n{noformat}\n[centos@cassandra-c-3]$ nodetool getsstables -- keyspace table 60ea4399-6b9f-4419-9ccb-ff2e6742de10\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-647146-big-Data.db\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-647147-big-Data.db\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-647145-big-Data.db\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-647152-big-Data.db\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-647157-big-Data.db\n/mnt/cassandra/data/data/keyspace/table-62f30431acf411e69a4ed7dd11246f8a/mc-648137-big-Data.db\n{noformat}\n\nWe can see that this key value appears in just 6 sstables. However, when we run a select against the table and key we get:\n\n{noformat}\nTracing session: a6c81330-d670-11e6-b00b-c1d403fd6e84\n\n activity | timestamp | source | source_elapsed\n-------------------------------------------------------------------------------------------------------------------+----------------------------+----------------+----------------\n Execute CQL3 query | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 0\n Parsing SELECT * FROM keyspace.table WHERE id = 60ea4399-6b9f-4419-9ccb-ff2e6742de10; [SharedPool-Worker-2] | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 104\n Preparing statement [SharedPool-Worker-2] | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 220\n Executing single-partition query on table [SharedPool-Worker-1] | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 450\n Acquiring sstable references [SharedPool-Worker-1] | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 477\n Bloom filter allows skipping sstable 648146 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419000 | 10.200.254.141 | 496\n Bloom filter allows skipping sstable 648145 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 503\n Key cache hit for sstable 648140 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 513\n Bloom filter allows skipping sstable 648135 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 520\n Bloom filter allows skipping sstable 648130 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 526\n Bloom filter allows skipping sstable 648048 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 530\n Bloom filter allows skipping sstable 647749 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 535\n Bloom filter allows skipping sstable 647404 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 540\n Key cache hit for sstable 647145 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 548\n Key cache hit for sstable 647146 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419001 | 10.200.254.141 | 556\n Key cache hit for sstable 647147 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 564\n Bloom filter allows skipping sstable 647148 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 570\n Bloom filter allows skipping sstable 647149 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 575\n Bloom filter allows skipping sstable 647150 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 580\n Bloom filter allows skipping sstable 647151 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 585\n Key cache hit for sstable 647152 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 591\n Bloom filter allows skipping sstable 647153 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 597\n Bloom filter allows skipping sstable 647154 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 601\n Bloom filter allows skipping sstable 647155 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 606\n Bloom filter allows skipping sstable 647156 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 611\n Key cache hit for sstable 647157 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419002 | 10.200.254.141 | 617\n Bloom filter allows skipping sstable 647158 [SharedPool-Worker-1] | 2017-01-09 13:36:40.419003 | 10.200.254.141 | 623\n Skipped 0/22 non-slice-intersecting sstables, included 0 due to tombstones [SharedPool-Worker-1] | 2017-01-09 13:36:40.419003 | 10.200.254.141 | 644\n Merging data from memtables and 22 sstables [SharedPool-Worker-1] | 2017-01-09 13:36:40.419003 | 10.200.254.141 | 654\n Read 9 live and 0 tombstone cells [SharedPool-Worker-1] | 2017-01-09 13:36:40.419003 | 10.200.254.141 | 732\n Request complete | 2017-01-09 13:36:40.419808 | 10.200.254.141 | 808\n{noformat}\n\nYou'll note we claim not to have skipped any files due to bloom filters - even though we know the data is only in 6 files.\n\nCFHistograms also report that we're hitting every sstables:\n\n{noformat}\nPercentile SSTables Write Latency Read Latency Partition Size Cell Count \n(micros) (micros) (bytes) \n50% 24.00 14.24 182.79 103 1 \n75% 24.00 17.08 315.85 149 2 \n95% 24.00 20.50 7007.51 372 7 \n98% 24.00 24.60 10090.81 642 12 \n99% 24.00 29.52 12108.97 770 14 \nMin 21.00 3.31 29.52 43 0 \nMax 29.00 1358.10 62479.63 1597 35\n{noformat}\n\nCode for the read is here:\n\nhttps://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/SinglePartitionReadCommand.java#L561\n\nWe seem to iterate over all the sstables and increment the metric as part of that iteration.\n\nEither the reporting is incorrect - or we should maybe check the bloom filters first and then iterate the tombstones after? \n\nIn this particular case we were using TWCS which makes the problem more apparent. TWCS guarantees that we'll keep more sstables in an un-merged state. With STCS we have to search them all, but most of them should be merged together if Compaction is keeping up. LCS the read path is restricted which will disguise the impact.\n","issue_id":"13034648","key":"CASSANDRA-13120","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-06-01T08:26:09.000+0000","role":"fixed_distractor","summary":"Trace and Histogram output misleading"} {"case_id":"13035206","cluster":"DISTRACTOR-CASSANDRA-13125","comments":[{"body":"h2. Investigations...\n\nAfter some debugging, I found interesting difference in serialized RangeTombstoneLists between 2.1.16 and 3.0.10.\n\n- I ran 3 Cassandra nodes with some debug prints.\n-- 127.0.0.1 (C* 3.0.10)\n-- 127.0.0.2 (C* 2.1.16)\n-- 127.0.0.3 (C* 2.1.16)\n- They have a keyspace and a table already created.\n-- CREATE KEYSPACE test WITH replication = {'class': 'SimpleStrategy', 'replication_factor': '1'}\n-- CREATE TABLE test.test ( a int PRIMARY KEY, b int, c set, d set, e int )\n- And I query a same INSERT (which mutation is sent to 127.0.0.2) query from 127.0.0.1(C*3.0) and 127.0.0.3(C*2.1) and see the difference.\n\nInsert a row from 127.0.0.1 and scan. ( inserted (a=14) row is broken)\n{code:sql}\ncqlsh> insert into test.test(a,b,c,d,e) values(14,1,{2,3},{4,5},6);\ncqlsh> select * from test.test;\n\n a | b | c | d | e\n----+------+--------+--------+------\n 14 | 1 | null | null | null\n 14 | null | {2, 3} | {4, 5} | 6\n\n(2 rows)\n{code}\n\nAnd then, I insert from 127.0.0.3 and scan. (neither a=5 nor a=14 are broken)\n{code:sql}\ncqlsh> insert into test.test(a,b,c,d,e)values(5,1,{2,3},{4,5},6);\ncqlsh> select * from test.test;\n\n a | b | c | d | e\n----+---+--------+--------+---\n 5 | 1 | {2, 3} | {4, 5} | 6\n 14 | 1 | {2, 3} | {4, 5} | 6\n{code}\n\nAnd back to 127.0.0.1 and scan the table. a=14 is broken but a=5 is not.\n{code:sql}\ncqlsh> select * from test.test;\n\n a | b | c | d | e\n----+------+--------+--------+------\n 5 | 1 | {2, 3} | {4, 5} | 6\n 14 | 1 | null | null | null\n 14 | null | {2, 3} | {4, 5} | 6\n{code}\n\nTherefore,It looks like that \"C*3 can't scan properly rows that is stored in C*2 but inserted from C*3.\";\n\nNext, I observed some incoming MUTATIONs in 127.0.0.2 like below. I saw that C*3.0 sent RangeTombstones like {{[c-c],[c-d]}}, but C*2.1 sent {{[c:_-c],[d:_-d]}}.\n\n{noformat}\n> insert into test.test(a,b,c,d,e) values(14,1,{2,3},{4,5},6); from 127.0.0.1\n\nDeletionInfo:{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484710273390930, localDeletion=1484710273][c-d:!, deletedAt=1484710273390930, localDeletion=1484710273]}\nfrom:/127.0.0.1, payload:Mutation(keyspace='test', key='0000000e', modifications=[ColumnFamily(test -{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484710273390930, localDeletion=1484710273][c-d:!, deletedAt=1484710273390930, localDeletion=1484710273]}- [:false:0@1484710273390931,b:false:4@1484710273390931,c:00000002:false:0@1484710273390931,c:00000003:false:0@1484710273390931,d:00000004:false:0@1484710273390931,d:00000005:false:0@1484710273390931,e:false:4@1484710273390931,])]), verb:MUTATION, version:8\n\n> insert into test.test(a,b,c,d,e) values(14,1,{2,3},{4,5},6); from 127.0.0.3\nDeletionInfo:{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c:_-c:!, deletedAt=1484710277987556, localDeletion=1484710277][d:_-d:!, deletedAt=1484710277987556, localDeletion=1484710277]}\nfrom:/127.0.0.3, payload:Mutation(keyspace='test', key='0000000e', modifications=[ColumnFamily(test -{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c:_-c:!, deletedAt=1484710277987556, localDeletion=1484710277][d:_-d:!, deletedAt=1484710277987556, localDeletion=1484710277]}- [:false:0@1484710277987557,b:false:4@1484710277987557,c:00000002:false:0@1484710277987557,c:00000003:false:0@1484710277987557,d:00000004:false:0@1484710277987557,d:00000005:false:0@1484710277987557,e:false:4@1484710277987557,])]), verb:MUTATION, version:8\n{noformat}\n\nh2. Workaround Plan-A\n\nBut, LegacyRangeTombstone remove {{collectionName}} from RangeTombStone which start.bound != end.bound like {{[c-d]}}\nhttps://github.com/apache/cassandra/blob/cassandra-3.0.10/src/java/org/apache/cassandra/db/LegacyLayout.java#L1592-L1599\nIt seems like that this deletions of collectionName corrupt the unmarshal of legacy tombstone. After I commentized these else-if block, I could scan the table correctly.\n\n\n{code:java}\n if ((start.collectionName == null) != (stop.collectionName == null))\n {\n if (start.collectionName == null)\n stop = new LegacyBound(stop.bound, stop.isStatic, null);\n else\n start = new LegacyBound(start.bound, start.isStatic, null);\n }\n /*else if (!Objects.equals(start.collectionName, stop.collectionName))\n {\n // We're in the similar but slightly more complex case where on top of the big tombstone\n // A, we have 2 (or more) collection tombstones B and C within A. So we also end up with\n // a tombstone that goes between the end of B and the start of C.\n start = new LegacyBound(start.bound, start.isStatic, null);\n stop = new LegacyBound(stop.bound, stop.isStatic, null);\n }\n */\n{code}\n\n{noformat}\ncqlsh> select * from test.test;\n\n a | b | c | d | e\n----+---+--------+--------+---\n 5 | 1 | {2, 3} | {4, 5} | 6\n 14 | 1 | {2, 3} | {4, 5} | 6\n{noformat}\n\nsee patch diff-a.patch.\n\nh2. Workaround Plan-B\nInstead of modify the LegacyLayout unmarshal code, commentizing the following line fixed the problem too. It changes the TombStoneRange which is serialized by LegacyLayout from {{[c-c][c-d]}} to {{[c-c][d-d]}}.\n\nhttps://github.com/apache/cassandra/blob/cassandra-3.0.10/src/java/org/apache/cassandra/db/LegacyLayout.java#L2099\n{code:java}\n// start = ends[i];\n{code}\n\n\n{noformat}\n\nDeletionInfo:{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484715120458008, localDeletion=1484715120][d-d:!, deletedAt=1484715120458008, localDeletion=1484715120]}\nfrom:/127.0.0.1, payload:Mutation(keyspace='test', key='0000000e', modifications=[ColumnFamily(test -{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484715120458008, localDeletion=1484715120][d-d:!, deletedAt=1484715120458008, localDeletion=1484715120]}- [:false:0@1484715120458009,b:false:4@1484715120458009,c:00000002:false:0@1484715120458009,c:00000003:false:0@1484715120458009,d:00000004:false:0@1484715120458009,d:00000005:false:0@1484715120458009,e:false:4@1484715120458009,])]), verb:MUTATION, version:8\n{noformat}\n\nsee patch diff-b.patch.\n\nI'm not sure if my solution cause any unexpected effects. But I attach my patches for reference.\nCould anyboody please review my patch?\n","created":"2017-01-19T01:55:09.602+0000"},{"body":"A patch for Plan-A on Cassandra 3.0.10","created":"2017-01-19T01:55:59.207+0000"},{"body":"A patch for Plan-B on Cassandra-3.0.10.","created":"2017-01-19T01:56:26.196+0000"},{"body":"I've submitted a brief patch for reference.","created":"2017-01-19T01:57:08.772+0000"},{"body":"Thanks a lot for the very detailed reproduction case and analysis. You are absolutely right that it's abnormal for 3.0 to transform a range like {{\\[c-c!\\]\\[d-d!\\]}} into {{\\[c-c!\\]\\[c-d!\\]}}. And having duplicate row entries is only one of the problem this triggers, as with more collections you could this ending up to delete data that it shouldn't. In any case, the problem is definitively in 3.0 generating this in the first place, so that's what we should fix, so fixing this in LegacyRangeTombstone constructor (Plan-A above) is too late.\n\nSo the problem is when building the {{LegacyRangeTombstonList}} on 3.0, though it's not in the {{insertFrom}} method and the Plan-B patch would likely create other problems. The problem is in the {{LegacyBoundComparator}}: it delegates to {{clusteringComparator.compare()}} before it even checks if the bounds it compares have a {{collectionName}}, but that's incorrect.\n\nAnyway, attaching below a fix for this as well as a simple unit test showing\nwhere the problem lies:\n| [13125-3.0|https://github.com/pcmanus/cassandra/commits/13125-3.0] | [utests|http://cassci.datastax.com/job/pcmanus-13125-3.0-testall] | [dtests|http://cassci.datastax.com/job/pcmanus-13125-3.0-dtest] |\n| [13125-3.11|https://github.com/pcmanus/cassandra/commits/13125-3.11] | [utests|http://cassci.datastax.com/job/pcmanus-13125-3.11-testall] | [dtests|http://cassci.datastax.com/job/pcmanus-13125-3.11-dtest] |\n\n[~Yasuharu], if you could check if this patch does properly work in your case, that would be much appreciated. [~thobbs], since you're familiar with 3.0 compatibility code, would you also mind having a quick look at the patch?","created":"2017-01-20T14:31:13.055+0000"},{"body":"Thank you [~slebresne]! I checked that your patches worked properly in my reproduce procedure. The result is below. Now I could see that C*3.0 generated {{[c-c!][d-d!]}} style range tombstones and the rows are not broken!\n\nh4. On 13125-3.0\n{noformat}\ncqlsh> insert into test.test(a,b,c,d,e) values(14,1,{2,3},{4,5},6);\ncqlsh> select * from test.test;\n\n a | b | c | d | e\n----+---+--------+--------+---\n 14 | 1 | {2, 3} | {4, 5} | 6\n\nRangeTombstone(0), start:org.apache.cassandra.db.composites.CompoundSparseCellName@78e3b54a, end:org.apache.cassandra.db.composites.BoundedComposite@6b517b57, markedAt:1484931972134334, delTime:1484931972\nRangeTombstone(1), start:org.apache.cassandra.db.composites.CompoundSparseCellName@78e3b54b, end:org.apache.cassandra.db.composites.BoundedComposite@6b517b58, markedAt:1484931972134334, delTime:1484931972\nDeletionInfo:{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484931972134334, localDeletion=1484931972][d-d:!, deletedAt=1484931972134334, localDeletion=1484931972]}\nfrom:/127.0.0.1, payload:Mutation(keyspace='test', key='0000000e', modifications=[ColumnFamily(test -{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484931972134334, localDeletion=1484931972][d-d:!, deletedAt=1484931972134334, localDeletion=1484931972]}- [:false:0@1484931972134335,b:false:4@1484931972134335,c:00000002:false:0@1484931972134335,c:00000003:false:0@1484931972134335,d:00000004:false:0@1484931972134335,d:00000005:false:0@1484931972134335,e:false:4@1484931972134335,])]), verb:MUTATION, version:8\n{noformat}\n\nh4. On 13125-3.11\n{noformat}\ncqlsh> insert into test.test(a,b,c,d,e) values(14,1,{2,3},{4,5},6);\ncqlsh> select * from test.test; \n a | b | c | d | e\n----+---+--------+--------+---\n 14 | 1 | {2, 3} | {4, 5} | 6\n\nMutation.deserialize() size==1\nRangeTombstone(0), start:org.apache.cassandra.db.composites.CompoundSparseCellName@4316af5d, end:org.apache.cassandra.db.composites.BoundedComposite@256e93b0, markedAt:1484933162431359, delTime:1484933162\nRangeTombstone(1), start:org.apache.cassandra.db.composites.CompoundSparseCellName@4316af5e, end:org.apache.cassandra.db.composites.BoundedComposite@256e93b1, markedAt:1484933162431359, delTime:1484933162\nDeletionInfo:{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484933162431359, localDeletion=1484933162][d-d:!, deletedAt=1484933162431359, localDeletion=1484933162]}\nfrom:/127.0.0.1, payload:Mutation(keyspace='test', key='0000000e', modifications=[ColumnFamily(test -{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484933162431359, localDeletion=1484933162][d-d:!, deletedAt=1484933162431359, localDeletion=1484933162]}- [:false:0@1484933162431360,b:false:4@1484933162431360,c:00000002:false:0@1484933162431360,c:00000003:false:0@1484933162431360,d:00000004:false:0@1484933162431360,d:00000005:false:0@1484933162431360,e:false:4@1484933162431360,])]), verb:MUTATION, version:8\n{noformat}\n\n","created":"2017-01-20T17:48:18.633+0000"},{"body":"[~slebresne] the patch and new test look good to me. For some reason related to sigar your new unit test is erroring on Jenkins in the 3.11 version -- maybe [~mshuler] knows what's up with that?","created":"2017-01-20T20:51:09.515+0000"},{"body":"Would this have something to do with it?\n{noformat}\n(cassandra-3.11)mshuler@hana:~/git/cassandra$ grep -r sigar-bin conf/\nconf/cassandra-env.ps1: $env:JVM_OPTS = \"$env:JVM_OPTS -Djava.library.path=\"\"$env:CASSANDRA_HOME\\lib\\sigar-bin\"\"\"\nconf/cassandra-env.sh:JVM_OPTS=\"$JVM_OPTS -Djava.library.path=$CASSANDRA_HOME/lib/sigar-bin\"\n\n(cassandra-3.11)mshuler@hana:~/git/cassandra$ grep -r sigar-bin test/conf/\n(cassandra-3.11)mshuler@hana:~/git/cassandra$\n{noformat}","created":"2017-01-20T21:12:17.544+0000"},{"body":"bq. For some reason related to sigar your new unit test is erroring on Jenkins in the 3.11 version\n\nThe sigar stack is unrelated noise: it's not being properly set by Jenkins, but that shouldn't impact any test, and while it could be nice to fix it, let's please leave that to some other venue.\n\nThe failure seems to be due to the fact that the test don't initialize {{DatabaseDescriptor}} and so {{CFMetaData.Builder.build()}} complains about not knowing the partitioner. Pushed a trivial fix that forces a partitioner (we could also force daemon initialization I suppose but no point is making the test more heavy than it has to be). Restarted the unit test on 3.11 to check it fixes it.\n","created":"2017-01-23T10:06:16.266+0000"},{"body":"The test fix and latest test run look good, so +1 on committing.","created":"2017-01-24T17:22:57.469+0000"},{"body":"Committed, thanks.","created":"2017-02-08T10:43:19.739+0000"},{"body":"Thank you guys! :)","created":"2017-02-08T10:53:15.016+0000"},{"body":"Thank you!","created":"2017-02-09T01:03:51.286+0000"},{"body":"Thank you for the fix !\n\nJust a small question, how can I repair my table now ? I have rows with null because of that issue and I cannot delete them:\n\n(primary_id and secondary_id are my primary key)\n\n{noformat}\nselect primary_id,secondary_id, access_rights, date_modification from access WHERE primary_id = 92a9a568-e585-45cd-a1f6-96376809098c;\n\n primary_id | secondary_id | access_rights | date_modification\n--------------------------------------+--------------------------------------+-----------------+--------------------------\n 92a9a568-e585-45cd-a3f6-96376819098c | 078ae733-5be5-402e-b9b2-5f989550787e | ['b', 'u', 'r'] | 2017-02-20 12:35:02+0000\n 92a9a568-e585-45cd-a3f6-96376819098c | 078ae733-5be5-402e-b9b2-5f989550787e | null | 2017-02-20 12:35:02+0000\n{noformat}","created":"2017-02-21T16:28:26.200+0000"},{"body":"[~mnantern] In our case, {{nodetool scrub}} (C* 3.0.9 or later) fixed our broken sstables.","created":"2017-02-22T03:16:06.336+0000"}],"conversations":[{"body":"I found that rows are splitting and duplicated after upgrading the cluster from 2.1.x to 3.0.x.\n\nI found the way to reproduce the problem as below.\n{code}\n$ ccm create test -v 2.1.16 -n 3 -s \nCurrent cluster is now: test\n$ ccm node1 cqlsh -e \"CREATE KEYSPACE test WITH replication = {'class':'SimpleStrategy', 'replication_factor':3}\"\n$ ccm node1 cqlsh -e \"CREATE TABLE test.test (id text PRIMARY KEY, value1 set, value2 set);\"\n\n# Upgrade node1\n$ for i in 1; do ccm node${i} stop; ccm node${i} setdir -v3.0.10; ccm node${i} start;ccm node${i} nodetool upgradesstables; done\n\n# Insert a row through node1(3.0.10)\n$ ccm node1 cqlsh -e \"INSERT INTO test.test (id, value1, value2) values ('aaa', {'aaa', 'bbb'}, {'ccc', 'ddd'});\" \n\n# Insert a row through node2(2.1.16)\n$ ccm node2 cqlsh -e \"INSERT INTO test.test (id, value1, value2) values ('bbb', {'aaa', 'bbb'}, {'ccc', 'ddd'});\" \n\n# The row inserted from node1 is splitting\n$ ccm node1 cqlsh -e \"SELECT * FROM test.test ;\"\n\n id | value1 | value2\n-----+----------------+----------------\n aaa | null | null\n aaa | {'aaa', 'bbb'} | {'ccc', 'ddd'}\n bbb | {'aaa', 'bbb'} | {'ccc', 'ddd'}\n\n$ for i in 1 2; do ccm node${i} nodetool flush; done\n\n# Results of sstable2json of node2. The row inserted from node1(3.0.10) is different from the row inserted from node2(2.1.16).\n$ ccm node2 json -k test -c test\nrunning\n['/home/zzheng/.ccm/test/node2/data0/test/test-5406ee80dbdb11e6a175f57c4c7c85f3/test-test-ka-1-Data.db']\n-- test-test-ka-1-Data.db -----\n[\n{\"key\": \"aaa\",\n \"cells\": [[\"\",\"\",1484564624769577],\n [\"value1\",\"value2:!\",1484564624769576,\"t\",1484564624],\n [\"value1:616161\",\"\",1484564624769577],\n [\"value1:626262\",\"\",1484564624769577],\n [\"value2:636363\",\"\",1484564624769577],\n [\"value2:646464\",\"\",1484564624769577]]},\n{\"key\": \"bbb\",\n \"cells\": [[\"\",\"\",1484564634508029],\n [\"value1:_\",\"value1:!\",1484564634508028,\"t\",1484564634],\n [\"value1:616161\",\"\",1484564634508029],\n [\"value1:626262\",\"\",1484564634508029],\n [\"value2:_\",\"value2:!\",1484564634508028,\"t\",1484564634],\n [\"value2:636363\",\"\",1484564634508029],\n [\"value2:646464\",\"\",1484564634508029]]}\n]\n\n# Upgrade node2,3\n$ for i in `seq 2 3`; do ccm node${i} stop; ccm node${i} setdir -v3.0.10; ccm node${i} start;ccm node${i} nodetool upgradesstables; done\n\n# After upgrade node2,3, the row inserted from node1 is splitting in node2,3\n$ ccm node2 cqlsh -e \"SELECT * FROM test.test ;\" \n\n id | value1 | value2\n-----+----------------+----------------\n aaa | null | null\n aaa | {'aaa', 'bbb'} | {'ccc', 'ddd'}\n bbb | {'aaa', 'bbb'} | {'ccc', 'ddd'}\n\n(3 rows)\n\n# Results of sstabledump\n# node1\n[\n {\n \"partition\" : {\n \"key\" : [ \"aaa\" ],\n \"position\" : 0\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 17,\n \"liveness_info\" : { \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" },\n \"cells\" : [\n { \"name\" : \"value1\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:44.769576Z\", \"local_delete_time\" : \"2017-01-16T11:03:44Z\" } },\n { \"name\" : \"value1\", \"path\" : [ \"aaa\" ], \"value\" : \"\" },\n { \"name\" : \"value1\", \"path\" : [ \"bbb\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:44.769576Z\", \"local_delete_time\" : \"2017-01-16T11:03:44Z\" } },\n { \"name\" : \"value2\", \"path\" : [ \"ccc\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"path\" : [ \"ddd\" ], \"value\" : \"\" }\n ]\n }\n ]\n },\n {\n \"partition\" : {\n \"key\" : [ \"bbb\" ],\n \"position\" : 48\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 65,\n \"liveness_info\" : { \"tstamp\" : \"2017-01-16T11:03:54.508029Z\" },\n \"cells\" : [\n { \"name\" : \"value1\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:54.508028Z\", \"local_delete_time\" : \"2017-01-16T11:03:54Z\" } },\n { \"name\" : \"value1\", \"path\" : [ \"aaa\" ], \"value\" : \"\" },\n { \"name\" : \"value1\", \"path\" : [ \"bbb\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:54.508028Z\", \"local_delete_time\" : \"2017-01-16T11:03:54Z\" } },\n { \"name\" : \"value2\", \"path\" : [ \"ccc\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"path\" : [ \"ddd\" ], \"value\" : \"\" }\n ]\n }\n ]\n }\n] \n\n# node2\n[\n {\n \"partition\" : {\n \"key\" : [ \"aaa\" ],\n \"position\" : 0\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 17,\n \"liveness_info\" : { \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" },\n \"cells\" : [ ]\n },\n {\n \"type\" : \"row\",\n \"position\" : 22,\n \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:44.769576Z\", \"local_delete_time\" : \"2017-01-16T11:03:44Z\" },\n \"cells\" : [\n { \"name\" : \"value1\", \"path\" : [ \"aaa\" ], \"value\" : \"\", \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" },\n { \"name\" : \"value1\", \"path\" : [ \"bbb\" ], \"value\" : \"\", \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" },\n { \"name\" : \"value2\", \"path\" : [ \"ccc\" ], \"value\" : \"\", \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" },\n { \"name\" : \"value2\", \"path\" : [ \"ddd\" ], \"value\" : \"\", \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" }\n ]\n }\n ]\n },\n {\n \"partition\" : {\n \"key\" : [ \"bbb\" ],\n \"position\" : 57\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 74,\n \"liveness_info\" : { \"tstamp\" : \"2017-01-16T11:03:54.508029Z\" },\n \"cells\" : [\n { \"name\" : \"value1\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:54.508028Z\", \"local_delete_time\" : \"2017-01-16T11:03:54Z\" } },\n { \"name\" : \"value1\", \"path\" : [ \"aaa\" ], \"value\" : \"\" },\n { \"name\" : \"value1\", \"path\" : [ \"bbb\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:54.508028Z\", \"local_delete_time\" : \"2017-01-16T11:03:54Z\" } },\n { \"name\" : \"value2\", \"path\" : [ \"ccc\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"path\" : [ \"ddd\" ], \"value\" : \"\" }\n ]\n }\n ]\n }\n]\n{code}\n\nAnother example of row splitting is as follows.\n{code}\n$ ccm create test2 -v 2.1.16 -n 3 -s \nCurrent cluster is now: test2\n$ ccm node1 cqlsh -e \"CREATE KEYSPACE test WITH replication = {'class':'SimpleStrategy', 'replication_factor':3}\" \n$ ccm node1 cqlsh -e \"CREATE TABLE test.text_set_set (id text PRIMARY KEY, value1 text, value2 set, value3 set);\" \n$ for i in `seq 1`; do ccm node${i} stop; ccm node${i} setdir -v3.0.10; ccm node${i} start;ccm node${i} nodetool upgradesstables; done \n$ ccm node1 cqlsh -e \"INSERT INTO test.text_set_set (id, value1, value2, value3) values ('aaa', 'aaa', {'aaa', 'bbb'}, {'ccc', 'ddd'});\"\n$ ccm node1 cqlsh -e \"SELECT * FROM test.text_set_set;\" \n\n id | value1 | value2 | value3\n-----+--------+----------------+----------------\n aaa | aaa | null | null\n aaa | null | {'aaa', 'bbb'} | {'ccc', 'ddd'}\n\n(2 rows)\n{code}\n\nAs far as I investigated, the occurrence conditions are as follows.\n* Table schema contains multiple collections.\n* Insert a row, which values of the collection column are not null through 3.x node while both 2.1 and 3.x nodes exist in a cluster.\n* Rows in sstables of node which version was 2.1 at the time the row was inserted is splitting after upgrading to 3.x.\n\nThanks.","from":"reporter","subject":"Duplicate rows after upgrading from 2.1.16 to 3.0.10/3.9"},{"body":"h2. Investigations...\n\nAfter some debugging, I found interesting difference in serialized RangeTombstoneLists between 2.1.16 and 3.0.10.\n\n- I ran 3 Cassandra nodes with some debug prints.\n-- 127.0.0.1 (C* 3.0.10)\n-- 127.0.0.2 (C* 2.1.16)\n-- 127.0.0.3 (C* 2.1.16)\n- They have a keyspace and a table already created.\n-- CREATE KEYSPACE test WITH replication = {'class': 'SimpleStrategy', 'replication_factor': '1'}\n-- CREATE TABLE test.test ( a int PRIMARY KEY, b int, c set, d set, e int )\n- And I query a same INSERT (which mutation is sent to 127.0.0.2) query from 127.0.0.1(C*3.0) and 127.0.0.3(C*2.1) and see the difference.\n\nInsert a row from 127.0.0.1 and scan. ( inserted (a=14) row is broken)\n{code:sql}\ncqlsh> insert into test.test(a,b,c,d,e) values(14,1,{2,3},{4,5},6);\ncqlsh> select * from test.test;\n\n a | b | c | d | e\n----+------+--------+--------+------\n 14 | 1 | null | null | null\n 14 | null | {2, 3} | {4, 5} | 6\n\n(2 rows)\n{code}\n\nAnd then, I insert from 127.0.0.3 and scan. (neither a=5 nor a=14 are broken)\n{code:sql}\ncqlsh> insert into test.test(a,b,c,d,e)values(5,1,{2,3},{4,5},6);\ncqlsh> select * from test.test;\n\n a | b | c | d | e\n----+---+--------+--------+---\n 5 | 1 | {2, 3} | {4, 5} | 6\n 14 | 1 | {2, 3} | {4, 5} | 6\n{code}\n\nAnd back to 127.0.0.1 and scan the table. a=14 is broken but a=5 is not.\n{code:sql}\ncqlsh> select * from test.test;\n\n a | b | c | d | e\n----+------+--------+--------+------\n 5 | 1 | {2, 3} | {4, 5} | 6\n 14 | 1 | null | null | null\n 14 | null | {2, 3} | {4, 5} | 6\n{code}\n\nTherefore,It looks like that \"C*3 can't scan properly rows that is stored in C*2 but inserted from C*3.\";\n\nNext, I observed some incoming MUTATIONs in 127.0.0.2 like below. I saw that C*3.0 sent RangeTombstones like {{[c-c],[c-d]}}, but C*2.1 sent {{[c:_-c],[d:_-d]}}.\n\n{noformat}\n> insert into test.test(a,b,c,d,e) values(14,1,{2,3},{4,5},6); from 127.0.0.1\n\nDeletionInfo:{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484710273390930, localDeletion=1484710273][c-d:!, deletedAt=1484710273390930, localDeletion=1484710273]}\nfrom:/127.0.0.1, payload:Mutation(keyspace='test', key='0000000e', modifications=[ColumnFamily(test -{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484710273390930, localDeletion=1484710273][c-d:!, deletedAt=1484710273390930, localDeletion=1484710273]}- [:false:0@1484710273390931,b:false:4@1484710273390931,c:00000002:false:0@1484710273390931,c:00000003:false:0@1484710273390931,d:00000004:false:0@1484710273390931,d:00000005:false:0@1484710273390931,e:false:4@1484710273390931,])]), verb:MUTATION, version:8\n\n> insert into test.test(a,b,c,d,e) values(14,1,{2,3},{4,5},6); from 127.0.0.3\nDeletionInfo:{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c:_-c:!, deletedAt=1484710277987556, localDeletion=1484710277][d:_-d:!, deletedAt=1484710277987556, localDeletion=1484710277]}\nfrom:/127.0.0.3, payload:Mutation(keyspace='test', key='0000000e', modifications=[ColumnFamily(test -{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c:_-c:!, deletedAt=1484710277987556, localDeletion=1484710277][d:_-d:!, deletedAt=1484710277987556, localDeletion=1484710277]}- [:false:0@1484710277987557,b:false:4@1484710277987557,c:00000002:false:0@1484710277987557,c:00000003:false:0@1484710277987557,d:00000004:false:0@1484710277987557,d:00000005:false:0@1484710277987557,e:false:4@1484710277987557,])]), verb:MUTATION, version:8\n{noformat}\n\nh2. Workaround Plan-A\n\nBut, LegacyRangeTombstone remove {{collectionName}} from RangeTombStone which start.bound != end.bound like {{[c-d]}}\nhttps://github.com/apache/cassandra/blob/cassandra-3.0.10/src/java/org/apache/cassandra/db/LegacyLayout.java#L1592-L1599\nIt seems like that this deletions of collectionName corrupt the unmarshal of legacy tombstone. After I commentized these else-if block, I could scan the table correctly.\n\n\n{code:java}\n if ((start.collectionName == null) != (stop.collectionName == null))\n {\n if (start.collectionName == null)\n stop = new LegacyBound(stop.bound, stop.isStatic, null);\n else\n start = new LegacyBound(start.bound, start.isStatic, null);\n }\n /*else if (!Objects.equals(start.collectionName, stop.collectionName))\n {\n // We're in the similar but slightly more complex case where on top of the big tombstone\n // A, we have 2 (or more) collection tombstones B and C within A. So we also end up with\n // a tombstone that goes between the end of B and the start of C.\n start = new LegacyBound(start.bound, start.isStatic, null);\n stop = new LegacyBound(stop.bound, stop.isStatic, null);\n }\n */\n{code}\n\n{noformat}\ncqlsh> select * from test.test;\n\n a | b | c | d | e\n----+---+--------+--------+---\n 5 | 1 | {2, 3} | {4, 5} | 6\n 14 | 1 | {2, 3} | {4, 5} | 6\n{noformat}\n\nsee patch diff-a.patch.\n\nh2. Workaround Plan-B\nInstead of modify the LegacyLayout unmarshal code, commentizing the following line fixed the problem too. It changes the TombStoneRange which is serialized by LegacyLayout from {{[c-c][c-d]}} to {{[c-c][d-d]}}.\n\nhttps://github.com/apache/cassandra/blob/cassandra-3.0.10/src/java/org/apache/cassandra/db/LegacyLayout.java#L2099\n{code:java}\n// start = ends[i];\n{code}\n\n\n{noformat}\n\nDeletionInfo:{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484715120458008, localDeletion=1484715120][d-d:!, deletedAt=1484715120458008, localDeletion=1484715120]}\nfrom:/127.0.0.1, payload:Mutation(keyspace='test', key='0000000e', modifications=[ColumnFamily(test -{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484715120458008, localDeletion=1484715120][d-d:!, deletedAt=1484715120458008, localDeletion=1484715120]}- [:false:0@1484715120458009,b:false:4@1484715120458009,c:00000002:false:0@1484715120458009,c:00000003:false:0@1484715120458009,d:00000004:false:0@1484715120458009,d:00000005:false:0@1484715120458009,e:false:4@1484715120458009,])]), verb:MUTATION, version:8\n{noformat}\n\nsee patch diff-b.patch.\n\nI'm not sure if my solution cause any unexpected effects. But I attach my patches for reference.\nCould anyboody please review my patch?\n","from":"developer"},{"body":"A patch for Plan-A on Cassandra 3.0.10","from":"developer"},{"body":"A patch for Plan-B on Cassandra-3.0.10.","from":"developer"},{"body":"I've submitted a brief patch for reference.","from":"developer"},{"body":"Thanks a lot for the very detailed reproduction case and analysis. You are absolutely right that it's abnormal for 3.0 to transform a range like {{\\[c-c!\\]\\[d-d!\\]}} into {{\\[c-c!\\]\\[c-d!\\]}}. And having duplicate row entries is only one of the problem this triggers, as with more collections you could this ending up to delete data that it shouldn't. In any case, the problem is definitively in 3.0 generating this in the first place, so that's what we should fix, so fixing this in LegacyRangeTombstone constructor (Plan-A above) is too late.\n\nSo the problem is when building the {{LegacyRangeTombstonList}} on 3.0, though it's not in the {{insertFrom}} method and the Plan-B patch would likely create other problems. The problem is in the {{LegacyBoundComparator}}: it delegates to {{clusteringComparator.compare()}} before it even checks if the bounds it compares have a {{collectionName}}, but that's incorrect.\n\nAnyway, attaching below a fix for this as well as a simple unit test showing\nwhere the problem lies:\n| [13125-3.0|https://github.com/pcmanus/cassandra/commits/13125-3.0] | [utests|http://cassci.datastax.com/job/pcmanus-13125-3.0-testall] | [dtests|http://cassci.datastax.com/job/pcmanus-13125-3.0-dtest] |\n| [13125-3.11|https://github.com/pcmanus/cassandra/commits/13125-3.11] | [utests|http://cassci.datastax.com/job/pcmanus-13125-3.11-testall] | [dtests|http://cassci.datastax.com/job/pcmanus-13125-3.11-dtest] |\n\n[~Yasuharu], if you could check if this patch does properly work in your case, that would be much appreciated. [~thobbs], since you're familiar with 3.0 compatibility code, would you also mind having a quick look at the patch?","from":"developer"},{"body":"Thank you [~slebresne]! I checked that your patches worked properly in my reproduce procedure. The result is below. Now I could see that C*3.0 generated {{[c-c!][d-d!]}} style range tombstones and the rows are not broken!\n\nh4. On 13125-3.0\n{noformat}\ncqlsh> insert into test.test(a,b,c,d,e) values(14,1,{2,3},{4,5},6);\ncqlsh> select * from test.test;\n\n a | b | c | d | e\n----+---+--------+--------+---\n 14 | 1 | {2, 3} | {4, 5} | 6\n\nRangeTombstone(0), start:org.apache.cassandra.db.composites.CompoundSparseCellName@78e3b54a, end:org.apache.cassandra.db.composites.BoundedComposite@6b517b57, markedAt:1484931972134334, delTime:1484931972\nRangeTombstone(1), start:org.apache.cassandra.db.composites.CompoundSparseCellName@78e3b54b, end:org.apache.cassandra.db.composites.BoundedComposite@6b517b58, markedAt:1484931972134334, delTime:1484931972\nDeletionInfo:{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484931972134334, localDeletion=1484931972][d-d:!, deletedAt=1484931972134334, localDeletion=1484931972]}\nfrom:/127.0.0.1, payload:Mutation(keyspace='test', key='0000000e', modifications=[ColumnFamily(test -{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484931972134334, localDeletion=1484931972][d-d:!, deletedAt=1484931972134334, localDeletion=1484931972]}- [:false:0@1484931972134335,b:false:4@1484931972134335,c:00000002:false:0@1484931972134335,c:00000003:false:0@1484931972134335,d:00000004:false:0@1484931972134335,d:00000005:false:0@1484931972134335,e:false:4@1484931972134335,])]), verb:MUTATION, version:8\n{noformat}\n\nh4. On 13125-3.11\n{noformat}\ncqlsh> insert into test.test(a,b,c,d,e) values(14,1,{2,3},{4,5},6);\ncqlsh> select * from test.test; \n a | b | c | d | e\n----+---+--------+--------+---\n 14 | 1 | {2, 3} | {4, 5} | 6\n\nMutation.deserialize() size==1\nRangeTombstone(0), start:org.apache.cassandra.db.composites.CompoundSparseCellName@4316af5d, end:org.apache.cassandra.db.composites.BoundedComposite@256e93b0, markedAt:1484933162431359, delTime:1484933162\nRangeTombstone(1), start:org.apache.cassandra.db.composites.CompoundSparseCellName@4316af5e, end:org.apache.cassandra.db.composites.BoundedComposite@256e93b1, markedAt:1484933162431359, delTime:1484933162\nDeletionInfo:{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484933162431359, localDeletion=1484933162][d-d:!, deletedAt=1484933162431359, localDeletion=1484933162]}\nfrom:/127.0.0.1, payload:Mutation(keyspace='test', key='0000000e', modifications=[ColumnFamily(test -{deletedAt=-9223372036854775808, localDeletion=2147483647, ranges=[c-c:!, deletedAt=1484933162431359, localDeletion=1484933162][d-d:!, deletedAt=1484933162431359, localDeletion=1484933162]}- [:false:0@1484933162431360,b:false:4@1484933162431360,c:00000002:false:0@1484933162431360,c:00000003:false:0@1484933162431360,d:00000004:false:0@1484933162431360,d:00000005:false:0@1484933162431360,e:false:4@1484933162431360,])]), verb:MUTATION, version:8\n{noformat}\n\n","from":"developer"},{"body":"[~slebresne] the patch and new test look good to me. For some reason related to sigar your new unit test is erroring on Jenkins in the 3.11 version -- maybe [~mshuler] knows what's up with that?","from":"developer"},{"body":"Would this have something to do with it?\n{noformat}\n(cassandra-3.11)mshuler@hana:~/git/cassandra$ grep -r sigar-bin conf/\nconf/cassandra-env.ps1: $env:JVM_OPTS = \"$env:JVM_OPTS -Djava.library.path=\"\"$env:CASSANDRA_HOME\\lib\\sigar-bin\"\"\"\nconf/cassandra-env.sh:JVM_OPTS=\"$JVM_OPTS -Djava.library.path=$CASSANDRA_HOME/lib/sigar-bin\"\n\n(cassandra-3.11)mshuler@hana:~/git/cassandra$ grep -r sigar-bin test/conf/\n(cassandra-3.11)mshuler@hana:~/git/cassandra$\n{noformat}","from":"developer"},{"body":"bq. For some reason related to sigar your new unit test is erroring on Jenkins in the 3.11 version\n\nThe sigar stack is unrelated noise: it's not being properly set by Jenkins, but that shouldn't impact any test, and while it could be nice to fix it, let's please leave that to some other venue.\n\nThe failure seems to be due to the fact that the test don't initialize {{DatabaseDescriptor}} and so {{CFMetaData.Builder.build()}} complains about not knowing the partitioner. Pushed a trivial fix that forces a partitioner (we could also force daemon initialization I suppose but no point is making the test more heavy than it has to be). Restarted the unit test on 3.11 to check it fixes it.\n","from":"developer"},{"body":"The test fix and latest test run look good, so +1 on committing.","from":"developer"},{"body":"Committed, thanks.","from":"developer"},{"body":"Thank you guys! :)","from":"developer"},{"body":"Thank you!","from":"developer"},{"body":"Thank you for the fix !\n\nJust a small question, how can I repair my table now ? I have rows with null because of that issue and I cannot delete them:\n\n(primary_id and secondary_id are my primary key)\n\n{noformat}\nselect primary_id,secondary_id, access_rights, date_modification from access WHERE primary_id = 92a9a568-e585-45cd-a1f6-96376809098c;\n\n primary_id | secondary_id | access_rights | date_modification\n--------------------------------------+--------------------------------------+-----------------+--------------------------\n 92a9a568-e585-45cd-a3f6-96376819098c | 078ae733-5be5-402e-b9b2-5f989550787e | ['b', 'u', 'r'] | 2017-02-20 12:35:02+0000\n 92a9a568-e585-45cd-a3f6-96376819098c | 078ae733-5be5-402e-b9b2-5f989550787e | null | 2017-02-20 12:35:02+0000\n{noformat}","from":"developer"},{"body":"[~mnantern] In our case, {{nodetool scrub}} (C* 3.0.9 or later) fixed our broken sstables.","from":"developer"}],"created":"2017-01-16T12:29:01.000+0000","description":"I found that rows are splitting and duplicated after upgrading the cluster from 2.1.x to 3.0.x.\n\nI found the way to reproduce the problem as below.\n{code}\n$ ccm create test -v 2.1.16 -n 3 -s \nCurrent cluster is now: test\n$ ccm node1 cqlsh -e \"CREATE KEYSPACE test WITH replication = {'class':'SimpleStrategy', 'replication_factor':3}\"\n$ ccm node1 cqlsh -e \"CREATE TABLE test.test (id text PRIMARY KEY, value1 set, value2 set);\"\n\n# Upgrade node1\n$ for i in 1; do ccm node${i} stop; ccm node${i} setdir -v3.0.10; ccm node${i} start;ccm node${i} nodetool upgradesstables; done\n\n# Insert a row through node1(3.0.10)\n$ ccm node1 cqlsh -e \"INSERT INTO test.test (id, value1, value2) values ('aaa', {'aaa', 'bbb'}, {'ccc', 'ddd'});\" \n\n# Insert a row through node2(2.1.16)\n$ ccm node2 cqlsh -e \"INSERT INTO test.test (id, value1, value2) values ('bbb', {'aaa', 'bbb'}, {'ccc', 'ddd'});\" \n\n# The row inserted from node1 is splitting\n$ ccm node1 cqlsh -e \"SELECT * FROM test.test ;\"\n\n id | value1 | value2\n-----+----------------+----------------\n aaa | null | null\n aaa | {'aaa', 'bbb'} | {'ccc', 'ddd'}\n bbb | {'aaa', 'bbb'} | {'ccc', 'ddd'}\n\n$ for i in 1 2; do ccm node${i} nodetool flush; done\n\n# Results of sstable2json of node2. The row inserted from node1(3.0.10) is different from the row inserted from node2(2.1.16).\n$ ccm node2 json -k test -c test\nrunning\n['/home/zzheng/.ccm/test/node2/data0/test/test-5406ee80dbdb11e6a175f57c4c7c85f3/test-test-ka-1-Data.db']\n-- test-test-ka-1-Data.db -----\n[\n{\"key\": \"aaa\",\n \"cells\": [[\"\",\"\",1484564624769577],\n [\"value1\",\"value2:!\",1484564624769576,\"t\",1484564624],\n [\"value1:616161\",\"\",1484564624769577],\n [\"value1:626262\",\"\",1484564624769577],\n [\"value2:636363\",\"\",1484564624769577],\n [\"value2:646464\",\"\",1484564624769577]]},\n{\"key\": \"bbb\",\n \"cells\": [[\"\",\"\",1484564634508029],\n [\"value1:_\",\"value1:!\",1484564634508028,\"t\",1484564634],\n [\"value1:616161\",\"\",1484564634508029],\n [\"value1:626262\",\"\",1484564634508029],\n [\"value2:_\",\"value2:!\",1484564634508028,\"t\",1484564634],\n [\"value2:636363\",\"\",1484564634508029],\n [\"value2:646464\",\"\",1484564634508029]]}\n]\n\n# Upgrade node2,3\n$ for i in `seq 2 3`; do ccm node${i} stop; ccm node${i} setdir -v3.0.10; ccm node${i} start;ccm node${i} nodetool upgradesstables; done\n\n# After upgrade node2,3, the row inserted from node1 is splitting in node2,3\n$ ccm node2 cqlsh -e \"SELECT * FROM test.test ;\" \n\n id | value1 | value2\n-----+----------------+----------------\n aaa | null | null\n aaa | {'aaa', 'bbb'} | {'ccc', 'ddd'}\n bbb | {'aaa', 'bbb'} | {'ccc', 'ddd'}\n\n(3 rows)\n\n# Results of sstabledump\n# node1\n[\n {\n \"partition\" : {\n \"key\" : [ \"aaa\" ],\n \"position\" : 0\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 17,\n \"liveness_info\" : { \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" },\n \"cells\" : [\n { \"name\" : \"value1\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:44.769576Z\", \"local_delete_time\" : \"2017-01-16T11:03:44Z\" } },\n { \"name\" : \"value1\", \"path\" : [ \"aaa\" ], \"value\" : \"\" },\n { \"name\" : \"value1\", \"path\" : [ \"bbb\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:44.769576Z\", \"local_delete_time\" : \"2017-01-16T11:03:44Z\" } },\n { \"name\" : \"value2\", \"path\" : [ \"ccc\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"path\" : [ \"ddd\" ], \"value\" : \"\" }\n ]\n }\n ]\n },\n {\n \"partition\" : {\n \"key\" : [ \"bbb\" ],\n \"position\" : 48\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 65,\n \"liveness_info\" : { \"tstamp\" : \"2017-01-16T11:03:54.508029Z\" },\n \"cells\" : [\n { \"name\" : \"value1\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:54.508028Z\", \"local_delete_time\" : \"2017-01-16T11:03:54Z\" } },\n { \"name\" : \"value1\", \"path\" : [ \"aaa\" ], \"value\" : \"\" },\n { \"name\" : \"value1\", \"path\" : [ \"bbb\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:54.508028Z\", \"local_delete_time\" : \"2017-01-16T11:03:54Z\" } },\n { \"name\" : \"value2\", \"path\" : [ \"ccc\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"path\" : [ \"ddd\" ], \"value\" : \"\" }\n ]\n }\n ]\n }\n] \n\n# node2\n[\n {\n \"partition\" : {\n \"key\" : [ \"aaa\" ],\n \"position\" : 0\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 17,\n \"liveness_info\" : { \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" },\n \"cells\" : [ ]\n },\n {\n \"type\" : \"row\",\n \"position\" : 22,\n \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:44.769576Z\", \"local_delete_time\" : \"2017-01-16T11:03:44Z\" },\n \"cells\" : [\n { \"name\" : \"value1\", \"path\" : [ \"aaa\" ], \"value\" : \"\", \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" },\n { \"name\" : \"value1\", \"path\" : [ \"bbb\" ], \"value\" : \"\", \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" },\n { \"name\" : \"value2\", \"path\" : [ \"ccc\" ], \"value\" : \"\", \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" },\n { \"name\" : \"value2\", \"path\" : [ \"ddd\" ], \"value\" : \"\", \"tstamp\" : \"2017-01-16T11:03:44.769577Z\" }\n ]\n }\n ]\n },\n {\n \"partition\" : {\n \"key\" : [ \"bbb\" ],\n \"position\" : 57\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 74,\n \"liveness_info\" : { \"tstamp\" : \"2017-01-16T11:03:54.508029Z\" },\n \"cells\" : [\n { \"name\" : \"value1\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:54.508028Z\", \"local_delete_time\" : \"2017-01-16T11:03:54Z\" } },\n { \"name\" : \"value1\", \"path\" : [ \"aaa\" ], \"value\" : \"\" },\n { \"name\" : \"value1\", \"path\" : [ \"bbb\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"deletion_info\" : { \"marked_deleted\" : \"2017-01-16T11:03:54.508028Z\", \"local_delete_time\" : \"2017-01-16T11:03:54Z\" } },\n { \"name\" : \"value2\", \"path\" : [ \"ccc\" ], \"value\" : \"\" },\n { \"name\" : \"value2\", \"path\" : [ \"ddd\" ], \"value\" : \"\" }\n ]\n }\n ]\n }\n]\n{code}\n\nAnother example of row splitting is as follows.\n{code}\n$ ccm create test2 -v 2.1.16 -n 3 -s \nCurrent cluster is now: test2\n$ ccm node1 cqlsh -e \"CREATE KEYSPACE test WITH replication = {'class':'SimpleStrategy', 'replication_factor':3}\" \n$ ccm node1 cqlsh -e \"CREATE TABLE test.text_set_set (id text PRIMARY KEY, value1 text, value2 set, value3 set);\" \n$ for i in `seq 1`; do ccm node${i} stop; ccm node${i} setdir -v3.0.10; ccm node${i} start;ccm node${i} nodetool upgradesstables; done \n$ ccm node1 cqlsh -e \"INSERT INTO test.text_set_set (id, value1, value2, value3) values ('aaa', 'aaa', {'aaa', 'bbb'}, {'ccc', 'ddd'});\"\n$ ccm node1 cqlsh -e \"SELECT * FROM test.text_set_set;\" \n\n id | value1 | value2 | value3\n-----+--------+----------------+----------------\n aaa | aaa | null | null\n aaa | null | {'aaa', 'bbb'} | {'ccc', 'ddd'}\n\n(2 rows)\n{code}\n\nAs far as I investigated, the occurrence conditions are as follows.\n* Table schema contains multiple collections.\n* Insert a row, which values of the collection column are not null through 3.x node while both 2.1 and 3.x nodes exist in a cluster.\n* Rows in sstables of node which version was 2.1 at the time the row was inserted is splitting after upgrading to 3.x.\n\nThanks.","issue_id":"13035206","key":"CASSANDRA-13125","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-02-08T10:43:19.000+0000","role":"fixed_distractor","summary":"Duplicate rows after upgrading from 2.1.16 to 3.0.10/3.9"} {"case_id":"13035608","cluster":"DISTRACTOR-CASSANDRA-13130","comments":[{"body":"There is clearly a bug. The {{7}} should not appear.\nAt the same time you should be carefull with those type of queries. When you issue a single query with two updates, the two updates are actually having the same timestamp which means that the update with the higher value will win. As if you had send 2 {{UPDATE}} queries with exactly the same timestamp.\n\nSo, {code}UPDATE t SET listColumn[2] = 8, listColumn[2] = 7 WHERE id = 1;{code} should also return {{listColumn=(1,2,8,4)}}.","created":"2017-02-16T16:59:27.435+0000"},{"body":"{quote}When you issue a single query with two updates, the two updates are actually having the same timestamp which means that the update with the higher value will win.{quote}\nCould you please clarify what does \"the higher value will win\" mean?\nDoes it mean that it's not defined which of two updates will be actually applied?","created":"2017-02-16T19:45:39.975+0000"},{"body":"bq. Could you please clarify what does \"the higher value will win\" mean?\n\nSorry, I meant the {{greater}} value win.\n\nbq. Does it mean that it's not defined which of two updates will be actually applied?\n\nIt is clearly defined which value will win. It might not just be the one that you expect.\n\nIf you look at the following query: {{UPDATE t SET listColumn\\[2\\] = 8, listColumn\\[2\\] = 7 WHERE id = 1;}}\nYou might expect that the second value will win because it comes last.\n\nIn reality C* will consider that it has received 2 updates with exactly the same {{timestamp}}: {{UPDATE t SET listColumn\\[2\\] = 8 WHERE id = 1;}} and {{UPDATE t SET listColumn\\[2\\] = 7 WHERE id = 1;}} and will reconcile the data.\n\nThe data will be reconciled as follow:\n# if one of the modifications is a deletion (tomstone), it wins and the data is marked as deleted\n# if none of the modifications is a deletion, the update with the greater value win. If {{A > B}} then A win. If {{B > A}} then B win.\n\nYou will face the same issue with batches:\n{code}\nBEGIN BATCH\nUPDATE t SET listColumn[2] = 8 WHERE id = 1;\nUPDATE t SET listColumn[2] = 7 WHERE id = 1;\nAPPLY BATCH;\n{code}\nwill also result in {{listColumn=\\{1,2,8,4\\}}}.\n \nI you want the second update to always be the winner then you should send to separate updates.\n","created":"2017-02-17T12:08:33.597+0000"},{"body":"Thank you for the clarification!\nSuch behaviour looks like not intuitive - will be very useful to have the description somewhere in Cassandra's docs. \n\n1) Are there any plans to make the behaviour more intuitive like \"the second value will win because it comes last\"?\n2) How the following query will work (does an order matter?)?\n{code}UPDATE t SET listColumn[2] = 8, listColumn = 7 + listColumn WHERE id = 1;{code}\n3) Is there a general recommendation not to combine several updates to one column in the single query?","created":"2017-02-17T15:27:53.469+0000"},{"body":"bq. 1) Are there any plans to make the behaviour more intuitive like \"the second value will win because it comes last\"?\n\nNot for the moment. Feel free to open an improvement ticket if you want to (but make sure that you have looked at my answer to question 3))\n\nbq. How the following query will work (does an order matter?)?\n\nI had to fix the behavior for this type of queries as part of this ticket patch.\nAs long as the 2 operations do not affect the same column, the results should be the same as if they were done in 2 separate statements.\nIf they affect the same column then the data will be reconciled using the rules that I previously mentioned .\n\nFor examples you can look [here|https://github.com/blerer/cassandra/blob/13130-3.0/test/unit/org/apache/cassandra/cql3/validation/entities/CollectionsTest.java#L911]. \n\nbq. 3) Is there a general recommendation not to combine several updates to one column in the single query?\n\nOur recommendation is: *Do not do it*.\nThe output is difficult to predict and it is inefficient from the performance point of view.\nEven if the output was predictable, the update would require unecessary transfert of data between the client and the server and unecessary computing. Due to that our advise would still be: *Do not do it* :-) \n \n\n ","created":"2017-02-20T09:38:08.360+0000"},{"body":"||[2.2|https://github.com/apache/cassandra/compare/trunk...blerer:13130-2.2]|[utests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-2.2-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-2.2-dtest/]|\n||[3.0|https://github.com/apache/cassandra/compare/trunk...blerer:13130-3.0]|[utests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-3.0-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-3.0-dtest/]|\n||[3.11|https://github.com/apache/cassandra/compare/trunk...blerer:13130-3.11]|[utests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-3.11-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-3.11-dtest/]|\n||[trunk|https://github.com/apache/cassandra/compare/trunk...blerer:trunk]|[utests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-trunk-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-trunk-dtest/]|\n\nThe patches fix 2 problems:\n# For lists, the previous operations were not taken into account as the code was only looking at the prefetched list.\n# In 3.0 and after, the reconciliation of the Cells was not performed correctly for complex columns","created":"2017-02-20T09:50:18.177+0000"},{"body":"Sorry, it appears I missed that one and so it may require rebase, but +1 on the patches otherwise.","created":"2017-03-09T15:51:52.833+0000"},{"body":"Committed into 2.2 at 5ef8a8b408d4c492f7f2ffbbbe6fce237140c7cb and merged into 3.0, 3.11 and trunk","created":"2017-03-10T20:45:07.142+0000"},{"body":"Benjamin, thank you for the clarifications.\nWill try to follow your recommendations.","created":"2017-03-13T17:27:09.228+0000"}],"conversations":[{"body":"Let's assume that we have a row with the 'listColumn' column and value \\{1,2,3,4\\}.\nFor me it looks logical to expect that the following two pieces of code will ends up with the same result but it isn't so.\nCode1:\n{code}\nUPDATE t SET listColumn[2] = 7, listColumn[2] = 8 WHERE id = 1;\n{code}\nExpected result: listColumn=\\{1,2,8,4\\} \nActual result: listColumn=\\{1,2,7,8,4\\}\n\nCode2:\n{code}\nUPDATE t SET listColumn[2] = 7 WHERE id = 1;\nUPDATE t SET listColumn[2] = 8 WHERE id = 1;\n{code}\nExpected result: listColumn=\\{1,2,8,4\\} \nActual result: listColumn=\\{1,2,8,4\\}\n\nSo the question is why Code1 and Code2 give different results?\nLooks like Code1 should give the same result as Code2.","from":"reporter","subject":"Strange result of several list updates in a single request"},{"body":"There is clearly a bug. The {{7}} should not appear.\nAt the same time you should be carefull with those type of queries. When you issue a single query with two updates, the two updates are actually having the same timestamp which means that the update with the higher value will win. As if you had send 2 {{UPDATE}} queries with exactly the same timestamp.\n\nSo, {code}UPDATE t SET listColumn[2] = 8, listColumn[2] = 7 WHERE id = 1;{code} should also return {{listColumn=(1,2,8,4)}}.","from":"developer"},{"body":"{quote}When you issue a single query with two updates, the two updates are actually having the same timestamp which means that the update with the higher value will win.{quote}\nCould you please clarify what does \"the higher value will win\" mean?\nDoes it mean that it's not defined which of two updates will be actually applied?","from":"developer"},{"body":"bq. Could you please clarify what does \"the higher value will win\" mean?\n\nSorry, I meant the {{greater}} value win.\n\nbq. Does it mean that it's not defined which of two updates will be actually applied?\n\nIt is clearly defined which value will win. It might not just be the one that you expect.\n\nIf you look at the following query: {{UPDATE t SET listColumn\\[2\\] = 8, listColumn\\[2\\] = 7 WHERE id = 1;}}\nYou might expect that the second value will win because it comes last.\n\nIn reality C* will consider that it has received 2 updates with exactly the same {{timestamp}}: {{UPDATE t SET listColumn\\[2\\] = 8 WHERE id = 1;}} and {{UPDATE t SET listColumn\\[2\\] = 7 WHERE id = 1;}} and will reconcile the data.\n\nThe data will be reconciled as follow:\n# if one of the modifications is a deletion (tomstone), it wins and the data is marked as deleted\n# if none of the modifications is a deletion, the update with the greater value win. If {{A > B}} then A win. If {{B > A}} then B win.\n\nYou will face the same issue with batches:\n{code}\nBEGIN BATCH\nUPDATE t SET listColumn[2] = 8 WHERE id = 1;\nUPDATE t SET listColumn[2] = 7 WHERE id = 1;\nAPPLY BATCH;\n{code}\nwill also result in {{listColumn=\\{1,2,8,4\\}}}.\n \nI you want the second update to always be the winner then you should send to separate updates.\n","from":"developer"},{"body":"Thank you for the clarification!\nSuch behaviour looks like not intuitive - will be very useful to have the description somewhere in Cassandra's docs. \n\n1) Are there any plans to make the behaviour more intuitive like \"the second value will win because it comes last\"?\n2) How the following query will work (does an order matter?)?\n{code}UPDATE t SET listColumn[2] = 8, listColumn = 7 + listColumn WHERE id = 1;{code}\n3) Is there a general recommendation not to combine several updates to one column in the single query?","from":"developer"},{"body":"bq. 1) Are there any plans to make the behaviour more intuitive like \"the second value will win because it comes last\"?\n\nNot for the moment. Feel free to open an improvement ticket if you want to (but make sure that you have looked at my answer to question 3))\n\nbq. How the following query will work (does an order matter?)?\n\nI had to fix the behavior for this type of queries as part of this ticket patch.\nAs long as the 2 operations do not affect the same column, the results should be the same as if they were done in 2 separate statements.\nIf they affect the same column then the data will be reconciled using the rules that I previously mentioned .\n\nFor examples you can look [here|https://github.com/blerer/cassandra/blob/13130-3.0/test/unit/org/apache/cassandra/cql3/validation/entities/CollectionsTest.java#L911]. \n\nbq. 3) Is there a general recommendation not to combine several updates to one column in the single query?\n\nOur recommendation is: *Do not do it*.\nThe output is difficult to predict and it is inefficient from the performance point of view.\nEven if the output was predictable, the update would require unecessary transfert of data between the client and the server and unecessary computing. Due to that our advise would still be: *Do not do it* :-) \n \n\n ","from":"developer"},{"body":"||[2.2|https://github.com/apache/cassandra/compare/trunk...blerer:13130-2.2]|[utests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-2.2-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-2.2-dtest/]|\n||[3.0|https://github.com/apache/cassandra/compare/trunk...blerer:13130-3.0]|[utests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-3.0-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-3.0-dtest/]|\n||[3.11|https://github.com/apache/cassandra/compare/trunk...blerer:13130-3.11]|[utests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-3.11-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-3.11-dtest/]|\n||[trunk|https://github.com/apache/cassandra/compare/trunk...blerer:trunk]|[utests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-trunk-testall/]|[dtests|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13130-trunk-dtest/]|\n\nThe patches fix 2 problems:\n# For lists, the previous operations were not taken into account as the code was only looking at the prefetched list.\n# In 3.0 and after, the reconciliation of the Cells was not performed correctly for complex columns","from":"developer"},{"body":"Sorry, it appears I missed that one and so it may require rebase, but +1 on the patches otherwise.","from":"developer"},{"body":"Committed into 2.2 at 5ef8a8b408d4c492f7f2ffbbbe6fce237140c7cb and merged into 3.0, 3.11 and trunk","from":"developer"},{"body":"Benjamin, thank you for the clarifications.\nWill try to follow your recommendations.","from":"developer"}],"created":"2017-01-17T17:41:00.000+0000","description":"Let's assume that we have a row with the 'listColumn' column and value \\{1,2,3,4\\}.\nFor me it looks logical to expect that the following two pieces of code will ends up with the same result but it isn't so.\nCode1:\n{code}\nUPDATE t SET listColumn[2] = 7, listColumn[2] = 8 WHERE id = 1;\n{code}\nExpected result: listColumn=\\{1,2,8,4\\} \nActual result: listColumn=\\{1,2,7,8,4\\}\n\nCode2:\n{code}\nUPDATE t SET listColumn[2] = 7 WHERE id = 1;\nUPDATE t SET listColumn[2] = 8 WHERE id = 1;\n{code}\nExpected result: listColumn=\\{1,2,8,4\\} \nActual result: listColumn=\\{1,2,8,4\\}\n\nSo the question is why Code1 and Code2 give different results?\nLooks like Code1 should give the same result as Code2.","issue_id":"13035608","key":"CASSANDRA-13130","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-03-10T20:45:07.000+0000","role":"fixed_distractor","summary":"Strange result of several list updates in a single request"} {"case_id":"13044989","cluster":"DISTRACTOR-CASSANDRA-13246","comments":[{"body":"Attaching the file so its easier to apply the patch.","created":"2017-03-14T19:56:08.839+0000"},{"body":"I hope is easy for you to review and apply otherwise please contact me.","created":"2017-03-14T19:57:42.204+0000"},{"body":"Thanks for the patch. It looks good to me.\nI am just waiting for the CI results.","created":"2017-03-17T15:23:00.682+0000"},{"body":"[~mikkel.t.andersen@gmail.com] Be carefull, the patch should be marked as {{Ready To Commit}} only when the reviewer give his green light. Which I did not do yet as I was waiting for the CI results as mentioned in my comment.","created":"2017-03-20T08:40:15.533+0000"},{"body":"Sorry Benjamin - did not mean to cause problems... could you send me the\nlink to where the workflow is described?\n\nOn Mon, Mar 20, 2017 at 9:40 AM, Benjamin Lerer (JIRA) ","created":"2017-03-20T08:49:40.626+0000"},{"body":"Do not worry, I am just mentionning it for you to know (for your next patches ;-) )\n{{Ready to commit}} should normally be set by the reviewer to tell that the review is complete. If the reviewer is also a committer, he might simply commit the patch and skip that state.\nSo the workflow is simply: \n\n{noformat}\n +-----------------+ +----------+\n +--- ok ----> | READY TO COMMIT | --> | RESOLVED | or something like that\n+------+ +------------------+ | +-----------------+ +----------+\n| OPEN | -----> | PATCH AVAILABLE | ---+\n+------+ +------------------+ | +------+ \n +-- not ok --> | OPEN |\n +------+ \n{noformat}\n \n ","created":"2017-03-20T09:25:16.643+0000"},{"body":"CI results look good.\n|| 3.0 | [utests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-3.0-testall/] | [dtests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-3.0-dtest/]|\n|| 3.11 | [utests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-3.11-testall/] | [dtests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-3.11-dtest/]|\n|| trunk | [utests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-trunk-testall/] | [dtests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-trunk-dtest/]|\n\nI just added an extra unit test to test filtering without secondary index.","created":"2017-03-20T09:50:07.892+0000"},{"body":"Committed into 3.0 at 0eebc6e6b7cd7fc801579e57701608e7bf155ee0 and merged into 3.11 and trunk.","created":"2017-03-20T11:43:37.360+0000"}],"conversations":[{"body":"Not sure if this is the absolute minimal case that produces the bug, but here are the steps for reproducing.\n\n1. Create table\n{code}\nCREATE TABLE test (\nid text,\nck1 text,\nck2 text,\nstatic_value text static,\nset_value set,\nprimary key (id, ck1, ck2)\n);\n{code}\n2. Create secondary indices on the clustering columns, static column, and collection column\n{code}\ncreate index on test (set_value);\ncreate index on test (static_value);\ncreate index on test (ck1);\ncreate index on test (ck2);\n{code}\n3. Insert a null value into the `set_value` column\n{code}\ninsert into test (id, ck1, ck2, static_value, set_value) values ('id', 'key1', 'key2', 'static', {'one', 'two'} );\n{code}\nSanity check: \n{code}\nselect * from test;\n\n id | ck1 | ck2 | static_value | set_value\n----+------+------+--------------+----------------\n id | key1 | key2 | static | {'one', 'two'}\n{code}\n4. Set the set_value to be empty\n{code}\nupdate test set set_value = {} where id = 'id' and ck1 = 'key1' and ck2 = 'key2';\n{code}\n5. Make a select query that uses `CONTAINS` in the `set_value` column\n{code}\nselect * from test where ck2 = 'key2' and static_value = 'static' and set_value contains 'one' allow filtering;\n{code}\nHere we get a ReadFailure:\n{code}\nReadFailure: Error from server: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 1 failures\" info={'failures': 1, 'received_responses': 0, 'required_responses': 1, 'consistency': 'ONE'}\n{code}\nLogs show a NullPointerException\n{code}\njava.lang.RuntimeException: java.lang.NullPointerException\n \tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2470) ~[apache-cassandra-3.7.jar:3.7]\n \tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_101]\n \tat org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [apache-cassandra-3.7.jar:3.7]\n \tat java.lang.Thread.run(Thread.java:745) [na:1.8.0_101]\nCaused by: java.lang.NullPointerException: null\n \tat org.apache.cassandra.db.filter.RowFilter$SimpleExpression.isSatisfiedBy(RowFilter.java:720) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.filter.RowFilter$CQLFilter$1IsSatisfiedFilter.applyToRow(RowFilter.java:303) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.transform.BaseRows.hasNext(BaseRows.java:120) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.filter.RowFilter$CQLFilter$1IsSatisfiedFilter.applyToPartition(RowFilter.java:293) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.filter.RowFilter$CQLFilter$1IsSatisfiedFilter.applyToPartition(RowFilter.java:281) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.transform.BasePartitions.hasNext(BasePartitions.java:76) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:289) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:134) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:127) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:123) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:65) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:292) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1799) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2466) ~[apache-cassandra-3.7.jar:3.7]\n \t... 5 common frames omitted\n{code}","from":"reporter","subject":"Querying by secondary index on collection column returns NullPointerException sometimes"},{"body":"Attaching the file so its easier to apply the patch.","from":"developer"},{"body":"I hope is easy for you to review and apply otherwise please contact me.","from":"developer"},{"body":"Thanks for the patch. It looks good to me.\nI am just waiting for the CI results.","from":"developer"},{"body":"[~mikkel.t.andersen@gmail.com] Be carefull, the patch should be marked as {{Ready To Commit}} only when the reviewer give his green light. Which I did not do yet as I was waiting for the CI results as mentioned in my comment.","from":"developer"},{"body":"Sorry Benjamin - did not mean to cause problems... could you send me the\nlink to where the workflow is described?\n\nOn Mon, Mar 20, 2017 at 9:40 AM, Benjamin Lerer (JIRA) ","from":"developer"},{"body":"Do not worry, I am just mentionning it for you to know (for your next patches ;-) )\n{{Ready to commit}} should normally be set by the reviewer to tell that the review is complete. If the reviewer is also a committer, he might simply commit the patch and skip that state.\nSo the workflow is simply: \n\n{noformat}\n +-----------------+ +----------+\n +--- ok ----> | READY TO COMMIT | --> | RESOLVED | or something like that\n+------+ +------------------+ | +-----------------+ +----------+\n| OPEN | -----> | PATCH AVAILABLE | ---+\n+------+ +------------------+ | +------+ \n +-- not ok --> | OPEN |\n +------+ \n{noformat}\n \n ","from":"developer"},{"body":"CI results look good.\n|| 3.0 | [utests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-3.0-testall/] | [dtests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-3.0-dtest/]|\n|| 3.11 | [utests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-3.11-testall/] | [dtests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-3.11-dtest/]|\n|| trunk | [utests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-trunk-testall/] | [dtests| http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-13246-trunk-dtest/]|\n\nI just added an extra unit test to test filtering without secondary index.","from":"developer"},{"body":"Committed into 3.0 at 0eebc6e6b7cd7fc801579e57701608e7bf155ee0 and merged into 3.11 and trunk.","from":"developer"}],"created":"2017-02-21T22:24:37.000+0000","description":"Not sure if this is the absolute minimal case that produces the bug, but here are the steps for reproducing.\n\n1. Create table\n{code}\nCREATE TABLE test (\nid text,\nck1 text,\nck2 text,\nstatic_value text static,\nset_value set,\nprimary key (id, ck1, ck2)\n);\n{code}\n2. Create secondary indices on the clustering columns, static column, and collection column\n{code}\ncreate index on test (set_value);\ncreate index on test (static_value);\ncreate index on test (ck1);\ncreate index on test (ck2);\n{code}\n3. Insert a null value into the `set_value` column\n{code}\ninsert into test (id, ck1, ck2, static_value, set_value) values ('id', 'key1', 'key2', 'static', {'one', 'two'} );\n{code}\nSanity check: \n{code}\nselect * from test;\n\n id | ck1 | ck2 | static_value | set_value\n----+------+------+--------------+----------------\n id | key1 | key2 | static | {'one', 'two'}\n{code}\n4. Set the set_value to be empty\n{code}\nupdate test set set_value = {} where id = 'id' and ck1 = 'key1' and ck2 = 'key2';\n{code}\n5. Make a select query that uses `CONTAINS` in the `set_value` column\n{code}\nselect * from test where ck2 = 'key2' and static_value = 'static' and set_value contains 'one' allow filtering;\n{code}\nHere we get a ReadFailure:\n{code}\nReadFailure: Error from server: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 1 failures\" info={'failures': 1, 'received_responses': 0, 'required_responses': 1, 'consistency': 'ONE'}\n{code}\nLogs show a NullPointerException\n{code}\njava.lang.RuntimeException: java.lang.NullPointerException\n \tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2470) ~[apache-cassandra-3.7.jar:3.7]\n \tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_101]\n \tat org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [apache-cassandra-3.7.jar:3.7]\n \tat java.lang.Thread.run(Thread.java:745) [na:1.8.0_101]\nCaused by: java.lang.NullPointerException: null\n \tat org.apache.cassandra.db.filter.RowFilter$SimpleExpression.isSatisfiedBy(RowFilter.java:720) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.filter.RowFilter$CQLFilter$1IsSatisfiedFilter.applyToRow(RowFilter.java:303) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.transform.BaseRows.hasNext(BaseRows.java:120) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.filter.RowFilter$CQLFilter$1IsSatisfiedFilter.applyToPartition(RowFilter.java:293) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.filter.RowFilter$CQLFilter$1IsSatisfiedFilter.applyToPartition(RowFilter.java:281) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.transform.BasePartitions.hasNext(BasePartitions.java:76) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.partitions.UnfilteredPartitionIterators$Serializer.serialize(UnfilteredPartitionIterators.java:289) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.ReadResponse$LocalDataResponse.build(ReadResponse.java:134) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:127) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.ReadResponse$LocalDataResponse.(ReadResponse.java:123) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.ReadResponse.createDataResponse(ReadResponse.java:65) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:292) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1799) ~[apache-cassandra-3.7.jar:3.7]\n \tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2466) ~[apache-cassandra-3.7.jar:3.7]\n \t... 5 common frames omitted\n{code}","issue_id":"13044989","key":"CASSANDRA-13246","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-03-20T11:43:37.000+0000","role":"fixed_distractor","summary":"Querying by secondary index on collection column returns NullPointerException sometimes"} {"case_id":"13067657","cluster":"DISTRACTOR-CASSANDRA-13482","comments":[{"body":"I have same error on cassandra 3.10:\n\nquery fail with row_cache on, and fine with cache off.\nconsistency level doesn't matter - one/quorum/all -> same result.\nswitch cache off and when on and/or restart cassandra doesn't fix problem.\n\n\njava.lang.AssertionError: null\n at org.apache.cassandra.db.rows.UnfilteredRowIterators.concat(UnfilteredRowIterators.java:210) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.db.SinglePartitionReadCommand.getThroughCache(SinglePartitionReadCommand.java:474) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.db.SinglePartitionReadCommand.queryStorage(SinglePartitionReadCommand.java:374) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.db.ReadCommand.executeLocally(ReadCommand.java:407) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.db.ReadCommandVerbHandler.doVerb(ReadCommandVerbHandler.java:48) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:66) ~[apache-cassandra-3.10.jar:3.10]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_131]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:162) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:134) [apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:109) [apache-cassandra-3.10.jar:3.10]\n at java.lang.Thread.run(Thread.java:748) [na:1.8.0_131]\n\n\n---\n\ncassandra@cqlsh:sb> select * from session where agent_id = 846bed6c-978c-4cdb-958a-6a6155e9cdb5;\nReadFailure: Error from server: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 2 failures\" info={'failures': 2, 'received_responses': 0, 'required_responses': 2, 'consistency': 'QUORUM'}\n\ncassandra@cqlsh:sb> ALTER TABLE session WITH caching = {'keys': 'ALL'};\ncassandra@cqlsh:sb> select * from session where agent_id = 846bed6c-978c-4cdb-958a-6a6155e9cdb5;\n\n agent_id | session_id | all_isps | all_locations | all_logins | last_ip | last_login | last_page | last_time | legacy_id\n----------+------------+----------+---------------+------------+---------+------------+-----------+-----------+-----------\n\n(0 rows)\n\ncassandra@cqlsh:sb> ALTER TABLE session WITH compaction = {'class': 'LeveledCompactionStrategy'} and caching = {'keys': 'ALL','rows_per_partition': 1};\ncassandra@cqlsh:sb> select * from session where agent_id = 846bed6c-978c-4cdb-958a-6a6155e9cdb5;\nReadFailure: Error from server: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 2 failures\" info={'failures': 2, 'received_responses': 0, 'required_responses': 2, 'consistency': 'QUORUM'}\n\n---\n\nif i do select with another agent_id - all work fine! only one key fail...","created":"2017-05-18T16:00:30.418+0000"},{"body":"I've composed a patch to mitigate the problem. \n\nIn order to fix it, we have to allow for concatenating the iterators with different amounts of columns, although make sure that cases with wrapped iterators, limits, stopping and empty iterators all have some predictable behaviour. \n\nWhile working on this problem I have also discovered the slight inconsistency in the way concatenation and {{MoreRows}} is working right now: {{DataLimits}} filter [here|https://github.com/apache/cassandra/blob/a87b15d1d6c42f4247c84b460ed39899d8813a6f/src/java/org/apache/cassandra/db/SinglePartitionReadCommand.java#L423] will call {{stopInPartition}}, which would effectively \"stop\" the original {{iter}} iterator, and make it return {{hasNext() => false}}, even though only one row was consumed. \n\nRight now, it works only because internally {{concat}} would take {{input}} from the {{BaseIterator}} and discard this {{isStopped}} in [tryGetMoreContent|https://github.com/apache/cassandra/blob/81f6c784ce967fadb6ed7f58de1328e713eaf53c/src/java/org/apache/cassandra/db/transform/BaseIterator.java#L124]. In the context of this patch this would mean that we'd get only cached results and avoid reading mem/sstable. \n\nIn other words, current behaviour can be described as:\n\n{code}\n iter1 = /* some iterator yielding: 1, 2, 3 */;\n iter2 = /* some iterator yielding: 3, 4, 5, 6, 7, 8, 9 */;\n concatenated = concat(iter1, DataLimits.cqlLimits(3).filter(iter2));\n{code}\n\nwould result in {{concatenated}} yielding 1 through 9, which is incorrect.\n\nThe patch implements the following changes:\n\n * {{concat}} [now allows|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:13482-trunk#diff-57d0dfa95504bfd17d30539b3b338c0cL205] different amount of columns from iterators, but returned columns will be a union of the two iterators\n * {{BaseIterator}} [would now|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:13482-trunk#diff-d14a4b314544d3720010343e330e7e3cR125] take the {{stop}} from the child iterator, which means that the iterator was stopped, even if it might have contents, it will not yield additional data. Unfortunately, since the iterator itself had a stopping transformation, and it's own {{stop}} reference is \"leaked\" during the creation of the transformation, we have to save (and check) both the original stop and the child one. Here, the naming might be a bit off.\n * {{cacheIterator}} [is now|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:13482-trunk#diff-2e17efa5977a71330df6651d3bec0d12R424] using a custom wrapped iterator that will not call {{stop}} on the wrapping iterator, but previous problem with {{Unfiltered}} is now fixed\n\n|[3.0|https://github.com/apache/cassandra/compare/cassandra-3.0...ifesdjeen:13482-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-3.0-testall/]|[dtest|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-3.0-dtest/]|\n|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...ifesdjeen:13482-3.11]|[testall|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-3.11-testall/]|[dtest|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-3.11-dtest/]|\n|[trunk|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:13482-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-trunk-testall/]|[dtest|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-trunk-dtest/]|","created":"2017-06-09T09:23:31.896+0000"},{"body":"lgtm, +1.","created":"2017-07-12T07:39:56.008+0000"},{"body":"Thank you for the review,\n\nCommitted to 3.0 with [7251c9559805d83423ca5ddbe4f955ce668c3d9a|https://github.com/apache/cassandra/commit/7251c9559805d83423ca5ddbe4f955ce668c3d9a] and merged up to [3.11|https://github.com/apache/cassandra/commit/29db2511621e420b8d64c867a16e317589397d36] and [trunk|https://github.com/apache/cassandra/commit/f48a319ac884ef8d6eb54db3176ea2acf627bb89].","created":"2017-07-12T13:55:44.702+0000"}],"conversations":[{"body":"The problem is reproducible on 3.0 with:\n\n{code}\n-# row_cache_class_name: org.apache.cassandra.cache.OHCProvider\n+row_cache_class_name: org.apache.cassandra.cache.OHCProvider\n\n-row_cache_size_in_mb: 0\n+row_cache_size_in_mb: 100\n{code}\n\nTable setup:\n\n{code}\nCREATE TABLE cache_tables (pk int, v1 int, v2 int, v3 int, primary key (pk, v1)) WITH CACHING = { 'keys': 'ALL', 'rows_per_partition': '1' } ;\n{code}\n\nNo data is required, only a head query (or any pk/ck query but with full partitions cached). \n\n{code}\nselect * from cross_page_queries where pk = 10000 ;\n{code}\n\n{code}\njava.lang.AssertionError: null\n at org.apache.cassandra.db.rows.UnfilteredRowIterators.concat(UnfilteredRowIterators.java:193) ~[main/:na]\n at org.apache.cassandra.db.SinglePartitionReadCommand.getThroughCache(SinglePartitionReadCommand.java:461) ~[main/:na]\n at org.apache.cassandra.db.SinglePartitionReadCommand.queryStorage(SinglePartitionReadCommand.java:358) ~[main/:na]\n at org.apache.cassandra.db.ReadCommand.executeLocally(ReadCommand.java:395) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1794) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2472) ~[main/:na]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_121]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[main/:na]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [main/:na]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_121]\n{code}","from":"reporter","subject":"NPE on non-existing row read when row cache is enabled"},{"body":"I have same error on cassandra 3.10:\n\nquery fail with row_cache on, and fine with cache off.\nconsistency level doesn't matter - one/quorum/all -> same result.\nswitch cache off and when on and/or restart cassandra doesn't fix problem.\n\n\njava.lang.AssertionError: null\n at org.apache.cassandra.db.rows.UnfilteredRowIterators.concat(UnfilteredRowIterators.java:210) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.db.SinglePartitionReadCommand.getThroughCache(SinglePartitionReadCommand.java:474) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.db.SinglePartitionReadCommand.queryStorage(SinglePartitionReadCommand.java:374) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.db.ReadCommand.executeLocally(ReadCommand.java:407) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.db.ReadCommandVerbHandler.doVerb(ReadCommandVerbHandler.java:48) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:66) ~[apache-cassandra-3.10.jar:3.10]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_131]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:162) ~[apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:134) [apache-cassandra-3.10.jar:3.10]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:109) [apache-cassandra-3.10.jar:3.10]\n at java.lang.Thread.run(Thread.java:748) [na:1.8.0_131]\n\n\n---\n\ncassandra@cqlsh:sb> select * from session where agent_id = 846bed6c-978c-4cdb-958a-6a6155e9cdb5;\nReadFailure: Error from server: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 2 failures\" info={'failures': 2, 'received_responses': 0, 'required_responses': 2, 'consistency': 'QUORUM'}\n\ncassandra@cqlsh:sb> ALTER TABLE session WITH caching = {'keys': 'ALL'};\ncassandra@cqlsh:sb> select * from session where agent_id = 846bed6c-978c-4cdb-958a-6a6155e9cdb5;\n\n agent_id | session_id | all_isps | all_locations | all_logins | last_ip | last_login | last_page | last_time | legacy_id\n----------+------------+----------+---------------+------------+---------+------------+-----------+-----------+-----------\n\n(0 rows)\n\ncassandra@cqlsh:sb> ALTER TABLE session WITH compaction = {'class': 'LeveledCompactionStrategy'} and caching = {'keys': 'ALL','rows_per_partition': 1};\ncassandra@cqlsh:sb> select * from session where agent_id = 846bed6c-978c-4cdb-958a-6a6155e9cdb5;\nReadFailure: Error from server: code=1300 [Replica(s) failed to execute read] message=\"Operation failed - received 0 responses and 2 failures\" info={'failures': 2, 'received_responses': 0, 'required_responses': 2, 'consistency': 'QUORUM'}\n\n---\n\nif i do select with another agent_id - all work fine! only one key fail...","from":"developer"},{"body":"I've composed a patch to mitigate the problem. \n\nIn order to fix it, we have to allow for concatenating the iterators with different amounts of columns, although make sure that cases with wrapped iterators, limits, stopping and empty iterators all have some predictable behaviour. \n\nWhile working on this problem I have also discovered the slight inconsistency in the way concatenation and {{MoreRows}} is working right now: {{DataLimits}} filter [here|https://github.com/apache/cassandra/blob/a87b15d1d6c42f4247c84b460ed39899d8813a6f/src/java/org/apache/cassandra/db/SinglePartitionReadCommand.java#L423] will call {{stopInPartition}}, which would effectively \"stop\" the original {{iter}} iterator, and make it return {{hasNext() => false}}, even though only one row was consumed. \n\nRight now, it works only because internally {{concat}} would take {{input}} from the {{BaseIterator}} and discard this {{isStopped}} in [tryGetMoreContent|https://github.com/apache/cassandra/blob/81f6c784ce967fadb6ed7f58de1328e713eaf53c/src/java/org/apache/cassandra/db/transform/BaseIterator.java#L124]. In the context of this patch this would mean that we'd get only cached results and avoid reading mem/sstable. \n\nIn other words, current behaviour can be described as:\n\n{code}\n iter1 = /* some iterator yielding: 1, 2, 3 */;\n iter2 = /* some iterator yielding: 3, 4, 5, 6, 7, 8, 9 */;\n concatenated = concat(iter1, DataLimits.cqlLimits(3).filter(iter2));\n{code}\n\nwould result in {{concatenated}} yielding 1 through 9, which is incorrect.\n\nThe patch implements the following changes:\n\n * {{concat}} [now allows|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:13482-trunk#diff-57d0dfa95504bfd17d30539b3b338c0cL205] different amount of columns from iterators, but returned columns will be a union of the two iterators\n * {{BaseIterator}} [would now|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:13482-trunk#diff-d14a4b314544d3720010343e330e7e3cR125] take the {{stop}} from the child iterator, which means that the iterator was stopped, even if it might have contents, it will not yield additional data. Unfortunately, since the iterator itself had a stopping transformation, and it's own {{stop}} reference is \"leaked\" during the creation of the transformation, we have to save (and check) both the original stop and the child one. Here, the naming might be a bit off.\n * {{cacheIterator}} [is now|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:13482-trunk#diff-2e17efa5977a71330df6651d3bec0d12R424] using a custom wrapped iterator that will not call {{stop}} on the wrapping iterator, but previous problem with {{Unfiltered}} is now fixed\n\n|[3.0|https://github.com/apache/cassandra/compare/cassandra-3.0...ifesdjeen:13482-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-3.0-testall/]|[dtest|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-3.0-dtest/]|\n|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...ifesdjeen:13482-3.11]|[testall|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-3.11-testall/]|[dtest|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-3.11-dtest/]|\n|[trunk|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:13482-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-trunk-testall/]|[dtest|http://cassci.datastax.com/view/Dev/view/ifesdjeen/job/ifesdjeen-13482-trunk-dtest/]|","from":"developer"},{"body":"lgtm, +1.","from":"developer"},{"body":"Thank you for the review,\n\nCommitted to 3.0 with [7251c9559805d83423ca5ddbe4f955ce668c3d9a|https://github.com/apache/cassandra/commit/7251c9559805d83423ca5ddbe4f955ce668c3d9a] and merged up to [3.11|https://github.com/apache/cassandra/commit/29db2511621e420b8d64c867a16e317589397d36] and [trunk|https://github.com/apache/cassandra/commit/f48a319ac884ef8d6eb54db3176ea2acf627bb89].","from":"developer"}],"created":"2017-04-28T10:52:21.000+0000","description":"The problem is reproducible on 3.0 with:\n\n{code}\n-# row_cache_class_name: org.apache.cassandra.cache.OHCProvider\n+row_cache_class_name: org.apache.cassandra.cache.OHCProvider\n\n-row_cache_size_in_mb: 0\n+row_cache_size_in_mb: 100\n{code}\n\nTable setup:\n\n{code}\nCREATE TABLE cache_tables (pk int, v1 int, v2 int, v3 int, primary key (pk, v1)) WITH CACHING = { 'keys': 'ALL', 'rows_per_partition': '1' } ;\n{code}\n\nNo data is required, only a head query (or any pk/ck query but with full partitions cached). \n\n{code}\nselect * from cross_page_queries where pk = 10000 ;\n{code}\n\n{code}\njava.lang.AssertionError: null\n at org.apache.cassandra.db.rows.UnfilteredRowIterators.concat(UnfilteredRowIterators.java:193) ~[main/:na]\n at org.apache.cassandra.db.SinglePartitionReadCommand.getThroughCache(SinglePartitionReadCommand.java:461) ~[main/:na]\n at org.apache.cassandra.db.SinglePartitionReadCommand.queryStorage(SinglePartitionReadCommand.java:358) ~[main/:na]\n at org.apache.cassandra.db.ReadCommand.executeLocally(ReadCommand.java:395) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$LocalReadRunnable.runMayThrow(StorageProxy.java:1794) ~[main/:na]\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:2472) ~[main/:na]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_121]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:164) ~[main/:na]\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:136) [main/:na]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_121]\n{code}","issue_id":"13067657","key":"CASSANDRA-13482","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-07-12T13:55:52.000+0000","role":"fixed_distractor","summary":"NPE on non-existing row read when row cache is enabled"} {"case_id":"13088570","cluster":"DISTRACTOR-CASSANDRA-13700","comments":[{"body":"[~jkni] Fantastic debugging here, Joel. We have seen this problem, as well, with missing STATUS and TOKENS entries.\n\nI followed this through, and I believe you are correct. Just to point out (because I had to dig and reason through it), the key problem is (as Joel points out) the shared mutable state of {{HeartBeatState}}. In {{Gossiper.getStateForVersionBiggerThan}}, when the local node is building up the {{Map states}} about itself, if any states are added after the function returns *and* the heartbeat is incremented before serialization, the peer will get the updated heartbeat value but not the updated states (as we the set of states for the local node that we're sending over was already constructed a priori the serialization).\n\nOff the top of my head, I think there are at least two possible ways to fix this:\n\n- clone the {{HeartBeatState}} when constructing the {{EndpointState}} to return from {{Gossiper.getStateForVersionBiggerThan}}. That way it's not referencing mutable heartbeat state.\n- execute the {{GossipTask}} on the same thread the we receive the gossip syn/ack/ack2 messages (on the {{Stage.GOSSIP}} thread). That way we force (almost) all references to gossip's stated mutable state into one thread.\n\nThe first option is simpler, smaller in scope, and certainly safer.\nThe second option is has performance implications, especially if the {{GossipTask}} takes a while to execute, then we could start backing up the tasks on the stage. This option, though, has the \"possibility\" of eliminating more of the state race bugs that we seems to continually uncover as time goes on. (Side note: there are still some updates to local Gossip state from the main thread (via {{StorageService}}) at startup, and the response to the {{EchoMessage}} is on the wrong thread, as well.)\n\nJoel, can you share the method of how you are able to reproduce this?","created":"2017-07-20T14:14:10.701+0000"},{"body":"We also might need to make {{HeartBeatState.version}} volatile, but I'm still thinking about it (just adding it here for discussion)","created":"2017-07-20T14:17:02.593+0000"},{"body":"Thanks, Jason! In this case, I agree the first option is safer for this issue. Something like the second likely makes sense eventually, at least as part of a larger audit of correctness issues in gossip. I believe your volatile suggestion is correct.\n\nI don't have a lot of helpful information to reproduce this; it reproduces in larger clusters, particularly with higher latency levels. We can see the effects locally with a few well-timed sleeps in MessagingService, but that isn't terribly representative.\n\nBranches pushed here:\n||branch||\n|[13700-2.1|https://github.com/jkni/cassandra/tree/13700-2.1]||\n|[13700-2.2|https://github.com/jkni/cassandra/tree/13700-2.2]||\n|[13700-3.0|https://github.com/jkni/cassandra/tree/13700-3.0]||\n|[13700-3.11|https://github.com/jkni/cassandra/tree/13700-3.11]||\n|[13700-trunk|https://github.com/jkni/cassandra/tree/13700-trunk]||\n\nThere's a somewhat conceptually similar issue when we bump the gossip generation in the middle of constructing a reply - I believe that's the cause in [CASSANDRA-11825], which presents similar problems. I'm choosing to address them separately because they're indeed distinct problems and 11825 requires an additional trigger (enabling and disabling gossip during runtime).","created":"2017-07-20T16:25:03.765+0000"},{"body":"+1","created":"2017-07-20T17:58:31.865+0000"},{"body":"re: CASSANDRA-11825. Yes, shared mutable state strikes again, and thanks for addressing them separately ;)","created":"2017-07-20T17:59:57.803+0000"},{"body":"Thanks! Tests looked good on all branches.\n\nCommitted to 2.1 as {{2290c0d4b0c20ce3407ae2c542e580c75a5ab337}} and merged forward through 2.2, 3.0, 3.11, and trunk.","created":"2017-07-24T20:29:30.433+0000"},{"body":"[~jkni] In our internal review of this change, we discovered that that patch can be made incrementally safer. In {{Gossiper#getStateForVersionBiggerThan()}}, we [get the generation|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/gms/Gossiper.java#L899], but it's possible the {{epState.getHeartBeatState()}} could have been swapped out before the next line executes (where we get the version).\n\nAre you ok if I make the following small change to {{Gossiper#getStateForVersionBiggerThan()}}:\n{code}\n- int localHbGeneration = epState.getHeartBeatState().getGeneration();\n- int localHbVersion = epState.getHeartBeatState().getHeartBeatVersion();\n+ HeartBeatState heartBeatState = epState.getHeartBeatState();\n+ int localHbGeneration = heartBeatState.getGeneration();\n+ int localHbVersion = heartBeatState.getHeartBeatVersion();\n{code}\n\nBasically, just grab a reference the the same `HeartBeatState` that we'll use for both the generation and version. Of course, this does not protect against that `HeartBeatState` instance being mutated, but at least we can be smarter about what we reference from the `epState`.","created":"2017-07-27T18:36:01.224+0000"},{"body":"I'm not sure of any places in the current codebase where the distinction matters in practice, but the change is cleaner and makes the code more tolerant of changes elsewhere, so +1.","created":"2017-08-01T14:32:05.535+0000"},{"body":"pushed the minor fix as sha {{6b927783ba0777d3dd7c2c3311b246a8dbce5b59}} to 2.1+.","created":"2017-08-29T15:54:34.702+0000"}],"conversations":[{"body":"In {{Gossiper.getStateForVersionBiggerThan}}, we add the {{HeartBeatState}} from the corresponding {{EndpointState}} to the {{EndpointState}} to send. When we're getting state for ourselves, this means that we add a reference to the local {{HeartBeatState}}. Then, once we've built a message (in either the Syn or Ack handler), we send it through the {{MessagingService}}. In the case that the {{MessagingService}} is sufficiently slow, the {{GossipTask}} may run before serialization of the Syn or Ack. This means that when the {{GossipTask}} acquires the gossip {{taskLock}}, it may increment the {{HeartBeatState}} version of the local node as stored in the endpoint state map. Then, when we finally serialize the Syn or Ack, we'll follow the reference to the {{HeartBeatState}} and serialize it with a higher version than we saw when constructing the Ack or Ack2.\n\nConsider the case where we see {{HeartBeatState}} with version 4 when constructing an Ack and send it through the {{MessagingService}}. Then, we add some piece of state with version 5 to our local {{EndpointState}}. If {{GossipTask}} runs and increases the {{HeartBeatState}} version to 6 before the {{MessageOut}} containing the Ack is serialized, the node receiving the Ack will believe it is current to version 6, despite the fact that it has never received a message containing the {{ApplicationState}} tagged with version 5.\n\nI've reproduced in this in several versions; so far, I believe this is possible in all versions.","from":"reporter","subject":"Heartbeats can cause gossip information to go permanently missing on certain nodes"},{"body":"[~jkni] Fantastic debugging here, Joel. We have seen this problem, as well, with missing STATUS and TOKENS entries.\n\nI followed this through, and I believe you are correct. Just to point out (because I had to dig and reason through it), the key problem is (as Joel points out) the shared mutable state of {{HeartBeatState}}. In {{Gossiper.getStateForVersionBiggerThan}}, when the local node is building up the {{Map states}} about itself, if any states are added after the function returns *and* the heartbeat is incremented before serialization, the peer will get the updated heartbeat value but not the updated states (as we the set of states for the local node that we're sending over was already constructed a priori the serialization).\n\nOff the top of my head, I think there are at least two possible ways to fix this:\n\n- clone the {{HeartBeatState}} when constructing the {{EndpointState}} to return from {{Gossiper.getStateForVersionBiggerThan}}. That way it's not referencing mutable heartbeat state.\n- execute the {{GossipTask}} on the same thread the we receive the gossip syn/ack/ack2 messages (on the {{Stage.GOSSIP}} thread). That way we force (almost) all references to gossip's stated mutable state into one thread.\n\nThe first option is simpler, smaller in scope, and certainly safer.\nThe second option is has performance implications, especially if the {{GossipTask}} takes a while to execute, then we could start backing up the tasks on the stage. This option, though, has the \"possibility\" of eliminating more of the state race bugs that we seems to continually uncover as time goes on. (Side note: there are still some updates to local Gossip state from the main thread (via {{StorageService}}) at startup, and the response to the {{EchoMessage}} is on the wrong thread, as well.)\n\nJoel, can you share the method of how you are able to reproduce this?","from":"developer"},{"body":"We also might need to make {{HeartBeatState.version}} volatile, but I'm still thinking about it (just adding it here for discussion)","from":"developer"},{"body":"Thanks, Jason! In this case, I agree the first option is safer for this issue. Something like the second likely makes sense eventually, at least as part of a larger audit of correctness issues in gossip. I believe your volatile suggestion is correct.\n\nI don't have a lot of helpful information to reproduce this; it reproduces in larger clusters, particularly with higher latency levels. We can see the effects locally with a few well-timed sleeps in MessagingService, but that isn't terribly representative.\n\nBranches pushed here:\n||branch||\n|[13700-2.1|https://github.com/jkni/cassandra/tree/13700-2.1]||\n|[13700-2.2|https://github.com/jkni/cassandra/tree/13700-2.2]||\n|[13700-3.0|https://github.com/jkni/cassandra/tree/13700-3.0]||\n|[13700-3.11|https://github.com/jkni/cassandra/tree/13700-3.11]||\n|[13700-trunk|https://github.com/jkni/cassandra/tree/13700-trunk]||\n\nThere's a somewhat conceptually similar issue when we bump the gossip generation in the middle of constructing a reply - I believe that's the cause in [CASSANDRA-11825], which presents similar problems. I'm choosing to address them separately because they're indeed distinct problems and 11825 requires an additional trigger (enabling and disabling gossip during runtime).","from":"developer"},{"body":"+1","from":"developer"},{"body":"re: CASSANDRA-11825. Yes, shared mutable state strikes again, and thanks for addressing them separately ;)","from":"developer"},{"body":"Thanks! Tests looked good on all branches.\n\nCommitted to 2.1 as {{2290c0d4b0c20ce3407ae2c542e580c75a5ab337}} and merged forward through 2.2, 3.0, 3.11, and trunk.","from":"developer"},{"body":"[~jkni] In our internal review of this change, we discovered that that patch can be made incrementally safer. In {{Gossiper#getStateForVersionBiggerThan()}}, we [get the generation|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/gms/Gossiper.java#L899], but it's possible the {{epState.getHeartBeatState()}} could have been swapped out before the next line executes (where we get the version).\n\nAre you ok if I make the following small change to {{Gossiper#getStateForVersionBiggerThan()}}:\n{code}\n- int localHbGeneration = epState.getHeartBeatState().getGeneration();\n- int localHbVersion = epState.getHeartBeatState().getHeartBeatVersion();\n+ HeartBeatState heartBeatState = epState.getHeartBeatState();\n+ int localHbGeneration = heartBeatState.getGeneration();\n+ int localHbVersion = heartBeatState.getHeartBeatVersion();\n{code}\n\nBasically, just grab a reference the the same `HeartBeatState` that we'll use for both the generation and version. Of course, this does not protect against that `HeartBeatState` instance being mutated, but at least we can be smarter about what we reference from the `epState`.","from":"developer"},{"body":"I'm not sure of any places in the current codebase where the distinction matters in practice, but the change is cleaner and makes the code more tolerant of changes elsewhere, so +1.","from":"developer"},{"body":"pushed the minor fix as sha {{6b927783ba0777d3dd7c2c3311b246a8dbce5b59}} to 2.1+.","from":"developer"}],"created":"2017-07-19T21:50:15.000+0000","description":"In {{Gossiper.getStateForVersionBiggerThan}}, we add the {{HeartBeatState}} from the corresponding {{EndpointState}} to the {{EndpointState}} to send. When we're getting state for ourselves, this means that we add a reference to the local {{HeartBeatState}}. Then, once we've built a message (in either the Syn or Ack handler), we send it through the {{MessagingService}}. In the case that the {{MessagingService}} is sufficiently slow, the {{GossipTask}} may run before serialization of the Syn or Ack. This means that when the {{GossipTask}} acquires the gossip {{taskLock}}, it may increment the {{HeartBeatState}} version of the local node as stored in the endpoint state map. Then, when we finally serialize the Syn or Ack, we'll follow the reference to the {{HeartBeatState}} and serialize it with a higher version than we saw when constructing the Ack or Ack2.\n\nConsider the case where we see {{HeartBeatState}} with version 4 when constructing an Ack and send it through the {{MessagingService}}. Then, we add some piece of state with version 5 to our local {{EndpointState}}. If {{GossipTask}} runs and increases the {{HeartBeatState}} version to 6 before the {{MessageOut}} containing the Ack is serialized, the node receiving the Ack will believe it is current to version 6, despite the fact that it has never received a message containing the {{ApplicationState}} tagged with version 5.\n\nI've reproduced in this in several versions; so far, I believe this is possible in all versions.","issue_id":"13088570","key":"CASSANDRA-13700","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-07-24T20:29:30.000+0000","role":"fixed_distractor","summary":"Heartbeats can cause gossip information to go permanently missing on certain nodes"} {"case_id":"13093846","cluster":"DISTRACTOR-CASSANDRA-13754","comments":[{"body":"What version are you on [~urandom] (or really, which version of netty is in the classpath) ? \n","created":"2017-08-10T18:08:49.389+0000"},{"body":"Cassandra 3.11.0, Netty 4.0.44.Final","created":"2017-08-10T21:43:04.107+0000"},{"body":"[~jjirsa], we are experiencing the same issue on Cassandra 3.11.0 and Netty 4.0.44. We have some custom triggers which create additional rows when {{INSERT}} s on specific CFs are executed. When inserting a lot of data, the Cassandra would constantly crash with {{OutOfMemoryError}} s. We analyzed a heap dump and came to the same conclusion, which is that the instances of {{FastThreadLocalThread}} were responsible for this behaviour. As a quick workaround, we added a call to {{FastThreadLocal.removeAll()}} to our triggers, which alleviates the problem for us.\n\nIt seems that there already was an issue once with {{FastThreadLocal}} s, as described in CASSANDRA-13033, but apparently the change by [~snazy] in CASSANDRA-13034 resurfaced the issue. For a proper fix, I think all {{FastThreadLocalThread}} instances should call {{FastThreadLocal.removeAll()}} when they are done with their work. A quick glance at the code indicates that this affects the classes {{CompressedInputStream}} , {{StreamingInboundHandler}} , {{NamedThreadFactory}} and {{SEPWorker}}.","created":"2017-08-31T07:38:07.961+0000"},{"body":"Thanks for the investigation, [~markusdlugi], I'm going to check what's happening. All {{Runnables}} should have been decorated with a {{try-finally}} to call {{FastThreadLocal.removeAll()}} to prevent exactly that.","created":"2017-08-31T10:03:34.122+0000"},{"body":"I see, you are talking about {{NamedTheadFactory.threadLocalDeallocator()}}, right? That should help for any threads created there, but as stated above, {{FastThreadLocalThread}} is also used in some other classes. Heap-wise, the most problematic one for us was the usage in {{SEPWorker}} I think, and this seems to have only been introduced with commit [1e92ce43a5a730f81d3f6cfd72e7f4b126db788a|https://github.com/apache/cassandra/commit/1e92ce43a5a730f81d3f6cfd72e7f4b126db788a] by [~tjake]. Maybe something similar can be done there and in all other classes making use of {{FastThreadLocalThread}}.","created":"2017-08-31T10:55:22.477+0000"},{"body":"Yea, that {{SEPWorker}} plus a couple more places (see [here|https://github.com/apache/cassandra/compare/cassandra-3.11...snazy:13754-ftl-leak-3.11?expand=1] and [here|https://github.com/apache/cassandra/compare/trunk...snazy:13754-ftl-leak-trunk?expand=1]). It's actually difficult to tell whether there was really something broken in all these places in the patch, but those changes don't (should not) harm. But if someone introduces a {{FastThreadLocal}} there in the future, it should have no negative impact.\nOur CI's currently checking the patches.\n[~markusdlugi], do you have any chance to verify the patches?","created":"2017-08-31T11:33:09.627+0000"},{"body":"[~snazy], I checked out the 13754-ftl-leak-3.11 branch and ran our load test with your patches. Unfortunately, the memory leak is still there, Cassandra first starts GCing like crazy and then throws an {{OutOfMemoryError}} after a while. So it seems like the call to {{FastThreadLocal.removeAll()}} is not properly executed.\n\nActually, while writing this it just struck me why that is the case. The thread created in {{SEPWorker}} does not stop after it finishes, but is instead returned to the pool and picks up new tasks. So it will never leave the {{run()}} method, and therefore also never execute {{FastThreadLocal.removeAll()}}. So I think we will have to include that call in the {{SEPWorker}} itself.","created":"2017-08-31T13:53:55.009+0000"},{"body":"[~markusdlugi] do you have a heap dump for me or references to what's actually in these {{FastThreadLocal}} s or which FTLs actually cause this? I suspect these are FTLs created in non-static fields, which probably need a different handling.\n\nCalling {{FastThreadLocal.removeAll()}} after every iteration in {{SEPWorker}} is not a good solution, because that would effectively kill the benefit of all static FTLs.","created":"2017-08-31T14:05:26.224+0000"},{"body":"Well, yea. Looking at the heap dump, that [~markusdlugi] provided, is looks like the node is \"just\" overloaded with too many and maybe too big writes in combination with a small heap. There are lots of {{BTree$Builder}} instances with live references in their {{Object[] values}} array to {{HeapByteBuffer}} instances, each holding a 1MB {{byte[]}}.\n{{BTree$Builder}} instances reset the {{Object[] values}} when finished - i.e. those builders are actively doing something (writes are happening at that time).\nTL;DR I don't think this is actually related to the issue that [~urandom] describes.\n\n[~urandom], can you explain what actually what these {{ThreadLocal}} instances referenced?","created":"2017-08-31T16:24:22.622+0000"},{"body":"[~snazy], I don't think the node is overloaded. I originally thought so as well, so I made a little experiment where I included a cap in our load test limiting the {{INSERT}} s per minute from ~25,000 to ~10,000. As a consequence, the node survived a little longer, but in the end it still died with an {{OutOfMemoryError}} after more data had been inserted. So it's not that there are too many active writes, it's just that the node fails after a certain amount of total writes, which indicates to me that a memory leak is indeed happening.\n\nI also had another look into the heap dump I sent you, and you are correct that the heap is mostly filled with {{BTree$Builder}} instances that still have stuff in their {{values}} array. However, if you look closer, you will notice that for each of these instances, the {{values}} array always contains {{null}} for the first couple of entries, and only after those there is still actual content. For some reason, the actual content always starts at index 28, whereas indices 0 - 27 are {{null}} - not sure if this is a coincidence? But you can also see that for all the {{BTree$Builder}} objects, the {{count}} attribute is 0, which also indicates to me that {{BTree$Builder.cleanup()}} has already run and those are not active writes. This theory is supported by the fact that my little workaround of manually calling {{FastThreadLocal.removeAll()}} actually works, because this means that no other objects except the {{FastThreadLocal}} s still have references to the builders.\n\nTherefore, I think we have two issues here:\n\n# {{SEPWorker}} is never cleaning the {{FastThreadLocal}} s, therefore accumulating references to otherwise dead objects - maybe we can include something to at least remove non-static entries regularly?\n# {{BTree$Builder}} seems to have an issue properly cleaning up after building, so the objects referenced by the {{FastThreadLocal}} s of the {{SEPWorker}} threads are very large and thus ultimately lead to the {{OutOfMemoryError}} s","created":"2017-09-01T07:41:34.501+0000"},{"body":"Your observation regarding {{BTree.Builder.values[]}} seems correct.\nHowever, {{SEPWorker}} must *not* remove the thread locals - it's the intention of these thread-locals to be kept for reuse.","created":"2017-09-01T10:21:21.256+0000"},{"body":"Your latest patch which resets the entire {{BTree$Builder.values}} array seemed to do the trick, entire load test is now running smoothly. No more crazy GCing and most importantly no {{OutOfMemoryError}} s. Thanks a lot for the fast support and help!","created":"2017-09-01T12:08:56.665+0000"},{"body":"Given that the FTL changes apparently do not have any influence to the OOM issuse, but look serious enough to fix, I've split them out into CASSANDRA-13838.\n\nPatch for this ticket is reduced to the BTree change.\nCI looks good.","created":"2017-09-01T14:19:33.865+0000"},{"body":"+1 for just https://github.com/apache/cassandra/commit/2cafd0b6b4bbc5a6ec5726d47d0093bdac3af19c to fix this and splitting out the other changes to a new ticket.","created":"2017-09-01T14:21:32.172+0000"},{"body":"and +1 for the patch.","created":"2017-09-01T14:22:06.721+0000"},{"body":"Committed as [bed7fa5ef8492d1ff3852cf299622a5ad4e0b621|https://github.com/apache/cassandra/commit/bed7fa5ef8492d1ff3852cf299622a5ad4e0b621] to [cassandra-3.11|https://github.com/apache/cassandra/tree/cassandra-3.11] and merged to trunk.\n","created":"2017-09-01T17:49:54.609+0000"},{"body":"I apologize for chiming in here so late, but I'm not sure this addresses what I was seeing. In my dumps, all of the heap is tied up in the {{ThreadLocalMap}} s of instances of {{FastThreadLocalThread}} for _Native-Transport-Requests_, _RequestResponseStage_, _ReadStage_, etc; I think what I was seeing is different than [~markusdlugi].\n\nSee the attached screenshot of the dominator tree view in MemoryAnalyzer.\n\nI can make a dump available, but be warned, it is 12G in size.","created":"2017-09-11T22:13:58.872+0000"},{"body":"[~urandom], I think it might still be the same issue. The threads you mentioned are all created by the {{SEPWorker}} as well, as you can also see in your screenshot where your {{FastThreadLocalThread}} has a reference to an instance of that class. Now I'm not sure whether the actual content of your {{ThreadLocalMap}} s is the same as in my heap dump - in my case, the maps mostly held instances of {{BTree$Builder}} , which then had references to many {{byte[]}} arrays. Maybe you can check if this is the case for you as well?\n\nOther than that, you could also try and see if the patches created by [~snazy] alleviate your issue.","created":"2017-09-12T06:29:56.424+0000"},{"body":"\n{quote}\nI think it might still be the same issue. The threads you mentioned are all created by the SEPWorker as well, as you can also see in your screenshot where your FastThreadLocalThread has a reference to an instance of that class. Now I'm not sure whether the actual content of your ThreadLocalMap s is the same as in my heap dump - in my case, the maps mostly held instances of BTree$Builder , which then had references to many byte[] arrays. Maybe you can check if this is the case for you as well?\n{quote}\n\nThere are no instances of {{BTree}} here, (see new screenshot attached).\n\n{quote}\nOther than that, you could also try and see if the patches created by Robert Stupp alleviate your issue.\n{quote}\n\nDo you mean [bed7fa5|https://github.com/apache/cassandra/commit/bed7fa5ef8492d1ff3852cf299622a5ad4e0b621]? I haven't applied that, no, but it doesn't look like I'm leaking anything {{BTree}} so I don't think that would help.","created":"2017-09-13T15:47:19.694+0000"},{"body":"[~urandom], it's not about the {{BTree.Builder}} instances, it's about what's kept referenced by those - and that is the bunch of {{HeapByteBuffer}} instances, as reported by [~markusdlugi] and fixed by the patch for this ticket. The screenshot you attached, matches the fixed issue. Therefore, I recommend to try the patch or a build from the recent 3.11/trunk branches and test again. Going go resolve this issue. If it's really something else that's causing the issue, I'd prefer to open another ticket, because there is already something committed for this ticket that addresses a very particular issue.","created":"2017-09-13T17:05:05.438+0000"},{"body":"We have deployed Cassandra 3.11.1 (snapshot build from September 25, 2017) into our 9-node loadtest cluster. We still see increasing heap utilization over time with ~ 140 Recycler$Stack instances consuming ~1,8 GB heap. See attached screen (Eclipse memory analyzer screen): cassandra_3.11.1_Recycler_memleak.png","created":"2017-10-01T11:48:56.088+0000"},{"body":"72hrs heap utilization increase with 3.11.1 snapshot build from Sept. 25, 2017 + cluster rolling restart marker => cassandra_3.11.1_snapshot_heaputilization.png","created":"2017-10-02T20:27:59.022+0000"},{"body":"I'll note that the fixed issue and what you're describing are probably two different things.\nIt might also be that the combination of recycling the btree-builders _and_ many cells in a partition. This is technically different from what's been fixed.\nIt would help a lot, if someone can come up with steps (ideally some code) to reproduce the issue as mat screenshots show that something happened but not why.","created":"2017-10-02T21:44:23.324+0000"},{"body":"[~snazy]: Created CASSANDRA-13929 and discussed a potential fix.","created":"2017-10-03T08:09:14.586+0000"}],"conversations":[{"body":"After a chronic bout of {{OutOfMemoryError}} in our development environment, a heap analysis is showing that more than 10G of our 12G heaps are consumed by the {{threadLocals}} members (instances of {{java.lang.ThreadLocalMap}}) of various {{io.netty.util.concurrent.FastThreadLocalThread}} instances. Reverting [cecbe17|https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=cecbe17e3eafc052acc13950494f7dddf026aa54] fixes the issue.","from":"reporter","subject":"BTree.Builder memory leak"},{"body":"What version are you on [~urandom] (or really, which version of netty is in the classpath) ? \n","from":"developer"},{"body":"Cassandra 3.11.0, Netty 4.0.44.Final","from":"developer"},{"body":"[~jjirsa], we are experiencing the same issue on Cassandra 3.11.0 and Netty 4.0.44. We have some custom triggers which create additional rows when {{INSERT}} s on specific CFs are executed. When inserting a lot of data, the Cassandra would constantly crash with {{OutOfMemoryError}} s. We analyzed a heap dump and came to the same conclusion, which is that the instances of {{FastThreadLocalThread}} were responsible for this behaviour. As a quick workaround, we added a call to {{FastThreadLocal.removeAll()}} to our triggers, which alleviates the problem for us.\n\nIt seems that there already was an issue once with {{FastThreadLocal}} s, as described in CASSANDRA-13033, but apparently the change by [~snazy] in CASSANDRA-13034 resurfaced the issue. For a proper fix, I think all {{FastThreadLocalThread}} instances should call {{FastThreadLocal.removeAll()}} when they are done with their work. A quick glance at the code indicates that this affects the classes {{CompressedInputStream}} , {{StreamingInboundHandler}} , {{NamedThreadFactory}} and {{SEPWorker}}.","from":"developer"},{"body":"Thanks for the investigation, [~markusdlugi], I'm going to check what's happening. All {{Runnables}} should have been decorated with a {{try-finally}} to call {{FastThreadLocal.removeAll()}} to prevent exactly that.","from":"developer"},{"body":"I see, you are talking about {{NamedTheadFactory.threadLocalDeallocator()}}, right? That should help for any threads created there, but as stated above, {{FastThreadLocalThread}} is also used in some other classes. Heap-wise, the most problematic one for us was the usage in {{SEPWorker}} I think, and this seems to have only been introduced with commit [1e92ce43a5a730f81d3f6cfd72e7f4b126db788a|https://github.com/apache/cassandra/commit/1e92ce43a5a730f81d3f6cfd72e7f4b126db788a] by [~tjake]. Maybe something similar can be done there and in all other classes making use of {{FastThreadLocalThread}}.","from":"developer"},{"body":"Yea, that {{SEPWorker}} plus a couple more places (see [here|https://github.com/apache/cassandra/compare/cassandra-3.11...snazy:13754-ftl-leak-3.11?expand=1] and [here|https://github.com/apache/cassandra/compare/trunk...snazy:13754-ftl-leak-trunk?expand=1]). It's actually difficult to tell whether there was really something broken in all these places in the patch, but those changes don't (should not) harm. But if someone introduces a {{FastThreadLocal}} there in the future, it should have no negative impact.\nOur CI's currently checking the patches.\n[~markusdlugi], do you have any chance to verify the patches?","from":"developer"},{"body":"[~snazy], I checked out the 13754-ftl-leak-3.11 branch and ran our load test with your patches. Unfortunately, the memory leak is still there, Cassandra first starts GCing like crazy and then throws an {{OutOfMemoryError}} after a while. So it seems like the call to {{FastThreadLocal.removeAll()}} is not properly executed.\n\nActually, while writing this it just struck me why that is the case. The thread created in {{SEPWorker}} does not stop after it finishes, but is instead returned to the pool and picks up new tasks. So it will never leave the {{run()}} method, and therefore also never execute {{FastThreadLocal.removeAll()}}. So I think we will have to include that call in the {{SEPWorker}} itself.","from":"developer"},{"body":"[~markusdlugi] do you have a heap dump for me or references to what's actually in these {{FastThreadLocal}} s or which FTLs actually cause this? I suspect these are FTLs created in non-static fields, which probably need a different handling.\n\nCalling {{FastThreadLocal.removeAll()}} after every iteration in {{SEPWorker}} is not a good solution, because that would effectively kill the benefit of all static FTLs.","from":"developer"},{"body":"Well, yea. Looking at the heap dump, that [~markusdlugi] provided, is looks like the node is \"just\" overloaded with too many and maybe too big writes in combination with a small heap. There are lots of {{BTree$Builder}} instances with live references in their {{Object[] values}} array to {{HeapByteBuffer}} instances, each holding a 1MB {{byte[]}}.\n{{BTree$Builder}} instances reset the {{Object[] values}} when finished - i.e. those builders are actively doing something (writes are happening at that time).\nTL;DR I don't think this is actually related to the issue that [~urandom] describes.\n\n[~urandom], can you explain what actually what these {{ThreadLocal}} instances referenced?","from":"developer"},{"body":"[~snazy], I don't think the node is overloaded. I originally thought so as well, so I made a little experiment where I included a cap in our load test limiting the {{INSERT}} s per minute from ~25,000 to ~10,000. As a consequence, the node survived a little longer, but in the end it still died with an {{OutOfMemoryError}} after more data had been inserted. So it's not that there are too many active writes, it's just that the node fails after a certain amount of total writes, which indicates to me that a memory leak is indeed happening.\n\nI also had another look into the heap dump I sent you, and you are correct that the heap is mostly filled with {{BTree$Builder}} instances that still have stuff in their {{values}} array. However, if you look closer, you will notice that for each of these instances, the {{values}} array always contains {{null}} for the first couple of entries, and only after those there is still actual content. For some reason, the actual content always starts at index 28, whereas indices 0 - 27 are {{null}} - not sure if this is a coincidence? But you can also see that for all the {{BTree$Builder}} objects, the {{count}} attribute is 0, which also indicates to me that {{BTree$Builder.cleanup()}} has already run and those are not active writes. This theory is supported by the fact that my little workaround of manually calling {{FastThreadLocal.removeAll()}} actually works, because this means that no other objects except the {{FastThreadLocal}} s still have references to the builders.\n\nTherefore, I think we have two issues here:\n\n# {{SEPWorker}} is never cleaning the {{FastThreadLocal}} s, therefore accumulating references to otherwise dead objects - maybe we can include something to at least remove non-static entries regularly?\n# {{BTree$Builder}} seems to have an issue properly cleaning up after building, so the objects referenced by the {{FastThreadLocal}} s of the {{SEPWorker}} threads are very large and thus ultimately lead to the {{OutOfMemoryError}} s","from":"developer"},{"body":"Your observation regarding {{BTree.Builder.values[]}} seems correct.\nHowever, {{SEPWorker}} must *not* remove the thread locals - it's the intention of these thread-locals to be kept for reuse.","from":"developer"},{"body":"Your latest patch which resets the entire {{BTree$Builder.values}} array seemed to do the trick, entire load test is now running smoothly. No more crazy GCing and most importantly no {{OutOfMemoryError}} s. Thanks a lot for the fast support and help!","from":"developer"},{"body":"Given that the FTL changes apparently do not have any influence to the OOM issuse, but look serious enough to fix, I've split them out into CASSANDRA-13838.\n\nPatch for this ticket is reduced to the BTree change.\nCI looks good.","from":"developer"},{"body":"+1 for just https://github.com/apache/cassandra/commit/2cafd0b6b4bbc5a6ec5726d47d0093bdac3af19c to fix this and splitting out the other changes to a new ticket.","from":"developer"},{"body":"and +1 for the patch.","from":"developer"},{"body":"Committed as [bed7fa5ef8492d1ff3852cf299622a5ad4e0b621|https://github.com/apache/cassandra/commit/bed7fa5ef8492d1ff3852cf299622a5ad4e0b621] to [cassandra-3.11|https://github.com/apache/cassandra/tree/cassandra-3.11] and merged to trunk.\n","from":"developer"},{"body":"I apologize for chiming in here so late, but I'm not sure this addresses what I was seeing. In my dumps, all of the heap is tied up in the {{ThreadLocalMap}} s of instances of {{FastThreadLocalThread}} for _Native-Transport-Requests_, _RequestResponseStage_, _ReadStage_, etc; I think what I was seeing is different than [~markusdlugi].\n\nSee the attached screenshot of the dominator tree view in MemoryAnalyzer.\n\nI can make a dump available, but be warned, it is 12G in size.","from":"developer"},{"body":"[~urandom], I think it might still be the same issue. The threads you mentioned are all created by the {{SEPWorker}} as well, as you can also see in your screenshot where your {{FastThreadLocalThread}} has a reference to an instance of that class. Now I'm not sure whether the actual content of your {{ThreadLocalMap}} s is the same as in my heap dump - in my case, the maps mostly held instances of {{BTree$Builder}} , which then had references to many {{byte[]}} arrays. Maybe you can check if this is the case for you as well?\n\nOther than that, you could also try and see if the patches created by [~snazy] alleviate your issue.","from":"developer"},{"body":"\n{quote}\nI think it might still be the same issue. The threads you mentioned are all created by the SEPWorker as well, as you can also see in your screenshot where your FastThreadLocalThread has a reference to an instance of that class. Now I'm not sure whether the actual content of your ThreadLocalMap s is the same as in my heap dump - in my case, the maps mostly held instances of BTree$Builder , which then had references to many byte[] arrays. Maybe you can check if this is the case for you as well?\n{quote}\n\nThere are no instances of {{BTree}} here, (see new screenshot attached).\n\n{quote}\nOther than that, you could also try and see if the patches created by Robert Stupp alleviate your issue.\n{quote}\n\nDo you mean [bed7fa5|https://github.com/apache/cassandra/commit/bed7fa5ef8492d1ff3852cf299622a5ad4e0b621]? I haven't applied that, no, but it doesn't look like I'm leaking anything {{BTree}} so I don't think that would help.","from":"developer"},{"body":"[~urandom], it's not about the {{BTree.Builder}} instances, it's about what's kept referenced by those - and that is the bunch of {{HeapByteBuffer}} instances, as reported by [~markusdlugi] and fixed by the patch for this ticket. The screenshot you attached, matches the fixed issue. Therefore, I recommend to try the patch or a build from the recent 3.11/trunk branches and test again. Going go resolve this issue. If it's really something else that's causing the issue, I'd prefer to open another ticket, because there is already something committed for this ticket that addresses a very particular issue.","from":"developer"},{"body":"We have deployed Cassandra 3.11.1 (snapshot build from September 25, 2017) into our 9-node loadtest cluster. We still see increasing heap utilization over time with ~ 140 Recycler$Stack instances consuming ~1,8 GB heap. See attached screen (Eclipse memory analyzer screen): cassandra_3.11.1_Recycler_memleak.png","from":"developer"},{"body":"72hrs heap utilization increase with 3.11.1 snapshot build from Sept. 25, 2017 + cluster rolling restart marker => cassandra_3.11.1_snapshot_heaputilization.png","from":"developer"},{"body":"I'll note that the fixed issue and what you're describing are probably two different things.\nIt might also be that the combination of recycling the btree-builders _and_ many cells in a partition. This is technically different from what's been fixed.\nIt would help a lot, if someone can come up with steps (ideally some code) to reproduce the issue as mat screenshots show that something happened but not why.","from":"developer"},{"body":"[~snazy]: Created CASSANDRA-13929 and discussed a potential fix.","from":"developer"}],"created":"2017-08-10T16:15:08.000+0000","description":"After a chronic bout of {{OutOfMemoryError}} in our development environment, a heap analysis is showing that more than 10G of our 12G heaps are consumed by the {{threadLocals}} members (instances of {{java.lang.ThreadLocalMap}}) of various {{io.netty.util.concurrent.FastThreadLocalThread}} instances. Reverting [cecbe17|https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=cecbe17e3eafc052acc13950494f7dddf026aa54] fixes the issue.","issue_id":"13093846","key":"CASSANDRA-13754","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-09-13T17:05:05.000+0000","role":"fixed_distractor","summary":"BTree.Builder memory leak"} {"case_id":"13137703","cluster":"DISTRACTOR-CASSANDRA-14227","comments":[{"body":"[~VincentWhite] feel free to add your PoC patch from CASSANDRA-14092. :)","created":"2018-02-11T13:46:25.252+0000"},{"body":"It would be great to get this in for 4.0. The simplest approach would be to change the {{localDeletionTime}} representation to {{long}} as proposed by [~VincentWhite], but we need to verify the impact of this on heap for TTL and non-TTL workloads.\r\n\r\nAlternatives include using a unsigned int32 to represent {{localDeletionTime}} and/or using a later start epoch (maybe 2000 or so) - but this would require some considerate work to maintain backward compatibility and would difficult interoperability with other systems - even though it's an internal structure.\r\n\r\n[~VincentWhite] would you be willing to work on this and maybe perform some stress tests to check impact on heap and GC of changing the {{localDeletionTime}} representation from {{int}} to {{long}} ?","created":"2018-04-05T20:30:18.237+0000"},{"body":"[~pauloricardomg] Upgrade Sstable change to resume the original ttl should be in scope of this bug or should we open another jira ticket for this?\r\nNote: This change is important for those users who are currently using CAP or CAP_NOWARN to avoid request rejection but want to resume to originally  set TTL when they upgrade to new cassandra version with fix.","created":"2019-06-04T05:49:20.933+0000"},{"body":"Resuming the {{localDeletionTime}}  value as discussed [here|https://jira.apache.org/jira/browse/CASSANDRA-14092?focusedCommentId=16341749&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-16341749]  after an update operation on a row with TTL does not look feasible. Please see the below test case:\r\n\r\n*Step 1:*\r\n\r\ninsert into tab3(id,key,value) VALUES('id2', 'key1', 2) USING TTL 630720000; (20 years).\r\n{code:java}\r\nsstabletojson: \r\n {\r\n \"partition\" : {\r\n \"key\" : [ \"id2\" ],\r\n \"position\" : 33\r\n },\r\n \"rows\" : [\r\n {\r\n \"type\" : \"row\",\r\n \"position\" : 75,\r\n \"clustering\" : [ \"key1\" ],\r\n \"liveness_info\" : { \"tstamp\" : \"2019-07-18T07:03:22.198Z\", \"ttl\" : 630720000, \"expires_at\" : \"2038-01-19T03:14:06Z\", \"expired\" : false },\r\n \"cells\" : [\r\n { \"name\" : \"value\", \"value\" : 2 }\r\n ]\r\n }\r\n ]\r\n }{code}\r\n*Step 2:* \r\n 1. select ttl(value) from tab3 where id='id2';\r\n\r\n*This results 584043668* *and not 630720000 !* and then updating the row with fetched ttl loses the original ttl value (20 years) so resuming {{localDeletionTime}}  based on ttl the value will not be feasible if someone chooses CAP or CAP_NOWARN option as in sstable, TTL value is no more 20 years (630720000).\r\n\r\n2. insert into tab3(id,key,value) VALUES('id2', 'key1', 3) USING TTL 584043668;\r\n{code:java}\r\nsstabletojson: \r\n \r\n {\r\n \"partition\" : {\r\n \"key\" : [ \"id2\" ],\r\n \"position\" : 33\r\n },\r\n \"rows\" : [\r\n {\r\n \"type\" : \"row\",\r\n \"position\" : 75,\r\n \"clustering\" : [ \"key1\" ],\r\n \"liveness_info\" : { \"tstamp\" : \"2019-07-18T08:54:27.921Z\", \"ttl\" : 584043668, \"expires_at\" : \"2038-01-19T03:14:06Z\", \"expired\" : false },\r\n \"cells\" : [\r\n { \"name\" : \"value\", \"value\" : 3 }\r\n ]\r\n }\r\n ]\r\n }{code}\r\n ","created":"2019-07-18T10:33:06.682+0000"},{"body":"{quote}This results 584043668 and not 630720000 ! and then updating the row with fetched ttl loses the original ttl value (20 years) so resuming localDeletionTime based on ttl the value will not be feasible if someone chooses CAP or CAP_NOWARN option as in sstable, TTL value is no more 20 years (630720000).\r\n{quote}\r\nThe fact that {{SELECT TTL(*)}} does not return the original TTL does not mean we cannot resume the original TTL once this is fixed. Data will need to be rewritten via \"upgradesstables\" to restore the original TTL once this limitation is fixed. If you are updating the row with fetched TTL you are overwriting the original value, so it's expected behavior that you lose the original TTL.\r\n\r\nIf you are willing to take a stab at fixing this, I think the simplest approach is to change the {{localDeletionTime}} field from integer to long.","created":"2019-07-23T13:23:30.081+0000"},{"body":"[~pauloricardomg] I believe there are many use-cases where users maintain the TTL of a row while updating it. Hence in all such cases original TTL will be lost. So, is it worth giving TTL recalculation feature in upgradesstable which won't be applicable to all users ? \r\n\r\nI agree that a permanent fix, changing {{localDeletionTime}} field from integer to long is simplest way to go but this change will be big in terms changes in number of classes and we need to carefully analyse the impact as well.","created":"2019-07-25T05:23:37.462+0000"},{"body":"Set this Jira to Severity = Critical and Complexity = Challenging to reflect the state of this problem, as time goes on, as well as the nature of the fix.","created":"2020-06-12T15:08:12.272+0000"},{"body":"This is a critical issue, that severely limits TTL functionality; the CAP_NOWARN workaround of CASSANDRA-14092 is already limiting our data lifecycle.  I'm not sure why this is high-severity issue was not planned for 4.0, what's missing?","created":"2021-11-29T10:18:52.062+0000"},{"body":"Removed assignee. I only found a few random state changes by the user in Jira, so this assignment also looked random. If this was incorrect, please reassign!","created":"2022-04-11T18:55:33.659+0000"},{"body":"Hi,\r\n\r\nmoving to this to review. Some comments:\r\n- Max TTL is now 68y and there is a new default NONE overflow policy, the other ones are kept for backwards compatibility.\r\n- All tests pass but for one upgrade rolling failing 10% of times. I am investigating it but given it is taking a lot of time and I can't repro locally I prefer to start with the review already.\r\n- The code has a few 'TODO 14227' comments which are points I would like to discuss with the reviewer such as {{StreamingTombstoneHistogramBuilder#mergeNearestPoints()}}\r\n- Most files only change int to longs, think about sentinel values while reviewing and possible side-effects\r\n- The PR has a few self-contained commits for an easier review.\r\n\r\n","created":"2022-09-29T07:26:10.400+0000"},{"body":"Hi,\r\n\r\nI have been asked to run some profiling to see the impact 'longs' would have. I have tested on a single node doing inserts with a TTL + selects against that for 4m at around 100Kops/s after a warm up run. The results are virtual identical to both trunk and 14227 both at jfr and stress tool perf reporting\r\n\r\nTrunk:\r\n\r\n !screenshot-1.png! \r\n\r\n14227\r\n\r\n !screenshot-2.png! \r\n\r\nIf anything 14227's number are slightly better but probably just test env noise. GC pauses, total times, latencies, etc are all identical as well. The test CQL was:\r\n\r\n{noformat}\r\nCQL|INSERT INTO test.test (id, type, text) VALUES ($RANDOM_20000, $RANDOM_10, 'TTL Profiling') USING TTL $RANDOM_500000\r\nCQL|SELECT id, type, text FROM test.test WHERE id=$RANDOM_20000\r\n{noformat}\r\n\r\nwhere RANDOM_X means a random number up to X. Here we can see inserts, inserts colliding and selects. I will be happy to repeat the test if anybody has any suggestions. I am not posting the jfr for security reasons as the env vars section contains sensitive data (call me paranoid) but I can share them with known people on request. \r\n\r\nEDIT: I have been asked to confirm both runs are under {{memtable_allocation_type: heap_buffers}} so it is on heap.","created":"2022-10-18T08:51:02.577+0000"},{"body":"After some further perf testing we've found out there is a 3% disk size increase hit we've like to avoid as discussed in the [ML|https://lists.apache.org/thread/fyf2d9jlrsor3hmz46qn282o5goooowd]. We're going to experiment and POC a bit a uint encoding to get rid of that and any extra flushes as well.","created":"2022-11-21T07:54:38.113+0000"},{"body":"Update: The uint approach is giving good results in the POC on disk size, memtable, etc. I need to rebase, put capping back and revisit all perf.","created":"2023-01-12T10:04:13.419+0000"},{"body":"Uint POC seems to be OK as well on the latency side. There are latencies on 2 completely different tests with even 2 diff test tooling:\r\n\r\n10 averaged runs latency\r\n[^screenshot-3.png]\r\n\r\n!screenshot-3.png!\r\n\r\nAveraged latency on a 1h test\r\n[^screenshot-4.png]\r\n\r\n!screenshot-4.png!","created":"2023-01-26T05:45:21.970+0000"},{"body":"Putting this back for a first pass review with the new Uint approach. I have done a final round of perf testing:\r\n * Perf testing revolved around 10M rows and runs with random/fixed TTL for periods of 2m, 5m and 10m and averaging 10 runs. The exception being the much longer latency test.\r\n * Longer tests tend to give less noisy results. 14227 vs trunk may randomly appear one slightly better than the other which I attribute to env noise and lack of dedicated perf testing HW.\r\n * Latency seems to be aligned as per comment above\r\n * JFRs available on request\r\n * Disk size, flushes etc seem to be aligned. Here is some table stats where sometimes 14227 is slightly better and sometimes the other way around !unnamed-1.png!\r\n\r\nThe PR has several commits to make it easier to review. Some refactoring of an earlier commit might be found in a later one but nothing too big and quite straightforward. I have left a bunch of 14227 TODO comments to bring the attention of the reviewer just for feedback. CI can be found in the PR, yet repeat tests haven't been run yet due to the size. We can do it at a later time once we're happy with the approach\r\n\r\nThe only thing to note is the addition of the NONE policy where it's later removed. That is the only bit it can't be easily collapsed, apologies in advance.","created":"2023-02-03T07:45:25.299+0000"},{"body":"[~jlewandowski] helpfully pointed out we should consider downgrade for this patch. It looks like we have an option to CAP to 2038, and if we default to this on upgrade I think a downgrade should be smooth. I'm not certain we even need to change the sstable version? Though I haven't looked closely.","created":"2023-02-14T11:23:22.931+0000"},{"body":"Seems like downgradability, at the time of writing, is still being discussed iiuc.\r\n\r\nNow we're also writing an int, with the uint approach, and the sstable version change is to signal a negative int being 'undefined data' vs 'a valid uint'. Regardless of the downgradability discussion _I think_ (not tested) an sstable scrub would suffice as scrub will either REJECT or CAP those entries. That is pending the downgradibility discussion detailed requirements: forward compatibility?, flag?, scrub?\r\n\r\nSthg to take into account though, thx for bringing it up.","created":"2023-02-15T06:51:51.651+0000"},{"body":"Downgradeability isn't under continuing discussion at present, no, but you are welcome to raise it for discussion.\r\n\r\nEither way, if it makes it simpler: this patch can easily avoid breaking downgrade, so my binding review feedback is that it *must* not.","created":"2023-02-15T08:28:39.137+0000"},{"body":"{quote} It looks like we have an option to CAP to 2038, and if we default to this on upgrade I think a downgrade should be smooth.{quote}\r\n[~benedict] Could you elaborate a bit what you mean? I am not sure that I am understanding what you have in mind. ","created":"2023-02-20T15:00:47.830+0000"},{"body":"By default we should reject TTLs past 2038, and while this setting is in place we should continue to write \\-nc\\- format sstables. Once the operator configures longer TTLs, we can write \\-oa\\- format sstables.","created":"2023-02-20T15:20:01.913+0000"},{"body":"I think, that I am misunderstanding what we call downgradability. What you propose is some kind of feature flag that delay the problem, no? If a problem occurs once we switch to the new sstable format then we cannot downgrade anymore.","created":"2023-02-21T13:27:38.978+0000"},{"body":"This is the canonical example that was given in the original thread more than a year ago, i.e. that a cluster should not write files that are incompatible with the version we are upgrading from until the operator agrees the upgrade has been successful, yes. If there are performance or other regressions encountered during an upgrade, recovery should be as simple as restarting with the prior version (until the operator decides the new version is performing adequately). This minimises risks, without fully eliminating them. \r\n\r\n","created":"2023-02-21T13:42:56.259+0000"},{"body":"I decided to have a quick look at adding support for dynamically determining the output format based on whether the TTL data requires it, and I noticed that EncodingStats and EncodingStats.Collector likely treat TTL incorrectly, using simple min/max on int. This might not cause any bugs, but it might, and we should fix it - we should probably consider what additional testing we might need to catch this kind of error elsewhere.","created":"2023-02-23T19:28:45.998+0000"},{"body":"Perf check report post review and big rebase [^C14227 Perf check 2023.03.21.pdf] LGTM","created":"2023-03-22T07:53:02.045+0000"},{"body":"[~benedict] a feature flag with upgrade tests etc has been added. I have provided links to the key bits of code [here|https://github.com/apache/cassandra/pull/1891#issuecomment-1541352411]. If we could +1 I would start getting ready for the merge :-)","created":"2023-05-10T05:00:26.101+0000"},{"body":"This looks good to me. I have left some final minor nits on the ticket.\r\n\r\nIf [~benedict] agrees with the approach for compatibility we can probably rebase and solve [those naming details|https://github.com/apache/cassandra/pull/1891/files#r1188414272] on the new CircleCI jobs.","created":"2023-05-17T16:56:06.746+0000"},{"body":"^The diff to the commit a few days ago is that the property has been moved to yaml as suggested by [~benedict]","created":"2023-05-18T05:14:34.903+0000"},{"body":"Hi [~benedict] we're a bit blocked here, if you could find a gap it should be a quick review for the FF having been moved to yaml and it would be highly appreciated. Thx.","created":"2023-05-22T05:00:53.511+0000"},{"body":"Sorry, the downside of lots of Jira traffic (incl from GitHub comments) is that I don't check the email notifications for a high traffic ticket.\r\n\r\nI won't have time to look at the code soon, but I trust you to have addressed my concerns given what you describe above. Feel free to proceed.","created":"2023-05-22T09:02:22.572+0000"},{"body":"Thx, it should be pretty non-controversial and we can always tweak the name etc.","created":"2023-05-22T09:51:02.764+0000"},{"body":"Latest perf check after rebases etc LGTM [^C14227 Perf check 2023.05.26.pdf]\r\n\r\nThe merge is coming...","created":"2023-05-26T09:53:21.076+0000"},{"body":"Looks good to me after rebase, and CI also looks good. +1 from me. The few remaining, trivial suggestions about comments can be addressed on commit.","created":"2023-05-30T14:37:39.326+0000"},{"body":"[~benedict] Everything looks good as far as we can see. Do you think you could find a gap see if we can +1 and merge :-) You're probably most interested in the yaml feature flag commit [here|https://github.com/apache/cassandra/pull/1891/commits/159eabbc7a0dec932864447ecffbaae5ffe66c4c]","created":"2023-05-30T14:42:56.523+0000"},{"body":"Merged. Thx everyone for all the help!","created":"2023-06-05T05:27:18.671+0000"},{"body":"Shouldn't CASSANDRA_4 mode be using the `nb` format ? \r\n\r\nIt's currently using `nc`, but looking through 4.0 and 4.1 latest versions they are on `nb`.\r\n(See CASSANDRA-18933)","created":"2023-11-01T08:51:19.754+0000"},{"body":"That is weird because we always talked about 'nc' and I remember testing 4.1 with nc. I will have to check...","created":"2023-11-06T07:48:55.694+0000"}],"conversations":[{"body":"The maximum expiration timestamp that can be represented by the storage engine is\r\n2038-01-19T03:14:06+00:00 due to the encoding of {{localExpirationTime}} as an int32.\r\n\r\nOn CASSANDRA-14092 we added an overflow policy which rejects requests with expiration above the maximum date as a temporary measure, but we should remove this limitation by updating the storage engine to support at least the maximum allowed TTL of 20 years.","from":"reporter","subject":"Extend maximum expiration date"},{"body":"[~VincentWhite] feel free to add your PoC patch from CASSANDRA-14092. :)","from":"developer"},{"body":"It would be great to get this in for 4.0. The simplest approach would be to change the {{localDeletionTime}} representation to {{long}} as proposed by [~VincentWhite], but we need to verify the impact of this on heap for TTL and non-TTL workloads.\r\n\r\nAlternatives include using a unsigned int32 to represent {{localDeletionTime}} and/or using a later start epoch (maybe 2000 or so) - but this would require some considerate work to maintain backward compatibility and would difficult interoperability with other systems - even though it's an internal structure.\r\n\r\n[~VincentWhite] would you be willing to work on this and maybe perform some stress tests to check impact on heap and GC of changing the {{localDeletionTime}} representation from {{int}} to {{long}} ?","from":"developer"},{"body":"[~pauloricardomg] Upgrade Sstable change to resume the original ttl should be in scope of this bug or should we open another jira ticket for this?\r\nNote: This change is important for those users who are currently using CAP or CAP_NOWARN to avoid request rejection but want to resume to originally  set TTL when they upgrade to new cassandra version with fix.","from":"developer"},{"body":"Resuming the {{localDeletionTime}}  value as discussed [here|https://jira.apache.org/jira/browse/CASSANDRA-14092?focusedCommentId=16341749&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-16341749]  after an update operation on a row with TTL does not look feasible. Please see the below test case:\r\n\r\n*Step 1:*\r\n\r\ninsert into tab3(id,key,value) VALUES('id2', 'key1', 2) USING TTL 630720000; (20 years).\r\n{code:java}\r\nsstabletojson: \r\n {\r\n \"partition\" : {\r\n \"key\" : [ \"id2\" ],\r\n \"position\" : 33\r\n },\r\n \"rows\" : [\r\n {\r\n \"type\" : \"row\",\r\n \"position\" : 75,\r\n \"clustering\" : [ \"key1\" ],\r\n \"liveness_info\" : { \"tstamp\" : \"2019-07-18T07:03:22.198Z\", \"ttl\" : 630720000, \"expires_at\" : \"2038-01-19T03:14:06Z\", \"expired\" : false },\r\n \"cells\" : [\r\n { \"name\" : \"value\", \"value\" : 2 }\r\n ]\r\n }\r\n ]\r\n }{code}\r\n*Step 2:* \r\n 1. select ttl(value) from tab3 where id='id2';\r\n\r\n*This results 584043668* *and not 630720000 !* and then updating the row with fetched ttl loses the original ttl value (20 years) so resuming {{localDeletionTime}}  based on ttl the value will not be feasible if someone chooses CAP or CAP_NOWARN option as in sstable, TTL value is no more 20 years (630720000).\r\n\r\n2. insert into tab3(id,key,value) VALUES('id2', 'key1', 3) USING TTL 584043668;\r\n{code:java}\r\nsstabletojson: \r\n \r\n {\r\n \"partition\" : {\r\n \"key\" : [ \"id2\" ],\r\n \"position\" : 33\r\n },\r\n \"rows\" : [\r\n {\r\n \"type\" : \"row\",\r\n \"position\" : 75,\r\n \"clustering\" : [ \"key1\" ],\r\n \"liveness_info\" : { \"tstamp\" : \"2019-07-18T08:54:27.921Z\", \"ttl\" : 584043668, \"expires_at\" : \"2038-01-19T03:14:06Z\", \"expired\" : false },\r\n \"cells\" : [\r\n { \"name\" : \"value\", \"value\" : 3 }\r\n ]\r\n }\r\n ]\r\n }{code}\r\n ","from":"developer"},{"body":"{quote}This results 584043668 and not 630720000 ! and then updating the row with fetched ttl loses the original ttl value (20 years) so resuming localDeletionTime based on ttl the value will not be feasible if someone chooses CAP or CAP_NOWARN option as in sstable, TTL value is no more 20 years (630720000).\r\n{quote}\r\nThe fact that {{SELECT TTL(*)}} does not return the original TTL does not mean we cannot resume the original TTL once this is fixed. Data will need to be rewritten via \"upgradesstables\" to restore the original TTL once this limitation is fixed. If you are updating the row with fetched TTL you are overwriting the original value, so it's expected behavior that you lose the original TTL.\r\n\r\nIf you are willing to take a stab at fixing this, I think the simplest approach is to change the {{localDeletionTime}} field from integer to long.","from":"developer"},{"body":"[~pauloricardomg] I believe there are many use-cases where users maintain the TTL of a row while updating it. Hence in all such cases original TTL will be lost. So, is it worth giving TTL recalculation feature in upgradesstable which won't be applicable to all users ? \r\n\r\nI agree that a permanent fix, changing {{localDeletionTime}} field from integer to long is simplest way to go but this change will be big in terms changes in number of classes and we need to carefully analyse the impact as well.","from":"developer"},{"body":"Set this Jira to Severity = Critical and Complexity = Challenging to reflect the state of this problem, as time goes on, as well as the nature of the fix.","from":"developer"},{"body":"This is a critical issue, that severely limits TTL functionality; the CAP_NOWARN workaround of CASSANDRA-14092 is already limiting our data lifecycle.  I'm not sure why this is high-severity issue was not planned for 4.0, what's missing?","from":"developer"},{"body":"Removed assignee. I only found a few random state changes by the user in Jira, so this assignment also looked random. If this was incorrect, please reassign!","from":"developer"},{"body":"Hi,\r\n\r\nmoving to this to review. Some comments:\r\n- Max TTL is now 68y and there is a new default NONE overflow policy, the other ones are kept for backwards compatibility.\r\n- All tests pass but for one upgrade rolling failing 10% of times. I am investigating it but given it is taking a lot of time and I can't repro locally I prefer to start with the review already.\r\n- The code has a few 'TODO 14227' comments which are points I would like to discuss with the reviewer such as {{StreamingTombstoneHistogramBuilder#mergeNearestPoints()}}\r\n- Most files only change int to longs, think about sentinel values while reviewing and possible side-effects\r\n- The PR has a few self-contained commits for an easier review.\r\n\r\n","from":"developer"},{"body":"Hi,\r\n\r\nI have been asked to run some profiling to see the impact 'longs' would have. I have tested on a single node doing inserts with a TTL + selects against that for 4m at around 100Kops/s after a warm up run. The results are virtual identical to both trunk and 14227 both at jfr and stress tool perf reporting\r\n\r\nTrunk:\r\n\r\n !screenshot-1.png! \r\n\r\n14227\r\n\r\n !screenshot-2.png! \r\n\r\nIf anything 14227's number are slightly better but probably just test env noise. GC pauses, total times, latencies, etc are all identical as well. The test CQL was:\r\n\r\n{noformat}\r\nCQL|INSERT INTO test.test (id, type, text) VALUES ($RANDOM_20000, $RANDOM_10, 'TTL Profiling') USING TTL $RANDOM_500000\r\nCQL|SELECT id, type, text FROM test.test WHERE id=$RANDOM_20000\r\n{noformat}\r\n\r\nwhere RANDOM_X means a random number up to X. Here we can see inserts, inserts colliding and selects. I will be happy to repeat the test if anybody has any suggestions. I am not posting the jfr for security reasons as the env vars section contains sensitive data (call me paranoid) but I can share them with known people on request. \r\n\r\nEDIT: I have been asked to confirm both runs are under {{memtable_allocation_type: heap_buffers}} so it is on heap.","from":"developer"},{"body":"After some further perf testing we've found out there is a 3% disk size increase hit we've like to avoid as discussed in the [ML|https://lists.apache.org/thread/fyf2d9jlrsor3hmz46qn282o5goooowd]. We're going to experiment and POC a bit a uint encoding to get rid of that and any extra flushes as well.","from":"developer"},{"body":"Update: The uint approach is giving good results in the POC on disk size, memtable, etc. I need to rebase, put capping back and revisit all perf.","from":"developer"},{"body":"Uint POC seems to be OK as well on the latency side. There are latencies on 2 completely different tests with even 2 diff test tooling:\r\n\r\n10 averaged runs latency\r\n[^screenshot-3.png]\r\n\r\n!screenshot-3.png!\r\n\r\nAveraged latency on a 1h test\r\n[^screenshot-4.png]\r\n\r\n!screenshot-4.png!","from":"developer"},{"body":"Putting this back for a first pass review with the new Uint approach. I have done a final round of perf testing:\r\n * Perf testing revolved around 10M rows and runs with random/fixed TTL for periods of 2m, 5m and 10m and averaging 10 runs. The exception being the much longer latency test.\r\n * Longer tests tend to give less noisy results. 14227 vs trunk may randomly appear one slightly better than the other which I attribute to env noise and lack of dedicated perf testing HW.\r\n * Latency seems to be aligned as per comment above\r\n * JFRs available on request\r\n * Disk size, flushes etc seem to be aligned. Here is some table stats where sometimes 14227 is slightly better and sometimes the other way around !unnamed-1.png!\r\n\r\nThe PR has several commits to make it easier to review. Some refactoring of an earlier commit might be found in a later one but nothing too big and quite straightforward. I have left a bunch of 14227 TODO comments to bring the attention of the reviewer just for feedback. CI can be found in the PR, yet repeat tests haven't been run yet due to the size. We can do it at a later time once we're happy with the approach\r\n\r\nThe only thing to note is the addition of the NONE policy where it's later removed. That is the only bit it can't be easily collapsed, apologies in advance.","from":"developer"},{"body":"[~jlewandowski] helpfully pointed out we should consider downgrade for this patch. It looks like we have an option to CAP to 2038, and if we default to this on upgrade I think a downgrade should be smooth. I'm not certain we even need to change the sstable version? Though I haven't looked closely.","from":"developer"},{"body":"Seems like downgradability, at the time of writing, is still being discussed iiuc.\r\n\r\nNow we're also writing an int, with the uint approach, and the sstable version change is to signal a negative int being 'undefined data' vs 'a valid uint'. Regardless of the downgradability discussion _I think_ (not tested) an sstable scrub would suffice as scrub will either REJECT or CAP those entries. That is pending the downgradibility discussion detailed requirements: forward compatibility?, flag?, scrub?\r\n\r\nSthg to take into account though, thx for bringing it up.","from":"developer"},{"body":"Downgradeability isn't under continuing discussion at present, no, but you are welcome to raise it for discussion.\r\n\r\nEither way, if it makes it simpler: this patch can easily avoid breaking downgrade, so my binding review feedback is that it *must* not.","from":"developer"},{"body":"{quote} It looks like we have an option to CAP to 2038, and if we default to this on upgrade I think a downgrade should be smooth.{quote}\r\n[~benedict] Could you elaborate a bit what you mean? I am not sure that I am understanding what you have in mind. ","from":"developer"},{"body":"By default we should reject TTLs past 2038, and while this setting is in place we should continue to write \\-nc\\- format sstables. Once the operator configures longer TTLs, we can write \\-oa\\- format sstables.","from":"developer"},{"body":"I think, that I am misunderstanding what we call downgradability. What you propose is some kind of feature flag that delay the problem, no? If a problem occurs once we switch to the new sstable format then we cannot downgrade anymore.","from":"developer"},{"body":"This is the canonical example that was given in the original thread more than a year ago, i.e. that a cluster should not write files that are incompatible with the version we are upgrading from until the operator agrees the upgrade has been successful, yes. If there are performance or other regressions encountered during an upgrade, recovery should be as simple as restarting with the prior version (until the operator decides the new version is performing adequately). This minimises risks, without fully eliminating them. \r\n\r\n","from":"developer"},{"body":"I decided to have a quick look at adding support for dynamically determining the output format based on whether the TTL data requires it, and I noticed that EncodingStats and EncodingStats.Collector likely treat TTL incorrectly, using simple min/max on int. This might not cause any bugs, but it might, and we should fix it - we should probably consider what additional testing we might need to catch this kind of error elsewhere.","from":"developer"},{"body":"Perf check report post review and big rebase [^C14227 Perf check 2023.03.21.pdf] LGTM","from":"developer"},{"body":"[~benedict] a feature flag with upgrade tests etc has been added. I have provided links to the key bits of code [here|https://github.com/apache/cassandra/pull/1891#issuecomment-1541352411]. If we could +1 I would start getting ready for the merge :-)","from":"developer"},{"body":"This looks good to me. I have left some final minor nits on the ticket.\r\n\r\nIf [~benedict] agrees with the approach for compatibility we can probably rebase and solve [those naming details|https://github.com/apache/cassandra/pull/1891/files#r1188414272] on the new CircleCI jobs.","from":"developer"},{"body":"^The diff to the commit a few days ago is that the property has been moved to yaml as suggested by [~benedict]","from":"developer"},{"body":"Hi [~benedict] we're a bit blocked here, if you could find a gap it should be a quick review for the FF having been moved to yaml and it would be highly appreciated. Thx.","from":"developer"},{"body":"Sorry, the downside of lots of Jira traffic (incl from GitHub comments) is that I don't check the email notifications for a high traffic ticket.\r\n\r\nI won't have time to look at the code soon, but I trust you to have addressed my concerns given what you describe above. Feel free to proceed.","from":"developer"},{"body":"Thx, it should be pretty non-controversial and we can always tweak the name etc.","from":"developer"},{"body":"Latest perf check after rebases etc LGTM [^C14227 Perf check 2023.05.26.pdf]\r\n\r\nThe merge is coming...","from":"developer"},{"body":"Looks good to me after rebase, and CI also looks good. +1 from me. The few remaining, trivial suggestions about comments can be addressed on commit.","from":"developer"},{"body":"[~benedict] Everything looks good as far as we can see. Do you think you could find a gap see if we can +1 and merge :-) You're probably most interested in the yaml feature flag commit [here|https://github.com/apache/cassandra/pull/1891/commits/159eabbc7a0dec932864447ecffbaae5ffe66c4c]","from":"developer"},{"body":"Merged. Thx everyone for all the help!","from":"developer"},{"body":"Shouldn't CASSANDRA_4 mode be using the `nb` format ? \r\n\r\nIt's currently using `nc`, but looking through 4.0 and 4.1 latest versions they are on `nb`.\r\n(See CASSANDRA-18933)","from":"developer"},{"body":"That is weird because we always talked about 'nc' and I remember testing 4.1 with nc. I will have to check...","from":"developer"}],"created":"2018-02-11T13:44:31.000+0000","description":"The maximum expiration timestamp that can be represented by the storage engine is\r\n2038-01-19T03:14:06+00:00 due to the encoding of {{localExpirationTime}} as an int32.\r\n\r\nOn CASSANDRA-14092 we added an overflow policy which rejects requests with expiration above the maximum date as a temporary measure, but we should remove this limitation by updating the storage engine to support at least the maximum allowed TTL of 20 years.","issue_id":"13137703","key":"CASSANDRA-14227","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2023-06-05T05:23:08.000+0000","role":"fixed_distractor","summary":"Extend maximum expiration date"} {"case_id":"13143384","cluster":"DISTRACTOR-CASSANDRA-14297","comments":[{"body":"First implementation up on github along with a lot of unit tests. I'll start doing some more e2e testing using ccm just to make sure all the edge cases are covered but if someone ([~aweisberg] or [~jasobrown] perhaps) wants to review that would be excellent.\r\n||trunk||\r\n|[pull request|https://github.com/apache/cassandra/pull/212]|\r\n|[!https://circleci.com/gh/jolynch/cassandra/tree/CASSANDRA-14297.png?circle-token= 1102a59698d04899ec971dd36e925928f7b521f5!|https://circleci.com/gh/jolynch/cassandra/tree/CASSANDRA-14297] |\r\n\r\n ","created":"2018-03-27T01:05:00.293+0000"},{"body":"LGTM. +1","created":"2018-05-23T16:01:17.523+0000"},{"body":"Looks like CASSANDRA-14447 refactoring has given me a nasty merge conflict. I'll work on rebasing the patchset but if a someone has time to give me quick feedback on if this idea is mergeable I'd appreciate it before investing more time into it. I do think this change makes the connectivity checker feature much more useful for operators trying to restart their databases without dropping traffic.","created":"2018-05-26T00:17:17.366+0000"},{"body":"I updated the patch to fix the merge conflicts, and reduced it two just two options to make life easier (the default is tuned to wait for all but 2 local DC nodes and not care about non local DC == local_quorum).\r\n\r\nThis is ready for review. I hope we get it in before 4.0 because the user interface of a percentage is something users can't reliably set (compared to the count where there are correct answers for each use case).","created":"2018-08-24T00:46:34.259+0000"},{"body":"I'm changing this to a bug since I think the current user interface is not possible for users to correctly configure and I hope we don't ship 4.0 with the percentage option instead of a count. If someone thinks that there are plausible settings of the existing configuration options users can use we can change this back to an improvement.","created":"2018-09-26T17:18:56.809+0000"},{"body":"Alright, per the discussion on [IRC|https://wilderness.apache.org/channels/?f=cassandra-dev/2018-10-17#1539793033] with Ariel and Jason, we've decided that instead of counts we should always wait for all but a single local DC node and replace the percentage option with:\r\n{noformat}\r\nblock_for_remote_dcs: \r\n{noformat}\r\nThe startup connectivity checker will wait for all but a single node in the local datacenter, and if you want to block startup on every datacenter having only a single node down you can set this to true.\r\n\r\nThe timeout will be the fallback for when multiple nodes are down in a local DC.","created":"2018-10-25T00:16:23.651+0000"},{"body":"Ok I've uploaded a patch to my branch that does what was asked in IRC I believe. Let me know if it looks good and I can run dtests and such against it.","created":"2018-10-26T01:59:28.987+0000"},{"body":"I left some comments on the PR. Looks good.","created":"2018-11-01T14:51:03.403+0000"},{"body":"[~aweisberg] Awesome I just rebased and merged in your suggestions:\r\n\r\n||trunk||\r\n|[fdd8173f|https://github.com/apache/cassandra/pull/212/commits/fdd8173f]|\r\n|[!https://circleci.com/gh/jolynch/cassandra/tree/CASSANDRA-14297.png?circle-token= 1102a59698d04899ec971dd36e925928f7b521f5!|https://circleci.com/gh/jolynch/cassandra/tree/CASSANDRA-14297] |\r\n\r\nDtests are running.","created":"2018-11-07T23:01:59.183+0000"},{"body":"+1 🚢 it.\r\n\r\nThere is one unused import in [StartupClusterConnectivityCheckerTest.java|https://github.com/apache/cassandra/pull/212/files#diff-c74adeeae072ee4af35c12a157cd7d61L26] I'll fix on commit.","created":"2018-11-12T17:27:46.252+0000"},{"body":"Committed as [801cb70ee811c956e987718a00695638d5bec1b6|https://github.com/apache/cassandra/commit/801cb70ee811c956e987718a00695638d5bec1b6] thanks!\r\n\r\nI also added a NEWS.txt and CHANGES.txt entries. I also added Patchy by XYZ; Reviewed by XYZ for CASSANDRA-14297 to the commit message.","created":"2018-11-12T17:47:16.899+0000"},{"body":"Sweet, thanks! Yea I was holding off on adding the NEWs/CHANGES entries until you marked it ready to commit. In the future I'll include it with the dtest run.\r\n\r\nThanks for all the great feedback, I think this feature is much more valuable now to users.","created":"2018-11-12T18:45:42.077+0000"}],"conversations":[{"body":"As I commented in CASSANDRA-13993, the current wait for functionality is a great step in the right direction, but I don't think that the current setting (70% of nodes in the cluster) is the right configuration option. First I think this because 70% will not protect against errors as if you wait for 70% of the cluster you could still very easily have {{UnavailableException}} or {{ReadTimeoutException}} exceptions. This is because if you have even two nodes down in different racks in a Cassandra cluster these exceptions are possible (or with the default {{num_tokens}} setting of 256 it is basically guaranteed). Second I think this option is not easy for operators to set, the only setting I could think of that would \"just work\" is 100%.\r\n\r\nI proposed in that ticket instead of having `block_for_peers_percentage` defaulting to 70%, we instead have `block_for_peers` as a count of nodes that are allowed to be down before the starting node makes itself available as a coordinator. Of course, we would still have the timeout to limit startup time and deal with really extreme situations (whole datacenters down etc).\r\n\r\nI started working on a patch for this change [on github|https://github.com/jasobrown/cassandra/compare/13993...jolynch:13993], and am happy to finish it up with unit tests and such if someone can review/commit it (maybe [~aweisberg]?).\r\n\r\nI think the short version of my proposal is we replace:\r\n{noformat}\r\nblock_for_peers_percentage: \r\n{noformat}\r\n\r\nwith either\r\n{noformat}\r\nblock_for_peers: \r\n{noformat}\r\n\r\nor, if we want to do even better imo and enable advanced operators to finely tune this behavior (while still having good defaults that work for almost everyone):\r\n{noformat}\r\nblock_for_peers_local_dc: \r\nblock_for_peers_each_dc: \r\nblock_for_peers_all_dcs: \r\n{noformat}\r\n\r\nFor example if an operator knows that they must be available at {{LOCAL_QUORUM}} they would set {{block_for_peers_local_dc=1}}, if they use {{EACH_QUOURM}} they would set {{block_for_peers_local_dc=1}}, if they use {{QUORUM}} (RF=3, dcs=2) they would set {{block_for_peers_all_dcs=2}}. Naturally everything would of course have a timeout to prevent startup taking too long.\r\n","from":"reporter","subject":"Startup checker should wait for count rather than percentage"},{"body":"First implementation up on github along with a lot of unit tests. I'll start doing some more e2e testing using ccm just to make sure all the edge cases are covered but if someone ([~aweisberg] or [~jasobrown] perhaps) wants to review that would be excellent.\r\n||trunk||\r\n|[pull request|https://github.com/apache/cassandra/pull/212]|\r\n|[!https://circleci.com/gh/jolynch/cassandra/tree/CASSANDRA-14297.png?circle-token= 1102a59698d04899ec971dd36e925928f7b521f5!|https://circleci.com/gh/jolynch/cassandra/tree/CASSANDRA-14297] |\r\n\r\n ","from":"developer"},{"body":"LGTM. +1","from":"developer"},{"body":"Looks like CASSANDRA-14447 refactoring has given me a nasty merge conflict. I'll work on rebasing the patchset but if a someone has time to give me quick feedback on if this idea is mergeable I'd appreciate it before investing more time into it. I do think this change makes the connectivity checker feature much more useful for operators trying to restart their databases without dropping traffic.","from":"developer"},{"body":"I updated the patch to fix the merge conflicts, and reduced it two just two options to make life easier (the default is tuned to wait for all but 2 local DC nodes and not care about non local DC == local_quorum).\r\n\r\nThis is ready for review. I hope we get it in before 4.0 because the user interface of a percentage is something users can't reliably set (compared to the count where there are correct answers for each use case).","from":"developer"},{"body":"I'm changing this to a bug since I think the current user interface is not possible for users to correctly configure and I hope we don't ship 4.0 with the percentage option instead of a count. If someone thinks that there are plausible settings of the existing configuration options users can use we can change this back to an improvement.","from":"developer"},{"body":"Alright, per the discussion on [IRC|https://wilderness.apache.org/channels/?f=cassandra-dev/2018-10-17#1539793033] with Ariel and Jason, we've decided that instead of counts we should always wait for all but a single local DC node and replace the percentage option with:\r\n{noformat}\r\nblock_for_remote_dcs: \r\n{noformat}\r\nThe startup connectivity checker will wait for all but a single node in the local datacenter, and if you want to block startup on every datacenter having only a single node down you can set this to true.\r\n\r\nThe timeout will be the fallback for when multiple nodes are down in a local DC.","from":"developer"},{"body":"Ok I've uploaded a patch to my branch that does what was asked in IRC I believe. Let me know if it looks good and I can run dtests and such against it.","from":"developer"},{"body":"I left some comments on the PR. Looks good.","from":"developer"},{"body":"[~aweisberg] Awesome I just rebased and merged in your suggestions:\r\n\r\n||trunk||\r\n|[fdd8173f|https://github.com/apache/cassandra/pull/212/commits/fdd8173f]|\r\n|[!https://circleci.com/gh/jolynch/cassandra/tree/CASSANDRA-14297.png?circle-token= 1102a59698d04899ec971dd36e925928f7b521f5!|https://circleci.com/gh/jolynch/cassandra/tree/CASSANDRA-14297] |\r\n\r\nDtests are running.","from":"developer"},{"body":"+1 🚢 it.\r\n\r\nThere is one unused import in [StartupClusterConnectivityCheckerTest.java|https://github.com/apache/cassandra/pull/212/files#diff-c74adeeae072ee4af35c12a157cd7d61L26] I'll fix on commit.","from":"developer"},{"body":"Committed as [801cb70ee811c956e987718a00695638d5bec1b6|https://github.com/apache/cassandra/commit/801cb70ee811c956e987718a00695638d5bec1b6] thanks!\r\n\r\nI also added a NEWS.txt and CHANGES.txt entries. I also added Patchy by XYZ; Reviewed by XYZ for CASSANDRA-14297 to the commit message.","from":"developer"},{"body":"Sweet, thanks! Yea I was holding off on adding the NEWs/CHANGES entries until you marked it ready to commit. In the future I'll include it with the dtest run.\r\n\r\nThanks for all the great feedback, I think this feature is much more valuable now to users.","from":"developer"}],"created":"2018-03-07T22:52:16.000+0000","description":"As I commented in CASSANDRA-13993, the current wait for functionality is a great step in the right direction, but I don't think that the current setting (70% of nodes in the cluster) is the right configuration option. First I think this because 70% will not protect against errors as if you wait for 70% of the cluster you could still very easily have {{UnavailableException}} or {{ReadTimeoutException}} exceptions. This is because if you have even two nodes down in different racks in a Cassandra cluster these exceptions are possible (or with the default {{num_tokens}} setting of 256 it is basically guaranteed). Second I think this option is not easy for operators to set, the only setting I could think of that would \"just work\" is 100%.\r\n\r\nI proposed in that ticket instead of having `block_for_peers_percentage` defaulting to 70%, we instead have `block_for_peers` as a count of nodes that are allowed to be down before the starting node makes itself available as a coordinator. Of course, we would still have the timeout to limit startup time and deal with really extreme situations (whole datacenters down etc).\r\n\r\nI started working on a patch for this change [on github|https://github.com/jasobrown/cassandra/compare/13993...jolynch:13993], and am happy to finish it up with unit tests and such if someone can review/commit it (maybe [~aweisberg]?).\r\n\r\nI think the short version of my proposal is we replace:\r\n{noformat}\r\nblock_for_peers_percentage: \r\n{noformat}\r\n\r\nwith either\r\n{noformat}\r\nblock_for_peers: \r\n{noformat}\r\n\r\nor, if we want to do even better imo and enable advanced operators to finely tune this behavior (while still having good defaults that work for almost everyone):\r\n{noformat}\r\nblock_for_peers_local_dc: \r\nblock_for_peers_each_dc: \r\nblock_for_peers_all_dcs: \r\n{noformat}\r\n\r\nFor example if an operator knows that they must be available at {{LOCAL_QUORUM}} they would set {{block_for_peers_local_dc=1}}, if they use {{EACH_QUOURM}} they would set {{block_for_peers_local_dc=1}}, if they use {{QUORUM}} (RF=3, dcs=2) they would set {{block_for_peers_all_dcs=2}}. Naturally everything would of course have a timeout to prevent startup taking too long.\r\n","issue_id":"13143384","key":"CASSANDRA-14297","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2018-11-12T17:47:16.000+0000","role":"fixed_distractor","summary":"Startup checker should wait for count rather than percentage"} {"case_id":"13145729","cluster":"DISTRACTOR-CASSANDRA-14318","comments":[{"body":"Here's the [patch for 2.2|https://github.com/apache/cassandra/compare/cassandra-2.2...thelastpickle:disable-debug-logging-by-default]\r\n\r\nIt should be mergeable in 3.0/3.11/4.0 without a problem.","created":"2018-03-16T14:31:52.895+0000"},{"body":"+1 on your perf test results. Patch LGTM.","created":"2018-03-16T20:59:10.357+0000"},{"body":"CASSANDRA-10241 created that debug.log to be a “production” debug log. If there are things being logged at DEBUG which cause performance issues we should disable those or move them to TRACE, not turn off the debug.log.","created":"2018-03-17T18:21:00.714+0000"},{"body":"Fair enough, I'll go down the read and write path to see if there's debug logging in v3 and above, and switch debug to trace in the query pager in 2.2.","created":"2018-03-17T19:18:41.455+0000"},{"body":"Hey Alexander, would you care to re-run the experiments with the query pager DEBUG logs changed to TRACE level to check if debug.log logging overhead will still persist? Thanks!","created":"2018-03-18T19:25:08.526+0000"},{"body":"[~pauloricardomg], thanks for assigning the ticket to me.\r\n\r\nWhatever the implementation, I'll run new performance tests and generate flame graphs to verify the impact of the changes, no worries.\r\n\r\nI'll wait a bit for a consensus to come out of the discussion on the dev ML before moving on, as there seem to be some conflicting views on what should be done.","created":"2018-03-19T13:22:18.675+0000"},{"body":"No matter what is decided we should move those log messages to TRACE. So I think we can proceed here no matter what.","created":"2018-03-19T13:41:52.791+0000"},{"body":"Fair point. I'll write the patch and run the benchmarks.","created":"2018-03-19T16:10:27.898+0000"},{"body":"If you havent yet started the second run of benchmarks, you may also want to move the debug log statements from ReadCallback.java (one deals with digest mismatch exceptions, forgot what the other is) to trace as well. ","created":"2018-03-19T23:21:48.600+0000"},{"body":"[~jjirsa]: apparently the ReadCallback class already logs at TRACE and not DEBUG on the latest 2.2.\r\n\r\nI've created the fix that downgrades debug logging to trace logging in the query pager classes, and here are the results : \r\n\r\ndebug on - no fix :\r\n{noformat}\r\nResults:\r\nop rate : 6681 [read_event_1:1109, read_event_2:1119, read_event_3:4452]\r\npartition rate : 6681 [read_event_1:1109, read_event_2:1119, read_event_3:4452]\r\nrow rate : 6681 [read_event_1:1109, read_event_2:1119, read_event_3:4452]\r\nlatency mean : 19,1 [read_event_1:15,4, read_event_2:15,4, read_event_3:21,0]\r\nlatency median : 15,6 [read_event_1:14,2, read_event_2:14,0, read_event_3:16,3]\r\nlatency 95th percentile : 39,1 [read_event_1:28,4, read_event_2:28,6, read_event_3:44,2]\r\nlatency 99th percentile : 75,6 [read_event_1:52,9, read_event_2:53,6, read_event_3:87,7]\r\nlatency 99.9th percentile : 315,7 [read_event_1:101,0, read_event_2:110,1, read_event_3:361,1]\r\nlatency max : 609,1 [read_event_1:319,6, read_event_2:315,9, read_event_3:609,1]\r\nTotal partitions : 993050 [read_event_1:164882, read_event_2:166381, read_event_3:661787]\r\nTotal errors : 0 [read_event_1:0, read_event_2:0, read_event_3:0]\r\ntotal gc count : 189\r\ntotal gc mb : 56464\r\ntotal gc time (s) : 7\r\navg gc time(ms) : 37\r\nstdev gc time(ms) : 8\r\nTotal operation time : 00:02:28{noformat}\r\n \r\n\r\n \r\n\r\ndebug off - no fix :\r\n{noformat}\r\nResults:\r\nop rate : 12655 [read_event_1:2141, read_event_2:2093, read_event_3:8422]\r\npartition rate : 12655 [read_event_1:2141, read_event_2:2093, read_event_3:8422]\r\nrow rate : 12655 [read_event_1:2141, read_event_2:2093, read_event_3:8422]\r\nlatency mean : 10,1 [read_event_1:10,1, read_event_2:10,1, read_event_3:10,1]\r\nlatency median : 9,2 [read_event_1:9,2, read_event_2:9,2, read_event_3:9,3]\r\nlatency 95th percentile : 15,2 [read_event_1:15,8, read_event_2:15,9, read_event_3:15,7]\r\nlatency 99th percentile : 29,3 [read_event_1:44,5, read_event_2:45,1, read_event_3:41,3]\r\nlatency 99.9th percentile : 52,7 [read_event_1:67,9, read_event_2:66,9, read_event_3:67,1]\r\nlatency max : 268,0 [read_event_1:257,1, read_event_2:263,3, read_event_3:268,0]\r\nTotal partitions : 983056 [read_event_1:166311, read_event_2:162570, read_event_3:654175]\r\nTotal errors : 0 [read_event_1:0, read_event_2:0, read_event_3:0]\r\ntotal gc count : 100\r\ntotal gc mb : 31529\r\ntotal gc time (s) : 4\r\navg gc time(ms) : 37\r\nstdev gc time(ms) : 5\r\nTotal operation time : 00:01:17{noformat}\r\n \r\n\r\n \r\n\r\ndebug on - with fix :\r\n{noformat}\r\nResults:\r\nop rate : 12289 [read_event_1:2058, read_event_2:2051, read_event_3:8181]\r\npartition rate : 12289 [read_event_1:2058, read_event_2:2051, read_event_3:8181]\r\nrow rate : 12289 [read_event_1:2058, read_event_2:2051, read_event_3:8181]\r\nlatency mean : 10,4 [read_event_1:10,4, read_event_2:10,4, read_event_3:10,4]\r\nlatency median : 9,4 [read_event_1:9,4, read_event_2:9,4, read_event_3:9,4]\r\nlatency 95th percentile : 16,3 [read_event_1:16,8, read_event_2:17,3, read_event_3:16,2]\r\nlatency 99th percentile : 36,6 [read_event_1:44,3, read_event_2:46,6, read_event_3:37,2]\r\nlatency 99.9th percentile : 62,2 [read_event_1:78,0, read_event_2:77,1, read_event_3:80,8]\r\nlatency max : 251,2 [read_event_1:246,9, read_event_2:249,9, read_event_3:251,2]\r\nTotal partitions : 1000000 [read_event_1:167422, read_event_2:166861, read_event_3:665717]\r\nTotal errors : 0 [read_event_1:0, read_event_2:0, read_event_3:0]\r\ntotal gc count : 102\r\ntotal gc mb : 31843\r\ntotal gc time (s) : 4\r\navg gc time(ms) : 38\r\nstdev gc time(ms) : 6\r\nTotal operation time : 00:01:21{noformat}\r\n \r\n\r\n \r\n\r\nSo we have similar performance with debug logging off and with the fix and debug on.\r\n The difference in throughput is pretty massive as we roughly get *twice the read throughput* with the fix.\r\n\r\nLatencies without the fix and with the fix : \r\n\r\np95 : 35ms -> 16ms\r\n p99 : 75ms -> 36ms\r\n\r\nI've ran all tests several times, alternating with and without the fix to make sure caches were not making a difference, and results were consistent with what's pasted above.\r\n It's been running on a single node using an i3.xlarge instance for Cassandra and another i3.large for running cassandra-stress.\r\n\r\n \r\n\r\n*One pretty interesting thing to note* is that when I tested with the predefined mode of cassandra-stress, no paging occurred and the performance difference was not noticeable. This is due to the fact that the predefined mode generates COMPACT STORAGE tables, which involve a different read path (apparently). I think anyone performing benchmarks for Cassandra changes should be aware that the predefined mode isn't relevant and that a user defined test should be used (maybe we should create one that would be used as standard benchmark). \r\n Here's the one I used : [^cassandra-2.2-debug.yaml]\r\n\r\nWith the following commands for writing : \r\n\r\n \r\n{noformat}\r\n/usr/bin/cassandra-stress user profile=/home/ec2-user/cassandra-2.2-debug.yaml n=1000000 'ops(insert=1)' cl=LOCAL_ONE no-warmup -node 172.31.31.42 -mode native cql3 compression=lz4 -rate threads=128{noformat}\r\n \r\n\r\nAnd for reading : \r\n\r\n \r\n{noformat}\r\n/usr/bin/cassandra-stress user profile=/home/ec2-user/cassandra-2.2-debug.yaml n=1000000 'ops(read_event_1=1,read_event_2=1,read_event_3=4)' cl=LOCAL_ONE no-warmup -node 172.31.31.42 -mode native cql3 compression=lz4 -rate threads=128{noformat}\r\n \r\n\r\n[~pauloricardomg] : here's the [patch|https://github.com/apache/cassandra/compare/cassandra-2.2...thelastpickle:CASSANDRA-14318] if you're willing to review/commit it, and the unit test results in [CircleCI|https://circleci.com/gh/thelastpickle/cassandra/178].","created":"2018-03-27T16:10:33.245+0000"},{"body":"For the record, the same tests on 3.11.2 didn't show any notable performance difference between debug on and off : \r\n\r\nCassandra 3.11.2 debug on : \r\n{noformat}\r\nResults:\r\nOp rate : 18 777 op/s [read_event_1: 3 165 op/s, read_event_2: 3 109 op/s, read_event_3: 12 562 op/s]\r\nPartition rate : 6 215 pk/s [read_event_1: 3 165 pk/s, read_event_2: 3 109 pk/s, read_event_3: 0 pk/s]\r\nRow rate : 6 215 row/s [read_event_1: 3 165 row/s, read_event_2: 3 109 row/s, read_event_3: 0 row/s]\r\nLatency mean : 6,7 ms [read_event_1: 6,7 ms, read_event_2: 6,7 ms, read_event_3: 6,6 ms]\r\nLatency median : 5,0 ms [read_event_1: 5,0 ms, read_event_2: 5,0 ms, read_event_3: 4,9 ms]\r\nLatency 95th percentile : 15,6 ms [read_event_1: 15,5 ms, read_event_2: 15,9 ms, read_event_3: 15,5 ms]\r\nLatency 99th percentile : 43,3 ms [read_event_1: 42,7 ms, read_event_2: 44,2 ms, read_event_3: 43,2 ms]\r\nLatency 99.9th percentile : 82,0 ms [read_event_1: 80,3 ms, read_event_2: 82,4 ms, read_event_3: 82,1 ms]\r\nLatency max : 272,4 ms [read_event_1: 272,4 ms, read_event_2: 268,7 ms, read_event_3: 245,1 ms]\r\nTotal partitions : 330 970 [read_event_1: 165 386, read_event_2: 165 584, read_event_3: 0]\r\nTotal errors : 0 [read_event_1: 0, read_event_2: 0, read_event_3: 0]\r\nTotal GC count : 42\r\nTotal GC memory : 13,102 GiB\r\nTotal GC time : 1,8 seconds\r\nAvg GC time : 42,4 ms\r\nStdDev GC time : 1,3 ms\r\nTotal operation time : 00:00:53{noformat}\r\n \r\n\r\n\r\nCassandra 3.11.2 debug off : \r\n{noformat}\r\nResults:\r\nOp rate : 18 853 op/s [read_event_1: 3 138 op/s, read_event_2: 3 137 op/s, read_event_3: 12 578 op/s]\r\nPartition rate : 6 275 pk/s [read_event_1: 3 138 pk/s, read_event_2: 3 137 pk/s, read_event_3: 0 pk/s]\r\nRow rate : 6 275 row/s [read_event_1: 3 138 row/s, read_event_2: 3 137 row/s, read_event_3: 0 row/s]\r\nLatency mean : 6,7 ms [read_event_1: 6,7 ms, read_event_2: 6,7 ms, read_event_3: 6,7 ms]\r\nLatency median : 5,0 ms [read_event_1: 5,1 ms, read_event_2: 5,1 ms, read_event_3: 5,0 ms]\r\nLatency 95th percentile : 15,5 ms [read_event_1: 15,5 ms, read_event_2: 15,6 ms, read_event_3: 15,4 ms]\r\nLatency 99th percentile : 39,9 ms [read_event_1: 41,0 ms, read_event_2: 39,6 ms, read_event_3: 39,6 ms]\r\nLatency 99.9th percentile : 73,3 ms [read_event_1: 73,4 ms, read_event_2: 71,6 ms, read_event_3: 73,6 ms]\r\nLatency max : 367,0 ms [read_event_1: 240,5 ms, read_event_2: 250,3 ms, read_event_3: 367,0 ms]\r\nTotal partitions : 332 852 [read_event_1: 166 447, read_event_2: 166 405, read_event_3: 0]\r\nTotal errors : 0 [read_event_1: 0, read_event_2: 0, read_event_3: 0]\r\nTotal GC count : 46\r\nTotal GC memory : 14,024 GiB\r\nTotal GC time : 2,0 seconds\r\nAvg GC time : 42,7 ms\r\nStdDev GC time : 3,9 ms\r\nTotal operation time : 00:00:53{noformat}\r\nThe improvement over 2.2 is nice though :)\r\n\r\n ","created":"2018-03-27T16:57:43.191+0000"},{"body":"{quote}I think anyone performing benchmarks for Cassandra changes should be aware that the predefined mode isn't relevant and that a user defined test should be used (maybe we should create one that would be used as standard benchmark).\r\n{quote}\r\nGood find! Can you check if this is the case in trunk, and if so maybe open a lhf ticket to change that?\r\n{quote}For the record, the same tests on 3.11.2 didn't show any notable performance difference between debug on and off\r\n{quote}\r\nNice to know we managed to handle all debug/verbose log leaks there. It will be easier to maintain this after CASSANDRA-14326.\r\n{quote}here's the patch if you're willing to review/commit it, and the unit test results in CircleCI.\r\n{quote}\r\nThanks for the patch, experiments and analysis! Even though 2.2 is on critical fixes only mode, 50% is a significant performance hit on throughput for this workload, and since the patch is pretty simple I don't see a reason not to commit it.\r\n\r\nCI looks good. I added a CHANGES.txt not and committed as {{ac77e5e7742548f7c7c25da3923841f59d4b2713}} to cassandra-2.2 branch.","created":"2018-03-30T15:18:23.848+0000"},{"body":"Thanks for reviewing and merging [~pauloricardomg] !\r\n\r\nCASSANDRA-10857 removed compact storage options in trunk and the standard1&counter1 tables are no longer using it : [https://github.com/apache/cassandra/commit/07fbd8ee6042797aaade90357d625ba9d79c31e0#diff-e5d5cb263c5c84c322cd09391af46d7dL141] ","created":"2018-04-03T12:45:59.149+0000"}],"conversations":[{"body":"Debug logging can involve in many cases (especially very low latency ones) a very important overhead on the read path in 2.2 as we've seen when upgrading clusters from 2.0 to 2.2.\r\n\r\nThe performance impact was especially noticeable on the client side metrics, where p99 could go up to 10 times higher, while ClientRequest metrics recorded by Cassandra didn't show any overhead.\r\n\r\nBelow shows latencies recorded on the client side with debug logging on first, and then without it :\r\n\r\n!debuglogging.png!  \r\n\r\nWe generated a flame graph before turning off debug logging that shows the read call stack is dominated by debug logging : \r\n\r\n!flame_graph_snapshot.png!\r\n\r\nI've attached the original flame graph for exploration.\r\n\r\nOnce disabled, the new flame graph shows that the read call stack gets extremely thin, which is further confirmed by client recorded metrics : \r\n\r\n!flame22 nodebug sjk svg.png!\r\n\r\nThe query pager code has been reworked since 3.0 and it looks like log.debug() calls are gone there, but for 2.2 users and to prevent such issues to appear with default settings, I really think debug logging should be disabled by default.","from":"reporter","subject":"Fix query pager DEBUG log leak causing hit in paged reads throughput"},{"body":"Here's the [patch for 2.2|https://github.com/apache/cassandra/compare/cassandra-2.2...thelastpickle:disable-debug-logging-by-default]\r\n\r\nIt should be mergeable in 3.0/3.11/4.0 without a problem.","from":"developer"},{"body":"+1 on your perf test results. Patch LGTM.","from":"developer"},{"body":"CASSANDRA-10241 created that debug.log to be a “production” debug log. If there are things being logged at DEBUG which cause performance issues we should disable those or move them to TRACE, not turn off the debug.log.","from":"developer"},{"body":"Fair enough, I'll go down the read and write path to see if there's debug logging in v3 and above, and switch debug to trace in the query pager in 2.2.","from":"developer"},{"body":"Hey Alexander, would you care to re-run the experiments with the query pager DEBUG logs changed to TRACE level to check if debug.log logging overhead will still persist? Thanks!","from":"developer"},{"body":"[~pauloricardomg], thanks for assigning the ticket to me.\r\n\r\nWhatever the implementation, I'll run new performance tests and generate flame graphs to verify the impact of the changes, no worries.\r\n\r\nI'll wait a bit for a consensus to come out of the discussion on the dev ML before moving on, as there seem to be some conflicting views on what should be done.","from":"developer"},{"body":"No matter what is decided we should move those log messages to TRACE. So I think we can proceed here no matter what.","from":"developer"},{"body":"Fair point. I'll write the patch and run the benchmarks.","from":"developer"},{"body":"If you havent yet started the second run of benchmarks, you may also want to move the debug log statements from ReadCallback.java (one deals with digest mismatch exceptions, forgot what the other is) to trace as well. ","from":"developer"},{"body":"[~jjirsa]: apparently the ReadCallback class already logs at TRACE and not DEBUG on the latest 2.2.\r\n\r\nI've created the fix that downgrades debug logging to trace logging in the query pager classes, and here are the results : \r\n\r\ndebug on - no fix :\r\n{noformat}\r\nResults:\r\nop rate : 6681 [read_event_1:1109, read_event_2:1119, read_event_3:4452]\r\npartition rate : 6681 [read_event_1:1109, read_event_2:1119, read_event_3:4452]\r\nrow rate : 6681 [read_event_1:1109, read_event_2:1119, read_event_3:4452]\r\nlatency mean : 19,1 [read_event_1:15,4, read_event_2:15,4, read_event_3:21,0]\r\nlatency median : 15,6 [read_event_1:14,2, read_event_2:14,0, read_event_3:16,3]\r\nlatency 95th percentile : 39,1 [read_event_1:28,4, read_event_2:28,6, read_event_3:44,2]\r\nlatency 99th percentile : 75,6 [read_event_1:52,9, read_event_2:53,6, read_event_3:87,7]\r\nlatency 99.9th percentile : 315,7 [read_event_1:101,0, read_event_2:110,1, read_event_3:361,1]\r\nlatency max : 609,1 [read_event_1:319,6, read_event_2:315,9, read_event_3:609,1]\r\nTotal partitions : 993050 [read_event_1:164882, read_event_2:166381, read_event_3:661787]\r\nTotal errors : 0 [read_event_1:0, read_event_2:0, read_event_3:0]\r\ntotal gc count : 189\r\ntotal gc mb : 56464\r\ntotal gc time (s) : 7\r\navg gc time(ms) : 37\r\nstdev gc time(ms) : 8\r\nTotal operation time : 00:02:28{noformat}\r\n \r\n\r\n \r\n\r\ndebug off - no fix :\r\n{noformat}\r\nResults:\r\nop rate : 12655 [read_event_1:2141, read_event_2:2093, read_event_3:8422]\r\npartition rate : 12655 [read_event_1:2141, read_event_2:2093, read_event_3:8422]\r\nrow rate : 12655 [read_event_1:2141, read_event_2:2093, read_event_3:8422]\r\nlatency mean : 10,1 [read_event_1:10,1, read_event_2:10,1, read_event_3:10,1]\r\nlatency median : 9,2 [read_event_1:9,2, read_event_2:9,2, read_event_3:9,3]\r\nlatency 95th percentile : 15,2 [read_event_1:15,8, read_event_2:15,9, read_event_3:15,7]\r\nlatency 99th percentile : 29,3 [read_event_1:44,5, read_event_2:45,1, read_event_3:41,3]\r\nlatency 99.9th percentile : 52,7 [read_event_1:67,9, read_event_2:66,9, read_event_3:67,1]\r\nlatency max : 268,0 [read_event_1:257,1, read_event_2:263,3, read_event_3:268,0]\r\nTotal partitions : 983056 [read_event_1:166311, read_event_2:162570, read_event_3:654175]\r\nTotal errors : 0 [read_event_1:0, read_event_2:0, read_event_3:0]\r\ntotal gc count : 100\r\ntotal gc mb : 31529\r\ntotal gc time (s) : 4\r\navg gc time(ms) : 37\r\nstdev gc time(ms) : 5\r\nTotal operation time : 00:01:17{noformat}\r\n \r\n\r\n \r\n\r\ndebug on - with fix :\r\n{noformat}\r\nResults:\r\nop rate : 12289 [read_event_1:2058, read_event_2:2051, read_event_3:8181]\r\npartition rate : 12289 [read_event_1:2058, read_event_2:2051, read_event_3:8181]\r\nrow rate : 12289 [read_event_1:2058, read_event_2:2051, read_event_3:8181]\r\nlatency mean : 10,4 [read_event_1:10,4, read_event_2:10,4, read_event_3:10,4]\r\nlatency median : 9,4 [read_event_1:9,4, read_event_2:9,4, read_event_3:9,4]\r\nlatency 95th percentile : 16,3 [read_event_1:16,8, read_event_2:17,3, read_event_3:16,2]\r\nlatency 99th percentile : 36,6 [read_event_1:44,3, read_event_2:46,6, read_event_3:37,2]\r\nlatency 99.9th percentile : 62,2 [read_event_1:78,0, read_event_2:77,1, read_event_3:80,8]\r\nlatency max : 251,2 [read_event_1:246,9, read_event_2:249,9, read_event_3:251,2]\r\nTotal partitions : 1000000 [read_event_1:167422, read_event_2:166861, read_event_3:665717]\r\nTotal errors : 0 [read_event_1:0, read_event_2:0, read_event_3:0]\r\ntotal gc count : 102\r\ntotal gc mb : 31843\r\ntotal gc time (s) : 4\r\navg gc time(ms) : 38\r\nstdev gc time(ms) : 6\r\nTotal operation time : 00:01:21{noformat}\r\n \r\n\r\n \r\n\r\nSo we have similar performance with debug logging off and with the fix and debug on.\r\n The difference in throughput is pretty massive as we roughly get *twice the read throughput* with the fix.\r\n\r\nLatencies without the fix and with the fix : \r\n\r\np95 : 35ms -> 16ms\r\n p99 : 75ms -> 36ms\r\n\r\nI've ran all tests several times, alternating with and without the fix to make sure caches were not making a difference, and results were consistent with what's pasted above.\r\n It's been running on a single node using an i3.xlarge instance for Cassandra and another i3.large for running cassandra-stress.\r\n\r\n \r\n\r\n*One pretty interesting thing to note* is that when I tested with the predefined mode of cassandra-stress, no paging occurred and the performance difference was not noticeable. This is due to the fact that the predefined mode generates COMPACT STORAGE tables, which involve a different read path (apparently). I think anyone performing benchmarks for Cassandra changes should be aware that the predefined mode isn't relevant and that a user defined test should be used (maybe we should create one that would be used as standard benchmark). \r\n Here's the one I used : [^cassandra-2.2-debug.yaml]\r\n\r\nWith the following commands for writing : \r\n\r\n \r\n{noformat}\r\n/usr/bin/cassandra-stress user profile=/home/ec2-user/cassandra-2.2-debug.yaml n=1000000 'ops(insert=1)' cl=LOCAL_ONE no-warmup -node 172.31.31.42 -mode native cql3 compression=lz4 -rate threads=128{noformat}\r\n \r\n\r\nAnd for reading : \r\n\r\n \r\n{noformat}\r\n/usr/bin/cassandra-stress user profile=/home/ec2-user/cassandra-2.2-debug.yaml n=1000000 'ops(read_event_1=1,read_event_2=1,read_event_3=4)' cl=LOCAL_ONE no-warmup -node 172.31.31.42 -mode native cql3 compression=lz4 -rate threads=128{noformat}\r\n \r\n\r\n[~pauloricardomg] : here's the [patch|https://github.com/apache/cassandra/compare/cassandra-2.2...thelastpickle:CASSANDRA-14318] if you're willing to review/commit it, and the unit test results in [CircleCI|https://circleci.com/gh/thelastpickle/cassandra/178].","from":"developer"},{"body":"For the record, the same tests on 3.11.2 didn't show any notable performance difference between debug on and off : \r\n\r\nCassandra 3.11.2 debug on : \r\n{noformat}\r\nResults:\r\nOp rate : 18 777 op/s [read_event_1: 3 165 op/s, read_event_2: 3 109 op/s, read_event_3: 12 562 op/s]\r\nPartition rate : 6 215 pk/s [read_event_1: 3 165 pk/s, read_event_2: 3 109 pk/s, read_event_3: 0 pk/s]\r\nRow rate : 6 215 row/s [read_event_1: 3 165 row/s, read_event_2: 3 109 row/s, read_event_3: 0 row/s]\r\nLatency mean : 6,7 ms [read_event_1: 6,7 ms, read_event_2: 6,7 ms, read_event_3: 6,6 ms]\r\nLatency median : 5,0 ms [read_event_1: 5,0 ms, read_event_2: 5,0 ms, read_event_3: 4,9 ms]\r\nLatency 95th percentile : 15,6 ms [read_event_1: 15,5 ms, read_event_2: 15,9 ms, read_event_3: 15,5 ms]\r\nLatency 99th percentile : 43,3 ms [read_event_1: 42,7 ms, read_event_2: 44,2 ms, read_event_3: 43,2 ms]\r\nLatency 99.9th percentile : 82,0 ms [read_event_1: 80,3 ms, read_event_2: 82,4 ms, read_event_3: 82,1 ms]\r\nLatency max : 272,4 ms [read_event_1: 272,4 ms, read_event_2: 268,7 ms, read_event_3: 245,1 ms]\r\nTotal partitions : 330 970 [read_event_1: 165 386, read_event_2: 165 584, read_event_3: 0]\r\nTotal errors : 0 [read_event_1: 0, read_event_2: 0, read_event_3: 0]\r\nTotal GC count : 42\r\nTotal GC memory : 13,102 GiB\r\nTotal GC time : 1,8 seconds\r\nAvg GC time : 42,4 ms\r\nStdDev GC time : 1,3 ms\r\nTotal operation time : 00:00:53{noformat}\r\n \r\n\r\n\r\nCassandra 3.11.2 debug off : \r\n{noformat}\r\nResults:\r\nOp rate : 18 853 op/s [read_event_1: 3 138 op/s, read_event_2: 3 137 op/s, read_event_3: 12 578 op/s]\r\nPartition rate : 6 275 pk/s [read_event_1: 3 138 pk/s, read_event_2: 3 137 pk/s, read_event_3: 0 pk/s]\r\nRow rate : 6 275 row/s [read_event_1: 3 138 row/s, read_event_2: 3 137 row/s, read_event_3: 0 row/s]\r\nLatency mean : 6,7 ms [read_event_1: 6,7 ms, read_event_2: 6,7 ms, read_event_3: 6,7 ms]\r\nLatency median : 5,0 ms [read_event_1: 5,1 ms, read_event_2: 5,1 ms, read_event_3: 5,0 ms]\r\nLatency 95th percentile : 15,5 ms [read_event_1: 15,5 ms, read_event_2: 15,6 ms, read_event_3: 15,4 ms]\r\nLatency 99th percentile : 39,9 ms [read_event_1: 41,0 ms, read_event_2: 39,6 ms, read_event_3: 39,6 ms]\r\nLatency 99.9th percentile : 73,3 ms [read_event_1: 73,4 ms, read_event_2: 71,6 ms, read_event_3: 73,6 ms]\r\nLatency max : 367,0 ms [read_event_1: 240,5 ms, read_event_2: 250,3 ms, read_event_3: 367,0 ms]\r\nTotal partitions : 332 852 [read_event_1: 166 447, read_event_2: 166 405, read_event_3: 0]\r\nTotal errors : 0 [read_event_1: 0, read_event_2: 0, read_event_3: 0]\r\nTotal GC count : 46\r\nTotal GC memory : 14,024 GiB\r\nTotal GC time : 2,0 seconds\r\nAvg GC time : 42,7 ms\r\nStdDev GC time : 3,9 ms\r\nTotal operation time : 00:00:53{noformat}\r\nThe improvement over 2.2 is nice though :)\r\n\r\n ","from":"developer"},{"body":"{quote}I think anyone performing benchmarks for Cassandra changes should be aware that the predefined mode isn't relevant and that a user defined test should be used (maybe we should create one that would be used as standard benchmark).\r\n{quote}\r\nGood find! Can you check if this is the case in trunk, and if so maybe open a lhf ticket to change that?\r\n{quote}For the record, the same tests on 3.11.2 didn't show any notable performance difference between debug on and off\r\n{quote}\r\nNice to know we managed to handle all debug/verbose log leaks there. It will be easier to maintain this after CASSANDRA-14326.\r\n{quote}here's the patch if you're willing to review/commit it, and the unit test results in CircleCI.\r\n{quote}\r\nThanks for the patch, experiments and analysis! Even though 2.2 is on critical fixes only mode, 50% is a significant performance hit on throughput for this workload, and since the patch is pretty simple I don't see a reason not to commit it.\r\n\r\nCI looks good. I added a CHANGES.txt not and committed as {{ac77e5e7742548f7c7c25da3923841f59d4b2713}} to cassandra-2.2 branch.","from":"developer"},{"body":"Thanks for reviewing and merging [~pauloricardomg] !\r\n\r\nCASSANDRA-10857 removed compact storage options in trunk and the standard1&counter1 tables are no longer using it : [https://github.com/apache/cassandra/commit/07fbd8ee6042797aaade90357d625ba9d79c31e0#diff-e5d5cb263c5c84c322cd09391af46d7dL141] ","from":"developer"}],"created":"2018-03-16T14:25:43.000+0000","description":"Debug logging can involve in many cases (especially very low latency ones) a very important overhead on the read path in 2.2 as we've seen when upgrading clusters from 2.0 to 2.2.\r\n\r\nThe performance impact was especially noticeable on the client side metrics, where p99 could go up to 10 times higher, while ClientRequest metrics recorded by Cassandra didn't show any overhead.\r\n\r\nBelow shows latencies recorded on the client side with debug logging on first, and then without it :\r\n\r\n!debuglogging.png!  \r\n\r\nWe generated a flame graph before turning off debug logging that shows the read call stack is dominated by debug logging : \r\n\r\n!flame_graph_snapshot.png!\r\n\r\nI've attached the original flame graph for exploration.\r\n\r\nOnce disabled, the new flame graph shows that the read call stack gets extremely thin, which is further confirmed by client recorded metrics : \r\n\r\n!flame22 nodebug sjk svg.png!\r\n\r\nThe query pager code has been reworked since 3.0 and it looks like log.debug() calls are gone there, but for 2.2 users and to prevent such issues to appear with default settings, I really think debug logging should be disabled by default.","issue_id":"13145729","key":"CASSANDRA-14318","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2018-03-30T15:18:23.000+0000","role":"fixed_distractor","summary":"Fix query pager DEBUG log leak causing hit in paged reads throughput"} {"case_id":"13162924","cluster":"DISTRACTOR-CASSANDRA-14481","comments":[{"body":"This is my first \"PR\" to this project. I read the contributions document, but please let me know if I've missed any required tags or otherwise need to edit my submission. Thank you very much.","created":"2018-05-30T14:54:09.015+0000"},{"body":"Thanks for your contribution.\r\n\r\nI've verified the issue and the proposed fix on 3.11.2 and trunk.\r\n\r\nSome comments on your proposed changes.\r\n{quote}GRANT EXECUTE ON MBEAN ‘org.apache.cassandra.db:type=EndpointSnitchInfo’ TO jmx;\r\n GRANT SELECT, EXECUTE ON MBEAN ‘org.apache.cassandra.db:type=StorageService’ TO jmx;\r\n{quote}\r\nPlease update your patch such that it will contain relevant lines only. Right now you're duplicating parts of the example. \r\n\r\nMake sure to use a straight quotes {{'}}, not curved quotes {{’}} around the mbean names.\r\n\r\nThere is no need to grant both {{SELECT}} and {{EXECUTE}} on the {{StorageService}} as {{SELECT}} is granted to {{ALL MBEANS}} already in the example. And CQL don't let you grant two permissions in one statement anyway.\r\n\r\nFor small fixes like this, committers seem to prefer to get proposed fixes [like this|https://cassandra.apache.org/doc/latest/development/documentation.html#github-based-work-flow]. It is not clear to me how documentation is maintained and published on different branches/versions of Cassandra, but perhaps someone else can give advice on that.","created":"2018-06-20T10:00:17.834+0000"},{"body":"Thank you very much for this information, [~eperott]. I will edit the PR as you have suggested and resubmit.","created":"2018-06-20T13:58:33.711+0000"},{"body":"I've edited the file to support this PR in the following ways:\r\n\r\n- changed curly quotes to flat quotes\r\n- removed repeated lines\r\n- moved changes to a branch called docs_operating_security","created":"2018-06-20T18:55:16.661+0000"},{"body":"Thanks! Committed as {{4f02db5c45ece38dda48e0d19667888e9f46536e}} with [~eperott] as reviewer. The site won't update immediately, but will take effect on next rebuild.\r\n","created":"2018-06-21T03:08:56.706+0000"}],"conversations":[{"body":"Using the documentation here:\r\n\r\n[https://cassandra.apache.org/doc/latest/operating/security.html#cassandra-integrated-auth]\r\n\r\nRunning `nodetool status` on a cluster fails as follows:\r\n{noformat}\r\nerror: Access Denied\r\n-- StackTrace --\r\njava.lang.SecurityException: Access Denied\r\nat org.apache.cassandra.auth.jmx.AuthorizationProxy.invoke(AuthorizationProxy.java:172)\r\nat com.sun.proxy.$Proxy4.invoke(Unknown Source)\r\nat javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1468)\r\nat javax.management.remote.rmi.RMIConnectionImpl.access$300(RMIConnectionImpl.java:76)\r\nat javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1309)\r\nat java.security.AccessController.doPrivileged(Native Method)\r\nat javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1408)\r\nat javax.management.remote.rmi.RMIConnectionImpl.invoke(RMIConnectionImpl.java:829)\r\nat sun.reflect.GeneratedMethodAccessor24.invoke(Unknown Source)\r\nat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\nat java.lang.reflect.Method.invoke(Method.java:498)\r\nat sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:357)\r\nat sun.rmi.transport.Transport$1.run(Transport.java:200)\r\nat sun.rmi.transport.Transport$1.run(Transport.java:197)\r\nat java.security.AccessController.doPrivileged(Native Method)\r\nat sun.rmi.transport.Transport.serviceCall(Transport.java:196)\r\nat sun.rmi.transport.tcp.TCPTransport.handleMessages(TCPTransport.java:573)\r\nat sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run0(TCPTransport.java:835)\r\nat sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.lambda$run$0(TCPTransport.java:688)\r\nat java.security.AccessController.doPrivileged(Native Method)\r\nat sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run(TCPTransport.java:687)\r\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\nat java.lang.Thread.run(Thread.java:748)\r\nat sun.rmi.transport.StreamRemoteCall.exceptionReceivedFromServer(StreamRemoteCall.java:283)\r\nat sun.rmi.transport.StreamRemoteCall.executeCall(StreamRemoteCall.java:260)\r\nat sun.rmi.server.UnicastRef.invoke(UnicastRef.java:161)\r\nat com.sun.jmx.remote.internal.PRef.invoke(Unknown Source)\r\nat javax.management.remote.rmi.RMIConnectionImpl_Stub.invoke(Unknown Source)\r\nat javax.management.remote.rmi.RMIConnector$RemoteMBeanServerConnection.invoke(RMIConnector.java:1020)\r\nat javax.management.MBeanServerInvocationHandler.invoke(MBeanServerInvocationHandler.java:298)\r\nat com.sun.proxy.$Proxy7.effectiveOwnership(Unknown Source)\r\nat org.apache.cassandra.tools.NodeProbe.effectiveOwnership(NodeProbe.java:489)\r\nat org.apache.cassandra.tools.nodetool.Status.execute(Status.java:74)\r\nat org.apache.cassandra.tools.NodeTool$NodeToolCmd.run(NodeTool.java:255)\r\nat org.apache.cassandra.tools.NodeTool.main(NodeTool.java:169) {noformat}\r\nPermissions on two additional mbeans were required:\r\n{noformat}\r\nGRANT EXECUTE ON MBEAN 'org.apache.cassandra.db:type=StorageService' TO jmx;\r\nGRANT EXECUTE ON MBEAN 'org.apache.cassandra.db:type=EndpointSnitchInfo' TO jmx;\r\n{noformat}\r\nI've updated the documentation in my fork here and would like to do a pull request for the addition:\r\n\r\n[https://github.com/dataindataout/cassandra/blob/docs_operating_security/doc/source/operating/security.rst]\r\n\r\n ","from":"reporter","subject":"Using nodetool status after enabling Cassandra internal auth for JMX access fails with currently documented permissions"},{"body":"This is my first \"PR\" to this project. I read the contributions document, but please let me know if I've missed any required tags or otherwise need to edit my submission. Thank you very much.","from":"developer"},{"body":"Thanks for your contribution.\r\n\r\nI've verified the issue and the proposed fix on 3.11.2 and trunk.\r\n\r\nSome comments on your proposed changes.\r\n{quote}GRANT EXECUTE ON MBEAN ‘org.apache.cassandra.db:type=EndpointSnitchInfo’ TO jmx;\r\n GRANT SELECT, EXECUTE ON MBEAN ‘org.apache.cassandra.db:type=StorageService’ TO jmx;\r\n{quote}\r\nPlease update your patch such that it will contain relevant lines only. Right now you're duplicating parts of the example. \r\n\r\nMake sure to use a straight quotes {{'}}, not curved quotes {{’}} around the mbean names.\r\n\r\nThere is no need to grant both {{SELECT}} and {{EXECUTE}} on the {{StorageService}} as {{SELECT}} is granted to {{ALL MBEANS}} already in the example. And CQL don't let you grant two permissions in one statement anyway.\r\n\r\nFor small fixes like this, committers seem to prefer to get proposed fixes [like this|https://cassandra.apache.org/doc/latest/development/documentation.html#github-based-work-flow]. It is not clear to me how documentation is maintained and published on different branches/versions of Cassandra, but perhaps someone else can give advice on that.","from":"developer"},{"body":"Thank you very much for this information, [~eperott]. I will edit the PR as you have suggested and resubmit.","from":"developer"},{"body":"I've edited the file to support this PR in the following ways:\r\n\r\n- changed curly quotes to flat quotes\r\n- removed repeated lines\r\n- moved changes to a branch called docs_operating_security","from":"developer"},{"body":"Thanks! Committed as {{4f02db5c45ece38dda48e0d19667888e9f46536e}} with [~eperott] as reviewer. The site won't update immediately, but will take effect on next rebuild.\r\n","from":"developer"}],"created":"2018-05-30T14:53:25.000+0000","description":"Using the documentation here:\r\n\r\n[https://cassandra.apache.org/doc/latest/operating/security.html#cassandra-integrated-auth]\r\n\r\nRunning `nodetool status` on a cluster fails as follows:\r\n{noformat}\r\nerror: Access Denied\r\n-- StackTrace --\r\njava.lang.SecurityException: Access Denied\r\nat org.apache.cassandra.auth.jmx.AuthorizationProxy.invoke(AuthorizationProxy.java:172)\r\nat com.sun.proxy.$Proxy4.invoke(Unknown Source)\r\nat javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1468)\r\nat javax.management.remote.rmi.RMIConnectionImpl.access$300(RMIConnectionImpl.java:76)\r\nat javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1309)\r\nat java.security.AccessController.doPrivileged(Native Method)\r\nat javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1408)\r\nat javax.management.remote.rmi.RMIConnectionImpl.invoke(RMIConnectionImpl.java:829)\r\nat sun.reflect.GeneratedMethodAccessor24.invoke(Unknown Source)\r\nat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\nat java.lang.reflect.Method.invoke(Method.java:498)\r\nat sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:357)\r\nat sun.rmi.transport.Transport$1.run(Transport.java:200)\r\nat sun.rmi.transport.Transport$1.run(Transport.java:197)\r\nat java.security.AccessController.doPrivileged(Native Method)\r\nat sun.rmi.transport.Transport.serviceCall(Transport.java:196)\r\nat sun.rmi.transport.tcp.TCPTransport.handleMessages(TCPTransport.java:573)\r\nat sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run0(TCPTransport.java:835)\r\nat sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.lambda$run$0(TCPTransport.java:688)\r\nat java.security.AccessController.doPrivileged(Native Method)\r\nat sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run(TCPTransport.java:687)\r\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\nat java.lang.Thread.run(Thread.java:748)\r\nat sun.rmi.transport.StreamRemoteCall.exceptionReceivedFromServer(StreamRemoteCall.java:283)\r\nat sun.rmi.transport.StreamRemoteCall.executeCall(StreamRemoteCall.java:260)\r\nat sun.rmi.server.UnicastRef.invoke(UnicastRef.java:161)\r\nat com.sun.jmx.remote.internal.PRef.invoke(Unknown Source)\r\nat javax.management.remote.rmi.RMIConnectionImpl_Stub.invoke(Unknown Source)\r\nat javax.management.remote.rmi.RMIConnector$RemoteMBeanServerConnection.invoke(RMIConnector.java:1020)\r\nat javax.management.MBeanServerInvocationHandler.invoke(MBeanServerInvocationHandler.java:298)\r\nat com.sun.proxy.$Proxy7.effectiveOwnership(Unknown Source)\r\nat org.apache.cassandra.tools.NodeProbe.effectiveOwnership(NodeProbe.java:489)\r\nat org.apache.cassandra.tools.nodetool.Status.execute(Status.java:74)\r\nat org.apache.cassandra.tools.NodeTool$NodeToolCmd.run(NodeTool.java:255)\r\nat org.apache.cassandra.tools.NodeTool.main(NodeTool.java:169) {noformat}\r\nPermissions on two additional mbeans were required:\r\n{noformat}\r\nGRANT EXECUTE ON MBEAN 'org.apache.cassandra.db:type=StorageService' TO jmx;\r\nGRANT EXECUTE ON MBEAN 'org.apache.cassandra.db:type=EndpointSnitchInfo' TO jmx;\r\n{noformat}\r\nI've updated the documentation in my fork here and would like to do a pull request for the addition:\r\n\r\n[https://github.com/dataindataout/cassandra/blob/docs_operating_security/doc/source/operating/security.rst]\r\n\r\n ","issue_id":"13162924","key":"CASSANDRA-14481","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2018-06-21T03:08:56.000+0000","role":"fixed_distractor","summary":"Using nodetool status after enabling Cassandra internal auth for JMX access fails with currently documented permissions"} {"case_id":"13185075","cluster":"DISTRACTOR-CASSANDRA-14752","comments":[{"body":"bq. What will happen if the position of these Bytebuffers is being changed by some other operations?\r\n\r\nIn C* you should not change the position unless you have duplicated your `ByteBuffer` first or you really know what you are doing. All the data being stored in the memtables for example are shared so if you change the position on a `ByteBuffer` coming from there you can corrupt the data in a much worst way.\r\n\r\nPersonally, I would not try to change that code as its effect would simply be to cause more garbage.","created":"2018-09-17T07:48:09.668+0000"},{"body":" \r\n\r\n[~blerer] Thanks for your reply. In one of our tool, we use below code to generate the DecoratedKey from String and In the case of boolean type, we are facing this issue.\r\n{code:java}\r\nDatabaseDescriptor.getPartitioner().decorateKey(getKeyValidator(row.getColumnFamily())\r\n.fromString(stringKey));\r\n{code}\r\n\r\n[https://github.com/apache/cassandra/blob/cassandra-2.1.13/src/java/org/apache/cassandra/db/marshal/AbstractCompositeType.java#L255] `byteBuffer.put` changes the position. Though it has a comment: *// it's ok to consume component as we won't use it anymore.*\r\n\r\n ","created":"2018-09-17T08:40:43.081+0000"},{"body":"I found there are too many usages of `AbstractCompositeType#fromString()`. \r\nOne way to corrupt the data:-\r\n\r\nTable schema:-\r\n{code:java}\r\nCREATE TABLE ks1.table1 (\r\nt_id boolean,\r\nid boolean,\r\nck boolean,\r\nnk boolean,\r\nPRIMARY KEY ((t_id,id),ck)\r\n);{code}\r\nInsert statement:-\r\n{code:java}\r\ninsert into ks1.table1 (t_id, ck, id, nk)\r\nVALUES (true, false, false, true);\r\n{code}\r\nNow run nodetool command to get the SSTable for given key:-\r\n{code:java}\r\nbin/nodetool getsstables  ks1 table1 \"false:true\"\r\n{code}\r\nBasically, this operation will modify the positions.\r\n\r\nInsert again:-\r\n{code:java}\r\ninsert into ks1.table1 (t_id, ck, id, nk)\r\nVALUES (true, true, false, true);\r\n{code}\r\nselect data from this table:-\r\n{code:java}\r\ntrue,false,false,true\r\nnull,null,null,null\r\n{code}\r\nSo now all boolean type data will be written as null.","created":"2018-09-17T09:14:56.870+0000"},{"body":" \r\n||MR||\r\n|[trunk\\|https://github.com/Barala/cassandra/commits/CASSANDRA-14752-trunk]|\r\n\r\nI came up with a different approach where BooleanSerializer's static ByteBuffer can be detected using a reference equality check. This approach will avoid new object creations of ByteBuffers.\r\n\r\nI raised MR for trunk. If it passes the review then I'll raise MR to patch other affected versions.","created":"2018-12-12T11:38:21.351+0000"},{"body":"Hey [~varuna], apologize for the very late response.\r\n\r\nI think we can simplify a bit by reducing the patch to changing [this line|https://github.com/Barala/cassandra/commit/2328cda4d4e12e75bc1e9606969f18dd72d4fc27#diff-821957692da5432eb814312397d145fc6445ac35856563b21895e0c1df82ca3aL235]\r\n\r\nto \r\n{code:java}\r\nbb.put(component.duplicate()); // it's not ok to consume component as we did not create it{code}\r\nWhat do you think?\r\n[~blerer] , do you mind also to review it? I see it is also marked for 4.x but I think we can port it also to at least 4.0.x.\r\n\r\n ","created":"2021-09-09T13:12:56.202+0000"},{"body":"I reworked 3.11 and 4.0 and pushed additional changes for trunk based on Stefania Alborghetti's patch. (I will add her as author at the end before commit)\r\n\r\n*trunk:* Byte buffers of {{BooleanSerializer}} are now read-only. We cannot make them on-heap read-only, as we would need to change our code in several key places which seems not worth it at this point. However, buffers can be off-heap read-only. Note that this only offers partial protection against put calls done on the buffer itself. It will not protect, amongst several cases, for put calls where the read-only buffer is the source, as it was the case in {{AbstractCompositeType}}. In this case, the position will still be advanced. To make these buffers completely safe we would need to duplicate them and accept the additional GC pressure.\r\n||Patch||CI||\r\n|​[3.11|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-3.11?expand=1]|[Jenkins|https://ci-cassandra.apache.org/job/Cassandra-devbranch/1171/#showFailuresLink]|\r\n|​[4.0|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-4.0?expand=1]|[Jenkins|https://ci-cassandra.apache.org/job/Cassandra-devbranch/1172/#showFailuresLink]|\r\n|​[trunk|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-trunk?expand=1]|[Jenkins|https://ci-cassandra.apache.org/job/Cassandra-devbranch/1183/#showFailuresLink]|\r\n\r\nI don't see any related issues in the CI runs.\r\n\r\n[~blerer], [~ifesdjeen], as you are already familiar with this, does anyone of you have time for review?","created":"2021-10-06T17:35:44.688+0000"},{"body":"lgtm, +1","created":"2021-12-10T08:05:59.340+0000"},{"body":"This patch could potentially have serious performance consequences, by making many call-sites megamorphic that were previously bimorphic for clusters using e.g. offheap_buffers (and perhaps offheap_objects)...","created":"2021-12-10T12:20:43.638+0000"},{"body":"yeah my bad, only checked the 3.11 patch, assumed they were the same","created":"2021-12-10T12:25:39.242+0000"},{"body":"Thank you both, actually Benjamin made a pass and there was also another issue with the trunk patch that the tests didn't catch but I got side-tracked and left to follow up on this when there is more time to work on it and not jumping in between other tasks as it is important to do it right. \r\n\r\nQuick suggestion - should we apply the 3.11 patch and fix the bug for our users and open follow up ticket if you think it is really worth it to pursue something more at this point? [~marcuse] , [~blerer] , [~benedict] , what do you think about that?","created":"2021-12-10T15:32:36.303+0000"},{"body":"bq. should we apply the 3.11 patch and fix the bug for our users and open follow up ticket if you think it is really worth it to pursue something more at this point?\r\nyes this sounds good to me","created":"2021-12-10T15:49:51.741+0000"},{"body":"Based on feedback and also Slack discussion with [~marcuse], I rebased and updated all branches.\r\n\r\nAdded one more test by [~marcuse], CI started:\r\n\r\n[3.11| https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-3.11?expand=1]|[Jenkins CI| https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1587/]\r\n\r\n[4.0| https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-4.0?expand=1]|[Jenkins CI| https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1588/]\r\n\r\n[trunk| https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-trunk?expand=1]|[Jenkins CI| https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1589/]  ","created":"2022-04-08T20:31:12.223+0000"},{"body":"Jenkins 3.11 results look awful, seems like I ran it by mistake with the old not rebased branch.... Pushed it again now [here|https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1591/]","created":"2022-04-11T02:22:08.968+0000"},{"body":"4.0 has one new failure but I don't see how it can be related - [https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1588/testReport/junit/org.apache.cassandra.distributed.test/RepairTest/testForcedNormalRepairWithOneNodeDown/]\r\n\r\nI will run it with the multiplexer in a loop tomorrow. \r\n\r\nThe trunk version has suffered from full disk... I will run these tests again probably in Circle tomorrow morning. ","created":"2022-04-11T02:28:39.317+0000"},{"body":"So Jenkins failed with some weird rpm errors. I realized I didn't use the image used for CI but one built more recently so there might be some differences causing troubles. Pushed with the right image, nightlies connection issues but those errors I saw before are gone. I wanted to be sure about them as CircleCI does not test our packaging, only Jenkins.\r\n\r\nI pushed CircleCI runs for all three branches, still running: \r\n\r\n [3.11 CircleCI|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1525/workflows/86e5bb0d-2e14-472e-a9f3-312b24ac1a38]\r\n\r\n4.0 [j8|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1526/workflows/9855f9f5-4543-4fa6-a2a9-11ec0718f1e4], [j11|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1526/workflows/e4dbf043-aaf3-471a-bae6-931676df893c]\r\n\r\ntrunk [j8|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1527/workflows/81685d1c-5349-4f44-b39b-d655ccb00cee], [j11|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1527/workflows/efbb02f0-f6fd-4d8f-92d1-404f1699d7b0]","created":"2022-04-11T17:37:43.894+0000"},{"body":"The upgrade_udtfix_test look suspicious from the perspective they don't fail in Jenkins (talking about Cassandra-3.11).\r\n\r\nI pushed in a loop with the patch:\r\n\r\n[https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1532/workflows/dc444e1f-42ca-4993-b053-5a9c69ba1c2d]\r\n\r\nWithout the patch:\r\n\r\n[https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1533/workflows/118fff4b-1d4f-4fd6-8b2d-c33dd8f76bde]\r\n\r\n \r\n\r\nEverything is green and when you try to look at the logs - test runs were just skipped.... this made me check Jenkins, it seems those tests are skipped also there... \r\n\r\nMaybe [~jlewandowski], [~brandon.williams] or [~mck] will know something? I see the three of you were interacting with those tests before.  \r\n\r\n \r\n\r\nThe rest of the failures are all known with associated tickets.","created":"2022-04-11T23:26:18.003+0000"},{"body":"So I reran the upgrade tests with 3.11 in CircleCI without the patch. - https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1533/workflows/118fff4b-1d4f-4fd6-8b2d-c33dd8f76bde/jobs/9933/tests\r\nThe only two tests that failed with my patch and I don't see in the list of failures without are:\r\nupgrade_tests.cql_tests.TestCQLNodes2RF1_Upgrade_indev_3_11_x_To_indev_trunk - I found it in Butler, trunk (https://ci-cassandra.apache.org/job/Cassandra-trunk/1076/testReport/dtest-upgrade.upgrade_tests.cql_tests/TestCQLNodes2RF1_Upgrade_indev_3_0_x_To_indev_trunk/test_noncomposite_static_cf/) \r\n \r\ntest_noncomposite_static_cf - found it again on trunk... https://ci-cassandra.apache.org/job/Cassandra-trunk/1076/testReport/dtest-upgrade.upgrade_tests.cql_tests/cls/test_noncomposite_static_cf/\r\n\r\nSo to conclude, I see the same failures, except two which for some reason in Butler appear under trunk and not 3.11... I am not sure why I see what I see...\r\n\r\nI got +1 on the patch on green CI from [~blerer] in Slack.\r\nI consider all failures known and I will open a ticket to align how we run the upgrade tests in Circle and in Jenkins to ensure we don't miss anything... \r\n\r\nI will wait until tomorrow to commit the patch as there is some Jenkins issue and I see the last runs marked in red. Infra is checking. \r\n","created":"2022-04-13T01:36:19.475+0000"},{"body":"While waiting for Jenkins issues to be cleared I realized the issue was not reported for 3.0 but it also exists so pushed there the same patch and tests. 3.0 will be supported one more year so let's be good citizens.\r\n\r\n[3.0 patch|https://github.com/apache/cassandra/compare/cassandra-3.0...ekaterinadimitrova2:14752-3.0]| [CI|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1534/workflows/9ceb0797-e623-4b1e-b3f6-f8ef9fa29f6d]\r\n\r\nCI only known issues and the same upgrade tests issue I see with 3.11 and not related to the patch","created":"2022-04-13T16:06:55.320+0000"},{"body":"[~blerer] , [~marcuse]  do you also agree on the 3.0 patch? I can commit all branches tomorrow","created":"2022-04-14T02:33:46.981+0000"},{"body":"+1 :)","created":"2022-04-14T09:01:15.308+0000"},{"body":"Starting commit, pending CI:\r\n\r\n[3.0|https://github.com/apache/cassandra/compare/cassandra-3.0...ekaterinadimitrova2:14752-3.0] | [CI|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1539/workflows/45e492b7-bdb7-43d3-ac9c-194675661b86]\r\n\r\n[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...ekaterinadimitrova2:14752-3.11?expand=1] | [CI|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1540/workflows/92dd6163-c11a-41e7-8156-acfff45010db] \r\n\r\n[4.0|https://github.com/apache/cassandra/compare/cassandra-4.0...ekaterinadimitrova2:14752-4.0?expand=1] | [CI J8|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1541/workflows/5a606780-788f-4ce7-a769-47f2c91e9398] | [CI J11|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1541/workflows/cc29f66c-3679-48ca-9c49-388106a02162] \r\n\r\n[trunk|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-trunk?expand=1] | [CI J8|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1542/workflows/823731cb-54f6-4208-a619-6e77af751fa0] | [CI J11|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1542/workflows/11215d18-48d0-4b7a-938b-0b2c0e4dc3f8]","created":"2022-04-15T23:17:35.911+0000"},{"body":"I hate this CI game :( I think there was some environmental issue with 4.0 and there are 6 tests failing with J11 on J8.\r\n\r\nI reran the suite now - [https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1541/workflows/5a606780-788f-4ce7-a769-47f2c91e9398/jobs/10057.] All green\r\n\r\nMost of the failures were timeouts and busy addresses. There were two failures which I've not seen before and looked suspicious; not from the perspective of this patch but in general.\r\n\r\nI cannot reproduce with or without the patch in 100 runs and it seems possible to be related to the other environmental twists:\r\n\r\n1) [https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1544/workflows/0e743bbc-694b-4105-811a-504822334f97/jobs/10060/parallel-runs/13?filterBy=ALL]\r\n\r\n[https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1545/workflows/864bee95-d173-41aa-9702-889e39ee7117/jobs/10061]\r\n\r\n2) test_expiration_overflow_policy_capnowarn\r\n\r\nwith the patch\r\n\r\nhttps://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1546/workflows/4c16d1d4-309a-4623-b4a0-2eccd9569dd5\r\n\r\nwithout the patch\r\n\r\nhttps://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1547/workflows/13c1e0a1-aee2-462a-93ea-6132aa9e7c84\r\n\r\n \r\n\r\nCommitting this in a bit... ","created":"2022-04-17T17:50:32.492+0000"},{"body":"Committed to all branches:\r\n\r\nTo https://github.com/apache/cassandra.git\r\n\r\n   84bc0e8c3b..fb66800a00  cassandra-3.0 -> cassandra-3.0\r\n\r\n   cde152cdb7..eeb89ea53b  cassandra-3.11 -> cassandra-3.11\r\n\r\n   d1270c204f..a46e467737  cassandra-4.0 -> cassandra-4.0\r\n\r\n   74bb6d8496..03ef67c9d5  trunk -> trunk","created":"2022-04-17T18:49:58.498+0000"}],"conversations":[{"body":"[https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/serializers/BooleanSerializer.java#L26] It has two static Bytebuffer variables:-\r\n{code:java}\r\nprivate static final ByteBuffer TRUE = ByteBuffer.wrap(new byte[]{1});\r\nprivate static final ByteBuffer FALSE = ByteBuffer.wrap(new byte[]{0});{code}\r\nWhat will happen if the position of these Bytebuffers is being changed by some other operations? It'll affect other subsequent operations. -IMO Using static is not a good idea here.-\r\n\r\nA potential place where it can become problematic: [https://github.com/apache/cassandra/blob/cassandra-2.1.13/src/java/org/apache/cassandra/db/marshal/AbstractCompositeType.java#L243] Since we are calling *`.remaining()`* It may give wrong results _i.e 0_ if these Bytebuffers have been used previously.\r\n\r\nSolution: \r\n [https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/serializers/BooleanSerializer.java#L42] Every time we return new bytebuffer object. Please do let me know If there is a better way. I'd like to contribute. Thanks!!\r\n{code:java}\r\npublic ByteBuffer serialize(Boolean value)\r\n{\r\nreturn (value == null) ? ByteBufferUtil.EMPTY_BYTE_BUFFER\r\n: value ? ByteBuffer.wrap(new byte[] {1}) : ByteBuffer.wrap(new byte[] {0}); // false\r\n}\r\n{code}","from":"reporter","subject":"serializers/BooleanSerializer.java is using static bytebuffers which may cause problem for subsequent operations"},{"body":"bq. What will happen if the position of these Bytebuffers is being changed by some other operations?\r\n\r\nIn C* you should not change the position unless you have duplicated your `ByteBuffer` first or you really know what you are doing. All the data being stored in the memtables for example are shared so if you change the position on a `ByteBuffer` coming from there you can corrupt the data in a much worst way.\r\n\r\nPersonally, I would not try to change that code as its effect would simply be to cause more garbage.","from":"developer"},{"body":" \r\n\r\n[~blerer] Thanks for your reply. In one of our tool, we use below code to generate the DecoratedKey from String and In the case of boolean type, we are facing this issue.\r\n{code:java}\r\nDatabaseDescriptor.getPartitioner().decorateKey(getKeyValidator(row.getColumnFamily())\r\n.fromString(stringKey));\r\n{code}\r\n\r\n[https://github.com/apache/cassandra/blob/cassandra-2.1.13/src/java/org/apache/cassandra/db/marshal/AbstractCompositeType.java#L255] `byteBuffer.put` changes the position. Though it has a comment: *// it's ok to consume component as we won't use it anymore.*\r\n\r\n ","from":"developer"},{"body":"I found there are too many usages of `AbstractCompositeType#fromString()`. \r\nOne way to corrupt the data:-\r\n\r\nTable schema:-\r\n{code:java}\r\nCREATE TABLE ks1.table1 (\r\nt_id boolean,\r\nid boolean,\r\nck boolean,\r\nnk boolean,\r\nPRIMARY KEY ((t_id,id),ck)\r\n);{code}\r\nInsert statement:-\r\n{code:java}\r\ninsert into ks1.table1 (t_id, ck, id, nk)\r\nVALUES (true, false, false, true);\r\n{code}\r\nNow run nodetool command to get the SSTable for given key:-\r\n{code:java}\r\nbin/nodetool getsstables  ks1 table1 \"false:true\"\r\n{code}\r\nBasically, this operation will modify the positions.\r\n\r\nInsert again:-\r\n{code:java}\r\ninsert into ks1.table1 (t_id, ck, id, nk)\r\nVALUES (true, true, false, true);\r\n{code}\r\nselect data from this table:-\r\n{code:java}\r\ntrue,false,false,true\r\nnull,null,null,null\r\n{code}\r\nSo now all boolean type data will be written as null.","from":"developer"},{"body":" \r\n||MR||\r\n|[trunk\\|https://github.com/Barala/cassandra/commits/CASSANDRA-14752-trunk]|\r\n\r\nI came up with a different approach where BooleanSerializer's static ByteBuffer can be detected using a reference equality check. This approach will avoid new object creations of ByteBuffers.\r\n\r\nI raised MR for trunk. If it passes the review then I'll raise MR to patch other affected versions.","from":"developer"},{"body":"Hey [~varuna], apologize for the very late response.\r\n\r\nI think we can simplify a bit by reducing the patch to changing [this line|https://github.com/Barala/cassandra/commit/2328cda4d4e12e75bc1e9606969f18dd72d4fc27#diff-821957692da5432eb814312397d145fc6445ac35856563b21895e0c1df82ca3aL235]\r\n\r\nto \r\n{code:java}\r\nbb.put(component.duplicate()); // it's not ok to consume component as we did not create it{code}\r\nWhat do you think?\r\n[~blerer] , do you mind also to review it? I see it is also marked for 4.x but I think we can port it also to at least 4.0.x.\r\n\r\n ","from":"developer"},{"body":"I reworked 3.11 and 4.0 and pushed additional changes for trunk based on Stefania Alborghetti's patch. (I will add her as author at the end before commit)\r\n\r\n*trunk:* Byte buffers of {{BooleanSerializer}} are now read-only. We cannot make them on-heap read-only, as we would need to change our code in several key places which seems not worth it at this point. However, buffers can be off-heap read-only. Note that this only offers partial protection against put calls done on the buffer itself. It will not protect, amongst several cases, for put calls where the read-only buffer is the source, as it was the case in {{AbstractCompositeType}}. In this case, the position will still be advanced. To make these buffers completely safe we would need to duplicate them and accept the additional GC pressure.\r\n||Patch||CI||\r\n|​[3.11|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-3.11?expand=1]|[Jenkins|https://ci-cassandra.apache.org/job/Cassandra-devbranch/1171/#showFailuresLink]|\r\n|​[4.0|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-4.0?expand=1]|[Jenkins|https://ci-cassandra.apache.org/job/Cassandra-devbranch/1172/#showFailuresLink]|\r\n|​[trunk|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-trunk?expand=1]|[Jenkins|https://ci-cassandra.apache.org/job/Cassandra-devbranch/1183/#showFailuresLink]|\r\n\r\nI don't see any related issues in the CI runs.\r\n\r\n[~blerer], [~ifesdjeen], as you are already familiar with this, does anyone of you have time for review?","from":"developer"},{"body":"lgtm, +1","from":"developer"},{"body":"This patch could potentially have serious performance consequences, by making many call-sites megamorphic that were previously bimorphic for clusters using e.g. offheap_buffers (and perhaps offheap_objects)...","from":"developer"},{"body":"yeah my bad, only checked the 3.11 patch, assumed they were the same","from":"developer"},{"body":"Thank you both, actually Benjamin made a pass and there was also another issue with the trunk patch that the tests didn't catch but I got side-tracked and left to follow up on this when there is more time to work on it and not jumping in between other tasks as it is important to do it right. \r\n\r\nQuick suggestion - should we apply the 3.11 patch and fix the bug for our users and open follow up ticket if you think it is really worth it to pursue something more at this point? [~marcuse] , [~blerer] , [~benedict] , what do you think about that?","from":"developer"},{"body":"bq. should we apply the 3.11 patch and fix the bug for our users and open follow up ticket if you think it is really worth it to pursue something more at this point?\r\nyes this sounds good to me","from":"developer"},{"body":"Based on feedback and also Slack discussion with [~marcuse], I rebased and updated all branches.\r\n\r\nAdded one more test by [~marcuse], CI started:\r\n\r\n[3.11| https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-3.11?expand=1]|[Jenkins CI| https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1587/]\r\n\r\n[4.0| https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-4.0?expand=1]|[Jenkins CI| https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1588/]\r\n\r\n[trunk| https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-trunk?expand=1]|[Jenkins CI| https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1589/]  ","from":"developer"},{"body":"Jenkins 3.11 results look awful, seems like I ran it by mistake with the old not rebased branch.... Pushed it again now [here|https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1591/]","from":"developer"},{"body":"4.0 has one new failure but I don't see how it can be related - [https://jenkins-cm4.apache.org/job/Cassandra-devbranch/1588/testReport/junit/org.apache.cassandra.distributed.test/RepairTest/testForcedNormalRepairWithOneNodeDown/]\r\n\r\nI will run it with the multiplexer in a loop tomorrow. \r\n\r\nThe trunk version has suffered from full disk... I will run these tests again probably in Circle tomorrow morning. ","from":"developer"},{"body":"So Jenkins failed with some weird rpm errors. I realized I didn't use the image used for CI but one built more recently so there might be some differences causing troubles. Pushed with the right image, nightlies connection issues but those errors I saw before are gone. I wanted to be sure about them as CircleCI does not test our packaging, only Jenkins.\r\n\r\nI pushed CircleCI runs for all three branches, still running: \r\n\r\n [3.11 CircleCI|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1525/workflows/86e5bb0d-2e14-472e-a9f3-312b24ac1a38]\r\n\r\n4.0 [j8|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1526/workflows/9855f9f5-4543-4fa6-a2a9-11ec0718f1e4], [j11|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1526/workflows/e4dbf043-aaf3-471a-bae6-931676df893c]\r\n\r\ntrunk [j8|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1527/workflows/81685d1c-5349-4f44-b39b-d655ccb00cee], [j11|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1527/workflows/efbb02f0-f6fd-4d8f-92d1-404f1699d7b0]","from":"developer"},{"body":"The upgrade_udtfix_test look suspicious from the perspective they don't fail in Jenkins (talking about Cassandra-3.11).\r\n\r\nI pushed in a loop with the patch:\r\n\r\n[https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1532/workflows/dc444e1f-42ca-4993-b053-5a9c69ba1c2d]\r\n\r\nWithout the patch:\r\n\r\n[https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1533/workflows/118fff4b-1d4f-4fd6-8b2d-c33dd8f76bde]\r\n\r\n \r\n\r\nEverything is green and when you try to look at the logs - test runs were just skipped.... this made me check Jenkins, it seems those tests are skipped also there... \r\n\r\nMaybe [~jlewandowski], [~brandon.williams] or [~mck] will know something? I see the three of you were interacting with those tests before.  \r\n\r\n \r\n\r\nThe rest of the failures are all known with associated tickets.","from":"developer"},{"body":"So I reran the upgrade tests with 3.11 in CircleCI without the patch. - https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1533/workflows/118fff4b-1d4f-4fd6-8b2d-c33dd8f76bde/jobs/9933/tests\r\nThe only two tests that failed with my patch and I don't see in the list of failures without are:\r\nupgrade_tests.cql_tests.TestCQLNodes2RF1_Upgrade_indev_3_11_x_To_indev_trunk - I found it in Butler, trunk (https://ci-cassandra.apache.org/job/Cassandra-trunk/1076/testReport/dtest-upgrade.upgrade_tests.cql_tests/TestCQLNodes2RF1_Upgrade_indev_3_0_x_To_indev_trunk/test_noncomposite_static_cf/) \r\n \r\ntest_noncomposite_static_cf - found it again on trunk... https://ci-cassandra.apache.org/job/Cassandra-trunk/1076/testReport/dtest-upgrade.upgrade_tests.cql_tests/cls/test_noncomposite_static_cf/\r\n\r\nSo to conclude, I see the same failures, except two which for some reason in Butler appear under trunk and not 3.11... I am not sure why I see what I see...\r\n\r\nI got +1 on the patch on green CI from [~blerer] in Slack.\r\nI consider all failures known and I will open a ticket to align how we run the upgrade tests in Circle and in Jenkins to ensure we don't miss anything... \r\n\r\nI will wait until tomorrow to commit the patch as there is some Jenkins issue and I see the last runs marked in red. Infra is checking. \r\n","from":"developer"},{"body":"While waiting for Jenkins issues to be cleared I realized the issue was not reported for 3.0 but it also exists so pushed there the same patch and tests. 3.0 will be supported one more year so let's be good citizens.\r\n\r\n[3.0 patch|https://github.com/apache/cassandra/compare/cassandra-3.0...ekaterinadimitrova2:14752-3.0]| [CI|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1534/workflows/9ceb0797-e623-4b1e-b3f6-f8ef9fa29f6d]\r\n\r\nCI only known issues and the same upgrade tests issue I see with 3.11 and not related to the patch","from":"developer"},{"body":"[~blerer] , [~marcuse]  do you also agree on the 3.0 patch? I can commit all branches tomorrow","from":"developer"},{"body":"+1 :)","from":"developer"},{"body":"Starting commit, pending CI:\r\n\r\n[3.0|https://github.com/apache/cassandra/compare/cassandra-3.0...ekaterinadimitrova2:14752-3.0] | [CI|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1539/workflows/45e492b7-bdb7-43d3-ac9c-194675661b86]\r\n\r\n[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...ekaterinadimitrova2:14752-3.11?expand=1] | [CI|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1540/workflows/92dd6163-c11a-41e7-8156-acfff45010db] \r\n\r\n[4.0|https://github.com/apache/cassandra/compare/cassandra-4.0...ekaterinadimitrova2:14752-4.0?expand=1] | [CI J8|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1541/workflows/5a606780-788f-4ce7-a769-47f2c91e9398] | [CI J11|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1541/workflows/cc29f66c-3679-48ca-9c49-388106a02162] \r\n\r\n[trunk|https://github.com/apache/cassandra/compare/trunk...ekaterinadimitrova2:14752-trunk?expand=1] | [CI J8|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1542/workflows/823731cb-54f6-4208-a619-6e77af751fa0] | [CI J11|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1542/workflows/11215d18-48d0-4b7a-938b-0b2c0e4dc3f8]","from":"developer"},{"body":"I hate this CI game :( I think there was some environmental issue with 4.0 and there are 6 tests failing with J11 on J8.\r\n\r\nI reran the suite now - [https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1541/workflows/5a606780-788f-4ce7-a769-47f2c91e9398/jobs/10057.] All green\r\n\r\nMost of the failures were timeouts and busy addresses. There were two failures which I've not seen before and looked suspicious; not from the perspective of this patch but in general.\r\n\r\nI cannot reproduce with or without the patch in 100 runs and it seems possible to be related to the other environmental twists:\r\n\r\n1) [https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1544/workflows/0e743bbc-694b-4105-811a-504822334f97/jobs/10060/parallel-runs/13?filterBy=ALL]\r\n\r\n[https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1545/workflows/864bee95-d173-41aa-9702-889e39ee7117/jobs/10061]\r\n\r\n2) test_expiration_overflow_policy_capnowarn\r\n\r\nwith the patch\r\n\r\nhttps://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1546/workflows/4c16d1d4-309a-4623-b4a0-2eccd9569dd5\r\n\r\nwithout the patch\r\n\r\nhttps://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1547/workflows/13c1e0a1-aee2-462a-93ea-6132aa9e7c84\r\n\r\n \r\n\r\nCommitting this in a bit... ","from":"developer"},{"body":"Committed to all branches:\r\n\r\nTo https://github.com/apache/cassandra.git\r\n\r\n   84bc0e8c3b..fb66800a00  cassandra-3.0 -> cassandra-3.0\r\n\r\n   cde152cdb7..eeb89ea53b  cassandra-3.11 -> cassandra-3.11\r\n\r\n   d1270c204f..a46e467737  cassandra-4.0 -> cassandra-4.0\r\n\r\n   74bb6d8496..03ef67c9d5  trunk -> trunk","from":"developer"}],"created":"2018-09-14T07:50:38.000+0000","description":"[https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/serializers/BooleanSerializer.java#L26] It has two static Bytebuffer variables:-\r\n{code:java}\r\nprivate static final ByteBuffer TRUE = ByteBuffer.wrap(new byte[]{1});\r\nprivate static final ByteBuffer FALSE = ByteBuffer.wrap(new byte[]{0});{code}\r\nWhat will happen if the position of these Bytebuffers is being changed by some other operations? It'll affect other subsequent operations. -IMO Using static is not a good idea here.-\r\n\r\nA potential place where it can become problematic: [https://github.com/apache/cassandra/blob/cassandra-2.1.13/src/java/org/apache/cassandra/db/marshal/AbstractCompositeType.java#L243] Since we are calling *`.remaining()`* It may give wrong results _i.e 0_ if these Bytebuffers have been used previously.\r\n\r\nSolution: \r\n [https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/serializers/BooleanSerializer.java#L42] Every time we return new bytebuffer object. Please do let me know If there is a better way. I'd like to contribute. Thanks!!\r\n{code:java}\r\npublic ByteBuffer serialize(Boolean value)\r\n{\r\nreturn (value == null) ? ByteBufferUtil.EMPTY_BYTE_BUFFER\r\n: value ? ByteBuffer.wrap(new byte[] {1}) : ByteBuffer.wrap(new byte[] {0}); // false\r\n}\r\n{code}","issue_id":"13185075","key":"CASSANDRA-14752","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2022-04-17T18:51:17.000+0000","role":"fixed_distractor","summary":"serializers/BooleanSerializer.java is using static bytebuffers which may cause problem for subsequent operations"} {"case_id":"13186708","cluster":"DISTRACTOR-CASSANDRA-14781","comments":[{"body":"Related: https://issues.apache.org/jira/browse/CASSANDRA-12231\r\n\r\n[^CASSANDRA-14781.patch]\r\n\r\n^The attached trivial patch (built off of 2.2.X) is an easy way to get more info^","created":"2018-09-24T16:46:59.578+0000"},{"body":"[~tpetracca] thanks for the patch! {{Mutation::getColumnFamilies}} was removed in Cassandra 3.0. Would you be interested in updating your patch for 3.0+ and I can review?","created":"2018-09-25T01:49:05.894+0000"},{"body":"I'm significantly less familiar with the 3.0+ code but at a quick glance it looks like every partition update should actually contain all this information at this point.\r\n\r\n[^CASSANDRA-14781_3.0.patch]","created":"2018-09-25T02:37:23.326+0000"},{"body":"To cover extreme cases, can you add a limit of like 10 tables then add a \" ... and X more\" or something? Also can you include the partition keys with a similar cap to prevent the string from getting too large. For the common case where theres only 1 large mutation or a bunch to a single partition that would be monumentally helpful.","created":"2018-10-18T14:31:31.047+0000"},{"body":"[^CASSANDRA-14781_3.11.patch]\r\n\r\nHey [~cnlwsu] [~jrwest], I've attached a patch for 3.11 (should be the same for 3.0) that prints (cf, key) pairs with a limit of 10.","created":"2019-10-04T18:35:30.324+0000"},{"body":"Thanks for the update! A few more things things:\r\n\r\n* Lets make this work for 4.0 first, I am not sure about backporting it to 3.11 or 3.0 at this point\r\n* Can you move the {{limit}} up to a constant? Be great if that was something we can override too in yaml/jmx override incase need to to see more than 10 and accept risk but that can come later if dont want to include it in this patch. \r\n* If limited, can you make sure to show largest keys first instead of just taking first off the updates.\r\n* If we have metadata and the key, instead of doing a toString on the decorated key which gives a large hex dump, you can use {{partitionKeyType.getString(keybytes)}} to get the human readable version.\r\n\r\nsomething like (this is untested, just example):\r\n{code}\r\n mutation.getPartitionUpdates().stream()\r\n .sorted(Comparator.comparingInt(PartitionUpdate::dataSize).reversed())\r\n .limit(LIMIT)\r\n .map(upd -> String.format(\"%s[%s]\",\r\n upd.metadata().name,\r\n upd.metadata().partitionKeyType.getString(upd.partitionKey().getKey())))\r\n .collect(Collectors.joining(\", \"));\r\n{code}\r\n","created":"2019-10-08T18:20:29.242+0000"},{"body":"I'd say this is a little insufficient. You have the exact same situation with writing hints in {{HintsBuffer.allocate()}}.\r\n\r\nWhat you want is to add an extra validation downstream all the way to {{ModificationStatement}}, so that you can return a meaningful exception to the client immediately - rather than ending up timing out the response.","created":"2019-11-12T16:30:48.452+0000"},{"body":"I would also not bother with listing individual tables; keyspace and partition key should hopefully be sufficient enough, and memoise calculated {{Mutation}} size in the {{Mutation}} object (see {{serializedSize*}} fields in {{Message}} in 4.0) to prevent redundant calculations by subsequent stages.","created":"2019-11-12T16:34:36.562+0000"},{"body":"[~tpetracca] Would you mind if I submit a patch for this? I ran into this requirement and made some progress on the patch.","created":"2020-01-02T18:50:47.787+0000"},{"body":"[~jrwest] Made a patch with the changes. [~tpetracca] apologies if I am overtaking you.\r\n\r\nHere is the [patch|https://github.com/nvharikrishna/cassandra/commit/5b3af390ce64860505dfeb3a3549cc9897987771] and [CI|https://app.circleci.com/jobs/github/nvharikrishna/cassandra/95].\r\n\r\nI have a small question though. In [CommitLog.java|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/db/commitlog/CommitLog.java#L274], it considers commit log entry overhead with mutation size to compare with max mutation size. Shouldn't it consider only the mutation size? I made changes to be compatible with overhead (I can change based on your comments).\r\n\r\n{{SizeHolder}} introduced in this patch can be used in {{Message}} as well. I can make this change too if it is okay to mixup.","created":"2020-01-06T08:25:57.608+0000"},{"body":"Thanks for picking this up [~n.v.harikrishna]. Comments below. \r\n\r\n \r\n\r\nGeneral:\r\n * Since this change logs keys which in some use cases may contain PII that users may not want logged, should we add a flag to gate this feature? \r\n * I don’t believe this address [~aleksey]'s concern that the exception be propagated to the client vs. timing out. At a minimum, we should add an in-jvm dtest to show it does (I’m happy to help contribute this unless you’d like to). Its possible to split the patch into two parts (one that changes the logging and one that propagates the exception), however, in prod I’ve seen the lack of this logging and the timeouts be an issue so they are definitely both worth addressing. \r\n\r\nCommitLog:\r\n * With your change to validate the size using serializedSize, it looks like its no longer be necessary to use the scratch buffer (which seems to be used primarily to get the length w/o calculating serialized size). However, I wonder if the original approach is better performance wise since we are already serializing the mutation. It would be good to benchmark workloads with all mutations exceeding max size, no mutations exceeding max size, and mixed to check because I can imagine how what I said is wrong as well. We should check if there is any difference or if its negligible. \r\n\r\nSizeHolder:\r\n * Since the only call to {{#getSize()}} currently is {{IMutation#validateSize}} which is only called once from CommitLog.java. Consider removing the memoization. \r\n * If the memoization is still desired, I don’t think its necessary to have fields for each version since a mutation is executed within a single version of Cassandra (it may cross network boundaries and thus versions but at this point we have a new SizeHolder). Using a single variable would cleanup {{#getSize}} considerably. Further, it would be cleaner/less error prone to memoize the size in the mutation since nothing maps the object passed to {{#getSize}} to the stored value. \r\n\r\nMutationExceededMaxSizeException:\r\n * {{prepareMessage}} recalculates the mutation size. While this is an exceptional case it is still common. Perhaps a place where the memoization could be helpful? \r\n * Spelling error: maxMutaionSize\r\n * I almost always prefer a configurable limit to something hardcoded and requiring recompilation to change \r\n\r\nOther:\r\n * I don’t think {{IMutation.getMaxMutationSize()}} is necessary. Use {{DatabaseDescriptor}} instead. This is more typical of Cassandra config code. Further, the original code only called {{DatabaseDescriptor.getMaxMutationSize()}} once instead of for every write. Unless we are going to make it a hot-prop its worth putting back to how it was. \r\n * We have a few {{Serializer}} interfaces in C* already, consider renaming {{SizeHolder.Serializer}} if its kept","created":"2020-01-06T21:28:57.864+0000"},{"body":"For what it's worth, I would strongly prefer the memoised fields to be inlined into {{Mutation}}, same way they are in {{Message}}. It's a very common object, and the bloat of the overhead of an extra {{SizeHolder}} object vs. just the 3 inlined fields matters.\r\n\r\nWe do want the memoisation itself, as otherwise we would be performing this calculation at least twice - once when validation, once when calculating message size for internode.","created":"2020-01-07T12:51:48.751+0000"},{"body":"{quote}Since the only call to #getSize() currently is IMutation#validateSize which is only called once from CommitLog.java. Consider removing the memoization.\r\n{quote}\r\nSize is validated at org.apache.cassandra.service.reads.repair.BlockingReadRepairs#createRepairMutation method too. It may use different version for serialization based on destination. I feel memoization is required. Somehow I missed to include changes for this method in the patch. Including this change while addressing other review comments. Will update the patch asap.","created":"2020-01-07T18:53:45.872+0000"},{"body":"Thanks [~n.v.harikrishna] . My comment re: memoization was more to give an option if we didn't need it. Since we do (based on my later comments, your comments, and [~aleksey] comments), I agree w/ [~aleksey] suggestion re: removing {{SizeHolder}}. ","created":"2020-01-07T19:26:23.340+0000"},{"body":"[~jrwest] Made the changes. \r\n\r\nHere is the [updated patch|https://github.com/nvharikrishna/cassandra/commit/1eb9a9846187f669516c88c85fa3550e4efb08f7] and [CI|https://app.circleci.com/jobs/github/nvharikrishna/cassandra/120]. \r\n\r\nSummary of changes:\r\n * Removed SizeHolder\r\n * MutationExceededMaxSizeException\r\n ** Avoided calculating size again.\r\n ** Changed constant for limiting no.of keys to size of the message. We are mostly concerned about dumping huge message to log. No.of keys to log has to vary based on its size and there won't be an ideal config by no.of keys. So changed it to message size (1kb for now, we can increase it further).\r\n * IMutation\r\n ** removed getMaxMutationSize and replaced it with constant from CommitLog. \r\n * Replaced Mutation.serializer.serializedSize with mutation.serializedSize.","created":"2020-01-08T19:18:33.520+0000"},{"body":"Could you flip the size fields to {{int}} from {{long}}? (not a full review, just a super quick skim related to my only previous point)","created":"2020-01-10T14:24:08.244+0000"},{"body":"A few code review comments below. I did want to discuss if we are going to address the user facing concerns Aleksey brought up in this ticket? The patch addresses the operators lack of visibility into keyspace/table/partitions but still results in timeouts for the user. Are we going to address those in a separate ticket? My thought is that something for the operators is better than no patch (having been blind in this situation before besides custom tools) but if the user facing changes require protocol changes we should probably fix it pre-4.0 like we have or plan to w/ other similar tickets – but that could still be in a separate ticket.\r\n\r\n \r\n\r\nCode Comments:  \r\n * -{{Mutation#serializedSize}}: you should only need one field to memoize the size and can you pass version directly to {{Serializer#serializedSize}} instead of the switch afterwards?- never mind [~n.v.harikrishna] pointed out code in {{BlockingReadRepairs}} that makes this statement false. \r\n * We also shouldn’t duplicate the implementations between counter and regular mutations\r\n * {{validateSize}}: since the two implementations are identical you could move them to a \\{{default}} implementation in {{IMutation}}\r\n * {\\{MaxMutationExceededException}}: the sort in {{#prepareMessage}} could get pretty expensive, is it necessary? It also looks like there is an edge case where “and more” will be added even when there aren’t more. Using {{listIterator.hasNext()}} instead of {{topPartitions.size() > 0}} should fix that","created":"2020-01-10T15:46:00.656+0000"},{"body":"[~jrwest]\r\n{quote}A few code review comments below. I did want to discuss if we are going to address the user facing concerns Aleksey brought up in this ticket? The patch addresses the operators lack of visibility into keyspace/table/partitions but still results in timeouts for the user. Are we going to address those in a separate ticket? My thought is that something for the operators is better than no patch (having been blind in this situation before besides custom tools) but if the user facing changes require protocol changes we should probably fix it pre-4.0 like we have or plan to w/ other similar tickets – but that could still be in a separate ticket.\r\n{quote}\r\nI would prefer to have a separate ticket. +1 on having something better than no patch.\r\n{quote}\r\n* We also shouldn’t duplicate the implementations between counter and regular mutations\r\n* validateSize: since the two implementations are identical you could move them to a default implementation in IMutation\r\n{quote}\r\nvalidateSize methods implementation looks similar, but they use different serialisers (Mutation uses MutationSerializer and CounterMutation uses CounterMutationSerializer) which are not visible in IMutation interface. serializedSize() methods needs memoization (needs serializedSize* fields) be cause of which we cannot move to interface as fields will be final. We had ruled out option having a separate class (i.e. SizeHolder). Now it makes me think about 2 options.\r\n\r\n1. I could not find any validation of size for {{CounterMutation}}. If it is expected or not required, then can remove {{CounterMutaiton}} changes and use\\{{ Mutation.validateSize}} directly (instead of defining it in {{IMutation}}). The disadvantage I see with this approach is caller has to be aware of implementation and it makes things hard to abstract (code has to be aware of implementation instead of {{IMutation}}).\r\n\r\n2. Expect {{VirtualMutaiton}}, I see mutations are expected to be serialized and/or deserialized. Provide serialize, serialziedSize and deserialize methods as part of {{IMutaiton}} (so that we can abstract out direct usages of {{Mutation.serializer}} and {{CounterMutation.serializer}}) with an abstract class in between having common functionality.\r\n\r\nOr else pay the price of duplicate code. What do you think?\r\n{quote}MaxMutationExceededException: the sort in #prepareMessage could get pretty expensive, is it necessary?\r\n{quote}\r\nIn Mutation, I see that there is only one PartitionUpdate per Table, and according to Mutation.merge() logic, a mutation can have changes related to only one keyspace and one key. Even if there are multiple updates for different rows of same partition, they are merged into single PartitionUpdate.\r\n\r\nWhen I ran a small test for sorting list of Longs (laptop having i7, 6 core and 16gb ram) it took approximately 33ms, 6ms and 1ms for 100K, 10k and 1k respectively.\r\n\r\nAccording to merge logic and test numbers, unless there are thousands of tables in a key space and trying to update all of them at once, I dont see a scenario where sorting can hurt (time taken for sort > 1 or 2ms).\r\n{quote}It also looks like there is an edge case where “and more” will be added even when there aren’t more. Using listIterator.hasNext() instead of topPartitions.size() > 0 should fix that\r\n{quote}\r\nI had moved the code into separte funtion and added unit test cases. It is working as expected. Using listIterator.hasNext() caused few tests to fail. Did I miss any scenario to test?\r\n\r\nConverted serializedSize* long fields to int as suggested by Aleksey. Changes are here: https://github.com/apache/cassandra/compare/trunk...nvharikrishna:14781-trunk?expand=1\r\n\r\n ","created":"2020-01-14T20:19:04.666+0000"},{"body":"Thanks [~n.v.harikrishna]. I think the duplicate code is probably simpler than making things more nuanced and thanks for testing the sorting performance. The code looks ok me and I am ok w/ the separate ticket approach so that there is some improvements for operators at a minimum but I would like to get [~aleksey]'s input as well. ","created":"2020-01-17T15:52:23.431+0000"},{"body":"I'm perfectly fine with this level of duplication. Don't mind a separate ticket, either. Although FWIW that other ticket is the important one - we should never ever get to the point when an oversized mutation makes it to the commitlog at all, outside of potentially boundary mixed mode conditions - when a mutation validates for the current messaging version but ends up being slightly over the size for legacy.","created":"2020-01-22T14:04:15.585+0000"},{"body":"Was additional ticket opened?\r\nWhat shall we do with this one?","created":"2020-03-23T21:22:49.848+0000"},{"body":"I believe this patch is ready (and has one +1) but needs a committer to review. [~n.v.harikrishna] was another ticket opened? ","created":"2020-04-08T21:35:01.057+0000"},{"body":"[~jrwest] raised CASSANDRA-15741 for validation and/or fixing client timeout when mutation exceeds max size.","created":"2020-04-20T05:51:54.121+0000"},{"body":"Hi [~n.v.harikrishna]. I've picked this back up and am getting it ready to commit. Thanks for your patience. \r\n\r\n \r\n\r\nI've squashed your branch here: [https://github.com/jrwest/cassandra/commits/14781-trunk.] I made a few minor changes along the way (I also re-reviewed since it had been a little bit since I had read the patch):\r\n\r\n \r\n * Modified {{CHANGES.txt}}\r\n * Modified {{IMutation#validateSize}} javadoc\r\n * Moved call to {{Keyspace.open}} into the catch block of {{BlockingReadRepairs#createRepairMutation}}. It was only used if we reached that block anyways.\r\n * Fixed whitespace formatting in {{MutationExceededMaxSizeException#prepareMessage}}\r\n\r\n \r\n\r\nI ran a build prior to these changes. The build looked good (better than trunk actually) and any failures do not seem related: [https://app.circleci.com/pipelines/github/jrwest/cassandra/4/workflows/e43918eb-40d2-45ad-80c3-dbeaa5ee186b]\r\n\r\n \r\n\r\nI've kicked off a new build with the changes above and with the squash performed: [https://app.circleci.com/pipelines/github/jrwest/cassandra/6/workflows/3c3f674e-db89-488a-bbd5-98f04de4fd0d]\r\n\r\nEDIT:\r\n\r\nI was slightly concerned about the failure in {{read_repair_test.py}}'s {{test_speculative_data_request}}. Looking closer at the test runs, its flaky and doesn't look like that flakiness could be related to the changes here (since the mutation sizes are static). \r\n\r\n \r\n\r\nI've also kicked off a Jenkins build for good measure: [https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/107/] \r\n\r\n \r\n\r\n ","created":"2020-04-30T14:39:41.766+0000"},{"body":"Tests looked good. Any failures are suspected flaky tests and are unrelated. Committed as d3dadcd6f3bbde471e972f8332eb62de0f2d4aae. ","created":"2020-05-04T18:20:02.259+0000"},{"body":"Thanks a lot [~jwest] for taking it forward!!","created":"2020-05-05T16:57:53.724+0000"}],"conversations":[{"body":"When hitting [https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/commitlog/CommitLog.java#L256-L257], the log message produced does not help the operator track down what data is being written. At a minimum the keyspace and cfIds involved would be useful (and are available) – more detail might not be reasonable to include. ","from":"reporter","subject":"Log message when mutation passed to CommitLog#add(Mutation) is too large is not descriptive enough"},{"body":"Related: https://issues.apache.org/jira/browse/CASSANDRA-12231\r\n\r\n[^CASSANDRA-14781.patch]\r\n\r\n^The attached trivial patch (built off of 2.2.X) is an easy way to get more info^","from":"developer"},{"body":"[~tpetracca] thanks for the patch! {{Mutation::getColumnFamilies}} was removed in Cassandra 3.0. Would you be interested in updating your patch for 3.0+ and I can review?","from":"developer"},{"body":"I'm significantly less familiar with the 3.0+ code but at a quick glance it looks like every partition update should actually contain all this information at this point.\r\n\r\n[^CASSANDRA-14781_3.0.patch]","from":"developer"},{"body":"To cover extreme cases, can you add a limit of like 10 tables then add a \" ... and X more\" or something? Also can you include the partition keys with a similar cap to prevent the string from getting too large. For the common case where theres only 1 large mutation or a bunch to a single partition that would be monumentally helpful.","from":"developer"},{"body":"[^CASSANDRA-14781_3.11.patch]\r\n\r\nHey [~cnlwsu] [~jrwest], I've attached a patch for 3.11 (should be the same for 3.0) that prints (cf, key) pairs with a limit of 10.","from":"developer"},{"body":"Thanks for the update! A few more things things:\r\n\r\n* Lets make this work for 4.0 first, I am not sure about backporting it to 3.11 or 3.0 at this point\r\n* Can you move the {{limit}} up to a constant? Be great if that was something we can override too in yaml/jmx override incase need to to see more than 10 and accept risk but that can come later if dont want to include it in this patch. \r\n* If limited, can you make sure to show largest keys first instead of just taking first off the updates.\r\n* If we have metadata and the key, instead of doing a toString on the decorated key which gives a large hex dump, you can use {{partitionKeyType.getString(keybytes)}} to get the human readable version.\r\n\r\nsomething like (this is untested, just example):\r\n{code}\r\n mutation.getPartitionUpdates().stream()\r\n .sorted(Comparator.comparingInt(PartitionUpdate::dataSize).reversed())\r\n .limit(LIMIT)\r\n .map(upd -> String.format(\"%s[%s]\",\r\n upd.metadata().name,\r\n upd.metadata().partitionKeyType.getString(upd.partitionKey().getKey())))\r\n .collect(Collectors.joining(\", \"));\r\n{code}\r\n","from":"developer"},{"body":"I'd say this is a little insufficient. You have the exact same situation with writing hints in {{HintsBuffer.allocate()}}.\r\n\r\nWhat you want is to add an extra validation downstream all the way to {{ModificationStatement}}, so that you can return a meaningful exception to the client immediately - rather than ending up timing out the response.","from":"developer"},{"body":"I would also not bother with listing individual tables; keyspace and partition key should hopefully be sufficient enough, and memoise calculated {{Mutation}} size in the {{Mutation}} object (see {{serializedSize*}} fields in {{Message}} in 4.0) to prevent redundant calculations by subsequent stages.","from":"developer"},{"body":"[~tpetracca] Would you mind if I submit a patch for this? I ran into this requirement and made some progress on the patch.","from":"developer"},{"body":"[~jrwest] Made a patch with the changes. [~tpetracca] apologies if I am overtaking you.\r\n\r\nHere is the [patch|https://github.com/nvharikrishna/cassandra/commit/5b3af390ce64860505dfeb3a3549cc9897987771] and [CI|https://app.circleci.com/jobs/github/nvharikrishna/cassandra/95].\r\n\r\nI have a small question though. In [CommitLog.java|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/db/commitlog/CommitLog.java#L274], it considers commit log entry overhead with mutation size to compare with max mutation size. Shouldn't it consider only the mutation size? I made changes to be compatible with overhead (I can change based on your comments).\r\n\r\n{{SizeHolder}} introduced in this patch can be used in {{Message}} as well. I can make this change too if it is okay to mixup.","from":"developer"},{"body":"Thanks for picking this up [~n.v.harikrishna]. Comments below. \r\n\r\n \r\n\r\nGeneral:\r\n * Since this change logs keys which in some use cases may contain PII that users may not want logged, should we add a flag to gate this feature? \r\n * I don’t believe this address [~aleksey]'s concern that the exception be propagated to the client vs. timing out. At a minimum, we should add an in-jvm dtest to show it does (I’m happy to help contribute this unless you’d like to). Its possible to split the patch into two parts (one that changes the logging and one that propagates the exception), however, in prod I’ve seen the lack of this logging and the timeouts be an issue so they are definitely both worth addressing. \r\n\r\nCommitLog:\r\n * With your change to validate the size using serializedSize, it looks like its no longer be necessary to use the scratch buffer (which seems to be used primarily to get the length w/o calculating serialized size). However, I wonder if the original approach is better performance wise since we are already serializing the mutation. It would be good to benchmark workloads with all mutations exceeding max size, no mutations exceeding max size, and mixed to check because I can imagine how what I said is wrong as well. We should check if there is any difference or if its negligible. \r\n\r\nSizeHolder:\r\n * Since the only call to {{#getSize()}} currently is {{IMutation#validateSize}} which is only called once from CommitLog.java. Consider removing the memoization. \r\n * If the memoization is still desired, I don’t think its necessary to have fields for each version since a mutation is executed within a single version of Cassandra (it may cross network boundaries and thus versions but at this point we have a new SizeHolder). Using a single variable would cleanup {{#getSize}} considerably. Further, it would be cleaner/less error prone to memoize the size in the mutation since nothing maps the object passed to {{#getSize}} to the stored value. \r\n\r\nMutationExceededMaxSizeException:\r\n * {{prepareMessage}} recalculates the mutation size. While this is an exceptional case it is still common. Perhaps a place where the memoization could be helpful? \r\n * Spelling error: maxMutaionSize\r\n * I almost always prefer a configurable limit to something hardcoded and requiring recompilation to change \r\n\r\nOther:\r\n * I don’t think {{IMutation.getMaxMutationSize()}} is necessary. Use {{DatabaseDescriptor}} instead. This is more typical of Cassandra config code. Further, the original code only called {{DatabaseDescriptor.getMaxMutationSize()}} once instead of for every write. Unless we are going to make it a hot-prop its worth putting back to how it was. \r\n * We have a few {{Serializer}} interfaces in C* already, consider renaming {{SizeHolder.Serializer}} if its kept","from":"developer"},{"body":"For what it's worth, I would strongly prefer the memoised fields to be inlined into {{Mutation}}, same way they are in {{Message}}. It's a very common object, and the bloat of the overhead of an extra {{SizeHolder}} object vs. just the 3 inlined fields matters.\r\n\r\nWe do want the memoisation itself, as otherwise we would be performing this calculation at least twice - once when validation, once when calculating message size for internode.","from":"developer"},{"body":"{quote}Since the only call to #getSize() currently is IMutation#validateSize which is only called once from CommitLog.java. Consider removing the memoization.\r\n{quote}\r\nSize is validated at org.apache.cassandra.service.reads.repair.BlockingReadRepairs#createRepairMutation method too. It may use different version for serialization based on destination. I feel memoization is required. Somehow I missed to include changes for this method in the patch. Including this change while addressing other review comments. Will update the patch asap.","from":"developer"},{"body":"Thanks [~n.v.harikrishna] . My comment re: memoization was more to give an option if we didn't need it. Since we do (based on my later comments, your comments, and [~aleksey] comments), I agree w/ [~aleksey] suggestion re: removing {{SizeHolder}}. ","from":"developer"},{"body":"[~jrwest] Made the changes. \r\n\r\nHere is the [updated patch|https://github.com/nvharikrishna/cassandra/commit/1eb9a9846187f669516c88c85fa3550e4efb08f7] and [CI|https://app.circleci.com/jobs/github/nvharikrishna/cassandra/120]. \r\n\r\nSummary of changes:\r\n * Removed SizeHolder\r\n * MutationExceededMaxSizeException\r\n ** Avoided calculating size again.\r\n ** Changed constant for limiting no.of keys to size of the message. We are mostly concerned about dumping huge message to log. No.of keys to log has to vary based on its size and there won't be an ideal config by no.of keys. So changed it to message size (1kb for now, we can increase it further).\r\n * IMutation\r\n ** removed getMaxMutationSize and replaced it with constant from CommitLog. \r\n * Replaced Mutation.serializer.serializedSize with mutation.serializedSize.","from":"developer"},{"body":"Could you flip the size fields to {{int}} from {{long}}? (not a full review, just a super quick skim related to my only previous point)","from":"developer"},{"body":"A few code review comments below. I did want to discuss if we are going to address the user facing concerns Aleksey brought up in this ticket? The patch addresses the operators lack of visibility into keyspace/table/partitions but still results in timeouts for the user. Are we going to address those in a separate ticket? My thought is that something for the operators is better than no patch (having been blind in this situation before besides custom tools) but if the user facing changes require protocol changes we should probably fix it pre-4.0 like we have or plan to w/ other similar tickets – but that could still be in a separate ticket.\r\n\r\n \r\n\r\nCode Comments:  \r\n * -{{Mutation#serializedSize}}: you should only need one field to memoize the size and can you pass version directly to {{Serializer#serializedSize}} instead of the switch afterwards?- never mind [~n.v.harikrishna] pointed out code in {{BlockingReadRepairs}} that makes this statement false. \r\n * We also shouldn’t duplicate the implementations between counter and regular mutations\r\n * {{validateSize}}: since the two implementations are identical you could move them to a \\{{default}} implementation in {{IMutation}}\r\n * {\\{MaxMutationExceededException}}: the sort in {{#prepareMessage}} could get pretty expensive, is it necessary? It also looks like there is an edge case where “and more” will be added even when there aren’t more. Using {{listIterator.hasNext()}} instead of {{topPartitions.size() > 0}} should fix that","from":"developer"},{"body":"[~jrwest]\r\n{quote}A few code review comments below. I did want to discuss if we are going to address the user facing concerns Aleksey brought up in this ticket? The patch addresses the operators lack of visibility into keyspace/table/partitions but still results in timeouts for the user. Are we going to address those in a separate ticket? My thought is that something for the operators is better than no patch (having been blind in this situation before besides custom tools) but if the user facing changes require protocol changes we should probably fix it pre-4.0 like we have or plan to w/ other similar tickets – but that could still be in a separate ticket.\r\n{quote}\r\nI would prefer to have a separate ticket. +1 on having something better than no patch.\r\n{quote}\r\n* We also shouldn’t duplicate the implementations between counter and regular mutations\r\n* validateSize: since the two implementations are identical you could move them to a default implementation in IMutation\r\n{quote}\r\nvalidateSize methods implementation looks similar, but they use different serialisers (Mutation uses MutationSerializer and CounterMutation uses CounterMutationSerializer) which are not visible in IMutation interface. serializedSize() methods needs memoization (needs serializedSize* fields) be cause of which we cannot move to interface as fields will be final. We had ruled out option having a separate class (i.e. SizeHolder). Now it makes me think about 2 options.\r\n\r\n1. I could not find any validation of size for {{CounterMutation}}. If it is expected or not required, then can remove {{CounterMutaiton}} changes and use\\{{ Mutation.validateSize}} directly (instead of defining it in {{IMutation}}). The disadvantage I see with this approach is caller has to be aware of implementation and it makes things hard to abstract (code has to be aware of implementation instead of {{IMutation}}).\r\n\r\n2. Expect {{VirtualMutaiton}}, I see mutations are expected to be serialized and/or deserialized. Provide serialize, serialziedSize and deserialize methods as part of {{IMutaiton}} (so that we can abstract out direct usages of {{Mutation.serializer}} and {{CounterMutation.serializer}}) with an abstract class in between having common functionality.\r\n\r\nOr else pay the price of duplicate code. What do you think?\r\n{quote}MaxMutationExceededException: the sort in #prepareMessage could get pretty expensive, is it necessary?\r\n{quote}\r\nIn Mutation, I see that there is only one PartitionUpdate per Table, and according to Mutation.merge() logic, a mutation can have changes related to only one keyspace and one key. Even if there are multiple updates for different rows of same partition, they are merged into single PartitionUpdate.\r\n\r\nWhen I ran a small test for sorting list of Longs (laptop having i7, 6 core and 16gb ram) it took approximately 33ms, 6ms and 1ms for 100K, 10k and 1k respectively.\r\n\r\nAccording to merge logic and test numbers, unless there are thousands of tables in a key space and trying to update all of them at once, I dont see a scenario where sorting can hurt (time taken for sort > 1 or 2ms).\r\n{quote}It also looks like there is an edge case where “and more” will be added even when there aren’t more. Using listIterator.hasNext() instead of topPartitions.size() > 0 should fix that\r\n{quote}\r\nI had moved the code into separte funtion and added unit test cases. It is working as expected. Using listIterator.hasNext() caused few tests to fail. Did I miss any scenario to test?\r\n\r\nConverted serializedSize* long fields to int as suggested by Aleksey. Changes are here: https://github.com/apache/cassandra/compare/trunk...nvharikrishna:14781-trunk?expand=1\r\n\r\n ","from":"developer"},{"body":"Thanks [~n.v.harikrishna]. I think the duplicate code is probably simpler than making things more nuanced and thanks for testing the sorting performance. The code looks ok me and I am ok w/ the separate ticket approach so that there is some improvements for operators at a minimum but I would like to get [~aleksey]'s input as well. ","from":"developer"},{"body":"I'm perfectly fine with this level of duplication. Don't mind a separate ticket, either. Although FWIW that other ticket is the important one - we should never ever get to the point when an oversized mutation makes it to the commitlog at all, outside of potentially boundary mixed mode conditions - when a mutation validates for the current messaging version but ends up being slightly over the size for legacy.","from":"developer"},{"body":"Was additional ticket opened?\r\nWhat shall we do with this one?","from":"developer"},{"body":"I believe this patch is ready (and has one +1) but needs a committer to review. [~n.v.harikrishna] was another ticket opened? ","from":"developer"},{"body":"[~jrwest] raised CASSANDRA-15741 for validation and/or fixing client timeout when mutation exceeds max size.","from":"developer"},{"body":"Hi [~n.v.harikrishna]. I've picked this back up and am getting it ready to commit. Thanks for your patience. \r\n\r\n \r\n\r\nI've squashed your branch here: [https://github.com/jrwest/cassandra/commits/14781-trunk.] I made a few minor changes along the way (I also re-reviewed since it had been a little bit since I had read the patch):\r\n\r\n \r\n * Modified {{CHANGES.txt}}\r\n * Modified {{IMutation#validateSize}} javadoc\r\n * Moved call to {{Keyspace.open}} into the catch block of {{BlockingReadRepairs#createRepairMutation}}. It was only used if we reached that block anyways.\r\n * Fixed whitespace formatting in {{MutationExceededMaxSizeException#prepareMessage}}\r\n\r\n \r\n\r\nI ran a build prior to these changes. The build looked good (better than trunk actually) and any failures do not seem related: [https://app.circleci.com/pipelines/github/jrwest/cassandra/4/workflows/e43918eb-40d2-45ad-80c3-dbeaa5ee186b]\r\n\r\n \r\n\r\nI've kicked off a new build with the changes above and with the squash performed: [https://app.circleci.com/pipelines/github/jrwest/cassandra/6/workflows/3c3f674e-db89-488a-bbd5-98f04de4fd0d]\r\n\r\nEDIT:\r\n\r\nI was slightly concerned about the failure in {{read_repair_test.py}}'s {{test_speculative_data_request}}. Looking closer at the test runs, its flaky and doesn't look like that flakiness could be related to the changes here (since the mutation sizes are static). \r\n\r\n \r\n\r\nI've also kicked off a Jenkins build for good measure: [https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/107/] \r\n\r\n \r\n\r\n ","from":"developer"},{"body":"Tests looked good. Any failures are suspected flaky tests and are unrelated. Committed as d3dadcd6f3bbde471e972f8332eb62de0f2d4aae. ","from":"developer"},{"body":"Thanks a lot [~jwest] for taking it forward!!","from":"developer"}],"created":"2018-09-21T19:29:34.000+0000","description":"When hitting [https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/commitlog/CommitLog.java#L256-L257], the log message produced does not help the operator track down what data is being written. At a minimum the keyspace and cfIds involved would be useful (and are available) – more detail might not be reasonable to include. ","issue_id":"13186708","key":"CASSANDRA-14781","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2020-05-04T18:33:40.000+0000","role":"fixed_distractor","summary":"Log message when mutation passed to CommitLog#add(Mutation) is too large is not descriptive enough"} {"case_id":"13189134","cluster":"DISTRACTOR-CASSANDRA-14801","comments":[{"body":"[~benedict] do you think this should block the first alpha or it can wait for beta?","created":"2019-08-29T20:59:24.746+0000"},{"body":"Is anyone working on this ticket? If not, I would like to work on it.","created":"2020-03-09T10:14:12.074+0000"},{"body":"Nobody is actively working on it, but this is one of the most deceptively complex tickets that needs to be accomplished before 4.0 is released. I can see you work at DataStax, so perhaps you have the time and skill to dedicate to this, but please be confident before you address it, and be willing to wait a while for a sufficient review. The class in which the change is needed has had numerous bugs (and in fact has inherent conceptually bugs wrt range movements that are mostly out of scope to address here), so a great deal of care is needed. Ideally this ticket would attempt to address some of the ugliness that permitted the bug, and _certainly_ needs to be accompanied by a sophisticated-ish randomised correctness test.","created":"2020-03-09T10:21:53.772+0000"},{"body":"Thank you for a comprehensive response, [~benedict]! I am quite new to C* 4.0 code, so it will take me some time to ramp up. If anyone has planned to work on this issue in the next 1-2 weeks, it probably makes sense for me to work on something else. Otherwise, I'd be happy to contribute.\r\n\r\nIn the latter case, in the next couple of days, I plan to read on how pending ranges are calculated and what changes since 3.11 introduced the bug. Then I'll write a test case that reproduces the issue.","created":"2020-03-09T12:28:10.439+0000"},{"body":"I reproduced the bug using simple [randomized test|https://gist.github.com/Gerrrr/f59dc5dedaf6501bb31d79b068244213] that creates a cluster of N nodes and adds, moves, and removes nodes from it. I also found another bug where {{TokenMetadata#calculatePendingRanges}} fails during move-affected replica calculation.\r\n\r\nPatch that addresses both problems ([link to PR|https://github.com/apache/cassandra/pull/495]):\r\n\r\n* [a534b2|https://github.com/Gerrrr/cassandra/commit/a534b2be9a653fb0cdda75043c0b79e481ca1701] adds tests that reproduce both bugs.\r\n* [bb36cd|https://github.com/Gerrrr/cassandra/commit/bb36cd09c164840ad9ab3231f1039c443f8f040c] fixes the newly discovered bug ({{test1Leave1Move}} in [a534b2|https://github.com/Gerrrr/cassandra/commit/a534b2be9a653fb0cdda75043c0b79e481ca1701]). There, the failure happens because for the same pending range we can include 2 replicas with the same endpoint and different ranges (violation of {{PendingRangeMaps Conflict.DUPLICATE}} policy). This happens because right now we include in {{PendingRangeMaps}} the entire new replica after leave for leave-affected ranges and only the pending part of it for move-affected ranges. This commit marks as pending only the missing part of the new replica.\r\n* [a33b03|https://github.com/Gerrrr/cassandra/commit/a33b032cd75044db3475505cd32e6afd6f98ad40] solves the original issue. Without this commit {{getPendingRanges}} can fail if the same replica covers more than 1 pending range. I think that this is a valid situation. Consider {{testLeave2}} in [a534b2|https://github.com/Gerrrr/cassandra/commit/a534b2be9a653fb0cdda75043c0b79e481ca1701]. In this case replica {{Full(127.0.0.1:7012,(-9,0])}} covers 2 ranges - {{(-9,-4]}} and {{(-4,0]}}. If we run the same test against 3.11 {{(-9, 0]}} is represented as {{\\{(-9, -4], (-4, 0]\\}}} and that possible representation would not trigger the bug in 4.0. However, I think that we should not base safety of {{getPendingRanges}} execution on the way a particular {{AbstractReplicationStrategy}} represents a range and we should allow duplicate entries while building pending {{RangesAtEndpoint}}.\r\n\r\nWith these changes, the randomized test hasn't failed after running for a while. As I mentioned, it is a very simplistic test, so any suggestions for improvement are welcome. I haven't included it in the PR, as it seems to be a good candidate to become flaky. Maybe it is worth adding as a {{long}} test?","created":"2020-03-27T14:10:55.953+0000"},{"body":"||branch||circleci||jenkins||\r\n|[trunk_14801|https://github.com/apache/cassandra/compare/trunk...Gerrrr:14801-4.0]|[circleci|https://circleci.com/gh/Gerrrr/workflows/cassandra/tree/14801-4.0]|[!https://ci-cassandra.apache.org/job/Cassandra-devbranch/13/badge/icon!|https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/13]|","created":"2020-04-03T07:15:36.150+0000"},{"body":"Thanks for raising that this is a thorny area of the code-base where specific care needs to be taken. There would probably be a lot of value in a class level comment on TokenMetaData.java indicating this for future people as they ramp on the code-base (and likely other areas of the code-base as well!).\r\n{quote}I can see you work at DataStax, so perhaps you have the time and skill to dedicate to this, but please be confident before you address it, and be willing to wait a while for a sufficient review.\r\n{quote}\r\nRegarding [earned authority|https://www.apache.org/theapacheway/]:\r\n{quote}their influence is based on publicly earned merit – what they contribute to the community. Merit lies with the individual, does not expire, is not influenced by employment status or employer\r\n{quote}\r\nPlease refrain from focusing on who is employed by whom and instead focus on their individual merit and contribution to the code-base in the future.","created":"2020-04-04T16:42:11.805+0000"},{"body":"bq. Please refrain from focusing on who is employed by whom and instead focus on their individual merit and contribution to the code-base in the future.\r\n\r\nIt is the case that new contributors can generally be assumed to have no knowledge of complex areas of the codebase. There is however precisely one company I am aware of with people who may secretly have the requisite knowledge and skill. There is also only one company with _full time_ contributors that were previously unknown to the community, who may have the time to either obtain the necessary skill, or to have weeks to dedicate to difficult tasks such as this.\r\n\r\nWere this to be a random individual, I could quite reasonably just direct them to another ticket. In this case I thought it was best only to issue a fair warning that this _might_ not be suitable but that I do not have sufficient knowledge to say for sure.\r\n\r\nI am perceiving a pattern on your part, of choosing to interpret my actions in an unwarrantedly negative light. Please check yourself, before you check others.","created":"2020-04-04T22:09:53.647+0000"},{"body":"[~fcofdezc] reviewed the PR and based on his suggestion I re-wrote the randomized test using QuickTheories. As the run time of that test ended comparable to the other unit tests, I included it in the suite as well ([a0246e|https://github.com/apache/cassandra/pull/495/commits/a0246e1e8ba481b98f76896e5005e9a8cc277586]).","created":"2020-04-14T16:29:51.766+0000"},{"body":"Thanks [~Gerrrr], this looks pretty good to me, with a few caveats regarding the QT test.\r\n\r\n * Move operations are not permitted when nodes have multiple tokens, so I think we can split the test into:\r\n ** multiple tokens per node with leave and bootstrap operations.\r\n ** single token nodes with leave, bootstrap and move operations.\r\n * Any node should be moved at most once per {{Cluster}}\r\n * The rf for each cluster was being hardcoded to 2, we ought to use {{Input.rf}} when constructing {{Cluster}}\r\n\r\nI've pushed a commit with those suggestions [here|https://github.com/beobal/cassandra/tree/14801-trunk] and kicked off CI [here|https://app.circleci.com/pipelines/github/beobal/cassandra?branch=14801-trunk].\r\n\r\nWould you mind taking a look at the suggestions?","created":"2020-09-02T18:30:33.802+0000"},{"body":"[~samt] Your suggestions look good! As a nitpick, I propose to replace the while-true loop with filtering out moving nodes and selecting one from what's left ([8d6dfc090da2a|https://github.com/Gerrrr/cassandra/commit/8d6dfc090da2a49dcc07c65029f33deeff7e186b]).","created":"2020-09-03T11:40:50.682+0000"},{"body":"Thanks [~Gerrrr], looks good as far as I'm concerned. I've pulled in your patch and I'm happy to commit pending CI and a second +1.\r\n\r\n\r\n||branch||Circle||\r\n|​[14801-trunk|https://github.com/beobal/cassandra/tree/14801-trunk]|[circle|https://app.circleci.com/pipelines/github/beobal/cassandra?branch=14801-trunk]​|","created":"2020-09-03T17:37:28.477+0000"},{"body":"+1","created":"2020-09-14T11:59:38.593+0000"},{"body":"and committed, thanks!","created":"2020-09-14T12:11:58.161+0000"}],"conversations":[{"body":"Correctness depended upon the narrowing to a {{Set}}, which we no longer do - we maintain a collection of all {{Replica}}. Our {{RangesAtEndpoint}} collection built by {{getPendingRanges}} can as a result contain the same endpoint multiple times; and our {{EndpointsForToken}} obtained by {{TokenMetadata.pendingEndpointsFor}} may fail to be constructed, resulting in cluster-wide failures for writes to the affected token ranges for the duration of the range movement.\r\n","from":"reporter","subject":"calculatePendingRanges no longer safe for multiple adjacent range movements"},{"body":"[~benedict] do you think this should block the first alpha or it can wait for beta?","from":"developer"},{"body":"Is anyone working on this ticket? If not, I would like to work on it.","from":"developer"},{"body":"Nobody is actively working on it, but this is one of the most deceptively complex tickets that needs to be accomplished before 4.0 is released. I can see you work at DataStax, so perhaps you have the time and skill to dedicate to this, but please be confident before you address it, and be willing to wait a while for a sufficient review. The class in which the change is needed has had numerous bugs (and in fact has inherent conceptually bugs wrt range movements that are mostly out of scope to address here), so a great deal of care is needed. Ideally this ticket would attempt to address some of the ugliness that permitted the bug, and _certainly_ needs to be accompanied by a sophisticated-ish randomised correctness test.","from":"developer"},{"body":"Thank you for a comprehensive response, [~benedict]! I am quite new to C* 4.0 code, so it will take me some time to ramp up. If anyone has planned to work on this issue in the next 1-2 weeks, it probably makes sense for me to work on something else. Otherwise, I'd be happy to contribute.\r\n\r\nIn the latter case, in the next couple of days, I plan to read on how pending ranges are calculated and what changes since 3.11 introduced the bug. Then I'll write a test case that reproduces the issue.","from":"developer"},{"body":"I reproduced the bug using simple [randomized test|https://gist.github.com/Gerrrr/f59dc5dedaf6501bb31d79b068244213] that creates a cluster of N nodes and adds, moves, and removes nodes from it. I also found another bug where {{TokenMetadata#calculatePendingRanges}} fails during move-affected replica calculation.\r\n\r\nPatch that addresses both problems ([link to PR|https://github.com/apache/cassandra/pull/495]):\r\n\r\n* [a534b2|https://github.com/Gerrrr/cassandra/commit/a534b2be9a653fb0cdda75043c0b79e481ca1701] adds tests that reproduce both bugs.\r\n* [bb36cd|https://github.com/Gerrrr/cassandra/commit/bb36cd09c164840ad9ab3231f1039c443f8f040c] fixes the newly discovered bug ({{test1Leave1Move}} in [a534b2|https://github.com/Gerrrr/cassandra/commit/a534b2be9a653fb0cdda75043c0b79e481ca1701]). There, the failure happens because for the same pending range we can include 2 replicas with the same endpoint and different ranges (violation of {{PendingRangeMaps Conflict.DUPLICATE}} policy). This happens because right now we include in {{PendingRangeMaps}} the entire new replica after leave for leave-affected ranges and only the pending part of it for move-affected ranges. This commit marks as pending only the missing part of the new replica.\r\n* [a33b03|https://github.com/Gerrrr/cassandra/commit/a33b032cd75044db3475505cd32e6afd6f98ad40] solves the original issue. Without this commit {{getPendingRanges}} can fail if the same replica covers more than 1 pending range. I think that this is a valid situation. Consider {{testLeave2}} in [a534b2|https://github.com/Gerrrr/cassandra/commit/a534b2be9a653fb0cdda75043c0b79e481ca1701]. In this case replica {{Full(127.0.0.1:7012,(-9,0])}} covers 2 ranges - {{(-9,-4]}} and {{(-4,0]}}. If we run the same test against 3.11 {{(-9, 0]}} is represented as {{\\{(-9, -4], (-4, 0]\\}}} and that possible representation would not trigger the bug in 4.0. However, I think that we should not base safety of {{getPendingRanges}} execution on the way a particular {{AbstractReplicationStrategy}} represents a range and we should allow duplicate entries while building pending {{RangesAtEndpoint}}.\r\n\r\nWith these changes, the randomized test hasn't failed after running for a while. As I mentioned, it is a very simplistic test, so any suggestions for improvement are welcome. I haven't included it in the PR, as it seems to be a good candidate to become flaky. Maybe it is worth adding as a {{long}} test?","from":"developer"},{"body":"||branch||circleci||jenkins||\r\n|[trunk_14801|https://github.com/apache/cassandra/compare/trunk...Gerrrr:14801-4.0]|[circleci|https://circleci.com/gh/Gerrrr/workflows/cassandra/tree/14801-4.0]|[!https://ci-cassandra.apache.org/job/Cassandra-devbranch/13/badge/icon!|https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/13]|","from":"developer"},{"body":"Thanks for raising that this is a thorny area of the code-base where specific care needs to be taken. There would probably be a lot of value in a class level comment on TokenMetaData.java indicating this for future people as they ramp on the code-base (and likely other areas of the code-base as well!).\r\n{quote}I can see you work at DataStax, so perhaps you have the time and skill to dedicate to this, but please be confident before you address it, and be willing to wait a while for a sufficient review.\r\n{quote}\r\nRegarding [earned authority|https://www.apache.org/theapacheway/]:\r\n{quote}their influence is based on publicly earned merit – what they contribute to the community. Merit lies with the individual, does not expire, is not influenced by employment status or employer\r\n{quote}\r\nPlease refrain from focusing on who is employed by whom and instead focus on their individual merit and contribution to the code-base in the future.","from":"developer"},{"body":"bq. Please refrain from focusing on who is employed by whom and instead focus on their individual merit and contribution to the code-base in the future.\r\n\r\nIt is the case that new contributors can generally be assumed to have no knowledge of complex areas of the codebase. There is however precisely one company I am aware of with people who may secretly have the requisite knowledge and skill. There is also only one company with _full time_ contributors that were previously unknown to the community, who may have the time to either obtain the necessary skill, or to have weeks to dedicate to difficult tasks such as this.\r\n\r\nWere this to be a random individual, I could quite reasonably just direct them to another ticket. In this case I thought it was best only to issue a fair warning that this _might_ not be suitable but that I do not have sufficient knowledge to say for sure.\r\n\r\nI am perceiving a pattern on your part, of choosing to interpret my actions in an unwarrantedly negative light. Please check yourself, before you check others.","from":"developer"},{"body":"[~fcofdezc] reviewed the PR and based on his suggestion I re-wrote the randomized test using QuickTheories. As the run time of that test ended comparable to the other unit tests, I included it in the suite as well ([a0246e|https://github.com/apache/cassandra/pull/495/commits/a0246e1e8ba481b98f76896e5005e9a8cc277586]).","from":"developer"},{"body":"Thanks [~Gerrrr], this looks pretty good to me, with a few caveats regarding the QT test.\r\n\r\n * Move operations are not permitted when nodes have multiple tokens, so I think we can split the test into:\r\n ** multiple tokens per node with leave and bootstrap operations.\r\n ** single token nodes with leave, bootstrap and move operations.\r\n * Any node should be moved at most once per {{Cluster}}\r\n * The rf for each cluster was being hardcoded to 2, we ought to use {{Input.rf}} when constructing {{Cluster}}\r\n\r\nI've pushed a commit with those suggestions [here|https://github.com/beobal/cassandra/tree/14801-trunk] and kicked off CI [here|https://app.circleci.com/pipelines/github/beobal/cassandra?branch=14801-trunk].\r\n\r\nWould you mind taking a look at the suggestions?","from":"developer"},{"body":"[~samt] Your suggestions look good! As a nitpick, I propose to replace the while-true loop with filtering out moving nodes and selecting one from what's left ([8d6dfc090da2a|https://github.com/Gerrrr/cassandra/commit/8d6dfc090da2a49dcc07c65029f33deeff7e186b]).","from":"developer"},{"body":"Thanks [~Gerrrr], looks good as far as I'm concerned. I've pulled in your patch and I'm happy to commit pending CI and a second +1.\r\n\r\n\r\n||branch||Circle||\r\n|​[14801-trunk|https://github.com/beobal/cassandra/tree/14801-trunk]|[circle|https://app.circleci.com/pipelines/github/beobal/cassandra?branch=14801-trunk]​|","from":"developer"},{"body":"+1","from":"developer"},{"body":"and committed, thanks!","from":"developer"}],"created":"2018-10-03T11:21:54.000+0000","description":"Correctness depended upon the narrowing to a {{Set}}, which we no longer do - we maintain a collection of all {{Replica}}. Our {{RangesAtEndpoint}} collection built by {{getPendingRanges}} can as a result contain the same endpoint multiple times; and our {{EndpointsForToken}} obtained by {{TokenMetadata.pendingEndpointsFor}} may fail to be constructed, resulting in cluster-wide failures for writes to the affected token ranges for the duration of the range movement.\r\n","issue_id":"13189134","key":"CASSANDRA-14801","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2020-09-14T12:11:58.000+0000","role":"fixed_distractor","summary":"calculatePendingRanges no longer safe for multiple adjacent range movements"} {"case_id":"13225337","cluster":"DISTRACTOR-CASSANDRA-15072","comments":[{"body":"Are you seeing incomplete results like this in a real cluster? If so, what consistency level are you reading and writing at?\r\n\r\nThe ccm script you have here _does_ return incomplete results, but it’s also writing and reading at CL ONE (the cqlsh default), so that’s not unexpected. I modified the script here to read and write at QUORUM, and haven't gotten any incomplete results.","created":"2019-04-01T20:55:47.577+0000"},{"body":"Yes, we saw a lot of incomplete results in a real cluster. We read and write at quorum.\r\n\r\nOops, you are right about my repro. I modified the steps to reproduce it at quorum (I upgraded two out of three nodes instead of just one, changed the reads/writes to be quorum, and connected to node 3 to perform the reproduction query).","created":"2019-04-01T21:49:32.057+0000"},{"body":"It seems with my updated steps that only the first query against node3 reproduces it. After that it returns both rows. If you restart node3, it reproduces it again for one query. This is not the behavior we experienced in production (i.e. the problem did not go away). I wonder if I have actually reproduced our issue or not...","created":"2019-04-01T21:54:02.984+0000"},{"body":"Muir's colleague here:\r\n\r\nI get 100% reproducibility for repeated queries with the following changes:\r\n # Create the keyspace with replication_factor 2\r\n # Do the inserts with CONSISTENCY ALL\r\n # Upgrade the two nodes that contain data (node1, node3); keep the node that does not contain any sstables for test.test (node2) back at 2.1.17\r\n\r\nAfter those steps, I get full results 100% of the time when querying node1 and node3, and truncated results 100% of the time when querying node2.\r\n\r\nThis is using cqlsh as packaged with 3.11.4.\r\n\r\n ","created":"2019-04-02T00:30:00.914+0000"},{"body":"Ok, I can repro your issue with the updated script. It looks like you’re hitting a commit log bug that was introduced in 2.1 and fixed in 3.0 (CASSANDRA-13987) If you drain node 1 & 2 before shutting them down, this should stop happening.\r\n\r\nI’d also expected putting a sleep larger than the commit log sync interval before shutting down node 1 would fix the problem, but it didn’t. I’m still looking at why that is.\r\n\r\nWhen you say:\r\n{quote}When all nodes were upgraded (before upgrading sstables), we stopped getting incomplete results\r\n{quote}\r\ndo you mean data you'd inserted before the upgrade reappeared?","created":"2019-04-02T00:47:29.654+0000"},{"body":"Please see the attached [^eriksw-repro.sh], which includes aggressive flushing, draining, and deleting commit logs when stopped to ensure they play no part. With these steps, the truncated results behavior when querying node2 (the un-upgraded node) is 100% reproducible for me for an unlimited number of queries with CONSISTENCY ALL.\r\n{quote}do you mean data you'd inserted before the upgrade reappeared?\r\n{quote}\r\nYes. After the last node was upgraded to Cassandra 3.11.4, we no longer saw truncated results regardless of which node we queried. All data that should have been in the results was correctly returned for all queries after that point.","created":"2019-04-02T02:08:32.894+0000"},{"body":"This is a great repro script, thanks. \r\n\r\nA couple of observations:\r\n * test.test has 2 columns, and uses compact storage, which shouldn’t be possible\r\n * node1 & node3 are the replicas of the missing partition (we’re querying from the un-upgraded node2, for those following along).\r\n * doing a point read ({{select * from test.test where id=‘1’;}}) returns the expected partition\r\n * using LIMIT 2 instead of PAGING 2 has the same problem\r\n * LIMIT 3 returns a partial row: {{1 | there | null}}\r\n * LIMIT 4 returns the entire row: {{1 | there |  hi}}\r\n\r\nTables with compact storage can only have a single column, so you shouldn’t be able to create a compact storage table with 2 columns. Instead of throwing an error though, it seems like it just silently treats the table as a normal table. This might be why no one has noticed that our ddl validation is broken.\r\n\r\nIt looks like the mixed mode read path is treating the table as a proper compact storage table though, and treating each cell as a row, which is why you see partial rows start to appear as you increase the limit. If you remove compact storage from the ddl, or only use a single column, everything works normally.\r\n\r\nI'll think on the best way to address this.","created":"2019-04-02T18:39:48.670+0000"},{"body":"{quote}Tables with compact storage can only have a single column, so you shouldn’t be able to create a compact storage table with 2 columns.\r\n{quote}\r\nAccording to [http://cassandra.apache.org/doc/latest/cql/ddl.html] that restriction is only for tables with clustering columns:\r\n{quote}if a compact table has at least one clustering column, then it must have exactly one column outside of the primary key ones.\r\n{quote}\r\nWe have a lot of tables created from thrift (compact storage) that do not have clustering columns and have > 1 column in the CQL schema.","created":"2019-04-02T19:07:02.405+0000"},{"body":"[https://docs.datastax.com/en/cql/3.3/cql/cql_using/useCompactStorage.html] also explicitly states the implied inverse:\r\n{quote}\r\nA compact table with a primary key that is not compound can have multiple columns that are not part of the primary key.\r\n{quote}","created":"2019-04-02T19:20:11.946+0000"},{"body":"Huh, I did not know that. I guess that makes sense though. So then this is just an upgrade bug.","created":"2019-04-02T20:07:34.107+0000"},{"body":"Thanks for helping us investigate this issue. Do you think you understand the exact cause at this point?\r\n\r\n{quote}It looks like the mixed mode read path is treating the table as a proper compact storage table though, and treating each cell as a row\r\n{quote}\r\nDoes \"mixed mode\" refer to the mixed 2.X <=> 3.X cassandra versions?\r\n\r\nFrom a high level it sounds like a 2.X coordinator and a 3.X replica have some confusion regarding compact storage cells vs. rows, and how many are needed to satisfy a limit or page quota. Is that still what you think is going on?","created":"2019-04-02T20:40:55.512+0000"},{"body":"No problem. Yes mixed mode just means you're upgrading your cluster.\r\n\r\nI don't know the exact cause, but you've summarized what I think is probably happening. Specifically the legacy read path on the 3.0 nodes is probably always interpreting single cells as rows for compact storage tables, even ones without clustering columns.","created":"2019-04-02T22:15:22.546+0000"},{"body":"|[3.0|https://github.com/bdeggleston/cassandra/tree/15072-3.0]|[tests|https://circleci.com/gh/bdeggleston/workflows/cassandra/tree/15072-3.0]|\r\n|[3.11|https://github.com/bdeggleston/cassandra/tree/15072-3.11]|[tests|https://circleci.com/workflow-run/4567dbed-be97-49e5-8c82-66e320e074ca]|\r\n\r\n[~beobal] do you have time to review this? It seems to be related to CASSANDRA-11087.\r\n\r\nA few notes:\r\n * From what I can tell, returning a row per cell is the right thing to do in this case, so I'm using a modified result counter to only going entire partitions in this specific case. However, I'm not familiar enough with all the dark corners of the 2.1 storage engine to be sure that's appropriate, or won't break something else.\r\n * Doing a point read with the partition key also returns a row per cell, but works correctly because the 2.2 coordinator seems to just discard the limit in that case.\r\n * If you're not familiar with the in-jvm dtests yet, and want to run the one in this patch, you'll want to run {{ant dtest-jar}} on this branch and [this 2.2 branch|https://github.com/bdeggleston/cassandra/tree/15078-2.2], and put the 2.2 dtest jar in the 3.0 build directory.\r\n * -CircleCI seems to be behind picking up new branches to test, but I'll update this with links to the workflows once it catches up.-","created":"2019-04-04T23:28:45.514+0000"},{"body":"[~bdeggleston] sure, I'll review asap","created":"2019-04-05T07:34:52.559+0000"},{"body":"This looks safe to me wrt to \"the dark corners\" as the new counter is only used in this very specific use case, so if the CI looks good I'm +1 on the patch. ","created":"2019-04-05T15:16:27.235+0000"},{"body":"Committed to 3.0 as [d27c3ad0d2d006a5f156f0a2f2a24286d31c5069|https://github.com/apache/cassandra/commit/d27c3ad0d2d006a5f156f0a2f2a24286d31c5069] and merged up. Thanks!","created":"2019-04-24T22:39:17.703+0000"}],"conversations":[{"body":"Hello\r\n\r\nDuring an upgrade from 2.1.17 to 3.11.4, our application starting getting back incomplete results for range queries. When all nodes were upgraded (before upgrading sstables), we stopped getting incomplete results. I was able to reproduce it and listed steps below. It seems to require the random partitioner and compact storage to reproduce reliably. It also reproduces coming from 2.1.21 and 2.2.14. You seem to get the bad behavior when an old node is your coordinator and it has to talk to an upgraded replica.\r\n{noformat}\r\nccm create test -v 2.1.17 -n 3\r\nccm updateconf 'partitioner: org.apache.cassandra.dht.RandomPartitioner'\r\nccm node1 updateconf 'initial_token: 0'\r\nccm node2 updateconf 'initial_token: 56713727820156410577229101238628035242'\r\nccm node3 updateconf 'initial_token: 113427455640312821154458202477256070484'\r\nccm start\r\n\r\nccm node1 cqlsh < 3.11.4 upgrade"},{"body":"Are you seeing incomplete results like this in a real cluster? If so, what consistency level are you reading and writing at?\r\n\r\nThe ccm script you have here _does_ return incomplete results, but it’s also writing and reading at CL ONE (the cqlsh default), so that’s not unexpected. I modified the script here to read and write at QUORUM, and haven't gotten any incomplete results.","from":"developer"},{"body":"Yes, we saw a lot of incomplete results in a real cluster. We read and write at quorum.\r\n\r\nOops, you are right about my repro. I modified the steps to reproduce it at quorum (I upgraded two out of three nodes instead of just one, changed the reads/writes to be quorum, and connected to node 3 to perform the reproduction query).","from":"developer"},{"body":"It seems with my updated steps that only the first query against node3 reproduces it. After that it returns both rows. If you restart node3, it reproduces it again for one query. This is not the behavior we experienced in production (i.e. the problem did not go away). I wonder if I have actually reproduced our issue or not...","from":"developer"},{"body":"Muir's colleague here:\r\n\r\nI get 100% reproducibility for repeated queries with the following changes:\r\n # Create the keyspace with replication_factor 2\r\n # Do the inserts with CONSISTENCY ALL\r\n # Upgrade the two nodes that contain data (node1, node3); keep the node that does not contain any sstables for test.test (node2) back at 2.1.17\r\n\r\nAfter those steps, I get full results 100% of the time when querying node1 and node3, and truncated results 100% of the time when querying node2.\r\n\r\nThis is using cqlsh as packaged with 3.11.4.\r\n\r\n ","from":"developer"},{"body":"Ok, I can repro your issue with the updated script. It looks like you’re hitting a commit log bug that was introduced in 2.1 and fixed in 3.0 (CASSANDRA-13987) If you drain node 1 & 2 before shutting them down, this should stop happening.\r\n\r\nI’d also expected putting a sleep larger than the commit log sync interval before shutting down node 1 would fix the problem, but it didn’t. I’m still looking at why that is.\r\n\r\nWhen you say:\r\n{quote}When all nodes were upgraded (before upgrading sstables), we stopped getting incomplete results\r\n{quote}\r\ndo you mean data you'd inserted before the upgrade reappeared?","from":"developer"},{"body":"Please see the attached [^eriksw-repro.sh], which includes aggressive flushing, draining, and deleting commit logs when stopped to ensure they play no part. With these steps, the truncated results behavior when querying node2 (the un-upgraded node) is 100% reproducible for me for an unlimited number of queries with CONSISTENCY ALL.\r\n{quote}do you mean data you'd inserted before the upgrade reappeared?\r\n{quote}\r\nYes. After the last node was upgraded to Cassandra 3.11.4, we no longer saw truncated results regardless of which node we queried. All data that should have been in the results was correctly returned for all queries after that point.","from":"developer"},{"body":"This is a great repro script, thanks. \r\n\r\nA couple of observations:\r\n * test.test has 2 columns, and uses compact storage, which shouldn’t be possible\r\n * node1 & node3 are the replicas of the missing partition (we’re querying from the un-upgraded node2, for those following along).\r\n * doing a point read ({{select * from test.test where id=‘1’;}}) returns the expected partition\r\n * using LIMIT 2 instead of PAGING 2 has the same problem\r\n * LIMIT 3 returns a partial row: {{1 | there | null}}\r\n * LIMIT 4 returns the entire row: {{1 | there |  hi}}\r\n\r\nTables with compact storage can only have a single column, so you shouldn’t be able to create a compact storage table with 2 columns. Instead of throwing an error though, it seems like it just silently treats the table as a normal table. This might be why no one has noticed that our ddl validation is broken.\r\n\r\nIt looks like the mixed mode read path is treating the table as a proper compact storage table though, and treating each cell as a row, which is why you see partial rows start to appear as you increase the limit. If you remove compact storage from the ddl, or only use a single column, everything works normally.\r\n\r\nI'll think on the best way to address this.","from":"developer"},{"body":"{quote}Tables with compact storage can only have a single column, so you shouldn’t be able to create a compact storage table with 2 columns.\r\n{quote}\r\nAccording to [http://cassandra.apache.org/doc/latest/cql/ddl.html] that restriction is only for tables with clustering columns:\r\n{quote}if a compact table has at least one clustering column, then it must have exactly one column outside of the primary key ones.\r\n{quote}\r\nWe have a lot of tables created from thrift (compact storage) that do not have clustering columns and have > 1 column in the CQL schema.","from":"developer"},{"body":"[https://docs.datastax.com/en/cql/3.3/cql/cql_using/useCompactStorage.html] also explicitly states the implied inverse:\r\n{quote}\r\nA compact table with a primary key that is not compound can have multiple columns that are not part of the primary key.\r\n{quote}","from":"developer"},{"body":"Huh, I did not know that. I guess that makes sense though. So then this is just an upgrade bug.","from":"developer"},{"body":"Thanks for helping us investigate this issue. Do you think you understand the exact cause at this point?\r\n\r\n{quote}It looks like the mixed mode read path is treating the table as a proper compact storage table though, and treating each cell as a row\r\n{quote}\r\nDoes \"mixed mode\" refer to the mixed 2.X <=> 3.X cassandra versions?\r\n\r\nFrom a high level it sounds like a 2.X coordinator and a 3.X replica have some confusion regarding compact storage cells vs. rows, and how many are needed to satisfy a limit or page quota. Is that still what you think is going on?","from":"developer"},{"body":"No problem. Yes mixed mode just means you're upgrading your cluster.\r\n\r\nI don't know the exact cause, but you've summarized what I think is probably happening. Specifically the legacy read path on the 3.0 nodes is probably always interpreting single cells as rows for compact storage tables, even ones without clustering columns.","from":"developer"},{"body":"|[3.0|https://github.com/bdeggleston/cassandra/tree/15072-3.0]|[tests|https://circleci.com/gh/bdeggleston/workflows/cassandra/tree/15072-3.0]|\r\n|[3.11|https://github.com/bdeggleston/cassandra/tree/15072-3.11]|[tests|https://circleci.com/workflow-run/4567dbed-be97-49e5-8c82-66e320e074ca]|\r\n\r\n[~beobal] do you have time to review this? It seems to be related to CASSANDRA-11087.\r\n\r\nA few notes:\r\n * From what I can tell, returning a row per cell is the right thing to do in this case, so I'm using a modified result counter to only going entire partitions in this specific case. However, I'm not familiar enough with all the dark corners of the 2.1 storage engine to be sure that's appropriate, or won't break something else.\r\n * Doing a point read with the partition key also returns a row per cell, but works correctly because the 2.2 coordinator seems to just discard the limit in that case.\r\n * If you're not familiar with the in-jvm dtests yet, and want to run the one in this patch, you'll want to run {{ant dtest-jar}} on this branch and [this 2.2 branch|https://github.com/bdeggleston/cassandra/tree/15078-2.2], and put the 2.2 dtest jar in the 3.0 build directory.\r\n * -CircleCI seems to be behind picking up new branches to test, but I'll update this with links to the workflows once it catches up.-","from":"developer"},{"body":"[~bdeggleston] sure, I'll review asap","from":"developer"},{"body":"This looks safe to me wrt to \"the dark corners\" as the new counter is only used in this very specific use case, so if the CI looks good I'm +1 on the patch. ","from":"developer"},{"body":"Committed to 3.0 as [d27c3ad0d2d006a5f156f0a2f2a24286d31c5069|https://github.com/apache/cassandra/commit/d27c3ad0d2d006a5f156f0a2f2a24286d31c5069] and merged up. Thanks!","from":"developer"}],"created":"2019-04-01T17:35:19.000+0000","description":"Hello\r\n\r\nDuring an upgrade from 2.1.17 to 3.11.4, our application starting getting back incomplete results for range queries. When all nodes were upgraded (before upgrading sstables), we stopped getting incomplete results. I was able to reproduce it and listed steps below. It seems to require the random partitioner and compact storage to reproduce reliably. It also reproduces coming from 2.1.21 and 2.2.14. You seem to get the bad behavior when an old node is your coordinator and it has to talk to an upgraded replica.\r\n{noformat}\r\nccm create test -v 2.1.17 -n 3\r\nccm updateconf 'partitioner: org.apache.cassandra.dht.RandomPartitioner'\r\nccm node1 updateconf 'initial_token: 0'\r\nccm node2 updateconf 'initial_token: 56713727820156410577229101238628035242'\r\nccm node3 updateconf 'initial_token: 113427455640312821154458202477256070484'\r\nccm start\r\n\r\nccm node1 cqlsh < 3.11.4 upgrade"} {"case_id":"13239212","cluster":"DISTRACTOR-CASSANDRA-15158","comments":[{"body":"So just skimming this, it looks like it's the right approach. We should add some tests though, think through what we want to do if we can't get the schema to converge, as well as leave an \"escape hatch\" if we need to start up and are willing to skip waiting on schema agreement.","created":"2020-04-23T20:27:04.335+0000"},{"body":"I also wanted to point out that the underlying issues causing this problem can lead to correctness issues or data loss in some scenarios. Since we don't currently confirm that any of the in flight migration tasks have been completed and applied, it's possible for us to not receive _any_ schema responses, and begin bootstrapping without a schema. In this case, bootstrap would complete immediately, because the node believes there is nothing to stream. When the node later receives the schema, it will begin serving reads and writes with no data. ","created":"2020-04-23T20:30:13.054+0000"},{"body":"Hi [~bdeggleston],\r\n\r\nplease review this one [https://github.com/smiklosovic/cassandra/tree/CASSANDRA-15158]\r\n\r\nSorry for the fuss but this seems to be more problematic than I was thinking initially. \r\n\r\nI do not think that we have to _check_ that respective migration request is successful or not, we should just check it all schemas match ...\r\n\r\nThe logic is based on repeated checking if schemas agree or not and only in case all schemas are equal we proceed, otherwise an exception is thrown (this might be reworked in such sense that we just log this). There is a global timeout for this check, it will be obvious from reading the code.\r\n\r\nThere is a test added too, because of the nature of the test, I had to \"inject\" these callbacks into respective methods so I could modify their default behaviour and test their state. \r\n\r\nThis is against 3.11, we need to backport this to 3.0 for our customer.","created":"2020-04-29T16:29:41.841+0000"},{"body":"Thanks [~stefan.miklosovic], I'm _should_ have some time to review next week.","created":"2020-04-29T16:38:46.692+0000"},{"body":"code for 3.0 [https://github.com/smiklosovic/cassandra/tree/CASSANDRA-15158-cassandra-3]","created":"2020-05-04T13:54:02.490+0000"},{"body":"for 3.11\r\n\r\n \r\n\r\n[https://github.com/smiklosovic/cassandra/tree/CASSANDRA-15158]\r\n\r\n \r\n\r\nfor 3.0\r\n\r\n \r\n\r\nhttps://github.com/smiklosovic/cassandra/tree/CASSANDRA-15158-cassandra-3","created":"2020-05-07T06:28:05.999+0000"},{"body":"It's likely we'll want to fix this in 3.0 and up, so I'll review the 3.0 version to start with and we can go from there.\r\n\r\nI haven't completed my review yet, but there are some structural and design issues we should address up front.\r\n\r\n*Structural issues*\r\n\r\nFirst, the {{waitForSchema*}} methods should live in the MigrationManager. This will prevent you from setting an updated status if you're sending out additional schema pulls, but we can revisit that later if we think it's neccesary.\r\n\r\nSecond, instantiating MigrationTaskCallbacks in StorageService/MigrationManager and passing it into MigrationTask is a little awkward. I'd prefer if the callback remained an anonymous class. We can communicate endpoint to send schema pulls to with an inetaddress argument, and we need to rethink what `isRunningForcibly` is doing and why. First, we shouldn't be adding it to {{IAsyncCallback}} for this narrow use case. Next, it's use seems to be changing how schema pulls actually work, but only in a test environment, which is something we should avoid.\r\n\r\n*Design issues*\r\n\r\nThis doesn't deal with multiple schema versions. If a node joins, and there are 2 or more schema versions floating around, it will only wait until it has _some_ schema to begin bootstrapping, not all. Related to this, we also need a plan for unreachable schema versions. For instance, if a single node is reporting a schema version that no one else has, but the node is unreachable, what do we do?\r\n\r\nNext, I like how this limits the number of messages sent to a given endpoint, but we should also limit the number of messages we send out for a given schema version. If we have a large cluster, and all nodes are reporting the same version, we don't need to ask every node for it's schema.","created":"2020-05-07T21:17:31.088+0000"},{"body":"Hi [~bdeggleston],\r\n\r\ncommenting on design issues, I am not completely sure if these issues you are talking about are related to this patch or they are already existing? We could indeed focus on the points you raised but it seems to me that the current (comitted) code is worse without this patch than with as I guess these problems are already there?\r\n\r\nIsn't the goal here to have all nodes on same versions? Isn't the very fact that there are multiple versions pretty strange to begin with so we should not even try to join a node if they mismatch hence there is nothing to deal with in the first place? \r\n{quote}It will only wait until it has _some_ schema to begin bootstrapping, not all\r\n{quote}\r\nThis is the most likely not true unless I am not getting something. The node to be bootstrapped will never advance in doing so unless all nodes have same versions. \r\n{quote} For instance, if a single node is reporting a schema version that no one else has, but the node is unreachable, what do we do?\r\n{quote}\r\nWe should fail whole bootstrapping and one should go and fix it.\r\n{quote}For instance, if a single node is reporting a schema version that no one else has, but the node is unreachable, what do we do?\r\n{quote}\r\nHow can a node report its schema while being unreachable?\r\n{quote}Next, I like how this limits the number of messages sent to a given endpoint, but we should also limit the number of messages we send out for a given schema version. If we have a large cluster, and all nodes are reporting the same version, we don't need to ask every node for it's schema.\r\n{quote}\r\n \r\n\r\nGot you, this might be tracked.\r\n\r\n \r\n\r\nWhen it comes to testing, I admit that adding isRunningForcibly method feels like a hack but I had very hard time to test this stuff out. It was basically the only reasonable way possible at the time I was coding it, if you know of more better version, please tell me otherwise I am not sure what might be better here and we could stick with this for a time being? The whole testing methodology was based on these callbacks and checking their inner state which results into having a methods which are accepting them so we can elaborate on their state. Without \"injecting\" them from outside, I would not be able to do that.","created":"2020-05-08T08:32:47.418+0000"},{"body":"It seems to me that one aspect of the PR was overlooked so I just iterate on that one. The mechanim how to not flood nodes with schema pull messages is incorporated in the loop over callbacks. If you notice it, there are sleeps of various lenghts based on a request being already sent or not. This sleep will actually \"delay\" the next schema pull from the other node because during this time of a sleep, some schema could come from the node we just sent a message to so on the next iteration when another node is compared on schema equality, it may happen that there is not any need to pull it anymore because they are on par. Hence we are not blindly sending messages to all nodes.\r\n If there are some discrepancies, there is the global timeout set after which whole bootstrapping process will be evaluated as errorneous and (in the current code) we throw a ConfigurationException. This behaviour might be relaxed but I consider it more appropriate to just throw it there.","created":"2020-05-08T12:48:19.232+0000"},{"body":"{quote}commenting on design issues, I am not completely sure if these issues you are talking about are related to this patch or they are already existing? We could indeed focus on the points you raised but it seems to me that the current (comitted) code is worse without this patch than with as I guess these problems are already there?\r\nIsn't the goal here to have all nodes on same versions? Isn't the very fact that there are multiple versions pretty strange to begin with so we should not even try to join a node if they mismatch hence there is nothing to deal with in the first place?\r\n{quote}\r\nWhen there are schema changes, it's not strange at all for there to be multiple schema versions in the cluster before they converge. We also don't forbid making schema changes while changing cluster topology, so this would be something we should expect to encounter, although I would expect it to happen infrequently. Since bootstrap doesn't stream keyspaces it doesn't know about, this could create a window of data loss. Since the goal of this ticket is to wait for schema to converge before starting bootstrap, we should deal with edge cases like this. Also, I believe there have been bugs that caused a lot of schema change activity when nodes bootstrap, so depending on what exactly you're doing\r\n{quote}How can a node report its schema while being unreachable?\r\n{quote}\r\nSchema versions are gossiped. So a node might gossip a new schema version then become unreachable. The bootstrapping node would learn about this new version via gossip, but be unable to contact it.\r\n{quote}> admit that adding isRunningForcibly method feels like a hack but I had very hard time to test this stuff out.\r\n{quote}\r\nI'll look into how testing can be improved.\r\n{quote}> This is the most likely not true unless I am not getting something. The node to be bootstrapped will never advance in doing so unless all nodes have same versions.\r\n{quote}\r\nAh, yes you're right. Althought waiting for all nodes to arrive at the same schema version isn't neccesary, we just need to receive and merge at least one schema pull from every schema version in the cluster.","created":"2020-05-08T17:03:57.768+0000"},{"body":"Hi Blake,\r\n\r\n \r\n\r\nbecause of your very helpful explanation I was able to put together yet another version of the solution to this problem. You will find it here\r\n\r\n[https://github.com/apache/cassandra/pull/628]\r\n\r\nThanks for the review in advance","created":"2020-06-11T21:44:34.307+0000"},{"body":"I've reworked this a bit more [here|https://github.com/bdeggleston/cassandra/tree/15158-coordinator]. It's now pretty self contained, has some fairly granular unit tests, and fixes a few functional things. Can you take a look [~stefan.miklosovic] and let me know what you think? [~aleksey], can you also take a look / review? The basic idea is that we now track which schema versions exist in gossip, and which endpoints are reporting them, then block bootstrap until we've received a schema for each version. It also tracks how many outstanding migration requests we have per version so we don't send out thousands.\r\n\r\nAlso, what are your opinions of this going into 3.x? I think I'd lean towards putting it in, since it eliminates a scenario where data loss could occur and will shave a few hours off of adding/replacing nodes in large clusters. On the other hand, since it reduces the amount of migration requests sent out on bootstrap, any bugs determining if we've received sufficient schema data could make data loss _more_ likely.","created":"2020-07-22T19:07:33.791+0000"},{"body":"In general looks good to me in spite of getting lost a bit on the actual schema migration response:\r\n\r\n \r\n\r\n \r\n{code:java}\r\nFuture response(Collection mutations)\r\n{\r\n synchronized (info)\r\n {\r\n if (shouldApplySchemaFrom(endpoint, info))\r\n mergeSchemaFrom(endpoint, mutations);\r\n return pullComplete(endpoint, info, true); /// why?\r\n }\r\n}\r\n{code}\r\n \r\n\r\nI am not completely sure why are we pulling again here. I would rewrite the whole solution in a such way that this Callable just does one thing on a successful response (merging of a schema) and the actual \"retry\" would be handled from outside. The reader has to make quite a mental exercise to visualise that this callback might actually call another callback in it until some \"version\" is completed etc ... At least for me, it was quite tedious to track.\r\n\r\nI understand the motivation behind that but callback should just do one task and thats it, it shouldnt be responsible for potentially scheduling another callback recursively until some conditions on some signals or what have you are met ... just my 2 cents.\r\n\r\nThere is also this:\r\n\r\n \r\n{code:java}\r\nsynchronized Future reportEndpointVersion(InetAddress endpoint, UUID version)\r\n{\r\n UUID current = endpointVersions.get(endpoint);\r\n if (current != null && current.equals(version))\r\n return FINISHED_FUTURE;\r\n\r\n VersionInfo info = versionInfo.computeIfAbsent(version, VersionInfo::new);\r\n if (isLocalVersion(version))\r\n info.markReceived();\r\n info.endpoints.add(endpoint);\r\n info.requestQueue.add(endpoint);\r\n endpointVersions.put(endpoint, version);\r\n\r\n removeEndpointFromVersion(endpoint, current); ///// why?\r\n return maybePullSchema(info);\r\n}\r\n{code}\r\n \r\n\r\nTBH that is quite counterintuitive too, maybe renaming of that method would help.\r\n\r\n \r\n\r\nThe test has failed for me (repeatedly):\r\n\r\n \r\n\r\n \r\n{code:java}\r\njava.lang.AssertionError: java.lang.AssertionError: \r\nExpected :2\r\nActual   :0\r\n at org.junit.Assert.fail(Assert.java:92) at org.junit.Assert.failNotEquals(Assert.java:689) at org.junit.Assert.assertEquals(Assert.java:127) at org.junit.Assert.assertEquals(Assert.java:514) at org.junit.Assert.assertEquals(Assert.java:498) at org.apache.cassandra.service.MigrationCoordinatorTest.testWeKeepSendingRequests(MigrationCoordinatorTest.java:278) at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method) at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62) at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43) at java.lang.reflect.Method.invoke(Method.java:498) at org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:44) at org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15) at org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:41) at org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:20) at org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:28) at org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:31) at org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:70) at org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:44) at org.junit.runners.ParentRunner.runChildren(ParentRunner.java:180) at org.junit.runners.ParentRunner.access$000(ParentRunner.java:41) at org.junit.runners.ParentRunner$1.evaluate(ParentRunner.java:173) at org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:28) at org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:31) at org.junit.runners.ParentRunner.run(ParentRunner.java:220) at org.junit.runner.JUnitCore.run(JUnitCore.java:159) at com.intellij.junit4.JUnit4IdeaTestRunner.startRunnerWithArgs(JUnit4IdeaTestRunner.java:68) at com.intellij.rt.junit.IdeaTestRunner$Repeater.startRunnerWithArgs(IdeaTestRunner.java:33) at com.intellij.rt.junit.JUnitStarter.prepareStreamsAndStart(JUnitStarter.java:230) at com.intellij.rt.junit.JUnitStarter.main(JUnitStarter.java:58)\r\n{code}\r\n ","created":"2020-07-27T11:00:43.527+0000"},{"body":"hi [~bdeggleston], have you had a chance to reflect the above issues?","created":"2020-07-31T19:28:58.619+0000"},{"body":"{quote}\r\nI am not completely sure why are we pulling again here. I would rewrite the whole solution in a such way that this Callable just does one thing on a successful response (merging of a schema) and the actual \"retry\" would be handled from outside. The reader has to make quite a mental exercise to visualise that this callback might actually call another callback in it until some \"version\" is completed etc ... At least for me, it was quite tedious to track.\r\n\r\n\r\n{quote}\r\nIn the case of a successful pull, we won't pull again. Response and fail both call pullComplete, but an additional pull is only called if it's called from fail.\r\n\r\nI get that this can be a bit difficult to follow, but I'm not sure there's a better approach, given the schema pulls are completely event driven during normal runtime. If we miss a schema change during normal runtime (not bootstrap), there's nothing waiting on schema convergence that would enable us to retry from the outside.\r\n\r\nThere is a periodic task that pulls schema for outstanding versions that don't have any in flight requests^[1]^, but it only runs once a minute, and we need to be more proactive about learning about schema updates since we'll be unable to serve some reads and writes until we're up to date.\r\n{quote}TBH that is quite counterintuitive too\r\n{quote}\r\nCould you expand on what's counterintuitive about it? If the endpoint's schema version has changed, we need to disassociate it with it's previously reported version. I have added a comment saying as much.\r\n{quote}The test has failed for me (repeatedly):\r\n{quote}\r\nThanks, it should be passing now.\r\n\r\n[1] This handles the case where all nodes reporting a given version are on a different version so we can't pull schema from them, and acts as a hedge against any bugs in this implementation that might cause us to not schedule schema pulls as intended","created":"2020-08-04T21:29:31.756+0000"},{"body":"[~bdeggleston] Thanks for the explanation. Maybe [~aleksey] might take a look before moving this forward?\r\n\r\n ","created":"2020-08-06T11:32:56.236+0000"},{"body":"Pushed some minor tweaks [here|https://github.com/iamaleksey/cassandra/commits/15158-review]. Made some bits more idiomatic, and changed the way in-flight requests are being kept track of.\r\n\r\nIn general, this does the job and solves the problem in the description. It doesn't, however, fully deal with storms in large clusters caused by a sequence of updates in quick succession, but, it's not intended to, either.\r\n\r\nEDIT: the amount of synchronisation here bothers me a tiny bit, as all of it will likely have to be eventually gotten rid of, when and if TPC happens, but I can live with it.","created":"2020-09-01T15:48:53.075+0000"},{"body":"Thanks Aleksey.\r\n\r\nRemoving requestQueue is going to cause some problems, We need to have some way of cycling requests through the entire set of endpoints reporting a given version. Without one, if the first few endpoints to come out of the iterator are down, we'll get stuck in a loop as we schedule pulls which will immediately fail causing new pulls to be scheduled for the same node.\r\n\r\nI've updated my branch with your changes and added the request queue back. Since the storage service doesn't really need to fire off migration requests, I also added a commit making the MigrationCoordinator an endpoint change subscriber. So StorageService is responsible for a little less.","created":"2020-09-02T16:32:17.707+0000"},{"body":"Oh, I brainfarted that node liveness check was a part of {{shouldPullFromEndpoint()}}. Could just extend that condition then to add the liveness check in addition to {{shouldPullFromEndpoint()}}?","created":"2020-09-02T18:09:55.583+0000"},{"body":"Possibly, you still need to check in the submission task in case the node has died in the meantime. There would still be an intersection of node flapping rate and unfortunate scheduling where the lockup could occur though.\r\n\r\nThe queue, while a little awkward, also makes us a bit more resilient against other unanticipated states and/or bugs.","created":"2020-09-02T19:16:30.680+0000"},{"body":"I am getting this exception on totally clean node, I am bootstrapping a cluster of 3 nodes:\r\n\r\n\r\n{code:java}\r\ncassandra_node_1 | INFO [ScheduledTasks:1] 2020-09-07 15:10:13,037 TokenMetadata.java:517 - Updating topology for all endpoints that have changed\r\ncassandra_node_1 | INFO [HANDSHAKE-spark-master-1/172.19.0.5] 2020-09-07 15:10:13,311 OutboundTcpConnection.java:561 - Handshaking version with spark-master-1/172.19.0.5\r\ncassandra_node_1 | INFO [GossipStage:1] 2020-09-07 15:10:13,870 Gossiper.java:1141 - Node /172.19.0.5 is now part of the cluster\r\ncassandra_node_1 | INFO [GossipStage:1] 2020-09-07 15:10:13,904 TokenMetadata.java:497 - Updating topology for /172.19.0.5\r\ncassandra_node_1 | INFO [GossipStage:1] 2020-09-07 15:10:13,907 TokenMetadata.java:497 - Updating topology for /172.19.0.5\r\ncassandra_node_1 | INFO [GossipStage:1] 2020-09-07 15:10:14,052 Gossiper.java:1103 - InetAddress /172.19.0.5 is now UP\r\ncassandra_node_1 | WARN [MessagingService-Incoming-/172.19.0.5] 2020-09-07 15:10:14,119 IncomingTcpConnection.java:103 - UnknownColumnFamilyException reading from socket; closing\r\ncassandra_node_1 | org.apache.cassandra.db.UnknownColumnFamilyException: Couldn't find table for cfId 5bc52802-de25-35ed-aeab-188eecebb090. If a table was just created, this is likely due to the schema not being fully propagated. Please wait for schema agreement on table creation.\r\ncassandra_node_1 | \tat org.apache.cassandra.config.CFMetaData$Serializer.deserialize(CFMetaData.java:1578) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.db.partitions.PartitionUpdate$PartitionUpdateSerializer.deserialize30(PartitionUpdate.java:899) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.db.partitions.PartitionUpdate$PartitionUpdateSerializer.deserialize(PartitionUpdate.java:874) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.db.Mutation$MutationSerializer.deserialize(Mutation.java:415) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.db.Mutation$MutationSerializer.deserialize(Mutation.java:434) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.db.Mutation$MutationSerializer.deserialize(Mutation.java:371) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.net.MessageIn.read(MessageIn.java:123) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.net.IncomingTcpConnection.receiveMessage(IncomingTcpConnection.java:195) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.net.IncomingTcpConnection.receiveMessages(IncomingTcpConnection.java:183) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\n\r\n{code}\r\n\r\nThat cfId stands for system_auth/roles. It seems like we are applying changes before schema agreement has occured so that table is not there yet to apply mutations against.\r\n\r\nThis is the log from the second node. The first one booted fine, the second one throws this, the third one boots fine. It seems like eventually everything is just fine however that exception is ... concerning.\r\n","created":"2020-09-07T13:16:30.794+0000"},{"body":"There is also a runtime error as that concurrent hash map from that package is not on the class path. I removed it here, I just squashed all changes in Blakes branch + this one fix:\r\n\r\nhttps://github.com/instaclustr/cassandra/commit/e23677deeb7c836b4b7c80f98009353668351620","created":"2020-09-07T13:36:56.715+0000"},{"body":"These tests are failing\r\n\r\nhttps://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/4/#showFailuresLink","created":"2020-09-07T16:58:53.491+0000"},{"body":"I have improved the original work of mine and I wrote a test for that. Jenkins build does not fail anymore so I believe I have totally on par solution when it comes to dtests as I do not have time to fix dtests which the other solution breaks. While I admit that the improved version is technicaly more superior, the necessity to have clean build and same behaviour when it comes to dtests is more important to me at this moment. It would be awesome if dtests and issues I spotted are resolved though.\r\n\r\nThe test is here (1), the main logic is that a cluster of two nodes is started, the third node is started afterwards and I am dropping all migration messages to the other two, simulating some communication error between them. After some time, migration messages starts to flow again. So by doing this, I ll test the internals of the logic I wrote and it seems to do its job.\r\n\r\nOne issue I am little bit concerned of is that StorageService is issuing schema migration requests on \"onAlive, onJoin ...\" in StorageService and these requests are not part of the waitForSchema() logic. It is understandable that it is like that as we need to track migration requests after a node fully bootstraps but we should skip this from happening when a node is under bootstrapping. I wrapped the bodies of these methods into \"if (hasJoined())\" but it was invoked anyway. However, it does not matter too much if this is outside of the logic I did because if schema migration was sucessful, the rewritten logic in waitForSchema does not have anything to deal with so we are done anyway. For skipping this in test, I used ByteBuddy to intercept MigrationManager#scheduleSchemaPull to do nothing hence I effectively skip migration schemas to be sent outside of the change I did.\r\n\r\nonChange method in onJoin merges schemas again too, in case state is SCHEMA so I am not completely sure why we are merging schemas on a join anyway?\r\n\r\n{code:java}\r\n public void onJoin(InetAddress endpoint, EndpointState epState)\r\n {\r\n for (Map.Entry entry : epState.states())\r\n {\r\n onChange(endpoint, entry.getKey(), entry.getValue());\r\n }\r\n\r\n // this is weird\r\n MigrationManager.instance.scheduleSchemaPull(endpoint, epState);\r\n }\r\n\r\n public void onAlive(InetAddress endpoint, EndpointState state)\r\n {\r\n // this is weird as well\r\n MigrationManager.instance.scheduleSchemaPull(endpoint, state);\r\n if (tokenMetadata.isMember(endpoint))\r\n notifyUp(endpoint);\r\n }\r\n{code}\r\n\r\n\r\n(1) https://github.com/instaclustr/cassandra/blob/15158-original-fix/test/distributed/org/apache/cassandra/distributed/test/BootstrappingSchemaAgreementTest.java\r\n\r\n","created":"2020-09-09T16:18:15.888+0000"},{"body":"Ok, I've fixed all the dtest issues and ported the 3.11 fix to 3.0 and trunk.\r\n\r\nThe 3.11 branch with misc fixes is here: https://github.com/bdeggleston/cassandra/tree/15158-3.11\r\n\r\nMost of the fixes are self explanatory, the less obvious ones are:\r\n\r\n* wait for gossip to settle before waiting on schemas. The original patch accidentally removed the wait on the schema version to be non-empty, so this both a fix and a change. Waiting for gossip makes it more likely that we've seen all current schema versions before we begin waiting instead of just waiting for the first schema to be received, which was the effect of waiting on Schema.instance.isEmpty\r\n* don't check liveness in shouldPullFromEndpoint. There were cases where the node wasn't considered alive until after `reportEndpointVersion` had been called (because it happens as part of the current gossip update).\r\n* move migration start inside StorageService.joinRing because in-jvm dtests don't use CassandraDaemon\r\n* don't fail maybePullSchema if version info has no endpoints. We automatically call that after completing a pull, so a version will eventually have no endpoints left on it\r\n\r\nThe squashed ports are here:\r\n| [3.0|https://github.com/bdeggleston/cassandra/tree/15158-3.0] | [circle|https://app.circleci.com/pipelines/github/bdeggleston/cassandra?branch=15158-3.0] |\r\n| [3.11|https://github.com/bdeggleston/cassandra/tree/15158-3.11-squashed] | [circle|https://app.circleci.com/pipelines/github/bdeggleston/cassandra?branch=15158-3.11-squashed] |\r\n| [trunk|https://github.com/bdeggleston/cassandra/tree/15158-trunk] | [circle|https://app.circleci.com/pipelines/github/bdeggleston/cassandra?branch=15158-trunk] |\r\n","created":"2020-10-09T21:11:31.865+0000"},{"body":"Thanks for this a lot! I ll review right at the beginning of the next week.","created":"2020-10-09T23:30:10.108+0000"},{"body":"I have left my comments in 3.0 branch - probably applicable to 3.11 and trunk too.","created":"2020-10-10T08:51:24.698+0000"},{"body":"Left a small comment on the 3.0 branch. Also, the following nits for {{MigrationCoordinator}}:\r\n1. A bunch of unused imports\r\n2. {{shouldApplySchemaFrom()}} has an unused argument\r\n3. {{requestQueue}} could be an {{ArrayDequeue}} instead of a {{LinkedList}} - should set a good example for anyone randomly reading this code, even if it's not critical to do the right thing in this context \r\n\r\nEDIT: LGTM, +1, ship it","created":"2020-10-19T15:27:17.766+0000"},{"body":"Addressed all review comments and rebased onto the latest branches","created":"2020-10-29T22:09:31.280+0000"},{"body":"I am +1 too (it if counts :) )","created":"2020-11-02T23:21:22.546+0000"},{"body":"[~stefan.miklosovic] it does, thanks to you and [~aleksey] for the reviews. Committed to 3.0 and merged up to trunk.","created":"2020-11-09T20:26:43.919+0000"},{"body":"Found a small set of typos which cause us to wait for schemas for 8h20m rather than 30s, going to submit a patch here and fix in all 3 branches...","created":"2020-11-13T18:52:24.841+0000"},{"body":"Patches:\r\n\r\n3.0: https://github.com/dcapwell/cassandra/tree/patchfix/CASSANDRA-15158-3.0\r\n3.11: https://github.com/dcapwell/cassandra/tree/patchfix/CASSANDRA-15158-3.11\r\ntrunk: https://github.com/dcapwell/cassandra/tree/patchfix/CASSANDRA-15158-trunk\r\n\r\nA test was added in CASSANDRA-16213 which showed this issue, its only stable once this patch is applied (and disable failing)","created":"2020-11-13T19:14:05.784+0000"},{"body":"Starting commit\r\n\r\nCI Results: Yellow. 3.1 org.apache.cassandra.service.MigrationCoordinatorTest but passes locally, -trunk org.apache.cassandra.distributed.test.ring.BootstrapTest fails frequently due to schemas not present added commit which increases timeout from 30s to 90s-, and other expected issues.\r\n||Branch||Source||Circle CI||Jenkins||\r\n|cassandra-3.0|[branch|https://github.com/dcapwell/cassandra/tree/commit_remote_branch/CASSANDRA-15158-cassandra-3.0-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://app.circleci.com/pipelines/github/dcapwell/cassandra?branch=commit_remote_branch%2FCASSANDRA-15158-cassandra-3.0-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://ci-cassandra.apache.org/job/Cassandra-devbranch/200/]|\r\n|cassandra-3.11|[branch|https://github.com/dcapwell/cassandra/tree/commit_remote_branch/CASSANDRA-15158-cassandra-3.11-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://app.circleci.com/pipelines/github/dcapwell/cassandra?branch=commit_remote_branch%2FCASSANDRA-15158-cassandra-3.11-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://ci-cassandra.apache.org/job/Cassandra-devbranch/201/]|\r\n|trunk|[branch|https://github.com/dcapwell/cassandra/tree/commit_remote_branch/CASSANDRA-15158-trunk-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://app.circleci.com/pipelines/github/dcapwell/cassandra?branch=commit_remote_branch%2FCASSANDRA-15158-trunk-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://ci-cassandra.apache.org/job/Cassandra-devbranch/202/]|\r\n","created":"2020-11-13T19:25:50.174+0000"},{"body":"Thanks David, +1","created":"2020-11-13T20:26:44.027+0000"},{"body":"Committed https://github.com/apache/cassandra/commit/7d6f9b94dd0d00bfd29374d7a645e650f451023d","created":"2020-11-13T20:39:30.648+0000"}],"conversations":[{"body":"Currently when a node is bootstrapping we use a set of latches (org.apache.cassandra.service.MigrationTask#inflightTasks) to keep track of in-flight schema pull requests, and we don't proceed with bootstrapping/stream until all the latches are released (or we timeout waiting for each one). One issue with this is that if we have a large schema, or the retrieval of the schema from the other nodes was unexpectedly slow then we have no explicit check in place to ensure we have actually received a schema before we proceed.\r\n\r\nWhile it's possible to increase \"migration_task_wait_in_seconds\" to force the node to wait on each latche longer, there are cases where this doesn't help because the callbacks for the schema pull requests have expired off the messaging service's callback map (org.apache.cassandra.net.MessagingService#callbacks) after request_timeout_in_ms (default 10 seconds) before the other nodes were able to respond to the new node.\r\n\r\nThis patch checks for schema agreement between the bootstrapping node and the rest of the live nodes before proceeding with bootstrapping. It also adds a check to prevent the new node from flooding existing nodes with simultaneous schema pull requests as can happen in large clusters.\r\n\r\nRemoving the latch system should also prevent new nodes in large clusters getting stuck for extended amounts of time as they wait `migration_task_wait_in_seconds` on each of the latches left orphaned by the timed out callbacks.\r\n\r\n \r\n||3.11||\r\n|[PoC|https://github.com/apache/cassandra/compare/cassandra-3.11...vincewhite:check_for_schema]|\r\n|[dtest|https://github.com/apache/cassandra-dtest/compare/master...vincewhite:wait_for_schema_agreement]|\r\n\r\n ","from":"reporter","subject":"Wait for schema agreement rather than in flight schema requests when bootstrapping"},{"body":"So just skimming this, it looks like it's the right approach. We should add some tests though, think through what we want to do if we can't get the schema to converge, as well as leave an \"escape hatch\" if we need to start up and are willing to skip waiting on schema agreement.","from":"developer"},{"body":"I also wanted to point out that the underlying issues causing this problem can lead to correctness issues or data loss in some scenarios. Since we don't currently confirm that any of the in flight migration tasks have been completed and applied, it's possible for us to not receive _any_ schema responses, and begin bootstrapping without a schema. In this case, bootstrap would complete immediately, because the node believes there is nothing to stream. When the node later receives the schema, it will begin serving reads and writes with no data. ","from":"developer"},{"body":"Hi [~bdeggleston],\r\n\r\nplease review this one [https://github.com/smiklosovic/cassandra/tree/CASSANDRA-15158]\r\n\r\nSorry for the fuss but this seems to be more problematic than I was thinking initially. \r\n\r\nI do not think that we have to _check_ that respective migration request is successful or not, we should just check it all schemas match ...\r\n\r\nThe logic is based on repeated checking if schemas agree or not and only in case all schemas are equal we proceed, otherwise an exception is thrown (this might be reworked in such sense that we just log this). There is a global timeout for this check, it will be obvious from reading the code.\r\n\r\nThere is a test added too, because of the nature of the test, I had to \"inject\" these callbacks into respective methods so I could modify their default behaviour and test their state. \r\n\r\nThis is against 3.11, we need to backport this to 3.0 for our customer.","from":"developer"},{"body":"Thanks [~stefan.miklosovic], I'm _should_ have some time to review next week.","from":"developer"},{"body":"code for 3.0 [https://github.com/smiklosovic/cassandra/tree/CASSANDRA-15158-cassandra-3]","from":"developer"},{"body":"for 3.11\r\n\r\n \r\n\r\n[https://github.com/smiklosovic/cassandra/tree/CASSANDRA-15158]\r\n\r\n \r\n\r\nfor 3.0\r\n\r\n \r\n\r\nhttps://github.com/smiklosovic/cassandra/tree/CASSANDRA-15158-cassandra-3","from":"developer"},{"body":"It's likely we'll want to fix this in 3.0 and up, so I'll review the 3.0 version to start with and we can go from there.\r\n\r\nI haven't completed my review yet, but there are some structural and design issues we should address up front.\r\n\r\n*Structural issues*\r\n\r\nFirst, the {{waitForSchema*}} methods should live in the MigrationManager. This will prevent you from setting an updated status if you're sending out additional schema pulls, but we can revisit that later if we think it's neccesary.\r\n\r\nSecond, instantiating MigrationTaskCallbacks in StorageService/MigrationManager and passing it into MigrationTask is a little awkward. I'd prefer if the callback remained an anonymous class. We can communicate endpoint to send schema pulls to with an inetaddress argument, and we need to rethink what `isRunningForcibly` is doing and why. First, we shouldn't be adding it to {{IAsyncCallback}} for this narrow use case. Next, it's use seems to be changing how schema pulls actually work, but only in a test environment, which is something we should avoid.\r\n\r\n*Design issues*\r\n\r\nThis doesn't deal with multiple schema versions. If a node joins, and there are 2 or more schema versions floating around, it will only wait until it has _some_ schema to begin bootstrapping, not all. Related to this, we also need a plan for unreachable schema versions. For instance, if a single node is reporting a schema version that no one else has, but the node is unreachable, what do we do?\r\n\r\nNext, I like how this limits the number of messages sent to a given endpoint, but we should also limit the number of messages we send out for a given schema version. If we have a large cluster, and all nodes are reporting the same version, we don't need to ask every node for it's schema.","from":"developer"},{"body":"Hi [~bdeggleston],\r\n\r\ncommenting on design issues, I am not completely sure if these issues you are talking about are related to this patch or they are already existing? We could indeed focus on the points you raised but it seems to me that the current (comitted) code is worse without this patch than with as I guess these problems are already there?\r\n\r\nIsn't the goal here to have all nodes on same versions? Isn't the very fact that there are multiple versions pretty strange to begin with so we should not even try to join a node if they mismatch hence there is nothing to deal with in the first place? \r\n{quote}It will only wait until it has _some_ schema to begin bootstrapping, not all\r\n{quote}\r\nThis is the most likely not true unless I am not getting something. The node to be bootstrapped will never advance in doing so unless all nodes have same versions. \r\n{quote} For instance, if a single node is reporting a schema version that no one else has, but the node is unreachable, what do we do?\r\n{quote}\r\nWe should fail whole bootstrapping and one should go and fix it.\r\n{quote}For instance, if a single node is reporting a schema version that no one else has, but the node is unreachable, what do we do?\r\n{quote}\r\nHow can a node report its schema while being unreachable?\r\n{quote}Next, I like how this limits the number of messages sent to a given endpoint, but we should also limit the number of messages we send out for a given schema version. If we have a large cluster, and all nodes are reporting the same version, we don't need to ask every node for it's schema.\r\n{quote}\r\n \r\n\r\nGot you, this might be tracked.\r\n\r\n \r\n\r\nWhen it comes to testing, I admit that adding isRunningForcibly method feels like a hack but I had very hard time to test this stuff out. It was basically the only reasonable way possible at the time I was coding it, if you know of more better version, please tell me otherwise I am not sure what might be better here and we could stick with this for a time being? The whole testing methodology was based on these callbacks and checking their inner state which results into having a methods which are accepting them so we can elaborate on their state. Without \"injecting\" them from outside, I would not be able to do that.","from":"developer"},{"body":"It seems to me that one aspect of the PR was overlooked so I just iterate on that one. The mechanim how to not flood nodes with schema pull messages is incorporated in the loop over callbacks. If you notice it, there are sleeps of various lenghts based on a request being already sent or not. This sleep will actually \"delay\" the next schema pull from the other node because during this time of a sleep, some schema could come from the node we just sent a message to so on the next iteration when another node is compared on schema equality, it may happen that there is not any need to pull it anymore because they are on par. Hence we are not blindly sending messages to all nodes.\r\n If there are some discrepancies, there is the global timeout set after which whole bootstrapping process will be evaluated as errorneous and (in the current code) we throw a ConfigurationException. This behaviour might be relaxed but I consider it more appropriate to just throw it there.","from":"developer"},{"body":"{quote}commenting on design issues, I am not completely sure if these issues you are talking about are related to this patch or they are already existing? We could indeed focus on the points you raised but it seems to me that the current (comitted) code is worse without this patch than with as I guess these problems are already there?\r\nIsn't the goal here to have all nodes on same versions? Isn't the very fact that there are multiple versions pretty strange to begin with so we should not even try to join a node if they mismatch hence there is nothing to deal with in the first place?\r\n{quote}\r\nWhen there are schema changes, it's not strange at all for there to be multiple schema versions in the cluster before they converge. We also don't forbid making schema changes while changing cluster topology, so this would be something we should expect to encounter, although I would expect it to happen infrequently. Since bootstrap doesn't stream keyspaces it doesn't know about, this could create a window of data loss. Since the goal of this ticket is to wait for schema to converge before starting bootstrap, we should deal with edge cases like this. Also, I believe there have been bugs that caused a lot of schema change activity when nodes bootstrap, so depending on what exactly you're doing\r\n{quote}How can a node report its schema while being unreachable?\r\n{quote}\r\nSchema versions are gossiped. So a node might gossip a new schema version then become unreachable. The bootstrapping node would learn about this new version via gossip, but be unable to contact it.\r\n{quote}> admit that adding isRunningForcibly method feels like a hack but I had very hard time to test this stuff out.\r\n{quote}\r\nI'll look into how testing can be improved.\r\n{quote}> This is the most likely not true unless I am not getting something. The node to be bootstrapped will never advance in doing so unless all nodes have same versions.\r\n{quote}\r\nAh, yes you're right. Althought waiting for all nodes to arrive at the same schema version isn't neccesary, we just need to receive and merge at least one schema pull from every schema version in the cluster.","from":"developer"},{"body":"Hi Blake,\r\n\r\n \r\n\r\nbecause of your very helpful explanation I was able to put together yet another version of the solution to this problem. You will find it here\r\n\r\n[https://github.com/apache/cassandra/pull/628]\r\n\r\nThanks for the review in advance","from":"developer"},{"body":"I've reworked this a bit more [here|https://github.com/bdeggleston/cassandra/tree/15158-coordinator]. It's now pretty self contained, has some fairly granular unit tests, and fixes a few functional things. Can you take a look [~stefan.miklosovic] and let me know what you think? [~aleksey], can you also take a look / review? The basic idea is that we now track which schema versions exist in gossip, and which endpoints are reporting them, then block bootstrap until we've received a schema for each version. It also tracks how many outstanding migration requests we have per version so we don't send out thousands.\r\n\r\nAlso, what are your opinions of this going into 3.x? I think I'd lean towards putting it in, since it eliminates a scenario where data loss could occur and will shave a few hours off of adding/replacing nodes in large clusters. On the other hand, since it reduces the amount of migration requests sent out on bootstrap, any bugs determining if we've received sufficient schema data could make data loss _more_ likely.","from":"developer"},{"body":"In general looks good to me in spite of getting lost a bit on the actual schema migration response:\r\n\r\n \r\n\r\n \r\n{code:java}\r\nFuture response(Collection mutations)\r\n{\r\n synchronized (info)\r\n {\r\n if (shouldApplySchemaFrom(endpoint, info))\r\n mergeSchemaFrom(endpoint, mutations);\r\n return pullComplete(endpoint, info, true); /// why?\r\n }\r\n}\r\n{code}\r\n \r\n\r\nI am not completely sure why are we pulling again here. I would rewrite the whole solution in a such way that this Callable just does one thing on a successful response (merging of a schema) and the actual \"retry\" would be handled from outside. The reader has to make quite a mental exercise to visualise that this callback might actually call another callback in it until some \"version\" is completed etc ... At least for me, it was quite tedious to track.\r\n\r\nI understand the motivation behind that but callback should just do one task and thats it, it shouldnt be responsible for potentially scheduling another callback recursively until some conditions on some signals or what have you are met ... just my 2 cents.\r\n\r\nThere is also this:\r\n\r\n \r\n{code:java}\r\nsynchronized Future reportEndpointVersion(InetAddress endpoint, UUID version)\r\n{\r\n UUID current = endpointVersions.get(endpoint);\r\n if (current != null && current.equals(version))\r\n return FINISHED_FUTURE;\r\n\r\n VersionInfo info = versionInfo.computeIfAbsent(version, VersionInfo::new);\r\n if (isLocalVersion(version))\r\n info.markReceived();\r\n info.endpoints.add(endpoint);\r\n info.requestQueue.add(endpoint);\r\n endpointVersions.put(endpoint, version);\r\n\r\n removeEndpointFromVersion(endpoint, current); ///// why?\r\n return maybePullSchema(info);\r\n}\r\n{code}\r\n \r\n\r\nTBH that is quite counterintuitive too, maybe renaming of that method would help.\r\n\r\n \r\n\r\nThe test has failed for me (repeatedly):\r\n\r\n \r\n\r\n \r\n{code:java}\r\njava.lang.AssertionError: java.lang.AssertionError: \r\nExpected :2\r\nActual   :0\r\n at org.junit.Assert.fail(Assert.java:92) at org.junit.Assert.failNotEquals(Assert.java:689) at org.junit.Assert.assertEquals(Assert.java:127) at org.junit.Assert.assertEquals(Assert.java:514) at org.junit.Assert.assertEquals(Assert.java:498) at org.apache.cassandra.service.MigrationCoordinatorTest.testWeKeepSendingRequests(MigrationCoordinatorTest.java:278) at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method) at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62) at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43) at java.lang.reflect.Method.invoke(Method.java:498) at org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:44) at org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15) at org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:41) at org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:20) at org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:28) at org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:31) at org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:70) at org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:44) at org.junit.runners.ParentRunner.runChildren(ParentRunner.java:180) at org.junit.runners.ParentRunner.access$000(ParentRunner.java:41) at org.junit.runners.ParentRunner$1.evaluate(ParentRunner.java:173) at org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:28) at org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:31) at org.junit.runners.ParentRunner.run(ParentRunner.java:220) at org.junit.runner.JUnitCore.run(JUnitCore.java:159) at com.intellij.junit4.JUnit4IdeaTestRunner.startRunnerWithArgs(JUnit4IdeaTestRunner.java:68) at com.intellij.rt.junit.IdeaTestRunner$Repeater.startRunnerWithArgs(IdeaTestRunner.java:33) at com.intellij.rt.junit.JUnitStarter.prepareStreamsAndStart(JUnitStarter.java:230) at com.intellij.rt.junit.JUnitStarter.main(JUnitStarter.java:58)\r\n{code}\r\n ","from":"developer"},{"body":"hi [~bdeggleston], have you had a chance to reflect the above issues?","from":"developer"},{"body":"{quote}\r\nI am not completely sure why are we pulling again here. I would rewrite the whole solution in a such way that this Callable just does one thing on a successful response (merging of a schema) and the actual \"retry\" would be handled from outside. The reader has to make quite a mental exercise to visualise that this callback might actually call another callback in it until some \"version\" is completed etc ... At least for me, it was quite tedious to track.\r\n\r\n\r\n{quote}\r\nIn the case of a successful pull, we won't pull again. Response and fail both call pullComplete, but an additional pull is only called if it's called from fail.\r\n\r\nI get that this can be a bit difficult to follow, but I'm not sure there's a better approach, given the schema pulls are completely event driven during normal runtime. If we miss a schema change during normal runtime (not bootstrap), there's nothing waiting on schema convergence that would enable us to retry from the outside.\r\n\r\nThere is a periodic task that pulls schema for outstanding versions that don't have any in flight requests^[1]^, but it only runs once a minute, and we need to be more proactive about learning about schema updates since we'll be unable to serve some reads and writes until we're up to date.\r\n{quote}TBH that is quite counterintuitive too\r\n{quote}\r\nCould you expand on what's counterintuitive about it? If the endpoint's schema version has changed, we need to disassociate it with it's previously reported version. I have added a comment saying as much.\r\n{quote}The test has failed for me (repeatedly):\r\n{quote}\r\nThanks, it should be passing now.\r\n\r\n[1] This handles the case where all nodes reporting a given version are on a different version so we can't pull schema from them, and acts as a hedge against any bugs in this implementation that might cause us to not schedule schema pulls as intended","from":"developer"},{"body":"[~bdeggleston] Thanks for the explanation. Maybe [~aleksey] might take a look before moving this forward?\r\n\r\n ","from":"developer"},{"body":"Pushed some minor tweaks [here|https://github.com/iamaleksey/cassandra/commits/15158-review]. Made some bits more idiomatic, and changed the way in-flight requests are being kept track of.\r\n\r\nIn general, this does the job and solves the problem in the description. It doesn't, however, fully deal with storms in large clusters caused by a sequence of updates in quick succession, but, it's not intended to, either.\r\n\r\nEDIT: the amount of synchronisation here bothers me a tiny bit, as all of it will likely have to be eventually gotten rid of, when and if TPC happens, but I can live with it.","from":"developer"},{"body":"Thanks Aleksey.\r\n\r\nRemoving requestQueue is going to cause some problems, We need to have some way of cycling requests through the entire set of endpoints reporting a given version. Without one, if the first few endpoints to come out of the iterator are down, we'll get stuck in a loop as we schedule pulls which will immediately fail causing new pulls to be scheduled for the same node.\r\n\r\nI've updated my branch with your changes and added the request queue back. Since the storage service doesn't really need to fire off migration requests, I also added a commit making the MigrationCoordinator an endpoint change subscriber. So StorageService is responsible for a little less.","from":"developer"},{"body":"Oh, I brainfarted that node liveness check was a part of {{shouldPullFromEndpoint()}}. Could just extend that condition then to add the liveness check in addition to {{shouldPullFromEndpoint()}}?","from":"developer"},{"body":"Possibly, you still need to check in the submission task in case the node has died in the meantime. There would still be an intersection of node flapping rate and unfortunate scheduling where the lockup could occur though.\r\n\r\nThe queue, while a little awkward, also makes us a bit more resilient against other unanticipated states and/or bugs.","from":"developer"},{"body":"I am getting this exception on totally clean node, I am bootstrapping a cluster of 3 nodes:\r\n\r\n\r\n{code:java}\r\ncassandra_node_1 | INFO [ScheduledTasks:1] 2020-09-07 15:10:13,037 TokenMetadata.java:517 - Updating topology for all endpoints that have changed\r\ncassandra_node_1 | INFO [HANDSHAKE-spark-master-1/172.19.0.5] 2020-09-07 15:10:13,311 OutboundTcpConnection.java:561 - Handshaking version with spark-master-1/172.19.0.5\r\ncassandra_node_1 | INFO [GossipStage:1] 2020-09-07 15:10:13,870 Gossiper.java:1141 - Node /172.19.0.5 is now part of the cluster\r\ncassandra_node_1 | INFO [GossipStage:1] 2020-09-07 15:10:13,904 TokenMetadata.java:497 - Updating topology for /172.19.0.5\r\ncassandra_node_1 | INFO [GossipStage:1] 2020-09-07 15:10:13,907 TokenMetadata.java:497 - Updating topology for /172.19.0.5\r\ncassandra_node_1 | INFO [GossipStage:1] 2020-09-07 15:10:14,052 Gossiper.java:1103 - InetAddress /172.19.0.5 is now UP\r\ncassandra_node_1 | WARN [MessagingService-Incoming-/172.19.0.5] 2020-09-07 15:10:14,119 IncomingTcpConnection.java:103 - UnknownColumnFamilyException reading from socket; closing\r\ncassandra_node_1 | org.apache.cassandra.db.UnknownColumnFamilyException: Couldn't find table for cfId 5bc52802-de25-35ed-aeab-188eecebb090. If a table was just created, this is likely due to the schema not being fully propagated. Please wait for schema agreement on table creation.\r\ncassandra_node_1 | \tat org.apache.cassandra.config.CFMetaData$Serializer.deserialize(CFMetaData.java:1578) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.db.partitions.PartitionUpdate$PartitionUpdateSerializer.deserialize30(PartitionUpdate.java:899) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.db.partitions.PartitionUpdate$PartitionUpdateSerializer.deserialize(PartitionUpdate.java:874) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.db.Mutation$MutationSerializer.deserialize(Mutation.java:415) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.db.Mutation$MutationSerializer.deserialize(Mutation.java:434) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.db.Mutation$MutationSerializer.deserialize(Mutation.java:371) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.net.MessageIn.read(MessageIn.java:123) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.net.IncomingTcpConnection.receiveMessage(IncomingTcpConnection.java:195) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\ncassandra_node_1 | \tat org.apache.cassandra.net.IncomingTcpConnection.receiveMessages(IncomingTcpConnection.java:183) ~[apache-cassandra-3.11.9-SNAPSHOT.jar:3.11.9-SNAPSHOT]\r\n\r\n{code}\r\n\r\nThat cfId stands for system_auth/roles. It seems like we are applying changes before schema agreement has occured so that table is not there yet to apply mutations against.\r\n\r\nThis is the log from the second node. The first one booted fine, the second one throws this, the third one boots fine. It seems like eventually everything is just fine however that exception is ... concerning.\r\n","from":"developer"},{"body":"There is also a runtime error as that concurrent hash map from that package is not on the class path. I removed it here, I just squashed all changes in Blakes branch + this one fix:\r\n\r\nhttps://github.com/instaclustr/cassandra/commit/e23677deeb7c836b4b7c80f98009353668351620","from":"developer"},{"body":"These tests are failing\r\n\r\nhttps://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/4/#showFailuresLink","from":"developer"},{"body":"I have improved the original work of mine and I wrote a test for that. Jenkins build does not fail anymore so I believe I have totally on par solution when it comes to dtests as I do not have time to fix dtests which the other solution breaks. While I admit that the improved version is technicaly more superior, the necessity to have clean build and same behaviour when it comes to dtests is more important to me at this moment. It would be awesome if dtests and issues I spotted are resolved though.\r\n\r\nThe test is here (1), the main logic is that a cluster of two nodes is started, the third node is started afterwards and I am dropping all migration messages to the other two, simulating some communication error between them. After some time, migration messages starts to flow again. So by doing this, I ll test the internals of the logic I wrote and it seems to do its job.\r\n\r\nOne issue I am little bit concerned of is that StorageService is issuing schema migration requests on \"onAlive, onJoin ...\" in StorageService and these requests are not part of the waitForSchema() logic. It is understandable that it is like that as we need to track migration requests after a node fully bootstraps but we should skip this from happening when a node is under bootstrapping. I wrapped the bodies of these methods into \"if (hasJoined())\" but it was invoked anyway. However, it does not matter too much if this is outside of the logic I did because if schema migration was sucessful, the rewritten logic in waitForSchema does not have anything to deal with so we are done anyway. For skipping this in test, I used ByteBuddy to intercept MigrationManager#scheduleSchemaPull to do nothing hence I effectively skip migration schemas to be sent outside of the change I did.\r\n\r\nonChange method in onJoin merges schemas again too, in case state is SCHEMA so I am not completely sure why we are merging schemas on a join anyway?\r\n\r\n{code:java}\r\n public void onJoin(InetAddress endpoint, EndpointState epState)\r\n {\r\n for (Map.Entry entry : epState.states())\r\n {\r\n onChange(endpoint, entry.getKey(), entry.getValue());\r\n }\r\n\r\n // this is weird\r\n MigrationManager.instance.scheduleSchemaPull(endpoint, epState);\r\n }\r\n\r\n public void onAlive(InetAddress endpoint, EndpointState state)\r\n {\r\n // this is weird as well\r\n MigrationManager.instance.scheduleSchemaPull(endpoint, state);\r\n if (tokenMetadata.isMember(endpoint))\r\n notifyUp(endpoint);\r\n }\r\n{code}\r\n\r\n\r\n(1) https://github.com/instaclustr/cassandra/blob/15158-original-fix/test/distributed/org/apache/cassandra/distributed/test/BootstrappingSchemaAgreementTest.java\r\n\r\n","from":"developer"},{"body":"Ok, I've fixed all the dtest issues and ported the 3.11 fix to 3.0 and trunk.\r\n\r\nThe 3.11 branch with misc fixes is here: https://github.com/bdeggleston/cassandra/tree/15158-3.11\r\n\r\nMost of the fixes are self explanatory, the less obvious ones are:\r\n\r\n* wait for gossip to settle before waiting on schemas. The original patch accidentally removed the wait on the schema version to be non-empty, so this both a fix and a change. Waiting for gossip makes it more likely that we've seen all current schema versions before we begin waiting instead of just waiting for the first schema to be received, which was the effect of waiting on Schema.instance.isEmpty\r\n* don't check liveness in shouldPullFromEndpoint. There were cases where the node wasn't considered alive until after `reportEndpointVersion` had been called (because it happens as part of the current gossip update).\r\n* move migration start inside StorageService.joinRing because in-jvm dtests don't use CassandraDaemon\r\n* don't fail maybePullSchema if version info has no endpoints. We automatically call that after completing a pull, so a version will eventually have no endpoints left on it\r\n\r\nThe squashed ports are here:\r\n| [3.0|https://github.com/bdeggleston/cassandra/tree/15158-3.0] | [circle|https://app.circleci.com/pipelines/github/bdeggleston/cassandra?branch=15158-3.0] |\r\n| [3.11|https://github.com/bdeggleston/cassandra/tree/15158-3.11-squashed] | [circle|https://app.circleci.com/pipelines/github/bdeggleston/cassandra?branch=15158-3.11-squashed] |\r\n| [trunk|https://github.com/bdeggleston/cassandra/tree/15158-trunk] | [circle|https://app.circleci.com/pipelines/github/bdeggleston/cassandra?branch=15158-trunk] |\r\n","from":"developer"},{"body":"Thanks for this a lot! I ll review right at the beginning of the next week.","from":"developer"},{"body":"I have left my comments in 3.0 branch - probably applicable to 3.11 and trunk too.","from":"developer"},{"body":"Left a small comment on the 3.0 branch. Also, the following nits for {{MigrationCoordinator}}:\r\n1. A bunch of unused imports\r\n2. {{shouldApplySchemaFrom()}} has an unused argument\r\n3. {{requestQueue}} could be an {{ArrayDequeue}} instead of a {{LinkedList}} - should set a good example for anyone randomly reading this code, even if it's not critical to do the right thing in this context \r\n\r\nEDIT: LGTM, +1, ship it","from":"developer"},{"body":"Addressed all review comments and rebased onto the latest branches","from":"developer"},{"body":"I am +1 too (it if counts :) )","from":"developer"},{"body":"[~stefan.miklosovic] it does, thanks to you and [~aleksey] for the reviews. Committed to 3.0 and merged up to trunk.","from":"developer"},{"body":"Found a small set of typos which cause us to wait for schemas for 8h20m rather than 30s, going to submit a patch here and fix in all 3 branches...","from":"developer"},{"body":"Patches:\r\n\r\n3.0: https://github.com/dcapwell/cassandra/tree/patchfix/CASSANDRA-15158-3.0\r\n3.11: https://github.com/dcapwell/cassandra/tree/patchfix/CASSANDRA-15158-3.11\r\ntrunk: https://github.com/dcapwell/cassandra/tree/patchfix/CASSANDRA-15158-trunk\r\n\r\nA test was added in CASSANDRA-16213 which showed this issue, its only stable once this patch is applied (and disable failing)","from":"developer"},{"body":"Starting commit\r\n\r\nCI Results: Yellow. 3.1 org.apache.cassandra.service.MigrationCoordinatorTest but passes locally, -trunk org.apache.cassandra.distributed.test.ring.BootstrapTest fails frequently due to schemas not present added commit which increases timeout from 30s to 90s-, and other expected issues.\r\n||Branch||Source||Circle CI||Jenkins||\r\n|cassandra-3.0|[branch|https://github.com/dcapwell/cassandra/tree/commit_remote_branch/CASSANDRA-15158-cassandra-3.0-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://app.circleci.com/pipelines/github/dcapwell/cassandra?branch=commit_remote_branch%2FCASSANDRA-15158-cassandra-3.0-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://ci-cassandra.apache.org/job/Cassandra-devbranch/200/]|\r\n|cassandra-3.11|[branch|https://github.com/dcapwell/cassandra/tree/commit_remote_branch/CASSANDRA-15158-cassandra-3.11-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://app.circleci.com/pipelines/github/dcapwell/cassandra?branch=commit_remote_branch%2FCASSANDRA-15158-cassandra-3.11-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://ci-cassandra.apache.org/job/Cassandra-devbranch/201/]|\r\n|trunk|[branch|https://github.com/dcapwell/cassandra/tree/commit_remote_branch/CASSANDRA-15158-trunk-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://app.circleci.com/pipelines/github/dcapwell/cassandra?branch=commit_remote_branch%2FCASSANDRA-15158-trunk-7E401495-E38F-4857-80C1-2C27028F572E]|[build|https://ci-cassandra.apache.org/job/Cassandra-devbranch/202/]|\r\n","from":"developer"},{"body":"Thanks David, +1","from":"developer"},{"body":"Committed https://github.com/apache/cassandra/commit/7d6f9b94dd0d00bfd29374d7a645e650f451023d","from":"developer"}],"created":"2019-06-13T09:18:30.000+0000","description":"Currently when a node is bootstrapping we use a set of latches (org.apache.cassandra.service.MigrationTask#inflightTasks) to keep track of in-flight schema pull requests, and we don't proceed with bootstrapping/stream until all the latches are released (or we timeout waiting for each one). One issue with this is that if we have a large schema, or the retrieval of the schema from the other nodes was unexpectedly slow then we have no explicit check in place to ensure we have actually received a schema before we proceed.\r\n\r\nWhile it's possible to increase \"migration_task_wait_in_seconds\" to force the node to wait on each latche longer, there are cases where this doesn't help because the callbacks for the schema pull requests have expired off the messaging service's callback map (org.apache.cassandra.net.MessagingService#callbacks) after request_timeout_in_ms (default 10 seconds) before the other nodes were able to respond to the new node.\r\n\r\nThis patch checks for schema agreement between the bootstrapping node and the rest of the live nodes before proceeding with bootstrapping. It also adds a check to prevent the new node from flooding existing nodes with simultaneous schema pull requests as can happen in large clusters.\r\n\r\nRemoving the latch system should also prevent new nodes in large clusters getting stuck for extended amounts of time as they wait `migration_task_wait_in_seconds` on each of the latches left orphaned by the timed out callbacks.\r\n\r\n \r\n||3.11||\r\n|[PoC|https://github.com/apache/cassandra/compare/cassandra-3.11...vincewhite:check_for_schema]|\r\n|[dtest|https://github.com/apache/cassandra-dtest/compare/master...vincewhite:wait_for_schema_agreement]|\r\n\r\n ","issue_id":"13239212","key":"CASSANDRA-15158","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2020-11-09T20:26:43.000+0000","role":"fixed_distractor","summary":"Wait for schema agreement rather than in flight schema requests when bootstrapping"} {"case_id":"13240381","cluster":"DISTRACTOR-CASSANDRA-15172","comments":[{"body":"Trying to push this up a bit because I still want to upgrade to 3.11.4 but I fear that this issue may recur.\r\n\r\nIf someone has any idea on what happened here or how to mitigate it, that'd be awesome!\r\n\r\n \r\n\r\nThanks!","created":"2019-07-18T14:46:36.573+0000"},{"body":"We have also hit this problem today while upgrading from 2.1.16 to 3.11.4/ \r\n\r\nwe encountered this as soon as node started up with 3.11.4 \r\n\r\n \r\n\r\nWARN [ReadStage-4] 2019-08-06 02:57:57,408 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-4,5,main]: {}\r\njava.lang.NullPointerException: null\r\n\r\n \r\n\r\nERROR [Native-Transport-Requests-32] 2019-08-06 02:14:20,353 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n\r\nand the below errors continued in the logfile as long as the process was up.\r\n\r\nERROR [Native-Transport-Requests-12] 2019-08-06 03:00:47,135 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-8] 2019-08-06 03:00:48,778 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-13] 2019-08-06 03:00:57,454 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-11] 2019-08-06 03:00:57,482 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-2] 2019-08-06 03:00:58,543 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-8] 2019-08-06 03:00:58,899 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-17] 2019-08-06 03:00:59,074 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-12] 2019-08-06 03:01:08,123 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-17] 2019-08-06 03:01:19,055 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-4] 2019-08-06 03:01:20,880 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n WARN [ReadStage-13] 2019-08-06 03:01:29,983 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-13,5,main]: {}\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-2] 2019-08-06 03:01:31,119 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-6] 2019-08-06 03:01:46,262 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-15] 2019-08-06 03:01:46,520 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n WARN [ReadStage-2] 2019-08-06 03:01:48,842 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-2,5,main]: {}\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-1] 2019-08-06 03:01:50,351 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-5] 2019-08-06 03:02:06,061 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n WARN [ReadStage-8] 2019-08-06 03:02:07,616 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-8,5,main]: {}\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-17] 2019-08-06 03:02:08,384 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-5] 2019-08-06 03:02:10,244 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n\r\n \r\n\r\nThe nodetool version says 3.11.4 and the no of connections on 9042 was similar to other nodes. The exceptions were scary that we had to call off the change. Any help and insights to this problem from the community is appreciated.","created":"2019-08-06T03:35:37.497+0000"},{"body":"Full stack trace is as below:\r\n\r\nWARN [ReadStage-4] 2019-08-06 02:57:57,408 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-4,5,main]: {}\r\njava.lang.NullPointerException: null\r\n at org.apache.cassandra.db.LegacyLayout$LegacyRangeTombstoneList.updateDigest(LegacyLayout.java:2433) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.LegacyLayout$LegacyUnfilteredPartition.digest(LegacyLayout.java:1479) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.rows.UnfilteredRowIterators.digest(UnfilteredRowIterators.java:182) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators.digest(UnfilteredPartitionIterators.java:263) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.ReadResponse.makeDigest(ReadResponse.java:140) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.ReadResponse.createDigestResponse(ReadResponse.java:87) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:352) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.ReadCommandVerbHandler.doVerb(ReadCommandVerbHandler.java:50) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:66) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_131]\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:162) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:134) [apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:114) [apache-cassandra-3.11.4.jar:3.11.4]\r\n at java.lang.Thread.run(Thread.java:748) [na:1.8.0_131]","created":"2019-08-06T10:41:02.159+0000"},{"body":"What we also know is that the nodes suffered some heavy tombstones recently.. could it be under specific condition like \"tombstones\" plus reading legacy version files is hitting this problem. [~Sagges] did you cluster have any tombstone references at all?","created":"2019-08-06T10:42:57.037+0000"},{"body":"We have currently isolated this node (cut off thrift , native protocols) and trying to run upgradesstables to see if it can re-write all the files and stop logging the below message.\r\n\r\n\"WARN [ReadStage-6] 2019-08-06 10:44:09,773 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-6,5,main]: {}\r\njava.lang.NullPointerException: null\"\r\n\r\nI will keep the thread posted about its outcome, \r\n\r\n ","created":"2019-08-06T10:45:24.016+0000"},{"body":"[~Sagges], sorry for the slow response - I missed the original filing of this ticket.\r\n\r\n[~ferozshaik552@gmail.com] it looks like your bug, while very similar, presents differently. It would be great if you could file a separate ticket.\r\n\r\nBoth of these look to be among the category of 2.1->3.0 upgrade bugs involving range deletions. To best investigate and diagnose, it would be great to start with information about the affected schema, the kinds of range tombstone deletes you perform, and preferably if you could pin down sstables that are affected and upload them somewhere private for us to access. This would help us investigate much more readily.\r\n\r\nCould you also confirm if you utilise thrift, or CQL schema? It's possible this is a compatibility issue specific to thrift.","created":"2019-08-06T12:58:00.594+0000"},{"body":"Sure, I shall raise a separate request.","created":"2019-08-06T13:25:26.564+0000"},{"body":"This bug appears to be similar to CASSANDRA-15263, in that a reverse query with the RTBoundCloser is the likely source of asymmetric range tombstone bounds. However in this case the problem is much easier to solve; we simply have to not assume the bounds have the same length.\r\n\r\nI have pushed a patch [here|https://github.com/belliottsmith/cassandra/tree/15172-3.0]","created":"2019-08-12T13:50:06.330+0000"},{"body":"[~benedict], we've seen this in the wild as well, with an upgrade from 2.2.14 to 3.11.4.\r\n I am jumping in to test and review it.","created":"2019-08-21T14:45:32.458+0000"},{"body":"\r\n||branch||circleci||asf jenkins testall||asf jenkins dtests||\r\n|[15172-3.0|https://github.com/apache/cassandra/compare/trunk...belliottsmith:15172-3.0]|[circleci|https://circleci.com/gh/belliottsmith/workflows/cassandra/tree/15172-3.0]|[!https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-testall/44//badge/icon!|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-testall/44/]|[!https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/679//badge/icon!|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/679]|\r\n\r\n","created":"2019-08-21T21:44:11.481+0000"},{"body":"Committed as 2b10a5f2b5e62f2900119a37e91637916e8b23df","created":"2019-08-22T11:50:27.026+0000"},{"body":"Hi [~benedict]\r\n\r\nSorry for the late reply...\r\n\r\n \r\n\r\nCould you also confirm if you utilise thrift, or CQL schema? It's possible this is a compatibility issue specific to thrift.\r\n\r\nI do see there are a few thrift counter tables that were created with the WITH COMPACT STORAGE attribute.\r\n\r\n \r\n\r\nI also see that [~ferozshaik552@gmail.com] had seen WARN [ReadStage-4] 2019-08-06 02:57:57,408 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-4,5,main]: {}\r\n java.lang.NullPointerException: null\r\n\r\n\r\n\r\nWhat I mainly saw was AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-9,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: null\r\n\r\nDoes the fix fit the same scenario as Feroz's? Could my issue be thrift related?\r\n\r\n \r\n\r\nSorry again for the late reply.\r\n\r\n ","created":"2019-08-25T09:02:22.682+0000"},{"body":"Hi [~benedict] ,\r\n\r\nI'm not sure if my comments/tickets are somehow not getting filed, but I hope not :)\r\n\r\nIs the above error (previous comment) related to this bug or is it a different one?\r\n\r\n \r\n\r\nThanks!","created":"2019-08-27T17:40:05.213+0000"},{"body":"[~Sagges],\r\n\r\n this bug comes from existing thrift (legacy) tables, where range tombstones were used. \r\n\r\nThe NPE [~ferozshaik552@gmail.com] reported is a separate bug, despite it also being coming from legacy thrift tables with range tombstones. Unfortunately though, the fix for this ticket will not solve the NPE bug. ","created":"2019-08-27T17:50:37.901+0000"},{"body":"Thanks for clarifying [~mck]!","created":"2019-08-27T17:56:09.308+0000"},{"body":"Hi [~Sagges], I'm afraid I was on holiday so was unable to respond, and am now otherwise engaged for the next few weeks, but [~mck] is mostly correct. However, to clarify, the likely cause of this bug that I have established is not related to thrift or legacy tables (though atypical range tombstone use with thrift could cause it), but to communication from a 3.0 node to a 2.2 or 2.1 node, in the face of range tombstones that cover a primary key prefix.\r\n\r\nThat is to say, a schema of the form (pk, c1, c2, v), with a deletion on (pk, c1)","created":"2019-08-27T18:47:37.183+0000"},{"body":"Thanks a lot for further clarifying [~benedict]. (I hope you enjoyed your vacation :) )\r\n\r\nJust to set my mind straight, the issue is when there are mixed versions in the cluster, so if I upgraded all binaries to 3.11, it won't recur even if I haven't upgraded the SSTables yet. Is my assumption correct?","created":"2019-08-29T12:04:08.512+0000"},{"body":"Hi [~benedict]\r\n\r\nI tried the patch. It didn't do the trick, however, I was able to fully reproduce the bug.\r\n\r\ntl;dr\r\n\r\nThe bug does occur when running queries on range tombstones, but on top of that, those queries have to specifically be range queries.\r\n\r\n \r\n\r\n*+Steps to Reproduce:\r\n+* \r\n\r\nCREATE KEYSPACE ks1 WITH replication = \\{'class': 'NetworkTopologyStrategy', 'DC1': '3'} AND durable_writes = true;\r\n\r\n+*TABLE:*+ \r\nCREATE TABLE ks1.table1 (\r\n col1 text,\r\n col2 text,\r\n col3 text,\r\n col4 text,\r\n col5 text,\r\n col6 timestamp,\r\n data text,\r\n PRIMARY KEY ((col1, col2, col3), col4, col5, col6)\r\n);\r\n\r\n \r\n\r\nInserted ~4 million rows and created range tombstones by deleting ~1 million rows.\r\n\r\n \r\n\r\n+*Create Data*+\r\n\r\n_insert into ks1.table1 (col1, col2 , col3 , col4 , col5 , col6 , data ) VALUES ( '1', '11', '21', '1', 'a', 12312312300000, 'data');_\r\n_insert into ks1.table1 (col1, col2 , col3 , col4 , col5 , col6 , data ) VALUES ( '1', '11', '21', '2', 'a', 12312312300000, 'data');_\r\n_insert into ks1.table1 (col1, col2 , col3 , col4 , col5 , col6 , data ) VALUES ( '1', '11', '21', '3', 'a', 12312312300000, 'data');_\r\n_insert into ks1.table1 (col1, col2 , col3 , col4 , col5 , col6 , data ) VALUES ( '1', '11', '21', '4', 'a', 12312312300000, 'data');_\r\n_insert into ks1.table1 (col1, col2 , col3 , col4 , col5 , col6 , data ) VALUES ( '1', '11', '21', '5', 'a', 12312312300000, 'data');_\r\n\r\n \r\n\r\n+*Create Range Tombstones*+\r\n\r\ndelete from ks1.table1 where col1='1' and col2='11' and col3='21' and col4='1';\r\n\r\n \r\n\r\n+*Query Live Rows (no tombstones)*+\r\n\r\n_select * from ks1.table1 where col1='1' and col2='201' and col3='21' and col4='1' and col5='a' and *col6>12312312300000*;_\r\n\r\nNo issues found, everything is running properly.\r\n\r\n \r\n\r\n+*Query Range Tombstones*+\r\n\r\n_select * from ks1.table1 where col1='1' and col2='11' and col3='21' and col4='1' and col5='a' and *col6=12312312300000*;_\r\n\r\nNo issues found, everything is running properly.\r\n\r\n \r\n\r\n+BUT when running range queries:+\r\n\r\n_select * from ks1.table1 where col1='1' and col2='11' and col3='21' and col4='1' and col5='a' and *col6>12312312200000;*_\r\n\r\nWARN [ReadStage-1] 2019-09-23 14:17:10,281 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-1,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: 2\r\n at org.apache.cassandra.db.AbstractBufferClusteringPrefix.get(AbstractBufferClusteringPrefix.java:55)\r\n at org.apache.cassandra.db.LegacyLayout$LegacyRangeTombstoneList.serializedSizeCompound(LegacyLayout.java:2545)\r\n at org.apache.cassandra.db.LegacyLayout$LegacyRangeTombstoneList.serializedSize(LegacyLayout.java:2522)\r\n at org.apache.cassandra.db.LegacyLayout.serializedSizeAsLegacyPartition(LegacyLayout.java:565)\r\n at org.apache.cassandra.db.ReadResponse$Serializer.serializedSize(ReadResponse.java:446)\r\n at org.apache.cassandra.db.ReadResponse$Serializer.serializedSize(ReadResponse.java:352)\r\n at org.apache.cassandra.net.MessageOut.payloadSize(MessageOut.java:171)\r\n at org.apache.cassandra.net.OutboundTcpConnectionPool.getConnection(OutboundTcpConnectionPool.java:77)\r\n at org.apache.cassandra.net.MessagingService.getConnection(MessagingService.java:802)\r\n at org.apache.cassandra.net.MessagingService.sendOneWay(MessagingService.java:953)\r\n at org.apache.cassandra.net.MessagingService.sendReply(MessagingService.java:929)\r\n at org.apache.cassandra.db.ReadCommandVerbHandler.doVerb(ReadCommandVerbHandler.java:62)\r\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:66)\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:162)\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:134)\r\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:114)\r\n at java.lang.Thread.run(Thread.java:745)\r\n\r\n \r\n\r\nThis WARN is constantly generated until I stop the range queries script.\r\n\r\nHope this helps..\r\n\r\nThanks!\r\n\r\n ","created":"2019-09-23T20:23:35.963+0000"},{"body":"For users that happen to stumble upon this use case too, I'd like to emphasize that the bug occurs only during mixed versions.\r\n\r\nOnce the cluster got fully upgraded (binaries only), the issue was gone.\r\n\r\n ","created":"2019-09-25T08:41:07.081+0000"},{"body":"Hi [~Sagges], I'm on holiday (again!) but please file a new bug report for this and assign it to me so that I remember when I return, as it is presumably a different (but very similar) bug.","created":"2019-09-25T09:19:36.637+0000"},{"body":"Hi [~benedict]\r\n\r\nDone and assigned to you! https://issues.apache.org/jira/browse/CASSANDRA-15336\r\n\r\nEnjoy your holiday. (the more the merrier) :)","created":"2019-09-25T16:39:47.865+0000"}],"conversations":[{"body":"Hi All,\r\n\r\nThis is the first time I open an issue, so apologies if I'm not following the rules properly.\r\n\r\n \r\n\r\nAfter upgrading a node from version 2.1.21 to 3.11.4, we've started seeing a lot of AbstractLocalAwareExecutorService exceptions. This happened right after the node successfully started up with the new 3.11.4 binaries. \r\n{noformat}\r\nINFO  [main] 2019-06-05 04:41:37,730 Gossiper.java:1715 - No gossip backlog; proceeding\r\nINFO  [main] 2019-06-05 04:41:38,036 NativeTransportService.java:70 - Netty using native Epoll event loop\r\nINFO  [main] 2019-06-05 04:41:38,117 Server.java:155 - Using Netty Version: [netty-buffer=netty-buffer-4.0.44.Final.452812a, netty-codec=netty-codec-4.0.44.Final.452812a, netty-codec-haproxy=netty-codec-haproxy-4.0.44.Final.452812a, netty-codec-http=netty-codec-http-4.0.44.Final.452812a, netty-codec-socks=netty-codec-socks-4.0.44.Final.452812a, netty-common=netty-common-4.0.44.Final.452812a, netty-handler=netty-handler-4.0.44.Final.452812a, netty-tcnative=netty-tcnative-1.1.33.Fork26.142ecbb, netty-transport=netty-transport-4.0.44.Final.452812a, netty-transport-native-epoll=netty-transport-native-epoll-4.0.44.Final.452812a, netty-transport-rxtx=netty-transport-rxtx-4.0.44.Final.452812a, netty-transport-sctp=netty-transport-sctp-4.0.44.Final.452812a, netty-transport-udt=netty-transport-udt-4.0.44.Final.452812a]\r\nINFO  [main] 2019-06-05 04:41:38,118 Server.java:156 - Starting listening for CQL clients on /0.0.0.0:9042 (unencrypted)...\r\nINFO  [main] 2019-06-05 04:41:38,179 CassandraDaemon.java:556 - Not starting RPC server as requested. Use JMX (StorageService->startRPCServer()) or nodetool (enablethrift) to start it\r\nINFO  [Native-Transport-Requests-21] 2019-06-05 04:41:39,145 AuthCache.java:161 - (Re)initializing PermissionsCache (validity period/update interval/max entries) (2000/2000/1000)\r\nINFO  [OptionalTasks:1] 2019-06-05 04:41:39,729 CassandraAuthorizer.java:409 - Converting legacy permissions data\r\nINFO  [HANDSHAKE-/10.10.10.8] 2019-06-05 04:41:39,808 OutboundTcpConnection.java:561 - Handshaking version with /10.10.10.8\r\nINFO  [HANDSHAKE-/10.10.10.9] 2019-06-05 04:41:39,808 OutboundTcpConnection.java:561 - Handshaking version with /10.10.10.9\r\nINFO  [HANDSHAKE-dc1_02/10.10.10.6] 2019-06-05 04:41:39,809 OutboundTcpConnection.java:561 - Handshaking version with dc1_02/10.10.10.6\r\n\r\nWARN  [ReadStage-2] 2019-06-05 04:41:39,857 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-2,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: 1\r\n        at org.apache.cassandra.db.AbstractBufferClusteringPrefix.get(AbstractBufferClusteringPrefix.java:55)\r\n        at org.apache.cassandra.db.LegacyLayout$LegacyRangeTombstoneList.serializedSizeCompound(LegacyLayout.java:2545)\r\n        at org.apache.cassandra.db.LegacyLayout$LegacyRangeTombstoneList.serializedSize(LegacyLayout.java:2522)\r\n        at org.apache.cassandra.db.LegacyLayout.serializedSizeAsLegacyPartition(LegacyLayout.java:565)\r\n        at org.apache.cassandra.db.ReadResponse$Serializer.serializedSize(ReadResponse.java:446)\r\n        at org.apache.cassandra.db.ReadResponse$Serializer.serializedSize(ReadResponse.java:352)\r\n        at org.apache.cassandra.net.MessageOut.payloadSize(MessageOut.java:171)\r\n        at org.apache.cassandra.net.OutboundTcpConnectionPool.getConnection(OutboundTcpConnectionPool.java:77)\r\n        at org.apache.cassandra.net.MessagingService.getConnection(MessagingService.java:802)\r\n        at org.apache.cassandra.net.MessagingService.sendOneWay(MessagingService.java:953)\r\n        at org.apache.cassandra.net.MessagingService.sendReply(MessagingService.java:929)\r\n        at org.apache.cassandra.db.ReadCommandVerbHandler.doVerb(ReadCommandVerbHandler.java:62)\r\n        at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:66)\r\n        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\r\n        at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:162)\r\n        at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:134)\r\n        at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:114)\r\n        at java.lang.Thread.run(Thread.java:745)\r\n {noformat}\r\n\r\n \r\n\r\nAfter several of the above warnings, the following warning appeared as well:\r\n\r\n {noformat}\r\nWARN  [ReadStage-9] 2019-06-05 04:42:04,369 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-9,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: null\r\nWARN  [ReadStage-11] 2019-06-05 04:42:04,381 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-11,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: null\r\nWARN  [ReadStage-10] 2019-06-05 04:42:04,396 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-10,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: null\r\nWARN  [ReadStage-2] 2019-06-05 04:42:04,443 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-2,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: null\r\n\r\n {noformat}\r\n \r\n\r\nThen suddenly, Validation errors appeared although *no repair was running on any of the nodes*! Checked with {{ps -ef}} command and {{nodetool compactionstats}} on the entire cluster.\r\n\r\n \r\n\r\n {noformat}\r\nERROR [ValidationExecutor:2] 2019-06-05 04:42:47,979 Validator.java:268 - Failed creating a merkle tree for [repair #e54b4090-876d-11e9-a3f4-c33d22c45471 on ks1/table1, []], /\r\n10.10.10.6 (see log for details)\r\nERROR [ValidationExecutor:2] 2019-06-05 04:42:47,979 CassandraDaemon.java:228 - Exception in thread Thread[ValidationExecutor:2,1,main]\r\njava.lang.NullPointerException: null\r\n        at org.apache.cassandra.db.compaction.CompactionManager.doValidationCompaction(CompactionManager.java:1363)\r\n        at org.apache.cassandra.db.compaction.CompactionManager.access$600(CompactionManager.java:83)\r\n        at org.apache.cassandra.db.compaction.CompactionManager$13.call(CompactionManager.java:977)\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n        at org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:81)\r\n        at java.lang.Thread.run(Thread.java:745)\r\n {noformat}\r\n \r\n\r\nFollowing those, client requests started to fail and NTR tasks started to pile up and get blocked and GC was impacted.\r\n\r\n {noformat}\r\nINFO  [ScheduledTasks:1] 2019-06-05 04:43:11,660 StatusLogger.java:51 - Native-Transport-Requests       128       197         594810        65              2725\r\n {noformat}\r\n\r\n \r\n\r\nFWIW, these are the warnings I found during startup: \r\n\r\n {noformat}\r\n-WARN in net.logstash.logback.encoder.LogstashEncoder@140e5a13 - Logback version is prior to 1.2.0.  Enabling backwards compatible encoding.  Logback 1.2.1 or greater is recommended.\r\n {noformat}\r\n\r\n \r\n\r\n {noformat}\r\nWARN  [main] 2019-06-05 08:44:18,568 NativeLibrary.java:187 - Unable to lock JVM memory (ENOMEM). This can result in part of the JVM being swapped out, especially with mmapped I/O enabled. Increase RLIMIT_MEMLOCK or run Cassandra as root.\r\nWARN  [main] 2019-06-05 08:44:18,569 StartupChecks.java:136 - jemalloc shared library could not be preloaded to speed up memory allocations\r\n\r\n \r\n\r\nWARN  [main] 2019-06-05 08:44:20,225 Optional.java:159 - Legacy auth tables credentials, users, permissions in keyspace system_auth still exist and have not been properly migrated.\r\n\r\nWARN  [MessagingService-Outgoing-dc1_03/10.10.10.4-Gossip] 2019-06-05 08:44:49,582 OutboundTcpConnection.java:486 - Seed gossip version is 8; will not connect with that version\r\nWARN  [MessagingService-Outgoing-dc2_02/10.20.20.4-Gossip] 2019-06-05 08:44:49,620 OutboundTcpConnection.java:486 - Seed gossip version is 8; will not connect with that version\r\nWARN  [MessagingService-Outgoing-dc2_01/10.20.20.1-Gossip] 2019-06-05 08:44:49,621 OutboundTcpConnection.java:486 - Seed gossip version is 8; will not connect with that version\r\nWARN  [MessagingService-Outgoing-dc2_03/10.20.20.5-Gossip] 2019-06-05 08:44:49,621 OutboundTcpConnection.java:486 - Seed gossip version is 8; will not connect with that version\r\nWARN  [GossipTasks:1] 2019-06-05 08:44:51,631 FailureDetector.java:278 - Not marking nodes down due to local pause of 30943606906 > 5000000000\r\n {noformat}\r\n\r\n \r\n\r\nWe've naturally stopped the upgrade but we still wish to upgrade from 2.1.21 and hopefully find the root cause of this matter. \r\nI'll be happy to provide additional details if needs be.\r\n\r\n \r\n\r\n ","from":"reporter","subject":"LegacyLayout RangeTombstoneList throws IndexOutOfBoundsException"},{"body":"Trying to push this up a bit because I still want to upgrade to 3.11.4 but I fear that this issue may recur.\r\n\r\nIf someone has any idea on what happened here or how to mitigate it, that'd be awesome!\r\n\r\n \r\n\r\nThanks!","from":"developer"},{"body":"We have also hit this problem today while upgrading from 2.1.16 to 3.11.4/ \r\n\r\nwe encountered this as soon as node started up with 3.11.4 \r\n\r\n \r\n\r\nWARN [ReadStage-4] 2019-08-06 02:57:57,408 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-4,5,main]: {}\r\njava.lang.NullPointerException: null\r\n\r\n \r\n\r\nERROR [Native-Transport-Requests-32] 2019-08-06 02:14:20,353 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n\r\nand the below errors continued in the logfile as long as the process was up.\r\n\r\nERROR [Native-Transport-Requests-12] 2019-08-06 03:00:47,135 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-8] 2019-08-06 03:00:48,778 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-13] 2019-08-06 03:00:57,454 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-11] 2019-08-06 03:00:57,482 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-2] 2019-08-06 03:00:58,543 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-8] 2019-08-06 03:00:58,899 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-17] 2019-08-06 03:00:59,074 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-12] 2019-08-06 03:01:08,123 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-17] 2019-08-06 03:01:19,055 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-4] 2019-08-06 03:01:20,880 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n WARN [ReadStage-13] 2019-08-06 03:01:29,983 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-13,5,main]: {}\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-2] 2019-08-06 03:01:31,119 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-6] 2019-08-06 03:01:46,262 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-15] 2019-08-06 03:01:46,520 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n WARN [ReadStage-2] 2019-08-06 03:01:48,842 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-2,5,main]: {}\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-1] 2019-08-06 03:01:50,351 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-5] 2019-08-06 03:02:06,061 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n WARN [ReadStage-8] 2019-08-06 03:02:07,616 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-8,5,main]: {}\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-17] 2019-08-06 03:02:08,384 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n ERROR [Native-Transport-Requests-5] 2019-08-06 03:02:10,244 ErrorMessage.java:384 - Unexpected exception during request\r\n java.lang.NullPointerException: null\r\n\r\n \r\n\r\nThe nodetool version says 3.11.4 and the no of connections on 9042 was similar to other nodes. The exceptions were scary that we had to call off the change. Any help and insights to this problem from the community is appreciated.","from":"developer"},{"body":"Full stack trace is as below:\r\n\r\nWARN [ReadStage-4] 2019-08-06 02:57:57,408 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-4,5,main]: {}\r\njava.lang.NullPointerException: null\r\n at org.apache.cassandra.db.LegacyLayout$LegacyRangeTombstoneList.updateDigest(LegacyLayout.java:2433) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.LegacyLayout$LegacyUnfilteredPartition.digest(LegacyLayout.java:1479) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.rows.UnfilteredRowIterators.digest(UnfilteredRowIterators.java:182) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.partitions.UnfilteredPartitionIterators.digest(UnfilteredPartitionIterators.java:263) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.ReadResponse.makeDigest(ReadResponse.java:140) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.ReadResponse.createDigestResponse(ReadResponse.java:87) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.ReadCommand.createResponse(ReadCommand.java:352) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.db.ReadCommandVerbHandler.doVerb(ReadCommandVerbHandler.java:50) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:66) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) ~[na:1.8.0_131]\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:162) ~[apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:134) [apache-cassandra-3.11.4.jar:3.11.4]\r\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:114) [apache-cassandra-3.11.4.jar:3.11.4]\r\n at java.lang.Thread.run(Thread.java:748) [na:1.8.0_131]","from":"developer"},{"body":"What we also know is that the nodes suffered some heavy tombstones recently.. could it be under specific condition like \"tombstones\" plus reading legacy version files is hitting this problem. [~Sagges] did you cluster have any tombstone references at all?","from":"developer"},{"body":"We have currently isolated this node (cut off thrift , native protocols) and trying to run upgradesstables to see if it can re-write all the files and stop logging the below message.\r\n\r\n\"WARN [ReadStage-6] 2019-08-06 10:44:09,773 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-6,5,main]: {}\r\njava.lang.NullPointerException: null\"\r\n\r\nI will keep the thread posted about its outcome, \r\n\r\n ","from":"developer"},{"body":"[~Sagges], sorry for the slow response - I missed the original filing of this ticket.\r\n\r\n[~ferozshaik552@gmail.com] it looks like your bug, while very similar, presents differently. It would be great if you could file a separate ticket.\r\n\r\nBoth of these look to be among the category of 2.1->3.0 upgrade bugs involving range deletions. To best investigate and diagnose, it would be great to start with information about the affected schema, the kinds of range tombstone deletes you perform, and preferably if you could pin down sstables that are affected and upload them somewhere private for us to access. This would help us investigate much more readily.\r\n\r\nCould you also confirm if you utilise thrift, or CQL schema? It's possible this is a compatibility issue specific to thrift.","from":"developer"},{"body":"Sure, I shall raise a separate request.","from":"developer"},{"body":"This bug appears to be similar to CASSANDRA-15263, in that a reverse query with the RTBoundCloser is the likely source of asymmetric range tombstone bounds. However in this case the problem is much easier to solve; we simply have to not assume the bounds have the same length.\r\n\r\nI have pushed a patch [here|https://github.com/belliottsmith/cassandra/tree/15172-3.0]","from":"developer"},{"body":"[~benedict], we've seen this in the wild as well, with an upgrade from 2.2.14 to 3.11.4.\r\n I am jumping in to test and review it.","from":"developer"},{"body":"\r\n||branch||circleci||asf jenkins testall||asf jenkins dtests||\r\n|[15172-3.0|https://github.com/apache/cassandra/compare/trunk...belliottsmith:15172-3.0]|[circleci|https://circleci.com/gh/belliottsmith/workflows/cassandra/tree/15172-3.0]|[!https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-testall/44//badge/icon!|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-testall/44/]|[!https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/679//badge/icon!|https://builds.apache.org/view/A-D/view/Cassandra/job/Cassandra-devbranch-dtest/679]|\r\n\r\n","from":"developer"},{"body":"Committed as 2b10a5f2b5e62f2900119a37e91637916e8b23df","from":"developer"},{"body":"Hi [~benedict]\r\n\r\nSorry for the late reply...\r\n\r\n \r\n\r\nCould you also confirm if you utilise thrift, or CQL schema? It's possible this is a compatibility issue specific to thrift.\r\n\r\nI do see there are a few thrift counter tables that were created with the WITH COMPACT STORAGE attribute.\r\n\r\n \r\n\r\nI also see that [~ferozshaik552@gmail.com] had seen WARN [ReadStage-4] 2019-08-06 02:57:57,408 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-4,5,main]: {}\r\n java.lang.NullPointerException: null\r\n\r\n\r\n\r\nWhat I mainly saw was AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-9,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: null\r\n\r\nDoes the fix fit the same scenario as Feroz's? Could my issue be thrift related?\r\n\r\n \r\n\r\nSorry again for the late reply.\r\n\r\n ","from":"developer"},{"body":"Hi [~benedict] ,\r\n\r\nI'm not sure if my comments/tickets are somehow not getting filed, but I hope not :)\r\n\r\nIs the above error (previous comment) related to this bug or is it a different one?\r\n\r\n \r\n\r\nThanks!","from":"developer"},{"body":"[~Sagges],\r\n\r\n this bug comes from existing thrift (legacy) tables, where range tombstones were used. \r\n\r\nThe NPE [~ferozshaik552@gmail.com] reported is a separate bug, despite it also being coming from legacy thrift tables with range tombstones. Unfortunately though, the fix for this ticket will not solve the NPE bug. ","from":"developer"},{"body":"Thanks for clarifying [~mck]!","from":"developer"},{"body":"Hi [~Sagges], I'm afraid I was on holiday so was unable to respond, and am now otherwise engaged for the next few weeks, but [~mck] is mostly correct. However, to clarify, the likely cause of this bug that I have established is not related to thrift or legacy tables (though atypical range tombstone use with thrift could cause it), but to communication from a 3.0 node to a 2.2 or 2.1 node, in the face of range tombstones that cover a primary key prefix.\r\n\r\nThat is to say, a schema of the form (pk, c1, c2, v), with a deletion on (pk, c1)","from":"developer"},{"body":"Thanks a lot for further clarifying [~benedict]. (I hope you enjoyed your vacation :) )\r\n\r\nJust to set my mind straight, the issue is when there are mixed versions in the cluster, so if I upgraded all binaries to 3.11, it won't recur even if I haven't upgraded the SSTables yet. Is my assumption correct?","from":"developer"},{"body":"Hi [~benedict]\r\n\r\nI tried the patch. It didn't do the trick, however, I was able to fully reproduce the bug.\r\n\r\ntl;dr\r\n\r\nThe bug does occur when running queries on range tombstones, but on top of that, those queries have to specifically be range queries.\r\n\r\n \r\n\r\n*+Steps to Reproduce:\r\n+* \r\n\r\nCREATE KEYSPACE ks1 WITH replication = \\{'class': 'NetworkTopologyStrategy', 'DC1': '3'} AND durable_writes = true;\r\n\r\n+*TABLE:*+ \r\nCREATE TABLE ks1.table1 (\r\n col1 text,\r\n col2 text,\r\n col3 text,\r\n col4 text,\r\n col5 text,\r\n col6 timestamp,\r\n data text,\r\n PRIMARY KEY ((col1, col2, col3), col4, col5, col6)\r\n);\r\n\r\n \r\n\r\nInserted ~4 million rows and created range tombstones by deleting ~1 million rows.\r\n\r\n \r\n\r\n+*Create Data*+\r\n\r\n_insert into ks1.table1 (col1, col2 , col3 , col4 , col5 , col6 , data ) VALUES ( '1', '11', '21', '1', 'a', 12312312300000, 'data');_\r\n_insert into ks1.table1 (col1, col2 , col3 , col4 , col5 , col6 , data ) VALUES ( '1', '11', '21', '2', 'a', 12312312300000, 'data');_\r\n_insert into ks1.table1 (col1, col2 , col3 , col4 , col5 , col6 , data ) VALUES ( '1', '11', '21', '3', 'a', 12312312300000, 'data');_\r\n_insert into ks1.table1 (col1, col2 , col3 , col4 , col5 , col6 , data ) VALUES ( '1', '11', '21', '4', 'a', 12312312300000, 'data');_\r\n_insert into ks1.table1 (col1, col2 , col3 , col4 , col5 , col6 , data ) VALUES ( '1', '11', '21', '5', 'a', 12312312300000, 'data');_\r\n\r\n \r\n\r\n+*Create Range Tombstones*+\r\n\r\ndelete from ks1.table1 where col1='1' and col2='11' and col3='21' and col4='1';\r\n\r\n \r\n\r\n+*Query Live Rows (no tombstones)*+\r\n\r\n_select * from ks1.table1 where col1='1' and col2='201' and col3='21' and col4='1' and col5='a' and *col6>12312312300000*;_\r\n\r\nNo issues found, everything is running properly.\r\n\r\n \r\n\r\n+*Query Range Tombstones*+\r\n\r\n_select * from ks1.table1 where col1='1' and col2='11' and col3='21' and col4='1' and col5='a' and *col6=12312312300000*;_\r\n\r\nNo issues found, everything is running properly.\r\n\r\n \r\n\r\n+BUT when running range queries:+\r\n\r\n_select * from ks1.table1 where col1='1' and col2='11' and col3='21' and col4='1' and col5='a' and *col6>12312312200000;*_\r\n\r\nWARN [ReadStage-1] 2019-09-23 14:17:10,281 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-1,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: 2\r\n at org.apache.cassandra.db.AbstractBufferClusteringPrefix.get(AbstractBufferClusteringPrefix.java:55)\r\n at org.apache.cassandra.db.LegacyLayout$LegacyRangeTombstoneList.serializedSizeCompound(LegacyLayout.java:2545)\r\n at org.apache.cassandra.db.LegacyLayout$LegacyRangeTombstoneList.serializedSize(LegacyLayout.java:2522)\r\n at org.apache.cassandra.db.LegacyLayout.serializedSizeAsLegacyPartition(LegacyLayout.java:565)\r\n at org.apache.cassandra.db.ReadResponse$Serializer.serializedSize(ReadResponse.java:446)\r\n at org.apache.cassandra.db.ReadResponse$Serializer.serializedSize(ReadResponse.java:352)\r\n at org.apache.cassandra.net.MessageOut.payloadSize(MessageOut.java:171)\r\n at org.apache.cassandra.net.OutboundTcpConnectionPool.getConnection(OutboundTcpConnectionPool.java:77)\r\n at org.apache.cassandra.net.MessagingService.getConnection(MessagingService.java:802)\r\n at org.apache.cassandra.net.MessagingService.sendOneWay(MessagingService.java:953)\r\n at org.apache.cassandra.net.MessagingService.sendReply(MessagingService.java:929)\r\n at org.apache.cassandra.db.ReadCommandVerbHandler.doVerb(ReadCommandVerbHandler.java:62)\r\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:66)\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:162)\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:134)\r\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:114)\r\n at java.lang.Thread.run(Thread.java:745)\r\n\r\n \r\n\r\nThis WARN is constantly generated until I stop the range queries script.\r\n\r\nHope this helps..\r\n\r\nThanks!\r\n\r\n ","from":"developer"},{"body":"For users that happen to stumble upon this use case too, I'd like to emphasize that the bug occurs only during mixed versions.\r\n\r\nOnce the cluster got fully upgraded (binaries only), the issue was gone.\r\n\r\n ","from":"developer"},{"body":"Hi [~Sagges], I'm on holiday (again!) but please file a new bug report for this and assign it to me so that I remember when I return, as it is presumably a different (but very similar) bug.","from":"developer"},{"body":"Hi [~benedict]\r\n\r\nDone and assigned to you! https://issues.apache.org/jira/browse/CASSANDRA-15336\r\n\r\nEnjoy your holiday. (the more the merrier) :)","from":"developer"}],"created":"2019-06-19T09:32:37.000+0000","description":"Hi All,\r\n\r\nThis is the first time I open an issue, so apologies if I'm not following the rules properly.\r\n\r\n \r\n\r\nAfter upgrading a node from version 2.1.21 to 3.11.4, we've started seeing a lot of AbstractLocalAwareExecutorService exceptions. This happened right after the node successfully started up with the new 3.11.4 binaries. \r\n{noformat}\r\nINFO  [main] 2019-06-05 04:41:37,730 Gossiper.java:1715 - No gossip backlog; proceeding\r\nINFO  [main] 2019-06-05 04:41:38,036 NativeTransportService.java:70 - Netty using native Epoll event loop\r\nINFO  [main] 2019-06-05 04:41:38,117 Server.java:155 - Using Netty Version: [netty-buffer=netty-buffer-4.0.44.Final.452812a, netty-codec=netty-codec-4.0.44.Final.452812a, netty-codec-haproxy=netty-codec-haproxy-4.0.44.Final.452812a, netty-codec-http=netty-codec-http-4.0.44.Final.452812a, netty-codec-socks=netty-codec-socks-4.0.44.Final.452812a, netty-common=netty-common-4.0.44.Final.452812a, netty-handler=netty-handler-4.0.44.Final.452812a, netty-tcnative=netty-tcnative-1.1.33.Fork26.142ecbb, netty-transport=netty-transport-4.0.44.Final.452812a, netty-transport-native-epoll=netty-transport-native-epoll-4.0.44.Final.452812a, netty-transport-rxtx=netty-transport-rxtx-4.0.44.Final.452812a, netty-transport-sctp=netty-transport-sctp-4.0.44.Final.452812a, netty-transport-udt=netty-transport-udt-4.0.44.Final.452812a]\r\nINFO  [main] 2019-06-05 04:41:38,118 Server.java:156 - Starting listening for CQL clients on /0.0.0.0:9042 (unencrypted)...\r\nINFO  [main] 2019-06-05 04:41:38,179 CassandraDaemon.java:556 - Not starting RPC server as requested. Use JMX (StorageService->startRPCServer()) or nodetool (enablethrift) to start it\r\nINFO  [Native-Transport-Requests-21] 2019-06-05 04:41:39,145 AuthCache.java:161 - (Re)initializing PermissionsCache (validity period/update interval/max entries) (2000/2000/1000)\r\nINFO  [OptionalTasks:1] 2019-06-05 04:41:39,729 CassandraAuthorizer.java:409 - Converting legacy permissions data\r\nINFO  [HANDSHAKE-/10.10.10.8] 2019-06-05 04:41:39,808 OutboundTcpConnection.java:561 - Handshaking version with /10.10.10.8\r\nINFO  [HANDSHAKE-/10.10.10.9] 2019-06-05 04:41:39,808 OutboundTcpConnection.java:561 - Handshaking version with /10.10.10.9\r\nINFO  [HANDSHAKE-dc1_02/10.10.10.6] 2019-06-05 04:41:39,809 OutboundTcpConnection.java:561 - Handshaking version with dc1_02/10.10.10.6\r\n\r\nWARN  [ReadStage-2] 2019-06-05 04:41:39,857 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-2,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: 1\r\n        at org.apache.cassandra.db.AbstractBufferClusteringPrefix.get(AbstractBufferClusteringPrefix.java:55)\r\n        at org.apache.cassandra.db.LegacyLayout$LegacyRangeTombstoneList.serializedSizeCompound(LegacyLayout.java:2545)\r\n        at org.apache.cassandra.db.LegacyLayout$LegacyRangeTombstoneList.serializedSize(LegacyLayout.java:2522)\r\n        at org.apache.cassandra.db.LegacyLayout.serializedSizeAsLegacyPartition(LegacyLayout.java:565)\r\n        at org.apache.cassandra.db.ReadResponse$Serializer.serializedSize(ReadResponse.java:446)\r\n        at org.apache.cassandra.db.ReadResponse$Serializer.serializedSize(ReadResponse.java:352)\r\n        at org.apache.cassandra.net.MessageOut.payloadSize(MessageOut.java:171)\r\n        at org.apache.cassandra.net.OutboundTcpConnectionPool.getConnection(OutboundTcpConnectionPool.java:77)\r\n        at org.apache.cassandra.net.MessagingService.getConnection(MessagingService.java:802)\r\n        at org.apache.cassandra.net.MessagingService.sendOneWay(MessagingService.java:953)\r\n        at org.apache.cassandra.net.MessagingService.sendReply(MessagingService.java:929)\r\n        at org.apache.cassandra.db.ReadCommandVerbHandler.doVerb(ReadCommandVerbHandler.java:62)\r\n        at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:66)\r\n        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\r\n        at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:162)\r\n        at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:134)\r\n        at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:114)\r\n        at java.lang.Thread.run(Thread.java:745)\r\n {noformat}\r\n\r\n \r\n\r\nAfter several of the above warnings, the following warning appeared as well:\r\n\r\n {noformat}\r\nWARN  [ReadStage-9] 2019-06-05 04:42:04,369 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-9,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: null\r\nWARN  [ReadStage-11] 2019-06-05 04:42:04,381 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-11,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: null\r\nWARN  [ReadStage-10] 2019-06-05 04:42:04,396 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-10,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: null\r\nWARN  [ReadStage-2] 2019-06-05 04:42:04,443 AbstractLocalAwareExecutorService.java:167 - Uncaught exception on thread Thread[ReadStage-2,5,main]: {}\r\njava.lang.ArrayIndexOutOfBoundsException: null\r\n\r\n {noformat}\r\n \r\n\r\nThen suddenly, Validation errors appeared although *no repair was running on any of the nodes*! Checked with {{ps -ef}} command and {{nodetool compactionstats}} on the entire cluster.\r\n\r\n \r\n\r\n {noformat}\r\nERROR [ValidationExecutor:2] 2019-06-05 04:42:47,979 Validator.java:268 - Failed creating a merkle tree for [repair #e54b4090-876d-11e9-a3f4-c33d22c45471 on ks1/table1, []], /\r\n10.10.10.6 (see log for details)\r\nERROR [ValidationExecutor:2] 2019-06-05 04:42:47,979 CassandraDaemon.java:228 - Exception in thread Thread[ValidationExecutor:2,1,main]\r\njava.lang.NullPointerException: null\r\n        at org.apache.cassandra.db.compaction.CompactionManager.doValidationCompaction(CompactionManager.java:1363)\r\n        at org.apache.cassandra.db.compaction.CompactionManager.access$600(CompactionManager.java:83)\r\n        at org.apache.cassandra.db.compaction.CompactionManager$13.call(CompactionManager.java:977)\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n        at org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:81)\r\n        at java.lang.Thread.run(Thread.java:745)\r\n {noformat}\r\n \r\n\r\nFollowing those, client requests started to fail and NTR tasks started to pile up and get blocked and GC was impacted.\r\n\r\n {noformat}\r\nINFO  [ScheduledTasks:1] 2019-06-05 04:43:11,660 StatusLogger.java:51 - Native-Transport-Requests       128       197         594810        65              2725\r\n {noformat}\r\n\r\n \r\n\r\nFWIW, these are the warnings I found during startup: \r\n\r\n {noformat}\r\n-WARN in net.logstash.logback.encoder.LogstashEncoder@140e5a13 - Logback version is prior to 1.2.0.  Enabling backwards compatible encoding.  Logback 1.2.1 or greater is recommended.\r\n {noformat}\r\n\r\n \r\n\r\n {noformat}\r\nWARN  [main] 2019-06-05 08:44:18,568 NativeLibrary.java:187 - Unable to lock JVM memory (ENOMEM). This can result in part of the JVM being swapped out, especially with mmapped I/O enabled. Increase RLIMIT_MEMLOCK or run Cassandra as root.\r\nWARN  [main] 2019-06-05 08:44:18,569 StartupChecks.java:136 - jemalloc shared library could not be preloaded to speed up memory allocations\r\n\r\n \r\n\r\nWARN  [main] 2019-06-05 08:44:20,225 Optional.java:159 - Legacy auth tables credentials, users, permissions in keyspace system_auth still exist and have not been properly migrated.\r\n\r\nWARN  [MessagingService-Outgoing-dc1_03/10.10.10.4-Gossip] 2019-06-05 08:44:49,582 OutboundTcpConnection.java:486 - Seed gossip version is 8; will not connect with that version\r\nWARN  [MessagingService-Outgoing-dc2_02/10.20.20.4-Gossip] 2019-06-05 08:44:49,620 OutboundTcpConnection.java:486 - Seed gossip version is 8; will not connect with that version\r\nWARN  [MessagingService-Outgoing-dc2_01/10.20.20.1-Gossip] 2019-06-05 08:44:49,621 OutboundTcpConnection.java:486 - Seed gossip version is 8; will not connect with that version\r\nWARN  [MessagingService-Outgoing-dc2_03/10.20.20.5-Gossip] 2019-06-05 08:44:49,621 OutboundTcpConnection.java:486 - Seed gossip version is 8; will not connect with that version\r\nWARN  [GossipTasks:1] 2019-06-05 08:44:51,631 FailureDetector.java:278 - Not marking nodes down due to local pause of 30943606906 > 5000000000\r\n {noformat}\r\n\r\n \r\n\r\nWe've naturally stopped the upgrade but we still wish to upgrade from 2.1.21 and hopefully find the root cause of this matter. \r\nI'll be happy to provide additional details if needs be.\r\n\r\n \r\n\r\n ","issue_id":"13240381","key":"CASSANDRA-15172","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2019-08-22T11:50:27.000+0000","role":"fixed_distractor","summary":"LegacyLayout RangeTombstoneList throws IndexOutOfBoundsException"} {"case_id":"13250196","cluster":"DISTRACTOR-CASSANDRA-15273","comments":[{"body":"I think you should attach some logs then we can get more information for this question. the older version of cassandra should also be attached .","created":"2019-08-13T02:40:07.887+0000"},{"body":"The cassandra starts, but the systemd cannot control it. The cause is that when the cassandra starts, the old initialization SysV script is used, in which it is obviously impossible to specify the user and group to start the service.\r\n\r\nIt's about user/group options for systemd:\r\n -----------------\r\n _[Service]_\r\n _User=cassandra_\r\n _Group=cassandra_\r\n -----------------\r\n\r\nBut since the process pid is created with the permissions of the cassandra user, and the user and group are not specified in the initialization script, the systemd consider that it uses the root to start the service (by default) and does not allow creating the pid with cassandra user permissions.\r\n ------------------\r\n _systemd[1]: New main PID 2545 does not belong to service, and PID file is not owned by root. Refusing._\r\n ------------------\r\n\r\nMore details in CVE-2018-16888 ([https://access.redhat.com/security/cve/cve-2018-16888])\r\n ------------------\r\n _It was discovered systemd does not correctly check the content of PIDFile files before using it to kill processes. When a service is run from an unprivileged user (e.g. User field set in the service file), a local attacker who is able to write to the PIDFile of the mentioned service may use this flaw to trick systemd into killing other services and/or privileged processes._\r\n ------------------","created":"2019-08-13T06:22:06.869+0000"},{"body":"We are experiencing the same! We are on  cassandra-version 3.11.1-1.\r\n\r\nWhat is the status on this?","created":"2019-08-21T07:11:37.157+0000"},{"body":"Hi,\r\n\r\nIs there any update on this? This is still a big issue.\r\n\r\nWe cannot patch our systems due to this bug. :/","created":"2019-09-26T11:12:13.167+0000"},{"body":"[~jolynch] [~jasobrown], can any of you guys assist with this issue? :)","created":"2019-09-26T11:15:37.393+0000"},{"body":"For anyone who wants a quick and dirty patch to the SysV init script, this should work.\r\n{noformat}\r\n --- /etc/rc.d/init.d/cassandra\t2017-06-19 20:09:05.000000000 +0000\r\n+++ cassandra\t2019-10-16 13:19:48.527564181 +0000\r\n@@ -69,7 +69,8 @@\r\n echo -n \"Starting Cassandra: \"\r\n [ -d `dirname \"$pid_file\"` ] || \\\r\n install -m 755 -o $CASSANDRA_OWNR -g $CASSANDRA_OWNR -d `dirname $pid_file`\r\n- su $CASSANDRA_OWNR -c \"$CASSANDRA_PROG -p $pid_file\" > $log_file 2>&1\r\n+ runuser -u $CASSANDRA_OWNR -- $CASSANDRA_PROG -p $pid_file > $log_file 2>&1 \r\n+ chown root.root $pid_file\r\n retval=$?\r\n [ $retval -eq 0 ] && touch $lock_file\r\n echo \"OK\"\r\n{noformat}\r\n\r\n","created":"2019-10-16T13:22:40.540+0000"},{"body":"Thanks [~musinsky], that fixed it for me!","created":"2019-10-16T15:00:35.169+0000"},{"body":"h5. Shouldn't cassandra register services in the systemd like other standard services?\r\n\r\nFor example in RHEL 7:\r\n{{/usr/lib/systemd/system/sshd.service}}\r\n{code:bash}\r\n[Unit]\r\nDescription=OpenSSH server daemon\r\nDocumentation=man:sshd(8) man:sshd_config(5)\r\nAfter=network.target sshd-keygen.service\r\nWants=sshd-keygen.service\r\n\r\n[Service]\r\nType=notify\r\nEnvironmentFile=/etc/sysconfig/sshd\r\nExecStart=/usr/sbin/sshd -D $OPTIONS\r\nExecReload=/bin/kill -HUP $MAINPID\r\nKillMode=process\r\nRestart=on-failure\r\nRestartSec=42s\r\n\r\n[Install]\r\nWantedBy=multi-user.target\r\n{code}\r\n{{Of course /etc/init.d/cassandra is then superfluous :)}}\r\n\r\n \r\n\r\nAs example it is {{/usr/lib/systemd/system/cassandra.service}} from OpenSUSE 15.1\r\n{code:bash}\r\n[Unit]\r\nDescription=Cassandra\r\nAfter=network.target\r\n\r\n[Service]\r\nEnvironment=CASSANDRA_HOME=/usr/share/cassandra CASSANDRA_CONF=/etc/cassandra/conf CASSANDRA_INCLUDE=/usr/share/cassandra/cassandra.in.sh\r\nEnvironmentFile=/etc/sysconfig/cassandra\r\nUser=cassandra\r\nExecStart=/usr/sbin/cassandra -f\r\nExecStopPost=/usr/bin/sleep 5 ; /usr/bin/rm -f /var/lock/subsys/cassandra\r\nStandardOutput=journal\r\nStandardError=journal\r\nLimitNOFILE=100000\r\nLimitMEMLOCK=infinity\r\nLimitNPROC=32768\r\nLimitAS=infinity\r\nSuccessExitStatus=143\r\nTimeoutStopSec=60\r\nRestart=on-failure\r\n\r\n[Install]\r\nWantedBy=multi-user.target\r\n{code}\r\n ","created":"2019-11-07T15:08:04.670+0000"},{"body":"Hi Guys, I have the same problem, in Centos7 and Redhat7 with CASSANDRA. could you help me to execute the scrip I tried but it did not work, where I must run it, first I must create a file with the scrip and then execute it so it would be the procedure .. thanks.\r\n\r\n \r\n\r\n[root@localhost ~]# service cassandra start\r\nStarting cassandra (via systemctl): Job for cassandra.service failed because a configured resource limit was exceeded. See \"systemctl status cassandra.service\" and \"journalctl -xe\" for details.\r\n[FAILED]\r\n\r\n\r\n[root@localhost ~]# systemctl status cassandra.service\r\n● cassandra.service - LSB: distributed storage system for structured data\r\n Loaded: loaded (/etc/rc.d/init.d/cassandra; bad; vendor preset: disabled)\r\n Active: failed (Result: resources) since Fri 2019-11-22 10:19:29 -05; 8s ago\r\n Docs: man:systemd-sysv-generator(8)\r\n Process: 10565 ExecStart=/etc/rc.d/init.d/cassandra start (code=exited, status=0/SUCCESS)\r\n\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: Starting LSB: distributed storage system for structured data...\r\nNov 22 10:19:29 localhost.localdomain su[10575]: (to cassandra) root on none\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: New main PID 10649 does not belong to service, and PID file is not owned by root. Refusing.\r\nNov 22 10:19:29 localhost.localdomain cassandra[10565]: Starting Cassandra: OK\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: New main PID 10649 does not belong to service, and PID file is not owned by root. Refusing.\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: Failed to start LSB: distributed storage system for structured data.\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: Unit cassandra.service entered failed state.\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: cassandra.service failed.\r\n[root@localhost ~]#","created":"2019-11-22T15:37:01.287+0000"},{"body":"The fix from [~musinsky] doesn't seem to have helped for me. I'm now seeing:\r\n\r\n{noformat}\r\nNov 25 18:19:31 test.example.com systemd[1]: Starting LSB: distributed storage system for structured data...\r\n-- Subject: Unit cassandra.service has begun start-up\r\n-- Defined-By: systemd\r\n-- Support: http://lists.freedesktop.org/mailman/listinfo/systemd-devel\r\n--\r\n-- Unit cassandra.service has begun starting up.\r\nNov 25 18:19:31 test.example.com runuser[7211]: pam_unix(runuser:session): session opened for user cassandra by (uid=0)\r\nNov 25 18:19:32 test.example.com cassandra[7201]: Starting Cassandra: OK\r\nNov 25 18:19:32 test.example.com systemd[1]: New main PID 7289 does not exist or is a zombie.\r\nNov 25 18:19:32 test.example.com systemd[1]: Failed to start LSB: distributed storage system for structured data.\r\n-- Subject: Unit cassandra.service has failed\r\n-- Defined-By: systemd\r\n-- Support: http://lists.freedesktop.org/mailman/listinfo/systemd-devel\r\n--\r\n-- Unit cassandra.service has failed.\r\n--\r\n-- The result is failed.\r\nNov 25 18:19:32 test.example.com systemd[1]: Unit cassandra.service entered failed state.\r\nNov 25 18:19:32 test.example.com systemd[1]: cassandra.service failed.\r\n{noformat}","created":"2019-11-25T23:21:55.965+0000"},{"body":"Nevermind, seems I had a different issue, my JAVA_HOME globally had changed to point to JDK 11.\r\n\r\nAdding {{JAVA_HOME=/usr/lib/jvm/java-1.8.0-openjdk}} to {{/etc/default/cassandra}}, along with the changes suggested by [~musinsky], were sufficient to get Cassandra running again.\r\n\r\nThanks!","created":"2019-11-25T23:34:36.473+0000"},{"body":"Hello, the solution for the start of the service was good, but I had to modify the command for RHEL 7.7 but when I execute the stop, the message status=1/FAILURE appears and that I also changed the command syntax.\r\n\r\nI have changed the command in the script this way:\r\n\r\nrunuser $ CASSANDRA_OWNR -c \"kill` cat $ pid_file` \"\r\n\r\nI have tried to make changes in the scritp but it has not worked, I just have to change the value from 1 to 0 in this part:\r\n\r\nif [$ retval -eq 3]; then\r\nI miss \"OK\"\r\nelse\r\necho \"ERROR: could not stop $ NAME\"\r\n# exit 1\r\nexit 0\r\n\r\nIf you already have the solution please let her know, I remain attentive to your comments.\r\n\r\n ","created":"2019-11-26T22:55:27.837+0000"},{"body":"As a followup to [~musinsky] and as an Answer for [~acerna], I fixed both {{start}} and {{stop}} as following:\r\n{noformat}\r\n--- /etc/rc.d/init.d/cassandra.old 2020-02-10 15:23:38.852120933 +0100\r\n+++ /etc/rc.d/init.d/cassandra.new 2020-02-10 15:23:10.610076622 +0100\r\n@@ -71,7 +71,9 @@\r\n chown -R cassandra:cassandra \"$(dirname \"$log_file\")\"\r\n [ -d `dirname \"$pid_file\"` ] || \\\r\n install -m 755 -o $CASSANDRA_OWNR -g $CASSANDRA_OWNR -d `dirname $pid_file`\r\n- su $CASSANDRA_OWNR -c \"$CASSANDRA_PROG -p $pid_file\" > $log_file 2>&1\r\n+ # su $CASSANDRA_OWNR -c \"$CASSANDRA_PROG -p $pid_file\" > $log_file 2>&1\r\n+ runuser -u $CASSANDRA_OWNR -- $CASSANDRA_PROG -p $pid_file > $log_file 2>&1\r\n+ chown root.root $pid_file\r\n retval=$?\r\n [ $retval -eq 0 ] && touch $lock_file\r\n echo \"OK\"\r\n@@ -79,7 +81,8 @@\r\n stop)\r\n # Cassandra shutdown\r\n echo -n \"Shutdown Cassandra: \"\r\n- su $CASSANDRA_OWNR -c \"kill `cat $pid_file`\"\r\n+ # su $CASSANDRA_OWNR -c \"kill `cat $pid_file`\"\r\n+ runuser -u $CASSANDRA_OWNR -- kill `cat $pid_file`\r\n retval=$?\r\n [ $retval -eq 0 ] && rm -f $lock_file\r\n for t in `seq 40`; do\r\n{noformat}","created":"2020-02-10T15:03:07.529+0000"},{"body":"If someone wants to distill this into a patch, we would be interested! cc/ [~mshuler] and [~mck2] sense they are wrestling some build stuff right now.","created":"2020-02-12T04:19:45.555+0000"},{"body":"I've attached a patch against the {{cassandra-3.11}} branch which should work.\r\n\r\n","created":"2020-02-19T18:56:16.474+0000"},{"body":"Thank you [~pioto]\r\n\r\nI followed the Code mentioned in [^0001-Fix-Red-Hat-init-script-on-newer-systemd-versions.patch]\r\n\r\n^Then run the command \"systemctl daemon-reload\" and it worked.^\r\n\r\nRegards,\r\n\r\nSwapnil","created":"2020-02-25T01:55:35.471+0000"},{"body":"||branch||circleci||\r\n|[cassandra_2.2_15273|https://github.com/apache/cassandra/compare/cassandra-2.2...thelastpickle:mck/cassandra-2.2_15273]|[circleci|https://circleci.com/gh/thelastpickle/workflows/cassandra/tree/mck%2Fcassandra-2.2_15273]|\r\n|[cassandra_3.0_15273|https://github.com/apache/cassandra/compare/cassandra-3.0...thelastpickle:mck/cassandra-3.0_15273]|[circleci|https://circleci.com/gh/thelastpickle/workflows/cassandra/tree/mck%2Fcassandra-3.0_15273]|\r\n|[cassandra_3.11_15273|https://github.com/apache/cassandra/compare/cassandra-3.11...thelastpickle:mck/cassandra-3.11_15273]|[circleci|https://circleci.com/gh/thelastpickle/workflows/cassandra/tree/mck%2Fcassandra-3.11_15273]|\r\n|[trunk_15273|https://github.com/apache/cassandra/compare/trunk...thelastpickle:mck/trunk_15273]|[circleci|https://circleci.com/gh/thelastpickle/workflows/cassandra/tree/mck%2Ftrunk_15273]|\r\n\r\nTo test this, rather than following the instructions in https://github.com/apache/cassandra/tree/trunk/redhat , I've taken the following approach:\r\n{code}\r\ndocker image rm -f `docker images -f label=org.cassandra.buildenv=centos -q`\r\ndocker build --build-arg CASSANDRA_GIT_URL=${CASSANDRA_GIT_URL} -f docker/centos7-image.docker docker/\r\nmkdir /tmp/cassandra-rpms\r\ndocker run --rm -v /tmp/cassandra-rpms:/dist `docker images -f label=org.cassandra.buildenv=centos -q` /home/build/build-rpms.sh ${CASSANDRA_GIT_BRANCH}\r\n\r\ndocker run -v /tmp/cassandra-rpms:/dist -it centos /bin/sh\r\nyum install java-1.8.0-openjdk\r\nrpm -ivh /dist/cassandra-4.0~alpha4-20200303git90a391e.noarch.rpm\r\nrm /etc/security/limits.d/cassandra.conf # this should be fixed\r\n/etc/init.d/cassandra start\r\ntail -F /var/log/cassandra/system.log\r\n{code}\r\n\r\nThis is the same approach as we used when cutting and staging releases, ref https://github.com/thelastpickle/cassandra-builds/blob/mck/14970_sha512-checksums/cassandra-release/prepare_release.sh#L299-L302\r\n\r\nNote, we are short of packaging experience in the community and really do need extra contributors on this front. Replacing the old initialisation SysV script with a service script (CASSANDRA-13148), and testing the building of packages in the CI pipeline, as well as others like CASSANDRA-13433 , are obvious subsequent steps that are needed. ","created":"2020-03-03T16:04:14.421+0000"},{"body":"Committed as 9105dcd99e61537e8d177b41b7d38c5569412230","created":"2020-03-03T21:55:46.236+0000"},{"body":"It would be appropriate to introduce systemd support for 4.0 instead \r\nSee also:\r\nhttps://issues.apache.org/jira/browse/CASSANDRA-13148\r\n\r\nhttps://github.com/apache/cassandra/pull/398\r\n\r\n ","created":"2020-03-10T09:08:35.158+0000"}],"conversations":[{"body":"After update systemd with  fixed vulnerability https://access.redhat.com/security/cve/cve-2018-16888, the cassandra service does not start correctly.\r\n\r\nEnvironment: RHEL 7, systemd-219-67.el7_7.1, cassandra-3.11.4-1 (https://www.apache.org/dist/cassandra/redhat/311x/cassandra-3.11.4-1.noarch.rpm)\r\n\r\n---------------------------------------------------------------\r\n\r\nsystemctl status cassandra\r\n● cassandra.service - LSB: distributed storage system for structured data\r\n Loaded: loaded (/etc/rc.d/init.d/cassandra; bad; vendor preset: disabled)\r\n Active: failed (Result: resources) since Fri 2019-08-09 17:20:26 MSK; 1s ago\r\n Docs: man:systemd-sysv-generator(8)\r\n Process: 2414 ExecStop=/etc/rc.d/init.d/cassandra stop (code=exited, status=0/SUCCESS)\r\n Process: 2463 ExecStart=/etc/rc.d/init.d/cassandra start (code=exited, status=0/SUCCESS)\r\n Main PID: 1884 (code=exited, status=143)\r\n\r\nAug 09 17:20:23 desktop43.example.com systemd[1]: Unit cassandra.service entered failed state.\r\nAug 09 17:20:23 desktop43.example.com systemd[1]: cassandra.service failed.\r\nAug 09 17:20:23 desktop43.example.com systemd[1]: Starting LSB: distributed storage system for structured data...\r\nAug 09 17:20:23 desktop43.example.com su[2473]: (to cassandra) root on none\r\nAug 09 17:20:26 desktop43.example.com cassandra[2463]: Starting Cassandra: OK\r\nAug 09 17:20:26 desktop43.example.com systemd[1]: New main PID 2545 does not belong to service, and PID file is not owned by root. Refusing.\r\nAug 09 17:20:26 desktop43.example.com systemd[1]: New main PID 2545 does not belong to service, and PID file is not owned by root. Refusing.\r\nAug 09 17:20:26 desktop43.example.com systemd[1]: Failed to start LSB: distributed storage system for structured data.\r\nAug 09 17:20:26 desktop43.example.com systemd[1]: Unit cassandra.service entered failed state.\r\nAug 09 17:20:26 desktop43.example.com systemd[1]: cassandra.service failed.","from":"reporter","subject":"cassandra does not start with new systemd version"},{"body":"I think you should attach some logs then we can get more information for this question. the older version of cassandra should also be attached .","from":"developer"},{"body":"The cassandra starts, but the systemd cannot control it. The cause is that when the cassandra starts, the old initialization SysV script is used, in which it is obviously impossible to specify the user and group to start the service.\r\n\r\nIt's about user/group options for systemd:\r\n -----------------\r\n _[Service]_\r\n _User=cassandra_\r\n _Group=cassandra_\r\n -----------------\r\n\r\nBut since the process pid is created with the permissions of the cassandra user, and the user and group are not specified in the initialization script, the systemd consider that it uses the root to start the service (by default) and does not allow creating the pid with cassandra user permissions.\r\n ------------------\r\n _systemd[1]: New main PID 2545 does not belong to service, and PID file is not owned by root. Refusing._\r\n ------------------\r\n\r\nMore details in CVE-2018-16888 ([https://access.redhat.com/security/cve/cve-2018-16888])\r\n ------------------\r\n _It was discovered systemd does not correctly check the content of PIDFile files before using it to kill processes. When a service is run from an unprivileged user (e.g. User field set in the service file), a local attacker who is able to write to the PIDFile of the mentioned service may use this flaw to trick systemd into killing other services and/or privileged processes._\r\n ------------------","from":"developer"},{"body":"We are experiencing the same! We are on  cassandra-version 3.11.1-1.\r\n\r\nWhat is the status on this?","from":"developer"},{"body":"Hi,\r\n\r\nIs there any update on this? This is still a big issue.\r\n\r\nWe cannot patch our systems due to this bug. :/","from":"developer"},{"body":"[~jolynch] [~jasobrown], can any of you guys assist with this issue? :)","from":"developer"},{"body":"For anyone who wants a quick and dirty patch to the SysV init script, this should work.\r\n{noformat}\r\n --- /etc/rc.d/init.d/cassandra\t2017-06-19 20:09:05.000000000 +0000\r\n+++ cassandra\t2019-10-16 13:19:48.527564181 +0000\r\n@@ -69,7 +69,8 @@\r\n echo -n \"Starting Cassandra: \"\r\n [ -d `dirname \"$pid_file\"` ] || \\\r\n install -m 755 -o $CASSANDRA_OWNR -g $CASSANDRA_OWNR -d `dirname $pid_file`\r\n- su $CASSANDRA_OWNR -c \"$CASSANDRA_PROG -p $pid_file\" > $log_file 2>&1\r\n+ runuser -u $CASSANDRA_OWNR -- $CASSANDRA_PROG -p $pid_file > $log_file 2>&1 \r\n+ chown root.root $pid_file\r\n retval=$?\r\n [ $retval -eq 0 ] && touch $lock_file\r\n echo \"OK\"\r\n{noformat}\r\n\r\n","from":"developer"},{"body":"Thanks [~musinsky], that fixed it for me!","from":"developer"},{"body":"h5. Shouldn't cassandra register services in the systemd like other standard services?\r\n\r\nFor example in RHEL 7:\r\n{{/usr/lib/systemd/system/sshd.service}}\r\n{code:bash}\r\n[Unit]\r\nDescription=OpenSSH server daemon\r\nDocumentation=man:sshd(8) man:sshd_config(5)\r\nAfter=network.target sshd-keygen.service\r\nWants=sshd-keygen.service\r\n\r\n[Service]\r\nType=notify\r\nEnvironmentFile=/etc/sysconfig/sshd\r\nExecStart=/usr/sbin/sshd -D $OPTIONS\r\nExecReload=/bin/kill -HUP $MAINPID\r\nKillMode=process\r\nRestart=on-failure\r\nRestartSec=42s\r\n\r\n[Install]\r\nWantedBy=multi-user.target\r\n{code}\r\n{{Of course /etc/init.d/cassandra is then superfluous :)}}\r\n\r\n \r\n\r\nAs example it is {{/usr/lib/systemd/system/cassandra.service}} from OpenSUSE 15.1\r\n{code:bash}\r\n[Unit]\r\nDescription=Cassandra\r\nAfter=network.target\r\n\r\n[Service]\r\nEnvironment=CASSANDRA_HOME=/usr/share/cassandra CASSANDRA_CONF=/etc/cassandra/conf CASSANDRA_INCLUDE=/usr/share/cassandra/cassandra.in.sh\r\nEnvironmentFile=/etc/sysconfig/cassandra\r\nUser=cassandra\r\nExecStart=/usr/sbin/cassandra -f\r\nExecStopPost=/usr/bin/sleep 5 ; /usr/bin/rm -f /var/lock/subsys/cassandra\r\nStandardOutput=journal\r\nStandardError=journal\r\nLimitNOFILE=100000\r\nLimitMEMLOCK=infinity\r\nLimitNPROC=32768\r\nLimitAS=infinity\r\nSuccessExitStatus=143\r\nTimeoutStopSec=60\r\nRestart=on-failure\r\n\r\n[Install]\r\nWantedBy=multi-user.target\r\n{code}\r\n ","from":"developer"},{"body":"Hi Guys, I have the same problem, in Centos7 and Redhat7 with CASSANDRA. could you help me to execute the scrip I tried but it did not work, where I must run it, first I must create a file with the scrip and then execute it so it would be the procedure .. thanks.\r\n\r\n \r\n\r\n[root@localhost ~]# service cassandra start\r\nStarting cassandra (via systemctl): Job for cassandra.service failed because a configured resource limit was exceeded. See \"systemctl status cassandra.service\" and \"journalctl -xe\" for details.\r\n[FAILED]\r\n\r\n\r\n[root@localhost ~]# systemctl status cassandra.service\r\n● cassandra.service - LSB: distributed storage system for structured data\r\n Loaded: loaded (/etc/rc.d/init.d/cassandra; bad; vendor preset: disabled)\r\n Active: failed (Result: resources) since Fri 2019-11-22 10:19:29 -05; 8s ago\r\n Docs: man:systemd-sysv-generator(8)\r\n Process: 10565 ExecStart=/etc/rc.d/init.d/cassandra start (code=exited, status=0/SUCCESS)\r\n\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: Starting LSB: distributed storage system for structured data...\r\nNov 22 10:19:29 localhost.localdomain su[10575]: (to cassandra) root on none\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: New main PID 10649 does not belong to service, and PID file is not owned by root. Refusing.\r\nNov 22 10:19:29 localhost.localdomain cassandra[10565]: Starting Cassandra: OK\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: New main PID 10649 does not belong to service, and PID file is not owned by root. Refusing.\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: Failed to start LSB: distributed storage system for structured data.\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: Unit cassandra.service entered failed state.\r\nNov 22 10:19:29 localhost.localdomain systemd[1]: cassandra.service failed.\r\n[root@localhost ~]#","from":"developer"},{"body":"The fix from [~musinsky] doesn't seem to have helped for me. I'm now seeing:\r\n\r\n{noformat}\r\nNov 25 18:19:31 test.example.com systemd[1]: Starting LSB: distributed storage system for structured data...\r\n-- Subject: Unit cassandra.service has begun start-up\r\n-- Defined-By: systemd\r\n-- Support: http://lists.freedesktop.org/mailman/listinfo/systemd-devel\r\n--\r\n-- Unit cassandra.service has begun starting up.\r\nNov 25 18:19:31 test.example.com runuser[7211]: pam_unix(runuser:session): session opened for user cassandra by (uid=0)\r\nNov 25 18:19:32 test.example.com cassandra[7201]: Starting Cassandra: OK\r\nNov 25 18:19:32 test.example.com systemd[1]: New main PID 7289 does not exist or is a zombie.\r\nNov 25 18:19:32 test.example.com systemd[1]: Failed to start LSB: distributed storage system for structured data.\r\n-- Subject: Unit cassandra.service has failed\r\n-- Defined-By: systemd\r\n-- Support: http://lists.freedesktop.org/mailman/listinfo/systemd-devel\r\n--\r\n-- Unit cassandra.service has failed.\r\n--\r\n-- The result is failed.\r\nNov 25 18:19:32 test.example.com systemd[1]: Unit cassandra.service entered failed state.\r\nNov 25 18:19:32 test.example.com systemd[1]: cassandra.service failed.\r\n{noformat}","from":"developer"},{"body":"Nevermind, seems I had a different issue, my JAVA_HOME globally had changed to point to JDK 11.\r\n\r\nAdding {{JAVA_HOME=/usr/lib/jvm/java-1.8.0-openjdk}} to {{/etc/default/cassandra}}, along with the changes suggested by [~musinsky], were sufficient to get Cassandra running again.\r\n\r\nThanks!","from":"developer"},{"body":"Hello, the solution for the start of the service was good, but I had to modify the command for RHEL 7.7 but when I execute the stop, the message status=1/FAILURE appears and that I also changed the command syntax.\r\n\r\nI have changed the command in the script this way:\r\n\r\nrunuser $ CASSANDRA_OWNR -c \"kill` cat $ pid_file` \"\r\n\r\nI have tried to make changes in the scritp but it has not worked, I just have to change the value from 1 to 0 in this part:\r\n\r\nif [$ retval -eq 3]; then\r\nI miss \"OK\"\r\nelse\r\necho \"ERROR: could not stop $ NAME\"\r\n# exit 1\r\nexit 0\r\n\r\nIf you already have the solution please let her know, I remain attentive to your comments.\r\n\r\n ","from":"developer"},{"body":"As a followup to [~musinsky] and as an Answer for [~acerna], I fixed both {{start}} and {{stop}} as following:\r\n{noformat}\r\n--- /etc/rc.d/init.d/cassandra.old 2020-02-10 15:23:38.852120933 +0100\r\n+++ /etc/rc.d/init.d/cassandra.new 2020-02-10 15:23:10.610076622 +0100\r\n@@ -71,7 +71,9 @@\r\n chown -R cassandra:cassandra \"$(dirname \"$log_file\")\"\r\n [ -d `dirname \"$pid_file\"` ] || \\\r\n install -m 755 -o $CASSANDRA_OWNR -g $CASSANDRA_OWNR -d `dirname $pid_file`\r\n- su $CASSANDRA_OWNR -c \"$CASSANDRA_PROG -p $pid_file\" > $log_file 2>&1\r\n+ # su $CASSANDRA_OWNR -c \"$CASSANDRA_PROG -p $pid_file\" > $log_file 2>&1\r\n+ runuser -u $CASSANDRA_OWNR -- $CASSANDRA_PROG -p $pid_file > $log_file 2>&1\r\n+ chown root.root $pid_file\r\n retval=$?\r\n [ $retval -eq 0 ] && touch $lock_file\r\n echo \"OK\"\r\n@@ -79,7 +81,8 @@\r\n stop)\r\n # Cassandra shutdown\r\n echo -n \"Shutdown Cassandra: \"\r\n- su $CASSANDRA_OWNR -c \"kill `cat $pid_file`\"\r\n+ # su $CASSANDRA_OWNR -c \"kill `cat $pid_file`\"\r\n+ runuser -u $CASSANDRA_OWNR -- kill `cat $pid_file`\r\n retval=$?\r\n [ $retval -eq 0 ] && rm -f $lock_file\r\n for t in `seq 40`; do\r\n{noformat}","from":"developer"},{"body":"If someone wants to distill this into a patch, we would be interested! cc/ [~mshuler] and [~mck2] sense they are wrestling some build stuff right now.","from":"developer"},{"body":"I've attached a patch against the {{cassandra-3.11}} branch which should work.\r\n\r\n","from":"developer"},{"body":"Thank you [~pioto]\r\n\r\nI followed the Code mentioned in [^0001-Fix-Red-Hat-init-script-on-newer-systemd-versions.patch]\r\n\r\n^Then run the command \"systemctl daemon-reload\" and it worked.^\r\n\r\nRegards,\r\n\r\nSwapnil","from":"developer"},{"body":"||branch||circleci||\r\n|[cassandra_2.2_15273|https://github.com/apache/cassandra/compare/cassandra-2.2...thelastpickle:mck/cassandra-2.2_15273]|[circleci|https://circleci.com/gh/thelastpickle/workflows/cassandra/tree/mck%2Fcassandra-2.2_15273]|\r\n|[cassandra_3.0_15273|https://github.com/apache/cassandra/compare/cassandra-3.0...thelastpickle:mck/cassandra-3.0_15273]|[circleci|https://circleci.com/gh/thelastpickle/workflows/cassandra/tree/mck%2Fcassandra-3.0_15273]|\r\n|[cassandra_3.11_15273|https://github.com/apache/cassandra/compare/cassandra-3.11...thelastpickle:mck/cassandra-3.11_15273]|[circleci|https://circleci.com/gh/thelastpickle/workflows/cassandra/tree/mck%2Fcassandra-3.11_15273]|\r\n|[trunk_15273|https://github.com/apache/cassandra/compare/trunk...thelastpickle:mck/trunk_15273]|[circleci|https://circleci.com/gh/thelastpickle/workflows/cassandra/tree/mck%2Ftrunk_15273]|\r\n\r\nTo test this, rather than following the instructions in https://github.com/apache/cassandra/tree/trunk/redhat , I've taken the following approach:\r\n{code}\r\ndocker image rm -f `docker images -f label=org.cassandra.buildenv=centos -q`\r\ndocker build --build-arg CASSANDRA_GIT_URL=${CASSANDRA_GIT_URL} -f docker/centos7-image.docker docker/\r\nmkdir /tmp/cassandra-rpms\r\ndocker run --rm -v /tmp/cassandra-rpms:/dist `docker images -f label=org.cassandra.buildenv=centos -q` /home/build/build-rpms.sh ${CASSANDRA_GIT_BRANCH}\r\n\r\ndocker run -v /tmp/cassandra-rpms:/dist -it centos /bin/sh\r\nyum install java-1.8.0-openjdk\r\nrpm -ivh /dist/cassandra-4.0~alpha4-20200303git90a391e.noarch.rpm\r\nrm /etc/security/limits.d/cassandra.conf # this should be fixed\r\n/etc/init.d/cassandra start\r\ntail -F /var/log/cassandra/system.log\r\n{code}\r\n\r\nThis is the same approach as we used when cutting and staging releases, ref https://github.com/thelastpickle/cassandra-builds/blob/mck/14970_sha512-checksums/cassandra-release/prepare_release.sh#L299-L302\r\n\r\nNote, we are short of packaging experience in the community and really do need extra contributors on this front. Replacing the old initialisation SysV script with a service script (CASSANDRA-13148), and testing the building of packages in the CI pipeline, as well as others like CASSANDRA-13433 , are obvious subsequent steps that are needed. ","from":"developer"},{"body":"Committed as 9105dcd99e61537e8d177b41b7d38c5569412230","from":"developer"},{"body":"It would be appropriate to introduce systemd support for 4.0 instead \r\nSee also:\r\nhttps://issues.apache.org/jira/browse/CASSANDRA-13148\r\n\r\nhttps://github.com/apache/cassandra/pull/398\r\n\r\n ","from":"developer"}],"created":"2019-08-12T07:04:51.000+0000","description":"After update systemd with  fixed vulnerability https://access.redhat.com/security/cve/cve-2018-16888, the cassandra service does not start correctly.\r\n\r\nEnvironment: RHEL 7, systemd-219-67.el7_7.1, cassandra-3.11.4-1 (https://www.apache.org/dist/cassandra/redhat/311x/cassandra-3.11.4-1.noarch.rpm)\r\n\r\n---------------------------------------------------------------\r\n\r\nsystemctl status cassandra\r\n● cassandra.service - LSB: distributed storage system for structured data\r\n Loaded: loaded (/etc/rc.d/init.d/cassandra; bad; vendor preset: disabled)\r\n Active: failed (Result: resources) since Fri 2019-08-09 17:20:26 MSK; 1s ago\r\n Docs: man:systemd-sysv-generator(8)\r\n Process: 2414 ExecStop=/etc/rc.d/init.d/cassandra stop (code=exited, status=0/SUCCESS)\r\n Process: 2463 ExecStart=/etc/rc.d/init.d/cassandra start (code=exited, status=0/SUCCESS)\r\n Main PID: 1884 (code=exited, status=143)\r\n\r\nAug 09 17:20:23 desktop43.example.com systemd[1]: Unit cassandra.service entered failed state.\r\nAug 09 17:20:23 desktop43.example.com systemd[1]: cassandra.service failed.\r\nAug 09 17:20:23 desktop43.example.com systemd[1]: Starting LSB: distributed storage system for structured data...\r\nAug 09 17:20:23 desktop43.example.com su[2473]: (to cassandra) root on none\r\nAug 09 17:20:26 desktop43.example.com cassandra[2463]: Starting Cassandra: OK\r\nAug 09 17:20:26 desktop43.example.com systemd[1]: New main PID 2545 does not belong to service, and PID file is not owned by root. Refusing.\r\nAug 09 17:20:26 desktop43.example.com systemd[1]: New main PID 2545 does not belong to service, and PID file is not owned by root. Refusing.\r\nAug 09 17:20:26 desktop43.example.com systemd[1]: Failed to start LSB: distributed storage system for structured data.\r\nAug 09 17:20:26 desktop43.example.com systemd[1]: Unit cassandra.service entered failed state.\r\nAug 09 17:20:26 desktop43.example.com systemd[1]: cassandra.service failed.","issue_id":"13250196","key":"CASSANDRA-15273","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2020-03-03T21:55:46.000+0000","role":"fixed_distractor","summary":"cassandra does not start with new systemd version"} {"case_id":"13268261","cluster":"DISTRACTOR-CASSANDRA-15426","comments":[{"body":"Seems like the problem was introduced with CASSANDRA-15053 patch. Note that {{NativeLibrary.tryOpenDirectory()}} is not implemented on Windows and always returns -1.","created":"2019-11-14T10:41:39.318+0000"},{"body":"Jeff supplied a patch that I imported [here|https://github.com/dineshjoshi/cassandra/tree/15426-3.0] for Cassandra 3.0. Circle is running [here|https://circleci.com/workflow-run/19a0c18c-ff2a-4131-8d44-63dd21bb84fd]. Once I've confirmed that this fixes the issue in 3.0 and doesn't cause regressions, I will forward port it to 3.11 and trunk.\r\n\r\n[~ashish_singh2@bmc.com] [~Andrew.Kostousov@gmail.com] any chance I could get both of you to try this out in your respective environments. I don't have Windows Server installation at hand.","created":"2020-01-28T19:46:16.729+0000"},{"body":"I verified that the patch works on 3.0, 3.11 and trunk. +1.","created":"2020-01-29T02:37:31.527+0000"},{"body":"Committed as 4e6c0fadfe69419e0f6f7c3fa5960bdcd90d510a and merged up to 3.11 and trunk.","created":"2020-01-29T02:39:24.017+0000"},{"body":"[~djoshi]: HI Dinesh , Where can i get the patch to try out the fix, Also, Please share steps to apply the patch to existing installation.","created":"2020-02-12T05:15:35.002+0000"},{"body":"[~ashish_singh2@bmc.com], this fix is in 3.11.6 which will be released in a few days. If you cannot wait that long, you can apply my patch from the [3.11 branch|https://github.com/apache/cassandra/commit/f2a36e17cd7977044e461e2ff134e2a703f6843d].","created":"2020-02-12T19:44:01.598+0000"}],"conversations":[{"body":"Cassandra 3.11.5 fails to start on Windows server 2012 R2. with following error trace.\r\n\r\nCassandra 3.11.4 doesn't fail on Windows 2012 R2. \r\n\r\n   \r\n\r\norg.apache.cassandra.io.FSReadError: java.io.IOException: Invalid folder descriptor trying to create log replica C:\\Users\\Administrator\\Downloads\\apache-cassandra-3.11.5-bin.tar\\apache-cassandra-3.11.5-bin\\apache-cassandra-3.11.5\\data\\data\\system\\local-7ad54392bcdd35a684174e047860b377\r\n at org.apache.cassandra.db.lifecycle.LogReplica.create(LogReplica.java:58) ~[apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.lifecycle.LogReplicaSet.maybeCreateReplica(LogReplicaSet.java:86) ~[apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.lifecycle.LogFile.makeRecord(LogFile.java:311) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.lifecycle.LogFile.add(LogFile.java:283) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.lifecycle.LogTransaction.trackNew(LogTransaction.java:139) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.lifecycle.LifecycleTransaction.trackNew(LifecycleTransaction.java:528) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.(BigTableWriter.java:81) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.io.sstable.format.big.BigFormat$WriterFactory.open(BigFormat.java:92) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.io.sstable.format.SSTableWriter.create(SSTableWriter.java:102) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.io.sstable.SimpleSSTableMultiWriter.create(SimpleSSTableMultiWriter.java:119) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.compaction.AbstractCompactionStrategy.createSSTableMultiWriter(AbstractCompactionStrategy.java:588) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.compaction.CompactionStrategyManager.createSSTableMultiWriter(CompactionStrategyManager.java:1027) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.ColumnFamilyStore.createSSTableMultiWriter(ColumnFamilyStore.java:532) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.Memtable$FlushRunnable.createFlushWriter(Memtable.java:504) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.Memtable$FlushRunnable.(Memtable.java:443) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.Memtable$FlushRunnable.(Memtable.java:420) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.Memtable.createFlushRunnables(Memtable.java:307) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.Memtable.flushRunnables(Memtable.java:298) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1153) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1118) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149) [na:1.8.0_161]\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624) [na:1.8.0_161]\r\n at org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:84) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at java.lang.Thread.run(Thread.java:748) ~[na:1.8.0_161]\r\nCaused by: java.io.IOException: Invalid folder descriptor trying to","from":"reporter","subject":"Cassandra 3.11.5 fails to start on Windows "},{"body":"Seems like the problem was introduced with CASSANDRA-15053 patch. Note that {{NativeLibrary.tryOpenDirectory()}} is not implemented on Windows and always returns -1.","from":"developer"},{"body":"Jeff supplied a patch that I imported [here|https://github.com/dineshjoshi/cassandra/tree/15426-3.0] for Cassandra 3.0. Circle is running [here|https://circleci.com/workflow-run/19a0c18c-ff2a-4131-8d44-63dd21bb84fd]. Once I've confirmed that this fixes the issue in 3.0 and doesn't cause regressions, I will forward port it to 3.11 and trunk.\r\n\r\n[~ashish_singh2@bmc.com] [~Andrew.Kostousov@gmail.com] any chance I could get both of you to try this out in your respective environments. I don't have Windows Server installation at hand.","from":"developer"},{"body":"I verified that the patch works on 3.0, 3.11 and trunk. +1.","from":"developer"},{"body":"Committed as 4e6c0fadfe69419e0f6f7c3fa5960bdcd90d510a and merged up to 3.11 and trunk.","from":"developer"},{"body":"[~djoshi]: HI Dinesh , Where can i get the patch to try out the fix, Also, Please share steps to apply the patch to existing installation.","from":"developer"},{"body":"[~ashish_singh2@bmc.com], this fix is in 3.11.6 which will be released in a few days. If you cannot wait that long, you can apply my patch from the [3.11 branch|https://github.com/apache/cassandra/commit/f2a36e17cd7977044e461e2ff134e2a703f6843d].","from":"developer"}],"created":"2019-11-14T09:47:54.000+0000","description":"Cassandra 3.11.5 fails to start on Windows server 2012 R2. with following error trace.\r\n\r\nCassandra 3.11.4 doesn't fail on Windows 2012 R2. \r\n\r\n   \r\n\r\norg.apache.cassandra.io.FSReadError: java.io.IOException: Invalid folder descriptor trying to create log replica C:\\Users\\Administrator\\Downloads\\apache-cassandra-3.11.5-bin.tar\\apache-cassandra-3.11.5-bin\\apache-cassandra-3.11.5\\data\\data\\system\\local-7ad54392bcdd35a684174e047860b377\r\n at org.apache.cassandra.db.lifecycle.LogReplica.create(LogReplica.java:58) ~[apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.lifecycle.LogReplicaSet.maybeCreateReplica(LogReplicaSet.java:86) ~[apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.lifecycle.LogFile.makeRecord(LogFile.java:311) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.lifecycle.LogFile.add(LogFile.java:283) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.lifecycle.LogTransaction.trackNew(LogTransaction.java:139) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.lifecycle.LifecycleTransaction.trackNew(LifecycleTransaction.java:528) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.io.sstable.format.big.BigTableWriter.(BigTableWriter.java:81) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.io.sstable.format.big.BigFormat$WriterFactory.open(BigFormat.java:92) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.io.sstable.format.SSTableWriter.create(SSTableWriter.java:102) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.io.sstable.SimpleSSTableMultiWriter.create(SimpleSSTableMultiWriter.java:119) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.compaction.AbstractCompactionStrategy.createSSTableMultiWriter(AbstractCompactionStrategy.java:588) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.compaction.CompactionStrategyManager.createSSTableMultiWriter(CompactionStrategyManager.java:1027) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.ColumnFamilyStore.createSSTableMultiWriter(ColumnFamilyStore.java:532) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.Memtable$FlushRunnable.createFlushWriter(Memtable.java:504) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.Memtable$FlushRunnable.(Memtable.java:443) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.Memtable$FlushRunnable.(Memtable.java:420) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.Memtable.createFlushRunnables(Memtable.java:307) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.Memtable.flushRunnables(Memtable.java:298) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1153) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1118) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149) [na:1.8.0_161]\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624) [na:1.8.0_161]\r\n at org.apache.cassandra.concurrent.NamedThreadFactory.lambda$threadLocalDeallocator$0(NamedThreadFactory.java:84) [apache-cassandra-3.11.5.jar:3.11.5]\r\n at java.lang.Thread.run(Thread.java:748) ~[na:1.8.0_161]\r\nCaused by: java.io.IOException: Invalid folder descriptor trying to","issue_id":"13268261","key":"CASSANDRA-15426","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2020-01-29T02:39:24.000+0000","role":"fixed_distractor","summary":"Cassandra 3.11.5 fails to start on Windows "} {"case_id":"13304387","cluster":"DISTRACTOR-CASSANDRA-15805","comments":[{"body":"To understand why this happens, let me write down the atoms that the example of the description generates on 2.X (using a simplified representation that I hope is clear enough):\r\n{noformat}\r\natom1: RT([A:_, A:X:b:_])@1, // beginning of all 'A' rows to beginning of A:X's b column\r\natom2: Cell(A:X:)@2, // row marker for A:X\r\natom3: Cell(A:X:a=foo)@2, // value of a in A:X\r\natom4: RT([A:X:b:_, A:X:b:!])@3, // collection tombstone for b in A:X's\r\natom5: RT([A:X:b:!, A:!])@1, // remainder of covering RT, from end of b in A:X to end of all 'A' rows\r\natom6: Cell(A:X:c=bar)@2 // value of c in A:X\r\n{noformat}\r\nThose atoms are deserialized into {{LegacyCell}} and {{LegacyRangeTombstone}} on 3.X as:\r\n{noformat}\r\natom1: RT(Bound(INCL_START_BOUND(A), collection=null)-Bound(EXCL_END_BOUND(A:B), collection=null), deletedAt=1, localDeletion=1589204864)\r\natom2: LegacyCell(REGULAR, name=Cellname(clustering=A:X, column=null, collElt=null), v=, ts=2, ldt=2147483647, ttl=0)\r\natom3: LegacyCell(REGULAR, name=Cellname(clustering=A:X, column=a, collElt=null), v=foo, ts=2, ldt=2147483647, ttl=0)\r\natom4: RT(Bound(INCL_START_BOUND(A:X), collection=b)-Bound(INCL_END_BOUND(A:X), collection=b), deletedAt=3, localDeletion=1589204864)\r\natom5: RT(Bound(EXCL_START_BOUND(A:X), collection=null)-Bound(INCL_END_BOUND(A), collection=null), deletedAt=1, localDeletion=1589204864)\r\natom6: LegacyCell(REGULAR, name=Cellname(clustering=A:X, column=c, collElt=null), v=bar, ts=2, ldt=2147483647, ttl=0)\r\n{noformat}\r\n\r\nI'll point out that those are a direct translation of the 2.X atoms except for {{atom1}} and {{atom5}} that are slightly different:\r\n* instead of {{atom1}} stopping at the beginning of the row {{b}} column, it extends to the end of the row.\r\n* and instead of {{atom5}} staring after that {{b}} column, it starts after the row. Do note however that the order of atoms is still the one above, so that atom is effectively out-of-order.\r\n\r\nThe reason for those differences is the logic [at the beginning of {{LegacyLayout.RangeTombstone}}|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/LegacyLayout.java#L1883], whose comment is trying to explain, but is basically due to the legacy layer having to map all 2.X RTs into either a 3.X range tombstone (so one over multiple rows), a row tombstone or a collection one.\r\n\r\nAnyway, as mentioned above, the problem is that {{atom5}} is out of order. What currently happens is that when {{atom5}} is encountered by {{UnfilteredDeserialized.OldFormatDeserializer}}, it will be passed to the {{CellGrouper}} currently grouping the row, and will end up in the [{{CellGrouper#addGenericTombstone}} method|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/LegacyLayout.java#L1544]. But, because that atom starts strictly after the row being grouped, the method returns {{false}} and the row is generated a first time. Later, we get {{atom6}} which restarts the row with the value of column {{c}}, after which it is generated a second time.\r\n","created":"2020-05-12T15:44:15.195+0000"},{"body":"There is probably a few variations for how to fix this, but what feels the more intuitive to me is to:\r\n # slightly modify the ctor of {{LegacyRangeTombstone}} so the {{atom5}} of my previous comment use an inclusive start. Mostly because that make the rest of the logic a bit simpler imo (we can still assume that when we get a atom whose cluster strictly sort after the currently grouper row, we're done with that row).\r\n # modify {{UnfilteredDeserializer.OldFormatDeserializer}} so that when, while grouping a row, it encounters a RT that covers it, is \"splits\" that into a row tombstone covering the row, and push back the handling of the rest of the tombstone to when the row is truly finished.\r\n\r\nI've pushed a patch doing so for 3.0 below (thanks to [~marcuse] for triggering CI on this):\r\n||branch||unit tests||dtests||jvm dtests||jvm upgrade dtest||\r\n|[https://github.com/pcmanus/cassandra/commits/C-15805-3.0]|[utests|https://circleci.com/gh/krummas/cassandra/3289]|[vnodes|https://circleci.com/gh/krummas/cassandra/3292] [no-vnodes|https://circleci.com/gh/krummas/cassandra/3293]|[jvm dtests|https://circleci.com/gh/krummas/cassandra/3290]|[upgrade dtests|https://circleci.com/gh/krummas/cassandra/3294]|\r\n\r\nI'll note that the branch contains another small fix, that is a lot less important. Namely, [the return at the beginning of {{CellGrouper#addCollectionTombstone}}|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/LegacyLayout.java#L1516] should be {{true}}, not {{false}} as it currently is. The test in question checks if a collection tombstone happens to not be selected by the query we're decoding data for. If it isn't included, we can ignore the tombstone, so we can/should return, but not with {{false}} as that imply the row is finished, which it probably isn't.\r\n\r\nNow the reason I say that last problem is less important is that in practice, only thrift queries should run into this (since CQL queries queries all column effectively) so even if we duplicate the row here, it won't matter when the result is converted back to thrift (besides, having a collection tombstone implies that this is a thrift query on a CQL table, which is dodgy in the first place). Anyway, the code is still obviously wrong and the fix trivial, so included it it nonetheless (in a separate commit).","created":"2020-05-13T12:51:32.842+0000"},{"body":"This LGTM, two minor ignoreable comments;\r\n\r\n* lazily initialise the {{outOfOrderAtom}} queue in {{AtomIterator}} as it will be infrequently used (and maybe name it {{outOfOrderAtoms}} )\r\n* consume the peeked atom ({{atoms.next()}}) before calling {{atoms.pushOutOfOrder}} in {{UnfilteredIterator#readRow}} to make it more obvious we can't consume the atom we just pushed","created":"2020-05-15T10:52:10.598+0000"},{"body":"Nit, feel free to ignore: in {{UnfilteredDeserializer}} on line 701, redundant {{this}}. Otherwise pretty straightforward and LGTM.","created":"2020-05-18T18:26:43.188+0000"},{"body":"Thanks for the review. I addressed the comments, squash-cleaned, 'merged' into 3.11 and started CI (first try at https://ci-cassandra.apache.org, not sure how that will go).\r\n\r\n||branch||CI||\r\n| [3.0|https://github.com/pcmanus/cassandra/commits/C-15805-3.0] | [ci-cassandra #134|https://ci-cassandra.apache.org/job/Cassandra-devbranch/134/] |\r\n| [3.11|https://github.com/pcmanus/cassandra/commits/C-15805-3.11] | [ci-cassandra #135|https://ci-cassandra.apache.org/job/Cassandra-devbranch/135/] |\r\n","created":"2020-05-19T09:42:39.111+0000"},{"body":"[~slebresne], i've re-run #131 for you as https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/134/pipeline (taking away the stress and cdc stages, as they don't exist in 3.0) and #132 as https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/135/pipeline (just a retry).\r\n\r\n(The CI failure in #132 is addressed in CASSANDRA-15826.)","created":"2020-05-26T09:42:53.287+0000"}],"conversations":[{"body":"The legacy reading code ({{LegacyLayout}} and {{UnfilteredDeserializer.OldFormatDeserializer}}) does not handle correctly the case where a range tombstone covering multiple rows interacts with a collection tombstone.\r\n\r\nA simple example of this problem is if one runs on 2.X:\r\n{noformat}\r\nCREATE TABLE t (\r\n k int,\r\n c1 text,\r\n c2 text,\r\n a text,\r\n b set,\r\n c text,\r\n PRIMARY KEY((k), c1, c2)\r\n);\r\n\r\n// Delete all rows where c1 is 'A'\r\nDELETE FROM t USING TIMESTAMP 1 WHERE k = 0 AND c1 = 'A';\r\n// Inserts a row covered by that previous range tombstone\r\nINSERT INTO t(k, c1, c2, a, b, c) VALUES (0, 'A', 'X', 'foo', {'whatever'}, 'bar') USING TIMESTAMP 2;\r\n// Delete the collection of that previously inserted row\r\nDELETE b FROM t USING TIMESTAMP 3 WHERE k = 0 AND c1 = 'A' and c2 = 'X';\r\n{noformat}\r\n\r\nIf the following is ran on 2.X (with everything either flushed in the same table or compacted together), then this will result in the inserted row being duplicated (one part containing the {{a}} column, the other the {{c}} one).\r\n\r\nI will note that this is _not_ a duplicate of CASSANDRA-15789 and this reproduce even with the fix to {{LegacyLayout}} of this ticket. That said, the additional code added to CASSANDRA-15789 to force merging duplicated rows if they are produced _will_ end up fixing this as a consequence (assuming there is no variation of this problem that leads to other visible issues than duplicated rows). That said, I \"think\" we'd still rather fix the source of the issue.\r\n","from":"reporter","subject":"Potential duplicate rows on 2.X->3.X upgrade when multi-rows range tombstones interacts with collection tombstones"},{"body":"To understand why this happens, let me write down the atoms that the example of the description generates on 2.X (using a simplified representation that I hope is clear enough):\r\n{noformat}\r\natom1: RT([A:_, A:X:b:_])@1, // beginning of all 'A' rows to beginning of A:X's b column\r\natom2: Cell(A:X:)@2, // row marker for A:X\r\natom3: Cell(A:X:a=foo)@2, // value of a in A:X\r\natom4: RT([A:X:b:_, A:X:b:!])@3, // collection tombstone for b in A:X's\r\natom5: RT([A:X:b:!, A:!])@1, // remainder of covering RT, from end of b in A:X to end of all 'A' rows\r\natom6: Cell(A:X:c=bar)@2 // value of c in A:X\r\n{noformat}\r\nThose atoms are deserialized into {{LegacyCell}} and {{LegacyRangeTombstone}} on 3.X as:\r\n{noformat}\r\natom1: RT(Bound(INCL_START_BOUND(A), collection=null)-Bound(EXCL_END_BOUND(A:B), collection=null), deletedAt=1, localDeletion=1589204864)\r\natom2: LegacyCell(REGULAR, name=Cellname(clustering=A:X, column=null, collElt=null), v=, ts=2, ldt=2147483647, ttl=0)\r\natom3: LegacyCell(REGULAR, name=Cellname(clustering=A:X, column=a, collElt=null), v=foo, ts=2, ldt=2147483647, ttl=0)\r\natom4: RT(Bound(INCL_START_BOUND(A:X), collection=b)-Bound(INCL_END_BOUND(A:X), collection=b), deletedAt=3, localDeletion=1589204864)\r\natom5: RT(Bound(EXCL_START_BOUND(A:X), collection=null)-Bound(INCL_END_BOUND(A), collection=null), deletedAt=1, localDeletion=1589204864)\r\natom6: LegacyCell(REGULAR, name=Cellname(clustering=A:X, column=c, collElt=null), v=bar, ts=2, ldt=2147483647, ttl=0)\r\n{noformat}\r\n\r\nI'll point out that those are a direct translation of the 2.X atoms except for {{atom1}} and {{atom5}} that are slightly different:\r\n* instead of {{atom1}} stopping at the beginning of the row {{b}} column, it extends to the end of the row.\r\n* and instead of {{atom5}} staring after that {{b}} column, it starts after the row. Do note however that the order of atoms is still the one above, so that atom is effectively out-of-order.\r\n\r\nThe reason for those differences is the logic [at the beginning of {{LegacyLayout.RangeTombstone}}|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/LegacyLayout.java#L1883], whose comment is trying to explain, but is basically due to the legacy layer having to map all 2.X RTs into either a 3.X range tombstone (so one over multiple rows), a row tombstone or a collection one.\r\n\r\nAnyway, as mentioned above, the problem is that {{atom5}} is out of order. What currently happens is that when {{atom5}} is encountered by {{UnfilteredDeserialized.OldFormatDeserializer}}, it will be passed to the {{CellGrouper}} currently grouping the row, and will end up in the [{{CellGrouper#addGenericTombstone}} method|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/LegacyLayout.java#L1544]. But, because that atom starts strictly after the row being grouped, the method returns {{false}} and the row is generated a first time. Later, we get {{atom6}} which restarts the row with the value of column {{c}}, after which it is generated a second time.\r\n","from":"developer"},{"body":"There is probably a few variations for how to fix this, but what feels the more intuitive to me is to:\r\n # slightly modify the ctor of {{LegacyRangeTombstone}} so the {{atom5}} of my previous comment use an inclusive start. Mostly because that make the rest of the logic a bit simpler imo (we can still assume that when we get a atom whose cluster strictly sort after the currently grouper row, we're done with that row).\r\n # modify {{UnfilteredDeserializer.OldFormatDeserializer}} so that when, while grouping a row, it encounters a RT that covers it, is \"splits\" that into a row tombstone covering the row, and push back the handling of the rest of the tombstone to when the row is truly finished.\r\n\r\nI've pushed a patch doing so for 3.0 below (thanks to [~marcuse] for triggering CI on this):\r\n||branch||unit tests||dtests||jvm dtests||jvm upgrade dtest||\r\n|[https://github.com/pcmanus/cassandra/commits/C-15805-3.0]|[utests|https://circleci.com/gh/krummas/cassandra/3289]|[vnodes|https://circleci.com/gh/krummas/cassandra/3292] [no-vnodes|https://circleci.com/gh/krummas/cassandra/3293]|[jvm dtests|https://circleci.com/gh/krummas/cassandra/3290]|[upgrade dtests|https://circleci.com/gh/krummas/cassandra/3294]|\r\n\r\nI'll note that the branch contains another small fix, that is a lot less important. Namely, [the return at the beginning of {{CellGrouper#addCollectionTombstone}}|https://github.com/apache/cassandra/blob/cassandra-3.0/src/java/org/apache/cassandra/db/LegacyLayout.java#L1516] should be {{true}}, not {{false}} as it currently is. The test in question checks if a collection tombstone happens to not be selected by the query we're decoding data for. If it isn't included, we can ignore the tombstone, so we can/should return, but not with {{false}} as that imply the row is finished, which it probably isn't.\r\n\r\nNow the reason I say that last problem is less important is that in practice, only thrift queries should run into this (since CQL queries queries all column effectively) so even if we duplicate the row here, it won't matter when the result is converted back to thrift (besides, having a collection tombstone implies that this is a thrift query on a CQL table, which is dodgy in the first place). Anyway, the code is still obviously wrong and the fix trivial, so included it it nonetheless (in a separate commit).","from":"developer"},{"body":"This LGTM, two minor ignoreable comments;\r\n\r\n* lazily initialise the {{outOfOrderAtom}} queue in {{AtomIterator}} as it will be infrequently used (and maybe name it {{outOfOrderAtoms}} )\r\n* consume the peeked atom ({{atoms.next()}}) before calling {{atoms.pushOutOfOrder}} in {{UnfilteredIterator#readRow}} to make it more obvious we can't consume the atom we just pushed","from":"developer"},{"body":"Nit, feel free to ignore: in {{UnfilteredDeserializer}} on line 701, redundant {{this}}. Otherwise pretty straightforward and LGTM.","from":"developer"},{"body":"Thanks for the review. I addressed the comments, squash-cleaned, 'merged' into 3.11 and started CI (first try at https://ci-cassandra.apache.org, not sure how that will go).\r\n\r\n||branch||CI||\r\n| [3.0|https://github.com/pcmanus/cassandra/commits/C-15805-3.0] | [ci-cassandra #134|https://ci-cassandra.apache.org/job/Cassandra-devbranch/134/] |\r\n| [3.11|https://github.com/pcmanus/cassandra/commits/C-15805-3.11] | [ci-cassandra #135|https://ci-cassandra.apache.org/job/Cassandra-devbranch/135/] |\r\n","from":"developer"},{"body":"[~slebresne], i've re-run #131 for you as https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/134/pipeline (taking away the stress and cdc stages, as they don't exist in 3.0) and #132 as https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/135/pipeline (just a retry).\r\n\r\n(The CI failure in #132 is addressed in CASSANDRA-15826.)","from":"developer"}],"created":"2020-05-12T15:12:55.000+0000","description":"The legacy reading code ({{LegacyLayout}} and {{UnfilteredDeserializer.OldFormatDeserializer}}) does not handle correctly the case where a range tombstone covering multiple rows interacts with a collection tombstone.\r\n\r\nA simple example of this problem is if one runs on 2.X:\r\n{noformat}\r\nCREATE TABLE t (\r\n k int,\r\n c1 text,\r\n c2 text,\r\n a text,\r\n b set,\r\n c text,\r\n PRIMARY KEY((k), c1, c2)\r\n);\r\n\r\n// Delete all rows where c1 is 'A'\r\nDELETE FROM t USING TIMESTAMP 1 WHERE k = 0 AND c1 = 'A';\r\n// Inserts a row covered by that previous range tombstone\r\nINSERT INTO t(k, c1, c2, a, b, c) VALUES (0, 'A', 'X', 'foo', {'whatever'}, 'bar') USING TIMESTAMP 2;\r\n// Delete the collection of that previously inserted row\r\nDELETE b FROM t USING TIMESTAMP 3 WHERE k = 0 AND c1 = 'A' and c2 = 'X';\r\n{noformat}\r\n\r\nIf the following is ran on 2.X (with everything either flushed in the same table or compacted together), then this will result in the inserted row being duplicated (one part containing the {{a}} column, the other the {{c}} one).\r\n\r\nI will note that this is _not_ a duplicate of CASSANDRA-15789 and this reproduce even with the fix to {{LegacyLayout}} of this ticket. That said, the additional code added to CASSANDRA-15789 to force merging duplicated rows if they are produced _will_ end up fixing this as a consequence (assuming there is no variation of this problem that leads to other visible issues than duplicated rows). That said, I \"think\" we'd still rather fix the source of the issue.\r\n","issue_id":"13304387","key":"CASSANDRA-15805","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2020-05-27T15:23:57.000+0000","role":"fixed_distractor","summary":"Potential duplicate rows on 2.X->3.X upgrade when multi-rows range tombstones interacts with collection tombstones"} {"case_id":"13311749","cluster":"DISTRACTOR-CASSANDRA-15878","comments":[{"body":"I've pushed a commit with a potential fix and updated unit tests: [https://github.com/apache/cassandra/pull/653/commits/7a53846a217102143ae56416ebcf534c59de93e6]\r\n\r\n[~jolynch], I'd love to have your input on this since you reviewed the original ticket that brought this change.\r\n\r\nAre there cases I'm not seeing where the dc name would be useful to check?","created":"2020-06-26T16:09:45.342+0000"},{"body":"Hi [~adejanovski] thanks for the mention! I will try to page back in context on this change and review this weekend, feel free to put me as a reviewer (although I see mck already got to it :-) ).","created":"2020-06-27T21:14:41.107+0000"},{"body":"bq. feel free to put me as a reviewer (although I see mck already got to it ).\r\n\r\nBest with your input on this [~jolynch], as you've got a better idea on the original work.\r\nmy two cents, is it would help if the unit tests could cover all scenarios, including those where we're dependent on racks to identify legacy, e.g. \"us-west-2\". ","created":"2020-06-28T10:17:53.102+0000"},{"body":"[~adejanovski] I agree with you that the rack check is sufficient to determine that we have an accidental mixed mode cluster, but I think we can salvage the datacenter check as well which imo provides a nice second layer of defense against a very easy to make mistake. Specifically, if we see a datacenter with no number (e.g. just \"us-east\" or just \"eu-west\", and usingLegacyNaming is set to False then we could reject it as invalid and raise an exception. Yes, if users had custom suffixes that included numbers we would miss it, but then the rack verification would catch it.\r\n\r\nI'm fine with the proposed patch as I agree the rack checks are sufficient, but I'd also be fine with fixing the datacenter check to be:\r\n{noformat}\r\nboolean dcUsesLegacyFormat = dc.matches(\"[a-z]+-[a-z]+$\");\r\nif (dcUsesLegacyFormat != usingLegacyNaming)\r\n valid = false;\r\n{noformat}\r\nSo for example if someone started up with usingLegacyNaming=false, and the node sees another node with datacenter \"us-east\" or \"eu-west\", then we fail with the error to prevent a split brained cluster. If someone had a custom datacenter suffix of -1, we would be foiled, and the rack check would have to catch the operator error.","created":"2020-06-28T23:58:53.217+0000"},{"body":"Thanks for the feedback [~jolynch].\r\n\r\nI've added back the DC name check and adjusted it as suggested. I provided accurate informations on which case we're actually covering now with this check.\r\n\r\n[~mck], I've reintroduced the unit tests that I had deleted and changed the assertions where needed.\r\n\r\nYou can check the changes here: [https://github.com/apache/cassandra/compare/trunk...thelastpickle:CASSANDRA-15878]\r\n\r\nLet me know what you think.","created":"2020-06-29T06:00:57.762+0000"},{"body":"Patch looks good to me. Move the ticket into 'patch submitted' if you're ready for me to commit it.","created":"2020-06-29T15:46:40.937+0000"},{"body":"Committed as [6fc8920889e8537a1f56f45e6c966b3d18325fbb |https://github.com/apache/cassandra/commit/6fc8920889e8537a1f56f45e6c966b3d18325fbb].","created":"2020-06-29T16:26:56.113+0000"}],"conversations":[{"body":"CASSANDRA-7839 changed the way the EC2 DC/Rack naming was handled in the Ec2Snitch to match AWS conventions.\r\n\r\nThe \"legacy\" mode was introduced to allow upgrades from Cassandra 3.0/3.x and keep the same naming as before (while the \"standard\" mode uses the new naming convention).\r\n\r\nWhen performing an upgrade in the us-west-2 region, the second node failed to start with the following exception:\r\n\r\n \r\n{code:java}\r\nERROR [main] 2020-06-16 09:14:42,218 Ec2Snitch.java:210 - This ec2-enabled snitch appears to be using the legacy naming scheme for regions, but existing nodes in cluster are using the opposite: region(s) = [us-west-2], availability zone(s) = [2a]. Please check the ec2_naming_scheme property in the cassandra-rackdc.properties configuration file for more details.\r\nERROR [main] 2020-06-16 09:14:42,219 CassandraDaemon.java:789 - Exception encountered during startup\r\njava.lang.IllegalStateException: null\r\n\tat org.apache.cassandra.service.StorageService.validateEndpointSnitch(StorageService.java:573)\r\n\tat org.apache.cassandra.service.StorageService.checkForEndpointCollision(StorageService.java:530)\r\n\tat org.apache.cassandra.service.StorageService.prepareToJoin(StorageService.java:800)\r\n\tat org.apache.cassandra.service.StorageService.initServer(StorageService.java:659)\r\n\tat org.apache.cassandra.service.StorageService.initServer(StorageService.java:610)\r\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:373)\r\n\tat org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:650)\r\n\tat org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:767)\r\n{code}\r\n \r\n\r\nThe exception leads back to [this piece of code|https://github.com/apache/cassandra/blob/cassandra-4.0-alpha4/src/java/org/apache/cassandra/locator/Ec2Snitch.java#L183-L185].\r\n\r\nAfter adding some logging, it turned out the DC name of the first upgraded node was considered invalid as a legacy one:\r\n{code:java}\r\nINFO [main] 2020-06-16 09:14:42,216 Ec2Snitch.java:183 - Detected DC us-west-2\r\nINFO [main] 2020-06-16 09:14:42,217 Ec2Snitch.java:185 - dcUsesLegacyFormat=false / usingLegacyNaming=true\r\nERROR [main] 2020-06-16 09:14:42,217 Ec2Snitch.java:188 - Invalid DC name us-west-2\r\n{code}\r\n \r\n\r\nThe problem is that the regex that's used to identify legacy dc names will match both old and new names : \r\n{code:java}\r\nboolean dcUsesLegacyFormat = !dc.matches(\"[a-z]+-[a-z].+-[\\\\d].*\");\r\n{code}\r\nKnowing that some dc names didn't change between the two modes (us-west-2 for example), I don't see how we can use the dc names to detect if the legacy mode is being used by other nodes in the cluster.\r\n  \r\n The rack names on the other hand are totally different in the legacy and standard modes and can be used to detect mismatching settings.\r\n  \r\n My go to fix would be to drop the check on datacenters by removing the following lines: [https://github.com/apache/cassandra/blob/cassandra-4.0-alpha4/src/java/org/apache/cassandra/locator/Ec2Snitch.java#L172-L186]","from":"reporter","subject":"Ec2Snitch fails on upgrade in legacy mode"},{"body":"I've pushed a commit with a potential fix and updated unit tests: [https://github.com/apache/cassandra/pull/653/commits/7a53846a217102143ae56416ebcf534c59de93e6]\r\n\r\n[~jolynch], I'd love to have your input on this since you reviewed the original ticket that brought this change.\r\n\r\nAre there cases I'm not seeing where the dc name would be useful to check?","from":"developer"},{"body":"Hi [~adejanovski] thanks for the mention! I will try to page back in context on this change and review this weekend, feel free to put me as a reviewer (although I see mck already got to it :-) ).","from":"developer"},{"body":"bq. feel free to put me as a reviewer (although I see mck already got to it ).\r\n\r\nBest with your input on this [~jolynch], as you've got a better idea on the original work.\r\nmy two cents, is it would help if the unit tests could cover all scenarios, including those where we're dependent on racks to identify legacy, e.g. \"us-west-2\". ","from":"developer"},{"body":"[~adejanovski] I agree with you that the rack check is sufficient to determine that we have an accidental mixed mode cluster, but I think we can salvage the datacenter check as well which imo provides a nice second layer of defense against a very easy to make mistake. Specifically, if we see a datacenter with no number (e.g. just \"us-east\" or just \"eu-west\", and usingLegacyNaming is set to False then we could reject it as invalid and raise an exception. Yes, if users had custom suffixes that included numbers we would miss it, but then the rack verification would catch it.\r\n\r\nI'm fine with the proposed patch as I agree the rack checks are sufficient, but I'd also be fine with fixing the datacenter check to be:\r\n{noformat}\r\nboolean dcUsesLegacyFormat = dc.matches(\"[a-z]+-[a-z]+$\");\r\nif (dcUsesLegacyFormat != usingLegacyNaming)\r\n valid = false;\r\n{noformat}\r\nSo for example if someone started up with usingLegacyNaming=false, and the node sees another node with datacenter \"us-east\" or \"eu-west\", then we fail with the error to prevent a split brained cluster. If someone had a custom datacenter suffix of -1, we would be foiled, and the rack check would have to catch the operator error.","from":"developer"},{"body":"Thanks for the feedback [~jolynch].\r\n\r\nI've added back the DC name check and adjusted it as suggested. I provided accurate informations on which case we're actually covering now with this check.\r\n\r\n[~mck], I've reintroduced the unit tests that I had deleted and changed the assertions where needed.\r\n\r\nYou can check the changes here: [https://github.com/apache/cassandra/compare/trunk...thelastpickle:CASSANDRA-15878]\r\n\r\nLet me know what you think.","from":"developer"},{"body":"Patch looks good to me. Move the ticket into 'patch submitted' if you're ready for me to commit it.","from":"developer"},{"body":"Committed as [6fc8920889e8537a1f56f45e6c966b3d18325fbb |https://github.com/apache/cassandra/commit/6fc8920889e8537a1f56f45e6c966b3d18325fbb].","from":"developer"}],"created":"2020-06-16T15:12:49.000+0000","description":"CASSANDRA-7839 changed the way the EC2 DC/Rack naming was handled in the Ec2Snitch to match AWS conventions.\r\n\r\nThe \"legacy\" mode was introduced to allow upgrades from Cassandra 3.0/3.x and keep the same naming as before (while the \"standard\" mode uses the new naming convention).\r\n\r\nWhen performing an upgrade in the us-west-2 region, the second node failed to start with the following exception:\r\n\r\n \r\n{code:java}\r\nERROR [main] 2020-06-16 09:14:42,218 Ec2Snitch.java:210 - This ec2-enabled snitch appears to be using the legacy naming scheme for regions, but existing nodes in cluster are using the opposite: region(s) = [us-west-2], availability zone(s) = [2a]. Please check the ec2_naming_scheme property in the cassandra-rackdc.properties configuration file for more details.\r\nERROR [main] 2020-06-16 09:14:42,219 CassandraDaemon.java:789 - Exception encountered during startup\r\njava.lang.IllegalStateException: null\r\n\tat org.apache.cassandra.service.StorageService.validateEndpointSnitch(StorageService.java:573)\r\n\tat org.apache.cassandra.service.StorageService.checkForEndpointCollision(StorageService.java:530)\r\n\tat org.apache.cassandra.service.StorageService.prepareToJoin(StorageService.java:800)\r\n\tat org.apache.cassandra.service.StorageService.initServer(StorageService.java:659)\r\n\tat org.apache.cassandra.service.StorageService.initServer(StorageService.java:610)\r\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:373)\r\n\tat org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:650)\r\n\tat org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:767)\r\n{code}\r\n \r\n\r\nThe exception leads back to [this piece of code|https://github.com/apache/cassandra/blob/cassandra-4.0-alpha4/src/java/org/apache/cassandra/locator/Ec2Snitch.java#L183-L185].\r\n\r\nAfter adding some logging, it turned out the DC name of the first upgraded node was considered invalid as a legacy one:\r\n{code:java}\r\nINFO [main] 2020-06-16 09:14:42,216 Ec2Snitch.java:183 - Detected DC us-west-2\r\nINFO [main] 2020-06-16 09:14:42,217 Ec2Snitch.java:185 - dcUsesLegacyFormat=false / usingLegacyNaming=true\r\nERROR [main] 2020-06-16 09:14:42,217 Ec2Snitch.java:188 - Invalid DC name us-west-2\r\n{code}\r\n \r\n\r\nThe problem is that the regex that's used to identify legacy dc names will match both old and new names : \r\n{code:java}\r\nboolean dcUsesLegacyFormat = !dc.matches(\"[a-z]+-[a-z].+-[\\\\d].*\");\r\n{code}\r\nKnowing that some dc names didn't change between the two modes (us-west-2 for example), I don't see how we can use the dc names to detect if the legacy mode is being used by other nodes in the cluster.\r\n  \r\n The rack names on the other hand are totally different in the legacy and standard modes and can be used to detect mismatching settings.\r\n  \r\n My go to fix would be to drop the check on datacenters by removing the following lines: [https://github.com/apache/cassandra/blob/cassandra-4.0-alpha4/src/java/org/apache/cassandra/locator/Ec2Snitch.java#L172-L186]","issue_id":"13311749","key":"CASSANDRA-15878","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2020-06-29T16:26:56.000+0000","role":"fixed_distractor","summary":"Ec2Snitch fails on upgrade in legacy mode"} {"case_id":"13320752","cluster":"DISTRACTOR-CASSANDRA-16008","comments":[{"body":"Would be very helpful to have this addressed as it is currently blocking new Docker containers of the Cassandra 2.2 tag - https://github.com/docker-library/cassandra/issues/215","created":"2020-09-23T10:50:04.724+0000"},{"body":"Applied CASSANDRA-7066 fix in C* 3.0/3.11 to {{DatabaseDescriptor.java}} broken by commit {{257fb03}}:\r\n\r\n* 2.2 PR - https://github.com/apache/cassandra/pull/757\r\n* 2.2 Patch - [^CASSANDRA-16008-2.2.txt] \r\n\r\nThis only applies to the 2.2 branch since it's already fixed in C* 3.x/4.0.","created":"2020-09-23T11:28:43.897+0000"},{"body":"Committed, thanks.","created":"2020-09-23T15:08:57.364+0000"}],"conversations":[{"body":"After the upgrade to 2.2.17, Cassandra fails to start with the following error:\r\n{noformat}\r\nINFO 20:28:57 JVM Arguments: [-Dcom.sun.management.jmxremote.port=7199, -Dcom.sun.management.jmxremote.ssl=false, -Dcom.sun.management.jmxremote.authenticate=false, -ea, -javaagent:/opt/cassandra/lib/jamm-0.3.0.jar, -XX:+CMSClassUnloadingEnabled, -XX:+UseThreadPriorities, -XX:ThreadPriorityPolicy=42, -Xms128m, -Xmx128m, -Xmn32m, -XX:+HeapDumpOnOutOfMemoryError, -Xss256k, -XX:StringTableSize=1000003, -XX:+UseParNewGC, -XX:+UseConcMarkSweepGC, -XX:+CMSParallelRemarkEnabled, -XX:SurvivorRatio=8, -XX:MaxTenuringThreshold=1, -XX:CMSInitiatingOccupancyFraction=75, -XX:+UseCMSInitiatingOccupancyOnly, -XX:+UseTLAB, -XX:+PerfDisableSharedMem, -XX:CompileCommandFile=/etc/cassandra/hotspot_compiler, -XX:CMSWaitDuration=10000, -XX:+CMSParallelInitialMarkEnabled, -XX:+CMSEdenChunksRecordAlways, -XX:CMSWaitDuration=10000, -XX:+PrintGCDetails, -XX:+PrintGCDateStamps, -XX:+PrintHeapAtGC, -XX:+PrintTenuringDistribution, -XX:+PrintGCApplicationStoppedTime, -XX:+PrintPromotionFailure, -Xloggc:/opt/cassandra/logs/gc.log, -XX:+UseGCLogFileRotation, -XX:NumberOfGCLogFiles=10, -XX:GCLogFileSize=10M, -Djava.net.preferIPv4Stack=true, -Dcassandra.jmx.local.port=7199, -XX:+DisableExplicitGC, -Djava.library.path=/opt/cassandra/lib/sigar-bin, -Dcassandra.libjemalloc=/usr/lib/x86_64-linux-gnu/libjemalloc.so.1, -XX:OnOutOfMemoryError=kill -9 %p, -Dlogback.configurationFile=logback.xml, -Dcassandra.logdir=/opt/cassandra/logs, -Dcassandra.storagedir=/opt/cassandra/data, -Dcassandra-foreground=yes]\r\nWARN 20:28:57 Unable to lock JVM memory (ENOMEM). This can result in part of the JVM being swapped out, especially with mmapped I/O enabled. Increase RLIMIT_MEMLOCK or run Cassandra as root.\r\nINFO 20:28:57 jemalloc seems to be preloaded from /usr/lib/x86_64-linux-gnu/libjemalloc.so.1\r\nINFO 20:28:57 JMX is enabled to receive remote connections on port: 7199\r\nWARN 20:28:57 OpenJDK is not recommended. Please upgrade to the newest Oracle Java release\r\nINFO 20:28:57 Initializing SIGAR library\r\nINFO 20:28:57 Checked OS settings and found them configured for optimal performance.\r\nWARN 20:28:57 Directory /opt/cassandra/data/commitlog doesn't exist\r\nWARN 20:28:57 Directory /opt/cassandra/data/saved_caches doesn't exist\r\nException (java.lang.ExceptionInInitializerError) encountered during startup: null\r\njava.lang.ExceptionInInitializerError\r\n\tat org.apache.cassandra.db.SystemKeyspace.checkHealth(SystemKeyspace.java:709)\r\n\tat org.apache.cassandra.service.StartupChecks$9.execute(StartupChecks.java:351)\r\n\tat org.apache.cassandra.service.StartupChecks.verify(StartupChecks.java:109)\r\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:188)\r\n\tat org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:607)\r\n\tat org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:717)\r\nCaused by: java.lang.IllegalArgumentException: Bad configuration; unable to start server: At least one DataFileDirectory must be specified\r\n\tat org.apache.cassandra.config.DatabaseDescriptor.createAllDirectories(DatabaseDescriptor.java:846)\r\n\tat org.apache.cassandra.db.Keyspace.(Keyspace.java:66)\r\n\t... 6 more\r\nERROR 20:28:58 Exception encountered during startup\r\njava.lang.ExceptionInInitializerError: null\r\n\tat org.apache.cassandra.db.SystemKeyspace.checkHealth(SystemKeyspace.java:709) ~[apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.service.StartupChecks$9.execute(StartupChecks.java:351) ~[apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.service.StartupChecks.verify(StartupChecks.java:109) ~[apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:188) [apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:607) [apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:717) [apache-cassandra-2.2.17.jar:2.2.17]\r\nCaused by: java.lang.IllegalArgumentException: Bad configuration; unable to start server: At least one DataFileDirectory must be specified\r\n\tat org.apache.cassandra.config.DatabaseDescriptor.createAllDirectories(DatabaseDescriptor.java:846) ~[apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.db.Keyspace.(Keyspace.java:66) ~[apache-cassandra-2.2.17.jar:2.2.17]\r\n\t... 6 common frames omitted\r\n\tat org.apache.cassandra.service.StartupChecks.verify(StartupChecks.java:109)\r\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:188)\r\n\tat org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:607)\r\n\tat org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:717)\r\nCaused by: java.lang.IllegalArgumentException: Bad configuration; unable to start server: At least one DataFileDirectory must be specified\r\n\tat org.apache.cassandra.config.DatabaseDescriptor.createAllDirectories(DatabaseDescriptor.java:846)\r\n\tat org.apache.cassandra.db.Keyspace.(Keyspace.java:66)\r\n\t... 6 more\r\n{noformat}\r\nI've traced this down to what I believe is the issue in [https://github.com/apache/cassandra/commit/257fb0377343cbfdb58327da17f31d4eaed940f5], specifically [https://github.com/apache/cassandra/commit/257fb0377343cbfdb58327da17f31d4eaed940f5#r40944000] – the addition of an empty value for {{data_file_directories}} needs to be accompanied with a change to {{DatabaseDescriptor.java}} to handle an empty array the same way as the previous nil value was (and seed the value of {{cassandra.storagedir}} into that empty array), as was done in [https://github.com/apache/cassandra/commit/b09e60f72bb2f37235d9e9190c25db36371b3c18#diff-b66584c9ce7b64019b5db5a531deeda1] (which I believe is the origin of this change).","from":"reporter","subject":"2.2.17 fails to start up with ExceptionInInitializerError"},{"body":"Would be very helpful to have this addressed as it is currently blocking new Docker containers of the Cassandra 2.2 tag - https://github.com/docker-library/cassandra/issues/215","from":"developer"},{"body":"Applied CASSANDRA-7066 fix in C* 3.0/3.11 to {{DatabaseDescriptor.java}} broken by commit {{257fb03}}:\r\n\r\n* 2.2 PR - https://github.com/apache/cassandra/pull/757\r\n* 2.2 Patch - [^CASSANDRA-16008-2.2.txt] \r\n\r\nThis only applies to the 2.2 branch since it's already fixed in C* 3.x/4.0.","from":"developer"},{"body":"Committed, thanks.","from":"developer"}],"created":"2020-08-03T22:06:08.000+0000","description":"After the upgrade to 2.2.17, Cassandra fails to start with the following error:\r\n{noformat}\r\nINFO 20:28:57 JVM Arguments: [-Dcom.sun.management.jmxremote.port=7199, -Dcom.sun.management.jmxremote.ssl=false, -Dcom.sun.management.jmxremote.authenticate=false, -ea, -javaagent:/opt/cassandra/lib/jamm-0.3.0.jar, -XX:+CMSClassUnloadingEnabled, -XX:+UseThreadPriorities, -XX:ThreadPriorityPolicy=42, -Xms128m, -Xmx128m, -Xmn32m, -XX:+HeapDumpOnOutOfMemoryError, -Xss256k, -XX:StringTableSize=1000003, -XX:+UseParNewGC, -XX:+UseConcMarkSweepGC, -XX:+CMSParallelRemarkEnabled, -XX:SurvivorRatio=8, -XX:MaxTenuringThreshold=1, -XX:CMSInitiatingOccupancyFraction=75, -XX:+UseCMSInitiatingOccupancyOnly, -XX:+UseTLAB, -XX:+PerfDisableSharedMem, -XX:CompileCommandFile=/etc/cassandra/hotspot_compiler, -XX:CMSWaitDuration=10000, -XX:+CMSParallelInitialMarkEnabled, -XX:+CMSEdenChunksRecordAlways, -XX:CMSWaitDuration=10000, -XX:+PrintGCDetails, -XX:+PrintGCDateStamps, -XX:+PrintHeapAtGC, -XX:+PrintTenuringDistribution, -XX:+PrintGCApplicationStoppedTime, -XX:+PrintPromotionFailure, -Xloggc:/opt/cassandra/logs/gc.log, -XX:+UseGCLogFileRotation, -XX:NumberOfGCLogFiles=10, -XX:GCLogFileSize=10M, -Djava.net.preferIPv4Stack=true, -Dcassandra.jmx.local.port=7199, -XX:+DisableExplicitGC, -Djava.library.path=/opt/cassandra/lib/sigar-bin, -Dcassandra.libjemalloc=/usr/lib/x86_64-linux-gnu/libjemalloc.so.1, -XX:OnOutOfMemoryError=kill -9 %p, -Dlogback.configurationFile=logback.xml, -Dcassandra.logdir=/opt/cassandra/logs, -Dcassandra.storagedir=/opt/cassandra/data, -Dcassandra-foreground=yes]\r\nWARN 20:28:57 Unable to lock JVM memory (ENOMEM). This can result in part of the JVM being swapped out, especially with mmapped I/O enabled. Increase RLIMIT_MEMLOCK or run Cassandra as root.\r\nINFO 20:28:57 jemalloc seems to be preloaded from /usr/lib/x86_64-linux-gnu/libjemalloc.so.1\r\nINFO 20:28:57 JMX is enabled to receive remote connections on port: 7199\r\nWARN 20:28:57 OpenJDK is not recommended. Please upgrade to the newest Oracle Java release\r\nINFO 20:28:57 Initializing SIGAR library\r\nINFO 20:28:57 Checked OS settings and found them configured for optimal performance.\r\nWARN 20:28:57 Directory /opt/cassandra/data/commitlog doesn't exist\r\nWARN 20:28:57 Directory /opt/cassandra/data/saved_caches doesn't exist\r\nException (java.lang.ExceptionInInitializerError) encountered during startup: null\r\njava.lang.ExceptionInInitializerError\r\n\tat org.apache.cassandra.db.SystemKeyspace.checkHealth(SystemKeyspace.java:709)\r\n\tat org.apache.cassandra.service.StartupChecks$9.execute(StartupChecks.java:351)\r\n\tat org.apache.cassandra.service.StartupChecks.verify(StartupChecks.java:109)\r\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:188)\r\n\tat org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:607)\r\n\tat org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:717)\r\nCaused by: java.lang.IllegalArgumentException: Bad configuration; unable to start server: At least one DataFileDirectory must be specified\r\n\tat org.apache.cassandra.config.DatabaseDescriptor.createAllDirectories(DatabaseDescriptor.java:846)\r\n\tat org.apache.cassandra.db.Keyspace.(Keyspace.java:66)\r\n\t... 6 more\r\nERROR 20:28:58 Exception encountered during startup\r\njava.lang.ExceptionInInitializerError: null\r\n\tat org.apache.cassandra.db.SystemKeyspace.checkHealth(SystemKeyspace.java:709) ~[apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.service.StartupChecks$9.execute(StartupChecks.java:351) ~[apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.service.StartupChecks.verify(StartupChecks.java:109) ~[apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:188) [apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:607) [apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:717) [apache-cassandra-2.2.17.jar:2.2.17]\r\nCaused by: java.lang.IllegalArgumentException: Bad configuration; unable to start server: At least one DataFileDirectory must be specified\r\n\tat org.apache.cassandra.config.DatabaseDescriptor.createAllDirectories(DatabaseDescriptor.java:846) ~[apache-cassandra-2.2.17.jar:2.2.17]\r\n\tat org.apache.cassandra.db.Keyspace.(Keyspace.java:66) ~[apache-cassandra-2.2.17.jar:2.2.17]\r\n\t... 6 common frames omitted\r\n\tat org.apache.cassandra.service.StartupChecks.verify(StartupChecks.java:109)\r\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:188)\r\n\tat org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:607)\r\n\tat org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:717)\r\nCaused by: java.lang.IllegalArgumentException: Bad configuration; unable to start server: At least one DataFileDirectory must be specified\r\n\tat org.apache.cassandra.config.DatabaseDescriptor.createAllDirectories(DatabaseDescriptor.java:846)\r\n\tat org.apache.cassandra.db.Keyspace.(Keyspace.java:66)\r\n\t... 6 more\r\n{noformat}\r\nI've traced this down to what I believe is the issue in [https://github.com/apache/cassandra/commit/257fb0377343cbfdb58327da17f31d4eaed940f5], specifically [https://github.com/apache/cassandra/commit/257fb0377343cbfdb58327da17f31d4eaed940f5#r40944000] – the addition of an empty value for {{data_file_directories}} needs to be accompanied with a change to {{DatabaseDescriptor.java}} to handle an empty array the same way as the previous nil value was (and seed the value of {{cassandra.storagedir}} into that empty array), as was done in [https://github.com/apache/cassandra/commit/b09e60f72bb2f37235d9e9190c25db36371b3c18#diff-b66584c9ce7b64019b5db5a531deeda1] (which I believe is the origin of this change).","issue_id":"13320752","key":"CASSANDRA-16008","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2020-09-23T15:08:57.000+0000","role":"fixed_distractor","summary":"2.2.17 fails to start up with ExceptionInInitializerError"} {"case_id":"13325356","cluster":"DISTRACTOR-CASSANDRA-16085","comments":[{"body":"Add the ability to catch UnsatisfiedLinkError in WindowsTimer and if not available to allow startup to continue. I get the following on Windows 2019 Server running Cassandra 3.11.3 on Azul Zulu Java 1.8u302:\r\n{code:java}\r\nERROR [main] 2021-09-10 14:25:29,064 CassandraDaemon.java:708 - Exception encountered during startup\r\njava.lang.UnsatisfiedLinkError: C:\\Windows\\Temp\\jna-1617058420\\jna439724638011461000.dll: Can't find dependent libraries\r\n at java.lang.ClassLoader$NativeLibrary.load(Native Method) ~[na:1.8.0_302]\r\n at java.lang.ClassLoader.loadLibrary0(ClassLoader.java:1950) ~[na:1.8.0_302]\r\n at java.lang.ClassLoader.loadLibrary(ClassLoader.java:1832) ~[na:1.8.0_302]\r\n at java.lang.Runtime.load0(Runtime.java:811) ~[na:1.8.0_302]\r\n at java.lang.System.load(System.java:1088) ~[na:1.8.0_302]\r\n at com.sun.jna.Native.loadNativeDispatchLibraryFromClasspath(Native.java:851) ~[jna-4.2.2.jar:4.2.2 (b0)]\r\n at com.sun.jna.Native.loadNativeDispatchLibrary(Native.java:826) ~[jna-4.2.2.jar:4.2.2 (b0)]\r\n at com.sun.jna.Native.(Native.java:140) ~[jna-4.2.2.jar:4.2.2 (b0)]\r\n at org.apache.cassandra.utils.WindowsTimer.(WindowsTimer.java:35) ~[apache-cassandra-3.11.3.jar:3.11.3]\r\n at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:599) [apache-cassandra-3.11.3.jar:3.11.3]\r\n at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:691) [apache-cassandra-3.11.3.jar:3.11.3]\r\n{code}\r\nThe fix implemented is similar to fix done in CASSANDRA-13333. In that issue  WindowsTimer was mentioned but further down the question of  moving WindowsTimer is asked so it could be fixed and answered not to move.\r\n \r\nI have attached the patch files for branches cassandra-3.0, cassandra-3.11, and trunk.\r\nBelow are the links to the 3 github branches (since I am new at this; I hope at least 1 set of patches or branches will work).\r\n[CASSANDRA-16085-3.0|https://github.com/sldr/cassandra/tree/CASSANDRA-16085-3.0]\r\n[CASSANDRA-16085-3.11|https://github.com/sldr/cassandra/tree/CASSANDRA-16085-3.11]\r\n[CASSANDRA-16085-trunk|https://github.com/sldr/cassandra/tree/CASSANDRA-16085-trunk]\r\n \r\nLet me know if you need anything else.\r\nThanks!","created":"2021-09-15T04:00:23.805+0000"},{"body":"Thanks for the patches! Unfortunately, since [Windows support has ended|https://issues.apache.org/jira/browse/CASSANDRA-16171] with 4.0, this is a bit of a conundrum, since I don't know any committers with a windows machine to test this. That said, since the patch is very straightforward and only touches WindowsTimer, I think we could commit this (except for the trunk patch - CASSANDRA-16956) [~jmckenzie] what do you think?","created":"2021-09-15T13:57:00.523+0000"},{"body":"Can't test it, but visually +1 to the patch. The default coalescing value performs reasonably well; think it was the memory mapping that really made the difference. Yeah, checking [here|https://web.archive.org/web/20160316234320/http://www.datastax.com/dev/blog/windows-and-cassandra-2-2-objects-in-mirror-are-closer-than-they-appear] it was between 9 and 14% from kernel timer 6 years ago, so I think we're 100% safe to continue if it fails to register.","created":"2021-09-15T17:40:10.350+0000"},{"body":"I think [~blerer] might be able to test it :) ","created":"2021-09-15T18:31:08.153+0000"},{"body":"I had already started the commit, so I guess now testing will be easier. Committed to 3.0 and 3.11 only.","created":"2021-09-15T18:44:53.410+0000"},{"body":"THANK YOU!!!\r\n\r\nWe will be looking forward to updating to 3.11.12.","created":"2021-09-15T20:00:29.069+0000"},{"body":"*Workaround:*\r\n\r\nCassandra is using jna-4.2.2.2.jar that has a native dll in it (jnidispatch.dll) that needs msvcr100.dll that is part of Microsoft Visual C++ 2010 Redistributable Package (x64). If you install that then Cassandra will run without crashing. NOTE: It may be hard to find \"Microsoft Visual C++ 2010 Redistributable Package (x64)\" from Microsoft.\r\n\r\nHere is a URL where I found it (url from Installshield prerequisite file for \"Microsoft Visual C++ 2010 Redistributable Package (x64)\"):\r\n\r\nhttp://download.microsoft.com/download/3/2/2/3224B87F-CFA0-4E70-BDA3-3DE650EFEBA5/vcredist_x64.exe\r\n\r\n*Better Option:*\r\n\r\nThe other option is getting Cassandra 3.11.12 released but I am not sure if and when that is going to happen.","created":"2022-01-19T19:54:57.555+0000"}],"conversations":[{"body":"When using JRE 8u261 to start cassandra it fails and gives this error: \\njava.lang.UnsatisfiedLinkError: WindowsTimer ","from":"reporter","subject":"Cassandra fails to start using Java SE 8 Update 261"},{"body":"Add the ability to catch UnsatisfiedLinkError in WindowsTimer and if not available to allow startup to continue. I get the following on Windows 2019 Server running Cassandra 3.11.3 on Azul Zulu Java 1.8u302:\r\n{code:java}\r\nERROR [main] 2021-09-10 14:25:29,064 CassandraDaemon.java:708 - Exception encountered during startup\r\njava.lang.UnsatisfiedLinkError: C:\\Windows\\Temp\\jna-1617058420\\jna439724638011461000.dll: Can't find dependent libraries\r\n at java.lang.ClassLoader$NativeLibrary.load(Native Method) ~[na:1.8.0_302]\r\n at java.lang.ClassLoader.loadLibrary0(ClassLoader.java:1950) ~[na:1.8.0_302]\r\n at java.lang.ClassLoader.loadLibrary(ClassLoader.java:1832) ~[na:1.8.0_302]\r\n at java.lang.Runtime.load0(Runtime.java:811) ~[na:1.8.0_302]\r\n at java.lang.System.load(System.java:1088) ~[na:1.8.0_302]\r\n at com.sun.jna.Native.loadNativeDispatchLibraryFromClasspath(Native.java:851) ~[jna-4.2.2.jar:4.2.2 (b0)]\r\n at com.sun.jna.Native.loadNativeDispatchLibrary(Native.java:826) ~[jna-4.2.2.jar:4.2.2 (b0)]\r\n at com.sun.jna.Native.(Native.java:140) ~[jna-4.2.2.jar:4.2.2 (b0)]\r\n at org.apache.cassandra.utils.WindowsTimer.(WindowsTimer.java:35) ~[apache-cassandra-3.11.3.jar:3.11.3]\r\n at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:599) [apache-cassandra-3.11.3.jar:3.11.3]\r\n at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:691) [apache-cassandra-3.11.3.jar:3.11.3]\r\n{code}\r\nThe fix implemented is similar to fix done in CASSANDRA-13333. In that issue  WindowsTimer was mentioned but further down the question of  moving WindowsTimer is asked so it could be fixed and answered not to move.\r\n \r\nI have attached the patch files for branches cassandra-3.0, cassandra-3.11, and trunk.\r\nBelow are the links to the 3 github branches (since I am new at this; I hope at least 1 set of patches or branches will work).\r\n[CASSANDRA-16085-3.0|https://github.com/sldr/cassandra/tree/CASSANDRA-16085-3.0]\r\n[CASSANDRA-16085-3.11|https://github.com/sldr/cassandra/tree/CASSANDRA-16085-3.11]\r\n[CASSANDRA-16085-trunk|https://github.com/sldr/cassandra/tree/CASSANDRA-16085-trunk]\r\n \r\nLet me know if you need anything else.\r\nThanks!","from":"developer"},{"body":"Thanks for the patches! Unfortunately, since [Windows support has ended|https://issues.apache.org/jira/browse/CASSANDRA-16171] with 4.0, this is a bit of a conundrum, since I don't know any committers with a windows machine to test this. That said, since the patch is very straightforward and only touches WindowsTimer, I think we could commit this (except for the trunk patch - CASSANDRA-16956) [~jmckenzie] what do you think?","from":"developer"},{"body":"Can't test it, but visually +1 to the patch. The default coalescing value performs reasonably well; think it was the memory mapping that really made the difference. Yeah, checking [here|https://web.archive.org/web/20160316234320/http://www.datastax.com/dev/blog/windows-and-cassandra-2-2-objects-in-mirror-are-closer-than-they-appear] it was between 9 and 14% from kernel timer 6 years ago, so I think we're 100% safe to continue if it fails to register.","from":"developer"},{"body":"I think [~blerer] might be able to test it :) ","from":"developer"},{"body":"I had already started the commit, so I guess now testing will be easier. Committed to 3.0 and 3.11 only.","from":"developer"},{"body":"THANK YOU!!!\r\n\r\nWe will be looking forward to updating to 3.11.12.","from":"developer"},{"body":"*Workaround:*\r\n\r\nCassandra is using jna-4.2.2.2.jar that has a native dll in it (jnidispatch.dll) that needs msvcr100.dll that is part of Microsoft Visual C++ 2010 Redistributable Package (x64). If you install that then Cassandra will run without crashing. NOTE: It may be hard to find \"Microsoft Visual C++ 2010 Redistributable Package (x64)\" from Microsoft.\r\n\r\nHere is a URL where I found it (url from Installshield prerequisite file for \"Microsoft Visual C++ 2010 Redistributable Package (x64)\"):\r\n\r\nhttp://download.microsoft.com/download/3/2/2/3224B87F-CFA0-4E70-BDA3-3DE650EFEBA5/vcredist_x64.exe\r\n\r\n*Better Option:*\r\n\r\nThe other option is getting Cassandra 3.11.12 released but I am not sure if and when that is going to happen.","from":"developer"}],"created":"2020-08-31T15:46:20.000+0000","description":"When using JRE 8u261 to start cassandra it fails and gives this error: \\njava.lang.UnsatisfiedLinkError: WindowsTimer ","issue_id":"13325356","key":"CASSANDRA-16085","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2021-09-15T18:44:53.000+0000","role":"fixed_distractor","summary":"Cassandra fails to start using Java SE 8 Update 261"} {"case_id":"13343356","cluster":"DISTRACTOR-CASSANDRA-16307","comments":[{"body":"Any idea if this impacts 3.x as well?","created":"2020-11-30T23:58:52.309+0000"},{"body":"[~dcapwell] Yes, it affects 3.11 and trunk, but not 3.0 because it doesn't have group-by queries.","created":"2020-12-01T00:39:42.558+0000"},{"body":"[~adelapena], [~maedhroz] I let my backlog grow to big. I will unassign myself from this ticket. I anybody has some available cycles he can feel free to pick it up.","created":"2021-01-11T08:06:24.531+0000"},{"body":"Good news: this is just an artifact of faulty pager in in-jvm dtests.\r\n\r\n|[trunk patch|https://github.com/apache/cassandra/pull/884]|[CI|https://app.circleci.com/pipelines/github/ifesdjeen/cassandra?branch=16307-trunk]|","created":"2021-02-02T16:25:44.762+0000"},{"body":"ci-cassandra runs for [jvm-dtest|https://ci-cassandra.apache.org/job/Cassandra-devbranch-jvm-dtest/207/] and [jvm-dtest-upgrade|https://ci-cassandra.apache.org/job/Cassandra-devbranch-jvm-dtest-upgrade/205/]. (No idea if the latter was needed.)","created":"2021-02-02T20:10:33.716+0000"},{"body":"That's a nice finding! However I think that it only solves some of the problems with paging and {{GROUP BY}}. The test above was an attempt to simplify [the original tests|https://github.com/apache/cassandra/blob/16075ed5c7429b3dd30d1e894b20e3598b8b9a10/test/distributed/org/apache/cassandra/distributed/test/ShortReadProtectionTest.java#L301-L347] in CASSANDRA-16180 where we found the problem. I'm afraid that a test more faithful to the original still fails:\r\n{code:java}\r\ntry (Cluster cluster = init(Cluster.create(2)))\r\n{\r\n cluster.schemaChange(withKeyspace(\"CREATE TABLE %s.t (pk int, ck int, PRIMARY KEY (pk, ck))\"));\r\n\r\n cluster.get(1).executeInternal(withKeyspace(\"INSERT INTO %s.t (pk, ck) VALUES (1, 1) USING TIMESTAMP 0\"));\r\n cluster.get(1).executeInternal(withKeyspace(\"DELETE FROM %s.t WHERE pk=0 AND ck=0\"));\r\n cluster.get(1).executeInternal(withKeyspace(\"INSERT INTO %s.t (pk, ck) VALUES (2, 2) USING TIMESTAMP 0\"));\r\n\r\n cluster.get(2).executeInternal(withKeyspace(\"DELETE FROM %s.t WHERE pk=1 AND ck=1\"));\r\n cluster.get(2).executeInternal(withKeyspace(\"INSERT INTO %s.t (pk, ck) VALUES (0, 0) USING TIMESTAMP 0\"));\r\n cluster.get(2).executeInternal(withKeyspace(\"DELETE FROM %s.t WHERE pk=2 AND ck=2\"));\r\n\r\n String query = withKeyspace(\"SELECT * FROM %s.t GROUP BY pk\");\r\n Iterator rows = cluster.coordinator(1).executeWithPaging(query, ConsistencyLevel.ALL, 1);\r\n assertRows(Iterators.toArray(rows, Object[].class));\r\n}\r\n{code}\r\nApparently this can also be reproduced with a Python dtest, suggesting that the problem is not only caused by in-jvm dtests:\r\n{code:python}\r\ncluster = self.cluster\r\ncluster.populate(2).start()\r\nnode1, node2 = cluster.nodelist()\r\n\r\nsession = self.patient_exclusive_cql_connection(node1, consistency_level=ConsistencyLevel.ONE)\r\ncreate_ks(session, 'ks', 2)\r\nsession.execute('CREATE TABLE ks.t (pk int, ck int, PRIMARY KEY (pk, ck))')\r\n\r\nnode2.stop()\r\nsession.execute('INSERT INTO ks.t (pk, ck) VALUES (1, 1) USING TIMESTAMP 0')\r\nsession.execute('DELETE FROM ks.t WHERE pk=0 AND ck=0')\r\nsession.execute('INSERT INTO ks.t (pk, ck) VALUES (2, 2) USING TIMESTAMP 0')\r\nnode2.start()\r\n\r\nnode1.stop()\r\nsession = self.patient_exclusive_cql_connection(node2, consistency_level=ConsistencyLevel.ONE)\r\nsession.execute('DELETE FROM ks.t WHERE pk=1 AND ck=1')\r\nsession.execute('INSERT INTO ks.t (pk, ck) VALUES (0, 0) USING TIMESTAMP 0')\r\nsession.execute('DELETE FROM ks.t WHERE pk=2 AND ck=2')\r\nnode1.start()\r\n\r\nsession = self.patient_exclusive_cql_connection(node1, consistency_level=ConsistencyLevel.ALL)\r\nsession.default_fetch_size = 1\r\nassert_none(session, 'SELECT * FROM ks.t GROUP BY pk')\r\n{code}","created":"2021-02-02T20:42:42.188+0000"},{"body":"[~adelapena] thanks for elaborating. This looks like a separate problem. I'll open a separate jira for one of the bugs. I'm sure that testing/fixing the new problem will be much easier now when paging is working. ","created":"2021-02-02T21:54:58.357+0000"},{"body":"Just started to take a look at the trunk patch, but before I get too far, I see a failure for [CompactStorageUpgradeTest#compactStoragePagingTest |https://app.circleci.com/pipelines/github/ifesdjeen/cassandra/148/workflows/c586d0ab-7fa0-4c13-a6d1-c927bb55342a/jobs/3153] that looks new (and might be related, given it touches paging).","created":"2021-02-03T20:05:45.061+0000"},{"body":"I've just confirmed that {{CompactStorageUpgradeTest#compactStoragePagingTest}} consistently fails with the patch, and it works by reverting it.","created":"2021-02-04T18:13:59.945+0000"},{"body":"[~maedhroz][~adelapena] I've fixed the test right away, just didn't expect the patch is going to get attention right away, and was focusing on fixing the real group by/paging/srp bug. The test failure in compact storage upgrade test is just a conseqeunce of the test not complying to {{Iterator}} interface and not calling {{hasNext}} before calling {{next}}. \r\n\r\nThat said, I'm pretty close to the fix of the actual issue, just running some more Harry tests to verify the fix. It is caused by the fact that for group by, we're increasing data limit counters only after the next row was seen, which is why SRP thinks there's no need to try to fetch more contents.\r\n\r\nUPD: I'll move the in-jvm dtest paging problem to a [CASSANDRA-16427], and will put [~maedhroz] and [~adelapena] as reviewers there. Thank you!","created":"2021-02-04T18:48:47.710+0000"},{"body":"The problem here turned out to be that the Group By pager incorrectly interacts with SRP. \r\n\r\nIn short, we have multiple implementations of paging counters: one is regular CQL one and, the other is Group By. CQL one is counting rows as it sees them. Group By is counting rows its going to return (in other words, groups) only when it encounters the next group (or partition/data end).\r\n\r\nThe problem is that when issuing SRP, we rely on correctness of these counters. And if counter says that iterator couldn’t give enough data to exhaust limits, we’re not going to ask for more data. In case of regular counters it all works as expected, but with group by counters, we’ll say that we have seen 1 row less every time, and won’t issue SRP. In other words, if one of the nodes holds a row followed by the tombstone, current behaviour will prevent SRP from getting triggered and tombstone from being visible, since group will be unfinished, and it doesn’t make sense to ask the node for more data since it couldn’t even satisfy the current data limit.\r\n\r\nA straightforward solution to this would have been to just start counting groups as we’re encountering them. However, here the problem is that if we say that we have a group as soon as we see the first row that belongs to it, internal pager will think that there’s no need to continue, since data limits are satisfied.\r\n\r\nThe patch is introducing a distinction between {{needsMoreContents}} (better name pending) and {{isDone}}. This allows us to know we need to fetch more rows in case we haven’t seen either end of data, or next row. I think we can even start counting groups right away instead of delaying counting till we see the next group to (potentially) simplify the logic, however I wanted to leave a larger refactoring to a separate patch.\r\n\r\n|[trunk|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:16307-trunk]|[CI|https://app.circleci.com/pipelines/github/ifesdjeen/cassandra?branch=16307-trunk]|\r\n|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...ifesdjeen:16307-3.11]|[CI|https://app.circleci.com/pipelines/github/ifesdjeen/cassandra?branch=16307-3.11]|","created":"2021-02-08T08:34:16.814+0000"},{"body":"Looks good to me, assuming that you're planning to include the logic inversion patch on the 3.11 branch too. One minor nit is the naming of the new methods. They feel slightly off because counters/limits don't have content, so maybe a better options could be:\r\n\r\n{{maybeUnderCounted / maybeUnderCountedForPartition}} or {{maybeShort / maybeShortForPartition}} ?","created":"2021-02-08T12:35:08.489+0000"},{"body":"[~samt] I've tried out a slightly different approach, suggested by [~blerer], which is to use {{command.isExhausted}}. Since it has a logic that checks whether or not the counter is exhausted (in other words, we've reached _either_ row limit, or group limit), it works for this case. Works for rows, since we count them as we see, and works for groups since we'll only count the group as soon as iterator is closed or we encounter the next group. One of the advantages of using this approach is that we will avoid an extra page request for cases when group limit coincides with row limit. Would you be able to take another look?","created":"2021-02-11T13:18:12.113+0000"},{"body":"The patches look good to me. ","created":"2021-02-12T08:24:48.714+0000"},{"body":"+1 LGTM too.","created":"2021-02-12T15:07:12.113+0000"},{"body":"Thank you for the review! \r\n\r\nCommitted to 3.11 with [f258ae67516d53752c8d1f0a2576d72471ed427f|https://github.com/apache/cassandra/commit/f258ae67516d53752c8d1f0a2576d72471ed427f] and merged up to [trunk|https://github.com/apache/cassandra/commit/7cddbd40ce6b326df533fd6d3c4131ef70b3b068].","created":"2021-02-15T11:56:33.234+0000"}],"conversations":[{"body":"{{GROUP BY}} queries using paging and CL>ONE/LOCAL_ONE. This dtest reproduces the problem:\r\n{code:java}\r\ntry (Cluster cluster = init(Cluster.create(2)))\r\n{\r\n cluster.schemaChange(withKeyspace(\"CREATE TABLE %s.t (pk int, ck int, PRIMARY KEY (pk, ck))\"));\r\n ICoordinator coordinator = cluster.coordinator(1);\r\n coordinator.execute(withKeyspace(\"INSERT INTO %s.t (pk, ck) VALUES (0, 0)\"), ConsistencyLevel.ALL);\r\n coordinator.execute(withKeyspace(\"INSERT INTO %s.t (pk, ck) VALUES (1, 1)\"), ConsistencyLevel.ALL);\r\n \r\n cluster.get(1).executeInternal(withKeyspace(\"DELETE FROM %s.t WHERE pk=0 AND ck=0\"));\r\n cluster.get(2).executeInternal(withKeyspace(\"DELETE FROM %s.t WHERE pk=1 AND ck=1\"));\r\n String query = withKeyspace(\"SELECT * FROM %s.t GROUP BY pk\");\r\n Iterator rows = coordinator.executeWithPaging(query, ConsistencyLevel.ALL, 1);\r\n assertRows(Iterators.toArray(rows, Object[].class));\r\n}\r\n{code}\r\nUsing a 2-node cluster and RF=2, the test inserts two partitions in both nodes. Then it locally deletes each row in a separate node, so each node sees a different partition alive, but reconciliation should produce no alive partitions. However, a {{GROUP BY}} query using a page size of 1 wrongly returns one of the rows.\r\n\r\nThis has been detected during CASSANDRA-16180, and it is probably related to CASSANDRA-15459, which solved a similar problem for group-by queries with limit, instead of paging.","from":"reporter","subject":"GROUP BY queries with paging can return deleted data"},{"body":"Any idea if this impacts 3.x as well?","from":"developer"},{"body":"[~dcapwell] Yes, it affects 3.11 and trunk, but not 3.0 because it doesn't have group-by queries.","from":"developer"},{"body":"[~adelapena], [~maedhroz] I let my backlog grow to big. I will unassign myself from this ticket. I anybody has some available cycles he can feel free to pick it up.","from":"developer"},{"body":"Good news: this is just an artifact of faulty pager in in-jvm dtests.\r\n\r\n|[trunk patch|https://github.com/apache/cassandra/pull/884]|[CI|https://app.circleci.com/pipelines/github/ifesdjeen/cassandra?branch=16307-trunk]|","from":"developer"},{"body":"ci-cassandra runs for [jvm-dtest|https://ci-cassandra.apache.org/job/Cassandra-devbranch-jvm-dtest/207/] and [jvm-dtest-upgrade|https://ci-cassandra.apache.org/job/Cassandra-devbranch-jvm-dtest-upgrade/205/]. (No idea if the latter was needed.)","from":"developer"},{"body":"That's a nice finding! However I think that it only solves some of the problems with paging and {{GROUP BY}}. The test above was an attempt to simplify [the original tests|https://github.com/apache/cassandra/blob/16075ed5c7429b3dd30d1e894b20e3598b8b9a10/test/distributed/org/apache/cassandra/distributed/test/ShortReadProtectionTest.java#L301-L347] in CASSANDRA-16180 where we found the problem. I'm afraid that a test more faithful to the original still fails:\r\n{code:java}\r\ntry (Cluster cluster = init(Cluster.create(2)))\r\n{\r\n cluster.schemaChange(withKeyspace(\"CREATE TABLE %s.t (pk int, ck int, PRIMARY KEY (pk, ck))\"));\r\n\r\n cluster.get(1).executeInternal(withKeyspace(\"INSERT INTO %s.t (pk, ck) VALUES (1, 1) USING TIMESTAMP 0\"));\r\n cluster.get(1).executeInternal(withKeyspace(\"DELETE FROM %s.t WHERE pk=0 AND ck=0\"));\r\n cluster.get(1).executeInternal(withKeyspace(\"INSERT INTO %s.t (pk, ck) VALUES (2, 2) USING TIMESTAMP 0\"));\r\n\r\n cluster.get(2).executeInternal(withKeyspace(\"DELETE FROM %s.t WHERE pk=1 AND ck=1\"));\r\n cluster.get(2).executeInternal(withKeyspace(\"INSERT INTO %s.t (pk, ck) VALUES (0, 0) USING TIMESTAMP 0\"));\r\n cluster.get(2).executeInternal(withKeyspace(\"DELETE FROM %s.t WHERE pk=2 AND ck=2\"));\r\n\r\n String query = withKeyspace(\"SELECT * FROM %s.t GROUP BY pk\");\r\n Iterator rows = cluster.coordinator(1).executeWithPaging(query, ConsistencyLevel.ALL, 1);\r\n assertRows(Iterators.toArray(rows, Object[].class));\r\n}\r\n{code}\r\nApparently this can also be reproduced with a Python dtest, suggesting that the problem is not only caused by in-jvm dtests:\r\n{code:python}\r\ncluster = self.cluster\r\ncluster.populate(2).start()\r\nnode1, node2 = cluster.nodelist()\r\n\r\nsession = self.patient_exclusive_cql_connection(node1, consistency_level=ConsistencyLevel.ONE)\r\ncreate_ks(session, 'ks', 2)\r\nsession.execute('CREATE TABLE ks.t (pk int, ck int, PRIMARY KEY (pk, ck))')\r\n\r\nnode2.stop()\r\nsession.execute('INSERT INTO ks.t (pk, ck) VALUES (1, 1) USING TIMESTAMP 0')\r\nsession.execute('DELETE FROM ks.t WHERE pk=0 AND ck=0')\r\nsession.execute('INSERT INTO ks.t (pk, ck) VALUES (2, 2) USING TIMESTAMP 0')\r\nnode2.start()\r\n\r\nnode1.stop()\r\nsession = self.patient_exclusive_cql_connection(node2, consistency_level=ConsistencyLevel.ONE)\r\nsession.execute('DELETE FROM ks.t WHERE pk=1 AND ck=1')\r\nsession.execute('INSERT INTO ks.t (pk, ck) VALUES (0, 0) USING TIMESTAMP 0')\r\nsession.execute('DELETE FROM ks.t WHERE pk=2 AND ck=2')\r\nnode1.start()\r\n\r\nsession = self.patient_exclusive_cql_connection(node1, consistency_level=ConsistencyLevel.ALL)\r\nsession.default_fetch_size = 1\r\nassert_none(session, 'SELECT * FROM ks.t GROUP BY pk')\r\n{code}","from":"developer"},{"body":"[~adelapena] thanks for elaborating. This looks like a separate problem. I'll open a separate jira for one of the bugs. I'm sure that testing/fixing the new problem will be much easier now when paging is working. ","from":"developer"},{"body":"Just started to take a look at the trunk patch, but before I get too far, I see a failure for [CompactStorageUpgradeTest#compactStoragePagingTest |https://app.circleci.com/pipelines/github/ifesdjeen/cassandra/148/workflows/c586d0ab-7fa0-4c13-a6d1-c927bb55342a/jobs/3153] that looks new (and might be related, given it touches paging).","from":"developer"},{"body":"I've just confirmed that {{CompactStorageUpgradeTest#compactStoragePagingTest}} consistently fails with the patch, and it works by reverting it.","from":"developer"},{"body":"[~maedhroz][~adelapena] I've fixed the test right away, just didn't expect the patch is going to get attention right away, and was focusing on fixing the real group by/paging/srp bug. The test failure in compact storage upgrade test is just a conseqeunce of the test not complying to {{Iterator}} interface and not calling {{hasNext}} before calling {{next}}. \r\n\r\nThat said, I'm pretty close to the fix of the actual issue, just running some more Harry tests to verify the fix. It is caused by the fact that for group by, we're increasing data limit counters only after the next row was seen, which is why SRP thinks there's no need to try to fetch more contents.\r\n\r\nUPD: I'll move the in-jvm dtest paging problem to a [CASSANDRA-16427], and will put [~maedhroz] and [~adelapena] as reviewers there. Thank you!","from":"developer"},{"body":"The problem here turned out to be that the Group By pager incorrectly interacts with SRP. \r\n\r\nIn short, we have multiple implementations of paging counters: one is regular CQL one and, the other is Group By. CQL one is counting rows as it sees them. Group By is counting rows its going to return (in other words, groups) only when it encounters the next group (or partition/data end).\r\n\r\nThe problem is that when issuing SRP, we rely on correctness of these counters. And if counter says that iterator couldn’t give enough data to exhaust limits, we’re not going to ask for more data. In case of regular counters it all works as expected, but with group by counters, we’ll say that we have seen 1 row less every time, and won’t issue SRP. In other words, if one of the nodes holds a row followed by the tombstone, current behaviour will prevent SRP from getting triggered and tombstone from being visible, since group will be unfinished, and it doesn’t make sense to ask the node for more data since it couldn’t even satisfy the current data limit.\r\n\r\nA straightforward solution to this would have been to just start counting groups as we’re encountering them. However, here the problem is that if we say that we have a group as soon as we see the first row that belongs to it, internal pager will think that there’s no need to continue, since data limits are satisfied.\r\n\r\nThe patch is introducing a distinction between {{needsMoreContents}} (better name pending) and {{isDone}}. This allows us to know we need to fetch more rows in case we haven’t seen either end of data, or next row. I think we can even start counting groups right away instead of delaying counting till we see the next group to (potentially) simplify the logic, however I wanted to leave a larger refactoring to a separate patch.\r\n\r\n|[trunk|https://github.com/apache/cassandra/compare/trunk...ifesdjeen:16307-trunk]|[CI|https://app.circleci.com/pipelines/github/ifesdjeen/cassandra?branch=16307-trunk]|\r\n|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...ifesdjeen:16307-3.11]|[CI|https://app.circleci.com/pipelines/github/ifesdjeen/cassandra?branch=16307-3.11]|","from":"developer"},{"body":"Looks good to me, assuming that you're planning to include the logic inversion patch on the 3.11 branch too. One minor nit is the naming of the new methods. They feel slightly off because counters/limits don't have content, so maybe a better options could be:\r\n\r\n{{maybeUnderCounted / maybeUnderCountedForPartition}} or {{maybeShort / maybeShortForPartition}} ?","from":"developer"},{"body":"[~samt] I've tried out a slightly different approach, suggested by [~blerer], which is to use {{command.isExhausted}}. Since it has a logic that checks whether or not the counter is exhausted (in other words, we've reached _either_ row limit, or group limit), it works for this case. Works for rows, since we count them as we see, and works for groups since we'll only count the group as soon as iterator is closed or we encounter the next group. One of the advantages of using this approach is that we will avoid an extra page request for cases when group limit coincides with row limit. Would you be able to take another look?","from":"developer"},{"body":"The patches look good to me. ","from":"developer"},{"body":"+1 LGTM too.","from":"developer"},{"body":"Thank you for the review! \r\n\r\nCommitted to 3.11 with [f258ae67516d53752c8d1f0a2576d72471ed427f|https://github.com/apache/cassandra/commit/f258ae67516d53752c8d1f0a2576d72471ed427f] and merged up to [trunk|https://github.com/apache/cassandra/commit/7cddbd40ce6b326df533fd6d3c4131ef70b3b068].","from":"developer"}],"created":"2020-11-30T16:57:25.000+0000","description":"{{GROUP BY}} queries using paging and CL>ONE/LOCAL_ONE. This dtest reproduces the problem:\r\n{code:java}\r\ntry (Cluster cluster = init(Cluster.create(2)))\r\n{\r\n cluster.schemaChange(withKeyspace(\"CREATE TABLE %s.t (pk int, ck int, PRIMARY KEY (pk, ck))\"));\r\n ICoordinator coordinator = cluster.coordinator(1);\r\n coordinator.execute(withKeyspace(\"INSERT INTO %s.t (pk, ck) VALUES (0, 0)\"), ConsistencyLevel.ALL);\r\n coordinator.execute(withKeyspace(\"INSERT INTO %s.t (pk, ck) VALUES (1, 1)\"), ConsistencyLevel.ALL);\r\n \r\n cluster.get(1).executeInternal(withKeyspace(\"DELETE FROM %s.t WHERE pk=0 AND ck=0\"));\r\n cluster.get(2).executeInternal(withKeyspace(\"DELETE FROM %s.t WHERE pk=1 AND ck=1\"));\r\n String query = withKeyspace(\"SELECT * FROM %s.t GROUP BY pk\");\r\n Iterator rows = coordinator.executeWithPaging(query, ConsistencyLevel.ALL, 1);\r\n assertRows(Iterators.toArray(rows, Object[].class));\r\n}\r\n{code}\r\nUsing a 2-node cluster and RF=2, the test inserts two partitions in both nodes. Then it locally deletes each row in a separate node, so each node sees a different partition alive, but reconciliation should produce no alive partitions. However, a {{GROUP BY}} query using a page size of 1 wrongly returns one of the rows.\r\n\r\nThis has been detected during CASSANDRA-16180, and it is probably related to CASSANDRA-15459, which solved a similar problem for group-by queries with limit, instead of paging.","issue_id":"13343356","key":"CASSANDRA-16307","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2021-02-15T11:56:33.000+0000","role":"fixed_distractor","summary":"GROUP BY queries with paging can return deleted data"} {"case_id":"12478581","cluster":"DISTRACTOR-CASSANDRA-1674","comments":[{"body":"(this is rc1 snapshot, btw)","created":"2010-10-28T14:38:28.741+0000"},{"body":"bq. WARN [ScheduledTasks:1] 2010-10-28 08:31:14,305 MessagingService.java (line 515) Dropped 7 messages in the last 1000ms\n\nCreated CASSANDRA-1676 for this.","created":"2010-10-28T15:40:36.744+0000"},{"body":"I haven't tested the theory, but various messages are executed in the single threaded MISC stage.","created":"2010-10-28T15:47:41.398+0000"},{"body":"it doesn't have to be single threaded to drop messages, it just has to not run the verbhandler until RPC_TIMEOUT after the message was received.\n\nIn practice though I don't see any place where the verbhandler would be blocking or otherwise slow, including the ones running on MISC.","created":"2010-10-28T16:05:03.832+0000"},{"body":"so we have CASSANDRA-1685 and CASSANDRA-1676 committed now, but I don't think those are the root cause here since for them to cause lots of disk space usage repair would have to retry dropped requests which I don't see it doing. (is this correct?)","created":"2010-11-01T04:08:51.038+0000"},{"body":"Here is one way to reproduce something similar, at least.\n\nI start with 1 node, and put 1M rows on it. Then I add a 2nd node, then a third, then run cleanup. So I have 25% 50% 25% of the 1M rows on those 3 machines:\n\n{code}\n$ nodetool -h localhost -p 8080 ring\nAddress Status State Load Token \n 164074424718159380631425626216484638578 \n127.0.0.2 Up Normal 95.39 MB 36509681143663337709904175976843745575 \n127.0.0.1 Up Normal 190.72 MB 121466565577718822332569289829746132598 \n127.0.0.3 Up Normal 95.4 MB 164074424718159380631425626216484638578 \n{code}\n\nI increase the RF to 2, and run repair against .2:\n\n{code}\nAddress Status State Load Token \n 164074424718159380631425626216484638578 \n127.0.0.2 Up Normal 381.41 MB 36509681143663337709904175976843745575 \n127.0.0.1 Up Normal 381.41 MB 121466565577718822332569289829746132598 \n127.0.0.3 Up Normal 190.75 MB 164074424718159380631425626216484638578 \n{code}\n\nLet's call the ranges of data that .2, .1, and .3 originally had R, G, and B. Post repair, .2 should have RB but it has RGB. Node 1 should have GB but it also has RGB.\n\nCleanup puts things in the expected state:\n{code}\nAddress Status State Load Token \n 164074424718159380631425626216484638578 \n127.0.0.2 Up Normal 190.74 MB 36509681143663337709904175976843745575 \n127.0.0.1 Up Normal 285.61 MB 121466565577718822332569289829746132598 \n127.0.0.3 Up Normal 190.75 MB 164074424718159380631425626216484638578 \n{code}\n\nThis is reproducible against 0.6 as well as 0.7.","created":"2010-11-12T03:10:35.293+0000"},{"body":"After finding the differences for all data held on both nodes, repair was not limiting the repaired area to the ranges that the nodes had in common. jbellis: In your example above, because the nodes had no data in common to start with, the MerkleTree was determining that they were different for (0,0], but rather than intersecting the different range with the range of responsibility they had in common, it was repairing all of (0,0].\n\nAttaching a patch for 0.7 to only repair the intersecting ranges.","created":"2010-11-17T02:03:34.160+0000"},{"body":"I started rebasing this for 0.6, but noticed some odd behaviour in the 0.6 branch: in what cases would SS.getLocalToken not equal the RHS of SS.getLocalPrimaryRange?","created":"2010-11-17T02:04:49.643+0000"},{"body":"Surely something that can overflow disks, even when below 50% usage, is worthy of fix before 0.7.0 is released...","created":"2010-11-17T02:35:55.245+0000"},{"body":"Chip: agreed. While it was broken before 0.7, it's a small enough fix that we should try to get it in 0.7.0","created":"2010-11-17T03:10:24.932+0000"},{"body":"bq. in what cases would SS.getLocalToken not equal the RHS of SS.getLocalPrimaryRange? \n\nat a guess, some test is mocking up a TokenMetadata that isn't consistent w/ localtoken","created":"2010-11-17T08:56:52.133+0000"},{"body":"Patches for 0.6 and 0.7/trunk: ready for review.","created":"2010-11-17T19:30:56.994+0000"},{"body":"committed, but I think there is a bug. With a similar setup to the above (200K keys instead of 1M), the pre-RF change setup is\n\n{code}\nAddress Status State Load Token \n 106239986353888428655683112465158427815 \n127.0.0.2 Up Normal 37.97 MB 21212647344528771789748883276744400257 \n127.0.0.3 Up Normal 18.98 MB 63523312719601176253752035031089272162 \n127.0.0.1 Up Normal 19.05 MB 106239986353888428655683112465158427815 \n{code}\n\npost-repair is\n{code}\nAddress Status State Load Token \n 106239986353888428655683112465158427815 \n127.0.0.2 Up Normal 57.01 MB 21212647344528771789748883276744400257 \n127.0.0.3 Up Normal 56.94 MB 63523312719601176253752035031089272162 \n127.0.0.1 Up Normal 19.07 MB 106239986353888428655683112465158427815 \n{code}\n\nSo eyeballing it looks reasonable. But when I kill node 2 and run\n\n{code}$ python contrib/py_stress/stress.py -n 200000 -o read{code}\n\nI get a ton of key-not-found exceptions, indicating that not all the data on 2 got replicated to 3.","created":"2010-11-17T20:08:15.434+0000"},{"body":"the not-found was because i only repaired .2, but the RF increase meant requests to .1 thought they could find the .3 data locally. not a bug. Thanks to Stu for pointing this out on IRC.","created":"2010-11-18T01:06:25.958+0000"},{"body":"This still needs to be committed to trunk.","created":"2010-11-18T01:46:01.309+0000"}],"conversations":[{"body":"I'm watching a repair on a 7 node cluster. Repair was sent to one node; the node had 18G of data. No other node has more than 28G. The node where the repair initiated is now up to 261G with 53/60 AES tasks outstanding.\n\nI have seen repair take more space than expected on 0.6 but nothing this extreme.\n\nOther nodes in the cluster are occasionally logging\nWARN [ScheduledTasks:1] 2010-10-28 08:31:14,305 MessagingService.java (line 515) Dropped 7 messages in the last 1000ms\n\nThe cluster is quiesced except for the repair. Not sure if the dropped messages are contributing the the disk space (b/c of retries?).","from":"reporter","subject":"Repair using abnormally large amounts of disk space"},{"body":"(this is rc1 snapshot, btw)","from":"developer"},{"body":"bq. WARN [ScheduledTasks:1] 2010-10-28 08:31:14,305 MessagingService.java (line 515) Dropped 7 messages in the last 1000ms\n\nCreated CASSANDRA-1676 for this.","from":"developer"},{"body":"I haven't tested the theory, but various messages are executed in the single threaded MISC stage.","from":"developer"},{"body":"it doesn't have to be single threaded to drop messages, it just has to not run the verbhandler until RPC_TIMEOUT after the message was received.\n\nIn practice though I don't see any place where the verbhandler would be blocking or otherwise slow, including the ones running on MISC.","from":"developer"},{"body":"so we have CASSANDRA-1685 and CASSANDRA-1676 committed now, but I don't think those are the root cause here since for them to cause lots of disk space usage repair would have to retry dropped requests which I don't see it doing. (is this correct?)","from":"developer"},{"body":"Here is one way to reproduce something similar, at least.\n\nI start with 1 node, and put 1M rows on it. Then I add a 2nd node, then a third, then run cleanup. So I have 25% 50% 25% of the 1M rows on those 3 machines:\n\n{code}\n$ nodetool -h localhost -p 8080 ring\nAddress Status State Load Token \n 164074424718159380631425626216484638578 \n127.0.0.2 Up Normal 95.39 MB 36509681143663337709904175976843745575 \n127.0.0.1 Up Normal 190.72 MB 121466565577718822332569289829746132598 \n127.0.0.3 Up Normal 95.4 MB 164074424718159380631425626216484638578 \n{code}\n\nI increase the RF to 2, and run repair against .2:\n\n{code}\nAddress Status State Load Token \n 164074424718159380631425626216484638578 \n127.0.0.2 Up Normal 381.41 MB 36509681143663337709904175976843745575 \n127.0.0.1 Up Normal 381.41 MB 121466565577718822332569289829746132598 \n127.0.0.3 Up Normal 190.75 MB 164074424718159380631425626216484638578 \n{code}\n\nLet's call the ranges of data that .2, .1, and .3 originally had R, G, and B. Post repair, .2 should have RB but it has RGB. Node 1 should have GB but it also has RGB.\n\nCleanup puts things in the expected state:\n{code}\nAddress Status State Load Token \n 164074424718159380631425626216484638578 \n127.0.0.2 Up Normal 190.74 MB 36509681143663337709904175976843745575 \n127.0.0.1 Up Normal 285.61 MB 121466565577718822332569289829746132598 \n127.0.0.3 Up Normal 190.75 MB 164074424718159380631425626216484638578 \n{code}\n\nThis is reproducible against 0.6 as well as 0.7.","from":"developer"},{"body":"After finding the differences for all data held on both nodes, repair was not limiting the repaired area to the ranges that the nodes had in common. jbellis: In your example above, because the nodes had no data in common to start with, the MerkleTree was determining that they were different for (0,0], but rather than intersecting the different range with the range of responsibility they had in common, it was repairing all of (0,0].\n\nAttaching a patch for 0.7 to only repair the intersecting ranges.","from":"developer"},{"body":"I started rebasing this for 0.6, but noticed some odd behaviour in the 0.6 branch: in what cases would SS.getLocalToken not equal the RHS of SS.getLocalPrimaryRange?","from":"developer"},{"body":"Surely something that can overflow disks, even when below 50% usage, is worthy of fix before 0.7.0 is released...","from":"developer"},{"body":"Chip: agreed. While it was broken before 0.7, it's a small enough fix that we should try to get it in 0.7.0","from":"developer"},{"body":"bq. in what cases would SS.getLocalToken not equal the RHS of SS.getLocalPrimaryRange? \n\nat a guess, some test is mocking up a TokenMetadata that isn't consistent w/ localtoken","from":"developer"},{"body":"Patches for 0.6 and 0.7/trunk: ready for review.","from":"developer"},{"body":"committed, but I think there is a bug. With a similar setup to the above (200K keys instead of 1M), the pre-RF change setup is\n\n{code}\nAddress Status State Load Token \n 106239986353888428655683112465158427815 \n127.0.0.2 Up Normal 37.97 MB 21212647344528771789748883276744400257 \n127.0.0.3 Up Normal 18.98 MB 63523312719601176253752035031089272162 \n127.0.0.1 Up Normal 19.05 MB 106239986353888428655683112465158427815 \n{code}\n\npost-repair is\n{code}\nAddress Status State Load Token \n 106239986353888428655683112465158427815 \n127.0.0.2 Up Normal 57.01 MB 21212647344528771789748883276744400257 \n127.0.0.3 Up Normal 56.94 MB 63523312719601176253752035031089272162 \n127.0.0.1 Up Normal 19.07 MB 106239986353888428655683112465158427815 \n{code}\n\nSo eyeballing it looks reasonable. But when I kill node 2 and run\n\n{code}$ python contrib/py_stress/stress.py -n 200000 -o read{code}\n\nI get a ton of key-not-found exceptions, indicating that not all the data on 2 got replicated to 3.","from":"developer"},{"body":"the not-found was because i only repaired .2, but the RF increase meant requests to .1 thought they could find the .3 data locally. not a bug. Thanks to Stu for pointing this out on IRC.","from":"developer"},{"body":"This still needs to be committed to trunk.","from":"developer"}],"created":"2010-10-28T14:38:05.000+0000","description":"I'm watching a repair on a 7 node cluster. Repair was sent to one node; the node had 18G of data. No other node has more than 28G. The node where the repair initiated is now up to 261G with 53/60 AES tasks outstanding.\n\nI have seen repair take more space than expected on 0.6 but nothing this extreme.\n\nOther nodes in the cluster are occasionally logging\nWARN [ScheduledTasks:1] 2010-10-28 08:31:14,305 MessagingService.java (line 515) Dropped 7 messages in the last 1000ms\n\nThe cluster is quiesced except for the repair. Not sure if the dropped messages are contributing the the disk space (b/c of retries?).","issue_id":"12478581","key":"CASSANDRA-1674","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2010-11-18T01:06:26.000+0000","role":"fixed_distractor","summary":"Repair using abnormally large amounts of disk space"} {"case_id":"13390199","cluster":"DISTRACTOR-CASSANDRA-16808","comments":[{"body":"CI [!https://ci-cassandra.apache.org/job/Cassandra-devbranch/949/badge/icon!|https://ci-cassandra.apache.org/job/Cassandra-devbranch/949/]\r\n(jvm upgrade dtests known broken on the [cassandra-4.0.0|https://ci-cassandra.apache.org/job/Cassandra-4.0.0/52/#showFailuresLink] branch)","created":"2021-07-18T16:50:39.269+0000"},{"body":"I've left some nits and structural thoughts in the PR, but the fixes and new test look good.\r\n\r\n+1","created":"2021-07-19T21:24:05.234+0000"},{"body":"+1","created":"2021-07-20T20:06:11.808+0000"},{"body":"Here are the pre-commit branches and in-progress Circle CI runs...\r\n\r\n||[4.0.0|https://github.com/apache/cassandra/pull/1112]|[Circle CI|https://app.circleci.com/pipelines/github/maedhroz/cassandra?branch=CASSANDRA-16808-4.0.0]||\r\n||[4.0|https://github.com/apache/cassandra/pull/1113]|[Circle CI|https://app.circleci.com/pipelines/github/maedhroz/cassandra?branch=CASSANDRA-16808-4.0]||\r\n||[trunk|https://github.com/apache/cassandra/pull/1114]|[Circle CI|https://app.circleci.com/pipelines/github/maedhroz/cassandra?branch=CASSANDRA-16808-trunk]|","created":"2021-07-20T21:58:53.157+0000"},{"body":"The 4.0.0 test run is perfect.\r\n\r\n[~jmeredithco] 4.0 and trunk aren't quite there. We're seeing CASSANDRA-16803 and some other assorted timeout and heap space problems, but those are almost certainly unrelated. The troubling bit is what looks like {{MixedModeMessageForwardTest.checkWritesForwardedToOtherDcTest}}, which is new, [failing|https://app.circleci.com/pipelines/github/maedhroz/cassandra/295/workflows/1e5bff9f-36db-46c8-bdf2-419f830a4b3f/jobs/1792/tests#failed-test-1].\r\n\r\nUPDATE: This is a 3.0 -> 3.11 problem, so I'm slightly refactoring the 4.0 and trunk patches to mimic the single upgrade path test in 4.0.0, which is all we should need...","created":"2021-07-21T03:44:21.451+0000"},{"body":"+1\r\n\r\nIt's a shame we didn't actually maintain the serialisation behaviour with the introduction of ports as we had previously in 3.0, but what's a byte here or there amongst friends, I suppose.","created":"2021-07-21T07:54:32.438+0000"},{"body":"Committed.\r\n\r\n4.0.0: https://github.com/apache/cassandra/commit/4d4e1e88d095e10d53b59bf004a59709c3cee186\r\n\r\n4.0: https://github.com/apache/cassandra/commit/cdf68e39ba9ab399936d3359bc0d24b4f38dbae8\r\n\r\ntrunk: https://github.com/apache/cassandra/commit/c1f235f71d013318b86bddee3121beb4923ae0f9","created":"2021-07-21T17:21:45.077+0000"}],"conversations":[{"body":"Fixing CASSANDRA-16797 has exposed an issue with the way {{FWD_FRM}} is serialized.\r\n\r\nIn the code cleanup during the internode messaging refactor, the serialization for {{FWD_FRM}} (the endpoint to respond to for forwarded messages) was implemented using the same serialization format as CompactEndpointSerializationHelper which prefixes the address bytes with their length, however the FWD_FRM parameter value does not include a length and just converts the parameter value to an InetAddress.\r\n\r\nIn a mixed version cluster this causes the pre-4.0 nodes to fail when deserializing the mutation\r\n{code:java}\r\njava.lang.RuntimeException: java.net.UnknownHostException: addr is of illegal length\r\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:72) ~[dtest-3.0.25.jar:na]\r\n at java.base/java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:515) ~[na:na]\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:162) ~[dtest-3.0.25.jar:na]\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:134) ~[dtest-3.0.25.jar:na]\r\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:109) ~[dtest-3.0.25.jar:na]\r\n at java.base/java.lang.Thread.run(Thread.java:834) ~[na:na]\r\nCaused by: java.net.UnknownHostException: addr is of illegal length\r\n at java.base/java.net.InetAddress.getByAddress(InetAddress.java:1208) ~[na:na]\r\n at java.base/java.net.InetAddress.getByAddress(InetAddress.java:1571) ~[na:na]\r\n at org.apache.cassandra.db.MutationVerbHandler.doVerb(MutationVerbHandler.java:57) ~[dtest-3.0.25.jar:na]\r\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:67) ~[dtest-3.0.25.jar:na]\r\n ... 5 common frames omitted\r\n{code}\r\nUnfortunately there isn't a clean fix I can see as {{org.apache.cassandra.io.IVersionedAsymmetricSerializer#deserialize}} used to deserialize the FWD_FRM address does not take a maximum length to deserialize and it's impossible to tell definitely know if it's an IPv4 or IPv6 address from the first four bytes.\r\n\r\nThe patch I'm submitting special-cases the deserializing pre-4.0 {{FWD_FRM}} parameters in the {{Message}} deserializer. That seems preferable to extending the deserialization interface or creating a new {{DataInputBuffer}} limited by the parameter value length.\r\n\r\nOnce that was fixed, the INSERT statements were still failing which I tracked down to the 4.0 optimization of serializing the forwarded message once if the message id is the same\r\n [https://github.com/apache/cassandra/blob/cassandra-4.0/src/java/org/apache/cassandra/db/MutationVerbHandler.java#L76]\r\n\r\nIn the test case I wrote, only one message was being forwarded and that had a different id to the original forwarded message. The {{useSameMessageID}} method only checked message Ids within the forwarded messages.\r\n\r\n \r\n\r\nCode Details:\r\n\r\nWhen MutationVerbHandler.forwardToLocalNodes is constructing the forwarding message it just stores the the byte array representing the IPv4 or IPv6 address in the parameter array.\r\n\r\n(link [https://github.com/apache/cassandra/blob/44604b7316fcbfd7d0d7425e75cd7ebe267e3247/src/java/org/apache/cassandra/db/MutationVerbHandler.java#L90] )\r\n{code:java}\r\n private static void forwardToLocalNodes(Mutation mutation, MessagingService.Verb verb, byte[] forwardBytes, InetAddress from) throws IOException\r\n {\r\n try (DataInputStream in = new DataInputStream(new FastByteArrayInputStream(forwardBytes)))\r\n {\r\n int size = in.readInt();\r\n\r\n // tell the recipients who to send their ack to\r\n MessageOut message = new MessageOut<>(verb, mutation, Mutation.serializer).withParameter(Mutation.FORWARD_FROM, from.getAddress());\r\n{code}\r\nWhen the message is serialized in 3.0 MessageOut.serialize, that raw entry of bytes is written with the length\r\n\r\n(link [https://github.com/apache/cassandra/blob/44604b7316fcbfd7d0d7425e75cd7ebe267e3247/src/java/org/apache/cassandra/net/MessageOut.java#L119] )\r\n{code:java}\r\n public void serialize(DataOutputPlus out, int version) throws IOException\r\n {\r\n CompactEndpointSerializationHelper.serialize(from, out);\r\n\r\n out.writeInt(MessagingService.Verb.convertForMessagingServiceVersion(verb, version).getId());\r\n out.writeInt(parameters.size());\r\n for (Map.Entry entry : parameters.entrySet())\r\n {\r\n out.writeUTF(entry.getKey());\r\n out.writeInt(entry.getValue().length);\r\n out.write(entry.getValue());\r\n }\r\n ....\r\n }\r\n{code}\r\nAnd we do the same on 4.0, however in 4.0 the parameter is serialized using the ParamType enum\r\n\r\n(link [https://github.com/apache/cassandra/blob/fcd30b6e0db3622a8e78e9aa35221f630c77f6de/src/java/org/apache/cassandra/net/Message.java#L1154] )\r\n{code:java}\r\n for (int i = 0; i < count; i++)\r\n {\r\n ParamType type = version >= VERSION_40\r\n ? ParamType.lookUpById(Ints.checkedCast(in.readUnsignedVInt()))\r\n : ParamType.lookUpByAlias(in.readUTF());\r\n\r\n int length = version >= VERSION_40\r\n ? Ints.checkedCast(in.readUnsignedVInt())\r\n : in.readInt();\r\n\r\n if (null != type)\r\n params.put(type, type.serializer.deserialize(in, version));\r\n else\r\n in.skipBytesFully(length); // forward compatibiliy with minor version changes\r\n }\r\n{code}\r\n(link [https://github.com/apache/cassandra/blob/fcd30b6e0db3622a8e78e9aa35221f630c77f6de/src/java/org/apache/cassandra/net/ParamType.java#L45] )\r\n{code:java}\r\npublic enum ParamType\r\n{\r\n FORWARD_TO (0, \"FWD_TO\", ForwardingInfo.serializer),\r\n RESPOND_TO (1, \"FWD_FRM\", inetAddressAndPortSerializer),\r\n ...\r\n}\r\n{code}\r\nThe {{InetAddressAndPortSerializer}} has been based on the 3.0 {{CompactEndpointSerializationHelper}} encoding used in the message header,\r\n however that format includes a single byte with the length of the address when pre-4.0 nodes are just expecting the parameter value\r\n to contain the raw address bytes.\r\n\r\n(link [https://github.com/apache/cassandra/blob/fcd30b6e0db3622a8e78e9aa35221f630c77f6de/src/java/org/apache/cassandra/locator/InetAddressAndPort.java#L308] )\r\n{code:java}\r\n if (version >= MessagingService.VERSION_40)\r\n {\r\n out.writeByte(buf.length + 2);\r\n out.write(buf);\r\n out.writeShort(endpoint.port);\r\n }\r\n else\r\n {\r\n out.writeByte(buf.length); //// Surprise! Bonus byte!\r\n out.write(buf);\r\n }\r\n{code}","from":"reporter","subject":"Pre-4.0 FWD_FRM message parameter serialization and message-id forwarding is incorrect"},{"body":"CI [!https://ci-cassandra.apache.org/job/Cassandra-devbranch/949/badge/icon!|https://ci-cassandra.apache.org/job/Cassandra-devbranch/949/]\r\n(jvm upgrade dtests known broken on the [cassandra-4.0.0|https://ci-cassandra.apache.org/job/Cassandra-4.0.0/52/#showFailuresLink] branch)","from":"developer"},{"body":"I've left some nits and structural thoughts in the PR, but the fixes and new test look good.\r\n\r\n+1","from":"developer"},{"body":"+1","from":"developer"},{"body":"Here are the pre-commit branches and in-progress Circle CI runs...\r\n\r\n||[4.0.0|https://github.com/apache/cassandra/pull/1112]|[Circle CI|https://app.circleci.com/pipelines/github/maedhroz/cassandra?branch=CASSANDRA-16808-4.0.0]||\r\n||[4.0|https://github.com/apache/cassandra/pull/1113]|[Circle CI|https://app.circleci.com/pipelines/github/maedhroz/cassandra?branch=CASSANDRA-16808-4.0]||\r\n||[trunk|https://github.com/apache/cassandra/pull/1114]|[Circle CI|https://app.circleci.com/pipelines/github/maedhroz/cassandra?branch=CASSANDRA-16808-trunk]|","from":"developer"},{"body":"The 4.0.0 test run is perfect.\r\n\r\n[~jmeredithco] 4.0 and trunk aren't quite there. We're seeing CASSANDRA-16803 and some other assorted timeout and heap space problems, but those are almost certainly unrelated. The troubling bit is what looks like {{MixedModeMessageForwardTest.checkWritesForwardedToOtherDcTest}}, which is new, [failing|https://app.circleci.com/pipelines/github/maedhroz/cassandra/295/workflows/1e5bff9f-36db-46c8-bdf2-419f830a4b3f/jobs/1792/tests#failed-test-1].\r\n\r\nUPDATE: This is a 3.0 -> 3.11 problem, so I'm slightly refactoring the 4.0 and trunk patches to mimic the single upgrade path test in 4.0.0, which is all we should need...","from":"developer"},{"body":"+1\r\n\r\nIt's a shame we didn't actually maintain the serialisation behaviour with the introduction of ports as we had previously in 3.0, but what's a byte here or there amongst friends, I suppose.","from":"developer"},{"body":"Committed.\r\n\r\n4.0.0: https://github.com/apache/cassandra/commit/4d4e1e88d095e10d53b59bf004a59709c3cee186\r\n\r\n4.0: https://github.com/apache/cassandra/commit/cdf68e39ba9ab399936d3359bc0d24b4f38dbae8\r\n\r\ntrunk: https://github.com/apache/cassandra/commit/c1f235f71d013318b86bddee3121beb4923ae0f9","from":"developer"}],"created":"2021-07-17T04:32:36.000+0000","description":"Fixing CASSANDRA-16797 has exposed an issue with the way {{FWD_FRM}} is serialized.\r\n\r\nIn the code cleanup during the internode messaging refactor, the serialization for {{FWD_FRM}} (the endpoint to respond to for forwarded messages) was implemented using the same serialization format as CompactEndpointSerializationHelper which prefixes the address bytes with their length, however the FWD_FRM parameter value does not include a length and just converts the parameter value to an InetAddress.\r\n\r\nIn a mixed version cluster this causes the pre-4.0 nodes to fail when deserializing the mutation\r\n{code:java}\r\njava.lang.RuntimeException: java.net.UnknownHostException: addr is of illegal length\r\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:72) ~[dtest-3.0.25.jar:na]\r\n at java.base/java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:515) ~[na:na]\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$FutureTask.run(AbstractLocalAwareExecutorService.java:162) ~[dtest-3.0.25.jar:na]\r\n at org.apache.cassandra.concurrent.AbstractLocalAwareExecutorService$LocalSessionFutureTask.run(AbstractLocalAwareExecutorService.java:134) ~[dtest-3.0.25.jar:na]\r\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:109) ~[dtest-3.0.25.jar:na]\r\n at java.base/java.lang.Thread.run(Thread.java:834) ~[na:na]\r\nCaused by: java.net.UnknownHostException: addr is of illegal length\r\n at java.base/java.net.InetAddress.getByAddress(InetAddress.java:1208) ~[na:na]\r\n at java.base/java.net.InetAddress.getByAddress(InetAddress.java:1571) ~[na:na]\r\n at org.apache.cassandra.db.MutationVerbHandler.doVerb(MutationVerbHandler.java:57) ~[dtest-3.0.25.jar:na]\r\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:67) ~[dtest-3.0.25.jar:na]\r\n ... 5 common frames omitted\r\n{code}\r\nUnfortunately there isn't a clean fix I can see as {{org.apache.cassandra.io.IVersionedAsymmetricSerializer#deserialize}} used to deserialize the FWD_FRM address does not take a maximum length to deserialize and it's impossible to tell definitely know if it's an IPv4 or IPv6 address from the first four bytes.\r\n\r\nThe patch I'm submitting special-cases the deserializing pre-4.0 {{FWD_FRM}} parameters in the {{Message}} deserializer. That seems preferable to extending the deserialization interface or creating a new {{DataInputBuffer}} limited by the parameter value length.\r\n\r\nOnce that was fixed, the INSERT statements were still failing which I tracked down to the 4.0 optimization of serializing the forwarded message once if the message id is the same\r\n [https://github.com/apache/cassandra/blob/cassandra-4.0/src/java/org/apache/cassandra/db/MutationVerbHandler.java#L76]\r\n\r\nIn the test case I wrote, only one message was being forwarded and that had a different id to the original forwarded message. The {{useSameMessageID}} method only checked message Ids within the forwarded messages.\r\n\r\n \r\n\r\nCode Details:\r\n\r\nWhen MutationVerbHandler.forwardToLocalNodes is constructing the forwarding message it just stores the the byte array representing the IPv4 or IPv6 address in the parameter array.\r\n\r\n(link [https://github.com/apache/cassandra/blob/44604b7316fcbfd7d0d7425e75cd7ebe267e3247/src/java/org/apache/cassandra/db/MutationVerbHandler.java#L90] )\r\n{code:java}\r\n private static void forwardToLocalNodes(Mutation mutation, MessagingService.Verb verb, byte[] forwardBytes, InetAddress from) throws IOException\r\n {\r\n try (DataInputStream in = new DataInputStream(new FastByteArrayInputStream(forwardBytes)))\r\n {\r\n int size = in.readInt();\r\n\r\n // tell the recipients who to send their ack to\r\n MessageOut message = new MessageOut<>(verb, mutation, Mutation.serializer).withParameter(Mutation.FORWARD_FROM, from.getAddress());\r\n{code}\r\nWhen the message is serialized in 3.0 MessageOut.serialize, that raw entry of bytes is written with the length\r\n\r\n(link [https://github.com/apache/cassandra/blob/44604b7316fcbfd7d0d7425e75cd7ebe267e3247/src/java/org/apache/cassandra/net/MessageOut.java#L119] )\r\n{code:java}\r\n public void serialize(DataOutputPlus out, int version) throws IOException\r\n {\r\n CompactEndpointSerializationHelper.serialize(from, out);\r\n\r\n out.writeInt(MessagingService.Verb.convertForMessagingServiceVersion(verb, version).getId());\r\n out.writeInt(parameters.size());\r\n for (Map.Entry entry : parameters.entrySet())\r\n {\r\n out.writeUTF(entry.getKey());\r\n out.writeInt(entry.getValue().length);\r\n out.write(entry.getValue());\r\n }\r\n ....\r\n }\r\n{code}\r\nAnd we do the same on 4.0, however in 4.0 the parameter is serialized using the ParamType enum\r\n\r\n(link [https://github.com/apache/cassandra/blob/fcd30b6e0db3622a8e78e9aa35221f630c77f6de/src/java/org/apache/cassandra/net/Message.java#L1154] )\r\n{code:java}\r\n for (int i = 0; i < count; i++)\r\n {\r\n ParamType type = version >= VERSION_40\r\n ? ParamType.lookUpById(Ints.checkedCast(in.readUnsignedVInt()))\r\n : ParamType.lookUpByAlias(in.readUTF());\r\n\r\n int length = version >= VERSION_40\r\n ? Ints.checkedCast(in.readUnsignedVInt())\r\n : in.readInt();\r\n\r\n if (null != type)\r\n params.put(type, type.serializer.deserialize(in, version));\r\n else\r\n in.skipBytesFully(length); // forward compatibiliy with minor version changes\r\n }\r\n{code}\r\n(link [https://github.com/apache/cassandra/blob/fcd30b6e0db3622a8e78e9aa35221f630c77f6de/src/java/org/apache/cassandra/net/ParamType.java#L45] )\r\n{code:java}\r\npublic enum ParamType\r\n{\r\n FORWARD_TO (0, \"FWD_TO\", ForwardingInfo.serializer),\r\n RESPOND_TO (1, \"FWD_FRM\", inetAddressAndPortSerializer),\r\n ...\r\n}\r\n{code}\r\nThe {{InetAddressAndPortSerializer}} has been based on the 3.0 {{CompactEndpointSerializationHelper}} encoding used in the message header,\r\n however that format includes a single byte with the length of the address when pre-4.0 nodes are just expecting the parameter value\r\n to contain the raw address bytes.\r\n\r\n(link [https://github.com/apache/cassandra/blob/fcd30b6e0db3622a8e78e9aa35221f630c77f6de/src/java/org/apache/cassandra/locator/InetAddressAndPort.java#L308] )\r\n{code:java}\r\n if (version >= MessagingService.VERSION_40)\r\n {\r\n out.writeByte(buf.length + 2);\r\n out.write(buf);\r\n out.writeShort(endpoint.port);\r\n }\r\n else\r\n {\r\n out.writeByte(buf.length); //// Surprise! Bonus byte!\r\n out.write(buf);\r\n }\r\n{code}","issue_id":"13390199","key":"CASSANDRA-16808","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2021-07-21T17:21:45.000+0000","role":"fixed_distractor","summary":"Pre-4.0 FWD_FRM message parameter serialization and message-id forwarding is incorrect"} {"case_id":"13394438","cluster":"DISTRACTOR-CASSANDRA-16839","comments":[{"body":"I can confirm that I'm experiencing the same issue on a Cassandra 4.0.1 cluster on production. New snapshot is created for those system tables when the server is rebooted or the Cassandra service is restarted. The size of the snapshot is small, but if left untreated, the filesystem may run out of free inodes eventually.","created":"2021-09-25T16:55:06.743+0000"},{"body":"This is caused by CASSANDRA-15776 truncating these tables via CQL, which does not have a mechanism for truncation without snapshotting. I actually couldn't find a way to truncate without snapshotting, so this patch adds one, uses it instead, and additionally replaces the other two truncateBlocking instances that generate snapshots for prepared statements and available ranges.\r\n\r\n\r\n||Branch||CI||\r\n|[4.0|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-4.0]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-4.0], [!https://ci-cassandra.apache.org/job/Cassandra-devbranch/1184/badge/icon!|https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/1184/pipeline]|\r\n|[trunk|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-trunk]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-trunk], [!https://ci-cassandra.apache.org/job/Cassandra-devbranch/1185/badge/icon!|https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/1185/pipeline]|","created":"2021-10-06T16:21:16.477+0000"},{"body":"The fix looks good to me. As for using the new {{truncateBlockingWithoutSnapshot}} method in {{resetAvailableRanges}} and {{resetPreparedStatements}} to prevent the unwanted snapshot, shouldn't we backport that to 3.0 and 3.11?","created":"2021-10-08T12:30:13.011+0000"},{"body":"The patch looks good to me.","created":"2021-10-08T12:33:57.995+0000"},{"body":"bq. shouldn't we backport that to 3.0 and 3.11?\r\n\r\nGood catch. I've fixed the nits, backported this to 3.11 and 3.0 (which didn't have CASSANDRA-13641), and also renamed the variable for prepared statements correctly.\r\n\r\n||Branch||CI||\r\n|[3.0|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-3.0]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-3.0], [!https://ci-cassandra.apache.org/job/Cassandra-devbranch/1197/badge/icon!|https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/1197/pipeline]|\r\n|[3.11|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-3.11]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-3.11], [!https://ci-cassandra.apache.org/job/Cassandra-devbranch/1198/badge/icon!|https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/1198/pipeline]|\r\n|[4.0|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-4.0]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-4.0]|\r\n|[trunk|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-trunk]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-trunk]|\r\n","created":"2021-10-08T16:16:14.643+0000"},{"body":"Great, looks good to me, +1 if CI is ok. Nice catch on the var name.","created":"2021-10-08T16:39:44.276+0000"},{"body":"I don't think there are any related failures, committed. Thank you for the reviews.","created":"2021-10-11T17:12:48.092+0000"}],"conversations":[{"body":"When testing cassandra 4.0 on ccm I noticed that everytime I restart a node, truncation snapshots are created for the tables {{system.table_estimates}} and {{system.size_estimates}}:\r\n\r\n{noformat}\r\n$ ccm create -n 1 test -s\r\n$ ccm node1 stop\r\n$ ccm node1 start\r\n$ ccm node1 stop\r\n$ ccm node1 start\r\n$ ccm node1 nodetool listsnapshots\r\n\r\nSnapshot Details:\r\nSnapshot name Keyspace name Column family name True size Size on disk\r\ntruncated-1628599001857-table_estimates system table_estimates 0 bytes 13 bytes\r\ntruncated-1628599099560-table_estimates system table_estimates 0 bytes 13 bytes\r\ntruncated-1628599001736-size_estimates system size_estimates 0 bytes 13 bytes\r\ntruncated-1628599057438-table_estimates system table_estimates 6.16 KiB 6.19 KiB\r\ntruncated-1628599099458-size_estimates system size_estimates 0 bytes 13 bytes\r\ntruncated-1628599057340-size_estimates system size_estimates 5.73 KiB 5.76 KiB\r\n\r\nTotal TrueDiskSpaceUsed: 0 bytes\r\n{noformat}\r\n\r\nNot sure if this is expected behavior, but feels like a bug to me.\r\n\r\nReproduced on 4.0, not sure if it reproduces on lower versions.","from":"reporter","subject":"Truncation snapshots unnecessarily created on node startup"},{"body":"I can confirm that I'm experiencing the same issue on a Cassandra 4.0.1 cluster on production. New snapshot is created for those system tables when the server is rebooted or the Cassandra service is restarted. The size of the snapshot is small, but if left untreated, the filesystem may run out of free inodes eventually.","from":"developer"},{"body":"This is caused by CASSANDRA-15776 truncating these tables via CQL, which does not have a mechanism for truncation without snapshotting. I actually couldn't find a way to truncate without snapshotting, so this patch adds one, uses it instead, and additionally replaces the other two truncateBlocking instances that generate snapshots for prepared statements and available ranges.\r\n\r\n\r\n||Branch||CI||\r\n|[4.0|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-4.0]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-4.0], [!https://ci-cassandra.apache.org/job/Cassandra-devbranch/1184/badge/icon!|https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/1184/pipeline]|\r\n|[trunk|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-trunk]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-trunk], [!https://ci-cassandra.apache.org/job/Cassandra-devbranch/1185/badge/icon!|https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/1185/pipeline]|","from":"developer"},{"body":"The fix looks good to me. As for using the new {{truncateBlockingWithoutSnapshot}} method in {{resetAvailableRanges}} and {{resetPreparedStatements}} to prevent the unwanted snapshot, shouldn't we backport that to 3.0 and 3.11?","from":"developer"},{"body":"The patch looks good to me.","from":"developer"},{"body":"bq. shouldn't we backport that to 3.0 and 3.11?\r\n\r\nGood catch. I've fixed the nits, backported this to 3.11 and 3.0 (which didn't have CASSANDRA-13641), and also renamed the variable for prepared statements correctly.\r\n\r\n||Branch||CI||\r\n|[3.0|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-3.0]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-3.0], [!https://ci-cassandra.apache.org/job/Cassandra-devbranch/1197/badge/icon!|https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/1197/pipeline]|\r\n|[3.11|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-3.11]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-3.11], [!https://ci-cassandra.apache.org/job/Cassandra-devbranch/1198/badge/icon!|https://ci-cassandra.apache.org/blue/organizations/jenkins/Cassandra-devbranch/detail/Cassandra-devbranch/1198/pipeline]|\r\n|[4.0|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-4.0]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-4.0]|\r\n|[trunk|https://github.com/driftx/cassandra/tree/CASSANDRA-16839-trunk]|[circle|https://app.circleci.com/pipelines/github/driftx/cassandra?branch=CASSANDRA-16839-trunk]|\r\n","from":"developer"},{"body":"Great, looks good to me, +1 if CI is ok. Nice catch on the var name.","from":"developer"},{"body":"I don't think there are any related failures, committed. Thank you for the reviews.","from":"developer"}],"created":"2021-08-10T12:46:55.000+0000","description":"When testing cassandra 4.0 on ccm I noticed that everytime I restart a node, truncation snapshots are created for the tables {{system.table_estimates}} and {{system.size_estimates}}:\r\n\r\n{noformat}\r\n$ ccm create -n 1 test -s\r\n$ ccm node1 stop\r\n$ ccm node1 start\r\n$ ccm node1 stop\r\n$ ccm node1 start\r\n$ ccm node1 nodetool listsnapshots\r\n\r\nSnapshot Details:\r\nSnapshot name Keyspace name Column family name True size Size on disk\r\ntruncated-1628599001857-table_estimates system table_estimates 0 bytes 13 bytes\r\ntruncated-1628599099560-table_estimates system table_estimates 0 bytes 13 bytes\r\ntruncated-1628599001736-size_estimates system size_estimates 0 bytes 13 bytes\r\ntruncated-1628599057438-table_estimates system table_estimates 6.16 KiB 6.19 KiB\r\ntruncated-1628599099458-size_estimates system size_estimates 0 bytes 13 bytes\r\ntruncated-1628599057340-size_estimates system size_estimates 5.73 KiB 5.76 KiB\r\n\r\nTotal TrueDiskSpaceUsed: 0 bytes\r\n{noformat}\r\n\r\nNot sure if this is expected behavior, but feels like a bug to me.\r\n\r\nReproduced on 4.0, not sure if it reproduces on lower versions.","issue_id":"13394438","key":"CASSANDRA-16839","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2021-10-11T17:12:48.000+0000","role":"fixed_distractor","summary":"Truncation snapshots unnecessarily created on node startup"} {"case_id":"13408819","cluster":"DISTRACTOR-CASSANDRA-17070","comments":[{"body":"Hardening by trying to rerun the cleanup logic on any leftovers in case the first attempt failed. Trunk PR will be identical.","created":"2021-10-28T06:16:40.617+0000"},{"body":"I'm not sure I understand why the MV cleanup can fail. Would the second round of cleanup guarantee that the MV cleanups succeed, or would it reduce the chances of having leftovers because it's unlikely to have the same error twice? Maybe it's just giving more time for the first async cleanup to happen?\r\n\r\nIn any case, here are 500 runs of each ViewComplex test:\r\n * [ViewComplexDeletionsTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1090/workflows/4362eefe-a13a-4a8c-a8eb-9e543a74623a/jobs/10187]\r\n * [ViewComplexLivenessTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1091/workflows/a9185c89-ceff-4ddd-90c0-960281233fe3/jobs/10189]\r\n * [ViewComplexTTLTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1092/workflows/1f74edb4-3a41-47fd-b95c-0edc1d514a20/jobs/10191]\r\n * [ViewComplexTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1093/workflows/f2e69851-19ea-4ab6-9315-5af9d721ec3d/jobs/10193]\r\n * [ViewComplexUpdatesTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1094/workflows/0d25bf72-7c90-4287-92d6-21a751082e21/jobs/10196]\r\n\r\nFrom these runs only a test runner for {{ViewComplexTest}} is failing due to timeouts.\r\n\r\nIt's not strictly related to the fix, but it seems that there is some repeated code among the five tests in which {{ViewComplexTest}} was split by CASSANDRA-16670. Probably we should do the same thing that we did with {{ViewTest}} in CASSANDRA-16777 and extract a superclass to avoid code duplication. This would make changes like this one a bit easier to do. I gave it a go in [this commit|https://github.com/adelapena/cassandra/commit/204f3bbe69aae9165c0570224bfd811c8bbd4d07], and it saves us some 176 lines of code, which I think is always nice.\r\n\r\nHere are some repeated runs for the tests using a common superclass, with no failures so far:\r\n * [ViewComplexDeletionsTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1096/workflows/ded0dfe1-e736-468a-babb-0c007a6bf3c9/jobs/10201]\r\n * [ViewComplexLivenessTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1097/workflows/544b429d-b712-428c-b97d-edd50076992a/jobs/10203]\r\n * [ViewComplexTTLTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1098/workflows/c062d815-2e26-4a75-aa64-656bfc48c8cc/jobs/10205]\r\n * [ViewComplexTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1099/workflows/1f834191-b30b-4380-87ce-da257fac731e/jobs/10207]\r\n * [ViewComplexUpdatesTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1095/workflows/3df8138d-1d7b-4283-be3c-a23cb2c0bc40]\r\n\r\nAn alternative/complementary approach to avoid problems with leftovers from previous runs is making sure that all MVs are created with a new, unique name. This way the tests don't collide with any surviving MVs if the cleanup has failed or if it hasn't finished. {{CQLTester}} does this with keyspace, table, function and aggregate names.\r\n\r\nI gave a go to this generation of unique names [here|https://github.com/adelapena/cassandra/commit/e1928447e4139f1ad4126d9ed34e3de9d3bceb10]. It is just for {{ViewComplexTest}}, but we may consider having a more generic and reusable solution for creating/cleaning MVs in {{CQLTester}}, although I think we don't have to do it here. wdyt?","created":"2021-10-28T16:43:07.939+0000"},{"body":"I forgot to mention that I found an unrelated mistake [here|https://github.com/apache/cassandra/blob/trunk/test/unit/org/apache/cassandra/cql3/ViewComplexTest.java#L142-L148], where two tests are supposed to run a scenario with and without flushing the table, but both of them run the scenario with flush. It's fixed in the commit extracting a superclass, [here|https://github.com/adelapena/cassandra/blob/204f3bbe69aae9165c0570224bfd811c8bbd4d07/test/unit/org/apache/cassandra/cql3/ViewComplexTest.java#L55-L61].","created":"2021-10-28T17:01:08.950+0000"},{"body":"+1 on the proposed change to use a super class similar to #16777, it was on my list for some time but it was a lower priority. Honestly, we should have done it at first place like that. Thanks [~adelapena]!","created":"2021-10-28T19:05:45.440+0000"},{"body":"[~adelapena] the reason to run that code twice was that the test keeps a list of the views. So if a cleanup failed the next one would pick up from the failing point\r\n\r\nBut +1 to the abstract class and unique MV name. I ported your changes to my PR but: I deleted the extra cleanup which is not needed anymore, I changed the config to run {{ViewComplex*Test}} instead of having one commit per repeat test and I attached results to the PR with 100 repeats and 0 failures (CASSANDRA-17083).\r\n\r\nI am only on the fence about the cleanup logic, which we could remove now there are unique MV names, but on the other hand it's nice to have it. I think I am +1 on how things are right now. Wdyt?","created":"2021-10-29T08:07:05.319+0000"},{"body":"It seems that the multiplexer runs in the PR have failed due to the wildcard. I think it would work if we use {{ant test}} instead of {{ant testsome}}, with the unqualified class name:\r\n{code:java}\r\n.circleci/generate.sh -m \\\r\n -e REPEATED_UTEST_TARGET=test \\\r\n -e REPEATED_UTEST_COUNT=500 \\\r\n -e REPEATED_UTEST_CLASS=ViewComplex*Test\r\n{code}\r\nThe wildcard is a great idea, it would have saved me time with the above runs :)\r\n\r\nI think we can try some more runs without the duplicated cleanup logic and see how it goes, wdyt?","created":"2021-10-29T10:24:34.734+0000"},{"body":"They didn't fail sort to speak... They show the 13200 passed tests with no _test_ failures. It failed the run bc of the wilcard. It is easy to remove the noise by switching to {{test}} instead of {{testsome}}.\r\n\r\nRemoving the cleanup logic breaks the {{CQLTester}} teardown when it tries to drop tables that have attached MVs. So I want to give that a thought...","created":"2021-10-29T10:45:55.236+0000"},{"body":"[Here|https://app.circleci.com/pipelines/github/adelapena/cassandra/1101/workflows/c335ffa0-1513-40a1-b925-292e0d03cee3/jobs/10214] are other 500 rounds of the tests without the duplicated cleanup. There is a single error in {{ViewComplexUpdatesTest.testUpdateWithColumnTimestampSmallerThanPk}}, apparently due to a query timeout (see the bottom of the [stdout|[REDACTED_AWS_PRESIGNED_URL]] file).","created":"2021-10-29T14:55:31.299+0000"},{"body":"Right. But it's on view creation, not the cleanup logic. Seems like a legit timeout. I've had the feeling timeouts have to be raised for some time. I still think we're good to merge. Or are you suggesting we put the cleanup back in?\r\n\r\n{noformat}\r\n[junit-timeout] Testcase: testUpdateWithColumnTimestampSmallerThanPkWithFlush[3](org.apache.cassandra.cql3.ViewComplexUpdatesTest):\tCaused an ERROR\r\n[junit-timeout] [localhost/127.0.0.1:43739] Timed out waiting for server response\r\n[junit-timeout] com.datastax.driver.core.exceptions.OperationTimedOutException: [localhost/127.0.0.1:43739] Timed out waiting for server response\r\n[junit-timeout] \tat com.datastax.driver.core.exceptions.OperationTimedOutException.copy(OperationTimedOutException.java:43)\r\n[junit-timeout] \tat com.datastax.driver.core.exceptions.OperationTimedOutException.copy(OperationTimedOutException.java:25)\r\n[junit-timeout] \tat com.datastax.driver.core.DriverThrowables.propagateCause(DriverThrowables.java:35)\r\n[junit-timeout] \tat com.datastax.driver.core.DefaultResultSetFuture.getUninterruptibly(DefaultResultSetFuture.java:293)\r\n[junit-timeout] \tat com.datastax.driver.core.AbstractSession.execute(AbstractSession.java:58)\r\n[junit-timeout] \tat com.datastax.driver.core.AbstractSession.execute(AbstractSession.java:45)\r\n[junit-timeout] \tat org.apache.cassandra.cql3.CQLTester.executeNet(CQLTester.java:972)\r\n[junit-timeout] \tat org.apache.cassandra.cql3.ViewComplexTester.createView(ViewComplexTester.java:83)\r\n[junit-timeout] \tat org.apache.cassandra.cql3.ViewComplexUpdatesTest.testUpdateWithColumnTimestampSmallerThanPk(ViewComplexUpdatesTest.java:224)\r\n[junit-timeout] \tat org.apache.cassandra.cql3.ViewComplexUpdatesTest.testUpdateWithColumnTimestampSmallerThanPkWithFlush(ViewComplexUpdatesTest.java:207)\r\n[junit-timeout] Caused by: com.datastax.driver.core.exceptions.OperationTimedOutException: [localhost/127.0.0.1:43739] Timed out waiting for server response\r\n[junit-timeout] \tat com.datastax.driver.core.RequestHandler$SpeculativeExecution.onTimeout(RequestHandler.java:979)\r\n[junit-timeout] \tat com.datastax.driver.core.Connection$ResponseHandler$1.run(Connection.java:1636)\r\n[junit-timeout] \tat com.datastax.shaded.netty.util.HashedWheelTimer$HashedWheelTimeout.expire(HashedWheelTimer.java:663)\r\n[junit-timeout] \tat com.datastax.shaded.netty.util.HashedWheelTimer$HashedWheelBucket.expireTimeouts(HashedWheelTimer.java:738)\r\n[junit-timeout] \tat com.datastax.shaded.netty.util.HashedWheelTimer$Worker.run(HashedWheelTimer.java:466)\r\n[junit-timeout] \tat com.datastax.shaded.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n[junit-timeout] \tat java.lang.Thread.run(Thread.java:748)\r\n{noformat}\r\n","created":"2021-11-02T08:53:11.965+0000"},{"body":"I don't think we should put the duplicated cleanup logic back, in principle not reusing MV names should be enough.\r\n\r\nI mentioned the query timeout to indicate that the failure is not related to what we are addressing here, so I think the run can be considered successful. By the way, [here|https://app.circleci.com/pipelines/github/adelapena/cassandra/1101/workflows/63014008-3d52-4d07-a5aa-748fad389ec7] is the repeated run for j11, all green. I think we are ready to merge.","created":"2021-11-02T14:55:25.029+0000"},{"body":"I added the trunk PR with CI. There's no rush so this will give you time to take a look at it. Let me know if you're still ok I merge. ","created":"2021-11-03T11:56:43.021+0000"},{"body":"The PR for trunk looks good to me. Here are 200 runs of {{ViewComplex*Test}} in the PR for trunk:\r\n * [j8|https://app.circleci.com/pipelines/github/adelapena/cassandra/1114/workflows/2a6ea982-7109-4cf2-b5af-7f06188f2d4f]\r\n * [j11|https://app.circleci.com/pipelines/github/adelapena/cassandra/1114/workflows/12ab59cb-24d5-4795-a309-21b7eb6b187c]\r\n\r\nIf those succeed I think we'll be definitively ready to merge.","created":"2021-11-03T17:11:03.220+0000"},{"body":"Yep wall of green :-)","created":"2021-11-04T07:25:36.380+0000"}],"conversations":[{"body":"I have seen a number of times already the {{ViewComplexTest}} family timeout on test method teardown. This leaves a dirty env behind triggering the following test methods to fail on it. This ticket aims at hardening them.","from":"reporter","subject":"ViewComplexTest hardening"},{"body":"Hardening by trying to rerun the cleanup logic on any leftovers in case the first attempt failed. Trunk PR will be identical.","from":"developer"},{"body":"I'm not sure I understand why the MV cleanup can fail. Would the second round of cleanup guarantee that the MV cleanups succeed, or would it reduce the chances of having leftovers because it's unlikely to have the same error twice? Maybe it's just giving more time for the first async cleanup to happen?\r\n\r\nIn any case, here are 500 runs of each ViewComplex test:\r\n * [ViewComplexDeletionsTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1090/workflows/4362eefe-a13a-4a8c-a8eb-9e543a74623a/jobs/10187]\r\n * [ViewComplexLivenessTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1091/workflows/a9185c89-ceff-4ddd-90c0-960281233fe3/jobs/10189]\r\n * [ViewComplexTTLTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1092/workflows/1f74edb4-3a41-47fd-b95c-0edc1d514a20/jobs/10191]\r\n * [ViewComplexTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1093/workflows/f2e69851-19ea-4ab6-9315-5af9d721ec3d/jobs/10193]\r\n * [ViewComplexUpdatesTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1094/workflows/0d25bf72-7c90-4287-92d6-21a751082e21/jobs/10196]\r\n\r\nFrom these runs only a test runner for {{ViewComplexTest}} is failing due to timeouts.\r\n\r\nIt's not strictly related to the fix, but it seems that there is some repeated code among the five tests in which {{ViewComplexTest}} was split by CASSANDRA-16670. Probably we should do the same thing that we did with {{ViewTest}} in CASSANDRA-16777 and extract a superclass to avoid code duplication. This would make changes like this one a bit easier to do. I gave it a go in [this commit|https://github.com/adelapena/cassandra/commit/204f3bbe69aae9165c0570224bfd811c8bbd4d07], and it saves us some 176 lines of code, which I think is always nice.\r\n\r\nHere are some repeated runs for the tests using a common superclass, with no failures so far:\r\n * [ViewComplexDeletionsTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1096/workflows/ded0dfe1-e736-468a-babb-0c007a6bf3c9/jobs/10201]\r\n * [ViewComplexLivenessTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1097/workflows/544b429d-b712-428c-b97d-edd50076992a/jobs/10203]\r\n * [ViewComplexTTLTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1098/workflows/c062d815-2e26-4a75-aa64-656bfc48c8cc/jobs/10205]\r\n * [ViewComplexTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1099/workflows/1f834191-b30b-4380-87ce-da257fac731e/jobs/10207]\r\n * [ViewComplexUpdatesTest|https://app.circleci.com/pipelines/github/adelapena/cassandra/1095/workflows/3df8138d-1d7b-4283-be3c-a23cb2c0bc40]\r\n\r\nAn alternative/complementary approach to avoid problems with leftovers from previous runs is making sure that all MVs are created with a new, unique name. This way the tests don't collide with any surviving MVs if the cleanup has failed or if it hasn't finished. {{CQLTester}} does this with keyspace, table, function and aggregate names.\r\n\r\nI gave a go to this generation of unique names [here|https://github.com/adelapena/cassandra/commit/e1928447e4139f1ad4126d9ed34e3de9d3bceb10]. It is just for {{ViewComplexTest}}, but we may consider having a more generic and reusable solution for creating/cleaning MVs in {{CQLTester}}, although I think we don't have to do it here. wdyt?","from":"developer"},{"body":"I forgot to mention that I found an unrelated mistake [here|https://github.com/apache/cassandra/blob/trunk/test/unit/org/apache/cassandra/cql3/ViewComplexTest.java#L142-L148], where two tests are supposed to run a scenario with and without flushing the table, but both of them run the scenario with flush. It's fixed in the commit extracting a superclass, [here|https://github.com/adelapena/cassandra/blob/204f3bbe69aae9165c0570224bfd811c8bbd4d07/test/unit/org/apache/cassandra/cql3/ViewComplexTest.java#L55-L61].","from":"developer"},{"body":"+1 on the proposed change to use a super class similar to #16777, it was on my list for some time but it was a lower priority. Honestly, we should have done it at first place like that. Thanks [~adelapena]!","from":"developer"},{"body":"[~adelapena] the reason to run that code twice was that the test keeps a list of the views. So if a cleanup failed the next one would pick up from the failing point\r\n\r\nBut +1 to the abstract class and unique MV name. I ported your changes to my PR but: I deleted the extra cleanup which is not needed anymore, I changed the config to run {{ViewComplex*Test}} instead of having one commit per repeat test and I attached results to the PR with 100 repeats and 0 failures (CASSANDRA-17083).\r\n\r\nI am only on the fence about the cleanup logic, which we could remove now there are unique MV names, but on the other hand it's nice to have it. I think I am +1 on how things are right now. Wdyt?","from":"developer"},{"body":"It seems that the multiplexer runs in the PR have failed due to the wildcard. I think it would work if we use {{ant test}} instead of {{ant testsome}}, with the unqualified class name:\r\n{code:java}\r\n.circleci/generate.sh -m \\\r\n -e REPEATED_UTEST_TARGET=test \\\r\n -e REPEATED_UTEST_COUNT=500 \\\r\n -e REPEATED_UTEST_CLASS=ViewComplex*Test\r\n{code}\r\nThe wildcard is a great idea, it would have saved me time with the above runs :)\r\n\r\nI think we can try some more runs without the duplicated cleanup logic and see how it goes, wdyt?","from":"developer"},{"body":"They didn't fail sort to speak... They show the 13200 passed tests with no _test_ failures. It failed the run bc of the wilcard. It is easy to remove the noise by switching to {{test}} instead of {{testsome}}.\r\n\r\nRemoving the cleanup logic breaks the {{CQLTester}} teardown when it tries to drop tables that have attached MVs. So I want to give that a thought...","from":"developer"},{"body":"[Here|https://app.circleci.com/pipelines/github/adelapena/cassandra/1101/workflows/c335ffa0-1513-40a1-b925-292e0d03cee3/jobs/10214] are other 500 rounds of the tests without the duplicated cleanup. There is a single error in {{ViewComplexUpdatesTest.testUpdateWithColumnTimestampSmallerThanPk}}, apparently due to a query timeout (see the bottom of the [stdout|[REDACTED_AWS_PRESIGNED_URL]] file).","from":"developer"},{"body":"Right. But it's on view creation, not the cleanup logic. Seems like a legit timeout. I've had the feeling timeouts have to be raised for some time. I still think we're good to merge. Or are you suggesting we put the cleanup back in?\r\n\r\n{noformat}\r\n[junit-timeout] Testcase: testUpdateWithColumnTimestampSmallerThanPkWithFlush[3](org.apache.cassandra.cql3.ViewComplexUpdatesTest):\tCaused an ERROR\r\n[junit-timeout] [localhost/127.0.0.1:43739] Timed out waiting for server response\r\n[junit-timeout] com.datastax.driver.core.exceptions.OperationTimedOutException: [localhost/127.0.0.1:43739] Timed out waiting for server response\r\n[junit-timeout] \tat com.datastax.driver.core.exceptions.OperationTimedOutException.copy(OperationTimedOutException.java:43)\r\n[junit-timeout] \tat com.datastax.driver.core.exceptions.OperationTimedOutException.copy(OperationTimedOutException.java:25)\r\n[junit-timeout] \tat com.datastax.driver.core.DriverThrowables.propagateCause(DriverThrowables.java:35)\r\n[junit-timeout] \tat com.datastax.driver.core.DefaultResultSetFuture.getUninterruptibly(DefaultResultSetFuture.java:293)\r\n[junit-timeout] \tat com.datastax.driver.core.AbstractSession.execute(AbstractSession.java:58)\r\n[junit-timeout] \tat com.datastax.driver.core.AbstractSession.execute(AbstractSession.java:45)\r\n[junit-timeout] \tat org.apache.cassandra.cql3.CQLTester.executeNet(CQLTester.java:972)\r\n[junit-timeout] \tat org.apache.cassandra.cql3.ViewComplexTester.createView(ViewComplexTester.java:83)\r\n[junit-timeout] \tat org.apache.cassandra.cql3.ViewComplexUpdatesTest.testUpdateWithColumnTimestampSmallerThanPk(ViewComplexUpdatesTest.java:224)\r\n[junit-timeout] \tat org.apache.cassandra.cql3.ViewComplexUpdatesTest.testUpdateWithColumnTimestampSmallerThanPkWithFlush(ViewComplexUpdatesTest.java:207)\r\n[junit-timeout] Caused by: com.datastax.driver.core.exceptions.OperationTimedOutException: [localhost/127.0.0.1:43739] Timed out waiting for server response\r\n[junit-timeout] \tat com.datastax.driver.core.RequestHandler$SpeculativeExecution.onTimeout(RequestHandler.java:979)\r\n[junit-timeout] \tat com.datastax.driver.core.Connection$ResponseHandler$1.run(Connection.java:1636)\r\n[junit-timeout] \tat com.datastax.shaded.netty.util.HashedWheelTimer$HashedWheelTimeout.expire(HashedWheelTimer.java:663)\r\n[junit-timeout] \tat com.datastax.shaded.netty.util.HashedWheelTimer$HashedWheelBucket.expireTimeouts(HashedWheelTimer.java:738)\r\n[junit-timeout] \tat com.datastax.shaded.netty.util.HashedWheelTimer$Worker.run(HashedWheelTimer.java:466)\r\n[junit-timeout] \tat com.datastax.shaded.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n[junit-timeout] \tat java.lang.Thread.run(Thread.java:748)\r\n{noformat}\r\n","from":"developer"},{"body":"I don't think we should put the duplicated cleanup logic back, in principle not reusing MV names should be enough.\r\n\r\nI mentioned the query timeout to indicate that the failure is not related to what we are addressing here, so I think the run can be considered successful. By the way, [here|https://app.circleci.com/pipelines/github/adelapena/cassandra/1101/workflows/63014008-3d52-4d07-a5aa-748fad389ec7] is the repeated run for j11, all green. I think we are ready to merge.","from":"developer"},{"body":"I added the trunk PR with CI. There's no rush so this will give you time to take a look at it. Let me know if you're still ok I merge. ","from":"developer"},{"body":"The PR for trunk looks good to me. Here are 200 runs of {{ViewComplex*Test}} in the PR for trunk:\r\n * [j8|https://app.circleci.com/pipelines/github/adelapena/cassandra/1114/workflows/2a6ea982-7109-4cf2-b5af-7f06188f2d4f]\r\n * [j11|https://app.circleci.com/pipelines/github/adelapena/cassandra/1114/workflows/12ab59cb-24d5-4795-a309-21b7eb6b187c]\r\n\r\nIf those succeed I think we'll be definitively ready to merge.","from":"developer"},{"body":"Yep wall of green :-)","from":"developer"}],"created":"2021-10-28T05:24:33.000+0000","description":"I have seen a number of times already the {{ViewComplexTest}} family timeout on test method teardown. This leaves a dirty env behind triggering the following test methods to fail on it. This ticket aims at hardening them.","issue_id":"13408819","key":"CASSANDRA-17070","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2021-11-04T07:26:20.000+0000","role":"fixed_distractor","summary":"ViewComplexTest hardening"} {"case_id":"13423259","cluster":"DISTRACTOR-CASSANDRA-17266","comments":[{"body":"Tagged this 4.0.x but versions should be examined.","created":"2022-01-18T13:03:03.080+0000"},{"body":"Hi [~taiga-elephant] I would like to work on this issue as it will be a good task for me to get to know the code.","created":"2022-01-26T04:19:23.688+0000"},{"body":"PR : https://github.com/apache/cassandra/pull/1427","created":"2022-01-26T05:27:50.436+0000"},{"body":"CI results for [4.0|https://app.circleci.com/pipelines/github/blerer/cassandra/259/workflows/01af39d8-847c-4c5a-8418-80e444cb7afb] and [trunk|https://app.circleci.com/pipelines/github/blerer/cassandra/258/workflows/03a953ed-8df6-41ef-8b2f-1816c4d242b1].","created":"2022-01-31T17:26:34.611+0000"},{"body":"[~yashLadha] One of the DTest [failed|https://app.circleci.com/pipelines/github/blerer/cassandra/258/workflows/03a953ed-8df6-41ef-8b2f-1816c4d242b1/jobs/2357/tests#failed-test-0] after looking into I discover that it was we were still returning the correct value for DESCRIBE statements. What changed is that we stop accepting {{defaultTimeToLive = 0}} in 4.0 (CASSANDRA-13426). We need to revert that behavior and accept {{defaultTimeToLive = 0}} to allow people to restore old backup but we should keep the new logic and modify the {{cqlsh_tests.test_cqlsh.TestCqlsh}} test. Could you do that? \r\n","created":"2022-02-01T10:21:57.187+0000"},{"body":"Sure [~blerer] let me take a look at this, sorry for the delayed response wasn't able to look into this.","created":"2022-02-03T14:24:06.412+0000"},{"body":"[~yashLadha] any progress on this?","created":"2022-03-16T07:36:05.325+0000"},{"body":"Assigning to myself due to Yash inactivity.","created":"2022-03-21T18:09:35.182+0000"},{"body":"4.0 [https://github.com/instaclustr/cassandra/tree/CASSANDRA-17266-4.0]\r\nmodified dtest - [https://github.com/smiklosovic/cassandra-dtest/tree/CASSANDRA-17266]\r\nbuild 4.0: [https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1538/]","created":"2022-03-22T15:23:15.895+0000"},{"body":"[~brandon.williams] [~blerer] would you mind to review? The build I got is quite nice. ","created":"2022-03-24T07:17:30.020+0000"},{"body":"maybe [~bereng] could take a look?","created":"2022-03-25T07:02:08.052+0000"},{"body":"Overall the patches look good to me, I just have a few minor comments/questions:\r\n* Why do we need those 2 USE queries in the {{ViewTimesTest}}? It seems to me that they could be avoided by providing the view name qualified with the keyspace. We do not seems to use them anymore in trunk. \r\n* In {{testCreateMvWithTTL}} we do not check the exception being thrown as it is done in {{testAlterMvWithNoZeroTTL}}. This has apparently been fixed in trunk.\r\n* It would be good to have a branch for trunk as the {{ViewTimesTest}} has apparently changed.\r\n* [~yashLadha] has contributed to the patch therefore it should be mentioned as one of the author of the patch. ","created":"2022-03-25T08:52:25.173+0000"},{"body":"Thanks [~blerer], all addressed. I ll provide trunk branch too.","created":"2022-03-25T09:37:57.623+0000"},{"body":"4.0: https://github.com/instaclustr/cassandra/tree/CASSANDRA-17266-4.0\r\n4.0 build: https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1538/\r\ntrunk: https://github.com/instaclustr/cassandra/tree/CASSANDRA-17266-trunk\r\ntrunk build: https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1545\r\n\r\n[~blerer] I am merging on your +1.","created":"2022-03-26T12:35:53.323+0000"},{"body":"Thanks [~smiklosovic] It looks good to me.","created":"2022-03-28T11:24:51.924+0000"}],"conversations":[{"body":"Materialized views do not allow default_time_to_live option in CQL (see CASSANDRA-12868).\r\n\r\nBut, if the MV was created without this option, DESCRIBE KEYSPACE / MATERIALIZED VIEW command generates CQL that includes it.\r\n\r\nE.g.\r\n{code:java}\r\nCREATE KEYSPACE test WITH replication = {'class': 'SimpleStrategy', 'replication_factor': '1'};\r\n\r\nUSE test;\r\n\r\nCREATE TABLE test_table(\r\n id text,\r\n date text,\r\n col1 text,\r\n col2 text,\r\n PRIMARY KEY(id,date)\r\n) WITH default_time_to_live = 60 AND CLUSTERING ORDER BY (date DESC);\r\n\r\nCREATE MATERIALIZED VIEW test_view AS\r\nSELECT id, date, col1\r\nFROM test_table\r\nWHERE id IS NOT NULL AND date IS NOT NULL\r\nPRIMARY KEY(id, date);{code}\r\nIt is OK. \r\n{code:java}\r\nDESCRIBE MATERIALIZED VIEW test_view; {code}\r\nreturns the same CQL + all default options:\r\n{code:java}\r\nCREATE MATERIALIZED VIEW test.test_view AS\r\n    SELECT id, date, col1\r\n    FROM test.test_table\r\n    WHERE id IS NOT NULL AND date IS NOT NULL\r\n    PRIMARY KEY (id, date)\r\n WITH CLUSTERING ORDER BY (date ASC)\r\n    AND additional_write_policy = '99p'\r\n    AND bloom_filter_fp_chance = 0.01\r\n    AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\r\n    AND cdc = false\r\n    AND comment = ''\r\n    AND compaction = {'class': 'org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy', 'max_threshold': '32', 'min_threshold': '4'}\r\n    AND compression = {'chunk_length_in_kb': '16', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}\r\n    AND crc_check_chance = 1.0\r\n    AND default_time_to_live = 0\r\n    AND extensions = {}\r\n    AND gc_grace_seconds = 864000\r\n    AND max_index_interval = 2048\r\n    AND memtable_flush_period_in_ms = 0\r\n    AND min_index_interval = 128\r\n    AND read_repair = 'BLOCKING'\r\n    AND speculative_retry = '99p';\r\n{code}\r\nNote the 'default_time_to_live = 0' clause! If this veiw is dropped, re-creating it using DESCRIBE output would fail with \r\n{noformat}\r\nCannot set default_time_to_live for a materialized view. Data in a materialized view always expire at the same time than the corresponding data in the parent table.{noformat}\r\n\r\n+Additional Information for newcomers:+\r\n\r\nThe code writting the table parameters is in {{TableParams.appendCqlTo}} and is called through {{TableMetadata.appendTableOptions}}. Those method will need to have a new parameter specifying if the call is for a table or a materialized view.\r\nSome unit test need to be adapted in {{DescribeStatementTest}} ","from":"reporter","subject":"DESCRIBE KEYSPACE / MATERIALIZED VIEW generates invalid CQL for views"},{"body":"Tagged this 4.0.x but versions should be examined.","from":"developer"},{"body":"Hi [~taiga-elephant] I would like to work on this issue as it will be a good task for me to get to know the code.","from":"developer"},{"body":"PR : https://github.com/apache/cassandra/pull/1427","from":"developer"},{"body":"CI results for [4.0|https://app.circleci.com/pipelines/github/blerer/cassandra/259/workflows/01af39d8-847c-4c5a-8418-80e444cb7afb] and [trunk|https://app.circleci.com/pipelines/github/blerer/cassandra/258/workflows/03a953ed-8df6-41ef-8b2f-1816c4d242b1].","from":"developer"},{"body":"[~yashLadha] One of the DTest [failed|https://app.circleci.com/pipelines/github/blerer/cassandra/258/workflows/03a953ed-8df6-41ef-8b2f-1816c4d242b1/jobs/2357/tests#failed-test-0] after looking into I discover that it was we were still returning the correct value for DESCRIBE statements. What changed is that we stop accepting {{defaultTimeToLive = 0}} in 4.0 (CASSANDRA-13426). We need to revert that behavior and accept {{defaultTimeToLive = 0}} to allow people to restore old backup but we should keep the new logic and modify the {{cqlsh_tests.test_cqlsh.TestCqlsh}} test. Could you do that? \r\n","from":"developer"},{"body":"Sure [~blerer] let me take a look at this, sorry for the delayed response wasn't able to look into this.","from":"developer"},{"body":"[~yashLadha] any progress on this?","from":"developer"},{"body":"Assigning to myself due to Yash inactivity.","from":"developer"},{"body":"4.0 [https://github.com/instaclustr/cassandra/tree/CASSANDRA-17266-4.0]\r\nmodified dtest - [https://github.com/smiklosovic/cassandra-dtest/tree/CASSANDRA-17266]\r\nbuild 4.0: [https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1538/]","from":"developer"},{"body":"[~brandon.williams] [~blerer] would you mind to review? The build I got is quite nice. ","from":"developer"},{"body":"maybe [~bereng] could take a look?","from":"developer"},{"body":"Overall the patches look good to me, I just have a few minor comments/questions:\r\n* Why do we need those 2 USE queries in the {{ViewTimesTest}}? It seems to me that they could be avoided by providing the view name qualified with the keyspace. We do not seems to use them anymore in trunk. \r\n* In {{testCreateMvWithTTL}} we do not check the exception being thrown as it is done in {{testAlterMvWithNoZeroTTL}}. This has apparently been fixed in trunk.\r\n* It would be good to have a branch for trunk as the {{ViewTimesTest}} has apparently changed.\r\n* [~yashLadha] has contributed to the patch therefore it should be mentioned as one of the author of the patch. ","from":"developer"},{"body":"Thanks [~blerer], all addressed. I ll provide trunk branch too.","from":"developer"},{"body":"4.0: https://github.com/instaclustr/cassandra/tree/CASSANDRA-17266-4.0\r\n4.0 build: https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1538/\r\ntrunk: https://github.com/instaclustr/cassandra/tree/CASSANDRA-17266-trunk\r\ntrunk build: https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1545\r\n\r\n[~blerer] I am merging on your +1.","from":"developer"},{"body":"Thanks [~smiklosovic] It looks good to me.","from":"developer"}],"created":"2022-01-18T11:58:16.000+0000","description":"Materialized views do not allow default_time_to_live option in CQL (see CASSANDRA-12868).\r\n\r\nBut, if the MV was created without this option, DESCRIBE KEYSPACE / MATERIALIZED VIEW command generates CQL that includes it.\r\n\r\nE.g.\r\n{code:java}\r\nCREATE KEYSPACE test WITH replication = {'class': 'SimpleStrategy', 'replication_factor': '1'};\r\n\r\nUSE test;\r\n\r\nCREATE TABLE test_table(\r\n id text,\r\n date text,\r\n col1 text,\r\n col2 text,\r\n PRIMARY KEY(id,date)\r\n) WITH default_time_to_live = 60 AND CLUSTERING ORDER BY (date DESC);\r\n\r\nCREATE MATERIALIZED VIEW test_view AS\r\nSELECT id, date, col1\r\nFROM test_table\r\nWHERE id IS NOT NULL AND date IS NOT NULL\r\nPRIMARY KEY(id, date);{code}\r\nIt is OK. \r\n{code:java}\r\nDESCRIBE MATERIALIZED VIEW test_view; {code}\r\nreturns the same CQL + all default options:\r\n{code:java}\r\nCREATE MATERIALIZED VIEW test.test_view AS\r\n    SELECT id, date, col1\r\n    FROM test.test_table\r\n    WHERE id IS NOT NULL AND date IS NOT NULL\r\n    PRIMARY KEY (id, date)\r\n WITH CLUSTERING ORDER BY (date ASC)\r\n    AND additional_write_policy = '99p'\r\n    AND bloom_filter_fp_chance = 0.01\r\n    AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\r\n    AND cdc = false\r\n    AND comment = ''\r\n    AND compaction = {'class': 'org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy', 'max_threshold': '32', 'min_threshold': '4'}\r\n    AND compression = {'chunk_length_in_kb': '16', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}\r\n    AND crc_check_chance = 1.0\r\n    AND default_time_to_live = 0\r\n    AND extensions = {}\r\n    AND gc_grace_seconds = 864000\r\n    AND max_index_interval = 2048\r\n    AND memtable_flush_period_in_ms = 0\r\n    AND min_index_interval = 128\r\n    AND read_repair = 'BLOCKING'\r\n    AND speculative_retry = '99p';\r\n{code}\r\nNote the 'default_time_to_live = 0' clause! If this veiw is dropped, re-creating it using DESCRIBE output would fail with \r\n{noformat}\r\nCannot set default_time_to_live for a materialized view. Data in a materialized view always expire at the same time than the corresponding data in the parent table.{noformat}\r\n\r\n+Additional Information for newcomers:+\r\n\r\nThe code writting the table parameters is in {{TableParams.appendCqlTo}} and is called through {{TableMetadata.appendTableOptions}}. Those method will need to have a new parameter specifying if the call is for a table or a materialized view.\r\nSome unit test need to be adapted in {{DescribeStatementTest}} ","issue_id":"13423259","key":"CASSANDRA-17266","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2022-03-28T12:31:16.000+0000","role":"fixed_distractor","summary":"DESCRIBE KEYSPACE / MATERIALIZED VIEW generates invalid CQL for views"} {"case_id":"13423390","cluster":"DISTRACTOR-CASSANDRA-17267","comments":[{"body":"The snapshot true size is calculated by [Directories.getTrueAllocatedSizeIn|https://github.com/apache/cassandra/blob/cassandra-3.11/src/java/org/apache/cassandra/db/Directories.java#L960].\r\n\r\nThis method creates a [SSTableSizeSummer|https://github.com/apache/cassandra/blob/cassandra-3.11/src/java/org/apache/cassandra/db/Directories.java#L1054] using the snapshot folder as the list of files to be iterated/counted and the list of live sstables as the list of files to be skipped (toSkip set).\r\n\r\nThe [isAcceptable|https://github.com/apache/cassandra/blob/cassandra-3.11/src/java/org/apache/cassandra/db/Directories.java#L1064] method decides whether the snapshot file size must be counted by checking if it's an sstable component and if it's not present on the \"toSkip\" set.\r\n\r\nHowever the snapshot files will never be present in the \"toSkip\" set, causing the snapshot file sizes to always be accounted - whether or not a \"corresponding\" live sstable is found.\r\n\r\nI believe the original implementer's intent was to verify that the \"corresponding\" sstable file is present in the \"toSkip\" set, but it doesn't reconstruct the original sstable file from the snapshot file before checking it's present on the set.\r\n\r\nI created a [PR|https://github.com/apache/cassandra/pull/1408] with a reproduction and preliminary fix.\r\n\r\nThe reproduction can be found [on this test|https://github.com/apache/cassandra/pull/1408/files#diff-ef5be0b69d0440b76021282c4b24bad69770ef9419be260df2169f49921db377R346].\r\n\r\n[The fix|https://github.com/apache/cassandra/pull/1408/files#diff-bb20d0c655884c2211213190ae4787ace619cdff4c0235f147db7dfbf1e7a869R1067] only counts the snapshot file size if the file is an sstable component *AND* if a corresponding live sstable component can *not* be found on \"snapshot_dir/../../file_name\" (since the snapshot file is found on /snapshots//file).\r\n\r\nThe same snapshot of the ticket description is displayed as following after the fix:\r\n{noformat}\r\n$ nodetool listsnapshots\r\nSnapshot Details:\r\nSnapshot name Keyspace name Column family name True size Size on disk\r\ntest ks1 tbl1 0 bytes 5.69 KiB\r\n\r\nTotal TrueDiskSpaceUsed: 0 bytes\r\n{noformat}","created":"2022-01-19T02:06:06.202+0000"},{"body":"[~brandon.williams] I found this while working on CASSANDRA-16843. Can you take a look since it's somewhat related?\r\n\r\nIf the approach looks good I will work on more tests (secondary index) and fix existing tests, as well as port to other branches - this might affect 3.0 too.","created":"2022-01-19T02:09:49.249+0000"},{"body":"Nice catch. This looks like a good approach to me, and I suspect 3.0 has a problem as well.","created":"2022-01-19T13:08:02.927+0000"},{"body":"Curiously this did not reproduce on 3.0, so I used the same approach of comparing the names of the snapshot files with the files present in the live set to skip accounting live sstables during snapshot true size calculation.\r\n\r\nccm repro after fix:\r\n{noformat}\r\n% ccm node1 nodetool -- snapshot -t test test_ks\r\n\r\nRequested creating snapshot(s) for [test_ks] with snapshot name [test]\r\nSnapshot directory: test\r\n\r\n% ccm node1 nodetool tablestats test_ks.tbl | grep -i snapshot\r\n Space used by snapshots (total): 0\r\n\r\n% ccm node1 nodetool listsnapshots\r\n\r\nSnapshot Details:\r\nSnapshot name Keyspace name Column family name True size Size on disk\r\ntest test_ks tbl 0 bytes 5.74 KB\r\n\r\nTotal TrueDiskSpaceUsed: 0 bytes\r\n\r\n% ccm node1 nodetool compact test_ks tbl\r\n\r\n% ccm node1 nodetool tablestats test_ks.tbl | grep -i snapshot\r\n Space used by snapshots (total): 5044\r\n\r\n% ccm node1 nodetool listsnapshots\r\n\r\nSnapshot Details:\r\nSnapshot name Keyspace name Column family name True size Size on disk\r\ntest test_ks tbl 4.93 KB 5.74 KB\r\n\r\nTotal TrueDiskSpaceUsed: 4.93 KB\r\n{noformat}\r\n\r\nI will use the new approach of using only the directory structure to decide whether a snapshot file is present in the live set when decoupling snapshot size computation from {{ColumnFamilyStore}} on CASSANDRA-16843.\r\n\r\nWhile working on this I noticed that secondary indexes are not included in the computation of the true size so I created CASSANDRA-17357 to address this separately.\r\n\r\n3.11+ patches and CI below:\r\n\r\n|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...pauloricardomg:CASSANDRA-17267-3.11]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1414/]|\r\n|[4.0|https://github.com/apache/cassandra/compare/cassandra-4.0...pauloricardomg:CASSANDRA-17267-4.0]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1415/]|\r\n|[trunk|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:CASSANDRA-17267-trunk]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1422/]|","created":"2022-02-07T23:53:32.157+0000"},{"body":"Thanks [~paulo]. The patch looks good to me but the tests needs to be re-run.","created":"2022-02-11T09:17:37.360+0000"},{"body":"In the [previous test run|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1422/] {{org.apache.cassandra.index.sasi.SASIIndexTest.testSASIComponentsAddedToSnapshot}} was getting stuck when running within the suite (worked when executed individually).\r\n\r\nI tracked down the reason to the {{ReadExecutionController}} not being closed properly on other tests, causing operations to block indefinitely on the {{{}OpOrder{}}}. Fixed [on this commit|https://github.com/apache/cassandra/commit/77f688e75ff403875755f34dc31ab75401bcaa3d] on all branches.\r\n\r\nI created CASSANDRA-17400 to add a checker to verify resources are being properly closed to avoid stuck tests in the future.\r\n\r\nResubmitted CI:\r\n|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...pauloricardomg:CASSANDRA-17267-3.11]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1440/]|\r\n|[4.0|https://github.com/apache/cassandra/compare/cassandra-4.0...pauloricardomg:CASSANDRA-17267-4.0]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1441/]|\r\n|[trunk|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:CASSANDRA-17267-trunk]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1442/]|","created":"2022-02-22T13:52:58.854+0000"},{"body":"+1","created":"2022-02-28T09:33:21.805+0000"},{"body":"[~paulo] would you mind if I take over as you are swamped these days?","created":"2022-03-09T11:03:42.978+0000"},{"body":"Committed to cassandra-3.11 branch and merged up to {{trunk}} as {{{}95a622305722889c321204c4bca68a3517a29aab{}}}.","created":"2022-03-14T20:26:59.601+0000"},{"body":"Doubt any of the \"only failed once\" below are related to this ticket.\r\n\r\n \r\n\r\n[CI Results]\r\nBranch: trunk, build number: 1010\r\njenkins url: [https://ci-cassandra.apache.org/job/Cassandra-trunk/1010/]\r\nJIRA: CASSANDRA-17267\r\ncommit url: [https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=95a622305722889c321204c4bca68a3517a29aab]\r\naffected paths:\r\n * test/unit/org/apache/cassandra/db/ColumnFamilyStoreTest.java\r\n * test/unit/org/apache/cassandra/index/sasi/SASIIndexTest.java\r\n * CHANGES.txt\r\n * src/java/org/apache/cassandra/db/Directories.java\r\n\r\nBuild Result: UNSTABLE\r\nPassing Tests: 44358\r\nFailing Tests: 17\r\n||Test|Failures|JIRA|\r\n|dtest-upgrade.upgrade_tests.cql_tests.TestCQLNodes2RF1_Upgrade_indev_4_0_x_To_indev_trunk.test_static_cf|4 of 62|CASSANDRA-17309?|\r\n|dtest.write_failures_test.TestMultiDCWriteFailures.test_oversized_mutation|8 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|org.apache.cassandra.db.commitlog.BatchCommitLogTest.testOutOfOrderLogDiscard[3]|1 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|junit.framework.TestSuite.org.apache.cassandra.audit.BinAuditLoggerTest|1 of 62|[Multiple JIRAs found|https://issues.apache.org/jira/issues/?jql=project%20%3D%20CASSANDRA%20and%20resolution%20%3D%20unresolved%20and%20summary%20~%20%22*audit*%22]|\r\n|org.apache.cassandra.distributed.test.CASTest.testSuccessfulWriteDuringRangeMovementFollowedByConflicting|4 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|dtest-offheap.write_failures_test.TestMultiDCWriteFailures.test_oversized_mutation|8 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|dtest-upgrade.upgrade_tests.cql_tests.TestCQLNodes2RF1_Upgrade_indev_4_0_x_To_indev_trunk.test_noncomposite_static_cf|1 of 62|CASSANDRA-17309?|\r\n|org.apache.cassandra.distributed.test.CasCriticalSectionTest.criticalSectionTest|8 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV3Upgrade_AllVersions_RandomPartitioner_EndsAt_3_11_X_HEAD.test_rolling_upgrade_with_internode_ssl|27 of 62|CASSANDRA-17305?|\r\n|org.apache.cassandra.db.lifecycle.LogTransactionTest.testGetTemporaryFilesSafeAfterObsoletion-cdc|11 of 62|CASSANDRA-17286?|\r\n|org.apache.cassandra.cql3.validation.operations.CompactStorageTest.testCfmCounterCQL|1 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV3Upgrade_AllVersions_RandomPartitioner_EndsAt_3_11_X_HEAD.test_parallel_upgrade|2 of 62|CASSANDRA-17305?|\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV3Upgrade_AllVersions_EndsAt_3_11_X.test_rolling_upgrade|28 of 62|CASSANDRA-17305?|\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV3Upgrade_AllVersions_RandomPartitioner_EndsAt_3_11_X_HEAD.test_rolling_upgrade|28 of 62|CASSANDRA-17305?|\r\n|org.apache.cassandra.cql3.KeywordTest.test[keyword CLUSTERING isReserved false]|2 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV3Upgrade_AllVersions_EndsAt_3_11_X.test_rolling_upgrade_with_internode_ssl|27 of 62|CASSANDRA-17305?|\r\n|dtest-novnode.write_failures_test.TestMultiDCWriteFailures.test_oversized_mutation|8 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|","created":"2022-03-15T18:47:41.792+0000"},{"body":"[~jmckenzie] I've checked and don't think these failures are related to this change. Did you trigger a re-run?","created":"2022-03-15T18:58:37.427+0000"},{"body":"{quote}Did you trigger a re-run?\r\n{quote}\r\nI did not. I'm updating tickets with the results of CI based on the commit, however CI is in a pretty rough spot right now so this is likely minimally valuable outside checking tests that have only failed once.","created":"2022-03-17T15:52:01.135+0000"}],"conversations":[{"body":"As far as I understand, the snapshot \"size on disk\" is the total size of the snapshot, while the \"true size\" is the (size_on_disk - size_of_live_sstables).\r\n\r\nI created a snapshot on a 3.11 node without traffic and I expected the \"true size\" to be 0KB since the original sstables were still present, but this didn't seem to be the case:\r\n{noformat}\r\n$ nodetool listsnapshots\r\nSnapshot Details:\r\nSnapshot name Keyspace name Column family name True size Size on disk\r\ntest ks1 tbl1 4.86 KiB 5.69 KiB\r\n\r\nTotal TrueDiskSpaceUsed: 4.86 KiB\r\n{noformat}","from":"reporter","subject":"Snapshot true size is miscalculated"},{"body":"The snapshot true size is calculated by [Directories.getTrueAllocatedSizeIn|https://github.com/apache/cassandra/blob/cassandra-3.11/src/java/org/apache/cassandra/db/Directories.java#L960].\r\n\r\nThis method creates a [SSTableSizeSummer|https://github.com/apache/cassandra/blob/cassandra-3.11/src/java/org/apache/cassandra/db/Directories.java#L1054] using the snapshot folder as the list of files to be iterated/counted and the list of live sstables as the list of files to be skipped (toSkip set).\r\n\r\nThe [isAcceptable|https://github.com/apache/cassandra/blob/cassandra-3.11/src/java/org/apache/cassandra/db/Directories.java#L1064] method decides whether the snapshot file size must be counted by checking if it's an sstable component and if it's not present on the \"toSkip\" set.\r\n\r\nHowever the snapshot files will never be present in the \"toSkip\" set, causing the snapshot file sizes to always be accounted - whether or not a \"corresponding\" live sstable is found.\r\n\r\nI believe the original implementer's intent was to verify that the \"corresponding\" sstable file is present in the \"toSkip\" set, but it doesn't reconstruct the original sstable file from the snapshot file before checking it's present on the set.\r\n\r\nI created a [PR|https://github.com/apache/cassandra/pull/1408] with a reproduction and preliminary fix.\r\n\r\nThe reproduction can be found [on this test|https://github.com/apache/cassandra/pull/1408/files#diff-ef5be0b69d0440b76021282c4b24bad69770ef9419be260df2169f49921db377R346].\r\n\r\n[The fix|https://github.com/apache/cassandra/pull/1408/files#diff-bb20d0c655884c2211213190ae4787ace619cdff4c0235f147db7dfbf1e7a869R1067] only counts the snapshot file size if the file is an sstable component *AND* if a corresponding live sstable component can *not* be found on \"snapshot_dir/../../file_name\" (since the snapshot file is found on /snapshots//file).\r\n\r\nThe same snapshot of the ticket description is displayed as following after the fix:\r\n{noformat}\r\n$ nodetool listsnapshots\r\nSnapshot Details:\r\nSnapshot name Keyspace name Column family name True size Size on disk\r\ntest ks1 tbl1 0 bytes 5.69 KiB\r\n\r\nTotal TrueDiskSpaceUsed: 0 bytes\r\n{noformat}","from":"developer"},{"body":"[~brandon.williams] I found this while working on CASSANDRA-16843. Can you take a look since it's somewhat related?\r\n\r\nIf the approach looks good I will work on more tests (secondary index) and fix existing tests, as well as port to other branches - this might affect 3.0 too.","from":"developer"},{"body":"Nice catch. This looks like a good approach to me, and I suspect 3.0 has a problem as well.","from":"developer"},{"body":"Curiously this did not reproduce on 3.0, so I used the same approach of comparing the names of the snapshot files with the files present in the live set to skip accounting live sstables during snapshot true size calculation.\r\n\r\nccm repro after fix:\r\n{noformat}\r\n% ccm node1 nodetool -- snapshot -t test test_ks\r\n\r\nRequested creating snapshot(s) for [test_ks] with snapshot name [test]\r\nSnapshot directory: test\r\n\r\n% ccm node1 nodetool tablestats test_ks.tbl | grep -i snapshot\r\n Space used by snapshots (total): 0\r\n\r\n% ccm node1 nodetool listsnapshots\r\n\r\nSnapshot Details:\r\nSnapshot name Keyspace name Column family name True size Size on disk\r\ntest test_ks tbl 0 bytes 5.74 KB\r\n\r\nTotal TrueDiskSpaceUsed: 0 bytes\r\n\r\n% ccm node1 nodetool compact test_ks tbl\r\n\r\n% ccm node1 nodetool tablestats test_ks.tbl | grep -i snapshot\r\n Space used by snapshots (total): 5044\r\n\r\n% ccm node1 nodetool listsnapshots\r\n\r\nSnapshot Details:\r\nSnapshot name Keyspace name Column family name True size Size on disk\r\ntest test_ks tbl 4.93 KB 5.74 KB\r\n\r\nTotal TrueDiskSpaceUsed: 4.93 KB\r\n{noformat}\r\n\r\nI will use the new approach of using only the directory structure to decide whether a snapshot file is present in the live set when decoupling snapshot size computation from {{ColumnFamilyStore}} on CASSANDRA-16843.\r\n\r\nWhile working on this I noticed that secondary indexes are not included in the computation of the true size so I created CASSANDRA-17357 to address this separately.\r\n\r\n3.11+ patches and CI below:\r\n\r\n|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...pauloricardomg:CASSANDRA-17267-3.11]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1414/]|\r\n|[4.0|https://github.com/apache/cassandra/compare/cassandra-4.0...pauloricardomg:CASSANDRA-17267-4.0]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1415/]|\r\n|[trunk|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:CASSANDRA-17267-trunk]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1422/]|","from":"developer"},{"body":"Thanks [~paulo]. The patch looks good to me but the tests needs to be re-run.","from":"developer"},{"body":"In the [previous test run|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1422/] {{org.apache.cassandra.index.sasi.SASIIndexTest.testSASIComponentsAddedToSnapshot}} was getting stuck when running within the suite (worked when executed individually).\r\n\r\nI tracked down the reason to the {{ReadExecutionController}} not being closed properly on other tests, causing operations to block indefinitely on the {{{}OpOrder{}}}. Fixed [on this commit|https://github.com/apache/cassandra/commit/77f688e75ff403875755f34dc31ab75401bcaa3d] on all branches.\r\n\r\nI created CASSANDRA-17400 to add a checker to verify resources are being properly closed to avoid stuck tests in the future.\r\n\r\nResubmitted CI:\r\n|[3.11|https://github.com/apache/cassandra/compare/cassandra-3.11...pauloricardomg:CASSANDRA-17267-3.11]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1440/]|\r\n|[4.0|https://github.com/apache/cassandra/compare/cassandra-4.0...pauloricardomg:CASSANDRA-17267-4.0]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1441/]|\r\n|[trunk|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:CASSANDRA-17267-trunk]|[tests|https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1442/]|","from":"developer"},{"body":"+1","from":"developer"},{"body":"[~paulo] would you mind if I take over as you are swamped these days?","from":"developer"},{"body":"Committed to cassandra-3.11 branch and merged up to {{trunk}} as {{{}95a622305722889c321204c4bca68a3517a29aab{}}}.","from":"developer"},{"body":"Doubt any of the \"only failed once\" below are related to this ticket.\r\n\r\n \r\n\r\n[CI Results]\r\nBranch: trunk, build number: 1010\r\njenkins url: [https://ci-cassandra.apache.org/job/Cassandra-trunk/1010/]\r\nJIRA: CASSANDRA-17267\r\ncommit url: [https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=95a622305722889c321204c4bca68a3517a29aab]\r\naffected paths:\r\n * test/unit/org/apache/cassandra/db/ColumnFamilyStoreTest.java\r\n * test/unit/org/apache/cassandra/index/sasi/SASIIndexTest.java\r\n * CHANGES.txt\r\n * src/java/org/apache/cassandra/db/Directories.java\r\n\r\nBuild Result: UNSTABLE\r\nPassing Tests: 44358\r\nFailing Tests: 17\r\n||Test|Failures|JIRA|\r\n|dtest-upgrade.upgrade_tests.cql_tests.TestCQLNodes2RF1_Upgrade_indev_4_0_x_To_indev_trunk.test_static_cf|4 of 62|CASSANDRA-17309?|\r\n|dtest.write_failures_test.TestMultiDCWriteFailures.test_oversized_mutation|8 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|org.apache.cassandra.db.commitlog.BatchCommitLogTest.testOutOfOrderLogDiscard[3]|1 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|junit.framework.TestSuite.org.apache.cassandra.audit.BinAuditLoggerTest|1 of 62|[Multiple JIRAs found|https://issues.apache.org/jira/issues/?jql=project%20%3D%20CASSANDRA%20and%20resolution%20%3D%20unresolved%20and%20summary%20~%20%22*audit*%22]|\r\n|org.apache.cassandra.distributed.test.CASTest.testSuccessfulWriteDuringRangeMovementFollowedByConflicting|4 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|dtest-offheap.write_failures_test.TestMultiDCWriteFailures.test_oversized_mutation|8 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|dtest-upgrade.upgrade_tests.cql_tests.TestCQLNodes2RF1_Upgrade_indev_4_0_x_To_indev_trunk.test_noncomposite_static_cf|1 of 62|CASSANDRA-17309?|\r\n|org.apache.cassandra.distributed.test.CasCriticalSectionTest.criticalSectionTest|8 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV3Upgrade_AllVersions_RandomPartitioner_EndsAt_3_11_X_HEAD.test_rolling_upgrade_with_internode_ssl|27 of 62|CASSANDRA-17305?|\r\n|org.apache.cassandra.db.lifecycle.LogTransactionTest.testGetTemporaryFilesSafeAfterObsoletion-cdc|11 of 62|CASSANDRA-17286?|\r\n|org.apache.cassandra.cql3.validation.operations.CompactStorageTest.testCfmCounterCQL|1 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV3Upgrade_AllVersions_RandomPartitioner_EndsAt_3_11_X_HEAD.test_parallel_upgrade|2 of 62|CASSANDRA-17305?|\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV3Upgrade_AllVersions_EndsAt_3_11_X.test_rolling_upgrade|28 of 62|CASSANDRA-17305?|\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV3Upgrade_AllVersions_RandomPartitioner_EndsAt_3_11_X_HEAD.test_rolling_upgrade|28 of 62|CASSANDRA-17305?|\r\n|org.apache.cassandra.cql3.KeywordTest.test[keyword CLUSTERING isReserved false]|2 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV3Upgrade_AllVersions_EndsAt_3_11_X.test_rolling_upgrade_with_internode_ssl|27 of 62|CASSANDRA-17305?|\r\n|dtest-novnode.write_failures_test.TestMultiDCWriteFailures.test_oversized_mutation|8 of 62|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]|","from":"developer"},{"body":"[~jmckenzie] I've checked and don't think these failures are related to this change. Did you trigger a re-run?","from":"developer"},{"body":"{quote}Did you trigger a re-run?\r\n{quote}\r\nI did not. I'm updating tickets with the results of CI based on the commit, however CI is in a pretty rough spot right now so this is likely minimally valuable outside checking tests that have only failed once.","from":"developer"}],"created":"2022-01-19T01:38:49.000+0000","description":"As far as I understand, the snapshot \"size on disk\" is the total size of the snapshot, while the \"true size\" is the (size_on_disk - size_of_live_sstables).\r\n\r\nI created a snapshot on a 3.11 node without traffic and I expected the \"true size\" to be 0KB since the original sstables were still present, but this didn't seem to be the case:\r\n{noformat}\r\n$ nodetool listsnapshots\r\nSnapshot Details:\r\nSnapshot name Keyspace name Column family name True size Size on disk\r\ntest ks1 tbl1 4.86 KiB 5.69 KiB\r\n\r\nTotal TrueDiskSpaceUsed: 4.86 KiB\r\n{noformat}","issue_id":"13423390","key":"CASSANDRA-17267","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2022-03-14T20:26:59.000+0000","role":"fixed_distractor","summary":"Snapshot true size is miscalculated"} {"case_id":"13424738","cluster":"DISTRACTOR-CASSANDRA-17291","comments":[{"body":"[~jmckenzie] , I was wondering what is the reason to leave in triage some of the test failures tickets like this one? Did you think it needs further testing and it might have been random failure? \r\n\r\nI am opening this one as I see it consistently failing in CircleCI. Also, my understanding is that we use the offline token allocator for the Python DTests so it is good to verify there is no real problem there.\r\n\r\nCC [~mck] ","created":"2022-03-28T13:24:02.449+0000"},{"body":"Just missed hitting open when creating as build lead. Not intentional.","created":"2022-03-28T13:35:24.173+0000"},{"body":"100 runs on JDK8 w/this passing in circle; going to try 100 on JDK11 then maybe up the repeat count.","created":"2022-05-16T16:04:19.593+0000"},{"body":"I see it almost all the time with MIDRES in CircleCI ","created":"2022-05-16T16:30:12.202+0000"},{"body":"Yeah, this is somewhat odd. using `generate.sh` to multiplex it it doesn't look like it's showing up; may have something to do with timing and other tests parallelizing and sharing state and stomping each other.\r\n\r\n \r\n\r\nI assume generate.sh is going to keep provisioning at the effectively \"LOWRES\" footprint on hardware alloc right?","created":"2022-05-16T17:08:22.457+0000"},{"body":"generate.sh accepts flags for medium, high, etc. See [here|https://github.com/apache/cassandra/tree/trunk/.circleci#setting-environment-variables] i.e.","created":"2022-05-17T08:05:05.253+0000"},{"body":"bq. generate.sh accepts flags for medium, high, etc. See here i.e.\r\nIndeed. And when you don't run it with a flag it stays at the default, which in this case is LOWRES.","created":"2022-05-17T15:36:54.057+0000"},{"body":"Yes, and I checked currently LOWRES and MIDRES use the same resources for unit tests in particular.\r\n\r\nI ran this one  500 times, the whole class in  [here|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1622/workflows/b69d6e45-b012-4fc7-a839-7320b86a319f/jobs/11106/steps] and weirdly it doesn't reproduce. But then I see it all the time in my MIDRES runs.\r\n\r\nI will experiment a bit more. \r\n\r\n ","created":"2022-05-17T16:55:50.439+0000"},{"body":"Ok. So this was annoying but I think I got to the bottom of it.\r\n\r\n[PR|https://github.com/apache/cassandra/pull/1638]\r\n[JDK8 CI|https://app.circleci.com/pipelines/github/josh-mckenzie/cassandra/232/workflows/6c742dd9-e1b3-4056-aac8-9f0b95ae4e84]\r\n[JDK11 CI|https://app.circleci.com/pipelines/github/josh-mckenzie/cassandra/232/workflows/5fcce28f-a258-485f-906e-45957c2b3b08]\r\n\r\nThe offending bit of code:\r\n{code:java}\r\n for (int numTokens = 1; numTokens <= 16 ; ++numTokens)\r\n {\r\n for (int rf = 1; rf <=5; ++rf)\r\n {\r\n int nodeCount = 32;\r\n for (int racks = 1; racks <= 10; ++racks)\r\n {\r\n int[] nodeToRack = makeRackCountArray(nodeCount, racks);\r\n for (IPartitioner partitioner : new IPartitioner[] { Murmur3Partitioner.instance, RandomPartitioner.instance })\r\n {\r\n{code}\r\n{{void testTokenGenerations()}} was timing out due to this pretty beastly nesting of different permutations it runs combined with it logging output, etc. Effectively we have to run 1600 tests in 900 seconds sans whatever the other tests were consuming for time, and with all the logging about tokens and iteration through it looks like we tipped over an edge.\r\n\r\nI pared the test down to the following (about a third of the combinations with what I think is *logically* comparably useful coverage):\r\n{code:java}\r\n private final int[] racks = { 1, 2, 3, 5, 6, 9, 10 };\r\n private final int[] rfs = { 1, 2, 3, 5 };\r\n private final int[] tokens = { 1, 2, 3, 5, 6, 9, 10, 13, 15, 16 };\r\n{code}\r\nI toyed with getting rid of the logging inside the test class but the lion's share of what's spamming is in a variety of other classes. I also looked into disabling logging in the {{SystemOutputImpl}} in {{OfflineTokenAllocatorTestUtils}} but that ended up being more pain than it was worth.\r\n\r\nRather than taking 15+ minutes and timing out my laptop it takes about 20 seconds; should be ok in our CI env. Also split out this test method to its own file entirely so it doesn't stomp on the other offline token allocation tests and throw a red herring of timeout like this did (parallelization and method timeouts within a class are... not fun).\r\n ","created":"2022-05-19T19:29:21.620+0000"},{"body":"change LGTM but the new test class doesn't end in Test so it doesn't run in CI; once that is fixed I am +1","created":"2022-05-19T21:12:33.643+0000"},{"body":"bq. the new test class doesn't end in Test\r\nBlargh! It's always something. :)\r\n\r\nFigured I'd multiplex both the classes for 100 runs as well just to double check they're solid; I'll do that, fix the file name, and assuming all green, merge in.\r\n\r\nThanks [~dcapwell]!","created":"2022-05-20T14:55:55.077+0000"},{"body":"Multiplex on both tests looked clean; merged up.","created":"2022-05-23T17:04:17.002+0000"},{"body":"[CI Results]\r\nBranch: 4.1, build number: 35\r\n butler url: https://butler.cassandra.apache.org/#/ci/upstream/compare/Cassandra-4.1/Cassandra-4.1\r\n jenkins url: https://ci-cassandra.apache.org/job/Cassandra-4.1/35/\r\n JIRA: CASSANDRA-17291\r\n commit url: https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=951aff25a1183f41fd146d674136399f3f25b3f0\r\n affected paths:\r\n* test/unit/org/apache/cassandra/dht/tokenallocator/OfflineTokenAllocatorGenerationsTest.java\r\n* test/unit/org/apache/cassandra/dht/tokenallocator/OfflineTokenAllocatorTestUtils.java\r\n* test/unit/org/apache/cassandra/dht/tokenallocator/OfflineTokenAllocatorTest.java\r\n\r\n Build Result: UNSTABLE\r\n Passing Tests: 47094\r\n Failing Tests: 27\r\n\r\n||Test|Failures|JIRA||\r\n|org.apache.cassandra.cql3.validation.operations.CompactStorageTest.testCounterAndColumnSelection|1 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.distributed.upgrade.MixedModeAvailabilityV30Test.testAvailability|5 of 33|[CASSANDRA-17307?|https://issues.apache.org/jira/browse/CASSANDRA-17307]|\r\n|org.apache.cassandra.distributed.test.SchemaTest.readRepairWithCompaction|4 of 33|[CASSANDRA-17641?|https://issues.apache.org/jira/browse/CASSANDRA-17641]|\r\n|org.apache.cassandra.distributed.test.MessageForwardingTest.mutationsForwardedToAllReplicasTest|1 of 33|[CASSANDRA-17583?|https://issues.apache.org/jira/browse/CASSANDRA-17583]|\r\n|org.apache.cassandra.cql3.validation.operations.CompactStorageTest.testAlterWithCompactNonStaticFormat|2 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.distributed.test.ring.BootstrapTest.readWriteDuringBootstrapTest|3 of 33|[Multiple JIRAs found|https://issues.apache.org/jira/issues/?jql=project%20%3D%20CASSANDRA%20and%20resolution%20%3D%20unresolved%20and%20summary%20~%20\"*BootstrapTest*\"]\r\n|org.apache.cassandra.distributed.upgrade.CompactStorageUpgradeTest.compactStorageImplicitNullInClusteringTest|5 of 33|[CASSANDRA-17213?|https://issues.apache.org/jira/browse/CASSANDRA-17213]|\r\n|junit.framework.TestSuite.org.apache.cassandra.distributed.test.CASMultiDCTest|2 of 33|[Multiple JIRAs found|https://issues.apache.org/jira/issues/?jql=project%20%3D%20CASSANDRA%20and%20resolution%20%3D%20unresolved%20and%20summary%20~%20\"*test*\"]\r\n|dtest-upgrade.upgrade_tests.drop_compact_storage_upgrade_test.TestDropCompactStorage.test_drop_compact_storage_mixed_cluster|4 of 33|[CASSANDRA-17634?|https://issues.apache.org/jira/browse/CASSANDRA-17634]|\r\n|org.apache.cassandra.tools.TopPartitionsTest.testServiceTopPartitionsSingleTable-cdc|5 of 33|[CASSANDRA-17649?|https://issues.apache.org/jira/browse/CASSANDRA-17649]|\r\n|org.apache.cassandra.cql3.validation.entities.SecondaryIndexTest.testIndexOnRegularColumnInsertExpiringColumnWithFlush|3 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.cql3.validation.operations.CompactStorageTest.testStaticCompactWithCounters|2 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.distributed.upgrade.BatchUpgradeTest.batchTest|6 of 33|[CASSANDRA-17651?|https://issues.apache.org/jira/browse/CASSANDRA-17651]|\r\n|org.apache.cassandra.cql3.ViewFilteringClustering1Test.testClusteringKeySliceRestrictions[3]|6 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV4Upgrade_AllVersions_EndsAt_Trunk_HEAD.test_parallel_upgrade|1 of 33|[CASSANDRA-17296?|https://issues.apache.org/jira/browse/CASSANDRA-17296]|\r\n|org.apache.cassandra.cql3.validation.operations.SelectTest.filteringWithOrderClause|3 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.cql3.validation.entities.SecondaryIndexTest.testIndexOnNonFrozenCollectionOfFrozenUDT|1 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.distributed.upgrade.CompactStorageUpgradeTest.compactStoragePagingTest|3 of 33|[CASSANDRA-17213?|https://issues.apache.org/jira/browse/CASSANDRA-17213]|\r\n|org.apache.cassandra.db.VerifyTest.testMutateRepair-cdc|1 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestUpgrade_current_4_0_x_To_indev_4_1_x.test_parallel_upgrade_with_internode_ssl|1 of 33|[Multiple JIRAs found|https://issues.apache.org/jira/issues/?jql=project%20%3D%20CASSANDRA%20and%20resolution%20%3D%20unresolved%20and%20summary%20~%20\"*TestUpgrade*\"]\r\n|org.apache.cassandra.distributed.test.CASTest.testIncompleteWriteFollowedBySuccessfulWriteWithStaleRingDuringRangeMovementFollowedByRead|1 of 33|[CASSANDRA-17461?|https://issues.apache.org/jira/browse/CASSANDRA-17461]|\r\n|org.apache.cassandra.db.virtual.GossipInfoTableTest.testSelectAllWithStateTransitions|1 of 33|[CASSANDRA-17584?|https://issues.apache.org/jira/browse/CASSANDRA-17584]|\r\n|org.apache.cassandra.distributed.upgrade.CompactStorageUpgradeTest.compactStorageColumnDeleteTest|3 of 33|[CASSANDRA-17213?|https://issues.apache.org/jira/browse/CASSANDRA-17213]|\r\n|org.apache.cassandra.cql3.ViewComplexTTLTest.testUnselectedColumnsTTLWithFlush[3]|1 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|dtest-offheap.cqlsh_tests.test_cqlsh_copy.TestCqlshCopy.test_round_trip_with_rate_file|1 of 33|[CASSANDRA-17322?|https://issues.apache.org/jira/browse/CASSANDRA-17322]|\r\n","created":"2022-05-24T14:42:54.310+0000"}],"conversations":[{"body":"org.apache.cassandra.dht.tokenallocator.OfflineTokenAllocatorTest\r\n\r\n[https://app.circleci.com/pipelines/github/josh-mckenzie/cassandra/168/workflows/1d30a113-c14b-4cf4-a631-bedb9eb65762/jobs/1447]\r\n\r\nLooks like a pretty simple / straightforward timeout\r\n\r\n {code}\r\njunit.framework.AssertionFailedError: Timeout occurred. Please note the time in the report does not reflect the time until the timeout.\r\n\tat java.util.Vector.forEach(Vector.java:1277)\r\n\tat java.util.Vector.forEach(Vector.java:1277)\r\n\tat java.util.Vector.forEach(Vector.java:1277)\r\n\tat org.apache.cassandra.anttasks.TestHelper.execute(TestHelper.java:53)\r\n\tat java.util.Vector.forEach(Vector.java:1277)\r\n{code}\r\n\r\n ","from":"reporter","subject":"Test Failure: unit test compression: testTokenGenerator_single_rack_or_single_rf"},{"body":"[~jmckenzie] , I was wondering what is the reason to leave in triage some of the test failures tickets like this one? Did you think it needs further testing and it might have been random failure? \r\n\r\nI am opening this one as I see it consistently failing in CircleCI. Also, my understanding is that we use the offline token allocator for the Python DTests so it is good to verify there is no real problem there.\r\n\r\nCC [~mck] ","from":"developer"},{"body":"Just missed hitting open when creating as build lead. Not intentional.","from":"developer"},{"body":"100 runs on JDK8 w/this passing in circle; going to try 100 on JDK11 then maybe up the repeat count.","from":"developer"},{"body":"I see it almost all the time with MIDRES in CircleCI ","from":"developer"},{"body":"Yeah, this is somewhat odd. using `generate.sh` to multiplex it it doesn't look like it's showing up; may have something to do with timing and other tests parallelizing and sharing state and stomping each other.\r\n\r\n \r\n\r\nI assume generate.sh is going to keep provisioning at the effectively \"LOWRES\" footprint on hardware alloc right?","from":"developer"},{"body":"generate.sh accepts flags for medium, high, etc. See [here|https://github.com/apache/cassandra/tree/trunk/.circleci#setting-environment-variables] i.e.","from":"developer"},{"body":"bq. generate.sh accepts flags for medium, high, etc. See here i.e.\r\nIndeed. And when you don't run it with a flag it stays at the default, which in this case is LOWRES.","from":"developer"},{"body":"Yes, and I checked currently LOWRES and MIDRES use the same resources for unit tests in particular.\r\n\r\nI ran this one  500 times, the whole class in  [here|https://app.circleci.com/pipelines/github/ekaterinadimitrova2/cassandra/1622/workflows/b69d6e45-b012-4fc7-a839-7320b86a319f/jobs/11106/steps] and weirdly it doesn't reproduce. But then I see it all the time in my MIDRES runs.\r\n\r\nI will experiment a bit more. \r\n\r\n ","from":"developer"},{"body":"Ok. So this was annoying but I think I got to the bottom of it.\r\n\r\n[PR|https://github.com/apache/cassandra/pull/1638]\r\n[JDK8 CI|https://app.circleci.com/pipelines/github/josh-mckenzie/cassandra/232/workflows/6c742dd9-e1b3-4056-aac8-9f0b95ae4e84]\r\n[JDK11 CI|https://app.circleci.com/pipelines/github/josh-mckenzie/cassandra/232/workflows/5fcce28f-a258-485f-906e-45957c2b3b08]\r\n\r\nThe offending bit of code:\r\n{code:java}\r\n for (int numTokens = 1; numTokens <= 16 ; ++numTokens)\r\n {\r\n for (int rf = 1; rf <=5; ++rf)\r\n {\r\n int nodeCount = 32;\r\n for (int racks = 1; racks <= 10; ++racks)\r\n {\r\n int[] nodeToRack = makeRackCountArray(nodeCount, racks);\r\n for (IPartitioner partitioner : new IPartitioner[] { Murmur3Partitioner.instance, RandomPartitioner.instance })\r\n {\r\n{code}\r\n{{void testTokenGenerations()}} was timing out due to this pretty beastly nesting of different permutations it runs combined with it logging output, etc. Effectively we have to run 1600 tests in 900 seconds sans whatever the other tests were consuming for time, and with all the logging about tokens and iteration through it looks like we tipped over an edge.\r\n\r\nI pared the test down to the following (about a third of the combinations with what I think is *logically* comparably useful coverage):\r\n{code:java}\r\n private final int[] racks = { 1, 2, 3, 5, 6, 9, 10 };\r\n private final int[] rfs = { 1, 2, 3, 5 };\r\n private final int[] tokens = { 1, 2, 3, 5, 6, 9, 10, 13, 15, 16 };\r\n{code}\r\nI toyed with getting rid of the logging inside the test class but the lion's share of what's spamming is in a variety of other classes. I also looked into disabling logging in the {{SystemOutputImpl}} in {{OfflineTokenAllocatorTestUtils}} but that ended up being more pain than it was worth.\r\n\r\nRather than taking 15+ minutes and timing out my laptop it takes about 20 seconds; should be ok in our CI env. Also split out this test method to its own file entirely so it doesn't stomp on the other offline token allocation tests and throw a red herring of timeout like this did (parallelization and method timeouts within a class are... not fun).\r\n ","from":"developer"},{"body":"change LGTM but the new test class doesn't end in Test so it doesn't run in CI; once that is fixed I am +1","from":"developer"},{"body":"bq. the new test class doesn't end in Test\r\nBlargh! It's always something. :)\r\n\r\nFigured I'd multiplex both the classes for 100 runs as well just to double check they're solid; I'll do that, fix the file name, and assuming all green, merge in.\r\n\r\nThanks [~dcapwell]!","from":"developer"},{"body":"Multiplex on both tests looked clean; merged up.","from":"developer"},{"body":"[CI Results]\r\nBranch: 4.1, build number: 35\r\n butler url: https://butler.cassandra.apache.org/#/ci/upstream/compare/Cassandra-4.1/Cassandra-4.1\r\n jenkins url: https://ci-cassandra.apache.org/job/Cassandra-4.1/35/\r\n JIRA: CASSANDRA-17291\r\n commit url: https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=commit;h=951aff25a1183f41fd146d674136399f3f25b3f0\r\n affected paths:\r\n* test/unit/org/apache/cassandra/dht/tokenallocator/OfflineTokenAllocatorGenerationsTest.java\r\n* test/unit/org/apache/cassandra/dht/tokenallocator/OfflineTokenAllocatorTestUtils.java\r\n* test/unit/org/apache/cassandra/dht/tokenallocator/OfflineTokenAllocatorTest.java\r\n\r\n Build Result: UNSTABLE\r\n Passing Tests: 47094\r\n Failing Tests: 27\r\n\r\n||Test|Failures|JIRA||\r\n|org.apache.cassandra.cql3.validation.operations.CompactStorageTest.testCounterAndColumnSelection|1 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.distributed.upgrade.MixedModeAvailabilityV30Test.testAvailability|5 of 33|[CASSANDRA-17307?|https://issues.apache.org/jira/browse/CASSANDRA-17307]|\r\n|org.apache.cassandra.distributed.test.SchemaTest.readRepairWithCompaction|4 of 33|[CASSANDRA-17641?|https://issues.apache.org/jira/browse/CASSANDRA-17641]|\r\n|org.apache.cassandra.distributed.test.MessageForwardingTest.mutationsForwardedToAllReplicasTest|1 of 33|[CASSANDRA-17583?|https://issues.apache.org/jira/browse/CASSANDRA-17583]|\r\n|org.apache.cassandra.cql3.validation.operations.CompactStorageTest.testAlterWithCompactNonStaticFormat|2 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.distributed.test.ring.BootstrapTest.readWriteDuringBootstrapTest|3 of 33|[Multiple JIRAs found|https://issues.apache.org/jira/issues/?jql=project%20%3D%20CASSANDRA%20and%20resolution%20%3D%20unresolved%20and%20summary%20~%20\"*BootstrapTest*\"]\r\n|org.apache.cassandra.distributed.upgrade.CompactStorageUpgradeTest.compactStorageImplicitNullInClusteringTest|5 of 33|[CASSANDRA-17213?|https://issues.apache.org/jira/browse/CASSANDRA-17213]|\r\n|junit.framework.TestSuite.org.apache.cassandra.distributed.test.CASMultiDCTest|2 of 33|[Multiple JIRAs found|https://issues.apache.org/jira/issues/?jql=project%20%3D%20CASSANDRA%20and%20resolution%20%3D%20unresolved%20and%20summary%20~%20\"*test*\"]\r\n|dtest-upgrade.upgrade_tests.drop_compact_storage_upgrade_test.TestDropCompactStorage.test_drop_compact_storage_mixed_cluster|4 of 33|[CASSANDRA-17634?|https://issues.apache.org/jira/browse/CASSANDRA-17634]|\r\n|org.apache.cassandra.tools.TopPartitionsTest.testServiceTopPartitionsSingleTable-cdc|5 of 33|[CASSANDRA-17649?|https://issues.apache.org/jira/browse/CASSANDRA-17649]|\r\n|org.apache.cassandra.cql3.validation.entities.SecondaryIndexTest.testIndexOnRegularColumnInsertExpiringColumnWithFlush|3 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.cql3.validation.operations.CompactStorageTest.testStaticCompactWithCounters|2 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.distributed.upgrade.BatchUpgradeTest.batchTest|6 of 33|[CASSANDRA-17651?|https://issues.apache.org/jira/browse/CASSANDRA-17651]|\r\n|org.apache.cassandra.cql3.ViewFilteringClustering1Test.testClusteringKeySliceRestrictions[3]|6 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestProtoV4Upgrade_AllVersions_EndsAt_Trunk_HEAD.test_parallel_upgrade|1 of 33|[CASSANDRA-17296?|https://issues.apache.org/jira/browse/CASSANDRA-17296]|\r\n|org.apache.cassandra.cql3.validation.operations.SelectTest.filteringWithOrderClause|3 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.cql3.validation.entities.SecondaryIndexTest.testIndexOnNonFrozenCollectionOfFrozenUDT|1 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|org.apache.cassandra.distributed.upgrade.CompactStorageUpgradeTest.compactStoragePagingTest|3 of 33|[CASSANDRA-17213?|https://issues.apache.org/jira/browse/CASSANDRA-17213]|\r\n|org.apache.cassandra.db.VerifyTest.testMutateRepair-cdc|1 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|dtest-upgrade.upgrade_tests.upgrade_through_versions_test.TestUpgrade_current_4_0_x_To_indev_4_1_x.test_parallel_upgrade_with_internode_ssl|1 of 33|[Multiple JIRAs found|https://issues.apache.org/jira/issues/?jql=project%20%3D%20CASSANDRA%20and%20resolution%20%3D%20unresolved%20and%20summary%20~%20\"*TestUpgrade*\"]\r\n|org.apache.cassandra.distributed.test.CASTest.testIncompleteWriteFollowedBySuccessfulWriteWithStaleRingDuringRangeMovementFollowedByRead|1 of 33|[CASSANDRA-17461?|https://issues.apache.org/jira/browse/CASSANDRA-17461]|\r\n|org.apache.cassandra.db.virtual.GossipInfoTableTest.testSelectAllWithStateTransitions|1 of 33|[CASSANDRA-17584?|https://issues.apache.org/jira/browse/CASSANDRA-17584]|\r\n|org.apache.cassandra.distributed.upgrade.CompactStorageUpgradeTest.compactStorageColumnDeleteTest|3 of 33|[CASSANDRA-17213?|https://issues.apache.org/jira/browse/CASSANDRA-17213]|\r\n|org.apache.cassandra.cql3.ViewComplexTTLTest.testUnselectedColumnsTTLWithFlush[3]|1 of 33|[No JIRA found|https://issues.apache.org/jira/secure/RapidBoard.jspa?rapidView=496&quickFilter=2252]\r\n|dtest-offheap.cqlsh_tests.test_cqlsh_copy.TestCqlshCopy.test_round_trip_with_rate_file|1 of 33|[CASSANDRA-17322?|https://issues.apache.org/jira/browse/CASSANDRA-17322]|\r\n","from":"developer"}],"created":"2022-01-25T16:45:15.000+0000","description":"org.apache.cassandra.dht.tokenallocator.OfflineTokenAllocatorTest\r\n\r\n[https://app.circleci.com/pipelines/github/josh-mckenzie/cassandra/168/workflows/1d30a113-c14b-4cf4-a631-bedb9eb65762/jobs/1447]\r\n\r\nLooks like a pretty simple / straightforward timeout\r\n\r\n {code}\r\njunit.framework.AssertionFailedError: Timeout occurred. Please note the time in the report does not reflect the time until the timeout.\r\n\tat java.util.Vector.forEach(Vector.java:1277)\r\n\tat java.util.Vector.forEach(Vector.java:1277)\r\n\tat java.util.Vector.forEach(Vector.java:1277)\r\n\tat org.apache.cassandra.anttasks.TestHelper.execute(TestHelper.java:53)\r\n\tat java.util.Vector.forEach(Vector.java:1277)\r\n{code}\r\n\r\n ","issue_id":"13424738","key":"CASSANDRA-17291","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2022-05-23T17:03:58.000+0000","role":"fixed_distractor","summary":"Test Failure: unit test compression: testTokenGenerator_single_rack_or_single_rf"} {"case_id":"13430156","cluster":"DISTRACTOR-CASSANDRA-17401","comments":[{"body":"[~ifesdjeen] , [~marcuse] , do you mind to take a look, please? Looking at the links provided, seems related to the issues you were working on?","created":"2022-02-23T14:36:08.021+0000"},{"body":"[~ivansenic] thank you for reporting this. While I can imagine how something like this may happen, I do not understand how it is bad or dangerous, could you elaborate? It is not in the cache, so it will get re-prepared. I'm fairly certain the fuzz test included with the patch covers the this. What specifically are you suggesting to do, since the issue just describes a behaviour and calls it a bug but gives little detail besides that. ","created":"2022-02-23T16:16:38.221+0000"},{"body":"[~ifesdjeen] I have to admin that internals of Cassandra are not 100% familiar to me. However, wouldn't there be a _PreparedQueryNotFoundException_ thrown in the _ExecuteMessage_ or the _BatchMessage_ if the _QueryHandler.getPrepared_ returns null? I don't know how this is exception handled or how is the re-preparing actually working.\r\n\r\nMaybe I was not clear in my explanation above, but bottom line is that if the race condition is uncovering the issue described, the _QueryHandler.getPrepared_  will return null for the MD5 that was just returned by the same class prepare method call. If you think this is acceptable, then you can close the issue.","created":"2022-02-23T17:08:01.340+0000"},{"body":"It seems to me that returning `null` – if that is to be expected – definitely needs to be well documented. It would also seem quite awkward for caller to essentially have to retry the operation, theoretically multiple times (more than 2 threads).\r\nOr perhaps I have misunderstood something.","created":"2022-02-23T17:14:48.010+0000"},{"body":"[~ifesdjeen] [~e.dimitrova] You can close the issue. I do understand the on execution of the prepared statement it will be re-prepared if the cache does not contain it. We ll adapt usage in Stargate.","created":"2022-02-24T13:06:09.197+0000"},{"body":"We have been facing the exact same problem in our production environment. \r\n\r\nAs part of CASSANDRA-17248 (in C* 3.0.26), the following code was introduced in [QueryProcessor.java|https://github.com/apache/cassandra/commit/242f7f9b18db77bce36c9bba00b2acda4ff3209e#r137491766]\r\n{code:java}\r\n// Make sure the missing one is going to be eventually re-prepared \r\nevictPrepared(hashWithKeyspace); \r\nevictPrepared(hashWithoutKeyspace); {code}\r\nThis code could very well create a race condition between two calls of [_QueryProcessor::prepare_|https://github.com/apache/cassandra/blob/cassandra-4.0/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L575] call in which one thread is adding and another thread is silently removing from the cache. Imagine there are thousands of threads calling the API, and then it might be possible for one thread to update the cache and another thread to remove it.\r\n\r\nIf we look at the code of [Cassandra 3.0.25|https://github.com/apache/cassandra/blob/cassandra-3.0.25/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L391], then such eviction was not present. Hence, this seems like a regression since the 3.0.26 version of the Cassandra.\r\n\r\n\r\nTo fix this issue, we should not *evict* the cache entries, i.e., the above-mentioned code path introduced since C* 3.0.26 is no longer required.\r\n\r\n[~ifesdjeen] , Could you please take a look at it?\r\n\r\n \r\n\r\n ","created":"2024-01-20T16:24:19.703+0000"},{"body":"Hi [~chovatia.jaydeep@gmail.com] can you provide a regression test case reproducing this issue and a patch with a proposed fix ?","created":"2024-01-21T16:46:29.523+0000"},{"body":"Sure, [~paulo] \r\n\r\nPlease note that reproducing this issue is extremely tricky as it depends on the thread contention, timing, etc., so I have injected some delay in preparing the statement code path to reproduce. If we run this test case on 3.0.25 or before, it does not reproduce, i.e., it is a regression since 3.0.26. Here is the PR that reproduces this issue: [https://github.com/apache/cassandra/pull/3058]\r\n\r\n \r\n\r\nHere is the proposed fix: [https://github.com/apache/cassandra/pull/3059]\r\n\r\n \r\n\r\nPlease take a look at it and let me know your comments.","created":"2024-01-21T20:59:38.338+0000"},{"body":"Ok thanks [~chovatia.jaydeep@gmail.com]! I'm not familiar with this area but will try to look at it if I find cycles in the next few days and nobody beats me to it. :)\r\n\r\nBtw did you observe a single occurrence of this issue, or is it recurrent?","created":"2024-01-22T01:53:50.545+0000"},{"body":"Thank you, [~paulo] for your help with this ticket.\r\n\r\nYes, it has occurred one time in production. After that, we could successfully reproduce a couple of times by injecting some delay in thread scheduling and by running a very high concurrency on _prepare_ code. ","created":"2024-01-22T04:43:37.219+0000"},{"body":"[~chovatia.jaydeep@gmail.com] and I managed to reproduce the issue. Basically a large number of QPS and client connections are necessary to reproduce it. Here are the steps to reproduce:\r\n\r\n*Server Setup (Cassandra 4.0.6):*\r\nA 3-node Cassandra cluster. Each node with 64GB mem (16GB heap), 7 CPU cores. ({{{}native_transport_max_threads = 1024).{}}}\r\n\r\n*Keyspace/Table:*\r\nCREATE KEYSPACE test_ks WITH REPLICATION = \\{ ‘class’ : ‘NetworkTopologyStrategy’, ‘datacenter1’ : 3 } ;\r\nCREATE TABLE test_ks.table1 ( p_id text, c_id text, v text);\r\n\r\n*Client Setup:*\r\n30 hosts (12 CPU cores per host). Each host run the following pseudo-code, using *GoCql* client:\r\n{code:java}\r\n cluster.CQLVersion = \"3.4.0\"\r\n cluster.ProtoVersion = 4\r\n cluster.Timeout = 5s\r\n cluster.ConnectTimeout = 10s\r\n cluster.NumConns = 3\r\n cluster.Consistency = LocalQuorum\r\n cluster.RetryPolicy = SimpleRetryPolicy{NumRetries: 1}\r\n cluster.SocketKeepalive = 20s\r\n cluster.HostSelectionPolicy = RoundRobinHostPolicy\r\n\r\n sessionCount = 30\r\n qpsPerSession = 30\r\n cqlQuery = \"SELECT p_id,c_id,v FROM test_ks.table1 WHERE p_id = ? AND c_id = ?\"\r\n for (i = 0; i < sessionCount; i++) {\r\n session = cluster.createSession\r\n rateLimiter = NewRateLimiter(qpsPerSession)\r\n newGoRoutine.run( sendReads(session, rateLimiter) )\r\n }\r\n\r\n / *\r\n sendReads(session, rateLimiter) {\r\n for {\r\n newGoRoutine.run (\r\n if (rateLimiter.allow) {\r\n session.execute(cqlQuery, randomString, randomString)\r\n }\r\n )\r\n }\r\n }\r\n */ {code}\r\nTraffic generated this way will result in ~10K coordiator QPS and ~3k client connections per Cassandra node.\r\n\r\n*Trigger Point:*\r\nManually issue a CQL query to add a column in the table: “ALTER TABLE test_ks.table1 ADD new_col text;”\r\n\r\n{*}Symmpton{*}:\r\nSeconds after the trigger point, one or more Cassandra nodes will show number of native_transport threads reaching {{{}native_transport_max_threads{}}}, and pending native transport tasks grow endlessly.","created":"2024-01-23T06:34:33.752+0000"},{"body":"Hi [~paulo],\r\n\r\n \r\n\r\nWould appreciate it if you could spare time to review the pull request.\r\n\r\n \r\n\r\nJaydeep","created":"2024-01-29T18:33:39.021+0000"},{"body":"[~chovatia.jaydeep@gmail.com] I was not able to review this yet, will send an update when I get a chance to review it. If anyone else subscribed wants to review this on the meantime feel free to take it.","created":"2024-01-29T19:26:40.882+0000"},{"body":"Thanks for the detailed reports and repro steps. I've taken a look and this looks to me to be a legitimate race condition that can cause a re-prepare storm under large concurrency and unlucky timing.\r\n\r\nMy understanding is that [these evict statements|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L735] are not required for the correctness of the upgrade compatibility logic and can be safely removed. Would you have some cycles to confirm this [~ifesdjeen] ?\r\n\r\nIn addition to this, I think there's a pending issue from CASSANDRA-17248 that can leak prepared statements between keyspaces during mixed upgrade mode. Since these issues are in a related area I think it makes sense to address them together (in separate commits) to ensure these changes are tested together.\r\n\r\nI think the {{PreparedStatementCollisionTest}} suite from [this commit|https://github.com/apache/cassandra/pull/1872/commits/758bc4a89d7ca9d0bfe27e6f41000484724261bc] can help improve the validation coverage of this logic. That change looks correct to me but may need some cleanup. We should probably keep the metric changes out of this to keep the scope of this patch to a minimum.\r\n\r\nAfter proper review and validation I think there's value in including these fixes in the final 3.X releases to address these outstanding issues as users will still do upgrade cycles as 5.x release approaches. This will make resolution more laborious as we will need to provide patches for 3.x all the way up to trunk + CI for all branches. What do you think [~brandon.williams] [~stefan.miklosovic]  ?","created":"2024-02-28T02:07:26.941+0000"},{"body":"[~paulo] I have ran both regular and mixed mode fuzz tests and could not make them fail with this change. I've also looked at the code again and it looks like in cases where we would previously evict, in the new version we would just insert the missing one. That said, I do not see this as a correctness issue (not that it should not be still improved upon), so I would probably focus on 17248.","created":"2024-02-28T16:17:00.476+0000"},{"body":"Hi, [~paulo] [~ifesdjeen] \r\n\r\nIt appears that we've encountered the exact issue [~chovatia.jaydeep@gmail.com] described here. During our production migration involving an {{ALTER TABLE ... ADD COLUMN}} operation on Cassandra 4.1.3, we experienced a severe reprepare storm — a situation where prepared statements are repeatedly invalidated and reprepared across cluster nodes, leading to significant performance degradation.\r\n \r\nCASSANDRA-17248, which was mentioned in the discussion, does not fully address the root cause of this problem (we did not perform a Cassandra version upgrade). After analyzing the behaviour, I believe the most effective long-term solution is to stop evicting the prepared statement cache.\r\n \r\nCan we revisit the proposed fix that implements this approach?\r\n[https://github.com/apache/cassandra/pull/3059]\r\n\r\n \r\n\r\n ","created":"2025-12-10T14:54:23.373+0000"},{"body":"[~norkandrei] I have just raised this to the dev mailing list so someone can prioritize the review. In the meantime, you may want to continue with the private fix ([https://github.com/apache/cassandra/pull/3059]) until it is officially available in the repo. I have been using this fix in our setup, and it has been working fine.","created":"2025-12-14T20:05:18.222+0000"},{"body":"I have looked around, and it looks like we rely on having both statements during upgrades to check if statements are cached correctly (perhaps we need to add a corresponding comment).\r\n\r\nI think we can still make it work though. I think we should just use {{preparedStatements.invalidateAll()}} instead of regular {{invalidate}}. Unfortunately we will be able to somehow synchronize deletes from stable storage, but I think this is fine, we can allow races there, since they're not correctness-impacting. Eventually we will save both statements with correct hashes.\r\n\r\nAnother addition to the patch I would do is adding an if statement when retrieving {{cachedWithKeyspace}} after {{cachedWithoutKeyspace}}, and only going for a second {{get}} if hashes are actually different. \r\n\r\nLMK if this approach still fixes the issue for you.","created":"2025-12-16T11:46:13.376+0000"},{"body":"[~ifesdjeen] As I understand it, this eviction logic was introduced to handle a hash calculation change in 4.0.2 and is only needed during upgrades from pre-4.0.2 to 4.0.2. If all nodes are already on 4.0.2+, the eviction is unnecessary.\r\n\r\nReplacing _invalidate_ with _preparedStatements.invalidateAll_ does not address the issue in this ticket. Given this, we should apply the eviction logic only when pre-4.0.2 nodes are present and skip it otherwise. Happy to hear your thoughts.","created":"2025-12-16T23:58:55.784+0000"},{"body":"I am personally not a big fan of adding version-dependent logic, since it complicates maintenance and testing substantially. I think we do need to replace invalidate with invalidateAll regardless of this ticket, since this is a race, too. Another thing I'd do is I'd check if removal listener is sufficient to clean up the prep statements table, and avoid explicitly calling {{removePreparedStatement}}. \r\n\r\nAnother approach is to evict only if we see that one of the statements is cached but not the other (i.e. {{ (a == null && b != null) && (b == null && a != null) }}), which is going to substantially reduce the window of race possibility. This, combined with what I wrote earlier, namely\r\n\r\nbq. Another addition to the patch I would do is adding an if statement when retrieving {{cachedWithKeyspace}} after {{cachedWithoutKeyspace}}, and only going for a second get if hashes are actually different.\r\n\r\nshould be sufficient to trigger only a handful of spurious re-prepares. WDYT?","created":"2025-12-17T07:39:26.205+0000"},{"body":"> I am personally not a big fan of adding version-dependent logic, since it complicates maintenance and testing substantially.\r\n\r\nAs of today, in the trunk, we already have 4.0.1 -> 4.0.2 version-specific [logic|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L803]; please see the \"else\" portion.\r\n{code:java}\r\n            if (useNewPreparedStatementBehaviour)\r\n            {\r\n                if (cachedWithoutKeyspace.fullyQualified) // For fully qualified statements, we always skip keyspace to avoid digest switching\r\n                    return createResultMessage(hashWithoutKeyspace, cachedWithoutKeyspace);                if (clientState.getRawKeyspace() != null && !cachedWithKeyspace.fullyQualified) // For non-fully qualified statements, we always include keyspace to avoid ambiguity\r\n                    return createResultMessage(hashWithKeyspace, cachedWithKeyspace);            }\r\n            else // legacy caches, pre-CASSANDRA-15252 behaviour\r\n            {\r\n                return createResultMessage(hashWithKeyspace, cachedWithKeyspace);\r\n            } {code}\r\nThis eviction logic became necessary for 4.0.1 -> 4.0.2, and is not needed afterwards. Therefore, the proposed PR on _trunk_ removes the \"eviction\" logic entirely.\r\n\r\nAnd we can selectively decide till what version we want to land the proposed PR - I propose to land it on trunk and 5.0, maybe 4.1, but not land on 4.x.\r\n\r\n \r\n\r\n>Another approach is to evict only if we see that one of the statements is cached but not the other (i.e., \\{{ (a == null && b != null) && (b == null && a != null) }}), which is going to reduce the window of race possibility substantially. \r\n\r\nI agree that this logic reduces the possibility of a race, but it does not eliminate it. The point I am trying to make is not to keep unnecessary code/logic that is only needed for the transition from version 4.0.1 to 4.0.2.\r\n\r\nWDYT?","created":"2025-12-18T00:04:18.236+0000"},{"body":"Alright, let's do that. I still think we should apply the patch to reduce the race window if you got time for that. ","created":"2025-12-18T07:52:10.220+0000"},{"body":"Just to be clear, you are ok with the proposed PR [https://github.com/apache/cassandra/pull/3059] for trunk/5.0, i.e., no eviction.\r\n\r\nHowever, for the 4.0 branch, keep the eviction logic, and on top of that, you want me to reduce the probability window. Is that accurate? ","created":"2025-12-19T01:17:24.957+0000"},{"body":"I fully support Jaydeepkumar Chovatia's proposal to remove the invalidation logic in versions 4.1, 5.0, and trunk, as implemented in PR [https://github.com/apache/cassandra/pull/3059 |https://github.com/apache/cassandra/pull/3059]. Although the idea to reduce the window of race possibility mitigates the issue, it does not completely eliminate the underlying risk.","created":"2025-12-19T09:37:14.927+0000"},{"body":"[~chovatia.jaydeep@gmail.com] that's right: no eviction in trunk, in 4.0 predicate eviction on presence of pre-4.0.2 nodes (as you suggested), and, if you are not opposed, to add the suggestions that would improve the code in both situations.\r\n\r\nThank you!","created":"2025-12-19T10:33:08.723+0000"},{"body":"Thanks [~ifesdjeen]. In that case, are you okay with reviewing/approving the proposed PR for trunk so I can formally merge it into the {_}trunk{_}? [https://github.com/apache/cassandra/pull/3059]\r\n\r\nI will create a new PR for the 4.0 branch and keep eviction + extra safety checks to minimize the issue mentioned in this ticket on the 4.0 branch.","created":"2025-12-22T05:43:02.858+0000"},{"body":"[~ifesdjeen] I have also created a PR for the 4.0 branch (tests are currently running). Could you please take a look at it? [https://github.com/apache/cassandra/pull/4542]","created":"2025-12-23T06:15:13.910+0000"},{"body":"Additionally, I request that the update removing the cache-clearing functionality ([https://github.com/apache/cassandra/pull/3059 |https://github.com/apache/cassandra/pull/3059]) be included in version 4.1.","created":"2025-12-26T06:11:11.314+0000"},{"body":"[~ifesdjeen] I'd really appreciate it if you could find some time to review the PR and merge the PR — it addresses a much-needed fix.","created":"2026-01-12T09:25:43.965+0000"},{"body":"> update removing the cache-clearing functionality be included in version 4.1.\r\nFrom what I understand, this is the version where it is actually needed: we need to clean up both hashes to preclude an incorrect version from being loaded, in case someone upgrades from 4.0.1 to 4.1.","created":"2026-01-12T09:54:19.114+0000"},{"body":"I suppose that the possibility of upgrading from version 4.0.1 to 4.1 is significantly lower compared to the risk of encountering the reprepare storm problem in all 4.1.x versions. Moreover, adding an additional check before invalidation does not fully eliminate the root cause of the repreparation issue. Therefore, in my opinion, the best solution is to make this logic applicable only to version 4.0.","created":"2026-01-13T14:17:59.470+0000"},{"body":"I have just committed the approved PR to {_}trunk{_}. [https://github.com/apache/cassandra/commit/a06df099f4e1c6265d25c94336c493aded3aa644] & https://github.com/apache/cassandra/commit/aa4da838ec0d05f7c7494536082271a6ed135fa1\r\n\r\nHere is the PR for 5.0 [https://github.com/jaydeepkumar1984/cassandra/pull/67] - please review.","created":"2026-01-19T06:10:20.234+0000"},{"body":"[~chovatia.jaydeep@gmail.com], unless there's reason, PRs for all branches should be merged in one go, and pushed atomically.\r\nref: https://cassandra.apache.org/_/development/how_to_commit.html\r\n(there's no need to revert, just fyi on the project precedence)","created":"2026-01-19T17:47:23.993+0000"},{"body":"Sure [~mck] - I will follow this practice going forward.","created":"2026-01-19T18:43:03.218+0000"},{"body":"PR for 5.0 [https://github.com/jaydeepkumar1984/cassandra/pull/67] - please review and merge. Kindly reminder :)","created":"2026-01-26T09:49:31.930+0000"},{"body":"[~ifesdjeen] Plz review PR for 5.0 [https://github.com/jaydeepkumar1984/cassandra/pull/67]","created":"2026-02-02T06:44:39.027+0000"},{"body":"[~norkandrei] I have already reviewed this patch for trunk. I think committer can just merge it upwards. ","created":"2026-02-03T09:16:51.466+0000"},{"body":"Thanks, Alex.\r\n\r\n[~norkandrei] I will port the patch to 5.0 sometime this weekend.","created":"2026-02-03T19:26:24.093+0000"},{"body":"[~ifesdjeen] [~chovatia.jaydeep@gmail.com] Thank you!","created":"2026-02-04T11:18:55.457+0000"},{"body":"The fix has also been landed in 5.0. SHA: [https://github.com/apache/cassandra/commit/d90fe76c8efb374c3b1b09e95f6e9925090fe638]\r\n\r\n[https://github.com/apache/cassandra/commit/4a2bf29a34cea273c5d23e0b9cd3c19ef0954853]\r\n\r\n \r\n\r\nFor the 4.x, I responded to the review comment from [~ifesdjeen] as 4.x change will be different than 5.0 and trunk; once that is approved, I will take care of 4.x too. [https://github.com/apache/cassandra/pull/4542/changes] ","created":"2026-02-08T22:01:37.926+0000"},{"body":"[~chovatia.jaydeep@gmail.com] just looked at it again, and I agree with you. Thanks for elaborating! +1","created":"2026-02-09T08:28:44.535+0000"},{"body":"Committed to 4.1 ([SHA|https://github.com/apache/cassandra/commit/4ada37e940431686cbc82e2d0b01f482f0041a3f]) & 4.0 ([SHA|https://github.com/apache/cassandra/commit/077b7ebe22c750b9f0c4a83f7979356b5095d6a5])","created":"2026-02-11T01:30:13.220+0000"},{"body":"This fix has been landed to all the Cassandra versions: 4.0.20, 4.1.11, 5.0.7, and 5.1\r\n\r\nThanks [~ifesdjeen]  for your help with the code review!","created":"2026-02-11T01:32:33.735+0000"}],"conversations":[{"body":"The changes in the [QueryProcessor#prepare|https://github.com/apache/cassandra/blame/cassandra-4.0.2/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L575-L638] method that were introduced in versions *4.0.2* and *3.11.12* can cause a race condition between two threads trying to concurrently prepare the same statement. This race condition can cause removing of a prepared statement from the cache, after one of the threads has received the result of the prepare and eventually uses MD5Digest to call [QueryProcessor#getPrepared|https://github.com/apache/cassandra/blame/cassandra-4.0.2/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L212-L215].\r\n\r\nThe race condition looks like this:\r\n * Thread1 enters _prepare_ method and resolves _safeToReturnCached_ as false\r\n * Thread1 executes eviction of hashes\r\n * Thread2 enters _prepare_ method and resolves _safeToReturnCached_ as false\r\n * Thread1 prepares the statement and caches it\r\n * Thread1 returns the result of the prepare\r\n * Thread2 executes eviction of hashes\r\n * Thread1 tries to execute the prepared statement with the received MD5Digest, but statement is not in the cache as it was evicted by Thread2\r\n\r\nI tried to reproduce this by using a Java driver, but hitting this case from a client side is highly unlikely and I can not simulate the needed race condition. However, we can easily reproduce this in Stargate (details [here|https://github.com/stargate/stargate/pull/1647]), as it's closer to QueryProcessor.\r\n\r\nReproducing this in a unit test is fairly easy. I am happy to showcase this if needed.\r\n\r\nNote that the issue can occur only when  safeToReturnCached is resolved as false.","from":"reporter","subject":"Race condition in QueryProcessor causes just prepared statement not to be in the prepared statements cache"},{"body":"[~ifesdjeen] , [~marcuse] , do you mind to take a look, please? Looking at the links provided, seems related to the issues you were working on?","from":"developer"},{"body":"[~ivansenic] thank you for reporting this. While I can imagine how something like this may happen, I do not understand how it is bad or dangerous, could you elaborate? It is not in the cache, so it will get re-prepared. I'm fairly certain the fuzz test included with the patch covers the this. What specifically are you suggesting to do, since the issue just describes a behaviour and calls it a bug but gives little detail besides that. ","from":"developer"},{"body":"[~ifesdjeen] I have to admin that internals of Cassandra are not 100% familiar to me. However, wouldn't there be a _PreparedQueryNotFoundException_ thrown in the _ExecuteMessage_ or the _BatchMessage_ if the _QueryHandler.getPrepared_ returns null? I don't know how this is exception handled or how is the re-preparing actually working.\r\n\r\nMaybe I was not clear in my explanation above, but bottom line is that if the race condition is uncovering the issue described, the _QueryHandler.getPrepared_  will return null for the MD5 that was just returned by the same class prepare method call. If you think this is acceptable, then you can close the issue.","from":"developer"},{"body":"It seems to me that returning `null` – if that is to be expected – definitely needs to be well documented. It would also seem quite awkward for caller to essentially have to retry the operation, theoretically multiple times (more than 2 threads).\r\nOr perhaps I have misunderstood something.","from":"developer"},{"body":"[~ifesdjeen] [~e.dimitrova] You can close the issue. I do understand the on execution of the prepared statement it will be re-prepared if the cache does not contain it. We ll adapt usage in Stargate.","from":"developer"},{"body":"We have been facing the exact same problem in our production environment. \r\n\r\nAs part of CASSANDRA-17248 (in C* 3.0.26), the following code was introduced in [QueryProcessor.java|https://github.com/apache/cassandra/commit/242f7f9b18db77bce36c9bba00b2acda4ff3209e#r137491766]\r\n{code:java}\r\n// Make sure the missing one is going to be eventually re-prepared \r\nevictPrepared(hashWithKeyspace); \r\nevictPrepared(hashWithoutKeyspace); {code}\r\nThis code could very well create a race condition between two calls of [_QueryProcessor::prepare_|https://github.com/apache/cassandra/blob/cassandra-4.0/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L575] call in which one thread is adding and another thread is silently removing from the cache. Imagine there are thousands of threads calling the API, and then it might be possible for one thread to update the cache and another thread to remove it.\r\n\r\nIf we look at the code of [Cassandra 3.0.25|https://github.com/apache/cassandra/blob/cassandra-3.0.25/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L391], then such eviction was not present. Hence, this seems like a regression since the 3.0.26 version of the Cassandra.\r\n\r\n\r\nTo fix this issue, we should not *evict* the cache entries, i.e., the above-mentioned code path introduced since C* 3.0.26 is no longer required.\r\n\r\n[~ifesdjeen] , Could you please take a look at it?\r\n\r\n \r\n\r\n ","from":"developer"},{"body":"Hi [~chovatia.jaydeep@gmail.com] can you provide a regression test case reproducing this issue and a patch with a proposed fix ?","from":"developer"},{"body":"Sure, [~paulo] \r\n\r\nPlease note that reproducing this issue is extremely tricky as it depends on the thread contention, timing, etc., so I have injected some delay in preparing the statement code path to reproduce. If we run this test case on 3.0.25 or before, it does not reproduce, i.e., it is a regression since 3.0.26. Here is the PR that reproduces this issue: [https://github.com/apache/cassandra/pull/3058]\r\n\r\n \r\n\r\nHere is the proposed fix: [https://github.com/apache/cassandra/pull/3059]\r\n\r\n \r\n\r\nPlease take a look at it and let me know your comments.","from":"developer"},{"body":"Ok thanks [~chovatia.jaydeep@gmail.com]! I'm not familiar with this area but will try to look at it if I find cycles in the next few days and nobody beats me to it. :)\r\n\r\nBtw did you observe a single occurrence of this issue, or is it recurrent?","from":"developer"},{"body":"Thank you, [~paulo] for your help with this ticket.\r\n\r\nYes, it has occurred one time in production. After that, we could successfully reproduce a couple of times by injecting some delay in thread scheduling and by running a very high concurrency on _prepare_ code. ","from":"developer"},{"body":"[~chovatia.jaydeep@gmail.com] and I managed to reproduce the issue. Basically a large number of QPS and client connections are necessary to reproduce it. Here are the steps to reproduce:\r\n\r\n*Server Setup (Cassandra 4.0.6):*\r\nA 3-node Cassandra cluster. Each node with 64GB mem (16GB heap), 7 CPU cores. ({{{}native_transport_max_threads = 1024).{}}}\r\n\r\n*Keyspace/Table:*\r\nCREATE KEYSPACE test_ks WITH REPLICATION = \\{ ‘class’ : ‘NetworkTopologyStrategy’, ‘datacenter1’ : 3 } ;\r\nCREATE TABLE test_ks.table1 ( p_id text, c_id text, v text);\r\n\r\n*Client Setup:*\r\n30 hosts (12 CPU cores per host). Each host run the following pseudo-code, using *GoCql* client:\r\n{code:java}\r\n cluster.CQLVersion = \"3.4.0\"\r\n cluster.ProtoVersion = 4\r\n cluster.Timeout = 5s\r\n cluster.ConnectTimeout = 10s\r\n cluster.NumConns = 3\r\n cluster.Consistency = LocalQuorum\r\n cluster.RetryPolicy = SimpleRetryPolicy{NumRetries: 1}\r\n cluster.SocketKeepalive = 20s\r\n cluster.HostSelectionPolicy = RoundRobinHostPolicy\r\n\r\n sessionCount = 30\r\n qpsPerSession = 30\r\n cqlQuery = \"SELECT p_id,c_id,v FROM test_ks.table1 WHERE p_id = ? AND c_id = ?\"\r\n for (i = 0; i < sessionCount; i++) {\r\n session = cluster.createSession\r\n rateLimiter = NewRateLimiter(qpsPerSession)\r\n newGoRoutine.run( sendReads(session, rateLimiter) )\r\n }\r\n\r\n / *\r\n sendReads(session, rateLimiter) {\r\n for {\r\n newGoRoutine.run (\r\n if (rateLimiter.allow) {\r\n session.execute(cqlQuery, randomString, randomString)\r\n }\r\n )\r\n }\r\n }\r\n */ {code}\r\nTraffic generated this way will result in ~10K coordiator QPS and ~3k client connections per Cassandra node.\r\n\r\n*Trigger Point:*\r\nManually issue a CQL query to add a column in the table: “ALTER TABLE test_ks.table1 ADD new_col text;”\r\n\r\n{*}Symmpton{*}:\r\nSeconds after the trigger point, one or more Cassandra nodes will show number of native_transport threads reaching {{{}native_transport_max_threads{}}}, and pending native transport tasks grow endlessly.","from":"developer"},{"body":"Hi [~paulo],\r\n\r\n \r\n\r\nWould appreciate it if you could spare time to review the pull request.\r\n\r\n \r\n\r\nJaydeep","from":"developer"},{"body":"[~chovatia.jaydeep@gmail.com] I was not able to review this yet, will send an update when I get a chance to review it. If anyone else subscribed wants to review this on the meantime feel free to take it.","from":"developer"},{"body":"Thanks for the detailed reports and repro steps. I've taken a look and this looks to me to be a legitimate race condition that can cause a re-prepare storm under large concurrency and unlucky timing.\r\n\r\nMy understanding is that [these evict statements|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L735] are not required for the correctness of the upgrade compatibility logic and can be safely removed. Would you have some cycles to confirm this [~ifesdjeen] ?\r\n\r\nIn addition to this, I think there's a pending issue from CASSANDRA-17248 that can leak prepared statements between keyspaces during mixed upgrade mode. Since these issues are in a related area I think it makes sense to address them together (in separate commits) to ensure these changes are tested together.\r\n\r\nI think the {{PreparedStatementCollisionTest}} suite from [this commit|https://github.com/apache/cassandra/pull/1872/commits/758bc4a89d7ca9d0bfe27e6f41000484724261bc] can help improve the validation coverage of this logic. That change looks correct to me but may need some cleanup. We should probably keep the metric changes out of this to keep the scope of this patch to a minimum.\r\n\r\nAfter proper review and validation I think there's value in including these fixes in the final 3.X releases to address these outstanding issues as users will still do upgrade cycles as 5.x release approaches. This will make resolution more laborious as we will need to provide patches for 3.x all the way up to trunk + CI for all branches. What do you think [~brandon.williams] [~stefan.miklosovic]  ?","from":"developer"},{"body":"[~paulo] I have ran both regular and mixed mode fuzz tests and could not make them fail with this change. I've also looked at the code again and it looks like in cases where we would previously evict, in the new version we would just insert the missing one. That said, I do not see this as a correctness issue (not that it should not be still improved upon), so I would probably focus on 17248.","from":"developer"},{"body":"Hi, [~paulo] [~ifesdjeen] \r\n\r\nIt appears that we've encountered the exact issue [~chovatia.jaydeep@gmail.com] described here. During our production migration involving an {{ALTER TABLE ... ADD COLUMN}} operation on Cassandra 4.1.3, we experienced a severe reprepare storm — a situation where prepared statements are repeatedly invalidated and reprepared across cluster nodes, leading to significant performance degradation.\r\n \r\nCASSANDRA-17248, which was mentioned in the discussion, does not fully address the root cause of this problem (we did not perform a Cassandra version upgrade). After analyzing the behaviour, I believe the most effective long-term solution is to stop evicting the prepared statement cache.\r\n \r\nCan we revisit the proposed fix that implements this approach?\r\n[https://github.com/apache/cassandra/pull/3059]\r\n\r\n \r\n\r\n ","from":"developer"},{"body":"[~norkandrei] I have just raised this to the dev mailing list so someone can prioritize the review. In the meantime, you may want to continue with the private fix ([https://github.com/apache/cassandra/pull/3059]) until it is officially available in the repo. I have been using this fix in our setup, and it has been working fine.","from":"developer"},{"body":"I have looked around, and it looks like we rely on having both statements during upgrades to check if statements are cached correctly (perhaps we need to add a corresponding comment).\r\n\r\nI think we can still make it work though. I think we should just use {{preparedStatements.invalidateAll()}} instead of regular {{invalidate}}. Unfortunately we will be able to somehow synchronize deletes from stable storage, but I think this is fine, we can allow races there, since they're not correctness-impacting. Eventually we will save both statements with correct hashes.\r\n\r\nAnother addition to the patch I would do is adding an if statement when retrieving {{cachedWithKeyspace}} after {{cachedWithoutKeyspace}}, and only going for a second {{get}} if hashes are actually different. \r\n\r\nLMK if this approach still fixes the issue for you.","from":"developer"},{"body":"[~ifesdjeen] As I understand it, this eviction logic was introduced to handle a hash calculation change in 4.0.2 and is only needed during upgrades from pre-4.0.2 to 4.0.2. If all nodes are already on 4.0.2+, the eviction is unnecessary.\r\n\r\nReplacing _invalidate_ with _preparedStatements.invalidateAll_ does not address the issue in this ticket. Given this, we should apply the eviction logic only when pre-4.0.2 nodes are present and skip it otherwise. Happy to hear your thoughts.","from":"developer"},{"body":"I am personally not a big fan of adding version-dependent logic, since it complicates maintenance and testing substantially. I think we do need to replace invalidate with invalidateAll regardless of this ticket, since this is a race, too. Another thing I'd do is I'd check if removal listener is sufficient to clean up the prep statements table, and avoid explicitly calling {{removePreparedStatement}}. \r\n\r\nAnother approach is to evict only if we see that one of the statements is cached but not the other (i.e. {{ (a == null && b != null) && (b == null && a != null) }}), which is going to substantially reduce the window of race possibility. This, combined with what I wrote earlier, namely\r\n\r\nbq. Another addition to the patch I would do is adding an if statement when retrieving {{cachedWithKeyspace}} after {{cachedWithoutKeyspace}}, and only going for a second get if hashes are actually different.\r\n\r\nshould be sufficient to trigger only a handful of spurious re-prepares. WDYT?","from":"developer"},{"body":"> I am personally not a big fan of adding version-dependent logic, since it complicates maintenance and testing substantially.\r\n\r\nAs of today, in the trunk, we already have 4.0.1 -> 4.0.2 version-specific [logic|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L803]; please see the \"else\" portion.\r\n{code:java}\r\n            if (useNewPreparedStatementBehaviour)\r\n            {\r\n                if (cachedWithoutKeyspace.fullyQualified) // For fully qualified statements, we always skip keyspace to avoid digest switching\r\n                    return createResultMessage(hashWithoutKeyspace, cachedWithoutKeyspace);                if (clientState.getRawKeyspace() != null && !cachedWithKeyspace.fullyQualified) // For non-fully qualified statements, we always include keyspace to avoid ambiguity\r\n                    return createResultMessage(hashWithKeyspace, cachedWithKeyspace);            }\r\n            else // legacy caches, pre-CASSANDRA-15252 behaviour\r\n            {\r\n                return createResultMessage(hashWithKeyspace, cachedWithKeyspace);\r\n            } {code}\r\nThis eviction logic became necessary for 4.0.1 -> 4.0.2, and is not needed afterwards. Therefore, the proposed PR on _trunk_ removes the \"eviction\" logic entirely.\r\n\r\nAnd we can selectively decide till what version we want to land the proposed PR - I propose to land it on trunk and 5.0, maybe 4.1, but not land on 4.x.\r\n\r\n \r\n\r\n>Another approach is to evict only if we see that one of the statements is cached but not the other (i.e., \\{{ (a == null && b != null) && (b == null && a != null) }}), which is going to reduce the window of race possibility substantially. \r\n\r\nI agree that this logic reduces the possibility of a race, but it does not eliminate it. The point I am trying to make is not to keep unnecessary code/logic that is only needed for the transition from version 4.0.1 to 4.0.2.\r\n\r\nWDYT?","from":"developer"},{"body":"Alright, let's do that. I still think we should apply the patch to reduce the race window if you got time for that. ","from":"developer"},{"body":"Just to be clear, you are ok with the proposed PR [https://github.com/apache/cassandra/pull/3059] for trunk/5.0, i.e., no eviction.\r\n\r\nHowever, for the 4.0 branch, keep the eviction logic, and on top of that, you want me to reduce the probability window. Is that accurate? ","from":"developer"},{"body":"I fully support Jaydeepkumar Chovatia's proposal to remove the invalidation logic in versions 4.1, 5.0, and trunk, as implemented in PR [https://github.com/apache/cassandra/pull/3059 |https://github.com/apache/cassandra/pull/3059]. Although the idea to reduce the window of race possibility mitigates the issue, it does not completely eliminate the underlying risk.","from":"developer"},{"body":"[~chovatia.jaydeep@gmail.com] that's right: no eviction in trunk, in 4.0 predicate eviction on presence of pre-4.0.2 nodes (as you suggested), and, if you are not opposed, to add the suggestions that would improve the code in both situations.\r\n\r\nThank you!","from":"developer"},{"body":"Thanks [~ifesdjeen]. In that case, are you okay with reviewing/approving the proposed PR for trunk so I can formally merge it into the {_}trunk{_}? [https://github.com/apache/cassandra/pull/3059]\r\n\r\nI will create a new PR for the 4.0 branch and keep eviction + extra safety checks to minimize the issue mentioned in this ticket on the 4.0 branch.","from":"developer"},{"body":"[~ifesdjeen] I have also created a PR for the 4.0 branch (tests are currently running). Could you please take a look at it? [https://github.com/apache/cassandra/pull/4542]","from":"developer"},{"body":"Additionally, I request that the update removing the cache-clearing functionality ([https://github.com/apache/cassandra/pull/3059 |https://github.com/apache/cassandra/pull/3059]) be included in version 4.1.","from":"developer"},{"body":"[~ifesdjeen] I'd really appreciate it if you could find some time to review the PR and merge the PR — it addresses a much-needed fix.","from":"developer"},{"body":"> update removing the cache-clearing functionality be included in version 4.1.\r\nFrom what I understand, this is the version where it is actually needed: we need to clean up both hashes to preclude an incorrect version from being loaded, in case someone upgrades from 4.0.1 to 4.1.","from":"developer"},{"body":"I suppose that the possibility of upgrading from version 4.0.1 to 4.1 is significantly lower compared to the risk of encountering the reprepare storm problem in all 4.1.x versions. Moreover, adding an additional check before invalidation does not fully eliminate the root cause of the repreparation issue. Therefore, in my opinion, the best solution is to make this logic applicable only to version 4.0.","from":"developer"},{"body":"I have just committed the approved PR to {_}trunk{_}. [https://github.com/apache/cassandra/commit/a06df099f4e1c6265d25c94336c493aded3aa644] & https://github.com/apache/cassandra/commit/aa4da838ec0d05f7c7494536082271a6ed135fa1\r\n\r\nHere is the PR for 5.0 [https://github.com/jaydeepkumar1984/cassandra/pull/67] - please review.","from":"developer"},{"body":"[~chovatia.jaydeep@gmail.com], unless there's reason, PRs for all branches should be merged in one go, and pushed atomically.\r\nref: https://cassandra.apache.org/_/development/how_to_commit.html\r\n(there's no need to revert, just fyi on the project precedence)","from":"developer"},{"body":"Sure [~mck] - I will follow this practice going forward.","from":"developer"},{"body":"PR for 5.0 [https://github.com/jaydeepkumar1984/cassandra/pull/67] - please review and merge. Kindly reminder :)","from":"developer"},{"body":"[~ifesdjeen] Plz review PR for 5.0 [https://github.com/jaydeepkumar1984/cassandra/pull/67]","from":"developer"},{"body":"[~norkandrei] I have already reviewed this patch for trunk. I think committer can just merge it upwards. ","from":"developer"},{"body":"Thanks, Alex.\r\n\r\n[~norkandrei] I will port the patch to 5.0 sometime this weekend.","from":"developer"},{"body":"[~ifesdjeen] [~chovatia.jaydeep@gmail.com] Thank you!","from":"developer"},{"body":"The fix has also been landed in 5.0. SHA: [https://github.com/apache/cassandra/commit/d90fe76c8efb374c3b1b09e95f6e9925090fe638]\r\n\r\n[https://github.com/apache/cassandra/commit/4a2bf29a34cea273c5d23e0b9cd3c19ef0954853]\r\n\r\n \r\n\r\nFor the 4.x, I responded to the review comment from [~ifesdjeen] as 4.x change will be different than 5.0 and trunk; once that is approved, I will take care of 4.x too. [https://github.com/apache/cassandra/pull/4542/changes] ","from":"developer"},{"body":"[~chovatia.jaydeep@gmail.com] just looked at it again, and I agree with you. Thanks for elaborating! +1","from":"developer"},{"body":"Committed to 4.1 ([SHA|https://github.com/apache/cassandra/commit/4ada37e940431686cbc82e2d0b01f482f0041a3f]) & 4.0 ([SHA|https://github.com/apache/cassandra/commit/077b7ebe22c750b9f0c4a83f7979356b5095d6a5])","from":"developer"},{"body":"This fix has been landed to all the Cassandra versions: 4.0.20, 4.1.11, 5.0.7, and 5.1\r\n\r\nThanks [~ifesdjeen]  for your help with the code review!","from":"developer"}],"created":"2022-02-23T09:15:21.000+0000","description":"The changes in the [QueryProcessor#prepare|https://github.com/apache/cassandra/blame/cassandra-4.0.2/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L575-L638] method that were introduced in versions *4.0.2* and *3.11.12* can cause a race condition between two threads trying to concurrently prepare the same statement. This race condition can cause removing of a prepared statement from the cache, after one of the threads has received the result of the prepare and eventually uses MD5Digest to call [QueryProcessor#getPrepared|https://github.com/apache/cassandra/blame/cassandra-4.0.2/src/java/org/apache/cassandra/cql3/QueryProcessor.java#L212-L215].\r\n\r\nThe race condition looks like this:\r\n * Thread1 enters _prepare_ method and resolves _safeToReturnCached_ as false\r\n * Thread1 executes eviction of hashes\r\n * Thread2 enters _prepare_ method and resolves _safeToReturnCached_ as false\r\n * Thread1 prepares the statement and caches it\r\n * Thread1 returns the result of the prepare\r\n * Thread2 executes eviction of hashes\r\n * Thread1 tries to execute the prepared statement with the received MD5Digest, but statement is not in the cache as it was evicted by Thread2\r\n\r\nI tried to reproduce this by using a Java driver, but hitting this case from a client side is highly unlikely and I can not simulate the needed race condition. However, we can easily reproduce this in Stargate (details [here|https://github.com/stargate/stargate/pull/1647]), as it's closer to QueryProcessor.\r\n\r\nReproducing this in a unit test is fairly easy. I am happy to showcase this if needed.\r\n\r\nNote that the issue can occur only when  safeToReturnCached is resolved as false.","issue_id":"13430156","key":"CASSANDRA-17401","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2026-02-11T01:32:33.000+0000","role":"fixed_distractor","summary":"Race condition in QueryProcessor causes just prepared statement not to be in the prepared statements cache"} {"case_id":"13439525","cluster":"DISTRACTOR-CASSANDRA-17556","comments":[{"body":"||Branch||CI||\r\n|[3.11|https://github.com/driftx/cassandra/tree/CASSANDRA-17556-3.11]|[j8|https://app.circleci.com/pipelines/github/driftx/cassandra/435/workflows/bc986198-1eea-4ba4-9341-686e3ff68ccc]|\r\n|[4.0|https://github.com/driftx/cassandra/tree/CASSANDRA-17556-4.0]|[j8|https://app.circleci.com/pipelines/github/driftx/cassandra/434/workflows/77804032-ff82-4a01-9c6a-bfbbc4166f0b], [j11|https://app.circleci.com/pipelines/github/driftx/cassandra/434/workflows/eb11b25f-5802-47c8-a655-03cfda6feb36]|\r\n|[trunk|https://github.com/driftx/cassandra/tree/CASSANDRA-17556-trunk]|[j8|https://app.circleci.com/pipelines/github/driftx/cassandra/433/workflows/1908788a-7b92-4791-b9ea-946c2fe105cb], [j11|https://app.circleci.com/pipelines/github/driftx/cassandra/433/workflows/7e083a05-395d-4294-986f-a7c7985ffb3f]|\r\n","created":"2022-04-13T19:38:29.548+0000"},{"body":"Considering CASSANDRA-16851   maybe it will be a good idea to double-check if we have something to consider with [~tatu-at-datastax] / [~cowtowncoder] (not sure which user he uses now so pinged both :D )","created":"2022-04-13T21:47:05.927+0000"},{"body":"Alas, I don't think that ticket has bearing any longer since we forgot about it and upgraded to 2.13 in CASSANDRA-17492 without issue, also for security reasons.","created":"2022-04-14T10:43:10.420+0000"},{"body":"Yeah this would be covered by either tiny bump to 2.12.6.1 of jackson-databind, or, with 2.13.2.1 (or 2.13.2.2) of jackson-databind. I don't think Cassandra would actually be exposed by this CVE (unless I misremember limited usage there is for Jackson databind), but I know that sec scan tools have no concept of actual applicability to tell it \"no, it is actually not a problem\" so it tends to be easier to Just Upgrade That Dep to silence bogus warnings.\r\n\r\nI do check both of handles but it doesn't hurt using both so that's fine too :)\r\n\r\n ","created":"2022-04-18T18:05:11.289+0000"},{"body":"Thanks [~cowtowncoder],  the CVE will be covered by those versions, the question was whether again we might run into changes in those areas we use Jackson that can affect our performance while fixing security bugs. (Referring back to issues like the one we hit in  CASSANDRA-16851)\r\n\r\nAny changes we might want to know/consider?","created":"2022-04-18T19:39:16.792+0000"},{"body":"Performance changes (or at least negatives ones) are rather rare so that issue was an outlier.\r\nAlthough I can see how this specific type of change could seem suspicious to be sure.\r\n\r\nBut I do not think fix here should have any measurable performance impact; and since for 2.12.6.1 it is literally the only change beyond 2.12.6, that should be particularly safe.\r\n\r\nI wouldn't expect 2.12 -> 2.13 have meaningful performance difference either, but there are more changes that being a minor version bump.\r\n\r\nPut another way: I am not aware of any performance degradation between 2.12(.6) and 2.13(.2), nor expect there to be anything. But it is not possible to prove something does not exist, in general.\r\n\r\n \r\n\r\n ","created":"2022-04-18T20:25:58.569+0000"},{"body":"If I am reading this correctly we have to choose between a certain security vulnerability to C* json support ([here|https://cassandra.apache.org/doc/latest/cassandra/cql/json.html] and [here|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/cql3/Json.java]) vs an improbable performance impact. Even if the perf impact was a certain thing we'd still have to upgrade and fix it afterwards. Did I miss anything?","created":"2022-04-26T05:53:40.086+0000"},{"body":"I don't think so. We cannot release with OWASP failing, so there are only two paths here: upgrade, or determine that this isn't a problem and add a suppression. The latter is a lot simpler for when the component is something we clearly don't use, like Netty's HTTP implementations, where here I think there would always be a shadow of doubt. So in my view, upgrading is the likely future.","created":"2022-04-26T11:01:54.665+0000"},{"body":"I'm +1 on the upgrade as well. But happy to hear more opinions if there are any.","created":"2022-04-26T12:36:18.635+0000"},{"body":"My only question really was - is there anything we need to know around the usages or test and to remind of the old issue on the hot path. That’s it","created":"2022-04-26T13:27:00.946+0000"},{"body":"The usages I've found are auxiliary for logging with the CQL JSON support being the only one worth mentioning. I'm trying to think of more things we could do or check but nothing is coming up so far.","created":"2022-04-26T13:54:30.543+0000"},{"body":"Committed, thank you.","created":"2022-04-28T14:13:41.149+0000"},{"body":"I am late to the party but I really don't think there is a big change for significant performance degradation; beyond likelihood that any upgrade could have it.\r\nBut I understand that since there was one last time – highly unusual case fwtw – there is desire to be doubly sure there is due diligence before upgrade.\r\n\r\n ","created":"2022-04-28T21:28:15.660+0000"}],"conversations":[{"body":"Seems like it's technically possible to cause a DoS with nested json.","from":"reporter","subject":"jackson-databind 2.13.2 is vulnerable to CVE-2020-36518"},{"body":"||Branch||CI||\r\n|[3.11|https://github.com/driftx/cassandra/tree/CASSANDRA-17556-3.11]|[j8|https://app.circleci.com/pipelines/github/driftx/cassandra/435/workflows/bc986198-1eea-4ba4-9341-686e3ff68ccc]|\r\n|[4.0|https://github.com/driftx/cassandra/tree/CASSANDRA-17556-4.0]|[j8|https://app.circleci.com/pipelines/github/driftx/cassandra/434/workflows/77804032-ff82-4a01-9c6a-bfbbc4166f0b], [j11|https://app.circleci.com/pipelines/github/driftx/cassandra/434/workflows/eb11b25f-5802-47c8-a655-03cfda6feb36]|\r\n|[trunk|https://github.com/driftx/cassandra/tree/CASSANDRA-17556-trunk]|[j8|https://app.circleci.com/pipelines/github/driftx/cassandra/433/workflows/1908788a-7b92-4791-b9ea-946c2fe105cb], [j11|https://app.circleci.com/pipelines/github/driftx/cassandra/433/workflows/7e083a05-395d-4294-986f-a7c7985ffb3f]|\r\n","from":"developer"},{"body":"Considering CASSANDRA-16851   maybe it will be a good idea to double-check if we have something to consider with [~tatu-at-datastax] / [~cowtowncoder] (not sure which user he uses now so pinged both :D )","from":"developer"},{"body":"Alas, I don't think that ticket has bearing any longer since we forgot about it and upgraded to 2.13 in CASSANDRA-17492 without issue, also for security reasons.","from":"developer"},{"body":"Yeah this would be covered by either tiny bump to 2.12.6.1 of jackson-databind, or, with 2.13.2.1 (or 2.13.2.2) of jackson-databind. I don't think Cassandra would actually be exposed by this CVE (unless I misremember limited usage there is for Jackson databind), but I know that sec scan tools have no concept of actual applicability to tell it \"no, it is actually not a problem\" so it tends to be easier to Just Upgrade That Dep to silence bogus warnings.\r\n\r\nI do check both of handles but it doesn't hurt using both so that's fine too :)\r\n\r\n ","from":"developer"},{"body":"Thanks [~cowtowncoder],  the CVE will be covered by those versions, the question was whether again we might run into changes in those areas we use Jackson that can affect our performance while fixing security bugs. (Referring back to issues like the one we hit in  CASSANDRA-16851)\r\n\r\nAny changes we might want to know/consider?","from":"developer"},{"body":"Performance changes (or at least negatives ones) are rather rare so that issue was an outlier.\r\nAlthough I can see how this specific type of change could seem suspicious to be sure.\r\n\r\nBut I do not think fix here should have any measurable performance impact; and since for 2.12.6.1 it is literally the only change beyond 2.12.6, that should be particularly safe.\r\n\r\nI wouldn't expect 2.12 -> 2.13 have meaningful performance difference either, but there are more changes that being a minor version bump.\r\n\r\nPut another way: I am not aware of any performance degradation between 2.12(.6) and 2.13(.2), nor expect there to be anything. But it is not possible to prove something does not exist, in general.\r\n\r\n \r\n\r\n ","from":"developer"},{"body":"If I am reading this correctly we have to choose between a certain security vulnerability to C* json support ([here|https://cassandra.apache.org/doc/latest/cassandra/cql/json.html] and [here|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/cql3/Json.java]) vs an improbable performance impact. Even if the perf impact was a certain thing we'd still have to upgrade and fix it afterwards. Did I miss anything?","from":"developer"},{"body":"I don't think so. We cannot release with OWASP failing, so there are only two paths here: upgrade, or determine that this isn't a problem and add a suppression. The latter is a lot simpler for when the component is something we clearly don't use, like Netty's HTTP implementations, where here I think there would always be a shadow of doubt. So in my view, upgrading is the likely future.","from":"developer"},{"body":"I'm +1 on the upgrade as well. But happy to hear more opinions if there are any.","from":"developer"},{"body":"My only question really was - is there anything we need to know around the usages or test and to remind of the old issue on the hot path. That’s it","from":"developer"},{"body":"The usages I've found are auxiliary for logging with the CQL JSON support being the only one worth mentioning. I'm trying to think of more things we could do or check but nothing is coming up so far.","from":"developer"},{"body":"Committed, thank you.","from":"developer"},{"body":"I am late to the party but I really don't think there is a big change for significant performance degradation; beyond likelihood that any upgrade could have it.\r\nBut I understand that since there was one last time – highly unusual case fwtw – there is desire to be doubly sure there is due diligence before upgrade.\r\n\r\n ","from":"developer"}],"created":"2022-04-13T19:22:06.000+0000","description":"Seems like it's technically possible to cause a DoS with nested json.","issue_id":"13439525","key":"CASSANDRA-17556","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2022-04-28T14:13:41.000+0000","role":"fixed_distractor","summary":"jackson-databind 2.13.2 is vulnerable to CVE-2020-36518"} {"case_id":"13485353","cluster":"DISTRACTOR-CASSANDRA-17955","comments":[{"body":"This might be the solution (1). Basically, we need to make sure that a new snapshot is not taken until old snapshot is cleared. Snapshot name is UUID.toString() of \"parentSessionId\". Executor in ActiveRepairService is running snapshot cleanup in a non-blocking way. That executor can run 1 thread only at any given time. \r\n\r\nCassandraTableRepairManager takes an emphemeral snapshot and it might race in ActiveRepairService as a snapshot is being cleared but it expects the directory to be empty - but it is not, because CassandraTableRepairManager created a snapshot in it. \r\n\r\n(1) https://github.com/apache/cassandra/pull/1903/files","created":"2022-10-10T11:03:00.224+0000"},{"body":"[~marcuse] [~dcapwell] would you mind to take a look? I see you were involved in that part of the code lastly via git blame.\r\n\r\nThis is reported to be quite a big issue, we have a case of 200 nodes cluster where 80 nodes across 3 dcs hit this problem.\r\n\r\nHaving this in 4.1 GA would be really great. Isn't this actually a blocker?","created":"2022-10-10T11:06:28.552+0000"},{"body":"https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1990/","created":"2022-10-10T20:47:50.920+0000"},{"body":"Patch makes sense to me, we have a single threaded executor that would do cleanup, but now also added snapshot, so this avoids the 2 actions overlapping...\r\n\r\nI am not sure how I feel about snapshots now being single threaded... this could be an issue for some users maybe... need to think more about it.\r\n\r\n[~smiklosovic] can you add a test to replicate the race condition? ","created":"2022-10-10T21:19:48.355+0000"},{"body":"I do not have a reproducer yet. I will try to do one but my gut feeling is that it wont be so easy. I checked the tests and I think you did one for repairs when one node went down. I might take that as a base and refactor it maybe.\n\n\nSent from ProtonMail mobile\n\n\n\n\\","created":"2022-10-10T23:10:00.022+0000"},{"body":"[~dcapwell] I am letting you know that I am not trying to develop the reproducer at the moment. I think we need to go without if somebody else does not write it.","created":"2022-10-13T10:31:10.025+0000"},{"body":"[~dcapwell] any progress here? Have you been thinking about consequences of having single threaded snapshots during repairs?","created":"2022-10-21T09:37:18.241+0000"},{"body":"I don't see any other than its slow if you try to do a lot at once... so LGTM\r\n\r\n+1","created":"2022-10-21T20:37:03.266+0000"},{"body":"Snaps are just hard links so it should still be sufficiently fast imo.","created":"2022-10-21T20:41:37.536+0000"},{"body":"So clearSnapshotExecutor was added as part of CASSANDRA-17168 . Instead of adding more tasks to this executor why not pass it back to \r\n\r\nANTI_ENTROPY stage. Snapshotting messages already take place through this stage, and I think its cleaner and makes more sense to clear the snapshot there. Note this stage already is only 1 thread so concerns about single threaded snapshots has always been the case. As [~brandon.williams] mentioned its mostly just hard linking (other then it having to block for a flush). The argument for having more threads could be made but I think that is seperate to getting clearing and creating snapshots happening in the same thread pool.","created":"2022-10-21T22:48:27.504+0000"},{"body":"Tagging [~marcuse] since you are author of CASSANDRA-17168 ","created":"2022-10-21T23:24:39.980+0000"},{"body":"Haven't looked at this in detail, but purpose of 17168 was just to avoid doing this on GossipTasks, so as long as its not there we're fine","created":"2022-10-24T11:46:50.161+0000"},{"body":"I run pipeline with the proposed approach (running it in Stage) and while it seems to be ok in 4.0, it does not work in 4.1, there are dozens of errors, reproducible locally too. It seems like it is stuck and then it just timeouts. \r\n\r\nI do not have a lot of time to investigate what is going on here, I run the original patch and passes the pipeline just fine so I will go with that one.\r\n\r\nI ll formally prepare all the branches and builds and merge the original and already approved approach.","created":"2022-10-25T11:10:18.674+0000"},{"body":"4.0\r\nj11 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1488/workflows/02fccb66-30fb-4ce3-ac75-6f275221af94\r\nj8 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1488/workflows/08e02b33-6be5-4de9-8801-f725077bdca1\r\n4.1\r\nj11 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1487/workflows/ad3a7fb8-c84d-48d7-ae0c-6094cc6a1a21\r\nj8 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1487/workflows/d5c56237-776e-4f81-92be-87bec32cfe1a\r\ntrunk\r\nj11 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1486/workflows/9a714592-24e4-415a-951d-e143ec73ff50\r\nj8 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1486/workflows/a1b9918c-e055-41b4-94c8-d47d90068aa1\r\n\r\nEverything is green, I ve never seen 6 green builds in a row tbh.\r\n\r\nWhat I want to do now is to do run 500x on all repair related tests (python dtests included).","created":"2022-10-26T12:14:27.945+0000"},{"body":"As mentioned, I tried to run multiplexer on all repair tests, I can use only 20 runners and CI job timeouts after 1 hour so I tried to measure the maximum amount of repeats over all repair tests. I think I run all repair unit tests around 120 times and all dtests 20 times it went all fine. This is only the build of trunk (1).\r\n\r\nThis patch is not introducing any new test nor it modifies any but I still tried to run all repair tests in a loop to see if it is stable, which it seems it is. Due to limited resources and time constraints I consider this kind of testing enough (on top of regular and mandatory 6 jobs above, 2 per branch (8 and 11 pre-commit)\r\n\r\nhttps://app.circleci.com/pipelines/github/instaclustr/cassandra/1498/workflows/e0f4e61a-cf3b-4ef7-a0c3-d02a25778bb8","created":"2022-10-27T13:15:20.855+0000"}],"conversations":[{"body":"If an endpoint is convicted and that endpoint is a coordinator then ActiveRepairService::removeParentRepairSession is called.\r\n\r\nThe issue is that this occurs on clearSnapshotExecutor and can happen while RepairMessageVerbHandler is in process of taking a snapshot. So then you get a race condition and clearSnapshot will throw a java.nio.file.DirectoryNotEmptyException\r\n\r\n \r\n{code:java}\r\npublic static void deleteRecursiveWithThrottle(File dir, RateLimiter rateLimiter)\r\n{\r\n if (dir.isDirectory())\r\n {\r\n String[] children = dir.list();\r\n for (String child : children)\r\n deleteRecursiveWithThrottle(new File(dir, child), rateLimiter);\r\n }\r\n\r\n // The directory is now empty so now it can be smoked\r\n deleteWithConfirmWithThrottle(dir, rateLimiter);\r\n} {code}\r\nDue to the directory not being empty when it goes to remove the directory at the end.","from":"reporter","subject":"Race condition on repair snapshots"},{"body":"This might be the solution (1). Basically, we need to make sure that a new snapshot is not taken until old snapshot is cleared. Snapshot name is UUID.toString() of \"parentSessionId\". Executor in ActiveRepairService is running snapshot cleanup in a non-blocking way. That executor can run 1 thread only at any given time. \r\n\r\nCassandraTableRepairManager takes an emphemeral snapshot and it might race in ActiveRepairService as a snapshot is being cleared but it expects the directory to be empty - but it is not, because CassandraTableRepairManager created a snapshot in it. \r\n\r\n(1) https://github.com/apache/cassandra/pull/1903/files","from":"developer"},{"body":"[~marcuse] [~dcapwell] would you mind to take a look? I see you were involved in that part of the code lastly via git blame.\r\n\r\nThis is reported to be quite a big issue, we have a case of 200 nodes cluster where 80 nodes across 3 dcs hit this problem.\r\n\r\nHaving this in 4.1 GA would be really great. Isn't this actually a blocker?","from":"developer"},{"body":"https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch/1990/","from":"developer"},{"body":"Patch makes sense to me, we have a single threaded executor that would do cleanup, but now also added snapshot, so this avoids the 2 actions overlapping...\r\n\r\nI am not sure how I feel about snapshots now being single threaded... this could be an issue for some users maybe... need to think more about it.\r\n\r\n[~smiklosovic] can you add a test to replicate the race condition? ","from":"developer"},{"body":"I do not have a reproducer yet. I will try to do one but my gut feeling is that it wont be so easy. I checked the tests and I think you did one for repairs when one node went down. I might take that as a base and refactor it maybe.\n\n\nSent from ProtonMail mobile\n\n\n\n\\","from":"developer"},{"body":"[~dcapwell] I am letting you know that I am not trying to develop the reproducer at the moment. I think we need to go without if somebody else does not write it.","from":"developer"},{"body":"[~dcapwell] any progress here? Have you been thinking about consequences of having single threaded snapshots during repairs?","from":"developer"},{"body":"I don't see any other than its slow if you try to do a lot at once... so LGTM\r\n\r\n+1","from":"developer"},{"body":"Snaps are just hard links so it should still be sufficiently fast imo.","from":"developer"},{"body":"So clearSnapshotExecutor was added as part of CASSANDRA-17168 . Instead of adding more tasks to this executor why not pass it back to \r\n\r\nANTI_ENTROPY stage. Snapshotting messages already take place through this stage, and I think its cleaner and makes more sense to clear the snapshot there. Note this stage already is only 1 thread so concerns about single threaded snapshots has always been the case. As [~brandon.williams] mentioned its mostly just hard linking (other then it having to block for a flush). The argument for having more threads could be made but I think that is seperate to getting clearing and creating snapshots happening in the same thread pool.","from":"developer"},{"body":"Tagging [~marcuse] since you are author of CASSANDRA-17168 ","from":"developer"},{"body":"Haven't looked at this in detail, but purpose of 17168 was just to avoid doing this on GossipTasks, so as long as its not there we're fine","from":"developer"},{"body":"I run pipeline with the proposed approach (running it in Stage) and while it seems to be ok in 4.0, it does not work in 4.1, there are dozens of errors, reproducible locally too. It seems like it is stuck and then it just timeouts. \r\n\r\nI do not have a lot of time to investigate what is going on here, I run the original patch and passes the pipeline just fine so I will go with that one.\r\n\r\nI ll formally prepare all the branches and builds and merge the original and already approved approach.","from":"developer"},{"body":"4.0\r\nj11 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1488/workflows/02fccb66-30fb-4ce3-ac75-6f275221af94\r\nj8 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1488/workflows/08e02b33-6be5-4de9-8801-f725077bdca1\r\n4.1\r\nj11 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1487/workflows/ad3a7fb8-c84d-48d7-ae0c-6094cc6a1a21\r\nj8 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1487/workflows/d5c56237-776e-4f81-92be-87bec32cfe1a\r\ntrunk\r\nj11 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1486/workflows/9a714592-24e4-415a-951d-e143ec73ff50\r\nj8 precommit https://app.circleci.com/pipelines/github/instaclustr/cassandra/1486/workflows/a1b9918c-e055-41b4-94c8-d47d90068aa1\r\n\r\nEverything is green, I ve never seen 6 green builds in a row tbh.\r\n\r\nWhat I want to do now is to do run 500x on all repair related tests (python dtests included).","from":"developer"},{"body":"As mentioned, I tried to run multiplexer on all repair tests, I can use only 20 runners and CI job timeouts after 1 hour so I tried to measure the maximum amount of repeats over all repair tests. I think I run all repair unit tests around 120 times and all dtests 20 times it went all fine. This is only the build of trunk (1).\r\n\r\nThis patch is not introducing any new test nor it modifies any but I still tried to run all repair tests in a loop to see if it is stable, which it seems it is. Due to limited resources and time constraints I consider this kind of testing enough (on top of regular and mandatory 6 jobs above, 2 per branch (8 and 11 pre-commit)\r\n\r\nhttps://app.circleci.com/pipelines/github/instaclustr/cassandra/1498/workflows/e0f4e61a-cf3b-4ef7-a0c3-d02a25778bb8","from":"developer"}],"created":"2022-10-10T00:39:55.000+0000","description":"If an endpoint is convicted and that endpoint is a coordinator then ActiveRepairService::removeParentRepairSession is called.\r\n\r\nThe issue is that this occurs on clearSnapshotExecutor and can happen while RepairMessageVerbHandler is in process of taking a snapshot. So then you get a race condition and clearSnapshot will throw a java.nio.file.DirectoryNotEmptyException\r\n\r\n \r\n{code:java}\r\npublic static void deleteRecursiveWithThrottle(File dir, RateLimiter rateLimiter)\r\n{\r\n if (dir.isDirectory())\r\n {\r\n String[] children = dir.list();\r\n for (String child : children)\r\n deleteRecursiveWithThrottle(new File(dir, child), rateLimiter);\r\n }\r\n\r\n // The directory is now empty so now it can be smoked\r\n deleteWithConfirmWithThrottle(dir, rateLimiter);\r\n} {code}\r\nDue to the directory not being empty when it goes to remove the directory at the end.","issue_id":"13485353","key":"CASSANDRA-17955","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2022-10-27T14:52:09.000+0000","role":"fixed_distractor","summary":"Race condition on repair snapshots"} {"case_id":"13520058","cluster":"DISTRACTOR-CASSANDRA-18176","comments":[{"body":"You could enable {{-Dcassandra.debugrefcount=true}} - this would have some process overhead, but it should find any resource leak of this kind.","created":"2023-01-18T13:48:07.049+0000"},{"body":"Thanks for the pointer [~benedict] , I'll see if I can add it \"safely\".","created":"2023-01-18T14:01:15.353+0000"},{"body":"Finally managed to enable it in one node, keeping an eye out for leaks. I'd say though, _this would have some process overhead_ may be a bit of an understatement, load average jumped from 2-3 to 10-11 in a 14 core machine","created":"2023-02-09T15:51:02.380+0000"},{"body":"Found quite a few of these, anything else I should look out for that may be interesting [~benedict]?\r\n{noformat}\r\nERROR [Strong-Reference-Leak-Detector:1] 2023-02-10 09:24:08,119 NoSpamLogger.java:98 - Strong self-ref loop detected \r\n[/cassandra-data/data/system_distributed/repair_history-759fffad624b318180eefa9a52d1f627/nb-628647-big, private org.apache.cassandra.utils.concurrent.Ref org.apache.cassandra.io.sstable.format.SSTableReader$InstanceTidier.globalRef-org.apache.cassandra.utils.concurrent.Ref, final org.apache.cassandra.utils.concurrent.Ref$State org.apache.cassandra.utils.concurrent.Ref.state-org.apache.cassandra.utils.concurrent.Ref$State, final org.apache.cassandra.utils.concurrent.Ref$GlobalState org.apache.cassandra.utils.concurrent.Ref$State.globalState-org.apache.cassandra.utils.concurrent.Ref$GlobalState, private final org.apache.cassandra.utils.concurrent.RefCounted$Tidy org.apache.cassandra.utils.concurrent.Ref$GlobalState.tidy-org.apache.cassandra.io.sstable.format.SSTableReader$GlobalTidy, private volatile java.lang.Runnable org.apache.cassandra.io.sstable.format.SSTableReader$GlobalTidy.obsoletion-org.apache.cassandra.db.lifecycle.LogTransaction$SSTableTidier, private final org.apache.cassandra.db.lifecycle.Tracker org.apache.cassandra.db.lifecycle.LogTransaction$SSTableTidier.tracker-org.apache.cassandra.db.lifecycle.Tracker, private final java.util.Collection org.apache.cassandra.db.lifecycle.Tracker.subscribers-java.util.concurrent.CopyOnWriteArrayList, private final java.util.Collection org.apache.cassandra.db.lifecycle.Tracker.subscribers-org.apache.cassandra.service.SSTablesGlobalTracker, private final java.util.Set org.apache.cassandra.service.SSTablesGlobalTracker.subscribers-java.util.concurrent.CopyOnWriteArraySet, private final java.util.Set org.apache.cassandra.service.SSTablesGlobalTracker.subscribers-org.apache.cassandra.service.StorageService$$Lambda$883/146242186, private final org.apache.cassandra.service.StorageService org.apache.cassandra.service.StorageService$$Lambda$883/146242186.arg$1-org.apache.cassandra.service.StorageService, private java.lang.Thread org.apache.cassandra.service.StorageService.drainOnShutdown-io.netty.util.concurrent.FastThreadLocalThread, private java.lang.ThreadGroup java.lang.Thread.group-java.lang.ThreadGroup, private final java.lang.ThreadGroup java.lang.ThreadGroup.parent-java.lang.ThreadGroup, java.lang.Thread[] java.lang.ThreadGroup.threads-[Ljava.lang.Thread;, java.lang.Thread[] java.lang.ThreadGroup.threads-java.lang.Thread, private java.lang.Runnable java.lang.Thread.target-java.util.concurrent.ThreadPoolExecutor$Worker, final java.util.concurrent.ThreadPoolExecutor java.util.concurrent.ThreadPoolExecutor$Worker.this$0-java.util.concurrent.ScheduledThreadPoolExecutor, private final java.util.concurrent.BlockingQueue java.util.concurrent.ThreadPoolExecutor.workQueue-java.util.concurrent.ScheduledThreadPoolExecutor$DelayedWorkQueue, private final java.util.concurrent.BlockingQueue java.util.concurrent.ThreadPoolExecutor.workQueue-java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask, private java.util.concurrent.Callable java.util.concurrent.FutureTask.callable-java.util.concurrent.Executors$RunnableAdapter, final java.lang.Runnable java.util.concurrent.Executors$RunnableAdapter.task-sun.rmi.transport.DGCImpl$1, final sun.rmi.transport.DGCImpl sun.rmi.transport.DGCImpl$1.this$0-sun.rmi.transport.DGCImpl, private java.util.Map sun.rmi.transport.DGCImpl.leaseTable-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-sun.rmi.transport.DGCImpl$LeaseInfo, java.util.Set sun.rmi.transport.DGCImpl$LeaseInfo.notifySet-java.util.HashSet, private transient java.util.HashMap java.util.HashSet.map-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, final java.lang.Object java.util.HashMap$Node.key-sun.rmi.transport.Target, private final sun.rmi.transport.WeakRef sun.rmi.transport.Target.weakImpl-sun.rmi.transport.WeakRef, private java.lang.Object sun.rmi.transport.WeakRef.strongRef-javax.management.remote.rmi.RMIJRMPServerImpl, private javax.management.MBeanServer javax.management.remote.rmi.RMIServerImpl.mbeanServer-com.sun.jmx.mbeanserver.JmxMBeanServer, private volatile javax.management.MBeanServer com.sun.jmx.mbeanserver.JmxMBeanServer.mbsInterceptor-com.zegelin.cassandra.exporter.MBeanServerInterceptorHarvester$MBeanServerInterceptor, private final javax.management.MBeanServer com.zegelin.jmx.DelegatingMBeanServerInterceptor.delegate-com.sun.jmx.interceptor.DefaultMBeanServerInterceptor, private final transient com.sun.jmx.mbeanserver.Repository com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.repository-com.sun.jmx.mbeanserver.Repository, private final java.util.Map com.sun.jmx.mbeanserver.Repository.domainTb-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.util.HashMap$Node java.util.HashMap$Node.next-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-com.sun.jmx.mbeanserver.NamedObject, private final javax.management.DynamicMBean com.sun.jmx.mbeanserver.NamedObject.object-com.sun.jmx.mbeanserver.StandardMBeanSupport, private final java.lang.Object com.sun.jmx.mbeanserver.MBeanSupport.resource-org.apache.cassandra.db.compaction.CompactionManager, public final org.apache.cassandra.db.compaction.ActiveCompactions org.apache.cassandra.db.compaction.CompactionManager.active-org.apache.cassandra.db.compaction.ActiveCompactions, private final java.util.Set org.apache.cassandra.db.compaction.ActiveCompactions.compactions-java.util.Collections$SynchronizedSet, final java.util.Collection java.util.Collections$SynchronizedCollection.c-java.util.Collections$SetFromMap, private final java.util.Map java.util.Collections$SetFromMap.m-java.util.IdentityHashMap, transient java.lang.Object[] java.util.IdentityHashMap.table-[Ljava.lang.Object;, transient java.lang.Object[] java.util.IdentityHashMap.table-org.apache.cassandra.db.compaction.CompactionIterator, private final org.apache.cassandra.db.AbstractCompactionController org.apache.cassandra.db.compaction.CompactionIterator.controller-org.apache.cassandra.db.compaction.TimeWindowCompactionController, private org.apache.cassandra.utils.concurrent.Refs org.apache.cassandra.db.compaction.CompactionController.overlappingSSTables-org.apache.cassandra.utils.concurrent.Refs, private final java.util.Map org.apache.cassandra.utils.concurrent.Refs.references-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, final java.lang.Object java.util.HashMap$Node.key-org.apache.cassandra.io.sstable.format.big.BigTableReader, private final org.apache.cassandra.utils.concurrent.Ref org.apache.cassandra.io.sstable.format.SSTableReader.selfRef-org.apache.cassandra.utils.concurrent.Ref]\r\n{noformat}\r\n{noformat}\r\nERROR [Strong-Reference-Leak-Detector:1] 2023-02-10 09:26:30,587 NoSpamLogger.java:98 - Strong self-ref loop detected \r\n[/cassandra-data/data/system_distributed/repair_history-759fffad624b318180eefa9a52d1f627/nb-628650-big, private volatile java.lang.Runnable org.apache.cassandra.io.sstable.format.SSTableReader$GlobalTidy.obsoletion-org.apache.cassandra.db.lifecycle.LogTransaction$SSTableTidier, private final org.apache.cassandra.db.lifecycle.Tracker org.apache.cassandra.db.lifecycle.LogTransaction$SSTableTidier.tracker-org.apache.cassandra.db.lifecycle.Tracker, private final java.util.Collection org.apache.cassandra.db.lifecycle.Tracker.subscribers-java.util.concurrent.CopyOnWriteArrayList, private final java.util.Collection org.apache.cassandra.db.lifecycle.Tracker.subscribers-org.apache.cassandra.service.SSTablesGlobalTracker, private final java.util.Set org.apache.cassandra.service.SSTablesGlobalTracker.subscribers-java.util.concurrent.CopyOnWriteArraySet, private final java.util.Set org.apache.cassandra.service.SSTablesGlobalTracker.subscribers-org.apache.cassandra.service.StorageService$$Lambda$883/146242186, private final org.apache.cassandra.service.StorageService org.apache.cassandra.service.StorageService$$Lambda$883/146242186.arg$1-org.apache.cassandra.service.StorageService, private java.lang.Thread org.apache.cassandra.service.StorageService.drainOnShutdown-io.netty.util.concurrent.FastThreadLocalThread, private java.lang.ThreadGroup java.lang.Thread.group-java.lang.ThreadGroup, private final java.lang.ThreadGroup java.lang.ThreadGroup.parent-java.lang.ThreadGroup, java.lang.Thread[] java.lang.ThreadGroup.threads-[Ljava.lang.Thread;, java.lang.Thread[] java.lang.ThreadGroup.threads-java.lang.Thread, private java.lang.Runnable java.lang.Thread.target-java.util.concurrent.ThreadPoolExecutor$Worker, final java.util.concurrent.ThreadPoolExecutor java.util.concurrent.ThreadPoolExecutor$Worker.this$0-java.util.concurrent.ScheduledThreadPoolExecutor, private final java.util.concurrent.BlockingQueue java.util.concurrent.ThreadPoolExecutor.workQueue-java.util.concurrent.ScheduledThreadPoolExecutor$DelayedWorkQueue, private final java.util.concurrent.BlockingQueue java.util.concurrent.ThreadPoolExecutor.workQueue-java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask, private java.util.concurrent.Callable java.util.concurrent.FutureTask.callable-java.util.concurrent.Executors$RunnableAdapter, final java.lang.Runnable java.util.concurrent.Executors$RunnableAdapter.task-sun.rmi.transport.DGCImpl$1, final sun.rmi.transport.DGCImpl sun.rmi.transport.DGCImpl$1.this$0-sun.rmi.transport.DGCImpl, private java.util.Map sun.rmi.transport.DGCImpl.leaseTable-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-sun.rmi.transport.DGCImpl$LeaseInfo, java.util.Set sun.rmi.transport.DGCImpl$LeaseInfo.notifySet-java.util.HashSet, private transient java.util.HashMap java.util.HashSet.map-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, final java.lang.Object java.util.HashMap$Node.key-sun.rmi.transport.Target, private final sun.rmi.transport.WeakRef sun.rmi.transport.Target.weakImpl-sun.rmi.transport.WeakRef, private java.lang.Object sun.rmi.transport.WeakRef.strongRef-javax.management.remote.rmi.RMIJRMPServerImpl, private javax.management.MBeanServer javax.management.remote.rmi.RMIServerImpl.mbeanServer-com.sun.jmx.mbeanserver.JmxMBeanServer, private volatile javax.management.MBeanServer com.sun.jmx.mbeanserver.JmxMBeanServer.mbsInterceptor-com.zegelin.cassandra.exporter.MBeanServerInterceptorHarvester$MBeanServerInterceptor, private final javax.management.MBeanServer com.zegelin.jmx.DelegatingMBeanServerInterceptor.delegate-com.sun.jmx.interceptor.DefaultMBeanServerInterceptor, private final transient com.sun.jmx.mbeanserver.Repository com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.repository-com.sun.jmx.mbeanserver.Repository, private final java.util.Map com.sun.jmx.mbeanserver.Repository.domainTb-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.util.HashMap$Node java.util.HashMap$Node.next-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-com.sun.jmx.mbeanserver.NamedObject, private final javax.management.DynamicMBean com.sun.jmx.mbeanserver.NamedObject.object-com.sun.jmx.mbeanserver.StandardMBeanSupport, private final java.lang.Object com.sun.jmx.mbeanserver.MBeanSupport.resource-org.apache.cassandra.db.compaction.CompactionManager, public final org.apache.cassandra.db.compaction.ActiveCompactions org.apache.cassandra.db.compaction.CompactionManager.active-org.apache.cassandra.db.compaction.ActiveCompactions, private final java.util.Set org.apache.cassandra.db.compaction.ActiveCompactions.compactions-java.util.Collections$SynchronizedSet, final java.util.Collection java.util.Collections$SynchronizedCollection.c-java.util.Collections$SetFromMap, private final java.util.Map java.util.Collections$SetFromMap.m-java.util.IdentityHashMap, transient java.lang.Object[] java.util.IdentityHashMap.table-[Ljava.lang.Object;, transient java.lang.Object[] java.util.IdentityHashMap.table-org.apache.cassandra.db.compaction.CompactionIterator, private final org.apache.cassandra.db.AbstractCompactionController org.apache.cassandra.db.compaction.CompactionIterator.controller-org.apache.cassandra.db.compaction.TimeWindowCompactionController, private org.apache.cassandra.utils.concurrent.Refs org.apache.cassandra.db.compaction.CompactionController.overlappingSSTables-org.apache.cassandra.utils.concurrent.Refs, private final java.util.Map org.apache.cassandra.utils.concurrent.Refs.references-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, final java.lang.Object java.util.HashMap$Node.key-org.apache.cassandra.io.sstable.format.big.BigTableReader, private final org.apache.cassandra.io.sstable.format.SSTableReader$InstanceTidier org.apache.cassandra.io.sstable.format.SSTableReader.tidy-org.apache.cassandra.io.sstable.format.SSTableReader$InstanceTidier, private org.apache.cassandra.utils.concurrent.Ref org.apache.cassandra.io.sstable.format.SSTableReader$InstanceTidier.globalRef-org.apache.cassandra.utils.concurrent.Ref]\r\n {noformat}","created":"2023-02-10T11:17:07.551+0000"},{"body":"Thanks [~pbalaguer]. I think probably this bug is a combination of a strong ref leak and a regular leak, with the former masking the latter as it prevents our regular leak detection logic kicking in. \r\n\r\nHowever, once we fix the former, we will hopefully see the latter in our usual test suite (though we should really be finding the former in our test suites too, so we need to look into why we aren't seeing that).","created":"2023-02-10T11:26:35.915+0000"},{"body":"I would say you can definitely disable the additional logging on your production systems at least for now. Thanks.","created":"2023-02-10T11:40:59.670+0000"},{"body":"Ok, this particular (strong ref loop) issue is a duplicate of CASSANDRA-17205. [~jmckenzie] from that ticket it looks like we didn't backport the fix for some reason?\r\n\r\n","created":"2023-02-10T11:44:15.502+0000"},{"body":"Understood, I'll proceed to disable it. Given the automation has been updated, if this needs to be re-enabled it'll be a matter of minutes.","created":"2023-02-10T11:49:05.582+0000"},{"body":"{quote}Josh McKenzie from that ticket it looks like we didn't backport the fix for some reason?\r\n{quote}\r\n... I'm at a loss. I don't even remember working on that _at all_ much less why I didn't backport it.\r\n\r\nI'll massage this ticket's metadata in a bit, take ownership, and backport CASSANDRA-17205 into the 4.0 line.","created":"2023-02-10T16:23:44.763+0000"},{"body":"For posterity (i.e. my future Not Forgetting-ness), Brandon pointed out this might need to be backported to all supported branches. I'll take a look when I get to this and handle that.","created":"2023-02-10T16:30:12.752+0000"},{"body":"I was probably unclear, but I do not think this bug report is a duplicate of CASSANDRA-17205, only that this earlier bug is masking whatever the underlying bug is here.","created":"2023-02-10T16:30:22.807+0000"},{"body":"Ah. Hm. So another JIRA for the backports as a blocker for this may well be the appropriate move here. Sorry for the noise / pollution.","created":"2023-02-10T16:33:03.462+0000"},{"body":"As part of any backport work, do you want to perhaps take a look at why we aren't seeing the strong loop detection errors anywhere in our CI, or if we are why they aren't being surfaced? Or perhaps just create a dedicated strong loop test that makes sure to run the detector at least once after running some basic tasks?\r\n\r\nThis is a pretty serious issue really, as it completely disables all of our leak detection which is obviously very important. So we want to catch mistakes here as early as possible.","created":"2023-02-10T16:38:44.761+0000"},{"body":"bq. do you want to perhaps take a look at why we aren't seeing the strong loop detection errors anywhere in our CI, or if we are why they aren't being surfaced?\r\nJust so happens that I've had a task in my backlog for awhile to surface and fail out tests if we have ref leaks on them, just haven't gotten to that yet.\r\n\r\nI'll be pivoting to focus on CI related things in a month or so so I'll pencil in rolling that up w/the rest of the effort. Completely agree on the importance of it.","created":"2023-02-10T16:49:16.105+0000"},{"body":"Perfect. Thanks","created":"2023-02-10T16:50:48.107+0000"},{"body":"Anything logged at ERROR would surface in CI (at least off the top of my head I know the python dtests will complain) so I have to think they aren't being logged or the problem isn't being triggered.","created":"2023-02-10T16:52:07.350+0000"},{"body":"Looks like strong leaks are at WARN. We should bump to ERROR I guess.\r\n\r\nAlso, the strong loop detector can take a while to run as it depends on reflection, and doesn't really know where to focus its efforts. So it might be many tests are any way much to brief. I'm not 100% certain we run with debugrefcount for all our tests either. We should make sure that we do.","created":"2023-02-10T16:54:37.818+0000"},{"body":"bq. I'm not 100% certain we run with debugrefcount for all our tests \r\n\r\nIf we do, I don't see it anywhere.","created":"2023-02-10T16:56:57.562+0000"},{"body":"It's in {{testmacrohelper}} as one of the jvmargs. So we should be running with it ({{-Dcassandra.debugrefcount=true}})","created":"2023-02-10T16:58:17.432+0000"},{"body":"bq. Looks like strong leaks are at WARN. We should bump to ERROR I guess.\r\n\r\nI think we do log at [ERROR|https://github.com/apache/cassandra/blob/cassandra-4.0/src/java/org/apache/cassandra/utils/concurrent/Ref.java#L600] so something else is going on here.","created":"2023-02-10T16:59:51.610+0000"},{"body":"That's the reference leak, not a \"strong reference leak\" - which is perhaps a misnomer. There's a separate feature to search the object hierarchy for strong reference loops that prevent {{Ref}} from printing those ERROR messages, i.e. it looks for us \"leaking\" a strong reference to a {{Ref}} into its {{Tidy}}, so that it is not garbage collected (and so, if we fail to release it, our leak detection will not inform us).\r\n\r\nThis feature logs at [WARN|https://github.com/apache/cassandra/blob/5d3c747719f01d87c9086c405c806317405a8e43/src/java/org/apache/cassandra/utils/concurrent/Ref.java#L711]\r\n\r\nIt's worth noting that it's an assumption that there's a regular leak in this ticket. It might be something else going on. But we won't know until our leak detection is working.","created":"2023-02-10T17:04:32.355+0000"},{"body":"bq. This feature logs at WARN\r\n\r\nI see. I agree we should bump that up to ERROR.","created":"2023-02-10T17:44:48.920+0000"},{"body":"Hey folks, sorry for bumping this but the issue is rather annoying when there are a lot of ongoing ring movements as it triggers quite often forcing node restarts which in turn delay everything. Do you happen to have an idea on when can this be worked on and if I can do something to contribute a bit more to make it happen sooner? (not so sure about deep down diving and coding but hey, maybe).","created":"2023-03-13T09:51:18.490+0000"},{"body":"Sorry [~pbalaguer], I have a lot on and had mentally handed this off to others for the moment.\r\n\r\nI could very easily back port this to a version you are running (please nominate one), and post a branch to GitHub that you could build and deploy with the strong reference loop only fixed, so that you could enable logging so that we might be able to track down the underlying resource leak.\r\n\r\nIf you're happy with this course of action, I'll post a branch either today or tomorrow. Otherwise, we can prod [~jmckenzie] and see about when we can get it properly backported and merged into a release.\r\n\r\nI agree that we should get to the bottom of this sooner than later, for the project as well as yourself.","created":"2023-03-13T09:56:10.863+0000"},{"body":"Hey [~benedict] , if that's the situation on your end don't worry, I'll try to backport it myself and see how it goes from there. If I can't figure it out I'll ping here again.\r\n\r\nThis effort is not meant to stop nor delay [~jmckenzie] but to happen in parallel.","created":"2023-03-13T12:12:31.235+0000"},{"body":"Yikes; this one slipped through the cracks on my todo list. Created CASSANDRA-18332 for the backport; should be trivial to do tomorrow morning.\r\n\r\nSorry about that.","created":"2023-03-14T20:45:14.981+0000"},{"body":"CASSANDRA-18332 will be in 4.0.9. Do we need to do anything else in this ticket, [~pbalaguer] ? ","created":"2023-04-12T13:18:24.858+0000"},{"body":"I don't think we've definitively solved anything here, but perhaps only added debugging that can tell us.","created":"2023-04-12T14:03:48.676+0000"},{"body":"Let's wait for more debugging then. Looking forward what Pere reports.","created":"2023-04-12T14:05:51.681+0000"},{"body":" Hey peeps, apologies for going radio silent, I've been context switching like crazy lately and this fell through.\r\n\r\nI back ported Josh changes onto our own 4.0.5 (as in, putting https://github.com/apache/cassandra/commit/f4a2135c5ba442aafd27bb7c12c85b376d5a2b87 on top of 4.0.5) and ran it for a few days. I captured some traces which again talk about self-ref loops, not sure if they're caused by a bad port on my end or are legitimate. Additionally, I may be looking at the wrong place (Ref.java) so if you have any suggestion on which logger class I should keep an eye out for, I'd be grateful.\r\n\r\n [^cassandra_18332-20230415.txt] \r\n\r\nUnless you deem the information sufficient, I'll try to roll out 4.0.9 in the upcoming week(s)","created":"2023-04-15T08:05:49.296+0000"},{"body":"Hey [~brandon.williams], [~jmckenzie] and [~smiklosovic],\r\n\r\nApologies for going M.i.A so long but I do have an update: it looks like with 4.0.9 the issue is not popping up again.\r\n\r\nWe've been running 4.0.9 in hundreds of nodes and multiple access patterns for over a month now and we've not seen any unreleased file descriptors. From this I'd say the issue can be closed","created":"2023-07-03T14:07:29.695+0000"}],"conversations":[{"body":"(EDIT: Looks like this is masked by the lack of backport of CASSANDRA-17205 just on the 4.0 line - will backport that and block here on it)\r\n\r\nAfter upgrading to Cassandra 4.x (4.0.1 and 4.0.5) we've noticed that at times after a compaction deleted sstables diskspace doesn't get reclaimed by the OS until the cassandra process is restarted (which kinda points at some sort of resource leak), I do not recall this happening in cassandra 3, at least not to such degree.\r\n\r\nWe've seen the behavior in multiple clusters with different schemas, access patterns and consistency levels at somewhat \"random\" points in time, the only interesting thing is that there were active repair sessions at the time affecting the node, keyspace and table.\r\n{noformat}\r\n$ date +%Y-%m-%d\r\n2023-01-17\r\n\r\n$ nodetool version\r\nReleaseVersion: 4.0.5\r\n{noformat}\r\n{noformat}\r\n$ lsof +L1 | grep cassandra | grep myawesomecluster | wc -l\r\n2772\r\n\r\n$ lsof +L1 | grep cassandra | grep myawesomecluster | tail -n1\r\njava 59003 cassandra *979u REG 253,8 10 0 1208053768 /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Digest.crc32 (deleted)\r\n{noformat}\r\n{noformat}\r\n$ grep 2274426 /cassandra/systemlog.log\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,969 BigTableZeroCopyWriter.java:203 - Writing component DATA to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Data.db length 2.900KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,969 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Data.db length 2.900KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,970 BigTableZeroCopyWriter.java:203 - Writing component PRIMARY_INDEX to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Index.db length 3.739KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,970 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Index.db length 3.739KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,970 BigTableZeroCopyWriter.java:203 - Writing component STATS to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Statistics.db length 5.062KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,970 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Statistics.db length 5.062KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,971 BigTableZeroCopyWriter.java:203 - Writing component COMPRESSION_INFO to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-CompressionInfo.db length 0.054KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,971 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-CompressionInfo.db length 0.054KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,971 BigTableZeroCopyWriter.java:203 - Writing component FILTER to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Filter.db length 0.031KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,971 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Filter.db length 0.031KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,972 BigTableZeroCopyWriter.java:203 - Writing component SUMMARY to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Summary.db length 0.436KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,972 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Summary.db length 0.436KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,972 BigTableZeroCopyWriter.java:203 - Writing component DIGEST to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Digest.crc32 length 0.010KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,972 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Digest.crc32 length 0.010KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,974 SSTableReaderBuilder.java:351 - Opening /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big (2.900KiB)\r\nINFO [CompactionExecutor:54141] 2023-01-15 13:06:24,978 CompactionTask.java:150 - Compacting (6a3e8320-94d5-11ed-b2c8-7b967b642f39) [/cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274422-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274411-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274423-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274410-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274412-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274413-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274414-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274419-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274425-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274424-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274418-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274421-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274420-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274416-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274417-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274415-big-Data.db:level=0, ]\r\nINFO [NonPeriodicTasks:1] 2023-01-15 13:06:25,254 SSTable.java:111 - Deleting sstable: /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big\r\n{noformat}\r\n{noformat}\r\n$ nodetool compactionhistory | grep '2023-01-15T13:06' \r\n75032a40-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:43.044 53958242 52853746 {1:478098, 2:11313, 3:1765, 4:287, 5:185, 6:88, 7:39, 8:22, 9:13, 10:7, 11:3, 14:1}\r\n6a4539e0-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:25.022 107577 30110 {1:25, 2:66, 3:124, 4:16, 5:66, 6:15, 9:1}\r\n6a1b43b0-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:24.747 91018 29868 {1:16, 2:127, 3:148, 4:5, 5:7, 6:6, 7:3, 8:1, 10:1}\r\n6a063510-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:24.609 767 366 {2:2}\r\n6a0523a0-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:24.602 87020 27917 {1:73, 2:38, 3:96, 4:15, 5:67, 6:1}\r\n69c5a9a0-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:24.186 45345 25814 {1:117, 2:150, 3:3}\r\n6956bb30-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_two 2023-01-15T13:06:23.459 8925831 8904158 {1:102662, 2:120}\r\n5d886c40-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_two 2023-01-15T13:06:03.652 8901698 8901269 {1:102606, 2:9}\r\n5d02c180-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_two 2023-01-15T13:06:02.776 8976310 8900384 {1:102844, 2:324}\r\n5bf9dd00-94d5-11ed-b2c8-7b967b642f39 system_distributed repair_history 2023-01-15T13:06:01.040 4848124 4847506 {4:2}\r\n5bc27950-94d5-11ed-b2c8-7b967b642f39 system_distributed parent_repair_history 2023-01-15T13:06:00.677 351877 351135 {1:2383, 2:1}\r\n{noformat}","from":"reporter","subject":"Merged SSTable files not reclaimed by OS"},{"body":"You could enable {{-Dcassandra.debugrefcount=true}} - this would have some process overhead, but it should find any resource leak of this kind.","from":"developer"},{"body":"Thanks for the pointer [~benedict] , I'll see if I can add it \"safely\".","from":"developer"},{"body":"Finally managed to enable it in one node, keeping an eye out for leaks. I'd say though, _this would have some process overhead_ may be a bit of an understatement, load average jumped from 2-3 to 10-11 in a 14 core machine","from":"developer"},{"body":"Found quite a few of these, anything else I should look out for that may be interesting [~benedict]?\r\n{noformat}\r\nERROR [Strong-Reference-Leak-Detector:1] 2023-02-10 09:24:08,119 NoSpamLogger.java:98 - Strong self-ref loop detected \r\n[/cassandra-data/data/system_distributed/repair_history-759fffad624b318180eefa9a52d1f627/nb-628647-big, private org.apache.cassandra.utils.concurrent.Ref org.apache.cassandra.io.sstable.format.SSTableReader$InstanceTidier.globalRef-org.apache.cassandra.utils.concurrent.Ref, final org.apache.cassandra.utils.concurrent.Ref$State org.apache.cassandra.utils.concurrent.Ref.state-org.apache.cassandra.utils.concurrent.Ref$State, final org.apache.cassandra.utils.concurrent.Ref$GlobalState org.apache.cassandra.utils.concurrent.Ref$State.globalState-org.apache.cassandra.utils.concurrent.Ref$GlobalState, private final org.apache.cassandra.utils.concurrent.RefCounted$Tidy org.apache.cassandra.utils.concurrent.Ref$GlobalState.tidy-org.apache.cassandra.io.sstable.format.SSTableReader$GlobalTidy, private volatile java.lang.Runnable org.apache.cassandra.io.sstable.format.SSTableReader$GlobalTidy.obsoletion-org.apache.cassandra.db.lifecycle.LogTransaction$SSTableTidier, private final org.apache.cassandra.db.lifecycle.Tracker org.apache.cassandra.db.lifecycle.LogTransaction$SSTableTidier.tracker-org.apache.cassandra.db.lifecycle.Tracker, private final java.util.Collection org.apache.cassandra.db.lifecycle.Tracker.subscribers-java.util.concurrent.CopyOnWriteArrayList, private final java.util.Collection org.apache.cassandra.db.lifecycle.Tracker.subscribers-org.apache.cassandra.service.SSTablesGlobalTracker, private final java.util.Set org.apache.cassandra.service.SSTablesGlobalTracker.subscribers-java.util.concurrent.CopyOnWriteArraySet, private final java.util.Set org.apache.cassandra.service.SSTablesGlobalTracker.subscribers-org.apache.cassandra.service.StorageService$$Lambda$883/146242186, private final org.apache.cassandra.service.StorageService org.apache.cassandra.service.StorageService$$Lambda$883/146242186.arg$1-org.apache.cassandra.service.StorageService, private java.lang.Thread org.apache.cassandra.service.StorageService.drainOnShutdown-io.netty.util.concurrent.FastThreadLocalThread, private java.lang.ThreadGroup java.lang.Thread.group-java.lang.ThreadGroup, private final java.lang.ThreadGroup java.lang.ThreadGroup.parent-java.lang.ThreadGroup, java.lang.Thread[] java.lang.ThreadGroup.threads-[Ljava.lang.Thread;, java.lang.Thread[] java.lang.ThreadGroup.threads-java.lang.Thread, private java.lang.Runnable java.lang.Thread.target-java.util.concurrent.ThreadPoolExecutor$Worker, final java.util.concurrent.ThreadPoolExecutor java.util.concurrent.ThreadPoolExecutor$Worker.this$0-java.util.concurrent.ScheduledThreadPoolExecutor, private final java.util.concurrent.BlockingQueue java.util.concurrent.ThreadPoolExecutor.workQueue-java.util.concurrent.ScheduledThreadPoolExecutor$DelayedWorkQueue, private final java.util.concurrent.BlockingQueue java.util.concurrent.ThreadPoolExecutor.workQueue-java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask, private java.util.concurrent.Callable java.util.concurrent.FutureTask.callable-java.util.concurrent.Executors$RunnableAdapter, final java.lang.Runnable java.util.concurrent.Executors$RunnableAdapter.task-sun.rmi.transport.DGCImpl$1, final sun.rmi.transport.DGCImpl sun.rmi.transport.DGCImpl$1.this$0-sun.rmi.transport.DGCImpl, private java.util.Map sun.rmi.transport.DGCImpl.leaseTable-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-sun.rmi.transport.DGCImpl$LeaseInfo, java.util.Set sun.rmi.transport.DGCImpl$LeaseInfo.notifySet-java.util.HashSet, private transient java.util.HashMap java.util.HashSet.map-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, final java.lang.Object java.util.HashMap$Node.key-sun.rmi.transport.Target, private final sun.rmi.transport.WeakRef sun.rmi.transport.Target.weakImpl-sun.rmi.transport.WeakRef, private java.lang.Object sun.rmi.transport.WeakRef.strongRef-javax.management.remote.rmi.RMIJRMPServerImpl, private javax.management.MBeanServer javax.management.remote.rmi.RMIServerImpl.mbeanServer-com.sun.jmx.mbeanserver.JmxMBeanServer, private volatile javax.management.MBeanServer com.sun.jmx.mbeanserver.JmxMBeanServer.mbsInterceptor-com.zegelin.cassandra.exporter.MBeanServerInterceptorHarvester$MBeanServerInterceptor, private final javax.management.MBeanServer com.zegelin.jmx.DelegatingMBeanServerInterceptor.delegate-com.sun.jmx.interceptor.DefaultMBeanServerInterceptor, private final transient com.sun.jmx.mbeanserver.Repository com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.repository-com.sun.jmx.mbeanserver.Repository, private final java.util.Map com.sun.jmx.mbeanserver.Repository.domainTb-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.util.HashMap$Node java.util.HashMap$Node.next-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-com.sun.jmx.mbeanserver.NamedObject, private final javax.management.DynamicMBean com.sun.jmx.mbeanserver.NamedObject.object-com.sun.jmx.mbeanserver.StandardMBeanSupport, private final java.lang.Object com.sun.jmx.mbeanserver.MBeanSupport.resource-org.apache.cassandra.db.compaction.CompactionManager, public final org.apache.cassandra.db.compaction.ActiveCompactions org.apache.cassandra.db.compaction.CompactionManager.active-org.apache.cassandra.db.compaction.ActiveCompactions, private final java.util.Set org.apache.cassandra.db.compaction.ActiveCompactions.compactions-java.util.Collections$SynchronizedSet, final java.util.Collection java.util.Collections$SynchronizedCollection.c-java.util.Collections$SetFromMap, private final java.util.Map java.util.Collections$SetFromMap.m-java.util.IdentityHashMap, transient java.lang.Object[] java.util.IdentityHashMap.table-[Ljava.lang.Object;, transient java.lang.Object[] java.util.IdentityHashMap.table-org.apache.cassandra.db.compaction.CompactionIterator, private final org.apache.cassandra.db.AbstractCompactionController org.apache.cassandra.db.compaction.CompactionIterator.controller-org.apache.cassandra.db.compaction.TimeWindowCompactionController, private org.apache.cassandra.utils.concurrent.Refs org.apache.cassandra.db.compaction.CompactionController.overlappingSSTables-org.apache.cassandra.utils.concurrent.Refs, private final java.util.Map org.apache.cassandra.utils.concurrent.Refs.references-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, final java.lang.Object java.util.HashMap$Node.key-org.apache.cassandra.io.sstable.format.big.BigTableReader, private final org.apache.cassandra.utils.concurrent.Ref org.apache.cassandra.io.sstable.format.SSTableReader.selfRef-org.apache.cassandra.utils.concurrent.Ref]\r\n{noformat}\r\n{noformat}\r\nERROR [Strong-Reference-Leak-Detector:1] 2023-02-10 09:26:30,587 NoSpamLogger.java:98 - Strong self-ref loop detected \r\n[/cassandra-data/data/system_distributed/repair_history-759fffad624b318180eefa9a52d1f627/nb-628650-big, private volatile java.lang.Runnable org.apache.cassandra.io.sstable.format.SSTableReader$GlobalTidy.obsoletion-org.apache.cassandra.db.lifecycle.LogTransaction$SSTableTidier, private final org.apache.cassandra.db.lifecycle.Tracker org.apache.cassandra.db.lifecycle.LogTransaction$SSTableTidier.tracker-org.apache.cassandra.db.lifecycle.Tracker, private final java.util.Collection org.apache.cassandra.db.lifecycle.Tracker.subscribers-java.util.concurrent.CopyOnWriteArrayList, private final java.util.Collection org.apache.cassandra.db.lifecycle.Tracker.subscribers-org.apache.cassandra.service.SSTablesGlobalTracker, private final java.util.Set org.apache.cassandra.service.SSTablesGlobalTracker.subscribers-java.util.concurrent.CopyOnWriteArraySet, private final java.util.Set org.apache.cassandra.service.SSTablesGlobalTracker.subscribers-org.apache.cassandra.service.StorageService$$Lambda$883/146242186, private final org.apache.cassandra.service.StorageService org.apache.cassandra.service.StorageService$$Lambda$883/146242186.arg$1-org.apache.cassandra.service.StorageService, private java.lang.Thread org.apache.cassandra.service.StorageService.drainOnShutdown-io.netty.util.concurrent.FastThreadLocalThread, private java.lang.ThreadGroup java.lang.Thread.group-java.lang.ThreadGroup, private final java.lang.ThreadGroup java.lang.ThreadGroup.parent-java.lang.ThreadGroup, java.lang.Thread[] java.lang.ThreadGroup.threads-[Ljava.lang.Thread;, java.lang.Thread[] java.lang.ThreadGroup.threads-java.lang.Thread, private java.lang.Runnable java.lang.Thread.target-java.util.concurrent.ThreadPoolExecutor$Worker, final java.util.concurrent.ThreadPoolExecutor java.util.concurrent.ThreadPoolExecutor$Worker.this$0-java.util.concurrent.ScheduledThreadPoolExecutor, private final java.util.concurrent.BlockingQueue java.util.concurrent.ThreadPoolExecutor.workQueue-java.util.concurrent.ScheduledThreadPoolExecutor$DelayedWorkQueue, private final java.util.concurrent.BlockingQueue java.util.concurrent.ThreadPoolExecutor.workQueue-java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask, private java.util.concurrent.Callable java.util.concurrent.FutureTask.callable-java.util.concurrent.Executors$RunnableAdapter, final java.lang.Runnable java.util.concurrent.Executors$RunnableAdapter.task-sun.rmi.transport.DGCImpl$1, final sun.rmi.transport.DGCImpl sun.rmi.transport.DGCImpl$1.this$0-sun.rmi.transport.DGCImpl, private java.util.Map sun.rmi.transport.DGCImpl.leaseTable-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-sun.rmi.transport.DGCImpl$LeaseInfo, java.util.Set sun.rmi.transport.DGCImpl$LeaseInfo.notifySet-java.util.HashSet, private transient java.util.HashMap java.util.HashSet.map-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, final java.lang.Object java.util.HashMap$Node.key-sun.rmi.transport.Target, private final sun.rmi.transport.WeakRef sun.rmi.transport.Target.weakImpl-sun.rmi.transport.WeakRef, private java.lang.Object sun.rmi.transport.WeakRef.strongRef-javax.management.remote.rmi.RMIJRMPServerImpl, private javax.management.MBeanServer javax.management.remote.rmi.RMIServerImpl.mbeanServer-com.sun.jmx.mbeanserver.JmxMBeanServer, private volatile javax.management.MBeanServer com.sun.jmx.mbeanserver.JmxMBeanServer.mbsInterceptor-com.zegelin.cassandra.exporter.MBeanServerInterceptorHarvester$MBeanServerInterceptor, private final javax.management.MBeanServer com.zegelin.jmx.DelegatingMBeanServerInterceptor.delegate-com.sun.jmx.interceptor.DefaultMBeanServerInterceptor, private final transient com.sun.jmx.mbeanserver.Repository com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.repository-com.sun.jmx.mbeanserver.Repository, private final java.util.Map com.sun.jmx.mbeanserver.Repository.domainTb-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, java.util.HashMap$Node java.util.HashMap$Node.next-java.util.HashMap$Node, java.lang.Object java.util.HashMap$Node.value-com.sun.jmx.mbeanserver.NamedObject, private final javax.management.DynamicMBean com.sun.jmx.mbeanserver.NamedObject.object-com.sun.jmx.mbeanserver.StandardMBeanSupport, private final java.lang.Object com.sun.jmx.mbeanserver.MBeanSupport.resource-org.apache.cassandra.db.compaction.CompactionManager, public final org.apache.cassandra.db.compaction.ActiveCompactions org.apache.cassandra.db.compaction.CompactionManager.active-org.apache.cassandra.db.compaction.ActiveCompactions, private final java.util.Set org.apache.cassandra.db.compaction.ActiveCompactions.compactions-java.util.Collections$SynchronizedSet, final java.util.Collection java.util.Collections$SynchronizedCollection.c-java.util.Collections$SetFromMap, private final java.util.Map java.util.Collections$SetFromMap.m-java.util.IdentityHashMap, transient java.lang.Object[] java.util.IdentityHashMap.table-[Ljava.lang.Object;, transient java.lang.Object[] java.util.IdentityHashMap.table-org.apache.cassandra.db.compaction.CompactionIterator, private final org.apache.cassandra.db.AbstractCompactionController org.apache.cassandra.db.compaction.CompactionIterator.controller-org.apache.cassandra.db.compaction.TimeWindowCompactionController, private org.apache.cassandra.utils.concurrent.Refs org.apache.cassandra.db.compaction.CompactionController.overlappingSSTables-org.apache.cassandra.utils.concurrent.Refs, private final java.util.Map org.apache.cassandra.utils.concurrent.Refs.references-java.util.HashMap, transient java.util.HashMap$Node[] java.util.HashMap.table-[Ljava.util.HashMap$Node;, transient java.util.HashMap$Node[] java.util.HashMap.table-java.util.HashMap$Node, final java.lang.Object java.util.HashMap$Node.key-org.apache.cassandra.io.sstable.format.big.BigTableReader, private final org.apache.cassandra.io.sstable.format.SSTableReader$InstanceTidier org.apache.cassandra.io.sstable.format.SSTableReader.tidy-org.apache.cassandra.io.sstable.format.SSTableReader$InstanceTidier, private org.apache.cassandra.utils.concurrent.Ref org.apache.cassandra.io.sstable.format.SSTableReader$InstanceTidier.globalRef-org.apache.cassandra.utils.concurrent.Ref]\r\n {noformat}","from":"developer"},{"body":"Thanks [~pbalaguer]. I think probably this bug is a combination of a strong ref leak and a regular leak, with the former masking the latter as it prevents our regular leak detection logic kicking in. \r\n\r\nHowever, once we fix the former, we will hopefully see the latter in our usual test suite (though we should really be finding the former in our test suites too, so we need to look into why we aren't seeing that).","from":"developer"},{"body":"I would say you can definitely disable the additional logging on your production systems at least for now. Thanks.","from":"developer"},{"body":"Ok, this particular (strong ref loop) issue is a duplicate of CASSANDRA-17205. [~jmckenzie] from that ticket it looks like we didn't backport the fix for some reason?\r\n\r\n","from":"developer"},{"body":"Understood, I'll proceed to disable it. Given the automation has been updated, if this needs to be re-enabled it'll be a matter of minutes.","from":"developer"},{"body":"{quote}Josh McKenzie from that ticket it looks like we didn't backport the fix for some reason?\r\n{quote}\r\n... I'm at a loss. I don't even remember working on that _at all_ much less why I didn't backport it.\r\n\r\nI'll massage this ticket's metadata in a bit, take ownership, and backport CASSANDRA-17205 into the 4.0 line.","from":"developer"},{"body":"For posterity (i.e. my future Not Forgetting-ness), Brandon pointed out this might need to be backported to all supported branches. I'll take a look when I get to this and handle that.","from":"developer"},{"body":"I was probably unclear, but I do not think this bug report is a duplicate of CASSANDRA-17205, only that this earlier bug is masking whatever the underlying bug is here.","from":"developer"},{"body":"Ah. Hm. So another JIRA for the backports as a blocker for this may well be the appropriate move here. Sorry for the noise / pollution.","from":"developer"},{"body":"As part of any backport work, do you want to perhaps take a look at why we aren't seeing the strong loop detection errors anywhere in our CI, or if we are why they aren't being surfaced? Or perhaps just create a dedicated strong loop test that makes sure to run the detector at least once after running some basic tasks?\r\n\r\nThis is a pretty serious issue really, as it completely disables all of our leak detection which is obviously very important. So we want to catch mistakes here as early as possible.","from":"developer"},{"body":"bq. do you want to perhaps take a look at why we aren't seeing the strong loop detection errors anywhere in our CI, or if we are why they aren't being surfaced?\r\nJust so happens that I've had a task in my backlog for awhile to surface and fail out tests if we have ref leaks on them, just haven't gotten to that yet.\r\n\r\nI'll be pivoting to focus on CI related things in a month or so so I'll pencil in rolling that up w/the rest of the effort. Completely agree on the importance of it.","from":"developer"},{"body":"Perfect. Thanks","from":"developer"},{"body":"Anything logged at ERROR would surface in CI (at least off the top of my head I know the python dtests will complain) so I have to think they aren't being logged or the problem isn't being triggered.","from":"developer"},{"body":"Looks like strong leaks are at WARN. We should bump to ERROR I guess.\r\n\r\nAlso, the strong loop detector can take a while to run as it depends on reflection, and doesn't really know where to focus its efforts. So it might be many tests are any way much to brief. I'm not 100% certain we run with debugrefcount for all our tests either. We should make sure that we do.","from":"developer"},{"body":"bq. I'm not 100% certain we run with debugrefcount for all our tests \r\n\r\nIf we do, I don't see it anywhere.","from":"developer"},{"body":"It's in {{testmacrohelper}} as one of the jvmargs. So we should be running with it ({{-Dcassandra.debugrefcount=true}})","from":"developer"},{"body":"bq. Looks like strong leaks are at WARN. We should bump to ERROR I guess.\r\n\r\nI think we do log at [ERROR|https://github.com/apache/cassandra/blob/cassandra-4.0/src/java/org/apache/cassandra/utils/concurrent/Ref.java#L600] so something else is going on here.","from":"developer"},{"body":"That's the reference leak, not a \"strong reference leak\" - which is perhaps a misnomer. There's a separate feature to search the object hierarchy for strong reference loops that prevent {{Ref}} from printing those ERROR messages, i.e. it looks for us \"leaking\" a strong reference to a {{Ref}} into its {{Tidy}}, so that it is not garbage collected (and so, if we fail to release it, our leak detection will not inform us).\r\n\r\nThis feature logs at [WARN|https://github.com/apache/cassandra/blob/5d3c747719f01d87c9086c405c806317405a8e43/src/java/org/apache/cassandra/utils/concurrent/Ref.java#L711]\r\n\r\nIt's worth noting that it's an assumption that there's a regular leak in this ticket. It might be something else going on. But we won't know until our leak detection is working.","from":"developer"},{"body":"bq. This feature logs at WARN\r\n\r\nI see. I agree we should bump that up to ERROR.","from":"developer"},{"body":"Hey folks, sorry for bumping this but the issue is rather annoying when there are a lot of ongoing ring movements as it triggers quite often forcing node restarts which in turn delay everything. Do you happen to have an idea on when can this be worked on and if I can do something to contribute a bit more to make it happen sooner? (not so sure about deep down diving and coding but hey, maybe).","from":"developer"},{"body":"Sorry [~pbalaguer], I have a lot on and had mentally handed this off to others for the moment.\r\n\r\nI could very easily back port this to a version you are running (please nominate one), and post a branch to GitHub that you could build and deploy with the strong reference loop only fixed, so that you could enable logging so that we might be able to track down the underlying resource leak.\r\n\r\nIf you're happy with this course of action, I'll post a branch either today or tomorrow. Otherwise, we can prod [~jmckenzie] and see about when we can get it properly backported and merged into a release.\r\n\r\nI agree that we should get to the bottom of this sooner than later, for the project as well as yourself.","from":"developer"},{"body":"Hey [~benedict] , if that's the situation on your end don't worry, I'll try to backport it myself and see how it goes from there. If I can't figure it out I'll ping here again.\r\n\r\nThis effort is not meant to stop nor delay [~jmckenzie] but to happen in parallel.","from":"developer"},{"body":"Yikes; this one slipped through the cracks on my todo list. Created CASSANDRA-18332 for the backport; should be trivial to do tomorrow morning.\r\n\r\nSorry about that.","from":"developer"},{"body":"CASSANDRA-18332 will be in 4.0.9. Do we need to do anything else in this ticket, [~pbalaguer] ? ","from":"developer"},{"body":"I don't think we've definitively solved anything here, but perhaps only added debugging that can tell us.","from":"developer"},{"body":"Let's wait for more debugging then. Looking forward what Pere reports.","from":"developer"},{"body":" Hey peeps, apologies for going radio silent, I've been context switching like crazy lately and this fell through.\r\n\r\nI back ported Josh changes onto our own 4.0.5 (as in, putting https://github.com/apache/cassandra/commit/f4a2135c5ba442aafd27bb7c12c85b376d5a2b87 on top of 4.0.5) and ran it for a few days. I captured some traces which again talk about self-ref loops, not sure if they're caused by a bad port on my end or are legitimate. Additionally, I may be looking at the wrong place (Ref.java) so if you have any suggestion on which logger class I should keep an eye out for, I'd be grateful.\r\n\r\n [^cassandra_18332-20230415.txt] \r\n\r\nUnless you deem the information sufficient, I'll try to roll out 4.0.9 in the upcoming week(s)","from":"developer"},{"body":"Hey [~brandon.williams], [~jmckenzie] and [~smiklosovic],\r\n\r\nApologies for going M.i.A so long but I do have an update: it looks like with 4.0.9 the issue is not popping up again.\r\n\r\nWe've been running 4.0.9 in hundreds of nodes and multiple access patterns for over a month now and we've not seen any unreleased file descriptors. From this I'd say the issue can be closed","from":"developer"}],"created":"2023-01-18T13:19:34.000+0000","description":"(EDIT: Looks like this is masked by the lack of backport of CASSANDRA-17205 just on the 4.0 line - will backport that and block here on it)\r\n\r\nAfter upgrading to Cassandra 4.x (4.0.1 and 4.0.5) we've noticed that at times after a compaction deleted sstables diskspace doesn't get reclaimed by the OS until the cassandra process is restarted (which kinda points at some sort of resource leak), I do not recall this happening in cassandra 3, at least not to such degree.\r\n\r\nWe've seen the behavior in multiple clusters with different schemas, access patterns and consistency levels at somewhat \"random\" points in time, the only interesting thing is that there were active repair sessions at the time affecting the node, keyspace and table.\r\n{noformat}\r\n$ date +%Y-%m-%d\r\n2023-01-17\r\n\r\n$ nodetool version\r\nReleaseVersion: 4.0.5\r\n{noformat}\r\n{noformat}\r\n$ lsof +L1 | grep cassandra | grep myawesomecluster | wc -l\r\n2772\r\n\r\n$ lsof +L1 | grep cassandra | grep myawesomecluster | tail -n1\r\njava 59003 cassandra *979u REG 253,8 10 0 1208053768 /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Digest.crc32 (deleted)\r\n{noformat}\r\n{noformat}\r\n$ grep 2274426 /cassandra/systemlog.log\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,969 BigTableZeroCopyWriter.java:203 - Writing component DATA to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Data.db length 2.900KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,969 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Data.db length 2.900KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,970 BigTableZeroCopyWriter.java:203 - Writing component PRIMARY_INDEX to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Index.db length 3.739KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,970 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Index.db length 3.739KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,970 BigTableZeroCopyWriter.java:203 - Writing component STATS to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Statistics.db length 5.062KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,970 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Statistics.db length 5.062KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,971 BigTableZeroCopyWriter.java:203 - Writing component COMPRESSION_INFO to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-CompressionInfo.db length 0.054KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,971 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-CompressionInfo.db length 0.054KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,971 BigTableZeroCopyWriter.java:203 - Writing component FILTER to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Filter.db length 0.031KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,971 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Filter.db length 0.031KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,972 BigTableZeroCopyWriter.java:203 - Writing component SUMMARY to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Summary.db length 0.436KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,972 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Summary.db length 0.436KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,972 BigTableZeroCopyWriter.java:203 - Writing component DIGEST to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Digest.crc32 length 0.010KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,972 BigTableZeroCopyWriter.java:213 - Block Writing component to /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Digest.crc32 length 0.010KiB\r\nINFO [Stream-Deserializer-/10.214.79.62:randomport-b904af67] 2023-01-15 13:06:24,974 SSTableReaderBuilder.java:351 - Opening /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big (2.900KiB)\r\nINFO [CompactionExecutor:54141] 2023-01-15 13:06:24,978 CompactionTask.java:150 - Compacting (6a3e8320-94d5-11ed-b2c8-7b967b642f39) [/cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274422-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274411-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274423-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274410-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274412-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274413-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274414-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274419-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274425-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274424-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274418-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274421-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274420-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274416-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274417-big-Data.db:level=0, /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274415-big-Data.db:level=0, ]\r\nINFO [NonPeriodicTasks:1] 2023-01-15 13:06:25,254 SSTable.java:111 - Deleting sstable: /cassandra-data/data/myawesomecluster/schema_one-randomchecksum/nb-2274426-big\r\n{noformat}\r\n{noformat}\r\n$ nodetool compactionhistory | grep '2023-01-15T13:06' \r\n75032a40-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:43.044 53958242 52853746 {1:478098, 2:11313, 3:1765, 4:287, 5:185, 6:88, 7:39, 8:22, 9:13, 10:7, 11:3, 14:1}\r\n6a4539e0-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:25.022 107577 30110 {1:25, 2:66, 3:124, 4:16, 5:66, 6:15, 9:1}\r\n6a1b43b0-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:24.747 91018 29868 {1:16, 2:127, 3:148, 4:5, 5:7, 6:6, 7:3, 8:1, 10:1}\r\n6a063510-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:24.609 767 366 {2:2}\r\n6a0523a0-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:24.602 87020 27917 {1:73, 2:38, 3:96, 4:15, 5:67, 6:1}\r\n69c5a9a0-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_one 2023-01-15T13:06:24.186 45345 25814 {1:117, 2:150, 3:3}\r\n6956bb30-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_two 2023-01-15T13:06:23.459 8925831 8904158 {1:102662, 2:120}\r\n5d886c40-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_two 2023-01-15T13:06:03.652 8901698 8901269 {1:102606, 2:9}\r\n5d02c180-94d5-11ed-b2c8-7b967b642f39 myawesomecluster schema_two 2023-01-15T13:06:02.776 8976310 8900384 {1:102844, 2:324}\r\n5bf9dd00-94d5-11ed-b2c8-7b967b642f39 system_distributed repair_history 2023-01-15T13:06:01.040 4848124 4847506 {4:2}\r\n5bc27950-94d5-11ed-b2c8-7b967b642f39 system_distributed parent_repair_history 2023-01-15T13:06:00.677 351877 351135 {1:2383, 2:1}\r\n{noformat}","issue_id":"13520058","key":"CASSANDRA-18176","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2023-08-17T14:23:47.000+0000","role":"fixed_distractor","summary":"Merged SSTable files not reclaimed by OS"} {"case_id":"13546581","cluster":"DISTRACTOR-CASSANDRA-18736","comments":[{"body":"This ticket hasn't been updated for a while, is anything blocking it? ","created":"2024-03-13T15:15:55.389+0000"},{"body":"[~jonmeredith] , the fix versions mention 4.0, but I do not see a branch for it. Do you plan to apply it on that branch too?","created":"2024-04-15T14:14:27.118+0000"},{"body":"I'm double-checking at the moment. I think the bug became an issue after some restructuring for the simulator, but need to convince myself of that. \r\n\r\nAlso, I've deleted the original branches/PRs as I accidentally referenced ones from a different fix. I'll post new ones in a day or so.","created":"2024-04-15T15:58:35.624+0000"},{"body":"Hey [~jonmeredith], I haven't tried myself, but I have people telling me they reproduced the issue on 4.0.\r\n\r\nI can try to ask for more details if that will help you.  ","created":"2024-04-15T16:10:18.649+0000"},{"body":"As expected, the (modified) tests fail on 4.0 so I've backported there too. I'll try and post test runs tomorrow but wanted to share the changes in case anybody was waiting on it.\r\n\r\nPRs\r\nTrunk https://github.com/apache/cassandra/pull/3250\r\n5.0 https://github.com/apache/cassandra/pull/3251\r\n4.1 https://github.com/apache/cassandra/pull/3252\r\n4.0 https://github.com/apache/cassandra/pull/3253","created":"2024-04-16T02:48:51.231+0000"},{"body":"I left a few nits in the 4.0 PR, but also approved all 4 PRs.\r\n\r\n+1","created":"2024-04-18T20:29:19.988+0000"},{"body":"Starting commit. I believe all of the test failures are known flakes or unrelated issues. This patch has been running in production for me for many months without issue so I have high confidence in it.","created":"2024-04-25T17:56:37.914+0000"}],"conversations":[{"body":"On restart, Cassandra logs this message and terminates.\r\n{code:java}\r\nERROR 2023-07-17T17:17:22,931 [main] org.apache.cassandra.db.lifecycle.LogTransaction:561 - Unexpected disk state: failed to read transaction log [nb_txn_stream_39d5f6b0-fb81-11ed-8f46-e97b3f61511e.log in /datadir1/keyspace/table-c9527530a0d611e8813f034699fc9043]\r\nFiles and contents follow:\r\n/datadir1/keyspace/table-c9527530a0d611e8813f034699fc9043/nb_txn_stream_39d5f6b0-fb81-11ed-8f46-e97b3f61511e.log\r\n ABORT:[,0,0][737437348]\r\n ***This record should have been the last one in all replicas\r\n ADD:[/datadir1/keyspace/table-c9527530a0d611e8813f034699fc9043/nb-284490-big-,0,8][2493503833]\r\n{code}\r\nThe root cause is a race during streaming exception handling.\r\n\r\nAlthough concurrent modification of to the {{LogTransaction}} was added for CASSANDRA-16225, there is nothing to prevent usage after the transaction is completed (committed/aborted) once it has been processed by {{TransactionTidier}} (after the last reference is released). Before the transaction is tidied, the {{LogFile}} keeps a list of records that are checked for completion before adding new entries. In {{TransactionTidier}} {{LogFile.records}} are cleared as no longer needed, however the LogTransaction/LogFile is still accessible to the stream.\r\n\r\nThe changes in CASSANDRA-17273 added a parallel set of {{onDiskRecords}} that could be used to reliably recreate the transaction log at any new datadirs the same as the existing\r\ndatadirs - regardless of the effect of {{LogTransaction.untrackNew/LogFile.remove}}\r\n\r\nIf a streaming exception causes the LogTransaction to be aborted and tidied just before {{SimpleSSTableMultiWriter}} calls trackNew to add a new sstable. At the time of the call, the {{LogFile}} will not contain any {{LogReplicas}},\r\n{{LogFile.records}} will be empty, and {{LogFile.onDiskRecords}} will contain an {{ABORT}}.\r\n\r\nWhen {{LogTransaction.trackNew/LogFile.add}} is called, the check for completed transaction fails as records is empty, there are no replicas on the datadir, so {{maybeCreateReplicas}} creates a new txnlog file replica containing ABORT, then\r\nappends an ADD record.\r\n\r\nThe LogFile has already been tidied after the abort so the txnlog file is not removed and sits on disk until a restart, causing the faiulre.\r\n\r\nThere is a related exception caused with a different interleaving of aborts, after an sstable is added, however this is just a nuisance in the logs as the LogRelica is already created with an {{ADD}} record first.\r\n{code:java}\r\njava.lang.AssertionError: [ADD:[/datadir1/keyspace/table/nb-23314378-big-,0,8][1869379820]] is not tracked by 55be35b0-35d1-11ee-865d-8b1e3c48ca06\r\n at org.apache.cassandra.db.lifecycle.LogFile.remove(LogFile.java:388)\r\n at org.apache.cassandra.db.lifecycle.LogTransaction.untrackNew(LogTransaction.java:158)\r\n at org.apache.cassandra.db.lifecycle.LifecycleTransaction.untrackNew(LifecycleTransaction.java:577)\r\n at org.apache.cassandra.db.streaming.CassandraStreamReceiver$1.untrackNew(CassandraStreamReceiver.java:149)\r\n at org.apache.cassandra.io.sstable.SimpleSSTableMultiWriter.abort(SimpleSSTableMultiWriter.java:95)\r\n at org.apache.cassandra.io.sstable.format.RangeAwareSSTableWriter.abort(RangeAwareSSTableWriter.java:191)\r\n at org.apache.cassandra.db.streaming.CassandraCompressedStreamReader.read(CassandraCompressedStreamReader.java:115)\r\n at org.apache.cassandra.db.streaming.CassandraIncomingFile.read(CassandraIncomingFile.java:85)\r\n at org.apache.cassandra.streaming.messages.IncomingStreamMessage$1.deserialize(IncomingStreamMessage.java:53)\r\n at org.apache.cassandra.streaming.messages.IncomingStreamMessage$1.deserialize(IncomingStreamMessage.java:38)\r\n at org.apache.cassandra.streaming.messages.StreamMessage.deserialize(StreamMessage.java:53)\r\n at org.apache.cassandra.streaming.async.StreamingInboundHandler$StreamDeserializingTask.run(StreamingInboundHandler.java:172)\r\n at io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n at java.base/java.lang.Thread.run(Thread.java:829)\r\n{code}","from":"reporter","subject":"Streaming exception race creates corrupt transaction log files that prevent restart"},{"body":"This ticket hasn't been updated for a while, is anything blocking it? ","from":"developer"},{"body":"[~jonmeredith] , the fix versions mention 4.0, but I do not see a branch for it. Do you plan to apply it on that branch too?","from":"developer"},{"body":"I'm double-checking at the moment. I think the bug became an issue after some restructuring for the simulator, but need to convince myself of that. \r\n\r\nAlso, I've deleted the original branches/PRs as I accidentally referenced ones from a different fix. I'll post new ones in a day or so.","from":"developer"},{"body":"Hey [~jonmeredith], I haven't tried myself, but I have people telling me they reproduced the issue on 4.0.\r\n\r\nI can try to ask for more details if that will help you.  ","from":"developer"},{"body":"As expected, the (modified) tests fail on 4.0 so I've backported there too. I'll try and post test runs tomorrow but wanted to share the changes in case anybody was waiting on it.\r\n\r\nPRs\r\nTrunk https://github.com/apache/cassandra/pull/3250\r\n5.0 https://github.com/apache/cassandra/pull/3251\r\n4.1 https://github.com/apache/cassandra/pull/3252\r\n4.0 https://github.com/apache/cassandra/pull/3253","from":"developer"},{"body":"I left a few nits in the 4.0 PR, but also approved all 4 PRs.\r\n\r\n+1","from":"developer"},{"body":"Starting commit. I believe all of the test failures are known flakes or unrelated issues. This patch has been running in production for me for many months without issue so I have high confidence in it.","from":"developer"}],"created":"2023-08-08T19:59:55.000+0000","description":"On restart, Cassandra logs this message and terminates.\r\n{code:java}\r\nERROR 2023-07-17T17:17:22,931 [main] org.apache.cassandra.db.lifecycle.LogTransaction:561 - Unexpected disk state: failed to read transaction log [nb_txn_stream_39d5f6b0-fb81-11ed-8f46-e97b3f61511e.log in /datadir1/keyspace/table-c9527530a0d611e8813f034699fc9043]\r\nFiles and contents follow:\r\n/datadir1/keyspace/table-c9527530a0d611e8813f034699fc9043/nb_txn_stream_39d5f6b0-fb81-11ed-8f46-e97b3f61511e.log\r\n ABORT:[,0,0][737437348]\r\n ***This record should have been the last one in all replicas\r\n ADD:[/datadir1/keyspace/table-c9527530a0d611e8813f034699fc9043/nb-284490-big-,0,8][2493503833]\r\n{code}\r\nThe root cause is a race during streaming exception handling.\r\n\r\nAlthough concurrent modification of to the {{LogTransaction}} was added for CASSANDRA-16225, there is nothing to prevent usage after the transaction is completed (committed/aborted) once it has been processed by {{TransactionTidier}} (after the last reference is released). Before the transaction is tidied, the {{LogFile}} keeps a list of records that are checked for completion before adding new entries. In {{TransactionTidier}} {{LogFile.records}} are cleared as no longer needed, however the LogTransaction/LogFile is still accessible to the stream.\r\n\r\nThe changes in CASSANDRA-17273 added a parallel set of {{onDiskRecords}} that could be used to reliably recreate the transaction log at any new datadirs the same as the existing\r\ndatadirs - regardless of the effect of {{LogTransaction.untrackNew/LogFile.remove}}\r\n\r\nIf a streaming exception causes the LogTransaction to be aborted and tidied just before {{SimpleSSTableMultiWriter}} calls trackNew to add a new sstable. At the time of the call, the {{LogFile}} will not contain any {{LogReplicas}},\r\n{{LogFile.records}} will be empty, and {{LogFile.onDiskRecords}} will contain an {{ABORT}}.\r\n\r\nWhen {{LogTransaction.trackNew/LogFile.add}} is called, the check for completed transaction fails as records is empty, there are no replicas on the datadir, so {{maybeCreateReplicas}} creates a new txnlog file replica containing ABORT, then\r\nappends an ADD record.\r\n\r\nThe LogFile has already been tidied after the abort so the txnlog file is not removed and sits on disk until a restart, causing the faiulre.\r\n\r\nThere is a related exception caused with a different interleaving of aborts, after an sstable is added, however this is just a nuisance in the logs as the LogRelica is already created with an {{ADD}} record first.\r\n{code:java}\r\njava.lang.AssertionError: [ADD:[/datadir1/keyspace/table/nb-23314378-big-,0,8][1869379820]] is not tracked by 55be35b0-35d1-11ee-865d-8b1e3c48ca06\r\n at org.apache.cassandra.db.lifecycle.LogFile.remove(LogFile.java:388)\r\n at org.apache.cassandra.db.lifecycle.LogTransaction.untrackNew(LogTransaction.java:158)\r\n at org.apache.cassandra.db.lifecycle.LifecycleTransaction.untrackNew(LifecycleTransaction.java:577)\r\n at org.apache.cassandra.db.streaming.CassandraStreamReceiver$1.untrackNew(CassandraStreamReceiver.java:149)\r\n at org.apache.cassandra.io.sstable.SimpleSSTableMultiWriter.abort(SimpleSSTableMultiWriter.java:95)\r\n at org.apache.cassandra.io.sstable.format.RangeAwareSSTableWriter.abort(RangeAwareSSTableWriter.java:191)\r\n at org.apache.cassandra.db.streaming.CassandraCompressedStreamReader.read(CassandraCompressedStreamReader.java:115)\r\n at org.apache.cassandra.db.streaming.CassandraIncomingFile.read(CassandraIncomingFile.java:85)\r\n at org.apache.cassandra.streaming.messages.IncomingStreamMessage$1.deserialize(IncomingStreamMessage.java:53)\r\n at org.apache.cassandra.streaming.messages.IncomingStreamMessage$1.deserialize(IncomingStreamMessage.java:38)\r\n at org.apache.cassandra.streaming.messages.StreamMessage.deserialize(StreamMessage.java:53)\r\n at org.apache.cassandra.streaming.async.StreamingInboundHandler$StreamDeserializingTask.run(StreamingInboundHandler.java:172)\r\n at io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n at java.base/java.lang.Thread.run(Thread.java:829)\r\n{code}","issue_id":"13546581","key":"CASSANDRA-18736","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2024-04-25T20:18:28.000+0000","role":"fixed_distractor","summary":"Streaming exception race creates corrupt transaction log files that prevent restart"} {"case_id":"12494462","cluster":"DISTRACTOR-CASSANDRA-1927","comments":[{"body":"Client side (hadoop job):\n\njava.io.IOException: Could not get input splits\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.getSplits(ColumnFamilyInputFormat.java:127)\n\tat org.apache.hadoop.mapred.JobClient.writeNewSplits(JobClient.java:885)\n\tat org.apache.hadoop.mapred.JobClient.submitJobInternal(JobClient.java:779)\n\tat org.apache.hadoop.mapreduce.Job.submit(Job.java:432)\n\tat org.apache.hadoop.mapreduce.Job.waitForCompletion(Job.java:447)\n\tat no.finntech.countstats.reduce.FakeAdCounterTableReduce.run(FakeAdCounterTableReduce.java:421)\n\tat org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:65)\n\tat org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:79)\n\tat no.finntech.countstats.reduce.FakeAdCounterTableReduce.main(FakeAdCounterTableReduce.java:75)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.hadoop.util.RunJar.main(RunJar.java:156)\nCaused by: java.util.concurrent.ExecutionException: java.io.IOException: unable to connect to server\n\tat java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n\tat java.util.concurrent.FutureTask.get(FutureTask.java:83)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.getSplits(ColumnFamilyInputFormat.java:123)\n\t... 13 more\nCaused by: java.io.IOException: unable to connect to server\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.createConnection(ColumnFamilyInputFormat.java:212)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.getSubSplits(ColumnFamilyInputFormat.java:187)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.access$200(ColumnFamilyInputFormat.java:74)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat$SplitCallable.call(ColumnFamilyInputFormat.java:160)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat$SplitCallable.call(ColumnFamilyInputFormat.java:145)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:619)\nCaused by: org.apache.thrift.transport.TTransportException: java.net.ConnectException: Connection refused\n\tat org.apache.thrift.transport.TSocket.open(TSocket.java:185)\n\tat org.apache.thrift.transport.TFramedTransport.open(TFramedTransport.java:81)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.createConnection(ColumnFamilyInputFormat.java:208)\n\t... 9 more\nCaused by: java.net.ConnectException: Connection refused\n\tat java.net.PlainSocketImpl.socketConnect(Native Method)\n\tat java.net.PlainSocketImpl.doConnect(PlainSocketImpl.java:333)\n\tat java.net.PlainSocketImpl.connectToAddress(PlainSocketImpl.java:195)\n\tat java.net.PlainSocketImpl.connect(PlainSocketImpl.java:182)\n\tat java.net.SocksSocketImpl.connect(SocksSocketImpl.java:366)\n\tat java.net.Socket.connect(Socket.java:525)\n\tat org.apache.thrift.transport.TSocket.open(TSocket.java:180)\n\t... 11 more\n\n","created":"2011-01-03T08:25:34.548+0000"},{"body":"There's a todo comment in ColumnFamilyInputFormat\n // TODO handle failure of range replicas & retry\n\nline 198 only tries the first endpoint. a loop on the TException trying the next endpoint is needed.","created":"2011-01-03T08:27:16.207+0000"},{"body":"Utku: are you able to test this patch?\n\n( It didn't work for me because RF was never really set to 3. using cassandra-cli \"describe keyspace xxx\" reported \"Replication Factor: 1\" ) :-$\n\n","created":"2011-01-03T10:45:29.983+0000"},{"body":"Mck: Right now I can't access to our compilation server. However I can replace the running binaries and test them if I have the patched rc4. Can you somehow provide me the compiled package?","created":"2011-01-03T11:00:35.396+0000"},{"body":"Sent DM. If it doesn't work you should at minimum see the job's IOException stacktrace change from \"unable to connect to server\" to \"failed connecting to all endpoints\".","created":"2011-01-03T11:08:00.682+0000"},{"body":"I'll be testing it in a few hours. I'll write down the results. something urgent came up.","created":"2011-01-03T13:58:27.758+0000"},{"body":"After fixing my local RF problem this patch works for me.","created":"2011-01-03T15:53:31.683+0000"},{"body":"It looks like this patch includes the code from CASSANDRA-1921, which is causing conflicts b/c it's already applied on 0.7 and trunk. Can you create a patch for 1927 only?","created":"2011-01-03T17:25:12.112+0000"},{"body":"Putting Stu as reviewer since he was for CASSANDRA-342 (which the TODO comment in question was added under).","created":"2011-01-03T17:27:18.239+0000"},{"body":"Yeah, the patch had a lot of crap in it. sorry. will re-apply.","created":"2011-01-03T17:34:21.388+0000"},{"body":"correct patch & license grant","created":"2011-01-03T17:38:54.176+0000"},{"body":"third time lucky. removed unnecessary import.","created":"2011-01-03T17:41:20.536+0000"},{"body":"I've tested against the rc4+patch and it works.","created":"2011-01-03T17:54:09.054+0000"},{"body":"committed, thanks!","created":"2011-01-03T19:43:30.549+0000"}],"conversations":[{"body":"using the same directives in the sample code:\n\nWhen I start the CFInputFormat to read a CF in a keyspace of RF=3 on a 4-node cluster:\n- If all the nodes are all up, everything works fine and I don't have any problems walking through the all data in the CF, however\n- If there's a node down, the hadoop job does not even start, just dies without any errors or exceptions.\n\nSo I'm really sorry for not being able to post any errors or exceptions, though it's really easy to reproduce. Just startup a cluster and take one node down and you're there :)","from":"reporter","subject":"Hadoop Integration doesn't work when one node is down"},{"body":"Client side (hadoop job):\n\njava.io.IOException: Could not get input splits\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.getSplits(ColumnFamilyInputFormat.java:127)\n\tat org.apache.hadoop.mapred.JobClient.writeNewSplits(JobClient.java:885)\n\tat org.apache.hadoop.mapred.JobClient.submitJobInternal(JobClient.java:779)\n\tat org.apache.hadoop.mapreduce.Job.submit(Job.java:432)\n\tat org.apache.hadoop.mapreduce.Job.waitForCompletion(Job.java:447)\n\tat no.finntech.countstats.reduce.FakeAdCounterTableReduce.run(FakeAdCounterTableReduce.java:421)\n\tat org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:65)\n\tat org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:79)\n\tat no.finntech.countstats.reduce.FakeAdCounterTableReduce.main(FakeAdCounterTableReduce.java:75)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.hadoop.util.RunJar.main(RunJar.java:156)\nCaused by: java.util.concurrent.ExecutionException: java.io.IOException: unable to connect to server\n\tat java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n\tat java.util.concurrent.FutureTask.get(FutureTask.java:83)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.getSplits(ColumnFamilyInputFormat.java:123)\n\t... 13 more\nCaused by: java.io.IOException: unable to connect to server\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.createConnection(ColumnFamilyInputFormat.java:212)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.getSubSplits(ColumnFamilyInputFormat.java:187)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.access$200(ColumnFamilyInputFormat.java:74)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat$SplitCallable.call(ColumnFamilyInputFormat.java:160)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat$SplitCallable.call(ColumnFamilyInputFormat.java:145)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:619)\nCaused by: org.apache.thrift.transport.TTransportException: java.net.ConnectException: Connection refused\n\tat org.apache.thrift.transport.TSocket.open(TSocket.java:185)\n\tat org.apache.thrift.transport.TFramedTransport.open(TFramedTransport.java:81)\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.createConnection(ColumnFamilyInputFormat.java:208)\n\t... 9 more\nCaused by: java.net.ConnectException: Connection refused\n\tat java.net.PlainSocketImpl.socketConnect(Native Method)\n\tat java.net.PlainSocketImpl.doConnect(PlainSocketImpl.java:333)\n\tat java.net.PlainSocketImpl.connectToAddress(PlainSocketImpl.java:195)\n\tat java.net.PlainSocketImpl.connect(PlainSocketImpl.java:182)\n\tat java.net.SocksSocketImpl.connect(SocksSocketImpl.java:366)\n\tat java.net.Socket.connect(Socket.java:525)\n\tat org.apache.thrift.transport.TSocket.open(TSocket.java:180)\n\t... 11 more\n\n","from":"developer"},{"body":"There's a todo comment in ColumnFamilyInputFormat\n // TODO handle failure of range replicas & retry\n\nline 198 only tries the first endpoint. a loop on the TException trying the next endpoint is needed.","from":"developer"},{"body":"Utku: are you able to test this patch?\n\n( It didn't work for me because RF was never really set to 3. using cassandra-cli \"describe keyspace xxx\" reported \"Replication Factor: 1\" ) :-$\n\n","from":"developer"},{"body":"Mck: Right now I can't access to our compilation server. However I can replace the running binaries and test them if I have the patched rc4. Can you somehow provide me the compiled package?","from":"developer"},{"body":"Sent DM. If it doesn't work you should at minimum see the job's IOException stacktrace change from \"unable to connect to server\" to \"failed connecting to all endpoints\".","from":"developer"},{"body":"I'll be testing it in a few hours. I'll write down the results. something urgent came up.","from":"developer"},{"body":"After fixing my local RF problem this patch works for me.","from":"developer"},{"body":"It looks like this patch includes the code from CASSANDRA-1921, which is causing conflicts b/c it's already applied on 0.7 and trunk. Can you create a patch for 1927 only?","from":"developer"},{"body":"Putting Stu as reviewer since he was for CASSANDRA-342 (which the TODO comment in question was added under).","from":"developer"},{"body":"Yeah, the patch had a lot of crap in it. sorry. will re-apply.","from":"developer"},{"body":"correct patch & license grant","from":"developer"},{"body":"third time lucky. removed unnecessary import.","from":"developer"},{"body":"I've tested against the rc4+patch and it works.","from":"developer"},{"body":"committed, thanks!","from":"developer"}],"created":"2011-01-03T01:15:15.000+0000","description":"using the same directives in the sample code:\n\nWhen I start the CFInputFormat to read a CF in a keyspace of RF=3 on a 4-node cluster:\n- If all the nodes are all up, everything works fine and I don't have any problems walking through the all data in the CF, however\n- If there's a node down, the hadoop job does not even start, just dies without any errors or exceptions.\n\nSo I'm really sorry for not being able to post any errors or exceptions, though it's really easy to reproduce. Just startup a cluster and take one node down and you're there :)","issue_id":"12494462","key":"CASSANDRA-1927","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-01-03T19:43:30.000+0000","role":"fixed_distractor","summary":"Hadoop Integration doesn't work when one node is down"} {"case_id":"13580386","cluster":"DISTRACTOR-CASSANDRA-19661","comments":[{"body":"The schema alone is not enough to reproduce, but it looks like this should block 5.0-rc.","created":"2024-05-24T17:38:34.285+0000"},{"body":"[~adelapena] [~jbellis] [~mike_tr_adamson] It looks like perhaps {{computeRowIds()}} is being called multiple times when it shouldn't via {{OnHeapGraph#writeData()}}? Is it possible for multiple keys in {{postingsMap}} to point to the same {{VectorPostings}} instance?","created":"2024-05-24T18:12:45.189+0000"},{"body":"> It looks like perhaps {{computeRowIds()}} is being called multiple times when it shouldn't\r\n\r\nAgreed that it's not intended to be called multiple times.\r\n\r\n> Is it possible for multiple keys in {{postingsMap}} to point to the same {{VectorPostings}} instance?\r\n\r\nI don't think that's possible – the only mutation against `postingsMap` is done with a freshly instantiated `VectorPostings`.","created":"2024-05-24T19:23:14.077+0000"},{"body":"A couple of observations.\r\n\r\nFirstly, the stacktrace provided is not consistent with the latest 5.0 codebase across a number of files. This means it is quite old. It would be good to get a reproduction on the latest 5.0.\r\n\r\nSecondly, it would be good to get a more logging around the error to see what was going on at the time. \r\n\r\nI agree with [~jbellis] in that it only appears possible to call VectorPostings.computeRowIds once. This makes me think that we are attempting to flush the memtable index more than once but I've no idea how that could be happening either.","created":"2024-05-27T13:41:44.131+0000"},{"body":"I think I saw something like this very early in 5.0 development cycle but I can not reproduce that anymore.\r\n\r\nI wrote this trainer / prompt (1) which takes the sources or Cassandra, trains them and one can ask questions against Cassandra codebase. It uses cassio and langchain. I can not replicate the issue. After the training I can just restart it fine. \r\n\r\nIf you want to try that, you need to pay to OpenAI to train it and put the token into .env file where that script is like \"OPENAI_API_KEY=...\"\r\n\r\n[https://gist.github.com/smiklosovic/98b52502df8061ec1d8c3546cf489628]\r\n\r\nI am using the latest 5.0 branch.\r\n\r\nI am *not* saying this is not an issue. I am just saying that what I tried it on has not produced that error. ","created":"2024-05-29T09:53:55.164+0000"},{"body":"[~sergrua] Can you reproduce the problem with the latest C* version?\r\n","created":"2024-05-31T14:22:30.192+0000"},{"body":"I was hoping I could drive this forward by running through https://docs.llamaindex.ai/en/stable/examples/vector_stores/CassandraIndexDemo/ and getting a reproduction against beta1 or HEAD, but it does not reproduce. Adding all the text files from test/resources/tokenization/ does not help either.","created":"2024-06-04T15:54:56.132+0000"},{"body":"I'll try to find some time this week to write the code to reproduce it. One thing I failed to mention and it may be important is that it's a 3-node cluster with `NetworkTopologyStrategy`","created":"2024-06-04T16:54:59.396+0000"},{"body":"Thanks [~sergrua] , does this mean you were able to reproduce it also on HEAD and the issue was not fixed by chance since beta1?","created":"2024-06-04T16:56:43.028+0000"},{"body":"No, I'm still on beta-1 [~e.dimitrova] ","created":"2024-06-04T18:17:37.826+0000"},{"body":"Hi all. I haven't been able to reproduce it on the latest build. I'm attaching a cut down version of the code I used.","created":"2024-06-10T09:05:42.718+0000"},{"body":"[~sergrua] what documents are you loading from the ./data directory?","created":"2024-06-10T11:05:06.922+0000"},{"body":"[~brandon.williams] they are just txt files, nothing special.","created":"2024-06-10T11:20:57.684+0000"},{"body":"[~sergrua] is it possible to upload them? I think there must be something endemic to the data that caused this, since your script alone is not enough to reproduce against beta1.","created":"2024-06-10T16:14:59.526+0000"},{"body":"[~brandon.williams] Unfortunately it's not possible. It's not public data.","created":"2024-06-10T18:40:28.700+0000"},{"body":"I see. Well if you are certain this only affected beta1 and no longer reproduces then I think we can close this ticket.","created":"2024-06-10T22:08:17.903+0000"},{"body":"I'm reopening this ticket as I've reproduced exactly the same issue in 5.0.2.\r\n\r\nBasically, I also run a cluster of 3 nodes, I also use a vector index '', and my nodes are also unable to restart after data is written into that index.\r\n\r\nOne table is loaded with a large amount of vectors (millions), which seems to make Cassandra unstable, which eventually makes a node stop and unable to restart, presenting the same logs as above. Only solution is then to reset it.\r\n\r\nI haven't investigated Cassandra's code in detail, but in practice, if I don't write into the vector-indexed column, my cluster is suddenly very stable with no problem restarting.\r\n\r\nI've tried several implementations of Memtable and Compaction Strategies, which made no difference.\r\n\r\nI'm unfortunately also unable to provide data as it's non public, but I'll try to send some logs as soon as I can.","created":"2025-02-10T20:39:11.748+0000"},{"body":"bq. my nodes are also unable to restart\r\n\r\nAre you receiving the same error as in the description?","created":"2025-02-10T20:59:04.938+0000"},{"body":"I get the same error, yes.","created":"2025-02-10T21:08:27.486+0000"},{"body":"Here are some logs from last time I tried using vectors.\r\n[^5.0.2_fail_memtableflush_vector_full.txt]\r\n\r\nThe error is essentially the same.\r\n{noformat}\r\nERROR [MemtableFlushWriter:464] 2025-01-13 13:48:27,449 MemtableIndexWriter.java:157 - [default.docs.vector_search_index] Error while flushing index null\r\njava.lang.IllegalStateException: null\r\n\tat com.google.common.base.Preconditions.checkState(Preconditions.java:496)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.VectorPostings.computeRowIds(VectorPostings.java:76)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.OnHeapGraph.writeData(OnHeapGraph.java:313)\r\n\tat org.apache.cassandra.index.sai.memory.VectorMemoryIndex.writeDirect(VectorMemoryIndex.java:271)\r\n\tat org.apache.cassandra.index.sai.memory.MemtableIndex.writeDirect(MemtableIndex.java:113)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.flushVectorIndex(MemtableIndexWriter.java:200)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.complete(MemtableIndexWriter.java:141)\r\n\tat org.apache.cassandra.index.sai.disk.StorageAttachedIndexWriter.complete(StorageAttachedIndexWriter.java:185)\r\n\tat java.base/java.util.ArrayList.forEach(ArrayList.java:1541)\r\n\tat java.base/java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1085)\r\n\tat org.apache.cassandra.io.sstable.format.SSTableWriter.commit(SSTableWriter.java:289)\r\n\tat org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit(ShardedMultiWriter.java:219)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1354)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1253)\r\n\tat org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)\r\n\tat io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n\tat java.base/java.lang.Thread.run(Thread.java:829)\r\n{noformat}\r\nI'm currently setting up another cluster to test this, because I can't have nodes get in that state on my main cluster.","created":"2025-02-11T20:45:43.555+0000"},{"body":"I hadn't had a look at this for a few months, but last week we set up a new cluster and I decided to try again : after >10 millions vectors the problem reappeared, lost 2 nodes out of 8, unable to restart without a data wipe.\r\n\r\nI won't be able to test further as we're going to have to use another database for storing vectors, unfortunately.\r\nI'm surprised not more people have shared this issue. If any of you has successfully indexed 100 million - 1 billion vectors on Cassandra, I'd be very curious to know what settings you used.","created":"2025-05-19T18:13:53.443+0000"},{"body":"I can confirm that this reproduces on 5.0.6.  I attempted to load 100 million embeddings and upon startup there's a high risk of the same exception:\r\n\r\n{code}\r\nERROR [MemtableFlushWriter:4] 2026-01-28 18:29:16,976 StorageAttachedIndexWriter.java:220 - [vector_bench.vectors_cbp.*] Failed to complete an index build\r\njava.lang.IllegalStateException: null\r\n\tat com.google.common.base.Preconditions.checkState(Preconditions.java:496)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.VectorPostings.computeRowIds(VectorPostings.java:76)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.OnHeapGraph.writeData(OnHeapGraph.java:315)\r\n\tat org.apache.cassandra.index.sai.memory.VectorMemoryIndex.writeDirect(VectorMemoryIndex.java:272)\r\n\tat org.apache.cassandra.index.sai.memory.MemtableIndex.writeDirect(MemtableIndex.java:113)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.flushVectorIndex(MemtableIndexWriter.java:212)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.complete(MemtableIndexWriter.java:143)\r\n\tat org.apache.cassandra.index.sai.disk.StorageAttachedIndexWriter.complete(StorageAttachedIndexWriter.java:212)\r\n\tat java.base/java.util.ArrayList.forEach(ArrayList.java:1511)\r\n\tat java.base/java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1092)\r\n\tat org.apache.cassandra.io.sstable.format.SSTableWriter.commit(SSTableWriter.java:295)\r\n\tat org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit(ShardedMultiWriter.java:219)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1354)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1253)\r\n\tat org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)\r\n\tat io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n\tat java.base/java.lang.Thread.run(Thread.java:833)\r\nWARN [MemtableFlushWriter:4] 2026-01-28 18:29:16,976 MemtableIndexWriter.java:117 - [vector_bench.vectors_cbp.vectors_cbp_embedding_idx] Aborting index memtable flush for /app/cassandra/data/vector_bench/vectors_cbp-3c1d51e0fbcc11f08c7baf9115bb8d33/oa-3gxg_1fbu_5az002afoxgv570894-big...\r\njava.lang.IllegalStateException: null\r\n\tat com.google.common.base.Preconditions.checkState(Preconditions.java:496)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.VectorPostings.computeRowIds(VectorPostings.java:76)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.OnHeapGraph.writeData(OnHeapGraph.java:315)\r\n\tat org.apache.cassandra.index.sai.memory.VectorMemoryIndex.writeDirect(VectorMemoryIndex.java:272)\r\n\tat org.apache.cassandra.index.sai.memory.MemtableIndex.writeDirect(MemtableIndex.java:113)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.flushVectorIndex(MemtableIndexWriter.java:212)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.complete(MemtableIndexWriter.java:143)\r\n\tat org.apache.cassandra.index.sai.disk.StorageAttachedIndexWriter.complete(StorageAttachedIndexWriter.java:212)\r\n\tat java.base/java.util.ArrayList.forEach(ArrayList.java:1511)\r\n\tat java.base/java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1092)\r\n\tat org.apache.cassandra.io.sstable.format.SSTableWriter.commit(SSTableWriter.java:295)\r\n\tat org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit(ShardedMultiWriter.java:219)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1354)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1253)\r\n\tat org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)\r\n\tat io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n\tat java.base/java.lang.Thread.run(Thread.java:833)\r\n {code}\r\n\r\nEdit: In my case, wiping commit logs allowed the node to start up successfully as the error was thrown during significant commit log replay.","created":"2026-01-28T18:39:50.884+0000"},{"body":"[~andrewweaver] does it reproduce if you set memtable_flush_writers: 1 in conf/cassandra.yaml ?","created":"2026-01-28T20:41:30.722+0000"},{"body":"I just confirmed that it reproduces the {{java.lang.IllegalStateException}} with {{memtable_flush_writers: 1}}.","created":"2026-01-29T19:42:05.989+0000"},{"body":"If possible could you please share system.log + debug.log for a failed Cassandra startup since the beginning of the start and till the failure?","created":"2026-01-29T19:48:39.154+0000"},{"body":"Sure thing: [^logs.tar.gz] \r\n\r\n{{2026-01-29 09:57:29,176}} Crashed\r\n{{2026-01-29 09:58:59,215}} Restarted \r\n{{2026-01-29 11:36:44,951}} Error while replaying commitlogs, at 14:06, it finally exits with an error.","created":"2026-01-29T20:51:43.846+0000"},{"body":"Thank you for the logs.\r\nI think I have a theory about a root cause for this issue.\r\nInitially I thought that it could be a kind of concurrency issue when two different flusher threads are trying to write the same index but I've not found in the code how it could happen + based on the last comments the issue is reproduced with memtable_flush_writers=1 as well.\r\n\r\nThe second idea was about a cycle here when we invoke writer.commit for every Flushing.FlushRunnable\r\n{code:java}\r\nfor (SSTableMultiWriter writer : flushResults)\r\n{\r\n accumulate = writer.commit(accumulate);\r\n metric.flushSizeOnDisk.update(writer.getOnDiskBytesWritten());\r\n}\r\n{code}\r\nit could happen (and maybe it is a real issue too) if we have multiple data directories and a memtable if flushed to several SSTables to them.\r\nBut, based on the provided log it is not the case - we have a single data directory:\r\n{code:java}\r\nDEBUG [main] 2026-01-29 09:59:01,308 DiskBoundaryManager.java:57 - Updating boundaries from null to DiskBoundaries{directories=[DataDirectory{location=/app/cassandra/data}],\r\n{code}\r\nAfter looking to the stacktrace more precisely I've noticed that we use org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit\r\nwhich has multiple inner writers:\r\n{code:java}\r\n@Override\r\npublic Throwable commit(Throwable accumulate)\r\n{\r\n Throwable t = accumulate;\r\n for (SSTableWriter writer : writers)\r\n if (writer != null)\r\n t = writer.commit(t);\r\n return t;\r\n}\r\n{code}\r\nso, if we have more than 1 writer in our case here we will invoke the commit method several times for the same observers:\r\n{code:java}\r\nobservers.forEach(SSTableFlushObserver::complete);\r\n{code}\r\nthe observers are retrieved in the same way during org.apache.cassandra.io.sstable.format.SSTableWriter#SSTableWriter constructing:\r\n{code:java}\r\nSSTableFlushObserver observer = group.getFlushObserver(descriptor, lifecycleNewTracker, metadata.getLocal());\r\n{code}\r\nThe number of shards is calculated by UCS here: org.apache.cassandra.db.compaction.UnifiedCompactionStrategy#createSSTableMultiWriter\r\n\r\nand this logic is applicable to flushing writers too:\r\n{code:java}\r\ndouble flushDensity = cfs.metric.flushSizeOnDisk.get() * shardManager.shardSetCoverage() / shardManager.localSpaceCoverage();\r\nShardTracker boundaries = shardManager.boundaries(controller.getNumShards(flushDensity));\r\n{code}\r\nit is dynamic and depends on data distribution/density.\r\nWe have a debug log printed by controller.getNumShards, so we can find in the debug.log a proof for this idea:\r\n{code:java}\r\nINFO [NativePoolCleaner] 2026-01-29 11:34:52,296 ColumnFamilyStore.java:1052 - Enqueuing flush of vector_bench.vectors_cbp, Reason: MEMTABLE_LIMIT, Usage: 2.000GiB (20%) on-heap, 1.256GiB (13%) off-heap\r\nDEBUG [MemtableFlushWriter:5] 2026-01-29 11:34:53,071 Controller.java:348 - Shard count 4 for density 1.241 GiB, 1.8334467342069816 times target 693.273 MiB\r\n{code}\r\n=====\r\n\r\nAs a result, it looks like the issue is caused by combination of 2 features: UCS which can flush a memtable to multiple sharded SSTables and SAI vector index which does not support sharding (https://issues.apache.org/jira/browse/CASSANDRA-20752 - a related change but it looks like it does not cover Vector index scenario).\r\n\r\nSo, if the theory is correct then the possible WA can be switch compaction from UCS to another strategy for this table or to force UCS to use a single shard. Based on org.apache.cassandra.db.compaction.unified.Controller#getNumShards code we can get it if we set min_sstable_size to 0 and base_shard_count to 1 in the table settings or the same via Java system properties: -Dunified_compaction.min_sstable_size=0MiB -D\r\nunified_compaction.base_shard_count=1 (it is for all tables but does not require a schema change)\r\n[~maedhroz] [~blambov]  - what do you think, am I missing something?","created":"2026-01-30T13:54:07.605+0000"},{"body":"Thank you for the quick analysis [~dnk].\r\n\r\nI'll try with 1 UCS shard and STCS if necessary.\r\n\r\nI've consistently observed a pattern where nodes get into a state where pending mutations grow at a steady rate. Unusual as my test is only doing up to about 200 writes/s per node and the rest of the nodes keep up. This seems to happen after:\r\n{code}\r\nDEBUG [MemtablePostFlush:1] 2026-01-30 19:03:27,808 HeapUtils.java:133 - Heap dump creation on uncaught exceptions is disabled.\r\nERROR [MemtablePostFlush:1] 2026-01-30 19:03:27,808 JVMStabilityInspector.java:70 - Exception in thread Thread[MemtablePostFlush:1,5,MemtablePostFlush]\r\njava.lang.NullPointerException: Cannot invoke \"java.lang.Boolean.booleanValue()\" because \"res\" is null\r\n at org.apache.cassandra.utils.memory.MemtableCleanerThread$Clean.apply(MemtableCleanerThread.java:97)\r\n{code}","created":"2026-01-30T19:40:17.178+0000"},{"body":"I verified that it reproduces with {{base_shard_count}} of 1.\r\n\r\n{code}\r\ncompaction = {'base_shard_count': '1', 'class': 'org.apache.cassandra.db.compaction.UnifiedCompactionStrategy', 'max_sstables_to_compact': '64', 'min_sstable_size': '100MiB', 'scaling_parameters': 'T4', 'sstable_growth': '0.3333', 'target_sstable_size': '1GiB'}\r\n{code}\r\n\r\nTrying STCS.","created":"2026-01-31T02:34:14.877+0000"},{"body":"For UCS it should be such combination (I forgot to mentioned sstable_growth):\r\n{code:java}\r\n'min_sstable_size' : '0MiB', 'base_shard_count': '1', 'sstable_growth' : '1'{code}\r\na message like this should be printed in the log before a flush for the table:\r\n{code:java}\r\nShard count 1 for density {some number} in fixed shards mode {code}\r\n====\r\nRegarding\r\n{quote}I've consistently observed a pattern where nodes get into a state where pending mutations grow at a steady rate.\r\n{quote}\r\nIn the shared logs I see a combination of 3 subsequent errors:\r\n{code:java}\r\nERROR [MemtablePostFlush:1] 2026-01-29 11:38:04,159 JVMStabilityInspector.java:70 - Exception in thread Thread[MemtablePostFlush:1,5,MemtablePostFlush]\r\njava.lang.IllegalStateException: null\r\n\tat com.google.common.base.Preconditions.checkState(Preconditions.java:496)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.VectorPostings.computeRowIds(VectorPostings.java:76)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.OnHeapGraph.writeData(OnHeapGraph.java:315)\r\n\tat org.apache.cassandra.index.sai.memory.VectorMemoryIndex.writeDirect(VectorMemoryIndex.java:272)\r\n\tat org.apache.cassandra.index.sai.memory.MemtableIndex.writeDirect(MemtableIndex.java:113)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.flushVectorIndex(MemtableIndexWriter.java:212)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.complete(MemtableIndexWriter.java:143)\r\n\tat org.apache.cassandra.index.sai.disk.StorageAttachedIndexWriter.complete(StorageAttachedIndexWriter.java:212)\r\n\tat java.base/java.util.ArrayList.forEach(ArrayList.java:1511)\r\n\tat java.base/java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1092)\r\n\tat org.apache.cassandra.io.sstable.format.SSTableWriter.commit(SSTableWriter.java:295)\r\n\tat org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit(ShardedMultiWriter.java:219)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1354)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1253)\r\n\tat org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)\r\n\tat io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n\tat java.base/java.lang.Thread.run(Thread.java:833)\r\n\tSuppressed: java.lang.IllegalStateException: null\r\n\t\t... 19 common frames omitted\r\n\tSuppressed: java.lang.IllegalStateException: null\r\n\t\t... 19 common frames omitted\r\n\r\n{code}\r\nthen\r\n{code:java}\r\nERROR [MemtablePostFlush:1] 2026-01-29 11:38:04,159 JVMStabilityInspector.java:70 - Exception in thread Thread[MemtablePostFlush:1,5,MemtablePostFlush]\r\njava.lang.NullPointerException: Cannot invoke \"java.lang.Boolean.booleanValue()\" because \"res\" is null\r\n\tat org.apache.cassandra.utils.memory.MemtableCleanerThread$Clean.apply(MemtableCleanerThread.java:97)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList$CallbackBiConsumerListener.run(ListenerList.java:244)\r\n\tat org.apache.cassandra.concurrent.ImmediateExecutor.execute(ImmediateExecutor.java:140)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.safeExecute(ListenerList.java:166)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notifyListener(ListenerList.java:157)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList$CallbackBiConsumerListener.notifySelf(ListenerList.java:250)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.lambda$notifyExclusive$0(ListenerList.java:124)\r\n\tat org.apache.cassandra.utils.concurrent.IntrusiveStack.forEach(IntrusiveStack.java:195)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notifyExclusive(ListenerList.java:124)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notify(ListenerList.java:96)\r\n\tat org.apache.cassandra.utils.concurrent.AsyncFuture.trySet(AsyncFuture.java:104)\r\n\tat org.apache.cassandra.utils.concurrent.AbstractFuture.tryFailure(AbstractFuture.java:148)\r\n\tat org.apache.cassandra.utils.concurrent.AsyncPromise.tryFailure(AsyncPromise.java:139)\r\n\tat org.apache.cassandra.db.memtable.AbstractAllocatorMemtable.lambda$flushLargestMemtable$0(AbstractAllocatorMemtable.java:306)\r\n\tat org.apache.cassandra.concurrent.ImmediateExecutor.execute(ImmediateExecutor.java:140)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.safeExecute(ListenerList.java:166)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notifyListener(ListenerList.java:157)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList$RunnableWithExecutor.notifySelf(ListenerList.java:345)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.lambda$notifyExclusive$0(ListenerList.java:124)\r\n\tat org.apache.cassandra.utils.concurrent.IntrusiveStack.forEach(IntrusiveStack.java:195)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notifyExclusive(ListenerList.java:124)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notify(ListenerList.java:96)\r\n\tat org.apache.cassandra.utils.concurrent.AsyncFuture.trySet(AsyncFuture.java:104)\r\n\tat org.apache.cassandra.utils.concurrent.AbstractFuture.tryFailure(AbstractFuture.java:148)\r\n\tat org.apache.cassandra.concurrent.FutureTask.tryFailure(FutureTask.java:87)\r\n\tat org.apache.cassandra.concurrent.FutureTask.run(FutureTask.java:75)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)\r\n\tat io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n\tat java.base/java.lang.Thread.run(Thread.java:833)\r\n{code}\r\nand finally there are also messages about a resource leak:\r\n{code:java}\r\nERROR [Reference-Reaper] 2026-01-29 11:41:57,667 Ref.java:243 - LEAK DETECTED: a reference (class org.apache.cassandra.utils.concurrent.WrappedSharedCloseable$Tidy@344029653:[[OffHeapBitSet]]) to class org.apache.cassandra.utils.concurrent.WrappedSharedCloseable$Tidy@344029653:[[OffHeapBitSet]] was not released before the reference was garbage collected\r\nERROR [Reference-Reaper] 2026-01-29 11:41:57,669 Ref.java:243 - LEAK DETECTED: a reference (class org.apache.cassandra.io.util.MmappedRegions$Tidier@434613886:/app/cassandra/data/vector_bench/vectors_cbp-3c1d51e0fbcc11f08c7baf9115bb8d33/oa-3gxh_0w67_0rv342lhihvcwtdh2y-big-Index.db) to class org.apache.cassandra.io.util.MmappedRegions$Tidier@434613886:/app/cassandra/data/vector_bench/vectors_cbp-3c1d51e0fbcc11f08c7baf9115bb8d33/oa-3gxh_0w67_0rv342lhihvcwtdh2y-big-Index.db was not released before the reference was garbage collected\r\n{code}\r\nDo you observe the pending mutations grow issue only in such combination or there are other patterns too?\r\n\r\nAlso a thread dump would help here to see what are the mutation threads waiting for.","created":"2026-01-31T16:32:20.715+0000"},{"body":"Apologies about not being able to actively engage much here recently. (Trying to get some other work over the line.) The big thing I'm worried about is whether or not we can get into this state w/ UCS and vanilla SAI (say on an indexed text or numeric/scalar column). There is some work in progress to shard the Memtable indexes in CASSANDRA-18216, but I'm not sure if that is related here...","created":"2026-01-31T16:46:51.900+0000"},{"body":"For vanilla SAI, it looks like there was a fix here... - CASSANDRA-20752","created":"2026-01-31T16:51:17.095+0000"},{"body":"My impression was that CASSANDRA-20752 only touches compaction. In any case, the writer creation _should_ be creating a separate observer for each shard, i.e. a separate StorageAttachedIndexWriter/MemtableIndexWriter, and those should have commit() called independently on them...unless I'm missing something. ({{Index.Group#getFlushObserver()}} creates new {{SSTableFlushObserver}} instances.)","created":"2026-01-31T17:23:22.682+0000"},{"body":"So to elaborate on my previous comment, it feels like as long as the {{MemtableIndexWriter}} instances for the shards don't reuse anything important, they should be able to deal with a shared writer (even though the Memtable index itself isn't actually shared). Non-Vector, for instance, doesn't share {{RowMapping}} instances. {{MemtableIndexWriter#flushVectorIndex()}} does call into {{writeDirect()}}, which ultimately hits {{OnHeapGraph#writeData()}}. This touches the shared {{postingsMap}} to get {{VectorPostings}} and calls {{computeRowIds()}} on them, which doesn't look like it should happen twice, as {{rowIds}} is only supposed to be populated once. (It's not memoized. It's illegal to call it twice.) Am I getting warmer?","created":"2026-01-31T18:38:02.933+0000"},{"body":"Would it be appropriate to just make {{computeRowIds()}} idempotent?","created":"2026-01-31T18:43:19.507+0000"},{"body":"Confirmed this message with shard count set to 1: {{Shard count 1 for density 19.241 KiB, 1.8790509741811726E-4 times min size 100.000 MiB}}\r\n\r\nI tried again with STCS and it completed writing all 10 million embeddings without issue.\r\n\r\n{quote}\r\nDo you observe the pending mutations grow issue only in such combination or there are other patterns too?\r\n{quote}\r\nIt seems to be this pattern consistently.\r\n\r\nI can try to reproduce again and get a thread dump.\r\n\r\n","created":"2026-01-31T20:29:22.908+0000"},{"body":"I replicated the issue again and captured thread dumps. Note that I did have to undo some tuning that I did after switching to STCS to trigger it.\r\n{code:java}\r\nmemtable_flush_writers: 1 -> 2\r\nmemtable_offheap_space: default (~2GB) -> 1024MiB\r\nmemtable_cleanup_threshold: default (0.33) -> 0.15\r\n{code}\r\nMy goal was to get more frequent flushes.\r\n\r\n!screenshot-1.png|width=590,height=255!\r\n\r\nAttached are thread dumps taken every 5 minutes.\r\n[^10.103.220.89_thread_dump.tgz]","created":"2026-02-04T15:21:42.350+0000"},{"body":"So, am I right that the issue is observed only for the runs when you have observed the IllegalStateException mentioned before? ","created":"2026-02-04T21:03:52.269+0000"},{"body":"Here is a draft solution: https://github.com/apache/cassandra/pull/4605. It works by modifying UCS, while leaving the underlying vector design untouched.\r\n\r\nIn general, I see one key requirement: do not attempt to build a graph at flush time. It is expensive to build vector graphs, and we do not want to build up an excessive number of vectors in memory due to delays at flush time.\r\n\r\nTherefore, I see two options:\r\n\r\n1. Pre-shard the vector memtable index to align with flush shard bounds\r\n2. Prevent sharding at flush time. (This is the proposed solution in the draft PR. I haven't explicitly tested it yet because I want to get general feedback on the design before proceeding)\r\n\r\nFrom a vector perf perspective, we always want to build bigger graphs. The time complexity of graph search is `log( n )`. Many small graphs means we have `k * log( n )` performance and that significantly increases latency/reduces throughput. Given the benefit to read performance, I lean towards avoiding eager sharding.\r\n\r\nNote also that the `TrieMemoryIndex` differs from the `VectorMemoryIndex` in one significant way: one has a `synchronized` add method while the other does not. (Actually, I just noticed that the `VectorMemoryIndex` has the `synchronized` keyword, but doesn't need it! I just created this ticket as a follow up https://issues.apache.org/jira/browse/CASSANDRA-21160) I mention this because the primary benefit for pre-sharding the memtable index is to increase write throughput so that writes to different partitions can proceed independently. The jvector library allows for this, but the `TrieMemoryIndex` does not.","created":"2026-02-06T19:49:59.454+0000"},{"body":"JUnit test to reproduce the issue: https://github.com/netudima/cassandra/commit/94406f37f2f6c7b955358001a637d956adf5d064\r\n[~mmarshall] could you add it to your PR?","created":"2026-02-17T16:04:47.836+0000"},{"body":"An interim summary:\r\n * To fix the issue we have a patch from [~mmarshall] which enforces usage of 1 shard for UCS during a flushing if we have a vector index for a table. It should solve the reported issue.\r\n * An alternative could be a support of sharding for memory indexes (related story: CASSANDRA-18216) but it is much more complicated to implement + there are performance concerns regarding this way for vector indexes as it was mentioned before.\r\n * The fix does not cover a possible multiple data directories use case (when a flushing memtable is split  between multiple data directories) but it is much rare use case and I suppose it should be checked and fixed if needed separately, to not delay the current fix delivery.\r\n * A workaround for the issue is to switch to a non-UCS or to use UCS with the following configuration:\r\n{code:java}\r\ncompaction = {\r\n 'class': 'UnifiedCompactionStrategy',\r\n 'base_shard_count': '1',\r\n 'min_sstable_size' : '0MiB', \r\n 'sstable_growth' : '1'\r\n}\r\n{code}","created":"2026-02-17T16:19:12.212+0000"},{"body":"{quote}replicated the issue again and captured thread dumps.\r\n{quote}\r\nRegarding pending Mutations issue mentioned before. Assuming that we speak about the case when the IllegalStateException is thrown.\r\n\r\nThe unexpected IllegalStateException breaks reclaiming of memory used by the flushed memtable, so we consume all memory available for memtables and block on awaiting for it.\r\n\r\n[https://github.com/apache/cassandra/blob/cassandra-5.0.6/src/java/org/apache/cassandra/db/ColumnFamilyStore.java#L1377] - the place where reclaim logic is register to trigger at the end of flush. We don't reach this method invocation because the IllegalStateException is thrown before on writer commit step:\r\n\r\n[https://github.com/apache/cassandra/blob/cassandra-5.0.6/src/java/org/apache/cassandra/db/ColumnFamilyStore.java#L1354] \r\n\r\nHere is an example of a mutation thread  trying to allocate some memory for memtables:\r\n\r\n{{\"MutationStage-1\" #131 daemon prio=5 os_prio=0 cpu=949990.56ms elapsed=26753.90s tid=0x00007267aea3f2d0 nid=0xcc734 waiting on condition  [0x0000725ce31bc000]}}{{{}   java.lang.Thread.State: WAITING (parking){}}}{{{}at jdk.internal.misc.Unsafe.park(java.base@17.0.3/Native Method){}}}{{{}at java.util.concurrent.locks.LockSupport.park(java.base@17.0.3/LockSupport.java:341){}}}{{{}at org.apache.cassandra.utils.concurrent.WaitQueue$Standard$AbstractSignal.await(WaitQueue.java:321){}}}{{{}at org.apache.cassandra.utils.concurrent.WaitQueue$Standard$AbstractSignal.await(WaitQueue.java:299){}}}{{{}at org.apache.cassandra.utils.concurrent.Awaitable$Defaults.awaitThrowUncheckedOnInterrupt(Awaitable.java:131){}}}{{{}at *{color:#de350b}org.apache.cassandra.utils.concurrent.Awaitable$AbstractAwaitable.awaitThrowUncheckedOnInterrupt(Awaitable.java:235){color}*{}}}{*}{color:#de350b}{{at org.apache.cassandra.utils.memory.MemtableAllocator$SubAllocator.allocate(MemtableAllocator.java:195)}}{color}{*}{{{}*{color:#de350b}at{color}* org.apache.cassandra.db.memtable.AbstractAllocatorMemtable.markExtraOnHeapUsed(AbstractAllocatorMemtable.java:196){}}}{{{}at org.apache.cassandra.index.sai.StorageAttachedIndex$UpdateIndexer.adjustMemtableSize(StorageAttachedIndex.java:998){}}}{{{}at org.apache.cassandra.index.sai.StorageAttachedIndex$UpdateIndexer.updateRow(StorageAttachedIndex.java:990){}}}{{{}at org.apache.cassandra.index.sai.StorageAttachedIndexGroup$1.updateRow(StorageAttachedIndexGroup.java:191){}}}{{{}at org.apache.cassandra.index.SecondaryIndexManager$WriteTimeTransaction.onUpdated(SecondaryIndexManager.java:1570){}}}{{{}at org.apache.cassandra.db.partitions.BTreePartitionUpdater.merge(BTreePartitionUpdater.java:139){}}}{{{}at org.apache.cassandra.db.partitions.BTreePartitionUpdater.merge(BTreePartitionUpdater.java:39){}}}{{{}at org.apache.cassandra.utils.btree.BTree.updateLeaves(BTree.java:430){}}}{{{}at org.apache.cassandra.utils.btree.BTree.update(BTree.java:372){}}}{{{}at org.apache.cassandra.db.partitions.BTreePartitionUpdater.makeMergedPartition(BTreePartitionUpdater.java:89){}}}{{{}at org.apache.cassandra.db.partitions.BTreePartitionUpdater.mergePartitions(BTreePartitionUpdater.java:71){}}}{{{}at org.apache.cassandra.db.memtable.TrieMemtable$MemtableShard$$Lambda$1179/0x00000008011e2d58.apply(Unknown Source){}}}{{{}at org.apache.cassandra.db.tries.InMemoryTrie.applyContent(InMemoryTrie.java:930){}}}{{{}at org.apache.cassandra.db.tries.InMemoryTrie.putRecursive(InMemoryTrie.java:906){}}}{{{}at org.apache.cassandra.db.tries.InMemoryTrie.putRecursive(InMemoryTrie.java:910){}}}{{{}at {}}}\r\n\r\n{{...}}\r\n\r\n{{{}org.apache.cassandra.db.tries.InMemoryTrie.putRecursive(InMemoryTrie.java:897){}}}{{{}at org.apache.cassandra.db.tries.InMemoryTrie.putSingleton(InMemoryTrie.java:878){}}}{{{}at org.apache.cassandra.db.memtable.TrieMemtable$MemtableShard.put(TrieMemtable.java:480){}}}{{{}at org.apache.cassandra.db.memtable.TrieMemtable.put(TrieMemtable.java:190){}}}{{{}at org.apache.cassandra.db.ColumnFamilyStore.apply(ColumnFamilyStore.java:1474){}}}{{{}at org.apache.cassandra.db.CassandraTableWriteHandler.write(CassandraTableWriteHandler.java:38){}}}{{{}at org.apache.cassandra.db.Keyspace.applyInternal(Keyspace.java:653){}}}{{{}at org.apache.cassandra.db.Keyspace.applyFuture(Keyspace.java:474){}}}{{{}at org.apache.cassandra.db.Mutation.applyFuture(Mutation.java:244){}}}{{{}at org.apache.cassandra.hints.Hint.applyFuture(Hint.java:109){}}}{{{}at org.apache.cassandra.hints.HintVerbHandler.doVerb(HintVerbHandler.java:116){}}}{{{}at org.apache.cassandra.net.InboundSink.lambda$new$0(InboundSink.java:78){}}}{{{}at org.apache.cassandra.net.InboundSink$$Lambda$783/0x00000008010bf960.accept(Unknown Source){}}}{{{}at org.apache.cassandra.net.InboundSink.accept(InboundSink.java:97){}}}{{{}at org.apache.cassandra.net.InboundSink.accept(InboundSink.java:45){}}}{{{}at org.apache.cassandra.net.InboundMessageHandler$ProcessMessage.run(InboundMessageHandler.java:430){}}}{{{}at org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133){}}}{{{}at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:143){}}}{{{}at io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30){}}}{{{}at java.lang.Thread.run(java.base@17.0.3/Thread.java:833){}}}\r\n\r\n \r\n\r\nThe pending mutation issue should disappear once we fix the {{{}IllegalStateException{}}}. We could try to improve the error handling to better process such unexpected exceptions, but there may be no way to handle this completely correctly — we might even have to stop processing at all or risk data loss (at least for indexes). So I’m not sure whether investing significant effort into this will really pay off.","created":"2026-02-17T21:25:25.942+0000"},{"body":"+1 on the PR","created":"2026-02-25T17:19:26.967+0000"},{"body":"CI run for 5.0 is ok (its known flaky TestBootstrap only):\r\n* https://pre-ci.cassandra.apache.org/job/cassandra-5.0/86/\r\n* [^ci_summary_CASSANDRA-19661.html.zip] \r\n* [^results_details_CASSANDRA-19661.tar.xz] \r\n","created":"2026-02-25T20:55:34.923+0000"},{"body":"CI run for trunk is ok (known failed/flaky tests only, not related to the change):\r\n * [^ci_summary_CASSANDRA-19661_trunk.htm]\r\n * [^results_details_CASSANDRA-19661_trunk.tar.xz]","created":"2026-02-26T07:36:47.533+0000"}],"conversations":[{"body":"I'm using llama-index and llama3 to train a model. I'm using a very simple code that reads some *.txt files from local and uploads them to Cassandra and then creates the index:\r\n\r\n \r\n{code:java}\r\n# Create the index from documents\r\nindex = VectorStoreIndex.from_documents(\r\n documents,\r\n service_context=vector_store.service_context,\r\n storage_context=storage_context,\r\n show_progress=True,\r\n ) {code}\r\nThis works well and I'm able to use a Chat app to get responses from the Cassandra data. however, right after, I cannot restart Cassandra. It'll break with the following error:\r\n\r\n \r\n{code:java}\r\nINFO [PerDiskMemtableFlushWriter_0:7] 2024-05-23 08:23:20,102 Flushing.java:179 - Completed flushing /data/cassandra/data/gpt/docs_20240523-10c8eaa018d811ef8dadf75182f3e2b4/da-6-bti-Data.db (124.236MiB) for commitlog position CommitLogPosition(segmentId=1716452305636, position=15336)\r\n[...]\r\nWARN [MemtableFlushWriter:1] 2024-05-23 08:28:29,575 MemtableIndexWriter.java:92 - [gpt.docs.idx_vector_docs] Aborting index memtable flush for /data/cassandra/data/gpt/docs-aea77a80184b11ef8dadf75182f3e2b4/da-3-bti...{code}\r\n{code:java}\r\njava.lang.IllegalStateException: null\r\n        at com.google.common.base.Preconditions.checkState(Preconditions.java:496)\r\n        at org.apache.cassandra.index.sai.disk.v1.vector.VectorPostings.computeRowIds(VectorPostings.java:76)\r\n        at org.apache.cassandra.index.sai.disk.v1.vector.OnHeapGraph.writeData(OnHeapGraph.java:313)\r\n        at org.apache.cassandra.index.sai.memory.VectorMemoryIndex.writeDirect(VectorMemoryIndex.java:272)\r\n        at org.apache.cassandra.index.sai.memory.MemtableIndex.writeDirect(MemtableIndex.java:110)\r\n        at org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.flushVectorIndex(MemtableIndexWriter.java:192)\r\n        at org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.complete(MemtableIndexWriter.java:117)\r\n        at org.apache.cassandra.index.sai.disk.StorageAttachedIndexWriter.complete(StorageAttachedIndexWriter.java:185)\r\n        at java.base/java.util.ArrayList.forEach(ArrayList.java:1541)\r\n        at java.base/java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1085)\r\n        at org.apache.cassandra.io.sstable.format.SSTableWriter.commit(SSTableWriter.java:289)\r\n        at org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit(ShardedMultiWriter.java:219)\r\n        at org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1323)\r\n        at org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1222)\r\n        at org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133)\r\n        at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)\r\n        at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)\r\n        at io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n        at java.base/java.lang.Thread.run(Thread.java:829) {code}\r\nThe table created by the script is as follows:\r\n\r\n \r\n{noformat}\r\nCREATE TABLE gpt.docs (\r\n partition_id text,\r\n row_id text,\r\n attributes_blob text,\r\n body_blob text,\r\n vector vector,\r\n metadata_s map,\r\n PRIMARY KEY (partition_id, row_id)\r\n) WITH CLUSTERING ORDER BY (row_id ASC)\r\n AND additional_write_policy = '99p'\r\n AND allow_auto_snapshot = true\r\n AND bloom_filter_fp_chance = 0.01\r\n AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\r\n AND cdc = false\r\n AND comment = ''\r\n AND compaction = {'class': 'org.apache.cassandra.db.compaction.UnifiedCompactionStrategy', 'scaling_parameters': 'T4', 'target_sstable_size': '1GiB'}\r\n AND compression = {'chunk_length_in_kb': '16', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}\r\n AND memtable = 'default'\r\n AND crc_check_chance = 1.0\r\n AND default_time_to_live = 0\r\n AND extensions = {}\r\n AND gc_grace_seconds = 864000\r\n AND incremental_backups = true\r\n AND max_index_interval = 2048\r\n AND memtable_flush_period_in_ms = 0\r\n AND min_index_interval = 128\r\n AND read_repair = 'BLOCKING'\r\n AND speculative_retry = '99p';\r\n\r\nCREATE CUSTOM INDEX eidx_metadata_s_docs ON gpt.docs (entries(metadata_s)) USING 'org.apache.cassandra.index.sai.StorageAttachedIndex';\r\n\r\nCREATE CUSTOM INDEX idx_vector_docs ON gpt.docs (vector) USING 'org.apache.cassandra.index.sai.StorageAttachedIndex';{noformat}\r\n\r\n\r\nThank you\r\n\r\n ","from":"reporter","subject":"Dynamically skip sharding L0 when SAI Vector index present"},{"body":"The schema alone is not enough to reproduce, but it looks like this should block 5.0-rc.","from":"developer"},{"body":"[~adelapena] [~jbellis] [~mike_tr_adamson] It looks like perhaps {{computeRowIds()}} is being called multiple times when it shouldn't via {{OnHeapGraph#writeData()}}? Is it possible for multiple keys in {{postingsMap}} to point to the same {{VectorPostings}} instance?","from":"developer"},{"body":"> It looks like perhaps {{computeRowIds()}} is being called multiple times when it shouldn't\r\n\r\nAgreed that it's not intended to be called multiple times.\r\n\r\n> Is it possible for multiple keys in {{postingsMap}} to point to the same {{VectorPostings}} instance?\r\n\r\nI don't think that's possible – the only mutation against `postingsMap` is done with a freshly instantiated `VectorPostings`.","from":"developer"},{"body":"A couple of observations.\r\n\r\nFirstly, the stacktrace provided is not consistent with the latest 5.0 codebase across a number of files. This means it is quite old. It would be good to get a reproduction on the latest 5.0.\r\n\r\nSecondly, it would be good to get a more logging around the error to see what was going on at the time. \r\n\r\nI agree with [~jbellis] in that it only appears possible to call VectorPostings.computeRowIds once. This makes me think that we are attempting to flush the memtable index more than once but I've no idea how that could be happening either.","from":"developer"},{"body":"I think I saw something like this very early in 5.0 development cycle but I can not reproduce that anymore.\r\n\r\nI wrote this trainer / prompt (1) which takes the sources or Cassandra, trains them and one can ask questions against Cassandra codebase. It uses cassio and langchain. I can not replicate the issue. After the training I can just restart it fine. \r\n\r\nIf you want to try that, you need to pay to OpenAI to train it and put the token into .env file where that script is like \"OPENAI_API_KEY=...\"\r\n\r\n[https://gist.github.com/smiklosovic/98b52502df8061ec1d8c3546cf489628]\r\n\r\nI am using the latest 5.0 branch.\r\n\r\nI am *not* saying this is not an issue. I am just saying that what I tried it on has not produced that error. ","from":"developer"},{"body":"[~sergrua] Can you reproduce the problem with the latest C* version?\r\n","from":"developer"},{"body":"I was hoping I could drive this forward by running through https://docs.llamaindex.ai/en/stable/examples/vector_stores/CassandraIndexDemo/ and getting a reproduction against beta1 or HEAD, but it does not reproduce. Adding all the text files from test/resources/tokenization/ does not help either.","from":"developer"},{"body":"I'll try to find some time this week to write the code to reproduce it. One thing I failed to mention and it may be important is that it's a 3-node cluster with `NetworkTopologyStrategy`","from":"developer"},{"body":"Thanks [~sergrua] , does this mean you were able to reproduce it also on HEAD and the issue was not fixed by chance since beta1?","from":"developer"},{"body":"No, I'm still on beta-1 [~e.dimitrova] ","from":"developer"},{"body":"Hi all. I haven't been able to reproduce it on the latest build. I'm attaching a cut down version of the code I used.","from":"developer"},{"body":"[~sergrua] what documents are you loading from the ./data directory?","from":"developer"},{"body":"[~brandon.williams] they are just txt files, nothing special.","from":"developer"},{"body":"[~sergrua] is it possible to upload them? I think there must be something endemic to the data that caused this, since your script alone is not enough to reproduce against beta1.","from":"developer"},{"body":"[~brandon.williams] Unfortunately it's not possible. It's not public data.","from":"developer"},{"body":"I see. Well if you are certain this only affected beta1 and no longer reproduces then I think we can close this ticket.","from":"developer"},{"body":"I'm reopening this ticket as I've reproduced exactly the same issue in 5.0.2.\r\n\r\nBasically, I also run a cluster of 3 nodes, I also use a vector index '', and my nodes are also unable to restart after data is written into that index.\r\n\r\nOne table is loaded with a large amount of vectors (millions), which seems to make Cassandra unstable, which eventually makes a node stop and unable to restart, presenting the same logs as above. Only solution is then to reset it.\r\n\r\nI haven't investigated Cassandra's code in detail, but in practice, if I don't write into the vector-indexed column, my cluster is suddenly very stable with no problem restarting.\r\n\r\nI've tried several implementations of Memtable and Compaction Strategies, which made no difference.\r\n\r\nI'm unfortunately also unable to provide data as it's non public, but I'll try to send some logs as soon as I can.","from":"developer"},{"body":"bq. my nodes are also unable to restart\r\n\r\nAre you receiving the same error as in the description?","from":"developer"},{"body":"I get the same error, yes.","from":"developer"},{"body":"Here are some logs from last time I tried using vectors.\r\n[^5.0.2_fail_memtableflush_vector_full.txt]\r\n\r\nThe error is essentially the same.\r\n{noformat}\r\nERROR [MemtableFlushWriter:464] 2025-01-13 13:48:27,449 MemtableIndexWriter.java:157 - [default.docs.vector_search_index] Error while flushing index null\r\njava.lang.IllegalStateException: null\r\n\tat com.google.common.base.Preconditions.checkState(Preconditions.java:496)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.VectorPostings.computeRowIds(VectorPostings.java:76)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.OnHeapGraph.writeData(OnHeapGraph.java:313)\r\n\tat org.apache.cassandra.index.sai.memory.VectorMemoryIndex.writeDirect(VectorMemoryIndex.java:271)\r\n\tat org.apache.cassandra.index.sai.memory.MemtableIndex.writeDirect(MemtableIndex.java:113)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.flushVectorIndex(MemtableIndexWriter.java:200)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.complete(MemtableIndexWriter.java:141)\r\n\tat org.apache.cassandra.index.sai.disk.StorageAttachedIndexWriter.complete(StorageAttachedIndexWriter.java:185)\r\n\tat java.base/java.util.ArrayList.forEach(ArrayList.java:1541)\r\n\tat java.base/java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1085)\r\n\tat org.apache.cassandra.io.sstable.format.SSTableWriter.commit(SSTableWriter.java:289)\r\n\tat org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit(ShardedMultiWriter.java:219)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1354)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1253)\r\n\tat org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)\r\n\tat io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n\tat java.base/java.lang.Thread.run(Thread.java:829)\r\n{noformat}\r\nI'm currently setting up another cluster to test this, because I can't have nodes get in that state on my main cluster.","from":"developer"},{"body":"I hadn't had a look at this for a few months, but last week we set up a new cluster and I decided to try again : after >10 millions vectors the problem reappeared, lost 2 nodes out of 8, unable to restart without a data wipe.\r\n\r\nI won't be able to test further as we're going to have to use another database for storing vectors, unfortunately.\r\nI'm surprised not more people have shared this issue. If any of you has successfully indexed 100 million - 1 billion vectors on Cassandra, I'd be very curious to know what settings you used.","from":"developer"},{"body":"I can confirm that this reproduces on 5.0.6.  I attempted to load 100 million embeddings and upon startup there's a high risk of the same exception:\r\n\r\n{code}\r\nERROR [MemtableFlushWriter:4] 2026-01-28 18:29:16,976 StorageAttachedIndexWriter.java:220 - [vector_bench.vectors_cbp.*] Failed to complete an index build\r\njava.lang.IllegalStateException: null\r\n\tat com.google.common.base.Preconditions.checkState(Preconditions.java:496)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.VectorPostings.computeRowIds(VectorPostings.java:76)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.OnHeapGraph.writeData(OnHeapGraph.java:315)\r\n\tat org.apache.cassandra.index.sai.memory.VectorMemoryIndex.writeDirect(VectorMemoryIndex.java:272)\r\n\tat org.apache.cassandra.index.sai.memory.MemtableIndex.writeDirect(MemtableIndex.java:113)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.flushVectorIndex(MemtableIndexWriter.java:212)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.complete(MemtableIndexWriter.java:143)\r\n\tat org.apache.cassandra.index.sai.disk.StorageAttachedIndexWriter.complete(StorageAttachedIndexWriter.java:212)\r\n\tat java.base/java.util.ArrayList.forEach(ArrayList.java:1511)\r\n\tat java.base/java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1092)\r\n\tat org.apache.cassandra.io.sstable.format.SSTableWriter.commit(SSTableWriter.java:295)\r\n\tat org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit(ShardedMultiWriter.java:219)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1354)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1253)\r\n\tat org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)\r\n\tat io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n\tat java.base/java.lang.Thread.run(Thread.java:833)\r\nWARN [MemtableFlushWriter:4] 2026-01-28 18:29:16,976 MemtableIndexWriter.java:117 - [vector_bench.vectors_cbp.vectors_cbp_embedding_idx] Aborting index memtable flush for /app/cassandra/data/vector_bench/vectors_cbp-3c1d51e0fbcc11f08c7baf9115bb8d33/oa-3gxg_1fbu_5az002afoxgv570894-big...\r\njava.lang.IllegalStateException: null\r\n\tat com.google.common.base.Preconditions.checkState(Preconditions.java:496)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.VectorPostings.computeRowIds(VectorPostings.java:76)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.OnHeapGraph.writeData(OnHeapGraph.java:315)\r\n\tat org.apache.cassandra.index.sai.memory.VectorMemoryIndex.writeDirect(VectorMemoryIndex.java:272)\r\n\tat org.apache.cassandra.index.sai.memory.MemtableIndex.writeDirect(MemtableIndex.java:113)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.flushVectorIndex(MemtableIndexWriter.java:212)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.complete(MemtableIndexWriter.java:143)\r\n\tat org.apache.cassandra.index.sai.disk.StorageAttachedIndexWriter.complete(StorageAttachedIndexWriter.java:212)\r\n\tat java.base/java.util.ArrayList.forEach(ArrayList.java:1511)\r\n\tat java.base/java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1092)\r\n\tat org.apache.cassandra.io.sstable.format.SSTableWriter.commit(SSTableWriter.java:295)\r\n\tat org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit(ShardedMultiWriter.java:219)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1354)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1253)\r\n\tat org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)\r\n\tat io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n\tat java.base/java.lang.Thread.run(Thread.java:833)\r\n {code}\r\n\r\nEdit: In my case, wiping commit logs allowed the node to start up successfully as the error was thrown during significant commit log replay.","from":"developer"},{"body":"[~andrewweaver] does it reproduce if you set memtable_flush_writers: 1 in conf/cassandra.yaml ?","from":"developer"},{"body":"I just confirmed that it reproduces the {{java.lang.IllegalStateException}} with {{memtable_flush_writers: 1}}.","from":"developer"},{"body":"If possible could you please share system.log + debug.log for a failed Cassandra startup since the beginning of the start and till the failure?","from":"developer"},{"body":"Sure thing: [^logs.tar.gz] \r\n\r\n{{2026-01-29 09:57:29,176}} Crashed\r\n{{2026-01-29 09:58:59,215}} Restarted \r\n{{2026-01-29 11:36:44,951}} Error while replaying commitlogs, at 14:06, it finally exits with an error.","from":"developer"},{"body":"Thank you for the logs.\r\nI think I have a theory about a root cause for this issue.\r\nInitially I thought that it could be a kind of concurrency issue when two different flusher threads are trying to write the same index but I've not found in the code how it could happen + based on the last comments the issue is reproduced with memtable_flush_writers=1 as well.\r\n\r\nThe second idea was about a cycle here when we invoke writer.commit for every Flushing.FlushRunnable\r\n{code:java}\r\nfor (SSTableMultiWriter writer : flushResults)\r\n{\r\n accumulate = writer.commit(accumulate);\r\n metric.flushSizeOnDisk.update(writer.getOnDiskBytesWritten());\r\n}\r\n{code}\r\nit could happen (and maybe it is a real issue too) if we have multiple data directories and a memtable if flushed to several SSTables to them.\r\nBut, based on the provided log it is not the case - we have a single data directory:\r\n{code:java}\r\nDEBUG [main] 2026-01-29 09:59:01,308 DiskBoundaryManager.java:57 - Updating boundaries from null to DiskBoundaries{directories=[DataDirectory{location=/app/cassandra/data}],\r\n{code}\r\nAfter looking to the stacktrace more precisely I've noticed that we use org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit\r\nwhich has multiple inner writers:\r\n{code:java}\r\n@Override\r\npublic Throwable commit(Throwable accumulate)\r\n{\r\n Throwable t = accumulate;\r\n for (SSTableWriter writer : writers)\r\n if (writer != null)\r\n t = writer.commit(t);\r\n return t;\r\n}\r\n{code}\r\nso, if we have more than 1 writer in our case here we will invoke the commit method several times for the same observers:\r\n{code:java}\r\nobservers.forEach(SSTableFlushObserver::complete);\r\n{code}\r\nthe observers are retrieved in the same way during org.apache.cassandra.io.sstable.format.SSTableWriter#SSTableWriter constructing:\r\n{code:java}\r\nSSTableFlushObserver observer = group.getFlushObserver(descriptor, lifecycleNewTracker, metadata.getLocal());\r\n{code}\r\nThe number of shards is calculated by UCS here: org.apache.cassandra.db.compaction.UnifiedCompactionStrategy#createSSTableMultiWriter\r\n\r\nand this logic is applicable to flushing writers too:\r\n{code:java}\r\ndouble flushDensity = cfs.metric.flushSizeOnDisk.get() * shardManager.shardSetCoverage() / shardManager.localSpaceCoverage();\r\nShardTracker boundaries = shardManager.boundaries(controller.getNumShards(flushDensity));\r\n{code}\r\nit is dynamic and depends on data distribution/density.\r\nWe have a debug log printed by controller.getNumShards, so we can find in the debug.log a proof for this idea:\r\n{code:java}\r\nINFO [NativePoolCleaner] 2026-01-29 11:34:52,296 ColumnFamilyStore.java:1052 - Enqueuing flush of vector_bench.vectors_cbp, Reason: MEMTABLE_LIMIT, Usage: 2.000GiB (20%) on-heap, 1.256GiB (13%) off-heap\r\nDEBUG [MemtableFlushWriter:5] 2026-01-29 11:34:53,071 Controller.java:348 - Shard count 4 for density 1.241 GiB, 1.8334467342069816 times target 693.273 MiB\r\n{code}\r\n=====\r\n\r\nAs a result, it looks like the issue is caused by combination of 2 features: UCS which can flush a memtable to multiple sharded SSTables and SAI vector index which does not support sharding (https://issues.apache.org/jira/browse/CASSANDRA-20752 - a related change but it looks like it does not cover Vector index scenario).\r\n\r\nSo, if the theory is correct then the possible WA can be switch compaction from UCS to another strategy for this table or to force UCS to use a single shard. Based on org.apache.cassandra.db.compaction.unified.Controller#getNumShards code we can get it if we set min_sstable_size to 0 and base_shard_count to 1 in the table settings or the same via Java system properties: -Dunified_compaction.min_sstable_size=0MiB -D\r\nunified_compaction.base_shard_count=1 (it is for all tables but does not require a schema change)\r\n[~maedhroz] [~blambov]  - what do you think, am I missing something?","from":"developer"},{"body":"Thank you for the quick analysis [~dnk].\r\n\r\nI'll try with 1 UCS shard and STCS if necessary.\r\n\r\nI've consistently observed a pattern where nodes get into a state where pending mutations grow at a steady rate. Unusual as my test is only doing up to about 200 writes/s per node and the rest of the nodes keep up. This seems to happen after:\r\n{code}\r\nDEBUG [MemtablePostFlush:1] 2026-01-30 19:03:27,808 HeapUtils.java:133 - Heap dump creation on uncaught exceptions is disabled.\r\nERROR [MemtablePostFlush:1] 2026-01-30 19:03:27,808 JVMStabilityInspector.java:70 - Exception in thread Thread[MemtablePostFlush:1,5,MemtablePostFlush]\r\njava.lang.NullPointerException: Cannot invoke \"java.lang.Boolean.booleanValue()\" because \"res\" is null\r\n at org.apache.cassandra.utils.memory.MemtableCleanerThread$Clean.apply(MemtableCleanerThread.java:97)\r\n{code}","from":"developer"},{"body":"I verified that it reproduces with {{base_shard_count}} of 1.\r\n\r\n{code}\r\ncompaction = {'base_shard_count': '1', 'class': 'org.apache.cassandra.db.compaction.UnifiedCompactionStrategy', 'max_sstables_to_compact': '64', 'min_sstable_size': '100MiB', 'scaling_parameters': 'T4', 'sstable_growth': '0.3333', 'target_sstable_size': '1GiB'}\r\n{code}\r\n\r\nTrying STCS.","from":"developer"},{"body":"For UCS it should be such combination (I forgot to mentioned sstable_growth):\r\n{code:java}\r\n'min_sstable_size' : '0MiB', 'base_shard_count': '1', 'sstable_growth' : '1'{code}\r\na message like this should be printed in the log before a flush for the table:\r\n{code:java}\r\nShard count 1 for density {some number} in fixed shards mode {code}\r\n====\r\nRegarding\r\n{quote}I've consistently observed a pattern where nodes get into a state where pending mutations grow at a steady rate.\r\n{quote}\r\nIn the shared logs I see a combination of 3 subsequent errors:\r\n{code:java}\r\nERROR [MemtablePostFlush:1] 2026-01-29 11:38:04,159 JVMStabilityInspector.java:70 - Exception in thread Thread[MemtablePostFlush:1,5,MemtablePostFlush]\r\njava.lang.IllegalStateException: null\r\n\tat com.google.common.base.Preconditions.checkState(Preconditions.java:496)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.VectorPostings.computeRowIds(VectorPostings.java:76)\r\n\tat org.apache.cassandra.index.sai.disk.v1.vector.OnHeapGraph.writeData(OnHeapGraph.java:315)\r\n\tat org.apache.cassandra.index.sai.memory.VectorMemoryIndex.writeDirect(VectorMemoryIndex.java:272)\r\n\tat org.apache.cassandra.index.sai.memory.MemtableIndex.writeDirect(MemtableIndex.java:113)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.flushVectorIndex(MemtableIndexWriter.java:212)\r\n\tat org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.complete(MemtableIndexWriter.java:143)\r\n\tat org.apache.cassandra.index.sai.disk.StorageAttachedIndexWriter.complete(StorageAttachedIndexWriter.java:212)\r\n\tat java.base/java.util.ArrayList.forEach(ArrayList.java:1511)\r\n\tat java.base/java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1092)\r\n\tat org.apache.cassandra.io.sstable.format.SSTableWriter.commit(SSTableWriter.java:295)\r\n\tat org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit(ShardedMultiWriter.java:219)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1354)\r\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1253)\r\n\tat org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)\r\n\tat io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n\tat java.base/java.lang.Thread.run(Thread.java:833)\r\n\tSuppressed: java.lang.IllegalStateException: null\r\n\t\t... 19 common frames omitted\r\n\tSuppressed: java.lang.IllegalStateException: null\r\n\t\t... 19 common frames omitted\r\n\r\n{code}\r\nthen\r\n{code:java}\r\nERROR [MemtablePostFlush:1] 2026-01-29 11:38:04,159 JVMStabilityInspector.java:70 - Exception in thread Thread[MemtablePostFlush:1,5,MemtablePostFlush]\r\njava.lang.NullPointerException: Cannot invoke \"java.lang.Boolean.booleanValue()\" because \"res\" is null\r\n\tat org.apache.cassandra.utils.memory.MemtableCleanerThread$Clean.apply(MemtableCleanerThread.java:97)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList$CallbackBiConsumerListener.run(ListenerList.java:244)\r\n\tat org.apache.cassandra.concurrent.ImmediateExecutor.execute(ImmediateExecutor.java:140)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.safeExecute(ListenerList.java:166)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notifyListener(ListenerList.java:157)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList$CallbackBiConsumerListener.notifySelf(ListenerList.java:250)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.lambda$notifyExclusive$0(ListenerList.java:124)\r\n\tat org.apache.cassandra.utils.concurrent.IntrusiveStack.forEach(IntrusiveStack.java:195)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notifyExclusive(ListenerList.java:124)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notify(ListenerList.java:96)\r\n\tat org.apache.cassandra.utils.concurrent.AsyncFuture.trySet(AsyncFuture.java:104)\r\n\tat org.apache.cassandra.utils.concurrent.AbstractFuture.tryFailure(AbstractFuture.java:148)\r\n\tat org.apache.cassandra.utils.concurrent.AsyncPromise.tryFailure(AsyncPromise.java:139)\r\n\tat org.apache.cassandra.db.memtable.AbstractAllocatorMemtable.lambda$flushLargestMemtable$0(AbstractAllocatorMemtable.java:306)\r\n\tat org.apache.cassandra.concurrent.ImmediateExecutor.execute(ImmediateExecutor.java:140)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.safeExecute(ListenerList.java:166)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notifyListener(ListenerList.java:157)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList$RunnableWithExecutor.notifySelf(ListenerList.java:345)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.lambda$notifyExclusive$0(ListenerList.java:124)\r\n\tat org.apache.cassandra.utils.concurrent.IntrusiveStack.forEach(IntrusiveStack.java:195)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notifyExclusive(ListenerList.java:124)\r\n\tat org.apache.cassandra.utils.concurrent.ListenerList.notify(ListenerList.java:96)\r\n\tat org.apache.cassandra.utils.concurrent.AsyncFuture.trySet(AsyncFuture.java:104)\r\n\tat org.apache.cassandra.utils.concurrent.AbstractFuture.tryFailure(AbstractFuture.java:148)\r\n\tat org.apache.cassandra.concurrent.FutureTask.tryFailure(FutureTask.java:87)\r\n\tat org.apache.cassandra.concurrent.FutureTask.run(FutureTask.java:75)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)\r\n\tat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)\r\n\tat io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n\tat java.base/java.lang.Thread.run(Thread.java:833)\r\n{code}\r\nand finally there are also messages about a resource leak:\r\n{code:java}\r\nERROR [Reference-Reaper] 2026-01-29 11:41:57,667 Ref.java:243 - LEAK DETECTED: a reference (class org.apache.cassandra.utils.concurrent.WrappedSharedCloseable$Tidy@344029653:[[OffHeapBitSet]]) to class org.apache.cassandra.utils.concurrent.WrappedSharedCloseable$Tidy@344029653:[[OffHeapBitSet]] was not released before the reference was garbage collected\r\nERROR [Reference-Reaper] 2026-01-29 11:41:57,669 Ref.java:243 - LEAK DETECTED: a reference (class org.apache.cassandra.io.util.MmappedRegions$Tidier@434613886:/app/cassandra/data/vector_bench/vectors_cbp-3c1d51e0fbcc11f08c7baf9115bb8d33/oa-3gxh_0w67_0rv342lhihvcwtdh2y-big-Index.db) to class org.apache.cassandra.io.util.MmappedRegions$Tidier@434613886:/app/cassandra/data/vector_bench/vectors_cbp-3c1d51e0fbcc11f08c7baf9115bb8d33/oa-3gxh_0w67_0rv342lhihvcwtdh2y-big-Index.db was not released before the reference was garbage collected\r\n{code}\r\nDo you observe the pending mutations grow issue only in such combination or there are other patterns too?\r\n\r\nAlso a thread dump would help here to see what are the mutation threads waiting for.","from":"developer"},{"body":"Apologies about not being able to actively engage much here recently. (Trying to get some other work over the line.) The big thing I'm worried about is whether or not we can get into this state w/ UCS and vanilla SAI (say on an indexed text or numeric/scalar column). There is some work in progress to shard the Memtable indexes in CASSANDRA-18216, but I'm not sure if that is related here...","from":"developer"},{"body":"For vanilla SAI, it looks like there was a fix here... - CASSANDRA-20752","from":"developer"},{"body":"My impression was that CASSANDRA-20752 only touches compaction. In any case, the writer creation _should_ be creating a separate observer for each shard, i.e. a separate StorageAttachedIndexWriter/MemtableIndexWriter, and those should have commit() called independently on them...unless I'm missing something. ({{Index.Group#getFlushObserver()}} creates new {{SSTableFlushObserver}} instances.)","from":"developer"},{"body":"So to elaborate on my previous comment, it feels like as long as the {{MemtableIndexWriter}} instances for the shards don't reuse anything important, they should be able to deal with a shared writer (even though the Memtable index itself isn't actually shared). Non-Vector, for instance, doesn't share {{RowMapping}} instances. {{MemtableIndexWriter#flushVectorIndex()}} does call into {{writeDirect()}}, which ultimately hits {{OnHeapGraph#writeData()}}. This touches the shared {{postingsMap}} to get {{VectorPostings}} and calls {{computeRowIds()}} on them, which doesn't look like it should happen twice, as {{rowIds}} is only supposed to be populated once. (It's not memoized. It's illegal to call it twice.) Am I getting warmer?","from":"developer"},{"body":"Would it be appropriate to just make {{computeRowIds()}} idempotent?","from":"developer"},{"body":"Confirmed this message with shard count set to 1: {{Shard count 1 for density 19.241 KiB, 1.8790509741811726E-4 times min size 100.000 MiB}}\r\n\r\nI tried again with STCS and it completed writing all 10 million embeddings without issue.\r\n\r\n{quote}\r\nDo you observe the pending mutations grow issue only in such combination or there are other patterns too?\r\n{quote}\r\nIt seems to be this pattern consistently.\r\n\r\nI can try to reproduce again and get a thread dump.\r\n\r\n","from":"developer"},{"body":"I replicated the issue again and captured thread dumps. Note that I did have to undo some tuning that I did after switching to STCS to trigger it.\r\n{code:java}\r\nmemtable_flush_writers: 1 -> 2\r\nmemtable_offheap_space: default (~2GB) -> 1024MiB\r\nmemtable_cleanup_threshold: default (0.33) -> 0.15\r\n{code}\r\nMy goal was to get more frequent flushes.\r\n\r\n!screenshot-1.png|width=590,height=255!\r\n\r\nAttached are thread dumps taken every 5 minutes.\r\n[^10.103.220.89_thread_dump.tgz]","from":"developer"},{"body":"So, am I right that the issue is observed only for the runs when you have observed the IllegalStateException mentioned before? ","from":"developer"},{"body":"Here is a draft solution: https://github.com/apache/cassandra/pull/4605. It works by modifying UCS, while leaving the underlying vector design untouched.\r\n\r\nIn general, I see one key requirement: do not attempt to build a graph at flush time. It is expensive to build vector graphs, and we do not want to build up an excessive number of vectors in memory due to delays at flush time.\r\n\r\nTherefore, I see two options:\r\n\r\n1. Pre-shard the vector memtable index to align with flush shard bounds\r\n2. Prevent sharding at flush time. (This is the proposed solution in the draft PR. I haven't explicitly tested it yet because I want to get general feedback on the design before proceeding)\r\n\r\nFrom a vector perf perspective, we always want to build bigger graphs. The time complexity of graph search is `log( n )`. Many small graphs means we have `k * log( n )` performance and that significantly increases latency/reduces throughput. Given the benefit to read performance, I lean towards avoiding eager sharding.\r\n\r\nNote also that the `TrieMemoryIndex` differs from the `VectorMemoryIndex` in one significant way: one has a `synchronized` add method while the other does not. (Actually, I just noticed that the `VectorMemoryIndex` has the `synchronized` keyword, but doesn't need it! I just created this ticket as a follow up https://issues.apache.org/jira/browse/CASSANDRA-21160) I mention this because the primary benefit for pre-sharding the memtable index is to increase write throughput so that writes to different partitions can proceed independently. The jvector library allows for this, but the `TrieMemoryIndex` does not.","from":"developer"},{"body":"JUnit test to reproduce the issue: https://github.com/netudima/cassandra/commit/94406f37f2f6c7b955358001a637d956adf5d064\r\n[~mmarshall] could you add it to your PR?","from":"developer"},{"body":"An interim summary:\r\n * To fix the issue we have a patch from [~mmarshall] which enforces usage of 1 shard for UCS during a flushing if we have a vector index for a table. It should solve the reported issue.\r\n * An alternative could be a support of sharding for memory indexes (related story: CASSANDRA-18216) but it is much more complicated to implement + there are performance concerns regarding this way for vector indexes as it was mentioned before.\r\n * The fix does not cover a possible multiple data directories use case (when a flushing memtable is split  between multiple data directories) but it is much rare use case and I suppose it should be checked and fixed if needed separately, to not delay the current fix delivery.\r\n * A workaround for the issue is to switch to a non-UCS or to use UCS with the following configuration:\r\n{code:java}\r\ncompaction = {\r\n 'class': 'UnifiedCompactionStrategy',\r\n 'base_shard_count': '1',\r\n 'min_sstable_size' : '0MiB', \r\n 'sstable_growth' : '1'\r\n}\r\n{code}","from":"developer"},{"body":"{quote}replicated the issue again and captured thread dumps.\r\n{quote}\r\nRegarding pending Mutations issue mentioned before. Assuming that we speak about the case when the IllegalStateException is thrown.\r\n\r\nThe unexpected IllegalStateException breaks reclaiming of memory used by the flushed memtable, so we consume all memory available for memtables and block on awaiting for it.\r\n\r\n[https://github.com/apache/cassandra/blob/cassandra-5.0.6/src/java/org/apache/cassandra/db/ColumnFamilyStore.java#L1377] - the place where reclaim logic is register to trigger at the end of flush. We don't reach this method invocation because the IllegalStateException is thrown before on writer commit step:\r\n\r\n[https://github.com/apache/cassandra/blob/cassandra-5.0.6/src/java/org/apache/cassandra/db/ColumnFamilyStore.java#L1354] \r\n\r\nHere is an example of a mutation thread  trying to allocate some memory for memtables:\r\n\r\n{{\"MutationStage-1\" #131 daemon prio=5 os_prio=0 cpu=949990.56ms elapsed=26753.90s tid=0x00007267aea3f2d0 nid=0xcc734 waiting on condition  [0x0000725ce31bc000]}}{{{}   java.lang.Thread.State: WAITING (parking){}}}{{{}at jdk.internal.misc.Unsafe.park(java.base@17.0.3/Native Method){}}}{{{}at java.util.concurrent.locks.LockSupport.park(java.base@17.0.3/LockSupport.java:341){}}}{{{}at org.apache.cassandra.utils.concurrent.WaitQueue$Standard$AbstractSignal.await(WaitQueue.java:321){}}}{{{}at org.apache.cassandra.utils.concurrent.WaitQueue$Standard$AbstractSignal.await(WaitQueue.java:299){}}}{{{}at org.apache.cassandra.utils.concurrent.Awaitable$Defaults.awaitThrowUncheckedOnInterrupt(Awaitable.java:131){}}}{{{}at *{color:#de350b}org.apache.cassandra.utils.concurrent.Awaitable$AbstractAwaitable.awaitThrowUncheckedOnInterrupt(Awaitable.java:235){color}*{}}}{*}{color:#de350b}{{at org.apache.cassandra.utils.memory.MemtableAllocator$SubAllocator.allocate(MemtableAllocator.java:195)}}{color}{*}{{{}*{color:#de350b}at{color}* org.apache.cassandra.db.memtable.AbstractAllocatorMemtable.markExtraOnHeapUsed(AbstractAllocatorMemtable.java:196){}}}{{{}at org.apache.cassandra.index.sai.StorageAttachedIndex$UpdateIndexer.adjustMemtableSize(StorageAttachedIndex.java:998){}}}{{{}at org.apache.cassandra.index.sai.StorageAttachedIndex$UpdateIndexer.updateRow(StorageAttachedIndex.java:990){}}}{{{}at org.apache.cassandra.index.sai.StorageAttachedIndexGroup$1.updateRow(StorageAttachedIndexGroup.java:191){}}}{{{}at org.apache.cassandra.index.SecondaryIndexManager$WriteTimeTransaction.onUpdated(SecondaryIndexManager.java:1570){}}}{{{}at org.apache.cassandra.db.partitions.BTreePartitionUpdater.merge(BTreePartitionUpdater.java:139){}}}{{{}at org.apache.cassandra.db.partitions.BTreePartitionUpdater.merge(BTreePartitionUpdater.java:39){}}}{{{}at org.apache.cassandra.utils.btree.BTree.updateLeaves(BTree.java:430){}}}{{{}at org.apache.cassandra.utils.btree.BTree.update(BTree.java:372){}}}{{{}at org.apache.cassandra.db.partitions.BTreePartitionUpdater.makeMergedPartition(BTreePartitionUpdater.java:89){}}}{{{}at org.apache.cassandra.db.partitions.BTreePartitionUpdater.mergePartitions(BTreePartitionUpdater.java:71){}}}{{{}at org.apache.cassandra.db.memtable.TrieMemtable$MemtableShard$$Lambda$1179/0x00000008011e2d58.apply(Unknown Source){}}}{{{}at org.apache.cassandra.db.tries.InMemoryTrie.applyContent(InMemoryTrie.java:930){}}}{{{}at org.apache.cassandra.db.tries.InMemoryTrie.putRecursive(InMemoryTrie.java:906){}}}{{{}at org.apache.cassandra.db.tries.InMemoryTrie.putRecursive(InMemoryTrie.java:910){}}}{{{}at {}}}\r\n\r\n{{...}}\r\n\r\n{{{}org.apache.cassandra.db.tries.InMemoryTrie.putRecursive(InMemoryTrie.java:897){}}}{{{}at org.apache.cassandra.db.tries.InMemoryTrie.putSingleton(InMemoryTrie.java:878){}}}{{{}at org.apache.cassandra.db.memtable.TrieMemtable$MemtableShard.put(TrieMemtable.java:480){}}}{{{}at org.apache.cassandra.db.memtable.TrieMemtable.put(TrieMemtable.java:190){}}}{{{}at org.apache.cassandra.db.ColumnFamilyStore.apply(ColumnFamilyStore.java:1474){}}}{{{}at org.apache.cassandra.db.CassandraTableWriteHandler.write(CassandraTableWriteHandler.java:38){}}}{{{}at org.apache.cassandra.db.Keyspace.applyInternal(Keyspace.java:653){}}}{{{}at org.apache.cassandra.db.Keyspace.applyFuture(Keyspace.java:474){}}}{{{}at org.apache.cassandra.db.Mutation.applyFuture(Mutation.java:244){}}}{{{}at org.apache.cassandra.hints.Hint.applyFuture(Hint.java:109){}}}{{{}at org.apache.cassandra.hints.HintVerbHandler.doVerb(HintVerbHandler.java:116){}}}{{{}at org.apache.cassandra.net.InboundSink.lambda$new$0(InboundSink.java:78){}}}{{{}at org.apache.cassandra.net.InboundSink$$Lambda$783/0x00000008010bf960.accept(Unknown Source){}}}{{{}at org.apache.cassandra.net.InboundSink.accept(InboundSink.java:97){}}}{{{}at org.apache.cassandra.net.InboundSink.accept(InboundSink.java:45){}}}{{{}at org.apache.cassandra.net.InboundMessageHandler$ProcessMessage.run(InboundMessageHandler.java:430){}}}{{{}at org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133){}}}{{{}at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:143){}}}{{{}at io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30){}}}{{{}at java.lang.Thread.run(java.base@17.0.3/Thread.java:833){}}}\r\n\r\n \r\n\r\nThe pending mutation issue should disappear once we fix the {{{}IllegalStateException{}}}. We could try to improve the error handling to better process such unexpected exceptions, but there may be no way to handle this completely correctly — we might even have to stop processing at all or risk data loss (at least for indexes). So I’m not sure whether investing significant effort into this will really pay off.","from":"developer"},{"body":"+1 on the PR","from":"developer"},{"body":"CI run for 5.0 is ok (its known flaky TestBootstrap only):\r\n* https://pre-ci.cassandra.apache.org/job/cassandra-5.0/86/\r\n* [^ci_summary_CASSANDRA-19661.html.zip] \r\n* [^results_details_CASSANDRA-19661.tar.xz] \r\n","from":"developer"},{"body":"CI run for trunk is ok (known failed/flaky tests only, not related to the change):\r\n * [^ci_summary_CASSANDRA-19661_trunk.htm]\r\n * [^results_details_CASSANDRA-19661_trunk.tar.xz]","from":"developer"}],"created":"2024-05-24T16:27:31.000+0000","description":"I'm using llama-index and llama3 to train a model. I'm using a very simple code that reads some *.txt files from local and uploads them to Cassandra and then creates the index:\r\n\r\n \r\n{code:java}\r\n# Create the index from documents\r\nindex = VectorStoreIndex.from_documents(\r\n documents,\r\n service_context=vector_store.service_context,\r\n storage_context=storage_context,\r\n show_progress=True,\r\n ) {code}\r\nThis works well and I'm able to use a Chat app to get responses from the Cassandra data. however, right after, I cannot restart Cassandra. It'll break with the following error:\r\n\r\n \r\n{code:java}\r\nINFO [PerDiskMemtableFlushWriter_0:7] 2024-05-23 08:23:20,102 Flushing.java:179 - Completed flushing /data/cassandra/data/gpt/docs_20240523-10c8eaa018d811ef8dadf75182f3e2b4/da-6-bti-Data.db (124.236MiB) for commitlog position CommitLogPosition(segmentId=1716452305636, position=15336)\r\n[...]\r\nWARN [MemtableFlushWriter:1] 2024-05-23 08:28:29,575 MemtableIndexWriter.java:92 - [gpt.docs.idx_vector_docs] Aborting index memtable flush for /data/cassandra/data/gpt/docs-aea77a80184b11ef8dadf75182f3e2b4/da-3-bti...{code}\r\n{code:java}\r\njava.lang.IllegalStateException: null\r\n        at com.google.common.base.Preconditions.checkState(Preconditions.java:496)\r\n        at org.apache.cassandra.index.sai.disk.v1.vector.VectorPostings.computeRowIds(VectorPostings.java:76)\r\n        at org.apache.cassandra.index.sai.disk.v1.vector.OnHeapGraph.writeData(OnHeapGraph.java:313)\r\n        at org.apache.cassandra.index.sai.memory.VectorMemoryIndex.writeDirect(VectorMemoryIndex.java:272)\r\n        at org.apache.cassandra.index.sai.memory.MemtableIndex.writeDirect(MemtableIndex.java:110)\r\n        at org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.flushVectorIndex(MemtableIndexWriter.java:192)\r\n        at org.apache.cassandra.index.sai.disk.v1.MemtableIndexWriter.complete(MemtableIndexWriter.java:117)\r\n        at org.apache.cassandra.index.sai.disk.StorageAttachedIndexWriter.complete(StorageAttachedIndexWriter.java:185)\r\n        at java.base/java.util.ArrayList.forEach(ArrayList.java:1541)\r\n        at java.base/java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1085)\r\n        at org.apache.cassandra.io.sstable.format.SSTableWriter.commit(SSTableWriter.java:289)\r\n        at org.apache.cassandra.db.compaction.unified.ShardedMultiWriter.commit(ShardedMultiWriter.java:219)\r\n        at org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1323)\r\n        at org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1222)\r\n        at org.apache.cassandra.concurrent.ExecutionFailure$1.run(ExecutionFailure.java:133)\r\n        at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)\r\n        at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)\r\n        at io.netty.util.concurrent.FastThreadLocalRunnable.run(FastThreadLocalRunnable.java:30)\r\n        at java.base/java.lang.Thread.run(Thread.java:829) {code}\r\nThe table created by the script is as follows:\r\n\r\n \r\n{noformat}\r\nCREATE TABLE gpt.docs (\r\n partition_id text,\r\n row_id text,\r\n attributes_blob text,\r\n body_blob text,\r\n vector vector,\r\n metadata_s map,\r\n PRIMARY KEY (partition_id, row_id)\r\n) WITH CLUSTERING ORDER BY (row_id ASC)\r\n AND additional_write_policy = '99p'\r\n AND allow_auto_snapshot = true\r\n AND bloom_filter_fp_chance = 0.01\r\n AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\r\n AND cdc = false\r\n AND comment = ''\r\n AND compaction = {'class': 'org.apache.cassandra.db.compaction.UnifiedCompactionStrategy', 'scaling_parameters': 'T4', 'target_sstable_size': '1GiB'}\r\n AND compression = {'chunk_length_in_kb': '16', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}\r\n AND memtable = 'default'\r\n AND crc_check_chance = 1.0\r\n AND default_time_to_live = 0\r\n AND extensions = {}\r\n AND gc_grace_seconds = 864000\r\n AND incremental_backups = true\r\n AND max_index_interval = 2048\r\n AND memtable_flush_period_in_ms = 0\r\n AND min_index_interval = 128\r\n AND read_repair = 'BLOCKING'\r\n AND speculative_retry = '99p';\r\n\r\nCREATE CUSTOM INDEX eidx_metadata_s_docs ON gpt.docs (entries(metadata_s)) USING 'org.apache.cassandra.index.sai.StorageAttachedIndex';\r\n\r\nCREATE CUSTOM INDEX idx_vector_docs ON gpt.docs (vector) USING 'org.apache.cassandra.index.sai.StorageAttachedIndex';{noformat}\r\n\r\n\r\nThank you\r\n\r\n ","issue_id":"13580386","key":"CASSANDRA-19661","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2026-02-26T13:18:27.000+0000","role":"fixed_distractor","summary":"Dynamically skip sharding L0 when SAI Vector index present"} {"case_id":"13618651","cluster":"DISTRACTOR-CASSANDRA-20664","comments":[{"body":"I have faced this issue as well, my coworker has a patch for it, I will ask him to share it to review and merge ...","created":"2025-05-20T12:26:50.441+0000"},{"body":"Good to hear that, thank you [~dnk] ! Please keep me posted.","created":"2025-06-02T09:29:20.653+0000"},{"body":"Hi team,\r\nI've started work on this ticket. Local dev environment is ready, branch {{CASSANDRA-20664-dev}} created, and Ant build initiated. ","created":"2025-06-19T06:13:55.490+0000"},{"body":"After testing and validating the latest Cassandra run using `./bin/cassandra -f`, I can confirm that the previously encountered issue with infinite commitlog replay and stuck startup is no longer occurring.\r\n\r\nThe logs indicate that all components are functioning normally:\r\n- Commitlog segments are being read and flushed without error.\r\n- Memtable flushes and SSTable compactions are executing as expected.\r\n- Schema changes (CREATE KEYSPACE and CREATE TABLE) are being successfully applied and committed.\r\n- No unexpected loops, warnings, or stack traces are observed.\r\n\r\nHence, the issue appears to be resolved. I will continue to monitor further executions to ensure consistency, but as of now, Cassandra is running stable on the latest configuration.","created":"2025-06-20T13:42:58.535+0000"},{"body":"Thanks for the analysis. I guess you tested trunk? Was there a bigger change? Can we fix this for Cassandra 4.x as well please?","created":"2025-06-23T06:48:31.783+0000"},{"body":"I’ve completed the verification using a custom build from the latest code base with the applied fix:\r\n1.Cassandra starts cleanly without encountering the infinite commit log replay loop or startup issues.\r\n2.Verified connectivity via cqlsh (connected to 5.1-SNAPSHOT), executed schema creation and data insertion, and performed a clean restart.\r\n3.Logs indicate normal startup, with no errors or unexpected behaviour.\r\n\r\nConclusion: The issue appears to be fully resolved.\r\n\r\nI will now raise a pull request with the fix against the trunk branch and proceed to get it merged. Once done, I’ll prepare a back port for the 4.1 branch as well.","created":"2025-06-23T16:34:55.538+0000"},{"body":"Hi [~christophschnepf],\r\n\r\nI've opened a PR to fix the issue described in this ticket:\r\n\r\n🔗 PR: https://github.com/apache/cassandra/pull/4207\r\n\r\nSummary of the fix:\r\n- Handled replay errors in `CommitLogReadHandler.java` to prevent infinite loops on unreadable commit log entries.\r\n- Updated `CommitLogReader.java` to integrate improved error handling at a higher level.\r\n- Verified the fix by building Cassandra locally and confirming the issue no longer occurs.\r\n- Committed clean changes from a dedicated feature branch (`CASSANDRA-20664-fix`).\r\n- Ready to backport to 4.x branches after review and merge.\r\n\r\nKindly review the PR when convenient. Happy to address any feedback.\r\n\r\nThanks!\r\n","created":"2025-06-23T19:31:56.196+0000"},{"body":"Hello [~christophschnepf]\r\nThe development work for this ticket is now complete from my end. The issue of the endless loop during commit log replay has been resolved.\r\nAll relevant changes have been pushed, and PR #4207 is updated with the fix and corresponding tests.\r\n\r\nKindly review the PR at your convenience and let me know if any further updates are needed.\r\n\r\nThank you!","created":"2025-06-26T11:58:48.919+0000"},{"body":"Thank you for your work on this topic and the fast fix. Unfortunately I do not know the Cassandra code well enough for a meaningful review. However, I see that there are already comments and questions on the PR.","created":"2025-06-30T12:40:14.392+0000"},{"body":"[~christophschnepf]\r\nThank you for taking the time to check the PR. I understand if you're not deeply familiar with the Cassandra code, I appreciate your note.\r\nI've seen the comments and questions on the PR and will review and address them accordingly. Please let me know if there's anything specific you'd like me to clarify or update.","created":"2025-06-30T14:48:33.503+0000"},{"body":"Implemented the fix for the issue described in this ticket, Verified the fix locally by building Cassandra and performing relevant tests , everything is working as expected,Please review the patch and share feedback. I’m open to suggestions or changes.","created":"2025-06-30T15:14:26.740+0000"},{"body":"The PR has been successfully merged by [~ifesdjeen]n. Marking this as Resolved from my side. Please let me know if anything else is needed.","created":"2025-07-01T09:14:30.276+0000"},{"body":"To me this looks like the additional PR (https://github.com/apache/cassandra/pull/4170) was deleted, and not merged.\r\n\r\nAnd I guess we would need https://github.com/apache/cassandra/pull/4207? \r\nPlus: would this be backported to 4.1 then as well?","created":"2025-07-01T09:27:04.015+0000"},{"body":"You're right ,PR #4170 was closed and not merged.\r\nI have created a new PR with the updated fix: https://github.com/apache/cassandra/pull/4207?,\r\nOnce this is merged into trunk, I’ll be happy to prepare a clean backport for the 4.1 branch as well.","created":"2025-07-01T09:41:00.523+0000"},{"body":"MR for 4.1: https://github.com/apache/cassandra/pull/4276","created":"2025-07-28T11:13:23.845+0000"},{"body":"Test results for 4.1: https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch-before-5/2690/testReport/\r\nthere is only one failure: dtest-upgrade.upgrade_tests.cql_tests.TestCQLNodes3RF3_Upgrade_indev_4_0_x_To_indev_4_1_x.test_multi_list_set (from dtest-upgrade_jdk1_python_cython_x86_64)\r\n{code}\r\nfailed on teardown with \"Unexpected error found in node logs (see stdout for full details). Errors: [[node1] 'WARN [epollEventLoopGroup-5-3] 2025-07-30 12:17:26,748 PreV5Handlers.java:261 - Unknown exception in client networking\\nio.netty.channel.unix.Errors$NativeIoException: writevAddresses(..) failed: Broken pipe', [node1] 'WARN [epollEventLoopGroup-5-3] 2025-07-30 12:17:26,751 PreV5Handlers.java:261 - Unknown exception in client networking\\nio.netty.channel.unix.Errors$NativeIoException: writevAddresses(..) failed: Broken pipe']\"\r\n{code}\r\nwhich looks like env issue and not related to the changes.","created":"2025-07-30T17:23:58.935+0000"},{"body":"MRs:\r\n * 4.1: [https://github.com/apache/cassandra/pull/4276]\r\n * 5.0: [https://github.com/apache/cassandra/pull/4284]\r\n * trunk: [https://github.com/apache/cassandra/pull/4283] \r\n\r\nCI results for 5.0:\r\n * [^5.0_ci_summary.htm] [^5.0_results_details.tar.xz]\r\n * 1 flaky test: transport.AuthMessageSizeLimitTest\r\n\r\nCI results for trunk: \r\n * [^trunk_ci_summary.htm] [^trunk_results_details.tar.xz]\r\n * 4 failures, all of the were observed before and not related to the changed logic\r\nTests / test jdk11 16/16 / org.apache.cassandra.db.virtual.AccordDebugKeyspaceTest.patchJournalVestigialTest-_jdk11_x86_64\r\nTests / test-latest jdk11 16/16 / org.apache.cassandra.db.virtual.AccordDebugKeyspaceTest.patchJournalVestigialTest-latest_jdk11_x86_64\r\nTests / jvm-dtest jdk11 3/12 / org.apache.cassandra.distributed.test.SSTableLoaderEncryptionOptionsTest.bulkLoaderSuccessfullyStreamsOverSslWithDeprecatedSslStoragePort-_jdk11_x86_64\r\nTests / simulator-dtest jdk11 / org.apache.cassandra.simulator.test.AccordHarrySimulationTest.test-_jdk11_x86_64","created":"2025-07-30T22:18:30.460+0000"},{"body":"so, tests are good for commit, [~smiklosovic] could you please take a look at:\r\n * 5.0: [https://github.com/apache/cassandra/pull/4284]\r\n * trunk: [https://github.com/apache/cassandra/pull/4283] \r\n\r\nthe logic there is the same, the only difference compared to 4.1 is the way how commitlog.ignorereplayerrorssystem property is configured via WithProperties","created":"2025-07-30T22:46:56.855+0000"}],"conversations":[{"body":"Hi,\r\nWe're using Cassandra 4.1.8 and specify the option {_}-Dcassandra.commitlog.ignorereplayerrors=true{_}, however we see an endless loop on starting Cassandra when there are corrupt commit log files found.\r\n\r\nThe stacktrace which is printed over and over again is: \r\n{code:java}\r\nINFO  [main] 2025-05-19 19:25:22,658 UTC CommitLogReader.java:257 - Finished reading /data/cassandra/commitlog/CommitLog-7-1745459535901.log\r\nINFO  [main] 2025-05-19 19:25:23,614 UTC CommitLogReader.java:257 - Finished reading /data/cassandra/commitlog/CommitLog-7-1745459535902.log\r\nERROR [main] 2025-05-19 19:25:24,572 UTC CommitLogReplayer.java:501 - Ignoring commit log replay error\r\norg.apache.cassandra.db.commitlog.CommitLogReadHandler$CommitLogReadException: Mutation checksum failure at 60807439 in Next section at 60745241 in CommitLog-7-1745459535903.log\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readSection(CommitLogReader.java:387)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:244)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:147)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReplayer.replayFiles(CommitLogReplayer.java:200)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverFiles(CommitLog.java:223)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverSegmentsOnDisk(CommitLog.java:204)\r\n    at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:353)\r\n    at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:744)\r\n    at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:878)\r\nERROR [main] 2025-05-19 19:25:24,572 UTC CommitLogReplayer.java:501 - Ignoring commit log replay error\r\norg.apache.cassandra.db.commitlog.CommitLogReadHandler$CommitLogReadException: Mutation size checksum failure at 60838538 in Next section at 60745241 in CommitLog-7-1745459535903.log\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readSection(CommitLogReader.java:356)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:244)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:147)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReplayer.replayFiles(CommitLogReplayer.java:200)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverFiles(CommitLog.java:223)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverSegmentsOnDisk(CommitLog.java:204)\r\n    at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:353)\r\n    at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:744)\r\n    at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:878)\r\nERROR [main] 2025-05-19 19:25:24,573 UTC CommitLogReplayer.java:501 - Ignoring commit log replay error\r\norg.apache.cassandra.db.commitlog.CommitLogReadHandler$CommitLogReadException: Encountered bad header at position 60865611 of commit log /data/cassandra/commitlog/CommitLog-7-1745459535903.log, with invalid CRC. The end of segment marker should be zero.\r\n    at org.apache.cassandra.db.commitlog.CommitLogSegmentReader$SegmentIterator.computeNext(CommitLogSegmentReader.java:127)\r\n    at org.apache.cassandra.db.commitlog.CommitLogSegmentReader$SegmentIterator.computeNext(CommitLogSegmentReader.java:98)\r\n    at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:141)\r\n    at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:136)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:233)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:147)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReplayer.replayFiles(CommitLogReplayer.java:200)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverFiles(CommitLog.java:223)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverSegmentsOnDisk(CommitLog.java:204)\r\n    at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:353)\r\n    at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:744)\r\n    at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:878)\r\nERROR [main] 2025-05-19 19:25:24,573 UTC CommitLogReplayer.java:501 - Ignoring commit log replay error\r\norg.apache.cassandra.db.commitlog.CommitLogReadHandler$CommitLogReadException: Encountered bad header at position 60865611 of commit log /data/cassandra/commitlog/CommitLog-7-1745459535903.log, with invalid CRC. The end of segment marker should be zero.\r\n    at org.apache.cassandra.db.commitlog.CommitLogSegmentReader$SegmentIterator.computeNext(CommitLogSegmentReader.java:127)\r\n    at org.apache.cassandra.db.commitlog.CommitLogSegmentReader$SegmentIterator.computeNext(CommitLogSegmentReader.java:98)\r\n    at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:141)\r\n    at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:136)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:233)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:147)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReplayer.replayFiles(CommitLogReplayer.java:200)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverFiles(CommitLog.java:223)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverSegmentsOnDisk(CommitLog.java:204)\r\n    at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:353)\r\n    at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:744)\r\n    at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:878)\r\nERROR [main] 2025-05-19 19:25:24,573 UTC CommitLogReplayer.java:501 - Ignoring commit log replay error {code}\r\nThis prevents the Cassandra startup on this node and it writes 50 MB to the _system.log_ in about 2 seconds.","from":"reporter","subject":"Endless loop on reading commitlogs when it should ignore replay errors"},{"body":"I have faced this issue as well, my coworker has a patch for it, I will ask him to share it to review and merge ...","from":"developer"},{"body":"Good to hear that, thank you [~dnk] ! Please keep me posted.","from":"developer"},{"body":"Hi team,\r\nI've started work on this ticket. Local dev environment is ready, branch {{CASSANDRA-20664-dev}} created, and Ant build initiated. ","from":"developer"},{"body":"After testing and validating the latest Cassandra run using `./bin/cassandra -f`, I can confirm that the previously encountered issue with infinite commitlog replay and stuck startup is no longer occurring.\r\n\r\nThe logs indicate that all components are functioning normally:\r\n- Commitlog segments are being read and flushed without error.\r\n- Memtable flushes and SSTable compactions are executing as expected.\r\n- Schema changes (CREATE KEYSPACE and CREATE TABLE) are being successfully applied and committed.\r\n- No unexpected loops, warnings, or stack traces are observed.\r\n\r\nHence, the issue appears to be resolved. I will continue to monitor further executions to ensure consistency, but as of now, Cassandra is running stable on the latest configuration.","from":"developer"},{"body":"Thanks for the analysis. I guess you tested trunk? Was there a bigger change? Can we fix this for Cassandra 4.x as well please?","from":"developer"},{"body":"I’ve completed the verification using a custom build from the latest code base with the applied fix:\r\n1.Cassandra starts cleanly without encountering the infinite commit log replay loop or startup issues.\r\n2.Verified connectivity via cqlsh (connected to 5.1-SNAPSHOT), executed schema creation and data insertion, and performed a clean restart.\r\n3.Logs indicate normal startup, with no errors or unexpected behaviour.\r\n\r\nConclusion: The issue appears to be fully resolved.\r\n\r\nI will now raise a pull request with the fix against the trunk branch and proceed to get it merged. Once done, I’ll prepare a back port for the 4.1 branch as well.","from":"developer"},{"body":"Hi [~christophschnepf],\r\n\r\nI've opened a PR to fix the issue described in this ticket:\r\n\r\n🔗 PR: https://github.com/apache/cassandra/pull/4207\r\n\r\nSummary of the fix:\r\n- Handled replay errors in `CommitLogReadHandler.java` to prevent infinite loops on unreadable commit log entries.\r\n- Updated `CommitLogReader.java` to integrate improved error handling at a higher level.\r\n- Verified the fix by building Cassandra locally and confirming the issue no longer occurs.\r\n- Committed clean changes from a dedicated feature branch (`CASSANDRA-20664-fix`).\r\n- Ready to backport to 4.x branches after review and merge.\r\n\r\nKindly review the PR when convenient. Happy to address any feedback.\r\n\r\nThanks!\r\n","from":"developer"},{"body":"Hello [~christophschnepf]\r\nThe development work for this ticket is now complete from my end. The issue of the endless loop during commit log replay has been resolved.\r\nAll relevant changes have been pushed, and PR #4207 is updated with the fix and corresponding tests.\r\n\r\nKindly review the PR at your convenience and let me know if any further updates are needed.\r\n\r\nThank you!","from":"developer"},{"body":"Thank you for your work on this topic and the fast fix. Unfortunately I do not know the Cassandra code well enough for a meaningful review. However, I see that there are already comments and questions on the PR.","from":"developer"},{"body":"[~christophschnepf]\r\nThank you for taking the time to check the PR. I understand if you're not deeply familiar with the Cassandra code, I appreciate your note.\r\nI've seen the comments and questions on the PR and will review and address them accordingly. Please let me know if there's anything specific you'd like me to clarify or update.","from":"developer"},{"body":"Implemented the fix for the issue described in this ticket, Verified the fix locally by building Cassandra and performing relevant tests , everything is working as expected,Please review the patch and share feedback. I’m open to suggestions or changes.","from":"developer"},{"body":"The PR has been successfully merged by [~ifesdjeen]n. Marking this as Resolved from my side. Please let me know if anything else is needed.","from":"developer"},{"body":"To me this looks like the additional PR (https://github.com/apache/cassandra/pull/4170) was deleted, and not merged.\r\n\r\nAnd I guess we would need https://github.com/apache/cassandra/pull/4207? \r\nPlus: would this be backported to 4.1 then as well?","from":"developer"},{"body":"You're right ,PR #4170 was closed and not merged.\r\nI have created a new PR with the updated fix: https://github.com/apache/cassandra/pull/4207?,\r\nOnce this is merged into trunk, I’ll be happy to prepare a clean backport for the 4.1 branch as well.","from":"developer"},{"body":"MR for 4.1: https://github.com/apache/cassandra/pull/4276","from":"developer"},{"body":"Test results for 4.1: https://ci-cassandra.apache.org/view/patches/job/Cassandra-devbranch-before-5/2690/testReport/\r\nthere is only one failure: dtest-upgrade.upgrade_tests.cql_tests.TestCQLNodes3RF3_Upgrade_indev_4_0_x_To_indev_4_1_x.test_multi_list_set (from dtest-upgrade_jdk1_python_cython_x86_64)\r\n{code}\r\nfailed on teardown with \"Unexpected error found in node logs (see stdout for full details). Errors: [[node1] 'WARN [epollEventLoopGroup-5-3] 2025-07-30 12:17:26,748 PreV5Handlers.java:261 - Unknown exception in client networking\\nio.netty.channel.unix.Errors$NativeIoException: writevAddresses(..) failed: Broken pipe', [node1] 'WARN [epollEventLoopGroup-5-3] 2025-07-30 12:17:26,751 PreV5Handlers.java:261 - Unknown exception in client networking\\nio.netty.channel.unix.Errors$NativeIoException: writevAddresses(..) failed: Broken pipe']\"\r\n{code}\r\nwhich looks like env issue and not related to the changes.","from":"developer"},{"body":"MRs:\r\n * 4.1: [https://github.com/apache/cassandra/pull/4276]\r\n * 5.0: [https://github.com/apache/cassandra/pull/4284]\r\n * trunk: [https://github.com/apache/cassandra/pull/4283] \r\n\r\nCI results for 5.0:\r\n * [^5.0_ci_summary.htm] [^5.0_results_details.tar.xz]\r\n * 1 flaky test: transport.AuthMessageSizeLimitTest\r\n\r\nCI results for trunk: \r\n * [^trunk_ci_summary.htm] [^trunk_results_details.tar.xz]\r\n * 4 failures, all of the were observed before and not related to the changed logic\r\nTests / test jdk11 16/16 / org.apache.cassandra.db.virtual.AccordDebugKeyspaceTest.patchJournalVestigialTest-_jdk11_x86_64\r\nTests / test-latest jdk11 16/16 / org.apache.cassandra.db.virtual.AccordDebugKeyspaceTest.patchJournalVestigialTest-latest_jdk11_x86_64\r\nTests / jvm-dtest jdk11 3/12 / org.apache.cassandra.distributed.test.SSTableLoaderEncryptionOptionsTest.bulkLoaderSuccessfullyStreamsOverSslWithDeprecatedSslStoragePort-_jdk11_x86_64\r\nTests / simulator-dtest jdk11 / org.apache.cassandra.simulator.test.AccordHarrySimulationTest.test-_jdk11_x86_64","from":"developer"},{"body":"so, tests are good for commit, [~smiklosovic] could you please take a look at:\r\n * 5.0: [https://github.com/apache/cassandra/pull/4284]\r\n * trunk: [https://github.com/apache/cassandra/pull/4283] \r\n\r\nthe logic there is the same, the only difference compared to 4.1 is the way how commitlog.ignorereplayerrorssystem property is configured via WithProperties","from":"developer"}],"created":"2025-05-20T12:18:42.000+0000","description":"Hi,\r\nWe're using Cassandra 4.1.8 and specify the option {_}-Dcassandra.commitlog.ignorereplayerrors=true{_}, however we see an endless loop on starting Cassandra when there are corrupt commit log files found.\r\n\r\nThe stacktrace which is printed over and over again is: \r\n{code:java}\r\nINFO  [main] 2025-05-19 19:25:22,658 UTC CommitLogReader.java:257 - Finished reading /data/cassandra/commitlog/CommitLog-7-1745459535901.log\r\nINFO  [main] 2025-05-19 19:25:23,614 UTC CommitLogReader.java:257 - Finished reading /data/cassandra/commitlog/CommitLog-7-1745459535902.log\r\nERROR [main] 2025-05-19 19:25:24,572 UTC CommitLogReplayer.java:501 - Ignoring commit log replay error\r\norg.apache.cassandra.db.commitlog.CommitLogReadHandler$CommitLogReadException: Mutation checksum failure at 60807439 in Next section at 60745241 in CommitLog-7-1745459535903.log\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readSection(CommitLogReader.java:387)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:244)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:147)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReplayer.replayFiles(CommitLogReplayer.java:200)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverFiles(CommitLog.java:223)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverSegmentsOnDisk(CommitLog.java:204)\r\n    at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:353)\r\n    at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:744)\r\n    at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:878)\r\nERROR [main] 2025-05-19 19:25:24,572 UTC CommitLogReplayer.java:501 - Ignoring commit log replay error\r\norg.apache.cassandra.db.commitlog.CommitLogReadHandler$CommitLogReadException: Mutation size checksum failure at 60838538 in Next section at 60745241 in CommitLog-7-1745459535903.log\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readSection(CommitLogReader.java:356)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:244)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:147)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReplayer.replayFiles(CommitLogReplayer.java:200)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverFiles(CommitLog.java:223)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverSegmentsOnDisk(CommitLog.java:204)\r\n    at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:353)\r\n    at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:744)\r\n    at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:878)\r\nERROR [main] 2025-05-19 19:25:24,573 UTC CommitLogReplayer.java:501 - Ignoring commit log replay error\r\norg.apache.cassandra.db.commitlog.CommitLogReadHandler$CommitLogReadException: Encountered bad header at position 60865611 of commit log /data/cassandra/commitlog/CommitLog-7-1745459535903.log, with invalid CRC. The end of segment marker should be zero.\r\n    at org.apache.cassandra.db.commitlog.CommitLogSegmentReader$SegmentIterator.computeNext(CommitLogSegmentReader.java:127)\r\n    at org.apache.cassandra.db.commitlog.CommitLogSegmentReader$SegmentIterator.computeNext(CommitLogSegmentReader.java:98)\r\n    at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:141)\r\n    at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:136)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:233)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:147)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReplayer.replayFiles(CommitLogReplayer.java:200)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverFiles(CommitLog.java:223)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverSegmentsOnDisk(CommitLog.java:204)\r\n    at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:353)\r\n    at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:744)\r\n    at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:878)\r\nERROR [main] 2025-05-19 19:25:24,573 UTC CommitLogReplayer.java:501 - Ignoring commit log replay error\r\norg.apache.cassandra.db.commitlog.CommitLogReadHandler$CommitLogReadException: Encountered bad header at position 60865611 of commit log /data/cassandra/commitlog/CommitLog-7-1745459535903.log, with invalid CRC. The end of segment marker should be zero.\r\n    at org.apache.cassandra.db.commitlog.CommitLogSegmentReader$SegmentIterator.computeNext(CommitLogSegmentReader.java:127)\r\n    at org.apache.cassandra.db.commitlog.CommitLogSegmentReader$SegmentIterator.computeNext(CommitLogSegmentReader.java:98)\r\n    at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:141)\r\n    at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:136)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:233)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReader.readCommitLogSegment(CommitLogReader.java:147)\r\n    at org.apache.cassandra.db.commitlog.CommitLogReplayer.replayFiles(CommitLogReplayer.java:200)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverFiles(CommitLog.java:223)\r\n    at org.apache.cassandra.db.commitlog.CommitLog.recoverSegmentsOnDisk(CommitLog.java:204)\r\n    at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:353)\r\n    at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:744)\r\n    at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:878)\r\nERROR [main] 2025-05-19 19:25:24,573 UTC CommitLogReplayer.java:501 - Ignoring commit log replay error {code}\r\nThis prevents the Cassandra startup on this node and it writes 50 MB to the _system.log_ in about 2 seconds.","issue_id":"13618651","key":"CASSANDRA-20664","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2025-07-31T09:31:01.000+0000","role":"fixed_distractor","summary":"Endless loop on reading commitlogs when it should ignore replay errors"} {"case_id":"12497286","cluster":"DISTRACTOR-CASSANDRA-2088","comments":[{"body":"Regarding repair: http://www.mail-archive.com/user@cassandra.apache.org/msg09259.html\nAnd compaction: CASSANDRA-2084","created":"2011-02-01T05:53:20.376+0000"},{"body":"I'm keen to try this ticket (to learn more about compaction and repair) if it's not already been worked on. Also if it's ok for me to take a couple of days while I dig into this.\n\nFor compaction I'm looking in\n- CompactionManager.doCompaction where it creates a new SSTableWriter via cfs.createCompactionWriter() \n- CompactionManager.doCleanupCompaction() also uses an SSTableWriter\n\nAre the sorts of failures we're considering for compaction ones that come from the CompactionIterator or SSTableScanner ?\n\nFor repair I'm looking in:\n- IncomingStreamReader appears to clean up the temporary pending file in some error situations. Do we have any more info on the sorts of failures here? e.g. If there is an IOException sending the re-stream message, or a non checked exception it will fail to cleaup the file. \n- I'm looking into what happens in StreamInSession.finished() closeIfFinished()\n- Are we considering failures during the streaming or when processing the data after the stream has finished?\n\nAny guidance welcome. ","created":"2011-03-07T05:23:40.889+0000"},{"body":"bq. Are the sorts of failures we're considering for compaction ones that come from the CompactionIterator or SSTableScanner ?\n\nBoth. Also I suppose it's possible for the writer to error out from lack of disk space since it only checks at the beginning for space and doesn't \"reserve\" it vs flushes.\n\nbq. Are we considering failures during the streaming or when processing the data after the stream has finished?\n\nThe former is much more common (I've never seen the latter reported), so I'd start with that.","created":"2011-03-07T16:59:43.931+0000"},{"body":"patch 0001 tracks failures during AES streaming, files for failed Stream sessions are cleaned up and repair is allowed to continue. Failed files are logged at the StreamSession, TreeRequest, and RepairSession level. \n\npatch 0002 handle exceptions when doing a (normal) compaction and deletes the temp SSTable. The SSTableWriter components are closed before deletion so that windows will delete correctly. ","created":"2011-04-11T05:50:45.431+0000"},{"body":"I think there is a few different things here and I think we should separate them somehow.\n\nFixing the fact that streaming leave tmp files around when it fails is a 2 lines fix and I think this is simple enough that it could go to 0.7. I'm attaching a patch against 0.7. It's extracted from Aaron first patch, although rebased on 0.7 (and fix a bug).\n\nMaking repair aware that there has been some failures is actually more complicated so that should go in 0.8.1 or something (and should go to CASSANDRA-2433 or another ticket that describe the problem better). ","created":"2011-04-12T17:45:13.702+0000"},{"body":"bq. I'm attaching a patch against 0.7\n\nIs that 0001-Better-detect-failures-from-the-other-side-in-Incomi.patch? I don't see the connection to .tmp files. (Also: have you verified that the channel will actually infinite-loop returning 0? Kind of odd behavior, although I guess it's technically within-spec.)","created":"2011-04-12T17:51:42.606+0000"},{"body":"bq. Is that 0001-Better-detect-failures-from-the-other-side-in-Incomi.patch? I don't see the connection to .tmp files. (Also: have you verified that the channel will actually infinite-loop returning 0? Kind of odd behavior, although I guess it's technically within-spec.)\n\nYes. IncomingStreamReader does clean the tmp file when there is an expection (there's an enclosing 'try catch'). The problem is that no exception is raised if the other side of the connection dies. What will happen then is the read will infinitely read 0 bytes. So this actually avoid the infinite loop returning 0 (and so I think answered your second question, so it wasn't very clear).\n\nNote that without this patch, there is an infinite loop that will hold a socket open forever (and consume cpu, though very few probably in that case). So this is not just merely a fix of deleting the tmp files. But it does as a consequence of correctly raising an exception when should be.","created":"2011-04-12T18:21:40.500+0000"},{"body":"+1, and can you move some of that explanation inline as a comment?","created":"2011-04-12T18:25:39.586+0000"},{"body":"Committed that first part. I think we should keep that open to fix the tmp files for failed compaction and move the rest to another ticket (like CASSANDRA-2433 for instance).\n\nAbout the attached patch on cleaning up failed compaction:\n * We should also handle cleanup and scrub\n * We should handle SSTableWriter.Builder as it is yet another place where we could miss to cleanup a tmp file on error.\n * In theory a failed flush could leave a tmp file behind. If that happens having a tmp file would be the least of your problem but for completeness sake we could handle it.\n * The logging when failing to close iwriter and dataFile in SSTableWriter could probably go at error (we should not be failing there, if we do something is wrong)\n * That's nitpick but I'm not a huge fan of catching RuntimeException in this case as this pollute the code for something that would be a programming error (that's probably debatable though). Maybe another solution would be to have this in the final block. It means making sure closeAndDelete() is ok with the file being already closed and/or deleted and having this final block *after* the closeAndOpenReader call.\n","created":"2011-04-12T19:39:50.523+0000"},{"body":"Thanks will take another look at the cleanup for compaction. \n","created":"2011-04-13T03:08:25.557+0000"},{"body":"Created CASSANDRA-2468 for compaction cleanup. Will close this one for streaming.","created":"2011-04-13T21:06:38.012+0000"}],"conversations":[{"body":"","from":"reporter","subject":"Clean up after failed (repair) streaming operation"},{"body":"Regarding repair: http://www.mail-archive.com/user@cassandra.apache.org/msg09259.html\nAnd compaction: CASSANDRA-2084","from":"developer"},{"body":"I'm keen to try this ticket (to learn more about compaction and repair) if it's not already been worked on. Also if it's ok for me to take a couple of days while I dig into this.\n\nFor compaction I'm looking in\n- CompactionManager.doCompaction where it creates a new SSTableWriter via cfs.createCompactionWriter() \n- CompactionManager.doCleanupCompaction() also uses an SSTableWriter\n\nAre the sorts of failures we're considering for compaction ones that come from the CompactionIterator or SSTableScanner ?\n\nFor repair I'm looking in:\n- IncomingStreamReader appears to clean up the temporary pending file in some error situations. Do we have any more info on the sorts of failures here? e.g. If there is an IOException sending the re-stream message, or a non checked exception it will fail to cleaup the file. \n- I'm looking into what happens in StreamInSession.finished() closeIfFinished()\n- Are we considering failures during the streaming or when processing the data after the stream has finished?\n\nAny guidance welcome. ","from":"developer"},{"body":"bq. Are the sorts of failures we're considering for compaction ones that come from the CompactionIterator or SSTableScanner ?\n\nBoth. Also I suppose it's possible for the writer to error out from lack of disk space since it only checks at the beginning for space and doesn't \"reserve\" it vs flushes.\n\nbq. Are we considering failures during the streaming or when processing the data after the stream has finished?\n\nThe former is much more common (I've never seen the latter reported), so I'd start with that.","from":"developer"},{"body":"patch 0001 tracks failures during AES streaming, files for failed Stream sessions are cleaned up and repair is allowed to continue. Failed files are logged at the StreamSession, TreeRequest, and RepairSession level. \n\npatch 0002 handle exceptions when doing a (normal) compaction and deletes the temp SSTable. The SSTableWriter components are closed before deletion so that windows will delete correctly. ","from":"developer"},{"body":"I think there is a few different things here and I think we should separate them somehow.\n\nFixing the fact that streaming leave tmp files around when it fails is a 2 lines fix and I think this is simple enough that it could go to 0.7. I'm attaching a patch against 0.7. It's extracted from Aaron first patch, although rebased on 0.7 (and fix a bug).\n\nMaking repair aware that there has been some failures is actually more complicated so that should go in 0.8.1 or something (and should go to CASSANDRA-2433 or another ticket that describe the problem better). ","from":"developer"},{"body":"bq. I'm attaching a patch against 0.7\n\nIs that 0001-Better-detect-failures-from-the-other-side-in-Incomi.patch? I don't see the connection to .tmp files. (Also: have you verified that the channel will actually infinite-loop returning 0? Kind of odd behavior, although I guess it's technically within-spec.)","from":"developer"},{"body":"bq. Is that 0001-Better-detect-failures-from-the-other-side-in-Incomi.patch? I don't see the connection to .tmp files. (Also: have you verified that the channel will actually infinite-loop returning 0? Kind of odd behavior, although I guess it's technically within-spec.)\n\nYes. IncomingStreamReader does clean the tmp file when there is an expection (there's an enclosing 'try catch'). The problem is that no exception is raised if the other side of the connection dies. What will happen then is the read will infinitely read 0 bytes. So this actually avoid the infinite loop returning 0 (and so I think answered your second question, so it wasn't very clear).\n\nNote that without this patch, there is an infinite loop that will hold a socket open forever (and consume cpu, though very few probably in that case). So this is not just merely a fix of deleting the tmp files. But it does as a consequence of correctly raising an exception when should be.","from":"developer"},{"body":"+1, and can you move some of that explanation inline as a comment?","from":"developer"},{"body":"Committed that first part. I think we should keep that open to fix the tmp files for failed compaction and move the rest to another ticket (like CASSANDRA-2433 for instance).\n\nAbout the attached patch on cleaning up failed compaction:\n * We should also handle cleanup and scrub\n * We should handle SSTableWriter.Builder as it is yet another place where we could miss to cleanup a tmp file on error.\n * In theory a failed flush could leave a tmp file behind. If that happens having a tmp file would be the least of your problem but for completeness sake we could handle it.\n * The logging when failing to close iwriter and dataFile in SSTableWriter could probably go at error (we should not be failing there, if we do something is wrong)\n * That's nitpick but I'm not a huge fan of catching RuntimeException in this case as this pollute the code for something that would be a programming error (that's probably debatable though). Maybe another solution would be to have this in the final block. It means making sure closeAndDelete() is ok with the file being already closed and/or deleted and having this final block *after* the closeAndOpenReader call.\n","from":"developer"},{"body":"Thanks will take another look at the cleanup for compaction. \n","from":"developer"},{"body":"Created CASSANDRA-2468 for compaction cleanup. Will close this one for streaming.","from":"developer"}],"created":"2011-02-01T05:49:53.000+0000","description":"","issue_id":"12497286","key":"CASSANDRA-2088","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-04-13T21:06:46.000+0000","role":"fixed_distractor","summary":"Clean up after failed (repair) streaming operation"} {"case_id":"12501390","cluster":"DISTRACTOR-CASSANDRA-2324","comments":[{"body":"Key out of order is because sstableexport doesn't know that index sstables use LocalPartitioner instead of the cluster partitioner RP or BOPP.","created":"2011-03-14T19:44:51.302+0000"},{"body":"It looks like INDEXED_RANGE_SLICE is broken in stress.java, so the only problem here is repair doing superfluous work.","created":"2011-03-14T21:24:15.321+0000"},{"body":"The problem is, the ranges repair hashes are not actual node ranges.\n\nLet's consider the following ring (RF=2), where I consider token being in [0..12] to simplify, and where everything is consistent:\n{noformat}\n _.-\"\"\"\"-._\n C (token: 11) .' `.\n [11,3][3,7] / \\\n | |\n | | A(token: 3)\n | | [3,7],[7,11]\n \\ /\n `._ _.'\n B (token: 7)`-....-'\n [7,11],[11,3]\n{noformat}\nNow say I run a repair on node A. The problem is that the Merkle tree ranges are built by dividing the full range by 2 recursively. This means that in this example, the ranges in the tree will for instance be [0,2], [2, 4], [4, 6], [6, 8], [8,10] and [10,12].\n\nIf you compare the hashes for A and B on those ranges, changes are you'll find mismatches for [6,8] and [10,12] (because A don't have anyone on [11, 12] while B have, and B don't have anyone on [6, 7] while A have). As a consequence, the range [7,8] and [10,11] will be repaired, even though there is no inconsistencies.\n\nWhat that means in practice is that it will be very rare for anti-antropy to actually consider the nodes in sync, it will almost surely \"repair\" something, even if the nodes are perfectly consistent. It's Very easy to check btw: with a cluster right the one above (3 nodes, RF=2), with as few as 5 keys for the whole cluster I'm able to have a repair do repairs over and over again.\n\nNow the good question is: how bad is it ? I'm not sure, I depends a bit.\n\nOn a 3 nodes cluster (RF=2), I tried inserting 1M keys with stress (stress -l 2) and triggered repair afterwards. The amount of (unnecessarily) repaired keys was around 150 keys for a given node (it varies slightly for run to run because there is some randomness in the creation of the Merkle tree), corresponding to ~44KB streamed (that is the amount transfered to the node where repair has been ran, so for the total operation its twice this, since we stream in both ways). That's ~0.02% of keys (a given node have ~666 666 keys). It's bad to do useless work, but not a really big deal.\n\nHowever, the less keys we'll have, the worst it gets (and the bigger our rows are, the more useless transfer we do). With the same experiment inserting only 10K keys, there is 190 keys uselessly repaired. That's now close to 3% of the load. It also gets worst with increasing replication factor.\n\n\nTo fix this, we would need for the range in the Merkle tree to \"share\" the node range boundaries. An interesting way to do this would be to have the coordinating node give a list a range for which to calculate Merkle trees, and the node would compute one tree by range (for the coordinating node, that would be #RF's tree). A nice think with this is that it would leave room to optimizing repair since a node would need to do a validation compaction only on the range asked for, which means that only the coordinator node would validate all its data. The neighbors would do less work.\n","created":"2011-03-28T12:57:17.349+0000"},{"body":"bq. To fix this, we would need for the range in the Merkle tree to \"share\" the node range boundaries\n\ncouldn't we just take the interesection of the computed ranges w/ the range actually being repaired? ","created":"2011-03-28T13:44:54.480+0000"},{"body":"bq. couldn't we just take the interesection of the computed ranges w/ the range actually being repaired?\n\nWe do that. But the problem is: you're node A and you receive a merkle tree from B that in particular says that for the range [0..10] the hash is x. And on [0..10] your has is x'. The problem is when [0..10] is partly one of your range, partly not. For instance it can be that you're a replica for [8..10] but not at all for [0..8].\nThis is due to the fact that the ranges for which the hashes are computed are computed without concern for actual node ranges. So now you know there is some inconsistency on [0..10] but it may just be that B is responsible for [0..8] and have data for it (and we don't since we are not in charge of that).\nIn that case, the code do take the intersection of [0..10] with the local range and will stream only [8..10]. But it's still useless.","created":"2011-03-28T13:58:08.902+0000"},{"body":"I thought repair is per-token-range, i.e., if I say \"nodetool repair A\" then range (11, 3] and (3, 7] will be repaired independently.","created":"2011-03-28T14:14:15.055+0000"},{"body":"No, not if I read this code correctly (but I think it should, that's roughly what I'm proposing to do).\n\nActually thinking about it, there is probably no need to construct multiple merkle trees, it will be enough for neighbors to only add to the tree the keys that are in the range of the node asking for the tree.","created":"2011-03-28T14:34:42.487+0000"},{"body":"So what about this:\n\n- change the atom of repair (in nodetool + StorageService) to be a single token range, so it's unambiguous what we're repairing. This has the side benefit of making it enormously easier to repair an entire cluster w/o doing redundant work.\n- provide backwards compatibility w/ existing repair command by splitting it into RF repair ranges and waiting on each of those futures in StorageService mbean\n","created":"2011-03-28T14:47:23.583+0000"},{"body":"Sounds good, will do.","created":"2011-03-28T15:00:30.157+0000"},{"body":"Attached patch modify repair to operate on one token range at a time. Nodetool repair schedule as many repair session than the node have ranges to perform a full node repair. Note that this is more efficient than previously, since the neighbors of the node will only do a validation compaction on the range they have in common with the node coordinating the repair (instead of validating everything).\n\nThis moreover makes it trivial to add an option to nodetool so that the node only repair it's primary range. That way, you can repair a full cluster by calling this operation on every node and there is no duplication of work. The patch doesn't add this option yet though.\n\nThe patch is against trunk. Because the way we construct the merkleTree is fundamentally different, the trees created by 0.7 cannot be compared to the ones created with this patch. The strategy this patch adopts with respect to talking to 0.7 nodes is this:\n * If a 0.7 node asks for a merkleTree, since we are still able to do a full compaction validation, we do it and answer with that.\n * Since a 0.7 node cannot do a merkleTree that would be ok for us, we simply exclude 0.7 nodes from the endpoints we ask merkleTree from.\n\nI don't feel this is a trivially enough patch to go to the 0.7 branch.","created":"2011-04-01T15:00:07.863+0000"},{"body":"This change definitely makes sense: thanks for tackling it. The original implementation was intended to take advantage of naturally occurring compactions: I would still like to get in a position where that is possible, but living with the existing implementation until then isn't worth it.\n\nFrom a quick skim: forceTableRepair incorrectly reports that the session has failed if the client thread dies: the repair will continue in the background (or used to).","created":"2011-04-01T22:24:35.104+0000"},{"body":"I'll give this a more complete review over the weekend.","created":"2011-04-01T23:03:43.791+0000"},{"body":"* SSTableBoundScanner might be much simpler if it iterates within a list of file offsets, as returned by SSTableReader.getPositionsForRanges\n* SSTableReader.getKeySamples could perform two binary searches for min and max rather than doing sequential comparisons to the keys\n\nThanks again Sylvain: this is great!","created":"2011-04-07T08:22:56.538+0000"},{"body":"bq. SSTableBoundScanner might be much simpler if it iterates within a list of file offsets, as returned by SSTableReader.getPositionsForRanges\n\nGood call, that's much simpler. Thanks.\n\nbq. SSTableReader.getKeySamples could perform two binary searches for min and max rather than doing sequential comparisons to the keys\n\nYeah, realized getKeySamples was buggy anyway since it wasn't handling wrapping ranges correctly.\n\nAttaching patch that simplify the bounded scanner and fixes getKeySamples.\n","created":"2011-04-08T01:29:53.274+0000"},{"body":"+1\nThanks!","created":"2011-04-08T04:45:03.354+0000"},{"body":"Committed as r1090840. Thanks.","created":"2011-04-10T18:16:27.258+0000"}],"conversations":[{"body":"To repro: 3 node cluster, stress.java 1M rows with -x KEYS and -l 2. The index is enough to make some mutations drop (about 20-30k total in my tests). Repair afterwards will repair a large amount of ranges the first time. However, each subsequent run will repair the same set of small ranges every time. INDEXED_RANGE_SLICE in stress never fully works. Counting rows with sstablekeys shows there are 2M rows total as expected, however when trying to count the indexed keys, I get exceptions like:\n{noformat}\nException in thread \"main\" java.io.IOException: Key out of order! DecoratedKey(101571366040797913119296586470838356016, 0707ab782c5b5029d28a5e6d508ef72f0222528b5e28da3b7787492679dc51b96f868e0746073e54bc173be927049d0f51e25a6a95b3268213b8969abf40cea7d7) > DecoratedKey(12639574763031545147067490818595764132, 0bc414be3093348a2ad389ed28f18f0cc9a044b2e98587848a0d289dae13ed0ad479c74654900eeffc6236)\n at org.apache.cassandra.tools.SSTableExport.enumeratekeys(SSTableExport.java:206)\n at org.apache.cassandra.tools.SSTableExport.main(SSTableExport.java:388)\n{noformat}","from":"reporter","subject":"Repair transfers more data than necessary"},{"body":"Key out of order is because sstableexport doesn't know that index sstables use LocalPartitioner instead of the cluster partitioner RP or BOPP.","from":"developer"},{"body":"It looks like INDEXED_RANGE_SLICE is broken in stress.java, so the only problem here is repair doing superfluous work.","from":"developer"},{"body":"The problem is, the ranges repair hashes are not actual node ranges.\n\nLet's consider the following ring (RF=2), where I consider token being in [0..12] to simplify, and where everything is consistent:\n{noformat}\n _.-\"\"\"\"-._\n C (token: 11) .' `.\n [11,3][3,7] / \\\n | |\n | | A(token: 3)\n | | [3,7],[7,11]\n \\ /\n `._ _.'\n B (token: 7)`-....-'\n [7,11],[11,3]\n{noformat}\nNow say I run a repair on node A. The problem is that the Merkle tree ranges are built by dividing the full range by 2 recursively. This means that in this example, the ranges in the tree will for instance be [0,2], [2, 4], [4, 6], [6, 8], [8,10] and [10,12].\n\nIf you compare the hashes for A and B on those ranges, changes are you'll find mismatches for [6,8] and [10,12] (because A don't have anyone on [11, 12] while B have, and B don't have anyone on [6, 7] while A have). As a consequence, the range [7,8] and [10,11] will be repaired, even though there is no inconsistencies.\n\nWhat that means in practice is that it will be very rare for anti-antropy to actually consider the nodes in sync, it will almost surely \"repair\" something, even if the nodes are perfectly consistent. It's Very easy to check btw: with a cluster right the one above (3 nodes, RF=2), with as few as 5 keys for the whole cluster I'm able to have a repair do repairs over and over again.\n\nNow the good question is: how bad is it ? I'm not sure, I depends a bit.\n\nOn a 3 nodes cluster (RF=2), I tried inserting 1M keys with stress (stress -l 2) and triggered repair afterwards. The amount of (unnecessarily) repaired keys was around 150 keys for a given node (it varies slightly for run to run because there is some randomness in the creation of the Merkle tree), corresponding to ~44KB streamed (that is the amount transfered to the node where repair has been ran, so for the total operation its twice this, since we stream in both ways). That's ~0.02% of keys (a given node have ~666 666 keys). It's bad to do useless work, but not a really big deal.\n\nHowever, the less keys we'll have, the worst it gets (and the bigger our rows are, the more useless transfer we do). With the same experiment inserting only 10K keys, there is 190 keys uselessly repaired. That's now close to 3% of the load. It also gets worst with increasing replication factor.\n\n\nTo fix this, we would need for the range in the Merkle tree to \"share\" the node range boundaries. An interesting way to do this would be to have the coordinating node give a list a range for which to calculate Merkle trees, and the node would compute one tree by range (for the coordinating node, that would be #RF's tree). A nice think with this is that it would leave room to optimizing repair since a node would need to do a validation compaction only on the range asked for, which means that only the coordinator node would validate all its data. The neighbors would do less work.\n","from":"developer"},{"body":"bq. To fix this, we would need for the range in the Merkle tree to \"share\" the node range boundaries\n\ncouldn't we just take the interesection of the computed ranges w/ the range actually being repaired? ","from":"developer"},{"body":"bq. couldn't we just take the interesection of the computed ranges w/ the range actually being repaired?\n\nWe do that. But the problem is: you're node A and you receive a merkle tree from B that in particular says that for the range [0..10] the hash is x. And on [0..10] your has is x'. The problem is when [0..10] is partly one of your range, partly not. For instance it can be that you're a replica for [8..10] but not at all for [0..8].\nThis is due to the fact that the ranges for which the hashes are computed are computed without concern for actual node ranges. So now you know there is some inconsistency on [0..10] but it may just be that B is responsible for [0..8] and have data for it (and we don't since we are not in charge of that).\nIn that case, the code do take the intersection of [0..10] with the local range and will stream only [8..10]. But it's still useless.","from":"developer"},{"body":"I thought repair is per-token-range, i.e., if I say \"nodetool repair A\" then range (11, 3] and (3, 7] will be repaired independently.","from":"developer"},{"body":"No, not if I read this code correctly (but I think it should, that's roughly what I'm proposing to do).\n\nActually thinking about it, there is probably no need to construct multiple merkle trees, it will be enough for neighbors to only add to the tree the keys that are in the range of the node asking for the tree.","from":"developer"},{"body":"So what about this:\n\n- change the atom of repair (in nodetool + StorageService) to be a single token range, so it's unambiguous what we're repairing. This has the side benefit of making it enormously easier to repair an entire cluster w/o doing redundant work.\n- provide backwards compatibility w/ existing repair command by splitting it into RF repair ranges and waiting on each of those futures in StorageService mbean\n","from":"developer"},{"body":"Sounds good, will do.","from":"developer"},{"body":"Attached patch modify repair to operate on one token range at a time. Nodetool repair schedule as many repair session than the node have ranges to perform a full node repair. Note that this is more efficient than previously, since the neighbors of the node will only do a validation compaction on the range they have in common with the node coordinating the repair (instead of validating everything).\n\nThis moreover makes it trivial to add an option to nodetool so that the node only repair it's primary range. That way, you can repair a full cluster by calling this operation on every node and there is no duplication of work. The patch doesn't add this option yet though.\n\nThe patch is against trunk. Because the way we construct the merkleTree is fundamentally different, the trees created by 0.7 cannot be compared to the ones created with this patch. The strategy this patch adopts with respect to talking to 0.7 nodes is this:\n * If a 0.7 node asks for a merkleTree, since we are still able to do a full compaction validation, we do it and answer with that.\n * Since a 0.7 node cannot do a merkleTree that would be ok for us, we simply exclude 0.7 nodes from the endpoints we ask merkleTree from.\n\nI don't feel this is a trivially enough patch to go to the 0.7 branch.","from":"developer"},{"body":"This change definitely makes sense: thanks for tackling it. The original implementation was intended to take advantage of naturally occurring compactions: I would still like to get in a position where that is possible, but living with the existing implementation until then isn't worth it.\n\nFrom a quick skim: forceTableRepair incorrectly reports that the session has failed if the client thread dies: the repair will continue in the background (or used to).","from":"developer"},{"body":"I'll give this a more complete review over the weekend.","from":"developer"},{"body":"* SSTableBoundScanner might be much simpler if it iterates within a list of file offsets, as returned by SSTableReader.getPositionsForRanges\n* SSTableReader.getKeySamples could perform two binary searches for min and max rather than doing sequential comparisons to the keys\n\nThanks again Sylvain: this is great!","from":"developer"},{"body":"bq. SSTableBoundScanner might be much simpler if it iterates within a list of file offsets, as returned by SSTableReader.getPositionsForRanges\n\nGood call, that's much simpler. Thanks.\n\nbq. SSTableReader.getKeySamples could perform two binary searches for min and max rather than doing sequential comparisons to the keys\n\nYeah, realized getKeySamples was buggy anyway since it wasn't handling wrapping ranges correctly.\n\nAttaching patch that simplify the bounded scanner and fixes getKeySamples.\n","from":"developer"},{"body":"+1\nThanks!","from":"developer"},{"body":"Committed as r1090840. Thanks.","from":"developer"}],"created":"2011-03-14T19:43:08.000+0000","description":"To repro: 3 node cluster, stress.java 1M rows with -x KEYS and -l 2. The index is enough to make some mutations drop (about 20-30k total in my tests). Repair afterwards will repair a large amount of ranges the first time. However, each subsequent run will repair the same set of small ranges every time. INDEXED_RANGE_SLICE in stress never fully works. Counting rows with sstablekeys shows there are 2M rows total as expected, however when trying to count the indexed keys, I get exceptions like:\n{noformat}\nException in thread \"main\" java.io.IOException: Key out of order! DecoratedKey(101571366040797913119296586470838356016, 0707ab782c5b5029d28a5e6d508ef72f0222528b5e28da3b7787492679dc51b96f868e0746073e54bc173be927049d0f51e25a6a95b3268213b8969abf40cea7d7) > DecoratedKey(12639574763031545147067490818595764132, 0bc414be3093348a2ad389ed28f18f0cc9a044b2e98587848a0d289dae13ed0ad479c74654900eeffc6236)\n at org.apache.cassandra.tools.SSTableExport.enumeratekeys(SSTableExport.java:206)\n at org.apache.cassandra.tools.SSTableExport.main(SSTableExport.java:388)\n{noformat}","issue_id":"12501390","key":"CASSANDRA-2324","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-04-10T18:16:27.000+0000","role":"fixed_distractor","summary":"Repair transfers more data than necessary"} {"case_id":"12502232","cluster":"DISTRACTOR-CASSANDRA-2371","comments":[{"body":"My guess is what is happening is that after aVeryLongTime when the state is evicted, another gossip round occurs and the state is repopulated. Patch to re-quarantine on eviction to avoid this.","created":"2011-03-23T22:43:23.439+0000"},{"body":"This looks like it's changing the source code. Can we deploy this on a live cluster? ","created":"2011-03-23T23:30:13.039+0000"},{"body":"Did you get a chance to try a patched build?","created":"2011-04-01T04:09:28.048+0000"},{"body":"Actually the issue was resolved when we did a restart of the cluster, and ran the nodetool removetoken command when only a portion of the nodes had started up. The main reason I didn't apply the patch yet, is merely because I need to learn how to patch the code and integrate the new build first... Thank you!","created":"2011-04-01T16:48:20.443+0000"},{"body":"There is a second problem here. We don't populate the gossiper's application state to LEFT for removetoken, only for decommission. The problem with adding it, however, is that SS.handleStateLeft gets called every gossip round and runs through the hint removal process. One option may be to check if the node is locally persisted and if not, just ignore the message since we never knew about it anyway. Another is to just not remove hints when we see the LEFT state, because they'll expire anyway and large unaccessed rows aren't a problem, so this seems like a throwback from the <=0.6 days.","created":"2011-04-07T00:05:22.755+0000"},{"body":"I applied the patch and within a few days even with no restarting of any of the Cassandra nodes, the removed token came back. Just FYI.","created":"2011-04-10T09:03:30.789+0000"},{"body":"bq. The problem with adding it, however, is that SS.handleStateLeft gets called every gossip round and runs through the hint removal process\n\nBut it only runs hint removal once per node, right? Or is \"onChange\" not an accurate method name?\n\nbq. One option may be to check if the node is locally persisted and if not, just ignore the message since we never knew about it anyway.\n\nLocally persisted... with a token in SystemTable? Ignore the ... hint message?\n\nbq. Another is to just not remove hints when we see the LEFT state\n\nWe should issue a delete to the hints row, but we should not force a major compaction. The rule of thumb is, avoiding inflicting a performance hit on the cluster trumps immediate disk space cleanup.","created":"2011-04-11T16:11:27.749+0000"},{"body":"It looks like the bad news here is that gossip makes many assumptions about a node being alive when it hears about it, and has no provisions for keeping removed node state. This is bad, but highly exacerbated by keeping the removed node state for 3 days instead of 30s. I think the best solution right now is to revert CASSANDRA-2115 until we can overhaul gossip to handle this. The bug will still exist as it always has, but the window to trigger it is far shorter.","created":"2011-04-14T22:55:44.056+0000"},{"body":"Sounds reasonable.","created":"2011-04-15T00:50:16.810+0000"},{"body":"Reverted CASSANDRA-2115 and created CASSANDRA-2496 to fix this correctly.","created":"2011-04-18T18:59:23.202+0000"}],"conversations":[{"body":"The removetoken option does not seem to work. The original node 10.240.50.63 comes back into the ring, even after the EC2 instance is no longer in existence. Originally I tried to add a new node 10.214.103.224 with the same token, but there were some complications with that. I have pasted below all the INFO log entries found with greping the system log files.\n\nSeems to be a similar issue seen with http://cassandra-user-incubator-apache-org.3065146.n2.nabble.com/Ghost-node-showing-up-in-the-ring-td6198180.html \n\nINFO [GossipStage:1] 2011-03-16 00:54:31,590 StorageService.java (line 745) Nodes /10.214.103.224 and /10.240.50.63 have the same token 95704415696513900000000000000000000000. /10.214.103.224 is the new owner\n INFO [GossipStage:1] 2011-03-16 17:26:51,083 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.214.103.224\n INFO [GossipStage:1] 2011-03-19 17:27:24,767 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.214.103.224\n INFO [GossipStage:1] 2011-03-19 17:29:30,191 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.214.103.224\n INFO [GossipStage:1] 2011-03-19 17:31:35,609 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.214.103.224\n INFO [GossipStage:1] 2011-03-19 17:33:39,440 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.214.103.224\n INFO [GossipStage:1] 2011-03-23 17:22:55,520 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.240.50.63\n\n\n INFO [GossipStage:1] 2011-03-10 03:52:37,299 Gossiper.java (line 608) Node /10.240.50.63 is now part of the cluster\n INFO [GossipStage:1] 2011-03-10 03:52:37,545 Gossiper.java (line 600) InetAddress /10.240.50.63 is now UP\n INFO [HintedHandoff:1] 2011-03-10 03:53:36,168 HintedHandOffManager.java (line 304) Started hinted handoff for endpoint /10.240.50.63\n INFO [HintedHandoff:1] 2011-03-10 03:53:36,169 HintedHandOffManager.java (line 360) Finished hinted handoff of 0 rows to endpoint /10.240.50.63\n INFO [GossipStage:1] 2011-03-15 23:23:43,770 Gossiper.java (line 623) Node /10.240.50.63 has restarted, now UP again\n INFO [GossipStage:1] 2011-03-15 23:23:43,771 StorageService.java (line 726) Node /10.240.50.63 state jump to normal\n INFO [HintedHandoff:1] 2011-03-15 23:28:48,957 HintedHandOffManager.java (line 304) Started hinted handoff for endpoint /10.240.50.63\n INFO [HintedHandoff:1] 2011-03-15 23:28:48,958 HintedHandOffManager.java (line 360) Finished hinted handoff of 0 rows to endpoint /10.240.50.63\n INFO [ScheduledTasks:1] 2011-03-15 23:37:25,071 Gossiper.java (line 226) InetAddress /10.240.50.63 is now dead.\n INFO [GossipStage:1] 2011-03-16 00:54:31,590 StorageService.java (line 745) Nodes /10.214.103.224 and /10.240.50.63 have the same token 95704415696513900000000000000000000000. /10.214.103.224 is the new owner\n WARN [GossipStage:1] 2011-03-16 00:54:31,590 TokenMetadata.java (line 115) Token 95704415696513900000000000000000000000 changing ownership from /10.240.50.63 to /10.214.103.224\n INFO [GossipStage:1] 2011-03-18 23:37:09,158 Gossiper.java (line 610) Node /10.240.50.63 is now part of the cluster\n INFO [GossipStage:1] 2011-03-21 23:37:10,421 Gossiper.java (line 610) Node /10.240.50.63 is now part of the cluster\n INFO [GossipStage:1] 2011-03-21 23:37:10,421 StorageService.java (line 726) Node /10.240.50.63 state jump to normal\n INFO [GossipStage:1] 2011-03-23 17:22:55,520 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.240.50.63\n INFO [ScheduledTasks:1] 2011-03-23 17:22:55,521 HintedHandOffManager.java (line 210) Deleting any stored hints for 10.240.50.63\n","from":"reporter","subject":"Removed/Dead Node keeps reappearing"},{"body":"My guess is what is happening is that after aVeryLongTime when the state is evicted, another gossip round occurs and the state is repopulated. Patch to re-quarantine on eviction to avoid this.","from":"developer"},{"body":"This looks like it's changing the source code. Can we deploy this on a live cluster? ","from":"developer"},{"body":"Did you get a chance to try a patched build?","from":"developer"},{"body":"Actually the issue was resolved when we did a restart of the cluster, and ran the nodetool removetoken command when only a portion of the nodes had started up. The main reason I didn't apply the patch yet, is merely because I need to learn how to patch the code and integrate the new build first... Thank you!","from":"developer"},{"body":"There is a second problem here. We don't populate the gossiper's application state to LEFT for removetoken, only for decommission. The problem with adding it, however, is that SS.handleStateLeft gets called every gossip round and runs through the hint removal process. One option may be to check if the node is locally persisted and if not, just ignore the message since we never knew about it anyway. Another is to just not remove hints when we see the LEFT state, because they'll expire anyway and large unaccessed rows aren't a problem, so this seems like a throwback from the <=0.6 days.","from":"developer"},{"body":"I applied the patch and within a few days even with no restarting of any of the Cassandra nodes, the removed token came back. Just FYI.","from":"developer"},{"body":"bq. The problem with adding it, however, is that SS.handleStateLeft gets called every gossip round and runs through the hint removal process\n\nBut it only runs hint removal once per node, right? Or is \"onChange\" not an accurate method name?\n\nbq. One option may be to check if the node is locally persisted and if not, just ignore the message since we never knew about it anyway.\n\nLocally persisted... with a token in SystemTable? Ignore the ... hint message?\n\nbq. Another is to just not remove hints when we see the LEFT state\n\nWe should issue a delete to the hints row, but we should not force a major compaction. The rule of thumb is, avoiding inflicting a performance hit on the cluster trumps immediate disk space cleanup.","from":"developer"},{"body":"It looks like the bad news here is that gossip makes many assumptions about a node being alive when it hears about it, and has no provisions for keeping removed node state. This is bad, but highly exacerbated by keeping the removed node state for 3 days instead of 30s. I think the best solution right now is to revert CASSANDRA-2115 until we can overhaul gossip to handle this. The bug will still exist as it always has, but the window to trigger it is far shorter.","from":"developer"},{"body":"Sounds reasonable.","from":"developer"},{"body":"Reverted CASSANDRA-2115 and created CASSANDRA-2496 to fix this correctly.","from":"developer"}],"created":"2011-03-23T22:03:42.000+0000","description":"The removetoken option does not seem to work. The original node 10.240.50.63 comes back into the ring, even after the EC2 instance is no longer in existence. Originally I tried to add a new node 10.214.103.224 with the same token, but there were some complications with that. I have pasted below all the INFO log entries found with greping the system log files.\n\nSeems to be a similar issue seen with http://cassandra-user-incubator-apache-org.3065146.n2.nabble.com/Ghost-node-showing-up-in-the-ring-td6198180.html \n\nINFO [GossipStage:1] 2011-03-16 00:54:31,590 StorageService.java (line 745) Nodes /10.214.103.224 and /10.240.50.63 have the same token 95704415696513900000000000000000000000. /10.214.103.224 is the new owner\n INFO [GossipStage:1] 2011-03-16 17:26:51,083 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.214.103.224\n INFO [GossipStage:1] 2011-03-19 17:27:24,767 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.214.103.224\n INFO [GossipStage:1] 2011-03-19 17:29:30,191 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.214.103.224\n INFO [GossipStage:1] 2011-03-19 17:31:35,609 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.214.103.224\n INFO [GossipStage:1] 2011-03-19 17:33:39,440 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.214.103.224\n INFO [GossipStage:1] 2011-03-23 17:22:55,520 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.240.50.63\n\n\n INFO [GossipStage:1] 2011-03-10 03:52:37,299 Gossiper.java (line 608) Node /10.240.50.63 is now part of the cluster\n INFO [GossipStage:1] 2011-03-10 03:52:37,545 Gossiper.java (line 600) InetAddress /10.240.50.63 is now UP\n INFO [HintedHandoff:1] 2011-03-10 03:53:36,168 HintedHandOffManager.java (line 304) Started hinted handoff for endpoint /10.240.50.63\n INFO [HintedHandoff:1] 2011-03-10 03:53:36,169 HintedHandOffManager.java (line 360) Finished hinted handoff of 0 rows to endpoint /10.240.50.63\n INFO [GossipStage:1] 2011-03-15 23:23:43,770 Gossiper.java (line 623) Node /10.240.50.63 has restarted, now UP again\n INFO [GossipStage:1] 2011-03-15 23:23:43,771 StorageService.java (line 726) Node /10.240.50.63 state jump to normal\n INFO [HintedHandoff:1] 2011-03-15 23:28:48,957 HintedHandOffManager.java (line 304) Started hinted handoff for endpoint /10.240.50.63\n INFO [HintedHandoff:1] 2011-03-15 23:28:48,958 HintedHandOffManager.java (line 360) Finished hinted handoff of 0 rows to endpoint /10.240.50.63\n INFO [ScheduledTasks:1] 2011-03-15 23:37:25,071 Gossiper.java (line 226) InetAddress /10.240.50.63 is now dead.\n INFO [GossipStage:1] 2011-03-16 00:54:31,590 StorageService.java (line 745) Nodes /10.214.103.224 and /10.240.50.63 have the same token 95704415696513900000000000000000000000. /10.214.103.224 is the new owner\n WARN [GossipStage:1] 2011-03-16 00:54:31,590 TokenMetadata.java (line 115) Token 95704415696513900000000000000000000000 changing ownership from /10.240.50.63 to /10.214.103.224\n INFO [GossipStage:1] 2011-03-18 23:37:09,158 Gossiper.java (line 610) Node /10.240.50.63 is now part of the cluster\n INFO [GossipStage:1] 2011-03-21 23:37:10,421 Gossiper.java (line 610) Node /10.240.50.63 is now part of the cluster\n INFO [GossipStage:1] 2011-03-21 23:37:10,421 StorageService.java (line 726) Node /10.240.50.63 state jump to normal\n INFO [GossipStage:1] 2011-03-23 17:22:55,520 StorageService.java (line 865) Removing token 95704415696513900000000000000000000000 for /10.240.50.63\n INFO [ScheduledTasks:1] 2011-03-23 17:22:55,521 HintedHandOffManager.java (line 210) Deleting any stored hints for 10.240.50.63\n","issue_id":"12502232","key":"CASSANDRA-2371","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-04-18T18:59:23.000+0000","role":"fixed_distractor","summary":"Removed/Dead Node keeps reappearing"} {"case_id":"12502417","cluster":"DISTRACTOR-CASSANDRA-2388","comments":[{"body":"I'm not sure special casing NoRouteToHostException to be blacklisted is the best thing to do. I don't think connections are being setup so often that maintaining a blacklist for any reason is needed.","created":"2011-03-25T22:45:55.124+0000"},{"body":"Unfortunately, I thought of another problem here. If we go over the entire replica set, we're potentially going outside of the DC, which is bad since a lot of installations have a DC dedicated to analytics so it doesn't affect their app. It seems that the local address is preferred though, are your task trackers not on the same machines as Cassandra?","created":"2011-03-26T18:24:39.652+0000"},{"body":"Eldon, are you planning to take another stab at this?","created":"2011-05-17T21:07:16.610+0000"},{"body":"We need to return the list of replicas in the same DC","created":"2011-05-17T21:09:01.890+0000"},{"body":"Mck, do you want to take a stab at this?","created":"2011-05-23T14:45:44.358+0000"},{"body":"I'm having a go currently at CASSANDRA-1125 so i might as well look at this too. (but you've caught me on a holiday-week...)","created":"2011-05-23T16:27:46.513+0000"},{"body":"How do i obtain the DataCenter name for a given address?\n\nIEndpointSnitch.getDataCenter(inetAddress) would work nicely for me but how do i get the snitch client-side?","created":"2011-06-07T14:11:07.463+0000"},{"body":"Initial attempt at solution. Although I'm a little apprehensive to the additions to cassandra.thrift\n(describe_rack(..) isn't used anywhere, it just made sense to add describe_datacenter(..) and describe_rack(..) at the same time).\n\nI've tested that existing hadoop jobs work but the new functionality hasn't been tested (as i currently don't have any RF=2 data setup).\n\nThis patch does not include the required re-generated Cassandra.java","created":"2011-06-07T20:41:15.028+0000"},{"body":"Second attempt. (god only knows what i was trying to test last patch ;)\nthis patch:\n - adds describe_datacenter and describe_rack to cassandra.thrift\n - adds locations in ColumnFamilyRecordReader from the split's alternative endpoints if dc is the same\n\nThis patch does not include the required re-generated Cassandra.java\n","created":"2011-06-09T06:29:04.881+0000"},{"body":"I have tested this now on data w/ RF=2.\nSeems to work ~ok as far as i can see.\n\nOne side-effect of this patch is where once one could configure ConfigHelper.setInitialAddress(conf, \"localhost\") this will no longer work for tasks trying to run on the down node.\nColumnFamilyRecordReader.getLocations() will ConnectException trying to call describe_datacenter(..). This will lead to the task failing. Hadoop re-runs the task then on another node and eventually the job will complete. But the fall back to replica never is used.\n\nIf the initialAddress is hardcoded to one node then we no longer have a decentralised job.\n\nI would like to allow a comma-separated in initialAddress, for example it could be \"localhost, node01, node02, node03\". This would give preference to localhost and avoid any centralisation.\n\nI would also like to make ColumnFamilyRecordReader.getLocations() return an iterator instead of an array.\nThe createConnection(..) and client.describe_datacenter(..) calls are an unnecessary overhead when all nodes (or first endpoint location) are up, and could be avoided by lazy-loading the list.","created":"2011-06-09T10:22:57.601+0000"},{"body":"New patch. I think i'm at last happy with it.\n\ngetLocations() returns an iterator so client.describe_datacenter() is only called when necessary.\n\nRather than provide a list in initialAddress it was possible to use either the initialAddress OR the endpoint. This gave the benefit in not listing a location that can't actually be connected to.\n\nThe \"only use replica from same DC\" is an option now in ConfigHelper. By default it is true.\n\nAgain the re-generated Cassandra.java is not included in the patch.\n\nI have tested this on normal jobs, and RF=2 jobs with a node down.","created":"2011-06-11T09:15:23.698+0000"},{"body":"The get_rack seems unused so it should be removed.\n\nAlso, it might be better to pass all locations in the get_datacenter thrift call since you can get the results in one shot and sort them by the dynamic snitch, filtering out the dead nodes:\n\n{noformat}\n DatabaseDescriptor.getEndpointSnitch().sortByProximity(FBUtilities.getLocalAddress(), endpoints); \n{noformat}\n\n{noformat}\n FailureDetector.instance.isAlive(endpoint)\n{noformat}\n\n","created":"2011-06-13T15:23:24.480+0000"},{"body":"Then (if i understand you correctly) i would need in cassandra.thrift\n{noformat}\n /** returns alive endpoints, sorted by proximity, that belong in the same datacenter as the given endpoint */\n list get_endpoints_in_same_datacenter(1: string endpoint, 2: required list endpoints)\n throws (1:InvalidRequestException ire)\n{noformat}\n\nThen the API becomes quite specific to this usecase. Is the performance gain worth it? What's the cost of each client.describe_datacenter(..) call, and probably more important what is the lost performance of writing to the furthest node that's within the same datacenter?","created":"2011-06-13T16:15:18.976+0000"},{"body":"Just make sure i understand you T Jake, you would rather something like this in CassandraServer.java?\n(I've renamed from the previous comment get_endpoints_in_same_datacenter(..) to sort_endpoints_by_proximity(..))\n{noformat}\n public String[] sort_endpoints_by_proximity(String endpoint, String[] endpoints, boolean restrictToSameDC) \n throws TException, InvalidRequestException\n {\n try\n {\n List results = new ArrayList();\n InetAddress address = InetAddress.getByName(endpoint);\n String datacenter = DatabaseDescriptor.getEndpointSnitch().getDatacenter(address);\n List addresses = new ArrayList();\n for(String ep : endpoints)\n {\n addresses.add(InetAddress.getByName(ep));\n }\n DatabaseDescriptor.getEndpointSnitch().sortByProximity(address, addresses);\n for(InetAddress ep : addresses)\n {\n String dc = DatabaseDescriptor.getEndpointSnitch().getDatacenter(ep);\n if(FailureDetector.instance.isAlive(ep) && (!restrictToSameDC || datacenter.equals(dc)))\n {\n results.add(ep.getHostName());\n }\n }\n return results.toArray(new String[results.size()]);\n }\n catch (UnknownHostException e)\n {\n throw new InvalidRequestException(e.getMessage());\n }\n }\n{noformat}","created":"2011-06-13T16:40:04.268+0000"},{"body":"bq. what is the lost performance of writing to the furthest node that's within the same datacenter?\n\nThe benefit is really the DynamicSnitch. if a node it slow due to compaction then this would avoid sending requests there... \n\nbq. public String[] sort_endpoints_by_proximity(String endpoint, String[] endpoints, boolean restrictToSameDC)\n\nI don't think it makes sense to send the client endpoint to this call since the endpoint might not be a cassandra node. It's a reasonable assumption that the endpoint it's talking to is local enough to the client to use that.\n\n","created":"2011-06-13T16:59:28.003+0000"},{"body":"{quote}\nbq. public String[] sort_endpoints_by_proximity(String endpoint, String[] endpoints, boolean restrictToSameDC)\nI don't think it makes sense to send the client endpoint to this call since the endpoint might not be a cassandra node. It's a reasonable assumption that the endpoint it's talking to is local enough to the client to use that.\n{quote}\nFor the test set i was running against, RF=2, each split's has two endpoints always in different datacenters.\n\nIf the \"local\" endpoint is down then getLocations() will then call client.sort_endpoints_by_proximity(..) and this will fail (being the same endpoint).\nIt then makes a client connection through the \"other\" endpoint. \\[see CFRR.describeDatacenter(..)].\nThis will presume the wrong datacenter and return itself as a valid endpoint. \nI need some way to know what the original datacenter is, even when it is down.","created":"2011-06-13T17:52:38.898+0000"},{"body":"ok but why not change the response to map> where key is DC and value are proximity sorted endpoints?","created":"2011-06-13T18:19:43.689+0000"},{"body":"Won't the sorting still be wrong?\nFor the use-case above it will solve restricting to the correct datacenter, but the sorting will still be based on proximity to the wrong node?\n\nbq. I don't think it makes sense to send the client endpoint to this call since the endpoint might not be a cassandra node. \nIt might not be an alive cassandra node, but it should be a cassandra node. It comes from the split's list of endpoints. At least in this use-case, or are you referring to general usage for this new api?\nbq. It's a reasonable assumption that the endpoint it's talking to is local enough to the client to use that.\nI don't think so... The endpoint that it talks to is a completely random (just the next endpoint listed in the split's list). This is why i think that such sorting won't just be wrong but not even close. Does this make sense?","created":"2011-06-13T19:19:25.613+0000"},{"body":"I think the core issue is you can't assume the hadoop node is running on a cassandra node...\n\nIf it is then the logic is straight forward, if not then it's possible the connection could cross DC boundaries. One possibility is to use the ip octets like the RackInferringSnitch. \n\nHow's this proposal then? keep the sort_endpoints_by_proximity signature as is and pass the client endpoint along with the list of data endpoints and add the following logic:\n\n1) sort the endpoints using the endpoint_snitch.\n2) if client endpoint *is* a valid cassandra node get the nodes DC and prune nodes outside of this DC\n3) if client endpoint *is not* a valid cassandra node try to infer the DC from its ip and prune dataendpoint nodes in a different DC. If no cassandra nodes are in the DC list goto 3).\n4) all else fails return the sorted endpoint list\n","created":"2011-06-14T13:07:49.114+0000"},{"body":"bq. [snip] One possibility is to use the ip octets like the RackInferringSnitch. \n\nIn our usecase we have three nodes defined via PropertyFileSnitch:{noformat}152.90.241.22=DC1:RAC1 #node1\n152.90.241.23=DC2:RAC1 #node2\n152.90.241.24=DC1:RAC1 #node3{noformat}\nThe only way to infer here is even addresses belong to one dc, odd to the other. This is not how RackInferringSnithc works.\n\nWhen we make the connection through the \"other\" (node2) endpoint taking the rack inferring approach \"152.90.\" will say it's in DC2. (again) this is the wrong DC and will return itself as a valid endpoint....\n\nStep (3) seems to me to be too specific to be included here.\nIf i go only with steps (1),(2),and (4) we get this code:{noformat} public String[] sort_endpoints_by_proximity(String endpoint, String[] endpoints, boolean restrictToSameDC) \n throws TException, InvalidRequestException\n {\n try\n {\n List results = new ArrayList();\n InetAddress address = InetAddress.getByName(endpoint);\n boolean endpointValid = null != Gossiper.instance.getEndpointStateForEndpoint(address);\n String datacenter = DatabaseDescriptor\n .getEndpointSnitch().getDatacenter(endpointValid ? address : FBUtilities.getLocalAddress());\n List addresses = new ArrayList();\n for(String ep : endpoints)\n {\n addresses.add(InetAddress.getByName(endpoint));\n }\n DatabaseDescriptor.getEndpointSnitch().sortByProximity(address, addresses);\n for(InetAddress ep : addresses)\n {\n String dc = DatabaseDescriptor.getEndpointSnitch().getDatacenter(ep);\n if(FailureDetector.instance.isAlive(ep) && (!restrictToSameDC || datacenter.equals(dc)))\n {\n results.add(ep.getHostName());\n }\n }\n return results.toArray(new String[results.size()]);\n }\n catch (UnknownHostException e)\n {\n throw new InvalidRequestException(e.getMessage());\n }\n }{noformat}\n\nI'm happy with this (except that {{Gossiper.instance.getEndpointStateForEndpoint(address)}} is only my guess on how to tell if an endpoint is valid as such).","created":"2011-06-15T11:23:15.494+0000"},{"body":"Problem with the suggested approach is that sortByProximity(..) *only* works when address is the local address. See assert statement DynamicEndpointSnitch:134\n\nI could hack this and rewrite the line to\n{noformat}IEndpointSnitch snitch = DatabaseDescriptor.getEndpointSnitch();\nsnitch = snitch instanceof DynamicEndpointSnitch ? ((DynamicEndpointSnitch)snitch).subsnitch : snitch;\nsnitch.sortByProximity(address, addresses);{noformat}\nBut this of course means that we always bypass DynamicEndpointSnitch's \"scores\".","created":"2011-06-22T13:04:06.477+0000"},{"body":"Up to date patch.\nFollows T Jake's points (1),(2), and (4).\nAnd bypasses DynamicEndpointSnitch when sorting by proximity.","created":"2011-06-22T13:46:37.763+0000"},{"body":"committed with a change to use the dynamic snitch id the passed endpoint is valid.","created":"2011-06-24T15:43:18.670+0000"},{"body":"This patch applies to the current 0.7-branch with minimal problems - just some imports on CassandraServer that it couldn't resolve properly. Can this be committed against 0.7-branch for inclusion in 0.7.7?","created":"2011-06-25T00:50:01.457+0000"},{"body":"I've done basic testing with the word count and pig examples to make sure that the basic hadoop integration isn't negatively affected by this. I'll also try it against our dev cluster before and after the patch - killing one node to see if it fails over to another replica - to make sure it does what it should that way.","created":"2011-06-25T01:05:34.143+0000"},{"body":"Reopening for testing against 0.7.6.","created":"2011-06-25T01:15:43.006+0000"},{"body":"Took a look at this belatedly. I don't understand the contortions at all. It looks like there's a ton of effort put in to avoiding making sortByProximity work w/ non-local nodes. Why not just make that work instead?\n","created":"2011-06-25T02:44:39.328+0000"},{"body":"also: running hadoop on a non-cassandra node is dumb. i don't see a point in supporting that really. (yes, my fault it was written that way to begin with, mea culpa.)","created":"2011-06-25T02:48:46.487+0000"},{"body":"Jonathan - is it possible to attach an updated patch based on your changes to 0.8 branch? Not sure if that would be simple to extract.","created":"2011-06-25T02:58:15.155+0000"},{"body":"with svn: svn diff -r 1139323:1139483 and hack out the OutboundTcpConnection change in the middle manually from the output.\n\nwith git: create a branch, rip out the offending OTC change and squash the other two","created":"2011-06-25T03:07:16.985+0000"},{"body":"I think there's deep surgery to be done here still though. Backporting is probably premature.","created":"2011-06-25T03:07:59.703+0000"},{"body":"bq. It looks like there's a ton of effort put in to avoiding making sortByProximity work w/ non-local nodes\n\nWait, why do we even care? \"local node\" IS the right host to sort against -- we want the split that is closest to the node running the job, this is not the same as some other C* node we contact.","created":"2011-06-25T03:12:32.885+0000"},{"body":"bq. It looks like there's a ton of effort put in to avoiding making sortByProximity work w/ non-local nodes\nBecause it's only when that local node is down that we actually need to sort...\nWhen/if DynamicEndpointSnitch's limitation is fixed (and it can sort by non-local nodes) then CassandraServer.java need not bypass it. But this won't simplify the code in CFRR. Now that CFIF supports multiple initialAddresses the method CFRR.sortEndpointsByProximity(..) can be rewritten (ie any connection to any initialAddress is all we need, no need to mess around with trying to connect through replica's to find information about replicas...)\nbq. Wait, why do we even care? \"local node\" IS the right host to sort against\nDepends on this is CFRR's \"local node\" or CassandraServer's \"local node\"... \nCFRR's local node is the right and only node worth sorting against, it being the \"task tracker node\". \nBut when c* on the \"task tracker node\" is down, then we randomly connect to another c* node so to find out of the replica we know about which are 1) up, 2) closest, and 3) in the same dc. Then it is a random c* node that becomes the \"local node\" and the call needs to be {{snitch.sortByProximity(initialAddress, addresses)}}.\nBut yes... the CFRR code is contorted. In many ways i prefer the simplicity of the first patch (both in api and in implementation) despite it not being \"as correct\". i thought of this \"fallback to replica\" as a last resort to keep the m/r job running, rather than an actively used feature where DynamicEndpointSnitch's scores will maximise performance. But then i'm only thinking in terms of a small c* cluster and i certainly am naive about what performance gains these scores can give...","created":"2011-06-28T14:13:20.615+0000"},{"body":"bq. CFRR's local node is the right and only node worth sorting against, it being the \"task tracker node\". \n\nRight.\n\nbq. Then it is a random c* node that becomes the \"local node\"\n\nWe still want to sort by proxmity-to-TT, because CFRR connects directly to the split owner to do the reads. initialAddress isn't involved post-split-discovery.\n\nAgain, all the complexity goes away if we just embed the snitch into CFIF/TT.\n\nOne wrinkle: ec2snitch requires gossip, so TT would need a separate local ip to participate in the gossip ring. We could make that optional (and fall back to old \"recognize local data, otherwise you get a random replica\" behavior otherwise).\n","created":"2011-06-28T14:51:07.348+0000"},{"body":"Taking a step back: aren't we optimizing for (1) a corner case with (2) the wrong solution?\n\nHere's what I mean:\n\n1) CFRR already prioritizes the local replica. So if you have >= one TT for each replica, this only helps if the local C* node dies, BUT the TT does not. This doesn't happen often.\n\n2) If we ARE in that situation, the \"right\" solution would be to send the job to a TT whose local replica IS live, not to read the data from a nonlocal replica. How can we signal that? ","created":"2011-06-28T15:01:23.257+0000"},{"body":"CASSANDRA-2388-addition1.patch: Simplify CFRR now that multiple initialAddresses are supported.","created":"2011-06-28T15:02:16.670+0000"},{"body":"bq. If we ARE in that situation, the \"right\" solution would be to send the job to a TT whose local replica IS live, not to read the data from a nonlocal replica. How can we signal that?\n\nISTM the right thing to do in that situation is just fail and let the JT reschedule somewhere else.","created":"2011-06-28T21:11:28.623+0000"},{"body":" - This does happen already (i've seen it while testing initial patches that were no good).\nProblem is that the TT is blacklisted, reducing hadoop's throughput for all jobs running.\nI bet too that a fallback to a replica is faster than a fallback to another TT.\n\n - There is no guarantee that any given TT will have its split accessible via a local c* node - this is only a preference in CFRR. A failed task may just as likely go to a random c* node. At least now we can actually properly limit to the one DC and sort by proximity. \n\n - One thing we're not doing here is applying this same DC limit and sort by proximity in the case when there isn't a localhost preference. See CFRR.initialize(..)\nIt would make sense to rewrite CFRR.getLocations(..) to\n{noformat} private Iterator getLocations(final Configuration conf) throws IOException\n {\n return new SplitEndpointIterator(conf);\n }{noformat} and then to move the finding-a-preference-to-localhost code into SplitEndpointIterator...\n\n - A bug i can see in the patch that did get accepted already is in CassandraServer.java:763 when endpointValid is false and restrictToSameDC is true we end up restricting to a random DC. I could fix this so restrictToSameDC is disabled in such situations but this actually invalidates the previous point: we can't restrict to DC anymore and we can only sortByProximity to a random node... I think this supports Jonathan's point that it's overall a poor approach. I'm more and more in preference of my original approach using just client.getDatacenter(..) and not worrying about proximity within the datacenter.\n\n - Another bug is that, contray to my patch, the code committed\nbq. committed with a change to use the dynamic snitch id the passed endpoint is valid.\n can call {{DynamicEndpointSnitch.sortByProximity(..)}} with an address that is not localhost and this breaks the assertion in the method. ","created":"2011-06-29T04:58:43.694+0000"},{"body":"{quote}\nThis does happen already (i've seen it while testing initial patches that were no good).\nProblem is that the TT is blacklisted, reducing hadoop's throughput for all jobs running.\n{quote}\n\nIf the cassandra node where the TT resides isn't working, then throughput is reduced regardless.\n\n\nbq. I bet too that a fallback to a replica is faster than a fallback to another TT.\n\nI doubt that for any significant job. Locality is important. Move the job to the data, not the data to the job.\n\n{quote}\nThere is no guarantee that any given TT will have its split accessible via a local c* node - this is only a preference in CFRR. A failed task may just as likely go to a random c* node. At least now we can actually properly limit to the one DC and sort by proximity.\n{quote}\n\nThis sounds like the thing we need to fix, then. Ensuring that the TT assigned to the map has a local replica.","created":"2011-06-29T18:40:40.810+0000"},{"body":"bq. If the cassandra node where the TT resides isn't working, then throughput is reduced regardless.\n\nRight: we _want_ it to be blacklisted in that scenario.","created":"2011-06-29T19:25:02.737+0000"},{"body":"bq. This sounds like the thing we need to fix, then. Ensuring that the TT assigned to the map has a local replica.\n\nreverted 1139358, 1139483 to make a fresh start for this.\n\nhow do we \"ensure\" this? isn't that the JT's job, to send jobs to the splits we gave it from CFIF? (which does make sure that only nodes with the data, are included in the split source list.)","created":"2011-06-29T19:45:21.179+0000"},{"body":"{quote}If the cassandra node where the TT resides isn't working, then throughput is reduced regardless.\nbq. Right: we want it to be blacklisted in that scenario.{quote}\nThis is making the presumption that the hadoop cluster is only used with CFIF.\nThe TT could still be useful for other jobs submitted.\nFurthermore a blacklisted TT does't automatically come back - it needs to be manually restarted. Isn't this creating more headache for operations?","created":"2011-06-29T19:49:28.321+0000"},{"body":"I dont think we should require the TT to be running locally. The whole idea is to support access to Cassandra data from hadoop even if it's just an import. \n\nThis patch does spend a lot of time dealing with non local data for that reason. ","created":"2011-06-29T19:59:09.009+0000"},{"body":"{quote}\nThis is making the presumption that the hadoop cluster is only used with CFIF.\nThe TT could still be useful for other jobs submitted.\n{quote}\n\nI'm fine with that assumption. If you want to run other jobs, use a different cluster. Cassandra's JVM is eating wasteful memory at that point.\n\n{quote}\nFurthermore a blacklisted TT does't automatically come back - it needs to be manually restarted. Isn't this creating more headache for operations?\n{quote}\n\nI don't think this is actually the case, see HADOOP-4305\n\n\n{quote}\nI dont think we should require the TT to be running locally. The whole idea is to support access to Cassandra data from hadoop even if it's just an import.\n\nThis patch does spend a lot of time dealing with non local data for that reason.\n{quote}\n\nI'm fine with dropping support for non-colocated TTs, or at least saying there's no DC-specific support. Because frankly, that is a very suboptimal thing to do, transfer the data across the network all the time, and flies in the face of Hadoop's core principles.","created":"2011-06-29T20:47:10.645+0000"},{"body":"bq. a blacklisted TT does't automatically come back\n\ntlipcon says it comes back after 24h, fwiw. In any case it's still the case that we DO want to blacklist it while it's down. (Brisk could perhaps add a \"clear my tasktracker on restart\" operation as a further enhancement.)\n\nbq. I'm fine with dropping support for non-colocated TTs\n\n+1, it was a bad idea and I'm sorry I wrote it. :)","created":"2011-06-29T21:00:55.453+0000"},{"body":"bq. tlipcon says it comes back after 24h\njust to be clear about my concerns. \nthis means a dead c* node will bring down a TT. In a hadoop cluster with 3 nodes this means for 24hrs you're lost 33% throughput. (If less than 10% of hadoop jobs used CFIF i could well imagine some pissed users). (What if you have a temporarily problem with flapping c* nodes and you end up with a handful of blacklisted TTs? etc etc etc).\n\nAll this when using a replica, any replica, could have kept things going smoothly, the only slowdown being some of the data into CFIF had to go over the network instead...\n","created":"2011-06-29T21:16:35.867+0000"},{"body":"bq. this means a dead c* node will bring down a TT\n\nAgain: _this is what you want to happen_. As long as the C* process on the same node is down, you want the TT to be blacklisted and the jobs to go elsewhere.\n\nbq. In a hadoop cluster with 3 nodes this means for 24hrs you're lost 33% throughput\n\nRight, but the real cause is because the C* process is dead, not b/c the TT is blacklisted. Making the TT read from other nodes will only hurt your network, not fix the throughput problem, b/c i/o is the bottleneck.","created":"2011-06-30T00:32:40.402+0000"},{"body":"Then i would hope for two separate InputFormats. One optimised for local node connection, where cassandra is deemed the more important system over hadoop, and another where data can be read in from anywhere. I think the latter should be supported in some manner since users may not always have the possibility to install hadoop and cassandra on the same servers, or they might not think it to be so critical part (eg if CFIF is reading using a IndexClause the input data set might be quite small and the remaining code in the m/r be the bulk of the processing...)","created":"2011-06-30T07:50:51.761+0000"},{"body":"bq. another where data can be read in from anywhere\n\nThis is totally antithetical to how hadoop is designed to work. I don't think it's worth supporting in-tree.","created":"2011-06-30T13:17:37.805+0000"},{"body":"Is CASSANDRA-2388-local-nodes-only-rough-sketch the direction we want then?\n\nThis is very initial code, i can't get {{new JobClient(JobTracker.getAddress(conf), conf).getClusterStatus().getActiveTrackerNames()}} to work, need a little help here.\n(Also CFRR.getLocations() can be drastically reduced).","created":"2011-06-30T21:43:26.156+0000"},{"body":"+1 to CFRR changes\n\nwasn't immediately clear to me what CFIF changes are doing, can you elaborate?","created":"2011-07-01T20:05:13.278+0000"},{"body":"The idea is to setup splits to have only endpoints that are valid trackers. But now i see this is just a brainfart :-) Ofc the jobTracker will apply this match for us. And that CFIF was always 'restricted' to running on endpoints. Although the documentation on inputSplit.getLocations() is a little thin as to whether this restricts which trackers it should run on or whether is just a preference... I guess it doesn't matter, as you point out Jonathan all that's required here is the one line changed in CFRR.\n\n","created":"2011-07-02T22:07:26.952+0000"},{"body":"the new \"one-liner\" CASSANDRA-2388 attached. i'll \"submit patch\" once i've tested it some...","created":"2011-07-02T22:16:38.381+0000"},{"body":"Sounds good, thanks!","created":"2011-07-02T22:53:36.840+0000"},{"body":"{quote}2) If we ARE in that situation, the \"right\" solution would be to send the job to a TT whose local replica IS live, not to read the data from a nonlocal replica. How can we signal that?{quote}To /really/ solve this issue can we do the following? \nIn CFIF.getRangeMap() take out of each range any endpoints that are not alive. A client connection already exists in this method. This filtering out of dead endpoints wouldn't be difficult, and would move tasks *to* the data making use of replica. This approach does need a new method in cassandra.thrift, eg {{list describe_alive_nodes()}}","created":"2011-07-04T06:39:00.633+0000"},{"body":"Does that really fix things though? Because you could have a data node be reachable from the coordinator answering describe_alive_nodes, but unreachable from the client. So the client still needs to be able to skip unreachable endpoints itself, so describe_alive seems like gratuitous complexity.","created":"2011-08-15T18:47:28.517+0000"},{"body":"I'd like to point out the situation in which no node for a given range of keys is available. It can happen for example with keyspace set to RF=1 and a node goes down. I created a patch that gives a user a chance to ignore missing range/node and continue runnig the MapReduce job. The patch is here: http://pastebin.com/hhrr8m9P\n\nJonathan already replied to the ML with \"ignoring unavailable ranges is a misfeature, imo\".\n\nIn our case it's very usefull, although there may be another/smarter solution. We have a keyspace with RF=1 and the nature of our data allows us to ignore temporarily missing node. The current ColumnFamilyInputFormat fails with RuntimeException and AFAIK there is no way around.","created":"2011-08-19T19:39:51.171+0000"},{"body":"bq. Does that really fix things though? Because you could have a data node be reachable from the coordinator answering describe_alive_nodes, but unreachable from the client. So the client still needs to be able to skip unreachable endpoints itself, so describe_alive seems like gratuitous complexity.\n\nI agree, since the view is from the coordinator, describe_alive_nodes isn't very helpful, and also has to wait on the failure detector to mark the node down anyway.","created":"2011-08-19T20:59:36.186+0000"},{"body":"bq. I agree, since the view is from the coordinator, describe_alive_nodes isn't very helpful\n\nCommitted Mck's most recent patch.\n\nbq. We have a keyspace with RF=1 and the nature of our data allows us to ignore temporarily missing node\n\nThe \"right\" fix is to increase RF. Ignoring missing data is not a scenario we want to support.","created":"2011-08-19T22:44:52.048+0000"},{"body":"This approach isn't really working for me and was committed too quickly i believe.\n\nbq. Although the documentation on inputSplit.getLocations() is a little thin as to whether this restricts which trackers it should run on or whether is just a preference\n\nTasks are still being evenly distributed around the ring regardless of what the ColumnFamilySplit.locations is.\n\nThe chance of a task actually working is RF/N. Therefore the chances of a blacklisted node are high. Worse is that the whole ring can quickly become blacklisted.\n\nhttp://abel-perez.com/hadoop-task-assignment has an interesting section in it explaining how the task assignment is supposed to work (and that data locality is preferred but not a requirement). Could ColumnFamilySplit.locations be in the wrong format? (eg they should ip not hostname?).","created":"2011-08-30T20:55:15.529+0000"},{"body":"see last comment. (say if this should be a separate bug...)\n\nMaybe hadoop's task allocation isn't working properly because i've an unbalanced ring (i'm working in parallel to fix that).\nIf this is the case i think it's an unfortunate limitation (the ring must be balanced to get any decent hadoop performance).\nIt's also probably likely when using {{ConfigHelper.setInputRange(..)}} that the number of nodes involved is small (approaching RF).\nWith the default hadoop scheduler your hadoop cluster is occupied while just a few taskTrackers are busy. Of course switching to FairScheduler will help some here.\n\nI'll take a look into hadoop's task allocation code as well...","created":"2011-08-31T16:23:23.789+0000"},{"body":"In the meantime could we make this behavior configurable.\neg replace CFRR:176 with something like\n{noformat}\n if(ConfigHelper.isDataLocalityDisabled())\n {\n return split.getLocations()[0];\n }\n else\n {\n throw new UnsupportedOperationException(\"no local connection available\");\n }{noformat}\n","created":"2011-09-08T06:01:39.491+0000"},{"body":"Should we just revert the change for now?","created":"2011-09-08T12:55:51.306+0000"},{"body":"Well that would work for me, was only thinking you want to push a \"default behavior\" (especially for those using a RP). \nBut I think a better understanding (at least from me) of hadoop's task scheduling is required before enforcing data locality, as as-is it certainly doesn't work for all.","created":"2011-09-08T21:36:39.552+0000"},{"body":"I just want to confirm what this ticket is about.\n\nThe JT has a list of endpoints for a given split.\nWhen a task runs it may or may not be on one of those nodes \nIf other tasks are running on all those replicas the JT may put them on a remote node.\n\nSo we need to decide which endpoint to connect to given the chance that nodes are down.\n\n1. Check if the node running CFRR is one of the replicas (we have this) this means JT has assigned a data-local task (good)\n2. If none of these nodes are local then pick another.\n3. If connection fails try the one other nodes.\n4. Try to avoid endpoints in a different DC.\n\nThe biggest problem is 4. Maybe the way todo this is change getSplits logic to never return replicas in another DC. I think this would require adding DC info to the describe_ring call. Then we only need to worry about 1-3.\n\n\n\n\n\n\n","created":"2011-09-14T02:25:54.953+0000"},{"body":"Marking as minor since the job should get re-submitted, and it's very difficult to reproduce when the tasktrackers are colocated with cassandra nodes (the recommended configuration).","created":"2012-03-08T22:38:35.989+0000"},{"body":"Would very much like a fix to this. We have a 40 node ring running 2x hadoop clusters on 20 nodes each. One cluster is on systems that are more flaky than the other (bad batch of memory). When building a split on the first cluster if a ring node is down in the area of the second cluster we get timeouts with no way to blacklist the offending node even though we have replicas local to the first cluster.\n\nThe ring is partitioned into DC1:2, DC2:2 with a hadoop cluster over each DC.","created":"2012-10-31T16:33:01.666+0000"},{"body":"Jake's plan above seems like a reasonable approach, but let me back up a step. I'm just not convinced that the problem we're trying to solve is a real one. Why do we want to suck a split's worth of data off-node? If it's because you don't have TackTrackers running on your Cassandra nodes, well, go fix that.\n\nIf it's because Hadoop has created too many tasks and all the local replicas have their task queue full, won't assigning it to a non-local TT just cause more contention, than waiting for a local slot to free up?","created":"2012-11-21T11:06:57.724+0000"},{"body":"I have two distinct use-cases where running TaskTrackers alongside Cassandra nodes does not accomplish our goals:\n\n1. Joining data. We have a large data set in cassandra, true, but we have a *much* larger data set held in Hadoop itself (around 4 orders of magnitude larger in hadoop than in cassandra). We need to join the two datasets together, and use the output from that join to feed multiple systems, none of which are cassandra. Since the data in Hadoop is so much larger than that in Cassandra, we have to bring the Cassandra data to hadoop, not the other way around. Because of security concerns, we can't spread our hadoop data onto our cassandra nodes (even if that didn't screw with our capacity planning), so we have no other choice but to move the Cassandra data (in small chunks) onto Hadoop. Why not use HBase, you say? We needed Cassandra for its write performance for other problems than this one. \n\n1. Offline, incremental backups. We have a large volume of time-series data held in Cassandra, and taking nightly snapshots and moving them to our archival center is prohibitively slow--it turns out that moving RF copies of our entire dataset over a leased line every night is a pretty bad idea. Instead, I use MapReduce to take an incremental backup of a much smaller subset of the data, then move that. That way, we not only are not moving the entire data set, but we are also using Cassandra's consistency mechanisms to resolve all the replicas. The only efficient way I've found to do this is via MapReduce (we use the Random Partitioner), and since it's an offline backup, we need to move it over the network anyway--may as well use the optimized network connecting Hadoop and Cassandra instead of the tiny pipe connecting cassandra to our archival center. \n\nBoth of these reasons dictate that we *not* run a TT alongside our Cassandra nodes, no matter what the *recommended* approach is. In this case, we need a strong, fault-tolerant CFIF to serve our purposes.\n\n","created":"2012-11-21T14:42:53.248+0000"},{"body":"Jonathan,\n I can't say i'm in favour of enforcing data locality.\nBecause data locality in hadoop doesn't work this way… when a tasktracker through the next heartbeat announces that it has a task slot free the jobtracker will do its best to assign a task with data locality to it but failing this will assign it a random task. the number of these random tasks can be quite high, just like i mentioned above\n{quote} Tasks are still being evenly distributed around the ring regardless of what the ColumnFamilySplit.locations is. {quote}\n\nThis can be almost solved by upgrading to hadoop-0.21+, using the fair scheduler and setting the property {code}\n mapred.fairscheduler.locality.delay\n 360000000\n{code}.\n\nAt the end of the day while hadoop encourages data locality it does not enforce it.\nThe ideal approach would be to sort all locations by proximity.\nThe feasible approach hopefully is still [~tjake]'s above. In addition i'd be in favour of a setting in the job's configuration as to whether a location from another datacenter can be used.\n\nreferences:\n - http://www.infoq.com/articles/HadoopInputFormat\n - http://www.mentby.com/matei-zaharia/running-only-node-local-jobs.html\n - https://groups.google.com/a/cloudera.org/forum/?fromgroups#!topic/cdh-user/3ggnE5hV0PY\n - http://www.cs.berkeley.edu/~matei/papers/2010/eurosys_delay_scheduling.pdf","created":"2013-05-18T21:50:22.969+0000"},{"body":"bq. The feasible approach hopefully is still T Jake Luciani's above\n\nOkay. Referring back to Jake's comments,\n\nbq. The biggest problem is [avoiding endpoints in a different DC]. Maybe the way todo this is change getSplits logic to never return replicas in another DC. I think this would require adding DC info to the describe_ring call\n\nI note that we expose node snitch location in system.peers. So at worst we could \"join\" against that manually.","created":"2013-05-27T16:02:01.381+0000"},{"body":"{quote}The biggest problem is [avoiding endpoints in a different DC]. Maybe the way todo this is change getSplits logic to never return replicas in another DC. I think this would require adding DC info to the describe_ring call{quote}\n\nTasktrackers may have access to a set of datacenters, so this DC info needs contain a list of DCs.\n\nFor example, our setup separates datacenters by physical datacenter and hadoop-usage, like:{noformat}DC1 \"Production + Hadoop\"\n c*01 c*03\nDC2 \"Production + Hadoop\"\n c*02 c*04\nDC3 \"Production\"\n c*05\nDC4 \"Production\"\n c*06{noformat}\n\nSo here we'd pass to getSplits() a DC info like \"DC1,DC2\".\nBut the problem remain, given a task executing on c*01 that fails to connect to localhost, although we can now prevent a connection to DC3 or DC4, we can't favour a connection to any other split in DC1 over anything in DC2. Is this solvable? ","created":"2013-05-27T18:41:31.744+0000"},{"body":"Attached [2.0-CASSANDRA-2388.patch|https://issues.apache.org/jira/secure/attachment/12649379/2.0-CASSANDRA-2388.patch] (against 2.0), where CASSANDRA-6302 is ported to ColumnFamilyRecordReader, so the next replicas are tried when the first one fails.\n\nReading from a non-local-DC replica is not a problem anymore due to the introduction of describe_local_ring (CASSANDRA-6268), that limits the input splits to the local DC.\n\nI would like to acknowledge my colleague Danilo Penna Queiroz who paired with me on this patch. ","created":"2014-06-09T14:19:48.838+0000"},{"body":"[~pkolaczk] to review","created":"2014-06-09T14:41:04.762+0000"},{"body":"[~pkolaczk] any update on this? cheers! :)","created":"2014-06-22T21:05:47.337+0000"},{"body":"[~pauloricardomg] I can't promise, but I try to do that at the end of this week. ","created":"2014-06-23T07:44:09.867+0000"},{"body":"Attaching patch based on 1.2.16 and fixed patch (v2) for 2.0.\n\nMaybe the 1.2 patch can still make it to 1.2.17...","created":"2014-06-25T19:22:36.783+0000"},{"body":"[~pkolaczk] any update on this? this review has been roaming for quite some time now...\nsorry for bothering but would be nice to see this integrated. cheers!","created":"2014-08-07T19:19:57.524+0000"},{"body":"http://mail-archives.apache.org/mod_mbox/cassandra-dev/201410.mbox/%3CCALdd-zjmvp7JOtguZ_k951RQHDtFt1cthX=RnHQ332C=gAZbjw@mail.gmail.com%3E","created":"2014-10-06T20:27:28.323+0000"},{"body":" [~pkolaczk] + [~pauloricardomg], AFAIK everything thrift related is frozen, so i presume the patch isn't going to be applied to master.\nOtherwise it's +1 on the patch from me.","created":"2014-10-10T11:16:23.350+0000"},{"body":"+1","created":"2014-12-01T14:02:24.337+0000"},{"body":"Paulo, can you rebase to 2.1?","created":"2015-08-24T16:18:46.847+0000"},{"body":"Rebased 2.1 patch available [here|https://github.com/pauloricardomg/cassandra/tree/2388-2.1].\n\n2.1 tests:\n* [testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2388-2.1-testall/lastCompletedBuild/testReport/]\n* [dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2388-2.1-dtest/lastCompletedBuild/testReport/]\n\n2.2 tests (don't know if {{ColumnFamilyInputFormat}} and {{ColumnFamilyRecordReader}} should be deprecated by then, but the classes are still present there):\n* [testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2388-2.2-testall/lastCompletedBuild/testReport/]\n* [dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2388-2.2-dtest/lastCompletedBuild/testReport/]","created":"2015-08-25T13:30:59.007+0000"},{"body":"Committed, thanks.","created":"2015-11-20T13:54:28.501+0000"}],"conversations":[{"body":"ColumnFamilyRecordReader only tries the first location for a given split. We should try multiple locations for a given split.","from":"reporter","subject":"ColumnFamilyRecordReader fails for a given split because a host is down, even if records could reasonably be read from other replica."},{"body":"I'm not sure special casing NoRouteToHostException to be blacklisted is the best thing to do. I don't think connections are being setup so often that maintaining a blacklist for any reason is needed.","from":"developer"},{"body":"Unfortunately, I thought of another problem here. If we go over the entire replica set, we're potentially going outside of the DC, which is bad since a lot of installations have a DC dedicated to analytics so it doesn't affect their app. It seems that the local address is preferred though, are your task trackers not on the same machines as Cassandra?","from":"developer"},{"body":"Eldon, are you planning to take another stab at this?","from":"developer"},{"body":"We need to return the list of replicas in the same DC","from":"developer"},{"body":"Mck, do you want to take a stab at this?","from":"developer"},{"body":"I'm having a go currently at CASSANDRA-1125 so i might as well look at this too. (but you've caught me on a holiday-week...)","from":"developer"},{"body":"How do i obtain the DataCenter name for a given address?\n\nIEndpointSnitch.getDataCenter(inetAddress) would work nicely for me but how do i get the snitch client-side?","from":"developer"},{"body":"Initial attempt at solution. Although I'm a little apprehensive to the additions to cassandra.thrift\n(describe_rack(..) isn't used anywhere, it just made sense to add describe_datacenter(..) and describe_rack(..) at the same time).\n\nI've tested that existing hadoop jobs work but the new functionality hasn't been tested (as i currently don't have any RF=2 data setup).\n\nThis patch does not include the required re-generated Cassandra.java","from":"developer"},{"body":"Second attempt. (god only knows what i was trying to test last patch ;)\nthis patch:\n - adds describe_datacenter and describe_rack to cassandra.thrift\n - adds locations in ColumnFamilyRecordReader from the split's alternative endpoints if dc is the same\n\nThis patch does not include the required re-generated Cassandra.java\n","from":"developer"},{"body":"I have tested this now on data w/ RF=2.\nSeems to work ~ok as far as i can see.\n\nOne side-effect of this patch is where once one could configure ConfigHelper.setInitialAddress(conf, \"localhost\") this will no longer work for tasks trying to run on the down node.\nColumnFamilyRecordReader.getLocations() will ConnectException trying to call describe_datacenter(..). This will lead to the task failing. Hadoop re-runs the task then on another node and eventually the job will complete. But the fall back to replica never is used.\n\nIf the initialAddress is hardcoded to one node then we no longer have a decentralised job.\n\nI would like to allow a comma-separated in initialAddress, for example it could be \"localhost, node01, node02, node03\". This would give preference to localhost and avoid any centralisation.\n\nI would also like to make ColumnFamilyRecordReader.getLocations() return an iterator instead of an array.\nThe createConnection(..) and client.describe_datacenter(..) calls are an unnecessary overhead when all nodes (or first endpoint location) are up, and could be avoided by lazy-loading the list.","from":"developer"},{"body":"New patch. I think i'm at last happy with it.\n\ngetLocations() returns an iterator so client.describe_datacenter() is only called when necessary.\n\nRather than provide a list in initialAddress it was possible to use either the initialAddress OR the endpoint. This gave the benefit in not listing a location that can't actually be connected to.\n\nThe \"only use replica from same DC\" is an option now in ConfigHelper. By default it is true.\n\nAgain the re-generated Cassandra.java is not included in the patch.\n\nI have tested this on normal jobs, and RF=2 jobs with a node down.","from":"developer"},{"body":"The get_rack seems unused so it should be removed.\n\nAlso, it might be better to pass all locations in the get_datacenter thrift call since you can get the results in one shot and sort them by the dynamic snitch, filtering out the dead nodes:\n\n{noformat}\n DatabaseDescriptor.getEndpointSnitch().sortByProximity(FBUtilities.getLocalAddress(), endpoints); \n{noformat}\n\n{noformat}\n FailureDetector.instance.isAlive(endpoint)\n{noformat}\n\n","from":"developer"},{"body":"Then (if i understand you correctly) i would need in cassandra.thrift\n{noformat}\n /** returns alive endpoints, sorted by proximity, that belong in the same datacenter as the given endpoint */\n list get_endpoints_in_same_datacenter(1: string endpoint, 2: required list endpoints)\n throws (1:InvalidRequestException ire)\n{noformat}\n\nThen the API becomes quite specific to this usecase. Is the performance gain worth it? What's the cost of each client.describe_datacenter(..) call, and probably more important what is the lost performance of writing to the furthest node that's within the same datacenter?","from":"developer"},{"body":"Just make sure i understand you T Jake, you would rather something like this in CassandraServer.java?\n(I've renamed from the previous comment get_endpoints_in_same_datacenter(..) to sort_endpoints_by_proximity(..))\n{noformat}\n public String[] sort_endpoints_by_proximity(String endpoint, String[] endpoints, boolean restrictToSameDC) \n throws TException, InvalidRequestException\n {\n try\n {\n List results = new ArrayList();\n InetAddress address = InetAddress.getByName(endpoint);\n String datacenter = DatabaseDescriptor.getEndpointSnitch().getDatacenter(address);\n List addresses = new ArrayList();\n for(String ep : endpoints)\n {\n addresses.add(InetAddress.getByName(ep));\n }\n DatabaseDescriptor.getEndpointSnitch().sortByProximity(address, addresses);\n for(InetAddress ep : addresses)\n {\n String dc = DatabaseDescriptor.getEndpointSnitch().getDatacenter(ep);\n if(FailureDetector.instance.isAlive(ep) && (!restrictToSameDC || datacenter.equals(dc)))\n {\n results.add(ep.getHostName());\n }\n }\n return results.toArray(new String[results.size()]);\n }\n catch (UnknownHostException e)\n {\n throw new InvalidRequestException(e.getMessage());\n }\n }\n{noformat}","from":"developer"},{"body":"bq. what is the lost performance of writing to the furthest node that's within the same datacenter?\n\nThe benefit is really the DynamicSnitch. if a node it slow due to compaction then this would avoid sending requests there... \n\nbq. public String[] sort_endpoints_by_proximity(String endpoint, String[] endpoints, boolean restrictToSameDC)\n\nI don't think it makes sense to send the client endpoint to this call since the endpoint might not be a cassandra node. It's a reasonable assumption that the endpoint it's talking to is local enough to the client to use that.\n\n","from":"developer"},{"body":"{quote}\nbq. public String[] sort_endpoints_by_proximity(String endpoint, String[] endpoints, boolean restrictToSameDC)\nI don't think it makes sense to send the client endpoint to this call since the endpoint might not be a cassandra node. It's a reasonable assumption that the endpoint it's talking to is local enough to the client to use that.\n{quote}\nFor the test set i was running against, RF=2, each split's has two endpoints always in different datacenters.\n\nIf the \"local\" endpoint is down then getLocations() will then call client.sort_endpoints_by_proximity(..) and this will fail (being the same endpoint).\nIt then makes a client connection through the \"other\" endpoint. \\[see CFRR.describeDatacenter(..)].\nThis will presume the wrong datacenter and return itself as a valid endpoint. \nI need some way to know what the original datacenter is, even when it is down.","from":"developer"},{"body":"ok but why not change the response to map> where key is DC and value are proximity sorted endpoints?","from":"developer"},{"body":"Won't the sorting still be wrong?\nFor the use-case above it will solve restricting to the correct datacenter, but the sorting will still be based on proximity to the wrong node?\n\nbq. I don't think it makes sense to send the client endpoint to this call since the endpoint might not be a cassandra node. \nIt might not be an alive cassandra node, but it should be a cassandra node. It comes from the split's list of endpoints. At least in this use-case, or are you referring to general usage for this new api?\nbq. It's a reasonable assumption that the endpoint it's talking to is local enough to the client to use that.\nI don't think so... The endpoint that it talks to is a completely random (just the next endpoint listed in the split's list). This is why i think that such sorting won't just be wrong but not even close. Does this make sense?","from":"developer"},{"body":"I think the core issue is you can't assume the hadoop node is running on a cassandra node...\n\nIf it is then the logic is straight forward, if not then it's possible the connection could cross DC boundaries. One possibility is to use the ip octets like the RackInferringSnitch. \n\nHow's this proposal then? keep the sort_endpoints_by_proximity signature as is and pass the client endpoint along with the list of data endpoints and add the following logic:\n\n1) sort the endpoints using the endpoint_snitch.\n2) if client endpoint *is* a valid cassandra node get the nodes DC and prune nodes outside of this DC\n3) if client endpoint *is not* a valid cassandra node try to infer the DC from its ip and prune dataendpoint nodes in a different DC. If no cassandra nodes are in the DC list goto 3).\n4) all else fails return the sorted endpoint list\n","from":"developer"},{"body":"bq. [snip] One possibility is to use the ip octets like the RackInferringSnitch. \n\nIn our usecase we have three nodes defined via PropertyFileSnitch:{noformat}152.90.241.22=DC1:RAC1 #node1\n152.90.241.23=DC2:RAC1 #node2\n152.90.241.24=DC1:RAC1 #node3{noformat}\nThe only way to infer here is even addresses belong to one dc, odd to the other. This is not how RackInferringSnithc works.\n\nWhen we make the connection through the \"other\" (node2) endpoint taking the rack inferring approach \"152.90.\" will say it's in DC2. (again) this is the wrong DC and will return itself as a valid endpoint....\n\nStep (3) seems to me to be too specific to be included here.\nIf i go only with steps (1),(2),and (4) we get this code:{noformat} public String[] sort_endpoints_by_proximity(String endpoint, String[] endpoints, boolean restrictToSameDC) \n throws TException, InvalidRequestException\n {\n try\n {\n List results = new ArrayList();\n InetAddress address = InetAddress.getByName(endpoint);\n boolean endpointValid = null != Gossiper.instance.getEndpointStateForEndpoint(address);\n String datacenter = DatabaseDescriptor\n .getEndpointSnitch().getDatacenter(endpointValid ? address : FBUtilities.getLocalAddress());\n List addresses = new ArrayList();\n for(String ep : endpoints)\n {\n addresses.add(InetAddress.getByName(endpoint));\n }\n DatabaseDescriptor.getEndpointSnitch().sortByProximity(address, addresses);\n for(InetAddress ep : addresses)\n {\n String dc = DatabaseDescriptor.getEndpointSnitch().getDatacenter(ep);\n if(FailureDetector.instance.isAlive(ep) && (!restrictToSameDC || datacenter.equals(dc)))\n {\n results.add(ep.getHostName());\n }\n }\n return results.toArray(new String[results.size()]);\n }\n catch (UnknownHostException e)\n {\n throw new InvalidRequestException(e.getMessage());\n }\n }{noformat}\n\nI'm happy with this (except that {{Gossiper.instance.getEndpointStateForEndpoint(address)}} is only my guess on how to tell if an endpoint is valid as such).","from":"developer"},{"body":"Problem with the suggested approach is that sortByProximity(..) *only* works when address is the local address. See assert statement DynamicEndpointSnitch:134\n\nI could hack this and rewrite the line to\n{noformat}IEndpointSnitch snitch = DatabaseDescriptor.getEndpointSnitch();\nsnitch = snitch instanceof DynamicEndpointSnitch ? ((DynamicEndpointSnitch)snitch).subsnitch : snitch;\nsnitch.sortByProximity(address, addresses);{noformat}\nBut this of course means that we always bypass DynamicEndpointSnitch's \"scores\".","from":"developer"},{"body":"Up to date patch.\nFollows T Jake's points (1),(2), and (4).\nAnd bypasses DynamicEndpointSnitch when sorting by proximity.","from":"developer"},{"body":"committed with a change to use the dynamic snitch id the passed endpoint is valid.","from":"developer"},{"body":"This patch applies to the current 0.7-branch with minimal problems - just some imports on CassandraServer that it couldn't resolve properly. Can this be committed against 0.7-branch for inclusion in 0.7.7?","from":"developer"},{"body":"I've done basic testing with the word count and pig examples to make sure that the basic hadoop integration isn't negatively affected by this. I'll also try it against our dev cluster before and after the patch - killing one node to see if it fails over to another replica - to make sure it does what it should that way.","from":"developer"},{"body":"Reopening for testing against 0.7.6.","from":"developer"},{"body":"Took a look at this belatedly. I don't understand the contortions at all. It looks like there's a ton of effort put in to avoiding making sortByProximity work w/ non-local nodes. Why not just make that work instead?\n","from":"developer"},{"body":"also: running hadoop on a non-cassandra node is dumb. i don't see a point in supporting that really. (yes, my fault it was written that way to begin with, mea culpa.)","from":"developer"},{"body":"Jonathan - is it possible to attach an updated patch based on your changes to 0.8 branch? Not sure if that would be simple to extract.","from":"developer"},{"body":"with svn: svn diff -r 1139323:1139483 and hack out the OutboundTcpConnection change in the middle manually from the output.\n\nwith git: create a branch, rip out the offending OTC change and squash the other two","from":"developer"},{"body":"I think there's deep surgery to be done here still though. Backporting is probably premature.","from":"developer"},{"body":"bq. It looks like there's a ton of effort put in to avoiding making sortByProximity work w/ non-local nodes\n\nWait, why do we even care? \"local node\" IS the right host to sort against -- we want the split that is closest to the node running the job, this is not the same as some other C* node we contact.","from":"developer"},{"body":"bq. It looks like there's a ton of effort put in to avoiding making sortByProximity work w/ non-local nodes\nBecause it's only when that local node is down that we actually need to sort...\nWhen/if DynamicEndpointSnitch's limitation is fixed (and it can sort by non-local nodes) then CassandraServer.java need not bypass it. But this won't simplify the code in CFRR. Now that CFIF supports multiple initialAddresses the method CFRR.sortEndpointsByProximity(..) can be rewritten (ie any connection to any initialAddress is all we need, no need to mess around with trying to connect through replica's to find information about replicas...)\nbq. Wait, why do we even care? \"local node\" IS the right host to sort against\nDepends on this is CFRR's \"local node\" or CassandraServer's \"local node\"... \nCFRR's local node is the right and only node worth sorting against, it being the \"task tracker node\". \nBut when c* on the \"task tracker node\" is down, then we randomly connect to another c* node so to find out of the replica we know about which are 1) up, 2) closest, and 3) in the same dc. Then it is a random c* node that becomes the \"local node\" and the call needs to be {{snitch.sortByProximity(initialAddress, addresses)}}.\nBut yes... the CFRR code is contorted. In many ways i prefer the simplicity of the first patch (both in api and in implementation) despite it not being \"as correct\". i thought of this \"fallback to replica\" as a last resort to keep the m/r job running, rather than an actively used feature where DynamicEndpointSnitch's scores will maximise performance. But then i'm only thinking in terms of a small c* cluster and i certainly am naive about what performance gains these scores can give...","from":"developer"},{"body":"bq. CFRR's local node is the right and only node worth sorting against, it being the \"task tracker node\". \n\nRight.\n\nbq. Then it is a random c* node that becomes the \"local node\"\n\nWe still want to sort by proxmity-to-TT, because CFRR connects directly to the split owner to do the reads. initialAddress isn't involved post-split-discovery.\n\nAgain, all the complexity goes away if we just embed the snitch into CFIF/TT.\n\nOne wrinkle: ec2snitch requires gossip, so TT would need a separate local ip to participate in the gossip ring. We could make that optional (and fall back to old \"recognize local data, otherwise you get a random replica\" behavior otherwise).\n","from":"developer"},{"body":"Taking a step back: aren't we optimizing for (1) a corner case with (2) the wrong solution?\n\nHere's what I mean:\n\n1) CFRR already prioritizes the local replica. So if you have >= one TT for each replica, this only helps if the local C* node dies, BUT the TT does not. This doesn't happen often.\n\n2) If we ARE in that situation, the \"right\" solution would be to send the job to a TT whose local replica IS live, not to read the data from a nonlocal replica. How can we signal that? ","from":"developer"},{"body":"CASSANDRA-2388-addition1.patch: Simplify CFRR now that multiple initialAddresses are supported.","from":"developer"},{"body":"bq. If we ARE in that situation, the \"right\" solution would be to send the job to a TT whose local replica IS live, not to read the data from a nonlocal replica. How can we signal that?\n\nISTM the right thing to do in that situation is just fail and let the JT reschedule somewhere else.","from":"developer"},{"body":" - This does happen already (i've seen it while testing initial patches that were no good).\nProblem is that the TT is blacklisted, reducing hadoop's throughput for all jobs running.\nI bet too that a fallback to a replica is faster than a fallback to another TT.\n\n - There is no guarantee that any given TT will have its split accessible via a local c* node - this is only a preference in CFRR. A failed task may just as likely go to a random c* node. At least now we can actually properly limit to the one DC and sort by proximity. \n\n - One thing we're not doing here is applying this same DC limit and sort by proximity in the case when there isn't a localhost preference. See CFRR.initialize(..)\nIt would make sense to rewrite CFRR.getLocations(..) to\n{noformat} private Iterator getLocations(final Configuration conf) throws IOException\n {\n return new SplitEndpointIterator(conf);\n }{noformat} and then to move the finding-a-preference-to-localhost code into SplitEndpointIterator...\n\n - A bug i can see in the patch that did get accepted already is in CassandraServer.java:763 when endpointValid is false and restrictToSameDC is true we end up restricting to a random DC. I could fix this so restrictToSameDC is disabled in such situations but this actually invalidates the previous point: we can't restrict to DC anymore and we can only sortByProximity to a random node... I think this supports Jonathan's point that it's overall a poor approach. I'm more and more in preference of my original approach using just client.getDatacenter(..) and not worrying about proximity within the datacenter.\n\n - Another bug is that, contray to my patch, the code committed\nbq. committed with a change to use the dynamic snitch id the passed endpoint is valid.\n can call {{DynamicEndpointSnitch.sortByProximity(..)}} with an address that is not localhost and this breaks the assertion in the method. ","from":"developer"},{"body":"{quote}\nThis does happen already (i've seen it while testing initial patches that were no good).\nProblem is that the TT is blacklisted, reducing hadoop's throughput for all jobs running.\n{quote}\n\nIf the cassandra node where the TT resides isn't working, then throughput is reduced regardless.\n\n\nbq. I bet too that a fallback to a replica is faster than a fallback to another TT.\n\nI doubt that for any significant job. Locality is important. Move the job to the data, not the data to the job.\n\n{quote}\nThere is no guarantee that any given TT will have its split accessible via a local c* node - this is only a preference in CFRR. A failed task may just as likely go to a random c* node. At least now we can actually properly limit to the one DC and sort by proximity.\n{quote}\n\nThis sounds like the thing we need to fix, then. Ensuring that the TT assigned to the map has a local replica.","from":"developer"},{"body":"bq. If the cassandra node where the TT resides isn't working, then throughput is reduced regardless.\n\nRight: we _want_ it to be blacklisted in that scenario.","from":"developer"},{"body":"bq. This sounds like the thing we need to fix, then. Ensuring that the TT assigned to the map has a local replica.\n\nreverted 1139358, 1139483 to make a fresh start for this.\n\nhow do we \"ensure\" this? isn't that the JT's job, to send jobs to the splits we gave it from CFIF? (which does make sure that only nodes with the data, are included in the split source list.)","from":"developer"},{"body":"{quote}If the cassandra node where the TT resides isn't working, then throughput is reduced regardless.\nbq. Right: we want it to be blacklisted in that scenario.{quote}\nThis is making the presumption that the hadoop cluster is only used with CFIF.\nThe TT could still be useful for other jobs submitted.\nFurthermore a blacklisted TT does't automatically come back - it needs to be manually restarted. Isn't this creating more headache for operations?","from":"developer"},{"body":"I dont think we should require the TT to be running locally. The whole idea is to support access to Cassandra data from hadoop even if it's just an import. \n\nThis patch does spend a lot of time dealing with non local data for that reason. ","from":"developer"},{"body":"{quote}\nThis is making the presumption that the hadoop cluster is only used with CFIF.\nThe TT could still be useful for other jobs submitted.\n{quote}\n\nI'm fine with that assumption. If you want to run other jobs, use a different cluster. Cassandra's JVM is eating wasteful memory at that point.\n\n{quote}\nFurthermore a blacklisted TT does't automatically come back - it needs to be manually restarted. Isn't this creating more headache for operations?\n{quote}\n\nI don't think this is actually the case, see HADOOP-4305\n\n\n{quote}\nI dont think we should require the TT to be running locally. The whole idea is to support access to Cassandra data from hadoop even if it's just an import.\n\nThis patch does spend a lot of time dealing with non local data for that reason.\n{quote}\n\nI'm fine with dropping support for non-colocated TTs, or at least saying there's no DC-specific support. Because frankly, that is a very suboptimal thing to do, transfer the data across the network all the time, and flies in the face of Hadoop's core principles.","from":"developer"},{"body":"bq. a blacklisted TT does't automatically come back\n\ntlipcon says it comes back after 24h, fwiw. In any case it's still the case that we DO want to blacklist it while it's down. (Brisk could perhaps add a \"clear my tasktracker on restart\" operation as a further enhancement.)\n\nbq. I'm fine with dropping support for non-colocated TTs\n\n+1, it was a bad idea and I'm sorry I wrote it. :)","from":"developer"},{"body":"bq. tlipcon says it comes back after 24h\njust to be clear about my concerns. \nthis means a dead c* node will bring down a TT. In a hadoop cluster with 3 nodes this means for 24hrs you're lost 33% throughput. (If less than 10% of hadoop jobs used CFIF i could well imagine some pissed users). (What if you have a temporarily problem with flapping c* nodes and you end up with a handful of blacklisted TTs? etc etc etc).\n\nAll this when using a replica, any replica, could have kept things going smoothly, the only slowdown being some of the data into CFIF had to go over the network instead...\n","from":"developer"},{"body":"bq. this means a dead c* node will bring down a TT\n\nAgain: _this is what you want to happen_. As long as the C* process on the same node is down, you want the TT to be blacklisted and the jobs to go elsewhere.\n\nbq. In a hadoop cluster with 3 nodes this means for 24hrs you're lost 33% throughput\n\nRight, but the real cause is because the C* process is dead, not b/c the TT is blacklisted. Making the TT read from other nodes will only hurt your network, not fix the throughput problem, b/c i/o is the bottleneck.","from":"developer"},{"body":"Then i would hope for two separate InputFormats. One optimised for local node connection, where cassandra is deemed the more important system over hadoop, and another where data can be read in from anywhere. I think the latter should be supported in some manner since users may not always have the possibility to install hadoop and cassandra on the same servers, or they might not think it to be so critical part (eg if CFIF is reading using a IndexClause the input data set might be quite small and the remaining code in the m/r be the bulk of the processing...)","from":"developer"},{"body":"bq. another where data can be read in from anywhere\n\nThis is totally antithetical to how hadoop is designed to work. I don't think it's worth supporting in-tree.","from":"developer"},{"body":"Is CASSANDRA-2388-local-nodes-only-rough-sketch the direction we want then?\n\nThis is very initial code, i can't get {{new JobClient(JobTracker.getAddress(conf), conf).getClusterStatus().getActiveTrackerNames()}} to work, need a little help here.\n(Also CFRR.getLocations() can be drastically reduced).","from":"developer"},{"body":"+1 to CFRR changes\n\nwasn't immediately clear to me what CFIF changes are doing, can you elaborate?","from":"developer"},{"body":"The idea is to setup splits to have only endpoints that are valid trackers. But now i see this is just a brainfart :-) Ofc the jobTracker will apply this match for us. And that CFIF was always 'restricted' to running on endpoints. Although the documentation on inputSplit.getLocations() is a little thin as to whether this restricts which trackers it should run on or whether is just a preference... I guess it doesn't matter, as you point out Jonathan all that's required here is the one line changed in CFRR.\n\n","from":"developer"},{"body":"the new \"one-liner\" CASSANDRA-2388 attached. i'll \"submit patch\" once i've tested it some...","from":"developer"},{"body":"Sounds good, thanks!","from":"developer"},{"body":"{quote}2) If we ARE in that situation, the \"right\" solution would be to send the job to a TT whose local replica IS live, not to read the data from a nonlocal replica. How can we signal that?{quote}To /really/ solve this issue can we do the following? \nIn CFIF.getRangeMap() take out of each range any endpoints that are not alive. A client connection already exists in this method. This filtering out of dead endpoints wouldn't be difficult, and would move tasks *to* the data making use of replica. This approach does need a new method in cassandra.thrift, eg {{list describe_alive_nodes()}}","from":"developer"},{"body":"Does that really fix things though? Because you could have a data node be reachable from the coordinator answering describe_alive_nodes, but unreachable from the client. So the client still needs to be able to skip unreachable endpoints itself, so describe_alive seems like gratuitous complexity.","from":"developer"},{"body":"I'd like to point out the situation in which no node for a given range of keys is available. It can happen for example with keyspace set to RF=1 and a node goes down. I created a patch that gives a user a chance to ignore missing range/node and continue runnig the MapReduce job. The patch is here: http://pastebin.com/hhrr8m9P\n\nJonathan already replied to the ML with \"ignoring unavailable ranges is a misfeature, imo\".\n\nIn our case it's very usefull, although there may be another/smarter solution. We have a keyspace with RF=1 and the nature of our data allows us to ignore temporarily missing node. The current ColumnFamilyInputFormat fails with RuntimeException and AFAIK there is no way around.","from":"developer"},{"body":"bq. Does that really fix things though? Because you could have a data node be reachable from the coordinator answering describe_alive_nodes, but unreachable from the client. So the client still needs to be able to skip unreachable endpoints itself, so describe_alive seems like gratuitous complexity.\n\nI agree, since the view is from the coordinator, describe_alive_nodes isn't very helpful, and also has to wait on the failure detector to mark the node down anyway.","from":"developer"},{"body":"bq. I agree, since the view is from the coordinator, describe_alive_nodes isn't very helpful\n\nCommitted Mck's most recent patch.\n\nbq. We have a keyspace with RF=1 and the nature of our data allows us to ignore temporarily missing node\n\nThe \"right\" fix is to increase RF. Ignoring missing data is not a scenario we want to support.","from":"developer"},{"body":"This approach isn't really working for me and was committed too quickly i believe.\n\nbq. Although the documentation on inputSplit.getLocations() is a little thin as to whether this restricts which trackers it should run on or whether is just a preference\n\nTasks are still being evenly distributed around the ring regardless of what the ColumnFamilySplit.locations is.\n\nThe chance of a task actually working is RF/N. Therefore the chances of a blacklisted node are high. Worse is that the whole ring can quickly become blacklisted.\n\nhttp://abel-perez.com/hadoop-task-assignment has an interesting section in it explaining how the task assignment is supposed to work (and that data locality is preferred but not a requirement). Could ColumnFamilySplit.locations be in the wrong format? (eg they should ip not hostname?).","from":"developer"},{"body":"see last comment. (say if this should be a separate bug...)\n\nMaybe hadoop's task allocation isn't working properly because i've an unbalanced ring (i'm working in parallel to fix that).\nIf this is the case i think it's an unfortunate limitation (the ring must be balanced to get any decent hadoop performance).\nIt's also probably likely when using {{ConfigHelper.setInputRange(..)}} that the number of nodes involved is small (approaching RF).\nWith the default hadoop scheduler your hadoop cluster is occupied while just a few taskTrackers are busy. Of course switching to FairScheduler will help some here.\n\nI'll take a look into hadoop's task allocation code as well...","from":"developer"},{"body":"In the meantime could we make this behavior configurable.\neg replace CFRR:176 with something like\n{noformat}\n if(ConfigHelper.isDataLocalityDisabled())\n {\n return split.getLocations()[0];\n }\n else\n {\n throw new UnsupportedOperationException(\"no local connection available\");\n }{noformat}\n","from":"developer"},{"body":"Should we just revert the change for now?","from":"developer"},{"body":"Well that would work for me, was only thinking you want to push a \"default behavior\" (especially for those using a RP). \nBut I think a better understanding (at least from me) of hadoop's task scheduling is required before enforcing data locality, as as-is it certainly doesn't work for all.","from":"developer"},{"body":"I just want to confirm what this ticket is about.\n\nThe JT has a list of endpoints for a given split.\nWhen a task runs it may or may not be on one of those nodes \nIf other tasks are running on all those replicas the JT may put them on a remote node.\n\nSo we need to decide which endpoint to connect to given the chance that nodes are down.\n\n1. Check if the node running CFRR is one of the replicas (we have this) this means JT has assigned a data-local task (good)\n2. If none of these nodes are local then pick another.\n3. If connection fails try the one other nodes.\n4. Try to avoid endpoints in a different DC.\n\nThe biggest problem is 4. Maybe the way todo this is change getSplits logic to never return replicas in another DC. I think this would require adding DC info to the describe_ring call. Then we only need to worry about 1-3.\n\n\n\n\n\n\n","from":"developer"},{"body":"Marking as minor since the job should get re-submitted, and it's very difficult to reproduce when the tasktrackers are colocated with cassandra nodes (the recommended configuration).","from":"developer"},{"body":"Would very much like a fix to this. We have a 40 node ring running 2x hadoop clusters on 20 nodes each. One cluster is on systems that are more flaky than the other (bad batch of memory). When building a split on the first cluster if a ring node is down in the area of the second cluster we get timeouts with no way to blacklist the offending node even though we have replicas local to the first cluster.\n\nThe ring is partitioned into DC1:2, DC2:2 with a hadoop cluster over each DC.","from":"developer"},{"body":"Jake's plan above seems like a reasonable approach, but let me back up a step. I'm just not convinced that the problem we're trying to solve is a real one. Why do we want to suck a split's worth of data off-node? If it's because you don't have TackTrackers running on your Cassandra nodes, well, go fix that.\n\nIf it's because Hadoop has created too many tasks and all the local replicas have their task queue full, won't assigning it to a non-local TT just cause more contention, than waiting for a local slot to free up?","from":"developer"},{"body":"I have two distinct use-cases where running TaskTrackers alongside Cassandra nodes does not accomplish our goals:\n\n1. Joining data. We have a large data set in cassandra, true, but we have a *much* larger data set held in Hadoop itself (around 4 orders of magnitude larger in hadoop than in cassandra). We need to join the two datasets together, and use the output from that join to feed multiple systems, none of which are cassandra. Since the data in Hadoop is so much larger than that in Cassandra, we have to bring the Cassandra data to hadoop, not the other way around. Because of security concerns, we can't spread our hadoop data onto our cassandra nodes (even if that didn't screw with our capacity planning), so we have no other choice but to move the Cassandra data (in small chunks) onto Hadoop. Why not use HBase, you say? We needed Cassandra for its write performance for other problems than this one. \n\n1. Offline, incremental backups. We have a large volume of time-series data held in Cassandra, and taking nightly snapshots and moving them to our archival center is prohibitively slow--it turns out that moving RF copies of our entire dataset over a leased line every night is a pretty bad idea. Instead, I use MapReduce to take an incremental backup of a much smaller subset of the data, then move that. That way, we not only are not moving the entire data set, but we are also using Cassandra's consistency mechanisms to resolve all the replicas. The only efficient way I've found to do this is via MapReduce (we use the Random Partitioner), and since it's an offline backup, we need to move it over the network anyway--may as well use the optimized network connecting Hadoop and Cassandra instead of the tiny pipe connecting cassandra to our archival center. \n\nBoth of these reasons dictate that we *not* run a TT alongside our Cassandra nodes, no matter what the *recommended* approach is. In this case, we need a strong, fault-tolerant CFIF to serve our purposes.\n\n","from":"developer"},{"body":"Jonathan,\n I can't say i'm in favour of enforcing data locality.\nBecause data locality in hadoop doesn't work this way… when a tasktracker through the next heartbeat announces that it has a task slot free the jobtracker will do its best to assign a task with data locality to it but failing this will assign it a random task. the number of these random tasks can be quite high, just like i mentioned above\n{quote} Tasks are still being evenly distributed around the ring regardless of what the ColumnFamilySplit.locations is. {quote}\n\nThis can be almost solved by upgrading to hadoop-0.21+, using the fair scheduler and setting the property {code}\n mapred.fairscheduler.locality.delay\n 360000000\n{code}.\n\nAt the end of the day while hadoop encourages data locality it does not enforce it.\nThe ideal approach would be to sort all locations by proximity.\nThe feasible approach hopefully is still [~tjake]'s above. In addition i'd be in favour of a setting in the job's configuration as to whether a location from another datacenter can be used.\n\nreferences:\n - http://www.infoq.com/articles/HadoopInputFormat\n - http://www.mentby.com/matei-zaharia/running-only-node-local-jobs.html\n - https://groups.google.com/a/cloudera.org/forum/?fromgroups#!topic/cdh-user/3ggnE5hV0PY\n - http://www.cs.berkeley.edu/~matei/papers/2010/eurosys_delay_scheduling.pdf","from":"developer"},{"body":"bq. The feasible approach hopefully is still T Jake Luciani's above\n\nOkay. Referring back to Jake's comments,\n\nbq. The biggest problem is [avoiding endpoints in a different DC]. Maybe the way todo this is change getSplits logic to never return replicas in another DC. I think this would require adding DC info to the describe_ring call\n\nI note that we expose node snitch location in system.peers. So at worst we could \"join\" against that manually.","from":"developer"},{"body":"{quote}The biggest problem is [avoiding endpoints in a different DC]. Maybe the way todo this is change getSplits logic to never return replicas in another DC. I think this would require adding DC info to the describe_ring call{quote}\n\nTasktrackers may have access to a set of datacenters, so this DC info needs contain a list of DCs.\n\nFor example, our setup separates datacenters by physical datacenter and hadoop-usage, like:{noformat}DC1 \"Production + Hadoop\"\n c*01 c*03\nDC2 \"Production + Hadoop\"\n c*02 c*04\nDC3 \"Production\"\n c*05\nDC4 \"Production\"\n c*06{noformat}\n\nSo here we'd pass to getSplits() a DC info like \"DC1,DC2\".\nBut the problem remain, given a task executing on c*01 that fails to connect to localhost, although we can now prevent a connection to DC3 or DC4, we can't favour a connection to any other split in DC1 over anything in DC2. Is this solvable? ","from":"developer"},{"body":"Attached [2.0-CASSANDRA-2388.patch|https://issues.apache.org/jira/secure/attachment/12649379/2.0-CASSANDRA-2388.patch] (against 2.0), where CASSANDRA-6302 is ported to ColumnFamilyRecordReader, so the next replicas are tried when the first one fails.\n\nReading from a non-local-DC replica is not a problem anymore due to the introduction of describe_local_ring (CASSANDRA-6268), that limits the input splits to the local DC.\n\nI would like to acknowledge my colleague Danilo Penna Queiroz who paired with me on this patch. ","from":"developer"},{"body":"[~pkolaczk] to review","from":"developer"},{"body":"[~pkolaczk] any update on this? cheers! :)","from":"developer"},{"body":"[~pauloricardomg] I can't promise, but I try to do that at the end of this week. ","from":"developer"},{"body":"Attaching patch based on 1.2.16 and fixed patch (v2) for 2.0.\n\nMaybe the 1.2 patch can still make it to 1.2.17...","from":"developer"},{"body":"[~pkolaczk] any update on this? this review has been roaming for quite some time now...\nsorry for bothering but would be nice to see this integrated. cheers!","from":"developer"},{"body":"http://mail-archives.apache.org/mod_mbox/cassandra-dev/201410.mbox/%3CCALdd-zjmvp7JOtguZ_k951RQHDtFt1cthX=RnHQ332C=gAZbjw@mail.gmail.com%3E","from":"developer"},{"body":" [~pkolaczk] + [~pauloricardomg], AFAIK everything thrift related is frozen, so i presume the patch isn't going to be applied to master.\nOtherwise it's +1 on the patch from me.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Paulo, can you rebase to 2.1?","from":"developer"},{"body":"Rebased 2.1 patch available [here|https://github.com/pauloricardomg/cassandra/tree/2388-2.1].\n\n2.1 tests:\n* [testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2388-2.1-testall/lastCompletedBuild/testReport/]\n* [dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2388-2.1-dtest/lastCompletedBuild/testReport/]\n\n2.2 tests (don't know if {{ColumnFamilyInputFormat}} and {{ColumnFamilyRecordReader}} should be deprecated by then, but the classes are still present there):\n* [testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2388-2.2-testall/lastCompletedBuild/testReport/]\n* [dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2388-2.2-dtest/lastCompletedBuild/testReport/]","from":"developer"},{"body":"Committed, thanks.","from":"developer"}],"created":"2011-03-25T20:40:24.000+0000","description":"ColumnFamilyRecordReader only tries the first location for a given split. We should try multiple locations for a given split.","issue_id":"12502417","key":"CASSANDRA-2388","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-11-20T13:54:28.000+0000","role":"fixed_distractor","summary":"ColumnFamilyRecordReader fails for a given split because a host is down, even if records could reasonably be read from other replica."} {"case_id":"12503450","cluster":"DISTRACTOR-CASSANDRA-2419","comments":[{"body":"One solution I see to this problem would be to record along with the replay position the time when we last updated this replay position. The during recover, we would first look at all the sstables (for the CF) and if a sstable is freshly flushed (which implies that we have a marker to know that a sstable was never compacted) and have a modification time higher that the last time we updated the replay position, then we'll just remove the sstable since we know it will be fully replayed.\n\nNote that to work correctly we also need a way to mark a freshly flushed sstable as 'non compactable' during the time it takes to mark the commit log.\n\nWe would probably only do this for counter CF just to be on the safe side.\n\nOpinions ?","created":"2011-04-05T21:02:55.375+0000"},{"body":"What if instead of the CL \"header\" we record the CL context as part of an sstable footer? (footer is less likely to cause bugs w/ sstable math that assumes 0 = start of first row.) then there is no race.","created":"2011-04-05T21:25:59.077+0000"},{"body":"Hmm, I think we need both the CL header and this information, since this flush footer would only give us when we flushed which is not the same as \"do I need to replay.\"\n\nFor instance: if there is no flush marker for a commitlog segment in any existing sstable, that does not necessarily mean no data is in the commitlog for that CF.\n\nSo replay position would be max(dirty at from CL header, flushed at from sstable footers).\n\n(You would need to allow multiple flush contexts in a single sstable footer, to preserve them during compaction.)","created":"2011-04-05T21:34:41.135+0000"},{"body":"What about a new component .metadata for each sstable instead of a footer. I actually think we will have a use for other sstable metadata at some point anyway. For instance we could keep the file format version. That way we wouldn't rely so much on the data file name.","created":"2011-04-05T21:39:28.601+0000"},{"body":"A separate component is a better idea.","created":"2011-04-05T21:44:50.651+0000"},{"body":"Actually I don't think this really solves the race condition. We really shouldn't compact the newly flushed sstable until we marked the commit log, because if we compact it, even if we're able to detect that 'some parts' of a sstable will be replay during recover, there is nothing we can do about it.","created":"2011-04-05T21:52:53.347+0000"},{"body":"I don't understand the problem. Say we have this situation:\n\nCommitLog-1302036825548.log: [full of writes to CF Foo counters, up to position 100. header reads dirty-at 50, our last flush position]\nFoo-g-45-Metadata.db: [flushed at position (1302036825548, 50)]\nFoo-g-46-Metadata.db: [flushed at position (1302036825548, 100)]\n\nWe compact and get\nFoo-g-47-Metadata.db: [flushed at position (1302036825548, 100)]\n\nIf we die and restart here we will correctly start reply of Foo at position 100 in this segment.\n\n(we can combine to a single flushed-at entry in this case since they were from the same CL segment. if they were from different segments we would keep both.)","created":"2011-04-05T22:01:02.724+0000"},{"body":"Right, that was just me not getting you idea at first. Make sense, sorry.","created":"2011-04-05T22:10:46.122+0000"},{"body":"Attaching patch implementing Jonathan's idea to record the CL replay position along with the sstable.","created":"2011-04-07T00:38:40.047+0000"},{"body":"I'm wondering, couldn't we just drop the commit log header if we do that.","created":"2011-04-07T00:40:59.629+0000"},{"body":"Yes, I think we can.","created":"2011-04-07T01:56:56.580+0000"},{"body":"bq. I'm wondering, couldn't we just drop the commit log header if we do that.\n\nWere you planning to update w/ that change?","created":"2011-04-14T21:39:21.472+0000"},{"body":"v2 removes commit log header completely in favor of sstable metadata about where to replay (patch against 0.8).\n\nThis differs from v1 in that instead of keeping every (segment, replay_position) pair, we keep for a given sstable, only the position for the most recent segment (that is, we leverage the fact that we use increasing timestamps for commit logs).\n\nThe reason for this is twofold:\n # this more compact (and simple)\n # if we remove the commit log header, we need to be able to say if a given segment is dirty or not for a given column family. That is, we don't want to know if some replay position existed on this segment, but if a relevant one still exist. So for a given column family we really only care about the newest (segment, replay_position) pair.\n\nNow there is the question of the update path. With this patch, the (existing) commit log headers will be ignored. This means that ideally before updating to a version having this patch people would use drain. If they do not, then the commit logs will be fully replayed. Pre-0.8, it's not a big deal. With counters, this could mean over-counts (that's exactly what this ticket is about). So I would be in favor of putting this for 0.8.0, since it is a bug fix and it will avoids the problem of upgrading from a version already having counters. But I would admit this is not trivial patch, so ...\n","created":"2011-04-29T13:57:21.182+0000"},{"body":"v3 attached with some changes:\n\n- SSTableMetadata removed; replayposition becomes a field in SSTable that is serialized w/ statistics\n- RP is final and part of the SSTable constructor\n- RP implements Comparable instead of a one-off resolve API; RP.getReplayPosition encapsulates the find-replay-point logic\n- RP moved to a top-level class and replaces CLContext\n\nI'd like to make the metadata a json blob so we can extend it more easily, so I probably need to re-introduce SSTM. Consider v3 a work in progress.","created":"2011-05-02T19:53:12.011+0000"},{"body":"I tried two ways of storing metadata as yaml -- first with the metadata as java beans that were stored directly as yaml, and second half-manually serializing to a yaml Map -- and both feel clunkier than just using the version field to deal with adding things. (Especially when you need to do version checks anyway when modifying things rather than just adding new fields.)\n\nSo, v4 is substantially the same as v3 but Descriptor version is bumped to g and we use that instead of EOF to when reading RP. (Also, writeStatistics is renamed to writeMetadata.)","created":"2011-05-03T01:57:40.631+0000"},{"body":"v4 looks good, I like those changes (note: I committed r1099037 to fix CFSTest, it had a test with hardcoded sstable filenames using version 'f' and thus was failing)","created":"2011-05-03T12:32:19.208+0000"},{"body":"v5 updates CL replay to use RP.getReplayPosition as well","created":"2011-05-03T15:00:41.291+0000"},{"body":"v6 fixes a test failure in v5","created":"2011-05-03T15:37:34.278+0000"},{"body":"Minor nitpick: in CommitLog.java:recover(), was there a reason to create a new List positions instead of using cfPositions.values() ?\n\nbut other than that, +1 on v6.","created":"2011-05-03T17:05:17.474+0000"},{"body":"bq. in CommitLog.java:recover(), was there a reason to create a new List positions instead of using cfPositions.values()\n\nNope. I'll fix that before commit.\n\nAsked Paul Cannon to also review first though, since obviously we want to be extra sure not to cause regressions here.","created":"2011-05-03T17:35:15.150+0000"},{"body":"committed","created":"2011-05-08T01:42:50.289+0000"},{"body":"Thanks a ton for this work! Transactionality here we come.","created":"2011-05-09T19:12:32.062+0000"}],"conversations":[{"body":"When a memtable was flush, there is a small delay before the commit log replay position gets updated. If the node fails during this delay, all the updates of this memtable will be replay during commit log recovery and will end-up being over-counts.","from":"reporter","subject":"Risk of counter over-count when recovering commit log"},{"body":"One solution I see to this problem would be to record along with the replay position the time when we last updated this replay position. The during recover, we would first look at all the sstables (for the CF) and if a sstable is freshly flushed (which implies that we have a marker to know that a sstable was never compacted) and have a modification time higher that the last time we updated the replay position, then we'll just remove the sstable since we know it will be fully replayed.\n\nNote that to work correctly we also need a way to mark a freshly flushed sstable as 'non compactable' during the time it takes to mark the commit log.\n\nWe would probably only do this for counter CF just to be on the safe side.\n\nOpinions ?","from":"developer"},{"body":"What if instead of the CL \"header\" we record the CL context as part of an sstable footer? (footer is less likely to cause bugs w/ sstable math that assumes 0 = start of first row.) then there is no race.","from":"developer"},{"body":"Hmm, I think we need both the CL header and this information, since this flush footer would only give us when we flushed which is not the same as \"do I need to replay.\"\n\nFor instance: if there is no flush marker for a commitlog segment in any existing sstable, that does not necessarily mean no data is in the commitlog for that CF.\n\nSo replay position would be max(dirty at from CL header, flushed at from sstable footers).\n\n(You would need to allow multiple flush contexts in a single sstable footer, to preserve them during compaction.)","from":"developer"},{"body":"What about a new component .metadata for each sstable instead of a footer. I actually think we will have a use for other sstable metadata at some point anyway. For instance we could keep the file format version. That way we wouldn't rely so much on the data file name.","from":"developer"},{"body":"A separate component is a better idea.","from":"developer"},{"body":"Actually I don't think this really solves the race condition. We really shouldn't compact the newly flushed sstable until we marked the commit log, because if we compact it, even if we're able to detect that 'some parts' of a sstable will be replay during recover, there is nothing we can do about it.","from":"developer"},{"body":"I don't understand the problem. Say we have this situation:\n\nCommitLog-1302036825548.log: [full of writes to CF Foo counters, up to position 100. header reads dirty-at 50, our last flush position]\nFoo-g-45-Metadata.db: [flushed at position (1302036825548, 50)]\nFoo-g-46-Metadata.db: [flushed at position (1302036825548, 100)]\n\nWe compact and get\nFoo-g-47-Metadata.db: [flushed at position (1302036825548, 100)]\n\nIf we die and restart here we will correctly start reply of Foo at position 100 in this segment.\n\n(we can combine to a single flushed-at entry in this case since they were from the same CL segment. if they were from different segments we would keep both.)","from":"developer"},{"body":"Right, that was just me not getting you idea at first. Make sense, sorry.","from":"developer"},{"body":"Attaching patch implementing Jonathan's idea to record the CL replay position along with the sstable.","from":"developer"},{"body":"I'm wondering, couldn't we just drop the commit log header if we do that.","from":"developer"},{"body":"Yes, I think we can.","from":"developer"},{"body":"bq. I'm wondering, couldn't we just drop the commit log header if we do that.\n\nWere you planning to update w/ that change?","from":"developer"},{"body":"v2 removes commit log header completely in favor of sstable metadata about where to replay (patch against 0.8).\n\nThis differs from v1 in that instead of keeping every (segment, replay_position) pair, we keep for a given sstable, only the position for the most recent segment (that is, we leverage the fact that we use increasing timestamps for commit logs).\n\nThe reason for this is twofold:\n # this more compact (and simple)\n # if we remove the commit log header, we need to be able to say if a given segment is dirty or not for a given column family. That is, we don't want to know if some replay position existed on this segment, but if a relevant one still exist. So for a given column family we really only care about the newest (segment, replay_position) pair.\n\nNow there is the question of the update path. With this patch, the (existing) commit log headers will be ignored. This means that ideally before updating to a version having this patch people would use drain. If they do not, then the commit logs will be fully replayed. Pre-0.8, it's not a big deal. With counters, this could mean over-counts (that's exactly what this ticket is about). So I would be in favor of putting this for 0.8.0, since it is a bug fix and it will avoids the problem of upgrading from a version already having counters. But I would admit this is not trivial patch, so ...\n","from":"developer"},{"body":"v3 attached with some changes:\n\n- SSTableMetadata removed; replayposition becomes a field in SSTable that is serialized w/ statistics\n- RP is final and part of the SSTable constructor\n- RP implements Comparable instead of a one-off resolve API; RP.getReplayPosition encapsulates the find-replay-point logic\n- RP moved to a top-level class and replaces CLContext\n\nI'd like to make the metadata a json blob so we can extend it more easily, so I probably need to re-introduce SSTM. Consider v3 a work in progress.","from":"developer"},{"body":"I tried two ways of storing metadata as yaml -- first with the metadata as java beans that were stored directly as yaml, and second half-manually serializing to a yaml Map -- and both feel clunkier than just using the version field to deal with adding things. (Especially when you need to do version checks anyway when modifying things rather than just adding new fields.)\n\nSo, v4 is substantially the same as v3 but Descriptor version is bumped to g and we use that instead of EOF to when reading RP. (Also, writeStatistics is renamed to writeMetadata.)","from":"developer"},{"body":"v4 looks good, I like those changes (note: I committed r1099037 to fix CFSTest, it had a test with hardcoded sstable filenames using version 'f' and thus was failing)","from":"developer"},{"body":"v5 updates CL replay to use RP.getReplayPosition as well","from":"developer"},{"body":"v6 fixes a test failure in v5","from":"developer"},{"body":"Minor nitpick: in CommitLog.java:recover(), was there a reason to create a new List positions instead of using cfPositions.values() ?\n\nbut other than that, +1 on v6.","from":"developer"},{"body":"bq. in CommitLog.java:recover(), was there a reason to create a new List positions instead of using cfPositions.values()\n\nNope. I'll fix that before commit.\n\nAsked Paul Cannon to also review first though, since obviously we want to be extra sure not to cause regressions here.","from":"developer"},{"body":"committed","from":"developer"},{"body":"Thanks a ton for this work! Transactionality here we come.","from":"developer"}],"created":"2011-04-05T20:49:22.000+0000","description":"When a memtable was flush, there is a small delay before the commit log replay position gets updated. If the node fails during this delay, all the updates of this memtable will be replay during commit log recovery and will end-up being over-counts.","issue_id":"12503450","key":"CASSANDRA-2419","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-05-08T01:42:50.000+0000","role":"fixed_distractor","summary":"Risk of counter over-count when recovering commit log"} {"case_id":"12503646","cluster":"DISTRACTOR-CASSANDRA-2433","comments":[{"body":"Attached patches are against 0.8.\n\nThis tries to catch what can go wrong with repair and reports it back to the user by making the full repair throw an exception. More precisely:\n * patch 0001: add a method to repair for reporting failure and propagate that up to the repair session. This puts repair session on a specific stage (instead of having RepairSession be a Thread) and use a future to allow waiting on completion. This allows a cleaner API to deal with errors (the Future.get() simply throw an ExecutionException) and this add the advantage of stage management to repair sessions.\n * patch 0002: Make repair session register through gossip to be informed of node dying and failing the session when that happens.\n * patch 0003: Reports errors during streaming to the repair session. This actually introduces a generic way to handle streaming failures and after that we should probably update the other user of streaming to deal correctly with failure too.\n * patch 004: Catch errors during validation compaction and push them up to repair (whether those happens on the coordinator of the repair or not).\n\nNote that this includes streaming failures and thus includes stuffs from the patch of Aaron Morton attached on CASSANDRA-2088, but contrarily to that patch, it takes the approach of failing fast. This means that if streaming fails on a file, it fails the streaming altogether (same for repair). I think this is simpler code-wise and more useful from the point of view of the user, since a failure means the use will have to retry anyway.\n\nLast but not least, this makes some modification to messages. So either this goes into 0.8.0 (which I think it should, because this really is a bug fix and fixes something that is a pain for users), or we should had a new messaging version for 0.8.0 and modify this to take it into account (we should probably add a 0.8.0 version to the messaging service anyway).\n","created":"2011-04-21T11:45:12.903+0000"},{"body":"Attaching rebased patch (against 0.8.1). It also change the behavior a little bit so as to not fail repair right away if a problem occur (it still throw an exception at the end if any problem had occured). It turns out to be slightly simpler that way. Especially for CASSANDRA-1610.","created":"2011-05-17T20:01:42.646+0000"},{"body":"0001\n* Since we're not trying to control throughput or monitor sessions, could we just use Stage.MISC?\n\n0002\n* I think RepairSession.exception needs to be volatile to ensure that the awoken thread sees it\n* Would it be better if RepairSession implemented IEndpointStateChangeSubscriber directly?\n* The endpoint set needs to be threadsafe, since it will be modified by the endpoint state change thread, and the AE_STAGE thread\n\n0003\n* Should StreamInSession.retries be volatile/atomic? (likely they won't retry quickly enough for it to be a problem, but...)\n\n0004\n* Playing devil's advocate: would sending a half-built tree in case of failure still be useful?\n* success might need to be volatile as well\n\nThanks Sylvain!","created":"2011-05-23T20:09:40.110+0000"},{"body":"Attaching v3 rebased (on 0.8).\n\nbq. Since we're not trying to control throughput or monitor sessions, could we just use Stage.MISC?\n\nThe thing is that repair session are very long lived. And MISC is single threaded. So that would block other task that are not supposed to block. We could make MISC multi-threaded but even then it's not a good idea to mix short lived and long lived task on the same stage.\n\nbq. I think RepairSession.exception needs to be volatile to ensure that the awoken thread sees it\n\nDone in v3.\n\nbq. Would it be better if RepairSession implemented IEndpointStateChangeSubscriber directly?\n\nGood idea, it's slightly simpler, done in v3.\n\nbq. The endpoint set needs to be threadsafe, since it will be modified by the endpoint state change thread, and the AE_STAGE thread\n\nDone in v3. That will probably change with CASSANDRA-2610 anyway (which I have to update)\n\nbq. Should StreamInSession.retries be volatile/atomic? (likely they won't retry quickly enough for it to be a problem, but...)\n\nI did not change that, but if it's a problem for retries to not be volatile, I suspect having StreamInSession.current not volatile is also a problem. But really I'd be curious to see that be a problem.\n\nbq. Playing devil's advocate: would sending a half-built tree in case of failure still be useful?\n\nI don't think it is. Or more precisely, if you do send half-built tree, you'll have to be careful that the other doesn't consider what's missing as ranges not being in sync (I don't think people will be happy with tons of data being stream just because we happen to have a bug that make compaction throw an exception during the validation). So I think you cannot do much with a half-built tree, and it will add complication. For a case where people will need to restart a repair anyway once whatever happened is fixed\n\nbq. success might need to be volatile as well\n\nDone in v3.\n","created":"2011-06-09T14:09:31.210+0000"},{"body":"Attaching v4 that is rebased and simply set the reties variable in StreamInSession volatile after all (I've removed old version because it was a mess).","created":"2011-06-15T12:27:21.068+0000"},{"body":"Hey Sylvain: sorry it took me so long to get back to this one. Would you mind rebasing it?","created":"2011-07-18T08:08:22.155+0000"},{"body":"Attaching a rebase of the two previous first patches as '2433.patch'. That is, this patch adds registering in gossip so that repair fails and report it to the user when a node participating to the repair dies. Compared to the previous version, it fails fast because it's the easier thing to do now and a better option imho.\n\nI should mention that while it is lame that repair get stuck when a node dies and we should fix it, this means that if a node is wrongly marked down, we will fail repair for no reason (but I suppose it's a failure detector problem).\n\nAttached patch is against 0.8. This has no upgrade consequence of any sort and is a reasonably simple patch, so I think it could be worth committing in 0.8.\nThe rest of what was in previous patch 0003 and 0004 cannot go into 0.8 because it changes the wire protocol, so I will rebase against trunk directly, and maybe in another ticket. Having this first patch committed would help with that though :)","created":"2011-08-02T17:52:47.018+0000"},{"body":"Yuki, can you review this patch?","created":"2011-08-08T17:46:58.038+0000"},{"body":"Attached v2 is rebased and use a higher conviction threshold before deciding to fail the repair, as the goal here is to avoid having a repair getting stuck for hours, but we want to avoid stopping a repair just because a node got into a longer than usual GC pause.\n\nThe threshold used is twice the configured phi_convict_threshold. This give 16 by default, which if I trust the original 'phi accrual failure detection' should give an order of magnitude less false positive than 8 (for about an order of magnitude in the detection time though). It feels reasonable to me but if a FD specialist want to voice his opinion, please do.","created":"2011-08-30T15:25:15.955+0000"},{"body":"- Why do we need the new AE_SESSIONS stage?\n- I prefer using WrappedRunnable to a Callable when you want to allow exceptions but don't care about a return value\n- I think we can avoid a bunch of no-op onConvicts if RepairSession were to subscribe to FD directly instead of going through Gossip (i.e., leave IEndpointStateChangeSubscriber unchanged and expose convict in IFailureDetectionEventListener for when we need to go low-level). Gossip is about high-level \"events\" which doesn't really fit here.","created":"2011-08-30T15:42:06.829+0000"},{"body":"bq. Why do we need the new AE_SESSIONS stage?\n\nIf you mean \"why AE_SESSIONS when we already have the AE stage?\", then it is because repair push stuffs on the AE stage that it wait for, so we would deadlock. If you mean \"why a stage?\", it felt cleaner that just a Thread now that we want to check for exception at the end of the exception. If you mean \"why a stage rather than a simple ThreadExecutor?\", it is a good question. I guess it was just some reflex of mine to get a JMXEnabledThreadPool, but it's probably not worth a stage, not even the jmx enabledness maybe.\n\nbq. I prefer using WrappedRunnable to a Callable when you want to allow exceptions but don't care about a return value\n\nAgreed. I'll update the patch.\n\nbq. I think we can avoid a bunch of no-op onConvicts if RepairSession were to subscribe to FD directly instead of going through Gossip\n\nYeah, I kind of started with that but the problem is that we must deal with the case of a node restarting before it has been convicted (especially if the conviction threshold is higher), which the FD won't see. We could deal of that last situation separately and have Gossip call some trigger into AntiEntropy on a gossip generation change to indicate to stop every started session involving the given endpoint, but creating a dependency of gossip to anti-entropy didn't felt like a good idea a priori.","created":"2011-08-30T16:37:07.839+0000"},{"body":"bq. it's probably not worth a stage, not even the jmx enabledness maybe\n\nSomeone's probably going to want the JMX information but let's keep Stages for Verb-associated tasks.\n\nbq. the problem is that we must deal with the case of a node restarting before it has been convicted (especially if the conviction threshold is higher), which the FD won't see\n\nHow about splitting onDead and onRestart in EndpointStateChange, then? Then RS could implement convict and onRestart (ignoring onDead); other ESCS listeners could implement onRestart == onDead. That would maintain the \"ESCS is about events, FDEL is low-level convict information\" separation of roles.","created":"2011-08-30T16:44:22.826+0000"},{"body":"bq. Someone's probably going to want the JMX information but let's keep Stages for Verb-associated tasks\n\nSounds good, updated patch add a new executor directly into AntiEntropy.\n\nbq. How about splitting onDead and onRestart in EndpointStateChange, then?\n\nDone.\n\nbq. I prefer using WrappedRunnable to a Callable\n\nI changed to use WrappedRunnable. However, we still need to have access to both the repair session and the future from the executor so the implementation returns a pair of those two objects. I'm only marginally convinced this is cleaner than the previous solution...\n","created":"2011-08-31T14:46:03.010+0000"},{"body":"(Sorry, I had attached the wrong version of v3, corrected now)","created":"2011-08-31T14:52:17.076+0000"},{"body":"bq. we still need to have access to both the repair session and the future from the executor so the implementation returns a pair of those two objects\n\nYou can still use the RepairFuture approach, just use the FutureTask(Runnable, V) constructor","created":"2011-08-31T14:56:18.410+0000"},{"body":"You're right, don't know why I got carried away like that. v4 \"fixes\" this.","created":"2011-08-31T15:46:43.396+0000"},{"body":"+1","created":"2011-08-31T15:51:38.803+0000"},{"body":"Committed, thanks.\n\nThis probably solves most of the case where repair was hanging infinitely. I've created CASSANDRA-3112 to handle the remaining cases, but it is much less urgent imho. Marking that one as resolved","created":"2011-08-31T16:36:59.840+0000"}],"conversations":[{"body":"Running repair in cases where a stream fails we are seeing multiple problems.\n\n1. Although retry is initiated and completes, the old stream doesn't seem to clean itself up and repair hangs.\n2. The temp files are left behind and multiple failures can end up filling up the data partition.\n\nThese issues together are making repair very difficult for nearly everyone running repair on a non-trivial sized data set.\n\nThis issue is also being worked on w.r.t CASSANDRA-2088, however that was moved to 0.8 for a few reasons. This ticket is to fix the immediate issues that we are seeing in 0.7.","from":"reporter","subject":"Failed Streams Break Repair"},{"body":"Attached patches are against 0.8.\n\nThis tries to catch what can go wrong with repair and reports it back to the user by making the full repair throw an exception. More precisely:\n * patch 0001: add a method to repair for reporting failure and propagate that up to the repair session. This puts repair session on a specific stage (instead of having RepairSession be a Thread) and use a future to allow waiting on completion. This allows a cleaner API to deal with errors (the Future.get() simply throw an ExecutionException) and this add the advantage of stage management to repair sessions.\n * patch 0002: Make repair session register through gossip to be informed of node dying and failing the session when that happens.\n * patch 0003: Reports errors during streaming to the repair session. This actually introduces a generic way to handle streaming failures and after that we should probably update the other user of streaming to deal correctly with failure too.\n * patch 004: Catch errors during validation compaction and push them up to repair (whether those happens on the coordinator of the repair or not).\n\nNote that this includes streaming failures and thus includes stuffs from the patch of Aaron Morton attached on CASSANDRA-2088, but contrarily to that patch, it takes the approach of failing fast. This means that if streaming fails on a file, it fails the streaming altogether (same for repair). I think this is simpler code-wise and more useful from the point of view of the user, since a failure means the use will have to retry anyway.\n\nLast but not least, this makes some modification to messages. So either this goes into 0.8.0 (which I think it should, because this really is a bug fix and fixes something that is a pain for users), or we should had a new messaging version for 0.8.0 and modify this to take it into account (we should probably add a 0.8.0 version to the messaging service anyway).\n","from":"developer"},{"body":"Attaching rebased patch (against 0.8.1). It also change the behavior a little bit so as to not fail repair right away if a problem occur (it still throw an exception at the end if any problem had occured). It turns out to be slightly simpler that way. Especially for CASSANDRA-1610.","from":"developer"},{"body":"0001\n* Since we're not trying to control throughput or monitor sessions, could we just use Stage.MISC?\n\n0002\n* I think RepairSession.exception needs to be volatile to ensure that the awoken thread sees it\n* Would it be better if RepairSession implemented IEndpointStateChangeSubscriber directly?\n* The endpoint set needs to be threadsafe, since it will be modified by the endpoint state change thread, and the AE_STAGE thread\n\n0003\n* Should StreamInSession.retries be volatile/atomic? (likely they won't retry quickly enough for it to be a problem, but...)\n\n0004\n* Playing devil's advocate: would sending a half-built tree in case of failure still be useful?\n* success might need to be volatile as well\n\nThanks Sylvain!","from":"developer"},{"body":"Attaching v3 rebased (on 0.8).\n\nbq. Since we're not trying to control throughput or monitor sessions, could we just use Stage.MISC?\n\nThe thing is that repair session are very long lived. And MISC is single threaded. So that would block other task that are not supposed to block. We could make MISC multi-threaded but even then it's not a good idea to mix short lived and long lived task on the same stage.\n\nbq. I think RepairSession.exception needs to be volatile to ensure that the awoken thread sees it\n\nDone in v3.\n\nbq. Would it be better if RepairSession implemented IEndpointStateChangeSubscriber directly?\n\nGood idea, it's slightly simpler, done in v3.\n\nbq. The endpoint set needs to be threadsafe, since it will be modified by the endpoint state change thread, and the AE_STAGE thread\n\nDone in v3. That will probably change with CASSANDRA-2610 anyway (which I have to update)\n\nbq. Should StreamInSession.retries be volatile/atomic? (likely they won't retry quickly enough for it to be a problem, but...)\n\nI did not change that, but if it's a problem for retries to not be volatile, I suspect having StreamInSession.current not volatile is also a problem. But really I'd be curious to see that be a problem.\n\nbq. Playing devil's advocate: would sending a half-built tree in case of failure still be useful?\n\nI don't think it is. Or more precisely, if you do send half-built tree, you'll have to be careful that the other doesn't consider what's missing as ranges not being in sync (I don't think people will be happy with tons of data being stream just because we happen to have a bug that make compaction throw an exception during the validation). So I think you cannot do much with a half-built tree, and it will add complication. For a case where people will need to restart a repair anyway once whatever happened is fixed\n\nbq. success might need to be volatile as well\n\nDone in v3.\n","from":"developer"},{"body":"Attaching v4 that is rebased and simply set the reties variable in StreamInSession volatile after all (I've removed old version because it was a mess).","from":"developer"},{"body":"Hey Sylvain: sorry it took me so long to get back to this one. Would you mind rebasing it?","from":"developer"},{"body":"Attaching a rebase of the two previous first patches as '2433.patch'. That is, this patch adds registering in gossip so that repair fails and report it to the user when a node participating to the repair dies. Compared to the previous version, it fails fast because it's the easier thing to do now and a better option imho.\n\nI should mention that while it is lame that repair get stuck when a node dies and we should fix it, this means that if a node is wrongly marked down, we will fail repair for no reason (but I suppose it's a failure detector problem).\n\nAttached patch is against 0.8. This has no upgrade consequence of any sort and is a reasonably simple patch, so I think it could be worth committing in 0.8.\nThe rest of what was in previous patch 0003 and 0004 cannot go into 0.8 because it changes the wire protocol, so I will rebase against trunk directly, and maybe in another ticket. Having this first patch committed would help with that though :)","from":"developer"},{"body":"Yuki, can you review this patch?","from":"developer"},{"body":"Attached v2 is rebased and use a higher conviction threshold before deciding to fail the repair, as the goal here is to avoid having a repair getting stuck for hours, but we want to avoid stopping a repair just because a node got into a longer than usual GC pause.\n\nThe threshold used is twice the configured phi_convict_threshold. This give 16 by default, which if I trust the original 'phi accrual failure detection' should give an order of magnitude less false positive than 8 (for about an order of magnitude in the detection time though). It feels reasonable to me but if a FD specialist want to voice his opinion, please do.","from":"developer"},{"body":"- Why do we need the new AE_SESSIONS stage?\n- I prefer using WrappedRunnable to a Callable when you want to allow exceptions but don't care about a return value\n- I think we can avoid a bunch of no-op onConvicts if RepairSession were to subscribe to FD directly instead of going through Gossip (i.e., leave IEndpointStateChangeSubscriber unchanged and expose convict in IFailureDetectionEventListener for when we need to go low-level). Gossip is about high-level \"events\" which doesn't really fit here.","from":"developer"},{"body":"bq. Why do we need the new AE_SESSIONS stage?\n\nIf you mean \"why AE_SESSIONS when we already have the AE stage?\", then it is because repair push stuffs on the AE stage that it wait for, so we would deadlock. If you mean \"why a stage?\", it felt cleaner that just a Thread now that we want to check for exception at the end of the exception. If you mean \"why a stage rather than a simple ThreadExecutor?\", it is a good question. I guess it was just some reflex of mine to get a JMXEnabledThreadPool, but it's probably not worth a stage, not even the jmx enabledness maybe.\n\nbq. I prefer using WrappedRunnable to a Callable when you want to allow exceptions but don't care about a return value\n\nAgreed. I'll update the patch.\n\nbq. I think we can avoid a bunch of no-op onConvicts if RepairSession were to subscribe to FD directly instead of going through Gossip\n\nYeah, I kind of started with that but the problem is that we must deal with the case of a node restarting before it has been convicted (especially if the conviction threshold is higher), which the FD won't see. We could deal of that last situation separately and have Gossip call some trigger into AntiEntropy on a gossip generation change to indicate to stop every started session involving the given endpoint, but creating a dependency of gossip to anti-entropy didn't felt like a good idea a priori.","from":"developer"},{"body":"bq. it's probably not worth a stage, not even the jmx enabledness maybe\n\nSomeone's probably going to want the JMX information but let's keep Stages for Verb-associated tasks.\n\nbq. the problem is that we must deal with the case of a node restarting before it has been convicted (especially if the conviction threshold is higher), which the FD won't see\n\nHow about splitting onDead and onRestart in EndpointStateChange, then? Then RS could implement convict and onRestart (ignoring onDead); other ESCS listeners could implement onRestart == onDead. That would maintain the \"ESCS is about events, FDEL is low-level convict information\" separation of roles.","from":"developer"},{"body":"bq. Someone's probably going to want the JMX information but let's keep Stages for Verb-associated tasks\n\nSounds good, updated patch add a new executor directly into AntiEntropy.\n\nbq. How about splitting onDead and onRestart in EndpointStateChange, then?\n\nDone.\n\nbq. I prefer using WrappedRunnable to a Callable\n\nI changed to use WrappedRunnable. However, we still need to have access to both the repair session and the future from the executor so the implementation returns a pair of those two objects. I'm only marginally convinced this is cleaner than the previous solution...\n","from":"developer"},{"body":"(Sorry, I had attached the wrong version of v3, corrected now)","from":"developer"},{"body":"bq. we still need to have access to both the repair session and the future from the executor so the implementation returns a pair of those two objects\n\nYou can still use the RepairFuture approach, just use the FutureTask(Runnable, V) constructor","from":"developer"},{"body":"You're right, don't know why I got carried away like that. v4 \"fixes\" this.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed, thanks.\n\nThis probably solves most of the case where repair was hanging infinitely. I've created CASSANDRA-3112 to handle the remaining cases, but it is much less urgent imho. Marking that one as resolved","from":"developer"}],"created":"2011-04-07T15:40:10.000+0000","description":"Running repair in cases where a stream fails we are seeing multiple problems.\n\n1. Although retry is initiated and completes, the old stream doesn't seem to clean itself up and repair hangs.\n2. The temp files are left behind and multiple failures can end up filling up the data partition.\n\nThese issues together are making repair very difficult for nearly everyone running repair on a non-trivial sized data set.\n\nThis issue is also being worked on w.r.t CASSANDRA-2088, however that was moved to 0.8 for a few reasons. This ticket is to fix the immediate issues that we are seeing in 0.7.","issue_id":"12503646","key":"CASSANDRA-2433","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-08-31T16:36:59.000+0000","role":"fixed_distractor","summary":"Failed Streams Break Repair"} {"case_id":"12504486","cluster":"DISTRACTOR-CASSANDRA-2494","comments":[{"body":"As far as I can tell the consistency being asked for was never promised by Cassandra is in fact not expected.\n\nThe expected behavior of writes is that they propagate; the difference between ONE and QUORUM is just how many are required to receive a write prior to a return to the client with a successful error code. For reads, that means you may get lucky at ONE or you may get lucky at QUORUM; the positive guarantee is in the case of a *completing* QUORUM write followed by a QUORUM read.\n\nSo just to be clear, although I don't think this is what is being asked for: As far as I know, it has never been the case, nor the intent to promise, that a write which fails is guaranteed not to eventually complete. Simply \"fixing\" reads is not enough; by design the data will be replicated during read-repair and AES - this is how consistency is achieved in Cassandra.\n\nHowever, it sounds like what is being asked for is not that they don't propagate in the event of a write \"failure\", but just that reads don't see the writes until they are sufficiently propagated to guarantee that any future QUORUM read will also see the data. I can understand that is desirable, in the sense of achieving monotonically forward-moving data as the benchmark/test from the e-mail thread does. Another way to look at is that maybe you never want to read data successfully prior to achieving a certain level of replication, in order to avoid a client ever seeing data that may suddenly go away due to e.g. a node failure in spite of said failure not exceeding the number of failures the cluster was designed to survive.\n\nSo the key point would be the bit about guaranteeing that any \"future QUORUM read will see the data or data subsequently overwritten\", and actively read-repairing and waiting for it to happen would take care of that. It would be important to ensure that the act of ensuring a quorum of nodes have seen the data is the important part; one should not await for a quorum to agree on the *current* version of the data as that would create potentially unbounded round-trips on hotly contended data.\n\nThing to consider: One might think about cases where read-repair is currently not done, like range slices, and how an implementation that requires read repair for consistency affects that.\n\n","created":"2011-04-17T21:01:09.146+0000"},{"body":"Peter Shuller wrote,\n\n\"However, it sounds like what is being asked for is not that they don't propagate in the event of a write \"failure\", but just that reads don't see the writes until they are sufficiently propagated to guarantee that any future QUORUM read will also see the data.\"\n\nYes, that is the issue. The comment in the bug about writing at ONE and reading at QUORUM is just a way of testing this new guarantee in a distributed test, if Cassandra has those.","created":"2011-04-18T04:20:10.172+0000"},{"body":"I would think that reads at QUORUM should never go backwards. Even if the Write was at ZERO. If there were writes to the cluster of a=1 time=5, a=2 time=10, a=3 time=15, and I do a read at QUORUM which tells me a=3 time=15, I should not be able to do another read at QUORUM and get a=2 time=10.","created":"2011-04-18T20:41:27.221+0000"},{"body":"W plus R must be _greater than_ N for consistency.\n\nEDIT: And adding a blocking implicit write step to QUORUM reads by waiting for read repair is not reasonable.","created":"2011-04-22T05:17:00.824+0000"},{"body":"I don't think anyone is claiming otherwise, unless I'm misunderstanding. The problem is that while the \"if sucessfully written to quorum, subsequent quorum reads will see it\" guarantee is indeed maintained, it is possible for quorum reads to see data go backwards (on a timeline) in the event of a *failed* attempted quorum write. This includes the possibility of reads seeing data that then permanently vanishes, even though you only lost say 1 node that you designed your cluster for surviving (RF >= 3, QUORUM). (\"lost 1 node\" can be substituted with \"killed 1 node in periodic commit mode\")\n\nI still don't think this is a violation of what was promised, but I can see how making the further guarantee would make for more useful consistency semantics in some cases.\n\nWith respect to implicit write: An alternative is to adjust reconciliation logic when applied as part of reads (as opposed to AES, hinted hand-off, writes) to take consistency level into account and only consider columns whose timestamp is >= the greatest timestamp that has quorum (off the top of my head I think that should be correct in call cases, but I didn't think this through terribly).\n","created":"2011-04-22T08:26:09.087+0000"},{"body":"Ok, so my last suggestion is in fact broken. A counter example is:\n\n A: column @ t1\n B: column @ t2\n C: column @ t3\n\nIf A + B is participating, A's column @ t1 has timestamp quorum and would be selected. If B + C is participating, B's column is picked. Thus, a read where B + C participates will see data that will be reverted once A + B happens to be picked.\n\nNote to self: Think before posting.\n","created":"2011-04-22T08:34:42.861+0000"},{"body":"I think the guarantee of quorum reads not seeing old writes once a quorum read sees a new write is very useful. I suspect most people already think that this guarantee occurs, including, it seems, Jonathan Ellis whose quote can be found in the email thread linked to in the bug,\n\n\"The important guarantee this gives you is that once one quorum read sees the new value, all others will too. You can't see the newest version, then see an older version on a subsequent write [sic, I\nassume he meant read], which is the characteristic of non-strong consistency\"\n\n\n\n","created":"2011-04-22T14:45:41.755+0000"},{"body":"The problem is you are considering the consistency of reads but not write. The guarantee is: \"quorum reads will not see old quorum write once a quorum read sees a new quorum\". Period. I you don't consider the consistency of a write, consider the case of a CL.ANY write. In this case, the update may not be at all on any replica. How can we ensure the quorum read property that you want ? We query all nodes for quorum reads in case there is an hint somewhere ?\n\nIf you look at the Consistency part of http://wiki.apache.org/cassandra/ArchitectureOverview, it seems to me that it is pretty clear that the consistency of reads *and* writes is involved to achieve strong consistency. So I would hope 'most people' are aware of that.","created":"2011-04-22T15:01:25.371+0000"},{"body":"The issue is that of *failed* QUORUM writes. I.e., you design your system to use QUORUM writes and QUORUM reads, and expect that once a QUORUM read sees a given piece of data a subsequent QUORUM read will also see it (or a later data). A *failed* QUORUM write that was replicated to less than a QUORUM would be visible as part of QUORUM reads that happen to touch one of those replicas, but there is no guarantee that subsequent reads see it.\n\nI was under the impression this was never an intended guarantee. Apparently I may be wrong about that given the jbellis quote above. In either case, it is certainly not an *actual* guarantee given by the current implementation.\n\nThe guarantee that a *successful* QUORUM write is seen by a subsequent QUORUM read is, as far as I can tell, not in question here.\n\n","created":"2011-04-22T15:11:55.698+0000"},{"body":"To be clear, this is a new guarantee. The current guarantee is R+W>N gives you consistency. This bug is asking that a successful quorum read of A means that A has been committed to a quorum of nodes.\n\n\"How can we ensure the quorum read property that you want ?\"\n\nIf when reading at quorum, and no quorum can be found which agrees on a particular value, then the coordinator ( ? ) will wait for acks of read repair writes (or perhaps just do normal writes) to be returned from a sufficient number of nodes to ensure that the value has been committed to a quorum of nodes.\n\nWithout this new guarantee it is hard for readers to function correctly. The reader does not know that the quorum write failed, or is still in progress, so without reading at ALL, the R+W>N guarantee does not help the reader.\n\n\n\n","created":"2011-04-22T15:22:29.455+0000"},{"body":"I don't see any reason not to guarantee that the replicas we read from provide monotonic read consistency.\n\nPatch attached to do this.\n","created":"2011-07-27T01:46:55.101+0000"},{"body":"Ok, I now see what you mean :)\nMakes perfect sense.\n\nComments on the patch:\n* There is a number of case where scheduleRepairs may not have been called (if the read for repair timeout and/or we had no or only 1 response or we have the situation were removeDeleted removes everything), so repairResults will be null in those cases. In SP, we should check for it.\n* That new wait can extend the rpc timeout to almost twice what it should be. I agree that it is not a huge deal, but by exposing the 'startTime' stored in RepairCallback we can make it so we don't extend it that way.\n* Shouldn't we give the same love to range requests, now that we do repairs there too ?","created":"2011-07-27T11:08:35.136+0000"},{"body":"bq. There is a number of case where scheduleRepairs may not have been called\n\ndefaulted repairResults to emptyList.\n\nbq. That new wait can extend the rpc timeout to almost twice what it should be\n\nWell, sort of -- rpctimeout is working exactly as intended, i.e., to prevent waiting indefinitely for a node that died after we sent it a request. Treating it as \"max time to respond to client\" has never really been correct. (E.g., in the CL > ONE case we can already wait up to rpctimeout twice, one for the original digest read set, and again for the data read after mismatch.) So I don't think we should try to be clever with that here.\n\nbq. Shouldn't we give the same love to range requests, now that we do repairs there too\n\ndone.","created":"2011-08-03T16:34:56.698+0000"},{"body":"bq. Well, sort of – rpctimeout is working exactly as intended, i.e., to prevent waiting indefinitely for a node that died after we sent it a request. Treating it as \"max time to respond to client\" has never really been correct. (E.g., in the CL > ONE case we can already wait up to rpctimeout twice, one for the original digest read set, and again for the data read after mismatch.) So I don't think we should try to be clever with that here.\n\nFair enough. It would probably be useful to make rpctimeout meaning closer to \"max time to respond to client\". Created CASSANDRA-3018 for that though.\n\n+1 on v2.","created":"2011-08-11T15:34:34.068+0000"},{"body":"committed","created":"2011-08-11T19:29:40.295+0000"},{"body":"The relevant code in the patch has changed significantly. Is the monotonic read consistency guarantee still provided?","created":"2015-09-24T20:45:58.942+0000"}],"conversations":[{"body":"As discussed in this thread,\n\nhttp://www.mail-archive.com/user@cassandra.apache.org/msg12421.html\n\nQuorum reads should be consistent. Assume we have a cluster of 3 nodes (X,Y,Z) and a replication factor of 3. If a write of N is committed to X, but not Y and Z, then a read from X should not return N unless the read is committed to at least two nodes. To ensure this, a read from X should wait for an ack of the read repair write from either Y or Z before returning.\n\nAre there system tests for cassandra? If so, there should be a test similar to the original post in the email thread. One thread should write 1,2,3... at consistency level ONE. Another thread should read at consistency level QUORUM from a random host, and verify that each read is >= the last read.","from":"reporter","subject":"Quorum reads are not monotonically consistent"},{"body":"As far as I can tell the consistency being asked for was never promised by Cassandra is in fact not expected.\n\nThe expected behavior of writes is that they propagate; the difference between ONE and QUORUM is just how many are required to receive a write prior to a return to the client with a successful error code. For reads, that means you may get lucky at ONE or you may get lucky at QUORUM; the positive guarantee is in the case of a *completing* QUORUM write followed by a QUORUM read.\n\nSo just to be clear, although I don't think this is what is being asked for: As far as I know, it has never been the case, nor the intent to promise, that a write which fails is guaranteed not to eventually complete. Simply \"fixing\" reads is not enough; by design the data will be replicated during read-repair and AES - this is how consistency is achieved in Cassandra.\n\nHowever, it sounds like what is being asked for is not that they don't propagate in the event of a write \"failure\", but just that reads don't see the writes until they are sufficiently propagated to guarantee that any future QUORUM read will also see the data. I can understand that is desirable, in the sense of achieving monotonically forward-moving data as the benchmark/test from the e-mail thread does. Another way to look at is that maybe you never want to read data successfully prior to achieving a certain level of replication, in order to avoid a client ever seeing data that may suddenly go away due to e.g. a node failure in spite of said failure not exceeding the number of failures the cluster was designed to survive.\n\nSo the key point would be the bit about guaranteeing that any \"future QUORUM read will see the data or data subsequently overwritten\", and actively read-repairing and waiting for it to happen would take care of that. It would be important to ensure that the act of ensuring a quorum of nodes have seen the data is the important part; one should not await for a quorum to agree on the *current* version of the data as that would create potentially unbounded round-trips on hotly contended data.\n\nThing to consider: One might think about cases where read-repair is currently not done, like range slices, and how an implementation that requires read repair for consistency affects that.\n\n","from":"developer"},{"body":"Peter Shuller wrote,\n\n\"However, it sounds like what is being asked for is not that they don't propagate in the event of a write \"failure\", but just that reads don't see the writes until they are sufficiently propagated to guarantee that any future QUORUM read will also see the data.\"\n\nYes, that is the issue. The comment in the bug about writing at ONE and reading at QUORUM is just a way of testing this new guarantee in a distributed test, if Cassandra has those.","from":"developer"},{"body":"I would think that reads at QUORUM should never go backwards. Even if the Write was at ZERO. If there were writes to the cluster of a=1 time=5, a=2 time=10, a=3 time=15, and I do a read at QUORUM which tells me a=3 time=15, I should not be able to do another read at QUORUM and get a=2 time=10.","from":"developer"},{"body":"W plus R must be _greater than_ N for consistency.\n\nEDIT: And adding a blocking implicit write step to QUORUM reads by waiting for read repair is not reasonable.","from":"developer"},{"body":"I don't think anyone is claiming otherwise, unless I'm misunderstanding. The problem is that while the \"if sucessfully written to quorum, subsequent quorum reads will see it\" guarantee is indeed maintained, it is possible for quorum reads to see data go backwards (on a timeline) in the event of a *failed* attempted quorum write. This includes the possibility of reads seeing data that then permanently vanishes, even though you only lost say 1 node that you designed your cluster for surviving (RF >= 3, QUORUM). (\"lost 1 node\" can be substituted with \"killed 1 node in periodic commit mode\")\n\nI still don't think this is a violation of what was promised, but I can see how making the further guarantee would make for more useful consistency semantics in some cases.\n\nWith respect to implicit write: An alternative is to adjust reconciliation logic when applied as part of reads (as opposed to AES, hinted hand-off, writes) to take consistency level into account and only consider columns whose timestamp is >= the greatest timestamp that has quorum (off the top of my head I think that should be correct in call cases, but I didn't think this through terribly).\n","from":"developer"},{"body":"Ok, so my last suggestion is in fact broken. A counter example is:\n\n A: column @ t1\n B: column @ t2\n C: column @ t3\n\nIf A + B is participating, A's column @ t1 has timestamp quorum and would be selected. If B + C is participating, B's column is picked. Thus, a read where B + C participates will see data that will be reverted once A + B happens to be picked.\n\nNote to self: Think before posting.\n","from":"developer"},{"body":"I think the guarantee of quorum reads not seeing old writes once a quorum read sees a new write is very useful. I suspect most people already think that this guarantee occurs, including, it seems, Jonathan Ellis whose quote can be found in the email thread linked to in the bug,\n\n\"The important guarantee this gives you is that once one quorum read sees the new value, all others will too. You can't see the newest version, then see an older version on a subsequent write [sic, I\nassume he meant read], which is the characteristic of non-strong consistency\"\n\n\n\n","from":"developer"},{"body":"The problem is you are considering the consistency of reads but not write. The guarantee is: \"quorum reads will not see old quorum write once a quorum read sees a new quorum\". Period. I you don't consider the consistency of a write, consider the case of a CL.ANY write. In this case, the update may not be at all on any replica. How can we ensure the quorum read property that you want ? We query all nodes for quorum reads in case there is an hint somewhere ?\n\nIf you look at the Consistency part of http://wiki.apache.org/cassandra/ArchitectureOverview, it seems to me that it is pretty clear that the consistency of reads *and* writes is involved to achieve strong consistency. So I would hope 'most people' are aware of that.","from":"developer"},{"body":"The issue is that of *failed* QUORUM writes. I.e., you design your system to use QUORUM writes and QUORUM reads, and expect that once a QUORUM read sees a given piece of data a subsequent QUORUM read will also see it (or a later data). A *failed* QUORUM write that was replicated to less than a QUORUM would be visible as part of QUORUM reads that happen to touch one of those replicas, but there is no guarantee that subsequent reads see it.\n\nI was under the impression this was never an intended guarantee. Apparently I may be wrong about that given the jbellis quote above. In either case, it is certainly not an *actual* guarantee given by the current implementation.\n\nThe guarantee that a *successful* QUORUM write is seen by a subsequent QUORUM read is, as far as I can tell, not in question here.\n\n","from":"developer"},{"body":"To be clear, this is a new guarantee. The current guarantee is R+W>N gives you consistency. This bug is asking that a successful quorum read of A means that A has been committed to a quorum of nodes.\n\n\"How can we ensure the quorum read property that you want ?\"\n\nIf when reading at quorum, and no quorum can be found which agrees on a particular value, then the coordinator ( ? ) will wait for acks of read repair writes (or perhaps just do normal writes) to be returned from a sufficient number of nodes to ensure that the value has been committed to a quorum of nodes.\n\nWithout this new guarantee it is hard for readers to function correctly. The reader does not know that the quorum write failed, or is still in progress, so without reading at ALL, the R+W>N guarantee does not help the reader.\n\n\n\n","from":"developer"},{"body":"I don't see any reason not to guarantee that the replicas we read from provide monotonic read consistency.\n\nPatch attached to do this.\n","from":"developer"},{"body":"Ok, I now see what you mean :)\nMakes perfect sense.\n\nComments on the patch:\n* There is a number of case where scheduleRepairs may not have been called (if the read for repair timeout and/or we had no or only 1 response or we have the situation were removeDeleted removes everything), so repairResults will be null in those cases. In SP, we should check for it.\n* That new wait can extend the rpc timeout to almost twice what it should be. I agree that it is not a huge deal, but by exposing the 'startTime' stored in RepairCallback we can make it so we don't extend it that way.\n* Shouldn't we give the same love to range requests, now that we do repairs there too ?","from":"developer"},{"body":"bq. There is a number of case where scheduleRepairs may not have been called\n\ndefaulted repairResults to emptyList.\n\nbq. That new wait can extend the rpc timeout to almost twice what it should be\n\nWell, sort of -- rpctimeout is working exactly as intended, i.e., to prevent waiting indefinitely for a node that died after we sent it a request. Treating it as \"max time to respond to client\" has never really been correct. (E.g., in the CL > ONE case we can already wait up to rpctimeout twice, one for the original digest read set, and again for the data read after mismatch.) So I don't think we should try to be clever with that here.\n\nbq. Shouldn't we give the same love to range requests, now that we do repairs there too\n\ndone.","from":"developer"},{"body":"bq. Well, sort of – rpctimeout is working exactly as intended, i.e., to prevent waiting indefinitely for a node that died after we sent it a request. Treating it as \"max time to respond to client\" has never really been correct. (E.g., in the CL > ONE case we can already wait up to rpctimeout twice, one for the original digest read set, and again for the data read after mismatch.) So I don't think we should try to be clever with that here.\n\nFair enough. It would probably be useful to make rpctimeout meaning closer to \"max time to respond to client\". Created CASSANDRA-3018 for that though.\n\n+1 on v2.","from":"developer"},{"body":"committed","from":"developer"},{"body":"The relevant code in the patch has changed significantly. Is the monotonic read consistency guarantee still provided?","from":"developer"}],"created":"2011-04-17T14:29:37.000+0000","description":"As discussed in this thread,\n\nhttp://www.mail-archive.com/user@cassandra.apache.org/msg12421.html\n\nQuorum reads should be consistent. Assume we have a cluster of 3 nodes (X,Y,Z) and a replication factor of 3. If a write of N is committed to X, but not Y and Z, then a read from X should not return N unless the read is committed to at least two nodes. To ensure this, a read from X should wait for an ack of the read repair write from either Y or Z before returning.\n\nAre there system tests for cassandra? If so, there should be a test similar to the original post in the email thread. One thread should write 1,2,3... at consistency level ONE. Another thread should read at consistency level QUORUM from a random host, and verify that each read is >= the last read.","issue_id":"12504486","key":"CASSANDRA-2494","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-08-11T19:29:40.000+0000","role":"fixed_distractor","summary":"Quorum reads are not monotonically consistent"} {"case_id":"12504582","cluster":"DISTRACTOR-CASSANDRA-2496","comments":[{"body":"The first patch allows gossip to track dead states and completely changes how removetoken works, since it is the problem with keeping dead state around, and currently broken in many scenarios. The second patch adds the previously reverted CASSANDRA-2115 back now that removetoken is more resilient.","created":"2011-06-01T00:08:52.445+0000"},{"body":"Some explanation of what changed and why it was necessary:\n\nConsider nodes A through D. D is partitioned, and C is dead and needs to be removed. A removetoken will be issued to A for this.\n\nIn the current way we do things, A will modify it's own state by appending information to its status indicating that it will be removing C. B will see this, re-replicate as needed, then report to A that is is done. The problem however, is that since A modified its own state, A is also free to wipe that state out, either by restarting, or simple remove another token, because there's only space for one. If A reboots and then D's partition heals, D will never know C was removed. Worse, it will still have state for C that neither A nor B do, and so it will repopulate the ring with C again.\n\nThis patch changes this by instead having A sleep for RING_DELAY to make sure the generation for C is stable, and then it modifies C's state to indicate it is being removed, just as if C itself had done this. It also appends some extra state to indicate that A will be the removal coordinator. The others nodes see this, re-replicate and report back to A, which then modifies C's state once more to indicate it is completely removed. At this point, it doesn't matter if A dies completely and D's partition heals, since the state is stored in C's gossip information. If A reboots, it will be able to get the correct state information from B, or any other node.\n\nIf A fails while the other nodes are re-replicating, a new removetoken can be started elsewhere, or in the case of other replicas being down preventing removetoken from completing, a removetoken force will remove the node and then repair can be run to restore the replica count.","created":"2011-06-01T16:02:19.105+0000"},{"body":"I see two more things to be done with this patch. First, when re-replicating nodes report back to the removal coordinator, if the coordinator has restarted it won't understand them, and they will infinitely loop retrying the confirmation. Second, since we're holding dead states, we need to make sure that bootstrapping/moving nodes can take over these dead tokens.","created":"2011-07-14T17:40:25.936+0000"},{"body":"These small patches build on the others.\n\n0003-update-gossip-related-comments.patch.txt: updates gossip-related comments derp derp.","created":"2011-07-14T20:54:34.916+0000"},{"body":"0004-do-REMOVING_TOKEN-REMOVED_TOKEN.patch.txt: use REMOVED_TOKEN instead of STATUS_LEFT (would probably be ok either way, but otherwise, the REMOVED_TOKEN state would not be used). Seems this is more the way it was intended.","created":"2011-07-14T20:56:12.473+0000"},{"body":"0005-drain-self-if-removetoken-d-elsewhere.patch.txt : when node X was partitioned and removetoken'd but then it shows up again, it should shut itself down, rather than becoming a zombie","created":"2011-07-14T20:57:51.730+0000"},{"body":"I'll see what I can do to test the \"infinitely loop retrying the confirmation\" and \"bootstrapping/moving nodes can take over these dead tokens\" situations.","created":"2011-07-14T20:58:50.195+0000"},{"body":"Ok, nodes do indeed infinitely retry the replication confirmation in some cases, but it appears it's not just when the former removal coordinator has restarted in the interim- it seems to be when the removetoken is reissued to another, new removal coordinator. In this case, I get this traceback every 10 seconds:\n\n{noformat}\nERROR [MiscStage:9] 2011-07-20 23:42:06,599 AbstractCassandraDaemon.java (line 113) Fatal exception in thread Thread[MiscStage:9,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.service.StorageService.confirmReplication(StorageService.java:2088)\n at org.apache.cassandra.streaming.ReplicationFinishedVerbHandler.doVerb(ReplicationFinishedVerbHandler.java:38)\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:72)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n{noformat}\n\nI'll look into this.\n\nSecond, it seems that moving/joining nodes can take over the removed token fine, once the removetoken is complete. I haven't tried having a node take over the removed token while the removal is ongoing- I assume we can just document that that probably isn't a great idea?","created":"2011-07-20T23:48:29.381+0000"},{"body":"0006-acknowledge-unexpected-repl-fins.patch.txt: don't assert and drop the message when we see an unexpected REPLICATION_FINISHED. Ack it instead, so the sender doesn't continually retry.","created":"2011-07-21T19:50:45.144+0000"},{"body":"0006-acknowledge-unexpected-repl-fins.patch.txt (updated): also log at info when acknowledging the unexpected messages","created":"2011-07-21T20:09:27.489+0000"},{"body":"ok, +1 with these patches.","created":"2011-07-21T20:52:57.386+0000"},{"body":"0007 handles problems when a node has been down longer than aVeryLongTime, and ensures that we advertise the new token states long enough.\n\n0008 makes sure that SS only get involved with removal if the token is a member.","created":"2011-07-23T00:20:53.440+0000"},{"body":"+1","created":"2011-07-25T18:33:52.903+0000"},{"body":"Committed.","created":"2011-07-25T18:58:42.902+0000"}],"conversations":[{"body":"For background, see CASSANDRA-2371","from":"reporter","subject":"Gossip should handle 'dead' states"},{"body":"The first patch allows gossip to track dead states and completely changes how removetoken works, since it is the problem with keeping dead state around, and currently broken in many scenarios. The second patch adds the previously reverted CASSANDRA-2115 back now that removetoken is more resilient.","from":"developer"},{"body":"Some explanation of what changed and why it was necessary:\n\nConsider nodes A through D. D is partitioned, and C is dead and needs to be removed. A removetoken will be issued to A for this.\n\nIn the current way we do things, A will modify it's own state by appending information to its status indicating that it will be removing C. B will see this, re-replicate as needed, then report to A that is is done. The problem however, is that since A modified its own state, A is also free to wipe that state out, either by restarting, or simple remove another token, because there's only space for one. If A reboots and then D's partition heals, D will never know C was removed. Worse, it will still have state for C that neither A nor B do, and so it will repopulate the ring with C again.\n\nThis patch changes this by instead having A sleep for RING_DELAY to make sure the generation for C is stable, and then it modifies C's state to indicate it is being removed, just as if C itself had done this. It also appends some extra state to indicate that A will be the removal coordinator. The others nodes see this, re-replicate and report back to A, which then modifies C's state once more to indicate it is completely removed. At this point, it doesn't matter if A dies completely and D's partition heals, since the state is stored in C's gossip information. If A reboots, it will be able to get the correct state information from B, or any other node.\n\nIf A fails while the other nodes are re-replicating, a new removetoken can be started elsewhere, or in the case of other replicas being down preventing removetoken from completing, a removetoken force will remove the node and then repair can be run to restore the replica count.","from":"developer"},{"body":"I see two more things to be done with this patch. First, when re-replicating nodes report back to the removal coordinator, if the coordinator has restarted it won't understand them, and they will infinitely loop retrying the confirmation. Second, since we're holding dead states, we need to make sure that bootstrapping/moving nodes can take over these dead tokens.","from":"developer"},{"body":"These small patches build on the others.\n\n0003-update-gossip-related-comments.patch.txt: updates gossip-related comments derp derp.","from":"developer"},{"body":"0004-do-REMOVING_TOKEN-REMOVED_TOKEN.patch.txt: use REMOVED_TOKEN instead of STATUS_LEFT (would probably be ok either way, but otherwise, the REMOVED_TOKEN state would not be used). Seems this is more the way it was intended.","from":"developer"},{"body":"0005-drain-self-if-removetoken-d-elsewhere.patch.txt : when node X was partitioned and removetoken'd but then it shows up again, it should shut itself down, rather than becoming a zombie","from":"developer"},{"body":"I'll see what I can do to test the \"infinitely loop retrying the confirmation\" and \"bootstrapping/moving nodes can take over these dead tokens\" situations.","from":"developer"},{"body":"Ok, nodes do indeed infinitely retry the replication confirmation in some cases, but it appears it's not just when the former removal coordinator has restarted in the interim- it seems to be when the removetoken is reissued to another, new removal coordinator. In this case, I get this traceback every 10 seconds:\n\n{noformat}\nERROR [MiscStage:9] 2011-07-20 23:42:06,599 AbstractCassandraDaemon.java (line 113) Fatal exception in thread Thread[MiscStage:9,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.service.StorageService.confirmReplication(StorageService.java:2088)\n at org.apache.cassandra.streaming.ReplicationFinishedVerbHandler.doVerb(ReplicationFinishedVerbHandler.java:38)\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:72)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n{noformat}\n\nI'll look into this.\n\nSecond, it seems that moving/joining nodes can take over the removed token fine, once the removetoken is complete. I haven't tried having a node take over the removed token while the removal is ongoing- I assume we can just document that that probably isn't a great idea?","from":"developer"},{"body":"0006-acknowledge-unexpected-repl-fins.patch.txt: don't assert and drop the message when we see an unexpected REPLICATION_FINISHED. Ack it instead, so the sender doesn't continually retry.","from":"developer"},{"body":"0006-acknowledge-unexpected-repl-fins.patch.txt (updated): also log at info when acknowledging the unexpected messages","from":"developer"},{"body":"ok, +1 with these patches.","from":"developer"},{"body":"0007 handles problems when a node has been down longer than aVeryLongTime, and ensures that we advertise the new token states long enough.\n\n0008 makes sure that SS only get involved with removal if the token is a member.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed.","from":"developer"}],"created":"2011-04-18T18:57:53.000+0000","description":"For background, see CASSANDRA-2371","issue_id":"12504582","key":"CASSANDRA-2496","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-07-25T18:58:42.000+0000","role":"fixed_distractor","summary":"Gossip should handle 'dead' states"} {"case_id":"12504930","cluster":"DISTRACTOR-CASSANDRA-2536","comments":[{"body":"I feel like a better, more critical sounding explanation, is: create a keyspace on node1, create a cf in that keyspace on node2 = hang + schema disagreement.","created":"2011-04-21T23:20:47.356+0000"},{"body":"And this is fine in 0.7?","created":"2011-04-22T01:37:04.529+0000"},{"body":"Actually, I can also reproduce this with a two node 0.7.4 cluster. I'm pretty sure that this does not happen with 0.7.3, but I'll go ahead and verify that.","created":"2011-04-22T16:03:49.807+0000"},{"body":"Nevermind my thoughts that it doesn't happen in 0.7.3 -- it seems to happen there too. It appears this is not a recent problem.","created":"2011-04-22T16:20:22.197+0000"},{"body":"Just bumped into this on a fresh 0.7.4 install on our test cluster. Does this only happen in a 2 node ring?","created":"2011-04-25T14:09:58.832+0000"},{"body":"encountered this on fresh 0.7.4 - 5 nodes - 100G+ per node. \n\ndecommissioning the bad node and rejoin fixed the problem.","created":"2011-04-27T18:38:05.810+0000"},{"body":"Gary, any thoughts on where to start looking?","created":"2011-04-28T02:10:29.671+0000"},{"body":"bq. any thoughts...\nI was going to add some jmx to get the last N schema versions (seems like it would be handy anyway and will be necessary if we ever get the rollback pony). Send schema to node A, verify that schema is propagated to B, send schema to B and watch the problem happen. The code to start looking at are the Definitions*VerbHandlers.\n\nSchema version is tracked in two places: gossip and in DatabaseDescriptor.defsVersion. Make sure those are reasonably in sync (was the sourced of one bug in the past). ","created":"2011-04-28T12:42:55.991+0000"},{"body":"The issue is the clocks being out of sync between nodes. Sometimes the v1 UUID generated by the second node has an earlier timestamp than the current schema UUID has.\n\nThere are a couple of things that could be fixed here:\n\n1. A node shouldn't accept a schema change if the timestamp for the new schema would be earlier than its current schema.\n2. Schema modification calls should accept an optional client-side timestamp that will be used for the v1 UUID.","created":"2011-04-28T20:58:51.932+0000"},{"body":"bq. Sometimes the v1 UUID generated by the second node has an earlier timestamp than the current schema UUID has.\nWouldn't that update be DOA then? I thought we checked to make sure the new migration compared after the current migration (as well as making sure the new migration's previous version matches with the current version).\n\nbq. A node shouldn't accept a schema change if the timestamp for the new schema would be earlier than its current schema.\nIf the clocks are *that* far off sync, I think the cluster has bigger problems (like writes not being applied). Plus, it would be easy for a node whose clock is way head to 'poison' schema updates from the rest of the cluster who are, in effect, behind the times.\n\nbq. Schema modification calls should accept an optional client-side timestamp that will be used for the v1 UUID.\nSeems like a better approach.\n\n","created":"2011-04-28T21:08:27.675+0000"},{"body":"bq. A node shouldn't accept a schema change if the timestamp for the new schema would be earlier than its current schema.\n\nYou need this with or without the client-side timestamp, though; there's no sense in letting people blow their leg off.\n\nAnd once you have that you don't need to add a client-side timestamp with all the PITA-ness that involves.\n\n(And unlike with data modification, I can't think of a use for doing \"clever\" things w/ a client side timesamp. So pushing it to the client doesn't really solve anything, just means you need to sync clocks across more machines.)","created":"2011-04-28T21:21:14.068+0000"},{"body":"{quote}\nWouldn't that update be DOA then? I thought we checked to make sure the new migration compared after the current migration (as well as making sure the new migration's previous version matches with the current version).\n{quote}\nWe do check that the previous version matches, but the migration is applied locally without comparing the current and new uuids.\n\n{quote}\nIf the clocks are that far off sync, I think the cluster has bigger problems (like writes not being applied).\n{quote}\nThis can theoretically happen with clocks being off by only tens of milliseconds.\n\n{quote}\nAnd unlike with data modification, I can't think of a use for doing \"clever\" things w/ a client side timesamp. So pushing it to the client doesn't really solve anything, just means you need to sync clocks across more machines.\n{quote}\nNot for clever purposes -- it seems to me that clients making schema modifications are more likely to be centralized, so schema changes coming from a single client will (almost) always have increasing timestamps.","created":"2011-04-28T21:33:33.927+0000"},{"body":"Attached patch compares version timestamps before applying migration locally.","created":"2011-04-28T22:26:27.175+0000"},{"body":"I personally think the timestamp comparison is good enough for now. Any interest in opening a new ticket for client-side timestamps?","created":"2011-04-28T23:08:39.469+0000"},{"body":"bq. it seems to me that clients making schema modifications are more likely to be centralized\n\nI would have also argued that they are likely to use the same connection (to the same server), and look where that got us. :)\n\nbq. I personally think the timestamp comparison is good enough for now\n\nI am okay with this. What do you think, Gary?\n\n(Nit: the exception message says \"older\" but the comparison is \"older or equal.\")","created":"2011-04-29T01:11:52.605+0000"},{"body":"I'll hijack this conversation by saying that I think we should start advertising that people should try to keep their server clocks in sync unless they have a good reason not too (which would legitimize the fact that \"timestamp comparison is good enough\"). Counter removes for instance use server side timestamps and would be screwed up by diverging clocks (and by that I mean more screwed up than they already are by design). And really, is there any reason not to install a ntpd server in the first place anyway?","created":"2011-04-29T07:38:44.246+0000"},{"body":"I think timestamp comparisons will be fine.","created":"2011-04-29T12:35:33.542+0000"},{"body":"committed, thanks!","created":"2011-04-29T15:38:01.525+0000"}],"conversations":[{"body":"If you have two thrift connections open to different nodes and you create a KS using the first, then a CF in that KS using the second, you wind up with a schema disagreement even if you wait/sleep after creating the KS.\n\nThe attached script reproduces the issue using pycassa (1.0.6 should work fine, although it has the 0.7 thrift-gen code). It's also reproducible by hand with two cassandra-cli sessions.","from":"reporter","subject":"Schema disagreements when using connections to multiple hosts"},{"body":"I feel like a better, more critical sounding explanation, is: create a keyspace on node1, create a cf in that keyspace on node2 = hang + schema disagreement.","from":"developer"},{"body":"And this is fine in 0.7?","from":"developer"},{"body":"Actually, I can also reproduce this with a two node 0.7.4 cluster. I'm pretty sure that this does not happen with 0.7.3, but I'll go ahead and verify that.","from":"developer"},{"body":"Nevermind my thoughts that it doesn't happen in 0.7.3 -- it seems to happen there too. It appears this is not a recent problem.","from":"developer"},{"body":"Just bumped into this on a fresh 0.7.4 install on our test cluster. Does this only happen in a 2 node ring?","from":"developer"},{"body":"encountered this on fresh 0.7.4 - 5 nodes - 100G+ per node. \n\ndecommissioning the bad node and rejoin fixed the problem.","from":"developer"},{"body":"Gary, any thoughts on where to start looking?","from":"developer"},{"body":"bq. any thoughts...\nI was going to add some jmx to get the last N schema versions (seems like it would be handy anyway and will be necessary if we ever get the rollback pony). Send schema to node A, verify that schema is propagated to B, send schema to B and watch the problem happen. The code to start looking at are the Definitions*VerbHandlers.\n\nSchema version is tracked in two places: gossip and in DatabaseDescriptor.defsVersion. Make sure those are reasonably in sync (was the sourced of one bug in the past). ","from":"developer"},{"body":"The issue is the clocks being out of sync between nodes. Sometimes the v1 UUID generated by the second node has an earlier timestamp than the current schema UUID has.\n\nThere are a couple of things that could be fixed here:\n\n1. A node shouldn't accept a schema change if the timestamp for the new schema would be earlier than its current schema.\n2. Schema modification calls should accept an optional client-side timestamp that will be used for the v1 UUID.","from":"developer"},{"body":"bq. Sometimes the v1 UUID generated by the second node has an earlier timestamp than the current schema UUID has.\nWouldn't that update be DOA then? I thought we checked to make sure the new migration compared after the current migration (as well as making sure the new migration's previous version matches with the current version).\n\nbq. A node shouldn't accept a schema change if the timestamp for the new schema would be earlier than its current schema.\nIf the clocks are *that* far off sync, I think the cluster has bigger problems (like writes not being applied). Plus, it would be easy for a node whose clock is way head to 'poison' schema updates from the rest of the cluster who are, in effect, behind the times.\n\nbq. Schema modification calls should accept an optional client-side timestamp that will be used for the v1 UUID.\nSeems like a better approach.\n\n","from":"developer"},{"body":"bq. A node shouldn't accept a schema change if the timestamp for the new schema would be earlier than its current schema.\n\nYou need this with or without the client-side timestamp, though; there's no sense in letting people blow their leg off.\n\nAnd once you have that you don't need to add a client-side timestamp with all the PITA-ness that involves.\n\n(And unlike with data modification, I can't think of a use for doing \"clever\" things w/ a client side timesamp. So pushing it to the client doesn't really solve anything, just means you need to sync clocks across more machines.)","from":"developer"},{"body":"{quote}\nWouldn't that update be DOA then? I thought we checked to make sure the new migration compared after the current migration (as well as making sure the new migration's previous version matches with the current version).\n{quote}\nWe do check that the previous version matches, but the migration is applied locally without comparing the current and new uuids.\n\n{quote}\nIf the clocks are that far off sync, I think the cluster has bigger problems (like writes not being applied).\n{quote}\nThis can theoretically happen with clocks being off by only tens of milliseconds.\n\n{quote}\nAnd unlike with data modification, I can't think of a use for doing \"clever\" things w/ a client side timesamp. So pushing it to the client doesn't really solve anything, just means you need to sync clocks across more machines.\n{quote}\nNot for clever purposes -- it seems to me that clients making schema modifications are more likely to be centralized, so schema changes coming from a single client will (almost) always have increasing timestamps.","from":"developer"},{"body":"Attached patch compares version timestamps before applying migration locally.","from":"developer"},{"body":"I personally think the timestamp comparison is good enough for now. Any interest in opening a new ticket for client-side timestamps?","from":"developer"},{"body":"bq. it seems to me that clients making schema modifications are more likely to be centralized\n\nI would have also argued that they are likely to use the same connection (to the same server), and look where that got us. :)\n\nbq. I personally think the timestamp comparison is good enough for now\n\nI am okay with this. What do you think, Gary?\n\n(Nit: the exception message says \"older\" but the comparison is \"older or equal.\")","from":"developer"},{"body":"I'll hijack this conversation by saying that I think we should start advertising that people should try to keep their server clocks in sync unless they have a good reason not too (which would legitimize the fact that \"timestamp comparison is good enough\"). Counter removes for instance use server side timestamps and would be screwed up by diverging clocks (and by that I mean more screwed up than they already are by design). And really, is there any reason not to install a ntpd server in the first place anyway?","from":"developer"},{"body":"I think timestamp comparisons will be fine.","from":"developer"},{"body":"committed, thanks!","from":"developer"}],"created":"2011-04-21T22:55:36.000+0000","description":"If you have two thrift connections open to different nodes and you create a KS using the first, then a CF in that KS using the second, you wind up with a schema disagreement even if you wait/sleep after creating the KS.\n\nThe attached script reproduces the issue using pycassa (1.0.6 should work fine, although it has the 0.7 thrift-gen code). It's also reproducible by hand with two cassandra-cli sessions.","issue_id":"12504930","key":"CASSANDRA-2536","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-04-29T15:38:01.000+0000","role":"fixed_distractor","summary":"Schema disagreements when using connections to multiple hosts"} {"case_id":"12509880","cluster":"DISTRACTOR-CASSANDRA-2761","comments":[{"body":"Moved from CASSANDRA-2754.\n\nJust a few clarifications as I am not that familiar with the C* infrastructure and build policies etc...\n\n- I'm a Maven guy, but I guess that is out of the question?\n- This needs to be Jenkins ready?\n- This should build Source and Javadocs artifacts ready for Maven Repo deployment right?\n- The resulting JDBC jar file would contain the C* dependancies from the build directory pointed to by the properties file/system property? It would be self contained?\n- The testing part of the build is going to have to reach into the C* build area too I think?","created":"2011-06-12T15:04:04.424+0000"},{"body":"bq. I'm a Maven guy, but I guess that is out of the question?\n\nI don't know, but if I were a Maven Guy (I am not), I would use Ant anyway. The vast majority of the active committers wish that Maven would die in a fire, so using it would be the fastest way to engender apathy. \n\nbq. This needs to be Jenkins ready?\n\nNot sure what this means.\n\nbq. This should build Source and Javadocs artifacts ready for Maven Repo deployment right?\n\nYes.\n\nbq. The resulting JDBC jar file would contain the C* dependancies from the build directory pointed to by the properties file/system property? It would be self contained?\n\nYes.\n\nbq. The testing part of the build is going to have to reach into the C* build area too I think?\n\nThe classpath for tests will probably contain a few additional items that the runtime classpath does not.","created":"2011-06-12T17:41:27.455+0000"},{"body":"OK, Patch for an ANT Build for the JDBC Driver suite is attached. \n\nIt builds all the artifacts, but at the moment it does NOT include any other classes but the ones in the o.a.c.cql.jdbc package.\n\nI will update when a few more clarifications of exactly what dependancies we want to declare for the driver and what we want to include in the jar.\n\nI think I got the Eclipse stuff in ther as well, but that could probably use some wringing out.. I was able to successfully run the junit tests from Eclipse.","created":"2011-06-13T03:57:04.913+0000"},{"body":"Thanks Rick, this is a good start, and a lot better than having no build at all, so I went ahead and committed it.\n\nI made a fewer minor changes (mostly replacing tabs for spaces), so you'll need to rebase against svn before continuing.\n\nbq. I will update when a few more clarifications of exactly what dependancies we want to declare for the driver and what we want to include in the jar.\n\nThere is probably some handy tool for this, but I don't know what it is. Just looking at it I'd say you need to copy-in:\n\n* {{o.a.c.db.marshall.*}}\n* {{o.a.c.utils.ByteBufferUtil}}\n* {{o.a.c.config.ColumnDefinition}}\n* {{o.a.c.config.CFMetaData}}\n* {{o.a.c.config.ConfigurationException}}\n\nThere might be some transitive dependencies though.","created":"2011-06-13T19:39:02.008+0000"},{"body":"Thanks Eric. \n\nActually you list is the same as mine so we are probably close. Note you will still need the whole C*-Thrift jar. I will build a new driver with the additional classes and test it without depending on the full C*. I'll document the full dependancy list here along with the next patch.","created":"2011-06-13T20:08:11.912+0000"},{"body":"bq. Note you will still need the whole C*-Thrift jar.\n\nYeah, but don't copy those. The driver jar will need to depend on the thrift jar.","created":"2011-06-13T20:12:25.376+0000"},{"body":"The v2 patch adds inclusion of classes that are required on the client side from the main C* build. ","created":"2011-06-15T15:27:18.794+0000"},{"body":"I put together a patch that covered the obvious but I had strange problems in test. So I dug deeper and found a ton (over 50 I think) additional transitive dependencies that I had missed... (silly me) \n\nI can try an get them all in but I think we should seriously rethink that strategy for now. With this many dependancies they should probably be put in their own jar(s). It would really make the driver jar have to keep up with detailed dependency changes in the server code base. Miss one and it's messy and the errors are non-obvious.\n\nI was a big whiner about not carrying around the whole server just for client access, but until it is re-factored, I can now appreciate the horror that will commence if we piece-meal drag over classes from the server into the JAR.\n\nSo fo now I suggest we keep the jar the way it is with just the o.a.c.cql.jdbc.* classes.","created":"2011-06-15T20:02:22.536+0000"},{"body":"{quote}\nI can try an get them all in but I think we should seriously rethink that strategy for now. With this many dependancies they should probably be put in their own jar(s). It would really make the driver jar have to keep up with detailed dependency changes in the server code base. Miss one and it's messy and the errors are non-obvious.\n{quote}\n\nWhat would you call such a jar? When I looked at this, there didn't seem to be any delineation that made sense.\n\n{quote}\nI was a big whiner about not carrying around the whole server just for client access, but until it is re-factored, I can now appreciate the horror that will commence if we piece-meal drag over classes from the server into the JAR.\n{quote}\n\nWhat are these 50 other class dependencies? Where are they being drug in?","created":"2011-06-16T16:17:40.350+0000"},{"body":"{quote}\nWhat are these 50 other class dependencies? Where are they being drug in?\n{quote}\n\n- o.a.c.db.marshall.*\n-- 23 various classes\n- o.a.c.utils.ByteBufferUtil\n-- org.apache.cassandra.io.util.FileDataInput\n-- org.apache.cassandra.io.util.FileUtils\n- o.a.c.config.ColumnDefinition\n- o.a.c.config.CFMetaData\n-- org.apache.cassandra.cache.IRowCacheProvider;\n-- org.apache.cassandra.db.migration.avro.ColumnDef\n-- org.apache.cassandra.db.ColumnFamilyType;\n--- org.apache.cassandra.concurrent.JMXEnabledThreadPoolExecutor\n--- org.apache.cassandra.config.DatabaseDescriptor\n---- org.apache.cassandra.auth.AllowAllAuthenticator\n---- org.apache.cassandra.auth.AllowAllAuthority\n---- org.apache.cassandra.auth.IAuthenticator\n---- org.apache.cassandra.auth.IAuthority\n---- org.apache.cassandra.config.Config.RequestSchedulerId\n---- org.apache.cassandra.db.ColumnFamilyStore\n----- 13 new items\n---- org.apache.cassandra.db.ColumnFamilyType\n---- org.apache.cassandra.db.DefsTable\n----- org.apache.cassandra.config.DatabaseDescriptor;\n----- org.apache.cassandra.config.KSMetaData;\n----- org.apache.cassandra.db.filter.QueryFilter;\n----- org.apache.cassandra.db.filter.QueryPath;\n----- org.apache.cassandra.db.migration.Migration;\n----- org.apache.cassandra.io.SerDeUtils;\n----- org.apache.cassandra.service.StorageService;\n----- org.apache.cassandra.utils.ByteBufferUtil;\n----- org.apache.cassandra.utils.UUIDGen;\n---- org.apache.cassandra.db.migration.Migration\n---- org.apache.cassandra.dht.IPartitioner\n---- org.apache.cassandra.io.sstable.Descriptor\n---- org.apache.cassandra.io.util.FileUtils\n---- org.apache.cassandra.locator.*\n---- org.apache.cassandra.scheduler.IRequestSchedule;\n---- org.apache.cassandra.scheduler.NoScheduler\n--- org.apache.cassandra.db.compaction.CompactionManager\n--- org.apache.cassandra.db.filter.QueryFilter\n--- org.apache.cassandra.db.filter.QueryPath\n--- org.apache.cassandra.dht.IPartitioner\n--- org.apache.cassandra.dht.Range\n--- org.apache.cassandra.gms.FailureDetector\n--- org.apache.cassandra.gms.Gossiper\n--- org.apache.cassandra.gms.ApplicationState\n--- org.apache.cassandra.net.MessagingService\n---- 10 new classes\n--- org.apache.cassandra.service.*\n--- org.apache.cassandra.utils.WrappedRunnable\n-- org.apache.cassandra.db.HintedHandOffManager;\n-- org.apache.cassandra.db.SystemTable;\n-- org.apache.cassandra.db.Table;\n-- org.apache.cassandra.db.ColumnFamilyStore;\n-- org.apache.cassandra.db.migration.Migration;\n-- org.apache.cassandra.db.compaction.AbstractCompactionStrategy;\n-- org.apache.cassandra.io.SerDeUtils;\n-- org.apache.cassandra.utils.Pair;\n- o.a.c.config.ConfigurationException\n\n\nThe point is there are lots and they are scattered all over the various packages; It will be very difficult to manage when they change from the driver package (client side), which is supposed to be able to change independent of the server code. If a subset of the server code is to be a dependency then that subset (jar/s) must be managed in the main build not the driver build. \n\n\n{quote}\nWhat would you call such a jar? When I looked at this, there didn't seem to be any delineation that made sense.\n{quote}\n\nI agree it is not any clear set of packages. They are scattered all over.\n\nAs to a name for the jar... I'm not a good namer in the best of circumstances but I think the intent is to pick those files that are used in common between client and server. I guess I'd use that as the basis for the name.\n\n\n","created":"2011-06-17T03:30:44.631+0000"},{"body":"{quote}\nThe point is there are lots and they are scattered all over the various packages; It will be very difficult to manage when they change from the driver package (client side), which is supposed to be able to change independent of the server code. If a subset of the server code is to be a dependency then that subset (jar/s) must be managed in the main build not the driver build.\n{quote}\n\nRight, I was curious to see the list of classes (that list is fantastic btw, thanks for that), to see if there was one point in the graph where breaking a dependency would drastically change the scope of the problem. It looks like the answer is Yes, and the dependency is {{o.a.c.config.CFMetaData}}, (needed by {{ColumnDecoder}}).\n\nJust skimming through the code, I don't think it would be hard to either re-implement the needed parts of CFMetaData, or refactor CFMetaData to limit what it pulled in.","created":"2011-06-17T15:31:21.770+0000"},{"body":"What's the status here? Not having \"ant test\" for the drivers anymore is causing me pain and suffering. :(","created":"2011-07-05T19:13:51.986+0000"},{"body":"Sorry. I didn't catch that there was no ant task to RUN them... :) {{build-test}} builds the included tests just fine. And I have been running them from Eclipse once built. I'll look into it. ","created":"2011-07-05T20:16:45.355+0000"},{"body":"Attached (v1-0001-CASSANDRA-2761-cleanup-nits.txt) is a patch to do a little cleanup (clearer variable naming, some simplification of classpaths, etc), basically stuff that set my OCD off and seemed better to get in *before* diving into test running and dependency management.","created":"2011-07-14T03:29:44.102+0000"},{"body":"To summarize, it is now possible to build and test w/ ant. This is currently done by pointing to a local (built) working copy of Cassandra (a site config). What's left, and seems reasonable to scope with this issue:\n\n* Create an alternate mechanism for specifying the version of Cassandra to build/test against (in order to run the tests against prior releases). I'm thinking Ivy could be used here to automatically download artifacts when a property is passed (-Dcassandra.release=0.8.0 for example).\n* (Re)build Cassandra as needed from the drivers Ant build, or at the very least, handle the case when a build is needed.\n* Fix the {{generate-eclipse-files}} target if possible, or remove it otherwise.\n\nWork should also continue to reduce the cross-section of Cassandra that this driver depends on, but I'll open another issue for that.","created":"2011-07-21T22:20:39.095+0000"},{"body":"+1 cleanup patch","created":"2011-07-22T00:27:45.329+0000"},{"body":"+1 for the cleanup patch\n\nThe {{generate-eclipse-files}} seems to be working for me? How does it fail?\n\n","created":"2011-07-22T02:07:45.853+0000"},{"body":"Note that CASSANDRA-3010 makes moving drivers out-of-\"tree\" even sillier: as things stand, we'll need to grab cassandra from the maven repo to build jdbc, then back in tree we'll need to grab jdbc from the maven repo to build the cql shell.\n\nI'm all for purity but at this point I'm ready to just put things back the way they were* so we can spend time before the 1.0 freeze building things instead of wrestling with maven.\n","created":"2011-08-10T15:35:10.124+0000"},{"body":"The lib/ directory is full of dependency jars, some which exist solely for the cli. ","created":"2011-08-10T15:53:10.636+0000"},{"body":"bq. I'm all for purity but at this point I'm ready to just put things back the way they were* so we can spend time before the 1.0 freeze building things instead of wrestling with maven.\n\nI for one would be delighted if I didn't have to use svn to get at the drivers or cqlsh.","created":"2011-08-10T16:01:47.141+0000"},{"body":"bq. The lib/ directory is full of dependency jars, some which exist solely for the cli.\n\nBut we are updating none of those on a daily or weekly basis.","created":"2011-08-10T16:04:30.858+0000"},{"body":"bq. But we are updating none of those on a daily or weekly basis\n\nI certainly hope we're not updating the JDBC driver that often, particularly in ways that would impact a simple application like a shell.\n\nOur experience using the JDBC driver as an application dependency shouldn't be any better or worse than it would be for other users.","created":"2011-08-10T16:41:09.211+0000"},{"body":"bq. I certainly hope we're not updating the JDBC driver that often\n\nThen you're not very familiar with the current stability or lack there of of the JDBC driver. :)","created":"2011-08-10T17:30:07.918+0000"},{"body":"bq. Then you're not very familiar with the current stability or lack there of of the JDBC driver. \n\nNo, I guess not, but it can't be any worse on us than for anyone else using it. A solution that only benefits us doesn't seem very... friendly.\n\n----\n\nI think the difference in opinion here comes from what determines whether the server and driver should be considered different projects. I can see where people who feel a high degree of ownership over the code of both, and who by virtue of infrastructure have write access to both, might consider it contrived to treat them as separate projects. I'm not discounting that point of view, but I do think that's more Social than Technical, and is limited to a relatively small group of Cassandra hackers.\n\nAs I've said elsewhere, I think it is very important that these drivers be allowed to evolve on their own release schedule with their own versioning that reflects compatibility with a CQL version and not any particular Cassandra version(s). It's also quite likely, particularly where the driver language != Java that the group of developers is entirely different from those working on the server. And there is no hard dependency between drivers and Cassandra in either direction (the JDBC->Cassandra dependency is one of convenience). To me, this pretty solidly points to them being separate projects.\n\nKeeping them separate won't be as convenient as treating them as one monolithic project, but it's no worse an experience than what other application developers are subjected to. We should be able to eat our own dog food.\n\nI also realize that there is more work needed here to decouple the JDBC driver and make this all work better, work that I've volunteered to do. I haven't had as much time to spend on this lately, but that should be changing RSN.\n\n","created":"2011-08-10T20:53:34.321+0000"},{"body":"Is their someway to get a git repo for just drivers? like cassandra.git but cassandra-drivers.git? This is causing major pain.","created":"2011-08-24T13:17:13.746+0000"},{"body":"bq. I also realize that there is more work needed here to decouple the JDBC driver and make this all work better, work that I've volunteered to do. I haven't had as much time to spend on this lately, but that should be changing RSN.\n\nIt's been another two weeks and we're still in this no-man's land where git can't see drivers, the jdbc build is a big TODO, and anything that depends on jdbc like 3010 is basically SOL.\n\nAgain, I'm not against purity, but it's clear that the original breaking out of drivers/ was done prematurely (was there even a Jira ticket with a patch? It looks like we went from \"that's a good idea\" on -dev to ripping apart svn). I've reverted things until the breakout can be done properly.","created":"2011-08-24T18:19:38.970+0000"},{"body":"I don't see any driver code in the move? The {{/driver}} sub-directory now exists but there is nothing underneath?","created":"2011-08-24T19:56:23.613+0000"},{"body":"https://svn.apache.org/repos/asf/cassandra/trunk/drivers/","created":"2011-08-24T20:10:11.969+0000"},{"body":"{quote}\nIt's been another two weeks and we're still in this no-man's land where git can't see drivers\n{quote}\n\nSo does this mean you've decided Git is hard a requirement? Our project is (unfortunately )managed in SVN. You know if it were up to me we'd be using Git for source control, but it's not up to me and we don't.\n\n{quote}\n...anything that depends on jdbc like 3010 is basically SOL.\n{quote}\n\nAs I've already mentioned, CASSANDAR-3010 is for a JDBC using _application_. If it's SOL unless the driver is embedded in the tree, then so is every other application that would make use of the driver. If the driver is so bad, maybe no one should be using it. Why are we giving ourselves special treatment?\n\n{quote}\nAgain, I'm not against purity...\n{quote}\n\nBy saying \"purity\" it sounds like you're dismissing it as something that's based on aesthetics. I assure that's not the case, it's about avoiding the inevitable lock-step relationship release-wise. \n\n{quote}\n...but it's clear that the original breaking out of drivers/ was done prematurely (was there even a Jira ticket with a patch?\n{quote}\n\nThere was a Jira yes, though I'm pretty sure it done as part of larger set of tasks, something with a description other than \"relocate drivers out of tree\". And, I don't think it was in a patch, no, due to the fact that a large move with some minor changes was (a) straightforward, and (b) would have needed rebasing several times a day.\n\n{quote}\nIt looks like we went from \"that's a good idea\" on -dev to ripping apart svn). I've reverted things until the breakout can be done properly.\n{quote}\n\nSo all I need to do is find a ticket and a +1 and I can -1 this, and revert your revert?\n","created":"2011-08-24T20:43:34.039+0000"},{"body":"bq. Our project is (unfortunately )managed in SVN\n\nputting a sub-project in a svn root isn't the way to do it for SVN.\n\nAnother issue is the drivers build depends on cassandra. That makes this even more strange. Can't drivers live in /trunk and still have it's own releases?\n\n","created":"2011-08-24T21:00:26.955+0000"},{"body":"bq. There was a Jira yes\n\nI couldn't find one, and the svn commit messages didn't mention it. I still can't find one.\n\nbq. So all I need to do is find a ticket and a +1 and I can -1 this, and revert your revert?\n\nThe point of making a ticket with a patch (or a shell script if you're throwing svn refactoring around -- you can do mv's from a working copy, not just on the repository, btw) is that people can review it and see if the result matches what they thought they were going to get out of it. In this case it's clear that it didn't, and we ended up with making things worse in exchange for a promise to clean it up eventually.\n\n(Of course, even with proper review, sometimes we've needed to revert things after unexpected problems arose. But skipping the review makes that more likely.)\n \nIt's been *over two months.* So let's reboot, and do it right. Note that this time around I got everything in one commit (well, one for trunk, and one for drivers) so \"reverting the revert\" will be easy.","created":"2011-08-24T21:13:22.007+0000"},{"body":"bq. putting a sub-project in a svn root isn't the way to do it for SVN.\n\nWhat is?\n\nbq. Another issue is the drivers build depends on cassandra. That makes this even more strange. Can't drivers live in /trunk and still have it's own releases?\n\nWe had it that way originally. It made releasing a pain, and left people with the expectation that the driver version to use was the one that corresponded to source it was with (which is unavoidable).","created":"2011-08-24T21:16:50.668+0000"},{"body":"bq. Can't drivers live in /trunk and still have it's own releases?\n\nRight, that seems like the best interim solution to me. (I thought we could even delete it from old branches to make it clear that it's not in lockstep with the rest of the tree, but Eric pointed out that if something like 3010 is depending on it, we can't do that. Still, having it present but \"frozen\" in the branches feels like a relatively minor downside.)","created":"2011-08-24T21:17:07.237+0000"},{"body":"How do we reboot and do it right? Your requirements as I understand them are structured in such a way that keeping them in-tree is the only solution. ","created":"2011-08-24T21:22:04.409+0000"},{"body":"bq. left people with the expectation that the driver version to use was the one that corresponded to source it was with\n\nI suspect this was when svn was effectively the only way to get a driver. My experience is that even most people building \"from source\" use a tarball, not svn. So the combination of making drivers available on maven, PyPI, etc., with their own version numbers, should address this.","created":"2011-08-24T21:22:05.945+0000"},{"body":"bq. Right, that seems like the best interim solution to me. (I thought we could even delete it from old branches to make it clear that it's not in lockstep with the rest of the tree, but Eric pointed out that if something like 3010 is depending on it, we can't do that. Still, having it present but \"frozen\" in the branches feels like a relatively minor downside.)\n\nExcept that people's expectation will always be that the version they need is the one that came with their software (which will likely be something pre-release). It is completely unrealistic to expect folks to just Know Better here. ","created":"2011-08-24T21:24:05.661+0000"},{"body":"bq. people's expectation will always be that the version they need is the one that came with their software\n\nI don't understand. What version will \"come with their [server]?\"\n\nConsider postgresql: the server is distributed on postgresql.org, and psycopg is distributed over PyPI. Nobody gets confused about staying on an obsolete version of psycopg.\n\nThat's the model I see us moving towards. (As you know, we recently got the Python cql driver on PyPI.)","created":"2011-08-24T21:38:45.472+0000"},{"body":"One final thought: by coincidence, we broke the JDBC build again today with CASSANDRA-3039. This isn't the first time this has happened. It's a minor win but not negligible to catch those before commit because \"ant test\" runs both suites.","created":"2011-08-24T21:52:04.424+0000"},{"body":"{quote}\nI don't understand. What version will \"come with their [server]?\"\n{quote}\n\nWherever we publish the source, be it an SVN checkout with an SVN revision ID, a date-based development snapshot, or a full-on release artifact, if there is driver source contained within then people are going to be encouraged to think of those drivers as being the same version as the node (e.g. 0.8.9). They're also going to be encouraged to think that those drivers are the best choice to use with the corresponding node. It's futile to think they won't, and having other vectors (www.a.o, PyPI, etc), will only add to the confusion. \n\n{quote}\nConsider postgresql: the server is distributed on postgresql.org, and psycopg is distributed over PyPI. Nobody gets confused about staying on an obsolete version of psycopg.\n\nThat's the model I see us moving towards. (As you know, we recently got the Python cql driver on PyPI.)\n{quote}\n\nYou make an excellent point. Pyscopg is maintained in a completely different repository (git://luna.dndg.it/public/psycopg2.git) than Postgesql (git://git.postgresql.org/git/postgresql.git) and is released (and published) separately, (which is in fact consistent with best practice elsewhere).","created":"2011-08-24T22:07:11.921+0000"},{"body":"bq. One final thought: by coincidence, we broke the JDBC build again today with CASSANDRA-3039. This isn't the first time this has happened. It's a minor win but not negligible to catch those before commit because \"ant test\" runs both suites.\n\nBroke why? Because of Cassandra code that the driver depends on? That would be an argument in favor of stabalizing the dependent code.\n\nAnd, the inverse of this is that the drivers should ultimately be getting tested against trunk, each active branch, and all past released versions (limited of course to CQL availability). That is only going to be practical through CI, and is made harder by your monolithic approach.","created":"2011-08-24T22:12:48.084+0000"},{"body":"bq. if there is driver source contained\n\nAlmost all non-developers get the source from a release tarball, not svn. Do we publish drivers/ in the source tarballs? If so that's easy enough to fix.\n\nAnyone actually using svn I'm willing to educate. You can give them my email. :)\n\nbq. Pyscopg is maintained in a completely different repository\n\nBut as far as users are concerned that is irrelevant.\n\nIf you want examples of drivers in the same tree that are also not causing confusion, I can point you to https://github.com/mongodb/mongo and https://github.com/mongodb/mongo/tree/master/client. Also http://hg.basho.com/riak/src/5ffa6ae7e699 (http://hg.basho.com/riak/src/5ffa6ae7e699/client_lib/) and https://github.com/voldemort/voldemort (https://github.com/voldemort/voldemort/tree/master/clients).","created":"2011-08-24T22:19:09.425+0000"},{"body":"So to summarize: You've decided.\n\nI know that sounds a bit snarky, but you summarily reverted the change, in part based on reasoning that wasn't true (it doesn't build), and in part pending a \"reboot\" to meet unmeetable requirements.","created":"2011-08-24T22:44:39.387+0000"},{"body":"If waiting for two months of \"we'll fix it real soon now\" is \"summarily,\" then yeah, I guess guilty as charged.\n\nBut it looks like you've restarted discussion on -dev, so I'll move there.","created":"2011-08-25T02:27:33.079+0000"},{"body":"Up until the July 21st, the scope of the ticket was building and running the unit tests (and as of July 21st that much was working). It wasn't until Aug 10th (and the creation of CASSANDRA-3010) that the discussion (and scope) changed. Between the 10th and today there was no response to my last, then 4 hours after Jake's comment the revert was made.\n\nI already told you once today that I ascribed no malice in this, and I don't, but I think \"summarily\" is a fair assessment.","created":"2011-08-25T02:42:17.769+0000"}],"conversations":[{"body":"Need a way to build (and run tests for) the Java driver.\n\nAlso: still some vestigal references to drivers/ in trunk build.xml.\n\nShould we remove drivers/ from the 0.8 branch as well?","from":"reporter","subject":"JDBC driver does not build"},{"body":"Moved from CASSANDRA-2754.\n\nJust a few clarifications as I am not that familiar with the C* infrastructure and build policies etc...\n\n- I'm a Maven guy, but I guess that is out of the question?\n- This needs to be Jenkins ready?\n- This should build Source and Javadocs artifacts ready for Maven Repo deployment right?\n- The resulting JDBC jar file would contain the C* dependancies from the build directory pointed to by the properties file/system property? It would be self contained?\n- The testing part of the build is going to have to reach into the C* build area too I think?","from":"developer"},{"body":"bq. I'm a Maven guy, but I guess that is out of the question?\n\nI don't know, but if I were a Maven Guy (I am not), I would use Ant anyway. The vast majority of the active committers wish that Maven would die in a fire, so using it would be the fastest way to engender apathy. \n\nbq. This needs to be Jenkins ready?\n\nNot sure what this means.\n\nbq. This should build Source and Javadocs artifacts ready for Maven Repo deployment right?\n\nYes.\n\nbq. The resulting JDBC jar file would contain the C* dependancies from the build directory pointed to by the properties file/system property? It would be self contained?\n\nYes.\n\nbq. The testing part of the build is going to have to reach into the C* build area too I think?\n\nThe classpath for tests will probably contain a few additional items that the runtime classpath does not.","from":"developer"},{"body":"OK, Patch for an ANT Build for the JDBC Driver suite is attached. \n\nIt builds all the artifacts, but at the moment it does NOT include any other classes but the ones in the o.a.c.cql.jdbc package.\n\nI will update when a few more clarifications of exactly what dependancies we want to declare for the driver and what we want to include in the jar.\n\nI think I got the Eclipse stuff in ther as well, but that could probably use some wringing out.. I was able to successfully run the junit tests from Eclipse.","from":"developer"},{"body":"Thanks Rick, this is a good start, and a lot better than having no build at all, so I went ahead and committed it.\n\nI made a fewer minor changes (mostly replacing tabs for spaces), so you'll need to rebase against svn before continuing.\n\nbq. I will update when a few more clarifications of exactly what dependancies we want to declare for the driver and what we want to include in the jar.\n\nThere is probably some handy tool for this, but I don't know what it is. Just looking at it I'd say you need to copy-in:\n\n* {{o.a.c.db.marshall.*}}\n* {{o.a.c.utils.ByteBufferUtil}}\n* {{o.a.c.config.ColumnDefinition}}\n* {{o.a.c.config.CFMetaData}}\n* {{o.a.c.config.ConfigurationException}}\n\nThere might be some transitive dependencies though.","from":"developer"},{"body":"Thanks Eric. \n\nActually you list is the same as mine so we are probably close. Note you will still need the whole C*-Thrift jar. I will build a new driver with the additional classes and test it without depending on the full C*. I'll document the full dependancy list here along with the next patch.","from":"developer"},{"body":"bq. Note you will still need the whole C*-Thrift jar.\n\nYeah, but don't copy those. The driver jar will need to depend on the thrift jar.","from":"developer"},{"body":"The v2 patch adds inclusion of classes that are required on the client side from the main C* build. ","from":"developer"},{"body":"I put together a patch that covered the obvious but I had strange problems in test. So I dug deeper and found a ton (over 50 I think) additional transitive dependencies that I had missed... (silly me) \n\nI can try an get them all in but I think we should seriously rethink that strategy for now. With this many dependancies they should probably be put in their own jar(s). It would really make the driver jar have to keep up with detailed dependency changes in the server code base. Miss one and it's messy and the errors are non-obvious.\n\nI was a big whiner about not carrying around the whole server just for client access, but until it is re-factored, I can now appreciate the horror that will commence if we piece-meal drag over classes from the server into the JAR.\n\nSo fo now I suggest we keep the jar the way it is with just the o.a.c.cql.jdbc.* classes.","from":"developer"},{"body":"{quote}\nI can try an get them all in but I think we should seriously rethink that strategy for now. With this many dependancies they should probably be put in their own jar(s). It would really make the driver jar have to keep up with detailed dependency changes in the server code base. Miss one and it's messy and the errors are non-obvious.\n{quote}\n\nWhat would you call such a jar? When I looked at this, there didn't seem to be any delineation that made sense.\n\n{quote}\nI was a big whiner about not carrying around the whole server just for client access, but until it is re-factored, I can now appreciate the horror that will commence if we piece-meal drag over classes from the server into the JAR.\n{quote}\n\nWhat are these 50 other class dependencies? Where are they being drug in?","from":"developer"},{"body":"{quote}\nWhat are these 50 other class dependencies? Where are they being drug in?\n{quote}\n\n- o.a.c.db.marshall.*\n-- 23 various classes\n- o.a.c.utils.ByteBufferUtil\n-- org.apache.cassandra.io.util.FileDataInput\n-- org.apache.cassandra.io.util.FileUtils\n- o.a.c.config.ColumnDefinition\n- o.a.c.config.CFMetaData\n-- org.apache.cassandra.cache.IRowCacheProvider;\n-- org.apache.cassandra.db.migration.avro.ColumnDef\n-- org.apache.cassandra.db.ColumnFamilyType;\n--- org.apache.cassandra.concurrent.JMXEnabledThreadPoolExecutor\n--- org.apache.cassandra.config.DatabaseDescriptor\n---- org.apache.cassandra.auth.AllowAllAuthenticator\n---- org.apache.cassandra.auth.AllowAllAuthority\n---- org.apache.cassandra.auth.IAuthenticator\n---- org.apache.cassandra.auth.IAuthority\n---- org.apache.cassandra.config.Config.RequestSchedulerId\n---- org.apache.cassandra.db.ColumnFamilyStore\n----- 13 new items\n---- org.apache.cassandra.db.ColumnFamilyType\n---- org.apache.cassandra.db.DefsTable\n----- org.apache.cassandra.config.DatabaseDescriptor;\n----- org.apache.cassandra.config.KSMetaData;\n----- org.apache.cassandra.db.filter.QueryFilter;\n----- org.apache.cassandra.db.filter.QueryPath;\n----- org.apache.cassandra.db.migration.Migration;\n----- org.apache.cassandra.io.SerDeUtils;\n----- org.apache.cassandra.service.StorageService;\n----- org.apache.cassandra.utils.ByteBufferUtil;\n----- org.apache.cassandra.utils.UUIDGen;\n---- org.apache.cassandra.db.migration.Migration\n---- org.apache.cassandra.dht.IPartitioner\n---- org.apache.cassandra.io.sstable.Descriptor\n---- org.apache.cassandra.io.util.FileUtils\n---- org.apache.cassandra.locator.*\n---- org.apache.cassandra.scheduler.IRequestSchedule;\n---- org.apache.cassandra.scheduler.NoScheduler\n--- org.apache.cassandra.db.compaction.CompactionManager\n--- org.apache.cassandra.db.filter.QueryFilter\n--- org.apache.cassandra.db.filter.QueryPath\n--- org.apache.cassandra.dht.IPartitioner\n--- org.apache.cassandra.dht.Range\n--- org.apache.cassandra.gms.FailureDetector\n--- org.apache.cassandra.gms.Gossiper\n--- org.apache.cassandra.gms.ApplicationState\n--- org.apache.cassandra.net.MessagingService\n---- 10 new classes\n--- org.apache.cassandra.service.*\n--- org.apache.cassandra.utils.WrappedRunnable\n-- org.apache.cassandra.db.HintedHandOffManager;\n-- org.apache.cassandra.db.SystemTable;\n-- org.apache.cassandra.db.Table;\n-- org.apache.cassandra.db.ColumnFamilyStore;\n-- org.apache.cassandra.db.migration.Migration;\n-- org.apache.cassandra.db.compaction.AbstractCompactionStrategy;\n-- org.apache.cassandra.io.SerDeUtils;\n-- org.apache.cassandra.utils.Pair;\n- o.a.c.config.ConfigurationException\n\n\nThe point is there are lots and they are scattered all over the various packages; It will be very difficult to manage when they change from the driver package (client side), which is supposed to be able to change independent of the server code. If a subset of the server code is to be a dependency then that subset (jar/s) must be managed in the main build not the driver build. \n\n\n{quote}\nWhat would you call such a jar? When I looked at this, there didn't seem to be any delineation that made sense.\n{quote}\n\nI agree it is not any clear set of packages. They are scattered all over.\n\nAs to a name for the jar... I'm not a good namer in the best of circumstances but I think the intent is to pick those files that are used in common between client and server. I guess I'd use that as the basis for the name.\n\n\n","from":"developer"},{"body":"{quote}\nThe point is there are lots and they are scattered all over the various packages; It will be very difficult to manage when they change from the driver package (client side), which is supposed to be able to change independent of the server code. If a subset of the server code is to be a dependency then that subset (jar/s) must be managed in the main build not the driver build.\n{quote}\n\nRight, I was curious to see the list of classes (that list is fantastic btw, thanks for that), to see if there was one point in the graph where breaking a dependency would drastically change the scope of the problem. It looks like the answer is Yes, and the dependency is {{o.a.c.config.CFMetaData}}, (needed by {{ColumnDecoder}}).\n\nJust skimming through the code, I don't think it would be hard to either re-implement the needed parts of CFMetaData, or refactor CFMetaData to limit what it pulled in.","from":"developer"},{"body":"What's the status here? Not having \"ant test\" for the drivers anymore is causing me pain and suffering. :(","from":"developer"},{"body":"Sorry. I didn't catch that there was no ant task to RUN them... :) {{build-test}} builds the included tests just fine. And I have been running them from Eclipse once built. I'll look into it. ","from":"developer"},{"body":"Attached (v1-0001-CASSANDRA-2761-cleanup-nits.txt) is a patch to do a little cleanup (clearer variable naming, some simplification of classpaths, etc), basically stuff that set my OCD off and seemed better to get in *before* diving into test running and dependency management.","from":"developer"},{"body":"To summarize, it is now possible to build and test w/ ant. This is currently done by pointing to a local (built) working copy of Cassandra (a site config). What's left, and seems reasonable to scope with this issue:\n\n* Create an alternate mechanism for specifying the version of Cassandra to build/test against (in order to run the tests against prior releases). I'm thinking Ivy could be used here to automatically download artifacts when a property is passed (-Dcassandra.release=0.8.0 for example).\n* (Re)build Cassandra as needed from the drivers Ant build, or at the very least, handle the case when a build is needed.\n* Fix the {{generate-eclipse-files}} target if possible, or remove it otherwise.\n\nWork should also continue to reduce the cross-section of Cassandra that this driver depends on, but I'll open another issue for that.","from":"developer"},{"body":"+1 cleanup patch","from":"developer"},{"body":"+1 for the cleanup patch\n\nThe {{generate-eclipse-files}} seems to be working for me? How does it fail?\n\n","from":"developer"},{"body":"Note that CASSANDRA-3010 makes moving drivers out-of-\"tree\" even sillier: as things stand, we'll need to grab cassandra from the maven repo to build jdbc, then back in tree we'll need to grab jdbc from the maven repo to build the cql shell.\n\nI'm all for purity but at this point I'm ready to just put things back the way they were* so we can spend time before the 1.0 freeze building things instead of wrestling with maven.\n","from":"developer"},{"body":"The lib/ directory is full of dependency jars, some which exist solely for the cli. ","from":"developer"},{"body":"bq. I'm all for purity but at this point I'm ready to just put things back the way they were* so we can spend time before the 1.0 freeze building things instead of wrestling with maven.\n\nI for one would be delighted if I didn't have to use svn to get at the drivers or cqlsh.","from":"developer"},{"body":"bq. The lib/ directory is full of dependency jars, some which exist solely for the cli.\n\nBut we are updating none of those on a daily or weekly basis.","from":"developer"},{"body":"bq. But we are updating none of those on a daily or weekly basis\n\nI certainly hope we're not updating the JDBC driver that often, particularly in ways that would impact a simple application like a shell.\n\nOur experience using the JDBC driver as an application dependency shouldn't be any better or worse than it would be for other users.","from":"developer"},{"body":"bq. I certainly hope we're not updating the JDBC driver that often\n\nThen you're not very familiar with the current stability or lack there of of the JDBC driver. :)","from":"developer"},{"body":"bq. Then you're not very familiar with the current stability or lack there of of the JDBC driver. \n\nNo, I guess not, but it can't be any worse on us than for anyone else using it. A solution that only benefits us doesn't seem very... friendly.\n\n----\n\nI think the difference in opinion here comes from what determines whether the server and driver should be considered different projects. I can see where people who feel a high degree of ownership over the code of both, and who by virtue of infrastructure have write access to both, might consider it contrived to treat them as separate projects. I'm not discounting that point of view, but I do think that's more Social than Technical, and is limited to a relatively small group of Cassandra hackers.\n\nAs I've said elsewhere, I think it is very important that these drivers be allowed to evolve on their own release schedule with their own versioning that reflects compatibility with a CQL version and not any particular Cassandra version(s). It's also quite likely, particularly where the driver language != Java that the group of developers is entirely different from those working on the server. And there is no hard dependency between drivers and Cassandra in either direction (the JDBC->Cassandra dependency is one of convenience). To me, this pretty solidly points to them being separate projects.\n\nKeeping them separate won't be as convenient as treating them as one monolithic project, but it's no worse an experience than what other application developers are subjected to. We should be able to eat our own dog food.\n\nI also realize that there is more work needed here to decouple the JDBC driver and make this all work better, work that I've volunteered to do. I haven't had as much time to spend on this lately, but that should be changing RSN.\n\n","from":"developer"},{"body":"Is their someway to get a git repo for just drivers? like cassandra.git but cassandra-drivers.git? This is causing major pain.","from":"developer"},{"body":"bq. I also realize that there is more work needed here to decouple the JDBC driver and make this all work better, work that I've volunteered to do. I haven't had as much time to spend on this lately, but that should be changing RSN.\n\nIt's been another two weeks and we're still in this no-man's land where git can't see drivers, the jdbc build is a big TODO, and anything that depends on jdbc like 3010 is basically SOL.\n\nAgain, I'm not against purity, but it's clear that the original breaking out of drivers/ was done prematurely (was there even a Jira ticket with a patch? It looks like we went from \"that's a good idea\" on -dev to ripping apart svn). I've reverted things until the breakout can be done properly.","from":"developer"},{"body":"I don't see any driver code in the move? The {{/driver}} sub-directory now exists but there is nothing underneath?","from":"developer"},{"body":"https://svn.apache.org/repos/asf/cassandra/trunk/drivers/","from":"developer"},{"body":"{quote}\nIt's been another two weeks and we're still in this no-man's land where git can't see drivers\n{quote}\n\nSo does this mean you've decided Git is hard a requirement? Our project is (unfortunately )managed in SVN. You know if it were up to me we'd be using Git for source control, but it's not up to me and we don't.\n\n{quote}\n...anything that depends on jdbc like 3010 is basically SOL.\n{quote}\n\nAs I've already mentioned, CASSANDAR-3010 is for a JDBC using _application_. If it's SOL unless the driver is embedded in the tree, then so is every other application that would make use of the driver. If the driver is so bad, maybe no one should be using it. Why are we giving ourselves special treatment?\n\n{quote}\nAgain, I'm not against purity...\n{quote}\n\nBy saying \"purity\" it sounds like you're dismissing it as something that's based on aesthetics. I assure that's not the case, it's about avoiding the inevitable lock-step relationship release-wise. \n\n{quote}\n...but it's clear that the original breaking out of drivers/ was done prematurely (was there even a Jira ticket with a patch?\n{quote}\n\nThere was a Jira yes, though I'm pretty sure it done as part of larger set of tasks, something with a description other than \"relocate drivers out of tree\". And, I don't think it was in a patch, no, due to the fact that a large move with some minor changes was (a) straightforward, and (b) would have needed rebasing several times a day.\n\n{quote}\nIt looks like we went from \"that's a good idea\" on -dev to ripping apart svn). I've reverted things until the breakout can be done properly.\n{quote}\n\nSo all I need to do is find a ticket and a +1 and I can -1 this, and revert your revert?\n","from":"developer"},{"body":"bq. Our project is (unfortunately )managed in SVN\n\nputting a sub-project in a svn root isn't the way to do it for SVN.\n\nAnother issue is the drivers build depends on cassandra. That makes this even more strange. Can't drivers live in /trunk and still have it's own releases?\n\n","from":"developer"},{"body":"bq. There was a Jira yes\n\nI couldn't find one, and the svn commit messages didn't mention it. I still can't find one.\n\nbq. So all I need to do is find a ticket and a +1 and I can -1 this, and revert your revert?\n\nThe point of making a ticket with a patch (or a shell script if you're throwing svn refactoring around -- you can do mv's from a working copy, not just on the repository, btw) is that people can review it and see if the result matches what they thought they were going to get out of it. In this case it's clear that it didn't, and we ended up with making things worse in exchange for a promise to clean it up eventually.\n\n(Of course, even with proper review, sometimes we've needed to revert things after unexpected problems arose. But skipping the review makes that more likely.)\n \nIt's been *over two months.* So let's reboot, and do it right. Note that this time around I got everything in one commit (well, one for trunk, and one for drivers) so \"reverting the revert\" will be easy.","from":"developer"},{"body":"bq. putting a sub-project in a svn root isn't the way to do it for SVN.\n\nWhat is?\n\nbq. Another issue is the drivers build depends on cassandra. That makes this even more strange. Can't drivers live in /trunk and still have it's own releases?\n\nWe had it that way originally. It made releasing a pain, and left people with the expectation that the driver version to use was the one that corresponded to source it was with (which is unavoidable).","from":"developer"},{"body":"bq. Can't drivers live in /trunk and still have it's own releases?\n\nRight, that seems like the best interim solution to me. (I thought we could even delete it from old branches to make it clear that it's not in lockstep with the rest of the tree, but Eric pointed out that if something like 3010 is depending on it, we can't do that. Still, having it present but \"frozen\" in the branches feels like a relatively minor downside.)","from":"developer"},{"body":"How do we reboot and do it right? Your requirements as I understand them are structured in such a way that keeping them in-tree is the only solution. ","from":"developer"},{"body":"bq. left people with the expectation that the driver version to use was the one that corresponded to source it was with\n\nI suspect this was when svn was effectively the only way to get a driver. My experience is that even most people building \"from source\" use a tarball, not svn. So the combination of making drivers available on maven, PyPI, etc., with their own version numbers, should address this.","from":"developer"},{"body":"bq. Right, that seems like the best interim solution to me. (I thought we could even delete it from old branches to make it clear that it's not in lockstep with the rest of the tree, but Eric pointed out that if something like 3010 is depending on it, we can't do that. Still, having it present but \"frozen\" in the branches feels like a relatively minor downside.)\n\nExcept that people's expectation will always be that the version they need is the one that came with their software (which will likely be something pre-release). It is completely unrealistic to expect folks to just Know Better here. ","from":"developer"},{"body":"bq. people's expectation will always be that the version they need is the one that came with their software\n\nI don't understand. What version will \"come with their [server]?\"\n\nConsider postgresql: the server is distributed on postgresql.org, and psycopg is distributed over PyPI. Nobody gets confused about staying on an obsolete version of psycopg.\n\nThat's the model I see us moving towards. (As you know, we recently got the Python cql driver on PyPI.)","from":"developer"},{"body":"One final thought: by coincidence, we broke the JDBC build again today with CASSANDRA-3039. This isn't the first time this has happened. It's a minor win but not negligible to catch those before commit because \"ant test\" runs both suites.","from":"developer"},{"body":"{quote}\nI don't understand. What version will \"come with their [server]?\"\n{quote}\n\nWherever we publish the source, be it an SVN checkout with an SVN revision ID, a date-based development snapshot, or a full-on release artifact, if there is driver source contained within then people are going to be encouraged to think of those drivers as being the same version as the node (e.g. 0.8.9). They're also going to be encouraged to think that those drivers are the best choice to use with the corresponding node. It's futile to think they won't, and having other vectors (www.a.o, PyPI, etc), will only add to the confusion. \n\n{quote}\nConsider postgresql: the server is distributed on postgresql.org, and psycopg is distributed over PyPI. Nobody gets confused about staying on an obsolete version of psycopg.\n\nThat's the model I see us moving towards. (As you know, we recently got the Python cql driver on PyPI.)\n{quote}\n\nYou make an excellent point. Pyscopg is maintained in a completely different repository (git://luna.dndg.it/public/psycopg2.git) than Postgesql (git://git.postgresql.org/git/postgresql.git) and is released (and published) separately, (which is in fact consistent with best practice elsewhere).","from":"developer"},{"body":"bq. One final thought: by coincidence, we broke the JDBC build again today with CASSANDRA-3039. This isn't the first time this has happened. It's a minor win but not negligible to catch those before commit because \"ant test\" runs both suites.\n\nBroke why? Because of Cassandra code that the driver depends on? That would be an argument in favor of stabalizing the dependent code.\n\nAnd, the inverse of this is that the drivers should ultimately be getting tested against trunk, each active branch, and all past released versions (limited of course to CQL availability). That is only going to be practical through CI, and is made harder by your monolithic approach.","from":"developer"},{"body":"bq. if there is driver source contained\n\nAlmost all non-developers get the source from a release tarball, not svn. Do we publish drivers/ in the source tarballs? If so that's easy enough to fix.\n\nAnyone actually using svn I'm willing to educate. You can give them my email. :)\n\nbq. Pyscopg is maintained in a completely different repository\n\nBut as far as users are concerned that is irrelevant.\n\nIf you want examples of drivers in the same tree that are also not causing confusion, I can point you to https://github.com/mongodb/mongo and https://github.com/mongodb/mongo/tree/master/client. Also http://hg.basho.com/riak/src/5ffa6ae7e699 (http://hg.basho.com/riak/src/5ffa6ae7e699/client_lib/) and https://github.com/voldemort/voldemort (https://github.com/voldemort/voldemort/tree/master/clients).","from":"developer"},{"body":"So to summarize: You've decided.\n\nI know that sounds a bit snarky, but you summarily reverted the change, in part based on reasoning that wasn't true (it doesn't build), and in part pending a \"reboot\" to meet unmeetable requirements.","from":"developer"},{"body":"If waiting for two months of \"we'll fix it real soon now\" is \"summarily,\" then yeah, I guess guilty as charged.\n\nBut it looks like you've restarted discussion on -dev, so I'll move there.","from":"developer"},{"body":"Up until the July 21st, the scope of the ticket was building and running the unit tests (and as of July 21st that much was working). It wasn't until Aug 10th (and the creation of CASSANDRA-3010) that the discussion (and scope) changed. Between the 10th and today there was no response to my last, then 4 hours after Jake's comment the revert was made.\n\nI already told you once today that I ascribed no malice in this, and I don't, but I think \"summarily\" is a fair assessment.","from":"developer"}],"created":"2011-06-10T21:36:38.000+0000","description":"Need a way to build (and run tests for) the Java driver.\n\nAlso: still some vestigal references to drivers/ in trunk build.xml.\n\nShould we remove drivers/ from the 0.8 branch as well?","issue_id":"12509880","key":"CASSANDRA-2761","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-08-25T14:27:55.000+0000","role":"fixed_distractor","summary":"JDBC driver does not build"} {"case_id":"12510393","cluster":"DISTRACTOR-CASSANDRA-2773","comments":[{"body":"Like it says, you will have to send your deletes and mutations in different mutations.","created":"2011-06-15T03:33:29.608+0000"},{"body":"sure, I know what you mean.\n\nhowever, when the user uses it the wrong way, the server should at least response an error or exception... not just hold the connection and do nothing...\n\nIn addition, the server should not be dead forever, when this kind of exception occurs. ","created":"2011-06-15T03:45:46.026+0000"},{"body":"You're right, there's a validation problem here.","created":"2011-06-15T03:51:07.434+0000"},{"body":"we could add some valdation logic but it looks like it's almost as easy to just remove this limitation. patch to do this attached.","created":"2011-06-17T17:17:30.078+0000"},{"body":"Hum, we cannot remove the column from cf in ignoreObsoleteMutations() because cf is the original column family from the row mutation and that's racy with commit log write (à la CASSANDRA-2604). We should clone the column family, but maybe it's simpler to add validation logic after all ? In any case, it could be worth it adding some comment in Table.apply() or Table.ignoreObsoleteMutations(). ","created":"2011-06-20T09:19:58.028+0000"},{"body":"bq. we cannot remove the column from cf in ignoreObsoleteMutations()\n\nYou're right. Fortunately I don't think that's actually necessary. v2 attached.","created":"2011-06-20T16:49:14.091+0000"},{"body":"Right, +1. I still think we should add a comment somewhere saying we shouldn't change cf (and that there is no reason to change it anyway).","created":"2011-06-20T16:53:32.356+0000"},{"body":"Committed, with comment. (I wish CF objects could be immutable AND efficient...)","created":"2011-06-23T16:00:59.951+0000"},{"body":"We are experiencing this issue on a cluster running 0.7.6-2. Once this occurs on a server, the same message is repeated over and over again in system.log. When we attempted to restart the box it failed. I will attach the log to this issue.","created":"2011-06-23T21:32:28.164+0000"},{"body":"Log of failed restart attempt","created":"2011-06-23T21:34:46.015+0000"},{"body":"Marked as critical, because when this occurs the server will not restart without removing commitlog and attending loss of data.","created":"2011-06-24T17:04:55.809+0000"},{"body":"I explained on-list that I'm not comfortable committing this to 0.7. The critical thing we need to deliver in 0.7 at this stage is stability. \n\nThe rarity of the problem combined with the complexity of the code involved leads me to conclude that it's better to live with a bug we know how to avoid, than risk introducing new ones. Remember, the risk of new bugs affects _everyone_, while the fix can only benefit those who were generating the unusual mutation pattern here. It's our responsibility to take a balanced view.\n\nFor the rare people who do find themselves affected here, your options include\n- stay on < 0.7.6\n- drain the commitlog with an earlier version before upgrading, then re-upgrade to 0.7.6 after fixing your code to not generate problematic mutations\n- run an unofficial build with this patch included\n- upgrade to 0.8.2 when released","created":"2011-06-24T17:55:05.573+0000"},{"body":"I disagree. This is a bug which allows a client to make a cluster unresponsive by performing a seemingly innocuous series of operations. If that happens, the cluster is un-restartable without loss of data. I wouldn't call a release where this can occur \"stable\". So if the goal for 0.7 is stability...\n\nWRT \"fixing your code to not generate problematic mutations,\" this may be difficult to do. I have so far identified code that does deletes followed by updates in the same mutation, but I haven't yet found any updates followed by deletes. Are we sure that only the update-followed-by-delete scenario is problematic?\n\nIn any case, even after reviewing all our code for the relevant scenarios, I would not feel comfortable deploying an 0.7 release with this vulnerability to production. The risk of a catastrophic failure is too great.\n","created":"2011-06-24T19:03:21.511+0000"},{"body":"Actually, there isn't really much risk for data loss given that as Jonathan said, if you hit that, it's fairly easy to go back to 0.7.5, fix client code and upgrade again. Granted this is not user friendly and not something you should expect from a minor upgrade, but let at least set the record straight on the data loss part.\n\nThat being said, I don't think the patch on this ticket could screw up indexes more that we use to prior to 0.7.6, so maybe we can commit to 0.7.7 on that ground.\n\nI'd still suggest fixing client code in the meantime. ","created":"2011-06-27T15:54:14.712+0000"},{"body":"bq. I don't think the patch on this ticket could screw up indexes more that we use to prior to 0.7.6\n\nThat's a valid way to frame the issue.\n\nI'm good to commit for 0.7.7 if Jim can test the patch first, since he's the only one we've heard of hitting this in 0.7.x. (Specifically, we want to make sure that if we query \"WHERE foo = X\" we don't get results back where foo is something other than X. Ideally you'd start with an empty database, or at least drop + recreate indexes first to make sure the results aren't contaminated w/ corrupt entries from pre-0.7.6.)","created":"2011-06-27T16:02:32.503+0000"},{"body":"I applied the 0.8 patch and added a couple of tests to ColumnFamilyStoreTest. The tests trigger the UnsupportedOperationException in 0.7.6 and return the correct values with the patch applied. Would you like me to test the same scenario against an actual patched server?","created":"2011-06-27T20:03:06.129+0000"},{"body":"Thanks for the test case, Jim.\n\nIf by the same scenario, you mean the workload that left your commitlog throwing exceptions, then yes please.","created":"2011-06-27T21:27:58.807+0000"},{"body":"We have deployed and tested 0.7.6 plus this patch to the affected cluster. The cluster restarted successfully and the tests that caused the original failure ran successfully. In addition, functional tests of our applications show no regressions. I also reviewed the Cassandra system logs after the testing and saw no errors or obvious problems.\n","created":"2011-06-29T15:35:01.348+0000"},{"body":"committed. Thanks, Jim!","created":"2011-06-30T01:02:54.430+0000"}],"conversations":[{"body":"I use hector 0.8.0-1 and cassandra 0.8.\n\n1. create mutator by using hector api, \n2. Insert a few columns into the mutator for key \"key1\", cf \"standard\". \n3. add a deletion to the mutator to delete the record of \"key1\", cf \"standard\".\n4. repeat 2 and 3\n5. execute the mutator.\n\nthe result: the connection seems to be held by the sever forever, it never returns. when I tried to restart the cassandra I saw unsupportedexception : \"Index manager cannot support deleting and inserting into a row in the same mutation\". and the cassandra is dead forever, unless I delete the commitlog. \n\nI would expect to get an exception when I execute the mutator, not after I restart the cassandra.","from":"reporter","subject":"\"Index manager cannot support deleting and inserting into a row in the same mutation\""},{"body":"Like it says, you will have to send your deletes and mutations in different mutations.","from":"developer"},{"body":"sure, I know what you mean.\n\nhowever, when the user uses it the wrong way, the server should at least response an error or exception... not just hold the connection and do nothing...\n\nIn addition, the server should not be dead forever, when this kind of exception occurs. ","from":"developer"},{"body":"You're right, there's a validation problem here.","from":"developer"},{"body":"we could add some valdation logic but it looks like it's almost as easy to just remove this limitation. patch to do this attached.","from":"developer"},{"body":"Hum, we cannot remove the column from cf in ignoreObsoleteMutations() because cf is the original column family from the row mutation and that's racy with commit log write (à la CASSANDRA-2604). We should clone the column family, but maybe it's simpler to add validation logic after all ? In any case, it could be worth it adding some comment in Table.apply() or Table.ignoreObsoleteMutations(). ","from":"developer"},{"body":"bq. we cannot remove the column from cf in ignoreObsoleteMutations()\n\nYou're right. Fortunately I don't think that's actually necessary. v2 attached.","from":"developer"},{"body":"Right, +1. I still think we should add a comment somewhere saying we shouldn't change cf (and that there is no reason to change it anyway).","from":"developer"},{"body":"Committed, with comment. (I wish CF objects could be immutable AND efficient...)","from":"developer"},{"body":"We are experiencing this issue on a cluster running 0.7.6-2. Once this occurs on a server, the same message is repeated over and over again in system.log. When we attempted to restart the box it failed. I will attach the log to this issue.","from":"developer"},{"body":"Log of failed restart attempt","from":"developer"},{"body":"Marked as critical, because when this occurs the server will not restart without removing commitlog and attending loss of data.","from":"developer"},{"body":"I explained on-list that I'm not comfortable committing this to 0.7. The critical thing we need to deliver in 0.7 at this stage is stability. \n\nThe rarity of the problem combined with the complexity of the code involved leads me to conclude that it's better to live with a bug we know how to avoid, than risk introducing new ones. Remember, the risk of new bugs affects _everyone_, while the fix can only benefit those who were generating the unusual mutation pattern here. It's our responsibility to take a balanced view.\n\nFor the rare people who do find themselves affected here, your options include\n- stay on < 0.7.6\n- drain the commitlog with an earlier version before upgrading, then re-upgrade to 0.7.6 after fixing your code to not generate problematic mutations\n- run an unofficial build with this patch included\n- upgrade to 0.8.2 when released","from":"developer"},{"body":"I disagree. This is a bug which allows a client to make a cluster unresponsive by performing a seemingly innocuous series of operations. If that happens, the cluster is un-restartable without loss of data. I wouldn't call a release where this can occur \"stable\". So if the goal for 0.7 is stability...\n\nWRT \"fixing your code to not generate problematic mutations,\" this may be difficult to do. I have so far identified code that does deletes followed by updates in the same mutation, but I haven't yet found any updates followed by deletes. Are we sure that only the update-followed-by-delete scenario is problematic?\n\nIn any case, even after reviewing all our code for the relevant scenarios, I would not feel comfortable deploying an 0.7 release with this vulnerability to production. The risk of a catastrophic failure is too great.\n","from":"developer"},{"body":"Actually, there isn't really much risk for data loss given that as Jonathan said, if you hit that, it's fairly easy to go back to 0.7.5, fix client code and upgrade again. Granted this is not user friendly and not something you should expect from a minor upgrade, but let at least set the record straight on the data loss part.\n\nThat being said, I don't think the patch on this ticket could screw up indexes more that we use to prior to 0.7.6, so maybe we can commit to 0.7.7 on that ground.\n\nI'd still suggest fixing client code in the meantime. ","from":"developer"},{"body":"bq. I don't think the patch on this ticket could screw up indexes more that we use to prior to 0.7.6\n\nThat's a valid way to frame the issue.\n\nI'm good to commit for 0.7.7 if Jim can test the patch first, since he's the only one we've heard of hitting this in 0.7.x. (Specifically, we want to make sure that if we query \"WHERE foo = X\" we don't get results back where foo is something other than X. Ideally you'd start with an empty database, or at least drop + recreate indexes first to make sure the results aren't contaminated w/ corrupt entries from pre-0.7.6.)","from":"developer"},{"body":"I applied the 0.8 patch and added a couple of tests to ColumnFamilyStoreTest. The tests trigger the UnsupportedOperationException in 0.7.6 and return the correct values with the patch applied. Would you like me to test the same scenario against an actual patched server?","from":"developer"},{"body":"Thanks for the test case, Jim.\n\nIf by the same scenario, you mean the workload that left your commitlog throwing exceptions, then yes please.","from":"developer"},{"body":"We have deployed and tested 0.7.6 plus this patch to the affected cluster. The cluster restarted successfully and the tests that caused the original failure ran successfully. In addition, functional tests of our applications show no regressions. I also reviewed the Cassandra system logs after the testing and saw no errors or obvious problems.\n","from":"developer"},{"body":"committed. Thanks, Jim!","from":"developer"}],"created":"2011-06-15T03:10:33.000+0000","description":"I use hector 0.8.0-1 and cassandra 0.8.\n\n1. create mutator by using hector api, \n2. Insert a few columns into the mutator for key \"key1\", cf \"standard\". \n3. add a deletion to the mutator to delete the record of \"key1\", cf \"standard\".\n4. repeat 2 and 3\n5. execute the mutator.\n\nthe result: the connection seems to be held by the sever forever, it never returns. when I tried to restart the cassandra I saw unsupportedexception : \"Index manager cannot support deleting and inserting into a row in the same mutation\". and the cassandra is dead forever, unless I delete the commitlog. \n\nI would expect to get an exception when I execute the mutator, not after I restart the cassandra.","issue_id":"12510393","key":"CASSANDRA-2773","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-06-30T01:02:54.000+0000","role":"fixed_distractor","summary":"\"Index manager cannot support deleting and inserting into a row in the same mutation\""} {"case_id":"12510655","cluster":"DISTRACTOR-CASSANDRA-2786","comments":[{"body":"The java version would be really cool :)","created":"2011-06-17T14:51:22.513+0000"},{"body":"I included the Java version. You have to play a little bit with the numbers of rows to insert in order to get the correct compaction timings.","created":"2011-06-20T13:19:04.541+0000"},{"body":"We were wrongfully skipping deleted rows with no columns during compaction. This indeed don't affect 0.7 since this was due to a refactor of PrecompactedRow in 0.8. Patch attached with a unit test to catch the error.","created":"2011-06-21T10:56:43.019+0000"},{"body":"+1\n\n(can we make the \"testing\" constructor package-local?)","created":"2011-06-21T12:26:59.381+0000"},{"body":"Committed, thanks.\n\nI did not made the \"testing\" constructor package-local because it is used in the AntiEntropyTests which are not on the same package. But I agree it's not the cleanest thing ever.","created":"2011-06-21T13:48:40.891+0000"},{"body":"Tested with 0.8.1 but still doesn't work","created":"2011-06-29T16:44:37.370+0000"},{"body":"Yeah, turns out EchoedRow is also handling Row tombstones with no columns inside badly.\n\nAttaching patch with fix and unit test. 0.7 is not really impacted because it uses EchoedRow only for cleanup and don't use its isEmpty() function there (but I suppose we could make it throw an UnsupporteOperationException to be on the safe side).\n\nThe patch actually ship with two changes that are not strictly related to the issue:\n# It fixes testEchoedRow in CompactionsTest. It wasn't using EchoedRow anymore (i.e, the test was useless).\n# It always forces deserialization for user submitted compaction (by opposition to only when the user submits only 1 sstable). It is done because exposing the forceDeserialization flag was necessary to write the test for this issue. Following that change, it was trivial to do the user submitted compaction change. It also fix a bad comment (forcing deserialization is only useful for forcing expired column to become tombstones, not for purging since purging will happen without force deserialization if it can).","created":"2011-06-30T15:04:47.389+0000"},{"body":"Nit: wouldn't it be cleaner to just pass gcBefore rather than the entire controller to EchoedRow constructor?\n\n+1 otherwise.","created":"2011-07-01T20:02:12.452+0000"},{"body":"Committed, thanks.\n\nbq. Nit: wouldn't it be cleaner to just pass gcBefore rather than the entire controller to EchoedRow constructor?\n\nI passed the controller because Precompacted and LazilyCompacted do that too, so it felt slightly cleaner, and if we happen to need more info from the controller in the future, it'll be there. But really at the end I did not change it before committing out of laziness :)","created":"2011-07-06T12:19:26.043+0000"},{"body":"bq. Precompacted and LazilyCompacted do that too\n\nThat makes sense.","created":"2011-07-06T13:10:14.719+0000"},{"body":"Tested with 0.8.2 and 0.8.7, but still does not work. On 0.7.x it works fine.","created":"2011-11-21T16:44:08.345+0000"},{"body":"One note: I tested with several grace-periods. With a grace-period of one minute, it is easier to reproduce. On our production site (with grace-priod of 24 hours), the data resurrects after several days.","created":"2011-11-21T16:48:13.102+0000"},{"body":"bq. With a grace-period of one minute, it is easier to reproduce.\n\nThis seems to me to suggest you aren't running repair often enough and are encountering the same effect as CASSANDRA-1316.","created":"2011-11-21T21:14:42.784+0000"},{"body":"bq. This seems to me to suggest you aren't running repair often enough \n\nIf you can only reproduce on multiple nodes, that is probably the issue here.","created":"2011-11-21T21:17:18.020+0000"},{"body":"With the attached program I'm able to reproduce it on a single node.","created":"2011-11-21T23:50:27.655+0000"},{"body":"Per previous phone conversation, I tested this against 0.8.7 twice.\n\nBoth times the attached java test ran to completion against a single node and output \"Done\".\n\nWhen was this last reproduced?","created":"2011-11-22T00:59:32.458+0000"},{"body":"I should mention that I did change one line in the test. I changed {noformat}column.timestamp = getTimestamp();{noformat} to {noformat}column.setTimestamp(getTimestamp());{noformat} because otherwise thrift complained that the timestamp wasn't set.","created":"2011-11-22T01:03:37.731+0000"},{"body":"Could you please check again with grace_period of 60.\n\nI use the following to reproduce:\n\ncreate column family Customers\n with column_type = 'Super' \n and comparator = 'BytesType'\n\tand memtable_flush_after = 60\n\tand gc_grace = 60;\n\nOn my system, it crashes every time:\n\nC:\\Temp\\JavaIssue>java -jar CassandraIssue.jar 127.0.0.1 Traxis\nException in thread \"main\" java.lang.Exception: test row should be empty\n at cassandraissue.Main.start(Main.java:88)\n at cassandraissue.Main.main(Main.java:178)\n\t\t\nThanks","created":"2011-11-22T08:49:06.047+0000"},{"body":"I had to tweak the variables a little bit but I'm able to reproduce. Will look into it.","created":"2011-11-22T13:12:44.957+0000"},{"body":"Hopefully we get this right that time. The problem was that we were calling removeDeleted even in case where we shouldn't have been purging. Attaching 'part3' patch to fix, along with an updated unit test for that.","created":"2011-11-22T17:18:50.544+0000"},{"body":"This is a little subtle so I'm going to spell it out:\n\nThe purpose of AbstractCompactedRow.isEmpty is to skip rows that consist only of expired tombstones:\n\n{code}\n. writer = cfs.createCompactionWriter(expectedBloomFilterSize, compactionFileLocation, sstables);\n while (nni.hasNext())\n {\n AbstractCompactedRow row = nni.next();\n if (row.isEmpty())\n continue;\n ...\n }\n{code}\n\nHowever, we can't skip tombstones if we're only compacting some of the sstables for a row (CASSANDRA-1074). The bug here is that the isEmpty test doesn't check the CompactionController.shouldPurge, which is how the controller lets us know it's okay to drop tombstones. (In the PR case the shouldPurge check was done correctly during creation of compactedCf, but then we ignored it when checking a second time for isEmpty.)\n\nSylvain's patch fixes the bug. Here is a v2 that simplifies isEmpty further:\n\n- ER.isEmpty is actually trivial\n- PR doesn't need to do a second check of no columns + no row level tombstone (i.e.: there were expired column tombstones); this case would be taken care of by the removeDeleted in compactedCf creation\n","created":"2011-11-23T14:02:11.172+0000"},{"body":"+1 on v2","created":"2011-11-23T14:51:24.810+0000"},{"body":"committed","created":"2011-11-23T15:39:16.730+0000"}],"conversations":[{"body":"After a minor compaction, deleted key-slices are visible again.\n\nSteps to reproduce:\n\n1) Insert a row named \"test\".\n2) Insert 500000 rows. During this step, row \"test\" is included in a major compaction:\n file-1, file-2, file-3 and file-4 compacted to file-5 (includes \"test\").\n3) Delete row named \"test\".\n4) Insert 500000 rows. During this step, row \"test\" is included in a minor compaction:\n file-6, file-7, file-8 and file-9 compacted to file-10 (should include tombstoned \"test\").\nAfter step 4, row \"test\" is live again.\n\nTest environment:\n\nSingle node with empty database.\n\nStandard configured super-column-family (I see this behavior with several gc_grace settings (big and small values):\ncreate column family Customers with column_type = 'Super' and comparator = 'BytesType;\n\nIn Cassandra 0.7.6 I observe the expected behavior, i.e. after step 4, the row is still deleted.\n\nI've included a .NET program to reproduce the problem. I will add a Java version later on.","from":"reporter","subject":"After a minor compaction, deleted key-slices are visible again"},{"body":"The java version would be really cool :)","from":"developer"},{"body":"I included the Java version. You have to play a little bit with the numbers of rows to insert in order to get the correct compaction timings.","from":"developer"},{"body":"We were wrongfully skipping deleted rows with no columns during compaction. This indeed don't affect 0.7 since this was due to a refactor of PrecompactedRow in 0.8. Patch attached with a unit test to catch the error.","from":"developer"},{"body":"+1\n\n(can we make the \"testing\" constructor package-local?)","from":"developer"},{"body":"Committed, thanks.\n\nI did not made the \"testing\" constructor package-local because it is used in the AntiEntropyTests which are not on the same package. But I agree it's not the cleanest thing ever.","from":"developer"},{"body":"Tested with 0.8.1 but still doesn't work","from":"developer"},{"body":"Yeah, turns out EchoedRow is also handling Row tombstones with no columns inside badly.\n\nAttaching patch with fix and unit test. 0.7 is not really impacted because it uses EchoedRow only for cleanup and don't use its isEmpty() function there (but I suppose we could make it throw an UnsupporteOperationException to be on the safe side).\n\nThe patch actually ship with two changes that are not strictly related to the issue:\n# It fixes testEchoedRow in CompactionsTest. It wasn't using EchoedRow anymore (i.e, the test was useless).\n# It always forces deserialization for user submitted compaction (by opposition to only when the user submits only 1 sstable). It is done because exposing the forceDeserialization flag was necessary to write the test for this issue. Following that change, it was trivial to do the user submitted compaction change. It also fix a bad comment (forcing deserialization is only useful for forcing expired column to become tombstones, not for purging since purging will happen without force deserialization if it can).","from":"developer"},{"body":"Nit: wouldn't it be cleaner to just pass gcBefore rather than the entire controller to EchoedRow constructor?\n\n+1 otherwise.","from":"developer"},{"body":"Committed, thanks.\n\nbq. Nit: wouldn't it be cleaner to just pass gcBefore rather than the entire controller to EchoedRow constructor?\n\nI passed the controller because Precompacted and LazilyCompacted do that too, so it felt slightly cleaner, and if we happen to need more info from the controller in the future, it'll be there. But really at the end I did not change it before committing out of laziness :)","from":"developer"},{"body":"bq. Precompacted and LazilyCompacted do that too\n\nThat makes sense.","from":"developer"},{"body":"Tested with 0.8.2 and 0.8.7, but still does not work. On 0.7.x it works fine.","from":"developer"},{"body":"One note: I tested with several grace-periods. With a grace-period of one minute, it is easier to reproduce. On our production site (with grace-priod of 24 hours), the data resurrects after several days.","from":"developer"},{"body":"bq. With a grace-period of one minute, it is easier to reproduce.\n\nThis seems to me to suggest you aren't running repair often enough and are encountering the same effect as CASSANDRA-1316.","from":"developer"},{"body":"bq. This seems to me to suggest you aren't running repair often enough \n\nIf you can only reproduce on multiple nodes, that is probably the issue here.","from":"developer"},{"body":"With the attached program I'm able to reproduce it on a single node.","from":"developer"},{"body":"Per previous phone conversation, I tested this against 0.8.7 twice.\n\nBoth times the attached java test ran to completion against a single node and output \"Done\".\n\nWhen was this last reproduced?","from":"developer"},{"body":"I should mention that I did change one line in the test. I changed {noformat}column.timestamp = getTimestamp();{noformat} to {noformat}column.setTimestamp(getTimestamp());{noformat} because otherwise thrift complained that the timestamp wasn't set.","from":"developer"},{"body":"Could you please check again with grace_period of 60.\n\nI use the following to reproduce:\n\ncreate column family Customers\n with column_type = 'Super' \n and comparator = 'BytesType'\n\tand memtable_flush_after = 60\n\tand gc_grace = 60;\n\nOn my system, it crashes every time:\n\nC:\\Temp\\JavaIssue>java -jar CassandraIssue.jar 127.0.0.1 Traxis\nException in thread \"main\" java.lang.Exception: test row should be empty\n at cassandraissue.Main.start(Main.java:88)\n at cassandraissue.Main.main(Main.java:178)\n\t\t\nThanks","from":"developer"},{"body":"I had to tweak the variables a little bit but I'm able to reproduce. Will look into it.","from":"developer"},{"body":"Hopefully we get this right that time. The problem was that we were calling removeDeleted even in case where we shouldn't have been purging. Attaching 'part3' patch to fix, along with an updated unit test for that.","from":"developer"},{"body":"This is a little subtle so I'm going to spell it out:\n\nThe purpose of AbstractCompactedRow.isEmpty is to skip rows that consist only of expired tombstones:\n\n{code}\n. writer = cfs.createCompactionWriter(expectedBloomFilterSize, compactionFileLocation, sstables);\n while (nni.hasNext())\n {\n AbstractCompactedRow row = nni.next();\n if (row.isEmpty())\n continue;\n ...\n }\n{code}\n\nHowever, we can't skip tombstones if we're only compacting some of the sstables for a row (CASSANDRA-1074). The bug here is that the isEmpty test doesn't check the CompactionController.shouldPurge, which is how the controller lets us know it's okay to drop tombstones. (In the PR case the shouldPurge check was done correctly during creation of compactedCf, but then we ignored it when checking a second time for isEmpty.)\n\nSylvain's patch fixes the bug. Here is a v2 that simplifies isEmpty further:\n\n- ER.isEmpty is actually trivial\n- PR doesn't need to do a second check of no columns + no row level tombstone (i.e.: there were expired column tombstones); this case would be taken care of by the removeDeleted in compactedCf creation\n","from":"developer"},{"body":"+1 on v2","from":"developer"},{"body":"committed","from":"developer"}],"created":"2011-06-17T12:42:37.000+0000","description":"After a minor compaction, deleted key-slices are visible again.\n\nSteps to reproduce:\n\n1) Insert a row named \"test\".\n2) Insert 500000 rows. During this step, row \"test\" is included in a major compaction:\n file-1, file-2, file-3 and file-4 compacted to file-5 (includes \"test\").\n3) Delete row named \"test\".\n4) Insert 500000 rows. During this step, row \"test\" is included in a minor compaction:\n file-6, file-7, file-8 and file-9 compacted to file-10 (should include tombstoned \"test\").\nAfter step 4, row \"test\" is live again.\n\nTest environment:\n\nSingle node with empty database.\n\nStandard configured super-column-family (I see this behavior with several gc_grace settings (big and small values):\ncreate column family Customers with column_type = 'Super' and comparator = 'BytesType;\n\nIn Cassandra 0.7.6 I observe the expected behavior, i.e. after step 4, the row is still deleted.\n\nI've included a .NET program to reproduce the problem. I will add a Java version later on.","issue_id":"12510655","key":"CASSANDRA-2786","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-11-23T15:39:16.000+0000","role":"fixed_distractor","summary":"After a minor compaction, deleted key-slices are visible again"} {"case_id":"12511243","cluster":"DISTRACTOR-CASSANDRA-2810","comments":[{"body":"Patch to use a custom AbstractType in place of BytesType to nip this in the bud, rather than have a bunch of one-off checks. Also fixes a bug where the supercolumn name is never set.","created":"2011-06-24T14:45:15.352+0000"},{"body":"DataByteArray is some kind of Pig thing?","created":"2011-06-24T14:52:20.683+0000"},{"body":"Yes, basically a byte array, but it's the pig type.","created":"2011-06-24T15:07:17.960+0000"},{"body":"I try again after applying [^2810.txt] and the patch from bug [CASSANDRA-2777] and the bug is still here.\nWith the patch, you need to replace\n{code}\ntest = LOAD 'cassandra://Test/test' USING CassandraStorage() AS (rowkey:chararray, columns: bag {T: (name:long, value:int)});\n{code}\nby\n{code}\ntest = LOAD 'cassandra://Test/test' USING CassandraStorage() AS ();\n{code}\nbecause CassandraStorage takes care of the schema.\n\nI try:\n{code}\ngrunt> describe test;\ntest: {key: chararray,columns: {(name: long,value: int)}}\n{code}\nso we can see that the patch from bug 2777 works correctly (I also test with different types for value).\nBut when I dump test, I still have the same exception.","created":"2011-06-27T13:07:56.407+0000"},{"body":"After more test (with both patches), path [^2810.txt] doesn't seems to solve the bug.\nHere is a new test case:\nCreate a _Test_ keyspace and a _test_ column family with key_validation_class = 'AsciiType' and comparator = 'LongType' and default_validation_class = 'IntegerType' (don't use the cli because of [#CASSANDRA-2831]).\nInsert some data:\n{code}\nset test[ascii('row1')][long(1)]=integer(35);\nset test[ascii('row1')][long(2)]=integer(36);\nset test[ascii('row1')][long(3)]=integer(38);\nset test[ascii('row2')][long(1)]=integer(45);\nset test[ascii('row2')][long(2)]=integer(42);\nset test[ascii('row2')][long(3)]=integer(33);\n{code}\n\nIn Pig cli:\n{code}\ntest = LOAD 'cassandra://Test/test' USING CassandraStorage() AS ();\ndump test;\n{code}\nThe same exception as before is raised:\n{code}\n INFO [IPC Server handler 4 on 8012] 2011-06-27 16:40:28,562 TaskInProgress.java (line 551) Error from attempt_201106271436_0012_m_000000_1: java.lang.RuntimeException: Unexpected data type -1 found in stream.\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:478)\n\tat org.apache.pig.data.BinInterSedes.writeTuple(BinInterSedes.java:541)\n\tat org.apache.pig.data.BinInterSedes.writeBag(BinInterSedes.java:522)\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:361)\n\tat org.apache.pig.data.BinInterSedes.writeTuple(BinInterSedes.java:541)\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:357)\n\tat org.apache.pig.impl.io.InterRecordWriter.write(InterRecordWriter.java:73)\n\tat org.apache.pig.impl.io.InterStorage.putNext(InterStorage.java:87)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:138)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:97)\n\tat org.apache.hadoop.mapred.MapTask$NewDirectOutputCollector.write(MapTask.java:638)\n\tat org.apache.hadoop.mapreduce.TaskInputOutputContext.write(TaskInputOutputContext.java:80)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapOnly$Map.collect(PigMapOnly.java:48)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapBase.map(PigMapBase.java:224)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapBase.map(PigMapBase.java:53)\n\tat org.apache.hadoop.mapreduce.Mapper.run(Mapper.java:144)\n\tat org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:763)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:369)\n\tat org.apache.hadoop.mapred.Child$4.run(Child.java:259)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:396)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1059)\n\tat org.apache.hadoop.mapred.Child.main(Child.java:253)\n\n{code}","created":"2011-06-27T14:47:23.697+0000"},{"body":"So is the conclusion that this patch by itself works fine, but there is a problem with CASSANDRA-2777?","created":"2011-07-08T19:41:16.068+0000"},{"body":"No, from my test I arrived to the inverse conclusion: [#CASSANDRA-2777] seems to works fine (Pig has the good type for my column family) but the bug is still here despite the 2 patches.","created":"2011-07-12T13:54:27.771+0000"},{"body":"It looks like the final problem here is that IntegerType always returns a BigInteger, which pig does not like. This is unfortunate since IntegerType can't be easily subclassed and overridden to return ints.\n\nv2 instead adds a setTupleValue method that is always used for adding values to tuples, and houses all the special-casing currently needed and provides a spot for more in the future, rather than proliferating custom type converters since I'm sure IntegerType won't be alone here.","created":"2011-08-10T00:55:40.951+0000"},{"body":"v3 also removes decomposing the values before inserting and instead forces them into a ByteBuffer with objToBB, since we actually don't care about the type. (why did we ever change this?)\n\nThis means that a UDF that doesn't preserve the schema and hands us back DataByteArrays when we fed it specific types can't make us fail anymore.","created":"2011-08-23T16:25:53.289+0000"},{"body":"Fixed it for me on Pig 0.9 and Cassandra 0.8.6 (Brisk).","created":"2011-09-28T21:30:30.414+0000"},{"body":"+1 - if we find any issues with it in production, we'll submit bug reports.","created":"2011-09-28T21:35:39.694+0000"},{"body":"Committed.","created":"2011-09-28T22:01:22.070+0000"}],"conversations":[{"body":"This bug was previously report on [Brisk bug tracker|https://datastax.jira.com/browse/BRISK-232].\n\nIn cassandra-cli:\n{code}\n[default@unknown] create keyspace Test\n with placement_strategy = 'org.apache.cassandra.locator.SimpleStrategy'\n and strategy_options = [{replication_factor:1}];\n\n[default@unknown] use Test;\nAuthenticated to keyspace: Test\n\n[default@Test] create column family test;\n\n[default@Test] set test[ascii('row1')][long(1)]=integer(35);\nset test[ascii('row1')][long(2)]=integer(36);\nset test[ascii('row1')][long(3)]=integer(38);\nset test[ascii('row2')][long(1)]=integer(45);\nset test[ascii('row2')][long(2)]=integer(42);\nset test[ascii('row2')][long(3)]=integer(33);\n\n[default@Test] list test;\nUsing default limit of 100\n-------------------\nRowKey: 726f7731\n=> (column=0000000000000001, value=35, timestamp=1308744931122000)\n=> (column=0000000000000002, value=36, timestamp=1308744931124000)\n=> (column=0000000000000003, value=38, timestamp=1308744931125000)\n-------------------\nRowKey: 726f7732\n=> (column=0000000000000001, value=45, timestamp=1308744931127000)\n=> (column=0000000000000002, value=42, timestamp=1308744931128000)\n=> (column=0000000000000003, value=33, timestamp=1308744932722000)\n\n2 Rows Returned.\n\n[default@Test] describe keyspace;\nKeyspace: Test:\n Replication Strategy: org.apache.cassandra.locator.SimpleStrategy\n Durable Writes: true\n Options: [replication_factor:1]\n Column Families:\n ColumnFamily: test\n Key Validation Class: org.apache.cassandra.db.marshal.BytesType\n Default column value validator: org.apache.cassandra.db.marshal.BytesType\n Columns sorted by: org.apache.cassandra.db.marshal.BytesType\n Row cache size / save period in seconds: 0.0/0\n Key cache size / save period in seconds: 200000.0/14400\n Memtable thresholds: 0.571875/122/1440 (millions of ops/MB/minutes)\n GC grace seconds: 864000\n Compaction min/max thresholds: 4/32\n Read repair chance: 1.0\n Replicate on write: false\n Built indexes: []\n{code}\nIn Pig command line:\n{code}\ngrunt> test = LOAD 'cassandra://Test/test' USING CassandraStorage() AS (rowkey:chararray, columns: bag {T: (name:long, value:int)});\n\ngrunt> value_test = foreach test generate rowkey, columns.name, columns.value;\n\ngrunt> dump value_test;\n{code}\nIn /var/log/cassandra/system.log, I have severals time this exception:\n{code}\nINFO [IPC Server handler 3 on 8012] 2011-06-22 15:03:28,533 TaskInProgress.java (line 551) Error from attempt_201106210955_0051_m_000000_3: java.lang.RuntimeException: Unexpected data type -1 found in stream.\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:478)\n\tat org.apache.pig.data.BinInterSedes.writeTuple(BinInterSedes.java:541)\n\tat org.apache.pig.data.BinInterSedes.writeBag(BinInterSedes.java:522)\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:361)\n\tat org.apache.pig.data.BinInterSedes.writeTuple(BinInterSedes.java:541)\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:357)\n\tat org.apache.pig.impl.io.InterRecordWriter.write(InterRecordWriter.java:73)\n\tat org.apache.pig.impl.io.InterStorage.putNext(InterStorage.java:87)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:138)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:97)\n\tat org.apache.hadoop.mapred.MapTask$NewDirectOutputCollector.write(MapTask.java:638)\n\tat org.apache.hadoop.mapreduce.TaskInputOutputContext.write(TaskInputOutputContext.java:80)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapOnly$Map.collect(PigMapOnly.java:48)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapBase.runPipeline(PigMapBase.java:239)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapBase.map(PigMapBase.java:232)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapBase.map(PigMapBase.java:53)\n\tat org.apache.hadoop.mapreduce.Mapper.run(Mapper.java:144)\n\tat org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:763)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:369)\n\tat org.apache.hadoop.mapred.Child$4.run(Child.java:259)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:396)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1059)\n\tat org.apache.hadoop.mapred.Child.main(Child.java:253)\n{code}\nand the request failed.\n\n{code}\ngrunt> test = LOAD 'cassandra://Test/test' USING CassandraStorage() AS (rowkey:chararray, columns: bag {T: (name:long, value:int)});\n\ngrunt> value_test = foreach test generate rowkey, columns.value;\n\ngrunt> dump value_test;\n{code}\n\nThis time, without the column name, it's work (but the value are displayed as char instead of integer). Result:\n{code}\n(row1,{(#),($),(&)})\n(row2,{(-),(*),(!)})\n{code}\n\nNow we do the same test but we set a comparator to the CF.\n{code}\n[default@Test] create column family test with comparator = 'LongType';\n\n[default@Test] set test[ascii('row1')][long(1)]=integer(35);\nset test[ascii('row1')][long(2)]=integer(36);\nset test[ascii('row1')][long(3)]=integer(38);\nset test[ascii('row2')][long(1)]=integer(45);\nset test[ascii('row2')][long(2)]=integer(42);\nset test[ascii('row2')][long(3)]=integer(33);\n\n[default@Test] list test;\nUsing default limit of 100\n-------------------\nRowKey: 726f7731\n=> (column=1, value=35, timestamp=1308748643506000)\n=> (column=2, value=36, timestamp=1308748643508000)\n=> (column=3, value=38, timestamp=1308748643509000)\n-------------------\nRowKey: 726f7732\n=> (column=1, value=45, timestamp=1308748643510000)\n=> (column=2, value=42, timestamp=1308748643512000)\n=> (column=3, value=33, timestamp=1308748645138000)\n\n2 Rows Returned.\n\n[default@Test] describe keyspace;\nKeyspace: Test:\n Replication Strategy: org.apache.cassandra.locator.SimpleStrategy\n Durable Writes: true\n Options: [replication_factor:1]\n Column Families:\n ColumnFamily: test\n Key Validation Class: org.apache.cassandra.db.marshal.BytesType\n Default column value validator: org.apache.cassandra.db.marshal.BytesType\n Columns sorted by: org.apache.cassandra.db.marshal.LongType\n Row cache size / save period in seconds: 0.0/0\n Key cache size / save period in seconds: 200000.0/14400\n Memtable thresholds: 0.571875/122/1440 (millions of ops/MB/minutes)\n GC grace seconds: 864000\n Compaction min/max thresholds: 4/32\n Read repair chance: 1.0\n Replicate on write: false\n Built indexes: []\n{code}\n{code}\ngrunt> test = LOAD 'cassandra://Test/test' USING CassandraStorage() AS (rowkey:chararray, columns: bag {T: (name:long, value:int)});\n\ngrunt> value_test = foreach test generate rowkey, columns.name, columns.value;\n\ngrunt> dump value_test;\n{code}\nThis time it's work as expected (appart from the value displayed as char). Result:\n{code}\n(row1,{(1),(2),(3)},{(#),($),(&)})\n(row2,{(1),(2),(3)},{(-),(*),(!)})\n{code}\n","from":"reporter","subject":"RuntimeException in Pig when using \"dump\" command on column name"},{"body":"Patch to use a custom AbstractType in place of BytesType to nip this in the bud, rather than have a bunch of one-off checks. Also fixes a bug where the supercolumn name is never set.","from":"developer"},{"body":"DataByteArray is some kind of Pig thing?","from":"developer"},{"body":"Yes, basically a byte array, but it's the pig type.","from":"developer"},{"body":"I try again after applying [^2810.txt] and the patch from bug [CASSANDRA-2777] and the bug is still here.\nWith the patch, you need to replace\n{code}\ntest = LOAD 'cassandra://Test/test' USING CassandraStorage() AS (rowkey:chararray, columns: bag {T: (name:long, value:int)});\n{code}\nby\n{code}\ntest = LOAD 'cassandra://Test/test' USING CassandraStorage() AS ();\n{code}\nbecause CassandraStorage takes care of the schema.\n\nI try:\n{code}\ngrunt> describe test;\ntest: {key: chararray,columns: {(name: long,value: int)}}\n{code}\nso we can see that the patch from bug 2777 works correctly (I also test with different types for value).\nBut when I dump test, I still have the same exception.","from":"developer"},{"body":"After more test (with both patches), path [^2810.txt] doesn't seems to solve the bug.\nHere is a new test case:\nCreate a _Test_ keyspace and a _test_ column family with key_validation_class = 'AsciiType' and comparator = 'LongType' and default_validation_class = 'IntegerType' (don't use the cli because of [#CASSANDRA-2831]).\nInsert some data:\n{code}\nset test[ascii('row1')][long(1)]=integer(35);\nset test[ascii('row1')][long(2)]=integer(36);\nset test[ascii('row1')][long(3)]=integer(38);\nset test[ascii('row2')][long(1)]=integer(45);\nset test[ascii('row2')][long(2)]=integer(42);\nset test[ascii('row2')][long(3)]=integer(33);\n{code}\n\nIn Pig cli:\n{code}\ntest = LOAD 'cassandra://Test/test' USING CassandraStorage() AS ();\ndump test;\n{code}\nThe same exception as before is raised:\n{code}\n INFO [IPC Server handler 4 on 8012] 2011-06-27 16:40:28,562 TaskInProgress.java (line 551) Error from attempt_201106271436_0012_m_000000_1: java.lang.RuntimeException: Unexpected data type -1 found in stream.\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:478)\n\tat org.apache.pig.data.BinInterSedes.writeTuple(BinInterSedes.java:541)\n\tat org.apache.pig.data.BinInterSedes.writeBag(BinInterSedes.java:522)\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:361)\n\tat org.apache.pig.data.BinInterSedes.writeTuple(BinInterSedes.java:541)\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:357)\n\tat org.apache.pig.impl.io.InterRecordWriter.write(InterRecordWriter.java:73)\n\tat org.apache.pig.impl.io.InterStorage.putNext(InterStorage.java:87)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:138)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:97)\n\tat org.apache.hadoop.mapred.MapTask$NewDirectOutputCollector.write(MapTask.java:638)\n\tat org.apache.hadoop.mapreduce.TaskInputOutputContext.write(TaskInputOutputContext.java:80)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapOnly$Map.collect(PigMapOnly.java:48)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapBase.map(PigMapBase.java:224)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapBase.map(PigMapBase.java:53)\n\tat org.apache.hadoop.mapreduce.Mapper.run(Mapper.java:144)\n\tat org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:763)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:369)\n\tat org.apache.hadoop.mapred.Child$4.run(Child.java:259)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:396)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1059)\n\tat org.apache.hadoop.mapred.Child.main(Child.java:253)\n\n{code}","from":"developer"},{"body":"So is the conclusion that this patch by itself works fine, but there is a problem with CASSANDRA-2777?","from":"developer"},{"body":"No, from my test I arrived to the inverse conclusion: [#CASSANDRA-2777] seems to works fine (Pig has the good type for my column family) but the bug is still here despite the 2 patches.","from":"developer"},{"body":"It looks like the final problem here is that IntegerType always returns a BigInteger, which pig does not like. This is unfortunate since IntegerType can't be easily subclassed and overridden to return ints.\n\nv2 instead adds a setTupleValue method that is always used for adding values to tuples, and houses all the special-casing currently needed and provides a spot for more in the future, rather than proliferating custom type converters since I'm sure IntegerType won't be alone here.","from":"developer"},{"body":"v3 also removes decomposing the values before inserting and instead forces them into a ByteBuffer with objToBB, since we actually don't care about the type. (why did we ever change this?)\n\nThis means that a UDF that doesn't preserve the schema and hands us back DataByteArrays when we fed it specific types can't make us fail anymore.","from":"developer"},{"body":"Fixed it for me on Pig 0.9 and Cassandra 0.8.6 (Brisk).","from":"developer"},{"body":"+1 - if we find any issues with it in production, we'll submit bug reports.","from":"developer"},{"body":"Committed.","from":"developer"}],"created":"2011-06-22T14:49:36.000+0000","description":"This bug was previously report on [Brisk bug tracker|https://datastax.jira.com/browse/BRISK-232].\n\nIn cassandra-cli:\n{code}\n[default@unknown] create keyspace Test\n with placement_strategy = 'org.apache.cassandra.locator.SimpleStrategy'\n and strategy_options = [{replication_factor:1}];\n\n[default@unknown] use Test;\nAuthenticated to keyspace: Test\n\n[default@Test] create column family test;\n\n[default@Test] set test[ascii('row1')][long(1)]=integer(35);\nset test[ascii('row1')][long(2)]=integer(36);\nset test[ascii('row1')][long(3)]=integer(38);\nset test[ascii('row2')][long(1)]=integer(45);\nset test[ascii('row2')][long(2)]=integer(42);\nset test[ascii('row2')][long(3)]=integer(33);\n\n[default@Test] list test;\nUsing default limit of 100\n-------------------\nRowKey: 726f7731\n=> (column=0000000000000001, value=35, timestamp=1308744931122000)\n=> (column=0000000000000002, value=36, timestamp=1308744931124000)\n=> (column=0000000000000003, value=38, timestamp=1308744931125000)\n-------------------\nRowKey: 726f7732\n=> (column=0000000000000001, value=45, timestamp=1308744931127000)\n=> (column=0000000000000002, value=42, timestamp=1308744931128000)\n=> (column=0000000000000003, value=33, timestamp=1308744932722000)\n\n2 Rows Returned.\n\n[default@Test] describe keyspace;\nKeyspace: Test:\n Replication Strategy: org.apache.cassandra.locator.SimpleStrategy\n Durable Writes: true\n Options: [replication_factor:1]\n Column Families:\n ColumnFamily: test\n Key Validation Class: org.apache.cassandra.db.marshal.BytesType\n Default column value validator: org.apache.cassandra.db.marshal.BytesType\n Columns sorted by: org.apache.cassandra.db.marshal.BytesType\n Row cache size / save period in seconds: 0.0/0\n Key cache size / save period in seconds: 200000.0/14400\n Memtable thresholds: 0.571875/122/1440 (millions of ops/MB/minutes)\n GC grace seconds: 864000\n Compaction min/max thresholds: 4/32\n Read repair chance: 1.0\n Replicate on write: false\n Built indexes: []\n{code}\nIn Pig command line:\n{code}\ngrunt> test = LOAD 'cassandra://Test/test' USING CassandraStorage() AS (rowkey:chararray, columns: bag {T: (name:long, value:int)});\n\ngrunt> value_test = foreach test generate rowkey, columns.name, columns.value;\n\ngrunt> dump value_test;\n{code}\nIn /var/log/cassandra/system.log, I have severals time this exception:\n{code}\nINFO [IPC Server handler 3 on 8012] 2011-06-22 15:03:28,533 TaskInProgress.java (line 551) Error from attempt_201106210955_0051_m_000000_3: java.lang.RuntimeException: Unexpected data type -1 found in stream.\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:478)\n\tat org.apache.pig.data.BinInterSedes.writeTuple(BinInterSedes.java:541)\n\tat org.apache.pig.data.BinInterSedes.writeBag(BinInterSedes.java:522)\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:361)\n\tat org.apache.pig.data.BinInterSedes.writeTuple(BinInterSedes.java:541)\n\tat org.apache.pig.data.BinInterSedes.writeDatum(BinInterSedes.java:357)\n\tat org.apache.pig.impl.io.InterRecordWriter.write(InterRecordWriter.java:73)\n\tat org.apache.pig.impl.io.InterStorage.putNext(InterStorage.java:87)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:138)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:97)\n\tat org.apache.hadoop.mapred.MapTask$NewDirectOutputCollector.write(MapTask.java:638)\n\tat org.apache.hadoop.mapreduce.TaskInputOutputContext.write(TaskInputOutputContext.java:80)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapOnly$Map.collect(PigMapOnly.java:48)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapBase.runPipeline(PigMapBase.java:239)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapBase.map(PigMapBase.java:232)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigMapBase.map(PigMapBase.java:53)\n\tat org.apache.hadoop.mapreduce.Mapper.run(Mapper.java:144)\n\tat org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:763)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:369)\n\tat org.apache.hadoop.mapred.Child$4.run(Child.java:259)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:396)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1059)\n\tat org.apache.hadoop.mapred.Child.main(Child.java:253)\n{code}\nand the request failed.\n\n{code}\ngrunt> test = LOAD 'cassandra://Test/test' USING CassandraStorage() AS (rowkey:chararray, columns: bag {T: (name:long, value:int)});\n\ngrunt> value_test = foreach test generate rowkey, columns.value;\n\ngrunt> dump value_test;\n{code}\n\nThis time, without the column name, it's work (but the value are displayed as char instead of integer). Result:\n{code}\n(row1,{(#),($),(&)})\n(row2,{(-),(*),(!)})\n{code}\n\nNow we do the same test but we set a comparator to the CF.\n{code}\n[default@Test] create column family test with comparator = 'LongType';\n\n[default@Test] set test[ascii('row1')][long(1)]=integer(35);\nset test[ascii('row1')][long(2)]=integer(36);\nset test[ascii('row1')][long(3)]=integer(38);\nset test[ascii('row2')][long(1)]=integer(45);\nset test[ascii('row2')][long(2)]=integer(42);\nset test[ascii('row2')][long(3)]=integer(33);\n\n[default@Test] list test;\nUsing default limit of 100\n-------------------\nRowKey: 726f7731\n=> (column=1, value=35, timestamp=1308748643506000)\n=> (column=2, value=36, timestamp=1308748643508000)\n=> (column=3, value=38, timestamp=1308748643509000)\n-------------------\nRowKey: 726f7732\n=> (column=1, value=45, timestamp=1308748643510000)\n=> (column=2, value=42, timestamp=1308748643512000)\n=> (column=3, value=33, timestamp=1308748645138000)\n\n2 Rows Returned.\n\n[default@Test] describe keyspace;\nKeyspace: Test:\n Replication Strategy: org.apache.cassandra.locator.SimpleStrategy\n Durable Writes: true\n Options: [replication_factor:1]\n Column Families:\n ColumnFamily: test\n Key Validation Class: org.apache.cassandra.db.marshal.BytesType\n Default column value validator: org.apache.cassandra.db.marshal.BytesType\n Columns sorted by: org.apache.cassandra.db.marshal.LongType\n Row cache size / save period in seconds: 0.0/0\n Key cache size / save period in seconds: 200000.0/14400\n Memtable thresholds: 0.571875/122/1440 (millions of ops/MB/minutes)\n GC grace seconds: 864000\n Compaction min/max thresholds: 4/32\n Read repair chance: 1.0\n Replicate on write: false\n Built indexes: []\n{code}\n{code}\ngrunt> test = LOAD 'cassandra://Test/test' USING CassandraStorage() AS (rowkey:chararray, columns: bag {T: (name:long, value:int)});\n\ngrunt> value_test = foreach test generate rowkey, columns.name, columns.value;\n\ngrunt> dump value_test;\n{code}\nThis time it's work as expected (appart from the value displayed as char). Result:\n{code}\n(row1,{(1),(2),(3)},{(#),($),(&)})\n(row2,{(1),(2),(3)},{(-),(*),(!)})\n{code}\n","issue_id":"12511243","key":"CASSANDRA-2810","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-09-28T22:01:22.000+0000","role":"fixed_distractor","summary":"RuntimeException in Pig when using \"dump\" command on column name"} {"case_id":"12511343","cluster":"DISTRACTOR-CASSANDRA-2816","comments":[{"body":"I guess a dedicated validation executor is ok as long as it still obeys the global \"compaction\" i/o limit.","created":"2011-06-23T12:40:12.178+0000"},{"body":"I'm a fan of the snapshotting immediately after receiving the request approach. In general, polishing our snapshot support to allow for this kind of usecase is likely to open up other interesting possibilities.","created":"2011-06-24T04:27:16.500+0000"},{"body":"Supporting actual live-reading of snapshotted sstables is a little more than \"polishing.\" It would be cool, but I wouldn't want it to block fixing repair.","created":"2011-06-24T04:45:27.177+0000"},{"body":"I've thought about this problem too, and it is really significant for some use-cases. Again because so few writes are needed in order to trigger large amounts of data being sent given the merklee tree granularity.\n\nWhile I'm all for a fixing it by e.g. more immediate snapshotting, I would like to raise the issue that repairs overall have pretty significant side-effects; particularly ones that can self-magnify and cause further problems. Beyond the obvious \"it does disk I/O\" and \"It uses CPU\", we have:\n\n* Over-repair due to merklee tree granularity can cause jumps in CF sizes, killing cache locality\n* Combine that with concurrent repairs then repairing the \"size-jumped\" set of sstables and you can magnify that effect on other nodes causing huge size increases.\n* Up to recently, mixing large and small cf:s was a significant problem if you wanted to have different frequencies and different gc grace times, due to one repair blocking on another. But fixes to this and the other JIRA about concurrency, might disable the \"fix\" for that that was concurrent compaction - so back to square one.\n\nI guess overall, it seems very easy to shoot yourself in the foot with repair.\n\nAny opinions on CASSANDRA-2699 for longer term changes to repair?\n\n","created":"2011-06-24T07:41:16.463+0000"},{"body":"I'm not sure what you mean by \"snapshotting immediately\" or \"polishing our snapshot support\", but one approach that I think is equivalent to that (or maybe that is what you meant by 'snapshotting') would be to grab references to the sstables at the very beginning for each request and use those all throughout the repair. This has however a problem: this means we retain sstables from being deleted during repair, including sstables that are compacted in the meantime. Because repair can take a while, this will be bad. This will also require changes to the wire protocol (because we'll need a way to indicate during streaming the set of sstables to consider), and since we've kind of decided to not do that in minor releases (at least until we've discussed that), this means this cannot be released quickly. Which is bad, because I'm pretty sure this is a good part of the reason why some people with big data sets have had huge pain with repair.\n\nScheduling the validation one by one avoids those problems. In theory this means we'll do less work in parallel, but in practice I doubt this is a big since the goal is probably to have repair have less impact on the node rather than more. It will also make this more easy to reason about.","created":"2011-06-24T07:55:14.971+0000"},{"body":"Attaching patch against 0.8. The patch implements the idea of scheduling the merkle tree requests one by one, to make sure the tree are started as close as possible of \"the same time\". This also put validation compaction in their own executor (to avoid them to be queued up behind standard compactions). That specific executor is created with 2 core threads, to allow for Peter's use case of wanting to do multiple repairs at the same time. That is, by default, you can do 2 repairs involving the same node and be ok. More and you may experience crappy precision in repair. The new concurrent_validators parameter is exposed in case some would want more that 2. That being said, regular compactions and validations are not separated for everything and in particular throttling is shared.\n\nAs far as I can test, this successfully fixes the problems from CASSANDRA-2811 and CASSANDRA-2815. This also don't change anything on the on-wire protocol side, so I think we can target that for 0.8.2.\n\n","created":"2011-06-24T14:05:20.575+0000"},{"body":"This sounds very interesting.\n\nWe have also spotted very noticable issues with full GCs when the merkle trees are passed around. Hopefully this could fix that too.\n\nI will see if I can get this patch tested somewhere if it is ready for that.\n\nOn a side topic, given the importance of getting tombstones properly synchronized within GCGraceSeconds, would it be an potential interesting idea to separate tombstones in different sstables to reduce the need to scan the whole dataset very frequently in the first place?\n\nAnother thought may be to make compaction deterministic or synchronized by a master across nodes so for older data, all we needed was to compare pre-stored md5s of how whole sstables? \n\nThat is, while keeping the masterless design for updates, we could consider a master based design for how older data is being organized by the compactor. so it would be much easier to verify that \"old\" data is the same without any large regular scans and that data is really the same after big compactions etc.\n","created":"2011-06-25T09:03:45.169+0000"},{"body":"bq. We have also spotted very noticable issues with full GCs when the merkle trees are passed around. Hopefully this could fix that too.\n\nThis do make sure that we don't do multiple validation at the same time and that we keep a small number of merkle tree in memory at the same time. So I suppose this could help on the GC side. But overall I don't know if I am too optimistic about that, in part because I'm not sure what causes your issues. But this can't hurt on that side at least.\n\nbq. I will see if I can get this patch tested somewhere if it is ready for that.\n\nI believe it should be ready for that.\n\nbq. would it be an potential interesting idea to separate tombstones in different sstables.\n\nThe thing is that some tombstones may be irrelevant become some update supersedes it (this is specially true of row tombstones). Hence basing a repair on tombstone only may transfer irrelevant data. I suppose it may depend on the use case this will be more or less a big deal. Also, this means that a read will be impacted in that we will often have to hit twice as many sstables. Given that it's not a crazy idea either to want to repair data regularly (if only for durability guarantee), I don't know if it is worth the trouble (we would have to separate tombstones from data at flush time, we'll have to maintain the two separate set of data/tombstone sstables, etc...).\n\nbq. make compaction deterministic or synchronized by a master across nodes\n\nPretty sure we want to avoid going to a master architecture for everything if we can. Having master means that failure handling is more difficult (think network partition for instance) and require leader election and such, and the whole point of the fully distribution of Cassandra is to avoid those. Even without consider those, synchronizing compaction means synchronizing flush somehow and you want to be precise if you're going to use whole sstable md5s, which will be hard and quite probably inefficient.","created":"2011-06-27T09:24:55.891+0000"},{"body":"I don't know what causes GC when doing repairs either, but fire off repair on a few nodes with 100 million docs/node and there is a reasonable chance that a node here and there will log messages about reducing cache sizes due to memory pressure (I am not really sure it is a good idea to do this at all, reducing caches during stress rarely improves anything) or full GC.\n\nThe thought about the master controlled compaction would not really affect network splits etc.\n\nReconciliation after a network split is really as complex with or without a master. We need to get back to a state where all the nodes have the same data anyway which is a complex task anyway.\n\nThis is more a consideration of the fact that we do not necessarily need to live in quorum based world during compaction and we are free to use alternative approaches in the compaction without changing read/write path or affecting availability. Master selection is not really a problem here. Start compaction, talk to other nodes with the same token ranges, select a leader. \n\nDoes not even have to be the same master every time and could consider if we could make compaction part of a background read repair to reduce the amount of times we need to read/write data. \n\nFor instance, if we can verify that the oldest/biggest sstables is 100% in sync with data on other replicas when it is compacted (why not do it during compaction when we go through the data anyway rather than later?),can we use that info to optimize the scans done during repairs by only using data in sstables with data received after some checkpoint in time as the starting point for the consistency check?","created":"2011-06-27T10:36:34.498+0000"},{"body":"bq. I am not really sure it is a good idea to do this at all, reducing caches during stress rarely improves anything\n\n(This is on by default because the most common cause of OOMing is people configuring their caches too large.)\n\nIt sounds odd to me that repair would balloon memory usage dramatically. Do you have monitoring graphs that show the difference in heap usage between \"normal\" and \"repair in progress?\"","created":"2011-06-27T12:21:22.417+0000"},{"body":"This is what heap looks like when GC start slowing things down so much that even gossip gets delayed long enough for nodes to be down for some seconds.\n\n num #instances #bytes class name\n----------------------------------------------\n 1: 9453188 453753024 java.nio.HeapByteBuffer\n 2: 10081546 392167064 [B\n 3: 7616875 243740000 org.apache.cassandra.db.Column\n 4: 9739914 233757936 java.util.concurrent.ConcurrentSkipListMap$Node\n 5: 4131938 99166512 java.util.concurrent.ConcurrentSkipListMap$Index\n 6: 1549230 49575360 org.apache.cassandra.db.DeletedColumn\n\nI guess this really ends up maybe being the mix of everything going on in total and all the reading and writing that may occur when repair runs (valiadation compactions, streaming, normal compactions and regular traffic all at the same time and maybe many CFs at the same time).\n\nHowever, I have suspected for some time that our young size was a bit on the small side and after increasing it and giving the heap a few more GB to work with, it seems like things are behaving quite a bit better.\n\nI mentioned issues with this patch when testing for CASSANDRA-2521. That was a problem caused by me. Was playing around with git for the first time and I manage to apply 2816 to a different branch than the one I used for testing.... :(\n\nMy appologies. \n\nInitial testing with that corrected looks a lot better for my small scale test case, but I noticed one time where I deleted an sstable and restarted. It did not get repaired (repair scanned but did nothing).\n\nNot entirely sure what to make out of that, I then tested to delete another sstable and repair started running.\n\nI will test more over the next days. \n","created":"2011-06-29T16:42:57.542+0000"},{"body":"Things definitely seems to be improved overall, but weird things still happens.\n\nSo... 12 node cluster, this is maybe ugly, I know, but start repair on all of them.\nMost nodes are fine, but one goes crazy. Disk use is now 3-4 times what it was before the repair started, and it does not seem to be done yet.\n\nI have really no idea if this is the case, but I am getting the hunch that this node has ended up streaming out some of the data it is getting in. Would this be possible?\n","created":"2011-07-04T14:02:46.999+0000"},{"body":"bq. So... 12 node cluster, this is maybe ugly, I know, but start repair on all of them.\n\nIs it started on all of them ? If so, this is \"kind of\" expected in the sense that the patch assumes that each node does not do more than 2 repairs (for any column family) at the same time (this is configurable through the new concurrent_validators option, but it's probably better to stick to 2 and stagger the repair). If you do more than that (that is, if you did repair on all node at the same time and RF>2), then we're back on our old demons.\n\nbq. I have really no idea if this is the case, but I am getting the hunch that this node has ended up streaming out some of the data it is getting in. Would this be possible?\n\nNot really. That is, it could be that you create a merkle tree on some data and once you start streaming you, you're picking up data that was just streamed to you and wasn't there when computing the tree. This patch is suppose to fixes this in parts, but this can still happen if you do repairs in parallel on neighboring nodes. However, you shouldn't get into a situation where 2 nodes stream forever because they pick up what is just streamed to them for instance, because what is streaming is determined at the very beginning of the streaming session.\n\nSo my first question would be, was all those repair started in parallel. If yes, you shall not do this :). CASSANDRA-2606 and CASSANDRA-2610 are here to help making the repair of a full cluster much easier (and efficient), but right now it's more about getting patch in one at a time.\nIf the repairs were started one at a time in a rolling fashion, then we do have a unknown problem somewhere.","created":"2011-07-04T14:30:28.122+0000"},{"body":"Cool!\n\nThen you confirmed what I have sort of believed for a while, but my understanding of code has been a bit in conflict with:\nhttp://wiki.apache.org/cassandra/Operations\nwhich says:\n\"It is safe to run repair against multiple machines at the same time, but to minimize the impact on your application workload it is recommended to wait for it to complete on one node before invoking it against the next.\"\n\nI have always read that as \"if you have the HW, go for it!\"\n\nMay I change to:\n\"It is safe to run repair against multiple machines at the same time. However, to minimize the amount of data transferred during a repair, careful synchronization is required between the nodes taking part of the repair. \n\nThis is difficult to do if nodes with the same data replicas runs repair at the same time and doing so can in extreme cases generate excessive transfers of data. \n\nImprovements is being worked on, but for now, avoid scheduling repair on several nodes with replicas of the same data at the same time.\"\n\n","created":"2011-07-04T15:37:17.190+0000"},{"body":"Regardless of change of documentation however, I don't think it should be possible to actually trigger a scenario like this in the first place.\n\nThe system should protect the user from that.\n\nI also noticed that in this case, we have RF3. The node which is going somewhat crazy is number \"6\", however during the repair, it does log that it talks compares and streams data with node 4, 5, 7 and 8.\n\nSeems like a couple of nodes too many?","created":"2011-07-04T23:18:36.363+0000"},{"body":"Regardless of change of documentation however, I don't think it should be possible to actually trigger a scenario like this in the first place.\n\nThe system should protect the user from that.\n\nI also noticed that in this case, we have RF3. The node which is going somewhat crazy is number \"6\", however during the repair, it does log that it talks compares and streams data with node 4, 5, 7 and 8.\n\nSeems like a couple of nodes too many?","created":"2011-07-04T23:18:36.587+0000"},{"body":"bq. May I change to\n\nSure.\n\nbq. The system should protect the user from that\n\nI'm not sure that in a p2p design we can posit an omniscient \"the system.\"","created":"2011-07-05T00:05:38.024+0000"},{"body":"bq.I'm not sure that in a p2p design we can posit an omniscient \"the system.\"\n\nIs that a philosophical statement? :)\n\nAs Cassandra, at least for now, is a p2p network with fairly clearly defined boundaries, I will continue calling it a \"system\" for now :)\n\nHowever, looking at it from the p2p viewpoint, the user potentially have no clue about where replicas are stored and given this, it may be impossible for the user to issue repair manually on more than one node at a time without getting in trouble. Given a large enough p2p setup, it would also be non-trivial to actually schedule a complete repair without ending up with 2 or more repairs running on the same replica set.\n\nSince Cassandra do no checkpoint the synchronization so it is forced to rescan everything on every repair, repairs easily take so long that you are forced to run it on several nodes at a time if you are going to manage to finish repairing all nodes in 10 days...\n\nAnyway, this is way outside the scope of this jira :)","created":"2011-07-05T01:23:06.528+0000"},{"body":"bq. I also noticed that in this case, we have RF3. The node which is going somewhat crazy is number \"6\", however during the repair, it does log that it talks compares and streams data with node 4, 5, 7 and 8.\n\nThis is maybe correct. Node 7 will replicate to node 6 and 8 so 6 and 8 would share data.\n\nSo, to make things safe, even with this patch, every 4th node can run repair at the same time if RF=3?, but you still need to run repair on each of those 4 nodes to make sure it is all repaired?\n\nAs for the comment I made earlier.\n\nTo me, it looks like if the repair start triggering transfers on a large scale, the file the node get streamed in will not be streamed out, but this may get compacted before the repair finished and the compacted file I suspect gets streamed out and the repair just never finishe","created":"2011-07-05T02:07:40.423+0000"},{"body":"bq. The patch implements the idea of scheduling the merkle tree requests one by one, to make sure the tree are started as close as possible of \"the same time\". \n\nCan you point out where this happens in AES?\n\nbq. This also put validation compaction in their own executor (to avoid them to be queued up behind standard compactions). That specific executor is created with 2 core threads, to allow for Peter's use case of wanting to do multiple repairs at the same time. That is, by default, you can do 2 repairs involving the same node and be ok\n\nThat feels like the wrong default to me. I think you can make a case for one (minimal interference with the rest of the system) or unlimited (no weird \"cliff\" to catch the unwary repair operator). But two is weird. :)","created":"2011-07-05T15:42:38.849+0000"},{"body":"bq. Can you point out where this happens in AES?\n\nMostly in AES.rendezvous and AES.RepairSession. Basically, RepairSession creates a queue of jobs, a job representing the repair of a given column family (for a given range, but that comes from the session itself). AES.rendezvous is then call for each received merkleTree. It waits to have all the merkeTree for the first job in the queue. When that is done, it dequeue the job (computing the merkle tree differences and scheduling streaming accordingly) and send the tree request for the next job in the queue.\nMoreover, in StorageService.forceTableRepair(), when scheduling the repair for all the ranges of the node, we actually start the session for the first range and wait for all the \"jobs\" for this range to be done before starting the next session.\n\nbq. That feels like the wrong default to me. I think you can make a case for one (minimal interference with the rest of the system) or unlimited (no weird \"cliff\" to catch the unwary repair operator). But two is weird.\n\nWell the rational was the following one: if you set it to two, then you're saying that as soon as you start 2 repairs in parallel, they will start being inaccurate. But as Peter was suggesting (maybe in another ticket but anyway), if you have huge CF and tiny ones, it's nice to be able to repair on the tiny ones while a repair on the huge one(s) is running. Now, making it unlimited feels dangerous, because if you do so, it means that if the use start a lot of repair, all the validation compaction will start right away. This will kill the cluster (at least a few nodes if all those repair were started on the same node). It sounded better to have degraded precision for repair in those cases rather than basically killing the nodes. Maybe 2 or 4 may be a better default than 2, but 1 is a bit limited and unlimited is clearly much too dangerous.","created":"2011-07-05T16:32:29.536+0000"},{"body":"bq. making it unlimited feels dangerous, because if you do so, it means that if the use start a lot of repair, all the validation compaction will start right away\n\nBut the easy solution is \"don't do that.\"\n\nBy setting a finite number greater than one, you have to restart machines when you realize \"oh, I want to have 3 simultaneous now.\"\n\nI'd rather keep it simple: make it unbounded, no configuration settings. If you ignore the instructions to only run one repair at once, then either you know what you're doing (maybe you have SSDs) or you will find out very quickly and never do it again. :)","created":"2011-07-13T21:30:21.502+0000"},{"body":"I'm kinda +1 on the simple version w/o bounds but not too fussy since I can obviously set it very high for my use case. In any case, the most important part for mixed small/large type of situation is that concurrent repair is possible, even if configuration changes are needed.\n","created":"2011-07-13T21:57:30.981+0000"},{"body":"bq. I'm kinda +1 on the simple version w/o bounds\n\nMe too. +1 with that change.","created":"2011-07-18T13:09:59.120+0000"},{"body":"(I'll go ahead and submit a version with that change.)","created":"2011-07-18T13:24:09.845+0000"},{"body":"rebased and switched to unbounded executor for validations.\n\ntests do not compile but I believe that was already the case w/ v1 -- not sure what to do with blockUntilRunning, which was removed.","created":"2011-07-18T14:26:49.097+0000"},{"body":"bq. rebased\n\nYou rebased it against trunk while the fix version is still 0.8.2. I agree that this feel a bigger change that what we would want for a minor release, but repair is really a major pain point for users. And to tell the truth, it's worth in 0.8 than it is in 0.7 because even though the problem this patch solves exists in 0.7 as well, the splitting into range of repair made for 0.8 exacerbate those problems. Moreover this patch is fairly well delimited in what it changes, and it don't make any change to the wire protocol or anything that would make upgrades a problem. So I'm actually in favor of taking the small risk of pushing that in 0.8.2 (and be very vigilant to test repair extensively before the release). So for now, attaching a rebase with tests fixed against 0.8.\n\nbq. and switched to unbounded executor for validations.\n\nOk, I realize that I'm not sure I understand what you mean by unbounded executor. In you rebased version, the ValidationExecutor apparently use the first constructor of DebuggableThreadPoolExecutor that will construct a mono-threaded executor, which is not what we want. Sure the queue of the executor will be unbounded, but if that was the problem, there was a misunderstanding, because the queue has always been unbounded, even in my initial patch. What concurrent_validators was dealing with is the number of core threads. And we need multiple core threads if we want to allow multiple concurrent repairs to work correctly.\n\nNow the idea behind a default of 2 for the core threads was because I see a reason to want 2 concurrent repairs, but I don't see a very good reason to want more (and it's configurable if someone really need more). I'm glad to say it's not a marvelous default and the patch I've just attached used the same default than concurrent_compactors which is maybe less \"random\". Now we could have an executor with unbounded threads, that spawn a thread if needed making sure we never queue validation compaction, but that seems a little bit dangerous to me. It seems more reasonable to me to have a (configurable) reasonable number of threads and let validation queue up if the user start more than that number of concurrent repair (which will impact the precision of those repair, but it would be the user fault and it's a better way to deal with such fault than starting an unreasonable number of validation compaction that will starve memory (on likely more than one node btw)). But if you still think that it's better to have an unbounded number of threads, I won't fight over this.\n\nbq. tests do not compile but I believe that was already the case w/ v1\n\nYes, I completely forgot to update the unit tests, sorry. Attached patch fixes those.\n","created":"2011-07-20T10:05:25.175+0000"},{"body":"v4 attached against 0.8 with a corrected uncapped validation executor.","created":"2011-07-20T17:11:43.000+0000"},{"body":"I think that if we don't want validation executor of v4 to ever queue tasks (which is what we need), then we need the executor queue to be a bounded queue of size 0 (i.e. that doesn't accept element). Indeed, as per the documentation of ThreadPoolExecutor:\n{noformat}\nIf corePoolSize or more threads are running, the Executor always prefers queuing a request rather than adding a new thread.\n{noformat} ","created":"2011-07-20T17:45:56.199+0000"},{"body":"You're right. v5 attached.","created":"2011-07-20T18:08:00.796+0000"},{"body":"Alright, v5 looks good to me. Committed, thanks.","created":"2011-07-21T11:12:09.745+0000"}],"conversations":[{"body":"Being a little slow, I just realized after having opened CASSANDRA-2811 and CASSANDRA-2815 that there is a more general problem with repair.\n\nWhen a repair is started, it will send a number of merkle tree to its neighbor as well as himself and assume for correction that the building of those trees will be started on every node roughly at the same time (if not, we end up comparing data snapshot at different time and will thus mistakenly repair a lot of useless data). This is bogus for many reasons:\n* Because validation compaction runs on the same executor that other compaction, the start of the validation on the different node is subject to other compactions. 0.8 mitigates this in a way by being multi-threaded (and thus there is less change to be blocked a long time by a long running compaction), but the compaction executor being bounded, its still a problem)\n* if you run a nodetool repair without arguments, it will repair every CFs. As a consequence it will generate lots of merkle tree requests and all of those requests will be issued at the same time. Because even in 0.8 the compaction executor is bounded, some of those validations will end up being queued behind the first ones. Even assuming that the different validation are submitted in the same order on each node (which isn't guaranteed either), there is no guarantee that on all nodes, the first validation will take the same time, hence desynchronizing the queued ones.\n\nOverall, it is important for the precision of repair that for a given CF and range (which is the unit at which trees are computed), we make sure that all node will start the validation at the same time (or, since we can't do magic, as close as possible).\n\nOne (reasonably simple) proposition to fix this would be to have repair schedule validation compactions across nodes one by one (i.e, one CF/range at a time), waiting for all nodes to return their tree before submitting the next request. Then on each node, we should make sure that the node will start the validation compaction as soon as requested. For that, we probably want to have a specific executor for validation compaction and:\n* either we fail the whole repair whenever one node is not able to execute the validation compaction right away (because no thread are available right away).\n* we simply tell the user that if he start too many repairs in parallel, he may start seeing some of those repairing more data than it should.\n","from":"reporter","subject":"Repair doesn't synchronize merkle tree creation properly"},{"body":"I guess a dedicated validation executor is ok as long as it still obeys the global \"compaction\" i/o limit.","from":"developer"},{"body":"I'm a fan of the snapshotting immediately after receiving the request approach. In general, polishing our snapshot support to allow for this kind of usecase is likely to open up other interesting possibilities.","from":"developer"},{"body":"Supporting actual live-reading of snapshotted sstables is a little more than \"polishing.\" It would be cool, but I wouldn't want it to block fixing repair.","from":"developer"},{"body":"I've thought about this problem too, and it is really significant for some use-cases. Again because so few writes are needed in order to trigger large amounts of data being sent given the merklee tree granularity.\n\nWhile I'm all for a fixing it by e.g. more immediate snapshotting, I would like to raise the issue that repairs overall have pretty significant side-effects; particularly ones that can self-magnify and cause further problems. Beyond the obvious \"it does disk I/O\" and \"It uses CPU\", we have:\n\n* Over-repair due to merklee tree granularity can cause jumps in CF sizes, killing cache locality\n* Combine that with concurrent repairs then repairing the \"size-jumped\" set of sstables and you can magnify that effect on other nodes causing huge size increases.\n* Up to recently, mixing large and small cf:s was a significant problem if you wanted to have different frequencies and different gc grace times, due to one repair blocking on another. But fixes to this and the other JIRA about concurrency, might disable the \"fix\" for that that was concurrent compaction - so back to square one.\n\nI guess overall, it seems very easy to shoot yourself in the foot with repair.\n\nAny opinions on CASSANDRA-2699 for longer term changes to repair?\n\n","from":"developer"},{"body":"I'm not sure what you mean by \"snapshotting immediately\" or \"polishing our snapshot support\", but one approach that I think is equivalent to that (or maybe that is what you meant by 'snapshotting') would be to grab references to the sstables at the very beginning for each request and use those all throughout the repair. This has however a problem: this means we retain sstables from being deleted during repair, including sstables that are compacted in the meantime. Because repair can take a while, this will be bad. This will also require changes to the wire protocol (because we'll need a way to indicate during streaming the set of sstables to consider), and since we've kind of decided to not do that in minor releases (at least until we've discussed that), this means this cannot be released quickly. Which is bad, because I'm pretty sure this is a good part of the reason why some people with big data sets have had huge pain with repair.\n\nScheduling the validation one by one avoids those problems. In theory this means we'll do less work in parallel, but in practice I doubt this is a big since the goal is probably to have repair have less impact on the node rather than more. It will also make this more easy to reason about.","from":"developer"},{"body":"Attaching patch against 0.8. The patch implements the idea of scheduling the merkle tree requests one by one, to make sure the tree are started as close as possible of \"the same time\". This also put validation compaction in their own executor (to avoid them to be queued up behind standard compactions). That specific executor is created with 2 core threads, to allow for Peter's use case of wanting to do multiple repairs at the same time. That is, by default, you can do 2 repairs involving the same node and be ok. More and you may experience crappy precision in repair. The new concurrent_validators parameter is exposed in case some would want more that 2. That being said, regular compactions and validations are not separated for everything and in particular throttling is shared.\n\nAs far as I can test, this successfully fixes the problems from CASSANDRA-2811 and CASSANDRA-2815. This also don't change anything on the on-wire protocol side, so I think we can target that for 0.8.2.\n\n","from":"developer"},{"body":"This sounds very interesting.\n\nWe have also spotted very noticable issues with full GCs when the merkle trees are passed around. Hopefully this could fix that too.\n\nI will see if I can get this patch tested somewhere if it is ready for that.\n\nOn a side topic, given the importance of getting tombstones properly synchronized within GCGraceSeconds, would it be an potential interesting idea to separate tombstones in different sstables to reduce the need to scan the whole dataset very frequently in the first place?\n\nAnother thought may be to make compaction deterministic or synchronized by a master across nodes so for older data, all we needed was to compare pre-stored md5s of how whole sstables? \n\nThat is, while keeping the masterless design for updates, we could consider a master based design for how older data is being organized by the compactor. so it would be much easier to verify that \"old\" data is the same without any large regular scans and that data is really the same after big compactions etc.\n","from":"developer"},{"body":"bq. We have also spotted very noticable issues with full GCs when the merkle trees are passed around. Hopefully this could fix that too.\n\nThis do make sure that we don't do multiple validation at the same time and that we keep a small number of merkle tree in memory at the same time. So I suppose this could help on the GC side. But overall I don't know if I am too optimistic about that, in part because I'm not sure what causes your issues. But this can't hurt on that side at least.\n\nbq. I will see if I can get this patch tested somewhere if it is ready for that.\n\nI believe it should be ready for that.\n\nbq. would it be an potential interesting idea to separate tombstones in different sstables.\n\nThe thing is that some tombstones may be irrelevant become some update supersedes it (this is specially true of row tombstones). Hence basing a repair on tombstone only may transfer irrelevant data. I suppose it may depend on the use case this will be more or less a big deal. Also, this means that a read will be impacted in that we will often have to hit twice as many sstables. Given that it's not a crazy idea either to want to repair data regularly (if only for durability guarantee), I don't know if it is worth the trouble (we would have to separate tombstones from data at flush time, we'll have to maintain the two separate set of data/tombstone sstables, etc...).\n\nbq. make compaction deterministic or synchronized by a master across nodes\n\nPretty sure we want to avoid going to a master architecture for everything if we can. Having master means that failure handling is more difficult (think network partition for instance) and require leader election and such, and the whole point of the fully distribution of Cassandra is to avoid those. Even without consider those, synchronizing compaction means synchronizing flush somehow and you want to be precise if you're going to use whole sstable md5s, which will be hard and quite probably inefficient.","from":"developer"},{"body":"I don't know what causes GC when doing repairs either, but fire off repair on a few nodes with 100 million docs/node and there is a reasonable chance that a node here and there will log messages about reducing cache sizes due to memory pressure (I am not really sure it is a good idea to do this at all, reducing caches during stress rarely improves anything) or full GC.\n\nThe thought about the master controlled compaction would not really affect network splits etc.\n\nReconciliation after a network split is really as complex with or without a master. We need to get back to a state where all the nodes have the same data anyway which is a complex task anyway.\n\nThis is more a consideration of the fact that we do not necessarily need to live in quorum based world during compaction and we are free to use alternative approaches in the compaction without changing read/write path or affecting availability. Master selection is not really a problem here. Start compaction, talk to other nodes with the same token ranges, select a leader. \n\nDoes not even have to be the same master every time and could consider if we could make compaction part of a background read repair to reduce the amount of times we need to read/write data. \n\nFor instance, if we can verify that the oldest/biggest sstables is 100% in sync with data on other replicas when it is compacted (why not do it during compaction when we go through the data anyway rather than later?),can we use that info to optimize the scans done during repairs by only using data in sstables with data received after some checkpoint in time as the starting point for the consistency check?","from":"developer"},{"body":"bq. I am not really sure it is a good idea to do this at all, reducing caches during stress rarely improves anything\n\n(This is on by default because the most common cause of OOMing is people configuring their caches too large.)\n\nIt sounds odd to me that repair would balloon memory usage dramatically. Do you have monitoring graphs that show the difference in heap usage between \"normal\" and \"repair in progress?\"","from":"developer"},{"body":"This is what heap looks like when GC start slowing things down so much that even gossip gets delayed long enough for nodes to be down for some seconds.\n\n num #instances #bytes class name\n----------------------------------------------\n 1: 9453188 453753024 java.nio.HeapByteBuffer\n 2: 10081546 392167064 [B\n 3: 7616875 243740000 org.apache.cassandra.db.Column\n 4: 9739914 233757936 java.util.concurrent.ConcurrentSkipListMap$Node\n 5: 4131938 99166512 java.util.concurrent.ConcurrentSkipListMap$Index\n 6: 1549230 49575360 org.apache.cassandra.db.DeletedColumn\n\nI guess this really ends up maybe being the mix of everything going on in total and all the reading and writing that may occur when repair runs (valiadation compactions, streaming, normal compactions and regular traffic all at the same time and maybe many CFs at the same time).\n\nHowever, I have suspected for some time that our young size was a bit on the small side and after increasing it and giving the heap a few more GB to work with, it seems like things are behaving quite a bit better.\n\nI mentioned issues with this patch when testing for CASSANDRA-2521. That was a problem caused by me. Was playing around with git for the first time and I manage to apply 2816 to a different branch than the one I used for testing.... :(\n\nMy appologies. \n\nInitial testing with that corrected looks a lot better for my small scale test case, but I noticed one time where I deleted an sstable and restarted. It did not get repaired (repair scanned but did nothing).\n\nNot entirely sure what to make out of that, I then tested to delete another sstable and repair started running.\n\nI will test more over the next days. \n","from":"developer"},{"body":"Things definitely seems to be improved overall, but weird things still happens.\n\nSo... 12 node cluster, this is maybe ugly, I know, but start repair on all of them.\nMost nodes are fine, but one goes crazy. Disk use is now 3-4 times what it was before the repair started, and it does not seem to be done yet.\n\nI have really no idea if this is the case, but I am getting the hunch that this node has ended up streaming out some of the data it is getting in. Would this be possible?\n","from":"developer"},{"body":"bq. So... 12 node cluster, this is maybe ugly, I know, but start repair on all of them.\n\nIs it started on all of them ? If so, this is \"kind of\" expected in the sense that the patch assumes that each node does not do more than 2 repairs (for any column family) at the same time (this is configurable through the new concurrent_validators option, but it's probably better to stick to 2 and stagger the repair). If you do more than that (that is, if you did repair on all node at the same time and RF>2), then we're back on our old demons.\n\nbq. I have really no idea if this is the case, but I am getting the hunch that this node has ended up streaming out some of the data it is getting in. Would this be possible?\n\nNot really. That is, it could be that you create a merkle tree on some data and once you start streaming you, you're picking up data that was just streamed to you and wasn't there when computing the tree. This patch is suppose to fixes this in parts, but this can still happen if you do repairs in parallel on neighboring nodes. However, you shouldn't get into a situation where 2 nodes stream forever because they pick up what is just streamed to them for instance, because what is streaming is determined at the very beginning of the streaming session.\n\nSo my first question would be, was all those repair started in parallel. If yes, you shall not do this :). CASSANDRA-2606 and CASSANDRA-2610 are here to help making the repair of a full cluster much easier (and efficient), but right now it's more about getting patch in one at a time.\nIf the repairs were started one at a time in a rolling fashion, then we do have a unknown problem somewhere.","from":"developer"},{"body":"Cool!\n\nThen you confirmed what I have sort of believed for a while, but my understanding of code has been a bit in conflict with:\nhttp://wiki.apache.org/cassandra/Operations\nwhich says:\n\"It is safe to run repair against multiple machines at the same time, but to minimize the impact on your application workload it is recommended to wait for it to complete on one node before invoking it against the next.\"\n\nI have always read that as \"if you have the HW, go for it!\"\n\nMay I change to:\n\"It is safe to run repair against multiple machines at the same time. However, to minimize the amount of data transferred during a repair, careful synchronization is required between the nodes taking part of the repair. \n\nThis is difficult to do if nodes with the same data replicas runs repair at the same time and doing so can in extreme cases generate excessive transfers of data. \n\nImprovements is being worked on, but for now, avoid scheduling repair on several nodes with replicas of the same data at the same time.\"\n\n","from":"developer"},{"body":"Regardless of change of documentation however, I don't think it should be possible to actually trigger a scenario like this in the first place.\n\nThe system should protect the user from that.\n\nI also noticed that in this case, we have RF3. The node which is going somewhat crazy is number \"6\", however during the repair, it does log that it talks compares and streams data with node 4, 5, 7 and 8.\n\nSeems like a couple of nodes too many?","from":"developer"},{"body":"Regardless of change of documentation however, I don't think it should be possible to actually trigger a scenario like this in the first place.\n\nThe system should protect the user from that.\n\nI also noticed that in this case, we have RF3. The node which is going somewhat crazy is number \"6\", however during the repair, it does log that it talks compares and streams data with node 4, 5, 7 and 8.\n\nSeems like a couple of nodes too many?","from":"developer"},{"body":"bq. May I change to\n\nSure.\n\nbq. The system should protect the user from that\n\nI'm not sure that in a p2p design we can posit an omniscient \"the system.\"","from":"developer"},{"body":"bq.I'm not sure that in a p2p design we can posit an omniscient \"the system.\"\n\nIs that a philosophical statement? :)\n\nAs Cassandra, at least for now, is a p2p network with fairly clearly defined boundaries, I will continue calling it a \"system\" for now :)\n\nHowever, looking at it from the p2p viewpoint, the user potentially have no clue about where replicas are stored and given this, it may be impossible for the user to issue repair manually on more than one node at a time without getting in trouble. Given a large enough p2p setup, it would also be non-trivial to actually schedule a complete repair without ending up with 2 or more repairs running on the same replica set.\n\nSince Cassandra do no checkpoint the synchronization so it is forced to rescan everything on every repair, repairs easily take so long that you are forced to run it on several nodes at a time if you are going to manage to finish repairing all nodes in 10 days...\n\nAnyway, this is way outside the scope of this jira :)","from":"developer"},{"body":"bq. I also noticed that in this case, we have RF3. The node which is going somewhat crazy is number \"6\", however during the repair, it does log that it talks compares and streams data with node 4, 5, 7 and 8.\n\nThis is maybe correct. Node 7 will replicate to node 6 and 8 so 6 and 8 would share data.\n\nSo, to make things safe, even with this patch, every 4th node can run repair at the same time if RF=3?, but you still need to run repair on each of those 4 nodes to make sure it is all repaired?\n\nAs for the comment I made earlier.\n\nTo me, it looks like if the repair start triggering transfers on a large scale, the file the node get streamed in will not be streamed out, but this may get compacted before the repair finished and the compacted file I suspect gets streamed out and the repair just never finishe","from":"developer"},{"body":"bq. The patch implements the idea of scheduling the merkle tree requests one by one, to make sure the tree are started as close as possible of \"the same time\". \n\nCan you point out where this happens in AES?\n\nbq. This also put validation compaction in their own executor (to avoid them to be queued up behind standard compactions). That specific executor is created with 2 core threads, to allow for Peter's use case of wanting to do multiple repairs at the same time. That is, by default, you can do 2 repairs involving the same node and be ok\n\nThat feels like the wrong default to me. I think you can make a case for one (minimal interference with the rest of the system) or unlimited (no weird \"cliff\" to catch the unwary repair operator). But two is weird. :)","from":"developer"},{"body":"bq. Can you point out where this happens in AES?\n\nMostly in AES.rendezvous and AES.RepairSession. Basically, RepairSession creates a queue of jobs, a job representing the repair of a given column family (for a given range, but that comes from the session itself). AES.rendezvous is then call for each received merkleTree. It waits to have all the merkeTree for the first job in the queue. When that is done, it dequeue the job (computing the merkle tree differences and scheduling streaming accordingly) and send the tree request for the next job in the queue.\nMoreover, in StorageService.forceTableRepair(), when scheduling the repair for all the ranges of the node, we actually start the session for the first range and wait for all the \"jobs\" for this range to be done before starting the next session.\n\nbq. That feels like the wrong default to me. I think you can make a case for one (minimal interference with the rest of the system) or unlimited (no weird \"cliff\" to catch the unwary repair operator). But two is weird.\n\nWell the rational was the following one: if you set it to two, then you're saying that as soon as you start 2 repairs in parallel, they will start being inaccurate. But as Peter was suggesting (maybe in another ticket but anyway), if you have huge CF and tiny ones, it's nice to be able to repair on the tiny ones while a repair on the huge one(s) is running. Now, making it unlimited feels dangerous, because if you do so, it means that if the use start a lot of repair, all the validation compaction will start right away. This will kill the cluster (at least a few nodes if all those repair were started on the same node). It sounded better to have degraded precision for repair in those cases rather than basically killing the nodes. Maybe 2 or 4 may be a better default than 2, but 1 is a bit limited and unlimited is clearly much too dangerous.","from":"developer"},{"body":"bq. making it unlimited feels dangerous, because if you do so, it means that if the use start a lot of repair, all the validation compaction will start right away\n\nBut the easy solution is \"don't do that.\"\n\nBy setting a finite number greater than one, you have to restart machines when you realize \"oh, I want to have 3 simultaneous now.\"\n\nI'd rather keep it simple: make it unbounded, no configuration settings. If you ignore the instructions to only run one repair at once, then either you know what you're doing (maybe you have SSDs) or you will find out very quickly and never do it again. :)","from":"developer"},{"body":"I'm kinda +1 on the simple version w/o bounds but not too fussy since I can obviously set it very high for my use case. In any case, the most important part for mixed small/large type of situation is that concurrent repair is possible, even if configuration changes are needed.\n","from":"developer"},{"body":"bq. I'm kinda +1 on the simple version w/o bounds\n\nMe too. +1 with that change.","from":"developer"},{"body":"(I'll go ahead and submit a version with that change.)","from":"developer"},{"body":"rebased and switched to unbounded executor for validations.\n\ntests do not compile but I believe that was already the case w/ v1 -- not sure what to do with blockUntilRunning, which was removed.","from":"developer"},{"body":"bq. rebased\n\nYou rebased it against trunk while the fix version is still 0.8.2. I agree that this feel a bigger change that what we would want for a minor release, but repair is really a major pain point for users. And to tell the truth, it's worth in 0.8 than it is in 0.7 because even though the problem this patch solves exists in 0.7 as well, the splitting into range of repair made for 0.8 exacerbate those problems. Moreover this patch is fairly well delimited in what it changes, and it don't make any change to the wire protocol or anything that would make upgrades a problem. So I'm actually in favor of taking the small risk of pushing that in 0.8.2 (and be very vigilant to test repair extensively before the release). So for now, attaching a rebase with tests fixed against 0.8.\n\nbq. and switched to unbounded executor for validations.\n\nOk, I realize that I'm not sure I understand what you mean by unbounded executor. In you rebased version, the ValidationExecutor apparently use the first constructor of DebuggableThreadPoolExecutor that will construct a mono-threaded executor, which is not what we want. Sure the queue of the executor will be unbounded, but if that was the problem, there was a misunderstanding, because the queue has always been unbounded, even in my initial patch. What concurrent_validators was dealing with is the number of core threads. And we need multiple core threads if we want to allow multiple concurrent repairs to work correctly.\n\nNow the idea behind a default of 2 for the core threads was because I see a reason to want 2 concurrent repairs, but I don't see a very good reason to want more (and it's configurable if someone really need more). I'm glad to say it's not a marvelous default and the patch I've just attached used the same default than concurrent_compactors which is maybe less \"random\". Now we could have an executor with unbounded threads, that spawn a thread if needed making sure we never queue validation compaction, but that seems a little bit dangerous to me. It seems more reasonable to me to have a (configurable) reasonable number of threads and let validation queue up if the user start more than that number of concurrent repair (which will impact the precision of those repair, but it would be the user fault and it's a better way to deal with such fault than starting an unreasonable number of validation compaction that will starve memory (on likely more than one node btw)). But if you still think that it's better to have an unbounded number of threads, I won't fight over this.\n\nbq. tests do not compile but I believe that was already the case w/ v1\n\nYes, I completely forgot to update the unit tests, sorry. Attached patch fixes those.\n","from":"developer"},{"body":"v4 attached against 0.8 with a corrected uncapped validation executor.","from":"developer"},{"body":"I think that if we don't want validation executor of v4 to ever queue tasks (which is what we need), then we need the executor queue to be a bounded queue of size 0 (i.e. that doesn't accept element). Indeed, as per the documentation of ThreadPoolExecutor:\n{noformat}\nIf corePoolSize or more threads are running, the Executor always prefers queuing a request rather than adding a new thread.\n{noformat} ","from":"developer"},{"body":"You're right. v5 attached.","from":"developer"},{"body":"Alright, v5 looks good to me. Committed, thanks.","from":"developer"}],"created":"2011-06-23T11:35:30.000+0000","description":"Being a little slow, I just realized after having opened CASSANDRA-2811 and CASSANDRA-2815 that there is a more general problem with repair.\n\nWhen a repair is started, it will send a number of merkle tree to its neighbor as well as himself and assume for correction that the building of those trees will be started on every node roughly at the same time (if not, we end up comparing data snapshot at different time and will thus mistakenly repair a lot of useless data). This is bogus for many reasons:\n* Because validation compaction runs on the same executor that other compaction, the start of the validation on the different node is subject to other compactions. 0.8 mitigates this in a way by being multi-threaded (and thus there is less change to be blocked a long time by a long running compaction), but the compaction executor being bounded, its still a problem)\n* if you run a nodetool repair without arguments, it will repair every CFs. As a consequence it will generate lots of merkle tree requests and all of those requests will be issued at the same time. Because even in 0.8 the compaction executor is bounded, some of those validations will end up being queued behind the first ones. Even assuming that the different validation are submitted in the same order on each node (which isn't guaranteed either), there is no guarantee that on all nodes, the first validation will take the same time, hence desynchronizing the queued ones.\n\nOverall, it is important for the precision of repair that for a given CF and range (which is the unit at which trees are computed), we make sure that all node will start the validation at the same time (or, since we can't do magic, as close as possible).\n\nOne (reasonably simple) proposition to fix this would be to have repair schedule validation compactions across nodes one by one (i.e, one CF/range at a time), waiting for all nodes to return their tree before submitting the next request. Then on each node, we should make sure that the node will start the validation compaction as soon as requested. For that, we probably want to have a specific executor for validation compaction and:\n* either we fail the whole repair whenever one node is not able to execute the validation compaction right away (because no thread are available right away).\n* we simply tell the user that if he start too many repairs in parallel, he may start seeing some of those repairing more data than it should.\n","issue_id":"12511343","key":"CASSANDRA-2816","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-07-21T11:13:13.000+0000","role":"fixed_distractor","summary":"Repair doesn't synchronize merkle tree creation properly"} {"case_id":"12512969","cluster":"DISTRACTOR-CASSANDRA-2863","comments":[{"body":"I don't now if it's the same one, buy I got another during repair on another node:\n\nERROR [Thread-1710] 2011-07-08 21:21:00,514 AbstractCassandraDaemon.java (line 113) Fatal exception in thread Thread[Thread-1710,5,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n\tat org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:154)\n\tat org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:63)\n\tat org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:162)\n\tat org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:95)\nCaused by: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n\tat java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n\tat java.util.concurrent.FutureTask.get(FutureTask.java:83)\n\tat org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:138)\n\t... 3 more\nCaused by: java.lang.NullPointerException\n\tat org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n\tat org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n\tat org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n\tat org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1103)\n\tat org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1094)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\n","created":"2011-07-08T21:22:52.785+0000"},{"body":"I'm a little bit baffled by that one. Trusting the stack trace, apparently when SSTW.RowIndexer.close() is called, the iwriter field is null. But iwriter is set in prepareIndexing() that is called the line before index() in SSTW.Builder. Thus if an exception happens in prepareIndexing, we shouldn't arrive to the index() method (which is the one triggering the close()). And looking at the use of iwriter, no other line set it (so it can't be set back to null after prepareIndexing()).\n\nSo I mean we can add a {{if (iwriter != null)}} before calling the close, but the truth is I have no clue how it could ever be null at that point.\n\nHéctor: are you positive that you are using stock 0.8.1 ?","created":"2011-07-21T12:28:09.977+0000"},{"body":"I have a patch from CASSANDRA-2818 (2818-v4) applied, if that's of any help. The patch only touches messaging classes though.","created":"2011-07-21T12:35:30.240+0000"},{"body":"Doesn't make any sense to me, either. The only place close() is called is from index() [as seen in the stacktrace here] and the only place index() is called is after prepareIndexing, which sets iwriter to non-null:\n\n{code}\n long estimatedRows = indexer.prepareIndexing();\n\n // build the index and filter\n long rows = indexer.index();\n{code}\n","created":"2011-07-21T20:13:00.772+0000"},{"body":"I just got a similar-looking NPE stack trace. The node that raised the exception was receiving streams from a node being decommissioned (with \"nodetool decommission\"). Both nodes were running 0.8.6; both had been upgraded a few hours earlier from 0.8.4.\n\nThe first: \n\nERROR [CompactionExecutor:72] 2011-09-20 22:34:20,892 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[CompactionExecutor:72,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n\nThen half a minute later:\n\n INFO [CompactionExecutor:72] 2011-09-20 22:34:46,923 SSTableReader.java (line 162) Opening /mnt/cassandra/data/Keyspace/CF1-g-1536\nERROR [Thread-785] 2011-09-20 22:34:52,054 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[Thread-785,5,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:154)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:63)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:189)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:117)\nCaused by: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:138)\n ... 3 more\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n","created":"2011-09-20T23:19:45.573+0000"},{"body":"So I restarted the node making the NPE above and ran the decommission again (on the other node) and it successfully finished decommissioning. A few hours later I tried to decommission another node and received this error, on a separate node from the two mentioned before, which is a new node running 0.8.6 for just about a day now.\n\nINFO [COMMIT-LOG-WRITER] 2011-09-21 05:12:46,361 CommitLogSegment.java (line 58) Creating new commitlog segment /raid0/cassandra/commitlog/CommitLog-1316581966361.log\nERROR [Thread-224] 2011-09-21 05:12:52,738 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[Thread-224,5,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:154)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:63)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:189)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:117)\nCaused by: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:138)\n ... 3 more\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n","created":"2011-09-21T14:58:23.897+0000"},{"body":"Since my last two comments, I increased my RF and started a rolling repair on my nodes. This has caused this NPE to pop up on all the boxes over the last couple of days as they process SSTables. Again, all the nodes are fresh 0.8.6 installs from a few days ago using the ComboAMI. I've seen backtraces like the one below appear at least a couple of times on each node in my cluster as I was repairing..\n\n INFO [CompactionExecutor:648] 2011-09-22 04:35:51,086 SSTableReader.java (line 162) Opening /raid0/cassandra/data/Keyspace/CF1-g-535\n INFO [CompactionExecutor:648] 2011-09-22 04:35:51,172 SSTableReader.java (line 162) Opening /raid0/cassandra/data/Keyspace/CF2-g-350\n INFO [CompactionExecutor:648] 2011-09-22 04:36:01,721 SSTableReader.java (line 162) Opening /raid0/cassandra/data/Keyspace/CF3-g-456\nERROR [Thread-3658] 2011-09-22 04:36:04,821 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[Thread-3658,5,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:154)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:63)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:189)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:117)\nCaused by: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:138)\n ... 3 more\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nERROR [CompactionExecutor:648] 2011-09-22 04:36:04,823 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[CompactionExecutor:648,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n","created":"2011-09-23T15:17:58.652+0000"},{"body":"Here's another from a node that was streaming out:\n\n\n INFO [AntiEntropyStage:1] 2011-09-22 17:56:33,737 StreamOut.java (line 181) Stream context metadata [/raid0/cassandra/data/Keyspace/CF1-g-315-Data.db sections=1073 progress=0/358\n INFO [AntiEntropyStage:1] 2011-09-22 17:56:33,738 StreamOutSession.java (line 174) Streaming to /10.207.38.196\n INFO [ValidationExecutor:14] 2011-09-22 17:56:33,740 ColumnFamilyStore.java (line 1128) Enqueuing flush of Memtable-CF4@516894197(9505947/21294113 serialized/live bytes, 5684 ops)\n INFO [FlushWriter:621] 2011-09-22 17:56:33,743 Memtable.java (line 237) Writing Memtable-CF4@516894197(9505947/21294113 serialized/live bytes, 5684 ops)\n INFO [FlushWriter:621] 2011-09-22 17:56:33,880 Memtable.java (line 254) Completed flushing /raid0/cassandra/data/Keyspace/CF4-g-434-Data.db (9914202 bytes)\n INFO [CompactionExecutor:884] 2011-09-22 17:56:45,912 CompactionManager.java (line 608) Compacted to /raid0/cassandra/data/Keyspace/CF4-tmp-g-332-Data.db. 32,824,572,147 to 18,506,0\nERROR [CompactionExecutor:884] 2011-09-22 17:56:46,134 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[CompactionExecutor:884,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n","created":"2011-09-23T15:20:42.363+0000"},{"body":"Getting the following after trying to 'move' on 0.8.5:\n\n{code}\nERROR [CompactionExecutor:213] 2011-09-23 14:02:26,571\nAbstractCassandraDaemon.java (line 139) Fatal exception in thread\nThread[CompactionExecutor:213,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n{code}","created":"2011-09-23T21:12:26.786+0000"},{"body":"Also got the same NPE when doing a repair on a different node. This time, the NPE was preceded by a different NPE:\n\n\n{code}\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:154)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:63)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:177)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:114)\nCaused by: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:138)\n ... 3 more\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n\n{code}\n","created":"2011-09-24T05:25:56.794+0000"},{"body":"We've seen the same error via autobootstrap.","created":"2011-10-07T17:40:08.596+0000"},{"body":"Joaquin, can you give steps to reproduce?","created":"2011-10-07T18:01:11.475+0000"},{"body":"Alright, I've tried boostrapping nodes a few times and I'm still not able to reproduce, so it's likely a race of something.\n\nThat being said, I still have no clue how the iwriter field in SSTableWriter.RowIndexer can be null where the stack indicates it to be null. The assignment happens before the access, there is no other assignment of that field and the assignment/access are in the same thread.\n\nAnyway, what we can easily do is to make iwrite being final and assign it in the constructor. If that doesn't make that field being non-null, I don't know what will. Patch attached to make that change.","created":"2011-10-10T11:07:16.276+0000"},{"body":"First patch was buggy in that it was now creating the index file *before* we assert that this file doesn't exist.\n\nv2 just moves the maybeOpenIndexer call after the asserts.","created":"2011-10-10T13:07:11.798+0000"},{"body":"+1 v2","created":"2011-10-10T13:12:54.581+0000"},{"body":"Committed and marking resolved for now. If that errors come up in another form, we'll see then.","created":"2011-10-10T13:59:13.893+0000"}],"conversations":[{"body":"A NPE is generated during repair when closing an sstable generated via SSTable build. It doesn't happen always. The node had been scrubbed and compacted before calling repair.\n\n INFO [CompactionExecutor:2] 2011-07-06 11:11:32,640 SSTableReader.java (line 158) Opening /d2/cassandra/data/sbs/walf-g-730\nERROR [CompactionExecutor:2] 2011-07-06 11:11:34,327 AbstractCassandraDaemon.java (line 113) Fatal exception in thread Thread[CompactionExecutor:2,1,main] \njava.lang.NullPointerException\n\tat org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n\tat org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n\tat org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n\tat org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1103)\n\tat org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1094)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\n","from":"reporter","subject":"NPE when writing SSTable generated via repair"},{"body":"I don't now if it's the same one, buy I got another during repair on another node:\n\nERROR [Thread-1710] 2011-07-08 21:21:00,514 AbstractCassandraDaemon.java (line 113) Fatal exception in thread Thread[Thread-1710,5,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n\tat org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:154)\n\tat org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:63)\n\tat org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:162)\n\tat org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:95)\nCaused by: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n\tat java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n\tat java.util.concurrent.FutureTask.get(FutureTask.java:83)\n\tat org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:138)\n\t... 3 more\nCaused by: java.lang.NullPointerException\n\tat org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n\tat org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n\tat org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n\tat org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1103)\n\tat org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1094)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\n","from":"developer"},{"body":"I'm a little bit baffled by that one. Trusting the stack trace, apparently when SSTW.RowIndexer.close() is called, the iwriter field is null. But iwriter is set in prepareIndexing() that is called the line before index() in SSTW.Builder. Thus if an exception happens in prepareIndexing, we shouldn't arrive to the index() method (which is the one triggering the close()). And looking at the use of iwriter, no other line set it (so it can't be set back to null after prepareIndexing()).\n\nSo I mean we can add a {{if (iwriter != null)}} before calling the close, but the truth is I have no clue how it could ever be null at that point.\n\nHéctor: are you positive that you are using stock 0.8.1 ?","from":"developer"},{"body":"I have a patch from CASSANDRA-2818 (2818-v4) applied, if that's of any help. The patch only touches messaging classes though.","from":"developer"},{"body":"Doesn't make any sense to me, either. The only place close() is called is from index() [as seen in the stacktrace here] and the only place index() is called is after prepareIndexing, which sets iwriter to non-null:\n\n{code}\n long estimatedRows = indexer.prepareIndexing();\n\n // build the index and filter\n long rows = indexer.index();\n{code}\n","from":"developer"},{"body":"I just got a similar-looking NPE stack trace. The node that raised the exception was receiving streams from a node being decommissioned (with \"nodetool decommission\"). Both nodes were running 0.8.6; both had been upgraded a few hours earlier from 0.8.4.\n\nThe first: \n\nERROR [CompactionExecutor:72] 2011-09-20 22:34:20,892 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[CompactionExecutor:72,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n\nThen half a minute later:\n\n INFO [CompactionExecutor:72] 2011-09-20 22:34:46,923 SSTableReader.java (line 162) Opening /mnt/cassandra/data/Keyspace/CF1-g-1536\nERROR [Thread-785] 2011-09-20 22:34:52,054 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[Thread-785,5,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:154)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:63)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:189)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:117)\nCaused by: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:138)\n ... 3 more\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n","from":"developer"},{"body":"So I restarted the node making the NPE above and ran the decommission again (on the other node) and it successfully finished decommissioning. A few hours later I tried to decommission another node and received this error, on a separate node from the two mentioned before, which is a new node running 0.8.6 for just about a day now.\n\nINFO [COMMIT-LOG-WRITER] 2011-09-21 05:12:46,361 CommitLogSegment.java (line 58) Creating new commitlog segment /raid0/cassandra/commitlog/CommitLog-1316581966361.log\nERROR [Thread-224] 2011-09-21 05:12:52,738 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[Thread-224,5,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:154)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:63)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:189)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:117)\nCaused by: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:138)\n ... 3 more\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n","from":"developer"},{"body":"Since my last two comments, I increased my RF and started a rolling repair on my nodes. This has caused this NPE to pop up on all the boxes over the last couple of days as they process SSTables. Again, all the nodes are fresh 0.8.6 installs from a few days ago using the ComboAMI. I've seen backtraces like the one below appear at least a couple of times on each node in my cluster as I was repairing..\n\n INFO [CompactionExecutor:648] 2011-09-22 04:35:51,086 SSTableReader.java (line 162) Opening /raid0/cassandra/data/Keyspace/CF1-g-535\n INFO [CompactionExecutor:648] 2011-09-22 04:35:51,172 SSTableReader.java (line 162) Opening /raid0/cassandra/data/Keyspace/CF2-g-350\n INFO [CompactionExecutor:648] 2011-09-22 04:36:01,721 SSTableReader.java (line 162) Opening /raid0/cassandra/data/Keyspace/CF3-g-456\nERROR [Thread-3658] 2011-09-22 04:36:04,821 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[Thread-3658,5,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:154)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:63)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:189)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:117)\nCaused by: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:138)\n ... 3 more\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nERROR [CompactionExecutor:648] 2011-09-22 04:36:04,823 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[CompactionExecutor:648,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n","from":"developer"},{"body":"Here's another from a node that was streaming out:\n\n\n INFO [AntiEntropyStage:1] 2011-09-22 17:56:33,737 StreamOut.java (line 181) Stream context metadata [/raid0/cassandra/data/Keyspace/CF1-g-315-Data.db sections=1073 progress=0/358\n INFO [AntiEntropyStage:1] 2011-09-22 17:56:33,738 StreamOutSession.java (line 174) Streaming to /10.207.38.196\n INFO [ValidationExecutor:14] 2011-09-22 17:56:33,740 ColumnFamilyStore.java (line 1128) Enqueuing flush of Memtable-CF4@516894197(9505947/21294113 serialized/live bytes, 5684 ops)\n INFO [FlushWriter:621] 2011-09-22 17:56:33,743 Memtable.java (line 237) Writing Memtable-CF4@516894197(9505947/21294113 serialized/live bytes, 5684 ops)\n INFO [FlushWriter:621] 2011-09-22 17:56:33,880 Memtable.java (line 254) Completed flushing /raid0/cassandra/data/Keyspace/CF4-g-434-Data.db (9914202 bytes)\n INFO [CompactionExecutor:884] 2011-09-22 17:56:45,912 CompactionManager.java (line 608) Compacted to /raid0/cassandra/data/Keyspace/CF4-tmp-g-332-Data.db. 32,824,572,147 to 18,506,0\nERROR [CompactionExecutor:884] 2011-09-22 17:56:46,134 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[CompactionExecutor:884,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n","from":"developer"},{"body":"Getting the following after trying to 'move' on 0.8.5:\n\n{code}\nERROR [CompactionExecutor:213] 2011-09-23 14:02:26,571\nAbstractCassandraDaemon.java (line 139) Fatal exception in thread\nThread[CompactionExecutor:213,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n{code}","from":"developer"},{"body":"Also got the same NPE when doing a repair on a different node. This time, the NPE was preceded by a different NPE:\n\n\n{code}\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:154)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:63)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:177)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:114)\nCaused by: java.util.concurrent.ExecutionException: java.lang.NullPointerException\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:138)\n ... 3 more\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n at org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n at org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1108)\n at org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1099)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n\n{code}\n","from":"developer"},{"body":"We've seen the same error via autobootstrap.","from":"developer"},{"body":"Joaquin, can you give steps to reproduce?","from":"developer"},{"body":"Alright, I've tried boostrapping nodes a few times and I'm still not able to reproduce, so it's likely a race of something.\n\nThat being said, I still have no clue how the iwriter field in SSTableWriter.RowIndexer can be null where the stack indicates it to be null. The assignment happens before the access, there is no other assignment of that field and the assignment/access are in the same thread.\n\nAnyway, what we can easily do is to make iwrite being final and assign it in the constructor. If that doesn't make that field being non-null, I don't know what will. Patch attached to make that change.","from":"developer"},{"body":"First patch was buggy in that it was now creating the index file *before* we assert that this file doesn't exist.\n\nv2 just moves the maybeOpenIndexer call after the asserts.","from":"developer"},{"body":"+1 v2","from":"developer"},{"body":"Committed and marking resolved for now. If that errors come up in another form, we'll see then.","from":"developer"}],"created":"2011-07-06T11:32:18.000+0000","description":"A NPE is generated during repair when closing an sstable generated via SSTable build. It doesn't happen always. The node had been scrubbed and compacted before calling repair.\n\n INFO [CompactionExecutor:2] 2011-07-06 11:11:32,640 SSTableReader.java (line 158) Opening /d2/cassandra/data/sbs/walf-g-730\nERROR [CompactionExecutor:2] 2011-07-06 11:11:34,327 AbstractCassandraDaemon.java (line 113) Fatal exception in thread Thread[CompactionExecutor:2,1,main] \njava.lang.NullPointerException\n\tat org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.close(SSTableWriter.java:382)\n\tat org.apache.cassandra.io.sstable.SSTableWriter$RowIndexer.index(SSTableWriter.java:370)\n\tat org.apache.cassandra.io.sstable.SSTableWriter$Builder.build(SSTableWriter.java:315)\n\tat org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1103)\n\tat org.apache.cassandra.db.compaction.CompactionManager$9.call(CompactionManager.java:1094)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\n","issue_id":"12512969","key":"CASSANDRA-2863","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-10-10T13:58:31.000+0000","role":"fixed_distractor","summary":"NPE when writing SSTable generated via repair"} {"case_id":"12513118","cluster":"DISTRACTOR-CASSANDRA-2867","comments":[{"body":"upgrade from 0.8 to 0.8.1 \n(update: tried daily trunc build apache-cassandra-2011-07-20_16-02-23 with the same effect)\n - just put new libararies and started cassandra (CentOS 5.6 64bit)\nWe are not using any CF indexes (only counters if that matters), \nbut it won't start, so I don't see a way to upgrade to 0.8.1:\n\n{code}\n INFO [main] 2011-07-23 19:32:23,446 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/NodeIdInfo-g-1\n INFO [main] 2011-07-23 19:32:23,516 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/Schema-g-74\n INFO [main] 2011-07-23 19:32:23,529 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/Schema-g-73\n INFO [main] 2011-07-23 19:32:23,548 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/Migrations-g-73\n INFO [main] 2011-07-23 19:32:23,561 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/Migrations-g-74\n INFO [main] 2011-07-23 19:32:23,570 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/LocationInfo-g-46\n INFO [main] 2011-07-23 19:32:23,587 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/LocationInfo-g-45\n INFO [main] 2011-07-23 19:32:23,597 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/HintsColumnFamily-g-1\n INFO [main] 2011-07-23 19:32:23,656 DatabaseDescriptor.java (line 478) Loading schema version 1e26a500-b538-11e0-0000-f03fd2ec87ff\nERROR [main] 2011-07-23 19:32:23,889 AbstractCassandraDaemon.java (line 332) Exception encountered during startup.\norg.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 4\n at org.apache.cassandra.db.marshal.LongType.getString(LongType.java:72)\n at org.apache.cassandra.config.CFMetaData.getDefaultIndexName(CFMetaData.java:971)\n at org.apache.cassandra.config.CFMetaData.inflate(CFMetaData.java:381)\n at org.apache.cassandra.config.KSMetaData.inflate(KSMetaData.java:172)\n at org.apache.cassandra.db.DefsTable.loadFromStorage(DefsTable.java:99)\n at org.apache.cassandra.config.DatabaseDescriptor.loadSchemas(DatabaseDescriptor.java:479)\n at org.apache.cassandra.service.AbstractCassandraDaemon.setup(AbstractCassandraDaemon.java:139)\n at org.apache.cassandra.service.AbstractCassandraDaemon.activate(AbstractCassandraDaemon.java:315)\n at org.apache.cassandra.thrift.CassandraDaemon.main(CassandraDaemon.java:80)\n {code}","created":"2011-07-23T20:18:48.344+0000"},{"body":"You have an index name set on at least one CF. If you don't actually have an index on it, Cassandra shouldn't have let you set the index name, but it's possible an older version was not as rigorous about that.\n\nIf you can't figure out where the index name is, you can also drop your schema and migrations system CFs and rebuild your schema from scratch.","created":"2011-07-24T20:56:09.839+0000"},{"body":"The patch trunk-2867.txt fixes the MarshalException on startup.","created":"2011-07-26T13:48:55.413+0000"},{"body":"Hmm, you have an index name on a subcolumn? that's not supposed to be legal either :)","created":"2011-07-26T14:35:08.896+0000"},{"body":"No, just create the following column family and restart the server:\n\n create column family MailStat\n with column_type = Super\n and key_validation_class = UTF8Type\n and comparator = LongType\n and subcomparator = AsciiType\n and column_metadata = [\n {column_name : subj, validation_class : UTF8Type},\n {column_name : body, validation_class : UTF8Type},\n {column_name : comment, validation_class : UTF8Type},\n {column_name : categ, validation_class : AsciiType}\n ];\n","created":"2011-07-27T10:29:48.155+0000"},{"body":"This problem is not related to migration, just 0.8.1 - 0.8.2 cannot start when there is a super column family with different comparators for supercolumns and metadata.","created":"2011-07-27T10:32:41.281+0000"},{"body":"Thanks Taras for commenting on the issue, I just ran into the same problem with a super column family with a TimeUUIDType comparator, and UTF8Type subcomparators; I can insert, get data all day long from the cluster while it is running fresh. But the moment I restart the cassandra node, I get the exception saying that TimeUUIDTypes must be exactly 16 bytes.\n\nEdit: Just verified your fix and it fixes my case as well.","created":"2011-07-28T08:06:49.277+0000"},{"body":"Thanks, I see the problem now -- it was generating an index name even if there was no index present on the column (which is always the case for supercolumns).\n\nAlternate patch attached to fix this.","created":"2011-07-31T04:15:22.606+0000"},{"body":"The attached patch has a lot of unrelated changes I believe, but considering the fix is the one line change in CFMetadata, then +1.","created":"2011-08-02T15:27:43.788+0000"},{"body":"committed, minus the extra stuff from CASSANDRA-2894 that was in the diff","created":"2011-08-02T16:05:46.600+0000"}],"conversations":[{"body":"After upgrading the binaries to 0.8.1 I get an exception when starting cassandra:\n\n{noformat}\n[root@bserv2 local]# INFO 12:51:04,512 Logging initialized\n INFO 12:51:04,523 Heap size: 8329887744/8329887744\n INFO 12:51:04,524 JNA not found. Native methods will be disabled.\n INFO 12:51:04,531 Loading settings from file:/usr/local/apache-cassandra-0.8.1/conf/cassandra.yaml\n INFO 12:51:04,621 DiskAccessMode 'auto' determined to be mmap, indexAccessMode is mmap\n INFO 12:51:04,707 Global memtable threshold is enabled at 2648MB\n INFO 12:51:04,708 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,713 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,714 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,716 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,717 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,719 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,770 reading saved cache /vm1/cassandraDB/saved_caches/system-IndexInfo-KeyCache\n INFO 12:51:04,776 Opening /vm1/cassandraDB/data/system/IndexInfo-f-9\n INFO 12:51:04,792 reading saved cache /vm1/cassandraDB/saved_caches/system-Schema-KeyCache\n INFO 12:51:04,794 Opening /vm1/cassandraDB/data/system/Schema-f-194\n INFO 12:51:04,797 Opening /vm1/cassandraDB/data/system/Schema-f-195\n INFO 12:51:04,802 Opening /vm1/cassandraDB/data/system/Schema-f-193\n INFO 12:51:04,811 Opening /vm1/cassandraDB/data/system/Migrations-f-193\n INFO 12:51:04,814 reading saved cache /vm1/cassandraDB/saved_caches/system-LocationInfo-KeyCache\n INFO 12:51:04,815 Opening /vm1/cassandraDB/data/system/LocationInfo-f-292\n INFO 12:51:04,843 Loading schema version 586e70fd-a332-11e0-828e-34b74a661156\nERROR 12:51:04,996 Exception encountered during startup.\norg.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 15\n at org.apache.cassandra.db.marshal.LongType.getString(LongType.java:72)\n at org.apache.cassandra.config.CFMetaData.getDefaultIndexName(CFMetaData.java:971)\n at org.apache.cassandra.config.CFMetaData.inflate(CFMetaData.java:381)\n at org.apache.cassandra.config.KSMetaData.inflate(KSMetaData.java:172)\n at org.apache.cassandra.db.DefsTable.loadFromStorage(DefsTable.java:99)\n at org.apache.cassandra.config.DatabaseDescriptor.loadSchemas(DatabaseDescriptor.java:479)\n at org.apache.cassandra.service.AbstractCassandraDaemon.setup(AbstractCassandraDaemon.java:139)\n at org.apache.cassandra.service.AbstractCassandraDaemon.activate(AbstractCassandraDaemon.java:315)\n at org.apache.cassandra.thrift.CassandraDaemon.main(CassandraDaemon.java:80)\nException encountered during startup.\norg.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 15\n at org.apache.cassandra.db.marshal.LongType.getString(LongType.java:72)\n at org.apache.cassandra.config.CFMetaData.getDefaultIndexName(CFMetaData.java:971)\n at org.apache.cassandra.config.CFMetaData.inflate(CFMetaData.java:381)\n at org.apache.cassandra.config.KSMetaData.inflate(KSMetaData.java:172)\n at org.apache.cassandra.db.DefsTable.loadFromStorage(DefsTable.java:99)\n at org.apache.cassandra.config.DatabaseDescriptor.loadSchemas(DatabaseDescriptor.java:479)\n at org.apache.cassandra.service.AbstractCassandraDaemon.setup(AbstractCassandraDaemon.java:139)\n at org.apache.cassandra.service.AbstractCassandraDaemon.activate(AbstractCassandraDaemon.java:315)\n at org.apache.cassandra.thrift.CassandraDaemon.main(CassandraDaemon.java:80)\n{noformat}\n\nIt seems this has something to do with indexes, and I do have a CF with an index on it, but it is not used.\nI can try and remove the index with 0.7.x binaries, but I will wait a bit to see if anyone needs it to reproduce the bug.","from":"reporter","subject":"Starting 0.8.1 after upgrade from 0.7.6-2 fails"},{"body":"upgrade from 0.8 to 0.8.1 \n(update: tried daily trunc build apache-cassandra-2011-07-20_16-02-23 with the same effect)\n - just put new libararies and started cassandra (CentOS 5.6 64bit)\nWe are not using any CF indexes (only counters if that matters), \nbut it won't start, so I don't see a way to upgrade to 0.8.1:\n\n{code}\n INFO [main] 2011-07-23 19:32:23,446 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/NodeIdInfo-g-1\n INFO [main] 2011-07-23 19:32:23,516 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/Schema-g-74\n INFO [main] 2011-07-23 19:32:23,529 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/Schema-g-73\n INFO [main] 2011-07-23 19:32:23,548 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/Migrations-g-73\n INFO [main] 2011-07-23 19:32:23,561 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/Migrations-g-74\n INFO [main] 2011-07-23 19:32:23,570 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/LocationInfo-g-46\n INFO [main] 2011-07-23 19:32:23,587 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/LocationInfo-g-45\n INFO [main] 2011-07-23 19:32:23,597 SSTableReader.java (line 158) Opening /var/lib/cassandra/data/system/HintsColumnFamily-g-1\n INFO [main] 2011-07-23 19:32:23,656 DatabaseDescriptor.java (line 478) Loading schema version 1e26a500-b538-11e0-0000-f03fd2ec87ff\nERROR [main] 2011-07-23 19:32:23,889 AbstractCassandraDaemon.java (line 332) Exception encountered during startup.\norg.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 4\n at org.apache.cassandra.db.marshal.LongType.getString(LongType.java:72)\n at org.apache.cassandra.config.CFMetaData.getDefaultIndexName(CFMetaData.java:971)\n at org.apache.cassandra.config.CFMetaData.inflate(CFMetaData.java:381)\n at org.apache.cassandra.config.KSMetaData.inflate(KSMetaData.java:172)\n at org.apache.cassandra.db.DefsTable.loadFromStorage(DefsTable.java:99)\n at org.apache.cassandra.config.DatabaseDescriptor.loadSchemas(DatabaseDescriptor.java:479)\n at org.apache.cassandra.service.AbstractCassandraDaemon.setup(AbstractCassandraDaemon.java:139)\n at org.apache.cassandra.service.AbstractCassandraDaemon.activate(AbstractCassandraDaemon.java:315)\n at org.apache.cassandra.thrift.CassandraDaemon.main(CassandraDaemon.java:80)\n {code}","from":"developer"},{"body":"You have an index name set on at least one CF. If you don't actually have an index on it, Cassandra shouldn't have let you set the index name, but it's possible an older version was not as rigorous about that.\n\nIf you can't figure out where the index name is, you can also drop your schema and migrations system CFs and rebuild your schema from scratch.","from":"developer"},{"body":"The patch trunk-2867.txt fixes the MarshalException on startup.","from":"developer"},{"body":"Hmm, you have an index name on a subcolumn? that's not supposed to be legal either :)","from":"developer"},{"body":"No, just create the following column family and restart the server:\n\n create column family MailStat\n with column_type = Super\n and key_validation_class = UTF8Type\n and comparator = LongType\n and subcomparator = AsciiType\n and column_metadata = [\n {column_name : subj, validation_class : UTF8Type},\n {column_name : body, validation_class : UTF8Type},\n {column_name : comment, validation_class : UTF8Type},\n {column_name : categ, validation_class : AsciiType}\n ];\n","from":"developer"},{"body":"This problem is not related to migration, just 0.8.1 - 0.8.2 cannot start when there is a super column family with different comparators for supercolumns and metadata.","from":"developer"},{"body":"Thanks Taras for commenting on the issue, I just ran into the same problem with a super column family with a TimeUUIDType comparator, and UTF8Type subcomparators; I can insert, get data all day long from the cluster while it is running fresh. But the moment I restart the cassandra node, I get the exception saying that TimeUUIDTypes must be exactly 16 bytes.\n\nEdit: Just verified your fix and it fixes my case as well.","from":"developer"},{"body":"Thanks, I see the problem now -- it was generating an index name even if there was no index present on the column (which is always the case for supercolumns).\n\nAlternate patch attached to fix this.","from":"developer"},{"body":"The attached patch has a lot of unrelated changes I believe, but considering the fix is the one line change in CFMetadata, then +1.","from":"developer"},{"body":"committed, minus the extra stuff from CASSANDRA-2894 that was in the diff","from":"developer"}],"created":"2011-07-07T10:20:03.000+0000","description":"After upgrading the binaries to 0.8.1 I get an exception when starting cassandra:\n\n{noformat}\n[root@bserv2 local]# INFO 12:51:04,512 Logging initialized\n INFO 12:51:04,523 Heap size: 8329887744/8329887744\n INFO 12:51:04,524 JNA not found. Native methods will be disabled.\n INFO 12:51:04,531 Loading settings from file:/usr/local/apache-cassandra-0.8.1/conf/cassandra.yaml\n INFO 12:51:04,621 DiskAccessMode 'auto' determined to be mmap, indexAccessMode is mmap\n INFO 12:51:04,707 Global memtable threshold is enabled at 2648MB\n INFO 12:51:04,708 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,713 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,714 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,716 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,717 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,719 Removing compacted SSTable files (see http://wiki.apache.org/cassandra/MemtableSSTable)\n INFO 12:51:04,770 reading saved cache /vm1/cassandraDB/saved_caches/system-IndexInfo-KeyCache\n INFO 12:51:04,776 Opening /vm1/cassandraDB/data/system/IndexInfo-f-9\n INFO 12:51:04,792 reading saved cache /vm1/cassandraDB/saved_caches/system-Schema-KeyCache\n INFO 12:51:04,794 Opening /vm1/cassandraDB/data/system/Schema-f-194\n INFO 12:51:04,797 Opening /vm1/cassandraDB/data/system/Schema-f-195\n INFO 12:51:04,802 Opening /vm1/cassandraDB/data/system/Schema-f-193\n INFO 12:51:04,811 Opening /vm1/cassandraDB/data/system/Migrations-f-193\n INFO 12:51:04,814 reading saved cache /vm1/cassandraDB/saved_caches/system-LocationInfo-KeyCache\n INFO 12:51:04,815 Opening /vm1/cassandraDB/data/system/LocationInfo-f-292\n INFO 12:51:04,843 Loading schema version 586e70fd-a332-11e0-828e-34b74a661156\nERROR 12:51:04,996 Exception encountered during startup.\norg.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 15\n at org.apache.cassandra.db.marshal.LongType.getString(LongType.java:72)\n at org.apache.cassandra.config.CFMetaData.getDefaultIndexName(CFMetaData.java:971)\n at org.apache.cassandra.config.CFMetaData.inflate(CFMetaData.java:381)\n at org.apache.cassandra.config.KSMetaData.inflate(KSMetaData.java:172)\n at org.apache.cassandra.db.DefsTable.loadFromStorage(DefsTable.java:99)\n at org.apache.cassandra.config.DatabaseDescriptor.loadSchemas(DatabaseDescriptor.java:479)\n at org.apache.cassandra.service.AbstractCassandraDaemon.setup(AbstractCassandraDaemon.java:139)\n at org.apache.cassandra.service.AbstractCassandraDaemon.activate(AbstractCassandraDaemon.java:315)\n at org.apache.cassandra.thrift.CassandraDaemon.main(CassandraDaemon.java:80)\nException encountered during startup.\norg.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 15\n at org.apache.cassandra.db.marshal.LongType.getString(LongType.java:72)\n at org.apache.cassandra.config.CFMetaData.getDefaultIndexName(CFMetaData.java:971)\n at org.apache.cassandra.config.CFMetaData.inflate(CFMetaData.java:381)\n at org.apache.cassandra.config.KSMetaData.inflate(KSMetaData.java:172)\n at org.apache.cassandra.db.DefsTable.loadFromStorage(DefsTable.java:99)\n at org.apache.cassandra.config.DatabaseDescriptor.loadSchemas(DatabaseDescriptor.java:479)\n at org.apache.cassandra.service.AbstractCassandraDaemon.setup(AbstractCassandraDaemon.java:139)\n at org.apache.cassandra.service.AbstractCassandraDaemon.activate(AbstractCassandraDaemon.java:315)\n at org.apache.cassandra.thrift.CassandraDaemon.main(CassandraDaemon.java:80)\n{noformat}\n\nIt seems this has something to do with indexes, and I do have a CF with an index on it, but it is not used.\nI can try and remove the index with 0.7.x binaries, but I will wait a bit to see if anyone needs it to reproduce the bug.","issue_id":"12513118","key":"CASSANDRA-2867","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-08-02T16:05:46.000+0000","role":"fixed_distractor","summary":"Starting 0.8.1 after upgrade from 0.7.6-2 fails"} {"case_id":"12513152","cluster":"DISTRACTOR-CASSANDRA-2868","comments":[{"body":"Hm after 3 days checking a node that does not use mmaped files it looks like this:\n\nnativelib: 14128\nlocale-archive: 1492\nffiSwFShY(deleted): 8\njavajar: 2292\n[anon]: 3609388\n[stack]: 132\njava: 44\n7008: 32\njna534482390478104336.tmp: 92\n\nTotal RSS: 3627608\nTotal SST: 0\n\n\nCompared to start RSS increased by ~400MB. So it seems that this is not related to mem mapping.\n\nWe will deploy CASSANDRA-2654 this week. Will see if that changes anything but I suspect not ...","created":"2011-07-11T10:45:03.963+0000"},{"body":"At one point I was convinced this was a JVM bug and opened http://bugs.sun.com/bugdatabase/view_bug.do?bug_id=7037080 After seeing how totally broken NIO is after CASSANDRA-2654 I'm no longer sure of anything.\n\nI was going to start a survey on the user list after the summit to see if any OS/jvm level pattern could be found, since clearly it doesn't happen to everyone in all cases.","created":"2011-07-11T18:31:36.862+0000"},{"body":"Is your data size constant? If not you are probably seeing growth in the index samples and bloom filters.","created":"2011-07-11T20:33:24.830+0000"},{"body":"Next: [anon]: 3675224 (+47616KB in 1 day)\n\nbq. Is your data size constant? If not you are probably seeing growth in the index samples and bloom filters.\n\nWell no - the data size is increasing. But I thought that index and bf is good old plain java heap no? JVM heap stats are really relaxed. Yet I think that doesn't really matter because what we are seeing is an ever increasing rss mem consumption even though we have -Xmx3G and -Xms3G and mlockall (pmap shows these 3G as one block). So something seems to be constantly allocating native mem that has nothing to do with java heap.","created":"2011-07-12T08:30:05.200+0000"},{"body":"We call getLastGcInfo several times a second. http://twitter.com/#!/kimchy/status/90861039930970113\n\nYou could try turning GCInspector methods into a no-op and see if that makes it go away.","created":"2011-07-12T20:43:40.511+0000"},{"body":"http://bugs.sun.com/bugdatabase/view_bug.do?bug_id=7066129 will be the id when bugs.sun.com gets around to doing it's thing.\n\nI confirmed that -XX:MaxDirectMemorySize does not protect you from this (ie it's a native native leak, not some DirectByteBuffer thing). I'll be able to test this but not until the end of this week at the earliest (and it will then take at least another week to be sure).","created":"2011-07-13T00:11:52.073+0000"},{"body":"In case it is useful to anyone else this is what I intend to test with.","created":"2011-07-13T01:13:47.146+0000"},{"body":"Initial results. Graph of VmRSS from /proc/PID/status at 10 second intervals from my last comment to now. Box on the left has GCInspector disabled. These are on two test boxes under trivial load so this is all still *very* tentative. Will start testing under real load by early next week.","created":"2011-07-14T13:16:14.925+0000"},{"body":"Promising!","created":"2011-07-14T13:21:53.204+0000"},{"body":"It's indeed promising. We have been running this in production for 3 days now and rss increased only insignificantly by ~5MB a day. ","created":"2011-07-15T11:31:24.840+0000"},{"body":"{quote}\nWe have been running this in production for 3 days now and rss increased only insignificantly by ~5MB a day\n{quote}\n\nDo you mean it is very helpful to control RSS increasing by removing getLastGcInfo()? \n\nI have no idea why just some of us meets the problem.","created":"2011-07-15T15:33:04.188+0000"},{"body":"I interpreted Daniel's \"this\" to be the 2868-v1.txt patch (or something equivalent) with cassandra.enable_gc_inspector=false. I did not find -XX:MaxDirectMemorySize to be helpful.","created":"2011-07-15T16:24:32.890+0000"},{"body":"Yes - we did disable the GCInspector.","created":"2011-07-16T06:45:36.316+0000"},{"body":"Got it!\n\nDo you have any idea why only some of us reports the problem?","created":"2011-07-16T07:04:41.884+0000"},{"body":"Well either it's environment specific or (more likely) others didn't notice / care because they have enough memory and/or restart the nodes often enough.\n\nWe have 16GB of RAM and run Cassandra with 3GB. Within one month we loose ~3GB (13GB -> 10GB) files system cache because of the mem leak. Looking at our graphs I can't really tell a difference performance wise. So I guess only people with weaker servers (less memory headroom) will really notice. We noticed only because we got the system oom on a cluster that's not critical and which we didn't really monitor.","created":"2011-07-16T10:29:26.394+0000"},{"body":"Looks good to me. Guess cassandra should just disable the inspector for now (probably make it jmx'able to start it manually)\n\nThu Jul 14 09:39:26 CEST 2011: [anon]: 3234068\nThu Jul 14 17:22:45 CEST 2011: [anon]: 3266888\nFri Jul 15 09:33:53 CEST 2011: [anon]: 3269160\nMon Jul 18 09:54:29 CEST 2011: [anon]: 3270188","created":"2011-07-18T07:59:06.665+0000"},{"body":"We could try switching to what JConsole does, which is just log the total number and time spent for each compaction type. This uses a different API which hopefully does not leak: http://www.java2s.com/Open-Source/Java-Document/6.0-JDK-Modules-sun/tools/sun/tools/jconsole/MemoryTab.java.htm\n\nLogging the lifetime totals there with StatusLogger similar to what we do now for dropped messages would be better than nothing.","created":"2011-07-18T20:36:03.641+0000"},{"body":"48 hours under production load after C* had already been running for a few days. Two on the left have GCInspector enabled. The two on the right do not. (Note that the scale on the lower right one reflects a change of only 10s of bytes.)\n\nSo it looks like victory to me.","created":"2011-07-27T17:49:03.847+0000"},{"body":"Thanks, Chris. We'll work on rewriting GCInspector to use the java.lang.management api instead, unless you have time to take a stab at that.","created":"2011-07-27T18:09:46.534+0000"},{"body":"Depending how long the rewrite is going to take, can we get the config file option to disable gc inspector into a new 0.7.X and 0.8.X release?","created":"2011-07-29T20:15:52.688+0000"},{"body":"v2 switches the GCInspector to us java.lang.managment. I don't know if it too leaks or not yet.","created":"2011-07-29T21:10:13.991+0000"},{"body":"I created three isolated nodes, all with a hack of setting the inspector interval to 1ms applied (not the tightest loop, but good enough and easy.) One of the nodes had the inspector disabled entirely (the control), one was vanilla, and one had v2 applied. After starting them up with a 128M heap and letting them run for a few minutes, here are the results:\n\n||version||resident||\n|control|72M|\n|patched|72M|\n|vanilla|540M|\n\nI think it's safe to say java.lang.management doesn't share the leak.","created":"2011-08-01T00:56:22.323+0000"},{"body":"Comments on v2:\n* Couldn't we estimate the reclaimed size by recording the last memory used (that would need to be the first thing we do in logGCResults so that we record it each time) ?\n* Wouldn't it be worth indicating that how many collection have been done since last log message if it's > 1, since it can (be > 1).\n* Nit: especially if we decide to keep the last memory used, it may be more efficient (in cleaner imho) to have just one HashMap of string -> GCInfo where GCInfo would be a small struct with times, counts and usedMemory. Not that it is very performance sensitive... ","created":"2011-08-01T18:05:32.654+0000"},{"body":"bq. Couldn't we estimate the reclaimed size\n\nWell, not really, what we'd have is \"difference in size between last time it was called, and now\" which isn't all that close to \"amount reclaimed by a specific GC.\"\n\nbq. Wouldn't it be worth indicating that how many collection have been done since last log message\n\nIMO the duration-based thresholds are hard to reason about here, where we're dealing w/ summaries and not individual GC results. I think I'd rather have something like the dropped messages logger, where every N seconds we log the summary we get from the mbean.\n\nThe flushLargestMemtables/reduceCacheSizes stuff should probably be removed. :(","created":"2011-08-01T18:55:00.195+0000"},{"body":"bq. Wouldn't it be worth indicating that how many collection have been done since last log message if it's > 1, since it can (be > 1).\n\nThe only reason I added count tracking was to prevent it from firing when there were no GCs (the api is flakey.) I've never actually been able to get > 1 to happen, but we can add it to the logging.\n\nbq. IMO the duration-based thresholds are hard to reason about here, where we're dealing w/ summaries and not individual GC results.\n\nWe are dealing with individual GCs at least 99% of the time in practice. The worst case is >1 GC inflates the gctime enough that we errantly log when it's not needed, but I imagine to trigger that you would have to be in a gc pressure situation already.\n\nbq. I think I'd rather have something like the dropped messages logger, where every N seconds we log the summary we get from the mbean.\n\nThat seems like it could be a lot of noise since GC is constantly happening.\n\nbq. The flushLargestMemtables/reduceCacheSizes stuff should probably be removed. \n\nI think the logic there is still sound (\"Did we just do a CMS? Is the heap still 80% full?\") and it seems to work as well as it always has.\n\n","created":"2011-08-09T18:43:06.689+0000"},{"body":"bq. I've never actually been able to get > 1 to happen, but we can add it to the logging\n\nI'm sure it's possible w/ a small enough heap, especially since GCInspector is paused along w/ everything else for STW collections (including new gen).\n\nv3 attached to accomodate this and add durationPerCollection.","created":"2011-08-16T16:59:29.698+0000"},{"body":"Why is v3 touching compaction?","created":"2011-08-16T17:05:24.485+0000"},{"body":"dirty working directory. GCI is the only relevant file.","created":"2011-08-16T17:09:49.380+0000"},{"body":"+1 to GCI changes. Also, it is indeed possible to get >1 with a tiny heap.","created":"2011-08-16T23:36:31.839+0000"},{"body":"committed","created":"2011-08-17T02:20:53.855+0000"},{"body":"Can we get this in 0.7.X as well?","created":"2011-08-17T03:28:44.294+0000"},{"body":"Reopening to backport to 0.7","created":"2011-08-23T19:18:36.121+0000"},{"body":"Committed to 0.7 in r1160879","created":"2011-08-23T20:06:00.263+0000"}],"conversations":[{"body":"We have memory issues with long running servers. These have been confirmed by several users in the user list. That's why I report.\n\nThe memory consumption of the cassandra java process increases steadily until it's killed by the os because of oom (with no swap)\n\nOur server is started with -Xmx3000M and running for around 23 days.\n\npmap -x shows\n\nTotal SST: 1961616 (mem mapped data and index files)\nAnon RSS: 6499640\nTotal RSS: 8478376\n\nThis shows that > 3G are 'overallocated'.\n\nWe will use BRAF on one of our less important nodes to check wether it is related to mmap and report back.","from":"reporter","subject":"Native Memory Leak"},{"body":"Hm after 3 days checking a node that does not use mmaped files it looks like this:\n\nnativelib: 14128\nlocale-archive: 1492\nffiSwFShY(deleted): 8\njavajar: 2292\n[anon]: 3609388\n[stack]: 132\njava: 44\n7008: 32\njna534482390478104336.tmp: 92\n\nTotal RSS: 3627608\nTotal SST: 0\n\n\nCompared to start RSS increased by ~400MB. So it seems that this is not related to mem mapping.\n\nWe will deploy CASSANDRA-2654 this week. Will see if that changes anything but I suspect not ...","from":"developer"},{"body":"At one point I was convinced this was a JVM bug and opened http://bugs.sun.com/bugdatabase/view_bug.do?bug_id=7037080 After seeing how totally broken NIO is after CASSANDRA-2654 I'm no longer sure of anything.\n\nI was going to start a survey on the user list after the summit to see if any OS/jvm level pattern could be found, since clearly it doesn't happen to everyone in all cases.","from":"developer"},{"body":"Is your data size constant? If not you are probably seeing growth in the index samples and bloom filters.","from":"developer"},{"body":"Next: [anon]: 3675224 (+47616KB in 1 day)\n\nbq. Is your data size constant? If not you are probably seeing growth in the index samples and bloom filters.\n\nWell no - the data size is increasing. But I thought that index and bf is good old plain java heap no? JVM heap stats are really relaxed. Yet I think that doesn't really matter because what we are seeing is an ever increasing rss mem consumption even though we have -Xmx3G and -Xms3G and mlockall (pmap shows these 3G as one block). So something seems to be constantly allocating native mem that has nothing to do with java heap.","from":"developer"},{"body":"We call getLastGcInfo several times a second. http://twitter.com/#!/kimchy/status/90861039930970113\n\nYou could try turning GCInspector methods into a no-op and see if that makes it go away.","from":"developer"},{"body":"http://bugs.sun.com/bugdatabase/view_bug.do?bug_id=7066129 will be the id when bugs.sun.com gets around to doing it's thing.\n\nI confirmed that -XX:MaxDirectMemorySize does not protect you from this (ie it's a native native leak, not some DirectByteBuffer thing). I'll be able to test this but not until the end of this week at the earliest (and it will then take at least another week to be sure).","from":"developer"},{"body":"In case it is useful to anyone else this is what I intend to test with.","from":"developer"},{"body":"Initial results. Graph of VmRSS from /proc/PID/status at 10 second intervals from my last comment to now. Box on the left has GCInspector disabled. These are on two test boxes under trivial load so this is all still *very* tentative. Will start testing under real load by early next week.","from":"developer"},{"body":"Promising!","from":"developer"},{"body":"It's indeed promising. We have been running this in production for 3 days now and rss increased only insignificantly by ~5MB a day. ","from":"developer"},{"body":"{quote}\nWe have been running this in production for 3 days now and rss increased only insignificantly by ~5MB a day\n{quote}\n\nDo you mean it is very helpful to control RSS increasing by removing getLastGcInfo()? \n\nI have no idea why just some of us meets the problem.","from":"developer"},{"body":"I interpreted Daniel's \"this\" to be the 2868-v1.txt patch (or something equivalent) with cassandra.enable_gc_inspector=false. I did not find -XX:MaxDirectMemorySize to be helpful.","from":"developer"},{"body":"Yes - we did disable the GCInspector.","from":"developer"},{"body":"Got it!\n\nDo you have any idea why only some of us reports the problem?","from":"developer"},{"body":"Well either it's environment specific or (more likely) others didn't notice / care because they have enough memory and/or restart the nodes often enough.\n\nWe have 16GB of RAM and run Cassandra with 3GB. Within one month we loose ~3GB (13GB -> 10GB) files system cache because of the mem leak. Looking at our graphs I can't really tell a difference performance wise. So I guess only people with weaker servers (less memory headroom) will really notice. We noticed only because we got the system oom on a cluster that's not critical and which we didn't really monitor.","from":"developer"},{"body":"Looks good to me. Guess cassandra should just disable the inspector for now (probably make it jmx'able to start it manually)\n\nThu Jul 14 09:39:26 CEST 2011: [anon]: 3234068\nThu Jul 14 17:22:45 CEST 2011: [anon]: 3266888\nFri Jul 15 09:33:53 CEST 2011: [anon]: 3269160\nMon Jul 18 09:54:29 CEST 2011: [anon]: 3270188","from":"developer"},{"body":"We could try switching to what JConsole does, which is just log the total number and time spent for each compaction type. This uses a different API which hopefully does not leak: http://www.java2s.com/Open-Source/Java-Document/6.0-JDK-Modules-sun/tools/sun/tools/jconsole/MemoryTab.java.htm\n\nLogging the lifetime totals there with StatusLogger similar to what we do now for dropped messages would be better than nothing.","from":"developer"},{"body":"48 hours under production load after C* had already been running for a few days. Two on the left have GCInspector enabled. The two on the right do not. (Note that the scale on the lower right one reflects a change of only 10s of bytes.)\n\nSo it looks like victory to me.","from":"developer"},{"body":"Thanks, Chris. We'll work on rewriting GCInspector to use the java.lang.management api instead, unless you have time to take a stab at that.","from":"developer"},{"body":"Depending how long the rewrite is going to take, can we get the config file option to disable gc inspector into a new 0.7.X and 0.8.X release?","from":"developer"},{"body":"v2 switches the GCInspector to us java.lang.managment. I don't know if it too leaks or not yet.","from":"developer"},{"body":"I created three isolated nodes, all with a hack of setting the inspector interval to 1ms applied (not the tightest loop, but good enough and easy.) One of the nodes had the inspector disabled entirely (the control), one was vanilla, and one had v2 applied. After starting them up with a 128M heap and letting them run for a few minutes, here are the results:\n\n||version||resident||\n|control|72M|\n|patched|72M|\n|vanilla|540M|\n\nI think it's safe to say java.lang.management doesn't share the leak.","from":"developer"},{"body":"Comments on v2:\n* Couldn't we estimate the reclaimed size by recording the last memory used (that would need to be the first thing we do in logGCResults so that we record it each time) ?\n* Wouldn't it be worth indicating that how many collection have been done since last log message if it's > 1, since it can (be > 1).\n* Nit: especially if we decide to keep the last memory used, it may be more efficient (in cleaner imho) to have just one HashMap of string -> GCInfo where GCInfo would be a small struct with times, counts and usedMemory. Not that it is very performance sensitive... ","from":"developer"},{"body":"bq. Couldn't we estimate the reclaimed size\n\nWell, not really, what we'd have is \"difference in size between last time it was called, and now\" which isn't all that close to \"amount reclaimed by a specific GC.\"\n\nbq. Wouldn't it be worth indicating that how many collection have been done since last log message\n\nIMO the duration-based thresholds are hard to reason about here, where we're dealing w/ summaries and not individual GC results. I think I'd rather have something like the dropped messages logger, where every N seconds we log the summary we get from the mbean.\n\nThe flushLargestMemtables/reduceCacheSizes stuff should probably be removed. :(","from":"developer"},{"body":"bq. Wouldn't it be worth indicating that how many collection have been done since last log message if it's > 1, since it can (be > 1).\n\nThe only reason I added count tracking was to prevent it from firing when there were no GCs (the api is flakey.) I've never actually been able to get > 1 to happen, but we can add it to the logging.\n\nbq. IMO the duration-based thresholds are hard to reason about here, where we're dealing w/ summaries and not individual GC results.\n\nWe are dealing with individual GCs at least 99% of the time in practice. The worst case is >1 GC inflates the gctime enough that we errantly log when it's not needed, but I imagine to trigger that you would have to be in a gc pressure situation already.\n\nbq. I think I'd rather have something like the dropped messages logger, where every N seconds we log the summary we get from the mbean.\n\nThat seems like it could be a lot of noise since GC is constantly happening.\n\nbq. The flushLargestMemtables/reduceCacheSizes stuff should probably be removed. \n\nI think the logic there is still sound (\"Did we just do a CMS? Is the heap still 80% full?\") and it seems to work as well as it always has.\n\n","from":"developer"},{"body":"bq. I've never actually been able to get > 1 to happen, but we can add it to the logging\n\nI'm sure it's possible w/ a small enough heap, especially since GCInspector is paused along w/ everything else for STW collections (including new gen).\n\nv3 attached to accomodate this and add durationPerCollection.","from":"developer"},{"body":"Why is v3 touching compaction?","from":"developer"},{"body":"dirty working directory. GCI is the only relevant file.","from":"developer"},{"body":"+1 to GCI changes. Also, it is indeed possible to get >1 with a tiny heap.","from":"developer"},{"body":"committed","from":"developer"},{"body":"Can we get this in 0.7.X as well?","from":"developer"},{"body":"Reopening to backport to 0.7","from":"developer"},{"body":"Committed to 0.7 in r1160879","from":"developer"}],"created":"2011-07-07T14:05:41.000+0000","description":"We have memory issues with long running servers. These have been confirmed by several users in the user list. That's why I report.\n\nThe memory consumption of the cassandra java process increases steadily until it's killed by the os because of oom (with no swap)\n\nOur server is started with -Xmx3000M and running for around 23 days.\n\npmap -x shows\n\nTotal SST: 1961616 (mem mapped data and index files)\nAnon RSS: 6499640\nTotal RSS: 8478376\n\nThis shows that > 3G are 'overallocated'.\n\nWe will use BRAF on one of our less important nodes to check wether it is related to mmap and report back.","issue_id":"12513152","key":"CASSANDRA-2868","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-08-23T20:06:00.000+0000","role":"fixed_distractor","summary":"Native Memory Leak"} {"case_id":"12515455","cluster":"DISTRACTOR-CASSANDRA-2950","comments":[{"body":"This is a general issue with all CF's. updating bug.","created":"2011-07-27T01:37:15.165+0000"},{"body":"The other permutation of this bug looked like, assuming write with CL.Q:\n* Insert 50 (3 nodes up)\n* truncate CF (3 nodes up)\n* Insert 1 (3 nodes up)\n* Bring node3 down\n* Delete 1 (2 nodes up)\n* Bring up node3 and run repair\n* Take down node1 and node2.\n* Query node3 with CL.ONE: list Standard1; --- 30 rows returned\n\nNot sure, but this looked suspicious in my logs:\n{code}\n INFO 01:19:45,616 Streaming to /50.57.114.45\n INFO 01:19:45,689 Finished streaming session 698609583499991 from /50.57.107.176\n INFO 01:19:45,690 Finished streaming session 698609609994154 from /50.57.114.45\n INFO 01:19:46,501 Finished streaming repair with /50.57.114.45 for (0,56713727820156410577229101238628035242]: 0 oustanding to complete session\n INFO 01:19:46,531 Compacted to /var/lib/cassandra/data/Keyspace1/Standard1-tmp-g-106-Data.db. 16,646,523 to 16,646,352 (~99% of original) bytes for 30 keys. Time: 1,509ms.\n INFO 01:19:46,930 Finished streaming repair with /50.57.107.176 for (113427455640312821154458202477256070484,0]: 1 oustanding to complete session\n INFO 01:19:47,619 Finished streaming repair with /50.57.114.45 for (113427455640312821154458202477256070484,0]: 0 oustanding to complete session\n INFO 01:19:48,232 Finished streaming repair with /50.57.107.176 for (56713727820156410577229101238628035242,113427455640312821154458202477256070484]: 1 oustanding to complete session\n INFO 01:19:48,856 Finished streaming repair with /50.57.114.45 for (56713727820156410577229101238628035242,113427455640312821154458202477256070484]: 0 oustanding to complete session\n{code}","created":"2011-07-27T01:47:16.023+0000"},{"body":"Currently, truncate does:\n* force a flush\n* record the time\n* delete any sstables older than the time\n\nThis isn't quite enough if the machine crashes shortly afterward, however, since there can be mutations present in the commitlog that were previously truncated and are now resurrected by CL replay.\n\nOne thing we could do is record the truncate time for the CF in the system ks and then ignore mutations older than that, however this would require time synchronization between the client and the server to be accurate.\n","created":"2011-08-09T21:31:48.829+0000"},{"body":"but we record CL \"context\" at time of flush in the sstable it makes, and we on replay we ignore any mutations from before that position.\n\nchecked and we do wait for flush to complete in truncate.","created":"2011-08-09T22:12:00.074+0000"},{"body":"bq. but we record CL \"context\" at time of flush in the sstable it makes, and we on replay we ignore any mutations from before that position.\n\nI think there's something wrong with that, then:\n\n{noformat}\n INFO 21:25:15,274 Replaying /var/lib/cassandra/commitlog/CommitLog-1312924388053.log\nDEBUG 21:25:15,290 Replaying /var/lib/cassandra/commitlog/CommitLog-1312924388053.log starting at 0\nDEBUG 21:25:15,291 Reading mutation at 0\nDEBUG 21:25:15,295 replaying mutation for system.4c: {ColumnFamily(LocationInfo [47656e65726174696f6e:false:4@1312924388140000,])}\nDEBUG 21:25:15,321 Reading mutation at 89\nDEBUG 21:25:15,322 replaying mutation for system.426f6f747374726170: {ColumnFamily(LocationInfo [42:false:1@1312924388203,])}\nDEBUG 21:25:15,322 Reading mutation at 174\nDEBUG 21:25:15,322 replaying mutation for system.4c: {ColumnFamily(LocationInfo [546f6b656e:false:16@1312924388204,])}\nDEBUG 21:25:15,322 Reading mutation at 270\nDEBUG 21:25:15,324 replaying mutation for Keyspace1.3030: {ColumnFamily(Standard1 [C0:false:34@1312924813259,C1:false:34@1312924813260,C2:false:34@1312924813260,C3:false:34@1312924813260,C4:false:34@1312924813260,])}\n{noformat}\n\nThe last entry there is the first of many errant mutations.","created":"2011-08-09T22:19:12.442+0000"},{"body":"Ah, CASSANDRA-2419 keeps on giving...\n\nbq. but we record CL \"context\" at time of flush in the sstable it makes, and we on replay we ignore any mutations from before that position.\n\nThe obvious problem with this is that the point of truncate is to blow away such sstables... Patch attached. Comment explains the core fix:\n\n{noformat}\n// Bonus complication: since we store replay position in sstable metadata,\n// truncating those sstables means we will replay any CL segments from the\n// beginning if we restart before they are discarded for normal reasons\n// post-truncate. So we need to (a) force a new segment so the currently\n// active one can be discarded, and (b) flush *all* CFs so that unflushed\n// data in others don't keep any pre-truncate CL segments alive.\n{noformat}\n\nPatch also fixes the bug in ReplayManagerTruncateTest that made it miss this.\n","created":"2011-08-10T14:57:14.901+0000"},{"body":"I think the forceFlush of all the CF is not safe, because if for a given column family the memtable is clean, forceFlush will return immediately, even though there could be a memtable being flush at the same time (or pending flush). So we cannot be sure all the old segment are clean after the waitFutures (I know, it took me some time to figure out some problem with repair for this very reason when the repairs were not properly synchronized).\n\nWhat we would need is to add to the future we wait on the futures of all the flush being processed at that time. Sounds annoying though. ","created":"2011-08-10T16:06:36.189+0000"},{"body":"+1, though this patch is against trunk, not 0.8. Also mistakenly bumps the log4j level to debug.","created":"2011-08-10T16:09:03.417+0000"},{"body":"v2:\n\n{noformat}\n// Bonus bonus: simply forceFlush of all the CF is not enough, because if\n// for a given column family the memtable is clean, forceFlush will return\n// immediately, even though there could be a memtable being flush at the same\n// time. So to guarantee that all segments can be cleaned out, we need\n// \"waitForActiveFlushes\" after the new segment has been created.\n{noformat}","created":"2011-08-10T19:46:08.427+0000"},{"body":"Attaching a v3 that is rebased against 0.8. I've also slightly change the logic in Truncate to submit all the flushes and then call waitForActiveFlushes, as this is slightly simpler and should work equally well as far as I can tell.\nApart from that, this lgtm.","created":"2011-08-11T12:53:17.794+0000"},{"body":"committed","created":"2011-08-11T19:36:15.101+0000"}],"conversations":[{"body":"* Configure 3 node cluster\n* Ensure the java stress tool creates Keyspace1 with RF=3\n\n{code}\n// Run Stress Tool to generate 10 keys, 1 column\nstress --operation=INSERT -t 2 --num-keys=50 --columns=20 --consistency-level=QUORUM --average-size-values --replication-factor=3 --create-index=KEYS --nodes=cathy1,cathy2\n\n// Verify 50 keys in CLI\nuse Keyspace1; \nlist Standard1; \n\n// TRUNCATE CF in CLI\nuse Keyspace1;\ntruncate counter1;\nlist counter1;\n\n// Run stress tool and verify creation of 1 key with 10 columns\nstress --operation=INSERT -t 2 --num-keys=1 --columns=10 --consistency-level=QUORUM --average-size-values --replication-factor=3 --create-index=KEYS --nodes=cathy1,cathy2\n\n// Verify 1 key in CLI\nuse Keyspace1; \nlist Standard1; \n\n// Restart all three nodes\n\n// You will see 51 keys in CLI\nuse Keyspace1; \nlist Standard1; \n{code}\n\n\n","from":"reporter","subject":"Data from truncated CF reappears after server restart"},{"body":"This is a general issue with all CF's. updating bug.","from":"developer"},{"body":"The other permutation of this bug looked like, assuming write with CL.Q:\n* Insert 50 (3 nodes up)\n* truncate CF (3 nodes up)\n* Insert 1 (3 nodes up)\n* Bring node3 down\n* Delete 1 (2 nodes up)\n* Bring up node3 and run repair\n* Take down node1 and node2.\n* Query node3 with CL.ONE: list Standard1; --- 30 rows returned\n\nNot sure, but this looked suspicious in my logs:\n{code}\n INFO 01:19:45,616 Streaming to /50.57.114.45\n INFO 01:19:45,689 Finished streaming session 698609583499991 from /50.57.107.176\n INFO 01:19:45,690 Finished streaming session 698609609994154 from /50.57.114.45\n INFO 01:19:46,501 Finished streaming repair with /50.57.114.45 for (0,56713727820156410577229101238628035242]: 0 oustanding to complete session\n INFO 01:19:46,531 Compacted to /var/lib/cassandra/data/Keyspace1/Standard1-tmp-g-106-Data.db. 16,646,523 to 16,646,352 (~99% of original) bytes for 30 keys. Time: 1,509ms.\n INFO 01:19:46,930 Finished streaming repair with /50.57.107.176 for (113427455640312821154458202477256070484,0]: 1 oustanding to complete session\n INFO 01:19:47,619 Finished streaming repair with /50.57.114.45 for (113427455640312821154458202477256070484,0]: 0 oustanding to complete session\n INFO 01:19:48,232 Finished streaming repair with /50.57.107.176 for (56713727820156410577229101238628035242,113427455640312821154458202477256070484]: 1 oustanding to complete session\n INFO 01:19:48,856 Finished streaming repair with /50.57.114.45 for (56713727820156410577229101238628035242,113427455640312821154458202477256070484]: 0 oustanding to complete session\n{code}","from":"developer"},{"body":"Currently, truncate does:\n* force a flush\n* record the time\n* delete any sstables older than the time\n\nThis isn't quite enough if the machine crashes shortly afterward, however, since there can be mutations present in the commitlog that were previously truncated and are now resurrected by CL replay.\n\nOne thing we could do is record the truncate time for the CF in the system ks and then ignore mutations older than that, however this would require time synchronization between the client and the server to be accurate.\n","from":"developer"},{"body":"but we record CL \"context\" at time of flush in the sstable it makes, and we on replay we ignore any mutations from before that position.\n\nchecked and we do wait for flush to complete in truncate.","from":"developer"},{"body":"bq. but we record CL \"context\" at time of flush in the sstable it makes, and we on replay we ignore any mutations from before that position.\n\nI think there's something wrong with that, then:\n\n{noformat}\n INFO 21:25:15,274 Replaying /var/lib/cassandra/commitlog/CommitLog-1312924388053.log\nDEBUG 21:25:15,290 Replaying /var/lib/cassandra/commitlog/CommitLog-1312924388053.log starting at 0\nDEBUG 21:25:15,291 Reading mutation at 0\nDEBUG 21:25:15,295 replaying mutation for system.4c: {ColumnFamily(LocationInfo [47656e65726174696f6e:false:4@1312924388140000,])}\nDEBUG 21:25:15,321 Reading mutation at 89\nDEBUG 21:25:15,322 replaying mutation for system.426f6f747374726170: {ColumnFamily(LocationInfo [42:false:1@1312924388203,])}\nDEBUG 21:25:15,322 Reading mutation at 174\nDEBUG 21:25:15,322 replaying mutation for system.4c: {ColumnFamily(LocationInfo [546f6b656e:false:16@1312924388204,])}\nDEBUG 21:25:15,322 Reading mutation at 270\nDEBUG 21:25:15,324 replaying mutation for Keyspace1.3030: {ColumnFamily(Standard1 [C0:false:34@1312924813259,C1:false:34@1312924813260,C2:false:34@1312924813260,C3:false:34@1312924813260,C4:false:34@1312924813260,])}\n{noformat}\n\nThe last entry there is the first of many errant mutations.","from":"developer"},{"body":"Ah, CASSANDRA-2419 keeps on giving...\n\nbq. but we record CL \"context\" at time of flush in the sstable it makes, and we on replay we ignore any mutations from before that position.\n\nThe obvious problem with this is that the point of truncate is to blow away such sstables... Patch attached. Comment explains the core fix:\n\n{noformat}\n// Bonus complication: since we store replay position in sstable metadata,\n// truncating those sstables means we will replay any CL segments from the\n// beginning if we restart before they are discarded for normal reasons\n// post-truncate. So we need to (a) force a new segment so the currently\n// active one can be discarded, and (b) flush *all* CFs so that unflushed\n// data in others don't keep any pre-truncate CL segments alive.\n{noformat}\n\nPatch also fixes the bug in ReplayManagerTruncateTest that made it miss this.\n","from":"developer"},{"body":"I think the forceFlush of all the CF is not safe, because if for a given column family the memtable is clean, forceFlush will return immediately, even though there could be a memtable being flush at the same time (or pending flush). So we cannot be sure all the old segment are clean after the waitFutures (I know, it took me some time to figure out some problem with repair for this very reason when the repairs were not properly synchronized).\n\nWhat we would need is to add to the future we wait on the futures of all the flush being processed at that time. Sounds annoying though. ","from":"developer"},{"body":"+1, though this patch is against trunk, not 0.8. Also mistakenly bumps the log4j level to debug.","from":"developer"},{"body":"v2:\n\n{noformat}\n// Bonus bonus: simply forceFlush of all the CF is not enough, because if\n// for a given column family the memtable is clean, forceFlush will return\n// immediately, even though there could be a memtable being flush at the same\n// time. So to guarantee that all segments can be cleaned out, we need\n// \"waitForActiveFlushes\" after the new segment has been created.\n{noformat}","from":"developer"},{"body":"Attaching a v3 that is rebased against 0.8. I've also slightly change the logic in Truncate to submit all the flushes and then call waitForActiveFlushes, as this is slightly simpler and should work equally well as far as I can tell.\nApart from that, this lgtm.","from":"developer"},{"body":"committed","from":"developer"}],"created":"2011-07-26T20:56:33.000+0000","description":"* Configure 3 node cluster\n* Ensure the java stress tool creates Keyspace1 with RF=3\n\n{code}\n// Run Stress Tool to generate 10 keys, 1 column\nstress --operation=INSERT -t 2 --num-keys=50 --columns=20 --consistency-level=QUORUM --average-size-values --replication-factor=3 --create-index=KEYS --nodes=cathy1,cathy2\n\n// Verify 50 keys in CLI\nuse Keyspace1; \nlist Standard1; \n\n// TRUNCATE CF in CLI\nuse Keyspace1;\ntruncate counter1;\nlist counter1;\n\n// Run stress tool and verify creation of 1 key with 10 columns\nstress --operation=INSERT -t 2 --num-keys=1 --columns=10 --consistency-level=QUORUM --average-size-values --replication-factor=3 --create-index=KEYS --nodes=cathy1,cathy2\n\n// Verify 1 key in CLI\nuse Keyspace1; \nlist Standard1; \n\n// Restart all three nodes\n\n// You will see 51 keys in CLI\nuse Keyspace1; \nlist Standard1; \n{code}\n\n\n","issue_id":"12515455","key":"CASSANDRA-2950","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-08-11T19:36:15.000+0000","role":"fixed_distractor","summary":"Data from truncated CF reappears after server restart"} {"case_id":"12515766","cluster":"DISTRACTOR-CASSANDRA-2968","comments":[{"body":"Reproducible test case for this bug:\n\ncreate column family AffiliateActivity with default_validation_class=CounterColumnType\n and key_validation_class=LongType and comparator=AsciiType\n and memtable_throughput=30\n and memtable_operations=0.5\n and replicate_on_write=true;\n\nput AffiliateActivity-g-195-Data.db and AffiliateActivity-g-195-Index.db (attached) into cassandra 0.8.2 data directory for some keyspace, then run cassandra server to open the files and run nodetool scrub\n\nThose AffiliateActivity-g-195 data files were originally created with cassandra 0.8\n\njava.io.IOError: java.lang.AssertionError\n\tat org.apache.cassandra.db.compaction.CompactionManager.scrubOne(CompactionManager.java:775)\n\tat org.apache.cassandra.db.compaction.CompactionManager.doScrub(CompactionManager.java:631)\n\tat org.apache.cassandra.db.compaction.CompactionManager.access$600(CompactionManager.java:65)\n\tat org.apache.cassandra.db.compaction.CompactionManager$3.call(CompactionManager.java:251)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:166)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n\tat java.lang.Thread.run(Thread.java:679)\nCaused by: java.lang.AssertionError\n\tat org.apache.cassandra.db.context.CounterContext.removeOldShards(CounterContext.java:593)\n\tat org.apache.cassandra.db.CounterColumn.removeOldShards(CounterColumn.java:237)\n\tat org.apache.cassandra.db.CounterColumn.removeOldShards(CounterColumn.java:256)\n\tat org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:88)\n\tat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:140)\n\tat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:146)\n\tat org.apache.cassandra.db.compaction.CompactionManager.scrubOne(CompactionManager.java:719)\n\t... 8 more\n\n\n","created":"2011-07-30T13:18:57.597+0000"},{"body":"Debugger shows at the exception breakpoint in the test case described that\n\"state.getCount()\"\t-9132236803706623116\t\n\"state.getClock()\"\t-9132236803706623116\t\n\nand context.array() (exactly 100 bytes in size) is the following so the duplication is really there, \nand this record in sstable is likely corrupted - but how come!\n\n000100012335e4a0b2c911e000006bb62b336af7 8143c702fe516b74 8143c702fe516b74\n23d44780b2c911e00000f03fd2ec87ff000000000000b3c\n0000000000000b3c025188750b2c911e0000058b7809d76fff082e59147f68e9cf082e59147f68e9c\n","created":"2011-07-30T14:41:52.729+0000"},{"body":"yes, sstable likely corrupted but by which code? 0.8.0 or 0.8.2?\n\nsstable2json AffiliateActivity-g-195-Data.db\n(first row is parsed ok, second has state.getCount()==state.getClock())\n\n....\n\"00000000019d499f\": [[\"clicks\",\"000100012335e4a0b2c911e000006bb62b336af7000000000000285c000000000000285c23d44780b2c911e00000f03fd2ec87ff0000000000000002000000000000000225188750b\n2c911e0000058b7809d76ff00000000000000690000000000000069\",1311664683623,\"c\",-9223372036854775808]],\n\"0000000000748b45\": [[\"clicks\",\"000100012335e4a0b2c911e000006bb62b336af78143c702fe516b748143c702fe516b7423d44780b2c911e00000f03fd2ec87ff000000000000b3c0000000000000b3c025188750b\n2c911e0000058b7809d76fff082e59147f68e9cf082e59147f68e9c\",1311665756252,\"c\",-9223372036854775808]],\n....\n\ncassandra-cli:\nget AffiliateActivity[27085215]; \n=> (counter=clicks, value=10439) - reasonable value\n\nget AffiliateActivity[7637829]; \n=> (counter=clicks, value=8198429924508872144) - nonsense value\n\n","created":"2011-07-30T16:59:26.257+0000"},{"body":"actually 0.8.2 continues to write non-compactable sstables :-(\n\nAffiliateActivity-g-147-Data.db is a recent file written 3 days after upgrade from 0.8.0 to 0.8.2 and it cannot be even scrubbed\n\nThe pattern I can see is that we have only every other sstable file left in data directory - AffiliateActivity-g-133, 135, 137, 139, 141, 143, 145, 147. Not quite sure - but it seems 1 compaction is done to new files - otherwise where are the even numbered ones? - but those files which resulted from first compaction are broken?\n\nAnyway that's purely 0.8.2 problem, not any upgrade issue","created":"2011-07-31T12:43:09.659+0000"},{"body":"For what it's worth I've seen this same bug in 0.8.1.","created":"2011-07-31T17:09:05.100+0000"},{"body":"no, I was wrong, compaction is not the culprit\n\nI ran nodetool flush and produced 3 new small sstables - strangely they are all numbered with odd numbers only 153 155 157 - but the point is a freshly made sstable with only 8 rows is already unreadable by compaction/scrub.\n\nMaybe header in counter binary value is not written or what...\n\n{code}\n{\n\"0000000000cc71a4\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af700000000000000010000000000000001\",1312201208019,\"c\",-9223372036854775808]],\n\"0000000000748b45\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af7000000000000009f000000000000009f23d44780b2c911e00000f03fd2ec87ffa6cd8bac6eb8a32d2fd3c2059ade464925188750b2c911e0000058b7809d76ff9a8b961aed94ef799a8b961aed94ef79\",1312201220709,\"c\",-9223372036854775808]],\n\"00000000010a4465\": [[\"clicks\",\"00002335e4a0b2c911e000006bb62b336af700000000000045ad00000000000045ad23d44780b2c911e00000f03fd2ec87ff0000000000004bfd0000000001fae30325188750b2c911e0000058b7809d76ff000000000000038f000000000000038f\",1312201219737,\"c\",-9223372036854775808]],\n\"0000000001319592\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af70000000000000001000000000000000123d44780b2c911e00000f03fd2ec87ffdf1636876c1c1bc2df1636876c1c1bc225188750b2c911e0000058b7809d76ff0876b1020dc02ac97f12b88d9f9df751\",1312201215475,\"c\",-9223372036854775808]],\n\"0000000001b9a53e\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af700000000000000010000000000000001\",1312201202338,\"c\",-9223372036854775808]],\n\"00000000004dd84f\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af70000000000000001000000000000000123d44780b2c911e00000f03fd2ec87ff000000000017d1550000000059fb2ad925188750b2c911e0000058b7809d76ff00000000000173bf00000000000173bf\",1312201205636,\"c\",-9223372036854775808]],\n\"0000000000d1d52f\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af700000000000000020000000000000002\",1312201210515,\"c\",-9223372036854775808]],\n\"0000000001410d4a\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af700000000000000010000000000000001\",1312201220365,\"c\",-9223372036854775808]]\n}\n{code}","created":"2011-08-01T12:38:00.266+0000"},{"body":"This is actually a pretty stupid bug (not that there is smart bug): the old NodeId for the local node were read from the system table in reversed order while they shouldn't. The wrong path was then taken based on that mistake. No data was lost due to that (i.e, the total value of the counters is preserved), but non-sensical counter context were created (hence triggering the assertion).\n\nFixing the root cause is pretty straightforward. Fixing the nonsensical counter contexts is more subtle, but it is doable up to the fact that the local NodeId on the node(s) where the assertion is triggered will have to be renewed. Attaching a patch that does both (fixing root cause and repairing the bad data). Also add two unit tests, one for the root cause and one to check that the bad data repair code does what it is supposed to do.\n\nAfter applying that patch (or upgrading on a release shipping it), you will (potentially) need to restart the node with the -Dcassandra.renew_counter_id=true (compaction will still fail if you don't but with a message saying that you should restart with the startup flag).","created":"2011-08-01T17:22:37.441+0000"},{"body":"+1 w/ println removed :)","created":"2011-08-02T14:17:42.647+0000"},{"body":"Committed (without println), thanks","created":"2011-08-02T15:05:54.799+0000"}],"conversations":[{"body":"Having upgraded from 0.8.0 to 0.8.2 we ran nodetool compact and got\n\nError occured during compaction\njava.util.concurrent.ExecutionException: java.lang.AssertionError\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.db.compaction.CompactionManager.performMajor(CompactionManager.java:277)\n at org.apache.cassandra.db.ColumnFamilyStore.forceMajorCompaction(ColumnFamilyStore.java:1762)\n at org.apache.cassandra.service.StorageService.forceTableCompaction(StorageService.java:1358)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:93)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:27)\n at com.sun.jmx.mbeanserver.MBeanIntrospector.invokeM(MBeanIntrospector.java:208)\n at com.sun.jmx.mbeanserver.PerInterface.invoke(PerInterface.java:120)\n at com.sun.jmx.mbeanserver.MBeanSupport.invoke(MBeanSupport.java:262)\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.invoke(DefaultMBeanServerInterceptor.java:836)\n at com.sun.jmx.mbeanserver.JmxMBeanServer.invoke(JmxMBeanServer.java:761)\n at javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1427)\n at javax.management.remote.rmi.RMIConnectionImpl.access$200(RMIConnectionImpl.java:72)\n at javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1265)\n at javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1360)\n at javax.management.remote.rmi.RMIConnectionImpl.invoke(RMIConnectionImpl.java:788)\n at sun.reflect.GeneratedMethodAccessor24.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:305)\n at sun.rmi.transport.Transport$1.run(Transport.java:159)\n at java.security.AccessController.doPrivileged(Native Method)\n at sun.rmi.transport.Transport.serviceCall(Transport.java:155)\n at sun.rmi.transport.tcp.TCPTransport.handleMessages(TCPTransport.java:535)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run0(TCPTransport.java:790)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run(TCPTransport.java:649)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.lang.AssertionError \n at org.apache.cassandra.db.context.CounterContext.removeOldShards(CounterContext.java:593) \n at org.apache.cassandra.db.CounterColumn.removeOldShards(CounterColumn.java:237) \n at org.apache.cassandra.db.CounterColumn.removeOldShards(CounterColumn.java:256) \n at org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:88) \n at org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:140) \n at org.apache.cassandra.db.compaction.CompactionIterator.getReduced(CompactionIterator.java:123) \n at org.apache.cassandra.db.compaction.CompactionIterator.getReduced(CompactionIterator.java:43) \n at org.apache.cassandra.utils.ReducingIterator.computeNext(ReducingIterator.java:74) \n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:140) \n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:135) \n at org.apache.commons.collections.iterators.FilterIterator.setNextObject(FilterIterator.java:183) \n at org.apache.commons.collections.iterators.FilterIterator.hasNext(FilterIterator.java:94) \n at org.apache.cassandra.db.compaction.CompactionManager.doCompactionWithoutSizeEstimation(CompactionManager.java:569) \n at org.apache.cassandra.db.compaction.CompactionManager.doCompaction(CompactionManager.java:506) \n at org.apache.cassandra.db.compaction.CompactionManager$4.call(CompactionManager.java:319) \n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303) \n at java.util.concurrent.FutureTask.run(FutureTask.java:138) \n ... 3 more","from":"reporter","subject":"AssertionError during compaction of CF with counter columns"},{"body":"Reproducible test case for this bug:\n\ncreate column family AffiliateActivity with default_validation_class=CounterColumnType\n and key_validation_class=LongType and comparator=AsciiType\n and memtable_throughput=30\n and memtable_operations=0.5\n and replicate_on_write=true;\n\nput AffiliateActivity-g-195-Data.db and AffiliateActivity-g-195-Index.db (attached) into cassandra 0.8.2 data directory for some keyspace, then run cassandra server to open the files and run nodetool scrub\n\nThose AffiliateActivity-g-195 data files were originally created with cassandra 0.8\n\njava.io.IOError: java.lang.AssertionError\n\tat org.apache.cassandra.db.compaction.CompactionManager.scrubOne(CompactionManager.java:775)\n\tat org.apache.cassandra.db.compaction.CompactionManager.doScrub(CompactionManager.java:631)\n\tat org.apache.cassandra.db.compaction.CompactionManager.access$600(CompactionManager.java:65)\n\tat org.apache.cassandra.db.compaction.CompactionManager$3.call(CompactionManager.java:251)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:166)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n\tat java.lang.Thread.run(Thread.java:679)\nCaused by: java.lang.AssertionError\n\tat org.apache.cassandra.db.context.CounterContext.removeOldShards(CounterContext.java:593)\n\tat org.apache.cassandra.db.CounterColumn.removeOldShards(CounterColumn.java:237)\n\tat org.apache.cassandra.db.CounterColumn.removeOldShards(CounterColumn.java:256)\n\tat org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:88)\n\tat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:140)\n\tat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:146)\n\tat org.apache.cassandra.db.compaction.CompactionManager.scrubOne(CompactionManager.java:719)\n\t... 8 more\n\n\n","from":"developer"},{"body":"Debugger shows at the exception breakpoint in the test case described that\n\"state.getCount()\"\t-9132236803706623116\t\n\"state.getClock()\"\t-9132236803706623116\t\n\nand context.array() (exactly 100 bytes in size) is the following so the duplication is really there, \nand this record in sstable is likely corrupted - but how come!\n\n000100012335e4a0b2c911e000006bb62b336af7 8143c702fe516b74 8143c702fe516b74\n23d44780b2c911e00000f03fd2ec87ff000000000000b3c\n0000000000000b3c025188750b2c911e0000058b7809d76fff082e59147f68e9cf082e59147f68e9c\n","from":"developer"},{"body":"yes, sstable likely corrupted but by which code? 0.8.0 or 0.8.2?\n\nsstable2json AffiliateActivity-g-195-Data.db\n(first row is parsed ok, second has state.getCount()==state.getClock())\n\n....\n\"00000000019d499f\": [[\"clicks\",\"000100012335e4a0b2c911e000006bb62b336af7000000000000285c000000000000285c23d44780b2c911e00000f03fd2ec87ff0000000000000002000000000000000225188750b\n2c911e0000058b7809d76ff00000000000000690000000000000069\",1311664683623,\"c\",-9223372036854775808]],\n\"0000000000748b45\": [[\"clicks\",\"000100012335e4a0b2c911e000006bb62b336af78143c702fe516b748143c702fe516b7423d44780b2c911e00000f03fd2ec87ff000000000000b3c0000000000000b3c025188750b\n2c911e0000058b7809d76fff082e59147f68e9cf082e59147f68e9c\",1311665756252,\"c\",-9223372036854775808]],\n....\n\ncassandra-cli:\nget AffiliateActivity[27085215]; \n=> (counter=clicks, value=10439) - reasonable value\n\nget AffiliateActivity[7637829]; \n=> (counter=clicks, value=8198429924508872144) - nonsense value\n\n","from":"developer"},{"body":"actually 0.8.2 continues to write non-compactable sstables :-(\n\nAffiliateActivity-g-147-Data.db is a recent file written 3 days after upgrade from 0.8.0 to 0.8.2 and it cannot be even scrubbed\n\nThe pattern I can see is that we have only every other sstable file left in data directory - AffiliateActivity-g-133, 135, 137, 139, 141, 143, 145, 147. Not quite sure - but it seems 1 compaction is done to new files - otherwise where are the even numbered ones? - but those files which resulted from first compaction are broken?\n\nAnyway that's purely 0.8.2 problem, not any upgrade issue","from":"developer"},{"body":"For what it's worth I've seen this same bug in 0.8.1.","from":"developer"},{"body":"no, I was wrong, compaction is not the culprit\n\nI ran nodetool flush and produced 3 new small sstables - strangely they are all numbered with odd numbers only 153 155 157 - but the point is a freshly made sstable with only 8 rows is already unreadable by compaction/scrub.\n\nMaybe header in counter binary value is not written or what...\n\n{code}\n{\n\"0000000000cc71a4\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af700000000000000010000000000000001\",1312201208019,\"c\",-9223372036854775808]],\n\"0000000000748b45\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af7000000000000009f000000000000009f23d44780b2c911e00000f03fd2ec87ffa6cd8bac6eb8a32d2fd3c2059ade464925188750b2c911e0000058b7809d76ff9a8b961aed94ef799a8b961aed94ef79\",1312201220709,\"c\",-9223372036854775808]],\n\"00000000010a4465\": [[\"clicks\",\"00002335e4a0b2c911e000006bb62b336af700000000000045ad00000000000045ad23d44780b2c911e00000f03fd2ec87ff0000000000004bfd0000000001fae30325188750b2c911e0000058b7809d76ff000000000000038f000000000000038f\",1312201219737,\"c\",-9223372036854775808]],\n\"0000000001319592\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af70000000000000001000000000000000123d44780b2c911e00000f03fd2ec87ffdf1636876c1c1bc2df1636876c1c1bc225188750b2c911e0000058b7809d76ff0876b1020dc02ac97f12b88d9f9df751\",1312201215475,\"c\",-9223372036854775808]],\n\"0000000001b9a53e\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af700000000000000010000000000000001\",1312201202338,\"c\",-9223372036854775808]],\n\"00000000004dd84f\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af70000000000000001000000000000000123d44780b2c911e00000f03fd2ec87ff000000000017d1550000000059fb2ad925188750b2c911e0000058b7809d76ff00000000000173bf00000000000173bf\",1312201205636,\"c\",-9223372036854775808]],\n\"0000000000d1d52f\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af700000000000000020000000000000002\",1312201210515,\"c\",-9223372036854775808]],\n\"0000000001410d4a\": [[\"clicks\",\"000100002335e4a0b2c911e000006bb62b336af700000000000000010000000000000001\",1312201220365,\"c\",-9223372036854775808]]\n}\n{code}","from":"developer"},{"body":"This is actually a pretty stupid bug (not that there is smart bug): the old NodeId for the local node were read from the system table in reversed order while they shouldn't. The wrong path was then taken based on that mistake. No data was lost due to that (i.e, the total value of the counters is preserved), but non-sensical counter context were created (hence triggering the assertion).\n\nFixing the root cause is pretty straightforward. Fixing the nonsensical counter contexts is more subtle, but it is doable up to the fact that the local NodeId on the node(s) where the assertion is triggered will have to be renewed. Attaching a patch that does both (fixing root cause and repairing the bad data). Also add two unit tests, one for the root cause and one to check that the bad data repair code does what it is supposed to do.\n\nAfter applying that patch (or upgrading on a release shipping it), you will (potentially) need to restart the node with the -Dcassandra.renew_counter_id=true (compaction will still fail if you don't but with a message saying that you should restart with the startup flag).","from":"developer"},{"body":"+1 w/ println removed :)","from":"developer"},{"body":"Committed (without println), thanks","from":"developer"}],"created":"2011-07-29T12:00:40.000+0000","description":"Having upgraded from 0.8.0 to 0.8.2 we ran nodetool compact and got\n\nError occured during compaction\njava.util.concurrent.ExecutionException: java.lang.AssertionError\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.db.compaction.CompactionManager.performMajor(CompactionManager.java:277)\n at org.apache.cassandra.db.ColumnFamilyStore.forceMajorCompaction(ColumnFamilyStore.java:1762)\n at org.apache.cassandra.service.StorageService.forceTableCompaction(StorageService.java:1358)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:93)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:27)\n at com.sun.jmx.mbeanserver.MBeanIntrospector.invokeM(MBeanIntrospector.java:208)\n at com.sun.jmx.mbeanserver.PerInterface.invoke(PerInterface.java:120)\n at com.sun.jmx.mbeanserver.MBeanSupport.invoke(MBeanSupport.java:262)\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.invoke(DefaultMBeanServerInterceptor.java:836)\n at com.sun.jmx.mbeanserver.JmxMBeanServer.invoke(JmxMBeanServer.java:761)\n at javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1427)\n at javax.management.remote.rmi.RMIConnectionImpl.access$200(RMIConnectionImpl.java:72)\n at javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1265)\n at javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1360)\n at javax.management.remote.rmi.RMIConnectionImpl.invoke(RMIConnectionImpl.java:788)\n at sun.reflect.GeneratedMethodAccessor24.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:305)\n at sun.rmi.transport.Transport$1.run(Transport.java:159)\n at java.security.AccessController.doPrivileged(Native Method)\n at sun.rmi.transport.Transport.serviceCall(Transport.java:155)\n at sun.rmi.transport.tcp.TCPTransport.handleMessages(TCPTransport.java:535)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run0(TCPTransport.java:790)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run(TCPTransport.java:649)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.lang.AssertionError \n at org.apache.cassandra.db.context.CounterContext.removeOldShards(CounterContext.java:593) \n at org.apache.cassandra.db.CounterColumn.removeOldShards(CounterColumn.java:237) \n at org.apache.cassandra.db.CounterColumn.removeOldShards(CounterColumn.java:256) \n at org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:88) \n at org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:140) \n at org.apache.cassandra.db.compaction.CompactionIterator.getReduced(CompactionIterator.java:123) \n at org.apache.cassandra.db.compaction.CompactionIterator.getReduced(CompactionIterator.java:43) \n at org.apache.cassandra.utils.ReducingIterator.computeNext(ReducingIterator.java:74) \n at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:140) \n at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:135) \n at org.apache.commons.collections.iterators.FilterIterator.setNextObject(FilterIterator.java:183) \n at org.apache.commons.collections.iterators.FilterIterator.hasNext(FilterIterator.java:94) \n at org.apache.cassandra.db.compaction.CompactionManager.doCompactionWithoutSizeEstimation(CompactionManager.java:569) \n at org.apache.cassandra.db.compaction.CompactionManager.doCompaction(CompactionManager.java:506) \n at org.apache.cassandra.db.compaction.CompactionManager$4.call(CompactionManager.java:319) \n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303) \n at java.util.concurrent.FutureTask.run(FutureTask.java:138) \n ... 3 more","issue_id":"12515766","key":"CASSANDRA-2968","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-08-02T15:05:54.000+0000","role":"fixed_distractor","summary":"AssertionError during compaction of CF with counter columns"} {"case_id":"12525677","cluster":"DISTRACTOR-CASSANDRA-3306","comments":[{"body":"another problem. why not store data in some system CF? would be probably safer choice.\n\nERROR [CompactionExecutor:5] 2011-10-04 17:13:13,922 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[CompactionExecutor:5,5,main]\njava.io.IOError: java.io.IOException: Failed to rename \\var\\lib\\cassandra\\data\\test\\sipdb.json to \\var\\lib\\cassandra\\data\\test\\sipdb-old.json\n\tat org.apache.cassandra.db.compaction.LeveledManifest.serialize(LeveledManifest.java:382)\n\tat org.apache.cassandra.db.compaction.LeveledManifest.promote(LeveledManifest.java:182)\n\tat org.apache.cassandra.db.compaction.LeveledCompactionStrategy.handleNotification(LeveledCompactionStrategy.java:152)\n\tat org.apache.cassandra.db.DataTracker.notifySSTablesChanged(DataTracker.java:466)\n\tat org.apache.cassandra.db.DataTracker.replace(DataTracker.java:275)\n\tat org.apache.cassandra.db.DataTracker.replaceCompactedSSTables(DataTracker.java:232)\n\tat org.apache.cassandra.db.ColumnFamilyStore.replaceCompactedSSTables(ColumnFamilyStore.java:960)\n\tat org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:199)\n\tat org.apache.cassandra.db.compaction.LeveledCompactionTask.execute(LeveledCompactionTask.java:47)\n\tat org.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:131)\n\tat org.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:114)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\nCaused by: java.io.IOException: Failed to rename \\var\\lib\\cassandra\\data\\test\\sipdb.json to \\var\\lib\\cassandra\\data\\test\\sipdb-old.json\n\tat org.apache.cassandra.io.util.FileUtils.renameWithConfirm(FileUtils.java:64)\n\tat org.apache.cassandra.db.compaction.LeveledManifest.serialize(LeveledManifest.java:375)\n\t... 15 more\n","created":"2011-10-04T15:18:55.646+0000"},{"body":"Because then you get into hairy cyclical situations where you can't read the manifest until you replay the commitlog, but replaying the commitlog requires writing new sstables and thus knowing the manifest","created":"2011-10-04T15:27:07.848+0000"},{"body":"as i understand new flushed tables are placed at level 0. Just replay commitlog and put all new stuff in lvl 0. after comitlog is done, it can do voodoo shuffles.\n\nbut why not to rename tables like table-h-333-l1-Data.db?\n\nidea to have stables with non overlapping key ranges is interesting, but read performance is kinda slow (about 50% of normal) here. Its cassandra core modified to get advantage of leveled tables? i.e. search one sstable at level1, one at lvl2 using bloom filters for key?","created":"2011-10-04T15:52:40.139+0000"},{"body":"This isn't really a great place to rehash http://leveldb.googlecode.com/svn/trunk/doc/impl.html and CASSANDRA-1608.","created":"2011-10-04T16:01:27.932+0000"},{"body":"bq. why not store data in some system CF? would be probably safer choice.\n\nThis has historically been a bad idea, see CASSANDRA-1155, then CASSANDRA-1318 and finally CASSANDRA-1430.","created":"2011-10-04T16:04:26.256+0000"},{"body":"I don't suppose you were using column family truncation in your tests, where you? ","created":"2011-10-24T14:25:42.406+0000"},{"body":"no truncation, no supercolumns.","created":"2011-10-24T16:42:01.183+0000"},{"body":"Are you still able to reproduce reliably? Because we aren't and being able to would help considerably, so if you are and could share whatever script you're using to reproduce, that would be awesome.","created":"2011-10-24T16:45:52.815+0000"},{"body":"i tested it on 1.0 final and it worked without error for 1 test run. i will give it another test without index.","created":"2011-10-25T06:33:38.041+0000"},{"body":"I'll note that Ramesh Natarajan reported on the mailing list what clearly appears to be the same bug (http://www.mail-archive.com/user@cassandra.apache.org/msg18146.html), but while not using leveled compaction. I also think he was using the 1.0.0 final.","created":"2011-10-26T16:14:07.714+0000"},{"body":"I'll note that more info have been added to the messages thrown by the exception here in 1.0.1. So if someone can reproduce this issue on 1.0.1, it would be useful to get the stacktrace (the full system.log would actually be even better).","created":"2011-11-04T18:32:42.886+0000"},{"body":"This AssertionError happened always in cassandra1.0.0 ,not just only in LeveledCompactionStrategy","created":"2011-11-24T02:16:39.176+0000"},{"body":"I suppose it's a bug in DataTracker .","created":"2011-11-24T02:24:04.971+0000"},{"body":"As I already said, if you are able to reproduce this, please try reproducing with 1.0.3. And if you are still able to, please attach you system.log with the exception here because it will have more info on the error that should help. And if you're not able to reproduce with 1.0.3, then I guess it means we've fixed it without knowing.","created":"2011-11-24T08:10:53.294+0000"},{"body":"Running under 1.0.4, can easily reproduce this by just kicking of a repair of any LeveledCompactionStrategy CF. \n\nThe 'zero' on the assert indicates the value (added that to the code to see what the value was): \n\njava.lang.AssertionError: 0\n at org.apache.cassandra.db.compaction.LeveledManifest.promote(LeveledManifest.java:178)\n at org.apache.cassandra.db.compaction.LeveledCompactionStrategy.handleNotification(LeveledCompactionStrategy.java:141)\n at org.apache.cassandra.db.DataTracker.notifySSTablesChanged(DataTracker.java:481)\n at org.apache.cassandra.db.DataTracker.replace(DataTracker.java:275)\n at org.apache.cassandra.db.DataTracker.addSSTables(DataTracker.java:237)\n at org.apache.cassandra.db.DataTracker.addStreamedSSTable(DataTracker.java:242)\n at org.apache.cassandra.db.ColumnFamilyStore.addSSTable(ColumnFamilyStore.java:920)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:141)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:103)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:184)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:81)\n\n\nRelevant lines from system.log leading up the it: \n\n INFO [FlushWriter:794] 2011-12-01 14:23:22,966 Memtable.java (line 275) Completed flushing /var/lib/cassandra/data/sso/Sessions-hc-12524-Data.db (1119784 bytes)\n INFO [CompactionExecutor:2379] 2011-12-01 14:23:22,969 CompactionTask.java (line 112) Compacting [SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12501-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12517-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12513-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12512-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12502-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12507-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12519-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12500-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12508-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12504-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12510-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12515-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12509-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12524-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12514-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12518-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12505-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12516-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12511-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12506-Data.db')]\n INFO [AntiEntropyStage:1] 2011-12-01 14:25:06,321 AntiEntropyService.java (line 186) [repair #ea080b70-1c51-11e1-0000-692e0c239dfd] Received merkle tree for Sessions from /xxxxxxxx\nERROR [Thread-177] 2011-12-01 14:25:17,863 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[Thread-177,5,main]\njava.lang.AssertionError: 0 [see above]\n\nIf you want more let me know I can reproduce instantly.\n","created":"2011-12-01T19:42:01.631+0000"},{"body":"Joe, your assertion is the one in CASSANDRA-3536 (where I've attached a patch fixing it). Closing this other one as cantrepro.","created":"2011-12-02T05:36:16.848+0000"},{"body":"This error actually happens on 1.1. And I can easily reproduce with unit test(Test code attached).\n\n{code}\n [junit] ERROR 17:34:46,696 Fatal exception in thread Thread[CompactionExecutor:3,1,main]\n [junit] java.lang.AssertionError: Expecting new size of 2, got 1 while replacing [SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-1-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-5-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-4-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-2-Data.db')] by [SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-6-Data.db')] in View(pending_count=0, sstables=[SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-1-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-2-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-4-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-4-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-5-Data.db')], compacting=[SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-1-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-5-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-4-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-2-Data.db')])\n [junit] \tat org.apache.cassandra.db.DataTracker$View.newSSTables(DataTracker.java:651)\n [junit] \tat org.apache.cassandra.db.DataTracker$View.replace(DataTracker.java:616)\n [junit] \tat org.apache.cassandra.db.DataTracker.replace(DataTracker.java:320)\n [junit] \tat org.apache.cassandra.db.DataTracker.replaceCompactedSSTables(DataTracker.java:253)\n [junit] \tat org.apache.cassandra.db.ColumnFamilyStore.replaceCompactedSSTables(ColumnFamilyStore.java:994)\n [junit] \tat org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:200)\n [junit] \tat org.apache.cassandra.db.compaction.CompactionManager$1.runMayThrow(CompactionManager.java:154)\n [junit] \tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n [junit] \tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:441)\n [junit] \tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n [junit] \tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n [junit] \tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n [junit] \tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n [junit] \tat java.lang.Thread.run(Thread.java:680)\n{code}\n\nThe cause is actually in streaming. StreamInSession can add duplicate reference to SSTable to DataTracker when it is left even after stream session finishes. This typically happens when source node is marked as dead by FailureDetector during streaming session(GC storm is the one I saw) and keep sending file in same session after the node comes back.","created":"2012-10-23T22:44:08.200+0000"},{"body":"Test code attached. Compaction strategy is not related.","created":"2012-10-23T22:45:05.121+0000"},{"body":"Good analysis Yuki. I'm not really sure what is the right fix though. Given that this should very rarely happen (repair uses a much higher failure detection threshold than the normal one, though maybe we can increase it even more to make this even less likely) and that I don't seen any obvious way to avoid that kind of situation, maybe making DataTracker handle duplicate addition of a SSTableReader is the simplest thing to do. The obvious way to do that would be to change the View sstables List to a Set, which leads me to the current commentary in the code:\n{noformat}\n // We can't use a SortedSet here because \"the ordering maintained by a sorted set (whether or not an\n // explicit comparator is provided) must be consistent with equals.\" In particular,\n // ImmutableSortedSet will ignore any objects that compare equally with an existing Set member.\n // Obviously, dropping sstables whose max column timestamp happens to be equal to another's\n // is not acceptable for us. So, we use a List instead.\n{noformat}\nI think that comment is obsolete. Namely, it was added with CASSANDRA-2498 and at the time the list of sstable was kept in max timestamp order at all time. But since then, we've moved the sorting in max timestamp in CollationController directly (which is less fragile), so the order inside DataTracker doesn't matter anymore.","created":"2012-10-24T12:11:13.882+0000"},{"body":"bq. This typically happens when source node is marked as dead by FailureDetector during streaming session(GC storm is the one I saw) and keep sending file in same session after the node comes back\n\nBut we close the session on convict, so shouldn't it start a new one?","created":"2012-10-24T20:46:37.261+0000"},{"body":"bq. But we close the session on convict, so shouldn't it start a new one?\n\nYes, StreamInSession gets closed and removed on convict _once_. But if GC pause happens in the middle of streaming session, the node resumes streaming in the same session after GC. Since resumed stream carries session ID that is once closed on receiver side, StreamInSession is created again with the same old session ID and this time just 1 file to receive.\nThis continues again and again until source node's StreamingOutSession sends all files.\nYou can see this in receiver's log file like below:\n\n{code}\nINFO [Thread-50] 2012-10-20 13:13:26,574 StreamInSession.java (line 214) Finished streaming session 10 from /10.xx.xx.xx\nINFO [Thread-51] 2012-10-20 13:13:29,691 StreamInSession.java (line 214) Finished streaming session 10 from /10.xx.xx.xx\nINFO [Thread-52] 2012-10-20 13:13:32,957 StreamInSession.java (line 214) Finished streaming session 10 from /10.xx.xx.xx\n{code}\n\nDuplication happens during this partially broken streaming session. Because StreamInSession is removed after sending SESSION_FINISHED reply, and StreamOutSession keeps sending files, sometimes the same StreamInSession instance receives more than 1 file and calls closeIfFinished every time it received the file.\n(Sorry, this is hard to explain in words.\nhttps://github.com/apache/cassandra/blob/cassandra-1.1.6/src/java/org/apache/cassandra/streaming/StreamInSession.java#L181 this part is executed multiple times with _readers_ growing by received new file.)\n\nSo as Sylvain stated above, changing DataTracker.View's sstable to Set is one way to eliminate duplicate reference and we should do it. In addition, I'm thinking not to create duplicate StreamInSession by checking StreamHeader.pendingFiles because this field is only filled when initiating streaming.","created":"2012-10-24T22:00:17.754+0000"},{"body":"That code is a mess so let me give a shot at describing what happens for the record. Say node1 wants to stream files A, B and C to node2. If everything goes well what happens is:\n# node1 sends the first file A with a StreamHeader that says that A, B and C are pending files and A is the currently sent file. On node2, a new StreamInSession is created with those information.\n# Once A is finished, node2 remove A from the pending file in the StreamInSession send an acknowledgement to node1, and then node1 sends B with a StreamHeader with no pending files (basically the list of pending files is only sent the first time so that the StreamInSession on node2 knows when everything is finished) and B as current file. When node2 received that StreamHeader, it retrieve the StreamInSession, setting B as the current files.\n# Once B is finished, node2 removes it from pending files, acks to node1 and node1 sends C with a StreamHeader with no pending file and C as current file. Node2 retrieven the StreamInSession and modify it accordingly.\n# At last, once C is finished, node2 removes it from the pending files. Then it realizes the pending files are empty and so that the streaming is finished and at that point it adds all the SSTableReader created so far to the cfs (and acks to node1 the end of the streaming).\n\nNow, the problem is if say node1 is marked dead by mistake by node2 during say the streaming of A. I that happens, the only thing we do on node2 is to close the session and remove the streamInSession from the global sessions map. However we don't shutdown the stream or anything, so if node1 is in fact still alive, what will happen is:\n# A will finish his transfer correctly. Once that's done, node2 will still send an acknowledgement (probably the first mistake, we could check that the session has been closed and send an error instead).\n# Node1 getting it's acknowledgement will send B with a StreamHeader that has B as current file and no pending files as usual. On reception, node2 will not find any StreamInSession (it has been removed during the close), and so it will create a new one as if that was the beginning of a transfer. And that session will have no pending file (second mistake: if we have to create a new StreamInSession but there is no pending file at all something wrong has happened).\n# Once B is fully streamed, node2 will acknowledge it to node1 and remove it from it's streamInSession. But that session the new one we just created with no pending file. So the streamInSession will consider the streaming is finished, and it will thus add the SSTableReader for B to the cfs.\n# Because B has been acknowledged, node1 will start sending C (again, with no pending file in the StreamHeader). This will be done as soon as B was finished, and so concurrently with the streamInSession on node2 closing itself.\n# So when node2 receives the StreamHeader with C, it will try to retrieve the session and will find the previous session. And will happily add C as the current file for that session (third and fourth mistake: StreamInSession should not add a file as current unless it is a pending file for this session, and a session could detect that it's being reused even though it has just detected itself as finished).\n# Now when C transfer finishes, the seesion will be notify and since it still has no pending files, it will once again consider the streaming as complete. But since it's still the same session, it still has the SSTableReader for B in its list of created reader (as well as the one for C now). And that's when it adds B for a second time to the DataTracker.\n\nI also not that we end up without having ever add the SSTableReader for A to the cfs since the very first StreamInSession was never finished. This is not a big deal in that the stream itself has been indicated as failed to the client anyway, but just to say that it's not just a problem of duplicating a SSTableReader preference.\n\nAnyway, let me back on what I said earlier. We should definitively fix some if not all of the \"mistake\" above (and send a SESSION_FAILURE to node1 as soon as we detect something is wrong).\n\nBut that being said, my comment on the comment in DataTracker being obsolete still stand, and replacing the list by a set in there would have at least the advantage of slightly simplifying the code of DataTracker.View.newSSTables(), as well as being more resilient if a SSTableReader is added twice. Not a big deal though.\n","created":"2012-10-25T07:44:01.049+0000"},{"body":"Attaching first attempt.\nI changed DataTracker.View's sstables to Set, and made stream fail when file arrives after StreamInSession failed.\n\nChanging List to Set for sstables sometimes makes CollationControllerTest fail. It was introduced in CASSANDRA-4116, and I think the test and CollationController#collectAllData expect sstables to be ordered by timestamp. I'm not sure if the test is obsolete or we really need sstables to be sorted all the time.\n0002 patch alone will fix the issue, so we can apply that for now.","created":"2012-10-29T20:28:38.228+0000"},{"body":"For patch 0002, we shouldn't check the FailureDetector otherwise we don't really fix the issue. The only way we know this bug can happen is wher the FailureDetector *had* marked a node down while it shouldn't have (besides, we just got something from a node so it's fair to assume it is alive).\n\nbq. and I think the test and CollationController#collectAllData expect sstables to be ordered by timestamp\n\nIt doesn't seem to me that collectAllData needs sstable ordered. In fact, I think that it does a second pass over the sstables iterators just because it doesn't assume sstables are ordered by max timestamp. Moreover, I'm pretty sure it would be a bug to assume that. If you look at DataTracker.View.newSSTables, it ends by {{Iterables.addAll(newSSTables, replacements)}} which clearly won't maintain any specific ordering of sstables.\n\nbq. I'm not sure if the test is obsolete.\n\nI don't think the test is obsolete but I think we have a minor bug in CollationController. The test want to test that we correctly exclude sstable whose maxTimestamp is less than the most recent row tombstone we have. But that test checks controller.getSstablesIterated(), and for collectAllData, it will count every sstable it include in the first iteration of collectAllData but don't remove those that are remove by the second pass. In other words, I think the correct fix is to decrement stablesIterated in CollationController when in the second pass we remove a sstable (or more simply to set it to iterators.size() just before we collate everything).","created":"2012-10-30T08:15:53.189+0000"},{"body":"bq. we shouldn't check the FailureDetector otherwise we don't really fix the issue.\n\nOk. I've fixed this and reattached 0002.\n\nbq. The test want to test that we correctly exclude sstable whose maxTimestamp is less than the most recent row tombstone we have.\n\nRight. But I think the test assumes that SSTables are added to List in order of flush, and that's true as long as we use List. So what I suggest is to remove that part from the test since we no longer use List.\nAnd sstablesIterated counter in collectAllData is doing fine because we actually read the data from sstable when we go over\n{code}\nIColumnIterator iter = filter.getSSTableColumnIterator(sstable);\n{code}\nbefore incrementing counter.\n\nSo I removed that test from CollationControllerTest in 0001-change-DataTracker.View-s-sstables-from-List-to-Set.patch.","created":"2012-10-30T22:34:21.628+0000"},{"body":"bq. Ok. I've fixed this and reattached 0002.\n\nAlright, +1 on 0002. Let's commit that for now to 1.1/1.2 as this fix this ticket.\n\n{quote}\nthe test assumes that SSTables are added to List in order of flush\nAnd sstablesIterated counter in collectAllData is doing fine because we actually read the data\n{quote}\n\nRight. I guess what I meant is that what is tested right now is not really sensible. Relying on the order of flush is only valid for a small, controlled test, but in reality as soon as compaction kicks in, the order of sstable in DataTracker will be meaningless even with a List instead of a Set. Basically the guarantee collectAll gives us today is that it will eliminate sstables whose maxTimestamp < mostRecentTombstone with just having read the sstable row header, not the full data. But that's not what sstablesIterated counts so it's broken.\n\nThat being said, I think we can improve collectAll in the way described in CASSANDRA-4883. If we do so, the test will pass again without relying on any assumption of the order of sstables in DataTracker. So overall I suggest moving all of this to CASSANDRA-4883.","created":"2012-10-31T09:48:47.670+0000"},{"body":"Committed 0002 to 1.1 and trunk.","created":"2012-10-31T16:13:45.629+0000"},{"body":"I understand that this has been fixed in newer versions of Cassandra.\n\nBut I'm currently seeing this exact issue on a production 1.1.1 node in my cluster.\n\nWhat should be my next step?\n\nDo I simply restart it?\n\nRun cleanup? Scrub? Repair?\n\nSounds like repair would just fail with the same problem.\n\nAny advice would be appreciated.","created":"2015-03-11T19:44:40.964+0000"},{"body":"Yes, restarting the node will help.\nNo need to clean up/scrub.\n\nPlease use user@cassandra.apache.org mailing list for these type of questions.","created":"2015-03-11T19:56:20.704+0000"},{"body":"Thanks for the swift reply and I will use the mailing list in the future.","created":"2015-03-11T20:32:13.485+0000"}],"conversations":[{"body":"during stress testing, i always get this error making leveledcompaction strategy unusable. Should be easy to reproduce - just write fast.\n\nERROR [CompactionExecutor:6] 2011-10-04 15:48:52,179 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[CompactionExecutor:6,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.db.DataTracker$View.newSSTables(DataTracker.java:580)\n\tat org.apache.cassandra.db.DataTracker$View.replace(DataTracker.java:546)\n\tat org.apache.cassandra.db.DataTracker.replace(DataTracker.java:268)\n\tat org.apache.cassandra.db.DataTracker.replaceCompactedSSTables(DataTracker.java:232)\n\tat org.apache.cassandra.db.ColumnFamilyStore.replaceCompactedSSTables(ColumnFamilyStore.java:960)\n\tat org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:199)\n\tat org.apache.cassandra.db.compaction.LeveledCompactionTask.execute(LeveledCompactionTask.java:47)\n\tat org.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:131)\n\tat org.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:114)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\n\nand this is in json data for table:\n\n{\n \"generations\" : [ {\n \"generation\" : 0,\n \"members\" : [ 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484 ]\n }, {\n \"generation\" : 1,\n \"members\" : [ ]\n }, {\n \"generation\" : 2,\n \"members\" : [ ]\n }, {\n \"generation\" : 3,\n \"members\" : [ ]\n }, {\n \"generation\" : 4,\n \"members\" : [ ]\n }, {\n \"generation\" : 5,\n \"members\" : [ ]\n }, {\n \"generation\" : 6,\n \"members\" : [ ]\n }, {\n \"generation\" : 7,\n \"members\" : [ ]\n } ]\n}","from":"reporter","subject":"Failed streaming may cause duplicate SSTable reference"},{"body":"another problem. why not store data in some system CF? would be probably safer choice.\n\nERROR [CompactionExecutor:5] 2011-10-04 17:13:13,922 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[CompactionExecutor:5,5,main]\njava.io.IOError: java.io.IOException: Failed to rename \\var\\lib\\cassandra\\data\\test\\sipdb.json to \\var\\lib\\cassandra\\data\\test\\sipdb-old.json\n\tat org.apache.cassandra.db.compaction.LeveledManifest.serialize(LeveledManifest.java:382)\n\tat org.apache.cassandra.db.compaction.LeveledManifest.promote(LeveledManifest.java:182)\n\tat org.apache.cassandra.db.compaction.LeveledCompactionStrategy.handleNotification(LeveledCompactionStrategy.java:152)\n\tat org.apache.cassandra.db.DataTracker.notifySSTablesChanged(DataTracker.java:466)\n\tat org.apache.cassandra.db.DataTracker.replace(DataTracker.java:275)\n\tat org.apache.cassandra.db.DataTracker.replaceCompactedSSTables(DataTracker.java:232)\n\tat org.apache.cassandra.db.ColumnFamilyStore.replaceCompactedSSTables(ColumnFamilyStore.java:960)\n\tat org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:199)\n\tat org.apache.cassandra.db.compaction.LeveledCompactionTask.execute(LeveledCompactionTask.java:47)\n\tat org.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:131)\n\tat org.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:114)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\nCaused by: java.io.IOException: Failed to rename \\var\\lib\\cassandra\\data\\test\\sipdb.json to \\var\\lib\\cassandra\\data\\test\\sipdb-old.json\n\tat org.apache.cassandra.io.util.FileUtils.renameWithConfirm(FileUtils.java:64)\n\tat org.apache.cassandra.db.compaction.LeveledManifest.serialize(LeveledManifest.java:375)\n\t... 15 more\n","from":"developer"},{"body":"Because then you get into hairy cyclical situations where you can't read the manifest until you replay the commitlog, but replaying the commitlog requires writing new sstables and thus knowing the manifest","from":"developer"},{"body":"as i understand new flushed tables are placed at level 0. Just replay commitlog and put all new stuff in lvl 0. after comitlog is done, it can do voodoo shuffles.\n\nbut why not to rename tables like table-h-333-l1-Data.db?\n\nidea to have stables with non overlapping key ranges is interesting, but read performance is kinda slow (about 50% of normal) here. Its cassandra core modified to get advantage of leveled tables? i.e. search one sstable at level1, one at lvl2 using bloom filters for key?","from":"developer"},{"body":"This isn't really a great place to rehash http://leveldb.googlecode.com/svn/trunk/doc/impl.html and CASSANDRA-1608.","from":"developer"},{"body":"bq. why not store data in some system CF? would be probably safer choice.\n\nThis has historically been a bad idea, see CASSANDRA-1155, then CASSANDRA-1318 and finally CASSANDRA-1430.","from":"developer"},{"body":"I don't suppose you were using column family truncation in your tests, where you? ","from":"developer"},{"body":"no truncation, no supercolumns.","from":"developer"},{"body":"Are you still able to reproduce reliably? Because we aren't and being able to would help considerably, so if you are and could share whatever script you're using to reproduce, that would be awesome.","from":"developer"},{"body":"i tested it on 1.0 final and it worked without error for 1 test run. i will give it another test without index.","from":"developer"},{"body":"I'll note that Ramesh Natarajan reported on the mailing list what clearly appears to be the same bug (http://www.mail-archive.com/user@cassandra.apache.org/msg18146.html), but while not using leveled compaction. I also think he was using the 1.0.0 final.","from":"developer"},{"body":"I'll note that more info have been added to the messages thrown by the exception here in 1.0.1. So if someone can reproduce this issue on 1.0.1, it would be useful to get the stacktrace (the full system.log would actually be even better).","from":"developer"},{"body":"This AssertionError happened always in cassandra1.0.0 ,not just only in LeveledCompactionStrategy","from":"developer"},{"body":"I suppose it's a bug in DataTracker .","from":"developer"},{"body":"As I already said, if you are able to reproduce this, please try reproducing with 1.0.3. And if you are still able to, please attach you system.log with the exception here because it will have more info on the error that should help. And if you're not able to reproduce with 1.0.3, then I guess it means we've fixed it without knowing.","from":"developer"},{"body":"Running under 1.0.4, can easily reproduce this by just kicking of a repair of any LeveledCompactionStrategy CF. \n\nThe 'zero' on the assert indicates the value (added that to the code to see what the value was): \n\njava.lang.AssertionError: 0\n at org.apache.cassandra.db.compaction.LeveledManifest.promote(LeveledManifest.java:178)\n at org.apache.cassandra.db.compaction.LeveledCompactionStrategy.handleNotification(LeveledCompactionStrategy.java:141)\n at org.apache.cassandra.db.DataTracker.notifySSTablesChanged(DataTracker.java:481)\n at org.apache.cassandra.db.DataTracker.replace(DataTracker.java:275)\n at org.apache.cassandra.db.DataTracker.addSSTables(DataTracker.java:237)\n at org.apache.cassandra.db.DataTracker.addStreamedSSTable(DataTracker.java:242)\n at org.apache.cassandra.db.ColumnFamilyStore.addSSTable(ColumnFamilyStore.java:920)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:141)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:103)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:184)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:81)\n\n\nRelevant lines from system.log leading up the it: \n\n INFO [FlushWriter:794] 2011-12-01 14:23:22,966 Memtable.java (line 275) Completed flushing /var/lib/cassandra/data/sso/Sessions-hc-12524-Data.db (1119784 bytes)\n INFO [CompactionExecutor:2379] 2011-12-01 14:23:22,969 CompactionTask.java (line 112) Compacting [SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12501-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12517-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12513-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12512-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12502-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12507-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12519-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12500-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12508-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12504-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12510-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12515-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12509-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12524-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12514-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12518-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12505-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12516-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12511-Data.db'), SSTableReader(path='/var/lib/cassandra/data/sso/Sessions-hc-12506-Data.db')]\n INFO [AntiEntropyStage:1] 2011-12-01 14:25:06,321 AntiEntropyService.java (line 186) [repair #ea080b70-1c51-11e1-0000-692e0c239dfd] Received merkle tree for Sessions from /xxxxxxxx\nERROR [Thread-177] 2011-12-01 14:25:17,863 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[Thread-177,5,main]\njava.lang.AssertionError: 0 [see above]\n\nIf you want more let me know I can reproduce instantly.\n","from":"developer"},{"body":"Joe, your assertion is the one in CASSANDRA-3536 (where I've attached a patch fixing it). Closing this other one as cantrepro.","from":"developer"},{"body":"This error actually happens on 1.1. And I can easily reproduce with unit test(Test code attached).\n\n{code}\n [junit] ERROR 17:34:46,696 Fatal exception in thread Thread[CompactionExecutor:3,1,main]\n [junit] java.lang.AssertionError: Expecting new size of 2, got 1 while replacing [SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-1-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-5-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-4-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-2-Data.db')] by [SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-6-Data.db')] in View(pending_count=0, sstables=[SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-1-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-2-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-4-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-4-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-5-Data.db')], compacting=[SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-1-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-5-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-4-Data.db'), SSTableReader(path='build/test/cassandra/data/Keyspace1/Standard1/Keyspace1-Standard1-hf-2-Data.db')])\n [junit] \tat org.apache.cassandra.db.DataTracker$View.newSSTables(DataTracker.java:651)\n [junit] \tat org.apache.cassandra.db.DataTracker$View.replace(DataTracker.java:616)\n [junit] \tat org.apache.cassandra.db.DataTracker.replace(DataTracker.java:320)\n [junit] \tat org.apache.cassandra.db.DataTracker.replaceCompactedSSTables(DataTracker.java:253)\n [junit] \tat org.apache.cassandra.db.ColumnFamilyStore.replaceCompactedSSTables(ColumnFamilyStore.java:994)\n [junit] \tat org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:200)\n [junit] \tat org.apache.cassandra.db.compaction.CompactionManager$1.runMayThrow(CompactionManager.java:154)\n [junit] \tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n [junit] \tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:441)\n [junit] \tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n [junit] \tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n [junit] \tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n [junit] \tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n [junit] \tat java.lang.Thread.run(Thread.java:680)\n{code}\n\nThe cause is actually in streaming. StreamInSession can add duplicate reference to SSTable to DataTracker when it is left even after stream session finishes. This typically happens when source node is marked as dead by FailureDetector during streaming session(GC storm is the one I saw) and keep sending file in same session after the node comes back.","from":"developer"},{"body":"Test code attached. Compaction strategy is not related.","from":"developer"},{"body":"Good analysis Yuki. I'm not really sure what is the right fix though. Given that this should very rarely happen (repair uses a much higher failure detection threshold than the normal one, though maybe we can increase it even more to make this even less likely) and that I don't seen any obvious way to avoid that kind of situation, maybe making DataTracker handle duplicate addition of a SSTableReader is the simplest thing to do. The obvious way to do that would be to change the View sstables List to a Set, which leads me to the current commentary in the code:\n{noformat}\n // We can't use a SortedSet here because \"the ordering maintained by a sorted set (whether or not an\n // explicit comparator is provided) must be consistent with equals.\" In particular,\n // ImmutableSortedSet will ignore any objects that compare equally with an existing Set member.\n // Obviously, dropping sstables whose max column timestamp happens to be equal to another's\n // is not acceptable for us. So, we use a List instead.\n{noformat}\nI think that comment is obsolete. Namely, it was added with CASSANDRA-2498 and at the time the list of sstable was kept in max timestamp order at all time. But since then, we've moved the sorting in max timestamp in CollationController directly (which is less fragile), so the order inside DataTracker doesn't matter anymore.","from":"developer"},{"body":"bq. This typically happens when source node is marked as dead by FailureDetector during streaming session(GC storm is the one I saw) and keep sending file in same session after the node comes back\n\nBut we close the session on convict, so shouldn't it start a new one?","from":"developer"},{"body":"bq. But we close the session on convict, so shouldn't it start a new one?\n\nYes, StreamInSession gets closed and removed on convict _once_. But if GC pause happens in the middle of streaming session, the node resumes streaming in the same session after GC. Since resumed stream carries session ID that is once closed on receiver side, StreamInSession is created again with the same old session ID and this time just 1 file to receive.\nThis continues again and again until source node's StreamingOutSession sends all files.\nYou can see this in receiver's log file like below:\n\n{code}\nINFO [Thread-50] 2012-10-20 13:13:26,574 StreamInSession.java (line 214) Finished streaming session 10 from /10.xx.xx.xx\nINFO [Thread-51] 2012-10-20 13:13:29,691 StreamInSession.java (line 214) Finished streaming session 10 from /10.xx.xx.xx\nINFO [Thread-52] 2012-10-20 13:13:32,957 StreamInSession.java (line 214) Finished streaming session 10 from /10.xx.xx.xx\n{code}\n\nDuplication happens during this partially broken streaming session. Because StreamInSession is removed after sending SESSION_FINISHED reply, and StreamOutSession keeps sending files, sometimes the same StreamInSession instance receives more than 1 file and calls closeIfFinished every time it received the file.\n(Sorry, this is hard to explain in words.\nhttps://github.com/apache/cassandra/blob/cassandra-1.1.6/src/java/org/apache/cassandra/streaming/StreamInSession.java#L181 this part is executed multiple times with _readers_ growing by received new file.)\n\nSo as Sylvain stated above, changing DataTracker.View's sstable to Set is one way to eliminate duplicate reference and we should do it. In addition, I'm thinking not to create duplicate StreamInSession by checking StreamHeader.pendingFiles because this field is only filled when initiating streaming.","from":"developer"},{"body":"That code is a mess so let me give a shot at describing what happens for the record. Say node1 wants to stream files A, B and C to node2. If everything goes well what happens is:\n# node1 sends the first file A with a StreamHeader that says that A, B and C are pending files and A is the currently sent file. On node2, a new StreamInSession is created with those information.\n# Once A is finished, node2 remove A from the pending file in the StreamInSession send an acknowledgement to node1, and then node1 sends B with a StreamHeader with no pending files (basically the list of pending files is only sent the first time so that the StreamInSession on node2 knows when everything is finished) and B as current file. When node2 received that StreamHeader, it retrieve the StreamInSession, setting B as the current files.\n# Once B is finished, node2 removes it from pending files, acks to node1 and node1 sends C with a StreamHeader with no pending file and C as current file. Node2 retrieven the StreamInSession and modify it accordingly.\n# At last, once C is finished, node2 removes it from the pending files. Then it realizes the pending files are empty and so that the streaming is finished and at that point it adds all the SSTableReader created so far to the cfs (and acks to node1 the end of the streaming).\n\nNow, the problem is if say node1 is marked dead by mistake by node2 during say the streaming of A. I that happens, the only thing we do on node2 is to close the session and remove the streamInSession from the global sessions map. However we don't shutdown the stream or anything, so if node1 is in fact still alive, what will happen is:\n# A will finish his transfer correctly. Once that's done, node2 will still send an acknowledgement (probably the first mistake, we could check that the session has been closed and send an error instead).\n# Node1 getting it's acknowledgement will send B with a StreamHeader that has B as current file and no pending files as usual. On reception, node2 will not find any StreamInSession (it has been removed during the close), and so it will create a new one as if that was the beginning of a transfer. And that session will have no pending file (second mistake: if we have to create a new StreamInSession but there is no pending file at all something wrong has happened).\n# Once B is fully streamed, node2 will acknowledge it to node1 and remove it from it's streamInSession. But that session the new one we just created with no pending file. So the streamInSession will consider the streaming is finished, and it will thus add the SSTableReader for B to the cfs.\n# Because B has been acknowledged, node1 will start sending C (again, with no pending file in the StreamHeader). This will be done as soon as B was finished, and so concurrently with the streamInSession on node2 closing itself.\n# So when node2 receives the StreamHeader with C, it will try to retrieve the session and will find the previous session. And will happily add C as the current file for that session (third and fourth mistake: StreamInSession should not add a file as current unless it is a pending file for this session, and a session could detect that it's being reused even though it has just detected itself as finished).\n# Now when C transfer finishes, the seesion will be notify and since it still has no pending files, it will once again consider the streaming as complete. But since it's still the same session, it still has the SSTableReader for B in its list of created reader (as well as the one for C now). And that's when it adds B for a second time to the DataTracker.\n\nI also not that we end up without having ever add the SSTableReader for A to the cfs since the very first StreamInSession was never finished. This is not a big deal in that the stream itself has been indicated as failed to the client anyway, but just to say that it's not just a problem of duplicating a SSTableReader preference.\n\nAnyway, let me back on what I said earlier. We should definitively fix some if not all of the \"mistake\" above (and send a SESSION_FAILURE to node1 as soon as we detect something is wrong).\n\nBut that being said, my comment on the comment in DataTracker being obsolete still stand, and replacing the list by a set in there would have at least the advantage of slightly simplifying the code of DataTracker.View.newSSTables(), as well as being more resilient if a SSTableReader is added twice. Not a big deal though.\n","from":"developer"},{"body":"Attaching first attempt.\nI changed DataTracker.View's sstables to Set, and made stream fail when file arrives after StreamInSession failed.\n\nChanging List to Set for sstables sometimes makes CollationControllerTest fail. It was introduced in CASSANDRA-4116, and I think the test and CollationController#collectAllData expect sstables to be ordered by timestamp. I'm not sure if the test is obsolete or we really need sstables to be sorted all the time.\n0002 patch alone will fix the issue, so we can apply that for now.","from":"developer"},{"body":"For patch 0002, we shouldn't check the FailureDetector otherwise we don't really fix the issue. The only way we know this bug can happen is wher the FailureDetector *had* marked a node down while it shouldn't have (besides, we just got something from a node so it's fair to assume it is alive).\n\nbq. and I think the test and CollationController#collectAllData expect sstables to be ordered by timestamp\n\nIt doesn't seem to me that collectAllData needs sstable ordered. In fact, I think that it does a second pass over the sstables iterators just because it doesn't assume sstables are ordered by max timestamp. Moreover, I'm pretty sure it would be a bug to assume that. If you look at DataTracker.View.newSSTables, it ends by {{Iterables.addAll(newSSTables, replacements)}} which clearly won't maintain any specific ordering of sstables.\n\nbq. I'm not sure if the test is obsolete.\n\nI don't think the test is obsolete but I think we have a minor bug in CollationController. The test want to test that we correctly exclude sstable whose maxTimestamp is less than the most recent row tombstone we have. But that test checks controller.getSstablesIterated(), and for collectAllData, it will count every sstable it include in the first iteration of collectAllData but don't remove those that are remove by the second pass. In other words, I think the correct fix is to decrement stablesIterated in CollationController when in the second pass we remove a sstable (or more simply to set it to iterators.size() just before we collate everything).","from":"developer"},{"body":"bq. we shouldn't check the FailureDetector otherwise we don't really fix the issue.\n\nOk. I've fixed this and reattached 0002.\n\nbq. The test want to test that we correctly exclude sstable whose maxTimestamp is less than the most recent row tombstone we have.\n\nRight. But I think the test assumes that SSTables are added to List in order of flush, and that's true as long as we use List. So what I suggest is to remove that part from the test since we no longer use List.\nAnd sstablesIterated counter in collectAllData is doing fine because we actually read the data from sstable when we go over\n{code}\nIColumnIterator iter = filter.getSSTableColumnIterator(sstable);\n{code}\nbefore incrementing counter.\n\nSo I removed that test from CollationControllerTest in 0001-change-DataTracker.View-s-sstables-from-List-to-Set.patch.","from":"developer"},{"body":"bq. Ok. I've fixed this and reattached 0002.\n\nAlright, +1 on 0002. Let's commit that for now to 1.1/1.2 as this fix this ticket.\n\n{quote}\nthe test assumes that SSTables are added to List in order of flush\nAnd sstablesIterated counter in collectAllData is doing fine because we actually read the data\n{quote}\n\nRight. I guess what I meant is that what is tested right now is not really sensible. Relying on the order of flush is only valid for a small, controlled test, but in reality as soon as compaction kicks in, the order of sstable in DataTracker will be meaningless even with a List instead of a Set. Basically the guarantee collectAll gives us today is that it will eliminate sstables whose maxTimestamp < mostRecentTombstone with just having read the sstable row header, not the full data. But that's not what sstablesIterated counts so it's broken.\n\nThat being said, I think we can improve collectAll in the way described in CASSANDRA-4883. If we do so, the test will pass again without relying on any assumption of the order of sstables in DataTracker. So overall I suggest moving all of this to CASSANDRA-4883.","from":"developer"},{"body":"Committed 0002 to 1.1 and trunk.","from":"developer"},{"body":"I understand that this has been fixed in newer versions of Cassandra.\n\nBut I'm currently seeing this exact issue on a production 1.1.1 node in my cluster.\n\nWhat should be my next step?\n\nDo I simply restart it?\n\nRun cleanup? Scrub? Repair?\n\nSounds like repair would just fail with the same problem.\n\nAny advice would be appreciated.","from":"developer"},{"body":"Yes, restarting the node will help.\nNo need to clean up/scrub.\n\nPlease use user@cassandra.apache.org mailing list for these type of questions.","from":"developer"},{"body":"Thanks for the swift reply and I will use the mailing list in the future.","from":"developer"}],"created":"2011-10-04T14:39:28.000+0000","description":"during stress testing, i always get this error making leveledcompaction strategy unusable. Should be easy to reproduce - just write fast.\n\nERROR [CompactionExecutor:6] 2011-10-04 15:48:52,179 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[CompactionExecutor:6,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.db.DataTracker$View.newSSTables(DataTracker.java:580)\n\tat org.apache.cassandra.db.DataTracker$View.replace(DataTracker.java:546)\n\tat org.apache.cassandra.db.DataTracker.replace(DataTracker.java:268)\n\tat org.apache.cassandra.db.DataTracker.replaceCompactedSSTables(DataTracker.java:232)\n\tat org.apache.cassandra.db.ColumnFamilyStore.replaceCompactedSSTables(ColumnFamilyStore.java:960)\n\tat org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:199)\n\tat org.apache.cassandra.db.compaction.LeveledCompactionTask.execute(LeveledCompactionTask.java:47)\n\tat org.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:131)\n\tat org.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:114)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\n\nand this is in json data for table:\n\n{\n \"generations\" : [ {\n \"generation\" : 0,\n \"members\" : [ 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484 ]\n }, {\n \"generation\" : 1,\n \"members\" : [ ]\n }, {\n \"generation\" : 2,\n \"members\" : [ ]\n }, {\n \"generation\" : 3,\n \"members\" : [ ]\n }, {\n \"generation\" : 4,\n \"members\" : [ ]\n }, {\n \"generation\" : 5,\n \"members\" : [ ]\n }, {\n \"generation\" : 6,\n \"members\" : [ ]\n }, {\n \"generation\" : 7,\n \"members\" : [ ]\n } ]\n}","issue_id":"12525677","key":"CASSANDRA-3306","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2012-10-31T16:13:45.000+0000","role":"fixed_distractor","summary":"Failed streaming may cause duplicate SSTable reference"} {"case_id":"12527792","cluster":"DISTRACTOR-CASSANDRA-3385","comments":[{"body":"enabled assertion\n\n\n\n INFO [HintedHandoff:1] 2011-10-19 13:44:08,346 HintedHandOffManager.java (line 263) Started hinted handoff for token: 0 with IP: \n/10.196.37.187\nERROR [HintedHandoff:1] 2011-10-19 13:44:08,513 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[HintedHan\ndoff:1,1,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:285)\n at org.apache.cassandra.db.HintedHandOffManager.access$100(HintedHandOffManager.java:81)\n at org.apache.cassandra.db.HintedHandOffManager$2.runMayThrow(HintedHandOffManager.java:337)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n\n\n\nline 285 is:\n assert versionColumn != null;\nwell the other following fields could also be null too,\n assert versionColumn != null;\n assert tableColumn != null;\n assert keyColumn != null;\n assert mutationColumn != null;\n\n","created":"2011-10-19T17:49:06.597+0000"},{"body":"I just found that this contributes to another symptom I'm seeing: for RF=3, and a ring of 3 nodes, if I bring down 1 box, the remaining 2 still work fine for Quorum access, but the latency is 20x high.\n\nI can see from debugging that a lot of time is spent on storing hints into local system table on the coordinator. but this Table.apply is slow because a lot of time is spent on the lock, while it really should not happen since the lock is sharded into 4096 ones. it turns out that all the keys used in the hints writing are the same key, at least in the examples I looked at in the debugger, if I'm correct in this observation, this is a serious bug","created":"2011-10-19T19:48:10.989+0000"},{"body":"I see, the key in hinted table is the ID of the dead box. given that this leads to lock contention, would it be better to change the storage layout of hints? ---- I have never tried hinted handoff before, not sure if lock contention was a problem before\n","created":"2011-10-19T19:57:32.386+0000"},{"body":"So either new-style hints are being written without versionColumn by RowMutation.hintFor, or old style hints did not get cleaned out properly by SystemTable.purgeIncompatibleHints. But both of those look fine to me.","created":"2011-10-19T19:57:49.952+0000"},{"body":"\nthe hints code was from:\n\nhttps://github.com/apache/cassandra/commit/3893f24098c3d82dc31571f0b6841e2d5821ea74#diff-12\n\n#CASSANDRA-2034\n\nmaybe it should be better to NOT let the writer wait for hints to finish? right now the local hints write make the entire write slower in probably 2 ways : 1) the main write has to wait for hint write to finish, which is slow due to lock 2) hints writes are slow, which create a lot of jobs on MUTATION stage, so even if main write does not wait for them, the MUTATION stage could possibly be bogged down with hints writes, and not able to handle normal writes fast enough\n\n\n","created":"2011-10-19T20:03:40.890+0000"},{"body":"IMO the right fix to the lock contention is to simply not synchronize when there are no indexes present. But that's unrelated to the assertion failure here.","created":"2011-10-19T20:07:08.903+0000"},{"body":"that works for me too. but I guess you will finally have to handle cases where indexes are needed. in those cases, probably we need to change the hints format away from using IP as key","created":"2011-10-19T20:18:26.643+0000"},{"body":"If you're not going to use IPs as keys how are you going to replay hints efficiently? You need to consider the read path as well as the write when modeling something. :)","created":"2011-10-19T20:23:04.637+0000"},{"body":"Created CASSANDRA-3386 for the contention problem.","created":"2011-10-19T20:32:56.126+0000"},{"body":"I am using the RPM provided by datastax, I upgraded from 0.8.7 to 1.0.0, then performed a scrub and repair on each node as described in the documentation.\nThis worked fine for all nodes except one, which now throws exceptions whenever I start it up, and goes \"Down\" occassionly. All other nodes have been fine.\n\nThis is from the log, I think it could be related, but I'm not sure:\n\nINFO [HintedHandoff:4] 2011-10-25 15:30:08,867 HintedHandOffManager.java (line 263) Started hinted handoff for token: 0 with IP: /xx.xx.xx.xx\n INFO [HintedHandoff:4] 2011-10-25 15:30:08,868 HintedHandOffManager.java (line 318) Finished hinted handoff of 0 rows to endpoint /xx.xx.xx.xx\n INFO [HintedHandoff:4] 2011-10-25 15:30:09,998 HintedHandOffManager.java (line 263) Started hinted handoff for token: 148873535527910577765226390751398592512 with IP: /xx.xx.xx.xx\nERROR [HintedHandoff:4] 2011-10-25 15:30:09,999 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[HintedHandoff:4,1,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:285)\n at org.apache.cassandra.db.HintedHandOffManager.access$100(HintedHandOffManager.java:81)\n at org.apache.cassandra.db.HintedHandOffManager$2.runMayThrow(HintedHandOffManager.java:337)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\n INFO [HintedHandoff:5] 2011-10-25 15:30:59,893 HintedHandOffManager.java (line 263) Started hinted handoff for token: 106338239662793269832304564822427566080 with IP: /xx.xx.xx.xx\nERROR [HintedHandoff:5] 2011-10-25 15:30:59,894 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[HintedHandoff:5,1,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:285)\n at org.apache.cassandra.db.HintedHandOffManager.access$100(HintedHandOffManager.java:81)\n at org.apache.cassandra.db.HintedHandOffManager$2.runMayThrow(HintedHandOffManager.java:337)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\n INFO [HintedHandoff:6] 2011-10-25 15:31:41,194 HintedHandOffManager.java (line 263) Started hinted handoff for token: 42535295865117307932921825928971026432 with IP: /xx.xx.xx\n INFO [HintedHandoff:6] 2011-10-25 15:31:41,194 HintedHandOffManager.java (line 318) Finished hinted handoff of 0 rows to endpoint /xx.xx.xx.xx","created":"2011-10-25T15:21:44.298+0000"},{"body":"As Jonathan said, it looks like either old hints that haven't been cleaned or some corruption to the hints that are stored.\n\nWhile it's probably important to find the source of this problem, I think it's really bad that we use assertions to check for it. Hinted Handoff is an optimization, so for invalid hints shouldn't break things.\n\nThe attached patch removes the assertions and replaces them with a check to ensure the hint is valid. If it's not, a warning will be emitted (so you know something's not quite right) and everything will continue as normal.","created":"2011-10-26T12:58:52.347+0000"},{"body":"I'd say that reasoning makes this the perfect case for assertions -- it doesn't affect anything but hints for the assert to fail.","created":"2011-10-26T13:48:24.331+0000"},{"body":"Wouldn't this disrupt the delivery of hints that aren't corrupt?","created":"2011-10-26T13:52:06.366+0000"},{"body":"One possible avenue for this is that if startup takes long enough (due to CL replay + sstable index sampling, probably) that compactions start before the upgrade hint purge, compaction can generation \"new\" 0.8 hints after we try to delete them. Patch to switch to truncate to avoid this, although CASSANDRA-3399 is open to make truncate more bulletproof in that situation itself. In the meantime, removing the hints columnfamily manually before restarting should fix the problem.\n\nAlso added some debug logging to the hint purge in r1189221.","created":"2011-10-26T14:08:15.505+0000"},{"body":"bq. Wouldn't this disrupt the delivery of hints that aren't corrupt?\n\nSure. Point is, that's not a Big Deal. So it's worth having the big neon sign of an exception telling people \"this ain't right.\"","created":"2011-10-26T14:10:58.673+0000"},{"body":"new version of patch adds a second check for 0.8 hints during hint delivery","created":"2011-10-26T14:31:54.351+0000"},{"body":"I guess I can live with ignoring the known types of corruption (old hints) and leaving the assertions to flag up unknown forms of corruption.","created":"2011-10-26T15:31:58.680+0000"},{"body":"+1","created":"2011-10-26T21:50:56.094+0000"},{"body":"committed","created":"2011-10-27T16:43:20.599+0000"}],"conversations":[{"body":"I'm using the current HEAD of 1.0.0 github branch, and I'm still seeing this error, not sure if it's this bug or another one.\n\n\n\n INFO [HintedHandoff:1] 2011-10-19 12:43:17,674 HintedHandOffManager.java (line 263) Started hinted handoff for token: 11342745564\n0312821154458202477256070484 with IP: /10.39.85.140\nERROR [HintedHandoff:1] 2011-10-19 12:43:17,885 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[HintedHan\ndoff:1,1,main]\njava.lang.RuntimeException: java.lang.NullPointerException\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:34)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:289)\n at org.apache.cassandra.db.HintedHandOffManager.access$100(HintedHandOffManager.java:81)\n at org.apache.cassandra.db.HintedHandOffManager$2.runMayThrow(HintedHandOffManager.java:337)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n ... 3 more\nERROR [HintedHandoff:1] 2011-10-19 12:43:17,886 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.RuntimeException: java.lang.NullPointerException\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:34)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:289)\n at org.apache.cassandra.db.HintedHandOffManager.access$100(HintedHandOffManager.java:81)\n at org.apache.cassandra.db.HintedHandOffManager$2.runMayThrow(HintedHandOffManager.java:337)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n ... 3 more\n\n\nthis could possibly be related to #3291\n","from":"reporter","subject":"NPE in hinted handoff"},{"body":"enabled assertion\n\n\n\n INFO [HintedHandoff:1] 2011-10-19 13:44:08,346 HintedHandOffManager.java (line 263) Started hinted handoff for token: 0 with IP: \n/10.196.37.187\nERROR [HintedHandoff:1] 2011-10-19 13:44:08,513 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[HintedHan\ndoff:1,1,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:285)\n at org.apache.cassandra.db.HintedHandOffManager.access$100(HintedHandOffManager.java:81)\n at org.apache.cassandra.db.HintedHandOffManager$2.runMayThrow(HintedHandOffManager.java:337)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n\n\n\nline 285 is:\n assert versionColumn != null;\nwell the other following fields could also be null too,\n assert versionColumn != null;\n assert tableColumn != null;\n assert keyColumn != null;\n assert mutationColumn != null;\n\n","from":"developer"},{"body":"I just found that this contributes to another symptom I'm seeing: for RF=3, and a ring of 3 nodes, if I bring down 1 box, the remaining 2 still work fine for Quorum access, but the latency is 20x high.\n\nI can see from debugging that a lot of time is spent on storing hints into local system table on the coordinator. but this Table.apply is slow because a lot of time is spent on the lock, while it really should not happen since the lock is sharded into 4096 ones. it turns out that all the keys used in the hints writing are the same key, at least in the examples I looked at in the debugger, if I'm correct in this observation, this is a serious bug","from":"developer"},{"body":"I see, the key in hinted table is the ID of the dead box. given that this leads to lock contention, would it be better to change the storage layout of hints? ---- I have never tried hinted handoff before, not sure if lock contention was a problem before\n","from":"developer"},{"body":"So either new-style hints are being written without versionColumn by RowMutation.hintFor, or old style hints did not get cleaned out properly by SystemTable.purgeIncompatibleHints. But both of those look fine to me.","from":"developer"},{"body":"\nthe hints code was from:\n\nhttps://github.com/apache/cassandra/commit/3893f24098c3d82dc31571f0b6841e2d5821ea74#diff-12\n\n#CASSANDRA-2034\n\nmaybe it should be better to NOT let the writer wait for hints to finish? right now the local hints write make the entire write slower in probably 2 ways : 1) the main write has to wait for hint write to finish, which is slow due to lock 2) hints writes are slow, which create a lot of jobs on MUTATION stage, so even if main write does not wait for them, the MUTATION stage could possibly be bogged down with hints writes, and not able to handle normal writes fast enough\n\n\n","from":"developer"},{"body":"IMO the right fix to the lock contention is to simply not synchronize when there are no indexes present. But that's unrelated to the assertion failure here.","from":"developer"},{"body":"that works for me too. but I guess you will finally have to handle cases where indexes are needed. in those cases, probably we need to change the hints format away from using IP as key","from":"developer"},{"body":"If you're not going to use IPs as keys how are you going to replay hints efficiently? You need to consider the read path as well as the write when modeling something. :)","from":"developer"},{"body":"Created CASSANDRA-3386 for the contention problem.","from":"developer"},{"body":"I am using the RPM provided by datastax, I upgraded from 0.8.7 to 1.0.0, then performed a scrub and repair on each node as described in the documentation.\nThis worked fine for all nodes except one, which now throws exceptions whenever I start it up, and goes \"Down\" occassionly. All other nodes have been fine.\n\nThis is from the log, I think it could be related, but I'm not sure:\n\nINFO [HintedHandoff:4] 2011-10-25 15:30:08,867 HintedHandOffManager.java (line 263) Started hinted handoff for token: 0 with IP: /xx.xx.xx.xx\n INFO [HintedHandoff:4] 2011-10-25 15:30:08,868 HintedHandOffManager.java (line 318) Finished hinted handoff of 0 rows to endpoint /xx.xx.xx.xx\n INFO [HintedHandoff:4] 2011-10-25 15:30:09,998 HintedHandOffManager.java (line 263) Started hinted handoff for token: 148873535527910577765226390751398592512 with IP: /xx.xx.xx.xx\nERROR [HintedHandoff:4] 2011-10-25 15:30:09,999 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[HintedHandoff:4,1,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:285)\n at org.apache.cassandra.db.HintedHandOffManager.access$100(HintedHandOffManager.java:81)\n at org.apache.cassandra.db.HintedHandOffManager$2.runMayThrow(HintedHandOffManager.java:337)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\n INFO [HintedHandoff:5] 2011-10-25 15:30:59,893 HintedHandOffManager.java (line 263) Started hinted handoff for token: 106338239662793269832304564822427566080 with IP: /xx.xx.xx.xx\nERROR [HintedHandoff:5] 2011-10-25 15:30:59,894 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[HintedHandoff:5,1,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:285)\n at org.apache.cassandra.db.HintedHandOffManager.access$100(HintedHandOffManager.java:81)\n at org.apache.cassandra.db.HintedHandOffManager$2.runMayThrow(HintedHandOffManager.java:337)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\n INFO [HintedHandoff:6] 2011-10-25 15:31:41,194 HintedHandOffManager.java (line 263) Started hinted handoff for token: 42535295865117307932921825928971026432 with IP: /xx.xx.xx\n INFO [HintedHandoff:6] 2011-10-25 15:31:41,194 HintedHandOffManager.java (line 318) Finished hinted handoff of 0 rows to endpoint /xx.xx.xx.xx","from":"developer"},{"body":"As Jonathan said, it looks like either old hints that haven't been cleaned or some corruption to the hints that are stored.\n\nWhile it's probably important to find the source of this problem, I think it's really bad that we use assertions to check for it. Hinted Handoff is an optimization, so for invalid hints shouldn't break things.\n\nThe attached patch removes the assertions and replaces them with a check to ensure the hint is valid. If it's not, a warning will be emitted (so you know something's not quite right) and everything will continue as normal.","from":"developer"},{"body":"I'd say that reasoning makes this the perfect case for assertions -- it doesn't affect anything but hints for the assert to fail.","from":"developer"},{"body":"Wouldn't this disrupt the delivery of hints that aren't corrupt?","from":"developer"},{"body":"One possible avenue for this is that if startup takes long enough (due to CL replay + sstable index sampling, probably) that compactions start before the upgrade hint purge, compaction can generation \"new\" 0.8 hints after we try to delete them. Patch to switch to truncate to avoid this, although CASSANDRA-3399 is open to make truncate more bulletproof in that situation itself. In the meantime, removing the hints columnfamily manually before restarting should fix the problem.\n\nAlso added some debug logging to the hint purge in r1189221.","from":"developer"},{"body":"bq. Wouldn't this disrupt the delivery of hints that aren't corrupt?\n\nSure. Point is, that's not a Big Deal. So it's worth having the big neon sign of an exception telling people \"this ain't right.\"","from":"developer"},{"body":"new version of patch adds a second check for 0.8 hints during hint delivery","from":"developer"},{"body":"I guess I can live with ignoring the known types of corruption (old hints) and leaving the assertions to flag up unknown forms of corruption.","from":"developer"},{"body":"+1","from":"developer"},{"body":"committed","from":"developer"}],"created":"2011-10-19T17:44:53.000+0000","description":"I'm using the current HEAD of 1.0.0 github branch, and I'm still seeing this error, not sure if it's this bug or another one.\n\n\n\n INFO [HintedHandoff:1] 2011-10-19 12:43:17,674 HintedHandOffManager.java (line 263) Started hinted handoff for token: 11342745564\n0312821154458202477256070484 with IP: /10.39.85.140\nERROR [HintedHandoff:1] 2011-10-19 12:43:17,885 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[HintedHan\ndoff:1,1,main]\njava.lang.RuntimeException: java.lang.NullPointerException\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:34)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:289)\n at org.apache.cassandra.db.HintedHandOffManager.access$100(HintedHandOffManager.java:81)\n at org.apache.cassandra.db.HintedHandOffManager$2.runMayThrow(HintedHandOffManager.java:337)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n ... 3 more\nERROR [HintedHandoff:1] 2011-10-19 12:43:17,886 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.RuntimeException: java.lang.NullPointerException\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:34)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:289)\n at org.apache.cassandra.db.HintedHandOffManager.access$100(HintedHandOffManager.java:81)\n at org.apache.cassandra.db.HintedHandOffManager$2.runMayThrow(HintedHandOffManager.java:337)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n ... 3 more\n\n\nthis could possibly be related to #3291\n","issue_id":"12527792","key":"CASSANDRA-3385","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-10-27T16:43:20.000+0000","role":"fixed_distractor","summary":"NPE in hinted handoff"} {"case_id":"12529483","cluster":"DISTRACTOR-CASSANDRA-3427","comments":[{"body":"This is unfortunately a showstopper for our hadoop jobs querying our production cluster.\n\nWith 1.0.1 is there any workaround for this issue?\nIs it correct that this \"compressed block offsets\" totals to \n ( / ) * 8bytes\n\nTherefore a change to a higher chunk_length should be an intermediate workaround?","created":"2011-10-31T15:01:48.338+0000"},{"body":"Quite honestly, the best workaround is likely to apply the attached patch on top of 1.0.1 (you can wait for someone to review to get a bit more confidence).\n\nBecause yes, a bigger chunk_length would diminish the problem, but if you do enough range_queries you would likely still OOM and there is point after which a chunk_length too big is just counter-productive. ","created":"2011-10-31T15:20:51.558+0000"},{"body":"Committed.","created":"2011-10-31T15:44:26.696+0000"},{"body":"Rolled out into production. Works a charm! Even on 200Gb sstables.\n\nSincere appreciations on this one.","created":"2011-10-31T20:10:58.313+0000"},{"body":"Won't the cache here leak?\nMany (most?) sstables are transient (gone after the next minor compaction), but this cache will just grow...","created":"2011-11-02T21:25:22.276+0000"},{"body":"Oh yes, it seems like I missed that one - we should remove entry from the cache when SSTable gets compacted out. What do you think, Sylvain?","created":"2011-11-02T21:30:59.094+0000"},{"body":"Or use a ConcurrentLinkedHashMap w/ fixed capacity?","created":"2011-11-02T21:38:07.695+0000"},{"body":"yes but that will also imply that we should weight it in memory size instead of number of entries so need to use jamm which is calculation overhead, better just remove unused because we know precisely when to do that...","created":"2011-11-02T21:43:23.541+0000"},{"body":"patch according to Pavel's suggestion","created":"2011-11-03T04:55:37.078+0000"},{"body":"Ok, I think using an object cache is ugly (I know that it was my idea).\n\nAt first, I tried going with the natural idea, to add the compressionMetadata as a final field of SSTableReader and use that everywhere, ensuring we use one per sstable. Turned out that for SSTableWriter you need to have the metadata existing before the SSTableReader is created and that seemed like a bit of a mess so I backtracked and decided to go with an object cache in CompressionMetada, but more out of laziness than anything else.\n\nThat was wrong of me to be lazy. We don't need that object cache and if its use is going to leak out of CompressionMetada (like hard coding the addition of the notifier in the DataTracker constructor; which defeats the purpose of the notifier abstraction in the first place) then it's not even clean as far as code is concerned.\n\nAttaching a v2 patch that remove the cache and do a slight modification of the initial idea, that is it just let CompressedSegmentedFile create the CompressionMetada and have the rest of the code use that. Turns out that once I plug my brain, it's only a few lines of code.\n","created":"2011-11-03T10:47:52.325+0000"},{"body":"+1","created":"2011-11-03T14:09:39.321+0000"},{"body":"Alright, committed this new version, thanks","created":"2011-11-03T14:17:52.256+0000"},{"body":"Handling jvm memory since upgrading to cassandra-1.0 and enabling compression is still a headache.\nWhere i used to be able to run w/ Xmx8G i'm now struggling to run with Xmx20G (all caches are disabled) and during startup can hit\n{noformat}java.lang.OutOfMemoryError: Java heap space\n at org.apache.cassandra.utils.BigLongArray.(BigLongArray.java:53)\n at org.apache.cassandra.utils.BigLongArray.(BigLongArray.java:39)\n at org.apache.cassandra.io.compress.CompressionMetadata.readChunkOffsets(CompressionMetadata.java:122){noformat}\n\nI could keep increasing chunk_length (it's already at 256) but this seems awkward just to get a cluster running smoothly. At minimum the calculations for memory requirements for cassandra should be re-written if compression is to take such a large chunk of heap.","created":"2011-11-13T11:39:51.028+0000"},{"body":"Does heap usage stay high-post startup? Can you try forcing a full GC to check that?","created":"2011-11-13T13:33:31.113+0000"},{"body":"bq. Does heap usage stay high-post startup? Can you try forcing a full GC to check that?\nYes. GC doesn't seem to help, i'll attach a munin graph that shows it over time. It was running for a number of days just under 20G, but you can see from that how \"squeezed\" it was.\n\n(invoking full gc via jmx has no noticeable effect on heap used)","created":"2011-11-13T16:45:36.610+0000"},{"body":"0.8 was running on Xmx8G up until week 44.\nat that point we upgraded to 1.0 and enabled compression. The very high memory usage at the beginning of week 44 was to handle the change from chunk_length 16 to 256.\nThen for most of week 44 and week 45 we ran with Xmx16G, but this was very \"squeezed\". Now that's OOM, and raising it to 20G didn't help. Currently it's on 30G.\n\nAlso note we're always used -XX:CMSInitiatingOccupancyFraction=60 for this cluster.\n(full java opts are \"-Xss128k -XX:+UseThreadPriorities -XX:ThreadPriorityPolicy=42 -XX:SurvivorRatio=8 -XX:MaxTenuringThreshold=1 -XX:+UseParNewGC -XX:+UseConcMarkSweepGC -XX:+CMSParallelRemarkEnabled -XX:CMSInitiatingOccupancyFraction=60 -XX:+UseCMSInitiatingOccupancyOnly -Xmx30g -Xmx30g -Xmn800M -XX:ParallelCMSThreads=4 -XX:+CMSIncrementalMode -XX:+CMSIncrementalPacing -XX:CMSIncrementalDutyCycleMin=0 -XX:CMSIncrementalDutyCycle=10\". the last 5 were added during week 44 to try and help, ref http://blog.mikiobraun.de/2010/08/cassandra-gc-tuning.html)","created":"2011-11-13T16:55:09.148+0000"},{"body":"What version exactly are you running now? 1.0.2? 1.0.0 + 3427? Something else?","created":"2011-11-14T14:12:49.957+0000"},{"body":"I've attached a tiny patch (0001-debugging.patch) that prints in the log the size of the long array we allocate for the chunk offsets. Would you mind trying with this and attach a log of when startup hits one of the OOM you pasted earlier (feel free to use a 8GB heap if it's easier to reproduce). I'd like to know if those offsets are indeed the problem.","created":"2011-11-14T14:25:36.139+0000"},{"body":"I can get you a place to upload the heap dump from an OOM too. (8GB would be best, since heap analysis requires ram proportional to the heap size.)\n\nTo rule out the obvious, have you tried running 1.0.x w/o compression?","created":"2011-11-14T14:30:20.575+0000"},{"body":"version: 1.0.2 snapshot (pretty close to release date) + CASSANDRA-3197.\nw/o compression: that would require a full compact/scrub. that takes close to 24hrs :-(\npatch: i attach that and hopefully have output soon. a heap dump can be done at the same time...","created":"2011-11-14T18:14:33.191+0000"},{"body":"startup log with debug (off 1.0.2 release){noformat}INFO 48:34,688 DatabaseDescriptor: Loading settings from file:/iad/finn/countstatistics/conf/cassandra-prod.yaml\nINFO 48:34,782 DatabaseDescriptor: DiskAccessMode 'auto' determined to be mmap, indexAccessMode is mmap\nINFO 48:34,792 DatabaseDescriptor: Global memtable threshold is enabled at 512MB\nINFO 48:34,890 AbstractCassandraDaemon: JVM vendor/version: Java HotSpot(TM) 64-Bit Server VM/1.6.0_24\nINFO 48:34,891 AbstractCassandraDaemon: Heap size: 760414208/8506048512\nINFO 48:34,891 AbstractCassandraDaemon: Classpath: /iad/finn/countstatistics/jar/countstatistics.jar:/iad/common/apps/cassandra/lib/jamm-0.2.5.jar\nINFO 48:37,158 CLibrary: JNA mlockall successful\nINFO 48:37,879 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/Versions-h-42 (256 bytes)\nINFO 48:37,879 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/Versions-h-41 (256 bytes)\nINFO 48:37,879 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/Versions-h-40 (256 bytes)\nINFO 48:37,959 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/IndexInfo-h-3 (223 bytes)\nINFO 48:38,001 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/Schema-h-15 (34257 bytes)\nINFO 48:38,045 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/Migrations-h-15 (78524 bytes)\nINFO 48:38,096 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/LocationInfo-h-150 (80 bytes)\nINFO 48:38,096 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/LocationInfo-h-149 (628 bytes)\nINFO 48:38,096 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/LocationInfo-h-151 (163 bytes)\nINFO 48:38,192 DatabaseDescriptor: Loading schema version 1940c630-0be4-11e1-0000-d1695892b1ff\nINFO 51:35,136 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191473 (38646535 bytes)\nINFO 51:35,136 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190467 (2284524668 bytes)\nINFO 51:35,136 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191469 (254927460 bytes)\nINFO 51:35,136 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191475 (30477008 bytes)\nINFO 51:35,136 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-114136 (156044360682 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191294 (4585008988 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190415 (15857295280 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-183183 (196289440978 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191472 (1346076 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190736 (4626053255 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191435 (1188223188 bytes)\nINFO 51:35,187 CompressionMetadata: Allocating chunks index for 5745 chunks for uncompressed size of 1470519 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191472-CompressionInfo.db)\nINFO 51:35,421 CompressionMetadata: Allocating chunks index for 129646 chunks for uncompressed size of 33189311 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191475-CompressionInfo.db)\nINFO 51:35,544 CompressionMetadata: Allocating chunks index for 165602 chunks for uncompressed size of 42393918 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191473-CompressionInfo.db)\nINFO 51:37,171 CompressionMetadata: Allocating chunks index for 1091377 chunks for uncompressed size of 279392485 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191469-CompressionInfo.db)\nINFO 51:41,148 CompressionMetadata: Allocating chunks index for 5086138 chunks for uncompressed size of 1302051278 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191435-CompressionInfo.db)\nINFO 51:46,351 CompressionMetadata: Allocating chunks index for 9766541 chunks for uncompressed size of 2500234376 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190467-CompressionInfo.db)\nINFO 51:56,717 CompressionMetadata: Allocating chunks index for 19828434 chunks for uncompressed size of 5076078986 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190736-CompressionInfo.db)\nINFO 51:56,897 CompressionMetadata: Allocating chunks index for 19626358 chunks for uncompressed size of 5024347477 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191294-CompressionInfo.db)\nINFO 52:21,670 CompressionMetadata: Allocating chunks index for 67865822 chunks for uncompressed size of 17373650297 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190415-CompressionInfo.db)\nINFO 55:55,920 CompressionMetadata: Allocating chunks index for 666981588 chunks for uncompressed size of 170747286320 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-114136-CompressionInfo.db)\nINFO 56:49,620 CompressionMetadata: Allocating chunks index for 840404671 chunks for uncompressed size of 215143595584 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-183183-CompressionInfo.db)\nERROR 57:51,112 AbstractCassandraDaemon: Fatal exception in thread Thread[SSTableBatchOpen:8,5,main]\njava.lang.OutOfMemoryError: Java heap space\n\tat org.apache.cassandra.utils.BigLongArray.(BigLongArray.java:53)\n\tat org.apache.cassandra.utils.BigLongArray.(BigLongArray.java:39)\n\tat org.apache.cassandra.io.compress.CompressionMetadata.readChunkOffsets(CompressionMetadata.java:127)\n ...{noformat}","created":"2011-11-14T20:06:26.396+0000"},{"body":"That's the stupidest bug ever. It happens we interpret the chunk_length_in_kb not in kb but in bytes.\nAnyway, I've created CASSANDRA-3492 to address this.\nTurns out if you don't update the chunk_length you're fine because the default is ok, but I guess hitting this issue initially has put you in the wrong spot :(","created":"2011-11-14T21:01:45.847+0000"}],"conversations":[{"body":"The CompressionMetada holds the compressed block offsets in memory. Without being absolutely huge, this is still of non-negligible size as soon as you have a bit of data in the DB. Reallocating this for each read is a very bad idea.\n\nNote that this only affect range queries, since \"normal\" queries uses CompressedSegmentedFile that does reuse a unique CompressionMetadata instance.\n\n( Background: http://thread.gmane.org/gmane.comp.db.cassandra.user/21362 )","from":"reporter","subject":"CompressionMetadata is not shared across threads, we create a new one for each read"},{"body":"This is unfortunately a showstopper for our hadoop jobs querying our production cluster.\n\nWith 1.0.1 is there any workaround for this issue?\nIs it correct that this \"compressed block offsets\" totals to \n ( / ) * 8bytes\n\nTherefore a change to a higher chunk_length should be an intermediate workaround?","from":"developer"},{"body":"Quite honestly, the best workaround is likely to apply the attached patch on top of 1.0.1 (you can wait for someone to review to get a bit more confidence).\n\nBecause yes, a bigger chunk_length would diminish the problem, but if you do enough range_queries you would likely still OOM and there is point after which a chunk_length too big is just counter-productive. ","from":"developer"},{"body":"Committed.","from":"developer"},{"body":"Rolled out into production. Works a charm! Even on 200Gb sstables.\n\nSincere appreciations on this one.","from":"developer"},{"body":"Won't the cache here leak?\nMany (most?) sstables are transient (gone after the next minor compaction), but this cache will just grow...","from":"developer"},{"body":"Oh yes, it seems like I missed that one - we should remove entry from the cache when SSTable gets compacted out. What do you think, Sylvain?","from":"developer"},{"body":"Or use a ConcurrentLinkedHashMap w/ fixed capacity?","from":"developer"},{"body":"yes but that will also imply that we should weight it in memory size instead of number of entries so need to use jamm which is calculation overhead, better just remove unused because we know precisely when to do that...","from":"developer"},{"body":"patch according to Pavel's suggestion","from":"developer"},{"body":"Ok, I think using an object cache is ugly (I know that it was my idea).\n\nAt first, I tried going with the natural idea, to add the compressionMetadata as a final field of SSTableReader and use that everywhere, ensuring we use one per sstable. Turned out that for SSTableWriter you need to have the metadata existing before the SSTableReader is created and that seemed like a bit of a mess so I backtracked and decided to go with an object cache in CompressionMetada, but more out of laziness than anything else.\n\nThat was wrong of me to be lazy. We don't need that object cache and if its use is going to leak out of CompressionMetada (like hard coding the addition of the notifier in the DataTracker constructor; which defeats the purpose of the notifier abstraction in the first place) then it's not even clean as far as code is concerned.\n\nAttaching a v2 patch that remove the cache and do a slight modification of the initial idea, that is it just let CompressedSegmentedFile create the CompressionMetada and have the rest of the code use that. Turns out that once I plug my brain, it's only a few lines of code.\n","from":"developer"},{"body":"+1","from":"developer"},{"body":"Alright, committed this new version, thanks","from":"developer"},{"body":"Handling jvm memory since upgrading to cassandra-1.0 and enabling compression is still a headache.\nWhere i used to be able to run w/ Xmx8G i'm now struggling to run with Xmx20G (all caches are disabled) and during startup can hit\n{noformat}java.lang.OutOfMemoryError: Java heap space\n at org.apache.cassandra.utils.BigLongArray.(BigLongArray.java:53)\n at org.apache.cassandra.utils.BigLongArray.(BigLongArray.java:39)\n at org.apache.cassandra.io.compress.CompressionMetadata.readChunkOffsets(CompressionMetadata.java:122){noformat}\n\nI could keep increasing chunk_length (it's already at 256) but this seems awkward just to get a cluster running smoothly. At minimum the calculations for memory requirements for cassandra should be re-written if compression is to take such a large chunk of heap.","from":"developer"},{"body":"Does heap usage stay high-post startup? Can you try forcing a full GC to check that?","from":"developer"},{"body":"bq. Does heap usage stay high-post startup? Can you try forcing a full GC to check that?\nYes. GC doesn't seem to help, i'll attach a munin graph that shows it over time. It was running for a number of days just under 20G, but you can see from that how \"squeezed\" it was.\n\n(invoking full gc via jmx has no noticeable effect on heap used)","from":"developer"},{"body":"0.8 was running on Xmx8G up until week 44.\nat that point we upgraded to 1.0 and enabled compression. The very high memory usage at the beginning of week 44 was to handle the change from chunk_length 16 to 256.\nThen for most of week 44 and week 45 we ran with Xmx16G, but this was very \"squeezed\". Now that's OOM, and raising it to 20G didn't help. Currently it's on 30G.\n\nAlso note we're always used -XX:CMSInitiatingOccupancyFraction=60 for this cluster.\n(full java opts are \"-Xss128k -XX:+UseThreadPriorities -XX:ThreadPriorityPolicy=42 -XX:SurvivorRatio=8 -XX:MaxTenuringThreshold=1 -XX:+UseParNewGC -XX:+UseConcMarkSweepGC -XX:+CMSParallelRemarkEnabled -XX:CMSInitiatingOccupancyFraction=60 -XX:+UseCMSInitiatingOccupancyOnly -Xmx30g -Xmx30g -Xmn800M -XX:ParallelCMSThreads=4 -XX:+CMSIncrementalMode -XX:+CMSIncrementalPacing -XX:CMSIncrementalDutyCycleMin=0 -XX:CMSIncrementalDutyCycle=10\". the last 5 were added during week 44 to try and help, ref http://blog.mikiobraun.de/2010/08/cassandra-gc-tuning.html)","from":"developer"},{"body":"What version exactly are you running now? 1.0.2? 1.0.0 + 3427? Something else?","from":"developer"},{"body":"I've attached a tiny patch (0001-debugging.patch) that prints in the log the size of the long array we allocate for the chunk offsets. Would you mind trying with this and attach a log of when startup hits one of the OOM you pasted earlier (feel free to use a 8GB heap if it's easier to reproduce). I'd like to know if those offsets are indeed the problem.","from":"developer"},{"body":"I can get you a place to upload the heap dump from an OOM too. (8GB would be best, since heap analysis requires ram proportional to the heap size.)\n\nTo rule out the obvious, have you tried running 1.0.x w/o compression?","from":"developer"},{"body":"version: 1.0.2 snapshot (pretty close to release date) + CASSANDRA-3197.\nw/o compression: that would require a full compact/scrub. that takes close to 24hrs :-(\npatch: i attach that and hopefully have output soon. a heap dump can be done at the same time...","from":"developer"},{"body":"startup log with debug (off 1.0.2 release){noformat}INFO 48:34,688 DatabaseDescriptor: Loading settings from file:/iad/finn/countstatistics/conf/cassandra-prod.yaml\nINFO 48:34,782 DatabaseDescriptor: DiskAccessMode 'auto' determined to be mmap, indexAccessMode is mmap\nINFO 48:34,792 DatabaseDescriptor: Global memtable threshold is enabled at 512MB\nINFO 48:34,890 AbstractCassandraDaemon: JVM vendor/version: Java HotSpot(TM) 64-Bit Server VM/1.6.0_24\nINFO 48:34,891 AbstractCassandraDaemon: Heap size: 760414208/8506048512\nINFO 48:34,891 AbstractCassandraDaemon: Classpath: /iad/finn/countstatistics/jar/countstatistics.jar:/iad/common/apps/cassandra/lib/jamm-0.2.5.jar\nINFO 48:37,158 CLibrary: JNA mlockall successful\nINFO 48:37,879 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/Versions-h-42 (256 bytes)\nINFO 48:37,879 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/Versions-h-41 (256 bytes)\nINFO 48:37,879 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/Versions-h-40 (256 bytes)\nINFO 48:37,959 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/IndexInfo-h-3 (223 bytes)\nINFO 48:38,001 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/Schema-h-15 (34257 bytes)\nINFO 48:38,045 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/Migrations-h-15 (78524 bytes)\nINFO 48:38,096 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/LocationInfo-h-150 (80 bytes)\nINFO 48:38,096 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/LocationInfo-h-149 (628 bytes)\nINFO 48:38,096 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/system/LocationInfo-h-151 (163 bytes)\nINFO 48:38,192 DatabaseDescriptor: Loading schema version 1940c630-0be4-11e1-0000-d1695892b1ff\nINFO 51:35,136 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191473 (38646535 bytes)\nINFO 51:35,136 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190467 (2284524668 bytes)\nINFO 51:35,136 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191469 (254927460 bytes)\nINFO 51:35,136 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191475 (30477008 bytes)\nINFO 51:35,136 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-114136 (156044360682 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191294 (4585008988 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190415 (15857295280 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-183183 (196289440978 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191472 (1346076 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190736 (4626053255 bytes)\nINFO 51:35,137 SSTableReader: Opening /iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191435 (1188223188 bytes)\nINFO 51:35,187 CompressionMetadata: Allocating chunks index for 5745 chunks for uncompressed size of 1470519 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191472-CompressionInfo.db)\nINFO 51:35,421 CompressionMetadata: Allocating chunks index for 129646 chunks for uncompressed size of 33189311 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191475-CompressionInfo.db)\nINFO 51:35,544 CompressionMetadata: Allocating chunks index for 165602 chunks for uncompressed size of 42393918 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191473-CompressionInfo.db)\nINFO 51:37,171 CompressionMetadata: Allocating chunks index for 1091377 chunks for uncompressed size of 279392485 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191469-CompressionInfo.db)\nINFO 51:41,148 CompressionMetadata: Allocating chunks index for 5086138 chunks for uncompressed size of 1302051278 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191435-CompressionInfo.db)\nINFO 51:46,351 CompressionMetadata: Allocating chunks index for 9766541 chunks for uncompressed size of 2500234376 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190467-CompressionInfo.db)\nINFO 51:56,717 CompressionMetadata: Allocating chunks index for 19828434 chunks for uncompressed size of 5076078986 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190736-CompressionInfo.db)\nINFO 51:56,897 CompressionMetadata: Allocating chunks index for 19626358 chunks for uncompressed size of 5024347477 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-191294-CompressionInfo.db)\nINFO 52:21,670 CompressionMetadata: Allocating chunks index for 67865822 chunks for uncompressed size of 17373650297 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-190415-CompressionInfo.db)\nINFO 55:55,920 CompressionMetadata: Allocating chunks index for 666981588 chunks for uncompressed size of 170747286320 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-114136-CompressionInfo.db)\nINFO 56:49,620 CompressionMetadata: Allocating chunks index for 840404671 chunks for uncompressed size of 215143595584 (/iad/finn/countstatistics/cassandra-data/countstatisticsCount/thrift_no_finntech_countstats_count_Count_neg8589045746818385983-h-183183-CompressionInfo.db)\nERROR 57:51,112 AbstractCassandraDaemon: Fatal exception in thread Thread[SSTableBatchOpen:8,5,main]\njava.lang.OutOfMemoryError: Java heap space\n\tat org.apache.cassandra.utils.BigLongArray.(BigLongArray.java:53)\n\tat org.apache.cassandra.utils.BigLongArray.(BigLongArray.java:39)\n\tat org.apache.cassandra.io.compress.CompressionMetadata.readChunkOffsets(CompressionMetadata.java:127)\n ...{noformat}","from":"developer"},{"body":"That's the stupidest bug ever. It happens we interpret the chunk_length_in_kb not in kb but in bytes.\nAnyway, I've created CASSANDRA-3492 to address this.\nTurns out if you don't update the chunk_length you're fine because the default is ok, but I guess hitting this issue initially has put you in the wrong spot :(","from":"developer"}],"created":"2011-10-31T13:57:36.000+0000","description":"The CompressionMetada holds the compressed block offsets in memory. Without being absolutely huge, this is still of non-negligible size as soon as you have a bit of data in the DB. Reallocating this for each read is a very bad idea.\n\nNote that this only affect range queries, since \"normal\" queries uses CompressedSegmentedFile that does reuse a unique CompressionMetadata instance.\n\n( Background: http://thread.gmane.org/gmane.comp.db.cassandra.user/21362 )","issue_id":"12529483","key":"CASSANDRA-3427","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-10-31T15:44:26.000+0000","role":"fixed_distractor","summary":"CompressionMetadata is not shared across threads, we create a new one for each read"} {"case_id":"12531095","cluster":"DISTRACTOR-CASSANDRA-3483","comments":[{"body":"Some discussion from irc:\n\n{noformat}\n23:43 < goffinet> has datastax ever had a customer add a new datacenter to an existing cluster? No docs or info on web suggest anyone has done this before\n23:44 < driftx> yeah\n23:44 < goffinet> how is it done? we are running a case where if i modify strategy options before adding nodes, writes will fail since no endpoints for DC have been added\n23:44 < goffinet> we were expecting this might work because we want to bootstrap the new DC to the existing cluster\n23:44 < goffinet> take on writes + stream data with RF factor\n23:45 < driftx> general best practice is (jbellis can correct if I'm outdated) add the dc at rf:0, add the nodes/update snitch, repair\n23:45 < driftx> err, update rf, repair\n23:46 < goffinet> yeah mind if i open up a jira? that seems extreme to make the cluster do that .. ?\n23:46 < goffinet> or is repair smart enough to just stream ranges instead of AES?\n23:46 < driftx> 'instead of AES?' that's what repair is, but if just streams ranges\n23:46 < driftx> s/if/it/\n23:47 < goffinet> right but AES builds merkle tree, scans through all data ?\n23:47 < goffinet> isn't bootstrap a different operation?\n23:47 < goffinet> when streaming just sstables\n23:47 < driftx> yeah, it is\n23:47 < goffinet> yeah thats more heavy. dont understand why we couldnt use that instead\n23:47 < goffinet> like bootstrap\n23:48 < stuhood> now that i think about it, it doesn't really make sense that a CL.ONE write fails if a DC isn't available\n23:48 < stuhood> independent of the bootstrap case, that sounds like the real issue\n23:49 < stuhood> goffinet: ^\n23:50 < driftx> hmm, yeah that doesn't\n23:50 < driftx> but the problem with bootstrapping a dc is the first node you bootstrap gets everything\n23:50 < goffinet> stuhood: yeah. it was complaining about not enough endpoints \n23:50 < goffinet> driftx: why is that? if you are doubling the cluster, and assign the tokens manually ?\n23:51 < driftx> still have to do them 2 mins apart, and they're probably going to be part of the same replica set which I think is troublesome too\n23:51 < goffinet> driftx: maybe we can make repair a bit more intelligent? if no data exists on the node .. just stream the ranges instead of using AES\n23:52 < driftx> problem is we're pushing AES to do the entire replica set (which is nearly does now)\n23:52 < stuhood> goffinet: it shouldn't be as heavyweight as you're thinking\n23:53 < goffinet> stuhood: but we have a way currently that is less heavy\n23:53 < goffinet> i dont understand why we couldnt use that method\n23:53 < stuhood> not implemented =)\n23:53 < goffinet> don't cut corners :)\n23:53 < stuhood> human time vs cpu time =P\n23:54 < driftx> you could almost do something like #3452 and then have a jmx call to say 'ok, finish'\n23:54 < CassBotJr> https://issues.apache.org/jira/browse/CASSANDRA-3452 : Create an 'infinite bootstrap' mode for sampling live traffic\n23:54 < driftx> except the first one that tries is going to have every node pound it with all the writes\n23:54 < goffinet> driftx: ill make a jira ticket so we can discuss there, it doesn't seem like it would be too much trouble to support this use case\n23:54 < goffinet> we'd be happy to write the patch after some input\n23:55 < driftx> trickier than it sounds I'll bet, but sgtm\n23:57 < stuhood> alternatively, is now the right time to add back group bootstrap?\n23:58 < stuhood> so you'd 1) add the dc to the strategy, 2) do a group bootstrap of the entire dc\n23:58 < stuhood> would also have to fix the CL.ONE problem though.\n23:59 < goffinet> how did group bootstrap work again?\n23:59 < driftx> #2434 is relevant\n23:59 < CassBotJr> https://issues.apache.org/jira/browse/CASSANDRA-2434 : range movements can violate consistency\n--- Day changed Fri Nov 11 2011\n00:00 < stuhood> goffinet: bootstrapping many nodes at once without the 2 minute wait\n00:01 < goffinet> why was it removed?\n00:01 < stuhood> used zookeeper\n00:01 < goffinet> oh.\n00:01 < stuhood> but come to think of it, removing the 2 minute wait would seem to be relatively easy\n00:02 < goffinet> stuhood, i thought the 2 minute wait was just waiting for ring state to settle?\n00:02 < goffinet> before it streamed from nodes\n00:02 < stuhood> goffinet: yea: you could form a \"group\" bootstrap by inverting things and waiting until you -hadn't- seen a new node in 2-10 minutes before you chose a token and started bootstrapping\n00:03 < stuhood> so, not terribly simple, but.\n00:04 < stuhood> you'd basically have a bunch of nodes sitting around waiting until no new nodes started, and then they have to deterministically choose tokens.\n00:05 < goffinet> yes\n00:05 < stuhood> well, alternatively, you wouldn't need a new way to deterministically choose tokens\n00:05 < stuhood> (easier)\n00:05 < stuhood> no… scratch that. you would need a way\n00:05 < stuhood> for this DC case, all of the nodes are entering an empty ring\n00:06 < stuhood> so the group would need to choose something balanced\n00:06 < goffinet> empty ring?\n00:06 < stuhood> yea, essentially… there are no tokens in that dc\n00:06 < goffinet> but we were going to provide the tokens manually?\n00:06 < goffinet> were you thinking of making it automatic?\n00:07 < stuhood> yea. fixing bootstrapping groups of nodes would make automatic safe again\n00:08 < stuhood> so… whatever state a node is in when it is sitting and waiting for enough information to choose a token, it should just stay that way and watch what other nodes enter that state\n00:08 < goffinet> so i have a question about the 120 second window you have to wait..\n00:09 < stuhood> mm\n00:09 < driftx> hmm, what if they started up at rf:0 but stayed in some dead state (hibernate might work) without doing anything until you changed the rf, then actually bootstrapped?\n00:09 < goffinet> so imagine i startup all the nodes in DC2 at same time, does join_ring=false not grab gossip info at all? I was thinking it would be good if we could just start gossip on all nodes, but until operator says 'go' then i could bootstrap them all at same time\n00:09 < goffinet> since i would only have to wait at most 120 seconds before kicking them all off\n00:10 < stuhood> driftx: yea, that could work too… but you'd still need to choose tokens. (also, the rf=0 thing shouldn't be necessary, right? that's the CL.ONE bug)\n00:11 < driftx> well, you really want to choose tokens anyway\n00:11 < stuhood> goffinet: it does get gossip… i think that's basically equivalent to the pre-join state\n00:11 < driftx> I guess you don't need rf=0 if all the nodes are in hibernate\n00:12 < goffinet> yeah i think you do need hibernate in this case, because if i set tokens upfront, i want all nodes to know about ATL ones too\n00:12 < goffinet> before i kick off bootstrap\n00:12 < stuhood> driftx: i'm confused… what is the difference between rf=0 and not being there?\n00:12 < stuhood> is that a workaround for the CL.ONE bug?\n00:13 < driftx> you know there's a dc with rf:0, can add one with impacting anything\n00:13 < driftx> err, without\n00:14 ?? boaz__ (0819c319@gateway/web/freenode/ip.8.25.195.25) has joined #cassandra-dev\n00:14 < stuhood> so what was the point of adding it? that's why i'm confused...\n00:14 < goffinet> im fine with rf:0, its so you can add the nodes to the cluster before calling repair\n00:14 < goffinet> before you add nodes\n00:15 < driftx> because the dc is in the schema\n00:15 < driftx> so you need it there to have nodes be in it\n00:15 < stuhood> ah\n00:16 < goffinet> driftx: any reason why we couldnt just fix that? so dc2:3 wont throw an error if nodes are down?\n00:16 < goffinet> that way you would needed to do two steps\n00:16 < goffinet> dc2:0, add nodes, dc2:3\n00:16 < goffinet> wouldn't*\n00:16 < driftx> I don't understand, you can already do that\n00:17 < driftx> you just have to repair afterwards\n00:17 < goffinet> it throws an error currently? if you set dc2:3 and no nodes exist for dc2\n00:17 < goffinet> we'll double check on that\n00:18 < goffinet> for writes\n00:18 < driftx> oh, it does\n00:19 < driftx> but only for writes\n00:19 < goffinet> yeah\n00:19 < goffinet> so thats fine, thats fixable\n00:19 < goffinet> im just curious about a) how can we bootstrap nodes without 120s delays between N nodes b) stream from DC1 without AES\n00:21 < stuhood> goffinet: if you figure out a, i don't think b is necessary?\n00:22 < stuhood> assuming they are aware of the other joining nodes, and can all join the same range\n00:22 < stuhood> that would be the keystone for some kind of group bootstrap\n00:23 < goffinet> let me test out join_ring, because im curious. if join_ring=false still gossips but doesnt offically join.. it would be nice if node 2 in DC2 knew about that node too somehow?\n00:23 < driftx> that's why I proposed cheating, add them all as non-members, then ask them to bootstrap\n00:23 < goffinet> because then .. i could just run a command on each node at same time\n00:23 < goffinet> since they all know about each other in a hibernate state\n00:23 < goffinet> driftx: yes i like that\n00:24 < driftx> private void joinTokenRing(int delay) throws IOException, org.apache.cassandra.config.ConfigurationException\n00:24 < driftx> {\n00:24 < driftx> logger_.info(\"Starting up server gossip\");\n00:24 < driftx> they don't use gossip with join_ring off\n00:24 < stuhood> but will that actually allow them to all join the same range?\n00:24 < goffinet> okay cool, yeah we would need to make it join in that special state then\n00:25 < stuhood> i think there is an edgecase here… if multiple nodes are joining the same range, and one of them fails, then should they all fail?\n00:25 < driftx> no, it basically saves you server startup time that is not ring-related :)\n00:25 < goffinet> stuhood, they all know the tokens ahead of time?\n00:25 < goffinet> they just need to know the current global state of things\n00:25 < stuhood> goffinet: right, but if they are streaming the range that they will be responsible for...\n00:26 ?? mw1 (~Adium@8.25.195.29) has quit (Quit: Leaving.)\n00:26 < stuhood> Joining nodes don't stick around if they fail\n00:26 < goffinet> they shouldnt be allowed to do that until they joined ?\n00:26 < stuhood> nah, you stream while you are joining… unless you are talking about repair\n00:26 < goffinet> stuhood: was that removed? i thought u had to still remove the node\n00:26 < goffinet> using the new options in 1.0\n00:26 < stuhood> don't know about 1.0\n00:27 < driftx> no, a failed non-member is just a fat client and disappears\n00:27 < goffinet> but i thought there was a timeout for fat client ?\n00:27 < goffinet> is it 30s or something?\n00:27 < driftx> yes\n00:28 < goffinet> so nodes that arent fat clients, why might we remove them ? if we didnt..\n00:28 < goffinet> and let the operator do it\n00:28 < goffinet> or have a larger timeout\n00:28 < goffinet> might make this a non-issue\n00:28 < driftx> what does a larger timeout/keeping them around buy you?\n00:29 < goffinet> because if they go away, and i bootstrap after they failed, wont my view of ring be skewed?\n00:29 < stuhood> driftx: i guess in this case, the node would resume bootstrapping from where it left off\n00:29 < driftx> it would've missed writes in the meantime and require a repair afterwards anyway\n00:29 < stuhood> sorry… \"resume\" in the sense of \"start over\", but yea\n00:31 < stuhood> that would be a pretty big change, but it might make sense\n00:31 < goffinet> stuhood: what would you change\n00:31 < stuhood> what you said, about nodes in joining staying in joining\n00:31 < stuhood> so if the machine restarts, it begins joining at the same position again\n00:33 < goffinet> if we supported that + letting nodes gossip in hibernate, would allow us to add capacity at operator control\n{noformat}","created":"2011-11-11T05:43:51.771+0000"},{"body":"I think it wouldn't be crazy and actually very (very) simple to add a new nodetool command (rebuild?) that would basically have the node asks the other replicas to stream all there data to him (for the correct ranges obviously). In other words, a command that force the streaming part of bootstrap without all the join ring part. Or another way to say is to have the streaming part of a repair but without the validation part.\n\nThe method to add a new DC would be the same as today except that repair would be replaced by this new operation.","created":"2011-11-16T20:49:20.283+0000"},{"body":"We would still need to put the nodes into a 'bootstrap' state to get incoming writes forwarded to them, otherwise you have to repair in the end anyway.","created":"2011-11-16T20:52:07.504+0000"},{"body":"Yeah as Brandon mentioned, we would still want to go into the bootstrap state to get those writes. This would also allow us to add capacity in the same way, if we manually pick tokens (auto bootstrap is kinda worthless IMO) to existing DC as well. We can just fire off the bootstrap command from nodetool as needed.","created":"2011-11-17T04:27:03.336+0000"},{"body":"Sorry, I had to re-read what Sylvain said for it to 'click'. So the process he proposes is as follows with RF of 3:\n\n1. strategy options dc2:0\n2. bring up new nodes in dc2 with auto_bootstrap off and token set\n3. set strategy options dc2:3\n4. run 'rebuild' on each node in dc2\n\nthis would handle the writes part.\n\ni was kinda hoping though that we could modify the gossip state because I could very well see this playing into the case where you weren't adding DCs but wanted to add lots of nodes (60-100 like we do currently) ... and wanted to have them all added to existing DC.. having the bootstrap defined that way, would allow us to bootstrap nodes as we please in existing DCs, bring them all to the ring at once to have a consistent state without taking on traffic until they transitioned states (joining/normal). Where as this proposal wouldn't be able to satisfy that use case.\n\n","created":"2011-11-17T06:30:04.893+0000"},{"body":"So you're proposing to add support for bootstrapping multiple nodes together. I'm not against that, it would be nice and that would give you 90% of what this ticket is about (you'd have to add the ability to multi-boostrap *and* add a DC/augment the replication faction at the same time). But it is orders of magnitude more complicated than what I'm suggesting. Which is not a problem in itself given it's a broader solution, but it means we'll have multi-node boostrap at best for 1.1, while I'm pretty sure I can get the 'rebuild' command wrote in like an hour (and I see no reason why it couldn't be put in the 1.0 series).\n\n","created":"2011-11-17T09:18:43.775+0000"},{"body":"Sylvain, it looks like a state 'HIBERNATE' exists in GOSSIP already based on recovering a dead node in 1.0. Do you have a preference on the name of the new state if we attempted adding multiple nodes patch? WAITING? STANDBY?","created":"2011-11-18T00:22:00.871+0000"},{"body":"I was thinking of GHOST too after I realized we'll need a state where the node receives neither writes nor reads (more details here later).","created":"2011-11-18T04:22:10.421+0000"},{"body":"So, am I understanding correctly that we're talking about two different scenarios?\n\n- Add new DC without repair\n- Add many nodes to existing DC without RING_DELAY in between\n\nI think Sylvain's proposal addresses the first nicely. So what I need help with is understanding what problem you're trying to solve with the second part. Dealing with overlapping ranges in node movement basically requires a rewrite of that subsystem (CASSANDRA-2434). But I suspect there is a \"good enough\" solution that we could find if I understood better what your pain point is here.","created":"2011-11-21T21:03:49.970+0000"},{"body":"Jonathan,\n\nYou are correct. Sylvain's proposal does satisfy this ticket. It doesn't solve the case of (2) where if you want to add lots of nodes in an existing DC, and you know the tokens they should be at, and want to join them all at once.\n\nOur use case is, we actually add 60-100 nodes in one big capacity add. We would like to avoid the 120 second per node time frame. It's not a deal breaker though. We actually realized though if we are adding that many nodes to our cluster, with a large cluster already, we need to rebalance heavily anyway. Peter is almost done with the 'rebuild' patch, I'm assigning him to this ticket. \n\nOur next big focus is improving the rebalancing of a cluster. We have a very large cluster and after adding 100 nodes every month or so, this becomes painful. Almost all of our nodes have over 600GB+ each. We have an application that will require us to be rebalancing at all times to reduce our hot spots.\n\n","created":"2011-11-22T05:57:02.521+0000"},{"body":"We ended up going for the simpler rebuild patch as Chris hinted at. I'll quote myself:\n\n{quote}\nI've been looking at this some more.\n\nHere's the proposal so far:\n\nWe introduce a GHOST state, in which a node receives neither reads nor writes. This allows us to bring in a group of nodes in the ring without suffering any ill effects. It is completely invisible to reads and writes, and will never count towards e.g. consistency level.\n\nOnce all ghosts are ready, we can start bootstrapping, taking ghost nodes into accounts for purposes of determining which range we are responsible for, but streaming only from non-ghost nodes. This accomplishes the goal of not transferring more data than necessary.\n\nIn order to avoid a bootstrapping node from taking writes for more than it's eventual share, we'd have to make the write endpoints be aware of ghost nodes. This is dosable, but not critical since we're bisecting a range that was previously handled by a single node anyway so the traffic would be managable. It would just be cleaner to not have to cleanup afterwards.\n\nOnce we transition from bootstrapping to being \"up\", we have a bigger problem however. If the read paths and write paths are only aware of non-ghost nodes, the read/write paths would think that these nodes had more ownership than they really do.\n\nSo we must really be taking into account the other nodes in the read/write path as well - but only when determining ownership of a completely bootstrapped node that was previously a ghost. That means we must distinguish between a \"normally up\" node and one that's been bootstraped from ghost state (call it \"ghost strapped\").\n\nThis suddenly gets complex. The process of group bootstrap would then be:\n\nAdd a bunch of nodes in GHOST state\nBootstrap all of them, each of them going into GHOSTSTRAPPED state\nOnce all are GHOSTSTRAPPED, we can safely transntion from GHOSTSTRAPPED to normal/up\nIs there a simpler solution?\n{quote}\n\nAfter some additional discussion we felt this was adding to much complexity and potential edge cases/bugs that it became more cost-effective to just go with the simple rebuild for our immediate needs, hoping to address the problem of adding lots of nodes to a DC separately in some other way.\n\nA patch is forthcoming soon.\n","created":"2011-11-23T00:51:51.712+0000"},{"body":"Here is a patch rebased against 0.8 for cursory review. I do not expect this to go into 0.8, and in fact I have not tested this patch other than build against vanilla 0.8 (the original patch is tested, but against our internal 0.8).\n\nIf there are no concerns with the overall implementation, I'll submit a rebased version for 1.0/trunk.\n\nThere are two components of the change:\n\n* Breaking out the streaming part of BootStrapper into a separate RangeStreamer. Change BootStrapper to use that.\n\n* Implement the rebuild command on top of RangeStreamer.\n\nThere are two ways to invoke rebuild:\n\n{code}\nnodetool rebuild\nnodetool rebuild nameofdc\n{code}\n\nThe first form streams from nearest endpoints, while the latter streams from nearest endpoints in the specified data center.\n\n","created":"2011-11-30T18:29:06.179+0000"},{"body":"bq. If there are no concerns with the overall implementation, I'll submit a rebased version for 1.0/trunk.\n\nThis looks good to me, I like the RangeStreamer abstraction.","created":"2011-11-30T18:40:44.384+0000"},{"body":"Attached is a version rebased against 1.0 (and tested).","created":"2011-12-05T01:40:03.674+0000"},{"body":"Once this option is in, is this the procedure for running rebuild (with 4 changed to 'rebuild dc1')?\n\n{quote}\n1. strategy options dc2:0\n2. bring up new nodes in dc2 with auto_bootstrap off and token set\n3. set strategy options dc2:3\n4. run 'rebuild' on each node in dc2\n{quote}\n\nDo we need to stagger issuing the rebuild commands or can they be run all at once?\n","created":"2011-12-05T16:09:23.069+0000"},{"body":"Yes, that looks good.\n\nAnd yes, you can run rebuilds concurrently as long as you're comfortable with the amount of bandwidth you'll be pushing and the load you'll be putting on the source nodes.\n\nHowever, if you expect to see reasonable performance and streaming at full speed to all nodes, you also need CASSANDRA-3494.\n\nRegardless: I strongly recommend testing this with your exact version of Cassandra before trying it for real.\n","created":"2011-12-05T19:28:45.952+0000"},{"body":"I haven't applied the patch yet, it needs rebase and preferably against trunk since that is the likely target for this, but a few comments.\n\nWe could have more reuse of code between Boostrapper ant the rebuild command. Typically:\n* RangeStreamer.getAllRangeWithSourcesFor does essentially the same thing that Boostrapper.getRangesWithSources, so it would be nice to do some reuse.\n* In rebuild, we essentially have the code of Boostrapper.getWorkMap, again would be nice to do some code reuse.\n\nI think we should move all of those in RangeStreamer and ultimately Boostrapper.boostrap() should be just one call to rebuild with the right arguments (mostly the correct tokenMetada instance and the \"myRange\" collection).\n\nA few nits:\n* rebuild code could be simplified slightly by using StorageService.getLocalRanges()\n* rebuild doesn't fully respect the code style.\n","created":"2011-12-14T00:37:22.326+0000"},{"body":"I'll get it rebased once it's otherwise okay.\n\nAs for re-use: I had intermediate versions that tried to do this, but ever time I ended up realizing that it was exploding in verbosity at the point where I was using the abstraction so it didn't actually help. However, I think there were a few changes towards the end after which I didn't re-evaluate.\n\nI'll look at it again and see what I can do.","created":"2011-12-14T01:52:49.903+0000"},{"body":"I'll get it rebased once it's otherwise okay.\n\nAs for re-use: I had intermediate versions that tried to do this, but ever time I ended up realizing that it was exploding in verbosity at the point where I was using the abstraction so it didn't actually help. However, I think there were a few changes towards the end after which I didn't re-evaluate.\n\nI'll look at it again and see what I can do.","created":"2011-12-14T01:52:49.939+0000"},{"body":"Attaching version rebased to trunk but not yet re-factored.","created":"2011-12-24T04:08:43.339+0000"},{"body":"Peter, are you planning to follow up on Sylvain's comments still?","created":"2012-01-25T21:24:09.068+0000"},{"body":"I do. I'm sorry for the delay, this has been nagging me for quite some time. It's not forgotten, I have just been inundated with urgent stuff to do.\n\nI'm attaching a fresh rebase against current trunk and I hope to submit an improved version later tonight (keyword being \"hope\").","created":"2012-01-28T05:31:28.765+0000"},{"body":"{{CASSANDRA\\-3483\\-trunk\\-refactored\\-v1.txt}} addresses the duplication between BootStrapper and RangeStreamer.\n\nNext patch will address rebuild/getworkmap duplication.\n","created":"2012-01-28T07:06:18.998+0000"},{"body":"(It also contains the addition of a brace from CASSANDRA-3806; this is intentional to avoid pain.) ","created":"2012-01-28T07:07:04.438+0000"},{"body":"I borked the unit test, will address that too.","created":"2012-01-28T07:12:34.281+0000"},{"body":"{{CASSANDRA\\-3483\\-trunk\\-refactored\\-v2.txt}} I believe addresses the concerns, plus makes other improvements. I'm much more happy with this one.\n\nIt addresses CASSANDRA-3807 by supporting fetch \"consistency levels\" (though only ONE is currently usable without patching), and the filtering of hosts is abstracted out.\n\nThere is still some duplication between {{Bootstrapper.bootstrap()}} and {{StorageService.rebuild()}} in that both do the dance of iteration over tables to construct the final map. I am not really feeling that abstracting away that is a good idea to include in this ticket, though I think it's worthwhile doing at some point separately.\n\nThe unit test is fixed; my adjustment of it was wrong because I wasn't picking pending ranges (in the test).\n\nI've tested both rebuild and bootstrap in a 3 node cluster.\n\nI've added some more logging than what is typically the case; there have been several cases where I wished streaming was logged in more detail at INFO, particularly when bootstrapping or rebuilding. I think it's worthwhile to get that in while at it.","created":"2012-01-28T10:00:16.097+0000"},{"body":"bq. It addresses CASSANDRA-3807 by supporting fetch \"consistency levels\" (though only ONE is currently usable without patching)\n\nI think a first issue is that current bootstrap does not fail if no node is alive for a given range, which arguably it should. I'm good with doing that, though it would be worth backporting to 1.0 too so it may be worth splitting that to a separate patch (or rather just create one for the fix in 1.0).\n\nHowever, that does not solve the problem of bootstrap possibly breaking the consistency contract. The problem being that if we transfer a range from a node that happens to be lacking behind in term of consistency, and we end up replacing a node that was not lacking behind, we could break some consistency contracts. To fix that, I really only see only one solution right off the bat (which doesn't mean there isn't other): it is to ensure that for each range, we transfer it from (at least) the node we will replace for this range.\n\nI believe the FetchConsistencyLevel of this patch is making an attempt to fix this by allowing to fetch from more than one node. While it does make it less likely to break consistency, unless we fetch from all nodes (and thus the one we'll replace), we cannot be sure we won't break the consistency level for people that say write at CL.ALL and read at CL.ONE. Overall, I fully agree this is a problem that we should fix someone, but I'm not sure the FetchConsistencyLevel is the right solution and even if it is it's a complicated enough problem that it's worth it's own ticket. I would agree that the problem with rebuild is a little bit different, but since anyway the patch introduce FCL without using it, let's keep that for later if that's ok.\n\nbq. There is still some duplication between Bootstrapper.bootstrap() and StorageService.rebuild() in that both do the dance of iteration over tables to construct the final map. I am not really feeling that abstracting away that is a good idea to include in this ticket\n\nI think it is, at least for a good chunk of it. It's not very complicated, it clearly improves code readability and since the patch already refactor that code I don't see a good reason to push that to later, especially if we agree it's worthwhile.\n\nAttaching a v3 that 1) remove FetchConsistencyLevel for the reasons above and 2) move most of the details of creating the multimaps in RangeStreamer.\n","created":"2012-01-30T12:47:03.467+0000"},{"body":"bq. Overall, I fully agree this is a problem that we should fix someone, but I'm not sure the FetchConsistencyLevel is the right solution and even if it is it's a complicated enough problem that it's worth it's own \n\nThis is CASSANDRA-2434 isn't it?","created":"2012-01-30T14:04:40.419+0000"},{"body":"bq. This is CASSANDRA-2434 isn't it?\n\nit is.","created":"2012-01-30T14:09:01.514+0000"},{"body":"cleanup patch addressing mostly typos and style. only substantial code change was to RangeStreamer.getRangeFetchMap. Also, moved OperationType.REBUILD to the end of the enum to make sure we don't break anything depending on ordinal.\n\n+1 from me otherwise.","created":"2012-01-30T14:49:18.819+0000"},{"body":"Committed v3 + Jonathan's cleanups (and a fix to the unit test).","created":"2012-01-30T15:24:36.831+0000"},{"body":"For the record I never intended to fix the general problem of bootstrapping never violating consistency. But in retrospect it's obvious how my choice of naming would make it sound like I did :) I agree it's a problem for its own ticket.\n\nThanks!","created":"2012-01-30T18:12:02.434+0000"}],"conversations":[{"body":"Was talking to Brandon in irc, and we ran into a case where we want to bring up a new DC to an existing cluster. He suggested from jbellis the way to do it currently was set strategy options of dc2:0, then add the nodes. After the nodes are up, change the RF of dc2, and run repair. \n\nI'd like to avoid a repair as it runs AES and is a bit more intense than how bootstrap works currently by just streaming ranges from the SSTables. Would it be possible to improve this functionality (adding a new DC to existing cluster) than the proposed method? We'd be happy to do a patch if we got some input on the best way to go about it.\n","from":"reporter","subject":"Support bringing up a new datacenter to existing cluster without repair"},{"body":"Some discussion from irc:\n\n{noformat}\n23:43 < goffinet> has datastax ever had a customer add a new datacenter to an existing cluster? No docs or info on web suggest anyone has done this before\n23:44 < driftx> yeah\n23:44 < goffinet> how is it done? we are running a case where if i modify strategy options before adding nodes, writes will fail since no endpoints for DC have been added\n23:44 < goffinet> we were expecting this might work because we want to bootstrap the new DC to the existing cluster\n23:44 < goffinet> take on writes + stream data with RF factor\n23:45 < driftx> general best practice is (jbellis can correct if I'm outdated) add the dc at rf:0, add the nodes/update snitch, repair\n23:45 < driftx> err, update rf, repair\n23:46 < goffinet> yeah mind if i open up a jira? that seems extreme to make the cluster do that .. ?\n23:46 < goffinet> or is repair smart enough to just stream ranges instead of AES?\n23:46 < driftx> 'instead of AES?' that's what repair is, but if just streams ranges\n23:46 < driftx> s/if/it/\n23:47 < goffinet> right but AES builds merkle tree, scans through all data ?\n23:47 < goffinet> isn't bootstrap a different operation?\n23:47 < goffinet> when streaming just sstables\n23:47 < driftx> yeah, it is\n23:47 < goffinet> yeah thats more heavy. dont understand why we couldnt use that instead\n23:47 < goffinet> like bootstrap\n23:48 < stuhood> now that i think about it, it doesn't really make sense that a CL.ONE write fails if a DC isn't available\n23:48 < stuhood> independent of the bootstrap case, that sounds like the real issue\n23:49 < stuhood> goffinet: ^\n23:50 < driftx> hmm, yeah that doesn't\n23:50 < driftx> but the problem with bootstrapping a dc is the first node you bootstrap gets everything\n23:50 < goffinet> stuhood: yeah. it was complaining about not enough endpoints \n23:50 < goffinet> driftx: why is that? if you are doubling the cluster, and assign the tokens manually ?\n23:51 < driftx> still have to do them 2 mins apart, and they're probably going to be part of the same replica set which I think is troublesome too\n23:51 < goffinet> driftx: maybe we can make repair a bit more intelligent? if no data exists on the node .. just stream the ranges instead of using AES\n23:52 < driftx> problem is we're pushing AES to do the entire replica set (which is nearly does now)\n23:52 < stuhood> goffinet: it shouldn't be as heavyweight as you're thinking\n23:53 < goffinet> stuhood: but we have a way currently that is less heavy\n23:53 < goffinet> i dont understand why we couldnt use that method\n23:53 < stuhood> not implemented =)\n23:53 < goffinet> don't cut corners :)\n23:53 < stuhood> human time vs cpu time =P\n23:54 < driftx> you could almost do something like #3452 and then have a jmx call to say 'ok, finish'\n23:54 < CassBotJr> https://issues.apache.org/jira/browse/CASSANDRA-3452 : Create an 'infinite bootstrap' mode for sampling live traffic\n23:54 < driftx> except the first one that tries is going to have every node pound it with all the writes\n23:54 < goffinet> driftx: ill make a jira ticket so we can discuss there, it doesn't seem like it would be too much trouble to support this use case\n23:54 < goffinet> we'd be happy to write the patch after some input\n23:55 < driftx> trickier than it sounds I'll bet, but sgtm\n23:57 < stuhood> alternatively, is now the right time to add back group bootstrap?\n23:58 < stuhood> so you'd 1) add the dc to the strategy, 2) do a group bootstrap of the entire dc\n23:58 < stuhood> would also have to fix the CL.ONE problem though.\n23:59 < goffinet> how did group bootstrap work again?\n23:59 < driftx> #2434 is relevant\n23:59 < CassBotJr> https://issues.apache.org/jira/browse/CASSANDRA-2434 : range movements can violate consistency\n--- Day changed Fri Nov 11 2011\n00:00 < stuhood> goffinet: bootstrapping many nodes at once without the 2 minute wait\n00:01 < goffinet> why was it removed?\n00:01 < stuhood> used zookeeper\n00:01 < goffinet> oh.\n00:01 < stuhood> but come to think of it, removing the 2 minute wait would seem to be relatively easy\n00:02 < goffinet> stuhood, i thought the 2 minute wait was just waiting for ring state to settle?\n00:02 < goffinet> before it streamed from nodes\n00:02 < stuhood> goffinet: yea: you could form a \"group\" bootstrap by inverting things and waiting until you -hadn't- seen a new node in 2-10 minutes before you chose a token and started bootstrapping\n00:03 < stuhood> so, not terribly simple, but.\n00:04 < stuhood> you'd basically have a bunch of nodes sitting around waiting until no new nodes started, and then they have to deterministically choose tokens.\n00:05 < goffinet> yes\n00:05 < stuhood> well, alternatively, you wouldn't need a new way to deterministically choose tokens\n00:05 < stuhood> (easier)\n00:05 < stuhood> no… scratch that. you would need a way\n00:05 < stuhood> for this DC case, all of the nodes are entering an empty ring\n00:06 < stuhood> so the group would need to choose something balanced\n00:06 < goffinet> empty ring?\n00:06 < stuhood> yea, essentially… there are no tokens in that dc\n00:06 < goffinet> but we were going to provide the tokens manually?\n00:06 < goffinet> were you thinking of making it automatic?\n00:07 < stuhood> yea. fixing bootstrapping groups of nodes would make automatic safe again\n00:08 < stuhood> so… whatever state a node is in when it is sitting and waiting for enough information to choose a token, it should just stay that way and watch what other nodes enter that state\n00:08 < goffinet> so i have a question about the 120 second window you have to wait..\n00:09 < stuhood> mm\n00:09 < driftx> hmm, what if they started up at rf:0 but stayed in some dead state (hibernate might work) without doing anything until you changed the rf, then actually bootstrapped?\n00:09 < goffinet> so imagine i startup all the nodes in DC2 at same time, does join_ring=false not grab gossip info at all? I was thinking it would be good if we could just start gossip on all nodes, but until operator says 'go' then i could bootstrap them all at same time\n00:09 < goffinet> since i would only have to wait at most 120 seconds before kicking them all off\n00:10 < stuhood> driftx: yea, that could work too… but you'd still need to choose tokens. (also, the rf=0 thing shouldn't be necessary, right? that's the CL.ONE bug)\n00:11 < driftx> well, you really want to choose tokens anyway\n00:11 < stuhood> goffinet: it does get gossip… i think that's basically equivalent to the pre-join state\n00:11 < driftx> I guess you don't need rf=0 if all the nodes are in hibernate\n00:12 < goffinet> yeah i think you do need hibernate in this case, because if i set tokens upfront, i want all nodes to know about ATL ones too\n00:12 < goffinet> before i kick off bootstrap\n00:12 < stuhood> driftx: i'm confused… what is the difference between rf=0 and not being there?\n00:12 < stuhood> is that a workaround for the CL.ONE bug?\n00:13 < driftx> you know there's a dc with rf:0, can add one with impacting anything\n00:13 < driftx> err, without\n00:14 ?? boaz__ (0819c319@gateway/web/freenode/ip.8.25.195.25) has joined #cassandra-dev\n00:14 < stuhood> so what was the point of adding it? that's why i'm confused...\n00:14 < goffinet> im fine with rf:0, its so you can add the nodes to the cluster before calling repair\n00:14 < goffinet> before you add nodes\n00:15 < driftx> because the dc is in the schema\n00:15 < driftx> so you need it there to have nodes be in it\n00:15 < stuhood> ah\n00:16 < goffinet> driftx: any reason why we couldnt just fix that? so dc2:3 wont throw an error if nodes are down?\n00:16 < goffinet> that way you would needed to do two steps\n00:16 < goffinet> dc2:0, add nodes, dc2:3\n00:16 < goffinet> wouldn't*\n00:16 < driftx> I don't understand, you can already do that\n00:17 < driftx> you just have to repair afterwards\n00:17 < goffinet> it throws an error currently? if you set dc2:3 and no nodes exist for dc2\n00:17 < goffinet> we'll double check on that\n00:18 < goffinet> for writes\n00:18 < driftx> oh, it does\n00:19 < driftx> but only for writes\n00:19 < goffinet> yeah\n00:19 < goffinet> so thats fine, thats fixable\n00:19 < goffinet> im just curious about a) how can we bootstrap nodes without 120s delays between N nodes b) stream from DC1 without AES\n00:21 < stuhood> goffinet: if you figure out a, i don't think b is necessary?\n00:22 < stuhood> assuming they are aware of the other joining nodes, and can all join the same range\n00:22 < stuhood> that would be the keystone for some kind of group bootstrap\n00:23 < goffinet> let me test out join_ring, because im curious. if join_ring=false still gossips but doesnt offically join.. it would be nice if node 2 in DC2 knew about that node too somehow?\n00:23 < driftx> that's why I proposed cheating, add them all as non-members, then ask them to bootstrap\n00:23 < goffinet> because then .. i could just run a command on each node at same time\n00:23 < goffinet> since they all know about each other in a hibernate state\n00:23 < goffinet> driftx: yes i like that\n00:24 < driftx> private void joinTokenRing(int delay) throws IOException, org.apache.cassandra.config.ConfigurationException\n00:24 < driftx> {\n00:24 < driftx> logger_.info(\"Starting up server gossip\");\n00:24 < driftx> they don't use gossip with join_ring off\n00:24 < stuhood> but will that actually allow them to all join the same range?\n00:24 < goffinet> okay cool, yeah we would need to make it join in that special state then\n00:25 < stuhood> i think there is an edgecase here… if multiple nodes are joining the same range, and one of them fails, then should they all fail?\n00:25 < driftx> no, it basically saves you server startup time that is not ring-related :)\n00:25 < goffinet> stuhood, they all know the tokens ahead of time?\n00:25 < goffinet> they just need to know the current global state of things\n00:25 < stuhood> goffinet: right, but if they are streaming the range that they will be responsible for...\n00:26 ?? mw1 (~Adium@8.25.195.29) has quit (Quit: Leaving.)\n00:26 < stuhood> Joining nodes don't stick around if they fail\n00:26 < goffinet> they shouldnt be allowed to do that until they joined ?\n00:26 < stuhood> nah, you stream while you are joining… unless you are talking about repair\n00:26 < goffinet> stuhood: was that removed? i thought u had to still remove the node\n00:26 < goffinet> using the new options in 1.0\n00:26 < stuhood> don't know about 1.0\n00:27 < driftx> no, a failed non-member is just a fat client and disappears\n00:27 < goffinet> but i thought there was a timeout for fat client ?\n00:27 < goffinet> is it 30s or something?\n00:27 < driftx> yes\n00:28 < goffinet> so nodes that arent fat clients, why might we remove them ? if we didnt..\n00:28 < goffinet> and let the operator do it\n00:28 < goffinet> or have a larger timeout\n00:28 < goffinet> might make this a non-issue\n00:28 < driftx> what does a larger timeout/keeping them around buy you?\n00:29 < goffinet> because if they go away, and i bootstrap after they failed, wont my view of ring be skewed?\n00:29 < stuhood> driftx: i guess in this case, the node would resume bootstrapping from where it left off\n00:29 < driftx> it would've missed writes in the meantime and require a repair afterwards anyway\n00:29 < stuhood> sorry… \"resume\" in the sense of \"start over\", but yea\n00:31 < stuhood> that would be a pretty big change, but it might make sense\n00:31 < goffinet> stuhood: what would you change\n00:31 < stuhood> what you said, about nodes in joining staying in joining\n00:31 < stuhood> so if the machine restarts, it begins joining at the same position again\n00:33 < goffinet> if we supported that + letting nodes gossip in hibernate, would allow us to add capacity at operator control\n{noformat}","from":"developer"},{"body":"I think it wouldn't be crazy and actually very (very) simple to add a new nodetool command (rebuild?) that would basically have the node asks the other replicas to stream all there data to him (for the correct ranges obviously). In other words, a command that force the streaming part of bootstrap without all the join ring part. Or another way to say is to have the streaming part of a repair but without the validation part.\n\nThe method to add a new DC would be the same as today except that repair would be replaced by this new operation.","from":"developer"},{"body":"We would still need to put the nodes into a 'bootstrap' state to get incoming writes forwarded to them, otherwise you have to repair in the end anyway.","from":"developer"},{"body":"Yeah as Brandon mentioned, we would still want to go into the bootstrap state to get those writes. This would also allow us to add capacity in the same way, if we manually pick tokens (auto bootstrap is kinda worthless IMO) to existing DC as well. We can just fire off the bootstrap command from nodetool as needed.","from":"developer"},{"body":"Sorry, I had to re-read what Sylvain said for it to 'click'. So the process he proposes is as follows with RF of 3:\n\n1. strategy options dc2:0\n2. bring up new nodes in dc2 with auto_bootstrap off and token set\n3. set strategy options dc2:3\n4. run 'rebuild' on each node in dc2\n\nthis would handle the writes part.\n\ni was kinda hoping though that we could modify the gossip state because I could very well see this playing into the case where you weren't adding DCs but wanted to add lots of nodes (60-100 like we do currently) ... and wanted to have them all added to existing DC.. having the bootstrap defined that way, would allow us to bootstrap nodes as we please in existing DCs, bring them all to the ring at once to have a consistent state without taking on traffic until they transitioned states (joining/normal). Where as this proposal wouldn't be able to satisfy that use case.\n\n","from":"developer"},{"body":"So you're proposing to add support for bootstrapping multiple nodes together. I'm not against that, it would be nice and that would give you 90% of what this ticket is about (you'd have to add the ability to multi-boostrap *and* add a DC/augment the replication faction at the same time). But it is orders of magnitude more complicated than what I'm suggesting. Which is not a problem in itself given it's a broader solution, but it means we'll have multi-node boostrap at best for 1.1, while I'm pretty sure I can get the 'rebuild' command wrote in like an hour (and I see no reason why it couldn't be put in the 1.0 series).\n\n","from":"developer"},{"body":"Sylvain, it looks like a state 'HIBERNATE' exists in GOSSIP already based on recovering a dead node in 1.0. Do you have a preference on the name of the new state if we attempted adding multiple nodes patch? WAITING? STANDBY?","from":"developer"},{"body":"I was thinking of GHOST too after I realized we'll need a state where the node receives neither writes nor reads (more details here later).","from":"developer"},{"body":"So, am I understanding correctly that we're talking about two different scenarios?\n\n- Add new DC without repair\n- Add many nodes to existing DC without RING_DELAY in between\n\nI think Sylvain's proposal addresses the first nicely. So what I need help with is understanding what problem you're trying to solve with the second part. Dealing with overlapping ranges in node movement basically requires a rewrite of that subsystem (CASSANDRA-2434). But I suspect there is a \"good enough\" solution that we could find if I understood better what your pain point is here.","from":"developer"},{"body":"Jonathan,\n\nYou are correct. Sylvain's proposal does satisfy this ticket. It doesn't solve the case of (2) where if you want to add lots of nodes in an existing DC, and you know the tokens they should be at, and want to join them all at once.\n\nOur use case is, we actually add 60-100 nodes in one big capacity add. We would like to avoid the 120 second per node time frame. It's not a deal breaker though. We actually realized though if we are adding that many nodes to our cluster, with a large cluster already, we need to rebalance heavily anyway. Peter is almost done with the 'rebuild' patch, I'm assigning him to this ticket. \n\nOur next big focus is improving the rebalancing of a cluster. We have a very large cluster and after adding 100 nodes every month or so, this becomes painful. Almost all of our nodes have over 600GB+ each. We have an application that will require us to be rebalancing at all times to reduce our hot spots.\n\n","from":"developer"},{"body":"We ended up going for the simpler rebuild patch as Chris hinted at. I'll quote myself:\n\n{quote}\nI've been looking at this some more.\n\nHere's the proposal so far:\n\nWe introduce a GHOST state, in which a node receives neither reads nor writes. This allows us to bring in a group of nodes in the ring without suffering any ill effects. It is completely invisible to reads and writes, and will never count towards e.g. consistency level.\n\nOnce all ghosts are ready, we can start bootstrapping, taking ghost nodes into accounts for purposes of determining which range we are responsible for, but streaming only from non-ghost nodes. This accomplishes the goal of not transferring more data than necessary.\n\nIn order to avoid a bootstrapping node from taking writes for more than it's eventual share, we'd have to make the write endpoints be aware of ghost nodes. This is dosable, but not critical since we're bisecting a range that was previously handled by a single node anyway so the traffic would be managable. It would just be cleaner to not have to cleanup afterwards.\n\nOnce we transition from bootstrapping to being \"up\", we have a bigger problem however. If the read paths and write paths are only aware of non-ghost nodes, the read/write paths would think that these nodes had more ownership than they really do.\n\nSo we must really be taking into account the other nodes in the read/write path as well - but only when determining ownership of a completely bootstrapped node that was previously a ghost. That means we must distinguish between a \"normally up\" node and one that's been bootstraped from ghost state (call it \"ghost strapped\").\n\nThis suddenly gets complex. The process of group bootstrap would then be:\n\nAdd a bunch of nodes in GHOST state\nBootstrap all of them, each of them going into GHOSTSTRAPPED state\nOnce all are GHOSTSTRAPPED, we can safely transntion from GHOSTSTRAPPED to normal/up\nIs there a simpler solution?\n{quote}\n\nAfter some additional discussion we felt this was adding to much complexity and potential edge cases/bugs that it became more cost-effective to just go with the simple rebuild for our immediate needs, hoping to address the problem of adding lots of nodes to a DC separately in some other way.\n\nA patch is forthcoming soon.\n","from":"developer"},{"body":"Here is a patch rebased against 0.8 for cursory review. I do not expect this to go into 0.8, and in fact I have not tested this patch other than build against vanilla 0.8 (the original patch is tested, but against our internal 0.8).\n\nIf there are no concerns with the overall implementation, I'll submit a rebased version for 1.0/trunk.\n\nThere are two components of the change:\n\n* Breaking out the streaming part of BootStrapper into a separate RangeStreamer. Change BootStrapper to use that.\n\n* Implement the rebuild command on top of RangeStreamer.\n\nThere are two ways to invoke rebuild:\n\n{code}\nnodetool rebuild\nnodetool rebuild nameofdc\n{code}\n\nThe first form streams from nearest endpoints, while the latter streams from nearest endpoints in the specified data center.\n\n","from":"developer"},{"body":"bq. If there are no concerns with the overall implementation, I'll submit a rebased version for 1.0/trunk.\n\nThis looks good to me, I like the RangeStreamer abstraction.","from":"developer"},{"body":"Attached is a version rebased against 1.0 (and tested).","from":"developer"},{"body":"Once this option is in, is this the procedure for running rebuild (with 4 changed to 'rebuild dc1')?\n\n{quote}\n1. strategy options dc2:0\n2. bring up new nodes in dc2 with auto_bootstrap off and token set\n3. set strategy options dc2:3\n4. run 'rebuild' on each node in dc2\n{quote}\n\nDo we need to stagger issuing the rebuild commands or can they be run all at once?\n","from":"developer"},{"body":"Yes, that looks good.\n\nAnd yes, you can run rebuilds concurrently as long as you're comfortable with the amount of bandwidth you'll be pushing and the load you'll be putting on the source nodes.\n\nHowever, if you expect to see reasonable performance and streaming at full speed to all nodes, you also need CASSANDRA-3494.\n\nRegardless: I strongly recommend testing this with your exact version of Cassandra before trying it for real.\n","from":"developer"},{"body":"I haven't applied the patch yet, it needs rebase and preferably against trunk since that is the likely target for this, but a few comments.\n\nWe could have more reuse of code between Boostrapper ant the rebuild command. Typically:\n* RangeStreamer.getAllRangeWithSourcesFor does essentially the same thing that Boostrapper.getRangesWithSources, so it would be nice to do some reuse.\n* In rebuild, we essentially have the code of Boostrapper.getWorkMap, again would be nice to do some code reuse.\n\nI think we should move all of those in RangeStreamer and ultimately Boostrapper.boostrap() should be just one call to rebuild with the right arguments (mostly the correct tokenMetada instance and the \"myRange\" collection).\n\nA few nits:\n* rebuild code could be simplified slightly by using StorageService.getLocalRanges()\n* rebuild doesn't fully respect the code style.\n","from":"developer"},{"body":"I'll get it rebased once it's otherwise okay.\n\nAs for re-use: I had intermediate versions that tried to do this, but ever time I ended up realizing that it was exploding in verbosity at the point where I was using the abstraction so it didn't actually help. However, I think there were a few changes towards the end after which I didn't re-evaluate.\n\nI'll look at it again and see what I can do.","from":"developer"},{"body":"I'll get it rebased once it's otherwise okay.\n\nAs for re-use: I had intermediate versions that tried to do this, but ever time I ended up realizing that it was exploding in verbosity at the point where I was using the abstraction so it didn't actually help. However, I think there were a few changes towards the end after which I didn't re-evaluate.\n\nI'll look at it again and see what I can do.","from":"developer"},{"body":"Attaching version rebased to trunk but not yet re-factored.","from":"developer"},{"body":"Peter, are you planning to follow up on Sylvain's comments still?","from":"developer"},{"body":"I do. I'm sorry for the delay, this has been nagging me for quite some time. It's not forgotten, I have just been inundated with urgent stuff to do.\n\nI'm attaching a fresh rebase against current trunk and I hope to submit an improved version later tonight (keyword being \"hope\").","from":"developer"},{"body":"{{CASSANDRA\\-3483\\-trunk\\-refactored\\-v1.txt}} addresses the duplication between BootStrapper and RangeStreamer.\n\nNext patch will address rebuild/getworkmap duplication.\n","from":"developer"},{"body":"(It also contains the addition of a brace from CASSANDRA-3806; this is intentional to avoid pain.) ","from":"developer"},{"body":"I borked the unit test, will address that too.","from":"developer"},{"body":"{{CASSANDRA\\-3483\\-trunk\\-refactored\\-v2.txt}} I believe addresses the concerns, plus makes other improvements. I'm much more happy with this one.\n\nIt addresses CASSANDRA-3807 by supporting fetch \"consistency levels\" (though only ONE is currently usable without patching), and the filtering of hosts is abstracted out.\n\nThere is still some duplication between {{Bootstrapper.bootstrap()}} and {{StorageService.rebuild()}} in that both do the dance of iteration over tables to construct the final map. I am not really feeling that abstracting away that is a good idea to include in this ticket, though I think it's worthwhile doing at some point separately.\n\nThe unit test is fixed; my adjustment of it was wrong because I wasn't picking pending ranges (in the test).\n\nI've tested both rebuild and bootstrap in a 3 node cluster.\n\nI've added some more logging than what is typically the case; there have been several cases where I wished streaming was logged in more detail at INFO, particularly when bootstrapping or rebuilding. I think it's worthwhile to get that in while at it.","from":"developer"},{"body":"bq. It addresses CASSANDRA-3807 by supporting fetch \"consistency levels\" (though only ONE is currently usable without patching)\n\nI think a first issue is that current bootstrap does not fail if no node is alive for a given range, which arguably it should. I'm good with doing that, though it would be worth backporting to 1.0 too so it may be worth splitting that to a separate patch (or rather just create one for the fix in 1.0).\n\nHowever, that does not solve the problem of bootstrap possibly breaking the consistency contract. The problem being that if we transfer a range from a node that happens to be lacking behind in term of consistency, and we end up replacing a node that was not lacking behind, we could break some consistency contracts. To fix that, I really only see only one solution right off the bat (which doesn't mean there isn't other): it is to ensure that for each range, we transfer it from (at least) the node we will replace for this range.\n\nI believe the FetchConsistencyLevel of this patch is making an attempt to fix this by allowing to fetch from more than one node. While it does make it less likely to break consistency, unless we fetch from all nodes (and thus the one we'll replace), we cannot be sure we won't break the consistency level for people that say write at CL.ALL and read at CL.ONE. Overall, I fully agree this is a problem that we should fix someone, but I'm not sure the FetchConsistencyLevel is the right solution and even if it is it's a complicated enough problem that it's worth it's own ticket. I would agree that the problem with rebuild is a little bit different, but since anyway the patch introduce FCL without using it, let's keep that for later if that's ok.\n\nbq. There is still some duplication between Bootstrapper.bootstrap() and StorageService.rebuild() in that both do the dance of iteration over tables to construct the final map. I am not really feeling that abstracting away that is a good idea to include in this ticket\n\nI think it is, at least for a good chunk of it. It's not very complicated, it clearly improves code readability and since the patch already refactor that code I don't see a good reason to push that to later, especially if we agree it's worthwhile.\n\nAttaching a v3 that 1) remove FetchConsistencyLevel for the reasons above and 2) move most of the details of creating the multimaps in RangeStreamer.\n","from":"developer"},{"body":"bq. Overall, I fully agree this is a problem that we should fix someone, but I'm not sure the FetchConsistencyLevel is the right solution and even if it is it's a complicated enough problem that it's worth it's own \n\nThis is CASSANDRA-2434 isn't it?","from":"developer"},{"body":"bq. This is CASSANDRA-2434 isn't it?\n\nit is.","from":"developer"},{"body":"cleanup patch addressing mostly typos and style. only substantial code change was to RangeStreamer.getRangeFetchMap. Also, moved OperationType.REBUILD to the end of the enum to make sure we don't break anything depending on ordinal.\n\n+1 from me otherwise.","from":"developer"},{"body":"Committed v3 + Jonathan's cleanups (and a fix to the unit test).","from":"developer"},{"body":"For the record I never intended to fix the general problem of bootstrapping never violating consistency. But in retrospect it's obvious how my choice of naming would make it sound like I did :) I agree it's a problem for its own ticket.\n\nThanks!","from":"developer"}],"created":"2011-11-11T04:59:06.000+0000","description":"Was talking to Brandon in irc, and we ran into a case where we want to bring up a new DC to an existing cluster. He suggested from jbellis the way to do it currently was set strategy options of dc2:0, then add the nodes. After the nodes are up, change the RF of dc2, and run repair. \n\nI'd like to avoid a repair as it runs AES and is a bit more intense than how bootstrap works currently by just streaming ranges from the SSTables. Would it be possible to improve this functionality (adding a new DC to existing cluster) than the proposed method? We'd be happy to do a patch if we got some input on the best way to go about it.\n","issue_id":"12531095","key":"CASSANDRA-3483","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2012-01-30T15:24:36.000+0000","role":"fixed_distractor","summary":"Support bringing up a new datacenter to existing cluster without repair"} {"case_id":"12532766","cluster":"DISTRACTOR-CASSANDRA-3532","comments":[{"body":"Looks like leveled compaction means that sstable creation can be part of the critical path now:\n\n{noformat}\n. /**\n * Discovers existing components for the descriptor. Slow: only intended for use outside the critical path.\n */\n static Set componentsFor(final Descriptor desc, final Descriptor.TempState matchState)\n{noformat}\n\nbq. Is it feasible to keep track of the temp files and just delete them rather than searching for them for each SSTable using SSTable.componentsFor()?\n\nSimplest would be to just check File.exists on the limited set of possible temp file names. Next simplest and slightly more performant would be to move the cleanup out of the finally blocks, and into a catch block: the cleanup is a no-op if everything went well.","created":"2011-11-26T07:04:37.578+0000"},{"body":"Thanks.\nDo I need to try to delete ~all~ components, for that descriptor.asTemporary()? Or just specific ones?\nA Component of type BITMAP_INDEX requires an id, so I'm not sure how I'd find this out without listing the directory contents.","created":"2011-11-27T10:21:46.678+0000"},{"body":"BITMAP_INDEX type was never completed (CASSANDRA-1472) so I'm fine with removing that code and doing whatever is simplest. We can get more sophisticated if/when we get back to working on 1472.","created":"2011-11-28T14:32:15.756+0000"},{"body":"Here's a small patch, and a few notes:\n\n- I wasn't sure how to best document the interaction with BITMAP_INDEX, hope a TODO there is ok.\n- I created a copy of descriptor.asTemporary(true). Is descriptor always guaranteed to be temporary==true? It left me a little uneasy.\n- I used SSTable.delete (while checking ahead of time that the file exists), because I wasn't sure if ordering of deletes was important.\n- SSTableReader.open() is the other method that calls SSTable.componentsFor, I only mention this because it may be a suboptimal call as well.","created":"2011-11-28T19:00:11.899+0000"},{"body":"I applied the patch in our test environment against the 1.0.4 revision, and compaction is humming along nicely now.\nOn the (otherwise idle) node I'm watching that had 250000 data/ files, file count is decreasing by 1000-1900/minute.","created":"2011-11-29T00:44:18.980+0000"},{"body":"v2 attached, with the new approach moved into componentsFor, so that open() can take advantage of the improvement too. Also, componentsFor now respects the temporary-ness of the descriptor passed, so a separate TempState enum is unnecessary.\n\nAlso renamed cleanupIfNecessary to abort, and moved to catch block as discussed above.","created":"2011-12-02T03:49:12.698+0000"},{"body":"Thank you Jonathan -- I applied the patch and it works for me.\nCheers","created":"2011-12-02T17:37:05.508+0000"},{"body":"* Wrapping exceptions into RuntimeException() blindly will confuse the catcher of UserInterruptedException in DTPE.logExceptionsAfterExecute. We could make make that catcher unwrap the full exception, but truth is I'm not a fan of wrapping exception needlessly. Maybe we could just add a silence method somewhere:\n{noformat}\npublic void silence(Exception e)\n{\n if (e instanceof RuntimeException)\n throw (RuntimeException) e;\n else\n throw new RuntimeException(e);\n}\n{noformat}\nand use that (as we already do in WrappedRunnable actually).\n* We could remove the bitmap indexes type while were at it (it'd be one less type to check).\n","created":"2011-12-02T18:45:53.121+0000"},{"body":"added FBUtilities.unchecked(Exception) as suggested, and removed BITMAP component.","created":"2011-12-02T20:22:01.139+0000"},{"body":"+1","created":"2011-12-02T20:48:23.781+0000"},{"body":"committed","created":"2011-12-02T21:01:44.049+0000"},{"body":"Patch 3532-v3.txt applied here, works for me. Thanks again!","created":"2011-12-02T21:39:54.762+0000"},{"body":"Hmm, I might have spoken too soon. This could also be a separate bug however.\n\nThe nodes in my cluster are using a lot of file descriptors, holding open tmp files. A few are using 50K+, nearing their limit (on Solaris, of 64K).\n\nHere's a small snippet of lsof:\njava 828 appdeployer *146u VREG 181,65540 0 333376 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776429-Data.db\njava 828 appdeployer *147u VREG 181,65540 0 332952 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776359-Data.db\njava 828 appdeployer *148u VREG 181,65540 0 333079 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776380-Index.db\njava 828 appdeployer *149u VREG 181,65540 0 333080 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776380-Data.db\njava 828 appdeployer *150u VREG 181,65540 0 333224 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776403-Index.db\njava 828 appdeployer *151u VREG 181,65540 0 333025 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776372-Data.db\njava 828 appdeployer *152u VREG 181,65540 0 333225 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776403-Data.db\njava 828 appdeployer *154u VREG 181,65540 0 333858 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776514-Index.db\njava 828 appdeployer *155u VREG 181,65540 0 333426 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776438-Data.db\njava 828 appdeployer *156u VREG 181,65540 0 333326 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776421-Data.db\njava 828 appdeployer *157u VREG 181,65540 0 333553 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776460-Data.db\njava 828 appdeployer *158u VREG 181,65540 0 333501 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776452-Index.db\njava 828 appdeployer *159u VREG 181,65540 0 333597 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776468-Index.db\njava 828 appdeployer *160u VREG 181,65540 0 333598 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776468-Data.db\njava 828 appdeployer *162u VREG 181,65540 0 333884 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776518-Data.db\njava 828 appdeployer *163u VREG 181,65540 0 333502 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776452-Data.db\njava 828 appdeployer *165u VREG 181,65540 0 333929 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776527-Index.db\njava 828 appdeployer *166u VREG 181,65540 0 333859 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776514-Data.db\njava 828 appdeployer *167u VREG 181,65540 0 333663 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776480-Data.db\njava 828 appdeployer *168u VREG 181,65540 0 333812 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776506-Index.db\n\nI spot checked a few and found they still exist on the filesystem too:\n-rw-r--r-- 1 appdeployer appdeployer 0 Dec 12 07:16 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776506-Index.db","created":"2011-12-12T07:56:05.628+0000"},{"body":"That does sound like a separate problem. What do you see when you grep the log for messages_meta-tmp-hb-776506?","created":"2011-12-12T13:59:43.558+0000"},{"body":"Sure. Let me know if you'd prefer a separate ticket.\n\nI don't see anything in the logs matching \"776506\". Any suggestions as to which class(es) I could turn on DEBUG log level for (via JMX), if that would help troubleshoot?","created":"2011-12-12T16:07:06.192+0000"},{"body":"Yes, separate ticket.\n\norg.apache.cassandra.db.compaction, org.apache.cassandra.db.Memtable, org.apache.cassandra.db.DataTracker to start with","created":"2011-12-12T16:48:12.132+0000"},{"body":"Separate ticket created: CASSANDRA-3616","created":"2011-12-12T22:07:38.739+0000"}],"conversations":[{"body":"From what I can tell SSTableWriter.cleanupIfNecessary seems increasingly costly as the number of files in the data dir increases.\nIt calls SSTable.componentsFor(descriptor, Descriptor.TempState.TEMP) which lists all files in the data dir to find matching components.\n\nAm I roughly correct that (cleanupCost = SSTable count * data dir size)?\n\n\nWe had been doing write load testing with default compaction throttling (16MB/s) and LeveledCompaction.\nUnfortunately we haven't been keeping tabs on sstable counts and it grew out of control.\n\nOn a system with 300,000 sstables (!) here is an example of our compaction rate. Note that as you're probably aware cleanupIfNecessary is included in the timing:\n\n INFO [CompactionExecutor:48] 2011-11-25 22:25:30,353 CompactionTask.java (line 213) Compacted to [/data1/cassandra/data/MA_DDR/indexes_03-hc-5369-Data.db,]. 5,821,590 to 5,306,354 (~91% of original) bytes for 123 keys at 0.163755MB/s. Time: 30,903ms.\n\nHere's a slightly larger one:\n INFO [CompactionExecutor:43] 2011-11-25 22:23:28,956 CompactionTask.java (line 213) Compacted to [/data1/cassandra/data/MA_DDR/indexes_03-hc-5336-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5337-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5338-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5339-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5340-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5341-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5342-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5343-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5344-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5345-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5346-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5347-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5348-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5349-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5350-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5351-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5352-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5353-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5354-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5355-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5356-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5357-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5358-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5359-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5360-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5361-Data.db,]. 140,706,512 to 137,990,868 (~98% of original) bytes for 2,181 keys at 0.338627MB/s. Time: 388,623ms.\n\n\nThis is with compaction throttling set to 0 (Off).\n\n\nSo I believe because of this it's going to take a very long time to recover from having so many small sstables. \nIt might be notable that we're using Solaris 10, possibly listFiles() is faster on other platforms?\n\nIs it feasible to keep track of the temp files and just delete them rather than searching for them for each SSTable using SSTable.componentsFor()?\n\n\n\nHere's the stack trace for the CompactionExecutor:14 thread that appears to be occupying the majority of the cpu time on this node:\n\nName: CompactionExecutor:14\nState: RUNNABLE\nTotal blocked: 3 Total waited: 1,610,714\n\nStack trace: \n java.io.UnixFileSystem.getBooleanAttributes0(Native Method)\njava.io.UnixFileSystem.getBooleanAttributes(Unknown Source)\njava.io.File.isDirectory(Unknown Source)\norg.apache.cassandra.io.sstable.SSTable$3.accept(SSTable.java:204)\njava.io.File.listFiles(Unknown Source)\norg.apache.cassandra.io.sstable.SSTable.componentsFor(SSTable.java:200)\norg.apache.cassandra.io.sstable.SSTableWriter.cleanupIfNecessary(SSTableWriter.java:289)\norg.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:189)\norg.apache.cassandra.db.compaction.LeveledCompactionTask.execute(LeveledCompactionTask.java:57)\norg.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:134)\norg.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:114)\njava.util.concurrent.FutureTask$Sync.innerRun(Unknown Source)\njava.util.concurrent.FutureTask.run(Unknown Source)\njava.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\njava.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\njava.lang.Thread.run(Unknown Source)\n\nNo matter where I click in the busy Compaction thread timeline in YourKit it's in Running state and showing this above trace, except for short periods of time where it's actually compacting :)\n\nThanks,\nEric","from":"reporter","subject":"Compaction cleanupIfNecessary costly when many files in data dir"},{"body":"Looks like leveled compaction means that sstable creation can be part of the critical path now:\n\n{noformat}\n. /**\n * Discovers existing components for the descriptor. Slow: only intended for use outside the critical path.\n */\n static Set componentsFor(final Descriptor desc, final Descriptor.TempState matchState)\n{noformat}\n\nbq. Is it feasible to keep track of the temp files and just delete them rather than searching for them for each SSTable using SSTable.componentsFor()?\n\nSimplest would be to just check File.exists on the limited set of possible temp file names. Next simplest and slightly more performant would be to move the cleanup out of the finally blocks, and into a catch block: the cleanup is a no-op if everything went well.","from":"developer"},{"body":"Thanks.\nDo I need to try to delete ~all~ components, for that descriptor.asTemporary()? Or just specific ones?\nA Component of type BITMAP_INDEX requires an id, so I'm not sure how I'd find this out without listing the directory contents.","from":"developer"},{"body":"BITMAP_INDEX type was never completed (CASSANDRA-1472) so I'm fine with removing that code and doing whatever is simplest. We can get more sophisticated if/when we get back to working on 1472.","from":"developer"},{"body":"Here's a small patch, and a few notes:\n\n- I wasn't sure how to best document the interaction with BITMAP_INDEX, hope a TODO there is ok.\n- I created a copy of descriptor.asTemporary(true). Is descriptor always guaranteed to be temporary==true? It left me a little uneasy.\n- I used SSTable.delete (while checking ahead of time that the file exists), because I wasn't sure if ordering of deletes was important.\n- SSTableReader.open() is the other method that calls SSTable.componentsFor, I only mention this because it may be a suboptimal call as well.","from":"developer"},{"body":"I applied the patch in our test environment against the 1.0.4 revision, and compaction is humming along nicely now.\nOn the (otherwise idle) node I'm watching that had 250000 data/ files, file count is decreasing by 1000-1900/minute.","from":"developer"},{"body":"v2 attached, with the new approach moved into componentsFor, so that open() can take advantage of the improvement too. Also, componentsFor now respects the temporary-ness of the descriptor passed, so a separate TempState enum is unnecessary.\n\nAlso renamed cleanupIfNecessary to abort, and moved to catch block as discussed above.","from":"developer"},{"body":"Thank you Jonathan -- I applied the patch and it works for me.\nCheers","from":"developer"},{"body":"* Wrapping exceptions into RuntimeException() blindly will confuse the catcher of UserInterruptedException in DTPE.logExceptionsAfterExecute. We could make make that catcher unwrap the full exception, but truth is I'm not a fan of wrapping exception needlessly. Maybe we could just add a silence method somewhere:\n{noformat}\npublic void silence(Exception e)\n{\n if (e instanceof RuntimeException)\n throw (RuntimeException) e;\n else\n throw new RuntimeException(e);\n}\n{noformat}\nand use that (as we already do in WrappedRunnable actually).\n* We could remove the bitmap indexes type while were at it (it'd be one less type to check).\n","from":"developer"},{"body":"added FBUtilities.unchecked(Exception) as suggested, and removed BITMAP component.","from":"developer"},{"body":"+1","from":"developer"},{"body":"committed","from":"developer"},{"body":"Patch 3532-v3.txt applied here, works for me. Thanks again!","from":"developer"},{"body":"Hmm, I might have spoken too soon. This could also be a separate bug however.\n\nThe nodes in my cluster are using a lot of file descriptors, holding open tmp files. A few are using 50K+, nearing their limit (on Solaris, of 64K).\n\nHere's a small snippet of lsof:\njava 828 appdeployer *146u VREG 181,65540 0 333376 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776429-Data.db\njava 828 appdeployer *147u VREG 181,65540 0 332952 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776359-Data.db\njava 828 appdeployer *148u VREG 181,65540 0 333079 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776380-Index.db\njava 828 appdeployer *149u VREG 181,65540 0 333080 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776380-Data.db\njava 828 appdeployer *150u VREG 181,65540 0 333224 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776403-Index.db\njava 828 appdeployer *151u VREG 181,65540 0 333025 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776372-Data.db\njava 828 appdeployer *152u VREG 181,65540 0 333225 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776403-Data.db\njava 828 appdeployer *154u VREG 181,65540 0 333858 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776514-Index.db\njava 828 appdeployer *155u VREG 181,65540 0 333426 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776438-Data.db\njava 828 appdeployer *156u VREG 181,65540 0 333326 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776421-Data.db\njava 828 appdeployer *157u VREG 181,65540 0 333553 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776460-Data.db\njava 828 appdeployer *158u VREG 181,65540 0 333501 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776452-Index.db\njava 828 appdeployer *159u VREG 181,65540 0 333597 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776468-Index.db\njava 828 appdeployer *160u VREG 181,65540 0 333598 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776468-Data.db\njava 828 appdeployer *162u VREG 181,65540 0 333884 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776518-Data.db\njava 828 appdeployer *163u VREG 181,65540 0 333502 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776452-Data.db\njava 828 appdeployer *165u VREG 181,65540 0 333929 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776527-Index.db\njava 828 appdeployer *166u VREG 181,65540 0 333859 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776514-Data.db\njava 828 appdeployer *167u VREG 181,65540 0 333663 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776480-Data.db\njava 828 appdeployer *168u VREG 181,65540 0 333812 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776506-Index.db\n\nI spot checked a few and found they still exist on the filesystem too:\n-rw-r--r-- 1 appdeployer appdeployer 0 Dec 12 07:16 /data1/cassandra/data/MA_DDR/messages_meta-tmp-hb-776506-Index.db","from":"developer"},{"body":"That does sound like a separate problem. What do you see when you grep the log for messages_meta-tmp-hb-776506?","from":"developer"},{"body":"Sure. Let me know if you'd prefer a separate ticket.\n\nI don't see anything in the logs matching \"776506\". Any suggestions as to which class(es) I could turn on DEBUG log level for (via JMX), if that would help troubleshoot?","from":"developer"},{"body":"Yes, separate ticket.\n\norg.apache.cassandra.db.compaction, org.apache.cassandra.db.Memtable, org.apache.cassandra.db.DataTracker to start with","from":"developer"},{"body":"Separate ticket created: CASSANDRA-3616","from":"developer"}],"created":"2011-11-25T22:58:24.000+0000","description":"From what I can tell SSTableWriter.cleanupIfNecessary seems increasingly costly as the number of files in the data dir increases.\nIt calls SSTable.componentsFor(descriptor, Descriptor.TempState.TEMP) which lists all files in the data dir to find matching components.\n\nAm I roughly correct that (cleanupCost = SSTable count * data dir size)?\n\n\nWe had been doing write load testing with default compaction throttling (16MB/s) and LeveledCompaction.\nUnfortunately we haven't been keeping tabs on sstable counts and it grew out of control.\n\nOn a system with 300,000 sstables (!) here is an example of our compaction rate. Note that as you're probably aware cleanupIfNecessary is included in the timing:\n\n INFO [CompactionExecutor:48] 2011-11-25 22:25:30,353 CompactionTask.java (line 213) Compacted to [/data1/cassandra/data/MA_DDR/indexes_03-hc-5369-Data.db,]. 5,821,590 to 5,306,354 (~91% of original) bytes for 123 keys at 0.163755MB/s. Time: 30,903ms.\n\nHere's a slightly larger one:\n INFO [CompactionExecutor:43] 2011-11-25 22:23:28,956 CompactionTask.java (line 213) Compacted to [/data1/cassandra/data/MA_DDR/indexes_03-hc-5336-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5337-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5338-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5339-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5340-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5341-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5342-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5343-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5344-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5345-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5346-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5347-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5348-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5349-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5350-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5351-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5352-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5353-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5354-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5355-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5356-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5357-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5358-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5359-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5360-Data.db,/data1/cassandra/data/MA_DDR/indexes_03-hc-5361-Data.db,]. 140,706,512 to 137,990,868 (~98% of original) bytes for 2,181 keys at 0.338627MB/s. Time: 388,623ms.\n\n\nThis is with compaction throttling set to 0 (Off).\n\n\nSo I believe because of this it's going to take a very long time to recover from having so many small sstables. \nIt might be notable that we're using Solaris 10, possibly listFiles() is faster on other platforms?\n\nIs it feasible to keep track of the temp files and just delete them rather than searching for them for each SSTable using SSTable.componentsFor()?\n\n\n\nHere's the stack trace for the CompactionExecutor:14 thread that appears to be occupying the majority of the cpu time on this node:\n\nName: CompactionExecutor:14\nState: RUNNABLE\nTotal blocked: 3 Total waited: 1,610,714\n\nStack trace: \n java.io.UnixFileSystem.getBooleanAttributes0(Native Method)\njava.io.UnixFileSystem.getBooleanAttributes(Unknown Source)\njava.io.File.isDirectory(Unknown Source)\norg.apache.cassandra.io.sstable.SSTable$3.accept(SSTable.java:204)\njava.io.File.listFiles(Unknown Source)\norg.apache.cassandra.io.sstable.SSTable.componentsFor(SSTable.java:200)\norg.apache.cassandra.io.sstable.SSTableWriter.cleanupIfNecessary(SSTableWriter.java:289)\norg.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:189)\norg.apache.cassandra.db.compaction.LeveledCompactionTask.execute(LeveledCompactionTask.java:57)\norg.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:134)\norg.apache.cassandra.db.compaction.CompactionManager$1.call(CompactionManager.java:114)\njava.util.concurrent.FutureTask$Sync.innerRun(Unknown Source)\njava.util.concurrent.FutureTask.run(Unknown Source)\njava.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\njava.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\njava.lang.Thread.run(Unknown Source)\n\nNo matter where I click in the busy Compaction thread timeline in YourKit it's in Running state and showing this above trace, except for short periods of time where it's actually compacting :)\n\nThanks,\nEric","issue_id":"12532766","key":"CASSANDRA-3532","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-12-02T21:01:44.000+0000","role":"fixed_distractor","summary":"Compaction cleanupIfNecessary costly when many files in data dir"} {"case_id":"12532941","cluster":"DISTRACTOR-CASSANDRA-3536","comments":[{"body":"I may have a similar, or the same problem.\n\nOn a node where I was running repair (with no other load), I am seeing these exceptions:\n\nERROR [Thread-90] 2011-11-29 17:44:22,753 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[Thread-90,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.compaction.LeveledManifest.promote(LeveledManifest.java:178)\n at org.apache.cassandra.db.compaction.LeveledCompactionStrategy.handleNotification(LeveledCompactionStrategy.java:141)\n at org.apache.cassandra.db.DataTracker.notifySSTablesChanged(DataTracker.java:481)\n at org.apache.cassandra.db.DataTracker.replace(DataTracker.java:275)\n at org.apache.cassandra.db.DataTracker.addSSTables(DataTracker.java:237)\n at org.apache.cassandra.db.DataTracker.addStreamedSSTable(DataTracker.java:242)\n at org.apache.cassandra.db.ColumnFamilyStore.addSSTable(ColumnFamilyStore.java:920)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:141)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:103)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:184)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:81)","created":"2011-11-29T18:28:19.930+0000"},{"body":"LeveledManifest.promote only really knows how to handle compaction -- swapping out one set of sstables, for another. When adding a new sstable, we need to call LM.add instead. Refactored DataTracker to do this correctly for streaming (and load-new-sstables-from-jmx); in the process, got rid of addStreamedSSTable (becomes just addSSTables(singletonlist) and added addInitialSSTables, to add to the tracker with no notifications (since Strategy creation loads the manifest separately).","created":"2011-12-02T05:30:47.884+0000"},{"body":"+1","created":"2011-12-02T05:52:06.386+0000"},{"body":"Thanks Jonathan, the patch works for me.","created":"2011-12-02T15:32:36.672+0000"},{"body":"committed","created":"2011-12-02T15:47:18.010+0000"}],"conversations":[{"body":" I have a 3 node cassandra cluster. I have RF set to 3 and do reads\nand writes using QUORUM.\n\nHere is my initial ring configuration\n\n[root@CAP4-CNode1 ~]# /root/cassandra/bin/nodetool -h localhost ring\nAddress DC Rack Status State Load\nOwns Token\n\n 113427455640312821154458202477256070484\n10.19.104.11 datacenter1 rack1 Up Normal 1.66 GB\n33.33% 0\n10.19.104.12 datacenter1 rack1 Up Normal 1.06 GB\n33.33% 56713727820156410577229101238628035242\n10.19.104.13 datacenter1 rack1 Up Normal 1.61 GB\n33.33% 113427455640312821154458202477256070484\n\nI want to add 10.19.104.14 to the cluster.\n\nI edited the 10.19.104.14 cassandra.yaml file and set the token to\n127605887595351923798765477786913079296 and set auto_bootstrap to\ntrue.\n\nWhen I started cassandra I am getting Assertion Error. \n\nthanks\nRamesh\n\n\n\n\n[root@CAP4-CNode4 cassandra]# INFO 10:29:46,093 Logging initialized\n INFO 10:29:46,099 JVM vendor/version: Java HotSpot(TM) 64-Bit Server\nVM/1.6.0_25\n INFO 10:29:46,100 Heap size: 8304721920/8304721920\n INFO 10:29:46,100 Classpath:\nbin/../conf:bin/../build/classes/main:bin/../build/classes/thrift:bin/../lib/antlr-3.2.jar:bin/../lib/apache-cassandra-1.0.2.jar:bin/../lib/apache-cassandra-clientutil-1.0.2.jar:bin/../lib/apache-cassandra-thrift-1.0.2.jar:bin/../lib/avro-1.4.0-fixes.jar:bin/../lib/avro-1.4.0-sources-fixes.jar:bin/../lib/commons-cli-1.1.jar:bin/../lib/commons-codec-1.2.jar:bin/../lib/commons-lang-2.4.jar:bin/../lib/compress-lzf-0.8.4.jar:bin/../lib/concurrentlinkedhashmap-lru-1.2.jar:bin/../lib/guava-r08.jar:bin/../lib/high-scale-lib-1.1.2.jar:bin/../lib/jackson-core-asl-1.4.0.jar:bin/../lib/jackson-mapper-asl-1.4.0.jar:bin/../lib/jamm-0.2.5.jar:bin/../lib/jline-0.9.94.jar:bin/../lib/jna.jar:bin/../lib/json-simple-1.1.jar:bin/../lib/libthrift-0.6.jar:bin/../lib/log4j-1.2.16.jar:bin/../lib/mx4j-examples.jar:bin/../lib/mx4j-impl.jar:bin/../lib/mx4j.jar:bin/../lib/mx4j-jmx.jar:bin/../lib/mx4j-remote.jar:bin/../lib/mx4j-rimpl.jar:bin/../lib/mx4j-rjmx.jar:bin/../lib/mx4j-tools.jar:bin/../lib/servlet-api-2.5-20081211.jar:bin/../lib/slf4j-api-1.6.1.jar:bin/../lib/slf4j-log4j12-1.6.1.jar:bin/../lib/snakeyaml-1.6.jar:bin/../lib/snappy-java-1.0.4.1.jar:bin/../lib/jamm-0.2.5.jar\n INFO 10:29:48,713 JNA mlockall successful\n INFO 10:29:48,726 Loading settings from\nfile:/root/apache-cassandra-1.0.2/conf/cassandra.yaml\n INFO 10:29:48,883 DiskAccessMode 'auto' determined to be mmap,\nindexAccessMode is mmap\n INFO 10:29:48,898 Global memtable threshold is enabled at 2640MB\n INFO 10:29:49,203 Couldn't detect any schema definitions in local storage.\n INFO 10:29:49,204 Found table data in data directories. Consider\nusing the CLI to define your schema.\n INFO 10:29:49,220 Creating new commitlog segment\n/var/lib/cassandra/commitlog/CommitLog-1321979389220.log\n INFO 10:29:49,227 No commitlog files found; skipping replay\n INFO 10:29:49,230 Cassandra version: 1.0.2\n INFO 10:29:49,230 Thrift API version: 19.18.0\n INFO 10:29:49,230 Loading persisted ring state\n INFO 10:29:49,235 Starting up server gossip\n INFO 10:29:49,259 Enqueuing flush of\nMemtable-LocationInfo@122130810(192/240 serialized/live bytes, 4 ops)\n INFO 10:29:49,260 Writing Memtable-LocationInfo@122130810(192/240\nserialized/live bytes, 4 ops)\n INFO 10:29:49,317 Completed flushing\n/var/lib/cassandra/data/system/LocationInfo-h-1-Data.db (300 bytes)\n INFO 10:29:49,340 Starting Messaging Service on port 7000\n INFO 10:29:49,349 JOINING: waiting for ring and schema information\n INFO 10:29:50,759 Applying migration\n4b0e20f0-1511-11e1-0000-c11bc95834d7 Add keyspace: MSA, rep\nstrategy:SimpleStrategy{}, durable_writes: true\n INFO 10:29:50,761 Enqueuing flush of\nMemtable-Migrations@1507565381(6744/8430 serialized/live bytes, 1 ops)\n INFO 10:29:50,761 Writing Memtable-Migrations@1507565381(6744/8430\nserialized/live bytes, 1 ops)\n INFO 10:29:50,761 Enqueuing flush of\nMemtable-Schema@1498835564(2889/3611 serialized/live bytes, 3 ops)\n INFO 10:29:50,776 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-1-Data.db (6808 bytes)\n INFO 10:29:50,777 Writing Memtable-Schema@1498835564(2889/3611\nserialized/live bytes, 3 ops)\n INFO 10:29:50,797 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-1-Data.db (3039 bytes)\n INFO 10:29:50,814 Applying migration\n4b6f2cb0-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@1639d811[cfId=1000,ksName=MSA,cfName=modseq,cfType=Standard,comparator=org.apache.cassandra.db.marshal.ReversedType(org.apache.cassandra.db.marshal.BytesType),subcolumncomparator=,comment=,rowCacheSize=0.0,keyCacheSize=5000000.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=0,keyCacheSavePeriodInSeconds=14400,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@2f984f7d,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.LeveledCompactionStrategy,compactionStrategyOptions={sstable_size_in_mb=10},compressionOptions={}]\n INFO 10:29:50,815 Enqueuing flush of\nMemtable-Migrations@948613108(7482/9352 serialized/live bytes, 1 ops)\n INFO 10:29:50,816 Writing Memtable-Migrations@948613108(7482/9352\nserialized/live bytes, 1 ops)\n INFO 10:29:50,816 Enqueuing flush of\nMemtable-Schema@421910828(3294/4117 serialized/live bytes, 3 ops)\n INFO 10:29:50,831 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-2-Data.db (7546 bytes)\n INFO 10:29:50,832 Writing Memtable-Schema@421910828(3294/4117\nserialized/live bytes, 3 ops)\n INFO 10:29:50,846 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-2-Data.db (3444 bytes)\n INFO 10:29:50,854 Applying migration\n4b8c9fc0-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@1bd97d0d[cfId=1001,ksName=MSA,cfName=msgid,cfType=Standard,comparator=org.apache.cassandra.db.marshal.BytesType,subcolumncomparator=,comment=,rowCacheSize=0.0,keyCacheSize=1000000.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=0,keyCacheSavePeriodInSeconds=14400,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@63a0eec3,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.SizeTieredCompactionStrategy,compactionStrategyOptions={},compressionOptions={}]\n INFO 10:29:50,855 Enqueuing flush of\nMemtable-Migrations@1520138062(7750/9687 serialized/live bytes, 1 ops)\n INFO 10:29:50,856 Writing Memtable-Migrations@1520138062(7750/9687\nserialized/live bytes, 1 ops)\n INFO 10:29:50,856 Enqueuing flush of\nMemtable-Schema@347459675(3630/4537 serialized/live bytes, 3 ops)\n INFO 10:29:50,878 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-3-Data.db (7814 bytes)\n INFO 10:29:50,879 Writing Memtable-Schema@347459675(3630/4537\nserialized/live bytes, 3 ops)\n INFO 10:29:50,894 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-3-Data.db (3780 bytes)\n INFO 10:29:50,900 Applying migration\n4ba1ae60-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@6a095b8a[cfId=1002,ksName=MSA,cfName=participants,cfType=Standard,comparator=org.apache.cassandra.db.marshal.ReversedType(org.apache.cassandra.db.marshal.BytesType),subcolumncomparator=,comment=,rowCacheSize=0.0,keyCacheSize=1000000.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=0,keyCacheSavePeriodInSeconds=14400,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@c58f769,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.SizeTieredCompactionStrategy,compactionStrategyOptions={},compressionOptions={}]\n INFO 10:29:50,900 Enqueuing flush of\nMemtable-Migrations@618337492(8194/10242 serialized/live bytes, 1 ops)\n INFO 10:29:50,901 Writing Memtable-Migrations@618337492(8194/10242\nserialized/live bytes, 1 ops)\n INFO 10:29:50,902 Enqueuing flush of\nMemtable-Schema@724860211(4020/5025 serialized/live bytes, 3 ops)\n INFO 10:29:50,917 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-4-Data.db (8258 bytes)\n INFO 10:29:50,918 Writing Memtable-Schema@724860211(4020/5025\nserialized/live bytes, 3 ops)\n INFO 10:29:50,925 Compacting\n[SSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-1-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-2-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-4-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-3-Data.db')]\n INFO 10:29:50,934 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-4-Data.db (4170 bytes)\n INFO 10:29:50,935 Compacting\n[SSTableReader(path='/var/lib/cassandra/data/system/Schema-h-2-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-1-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-4-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-3-Data.db')]\n INFO 10:29:50,940 Applying migration\n4bb4e840-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@318c69a9[cfId=1003,ksName=MSA,cfName=subinfo,cfType=Standard,comparator=org.apache.cassandra.db.marshal.ReversedType(org.apache.cassandra.db.marshal.BytesType),subcolumncomparator=,comment=,rowCacheSize=5000.0,keyCacheSize=5000000.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=14400,keyCacheSavePeriodInSeconds=14400,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@796cefa8,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.LeveledCompactionStrategy,compactionStrategyOptions={sstable_size_in_mb=10},compressionOptions={}]\n INFO 10:29:50,941 Enqueuing flush of\nMemtable-Migrations@1682081063(8618/10772 serialized/live bytes, 1\nops)\n INFO 10:29:50,941 Writing Memtable-Migrations@1682081063(8618/10772\nserialized/live bytes, 1 ops)\n INFO 10:29:50,941 Enqueuing flush of\nMemtable-Schema@1083461053(4427/5533 serialized/live bytes, 3 ops)\n INFO 10:29:50,977 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-5-Data.db (8682 bytes)\n INFO 10:29:50,978 Writing Memtable-Schema@1083461053(4427/5533\nserialized/live bytes, 3 ops)\n INFO 10:29:50,991 Compacted to\n[/var/lib/cassandra/data/system/Schema-h-5-Data.db,]. 14,433 to\n14,106 (~97% of original) bytes for 5 keys at 0.269051MB/s. Time:\n50ms.\n INFO 10:29:50,995 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-7-Data.db (4577 bytes)\n INFO 10:29:51,000 Applying migration\n4bc6e9a0-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@20b00ec2[cfId=1004,ksName=MSA,cfName=transactions,cfType=Standard,comparator=org.apache.cassandra.db.marshal.ReversedType(org.apache.cassandra.db.marshal.BytesType),subcolumncomparator=,comment=,rowCacheSize=0.0,keyCacheSize=0.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=0,keyCacheSavePeriodInSeconds=0,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@698f352,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.LeveledCompactionStrategy,compactionStrategyOptions={sstable_size_in_mb=10},compressionOptions={}]\n INFO 10:29:51,001 Enqueuing flush of\nMemtable-Migrations@596545504(9027/11283 serialized/live bytes, 1 ops)\n INFO 10:29:51,002 Writing Memtable-Migrations@596545504(9027/11283\nserialized/live bytes, 1 ops)\n INFO 10:29:51,003 Enqueuing flush of\nMemtable-Schema@1686621532(4835/6043 serialized/live bytes, 3 ops)\n INFO 10:29:51,029 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-7-Data.db (9091 bytes)\n INFO 10:29:51,029 Writing Memtable-Schema@1686621532(4835/6043\nserialized/live bytes, 3 ops)\n INFO 10:29:51,031 Compacted to\n[/var/lib/cassandra/data/system/Migrations-h-6-Data.db,]. 30,426 to\n30,234 (~99% of original) bytes for 1 keys at 0.272013MB/s. Time:\n106ms.\n INFO 10:29:51,044 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-8-Data.db (4985 bytes)\n INFO 10:29:51,049 Applying migration\n4bd76460-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@4ab4faeb[cfId=1005,ksName=MSA,cfName=uid,cfType=Standard,comparator=org.apache.cassandra.db.marshal.ReversedType(org.apache.cassandra.db.marshal.BytesType),subcolumncomparator=,comment=,rowCacheSize=0.0,keyCacheSize=1500000.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=0,keyCacheSavePeriodInSeconds=14400,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@2fc5809e,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.LeveledCompactionStrategy,compactionStrategyOptions={sstable_size_in_mb=10},compressionOptions={}]\n INFO 10:29:51,050 Enqueuing flush of\nMemtable-Migrations@1333730706(9421/11776 serialized/live bytes, 1\nops)\n INFO 10:29:51,050 Writing Memtable-Migrations@1333730706(9421/11776\nserialized/live bytes, 1 ops)\n INFO 10:29:51,051 Enqueuing flush of\nMemtable-Schema@577668356(5236/6545 serialized/live bytes, 3 ops)\n INFO 10:29:51,065 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-9-Data.db (9485 bytes)\n INFO 10:29:51,066 Compacting\n[SSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-6-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-9-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-7-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-5-Data.db')]\n INFO 10:29:51,066 Writing Memtable-Schema@577668356(5236/6545\nserialized/live bytes, 3 ops)\n INFO 10:29:51,081 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-9-Data.db (5386 bytes)\n INFO 10:29:51,083 Compacting\n[SSTableReader(path='/var/lib/cassandra/data/system/Schema-h-5-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-9-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-8-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-7-Data.db')]\n INFO 10:29:51,114 Compacted to\n[/var/lib/cassandra/data/system/Schema-h-10-Data.db,]. 29,054 to\n28,727 (~98% of original) bytes for 8 keys at 0.913207MB/s. Time:\n30ms.\n INFO 10:29:51,144 Compacted to\n[/var/lib/cassandra/data/system/Migrations-h-10-Data.db,]. 57,492 to\n57,300 (~99% of original) bytes for 1 keys at 0.700584MB/s. Time:\n78ms.\n INFO 10:29:51,410 Node /10.19.104.13 is now part of the cluster\n INFO 10:29:51,412 InetAddress /10.19.104.13 is now UP\n INFO 10:29:51,414 Enqueuing flush of\nMemtable-LocationInfo@709342045(35/43 serialized/live bytes, 1 ops)\n INFO 10:29:51,415 Writing Memtable-LocationInfo@709342045(35/43\nserialized/live bytes, 1 ops)\n INFO 10:29:51,428 Completed flushing\n/var/lib/cassandra/data/system/LocationInfo-h-2-Data.db (89 bytes)\n INFO 10:29:51,439 Node /10.19.104.12 is now part of the cluster\n INFO 10:29:51,439 InetAddress /10.19.104.12 is now UP\n INFO 10:29:51,441 Enqueuing flush of\nMemtable-LocationInfo@1292444743(35/43 serialized/live bytes, 1 ops)\n INFO 10:29:51,441 Writing Memtable-LocationInfo@1292444743(35/43\nserialized/live bytes, 1 ops)\n INFO 10:29:51,455 Completed flushing\n/var/lib/cassandra/data/system/LocationInfo-h-3-Data.db (89 bytes)\n INFO 10:29:51,456 Node /10.19.104.11 is now part of the cluster\n INFO 10:29:51,457 InetAddress /10.19.104.11 is now UP\n INFO 10:29:51,459 Enqueuing flush of\nMemtable-LocationInfo@1891328597(20/25 serialized/live bytes, 1 ops)\n INFO 10:29:51,459 Writing Memtable-LocationInfo@1891328597(20/25\nserialized/live bytes, 1 ops)\n INFO 10:29:51,471 Completed flushing\n/var/lib/cassandra/data/system/LocationInfo-h-4-Data.db (74 bytes)\n INFO 10:29:51,473 Compacting\n[SSTableReader(path='/var/lib/cassandra/data/system/LocationInfo-h-2-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/LocationInfo-h-4-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/LocationInfo-h-1-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/LocationInfo-h-3-Data.db')]\n INFO 10:29:51,497 Compacted to\n[/var/lib/cassandra/data/system/LocationInfo-h-5-Data.db,]. 552 to\n444 (~80% of original) bytes for 3 keys at 0.018410MB/s. Time: 23ms.\n INFO 10:30:19,349 JOINING: getting bootstrap token\n INFO 10:30:19,352 Enqueuing flush of\nMemtable-LocationInfo@225265367(36/45 serialized/live bytes, 1 ops)\n INFO 10:30:19,353 Writing Memtable-LocationInfo@225265367(36/45\nserialized/live bytes, 1 ops)\n INFO 10:30:19,364 Completed flushing\n/var/lib/cassandra/data/system/LocationInfo-h-7-Data.db (87 bytes)\n INFO 10:30:19,374 JOINING: sleeping 30000 ms for pending range setup\n INFO 10:30:49,375 JOINING: Starting to bootstrap...\nERROR 10:31:13,444 Fatal exception in thread Thread[Thread-49,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.compaction.LeveledManifest.promote(LeveledManifest.java:178)\n at org.apache.cassandra.db.compaction.LeveledCompactionStrategy.handleNotification(LeveledCompactionStrategy.java:141)\n at org.apache.cassandra.db.DataTracker.notifySSTablesChanged(DataTracker.java:466)\n at org.apache.cassandra.db.DataTracker.replace(DataTracker.java:275)\n at org.apache.cassandra.db.DataTracker.addSSTables(DataTracker.java:237)\n at org.apache.cassandra.db.DataTracker.addStreamedSSTable(DataTracker.java:242)\n at org.apache.cassandra.db.ColumnFamilyStore.addSSTable(ColumnFamilyStore.java:922)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:141)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:102)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:184)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:81)","from":"reporter","subject":"Assertion error during bootstraping cassandra"},{"body":"I may have a similar, or the same problem.\n\nOn a node where I was running repair (with no other load), I am seeing these exceptions:\n\nERROR [Thread-90] 2011-11-29 17:44:22,753 AbstractCassandraDaemon.java (line 133) Fatal exception in thread Thread[Thread-90,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.compaction.LeveledManifest.promote(LeveledManifest.java:178)\n at org.apache.cassandra.db.compaction.LeveledCompactionStrategy.handleNotification(LeveledCompactionStrategy.java:141)\n at org.apache.cassandra.db.DataTracker.notifySSTablesChanged(DataTracker.java:481)\n at org.apache.cassandra.db.DataTracker.replace(DataTracker.java:275)\n at org.apache.cassandra.db.DataTracker.addSSTables(DataTracker.java:237)\n at org.apache.cassandra.db.DataTracker.addStreamedSSTable(DataTracker.java:242)\n at org.apache.cassandra.db.ColumnFamilyStore.addSSTable(ColumnFamilyStore.java:920)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:141)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:103)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:184)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:81)","from":"developer"},{"body":"LeveledManifest.promote only really knows how to handle compaction -- swapping out one set of sstables, for another. When adding a new sstable, we need to call LM.add instead. Refactored DataTracker to do this correctly for streaming (and load-new-sstables-from-jmx); in the process, got rid of addStreamedSSTable (becomes just addSSTables(singletonlist) and added addInitialSSTables, to add to the tracker with no notifications (since Strategy creation loads the manifest separately).","from":"developer"},{"body":"+1","from":"developer"},{"body":"Thanks Jonathan, the patch works for me.","from":"developer"},{"body":"committed","from":"developer"}],"created":"2011-11-28T17:44:23.000+0000","description":" I have a 3 node cassandra cluster. I have RF set to 3 and do reads\nand writes using QUORUM.\n\nHere is my initial ring configuration\n\n[root@CAP4-CNode1 ~]# /root/cassandra/bin/nodetool -h localhost ring\nAddress DC Rack Status State Load\nOwns Token\n\n 113427455640312821154458202477256070484\n10.19.104.11 datacenter1 rack1 Up Normal 1.66 GB\n33.33% 0\n10.19.104.12 datacenter1 rack1 Up Normal 1.06 GB\n33.33% 56713727820156410577229101238628035242\n10.19.104.13 datacenter1 rack1 Up Normal 1.61 GB\n33.33% 113427455640312821154458202477256070484\n\nI want to add 10.19.104.14 to the cluster.\n\nI edited the 10.19.104.14 cassandra.yaml file and set the token to\n127605887595351923798765477786913079296 and set auto_bootstrap to\ntrue.\n\nWhen I started cassandra I am getting Assertion Error. \n\nthanks\nRamesh\n\n\n\n\n[root@CAP4-CNode4 cassandra]# INFO 10:29:46,093 Logging initialized\n INFO 10:29:46,099 JVM vendor/version: Java HotSpot(TM) 64-Bit Server\nVM/1.6.0_25\n INFO 10:29:46,100 Heap size: 8304721920/8304721920\n INFO 10:29:46,100 Classpath:\nbin/../conf:bin/../build/classes/main:bin/../build/classes/thrift:bin/../lib/antlr-3.2.jar:bin/../lib/apache-cassandra-1.0.2.jar:bin/../lib/apache-cassandra-clientutil-1.0.2.jar:bin/../lib/apache-cassandra-thrift-1.0.2.jar:bin/../lib/avro-1.4.0-fixes.jar:bin/../lib/avro-1.4.0-sources-fixes.jar:bin/../lib/commons-cli-1.1.jar:bin/../lib/commons-codec-1.2.jar:bin/../lib/commons-lang-2.4.jar:bin/../lib/compress-lzf-0.8.4.jar:bin/../lib/concurrentlinkedhashmap-lru-1.2.jar:bin/../lib/guava-r08.jar:bin/../lib/high-scale-lib-1.1.2.jar:bin/../lib/jackson-core-asl-1.4.0.jar:bin/../lib/jackson-mapper-asl-1.4.0.jar:bin/../lib/jamm-0.2.5.jar:bin/../lib/jline-0.9.94.jar:bin/../lib/jna.jar:bin/../lib/json-simple-1.1.jar:bin/../lib/libthrift-0.6.jar:bin/../lib/log4j-1.2.16.jar:bin/../lib/mx4j-examples.jar:bin/../lib/mx4j-impl.jar:bin/../lib/mx4j.jar:bin/../lib/mx4j-jmx.jar:bin/../lib/mx4j-remote.jar:bin/../lib/mx4j-rimpl.jar:bin/../lib/mx4j-rjmx.jar:bin/../lib/mx4j-tools.jar:bin/../lib/servlet-api-2.5-20081211.jar:bin/../lib/slf4j-api-1.6.1.jar:bin/../lib/slf4j-log4j12-1.6.1.jar:bin/../lib/snakeyaml-1.6.jar:bin/../lib/snappy-java-1.0.4.1.jar:bin/../lib/jamm-0.2.5.jar\n INFO 10:29:48,713 JNA mlockall successful\n INFO 10:29:48,726 Loading settings from\nfile:/root/apache-cassandra-1.0.2/conf/cassandra.yaml\n INFO 10:29:48,883 DiskAccessMode 'auto' determined to be mmap,\nindexAccessMode is mmap\n INFO 10:29:48,898 Global memtable threshold is enabled at 2640MB\n INFO 10:29:49,203 Couldn't detect any schema definitions in local storage.\n INFO 10:29:49,204 Found table data in data directories. Consider\nusing the CLI to define your schema.\n INFO 10:29:49,220 Creating new commitlog segment\n/var/lib/cassandra/commitlog/CommitLog-1321979389220.log\n INFO 10:29:49,227 No commitlog files found; skipping replay\n INFO 10:29:49,230 Cassandra version: 1.0.2\n INFO 10:29:49,230 Thrift API version: 19.18.0\n INFO 10:29:49,230 Loading persisted ring state\n INFO 10:29:49,235 Starting up server gossip\n INFO 10:29:49,259 Enqueuing flush of\nMemtable-LocationInfo@122130810(192/240 serialized/live bytes, 4 ops)\n INFO 10:29:49,260 Writing Memtable-LocationInfo@122130810(192/240\nserialized/live bytes, 4 ops)\n INFO 10:29:49,317 Completed flushing\n/var/lib/cassandra/data/system/LocationInfo-h-1-Data.db (300 bytes)\n INFO 10:29:49,340 Starting Messaging Service on port 7000\n INFO 10:29:49,349 JOINING: waiting for ring and schema information\n INFO 10:29:50,759 Applying migration\n4b0e20f0-1511-11e1-0000-c11bc95834d7 Add keyspace: MSA, rep\nstrategy:SimpleStrategy{}, durable_writes: true\n INFO 10:29:50,761 Enqueuing flush of\nMemtable-Migrations@1507565381(6744/8430 serialized/live bytes, 1 ops)\n INFO 10:29:50,761 Writing Memtable-Migrations@1507565381(6744/8430\nserialized/live bytes, 1 ops)\n INFO 10:29:50,761 Enqueuing flush of\nMemtable-Schema@1498835564(2889/3611 serialized/live bytes, 3 ops)\n INFO 10:29:50,776 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-1-Data.db (6808 bytes)\n INFO 10:29:50,777 Writing Memtable-Schema@1498835564(2889/3611\nserialized/live bytes, 3 ops)\n INFO 10:29:50,797 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-1-Data.db (3039 bytes)\n INFO 10:29:50,814 Applying migration\n4b6f2cb0-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@1639d811[cfId=1000,ksName=MSA,cfName=modseq,cfType=Standard,comparator=org.apache.cassandra.db.marshal.ReversedType(org.apache.cassandra.db.marshal.BytesType),subcolumncomparator=,comment=,rowCacheSize=0.0,keyCacheSize=5000000.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=0,keyCacheSavePeriodInSeconds=14400,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@2f984f7d,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.LeveledCompactionStrategy,compactionStrategyOptions={sstable_size_in_mb=10},compressionOptions={}]\n INFO 10:29:50,815 Enqueuing flush of\nMemtable-Migrations@948613108(7482/9352 serialized/live bytes, 1 ops)\n INFO 10:29:50,816 Writing Memtable-Migrations@948613108(7482/9352\nserialized/live bytes, 1 ops)\n INFO 10:29:50,816 Enqueuing flush of\nMemtable-Schema@421910828(3294/4117 serialized/live bytes, 3 ops)\n INFO 10:29:50,831 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-2-Data.db (7546 bytes)\n INFO 10:29:50,832 Writing Memtable-Schema@421910828(3294/4117\nserialized/live bytes, 3 ops)\n INFO 10:29:50,846 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-2-Data.db (3444 bytes)\n INFO 10:29:50,854 Applying migration\n4b8c9fc0-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@1bd97d0d[cfId=1001,ksName=MSA,cfName=msgid,cfType=Standard,comparator=org.apache.cassandra.db.marshal.BytesType,subcolumncomparator=,comment=,rowCacheSize=0.0,keyCacheSize=1000000.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=0,keyCacheSavePeriodInSeconds=14400,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@63a0eec3,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.SizeTieredCompactionStrategy,compactionStrategyOptions={},compressionOptions={}]\n INFO 10:29:50,855 Enqueuing flush of\nMemtable-Migrations@1520138062(7750/9687 serialized/live bytes, 1 ops)\n INFO 10:29:50,856 Writing Memtable-Migrations@1520138062(7750/9687\nserialized/live bytes, 1 ops)\n INFO 10:29:50,856 Enqueuing flush of\nMemtable-Schema@347459675(3630/4537 serialized/live bytes, 3 ops)\n INFO 10:29:50,878 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-3-Data.db (7814 bytes)\n INFO 10:29:50,879 Writing Memtable-Schema@347459675(3630/4537\nserialized/live bytes, 3 ops)\n INFO 10:29:50,894 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-3-Data.db (3780 bytes)\n INFO 10:29:50,900 Applying migration\n4ba1ae60-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@6a095b8a[cfId=1002,ksName=MSA,cfName=participants,cfType=Standard,comparator=org.apache.cassandra.db.marshal.ReversedType(org.apache.cassandra.db.marshal.BytesType),subcolumncomparator=,comment=,rowCacheSize=0.0,keyCacheSize=1000000.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=0,keyCacheSavePeriodInSeconds=14400,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@c58f769,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.SizeTieredCompactionStrategy,compactionStrategyOptions={},compressionOptions={}]\n INFO 10:29:50,900 Enqueuing flush of\nMemtable-Migrations@618337492(8194/10242 serialized/live bytes, 1 ops)\n INFO 10:29:50,901 Writing Memtable-Migrations@618337492(8194/10242\nserialized/live bytes, 1 ops)\n INFO 10:29:50,902 Enqueuing flush of\nMemtable-Schema@724860211(4020/5025 serialized/live bytes, 3 ops)\n INFO 10:29:50,917 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-4-Data.db (8258 bytes)\n INFO 10:29:50,918 Writing Memtable-Schema@724860211(4020/5025\nserialized/live bytes, 3 ops)\n INFO 10:29:50,925 Compacting\n[SSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-1-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-2-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-4-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-3-Data.db')]\n INFO 10:29:50,934 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-4-Data.db (4170 bytes)\n INFO 10:29:50,935 Compacting\n[SSTableReader(path='/var/lib/cassandra/data/system/Schema-h-2-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-1-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-4-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-3-Data.db')]\n INFO 10:29:50,940 Applying migration\n4bb4e840-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@318c69a9[cfId=1003,ksName=MSA,cfName=subinfo,cfType=Standard,comparator=org.apache.cassandra.db.marshal.ReversedType(org.apache.cassandra.db.marshal.BytesType),subcolumncomparator=,comment=,rowCacheSize=5000.0,keyCacheSize=5000000.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=14400,keyCacheSavePeriodInSeconds=14400,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@796cefa8,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.LeveledCompactionStrategy,compactionStrategyOptions={sstable_size_in_mb=10},compressionOptions={}]\n INFO 10:29:50,941 Enqueuing flush of\nMemtable-Migrations@1682081063(8618/10772 serialized/live bytes, 1\nops)\n INFO 10:29:50,941 Writing Memtable-Migrations@1682081063(8618/10772\nserialized/live bytes, 1 ops)\n INFO 10:29:50,941 Enqueuing flush of\nMemtable-Schema@1083461053(4427/5533 serialized/live bytes, 3 ops)\n INFO 10:29:50,977 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-5-Data.db (8682 bytes)\n INFO 10:29:50,978 Writing Memtable-Schema@1083461053(4427/5533\nserialized/live bytes, 3 ops)\n INFO 10:29:50,991 Compacted to\n[/var/lib/cassandra/data/system/Schema-h-5-Data.db,]. 14,433 to\n14,106 (~97% of original) bytes for 5 keys at 0.269051MB/s. Time:\n50ms.\n INFO 10:29:50,995 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-7-Data.db (4577 bytes)\n INFO 10:29:51,000 Applying migration\n4bc6e9a0-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@20b00ec2[cfId=1004,ksName=MSA,cfName=transactions,cfType=Standard,comparator=org.apache.cassandra.db.marshal.ReversedType(org.apache.cassandra.db.marshal.BytesType),subcolumncomparator=,comment=,rowCacheSize=0.0,keyCacheSize=0.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=0,keyCacheSavePeriodInSeconds=0,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@698f352,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.LeveledCompactionStrategy,compactionStrategyOptions={sstable_size_in_mb=10},compressionOptions={}]\n INFO 10:29:51,001 Enqueuing flush of\nMemtable-Migrations@596545504(9027/11283 serialized/live bytes, 1 ops)\n INFO 10:29:51,002 Writing Memtable-Migrations@596545504(9027/11283\nserialized/live bytes, 1 ops)\n INFO 10:29:51,003 Enqueuing flush of\nMemtable-Schema@1686621532(4835/6043 serialized/live bytes, 3 ops)\n INFO 10:29:51,029 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-7-Data.db (9091 bytes)\n INFO 10:29:51,029 Writing Memtable-Schema@1686621532(4835/6043\nserialized/live bytes, 3 ops)\n INFO 10:29:51,031 Compacted to\n[/var/lib/cassandra/data/system/Migrations-h-6-Data.db,]. 30,426 to\n30,234 (~99% of original) bytes for 1 keys at 0.272013MB/s. Time:\n106ms.\n INFO 10:29:51,044 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-8-Data.db (4985 bytes)\n INFO 10:29:51,049 Applying migration\n4bd76460-1511-11e1-0000-c11bc95834d7 Add column family:\norg.apache.cassandra.config.CFMetaData@4ab4faeb[cfId=1005,ksName=MSA,cfName=uid,cfType=Standard,comparator=org.apache.cassandra.db.marshal.ReversedType(org.apache.cassandra.db.marshal.BytesType),subcolumncomparator=,comment=,rowCacheSize=0.0,keyCacheSize=1500000.0,readRepairChance=1.0,replicateOnWrite=true,gcGraceSeconds=3600,defaultValidator=org.apache.cassandra.db.marshal.BytesType,keyValidator=org.apache.cassandra.db.marshal.BytesType,minCompactionThreshold=4,maxCompactionThreshold=32,rowCacheSavePeriodInSeconds=0,keyCacheSavePeriodInSeconds=14400,rowCacheKeysToSave=2147483647,rowCacheProvider=org.apache.cassandra.cache.SerializingCacheProvider@2fc5809e,mergeShardsChance=0.1,keyAlias=,column_metadata={},compactionStrategyClass=class\norg.apache.cassandra.db.compaction.LeveledCompactionStrategy,compactionStrategyOptions={sstable_size_in_mb=10},compressionOptions={}]\n INFO 10:29:51,050 Enqueuing flush of\nMemtable-Migrations@1333730706(9421/11776 serialized/live bytes, 1\nops)\n INFO 10:29:51,050 Writing Memtable-Migrations@1333730706(9421/11776\nserialized/live bytes, 1 ops)\n INFO 10:29:51,051 Enqueuing flush of\nMemtable-Schema@577668356(5236/6545 serialized/live bytes, 3 ops)\n INFO 10:29:51,065 Completed flushing\n/var/lib/cassandra/data/system/Migrations-h-9-Data.db (9485 bytes)\n INFO 10:29:51,066 Compacting\n[SSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-6-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-9-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-7-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Migrations-h-5-Data.db')]\n INFO 10:29:51,066 Writing Memtable-Schema@577668356(5236/6545\nserialized/live bytes, 3 ops)\n INFO 10:29:51,081 Completed flushing\n/var/lib/cassandra/data/system/Schema-h-9-Data.db (5386 bytes)\n INFO 10:29:51,083 Compacting\n[SSTableReader(path='/var/lib/cassandra/data/system/Schema-h-5-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-9-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-8-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/Schema-h-7-Data.db')]\n INFO 10:29:51,114 Compacted to\n[/var/lib/cassandra/data/system/Schema-h-10-Data.db,]. 29,054 to\n28,727 (~98% of original) bytes for 8 keys at 0.913207MB/s. Time:\n30ms.\n INFO 10:29:51,144 Compacted to\n[/var/lib/cassandra/data/system/Migrations-h-10-Data.db,]. 57,492 to\n57,300 (~99% of original) bytes for 1 keys at 0.700584MB/s. Time:\n78ms.\n INFO 10:29:51,410 Node /10.19.104.13 is now part of the cluster\n INFO 10:29:51,412 InetAddress /10.19.104.13 is now UP\n INFO 10:29:51,414 Enqueuing flush of\nMemtable-LocationInfo@709342045(35/43 serialized/live bytes, 1 ops)\n INFO 10:29:51,415 Writing Memtable-LocationInfo@709342045(35/43\nserialized/live bytes, 1 ops)\n INFO 10:29:51,428 Completed flushing\n/var/lib/cassandra/data/system/LocationInfo-h-2-Data.db (89 bytes)\n INFO 10:29:51,439 Node /10.19.104.12 is now part of the cluster\n INFO 10:29:51,439 InetAddress /10.19.104.12 is now UP\n INFO 10:29:51,441 Enqueuing flush of\nMemtable-LocationInfo@1292444743(35/43 serialized/live bytes, 1 ops)\n INFO 10:29:51,441 Writing Memtable-LocationInfo@1292444743(35/43\nserialized/live bytes, 1 ops)\n INFO 10:29:51,455 Completed flushing\n/var/lib/cassandra/data/system/LocationInfo-h-3-Data.db (89 bytes)\n INFO 10:29:51,456 Node /10.19.104.11 is now part of the cluster\n INFO 10:29:51,457 InetAddress /10.19.104.11 is now UP\n INFO 10:29:51,459 Enqueuing flush of\nMemtable-LocationInfo@1891328597(20/25 serialized/live bytes, 1 ops)\n INFO 10:29:51,459 Writing Memtable-LocationInfo@1891328597(20/25\nserialized/live bytes, 1 ops)\n INFO 10:29:51,471 Completed flushing\n/var/lib/cassandra/data/system/LocationInfo-h-4-Data.db (74 bytes)\n INFO 10:29:51,473 Compacting\n[SSTableReader(path='/var/lib/cassandra/data/system/LocationInfo-h-2-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/LocationInfo-h-4-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/LocationInfo-h-1-Data.db'),\nSSTableReader(path='/var/lib/cassandra/data/system/LocationInfo-h-3-Data.db')]\n INFO 10:29:51,497 Compacted to\n[/var/lib/cassandra/data/system/LocationInfo-h-5-Data.db,]. 552 to\n444 (~80% of original) bytes for 3 keys at 0.018410MB/s. Time: 23ms.\n INFO 10:30:19,349 JOINING: getting bootstrap token\n INFO 10:30:19,352 Enqueuing flush of\nMemtable-LocationInfo@225265367(36/45 serialized/live bytes, 1 ops)\n INFO 10:30:19,353 Writing Memtable-LocationInfo@225265367(36/45\nserialized/live bytes, 1 ops)\n INFO 10:30:19,364 Completed flushing\n/var/lib/cassandra/data/system/LocationInfo-h-7-Data.db (87 bytes)\n INFO 10:30:19,374 JOINING: sleeping 30000 ms for pending range setup\n INFO 10:30:49,375 JOINING: Starting to bootstrap...\nERROR 10:31:13,444 Fatal exception in thread Thread[Thread-49,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.compaction.LeveledManifest.promote(LeveledManifest.java:178)\n at org.apache.cassandra.db.compaction.LeveledCompactionStrategy.handleNotification(LeveledCompactionStrategy.java:141)\n at org.apache.cassandra.db.DataTracker.notifySSTablesChanged(DataTracker.java:466)\n at org.apache.cassandra.db.DataTracker.replace(DataTracker.java:275)\n at org.apache.cassandra.db.DataTracker.addSSTables(DataTracker.java:237)\n at org.apache.cassandra.db.DataTracker.addStreamedSSTable(DataTracker.java:242)\n at org.apache.cassandra.db.ColumnFamilyStore.addSSTable(ColumnFamilyStore.java:922)\n at org.apache.cassandra.streaming.StreamInSession.closeIfFinished(StreamInSession.java:141)\n at org.apache.cassandra.streaming.IncomingStreamReader.read(IncomingStreamReader.java:102)\n at org.apache.cassandra.net.IncomingTcpConnection.stream(IncomingTcpConnection.java:184)\n at org.apache.cassandra.net.IncomingTcpConnection.run(IncomingTcpConnection.java:81)","issue_id":"12532941","key":"CASSANDRA-3536","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-12-02T15:47:17.000+0000","role":"fixed_distractor","summary":"Assertion error during bootstraping cassandra"} {"case_id":"12533530","cluster":"DISTRACTOR-CASSANDRA-3551","comments":[{"body":"any exceptions on the other nodes?","created":"2011-12-01T23:08:51.803+0000"},{"body":"I am seeing this too; switching to ConsistencyLevel.ONE helps, but does not solve the problem completely, i.e. queries fail less often.","created":"2011-12-02T09:57:37.486+0000"},{"body":"More infos on you respective setups could help. For instance:\n* you said, 'For some column families'. Is there something specific to those column families ? Are they using compression? leveled compaction?\n* Janne: you're seeing it too, but on which version exactly did you definitively not see this problem and on which are you definitively seeing it? Is it 1.0.2 and 1.0.5 respectively as for Zhong?\n* As Jonathan said, are you seeing any error in any node logs?","created":"2011-12-02T10:23:39.390+0000"},{"body":"1.0.5, RF 3, 3 node cluster on EC2. I upgraded just recently directly from 0.6.13, so I have not been on any earlier 1.0.x version. No compression, just a straightforward upgrade with minimal tuning to the cassandra.yaml file. 2GB heap, maybe ~1GB in use. Happens with column families which have 20 rows, CFs which have 10000 rows and more. Happens when trying to read 100 rows at a time, happens when trying to read 10k rows at a time. The only factor that I've noticed while trying to tune that has any effect is changing the CL.\n\nNo errors in node logs, no anomalies in system monitoring (like suddenly increased disk latency). Only cassandra's storageproxy latency goes way up (hundreds of milliseconds), before failure.\n\nHere is the exception from hector:\n\nCaused by: me.prettyprint.hector.api.exceptions.HTimedOutException: TimedOutException()\n\tat me.prettyprint.cassandra.service.ExceptionsTranslatorImpl.translate(ExceptionsTranslatorImpl.java:42)\n\tat me.prettyprint.cassandra.service.KeyspaceServiceImpl$3.execute(KeyspaceServiceImpl.java:163)\n\tat me.prettyprint.cassandra.service.KeyspaceServiceImpl$3.execute(KeyspaceServiceImpl.java:145)\n\tat me.prettyprint.cassandra.service.Operation.executeAndSetResult(Operation.java:101)\n\tat me.prettyprint.cassandra.connection.HConnectionManager.operateWithFailover(HConnectionManager.java:233)\n\tat me.prettyprint.cassandra.service.KeyspaceServiceImpl.operateWithFailover(KeyspaceServiceImpl.java:131)\n\tat me.prettyprint.cassandra.service.KeyspaceServiceImpl.getRangeSlices(KeyspaceServiceImpl.java:167)\n\tat me.prettyprint.cassandra.model.thrift.ThriftRangeSlicesQuery$1.doInKeyspace(ThriftRangeSlicesQuery.java:67)\n\tat me.prettyprint.cassandra.model.thrift.ThriftRangeSlicesQuery$1.doInKeyspace(ThriftRangeSlicesQuery.java:63)\n\tat me.prettyprint.cassandra.model.KeyspaceOperationCallback.doInKeyspaceAndMeasure(KeyspaceOperationCallback.java:20)\n\tat me.prettyprint.cassandra.model.ExecutingKeyspace.doExecute(ExecutingKeyspace.java:85)\n\tat me.prettyprint.cassandra.model.thrift.ThriftRangeSlicesQuery.execute(ThriftRangeSlicesQuery.java:62)\n\nHere's the CF definition:\n\n ColumnFamily: XXXX\n Key Validation Class: org.apache.cassandra.db.marshal.BytesType\n Default column value validator: org.apache.cassandra.db.marshal.BytesType\n Columns sorted by: org.apache.cassandra.db.marshal.UTF8Type\n Row cache size / save period in seconds / keys to save : 0.0/0/all\n Row Cache Provider: org.apache.cassandra.cache.ConcurrentLinkedHashCacheProvider\n Key cache size / save period in seconds: 200000.0/14400\n GC grace seconds: 864000\n Compaction min/max thresholds: 4/32\n Read repair chance: 1.0\n Replicate on write: true\n Built indexes: []\n Compaction Strategy: org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy\n","created":"2011-12-02T11:28:37.198+0000"},{"body":"There is no exceptions on other nodes. \n\nI might be wrong about 'For some column families'. I saw another column family failed with Range Slice too. It works for insert. It might work for others retrieve command. I need test more when I have time. There is no compression, one row data only.\n\n\n\n ColumnFamily: dataProvider\n Key Validation Class: org.apache.cassandra.db.marshal.UTF8Type\n Default column value validator: org.apache.cassandra.db.marshal.UTF8Type\n Columns sorted by: org.apache.cassandra.db.marshal.UTF8Type\n Row cache size / save period in seconds / keys to save : 1024.0/0/all\n Row Cache Provider: org.apache.cassandra.cache.SerializingCacheProvider\n Key cache size / save period in seconds: 1024.0/14400\n GC grace seconds: 432000\n Compaction min/max thresholds: 4/32\n Read repair chance: 1.0\n Replicate on write: true\n Column Metadata:\n Column Name: active\n Validation Class: org.apache.cassandra.db.marshal.UTF8Type\n Index Name: dataProvider_active_idx\n Index Type: KEYS\n Column Name: object\n Validation Class: org.apache.cassandra.db.marshal.BytesType\n Column Name: providerData\n Validation Class: org.apache.cassandra.db.marshal.UTF8Type\n Index Name: dataProvider_providerData_idx\n Index Type: KEYS\n Column Name: providerID\n Validation Class: org.apache.cassandra.db.marshal.UTF8Type\n Index Name: dataProvider_providerID_idx\n Index Type: KEYS\n Compaction Strategy: org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy\n ","created":"2011-12-02T18:31:10.815+0000"},{"body":"I hit a version of this problem...\n\nI upgraded a production cluster from 1.0.3 (from a non-official version patched for CASSANDRA-3510) to 1.0.5. The aim was to pass CASSANDRA-3440.\n\nThis generated a timeout storm on range slices and I have reverted. \n\nNotes:\n\n1/ The 1.0.5 node CPUs all showed tiny load - in fact, they seemed to be substantially less loaded than the 1.0.3 nodes were/are again\n\n2/ The system.log files on the 1.0.5 nodes didn't record any errors\n\n3/ range_slice timeout storm experienced in application layer. Example log trace below\n\norg.apache.thrift.transport.TTransportException: java.net.SocketTimeoutException: Read timed out\n at org.apache.thrift.transport.TIOStreamTransport.read(TIOStreamTransport.java:129) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.transport.TTransport.readAll(TTransport.java:84) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.transport.TFramedTransport.readFrame(TFramedTransport.java:129) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.transport.TFramedTransport.read(TFramedTransport.java:101) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.transport.TTransport.readAll(TTransport.java:84) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.protocol.TBinaryProtocol.readAll(TBinaryProtocol.java:378) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.protocol.TBinaryProtocol.readI32(TBinaryProtocol.java:297) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.protocol.TBinaryProtocol.readMessageBegin(TBinaryProtocol.java:204) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.cassandra.thrift.Cassandra$Client.recv_get_slice(Cassandra.java:560) ~[cassandra-thrift-1.0.1.jar:1.0.1]\n at org.apache.cassandra.thrift.Cassandra$Client.get_slice(Cassandra.java:542) ~[cassandra-thrift-1.0.1.jar:1.0.1]\n at org.scale7.cassandra.pelops.Selector$3.execute(Selector.java:683) ~[scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Selector$3.execute(Selector.java:680) ~[scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Operand.tryOperation(Operand.java:86) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Operand.tryOperation(Operand.java:66) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Selector.getColumnOrSuperColumnsFromRow(Selector.java:680) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Selector.getColumnsFromRow(Selector.java:689) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Selector.getColumnsFromRow(Selector.java:676) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Selector.getColumnsFromRow(Selector.java:562) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at com.fightmymonster.game.Monsters.getMonster(Monsters.java:92) [fmmServer.jar:na]\n at com.fightmymonster.rmi.monsters.GetMonster.doWork(GetMonster.java:25) [fmmServer.jar:na]\n at org.wyki.networking.starburst.SyncRmiOperation.run(SyncRmiOperation.java:50) [fmmServer.jar:na]\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886) [na:1.6.0_22]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908) [na:1.6.0_22]\n at java.lang.Thread.run(Thread.java:662) [na:1.6.0_22]\nCaused by: java.net.SocketTimeoutException: Read timed out\n at java.net.SocketInputStream.socketRead0(Native Method) ~[na:1.6.0_22]\n at java.net.SocketInputStream.read(SocketInputStream.java:129) ~[na:1.6.0_22]\n at org.apache.thrift.transport.TIOStreamTransport.read(TIOStreamTransport.java:127) ~[libthrift-0.6.1.jar:0.6.1]\n ... 23 common frames omitted","created":"2011-12-06T18:06:04.868+0000"},{"body":"This is due to CASSANDRA-3440. More precisely, the fact that in RowRepairResolver it has changed the message from the mutation verb to the read_repair one. The problem is that ReadRepairVerbHandler does not respond anything, but the RowRepairResolver is waiting for a response.\n\nAfter looking, I haven't found any part of the code using the read_repair verb handler except for the RowRepairResolver (which would mean that before CASSANDRA-3440 it wasn't used at all, so it's worth having someone else double checking I didn't missed anything), so a simple fix is to make the ReadRepairVerbHandler return an acknowledgment. Attaching a patch for that.","created":"2011-12-07T17:07:14.058+0000"},{"body":"To be explicit: this only affects queries at CL > ONE.\n\nbq. I haven't found any part of the code using the read_repair verb handler except for the RowRepairResolver, which would mean that before CASSANDRA-3440 it wasn't used at all\n\nRight, it was used for a while, then we switched to MUTATION Verb to get the reply for \"free\" when we changed StorageProxy to wait for repair acks, but we never cleared out the Verb or RRVH.\n\n+1 on the fix.","created":"2011-12-07T17:29:41.814+0000"},{"body":"Committed, thanks","created":"2011-12-07T17:35:34.019+0000"},{"body":"Confirmed fixed in 1.0.6. Thanks!","created":"2011-12-14T19:54:21.503+0000"}],"conversations":[{"body":"I upgraded from 1.0.2 to 1.0.5. For some column families always got TimeoutException. I turned on debug and increase rpc_timeout to 1 minute, but still got timeout. I believe it is bug on 1.0.5.\n\nConsistencyLevel is QUORUM, replicate factor is 3. \n\nHere are partial logs. \n\n\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,717 StorageProxy.java (line 813) RangeSliceCommand{keyspace='keyspaceLBSDATAPRODUS', column_family='dataProvider', super_column=null, predicate=SlicePre\ndicate(slice_range:SliceRange(start:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 72 00 0C 00 0\n2 0C 00 02 0B 00 01 00 00 00 00, finish:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 72 00 0C \n00 02 0C 00 02 0B 00 01 00 00 00 00 0B 00 02 00 00 00 00, reversed:false, count:1024)), range=[PROD/US/000/0,PROD/US/999/99999], max_keys=1024}\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,718 StorageProxy.java (line 1012) restricted ranges for query [PROD/US/000/0,PROD/US/999/99999] are [[PROD/US/000/0,PROD/US/300/~], (PROD/US/300/~,PROD/\nUS/600/~], (PROD/US/600/~,PROD/US/999/99999]]\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,720 VoxeoStrategy.java (line 157) ReplicationFactor 3\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,720 VoxeoStrategy.java (line 33) PROD/US/300/~\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,721 VoxeoStrategy.java (line 96) End region for token PROD/US/300/~ PROD/US/300/~ 10.92.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,721 VoxeoStrategy.java (line 96) End region for token PROD/US/300/~ PROD/US/600/~ 10.72.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,721 VoxeoStrategy.java (line 96) End region for token PROD/US/300/~ PROD/US/999/~ 10.8.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,723 VoxeoStrategy.java (line 157) ReplicationFactor 3\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,724 ReadCallback.java (line 77) Blockfor/repair is 2/false; setting up requests to /10.92.208.103,/10.72.208.103\nDEBUG [WRITE-/10.92.208.103] 2011-12-01 22:25:39,725 OutboundTcpConnection.java (line 206) attempting to connect to /10.92.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,726 StorageProxy.java (line 859) reading RangeSliceCommand{keyspace='keyspaceLBSDATAPRODUS', column_family='dataProvider', super_column=null, predicate=\nSlicePredicate(slice_range:SliceRange(start:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 72 00\n 0C 00 02 0C 00 02 0B 00 01 00 00 00 00, finish:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 7\n2 00 0C 00 02 0C 00 02 0B 00 01 00 00 00 00 0B 00 02 00 00 00 00, reversed:false, count:1024)), range=[PROD/US/000/0,PROD/US/300/~], max_keys=1024} from /10.92.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,726 StorageProxy.java (line 859) reading RangeSliceCommand{keyspace='keyspaceLBSDATAPRODUS', column_family='dataProvider', super_column=null, predicate=\nSlicePredicate(slice_range:SliceRange(start:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 72 00\n 0C 00 02 0C 00 02 0B 00 01 00 00 00 00, finish:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 7\n2 00 0C 00 02 0C 00 02 0B 00 01 00 00 00 00 0B 00 02 00 00 00 00, reversed:false, count:1024)), range=[PROD/US/000/0,PROD/US/300/~], max_keys=1024} from /10.72.208.103\nDEBUG [WRITE-/10.8.208.103] 2011-12-01 22:25:39,727 OutboundTcpConnection.java (line 206) attempting to connect to /10.8.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,727 StorageProxy.java (line 859) reading RangeSliceCommand{keyspace='keyspaceLBSDATAPRODUS', column_family='dataProvider', super_column=null, predicate=\nSlicePredicate(slice_range:SliceRange(start:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 72 00\n 0C 00 02 0C 00 02 0B 00 01 00 00 00 00, finish:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 7\n2 00 0C 00 02 0C 00 02 0B 00 01 00 00 00 00 0B 00 02 00 00 00 00, reversed:false, count:1024)), range=[PROD/US/000/0,PROD/US/300/~], max_keys=1024} from /10.8.208.103\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,731 SliceQueryFilter.java (line 123) collecting 0 of 1024: active:false:1@1322777621601000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,731 SliceQueryFilter.java (line 123) collecting 1 of 1024: name:false:4@1322777621601000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,731 SliceQueryFilter.java (line 123) collecting 2 of 1024: providerData:false:2283@1321549067179000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,731 SliceQueryFilter.java (line 123) collecting 3 of 1024: providerID:false:1@1322777621601000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,732 SliceQueryFilter.java (line 123) collecting 4 of 1024: timestamp:false:13@1322777621601000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,732 SliceQueryFilter.java (line 123) collecting 5 of 1024: vendorData:false:2364@1322777621601000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,733 ColumnFamilyStore.java (line 1331) scanned DecoratedKey(PROD/US/001/1, 50524f442f55532f3030312f31)\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,733 RangeSliceVerbHandler.java (line 55) Sending RangeSliceReply{rows=Row(key=DecoratedKey(PROD/US/001/1, 50524f442f55532f3030312f31), cf=ColumnFamily(dataP\nrovider [active:false:1@1322777621601000,name:false:4@1322777621601000,providerData:false:2283@1321549067179000,providerID:false:1@1322777621601000,timestamp:false:13@1322777621601000,vendorData:f\nalse:2364@1322777621601000,]))} to 72@/10.72.208.103\nDEBUG [RequestResponseStage:1] 2011-12-01 22:25:39,734 ResponseVerbHandler.java (line 44) Processing response on a callback from 72@/10.72.208.103\nDEBUG [RequestResponseStage:2] 2011-12-01 22:25:39,887 ResponseVerbHandler.java (line 44) Processing response on a callback from 71@/10.92.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,889 SliceQueryFilter.java (line 123) collecting 0 of 2147483647: active:false:1@1322777621601000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,890 SliceQueryFilter.java (line 123) collecting 1 of 2147483647: name:false:4@1322777621601000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,890 SliceQueryFilter.java (line 123) collecting 2 of 2147483647: providerData:false:2283@1321549067179000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,890 SliceQueryFilter.java (line 123) collecting 3 of 2147483647: providerID:false:1@1322777621601000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,890 SliceQueryFilter.java (line 123) collecting 4 of 2147483647: timestamp:false:13@1322777621601000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,891 SliceQueryFilter.java (line 123) collecting 5 of 2147483647: vendorData:false:2364@1322777621601000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,892 StorageProxy.java (line 867) range slices read DecoratedKey(PROD/US/001/1, 50524f442f55532f3030312f31)\nDEBUG [RequestResponseStage:3] 2011-12-01 22:25:39,936 ResponseVerbHandler.java (line 44) Processing response on a callback from 73@/10.8.208.103\nDEBUG [ScheduledTasks:1] 2011-12-01 22:26:19,788 LoadBroadcaster.java (line 86) Disseminating load info ...\nDEBUG [pool-2-thread-1] 2011-12-01 22:26:39,904 StorageProxy.java (line 874) Range slice timeout: java.util.concurrent.TimeoutException: Operation timed out.\n","from":"reporter","subject":"Timeout exception for quorum reads after upgrade from 1.0.2 to 1.0.5"},{"body":"any exceptions on the other nodes?","from":"developer"},{"body":"I am seeing this too; switching to ConsistencyLevel.ONE helps, but does not solve the problem completely, i.e. queries fail less often.","from":"developer"},{"body":"More infos on you respective setups could help. For instance:\n* you said, 'For some column families'. Is there something specific to those column families ? Are they using compression? leveled compaction?\n* Janne: you're seeing it too, but on which version exactly did you definitively not see this problem and on which are you definitively seeing it? Is it 1.0.2 and 1.0.5 respectively as for Zhong?\n* As Jonathan said, are you seeing any error in any node logs?","from":"developer"},{"body":"1.0.5, RF 3, 3 node cluster on EC2. I upgraded just recently directly from 0.6.13, so I have not been on any earlier 1.0.x version. No compression, just a straightforward upgrade with minimal tuning to the cassandra.yaml file. 2GB heap, maybe ~1GB in use. Happens with column families which have 20 rows, CFs which have 10000 rows and more. Happens when trying to read 100 rows at a time, happens when trying to read 10k rows at a time. The only factor that I've noticed while trying to tune that has any effect is changing the CL.\n\nNo errors in node logs, no anomalies in system monitoring (like suddenly increased disk latency). Only cassandra's storageproxy latency goes way up (hundreds of milliseconds), before failure.\n\nHere is the exception from hector:\n\nCaused by: me.prettyprint.hector.api.exceptions.HTimedOutException: TimedOutException()\n\tat me.prettyprint.cassandra.service.ExceptionsTranslatorImpl.translate(ExceptionsTranslatorImpl.java:42)\n\tat me.prettyprint.cassandra.service.KeyspaceServiceImpl$3.execute(KeyspaceServiceImpl.java:163)\n\tat me.prettyprint.cassandra.service.KeyspaceServiceImpl$3.execute(KeyspaceServiceImpl.java:145)\n\tat me.prettyprint.cassandra.service.Operation.executeAndSetResult(Operation.java:101)\n\tat me.prettyprint.cassandra.connection.HConnectionManager.operateWithFailover(HConnectionManager.java:233)\n\tat me.prettyprint.cassandra.service.KeyspaceServiceImpl.operateWithFailover(KeyspaceServiceImpl.java:131)\n\tat me.prettyprint.cassandra.service.KeyspaceServiceImpl.getRangeSlices(KeyspaceServiceImpl.java:167)\n\tat me.prettyprint.cassandra.model.thrift.ThriftRangeSlicesQuery$1.doInKeyspace(ThriftRangeSlicesQuery.java:67)\n\tat me.prettyprint.cassandra.model.thrift.ThriftRangeSlicesQuery$1.doInKeyspace(ThriftRangeSlicesQuery.java:63)\n\tat me.prettyprint.cassandra.model.KeyspaceOperationCallback.doInKeyspaceAndMeasure(KeyspaceOperationCallback.java:20)\n\tat me.prettyprint.cassandra.model.ExecutingKeyspace.doExecute(ExecutingKeyspace.java:85)\n\tat me.prettyprint.cassandra.model.thrift.ThriftRangeSlicesQuery.execute(ThriftRangeSlicesQuery.java:62)\n\nHere's the CF definition:\n\n ColumnFamily: XXXX\n Key Validation Class: org.apache.cassandra.db.marshal.BytesType\n Default column value validator: org.apache.cassandra.db.marshal.BytesType\n Columns sorted by: org.apache.cassandra.db.marshal.UTF8Type\n Row cache size / save period in seconds / keys to save : 0.0/0/all\n Row Cache Provider: org.apache.cassandra.cache.ConcurrentLinkedHashCacheProvider\n Key cache size / save period in seconds: 200000.0/14400\n GC grace seconds: 864000\n Compaction min/max thresholds: 4/32\n Read repair chance: 1.0\n Replicate on write: true\n Built indexes: []\n Compaction Strategy: org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy\n","from":"developer"},{"body":"There is no exceptions on other nodes. \n\nI might be wrong about 'For some column families'. I saw another column family failed with Range Slice too. It works for insert. It might work for others retrieve command. I need test more when I have time. There is no compression, one row data only.\n\n\n\n ColumnFamily: dataProvider\n Key Validation Class: org.apache.cassandra.db.marshal.UTF8Type\n Default column value validator: org.apache.cassandra.db.marshal.UTF8Type\n Columns sorted by: org.apache.cassandra.db.marshal.UTF8Type\n Row cache size / save period in seconds / keys to save : 1024.0/0/all\n Row Cache Provider: org.apache.cassandra.cache.SerializingCacheProvider\n Key cache size / save period in seconds: 1024.0/14400\n GC grace seconds: 432000\n Compaction min/max thresholds: 4/32\n Read repair chance: 1.0\n Replicate on write: true\n Column Metadata:\n Column Name: active\n Validation Class: org.apache.cassandra.db.marshal.UTF8Type\n Index Name: dataProvider_active_idx\n Index Type: KEYS\n Column Name: object\n Validation Class: org.apache.cassandra.db.marshal.BytesType\n Column Name: providerData\n Validation Class: org.apache.cassandra.db.marshal.UTF8Type\n Index Name: dataProvider_providerData_idx\n Index Type: KEYS\n Column Name: providerID\n Validation Class: org.apache.cassandra.db.marshal.UTF8Type\n Index Name: dataProvider_providerID_idx\n Index Type: KEYS\n Compaction Strategy: org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy\n ","from":"developer"},{"body":"I hit a version of this problem...\n\nI upgraded a production cluster from 1.0.3 (from a non-official version patched for CASSANDRA-3510) to 1.0.5. The aim was to pass CASSANDRA-3440.\n\nThis generated a timeout storm on range slices and I have reverted. \n\nNotes:\n\n1/ The 1.0.5 node CPUs all showed tiny load - in fact, they seemed to be substantially less loaded than the 1.0.3 nodes were/are again\n\n2/ The system.log files on the 1.0.5 nodes didn't record any errors\n\n3/ range_slice timeout storm experienced in application layer. Example log trace below\n\norg.apache.thrift.transport.TTransportException: java.net.SocketTimeoutException: Read timed out\n at org.apache.thrift.transport.TIOStreamTransport.read(TIOStreamTransport.java:129) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.transport.TTransport.readAll(TTransport.java:84) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.transport.TFramedTransport.readFrame(TFramedTransport.java:129) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.transport.TFramedTransport.read(TFramedTransport.java:101) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.transport.TTransport.readAll(TTransport.java:84) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.protocol.TBinaryProtocol.readAll(TBinaryProtocol.java:378) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.protocol.TBinaryProtocol.readI32(TBinaryProtocol.java:297) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.thrift.protocol.TBinaryProtocol.readMessageBegin(TBinaryProtocol.java:204) ~[libthrift-0.6.1.jar:0.6.1]\n at org.apache.cassandra.thrift.Cassandra$Client.recv_get_slice(Cassandra.java:560) ~[cassandra-thrift-1.0.1.jar:1.0.1]\n at org.apache.cassandra.thrift.Cassandra$Client.get_slice(Cassandra.java:542) ~[cassandra-thrift-1.0.1.jar:1.0.1]\n at org.scale7.cassandra.pelops.Selector$3.execute(Selector.java:683) ~[scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Selector$3.execute(Selector.java:680) ~[scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Operand.tryOperation(Operand.java:86) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Operand.tryOperation(Operand.java:66) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Selector.getColumnOrSuperColumnsFromRow(Selector.java:680) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Selector.getColumnsFromRow(Selector.java:689) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Selector.getColumnsFromRow(Selector.java:676) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at org.scale7.cassandra.pelops.Selector.getColumnsFromRow(Selector.java:562) [scale7-pelops-1.3-1.0.x-SNAPSHOT.jar:na]\n at com.fightmymonster.game.Monsters.getMonster(Monsters.java:92) [fmmServer.jar:na]\n at com.fightmymonster.rmi.monsters.GetMonster.doWork(GetMonster.java:25) [fmmServer.jar:na]\n at org.wyki.networking.starburst.SyncRmiOperation.run(SyncRmiOperation.java:50) [fmmServer.jar:na]\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886) [na:1.6.0_22]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908) [na:1.6.0_22]\n at java.lang.Thread.run(Thread.java:662) [na:1.6.0_22]\nCaused by: java.net.SocketTimeoutException: Read timed out\n at java.net.SocketInputStream.socketRead0(Native Method) ~[na:1.6.0_22]\n at java.net.SocketInputStream.read(SocketInputStream.java:129) ~[na:1.6.0_22]\n at org.apache.thrift.transport.TIOStreamTransport.read(TIOStreamTransport.java:127) ~[libthrift-0.6.1.jar:0.6.1]\n ... 23 common frames omitted","from":"developer"},{"body":"This is due to CASSANDRA-3440. More precisely, the fact that in RowRepairResolver it has changed the message from the mutation verb to the read_repair one. The problem is that ReadRepairVerbHandler does not respond anything, but the RowRepairResolver is waiting for a response.\n\nAfter looking, I haven't found any part of the code using the read_repair verb handler except for the RowRepairResolver (which would mean that before CASSANDRA-3440 it wasn't used at all, so it's worth having someone else double checking I didn't missed anything), so a simple fix is to make the ReadRepairVerbHandler return an acknowledgment. Attaching a patch for that.","from":"developer"},{"body":"To be explicit: this only affects queries at CL > ONE.\n\nbq. I haven't found any part of the code using the read_repair verb handler except for the RowRepairResolver, which would mean that before CASSANDRA-3440 it wasn't used at all\n\nRight, it was used for a while, then we switched to MUTATION Verb to get the reply for \"free\" when we changed StorageProxy to wait for repair acks, but we never cleared out the Verb or RRVH.\n\n+1 on the fix.","from":"developer"},{"body":"Committed, thanks","from":"developer"},{"body":"Confirmed fixed in 1.0.6. Thanks!","from":"developer"}],"created":"2011-12-01T22:58:16.000+0000","description":"I upgraded from 1.0.2 to 1.0.5. For some column families always got TimeoutException. I turned on debug and increase rpc_timeout to 1 minute, but still got timeout. I believe it is bug on 1.0.5.\n\nConsistencyLevel is QUORUM, replicate factor is 3. \n\nHere are partial logs. \n\n\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,717 StorageProxy.java (line 813) RangeSliceCommand{keyspace='keyspaceLBSDATAPRODUS', column_family='dataProvider', super_column=null, predicate=SlicePre\ndicate(slice_range:SliceRange(start:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 72 00 0C 00 0\n2 0C 00 02 0B 00 01 00 00 00 00, finish:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 72 00 0C \n00 02 0C 00 02 0B 00 01 00 00 00 00 0B 00 02 00 00 00 00, reversed:false, count:1024)), range=[PROD/US/000/0,PROD/US/999/99999], max_keys=1024}\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,718 StorageProxy.java (line 1012) restricted ranges for query [PROD/US/000/0,PROD/US/999/99999] are [[PROD/US/000/0,PROD/US/300/~], (PROD/US/300/~,PROD/\nUS/600/~], (PROD/US/600/~,PROD/US/999/99999]]\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,720 VoxeoStrategy.java (line 157) ReplicationFactor 3\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,720 VoxeoStrategy.java (line 33) PROD/US/300/~\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,721 VoxeoStrategy.java (line 96) End region for token PROD/US/300/~ PROD/US/300/~ 10.92.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,721 VoxeoStrategy.java (line 96) End region for token PROD/US/300/~ PROD/US/600/~ 10.72.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,721 VoxeoStrategy.java (line 96) End region for token PROD/US/300/~ PROD/US/999/~ 10.8.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,723 VoxeoStrategy.java (line 157) ReplicationFactor 3\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,724 ReadCallback.java (line 77) Blockfor/repair is 2/false; setting up requests to /10.92.208.103,/10.72.208.103\nDEBUG [WRITE-/10.92.208.103] 2011-12-01 22:25:39,725 OutboundTcpConnection.java (line 206) attempting to connect to /10.92.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,726 StorageProxy.java (line 859) reading RangeSliceCommand{keyspace='keyspaceLBSDATAPRODUS', column_family='dataProvider', super_column=null, predicate=\nSlicePredicate(slice_range:SliceRange(start:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 72 00\n 0C 00 02 0C 00 02 0B 00 01 00 00 00 00, finish:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 7\n2 00 0C 00 02 0C 00 02 0B 00 01 00 00 00 00 0B 00 02 00 00 00 00, reversed:false, count:1024)), range=[PROD/US/000/0,PROD/US/300/~], max_keys=1024} from /10.92.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,726 StorageProxy.java (line 859) reading RangeSliceCommand{keyspace='keyspaceLBSDATAPRODUS', column_family='dataProvider', super_column=null, predicate=\nSlicePredicate(slice_range:SliceRange(start:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 72 00\n 0C 00 02 0C 00 02 0B 00 01 00 00 00 00, finish:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 7\n2 00 0C 00 02 0C 00 02 0B 00 01 00 00 00 00 0B 00 02 00 00 00 00, reversed:false, count:1024)), range=[PROD/US/000/0,PROD/US/300/~], max_keys=1024} from /10.72.208.103\nDEBUG [WRITE-/10.8.208.103] 2011-12-01 22:25:39,727 OutboundTcpConnection.java (line 206) attempting to connect to /10.8.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,727 StorageProxy.java (line 859) reading RangeSliceCommand{keyspace='keyspaceLBSDATAPRODUS', column_family='dataProvider', super_column=null, predicate=\nSlicePredicate(slice_range:SliceRange(start:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 72 00\n 0C 00 02 0C 00 02 0B 00 01 00 00 00 00, finish:80 01 00 01 00 00 00 10 67 65 74 5F 72 61 6E 67 65 5F 73 6C 69 63 65 73 00 00 00 03 0C 00 01 0B 00 03 00 00 00 0C 64 61 74 61 50 72 6F 76 69 64 65 7\n2 00 0C 00 02 0C 00 02 0B 00 01 00 00 00 00 0B 00 02 00 00 00 00, reversed:false, count:1024)), range=[PROD/US/000/0,PROD/US/300/~], max_keys=1024} from /10.8.208.103\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,731 SliceQueryFilter.java (line 123) collecting 0 of 1024: active:false:1@1322777621601000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,731 SliceQueryFilter.java (line 123) collecting 1 of 1024: name:false:4@1322777621601000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,731 SliceQueryFilter.java (line 123) collecting 2 of 1024: providerData:false:2283@1321549067179000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,731 SliceQueryFilter.java (line 123) collecting 3 of 1024: providerID:false:1@1322777621601000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,732 SliceQueryFilter.java (line 123) collecting 4 of 1024: timestamp:false:13@1322777621601000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,732 SliceQueryFilter.java (line 123) collecting 5 of 1024: vendorData:false:2364@1322777621601000\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,733 ColumnFamilyStore.java (line 1331) scanned DecoratedKey(PROD/US/001/1, 50524f442f55532f3030312f31)\nDEBUG [ReadStage:1] 2011-12-01 22:25:39,733 RangeSliceVerbHandler.java (line 55) Sending RangeSliceReply{rows=Row(key=DecoratedKey(PROD/US/001/1, 50524f442f55532f3030312f31), cf=ColumnFamily(dataP\nrovider [active:false:1@1322777621601000,name:false:4@1322777621601000,providerData:false:2283@1321549067179000,providerID:false:1@1322777621601000,timestamp:false:13@1322777621601000,vendorData:f\nalse:2364@1322777621601000,]))} to 72@/10.72.208.103\nDEBUG [RequestResponseStage:1] 2011-12-01 22:25:39,734 ResponseVerbHandler.java (line 44) Processing response on a callback from 72@/10.72.208.103\nDEBUG [RequestResponseStage:2] 2011-12-01 22:25:39,887 ResponseVerbHandler.java (line 44) Processing response on a callback from 71@/10.92.208.103\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,889 SliceQueryFilter.java (line 123) collecting 0 of 2147483647: active:false:1@1322777621601000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,890 SliceQueryFilter.java (line 123) collecting 1 of 2147483647: name:false:4@1322777621601000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,890 SliceQueryFilter.java (line 123) collecting 2 of 2147483647: providerData:false:2283@1321549067179000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,890 SliceQueryFilter.java (line 123) collecting 3 of 2147483647: providerID:false:1@1322777621601000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,890 SliceQueryFilter.java (line 123) collecting 4 of 2147483647: timestamp:false:13@1322777621601000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,891 SliceQueryFilter.java (line 123) collecting 5 of 2147483647: vendorData:false:2364@1322777621601000\nDEBUG [pool-2-thread-1] 2011-12-01 22:25:39,892 StorageProxy.java (line 867) range slices read DecoratedKey(PROD/US/001/1, 50524f442f55532f3030312f31)\nDEBUG [RequestResponseStage:3] 2011-12-01 22:25:39,936 ResponseVerbHandler.java (line 44) Processing response on a callback from 73@/10.8.208.103\nDEBUG [ScheduledTasks:1] 2011-12-01 22:26:19,788 LoadBroadcaster.java (line 86) Disseminating load info ...\nDEBUG [pool-2-thread-1] 2011-12-01 22:26:39,904 StorageProxy.java (line 874) Range slice timeout: java.util.concurrent.TimeoutException: Operation timed out.\n","issue_id":"12533530","key":"CASSANDRA-3551","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-12-07T17:35:33.000+0000","role":"fixed_distractor","summary":"Timeout exception for quorum reads after upgrade from 1.0.2 to 1.0.5"} {"case_id":"12533568","cluster":"DISTRACTOR-CASSANDRA-3554","comments":[{"body":"Unclear how we should tell when it's a good idea to re-attempt delivery in this scenario.\n\nPossibly the best solution is to just make FD smarter and mark nodes as \"effectively down\" in this situation. A \"local\" FD as in CASSANDRA-3533 could address this.","created":"2011-12-02T05:48:46.917+0000"},{"body":"Couldn't B handles this? When it drops writes from A, it could record it. Then B could have a scheduled tasks that looks for locally dropped writes and decide if it's ok to get hints based on the mutation stage queue. It could then request the hint delivery from A.","created":"2011-12-02T08:40:15.931+0000"},{"body":"That could work, but I cringe at adding both \"pull\" and \"push\" modes for hint delivery, which has historically been a source of enough bugs that more complexity is counterindicated.","created":"2011-12-02T14:11:07.516+0000"},{"body":"At risk of heresy, would bringing the hourly scan back be so bad now that our hint model doesn't suck like it did the first time we did that? ","created":"2011-12-02T16:07:30.448+0000"},{"body":"Can we keep a running tally of hints per endpoint when they are written, when they reach a threshold we deliver them? + hourly scan :)","created":"2011-12-02T16:11:50.055+0000"},{"body":"The problem I have with a \"brute force\" hourly scan or hint threshold is that you're likely to run into the same overload scenario that caused the hinting in the first place.","created":"2011-12-02T16:31:56.456+0000"},{"body":"Could we use the badness detector in dynamic switch?","created":"2011-12-03T00:27:39.373+0000"},{"body":"The other problem is even if we fix the replay issue it's still terribly slow due to excessive throttling\n\nI like the idea of changing from a push to pull mode for hint delivery. Similar to how mysql replication is client pull. Clients know how swamped they are and can throttle their own delivery.\n\n\n","created":"2011-12-05T15:22:30.182+0000"},{"body":"The problem with a pull model is that the node doing the pulling usually won't know \"I was down, therefore I should ask for hints.\" (The exception is on restart, but hints due to overload conditions or GC pauses are much more common.)","created":"2011-12-05T17:03:23.180+0000"},{"body":"Right, however its never going to know if hints are on a coordinator node due to the coordinator needed to drop some messages (backpressure?)\n\nSo either the clients can poll all nodes slowly and fetch hints or we perhaps gossip hints available flag so nodes know when hints are there to read?","created":"2011-12-05T17:07:54.287+0000"},{"body":"Right. So, not clearly simpler than doing something to our existing push model, but much more code churn. I'd rather stick with push.","created":"2011-12-06T01:33:23.406+0000"},{"body":"How expensive is the process of \"1) Wake Up. 2)Check for hints. 3) try to deliver them.\" As for parts 1 & 2 I would think we can do this more often then hourly, maybe a background thread that sleeps for N seconds and attempts again. \n\n{quote}\nThe other problem is even if we fix the replay issue it's still terribly slow due to excessive throttling{quote} Side note. This throttle can not currently be adjusted at runtime. This should be JMX able. The default may be too low. Historically the problem was the hint sending node got hammered. In my mind the throttle was protecting that system. ","created":"2011-12-07T16:19:58.222+0000"},{"body":"Brute force fix attached to check for hints-to-deliver every 10 minutes. If a hint cannot be replayed because of timing out, we abort (to that target). Reduces default hint throttle delay to 1ms.","created":"2011-12-15T04:47:42.580+0000"},{"body":"0001 does some related cleanup, including moving the mbean method StorageService.deliverHints to HHOM.scheduleHintDelivery.","created":"2011-12-15T04:52:21.126+0000"},{"body":"combined patch against 1.0","created":"2011-12-19T22:57:02.021+0000"},{"body":"updated 1.0 rebase that is pre-1034 friendly","created":"2011-12-19T23:34:32.052+0000"},{"body":"I'm not sure how exactly, but obviously the keys being passed here are not quite what we think they are:\n\n{noformat}\nDEBUG 11:57:03,907 Started scheduleAllDeliveries\nDEBUG 11:57:03,907 deliverHints to /7fff:ffff:ffff:ffff:ffff:ffff:ffff:fffe\nDEBUG 11:57:03,908 deliverHints to /5555:5555:5555:5555:5555:5555:5555:5554\nDEBUG 11:57:03,908 Checking remote(/7fff:ffff:ffff:ffff:ffff:ffff:ffff:fffe) schema before delivering hints\nDEBUG 11:57:03,908 Finished scheduleAllDeliveries\nERROR 11:57:03,909 Fatal exception in thread Thread[HintedHandoff:3,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.db.HintedHandOffManager.waitForSchemaAgreement(HintedHandOffManager.java:206)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:238)\n at org.apache.cassandra.db.HintedHandOffManager.access$200(HintedHandOffManager.java:84)\n at org.apache.cassandra.db.HintedHandOffManager$3.runMayThrow(HintedHandOffManager.java:383)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n{noformat}\n","created":"2011-12-20T12:06:03.553+0000"},{"body":"You're right, we switched from InetAddress as the key to Tokens. v2 attached.","created":"2011-12-22T20:29:11.862+0000"},{"body":"+1","created":"2011-12-22T21:18:31.203+0000"},{"body":"committed","created":"2011-12-22T21:34:34.638+0000"},{"body":"+1 \n\"shedule deliver hints\" should have been more ealier brought in .\n","created":"2012-03-21T11:59:05.954+0000"}],"conversations":[{"body":"If B drops a write from A because it is overwhelmed (but not dead), A will hint the write. But it will never get notified that B is back up (since it was never down), so it will never attempt hint delivery.","from":"reporter","subject":"Hints are not replayed unless node was marked down"},{"body":"Unclear how we should tell when it's a good idea to re-attempt delivery in this scenario.\n\nPossibly the best solution is to just make FD smarter and mark nodes as \"effectively down\" in this situation. A \"local\" FD as in CASSANDRA-3533 could address this.","from":"developer"},{"body":"Couldn't B handles this? When it drops writes from A, it could record it. Then B could have a scheduled tasks that looks for locally dropped writes and decide if it's ok to get hints based on the mutation stage queue. It could then request the hint delivery from A.","from":"developer"},{"body":"That could work, but I cringe at adding both \"pull\" and \"push\" modes for hint delivery, which has historically been a source of enough bugs that more complexity is counterindicated.","from":"developer"},{"body":"At risk of heresy, would bringing the hourly scan back be so bad now that our hint model doesn't suck like it did the first time we did that? ","from":"developer"},{"body":"Can we keep a running tally of hints per endpoint when they are written, when they reach a threshold we deliver them? + hourly scan :)","from":"developer"},{"body":"The problem I have with a \"brute force\" hourly scan or hint threshold is that you're likely to run into the same overload scenario that caused the hinting in the first place.","from":"developer"},{"body":"Could we use the badness detector in dynamic switch?","from":"developer"},{"body":"The other problem is even if we fix the replay issue it's still terribly slow due to excessive throttling\n\nI like the idea of changing from a push to pull mode for hint delivery. Similar to how mysql replication is client pull. Clients know how swamped they are and can throttle their own delivery.\n\n\n","from":"developer"},{"body":"The problem with a pull model is that the node doing the pulling usually won't know \"I was down, therefore I should ask for hints.\" (The exception is on restart, but hints due to overload conditions or GC pauses are much more common.)","from":"developer"},{"body":"Right, however its never going to know if hints are on a coordinator node due to the coordinator needed to drop some messages (backpressure?)\n\nSo either the clients can poll all nodes slowly and fetch hints or we perhaps gossip hints available flag so nodes know when hints are there to read?","from":"developer"},{"body":"Right. So, not clearly simpler than doing something to our existing push model, but much more code churn. I'd rather stick with push.","from":"developer"},{"body":"How expensive is the process of \"1) Wake Up. 2)Check for hints. 3) try to deliver them.\" As for parts 1 & 2 I would think we can do this more often then hourly, maybe a background thread that sleeps for N seconds and attempts again. \n\n{quote}\nThe other problem is even if we fix the replay issue it's still terribly slow due to excessive throttling{quote} Side note. This throttle can not currently be adjusted at runtime. This should be JMX able. The default may be too low. Historically the problem was the hint sending node got hammered. In my mind the throttle was protecting that system. ","from":"developer"},{"body":"Brute force fix attached to check for hints-to-deliver every 10 minutes. If a hint cannot be replayed because of timing out, we abort (to that target). Reduces default hint throttle delay to 1ms.","from":"developer"},{"body":"0001 does some related cleanup, including moving the mbean method StorageService.deliverHints to HHOM.scheduleHintDelivery.","from":"developer"},{"body":"combined patch against 1.0","from":"developer"},{"body":"updated 1.0 rebase that is pre-1034 friendly","from":"developer"},{"body":"I'm not sure how exactly, but obviously the keys being passed here are not quite what we think they are:\n\n{noformat}\nDEBUG 11:57:03,907 Started scheduleAllDeliveries\nDEBUG 11:57:03,907 deliverHints to /7fff:ffff:ffff:ffff:ffff:ffff:ffff:fffe\nDEBUG 11:57:03,908 deliverHints to /5555:5555:5555:5555:5555:5555:5555:5554\nDEBUG 11:57:03,908 Checking remote(/7fff:ffff:ffff:ffff:ffff:ffff:ffff:fffe) schema before delivering hints\nDEBUG 11:57:03,908 Finished scheduleAllDeliveries\nERROR 11:57:03,909 Fatal exception in thread Thread[HintedHandoff:3,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.db.HintedHandOffManager.waitForSchemaAgreement(HintedHandOffManager.java:206)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:238)\n at org.apache.cassandra.db.HintedHandOffManager.access$200(HintedHandOffManager.java:84)\n at org.apache.cassandra.db.HintedHandOffManager$3.runMayThrow(HintedHandOffManager.java:383)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n{noformat}\n","from":"developer"},{"body":"You're right, we switched from InetAddress as the key to Tokens. v2 attached.","from":"developer"},{"body":"+1","from":"developer"},{"body":"committed","from":"developer"},{"body":"+1 \n\"shedule deliver hints\" should have been more ealier brought in .\n","from":"developer"}],"created":"2011-12-02T05:46:15.000+0000","description":"If B drops a write from A because it is overwhelmed (but not dead), A will hint the write. But it will never get notified that B is back up (since it was never down), so it will never attempt hint delivery.","issue_id":"12533568","key":"CASSANDRA-3554","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2011-12-22T21:34:34.000+0000","role":"fixed_distractor","summary":"Hints are not replayed unless node was marked down"} {"case_id":"12538279","cluster":"DISTRACTOR-CASSANDRA-3736","comments":[{"body":"Hi Jackson, Just a clarification is 50.56.58.55 up and cassandra is running?\n\n INFO [GossipStage:1] 2012-01-12 23:59:35,177 StorageService.java (line 1016) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n\nThis happens when the replaced node is running or resurrected. ","created":"2012-01-13T01:37:45.636+0000"},{"body":"Hi Vijay,\n\nI filed the ticket with DataStax that prompted this issue. I'm not 100% certain whether the node we replaced was fully and consistently offline from the point we performed the replacement. I *believe* it was, especially because the -Dreplace_token refuses to work if the node being replaced is online --- and we took no further action to bring the replaced node back (its VM wasn't initializing any network interfaces other than \"lo\").\n\nEven if the replaced node comes back, it shouldn't be allowed to re-join the ring with a token already owned by an \"Up\" node. It should be subjected to the same condition -Dreplace_token is, where the token being used by the new ring member must be owned by a \"Down\" node.\n\nDavid","created":"2012-01-13T01:50:56.970+0000"},{"body":"Alternatively, it would be good for Cassandra to provide a convenient (nodetool) way to drop the \"Down\" IP when a token is simultaneously occupied by one \"Up\" IP and at least one \"Down\" IP.","created":"2012-01-13T01:59:49.273+0000"},{"body":"bq. Alternatively, it would be good for Cassandra to provide a convenient (nodetool) way to drop the \"Down\" IP when a token is simultaneously occupied by one \"Up\" IP and at least one \"Down\" IP.\n\nCASSANDRA-3337 is designed to handle these kinds of situations (where gossip is not doing the right thing naturally)","created":"2012-01-13T02:03:16.469+0000"},{"body":"bq. is 50.56.58.55 up and cassandra is running?\n\nno. the Cassandra on 50.56.58.55 was not UP/had shutdown. But the IP is available, though i don't think that matters.\n\nso my test case was simply:\n1) start 2 nodes (A , B). With A being the seed, B bootstrap into it (by specifying a token)\n2) stop B (after B had successfully joined)\n3) start C with -Dcassandra.replace_token=\n\ncontinuing restarting C (without the replace_token param) could observe the behavior.","created":"2012-01-13T02:29:47.587+0000"},{"body":"Simple patch to remove from SYSTEM_TABLE/RING_KEY when token is replaced.","created":"2012-01-13T03:43:13.017+0000"},{"body":"fix no good.\n\nand to ensure fix is deployed, checked the compiled class:\n\n$ javap -c -private -classpath ./build/classes/main/ org.apache.cassandra.db.SystemTable | grep \"updateToken(java.net.Inet\" -A10\npublic static synchronized void updateToken(java.net.InetAddress, org.apache.cassandra.dht.Token);\n Code:\n 0: aload_0\n 1: invokestatic #51; //Method org/apache/cassandra/utils/FBUtilities.getLocalAddress:()Ljava/net/InetAddress;\n 4: if_acmpne 12\n 7: aload_1\n 8: invokestatic #52; //Method removeToken:(Lorg/apache/cassandra/dht/Token;)V\n 11: return\n\nto ensure removeToken is added (per the patch)\n\nand the classpath of the jvm is using it:\n\n INFO 20:32:57,083 Classpath: ./bin/../conf:./bin/*../build/classes/main*:./bin/../build/classes/thrift:./bin/../lib/antlr-3.2.jar:./bin/../lib/avro-1.4.0-fixes.jar:./bin/../lib/avro-1.4.0-sources-fixes.jar:./bin/../lib/commons-cli-1.1.jar:./bin/../lib/commons-codec-1.2.jar:./bin/../lib/commons-lang-2.4.jar:./bin/../lib/compress-lzf-0.8.4.jar:./bin/../lib/concurrentlinkedhashmap-lru-1.2.jar:./bin/../lib/guava-r08.jar:./bin/../lib/high-scale-lib-1.1.2.jar:./bin/../lib/jackson-core-asl-1.4.0.jar:./bin/../lib/jackson-mapper-asl-1.4.0.jar:./bin/../lib/jamm-0.2.5.jar:./bin/../lib/jline-0.9.94.jar:./bin/../lib/json-simple-1.1.jar:./bin/../lib/libthrift-0.6.jar:./bin/../lib/log4j-1.2.16.jar:./bin/../lib/servlet-api-2.5-20081211.jar:./bin/../lib/slf4j-api-1.6.1.jar:./bin/../lib/slf4j-log4j12-1.6.1.jar:./bin/../lib/snakeyaml-1.6.jar:./bin/../lib/snappy-java-1.0.4.1.jar:./bin/../lib/jamm-0.2.5.jar\n\nlog from the replacement node:\n{noformat}\n INFO 20:34:27,856 Listening for thrift clients...\n INFO 20:35:28,750 Node /50.56.58.55 is now part of the cluster\n INFO 20:35:28,750 InetAddress /50.56.58.55 is now UP\n INFO 20:35:28,751 Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO 20:35:38,841 InetAddress /50.56.58.55 is now dead.\n INFO 20:35:58,852 FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n INFO 20:36:59,786 Node /50.56.58.55 is now part of the cluster\n INFO 20:36:59,787 InetAddress /50.56.58.55 is now UP\n INFO 20:36:59,787 Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO 20:37:09,887 InetAddress /50.56.58.55 is now dead.\n INFO 20:37:29,898 FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n{noformat}\n\n","created":"2012-01-13T21:39:32.223+0000"},{"body":"I suspect we have the same issue as I outlined in CASSANDRA-3737","created":"2012-01-13T22:00:43.495+0000"},{"body":"looks like fix from CASSANDRA-3747 got the fix.\n\nthe replacement node would still get this once:\n INFO [GossipStage:1] 2012-01-18 23:45:56,412 Gossiper.java (line 834) Node /50.56.58.55 is now part of the cluster\n INFO [GossipStage:1] 2012-01-18 23:45:56,412 Gossiper.java (line 800) InetAddress /50.56.58.55 is now UP\n INFO [GossipStage:1] 2012-01-18 23:45:56,413 StorageService.java (line 1016) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO [GossipTasks:1] 2012-01-18 23:46:05,805 Gossiper.java (line 814) InetAddress /50.56.58.55 is now dead.\n INFO [GossipTasks:1] 2012-01-18 23:46:26,819 Gossiper.java (line 628) FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n\nbut its quiet after that.\n\nthe other node would receive the same info also:\n\n INFO [GossipTasks:1] 2012-01-18 23:45:57,486 Gossiper.java (line 628) FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n\nand the gossipinfo of those nodes are the matching:\n\n\n$ ./bin/nodetool -h 50.56.31.186 gossipinfo\n/50.56.59.68\n RELEASE_VERSION:1.0.7-SNAPSHOT\n LOAD:6820.0\n RPC_ADDRESS:50.56.59.68\n STATUS:NORMAL,0\n SCHEMA:00000000-0000-1000-0000-000000000000\naction-quick2/50.56.31.186\n RELEASE_VERSION:1.0.7-SNAPSHOT\n RPC_ADDRESS:50.56.31.186\n STATUS:NORMAL,85070591730234615865843651857942052864\n LOAD:11372.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n\n$ ./bin/nodetool -h 50.56.59.68 gossipinfo\naction-quick/50.56.59.68\n SCHEMA:00000000-0000-1000-0000-000000000000\n RELEASE_VERSION:1.0.7-SNAPSHOT\n LOAD:6820.0\n RPC_ADDRESS:50.56.59.68\n STATUS:NORMAL,0\n/50.56.31.186\n SCHEMA:00000000-0000-1000-0000-000000000000\n RELEASE_VERSION:1.0.7-SNAPSHOT\n LOAD:11372.0\n RPC_ADDRESS:50.56.31.186\n STATUS:NORMAL,85070591730234615865843651857942052864\n","created":"2012-01-18T23:53:29.593+0000"},{"body":"Yes and the fix attached with this ticket will also remove the node from the System table, while replacing hence you wont even see the following message...\n\n>>> INFO [GossipStage:1] 2012-01-18 23:45:56,412 Gossiper.java (line 800) InetAddress /50.56.58.55 is now UP\n\nThe problem is that we remove the node after 30 seconds.... Meanwhile the gossip will make the other node know about .55 and hence the message in the other node. \nThe patch will fix this by removing the information from the System table in the first place instead of restart which triggering it to reappear. Can you try redoing the test? it doesn't appear back in my tests.","created":"2012-01-19T04:33:36.175+0000"},{"body":"If CASSANDRA-3747 solved this, then I don't think there's any full solution here worth applying, since this is mostly just a cosmetic problem and not worth introducing a possibly destabilizing change over. Anyone running into this can use CASSANDRA-3337 to remove it, or avoid replacing tokens.\n\n+1 to this patch for 1.0 and trunk, though.","created":"2012-01-19T19:11:04.252+0000"},{"body":"Committed both in 1.0 and trunk. Thanks!","created":"2012-01-19T21:08:38.498+0000"},{"body":"I have the same issue with Cassandra 1.0.11 (used DataStax AMI).\nI thought it was supposed to be solved already.\nI see those messages on the node that was started using -Dcassandra.replace_token=.\n\nFrom time to time I also see \n{color:blue} \n INFO [GossipTasks:1] 2012-11-13 12:26:38,195 Gossiper.java (line 818) InetAddress / is now dead.\n INFO [GossipTasks:1] 2012-11-13 12:26:58,203 Gossiper.java (line 632) FatClient / has been silent for 30000ms, removing from gossip\n INFO [GossipStage:1] 2012-11-13 12:27:59,210 Gossiper.java (line 838) Node / is now part of the cluster\n INFO [GossipStage:1] 2012-11-13 12:27:59,210 Gossiper.java (line 804) InetAddress / is now UP\n INFO [GossipStage:1] 2012-11-13 12:27:59,210 StorageService.java (line 1017) Nodes / and / have the same token 113427455640312821154458202477256070484. Ignoring /\n\n{color}\n\n","created":"2012-11-13T12:16:49.193+0000"}],"conversations":[{"body":"https://issues.apache.org/jira/browse/CASSANDRA-957 introduce a -Dreplace_token,\n\nhowever, the replaced IP keeps on showing up in the Gossiper when starting the replacement node:\n\n{noformat}\n INFO [Thread-2] 2012-01-12 23:59:35,162 CassandraDaemon.java (line 213) Listening for thrift clients...\n INFO [GossipStage:1] 2012-01-12 23:59:35,173 Gossiper.java (line 836) Node /50.56.59.68 has restarted, now UP\n INFO [GossipStage:1] 2012-01-12 23:59:35,174 Gossiper.java (line 804) InetAddress /50.56.59.68 is now UP\n INFO [GossipStage:1] 2012-01-12 23:59:35,175 StorageService.java (line 988) Node /50.56.59.68 state jump to normal\n INFO [GossipStage:1] 2012-01-12 23:59:35,176 Gossiper.java (line 836) Node /50.56.58.55 has restarted, now UP\n INFO [GossipStage:1] 2012-01-12 23:59:35,176 Gossiper.java (line 804) InetAddress /50.56.58.55 is now UP\n INFO [GossipStage:1] 2012-01-12 23:59:35,177 StorageService.java (line 1016) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO [GossipTasks:1] 2012-01-12 23:59:45,048 Gossiper.java (line 818) InetAddress /50.56.58.55 is now dead.\n INFO [GossipTasks:1] 2012-01-13 00:00:06,062 Gossiper.java (line 632) FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n INFO [GossipStage:1] 2012-01-13 00:01:06,320 Gossiper.java (line 838) Node /50.56.58.55 is now part of the cluster\n INFO [GossipStage:1] 2012-01-13 00:01:06,320 Gossiper.java (line 804) InetAddress /50.56.58.55 is now UP\n INFO [GossipStage:1] 2012-01-13 00:01:06,321 StorageService.java (line 1016) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO [GossipTasks:1] 2012-01-13 00:01:16,106 Gossiper.java (line 818) InetAddress /50.56.58.55 is now dead.\n INFO [GossipTasks:1] 2012-01-13 00:01:37,121 Gossiper.java (line 632) FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n INFO [GossipStage:1] 2012-01-13 00:02:37,352 Gossiper.java (line 838) Node /50.56.58.55 is now part of the cluster\n INFO [GossipStage:1] 2012-01-13 00:02:37,353 Gossiper.java (line 804) InetAddress /50.56.58.55 is now UP\n INFO [GossipStage:1] 2012-01-13 00:02:37,353 StorageService.java (line 1016) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO [GossipTasks:1] 2012-01-13 00:02:47,158 Gossiper.java (line 818) InetAddress /50.56.58.55 is now dead.\n INFO [GossipStage:1] 2012-01-13 00:02:50,162 Gossiper.java (line 818) InetAddress /50.56.58.55 is now dead.\n INFO [GossipStage:1] 2012-01-13 00:02:50,163 StorageService.java (line 1156) Removing token 122029383590318827259508597176866581733 for /50.56.58.55\n{noformat}\n\nin the above, /50.56.58.55 was the replaced IP.\n\ntried adding the \"Gossiper.instance.removeEndpoint(endpoint);\" in the StorageService.java where the message 'Nodes %s and %s have the same token %s. Ignoring %s\",' seems only have fixed this temporary. Here is a ring output:\n\n{noformat}\nriptano@action-quick:~/work/cassandra$ ./bin/nodetool -h localhost ring\nAddress DC Rack Status State Load Owns Token \n 85070591730234615865843651857942052864 \n50.56.59.68 datacenter1 rack1 Up Normal 6.67 KB 85.56% 60502102442797279294142560823234402248 \n50.56.31.186 datacenter1 rack1 Up Normal 11.12 KB 14.44% 85070591730234615865843651857942052864 \n{noformat}\n\ngossipinfo:\n{noformat}\n$ ./bin/nodetool -h localhost gossipinfo\n/50.56.58.55\n LOAD:6835.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n RPC_ADDRESS:50.56.58.55\n STATUS:NORMAL,85070591730234615865843651857942052864\n RELEASE_VERSION:1.0.7-SNAPSHOT\n/50.56.59.68\n LOAD:6835.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n RPC_ADDRESS:50.56.59.68\n STATUS:NORMAL,60502102442797279294142560823234402248\n RELEASE_VERSION:1.0.7-SNAPSHOT\naction-quick2/50.56.31.186\n LOAD:11387.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n RPC_ADDRESS:50.56.31.186\n STATUS:NORMAL,85070591730234615865843651857942052864\n RELEASE_VERSION:1.0.7-SNAPSHOT\n{noformat}\n\nNote that at 1 point earlier it seems to have been removed:\n\n$ ./bin/nodetool -h localhost gossipinfo\n/50.56.59.68\n LOAD:13815.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n RPC_ADDRESS:50.56.59.68\n STATUS:NORMAL,60502102442797279294142560823234402248\n RELEASE_VERSION:1.0.7-SNAPSHOT\naction-quick2/50.56.31.186\n LOAD:13725.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n RPC_ADDRESS:50.56.31.186\n STATUS:NORMAL,85070591730234615865843651857942052864\n RELEASE_VERSION:1.0.7-SNAPSHOT\n\nriptano@action-quick2:~/work/cassandra$ INFO [GossipStage:1] 2012-01-13 01:03:30,073 Gossiper.java (line 838) Node /50.56.58.55 is now part of the cluster\n\n INFO [GossipStage:1] 2012-01-13 01:03:30,073 Gossiper.java (line 804) InetAddress /50.56.58.55 is now UP\n\n INFO [GossipStage:1] 2012-01-13 01:03:30,074 StorageService.java (line 1017) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55","from":"reporter","subject":"-Dreplace_token leaves old node (IP) in the gossip with the token."},{"body":"Hi Jackson, Just a clarification is 50.56.58.55 up and cassandra is running?\n\n INFO [GossipStage:1] 2012-01-12 23:59:35,177 StorageService.java (line 1016) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n\nThis happens when the replaced node is running or resurrected. ","from":"developer"},{"body":"Hi Vijay,\n\nI filed the ticket with DataStax that prompted this issue. I'm not 100% certain whether the node we replaced was fully and consistently offline from the point we performed the replacement. I *believe* it was, especially because the -Dreplace_token refuses to work if the node being replaced is online --- and we took no further action to bring the replaced node back (its VM wasn't initializing any network interfaces other than \"lo\").\n\nEven if the replaced node comes back, it shouldn't be allowed to re-join the ring with a token already owned by an \"Up\" node. It should be subjected to the same condition -Dreplace_token is, where the token being used by the new ring member must be owned by a \"Down\" node.\n\nDavid","from":"developer"},{"body":"Alternatively, it would be good for Cassandra to provide a convenient (nodetool) way to drop the \"Down\" IP when a token is simultaneously occupied by one \"Up\" IP and at least one \"Down\" IP.","from":"developer"},{"body":"bq. Alternatively, it would be good for Cassandra to provide a convenient (nodetool) way to drop the \"Down\" IP when a token is simultaneously occupied by one \"Up\" IP and at least one \"Down\" IP.\n\nCASSANDRA-3337 is designed to handle these kinds of situations (where gossip is not doing the right thing naturally)","from":"developer"},{"body":"bq. is 50.56.58.55 up and cassandra is running?\n\nno. the Cassandra on 50.56.58.55 was not UP/had shutdown. But the IP is available, though i don't think that matters.\n\nso my test case was simply:\n1) start 2 nodes (A , B). With A being the seed, B bootstrap into it (by specifying a token)\n2) stop B (after B had successfully joined)\n3) start C with -Dcassandra.replace_token=\n\ncontinuing restarting C (without the replace_token param) could observe the behavior.","from":"developer"},{"body":"Simple patch to remove from SYSTEM_TABLE/RING_KEY when token is replaced.","from":"developer"},{"body":"fix no good.\n\nand to ensure fix is deployed, checked the compiled class:\n\n$ javap -c -private -classpath ./build/classes/main/ org.apache.cassandra.db.SystemTable | grep \"updateToken(java.net.Inet\" -A10\npublic static synchronized void updateToken(java.net.InetAddress, org.apache.cassandra.dht.Token);\n Code:\n 0: aload_0\n 1: invokestatic #51; //Method org/apache/cassandra/utils/FBUtilities.getLocalAddress:()Ljava/net/InetAddress;\n 4: if_acmpne 12\n 7: aload_1\n 8: invokestatic #52; //Method removeToken:(Lorg/apache/cassandra/dht/Token;)V\n 11: return\n\nto ensure removeToken is added (per the patch)\n\nand the classpath of the jvm is using it:\n\n INFO 20:32:57,083 Classpath: ./bin/../conf:./bin/*../build/classes/main*:./bin/../build/classes/thrift:./bin/../lib/antlr-3.2.jar:./bin/../lib/avro-1.4.0-fixes.jar:./bin/../lib/avro-1.4.0-sources-fixes.jar:./bin/../lib/commons-cli-1.1.jar:./bin/../lib/commons-codec-1.2.jar:./bin/../lib/commons-lang-2.4.jar:./bin/../lib/compress-lzf-0.8.4.jar:./bin/../lib/concurrentlinkedhashmap-lru-1.2.jar:./bin/../lib/guava-r08.jar:./bin/../lib/high-scale-lib-1.1.2.jar:./bin/../lib/jackson-core-asl-1.4.0.jar:./bin/../lib/jackson-mapper-asl-1.4.0.jar:./bin/../lib/jamm-0.2.5.jar:./bin/../lib/jline-0.9.94.jar:./bin/../lib/json-simple-1.1.jar:./bin/../lib/libthrift-0.6.jar:./bin/../lib/log4j-1.2.16.jar:./bin/../lib/servlet-api-2.5-20081211.jar:./bin/../lib/slf4j-api-1.6.1.jar:./bin/../lib/slf4j-log4j12-1.6.1.jar:./bin/../lib/snakeyaml-1.6.jar:./bin/../lib/snappy-java-1.0.4.1.jar:./bin/../lib/jamm-0.2.5.jar\n\nlog from the replacement node:\n{noformat}\n INFO 20:34:27,856 Listening for thrift clients...\n INFO 20:35:28,750 Node /50.56.58.55 is now part of the cluster\n INFO 20:35:28,750 InetAddress /50.56.58.55 is now UP\n INFO 20:35:28,751 Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO 20:35:38,841 InetAddress /50.56.58.55 is now dead.\n INFO 20:35:58,852 FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n INFO 20:36:59,786 Node /50.56.58.55 is now part of the cluster\n INFO 20:36:59,787 InetAddress /50.56.58.55 is now UP\n INFO 20:36:59,787 Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO 20:37:09,887 InetAddress /50.56.58.55 is now dead.\n INFO 20:37:29,898 FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n{noformat}\n\n","from":"developer"},{"body":"I suspect we have the same issue as I outlined in CASSANDRA-3737","from":"developer"},{"body":"looks like fix from CASSANDRA-3747 got the fix.\n\nthe replacement node would still get this once:\n INFO [GossipStage:1] 2012-01-18 23:45:56,412 Gossiper.java (line 834) Node /50.56.58.55 is now part of the cluster\n INFO [GossipStage:1] 2012-01-18 23:45:56,412 Gossiper.java (line 800) InetAddress /50.56.58.55 is now UP\n INFO [GossipStage:1] 2012-01-18 23:45:56,413 StorageService.java (line 1016) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO [GossipTasks:1] 2012-01-18 23:46:05,805 Gossiper.java (line 814) InetAddress /50.56.58.55 is now dead.\n INFO [GossipTasks:1] 2012-01-18 23:46:26,819 Gossiper.java (line 628) FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n\nbut its quiet after that.\n\nthe other node would receive the same info also:\n\n INFO [GossipTasks:1] 2012-01-18 23:45:57,486 Gossiper.java (line 628) FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n\nand the gossipinfo of those nodes are the matching:\n\n\n$ ./bin/nodetool -h 50.56.31.186 gossipinfo\n/50.56.59.68\n RELEASE_VERSION:1.0.7-SNAPSHOT\n LOAD:6820.0\n RPC_ADDRESS:50.56.59.68\n STATUS:NORMAL,0\n SCHEMA:00000000-0000-1000-0000-000000000000\naction-quick2/50.56.31.186\n RELEASE_VERSION:1.0.7-SNAPSHOT\n RPC_ADDRESS:50.56.31.186\n STATUS:NORMAL,85070591730234615865843651857942052864\n LOAD:11372.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n\n$ ./bin/nodetool -h 50.56.59.68 gossipinfo\naction-quick/50.56.59.68\n SCHEMA:00000000-0000-1000-0000-000000000000\n RELEASE_VERSION:1.0.7-SNAPSHOT\n LOAD:6820.0\n RPC_ADDRESS:50.56.59.68\n STATUS:NORMAL,0\n/50.56.31.186\n SCHEMA:00000000-0000-1000-0000-000000000000\n RELEASE_VERSION:1.0.7-SNAPSHOT\n LOAD:11372.0\n RPC_ADDRESS:50.56.31.186\n STATUS:NORMAL,85070591730234615865843651857942052864\n","from":"developer"},{"body":"Yes and the fix attached with this ticket will also remove the node from the System table, while replacing hence you wont even see the following message...\n\n>>> INFO [GossipStage:1] 2012-01-18 23:45:56,412 Gossiper.java (line 800) InetAddress /50.56.58.55 is now UP\n\nThe problem is that we remove the node after 30 seconds.... Meanwhile the gossip will make the other node know about .55 and hence the message in the other node. \nThe patch will fix this by removing the information from the System table in the first place instead of restart which triggering it to reappear. Can you try redoing the test? it doesn't appear back in my tests.","from":"developer"},{"body":"If CASSANDRA-3747 solved this, then I don't think there's any full solution here worth applying, since this is mostly just a cosmetic problem and not worth introducing a possibly destabilizing change over. Anyone running into this can use CASSANDRA-3337 to remove it, or avoid replacing tokens.\n\n+1 to this patch for 1.0 and trunk, though.","from":"developer"},{"body":"Committed both in 1.0 and trunk. Thanks!","from":"developer"},{"body":"I have the same issue with Cassandra 1.0.11 (used DataStax AMI).\nI thought it was supposed to be solved already.\nI see those messages on the node that was started using -Dcassandra.replace_token=.\n\nFrom time to time I also see \n{color:blue} \n INFO [GossipTasks:1] 2012-11-13 12:26:38,195 Gossiper.java (line 818) InetAddress / is now dead.\n INFO [GossipTasks:1] 2012-11-13 12:26:58,203 Gossiper.java (line 632) FatClient / has been silent for 30000ms, removing from gossip\n INFO [GossipStage:1] 2012-11-13 12:27:59,210 Gossiper.java (line 838) Node / is now part of the cluster\n INFO [GossipStage:1] 2012-11-13 12:27:59,210 Gossiper.java (line 804) InetAddress / is now UP\n INFO [GossipStage:1] 2012-11-13 12:27:59,210 StorageService.java (line 1017) Nodes / and / have the same token 113427455640312821154458202477256070484. Ignoring /\n\n{color}\n\n","from":"developer"}],"created":"2012-01-13T01:13:21.000+0000","description":"https://issues.apache.org/jira/browse/CASSANDRA-957 introduce a -Dreplace_token,\n\nhowever, the replaced IP keeps on showing up in the Gossiper when starting the replacement node:\n\n{noformat}\n INFO [Thread-2] 2012-01-12 23:59:35,162 CassandraDaemon.java (line 213) Listening for thrift clients...\n INFO [GossipStage:1] 2012-01-12 23:59:35,173 Gossiper.java (line 836) Node /50.56.59.68 has restarted, now UP\n INFO [GossipStage:1] 2012-01-12 23:59:35,174 Gossiper.java (line 804) InetAddress /50.56.59.68 is now UP\n INFO [GossipStage:1] 2012-01-12 23:59:35,175 StorageService.java (line 988) Node /50.56.59.68 state jump to normal\n INFO [GossipStage:1] 2012-01-12 23:59:35,176 Gossiper.java (line 836) Node /50.56.58.55 has restarted, now UP\n INFO [GossipStage:1] 2012-01-12 23:59:35,176 Gossiper.java (line 804) InetAddress /50.56.58.55 is now UP\n INFO [GossipStage:1] 2012-01-12 23:59:35,177 StorageService.java (line 1016) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO [GossipTasks:1] 2012-01-12 23:59:45,048 Gossiper.java (line 818) InetAddress /50.56.58.55 is now dead.\n INFO [GossipTasks:1] 2012-01-13 00:00:06,062 Gossiper.java (line 632) FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n INFO [GossipStage:1] 2012-01-13 00:01:06,320 Gossiper.java (line 838) Node /50.56.58.55 is now part of the cluster\n INFO [GossipStage:1] 2012-01-13 00:01:06,320 Gossiper.java (line 804) InetAddress /50.56.58.55 is now UP\n INFO [GossipStage:1] 2012-01-13 00:01:06,321 StorageService.java (line 1016) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO [GossipTasks:1] 2012-01-13 00:01:16,106 Gossiper.java (line 818) InetAddress /50.56.58.55 is now dead.\n INFO [GossipTasks:1] 2012-01-13 00:01:37,121 Gossiper.java (line 632) FatClient /50.56.58.55 has been silent for 30000ms, removing from gossip\n INFO [GossipStage:1] 2012-01-13 00:02:37,352 Gossiper.java (line 838) Node /50.56.58.55 is now part of the cluster\n INFO [GossipStage:1] 2012-01-13 00:02:37,353 Gossiper.java (line 804) InetAddress /50.56.58.55 is now UP\n INFO [GossipStage:1] 2012-01-13 00:02:37,353 StorageService.java (line 1016) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55\n INFO [GossipTasks:1] 2012-01-13 00:02:47,158 Gossiper.java (line 818) InetAddress /50.56.58.55 is now dead.\n INFO [GossipStage:1] 2012-01-13 00:02:50,162 Gossiper.java (line 818) InetAddress /50.56.58.55 is now dead.\n INFO [GossipStage:1] 2012-01-13 00:02:50,163 StorageService.java (line 1156) Removing token 122029383590318827259508597176866581733 for /50.56.58.55\n{noformat}\n\nin the above, /50.56.58.55 was the replaced IP.\n\ntried adding the \"Gossiper.instance.removeEndpoint(endpoint);\" in the StorageService.java where the message 'Nodes %s and %s have the same token %s. Ignoring %s\",' seems only have fixed this temporary. Here is a ring output:\n\n{noformat}\nriptano@action-quick:~/work/cassandra$ ./bin/nodetool -h localhost ring\nAddress DC Rack Status State Load Owns Token \n 85070591730234615865843651857942052864 \n50.56.59.68 datacenter1 rack1 Up Normal 6.67 KB 85.56% 60502102442797279294142560823234402248 \n50.56.31.186 datacenter1 rack1 Up Normal 11.12 KB 14.44% 85070591730234615865843651857942052864 \n{noformat}\n\ngossipinfo:\n{noformat}\n$ ./bin/nodetool -h localhost gossipinfo\n/50.56.58.55\n LOAD:6835.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n RPC_ADDRESS:50.56.58.55\n STATUS:NORMAL,85070591730234615865843651857942052864\n RELEASE_VERSION:1.0.7-SNAPSHOT\n/50.56.59.68\n LOAD:6835.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n RPC_ADDRESS:50.56.59.68\n STATUS:NORMAL,60502102442797279294142560823234402248\n RELEASE_VERSION:1.0.7-SNAPSHOT\naction-quick2/50.56.31.186\n LOAD:11387.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n RPC_ADDRESS:50.56.31.186\n STATUS:NORMAL,85070591730234615865843651857942052864\n RELEASE_VERSION:1.0.7-SNAPSHOT\n{noformat}\n\nNote that at 1 point earlier it seems to have been removed:\n\n$ ./bin/nodetool -h localhost gossipinfo\n/50.56.59.68\n LOAD:13815.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n RPC_ADDRESS:50.56.59.68\n STATUS:NORMAL,60502102442797279294142560823234402248\n RELEASE_VERSION:1.0.7-SNAPSHOT\naction-quick2/50.56.31.186\n LOAD:13725.0\n SCHEMA:00000000-0000-1000-0000-000000000000\n RPC_ADDRESS:50.56.31.186\n STATUS:NORMAL,85070591730234615865843651857942052864\n RELEASE_VERSION:1.0.7-SNAPSHOT\n\nriptano@action-quick2:~/work/cassandra$ INFO [GossipStage:1] 2012-01-13 01:03:30,073 Gossiper.java (line 838) Node /50.56.58.55 is now part of the cluster\n\n INFO [GossipStage:1] 2012-01-13 01:03:30,073 Gossiper.java (line 804) InetAddress /50.56.58.55 is now UP\n\n INFO [GossipStage:1] 2012-01-13 01:03:30,074 StorageService.java (line 1017) Nodes /50.56.58.55 and action-quick2/50.56.31.186 have the same token 85070591730234615865843651857942052864. Ignoring /50.56.58.55","issue_id":"12538279","key":"CASSANDRA-3736","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2012-01-19T21:08:38.000+0000","role":"fixed_distractor","summary":"-Dreplace_token leaves old node (IP) in the gossip with the token."} {"case_id":"12544061","cluster":"DISTRACTOR-CASSANDRA-3957","comments":[{"body":"I note that this is the Periodic commitlog executor, which does NOT block for CommitLog.add. So, if the CF is serialized (later) while another thread modifies it, we could hit this error.\n\nWe don't modify the mutation CF directly in CFS.apply, since we clone into the arena allocator.\n\nThe only place I see where we modify the mutation CF is in ignoreObsoleteMutations... but that only happens for indexed columns, so that can't be the cause here since it involves SuperColumns.\n\n(Note that the mutation CF won't be changed by updateRowCache, since the mutation CF is never inserted into the cache; it's only used to add to an existing cache row, if one exists.)\n\nAm I missing something?\n\nPerhaps it would be best to start by creating a test to serialize the above RowMutation and see if it reproduces in the absence of concurrent activity, to rule out the possibility of serializedSize simply having a bug w/ SuperColumns.","created":"2012-02-24T22:04:59.125+0000"},{"body":"below is another similar stacktrace where only standard column (no super, no composite) was used.\n\n{panel}\nERROR 08:50:12,479 Fatal exception in thread Thread[COMMIT-LOG-WRITER,5,main]\njava.lang.AssertionError: Final buffer length 2090383 to accomodate data size of 736236 (predicted 2090383) for RowMutation(keyspace='cfs', key='6337343566396330363438373131653130303030653562646562313236356262', modifications=[ColumnFamily(sblocks [6337343839316430363438373131653130303030653562646562313236356262:false:736121@1330707012466,])])\nat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:682)\nat org.apache.cassandra.db.RowMutation.getSerializedBuffer(RowMutation.java:279)\nat org.apache.cassandra.db.commitlog.CommitLogSegment.write(CommitLogSegment.java:122)\nat org.apache.cassandra.db.commitlog.CommitLog$LogRecordAdder.run(CommitLog.java:599)\nat org.apache.cassandra.db.commitlog.PeriodicCommitLogExecutorService$1.runMayThrow(PeriodicCommitLogExecutorService.java:49)\nat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\nat java.lang.Thread.run(Thread.java:680)\n{panel}","created":"2012-03-02T20:42:59.887+0000"},{"body":"Jackson's stacktrace involving sblocks was caused by a bug in DSE's CFS code re-using a ByteBuffer it had handed off to Thrift.\n\nCould be that the original report had a similar problem.","created":"2012-03-06T18:57:40.267+0000"},{"body":"this comes from a difference use case. the client code is in hector, maybe we need to look into how hector is handling the bytebuffer too?\n\nERROR [COMMIT-LOG-WRITER] 2012-02-20 06:16:56,621 org.apache.cassandra.service.AbstractCassandraDaemon Fatal exception in thread Thread[COMMIT-LOG-WRITER,5,main]\njava.lang.AssertionError: Final buffer length 275 to accomodate data size of 271 (predicted 275) for RowMutation(keyspace='foo', key='984fb8106d3ff5043365736457c45085', modifications=[ColumnFa\nmily(Dora_la [SuperColumn(1329718356206 [5f656e64:false:1@1329718616576004,5f6c6173744576656e7454696d65:false:8@1329718616576003,5f6c61737450617468:false:8@1329718616576005,5f76697369744576\n656e74496e646578:false:4@1329718616576002,5f76697369744964:false:16@1329718616576000,5f7669736974496e646578:false:4@1329718616576001,]),])])\nat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:682)\nat org.apache.cassandra.db.RowMutation.getSerializedBuffer(RowMutation.java:279)\nat org.apache.cassandra.db.commitlog.CommitLogSegment.write(CommitLogSegment.java:122)\nat org.apache.cassandra.db.commitlog.CommitLog$LogRecordAdder.run(CommitLog.java:599)\nat org.apache.cassandra.db.commitlog.PeriodicCommitLogExecutorService$1.runMayThrow(PeriodicCommitLogExecutorService.java:49)\nat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\nat java.lang.Thread.run(Thread.java:662)\n\nERROR [COMMIT-LOG-WRITER] 2012-02-07 22:33:09,691 org.apache.cassandra.service.AbstractCassandraDaemon Fatal exception in thread Thread[COMMIT-LOG-WRITER,5,main]\njava.lang.AssertionError: Final buffer length 550 to accomodate data size of 275 (predicted 271) for RowMutation(keyspace='dyn', key='7836ab62d10ddb48d00668639ebdf6ce', modifications=[ColumnFa\nmily(Dora_la [SuperColumn(1328653749409 [5f656e64:false:1@1328653989655004,5f6c6173744576656e7454696d65:false:8@1328653989655003,5f6c61737450617468:false:12@1328653989655005,5f7669736974457\n6656e74496e646578:false:4@1328653989655002,5f76697369744964:false:16@1328653989655000,5f7669736974496e646578:false:4@1328653989655001,]),])])\nat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:682)\nat org.apache.cassandra.db.RowMutation.getSerializedBuffer(RowMutation.java:279)\nat org.apache.cassandra.db.commitlog.CommitLogSegment.write(CommitLogSegment.java:122)\nat org.apache.cassandra.db.commitlog.CommitLog$LogRecordAdder.run(CommitLog.java:599)\nat org.apache.cassandra.db.commitlog.PeriodicCommitLogExecutorService$1.runMayThrow(PeriodicCommitLogExecutorService.java:49)\nat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\nat java.lang.Thread.run(Thread.java:662)","created":"2012-03-06T22:39:03.180+0000"},{"body":"Here's my theory.\n\nThe stacktrace concerning the sblocks is due to a bug in DSE's CFS. This happened because CFS sometimes communicate with C* in the same JVM without going trough the wire, so the re-use of ByteBuffer was problematic.\n\nThe other stacktrace however can't be the same problem. Whatever Hector does, it goes over the wire and thus can't change the size of a ByteBuffer in C*.\n\nHowever, if we eliminate the sblocks trace, all the other trace uses SuperColumn. And as it turns out, I think we have a race with SC. As Jonathan pointed out, with the period commit log, if we happen to modify a mutation that was sent to the commit log at any time, that would be a bug. This means not modifying the CF object in the RowMutation, but for SuperColumn, this also mean we shouldn't modify those super columns. However, when we update the row cache (and if it's not a SerializingCache), we do potentially store references to the original SCs in the cached row, which could then lead to the stack trace on this issue. I'm attaching a patch that fixes this.\n\nTo make sure this could be the issue here, it would help to know if the stacktrace comes from column families where row cache was used and was not set to the serializing cache.\n\nI'll note however that the very small difference in the stacktraces between the predicted size and the actual serialized side support the theory of having say just one column of a super column updated, while in the sblocks case, the difference was much bigger, supporting the theory of the ByteBuffer content being changed to something completely different.\n","created":"2012-03-07T16:19:35.967+0000"},{"body":"+1","created":"2012-03-07T17:26:09.098+0000"},{"body":"Committed, thanks","created":"2012-03-07T17:39:30.256+0000"},{"body":"Thomas (the reporter of the original stack trace in the description) just confirmed on the mailing list that he was using row cache / CLHC.","created":"2012-03-07T18:22:52.142+0000"}],"conversations":[{"body":"As reported at http://mail-archives.apache.org/mod_mbox/cassandra-user/201202.mbox/%3CCADJL=w5kH5TEQXOwhTn5Jm3cmR4Rj=NfjcqLryXV7pLyASi95A@mail.gmail.com%3E,\n\n{noformat}\nERROR 10:51:44,282 Fatal exception in thread\nThread[COMMIT-LOG-WRITER,5,main]\njava.lang.AssertionError: Final buffer length 4690 to accomodate data size\nof 2347 (predicted 2344) for RowMutation(keyspace='Player',\nkey='36336138643338652d366162302d343334392d383466302d356166643863353133356465',\nmodifications=[ColumnFamily(PlayerCity [SuperColumn(owneditem_1019\n[]),SuperColumn(owneditem_1024 []),SuperColumn(owneditem_1026\n[]),SuperColumn(owneditem_1074 []),SuperColumn(owneditem_1077\n[]),SuperColumn(owneditem_1084 []),SuperColumn(owneditem_1094\n[]),SuperColumn(owneditem_1130 []),SuperColumn(owneditem_1136\n[]),SuperColumn(owneditem_1141 []),SuperColumn(owneditem_1142\n[]),SuperColumn(owneditem_1145 []),SuperColumn(owneditem_1218\n[636f6e6e6563746564:false:5@1329648704269002\n,63757272656e744865616c7468:false:3@1329648704269006\n,656e64436f6e737472756374696f6e54696d65:false:13@1329648704269007\n,6964:false:4@1329648704269000,6974656d4964:false:15@1329648704269001\n,6c61737444657374726f79656454696d65:false:1@1329648704269008\n,6c61737454696d65436f6c6c6563746564:false:13@1329648704269005\n,736b696e4964:false:7@1329648704269009,78:false:4@1329648704269003\n,79:false:3@1329648704269004,]),SuperColumn(owneditem_133\n[]),SuperColumn(owneditem_134 []),SuperColumn(owneditem_135\n[]),SuperColumn(owneditem_141 []),SuperColumn(owneditem_147\n[]),SuperColumn(owneditem_154 []),SuperColumn(owneditem_159\n[]),SuperColumn(owneditem_171 []),SuperColumn(owneditem_253\n[]),SuperColumn(owneditem_422 []),SuperColumn(owneditem_438\n[]),SuperColumn(owneditem_515 []),SuperColumn(owneditem_521\n[]),SuperColumn(owneditem_523 []),SuperColumn(owneditem_525\n[]),SuperColumn(owneditem_562 []),SuperColumn(owneditem_61\n[]),SuperColumn(owneditem_634 []),SuperColumn(owneditem_636\n[]),SuperColumn(owneditem_71 []),SuperColumn(owneditem_712\n[]),SuperColumn(owneditem_720 []),SuperColumn(owneditem_728\n[]),SuperColumn(owneditem_787 []),SuperColumn(owneditem_797\n[]),SuperColumn(owneditem_798 []),SuperColumn(owneditem_838\n[]),SuperColumn(owneditem_842 []),SuperColumn(owneditem_847\n[]),SuperColumn(owneditem_849 []),SuperColumn(owneditem_851\n[]),SuperColumn(owneditem_852 []),SuperColumn(owneditem_853\n[]),SuperColumn(owneditem_854 []),SuperColumn(owneditem_857\n[]),SuperColumn(owneditem_858 []),SuperColumn(owneditem_874\n[]),SuperColumn(owneditem_884 []),SuperColumn(owneditem_886\n[]),SuperColumn(owneditem_908 []),SuperColumn(owneditem_91\n[]),SuperColumn(owneditem_911 []),SuperColumn(owneditem_930\n[]),SuperColumn(owneditem_934 []),SuperColumn(owneditem_937\n[]),SuperColumn(owneditem_944 []),SuperColumn(owneditem_945\n[]),SuperColumn(owneditem_962 []),SuperColumn(owneditem_963\n[]),SuperColumn(owneditem_964 []),])])\n at org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:682)\n at org.apache.cassandra.db.RowMutation.getSerializedBuffer(RowMutation.java:279)\n at org.apache.cassandra.db.commitlog.CommitLogSegment.write(CommitLogSegment.java:122)\n at org.apache.cassandra.db.commitlog.CommitLog$LogRecordAdder.run(CommitLog.java:599)\n at org.apache.cassandra.db.commitlog.PeriodicCommitLogExecutorService$1.runMayThrow(PeriodicCommitLogExecutorService.java:49)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.lang.Thread.run(Thread.java:662)\n{noformat}\n","from":"reporter","subject":"Supercolumn serialization assertion failure"},{"body":"I note that this is the Periodic commitlog executor, which does NOT block for CommitLog.add. So, if the CF is serialized (later) while another thread modifies it, we could hit this error.\n\nWe don't modify the mutation CF directly in CFS.apply, since we clone into the arena allocator.\n\nThe only place I see where we modify the mutation CF is in ignoreObsoleteMutations... but that only happens for indexed columns, so that can't be the cause here since it involves SuperColumns.\n\n(Note that the mutation CF won't be changed by updateRowCache, since the mutation CF is never inserted into the cache; it's only used to add to an existing cache row, if one exists.)\n\nAm I missing something?\n\nPerhaps it would be best to start by creating a test to serialize the above RowMutation and see if it reproduces in the absence of concurrent activity, to rule out the possibility of serializedSize simply having a bug w/ SuperColumns.","from":"developer"},{"body":"below is another similar stacktrace where only standard column (no super, no composite) was used.\n\n{panel}\nERROR 08:50:12,479 Fatal exception in thread Thread[COMMIT-LOG-WRITER,5,main]\njava.lang.AssertionError: Final buffer length 2090383 to accomodate data size of 736236 (predicted 2090383) for RowMutation(keyspace='cfs', key='6337343566396330363438373131653130303030653562646562313236356262', modifications=[ColumnFamily(sblocks [6337343839316430363438373131653130303030653562646562313236356262:false:736121@1330707012466,])])\nat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:682)\nat org.apache.cassandra.db.RowMutation.getSerializedBuffer(RowMutation.java:279)\nat org.apache.cassandra.db.commitlog.CommitLogSegment.write(CommitLogSegment.java:122)\nat org.apache.cassandra.db.commitlog.CommitLog$LogRecordAdder.run(CommitLog.java:599)\nat org.apache.cassandra.db.commitlog.PeriodicCommitLogExecutorService$1.runMayThrow(PeriodicCommitLogExecutorService.java:49)\nat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\nat java.lang.Thread.run(Thread.java:680)\n{panel}","from":"developer"},{"body":"Jackson's stacktrace involving sblocks was caused by a bug in DSE's CFS code re-using a ByteBuffer it had handed off to Thrift.\n\nCould be that the original report had a similar problem.","from":"developer"},{"body":"this comes from a difference use case. the client code is in hector, maybe we need to look into how hector is handling the bytebuffer too?\n\nERROR [COMMIT-LOG-WRITER] 2012-02-20 06:16:56,621 org.apache.cassandra.service.AbstractCassandraDaemon Fatal exception in thread Thread[COMMIT-LOG-WRITER,5,main]\njava.lang.AssertionError: Final buffer length 275 to accomodate data size of 271 (predicted 275) for RowMutation(keyspace='foo', key='984fb8106d3ff5043365736457c45085', modifications=[ColumnFa\nmily(Dora_la [SuperColumn(1329718356206 [5f656e64:false:1@1329718616576004,5f6c6173744576656e7454696d65:false:8@1329718616576003,5f6c61737450617468:false:8@1329718616576005,5f76697369744576\n656e74496e646578:false:4@1329718616576002,5f76697369744964:false:16@1329718616576000,5f7669736974496e646578:false:4@1329718616576001,]),])])\nat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:682)\nat org.apache.cassandra.db.RowMutation.getSerializedBuffer(RowMutation.java:279)\nat org.apache.cassandra.db.commitlog.CommitLogSegment.write(CommitLogSegment.java:122)\nat org.apache.cassandra.db.commitlog.CommitLog$LogRecordAdder.run(CommitLog.java:599)\nat org.apache.cassandra.db.commitlog.PeriodicCommitLogExecutorService$1.runMayThrow(PeriodicCommitLogExecutorService.java:49)\nat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\nat java.lang.Thread.run(Thread.java:662)\n\nERROR [COMMIT-LOG-WRITER] 2012-02-07 22:33:09,691 org.apache.cassandra.service.AbstractCassandraDaemon Fatal exception in thread Thread[COMMIT-LOG-WRITER,5,main]\njava.lang.AssertionError: Final buffer length 550 to accomodate data size of 275 (predicted 271) for RowMutation(keyspace='dyn', key='7836ab62d10ddb48d00668639ebdf6ce', modifications=[ColumnFa\nmily(Dora_la [SuperColumn(1328653749409 [5f656e64:false:1@1328653989655004,5f6c6173744576656e7454696d65:false:8@1328653989655003,5f6c61737450617468:false:12@1328653989655005,5f7669736974457\n6656e74496e646578:false:4@1328653989655002,5f76697369744964:false:16@1328653989655000,5f7669736974496e646578:false:4@1328653989655001,]),])])\nat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:682)\nat org.apache.cassandra.db.RowMutation.getSerializedBuffer(RowMutation.java:279)\nat org.apache.cassandra.db.commitlog.CommitLogSegment.write(CommitLogSegment.java:122)\nat org.apache.cassandra.db.commitlog.CommitLog$LogRecordAdder.run(CommitLog.java:599)\nat org.apache.cassandra.db.commitlog.PeriodicCommitLogExecutorService$1.runMayThrow(PeriodicCommitLogExecutorService.java:49)\nat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\nat java.lang.Thread.run(Thread.java:662)","from":"developer"},{"body":"Here's my theory.\n\nThe stacktrace concerning the sblocks is due to a bug in DSE's CFS. This happened because CFS sometimes communicate with C* in the same JVM without going trough the wire, so the re-use of ByteBuffer was problematic.\n\nThe other stacktrace however can't be the same problem. Whatever Hector does, it goes over the wire and thus can't change the size of a ByteBuffer in C*.\n\nHowever, if we eliminate the sblocks trace, all the other trace uses SuperColumn. And as it turns out, I think we have a race with SC. As Jonathan pointed out, with the period commit log, if we happen to modify a mutation that was sent to the commit log at any time, that would be a bug. This means not modifying the CF object in the RowMutation, but for SuperColumn, this also mean we shouldn't modify those super columns. However, when we update the row cache (and if it's not a SerializingCache), we do potentially store references to the original SCs in the cached row, which could then lead to the stack trace on this issue. I'm attaching a patch that fixes this.\n\nTo make sure this could be the issue here, it would help to know if the stacktrace comes from column families where row cache was used and was not set to the serializing cache.\n\nI'll note however that the very small difference in the stacktraces between the predicted size and the actual serialized side support the theory of having say just one column of a super column updated, while in the sblocks case, the difference was much bigger, supporting the theory of the ByteBuffer content being changed to something completely different.\n","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed, thanks","from":"developer"},{"body":"Thomas (the reporter of the original stack trace in the description) just confirmed on the mailing list that he was using row cache / CLHC.","from":"developer"}],"created":"2012-02-24T21:51:28.000+0000","description":"As reported at http://mail-archives.apache.org/mod_mbox/cassandra-user/201202.mbox/%3CCADJL=w5kH5TEQXOwhTn5Jm3cmR4Rj=NfjcqLryXV7pLyASi95A@mail.gmail.com%3E,\n\n{noformat}\nERROR 10:51:44,282 Fatal exception in thread\nThread[COMMIT-LOG-WRITER,5,main]\njava.lang.AssertionError: Final buffer length 4690 to accomodate data size\nof 2347 (predicted 2344) for RowMutation(keyspace='Player',\nkey='36336138643338652d366162302d343334392d383466302d356166643863353133356465',\nmodifications=[ColumnFamily(PlayerCity [SuperColumn(owneditem_1019\n[]),SuperColumn(owneditem_1024 []),SuperColumn(owneditem_1026\n[]),SuperColumn(owneditem_1074 []),SuperColumn(owneditem_1077\n[]),SuperColumn(owneditem_1084 []),SuperColumn(owneditem_1094\n[]),SuperColumn(owneditem_1130 []),SuperColumn(owneditem_1136\n[]),SuperColumn(owneditem_1141 []),SuperColumn(owneditem_1142\n[]),SuperColumn(owneditem_1145 []),SuperColumn(owneditem_1218\n[636f6e6e6563746564:false:5@1329648704269002\n,63757272656e744865616c7468:false:3@1329648704269006\n,656e64436f6e737472756374696f6e54696d65:false:13@1329648704269007\n,6964:false:4@1329648704269000,6974656d4964:false:15@1329648704269001\n,6c61737444657374726f79656454696d65:false:1@1329648704269008\n,6c61737454696d65436f6c6c6563746564:false:13@1329648704269005\n,736b696e4964:false:7@1329648704269009,78:false:4@1329648704269003\n,79:false:3@1329648704269004,]),SuperColumn(owneditem_133\n[]),SuperColumn(owneditem_134 []),SuperColumn(owneditem_135\n[]),SuperColumn(owneditem_141 []),SuperColumn(owneditem_147\n[]),SuperColumn(owneditem_154 []),SuperColumn(owneditem_159\n[]),SuperColumn(owneditem_171 []),SuperColumn(owneditem_253\n[]),SuperColumn(owneditem_422 []),SuperColumn(owneditem_438\n[]),SuperColumn(owneditem_515 []),SuperColumn(owneditem_521\n[]),SuperColumn(owneditem_523 []),SuperColumn(owneditem_525\n[]),SuperColumn(owneditem_562 []),SuperColumn(owneditem_61\n[]),SuperColumn(owneditem_634 []),SuperColumn(owneditem_636\n[]),SuperColumn(owneditem_71 []),SuperColumn(owneditem_712\n[]),SuperColumn(owneditem_720 []),SuperColumn(owneditem_728\n[]),SuperColumn(owneditem_787 []),SuperColumn(owneditem_797\n[]),SuperColumn(owneditem_798 []),SuperColumn(owneditem_838\n[]),SuperColumn(owneditem_842 []),SuperColumn(owneditem_847\n[]),SuperColumn(owneditem_849 []),SuperColumn(owneditem_851\n[]),SuperColumn(owneditem_852 []),SuperColumn(owneditem_853\n[]),SuperColumn(owneditem_854 []),SuperColumn(owneditem_857\n[]),SuperColumn(owneditem_858 []),SuperColumn(owneditem_874\n[]),SuperColumn(owneditem_884 []),SuperColumn(owneditem_886\n[]),SuperColumn(owneditem_908 []),SuperColumn(owneditem_91\n[]),SuperColumn(owneditem_911 []),SuperColumn(owneditem_930\n[]),SuperColumn(owneditem_934 []),SuperColumn(owneditem_937\n[]),SuperColumn(owneditem_944 []),SuperColumn(owneditem_945\n[]),SuperColumn(owneditem_962 []),SuperColumn(owneditem_963\n[]),SuperColumn(owneditem_964 []),])])\n at org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:682)\n at org.apache.cassandra.db.RowMutation.getSerializedBuffer(RowMutation.java:279)\n at org.apache.cassandra.db.commitlog.CommitLogSegment.write(CommitLogSegment.java:122)\n at org.apache.cassandra.db.commitlog.CommitLog$LogRecordAdder.run(CommitLog.java:599)\n at org.apache.cassandra.db.commitlog.PeriodicCommitLogExecutorService$1.runMayThrow(PeriodicCommitLogExecutorService.java:49)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.lang.Thread.run(Thread.java:662)\n{noformat}\n","issue_id":"12544061","key":"CASSANDRA-3957","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2012-03-07T17:39:30.000+0000","role":"fixed_distractor","summary":"Supercolumn serialization assertion failure"} {"case_id":"12553475","cluster":"DISTRACTOR-CASSANDRA-4206","comments":[{"body":"Do you have {{multithreaded_compaction}} enabled?","created":"2012-05-02T20:53:57.510+0000"},{"body":"bq. It's the first time I've tried secondary indexes of other type than UTF8Type\n\nI think that's a red herring; the error occurs when compacting the internal hints columnfamily, which has no indexes on it.","created":"2012-05-02T20:55:16.064+0000"},{"body":"bq. After a few hours of inserting data in the afternoon and midnight repair+compact, the next day I couldn't find any row using the IntegerType secondary index\n\nThat part is definitely a [separate] bug, assuming the index completed building.","created":"2012-05-02T20:56:55.826+0000"},{"body":"I'm not sure how useful it is at this point, but just for the record I'm seeing what appears to be the same issue with 1.0.7 (which means it shouldn't be CASSANDRA-3579):\n\n{noformat}\n INFO [GossipTasks:1] 2012-08-27 14:22:50,242 Gossiper.java (line 818) InetAddress /xx.xx.178.59 is now dead.\n INFO [GossipStage:1] 2012-08-27 14:22:59,090 Gossiper.java (line 804) InetAddress /xx.xx.178.59 is now UP\n INFO [HintedHandoff:1] 2012-08-27 14:23:41,548 HintedHandOffManager.java (line 296) Started hinted handoff for token: 132332031580364958013534569556798748899 with IP: /xx.xx.178.59\n INFO [HintedHandoff:1] 2012-08-27 14:23:41,870 ColumnFamilyStore.java (line 704) Enqueuing flush of Memtable-HintsColumnFamily@2081164539(597050/47764000 serialized/live bytes, 857 ops)\n INFO [FlushWriter:181] 2012-08-27 14:23:41,870 Memtable.java (line 246) Writing Memtable-HintsColumnFamily@2081164539(597050/47764000 serialized/live bytes, 857 ops)\n INFO [FlushWriter:181] 2012-08-27 14:23:41,959 Memtable.java (line 283) Completed flushing /xx/xx/xx/cassandra/datafile/system/HintsColumnFamily-hc-6730-Data.db (624946 bytes)\n INFO [CompactionExecutor:884] 2012-08-27 14:23:41,961 CompactionTask.java (line 113) Compacting [SSTableReader(path='/xx/xx/xx/cassandra/datafile/system/HintsColumnFamily-hc-6729-Data.db'), SSTableReader(path='/ngs/app/xcardp/cassandra/datafile/system/HintsColumnFamily-hc-6730-Data.db')]\n INFO [CompactionExecutor:884] 2012-08-27 14:23:41,987 CompactionController.java (line 133) Compacting large row system/HintsColumnFamily:31372e33342e3137382e3534 (274816343 bytes) incrementally\nERROR [CompactionExecutor:884] 2012-08-27 14:23:56,322 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[CompactionExecutor:884,1,main]\njava.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:124)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:159)\n at org.apache.cassandra.db.compaction.CompactionManager$6.call(CompactionManager.java:275)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nERROR [HintedHandoff:1] 2012-08-27 14:23:56,323 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:369)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:248)\n at org.apache.cassandra.db.HintedHandOffManager.access$200(HintedHandOffManager.java:84)\n at org.apache.cassandra.db.HintedHandOffManager$3.runMayThrow(HintedHandOffManager.java:416)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:365)\n ... 7 more\nCaused by: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:124)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:159)\n at org.apache.cassandra.db.compaction.CompactionManager$6.call(CompactionManager.java:275)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n ... 3 more\nERROR [HintedHandoff:1] 2012-08-27 14:23:56,345 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:369)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:248)\n at org.apache.cassandra.db.HintedHandOffManager.access$200(HintedHandOffManager.java:84)\n at org.apache.cassandra.db.HintedHandOffManager$3.runMayThrow(HintedHandOffManager.java:416)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:365)\n ... 7 more\nCaused by: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:124)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:159)\n at org.apache.cassandra.db.compaction.CompactionManager$6.call(CompactionManager.java:275)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n ... 3 more\n INFO [COMMIT-LOG-WRITER] 2012-08-27 14:48:20,812 CommitLogSegment.java (line 60) Creating new commitlog segment /xxx/xxx/xxx/commitlog/CommitLog-1346078900812.log\n{noformat}\n\nI should note that this node was upgraded from 0.7.4 about three days prior.","created":"2012-08-31T02:02:34.532+0000"},{"body":"Just had this pop up on a 1.1.6 node after doing a 'nodetool removetoken' for a node that had collected a lot of hints.","created":"2012-10-30T23:38:28.579+0000"},{"body":"I am seeing what seems like the same issue on 1.2.2 on nodes that had collected a lot of hints after a node had been down for a couple hours.\n\n{noformat}\nERROR [CompactionExecutor:104] 2013-03-01 00:29:54,211 CassandraDaemon.java (line 132) Exception in thread Thread[CompactionExecutor:104,1,main]\njava.lang.AssertionError: originally calculated column size of 350273328 but now it is 350297055\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:163)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:59)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:62)\n at org.apache.cassandra.db.compaction.CompactionManager$7.runMayThrow(CompactionManager.java:422)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source)\n at java.util.concurrent.FutureTask$Sync.innerRun(Unknown Source)\n at java.util.concurrent.FutureTask.run(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\nERROR [HintedHandoff:1] 2013-03-01 00:29:54,211 CassandraDaemon.java (line 132) Exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 350273328 but now it is 350297055\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:406)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:252)\n at org.apache.cassandra.db.HintedHandOffManager.access$300(HintedHandOffManager.java:89)\n at org.apache.cassandra.db.HintedHandOffManager$4.runMayThrow(HintedHandOffManager.java:459)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\nCaused by: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 350273328 but now it is 350297055\n at java.util.concurrent.FutureTask$Sync.innerGet(Unknown Source)\n at java.util.concurrent.FutureTask.get(Unknown Source)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:402)\n ... 7 more\nCaused by: java.lang.AssertionError: originally calculated column size of 350273328 but now it is 350297055\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:163)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:59)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:62)\n at org.apache.cassandra.db.compaction.CompactionManager$7.runMayThrow(CompactionManager.java:422)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source)\n at java.util.concurrent.FutureTask$Sync.innerRun(Unknown Source)\n at java.util.concurrent.FutureTask.run(Unknown Source)\n ... 3 more\n{noformat}","created":"2013-03-01T14:32:03.693+0000"},{"body":"Version: 1.2.2 - 6 nodes, RF3\nOS: Centos 6.3\n\nLooking closer at the errors, I'm only seeing this for the system/hints:\n\n{noformat}\n INFO [CompactionExecutor:555] 2013-03-01 15:37:10,737 CompactionController.java (line 158) Compacting large row system/hints:384ac791\n-d0d7-4f5e-b66b-57af885341d5 (350297131 bytes) incrementally\nERROR [CompactionExecutor:555] 2013-03-01 15:37:28,377 CassandraDaemon.java (line 132) Exception in thread Thread[CompactionExecutor:5\n55,1,main]\n{noformat}","created":"2013-03-01T15:39:57.384+0000"},{"body":"I have 1600 compactions pending and no compaction is running, so they are all stuck. I also got this in the logs. And Jonathan, I have multi_threaded_compaction disabled. \n\nERROR [CompactionExecutor:111] 2013-05-02 13:54:05,913 CassandraDaemon.java (line 174) Exception in thread Thread[CompactionExecutor:111,1,main]java.lang.AssertionError: originally calculated column size of 1337269150 but now it is 1337269195 at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135) at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162) at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48) at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:188)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:439)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:895)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:918)\n at java.lang.Thread.run(Thread.java:662)\n\n","created":"2013-05-02T20:44:55.647+0000"},{"body":"Similar thing over here.\n\nC*1.2.2\n\nI have a lot of hints (3 & 4 GB) on 2 nodes out of 12 due to an unknown issue while growing to multiple datacenter (all the nodes of the new datacenter saw themselves as UNREACHABLE from the cassandra-cli - nodetool ring is ok...).\n\nOn these 2 nodes I see now :\n\n{code}\nERROR 08:37:16,878 Exception in thread Thread[CompactionExecutor:1540,1,main]\njava.lang.AssertionError: originally calculated column size of 5469343266 but now it is 5469343506\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:163)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:59)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:62)\n at org.apache.cassandra.db.compaction.CompactionManager$7.runMayThrow(CompactionManager.java:422)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:439)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nERROR 08:37:16,878 Exception in thread Thread[HintedHandoff:176,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 5469343266 but now it is 5469343506\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:406)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:252)\n at org.apache.cassandra.db.HintedHandOffManager.access$300(HintedHandOffManager.java:89)\n at org.apache.cassandra.db.HintedHandOffManager$4.runMayThrow(HintedHandOffManager.java:459)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 5469343266 but now it is 5469343506\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:402)\n ... 7 more\nCaused by: java.lang.AssertionError: originally calculated column size of 5469343266 but now it is 5469343506\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:163)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n{code}","created":"2013-06-05T08:43:38.579+0000"},{"body":"Seeing this on my 1.2.5 cluster since upgrading:\n\nubuntu@c2-1d:~$ nodetool -h localhost compact TRProd Timelines \nError occurred during compaction\njava.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 91281671 but now it is 91281729\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:252)\n at java.util.concurrent.FutureTask.get(FutureTask.java:111)\n at org.apache.cassandra.db.compaction.CompactionManager.performMaximal(CompactionManager.java:334)\n at org.apache.cassandra.db.ColumnFamilyStore.forceMajorCompaction(ColumnFamilyStore.java:1657)\n at org.apache.cassandra.service.StorageService.forceTableCompaction(StorageService.java:2146)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:601)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:111)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:45)\n at com.sun.jmx.mbeanserver.MBeanIntrospector.invokeM(MBeanIntrospector.java:235)\n at com.sun.jmx.mbeanserver.PerInterface.invoke(PerInterface.java:138)\n at com.sun.jmx.mbeanserver.MBeanSupport.invoke(MBeanSupport.java:250)\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.invoke(DefaultMBeanServerInterceptor.java:819)\n at com.sun.jmx.mbeanserver.JmxMBeanServer.invoke(JmxMBeanServer.java:791)\n at javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1447)\n at javax.management.remote.rmi.RMIConnectionImpl.access$200(RMIConnectionImpl.java:89)\n at javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1292)\n at javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1380)\n at javax.management.remote.rmi.RMIConnectionImpl.invoke(RMIConnectionImpl.java:812)\n at sun.reflect.GeneratedMethodAccessor40.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:601)\n at sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:322)\n at sun.rmi.transport.Transport$1.run(Transport.java:177)\n at sun.rmi.transport.Transport$1.run(Transport.java:174)\n at java.security.AccessController.doPrivileged(Native Method)\n at sun.rmi.transport.Transport.serviceCall(Transport.java:173)\n at sun.rmi.transport.tcp.TCPTransport.handleMessages(TCPTransport.java:553)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run0(TCPTransport.java:808)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run(TCPTransport.java:667)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n at java.lang.Thread.run(Thread.java:722)\nCaused by: java.lang.AssertionError: originally calculated column size of 91281671 but now it is 91281729\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n at org.apache.cassandra.db.compaction.CompactionManager$6.runMayThrow(CompactionManager.java:355)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)","created":"2013-06-17T15:58:26.898+0000"},{"body":"I should add that I was upgrading from a 1.1 cluster.","created":"2013-06-17T17:19:31.891+0000"},{"body":"Any clues? We are not able to get one of our critical CFs in good shape as we cannot run compactions on it and its performance is deteriorating. ","created":"2013-06-26T17:33:56.296+0000"},{"body":"I think these issues are the same, just different versions of C*.","created":"2013-06-26T17:44:05.247+0000"},{"body":"Here is our exact scenario:\n\nSome bad code, resulted in a 600Mb row in one of our CFs. It is spanned across 2 SSTables. We've deleted it but it does not go away and nodetool cfstats still reports the max row size being the size of this row. It is causing problems because anytime this SSTable is part of a compaction or we do range queries on this CF, we end up with slow 16 sec GCs. I have tried the following to try and clean it up, but it does not go away:\n\n1. nodetool compact with sizeTiered;\n2. LCS with large SSTables and wait till it decides to promote the 2 SSTables and merge them;\n3. sizeTiered with forceUserDefined compaction trying to compact the 2 SSTable with the deleted key.\n\nEvery time we end up getting the exception complaining about calculated columns size like above. scrub, repair, rebuild, upgradesstable, no luck!\n\nAny love? ","created":"2013-07-16T23:15:39.362+0000"},{"body":"I don't remember which bug but there was one bug where Johnathan and Sylvain were arguing about discrepancy of column calculation logic in various places in the code which depended on gc_grace, deleted columns, etc. I took that as a hint and applied it to my scenario. I increased the gc_grace for the CF in question to a large number like 1 year. Then ran user defined compactions and they happily ran. Afther that I change the gc_grace setting back to 10 days and ran compactions again and my tombsones got cleaned up without seeing that exception. So, although, it didn't make sense to me what happened exactly, but at least this could be worth a try for those who are stuck.","created":"2013-08-21T21:11:23.452+0000"},{"body":"FTR - seeing this currently in 1.2.8 on batchlog compaction attempt (unfortunately I can't easily modify the gc_grace for this). Stack trace:\n{code}\nERROR [CompactionExecutor:105] 2013-08-27 17:54:39,942 CassandraDaemon.java (line 192) Exception in thread Thread[CompactionExecutor:105,1,main]\njava.lang.AssertionError: originally calculated column size of 17391408 but now it is 17391426\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n at org.apache.cassandra.db.compaction.CompactionManager$7.runMayThrow(CompactionManager.java:445)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n at java.util.concurrent.FutureTask.run(FutureTask.java:166)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:724)\nERROR [OptionalTasks:1] 2013-08-27 17:54:39,942 CassandraDaemon.java (line 192) Exception in thread Thread[OptionalTasks:1,5,main]\n{code}","created":"2013-08-27T16:46:53.140+0000"},{"body":"TBH I'm not sure we're going to fix this in 1.2.x. If you have a snapshot set of sstables that can reproduce it, then we can dig in, but eyeballing the code hasn't fixed it yet (despite multiple efforts) and probably won't.\n\nThe good news is that 2.0 fixed it by always doing single-pass compaction.","created":"2013-08-27T17:17:37.563+0000"},{"body":"We are also seeing the same error during compaction. We can provide the sstables if that helps resolving the issues.\n\nINFO [CompactionExecutor:74953] 2013-09-26 14:23:53,978 CompactionController.java (line 166) Compacting large row iqtell/mail_folder_data_subject_withdate_asc:97995 (131986528 bytes) incrementally\nERROR [CompactionExecutor:74953] 2013-09-26 14:24:01,126 CassandraDaemon.java (line 174) Exception in thread Thread[CompactionExecutor:74953,1,main]\njava.lang.AssertionError: originally calculated column size of 131986437 but now it is 131986500\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:188)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n at java.util.concurrent.FutureTask.run(FutureTask.java:166)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1146)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:679)\n","created":"2013-09-26T14:45:24.837+0000"},{"body":"Seen the same error in 1.2.10 with multithreaded_compaction disabled. It causes \"nodetool rebuild\" to fail making it really hard to rebuild a DC fully.","created":"2013-11-24T20:20:50.934+0000"},{"body":"Seen with 1.2.11:\n\n{code}\njava.lang.AssertionError: originally calculated column size of 44470356 but now it is 44470410\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:208)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:441)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n{code}","created":"2013-12-06T19:31:16.557+0000"},{"body":"For starters, could we add the SSTable filenames to the Exception? That way users could submit broken SStable files to this ticket in the future.","created":"2013-12-06T19:49:05.511+0000"},{"body":"We're seeing the AssertionError in LazilyCompactedRow a lot too on 1.2.8. Not a problem with hinted handoffs for us. We just have a large CF (close to 1TB on disk) that has had most of its row keys deleted at this point. Lots of pending compactions on many nodes and these errors seem to stymie progress. Hoping to get those compactions to finish to reclaim disk...\n\nHas anybody figured out a decent workaround? Should we try disabling multithreaded_compaction? Looks like folks are still seeing the errors with that off (it's on for us).\n\nHow stupid would running Cassandra with assertions off be?\n\nAlternatively, has anybody who's had this problem attempted to upgrade to 2.0 and had the problem fixed?","created":"2013-12-21T20:43:58.591+0000"},{"body":"I'll add that this CF is using leveled compaction in case that's useful.","created":"2013-12-21T20:52:02.110+0000"},{"body":"Disabling assertions to \"fix\" this is a great way to corrupt your sstables and get errors at read time later on.\n\nThis is definitely fixed in 2.0.","created":"2013-12-21T21:12:35.934+0000"},{"body":"That's sort of what I figured. That's why I asked how stupid it would be. The answer is \"very\", clearly.\n\nI understand that it's fixed in 2.0 but many of us are still on 1.2.x. Speaking for myself, upgrading to 2.0 is on the roadmap but I'd prefer to do that as part of a staged rollout and not as an attempt to fix what seems like a bug.\n\nDo you think disabling multithreaded_compaction would help?","created":"2013-12-21T21:19:33.560+0000"},{"body":"Doubt it. Suspect only workaround is to increase in_memory_compaction limit.","created":"2013-12-22T00:21:42.593+0000"},{"body":"I can confirm that this error also occurs with multithreaded_compaction=false.","created":"2013-12-22T11:06:33.285+0000"},{"body":"We got desperate and tried multithreaded_compaction=false yesterday afternoon and it seems to have worked for us. Pending compactions across the cluster have dropped from over 20,000 to under 600 (we're not quite done yet). The numbers were so high because we added six new nodes and they were unable to compact quickly until we made this settings change.\n\nNo idea why it worked for us but thought I'd share our experience. If there's any other useful information I can share, let me know.","created":"2013-12-22T18:05:48.935+0000"},{"body":"We are seeing this as well with 1.2.11. As was mentioned above, knowing which CF would very useful here. It seems to be happening to hints. The only other major action we are seeing that is out of the ordinary is thousands of hint SSTables being transferred at a time. Here is the Java error:\n\n{quote}\nERROR [HintedHandoff:6] 2014-01-01 16:45:19,914 CassandraDaemon.java (line 191) Exception in thread Thread[HintedHandoff:6,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 1028119265 but now it is 1028119453\n\tat org.apache.cassandra.db.HintedHandOffManager.doDeliverHintsToEndpoint(HintedHandOffManager.java:436)\n\tat org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:282)\n\tat org.apache.cassandra.db.HintedHandOffManager.access$300(HintedHandOffManager.java:90)\n\tat org.apache.cassandra.db.HintedHandOffManager$4.run(HintedHandOffManager.java:502)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:724)\nCaused by: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 1028119265 but now it is 1028119453\n\tat java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:252)\n\tat java.util.concurrent.FutureTask.get(FutureTask.java:111)\n\tat org.apache.cassandra.db.HintedHandOffManager.doDeliverHintsToEndpoint(HintedHandOffManager.java:432)\n\t... 6 more\nCaused by: java.lang.AssertionError: originally calculated column size of 1028119265 but now it is 1028119453\n\tat org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n\tat org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n\tat org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162)\n\tat org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n\tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n\tat org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n\tat org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n\tat org.apache.cassandra.db.compaction.CompactionManager$7.runMayThrow(CompactionManager.java:442)\n\tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:166)\n\t... 3 more\n{quote}\n\nMultithreaded compactions are set to false in our cluster on all nodes. We also don't have any pending compactions in the cluster. Just seeing this error a lot in the logs. The error seems to happen more frequently during bootstraps or repairs that have a lot of work to do.","created":"2014-01-01T16:52:40.653+0000"},{"body":"I am also seeing this happening during compaction with Cassandra 1.2.13 on a non-hints column family. I have multithreaded compaction disabled. \n\nHere is one new piece of information: based on the thread name the exception occurred on, I searched back through the Cassandra log and found a 'Compacting large row' message. I recognized the row key from *that* message because it has come up before within the context of a bug in our client code which resulted in runaway writes to Cassandra. Its entirely possible that that particular row had more than 2^31 columns written to it. \n\nThis smells like an integer overflow bug to me although I have not had a chance to dig into the Cassandra code yet.","created":"2014-01-15T17:49:16.436+0000"},{"body":"We're seeing this with 1.2.15, and seems related with large rows being compacted incrementally and being changed at the same time.\n\nWe upgraded from 1.1.5 and did not have this assertion error before.","created":"2014-03-18T14:12:54.971+0000"},{"body":"This was present in 1.1 and earlier releases as well. The fix is to upgrade to 2.0.x.","created":"2014-03-18T14:35:23.995+0000"},{"body":"+1 just recently upgraded to 2.0.6 and all the errors went away.","created":"2014-03-18T14:38:04.481+0000"},{"body":"{quote}\nThis was present in 1.1 and earlier releases as well. The fix is to upgrade to 2.0.x.\n{quote}\nWhich change in the 2.0 series fixes this, and how? Do you have a JIRA number we could refer to?","created":"2014-03-18T15:09:00.335+0000"},{"body":"I think the fix that is mentioned is the fact that 2.0 does only single-pass compactions, so it will never be in the situation of having calculated a previous value that has changed.\n\nIt seems an architectural change, and not something easily backported.","created":"2014-03-18T15:13:00.697+0000"},{"body":"Same issues on v 1.2.12 and 1.2.18\nError occurred during compaction\njava.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 116397997 but now it is 116398382\n...","created":"2014-07-24T14:56:33.737+0000"},{"body":"The root cause of this in 1.2 is CASSANDRA-7808.","created":"2014-08-22T16:15:15.844+0000"}],"conversations":[{"body":"I've 4 node cluster of Cassandra 1.0.9. There is a rfTest3 keyspace with RF=3 and one CF with two secondary indexes. I'm importing data into this CF using Hadoop Mapreduce job, each row has less than 10 colkumns. From JMX:\nMaxRowSize: 1597\nMeanRowSize: 369\n\nAnd there are some tens of millions of rows.\n\nIt's write-heavy usage and there is a big pressure on each node, there are quite some dropped mutations on each node. After ~12 hours of inserting I see these assertion exceptiona on 3 out of four nodes:\n\n{noformat}\nERROR 06:25:40,124 Fatal exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException:\njava.lang.AssertionError: originally calculated column size of 629444349 but now it is 588008950\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:388)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:256)\n at org.apache.cassandra.db.HintedHandOffManager.access$300(HintedHandOffManager.java:84)\n at org.apache.cassandra.db.HintedHandOffManager$3.runMayThrow(HintedHandOffManager.java:437)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.util.concurrent.ExecutionException:\njava.lang.AssertionError: originally calculated column size of\n629444349 but now it is 588008950\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:384)\n ... 7 more\nCaused by: java.lang.AssertionError: originally calculated column size\nof 629444349 but now it is 588008950\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:124)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:161)\n at org.apache.cassandra.db.compaction.CompactionManager$7.call(CompactionManager.java:380)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n ... 3 more\n{noformat}\n\nFew lines regarding Hints from the output.log:\n\n{noformat}\n INFO 06:21:26,202 Compacting large row system/HintsColumnFamily:70000000000000000000000000000000 (1712834057 bytes) incrementally\n INFO 06:22:52,610 Compacting large row system/HintsColumnFamily:10000000000000000000000000000000 (2616073981 bytes) incrementally\n INFO 06:22:59,111 flushing high-traffic column family CFS(Keyspace='system', ColumnFamily='HintsColumnFamily') (estimated 305147360 bytes)\n INFO 06:22:59,813 Enqueuing flush of Memtable-HintsColumnFamily@833933926(3814342/305147360 serialized/live bytes, 7452 ops)\n INFO 06:22:59,814 Writing Memtable-HintsColumnFamily@833933926(3814342/305147360 serialized/live bytes, 7452 ops)\n{noformat}\n\nI think the problem may be somehow connected to an IntegerType secondary index. I had a different problem with CF with two secondary indexes, the first UTF8Type, the second IntegerType. After a few hours of inserting data in the afternoon and midnight repair+compact, the next day I couldn't find any row using the IntegerType secondary index. The output was like this:\n\n{noformat}\n[default@rfTest3] get IndexTest where col1 = '3230727:http://zaskolak.cz/download.php';\n-------------------\nRowKey: 3230727:8383582:http://zaskolak.cz/download.php\n=> (column=col1, value=3230727:http://zaskolak.cz/download.php, timestamp=1335348630332000)\n=> (column=col2, value=8383582, timestamp=1335348630332000)\n-------------------\nRowKey: 3230727:8383583:http://zaskolak.cz/download.php\n=> (column=col1, value=3230727:http://zaskolak.cz/download.php, timestamp=1335348449078000)\n=> (column=col2, value=8383583, timestamp=1335348449078000)\n-------------------\nRowKey: 3230727:8383579:http://zaskolak.cz/download.php\n=> (column=col1, value=3230727:http://zaskolak.cz/download.php, timestamp=1335348778577000)\n=> (column=col2, value=8383579, timestamp=1335348778577000)\n\n3 Rows Returned.\nElapsed time: 292 msec(s).\n\n[default@rfTest3] get IndexTest where col2 = 8383583;\n\n0 Row Returned.\nElapsed time: 7 msec(s\n{noformat}\n\nYou can see there really is an 8383583 in col2 in on of the listed rows, but the search by secondary index returns nothing.\n\nThe Assert Exception also happend only on CF with the secondary index of IntegerType. There were also secondary indexes of UTF8Type and\nLongType types. It's the first time I've tried secondary indexes of other type than UTF8Type.\n\nRegards,\nPatrik","from":"reporter","subject":"AssertionError: originally calculated column size of 629444349 but now it is 588008950"},{"body":"Do you have {{multithreaded_compaction}} enabled?","from":"developer"},{"body":"bq. It's the first time I've tried secondary indexes of other type than UTF8Type\n\nI think that's a red herring; the error occurs when compacting the internal hints columnfamily, which has no indexes on it.","from":"developer"},{"body":"bq. After a few hours of inserting data in the afternoon and midnight repair+compact, the next day I couldn't find any row using the IntegerType secondary index\n\nThat part is definitely a [separate] bug, assuming the index completed building.","from":"developer"},{"body":"I'm not sure how useful it is at this point, but just for the record I'm seeing what appears to be the same issue with 1.0.7 (which means it shouldn't be CASSANDRA-3579):\n\n{noformat}\n INFO [GossipTasks:1] 2012-08-27 14:22:50,242 Gossiper.java (line 818) InetAddress /xx.xx.178.59 is now dead.\n INFO [GossipStage:1] 2012-08-27 14:22:59,090 Gossiper.java (line 804) InetAddress /xx.xx.178.59 is now UP\n INFO [HintedHandoff:1] 2012-08-27 14:23:41,548 HintedHandOffManager.java (line 296) Started hinted handoff for token: 132332031580364958013534569556798748899 with IP: /xx.xx.178.59\n INFO [HintedHandoff:1] 2012-08-27 14:23:41,870 ColumnFamilyStore.java (line 704) Enqueuing flush of Memtable-HintsColumnFamily@2081164539(597050/47764000 serialized/live bytes, 857 ops)\n INFO [FlushWriter:181] 2012-08-27 14:23:41,870 Memtable.java (line 246) Writing Memtable-HintsColumnFamily@2081164539(597050/47764000 serialized/live bytes, 857 ops)\n INFO [FlushWriter:181] 2012-08-27 14:23:41,959 Memtable.java (line 283) Completed flushing /xx/xx/xx/cassandra/datafile/system/HintsColumnFamily-hc-6730-Data.db (624946 bytes)\n INFO [CompactionExecutor:884] 2012-08-27 14:23:41,961 CompactionTask.java (line 113) Compacting [SSTableReader(path='/xx/xx/xx/cassandra/datafile/system/HintsColumnFamily-hc-6729-Data.db'), SSTableReader(path='/ngs/app/xcardp/cassandra/datafile/system/HintsColumnFamily-hc-6730-Data.db')]\n INFO [CompactionExecutor:884] 2012-08-27 14:23:41,987 CompactionController.java (line 133) Compacting large row system/HintsColumnFamily:31372e33342e3137382e3534 (274816343 bytes) incrementally\nERROR [CompactionExecutor:884] 2012-08-27 14:23:56,322 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[CompactionExecutor:884,1,main]\njava.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:124)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:159)\n at org.apache.cassandra.db.compaction.CompactionManager$6.call(CompactionManager.java:275)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nERROR [HintedHandoff:1] 2012-08-27 14:23:56,323 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:369)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:248)\n at org.apache.cassandra.db.HintedHandOffManager.access$200(HintedHandOffManager.java:84)\n at org.apache.cassandra.db.HintedHandOffManager$3.runMayThrow(HintedHandOffManager.java:416)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:365)\n ... 7 more\nCaused by: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:124)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:159)\n at org.apache.cassandra.db.compaction.CompactionManager$6.call(CompactionManager.java:275)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n ... 3 more\nERROR [HintedHandoff:1] 2012-08-27 14:23:56,345 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:369)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:248)\n at org.apache.cassandra.db.HintedHandOffManager.access$200(HintedHandOffManager.java:84)\n at org.apache.cassandra.db.HintedHandOffManager$3.runMayThrow(HintedHandOffManager.java:416)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:365)\n ... 7 more\nCaused by: java.lang.AssertionError: originally calculated column size of 197713629 but now it is 197711561\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:124)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:159)\n at org.apache.cassandra.db.compaction.CompactionManager$6.call(CompactionManager.java:275)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n ... 3 more\n INFO [COMMIT-LOG-WRITER] 2012-08-27 14:48:20,812 CommitLogSegment.java (line 60) Creating new commitlog segment /xxx/xxx/xxx/commitlog/CommitLog-1346078900812.log\n{noformat}\n\nI should note that this node was upgraded from 0.7.4 about three days prior.","from":"developer"},{"body":"Just had this pop up on a 1.1.6 node after doing a 'nodetool removetoken' for a node that had collected a lot of hints.","from":"developer"},{"body":"I am seeing what seems like the same issue on 1.2.2 on nodes that had collected a lot of hints after a node had been down for a couple hours.\n\n{noformat}\nERROR [CompactionExecutor:104] 2013-03-01 00:29:54,211 CassandraDaemon.java (line 132) Exception in thread Thread[CompactionExecutor:104,1,main]\njava.lang.AssertionError: originally calculated column size of 350273328 but now it is 350297055\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:163)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:59)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:62)\n at org.apache.cassandra.db.compaction.CompactionManager$7.runMayThrow(CompactionManager.java:422)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source)\n at java.util.concurrent.FutureTask$Sync.innerRun(Unknown Source)\n at java.util.concurrent.FutureTask.run(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\nERROR [HintedHandoff:1] 2013-03-01 00:29:54,211 CassandraDaemon.java (line 132) Exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 350273328 but now it is 350297055\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:406)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:252)\n at org.apache.cassandra.db.HintedHandOffManager.access$300(HintedHandOffManager.java:89)\n at org.apache.cassandra.db.HintedHandOffManager$4.runMayThrow(HintedHandOffManager.java:459)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\nCaused by: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 350273328 but now it is 350297055\n at java.util.concurrent.FutureTask$Sync.innerGet(Unknown Source)\n at java.util.concurrent.FutureTask.get(Unknown Source)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:402)\n ... 7 more\nCaused by: java.lang.AssertionError: originally calculated column size of 350273328 but now it is 350297055\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:163)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:59)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:62)\n at org.apache.cassandra.db.compaction.CompactionManager$7.runMayThrow(CompactionManager.java:422)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source)\n at java.util.concurrent.FutureTask$Sync.innerRun(Unknown Source)\n at java.util.concurrent.FutureTask.run(Unknown Source)\n ... 3 more\n{noformat}","from":"developer"},{"body":"Version: 1.2.2 - 6 nodes, RF3\nOS: Centos 6.3\n\nLooking closer at the errors, I'm only seeing this for the system/hints:\n\n{noformat}\n INFO [CompactionExecutor:555] 2013-03-01 15:37:10,737 CompactionController.java (line 158) Compacting large row system/hints:384ac791\n-d0d7-4f5e-b66b-57af885341d5 (350297131 bytes) incrementally\nERROR [CompactionExecutor:555] 2013-03-01 15:37:28,377 CassandraDaemon.java (line 132) Exception in thread Thread[CompactionExecutor:5\n55,1,main]\n{noformat}","from":"developer"},{"body":"I have 1600 compactions pending and no compaction is running, so they are all stuck. I also got this in the logs. And Jonathan, I have multi_threaded_compaction disabled. \n\nERROR [CompactionExecutor:111] 2013-05-02 13:54:05,913 CassandraDaemon.java (line 174) Exception in thread Thread[CompactionExecutor:111,1,main]java.lang.AssertionError: originally calculated column size of 1337269150 but now it is 1337269195 at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135) at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162) at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48) at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:188)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:439)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:895)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:918)\n at java.lang.Thread.run(Thread.java:662)\n\n","from":"developer"},{"body":"Similar thing over here.\n\nC*1.2.2\n\nI have a lot of hints (3 & 4 GB) on 2 nodes out of 12 due to an unknown issue while growing to multiple datacenter (all the nodes of the new datacenter saw themselves as UNREACHABLE from the cassandra-cli - nodetool ring is ok...).\n\nOn these 2 nodes I see now :\n\n{code}\nERROR 08:37:16,878 Exception in thread Thread[CompactionExecutor:1540,1,main]\njava.lang.AssertionError: originally calculated column size of 5469343266 but now it is 5469343506\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:163)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:59)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:62)\n at org.apache.cassandra.db.compaction.CompactionManager$7.runMayThrow(CompactionManager.java:422)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:439)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nERROR 08:37:16,878 Exception in thread Thread[HintedHandoff:176,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 5469343266 but now it is 5469343506\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:406)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:252)\n at org.apache.cassandra.db.HintedHandOffManager.access$300(HintedHandOffManager.java:89)\n at org.apache.cassandra.db.HintedHandOffManager$4.runMayThrow(HintedHandOffManager.java:459)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 5469343266 but now it is 5469343506\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:402)\n ... 7 more\nCaused by: java.lang.AssertionError: originally calculated column size of 5469343266 but now it is 5469343506\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:163)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n{code}","from":"developer"},{"body":"Seeing this on my 1.2.5 cluster since upgrading:\n\nubuntu@c2-1d:~$ nodetool -h localhost compact TRProd Timelines \nError occurred during compaction\njava.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 91281671 but now it is 91281729\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:252)\n at java.util.concurrent.FutureTask.get(FutureTask.java:111)\n at org.apache.cassandra.db.compaction.CompactionManager.performMaximal(CompactionManager.java:334)\n at org.apache.cassandra.db.ColumnFamilyStore.forceMajorCompaction(ColumnFamilyStore.java:1657)\n at org.apache.cassandra.service.StorageService.forceTableCompaction(StorageService.java:2146)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:601)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:111)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:45)\n at com.sun.jmx.mbeanserver.MBeanIntrospector.invokeM(MBeanIntrospector.java:235)\n at com.sun.jmx.mbeanserver.PerInterface.invoke(PerInterface.java:138)\n at com.sun.jmx.mbeanserver.MBeanSupport.invoke(MBeanSupport.java:250)\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.invoke(DefaultMBeanServerInterceptor.java:819)\n at com.sun.jmx.mbeanserver.JmxMBeanServer.invoke(JmxMBeanServer.java:791)\n at javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1447)\n at javax.management.remote.rmi.RMIConnectionImpl.access$200(RMIConnectionImpl.java:89)\n at javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1292)\n at javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1380)\n at javax.management.remote.rmi.RMIConnectionImpl.invoke(RMIConnectionImpl.java:812)\n at sun.reflect.GeneratedMethodAccessor40.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:601)\n at sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:322)\n at sun.rmi.transport.Transport$1.run(Transport.java:177)\n at sun.rmi.transport.Transport$1.run(Transport.java:174)\n at java.security.AccessController.doPrivileged(Native Method)\n at sun.rmi.transport.Transport.serviceCall(Transport.java:173)\n at sun.rmi.transport.tcp.TCPTransport.handleMessages(TCPTransport.java:553)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run0(TCPTransport.java:808)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run(TCPTransport.java:667)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n at java.lang.Thread.run(Thread.java:722)\nCaused by: java.lang.AssertionError: originally calculated column size of 91281671 but now it is 91281729\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n at org.apache.cassandra.db.compaction.CompactionManager$6.runMayThrow(CompactionManager.java:355)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)","from":"developer"},{"body":"I should add that I was upgrading from a 1.1 cluster.","from":"developer"},{"body":"Any clues? We are not able to get one of our critical CFs in good shape as we cannot run compactions on it and its performance is deteriorating. ","from":"developer"},{"body":"I think these issues are the same, just different versions of C*.","from":"developer"},{"body":"Here is our exact scenario:\n\nSome bad code, resulted in a 600Mb row in one of our CFs. It is spanned across 2 SSTables. We've deleted it but it does not go away and nodetool cfstats still reports the max row size being the size of this row. It is causing problems because anytime this SSTable is part of a compaction or we do range queries on this CF, we end up with slow 16 sec GCs. I have tried the following to try and clean it up, but it does not go away:\n\n1. nodetool compact with sizeTiered;\n2. LCS with large SSTables and wait till it decides to promote the 2 SSTables and merge them;\n3. sizeTiered with forceUserDefined compaction trying to compact the 2 SSTable with the deleted key.\n\nEvery time we end up getting the exception complaining about calculated columns size like above. scrub, repair, rebuild, upgradesstable, no luck!\n\nAny love? ","from":"developer"},{"body":"I don't remember which bug but there was one bug where Johnathan and Sylvain were arguing about discrepancy of column calculation logic in various places in the code which depended on gc_grace, deleted columns, etc. I took that as a hint and applied it to my scenario. I increased the gc_grace for the CF in question to a large number like 1 year. Then ran user defined compactions and they happily ran. Afther that I change the gc_grace setting back to 10 days and ran compactions again and my tombsones got cleaned up without seeing that exception. So, although, it didn't make sense to me what happened exactly, but at least this could be worth a try for those who are stuck.","from":"developer"},{"body":"FTR - seeing this currently in 1.2.8 on batchlog compaction attempt (unfortunately I can't easily modify the gc_grace for this). Stack trace:\n{code}\nERROR [CompactionExecutor:105] 2013-08-27 17:54:39,942 CassandraDaemon.java (line 192) Exception in thread Thread[CompactionExecutor:105,1,main]\njava.lang.AssertionError: originally calculated column size of 17391408 but now it is 17391426\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n at org.apache.cassandra.db.compaction.CompactionManager$7.runMayThrow(CompactionManager.java:445)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n at java.util.concurrent.FutureTask.run(FutureTask.java:166)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:724)\nERROR [OptionalTasks:1] 2013-08-27 17:54:39,942 CassandraDaemon.java (line 192) Exception in thread Thread[OptionalTasks:1,5,main]\n{code}","from":"developer"},{"body":"TBH I'm not sure we're going to fix this in 1.2.x. If you have a snapshot set of sstables that can reproduce it, then we can dig in, but eyeballing the code hasn't fixed it yet (despite multiple efforts) and probably won't.\n\nThe good news is that 2.0 fixed it by always doing single-pass compaction.","from":"developer"},{"body":"We are also seeing the same error during compaction. We can provide the sstables if that helps resolving the issues.\n\nINFO [CompactionExecutor:74953] 2013-09-26 14:23:53,978 CompactionController.java (line 166) Compacting large row iqtell/mail_folder_data_subject_withdate_asc:97995 (131986528 bytes) incrementally\nERROR [CompactionExecutor:74953] 2013-09-26 14:24:01,126 CassandraDaemon.java (line 174) Exception in thread Thread[CompactionExecutor:74953,1,main]\njava.lang.AssertionError: originally calculated column size of 131986437 but now it is 131986500\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:159)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:188)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n at java.util.concurrent.FutureTask.run(FutureTask.java:166)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1146)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:679)\n","from":"developer"},{"body":"Seen the same error in 1.2.10 with multithreaded_compaction disabled. It causes \"nodetool rebuild\" to fail making it really hard to rebuild a DC fully.","from":"developer"},{"body":"Seen with 1.2.11:\n\n{code}\njava.lang.AssertionError: originally calculated column size of 44470356 but now it is 44470410\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:208)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:441)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\n{code}","from":"developer"},{"body":"For starters, could we add the SSTable filenames to the Exception? That way users could submit broken SStable files to this ticket in the future.","from":"developer"},{"body":"We're seeing the AssertionError in LazilyCompactedRow a lot too on 1.2.8. Not a problem with hinted handoffs for us. We just have a large CF (close to 1TB on disk) that has had most of its row keys deleted at this point. Lots of pending compactions on many nodes and these errors seem to stymie progress. Hoping to get those compactions to finish to reclaim disk...\n\nHas anybody figured out a decent workaround? Should we try disabling multithreaded_compaction? Looks like folks are still seeing the errors with that off (it's on for us).\n\nHow stupid would running Cassandra with assertions off be?\n\nAlternatively, has anybody who's had this problem attempted to upgrade to 2.0 and had the problem fixed?","from":"developer"},{"body":"I'll add that this CF is using leveled compaction in case that's useful.","from":"developer"},{"body":"Disabling assertions to \"fix\" this is a great way to corrupt your sstables and get errors at read time later on.\n\nThis is definitely fixed in 2.0.","from":"developer"},{"body":"That's sort of what I figured. That's why I asked how stupid it would be. The answer is \"very\", clearly.\n\nI understand that it's fixed in 2.0 but many of us are still on 1.2.x. Speaking for myself, upgrading to 2.0 is on the roadmap but I'd prefer to do that as part of a staged rollout and not as an attempt to fix what seems like a bug.\n\nDo you think disabling multithreaded_compaction would help?","from":"developer"},{"body":"Doubt it. Suspect only workaround is to increase in_memory_compaction limit.","from":"developer"},{"body":"I can confirm that this error also occurs with multithreaded_compaction=false.","from":"developer"},{"body":"We got desperate and tried multithreaded_compaction=false yesterday afternoon and it seems to have worked for us. Pending compactions across the cluster have dropped from over 20,000 to under 600 (we're not quite done yet). The numbers were so high because we added six new nodes and they were unable to compact quickly until we made this settings change.\n\nNo idea why it worked for us but thought I'd share our experience. If there's any other useful information I can share, let me know.","from":"developer"},{"body":"We are seeing this as well with 1.2.11. As was mentioned above, knowing which CF would very useful here. It seems to be happening to hints. The only other major action we are seeing that is out of the ordinary is thousands of hint SSTables being transferred at a time. Here is the Java error:\n\n{quote}\nERROR [HintedHandoff:6] 2014-01-01 16:45:19,914 CassandraDaemon.java (line 191) Exception in thread Thread[HintedHandoff:6,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 1028119265 but now it is 1028119453\n\tat org.apache.cassandra.db.HintedHandOffManager.doDeliverHintsToEndpoint(HintedHandOffManager.java:436)\n\tat org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:282)\n\tat org.apache.cassandra.db.HintedHandOffManager.access$300(HintedHandOffManager.java:90)\n\tat org.apache.cassandra.db.HintedHandOffManager$4.run(HintedHandOffManager.java:502)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:724)\nCaused by: java.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 1028119265 but now it is 1028119453\n\tat java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:252)\n\tat java.util.concurrent.FutureTask.get(FutureTask.java:111)\n\tat org.apache.cassandra.db.HintedHandOffManager.doDeliverHintsToEndpoint(HintedHandOffManager.java:432)\n\t... 6 more\nCaused by: java.lang.AssertionError: originally calculated column size of 1028119265 but now it is 1028119453\n\tat org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:135)\n\tat org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n\tat org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:162)\n\tat org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n\tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n\tat org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:58)\n\tat org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:60)\n\tat org.apache.cassandra.db.compaction.CompactionManager$7.runMayThrow(CompactionManager.java:442)\n\tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:166)\n\t... 3 more\n{quote}\n\nMultithreaded compactions are set to false in our cluster on all nodes. We also don't have any pending compactions in the cluster. Just seeing this error a lot in the logs. The error seems to happen more frequently during bootstraps or repairs that have a lot of work to do.","from":"developer"},{"body":"I am also seeing this happening during compaction with Cassandra 1.2.13 on a non-hints column family. I have multithreaded compaction disabled. \n\nHere is one new piece of information: based on the thread name the exception occurred on, I searched back through the Cassandra log and found a 'Compacting large row' message. I recognized the row key from *that* message because it has come up before within the context of a bug in our client code which resulted in runaway writes to Cassandra. Its entirely possible that that particular row had more than 2^31 columns written to it. \n\nThis smells like an integer overflow bug to me although I have not had a chance to dig into the Cassandra code yet.","from":"developer"},{"body":"We're seeing this with 1.2.15, and seems related with large rows being compacted incrementally and being changed at the same time.\n\nWe upgraded from 1.1.5 and did not have this assertion error before.","from":"developer"},{"body":"This was present in 1.1 and earlier releases as well. The fix is to upgrade to 2.0.x.","from":"developer"},{"body":"+1 just recently upgraded to 2.0.6 and all the errors went away.","from":"developer"},{"body":"{quote}\nThis was present in 1.1 and earlier releases as well. The fix is to upgrade to 2.0.x.\n{quote}\nWhich change in the 2.0 series fixes this, and how? Do you have a JIRA number we could refer to?","from":"developer"},{"body":"I think the fix that is mentioned is the fact that 2.0 does only single-pass compactions, so it will never be in the situation of having calculated a previous value that has changed.\n\nIt seems an architectural change, and not something easily backported.","from":"developer"},{"body":"Same issues on v 1.2.12 and 1.2.18\nError occurred during compaction\njava.util.concurrent.ExecutionException: java.lang.AssertionError: originally calculated column size of 116397997 but now it is 116398382\n...","from":"developer"},{"body":"The root cause of this in 1.2 is CASSANDRA-7808.","from":"developer"}],"created":"2012-05-01T11:28:44.000+0000","description":"I've 4 node cluster of Cassandra 1.0.9. There is a rfTest3 keyspace with RF=3 and one CF with two secondary indexes. I'm importing data into this CF using Hadoop Mapreduce job, each row has less than 10 colkumns. From JMX:\nMaxRowSize: 1597\nMeanRowSize: 369\n\nAnd there are some tens of millions of rows.\n\nIt's write-heavy usage and there is a big pressure on each node, there are quite some dropped mutations on each node. After ~12 hours of inserting I see these assertion exceptiona on 3 out of four nodes:\n\n{noformat}\nERROR 06:25:40,124 Fatal exception in thread Thread[HintedHandoff:1,1,main]\njava.lang.RuntimeException: java.util.concurrent.ExecutionException:\njava.lang.AssertionError: originally calculated column size of 629444349 but now it is 588008950\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:388)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:256)\n at org.apache.cassandra.db.HintedHandOffManager.access$300(HintedHandOffManager.java:84)\n at org.apache.cassandra.db.HintedHandOffManager$3.runMayThrow(HintedHandOffManager.java:437)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:30)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: java.util.concurrent.ExecutionException:\njava.lang.AssertionError: originally calculated column size of\n629444349 but now it is 588008950\n at java.util.concurrent.FutureTask$Sync.innerGet(FutureTask.java:222)\n at java.util.concurrent.FutureTask.get(FutureTask.java:83)\n at org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpointInternal(HintedHandOffManager.java:384)\n ... 7 more\nCaused by: java.lang.AssertionError: originally calculated column size\nof 629444349 but now it is 588008950\n at org.apache.cassandra.db.compaction.LazilyCompactedRow.write(LazilyCompactedRow.java:124)\n at org.apache.cassandra.io.sstable.SSTableWriter.append(SSTableWriter.java:160)\n at org.apache.cassandra.db.compaction.CompactionTask.execute(CompactionTask.java:161)\n at org.apache.cassandra.db.compaction.CompactionManager$7.call(CompactionManager.java:380)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n at java.util.concurrent.FutureTask.run(FutureTask.java:138)\n ... 3 more\n{noformat}\n\nFew lines regarding Hints from the output.log:\n\n{noformat}\n INFO 06:21:26,202 Compacting large row system/HintsColumnFamily:70000000000000000000000000000000 (1712834057 bytes) incrementally\n INFO 06:22:52,610 Compacting large row system/HintsColumnFamily:10000000000000000000000000000000 (2616073981 bytes) incrementally\n INFO 06:22:59,111 flushing high-traffic column family CFS(Keyspace='system', ColumnFamily='HintsColumnFamily') (estimated 305147360 bytes)\n INFO 06:22:59,813 Enqueuing flush of Memtable-HintsColumnFamily@833933926(3814342/305147360 serialized/live bytes, 7452 ops)\n INFO 06:22:59,814 Writing Memtable-HintsColumnFamily@833933926(3814342/305147360 serialized/live bytes, 7452 ops)\n{noformat}\n\nI think the problem may be somehow connected to an IntegerType secondary index. I had a different problem with CF with two secondary indexes, the first UTF8Type, the second IntegerType. After a few hours of inserting data in the afternoon and midnight repair+compact, the next day I couldn't find any row using the IntegerType secondary index. The output was like this:\n\n{noformat}\n[default@rfTest3] get IndexTest where col1 = '3230727:http://zaskolak.cz/download.php';\n-------------------\nRowKey: 3230727:8383582:http://zaskolak.cz/download.php\n=> (column=col1, value=3230727:http://zaskolak.cz/download.php, timestamp=1335348630332000)\n=> (column=col2, value=8383582, timestamp=1335348630332000)\n-------------------\nRowKey: 3230727:8383583:http://zaskolak.cz/download.php\n=> (column=col1, value=3230727:http://zaskolak.cz/download.php, timestamp=1335348449078000)\n=> (column=col2, value=8383583, timestamp=1335348449078000)\n-------------------\nRowKey: 3230727:8383579:http://zaskolak.cz/download.php\n=> (column=col1, value=3230727:http://zaskolak.cz/download.php, timestamp=1335348778577000)\n=> (column=col2, value=8383579, timestamp=1335348778577000)\n\n3 Rows Returned.\nElapsed time: 292 msec(s).\n\n[default@rfTest3] get IndexTest where col2 = 8383583;\n\n0 Row Returned.\nElapsed time: 7 msec(s\n{noformat}\n\nYou can see there really is an 8383583 in col2 in on of the listed rows, but the search by secondary index returns nothing.\n\nThe Assert Exception also happend only on CF with the secondary index of IntegerType. There were also secondary indexes of UTF8Type and\nLongType types. It's the first time I've tried secondary indexes of other type than UTF8Type.\n\nRegards,\nPatrik","issue_id":"12553475","key":"CASSANDRA-4206","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-03-18T14:35:23.000+0000","role":"fixed_distractor","summary":"AssertionError: originally calculated column size of 629444349 but now it is 588008950"} {"case_id":"12595934","cluster":"DISTRACTOR-CASSANDRA-4377","comments":[{"body":"What happens is that columns metadata for composite CQL3 columns refers to only one of the component of the column name. So internally they have a 'componentIndex' parameter which allow to handle them correctly (otherwise you don't know which comparator to use to display those column metadata). However, we decided that thrift users shouldn't have to care about this, so we don't expose those metadata to thrift at all. In other words, the CF does have metadata for 'sum' and 'uniques' internally, but they are not exposed to thrift.\n\nSo I guess it is more of a cqlsh problem that shouldn't use the thrift describe call for CQL3. Instead, it should directly query the system.keyspace, system.columnfamilies and system.columns table. However, it'd be much more easier to do that with CASSANDRA-4018, but that is not in 1.1.","created":"2012-06-27T07:00:03.317+0000"},{"body":"bq. So I guess it is more of a cqlsh problem that shouldn't use the thrift describe call for CQL3.\n\nThe main problem here from my perspective is that it was impossible to insert data into the column family except with cqlsh. Using a basic thrift batch mutate failed on the validation step because it tried to validate the value in the sum column as text (default_validation_class).","created":"2012-06-27T14:45:47.994+0000"},{"body":"True, I'm unable to save any data to Cassandra.\nTrying to save with Hector (uses Mutators):\n#","created":"2012-06-27T22:01:56.813+0000"},{"body":"bq. it was impossible to insert data into the column family except with cqlsh. Using a basic thrift batch mutate failed on the validation step because it tried to validate the value\n\nAlright, that part is not expected and is likely a bug.\n\nThough I note that even if we fix that, since we don't expose on the thrift side everything needed to interpret the value correctly, thrift client won't be able to work with those kind of table very well. In particular they won't know how to interpret the value of a get. So I guess we need to either say clearly that CQL3 table are not meant to be accessed from thrift, or we should probably start exposing all the column metatada on the thrift side (but we need to expose the componentIndex part of the metadata in particular and the discussion on CASSANDRA-4093 is relevant for that).","created":"2012-06-28T11:51:57.381+0000"},{"body":"Attaching simple patch to allow the insertion in the case above. Truth is, I'm not really satisfied (though I don't have a clearly better option) by such a patch for 2 reasons:\n# it will break the case for thrift where people were using compositeType and column_metadata on them. That might be 0 people we're talking about but it's still a bit annoying. An alternative here would be that for each composite column, we iterater over all column_metadata and check where one apply. This would work but this feels butt ugly.\n# it feels to me we supporting either not enough or too much in the thrift side. Even if we support this, we still don't expose the CQL3 metadata to thrift, so one has to know the CQL schema definition to be able to create the column in the first place (or deserialize it on read). In particular, most advanced thrift client will likely still break at one point or another.\n\nOverall I see two reasonable approaches:\n* Either we start exposing enough on the thirft side so that thrift client can work correctly with CQL3. While this may be doable reasonably easily now, this will be increasingly difficult with things like CASSANDRA-4179, CASSANDRA-3647, ... Besides, even if we make it possible to work with CQL3 table, it doesn't mean it will be convenient since thrift won't do the grouping of columns in sparse table.\n* Be clear that you cannot work with CQL3 created table from thrift.\n\nBut imo the in-the-middle approach that this patch would start takes the risk of polluting the thrift side without adding much.\n","created":"2012-07-10T15:40:00.520+0000"},{"body":"To be clear, if this issue is for instance a blocker for map-reduce for CQL3, I'm not against committing it, but on a more long term/general level, I do want to note that accessing CQL3 tables to thrift is imho a larger problem and we should be clear on what we want to guarantee.","created":"2012-07-10T15:48:32.567+0000"},{"body":"It doesn't sound like its going to be possible to make things work well in the case of accessing cql3 data from thrift.\n\nIf thats the case I'm in favor of making that incompatibility *very* explicit. I would say explicitly throwing an exception when trying to access a cql3 column family from thrift would be an acceptable solution.\n\nThe fact that it took me quite a bit of time as well as digging around in both client code and cassandra source code to figure this bug out makes me worried for other users hitting similar problems.","created":"2012-07-10T16:32:29.208+0000"},{"body":"bq. I would say explicitly throwing an exception when trying to access a cql3 column family from thrift would be an acceptable solution.\n\n-1 on that; it's very useful to have a lower-level access method to look at the \"storage engine\" layer instead of the logical CQL3 rows at times.","created":"2012-07-10T17:11:55.766+0000"},{"body":"so disable write access to cql3 cfs from thrift? And leave read access with the qualification that everything will be returned as BytesType?\n\nThe only other option I see is to say that we are going to identify and fix all bugs like this one.","created":"2012-07-10T17:14:51.875+0000"},{"body":"bq. disable write access to cql3 cfs from thrift? And leave read access with the qualification that everything will be returned as BytesType?\n\nSounds reasonable to me.","created":"2012-07-10T17:22:11.767+0000"},{"body":"bq. disable write access to cql3 cfs from thrift? And leave read access with the qualification that everything will be returned as BytesType?\n\nA variation in the same spirit could be to disable access to CQL3 table by default but add a thrift debug mode (either per-connection through or globally through JMX if we don't want to add a new thrift method) that would make thrift disable every validation. That way, by default the message that CQL3 tables should be access through CQL3 would be clear but we would have low-level read/write for debugging. ","created":"2012-07-10T17:44:09.355+0000"},{"body":"bq. disable access to CQL3 table by default but add a thrift debug mode\n\n+lots.","created":"2012-07-10T17:46:40.964+0000"},{"body":"sgtm","created":"2012-07-10T17:53:15.128+0000"},{"body":"+1\n\n(and I'd prefer JMX otherwise I foresee clients using this as an escape hatch and fubarring things up.)","created":"2012-07-10T18:15:42.307+0000"},{"body":"So, from an outside conversation, it sounds like there may be some confusion on exactly what is being discussed here. Are we talking about disabling Thrift access entirely to columnfamilies which use named metadata with composites, or are we just talking about not supporting Thrift addressing columns inside composites by their CQL3 names?\n\nThe second seems eminently reasonable. The former sounds crazy.","created":"2012-07-10T22:38:17.067+0000"},{"body":"bq. are we just talking about not supporting Thrift addressing columns inside composites by their CQL3 names\n\nThat is not what I was talking about, but I don't even understand how we could ever support that.\n\nbq. disabling Thrift access entirely to columnfamilies which use named metadata with composites\n\nThat is what I'm talking about (though to be precise, it would be for composites using named metadata *created through CQL3*). Anyway, I don't know if that's so crazy but in any case calling it crazy doesn't help solving the problem.\n\n","created":"2012-07-11T07:36:29.842+0000"},{"body":"So here is where I ended up on this.\n\nAssuming we aren't extremely interested in fixing bugs like these when the come up then I'm all for disabling thrift access by default to cql3 cfs. From what I can tell there isn't an immediately easy way to know a cf was created through cql3 at the moment but we could add that I guess.\n\nOn the other hand if we want to just go ahead and fix these bugs when they happen, I don't see much reason to go out of our way to disable thrift access to cql3 cfs.","created":"2012-07-30T23:26:39.192+0000"},{"body":"bq. one has to know the CQL schema definition to be able to create the column in the first place (or deserialize it on read)\n\nCan you remind me why we don't translate {{PRIMARY KEY(gid, period, tid)}} into a comparator of {{CompositeType(Int32Type, BytesType, UTF8Type)}}? (period -> int, tid -> bytes, sparse columns -> utf8)","created":"2012-08-09T15:17:39.599+0000"},{"body":"bq. Can you remind me why we don't translate PRIMARY KEY(gid, period, tid) into a comparator of CompositeType(Int32Type, BytesType, UTF8Type)\n\nWe do. The problem is actually with columns that are not part of the key. Let me try to sum that up:\n* Pre-CQL3, the ColumnDefinition name was always a full column name.\n* In CQL3, the ColumnDefinition name only correspond to one of the component of the column name, i.e. to the UTF8Type component above.\n* To be able to distinguish both case internally, we've introduce the componentIndex field in ColumnDefinition. However, we decided that we didn't wanted to expose this field to thrift, and so we don't expose to thrift the ColumnDefinition from CQL3 table.\n\nThe net result is that as far as thrift is concerned, the CQL3 tables have no columns_metadata whatsoever. It follows that thrift clients don't know what are the correct value for the last UTF8Type component, and don't know what is the type of the corresponding value (and thus cannot serialize/deserialize said value correctly). ","created":"2012-08-09T15:41:17.679+0000"},{"body":"So basically, the problem is we're trying to maintain compatibility with (non-cql3) wide-row named columns? If we're willing to break that scenario, does the problem go away?\n\nI still think that naming wide row columns is nonsensical, but we could add an extra layer of protection by warning on startup in 1.2 that you need to update your schema:\n\n- if you have named columns and non-utf8-or-bytes comparator\n- if you have named columns but metadata shows more columns than names","created":"2012-08-09T15:53:30.733+0000"},{"body":"I'm not sure I understand what's a named columns above to be honest.\n\nThere is basically two informations from CFMetadata you need to know to insert a column correctly in a table (CQL3 or no CQL3): the comparator and *all* of column_metadata. The comparator is necessary to know what is a valid column name and the column_metadata is necessary to know what is a valid column value (I'm simplifying a bit, I'm assuming that the key_validation and default_validator are BytesType but that doesn't matter for the problem at hand).\n\nNow the problem is that for any table created through CQL3 that doesn't use COMPACT STORAGE (let's call those CQL3 tables), all the ColumnDefinition of column_metada will have a componentIndex. So none of those ColumnDefinition are exposed in thrift. In practice it means that if I do:\n{noformat}\nCREATE TABLE user {\n user_id blob PRIMARY KEY,\n name text,\n age int\n}\n{noformat}\nthen if a thrift client do a describe, it will basically get:\n{noformat}\ncomparator = CompositeType(UTF8Type) // it's a composite so that we can add collection later on\ncolumn_metadata = []\n{noformat}\n\nAt that point we have two slightly separate problems:\n# Even if a user produces a valid column, with say a composite name being \"age\" and a value being an int, then currently the code throw an exception. Fixing that exception is the goal of the attached patch (though it would have to be updated to work with collections in 1.2). I'm fine fixing that, though I'm pointing that there is a second, more general problem.\n# Since the thrift client doesn't know about the actual column_metadata, how can we expect it to correctly insert data. In particular I'm pretty sure higher level clients like pycassa or astyanax will serialize data incorrectly if they don't know the right value validator. Besides, there is many way to be confused if you use a CQL3 table from thrift. For instance if you create the wrong column (i'ts enough to mess up the case), you'll be surprised to not be able to access it when you go back to CQL3. So be clear, I do am suggesting that we don't allow accessing table created from CQL3 *without* COMPACT STORAGE from thrift, because I think it will be more sane, even if it does mean that you're not coming back from CQL3 once you've start really using it.\n","created":"2012-08-09T16:41:37.983+0000"},{"body":"Right, so what I was saying was, if we're willing to say that columndefinition name is always the cql3 column name, then we don't need componentIndex, at the price of potentially breaking a corner case in non-cql3 schemas.","created":"2012-08-09T17:07:04.684+0000"},{"body":"But what the componentIndex give us is which of the composite component is the cql3 column name. Typically, with collections, it's not even necessarily the last of the component. Now, if you know the table is a CQL3 one, you could try to infer which component it is by saying that it's the last component, except if the last is a collection type, in which case that's the previous one, but that feels a bit messy. And besides, thrift client libraries don't have a simple automatic way to know if the table is a CQL3 one in the first place. I guess you could say that if you have a composite comparator *and* some column_metadata then you are likely a CQL3 table, but again, not very clean imo.\n\n\nAt least internally I would be in favor of keeping the componentIndex as it is cleaner. I guess we could start returning the ColumnDefinition from CQL3 table without the componentIndex and let thrift client infer what they can. As said, I still think using CQL3 table from thrift has other way to be confusing, but why not.","created":"2012-08-09T17:33:09.925+0000"},{"body":"Here's what I think our goals are, in order of importance:\n\n# Existing \"high level\" clients should have a reasonable upgrade path to read and update collections and cql3 CFs.\n# CLI and other tools that don't speak cql should have a way to tell that they can't cope with CQL3 CF definitions. This is the problem Nick described originally in this ticket description. Actually making such tools able to manipulate CQL3 definitions is NOT a goal, but we should, as Nick says, make that more obvious.\n# We should allow updates to CQL3 CFs from Thrift, if someone manually composes the correct CompositeType bytes. This is what most of the rest of the discussion here involves.\n\nAnalysis:\n\n# This we have done--they will have to use cql-over-thrift, but IMO this is reasonable. Thrift RPC methods to deal with collections have never been on the roadmap.\n# This is tough since if this involves new information (like exposing component_index, or even adding a cql3 boolean to CfDef) then old tools by definition won't know about it. *Proposal:* what if we omit CQL3 CfDefs from those we return to describe_keyspace[s] calls? Not quite as good as returning a CfDef that explicitly says \"I am here but you can't touch this,\" but compare to displaying incomplete information (that the cli doesn't know is incomplete) it's much more obvious that the cli and other old thrift-base schema manipulators can't cope with such.\n# Sylvain mentioned adding a thrift or JMX method to enable validation-free updates but I'd rather make ThriftValidation cql3-aware, which would let this work without any special flags. I don't see any downside to this except added complexity for ThriftValidation.","created":"2012-08-17T16:00:05.572+0000"},{"body":"I do think we should push this to 1.2 however, since I'm leery of changing the behavior or describe_keyspace[s] when we are fairly deep into 1.1's stable lifetime.","created":"2012-08-17T16:11:50.288+0000"},{"body":"CQL3 defined CFs won't be exposed to Thrift API anymore (with warning), I also checked CassandraServer and ThriftValidation and figured that column name validation already in place via ThriftValidation.validateColumnNames(...).","created":"2012-08-22T09:50:59.278+0000"},{"body":"Just glanced at the patch, but it looks like this just makes cql3 cfs not show up in describe_keyspaces calls right? Reading/writing to the cf would still go through and error?\n\nIf I'm looking at that right then it seems like we would want the opposite behavior. At least from my perspective, I would like to have an indication that the cf exists in my client even if i can't write to it.","created":"2012-08-31T15:25:26.789+0000"},{"body":"The idea was to do two things:\n\n- update ThriftValidation to be cql3-aware, so if a Thrift user manually composes a valid column, we accept it\n- but, we don't expose cql3 CFs via describe_keyspaces since clients do not have enough information to generate valid requests automatically","created":"2012-08-31T23:43:17.085+0000"},{"body":"I agree that the two points above (make ThriftValidation cql3-aware but don't expose cql3 CFs defs) make sense as a strategy.","created":"2012-09-03T09:21:07.768+0000"},{"body":"On Pavel's patch:\n* I'm not a fan of logging a warning (or to log anything really) if someone has CQL3 CFs. We're trying to push CQL3 as a good thing, let's not log anything that could be interpreted as if something was wrong/abnormal.\n* The check is excluding composite CF without any ColumnDefinition while it shouldn't.\n\nAttaching a v2 that:\n* Fix the two remarks above.\n* Rename the check as isThriftIncompatible. I think it's more about detecting CF definitions that cannot be exploided fully by thrift rather than discriminate between what is a thrift CF and CQL3 CF. Especially since the intersection between those two notions is not empty.\n* Ship the changes to ThriftValidation from my first patch, though modified a bit to be more generic and handle correctly collections. I'll note that this part makes validation potentially iterate over all ColumnDefinition for composite CF, but that's not really an issue since composite CF created on the thrift side are almost guaranteed to have no ColumnDefinition.\n","created":"2012-09-06T11:16:31.011+0000"},{"body":"bq. I'm not a fan of logging a warning (or to log anything really) if someone has CQL3 CFs. We're trying to push CQL3 as a good thing, let's not log anything that could be interpreted as if something was wrong/abnormal.\n\nWouldn't that create confusion when some CFs are visible through Thrift and some are not or do we rely on that users should know what is so special about CQL3 CFs?","created":"2012-09-06T11:42:47.483+0000"},{"body":"bq. Wouldn't that create confusion when some CFs are visible through Thrift and some are not\n\nI'm not saying this shouldn't be documented at all, but merely that it's a documentation issue and as such logging it at each startup is not the right place. ","created":"2012-09-06T11:52:59.458+0000"},{"body":"If everyone else is ok with that, lgtm.","created":"2012-09-06T11:57:21.532+0000"},{"body":"So, this is good to commit?","created":"2012-09-11T02:36:59.996+0000"},{"body":"Alright, committed, thanks.","created":"2012-09-11T09:03:56.115+0000"}],"conversations":[{"body":"{noformat}\ncqlsh> create keyspace test with strategy_class = 'SimpleStrategy' and strategy_options:replication_factor = 1;\ncqlsh> use test;\ncqlsh:test> CREATE TABLE stats (\n ... gid blob,\n ... period int,\n ... tid blob, \n ... sum int,\n ... uniques blob,\n ... PRIMARY KEY(gid, period, tid)\n ... );\ncqlsh:test> describe columnfamily stats;\n\nCREATE TABLE stats (\n gid blob PRIMARY KEY\n) WITH\n comment='' AND\n comparator='CompositeType(org.apache.cassandra.db.marshal.Int32Type,org.apache.cassandra.db.marshal.BytesType,org.apache.cassandra.db.marshal.UTF8Type)' AND\n read_repair_chance=0.100000 AND\n gc_grace_seconds=864000 AND\n default_validation=text AND\n min_compaction_threshold=4 AND\n max_compaction_threshold=32 AND\n replicate_on_write='true' AND\n compaction_strategy_class='SizeTieredCompactionStrategy' AND\n compression_parameters:sstable_compression='SnappyCompressor';\n{noformat}\n\nYou can see in the above output that the stats cf is created with the column validator set to text, but neither of the non primary key columns defined are text. It should either be setting metadata for those columns or not setting a default validator or some combination of the two.","from":"reporter","subject":"CQL3 column value validation bug"},{"body":"What happens is that columns metadata for composite CQL3 columns refers to only one of the component of the column name. So internally they have a 'componentIndex' parameter which allow to handle them correctly (otherwise you don't know which comparator to use to display those column metadata). However, we decided that thrift users shouldn't have to care about this, so we don't expose those metadata to thrift at all. In other words, the CF does have metadata for 'sum' and 'uniques' internally, but they are not exposed to thrift.\n\nSo I guess it is more of a cqlsh problem that shouldn't use the thrift describe call for CQL3. Instead, it should directly query the system.keyspace, system.columnfamilies and system.columns table. However, it'd be much more easier to do that with CASSANDRA-4018, but that is not in 1.1.","from":"developer"},{"body":"bq. So I guess it is more of a cqlsh problem that shouldn't use the thrift describe call for CQL3.\n\nThe main problem here from my perspective is that it was impossible to insert data into the column family except with cqlsh. Using a basic thrift batch mutate failed on the validation step because it tried to validate the value in the sum column as text (default_validation_class).","from":"developer"},{"body":"True, I'm unable to save any data to Cassandra.\nTrying to save with Hector (uses Mutators):\n#","from":"developer"},{"body":"bq. it was impossible to insert data into the column family except with cqlsh. Using a basic thrift batch mutate failed on the validation step because it tried to validate the value\n\nAlright, that part is not expected and is likely a bug.\n\nThough I note that even if we fix that, since we don't expose on the thrift side everything needed to interpret the value correctly, thrift client won't be able to work with those kind of table very well. In particular they won't know how to interpret the value of a get. So I guess we need to either say clearly that CQL3 table are not meant to be accessed from thrift, or we should probably start exposing all the column metatada on the thrift side (but we need to expose the componentIndex part of the metadata in particular and the discussion on CASSANDRA-4093 is relevant for that).","from":"developer"},{"body":"Attaching simple patch to allow the insertion in the case above. Truth is, I'm not really satisfied (though I don't have a clearly better option) by such a patch for 2 reasons:\n# it will break the case for thrift where people were using compositeType and column_metadata on them. That might be 0 people we're talking about but it's still a bit annoying. An alternative here would be that for each composite column, we iterater over all column_metadata and check where one apply. This would work but this feels butt ugly.\n# it feels to me we supporting either not enough or too much in the thrift side. Even if we support this, we still don't expose the CQL3 metadata to thrift, so one has to know the CQL schema definition to be able to create the column in the first place (or deserialize it on read). In particular, most advanced thrift client will likely still break at one point or another.\n\nOverall I see two reasonable approaches:\n* Either we start exposing enough on the thirft side so that thrift client can work correctly with CQL3. While this may be doable reasonably easily now, this will be increasingly difficult with things like CASSANDRA-4179, CASSANDRA-3647, ... Besides, even if we make it possible to work with CQL3 table, it doesn't mean it will be convenient since thrift won't do the grouping of columns in sparse table.\n* Be clear that you cannot work with CQL3 created table from thrift.\n\nBut imo the in-the-middle approach that this patch would start takes the risk of polluting the thrift side without adding much.\n","from":"developer"},{"body":"To be clear, if this issue is for instance a blocker for map-reduce for CQL3, I'm not against committing it, but on a more long term/general level, I do want to note that accessing CQL3 tables to thrift is imho a larger problem and we should be clear on what we want to guarantee.","from":"developer"},{"body":"It doesn't sound like its going to be possible to make things work well in the case of accessing cql3 data from thrift.\n\nIf thats the case I'm in favor of making that incompatibility *very* explicit. I would say explicitly throwing an exception when trying to access a cql3 column family from thrift would be an acceptable solution.\n\nThe fact that it took me quite a bit of time as well as digging around in both client code and cassandra source code to figure this bug out makes me worried for other users hitting similar problems.","from":"developer"},{"body":"bq. I would say explicitly throwing an exception when trying to access a cql3 column family from thrift would be an acceptable solution.\n\n-1 on that; it's very useful to have a lower-level access method to look at the \"storage engine\" layer instead of the logical CQL3 rows at times.","from":"developer"},{"body":"so disable write access to cql3 cfs from thrift? And leave read access with the qualification that everything will be returned as BytesType?\n\nThe only other option I see is to say that we are going to identify and fix all bugs like this one.","from":"developer"},{"body":"bq. disable write access to cql3 cfs from thrift? And leave read access with the qualification that everything will be returned as BytesType?\n\nSounds reasonable to me.","from":"developer"},{"body":"bq. disable write access to cql3 cfs from thrift? And leave read access with the qualification that everything will be returned as BytesType?\n\nA variation in the same spirit could be to disable access to CQL3 table by default but add a thrift debug mode (either per-connection through or globally through JMX if we don't want to add a new thrift method) that would make thrift disable every validation. That way, by default the message that CQL3 tables should be access through CQL3 would be clear but we would have low-level read/write for debugging. ","from":"developer"},{"body":"bq. disable access to CQL3 table by default but add a thrift debug mode\n\n+lots.","from":"developer"},{"body":"sgtm","from":"developer"},{"body":"+1\n\n(and I'd prefer JMX otherwise I foresee clients using this as an escape hatch and fubarring things up.)","from":"developer"},{"body":"So, from an outside conversation, it sounds like there may be some confusion on exactly what is being discussed here. Are we talking about disabling Thrift access entirely to columnfamilies which use named metadata with composites, or are we just talking about not supporting Thrift addressing columns inside composites by their CQL3 names?\n\nThe second seems eminently reasonable. The former sounds crazy.","from":"developer"},{"body":"bq. are we just talking about not supporting Thrift addressing columns inside composites by their CQL3 names\n\nThat is not what I was talking about, but I don't even understand how we could ever support that.\n\nbq. disabling Thrift access entirely to columnfamilies which use named metadata with composites\n\nThat is what I'm talking about (though to be precise, it would be for composites using named metadata *created through CQL3*). Anyway, I don't know if that's so crazy but in any case calling it crazy doesn't help solving the problem.\n\n","from":"developer"},{"body":"So here is where I ended up on this.\n\nAssuming we aren't extremely interested in fixing bugs like these when the come up then I'm all for disabling thrift access by default to cql3 cfs. From what I can tell there isn't an immediately easy way to know a cf was created through cql3 at the moment but we could add that I guess.\n\nOn the other hand if we want to just go ahead and fix these bugs when they happen, I don't see much reason to go out of our way to disable thrift access to cql3 cfs.","from":"developer"},{"body":"bq. one has to know the CQL schema definition to be able to create the column in the first place (or deserialize it on read)\n\nCan you remind me why we don't translate {{PRIMARY KEY(gid, period, tid)}} into a comparator of {{CompositeType(Int32Type, BytesType, UTF8Type)}}? (period -> int, tid -> bytes, sparse columns -> utf8)","from":"developer"},{"body":"bq. Can you remind me why we don't translate PRIMARY KEY(gid, period, tid) into a comparator of CompositeType(Int32Type, BytesType, UTF8Type)\n\nWe do. The problem is actually with columns that are not part of the key. Let me try to sum that up:\n* Pre-CQL3, the ColumnDefinition name was always a full column name.\n* In CQL3, the ColumnDefinition name only correspond to one of the component of the column name, i.e. to the UTF8Type component above.\n* To be able to distinguish both case internally, we've introduce the componentIndex field in ColumnDefinition. However, we decided that we didn't wanted to expose this field to thrift, and so we don't expose to thrift the ColumnDefinition from CQL3 table.\n\nThe net result is that as far as thrift is concerned, the CQL3 tables have no columns_metadata whatsoever. It follows that thrift clients don't know what are the correct value for the last UTF8Type component, and don't know what is the type of the corresponding value (and thus cannot serialize/deserialize said value correctly). ","from":"developer"},{"body":"So basically, the problem is we're trying to maintain compatibility with (non-cql3) wide-row named columns? If we're willing to break that scenario, does the problem go away?\n\nI still think that naming wide row columns is nonsensical, but we could add an extra layer of protection by warning on startup in 1.2 that you need to update your schema:\n\n- if you have named columns and non-utf8-or-bytes comparator\n- if you have named columns but metadata shows more columns than names","from":"developer"},{"body":"I'm not sure I understand what's a named columns above to be honest.\n\nThere is basically two informations from CFMetadata you need to know to insert a column correctly in a table (CQL3 or no CQL3): the comparator and *all* of column_metadata. The comparator is necessary to know what is a valid column name and the column_metadata is necessary to know what is a valid column value (I'm simplifying a bit, I'm assuming that the key_validation and default_validator are BytesType but that doesn't matter for the problem at hand).\n\nNow the problem is that for any table created through CQL3 that doesn't use COMPACT STORAGE (let's call those CQL3 tables), all the ColumnDefinition of column_metada will have a componentIndex. So none of those ColumnDefinition are exposed in thrift. In practice it means that if I do:\n{noformat}\nCREATE TABLE user {\n user_id blob PRIMARY KEY,\n name text,\n age int\n}\n{noformat}\nthen if a thrift client do a describe, it will basically get:\n{noformat}\ncomparator = CompositeType(UTF8Type) // it's a composite so that we can add collection later on\ncolumn_metadata = []\n{noformat}\n\nAt that point we have two slightly separate problems:\n# Even if a user produces a valid column, with say a composite name being \"age\" and a value being an int, then currently the code throw an exception. Fixing that exception is the goal of the attached patch (though it would have to be updated to work with collections in 1.2). I'm fine fixing that, though I'm pointing that there is a second, more general problem.\n# Since the thrift client doesn't know about the actual column_metadata, how can we expect it to correctly insert data. In particular I'm pretty sure higher level clients like pycassa or astyanax will serialize data incorrectly if they don't know the right value validator. Besides, there is many way to be confused if you use a CQL3 table from thrift. For instance if you create the wrong column (i'ts enough to mess up the case), you'll be surprised to not be able to access it when you go back to CQL3. So be clear, I do am suggesting that we don't allow accessing table created from CQL3 *without* COMPACT STORAGE from thrift, because I think it will be more sane, even if it does mean that you're not coming back from CQL3 once you've start really using it.\n","from":"developer"},{"body":"Right, so what I was saying was, if we're willing to say that columndefinition name is always the cql3 column name, then we don't need componentIndex, at the price of potentially breaking a corner case in non-cql3 schemas.","from":"developer"},{"body":"But what the componentIndex give us is which of the composite component is the cql3 column name. Typically, with collections, it's not even necessarily the last of the component. Now, if you know the table is a CQL3 one, you could try to infer which component it is by saying that it's the last component, except if the last is a collection type, in which case that's the previous one, but that feels a bit messy. And besides, thrift client libraries don't have a simple automatic way to know if the table is a CQL3 one in the first place. I guess you could say that if you have a composite comparator *and* some column_metadata then you are likely a CQL3 table, but again, not very clean imo.\n\n\nAt least internally I would be in favor of keeping the componentIndex as it is cleaner. I guess we could start returning the ColumnDefinition from CQL3 table without the componentIndex and let thrift client infer what they can. As said, I still think using CQL3 table from thrift has other way to be confusing, but why not.","from":"developer"},{"body":"Here's what I think our goals are, in order of importance:\n\n# Existing \"high level\" clients should have a reasonable upgrade path to read and update collections and cql3 CFs.\n# CLI and other tools that don't speak cql should have a way to tell that they can't cope with CQL3 CF definitions. This is the problem Nick described originally in this ticket description. Actually making such tools able to manipulate CQL3 definitions is NOT a goal, but we should, as Nick says, make that more obvious.\n# We should allow updates to CQL3 CFs from Thrift, if someone manually composes the correct CompositeType bytes. This is what most of the rest of the discussion here involves.\n\nAnalysis:\n\n# This we have done--they will have to use cql-over-thrift, but IMO this is reasonable. Thrift RPC methods to deal with collections have never been on the roadmap.\n# This is tough since if this involves new information (like exposing component_index, or even adding a cql3 boolean to CfDef) then old tools by definition won't know about it. *Proposal:* what if we omit CQL3 CfDefs from those we return to describe_keyspace[s] calls? Not quite as good as returning a CfDef that explicitly says \"I am here but you can't touch this,\" but compare to displaying incomplete information (that the cli doesn't know is incomplete) it's much more obvious that the cli and other old thrift-base schema manipulators can't cope with such.\n# Sylvain mentioned adding a thrift or JMX method to enable validation-free updates but I'd rather make ThriftValidation cql3-aware, which would let this work without any special flags. I don't see any downside to this except added complexity for ThriftValidation.","from":"developer"},{"body":"I do think we should push this to 1.2 however, since I'm leery of changing the behavior or describe_keyspace[s] when we are fairly deep into 1.1's stable lifetime.","from":"developer"},{"body":"CQL3 defined CFs won't be exposed to Thrift API anymore (with warning), I also checked CassandraServer and ThriftValidation and figured that column name validation already in place via ThriftValidation.validateColumnNames(...).","from":"developer"},{"body":"Just glanced at the patch, but it looks like this just makes cql3 cfs not show up in describe_keyspaces calls right? Reading/writing to the cf would still go through and error?\n\nIf I'm looking at that right then it seems like we would want the opposite behavior. At least from my perspective, I would like to have an indication that the cf exists in my client even if i can't write to it.","from":"developer"},{"body":"The idea was to do two things:\n\n- update ThriftValidation to be cql3-aware, so if a Thrift user manually composes a valid column, we accept it\n- but, we don't expose cql3 CFs via describe_keyspaces since clients do not have enough information to generate valid requests automatically","from":"developer"},{"body":"I agree that the two points above (make ThriftValidation cql3-aware but don't expose cql3 CFs defs) make sense as a strategy.","from":"developer"},{"body":"On Pavel's patch:\n* I'm not a fan of logging a warning (or to log anything really) if someone has CQL3 CFs. We're trying to push CQL3 as a good thing, let's not log anything that could be interpreted as if something was wrong/abnormal.\n* The check is excluding composite CF without any ColumnDefinition while it shouldn't.\n\nAttaching a v2 that:\n* Fix the two remarks above.\n* Rename the check as isThriftIncompatible. I think it's more about detecting CF definitions that cannot be exploided fully by thrift rather than discriminate between what is a thrift CF and CQL3 CF. Especially since the intersection between those two notions is not empty.\n* Ship the changes to ThriftValidation from my first patch, though modified a bit to be more generic and handle correctly collections. I'll note that this part makes validation potentially iterate over all ColumnDefinition for composite CF, but that's not really an issue since composite CF created on the thrift side are almost guaranteed to have no ColumnDefinition.\n","from":"developer"},{"body":"bq. I'm not a fan of logging a warning (or to log anything really) if someone has CQL3 CFs. We're trying to push CQL3 as a good thing, let's not log anything that could be interpreted as if something was wrong/abnormal.\n\nWouldn't that create confusion when some CFs are visible through Thrift and some are not or do we rely on that users should know what is so special about CQL3 CFs?","from":"developer"},{"body":"bq. Wouldn't that create confusion when some CFs are visible through Thrift and some are not\n\nI'm not saying this shouldn't be documented at all, but merely that it's a documentation issue and as such logging it at each startup is not the right place. ","from":"developer"},{"body":"If everyone else is ok with that, lgtm.","from":"developer"},{"body":"So, this is good to commit?","from":"developer"},{"body":"Alright, committed, thanks.","from":"developer"}],"created":"2012-06-26T16:44:06.000+0000","description":"{noformat}\ncqlsh> create keyspace test with strategy_class = 'SimpleStrategy' and strategy_options:replication_factor = 1;\ncqlsh> use test;\ncqlsh:test> CREATE TABLE stats (\n ... gid blob,\n ... period int,\n ... tid blob, \n ... sum int,\n ... uniques blob,\n ... PRIMARY KEY(gid, period, tid)\n ... );\ncqlsh:test> describe columnfamily stats;\n\nCREATE TABLE stats (\n gid blob PRIMARY KEY\n) WITH\n comment='' AND\n comparator='CompositeType(org.apache.cassandra.db.marshal.Int32Type,org.apache.cassandra.db.marshal.BytesType,org.apache.cassandra.db.marshal.UTF8Type)' AND\n read_repair_chance=0.100000 AND\n gc_grace_seconds=864000 AND\n default_validation=text AND\n min_compaction_threshold=4 AND\n max_compaction_threshold=32 AND\n replicate_on_write='true' AND\n compaction_strategy_class='SizeTieredCompactionStrategy' AND\n compression_parameters:sstable_compression='SnappyCompressor';\n{noformat}\n\nYou can see in the above output that the stats cf is created with the column validator set to text, but neither of the non primary key columns defined are text. It should either be setting metadata for those columns or not setting a default validator or some combination of the two.","issue_id":"12595934","key":"CASSANDRA-4377","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2012-09-11T09:03:56.000+0000","role":"fixed_distractor","summary":"CQL3 column value validation bug"} {"case_id":"12599347","cluster":"DISTRACTOR-CASSANDRA-4446","comments":[{"body":"In general, nodetool drain never seems to completely eliminate on-startup log replay. I observe this all the time on all clusters. It certainly cuts down the amount of replay done, but either never or fairly seldom eliminates it completely - at least not based on log messages indicating replay.\n\nNever had time to investigate.","created":"2012-09-08T04:02:32.275+0000"},{"body":"Also seeing this in an upgrade from 1.0.xx to 1.1.15:\n\n INFO 16:29:17,486 completed pre-loading (3 keys) key cache.\n INFO 16:29:17,495 Replaying /data2/commit-cassandra/CommitLog-1349727956484.log\n INFO 16:29:17,503 Replaying /data2/commit-cassandra/CommitLog-1349727956484.log\n INFO 16:29:18,495 GC for ParNew: 3506 ms for 4 collections, 1963062320 used; max is 17095983104\n INFO 16:29:18,498 Finished reading /data2/commit-cassandra/CommitLog-1349727956484.log\n INFO 16:29:18,499 Log replay complete, 0 replayed mutations\n\n\nThis is a standard upgrade process which includes a drain","created":"2012-10-08T23:40:42.027+0000"},{"body":"I also experience this every time I drain / restart (up until latest 1.1.6 but not on 1.1.6 itself any more) and getting this message in log:\n\n{quote}\n2012-10-12_15:50:36.92191 INFO 15:50:36,921 Log replay complete, N replayed mutations \n{quote}\n\nwith N being non-zero. I wonder if this is a cause of double-counts for Counter mutations.","created":"2012-10-15T15:41:13.759+0000"},{"body":"I had the same experience, when I upgraded my cluster from 1.0.9 to 1.0.11. I ran drain before the upgrade, upgrade on the node finished and node restarted at 2012-11-20 10:20:58, but then I see in the logs reply of commit log:\n{quote} \n INFO [main] 2012-11-20 09:41:13,918 CommitLog.java (line 172) Replaying /raid0/cassandra/commitlog/CommitLog-1353402218337.log\n INFO [main] 2012-11-20 09:41:20,360 CommitLog.java (line 179) Log replay complete, 0 replayed mutations\n INFO [main] 2012-11-20 10:11:35,635 CommitLog.java (line 167) No commitlog files found; skipping replay\n INFO [main] 2012-11-20 10:21:11,631 CommitLog.java (line 172) Replaying /raid0/cassandra/commitlog/CommitLog-1353404473899.log\n INFO [main] 2012-11-20 10:21:18,119 CommitLog.java (line 179) Log replay complete, 6413 replayed mutations\n INFO [main] 2012-11-20 10:55:46,435 CommitLog.java (line 172) Replaying /raid0/cassandra/commitlog/CommitLog-1353406871619.log\n INFO [main] 2012-11-20 10:55:54,139 CommitLog.java (line 179) Log replay complete, 3 replayed mutations\n{quote} \nThis caused over increment of counters\n","created":"2012-11-21T06:36:37.743+0000"},{"body":"This is going to stand as a known limitation with 1.0.x; so far it looks like it is fixed in latest 1.1.","created":"2012-11-21T12:00:06.795+0000"},{"body":"did a nodetool drain before 1.1.7 -> 1.2.0. upon starting 1.2.0 every node in my cluster still replayed the commit logs and created mutations.\n\n{code}\n INFO 13:17:06,529 DRAINING: starting drain process\n INFO 13:17:06,529 Stop listening to thrift clients\n INFO 13:17:06,532 Announcing shutdown\n INFO 13:17:07,536 Waiting for messaging service to quiesce\n INFO 13:17:07,537 MessagingService shutting down server thread.\n\n... normal startup stuff..\n INFO 13:20:20,182 Replaying /ssd/commitlog/CommitLog-1355265349912.log\n INFO 13:20:24,166 Finished reading /ssd/commitlog/CommitLog-1355265349912.log\n INFO 13:20:24,166 Replaying /ssd/commitlog/CommitLog-1355265349914.log\n INFO 13:20:26,700 Finished reading /ssd/commitlog/CommitLog-1355265349914.log\n INFO 13:20:26,701 Replaying /ssd/commitlog/CommitLog-1355265349915.log\n INFO 13:20:28,118 Finished reading /ssd/commitlog/CommitLog-1355265349915.log\n... more replay lines ...\n INFO 13:22:00,061 Log replay complete, 8052 replayed mutations\n INFO 13:22:00,358 Possible old-format hints found. Truncating\n INFO 13:22:00,370 Enqueuing flush of Memtable-local@1908923620(402/402 serialized/live bytes, 13 ops)\n INFO 13:22:00,372 Writing Memtable-local@1908923620(402/402 serialized/live bytes, 13 ops)\n INFO 13:22:00,494 Cassandra version: 1.2.0\n INFO 13:22:00,495 Thrift API version: 19.35.0\n INFO 13:22:00,495 CQL supported versions: 2.0.0,3.0.0 (default: 3.0.0)\n INFO 13:22:00,534 Loading persisted ring state\n INFO 13:22:00,537 Starting up server gossip\n WARN 13:22:00,557 No host ID found, created dd3a40e2-fef1-4574-87b8-e2929fd80235 (Note: This should happen exactly once per node).\n{code}","created":"2013-01-04T00:16:26.808+0000"},{"body":"Let's reopen then if it doesn't sound like it's fixed in recent releases.","created":"2013-01-04T10:11:04.707+0000"},{"body":"+1. Good to see this ticket reopen. Drain didn't work for a while. I remove all the commit logs files before a restart to avoid counters operations to replay.","created":"2013-01-04T10:59:36.391+0000"},{"body":"It always replays 3 mutations when I follow these steps:\n{code}\nbin/cassandra\ntools/bin/cassandra-stress --operation=INSERT --num-keys=100000\nbin/nodetool drain\nbin/cassandra\n{code}\n\nThis is on trunk, commit acf30622","created":"2013-01-04T20:11:21.690+0000"},{"body":"System tables were not getting flushed. This is the source of the extra replaying. Patch attached to fix this, and also parallelize flushing.","created":"2013-01-11T23:46:43.621+0000"},{"body":"+1","created":"2013-01-14T20:28:02.213+0000"},{"body":"Committed.\n\nNote that earlier releases can workaround by manually running flush against system KS before drain.","created":"2013-01-14T22:07:38.040+0000"},{"body":"While I'm sure that this does fix one real cause of drain not working in trunk (yay!), one of the symptoms I've heard reported in the 1.0.x - 1.1.5 timeframe is that \"my counters over-counted on upgrade, despite drain\". Most recent report was 1.0.12->1.1.8 with drain being run as part of the upgrade process.\n\nNEWS.txt says :\n\n\"If you using counters and upgrading from a version prior to 1.1.6, you should drain existing Cassandra nodes prior to the upgrade to prevent overcount during commitlog replay (see CASSANDRA-4782). For non-counter uses, drain is not required but is a good practice to minimize restart time.\"\n\nIf drain in these versions can't be counted on (heh) to actually work for this purpose (which reports suggest it cannot), then I propose changing this line to read \"drain existing nodes and remove their commitlog\".","created":"2013-01-18T00:23:07.431+0000"},{"body":"No counter mutations are double-counted, only the unflushed system changes are replayed.","created":"2013-01-18T00:34:30.725+0000"},{"body":"If only unflushed system changes are replayed, how do you account for :\n\n\"I upgraded from 1.0.12 to 1.1.8, using drain, and I noticed overcounting counters\" ?\n\nIt's quite possible that upgrading from 1.1.x to 1.1.y>x does not in fact replay anything other than system keyspace and does not incur double counting of counters. I am however pretty confident based on multiple reports of the above quoted issue that counter increments may be over-replayed if one uses drain (as NEWS.txt suggests) while upgrading from 1.0.x to >1.1.6.\n\nIf this is being dealt with as \"known limitation of 1.0.x\", then I continue to suggest the above change to NEWS.txt, as otherwise people using counters in 1.0.x WILL incur double-increment while upgrading per the instructions in NEWS.txt.","created":"2013-01-18T01:04:28.313+0000"},{"body":"Show me how to reproduce it and I will re-evaluate my position, but as near as I can tell the advice in NEWS is still best practice.","created":"2013-01-18T01:11:07.046+0000"},{"body":"How to reproduce it, from the multiple reports :\n\n1) Drain and stop cluster with counters on 1.0.x\n2) Start same cluster on 1.1.x\n3) Notice commitlog replay of the counter columnfamily and that your counters have over-counted\n\nAttached is a log from the latest reporter, CASSANDRA-4446--1.0.12_to_1.1.8.txt. It shows the following.\n\n1) Drain starts and completes on 1.0.12\n2) Cluster then starts on 1.1.8, and replays the commit log\n3) As part of commitlog replay, it flushes various CFs including titan3/RMEntityCount/, which is a counter columnfamily; machine has 4gb of heap and the flush is while thrift is down and the node has not jumped state to normal, so it seems reasonable to conjecture this flush is part of commitlog replay\n4) It then logs \"10698 replayed mutations\", which adds further support to the idea that these Counts are part of replay\n5) Operator then noticed a significant percentage of records had overcounted in this columnfamily","created":"2013-01-18T23:45:47.992+0000"}],"conversations":[{"body":"I recently wiped a customer's QA cluster. I drained each node and verified that they were drained. When I restarted the nodes, I saw the commitlog replay create a memtable and then flush it. I have attached a sanitized log snippet from a representative node at the time. \n\nIt appears to show the following :\n1) Drain begins\n2) Drain triggers flush\n3) Flush triggers compaction\n4) StorageService logs DRAINED message\n5) compaction thread excepts\n6) on restart, same CF creates a memtable\n7) and then flushes it [1]\n\nThe columnfamily involved in the replay in 7) is the CF for which the compaction thread excepted in 5). This seems to suggest a timing issue whereby the exception in 5) prevents the flush in 3) from marking all the segments flushed, causing them to replay after restart.\n\nIn case it might be relevant, I did an online change of compaction strategy from Leveled to SizeTiered during the uptime period preceding this drain.\n\n[1] Isn't commitlog replay not supposed to automatically trigger a flush in modern cassandra?","from":"reporter","subject":"nodetool drain sometimes doesn't mark commitlog fully flushed"},{"body":"In general, nodetool drain never seems to completely eliminate on-startup log replay. I observe this all the time on all clusters. It certainly cuts down the amount of replay done, but either never or fairly seldom eliminates it completely - at least not based on log messages indicating replay.\n\nNever had time to investigate.","from":"developer"},{"body":"Also seeing this in an upgrade from 1.0.xx to 1.1.15:\n\n INFO 16:29:17,486 completed pre-loading (3 keys) key cache.\n INFO 16:29:17,495 Replaying /data2/commit-cassandra/CommitLog-1349727956484.log\n INFO 16:29:17,503 Replaying /data2/commit-cassandra/CommitLog-1349727956484.log\n INFO 16:29:18,495 GC for ParNew: 3506 ms for 4 collections, 1963062320 used; max is 17095983104\n INFO 16:29:18,498 Finished reading /data2/commit-cassandra/CommitLog-1349727956484.log\n INFO 16:29:18,499 Log replay complete, 0 replayed mutations\n\n\nThis is a standard upgrade process which includes a drain","from":"developer"},{"body":"I also experience this every time I drain / restart (up until latest 1.1.6 but not on 1.1.6 itself any more) and getting this message in log:\n\n{quote}\n2012-10-12_15:50:36.92191 INFO 15:50:36,921 Log replay complete, N replayed mutations \n{quote}\n\nwith N being non-zero. I wonder if this is a cause of double-counts for Counter mutations.","from":"developer"},{"body":"I had the same experience, when I upgraded my cluster from 1.0.9 to 1.0.11. I ran drain before the upgrade, upgrade on the node finished and node restarted at 2012-11-20 10:20:58, but then I see in the logs reply of commit log:\n{quote} \n INFO [main] 2012-11-20 09:41:13,918 CommitLog.java (line 172) Replaying /raid0/cassandra/commitlog/CommitLog-1353402218337.log\n INFO [main] 2012-11-20 09:41:20,360 CommitLog.java (line 179) Log replay complete, 0 replayed mutations\n INFO [main] 2012-11-20 10:11:35,635 CommitLog.java (line 167) No commitlog files found; skipping replay\n INFO [main] 2012-11-20 10:21:11,631 CommitLog.java (line 172) Replaying /raid0/cassandra/commitlog/CommitLog-1353404473899.log\n INFO [main] 2012-11-20 10:21:18,119 CommitLog.java (line 179) Log replay complete, 6413 replayed mutations\n INFO [main] 2012-11-20 10:55:46,435 CommitLog.java (line 172) Replaying /raid0/cassandra/commitlog/CommitLog-1353406871619.log\n INFO [main] 2012-11-20 10:55:54,139 CommitLog.java (line 179) Log replay complete, 3 replayed mutations\n{quote} \nThis caused over increment of counters\n","from":"developer"},{"body":"This is going to stand as a known limitation with 1.0.x; so far it looks like it is fixed in latest 1.1.","from":"developer"},{"body":"did a nodetool drain before 1.1.7 -> 1.2.0. upon starting 1.2.0 every node in my cluster still replayed the commit logs and created mutations.\n\n{code}\n INFO 13:17:06,529 DRAINING: starting drain process\n INFO 13:17:06,529 Stop listening to thrift clients\n INFO 13:17:06,532 Announcing shutdown\n INFO 13:17:07,536 Waiting for messaging service to quiesce\n INFO 13:17:07,537 MessagingService shutting down server thread.\n\n... normal startup stuff..\n INFO 13:20:20,182 Replaying /ssd/commitlog/CommitLog-1355265349912.log\n INFO 13:20:24,166 Finished reading /ssd/commitlog/CommitLog-1355265349912.log\n INFO 13:20:24,166 Replaying /ssd/commitlog/CommitLog-1355265349914.log\n INFO 13:20:26,700 Finished reading /ssd/commitlog/CommitLog-1355265349914.log\n INFO 13:20:26,701 Replaying /ssd/commitlog/CommitLog-1355265349915.log\n INFO 13:20:28,118 Finished reading /ssd/commitlog/CommitLog-1355265349915.log\n... more replay lines ...\n INFO 13:22:00,061 Log replay complete, 8052 replayed mutations\n INFO 13:22:00,358 Possible old-format hints found. Truncating\n INFO 13:22:00,370 Enqueuing flush of Memtable-local@1908923620(402/402 serialized/live bytes, 13 ops)\n INFO 13:22:00,372 Writing Memtable-local@1908923620(402/402 serialized/live bytes, 13 ops)\n INFO 13:22:00,494 Cassandra version: 1.2.0\n INFO 13:22:00,495 Thrift API version: 19.35.0\n INFO 13:22:00,495 CQL supported versions: 2.0.0,3.0.0 (default: 3.0.0)\n INFO 13:22:00,534 Loading persisted ring state\n INFO 13:22:00,537 Starting up server gossip\n WARN 13:22:00,557 No host ID found, created dd3a40e2-fef1-4574-87b8-e2929fd80235 (Note: This should happen exactly once per node).\n{code}","from":"developer"},{"body":"Let's reopen then if it doesn't sound like it's fixed in recent releases.","from":"developer"},{"body":"+1. Good to see this ticket reopen. Drain didn't work for a while. I remove all the commit logs files before a restart to avoid counters operations to replay.","from":"developer"},{"body":"It always replays 3 mutations when I follow these steps:\n{code}\nbin/cassandra\ntools/bin/cassandra-stress --operation=INSERT --num-keys=100000\nbin/nodetool drain\nbin/cassandra\n{code}\n\nThis is on trunk, commit acf30622","from":"developer"},{"body":"System tables were not getting flushed. This is the source of the extra replaying. Patch attached to fix this, and also parallelize flushing.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed.\n\nNote that earlier releases can workaround by manually running flush against system KS before drain.","from":"developer"},{"body":"While I'm sure that this does fix one real cause of drain not working in trunk (yay!), one of the symptoms I've heard reported in the 1.0.x - 1.1.5 timeframe is that \"my counters over-counted on upgrade, despite drain\". Most recent report was 1.0.12->1.1.8 with drain being run as part of the upgrade process.\n\nNEWS.txt says :\n\n\"If you using counters and upgrading from a version prior to 1.1.6, you should drain existing Cassandra nodes prior to the upgrade to prevent overcount during commitlog replay (see CASSANDRA-4782). For non-counter uses, drain is not required but is a good practice to minimize restart time.\"\n\nIf drain in these versions can't be counted on (heh) to actually work for this purpose (which reports suggest it cannot), then I propose changing this line to read \"drain existing nodes and remove their commitlog\".","from":"developer"},{"body":"No counter mutations are double-counted, only the unflushed system changes are replayed.","from":"developer"},{"body":"If only unflushed system changes are replayed, how do you account for :\n\n\"I upgraded from 1.0.12 to 1.1.8, using drain, and I noticed overcounting counters\" ?\n\nIt's quite possible that upgrading from 1.1.x to 1.1.y>x does not in fact replay anything other than system keyspace and does not incur double counting of counters. I am however pretty confident based on multiple reports of the above quoted issue that counter increments may be over-replayed if one uses drain (as NEWS.txt suggests) while upgrading from 1.0.x to >1.1.6.\n\nIf this is being dealt with as \"known limitation of 1.0.x\", then I continue to suggest the above change to NEWS.txt, as otherwise people using counters in 1.0.x WILL incur double-increment while upgrading per the instructions in NEWS.txt.","from":"developer"},{"body":"Show me how to reproduce it and I will re-evaluate my position, but as near as I can tell the advice in NEWS is still best practice.","from":"developer"},{"body":"How to reproduce it, from the multiple reports :\n\n1) Drain and stop cluster with counters on 1.0.x\n2) Start same cluster on 1.1.x\n3) Notice commitlog replay of the counter columnfamily and that your counters have over-counted\n\nAttached is a log from the latest reporter, CASSANDRA-4446--1.0.12_to_1.1.8.txt. It shows the following.\n\n1) Drain starts and completes on 1.0.12\n2) Cluster then starts on 1.1.8, and replays the commit log\n3) As part of commitlog replay, it flushes various CFs including titan3/RMEntityCount/, which is a counter columnfamily; machine has 4gb of heap and the flush is while thrift is down and the node has not jumped state to normal, so it seems reasonable to conjecture this flush is part of commitlog replay\n4) It then logs \"10698 replayed mutations\", which adds further support to the idea that these Counts are part of replay\n5) Operator then noticed a significant percentage of records had overcounted in this columnfamily","from":"developer"}],"created":"2012-07-18T21:30:43.000+0000","description":"I recently wiped a customer's QA cluster. I drained each node and verified that they were drained. When I restarted the nodes, I saw the commitlog replay create a memtable and then flush it. I have attached a sanitized log snippet from a representative node at the time. \n\nIt appears to show the following :\n1) Drain begins\n2) Drain triggers flush\n3) Flush triggers compaction\n4) StorageService logs DRAINED message\n5) compaction thread excepts\n6) on restart, same CF creates a memtable\n7) and then flushes it [1]\n\nThe columnfamily involved in the replay in 7) is the CF for which the compaction thread excepted in 5). This seems to suggest a timing issue whereby the exception in 5) prevents the flush in 3) from marking all the segments flushed, causing them to replay after restart.\n\nIn case it might be relevant, I did an online change of compaction strategy from Leveled to SizeTiered during the uptime period preceding this drain.\n\n[1] Isn't commitlog replay not supposed to automatically trigger a flush in modern cassandra?","issue_id":"12599347","key":"CASSANDRA-4446","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-01-15T15:14:04.000+0000","role":"fixed_distractor","summary":"nodetool drain sometimes doesn't mark commitlog fully flushed"} {"case_id":"12604193","cluster":"DISTRACTOR-CASSANDRA-4561","comments":[{"body":"in logfile I see only this:\n INFO [MigrationStage:1] 2012-08-21 11:27:55,560 ColumnFamilyStore.java (line 659) Enqueuing flush of Memtable-schema_columnfamilies@970905946(1266/1582 serialized/live bytes, 20 ops)\n INFO [FlushWriter:5] 2012-08-21 11:27:55,561 Memtable.java (line 264) Writing Memtable-schema_columnfamilies@970905946(1266/1582 serialized/live bytes, 20 ops)\n INFO [FlushWriter:5] 2012-08-21 11:27:55,587 Memtable.java (line 305) Completed flushing /var/lib/cassandra/data/system/schema_columnfamilies/system-schema_columnfamilies-he-196-Data.db (1336 bytes) for commitlog position ReplayPosition(segmentId=4914817711083622, position=333055)","created":"2012-08-21T09:28:38.950+0000"},{"body":"oh, and it obviously happens when I'm trying to change anything else, not only compaction treshhold, ex: sstable_size_in_mb","created":"2012-08-21T12:58:14.996+0000"},{"body":"I saw a similar case. In my case, columns of schema_keyspaces and schema_columnfamilies in the system keyspace had future timestamp such as 2052-08-01 11:22:33. All operations of changing schemas were not affected. I do not know why timestamps had future values. I fixed it with rebuilding system keyspace.","created":"2012-08-22T03:04:14.026+0000"},{"body":"Hmmm, it's possible. How did you rebuilt system keyspace?","created":"2012-08-22T07:12:26.694+0000"},{"body":"Zenek, what version are you using? It's possible if you are on version less than 1.1.3 or if you have created your schema before 1.1.3 because of CASSANDRA-4432...","created":"2012-08-22T09:16:16.164+0000"},{"body":"I'm using 1.1.4, but schema was created in 0.8.5\n\nCould you help me solve my problem step by step?","created":"2012-08-22T10:21:29.233+0000"},{"body":"Did you upgrade directly to 1.1.4 or one of the previous 1.1 versions first? Please check timestamps in your schema_* ColumnFamilies and see if they have any future dates. The easiest way I see to fix this would be to re-create your schema if you have timestamp problems from CASSANDRA-4432 ","created":"2012-08-22T10:27:50.483+0000"},{"body":"I've upgraded cassandra from 8.7 to 1.0, then to 1.0.5 -> 1.0.6 -> 1.1.0 -> 1.1.1 -> 1.1.2 -> 1.1.0 -> 1.1.3 -> 1.1.4.\n","created":"2012-08-22T12:40:36.067+0000"},{"body":"aha, so check system.schema_* CFs column's timestamp values I bet they are from the future. The simplest fix for you would be to delete system/schema_* SSTables and re-create a schema after restart.","created":"2012-08-22T15:16:15.146+0000"},{"body":"bq. check system.schema_* CFs column's timestamp values I bet they are from the future\n\nCan we fix that on startup for other upgraders?","created":"2012-08-22T15:20:11.280+0000"},{"body":"Yes, I was about too :)","created":"2012-08-22T15:23:44.368+0000"},{"body":"+1 I just ended up seeing this problem in our cluster which was upgraded from 1.1.2 to 1.1.3. Will probably have to find a workaround since I have to change schema now.","created":"2012-08-23T23:40:09.939+0000"},{"body":"I took my cluster down for 10 minutes. I took a snapshot of 'show schema' into a text file and removed system KS from it so I will only have my KS schema definition in it. Then I stop cassandra on all nodes. On each node, I removed system/schema_* folders from system's keyspace folder in cassandra data dir. I started all cassandra nodes. When I tried to reload the schema file using cli to recreate my CFs, I kept getting the message that CFs already exist. When I listed schema_columnfamilies in one of the node, I saws the same long timestamps like 2705487066780774 on columns of that CF. So, the procedure didn't quiet work out for me. What could have gone wrong here?\n\nPlease advice.","created":"2012-08-27T18:17:10.389+0000"},{"body":"What cassandra version are you running? Did you try to recreate a schema on all of the nodes or just one of them?","created":"2012-08-27T18:35:44.382+0000"},{"body":"I am running 1.1.3 which was upgraded from 1.1.2. When nodes came up, I created the schemas in one node only.","created":"2012-08-27T21:07:20.050+0000"},{"body":"As part of my process, I also removed saved_caches/system-*.","created":"2012-08-27T21:08:49.026+0000"},{"body":"if you still have such timestamps and messages that CF already exists means that schema wasn't empty on restart and was reloaded from one of the nodes. I'm currently working on the patch which would fix timestamp situation on node's start up.","created":"2012-08-27T21:14:11.095+0000"},{"body":"Do you have an ETA? I'll be happy to monkey apply it on my cluster and be the guinea pig for it.","created":"2012-08-27T21:44:43.586+0000"},{"body":"Not yet but I'm working on it and soon as it's ready I will attach it here.","created":"2012-08-27T23:40:58.029+0000"},{"body":"Don't you want to call cfs.truncate().get() to block for the truncate to finish? Rest LGTM.","created":"2012-09-04T19:38:53.861+0000"},{"body":"good point, I would add that and commit, thanks!","created":"2012-09-04T19:59:23.755+0000"},{"body":"Committed.","created":"2012-09-04T21:17:21.768+0000"},{"body":"After applying this patch (1.1.5-tentative) I see the future timestamps being detected and fixed:\n\n{code}\nINFO [main] 2012-09-07 02:24:00,090 DefsTable.java (line 202) Fixing timestamps of schema ColumnFamily schema_keyspaces...\nINFO [main] 2012-09-07 02:24:00,168 DefsTable.java (line 202) Fixing timestamps of schema ColumnFamily schema_columnfamilies...\n{code}\n\nIf I list schema_keyspaces I still see the old (future) timestamps.\n\nGiven the timestamps weren't actually updated I did go back and found the following error shortly after the restarted node joined the ring (this may or not be related):\n\n{code}ERROR [InternalResponseStage:1] 2012-09-07 02:24:16,555 AbstractCassandraDaemon.java (line 135) Exception in thread Thread[InternalResponseStage:1,5,main]\njava.lang.NullPointerException\n at org.apache.cassandra.db.DefsTable.mergeColumnFamilies(DefsTable.java:468)\n at org.apache.cassandra.db.DefsTable.mergeSchema(DefsTable.java:346)\n at org.apache.cassandra.db.DefsTable.mergeRemoteSchema(DefsTable.java:324)\n at org.apache.cassandra.service.MigrationManager$MigrationTask$1.response(MigrationManager.java:416)\n at org.apache.cassandra.net.ResponseVerbHandler.doVerb(ResponseVerbHandler.java:45)\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:59)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n at java.lang.Thread.run(Thread.java:679)\n{code}\n\nI have however assumed that the patch automatically fixes this across all servers in a multi dc ring once this patch is applied. If there is a specific series of steps required to fix this, such as the full ring outage, deletion of system/schema_* etc can it please be documented here, the wiki or in the NEWS.txt file under upgrading notes?","created":"2012-09-07T02:49:19.668+0000"},{"body":"Yes it does, but it assumes that you have schema in agreement across the nodes before it tries to fix timestamps. I can put that into NEWS.txt","created":"2012-09-07T09:01:06.112+0000"},{"body":"My schema is in agreement, but the timestamps aren't fixed, even though the log suggests otherwise (is that previously mentioned NPE related?). \n\n{code}\n[default@unknown] describe cluster;\nCluster Information:\n Snitch: org.apache.cassandra.locator.PropertyFileSnitch\n Partitioner: org.apache.cassandra.dht.RandomPartitioner\n Schema versions: \n 89b22434-5e34-381d-83d1-2a3cde1482fe: [x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x]\n\n[default@unknown]\n{code}\n\nSchema updates continue to silently fail.","created":"2012-09-07T09:09:23.855+0000"},{"body":"Interesting, why does it try to merge remote schema if all of your nodes are in agreement, can you run in debug mode as attach the log? The thing is timestamps are fixed long before the storage server is initialized and the process involves truncate of the system.schema_* CFs so I don't see any other reason why your schema still has old timestamps except some other node sent it migration request... \n\nNPE that you experience is also interesting, cfMetaData() is always not null as CFs are copied from one map to another on KS initialization, and KS itself could not be null or you would have the assertion error instead of NPE.","created":"2012-09-07T09:23:50.405+0000"},{"body":"Maybe we should reopen this if that didn't fully fixed the problem? I'd better hold on a bit on 1.1.5 if this is not fixed yet as this seem to hit quite a few people.","created":"2012-09-07T10:36:08.958+0000"},{"body":"We are not yet sure what is the problem so I think we should hold on reopening for a bit.","created":"2012-09-07T11:10:51.503+0000"},{"body":"I just thought that a good addition to existing patch could be change in ColumnSerializer.deserialize to fix timestamps from the future, that way migrations from the remote locations would be deserialized with correct timestamp even if they were sent with the wrong one...","created":"2012-09-07T11:19:32.068+0000"},{"body":"Well, Anton still sees timestamp in the future and is still unable to do schema updates, it does sound a lot like there is a problem and it's related to this ticket. I just meant that I prefer keeping that in mind before releasing 1.1.5 into the wild too quickly.","created":"2012-09-07T11:21:58.363+0000"},{"body":"I understand, but we don't know if that is caused by the fix not working or by something else as exception indicates that something was send to the node, this could be a separate issue. I will work on the ColumnSerializer change I mentioned and we'll see if it helps. ","created":"2012-09-07T11:27:39.209+0000"},{"body":"Debug log sent to [~xedin] privately.","created":"2012-09-07T11:35:47.506+0000"},{"body":"adding a patch that fixes column timestamp on deserialization. Anton, can you please apply and see if that fixes the problem?","created":"2012-09-07T11:45:27.228+0000"},{"body":"As I see from the log - timestamps were actually fixed to \"7 september 2012 cl 10:58 GMT\" but then overriten by remote migration, latest patch should help with that.","created":"2012-09-07T11:59:19.982+0000"},{"body":"When cassandra 1.1.5 will be available?","created":"2012-09-07T12:08:34.942+0000"},{"body":"I've patched & deployed to a couple of nodes where I now see the corrected timestamps. The NPE appears to also be gone.\n\nIt will take me some time to deploy to the entire ring but it looks promising so far.","created":"2012-09-07T13:16:13.907+0000"},{"body":"Ring upgraded with the second patch and I am now able to perform schema updates. Thanks!","created":"2012-09-07T14:42:38.165+0000"},{"body":"+1 lgtm","created":"2012-09-07T15:52:59.534+0000"},{"body":"Committed to cassandra-1.1 branch.","created":"2012-09-07T16:08:14.353+0000"}],"conversations":[{"body":"[default@test] show schema;\ncreate column family Messages\n with column_type = 'Standard'\n and comparator = 'AsciiType'\n and default_validation_class = 'BytesType'\n and key_validation_class = 'AsciiType'\n and read_repair_chance = 0.1\n and dclocal_read_repair_chance = 0.0\n and gc_grace = 864000\n and min_compaction_threshold = 2\n and max_compaction_threshold = 4\n and replicate_on_write = true\n and compaction_strategy = 'org.apache.cassandra.db.compaction.LeveledCompactionStrategy'\n and caching = 'KEYS_ONLY'\n and compaction_strategy_options = {'sstable_size_in_mb' : '1024'}\n and compression_options = {'chunk_length_kb' : '64', 'sstable_compression' : 'org.apache.cassandra.io.compress.DeflateCompressor'};\n\n\n[default@test] update column family Messages with min_compaction_threshold = 4 and max_compaction_threshold = 32;\na5b7544e-1ef5-3bfd-8770-c09594e37ec2\nWaiting for schema agreement...\n... schemas agree across the cluster\n\n[default@test] show schema;\ncreate column family Messages\n with column_type = 'Standard'\n and comparator = 'AsciiType'\n and default_validation_class = 'BytesType'\n and key_validation_class = 'AsciiType'\n and read_repair_chance = 0.1\n and dclocal_read_repair_chance = 0.0\n and gc_grace = 864000\n and min_compaction_threshold = 2\n and max_compaction_threshold = 4\n and replicate_on_write = true\n and compaction_strategy = 'org.apache.cassandra.db.compaction.LeveledCompactionStrategy'\n and caching = 'KEYS_ONLY'\n and compaction_strategy_options = {'sstable_size_in_mb' : '1024'}\n and compression_options = {'chunk_length_kb' : '64', 'sstable_compression' : 'org.apache.cassandra.io.compress.DeflateCompressor'};","from":"reporter","subject":"update column family fails"},{"body":"in logfile I see only this:\n INFO [MigrationStage:1] 2012-08-21 11:27:55,560 ColumnFamilyStore.java (line 659) Enqueuing flush of Memtable-schema_columnfamilies@970905946(1266/1582 serialized/live bytes, 20 ops)\n INFO [FlushWriter:5] 2012-08-21 11:27:55,561 Memtable.java (line 264) Writing Memtable-schema_columnfamilies@970905946(1266/1582 serialized/live bytes, 20 ops)\n INFO [FlushWriter:5] 2012-08-21 11:27:55,587 Memtable.java (line 305) Completed flushing /var/lib/cassandra/data/system/schema_columnfamilies/system-schema_columnfamilies-he-196-Data.db (1336 bytes) for commitlog position ReplayPosition(segmentId=4914817711083622, position=333055)","from":"developer"},{"body":"oh, and it obviously happens when I'm trying to change anything else, not only compaction treshhold, ex: sstable_size_in_mb","from":"developer"},{"body":"I saw a similar case. In my case, columns of schema_keyspaces and schema_columnfamilies in the system keyspace had future timestamp such as 2052-08-01 11:22:33. All operations of changing schemas were not affected. I do not know why timestamps had future values. I fixed it with rebuilding system keyspace.","from":"developer"},{"body":"Hmmm, it's possible. How did you rebuilt system keyspace?","from":"developer"},{"body":"Zenek, what version are you using? It's possible if you are on version less than 1.1.3 or if you have created your schema before 1.1.3 because of CASSANDRA-4432...","from":"developer"},{"body":"I'm using 1.1.4, but schema was created in 0.8.5\n\nCould you help me solve my problem step by step?","from":"developer"},{"body":"Did you upgrade directly to 1.1.4 or one of the previous 1.1 versions first? Please check timestamps in your schema_* ColumnFamilies and see if they have any future dates. The easiest way I see to fix this would be to re-create your schema if you have timestamp problems from CASSANDRA-4432 ","from":"developer"},{"body":"I've upgraded cassandra from 8.7 to 1.0, then to 1.0.5 -> 1.0.6 -> 1.1.0 -> 1.1.1 -> 1.1.2 -> 1.1.0 -> 1.1.3 -> 1.1.4.\n","from":"developer"},{"body":"aha, so check system.schema_* CFs column's timestamp values I bet they are from the future. The simplest fix for you would be to delete system/schema_* SSTables and re-create a schema after restart.","from":"developer"},{"body":"bq. check system.schema_* CFs column's timestamp values I bet they are from the future\n\nCan we fix that on startup for other upgraders?","from":"developer"},{"body":"Yes, I was about too :)","from":"developer"},{"body":"+1 I just ended up seeing this problem in our cluster which was upgraded from 1.1.2 to 1.1.3. Will probably have to find a workaround since I have to change schema now.","from":"developer"},{"body":"I took my cluster down for 10 minutes. I took a snapshot of 'show schema' into a text file and removed system KS from it so I will only have my KS schema definition in it. Then I stop cassandra on all nodes. On each node, I removed system/schema_* folders from system's keyspace folder in cassandra data dir. I started all cassandra nodes. When I tried to reload the schema file using cli to recreate my CFs, I kept getting the message that CFs already exist. When I listed schema_columnfamilies in one of the node, I saws the same long timestamps like 2705487066780774 on columns of that CF. So, the procedure didn't quiet work out for me. What could have gone wrong here?\n\nPlease advice.","from":"developer"},{"body":"What cassandra version are you running? Did you try to recreate a schema on all of the nodes or just one of them?","from":"developer"},{"body":"I am running 1.1.3 which was upgraded from 1.1.2. When nodes came up, I created the schemas in one node only.","from":"developer"},{"body":"As part of my process, I also removed saved_caches/system-*.","from":"developer"},{"body":"if you still have such timestamps and messages that CF already exists means that schema wasn't empty on restart and was reloaded from one of the nodes. I'm currently working on the patch which would fix timestamp situation on node's start up.","from":"developer"},{"body":"Do you have an ETA? I'll be happy to monkey apply it on my cluster and be the guinea pig for it.","from":"developer"},{"body":"Not yet but I'm working on it and soon as it's ready I will attach it here.","from":"developer"},{"body":"Don't you want to call cfs.truncate().get() to block for the truncate to finish? Rest LGTM.","from":"developer"},{"body":"good point, I would add that and commit, thanks!","from":"developer"},{"body":"Committed.","from":"developer"},{"body":"After applying this patch (1.1.5-tentative) I see the future timestamps being detected and fixed:\n\n{code}\nINFO [main] 2012-09-07 02:24:00,090 DefsTable.java (line 202) Fixing timestamps of schema ColumnFamily schema_keyspaces...\nINFO [main] 2012-09-07 02:24:00,168 DefsTable.java (line 202) Fixing timestamps of schema ColumnFamily schema_columnfamilies...\n{code}\n\nIf I list schema_keyspaces I still see the old (future) timestamps.\n\nGiven the timestamps weren't actually updated I did go back and found the following error shortly after the restarted node joined the ring (this may or not be related):\n\n{code}ERROR [InternalResponseStage:1] 2012-09-07 02:24:16,555 AbstractCassandraDaemon.java (line 135) Exception in thread Thread[InternalResponseStage:1,5,main]\njava.lang.NullPointerException\n at org.apache.cassandra.db.DefsTable.mergeColumnFamilies(DefsTable.java:468)\n at org.apache.cassandra.db.DefsTable.mergeSchema(DefsTable.java:346)\n at org.apache.cassandra.db.DefsTable.mergeRemoteSchema(DefsTable.java:324)\n at org.apache.cassandra.service.MigrationManager$MigrationTask$1.response(MigrationManager.java:416)\n at org.apache.cassandra.net.ResponseVerbHandler.doVerb(ResponseVerbHandler.java:45)\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:59)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n at java.lang.Thread.run(Thread.java:679)\n{code}\n\nI have however assumed that the patch automatically fixes this across all servers in a multi dc ring once this patch is applied. If there is a specific series of steps required to fix this, such as the full ring outage, deletion of system/schema_* etc can it please be documented here, the wiki or in the NEWS.txt file under upgrading notes?","from":"developer"},{"body":"Yes it does, but it assumes that you have schema in agreement across the nodes before it tries to fix timestamps. I can put that into NEWS.txt","from":"developer"},{"body":"My schema is in agreement, but the timestamps aren't fixed, even though the log suggests otherwise (is that previously mentioned NPE related?). \n\n{code}\n[default@unknown] describe cluster;\nCluster Information:\n Snitch: org.apache.cassandra.locator.PropertyFileSnitch\n Partitioner: org.apache.cassandra.dht.RandomPartitioner\n Schema versions: \n 89b22434-5e34-381d-83d1-2a3cde1482fe: [x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x, x]\n\n[default@unknown]\n{code}\n\nSchema updates continue to silently fail.","from":"developer"},{"body":"Interesting, why does it try to merge remote schema if all of your nodes are in agreement, can you run in debug mode as attach the log? The thing is timestamps are fixed long before the storage server is initialized and the process involves truncate of the system.schema_* CFs so I don't see any other reason why your schema still has old timestamps except some other node sent it migration request... \n\nNPE that you experience is also interesting, cfMetaData() is always not null as CFs are copied from one map to another on KS initialization, and KS itself could not be null or you would have the assertion error instead of NPE.","from":"developer"},{"body":"Maybe we should reopen this if that didn't fully fixed the problem? I'd better hold on a bit on 1.1.5 if this is not fixed yet as this seem to hit quite a few people.","from":"developer"},{"body":"We are not yet sure what is the problem so I think we should hold on reopening for a bit.","from":"developer"},{"body":"I just thought that a good addition to existing patch could be change in ColumnSerializer.deserialize to fix timestamps from the future, that way migrations from the remote locations would be deserialized with correct timestamp even if they were sent with the wrong one...","from":"developer"},{"body":"Well, Anton still sees timestamp in the future and is still unable to do schema updates, it does sound a lot like there is a problem and it's related to this ticket. I just meant that I prefer keeping that in mind before releasing 1.1.5 into the wild too quickly.","from":"developer"},{"body":"I understand, but we don't know if that is caused by the fix not working or by something else as exception indicates that something was send to the node, this could be a separate issue. I will work on the ColumnSerializer change I mentioned and we'll see if it helps. ","from":"developer"},{"body":"Debug log sent to [~xedin] privately.","from":"developer"},{"body":"adding a patch that fixes column timestamp on deserialization. Anton, can you please apply and see if that fixes the problem?","from":"developer"},{"body":"As I see from the log - timestamps were actually fixed to \"7 september 2012 cl 10:58 GMT\" but then overriten by remote migration, latest patch should help with that.","from":"developer"},{"body":"When cassandra 1.1.5 will be available?","from":"developer"},{"body":"I've patched & deployed to a couple of nodes where I now see the corrected timestamps. The NPE appears to also be gone.\n\nIt will take me some time to deploy to the entire ring but it looks promising so far.","from":"developer"},{"body":"Ring upgraded with the second patch and I am now able to perform schema updates. Thanks!","from":"developer"},{"body":"+1 lgtm","from":"developer"},{"body":"Committed to cassandra-1.1 branch.","from":"developer"}],"created":"2012-08-21T09:24:22.000+0000","description":"[default@test] show schema;\ncreate column family Messages\n with column_type = 'Standard'\n and comparator = 'AsciiType'\n and default_validation_class = 'BytesType'\n and key_validation_class = 'AsciiType'\n and read_repair_chance = 0.1\n and dclocal_read_repair_chance = 0.0\n and gc_grace = 864000\n and min_compaction_threshold = 2\n and max_compaction_threshold = 4\n and replicate_on_write = true\n and compaction_strategy = 'org.apache.cassandra.db.compaction.LeveledCompactionStrategy'\n and caching = 'KEYS_ONLY'\n and compaction_strategy_options = {'sstable_size_in_mb' : '1024'}\n and compression_options = {'chunk_length_kb' : '64', 'sstable_compression' : 'org.apache.cassandra.io.compress.DeflateCompressor'};\n\n\n[default@test] update column family Messages with min_compaction_threshold = 4 and max_compaction_threshold = 32;\na5b7544e-1ef5-3bfd-8770-c09594e37ec2\nWaiting for schema agreement...\n... schemas agree across the cluster\n\n[default@test] show schema;\ncreate column family Messages\n with column_type = 'Standard'\n and comparator = 'AsciiType'\n and default_validation_class = 'BytesType'\n and key_validation_class = 'AsciiType'\n and read_repair_chance = 0.1\n and dclocal_read_repair_chance = 0.0\n and gc_grace = 864000\n and min_compaction_threshold = 2\n and max_compaction_threshold = 4\n and replicate_on_write = true\n and compaction_strategy = 'org.apache.cassandra.db.compaction.LeveledCompactionStrategy'\n and caching = 'KEYS_ONLY'\n and compaction_strategy_options = {'sstable_size_in_mb' : '1024'}\n and compression_options = {'chunk_length_kb' : '64', 'sstable_compression' : 'org.apache.cassandra.io.compress.DeflateCompressor'};","issue_id":"12604193","key":"CASSANDRA-4561","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2012-09-04T21:17:21.000+0000","role":"fixed_distractor","summary":"update column family fails"} {"case_id":"12437517","cluster":"DISTRACTOR-CASSANDRA-475","comments":[{"body":"We don't have the resources to devote to properly fixing THRIFT-601, so we'll have to close this as wontfix until or unless something changes there.\n\nIf you need to expose Cassandra to untrusted sources, I suggest helping with Avro integration. (Ask Eric for how, he is the closest to that.)","created":"2010-03-10T19:42:11.352+0000"},{"body":"THRIFT-601 has reportedly been fixed. Not sure how painful their fix is for us, we should have a look.","created":"2010-06-10T15:24:31.099+0000"},{"body":"Do you know if their patch has been committed to to cassandra's trunk?","created":"2010-07-02T14:49:26.482+0000"},{"body":"this ticket is open precisely to incorporate a newer thrift and make any other necessary changes to fix the problem","created":"2010-07-02T15:02:46.030+0000"},{"body":"looks like we need to set a transport max length and a protocol read length (besides upgrading to a new thrift jar).\n\nunclear if these would prohibit using large binary column values. may need to make them configurable. (would transport length be read length + some overhead? or is transport length just a buffer size that we can leave at some reasonably safe value and is transparent to the rest?)","created":"2010-07-09T05:25:11.898+0000"},{"body":"TFramedTransport maxLength is just a frame buffer size. It defaults to 2^31-1 in the latest thrift. Im not sure tweaking that buys us much given the following. \n\nTBinaryProtocol.readLength is what needs to be set. *NOTE*: The value provided here will effectively be the max column size (there is no overhead due to thrift's delimitation granularity). \n\nThis should be configurable with a sane default. How about thrift_max_message_length with a default of 1 (in MB). (This is the default max_packet_size in mysql, so might be easier to grok if we match that). \n\n\n","created":"2010-07-10T01:06:02.394+0000"},{"body":"so readLength is the max size of a single field, and maxLength the max size of the entire rpc call?\n\nis maxLength actually controlling a byte[] size somewhere, or not? if it is then we do need to set it, or if the first 4 bytes correspond to a large int we still OOM.","created":"2010-07-10T03:02:10.343+0000"},{"body":"Yes, maxLength does directly control the size of the input byte[] on TFramedTransport, so it is succeptible to the same issue. \n\nI would like to change thrift_framed_transport to thrift_framed_transport_size with 0 (in MB as well) as the default indicating TSocket over TFramedTransport","created":"2010-07-10T05:33:58.745+0000"},{"body":"Looks like the thrift folks have changed the parameterization of TBase: THRIFT-759\n\nUnfortunately, this beat our patch for THRIFT-601 by a couple of days. Thoughts?","created":"2010-07-10T06:30:33.643+0000"},{"body":"If you're asking \"is this worth having to rewrite a bunch of our code\" then the answer is \"yes, but this is why we upgrade thrift version so infrequently.\"","created":"2010-07-10T14:18:22.653+0000"},{"body":"We are agreed on the importance of this. The changes on the surface seem trivial, since Comparator is widely used already. My concerns here are strictly schedule related given the process overhead involved in touching so many files in active development. \n\nI'm going to have to put this down temporarily, until I can focus on this and make the change a one-shot deal (or as close to that as possible). \n\nEDIT: I just added CASSANDRA-1266 to track the jar upgrade separately so we don't muddy the waters on this issue.","created":"2010-07-10T16:50:22.091+0000"},{"body":"I can continue with the config modifications required if the changes mentioned above are cool","created":"2010-07-10T17:05:36.194+0000"},{"body":"- trunk-475-config.txt updates the yaml config\n- trunk-475-src.txt updates java src\n\nEDIT: I dropped in a default setting of 256mb in the Converter when someone had framed transport configured in storage-conf. Was not sure what all to do there except add a sane default and a warning message. ","created":"2010-07-11T01:05:16.392+0000"},{"body":"do we need a sanity check that frame size must be < message length?","created":"2010-07-14T02:22:58.353+0000"},{"body":"I thought about that, but it's pretty clear exception wise when you go over the frame size that it is a transport issue","created":"2010-07-14T03:42:14.485+0000"},{"body":"Come to think, nice to have completeness-wise and was easy enough to add. \n\ntrunk-475-src-2.txt superseeds trunk-475-src.txt ","created":"2010-07-14T03:54:35.350+0000"},{"body":"hmm, why isn't that check firing with the default frame of 256 and length of 1?","created":"2010-07-14T05:18:22.924+0000"},{"body":"I'm gonna blame that one on your central-TX weather melting my brain. \n\ntrunk-475-src-3.txt replaces trunk-475-src-2.txt, fixes comparisson induced by mild heat stroke. ","created":"2010-07-14T06:17:44.587+0000"},{"body":"is thrift smart enough to re-use those byte[] buffers for multiple requests? allocating byte[] is expensive in java.","created":"2010-07-14T14:01:10.093+0000"},{"body":"Both TBinaryProtocol and TFramedTransport use the limits in checks only. The size used to construct the byte[] is pulled from the first couple bytes of any thrift message. When junk was sent, this size was getting intepreted as extremely large values.","created":"2010-07-14T22:38:09.593+0000"},{"body":"That makes sense.\n\nThat would be an obvious optimization for them to make, then -- allocate a buffer at the max size, and re-use it.","created":"2010-07-15T01:22:10.216+0000"},{"body":"committed","created":"2010-07-15T01:22:49.374+0000"},{"body":"I'm occasionally getting this error with stress.py after doing a few million inserts:\n\nERROR 13:30:12,158 Thrift error occurred during processing of message.\norg.apache.thrift.TException: Message length exceeded: 32\n at org.apache.thrift.protocol.TBinaryProtocol.checkReadLength(TBinaryProtocol.java:384)\n at org.apache.thrift.protocol.TBinaryProtocol.readBinary(TBinaryProtocol.java:361)\n at org.apache.cassandra.thrift.Column.read(Column.java:498)\n at org.apache.cassandra.thrift.ColumnOrSuperColumn.read(ColumnOrSuperColumn.java:351)\n at org.apache.cassandra.thrift.Mutation.read(Mutation.java:346)\n at org.apache.cassandra.thrift.Cassandra$batch_mutate_args.read(Cassandra.java:16780)\n at org.apache.cassandra.thrift.Cassandra$Processor$batch_mutate.process(Cassandra.java:3041)\n at org.apache.cassandra.thrift.Cassandra$Processor.process(Cassandra.java:2531)\n at org.apache.cassandra.thrift.CustomTThreadPoolServer$WorkerProcess.run(CustomTThreadPoolServer.java:167)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:619)","created":"2010-07-15T13:42:57.858+0000"},{"body":"Hudson also found -this- an error in build 492.\nI assume the hadoop test that failed found it by doing a lot of inserts as above.","created":"2010-07-15T19:04:49.054+0000"},{"body":"checkReadLength in TBinaryProtocol looks like it is treating the readLength attribute as an instance variable and is therefore slowly getting decremented on each call!\n\nI'm verifying my findings w/ thrift folks now","created":"2010-07-15T20:34:40.171+0000"},{"body":"THRIFT-820 created and patch submitted. Waiting on their review and feedback.","created":"2010-07-15T20:46:12.707+0000"},{"body":"Resets the input and output protocol on each after each successful call to process in CustomTThreadPoolServer\n\npasses nosetests and stress.py with 5 million rows","created":"2010-07-20T18:52:50.117+0000"},{"body":"Verified this solves the issue and checked for any adverse performance effects. Committed.","created":"2010-07-20T20:30:09.915+0000"},{"body":"is this going to hurt performance? (not sure how heavyweight getProtocol is)","created":"2010-07-20T20:59:47.524+0000"},{"body":"I checked the performance and couldn't see any noticeable difference. The extra garbage might minorly exacerbate CASSANDRA-1014.","created":"2010-07-20T21:05:03.372+0000"},{"body":"Hello,\n\nI am using the latest source code from trunk. SVN 961952\n\nAfter few hundred thousand inserts cassandra crashes and is throwing 2 different types of exceptions:\nThe first one being:\norg.apache.thrift.transport.TTransportException\nat org.apache.thrift.transport.TIOStreamTransport.read(TIOStreamTransport.java:132)\nat org.apache.thrift.transport.TTransport.readAll(TTransport.java:84)\nat org.apache.thrift.transport.TFramedTransport.readFrame(TFramedTransport.java:129)\nat org.apache.thrift.transport.TFramedTransport.read(TFramedTransport.java:101)\nat org.apache.thrift.transport.TTransport.readAll(TTransport.java:84)\nat org.apache.thrift.protocol.TBinaryProtocol.readAll(TBinaryProtocol.java:369)\nat org.apache.thrift.protocol.TBinaryProtocol.readI32(TBinaryProtocol.java:295)\nat org.apache.thrift.protocol.TBinaryProtocol.readMessageBegin(TBinaryProtocol.java:202)\nat org.apache.cassandra.thrift.Cassandra$Client.recv_batch_mutate(Cassandra.java:960)\nat org.apache.cassandra.thrift.Cassandra$Client.batch_mutate(Cassandra.java:944)\nat com.cbsi.pi.rtss.data.cassandra.CassandraDataManager.insert(CassandraDataManager.java:107)\nat com.cbsi.pi.rtss.service.bulk.BulkThread.run(BulkThread.java:59)\nat java.lang.Thread.run(Unknown Source)\n\nand the second one being:\norg.apache.thrift.transport.TTransportException: java.net.SocketException: Software caused connection abort: socket write error\nat org.apache.thrift.transport.TIOStreamTransport.write(TIOStreamTransport.java:147)\nat org.apache.thrift.transport.TFramedTransport.flush(TFramedTransport.java:156)\nat org.apache.cassandra.thrift.Cassandra$Client.send_set_keyspace(Cassandra.java:441)\nat org.apache.cassandra.thrift.Cassandra$Client.set_keyspace(Cassandra.java:430)\nat com.cbsi.pi.rtss.data.cassandra.CassandraDataManager.insert(CassandraDataManager.java:106)\nat com.cbsi.pi.rtss.service.bulk.BulkThread.run(BulkThread.java:59)\nat java.lang.Thread.run(Unknown Source)\nCaused by: java.net.SocketException: Software caused connection abort: socket write error\nat java.net.SocketOutputStream.socketWrite0(Native Method)\nat java.net.SocketOutputStream.socketWrite(Unknown Source)\nat java.net.SocketOutputStream.write(Unknown Source)\nat java.io.BufferedOutputStream.flushBuffer(Unknown Source)\nat java.io.BufferedOutputStream.write(Unknown Source)\nat org.apache.thrift.transport.TIOStreamTransport.write(TIOStreamTransport.java:145)\n... 6 more\n\nI was not getting this error yesterday but this morning when I updated my svn trunk I got update for following 2 files:\nU src/java/org/apache/cassandra/thrift/CustomTThreadPoolServer.java\nU src/java/org/apache/cassandra/scheduler/RoundRobinScheduler.java\n\nOne more thing.\n\nBefore I did a SVN update this morning, I checkout thrift and applied the patch suggested by Nate in bug THRIFT-820 https://issues.apache.org/jira/browse/THRIFT-820. When I rebuild thrift jar file and used it, I had no crashes or exceptions yesterday.\n\nBut today with/without thrift patch, I am getting exceptions mentioned above.\n\nThanks,\nJignesh\n\nand that is causing the problem.","created":"2010-07-21T16:15:08.975+0000"},{"body":"Apologize for the confusion. I did another svn update and picked bunch of new files. I am not getting any cassandra crashes or exceptions any more.","created":"2010-07-21T16:27:41.300+0000"}],"conversations":[{"body":"Use dd if=/dev/urandom count=1 | nc $host 9160 as a handy recipe for shutting a cassandra instance down. \n\nThrift has spoken (see THRIFT-601), but \"Don't Do That\" is probably an insufficient answer for our users. ","from":"reporter","subject":"sending random data crashes thrift service"},{"body":"We don't have the resources to devote to properly fixing THRIFT-601, so we'll have to close this as wontfix until or unless something changes there.\n\nIf you need to expose Cassandra to untrusted sources, I suggest helping with Avro integration. (Ask Eric for how, he is the closest to that.)","from":"developer"},{"body":"THRIFT-601 has reportedly been fixed. Not sure how painful their fix is for us, we should have a look.","from":"developer"},{"body":"Do you know if their patch has been committed to to cassandra's trunk?","from":"developer"},{"body":"this ticket is open precisely to incorporate a newer thrift and make any other necessary changes to fix the problem","from":"developer"},{"body":"looks like we need to set a transport max length and a protocol read length (besides upgrading to a new thrift jar).\n\nunclear if these would prohibit using large binary column values. may need to make them configurable. (would transport length be read length + some overhead? or is transport length just a buffer size that we can leave at some reasonably safe value and is transparent to the rest?)","from":"developer"},{"body":"TFramedTransport maxLength is just a frame buffer size. It defaults to 2^31-1 in the latest thrift. Im not sure tweaking that buys us much given the following. \n\nTBinaryProtocol.readLength is what needs to be set. *NOTE*: The value provided here will effectively be the max column size (there is no overhead due to thrift's delimitation granularity). \n\nThis should be configurable with a sane default. How about thrift_max_message_length with a default of 1 (in MB). (This is the default max_packet_size in mysql, so might be easier to grok if we match that). \n\n\n","from":"developer"},{"body":"so readLength is the max size of a single field, and maxLength the max size of the entire rpc call?\n\nis maxLength actually controlling a byte[] size somewhere, or not? if it is then we do need to set it, or if the first 4 bytes correspond to a large int we still OOM.","from":"developer"},{"body":"Yes, maxLength does directly control the size of the input byte[] on TFramedTransport, so it is succeptible to the same issue. \n\nI would like to change thrift_framed_transport to thrift_framed_transport_size with 0 (in MB as well) as the default indicating TSocket over TFramedTransport","from":"developer"},{"body":"Looks like the thrift folks have changed the parameterization of TBase: THRIFT-759\n\nUnfortunately, this beat our patch for THRIFT-601 by a couple of days. Thoughts?","from":"developer"},{"body":"If you're asking \"is this worth having to rewrite a bunch of our code\" then the answer is \"yes, but this is why we upgrade thrift version so infrequently.\"","from":"developer"},{"body":"We are agreed on the importance of this. The changes on the surface seem trivial, since Comparator is widely used already. My concerns here are strictly schedule related given the process overhead involved in touching so many files in active development. \n\nI'm going to have to put this down temporarily, until I can focus on this and make the change a one-shot deal (or as close to that as possible). \n\nEDIT: I just added CASSANDRA-1266 to track the jar upgrade separately so we don't muddy the waters on this issue.","from":"developer"},{"body":"I can continue with the config modifications required if the changes mentioned above are cool","from":"developer"},{"body":"- trunk-475-config.txt updates the yaml config\n- trunk-475-src.txt updates java src\n\nEDIT: I dropped in a default setting of 256mb in the Converter when someone had framed transport configured in storage-conf. Was not sure what all to do there except add a sane default and a warning message. ","from":"developer"},{"body":"do we need a sanity check that frame size must be < message length?","from":"developer"},{"body":"I thought about that, but it's pretty clear exception wise when you go over the frame size that it is a transport issue","from":"developer"},{"body":"Come to think, nice to have completeness-wise and was easy enough to add. \n\ntrunk-475-src-2.txt superseeds trunk-475-src.txt ","from":"developer"},{"body":"hmm, why isn't that check firing with the default frame of 256 and length of 1?","from":"developer"},{"body":"I'm gonna blame that one on your central-TX weather melting my brain. \n\ntrunk-475-src-3.txt replaces trunk-475-src-2.txt, fixes comparisson induced by mild heat stroke. ","from":"developer"},{"body":"is thrift smart enough to re-use those byte[] buffers for multiple requests? allocating byte[] is expensive in java.","from":"developer"},{"body":"Both TBinaryProtocol and TFramedTransport use the limits in checks only. The size used to construct the byte[] is pulled from the first couple bytes of any thrift message. When junk was sent, this size was getting intepreted as extremely large values.","from":"developer"},{"body":"That makes sense.\n\nThat would be an obvious optimization for them to make, then -- allocate a buffer at the max size, and re-use it.","from":"developer"},{"body":"committed","from":"developer"},{"body":"I'm occasionally getting this error with stress.py after doing a few million inserts:\n\nERROR 13:30:12,158 Thrift error occurred during processing of message.\norg.apache.thrift.TException: Message length exceeded: 32\n at org.apache.thrift.protocol.TBinaryProtocol.checkReadLength(TBinaryProtocol.java:384)\n at org.apache.thrift.protocol.TBinaryProtocol.readBinary(TBinaryProtocol.java:361)\n at org.apache.cassandra.thrift.Column.read(Column.java:498)\n at org.apache.cassandra.thrift.ColumnOrSuperColumn.read(ColumnOrSuperColumn.java:351)\n at org.apache.cassandra.thrift.Mutation.read(Mutation.java:346)\n at org.apache.cassandra.thrift.Cassandra$batch_mutate_args.read(Cassandra.java:16780)\n at org.apache.cassandra.thrift.Cassandra$Processor$batch_mutate.process(Cassandra.java:3041)\n at org.apache.cassandra.thrift.Cassandra$Processor.process(Cassandra.java:2531)\n at org.apache.cassandra.thrift.CustomTThreadPoolServer$WorkerProcess.run(CustomTThreadPoolServer.java:167)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:619)","from":"developer"},{"body":"Hudson also found -this- an error in build 492.\nI assume the hadoop test that failed found it by doing a lot of inserts as above.","from":"developer"},{"body":"checkReadLength in TBinaryProtocol looks like it is treating the readLength attribute as an instance variable and is therefore slowly getting decremented on each call!\n\nI'm verifying my findings w/ thrift folks now","from":"developer"},{"body":"THRIFT-820 created and patch submitted. Waiting on their review and feedback.","from":"developer"},{"body":"Resets the input and output protocol on each after each successful call to process in CustomTThreadPoolServer\n\npasses nosetests and stress.py with 5 million rows","from":"developer"},{"body":"Verified this solves the issue and checked for any adverse performance effects. Committed.","from":"developer"},{"body":"is this going to hurt performance? (not sure how heavyweight getProtocol is)","from":"developer"},{"body":"I checked the performance and couldn't see any noticeable difference. The extra garbage might minorly exacerbate CASSANDRA-1014.","from":"developer"},{"body":"Hello,\n\nI am using the latest source code from trunk. SVN 961952\n\nAfter few hundred thousand inserts cassandra crashes and is throwing 2 different types of exceptions:\nThe first one being:\norg.apache.thrift.transport.TTransportException\nat org.apache.thrift.transport.TIOStreamTransport.read(TIOStreamTransport.java:132)\nat org.apache.thrift.transport.TTransport.readAll(TTransport.java:84)\nat org.apache.thrift.transport.TFramedTransport.readFrame(TFramedTransport.java:129)\nat org.apache.thrift.transport.TFramedTransport.read(TFramedTransport.java:101)\nat org.apache.thrift.transport.TTransport.readAll(TTransport.java:84)\nat org.apache.thrift.protocol.TBinaryProtocol.readAll(TBinaryProtocol.java:369)\nat org.apache.thrift.protocol.TBinaryProtocol.readI32(TBinaryProtocol.java:295)\nat org.apache.thrift.protocol.TBinaryProtocol.readMessageBegin(TBinaryProtocol.java:202)\nat org.apache.cassandra.thrift.Cassandra$Client.recv_batch_mutate(Cassandra.java:960)\nat org.apache.cassandra.thrift.Cassandra$Client.batch_mutate(Cassandra.java:944)\nat com.cbsi.pi.rtss.data.cassandra.CassandraDataManager.insert(CassandraDataManager.java:107)\nat com.cbsi.pi.rtss.service.bulk.BulkThread.run(BulkThread.java:59)\nat java.lang.Thread.run(Unknown Source)\n\nand the second one being:\norg.apache.thrift.transport.TTransportException: java.net.SocketException: Software caused connection abort: socket write error\nat org.apache.thrift.transport.TIOStreamTransport.write(TIOStreamTransport.java:147)\nat org.apache.thrift.transport.TFramedTransport.flush(TFramedTransport.java:156)\nat org.apache.cassandra.thrift.Cassandra$Client.send_set_keyspace(Cassandra.java:441)\nat org.apache.cassandra.thrift.Cassandra$Client.set_keyspace(Cassandra.java:430)\nat com.cbsi.pi.rtss.data.cassandra.CassandraDataManager.insert(CassandraDataManager.java:106)\nat com.cbsi.pi.rtss.service.bulk.BulkThread.run(BulkThread.java:59)\nat java.lang.Thread.run(Unknown Source)\nCaused by: java.net.SocketException: Software caused connection abort: socket write error\nat java.net.SocketOutputStream.socketWrite0(Native Method)\nat java.net.SocketOutputStream.socketWrite(Unknown Source)\nat java.net.SocketOutputStream.write(Unknown Source)\nat java.io.BufferedOutputStream.flushBuffer(Unknown Source)\nat java.io.BufferedOutputStream.write(Unknown Source)\nat org.apache.thrift.transport.TIOStreamTransport.write(TIOStreamTransport.java:145)\n... 6 more\n\nI was not getting this error yesterday but this morning when I updated my svn trunk I got update for following 2 files:\nU src/java/org/apache/cassandra/thrift/CustomTThreadPoolServer.java\nU src/java/org/apache/cassandra/scheduler/RoundRobinScheduler.java\n\nOne more thing.\n\nBefore I did a SVN update this morning, I checkout thrift and applied the patch suggested by Nate in bug THRIFT-820 https://issues.apache.org/jira/browse/THRIFT-820. When I rebuild thrift jar file and used it, I had no crashes or exceptions yesterday.\n\nBut today with/without thrift patch, I am getting exceptions mentioned above.\n\nThanks,\nJignesh\n\nand that is causing the problem.","from":"developer"},{"body":"Apologize for the confusion. I did another svn update and picked bunch of new files. I am not getting any cassandra crashes or exceptions any more.","from":"developer"}],"created":"2009-10-07T14:44:21.000+0000","description":"Use dd if=/dev/urandom count=1 | nc $host 9160 as a handy recipe for shutting a cassandra instance down. \n\nThrift has spoken (see THRIFT-601), but \"Don't Do That\" is probably an insufficient answer for our users. ","issue_id":"12437517","key":"CASSANDRA-475","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2010-07-21T16:27:41.000+0000","role":"fixed_distractor","summary":"sending random data crashes thrift service"} {"case_id":"12626383","cluster":"DISTRACTOR-CASSANDRA-5125","comments":[{"body":"Attaching patches for that (also pushed to https://github.com/pcmanus/cassandra/commits/5125).\n\nThe first part of this ticket is about how we store the information that a clustering key column is indexed. Turns out that for \"regular\" columns we use ColumnDefinition and the indexing code also assumes that, so the probably best and simplest approach is to reuse ColumnDefinition for that too. But then it's easier to always store all primary key columns as ColumnDefinition, pretty much obsoleting the old key_aliases and column_aliases. There is a few related details worth noticing:\n# while this obsolete the aliases, those are not removed of the schema by the patch for compatibility sake. Truth is, I'm not sure there is a way to remove a field from the schema without breaking rolling upgrades at this point.\n# after this patch, CFDefinition becomes much less useful as CFMetadata + ColumnDefinition holds pretty much the same information in pretty much the same form. So we could slightly simplify things by removing CFDefinition. However, this is left to later (this won't be a 3 lines patch).\n\nAfter that, the patch adds a new type of composite indexes to handle indexing clustering keys (which share most code with the existing regular composite index) and update CQL3 to allow adding and querying the new indexes (in particular, it is slighty tricky in SelectStatment to recognize when a clustering key is restricted if 2ndary indexes should be used or not).\n\nThe last patch adds support for indexing components of the partition key (we don't allow indexing the first component of the partition key as it makes no sense (it's already primary indexed), but if the partition key is composite, secondary indexing the 2+ parts can be useful).\n\nLastly, I'll note that the patches only add theses news indexes for non compact tables. We should generalize to compact tables too, but that would require a bit of generalization that I'd rather add in a second phase.\n","created":"2013-02-18T05:25:01.462+0000"},{"body":"Rebased patches attached.","created":"2013-04-02T17:42:20.837+0000"},{"body":"Things move fast on trunk lately, so I've pushed a rebased version at https://github.com/pcmanus/cassandra/commits/5125-2 to avoid rebasing every day.","created":"2013-04-03T08:31:33.913+0000"},{"body":"A few errors when running the test suite:\n\nTestcase: testCli(org.apache.cassandra.cli.CliTest):\tCaused an ERROR\njava.lang.RuntimeException: org.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 4\n\tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1533)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:895)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:918)\n\tat java.lang.Thread.run(Thread.java:680)\nCaused by: org.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 4\n\tat org.apache.cassandra.db.marshal.LongType.getString(LongType.java:69)\n\tat org.apache.cassandra.db.index.AbstractSimplePerColumnSecondaryIndex.insert(AbstractSimplePerColumnSecondaryIndex.java:121)\n\tat org.apache.cassandra.db.index.SecondaryIndexManager$PerColumnIndexUpdater.update(SecondaryIndexManager.java:623)\n\tat org.apache.cassandra.db.AtomicSortedColumns$Holder.addColumn(AtomicSortedColumns.java:313)\n\tat org.apache.cassandra.db.AtomicSortedColumns.addAllWithSizeDelta(AtomicSortedColumns.java:168)\n\tat org.apache.cassandra.db.Memtable.resolve(Memtable.java:253)\n\tat org.apache.cassandra.db.Memtable.put(Memtable.java:169)\n\tat org.apache.cassandra.db.ColumnFamilyStore.apply(ColumnFamilyStore.java:852)\n\tat org.apache.cassandra.db.Table.apply(Table.java:379)\n\tat org.apache.cassandra.db.Table.apply(Table.java:342)\n\tat org.apache.cassandra.db.RowMutation.apply(RowMutation.java:189)\n\tat org.apache.cassandra.service.StorageProxy$6.runMayThrow(StorageProxy.java:667)\n\tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1529)\n\t... 3 more\n\t\nTestcase: testIndexDeletions(org.apache.cassandra.db.ColumnFamilyStoreTest):\tCaused an ERROR\nA long is exactly 8 bytes: 4\norg.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 4\n\tat org.apache.cassandra.db.marshal.LongType.getString(LongType.java:69)\n\tat org.apache.cassandra.db.index.AbstractSimplePerColumnSecondaryIndex.insert(AbstractSimplePerColumnSecondaryIndex.java:121)\n\tat org.apache.cassandra.db.index.SecondaryIndexManager$PerColumnIndexUpdater.update(SecondaryIndexManager.java:623)\n\tat org.apache.cassandra.db.AtomicSortedColumns$Holder.addColumn(AtomicSortedColumns.java:313)\n\tat org.apache.cassandra.db.AtomicSortedColumns.addAllWithSizeDelta(AtomicSortedColumns.java:168)\n\tat org.apache.cassandra.db.Memtable.resolve(Memtable.java:253)\n\tat org.apache.cassandra.db.Memtable.put(Memtable.java:169)\n\tat org.apache.cassandra.db.ColumnFamilyStore.apply(ColumnFamilyStore.java:852)\n\tat org.apache.cassandra.db.Table.apply(Table.java:379)\n\tat org.apache.cassandra.db.Table.apply(Table.java:342)\n\tat org.apache.cassandra.db.RowMutation.apply(RowMutation.java:189)\n\tat org.apache.cassandra.db.ColumnFamilyStoreTest.testIndexDeletions(ColumnFamilyStoreTest.java:301)\n\n\nTestcase: testIndexUpdate(org.apache.cassandra.db.ColumnFamilyStoreTest):\tCaused an ERROR\nIndex: 0, Size: 0\njava.lang.IndexOutOfBoundsException: Index: 0, Size: 0\n\tat java.util.ArrayList.RangeCheck(ArrayList.java:547)\n\tat java.util.ArrayList.get(ArrayList.java:322)\n\tat org.apache.cassandra.db.ColumnFamilyStoreTest.testIndexUpdate(ColumnFamilyStoreTest.java:398)","created":"2013-04-04T04:16:15.243+0000"},{"body":"My bad. That was due to a rebase typo that I had fixed in my CASSANDRA-5417 branch but not on that one. I've push the fix to the same github branch than above (https://github.com/pcmanus/cassandra/commits/5125-2).","created":"2013-04-04T07:10:32.289+0000"},{"body":"+1","created":"2013-04-04T09:35:43.521+0000"},{"body":"Committed, thanks","created":"2013-04-04T16:33:46.142+0000"},{"body":"??Lastly, I'll note that the patches only add theses news indexes for non compact tables. We should generalize to compact tables too, but that would require a bit of generalization that I'd rather add in a second phase.??\n\nWith 2.1 and *compact* tables it is possible to CREATE INDEX on composite primary key columns, but queries returns no results for the tests below.\nAdding this comment for now, can open a new ticket if you prefer.\n\n{code:SQL}\nCREATE TABLE users2 (\n userID uuid,\n fname text,\n zip int,\n state text,\n PRIMARY KEY ((userID, fname))\n) WITH COMPACT STORAGE;\n\nCREATE INDEX ON users2 (userID);\nCREATE INDEX ON users2 (fname);\n\nINSERT INTO users2 (userID, fname, zip, state) VALUES (b3e3bc33-b237-4b55-9337-3d41de9a5649, 'John', 10007, 'NY');\n\n// the following queries returns 0 rows, instead of 1 expected\nSELECT * FROM users2 WHERE fname='John'; \nSELECT * FROM users2 WHERE userid=b3e3bc33-b237-4b55-9337-3d41de9a5649;\nSELECT * FROM users2 WHERE userid=b3e3bc33-b237-4b55-9337-3d41de9a5649 AND fname='John';\n\n// dropping 2ndary indexes restore normal behavior\n{code}","created":"2014-10-20T20:45:47.860+0000"},{"body":"[~denis.angilella] Correct, the validation during index creation is broken. Do you mind creating a ticket indeed so we track the fix?","created":"2014-10-21T10:18:04.527+0000"},{"body":"[~slebresne]: I created CASSANDRA-8156 to track the fix.","created":"2014-10-21T11:07:08.337+0000"}],"conversations":[{"body":"Given\n\n{code}\nCREATE TABLE foo (\n a int,\n b int,\n c int,\n PRIMARY KEY (a, b)\n);\n{code}\n\nWe should support {{CREATE INDEX ON foo(b)}}.","from":"reporter","subject":"Support indexes on composite column components (clustered columns)"},{"body":"Attaching patches for that (also pushed to https://github.com/pcmanus/cassandra/commits/5125).\n\nThe first part of this ticket is about how we store the information that a clustering key column is indexed. Turns out that for \"regular\" columns we use ColumnDefinition and the indexing code also assumes that, so the probably best and simplest approach is to reuse ColumnDefinition for that too. But then it's easier to always store all primary key columns as ColumnDefinition, pretty much obsoleting the old key_aliases and column_aliases. There is a few related details worth noticing:\n# while this obsolete the aliases, those are not removed of the schema by the patch for compatibility sake. Truth is, I'm not sure there is a way to remove a field from the schema without breaking rolling upgrades at this point.\n# after this patch, CFDefinition becomes much less useful as CFMetadata + ColumnDefinition holds pretty much the same information in pretty much the same form. So we could slightly simplify things by removing CFDefinition. However, this is left to later (this won't be a 3 lines patch).\n\nAfter that, the patch adds a new type of composite indexes to handle indexing clustering keys (which share most code with the existing regular composite index) and update CQL3 to allow adding and querying the new indexes (in particular, it is slighty tricky in SelectStatment to recognize when a clustering key is restricted if 2ndary indexes should be used or not).\n\nThe last patch adds support for indexing components of the partition key (we don't allow indexing the first component of the partition key as it makes no sense (it's already primary indexed), but if the partition key is composite, secondary indexing the 2+ parts can be useful).\n\nLastly, I'll note that the patches only add theses news indexes for non compact tables. We should generalize to compact tables too, but that would require a bit of generalization that I'd rather add in a second phase.\n","from":"developer"},{"body":"Rebased patches attached.","from":"developer"},{"body":"Things move fast on trunk lately, so I've pushed a rebased version at https://github.com/pcmanus/cassandra/commits/5125-2 to avoid rebasing every day.","from":"developer"},{"body":"A few errors when running the test suite:\n\nTestcase: testCli(org.apache.cassandra.cli.CliTest):\tCaused an ERROR\njava.lang.RuntimeException: org.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 4\n\tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1533)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:895)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:918)\n\tat java.lang.Thread.run(Thread.java:680)\nCaused by: org.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 4\n\tat org.apache.cassandra.db.marshal.LongType.getString(LongType.java:69)\n\tat org.apache.cassandra.db.index.AbstractSimplePerColumnSecondaryIndex.insert(AbstractSimplePerColumnSecondaryIndex.java:121)\n\tat org.apache.cassandra.db.index.SecondaryIndexManager$PerColumnIndexUpdater.update(SecondaryIndexManager.java:623)\n\tat org.apache.cassandra.db.AtomicSortedColumns$Holder.addColumn(AtomicSortedColumns.java:313)\n\tat org.apache.cassandra.db.AtomicSortedColumns.addAllWithSizeDelta(AtomicSortedColumns.java:168)\n\tat org.apache.cassandra.db.Memtable.resolve(Memtable.java:253)\n\tat org.apache.cassandra.db.Memtable.put(Memtable.java:169)\n\tat org.apache.cassandra.db.ColumnFamilyStore.apply(ColumnFamilyStore.java:852)\n\tat org.apache.cassandra.db.Table.apply(Table.java:379)\n\tat org.apache.cassandra.db.Table.apply(Table.java:342)\n\tat org.apache.cassandra.db.RowMutation.apply(RowMutation.java:189)\n\tat org.apache.cassandra.service.StorageProxy$6.runMayThrow(StorageProxy.java:667)\n\tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1529)\n\t... 3 more\n\t\nTestcase: testIndexDeletions(org.apache.cassandra.db.ColumnFamilyStoreTest):\tCaused an ERROR\nA long is exactly 8 bytes: 4\norg.apache.cassandra.db.marshal.MarshalException: A long is exactly 8 bytes: 4\n\tat org.apache.cassandra.db.marshal.LongType.getString(LongType.java:69)\n\tat org.apache.cassandra.db.index.AbstractSimplePerColumnSecondaryIndex.insert(AbstractSimplePerColumnSecondaryIndex.java:121)\n\tat org.apache.cassandra.db.index.SecondaryIndexManager$PerColumnIndexUpdater.update(SecondaryIndexManager.java:623)\n\tat org.apache.cassandra.db.AtomicSortedColumns$Holder.addColumn(AtomicSortedColumns.java:313)\n\tat org.apache.cassandra.db.AtomicSortedColumns.addAllWithSizeDelta(AtomicSortedColumns.java:168)\n\tat org.apache.cassandra.db.Memtable.resolve(Memtable.java:253)\n\tat org.apache.cassandra.db.Memtable.put(Memtable.java:169)\n\tat org.apache.cassandra.db.ColumnFamilyStore.apply(ColumnFamilyStore.java:852)\n\tat org.apache.cassandra.db.Table.apply(Table.java:379)\n\tat org.apache.cassandra.db.Table.apply(Table.java:342)\n\tat org.apache.cassandra.db.RowMutation.apply(RowMutation.java:189)\n\tat org.apache.cassandra.db.ColumnFamilyStoreTest.testIndexDeletions(ColumnFamilyStoreTest.java:301)\n\n\nTestcase: testIndexUpdate(org.apache.cassandra.db.ColumnFamilyStoreTest):\tCaused an ERROR\nIndex: 0, Size: 0\njava.lang.IndexOutOfBoundsException: Index: 0, Size: 0\n\tat java.util.ArrayList.RangeCheck(ArrayList.java:547)\n\tat java.util.ArrayList.get(ArrayList.java:322)\n\tat org.apache.cassandra.db.ColumnFamilyStoreTest.testIndexUpdate(ColumnFamilyStoreTest.java:398)","from":"developer"},{"body":"My bad. That was due to a rebase typo that I had fixed in my CASSANDRA-5417 branch but not on that one. I've push the fix to the same github branch than above (https://github.com/pcmanus/cassandra/commits/5125-2).","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed, thanks","from":"developer"},{"body":"??Lastly, I'll note that the patches only add theses news indexes for non compact tables. We should generalize to compact tables too, but that would require a bit of generalization that I'd rather add in a second phase.??\n\nWith 2.1 and *compact* tables it is possible to CREATE INDEX on composite primary key columns, but queries returns no results for the tests below.\nAdding this comment for now, can open a new ticket if you prefer.\n\n{code:SQL}\nCREATE TABLE users2 (\n userID uuid,\n fname text,\n zip int,\n state text,\n PRIMARY KEY ((userID, fname))\n) WITH COMPACT STORAGE;\n\nCREATE INDEX ON users2 (userID);\nCREATE INDEX ON users2 (fname);\n\nINSERT INTO users2 (userID, fname, zip, state) VALUES (b3e3bc33-b237-4b55-9337-3d41de9a5649, 'John', 10007, 'NY');\n\n// the following queries returns 0 rows, instead of 1 expected\nSELECT * FROM users2 WHERE fname='John'; \nSELECT * FROM users2 WHERE userid=b3e3bc33-b237-4b55-9337-3d41de9a5649;\nSELECT * FROM users2 WHERE userid=b3e3bc33-b237-4b55-9337-3d41de9a5649 AND fname='John';\n\n// dropping 2ndary indexes restore normal behavior\n{code}","from":"developer"},{"body":"[~denis.angilella] Correct, the validation during index creation is broken. Do you mind creating a ticket indeed so we track the fix?","from":"developer"},{"body":"[~slebresne]: I created CASSANDRA-8156 to track the fix.","from":"developer"}],"created":"2013-01-07T18:33:56.000+0000","description":"Given\n\n{code}\nCREATE TABLE foo (\n a int,\n b int,\n c int,\n PRIMARY KEY (a, b)\n);\n{code}\n\nWe should support {{CREATE INDEX ON foo(b)}}.","issue_id":"12626383","key":"CASSANDRA-5125","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-04-04T16:33:46.000+0000","role":"fixed_distractor","summary":"Support indexes on composite column components (clustered columns)"} {"case_id":"12628677","cluster":"DISTRACTOR-CASSANDRA-5179","comments":[{"body":"Can you give us a debug log?","created":"2013-01-23T17:29:20.415+0000"},{"body":"It's written that it's fixed in 1.2.2\nDo you know if a bugfix corrected it ? ","created":"2013-02-04T16:02:58.819+0000"},{"body":"Fix-for is target fix version. The issue is not fixed until resolved.","created":"2013-02-04T16:23:50.352+0000"},{"body":"Ok thanks for this information …","created":"2013-02-04T16:25:47.095+0000"},{"body":"4881221363f984ab6610756cab38e1a016b79e15 (CASSANDRA-4761) broke it. Current pagination logic doesn't deal well with the fact that deleteHint is now being called from a response handler callback and can go through the same hint again and again and again until it's finally replaced with a tombstone. This becomes really visible in multi-dc setups, where it can take a while to complete the write, but I was able to trigger it with ccm as well (the same hint got sent up to 3 times).\n\nShould we reopen CASSANDRA-4761 or deal with it here?","created":"2013-03-12T04:14:39.282+0000"},{"body":"Let's fix it here.","created":"2013-03-12T04:33:17.895+0000"},{"body":"k. This is the problematic branch btw: https://github.com/apache/cassandra/blob/cassandra-1.2/src/java/org/apache/cassandra/db/HintedHandOffManager.java#L336","created":"2013-03-12T04:40:15.525+0000"},{"body":"https://github.com/iamaleksey/cassandra/compare/5179","created":"2013-03-12T05:41:09.530+0000"},{"body":"[~jbellis] patch looks good to me","created":"2013-03-18T15:20:37.533+0000"},{"body":"If we time out we should probably cease further delivery attempts, to avoid hammering a node that is behind.","created":"2013-03-18T15:44:16.616+0000"},{"body":"bq. If we time out we should probably cease further delivery attempts, to avoid hammering a node that is behind.\n\nProbably. But that's not what causing this particular issue - the patch only fixes that faulty pagination logic.\n\nI don't see an easy way to stop it on a timeout as we did in 1.1 now (yet), but that's a problem for another ticket anyway.","created":"2013-03-18T17:47:16.029+0000"},{"body":"We can't put that code in the empty catch block here?","created":"2013-03-18T18:17:25.735+0000"},{"body":"It'll be of no use - all the requests have been sent at that point. It's waiting for responses before forcing compaction.","created":"2013-03-18T18:20:07.200+0000"},{"body":"What if we made it wait per page, instead of all at once?\n\nSeems like that would be a good way to put a bound on how many callbacks we need to keep around too.","created":"2013-03-18T20:50:11.624+0000"},{"body":"bq. What if we made it wait per page, instead of all at once?\n\nThis goes against CASSANDRA-4761 somewhat, but I think it's a good compromise between sending hints one at a time and pouring everything out. Updated https://github.com/iamaleksey/cassandra/compare/5179","created":"2013-03-18T23:46:35.979+0000"},{"body":"You could use break-to-label instead of a bool in the get() loop, but shouldn't it just be a return instead of break?","created":"2013-03-19T01:08:19.133+0000"},{"body":"bq. You could use break-to-label instead of a bool in the get() loop, but shouldn't it just be a return instead of break?\n\nYes I could. But some requests after the timed-out one could actually have succeeded - with the ratelimiter delay and all, and since we sent them, we might as well wait for the replies before triggering compaction (hints are deleted in that callback). And with return instead of break 1) compaction wouldn't be triggered, event if it's the last page of many and 2) \"Finished hinted handoff ..\" message wouldn't be logged.","created":"2013-03-19T01:20:49.177+0000"},{"body":"I think the reasoning for return instead of break in the FD block was that, if we haven't finished sending hints then we probably don't want to force a major compaction that will rewrite a bunch of undelivered hints.","created":"2013-03-19T01:24:47.305+0000"},{"body":"Maybe. It was break in 1.1 though - https://github.com/apache/cassandra/blob/cassandra-1.1/src/java/org/apache/cassandra/db/HintedHandOffManager.java#L375","created":"2013-03-19T01:26:43.003+0000"},{"body":"True... and 1.0 had a \"if rowsReplayed > 0\" around it. Not sure why we removed that, except maybe that it would almost always be true.\n\nHow about we make the 1.0 check smarter and compact if we delivered over half the hints?","created":"2013-03-19T01:33:10.904+0000"},{"body":"bq. How about we make the 1.0 check smarter and compact if we delivered over half the hints?\n\nwfm","created":"2013-03-19T01:34:37.034+0000"},{"body":"Actually, it doesn't wfm. HHM is already complicated enough. I think returning completely is just fine there - as long as the log message mentions how many hints have been delivered.\n\nThere will be many more opportunities for compaction - when hints to another node are all sent, or after another attempt in <= 10 minutes.\n\nUpdated the branch (also did some *very* minor refactoring so that the paging logic would at least fit in a single screen, plus some even minorer changes that I just couldn't resist).","created":"2013-03-19T05:43:22.792+0000"},{"body":"+1","created":"2013-03-20T15:50:48.778+0000"},{"body":"Thanks, committed.","created":"2013-03-20T17:29:47.871+0000"}],"conversations":[{"body":"We have a small test environment with two datacenters (DC1 and DC2) running on Windows 7 laptops.\nBoth datacenters have one node. We use network topology strategy to replicate all data to both datacenters.\n\nWe started with empty db. \n1. Created a keyspace with strategy options [DC1:1, DC2:1]\n2. Added one row to a column family with CLI to DC1. Change was replicated to DC2.\n3. Disconnected network cable from DC2.\n4. Gossiper noticed, that other DC is dead.\n5. Added another row to DC1.\n6. Reconnected cable on DC2.\n7. DC1 started hinted handoff for DC2.\n8. Hinted handoff is finished with message: \"Finished hinted handoff of 1969 rows to endpoint \"\n\nWe repeated test with same results on Linux cluster with Cassandra 1.2.0. \n\nOn Cassandra 1.1.5 Linux cluster, only one row was sent to endpoint. \"Finished hinted handoff of 1 rows to endpoint \"","from":"reporter","subject":"Hinted handoff sends over 1000 rows for one column change"},{"body":"Can you give us a debug log?","from":"developer"},{"body":"It's written that it's fixed in 1.2.2\nDo you know if a bugfix corrected it ? ","from":"developer"},{"body":"Fix-for is target fix version. The issue is not fixed until resolved.","from":"developer"},{"body":"Ok thanks for this information …","from":"developer"},{"body":"4881221363f984ab6610756cab38e1a016b79e15 (CASSANDRA-4761) broke it. Current pagination logic doesn't deal well with the fact that deleteHint is now being called from a response handler callback and can go through the same hint again and again and again until it's finally replaced with a tombstone. This becomes really visible in multi-dc setups, where it can take a while to complete the write, but I was able to trigger it with ccm as well (the same hint got sent up to 3 times).\n\nShould we reopen CASSANDRA-4761 or deal with it here?","from":"developer"},{"body":"Let's fix it here.","from":"developer"},{"body":"k. This is the problematic branch btw: https://github.com/apache/cassandra/blob/cassandra-1.2/src/java/org/apache/cassandra/db/HintedHandOffManager.java#L336","from":"developer"},{"body":"https://github.com/iamaleksey/cassandra/compare/5179","from":"developer"},{"body":"[~jbellis] patch looks good to me","from":"developer"},{"body":"If we time out we should probably cease further delivery attempts, to avoid hammering a node that is behind.","from":"developer"},{"body":"bq. If we time out we should probably cease further delivery attempts, to avoid hammering a node that is behind.\n\nProbably. But that's not what causing this particular issue - the patch only fixes that faulty pagination logic.\n\nI don't see an easy way to stop it on a timeout as we did in 1.1 now (yet), but that's a problem for another ticket anyway.","from":"developer"},{"body":"We can't put that code in the empty catch block here?","from":"developer"},{"body":"It'll be of no use - all the requests have been sent at that point. It's waiting for responses before forcing compaction.","from":"developer"},{"body":"What if we made it wait per page, instead of all at once?\n\nSeems like that would be a good way to put a bound on how many callbacks we need to keep around too.","from":"developer"},{"body":"bq. What if we made it wait per page, instead of all at once?\n\nThis goes against CASSANDRA-4761 somewhat, but I think it's a good compromise between sending hints one at a time and pouring everything out. Updated https://github.com/iamaleksey/cassandra/compare/5179","from":"developer"},{"body":"You could use break-to-label instead of a bool in the get() loop, but shouldn't it just be a return instead of break?","from":"developer"},{"body":"bq. You could use break-to-label instead of a bool in the get() loop, but shouldn't it just be a return instead of break?\n\nYes I could. But some requests after the timed-out one could actually have succeeded - with the ratelimiter delay and all, and since we sent them, we might as well wait for the replies before triggering compaction (hints are deleted in that callback). And with return instead of break 1) compaction wouldn't be triggered, event if it's the last page of many and 2) \"Finished hinted handoff ..\" message wouldn't be logged.","from":"developer"},{"body":"I think the reasoning for return instead of break in the FD block was that, if we haven't finished sending hints then we probably don't want to force a major compaction that will rewrite a bunch of undelivered hints.","from":"developer"},{"body":"Maybe. It was break in 1.1 though - https://github.com/apache/cassandra/blob/cassandra-1.1/src/java/org/apache/cassandra/db/HintedHandOffManager.java#L375","from":"developer"},{"body":"True... and 1.0 had a \"if rowsReplayed > 0\" around it. Not sure why we removed that, except maybe that it would almost always be true.\n\nHow about we make the 1.0 check smarter and compact if we delivered over half the hints?","from":"developer"},{"body":"bq. How about we make the 1.0 check smarter and compact if we delivered over half the hints?\n\nwfm","from":"developer"},{"body":"Actually, it doesn't wfm. HHM is already complicated enough. I think returning completely is just fine there - as long as the log message mentions how many hints have been delivered.\n\nThere will be many more opportunities for compaction - when hints to another node are all sent, or after another attempt in <= 10 minutes.\n\nUpdated the branch (also did some *very* minor refactoring so that the paging logic would at least fit in a single screen, plus some even minorer changes that I just couldn't resist).","from":"developer"},{"body":"+1","from":"developer"},{"body":"Thanks, committed.","from":"developer"}],"created":"2013-01-22T08:15:55.000+0000","description":"We have a small test environment with two datacenters (DC1 and DC2) running on Windows 7 laptops.\nBoth datacenters have one node. We use network topology strategy to replicate all data to both datacenters.\n\nWe started with empty db. \n1. Created a keyspace with strategy options [DC1:1, DC2:1]\n2. Added one row to a column family with CLI to DC1. Change was replicated to DC2.\n3. Disconnected network cable from DC2.\n4. Gossiper noticed, that other DC is dead.\n5. Added another row to DC1.\n6. Reconnected cable on DC2.\n7. DC1 started hinted handoff for DC2.\n8. Hinted handoff is finished with message: \"Finished hinted handoff of 1969 rows to endpoint \"\n\nWe repeated test with same results on Linux cluster with Cassandra 1.2.0. \n\nOn Cassandra 1.1.5 Linux cluster, only one row was sent to endpoint. \"Finished hinted handoff of 1 rows to endpoint \"","issue_id":"12628677","key":"CASSANDRA-5179","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-03-20T17:29:47.000+0000","role":"fixed_distractor","summary":"Hinted handoff sends over 1000 rows for one column change"} {"case_id":"12629942","cluster":"DISTRACTOR-CASSANDRA-5201","comments":[{"body":"The current stable hadoop version is still the 1.0.x line according to http://hadoop.apache.org/releases.html#Download As long as the 1.0.x support isn't adversely affected, I would think a patch to provide support for all hadoop versions would be welcome.","created":"2013-01-30T17:29:56.738+0000"},{"body":"I think the last time I looked into this I couldn't find a way to support both versions, so we're either stuck with 0.21+ being broken, or breaking it for everyone that's on a lower version already.","created":"2013-01-31T00:47:56.692+0000"},{"body":"Do you have a patch?","created":"2013-02-02T19:42:26.806+0000"},{"body":"what if the hadoop directory was split out into two separate sub-projects which produced two separate jars, thus a jar for old hadoop support and one for new hadoop. then folks could choose which jar to add to the cassandra classpath?","created":"2013-03-05T02:11:31.246+0000"},{"body":"That's not a bad idea, but I'm not sure how it'd work entirely, care to formulate a patch? :)","created":"2013-03-05T03:43:42.428+0000"},{"body":"Actually, i only see 0.20.*, 1.0.* (both of which have JobContext as a class) on the central maven repository, and on apache's download page... \n\nDoes the 0.21+ jars still exist, (supported)?\n\n","created":"2013-03-05T06:40:45.844+0000"},{"body":"We are using the newer 0.23.x & are facing this integration issue. Any workaround?","created":"2013-03-07T20:23:15.449+0000"},{"body":"ah ok, here it is:\n\ngroupid: org.apache.hadoop\nartifactid: hadoop-mapreduce-client-core\n","created":"2013-03-08T05:04:42.098+0000"},{"body":"initial patch against 1.2\n - pulls hadoop code into src/hadoop/hadoop-legacy\n - replicates that into src/hadoop/hadoop-new with changes for api changes.\n\npulls new hadoop dependencies and builds separate jars for each.\n\nwhich jar you run against is still manual... command line switch?\n\nposted version a so i didn't lose stuff.\n\nThis uses 0.23.6 versions of jars, which i assume is compatible with the 2.0a versions of the api.","created":"2013-03-08T05:53:32.574+0000"},{"body":"I wonder if it would be preferable to have all of the code in tree in the same jar still, but have different package paths, like org.apache.cassandra.hadoop (for backwards compatibility) and org.apache.cassandra.hadoop2. Would that work?","created":"2013-03-08T10:35:37.605+0000"},{"body":"No workaround that I'm aware of. We don't actually have this problem; I just became aware of it by\nchange while looking into a problem report, and filed an issue.\n\nBrian\n\n\n\n","created":"2013-03-08T12:06:12.430+0000"},{"body":"{quote} wonder if it would be preferable to have all of the code in tree in the same jar still, but have different package paths, like org.apache.cassandra.hadoop (for backwards compatibility) and org.apache.cassandra.hadoop2. Would that work?{quote}\n\nprobably could be done. here's some 'issues' with that approach.\n\n1) external code won't work against both versions with the same code base. Obviously the modifications are insignificant, (packages) but still.\n2) Makes the build.xml file marginally more complicated as you have to use elements in the build\n3) Building in IDEs is complicated as you need to add resource excludes, since these packages need to be built as a separate project (at least one of them does).\n4) If you have two IDE projects targetting the same classes dir, cleaning one project cleans the other.\n\ni could do it either way... whatever people think.\n","created":"2013-03-09T01:12:47.430+0000"},{"body":"What about simply putting the hadoop2 package into a github project?\nit would become available for others to use, and c* can switch to it when they feel ready to drop support for hadoop-0.20\n\notherwise i'm in favour of separate jar files (apache-cassandra-hadoop-legacy-XXX.jar and apache-cassandra-hadoop-XXX.jar). c* already bundles too much into the one jar file IMHO.","created":"2013-05-19T07:49:08.363+0000"},{"body":"{quote}What about simply putting the hadoop2 package into a github project?{quote}\n\nDone @ https://github.com/michaelsembwever/cassandra-hadoop\n (i refactor the new package to hadoop1 instead of hadoop2, to better match the hadoop version we are actually supporting).","created":"2013-05-19T08:37:35.027+0000"},{"body":"I've updated the github project so to be a [patch|https://github.com/michaelsembwever/cassandra-hadoop/commit/6d7555ea205354a606907e40c16db35072004594] off the InputFormat and OutputFormat classes as found in cassandra-1.2.10\nIt works against hadoop-0.22.0","created":"2013-10-11T12:27:55.829+0000"},{"body":"How common is hadoop2 usage now? Can we drop hadoop1 for 2.1? /cc [~bcoverston]","created":"2013-10-13T16:55:27.169+0000"},{"body":"Re Hadoop 1.x vs. Hadoop 2.x - most companies I see are using a 1.x version of Hadoop. Usage of Hadoop 2.x is ramping up fast, but I'm guessing it will be at least another year (probably more) before you'd want to start thinking about phasing out support for Hadoop 1.x","created":"2013-10-25T00:55:55.406+0000"},{"body":"Hadoop-2 only just came out of alpha/beta with hadoop-2.2.0","created":"2013-10-25T07:07:11.221+0000"},{"body":"Why dropping support for old hadoop versions while you can just use a different package name (e.g org.apache.cassandra.hadoop2.)? Imo, the code in here:\nhttps://github.com/michaelsembwever/cassandra-hadoop/ should just be merged with mainstream implementation, so both 1.x and 2.x users can be happy.\n\nOr you can indeed make it a lot more complex by merging the 1.x and 2.x compatible implementations and deciding which one to use based on metadata obtained by polling the cluster, but that is a bit too sophisticated. For most people the first option, where they can choose which package to use, would be sufficient.","created":"2013-10-25T08:34:33.842+0000"},{"body":"[~mck] How could be used your patch to make work Cassandra + Hadoop 2.2 + Pig? I have your library compiled but cannot figure how to apply to Cassandra or Pig to make it work, because I still hit the exception from Cassandra:\njava.lang.IncompatibleClassChangeError: Found interface org.apache.hadoop.mapreduce.JobContext, but class was expected\n\tat org.apache.cassandra.hadoop.AbstractColumnFamilyInputFormat.getSplits(AbstractColumnFamilyInputFormat.java:115)","created":"2013-10-29T16:18:45.284+0000"},{"body":"[~claudio.romo.otto] \nIt is a bit complex but basically you need to use the org.apache.cassandra.hadoop2.ColumnFamilyInputFormat stuff from https://github.com/michaelsembwever/cassandra-hadoop/\nAs far as hadoop goes, I ran succesfully with a cdh 4.4.1 cloudera hadoop lib, and with a cassandra 1.2.10. ","created":"2013-10-29T16:32:28.928+0000"},{"body":"So the correct way is to update cassandra 1.2.10 using org.apache.cassandra.hadoop2.ColumnFamilyInputFormat, right?","created":"2013-10-29T19:36:39.025+0000"},{"body":"Poking around at other projects this generally gets solved in one of two ways: Ship two versions of their Hadoop integrations (one compiled for the old, and one compiled for the new), or use a little reflection to make things work across the board.\n\nI'm attaching a patch that uses the hadoopCompat subproject of elephantbird. This will allow us to compile a single binary and run with the new and old context objects.\n\nI've tested this patch with HDP 2.0, and Apache Hadoop 1.0.4 and it works fine with both (including Hive in DSE). With Pig I needed to compile our (optional) pig dependency with:\n\nbq. ant clean jar-withouthadoop -Dhadoopversion=23\n\nOnly really needed if you're using one of the current versions of thrift with the new JobContext.\n","created":"2013-12-03T22:58:16.682+0000"},{"body":"These changes also depend on CASSANDRA-6309 for anything to work.","created":"2013-12-03T23:36:16.401+0000"},{"body":"WDYT [~dbrosius]?","created":"2013-12-04T22:33:38.503+0000"},{"body":"[~mck]? [~jeromatron]?","created":"2013-12-17T03:31:30.545+0000"},{"body":"Seems reasonable if they're keeping their code up to date with all of the releases. The code appears lightweight to make to use it as well. I've also tried to reach out to [~dvryaboy] via twitter to see if he has any feedback about the approach of having elephant bird as a dependency, e.g. any hidden costs or pitfalls being more familiar with the project.","created":"2014-01-04T23:29:32.644+0000"},{"body":"The EB HadoopCompat is what we use in production at Twitter, and plan to keep up to date in the foreseeable future. Glad you are finding it useful.\n\nMaybe you can send that Reporter implementation as a pull request for hadoop-compat? Seems generally applicable.\n\nThanks to this ping, I looked around and noticed that our Parquet project uses a slightly different version of the same code -- we'll take a look and merge things. Shouldn't change anything significantly for this patch.\n\nAlso note that Tom White has a handy findbugs plugin to look for Hadoop incompatibility problems: https://github.com/tomwhite/hadoop-incompatibility-findbugs-detector \n\nHere's how you'd use it: https://github.com/Parquet/parquet-mr/pull/77/files","created":"2014-01-05T22:10:03.719+0000"},{"body":"Thanks [~dvryaboy]!\n\nJust for completeness the twitter thread is https://twitter.com/jeromatron/status/419607697588510721\n\n[~bcoverston] [~jbellis] what do you think?\n\nAs for me, it sounds like the EB (or if it makes more sense Parquet) dependency makes sense. The hadoop incompatibility findbugs detector also sounds great to include to catch anything before it is committed. So I'm +1 on this approach.","created":"2014-01-06T13:30:50.371+0000"},{"body":"I'll submit the reporter impl upstream. Thanks [~dvryaboy]!\n","created":"2014-01-06T18:02:48.944+0000"},{"body":"Are we good here [~brandon.williams]?","created":"2014-02-10T19:19:38.065+0000"},{"body":"Committed to 2.0. [~bcoverston] can you rebase a patch for trunk?","created":"2014-02-11T15:39:33.376+0000"},{"body":"Attached trunk-rebased patch.","created":"2014-02-11T21:26:30.179+0000"},{"body":"Committed.","created":"2014-02-11T22:10:59.249+0000"},{"body":"Hi [~brandon.williams]\n\nI'm using this patch and I keep having compatibility problems while trying to write something with pig using CqlStorage.\nI'm using Hadoop 2.2 and Cassandra 2.0 branch.\n\nThis is the full stack trace:\n\n{noformat}\njava.lang.IncompatibleClassChangeError: Found interface org.apache.hadoop.mapreduce.TaskAttemptContext, but class was expected\n at org.apache.cassandra.hadoop.Progressable.progress(Progressable.java:45)\n at org.apache.cassandra.hadoop.cql3.CqlRecordWriter.write(CqlRecordWriter.java:183) \n at org.apache.cassandra.hadoop.cql3.CqlRecordWriter.write(CqlRecordWriter.java:63) \n at org.apache.cassandra.hadoop.pig.CqlStorage.sendCqlQuery(CqlStorage.java:440) \n at org.apache.cassandra.hadoop.pig.CqlStorage.cqlQueryFromTuple(CqlStorage.java:414) \n at org.apache.cassandra.hadoop.pig.CqlStorage.putNext(CqlStorage.java:362) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:139) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:98) \n at org.apache.hadoop.mapred.ReduceTask$NewTrackingRecordWriter.write(ReduceTask.java:576) \n at org.apache.hadoop.mapreduce.task.TaskInputOutputContextImpl.write(TaskInputOutputContextImpl.java:89) \n at org.apache.hadoop.mapreduce.lib.reduce.WrappedReducer$Context.write(WrappedReducer.java:105) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigGenericMapReduce$Reduce.runPipeline(PigGenericMapReduce.java:467) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigGenericMapReduce$Reduce.processOnePackageOutput(PigGenericMapReduce.java:432) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigGenericMapReduce$Reduce.reduce(PigGenericMapReduce.java:412) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigGenericMapReduce$Reduce.reduce(PigGenericMapReduce.java:256) \n at org.apache.hadoop.mapreduce.Reducer.run(Reducer.java:171) \n at org.apache.hadoop.mapred.ReduceTask.runNewReducer(ReduceTask.java:645) \n at org.apache.hadoop.mapred.ReduceTask.run(ReduceTask.java:405) \n at org.apache.hadoop.mapred.YarnChild$2.run(YarnChild.java:162) \n at java.security.AccessController.doPrivileged(Native Method) \n at javax.security.auth.Subject.doAs(Subject.java:415) \n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1491) \n at org.apache.hadoop.mapred.YarnChild.main(YarnChild.java:157)\n{noformat}\n\n","created":"2014-02-27T13:37:25.780+0000"},{"body":"[~moliware], attaching a patch that will fix the issue with progressable when used on the jobcontext object. I've tested this with the latest HDP release, and it works with pig. [~brandon.williams] take a look.","created":"2014-02-28T16:57:16.733+0000"},{"body":"Thanks [~bcoverston] I will test the patch probably on monday and will confirm you but it looks good. Thanks!","created":"2014-02-28T17:51:12.151+0000"},{"body":"[~bcoverston] Tested and It worked! Thanks!","created":"2014-03-03T08:24:45.654+0000"},{"body":"Does this patch include deployment with maven classifiers, so that hadoop1 or hadoop2 are referenced dependencies?","created":"2014-03-03T08:44:00.341+0000"},{"body":"I also was able to make this patch work with HDP 2.0. Nevertheless before this ticket is closed I would like to make 2 suggestions:\n1. Use maven classifiers for deployment. If this is used in a hadoop2 project you don't want hadoop1 dependencies and vice versa. This approach is used by Parquet, Avro and EB. For detailed discussion please read [here|https://github.com/Parquet/parquet-mr/pull/32].\n2. Change {{hadoop-core}} to {{hadoop-client}} dependency. {{hadoop-client}} is the preferred dependency. Read [here|http://www.cloudera.com/content/cloudera-content/cloudera-docs/CDH4/latest/CDH-Version-and-Packaging-Information/cdhvd_topic_8_1.html] and the link above.","created":"2014-03-03T20:20:25.130+0000"},{"body":"[~hkropp] I agree that we need to do the second change, but I don't see a need to add deployment classifiers. Because we're using Hadoop-Compat (from EB) a single binary will work for both.\n\nThere's a niggle right now with Progressable that I'm working on, but a single binary will work for new and legacy deployments.","created":"2014-03-03T23:33:59.461+0000"},{"body":"I might be missing an important point, but Hadoop-Compat is nothing other than ContextUtil of hadoop2? It uses reflection to test what's on the classpath to decide how or better what to return. I made it work with just using ContextUtil before I found this patch here.\n\nTherefor the binary works fine for both (btw I would not consider it legacy!), if you manage the dependencies on your own.\n\nBut if you would start a new project you would import Cassandra for example like this:\n{code}\ndependencies {\n ...\n compile 'org.apache.cassandra:cassandra:2.0.6'\n ...\n}\n{code}\n\nWhat will happen now is, that this will load hadoop1 dependencies into my hadoop2 project for example, or not? To avoid this maven classifiers could be used, to call explicitly for hadoop2 dependencies:\n\n{code}\ndependencies {\n ....\n compile 'org.apache.cassandra:cassandra:2.0.6:hadoop2'\n ...\n}\n{code}","created":"2014-03-04T07:54:50.532+0000"},{"body":"Take a look at the code, it works for Hadoop1 and Hadoop2 without recompliation, and without shipping two sets of dependencies. Basically it detects the current version of Hadoop that you're running and dynamically determines which TaskAttemptContext to use, the Class, or the Interface. There's no need to use deployment classifiers to solve this particular problem.","created":"2014-03-04T16:36:34.524+0000"},{"body":"Attaching a patch that brings in hadoop-compat, and adds the progressable wrapper to it. I'm going to submit this upstream to the elephant bird project, so we should be able to remove this code and add the dependency in the future.","created":"2014-03-04T21:31:32.114+0000"},{"body":"Updated progressable-wrapper.patch to conform to current namespaces.","created":"2014-03-04T21:46:26.737+0000"},{"body":"I removed the hadoop-compat dependency, as maven couldn't resolve the dependencies, reverted to the old version. This won't stop our input/output formats from being compatible with past and future releases.","created":"2014-03-04T22:33:53.343+0000"},{"body":"Patch for 2.1 branch","created":"2014-03-05T18:55:39.003+0000"},{"body":"I had the same maven resolution problem with hadoop-compat. Committed.","created":"2014-03-05T19:00:44.503+0000"}],"conversations":[{"body":"Using Hadoop 0.22.0 with Cassandra results in the stack trace below.\nIt appears that version 0.21+ changed org.apache.hadoop.mapreduce.JobContext\nfrom a class to an interface.\n\n\nException in thread \"main\" java.lang.IncompatibleClassChangeError: Found interface org.apache.hadoop.mapreduce.JobContext, but class was expected\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.getSplits(ColumnFamilyInputFormat.java:103)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.writeNewSplits(JobSubmitter.java:445)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.writeSplits(JobSubmitter.java:462)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.submitJobInternal(JobSubmitter.java:357)\n\tat org.apache.hadoop.mapreduce.Job$2.run(Job.java:1045)\n\tat org.apache.hadoop.mapreduce.Job$2.run(Job.java:1042)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:415)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1153)\n\tat org.apache.hadoop.mapreduce.Job.submit(Job.java:1042)\n\tat org.apache.hadoop.mapreduce.Job.waitForCompletion(Job.java:1062)\n\tat MyHadoopApp.run(MyHadoopApp.java:163)\n\tat org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:69)\n\tat MyHadoopApp.main(MyHadoopApp.java:82)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:601)\n\tat org.apache.hadoop.util.RunJar.main(RunJar.java:192)\n","from":"reporter","subject":"Cassandra/Hadoop does not support current Hadoop releases"},{"body":"The current stable hadoop version is still the 1.0.x line according to http://hadoop.apache.org/releases.html#Download As long as the 1.0.x support isn't adversely affected, I would think a patch to provide support for all hadoop versions would be welcome.","from":"developer"},{"body":"I think the last time I looked into this I couldn't find a way to support both versions, so we're either stuck with 0.21+ being broken, or breaking it for everyone that's on a lower version already.","from":"developer"},{"body":"Do you have a patch?","from":"developer"},{"body":"what if the hadoop directory was split out into two separate sub-projects which produced two separate jars, thus a jar for old hadoop support and one for new hadoop. then folks could choose which jar to add to the cassandra classpath?","from":"developer"},{"body":"That's not a bad idea, but I'm not sure how it'd work entirely, care to formulate a patch? :)","from":"developer"},{"body":"Actually, i only see 0.20.*, 1.0.* (both of which have JobContext as a class) on the central maven repository, and on apache's download page... \n\nDoes the 0.21+ jars still exist, (supported)?\n\n","from":"developer"},{"body":"We are using the newer 0.23.x & are facing this integration issue. Any workaround?","from":"developer"},{"body":"ah ok, here it is:\n\ngroupid: org.apache.hadoop\nartifactid: hadoop-mapreduce-client-core\n","from":"developer"},{"body":"initial patch against 1.2\n - pulls hadoop code into src/hadoop/hadoop-legacy\n - replicates that into src/hadoop/hadoop-new with changes for api changes.\n\npulls new hadoop dependencies and builds separate jars for each.\n\nwhich jar you run against is still manual... command line switch?\n\nposted version a so i didn't lose stuff.\n\nThis uses 0.23.6 versions of jars, which i assume is compatible with the 2.0a versions of the api.","from":"developer"},{"body":"I wonder if it would be preferable to have all of the code in tree in the same jar still, but have different package paths, like org.apache.cassandra.hadoop (for backwards compatibility) and org.apache.cassandra.hadoop2. Would that work?","from":"developer"},{"body":"No workaround that I'm aware of. We don't actually have this problem; I just became aware of it by\nchange while looking into a problem report, and filed an issue.\n\nBrian\n\n\n\n","from":"developer"},{"body":"{quote} wonder if it would be preferable to have all of the code in tree in the same jar still, but have different package paths, like org.apache.cassandra.hadoop (for backwards compatibility) and org.apache.cassandra.hadoop2. Would that work?{quote}\n\nprobably could be done. here's some 'issues' with that approach.\n\n1) external code won't work against both versions with the same code base. Obviously the modifications are insignificant, (packages) but still.\n2) Makes the build.xml file marginally more complicated as you have to use elements in the build\n3) Building in IDEs is complicated as you need to add resource excludes, since these packages need to be built as a separate project (at least one of them does).\n4) If you have two IDE projects targetting the same classes dir, cleaning one project cleans the other.\n\ni could do it either way... whatever people think.\n","from":"developer"},{"body":"What about simply putting the hadoop2 package into a github project?\nit would become available for others to use, and c* can switch to it when they feel ready to drop support for hadoop-0.20\n\notherwise i'm in favour of separate jar files (apache-cassandra-hadoop-legacy-XXX.jar and apache-cassandra-hadoop-XXX.jar). c* already bundles too much into the one jar file IMHO.","from":"developer"},{"body":"{quote}What about simply putting the hadoop2 package into a github project?{quote}\n\nDone @ https://github.com/michaelsembwever/cassandra-hadoop\n (i refactor the new package to hadoop1 instead of hadoop2, to better match the hadoop version we are actually supporting).","from":"developer"},{"body":"I've updated the github project so to be a [patch|https://github.com/michaelsembwever/cassandra-hadoop/commit/6d7555ea205354a606907e40c16db35072004594] off the InputFormat and OutputFormat classes as found in cassandra-1.2.10\nIt works against hadoop-0.22.0","from":"developer"},{"body":"How common is hadoop2 usage now? Can we drop hadoop1 for 2.1? /cc [~bcoverston]","from":"developer"},{"body":"Re Hadoop 1.x vs. Hadoop 2.x - most companies I see are using a 1.x version of Hadoop. Usage of Hadoop 2.x is ramping up fast, but I'm guessing it will be at least another year (probably more) before you'd want to start thinking about phasing out support for Hadoop 1.x","from":"developer"},{"body":"Hadoop-2 only just came out of alpha/beta with hadoop-2.2.0","from":"developer"},{"body":"Why dropping support for old hadoop versions while you can just use a different package name (e.g org.apache.cassandra.hadoop2.)? Imo, the code in here:\nhttps://github.com/michaelsembwever/cassandra-hadoop/ should just be merged with mainstream implementation, so both 1.x and 2.x users can be happy.\n\nOr you can indeed make it a lot more complex by merging the 1.x and 2.x compatible implementations and deciding which one to use based on metadata obtained by polling the cluster, but that is a bit too sophisticated. For most people the first option, where they can choose which package to use, would be sufficient.","from":"developer"},{"body":"[~mck] How could be used your patch to make work Cassandra + Hadoop 2.2 + Pig? I have your library compiled but cannot figure how to apply to Cassandra or Pig to make it work, because I still hit the exception from Cassandra:\njava.lang.IncompatibleClassChangeError: Found interface org.apache.hadoop.mapreduce.JobContext, but class was expected\n\tat org.apache.cassandra.hadoop.AbstractColumnFamilyInputFormat.getSplits(AbstractColumnFamilyInputFormat.java:115)","from":"developer"},{"body":"[~claudio.romo.otto] \nIt is a bit complex but basically you need to use the org.apache.cassandra.hadoop2.ColumnFamilyInputFormat stuff from https://github.com/michaelsembwever/cassandra-hadoop/\nAs far as hadoop goes, I ran succesfully with a cdh 4.4.1 cloudera hadoop lib, and with a cassandra 1.2.10. ","from":"developer"},{"body":"So the correct way is to update cassandra 1.2.10 using org.apache.cassandra.hadoop2.ColumnFamilyInputFormat, right?","from":"developer"},{"body":"Poking around at other projects this generally gets solved in one of two ways: Ship two versions of their Hadoop integrations (one compiled for the old, and one compiled for the new), or use a little reflection to make things work across the board.\n\nI'm attaching a patch that uses the hadoopCompat subproject of elephantbird. This will allow us to compile a single binary and run with the new and old context objects.\n\nI've tested this patch with HDP 2.0, and Apache Hadoop 1.0.4 and it works fine with both (including Hive in DSE). With Pig I needed to compile our (optional) pig dependency with:\n\nbq. ant clean jar-withouthadoop -Dhadoopversion=23\n\nOnly really needed if you're using one of the current versions of thrift with the new JobContext.\n","from":"developer"},{"body":"These changes also depend on CASSANDRA-6309 for anything to work.","from":"developer"},{"body":"WDYT [~dbrosius]?","from":"developer"},{"body":"[~mck]? [~jeromatron]?","from":"developer"},{"body":"Seems reasonable if they're keeping their code up to date with all of the releases. The code appears lightweight to make to use it as well. I've also tried to reach out to [~dvryaboy] via twitter to see if he has any feedback about the approach of having elephant bird as a dependency, e.g. any hidden costs or pitfalls being more familiar with the project.","from":"developer"},{"body":"The EB HadoopCompat is what we use in production at Twitter, and plan to keep up to date in the foreseeable future. Glad you are finding it useful.\n\nMaybe you can send that Reporter implementation as a pull request for hadoop-compat? Seems generally applicable.\n\nThanks to this ping, I looked around and noticed that our Parquet project uses a slightly different version of the same code -- we'll take a look and merge things. Shouldn't change anything significantly for this patch.\n\nAlso note that Tom White has a handy findbugs plugin to look for Hadoop incompatibility problems: https://github.com/tomwhite/hadoop-incompatibility-findbugs-detector \n\nHere's how you'd use it: https://github.com/Parquet/parquet-mr/pull/77/files","from":"developer"},{"body":"Thanks [~dvryaboy]!\n\nJust for completeness the twitter thread is https://twitter.com/jeromatron/status/419607697588510721\n\n[~bcoverston] [~jbellis] what do you think?\n\nAs for me, it sounds like the EB (or if it makes more sense Parquet) dependency makes sense. The hadoop incompatibility findbugs detector also sounds great to include to catch anything before it is committed. So I'm +1 on this approach.","from":"developer"},{"body":"I'll submit the reporter impl upstream. Thanks [~dvryaboy]!\n","from":"developer"},{"body":"Are we good here [~brandon.williams]?","from":"developer"},{"body":"Committed to 2.0. [~bcoverston] can you rebase a patch for trunk?","from":"developer"},{"body":"Attached trunk-rebased patch.","from":"developer"},{"body":"Committed.","from":"developer"},{"body":"Hi [~brandon.williams]\n\nI'm using this patch and I keep having compatibility problems while trying to write something with pig using CqlStorage.\nI'm using Hadoop 2.2 and Cassandra 2.0 branch.\n\nThis is the full stack trace:\n\n{noformat}\njava.lang.IncompatibleClassChangeError: Found interface org.apache.hadoop.mapreduce.TaskAttemptContext, but class was expected\n at org.apache.cassandra.hadoop.Progressable.progress(Progressable.java:45)\n at org.apache.cassandra.hadoop.cql3.CqlRecordWriter.write(CqlRecordWriter.java:183) \n at org.apache.cassandra.hadoop.cql3.CqlRecordWriter.write(CqlRecordWriter.java:63) \n at org.apache.cassandra.hadoop.pig.CqlStorage.sendCqlQuery(CqlStorage.java:440) \n at org.apache.cassandra.hadoop.pig.CqlStorage.cqlQueryFromTuple(CqlStorage.java:414) \n at org.apache.cassandra.hadoop.pig.CqlStorage.putNext(CqlStorage.java:362) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:139) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigOutputFormat$PigRecordWriter.write(PigOutputFormat.java:98) \n at org.apache.hadoop.mapred.ReduceTask$NewTrackingRecordWriter.write(ReduceTask.java:576) \n at org.apache.hadoop.mapreduce.task.TaskInputOutputContextImpl.write(TaskInputOutputContextImpl.java:89) \n at org.apache.hadoop.mapreduce.lib.reduce.WrappedReducer$Context.write(WrappedReducer.java:105) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigGenericMapReduce$Reduce.runPipeline(PigGenericMapReduce.java:467) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigGenericMapReduce$Reduce.processOnePackageOutput(PigGenericMapReduce.java:432) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigGenericMapReduce$Reduce.reduce(PigGenericMapReduce.java:412) \n at org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigGenericMapReduce$Reduce.reduce(PigGenericMapReduce.java:256) \n at org.apache.hadoop.mapreduce.Reducer.run(Reducer.java:171) \n at org.apache.hadoop.mapred.ReduceTask.runNewReducer(ReduceTask.java:645) \n at org.apache.hadoop.mapred.ReduceTask.run(ReduceTask.java:405) \n at org.apache.hadoop.mapred.YarnChild$2.run(YarnChild.java:162) \n at java.security.AccessController.doPrivileged(Native Method) \n at javax.security.auth.Subject.doAs(Subject.java:415) \n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1491) \n at org.apache.hadoop.mapred.YarnChild.main(YarnChild.java:157)\n{noformat}\n\n","from":"developer"},{"body":"[~moliware], attaching a patch that will fix the issue with progressable when used on the jobcontext object. I've tested this with the latest HDP release, and it works with pig. [~brandon.williams] take a look.","from":"developer"},{"body":"Thanks [~bcoverston] I will test the patch probably on monday and will confirm you but it looks good. Thanks!","from":"developer"},{"body":"[~bcoverston] Tested and It worked! Thanks!","from":"developer"},{"body":"Does this patch include deployment with maven classifiers, so that hadoop1 or hadoop2 are referenced dependencies?","from":"developer"},{"body":"I also was able to make this patch work with HDP 2.0. Nevertheless before this ticket is closed I would like to make 2 suggestions:\n1. Use maven classifiers for deployment. If this is used in a hadoop2 project you don't want hadoop1 dependencies and vice versa. This approach is used by Parquet, Avro and EB. For detailed discussion please read [here|https://github.com/Parquet/parquet-mr/pull/32].\n2. Change {{hadoop-core}} to {{hadoop-client}} dependency. {{hadoop-client}} is the preferred dependency. Read [here|http://www.cloudera.com/content/cloudera-content/cloudera-docs/CDH4/latest/CDH-Version-and-Packaging-Information/cdhvd_topic_8_1.html] and the link above.","from":"developer"},{"body":"[~hkropp] I agree that we need to do the second change, but I don't see a need to add deployment classifiers. Because we're using Hadoop-Compat (from EB) a single binary will work for both.\n\nThere's a niggle right now with Progressable that I'm working on, but a single binary will work for new and legacy deployments.","from":"developer"},{"body":"I might be missing an important point, but Hadoop-Compat is nothing other than ContextUtil of hadoop2? It uses reflection to test what's on the classpath to decide how or better what to return. I made it work with just using ContextUtil before I found this patch here.\n\nTherefor the binary works fine for both (btw I would not consider it legacy!), if you manage the dependencies on your own.\n\nBut if you would start a new project you would import Cassandra for example like this:\n{code}\ndependencies {\n ...\n compile 'org.apache.cassandra:cassandra:2.0.6'\n ...\n}\n{code}\n\nWhat will happen now is, that this will load hadoop1 dependencies into my hadoop2 project for example, or not? To avoid this maven classifiers could be used, to call explicitly for hadoop2 dependencies:\n\n{code}\ndependencies {\n ....\n compile 'org.apache.cassandra:cassandra:2.0.6:hadoop2'\n ...\n}\n{code}","from":"developer"},{"body":"Take a look at the code, it works for Hadoop1 and Hadoop2 without recompliation, and without shipping two sets of dependencies. Basically it detects the current version of Hadoop that you're running and dynamically determines which TaskAttemptContext to use, the Class, or the Interface. There's no need to use deployment classifiers to solve this particular problem.","from":"developer"},{"body":"Attaching a patch that brings in hadoop-compat, and adds the progressable wrapper to it. I'm going to submit this upstream to the elephant bird project, so we should be able to remove this code and add the dependency in the future.","from":"developer"},{"body":"Updated progressable-wrapper.patch to conform to current namespaces.","from":"developer"},{"body":"I removed the hadoop-compat dependency, as maven couldn't resolve the dependencies, reverted to the old version. This won't stop our input/output formats from being compatible with past and future releases.","from":"developer"},{"body":"Patch for 2.1 branch","from":"developer"},{"body":"I had the same maven resolution problem with hadoop-compat. Committed.","from":"developer"}],"created":"2013-01-30T17:11:09.000+0000","description":"Using Hadoop 0.22.0 with Cassandra results in the stack trace below.\nIt appears that version 0.21+ changed org.apache.hadoop.mapreduce.JobContext\nfrom a class to an interface.\n\n\nException in thread \"main\" java.lang.IncompatibleClassChangeError: Found interface org.apache.hadoop.mapreduce.JobContext, but class was expected\n\tat org.apache.cassandra.hadoop.ColumnFamilyInputFormat.getSplits(ColumnFamilyInputFormat.java:103)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.writeNewSplits(JobSubmitter.java:445)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.writeSplits(JobSubmitter.java:462)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.submitJobInternal(JobSubmitter.java:357)\n\tat org.apache.hadoop.mapreduce.Job$2.run(Job.java:1045)\n\tat org.apache.hadoop.mapreduce.Job$2.run(Job.java:1042)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:415)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1153)\n\tat org.apache.hadoop.mapreduce.Job.submit(Job.java:1042)\n\tat org.apache.hadoop.mapreduce.Job.waitForCompletion(Job.java:1062)\n\tat MyHadoopApp.run(MyHadoopApp.java:163)\n\tat org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:69)\n\tat MyHadoopApp.main(MyHadoopApp.java:82)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:601)\n\tat org.apache.hadoop.util.RunJar.main(RunJar.java:192)\n","issue_id":"12629942","key":"CASSANDRA-5201","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-03-05T19:00:44.000+0000","role":"fixed_distractor","summary":"Cassandra/Hadoop does not support current Hadoop releases"} {"case_id":"12631394","cluster":"DISTRACTOR-CASSANDRA-5234","comments":[{"body":"This is not a bug - CQL3 tables are intentionally not included in thrift describe_keyspace(s) (CASSANDRA-4377).","created":"2013-02-08T16:12:14.497+0000"},{"body":"See CASSANDRA-4421","created":"2013-02-08T19:03:18.975+0000"},{"body":"It means that CQL3 column families are not accessible through thrift and for me it's an issue (I do not agree with your Resolution label). That's why Pig 0.11 cannot use it. Is there a way to solve it ?\nI can help you if necessary","created":"2013-03-27T13:15:02.530+0000"},{"body":"It should be fixed after [CASSANDRA-4421|https://issues.apache.org/jira/browse/CASSANDRA-4421]","created":"2013-05-13T08:48:14.367+0000"},{"body":"I will work on it once I am done with other assignments soon.","created":"2013-05-16T04:13:50.070+0000"},{"body":"thanks Alex !","created":"2013-05-16T05:44:27.321+0000"},{"body":"To fix it, we need modify CassandraStorage to get CF meta data from system table instead of thrift describe_keyspace because of the CQL3 table doesn't show up in thrift describe_keyspace call.","created":"2013-05-16T22:03:13.596+0000"},{"body":"Patch is attached. It's on top of the 4421 patch","created":"2013-05-30T18:24:18.004+0000"},{"body":"pull @ https://github.com/alexliu68/cassandra/pull/3\n\nUse CassandraStorage for any cql3 tables, you will have composite columns in \"columns\" bag\n\nUse CqlStorage for any cql3 table.\n{code}\ncassandra://[username:password@]/[?[page_size=]\n[&columns=][&output_query=]\n[&where_clause=][&split_size=][&partitioner=]]\n{code}\n\nwhere \n page_size is the number of cql3 rows per page (the default is 1000, it's optional)\n\n columns is the column names for the cql3 select query, it's optional\n \n where_clause is the user defined where clause on the indexed column, it's optional\n\n split_size is the number of C* rows per split which can be used to tune the number of mappers\n\n output_query is the prepared query for inserting data to cql3 table (replace the = by @ and ? by #,\n because Pig can't take = and ? as url parameter values)\n\nOutput row are in the following format\n{code}\n(((name, value), (name, value)), (value ... value), (value...value))\n{code}\n\nwhere the name and value tuples are key name and value pairs.\n\n\nThe input schema: ((name, value), (name, value), (name, value)) where keys are in the front.","created":"2013-05-30T18:32:06.924+0000"},{"body":"I attach 5234-1.2-patch.txt to patch 1.2 branch. It update the last patch with the latest 4421 changes.","created":"2013-06-14T05:51:18.672+0000"},{"body":"The patch resolves the following issues.\n\n1. allow access to cql3 type table through CassandraStorage.\n\n2. create new CqlStorage to easy access cql3 tables.","created":"2013-06-14T05:53:39.792+0000"},{"body":"Do you think I can test them now ?","created":"2013-06-14T09:23:41.995+0000"},{"body":"yes, I have done some testing.","created":"2013-06-14T15:34:21.017+0000"},{"body":"Okay, I'll give it a try :)\nthanks","created":"2013-06-14T17:46:17.136+0000"},{"body":"My Pig script doesn't work anymore. I suppose you changed the input format ? \nI get :\nProjected field [filtre] does not exist in schema: key:chararray,columns:bag{:tuple(name:tuple(),value:byte array)}\nwhen column family is created without COMPACT STORAGE\n\nSomething weird is that I get no values when I should get some :\ncqlsh:k1> SELECT * FROM cf1 WHERE ISE='XXXX';\n\n ise | filtre | value_1\n------+--------+---------\n XXXX | 1 | 81056\n\n2013-06-17 13:13:01,372 [main] INFO org.apache.pig.backend.hadoop.executionengine.util.MapRedUtil - Total input paths to process : 1\n(XXXX,{((),),((filtre),),((value_1), getQueryMap(String query) throws Exception\n {\n String[] params = query.split(\"&\");\n Map map = new HashMap();\n for (String param : params)\n {\n String[] keyValue = param.split(\"=\");\n map.put(keyValue[0], URLDecoder.decode(keyValue[1],\"UTF-8\"));\n\t\t\t\n }\n return map;\n }\nand now i can send the query with URL-Encoding character.\n","created":"2013-07-09T15:07:24.054+0000"},{"body":"Thx, I will update the patch for it.","created":"2013-07-09T15:55:00.124+0000"},{"body":"add the patch as a temporary fix","created":"2013-07-10T08:25:32.886+0000"},{"body":"I was trying it locally and I'm not sure whether it fully works. My stacktrace and cassandra table schema is here: http://pastebin.com/uPUAs9T2 (i tried using old cassandra:// ... - it worked for other table) and output when I used cql://.. is here: http://pastebin.com/b0bKd7G3 . I have worked on this with Pig in local mode. It might be important or not: this table contained counters.\n\nI was working on cassandra-1.2 with latest commit 27efded38d855b24f41e5332ffb29cd13d98f8da","created":"2013-07-19T14:00:39.128+0000"},{"body":"Just adding that I'm getting the same result as Konrad Kurdej. I'm using DSE 3.1, but I think this is the same bug. I was able to isolate the problem specifically to counter fields. Here's a simple set up and test case:\n\ncqlsh> create keyspace pigtest with REPLICATION = { 'class': 'SimpleStrategy', 'replication_factor': 1};\ncqlsh> use pigtest;\ncqlsh:pigtest> create table foo ( key_1 text primary key, value_1 bigint );\ncqlsh:pigtest> create table foo2 ( key_1 text primary key, value_1 counter );\ncqlsh:pigtest> update foo set value_1 = 1 where key_1 = 'foo';\ncqlsh:pigtest> update foo2 set value_1 = value_1 + 2 where key_1 = 'foo2';\ncqlsh:pigtest> select * from foo;\n\n key_1 | value_1\n-------+---------\n foo | 1\n\ncqlsh:pigtest> select * from foo2;\n\n key_1 | value_1\n-------+---------\n foo2 | 2\n\nNow, the following grunt commands:\n\ncounts = LOAD 'cassandra://pigtest/foo' USING CassandraStorage();\ndump counts;\n\nWill work, but:\n\ncounts = LOAD 'cassandra://pigtest/foo2' USING CassandraStorage();\ndump counts;\n\nWill fail with the same stack trace that Konrad mentioned.","created":"2013-07-29T03:41:49.413+0000"},{"body":"This patch against the head appears to have fixed the problem for me. I applied it to DSE 3.1 and it also worked for me.\n\nBasic deal is to use LongType for validation too.","created":"2013-07-29T14:33:30.314+0000"},{"body":"Use CqlStorage for your use case. CassandraStorage has some drawbacks.","created":"2013-07-29T17:47:31.049+0000"},{"body":"CassandraStorage is legacy for any none-CQL3 tables. Use CqlStorage for CQL3 tables.","created":"2013-07-29T18:06:51.847+0000"},{"body":"Cassandra 1.2.8 was released (after regression in 1.2.7), but this last patch for counters appears to be left. Anyone can confirm this?\n\nI downloaded the .deb from http://people.apache.org/~eevans/ but when in grunt i dump a table with counter column, get the error:\n\njava.lang.IndexOutOfBoundsException\n\tat java.nio.Buffer.checkIndex(Buffer.java:537)\n\tat java.nio.HeapByteBuffer.getLong(HeapByteBuffer.java:410)\n\tat org.apache.cassandra.db.context.CounterContext.total(CounterContext.java:477)\n\tat org.apache.cassandra.db.marshal.AbstractCommutativeType.compose(AbstractCommutativeType.java:34)\n\tat org.apache.cassandra.db.marshal.AbstractCommutativeType.compose(AbstractCommutativeType.java:25)\n\tat org.apache.cassandra.hadoop.pig.AbstractCassandraStorage.columnToTuple(AbstractCassandraStorage.java:137)\n\tat org.apache.cassandra.hadoop.pig.CqlStorage.getNext(CqlStorage.java:110)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigRecordReader.nextKeyValue(PigRecordReader.java:211)\n\tat org.apache.hadoop.mapred.MapTask$NewTrackingRecordReader.nextKeyValue(MapTask.java:532)\n\tat org.apache.hadoop.mapreduce.MapContext.nextKeyValue(MapContext.java:67)\n\tat org.apache.hadoop.mapreduce.Mapper.run(Mapper.java:143)\n\tat org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:764)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:370)\n\tat org.apache.hadoop.mapred.Child$4.run(Child.java:255)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:416)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1149)\n\tat org.apache.hadoop.mapred.Child.main(Child.java:249)\n\nMy table:\ncqlsh:test> desc table votes_count_period3;\n\nCREATE TABLE votes_count_period3 (\n period text,\n poll timeuuid,\n votes counter,\n PRIMARY KEY (period, poll)\n) WITH\n bloom_filter_fp_chance=0.010000 AND\n caching='KEYS_ONLY' AND\n comment='' AND\n dclocal_read_repair_chance=0.000000 AND\n gc_grace_seconds=864000 AND\n read_repair_chance=0.100000 AND\n replicate_on_write='true' AND\n populate_io_cache_on_flush='false' AND\n compaction={'class': 'SizeTieredCompactionStrategy'} AND\n compression={'sstable_compression': 'SnappyCompressor'};\n\n","created":"2013-07-30T01:55:21.285+0000"},{"body":"I confirmed this issue. I downloaded the src, applied the patch and after build the cassandra jar, pig works with counter. This will be patched in 1.2.9?","created":"2013-07-30T02:20:01.855+0000"},{"body":"+1","created":"2013-07-30T03:31:27.306+0000"},{"body":"[~alexliu68] The patch is to the base class shared by CqlStorage and CassandraStorage. They were both broken and they are now both fixed.","created":"2013-07-30T04:40:04.167+0000"},{"body":"The last patch for counters to AbstractCassandraStorage.java has not been applied either to cassandra-1.2 or to the trunk. Consequently, the problem still exists in 1.2.9, as I could verify myself today. I found another recent bug report on SO for the same problem: http://stackoverflow.com/questions/18553230/error-with-cassandra-pig-cql-counter-column.\n\nShould I open a new bug report for the counters bug even though we have a working patch or will you reopen the current issue?","created":"2013-09-05T18:35:48.356+0000"}],"conversations":[{"body":"Hi,\n i have faced a bug when creating table through CQL3 and trying to load data through pig 0.10 as follows:\njava.lang.RuntimeException: Column family 'abc' not found in keyspace 'XYZ'\n\tat org.apache.cassandra.hadoop.pig.CassandraStorage.initSchema(CassandraStorage.java:1112)\n\tat org.apache.cassandra.hadoop.pig.CassandraStorage.setLocation(CassandraStorage.java:615).\nThis effects from Simple table to table with compound key. ","from":"reporter","subject":"Table created through CQL3 are not accessble to Pig 0.10"},{"body":"This is not a bug - CQL3 tables are intentionally not included in thrift describe_keyspace(s) (CASSANDRA-4377).","from":"developer"},{"body":"See CASSANDRA-4421","from":"developer"},{"body":"It means that CQL3 column families are not accessible through thrift and for me it's an issue (I do not agree with your Resolution label). That's why Pig 0.11 cannot use it. Is there a way to solve it ?\nI can help you if necessary","from":"developer"},{"body":"It should be fixed after [CASSANDRA-4421|https://issues.apache.org/jira/browse/CASSANDRA-4421]","from":"developer"},{"body":"I will work on it once I am done with other assignments soon.","from":"developer"},{"body":"thanks Alex !","from":"developer"},{"body":"To fix it, we need modify CassandraStorage to get CF meta data from system table instead of thrift describe_keyspace because of the CQL3 table doesn't show up in thrift describe_keyspace call.","from":"developer"},{"body":"Patch is attached. It's on top of the 4421 patch","from":"developer"},{"body":"pull @ https://github.com/alexliu68/cassandra/pull/3\n\nUse CassandraStorage for any cql3 tables, you will have composite columns in \"columns\" bag\n\nUse CqlStorage for any cql3 table.\n{code}\ncassandra://[username:password@]/[?[page_size=]\n[&columns=][&output_query=]\n[&where_clause=][&split_size=][&partitioner=]]\n{code}\n\nwhere \n page_size is the number of cql3 rows per page (the default is 1000, it's optional)\n\n columns is the column names for the cql3 select query, it's optional\n \n where_clause is the user defined where clause on the indexed column, it's optional\n\n split_size is the number of C* rows per split which can be used to tune the number of mappers\n\n output_query is the prepared query for inserting data to cql3 table (replace the = by @ and ? by #,\n because Pig can't take = and ? as url parameter values)\n\nOutput row are in the following format\n{code}\n(((name, value), (name, value)), (value ... value), (value...value))\n{code}\n\nwhere the name and value tuples are key name and value pairs.\n\n\nThe input schema: ((name, value), (name, value), (name, value)) where keys are in the front.","from":"developer"},{"body":"I attach 5234-1.2-patch.txt to patch 1.2 branch. It update the last patch with the latest 4421 changes.","from":"developer"},{"body":"The patch resolves the following issues.\n\n1. allow access to cql3 type table through CassandraStorage.\n\n2. create new CqlStorage to easy access cql3 tables.","from":"developer"},{"body":"Do you think I can test them now ?","from":"developer"},{"body":"yes, I have done some testing.","from":"developer"},{"body":"Okay, I'll give it a try :)\nthanks","from":"developer"},{"body":"My Pig script doesn't work anymore. I suppose you changed the input format ? \nI get :\nProjected field [filtre] does not exist in schema: key:chararray,columns:bag{:tuple(name:tuple(),value:byte array)}\nwhen column family is created without COMPACT STORAGE\n\nSomething weird is that I get no values when I should get some :\ncqlsh:k1> SELECT * FROM cf1 WHERE ISE='XXXX';\n\n ise | filtre | value_1\n------+--------+---------\n XXXX | 1 | 81056\n\n2013-06-17 13:13:01,372 [main] INFO org.apache.pig.backend.hadoop.executionengine.util.MapRedUtil - Total input paths to process : 1\n(XXXX,{((),),((filtre),),((value_1), getQueryMap(String query) throws Exception\n {\n String[] params = query.split(\"&\");\n Map map = new HashMap();\n for (String param : params)\n {\n String[] keyValue = param.split(\"=\");\n map.put(keyValue[0], URLDecoder.decode(keyValue[1],\"UTF-8\"));\n\t\t\t\n }\n return map;\n }\nand now i can send the query with URL-Encoding character.\n","from":"developer"},{"body":"Thx, I will update the patch for it.","from":"developer"},{"body":"add the patch as a temporary fix","from":"developer"},{"body":"I was trying it locally and I'm not sure whether it fully works. My stacktrace and cassandra table schema is here: http://pastebin.com/uPUAs9T2 (i tried using old cassandra:// ... - it worked for other table) and output when I used cql://.. is here: http://pastebin.com/b0bKd7G3 . I have worked on this with Pig in local mode. It might be important or not: this table contained counters.\n\nI was working on cassandra-1.2 with latest commit 27efded38d855b24f41e5332ffb29cd13d98f8da","from":"developer"},{"body":"Just adding that I'm getting the same result as Konrad Kurdej. I'm using DSE 3.1, but I think this is the same bug. I was able to isolate the problem specifically to counter fields. Here's a simple set up and test case:\n\ncqlsh> create keyspace pigtest with REPLICATION = { 'class': 'SimpleStrategy', 'replication_factor': 1};\ncqlsh> use pigtest;\ncqlsh:pigtest> create table foo ( key_1 text primary key, value_1 bigint );\ncqlsh:pigtest> create table foo2 ( key_1 text primary key, value_1 counter );\ncqlsh:pigtest> update foo set value_1 = 1 where key_1 = 'foo';\ncqlsh:pigtest> update foo2 set value_1 = value_1 + 2 where key_1 = 'foo2';\ncqlsh:pigtest> select * from foo;\n\n key_1 | value_1\n-------+---------\n foo | 1\n\ncqlsh:pigtest> select * from foo2;\n\n key_1 | value_1\n-------+---------\n foo2 | 2\n\nNow, the following grunt commands:\n\ncounts = LOAD 'cassandra://pigtest/foo' USING CassandraStorage();\ndump counts;\n\nWill work, but:\n\ncounts = LOAD 'cassandra://pigtest/foo2' USING CassandraStorage();\ndump counts;\n\nWill fail with the same stack trace that Konrad mentioned.","from":"developer"},{"body":"This patch against the head appears to have fixed the problem for me. I applied it to DSE 3.1 and it also worked for me.\n\nBasic deal is to use LongType for validation too.","from":"developer"},{"body":"Use CqlStorage for your use case. CassandraStorage has some drawbacks.","from":"developer"},{"body":"CassandraStorage is legacy for any none-CQL3 tables. Use CqlStorage for CQL3 tables.","from":"developer"},{"body":"Cassandra 1.2.8 was released (after regression in 1.2.7), but this last patch for counters appears to be left. Anyone can confirm this?\n\nI downloaded the .deb from http://people.apache.org/~eevans/ but when in grunt i dump a table with counter column, get the error:\n\njava.lang.IndexOutOfBoundsException\n\tat java.nio.Buffer.checkIndex(Buffer.java:537)\n\tat java.nio.HeapByteBuffer.getLong(HeapByteBuffer.java:410)\n\tat org.apache.cassandra.db.context.CounterContext.total(CounterContext.java:477)\n\tat org.apache.cassandra.db.marshal.AbstractCommutativeType.compose(AbstractCommutativeType.java:34)\n\tat org.apache.cassandra.db.marshal.AbstractCommutativeType.compose(AbstractCommutativeType.java:25)\n\tat org.apache.cassandra.hadoop.pig.AbstractCassandraStorage.columnToTuple(AbstractCassandraStorage.java:137)\n\tat org.apache.cassandra.hadoop.pig.CqlStorage.getNext(CqlStorage.java:110)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigRecordReader.nextKeyValue(PigRecordReader.java:211)\n\tat org.apache.hadoop.mapred.MapTask$NewTrackingRecordReader.nextKeyValue(MapTask.java:532)\n\tat org.apache.hadoop.mapreduce.MapContext.nextKeyValue(MapContext.java:67)\n\tat org.apache.hadoop.mapreduce.Mapper.run(Mapper.java:143)\n\tat org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:764)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:370)\n\tat org.apache.hadoop.mapred.Child$4.run(Child.java:255)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:416)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1149)\n\tat org.apache.hadoop.mapred.Child.main(Child.java:249)\n\nMy table:\ncqlsh:test> desc table votes_count_period3;\n\nCREATE TABLE votes_count_period3 (\n period text,\n poll timeuuid,\n votes counter,\n PRIMARY KEY (period, poll)\n) WITH\n bloom_filter_fp_chance=0.010000 AND\n caching='KEYS_ONLY' AND\n comment='' AND\n dclocal_read_repair_chance=0.000000 AND\n gc_grace_seconds=864000 AND\n read_repair_chance=0.100000 AND\n replicate_on_write='true' AND\n populate_io_cache_on_flush='false' AND\n compaction={'class': 'SizeTieredCompactionStrategy'} AND\n compression={'sstable_compression': 'SnappyCompressor'};\n\n","from":"developer"},{"body":"I confirmed this issue. I downloaded the src, applied the patch and after build the cassandra jar, pig works with counter. This will be patched in 1.2.9?","from":"developer"},{"body":"+1","from":"developer"},{"body":"[~alexliu68] The patch is to the base class shared by CqlStorage and CassandraStorage. They were both broken and they are now both fixed.","from":"developer"},{"body":"The last patch for counters to AbstractCassandraStorage.java has not been applied either to cassandra-1.2 or to the trunk. Consequently, the problem still exists in 1.2.9, as I could verify myself today. I found another recent bug report on SO for the same problem: http://stackoverflow.com/questions/18553230/error-with-cassandra-pig-cql-counter-column.\n\nShould I open a new bug report for the counters bug even though we have a working patch or will you reopen the current issue?","from":"developer"}],"created":"2013-02-08T08:13:31.000+0000","description":"Hi,\n i have faced a bug when creating table through CQL3 and trying to load data through pig 0.10 as follows:\njava.lang.RuntimeException: Column family 'abc' not found in keyspace 'XYZ'\n\tat org.apache.cassandra.hadoop.pig.CassandraStorage.initSchema(CassandraStorage.java:1112)\n\tat org.apache.cassandra.hadoop.pig.CassandraStorage.setLocation(CassandraStorage.java:615).\nThis effects from Simple table to table with compound key. ","issue_id":"12631394","key":"CASSANDRA-5234","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-06-28T23:09:26.000+0000","role":"fixed_distractor","summary":"Table created through CQL3 are not accessble to Pig 0.10"} {"case_id":"12631964","cluster":"DISTRACTOR-CASSANDRA-5245","comments":[{"body":"That assertion is over 3 years old, so it's not as simple as \"1.2 added a bogus assert.\" (Which is not what you claimed, of course.)","created":"2013-02-12T15:55:48.863+0000"},{"body":"What do you mean? That assertion is from my logs yesterday? Been able to repro it 3-4 times during the last couple of days. I couldn't repro it on 1.1. (at least not so far). If it is a common assertion during repairs, then it should not be marked as closed.\n\nThe question was asked in the Cassandra User list to open a new bug or just edit the old one. Thats why I created a new one. (http://www.mail-archive.com/user@cassandra.apache.org/msg27686.html)","created":"2013-02-12T16:11:00.494+0000"},{"body":"We have the same problem even with Cassandra 1.2.2. ","created":"2013-02-27T13:50:47.862+0000"},{"body":"I have also run into this issue today.\n\n1.2.2, Vnodes, Murmur3Partitioner\n\n3 Nodes, RF3\n\nAll nodes have the following stacktrace 1+ times:\n\n{noformat}\nERROR [AntiEntropyStage:10] 2013-03-06 01:46:21,159 CassandraDaemon.java (line 132) Exception in thread Thread[AntiEntropyStage:10,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.utils.MerkleTree.inc(MerkleTree.java:137)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:245)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.difference(MerkleTree.java:227)\n\tat org.apache.cassandra.service.AntiEntropyService$RepairSession$Differencer.run(AntiEntropyService.java:983)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n\tat java.lang.Thread.run(Unknown Source)\n{noformat}\n\nAnother example of stacktrace:\n{noformat}\nERROR [AntiEntropyStage:9] 2013-03-05 22:24:41,730 CassandraDaemon.java (line 132) Exception in thread Thread[AntiEntropyStage:9,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.utils.MerkleTree.inc(MerkleTree.java:137)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:245)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n{noformat}","created":"2013-03-06T03:39:15.224+0000"},{"body":"I am able to reproduce this after restarting the cluster and trying to run nodetool repair -pr (instead of just nodetool repair)\n\n{noformat}\n \nERROR [AntiEntropyStage:8] 2013-03-06 07:56:24,794 CassandraDaemon.java (line 132) Exception in thread Thread[AntiEntropyStage:8,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.utils.MerkleTree.inc(MerkleTree.java:137)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:245)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.difference(MerkleTree.java:227)\n at org.apache.cassandra.service.AntiEntropyService$RepairSession$Differencer.run(AntiEntropyService.java:983)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\n{noformat}","created":"2013-03-06T13:55:44.704+0000"},{"body":"I just look at the code and found this part: https://github.com/apache/cassandra/blob/cassandra-1.2.2/src/java/org/apache/cassandra/service/AntiEntropyService.java#L299\n\nI don't test anything and I'm not sure this is related yet, but if we are using M3P, then we do manually splitting the tree based on the key samples between given range. We can have not enough splits initially.","created":"2013-03-06T14:23:24.691+0000"},{"body":"Using the key sample for M3P was definitively not the intention. That instanceof should be changed to a call to preserveOrder() really. I'm less surprised that the code for the ordering partition may get in an infinite recursion, though we should obviously fix it nonetheless.","created":"2013-03-06T15:51:43.724+0000"},{"body":"Patch to fix as Sylvain's suggestion.\nI think this is enough to fix AssertionError when using M3P.","created":"2013-03-06T18:50:29.247+0000"},{"body":"+1 on patch.\n\nReproduced repair issue with ccm and stress pretty easily with 256 tokens and M3. Calling preserveOrder() fixes this.","created":"2013-03-07T02:42:10.446+0000"},{"body":"Yes, +1 on Yuki's patch because that's a problem. But this doesn't fix the fact that differenceHelper might recurse one time too much if the tree ends up being Byte.MAX_VALUE-1 deep (which can happen with the method used to build the tree for OrderPreservingPartitioner, which was mistakenly used for M3P). So attaching a 2nd patch that I believe should fix the recursion problem.\n","created":"2013-03-07T14:00:31.207+0000"},{"body":"+1. Sylvain, could you commit both?","created":"2013-03-07T15:55:14.801+0000"},{"body":"Any chance these fixes will make it into the 1.2.3 release?","created":"2013-03-12T17:49:52.502+0000"},{"body":"Yep, committed, thanks.","created":"2013-03-12T18:03:47.490+0000"}],"conversations":[{"body":"We are seeing AntiEntropy errors when performing repair jobs in one of our Cassandra clusters. It seems to have started with 1.2. (maybe an issue with vnodes) The exceptions occur almost every time we try to do a repair on all column families in the cluster. Doing the same task on 1.1 does not trigger this.\n\n6 nodes cluster (vnodes, murmur3, rf:3)\nvery low activity\nrunning a nodetool repair -pr loop on the cluster nodes\nnodetool hangs, and same big stacktrace in logs.\n\nroot 11025 0.0 0.0 106100 1436 pts/0 S+ Feb11 0:00 _ /bin/sh /usr/bin/nodetool -h HOST -p 7199 -pr repair KEYSPACE COLUMN_FAMILY\n\nERROR [AntiEntropyStage:3] 2013-02-11 17:08:12,630 CassandraDaemon.java (line 133) Exception in thread Thread[AntiEntropyStage:3,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.utils.MerkleTree.inc(MerkleTree.java:137)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:245)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.difference(MerkleTree.java:227)\n\tat org.apache.cassandra.service.AntiEntropyService$RepairSession$Differencer.run(AntiEntropyService.java:982)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\n\nThis issue was partially solved earlier but seems to be back with vnodes: https://issues.apache.org/jira/browse/CASSANDRA-3014\n","from":"reporter","subject":"AnitEntropy/MerkleTree Error"},{"body":"That assertion is over 3 years old, so it's not as simple as \"1.2 added a bogus assert.\" (Which is not what you claimed, of course.)","from":"developer"},{"body":"What do you mean? That assertion is from my logs yesterday? Been able to repro it 3-4 times during the last couple of days. I couldn't repro it on 1.1. (at least not so far). If it is a common assertion during repairs, then it should not be marked as closed.\n\nThe question was asked in the Cassandra User list to open a new bug or just edit the old one. Thats why I created a new one. (http://www.mail-archive.com/user@cassandra.apache.org/msg27686.html)","from":"developer"},{"body":"We have the same problem even with Cassandra 1.2.2. ","from":"developer"},{"body":"I have also run into this issue today.\n\n1.2.2, Vnodes, Murmur3Partitioner\n\n3 Nodes, RF3\n\nAll nodes have the following stacktrace 1+ times:\n\n{noformat}\nERROR [AntiEntropyStage:10] 2013-03-06 01:46:21,159 CassandraDaemon.java (line 132) Exception in thread Thread[AntiEntropyStage:10,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.utils.MerkleTree.inc(MerkleTree.java:137)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:245)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.difference(MerkleTree.java:227)\n\tat org.apache.cassandra.service.AntiEntropyService$RepairSession$Differencer.run(AntiEntropyService.java:983)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n\tat java.lang.Thread.run(Unknown Source)\n{noformat}\n\nAnother example of stacktrace:\n{noformat}\nERROR [AntiEntropyStage:9] 2013-03-05 22:24:41,730 CassandraDaemon.java (line 132) Exception in thread Thread[AntiEntropyStage:9,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.utils.MerkleTree.inc(MerkleTree.java:137)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:245)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n{noformat}","from":"developer"},{"body":"I am able to reproduce this after restarting the cluster and trying to run nodetool repair -pr (instead of just nodetool repair)\n\n{noformat}\n \nERROR [AntiEntropyStage:8] 2013-03-06 07:56:24,794 CassandraDaemon.java (line 132) Exception in thread Thread[AntiEntropyStage:8,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.utils.MerkleTree.inc(MerkleTree.java:137)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:245)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n at org.apache.cassandra.utils.MerkleTree.difference(MerkleTree.java:227)\n at org.apache.cassandra.service.AntiEntropyService$RepairSession$Differencer.run(AntiEntropyService.java:983)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\n{noformat}","from":"developer"},{"body":"I just look at the code and found this part: https://github.com/apache/cassandra/blob/cassandra-1.2.2/src/java/org/apache/cassandra/service/AntiEntropyService.java#L299\n\nI don't test anything and I'm not sure this is related yet, but if we are using M3P, then we do manually splitting the tree based on the key samples between given range. We can have not enough splits initially.","from":"developer"},{"body":"Using the key sample for M3P was definitively not the intention. That instanceof should be changed to a call to preserveOrder() really. I'm less surprised that the code for the ordering partition may get in an infinite recursion, though we should obviously fix it nonetheless.","from":"developer"},{"body":"Patch to fix as Sylvain's suggestion.\nI think this is enough to fix AssertionError when using M3P.","from":"developer"},{"body":"+1 on patch.\n\nReproduced repair issue with ccm and stress pretty easily with 256 tokens and M3. Calling preserveOrder() fixes this.","from":"developer"},{"body":"Yes, +1 on Yuki's patch because that's a problem. But this doesn't fix the fact that differenceHelper might recurse one time too much if the tree ends up being Byte.MAX_VALUE-1 deep (which can happen with the method used to build the tree for OrderPreservingPartitioner, which was mistakenly used for M3P). So attaching a 2nd patch that I believe should fix the recursion problem.\n","from":"developer"},{"body":"+1. Sylvain, could you commit both?","from":"developer"},{"body":"Any chance these fixes will make it into the 1.2.3 release?","from":"developer"},{"body":"Yep, committed, thanks.","from":"developer"}],"created":"2013-02-12T15:43:03.000+0000","description":"We are seeing AntiEntropy errors when performing repair jobs in one of our Cassandra clusters. It seems to have started with 1.2. (maybe an issue with vnodes) The exceptions occur almost every time we try to do a repair on all column families in the cluster. Doing the same task on 1.1 does not trigger this.\n\n6 nodes cluster (vnodes, murmur3, rf:3)\nvery low activity\nrunning a nodetool repair -pr loop on the cluster nodes\nnodetool hangs, and same big stacktrace in logs.\n\nroot 11025 0.0 0.0 106100 1436 pts/0 S+ Feb11 0:00 _ /bin/sh /usr/bin/nodetool -h HOST -p 7199 -pr repair KEYSPACE COLUMN_FAMILY\n\nERROR [AntiEntropyStage:3] 2013-02-11 17:08:12,630 CassandraDaemon.java (line 133) Exception in thread Thread[AntiEntropyStage:3,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.utils.MerkleTree.inc(MerkleTree.java:137)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:245)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:256)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.differenceHelper(MerkleTree.java:267)\n\tat org.apache.cassandra.utils.MerkleTree.difference(MerkleTree.java:227)\n\tat org.apache.cassandra.service.AntiEntropyService$RepairSession$Differencer.run(AntiEntropyService.java:982)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\n\nThis issue was partially solved earlier but seems to be back with vnodes: https://issues.apache.org/jira/browse/CASSANDRA-3014\n","issue_id":"12631964","key":"CASSANDRA-5245","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-03-12T18:03:47.000+0000","role":"fixed_distractor","summary":"AnitEntropy/MerkleTree Error"} {"case_id":"12637020","cluster":"DISTRACTOR-CASSANDRA-5345","comments":[{"body":"Personally, I doubt this theory, the JVM has no reason to make PN disappear, and this code has been around for a long time with no similar reports. I think you might have a build problem.","created":"2013-03-14T19:46:54.037+0000"},{"body":"Running 1.2.14 on Ubuntu 12.04 with HotSpot Java 1.7.0_21-b11. \n\nFor the past couple of days our nodes started to produce this exception every second. This exception started to get produced after we decommissioned one DC, and it continued to get produced until the nodes started producing client errors as we got alerts. No suspicious other log entries were found and no full GCs were recorded in the GC logs. The GC logs looked normal however system.log was getting filled with this exception. \n\nWe decommissioned the DC by first altering our keyspace to not replication to that DC as we are using NetworkTopologyStrategy. Then we issued nodetool decommission on each node in that DC. \n\nWhy this exception started to get produced is not known. However, restarting the nodes fixed the problem. Another observation was that the nodetool info command which we use to collect heap size statistics was not working as well and was tossing this exception:\n\nException in thread \"main\" java.lang.IllegalArgumentException: javax.management.InstanceNotFoundException: java.lang:type=Memory\n at java.lang.management.ManagementFactory.newPlatformMXBeanProxy(ManagementFactory.java:610)\n at org.apache.cassandra.tools.NodeProbe.connect(NodeProbe.java:175)\n at org.apache.cassandra.tools.NodeProbe.(NodeProbe.java:116)\n at org.apache.cassandra.tools.NodeCmd.main(NodeCmd.java:1138)\nCaused by: javax.management.InstanceNotFoundException: java.lang:type=Memory\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.getMBean(DefaultMBeanServerInterceptor.java:1095)\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.isInstanceOf(DefaultMBeanServerInterceptor.java:1401)\n at com.sun.jmx.mbeanserver.JmxMBeanServer.isInstanceOf(JmxMBeanServer.java:1082)\n at javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1492)\n at javax.management.remote.rmi.RMIConnectionImpl.access$300(RMIConnectionImpl.java:96)\n at javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1327)\n at javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1419)\n at javax.management.remote.rmi.RMIConnectionImpl.isInstanceOf(RMIConnectionImpl.java:957)\n at sun.reflect.GeneratedMethodAccessor82.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:601)\n at sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:322)\n at sun.rmi.transport.Transport$1.run(Transport.java:177)","created":"2014-03-25T01:11:15.378+0000"},{"body":"Matt Byrd, please assign version and increase the priority as this did in fact bring our cluster to halt. ","created":"2014-03-25T01:13:19.633+0000"},{"body":"Ryan, can your team repro?","created":"2014-03-25T03:51:43.759+0000"},{"body":"The cluster in question was running on 1.0.7, however the code in question has remained static since well before that and doesn't look to have changed since. (though admittedly the problem could somehow be being caused elsewhere, jvm maybe?) \nI've upped the priority to major.\n\nHave you been able to reproduce? or seen the problem anywhere else?\nAny further details about your environment is set up and how you deploy may also help those trying to reproduce.\nSome common but perhaps co-incidental things about the two occurrences:\n1. virtual machines (though not both AWS)\n2. multi D.C , wouldn't have though this would be relevant but Arya does seem to see the problem after removing a d.c.\n3. Slightly old Jvm versions...\n\nI no longer have access to the cluster where we saw this previously but let me know if I can help in any other way.","created":"2014-03-25T06:30:46.875+0000"},{"body":"To give you more details about our setup:\n\nAWS\nus-east-1 3 nodes. \nus-west-2 3 nodes.\nClassic EC2 hi1.4xlarge machines\nNetworkTopologyStrategy us-east-1:3, us-west-2:3\nEffective load on each machine is 100% about 360Gb\nWe run a 24Gb Heap with MaxTenuringThreshold = 20\n\nAverage read 4k/sec on each node\nAverage write 1.5k/sec on each node\n\nWe have a mix of SizeTiered and Leveled CFs.\nWe turned off read repair. \nWe use mmap_index_only","created":"2014-03-25T19:56:41.084+0000"},{"body":"[~enigmacurry] - did you or your team have any luck reproducing this?\n\nIt should be trivial to throw a gc.isValid() check in the logGCResults loop and if invalid, flag to rebuild the List we're iterating across in that function as well as to introduce some exception handling for the UndeclaredThrowableException. I'm not finding much on the logic behind *when* these MXBeans can be removed from the system and I agree with Brandon that it seems incredibly odd for a JVM to punt and re-init a garbage collector MXBean on the fly with such infrequency that we haven't seen this more often.\n\nThe fact that Arya had a nodetool failure connecting to the Memory subsystem:\n{code:title=Memory failure}\nException in thread \"main\" java.lang.IllegalArgumentException: javax.management.InstanceNotFoundException: java.lang:type=Memory\nat java.lang.management.ManagementFactory.newPlatformMXBeanProxy(ManagementFactory.java:610)\nat org.apache.cassandra.tools.NodeProbe.connect(NodeProbe.java:175)\n{code}\nis concerning as it would seem to imply some deeper problems w/the JMX integration of the Memory subsystem on these JVM's given that we're querying the factory by String there rather than trying to use an invalid reference.","created":"2014-06-25T21:30:46.453+0000"},{"body":"[~JoshuaMcKenzie] no, I haven't reproduced it. :(","created":"2014-06-25T21:51:59.239+0000"},{"body":"Attaching a v1 that tries to gracefully check for GC MXBean validity on log and if an invalid GC is found, skip it and rebuild the list of cached MXBean's for next logging.\n\nGiven that I can't find any info on *when* these things are getting recycled, I also put in some exception handling in-case the change happens in the middle of our logging process.\n\nThis doesn't address the failure of nodetool info that Arya saw with nodetool info; that issue seems like a related but perhaps different (and less critical) effort than this ticket.","created":"2014-06-26T15:33:18.817+0000"},{"body":"I couldn't get this to apply to the repo as of Jun 26. Can you rebase against current 2.0 branch?","created":"2014-07-17T21:59:16.284+0000"},{"body":"rebased against cassandra-2.0","created":"2014-07-18T15:15:16.636+0000"},{"body":"Added a comment pointing to this issue and committed.","created":"2014-07-18T23:15:20.865+0000"}],"conversations":[{"body":"I am not certain this is definitely a bug, but I thought it might be worth posting to see if someone with more JVM//JMX knowledge could disprove my reasoning. Apologies if I've failed to understand something.\n\nWe've seen an intermittent problem where there is an uncaught exception in the scheduled task of logging gc results in GcInspector.java:\n\n{code}\n...\n ERROR [ScheduledTasks:1] 2013-03-08 01:09:06,335 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[ScheduledTasks:1,5,main]\njava.lang.reflect.UndeclaredThrowableException\n at $Proxy0.getName(Unknown Source)\n at org.apache.cassandra.service.GCInspector.logGCResults(GCInspector.java:95)\n at org.apache.cassandra.service.GCInspector.access$000(GCInspector.java:41)\n at org.apache.cassandra.service.GCInspector$1.run(GCInspector.java:85)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:441)\n at java.util.concurrent.FutureTask$Sync.innerRunAndReset(FutureTask.java:317)\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:150)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$101(ScheduledThreadPoolExecutor.java:98)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.runPeriodic(ScheduledThreadPoolExecutor.java:180)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:204)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: javax.management.InstanceNotFoundException: java.lang:name=ParNew,type=GarbageCollector\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.getMBean(DefaultMBeanServerInterceptor.java:1094)\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.getAttribute(DefaultMBeanServerInterceptor.java:662)\n at com.sun.jmx.mbeanserver.JmxMBeanServer.getAttribute(JmxMBeanServer.java:638)\n at com.sun.jmx.mbeanserver.MXBeanProxy$GetHandler.invoke(MXBeanProxy.java:106)\n at com.sun.jmx.mbeanserver.MXBeanProxy.invoke(MXBeanProxy.java:148)\n at javax.management.MBeanServerInvocationHandler.invoke(MBeanServerInvocationHandler.java:248)\n ... 13 more\n...\n{code}\n\nI think the problem, may be caused by the following reasoning:\n\nIn GcInspector we populate a list of mxbeans when the GcInspector instance is instantiated:\n\n{code}\n...\nList beans = new ArrayList();\n MBeanServer server = ManagementFactory.getPlatformMBeanServer();\n try\n {\n ObjectName gcName = new ObjectName(ManagementFactory.GARBAGE_COLLECTOR_MXBEAN_DOMAIN_TYPE + \",*\");\n for (ObjectName name : server.queryNames(gcName, null))\n {\n GarbageCollectorMXBean gc = ManagementFactory.newPlatformMXBeanProxy(server, name.getCanonicalName(), GarbageCollectorMXBean.class);\n beans.add(gc);\n }\n }\n catch (Exception e)\n {\n throw new RuntimeException(e);\n }\n...\n{code}\n\nCassandra then periodically calls:\n\n{code}\n...\n private void logGCResults()\n {\n for (GarbageCollectorMXBean gc : beans)\n {\n Long previousTotal = gctimes.get(gc.getName());\n...\n{code}\n\nIn the oracle javadocs, they seem to suggest that these beans could disappear at any time.(I'm not sure why when or how this might happen)\nhttp://docs.oracle.com/javase/6/docs/api/\nSee: getGarbageCollectorMXBeans\n\n{code}\n...\npublic static List getGarbageCollectorMXBeans()\nReturns a list of GarbageCollectorMXBean objects in the Java virtual machine. The Java virtual machine may have one or more GarbageCollectorMXBean objects. It may add or remove GarbageCollectorMXBean during execution.\nReturns:\na list of GarbageCollectorMXBean objects.\n...\n{code}\n\nCorrect me if I'm wrong, but do you think this might be causing the problem? That somehow the JVM decides to remove the GarbageCollectorMXBean temporarily or permanently (causing said exception) and if this is expected behaviour, should it be handled in some way?\nAlso I'd like to point out that this may be an issue on other versions as well as I don't believe this code has changed in quite a long time.\nUnfortunately I haven't been able to reproduce this outside of the production environment, if you have any tips, questions or are able to explain//disprove my concerns, I'd be very grateful.\n\nThanks,\nMatt","from":"reporter","subject":"Potential problem with GarbageCollectorMXBean"},{"body":"Personally, I doubt this theory, the JVM has no reason to make PN disappear, and this code has been around for a long time with no similar reports. I think you might have a build problem.","from":"developer"},{"body":"Running 1.2.14 on Ubuntu 12.04 with HotSpot Java 1.7.0_21-b11. \n\nFor the past couple of days our nodes started to produce this exception every second. This exception started to get produced after we decommissioned one DC, and it continued to get produced until the nodes started producing client errors as we got alerts. No suspicious other log entries were found and no full GCs were recorded in the GC logs. The GC logs looked normal however system.log was getting filled with this exception. \n\nWe decommissioned the DC by first altering our keyspace to not replication to that DC as we are using NetworkTopologyStrategy. Then we issued nodetool decommission on each node in that DC. \n\nWhy this exception started to get produced is not known. However, restarting the nodes fixed the problem. Another observation was that the nodetool info command which we use to collect heap size statistics was not working as well and was tossing this exception:\n\nException in thread \"main\" java.lang.IllegalArgumentException: javax.management.InstanceNotFoundException: java.lang:type=Memory\n at java.lang.management.ManagementFactory.newPlatformMXBeanProxy(ManagementFactory.java:610)\n at org.apache.cassandra.tools.NodeProbe.connect(NodeProbe.java:175)\n at org.apache.cassandra.tools.NodeProbe.(NodeProbe.java:116)\n at org.apache.cassandra.tools.NodeCmd.main(NodeCmd.java:1138)\nCaused by: javax.management.InstanceNotFoundException: java.lang:type=Memory\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.getMBean(DefaultMBeanServerInterceptor.java:1095)\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.isInstanceOf(DefaultMBeanServerInterceptor.java:1401)\n at com.sun.jmx.mbeanserver.JmxMBeanServer.isInstanceOf(JmxMBeanServer.java:1082)\n at javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1492)\n at javax.management.remote.rmi.RMIConnectionImpl.access$300(RMIConnectionImpl.java:96)\n at javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1327)\n at javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1419)\n at javax.management.remote.rmi.RMIConnectionImpl.isInstanceOf(RMIConnectionImpl.java:957)\n at sun.reflect.GeneratedMethodAccessor82.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:601)\n at sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:322)\n at sun.rmi.transport.Transport$1.run(Transport.java:177)","from":"developer"},{"body":"Matt Byrd, please assign version and increase the priority as this did in fact bring our cluster to halt. ","from":"developer"},{"body":"Ryan, can your team repro?","from":"developer"},{"body":"The cluster in question was running on 1.0.7, however the code in question has remained static since well before that and doesn't look to have changed since. (though admittedly the problem could somehow be being caused elsewhere, jvm maybe?) \nI've upped the priority to major.\n\nHave you been able to reproduce? or seen the problem anywhere else?\nAny further details about your environment is set up and how you deploy may also help those trying to reproduce.\nSome common but perhaps co-incidental things about the two occurrences:\n1. virtual machines (though not both AWS)\n2. multi D.C , wouldn't have though this would be relevant but Arya does seem to see the problem after removing a d.c.\n3. Slightly old Jvm versions...\n\nI no longer have access to the cluster where we saw this previously but let me know if I can help in any other way.","from":"developer"},{"body":"To give you more details about our setup:\n\nAWS\nus-east-1 3 nodes. \nus-west-2 3 nodes.\nClassic EC2 hi1.4xlarge machines\nNetworkTopologyStrategy us-east-1:3, us-west-2:3\nEffective load on each machine is 100% about 360Gb\nWe run a 24Gb Heap with MaxTenuringThreshold = 20\n\nAverage read 4k/sec on each node\nAverage write 1.5k/sec on each node\n\nWe have a mix of SizeTiered and Leveled CFs.\nWe turned off read repair. \nWe use mmap_index_only","from":"developer"},{"body":"[~enigmacurry] - did you or your team have any luck reproducing this?\n\nIt should be trivial to throw a gc.isValid() check in the logGCResults loop and if invalid, flag to rebuild the List we're iterating across in that function as well as to introduce some exception handling for the UndeclaredThrowableException. I'm not finding much on the logic behind *when* these MXBeans can be removed from the system and I agree with Brandon that it seems incredibly odd for a JVM to punt and re-init a garbage collector MXBean on the fly with such infrequency that we haven't seen this more often.\n\nThe fact that Arya had a nodetool failure connecting to the Memory subsystem:\n{code:title=Memory failure}\nException in thread \"main\" java.lang.IllegalArgumentException: javax.management.InstanceNotFoundException: java.lang:type=Memory\nat java.lang.management.ManagementFactory.newPlatformMXBeanProxy(ManagementFactory.java:610)\nat org.apache.cassandra.tools.NodeProbe.connect(NodeProbe.java:175)\n{code}\nis concerning as it would seem to imply some deeper problems w/the JMX integration of the Memory subsystem on these JVM's given that we're querying the factory by String there rather than trying to use an invalid reference.","from":"developer"},{"body":"[~JoshuaMcKenzie] no, I haven't reproduced it. :(","from":"developer"},{"body":"Attaching a v1 that tries to gracefully check for GC MXBean validity on log and if an invalid GC is found, skip it and rebuild the list of cached MXBean's for next logging.\n\nGiven that I can't find any info on *when* these things are getting recycled, I also put in some exception handling in-case the change happens in the middle of our logging process.\n\nThis doesn't address the failure of nodetool info that Arya saw with nodetool info; that issue seems like a related but perhaps different (and less critical) effort than this ticket.","from":"developer"},{"body":"I couldn't get this to apply to the repo as of Jun 26. Can you rebase against current 2.0 branch?","from":"developer"},{"body":"rebased against cassandra-2.0","from":"developer"},{"body":"Added a comment pointing to this issue and committed.","from":"developer"}],"created":"2013-03-14T13:45:41.000+0000","description":"I am not certain this is definitely a bug, but I thought it might be worth posting to see if someone with more JVM//JMX knowledge could disprove my reasoning. Apologies if I've failed to understand something.\n\nWe've seen an intermittent problem where there is an uncaught exception in the scheduled task of logging gc results in GcInspector.java:\n\n{code}\n...\n ERROR [ScheduledTasks:1] 2013-03-08 01:09:06,335 AbstractCassandraDaemon.java (line 139) Fatal exception in thread Thread[ScheduledTasks:1,5,main]\njava.lang.reflect.UndeclaredThrowableException\n at $Proxy0.getName(Unknown Source)\n at org.apache.cassandra.service.GCInspector.logGCResults(GCInspector.java:95)\n at org.apache.cassandra.service.GCInspector.access$000(GCInspector.java:41)\n at org.apache.cassandra.service.GCInspector$1.run(GCInspector.java:85)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:441)\n at java.util.concurrent.FutureTask$Sync.innerRunAndReset(FutureTask.java:317)\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:150)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$101(ScheduledThreadPoolExecutor.java:98)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.runPeriodic(ScheduledThreadPoolExecutor.java:180)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:204)\n at java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n at java.lang.Thread.run(Thread.java:662)\nCaused by: javax.management.InstanceNotFoundException: java.lang:name=ParNew,type=GarbageCollector\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.getMBean(DefaultMBeanServerInterceptor.java:1094)\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.getAttribute(DefaultMBeanServerInterceptor.java:662)\n at com.sun.jmx.mbeanserver.JmxMBeanServer.getAttribute(JmxMBeanServer.java:638)\n at com.sun.jmx.mbeanserver.MXBeanProxy$GetHandler.invoke(MXBeanProxy.java:106)\n at com.sun.jmx.mbeanserver.MXBeanProxy.invoke(MXBeanProxy.java:148)\n at javax.management.MBeanServerInvocationHandler.invoke(MBeanServerInvocationHandler.java:248)\n ... 13 more\n...\n{code}\n\nI think the problem, may be caused by the following reasoning:\n\nIn GcInspector we populate a list of mxbeans when the GcInspector instance is instantiated:\n\n{code}\n...\nList beans = new ArrayList();\n MBeanServer server = ManagementFactory.getPlatformMBeanServer();\n try\n {\n ObjectName gcName = new ObjectName(ManagementFactory.GARBAGE_COLLECTOR_MXBEAN_DOMAIN_TYPE + \",*\");\n for (ObjectName name : server.queryNames(gcName, null))\n {\n GarbageCollectorMXBean gc = ManagementFactory.newPlatformMXBeanProxy(server, name.getCanonicalName(), GarbageCollectorMXBean.class);\n beans.add(gc);\n }\n }\n catch (Exception e)\n {\n throw new RuntimeException(e);\n }\n...\n{code}\n\nCassandra then periodically calls:\n\n{code}\n...\n private void logGCResults()\n {\n for (GarbageCollectorMXBean gc : beans)\n {\n Long previousTotal = gctimes.get(gc.getName());\n...\n{code}\n\nIn the oracle javadocs, they seem to suggest that these beans could disappear at any time.(I'm not sure why when or how this might happen)\nhttp://docs.oracle.com/javase/6/docs/api/\nSee: getGarbageCollectorMXBeans\n\n{code}\n...\npublic static List getGarbageCollectorMXBeans()\nReturns a list of GarbageCollectorMXBean objects in the Java virtual machine. The Java virtual machine may have one or more GarbageCollectorMXBean objects. It may add or remove GarbageCollectorMXBean during execution.\nReturns:\na list of GarbageCollectorMXBean objects.\n...\n{code}\n\nCorrect me if I'm wrong, but do you think this might be causing the problem? That somehow the JVM decides to remove the GarbageCollectorMXBean temporarily or permanently (causing said exception) and if this is expected behaviour, should it be handled in some way?\nAlso I'd like to point out that this may be an issue on other versions as well as I don't believe this code has changed in quite a long time.\nUnfortunately I haven't been able to reproduce this outside of the production environment, if you have any tips, questions or are able to explain//disprove my concerns, I'd be very grateful.\n\nThanks,\nMatt","issue_id":"12637020","key":"CASSANDRA-5345","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-07-18T23:15:20.000+0000","role":"fixed_distractor","summary":"Potential problem with GarbageCollectorMXBean"} {"case_id":"12639405","cluster":"DISTRACTOR-CASSANDRA-5393","comments":[{"body":"We've got an idea we're testing out here, and will hopefully post a patch in a day or so.","created":"2013-04-02T07:34:06.585+0000"},{"body":"Yuki's ticket is more comprehensive than this one","created":"2013-04-03T23:48:07.557+0000"},{"body":"At the end of the day, this is what I see happening:\n\n{code}INFO [AntiEntropyStage:1] 2013-03-27 22:48:55,390 AntiEntropyService.java (line 239) repair #80fe25a0-9730-11e2-0000-ebe7011631ff Sending completed merkle tree to /54.246.XXX.YYY for (Geo,GeoCountryMetadata)\nDEBUG [WRITE-/54.246.XXX.YYY] 2013-03-27 22:48:55,392 OutboundTcpConnection.java (line 165) error writing to ec2-54-246-XXX.YYY.eu-west-1.compute.amazonaws.com/54.246.XXX.YYY\njava.net.SocketException: Connection timed out\nat java.net.SocketOutputStream.socketWrite0(Native Method)\nat java.net.SocketOutputStream.socketWrite(SocketOutputStream.java:92)\nat java.net.SocketOutputStream.write(SocketOutputStream.java:136)\nat com.sun.net.ssl.internal.ssl.OutputRecord.writeBuffer(OutputRecord.java:358)\nat com.sun.net.ssl.internal.ssl.OutputRecord.write(OutputRecord.java:346)\nat com.sun.net.ssl.internal.ssl.SSLSocketImpl.writeRecordInternal(SSLSocketImpl.java:781)\nat com.sun.net.ssl.internal.ssl.SSLSocketImpl.writeRecord(SSLSocketImpl.java:753)\nat com.sun.net.ssl.internal.ssl.AppOutputStream.write(AppOutputStream.java:100)\nat java.io.BufferedOutputStream.flushBuffer(BufferedOutputStream.java:65)\nat java.io.BufferedOutputStream.write(BufferedOutputStream.java:104)\nat java.io.DataOutputStream.write(DataOutputStream.java:90)\nat java.io.FilterOutputStream.write(FilterOutputStream.java:80)\nat org.apache.cassandra.net.OutboundTcpConnection.write(OutboundTcpConnection.java:200)\nat org.apache.cassandra.net.OutboundTcpConnection.writeConnected(OutboundTcpConnection.java:152)\nat org.apache.cassandra.net.OutboundTcpConnection.run(OutboundTcpConnection.java:126)\n{code}\n\nThe interesting thing is the \"Connection timed out\" exception message, rather than socket reset (or something similar). So, I'm thinking this might be to keepalive timing out after the connection is broken. I was able to reproduce this exception several times by having my test cluster setup in three ec2 regions (us-west-2, us-east-1, eu-west-1 - three nodes in each), and not sending any traffic for multiple hours. Basically, I'm waiting for the connection to get dropped. Thus, when I went to triggered repair on one of the nodes (usu. starting with us-west-2), I could see where the eu-west-1 nodes would get the request to build the merkle tree, but then failed on sending the tree response with the above exception. I was able to get similar problems when trying a schema update after many hours of cluster idleness.\n\nThe attached patch catches the exception when the socket is dead (for whatever reason), and attempts a simple retry by requeueing the message at the end of the backlog queue, with the hope that the next pass will successfully recreate the socket. Note that I'm excluding MessagingService.DROPPABLE_VERBS from retries as it's OK to drop reads/mutates, but it's really those AES and other schema-related messages that I think we'd want to retry.\n\nAdmittedly this is a simple mechanism that doesn't try to do anything fancy like exponential backoff, n-levels of configurable retrys, and so on. I'm open to discussion on that, but I'm not sure how much complexity we'd want to build in for that at this point. I think an incremental improvement would go a long way here as we're currently obscuring when messages can't be sent (which is OK for DROPPABLE_VERBS, but those other ones are ones are really important), so added visibility and a retry mechanism will help. \n\n\n ","created":"2013-04-18T00:01:24.756+0000"},{"body":"v2 addresses a potential race condition between disconnecting the bad socket and re-enqueueing the failed message.","created":"2013-04-18T00:08:28.588+0000"},{"body":"+1","created":"2013-04-18T00:42:54.369+0000"},{"body":"Can we make an Entry subclass instead of saddling each Entry with an extra field that will mostly be unused?\n\nAlso, space before open paren. :)","created":"2013-04-18T01:41:08.605+0000"},{"body":"v3 includes Jonathan's suggestions. Created RetryableEntry as a subclass of Entry. Added method shouldRetry() to Entry; Entry will always return false, and RetryableEntry will check it's member boolean.\n","created":"2013-04-18T04:57:14.653+0000"},{"body":"v4 attached -- easier to show than explain what I meant in English. :)","created":"2013-04-18T14:31:44.836+0000"},{"body":"sorry for the attachment churn, decided to improve the comments too :)","created":"2013-04-18T14:38:34.437+0000"},{"body":"I think your patch and my patch are rather similar, but I'm game either way :). However, there is a small bug in Entry.shouldRetry(); you have\n\n{code}return MessagingService.DROPPABLE_VERBS.contains(message.getVerb());{code}\n\nbut should be\n\n{code}return !MessagingService.DROPPABLE_VERBS.contains(message.getVerb());{code}\n\nOtherwise we would retry the DROPPABLE_VERBS, which we want to drop.\n\nWith that small fix, lgtm.","created":"2013-04-18T17:26:03.778+0000"},{"body":"Ship it!","created":"2013-04-18T18:18:23.247+0000"},{"body":"Changed name of ticket to better describe the change.\n\nCommitted to 1.1, 1.2, and trunk.","created":"2013-04-18T20:56:01.912+0000"}],"conversations":[{"body":"Can we add an Ack/Retry around passing merle tree's around in repair? If the following fails, the repair hangs for ever on the coordinating node.\n\nhttps://github.com/apache/cassandra/blob/cassandra-1.1.10/src/java/org/apache/cassandra/service/AntiEntropyService.java#L242\n\n{noformat}\n Message message = TreeResponseVerbHandler.makeVerb(local, validator);\n if (!validator.request.endpoint.equals(FBUtilities.getBroadcastAddress()))\n logger.info(String.format(\"[repair #%s] Sending completed merkle tree to %s for %s\", validator.request.sessionid, validator.request.endpoint, validator.request.cf));\n ms.sendOneWay(message, validator.request.endpoint);\n{noformat}\n\nIf the message asking for merkle tree's gets lost, coordinating node hangs for ever as well.","from":"reporter","subject":"Add retry mechanism to OTC for non-droppable_verbs"},{"body":"We've got an idea we're testing out here, and will hopefully post a patch in a day or so.","from":"developer"},{"body":"Yuki's ticket is more comprehensive than this one","from":"developer"},{"body":"At the end of the day, this is what I see happening:\n\n{code}INFO [AntiEntropyStage:1] 2013-03-27 22:48:55,390 AntiEntropyService.java (line 239) repair #80fe25a0-9730-11e2-0000-ebe7011631ff Sending completed merkle tree to /54.246.XXX.YYY for (Geo,GeoCountryMetadata)\nDEBUG [WRITE-/54.246.XXX.YYY] 2013-03-27 22:48:55,392 OutboundTcpConnection.java (line 165) error writing to ec2-54-246-XXX.YYY.eu-west-1.compute.amazonaws.com/54.246.XXX.YYY\njava.net.SocketException: Connection timed out\nat java.net.SocketOutputStream.socketWrite0(Native Method)\nat java.net.SocketOutputStream.socketWrite(SocketOutputStream.java:92)\nat java.net.SocketOutputStream.write(SocketOutputStream.java:136)\nat com.sun.net.ssl.internal.ssl.OutputRecord.writeBuffer(OutputRecord.java:358)\nat com.sun.net.ssl.internal.ssl.OutputRecord.write(OutputRecord.java:346)\nat com.sun.net.ssl.internal.ssl.SSLSocketImpl.writeRecordInternal(SSLSocketImpl.java:781)\nat com.sun.net.ssl.internal.ssl.SSLSocketImpl.writeRecord(SSLSocketImpl.java:753)\nat com.sun.net.ssl.internal.ssl.AppOutputStream.write(AppOutputStream.java:100)\nat java.io.BufferedOutputStream.flushBuffer(BufferedOutputStream.java:65)\nat java.io.BufferedOutputStream.write(BufferedOutputStream.java:104)\nat java.io.DataOutputStream.write(DataOutputStream.java:90)\nat java.io.FilterOutputStream.write(FilterOutputStream.java:80)\nat org.apache.cassandra.net.OutboundTcpConnection.write(OutboundTcpConnection.java:200)\nat org.apache.cassandra.net.OutboundTcpConnection.writeConnected(OutboundTcpConnection.java:152)\nat org.apache.cassandra.net.OutboundTcpConnection.run(OutboundTcpConnection.java:126)\n{code}\n\nThe interesting thing is the \"Connection timed out\" exception message, rather than socket reset (or something similar). So, I'm thinking this might be to keepalive timing out after the connection is broken. I was able to reproduce this exception several times by having my test cluster setup in three ec2 regions (us-west-2, us-east-1, eu-west-1 - three nodes in each), and not sending any traffic for multiple hours. Basically, I'm waiting for the connection to get dropped. Thus, when I went to triggered repair on one of the nodes (usu. starting with us-west-2), I could see where the eu-west-1 nodes would get the request to build the merkle tree, but then failed on sending the tree response with the above exception. I was able to get similar problems when trying a schema update after many hours of cluster idleness.\n\nThe attached patch catches the exception when the socket is dead (for whatever reason), and attempts a simple retry by requeueing the message at the end of the backlog queue, with the hope that the next pass will successfully recreate the socket. Note that I'm excluding MessagingService.DROPPABLE_VERBS from retries as it's OK to drop reads/mutates, but it's really those AES and other schema-related messages that I think we'd want to retry.\n\nAdmittedly this is a simple mechanism that doesn't try to do anything fancy like exponential backoff, n-levels of configurable retrys, and so on. I'm open to discussion on that, but I'm not sure how much complexity we'd want to build in for that at this point. I think an incremental improvement would go a long way here as we're currently obscuring when messages can't be sent (which is OK for DROPPABLE_VERBS, but those other ones are ones are really important), so added visibility and a retry mechanism will help. \n\n\n ","from":"developer"},{"body":"v2 addresses a potential race condition between disconnecting the bad socket and re-enqueueing the failed message.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Can we make an Entry subclass instead of saddling each Entry with an extra field that will mostly be unused?\n\nAlso, space before open paren. :)","from":"developer"},{"body":"v3 includes Jonathan's suggestions. Created RetryableEntry as a subclass of Entry. Added method shouldRetry() to Entry; Entry will always return false, and RetryableEntry will check it's member boolean.\n","from":"developer"},{"body":"v4 attached -- easier to show than explain what I meant in English. :)","from":"developer"},{"body":"sorry for the attachment churn, decided to improve the comments too :)","from":"developer"},{"body":"I think your patch and my patch are rather similar, but I'm game either way :). However, there is a small bug in Entry.shouldRetry(); you have\n\n{code}return MessagingService.DROPPABLE_VERBS.contains(message.getVerb());{code}\n\nbut should be\n\n{code}return !MessagingService.DROPPABLE_VERBS.contains(message.getVerb());{code}\n\nOtherwise we would retry the DROPPABLE_VERBS, which we want to drop.\n\nWith that small fix, lgtm.","from":"developer"},{"body":"Ship it!","from":"developer"},{"body":"Changed name of ticket to better describe the change.\n\nCommitted to 1.1, 1.2, and trunk.","from":"developer"}],"created":"2013-03-27T17:59:48.000+0000","description":"Can we add an Ack/Retry around passing merle tree's around in repair? If the following fails, the repair hangs for ever on the coordinating node.\n\nhttps://github.com/apache/cassandra/blob/cassandra-1.1.10/src/java/org/apache/cassandra/service/AntiEntropyService.java#L242\n\n{noformat}\n Message message = TreeResponseVerbHandler.makeVerb(local, validator);\n if (!validator.request.endpoint.equals(FBUtilities.getBroadcastAddress()))\n logger.info(String.format(\"[repair #%s] Sending completed merkle tree to %s for %s\", validator.request.sessionid, validator.request.endpoint, validator.request.cf));\n ms.sendOneWay(message, validator.request.endpoint);\n{noformat}\n\nIf the message asking for merkle tree's gets lost, coordinating node hangs for ever as well.","issue_id":"12639405","key":"CASSANDRA-5393","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-04-18T20:56:25.000+0000","role":"fixed_distractor","summary":"Add retry mechanism to OTC for non-droppable_verbs"} {"case_id":"12643918","cluster":"DISTRACTOR-CASSANDRA-5504","comments":[{"body":"I believe this issue also affects 1.1.11 and 1.2.4, I tried both had the same issue as on http://cassandra-user-incubator-apache-org.3065146.n2.nabble.com/Thrift-message-length-exceeded-td7587006.html and http://stackoverflow.com/questions/15487540/pig-cassandra-message-length-exceeded\n\nI then tried 1.1.9 (skipped .10) and it worked correctly","created":"2013-04-22T15:47:15.762+0000"},{"body":"patch doesn't fix my issue still get:\n\n{quote}\njava.lang.RuntimeException: org.apache.thrift.TException: Message length exceeded: 21\n\tat org.apache.cassandra.hadoop.ColumnFamilyRecordReader$StaticRowIterator.maybeInit(ColumnFamilyRecordReader.java:384)\n\tat org.apache.cassandra.hadoop.ColumnFamilyRecordReader$StaticRowIterator.computeNext(ColumnFamilyRecordReader.java:390)\n\tat org.apache.cassandra.hadoop.ColumnFamilyRecordReader$StaticRowIterator.computeNext(ColumnFamilyRecordReader.java:313)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n\tat org.apache.cassandra.hadoop.ColumnFamilyRecordReader.getProgress(ColumnFamilyRecordReader.java:103)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigRecordReader.getProgress(PigRecordReader.java:158)\n\tat org.apache.hadoop.mapred.MapTask$NewTrackingRecordReader.getProgress(MapTask.java:514)\n\tat org.apache.hadoop.mapred.MapTask$NewTrackingRecordReader.nextKeyValue(MapTask.java:539)\n\tat org.apache.hadoop.mapreduce.MapContext.nextKeyValue(MapContext.java:67)\n\tat org.apache.hadoop.mapreduce.Mapper.run(Mapper.java:143)\n\tat org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:764)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:370)\n\tat org.apache.hadoop.mapred.Child$4.run(Child.java:255)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Unknown Source)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1121)\n\tat org.apache.hadoop.mapred.Child.main(Child.java:249)\nCaused by: org.apache.thrift.TException: Message length exceeded: 21\n\tat org.apache.thrift.protocol.TBinaryProtocol.checkReadLength(TBinaryProtocol.java:393)\n\tat org.apache.thrift.protocol.TBinaryProtocol.readBinary(TBinaryProtocol.java:363)\n\tat org.apache.cassandra.thrift.Column.read(Column.java:528)\n\tat org.apache.cassandra.thrift.ColumnOrSuperColumn.read(ColumnOrSuperColumn.java:507)\n\tat org.apache.cassandra.thrift.KeySlice.read(KeySlice.java:408)\n\tat org.apache.cassandra.thrift.Cassandra$get_range_slices_result.read(Cassandra.java:12905)\n\tat org.apache.thrift.TServiceClient.receiveBase(TServiceClient.java:78)\n\tat org.apache.cassandra.thrift.Cassandra$Client.recv_get_range_slices(Cassandra.java:734)\n\tat org.apache.cassandra.thrift.Cassandra$Client.get_range_slices(Cassandra.java:718)\n\tat org.apache.cassandra.hadoop.ColumnFamilyRecordReader$StaticRowIterator.maybeInit(ColumnFamilyRecordReader.java:346)\n{quote}\n\nMaybe these are separate issues","created":"2013-04-22T16:48:05.148+0000"},{"body":"Sorry, my bad, I have forgotten to sync changes for static rows, give me a sec.\nFaced that one, too.\n","created":"2013-04-22T16:56:06.950+0000"},{"body":"Ben, I've been putting lastRowKey in a wrong place, now it's after get_range_slices occuring, which should be correct.\n\nI've had same exact stack trace.\nHope that solves issue for you.","created":"2013-04-22T17:11:05.784+0000"},{"body":"still not working for me. I can't really tell what is going on due to the magic thrift import line \n\nbq. import org.apache.cassandra.thrift.*;","created":"2013-04-22T17:59:32.179+0000"},{"body":"Hm... That's quite weird. \nMaybe there's something different with Thrift.\n\nI've seen people having trouble because of the message size, too, though. There're two settings, one for framed and one non-framed thrift.","created":"2013-04-22T18:15:01.284+0000"},{"body":"relevant changes since 1.1.9 https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=blobdiff;f=src/java/org/apache/cassandra/hadoop/ColumnFamilyRecordReader.java;h=dfeacc39a6ba49d7eaf6336251875e5b55528bb0;hp=a40e6c56c1fd0b52c482832d04e73387db001699;hb=4feb87d37544b9fde722786555475f2f790059ca;hpb=73d828e4e8023b9f7ca8fafd12becec34eb59211","created":"2013-04-22T18:18:47.959+0000"},{"body":"actually the diff between 1.1.11 and 1.1.9 is very simple \n\nbq. git diff cassandra-1.1.9 cassandra-1.1.11 -- src/java/org/apache/cassandra/hadoop/ColumnFamilyRecordReader.java\n\n{code}\n- TTransport transport = ConfigHelper.getInputTransportFactory(conf).openTransport(socket);\n- TBinaryProtocol binaryProtocol = new TBinaryProtocol(transport);\n+ TTransport transport = ConfigHelper.getInputTransportFactory(conf).openTransport(socket, conf);\n+ TBinaryProtocol binaryProtocol = new TBinaryProtocol(transport, ConfigHelper.getThriftMaxMessageLength(conf));\n{code}\n\nthat might be the difference","created":"2013-04-22T18:31:18.923+0000"},{"body":"reverting that change to TBinaryProtocol fixes it, now the question is do I just have something setup wrong since ConfigHelper.getThriftMaxMessageLength obviously returns some really low value","created":"2013-04-22T18:36:14.011+0000"},{"body":"You can configure it through the ConfigHelper.setThriftMaxMessageLength(), can you execute java code there? (e.q. not only pig queries)","created":"2013-04-22T18:52:11.239+0000"},{"body":"However, unfortunately, when using it with Cascading, I still get eternal iterations :/ so it'd still be good if someone could take a look at the patch :/","created":"2013-04-22T18:56:37.449+0000"},{"body":"yeah but i'm using pig, which means I shouldn't have to redo the whole hadoop backend to use pig","created":"2013-04-22T19:00:09.706+0000"},{"body":"I might be seeing the same bug you originally reported now that the job has started. I can't help but wonder if the reason none of this works is because the main codebase is now in DSE and its all modified to work with DSE and not with hadoop. ","created":"2013-04-22T19:10:25.819+0000"},{"body":"Instead of reverting the changes to TBinaryProtocol you probably need to use ConfigHelper to set the thrift_framed-transport_size_in_mb and thrift_max_message_length_in_mb to much larger values (if ConfigHelper is exposed for you). These values, prior to 1.10, were ignored (and a later version fixed a bug with getting them from ConfigHelper as well). Setting the values to 2047 and 2048 respectively got us working again.\n\nOleksandr -- patch2 works for us. Thanks!","created":"2013-04-24T19:48:10.154+0000"},{"body":"Thanks for the patch, Oleksandr.\n\nIt looks to me like the root of the problem is that {{key.put(this.getCurrentKey())}} destructively modifies currentKey. Attached is a patch to duplicate the buffer first.\n\nThis has the added benefit that we don't have to impose any overhead on the new mapreduce api to solve this problem in the old mapred one.","created":"2013-04-25T22:47:58.614+0000"},{"body":"While investigating whether this was also a problem in 1.1, I found that this was fixed for 1.1.7 in CASSANDRA-4834, with the same .duplicate() solution, but not merged forward. I've applied this fix to the 1.2 branch.","created":"2013-04-25T23:00:12.642+0000"},{"body":"Can anyone confirm if 1.2.4 contains the fix, too?\nIt seems to work, it's just not clear wether fix made it there or it's just a coincidence...\n\nUPDATE: sorry, I've tested against 1.2.5, so nevermind :)","created":"2013-06-29T20:55:00.280+0000"}],"conversations":[{"body":"Currently, when using newer hadoop versions, due to the call to \n\nnext(ByteBuffer key, SortedMap value)\n\nwithin ColumnFamilyRecordReader, because `key.clear();` is called, key is emptied. That causes the StaticRowIterator and WideRowIterator to glitch, namely, when Iterables.getLast(rows).key is called, key is already empty. This will cause Hadoop to request the same range again and again all the time.\n\nPlease see the attached patch/diff, it simply adds lastRowKey (ByteBuffer) and saves it for the next iteration along with all the rows, this allows query for the next range to be fully correct.\n\nThis patch is branched from 1.2.3 version.\n\nTested against Cassandra 1.2.3, with Hadoop 1.0.3, 1.0.4 and 0.20.2","from":"reporter","subject":"Eternal iteration when using older hadoop version due to next() call and empty key value"},{"body":"I believe this issue also affects 1.1.11 and 1.2.4, I tried both had the same issue as on http://cassandra-user-incubator-apache-org.3065146.n2.nabble.com/Thrift-message-length-exceeded-td7587006.html and http://stackoverflow.com/questions/15487540/pig-cassandra-message-length-exceeded\n\nI then tried 1.1.9 (skipped .10) and it worked correctly","from":"developer"},{"body":"patch doesn't fix my issue still get:\n\n{quote}\njava.lang.RuntimeException: org.apache.thrift.TException: Message length exceeded: 21\n\tat org.apache.cassandra.hadoop.ColumnFamilyRecordReader$StaticRowIterator.maybeInit(ColumnFamilyRecordReader.java:384)\n\tat org.apache.cassandra.hadoop.ColumnFamilyRecordReader$StaticRowIterator.computeNext(ColumnFamilyRecordReader.java:390)\n\tat org.apache.cassandra.hadoop.ColumnFamilyRecordReader$StaticRowIterator.computeNext(ColumnFamilyRecordReader.java:313)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n\tat org.apache.cassandra.hadoop.ColumnFamilyRecordReader.getProgress(ColumnFamilyRecordReader.java:103)\n\tat org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.PigRecordReader.getProgress(PigRecordReader.java:158)\n\tat org.apache.hadoop.mapred.MapTask$NewTrackingRecordReader.getProgress(MapTask.java:514)\n\tat org.apache.hadoop.mapred.MapTask$NewTrackingRecordReader.nextKeyValue(MapTask.java:539)\n\tat org.apache.hadoop.mapreduce.MapContext.nextKeyValue(MapContext.java:67)\n\tat org.apache.hadoop.mapreduce.Mapper.run(Mapper.java:143)\n\tat org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:764)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:370)\n\tat org.apache.hadoop.mapred.Child$4.run(Child.java:255)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Unknown Source)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1121)\n\tat org.apache.hadoop.mapred.Child.main(Child.java:249)\nCaused by: org.apache.thrift.TException: Message length exceeded: 21\n\tat org.apache.thrift.protocol.TBinaryProtocol.checkReadLength(TBinaryProtocol.java:393)\n\tat org.apache.thrift.protocol.TBinaryProtocol.readBinary(TBinaryProtocol.java:363)\n\tat org.apache.cassandra.thrift.Column.read(Column.java:528)\n\tat org.apache.cassandra.thrift.ColumnOrSuperColumn.read(ColumnOrSuperColumn.java:507)\n\tat org.apache.cassandra.thrift.KeySlice.read(KeySlice.java:408)\n\tat org.apache.cassandra.thrift.Cassandra$get_range_slices_result.read(Cassandra.java:12905)\n\tat org.apache.thrift.TServiceClient.receiveBase(TServiceClient.java:78)\n\tat org.apache.cassandra.thrift.Cassandra$Client.recv_get_range_slices(Cassandra.java:734)\n\tat org.apache.cassandra.thrift.Cassandra$Client.get_range_slices(Cassandra.java:718)\n\tat org.apache.cassandra.hadoop.ColumnFamilyRecordReader$StaticRowIterator.maybeInit(ColumnFamilyRecordReader.java:346)\n{quote}\n\nMaybe these are separate issues","from":"developer"},{"body":"Sorry, my bad, I have forgotten to sync changes for static rows, give me a sec.\nFaced that one, too.\n","from":"developer"},{"body":"Ben, I've been putting lastRowKey in a wrong place, now it's after get_range_slices occuring, which should be correct.\n\nI've had same exact stack trace.\nHope that solves issue for you.","from":"developer"},{"body":"still not working for me. I can't really tell what is going on due to the magic thrift import line \n\nbq. import org.apache.cassandra.thrift.*;","from":"developer"},{"body":"Hm... That's quite weird. \nMaybe there's something different with Thrift.\n\nI've seen people having trouble because of the message size, too, though. There're two settings, one for framed and one non-framed thrift.","from":"developer"},{"body":"relevant changes since 1.1.9 https://git-wip-us.apache.org/repos/asf?p=cassandra.git;a=blobdiff;f=src/java/org/apache/cassandra/hadoop/ColumnFamilyRecordReader.java;h=dfeacc39a6ba49d7eaf6336251875e5b55528bb0;hp=a40e6c56c1fd0b52c482832d04e73387db001699;hb=4feb87d37544b9fde722786555475f2f790059ca;hpb=73d828e4e8023b9f7ca8fafd12becec34eb59211","from":"developer"},{"body":"actually the diff between 1.1.11 and 1.1.9 is very simple \n\nbq. git diff cassandra-1.1.9 cassandra-1.1.11 -- src/java/org/apache/cassandra/hadoop/ColumnFamilyRecordReader.java\n\n{code}\n- TTransport transport = ConfigHelper.getInputTransportFactory(conf).openTransport(socket);\n- TBinaryProtocol binaryProtocol = new TBinaryProtocol(transport);\n+ TTransport transport = ConfigHelper.getInputTransportFactory(conf).openTransport(socket, conf);\n+ TBinaryProtocol binaryProtocol = new TBinaryProtocol(transport, ConfigHelper.getThriftMaxMessageLength(conf));\n{code}\n\nthat might be the difference","from":"developer"},{"body":"reverting that change to TBinaryProtocol fixes it, now the question is do I just have something setup wrong since ConfigHelper.getThriftMaxMessageLength obviously returns some really low value","from":"developer"},{"body":"You can configure it through the ConfigHelper.setThriftMaxMessageLength(), can you execute java code there? (e.q. not only pig queries)","from":"developer"},{"body":"However, unfortunately, when using it with Cascading, I still get eternal iterations :/ so it'd still be good if someone could take a look at the patch :/","from":"developer"},{"body":"yeah but i'm using pig, which means I shouldn't have to redo the whole hadoop backend to use pig","from":"developer"},{"body":"I might be seeing the same bug you originally reported now that the job has started. I can't help but wonder if the reason none of this works is because the main codebase is now in DSE and its all modified to work with DSE and not with hadoop. ","from":"developer"},{"body":"Instead of reverting the changes to TBinaryProtocol you probably need to use ConfigHelper to set the thrift_framed-transport_size_in_mb and thrift_max_message_length_in_mb to much larger values (if ConfigHelper is exposed for you). These values, prior to 1.10, were ignored (and a later version fixed a bug with getting them from ConfigHelper as well). Setting the values to 2047 and 2048 respectively got us working again.\n\nOleksandr -- patch2 works for us. Thanks!","from":"developer"},{"body":"Thanks for the patch, Oleksandr.\n\nIt looks to me like the root of the problem is that {{key.put(this.getCurrentKey())}} destructively modifies currentKey. Attached is a patch to duplicate the buffer first.\n\nThis has the added benefit that we don't have to impose any overhead on the new mapreduce api to solve this problem in the old mapred one.","from":"developer"},{"body":"While investigating whether this was also a problem in 1.1, I found that this was fixed for 1.1.7 in CASSANDRA-4834, with the same .duplicate() solution, but not merged forward. I've applied this fix to the 1.2 branch.","from":"developer"},{"body":"Can anyone confirm if 1.2.4 contains the fix, too?\nIt seems to work, it's just not clear wether fix made it there or it's just a coincidence...\n\nUPDATE: sorry, I've tested against 1.2.5, so nevermind :)","from":"developer"}],"created":"2013-04-22T11:50:35.000+0000","description":"Currently, when using newer hadoop versions, due to the call to \n\nnext(ByteBuffer key, SortedMap value)\n\nwithin ColumnFamilyRecordReader, because `key.clear();` is called, key is emptied. That causes the StaticRowIterator and WideRowIterator to glitch, namely, when Iterables.getLast(rows).key is called, key is already empty. This will cause Hadoop to request the same range again and again all the time.\n\nPlease see the attached patch/diff, it simply adds lastRowKey (ByteBuffer) and saves it for the next iteration along with all the rows, this allows query for the next range to be fully correct.\n\nThis patch is branched from 1.2.3 version.\n\nTested against Cassandra 1.2.3, with Hadoop 1.0.3, 1.0.4 and 0.20.2","issue_id":"12643918","key":"CASSANDRA-5504","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-04-25T23:02:28.000+0000","role":"fixed_distractor","summary":"Eternal iteration when using older hadoop version due to next() call and empty key value"} {"case_id":"12646363","cluster":"DISTRACTOR-CASSANDRA-5544","comments":[{"body":"same issue with Cassandra 1.2.3. I've tested with both RandomPartitioner and Murmur3Partitioner","created":"2013-05-07T07:57:12.032+0000"},{"body":"For more information, here is some threads from mail archive\n1) http://www.mail-archive.com/user@cassandra.apache.org/msg29663.html\n2) http://www.mail-archive.com/user@cassandra.apache.org/msg28016.html\n3) http://www.mail-archive.com/user@cassandra.apache.org/msg29425.html","created":"2013-05-07T08:54:59.314+0000"},{"body":"Does 1.1.11 have the same problem?","created":"2013-05-07T19:45:11.918+0000"},{"body":"Cassandra version 1.1.11 have no such problem. I have test in single node cluster and it's created 15 map. \nSee attach please.","created":"2013-05-26T12:56:59.849+0000"},{"body":"So something goes wrong with 1.2.x version","created":"2013-05-27T08:30:56.408+0000"},{"body":"Can you take a look, Alex? Nothing changed in pig as far I know.","created":"2013-05-27T17:27:40.950+0000"},{"body":"[~shamim] How many splits do you get for each hadoop node? You can set ConfigHelper.setInputSplitSize to a smaller number to get more mappers for your pig job. The existing CassandraStorage class doesn't set it, so it uses the defualt value of 64k. So if your nodes has less than 64k rows, it will have only one mapper.","created":"2013-05-27T21:42:48.099+0000"},{"body":"Some changes had been made to CassandraColumnInputFormat class since 1.1.5\n\ne.g.\nadd describe_splits_ex providing improved split size estimate\npatch by Piotr Kolaczkowski; reviewed by jbellis for CASSANDRA-4803","created":"2013-05-27T21:44:39.326+0000"},{"body":"[~alexliu68] I did some tests with more than 64k row and had only one mapper for the whole cluster. Even if we have less than 64k rows, why don't we have at least one mapper per node (in my case replication_factor=1) to work on rows using data locality. Vnodes are enabled on my cluster, can there be a relation with this option ?","created":"2013-05-27T22:08:00.634+0000"},{"body":"Yes, if vnode is enale, it creates a lot of smaller splits (which is not preferred, we will fix the vnode hadoop too many small splits issue later), so can you test it with vnode disable.","created":"2013-05-27T22:13:09.738+0000"},{"body":"But if there are many small splits it doesn't mean that we should have more mappers ? I'm saying that cause you propose to [~shamim_ru] to decrease ConfigHelper.setInputSplitSize exactly for that, right ?\nI need one more day to test without vnodes.","created":"2013-05-27T22:19:42.390+0000"},{"body":"Current implementation only matches one mapper to a split. Existing code doesn't set InputSplitSize (which means we can't change it to a smaller number unless we change the code at setLocation method to do it), so we need more than 64k rows to have more than one mapper per node.\n\nFor vnode we need to support a virtual split which combines multiple small splits. ","created":"2013-05-27T22:31:29.132+0000"},{"body":"okay. I'll test without vnodes and give you a feedback except if [~shamim_ru] confirms that he didn't use vnodes, which I suppose as he upgraded from C* 1.1.5 to 1.2.1","created":"2013-05-27T22:41:47.755+0000"},{"body":"[~alexliu68]\n1) I am using pig and actually don't know how many split i had (i am very curious to know how to calculate the split count). However i have had more than 30 million rows.\n2) I didn't use VNODES.\n3) SET mapred.min.split.size 12500000; \nSET mapred.max.split.size 12500000;\n doesn't help at all\n4) SET pig.noSplitCombination true; - did some magic trick, we got more than 100 maps but 2 of them (always two maps) got very large Map input records and runs more than hours. \n5) Observe one very interesting thing when used SET pig.noSplitCombination true, a lot of maps created with \t\nMap input records \t0 \n","created":"2013-05-28T07:04:54.611+0000"},{"body":"To get the splits for the node, call thrift API client.describe_splits_ex(cfName, range.start_token, range.end_token, splitsize) it returns the split for that node.\n\nwhere range.start_token and range.end_token is the start and end token of the node, and splitsize is 64 *1024","created":"2013-05-28T16:32:07.448+0000"},{"body":"[~shamim] I think you already found the answer, SET pig.noSplitCombination true, so Pig doesn't combine the small splits into one mapper. HBase internal code does it as well. I found that C*-1.2.1 update Pig from 0.9.0 version to 0.10.0 version which may cause the behavior changes.\n\nAs far as number 4) and number 5) concerns, I think the empty maps/big maps are due to data skewness. If you can first print out the splits, then you can check the rows for each split.\n\nI will add the following code to CassandraStorage.java\n\njob.getConfiguration().setBoolean(\"pig.noSplitCombination\", true);","created":"2013-05-28T17:04:12.379+0000"},{"body":"I attached the patch.","created":"2013-05-28T17:14:41.384+0000"},{"body":"AFAIK split combination is used to improve performance. Doesn't it mean the same for cassandra ?\nAnd if performance decreases without split combination, will the performance decrease much more with vnodes ?","created":"2013-05-28T19:46:25.022+0000"},{"body":"CassandraColumnInputFormat define the split size, so we don't want Pig to override it by combining splits. We can always tune the split size to tune the performance. Next step, we can open up a little bit so that Pig user can specify split size configuration.\n\nVnode hadoop performance generally decreases, we can do the split combination at Cassandra side to improve the performance, which could be another ticket.","created":"2013-05-28T20:20:01.217+0000"},{"body":"Version 2 patch is attached. It allows user to define PIG_INPUT_SPLIT_SIZE in the system env","created":"2013-05-28T20:45:48.661+0000"},{"body":"Alex, thank you very much for your quick response. \nHowever, i am afraid that above patch will not solve the problem i described \"we got more than 100 maps but 2 of them (always two maps) got very large Map input records and runs more than hours - point 4\" - this behavior is unexpected. This means Map input records is not evenly through cluster, most of the maps getting Map input records = 10000 but only two of them getting more than millions.\nCertainly i will do some test through thrift api as you described. \nOne more things, would you kindly allows user to define PIG_INPUT_SPLIT_SIZE through cassandra store URL as \"STORE updated INTO 'cassandra://KEYSPACE/CF?allow_deletes=true&PIG_INPUT_SPLIT_SIZE=xxxxxx' USING CassandraStorage()\" instead of system environment.","created":"2013-05-29T07:06:13.707+0000"},{"body":"Or maybe via a SET PIG_INPUT_SPLIT_SIZE in Pig script ?\n[~alexliu68] I open the second ticket to improve performance with vnodes except if you prefer to open it, which could be better :)","created":"2013-05-29T14:10:34.657+0000"},{"body":"Version 3 is attached. I add split_size as a parameter.","created":"2013-05-29T17:06:02.380+0000"},{"body":"[~cscetbon] please open it, someone else may already open it.","created":"2013-05-29T17:08:47.581+0000"},{"body":"Committed, with an update to the README to document split_size.","created":"2013-05-29T17:55:41.442+0000"},{"body":"Version 4 is attached, it removes getting split size as system env","created":"2013-05-30T18:10:11.899+0000"},{"body":"I'm fine with leaving that in for now.","created":"2013-05-31T12:46:24.198+0000"},{"body":"I have a plan to do some test in weekend ","created":"2013-05-31T13:36:57.285+0000"},{"body":"My tests confirm that I have multiple mappers (1025) and each mapper works on a range of my column family [http://pastebin.com/vL3uC5Ca]. Good job !","created":"2013-06-05T13:13:39.985+0000"},{"body":"did you run map reduce job through Pig?","created":"2013-06-05T13:44:50.523+0000"},{"body":"Yes. I used Pig 0.11.1, Hadoop 1.1.2 (as newer versions are not supported [CASSANDRA-5201|https://issues.apache.org/jira/browse/CASSANDRA-5201]) and cassandra 1.2.3 (I added the current patch from git commits and built sources) ","created":"2013-06-05T13:47:26.556+0000"},{"body":"At last , i could manage a few hours to try the fix. Definitely it's working now, every mapper works on their own range, however i have test in single node cluster with Hadoop 1.1.2 + Pig 0.11.1 and Cassandra 1.2.6. Thankx. ","created":"2013-06-28T07:27:57.720+0000"}],"conversations":[{"body":"We have got very strange beheviour of hadoop cluster after upgrading \nCassandra from 1.1.5 to Cassandra 1.2.1. We have 5 nodes cluster of Cassandra, where three of them are hodoop slaves. Now when we are submitting job through Pig script, only one map assigns in task running on one of the hadoop slaves regardless of \nvolume of data (already tried with more than million rows).\nConfigure of pig as follows:\nexport PIG_HOME=/oracle/pig-0.10.0\nexport PIG_CONF_DIR=${HADOOP_HOME}/conf\nexport PIG_INITIAL_ADDRESS=192.168.157.103\nexport PIG_RPC_PORT=9160\nexport PIG_PARTITIONER=org.apache.cassandra.dht.Murmur3Partitioner\n\n\nAlso we have these following properties in hadoop:\n \n mapred.tasktracker.map.tasks.maximum\n 10\n \n \n mapred.map.tasks\n 4\n ","from":"reporter","subject":"Hadoop jobs assigns only one mapper in task"},{"body":"same issue with Cassandra 1.2.3. I've tested with both RandomPartitioner and Murmur3Partitioner","from":"developer"},{"body":"For more information, here is some threads from mail archive\n1) http://www.mail-archive.com/user@cassandra.apache.org/msg29663.html\n2) http://www.mail-archive.com/user@cassandra.apache.org/msg28016.html\n3) http://www.mail-archive.com/user@cassandra.apache.org/msg29425.html","from":"developer"},{"body":"Does 1.1.11 have the same problem?","from":"developer"},{"body":"Cassandra version 1.1.11 have no such problem. I have test in single node cluster and it's created 15 map. \nSee attach please.","from":"developer"},{"body":"So something goes wrong with 1.2.x version","from":"developer"},{"body":"Can you take a look, Alex? Nothing changed in pig as far I know.","from":"developer"},{"body":"[~shamim] How many splits do you get for each hadoop node? You can set ConfigHelper.setInputSplitSize to a smaller number to get more mappers for your pig job. The existing CassandraStorage class doesn't set it, so it uses the defualt value of 64k. So if your nodes has less than 64k rows, it will have only one mapper.","from":"developer"},{"body":"Some changes had been made to CassandraColumnInputFormat class since 1.1.5\n\ne.g.\nadd describe_splits_ex providing improved split size estimate\npatch by Piotr Kolaczkowski; reviewed by jbellis for CASSANDRA-4803","from":"developer"},{"body":"[~alexliu68] I did some tests with more than 64k row and had only one mapper for the whole cluster. Even if we have less than 64k rows, why don't we have at least one mapper per node (in my case replication_factor=1) to work on rows using data locality. Vnodes are enabled on my cluster, can there be a relation with this option ?","from":"developer"},{"body":"Yes, if vnode is enale, it creates a lot of smaller splits (which is not preferred, we will fix the vnode hadoop too many small splits issue later), so can you test it with vnode disable.","from":"developer"},{"body":"But if there are many small splits it doesn't mean that we should have more mappers ? I'm saying that cause you propose to [~shamim_ru] to decrease ConfigHelper.setInputSplitSize exactly for that, right ?\nI need one more day to test without vnodes.","from":"developer"},{"body":"Current implementation only matches one mapper to a split. Existing code doesn't set InputSplitSize (which means we can't change it to a smaller number unless we change the code at setLocation method to do it), so we need more than 64k rows to have more than one mapper per node.\n\nFor vnode we need to support a virtual split which combines multiple small splits. ","from":"developer"},{"body":"okay. I'll test without vnodes and give you a feedback except if [~shamim_ru] confirms that he didn't use vnodes, which I suppose as he upgraded from C* 1.1.5 to 1.2.1","from":"developer"},{"body":"[~alexliu68]\n1) I am using pig and actually don't know how many split i had (i am very curious to know how to calculate the split count). However i have had more than 30 million rows.\n2) I didn't use VNODES.\n3) SET mapred.min.split.size 12500000; \nSET mapred.max.split.size 12500000;\n doesn't help at all\n4) SET pig.noSplitCombination true; - did some magic trick, we got more than 100 maps but 2 of them (always two maps) got very large Map input records and runs more than hours. \n5) Observe one very interesting thing when used SET pig.noSplitCombination true, a lot of maps created with \t\nMap input records \t0 \n","from":"developer"},{"body":"To get the splits for the node, call thrift API client.describe_splits_ex(cfName, range.start_token, range.end_token, splitsize) it returns the split for that node.\n\nwhere range.start_token and range.end_token is the start and end token of the node, and splitsize is 64 *1024","from":"developer"},{"body":"[~shamim] I think you already found the answer, SET pig.noSplitCombination true, so Pig doesn't combine the small splits into one mapper. HBase internal code does it as well. I found that C*-1.2.1 update Pig from 0.9.0 version to 0.10.0 version which may cause the behavior changes.\n\nAs far as number 4) and number 5) concerns, I think the empty maps/big maps are due to data skewness. If you can first print out the splits, then you can check the rows for each split.\n\nI will add the following code to CassandraStorage.java\n\njob.getConfiguration().setBoolean(\"pig.noSplitCombination\", true);","from":"developer"},{"body":"I attached the patch.","from":"developer"},{"body":"AFAIK split combination is used to improve performance. Doesn't it mean the same for cassandra ?\nAnd if performance decreases without split combination, will the performance decrease much more with vnodes ?","from":"developer"},{"body":"CassandraColumnInputFormat define the split size, so we don't want Pig to override it by combining splits. We can always tune the split size to tune the performance. Next step, we can open up a little bit so that Pig user can specify split size configuration.\n\nVnode hadoop performance generally decreases, we can do the split combination at Cassandra side to improve the performance, which could be another ticket.","from":"developer"},{"body":"Version 2 patch is attached. It allows user to define PIG_INPUT_SPLIT_SIZE in the system env","from":"developer"},{"body":"Alex, thank you very much for your quick response. \nHowever, i am afraid that above patch will not solve the problem i described \"we got more than 100 maps but 2 of them (always two maps) got very large Map input records and runs more than hours - point 4\" - this behavior is unexpected. This means Map input records is not evenly through cluster, most of the maps getting Map input records = 10000 but only two of them getting more than millions.\nCertainly i will do some test through thrift api as you described. \nOne more things, would you kindly allows user to define PIG_INPUT_SPLIT_SIZE through cassandra store URL as \"STORE updated INTO 'cassandra://KEYSPACE/CF?allow_deletes=true&PIG_INPUT_SPLIT_SIZE=xxxxxx' USING CassandraStorage()\" instead of system environment.","from":"developer"},{"body":"Or maybe via a SET PIG_INPUT_SPLIT_SIZE in Pig script ?\n[~alexliu68] I open the second ticket to improve performance with vnodes except if you prefer to open it, which could be better :)","from":"developer"},{"body":"Version 3 is attached. I add split_size as a parameter.","from":"developer"},{"body":"[~cscetbon] please open it, someone else may already open it.","from":"developer"},{"body":"Committed, with an update to the README to document split_size.","from":"developer"},{"body":"Version 4 is attached, it removes getting split size as system env","from":"developer"},{"body":"I'm fine with leaving that in for now.","from":"developer"},{"body":"I have a plan to do some test in weekend ","from":"developer"},{"body":"My tests confirm that I have multiple mappers (1025) and each mapper works on a range of my column family [http://pastebin.com/vL3uC5Ca]. Good job !","from":"developer"},{"body":"did you run map reduce job through Pig?","from":"developer"},{"body":"Yes. I used Pig 0.11.1, Hadoop 1.1.2 (as newer versions are not supported [CASSANDRA-5201|https://issues.apache.org/jira/browse/CASSANDRA-5201]) and cassandra 1.2.3 (I added the current patch from git commits and built sources) ","from":"developer"},{"body":"At last , i could manage a few hours to try the fix. Definitely it's working now, every mapper works on their own range, however i have test in single node cluster with Hadoop 1.1.2 + Pig 0.11.1 and Cassandra 1.2.6. Thankx. ","from":"developer"}],"created":"2013-05-07T07:00:42.000+0000","description":"We have got very strange beheviour of hadoop cluster after upgrading \nCassandra from 1.1.5 to Cassandra 1.2.1. We have 5 nodes cluster of Cassandra, where three of them are hodoop slaves. Now when we are submitting job through Pig script, only one map assigns in task running on one of the hadoop slaves regardless of \nvolume of data (already tried with more than million rows).\nConfigure of pig as follows:\nexport PIG_HOME=/oracle/pig-0.10.0\nexport PIG_CONF_DIR=${HADOOP_HOME}/conf\nexport PIG_INITIAL_ADDRESS=192.168.157.103\nexport PIG_RPC_PORT=9160\nexport PIG_PARTITIONER=org.apache.cassandra.dht.Murmur3Partitioner\n\n\nAlso we have these following properties in hadoop:\n \n mapred.tasktracker.map.tasks.maximum\n 10\n \n \n mapred.map.tasks\n 4\n ","issue_id":"12646363","key":"CASSANDRA-5544","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-05-29T17:55:41.000+0000","role":"fixed_distractor","summary":"Hadoop jobs assigns only one mapper in task"} {"case_id":"12650385","cluster":"DISTRACTOR-CASSANDRA-5605","comments":[{"body":"Do you have multiple data directories? If you only have one, it may have been blacklisted and marked read-only by a previous issue, can you check the logs for anything like that?","created":"2013-05-31T19:27:20.182+0000"},{"body":"We have only one data directory. There is nothing in the log about it being blacklisted.","created":"2013-05-31T19:53:05.661+0000"},{"body":"Am not sure if the following information helps but we too hit this issue in production today. We were running with cassandra 1.2.4 and two patches CASSANDRA-5554 & CASSANDRA-5418. \n\nWe were running with RF=3 and LCS. \n\nWe ran into this issue while using sstablelaoder to push data from remote 1.2.4 cluster nodes to another cluster\n\nWe cross checked using JMX if blacklisting is the cause of this bug and it looks like it is definitely not the case. \n\nWe however saw a pile up of pending compactions ~ 1800 pending compactions per node when node crashed. Surprising thing is that the \"Insufficient disk space to write xxxx bytes\" appears much before the node crashes. For us it started appearing aprrox 3 hours before the node crashed. \n\nThe cluster which showed this behavior was having loads of writes occurring ( We were using multiple SSTableLoaders to stream data into this cluster. ). We pushed in almost 15 TB worth data ( including the RF =3 ) in a matter of 16 hours. We were not serving any reads from this cluster as we were still migrating data to it. \n\nAnother interesting behavior observed that nodes were neighbors in most of the time. \n\nAm not sure if the above information helps but wanted to add it to the context of the ticket. ","created":"2013-07-11T03:56:41.173+0000"},{"body":"Apologies if this isn't directly relevant, but I seem to be experiencing the same issue using 1.2.8 launched via ccm for integration testing. One differentiating feature here is that this happened the first time the node was ever brought up (0 data). The integration tests attempting to use the ccm cluster never completed due to this hang, but all they do is test creation of a fairly simple schema and then attempt to write and then read back a single row. There is plenty of free disk space available...\n\nHere's what was in the log:\n INFO [main] 2013-08-07 14:56:46,763 CassandraDaemon.java (line 118) Logging initialized\n INFO [main] 2013-08-07 14:56:46,807 CassandraDaemon.java (line 145) JVM vendor/version: Java HotSpot(TM) 64-Bit Server VM/1.7.0_25\n INFO [main] 2013-08-07 14:56:46,808 CassandraDaemon.java (line 183) Heap size: 8248098816/8248098816\n INFO [main] 2013-08-07 14:56:46,808 CassandraDaemon.java (line 184) Classpath: /opt/ccmlib_cassandra/ccm/lwcdbng_test_cluster/node1/conf:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/build/classes/main:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/build/classes/thrift:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/antlr-3.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/avro-1.4.0-fixes.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/avro-1.4.0-sources-fixes.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/commons-cli-1.1.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/commons-codec-1.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/commons-lang-2.6.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/compress-lzf-0.8.4.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/concurrentlinkedhashmap-lru-1.3.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/guava-13.0.1.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/high-scale-lib-1.1.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jackson-core-asl-1.9.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jackson-mapper-asl-1.9.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jamm-0.2.5.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jbcrypt-0.3m.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jline-1.0.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/json-simple-1.1.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/libthrift-0.7.0.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/log4j-1.2.16.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/lz4-1.1.0.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/metrics-core-2.0.3.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/netty-3.5.9.Final.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/servlet-api-2.5-20081211.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/slf4j-api-1.7.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/slf4j-log4j12-1.7.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/snakeyaml-1.6.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/snappy-java-1.0.5.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/snaptree-0.1.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jamm-0.2.5.jar\n INFO [main] 2013-08-07 14:56:46,822 CLibrary.java (line 65) JNA not found. Native methods will be disabled.\n INFO [main] 2013-08-07 14:56:46,891 DatabaseDescriptor.java (line 132) Loading settings from file:/opt/ccmlib_cassandra/ccm/lwcdbng_test_cluster/node1/conf/cassandra.yaml\n INFO [main] 2013-08-07 14:56:47,821 DatabaseDescriptor.java (line 150) Data files directories: [/opt/ccmlib_cassandra/ccm/lwcdbng_test_cluster/node1/data]\n INFO [main] 2013-08-07 14:56:47,822 DatabaseDescriptor.java (line 151) Commit log directory: /opt/ccmlib_cassandra/ccm/lwcdbng_test_cluster/node1/commitlogs\n INFO [main] 2013-08-07 14:56:47,822 DatabaseDescriptor.java (line 191) DiskAccessMode 'auto' determined to be mmap, indexAccessMode is mmap\n INFO [main] 2013-08-07 14:56:47,822 DatabaseDescriptor.java (line 205) disk_failure_policy is stop\n INFO [main] 2013-08-07 14:56:48,000 DatabaseDescriptor.java (line 273) Global memtable threshold is enabled at 2622MB\n INFO [main] 2013-08-07 14:56:49,142 DatabaseDescriptor.java (line 401) Not using multi-threaded compaction\n INFO [main] 2013-08-07 14:56:52,610 CacheService.java (line 111) Initializing key cache with capacity of 100 MBs.\n INFO [main] 2013-08-07 14:56:52,675 CacheService.java (line 140) Scheduling key cache save to each 14400 seconds (going to save all keys).\n INFO [main] 2013-08-07 14:56:52,678 CacheService.java (line 154) Initializing row cache with capacity of 0 MBs and provider org.apache.cassandra.cache.SerializingCacheProvider\n INFO [main] 2013-08-07 14:56:52,698 CacheService.java (line 166) Scheduling row cache save to each 0 seconds (going to save all keys).\n INFO [main] 2013-08-07 14:56:57,163 DatabaseDescriptor.java (line 535) Couldn't detect any schema definitions in local storage.\n INFO [main] 2013-08-07 14:56:57,164 DatabaseDescriptor.java (line 540) To create keyspaces and column families, see 'help create' in cqlsh.\n INFO [main] 2013-08-07 14:56:57,258 CommitLog.java (line 120) No commitlog files found; skipping replay\n INFO [main] 2013-08-07 14:56:57,715 StorageService.java (line 456) Cassandra version: 1.2.8-SNAPSHOT\n INFO [main] 2013-08-07 14:56:57,715 StorageService.java (line 457) Thrift API version: 19.36.0\n INFO [main] 2013-08-07 14:56:57,716 StorageService.java (line 458) CQL supported versions: 2.0.0,3.0.5 (default: 3.0.5)\n INFO [main] 2013-08-07 14:56:57,795 StorageService.java (line 483) Loading persisted ring state\n INFO [main] 2013-08-07 14:56:57,805 StorageService.java (line 564) Starting up server gossip\n WARN [main] 2013-08-07 14:56:57,823 SystemTable.java (line 573) No host ID found, created 6abf69c6-8472-4607-96b7-4a532ed8e2a8 (Note: This should happen exactly once per node).\n INFO [main] 2013-08-07 14:56:57,852 ColumnFamilyStore.java (line 630) Enqueuing flush of Memtable-local@2038679449(367/367 serialized/live bytes, 15 ops)\nERROR [FlushWriter:1] 2013-08-07 14:56:57,860 CassandraDaemon.java (line 192) Exception in thread Thread[FlushWriter:1,5,main]\njava.lang.RuntimeException: Insufficient disk space to write 452 bytes\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:42)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:724)\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,243 GCInspector.java (line 119) GC for ParNew: 2786 ms for 1 collections, 230275096 used; max is 8248098816\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,250 StatusLogger.java (line 53) Pool Name Active Pending Blocked\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,291 StatusLogger.java (line 68) MemtablePostFlusher 1 1 0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,292 StatusLogger.java (line 68) FlushWriter 0 0 0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,293 StatusLogger.java (line 68) commitlog_archiver 0 0 0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,294 StatusLogger.java (line 73) CompactionManager 0 0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,522 StatusLogger.java (line 85) MessagingService n/a 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,523 StatusLogger.java (line 95) Cache Type Size Capacity KeysToSave Provider\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,524 StatusLogger.java (line 96) KeyCache 0 104857600 all \n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,525 StatusLogger.java (line 102) RowCache 0 0 all org.apache.cassandra.cache.SerializingCacheProvider\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,525 StatusLogger.java (line 109) ColumnFamily Memtable ops,data\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,526 StatusLogger.java (line 112) system.local 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,526 StatusLogger.java (line 112) system.peers 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,527 StatusLogger.java (line 112) system.batchlog 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,527 StatusLogger.java (line 112) system.NodeIdInfo 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,528 StatusLogger.java (line 112) system.LocationInfo 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,528 StatusLogger.java (line 112) system.Schema 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,529 StatusLogger.java (line 112) system.Migrations 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,529 StatusLogger.java (line 112) system.schema_keyspaces 8,251\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,530 StatusLogger.java (line 112) system.schema_columns 398,24717\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,530 StatusLogger.java (line 112) system.schema_columnfamilies 369,22187\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,530 StatusLogger.java (line 112) system.IndexInfo 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,531 StatusLogger.java (line 112) system.range_xfers 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,531 StatusLogger.java (line 112) system.peer_events 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,532 StatusLogger.java (line 112) system.hints 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,532 StatusLogger.java (line 112) system.HintsColumnFamily 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,532 StatusLogger.java (line 112) system_traces.sessions 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,533 StatusLogger.java (line 112) system_traces.events 0,0\n","created":"2013-08-07T15:27:41.152+0000"},{"body":"After CASSANDRA-4292 flushing checks the space reserved by compactions. So check if you have a bunch of pending compactions with a huge size.","created":"2013-09-16T14:46:11.569+0000"},{"body":"Have seen a lot of instances of people hitting this reported recently. I think we probably shouldn't fail stuff if C* \"thinks\" it might not have enough free space, because stuff is reserved. Probably a good idea to use that reserving information as a hint to which drive to pick for the JBOD case, but if no drive will fit the data when looking at reserved, pick the one with the most free space (like we used to).","created":"2013-09-16T16:33:00.127+0000"},{"body":"Patchset to avoid prematurely declaring ourselves out of space at https://github.com/jbellis/cassandra/commits/5605","created":"2013-09-17T18:55:20.292+0000"},{"body":"+1 to the patch.\n\nThough (maybe in separate ticket?) we still need to add error handler to FlushRunnable so postExecutor does not get blocked.","created":"2013-09-18T15:05:48.821+0000"},{"body":"bq. we still need to add error handler to FlushRunnable so postExecutor does not get blocked\n\nI'm not sure we want to unblock it -- if the flush errors out, then we definitely don't want commitlog segments getting cleaned up. What did you have in mind?","created":"2013-09-18T15:17:47.567+0000"},{"body":"Ah, right.\nJust wondered if we can do something about preventing filling up postExecutor queue.","created":"2013-09-18T15:32:52.641+0000"},{"body":"Committed. If you come up with a good idea there, let's open a new ticket.","created":"2013-09-18T17:45:39.267+0000"}],"conversations":[{"body":"A few times now I have seen our Cassandra nodes crash by running themselves out of memory. It starts with the following exception:\n\n{noformat}\nERROR [FlushWriter:13000] 2013-05-31 11:32:02,350 CassandraDaemon.java (line 164) Exception in thread Thread[FlushWriter:13000,5,main]\njava.lang.RuntimeException: Insufficient disk space to write 8042730 bytes\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:42)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:722)\n{noformat} \n\nAfter which, it seems the MemtablePostFlusher stage gets stuck and no further memtables get flushed: \n\n{noformat} \nINFO [ScheduledTasks:1] 2013-05-31 11:59:12,467 StatusLogger.java (line 68) MemtablePostFlusher 1 32 0\nINFO [ScheduledTasks:1] 2013-05-31 11:59:12,469 StatusLogger.java (line 73) CompactionManager 1 2\n{noformat} \n\nWhat makes this ridiculous is that, at the time, the data directory on this node had 981GB free disk space (as reported by du). We primarily use STCS and at the time the aforementioned exception occurred, at least one compaction task was executing which could have easily involved 981GB (or more) worth of input SSTables. Correct me if I am wrong but but Cassandra counts data currently being compacted against available disk space. In our case, this is a significant overestimation of the space required by compaction since a large portion of the data being compacted has expired or is an overwrite.\n\nMore to the point though, Cassandra should not crash because its out of disk space unless its really actually out of disk space (ie, dont consider 'phantom' compaction disk usage when flushing). I have seen one of our nodes die in this way before our alerts for disk space even went off.","from":"reporter","subject":"Crash caused by insufficient disk space to flush"},{"body":"Do you have multiple data directories? If you only have one, it may have been blacklisted and marked read-only by a previous issue, can you check the logs for anything like that?","from":"developer"},{"body":"We have only one data directory. There is nothing in the log about it being blacklisted.","from":"developer"},{"body":"Am not sure if the following information helps but we too hit this issue in production today. We were running with cassandra 1.2.4 and two patches CASSANDRA-5554 & CASSANDRA-5418. \n\nWe were running with RF=3 and LCS. \n\nWe ran into this issue while using sstablelaoder to push data from remote 1.2.4 cluster nodes to another cluster\n\nWe cross checked using JMX if blacklisting is the cause of this bug and it looks like it is definitely not the case. \n\nWe however saw a pile up of pending compactions ~ 1800 pending compactions per node when node crashed. Surprising thing is that the \"Insufficient disk space to write xxxx bytes\" appears much before the node crashes. For us it started appearing aprrox 3 hours before the node crashed. \n\nThe cluster which showed this behavior was having loads of writes occurring ( We were using multiple SSTableLoaders to stream data into this cluster. ). We pushed in almost 15 TB worth data ( including the RF =3 ) in a matter of 16 hours. We were not serving any reads from this cluster as we were still migrating data to it. \n\nAnother interesting behavior observed that nodes were neighbors in most of the time. \n\nAm not sure if the above information helps but wanted to add it to the context of the ticket. ","from":"developer"},{"body":"Apologies if this isn't directly relevant, but I seem to be experiencing the same issue using 1.2.8 launched via ccm for integration testing. One differentiating feature here is that this happened the first time the node was ever brought up (0 data). The integration tests attempting to use the ccm cluster never completed due to this hang, but all they do is test creation of a fairly simple schema and then attempt to write and then read back a single row. There is plenty of free disk space available...\n\nHere's what was in the log:\n INFO [main] 2013-08-07 14:56:46,763 CassandraDaemon.java (line 118) Logging initialized\n INFO [main] 2013-08-07 14:56:46,807 CassandraDaemon.java (line 145) JVM vendor/version: Java HotSpot(TM) 64-Bit Server VM/1.7.0_25\n INFO [main] 2013-08-07 14:56:46,808 CassandraDaemon.java (line 183) Heap size: 8248098816/8248098816\n INFO [main] 2013-08-07 14:56:46,808 CassandraDaemon.java (line 184) Classpath: /opt/ccmlib_cassandra/ccm/lwcdbng_test_cluster/node1/conf:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/build/classes/main:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/build/classes/thrift:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/antlr-3.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/avro-1.4.0-fixes.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/avro-1.4.0-sources-fixes.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/commons-cli-1.1.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/commons-codec-1.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/commons-lang-2.6.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/compress-lzf-0.8.4.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/concurrentlinkedhashmap-lru-1.3.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/guava-13.0.1.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/high-scale-lib-1.1.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jackson-core-asl-1.9.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jackson-mapper-asl-1.9.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jamm-0.2.5.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jbcrypt-0.3m.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jline-1.0.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/json-simple-1.1.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/libthrift-0.7.0.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/log4j-1.2.16.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/lz4-1.1.0.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/metrics-core-2.0.3.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/netty-3.5.9.Final.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/servlet-api-2.5-20081211.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/slf4j-api-1.7.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/slf4j-log4j12-1.7.2.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/snakeyaml-1.6.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/snappy-java-1.0.5.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/snaptree-0.1.jar:/opt/ccmlib_cassandra/apache-cassandra-1.2.8-src/lib/jamm-0.2.5.jar\n INFO [main] 2013-08-07 14:56:46,822 CLibrary.java (line 65) JNA not found. Native methods will be disabled.\n INFO [main] 2013-08-07 14:56:46,891 DatabaseDescriptor.java (line 132) Loading settings from file:/opt/ccmlib_cassandra/ccm/lwcdbng_test_cluster/node1/conf/cassandra.yaml\n INFO [main] 2013-08-07 14:56:47,821 DatabaseDescriptor.java (line 150) Data files directories: [/opt/ccmlib_cassandra/ccm/lwcdbng_test_cluster/node1/data]\n INFO [main] 2013-08-07 14:56:47,822 DatabaseDescriptor.java (line 151) Commit log directory: /opt/ccmlib_cassandra/ccm/lwcdbng_test_cluster/node1/commitlogs\n INFO [main] 2013-08-07 14:56:47,822 DatabaseDescriptor.java (line 191) DiskAccessMode 'auto' determined to be mmap, indexAccessMode is mmap\n INFO [main] 2013-08-07 14:56:47,822 DatabaseDescriptor.java (line 205) disk_failure_policy is stop\n INFO [main] 2013-08-07 14:56:48,000 DatabaseDescriptor.java (line 273) Global memtable threshold is enabled at 2622MB\n INFO [main] 2013-08-07 14:56:49,142 DatabaseDescriptor.java (line 401) Not using multi-threaded compaction\n INFO [main] 2013-08-07 14:56:52,610 CacheService.java (line 111) Initializing key cache with capacity of 100 MBs.\n INFO [main] 2013-08-07 14:56:52,675 CacheService.java (line 140) Scheduling key cache save to each 14400 seconds (going to save all keys).\n INFO [main] 2013-08-07 14:56:52,678 CacheService.java (line 154) Initializing row cache with capacity of 0 MBs and provider org.apache.cassandra.cache.SerializingCacheProvider\n INFO [main] 2013-08-07 14:56:52,698 CacheService.java (line 166) Scheduling row cache save to each 0 seconds (going to save all keys).\n INFO [main] 2013-08-07 14:56:57,163 DatabaseDescriptor.java (line 535) Couldn't detect any schema definitions in local storage.\n INFO [main] 2013-08-07 14:56:57,164 DatabaseDescriptor.java (line 540) To create keyspaces and column families, see 'help create' in cqlsh.\n INFO [main] 2013-08-07 14:56:57,258 CommitLog.java (line 120) No commitlog files found; skipping replay\n INFO [main] 2013-08-07 14:56:57,715 StorageService.java (line 456) Cassandra version: 1.2.8-SNAPSHOT\n INFO [main] 2013-08-07 14:56:57,715 StorageService.java (line 457) Thrift API version: 19.36.0\n INFO [main] 2013-08-07 14:56:57,716 StorageService.java (line 458) CQL supported versions: 2.0.0,3.0.5 (default: 3.0.5)\n INFO [main] 2013-08-07 14:56:57,795 StorageService.java (line 483) Loading persisted ring state\n INFO [main] 2013-08-07 14:56:57,805 StorageService.java (line 564) Starting up server gossip\n WARN [main] 2013-08-07 14:56:57,823 SystemTable.java (line 573) No host ID found, created 6abf69c6-8472-4607-96b7-4a532ed8e2a8 (Note: This should happen exactly once per node).\n INFO [main] 2013-08-07 14:56:57,852 ColumnFamilyStore.java (line 630) Enqueuing flush of Memtable-local@2038679449(367/367 serialized/live bytes, 15 ops)\nERROR [FlushWriter:1] 2013-08-07 14:56:57,860 CassandraDaemon.java (line 192) Exception in thread Thread[FlushWriter:1,5,main]\njava.lang.RuntimeException: Insufficient disk space to write 452 bytes\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:42)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:724)\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,243 GCInspector.java (line 119) GC for ParNew: 2786 ms for 1 collections, 230275096 used; max is 8248098816\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,250 StatusLogger.java (line 53) Pool Name Active Pending Blocked\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,291 StatusLogger.java (line 68) MemtablePostFlusher 1 1 0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,292 StatusLogger.java (line 68) FlushWriter 0 0 0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,293 StatusLogger.java (line 68) commitlog_archiver 0 0 0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,294 StatusLogger.java (line 73) CompactionManager 0 0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,522 StatusLogger.java (line 85) MessagingService n/a 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,523 StatusLogger.java (line 95) Cache Type Size Capacity KeysToSave Provider\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,524 StatusLogger.java (line 96) KeyCache 0 104857600 all \n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,525 StatusLogger.java (line 102) RowCache 0 0 all org.apache.cassandra.cache.SerializingCacheProvider\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,525 StatusLogger.java (line 109) ColumnFamily Memtable ops,data\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,526 StatusLogger.java (line 112) system.local 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,526 StatusLogger.java (line 112) system.peers 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,527 StatusLogger.java (line 112) system.batchlog 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,527 StatusLogger.java (line 112) system.NodeIdInfo 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,528 StatusLogger.java (line 112) system.LocationInfo 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,528 StatusLogger.java (line 112) system.Schema 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,529 StatusLogger.java (line 112) system.Migrations 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,529 StatusLogger.java (line 112) system.schema_keyspaces 8,251\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,530 StatusLogger.java (line 112) system.schema_columns 398,24717\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,530 StatusLogger.java (line 112) system.schema_columnfamilies 369,22187\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,530 StatusLogger.java (line 112) system.IndexInfo 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,531 StatusLogger.java (line 112) system.range_xfers 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,531 StatusLogger.java (line 112) system.peer_events 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,532 StatusLogger.java (line 112) system.hints 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,532 StatusLogger.java (line 112) system.HintsColumnFamily 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,532 StatusLogger.java (line 112) system_traces.sessions 0,0\n INFO [ScheduledTasks:1] 2013-08-07 15:14:08,533 StatusLogger.java (line 112) system_traces.events 0,0\n","from":"developer"},{"body":"After CASSANDRA-4292 flushing checks the space reserved by compactions. So check if you have a bunch of pending compactions with a huge size.","from":"developer"},{"body":"Have seen a lot of instances of people hitting this reported recently. I think we probably shouldn't fail stuff if C* \"thinks\" it might not have enough free space, because stuff is reserved. Probably a good idea to use that reserving information as a hint to which drive to pick for the JBOD case, but if no drive will fit the data when looking at reserved, pick the one with the most free space (like we used to).","from":"developer"},{"body":"Patchset to avoid prematurely declaring ourselves out of space at https://github.com/jbellis/cassandra/commits/5605","from":"developer"},{"body":"+1 to the patch.\n\nThough (maybe in separate ticket?) we still need to add error handler to FlushRunnable so postExecutor does not get blocked.","from":"developer"},{"body":"bq. we still need to add error handler to FlushRunnable so postExecutor does not get blocked\n\nI'm not sure we want to unblock it -- if the flush errors out, then we definitely don't want commitlog segments getting cleaned up. What did you have in mind?","from":"developer"},{"body":"Ah, right.\nJust wondered if we can do something about preventing filling up postExecutor queue.","from":"developer"},{"body":"Committed. If you come up with a good idea there, let's open a new ticket.","from":"developer"}],"created":"2013-05-31T19:08:54.000+0000","description":"A few times now I have seen our Cassandra nodes crash by running themselves out of memory. It starts with the following exception:\n\n{noformat}\nERROR [FlushWriter:13000] 2013-05-31 11:32:02,350 CassandraDaemon.java (line 164) Exception in thread Thread[FlushWriter:13000,5,main]\njava.lang.RuntimeException: Insufficient disk space to write 8042730 bytes\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:42)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:722)\n{noformat} \n\nAfter which, it seems the MemtablePostFlusher stage gets stuck and no further memtables get flushed: \n\n{noformat} \nINFO [ScheduledTasks:1] 2013-05-31 11:59:12,467 StatusLogger.java (line 68) MemtablePostFlusher 1 32 0\nINFO [ScheduledTasks:1] 2013-05-31 11:59:12,469 StatusLogger.java (line 73) CompactionManager 1 2\n{noformat} \n\nWhat makes this ridiculous is that, at the time, the data directory on this node had 981GB free disk space (as reported by du). We primarily use STCS and at the time the aforementioned exception occurred, at least one compaction task was executing which could have easily involved 981GB (or more) worth of input SSTables. Correct me if I am wrong but but Cassandra counts data currently being compacted against available disk space. In our case, this is a significant overestimation of the space required by compaction since a large portion of the data being compacted has expired or is an overwrite.\n\nMore to the point though, Cassandra should not crash because its out of disk space unless its really actually out of disk space (ie, dont consider 'phantom' compaction disk usage when flushing). I have seen one of our nodes die in this way before our alerts for disk space even went off.","issue_id":"12650385","key":"CASSANDRA-5605","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-09-18T17:45:37.000+0000","role":"fixed_distractor","summary":"Crash caused by insufficient disk space to flush"} {"case_id":"12652464","cluster":"DISTRACTOR-CASSANDRA-5631","comments":[{"body":"Please test 1.2.5","created":"2013-06-12T19:29:58.188+0000"},{"body":"Unfortunately we are using the Astyanax client which only supports up to 1.2.2. Is there anything I can do to detect this case and retry? or work around it?","created":"2013-06-12T21:21:59.506+0000"},{"body":"I was incorrect regarding Astyanax support. I just had a classpath issue. Anyway, I have tested in 1.2.5 and can no longer reproduce. Thanks!","created":"2013-06-13T19:46:03.805+0000"},{"body":"Great; thanks for the followup!","created":"2013-06-13T19:49:21.938+0000"},{"body":"I've seen this on Cassandra 1.2.11. It happens if you create a keyspace, following quickly by creating a column family within that keyspace. The NPE is thrown because Schema.instance.getTableDefinition returns null for the keyspace but it isn't checked.\n\nIn the case I saw, the node that threw the NPE had problems so it wasn't receiving many messages - it didn't get the create keyspace message but did get the create CF message. Even if a node doesn't have any problems, the ordering of these messages is not guaranteed. The node will get the create keyspace message some time later (probably about 60 seconds later when another node has noticed the schema version is wrong) but it won't attempt to recreate the CF unless there is a further CF change (create, update or delete) within that keyspace. Only then is the current cached schema compared with the on disk schema (in DefsTable.mergeColumnFamilies). It then notices the CF doesn't exist so creates it. This could never happen, so the node won't ever create the CF (unless it is restarted).\n\nI think a fix would be to catch the NPEs above, and then, on learning about a new keyspace, check to see if any CFs should have been created for that keyspace.\n\nI haven't tried to repro this on 2.0 but the code looks almost identical so I would expect it to still be present.\n\nCould someone reopen the ticket please?","created":"2014-01-24T14:10:59.953+0000"},{"body":"I'm just seeing the exact same thing on my 2 node cluster (cassandra 2.0.4).","created":"2014-01-30T18:37:41.506+0000"},{"body":"bq. I think a fix would be to catch the NPEs above, and then, on learning about a new keyspace, check to see if any CFs should have been created for that keyspace.\n\nThis sounds reasonable to me. Another way would be to send the keyspace mutation serialized along with any column families created/altered messages, so that there will never be an NPE there in the first place. This had actually come up before. Will have a look.","created":"2014-02-06T02:18:22.129+0000"},{"body":"If the issue is about the node not having gotten the create KS yet. Can you not just wait for schema agreement in your client before going on to the next create? That is how I do things to avoid these kinds of issues.","created":"2014-02-18T20:26:49.589+0000"},{"body":"The attached patch sends the serialized keyspace itself with any CF update/create/drop migration, making the NPE in question impossible - the keyspace will always be there now.","created":"2014-02-19T08:50:18.647+0000"},{"body":"Lgtm (nit: I'd rename serializeKeyspace to say addSerializedKeyspace).\n\nbq. Can you not just wait for schema agreement in your client before going on to the next create?\n\nFor the record, Jeremiah is right that clients are supposed to wait for schema agreement if they want to guarantee the table creation won't fail just after the keyspace one (or alternatively make sure both creation goes through the same coordinator node). Of course, we shouldn't NPE internally if a user don't respect that and that's just what this ticket is about.","created":"2014-02-19T13:18:17.627+0000"},{"body":"Committed with a nit, thanks.","created":"2014-02-19T19:28:18.474+0000"}],"conversations":[{"body":"I'm testing a 2-node cluster and creating a column family right after the nodes startup. I am using the Astyanax client. Sometimes column family creation fails and I see NPEs on the cassandra server:\n\n{noformat}\n2013-06-12 14:55:31,773 ERROR CassandraDaemon [MigrationStage:1] - Exception in thread Thread[MigrationStage:1,5,main]\njava.lang.NullPointerException\n\tat org.apache.cassandra.db.DefsTable.addColumnFamily(DefsTable.java:510)\n\tat org.apache.cassandra.db.DefsTable.mergeColumnFamilies(DefsTable.java:444)\n\tat org.apache.cassandra.db.DefsTable.mergeSchema(DefsTable.java:354)\n\tat org.apache.cassandra.db.DefinitionsUpdateVerbHandler$1.runMayThrow(DefinitionsUpdateVerbHandler.java:55)\n\tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:166)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:722)\n\n{noformat}\n\n{noformat}\n2013-06-12 14:55:31,880 ERROR CassandraDaemon [MigrationStage:1] - Exception in thread Thread[MigrationStage:1,5,main]\njava.lang.NullPointerException\n\tat org.apache.cassandra.db.DefsTable.mergeColumnFamilies(DefsTable.java:475)\n\tat org.apache.cassandra.db.DefsTable.mergeSchema(DefsTable.java:354)\n\tat org.apache.cassandra.db.DefinitionsUpdateVerbHandler$1.runMayThrow(DefinitionsUpdateVerbHandler.java:55)\n\tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:166)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:722)\n{noformat}\n","from":"reporter","subject":"NPE when creating column family shortly after multinode startup"},{"body":"Please test 1.2.5","from":"developer"},{"body":"Unfortunately we are using the Astyanax client which only supports up to 1.2.2. Is there anything I can do to detect this case and retry? or work around it?","from":"developer"},{"body":"I was incorrect regarding Astyanax support. I just had a classpath issue. Anyway, I have tested in 1.2.5 and can no longer reproduce. Thanks!","from":"developer"},{"body":"Great; thanks for the followup!","from":"developer"},{"body":"I've seen this on Cassandra 1.2.11. It happens if you create a keyspace, following quickly by creating a column family within that keyspace. The NPE is thrown because Schema.instance.getTableDefinition returns null for the keyspace but it isn't checked.\n\nIn the case I saw, the node that threw the NPE had problems so it wasn't receiving many messages - it didn't get the create keyspace message but did get the create CF message. Even if a node doesn't have any problems, the ordering of these messages is not guaranteed. The node will get the create keyspace message some time later (probably about 60 seconds later when another node has noticed the schema version is wrong) but it won't attempt to recreate the CF unless there is a further CF change (create, update or delete) within that keyspace. Only then is the current cached schema compared with the on disk schema (in DefsTable.mergeColumnFamilies). It then notices the CF doesn't exist so creates it. This could never happen, so the node won't ever create the CF (unless it is restarted).\n\nI think a fix would be to catch the NPEs above, and then, on learning about a new keyspace, check to see if any CFs should have been created for that keyspace.\n\nI haven't tried to repro this on 2.0 but the code looks almost identical so I would expect it to still be present.\n\nCould someone reopen the ticket please?","from":"developer"},{"body":"I'm just seeing the exact same thing on my 2 node cluster (cassandra 2.0.4).","from":"developer"},{"body":"bq. I think a fix would be to catch the NPEs above, and then, on learning about a new keyspace, check to see if any CFs should have been created for that keyspace.\n\nThis sounds reasonable to me. Another way would be to send the keyspace mutation serialized along with any column families created/altered messages, so that there will never be an NPE there in the first place. This had actually come up before. Will have a look.","from":"developer"},{"body":"If the issue is about the node not having gotten the create KS yet. Can you not just wait for schema agreement in your client before going on to the next create? That is how I do things to avoid these kinds of issues.","from":"developer"},{"body":"The attached patch sends the serialized keyspace itself with any CF update/create/drop migration, making the NPE in question impossible - the keyspace will always be there now.","from":"developer"},{"body":"Lgtm (nit: I'd rename serializeKeyspace to say addSerializedKeyspace).\n\nbq. Can you not just wait for schema agreement in your client before going on to the next create?\n\nFor the record, Jeremiah is right that clients are supposed to wait for schema agreement if they want to guarantee the table creation won't fail just after the keyspace one (or alternatively make sure both creation goes through the same coordinator node). Of course, we shouldn't NPE internally if a user don't respect that and that's just what this ticket is about.","from":"developer"},{"body":"Committed with a nit, thanks.","from":"developer"}],"created":"2013-06-12T19:27:30.000+0000","description":"I'm testing a 2-node cluster and creating a column family right after the nodes startup. I am using the Astyanax client. Sometimes column family creation fails and I see NPEs on the cassandra server:\n\n{noformat}\n2013-06-12 14:55:31,773 ERROR CassandraDaemon [MigrationStage:1] - Exception in thread Thread[MigrationStage:1,5,main]\njava.lang.NullPointerException\n\tat org.apache.cassandra.db.DefsTable.addColumnFamily(DefsTable.java:510)\n\tat org.apache.cassandra.db.DefsTable.mergeColumnFamilies(DefsTable.java:444)\n\tat org.apache.cassandra.db.DefsTable.mergeSchema(DefsTable.java:354)\n\tat org.apache.cassandra.db.DefinitionsUpdateVerbHandler$1.runMayThrow(DefinitionsUpdateVerbHandler.java:55)\n\tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:166)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:722)\n\n{noformat}\n\n{noformat}\n2013-06-12 14:55:31,880 ERROR CassandraDaemon [MigrationStage:1] - Exception in thread Thread[MigrationStage:1,5,main]\njava.lang.NullPointerException\n\tat org.apache.cassandra.db.DefsTable.mergeColumnFamilies(DefsTable.java:475)\n\tat org.apache.cassandra.db.DefsTable.mergeSchema(DefsTable.java:354)\n\tat org.apache.cassandra.db.DefinitionsUpdateVerbHandler$1.runMayThrow(DefinitionsUpdateVerbHandler.java:55)\n\tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:166)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:722)\n{noformat}\n","issue_id":"12652464","key":"CASSANDRA-5631","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-02-19T19:28:18.000+0000","role":"fixed_distractor","summary":"NPE when creating column family shortly after multinode startup"} {"case_id":"12656702","cluster":"DISTRACTOR-CASSANDRA-5732","comments":[{"body":"Hi,\n\nI tried the same set of tables, data, java code in Windows 7 jre 1.7.0_05 64-bit and it seems to not have the issue so far. Will try other environment using 1.7.0_x 32 then 64-bit to see if that solves the issue.\n\nRegards,\n-Tony","created":"2013-07-09T04:52:19.818+0000"},{"body":"Ok. I figured out it has to do with a config setting in cassandra.yaml file. If you set row_cache_size_in_mb to 200 instead of 0 and you are setting caching = \"ALL\" for a column family and using and secondary index for a query this issue occurs. If you set row_cache_size_in_mb this problem goes away.\nLet me know why this is. It would be nice to use row caching and column family caching.","created":"2013-07-10T04:54:46.236+0000"},{"body":"Is this the same as CASSANDRA-4785 and CASSANDRA-4973?","created":"2013-07-10T15:04:40.559+0000"},{"body":"Hi Janne,\n \nThere are simularities. Mine though is a solid failure and I narrowed it down to what I said so the Cassandra team should be able to solve the issue.\n \nBest Regards,\n-Tony\n\nFrom: Janne Jalkanen (JIRA) \nTo: adanecito@yahoo.com \nSent: Wednesday, July 10, 2013 9:05 AM\nSubject: [jira] [Commented] (CASSANDRA-5732) Can not query secondary index\n\n\n\n    [ https://issues.apache.org/jira/browse/CASSANDRA-5732?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13704633#comment-13704633 ] \n\nJanne Jalkanen commented on CASSANDRA-5732:\n-------------------------------------------\n\nIs this the same as CASSANDRA-4785?\n                \n\n--\nThis message is automatically generated by JIRA.\nIf you think it was sent incorrectly, please contact your JIRA administrators\nFor more information on JIRA, see: http://www.atlassian.com/software/jira\n","created":"2013-07-10T15:19:47.838+0000"},{"body":"This is definitly a bug in latest 1.2. Just reproduced in 1.2.10\n\nRepro is easy. Create single node cluster off of latest 1.2 branch and do the following:\n\ncqlsh:ks> CREATE TABLE test ( row text, name text, PRIMARY KEY (row) );\ncqlsh:ks> ALTER TABLE test WITH caching='all';\ncqlsh:ks> INSERT INTO test (row, name) VALUES ( 'row1', 'daniel' );\ncqlsh:ks> INSERT INTO test (row, name) VALUES ( 'row2', 'ryan' );\ncqlsh:ks> CREATE INDEX on test (name);\ncqlsh:ks> select * from test where name='daniel';\n\n row | name\n------+--------\n row1 | daniel\n#Notice how row is returned\n\nStop node and set:\nrow_cache_size_in_mb: 200\nStart node again and follow this procedure:\n\ncqlsh> use ks;\ncqlsh:ks> select * from test where name='daniel';\n#Nothing is returned from this query\n\n\nNow stop the node and set:\nrow_cache_size_in_mb: 0\n\nstart the node and do the following:\n\ncqlsh> use ks;\ncqlsh:ks> select * from test where name='daniel';\n\n row | name\n------+--------\n row1 | daniel\n\n#Notice how the row is returned.","created":"2013-09-19T16:54:57.825+0000"},{"body":"Suspect this might be up your alley, [~beobal].","created":"2013-09-19T16:57:54.213+0000"},{"body":"The reason for the missing results is that in CFS.getColumnFamily() we look up the cfs id from Schema to calculate the cache key. However, 2i CFSes are never loaded into the Schema, so Schema.instance.getId always returns null. Simply fixing this by calling Schema.instance.load() with the 2i CFMD when the index is initialized uncovers another issue. The cfid is now retrievable, but the deserialization of a cached 2i row fails as it depends on the 2i CFMD being present in the enclosing KSMD for the eventual call to Schema.getCFMD(). Once we start adding index CFs to Schema they then become involved in schema migrations which makes everything very messy. So rather than adding them directly to KSMD like regular CFs, I added a separate cfId->CFMD map to Schema, so as far as most things are concerned nothing has changed, just we have one further place to look when retrieving CFMD for a given cfId.\n\nThe attached patch is against the 1.2 branch, CASSANDRA-4875 is a duplicate of this, but has a fixver of 1.1 [~jbellis], do you want me to submit a patch against 1.1 also?\n\nI wrote a dtest for this, pull request for that here: https://github.com/riptano/cassandra-dtest/pull/22\n\nLooking at this, I also uncovered what I think is an issue with the setup of the 2i cache config. In AbstractSimplePerColumnSecondaryIndex (in 1.2, the same code is in KeysIndex in 1.1), the estimated key and mean column counts are used to gauge the index's cardinality then use that to decide whether or not to enable row caching. This calculation is first performed prior to the index actually being built, so there are no SSTables to provide the estimates, which results in row caching always being disabled until the next time the index is initialized when C* is restarted (this appears to be why the repro steps require a restart). If this is a genuine problem, I'll create a separate JIRA to address it. \n","created":"2013-10-10T14:02:11.527+0000"},{"body":"Hmm, it's not actually necessary to look up your id through the schema here. v2 is a simpler solution if that's the only problem.","created":"2013-10-10T18:24:18.485+0000"},{"body":"Yeah, unfortunately that failed lookup was only masking the other problem I mentioned. Even with the v2 fix, without the index cfm in Schema, you fall foul (silently, except for debug) of https://github.com/apache/cassandra/blob/cassandra-1.2/src/java/org/apache/cassandra/db/ColumnFamilySerializer.java#L183 and so the row cache is effectively bypassed for 2i cfs.","created":"2013-10-10T18:50:09.494+0000"},{"body":"I see.\n\nI think your patch is probably fine, but I'm super nervous about messing with Schema internals this late in 1.2.x (let alone 1.1).\n\nv3 just disables row cache entirely for 2i CFs.","created":"2013-10-10T19:49:26.589+0000"},{"body":"Sure, that's reasonable and pretty much expected. I'll attach a patch for trunk/2.0","created":"2013-10-10T19:59:19.282+0000"},{"body":"Committed v3 then.\n\nI don't think it's worth doing extra work for 2.0 since we're removing row cache for 2.1.","created":"2013-10-13T16:14:59.087+0000"},{"body":"FWIW, appears to also be a problem in Cassandra 1.1.","created":"2014-02-10T10:09:57.502+0000"}],"conversations":[{"body":"Noticed after taking a column family that already existed and assigning to an IntegerType index_type:KEYS and the caching was already set to 'ALL' that the prepared statement do not return rows neither did it throw an exception. Here is the sequence.\n1. Starting state query running with caching off for a Column Family with the query using the secondary index for te WHERE clause.\n2, Set Column Family caching to ALL using Cassandra-CLI and update CQL. Cassandra-cli Describe shows column family caching set to ALL\n3. Rerun query and it works.\n4. Restart Cassandra and run query and no rows returned. Cassandra-cli Describe shows column family caching set to ALL\n5. Set Column Family caching to NONE using Cassandra-cli and update CQL. Rerun query and no rows returned. Cassandra-cli Describe for column family shows caching set to NONE.\n6. Restart Cassandra. Rerun query and it is working again. We are now back to the starting state.\n\nBest Regards,\n-Tony","from":"reporter","subject":"Can not query secondary index"},{"body":"Hi,\n\nI tried the same set of tables, data, java code in Windows 7 jre 1.7.0_05 64-bit and it seems to not have the issue so far. Will try other environment using 1.7.0_x 32 then 64-bit to see if that solves the issue.\n\nRegards,\n-Tony","from":"developer"},{"body":"Ok. I figured out it has to do with a config setting in cassandra.yaml file. If you set row_cache_size_in_mb to 200 instead of 0 and you are setting caching = \"ALL\" for a column family and using and secondary index for a query this issue occurs. If you set row_cache_size_in_mb this problem goes away.\nLet me know why this is. It would be nice to use row caching and column family caching.","from":"developer"},{"body":"Is this the same as CASSANDRA-4785 and CASSANDRA-4973?","from":"developer"},{"body":"Hi Janne,\n \nThere are simularities. Mine though is a solid failure and I narrowed it down to what I said so the Cassandra team should be able to solve the issue.\n \nBest Regards,\n-Tony\n\nFrom: Janne Jalkanen (JIRA) \nTo: adanecito@yahoo.com \nSent: Wednesday, July 10, 2013 9:05 AM\nSubject: [jira] [Commented] (CASSANDRA-5732) Can not query secondary index\n\n\n\n    [ https://issues.apache.org/jira/browse/CASSANDRA-5732?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13704633#comment-13704633 ] \n\nJanne Jalkanen commented on CASSANDRA-5732:\n-------------------------------------------\n\nIs this the same as CASSANDRA-4785?\n                \n\n--\nThis message is automatically generated by JIRA.\nIf you think it was sent incorrectly, please contact your JIRA administrators\nFor more information on JIRA, see: http://www.atlassian.com/software/jira\n","from":"developer"},{"body":"This is definitly a bug in latest 1.2. Just reproduced in 1.2.10\n\nRepro is easy. Create single node cluster off of latest 1.2 branch and do the following:\n\ncqlsh:ks> CREATE TABLE test ( row text, name text, PRIMARY KEY (row) );\ncqlsh:ks> ALTER TABLE test WITH caching='all';\ncqlsh:ks> INSERT INTO test (row, name) VALUES ( 'row1', 'daniel' );\ncqlsh:ks> INSERT INTO test (row, name) VALUES ( 'row2', 'ryan' );\ncqlsh:ks> CREATE INDEX on test (name);\ncqlsh:ks> select * from test where name='daniel';\n\n row | name\n------+--------\n row1 | daniel\n#Notice how row is returned\n\nStop node and set:\nrow_cache_size_in_mb: 200\nStart node again and follow this procedure:\n\ncqlsh> use ks;\ncqlsh:ks> select * from test where name='daniel';\n#Nothing is returned from this query\n\n\nNow stop the node and set:\nrow_cache_size_in_mb: 0\n\nstart the node and do the following:\n\ncqlsh> use ks;\ncqlsh:ks> select * from test where name='daniel';\n\n row | name\n------+--------\n row1 | daniel\n\n#Notice how the row is returned.","from":"developer"},{"body":"Suspect this might be up your alley, [~beobal].","from":"developer"},{"body":"The reason for the missing results is that in CFS.getColumnFamily() we look up the cfs id from Schema to calculate the cache key. However, 2i CFSes are never loaded into the Schema, so Schema.instance.getId always returns null. Simply fixing this by calling Schema.instance.load() with the 2i CFMD when the index is initialized uncovers another issue. The cfid is now retrievable, but the deserialization of a cached 2i row fails as it depends on the 2i CFMD being present in the enclosing KSMD for the eventual call to Schema.getCFMD(). Once we start adding index CFs to Schema they then become involved in schema migrations which makes everything very messy. So rather than adding them directly to KSMD like regular CFs, I added a separate cfId->CFMD map to Schema, so as far as most things are concerned nothing has changed, just we have one further place to look when retrieving CFMD for a given cfId.\n\nThe attached patch is against the 1.2 branch, CASSANDRA-4875 is a duplicate of this, but has a fixver of 1.1 [~jbellis], do you want me to submit a patch against 1.1 also?\n\nI wrote a dtest for this, pull request for that here: https://github.com/riptano/cassandra-dtest/pull/22\n\nLooking at this, I also uncovered what I think is an issue with the setup of the 2i cache config. In AbstractSimplePerColumnSecondaryIndex (in 1.2, the same code is in KeysIndex in 1.1), the estimated key and mean column counts are used to gauge the index's cardinality then use that to decide whether or not to enable row caching. This calculation is first performed prior to the index actually being built, so there are no SSTables to provide the estimates, which results in row caching always being disabled until the next time the index is initialized when C* is restarted (this appears to be why the repro steps require a restart). If this is a genuine problem, I'll create a separate JIRA to address it. \n","from":"developer"},{"body":"Hmm, it's not actually necessary to look up your id through the schema here. v2 is a simpler solution if that's the only problem.","from":"developer"},{"body":"Yeah, unfortunately that failed lookup was only masking the other problem I mentioned. Even with the v2 fix, without the index cfm in Schema, you fall foul (silently, except for debug) of https://github.com/apache/cassandra/blob/cassandra-1.2/src/java/org/apache/cassandra/db/ColumnFamilySerializer.java#L183 and so the row cache is effectively bypassed for 2i cfs.","from":"developer"},{"body":"I see.\n\nI think your patch is probably fine, but I'm super nervous about messing with Schema internals this late in 1.2.x (let alone 1.1).\n\nv3 just disables row cache entirely for 2i CFs.","from":"developer"},{"body":"Sure, that's reasonable and pretty much expected. I'll attach a patch for trunk/2.0","from":"developer"},{"body":"Committed v3 then.\n\nI don't think it's worth doing extra work for 2.0 since we're removing row cache for 2.1.","from":"developer"},{"body":"FWIW, appears to also be a problem in Cassandra 1.1.","from":"developer"}],"created":"2013-07-08T21:35:28.000+0000","description":"Noticed after taking a column family that already existed and assigning to an IntegerType index_type:KEYS and the caching was already set to 'ALL' that the prepared statement do not return rows neither did it throw an exception. Here is the sequence.\n1. Starting state query running with caching off for a Column Family with the query using the secondary index for te WHERE clause.\n2, Set Column Family caching to ALL using Cassandra-CLI and update CQL. Cassandra-cli Describe shows column family caching set to ALL\n3. Rerun query and it works.\n4. Restart Cassandra and run query and no rows returned. Cassandra-cli Describe shows column family caching set to ALL\n5. Set Column Family caching to NONE using Cassandra-cli and update CQL. Rerun query and no rows returned. Cassandra-cli Describe for column family shows caching set to NONE.\n6. Restart Cassandra. Rerun query and it is working again. We are now back to the starting state.\n\nBest Regards,\n-Tony","issue_id":"12656702","key":"CASSANDRA-5732","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-10-13T16:14:59.000+0000","role":"fixed_distractor","summary":"Can not query secondary index"} {"case_id":"12659484","cluster":"DISTRACTOR-CASSANDRA-5797","comments":[{"body":"Not sure what CQL syntax for this is. Is it protocol level the way CL is?","created":"2013-07-24T03:03:25.900+0000"},{"body":"bq. Not sure what CQL syntax for this is. Is it protocol level the way CL is?\n\nThat's a good question and I'm not really sure what's the right answer.\n\nI think it may make the most sense to make it protocol level because of reads. For CAS writes, we do have a CQL syntax for it, so we could extends it with say:\n{noformat}\nUPDATE foo SET v1 = 2, v2 = 3 WHERE k = 1 IF v1 = 1 AND v2 = 1 IN LOCAL DC\n{noformat}\nBut for reads, we don't have any syntax, the consistency level (SERIAL) is the only thing that makes a read go through paxos, so I'm afraid adding some CQL syntax in that case would be confusing.\n\nBut even making it protocol level is not that easy. For thrift, on the read side, the only way I can see us supporting this DC-local CAS would be add a LOCAL_SERIAL consistency level (short of duplicating all read methods for CAS reads that is). But that doesn't really work for writes since the consistency level for writes is really the consistency of the paxos learn/commit phase.\n\nOne option (the best I can come up with so far) would be to add the LOCAL_SERIAL consistency level, and then to change CAS write to take 2 CL: the first one would be for the commit (learn) phase (as we have now, but we would refuse CL.SERIAL and CL.LOCAL_SERIAL in that case) and a 2nd CL that would control the \"Paxos consistency\" (and for that one, only CL.SERIAL or CL.LOCAL_SERIAL would be valid). It's not perfect however because the one thing you can't properly express is the ability to do CL.SERIAL for paxos but don't wait on any node for the learn phase. Unless we make CL.ANY for the \"commit consistency\" mean that, but that's a slight stretch.\n\nIn any case, we should probably make it sure to shove that in 2.0.0 because I don't want to change the thrift API nor break the native protocol in 2.0.1.\n\nAny better idea?\n","created":"2013-07-24T11:49:46.435+0000"},{"body":"Sounds reasonable, although I think it would be better to come up w/ a different enum for the Paxos phases than re-use CL, most of whose options are not appropriate.\n\nI actually think CL.ANY on commit is fine.","created":"2013-07-24T15:14:56.468+0000"},{"body":"bq. although I think it would be better to come up w/ a different enum for the Paxos phases than re-use CL\n\nThe thing is that for reads, we must have SERIAL and LOCAL_SERIAL in CL if we want thrift to support it. So once we have them in CL, is it really worth adding a separate enum for the write case? (honest question, I'm fine doing it, just wonder if it's worth bothering since things will be mixed up for reads anyway).","created":"2013-07-24T15:23:13.925+0000"},{"body":"I like LOCAL_SERIAL over ANY. It makes a closer match to LOCAL_QUORUM in that it's not meant to cross datacenter boundaries. There is enough confusion about ANY as it is and I think this would simplify things.","created":"2013-07-24T15:27:54.327+0000"},{"body":"bq. I like LOCAL_SERIAL over ANY\n\nI think there is some confusion. The suggestion of CL.ANY was for the commit part of Paxos. That part is basically a standard write (that happens after the paxos algorithm has unfolded but does impact the visibility of the CAS write by non-serial reads). For that, LOCAL_SERIAL don't really make sense imo (it's \"wrong\" even). ANY does is what match the most what happens, because you are guaranteed the write is replicated somewhere (paxos ensures that) but you may not be able to see your write right away with normal reads, even at CL.ALL (which is also something that CL.ANY). ","created":"2013-07-24T16:33:44.223+0000"},{"body":"bq. The thing is that for reads, we must have SERIAL and LOCAL_SERIAL in CL if we want thrift to support it. So once we have them in CL, is it really worth adding a separate enum for the write case?\n\nThe problem is that none of {ANY, ONE, TWO, THREE, LOCAL_QUORUM, EACH_QUORUM} are valid on writes, which isn't very clear if we reuse CL for everything.\n\nThen again we ANY is not a valid CL for read, and EACH_QUORUM is not valid for writes. I dunno.","created":"2013-07-24T23:51:40.212+0000"},{"body":"Attaching patch for this. I've currently stayed with the idea of 2 CL for writes. One small reason for which it is somewhat convenient is that when we throw a write timeout exception, we need to ship the consistency level. And when that timeout happens during the paxos prepare/propose phases, returning CL.SERIAL or CL.LOCAL_SERIAL make the most sense, so using CL as argument of the method in the first place is somewhat consistent. Anyway, it's certainly possible to add a new enum for that instead, but I don't think that' too horrible as is.\n\nThere's 3 patches: the first one is just the update to the thrift generated files and can be largely ignored. The 2nd one does the main change but does not change CQL. The 3rd patch is the CQL and native protocol change. That latter patch is a tad big because while adding the new \"serial consistency level\", I realized this was a pain with the current code and that in the protocol v2, QUERY and EXECUTE had basically the same parameters but they were set out in different order (in the protocol) which was killing all possibility of code reuse for no good reason. So I decided to go ahead and refactor that more cleanly since there's no point in making the life of client implementors harder for no good reason.\n\nAnyway, moving this issue to 2.0.0 because it changes the native protocol and thrift, so we really should have it in 2.0.\n","created":"2013-07-25T12:51:25.357+0000"},{"body":"+1 (mostly looked at patch 2)\n\nNit: rename isSameDCThan to isSameDCAs, or maybe better sameDCPredicateFor (since \"is\" implies it's doing a boolean evaluation)","created":"2013-07-26T14:14:24.504+0000"},{"body":"QueryOptions.Codec.encode() writes flags before writing the CL - should be reversed. Otherwise LGTM.","created":"2013-07-26T15:30:25.032+0000"},{"body":"Alright, committed (with nit fixed). Thanks","created":"2013-07-26T15:31:27.254+0000"}],"conversations":[{"body":"For two-datacenter deployments where the second DC is strictly for disaster failover, it would be useful to restrict CAS to a single DC to avoid cross-DC round trips.\n\n(This would require manually truncating {{system.paxos}} when failing over.)","from":"reporter","subject":"DC-local CAS"},{"body":"Not sure what CQL syntax for this is. Is it protocol level the way CL is?","from":"developer"},{"body":"bq. Not sure what CQL syntax for this is. Is it protocol level the way CL is?\n\nThat's a good question and I'm not really sure what's the right answer.\n\nI think it may make the most sense to make it protocol level because of reads. For CAS writes, we do have a CQL syntax for it, so we could extends it with say:\n{noformat}\nUPDATE foo SET v1 = 2, v2 = 3 WHERE k = 1 IF v1 = 1 AND v2 = 1 IN LOCAL DC\n{noformat}\nBut for reads, we don't have any syntax, the consistency level (SERIAL) is the only thing that makes a read go through paxos, so I'm afraid adding some CQL syntax in that case would be confusing.\n\nBut even making it protocol level is not that easy. For thrift, on the read side, the only way I can see us supporting this DC-local CAS would be add a LOCAL_SERIAL consistency level (short of duplicating all read methods for CAS reads that is). But that doesn't really work for writes since the consistency level for writes is really the consistency of the paxos learn/commit phase.\n\nOne option (the best I can come up with so far) would be to add the LOCAL_SERIAL consistency level, and then to change CAS write to take 2 CL: the first one would be for the commit (learn) phase (as we have now, but we would refuse CL.SERIAL and CL.LOCAL_SERIAL in that case) and a 2nd CL that would control the \"Paxos consistency\" (and for that one, only CL.SERIAL or CL.LOCAL_SERIAL would be valid). It's not perfect however because the one thing you can't properly express is the ability to do CL.SERIAL for paxos but don't wait on any node for the learn phase. Unless we make CL.ANY for the \"commit consistency\" mean that, but that's a slight stretch.\n\nIn any case, we should probably make it sure to shove that in 2.0.0 because I don't want to change the thrift API nor break the native protocol in 2.0.1.\n\nAny better idea?\n","from":"developer"},{"body":"Sounds reasonable, although I think it would be better to come up w/ a different enum for the Paxos phases than re-use CL, most of whose options are not appropriate.\n\nI actually think CL.ANY on commit is fine.","from":"developer"},{"body":"bq. although I think it would be better to come up w/ a different enum for the Paxos phases than re-use CL\n\nThe thing is that for reads, we must have SERIAL and LOCAL_SERIAL in CL if we want thrift to support it. So once we have them in CL, is it really worth adding a separate enum for the write case? (honest question, I'm fine doing it, just wonder if it's worth bothering since things will be mixed up for reads anyway).","from":"developer"},{"body":"I like LOCAL_SERIAL over ANY. It makes a closer match to LOCAL_QUORUM in that it's not meant to cross datacenter boundaries. There is enough confusion about ANY as it is and I think this would simplify things.","from":"developer"},{"body":"bq. I like LOCAL_SERIAL over ANY\n\nI think there is some confusion. The suggestion of CL.ANY was for the commit part of Paxos. That part is basically a standard write (that happens after the paxos algorithm has unfolded but does impact the visibility of the CAS write by non-serial reads). For that, LOCAL_SERIAL don't really make sense imo (it's \"wrong\" even). ANY does is what match the most what happens, because you are guaranteed the write is replicated somewhere (paxos ensures that) but you may not be able to see your write right away with normal reads, even at CL.ALL (which is also something that CL.ANY). ","from":"developer"},{"body":"bq. The thing is that for reads, we must have SERIAL and LOCAL_SERIAL in CL if we want thrift to support it. So once we have them in CL, is it really worth adding a separate enum for the write case?\n\nThe problem is that none of {ANY, ONE, TWO, THREE, LOCAL_QUORUM, EACH_QUORUM} are valid on writes, which isn't very clear if we reuse CL for everything.\n\nThen again we ANY is not a valid CL for read, and EACH_QUORUM is not valid for writes. I dunno.","from":"developer"},{"body":"Attaching patch for this. I've currently stayed with the idea of 2 CL for writes. One small reason for which it is somewhat convenient is that when we throw a write timeout exception, we need to ship the consistency level. And when that timeout happens during the paxos prepare/propose phases, returning CL.SERIAL or CL.LOCAL_SERIAL make the most sense, so using CL as argument of the method in the first place is somewhat consistent. Anyway, it's certainly possible to add a new enum for that instead, but I don't think that' too horrible as is.\n\nThere's 3 patches: the first one is just the update to the thrift generated files and can be largely ignored. The 2nd one does the main change but does not change CQL. The 3rd patch is the CQL and native protocol change. That latter patch is a tad big because while adding the new \"serial consistency level\", I realized this was a pain with the current code and that in the protocol v2, QUERY and EXECUTE had basically the same parameters but they were set out in different order (in the protocol) which was killing all possibility of code reuse for no good reason. So I decided to go ahead and refactor that more cleanly since there's no point in making the life of client implementors harder for no good reason.\n\nAnyway, moving this issue to 2.0.0 because it changes the native protocol and thrift, so we really should have it in 2.0.\n","from":"developer"},{"body":"+1 (mostly looked at patch 2)\n\nNit: rename isSameDCThan to isSameDCAs, or maybe better sameDCPredicateFor (since \"is\" implies it's doing a boolean evaluation)","from":"developer"},{"body":"QueryOptions.Codec.encode() writes flags before writing the CL - should be reversed. Otherwise LGTM.","from":"developer"},{"body":"Alright, committed (with nit fixed). Thanks","from":"developer"}],"created":"2013-07-24T03:02:53.000+0000","description":"For two-datacenter deployments where the second DC is strictly for disaster failover, it would be useful to restrict CAS to a single DC to avoid cross-DC round trips.\n\n(This would require manually truncating {{system.paxos}} when failing over.)","issue_id":"12659484","key":"CASSANDRA-5797","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-07-26T15:31:27.000+0000","role":"fixed_distractor","summary":"DC-local CAS"} {"case_id":"12662857","cluster":"DISTRACTOR-CASSANDRA-5867","comments":[{"body":"A tuple is the obvious choice, since a bag is cumbersome and there's no need to spill to disk.","created":"2013-08-09T14:26:56.699+0000"},{"body":"In the same vein as data types, as of PIG-2764 Pig has a BigInteger. So once 0.12 is out and mainstream, we could look at using that instead of the Pig integer type to avoid overflow. Just as a heads up.","created":"2013-08-09T14:31:21.378+0000"},{"body":"Patch on 1.2 branch is attached, which map ListType and SetType to tuple, MapType to map","created":"2013-08-15T02:08:33.394+0000"},{"body":"We may not want to use the MapType for maps.\n\nIn Pig, you can only have text based keys for maps. From Programming Pig: \"A map in Pig is a chararray to data element mapping, where that element can be any Pig type, including a complex type.\"\n\nIn Cassandra, you can have map keys of any type, like timestamp in the maps example on http://www.datastax.com/dev/blog/cql3_collections\n\nSo we may want to do a bag or tuple of tuples.","created":"2013-08-15T12:31:54.109+0000"},{"body":"The key of Cassandra map has been converted to string into Pig map. We need to decide whether use tuple of tuples vs map for Cassandra map type. Tuple of tuples is more general than map, and map is more specific. HBase uses map to map its row. So which one to use for map? Map vs Tuple of tuples?","created":"2013-08-15T16:50:28.610+0000"},{"body":"To store data to CQL3 table, we supports the following prepared statements\n\n{code}\nList\n e.g.\n UPDATE users SET top_places = ?\n UPDATE users SET top_places = [ 'rivendell', 'rohan' ] WHERE user_id = 'frodo';\n\n UPDATE users SET top_places = ? + top_places\n UPDATE users SET top_places = [ 'the shire' ] + top_places WHERE user_id = 'frodo';\n\n UPDATE users SET top_places = top_places - ?;\n UPDATE users SET top_places = top_places - ['riddermark'] WHERE user_id = 'frodo';\n\nSet statements are similar to List\n\nMap\n\n UPDATE users SET todo = ?\n UPDATE users\n SET todo = { '2012-9-24' : 'enter mordor',\n '2012-10-2 12:00' : 'throw ring into mount doom' }\n WHERE user_id = 'frodo';\n\n\nThe following queries are handled as a regular value instead of tuples\n UPDATE users SET top_places[2] = ?\n UPDATE users SET top_places[2] = 'riddermark' WHERE user_id = 'frodo';\n\n DELETE top_places[3] FROM users;\n DELETE top_places[3] FROM users WHERE user_id = 'frodo';\n\n UPDATE users SET todo[?] = ?\n UPDATE users SET todo['2012-10-2 12:10'] = 'die' WHERE user_id = 'frodo';\n{code}\n\nThe output schema for collections is as following\n\n{code}\n (((name, value), (name, value)), (value ... value), (value...value)\n If a value of tuple (value...value) is a tuple of (inner_value ...inner_value) \n and the first inner_value is the collection type. \n it is either \"set\", \"list\" or \"map\".\n\n e.g. (value ... value) as (value1, value2, (set, riddermark, tom))\n map (value1, value2, (map, (sfo, 12), (ny, 34))\n{code}\n","created":"2013-08-16T22:50:22.758+0000"},{"body":"5867-3-1.2-branch.txt is attached to support storing collections to Cassandra","created":"2013-08-19T20:14:23.242+0000"},{"body":"If the key of C* map is converted to string (chararray) in a Pig map on read, how will the keys be handled on write? Will the strings be auto-converted to appropriate C* types? In the past, I've seen issues with that - especially with timestamps and UUID values.","created":"2013-08-22T20:13:17.594+0000"},{"body":"The following is the data which will be auto converted to C* type, anything else will be bytes.\n\n{code}\n if (o == null)\n return (ByteBuffer)o;\n if (o instanceof java.lang.String)\n return ByteBuffer.wrap(new DataByteArray((String)o).get());\n if (o instanceof Integer)\n return Int32Type.instance.decompose((Integer)o);\n if (o instanceof Long)\n return LongType.instance.decompose((Long)o);\n if (o instanceof Float)\n return FloatType.instance.decompose((Float)o);\n if (o instanceof Double)\n return DoubleType.instance.decompose((Double)o);\n if (o instanceof UUID)\n return ByteBuffer.wrap(UUIDGen.decompose((UUID) o));\n\n return ByteBuffer.wrap(((DataByteArray) o).get());\n{code}\n\nyou need prepare the data for write different from read result.","created":"2013-08-22T20:49:27.525+0000"},{"body":"Alex, is this ready for review?","created":"2013-08-26T22:16:50.274+0000"},{"body":"Yes, But I map the C* MapType to Pig map, we may need map it to tuples of tuples.","created":"2013-08-26T23:04:31.704+0000"},{"body":"I also has a few bug fixes related to CqlStorage, so if it's possible, I can add those fix to this patch as well.","created":"2013-08-26T23:05:42.636+0000"},{"body":"Since pig's maps are limited to string keys, I think we'll need to use tuples. I'm ok with fixing bugs here, provided that's in a separate patch.","created":"2013-08-26T23:14:12.935+0000"},{"body":"First apply 5867-bug-fix-filter-push-down-1.2-branch.txt to fix issue with filter push down. (schema changes), then apply 5867-4-1.2-branch.txt for collection supports in CqlStorage which map SetType/ListType to tuple, MapType to tuple of tuples.\n\n","created":"2013-08-27T04:22:16.017+0000"},{"body":"5867-5-1.2-branch.txt is attached to fix a bug related to writing map collection to C*. It should be applied after 5867-bug-fix-filter-push-down-1.2-branch.txt","created":"2013-08-28T23:56:20.880+0000"},{"body":"e.g.\n{code}\n CREATE TABLE test(\n m text PRIMARY KEY,\n n map\n ) \n data format for insertion (((m,kk)),((map,(m,mm),(n,nn))))\n store recs into 'cql://test/test77?output_query=update+test.test+set+n+%3D+%3F' using CqlStorage();\n where output_query is url encoded\n{code}","created":"2013-08-29T00:07:17.003+0000"},{"body":"Committed, thanks","created":"2013-09-10T18:47:28.575+0000"}],"conversations":[{"body":"The CqlStorage class gets the Pig data type for values from the AbstractCassandraStorage class, in the getPigType method. If it isn't a known data type, it makes the value into a ByteArray. Currently there aren't any cases there for lists, maps, and sets.\nhttps://github.com/apache/cassandra/blob/cassandra-1.2.8/src/java/org/apache/cassandra/hadoop/pig/AbstractCassandraStorage.java#L336\n\nSee this describe output from the grunt shell:\n\n{code}\ngrunt> describe listdata ; \nlistdata: {id: (name: chararray,value: int),alist: (name: chararray,value: bytearray),amap: (name: chararray,value: bytearray),aset: (name: chararray,value: bytearray)}\n{code}\n\nwhere the cql data structures had this schema:\n\n{code}\nCREATE TABLE alltypes (\n id int PRIMARY KEY,\n alist list,\n amap map,\n aset set\n{code}\n\nIt turns out that if you cast the map in grunt to a pig map, then it sort of works, but I don't think we should probably use a pig map. Lists don't appear to work at all, as there is no Pig analogue. I *think* you could probably just do a UDF to cast these things, but we already have all of the type information, so we just need to change them to tuples or bags or whatever.","from":"reporter","subject":"The Pig CqlStorage/AbstractCassandraStorage classes don't handle collection types"},{"body":"A tuple is the obvious choice, since a bag is cumbersome and there's no need to spill to disk.","from":"developer"},{"body":"In the same vein as data types, as of PIG-2764 Pig has a BigInteger. So once 0.12 is out and mainstream, we could look at using that instead of the Pig integer type to avoid overflow. Just as a heads up.","from":"developer"},{"body":"Patch on 1.2 branch is attached, which map ListType and SetType to tuple, MapType to map","from":"developer"},{"body":"We may not want to use the MapType for maps.\n\nIn Pig, you can only have text based keys for maps. From Programming Pig: \"A map in Pig is a chararray to data element mapping, where that element can be any Pig type, including a complex type.\"\n\nIn Cassandra, you can have map keys of any type, like timestamp in the maps example on http://www.datastax.com/dev/blog/cql3_collections\n\nSo we may want to do a bag or tuple of tuples.","from":"developer"},{"body":"The key of Cassandra map has been converted to string into Pig map. We need to decide whether use tuple of tuples vs map for Cassandra map type. Tuple of tuples is more general than map, and map is more specific. HBase uses map to map its row. So which one to use for map? Map vs Tuple of tuples?","from":"developer"},{"body":"To store data to CQL3 table, we supports the following prepared statements\n\n{code}\nList\n e.g.\n UPDATE users SET top_places = ?\n UPDATE users SET top_places = [ 'rivendell', 'rohan' ] WHERE user_id = 'frodo';\n\n UPDATE users SET top_places = ? + top_places\n UPDATE users SET top_places = [ 'the shire' ] + top_places WHERE user_id = 'frodo';\n\n UPDATE users SET top_places = top_places - ?;\n UPDATE users SET top_places = top_places - ['riddermark'] WHERE user_id = 'frodo';\n\nSet statements are similar to List\n\nMap\n\n UPDATE users SET todo = ?\n UPDATE users\n SET todo = { '2012-9-24' : 'enter mordor',\n '2012-10-2 12:00' : 'throw ring into mount doom' }\n WHERE user_id = 'frodo';\n\n\nThe following queries are handled as a regular value instead of tuples\n UPDATE users SET top_places[2] = ?\n UPDATE users SET top_places[2] = 'riddermark' WHERE user_id = 'frodo';\n\n DELETE top_places[3] FROM users;\n DELETE top_places[3] FROM users WHERE user_id = 'frodo';\n\n UPDATE users SET todo[?] = ?\n UPDATE users SET todo['2012-10-2 12:10'] = 'die' WHERE user_id = 'frodo';\n{code}\n\nThe output schema for collections is as following\n\n{code}\n (((name, value), (name, value)), (value ... value), (value...value)\n If a value of tuple (value...value) is a tuple of (inner_value ...inner_value) \n and the first inner_value is the collection type. \n it is either \"set\", \"list\" or \"map\".\n\n e.g. (value ... value) as (value1, value2, (set, riddermark, tom))\n map (value1, value2, (map, (sfo, 12), (ny, 34))\n{code}\n","from":"developer"},{"body":"5867-3-1.2-branch.txt is attached to support storing collections to Cassandra","from":"developer"},{"body":"If the key of C* map is converted to string (chararray) in a Pig map on read, how will the keys be handled on write? Will the strings be auto-converted to appropriate C* types? In the past, I've seen issues with that - especially with timestamps and UUID values.","from":"developer"},{"body":"The following is the data which will be auto converted to C* type, anything else will be bytes.\n\n{code}\n if (o == null)\n return (ByteBuffer)o;\n if (o instanceof java.lang.String)\n return ByteBuffer.wrap(new DataByteArray((String)o).get());\n if (o instanceof Integer)\n return Int32Type.instance.decompose((Integer)o);\n if (o instanceof Long)\n return LongType.instance.decompose((Long)o);\n if (o instanceof Float)\n return FloatType.instance.decompose((Float)o);\n if (o instanceof Double)\n return DoubleType.instance.decompose((Double)o);\n if (o instanceof UUID)\n return ByteBuffer.wrap(UUIDGen.decompose((UUID) o));\n\n return ByteBuffer.wrap(((DataByteArray) o).get());\n{code}\n\nyou need prepare the data for write different from read result.","from":"developer"},{"body":"Alex, is this ready for review?","from":"developer"},{"body":"Yes, But I map the C* MapType to Pig map, we may need map it to tuples of tuples.","from":"developer"},{"body":"I also has a few bug fixes related to CqlStorage, so if it's possible, I can add those fix to this patch as well.","from":"developer"},{"body":"Since pig's maps are limited to string keys, I think we'll need to use tuples. I'm ok with fixing bugs here, provided that's in a separate patch.","from":"developer"},{"body":"First apply 5867-bug-fix-filter-push-down-1.2-branch.txt to fix issue with filter push down. (schema changes), then apply 5867-4-1.2-branch.txt for collection supports in CqlStorage which map SetType/ListType to tuple, MapType to tuple of tuples.\n\n","from":"developer"},{"body":"5867-5-1.2-branch.txt is attached to fix a bug related to writing map collection to C*. It should be applied after 5867-bug-fix-filter-push-down-1.2-branch.txt","from":"developer"},{"body":"e.g.\n{code}\n CREATE TABLE test(\n m text PRIMARY KEY,\n n map\n ) \n data format for insertion (((m,kk)),((map,(m,mm),(n,nn))))\n store recs into 'cql://test/test77?output_query=update+test.test+set+n+%3D+%3F' using CqlStorage();\n where output_query is url encoded\n{code}","from":"developer"},{"body":"Committed, thanks","from":"developer"}],"created":"2013-08-09T14:04:07.000+0000","description":"The CqlStorage class gets the Pig data type for values from the AbstractCassandraStorage class, in the getPigType method. If it isn't a known data type, it makes the value into a ByteArray. Currently there aren't any cases there for lists, maps, and sets.\nhttps://github.com/apache/cassandra/blob/cassandra-1.2.8/src/java/org/apache/cassandra/hadoop/pig/AbstractCassandraStorage.java#L336\n\nSee this describe output from the grunt shell:\n\n{code}\ngrunt> describe listdata ; \nlistdata: {id: (name: chararray,value: int),alist: (name: chararray,value: bytearray),amap: (name: chararray,value: bytearray),aset: (name: chararray,value: bytearray)}\n{code}\n\nwhere the cql data structures had this schema:\n\n{code}\nCREATE TABLE alltypes (\n id int PRIMARY KEY,\n alist list,\n amap map,\n aset set\n{code}\n\nIt turns out that if you cast the map in grunt to a pig map, then it sort of works, but I don't think we should probably use a pig map. Lists don't appear to work at all, as there is no Pig analogue. I *think* you could probably just do a UDF to cast these things, but we already have all of the type information, so we just need to change them to tuples or bags or whatever.","issue_id":"12662857","key":"CASSANDRA-5867","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-09-10T18:47:28.000+0000","role":"fixed_distractor","summary":"The Pig CqlStorage/AbstractCassandraStorage classes don't handle collection types"} {"case_id":"12665333","cluster":"DISTRACTOR-CASSANDRA-5930","comments":[{"body":"Can you have a look, Jason?","created":"2013-08-23T21:57:29.156+0000"},{"body":"We're seeing this too -- slightly different stack trace, which I'll include here in case it's of use.\n\n\nWARNING: Non-fatal error reading row (stacktrace follows)\nException in thread \"main\" java.io.IOError: java.lang.IllegalArgumentException\nat org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:244)\nat org.apache.cassandra.tools.StandaloneScrubber.main(StandaloneScrubber.java:125)\nCaused by: java.lang.IllegalArgumentException \nat java.nio.Buffer.limit(Buffer.java:247)\nat org.apache.cassandra.db.marshal.AbstractCompositeType.getBytes(AbstractCompositeType.java:51)\nat org.apache.cassandra.db.marshal.AbstractCompositeType.getWithShortLength(AbstractCompositeType.java:60)\nat org.apache.cassandra.db.marshal.AbstractCompositeType.compare(AbstractCompositeType.java:78) \nat org.apache.cassandra.db.marshal.AbstractCompositeType.compare(AbstractCompositeType.java:31)\nat org.apache.cassandra.db.ArrayBackedSortedColumns.addColumn(ArrayBackedSortedColumns.java:128)\nat org.apache.cassandra.db.AbstractColumnContainer.addColumn(AbstractColumnContainer.java:114)\nat org.apache.cassandra.db.AbstractColumnContainer.addColumn(AbstractColumnContainer.java:109) \nat org.apache.cassandra.db.ColumnFamily.addAtom(ColumnFamily.java:219)\nat org.apache.cassandra.db.ColumnFamilySerializer.deserializeColumnsFromSSTable(ColumnFamilySerializer.java:149)\nat org.apache.cassandra.io.sstable.SSTableIdentityIterator.getColumnFamilyWithColumns(SSTableIdentityIterator.java:234)\nat org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:114) \nat org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:98)\nat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:160)\nat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:166)\nat org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:173) \n... 1 more\n","created":"2013-09-02T16:49:04.099+0000"},{"body":"[~jeffpotter] what version of Cassandra were you running when you hit the above error?\n\nAs far as the original stacktrace for this ticket goes, it's unfortunately necessary for counter CFs. CASSANDRA-2759 explains the reasoning. I suppose I could make the error message mention that and point to the ticket.\n\nThe scrub code looks reasonably robust in general, so I think it's better to wait for individual bugs to get reported than to try to improve the code without any failure examples.","created":"2014-01-02T18:50:18.836+0000"},{"body":"Hi Tyler -- based on my notes, it should have been Cassandra 1.2.6.1 (DSE 3.1), at least, that's what other tickets we have filed at this same time suggest.","created":"2014-01-02T23:02:28.640+0000"},{"body":"Thanks, [~jeffpotter]. It looks like you also had a Counter table, in this case.\n\n5930-v1.patch (and [branch|https://github.com/thobbs/cassandra/tree/CASSANDRA-5930-2.0]) clarifies the error message for counter tables.","created":"2014-01-03T22:39:30.557+0000"},{"body":"It would be nice to be able to tell people how to fix it (realistically: what their options are) rather than just \"sorry, scrub can't help you.\" But I'm not sure what those options are. :) /cc [~slebresne] [~iamaleksey]","created":"2014-01-03T23:52:18.105+0000"},{"body":"[~jbellis] There are no options. That said, we should probably allow users to override this behavior, if they prefer losing some of the counters history to not scrubbing at all.\n\nAlso, with CASSANDRA-6504 in this becomes a non-issue (for the newly written 'global' 2.1 shards, at least - we *can* repair those after the scrub).","created":"2014-01-04T00:00:33.870+0000"},{"body":"You're right, it would be nice to give the user an option to skip the corrupted rows anyway.\n\n5930-v2.patch (the [branch|https://github.com/thobbs/cassandra/tree/CASSANDRA-5930-2.0] is still good) adds a {{--skip-corrupted}} option and a unit test to exercise it.","created":"2014-01-07T21:40:21.470+0000"},{"body":"LGTM, committed.","created":"2014-02-03T20:33:07.759+0000"}],"conversations":[{"body":"There are cases where offline scrub can hit an exception and die, like:\n\n{noformat}\nWARNING: Non-fatal error reading row (stacktrace follows)\nException in thread \"main\" java.io.IOError: java.io.IOError: java.io.EOFException\n\tat org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:242)\n\tat org.apache.cassandra.tools.StandaloneScrubber.main(StandaloneScrubber.java:121)\nCaused by: java.io.IOError: java.io.EOFException\n\tat org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:116)\n\tat org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:99)\n\tat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:176)\n\tat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:182)\n\tat org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:171)\n\t... 1 more\nCaused by: java.io.EOFException\n\tat java.io.RandomAccessFile.readFully(RandomAccessFile.java:399)\n\tat java.io.RandomAccessFile.readFully(RandomAccessFile.java:377)\n\tat org.apache.cassandra.utils.BytesReadTracker.readFully(BytesReadTracker.java:95)\n\tat org.apache.cassandra.utils.ByteBufferUtil.read(ByteBufferUtil.java:401)\n\tat org.apache.cassandra.utils.ByteBufferUtil.readWithLength(ByteBufferUtil.java:363)\n\tat org.apache.cassandra.db.ColumnSerializer.deserialize(ColumnSerializer.java:120)\n\tat org.apache.cassandra.db.ColumnSerializer.deserialize(ColumnSerializer.java:37)\n\tat org.apache.cassandra.db.ColumnFamilySerializer.deserializeColumns(ColumnFamilySerializer.java:144)\n\tat org.apache.cassandra.io.sstable.SSTableIdentityIterator.getColumnFamilyWithColumns(SSTableIdentityIterator.java:234)\n\tat org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:112)\n\t... 5 more\n{noformat}\n\nSince the purpose of offline scrub is to fix broken stuff, it should be more resilient to broken stuff...","from":"reporter","subject":"Offline scrubs can choke on broken files"},{"body":"Can you have a look, Jason?","from":"developer"},{"body":"We're seeing this too -- slightly different stack trace, which I'll include here in case it's of use.\n\n\nWARNING: Non-fatal error reading row (stacktrace follows)\nException in thread \"main\" java.io.IOError: java.lang.IllegalArgumentException\nat org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:244)\nat org.apache.cassandra.tools.StandaloneScrubber.main(StandaloneScrubber.java:125)\nCaused by: java.lang.IllegalArgumentException \nat java.nio.Buffer.limit(Buffer.java:247)\nat org.apache.cassandra.db.marshal.AbstractCompositeType.getBytes(AbstractCompositeType.java:51)\nat org.apache.cassandra.db.marshal.AbstractCompositeType.getWithShortLength(AbstractCompositeType.java:60)\nat org.apache.cassandra.db.marshal.AbstractCompositeType.compare(AbstractCompositeType.java:78) \nat org.apache.cassandra.db.marshal.AbstractCompositeType.compare(AbstractCompositeType.java:31)\nat org.apache.cassandra.db.ArrayBackedSortedColumns.addColumn(ArrayBackedSortedColumns.java:128)\nat org.apache.cassandra.db.AbstractColumnContainer.addColumn(AbstractColumnContainer.java:114)\nat org.apache.cassandra.db.AbstractColumnContainer.addColumn(AbstractColumnContainer.java:109) \nat org.apache.cassandra.db.ColumnFamily.addAtom(ColumnFamily.java:219)\nat org.apache.cassandra.db.ColumnFamilySerializer.deserializeColumnsFromSSTable(ColumnFamilySerializer.java:149)\nat org.apache.cassandra.io.sstable.SSTableIdentityIterator.getColumnFamilyWithColumns(SSTableIdentityIterator.java:234)\nat org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:114) \nat org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:98)\nat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:160)\nat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:166)\nat org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:173) \n... 1 more\n","from":"developer"},{"body":"[~jeffpotter] what version of Cassandra were you running when you hit the above error?\n\nAs far as the original stacktrace for this ticket goes, it's unfortunately necessary for counter CFs. CASSANDRA-2759 explains the reasoning. I suppose I could make the error message mention that and point to the ticket.\n\nThe scrub code looks reasonably robust in general, so I think it's better to wait for individual bugs to get reported than to try to improve the code without any failure examples.","from":"developer"},{"body":"Hi Tyler -- based on my notes, it should have been Cassandra 1.2.6.1 (DSE 3.1), at least, that's what other tickets we have filed at this same time suggest.","from":"developer"},{"body":"Thanks, [~jeffpotter]. It looks like you also had a Counter table, in this case.\n\n5930-v1.patch (and [branch|https://github.com/thobbs/cassandra/tree/CASSANDRA-5930-2.0]) clarifies the error message for counter tables.","from":"developer"},{"body":"It would be nice to be able to tell people how to fix it (realistically: what their options are) rather than just \"sorry, scrub can't help you.\" But I'm not sure what those options are. :) /cc [~slebresne] [~iamaleksey]","from":"developer"},{"body":"[~jbellis] There are no options. That said, we should probably allow users to override this behavior, if they prefer losing some of the counters history to not scrubbing at all.\n\nAlso, with CASSANDRA-6504 in this becomes a non-issue (for the newly written 'global' 2.1 shards, at least - we *can* repair those after the scrub).","from":"developer"},{"body":"You're right, it would be nice to give the user an option to skip the corrupted rows anyway.\n\n5930-v2.patch (the [branch|https://github.com/thobbs/cassandra/tree/CASSANDRA-5930-2.0] is still good) adds a {{--skip-corrupted}} option and a unit test to exercise it.","from":"developer"},{"body":"LGTM, committed.","from":"developer"}],"created":"2013-08-23T21:54:26.000+0000","description":"There are cases where offline scrub can hit an exception and die, like:\n\n{noformat}\nWARNING: Non-fatal error reading row (stacktrace follows)\nException in thread \"main\" java.io.IOError: java.io.IOError: java.io.EOFException\n\tat org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:242)\n\tat org.apache.cassandra.tools.StandaloneScrubber.main(StandaloneScrubber.java:121)\nCaused by: java.io.IOError: java.io.EOFException\n\tat org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:116)\n\tat org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:99)\n\tat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:176)\n\tat org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:182)\n\tat org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:171)\n\t... 1 more\nCaused by: java.io.EOFException\n\tat java.io.RandomAccessFile.readFully(RandomAccessFile.java:399)\n\tat java.io.RandomAccessFile.readFully(RandomAccessFile.java:377)\n\tat org.apache.cassandra.utils.BytesReadTracker.readFully(BytesReadTracker.java:95)\n\tat org.apache.cassandra.utils.ByteBufferUtil.read(ByteBufferUtil.java:401)\n\tat org.apache.cassandra.utils.ByteBufferUtil.readWithLength(ByteBufferUtil.java:363)\n\tat org.apache.cassandra.db.ColumnSerializer.deserialize(ColumnSerializer.java:120)\n\tat org.apache.cassandra.db.ColumnSerializer.deserialize(ColumnSerializer.java:37)\n\tat org.apache.cassandra.db.ColumnFamilySerializer.deserializeColumns(ColumnFamilySerializer.java:144)\n\tat org.apache.cassandra.io.sstable.SSTableIdentityIterator.getColumnFamilyWithColumns(SSTableIdentityIterator.java:234)\n\tat org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:112)\n\t... 5 more\n{noformat}\n\nSince the purpose of offline scrub is to fix broken stuff, it should be more resilient to broken stuff...","issue_id":"12665333","key":"CASSANDRA-5930","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-02-03T20:33:07.000+0000","role":"fixed_distractor","summary":"Offline scrubs can choke on broken files"} {"case_id":"12665353","cluster":"DISTRACTOR-CASSANDRA-5932","comments":[{"body":"[~enigmacurry], what was your collection interval for the metrics you've collected?\n\nI ask because I'm observing results similar to yours for \"3) Eager Reads tend to lessen the immediate performance impact of a node going down, but not consistently.\". However, I've polled metrics @ a 1 second granularity to see that it's actually a multi-second stress-client outage - not just poor and inconsistent performance.\n\nPolling metrics @ a 1 second interval, has observations that a ~20 second read operations starvation outage occurs for the stress client for all data in the cluster (even with the lowest phi_convict_threshold=6).\n\nAnalysis so far indicates that high-operations reads starve out all the C* client threads/connections, because they get stuck on awaiting for a server response whenever the key-space hits the node that is down (and by probability + high-operation reads, within 1 second each stress client thread will all hit the downed-node's key-space).\n\n\n\nSo I'm confirming that I'm also seeing this bug that Speculative reads (even with an ALWAYS setting). It isn't solving this outage for clients during high-operation reads, and based on what I understand of the feature, it should.\n\nThanks, guys!\n\n","created":"2013-09-16T16:03:55.944+0000"},{"body":"I'm using 5 second intervals in these charts. 'multi-second stress-client outage' is a good way to put it, for both the case of speculative retry and not, the drop in performance after a node goes down is a duration of complete non-responsiveness (not degraded performance.) The addition of speculative retry consistently shortens this duration (it's always better), but this duration itself is inconsistent. ","created":"2013-09-16T17:41:23.906+0000"},{"body":"SpeculateAlwaysExecutor - Here we are not reading from more endpoints than normal. We are only reading data from two endpoints. We should be reading from one more endpoint if possible. ","created":"2013-09-17T23:39:26.560+0000"},{"body":"Thanks, [~enigmacurry]. Sounds like you are seeing the same thing as us, so it's great to see it's getting attention!","created":"2013-09-18T19:28:52.643+0000"},{"body":"Hello [~iamaleksey],\n\nThanks for the link to this jira and for your very detailed testing results. It confirms what we have seen in our lab testing for the Cassandra 2.0.0-rc2 \"Speculative Execution for Reads\".\n\nWe have a very simple data center setup consisting of four Cassandra nodes running on four server machines. A testing application (Cassandra client) is interacting with Cassandra nodes 1, 2 and 3. That is, the testing app does not directly connected to the Cassandra node 4.\n\nThe keyspace Replication Factor is set to 3 and the client requested Consistency Level is set to CL_TWO.\n\nI have tested all of three configurations of the Speculative Execution for Reads ('ALWAYS', '85 PERCENTILE', '50 MS' / '100 MS'). It seems that none of them works as expected. From the test app log file point of view, they all give a 20-second window of outage immediately after the 4th node was killed. This behavior is consistent to Cassandra 1.2.4.\n\nI have done a quick code reading of the Cassandra Server implementation (Cassandra 2.0.0 tarball) and I have noticed some design issues. I would like to discuss them with you.\n\n*Issue 1* - StorageProxy.fetchRows() may still block for as long as conf.read_request_timeout_in_ms, though the speculative retry did fire correctly after the Cassandra node 4 was killed.\n\nTake the speculative configuration of 'PERCENTILE' / 'CUSTOM' as example, after the Cassandra node 4 was killed, SpeculativeReadExecutor.speculate() would block for responses. If timed out, it would send out one more read request to an alternative node (from {{unfiltered}}) and increment the speculativeRetry counter. This part should work.\n\nHowever, killing the 4th node would very likely cause inconsistency in the database and this will trigger the DigestMismatchException. In the fetchRows(), when handling DigestMismatchException, it uses handler.endpoints to send out digest mismatch retries and then block for responses. As we know that one of the endpoints was already killed, the handler.get() will block until it is timed out, which is 10 seconds.\n\n\n{noformat}\n catch (DigestMismatchException ex)\n {\n Tracing.trace(\"Digest mismatch: {}\", ex);\n\n ...\n\n MessageOut message = exec.command.createMessage();\n for (InetAddress endpoint : exec.handler.endpoints)\n {\n Tracing.trace(\"Enqueuing full data read to {}\", endpoint);\n MessagingService.instance().sendRR(message, endpoint, repairHandler);\n }\n }\n }\n\n ...\n\n // read the results for the digest mismatch retries\n if (repairResponseHandlers != null)\n {\n for (int i = 0; i < repairCommands.size(); i++)\n {\n ReadCommand command = repairCommands.get(i);\n ReadCallback handler = repairResponseHandlers.get(i);\n\n Row row;\n try\n {\n row = handler.get();\n }\n{noformat}\n\n\n\n*Issue 2* - The speculative 'ALWAYS' does NOT send out any more read requests. Thus, in face of the failure of node 4, it will not help at all.\n\nThe SpeculateAlwaysExecutor.executeAsync() only sends out handler.endpoints.size() number of read requests and it blocks for the responses to come back. If one of the nodes is killed, say node 4, this speculative retry 'ALWAYS' will work the same way as Cassandra 1.2.4, i.e. it will block until timed out, which is 10 seconds.\n\n??My understanding of this speculative retry 'ALWAYS' should ALWAYS send out \"handler.endpoints.size() + 1\" number of read requests and block for handler.endpoints.size() number of responses??.\n\n*Issue 3* - Since the ReadRepairDecison is determined by a Random() number, this speculative retry may not work as the ReadRepairDecision may be ??ReadRepairDecision.GLOBAL??\n\n*Issue 4* - For the ReadExecutor(s), the {{this.unfiltered}} and {{this.endpoints}} may not consistent. Thus, using {{this.unfiltered}} and {{this.endpoints}} for speculative retry may cause unexpected results. This is especially true when the Consistency Level is {{LOCAL_QUARUM}} and the ReadRepairDecision is {{DC_LOCAL}}.\n\n\n\n","created":"2013-09-24T21:53:19.233+0000"},{"body":"Hey [~lizou]. Yeah, I've fixed most of these already (rewritten most of the ARE code, actually). Specifically issues 2,3,4. Will look into 1 too.\n\nThanks.","created":"2013-09-24T21:58:37.755+0000"},{"body":"Attaching 5932.txt that will hopefully fix this ([~enigmacurry] could you run the tests again, please, with the patch applied?)\n\n1. As noted by [~lizou] and [~kohlisankalp], ALWAYS wasn't making an extra request, it was making an extra data request at the expense of one digest request. Fixed.\n\n2. SpecRetry wasn't working correctly with RRD.DC_LOCAL, as noted by @lizou, because the two lists will be in different order, and a retry might be sent to a node that already had a request sent to it. (Please note that LOCAL_QUORUM here does not affect anything - CL.filterForQuery() sorts in place, so the two lists would be in the same order, everything was working correct). RRD.DC_LOCAL handling was a legit issue though. Fixed.\n\n3. SpecRetry w/ RRD.GLOBAL is a noop, you can't speculate if you contact all the replicas in the first place. This is normal.\n\n4. The DME issue is semi-legit. Killing a node shouldn't trigger DME or increase the likelihood of DME happening. *HOWEVER* when shooting requests for repair, we were not considering the case where one of the replies satisfying the original CL came from a SpecRetry attempt. The patch includes the extra replica in repair commands if SpecRetry had been triggered by the original request.\n\n\n5. SP.getRangeSlice() is not SpectRetry-aware as of now. I don't know if this is an omission or by design, but for now, please don't include that in the benchmarks, since it would only be misleading. ","created":"2013-09-25T01:51:46.791+0000"},{"body":"Pushed some OCD of my own to https://github.com/jbellis/cassandra/commits/5932 on top of this.\n\nbq. ALWAYS wasn't making an extra request, it was making an extra data request at the expense of one digest request\n\nI'm not sure what the distinction is here. Do you mean that if we weren't read-repairing, there would be no extra data request at all?\n\nbq. SpecRetry w/ RRD.GLOBAL is a noop, you can't speculate if you contact all the replicas in the first place.\n\nI dunno, I think we should turn a digest into a data for redundancy the way ALWAYS used to.","created":"2013-09-25T16:07:35.218+0000"},{"body":"Pushed even more OCD to https://github.com/iamaleksey/cassandra/commits/5932 on top of yours.\n\nbq. I'm not sure what the distinction is here. Do you mean that if we weren't read-repairing, there would be no extra data request at all?\n\nInstead of making, for example 1 data request + 2 digest requests, ALWAYS was making 2 data requests + 1 digest request, instead of making 2 data requests + 2 digest requests, not really helping to satisfy the CL in case of node's failure.\n\nbq. I dunno, I think we should turn a digest into a data for redundancy the way ALWAYS used to.\n\nMaybe.","created":"2013-09-25T17:02:13.299+0000"},{"body":"Force-pushed the 'final' version to https://github.com/iamaleksey/cassandra/commits/5932.\n\nAmong other things, properly handles RRD.GLOBAL and RRD.DC_LOCAL in 1-DC scenario.","created":"2013-09-25T23:03:45.068+0000"},{"body":"Pushed one more set of changes to mine, not forced: https://github.com/jbellis/cassandra/commits/5932. Goal is to make SRE less fragile when doing RR.","created":"2013-09-26T17:54:24.042+0000"},{"body":"+1, I'm out of OCD juice.","created":"2013-09-26T20:30:52.826+0000"},{"body":"Hello [~iamaleksey] and [~jbellis],\n\nI took a quick look at the code changes. The new code looks very good to me. But I saw one potential issue in {{AlwaysSpeculatingReadExecutor.executeAsync()}}, in which it makes at least *two* data / digest requests. This will cause problems for a data center with only one Cassandra server node (e.g. bring up an embedded Cassandra node in JVM for JUnit test) or a deployed production data center of two Cassandra server nodes with one node shut down for maintenance. In the above mentioned two cases, {{AbstractReadExecutor.getReadExecutor()}} will return the {{AlwaysSpeculatingReadExecutor}} as condition {{(targetReplicas.size() == allReplicas.size())}} is met, though the tables may / may not be configured with ??Speculative ALWAYS??.\n\nIt is true for our legacy products we are considering to deploy each data center with only two Cassandra server nodes with RF = 2 and CL = 1.\n","created":"2013-09-26T20:43:02.768+0000"},{"body":"The logic looks like this:\n\n# Figure out how many replicas we need to contact to satisfy the desired consistencyLevel + Read Repair settings\n# If that ends up being all the replicas, then use ASRE to get some redundancy on the data reads. This will allow the read to succeed even if a digest for RR times out. Of course if you are reading at CL.ALL and a replica times out there's nothing we can do.\n# Otherwise, use SRE and make an \"extra\" request later, if it looks like one of the minimal set isn't going to respond in time\n\nNote that performing extra data requests does not affect handler.blockfor -- just makes it possible for the request to proceed if it gets enough responses back, no matter which replicas they come from.","created":"2013-09-26T20:55:01.295+0000"},{"body":"(Committed after Aleksey's +1, incidentally.)","created":"2013-09-26T20:55:41.621+0000"},{"body":"The logic for {{AlwaysSpeculatingReadExecutor}} is good. What I meant in my previous comment is that when {{targetReplicas.size() == allReplicas.size()}} and {{targetReplicas.size() == 1}}, then {{AlwaysSpeculatingReadExecutor.executeAsync()}} will throw an exception as there is only one endpoint in {{targetReplicas}}, but it tries to access two endpoints in {{targetReplicas}}.","created":"2013-09-26T21:04:25.213+0000"},{"body":"I see what you mean. Fixed in 7a87fc1186f39678382cf9b3e1dd224d9c71aead.","created":"2013-09-26T21:10:48.798+0000"},{"body":"The good news is that speculative read has improved across the board.\n\nHowever, this new batch of testing introduces some new mysteries.\n\nHere is all of the runs from 7a87fc1186f39678382cf9b3e1dd224d9c71aead:\n\n!5933-7a87fc11.png!\n\nAll of the speculative retry runs are better than with 2.0.0-rc1. However, I can't explain why sr=NONE did better than ALWAYS and 95percentile. There is no visible indication that a node went down for sr=NONE. I have double checked the logs, and it did, in fact, go down. \n\nCompare this to the baseline of 1.2.8 and 2.0.0-rc1 (redone last night on same hardware as above):\n\n!5933-128_and_200rc1.png!\n\nAll of these have clear indications of the node going down.\n\nYou can [see all the data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1] - you can double click the colored squares to toggle the visibility of the lines, as they do overlap.\n\nI've uploaded logs from all these runs as 5933-logs.tar.gz.","created":"2013-09-27T14:54:55.162+0000"},{"body":"Yeah. The graphs for ALWAYS and NONE look swapped from what I would expect.","created":"2013-09-27T15:00:13.221+0000"},{"body":"The other thing that I note, is that all of these runs are better than 1.2.8 :D (further evidence that CASSANDRA-5933 may be invalid)","created":"2013-09-27T15:02:06.068+0000"},{"body":"Hmm.\n\nI wonder if it's just luck of the draw as to which replica dsnitch is preferring. Here's a branch to randomize that, per-operation:\n\nhttps://github.com/jbellis/cassandra/commits/5932-randomized","created":"2013-09-27T15:07:02.812+0000"},{"body":"Couldn't we just do a run with the dsnitch disabled?","created":"2013-09-27T16:28:10.510+0000"},{"body":"That still gives you luck-of-the-draw as to which replica it prefers. (Unlikely to be evenly distributed.)","created":"2013-09-27T16:31:45.112+0000"},{"body":"This looks exactly like what I was expecting:\n\n!5933-randomized-dsnitch-replica.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.randomized-dsnitch.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]","created":"2013-09-27T17:14:29.654+0000"},{"body":"That's awesome.\n\nCan you test 90th and 75th percentile too?","created":"2013-09-27T17:29:07.078+0000"},{"body":"Seems like it's quite tunable, but not a lot of difference under 90%:\n\n!5933-randomized-dsnitch-replica.2.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.randomized-dsnitch.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]","created":"2013-09-27T18:44:44.375+0000"},{"body":"Yeah, that makes sense. 70th..95th are all pretty damn close to median still.\n\nWhat I'd like to do is get close to the 10ms performance hit (~none) as a percentile, and make that default in 2.1. Try 99th and 99.9th?","created":"2013-09-27T19:43:24.004+0000"},{"body":"Also, did that stress patch work to get you failed request counts? Would be good to get that too if we can show that even 10ms keeps requests from failing entirely.","created":"2013-09-27T19:54:22.517+0000"},{"body":"Here's 99th and 99.9th percentiles:\n\n!5933-randomized-dsnitch-replica.3.png!\n\nI'll look at that stress patch again, I seem to recall it not making a lot of sense to me when I last tried it, but will give it another go.","created":"2013-09-27T21:14:45.982+0000"},{"body":"Starting to think we still have a bug. 99.9 should be doing less retries than 10ms but the graph shows it doing more.","created":"2013-09-27T21:37:35.728+0000"},{"body":"I did some more code reading and noticed some potential issues and possible improvement. I've got to run now. I will get back to you guys Monday morning.\n\nMy guess is that the ??Speculative NONE?? is hit by the initial request reading path which is successfully resolved by the *Speculative Retry*. The observed throughput performance hit when Speculative Retry is enabled is caused by the ReadRepair path which has some coding / design issues. I will talk to you next Monday.","created":"2013-09-27T21:47:54.166+0000"},{"body":"Maybe the problem is that we're using CF-level latency instead of StorageProxy.\n\nWhat does cfhistograms give for read latency?","created":"2013-09-27T21:49:13.925+0000"},{"body":"Pretty sure that's our smoking gun. Pushed a commit to the -randomized branch that adds coordinator-level, per-cf latency tracking and uses that instead.\n\nCan you repeat the last test with that? (Maybe throw in ALL as well if you're feeling optimistic that we'll have a measurable difference between ALL and 90%. :)","created":"2013-09-27T22:40:34.906+0000"},{"body":"[~jbellis] Here's your two runs:\n\n[dea27f84f40|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.dea27f84f40.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n[ded39c7e1c2fa|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.ded39c7e1c2fa.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n\nLogs for the second run are attached as 5932.ded39c7e1c2fa.logs.tar.gz","created":"2013-09-29T16:15:40.781+0000"},{"body":"The code to convert 99 into 0.99 was buggy and was actually converting to 0.0099. Fix pushed, can you try it again?","created":"2013-09-29T17:48:46.849+0000"},{"body":"BINGO!\n\n!5932-6692c50412ef7d.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.6692c50412ef7d.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]","created":"2013-09-29T21:26:45.993+0000"},{"body":"Definitely the best results seen so far! Nice work!\n\nIn my mind, during the transition period right after the killing of the node, I expected \"ALWAYS\" to have negligible impact, or at least the smallest impact of all of the other values. However, red (90%), and purple (75%) is having a smaller impact. Seems fishy. Do I misunderstand the intention of the \"ALWAYS\" setting?\n\n(Edited to clarify the period I'm talking about.)","created":"2013-09-30T01:25:25.731+0000"},{"body":"It looks to me like 75/90/Always are about the same, with Always dropping from a lower baseline. Which makes sense; it's still doing a lot of unnecessary work compared to the others.","created":"2013-09-30T03:10:59.283+0000"},{"body":"[~enigmacurry], can you also re-test the uncapped compaction scenario with the same set of retry settings?","created":"2013-09-30T13:49:56.572+0000"},{"body":"Hello [~iamaleksey] and [~jbellis],\n\nIt appears to me that the testing results have suggested that the \"_data read + speculative retry_\" path work as expected. This \"_data read + speculative retry_\" path has greatly minimized the throughput impact caused by the failure of one of Cassandra server nodes.\n\nThe observed small degradation of throughput performance when _speculative retry_ is enabled is very likely to be caused by the \"*_read repair_*\" path. I did the code reading of this path last Friday and noticed some design / coding issues. I would like to discuss them with you.\n\nPlease note that my code base is still the Cassandra 2.0.0 tarball, not updated with the latest code changes.\n\n*Issue 1* -- When handling {{DigestMismatchException}} in {{StorageProxy.fetchRows()}}, all _data read requests_ are sent out using {{sendRR}} without distinguishing remote nodes from the local node.\n\nWill this cause an issue, as {{MessagingService.instance().sendRR()}} will send out enqueued messages for a specified remote node via its pre-established TCP socket connection. For local node, this should be done via {{LocalReadRunnable}}, i.e. {{StageManager.getStage(Stage.READ).execute(new LocalReadRunnable(command, handler))}}.\n\nIf this may cause an issue, the following wait may block.\n\n{noformat}\n // read the results for the digest mismatch retries\n if (repairResponseHandlers != null)\n {\n for (int i = 0; i < repairCommands.size(); i++)\n {\n ReadCommand command = repairCommands.get(i);\n ReadCallback handler = repairResponseHandlers.get(i);\n\n Row row;\n try\n {\n row = handler.get();\n }\n{noformat}\n\nFor two reasons.\n* The data read request for local node may never sent out\n* As one of the nodes is down (which triggered the Speculative Retry) will cause one missing response.\n\n*If missing two responses, this will block for 10 seconds*. \n\n*Issue 2* -- For _data repair_, {{RowDataResolver.resolve()}} has a similar issue as it calls {{scheduleRepairs()}} to send out messages using sendRR() without distinguishing remote nodes from the local node.\n\n*Issue 3* -- When handling _data repair_, {{StorageProxy.fetchRows()}} blocks waiting for acks to all of {{data repair}} requests sent out using sendRR(). This may cause the thread to block.\n\nFor _data repair_ path, *data requests* are sent out and then compare / merge the received responses; send out the merged / diff version and then block for acks.\n\nHow do we handle the case for _local node_? Does the sendRR() and the corresponding receive part can handle the case for local node? If not, then this may block for 10 seconds.\n\n{noformat}\n if (repairResponseHandlers != null)\n {\n for (int i = 0; i < repairCommands.size(); i++)\n {\n ReadCommand command = repairCommands.get(i);\n ReadCallback handler = repairResponseHandlers.get(i);\n\n Row row;\n try\n {\n row = handler.get();\n }\n catch (DigestMismatchException e)\n ...\n RowDataResolver resolver = (RowDataResolver)handler.resolver;\n try\n {\n // wait for the repair writes to be acknowledged, to minimize impact on any replica that's\n // behind on writes in case the out-of-sync row is read multiple times in quick succession\n FBUtilities.waitOnFutures(resolver.repairResults, DatabaseDescriptor.getWriteRpcTimeout());\n }\n catch (TimeoutException e)\n {\n Tracing.trace(\"Timed out on digest mismatch retries\");\n int blockFor = consistency_level.blockFor(Keyspace.open(command.getKeyspace()));\n throw new ReadTimeoutException(consistency_level, blockFor, blockFor, true);\n }\n{noformat}\n\n*Question for waiting for the ack* -- Do we really need to wait for the ack?\n\nWe should assume the best effort approach, i.e. do the data repair and then return. No need to block waiting for the acks for confirmation.\n\n*Question for the Randomized approach* -- Since the end points are randomized, the first node in the list is no likely the local node. This may cause a higher possibility of data repair.\n\nIn the *Randomized Approach*, the end points are reshuffled. Then, the first node in the list used for _data read request_ is not likely the local node. If this node happens to be the *DOWN* node, then, we end with all digest responses without the data, which will block and eventually timed out.\n\n","created":"2013-09-30T17:24:17.678+0000"},{"body":"First, let me thank you for your continued digging. Some of it helped. That said, you should probably look at the current cassandra-2.0 branch, and not the 2.0.0 tarball/branches here in the comments.\n\nbq. Issue 1 – When handling DigestMismatchException in StorageProxy.fetchRows(), all data read requests are sent out using sendRR without distinguishing remote nodes from the local node.\n\nThis is not an issue, and it's not spec retry related. Using LRR for local read requests is merely an optimisation - there is nothing wrong with sendRR (not that it isn't worth optimising here - just noting that it's not an issue). This is also the answer to \"How do we handle the case for local node? Does the sendRR() and the corresponding receive part can handle the case for local node? If not, then this may block for 10 seconds.\" Same goes for Issue 2 and Issue 3.\n\nbq. The data read request for local node may never sent out. As one of the nodes is down (which triggered the Speculative Retry) will cause one missing response.\n\nThe former is not true, the latter won't, since the current cassandra-2.0 code will send requests to all the contacted replicas. So if a node triggered spec retry, that extra speculated replica will get the request as well, and we can still satisfy the CL.\n\n{noformat}\n for (InetAddress endpoint : exec.getContactedReplicas())\n {\n Tracing.trace(\"Enqueuing full data read to {}\", endpoint);\n MessagingService.instance().sendRR(message, endpoint, repairHandler);\n }\n{noformat}\n\n\nbq. Question for the Randomized approach – Since the end points are randomized, the first node in the list is no likely the local node. This may cause a higher possibility of data repair.\n\nI don't see how the possibility of data repair is correlated with the locality of a target node, but, it doesn't matter. The 'randomised approach' was an experiment, it wasn't committed as part of the fix. See the latest cassandra-2.0 branch code.\n\nbq. In the Randomized Approach, the end points are reshuffled. Then, the first node in the list used for data read request is not likely the local node. If this node happens to be the DOWN node, then, we end with all digest responses without the data, which will block and eventually timed out.\n\nSee the above reply.\n\nTLDR: None of these seem to be issues, but we could optimise RR to use LRR for local reads to get slightly better performance for local requests (and to be consistent with the regular reads code).","created":"2013-09-30T20:46:05.962+0000"},{"body":"Thanks for the clarification of the sendRR issue.\n\nSince the Randomized approach is not checked in, let us skip over it.\n\nFor the _data repair_, do we need to block waiting for the acks?","created":"2013-09-30T21:32:59.959+0000"},{"body":"bq. For the data repair, do we need to block waiting for the acks?\n\nThe reasons are listed in the comments, as you've seen:\n{noformat}\n// wait for the repair writes to be acknowledged, to minimize impact on any replica that's\n// behind on writes in case the out-of-sync row is read multiple times in quick succession\n{noformat}\n\nTo reach that goal - yes, it's necessary. Is that scenario worth optimizing for or should we reconsider? Dunno. We are only writing to the replicas that we got the result from, though, so a known down replica wouldn't affect it.\n\n","created":"2013-09-30T21:53:54.586+0000"},{"body":"[~lizou] see CASSANDRA-4792 (TLDR: yes)","created":"2013-09-30T22:00:29.974+0000"},{"body":"I had to double the test length to get a good compaction graph. I'm not sure why it took so long, it didn't take as long in the original test.\n\n!5932.6692c50412ef7d.compaction.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.6692c50412ef7d.compaction.2.json&metric=interval_op_rate&operation=stress-read&smoothing=4]\n\n([~iamaleksey] your read_repair 0 / 1 tests are in progress...)","created":"2013-10-01T15:15:49.746+0000"},{"body":"Node killed while read_repair_chance=0. I accidentally left the test run to be 60M rows, so I chopped off the uninteresting bit.\n\n!5932.6692c50412ef7d.rr0.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.6692c50412ef7d.node_killed.rr0.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n\n(rr=1 is next..)","created":"2013-10-01T20:15:28.053+0000"},{"body":"The throughput is a wash in the compaction scenario, but the 99.9% latency looks a lot better with the retries.\n\nAny theories on why the percentile settings are posting better latency numbers than ALWAYS though?","created":"2013-10-02T03:39:23.416+0000"},{"body":"This testing result is reasonable and what is expected.\n\nFor PERCENTILE / CUSTOM configuration, the larger the {{cfs.sampleLatencyNanos}} the smaller the throughput impact for normal operations before the outage. However, during the outage period, the situation is reversed, i.e. the smaller {{cfs.sampleLatencyNanos}}, the smaller the throughput impact will be, as it times out quicker and triggers the speculative retries.\n\nFor the ALWAYS configuration, as it always sends out one speculative in addition to the usual read requests, the throughput performance should be lower than those of PERCENTILE / CUSTOM for normal operations before the outage. Since it always sends out the speculative retries, the throughput impact during the outage period should be the smallest. The testing result indicates that this is true.\n","created":"2013-10-02T15:46:48.843+0000"},{"body":"With read_repair_chance = 1\n\n!5932.6692c50412ef7d.rr1.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.6692c50412ef7d.node_killed.rr1.json&metric=interval_op_rate&operation=stress-read&smoothing=1]","created":"2013-10-02T15:50:04.065+0000"},{"body":"Have done some testing using today's trunk. Have observed following issues.\n\n*Issue 1* -- The first method {{MessagingService.addCallback()}} (i.e. without the ConsistencyLevel argument) asserts.\n\nCommenting out the assert statement seems to work. But the Cassandra servers themselves will produce 10-second outage (i.e. zero transactions from the client point of view) periodically.\n\n*Issue 2* -- The Speculative Retry seems stop retrying during the outage window.\n\nDuring the outage window triggered either by killing one of Cassandra nodes or produced by Cassandra servers themselves, the JConsole shows that the JMX stats, SpeculativeRetry counter stops incrementing until the gossip figures out the outage issue.\n\nWhat is the reason for this? The Speculative Retry is meant to help during the outage period. This observed behavior is consistent with Cassandra 2.0.0-rc2.\n","created":"2013-10-04T20:34:46.481+0000"},{"body":"bq. MessagingService.addCallback\n\nAre you sure you have the latest code? The only invocations of addCallback in 2.0/trunk include the consistencylevel argument as of late last night.","created":"2013-10-04T20:47:53.386+0000"},{"body":"The trunk load I used for testing was pulled this noon. It has two addCallback() methods. One of them (i.e. without the ConsistencyLevel) asserts. \n\nI checked the MessagingService.java, there are two addCallback() methods.\n* The one without ConsistencyLevel is called by sendRR()\n* The one with ConsistencyLevel is called by sendMessageToNonLocalDC()\n","created":"2013-10-04T20:58:05.523+0000"},{"body":"As for yesterday's trunk load, there were two addCallback() methods. But the one with ConsistencyLevel was not called by anyone. The one without ConsistencyLevel asserts.","created":"2013-10-04T21:06:48.497+0000"},{"body":"Are you doing counter updates? That's the only use of sendRR for updates I see.\n\nCan you post the stack trace of the assertion error you're getting?","created":"2013-10-04T21:27:15.521+0000"},{"body":"(Pushed fix for mutateCounter in 3da10f469d6a328bad209d723a5997c932284344.)","created":"2013-10-04T21:37:12.295+0000"},{"body":"[~jbellis], this morning's trunk load has a slightly different symptom, and is even more serious than last Friday's load, as this time just commenting out the assert statement in the {{MessagingService.addCallback()}} will not help.\n\nI copy the {{/var/log/cassandra/system.log}} exception errors below.\n\n{noformat}\nERROR [Thrift:12] 2013-10-07 14:42:39,396 Caller+0 at org.apache.cassandra.service.CassandraDaemon$2.uncaughtException(CassandraDaemon.java:134)\n - Exception in thread Thread[Thrift:12,5,main]\njava.lang.AssertionError: null\n at org.apache.cassandra.net.MessagingService.addCallback(MessagingService.java:543) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.net.MessagingService.sendRR(MessagingService.java:591) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.net.MessagingService.sendRR(MessagingService.java:571) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.service.StorageProxy.sendToHintedEndpoints(StorageProxy.java:869) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.service.StorageProxy$2.apply(StorageProxy.java:123) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.service.StorageProxy.performWrite(StorageProxy.java:739) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.service.StorageProxy.mutate(StorageProxy.java:511) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.service.StorageProxy.mutateWithTriggers(StorageProxy.java:581) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.cql3.statements.ModificationStatement.executeWithoutCondition(ModificationStatement.java:379) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.cql3.statements.ModificationStatement.execute(ModificationStatement.java:363) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.cql3.QueryProcessor.processStatement(QueryProcessor.java:126) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.cql3.QueryProcessor.processPrepared(QueryProcessor.java:267) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.thrift.CassandraServer.execute_prepared_cql3_query(CassandraServer.java:2061) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.thrift.Cassandra$Processor$execute_prepared_cql3_query.getResult(Cassandra.java:4502) ~[apache-cassandra-thrift-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.thrift.Cassandra$Processor$execute_prepared_cql3_query.getResult(Cassandra.java:4486) ~[apache-cassandra-thrift-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39) ~[libthrift-0.9.1.jar:0.9.1]\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39) ~[libthrift-0.9.1.jar:0.9.1]\n at org.apache.cassandra.thrift.CustomTThreadPoolServer$WorkerProcess.run(CustomTThreadPoolServer.java:194) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_25]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) ~[na:1.7.0_25]\n at java.lang.Thread.run(Thread.java:724) ~[na:1.7.0_25]\n\n{noformat}\n\n","created":"2013-10-07T18:56:51.265+0000"},{"body":"There's a ticket open for trunk over at CASSANDRA-6154.","created":"2013-10-07T19:35:58.090+0000"},{"body":"Hello,\n\nAs this ticket is already fixed in 2.0.2, where can I get the 2.0.2 source code?\n\nCurrently, my \"git tag\" only shows up to 2.0.1.\n","created":"2013-10-07T21:40:18.927+0000"},{"body":"the cassandra-2.0 branch is what will become 2.0.2","created":"2013-10-07T21:48:38.220+0000"},{"body":"I even cannot see the cassandra-2.0 branch.\nMy \"git tag\" gives a list including following branches.\n\n{noformat}\n$ git tag\n1.2.8\n1.2.8-tentative\ncassandra-0.3.0-final\ncassandra-0.3.0-rc1\ncassandra-0.3.0-rc2\n...\ncassandra-1.2.4\ncassandra-1.2.5\ncassandra-1.2.6\ncassandra-1.2.7\ncassandra-1.2.8\ncassandra-1.2.9\ncassandra-2.0.0\ncassandra-2.0.0-beta1\ncassandra-2.0.0-beta2\ncassandra-2.0.0-rc1\ncassandra-2.0.0-rc2\ncassandra-2.0.1\ndrivers\nlist\n{noformat}\n\nThere is no cassandra-2.0 branch. Where can I find it?\n","created":"2013-10-07T22:34:31.283+0000"},{"body":"Under {{git branch}}.","created":"2013-10-07T23:01:59.736+0000"}],"conversations":[{"body":"I've done a series of stress tests with eager retries enabled that show undesirable behavior. I'm grouping these behaviours into one ticket as they are most likely related.\n\n1) Killing off a node in a 4 node cluster actually increases performance.\n2) Compactions make nodes slow, even after the compaction is done.\n3) Eager Reads tend to lessen the *immediate* performance impact of a node going down, but not consistently.\n\nMy Environment:\n1 stress machine: node0\n4 C* nodes: node4, node5, node6, node7\n\nMy script:\nnode0 writes some data: stress -d node4 -F 30000000 -n 30000000 -i 5 -l 2 -K 20\nnode0 reads some data: stress -d node4 -n 30000000 -o read -i 5 -K 20\n\nh3. Examples:\n\nh5. A node going down increases performance:\n\n!node-down-increase-performance.png!\n\n[Data for this test here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.eager_retry.node_killed.just_20.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n\nAt 450s, I kill -9 one of the nodes. There is a brief decrease in performance as the snitch adapts, but then it recovers... to even higher performance than before.\n\nh5. Compactions make nodes permanently slow:\n\n!compaction-makes-slow.png!\n!compaction-makes-slow-stats.png!\n\nThe green and orange lines represent trials with eager retry enabled, they never recover their op-rate from before the compaction as the red and blue lines do.\n\n[Data for this test here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.eager_retry.compaction.2.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n\nh5. Speculative Read tends to lessen the *immediate* impact:\n\n!eager-read-looks-promising.png!\n!eager-read-looks-promising-stats.png!\n\nThis graph looked the most promising to me, the two trials with eager retry, the green and orange line, at 450s showed the smallest dip in performance. \n\n[Data for this test here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.eager_retry.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n\nh5. But not always:\n\n!eager-read-not-consistent.png!\n!eager-read-not-consistent-stats.png!\n\nThis is a retrial with the same settings as above, yet the 95percentile eager retry (red line) did poorly this time at 450s.\n\n[Data for this test here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.eager_retry.node_killed.just_20.rc1.try2.json&metric=interval_op_rate&operation=stress-read&smoothing=1]","from":"reporter","subject":"Speculative read performance data show unexpected results"},{"body":"[~enigmacurry], what was your collection interval for the metrics you've collected?\n\nI ask because I'm observing results similar to yours for \"3) Eager Reads tend to lessen the immediate performance impact of a node going down, but not consistently.\". However, I've polled metrics @ a 1 second granularity to see that it's actually a multi-second stress-client outage - not just poor and inconsistent performance.\n\nPolling metrics @ a 1 second interval, has observations that a ~20 second read operations starvation outage occurs for the stress client for all data in the cluster (even with the lowest phi_convict_threshold=6).\n\nAnalysis so far indicates that high-operations reads starve out all the C* client threads/connections, because they get stuck on awaiting for a server response whenever the key-space hits the node that is down (and by probability + high-operation reads, within 1 second each stress client thread will all hit the downed-node's key-space).\n\n\n\nSo I'm confirming that I'm also seeing this bug that Speculative reads (even with an ALWAYS setting). It isn't solving this outage for clients during high-operation reads, and based on what I understand of the feature, it should.\n\nThanks, guys!\n\n","from":"developer"},{"body":"I'm using 5 second intervals in these charts. 'multi-second stress-client outage' is a good way to put it, for both the case of speculative retry and not, the drop in performance after a node goes down is a duration of complete non-responsiveness (not degraded performance.) The addition of speculative retry consistently shortens this duration (it's always better), but this duration itself is inconsistent. ","from":"developer"},{"body":"SpeculateAlwaysExecutor - Here we are not reading from more endpoints than normal. We are only reading data from two endpoints. We should be reading from one more endpoint if possible. ","from":"developer"},{"body":"Thanks, [~enigmacurry]. Sounds like you are seeing the same thing as us, so it's great to see it's getting attention!","from":"developer"},{"body":"Hello [~iamaleksey],\n\nThanks for the link to this jira and for your very detailed testing results. It confirms what we have seen in our lab testing for the Cassandra 2.0.0-rc2 \"Speculative Execution for Reads\".\n\nWe have a very simple data center setup consisting of four Cassandra nodes running on four server machines. A testing application (Cassandra client) is interacting with Cassandra nodes 1, 2 and 3. That is, the testing app does not directly connected to the Cassandra node 4.\n\nThe keyspace Replication Factor is set to 3 and the client requested Consistency Level is set to CL_TWO.\n\nI have tested all of three configurations of the Speculative Execution for Reads ('ALWAYS', '85 PERCENTILE', '50 MS' / '100 MS'). It seems that none of them works as expected. From the test app log file point of view, they all give a 20-second window of outage immediately after the 4th node was killed. This behavior is consistent to Cassandra 1.2.4.\n\nI have done a quick code reading of the Cassandra Server implementation (Cassandra 2.0.0 tarball) and I have noticed some design issues. I would like to discuss them with you.\n\n*Issue 1* - StorageProxy.fetchRows() may still block for as long as conf.read_request_timeout_in_ms, though the speculative retry did fire correctly after the Cassandra node 4 was killed.\n\nTake the speculative configuration of 'PERCENTILE' / 'CUSTOM' as example, after the Cassandra node 4 was killed, SpeculativeReadExecutor.speculate() would block for responses. If timed out, it would send out one more read request to an alternative node (from {{unfiltered}}) and increment the speculativeRetry counter. This part should work.\n\nHowever, killing the 4th node would very likely cause inconsistency in the database and this will trigger the DigestMismatchException. In the fetchRows(), when handling DigestMismatchException, it uses handler.endpoints to send out digest mismatch retries and then block for responses. As we know that one of the endpoints was already killed, the handler.get() will block until it is timed out, which is 10 seconds.\n\n\n{noformat}\n catch (DigestMismatchException ex)\n {\n Tracing.trace(\"Digest mismatch: {}\", ex);\n\n ...\n\n MessageOut message = exec.command.createMessage();\n for (InetAddress endpoint : exec.handler.endpoints)\n {\n Tracing.trace(\"Enqueuing full data read to {}\", endpoint);\n MessagingService.instance().sendRR(message, endpoint, repairHandler);\n }\n }\n }\n\n ...\n\n // read the results for the digest mismatch retries\n if (repairResponseHandlers != null)\n {\n for (int i = 0; i < repairCommands.size(); i++)\n {\n ReadCommand command = repairCommands.get(i);\n ReadCallback handler = repairResponseHandlers.get(i);\n\n Row row;\n try\n {\n row = handler.get();\n }\n{noformat}\n\n\n\n*Issue 2* - The speculative 'ALWAYS' does NOT send out any more read requests. Thus, in face of the failure of node 4, it will not help at all.\n\nThe SpeculateAlwaysExecutor.executeAsync() only sends out handler.endpoints.size() number of read requests and it blocks for the responses to come back. If one of the nodes is killed, say node 4, this speculative retry 'ALWAYS' will work the same way as Cassandra 1.2.4, i.e. it will block until timed out, which is 10 seconds.\n\n??My understanding of this speculative retry 'ALWAYS' should ALWAYS send out \"handler.endpoints.size() + 1\" number of read requests and block for handler.endpoints.size() number of responses??.\n\n*Issue 3* - Since the ReadRepairDecison is determined by a Random() number, this speculative retry may not work as the ReadRepairDecision may be ??ReadRepairDecision.GLOBAL??\n\n*Issue 4* - For the ReadExecutor(s), the {{this.unfiltered}} and {{this.endpoints}} may not consistent. Thus, using {{this.unfiltered}} and {{this.endpoints}} for speculative retry may cause unexpected results. This is especially true when the Consistency Level is {{LOCAL_QUARUM}} and the ReadRepairDecision is {{DC_LOCAL}}.\n\n\n\n","from":"developer"},{"body":"Hey [~lizou]. Yeah, I've fixed most of these already (rewritten most of the ARE code, actually). Specifically issues 2,3,4. Will look into 1 too.\n\nThanks.","from":"developer"},{"body":"Attaching 5932.txt that will hopefully fix this ([~enigmacurry] could you run the tests again, please, with the patch applied?)\n\n1. As noted by [~lizou] and [~kohlisankalp], ALWAYS wasn't making an extra request, it was making an extra data request at the expense of one digest request. Fixed.\n\n2. SpecRetry wasn't working correctly with RRD.DC_LOCAL, as noted by @lizou, because the two lists will be in different order, and a retry might be sent to a node that already had a request sent to it. (Please note that LOCAL_QUORUM here does not affect anything - CL.filterForQuery() sorts in place, so the two lists would be in the same order, everything was working correct). RRD.DC_LOCAL handling was a legit issue though. Fixed.\n\n3. SpecRetry w/ RRD.GLOBAL is a noop, you can't speculate if you contact all the replicas in the first place. This is normal.\n\n4. The DME issue is semi-legit. Killing a node shouldn't trigger DME or increase the likelihood of DME happening. *HOWEVER* when shooting requests for repair, we were not considering the case where one of the replies satisfying the original CL came from a SpecRetry attempt. The patch includes the extra replica in repair commands if SpecRetry had been triggered by the original request.\n\n\n5. SP.getRangeSlice() is not SpectRetry-aware as of now. I don't know if this is an omission or by design, but for now, please don't include that in the benchmarks, since it would only be misleading. ","from":"developer"},{"body":"Pushed some OCD of my own to https://github.com/jbellis/cassandra/commits/5932 on top of this.\n\nbq. ALWAYS wasn't making an extra request, it was making an extra data request at the expense of one digest request\n\nI'm not sure what the distinction is here. Do you mean that if we weren't read-repairing, there would be no extra data request at all?\n\nbq. SpecRetry w/ RRD.GLOBAL is a noop, you can't speculate if you contact all the replicas in the first place.\n\nI dunno, I think we should turn a digest into a data for redundancy the way ALWAYS used to.","from":"developer"},{"body":"Pushed even more OCD to https://github.com/iamaleksey/cassandra/commits/5932 on top of yours.\n\nbq. I'm not sure what the distinction is here. Do you mean that if we weren't read-repairing, there would be no extra data request at all?\n\nInstead of making, for example 1 data request + 2 digest requests, ALWAYS was making 2 data requests + 1 digest request, instead of making 2 data requests + 2 digest requests, not really helping to satisfy the CL in case of node's failure.\n\nbq. I dunno, I think we should turn a digest into a data for redundancy the way ALWAYS used to.\n\nMaybe.","from":"developer"},{"body":"Force-pushed the 'final' version to https://github.com/iamaleksey/cassandra/commits/5932.\n\nAmong other things, properly handles RRD.GLOBAL and RRD.DC_LOCAL in 1-DC scenario.","from":"developer"},{"body":"Pushed one more set of changes to mine, not forced: https://github.com/jbellis/cassandra/commits/5932. Goal is to make SRE less fragile when doing RR.","from":"developer"},{"body":"+1, I'm out of OCD juice.","from":"developer"},{"body":"Hello [~iamaleksey] and [~jbellis],\n\nI took a quick look at the code changes. The new code looks very good to me. But I saw one potential issue in {{AlwaysSpeculatingReadExecutor.executeAsync()}}, in which it makes at least *two* data / digest requests. This will cause problems for a data center with only one Cassandra server node (e.g. bring up an embedded Cassandra node in JVM for JUnit test) or a deployed production data center of two Cassandra server nodes with one node shut down for maintenance. In the above mentioned two cases, {{AbstractReadExecutor.getReadExecutor()}} will return the {{AlwaysSpeculatingReadExecutor}} as condition {{(targetReplicas.size() == allReplicas.size())}} is met, though the tables may / may not be configured with ??Speculative ALWAYS??.\n\nIt is true for our legacy products we are considering to deploy each data center with only two Cassandra server nodes with RF = 2 and CL = 1.\n","from":"developer"},{"body":"The logic looks like this:\n\n# Figure out how many replicas we need to contact to satisfy the desired consistencyLevel + Read Repair settings\n# If that ends up being all the replicas, then use ASRE to get some redundancy on the data reads. This will allow the read to succeed even if a digest for RR times out. Of course if you are reading at CL.ALL and a replica times out there's nothing we can do.\n# Otherwise, use SRE and make an \"extra\" request later, if it looks like one of the minimal set isn't going to respond in time\n\nNote that performing extra data requests does not affect handler.blockfor -- just makes it possible for the request to proceed if it gets enough responses back, no matter which replicas they come from.","from":"developer"},{"body":"(Committed after Aleksey's +1, incidentally.)","from":"developer"},{"body":"The logic for {{AlwaysSpeculatingReadExecutor}} is good. What I meant in my previous comment is that when {{targetReplicas.size() == allReplicas.size()}} and {{targetReplicas.size() == 1}}, then {{AlwaysSpeculatingReadExecutor.executeAsync()}} will throw an exception as there is only one endpoint in {{targetReplicas}}, but it tries to access two endpoints in {{targetReplicas}}.","from":"developer"},{"body":"I see what you mean. Fixed in 7a87fc1186f39678382cf9b3e1dd224d9c71aead.","from":"developer"},{"body":"The good news is that speculative read has improved across the board.\n\nHowever, this new batch of testing introduces some new mysteries.\n\nHere is all of the runs from 7a87fc1186f39678382cf9b3e1dd224d9c71aead:\n\n!5933-7a87fc11.png!\n\nAll of the speculative retry runs are better than with 2.0.0-rc1. However, I can't explain why sr=NONE did better than ALWAYS and 95percentile. There is no visible indication that a node went down for sr=NONE. I have double checked the logs, and it did, in fact, go down. \n\nCompare this to the baseline of 1.2.8 and 2.0.0-rc1 (redone last night on same hardware as above):\n\n!5933-128_and_200rc1.png!\n\nAll of these have clear indications of the node going down.\n\nYou can [see all the data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1] - you can double click the colored squares to toggle the visibility of the lines, as they do overlap.\n\nI've uploaded logs from all these runs as 5933-logs.tar.gz.","from":"developer"},{"body":"Yeah. The graphs for ALWAYS and NONE look swapped from what I would expect.","from":"developer"},{"body":"The other thing that I note, is that all of these runs are better than 1.2.8 :D (further evidence that CASSANDRA-5933 may be invalid)","from":"developer"},{"body":"Hmm.\n\nI wonder if it's just luck of the draw as to which replica dsnitch is preferring. Here's a branch to randomize that, per-operation:\n\nhttps://github.com/jbellis/cassandra/commits/5932-randomized","from":"developer"},{"body":"Couldn't we just do a run with the dsnitch disabled?","from":"developer"},{"body":"That still gives you luck-of-the-draw as to which replica it prefers. (Unlikely to be evenly distributed.)","from":"developer"},{"body":"This looks exactly like what I was expecting:\n\n!5933-randomized-dsnitch-replica.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.randomized-dsnitch.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]","from":"developer"},{"body":"That's awesome.\n\nCan you test 90th and 75th percentile too?","from":"developer"},{"body":"Seems like it's quite tunable, but not a lot of difference under 90%:\n\n!5933-randomized-dsnitch-replica.2.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.randomized-dsnitch.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]","from":"developer"},{"body":"Yeah, that makes sense. 70th..95th are all pretty damn close to median still.\n\nWhat I'd like to do is get close to the 10ms performance hit (~none) as a percentile, and make that default in 2.1. Try 99th and 99.9th?","from":"developer"},{"body":"Also, did that stress patch work to get you failed request counts? Would be good to get that too if we can show that even 10ms keeps requests from failing entirely.","from":"developer"},{"body":"Here's 99th and 99.9th percentiles:\n\n!5933-randomized-dsnitch-replica.3.png!\n\nI'll look at that stress patch again, I seem to recall it not making a lot of sense to me when I last tried it, but will give it another go.","from":"developer"},{"body":"Starting to think we still have a bug. 99.9 should be doing less retries than 10ms but the graph shows it doing more.","from":"developer"},{"body":"I did some more code reading and noticed some potential issues and possible improvement. I've got to run now. I will get back to you guys Monday morning.\n\nMy guess is that the ??Speculative NONE?? is hit by the initial request reading path which is successfully resolved by the *Speculative Retry*. The observed throughput performance hit when Speculative Retry is enabled is caused by the ReadRepair path which has some coding / design issues. I will talk to you next Monday.","from":"developer"},{"body":"Maybe the problem is that we're using CF-level latency instead of StorageProxy.\n\nWhat does cfhistograms give for read latency?","from":"developer"},{"body":"Pretty sure that's our smoking gun. Pushed a commit to the -randomized branch that adds coordinator-level, per-cf latency tracking and uses that instead.\n\nCan you repeat the last test with that? (Maybe throw in ALL as well if you're feeling optimistic that we'll have a measurable difference between ALL and 90%. :)","from":"developer"},{"body":"[~jbellis] Here's your two runs:\n\n[dea27f84f40|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.dea27f84f40.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n[ded39c7e1c2fa|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.ded39c7e1c2fa.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n\nLogs for the second run are attached as 5932.ded39c7e1c2fa.logs.tar.gz","from":"developer"},{"body":"The code to convert 99 into 0.99 was buggy and was actually converting to 0.0099. Fix pushed, can you try it again?","from":"developer"},{"body":"BINGO!\n\n!5932-6692c50412ef7d.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.6692c50412ef7d.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]","from":"developer"},{"body":"Definitely the best results seen so far! Nice work!\n\nIn my mind, during the transition period right after the killing of the node, I expected \"ALWAYS\" to have negligible impact, or at least the smallest impact of all of the other values. However, red (90%), and purple (75%) is having a smaller impact. Seems fishy. Do I misunderstand the intention of the \"ALWAYS\" setting?\n\n(Edited to clarify the period I'm talking about.)","from":"developer"},{"body":"It looks to me like 75/90/Always are about the same, with Always dropping from a lower baseline. Which makes sense; it's still doing a lot of unnecessary work compared to the others.","from":"developer"},{"body":"[~enigmacurry], can you also re-test the uncapped compaction scenario with the same set of retry settings?","from":"developer"},{"body":"Hello [~iamaleksey] and [~jbellis],\n\nIt appears to me that the testing results have suggested that the \"_data read + speculative retry_\" path work as expected. This \"_data read + speculative retry_\" path has greatly minimized the throughput impact caused by the failure of one of Cassandra server nodes.\n\nThe observed small degradation of throughput performance when _speculative retry_ is enabled is very likely to be caused by the \"*_read repair_*\" path. I did the code reading of this path last Friday and noticed some design / coding issues. I would like to discuss them with you.\n\nPlease note that my code base is still the Cassandra 2.0.0 tarball, not updated with the latest code changes.\n\n*Issue 1* -- When handling {{DigestMismatchException}} in {{StorageProxy.fetchRows()}}, all _data read requests_ are sent out using {{sendRR}} without distinguishing remote nodes from the local node.\n\nWill this cause an issue, as {{MessagingService.instance().sendRR()}} will send out enqueued messages for a specified remote node via its pre-established TCP socket connection. For local node, this should be done via {{LocalReadRunnable}}, i.e. {{StageManager.getStage(Stage.READ).execute(new LocalReadRunnable(command, handler))}}.\n\nIf this may cause an issue, the following wait may block.\n\n{noformat}\n // read the results for the digest mismatch retries\n if (repairResponseHandlers != null)\n {\n for (int i = 0; i < repairCommands.size(); i++)\n {\n ReadCommand command = repairCommands.get(i);\n ReadCallback handler = repairResponseHandlers.get(i);\n\n Row row;\n try\n {\n row = handler.get();\n }\n{noformat}\n\nFor two reasons.\n* The data read request for local node may never sent out\n* As one of the nodes is down (which triggered the Speculative Retry) will cause one missing response.\n\n*If missing two responses, this will block for 10 seconds*. \n\n*Issue 2* -- For _data repair_, {{RowDataResolver.resolve()}} has a similar issue as it calls {{scheduleRepairs()}} to send out messages using sendRR() without distinguishing remote nodes from the local node.\n\n*Issue 3* -- When handling _data repair_, {{StorageProxy.fetchRows()}} blocks waiting for acks to all of {{data repair}} requests sent out using sendRR(). This may cause the thread to block.\n\nFor _data repair_ path, *data requests* are sent out and then compare / merge the received responses; send out the merged / diff version and then block for acks.\n\nHow do we handle the case for _local node_? Does the sendRR() and the corresponding receive part can handle the case for local node? If not, then this may block for 10 seconds.\n\n{noformat}\n if (repairResponseHandlers != null)\n {\n for (int i = 0; i < repairCommands.size(); i++)\n {\n ReadCommand command = repairCommands.get(i);\n ReadCallback handler = repairResponseHandlers.get(i);\n\n Row row;\n try\n {\n row = handler.get();\n }\n catch (DigestMismatchException e)\n ...\n RowDataResolver resolver = (RowDataResolver)handler.resolver;\n try\n {\n // wait for the repair writes to be acknowledged, to minimize impact on any replica that's\n // behind on writes in case the out-of-sync row is read multiple times in quick succession\n FBUtilities.waitOnFutures(resolver.repairResults, DatabaseDescriptor.getWriteRpcTimeout());\n }\n catch (TimeoutException e)\n {\n Tracing.trace(\"Timed out on digest mismatch retries\");\n int blockFor = consistency_level.blockFor(Keyspace.open(command.getKeyspace()));\n throw new ReadTimeoutException(consistency_level, blockFor, blockFor, true);\n }\n{noformat}\n\n*Question for waiting for the ack* -- Do we really need to wait for the ack?\n\nWe should assume the best effort approach, i.e. do the data repair and then return. No need to block waiting for the acks for confirmation.\n\n*Question for the Randomized approach* -- Since the end points are randomized, the first node in the list is no likely the local node. This may cause a higher possibility of data repair.\n\nIn the *Randomized Approach*, the end points are reshuffled. Then, the first node in the list used for _data read request_ is not likely the local node. If this node happens to be the *DOWN* node, then, we end with all digest responses without the data, which will block and eventually timed out.\n\n","from":"developer"},{"body":"First, let me thank you for your continued digging. Some of it helped. That said, you should probably look at the current cassandra-2.0 branch, and not the 2.0.0 tarball/branches here in the comments.\n\nbq. Issue 1 – When handling DigestMismatchException in StorageProxy.fetchRows(), all data read requests are sent out using sendRR without distinguishing remote nodes from the local node.\n\nThis is not an issue, and it's not spec retry related. Using LRR for local read requests is merely an optimisation - there is nothing wrong with sendRR (not that it isn't worth optimising here - just noting that it's not an issue). This is also the answer to \"How do we handle the case for local node? Does the sendRR() and the corresponding receive part can handle the case for local node? If not, then this may block for 10 seconds.\" Same goes for Issue 2 and Issue 3.\n\nbq. The data read request for local node may never sent out. As one of the nodes is down (which triggered the Speculative Retry) will cause one missing response.\n\nThe former is not true, the latter won't, since the current cassandra-2.0 code will send requests to all the contacted replicas. So if a node triggered spec retry, that extra speculated replica will get the request as well, and we can still satisfy the CL.\n\n{noformat}\n for (InetAddress endpoint : exec.getContactedReplicas())\n {\n Tracing.trace(\"Enqueuing full data read to {}\", endpoint);\n MessagingService.instance().sendRR(message, endpoint, repairHandler);\n }\n{noformat}\n\n\nbq. Question for the Randomized approach – Since the end points are randomized, the first node in the list is no likely the local node. This may cause a higher possibility of data repair.\n\nI don't see how the possibility of data repair is correlated with the locality of a target node, but, it doesn't matter. The 'randomised approach' was an experiment, it wasn't committed as part of the fix. See the latest cassandra-2.0 branch code.\n\nbq. In the Randomized Approach, the end points are reshuffled. Then, the first node in the list used for data read request is not likely the local node. If this node happens to be the DOWN node, then, we end with all digest responses without the data, which will block and eventually timed out.\n\nSee the above reply.\n\nTLDR: None of these seem to be issues, but we could optimise RR to use LRR for local reads to get slightly better performance for local requests (and to be consistent with the regular reads code).","from":"developer"},{"body":"Thanks for the clarification of the sendRR issue.\n\nSince the Randomized approach is not checked in, let us skip over it.\n\nFor the _data repair_, do we need to block waiting for the acks?","from":"developer"},{"body":"bq. For the data repair, do we need to block waiting for the acks?\n\nThe reasons are listed in the comments, as you've seen:\n{noformat}\n// wait for the repair writes to be acknowledged, to minimize impact on any replica that's\n// behind on writes in case the out-of-sync row is read multiple times in quick succession\n{noformat}\n\nTo reach that goal - yes, it's necessary. Is that scenario worth optimizing for or should we reconsider? Dunno. We are only writing to the replicas that we got the result from, though, so a known down replica wouldn't affect it.\n\n","from":"developer"},{"body":"[~lizou] see CASSANDRA-4792 (TLDR: yes)","from":"developer"},{"body":"I had to double the test length to get a good compaction graph. I'm not sure why it took so long, it didn't take as long in the original test.\n\n!5932.6692c50412ef7d.compaction.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.6692c50412ef7d.compaction.2.json&metric=interval_op_rate&operation=stress-read&smoothing=4]\n\n([~iamaleksey] your read_repair 0 / 1 tests are in progress...)","from":"developer"},{"body":"Node killed while read_repair_chance=0. I accidentally left the test run to be 60M rows, so I chopped off the uninteresting bit.\n\n!5932.6692c50412ef7d.rr0.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.6692c50412ef7d.node_killed.rr0.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n\n(rr=1 is next..)","from":"developer"},{"body":"The throughput is a wash in the compaction scenario, but the 99.9% latency looks a lot better with the retries.\n\nAny theories on why the percentile settings are posting better latency numbers than ALWAYS though?","from":"developer"},{"body":"This testing result is reasonable and what is expected.\n\nFor PERCENTILE / CUSTOM configuration, the larger the {{cfs.sampleLatencyNanos}} the smaller the throughput impact for normal operations before the outage. However, during the outage period, the situation is reversed, i.e. the smaller {{cfs.sampleLatencyNanos}}, the smaller the throughput impact will be, as it times out quicker and triggers the speculative retries.\n\nFor the ALWAYS configuration, as it always sends out one speculative in addition to the usual read requests, the throughput performance should be lower than those of PERCENTILE / CUSTOM for normal operations before the outage. Since it always sends out the speculative retries, the throughput impact during the outage period should be the smallest. The testing result indicates that this is true.\n","from":"developer"},{"body":"With read_repair_chance = 1\n\n!5932.6692c50412ef7d.rr1.png!\n\n[data here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.5933.6692c50412ef7d.node_killed.rr1.json&metric=interval_op_rate&operation=stress-read&smoothing=1]","from":"developer"},{"body":"Have done some testing using today's trunk. Have observed following issues.\n\n*Issue 1* -- The first method {{MessagingService.addCallback()}} (i.e. without the ConsistencyLevel argument) asserts.\n\nCommenting out the assert statement seems to work. But the Cassandra servers themselves will produce 10-second outage (i.e. zero transactions from the client point of view) periodically.\n\n*Issue 2* -- The Speculative Retry seems stop retrying during the outage window.\n\nDuring the outage window triggered either by killing one of Cassandra nodes or produced by Cassandra servers themselves, the JConsole shows that the JMX stats, SpeculativeRetry counter stops incrementing until the gossip figures out the outage issue.\n\nWhat is the reason for this? The Speculative Retry is meant to help during the outage period. This observed behavior is consistent with Cassandra 2.0.0-rc2.\n","from":"developer"},{"body":"bq. MessagingService.addCallback\n\nAre you sure you have the latest code? The only invocations of addCallback in 2.0/trunk include the consistencylevel argument as of late last night.","from":"developer"},{"body":"The trunk load I used for testing was pulled this noon. It has two addCallback() methods. One of them (i.e. without the ConsistencyLevel) asserts. \n\nI checked the MessagingService.java, there are two addCallback() methods.\n* The one without ConsistencyLevel is called by sendRR()\n* The one with ConsistencyLevel is called by sendMessageToNonLocalDC()\n","from":"developer"},{"body":"As for yesterday's trunk load, there were two addCallback() methods. But the one with ConsistencyLevel was not called by anyone. The one without ConsistencyLevel asserts.","from":"developer"},{"body":"Are you doing counter updates? That's the only use of sendRR for updates I see.\n\nCan you post the stack trace of the assertion error you're getting?","from":"developer"},{"body":"(Pushed fix for mutateCounter in 3da10f469d6a328bad209d723a5997c932284344.)","from":"developer"},{"body":"[~jbellis], this morning's trunk load has a slightly different symptom, and is even more serious than last Friday's load, as this time just commenting out the assert statement in the {{MessagingService.addCallback()}} will not help.\n\nI copy the {{/var/log/cassandra/system.log}} exception errors below.\n\n{noformat}\nERROR [Thrift:12] 2013-10-07 14:42:39,396 Caller+0 at org.apache.cassandra.service.CassandraDaemon$2.uncaughtException(CassandraDaemon.java:134)\n - Exception in thread Thread[Thrift:12,5,main]\njava.lang.AssertionError: null\n at org.apache.cassandra.net.MessagingService.addCallback(MessagingService.java:543) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.net.MessagingService.sendRR(MessagingService.java:591) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.net.MessagingService.sendRR(MessagingService.java:571) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.service.StorageProxy.sendToHintedEndpoints(StorageProxy.java:869) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.service.StorageProxy$2.apply(StorageProxy.java:123) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.service.StorageProxy.performWrite(StorageProxy.java:739) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.service.StorageProxy.mutate(StorageProxy.java:511) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.service.StorageProxy.mutateWithTriggers(StorageProxy.java:581) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.cql3.statements.ModificationStatement.executeWithoutCondition(ModificationStatement.java:379) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.cql3.statements.ModificationStatement.execute(ModificationStatement.java:363) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.cql3.QueryProcessor.processStatement(QueryProcessor.java:126) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.cql3.QueryProcessor.processPrepared(QueryProcessor.java:267) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.thrift.CassandraServer.execute_prepared_cql3_query(CassandraServer.java:2061) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.thrift.Cassandra$Processor$execute_prepared_cql3_query.getResult(Cassandra.java:4502) ~[apache-cassandra-thrift-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.cassandra.thrift.Cassandra$Processor$execute_prepared_cql3_query.getResult(Cassandra.java:4486) ~[apache-cassandra-thrift-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39) ~[libthrift-0.9.1.jar:0.9.1]\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39) ~[libthrift-0.9.1.jar:0.9.1]\n at org.apache.cassandra.thrift.CustomTThreadPoolServer$WorkerProcess.run(CustomTThreadPoolServer.java:194) ~[apache-cassandra-2.1-SNAPSHOT.jar:2.1-SNAPSHOT]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_25]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) ~[na:1.7.0_25]\n at java.lang.Thread.run(Thread.java:724) ~[na:1.7.0_25]\n\n{noformat}\n\n","from":"developer"},{"body":"There's a ticket open for trunk over at CASSANDRA-6154.","from":"developer"},{"body":"Hello,\n\nAs this ticket is already fixed in 2.0.2, where can I get the 2.0.2 source code?\n\nCurrently, my \"git tag\" only shows up to 2.0.1.\n","from":"developer"},{"body":"the cassandra-2.0 branch is what will become 2.0.2","from":"developer"},{"body":"I even cannot see the cassandra-2.0 branch.\nMy \"git tag\" gives a list including following branches.\n\n{noformat}\n$ git tag\n1.2.8\n1.2.8-tentative\ncassandra-0.3.0-final\ncassandra-0.3.0-rc1\ncassandra-0.3.0-rc2\n...\ncassandra-1.2.4\ncassandra-1.2.5\ncassandra-1.2.6\ncassandra-1.2.7\ncassandra-1.2.8\ncassandra-1.2.9\ncassandra-2.0.0\ncassandra-2.0.0-beta1\ncassandra-2.0.0-beta2\ncassandra-2.0.0-rc1\ncassandra-2.0.0-rc2\ncassandra-2.0.1\ndrivers\nlist\n{noformat}\n\nThere is no cassandra-2.0 branch. Where can I find it?\n","from":"developer"},{"body":"Under {{git branch}}.","from":"developer"}],"created":"2013-08-24T00:51:07.000+0000","description":"I've done a series of stress tests with eager retries enabled that show undesirable behavior. I'm grouping these behaviours into one ticket as they are most likely related.\n\n1) Killing off a node in a 4 node cluster actually increases performance.\n2) Compactions make nodes slow, even after the compaction is done.\n3) Eager Reads tend to lessen the *immediate* performance impact of a node going down, but not consistently.\n\nMy Environment:\n1 stress machine: node0\n4 C* nodes: node4, node5, node6, node7\n\nMy script:\nnode0 writes some data: stress -d node4 -F 30000000 -n 30000000 -i 5 -l 2 -K 20\nnode0 reads some data: stress -d node4 -n 30000000 -o read -i 5 -K 20\n\nh3. Examples:\n\nh5. A node going down increases performance:\n\n!node-down-increase-performance.png!\n\n[Data for this test here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.eager_retry.node_killed.just_20.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n\nAt 450s, I kill -9 one of the nodes. There is a brief decrease in performance as the snitch adapts, but then it recovers... to even higher performance than before.\n\nh5. Compactions make nodes permanently slow:\n\n!compaction-makes-slow.png!\n!compaction-makes-slow-stats.png!\n\nThe green and orange lines represent trials with eager retry enabled, they never recover their op-rate from before the compaction as the red and blue lines do.\n\n[Data for this test here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.eager_retry.compaction.2.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n\nh5. Speculative Read tends to lessen the *immediate* impact:\n\n!eager-read-looks-promising.png!\n!eager-read-looks-promising-stats.png!\n\nThis graph looked the most promising to me, the two trials with eager retry, the green and orange line, at 450s showed the smallest dip in performance. \n\n[Data for this test here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.eager_retry.node_killed.json&metric=interval_op_rate&operation=stress-read&smoothing=1]\n\nh5. But not always:\n\n!eager-read-not-consistent.png!\n!eager-read-not-consistent-stats.png!\n\nThis is a retrial with the same settings as above, yet the 95percentile eager retry (red line) did poorly this time at 450s.\n\n[Data for this test here|http://ryanmcguire.info/ds/graph/graph.html?stats=stats.eager_retry.node_killed.just_20.rc1.try2.json&metric=interval_op_rate&operation=stress-read&smoothing=1]","issue_id":"12665353","key":"CASSANDRA-5932","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-09-26T20:55:41.000+0000","role":"fixed_distractor","summary":"Speculative read performance data show unexpected results"} {"case_id":"12669215","cluster":"DISTRACTOR-CASSANDRA-6053","comments":[{"body":"Output of nodetool status and the contents of the system.peers tables.","created":"2013-09-18T09:07:16.965+0000"},{"body":"Just a heads up: I've just removed a broken node using 'nodetool removenode' and encountered the same problem again. The node wasn't removed from the other nodes' system.peers.","created":"2013-09-18T14:35:37.695+0000"},{"body":"We're seeing this on a 1.2.9 cluster as well.","created":"2013-09-26T21:28:41.175+0000"},{"body":"Is there a way to manually fix the system.peers table once it's in the messed up state? Like a nodetool command or simply using CQL to delete the bad rows?","created":"2013-09-27T18:56:30.582+0000"},{"body":"You can simply delete them all from system.peers if you want, they'll re-populate correctly.","created":"2013-09-27T19:11:58.018+0000"},{"body":"Not sure how this can happen, given this code:\n\n{noformat}\n private void removeEndpoint(InetAddress endpoint)\n {\n Gossiper.instance.removeEndpoint(endpoint);\n if (!isClientMode)\n SystemTable.removeEndpoint(endpoint);\n }\n{noformat}\n\nwhich is what decom and removetoken end up calling.","created":"2013-09-27T19:15:05.225+0000"},{"body":"I did a \"truncate peers\" and it hasn't repopulated (I waited about 10 mins). Do I need to restart one or more nodes to trigger repopulating of the table?","created":"2013-09-27T21:16:32.186+0000"},{"body":"Yes.","created":"2013-09-27T23:32:27.220+0000"},{"body":"The load_ring_state=false directive should probably also clear out the peers table because otherwise, the state that you're trying to get rid of is still persisted there.","created":"2013-10-03T11:12:44.593+0000"},{"body":"If a user knows enough to disable loading the state and that fixes the problem, they can clear the peers table manually.","created":"2013-10-07T16:56:19.451+0000"},{"body":"[~enigmacurry] can you reproduce?","created":"2013-12-18T17:02:35.874+0000"},{"body":"First attempt appears to work correctly on cassandra-2.0 HEAD and 1.2.9 : \n\n{code}\n12:53 PM:~$ ccm create -v git:cassandra-1.2.9 t\nFetching Cassandra updates...\nCurrent cluster is now: t\n12:53 PM:~$ ccm populate -n 5\n12:54 PM:~$ ccm start\n12:54 PM:~$ ccm node1 stress\nCreated keyspaces. Sleeping 1s for propagation.\ntotal,interval_op_rate,interval_key_rate,latency/95th/99th,elapsed_time\n24994,2499,2499,9.5,55.2,179.0,10\n103123,7812,7812,2.8,27.2,134.7,20\n236358,13323,13323,1.7,15.4,134.7,30\n329477,9311,9311,1.7,9.8,109.8,40\n405667,7619,7619,1.8,9.2,6591.9,50\n558989,15332,15332,1.5,6.6,6591.1,60\n^C12:55 PM:~$ ccm node1 cqlsh\nConnected to t at 127.0.0.1:9160.\n[cqlsh 3.1.7 | Cassandra 1.2.9-SNAPSHOT | CQL spec 3.0.0 | Thrift protocol 19.36.0]\nUse HELP for help.\ncqlsh> select peer from system.peers;\n\n peer\n-----------\n 127.0.0.3\n 127.0.0.2\n 127.0.0.5\n 127.0.0.4\n\ncqlsh>\n12:55 PM:~$ ccm node2 decommission\n12:57 PM:~$ ccm node1 cqlsh\nConnected to t at 127.0.0.1:9160.\n[cqlsh 3.1.7 | Cassandra 1.2.9-SNAPSHOT | CQL spec 3.0.0 | Thrift protocol 19.36.0]\nUse HELP for help.\ncqlsh> select peer from system.peers;\n\n peer\n-----------\n 127.0.0.3\n 127.0.0.5\n 127.0.0.4\n\ncqlsh>\n12:58 PM:~$\n{code}\n\nAll nodes show equivalent peers table.","created":"2013-12-18T18:03:07.391+0000"},{"body":"OK, reproduced this by killing -9 one of the nodes and then doing a 'nodetool removenode':\n\n{code}\n01:20 PM:~$ kill -9 18961 (PID of node1)\n01:21 PM:~$ ccm node1 status\nFailed to connect to '127.0.0.1:7100': Connection refused\n01:21 PM:~$ ccm node2 status\nDatacenter: datacenter1\n=======================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Owns Host ID Token Rack\nDN 127.0.0.1 62.93 KB 20.0% 896644af-8640-4be6-a3ff-e8ed559d851c -9223372036854775808 rack1\nUN 127.0.0.2 51.17 KB 20.0% d3801466-d36d-428c-b4e5-05ff69fe36c0 -5534023222112865485 rack1\nUN 127.0.0.3 62.78 KB 20.0% cb36c3ad-df45-4f77-bff5-ca93c504ec08 -1844674407370955162 rack1\nUN 127.0.0.4 51.17 KB 20.0% 89031a05-a3f6-4ac7-9d29-6caa0c609dbc 1844674407370955161 rack1\nUN 127.0.0.5 51.27 KB 20.0% 4909d856-a86e-493a-a7d0-7570d71eb9d8 5534023222112865484 rack1\n\n# Issue removenode on node3 :\n01:21 PM:~$ ~/.ccm/t/node1/bin/nodetool -p 7300 removenode 896644af-8640-4be6-a3ff-e8ed559d851c\n\n01:22 PM:~$ ccm node3 cqlsh\nConnected to t at 127.0.0.3:9160.\n[cqlsh 4.1.0 | Cassandra 2.0.3-SNAPSHOT | CQL spec 3.1.1 | Thrift protocol 19.39.0]\nUse HELP for help.\ncqlsh> select * from system.peers;\n\n peer | data_center | host_id | preferred_ip | rack | release_version | rpc_address | schema_version | tokens\n-----------+-------------+--------------------------------------+--------------+-------+-----------------+-------------+--------------------------------------+--------------------------\n 127.0.0.2 | datacenter1 | d3801466-d36d-428c-b4e5-05ff69fe36c0 | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.2 | d133398f-f287-3674-83af-a1b04ee29f1f | {'-5534023222112865485'}\n 127.0.0.5 | datacenter1 | 4909d856-a86e-493a-a7d0-7570d71eb9d8 | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.5 | d133398f-f287-3674-83af-a1b04ee29f1f | {'5534023222112865484'}\n 127.0.0.4 | datacenter1 | 89031a05-a3f6-4ac7-9d29-6caa0c609dbc | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.4 | d133398f-f287-3674-83af-a1b04ee29f1f | {'1844674407370955161'}\n\n(3 rows)\n\n# Check node2 peers table:\n\n01:23 PM:~$ ccm node2 cqlsh\nConnected to t at 127.0.0.2:9160.\n[cqlsh 4.1.0 | Cassandra 2.0.3-SNAPSHOT | CQL spec 3.1.1 | Thrift protocol 19.39.0]\nUse HELP for help.\ncqlsh> select * from system.peers;\n\n peer | data_center | host_id | preferred_ip | rack | release_version | rpc_address | schema_version | tokens\n-----------+-------------+--------------------------------------+--------------+-------+-----------------+-------------+--------------------------------------+--------------------------\n 127.0.0.3 | datacenter1 | cb36c3ad-df45-4f77-bff5-ca93c504ec08 | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.3 | d133398f-f287-3674-83af-a1b04ee29f1f | {'-1844674407370955162'}\n 127.0.0.1 | null | 896644af-8640-4be6-a3ff-e8ed559d851c | null | null | null | 127.0.0.1 | null | null\n 127.0.0.5 | datacenter1 | 4909d856-a86e-493a-a7d0-7570d71eb9d8 | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.5 | d133398f-f287-3674-83af-a1b04ee29f1f | {'5534023222112865484'}\n 127.0.0.4 | datacenter1 | 89031a05-a3f6-4ac7-9d29-6caa0c609dbc | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.4 | d133398f-f287-3674-83af-a1b04ee29f1f | {'1844674407370955161'}\n\n(4 rows)\n\n# oh noes!... node2 still has an entry for node1 in peers table.\n\n01:23 PM:~$ ccm node2 status\nDatacenter: datacenter1\n=======================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Owns Host ID Token Rack\nUN 127.0.0.2 51.17 KB 40.0% d3801466-d36d-428c-b4e5-05ff69fe36c0 -5534023222112865485 rack1\nUN 127.0.0.3 62.78 KB 20.0% cb36c3ad-df45-4f77-bff5-ca93c504ec08 -1844674407370955162 rack1\nUN 127.0.0.4 51.17 KB 20.0% 89031a05-a3f6-4ac7-9d29-6caa0c609dbc 1844674407370955161 rack1\nUN 127.0.0.5 51.27 KB 20.0% 4909d856-a86e-493a-a7d0-7570d71eb9d8 5534023222112865484 rack1\n\n{code}\n\nBy issuing the removenode on node3, node3 seems to know about the node being removed and it's peers table is correct. node2, although it's status output shows node1 going away, it's peers table has not been updated.","created":"2013-12-18T18:29:21.495+0000"},{"body":"Thanks, Ryan.","created":"2013-12-18T21:57:27.607+0000"},{"body":"Can you provide debug logs from node2?","created":"2013-12-30T13:54:39.832+0000"},{"body":"I was able to repro with Ryan's steps, so I can upload the logs if you'd like, but I should be able to figure this one out.","created":"2013-12-31T18:39:28.850+0000"},{"body":"The problem was that state changes for the removed node were being handled after the system.peers row was deleted. These state changes would result in update to the system.peers row, partially reviving it.\n\n6053-v1.patch (and [branch|https://github.com/thobbs/cassandra/tree/CASSANDRA-6053]) avoids updating the system.peers table if the node is unknown or is in one of the \"dead\" states and adds some basic unit test coverage.","created":"2013-12-31T22:49:05.566+0000"},{"body":"Committed.","created":"2014-01-10T18:52:04.507+0000"},{"body":"[~brandon.williams] FYI - Even though, this was fixed it surfaced in 2.0.14 release that we have in production. We are going to follow the work-around as mentioned above.","created":"2015-11-03T14:30:09.327+0000"}],"conversations":[{"body":"After decommissioning my cluster from 20 to 9 nodes using opscenter, I found all but one of the nodes had incorrect system.peers tables.\n\nThis became a problem (afaik) when using the python-driver, since this queries the peers table to set up its connection pool. Resulting in very slow startup times, because of timeouts.\n\nThe output of nodetool didn't seem to be affected. After removing the incorrect entries from the peers tables, the connection issues seem to have disappeared for us. \n\nWould like some feedback on if this was the right way to handle the issue or if I'm still left with a broken cluster.\n\nAttached is the output of nodetool status, which shows the correct 9 nodes. Below that the output of the system.peers tables on the individual nodes.\n","from":"reporter","subject":"system.peers table not updated after decommissioning nodes in C* 2.0"},{"body":"Output of nodetool status and the contents of the system.peers tables.","from":"developer"},{"body":"Just a heads up: I've just removed a broken node using 'nodetool removenode' and encountered the same problem again. The node wasn't removed from the other nodes' system.peers.","from":"developer"},{"body":"We're seeing this on a 1.2.9 cluster as well.","from":"developer"},{"body":"Is there a way to manually fix the system.peers table once it's in the messed up state? Like a nodetool command or simply using CQL to delete the bad rows?","from":"developer"},{"body":"You can simply delete them all from system.peers if you want, they'll re-populate correctly.","from":"developer"},{"body":"Not sure how this can happen, given this code:\n\n{noformat}\n private void removeEndpoint(InetAddress endpoint)\n {\n Gossiper.instance.removeEndpoint(endpoint);\n if (!isClientMode)\n SystemTable.removeEndpoint(endpoint);\n }\n{noformat}\n\nwhich is what decom and removetoken end up calling.","from":"developer"},{"body":"I did a \"truncate peers\" and it hasn't repopulated (I waited about 10 mins). Do I need to restart one or more nodes to trigger repopulating of the table?","from":"developer"},{"body":"Yes.","from":"developer"},{"body":"The load_ring_state=false directive should probably also clear out the peers table because otherwise, the state that you're trying to get rid of is still persisted there.","from":"developer"},{"body":"If a user knows enough to disable loading the state and that fixes the problem, they can clear the peers table manually.","from":"developer"},{"body":"[~enigmacurry] can you reproduce?","from":"developer"},{"body":"First attempt appears to work correctly on cassandra-2.0 HEAD and 1.2.9 : \n\n{code}\n12:53 PM:~$ ccm create -v git:cassandra-1.2.9 t\nFetching Cassandra updates...\nCurrent cluster is now: t\n12:53 PM:~$ ccm populate -n 5\n12:54 PM:~$ ccm start\n12:54 PM:~$ ccm node1 stress\nCreated keyspaces. Sleeping 1s for propagation.\ntotal,interval_op_rate,interval_key_rate,latency/95th/99th,elapsed_time\n24994,2499,2499,9.5,55.2,179.0,10\n103123,7812,7812,2.8,27.2,134.7,20\n236358,13323,13323,1.7,15.4,134.7,30\n329477,9311,9311,1.7,9.8,109.8,40\n405667,7619,7619,1.8,9.2,6591.9,50\n558989,15332,15332,1.5,6.6,6591.1,60\n^C12:55 PM:~$ ccm node1 cqlsh\nConnected to t at 127.0.0.1:9160.\n[cqlsh 3.1.7 | Cassandra 1.2.9-SNAPSHOT | CQL spec 3.0.0 | Thrift protocol 19.36.0]\nUse HELP for help.\ncqlsh> select peer from system.peers;\n\n peer\n-----------\n 127.0.0.3\n 127.0.0.2\n 127.0.0.5\n 127.0.0.4\n\ncqlsh>\n12:55 PM:~$ ccm node2 decommission\n12:57 PM:~$ ccm node1 cqlsh\nConnected to t at 127.0.0.1:9160.\n[cqlsh 3.1.7 | Cassandra 1.2.9-SNAPSHOT | CQL spec 3.0.0 | Thrift protocol 19.36.0]\nUse HELP for help.\ncqlsh> select peer from system.peers;\n\n peer\n-----------\n 127.0.0.3\n 127.0.0.5\n 127.0.0.4\n\ncqlsh>\n12:58 PM:~$\n{code}\n\nAll nodes show equivalent peers table.","from":"developer"},{"body":"OK, reproduced this by killing -9 one of the nodes and then doing a 'nodetool removenode':\n\n{code}\n01:20 PM:~$ kill -9 18961 (PID of node1)\n01:21 PM:~$ ccm node1 status\nFailed to connect to '127.0.0.1:7100': Connection refused\n01:21 PM:~$ ccm node2 status\nDatacenter: datacenter1\n=======================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Owns Host ID Token Rack\nDN 127.0.0.1 62.93 KB 20.0% 896644af-8640-4be6-a3ff-e8ed559d851c -9223372036854775808 rack1\nUN 127.0.0.2 51.17 KB 20.0% d3801466-d36d-428c-b4e5-05ff69fe36c0 -5534023222112865485 rack1\nUN 127.0.0.3 62.78 KB 20.0% cb36c3ad-df45-4f77-bff5-ca93c504ec08 -1844674407370955162 rack1\nUN 127.0.0.4 51.17 KB 20.0% 89031a05-a3f6-4ac7-9d29-6caa0c609dbc 1844674407370955161 rack1\nUN 127.0.0.5 51.27 KB 20.0% 4909d856-a86e-493a-a7d0-7570d71eb9d8 5534023222112865484 rack1\n\n# Issue removenode on node3 :\n01:21 PM:~$ ~/.ccm/t/node1/bin/nodetool -p 7300 removenode 896644af-8640-4be6-a3ff-e8ed559d851c\n\n01:22 PM:~$ ccm node3 cqlsh\nConnected to t at 127.0.0.3:9160.\n[cqlsh 4.1.0 | Cassandra 2.0.3-SNAPSHOT | CQL spec 3.1.1 | Thrift protocol 19.39.0]\nUse HELP for help.\ncqlsh> select * from system.peers;\n\n peer | data_center | host_id | preferred_ip | rack | release_version | rpc_address | schema_version | tokens\n-----------+-------------+--------------------------------------+--------------+-------+-----------------+-------------+--------------------------------------+--------------------------\n 127.0.0.2 | datacenter1 | d3801466-d36d-428c-b4e5-05ff69fe36c0 | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.2 | d133398f-f287-3674-83af-a1b04ee29f1f | {'-5534023222112865485'}\n 127.0.0.5 | datacenter1 | 4909d856-a86e-493a-a7d0-7570d71eb9d8 | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.5 | d133398f-f287-3674-83af-a1b04ee29f1f | {'5534023222112865484'}\n 127.0.0.4 | datacenter1 | 89031a05-a3f6-4ac7-9d29-6caa0c609dbc | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.4 | d133398f-f287-3674-83af-a1b04ee29f1f | {'1844674407370955161'}\n\n(3 rows)\n\n# Check node2 peers table:\n\n01:23 PM:~$ ccm node2 cqlsh\nConnected to t at 127.0.0.2:9160.\n[cqlsh 4.1.0 | Cassandra 2.0.3-SNAPSHOT | CQL spec 3.1.1 | Thrift protocol 19.39.0]\nUse HELP for help.\ncqlsh> select * from system.peers;\n\n peer | data_center | host_id | preferred_ip | rack | release_version | rpc_address | schema_version | tokens\n-----------+-------------+--------------------------------------+--------------+-------+-----------------+-------------+--------------------------------------+--------------------------\n 127.0.0.3 | datacenter1 | cb36c3ad-df45-4f77-bff5-ca93c504ec08 | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.3 | d133398f-f287-3674-83af-a1b04ee29f1f | {'-1844674407370955162'}\n 127.0.0.1 | null | 896644af-8640-4be6-a3ff-e8ed559d851c | null | null | null | 127.0.0.1 | null | null\n 127.0.0.5 | datacenter1 | 4909d856-a86e-493a-a7d0-7570d71eb9d8 | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.5 | d133398f-f287-3674-83af-a1b04ee29f1f | {'5534023222112865484'}\n 127.0.0.4 | datacenter1 | 89031a05-a3f6-4ac7-9d29-6caa0c609dbc | null | rack1 | 2.0.3-SNAPSHOT | 127.0.0.4 | d133398f-f287-3674-83af-a1b04ee29f1f | {'1844674407370955161'}\n\n(4 rows)\n\n# oh noes!... node2 still has an entry for node1 in peers table.\n\n01:23 PM:~$ ccm node2 status\nDatacenter: datacenter1\n=======================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Owns Host ID Token Rack\nUN 127.0.0.2 51.17 KB 40.0% d3801466-d36d-428c-b4e5-05ff69fe36c0 -5534023222112865485 rack1\nUN 127.0.0.3 62.78 KB 20.0% cb36c3ad-df45-4f77-bff5-ca93c504ec08 -1844674407370955162 rack1\nUN 127.0.0.4 51.17 KB 20.0% 89031a05-a3f6-4ac7-9d29-6caa0c609dbc 1844674407370955161 rack1\nUN 127.0.0.5 51.27 KB 20.0% 4909d856-a86e-493a-a7d0-7570d71eb9d8 5534023222112865484 rack1\n\n{code}\n\nBy issuing the removenode on node3, node3 seems to know about the node being removed and it's peers table is correct. node2, although it's status output shows node1 going away, it's peers table has not been updated.","from":"developer"},{"body":"Thanks, Ryan.","from":"developer"},{"body":"Can you provide debug logs from node2?","from":"developer"},{"body":"I was able to repro with Ryan's steps, so I can upload the logs if you'd like, but I should be able to figure this one out.","from":"developer"},{"body":"The problem was that state changes for the removed node were being handled after the system.peers row was deleted. These state changes would result in update to the system.peers row, partially reviving it.\n\n6053-v1.patch (and [branch|https://github.com/thobbs/cassandra/tree/CASSANDRA-6053]) avoids updating the system.peers table if the node is unknown or is in one of the \"dead\" states and adds some basic unit test coverage.","from":"developer"},{"body":"Committed.","from":"developer"},{"body":"[~brandon.williams] FYI - Even though, this was fixed it surfaced in 2.0.14 release that we have in production. We are going to follow the work-around as mentioned above.","from":"developer"}],"created":"2013-09-18T09:06:16.000+0000","description":"After decommissioning my cluster from 20 to 9 nodes using opscenter, I found all but one of the nodes had incorrect system.peers tables.\n\nThis became a problem (afaik) when using the python-driver, since this queries the peers table to set up its connection pool. Resulting in very slow startup times, because of timeouts.\n\nThe output of nodetool didn't seem to be affected. After removing the incorrect entries from the peers tables, the connection issues seem to have disappeared for us. \n\nWould like some feedback on if this was the right way to handle the issue or if I'm still left with a broken cluster.\n\nAttached is the output of nodetool status, which shows the correct 9 nodes. Below that the output of the system.peers tables on the individual nodes.\n","issue_id":"12669215","key":"CASSANDRA-6053","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-01-10T18:52:04.000+0000","role":"fixed_distractor","summary":"system.peers table not updated after decommissioning nodes in C* 2.0"} {"case_id":"12670272","cluster":"DISTRACTOR-CASSANDRA-6086","comments":[{"body":"The primary reason we have removeUnfinishedCompacitonLeftovers check is to not overcount counters stored.\nIf we do not stop at the original exception, then we will have both compacted SSTables and leftovers which may produce unexpected counter value.\n\nThough we have the report that \"this happens\"(CASSANDRA-6008), we should fix somehow.\nMaybe we can consider compaction is finished even if we have entry in compactions_in_progress but some or all of the input SSTables are missing.","created":"2013-09-24T15:31:46.891+0000"},{"body":"Well, the common with CASSANDRA-6008 is that in my case node was dead before restart due to OOM. However, unlike 6008, CFs on which it reported this error was never truncated.","created":"2013-09-24T16:06:46.145+0000"},{"body":"From the other hand, the leftovers about to remove are already missing (i.e. already removed) when exception is thrown. I am not sure how can it influence counter value.","created":"2013-09-24T16:09:58.607+0000"},{"body":"bq. From the other hand, the leftovers about to remove are already missing (i.e. already removed) when exception is thrown.\n\nIs this the case that, for example, you have SSTable 'C' that is produced from compacting SSTables 'A' and 'B'('C' has ancestors 'A' and 'B'), and 'A' and 'B' are already deleted, but removeUnfinishedCompactionLeftovers reports you have unfinished compaction of 'A' and 'B'?\n\nIn above case, you definitely don't want to proceed, since the code tries remove SSTable 'C' afterwards.\nPatch v2 is attached to prevent deleting SSTable if its ancestors are already missing after printing warning.\n\n(I'm still looking for the case why the above happened though. We supposed not to have entries in compaction_in_progress when those are removed.)\n","created":"2013-10-08T17:21:16.761+0000"},{"body":"Well, how would one repair this for now, beyond reconstructing a virgin node and then replication? If there are counters, certainly you could provide an option for an admin to reset them or accept invalid values and get access to the rest of the data, or mark it in some (tombstone-ish?) way that other nodes with more accurate values for the counters can then replicate the correct state?\n\nAnd yes, we just got this.","created":"2013-10-17T15:04:38.150+0000"},{"body":"@Constance,\n\nI was able to repair my node - see CASSANDRA-6008. Used the suggestion I received + had to add some flavor to it :)","created":"2013-10-17T15:08:13.767+0000"},{"body":"Hi, we are able to consistently reproduce this issue:\n{noformat}\nERROR 23:14:06,001 Exception encountered during startup\njava.lang.IllegalStateException: Unfinished compactions reference missing sstables. This should never happen since compactions are marked finished before we start removing the old sstables.\n at org.apache.cassandra.db.ColumnFamilyStore.removeUnfinishedCompactionLeftovers(ColumnFamilyStore.java:489)\n at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:264)\n at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:461)\n at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:504)\njava.lang.IllegalStateException: Unfinished compactions reference missing sstables. This should never happen since compactions are marked finished before we start removing the old sstables.\n at org.apache.cassandra.db.ColumnFamilyStore.removeUnfinishedCompactionLeftovers(ColumnFamilyStore.java:489)\n at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:264)\n at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:461)\n at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:504)\nException encountered during startup: Unfinished compactions reference missing sstables. This should never happen since compactions are marked finished before we start removing the old sstables.\n{noformat}\n\nHere are the two ways in which we have found to reproduce this issue:\n# Buffer a large amount of CL.LOCAL_QUORUM writes into a Cassandra CF, then drop the CF while the writes are still buffered.\n# Buffer a large amount of CL.LOCAL_QUORUM writes into a Cassandra CF, then restart some nodes while before the buffer has finished draining.","created":"2013-11-08T22:26:09.179+0000"},{"body":"Rebased patch against cassandra-2.0 attached.\n\nAgain, the patch changes log ERROR and stop starting C* to print WARN but prevent deleting SSTable if its ancestors are already missing.","created":"2013-11-19T17:14:31.317+0000"},{"body":"[~yukim] the patch looks pretty good. Can you add a test like {{ColumnFamilyStoreTest.testRemoveUnifinishedCompactionLeftovers()}} to cover this? Also, warning logs should either suggest an action or let the user know that no action is required; in this case, we should tell the user to report the issue but that no corrective action is needed.","created":"2013-12-11T20:21:33.095+0000"},{"body":"[~thobbs] Added test and slightly changed message: https://github.com/yukim/cassandra/commits/6086\n\nI'm not good at wording, so suggestion is welcome.","created":"2013-12-17T20:23:37.814+0000"},{"body":"The v3 patch (and [branch|https://github.com/thobbs/cassandra/tree/6086]) builds on [~yukim]'s v2 patch and adds incremental deletions of entries in {{compactions_in_progress}} as discussed in CASSANDRA-6008. Additionally, this move the warning log to debug level, since there's not anything the user can do or should do.","created":"2013-12-19T23:58:45.895+0000"},{"body":"Thanks Tyler, committed.\n(I removed LegacyLeveledManifest related change in trunk since it no longer exists.)","created":"2013-12-30T20:05:28.028+0000"}],"conversations":[{"body":"Node refuses to start with\n{code}\nCaused by: java.lang.IllegalStateException: Unfinished compactions reference missing sstables. This should never happen since compactions are marked finished before we start removing the old sstables.\n at org.apache.cassandra.db.ColumnFamilyStore.removeUnfinishedCompactionLeftovers(ColumnFamilyStore.java:544)\n at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:262)\n{code}\n\nIMO, there is no reason to refuse to start discivering files that must be removed are already removed. It looks like pure bug diagnostic code and mean nothing to operator (nor he can do anything about this).\n\nReplaced throw of excepion with dump of diagnostic warning and continue startup.\n","from":"reporter","subject":"Node refuses to start with exception in ColumnFamilyStore.removeUnfinishedCompactionLeftovers when find that some to be removed files are already removed"},{"body":"The primary reason we have removeUnfinishedCompacitonLeftovers check is to not overcount counters stored.\nIf we do not stop at the original exception, then we will have both compacted SSTables and leftovers which may produce unexpected counter value.\n\nThough we have the report that \"this happens\"(CASSANDRA-6008), we should fix somehow.\nMaybe we can consider compaction is finished even if we have entry in compactions_in_progress but some or all of the input SSTables are missing.","from":"developer"},{"body":"Well, the common with CASSANDRA-6008 is that in my case node was dead before restart due to OOM. However, unlike 6008, CFs on which it reported this error was never truncated.","from":"developer"},{"body":"From the other hand, the leftovers about to remove are already missing (i.e. already removed) when exception is thrown. I am not sure how can it influence counter value.","from":"developer"},{"body":"bq. From the other hand, the leftovers about to remove are already missing (i.e. already removed) when exception is thrown.\n\nIs this the case that, for example, you have SSTable 'C' that is produced from compacting SSTables 'A' and 'B'('C' has ancestors 'A' and 'B'), and 'A' and 'B' are already deleted, but removeUnfinishedCompactionLeftovers reports you have unfinished compaction of 'A' and 'B'?\n\nIn above case, you definitely don't want to proceed, since the code tries remove SSTable 'C' afterwards.\nPatch v2 is attached to prevent deleting SSTable if its ancestors are already missing after printing warning.\n\n(I'm still looking for the case why the above happened though. We supposed not to have entries in compaction_in_progress when those are removed.)\n","from":"developer"},{"body":"Well, how would one repair this for now, beyond reconstructing a virgin node and then replication? If there are counters, certainly you could provide an option for an admin to reset them or accept invalid values and get access to the rest of the data, or mark it in some (tombstone-ish?) way that other nodes with more accurate values for the counters can then replicate the correct state?\n\nAnd yes, we just got this.","from":"developer"},{"body":"@Constance,\n\nI was able to repair my node - see CASSANDRA-6008. Used the suggestion I received + had to add some flavor to it :)","from":"developer"},{"body":"Hi, we are able to consistently reproduce this issue:\n{noformat}\nERROR 23:14:06,001 Exception encountered during startup\njava.lang.IllegalStateException: Unfinished compactions reference missing sstables. This should never happen since compactions are marked finished before we start removing the old sstables.\n at org.apache.cassandra.db.ColumnFamilyStore.removeUnfinishedCompactionLeftovers(ColumnFamilyStore.java:489)\n at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:264)\n at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:461)\n at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:504)\njava.lang.IllegalStateException: Unfinished compactions reference missing sstables. This should never happen since compactions are marked finished before we start removing the old sstables.\n at org.apache.cassandra.db.ColumnFamilyStore.removeUnfinishedCompactionLeftovers(ColumnFamilyStore.java:489)\n at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:264)\n at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:461)\n at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:504)\nException encountered during startup: Unfinished compactions reference missing sstables. This should never happen since compactions are marked finished before we start removing the old sstables.\n{noformat}\n\nHere are the two ways in which we have found to reproduce this issue:\n# Buffer a large amount of CL.LOCAL_QUORUM writes into a Cassandra CF, then drop the CF while the writes are still buffered.\n# Buffer a large amount of CL.LOCAL_QUORUM writes into a Cassandra CF, then restart some nodes while before the buffer has finished draining.","from":"developer"},{"body":"Rebased patch against cassandra-2.0 attached.\n\nAgain, the patch changes log ERROR and stop starting C* to print WARN but prevent deleting SSTable if its ancestors are already missing.","from":"developer"},{"body":"[~yukim] the patch looks pretty good. Can you add a test like {{ColumnFamilyStoreTest.testRemoveUnifinishedCompactionLeftovers()}} to cover this? Also, warning logs should either suggest an action or let the user know that no action is required; in this case, we should tell the user to report the issue but that no corrective action is needed.","from":"developer"},{"body":"[~thobbs] Added test and slightly changed message: https://github.com/yukim/cassandra/commits/6086\n\nI'm not good at wording, so suggestion is welcome.","from":"developer"},{"body":"The v3 patch (and [branch|https://github.com/thobbs/cassandra/tree/6086]) builds on [~yukim]'s v2 patch and adds incremental deletions of entries in {{compactions_in_progress}} as discussed in CASSANDRA-6008. Additionally, this move the warning log to debug level, since there's not anything the user can do or should do.","from":"developer"},{"body":"Thanks Tyler, committed.\n(I removed LegacyLeveledManifest related change in trunk since it no longer exists.)","from":"developer"}],"created":"2013-09-24T13:30:48.000+0000","description":"Node refuses to start with\n{code}\nCaused by: java.lang.IllegalStateException: Unfinished compactions reference missing sstables. This should never happen since compactions are marked finished before we start removing the old sstables.\n at org.apache.cassandra.db.ColumnFamilyStore.removeUnfinishedCompactionLeftovers(ColumnFamilyStore.java:544)\n at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:262)\n{code}\n\nIMO, there is no reason to refuse to start discivering files that must be removed are already removed. It looks like pure bug diagnostic code and mean nothing to operator (nor he can do anything about this).\n\nReplaced throw of excepion with dump of diagnostic warning and continue startup.\n","issue_id":"12670272","key":"CASSANDRA-6086","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-12-30T20:05:27.000+0000","role":"fixed_distractor","summary":"Node refuses to start with exception in ColumnFamilyStore.removeUnfinishedCompactionLeftovers when find that some to be removed files are already removed"} {"case_id":"12672921","cluster":"DISTRACTOR-CASSANDRA-6168","comments":[{"body":"Since we need KS to pick the right replication strategy, I vote we just leave Owns out for no-KS situation. WDYT [~brandon.williams]?","created":"2013-10-21T04:09:15.841+0000"},{"body":"Well, we do issue a warning when no keyspace is specified.","created":"2013-10-21T04:11:30.341+0000"},{"body":"we actually don't for status, we need to make it act like ring","created":"2014-03-20T05:07:52.698+0000"},{"body":"Care to take a stab, Vijay?","created":"2014-03-20T05:14:10.432+0000"},{"body":"Hi Brandon, Sure, Thanks!","created":"2014-03-20T15:36:13.808+0000"},{"body":"One line change.","created":"2014-03-22T04:26:09.109+0000"},{"body":"+1","created":"2014-03-22T12:51:40.760+0000"},{"body":"Committed Thanks!","created":"2014-03-22T20:59:59.783+0000"}],"conversations":[{"body":"Seen in 1.2.10.\n\nApologies if this is expected behavior. Nodetool status reports 0% ownership unless I add a keyspace name.\n\nnodetool help docs says:\n...\" status - Print cluster information (state, load, IDs, ...)\"...\n\noutput without keyspace name\n{code}\nDatacenter: DC1\n===============\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns Host ID Rack\nUN 10.x.x.146 81.96 GB 256 0.0% a70c59b3-a667-4d76-ba5b-ba849ad672da r1\nUN 10.x.x.63 95.32 GB 256 0.0% f8cb7b10-4ebe-484a-a1c0-6cb2d053901b r1\nUN 10.x.x.184 89.54 GB 256 0.1% cd86c420-55e2-4d99-8ed9-d9ee8d6a9d9c r1\nUN 10.x.x.190 79.68 GB 256 0.0% 544c3906-bc02-400d-9fd2-1e39ecadd6ff r1\nUN 10.x.x.168 93.44 GB 256 0.7% 33be316f-1276-475d-90cf-2667950d3a2c r1\nUN 10.x.x.132 84.4 GB 256 0.0% b327d9f1-cab0-4583-8e5e-95c50b4074fd r1\nDatacenter: DCOFFLINE\n=====================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns Host ID Rack\nUN 10.x.x.62 56.09 GB 256 32.4% c8994d27-767b-431f-bdc2-9196eeeb6f44 r1\nUN 10.x.x.131 60.11 GB 256 32.8% 0b9d3314-039e-4f88-8ba6-d0f2885d9a30 r1\nUN 10.x.x.167 56.45 GB 256 34.0% ba76f4fe-4250-4839-a37d-c1a7c24e585d r1\n{code}\n\nand with keyspace. Example: nodetool status MYKSPS\n\n{code}\nDatacenter: DC1\n===============\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.x.x.184 89.51 GB 256 50.0% cd86c420-55e2-4d99-8ed9-d9ee8d6a9d9c r1\nUN 10.x.x.146 81.96 GB 256 50.0% a70c59b3-a667-4d76-ba5b-ba849ad672da r1\nUN 10.x.x.168 93.44 GB 256 50.0% 33be316f-1276-475d-90cf-2667950d3a2c r1\nUN 10.x.x.63 95.32 GB 256 50.0% f8cb7b10-4ebe-484a-a1c0-6cb2d053901b r1\nUN 10.x.x.190 79.68 GB 256 50.0% 544c3906-bc02-400d-9fd2-1e39ecadd6ff r1\nUN 10.x.x.132 84.4 GB 256 50.0% b327d9f1-cab0-4583-8e5e-95c50b4074fd r1\nDatacenter: DCOFFLINE\n=====================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.x.x.131 60.11 GB 256 32.8% 0b9d3314-039e-4f88-8ba6-d0f2885d9a30 r1\nUN 10.x.x.167 56.45 GB 256 34.7% ba76f4fe-4250-4839-a37d-c1a7c24e585d r1\nUN 10.x.x.62 56.09 GB 256 32.5% c8994d27-767b-431f-bdc2-9196eeeb6f44 r1\n{code}\n","from":"reporter","subject":"nodetool status should issue a warning when no keyspace is specified"},{"body":"Since we need KS to pick the right replication strategy, I vote we just leave Owns out for no-KS situation. WDYT [~brandon.williams]?","from":"developer"},{"body":"Well, we do issue a warning when no keyspace is specified.","from":"developer"},{"body":"we actually don't for status, we need to make it act like ring","from":"developer"},{"body":"Care to take a stab, Vijay?","from":"developer"},{"body":"Hi Brandon, Sure, Thanks!","from":"developer"},{"body":"One line change.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed Thanks!","from":"developer"}],"created":"2013-10-08T23:03:07.000+0000","description":"Seen in 1.2.10.\n\nApologies if this is expected behavior. Nodetool status reports 0% ownership unless I add a keyspace name.\n\nnodetool help docs says:\n...\" status - Print cluster information (state, load, IDs, ...)\"...\n\noutput without keyspace name\n{code}\nDatacenter: DC1\n===============\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns Host ID Rack\nUN 10.x.x.146 81.96 GB 256 0.0% a70c59b3-a667-4d76-ba5b-ba849ad672da r1\nUN 10.x.x.63 95.32 GB 256 0.0% f8cb7b10-4ebe-484a-a1c0-6cb2d053901b r1\nUN 10.x.x.184 89.54 GB 256 0.1% cd86c420-55e2-4d99-8ed9-d9ee8d6a9d9c r1\nUN 10.x.x.190 79.68 GB 256 0.0% 544c3906-bc02-400d-9fd2-1e39ecadd6ff r1\nUN 10.x.x.168 93.44 GB 256 0.7% 33be316f-1276-475d-90cf-2667950d3a2c r1\nUN 10.x.x.132 84.4 GB 256 0.0% b327d9f1-cab0-4583-8e5e-95c50b4074fd r1\nDatacenter: DCOFFLINE\n=====================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns Host ID Rack\nUN 10.x.x.62 56.09 GB 256 32.4% c8994d27-767b-431f-bdc2-9196eeeb6f44 r1\nUN 10.x.x.131 60.11 GB 256 32.8% 0b9d3314-039e-4f88-8ba6-d0f2885d9a30 r1\nUN 10.x.x.167 56.45 GB 256 34.0% ba76f4fe-4250-4839-a37d-c1a7c24e585d r1\n{code}\n\nand with keyspace. Example: nodetool status MYKSPS\n\n{code}\nDatacenter: DC1\n===============\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.x.x.184 89.51 GB 256 50.0% cd86c420-55e2-4d99-8ed9-d9ee8d6a9d9c r1\nUN 10.x.x.146 81.96 GB 256 50.0% a70c59b3-a667-4d76-ba5b-ba849ad672da r1\nUN 10.x.x.168 93.44 GB 256 50.0% 33be316f-1276-475d-90cf-2667950d3a2c r1\nUN 10.x.x.63 95.32 GB 256 50.0% f8cb7b10-4ebe-484a-a1c0-6cb2d053901b r1\nUN 10.x.x.190 79.68 GB 256 50.0% 544c3906-bc02-400d-9fd2-1e39ecadd6ff r1\nUN 10.x.x.132 84.4 GB 256 50.0% b327d9f1-cab0-4583-8e5e-95c50b4074fd r1\nDatacenter: DCOFFLINE\n=====================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.x.x.131 60.11 GB 256 32.8% 0b9d3314-039e-4f88-8ba6-d0f2885d9a30 r1\nUN 10.x.x.167 56.45 GB 256 34.7% ba76f4fe-4250-4839-a37d-c1a7c24e585d r1\nUN 10.x.x.62 56.09 GB 256 32.5% c8994d27-767b-431f-bdc2-9196eeeb6f44 r1\n{code}\n","issue_id":"12672921","key":"CASSANDRA-6168","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-03-22T20:59:59.000+0000","role":"fixed_distractor","summary":"nodetool status should issue a warning when no keyspace is specified"} {"case_id":"12676851","cluster":"DISTRACTOR-CASSANDRA-6276","comments":[{"body":"I think we're looking at CASSANDRA-5202 here.","created":"2013-10-31T13:15:04.158+0000"},{"body":"5202 is drop/recreate of an entire table.","created":"2013-10-31T13:39:45.189+0000"},{"body":"This is quite frustrating when developing and debugging with large portions of data that need to be re-inserted every time you have a schema change because you have to drop an entire CF to make minor schema changes like this...\n\nIs there a workaround one could use until this is resolved?\n\nThanks :)","created":"2014-01-23T10:23:12.227+0000"},{"body":"/cc [~slebresne]","created":"2014-06-05T20:52:51.628+0000"},{"body":"Unfortunately, we can't allow dropping a component from the comparator, including dropping individual collection columns from ColumnToCollectionType.\n\nIf we do allow that, and have pre-existing data of that type, C* simply wouldn't know how to compare those:\n\n{code:title=ColumnToCollectionType.java}\n public int compareCollectionMembers(ByteBuffer o1, ByteBuffer o2, ByteBuffer collectionName)\n {\n CollectionType t = defined.get(collectionName);\n if (t == null)\n throw new RuntimeException(ByteBufferUtil.bytesToHex(collectionName) + \" is not defined as a collection\");\n\n return t.nameComparator().compare(o1, o2);\n }\n{code}\n\nA simple algorithm to hit the RTE:\n1. create table test (id int primary key, col1 map, col2 set);\n2. insert into test (id, col1, col2) VALUES ( 0, \\{0:0, 1:1\\}, \\{0,1\\});\n3. flush\n4. update test set col1 = col1 + \\{2:2\\}, col2 = col2 + \\{2\\} where id = 0;\n5. flush\n6. select * from test;\n\n{noformat}\njava.lang.RuntimeException: 636f6c31 is not defined as a collection\n\tat org.apache.cassandra.db.marshal.ColumnToCollectionType.compareCollectionMembers(ColumnToCollectionType.java:79) ~[main/:na]\n\tat org.apache.cassandra.db.composites.CompoundSparseCellNameType$WithCollection.compare(CompoundSparseCellNameType.java:296) ~[main/:na]\n\tat org.apache.cassandra.db.composites.AbstractCellNameType$1.compare(AbstractCellNameType.java:61) ~[main/:na]\n\tat org.apache.cassandra.db.composites.AbstractCellNameType$1.compare(AbstractCellNameType.java:58) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$Candidate.compareTo(MergeIterator.java:154) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$Candidate.compareTo(MergeIterator.java:131) ~[main/:na]\n{noformat}\n\nFor this reason alone we can't allow getting rid of a comparator component.\n\nHowever, even if we did, and allowed to create a different collection with the same name, we'd hit a different issue: the new collection's comparator would be used to compare potentially incompatible types. Now, your unit tests aren't failing b/c most of our comparators assume valid values and don't perform extra validation, then use something like ByteBufferUtil.compareUnsigned() to compare the values, which doesn't fail and will just stop once the shortest BB gets exhausted. One exception is tuples/usertypes - they *do* expect at least length to be there, and will throw an exception.\n\nExample:\n1. create table test (id int primary key, col set);\n2. insert into test (id, col) values (0, \\{true,false\\});\n3. alter table test drop col;\n4. create type test (f1 int);\n5. alter table test add col set;\n6. update test set col = col + \\{ \\{f1 : 0 \\} \\} where id = 0;\n7. select * from test;\n\n{noformat}\njava.nio.BufferUnderflowException: null\n\tat java.nio.Buffer.nextGetIndex(Buffer.java:498) ~[na:1.7.0_65]\n\tat java.nio.HeapByteBuffer.getInt(HeapByteBuffer.java:355) ~[na:1.7.0_65]\n\tat org.apache.cassandra.db.marshal.TupleType.compare(TupleType.java:80) ~[main/:na]\n\tat org.apache.cassandra.db.marshal.TupleType.compare(TupleType.java:38) ~[main/:na]\n\tat org.apache.cassandra.db.marshal.ColumnToCollectionType.compareCollectionMembers(ColumnToCollectionType.java:81) ~[main/:na]\n\tat org.apache.cassandra.db.composites.CompoundSparseCellNameType$WithCollection.compare(CompoundSparseCellNameType.java:296) ~[main/:na]\n\tat org.apache.cassandra.db.composites.AbstractCellNameType$1.compare(AbstractCellNameType.java:61) ~[main/:na]\n\tat org.apache.cassandra.db.composites.AbstractCellNameType$1.compare(AbstractCellNameType.java:58) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$Candidate.compareTo(MergeIterator.java:154) ~[main/:na]\n{noformat}","created":"2014-07-31T12:53:55.604+0000"},{"body":"So the best we can do is provide a better exception message, unfortunately.\n\nThis also unfortunately complicates CASSANDRA-6717, since we don't store the comparator itself anymore and reconstruct it from the columns we have, and the current dropped_columns map only has names and drop timestamps in it. This means we must store the type of the dropped columns as well, in case some of those are collections, which is annoying.","created":"2014-07-31T12:58:32.486+0000"},{"body":"You're right, I'm not sure there is much we can do about this.\n\nbq. So the best we can do is provide a better exception message, unfortunately.\n\nI agree. Let's detect that case and provide a meaningful error message for now.\n\nIn the longer run, one way to maybe fix this would be to push dropped_columns down to the sstable reading level. If we were skipping cells as soon as they are deserialized before they ever hit any comparator we would be free to update the comparator. Probably messy/inefficient to do with the current code, but might become feasible with 3.0 storage engine changes, we'll see. ","created":"2014-08-04T17:16:01.508+0000"},{"body":"bq. In the longer run, one way to maybe fix this would be to push dropped_columns down to the sstable reading level. If we were skipping cells as soon as they are deserialized before they ever hit any comparator we would be free to update the comparator. Probably messy/inefficient to do with the current code, but might become feasible with 3.0 storage engine changes, we'll see. \n\nTrue. Would still be tricky somewhat, given that schema updates don't propagate instantly, but we'll see.","created":"2014-08-04T17:21:42.143+0000"},{"body":"3.0 storage engine won't be universal for a while (maybe never for thrift), but will index directly into columns (i.e. won't touch any not requested), so could trivially avoid retrieving data for dropped columns. The only problem is we'd need to track the range of sstables for which they were previously dropped (and maybe contains stale data), and which we now apply the new comparator too, which would be a bit ugly/annoying.","created":"2014-08-04T17:22:57.199+0000"},{"body":"This patch check if a collection of a different type and with the same name already existed and if it is the case it will send an error with the following message:\n\"Cannot add a collection with the name because a collection with the same name and a different type has already been used in the past\"\n","created":"2014-08-05T20:25:59.900+0000"},{"body":"Patch looks good, though adding a method to {{CellNameType}} feels a tad overkill to me, I'd rather keep the validation in {{AlterTableStatement}} (it's arguably a personal preference). The other reason being that we should fix this in 2.0 and we don't have {{CellNameType}} there. So anyway, attaching a slightly simpler alternative for 2.0. If we're good with that, I'll push a simple dtest so 2.0 is covered and I'll include the tests from [~blerer] patch while merging with 2.1 (since CqlTester is not in 2.0).","created":"2014-08-11T15:35:49.092+0000"},{"body":"2.0 patch LGTM","created":"2014-08-11T15:37:29.620+0000"},{"body":"Alright, committed, thanks","created":"2014-08-11T16:24:32.026+0000"},{"body":"Facing the same problem in cassandra v3.7. Btw, it works when you drop the column, recreate a non collection column such as int with the same name and then drop that again. After that you can add the same name column with a different collection type and cassandra allows the operation. Dirty workaround though!","created":"2016-08-04T13:06:01.700+0000"},{"body":"Dirty but also unsafe. There is a reason the limitation is there in the first place - try to 'work around it' and risk corruption.","created":"2016-09-20T01:36:43.916+0000"}],"conversations":[{"body":"If create a list, drop it and create a map with the same name, i get \"Bad Request: comparators do not match or are not compatible.\"\n\n{quote}\ncqlsh:os_test1> create table thetable(id timeuuid primary key, somevalue text);\ncqlsh:os_test1> alter table thetable add mycollection list; \ncqlsh:os_test1> alter table thetable drop mycollection;\ncqlsh:os_test1> alter table thetable add mycollection map; \nBad Request: comparators do not match or are not compatible.\n{quote}\n\n\n","from":"reporter","subject":"CQL: Map can not be created with the same name as a previously dropped list"},{"body":"I think we're looking at CASSANDRA-5202 here.","from":"developer"},{"body":"5202 is drop/recreate of an entire table.","from":"developer"},{"body":"This is quite frustrating when developing and debugging with large portions of data that need to be re-inserted every time you have a schema change because you have to drop an entire CF to make minor schema changes like this...\n\nIs there a workaround one could use until this is resolved?\n\nThanks :)","from":"developer"},{"body":"/cc [~slebresne]","from":"developer"},{"body":"Unfortunately, we can't allow dropping a component from the comparator, including dropping individual collection columns from ColumnToCollectionType.\n\nIf we do allow that, and have pre-existing data of that type, C* simply wouldn't know how to compare those:\n\n{code:title=ColumnToCollectionType.java}\n public int compareCollectionMembers(ByteBuffer o1, ByteBuffer o2, ByteBuffer collectionName)\n {\n CollectionType t = defined.get(collectionName);\n if (t == null)\n throw new RuntimeException(ByteBufferUtil.bytesToHex(collectionName) + \" is not defined as a collection\");\n\n return t.nameComparator().compare(o1, o2);\n }\n{code}\n\nA simple algorithm to hit the RTE:\n1. create table test (id int primary key, col1 map, col2 set);\n2. insert into test (id, col1, col2) VALUES ( 0, \\{0:0, 1:1\\}, \\{0,1\\});\n3. flush\n4. update test set col1 = col1 + \\{2:2\\}, col2 = col2 + \\{2\\} where id = 0;\n5. flush\n6. select * from test;\n\n{noformat}\njava.lang.RuntimeException: 636f6c31 is not defined as a collection\n\tat org.apache.cassandra.db.marshal.ColumnToCollectionType.compareCollectionMembers(ColumnToCollectionType.java:79) ~[main/:na]\n\tat org.apache.cassandra.db.composites.CompoundSparseCellNameType$WithCollection.compare(CompoundSparseCellNameType.java:296) ~[main/:na]\n\tat org.apache.cassandra.db.composites.AbstractCellNameType$1.compare(AbstractCellNameType.java:61) ~[main/:na]\n\tat org.apache.cassandra.db.composites.AbstractCellNameType$1.compare(AbstractCellNameType.java:58) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$Candidate.compareTo(MergeIterator.java:154) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$Candidate.compareTo(MergeIterator.java:131) ~[main/:na]\n{noformat}\n\nFor this reason alone we can't allow getting rid of a comparator component.\n\nHowever, even if we did, and allowed to create a different collection with the same name, we'd hit a different issue: the new collection's comparator would be used to compare potentially incompatible types. Now, your unit tests aren't failing b/c most of our comparators assume valid values and don't perform extra validation, then use something like ByteBufferUtil.compareUnsigned() to compare the values, which doesn't fail and will just stop once the shortest BB gets exhausted. One exception is tuples/usertypes - they *do* expect at least length to be there, and will throw an exception.\n\nExample:\n1. create table test (id int primary key, col set);\n2. insert into test (id, col) values (0, \\{true,false\\});\n3. alter table test drop col;\n4. create type test (f1 int);\n5. alter table test add col set;\n6. update test set col = col + \\{ \\{f1 : 0 \\} \\} where id = 0;\n7. select * from test;\n\n{noformat}\njava.nio.BufferUnderflowException: null\n\tat java.nio.Buffer.nextGetIndex(Buffer.java:498) ~[na:1.7.0_65]\n\tat java.nio.HeapByteBuffer.getInt(HeapByteBuffer.java:355) ~[na:1.7.0_65]\n\tat org.apache.cassandra.db.marshal.TupleType.compare(TupleType.java:80) ~[main/:na]\n\tat org.apache.cassandra.db.marshal.TupleType.compare(TupleType.java:38) ~[main/:na]\n\tat org.apache.cassandra.db.marshal.ColumnToCollectionType.compareCollectionMembers(ColumnToCollectionType.java:81) ~[main/:na]\n\tat org.apache.cassandra.db.composites.CompoundSparseCellNameType$WithCollection.compare(CompoundSparseCellNameType.java:296) ~[main/:na]\n\tat org.apache.cassandra.db.composites.AbstractCellNameType$1.compare(AbstractCellNameType.java:61) ~[main/:na]\n\tat org.apache.cassandra.db.composites.AbstractCellNameType$1.compare(AbstractCellNameType.java:58) ~[main/:na]\n\tat org.apache.cassandra.utils.MergeIterator$Candidate.compareTo(MergeIterator.java:154) ~[main/:na]\n{noformat}","from":"developer"},{"body":"So the best we can do is provide a better exception message, unfortunately.\n\nThis also unfortunately complicates CASSANDRA-6717, since we don't store the comparator itself anymore and reconstruct it from the columns we have, and the current dropped_columns map only has names and drop timestamps in it. This means we must store the type of the dropped columns as well, in case some of those are collections, which is annoying.","from":"developer"},{"body":"You're right, I'm not sure there is much we can do about this.\n\nbq. So the best we can do is provide a better exception message, unfortunately.\n\nI agree. Let's detect that case and provide a meaningful error message for now.\n\nIn the longer run, one way to maybe fix this would be to push dropped_columns down to the sstable reading level. If we were skipping cells as soon as they are deserialized before they ever hit any comparator we would be free to update the comparator. Probably messy/inefficient to do with the current code, but might become feasible with 3.0 storage engine changes, we'll see. ","from":"developer"},{"body":"bq. In the longer run, one way to maybe fix this would be to push dropped_columns down to the sstable reading level. If we were skipping cells as soon as they are deserialized before they ever hit any comparator we would be free to update the comparator. Probably messy/inefficient to do with the current code, but might become feasible with 3.0 storage engine changes, we'll see. \n\nTrue. Would still be tricky somewhat, given that schema updates don't propagate instantly, but we'll see.","from":"developer"},{"body":"3.0 storage engine won't be universal for a while (maybe never for thrift), but will index directly into columns (i.e. won't touch any not requested), so could trivially avoid retrieving data for dropped columns. The only problem is we'd need to track the range of sstables for which they were previously dropped (and maybe contains stale data), and which we now apply the new comparator too, which would be a bit ugly/annoying.","from":"developer"},{"body":"This patch check if a collection of a different type and with the same name already existed and if it is the case it will send an error with the following message:\n\"Cannot add a collection with the name because a collection with the same name and a different type has already been used in the past\"\n","from":"developer"},{"body":"Patch looks good, though adding a method to {{CellNameType}} feels a tad overkill to me, I'd rather keep the validation in {{AlterTableStatement}} (it's arguably a personal preference). The other reason being that we should fix this in 2.0 and we don't have {{CellNameType}} there. So anyway, attaching a slightly simpler alternative for 2.0. If we're good with that, I'll push a simple dtest so 2.0 is covered and I'll include the tests from [~blerer] patch while merging with 2.1 (since CqlTester is not in 2.0).","from":"developer"},{"body":"2.0 patch LGTM","from":"developer"},{"body":"Alright, committed, thanks","from":"developer"},{"body":"Facing the same problem in cassandra v3.7. Btw, it works when you drop the column, recreate a non collection column such as int with the same name and then drop that again. After that you can add the same name column with a different collection type and cassandra allows the operation. Dirty workaround though!","from":"developer"},{"body":"Dirty but also unsafe. There is a reason the limitation is there in the first place - try to 'work around it' and risk corruption.","from":"developer"}],"created":"2013-10-31T12:37:07.000+0000","description":"If create a list, drop it and create a map with the same name, i get \"Bad Request: comparators do not match or are not compatible.\"\n\n{quote}\ncqlsh:os_test1> create table thetable(id timeuuid primary key, somevalue text);\ncqlsh:os_test1> alter table thetable add mycollection list; \ncqlsh:os_test1> alter table thetable drop mycollection;\ncqlsh:os_test1> alter table thetable add mycollection map; \nBad Request: comparators do not match or are not compatible.\n{quote}\n\n\n","issue_id":"12676851","key":"CASSANDRA-6276","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-08-11T16:24:32.000+0000","role":"fixed_distractor","summary":"CQL: Map can not be created with the same name as a previously dropped list"} {"case_id":"12679044","cluster":"DISTRACTOR-CASSANDRA-6345","comments":[{"body":"Interesting. Could we optimize cOTM instead? That would definitely be the simplest solution.\n\nE.g. it looks like TreeMultimap.putAll actually loops over each entry and calls put one at a time which is the worst-case scenario for binary tree rebalancing -- quadratic time.","created":"2013-11-13T17:05:30.071+0000"},{"body":"bq. it looks like TreeMultimap.putAll actually loops over each entry and calls put one at a time\n\nI don't see a way around this: https://code.google.com/p/guava-libraries/issues/detail?id=1579","created":"2013-11-13T17:30:20.838+0000"},{"body":"A Treap? Can be cheaply built, cheaply merged and cheaply cloned.\n\nAlso, anything cheaply cloneable would work for that operation. A SnapTree that is wrapped to support multi-map functionality would also work.","created":"2013-11-13T17:34:59.896+0000"},{"body":"Yes. Too bad the implementation classes like AbstractSortedKeySortedSetMultimap are package-private.","created":"2013-11-13T18:01:51.064+0000"},{"body":"Actually I think a STM multimap would still get messy quickly since you need to do a \"deep\" clone -- cloning the top-level Map would leave the values (sub-collections) sharing a reference.","created":"2013-11-13T20:47:36.387+0000"},{"body":"Just have the wrapper make any updates to a collection replace the collection instead of modifying it.\n\n[NB: I haven't looked to see if this would have any negative performance implications on the update side, I'm assuming the reads are more frequent and/or collections small... if not, a Treap is probably the better choice as my old Jjoost implementation (IIRC) supports snapshotting and multiple values are dealt with inside the tree itself, not as a collection]","created":"2013-11-13T20:52:46.897+0000"},{"body":"bq. Proposal: In AbstractReplicationStrategy.getNaturalEndpoints(), cache the cloned TokenMetadata instance returned by TokenMetadata.cloneOnlyTokenMap(), wrapping it with a lock to prevent stampedes, and clearing it in clearEndpointCache().\n\nWhy not just use a sharded lock to prevent stampedes directly w/o the caching complexity, as in the attached?","created":"2013-11-13T21:06:52.464+0000"},{"body":"I see, with vnodes we have enough ranges that we can have a thundering herd even if each range only clones once.\n\nv2 attached with the approach you described originally.","created":"2013-11-13T21:26:40.735+0000"},{"body":"Attached a patch we deployed to production that fixed the issue.","created":"2013-11-13T21:44:04.090+0000"},{"body":"CPU user% graph during the rollout of the patch I attached on 1 DC (15 nodes) of the cluster. Around ~21:05 the patch starts to roll out and spikes are seen. The node in question receives the patch at ~21:30, and afterwards the spikes are gone. The rollout finishes at ~21:45.","created":"2013-11-13T21:48:40.349+0000"},{"body":"I have to admit I like it better without the custom wrapper class. :)","created":"2013-11-13T23:26:43.408+0000"},{"body":"Well, I started writing the patch this morning and I don't write multi-threaded Java code every day, so I'm overly careful ;) The only theoretical advantage to my patch is that it allows concurrent readers.","created":"2013-11-14T00:25:22.192+0000"},{"body":"Unfortunately both of the patches suffer from a deadlock, since the invalidation and fill are wrapped up in TokenMetadata's locks.\n\nT1 acquires cache read lock\nT2 acquires TokenMetadata write lock\nT1 acquires cache write lock on miss\nT2 is blocked on cache write lock trying to invalidate\nT1 is blocked on TokenMetadata read lock trying to cloneOnlyTokenMap to fill the cache\n\nTrying to work on a fix.","created":"2013-11-15T04:43:50.413+0000"},{"body":"Attached a new patch with the deadlock fixed. We're running this on a production cluster.\n\nThe primary issue was the callback for invalidation from TokenMetadata to all of the registered AbstractReplicationStrategy instances. This was asking for it anyway, so in the patch I replaced the \"push\" invalidation with simple versioning of the TokenMetadata endpoints. TokenMetadata bumps it's version number each time the cache would need to be invalidated, and AbstractReplicationStrategy checks it's version when it needs to do a read, invalidating if necessary. This gets the invalidation out of the gossip threads and into the RPC threads, which is probably a good thing. The only thing I'm not super crazy about is the extra hot path read lock acquisition on TokenMetadata.getEndpointVersion(), which might be avoidable.","created":"2013-11-15T07:28:35.789+0000"},{"body":"I think we can craft a simpler solution (v3) by using an AtomicReference to the TM clone. This removes the possibility of deadlock since clearEndpointCache now only makes non-blocking calls.\n\nI've also refined it to use a Striped per-keyToken, as well as synchronizing the TM clone itself, since concurrent endpoint computation is fine.","created":"2013-11-18T03:52:04.445+0000"},{"body":"I like the simpler approach. I still think the callbacks for invalidation are asking for it ;) I also think perhaps the stampede lock should be more explicit than a synchronized lock on \"this\" to prevent unintended blocking from future modifications.\n\nEither way, I think the only material concern I have is the order that TokenMetadata changes get applied to the caches in AbstractReplicationStrategy instances. Shouldn't the invalidation take place on all threads in all instances of AbstractReplicationStrategy before returning from an endpoint-mutating write operation in TokenMetadata? It seems as if just setting the cache to empty would allow a period of time where TokenMetadata write methods had returned but not all threads have seen the mutation yet because they are still holding onto the old clone of TM. This might be alright though, I'm not sure. Thoughts?","created":"2013-11-18T18:51:43.455+0000"},{"body":"bq. It seems as if just setting the cache to empty would allow a period of time where TokenMetadata write methods had returned but not all threads have seen the mutation yet\n\nI'm not 100% sure this is what you're talking about, but I see this problem with the existing code (and my v3):\n\n{noformat}\nThread 1 Thread 2 \ngetNaturalEndpoints \ncloneOnlyTokenMap \n invalidateCachedTokenEndpointValues\nendpoints = calculate\ncacheEndpoint [based on the now-invalidated token map]\n{noformat}\n\nSo it doesn't quite work. We'd need to introduce another AtomicReference on the cache, so that invalidate could create a new Map (so it doesn't matter if someone updates the old one). But I think you're right that getting rid of the callback approach entirely is better.","created":"2013-11-20T04:25:59.269+0000"},{"body":"v4 attached that uses a versioning approach like yours. I dropped the readLock acquire on version read since it's not necessary to block callers during the update. (A few extra over-broad replica set operations won't hurt.)","created":"2013-11-20T04:40:41.189+0000"},{"body":"+100 at removing those pub/sub callbacks :)\n\nThe concurrency issues I bring up are probably because I'm unfamiliar with the \"guarantees\" needed by TokenMetadata updates. It looks like the current release code is subject to the issue I brought up, where method calls on TokenMetadata that change state return successfully before all threads applying mutations have \"seen\" the update. There will be some mutations in progress that are using \"stale\" token data to apply writes even after TokenMetadata write methods returns as successful. So this does not appear to be a regression, but I'm just being overly cautious having been burned by these sort of double-caching scenarios before. You bring up the point that over-broad operations are ok, and I agree, but I'm more concerned about operations that are too narrow. It seems that unless I'm missing something either is possible with the current release code, and thus these patches as well (including mine).\n\nTokenMetadata#updateNormalTokens is (implicitly) relying on the removeFromMoving call to bump the version, but the tokenToEndpointMap is updated afterwards, which means internal data is updated after the version is bumped. IMHO to be defensive, any time the write lock is acquired in TokenMetadata, the version should be bumped in the finally block before the lock is released. I don't think this is exposing a bug in the existing patch though, because cloneOnlyTokenMap will be blocked until the write lock is released in the finally block.\n\nIs the idea with the striped lock on the endpoint cache in AbstractReplicationStrategy to help smooth out the stampede effect when the \"global\" lock on the cached TM gets released after the fill? How much do you think it's worth the extra complexity? FWIW, my v2 patch suffers from this issue and it hasn't reared itself in production. The write load for the machines in the cluster I've been looking at is comparatively low though compared to many others at 6-7k/sec peak on an 8-core box.","created":"2013-11-20T20:49:03.708+0000"},{"body":"bq. It seems that unless I'm missing something either is possible with the current release code, and thus these patches as well\n\nTechnically correct, but in practice we're in pretty good shape. The sequence is:\n\n# Add the changing node to pending ranges\n# Sleep for RING_DELAY so everyone else starts including the new target in their writes\n# Flush data to be transferred\n# Send over data for writes that happened before (1)\n\nStep 1 happens on every coordinator. 2-4 only happen on the node that is giving up a token range.\n\nThe guarantee we need is that any write that happens before the pending range change, completes before the subsequent flush.\n\nEven if we used TM.lock to protect the entire ARS sequence (guaranteeing that no local write is in progress once the PRC happens) we could still receive writes from other nodes that began their PRC change later. \n\nSo we rely on the RING_DELAY (30s) sleep. I suppose a GC pause for instance at just the wrong time could theoretically mean a mutation against the old state gets sent out late, but I don't see how we can improve it.\n\nbq. IMHO to be defensive, any time the write lock is acquired in TokenMetadata, the version should be bumped in the finally block before the lock is released\n\nHaven't thought this through as much. What are you saying we should bump that we weren't calling invalidate on before?\n\nbq. Is the idea with the striped lock on the endpoint cache in AbstractReplicationStrategy to help smooth out the stampede effect when the \"global\" lock on the cached TM gets released after the fill?\n\nI'm trying to avoid a minor stampede on calculateNaturalEndpoints (CASSANDRA-3881) but it's probably premature optimization. v5 attached w/o that.","created":"2013-11-22T14:29:57.941+0000"},{"body":"Thanks for taking the time to explain the consistency story. It makes perfect sense. \n\nMy defensiveness comment suggested bumping the version number each time the TM write lock is released, which would be in addition to the existing invalidations. You're probably a much better gauge on the usefulness of this, so up to you.\n\nReally nice that the v5 patch is so compact. Two minor comments: the endpointsLock declaration is still in there, and not to be all nitpicky but there are two typos in the comments (\"wo we keep\" and \"clone got invalidted\").","created":"2013-11-26T17:33:23.351+0000"},{"body":"bq. My defensiveness comment suggested bumping the version number each time the TM write lock is released, which would be in addition to the existing invalidations.\n\nOkay. I'm going to leave this be then, because I don't want to accidentally start invalidating the cache unnecessarily because one of those operations was more common than I thought. Could address in trunk if you want to open a ticket.\n\nCommitted v5 w/ nits fixed.","created":"2013-11-26T20:11:15.526+0000"},{"body":"LGTM!","created":"2013-11-26T22:26:43.416+0000"},{"body":"{noformat}\nprivate volatile long ringVersion = 0;\n\nringVersion++;\n{noformat}\n\nIf there is something tricky here that makes an increment on a volatile okay then it deserves a comment.","created":"2013-12-09T21:24:10.206+0000"},{"body":"We don't care about keeping an accurate count, only that once it's done with that block it's higher than it was before.","created":"2013-12-09T23:50:02.480+0000"}],"conversations":[{"body":"We've observed that events which cause invalidation of the endpoint cache (update keyspace, add/remove nodes, etc) in AbstractReplicationStrategy result in several seconds of thundering herd behavior on the entire cluster. \n\nA thread dump shows over a hundred threads (I stopped counting at that point) with a backtrace like this:\n\n at java.net.Inet4Address.getAddress(Inet4Address.java:288)\n at org.apache.cassandra.locator.TokenMetadata$1.compare(TokenMetadata.java:106)\n at org.apache.cassandra.locator.TokenMetadata$1.compare(TokenMetadata.java:103)\n at java.util.TreeMap.getEntryUsingComparator(TreeMap.java:351)\n at java.util.TreeMap.getEntry(TreeMap.java:322)\n at java.util.TreeMap.get(TreeMap.java:255)\n at com.google.common.collect.AbstractMultimap.put(AbstractMultimap.java:200)\n at com.google.common.collect.AbstractSetMultimap.put(AbstractSetMultimap.java:117)\n at com.google.common.collect.TreeMultimap.put(TreeMultimap.java:74)\n at com.google.common.collect.AbstractMultimap.putAll(AbstractMultimap.java:273)\n at com.google.common.collect.TreeMultimap.putAll(TreeMultimap.java:74)\n at org.apache.cassandra.utils.SortedBiMultiValMap.create(SortedBiMultiValMap.java:60)\n at org.apache.cassandra.locator.TokenMetadata.cloneOnlyTokenMap(TokenMetadata.java:598)\n at org.apache.cassandra.locator.AbstractReplicationStrategy.getNaturalEndpoints(AbstractReplicationStrategy.java:104)\n at org.apache.cassandra.service.StorageService.getNaturalEndpoints(StorageService.java:2671)\n at org.apache.cassandra.service.StorageProxy.performWrite(StorageProxy.java:375)\n\nIt looks like there's a large amount of cost in the TokenMetadata.cloneOnlyTokenMap that AbstractReplicationStrategy.getNaturalEndpoints is calling each time there is a cache miss for an endpoint. It seems as if this would only impact clusters with large numbers of tokens, so it's probably a vnodes-only issue.\n\nProposal: In AbstractReplicationStrategy.getNaturalEndpoints(), cache the cloned TokenMetadata instance returned by TokenMetadata.cloneOnlyTokenMap(), wrapping it with a lock to prevent stampedes, and clearing it in clearEndpointCache(). Thoughts?","from":"reporter","subject":"Endpoint cache invalidation causes CPU spike (on vnode rings?)"},{"body":"Interesting. Could we optimize cOTM instead? That would definitely be the simplest solution.\n\nE.g. it looks like TreeMultimap.putAll actually loops over each entry and calls put one at a time which is the worst-case scenario for binary tree rebalancing -- quadratic time.","from":"developer"},{"body":"bq. it looks like TreeMultimap.putAll actually loops over each entry and calls put one at a time\n\nI don't see a way around this: https://code.google.com/p/guava-libraries/issues/detail?id=1579","from":"developer"},{"body":"A Treap? Can be cheaply built, cheaply merged and cheaply cloned.\n\nAlso, anything cheaply cloneable would work for that operation. A SnapTree that is wrapped to support multi-map functionality would also work.","from":"developer"},{"body":"Yes. Too bad the implementation classes like AbstractSortedKeySortedSetMultimap are package-private.","from":"developer"},{"body":"Actually I think a STM multimap would still get messy quickly since you need to do a \"deep\" clone -- cloning the top-level Map would leave the values (sub-collections) sharing a reference.","from":"developer"},{"body":"Just have the wrapper make any updates to a collection replace the collection instead of modifying it.\n\n[NB: I haven't looked to see if this would have any negative performance implications on the update side, I'm assuming the reads are more frequent and/or collections small... if not, a Treap is probably the better choice as my old Jjoost implementation (IIRC) supports snapshotting and multiple values are dealt with inside the tree itself, not as a collection]","from":"developer"},{"body":"bq. Proposal: In AbstractReplicationStrategy.getNaturalEndpoints(), cache the cloned TokenMetadata instance returned by TokenMetadata.cloneOnlyTokenMap(), wrapping it with a lock to prevent stampedes, and clearing it in clearEndpointCache().\n\nWhy not just use a sharded lock to prevent stampedes directly w/o the caching complexity, as in the attached?","from":"developer"},{"body":"I see, with vnodes we have enough ranges that we can have a thundering herd even if each range only clones once.\n\nv2 attached with the approach you described originally.","from":"developer"},{"body":"Attached a patch we deployed to production that fixed the issue.","from":"developer"},{"body":"CPU user% graph during the rollout of the patch I attached on 1 DC (15 nodes) of the cluster. Around ~21:05 the patch starts to roll out and spikes are seen. The node in question receives the patch at ~21:30, and afterwards the spikes are gone. The rollout finishes at ~21:45.","from":"developer"},{"body":"I have to admit I like it better without the custom wrapper class. :)","from":"developer"},{"body":"Well, I started writing the patch this morning and I don't write multi-threaded Java code every day, so I'm overly careful ;) The only theoretical advantage to my patch is that it allows concurrent readers.","from":"developer"},{"body":"Unfortunately both of the patches suffer from a deadlock, since the invalidation and fill are wrapped up in TokenMetadata's locks.\n\nT1 acquires cache read lock\nT2 acquires TokenMetadata write lock\nT1 acquires cache write lock on miss\nT2 is blocked on cache write lock trying to invalidate\nT1 is blocked on TokenMetadata read lock trying to cloneOnlyTokenMap to fill the cache\n\nTrying to work on a fix.","from":"developer"},{"body":"Attached a new patch with the deadlock fixed. We're running this on a production cluster.\n\nThe primary issue was the callback for invalidation from TokenMetadata to all of the registered AbstractReplicationStrategy instances. This was asking for it anyway, so in the patch I replaced the \"push\" invalidation with simple versioning of the TokenMetadata endpoints. TokenMetadata bumps it's version number each time the cache would need to be invalidated, and AbstractReplicationStrategy checks it's version when it needs to do a read, invalidating if necessary. This gets the invalidation out of the gossip threads and into the RPC threads, which is probably a good thing. The only thing I'm not super crazy about is the extra hot path read lock acquisition on TokenMetadata.getEndpointVersion(), which might be avoidable.","from":"developer"},{"body":"I think we can craft a simpler solution (v3) by using an AtomicReference to the TM clone. This removes the possibility of deadlock since clearEndpointCache now only makes non-blocking calls.\n\nI've also refined it to use a Striped per-keyToken, as well as synchronizing the TM clone itself, since concurrent endpoint computation is fine.","from":"developer"},{"body":"I like the simpler approach. I still think the callbacks for invalidation are asking for it ;) I also think perhaps the stampede lock should be more explicit than a synchronized lock on \"this\" to prevent unintended blocking from future modifications.\n\nEither way, I think the only material concern I have is the order that TokenMetadata changes get applied to the caches in AbstractReplicationStrategy instances. Shouldn't the invalidation take place on all threads in all instances of AbstractReplicationStrategy before returning from an endpoint-mutating write operation in TokenMetadata? It seems as if just setting the cache to empty would allow a period of time where TokenMetadata write methods had returned but not all threads have seen the mutation yet because they are still holding onto the old clone of TM. This might be alright though, I'm not sure. Thoughts?","from":"developer"},{"body":"bq. It seems as if just setting the cache to empty would allow a period of time where TokenMetadata write methods had returned but not all threads have seen the mutation yet\n\nI'm not 100% sure this is what you're talking about, but I see this problem with the existing code (and my v3):\n\n{noformat}\nThread 1 Thread 2 \ngetNaturalEndpoints \ncloneOnlyTokenMap \n invalidateCachedTokenEndpointValues\nendpoints = calculate\ncacheEndpoint [based on the now-invalidated token map]\n{noformat}\n\nSo it doesn't quite work. We'd need to introduce another AtomicReference on the cache, so that invalidate could create a new Map (so it doesn't matter if someone updates the old one). But I think you're right that getting rid of the callback approach entirely is better.","from":"developer"},{"body":"v4 attached that uses a versioning approach like yours. I dropped the readLock acquire on version read since it's not necessary to block callers during the update. (A few extra over-broad replica set operations won't hurt.)","from":"developer"},{"body":"+100 at removing those pub/sub callbacks :)\n\nThe concurrency issues I bring up are probably because I'm unfamiliar with the \"guarantees\" needed by TokenMetadata updates. It looks like the current release code is subject to the issue I brought up, where method calls on TokenMetadata that change state return successfully before all threads applying mutations have \"seen\" the update. There will be some mutations in progress that are using \"stale\" token data to apply writes even after TokenMetadata write methods returns as successful. So this does not appear to be a regression, but I'm just being overly cautious having been burned by these sort of double-caching scenarios before. You bring up the point that over-broad operations are ok, and I agree, but I'm more concerned about operations that are too narrow. It seems that unless I'm missing something either is possible with the current release code, and thus these patches as well (including mine).\n\nTokenMetadata#updateNormalTokens is (implicitly) relying on the removeFromMoving call to bump the version, but the tokenToEndpointMap is updated afterwards, which means internal data is updated after the version is bumped. IMHO to be defensive, any time the write lock is acquired in TokenMetadata, the version should be bumped in the finally block before the lock is released. I don't think this is exposing a bug in the existing patch though, because cloneOnlyTokenMap will be blocked until the write lock is released in the finally block.\n\nIs the idea with the striped lock on the endpoint cache in AbstractReplicationStrategy to help smooth out the stampede effect when the \"global\" lock on the cached TM gets released after the fill? How much do you think it's worth the extra complexity? FWIW, my v2 patch suffers from this issue and it hasn't reared itself in production. The write load for the machines in the cluster I've been looking at is comparatively low though compared to many others at 6-7k/sec peak on an 8-core box.","from":"developer"},{"body":"bq. It seems that unless I'm missing something either is possible with the current release code, and thus these patches as well\n\nTechnically correct, but in practice we're in pretty good shape. The sequence is:\n\n# Add the changing node to pending ranges\n# Sleep for RING_DELAY so everyone else starts including the new target in their writes\n# Flush data to be transferred\n# Send over data for writes that happened before (1)\n\nStep 1 happens on every coordinator. 2-4 only happen on the node that is giving up a token range.\n\nThe guarantee we need is that any write that happens before the pending range change, completes before the subsequent flush.\n\nEven if we used TM.lock to protect the entire ARS sequence (guaranteeing that no local write is in progress once the PRC happens) we could still receive writes from other nodes that began their PRC change later. \n\nSo we rely on the RING_DELAY (30s) sleep. I suppose a GC pause for instance at just the wrong time could theoretically mean a mutation against the old state gets sent out late, but I don't see how we can improve it.\n\nbq. IMHO to be defensive, any time the write lock is acquired in TokenMetadata, the version should be bumped in the finally block before the lock is released\n\nHaven't thought this through as much. What are you saying we should bump that we weren't calling invalidate on before?\n\nbq. Is the idea with the striped lock on the endpoint cache in AbstractReplicationStrategy to help smooth out the stampede effect when the \"global\" lock on the cached TM gets released after the fill?\n\nI'm trying to avoid a minor stampede on calculateNaturalEndpoints (CASSANDRA-3881) but it's probably premature optimization. v5 attached w/o that.","from":"developer"},{"body":"Thanks for taking the time to explain the consistency story. It makes perfect sense. \n\nMy defensiveness comment suggested bumping the version number each time the TM write lock is released, which would be in addition to the existing invalidations. You're probably a much better gauge on the usefulness of this, so up to you.\n\nReally nice that the v5 patch is so compact. Two minor comments: the endpointsLock declaration is still in there, and not to be all nitpicky but there are two typos in the comments (\"wo we keep\" and \"clone got invalidted\").","from":"developer"},{"body":"bq. My defensiveness comment suggested bumping the version number each time the TM write lock is released, which would be in addition to the existing invalidations.\n\nOkay. I'm going to leave this be then, because I don't want to accidentally start invalidating the cache unnecessarily because one of those operations was more common than I thought. Could address in trunk if you want to open a ticket.\n\nCommitted v5 w/ nits fixed.","from":"developer"},{"body":"LGTM!","from":"developer"},{"body":"{noformat}\nprivate volatile long ringVersion = 0;\n\nringVersion++;\n{noformat}\n\nIf there is something tricky here that makes an increment on a volatile okay then it deserves a comment.","from":"developer"},{"body":"We don't care about keeping an accurate count, only that once it's done with that block it's higher than it was before.","from":"developer"}],"created":"2013-11-13T16:06:06.000+0000","description":"We've observed that events which cause invalidation of the endpoint cache (update keyspace, add/remove nodes, etc) in AbstractReplicationStrategy result in several seconds of thundering herd behavior on the entire cluster. \n\nA thread dump shows over a hundred threads (I stopped counting at that point) with a backtrace like this:\n\n at java.net.Inet4Address.getAddress(Inet4Address.java:288)\n at org.apache.cassandra.locator.TokenMetadata$1.compare(TokenMetadata.java:106)\n at org.apache.cassandra.locator.TokenMetadata$1.compare(TokenMetadata.java:103)\n at java.util.TreeMap.getEntryUsingComparator(TreeMap.java:351)\n at java.util.TreeMap.getEntry(TreeMap.java:322)\n at java.util.TreeMap.get(TreeMap.java:255)\n at com.google.common.collect.AbstractMultimap.put(AbstractMultimap.java:200)\n at com.google.common.collect.AbstractSetMultimap.put(AbstractSetMultimap.java:117)\n at com.google.common.collect.TreeMultimap.put(TreeMultimap.java:74)\n at com.google.common.collect.AbstractMultimap.putAll(AbstractMultimap.java:273)\n at com.google.common.collect.TreeMultimap.putAll(TreeMultimap.java:74)\n at org.apache.cassandra.utils.SortedBiMultiValMap.create(SortedBiMultiValMap.java:60)\n at org.apache.cassandra.locator.TokenMetadata.cloneOnlyTokenMap(TokenMetadata.java:598)\n at org.apache.cassandra.locator.AbstractReplicationStrategy.getNaturalEndpoints(AbstractReplicationStrategy.java:104)\n at org.apache.cassandra.service.StorageService.getNaturalEndpoints(StorageService.java:2671)\n at org.apache.cassandra.service.StorageProxy.performWrite(StorageProxy.java:375)\n\nIt looks like there's a large amount of cost in the TokenMetadata.cloneOnlyTokenMap that AbstractReplicationStrategy.getNaturalEndpoints is calling each time there is a cache miss for an endpoint. It seems as if this would only impact clusters with large numbers of tokens, so it's probably a vnodes-only issue.\n\nProposal: In AbstractReplicationStrategy.getNaturalEndpoints(), cache the cloned TokenMetadata instance returned by TokenMetadata.cloneOnlyTokenMap(), wrapping it with a lock to prevent stampedes, and clearing it in clearEndpointCache(). Thoughts?","issue_id":"12679044","key":"CASSANDRA-6345","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-11-26T20:11:15.000+0000","role":"fixed_distractor","summary":"Endpoint cache invalidation causes CPU spike (on vnode rings?)"} {"case_id":"12682745","cluster":"DISTRACTOR-CASSANDRA-6449","comments":[{"body":"From a customer:\n\nThe culprit is: / src / java / org / apache / cassandra / utils / FBUtilities.java\n\nFile historyDir = new File(System.getProperty(\"user.home\"), \".cassandra\");\n\nSetting an alternate environment variable HOME doesn't fix. I've tried patching the nodetool wrapper script to provide -Duser.home at runtime, but it seems when defining user.home, I get runtime errors with missing libraries. It would be nice if the tool just honoured $HOME (or let you specify a commandline override without hacking the script).","created":"2014-01-29T17:53:02.160+0000"},{"body":"This is the error that occurs when manually defining -Duser.home in the nodetool shell script:\n\n{code}\nException in thread \"main\" java.lang.NoClassDefFoundError: com/google/common/collect/AbstractMultimap$WrappedSortedSet \nat com.google.common.collect.AbstractMultimap.wrapCollection(AbstractMultimap.java:374) \nat com.google.common.collect.AbstractMultimap.get(AbstractMultimap.java:363) \nat com.google.common.collect.AbstractSetMultimap.get(AbstractSetMultimap.java:59) \nat com.google.common.collect.AbstractSortedSetMultimap.get(AbstractSortedSetMultimap.java:65) \nat com.google.common.collect.TreeMultimap.get(TreeMultimap.java:74) \nat com.google.common.collect.AbstractSortedSetMultimap.get(AbstractSortedSetMultimap.java:35) \nat com.google.common.collect.Multimaps$UnmodifiableMultimap.get(Multimaps.java:563) \nat org.apache.cassandra.locator.TokenMetadata.getTokens(TokenMetadata.java:507) \nat org.apache.cassandra.service.StorageService.getTokens(StorageService.java:2048) \nat org.apache.cassandra.service.StorageService.getTokens(StorageService.java:2042) \nat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method) \nat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39) \nat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25) \nat java.lang.reflect.Method.invoke(Method.java:597) \nat com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:93) \nat com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:27) \nat com.sun.jmx.mbeanserver.MBeanIntrospector.invokeM(MBeanIntrospector.java:208) \nat com.sun.jmx.mbeanserver.PerInterface.invoke(PerInterface.java:120) \nat com.sun.jmx.mbeanserver.MBeanSupport.invoke(MBeanSupport.java:264) \nat com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.invoke(DefaultMBeanServerInterceptor.java:836) \nat com.sun.jmx.mbeanserver.JmxMBeanServer.invoke(JmxMBeanServer.java:762) \nat javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1454) \nat javax.management.remote.rmi.RMIConnectionImpl.access$300(RMIConnectionImpl.java:74) \nat javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1295) \nat javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1387) \nat javax.management.remote.rmi.RMIConnectionImpl.invoke(RMIConnectionImpl.java:818) \nat sun.reflect.GeneratedMethodAccessor32.invoke(Unknown Source) \nat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25) \nat java.lang.reflect.Method.invoke(Method.java:597) \nat sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:303) \nat sun.rmi.transport.Transport$1.run(Transport.java:159) \nat java.security.AccessController.doPrivileged(Native Method) \nat sun.rmi.transport.Transport.serviceCall(Transport.java:155) \nat sun.rmi.transport.tcp.TCPTransport.handleMessages(TCPTransport.java:535) \nat sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run0(TCPTransport.java:790) \nat sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run(TCPTransport.java:649) \nat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:895) \nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:918) \nat java.lang.Thread.run(Thread.java:662)\n{code}","created":"2014-01-29T18:15:31.818+0000"},{"body":"NodeCmd honestly tries to ignore IOExceptions, the problem is that {{FBUtilities.getToolsOutputDirectory}} wraps IOExceptions in {{FSWriteError}},\n\n{code:title=org.apache.cassandra.tools.NodeCmd.printHistory}\n FileWriter writer = null;\n try\n {\n final String outputDir = FBUtilities.getToolsOutputDirectory().getCanonicalPath();\n.....\n }\n catch (IOException ioe)\n {\n //quietly ignore any errors about not being able to write out history\n }\n finally\n {\n FileUtils.closeQuietly(writer);\n }\n{code}","created":"2014-01-29T19:25:02.699+0000"},{"body":"I've had another user run into this issue who would like this functionality. I can reproduce the problem just by restricting write access to the users home dir that is running nodetool.","created":"2014-04-24T16:17:13.301+0000"},{"body":"Simply catching IOError in addition to IOException.","created":"2014-05-23T17:12:01.398+0000"},{"body":"We probably want to fix this in 1.2 or at least 2.0","created":"2014-06-18T19:05:25.329+0000"},{"body":"Patch applies to 2.0 if you point it at src/java/org/apache/cassandra/tools/NodeCmd.java.","created":"2014-06-18T19:10:40.413+0000"},{"body":"Good enough for me then, committed.","created":"2014-06-18T19:20:02.867+0000"}],"conversations":[{"body":"We shouldn't error out if we can't make the .cassandra folder for the new history stuff.\n\n{noformat}\nException in thread \"main\" FSWriteError in /usr/share/opscenter-agent/.cassandra\n\tat org.apache.cassandra.io.util.FileUtils.createDirectory(FileUtils.java:261)\n\tat org.apache.cassandra.utils.FBUtilities.getToolsOutputDirectory(FBUtilities.java:627)\n\tat org.apache.cassandra.tools.NodeCmd.printHistory(NodeCmd.java:1403)\n\tat org.apache.cassandra.tools.NodeCmd.main(NodeCmd.java:1122)\nCaused by: java.io.IOException: Failed to mkdirs /usr/share/opscenter-agent/.cassandra\n\t... 4 more\n{noformat}","from":"reporter","subject":"Tools error out if they can't make ~/.cassandra"},{"body":"From a customer:\n\nThe culprit is: / src / java / org / apache / cassandra / utils / FBUtilities.java\n\nFile historyDir = new File(System.getProperty(\"user.home\"), \".cassandra\");\n\nSetting an alternate environment variable HOME doesn't fix. I've tried patching the nodetool wrapper script to provide -Duser.home at runtime, but it seems when defining user.home, I get runtime errors with missing libraries. It would be nice if the tool just honoured $HOME (or let you specify a commandline override without hacking the script).","from":"developer"},{"body":"This is the error that occurs when manually defining -Duser.home in the nodetool shell script:\n\n{code}\nException in thread \"main\" java.lang.NoClassDefFoundError: com/google/common/collect/AbstractMultimap$WrappedSortedSet \nat com.google.common.collect.AbstractMultimap.wrapCollection(AbstractMultimap.java:374) \nat com.google.common.collect.AbstractMultimap.get(AbstractMultimap.java:363) \nat com.google.common.collect.AbstractSetMultimap.get(AbstractSetMultimap.java:59) \nat com.google.common.collect.AbstractSortedSetMultimap.get(AbstractSortedSetMultimap.java:65) \nat com.google.common.collect.TreeMultimap.get(TreeMultimap.java:74) \nat com.google.common.collect.AbstractSortedSetMultimap.get(AbstractSortedSetMultimap.java:35) \nat com.google.common.collect.Multimaps$UnmodifiableMultimap.get(Multimaps.java:563) \nat org.apache.cassandra.locator.TokenMetadata.getTokens(TokenMetadata.java:507) \nat org.apache.cassandra.service.StorageService.getTokens(StorageService.java:2048) \nat org.apache.cassandra.service.StorageService.getTokens(StorageService.java:2042) \nat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method) \nat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39) \nat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25) \nat java.lang.reflect.Method.invoke(Method.java:597) \nat com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:93) \nat com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:27) \nat com.sun.jmx.mbeanserver.MBeanIntrospector.invokeM(MBeanIntrospector.java:208) \nat com.sun.jmx.mbeanserver.PerInterface.invoke(PerInterface.java:120) \nat com.sun.jmx.mbeanserver.MBeanSupport.invoke(MBeanSupport.java:264) \nat com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.invoke(DefaultMBeanServerInterceptor.java:836) \nat com.sun.jmx.mbeanserver.JmxMBeanServer.invoke(JmxMBeanServer.java:762) \nat javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1454) \nat javax.management.remote.rmi.RMIConnectionImpl.access$300(RMIConnectionImpl.java:74) \nat javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1295) \nat javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1387) \nat javax.management.remote.rmi.RMIConnectionImpl.invoke(RMIConnectionImpl.java:818) \nat sun.reflect.GeneratedMethodAccessor32.invoke(Unknown Source) \nat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25) \nat java.lang.reflect.Method.invoke(Method.java:597) \nat sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:303) \nat sun.rmi.transport.Transport$1.run(Transport.java:159) \nat java.security.AccessController.doPrivileged(Native Method) \nat sun.rmi.transport.Transport.serviceCall(Transport.java:155) \nat sun.rmi.transport.tcp.TCPTransport.handleMessages(TCPTransport.java:535) \nat sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run0(TCPTransport.java:790) \nat sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run(TCPTransport.java:649) \nat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:895) \nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:918) \nat java.lang.Thread.run(Thread.java:662)\n{code}","from":"developer"},{"body":"NodeCmd honestly tries to ignore IOExceptions, the problem is that {{FBUtilities.getToolsOutputDirectory}} wraps IOExceptions in {{FSWriteError}},\n\n{code:title=org.apache.cassandra.tools.NodeCmd.printHistory}\n FileWriter writer = null;\n try\n {\n final String outputDir = FBUtilities.getToolsOutputDirectory().getCanonicalPath();\n.....\n }\n catch (IOException ioe)\n {\n //quietly ignore any errors about not being able to write out history\n }\n finally\n {\n FileUtils.closeQuietly(writer);\n }\n{code}","from":"developer"},{"body":"I've had another user run into this issue who would like this functionality. I can reproduce the problem just by restricting write access to the users home dir that is running nodetool.","from":"developer"},{"body":"Simply catching IOError in addition to IOException.","from":"developer"},{"body":"We probably want to fix this in 1.2 or at least 2.0","from":"developer"},{"body":"Patch applies to 2.0 if you point it at src/java/org/apache/cassandra/tools/NodeCmd.java.","from":"developer"},{"body":"Good enough for me then, committed.","from":"developer"}],"created":"2013-12-04T18:53:40.000+0000","description":"We shouldn't error out if we can't make the .cassandra folder for the new history stuff.\n\n{noformat}\nException in thread \"main\" FSWriteError in /usr/share/opscenter-agent/.cassandra\n\tat org.apache.cassandra.io.util.FileUtils.createDirectory(FileUtils.java:261)\n\tat org.apache.cassandra.utils.FBUtilities.getToolsOutputDirectory(FBUtilities.java:627)\n\tat org.apache.cassandra.tools.NodeCmd.printHistory(NodeCmd.java:1403)\n\tat org.apache.cassandra.tools.NodeCmd.main(NodeCmd.java:1122)\nCaused by: java.io.IOException: Failed to mkdirs /usr/share/opscenter-agent/.cassandra\n\t... 4 more\n{noformat}","issue_id":"12682745","key":"CASSANDRA-6449","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-06-18T19:20:02.000+0000","role":"fixed_distractor","summary":"Tools error out if they can't make ~/.cassandra"} {"case_id":"12683839","cluster":"DISTRACTOR-CASSANDRA-6470","comments":[{"body":"Can you reproduce, Ryan?","created":"2013-12-12T19:32:31.297+0000"},{"body":"[~enrico.scalavino] What version of Cassandra and datastax driver are you using? I'll try and recreate this, but if you have a test already written that is easily decoupled from your project can you post that too?","created":"2013-12-12T20:19:19.106+0000"},{"body":"[~enrico.scalavino] are you using version 2.0.X of C* and the driver? The limit being a bind parameter isn't supported unless you are. It may be confusing the driver if you are using the 1.0.X version of the driver (not sure what error gets thrown if you do that).","created":"2013-12-12T22:38:36.559+0000"},{"body":"I get the same error. I dont know when it has been started. I'm using Cassandra 2.0.2 and Datastax Java Driver 2.0.0-beta2. Query works in cqlsh but fail when running in the client. I tried to re-create (DROP/CREATE) the column family, but the error stills.\n\n=============\nTable layout:\n\ncqlsh:pollkan> desc table observed;\n\nCREATE TABLE observed (\n observed timeuuid,\n observer timeuuid,\n blocked boolean,\n PRIMARY KEY (observed, observer)\n) WITH\n bloom_filter_fp_chance=0.010000 AND\n caching='KEYS_ONLY' AND\n comment='' AND\n dclocal_read_repair_chance=0.000000 AND\n gc_grace_seconds=864000 AND\n index_interval=128 AND\n read_repair_chance=0.100000 AND\n replicate_on_write='true' AND\n populate_io_cache_on_flush='false' AND\n default_time_to_live=0 AND\n speculative_retry='99.0PERCENTILE' AND\n memtable_flush_period_in_ms=0 AND\n compaction={'class': 'SizeTieredCompactionStrategy'} AND\n compression={'sstable_compression': 'LZ4Compressor'};\n\nCREATE INDEX observedBlocked ON observed (blocked);\n\n=============\nQuery in the cqlsh:\n\ncqlsh:pollkan> SELECT observer FROM observed WHERE observed = fa93c210-4bff-11e3-b48f-5714d8c6f3b2 AND observer > 00000000-0000-1000-0000-000000000000 and blocked = false LIMIT 10000;\n\n observer\n--------------------------------------\n 43814f60-5bb1-11e3-97c8-ad396a9e8180\n\n(1 rows)\n\n=============\nQuery in the client log:\n\n2013-12-13/00:53:03.039/BRST [timeline_1] DEBUG br.com.pollkan.batch.CqlCommands Execute query [SELECT observer FROM observed WHERE observed = ? AND observer > ? and blocked = ? LIMIT 10000;] arguments [[fa93c210-4bff-11e3-b48f-5714d8c6f3b2][00000000-0000-1000-0000-000000000000][false]]\n\n=============\nError in cassandra:\n\nERROR [ReadStage:52] 2013-12-13 01:04:56,799 CassandraDaemon.java (line 187) Exception in thread Thread[ReadStage:52,5,main]\njava.lang.RuntimeException: java.lang.ArrayIndexOutOfBoundsException: 0\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1931)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:724)\nCaused by: java.lang.ArrayIndexOutOfBoundsException: 0\n at org.apache.cassandra.db.filter.SliceQueryFilter.start(SliceQueryFilter.java:261)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.makePrefix(CompositesSearcher.java:66)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.getIndexedIterator(CompositesSearcher.java:101)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.search(CompositesSearcher.java:53)\n at org.apache.cassandra.db.index.SecondaryIndexManager.search(SecondaryIndexManager.java:537)\n at org.apache.cassandra.db.ColumnFamilyStore.search(ColumnFamilyStore.java:1649)\n at org.apache.cassandra.db.PagedRangeCommand.executeLocally(PagedRangeCommand.java:109)\n at org.apache.cassandra.service.StorageProxy$LocalRangeSliceRunnable.runMayThrow(StorageProxy.java:1414)\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1927)\n ... 3 more\n\n=============\nError from driver log:\n\n2013-12-13/01:05:06.798/BRST [timeline_1] ERROR br.com.pollkan.batch.CqlCommands Exception! [Cassandra timeout during read query at consistency ONE (1 responses were required but only 0 replica responded)]\ncom.datastax.driver.core.exceptions.ReadTimeoutException: Cassandra timeout during read query at consistency ONE (1 responses were required but only 0 replica responded)\n at com.datastax.driver.core.exceptions.ReadTimeoutException.copy(ReadTimeoutException.java:69)\n at com.datastax.driver.core.ResultSetFuture.extractCauseFromExecutionException(ResultSetFuture.java:271)\n at com.datastax.driver.core.ResultSetFuture.getUninterruptibly(ResultSetFuture.java:187)\n at com.datastax.driver.core.Session.execute(Session.java:126)\n at br.com.pollkan.batch.CqlCommands.executeQuery(CqlCommands.java:149)\n at br.com.pollkan.batch.BaseBatch.processChild(BaseBatch.java:364)\n at br.com.pollkan.batch.BaseBatch.run(BaseBatch.java:640)\n at java.lang.Thread.run(Thread.java:722)\n\nIf need more information, please let me know. Tks","created":"2013-12-13T03:09:14.442+0000"},{"body":"[~jjordan] [~enigmacurry] I am using datastax 2.0.0-rc1 and Cassandra 2.0.3. The query works if I remove the where clause on either the read attribute or on the message_id attribute. \nUnfortunately the code calling the query is buried deep in the project, and the tests only test high level apis. ","created":"2013-12-13T11:22:46.875+0000"},{"body":"I get the same error too. I'm using Cassandra 2.0.2 and Datastax Java Driver 2.0.0-beta2","created":"2013-12-13T13:02:03.674+0000"},{"body":"I removed the \"blocked\" column from the query (indexed column) and now it works. This helps?","created":"2013-12-18T23:56:21.350+0000"},{"body":"No, also my query works if I remove one of the conditions. That still does not explain why that happens, and why there is an unmanaged exception on the server. ","created":"2013-12-19T10:47:33.416+0000"},{"body":"I uploaded an Eclipse project (6470-reproduced.tar.gz) that reproduces this issue.\n\nStart up a ccm cluster: \n\n{code}\nccm create -v git:cassandra-2.0.2\nccm create -v git:cassandra-2.0.2 test\nccm populate -n 3:3\nccm start\n{code}\n\nThen run Bug6470Client.\n\nThis code actually works with driver version 1.0.3, so I wonder if that's not the cause instead of C*...\n\nI can also confirm that the same queries issued in cqlsh do not repro this issue.","created":"2014-01-28T23:16:55.294+0000"},{"body":"I also tried from the latest python driver, no problem there.","created":"2014-01-28T23:24:14.908+0000"},{"body":"This is likely a problem with the paging over 2ndary indexes (which is why only the 2.0 version of the driver is running into it). I'll have a closer look.","created":"2014-01-29T09:07:17.327+0000"},{"body":"We were not handling empty bounds in DataRange.sliceForKey() (that is indeed used by paging calls) which was returning an empty slice array (which was incorrect but hence the error).\n\nAttached simple fix.","created":"2014-01-29T14:50:23.471+0000"},{"body":"+1","created":"2014-01-29T17:50:42.134+0000"},{"body":"Committed, thanks","created":"2014-01-29T18:25:44.664+0000"}],"conversations":[{"body":"schema: \n{noformat}\nCREATE TABLE inboxkeyspace.inboxes(user_id bigint, message_id bigint, thread_id bigint, network_id bigint, read boolean, PRIMARY KEY(user_id, message_id)) WITH CLUSTERING ORDER BY (message_id DESC);\nCREATE INDEX ON inboxkeyspace.inboxes(read);\n{noformat}\n\nquery: \n{noformat}\nSELECT thread_id, message_id, network_id FROM inboxkeyspace.inboxes WHERE user_id = ? AND message_id < ? AND read = ? LIMIT ? \n{noformat}\n\nThe query works if run via cqlsh. However, when run through the datastax client, on the client side we get a timeout exception and on the server side, the Cassandra log shows this exception: \n\n{noformat}\nERROR [ReadStage:4190] 2013-12-10 13:18:03,579 CassandraDaemon.java (line 187) Exception in thread Thread[ReadStage:4190,5,main]\njava.lang.RuntimeException: java.lang.ArrayIndexOutOfBoundsException: 0\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1940)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:722)\nCaused by: java.lang.ArrayIndexOutOfBoundsException: 0\n at org.apache.cassandra.db.filter.SliceQueryFilter.start(SliceQueryFilter.java:261)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.makePrefix(CompositesSearcher.java:66)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.getIndexedIterator(CompositesSearcher.java:101)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.search(CompositesSearcher.java:53)\n at org.apache.cassandra.db.index.SecondaryIndexManager.search(SecondaryIndexManager.java:537)\n at org.apache.cassandra.db.ColumnFamilyStore.search(ColumnFamilyStore.java:1669)\n at org.apache.cassandra.db.PagedRangeCommand.executeLocally(PagedRangeCommand.java:109)\n at org.apache.cassandra.service.StorageProxy$LocalRangeSliceRunnable.runMayThrow(StorageProxy.java:1423)\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1936)\n ... 3 more\n{noformat}","from":"reporter","subject":"ArrayIndexOutOfBoundsException on range query from client"},{"body":"Can you reproduce, Ryan?","from":"developer"},{"body":"[~enrico.scalavino] What version of Cassandra and datastax driver are you using? I'll try and recreate this, but if you have a test already written that is easily decoupled from your project can you post that too?","from":"developer"},{"body":"[~enrico.scalavino] are you using version 2.0.X of C* and the driver? The limit being a bind parameter isn't supported unless you are. It may be confusing the driver if you are using the 1.0.X version of the driver (not sure what error gets thrown if you do that).","from":"developer"},{"body":"I get the same error. I dont know when it has been started. I'm using Cassandra 2.0.2 and Datastax Java Driver 2.0.0-beta2. Query works in cqlsh but fail when running in the client. I tried to re-create (DROP/CREATE) the column family, but the error stills.\n\n=============\nTable layout:\n\ncqlsh:pollkan> desc table observed;\n\nCREATE TABLE observed (\n observed timeuuid,\n observer timeuuid,\n blocked boolean,\n PRIMARY KEY (observed, observer)\n) WITH\n bloom_filter_fp_chance=0.010000 AND\n caching='KEYS_ONLY' AND\n comment='' AND\n dclocal_read_repair_chance=0.000000 AND\n gc_grace_seconds=864000 AND\n index_interval=128 AND\n read_repair_chance=0.100000 AND\n replicate_on_write='true' AND\n populate_io_cache_on_flush='false' AND\n default_time_to_live=0 AND\n speculative_retry='99.0PERCENTILE' AND\n memtable_flush_period_in_ms=0 AND\n compaction={'class': 'SizeTieredCompactionStrategy'} AND\n compression={'sstable_compression': 'LZ4Compressor'};\n\nCREATE INDEX observedBlocked ON observed (blocked);\n\n=============\nQuery in the cqlsh:\n\ncqlsh:pollkan> SELECT observer FROM observed WHERE observed = fa93c210-4bff-11e3-b48f-5714d8c6f3b2 AND observer > 00000000-0000-1000-0000-000000000000 and blocked = false LIMIT 10000;\n\n observer\n--------------------------------------\n 43814f60-5bb1-11e3-97c8-ad396a9e8180\n\n(1 rows)\n\n=============\nQuery in the client log:\n\n2013-12-13/00:53:03.039/BRST [timeline_1] DEBUG br.com.pollkan.batch.CqlCommands Execute query [SELECT observer FROM observed WHERE observed = ? AND observer > ? and blocked = ? LIMIT 10000;] arguments [[fa93c210-4bff-11e3-b48f-5714d8c6f3b2][00000000-0000-1000-0000-000000000000][false]]\n\n=============\nError in cassandra:\n\nERROR [ReadStage:52] 2013-12-13 01:04:56,799 CassandraDaemon.java (line 187) Exception in thread Thread[ReadStage:52,5,main]\njava.lang.RuntimeException: java.lang.ArrayIndexOutOfBoundsException: 0\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1931)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:724)\nCaused by: java.lang.ArrayIndexOutOfBoundsException: 0\n at org.apache.cassandra.db.filter.SliceQueryFilter.start(SliceQueryFilter.java:261)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.makePrefix(CompositesSearcher.java:66)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.getIndexedIterator(CompositesSearcher.java:101)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.search(CompositesSearcher.java:53)\n at org.apache.cassandra.db.index.SecondaryIndexManager.search(SecondaryIndexManager.java:537)\n at org.apache.cassandra.db.ColumnFamilyStore.search(ColumnFamilyStore.java:1649)\n at org.apache.cassandra.db.PagedRangeCommand.executeLocally(PagedRangeCommand.java:109)\n at org.apache.cassandra.service.StorageProxy$LocalRangeSliceRunnable.runMayThrow(StorageProxy.java:1414)\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1927)\n ... 3 more\n\n=============\nError from driver log:\n\n2013-12-13/01:05:06.798/BRST [timeline_1] ERROR br.com.pollkan.batch.CqlCommands Exception! [Cassandra timeout during read query at consistency ONE (1 responses were required but only 0 replica responded)]\ncom.datastax.driver.core.exceptions.ReadTimeoutException: Cassandra timeout during read query at consistency ONE (1 responses were required but only 0 replica responded)\n at com.datastax.driver.core.exceptions.ReadTimeoutException.copy(ReadTimeoutException.java:69)\n at com.datastax.driver.core.ResultSetFuture.extractCauseFromExecutionException(ResultSetFuture.java:271)\n at com.datastax.driver.core.ResultSetFuture.getUninterruptibly(ResultSetFuture.java:187)\n at com.datastax.driver.core.Session.execute(Session.java:126)\n at br.com.pollkan.batch.CqlCommands.executeQuery(CqlCommands.java:149)\n at br.com.pollkan.batch.BaseBatch.processChild(BaseBatch.java:364)\n at br.com.pollkan.batch.BaseBatch.run(BaseBatch.java:640)\n at java.lang.Thread.run(Thread.java:722)\n\nIf need more information, please let me know. Tks","from":"developer"},{"body":"[~jjordan] [~enigmacurry] I am using datastax 2.0.0-rc1 and Cassandra 2.0.3. The query works if I remove the where clause on either the read attribute or on the message_id attribute. \nUnfortunately the code calling the query is buried deep in the project, and the tests only test high level apis. ","from":"developer"},{"body":"I get the same error too. I'm using Cassandra 2.0.2 and Datastax Java Driver 2.0.0-beta2","from":"developer"},{"body":"I removed the \"blocked\" column from the query (indexed column) and now it works. This helps?","from":"developer"},{"body":"No, also my query works if I remove one of the conditions. That still does not explain why that happens, and why there is an unmanaged exception on the server. ","from":"developer"},{"body":"I uploaded an Eclipse project (6470-reproduced.tar.gz) that reproduces this issue.\n\nStart up a ccm cluster: \n\n{code}\nccm create -v git:cassandra-2.0.2\nccm create -v git:cassandra-2.0.2 test\nccm populate -n 3:3\nccm start\n{code}\n\nThen run Bug6470Client.\n\nThis code actually works with driver version 1.0.3, so I wonder if that's not the cause instead of C*...\n\nI can also confirm that the same queries issued in cqlsh do not repro this issue.","from":"developer"},{"body":"I also tried from the latest python driver, no problem there.","from":"developer"},{"body":"This is likely a problem with the paging over 2ndary indexes (which is why only the 2.0 version of the driver is running into it). I'll have a closer look.","from":"developer"},{"body":"We were not handling empty bounds in DataRange.sliceForKey() (that is indeed used by paging calls) which was returning an empty slice array (which was incorrect but hence the error).\n\nAttached simple fix.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed, thanks","from":"developer"}],"created":"2013-12-10T13:39:18.000+0000","description":"schema: \n{noformat}\nCREATE TABLE inboxkeyspace.inboxes(user_id bigint, message_id bigint, thread_id bigint, network_id bigint, read boolean, PRIMARY KEY(user_id, message_id)) WITH CLUSTERING ORDER BY (message_id DESC);\nCREATE INDEX ON inboxkeyspace.inboxes(read);\n{noformat}\n\nquery: \n{noformat}\nSELECT thread_id, message_id, network_id FROM inboxkeyspace.inboxes WHERE user_id = ? AND message_id < ? AND read = ? LIMIT ? \n{noformat}\n\nThe query works if run via cqlsh. However, when run through the datastax client, on the client side we get a timeout exception and on the server side, the Cassandra log shows this exception: \n\n{noformat}\nERROR [ReadStage:4190] 2013-12-10 13:18:03,579 CassandraDaemon.java (line 187) Exception in thread Thread[ReadStage:4190,5,main]\njava.lang.RuntimeException: java.lang.ArrayIndexOutOfBoundsException: 0\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1940)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:722)\nCaused by: java.lang.ArrayIndexOutOfBoundsException: 0\n at org.apache.cassandra.db.filter.SliceQueryFilter.start(SliceQueryFilter.java:261)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.makePrefix(CompositesSearcher.java:66)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.getIndexedIterator(CompositesSearcher.java:101)\n at org.apache.cassandra.db.index.composites.CompositesSearcher.search(CompositesSearcher.java:53)\n at org.apache.cassandra.db.index.SecondaryIndexManager.search(SecondaryIndexManager.java:537)\n at org.apache.cassandra.db.ColumnFamilyStore.search(ColumnFamilyStore.java:1669)\n at org.apache.cassandra.db.PagedRangeCommand.executeLocally(PagedRangeCommand.java:109)\n at org.apache.cassandra.service.StorageProxy$LocalRangeSliceRunnable.runMayThrow(StorageProxy.java:1423)\n at org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1936)\n ... 3 more\n{noformat}","issue_id":"12683839","key":"CASSANDRA-6470","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-01-29T18:25:44.000+0000","role":"fixed_distractor","summary":"ArrayIndexOutOfBoundsException on range query from client"} {"case_id":"12685433","cluster":"DISTRACTOR-CASSANDRA-6503","comments":[{"body":"One thing I was thinking that might help with this is if we could leave the sstables named \"-tmp\" until we are ready to make them become active. That way when you reboot, any files hanging around will get removed on restart, instead of becoming active.","created":"2013-12-18T18:53:14.733+0000"},{"body":"I don't think we should be messing with repair code in 1.2.14, not for something that it took 3 years for someone to run across. Suggest targetting 2.0.","created":"2013-12-18T22:06:01.414+0000"},{"body":"(Not actually sure this is still an issue in 2.0.x but it seems possible.)","created":"2013-12-19T23:10:43.156+0000"},{"body":"In 1.2 as well as 2.0, every time the node receives SSTable file, it starts writing as 'tmp', but at the end of receiving, it calls SSTableWriter#closeAndOpenReader that moves received file out of 'tmp'. At this point, SSTables are still not added to ColumnFamilyStore, though if we shutdown the node and restart, the node would recognize received files even though streaming session was not finished successfully.\n\nWe need little tweak to defer renaming of SSTable files to do so after receiving all streamed files.\n ","created":"2013-12-19T23:42:55.026+0000"},{"body":"Attached patch, 6503_c1.2-v1, defers the release of the sstables to the CFS until the session is complete. Note: that patch is only for 1.2.\n\nFor c* 2.0, I'd like [~yukim]'s advice. I have a WIP here: https://github.com/jasobrown/cassandra/tree/6503_c2.0. The problem I'm running into is that FileMessage.sstable is of type SSTableReader, which we need on the sender side, but on the receiver side we want SSTableWriter (if we are going to defer the release of the sstables. For hacking things up sake, I've just changed FileMessage.sstable to a plain SSTable and let the users do the casting - which is only in two places, one of which is the FileMessage.Serializer.serailize() method. Not very extensive, but perhaps a bit sloppy.\n\nYuki, do you think it's worthwhile to split up the FileMessage object into two classes like OutFileMessage (which has a SSTR) and InFileMessage (which has SSTW)? ","created":"2014-01-06T14:48:28.672+0000"},{"body":"[~jasobrown] What I'm thinking to do is to closeAndOpenReader without renaming from tmp as we receive, and rename them at once at the end. And for renaming multiple files at once we probably want some kind of lockfile(also resolve CASSANDRA-2900?).\n\nThough finalizing SSTable write in closeAndOpenReader takes some time, so completely defer finalize as you do may be a good idea. I think splitting FileMessage is better than casting.","created":"2014-01-06T16:32:16.107+0000"},{"body":"bq. we probably want some kind of lockfile\n\nInteresting. What problems do you see this solving (I'm probably missing something in my understanding)?\n\nbq. splitting FileMessage is better than casting\n\nYeah, I knew you were gonna say that :)","created":"2014-01-07T01:28:57.510+0000"},{"body":"[~yukim] I can see how modifying closeAndOpenReader() would help here, but there's one wrinkle (I think): if we defer renaming the files from closeAndOpenReader(), we would need to rename the *open* SSTR, and there's a note in FileUtils.renameWithConfirm():\n\n{code} // this is not FSWE because usually when we see it it's because we didn't close the file before renaming it,\n // and Windows is picky about that.{code}\n\nSo I'm not sure about the deferred rename of the SSTR, assuming we care about Windows.\n","created":"2014-01-07T15:01:32.792+0000"},{"body":"I think you are right. We don't want to close SSTR again for renaming.\nLet's stick with deferring closeAndOpenReader as you did.\nFor the patch, I think it would be nice to call SSTW.abort() to discard already received file when something bad happened.\n\nbq. Interesting. What problems do you see this solving (I'm probably missing something in my understanding)?\n\nIf the node goes down during we are doing closeAndOpenReader to received files, there is a chance we have already renamed files.\nSo I wanted to make sure the node won't read those files when the node come up.\n","created":"2014-01-10T18:53:51.697+0000"},{"body":"bq. call SSTW.abort() ...\n\nMakes sense, will add that in.\n\nbq. If the node goes down during we are doing closeAndOpenReader to received files, there is a chance we have already renamed files\n\nOk, got it. Will come up with an idea and drop into the next patch. Thanks","created":"2014-01-10T19:30:12.817+0000"},{"body":"What is the \"since\" version for this ticket?\n\nIt seems rather serious that users were still exposed to zombie data even if following repair best practice, if a repair session ever stalled... especially because repair sessions tend to stall. In a way we are fortunate that the only way to restart a failed repair session is to restart the nodes (CASSANDRA-3486) because this workaround has protected some users from this bug.\n\n{quote}I don't think we should be messing with repair code in 1.2.14, not for something that it took 3 years for someone to run across. Suggest targetting 2.0.{quote}\nThe analysis of impact here seems confused. What has happened is not that it took 3 years for someone to encounter this bug; rather, it took 3 years for someone who ran across this very serious bug to notice and care enough to file a JIRA.\n\nBugs which violate guarantees in the way this one does (\"I repaired once every gc_grace_seconds and still ended up with zombie data!\") should IMO be considered Major or Critical and should be candidates for fixes at any point in the release cycle. If we can entirely re-write a broken feature like replace_node in 1.2.12, I fail to see what keeps us from fixing buggy code that erroneously treats temporary files as permanent, in 1.2.14.","created":"2014-01-10T21:24:18.704+0000"},{"body":"bq. What is the \"since\" version for this ticket?\n\nI suspect it's been around for a while - I only (knowingly) ran into the problem with 1.1.7.\n\nbq. \"3 years\"\n\nThe original descriptions says \"3 months\", and that is how long we took to run into the problem (at a serious level). I think perhaps [~jbellis] misspoke when he mentioned \"3 years\" in an earlier comment. \n\nAs for fix version, I would like to see this in 1.2, as well, but it depends on the changes needed. At this point, though, the patch I have now (still WIP) is not too dramatic, so it could reasonably go into 1.2, but we should judge when it's ready/accepted.","created":"2014-01-10T21:52:49.267+0000"},{"body":"I'm counting 3 years since 0.7 which is at least as long as repair has worked this way.","created":"2014-01-10T22:29:33.291+0000"},{"body":"[~jbellis] Ahh, OK. Thank you for clarifying.","created":"2014-01-10T22:48:44.747+0000"},{"body":"I think most people notice they have hung repairs, and restart the nodes to clear them, before it becomes a problem.","created":"2014-01-10T22:52:39.780+0000"},{"body":"Attached v2 patch has the following changes:\n\n- Changed StreamReceiveTask to keep a collection of SSTW rather than SSTR. This allows us to do the conversion of SSTW to SSTR all together after we've gotten all the streamed files. Also fixed up the code paths to here so they pass SSTW.\n\n- Also in StreamReceiveTask, added an abort() method, which will discard the SSTWs it has buffered up. Changed StreamSession so that when a session ends in failure, it calls the new STR.abort() method.\n\n- Split FileMessage out into IncomingFileMessage and OutgoingFileMessage. I needed to do this since as each one has a different subclass of SSTable, but also because java generics doesn't allow me to return different subclasses from StreamMessage.Serializer. This necessitated the changes in StreamMessage as I couldn't have one serializer for both IncomingFileMessage and OutgoingFileMessage. As it didn't seem best to create a new StreamMessage.Type (something like FILE_IN and FILE_OUT) just to represent the FILE message type's behavior on inbound vs. outbound, I instead split the SM.Type.serializer into two variables: inSerializer and outSerializer. For all the other Type's, the in and out serializers are the same class; in the case of Type.FILE, this is where I'm referencing IncomingFileMessage.serializer and OutgoingFileMessage.serializer, respectively. This seemed the cleanest way to introduce the now-bifurcated life of Type.FILE/FileMessage.\n\n- added StreamLockfile to satisfy [~yukim]'s request for a mechanism to remove, on restart, the subset of SSTRs that were successfully converted when others from it's stream session failed. Assumes the process crashed in the middle of converting the SSTWs to SSTRs.\n\nIn the first patch, I chose to write the lockfile out to the commitlog directory. I did this as it seems like overkill to add another yaml setting (and Config/DD change) just for this value. Thus, I wanted to piggyback off something else that we already have, and DD.getCommitLogDirectory seemed the least worst. I'm open to suggestions on this.\n\nOnce these changes are incorporated into 2.0 and trunk, I would still like to do something for 1.2 but I do not think we need to be as extensive as what we're doing for 2.0+. Perhaps leave out the lockfile and the abort(), and just leave the deferring of converting SSTW to SSTR until the end of the session (basically what the current 1.2 patch does, but I'll check it out again after the 2.0 stuff is good).\n","created":"2014-01-23T13:42:36.346+0000"},{"body":"bq. In the first patch, I chose to write the lockfile out to the commitlog directory. \n\nHow about using SSTable directory itself? You can access it using Directories object so it is easy to perform delete if we do it inside CFS.scrubDataDirectories. Attaching patch for this.\n\nOther than that, it seems good.\n\nbq. ...I would still like to do something for 1.2 but I do not think we need to be as extensive as what we're doing for 2.0+. Perhaps leave out the lockfile and the abort(), and just leave the deferring of converting SSTW to SSTR until the end of the session...\n\nI agree. We can commit attached 1.2 patch as well.","created":"2014-01-29T21:07:54.231+0000"},{"body":"Committed the 1.2 patch to 1.2, have a few questions for [~yukim] about 2.0 patch that I'll add here in a minute (after coffee)","created":"2014-01-30T18:01:50.325+0000"},{"body":"bq. How about using SSTable directory itself?\n\nI think that is legit as StreamReceiveTask is specific to a CF.","created":"2014-01-30T19:32:32.187+0000"},{"body":"[~jasobrown] Looks like we need to make a little tweak around complete message since we moved adding received SSTables to different thread.\nIt is breaking StreamingTransferTest.\n\nPatch attach for fix.","created":"2014-01-31T01:59:10.154+0000"},{"body":"[~yukim] i committed the followup patch to 2.0, revert if that was a bad idea :)","created":"2014-01-31T10:58:37.102+0000"},{"body":"The followup patch is fine, but I'm not quite sure why it is needed. How does being in a different thread affect the sending of the CompleteMessage, which doesn't look like it was there before?","created":"2014-01-31T14:23:52.956+0000"},{"body":"[~jasobrown] Complete message exchange was actually fragile before and could leave streaming session to WAIT_COMPLETE state on one side.\n\nOnly bellow pattern worked, and it worked because we were sending complete as we receive FileMessage in the same thread.\n\n{code}\n(A) ---> File ---> (B) ...1\n(A) <--- Complete <--- (B) ...2\n(A) ---> Complete ---> (B) ...3\n{code}\n\nBut now finalizing all received files moved to another thread. So sending receiving complete from A(3) gets first and B terminates it session without sending back complete, leaving A as WAIT_COMPLETE. Thus we needed to make sure to send complete message.\n","created":"2014-01-31T23:44:46.163+0000"}],"conversations":[{"body":"The sstables streamed in during a repair session don't become active until the session finishes. If something causes the repair session to hang for some reason, those sstables will hang around until the next reboot, and become active then. If you don't reboot for 3 months, this can cause data to resurrect, as GC grace has expired, so tombstones for the data in those sstables may have already been collected.","from":"reporter","subject":"sstables from stalled repair sessions become live after a reboot and can resurrect deleted data"},{"body":"One thing I was thinking that might help with this is if we could leave the sstables named \"-tmp\" until we are ready to make them become active. That way when you reboot, any files hanging around will get removed on restart, instead of becoming active.","from":"developer"},{"body":"I don't think we should be messing with repair code in 1.2.14, not for something that it took 3 years for someone to run across. Suggest targetting 2.0.","from":"developer"},{"body":"(Not actually sure this is still an issue in 2.0.x but it seems possible.)","from":"developer"},{"body":"In 1.2 as well as 2.0, every time the node receives SSTable file, it starts writing as 'tmp', but at the end of receiving, it calls SSTableWriter#closeAndOpenReader that moves received file out of 'tmp'. At this point, SSTables are still not added to ColumnFamilyStore, though if we shutdown the node and restart, the node would recognize received files even though streaming session was not finished successfully.\n\nWe need little tweak to defer renaming of SSTable files to do so after receiving all streamed files.\n ","from":"developer"},{"body":"Attached patch, 6503_c1.2-v1, defers the release of the sstables to the CFS until the session is complete. Note: that patch is only for 1.2.\n\nFor c* 2.0, I'd like [~yukim]'s advice. I have a WIP here: https://github.com/jasobrown/cassandra/tree/6503_c2.0. The problem I'm running into is that FileMessage.sstable is of type SSTableReader, which we need on the sender side, but on the receiver side we want SSTableWriter (if we are going to defer the release of the sstables. For hacking things up sake, I've just changed FileMessage.sstable to a plain SSTable and let the users do the casting - which is only in two places, one of which is the FileMessage.Serializer.serailize() method. Not very extensive, but perhaps a bit sloppy.\n\nYuki, do you think it's worthwhile to split up the FileMessage object into two classes like OutFileMessage (which has a SSTR) and InFileMessage (which has SSTW)? ","from":"developer"},{"body":"[~jasobrown] What I'm thinking to do is to closeAndOpenReader without renaming from tmp as we receive, and rename them at once at the end. And for renaming multiple files at once we probably want some kind of lockfile(also resolve CASSANDRA-2900?).\n\nThough finalizing SSTable write in closeAndOpenReader takes some time, so completely defer finalize as you do may be a good idea. I think splitting FileMessage is better than casting.","from":"developer"},{"body":"bq. we probably want some kind of lockfile\n\nInteresting. What problems do you see this solving (I'm probably missing something in my understanding)?\n\nbq. splitting FileMessage is better than casting\n\nYeah, I knew you were gonna say that :)","from":"developer"},{"body":"[~yukim] I can see how modifying closeAndOpenReader() would help here, but there's one wrinkle (I think): if we defer renaming the files from closeAndOpenReader(), we would need to rename the *open* SSTR, and there's a note in FileUtils.renameWithConfirm():\n\n{code} // this is not FSWE because usually when we see it it's because we didn't close the file before renaming it,\n // and Windows is picky about that.{code}\n\nSo I'm not sure about the deferred rename of the SSTR, assuming we care about Windows.\n","from":"developer"},{"body":"I think you are right. We don't want to close SSTR again for renaming.\nLet's stick with deferring closeAndOpenReader as you did.\nFor the patch, I think it would be nice to call SSTW.abort() to discard already received file when something bad happened.\n\nbq. Interesting. What problems do you see this solving (I'm probably missing something in my understanding)?\n\nIf the node goes down during we are doing closeAndOpenReader to received files, there is a chance we have already renamed files.\nSo I wanted to make sure the node won't read those files when the node come up.\n","from":"developer"},{"body":"bq. call SSTW.abort() ...\n\nMakes sense, will add that in.\n\nbq. If the node goes down during we are doing closeAndOpenReader to received files, there is a chance we have already renamed files\n\nOk, got it. Will come up with an idea and drop into the next patch. Thanks","from":"developer"},{"body":"What is the \"since\" version for this ticket?\n\nIt seems rather serious that users were still exposed to zombie data even if following repair best practice, if a repair session ever stalled... especially because repair sessions tend to stall. In a way we are fortunate that the only way to restart a failed repair session is to restart the nodes (CASSANDRA-3486) because this workaround has protected some users from this bug.\n\n{quote}I don't think we should be messing with repair code in 1.2.14, not for something that it took 3 years for someone to run across. Suggest targetting 2.0.{quote}\nThe analysis of impact here seems confused. What has happened is not that it took 3 years for someone to encounter this bug; rather, it took 3 years for someone who ran across this very serious bug to notice and care enough to file a JIRA.\n\nBugs which violate guarantees in the way this one does (\"I repaired once every gc_grace_seconds and still ended up with zombie data!\") should IMO be considered Major or Critical and should be candidates for fixes at any point in the release cycle. If we can entirely re-write a broken feature like replace_node in 1.2.12, I fail to see what keeps us from fixing buggy code that erroneously treats temporary files as permanent, in 1.2.14.","from":"developer"},{"body":"bq. What is the \"since\" version for this ticket?\n\nI suspect it's been around for a while - I only (knowingly) ran into the problem with 1.1.7.\n\nbq. \"3 years\"\n\nThe original descriptions says \"3 months\", and that is how long we took to run into the problem (at a serious level). I think perhaps [~jbellis] misspoke when he mentioned \"3 years\" in an earlier comment. \n\nAs for fix version, I would like to see this in 1.2, as well, but it depends on the changes needed. At this point, though, the patch I have now (still WIP) is not too dramatic, so it could reasonably go into 1.2, but we should judge when it's ready/accepted.","from":"developer"},{"body":"I'm counting 3 years since 0.7 which is at least as long as repair has worked this way.","from":"developer"},{"body":"[~jbellis] Ahh, OK. Thank you for clarifying.","from":"developer"},{"body":"I think most people notice they have hung repairs, and restart the nodes to clear them, before it becomes a problem.","from":"developer"},{"body":"Attached v2 patch has the following changes:\n\n- Changed StreamReceiveTask to keep a collection of SSTW rather than SSTR. This allows us to do the conversion of SSTW to SSTR all together after we've gotten all the streamed files. Also fixed up the code paths to here so they pass SSTW.\n\n- Also in StreamReceiveTask, added an abort() method, which will discard the SSTWs it has buffered up. Changed StreamSession so that when a session ends in failure, it calls the new STR.abort() method.\n\n- Split FileMessage out into IncomingFileMessage and OutgoingFileMessage. I needed to do this since as each one has a different subclass of SSTable, but also because java generics doesn't allow me to return different subclasses from StreamMessage.Serializer. This necessitated the changes in StreamMessage as I couldn't have one serializer for both IncomingFileMessage and OutgoingFileMessage. As it didn't seem best to create a new StreamMessage.Type (something like FILE_IN and FILE_OUT) just to represent the FILE message type's behavior on inbound vs. outbound, I instead split the SM.Type.serializer into two variables: inSerializer and outSerializer. For all the other Type's, the in and out serializers are the same class; in the case of Type.FILE, this is where I'm referencing IncomingFileMessage.serializer and OutgoingFileMessage.serializer, respectively. This seemed the cleanest way to introduce the now-bifurcated life of Type.FILE/FileMessage.\n\n- added StreamLockfile to satisfy [~yukim]'s request for a mechanism to remove, on restart, the subset of SSTRs that were successfully converted when others from it's stream session failed. Assumes the process crashed in the middle of converting the SSTWs to SSTRs.\n\nIn the first patch, I chose to write the lockfile out to the commitlog directory. I did this as it seems like overkill to add another yaml setting (and Config/DD change) just for this value. Thus, I wanted to piggyback off something else that we already have, and DD.getCommitLogDirectory seemed the least worst. I'm open to suggestions on this.\n\nOnce these changes are incorporated into 2.0 and trunk, I would still like to do something for 1.2 but I do not think we need to be as extensive as what we're doing for 2.0+. Perhaps leave out the lockfile and the abort(), and just leave the deferring of converting SSTW to SSTR until the end of the session (basically what the current 1.2 patch does, but I'll check it out again after the 2.0 stuff is good).\n","from":"developer"},{"body":"bq. In the first patch, I chose to write the lockfile out to the commitlog directory. \n\nHow about using SSTable directory itself? You can access it using Directories object so it is easy to perform delete if we do it inside CFS.scrubDataDirectories. Attaching patch for this.\n\nOther than that, it seems good.\n\nbq. ...I would still like to do something for 1.2 but I do not think we need to be as extensive as what we're doing for 2.0+. Perhaps leave out the lockfile and the abort(), and just leave the deferring of converting SSTW to SSTR until the end of the session...\n\nI agree. We can commit attached 1.2 patch as well.","from":"developer"},{"body":"Committed the 1.2 patch to 1.2, have a few questions for [~yukim] about 2.0 patch that I'll add here in a minute (after coffee)","from":"developer"},{"body":"bq. How about using SSTable directory itself?\n\nI think that is legit as StreamReceiveTask is specific to a CF.","from":"developer"},{"body":"[~jasobrown] Looks like we need to make a little tweak around complete message since we moved adding received SSTables to different thread.\nIt is breaking StreamingTransferTest.\n\nPatch attach for fix.","from":"developer"},{"body":"[~yukim] i committed the followup patch to 2.0, revert if that was a bad idea :)","from":"developer"},{"body":"The followup patch is fine, but I'm not quite sure why it is needed. How does being in a different thread affect the sending of the CompleteMessage, which doesn't look like it was there before?","from":"developer"},{"body":"[~jasobrown] Complete message exchange was actually fragile before and could leave streaming session to WAIT_COMPLETE state on one side.\n\nOnly bellow pattern worked, and it worked because we were sending complete as we receive FileMessage in the same thread.\n\n{code}\n(A) ---> File ---> (B) ...1\n(A) <--- Complete <--- (B) ...2\n(A) ---> Complete ---> (B) ...3\n{code}\n\nBut now finalizing all received files moved to another thread. So sending receiving complete from A(3) gets first and B terminates it session without sending back complete, leaving A as WAIT_COMPLETE. Thus we needed to make sure to send complete message.\n","from":"developer"}],"created":"2013-12-18T18:52:10.000+0000","description":"The sstables streamed in during a repair session don't become active until the session finishes. If something causes the repair session to hang for some reason, those sstables will hang around until the next reboot, and become active then. If you don't reboot for 3 months, this can cause data to resurrect, as GC grace has expired, so tombstones for the data in those sstables may have already been collected.","issue_id":"12685433","key":"CASSANDRA-6503","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-01-31T13:31:05.000+0000","role":"fixed_distractor","summary":"sstables from stalled repair sessions become live after a reboot and can resurrect deleted data"} {"case_id":"12444117","cluster":"DISTRACTOR-CASSANDRA-651","comments":[{"body":"More info:\n I do see that Node X.X.X.X is dead, and\nNode X.X.X.X has restarted.\n\nThis show up on all the 3 other servers:\n INFO [Timer-1] 2009-12-22 20:38:43,738 Gossiper.java (line 194)\nInetAddress /10.6.168.20 is now dead.\n\nNode /10.6.168.20 has restarted, now UP again\n INFO [GMFD:1] 2009-12-22 20:43:12,812 StorageService.java (line 475)\nNode /10.6.168.20 state jump to normal\n\n","created":"2009-12-23T17:39:27.522+0000"},{"body":"This is definitely a regression in 0.5. I tested out 0.4.2 and it works perfectly fine, and load goes back up to 100% on the restarted node. ","created":"2009-12-23T17:43:09.854+0000"},{"body":"I was able to reproduce this in a 4 node setup as well. The recovered node does not appear to receive any writes after rejoining the cluster, and I receive TimedOutExceptions in the client. I was able to write directly to the recovered node and things appeared to work, however after a while a different node OOM'd. I examined the dump in MAT and it shows org.apache.cassandra.net.MessagingService occupies 72.53% of the available heap, followed by java.util.concurrent.LinkedBlockingQueue using 11.87%.","created":"2009-12-28T22:00:42.573+0000"},{"body":"Confirm this issue by 8 nodes tests,\n\nAfter read system.log, I found after one node down and up again, some other nodes will not establish tcp connection to it(on tcp port 7000 ) forever! \nAnd read request sent to it (into Pending-Writes because socket channel is closed) will not sent to ethernet forever(from observing tcpdump).\n\nmaybe PendingWrites Queue consume lots memory and OOM (yes, it not the recovered node, but the node who try to send request to recovered node!)\n\nIt's seems when recovered node going down, some other node's socket channel was reset , after it come back, these socket channel remain closed, forever\n","created":"2009-12-29T03:34:49.896+0000"},{"body":"Brandon Williams said: \"I see what's happening with 651 -- TcpConnectionManager keeps trying to reuse a closed connection and never opens new ones to the recovered node\"\n\nsounds like a regression from CASSANDRA-488 to me.","created":"2009-12-29T14:26:00.148+0000"},{"body":"Patched into trunk. Link the gossip and messaging service so that invalid connection pools can be shutdown when a node goes offline.","created":"2009-12-29T18:45:44.259+0000"},{"body":"Patch is for trunk.","created":"2009-12-29T18:46:20.719+0000"},{"body":"relying on FD to notice is not going to work, though, since FD is not instantaneous (and cannot be made so). [edit: that is, a node could die and come back, or be partitioned and be available again, quickly enough that FD does not notice but old connections are still invalid.] \n\ncan we have it attempt to reconnect when it encounters an error sending instead?","created":"2009-12-29T18:51:48.351+0000"},{"body":"The last patch diffed in the wrong direction. This one is correct.","created":"2009-12-29T18:58:38.557+0000"},{"body":"that said, it could be good to have the FD *in addition* to the other, so that if a node goes down for a while that doesn't have much traffic, we don't lose the first attempted message once it's back up unnecessarily.","created":"2009-12-29T19:09:02.620+0000"},{"body":"I can still reproduce the issue with this patch applied. I'm receiving the following traceback:\n\nERROR - Fatal exception in thread Thread[TCP Selector Manager,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.net.TcpConnectionManager.destroy(TcpConnectionManager.java:85)\n at org.apache.cassandra.net.TcpConnection.errorClose(TcpConnection.java:319)\n at org.apache.cassandra.net.TcpConnection.connect(TcpConnection.java:364)\n at org.apache.cassandra.net.SelectorManager.doProcess(SelectorManager.java:143)\n at org.apache.cassandra.net.SelectorManager.run(SelectorManager.java:107)\n\nBecause ackCon is already null. With the assert removed, I'm unable to reproduce.","created":"2009-12-29T19:54:42.789+0000"},{"body":"Updated to include replacing calls to destroy() with shutdown(). Chances are if one TC is crappy, the other is not going to be useful either.","created":"2009-12-29T23:09:00.857+0000"},{"body":"+1, no longer reproducible with this patch. The recovered node begins receiving writes normally.","created":"2009-12-30T01:02:16.781+0000"},{"body":"Jaako and I had a discussion in which we agreed that it would be better to have MessagingService implement IEndPointStateChangeSubscriber and subscribe to the gossiper rather than just implement IFailureDetector. This would have provided a way to basically turn connection pools on and off and would allow writes to fail a bit faster.\n\nI implemented that this morning. It's a few more lines of code and all it really buys us is more descriptive error messages. We'll just have to put up with errored writes until gossip takes the failed node out of the ring.","created":"2009-12-30T13:40:49.679+0000"},{"body":"+1 from me too, this seems to have fixed the problem. ","created":"2009-12-30T18:51:52.406+0000"},{"body":"I still think we should not rely on FD, or we will still hit this bug with short-lived partitions (which do occur in the wild).\n\nsomething like brandon's throwing an exception if not connected or awaiting connection.\n\nI'm still baffled that write() apparently doesn't throw when the connection dies... Are we missing something there?","created":"2009-12-30T19:01:25.116+0000"},{"body":"My mistake: write _was_ throwing, but clearing the old conn out was not working. Still trying to understand why.","created":"2009-12-30T19:27:36.779+0000"},{"body":"MessagingService.sendOneWay inconveniently swallows errors, so StorageProxy is never wise about them.","created":"2009-12-30T19:30:28.873+0000"},{"body":"Got it: the problem is that sendOneWay calls shutdown on SocketException, not errorClose. so the part that really matters in gary's fix is adding the nulling out to shutdown.","created":"2009-12-30T20:28:47.298+0000"},{"body":"this version gets rid of the formely-problematic-and-now-redundant SocketException block, and renames TCM.shutdown to reset.","created":"2009-12-30T20:35:53.447+0000"},{"body":"Reviewed. +1 on the v4 patch.","created":"2009-12-30T22:06:42.277+0000"},{"body":"Fixed in trunk and 0.5 branch. Patch by Gary Dusbabek and Jonathan Ellis. Reviewed by same.","created":"2009-12-30T23:20:40.228+0000"}],"conversations":[{"body":"From the cassandra user message board: \n\"I just recently upgraded to latest in 0.5 branch, and I am running\ninto a serious issue. I have a cluster with 4 nodes, rackunaware\nstrategy, and using my own tokens distributed evenly over the hash\nspace. I am writing/reading equally to them at an equal rate of about\n230 reads/writes per second(and cfstats shows that). The first 3 nodes\nare seeds, the last one isn't. When I start all the nodes together at\nthe same time, they all receive equal amounts of reads/writes (about\n230).\nWhen I bring node 4 down and bring it back up again, node 4's load\nfluctuates between the 230 it used to get to sometimes no traffic at\nall. The other 3 still have the same amount of traffic. And no errors\nwhat so ever seen in logs. \" ","from":"reporter","subject":"cassandra 0.5 version throttles and sometimes kills traffic to a node if you restart it."},{"body":"More info:\n I do see that Node X.X.X.X is dead, and\nNode X.X.X.X has restarted.\n\nThis show up on all the 3 other servers:\n INFO [Timer-1] 2009-12-22 20:38:43,738 Gossiper.java (line 194)\nInetAddress /10.6.168.20 is now dead.\n\nNode /10.6.168.20 has restarted, now UP again\n INFO [GMFD:1] 2009-12-22 20:43:12,812 StorageService.java (line 475)\nNode /10.6.168.20 state jump to normal\n\n","from":"developer"},{"body":"This is definitely a regression in 0.5. I tested out 0.4.2 and it works perfectly fine, and load goes back up to 100% on the restarted node. ","from":"developer"},{"body":"I was able to reproduce this in a 4 node setup as well. The recovered node does not appear to receive any writes after rejoining the cluster, and I receive TimedOutExceptions in the client. I was able to write directly to the recovered node and things appeared to work, however after a while a different node OOM'd. I examined the dump in MAT and it shows org.apache.cassandra.net.MessagingService occupies 72.53% of the available heap, followed by java.util.concurrent.LinkedBlockingQueue using 11.87%.","from":"developer"},{"body":"Confirm this issue by 8 nodes tests,\n\nAfter read system.log, I found after one node down and up again, some other nodes will not establish tcp connection to it(on tcp port 7000 ) forever! \nAnd read request sent to it (into Pending-Writes because socket channel is closed) will not sent to ethernet forever(from observing tcpdump).\n\nmaybe PendingWrites Queue consume lots memory and OOM (yes, it not the recovered node, but the node who try to send request to recovered node!)\n\nIt's seems when recovered node going down, some other node's socket channel was reset , after it come back, these socket channel remain closed, forever\n","from":"developer"},{"body":"Brandon Williams said: \"I see what's happening with 651 -- TcpConnectionManager keeps trying to reuse a closed connection and never opens new ones to the recovered node\"\n\nsounds like a regression from CASSANDRA-488 to me.","from":"developer"},{"body":"Patched into trunk. Link the gossip and messaging service so that invalid connection pools can be shutdown when a node goes offline.","from":"developer"},{"body":"Patch is for trunk.","from":"developer"},{"body":"relying on FD to notice is not going to work, though, since FD is not instantaneous (and cannot be made so). [edit: that is, a node could die and come back, or be partitioned and be available again, quickly enough that FD does not notice but old connections are still invalid.] \n\ncan we have it attempt to reconnect when it encounters an error sending instead?","from":"developer"},{"body":"The last patch diffed in the wrong direction. This one is correct.","from":"developer"},{"body":"that said, it could be good to have the FD *in addition* to the other, so that if a node goes down for a while that doesn't have much traffic, we don't lose the first attempted message once it's back up unnecessarily.","from":"developer"},{"body":"I can still reproduce the issue with this patch applied. I'm receiving the following traceback:\n\nERROR - Fatal exception in thread Thread[TCP Selector Manager,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.net.TcpConnectionManager.destroy(TcpConnectionManager.java:85)\n at org.apache.cassandra.net.TcpConnection.errorClose(TcpConnection.java:319)\n at org.apache.cassandra.net.TcpConnection.connect(TcpConnection.java:364)\n at org.apache.cassandra.net.SelectorManager.doProcess(SelectorManager.java:143)\n at org.apache.cassandra.net.SelectorManager.run(SelectorManager.java:107)\n\nBecause ackCon is already null. With the assert removed, I'm unable to reproduce.","from":"developer"},{"body":"Updated to include replacing calls to destroy() with shutdown(). Chances are if one TC is crappy, the other is not going to be useful either.","from":"developer"},{"body":"+1, no longer reproducible with this patch. The recovered node begins receiving writes normally.","from":"developer"},{"body":"Jaako and I had a discussion in which we agreed that it would be better to have MessagingService implement IEndPointStateChangeSubscriber and subscribe to the gossiper rather than just implement IFailureDetector. This would have provided a way to basically turn connection pools on and off and would allow writes to fail a bit faster.\n\nI implemented that this morning. It's a few more lines of code and all it really buys us is more descriptive error messages. We'll just have to put up with errored writes until gossip takes the failed node out of the ring.","from":"developer"},{"body":"+1 from me too, this seems to have fixed the problem. ","from":"developer"},{"body":"I still think we should not rely on FD, or we will still hit this bug with short-lived partitions (which do occur in the wild).\n\nsomething like brandon's throwing an exception if not connected or awaiting connection.\n\nI'm still baffled that write() apparently doesn't throw when the connection dies... Are we missing something there?","from":"developer"},{"body":"My mistake: write _was_ throwing, but clearing the old conn out was not working. Still trying to understand why.","from":"developer"},{"body":"MessagingService.sendOneWay inconveniently swallows errors, so StorageProxy is never wise about them.","from":"developer"},{"body":"Got it: the problem is that sendOneWay calls shutdown on SocketException, not errorClose. so the part that really matters in gary's fix is adding the nulling out to shutdown.","from":"developer"},{"body":"this version gets rid of the formely-problematic-and-now-redundant SocketException block, and renames TCM.shutdown to reset.","from":"developer"},{"body":"Reviewed. +1 on the v4 patch.","from":"developer"},{"body":"Fixed in trunk and 0.5 branch. Patch by Gary Dusbabek and Jonathan Ellis. Reviewed by same.","from":"developer"}],"created":"2009-12-23T17:35:22.000+0000","description":"From the cassandra user message board: \n\"I just recently upgraded to latest in 0.5 branch, and I am running\ninto a serious issue. I have a cluster with 4 nodes, rackunaware\nstrategy, and using my own tokens distributed evenly over the hash\nspace. I am writing/reading equally to them at an equal rate of about\n230 reads/writes per second(and cfstats shows that). The first 3 nodes\nare seeds, the last one isn't. When I start all the nodes together at\nthe same time, they all receive equal amounts of reads/writes (about\n230).\nWhen I bring node 4 down and bring it back up again, node 4's load\nfluctuates between the 230 it used to get to sometimes no traffic at\nall. The other 3 still have the same amount of traffic. And no errors\nwhat so ever seen in logs. \" ","issue_id":"12444117","key":"CASSANDRA-651","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2009-12-30T23:20:42.000+0000","role":"fixed_distractor","summary":"cassandra 0.5 version throttles and sometimes kills traffic to a node if you restart it."} {"case_id":"12687035","cluster":"DISTRACTOR-CASSANDRA-6541","comments":[{"body":"I've seen at least 3 users in the field hit this issue, in each case enabling CMSClassUnloadingEnabled solved the issue.","created":"2014-01-02T20:45:11.806+0000"},{"body":"I think we've been seeing the same issue, as well, with JMX. Do you have any links or refs about the issue/resolution?","created":"2014-01-02T21:21:24.109+0000"},{"body":"[~benedict] tracked it down to most likely being this code change: http://hg.openjdk.java.net/icedtea/jdk7/jdk/rev/985b53122cf8","created":"2014-01-03T15:51:49.540+0000"},{"body":"A bit more context:\n\n# We have heap dumps full of javax.management.remote.rmi.RMIConnectionImpl$CombinedClassLoader\n# CombinedClassLoader appears to be sandboxing classes used by RMI/JMX and thus instantiates its own copies for each connection\n# If you only use a stable JMX connection pool you are fine, otherwise you \"leak\" CCL roots\n# If you have bad enough old-gen fragmentation to force periodic STW GC, you are also fine; perversely, this will only bite you if CMS is working as designed","created":"2014-01-03T16:03:05.323+0000"},{"body":"[~nickmbailey] did you ever file an openjdk bug, or are we convinced this is the new Working As Designed?","created":"2014-01-03T16:04:56.299+0000"},{"body":"[~jbellis] Yes, that is exactly what we have been seeing, as well. However, I'm not sure if our JVM is that 'recent' (/me weeps silently) - but, I'll try to confirm today across several of our clusters. ","created":"2014-01-03T16:35:26.627+0000"},{"body":"[~jasobrown] fwiw, even though the diff is from openjdk 1.7, we've seen it as far back as jdk 1.6.0 update 45 thus far.","created":"2014-01-03T16:38:55.997+0000"},{"body":"Yeah, a brute force check of some clusters confirms some are on 1.6.0_41, _37, and _45 - and, of course, they are mixed within the cluster, depending on when the instance was launched (if it was with an older AMI with older JVM).\n\nI haven't narrowed the problem down to which instances had troubles with which JVM, but some clusters with < u45 did not have any problems, so I suspect it's at least > u41 (our most highest version), but probably [~benedict] is correct with u45.","created":"2014-01-03T16:51:44.077+0000"},{"body":"I attempted to file a bug with oracle, but the end result was a notification that I would be emailed with either a bug number or further questions. I haven't heard anything from them yet though. It's been a little less than a month.","created":"2014-01-03T18:54:33.887+0000"},{"body":"This goes back at least as far as 1.6u43, just had someone hit it there.","created":"2014-01-09T14:05:27.674+0000"},{"body":"We could try emailing dmitry.samersoff@oracle.com who contributed the change, and andrew.gross@oracle.com who reviewed it, directly?","created":"2014-01-09T14:33:31.241+0000"},{"body":"FYI, I don't think I asserted this was u45. Only a specific changeset, which appears to have been in OpenJDK 7u12. It's not possible to see when this was included in Java 6, but could have been as early as u39 judging from release dates on Wikipedia.","created":"2014-01-09T15:45:28.246+0000"},{"body":"Yeah, I was adding 1.6 versions where we've seen it, since we can't know when they put that into 1.6.","created":"2014-01-09T15:55:18.267+0000"},{"body":"I have emailed dmitry and andrew, ccing cassandra-dev.","created":"2014-01-09T16:34:14.484+0000"},{"body":"I emailed jmx-dev a few weeks ago, and it's been tumbleweed. I intend to follow up soon, and try the core-dev list (see if I can get any response from a higher traffic list), but I expect this will be slow going. Nobody seems interested.","created":"2014-02-24T15:01:20.807+0000"},{"body":"Will {{-XX:+CMSClassUnloadingEnabled}} cause older VMs to bitch at us? If not let's add it as a default to 1.1+.","created":"2014-02-24T15:05:51.593+0000"},{"body":"We already do java version detection in shell and add flags based on that, so adding this shouldn't be a problem.","created":"2014-02-24T18:30:01.328+0000"},{"body":"Well, then someone needs to figure out when it was introduced. :)","created":"2014-02-24T19:25:34.931+0000"},{"body":"I volunteer [~mshuler] :)","created":"2014-02-25T13:41:56.503+0000"},{"body":"I get no complaints using -XX:+CMSClassUnloadingEnabled with 1.5.0_22-b03 and found an old doc that lists this VM option available in 1.4.1 1.4.2 1.5.0 1.6.0 - http://www.reins.altervista.org/java/A_Collection_of_JVM_Options_MP.html\nI'd say it's probably pretty safe to add it.","created":"2014-02-28T20:52:17.589+0000"},{"body":"I committed this.","created":"2014-02-28T21:19:37.856+0000"},{"body":"since I was there..\n$ /opt/jdk1.8.0/bin/java -XX:+CMSClassUnloadingEnabled -XX:+PrintFlagsFinal -version |grep CMSClassUnloadingEnabled\n bool CMSClassUnloadingEnabled := true {product}\njava version \"1.8.0\"\nJava(TM) SE Runtime Environment (build 1.8.0-b128)\nJava HotSpot(TM) 64-Bit Server VM (build 25.0-b69, mixed mode)","created":"2014-03-01T05:30:50.881+0000"},{"body":"Is Cassandra \"since\" on this issue \"the dawn of time\"?","created":"2014-03-11T01:08:09.017+0000"},{"body":"It's a HotSpot regression that we're working around, not a Cassandra bug.","created":"2014-03-11T05:16:22.245+0000"},{"body":"Just adding some more java versions here for people wondering if they are hitting this. Just saw this on 1.6.0_38, so it goes back at least that far on 1.6","created":"2014-03-21T15:01:19.030+0000"},{"body":"{quote}It's a HotSpot regression that we're working around, not a Cassandra bug.{quote}\nYes, so the statement \"this HotSpot regression is not-worked-around in all extant versions of Cassandra since the dawn of time\" is correct. Thanks for the clarification!","created":"2014-03-24T18:48:08.181+0000"},{"body":"Well, not really, since the problematic JVM versions didn't exist back then.","created":"2014-03-24T18:49:59.906+0000"},{"body":"bq. Yes, so the statement \"this HotSpot regression is not-worked-around in all extant versions of Cassandra since the dawn of time\" is correct. Thanks for the clarification!\n\nNo, it isn't, since the hot spot bug did not exist since the dawn of time. As such the from-version is ill-defined, but probably the most sensible definition requires determining the from-version-for-hotspot, which we have yet to manage, and aligning that with C* releases.","created":"2014-03-24T18:52:00.148+0000"},{"body":"C* from version really has no meaning here. But yes, if you run C* 0.3 (if that did JMX monitoring) with 1.6_45 you will hit this issue. Unless of course you add the given setting to your cassandra-env.sh. Which can be done without upgrading your C*.","created":"2014-03-24T19:22:13.482+0000"},{"body":"On DSE 4.5.1, Java 1.7, linux ami on aws encountered long running GC before dse was shutdown.This was a ring of 6 nodes, and no unexpected processes /load were run.The error log attached.\n\nI suspect to have hit that hotspot/java GC issue.But in this case the class unloading is enabled.\nJVM_OPTS=\"$JVM_OPTS -XX:+CMSClassUnloadingEnabled”\nXX:CMSInitiatingOccupancyFraction=75\n\nNow configured to have MAX_HEAP_SIZE=“8G”, HEAP_NEWSIZE=“2G” and GC logging enabled and is stable.However if there is a heap leak, it needs a better fix.\nSuggestion: If the CPU utilization is going up, or heap size leak/long running GC is sensed, the nodetool -h localhost flush can auto kick in?\n\nThanks\nRekha\n","created":"2015-02-28T05:16:17.981+0000"},{"body":"Dse system log with long running GC before DSE shutdown","created":"2015-02-28T05:17:47.831+0000"},{"body":"On newer systems with G1 GC is this option ignored? I'm assuming it would be.","created":"2015-09-20T08:22:46.915+0000"}],"conversations":[{"body":"Newer versions of Oracle's Hotspot JVM , post 6u43 (maybe earlier) and 7u25 (maybe earlier), are experiencing issues with GC and JMX where heap slowly fills up overtime until OOM or a full GC event occurs, specifically when CMS is leveraged. Adding:\n\n{noformat}\nJVM_OPTS=\"$JVM_OPTS -XX:+CMSClassUnloadingEnabled\"\n{noformat}\n\nThe the options in cassandra-env.sh alleviates the problem.","from":"reporter","subject":"New versions of Hotspot create new Class objects on every JMX connection causing the heap to fill up with them if CMSClassUnloadingEnabled isn't set."},{"body":"I've seen at least 3 users in the field hit this issue, in each case enabling CMSClassUnloadingEnabled solved the issue.","from":"developer"},{"body":"I think we've been seeing the same issue, as well, with JMX. Do you have any links or refs about the issue/resolution?","from":"developer"},{"body":"[~benedict] tracked it down to most likely being this code change: http://hg.openjdk.java.net/icedtea/jdk7/jdk/rev/985b53122cf8","from":"developer"},{"body":"A bit more context:\n\n# We have heap dumps full of javax.management.remote.rmi.RMIConnectionImpl$CombinedClassLoader\n# CombinedClassLoader appears to be sandboxing classes used by RMI/JMX and thus instantiates its own copies for each connection\n# If you only use a stable JMX connection pool you are fine, otherwise you \"leak\" CCL roots\n# If you have bad enough old-gen fragmentation to force periodic STW GC, you are also fine; perversely, this will only bite you if CMS is working as designed","from":"developer"},{"body":"[~nickmbailey] did you ever file an openjdk bug, or are we convinced this is the new Working As Designed?","from":"developer"},{"body":"[~jbellis] Yes, that is exactly what we have been seeing, as well. However, I'm not sure if our JVM is that 'recent' (/me weeps silently) - but, I'll try to confirm today across several of our clusters. ","from":"developer"},{"body":"[~jasobrown] fwiw, even though the diff is from openjdk 1.7, we've seen it as far back as jdk 1.6.0 update 45 thus far.","from":"developer"},{"body":"Yeah, a brute force check of some clusters confirms some are on 1.6.0_41, _37, and _45 - and, of course, they are mixed within the cluster, depending on when the instance was launched (if it was with an older AMI with older JVM).\n\nI haven't narrowed the problem down to which instances had troubles with which JVM, but some clusters with < u45 did not have any problems, so I suspect it's at least > u41 (our most highest version), but probably [~benedict] is correct with u45.","from":"developer"},{"body":"I attempted to file a bug with oracle, but the end result was a notification that I would be emailed with either a bug number or further questions. I haven't heard anything from them yet though. It's been a little less than a month.","from":"developer"},{"body":"This goes back at least as far as 1.6u43, just had someone hit it there.","from":"developer"},{"body":"We could try emailing dmitry.samersoff@oracle.com who contributed the change, and andrew.gross@oracle.com who reviewed it, directly?","from":"developer"},{"body":"FYI, I don't think I asserted this was u45. Only a specific changeset, which appears to have been in OpenJDK 7u12. It's not possible to see when this was included in Java 6, but could have been as early as u39 judging from release dates on Wikipedia.","from":"developer"},{"body":"Yeah, I was adding 1.6 versions where we've seen it, since we can't know when they put that into 1.6.","from":"developer"},{"body":"I have emailed dmitry and andrew, ccing cassandra-dev.","from":"developer"},{"body":"I emailed jmx-dev a few weeks ago, and it's been tumbleweed. I intend to follow up soon, and try the core-dev list (see if I can get any response from a higher traffic list), but I expect this will be slow going. Nobody seems interested.","from":"developer"},{"body":"Will {{-XX:+CMSClassUnloadingEnabled}} cause older VMs to bitch at us? If not let's add it as a default to 1.1+.","from":"developer"},{"body":"We already do java version detection in shell and add flags based on that, so adding this shouldn't be a problem.","from":"developer"},{"body":"Well, then someone needs to figure out when it was introduced. :)","from":"developer"},{"body":"I volunteer [~mshuler] :)","from":"developer"},{"body":"I get no complaints using -XX:+CMSClassUnloadingEnabled with 1.5.0_22-b03 and found an old doc that lists this VM option available in 1.4.1 1.4.2 1.5.0 1.6.0 - http://www.reins.altervista.org/java/A_Collection_of_JVM_Options_MP.html\nI'd say it's probably pretty safe to add it.","from":"developer"},{"body":"I committed this.","from":"developer"},{"body":"since I was there..\n$ /opt/jdk1.8.0/bin/java -XX:+CMSClassUnloadingEnabled -XX:+PrintFlagsFinal -version |grep CMSClassUnloadingEnabled\n bool CMSClassUnloadingEnabled := true {product}\njava version \"1.8.0\"\nJava(TM) SE Runtime Environment (build 1.8.0-b128)\nJava HotSpot(TM) 64-Bit Server VM (build 25.0-b69, mixed mode)","from":"developer"},{"body":"Is Cassandra \"since\" on this issue \"the dawn of time\"?","from":"developer"},{"body":"It's a HotSpot regression that we're working around, not a Cassandra bug.","from":"developer"},{"body":"Just adding some more java versions here for people wondering if they are hitting this. Just saw this on 1.6.0_38, so it goes back at least that far on 1.6","from":"developer"},{"body":"{quote}It's a HotSpot regression that we're working around, not a Cassandra bug.{quote}\nYes, so the statement \"this HotSpot regression is not-worked-around in all extant versions of Cassandra since the dawn of time\" is correct. Thanks for the clarification!","from":"developer"},{"body":"Well, not really, since the problematic JVM versions didn't exist back then.","from":"developer"},{"body":"bq. Yes, so the statement \"this HotSpot regression is not-worked-around in all extant versions of Cassandra since the dawn of time\" is correct. Thanks for the clarification!\n\nNo, it isn't, since the hot spot bug did not exist since the dawn of time. As such the from-version is ill-defined, but probably the most sensible definition requires determining the from-version-for-hotspot, which we have yet to manage, and aligning that with C* releases.","from":"developer"},{"body":"C* from version really has no meaning here. But yes, if you run C* 0.3 (if that did JMX monitoring) with 1.6_45 you will hit this issue. Unless of course you add the given setting to your cassandra-env.sh. Which can be done without upgrading your C*.","from":"developer"},{"body":"On DSE 4.5.1, Java 1.7, linux ami on aws encountered long running GC before dse was shutdown.This was a ring of 6 nodes, and no unexpected processes /load were run.The error log attached.\n\nI suspect to have hit that hotspot/java GC issue.But in this case the class unloading is enabled.\nJVM_OPTS=\"$JVM_OPTS -XX:+CMSClassUnloadingEnabled”\nXX:CMSInitiatingOccupancyFraction=75\n\nNow configured to have MAX_HEAP_SIZE=“8G”, HEAP_NEWSIZE=“2G” and GC logging enabled and is stable.However if there is a heap leak, it needs a better fix.\nSuggestion: If the CPU utilization is going up, or heap size leak/long running GC is sensed, the nodetool -h localhost flush can auto kick in?\n\nThanks\nRekha\n","from":"developer"},{"body":"Dse system log with long running GC before DSE shutdown","from":"developer"},{"body":"On newer systems with G1 GC is this option ignored? I'm assuming it would be.","from":"developer"}],"created":"2014-01-02T20:38:45.000+0000","description":"Newer versions of Oracle's Hotspot JVM , post 6u43 (maybe earlier) and 7u25 (maybe earlier), are experiencing issues with GC and JMX where heap slowly fills up overtime until OOM or a full GC event occurs, specifically when CMS is leveraged. Adding:\n\n{noformat}\nJVM_OPTS=\"$JVM_OPTS -XX:+CMSClassUnloadingEnabled\"\n{noformat}\n\nThe the options in cassandra-env.sh alleviates the problem.","issue_id":"12687035","key":"CASSANDRA-6541","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-03-26T20:50:28.000+0000","role":"fixed_distractor","summary":"New versions of Hotspot create new Class objects on every JMX connection causing the heap to fill up with them if CMSClassUnloadingEnabled isn't set."} {"case_id":"12697419","cluster":"DISTRACTOR-CASSANDRA-6774","comments":[{"body":"Additional info:\n\ncqlsh:system> select * from compactions_in_progress limit 10;\n\n id | columnfamily_name | inputs | keyspace_name\n--------------------------------------+------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+---------------\n 9a20d760-9f15-11e3-bfe7-fb46089edbb0 | global_user | {16796, 17068, 17073, 17077, 17118, 17135, 17164, 17193, 17219, 17281, 17293, 17365, 17553, 17575, 17606, 17668, 17695, 17701, 17906, 17983, 18016, 18017, 18069, 18089, 18098, 18099, 18108, 18114, 18119, 18123, 18188, 18449} | users\n 93778f70-9f16-11e3-bfe7-fb46089edbb0 | client_end_user_lookup | {3280, 3281, 3283, 3284, 3286, 3293, 3297, 3298, 3301, 3302, 3303, 3306, 3323, 3324, 3326, 3328, 3329, 3330, 3331, 3332, 3333, 3334} | users\n 96d23e10-9f14-11e3-bfe7-fb46089edbb0 | cookie_user_lookup | {9072, 9078, 9083, 9088, 9094, 9100, 9106, 9112, 9118, 9124, 9131, 9146, 9369, 9370, 9371, 9372, 9373, 9393} | users\n\n(3 rows)","created":"2014-02-26T18:54:21.713+0000"},{"body":"Looks like restarting the node corrected the issue thus far as cleanup is running.","created":"2014-02-26T19:14:04.996+0000"},{"body":"Same issue on LCS, but restarting doesn't help. The only way to remove data after decreasing RF is Cleanup, but it fails with the same exception.\nIs there a way to temporarily disable compactions on LCS?","created":"2014-03-12T08:40:56.019+0000"},{"body":"Same issue here on LCS - after decreasing RF cleanup fails with the same stacktrace. Restart didn't help, after running \"nodetool disableautocompaction\" command and waiting for compaction get finished cleanup fails with the following stacktrace:\n{quote}\nError occurred during cleanup\njava.util.concurrent.ExecutionException: java.lang.IndexOutOfBoundsException: Index: 1, Size: 1\n at java.util.concurrent.FutureTask.report(FutureTask.java:122)\n at java.util.concurrent.FutureTask.get(FutureTask.java:188)\n at org.apache.cassandra.db.compaction.CompactionManager.performAllSSTableOperation(CompactionManager.java:227)\n at org.apache.cassandra.db.compaction.CompactionManager.performCleanup(CompactionManager.java:265)\n at org.apache.cassandra.db.ColumnFamilyStore.forceCleanup(ColumnFamilyStore.java:1115)\n at org.apache.cassandra.service.StorageService.forceKeyspaceCleanup(StorageService.java:2152)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:606)\n at sun.reflect.misc.Trampoline.invoke(MethodUtil.java:75)\n at sun.reflect.GeneratedMethodAccessor12.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:606)\n at sun.reflect.misc.MethodUtil.invoke(MethodUtil.java:279)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:112)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:46)\n at com.sun.jmx.mbeanserver.MBeanIntrospector.invokeM(MBeanIntrospector.java:237)\n at com.sun.jmx.mbeanserver.PerInterface.invoke(PerInterface.java:138)\n at com.sun.jmx.mbeanserver.MBeanSupport.invoke(MBeanSupport.java:252)\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.invoke(DefaultMBeanServerInterceptor.java:819)\n at com.sun.jmx.mbeanserver.JmxMBeanServer.invoke(JmxMBeanServer.java:801)\n at javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1487)\n at javax.management.remote.rmi.RMIConnectionImpl.access$300(RMIConnectionImpl.java:97)\n at javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1328)\n at javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1420)\n at javax.management.remote.rmi.RMIConnectionImpl.invoke(RMIConnectionImpl.java:848)\n at sun.reflect.GeneratedMethodAccessor20.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:606)\n at sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:322)\n at sun.rmi.transport.Transport$1.run(Transport.java:177)\n at sun.rmi.transport.Transport$1.run(Transport.java:174)\n at java.security.AccessController.doPrivileged(Native Method)\n at sun.rmi.transport.Transport.serviceCall(Transport.java:173)\n at sun.rmi.transport.tcp.TCPTransport.handleMessages(TCPTransport.java:556)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run0(TCPTransport.java:811)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run(TCPTransport.java:670)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:744)\nCaused by: java.lang.IndexOutOfBoundsException: Index: 1, Size: 1\n at java.util.ArrayList.rangeCheck(ArrayList.java:635)\n at java.util.ArrayList.get(ArrayList.java:411)\n at org.apache.cassandra.db.compaction.CompactionManager.needsCleanup(CompactionManager.java:502)\n at org.apache.cassandra.db.compaction.CompactionManager.doCleanupCompaction(CompactionManager.java:540)\n at org.apache.cassandra.db.compaction.CompactionManager.access$400(CompactionManager.java:62)\n at org.apache.cassandra.db.compaction.CompactionManager$5.perform(CompactionManager.java:274)\n at org.apache.cassandra.db.compaction.CompactionManager$2.call(CompactionManager.java:222)\n at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n ... 3 more\n{quote}\n\n*Update:*\nIndexOutOfBoundsException exception issue has been fixed: [CASSANDRA-6845|https://issues.apache.org/jira/browse/CASSANDRA-6845]","created":"2014-03-12T09:11:43.288+0000"},{"body":"Looks like the cause here is that the node is so heavily overloaded that cassandra could not cancel in-progress compactions quickly enough, the exception is harmless, but perhaps not very nice. \n\nAttaching patch to not continue with the cleanup (or, all operations that need all sstables marked as compacting) after failing to cancel in-progress compactions.\n\nThe IndexOutOfBoundsException error is unrelated and fixed in CASSANDRA-6845 as mentioned in a previous comment.","created":"2014-03-13T08:25:44.165+0000"},{"body":"I don't think we want to drop an exception into the log for what is to some degree an expected situation. v2 refactors scrub/upgrade/cleanup to log an info message if they have to abort, instead. (Also runs unmarkCompacting in a finally block, which was not the case before.)","created":"2014-03-14T18:05:02.392+0000"},{"body":"hmm, i actually think we should throw an exception if we fail, so that the nodetool user gets some feedback without having to check logs, attached","created":"2014-03-17T08:25:51.219+0000"},{"body":"or.. maybe we should return a status code from scrub/upgradesstables/cleanup instead... i'll fix","created":"2014-03-17T08:35:26.479+0000"},{"body":"Attaching patch that;\n* Introduces an enum with the status for the operation (currently only aborted/successful)\n* Makes cfs.markAllCompacting() return empty list if there are no sstables to execute the operation on and null if we failed cancelling\n* Make nodetool output an error message and set exit code to 1 if we fail","created":"2014-03-17T12:12:51.072+0000"},{"body":"LGTM.\n\nAre you still comfortable with 2.0 for this or should we do it in 2.1?","created":"2014-03-17T13:16:53.455+0000"},{"body":"hmm i guess we shouldn't change the signatures of the exposed mbean methods in a minor rev\n\nlets go 2.1 and leave 2.0 with the assertion, as it is mostly a cosmetic change\n\n","created":"2014-03-17T18:03:24.502+0000"},{"body":"committed","created":"2014-03-18T08:48:53.158+0000"}],"conversations":[{"body":"I am stress testing a new 2.0.5 cluster and did the following:\n\n- start decommission during heavy write, moderate read load\n- trigger cleanup on non-decommissioning node (nodetool cleanup)\n- Started to see higher GC load stop stopped cleanup via nodetool stop CLEANUP\n- attempt to launch cleanup now fails with the following message in console.\n\nCassandra log shows: http://aep.appspot.com/display/cKmlMcDuKD72iYAcBykDuVZkRWY/","from":"reporter","subject":"Cleanup fails with assertion error after stopping previous run"},{"body":"Additional info:\n\ncqlsh:system> select * from compactions_in_progress limit 10;\n\n id | columnfamily_name | inputs | keyspace_name\n--------------------------------------+------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+---------------\n 9a20d760-9f15-11e3-bfe7-fb46089edbb0 | global_user | {16796, 17068, 17073, 17077, 17118, 17135, 17164, 17193, 17219, 17281, 17293, 17365, 17553, 17575, 17606, 17668, 17695, 17701, 17906, 17983, 18016, 18017, 18069, 18089, 18098, 18099, 18108, 18114, 18119, 18123, 18188, 18449} | users\n 93778f70-9f16-11e3-bfe7-fb46089edbb0 | client_end_user_lookup | {3280, 3281, 3283, 3284, 3286, 3293, 3297, 3298, 3301, 3302, 3303, 3306, 3323, 3324, 3326, 3328, 3329, 3330, 3331, 3332, 3333, 3334} | users\n 96d23e10-9f14-11e3-bfe7-fb46089edbb0 | cookie_user_lookup | {9072, 9078, 9083, 9088, 9094, 9100, 9106, 9112, 9118, 9124, 9131, 9146, 9369, 9370, 9371, 9372, 9373, 9393} | users\n\n(3 rows)","from":"developer"},{"body":"Looks like restarting the node corrected the issue thus far as cleanup is running.","from":"developer"},{"body":"Same issue on LCS, but restarting doesn't help. The only way to remove data after decreasing RF is Cleanup, but it fails with the same exception.\nIs there a way to temporarily disable compactions on LCS?","from":"developer"},{"body":"Same issue here on LCS - after decreasing RF cleanup fails with the same stacktrace. Restart didn't help, after running \"nodetool disableautocompaction\" command and waiting for compaction get finished cleanup fails with the following stacktrace:\n{quote}\nError occurred during cleanup\njava.util.concurrent.ExecutionException: java.lang.IndexOutOfBoundsException: Index: 1, Size: 1\n at java.util.concurrent.FutureTask.report(FutureTask.java:122)\n at java.util.concurrent.FutureTask.get(FutureTask.java:188)\n at org.apache.cassandra.db.compaction.CompactionManager.performAllSSTableOperation(CompactionManager.java:227)\n at org.apache.cassandra.db.compaction.CompactionManager.performCleanup(CompactionManager.java:265)\n at org.apache.cassandra.db.ColumnFamilyStore.forceCleanup(ColumnFamilyStore.java:1115)\n at org.apache.cassandra.service.StorageService.forceKeyspaceCleanup(StorageService.java:2152)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:606)\n at sun.reflect.misc.Trampoline.invoke(MethodUtil.java:75)\n at sun.reflect.GeneratedMethodAccessor12.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:606)\n at sun.reflect.misc.MethodUtil.invoke(MethodUtil.java:279)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:112)\n at com.sun.jmx.mbeanserver.StandardMBeanIntrospector.invokeM2(StandardMBeanIntrospector.java:46)\n at com.sun.jmx.mbeanserver.MBeanIntrospector.invokeM(MBeanIntrospector.java:237)\n at com.sun.jmx.mbeanserver.PerInterface.invoke(PerInterface.java:138)\n at com.sun.jmx.mbeanserver.MBeanSupport.invoke(MBeanSupport.java:252)\n at com.sun.jmx.interceptor.DefaultMBeanServerInterceptor.invoke(DefaultMBeanServerInterceptor.java:819)\n at com.sun.jmx.mbeanserver.JmxMBeanServer.invoke(JmxMBeanServer.java:801)\n at javax.management.remote.rmi.RMIConnectionImpl.doOperation(RMIConnectionImpl.java:1487)\n at javax.management.remote.rmi.RMIConnectionImpl.access$300(RMIConnectionImpl.java:97)\n at javax.management.remote.rmi.RMIConnectionImpl$PrivilegedOperation.run(RMIConnectionImpl.java:1328)\n at javax.management.remote.rmi.RMIConnectionImpl.doPrivilegedOperation(RMIConnectionImpl.java:1420)\n at javax.management.remote.rmi.RMIConnectionImpl.invoke(RMIConnectionImpl.java:848)\n at sun.reflect.GeneratedMethodAccessor20.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:606)\n at sun.rmi.server.UnicastServerRef.dispatch(UnicastServerRef.java:322)\n at sun.rmi.transport.Transport$1.run(Transport.java:177)\n at sun.rmi.transport.Transport$1.run(Transport.java:174)\n at java.security.AccessController.doPrivileged(Native Method)\n at sun.rmi.transport.Transport.serviceCall(Transport.java:173)\n at sun.rmi.transport.tcp.TCPTransport.handleMessages(TCPTransport.java:556)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run0(TCPTransport.java:811)\n at sun.rmi.transport.tcp.TCPTransport$ConnectionHandler.run(TCPTransport.java:670)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:744)\nCaused by: java.lang.IndexOutOfBoundsException: Index: 1, Size: 1\n at java.util.ArrayList.rangeCheck(ArrayList.java:635)\n at java.util.ArrayList.get(ArrayList.java:411)\n at org.apache.cassandra.db.compaction.CompactionManager.needsCleanup(CompactionManager.java:502)\n at org.apache.cassandra.db.compaction.CompactionManager.doCleanupCompaction(CompactionManager.java:540)\n at org.apache.cassandra.db.compaction.CompactionManager.access$400(CompactionManager.java:62)\n at org.apache.cassandra.db.compaction.CompactionManager$5.perform(CompactionManager.java:274)\n at org.apache.cassandra.db.compaction.CompactionManager$2.call(CompactionManager.java:222)\n at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n ... 3 more\n{quote}\n\n*Update:*\nIndexOutOfBoundsException exception issue has been fixed: [CASSANDRA-6845|https://issues.apache.org/jira/browse/CASSANDRA-6845]","from":"developer"},{"body":"Looks like the cause here is that the node is so heavily overloaded that cassandra could not cancel in-progress compactions quickly enough, the exception is harmless, but perhaps not very nice. \n\nAttaching patch to not continue with the cleanup (or, all operations that need all sstables marked as compacting) after failing to cancel in-progress compactions.\n\nThe IndexOutOfBoundsException error is unrelated and fixed in CASSANDRA-6845 as mentioned in a previous comment.","from":"developer"},{"body":"I don't think we want to drop an exception into the log for what is to some degree an expected situation. v2 refactors scrub/upgrade/cleanup to log an info message if they have to abort, instead. (Also runs unmarkCompacting in a finally block, which was not the case before.)","from":"developer"},{"body":"hmm, i actually think we should throw an exception if we fail, so that the nodetool user gets some feedback without having to check logs, attached","from":"developer"},{"body":"or.. maybe we should return a status code from scrub/upgradesstables/cleanup instead... i'll fix","from":"developer"},{"body":"Attaching patch that;\n* Introduces an enum with the status for the operation (currently only aborted/successful)\n* Makes cfs.markAllCompacting() return empty list if there are no sstables to execute the operation on and null if we failed cancelling\n* Make nodetool output an error message and set exit code to 1 if we fail","from":"developer"},{"body":"LGTM.\n\nAre you still comfortable with 2.0 for this or should we do it in 2.1?","from":"developer"},{"body":"hmm i guess we shouldn't change the signatures of the exposed mbean methods in a minor rev\n\nlets go 2.1 and leave 2.0 with the assertion, as it is mostly a cosmetic change\n\n","from":"developer"},{"body":"committed","from":"developer"}],"created":"2014-02-26T18:50:49.000+0000","description":"I am stress testing a new 2.0.5 cluster and did the following:\n\n- start decommission during heavy write, moderate read load\n- trigger cleanup on non-decommissioning node (nodetool cleanup)\n- Started to see higher GC load stop stopped cleanup via nodetool stop CLEANUP\n- attempt to launch cleanup now fails with the following message in console.\n\nCassandra log shows: http://aep.appspot.com/display/cKmlMcDuKD72iYAcBykDuVZkRWY/","issue_id":"12697419","key":"CASSANDRA-6774","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-03-18T08:48:53.000+0000","role":"fixed_distractor","summary":"Cleanup fails with assertion error after stopping previous run"} {"case_id":"12707157","cluster":"DISTRACTOR-CASSANDRA-6998","comments":[{"body":"See CASSANDRA-6666","created":"2014-04-08T13:02:21.822+0000"},{"body":"Thx [~iamaleksey] for your quick comment, I think this problem is not solved yet though. In the comments section of the issue referenced by you https://issues.apache.org/jira/browse/CASSANDRA-6666 [~gsanderson] asks:\n{quote}\nOne quick question... this patch will hopefully prevent us getting into this state in many cases, unless you have a huge number of hints for a node that is down for a very long time: since there is no auto-compaction, a large number of hints may have expired from TTL, thus preventing any further hint delivery.\n{quote}\n\nThat's the case I'm referring to. I got this Is it covered by 6666? If not, is it considered as a bug? I added stacktrace to the issue.\n","created":"2014-04-08T14:04:45.362+0000"},{"body":"Deleting undelivered tombstones can introduce problems with consistency. I don't think a hinted handoff row should ever be deleted before delivery even if it was a tombstone.\n\nThe exception is thrown when the number of deleted columns in the query result exceed the tombstonefailurethreshold.\nThis can be solved using in two ways:\n1-Instead of just removing the tombstones by compacting the table, we can decrease the page size when catching the exception and then try to perform the hinted handoff again.This will decrease the number of deleted columns below the tombstonefailurethreshold.\n2-a better design choice in my opinion is to use another type of queries, one that gets all rows with a given size regardless of being a live row or not.SliceQueries are not the suitable type for this task since hinted handoff deliver doesn't care for the status of the given row.","created":"2014-04-14T15:53:31.032+0000"},{"body":" Instead of ignoring tombstones when counting returned page size, and instead of ignoring irrelevant data, deliver both relevant and Irrelevant columns. This will increase data consistency when the recovering node was down for a time longer than the gc grace period.\n\nTo avoid performance problems, when choosing to return both relevant and irrelevant data, the slice query filter will count live, tombstones and irrelevant columns, compared to counting live columns only in the ordinary case. The check for tombstone threshold is ignored in this specific case.\n\nPatch submitted.","created":"2014-04-24T09:02:23.206+0000"},{"body":"Sorry, changed status to TESTING by error. Can someone of the moderators return the status back to Patch Available? there is no undo functionality","created":"2014-04-24T22:16:40.916+0000"},{"body":"I'm confused, both because I'm not sure how \"get all rows with a given size\" would fix the problem, and because this patch does not seem to match that proposal.","created":"2014-05-08T22:12:49.768+0000"},{"body":"Sorry if I didn't make myself clear enough. I will try to explain what this patch does:\n\n1-Normally, the page size passed to SliceQueryFilter represents number of live columns returned. The SliceQueryFilter returns a mix of live columns and tombstones which are not eligible for garbage collection yet. If the number of tombstones seen during constructing the query result exceeds the threshold, an exception is thrown. \n\n2-When retrieving hinted handoff data, use the SliceQueryFilter in a different way. The page size will represent the total number of columns returned: live columns count + tombstones count. Even tombstones that are eligible for garbage collection are returned. The check for the tombstones threshold is removed. No memory overload will take place since the size of data processed and returned is controlled and predictable (exactly the page size).\n\n3-This will send all data in the hinted handoff table to the recovering node allowing for higher data consistency even if the node was down for a time larger than the gc_grace period. On the other hand it can cause a higher network load.\n\nThis is a suggested way to deal with the issue of consistency of deleted hinted handoff data. It is not just an addressing for the exception problem, because I think the exception problem is a symptom for the more delicate issue of deleted hinted handoff. ","created":"2014-05-10T10:19:36.203+0000"},{"body":"I see.\n\nThat would fix the problem, but I'm not really a fan of adding extra layers of complexity to mask symptoms of a deeper underlying problem (CASSANDRA-6666). Let's first figure out *why* these tombstones aren't going away the way we expect them to; it's likely that we can then solve the problem without breaking the SQF contract (and then trying to make up for that at a higher level).","created":"2014-05-12T19:47:47.263+0000"},{"body":"Attaching a very simple v2 patch that fixes this issue, and fixed CASSANDRA-6666 in an arguably better way, too.\n\nCurrently, a large number of TTLd hints can really block future hint deliveries, b/c the mechanism of CASSANDRA-6666 won't even be able to kick in.\n\nThe attached patch moves the major compaction step in front of delivery, instead of afterwards, thus guaranteeing that we will not have a huge number of expired hints when we go replaying (and also cleans up delivered tombstoned hints).\n\nAs a nice side effect, we are saving *a lot* of compactions for the scheduled 'deliver all the hints' 10 minutes process, by having a single major compaction done for all the N nodes we are going to deliver hints to, instead of having a compaction at the end of each of those N deliveries.\n\nThe patch is small and non-invasive enough to go into 2.0.x.","created":"2014-10-07T00:25:49.143+0000"},{"body":"LGTM, +1","created":"2014-10-15T16:09:41.953+0000"},{"body":"Committed, thanks.","created":"2014-10-16T16:23:34.753+0000"}],"conversations":[{"body":"For tests purposes, DC2 was shut down for 1 day. The _hints_ table was filled with millions of rows. Now, when _HintedHandOffManager_ tries to _doDeliverHintsToEndpoint_ it queries the store with QueryFilter.getSliceFilter which counts deleted (TTLed) cells and throws org.apache.cassandra.db.filter.TombstoneOverwhelmingException. \nThrowing this exception stops the manager from running compaction as it is run only after successful handoff. This leaves the HH practically disabled till administrator runs truncateAllHints. \nWouldn't it be nicer if on org.apache.cassandra.db.filter.TombstoneOverwhelmingException run compaction? That would remove TTLed hints leaving whole HH mechanism in a healthy state.\n\nThe stacktrace is:\n{quote}\norg.apache.cassandra.db.filter.TombstoneOverwhelmingException\n\tat org.apache.cassandra.db.filter.SliceQueryFilter.collectReducedColumns(SliceQueryFilter.java:201)\n\tat org.apache.cassandra.db.filter.QueryFilter.collateColumns(QueryFilter.java:122)\n\tat org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:80)\n\tat org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:72)\n\tat org.apache.cassandra.db.CollationController.collectAllData(CollationController.java:297)\n\tat org.apache.cassandra.db.CollationController.getTopLevelColumns(CollationController.java:53)\n\tat org.apache.cassandra.db.ColumnFamilyStore.getTopLevelColumns(ColumnFamilyStore.java:1487)\n\tat org.apache.cassandra.db.ColumnFamilyStore.getColumnFamily(ColumnFamilyStore.java:1306)\n\tat org.apache.cassandra.db.HintedHandOffManager.doDeliverHintsToEndpoint(HintedHandOffManager.java:351)\n\tat org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:309)\n\tat org.apache.cassandra.db.HintedHandOffManager.access$300(HintedHandOffManager.java:92)\n\tat org.apache.cassandra.db.HintedHandOffManager$4.run(HintedHandOffManager.java:530)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n\tat java.lang.Thread.run(Thread.java:722)\n{quote}\n","from":"reporter","subject":"HintedHandoff - expired hints may block future hints deliveries"},{"body":"See CASSANDRA-6666","from":"developer"},{"body":"Thx [~iamaleksey] for your quick comment, I think this problem is not solved yet though. In the comments section of the issue referenced by you https://issues.apache.org/jira/browse/CASSANDRA-6666 [~gsanderson] asks:\n{quote}\nOne quick question... this patch will hopefully prevent us getting into this state in many cases, unless you have a huge number of hints for a node that is down for a very long time: since there is no auto-compaction, a large number of hints may have expired from TTL, thus preventing any further hint delivery.\n{quote}\n\nThat's the case I'm referring to. I got this Is it covered by 6666? If not, is it considered as a bug? I added stacktrace to the issue.\n","from":"developer"},{"body":"Deleting undelivered tombstones can introduce problems with consistency. I don't think a hinted handoff row should ever be deleted before delivery even if it was a tombstone.\n\nThe exception is thrown when the number of deleted columns in the query result exceed the tombstonefailurethreshold.\nThis can be solved using in two ways:\n1-Instead of just removing the tombstones by compacting the table, we can decrease the page size when catching the exception and then try to perform the hinted handoff again.This will decrease the number of deleted columns below the tombstonefailurethreshold.\n2-a better design choice in my opinion is to use another type of queries, one that gets all rows with a given size regardless of being a live row or not.SliceQueries are not the suitable type for this task since hinted handoff deliver doesn't care for the status of the given row.","from":"developer"},{"body":" Instead of ignoring tombstones when counting returned page size, and instead of ignoring irrelevant data, deliver both relevant and Irrelevant columns. This will increase data consistency when the recovering node was down for a time longer than the gc grace period.\n\nTo avoid performance problems, when choosing to return both relevant and irrelevant data, the slice query filter will count live, tombstones and irrelevant columns, compared to counting live columns only in the ordinary case. The check for tombstone threshold is ignored in this specific case.\n\nPatch submitted.","from":"developer"},{"body":"Sorry, changed status to TESTING by error. Can someone of the moderators return the status back to Patch Available? there is no undo functionality","from":"developer"},{"body":"I'm confused, both because I'm not sure how \"get all rows with a given size\" would fix the problem, and because this patch does not seem to match that proposal.","from":"developer"},{"body":"Sorry if I didn't make myself clear enough. I will try to explain what this patch does:\n\n1-Normally, the page size passed to SliceQueryFilter represents number of live columns returned. The SliceQueryFilter returns a mix of live columns and tombstones which are not eligible for garbage collection yet. If the number of tombstones seen during constructing the query result exceeds the threshold, an exception is thrown. \n\n2-When retrieving hinted handoff data, use the SliceQueryFilter in a different way. The page size will represent the total number of columns returned: live columns count + tombstones count. Even tombstones that are eligible for garbage collection are returned. The check for the tombstones threshold is removed. No memory overload will take place since the size of data processed and returned is controlled and predictable (exactly the page size).\n\n3-This will send all data in the hinted handoff table to the recovering node allowing for higher data consistency even if the node was down for a time larger than the gc_grace period. On the other hand it can cause a higher network load.\n\nThis is a suggested way to deal with the issue of consistency of deleted hinted handoff data. It is not just an addressing for the exception problem, because I think the exception problem is a symptom for the more delicate issue of deleted hinted handoff. ","from":"developer"},{"body":"I see.\n\nThat would fix the problem, but I'm not really a fan of adding extra layers of complexity to mask symptoms of a deeper underlying problem (CASSANDRA-6666). Let's first figure out *why* these tombstones aren't going away the way we expect them to; it's likely that we can then solve the problem without breaking the SQF contract (and then trying to make up for that at a higher level).","from":"developer"},{"body":"Attaching a very simple v2 patch that fixes this issue, and fixed CASSANDRA-6666 in an arguably better way, too.\n\nCurrently, a large number of TTLd hints can really block future hint deliveries, b/c the mechanism of CASSANDRA-6666 won't even be able to kick in.\n\nThe attached patch moves the major compaction step in front of delivery, instead of afterwards, thus guaranteeing that we will not have a huge number of expired hints when we go replaying (and also cleans up delivered tombstoned hints).\n\nAs a nice side effect, we are saving *a lot* of compactions for the scheduled 'deliver all the hints' 10 minutes process, by having a single major compaction done for all the N nodes we are going to deliver hints to, instead of having a compaction at the end of each of those N deliveries.\n\nThe patch is small and non-invasive enough to go into 2.0.x.","from":"developer"},{"body":"LGTM, +1","from":"developer"},{"body":"Committed, thanks.","from":"developer"}],"created":"2014-04-08T11:16:31.000+0000","description":"For tests purposes, DC2 was shut down for 1 day. The _hints_ table was filled with millions of rows. Now, when _HintedHandOffManager_ tries to _doDeliverHintsToEndpoint_ it queries the store with QueryFilter.getSliceFilter which counts deleted (TTLed) cells and throws org.apache.cassandra.db.filter.TombstoneOverwhelmingException. \nThrowing this exception stops the manager from running compaction as it is run only after successful handoff. This leaves the HH practically disabled till administrator runs truncateAllHints. \nWouldn't it be nicer if on org.apache.cassandra.db.filter.TombstoneOverwhelmingException run compaction? That would remove TTLed hints leaving whole HH mechanism in a healthy state.\n\nThe stacktrace is:\n{quote}\norg.apache.cassandra.db.filter.TombstoneOverwhelmingException\n\tat org.apache.cassandra.db.filter.SliceQueryFilter.collectReducedColumns(SliceQueryFilter.java:201)\n\tat org.apache.cassandra.db.filter.QueryFilter.collateColumns(QueryFilter.java:122)\n\tat org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:80)\n\tat org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:72)\n\tat org.apache.cassandra.db.CollationController.collectAllData(CollationController.java:297)\n\tat org.apache.cassandra.db.CollationController.getTopLevelColumns(CollationController.java:53)\n\tat org.apache.cassandra.db.ColumnFamilyStore.getTopLevelColumns(ColumnFamilyStore.java:1487)\n\tat org.apache.cassandra.db.ColumnFamilyStore.getColumnFamily(ColumnFamilyStore.java:1306)\n\tat org.apache.cassandra.db.HintedHandOffManager.doDeliverHintsToEndpoint(HintedHandOffManager.java:351)\n\tat org.apache.cassandra.db.HintedHandOffManager.deliverHintsToEndpoint(HintedHandOffManager.java:309)\n\tat org.apache.cassandra.db.HintedHandOffManager.access$300(HintedHandOffManager.java:92)\n\tat org.apache.cassandra.db.HintedHandOffManager$4.run(HintedHandOffManager.java:530)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n\tat java.lang.Thread.run(Thread.java:722)\n{quote}\n","issue_id":"12707157","key":"CASSANDRA-6998","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-10-16T16:23:34.000+0000","role":"fixed_distractor","summary":"HintedHandoff - expired hints may block future hints deliveries"} {"case_id":"12708554","cluster":"DISTRACTOR-CASSANDRA-7042","comments":[{"body":"What are you doing before the exception? Dropping tables? Running repair? Truncating?","created":"2014-04-22T04:34:54.420+0000"},{"body":"Doing nothing they just show up in the logs during normal use our tables are all ttl'd set to about a day with around 30GB or so of data churn. Upgraded to 2.0.7 and still having the issue, here is a error for the 2.0.7 version not even sure if the error is related to the space growth but it is the only bit of info in the logs that might suggest it. Willing to help you guys out anyway I can would love to get this fixed its pretty much our last major issue with keeping our cluster stable.\n\nERROR [CompactionExecutor:112] 2014-04-21 21:26:35,109 CassandraDaemon.java (line 198) Exception in thread Thread[CompactionExecutor:112,1,main]\njava.lang.RuntimeException: java.io.FileNotFoundException: /local-project/cassandra_data/data/wxgrid/grid/wxgrid-grid-jb-580812-Data.db (No such file or directory)\n at org.apache.cassandra.io.util.ThrottledReader.open(ThrottledReader.java:53)\n at org.apache.cassandra.io.sstable.SSTableReader.openDataReader(SSTableReader.java:1355)\n at org.apache.cassandra.io.sstable.SSTableScanner.(SSTableScanner.java:67)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1161)","created":"2014-04-22T04:56:10.484+0000"},{"body":"Here is some screen shots for datastax opscenter. The dips are related to either running cleanup or changing the gc_grace period down after learning that since we ttl all our data we don't have to run repairs within gc_grace. You can see the spikes in the total disk space they will basicly continue until it runs our of space or we restart cassandra.","created":"2014-04-22T05:01:13.074+0000"},{"body":"This is a dump of the directory for the cf that has the issues as you can see file count increases. It seems like files are not being deleted when they should be. The before file is taken when cassandra was in a stopped state the after is when cassandra had been started up again.","created":"2014-04-24T18:36:34.011+0000"},{"body":"Here is a schema in question\n\nCREATE TABLE grid (\n data_id text,\n cylinder text,\n value blob,\n PRIMARY KEY (data_id, cylinder)\n) WITH COMPACT STORAGE AND\n bloom_filter_fp_chance=0.100000 AND\n caching='ALL' AND\n comment='' AND\n dclocal_read_repair_chance=0.000000 AND\n gc_grace_seconds=0 AND\n index_interval=128 AND\n read_repair_chance=0.100000 AND\n populate_io_cache_on_flush='false' AND\n default_time_to_live=0 AND\n speculative_retry='99.0PERCENTILE' AND\n memtable_flush_period_in_ms=0 AND\n compaction={'sstable_size_in_mb': '160', 'class': 'LeveledCompactionStrategy'} AND\n compression={};","created":"2014-04-24T18:50:35.491+0000"},{"body":"Just thought I would also mention setup a brand new 3 node test cluster completely fresh and was able to reproduce.","created":"2014-04-28T20:09:35.699+0000"},{"body":"While trying to reproduce CASSANDRA-6525, I'm seeing similar stacktraces in the logs:\n\n{noformat}\nERROR 13:45:24,236 Exception in thread Thread[CompactionExecutor:7,1,main]\njava.lang.RuntimeException: java.io.FileNotFoundException: /var/lib/cassandra/data/test6981/t12/test6981-t12.t12_blob5_idx-jb-5-Data.db (No such file or directory)\n at com.google.common.base.Throwables.propagate(Throwables.java:160)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:32)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:60)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:59)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:197)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:744)\nCaused by: java.io.FileNotFoundException: /var/lib/cassandra/data/test6981/t12/test6981-t12.t12_blob5_idx-jb-5-Data.db (No such file or directory)\n at java.io.RandomAccessFile.open(Native Method)\n at java.io.RandomAccessFile.(RandomAccessFile.java:241)\n at java.io.RandomAccessFile.(RandomAccessFile.java:122)\n at org.apache.cassandra.io.sstable.SSTableReader.preheat(SSTableReader.java:820)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:245)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n ... 8 more\n{noformat}\n\nand\n\n{noformat}\nERROR 13:46:42,182 Exception in thread Thread[CompactionExecutor:5,1,main]\njava.lang.RuntimeException: java.io.FileNotFoundException: /var/lib/cassandra/data/test6981/t4/test6981-t4-jb-4-Data.db (No such file or directory)\n at org.apache.cassandra.io.compress.CompressedThrottledReader.open(CompressedThrottledReader.java:52)\n at org.apache.cassandra.io.sstable.SSTableReader.openDataReader(SSTableReader.java:1355)\n at org.apache.cassandra.io.sstable.SSTableScanner.(SSTableScanner.java:67)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1161)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1173)\n at org.apache.cassandra.db.compaction.AbstractCompactionStrategy.getScanners(AbstractCompactionStrategy.java:262)\n at org.apache.cassandra.db.compaction.AbstractCompactionStrategy.getScanners(AbstractCompactionStrategy.java:268)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:126)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:60)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:59)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:197)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:744)\nCaused by: java.io.FileNotFoundException: /var/lib/cassandra/data/test6981/t4/test6981-t4-jb-4-Data.db (No such file or directory)\n at java.io.RandomAccessFile.open(Native Method)\n at java.io.RandomAccessFile.(RandomAccessFile.java:241)\n at org.apache.cassandra.io.util.RandomAccessReader.(RandomAccessReader.java:58)\n at org.apache.cassandra.io.compress.CompressedRandomAccessReader.(CompressedRandomAccessReader.java:76)\n at org.apache.cassandra.io.compress.CompressedThrottledReader.(CompressedThrottledReader.java:34)\n at org.apache.cassandra.io.compress.CompressedThrottledReader.open(CompressedThrottledReader.java:48)\n ... 17 more\n{noformat}\n\nI didn't see any problems when running the 6525 repro script on my laptop, but in a VM with only 1GB of RAM, I got quite a few of these errors.","created":"2014-05-09T18:56:32.462+0000"},{"body":"Interesting, we are working on a script to reproduce that dose not require a bunch of dependancies and our whole infrastructure setup as a side note just recently deleted a key space and when looking at the data directories there where a bunch of files left behind on a few nodes that exhibited the issue.","created":"2014-05-14T20:03:22.903+0000"},{"body":"This is possibly CASSANDRA-7511. Are you truncating ever? Is your commit log or your sstable data what is taking up the space? The fact that it gets fixed on restart makes me think commit log.","created":"2014-08-21T17:02:49.044+0000"},{"body":"No, we never do a truncate, and our commit log folder is around 1GB in size. When we restart the data directory is what shrinks in size.","created":"2014-08-21T17:09:01.202+0000"},{"body":"I'm trying to figure out the file not found exceptions in CASSANDRA-7145 - don't think they are related to the disk space growth though, so not marking this as a duplicate (yet)","created":"2014-08-22T08:41:37.727+0000"},{"body":"At a guess, this is CASSANDRA-7139. Try raising your compaction throughput limit and/or lowering your concurrent_compactors (both in the cassandra.yaml)\n\n","created":"2014-08-22T08:59:27.220+0000"},{"body":"I've filed CASSANDRA-7819 with a more general fix, in case this turns out to be something different. ","created":"2014-08-22T09:36:20.447+0000"},{"body":"Our compaction limit is set to 0 and compaction memory 768MB. Concurrent compactors is set to the default and multithreading is off and compaction pre heat cache is also off. I am going to be getting a debug log today hopefully with this (log4j.logger.org.apache.cassandra.db.compaction=DEBUG) so we will see what the reveals.","created":"2014-08-22T13:08:16.297+0000"},{"body":"The default concurrent_compactors is most likely too high for machines with many CPUs and modest disk throughput. What disk layout do you have?","created":"2014-08-22T13:52:28.892+0000"},{"body":"We are using amazon ec2 i2.xlarge so one 800GB SSD","created":"2014-08-22T13:55:58.426+0000"},{"body":"Ok, so that box has a default of 32 concurrent compactors on 2.0 (on 2.1 it will default to 2). This is almost certainly too many, and I think this is highly likely to be your problem. What would be interesting is to try the patch I have posted in CASSANDRA-7819 to see if it fixes it, however if you want an immediate fix, lowering your concurrent compactors to <= 4 is probably best.","created":"2014-08-22T14:00:48.237+0000"},{"body":"Ok, I will still probably try to grab a debug log then after that try lowering concurrent compactors to see if it helps the issue.","created":"2014-08-22T14:08:22.944+0000"},{"body":"I uploaded the log now i will lower concurrent compactors to see if it helps and report back after its run for a while","created":"2014-08-25T04:38:30.250+0000"},{"body":"So question about current comapctors is the documentation wrong about it being set one per core because if that is the case then our box would be set to 4 correct?\n\nconcurrent_compactors¶\nSets the number of concurrent compaction processes allowed to run simultaneously on a node. Defaults to one compaction process per CPU core.\n\nper (http://www.datastax.com/documentation/cassandra/2.0/cassandra/configuration/configCassandra_yaml_r.html)","created":"2014-08-25T14:24:29.171+0000"},{"body":"Ah, my mistake. Somehow I read i2.8xlarge. Not i2.xlarge.","created":"2014-08-25T14:29:58.147+0000"},{"body":"[~zaller] Thanks for the log.\nI saw FNFE but I cannot track where the missing file comes from.\nCan you post the log before that?\n\nI think the cause of FNFE is CASSANDRA-7145, where the bug is L0 compaction accidentally catch already compacting SSTable from L1.\nIf you can try Marcus' patch attached to CASSANDRA-7145 and see if it fixes the problem, I appreciate.\n","created":"2014-08-26T17:23:15.251+0000"},{"body":"Is this still a concern? There's been no activity on this ticket in over a year.","created":"2015-09-15T17:01:27.237+0000"},{"body":"We are not experiencing the issue anymore it was fixed on some version in the 2.0.x line and we are now on 2.1.9 so it fixed in both versions. If my memory serves me right i could be way wrong i think it was the above patches that fixed it also i think i remember building cassandra with those and it worked.","created":"2015-09-15T17:38:58.705+0000"},{"body":"Great, thank you. Closing as fixed.","created":"2015-09-15T17:53:37.539+0000"}],"conversations":[{"body":"Cassandra will constantly eat disk space not sure whats causing it the only thing that seems to fix it is a restart of cassandra this happens about every 3-5 hrs we will grow from about 350GB to 650GB with no end in site. Once we restart cassandra it usually all clears itself up and disks return to normal for a while then something triggers its and starts climbing again. Sometimes when we restart compactions pending skyrocket and if we restart a second time the compactions pending drop off back to a normal level. One other thing to note is the space is not free'd until cassandra starts back up and not when shutdown.\n\nI will get a clean log of before and after restarting next time it happens and post it.\n\nHere is a common ERROR in our logs that might be related\n\n{noformat}\nERROR [CompactionExecutor:46] 2014-04-15 09:12:51,040 CassandraDaemon.java (line 196) Exception in thread Thread[CompactionExecutor:46,1,main]\njava.lang.RuntimeException: java.io.FileNotFoundException: /local-project/cassandra_data/data/wxgrid/grid/wxgrid-grid-jb-468677-Data.db (No such file or directory)\n at org.apache.cassandra.io.util.ThrottledReader.open(ThrottledReader.java:53)\n at org.apache.cassandra.io.sstable.SSTableReader.openDataReader(SSTableReader.java:1355)\n at org.apache.cassandra.io.sstable.SSTableScanner.(SSTableScanner.java:67)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1161)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1173)\n at org.apache.cassandra.db.compaction.LeveledCompactionStrategy.getScanners(LeveledCompactionStrategy.java:194)\n at org.apache.cassandra.db.compaction.AbstractCompactionStrategy.getScanners(AbstractCompactionStrategy.java:258)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:126)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:60)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:59)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:197)\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source)\n at java.util.concurrent.FutureTask.run(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\nCaused by: java.io.FileNotFoundException: /local-project/cassandra_data/data/wxgrid/grid/wxgrid-grid-jb-468677-Data.db (No such file or directory)\n at java.io.RandomAccessFile.open(Native Method)\n at java.io.RandomAccessFile.(Unknown Source)\n at org.apache.cassandra.io.util.RandomAccessReader.(RandomAccessReader.java:58)\n at org.apache.cassandra.io.util.ThrottledReader.(ThrottledReader.java:35)\n at org.apache.cassandra.io.util.ThrottledReader.open(ThrottledReader.java:49)\n ... 17 more\n{noformat}\n","from":"reporter","subject":"Disk space growth until restart"},{"body":"What are you doing before the exception? Dropping tables? Running repair? Truncating?","from":"developer"},{"body":"Doing nothing they just show up in the logs during normal use our tables are all ttl'd set to about a day with around 30GB or so of data churn. Upgraded to 2.0.7 and still having the issue, here is a error for the 2.0.7 version not even sure if the error is related to the space growth but it is the only bit of info in the logs that might suggest it. Willing to help you guys out anyway I can would love to get this fixed its pretty much our last major issue with keeping our cluster stable.\n\nERROR [CompactionExecutor:112] 2014-04-21 21:26:35,109 CassandraDaemon.java (line 198) Exception in thread Thread[CompactionExecutor:112,1,main]\njava.lang.RuntimeException: java.io.FileNotFoundException: /local-project/cassandra_data/data/wxgrid/grid/wxgrid-grid-jb-580812-Data.db (No such file or directory)\n at org.apache.cassandra.io.util.ThrottledReader.open(ThrottledReader.java:53)\n at org.apache.cassandra.io.sstable.SSTableReader.openDataReader(SSTableReader.java:1355)\n at org.apache.cassandra.io.sstable.SSTableScanner.(SSTableScanner.java:67)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1161)","from":"developer"},{"body":"Here is some screen shots for datastax opscenter. The dips are related to either running cleanup or changing the gc_grace period down after learning that since we ttl all our data we don't have to run repairs within gc_grace. You can see the spikes in the total disk space they will basicly continue until it runs our of space or we restart cassandra.","from":"developer"},{"body":"This is a dump of the directory for the cf that has the issues as you can see file count increases. It seems like files are not being deleted when they should be. The before file is taken when cassandra was in a stopped state the after is when cassandra had been started up again.","from":"developer"},{"body":"Here is a schema in question\n\nCREATE TABLE grid (\n data_id text,\n cylinder text,\n value blob,\n PRIMARY KEY (data_id, cylinder)\n) WITH COMPACT STORAGE AND\n bloom_filter_fp_chance=0.100000 AND\n caching='ALL' AND\n comment='' AND\n dclocal_read_repair_chance=0.000000 AND\n gc_grace_seconds=0 AND\n index_interval=128 AND\n read_repair_chance=0.100000 AND\n populate_io_cache_on_flush='false' AND\n default_time_to_live=0 AND\n speculative_retry='99.0PERCENTILE' AND\n memtable_flush_period_in_ms=0 AND\n compaction={'sstable_size_in_mb': '160', 'class': 'LeveledCompactionStrategy'} AND\n compression={};","from":"developer"},{"body":"Just thought I would also mention setup a brand new 3 node test cluster completely fresh and was able to reproduce.","from":"developer"},{"body":"While trying to reproduce CASSANDRA-6525, I'm seeing similar stacktraces in the logs:\n\n{noformat}\nERROR 13:45:24,236 Exception in thread Thread[CompactionExecutor:7,1,main]\njava.lang.RuntimeException: java.io.FileNotFoundException: /var/lib/cassandra/data/test6981/t12/test6981-t12.t12_blob5_idx-jb-5-Data.db (No such file or directory)\n at com.google.common.base.Throwables.propagate(Throwables.java:160)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:32)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:60)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:59)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:197)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:744)\nCaused by: java.io.FileNotFoundException: /var/lib/cassandra/data/test6981/t12/test6981-t12.t12_blob5_idx-jb-5-Data.db (No such file or directory)\n at java.io.RandomAccessFile.open(Native Method)\n at java.io.RandomAccessFile.(RandomAccessFile.java:241)\n at java.io.RandomAccessFile.(RandomAccessFile.java:122)\n at org.apache.cassandra.io.sstable.SSTableReader.preheat(SSTableReader.java:820)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:245)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n ... 8 more\n{noformat}\n\nand\n\n{noformat}\nERROR 13:46:42,182 Exception in thread Thread[CompactionExecutor:5,1,main]\njava.lang.RuntimeException: java.io.FileNotFoundException: /var/lib/cassandra/data/test6981/t4/test6981-t4-jb-4-Data.db (No such file or directory)\n at org.apache.cassandra.io.compress.CompressedThrottledReader.open(CompressedThrottledReader.java:52)\n at org.apache.cassandra.io.sstable.SSTableReader.openDataReader(SSTableReader.java:1355)\n at org.apache.cassandra.io.sstable.SSTableScanner.(SSTableScanner.java:67)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1161)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1173)\n at org.apache.cassandra.db.compaction.AbstractCompactionStrategy.getScanners(AbstractCompactionStrategy.java:262)\n at org.apache.cassandra.db.compaction.AbstractCompactionStrategy.getScanners(AbstractCompactionStrategy.java:268)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:126)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:60)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:59)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:197)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:744)\nCaused by: java.io.FileNotFoundException: /var/lib/cassandra/data/test6981/t4/test6981-t4-jb-4-Data.db (No such file or directory)\n at java.io.RandomAccessFile.open(Native Method)\n at java.io.RandomAccessFile.(RandomAccessFile.java:241)\n at org.apache.cassandra.io.util.RandomAccessReader.(RandomAccessReader.java:58)\n at org.apache.cassandra.io.compress.CompressedRandomAccessReader.(CompressedRandomAccessReader.java:76)\n at org.apache.cassandra.io.compress.CompressedThrottledReader.(CompressedThrottledReader.java:34)\n at org.apache.cassandra.io.compress.CompressedThrottledReader.open(CompressedThrottledReader.java:48)\n ... 17 more\n{noformat}\n\nI didn't see any problems when running the 6525 repro script on my laptop, but in a VM with only 1GB of RAM, I got quite a few of these errors.","from":"developer"},{"body":"Interesting, we are working on a script to reproduce that dose not require a bunch of dependancies and our whole infrastructure setup as a side note just recently deleted a key space and when looking at the data directories there where a bunch of files left behind on a few nodes that exhibited the issue.","from":"developer"},{"body":"This is possibly CASSANDRA-7511. Are you truncating ever? Is your commit log or your sstable data what is taking up the space? The fact that it gets fixed on restart makes me think commit log.","from":"developer"},{"body":"No, we never do a truncate, and our commit log folder is around 1GB in size. When we restart the data directory is what shrinks in size.","from":"developer"},{"body":"I'm trying to figure out the file not found exceptions in CASSANDRA-7145 - don't think they are related to the disk space growth though, so not marking this as a duplicate (yet)","from":"developer"},{"body":"At a guess, this is CASSANDRA-7139. Try raising your compaction throughput limit and/or lowering your concurrent_compactors (both in the cassandra.yaml)\n\n","from":"developer"},{"body":"I've filed CASSANDRA-7819 with a more general fix, in case this turns out to be something different. ","from":"developer"},{"body":"Our compaction limit is set to 0 and compaction memory 768MB. Concurrent compactors is set to the default and multithreading is off and compaction pre heat cache is also off. I am going to be getting a debug log today hopefully with this (log4j.logger.org.apache.cassandra.db.compaction=DEBUG) so we will see what the reveals.","from":"developer"},{"body":"The default concurrent_compactors is most likely too high for machines with many CPUs and modest disk throughput. What disk layout do you have?","from":"developer"},{"body":"We are using amazon ec2 i2.xlarge so one 800GB SSD","from":"developer"},{"body":"Ok, so that box has a default of 32 concurrent compactors on 2.0 (on 2.1 it will default to 2). This is almost certainly too many, and I think this is highly likely to be your problem. What would be interesting is to try the patch I have posted in CASSANDRA-7819 to see if it fixes it, however if you want an immediate fix, lowering your concurrent compactors to <= 4 is probably best.","from":"developer"},{"body":"Ok, I will still probably try to grab a debug log then after that try lowering concurrent compactors to see if it helps the issue.","from":"developer"},{"body":"I uploaded the log now i will lower concurrent compactors to see if it helps and report back after its run for a while","from":"developer"},{"body":"So question about current comapctors is the documentation wrong about it being set one per core because if that is the case then our box would be set to 4 correct?\n\nconcurrent_compactors¶\nSets the number of concurrent compaction processes allowed to run simultaneously on a node. Defaults to one compaction process per CPU core.\n\nper (http://www.datastax.com/documentation/cassandra/2.0/cassandra/configuration/configCassandra_yaml_r.html)","from":"developer"},{"body":"Ah, my mistake. Somehow I read i2.8xlarge. Not i2.xlarge.","from":"developer"},{"body":"[~zaller] Thanks for the log.\nI saw FNFE but I cannot track where the missing file comes from.\nCan you post the log before that?\n\nI think the cause of FNFE is CASSANDRA-7145, where the bug is L0 compaction accidentally catch already compacting SSTable from L1.\nIf you can try Marcus' patch attached to CASSANDRA-7145 and see if it fixes the problem, I appreciate.\n","from":"developer"},{"body":"Is this still a concern? There's been no activity on this ticket in over a year.","from":"developer"},{"body":"We are not experiencing the issue anymore it was fixed on some version in the 2.0.x line and we are now on 2.1.9 so it fixed in both versions. If my memory serves me right i could be way wrong i think it was the above patches that fixed it also i think i remember building cassandra with those and it worked.","from":"developer"},{"body":"Great, thank you. Closing as fixed.","from":"developer"}],"created":"2014-04-15T16:55:08.000+0000","description":"Cassandra will constantly eat disk space not sure whats causing it the only thing that seems to fix it is a restart of cassandra this happens about every 3-5 hrs we will grow from about 350GB to 650GB with no end in site. Once we restart cassandra it usually all clears itself up and disks return to normal for a while then something triggers its and starts climbing again. Sometimes when we restart compactions pending skyrocket and if we restart a second time the compactions pending drop off back to a normal level. One other thing to note is the space is not free'd until cassandra starts back up and not when shutdown.\n\nI will get a clean log of before and after restarting next time it happens and post it.\n\nHere is a common ERROR in our logs that might be related\n\n{noformat}\nERROR [CompactionExecutor:46] 2014-04-15 09:12:51,040 CassandraDaemon.java (line 196) Exception in thread Thread[CompactionExecutor:46,1,main]\njava.lang.RuntimeException: java.io.FileNotFoundException: /local-project/cassandra_data/data/wxgrid/grid/wxgrid-grid-jb-468677-Data.db (No such file or directory)\n at org.apache.cassandra.io.util.ThrottledReader.open(ThrottledReader.java:53)\n at org.apache.cassandra.io.sstable.SSTableReader.openDataReader(SSTableReader.java:1355)\n at org.apache.cassandra.io.sstable.SSTableScanner.(SSTableScanner.java:67)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1161)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1173)\n at org.apache.cassandra.db.compaction.LeveledCompactionStrategy.getScanners(LeveledCompactionStrategy.java:194)\n at org.apache.cassandra.db.compaction.AbstractCompactionStrategy.getScanners(AbstractCompactionStrategy.java:258)\n at org.apache.cassandra.db.compaction.CompactionTask.runWith(CompactionTask.java:126)\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48)\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n at org.apache.cassandra.db.compaction.CompactionTask.executeInternal(CompactionTask.java:60)\n at org.apache.cassandra.db.compaction.AbstractCompactionTask.execute(AbstractCompactionTask.java:59)\n at org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:197)\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source)\n at java.util.concurrent.FutureTask.run(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)\n at java.lang.Thread.run(Unknown Source)\nCaused by: java.io.FileNotFoundException: /local-project/cassandra_data/data/wxgrid/grid/wxgrid-grid-jb-468677-Data.db (No such file or directory)\n at java.io.RandomAccessFile.open(Native Method)\n at java.io.RandomAccessFile.(Unknown Source)\n at org.apache.cassandra.io.util.RandomAccessReader.(RandomAccessReader.java:58)\n at org.apache.cassandra.io.util.ThrottledReader.(ThrottledReader.java:35)\n at org.apache.cassandra.io.util.ThrottledReader.open(ThrottledReader.java:49)\n ... 17 more\n{noformat}\n","issue_id":"12708554","key":"CASSANDRA-7042","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-09-15T17:54:01.000+0000","role":"fixed_distractor","summary":"Disk space growth until restart"} {"case_id":"12712103","cluster":"DISTRACTOR-CASSANDRA-7144","comments":[{"body":"You didn't give a version so I'm not 100% sure but I think you're seeing this assert:\n\n{code}\n int size = rm.modifications.size();\n out.writeInt(size);\n assert size > 0;\n{code}\n\nwhich means that your client is sending a zero-length mutation to the server.\n\ndo we not validate that in the native protocol [~thobbs]?","created":"2014-05-05T18:45:42.909+0000"},{"body":"I had mentioned it, probably in the wrong field since it got removed. 2.0.7, latest at the moment. Quick update, I ended up decomissioning the node, wiping /var/lib/cassandra/* and then re-joining the cluster (same install, just wiped the dir). It has 100% solved my problem (so far anyway, about 2 days of uptime).\n\nSo it sounds like the data on disk was somehow corrupted somewhere between my client code update and repairs.","created":"2014-05-05T18:50:29.124+0000"},{"body":"[~nesnub] when you saw this, were you using prepared statements or unprepared statements? Were you using BATCH, INSERT, UPDATE, and/or DELETE?\n\nSo far I've tried a few things to reproduce this with no luck.","created":"2014-05-07T20:28:08.213+0000"},{"body":"I was not using prepared statements. I was doing INSERT and SELECT, no DELETE.\n\nI have some more info now. As I said I rebuilt the data on the node and it eliminated the problem. However, a day later, I found myself killing my ingestor (doing lots of INSERT), it's a python script using the new cassandra python-driver. When I did that I got the exception above. Thinking it was a \"runtime\" bug, I just kept going. Then the hinted-handoff started timing out on that box, so I restarted cassandra. From that point on, I would get the same exception without ever killing ingestors, at random interval. It seems as if killing the script during a query ended up sending data to cassandra that made it corrupt something on disk and that from that point on whenever it reached that part of the data on disk (I use that liberally, I just mean NOT directly from the script doing the ingestion) it would throw the exact same exception, leading to the timing out again. Restart of the cassandra node did nothing, I had to rebuild the data again. So now I'm very paranoid about killing my ingestion.\n\nThe ingestion uses UNLOGGED BATCH for some of the data for performance as well as normal INSERT.","created":"2014-05-08T15:22:48.038+0000"},{"body":"I'm Jason and I work at Stack Overflow. We're bit by this right now on our 2.0.6 cluster and we haven't wiped it out yet.\n\nWe got to this state on Friday after a power outage took out our entire datacenter (https://twitter.com/JasonPunyon/status/467290879569719296). We have a 3-node cluster and we interact with it using the CqlSharp library. We do our reads and writes at QUORUM, some of them are batches. The first symptom of the problem was that both read and write operations were timing out. When we dug in, we found that nodes 1 and 3 were throwing errors (not enough responses), but node 2 would execute the exact same operations fine at quorum. Nodes 1 and 3 would execute reads fine at ONE.\n\nWe're pretty interested in getting it fixed. Is there any way we can help? ","created":"2014-05-21T17:01:05.649+0000"},{"body":"Hi [~jason.punyon], thanks for your offer to help. The best way to help at the moment would be to help come up with a reproduceable test case, if you can. I have one other ticket on my queue in front of this one, so I should be able to give it proper attention pretty soon.","created":"2014-05-21T18:51:53.312+0000"},{"body":"[~thobbs] We racked our brains a bit, we can't really conceive of a direction to go in to replicate it again (remember it's still live on our cluster, we can run any diagnostics you want), aside from rebuilding and then randomly bringing power down (what we did the first time).","created":"2014-05-21T22:08:04.018+0000"},{"body":"[~jason.punyon] [~nesnub] can you provide any other details about the writes? I presume you were using prepared statements? How large were the batches? Roughly what do the individual inserts (within the batches or otherwise) look like?","created":"2014-05-22T21:55:35.504+0000"},{"body":"The writes were not using prepared statements. The batch contain average INSERT/UPDATEs, but I also have a counter update interleaved between the statements. Here is a sample:\n\n{code}\nself.queries.append( ( ( 'UPDATE met SET n = n + 1 WHERE id = %(id)s', { 'id' : k } ) )\nself.queries.append( ( '''BEGIN UNLOGGED BATCH\n INSERT INTO om ( id, obj, otype ) VALUES ( %(id)s, %(obj)s, %(type)s );\n UPDATE loc SET last = %(ts)s WHERE sid = %(sid)s AND otype = %(type)s AND id = %(id)s;\n INSERT INTO lid ( id, sid ) VALUES ( %(id)s, %(sid)s );\n APPLY BATCH;''', { 'id' : k,\n 'obj' : unicode( obj[ 0 ] ).lower(),\n 'type' : obj[ 1 ],\n 'ts' : ts,\n 'sid' : ag } ) )\n{code}","created":"2014-05-23T06:02:51.772+0000"},{"body":"Another point here is that one of these assertion errors which is caused in at OutboundTcpConnection.java:151 will cause that thread to die.\n\nThis will be very bad as it will not be recreated and it won't be able to talk to a node which had that connection. ","created":"2014-06-03T03:34:54.069+0000"},{"body":"[~nesnub] [~jason.punyon]\nAre you using CAS based inserts? ","created":"2014-06-03T04:12:49.398+0000"},{"body":"We saw this in 3 separate clusters at roughly the same time. In 2 clusters it happened on one node only. On 1 cluster it happened on 3 nodes.\n\nEvery time it happened, there were 3 errors in WRITE threads with stack trace:\n\n{code}\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.net.MessageOut.serialize(MessageOut.java:120)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeInternal(OutboundTcpConnection.java:251)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeConnected(OutboundTcpConnection.java:203)\n\tat org.apache.cassandra.net.OutboundTcpConnection.run(OutboundTcpConnection.java:151)\n{code}\n\nand the IPs in the thread names are all adjacent in the ring (we have replication 3+3 in 2 DCs). 2-3 seconds following that, there are then stack traces in the MutationStage like:\n\n{code}\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:654)\n\tat org.apache.cassandra.db.HintedHandOffManager.hintFor(HintedHandOffManager.java:137)\n\tat org.apache.cassandra.service.StorageProxy.writeHintForMutation(StorageProxy.java:897)\n\tat org.apache.cassandra.service.StorageProxy$6.runMayThrow(StorageProxy.java:870)\n\tat org.apache.cassandra.service.StorageProxy$HintRunnable.run(StorageProxy.java:1961)\n{code}\n\nThe number of these varies from 5 to 65.","created":"2014-06-03T11:09:59.512+0000"},{"body":"In fact, the number of MutationStage AssertionErrors always come in groups of 5 (within 1ms of each other).","created":"2014-06-03T12:25:27.640+0000"},{"body":"I have a reproduction. Not sure if it was the cause of our issue but it repros an identical situation.\n\nI have a 3 node cluster running on my box. I ran these cql commands:\n\n{code}\ncqlsh> create keyspace ks with replication = {'class':'SimpleStrategy', 'replication_factor':2};\ncqlsh> use ks;\ncqlsh:ks> create table counter_test (a text primary key, b counter) with gc_grace_seconds = 0;\ncqlsh:ks> consistency all;\nConsistency level set to ALL.\ncqlsh:ks> update counter_test set b = b + 1 where a = 'a';\ncqlsh:ks> delete from counter_test where a = 'a';\ncqlsh:ks> update counter_test set b = b + 1 where a = 'a';\nRequest did not complete within rpc_timeout.\n{code}\n\nAnd in the logs:\n\n{code}\nERROR [WRITE-/127.0.0.2] 2014-06-03 14:37:24,186 CassandraDaemon.java (line 196) Exception in thread Thread[WRITE-/127.0.0.2,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n at org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n at org.apache.cassandra.net.MessageOut.serialize(MessageOut.java:120)\n at org.apache.cassandra.net.OutboundTcpConnection.writeInternal(OutboundTcpConnection.java:251)\n at org.apache.cassandra.net.OutboundTcpConnection.writeConnected(OutboundTcpConnection.java:203)\n at org.apache.cassandra.net.OutboundTcpConnection.run(OutboundTcpConnection.java:151)\nERROR [MutationStage:156] 2014-06-03 14:37:26,301 CassandraDaemon.java (line 196) Exception in thread Thread[MutationStage:156,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n at org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n at org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:654)\n at org.apache.cassandra.db.HintedHandOffManager.hintFor(HintedHandOffManager.java:137)\n at org.apache.cassandra.service.StorageProxy.writeHintForMutation(StorageProxy.java:897)\n at org.apache.cassandra.service.StorageProxy$6.runMayThrow(StorageProxy.java:870)\n at org.apache.cassandra.service.StorageProxy$HintRunnable.run(StorageProxy.java:1961)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\nThe problem is that the row tombstone has microsecond timestamps but the counter has milliseconds (I just created CASSANDRA-7346 for this). When the local counter mutation is applied, it is immediately removed because the row tombstone has a later timestamp. When we read it back, once the row tombstone becomes GCable (that's why I set gc_grace to 0) we get null for the row, so we don't add a mutation to the RowMutation object (in CounterMutation.makeReplicaionMutation()). When this mutation gets sent it kills the connection thread, times out so we write a hint, which causes the second assertion failure. We now have to wait until the row tombstone becomes GCed to be able to inc anything in this row again.\n\nEven if there wasn't CASSANDRA-7346, it is legitimate to insert a row tombstone with a later timestamp so this could still happen. We need to detect this case and handle it so the empty counter increment doesn't cause trouble. We can just say it's done because the counter is deleted and not bother to replicate.\n\nI'll keep looking in case there are other code paths that trigger this.","created":"2014-06-03T13:58:54.342+0000"},{"body":"Attaching a simple 1.2 based patch. Only 1.2 and 2.0 are affected, the 2.1 implementation is not.","created":"2014-06-03T14:52:14.621+0000"},{"body":"Having looked at our writes, the were almost exclusively batches (sized 100-1000) and a counter batch (max size 1000). We used CAS for exactly one table (our migrations table) which was used very infrequently (in our build, a few times a day).","created":"2014-06-03T15:00:59.655+0000"},{"body":"[~jason.punyon] did you ever issue any row deletes to your counter table?\n\n[~iamaleksey] I don't see any patch","created":"2014-06-03T15:03:48.277+0000"},{"body":"[~rlow] Yes we did.","created":"2014-06-03T15:08:25.342+0000"},{"body":"[~jason.punyon] great, this is almost certainly the cause for you. You should check out CASSANDRA-7346 since row deletes on counter tables don't do at all what you expect.","created":"2014-06-03T15:11:13.232+0000"},{"body":"[~iamaleksey] I tried your patch (rebased for 2.0.6 where we're seeing the problem). It stops the AssertionError which is great, but I still get a timeout for the response. Do we need to poke the responseHandler too?","created":"2014-06-03T15:18:53.117+0000"},{"body":"bq. Do we need to poke the responseHandler too?\n\nWe do. Oh this is ugly-ish.","created":"2014-06-03T15:20:25.839+0000"},{"body":"Confirmed v2 works on 2.0.6. There is no AssertionError or timeout. I haven't verified it hasn't broken anything else though but it looks like it shouldn't.","created":"2014-06-03T15:41:06.739+0000"},{"body":"Committed, thanks. (fixed non-reusing the replicationMutation local var on commit).","created":"2014-06-03T16:04:33.069+0000"},{"body":"Confirmed that we issued counter deletes so it's the same root cause. Thanks for the fix!","created":"2014-06-03T16:49:23.404+0000"}],"conversations":[{"body":"First time reporting a bug here, apologies if I'm not posting it in the right space.\n\nAt what seems like random interval, on random nodes in random situations I will get the following exception. After this the hinted handoff start timing out and the node stops participating in the cluster.\n\nI started seeing these after switching to the Cassandra Python-Driver from the Python-CQL driver.\n\n{noformat}\nERROR [WRITE-/10.128.180.108] 2014-05-03 13:45:12,843 CassandraDaemon.java (line 198) Exception in thread Thread[WRITE-/10.128.180.108,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.net.MessageOut.serialize(MessageOut.java:120)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeInternal(OutboundTcpConnection.java:251)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeConnected(OutboundTcpConnection.java:203)\n\tat org.apache.cassandra.net.OutboundTcpConnection.run(OutboundTcpConnection.java:151)\nERROR [WRITE-/10.128.194.70] 2014-05-03 13:45:12,843 CassandraDaemon.java (line 198) Exception in thread Thread[WRITE-/10.128.194.70,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.net.MessageOut.serialize(MessageOut.java:120)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeInternal(OutboundTcpConnection.java:251)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeConnected(OutboundTcpConnection.java:203)\n\tat org.apache.cassandra.net.OutboundTcpConnection.run(OutboundTcpConnection.java:151)\nERROR [MutationStage:118] 2014-05-03 13:45:15,048 CassandraDaemon.java (line 198) Exception in thread Thread[MutationStage:118,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:654)\n\tat org.apache.cassandra.db.HintedHandOffManager.hintFor(HintedHandOffManager.java:137)\n\tat org.apache.cassandra.service.StorageProxy.writeHintForMutation(StorageProxy.java:908)\n\tat org.apache.cassandra.service.StorageProxy$6.runMayThrow(StorageProxy.java:881)\n\tat org.apache.cassandra.service.StorageProxy$HintRunnable.run(StorageProxy.java:1981)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:744)\nERROR [MutationStage:117] 2014-05-03 13:45:15,048 CassandraDaemon.java (line 198) Exception in thread Thread[MutationStage:117,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:654)\n\tat org.apache.cassandra.db.HintedHandOffManager.hintFor(HintedHandOffManager.java:137)\n\tat org.apache.cassandra.service.StorageProxy.writeHintForMutation(StorageProxy.java:908)\n\tat org.apache.cassandra.service.StorageProxy$6.runMayThrow(StorageProxy.java:881)\n\tat org.apache.cassandra.service.StorageProxy$HintRunnable.run(StorageProxy.java:1981)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:744)\n{noformat}\n\nThe service must be restarted for the node to come back online. Let me know any additional configuration details needed.","from":"reporter","subject":"Fix handling of empty counter replication mutations"},{"body":"You didn't give a version so I'm not 100% sure but I think you're seeing this assert:\n\n{code}\n int size = rm.modifications.size();\n out.writeInt(size);\n assert size > 0;\n{code}\n\nwhich means that your client is sending a zero-length mutation to the server.\n\ndo we not validate that in the native protocol [~thobbs]?","from":"developer"},{"body":"I had mentioned it, probably in the wrong field since it got removed. 2.0.7, latest at the moment. Quick update, I ended up decomissioning the node, wiping /var/lib/cassandra/* and then re-joining the cluster (same install, just wiped the dir). It has 100% solved my problem (so far anyway, about 2 days of uptime).\n\nSo it sounds like the data on disk was somehow corrupted somewhere between my client code update and repairs.","from":"developer"},{"body":"[~nesnub] when you saw this, were you using prepared statements or unprepared statements? Were you using BATCH, INSERT, UPDATE, and/or DELETE?\n\nSo far I've tried a few things to reproduce this with no luck.","from":"developer"},{"body":"I was not using prepared statements. I was doing INSERT and SELECT, no DELETE.\n\nI have some more info now. As I said I rebuilt the data on the node and it eliminated the problem. However, a day later, I found myself killing my ingestor (doing lots of INSERT), it's a python script using the new cassandra python-driver. When I did that I got the exception above. Thinking it was a \"runtime\" bug, I just kept going. Then the hinted-handoff started timing out on that box, so I restarted cassandra. From that point on, I would get the same exception without ever killing ingestors, at random interval. It seems as if killing the script during a query ended up sending data to cassandra that made it corrupt something on disk and that from that point on whenever it reached that part of the data on disk (I use that liberally, I just mean NOT directly from the script doing the ingestion) it would throw the exact same exception, leading to the timing out again. Restart of the cassandra node did nothing, I had to rebuild the data again. So now I'm very paranoid about killing my ingestion.\n\nThe ingestion uses UNLOGGED BATCH for some of the data for performance as well as normal INSERT.","from":"developer"},{"body":"I'm Jason and I work at Stack Overflow. We're bit by this right now on our 2.0.6 cluster and we haven't wiped it out yet.\n\nWe got to this state on Friday after a power outage took out our entire datacenter (https://twitter.com/JasonPunyon/status/467290879569719296). We have a 3-node cluster and we interact with it using the CqlSharp library. We do our reads and writes at QUORUM, some of them are batches. The first symptom of the problem was that both read and write operations were timing out. When we dug in, we found that nodes 1 and 3 were throwing errors (not enough responses), but node 2 would execute the exact same operations fine at quorum. Nodes 1 and 3 would execute reads fine at ONE.\n\nWe're pretty interested in getting it fixed. Is there any way we can help? ","from":"developer"},{"body":"Hi [~jason.punyon], thanks for your offer to help. The best way to help at the moment would be to help come up with a reproduceable test case, if you can. I have one other ticket on my queue in front of this one, so I should be able to give it proper attention pretty soon.","from":"developer"},{"body":"[~thobbs] We racked our brains a bit, we can't really conceive of a direction to go in to replicate it again (remember it's still live on our cluster, we can run any diagnostics you want), aside from rebuilding and then randomly bringing power down (what we did the first time).","from":"developer"},{"body":"[~jason.punyon] [~nesnub] can you provide any other details about the writes? I presume you were using prepared statements? How large were the batches? Roughly what do the individual inserts (within the batches or otherwise) look like?","from":"developer"},{"body":"The writes were not using prepared statements. The batch contain average INSERT/UPDATEs, but I also have a counter update interleaved between the statements. Here is a sample:\n\n{code}\nself.queries.append( ( ( 'UPDATE met SET n = n + 1 WHERE id = %(id)s', { 'id' : k } ) )\nself.queries.append( ( '''BEGIN UNLOGGED BATCH\n INSERT INTO om ( id, obj, otype ) VALUES ( %(id)s, %(obj)s, %(type)s );\n UPDATE loc SET last = %(ts)s WHERE sid = %(sid)s AND otype = %(type)s AND id = %(id)s;\n INSERT INTO lid ( id, sid ) VALUES ( %(id)s, %(sid)s );\n APPLY BATCH;''', { 'id' : k,\n 'obj' : unicode( obj[ 0 ] ).lower(),\n 'type' : obj[ 1 ],\n 'ts' : ts,\n 'sid' : ag } ) )\n{code}","from":"developer"},{"body":"Another point here is that one of these assertion errors which is caused in at OutboundTcpConnection.java:151 will cause that thread to die.\n\nThis will be very bad as it will not be recreated and it won't be able to talk to a node which had that connection. ","from":"developer"},{"body":"[~nesnub] [~jason.punyon]\nAre you using CAS based inserts? ","from":"developer"},{"body":"We saw this in 3 separate clusters at roughly the same time. In 2 clusters it happened on one node only. On 1 cluster it happened on 3 nodes.\n\nEvery time it happened, there were 3 errors in WRITE threads with stack trace:\n\n{code}\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.net.MessageOut.serialize(MessageOut.java:120)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeInternal(OutboundTcpConnection.java:251)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeConnected(OutboundTcpConnection.java:203)\n\tat org.apache.cassandra.net.OutboundTcpConnection.run(OutboundTcpConnection.java:151)\n{code}\n\nand the IPs in the thread names are all adjacent in the ring (we have replication 3+3 in 2 DCs). 2-3 seconds following that, there are then stack traces in the MutationStage like:\n\n{code}\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:654)\n\tat org.apache.cassandra.db.HintedHandOffManager.hintFor(HintedHandOffManager.java:137)\n\tat org.apache.cassandra.service.StorageProxy.writeHintForMutation(StorageProxy.java:897)\n\tat org.apache.cassandra.service.StorageProxy$6.runMayThrow(StorageProxy.java:870)\n\tat org.apache.cassandra.service.StorageProxy$HintRunnable.run(StorageProxy.java:1961)\n{code}\n\nThe number of these varies from 5 to 65.","from":"developer"},{"body":"In fact, the number of MutationStage AssertionErrors always come in groups of 5 (within 1ms of each other).","from":"developer"},{"body":"I have a reproduction. Not sure if it was the cause of our issue but it repros an identical situation.\n\nI have a 3 node cluster running on my box. I ran these cql commands:\n\n{code}\ncqlsh> create keyspace ks with replication = {'class':'SimpleStrategy', 'replication_factor':2};\ncqlsh> use ks;\ncqlsh:ks> create table counter_test (a text primary key, b counter) with gc_grace_seconds = 0;\ncqlsh:ks> consistency all;\nConsistency level set to ALL.\ncqlsh:ks> update counter_test set b = b + 1 where a = 'a';\ncqlsh:ks> delete from counter_test where a = 'a';\ncqlsh:ks> update counter_test set b = b + 1 where a = 'a';\nRequest did not complete within rpc_timeout.\n{code}\n\nAnd in the logs:\n\n{code}\nERROR [WRITE-/127.0.0.2] 2014-06-03 14:37:24,186 CassandraDaemon.java (line 196) Exception in thread Thread[WRITE-/127.0.0.2,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n at org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n at org.apache.cassandra.net.MessageOut.serialize(MessageOut.java:120)\n at org.apache.cassandra.net.OutboundTcpConnection.writeInternal(OutboundTcpConnection.java:251)\n at org.apache.cassandra.net.OutboundTcpConnection.writeConnected(OutboundTcpConnection.java:203)\n at org.apache.cassandra.net.OutboundTcpConnection.run(OutboundTcpConnection.java:151)\nERROR [MutationStage:156] 2014-06-03 14:37:26,301 CassandraDaemon.java (line 196) Exception in thread Thread[MutationStage:156,5,main]\njava.lang.AssertionError\n at org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n at org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n at org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:654)\n at org.apache.cassandra.db.HintedHandOffManager.hintFor(HintedHandOffManager.java:137)\n at org.apache.cassandra.service.StorageProxy.writeHintForMutation(StorageProxy.java:897)\n at org.apache.cassandra.service.StorageProxy$6.runMayThrow(StorageProxy.java:870)\n at org.apache.cassandra.service.StorageProxy$HintRunnable.run(StorageProxy.java:1961)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\nThe problem is that the row tombstone has microsecond timestamps but the counter has milliseconds (I just created CASSANDRA-7346 for this). When the local counter mutation is applied, it is immediately removed because the row tombstone has a later timestamp. When we read it back, once the row tombstone becomes GCable (that's why I set gc_grace to 0) we get null for the row, so we don't add a mutation to the RowMutation object (in CounterMutation.makeReplicaionMutation()). When this mutation gets sent it kills the connection thread, times out so we write a hint, which causes the second assertion failure. We now have to wait until the row tombstone becomes GCed to be able to inc anything in this row again.\n\nEven if there wasn't CASSANDRA-7346, it is legitimate to insert a row tombstone with a later timestamp so this could still happen. We need to detect this case and handle it so the empty counter increment doesn't cause trouble. We can just say it's done because the counter is deleted and not bother to replicate.\n\nI'll keep looking in case there are other code paths that trigger this.","from":"developer"},{"body":"Attaching a simple 1.2 based patch. Only 1.2 and 2.0 are affected, the 2.1 implementation is not.","from":"developer"},{"body":"Having looked at our writes, the were almost exclusively batches (sized 100-1000) and a counter batch (max size 1000). We used CAS for exactly one table (our migrations table) which was used very infrequently (in our build, a few times a day).","from":"developer"},{"body":"[~jason.punyon] did you ever issue any row deletes to your counter table?\n\n[~iamaleksey] I don't see any patch","from":"developer"},{"body":"[~rlow] Yes we did.","from":"developer"},{"body":"[~jason.punyon] great, this is almost certainly the cause for you. You should check out CASSANDRA-7346 since row deletes on counter tables don't do at all what you expect.","from":"developer"},{"body":"[~iamaleksey] I tried your patch (rebased for 2.0.6 where we're seeing the problem). It stops the AssertionError which is great, but I still get a timeout for the response. Do we need to poke the responseHandler too?","from":"developer"},{"body":"bq. Do we need to poke the responseHandler too?\n\nWe do. Oh this is ugly-ish.","from":"developer"},{"body":"Confirmed v2 works on 2.0.6. There is no AssertionError or timeout. I haven't verified it hasn't broken anything else though but it looks like it shouldn't.","from":"developer"},{"body":"Committed, thanks. (fixed non-reusing the replicationMutation local var on commit).","from":"developer"},{"body":"Confirmed that we issued counter deletes so it's the same root cause. Thanks for the fix!","from":"developer"}],"created":"2014-05-03T08:23:18.000+0000","description":"First time reporting a bug here, apologies if I'm not posting it in the right space.\n\nAt what seems like random interval, on random nodes in random situations I will get the following exception. After this the hinted handoff start timing out and the node stops participating in the cluster.\n\nI started seeing these after switching to the Cassandra Python-Driver from the Python-CQL driver.\n\n{noformat}\nERROR [WRITE-/10.128.180.108] 2014-05-03 13:45:12,843 CassandraDaemon.java (line 198) Exception in thread Thread[WRITE-/10.128.180.108,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.net.MessageOut.serialize(MessageOut.java:120)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeInternal(OutboundTcpConnection.java:251)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeConnected(OutboundTcpConnection.java:203)\n\tat org.apache.cassandra.net.OutboundTcpConnection.run(OutboundTcpConnection.java:151)\nERROR [WRITE-/10.128.194.70] 2014-05-03 13:45:12,843 CassandraDaemon.java (line 198) Exception in thread Thread[WRITE-/10.128.194.70,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.net.MessageOut.serialize(MessageOut.java:120)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeInternal(OutboundTcpConnection.java:251)\n\tat org.apache.cassandra.net.OutboundTcpConnection.writeConnected(OutboundTcpConnection.java:203)\n\tat org.apache.cassandra.net.OutboundTcpConnection.run(OutboundTcpConnection.java:151)\nERROR [MutationStage:118] 2014-05-03 13:45:15,048 CassandraDaemon.java (line 198) Exception in thread Thread[MutationStage:118,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:654)\n\tat org.apache.cassandra.db.HintedHandOffManager.hintFor(HintedHandOffManager.java:137)\n\tat org.apache.cassandra.service.StorageProxy.writeHintForMutation(StorageProxy.java:908)\n\tat org.apache.cassandra.service.StorageProxy$6.runMayThrow(StorageProxy.java:881)\n\tat org.apache.cassandra.service.StorageProxy$HintRunnable.run(StorageProxy.java:1981)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:744)\nERROR [MutationStage:117] 2014-05-03 13:45:15,048 CassandraDaemon.java (line 198) Exception in thread Thread[MutationStage:117,5,main]\njava.lang.AssertionError\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:271)\n\tat org.apache.cassandra.db.RowMutation$RowMutationSerializer.serialize(RowMutation.java:259)\n\tat org.apache.cassandra.utils.FBUtilities.serialize(FBUtilities.java:654)\n\tat org.apache.cassandra.db.HintedHandOffManager.hintFor(HintedHandOffManager.java:137)\n\tat org.apache.cassandra.service.StorageProxy.writeHintForMutation(StorageProxy.java:908)\n\tat org.apache.cassandra.service.StorageProxy$6.runMayThrow(StorageProxy.java:881)\n\tat org.apache.cassandra.service.StorageProxy$HintRunnable.run(StorageProxy.java:1981)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:744)\n{noformat}\n\nThe service must be restarted for the node to come back online. Let me know any additional configuration details needed.","issue_id":"12712103","key":"CASSANDRA-7144","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-06-03T16:04:33.000+0000","role":"fixed_distractor","summary":"Fix handling of empty counter replication mutations"} {"case_id":"12744762","cluster":"DISTRACTOR-CASSANDRA-8020","comments":[{"body":"It looks like there was an error while trying to create the snapshot. Can you look through the logs for \"Error occurred during snapshot phase\"? This will help pin down the original cause of the failed repair.","created":"2014-09-29T21:22:57.425+0000"},{"body":"Carl, no \"Error occurred during snapshot phase\". But I took another look (yes, I should have done this before) and the beginning of the \"nodetool repair\" points to the AntiEntropyStage class:\n\n{noformat}\nNFO [Thread-49681] 2014-09-29 11:25:22,787 StorageService.java:2595 - Starting repair command #5, repairing 743 ranges for keyspace test (seq=true, full=true)\nINFO [AntiEntropySessions:17] 2014-09-29 11:25:24,264 RepairSession.java:260 - [repair #736daa80-47e4-11e4-ba0f-c7788dc924ec] new session: will sync c1.test.net/192.168.33.248, /192.168.33.250, /192.168.33.252 on range (-7990332010750800111,-7986163865225197304] for test.[user]\nERROR [AntiEntropyStage:1487] 2014-09-29 11:25:24,265 CassandraDaemon.java:166 - Exception in thread Thread[AntiEntropyStage:1487,5,main]\njava.lang.ClassCastException: null\nERROR [RepairJobTask:3] 2014-09-29 11:25:24,266 RepairJob.java:127 - Error occurred during snapshot phase\njava.lang.RuntimeException: Could not create snapshot at /192.168.33.248\n\tat org.apache.cassandra.repair.SnapshotTask$SnapshotCallback.onFailure(SnapshotTask.java:77) ~[apache-cassandra-2.1.0.jar:2.1.0]\n\tat org.apache.cassandra.net.ResponseVerbHandler.doVerb(ResponseVerbHandler.java:48) ~[apache-cassandra-2.1.0.jar:2.1.0]\n\tat org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:62) ~[apache-cassandra-2.1.0.jar:2.1.0]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) [na:1.7.0_67]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_67]\n\tat java.lang.Thread.run(Thread.java:745) [na:1.7.0_67]\nERROR [AntiEntropySessions:17] 2014-09-29 11:25:24,266 RepairSession.java:303 - [repair #736daa80-47e4-11e4-ba0f-c7788dc924ec] session completed with the following error\njava.io.IOException: Failed during snapshot creation.\n\tat org.apache.cassandra.repair.RepairSession.failedSnapshot(RepairSession.java:344) ~[apache-cassandra-2.1.0.jar:2.1.0]\n\tat org.apache.cassandra.repair.RepairJob$2.onFailure(RepairJob.java:128) ~[apache-cassandra-2.1.0.jar:2.1.0]\n\tat com.google.common.util.concurrent.Futures$4.run(Futures.java:1172) ~[guava-16.0.jar:na]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) [na:1.7.0_67]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_67]\n\tat java.lang.Thread.run(Thread.java:745) [na:1.7.0_67]\n{noformat}\n\nI'm attaching the full log of the error.","created":"2014-09-29T21:53:43.180+0000"},{"body":"nodetool repair system.log","created":"2014-09-29T21:56:34.433+0000"},{"body":"Confirmed.\n\nWhen snapshotting, replica node is throwing the following on indexed table:\n\n{code}\nERROR [AntiEntropyStage:768] 2014-09-29 17:34:25,654 CassandraDaemon.java:166 - Exception in thread Thread[AntiEntropyStage:768,5,main]\njava.lang.ClassCastException: java.lang.Long cannot be cast to java.nio.ByteBuffer\n at org.apache.cassandra.db.marshal.BytesType.compare(BytesType.java:29) ~[main/:na]\n at org.apache.cassandra.dht.LocalToken.compareTo(LocalToken.java:44) ~[main/:na]\n at org.apache.cassandra.dht.LocalToken.compareTo(LocalToken.java:24) ~[main/:na]\n at org.apache.cassandra.dht.Range.contains(Range.java:71) ~[main/:na]\n at org.apache.cassandra.dht.Range.contains(Range.java:111) ~[main/:na]\n at org.apache.cassandra.dht.Range.intersects(Range.java:142) ~[main/:na]\n at org.apache.cassandra.dht.Range.intersects(Range.java:129) ~[main/:na]\n at org.apache.cassandra.dht.AbstractBounds.intersects(AbstractBounds.java:83) ~[main/:na]\n at org.apache.cassandra.repair.RepairMessageVerbHandler$1.apply(RepairMessageVerbHandler.java:83) ~[main/:na]\n at org.apache.cassandra.repair.RepairMessageVerbHandler$1.apply(RepairMessageVerbHandler.java:80) ~[main/:na]\n at org.apache.cassandra.db.ColumnFamilyStore.snapshotWithoutFlush(ColumnFamilyStore.java:2152) ~[main/:na]\n at org.apache.cassandra.db.ColumnFamilyStore.snapshot(ColumnFamilyStore.java:2215) ~[main/:na]\n at org.apache.cassandra.repair.RepairMessageVerbHandler.doVerb(RepairMessageVerbHandler.java:79) ~[main/:na]\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:62) ~[main/:na]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_51]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) ~[na:1.7.0_51]\n at java.lang.Thread.run(Thread.java:744) ~[na:1.7.0_51]\n{code}","created":"2014-09-29T22:38:01.326+0000"},{"body":"This is introduced in CASSANDRA-7024.","created":"2014-09-29T22:39:57.457+0000"},{"body":"Just bumped into this. Similar log lines found. Any workarounds until 2.1.1 is out?","created":"2014-10-01T11:34:42.140+0000"},{"body":"We're seeing this as well","created":"2014-10-01T14:16:45.723+0000"},{"body":"The crux of the problem seems to be that in {{org.apache.cassandra.repair.RepairMessageVerbHandler}}, the line 83\n\n{noformat}\nreturn sstable != null && new Bounds<>(sstable.first.getToken(), \n sstable.last.getToken()).intersects(Collections.singleton(repairingRange));\n{noformat}\n\nmight generate a comparison between a {{LocalToken}} and a {{LongToken}} which fails since the former wraps a ByteBuffer and the latter a Long\n\n","created":"2014-10-01T15:16:33.751+0000"},{"body":"Attached patch to exclude SSTables from 2i from snapshotting.\n","created":"2014-10-01T15:23:54.329+0000"},{"body":"+1\n\n(is this bug new in 2.1?)","created":"2014-10-01T15:37:11.817+0000"},{"body":"bq. is this bug new in 2.1?\n\nYes, it's from CASSANDRA-7024.","created":"2014-10-01T15:46:29.089+0000"},{"body":"Committed, thanks!","created":"2014-10-02T17:19:02.510+0000"}],"conversations":[{"body":"Running a nodetool repair on Cassandra 2.1.0 indexed tables returns java exception about creating snapshots:\n\nCommand line:\n{noformat}\n[2014-09-29 11:25:24,945] Repair session 73c0d390-47e4-11e4-ba0f-c7788dc924ec for range (-7298689860784559350,-7297558156602685286] failed with error java.io.IOException: Failed during snapshot creation.\n[2014-09-29 11:25:24,945] Repair command #5 finished\n{noformat}\nCassandra log:\n{noformat}\nERROR [Thread-49681] 2014-09-29 11:25:24,945 StorageService.java:2689 - Repair session 73c0d390-47e4-11e4-ba0f-c7788dc924ec for range (-7298689860784559350,-7297558156602685286] failed with error java.io.IOException: Failed during snapshot creation.\njava.util.concurrent.ExecutionException: java.lang.RuntimeException: java.io.IOException: Failed during snapshot creation.\n at java.util.concurrent.FutureTask.report(FutureTask.java:122) [na:1.7.0_67]\n at java.util.concurrent.FutureTask.get(FutureTask.java:188) [na:1.7.0_67]\n at org.apache.cassandra.service.StorageService$4.runMayThrow(StorageService.java:2680) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) [apache-cassandra-2.1.0.jar:2.1.0]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_67]\n at java.util.concurrent.FutureTask.run(FutureTask.java:262) [na:1.7.0_67]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_67]\nCaused by: java.lang.RuntimeException: java.io.IOException: Failed during snapshot creation.\n at com.google.common.base.Throwables.propagate(Throwables.java:160) ~[guava-16.0.jar:na]\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:32) [apache-cassandra-2.1.0.jar:2.1.0]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_67]\n at java.util.concurrent.FutureTask.run(FutureTask.java:262) [na:1.7.0_67]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_67]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) ~[na:1.7.0_67]\n ... 1 common frames omitted\nCaused by: java.io.IOException: Failed during snapshot creation.\n at org.apache.cassandra.repair.RepairSession.failedSnapshot(RepairSession.java:344) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at org.apache.cassandra.repair.RepairJob$2.onFailure(RepairJob.java:128) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at com.google.common.util.concurrent.Futures$4.run(Futures.java:1172) ~[guava-16.0.jar:na]\n ... 3 common frames omitted\n{noformat}\nIf the index is dropped, the repair returns no error:\n{noformat}\ncqlsh:test> drop INDEX user_pass_idx ;\n\nroot@test:~# nodetool repair test user\n[2014-09-29 11:27:29,668] Starting repair command #6, repairing 743 ranges for keyspace test (seq=true, full=true)\n.\n.\n[2014-09-29 11:28:38,030] Repair session e6d40e10-47e4-11e4-ba0f-c7788dc924ec for range (-7298689860784559350,-7297558156602685286] finished\n[2014-09-29 11:28:38,030] Repair command #6 finished\n{noformat}\nThe test table:\n{noformat}\nCREATE TABLE test.user (\n login text PRIMARY KEY,\n password text\n)\ncreate INDEX user_pass_idx on test.user (password) ;\n{noformat}","from":"reporter","subject":"nodetool repair on Cassandra 2.1.0 indexed tables returns java exception about creating snapshots"},{"body":"It looks like there was an error while trying to create the snapshot. Can you look through the logs for \"Error occurred during snapshot phase\"? This will help pin down the original cause of the failed repair.","from":"developer"},{"body":"Carl, no \"Error occurred during snapshot phase\". But I took another look (yes, I should have done this before) and the beginning of the \"nodetool repair\" points to the AntiEntropyStage class:\n\n{noformat}\nNFO [Thread-49681] 2014-09-29 11:25:22,787 StorageService.java:2595 - Starting repair command #5, repairing 743 ranges for keyspace test (seq=true, full=true)\nINFO [AntiEntropySessions:17] 2014-09-29 11:25:24,264 RepairSession.java:260 - [repair #736daa80-47e4-11e4-ba0f-c7788dc924ec] new session: will sync c1.test.net/192.168.33.248, /192.168.33.250, /192.168.33.252 on range (-7990332010750800111,-7986163865225197304] for test.[user]\nERROR [AntiEntropyStage:1487] 2014-09-29 11:25:24,265 CassandraDaemon.java:166 - Exception in thread Thread[AntiEntropyStage:1487,5,main]\njava.lang.ClassCastException: null\nERROR [RepairJobTask:3] 2014-09-29 11:25:24,266 RepairJob.java:127 - Error occurred during snapshot phase\njava.lang.RuntimeException: Could not create snapshot at /192.168.33.248\n\tat org.apache.cassandra.repair.SnapshotTask$SnapshotCallback.onFailure(SnapshotTask.java:77) ~[apache-cassandra-2.1.0.jar:2.1.0]\n\tat org.apache.cassandra.net.ResponseVerbHandler.doVerb(ResponseVerbHandler.java:48) ~[apache-cassandra-2.1.0.jar:2.1.0]\n\tat org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:62) ~[apache-cassandra-2.1.0.jar:2.1.0]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) [na:1.7.0_67]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_67]\n\tat java.lang.Thread.run(Thread.java:745) [na:1.7.0_67]\nERROR [AntiEntropySessions:17] 2014-09-29 11:25:24,266 RepairSession.java:303 - [repair #736daa80-47e4-11e4-ba0f-c7788dc924ec] session completed with the following error\njava.io.IOException: Failed during snapshot creation.\n\tat org.apache.cassandra.repair.RepairSession.failedSnapshot(RepairSession.java:344) ~[apache-cassandra-2.1.0.jar:2.1.0]\n\tat org.apache.cassandra.repair.RepairJob$2.onFailure(RepairJob.java:128) ~[apache-cassandra-2.1.0.jar:2.1.0]\n\tat com.google.common.util.concurrent.Futures$4.run(Futures.java:1172) ~[guava-16.0.jar:na]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) [na:1.7.0_67]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_67]\n\tat java.lang.Thread.run(Thread.java:745) [na:1.7.0_67]\n{noformat}\n\nI'm attaching the full log of the error.","from":"developer"},{"body":"nodetool repair system.log","from":"developer"},{"body":"Confirmed.\n\nWhen snapshotting, replica node is throwing the following on indexed table:\n\n{code}\nERROR [AntiEntropyStage:768] 2014-09-29 17:34:25,654 CassandraDaemon.java:166 - Exception in thread Thread[AntiEntropyStage:768,5,main]\njava.lang.ClassCastException: java.lang.Long cannot be cast to java.nio.ByteBuffer\n at org.apache.cassandra.db.marshal.BytesType.compare(BytesType.java:29) ~[main/:na]\n at org.apache.cassandra.dht.LocalToken.compareTo(LocalToken.java:44) ~[main/:na]\n at org.apache.cassandra.dht.LocalToken.compareTo(LocalToken.java:24) ~[main/:na]\n at org.apache.cassandra.dht.Range.contains(Range.java:71) ~[main/:na]\n at org.apache.cassandra.dht.Range.contains(Range.java:111) ~[main/:na]\n at org.apache.cassandra.dht.Range.intersects(Range.java:142) ~[main/:na]\n at org.apache.cassandra.dht.Range.intersects(Range.java:129) ~[main/:na]\n at org.apache.cassandra.dht.AbstractBounds.intersects(AbstractBounds.java:83) ~[main/:na]\n at org.apache.cassandra.repair.RepairMessageVerbHandler$1.apply(RepairMessageVerbHandler.java:83) ~[main/:na]\n at org.apache.cassandra.repair.RepairMessageVerbHandler$1.apply(RepairMessageVerbHandler.java:80) ~[main/:na]\n at org.apache.cassandra.db.ColumnFamilyStore.snapshotWithoutFlush(ColumnFamilyStore.java:2152) ~[main/:na]\n at org.apache.cassandra.db.ColumnFamilyStore.snapshot(ColumnFamilyStore.java:2215) ~[main/:na]\n at org.apache.cassandra.repair.RepairMessageVerbHandler.doVerb(RepairMessageVerbHandler.java:79) ~[main/:na]\n at org.apache.cassandra.net.MessageDeliveryTask.run(MessageDeliveryTask.java:62) ~[main/:na]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_51]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) ~[na:1.7.0_51]\n at java.lang.Thread.run(Thread.java:744) ~[na:1.7.0_51]\n{code}","from":"developer"},{"body":"This is introduced in CASSANDRA-7024.","from":"developer"},{"body":"Just bumped into this. Similar log lines found. Any workarounds until 2.1.1 is out?","from":"developer"},{"body":"We're seeing this as well","from":"developer"},{"body":"The crux of the problem seems to be that in {{org.apache.cassandra.repair.RepairMessageVerbHandler}}, the line 83\n\n{noformat}\nreturn sstable != null && new Bounds<>(sstable.first.getToken(), \n sstable.last.getToken()).intersects(Collections.singleton(repairingRange));\n{noformat}\n\nmight generate a comparison between a {{LocalToken}} and a {{LongToken}} which fails since the former wraps a ByteBuffer and the latter a Long\n\n","from":"developer"},{"body":"Attached patch to exclude SSTables from 2i from snapshotting.\n","from":"developer"},{"body":"+1\n\n(is this bug new in 2.1?)","from":"developer"},{"body":"bq. is this bug new in 2.1?\n\nYes, it's from CASSANDRA-7024.","from":"developer"},{"body":"Committed, thanks!","from":"developer"}],"created":"2014-09-29T19:30:06.000+0000","description":"Running a nodetool repair on Cassandra 2.1.0 indexed tables returns java exception about creating snapshots:\n\nCommand line:\n{noformat}\n[2014-09-29 11:25:24,945] Repair session 73c0d390-47e4-11e4-ba0f-c7788dc924ec for range (-7298689860784559350,-7297558156602685286] failed with error java.io.IOException: Failed during snapshot creation.\n[2014-09-29 11:25:24,945] Repair command #5 finished\n{noformat}\nCassandra log:\n{noformat}\nERROR [Thread-49681] 2014-09-29 11:25:24,945 StorageService.java:2689 - Repair session 73c0d390-47e4-11e4-ba0f-c7788dc924ec for range (-7298689860784559350,-7297558156602685286] failed with error java.io.IOException: Failed during snapshot creation.\njava.util.concurrent.ExecutionException: java.lang.RuntimeException: java.io.IOException: Failed during snapshot creation.\n at java.util.concurrent.FutureTask.report(FutureTask.java:122) [na:1.7.0_67]\n at java.util.concurrent.FutureTask.get(FutureTask.java:188) [na:1.7.0_67]\n at org.apache.cassandra.service.StorageService$4.runMayThrow(StorageService.java:2680) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) [apache-cassandra-2.1.0.jar:2.1.0]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_67]\n at java.util.concurrent.FutureTask.run(FutureTask.java:262) [na:1.7.0_67]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_67]\nCaused by: java.lang.RuntimeException: java.io.IOException: Failed during snapshot creation.\n at com.google.common.base.Throwables.propagate(Throwables.java:160) ~[guava-16.0.jar:na]\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:32) [apache-cassandra-2.1.0.jar:2.1.0]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_67]\n at java.util.concurrent.FutureTask.run(FutureTask.java:262) [na:1.7.0_67]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_67]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) ~[na:1.7.0_67]\n ... 1 common frames omitted\nCaused by: java.io.IOException: Failed during snapshot creation.\n at org.apache.cassandra.repair.RepairSession.failedSnapshot(RepairSession.java:344) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at org.apache.cassandra.repair.RepairJob$2.onFailure(RepairJob.java:128) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at com.google.common.util.concurrent.Futures$4.run(Futures.java:1172) ~[guava-16.0.jar:na]\n ... 3 common frames omitted\n{noformat}\nIf the index is dropped, the repair returns no error:\n{noformat}\ncqlsh:test> drop INDEX user_pass_idx ;\n\nroot@test:~# nodetool repair test user\n[2014-09-29 11:27:29,668] Starting repair command #6, repairing 743 ranges for keyspace test (seq=true, full=true)\n.\n.\n[2014-09-29 11:28:38,030] Repair session e6d40e10-47e4-11e4-ba0f-c7788dc924ec for range (-7298689860784559350,-7297558156602685286] finished\n[2014-09-29 11:28:38,030] Repair command #6 finished\n{noformat}\nThe test table:\n{noformat}\nCREATE TABLE test.user (\n login text PRIMARY KEY,\n password text\n)\ncreate INDEX user_pass_idx on test.user (password) ;\n{noformat}","issue_id":"12744762","key":"CASSANDRA-8020","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-10-02T17:19:02.000+0000","role":"fixed_distractor","summary":"nodetool repair on Cassandra 2.1.0 indexed tables returns java exception about creating snapshots"} {"case_id":"12761851","cluster":"DISTRACTOR-CASSANDRA-8487","comments":[{"body":"I am seeing the same problem on 2.1.0. The problem is reproducible in my test environment 1 out of 10 times.\n\nI have not been able to reproduce the problem in 2.1.3.","created":"2015-03-05T18:55:26.299+0000"},{"body":"[~danderson] Thanks. I'm going to close it as Cannot Reproduce (in 2.1.3). Can you please reopen if it happens again?","created":"2015-03-05T19:03:27.936+0000"},{"body":"Same goes to [~aholmber].","created":"2015-03-05T19:04:20.690+0000"},{"body":"I was able to reproduce this in C* 2.2.0-beta1 on Windows 64-bit. This is on a single-dc, 3 node cluster. In my case, I'm missing both system.schema_columnfamilies and system.schema_columns:\n\n{noformat}\ncqlsh> desc tables;\n\nKeyspace system_auth\n--------------------\nresource_role_permissons_index role_permissions role_members roles\n\nKeyspace system\n---------------\n\n\nKeyspace system_distributed\n---------------------------\nrepair_history parent_repair_history\n\nKeyspace system_traces\n----------------------\nevents sessions\n\ncqlsh> select * from system.schema_columnfamilies;\nColumn family 'schema_columnfamilies' not found\n{noformat}\n\nUsing python-driver 2.6.0-rc, I was able to do some further investigation:\n\n{noformat}\n>>> s.execute(\"select distinct keyspace_name from system.schema_columnfamilies\")\n[Row(keyspace_name=u'test3rf'), Row(keyspace_name=u'system_auth'), Row(keyspace_name=u'system_distributed'), Row(keyspac\ne_name=u'system'), Row(keyspace_name=u'system_traces')]\n>>> s.execute(\"select distinct keyspace_name from system.schema_columnfamilies\")\n[Row(keyspace_name=u'test3rf'), Row(keyspace_name=u'system_auth'), Row(keyspace_name=u'system_distributed'), Row(keyspac\ne_name=u'system'), Row(keyspace_name=u'system_traces')]\n>>> s.execute(\"select distinct keyspace_name from system.schema_columnfamilies\")\n[Row(keyspace_name=u'test3rf'), Row(keyspace_name=u'system_auth'), Row(keyspace_name=u'system_distributed'), Row(keyspac\ne_name=u'system_traces')]\n{noformat}\n\nSince the python-driver executes its queries with round-robin, we can see here that each SELECT query was routed to each node, and the 3rd one is missing the metadata about the keyspace \"system\". I'm not sure why I'm able to query against a \"nonexistent\" schema_columnfamilies table in the python-driver, but not in cqlsh. Does cqlsh run at CL.ALL or require metadata for formatting the results?","created":"2015-06-03T21:59:26.059+0000"},{"body":"bq. Does cqlsh run at CL.ALL or require metadata for formatting the results?\n\ncqlsh queries at ONE (by default), but does require metadata to format the results.","created":"2015-06-03T22:13:53.809+0000"},{"body":"The tables will always be there - they are hardcoded. You will always be able to query all system tables.\n\nPersistence of them into the on-disk data dictionary, though, for some reason seems to not work sometimes.\n\nSeems like all system.* tables are sometimes missing from columnfamilies/columns, which probably means some timestamp related issue.","created":"2015-06-03T22:19:08.849+0000"},{"body":"I'm able to repro the issue about 50% of the time on Windows when I start up a CCM cluster. Checking the system tables metadata in the python driver similarly just outputs an empty dict:\n{noformat}\nprint cluster.metadata.keyspaces['system'].tables\n{}\n{noformat}","created":"2015-06-03T22:45:11.630+0000"},{"body":"All right. I think I know what the issue is (and it is timestamp related). Let me cook up a patch for you to try.","created":"2015-06-03T22:52:20.900+0000"},{"body":"2.1 branch - https://github.com/iamaleksey/cassandra/commits/8487-2.1\n2.2 branch - https://github.com/iamaleksey/cassandra/commits/8487-2.2\n\nPlease let me know if you can reproduce there.","created":"2015-06-03T23:02:52.493+0000"},{"body":"I'm unable to reproduce on Windows against C* 2.1 and 2.2 with the patches above.","created":"2015-06-04T01:11:40.425+0000"},{"body":"cassci results will be available at:\n- http://cassci.datastax.com/view/Dev/view/iamaleksey/job/iamaleksey-8487-2.1-dtest/\n- http://cassci.datastax.com/view/Dev/view/iamaleksey/job/iamaleksey-8487-2.1-testall/\n- http://cassci.datastax.com/view/Dev/view/iamaleksey/job/iamaleksey-8487-2.2-dtest/\n- http://cassci.datastax.com/view/Dev/view/iamaleksey/job/iamaleksey-8487-2.2-testall/","created":"2015-06-04T10:51:45.230+0000"},{"body":"+1","created":"2015-06-04T15:18:32.996+0000"},{"body":"Committed to 2.1 as {{f1b22dfc4042cebb2de0e17d296ac4d7bc7d53df}} and merged into 2.2 and trunk. Thanks.","created":"2015-06-04T15:33:35.317+0000"}],"conversations":[{"body":"Occasionally a Cassandra node will have missing schema_columns information where keyspace_name='system'.\n\n{code}\ncqlsh> select * from system.schema_columns where keyspace_name='system';\n\n keyspace_name | columnfamily_name | column_name\n---------------+-------------------+-------------\n\n(0 rows)\n{code}\n\nAll keyspace and column family schema info is present for 'system' -- it's only the column information missing.\n\nThis can occur on an existing cluster following node restart. The data usually appears again after bouncing the node.\n\nThis is impactful to client drivers that expect column meta for configured tables.\n\nReproducible in 2.1.2. Have not seen it crop up in 2.0.11.","from":"reporter","subject":"system.schema_columns sometimes missing for 'system' keyspace"},{"body":"I am seeing the same problem on 2.1.0. The problem is reproducible in my test environment 1 out of 10 times.\n\nI have not been able to reproduce the problem in 2.1.3.","from":"developer"},{"body":"[~danderson] Thanks. I'm going to close it as Cannot Reproduce (in 2.1.3). Can you please reopen if it happens again?","from":"developer"},{"body":"Same goes to [~aholmber].","from":"developer"},{"body":"I was able to reproduce this in C* 2.2.0-beta1 on Windows 64-bit. This is on a single-dc, 3 node cluster. In my case, I'm missing both system.schema_columnfamilies and system.schema_columns:\n\n{noformat}\ncqlsh> desc tables;\n\nKeyspace system_auth\n--------------------\nresource_role_permissons_index role_permissions role_members roles\n\nKeyspace system\n---------------\n\n\nKeyspace system_distributed\n---------------------------\nrepair_history parent_repair_history\n\nKeyspace system_traces\n----------------------\nevents sessions\n\ncqlsh> select * from system.schema_columnfamilies;\nColumn family 'schema_columnfamilies' not found\n{noformat}\n\nUsing python-driver 2.6.0-rc, I was able to do some further investigation:\n\n{noformat}\n>>> s.execute(\"select distinct keyspace_name from system.schema_columnfamilies\")\n[Row(keyspace_name=u'test3rf'), Row(keyspace_name=u'system_auth'), Row(keyspace_name=u'system_distributed'), Row(keyspac\ne_name=u'system'), Row(keyspace_name=u'system_traces')]\n>>> s.execute(\"select distinct keyspace_name from system.schema_columnfamilies\")\n[Row(keyspace_name=u'test3rf'), Row(keyspace_name=u'system_auth'), Row(keyspace_name=u'system_distributed'), Row(keyspac\ne_name=u'system'), Row(keyspace_name=u'system_traces')]\n>>> s.execute(\"select distinct keyspace_name from system.schema_columnfamilies\")\n[Row(keyspace_name=u'test3rf'), Row(keyspace_name=u'system_auth'), Row(keyspace_name=u'system_distributed'), Row(keyspac\ne_name=u'system_traces')]\n{noformat}\n\nSince the python-driver executes its queries with round-robin, we can see here that each SELECT query was routed to each node, and the 3rd one is missing the metadata about the keyspace \"system\". I'm not sure why I'm able to query against a \"nonexistent\" schema_columnfamilies table in the python-driver, but not in cqlsh. Does cqlsh run at CL.ALL or require metadata for formatting the results?","from":"developer"},{"body":"bq. Does cqlsh run at CL.ALL or require metadata for formatting the results?\n\ncqlsh queries at ONE (by default), but does require metadata to format the results.","from":"developer"},{"body":"The tables will always be there - they are hardcoded. You will always be able to query all system tables.\n\nPersistence of them into the on-disk data dictionary, though, for some reason seems to not work sometimes.\n\nSeems like all system.* tables are sometimes missing from columnfamilies/columns, which probably means some timestamp related issue.","from":"developer"},{"body":"I'm able to repro the issue about 50% of the time on Windows when I start up a CCM cluster. Checking the system tables metadata in the python driver similarly just outputs an empty dict:\n{noformat}\nprint cluster.metadata.keyspaces['system'].tables\n{}\n{noformat}","from":"developer"},{"body":"All right. I think I know what the issue is (and it is timestamp related). Let me cook up a patch for you to try.","from":"developer"},{"body":"2.1 branch - https://github.com/iamaleksey/cassandra/commits/8487-2.1\n2.2 branch - https://github.com/iamaleksey/cassandra/commits/8487-2.2\n\nPlease let me know if you can reproduce there.","from":"developer"},{"body":"I'm unable to reproduce on Windows against C* 2.1 and 2.2 with the patches above.","from":"developer"},{"body":"cassci results will be available at:\n- http://cassci.datastax.com/view/Dev/view/iamaleksey/job/iamaleksey-8487-2.1-dtest/\n- http://cassci.datastax.com/view/Dev/view/iamaleksey/job/iamaleksey-8487-2.1-testall/\n- http://cassci.datastax.com/view/Dev/view/iamaleksey/job/iamaleksey-8487-2.2-dtest/\n- http://cassci.datastax.com/view/Dev/view/iamaleksey/job/iamaleksey-8487-2.2-testall/","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed to 2.1 as {{f1b22dfc4042cebb2de0e17d296ac4d7bc7d53df}} and merged into 2.2 and trunk. Thanks.","from":"developer"}],"created":"2014-12-15T20:42:11.000+0000","description":"Occasionally a Cassandra node will have missing schema_columns information where keyspace_name='system'.\n\n{code}\ncqlsh> select * from system.schema_columns where keyspace_name='system';\n\n keyspace_name | columnfamily_name | column_name\n---------------+-------------------+-------------\n\n(0 rows)\n{code}\n\nAll keyspace and column family schema info is present for 'system' -- it's only the column information missing.\n\nThis can occur on an existing cluster following node restart. The data usually appears again after bouncing the node.\n\nThis is impactful to client drivers that expect column meta for configured tables.\n\nReproducible in 2.1.2. Have not seen it crop up in 2.0.11.","issue_id":"12761851","key":"CASSANDRA-8487","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-06-04T15:34:07.000+0000","role":"fixed_distractor","summary":"system.schema_columns sometimes missing for 'system' keyspace"} {"case_id":"12762388","cluster":"DISTRACTOR-CASSANDRA-8505","comments":[{"body":"Is this too big of a change in existing behavior to put into 2.0 or 2.1?","created":"2015-02-03T23:26:31.927+0000"},{"body":"We probably need write the patch before deciding (in theory, it would be rather simple to have SelectStatement check if the index is built, but in practice doing so means a query to the a system table and that's probably not acceptable performance, so we'd need to cache the info (probably directly in ColumnDefinition?) but we'll have to check how hairy that is exactly). That said, even if we decide it's too hairy for 2.0, we should at least fix it in 2.1.","created":"2015-02-04T15:01:59.286+0000"},{"body":"Secondary index and their build/not build status are node-local. By consequence it is not possible to know on a coordinator node if the index is fully build. It can be built on the coordinator but still building on other nodes. Further more an index rebuild can be triggered at any time.\nTherefore the only moment where we can check if the index is ready is at query execution time.\n\nThe first problem of rejecting index queries at execution time is that some {{ALLOW FILTERING}} queries that could have been processed without an index will be rejected. As the {{ALLOW FILTERING}} information is not passed with the command we have no way to know if the query should be executed or not using filtering. On the other hand, currently, if an index exists but is not built Cassandra might silently return the wrong results. By consequence rejecting the query is still an improvement, in my opinion, and we can create a new ticket to improve the situation in the future.\n\nThe second problem if about communicating back the error to the coordinator node. CASSANDRA-7886 added a mechanism for that but it is not perfect. The user will receive a {{ReadFailureException}} but would have to look within the logs to find the root cause of the problem. Ideally this mechanism should be improved to be able to pass the error message to the {{ReadFailureException}}. The other problem of the mechanism is that it is only available since {{2.2}}, so I could not create a patch for {{2.1}}.\n\nThe patch for {{2.2}} is [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-2.2] and the patch for {{3.0}} is [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-3.0]\n\nBoth patches keep the index state in memory and throw an Exception if the index is not ready when a request arrive.\nThe paches also shortcut the building of a index if the base table is empty. This optimisation prevent a lot of the existing index tests to fail.\n\n*The unit test results for {{2.2}} are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-2.2-testall/3/] \n*The dtest results for {{2.2}} are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-2.2-dtest/3/] \n*The unit test results for {{3.0}} are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-3.0-testall/1/] \n*The dtest results for {{3.0}} are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-3.0-dtest/1/] \n\nThe {{secondary_indexes_test.TestSecondaryIndexesOnCollections.test_map_indexes}} dtest fails in {{2.2}} because it is not waiting for the index to be built before querying the index. I will provide a patch for the DTest. \n\n ","created":"2015-10-29T16:56:13.792+0000"},{"body":"I think the various index states can be reduced to a simple ready/not ready check. What's more unless we intend to change the established behaviour fairly significantly, once an index moves to a ready state it never moves back to being not ready. The only times when we modify the status in the system table are when the index is removed (in which case we have no problem with being able to query using it) or during a rebuild. In the latter case though, we probably shouldn't reject queries (and we don't currently), as an index rebuild is incremental. That is, we don't scrap the existing index tables and rebuild everything from scratch, just write new index SSTables to supercede the old ones. So although it's certainly possible to get incorrect results during a rebuild (because of missing/stale entries), the results only get more correct as the rebuild progresses. Changing this so that all queries against that index return errors until all rebuilds complete seems like a step backwards. It seems more reasonable to reject queries until the initial build has been performed, as per the example in the description, but this only requires a simple boolean to track state between instantiating/registering the index and its initial build task completing (if one is required). \n\nIt would be good to have some test coverage of this, although the best I could come up with is a dtest which inserts many rows, then adds the index and queries immediately expecting ReadFailureException, which is fairly lame and fragile.\n\nA couple of points specific to the 3.0 patch:\n\n* The fix for CASSANDRA-10595 has been lost. If an index doesn't register itself in {{createIndex}}, don't ask it for an initalization task, just set {{initialBuildTask == null}}. \n* {{SIM::reloadIndex}} has changed since the patch was created (due to CASSANDRA-10604) - I think that no changes to this method are now required. I did notice though that the current implementation actually makes a redundant call to {{getMetadataReloadTask}}, so if you could fix that while you're here, that'd be great.\n\nbq. Secondary index and their build/not build status are node-local. By consequence it is not possible to know on a coordinator node if the index is fully build. It can be built on the coordinator but still building on other nodes\n\nFor future reference on this point, we also have CASSANDRA-9967 which has a very similar intent.","created":"2015-11-11T17:46:20.245+0000"},{"body":"bq. It would be good to have some test coverage of this, although the best I could come up with is a dtest which inserts many rows, then adds the index and queries immediately expecting ReadFailureException, which is fairly lame and fragile.\n\nWhat about accepting either {{ReadFailureException}} or the complete, correct result? If index building got faster we might stop hitting the {{ReadFailureException}} case, but at least the test wouldn't flap.","created":"2015-11-11T19:22:57.753+0000"},{"body":"I think we do not really have the choice. I could easily reproduce the problem with a unit test on my machine but CI is usually much slower.\nFor reproducing it with cqlsh I add to put a breack point in the building task.","created":"2015-11-11T20:22:29.880+0000"},{"body":"bq. What about accepting either ReadFailureException or the complete, correct result? If index building got faster we might stop hitting the ReadFailureException case, but at least the test wouldn't flap.\n\nYes absolutely, even then though we're not going to be certain exactly what's being tested (if at all) - e.g. the index building gets faster but we also introduce a regression with the {{ReadFailureException}}, we'd never know. But like I say, I can't think of anything better.","created":"2015-11-11T20:44:54.996+0000"},{"body":"It occurred to me that we could also create a custom secondary index that delayed the build completion, either by waiting for some sort of signal or by sleeping.","created":"2015-11-11T21:40:47.312+0000"},{"body":"We could certainly do that in a utest (and we have plenty of tests with such custom indexes), but it not a dtest as it would require the custom index to be on the classpath. Naturally, a utest won't exercise the distributed side of things, but it's still better than no testing, so +1","created":"2015-11-11T21:55:23.422+0000"},{"body":"I have pushed the fixes for 2.2 and 3.0 [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-2.2] and [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-3.0].\n\n*The unit test results for 2.2 are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-2.2-testall/5/]\n*The dtest results for 2.2 are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-2.2-dtest/4/]\n*The unit test results for 3.0 are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-3.0-testall/3/]\n*The dtest results for 3.0 are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-3.0-dtest/3/]","created":"2015-11-18T16:24:13.714+0000"},{"body":"Looks pretty good to me, I just have a small nit & one question/suggestion about naming:\n\nFirstly, the catch blocks for {{TombstoneOverwhelmingException}} and {{IndexNotAvailableException}} in {{MessageDeliveryTask::run}} can be combined into one multi catch. Also, that catch for {{IndexNotAvailableException}} was only added in 3.0, and it seems it could also be done for the 2.2 branch.\n\nOn the question of naming, I wonder if perhaps the names of the new methods which check the state of the indexes ought to reflect the fact that the readiness is in regard to querying (i.e. an index will process updates as soon as it's created, the new flag just guards against it's use in queries). So, on the 2.2 branch, we could rename {{SI::isReady}} to {{SI::isQueryable}} and on 3.0 {{SIM::isIndexReady}} -> {{SIM::isIndexQueryable}}. \nDo you have any thoughts on that?\n","created":"2015-11-23T14:42:05.010+0000"},{"body":"{quote}Do you have any thoughts on that?{quote}\nIt makes sense to me.\n\nI have pushed the fixes for the 2 comments [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-2.2] and [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-3.0].\n\nI fixed the DTest for {{secondary_indexes_test.TestSecondaryIndexesOnCollections.test_map_indexes}} [here|https://github.com/riptano/cassandra-dtest/pull/685/files]\n\n[~philipthompson], [~mambocab] The changes of this ticket might cause some dtests to fail randomly (if a lot of data were inserted before the index was created). I had a look at the DTests but I might have missed some. If some tests using secondary index start failing once this ticket is committed do not hesitate to assign them to me.","created":"2015-11-25T10:46:31.426+0000"},{"body":"[~blerer] Noted, thanks for the warning.","created":"2015-11-25T14:34:43.964+0000"},{"body":"+1","created":"2015-11-26T10:11:07.923+0000"},{"body":"Thanks for the review.","created":"2015-11-27T10:20:31.302+0000"},{"body":"Committed in 2.2 at 61e0251a1d4ddc695382aee11e443506afd40899 and merged into 3.0, 3.1 and trunk","created":"2015-11-27T10:21:40.404+0000"}],"conversations":[{"body":"If you request an index creation and then execute a query that use the index the results returned might be invalid until the index is fully build. This is caused by the fact that the table column will be marked as indexed before the index is ready.\n\nThe following unit tests can be use to reproduce the problem:\n{code}\n @Test\n public void testIndexCreatedAfterInsert() throws Throwable\n {\n createTable(\"CREATE TABLE %s (a int, b int, c int, primary key((a, b)))\");\n\n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 0, 0);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 1, 1);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 2, 2);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (1, 0, 3);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (1, 1, 4);\");\n \n createIndex(\"CREATE INDEX ON %s(b)\");\n \n assertRows(execute(\"SELECT * FROM %s WHERE b = ?;\", 1),\n row(0, 1, 1),\n row(1, 1, 4));\n }\n \n @Test\n public void testIndexCreatedBeforeInsert() throws Throwable\n {\n createTable(\"CREATE TABLE %s (a int, b int, c int, primary key((a, b)))\");\n\n createIndex(\"CREATE INDEX ON %s(b)\");\n \n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 0, 0);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 1, 1);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 2, 2);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (1, 0, 3);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (1, 1, 4);\");\n\n assertRows(execute(\"SELECT * FROM %s WHERE b = ?;\", 1),\n row(0, 1, 1),\n row(1, 1, 4));\n }\n{code}\n\nThe first test will fail while the second will work. \n\nIn my opinion the first test should reject the request as invalid (as if the index was not existing) until the index is fully build.","from":"reporter","subject":"Invalid results are returned while secondary index are being build"},{"body":"Is this too big of a change in existing behavior to put into 2.0 or 2.1?","from":"developer"},{"body":"We probably need write the patch before deciding (in theory, it would be rather simple to have SelectStatement check if the index is built, but in practice doing so means a query to the a system table and that's probably not acceptable performance, so we'd need to cache the info (probably directly in ColumnDefinition?) but we'll have to check how hairy that is exactly). That said, even if we decide it's too hairy for 2.0, we should at least fix it in 2.1.","from":"developer"},{"body":"Secondary index and their build/not build status are node-local. By consequence it is not possible to know on a coordinator node if the index is fully build. It can be built on the coordinator but still building on other nodes. Further more an index rebuild can be triggered at any time.\nTherefore the only moment where we can check if the index is ready is at query execution time.\n\nThe first problem of rejecting index queries at execution time is that some {{ALLOW FILTERING}} queries that could have been processed without an index will be rejected. As the {{ALLOW FILTERING}} information is not passed with the command we have no way to know if the query should be executed or not using filtering. On the other hand, currently, if an index exists but is not built Cassandra might silently return the wrong results. By consequence rejecting the query is still an improvement, in my opinion, and we can create a new ticket to improve the situation in the future.\n\nThe second problem if about communicating back the error to the coordinator node. CASSANDRA-7886 added a mechanism for that but it is not perfect. The user will receive a {{ReadFailureException}} but would have to look within the logs to find the root cause of the problem. Ideally this mechanism should be improved to be able to pass the error message to the {{ReadFailureException}}. The other problem of the mechanism is that it is only available since {{2.2}}, so I could not create a patch for {{2.1}}.\n\nThe patch for {{2.2}} is [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-2.2] and the patch for {{3.0}} is [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-3.0]\n\nBoth patches keep the index state in memory and throw an Exception if the index is not ready when a request arrive.\nThe paches also shortcut the building of a index if the base table is empty. This optimisation prevent a lot of the existing index tests to fail.\n\n*The unit test results for {{2.2}} are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-2.2-testall/3/] \n*The dtest results for {{2.2}} are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-2.2-dtest/3/] \n*The unit test results for {{3.0}} are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-3.0-testall/1/] \n*The dtest results for {{3.0}} are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-3.0-dtest/1/] \n\nThe {{secondary_indexes_test.TestSecondaryIndexesOnCollections.test_map_indexes}} dtest fails in {{2.2}} because it is not waiting for the index to be built before querying the index. I will provide a patch for the DTest. \n\n ","from":"developer"},{"body":"I think the various index states can be reduced to a simple ready/not ready check. What's more unless we intend to change the established behaviour fairly significantly, once an index moves to a ready state it never moves back to being not ready. The only times when we modify the status in the system table are when the index is removed (in which case we have no problem with being able to query using it) or during a rebuild. In the latter case though, we probably shouldn't reject queries (and we don't currently), as an index rebuild is incremental. That is, we don't scrap the existing index tables and rebuild everything from scratch, just write new index SSTables to supercede the old ones. So although it's certainly possible to get incorrect results during a rebuild (because of missing/stale entries), the results only get more correct as the rebuild progresses. Changing this so that all queries against that index return errors until all rebuilds complete seems like a step backwards. It seems more reasonable to reject queries until the initial build has been performed, as per the example in the description, but this only requires a simple boolean to track state between instantiating/registering the index and its initial build task completing (if one is required). \n\nIt would be good to have some test coverage of this, although the best I could come up with is a dtest which inserts many rows, then adds the index and queries immediately expecting ReadFailureException, which is fairly lame and fragile.\n\nA couple of points specific to the 3.0 patch:\n\n* The fix for CASSANDRA-10595 has been lost. If an index doesn't register itself in {{createIndex}}, don't ask it for an initalization task, just set {{initialBuildTask == null}}. \n* {{SIM::reloadIndex}} has changed since the patch was created (due to CASSANDRA-10604) - I think that no changes to this method are now required. I did notice though that the current implementation actually makes a redundant call to {{getMetadataReloadTask}}, so if you could fix that while you're here, that'd be great.\n\nbq. Secondary index and their build/not build status are node-local. By consequence it is not possible to know on a coordinator node if the index is fully build. It can be built on the coordinator but still building on other nodes\n\nFor future reference on this point, we also have CASSANDRA-9967 which has a very similar intent.","from":"developer"},{"body":"bq. It would be good to have some test coverage of this, although the best I could come up with is a dtest which inserts many rows, then adds the index and queries immediately expecting ReadFailureException, which is fairly lame and fragile.\n\nWhat about accepting either {{ReadFailureException}} or the complete, correct result? If index building got faster we might stop hitting the {{ReadFailureException}} case, but at least the test wouldn't flap.","from":"developer"},{"body":"I think we do not really have the choice. I could easily reproduce the problem with a unit test on my machine but CI is usually much slower.\nFor reproducing it with cqlsh I add to put a breack point in the building task.","from":"developer"},{"body":"bq. What about accepting either ReadFailureException or the complete, correct result? If index building got faster we might stop hitting the ReadFailureException case, but at least the test wouldn't flap.\n\nYes absolutely, even then though we're not going to be certain exactly what's being tested (if at all) - e.g. the index building gets faster but we also introduce a regression with the {{ReadFailureException}}, we'd never know. But like I say, I can't think of anything better.","from":"developer"},{"body":"It occurred to me that we could also create a custom secondary index that delayed the build completion, either by waiting for some sort of signal or by sleeping.","from":"developer"},{"body":"We could certainly do that in a utest (and we have plenty of tests with such custom indexes), but it not a dtest as it would require the custom index to be on the classpath. Naturally, a utest won't exercise the distributed side of things, but it's still better than no testing, so +1","from":"developer"},{"body":"I have pushed the fixes for 2.2 and 3.0 [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-2.2] and [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-3.0].\n\n*The unit test results for 2.2 are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-2.2-testall/5/]\n*The dtest results for 2.2 are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-2.2-dtest/4/]\n*The unit test results for 3.0 are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-3.0-testall/3/]\n*The dtest results for 3.0 are [here|http://cassci.datastax.com/view/Dev/view/blerer/job/blerer-8505-3.0-dtest/3/]","from":"developer"},{"body":"Looks pretty good to me, I just have a small nit & one question/suggestion about naming:\n\nFirstly, the catch blocks for {{TombstoneOverwhelmingException}} and {{IndexNotAvailableException}} in {{MessageDeliveryTask::run}} can be combined into one multi catch. Also, that catch for {{IndexNotAvailableException}} was only added in 3.0, and it seems it could also be done for the 2.2 branch.\n\nOn the question of naming, I wonder if perhaps the names of the new methods which check the state of the indexes ought to reflect the fact that the readiness is in regard to querying (i.e. an index will process updates as soon as it's created, the new flag just guards against it's use in queries). So, on the 2.2 branch, we could rename {{SI::isReady}} to {{SI::isQueryable}} and on 3.0 {{SIM::isIndexReady}} -> {{SIM::isIndexQueryable}}. \nDo you have any thoughts on that?\n","from":"developer"},{"body":"{quote}Do you have any thoughts on that?{quote}\nIt makes sense to me.\n\nI have pushed the fixes for the 2 comments [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-2.2] and [here|https://github.com/apache/cassandra/compare/trunk...blerer:8505-3.0].\n\nI fixed the DTest for {{secondary_indexes_test.TestSecondaryIndexesOnCollections.test_map_indexes}} [here|https://github.com/riptano/cassandra-dtest/pull/685/files]\n\n[~philipthompson], [~mambocab] The changes of this ticket might cause some dtests to fail randomly (if a lot of data were inserted before the index was created). I had a look at the DTests but I might have missed some. If some tests using secondary index start failing once this ticket is committed do not hesitate to assign them to me.","from":"developer"},{"body":"[~blerer] Noted, thanks for the warning.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Thanks for the review.","from":"developer"},{"body":"Committed in 2.2 at 61e0251a1d4ddc695382aee11e443506afd40899 and merged into 3.0, 3.1 and trunk","from":"developer"}],"created":"2014-12-17T21:09:42.000+0000","description":"If you request an index creation and then execute a query that use the index the results returned might be invalid until the index is fully build. This is caused by the fact that the table column will be marked as indexed before the index is ready.\n\nThe following unit tests can be use to reproduce the problem:\n{code}\n @Test\n public void testIndexCreatedAfterInsert() throws Throwable\n {\n createTable(\"CREATE TABLE %s (a int, b int, c int, primary key((a, b)))\");\n\n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 0, 0);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 1, 1);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 2, 2);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (1, 0, 3);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (1, 1, 4);\");\n \n createIndex(\"CREATE INDEX ON %s(b)\");\n \n assertRows(execute(\"SELECT * FROM %s WHERE b = ?;\", 1),\n row(0, 1, 1),\n row(1, 1, 4));\n }\n \n @Test\n public void testIndexCreatedBeforeInsert() throws Throwable\n {\n createTable(\"CREATE TABLE %s (a int, b int, c int, primary key((a, b)))\");\n\n createIndex(\"CREATE INDEX ON %s(b)\");\n \n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 0, 0);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 1, 1);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (0, 2, 2);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (1, 0, 3);\");\n execute(\"INSERT INTO %s (a, b, c) VALUES (1, 1, 4);\");\n\n assertRows(execute(\"SELECT * FROM %s WHERE b = ?;\", 1),\n row(0, 1, 1),\n row(1, 1, 4));\n }\n{code}\n\nThe first test will fail while the second will work. \n\nIn my opinion the first test should reject the request as invalid (as if the index was not existing) until the index is fully build.","issue_id":"12762388","key":"CASSANDRA-8505","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-11-27T10:21:40.000+0000","role":"fixed_distractor","summary":"Invalid results are returned while secondary index are being build"} {"case_id":"12762978","cluster":"DISTRACTOR-CASSANDRA-8523","comments":[{"body":"I've checked with [~brandon.williams], and this is intended behavior, so I'm marking this as an Improvement, not a bug. Do note that the replacement node will receive all writes while it was streaming that were captured as hints, so if the stream takes less than the hint window, you should not see too much discrepancy.","created":"2015-01-09T16:13:50.489+0000"},{"body":"I completely agree this is an improvement, but it's going to be pretty tricky, especially since we can't use the FD to determine if the node has died, at least not in its current form since that would mark the node as UP.","created":"2015-01-09T16:16:59.956+0000"},{"body":"Marking related to CASSANDRA-8494, but may want to resolve as duplicate.","created":"2015-01-13T20:18:49.295+0000"},{"body":"I'm out of ideas here. Signalling that we want to send writes to the replacing node is easy enough, but knowing when to stop if it fails is tricky, since we can't use the FD and it's a singleton so we can't just grab a new instance. [~jkni], [~Stefania] do you have any ideas?","created":"2016-04-04T22:04:50.980+0000"},{"body":"I would consider a new, non dead, state for replacing a node rather than using hibernate. This should allow us to use the Failure Detector right?\n\nAlternatively, would a new Gossip property that indicates a \"shadow\" of an existing node be easier to implement? Could this property be made to expire unless it gets renewed by the shadow node? \n","created":"2016-04-05T01:20:30.011+0000"},{"body":"bq. I would consider a new, non dead, state for replacing a node rather than using hibernate. This should allow us to use the Failure Detector right?\n\nIt looks like this would work now, because the gossiper is the only thing that registers with the FD. In the past this wasn't always true, though.\n\nbq. Alternatively, would a new Gossip property that indicates a \"shadow\" of an existing node be easier to implement?\n\nHmm, can you can explain further?","created":"2016-04-05T01:26:23.089+0000"},{"body":"bq. Hmm, can you can explain further?\n\nWell, it would be a bit convoluted so a new Gossip state would probably be cleaner if we can add it. But the idea is to link a host to one or more hosts that act as 'shadows', that is they receive copies of all writes or hints but are not queried for reads. ","created":"2016-04-05T01:33:08.739+0000"},{"body":"There are two scenarios we should consider when replacing a node:\n1) The replacing node has the same IP as the previous node\n2) The replacing node has a different IP as the previous node\n\nOn CASSANDRA-9244 I have gotten pretty far in an [implementation|https://issues.apache.org/jira/browse/CASSANDRA-9244?focusedCommentId=15211202&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15211202] that adds a new non-dead gossip state {{BOOT_REPLACE}} and considers the replacing endpoint as a bootstrapping pending endpoint, solving case 2 transparently.\n\nCase 1 is trickier because when the replacing node enters gossip with a non-dead state, other nodes will think the previous node is back up and send reads to him (since he is a natural endpoint).\n\nA simple way to solve this is to special-case the read path and ignore nodes in \"NON-NORMAL\" state when sending reads to natural endpoints. While this will probably solve the problem, there are quite a few different paths we need to hack to make sure this is enforced correctly (paxos, read, hints, etc), so I'm not totally comfortable with that.\n\nA more transparent but a bit costlier approach to solve case 1 would be to change the {{TokenMetadata}} to keep nodes as {{(InetAddress, UUID)}} pairs, and create a new interface to the {{FailureDetector}} keyed by {{UUID}}. This way we could keep {{(IP=127.0.0.1,UUID=1)}} in {{TokenMetadata}} as a natural endpoint, and add a replacement node {{(IP=127.0.0.1,UUID=2)}} as a pending endpoint. So, during reads, {{FD.isAlive(UUID=1)}} would return false, and natural reads would not be sent to {{(IP=127.0.0.1,UUID=1)}}, while pending writes would be sent to {{(IP=127.0.0.1,UUID=2)}} because {{FD.isAlive(UUID=2)}} would return true.\n\nI'd be happy to continue working on this, so feedback on any of the above or alternative approaches would be greatly appreciated.","created":"2016-04-06T14:19:33.394+0000"},{"body":"I don't have a good idea for a low-hanging solution here - I think the cleanest is the combination of a new gossip state for replacing nodes + enhancement of the failure detector/token metadata. Since that's a bigger change, let's also make sure that it doesn't conflict with plans for strongly consistent membership. I'd really rather not couple gossip status to the internals of the read/write paths.\n\nI also agree that we should make sure that we kill two birds with one stone and handle this and [CASSANDRA-9244] at the same time.\n\nLet me know if I can help in design and/or review.","created":"2016-04-06T15:34:40.839+0000"},{"body":"I agree with [~jkni] - a new gossip state and updating FD/TMD is the best way to go. I'll follow up on #9244 to get better grasp of what's going on there, as well.\n\nwrt strongly consistent membership, that effort focuses primarily on correct and linearizable changes to the cluster state machine, and not any behaviors that should be triggered due to those changes when they arrive at cluster nodes. So, I don't *think* we'd run into problems between these two efforts, but good to keep the same players involved :)","created":"2016-04-15T23:21:04.204+0000"},{"body":"[~pauloricardomg] assigning to you, please do continue working on this.","created":"2016-05-09T21:22:54.839+0000"},{"body":"Initial route I will take is to create a new dead state for replace, that adds the node as replacing endpoint to TokenMetadata. When a replacement endpoint is added to TokenMetadata, if there's an existing down node with the same IP, then we remove it from natural endpoints so the replacement node no longer receive reads when it becomes alive in the FD and include the replacement node as pending joining endpoint so writes are forwarded to it. After pending ranges for replace are calculated the final step is to set the node as alive in the FD and change the dead state logic to not mark \"alive\" dead state nodes as DOWN automatically (we will probably rename this nomenclature to a better name, like invisible or whatever), so the FD will work as usual and remove the replacing endpoint from TokenMetadata (and restore it as a natural endpoint) if it becomes down. The current streaming logic will probably be unaffected by this, and the node will change its state to NORMAL after stream is finished completing the replace procedure.\n\nThe only downside I can think of this approach is that we will lose hints during a failed replace, but this is not a big deal as hints are an optimization, and replace will probably take longer than max_hint_window anyway.\n\nI will start going this route, feel free to give any feedback or let me know if I'm missing something on this high level flow.","created":"2016-05-27T16:29:21.515+0000"},{"body":"Without understanding the FD details, this sounds good. Losing hints isn't an issue, as you say.","created":"2016-06-02T01:04:53.886+0000"},{"body":"Solving this when the replacing node has a different ip address is quite simple:\n* Add a new non-dead gossip state for replace BOOT_REPLACE\n* When receiving BOOT_REPLACE, other node adds the replacing node as bootstrapping AND replacing endpoint\n* Pending ranges are calculated, and writes are sent to the replacing node during replace\n* When replacing node changes state to NORMAL, the old node is removed and the new node becomes a natural endpoint on {{TokenMetadata}}\n\nI have a WIP patch with this approach [here|https://github.com/pauloricardomg/cassandra/commit/a4be7f7a0f298b1d52821f67ff2e927a1d1f1c5e] that is functional. The tricky part is when the replacing node has the same IP as the old node, because if this node ever joins gossip with a non-dead state, {{FailureDetector.isAlive(InetAddress)}} return true and reads are sent to the replacing node, since the original node is a natural endpoint.\n\nMy initial idea of removing the old node as a natural endpoint from {{TokenMetadata}} obviously does not work because this changes range assignments making reads go the following node in the ring, which has no data for that range. I also considered replacing the node IP in TokenMetadata with a special marker, but this will also not play along well with snitches and topologies.\n\nA clean fix to this is to enhance the node representation on {{TokenMetadata}}, which is currently a plain {{InetAddresss}}, to include the information if an endpoint is available for reads (CASSANDRA-11559), but this cannot be done before cassandra 4.0 because it will change {{TokenMetadata}} and {{AbstractReplicationStrategy}} public interfaces and other related classes, so it's a major refactoring in the codebase. One step further is to change the {{FailureDetector}} interface to query nodes by {{UUIDs}} instead of {{InetAddress}}, so a replacing endpoint will not be confused with the original node since they will have different ids.\n\nOne workaround before that would be to send a read to a normal endpoint if both {{FailureDetector.isAlive(endpoint)}} is true and the endpoint is on {{NORMAL}} state, so we avoid sending reads to a replacing endpoint with the same IP address. A big downside is that this check should also be peformed for other operations such as bootstrap, repairs, hints, batches, etc. I experimented with that route [in this commit|https://github.com/pauloricardomg/cassandra/commit/a439edcbd8d4186ea2e3301681b7cc0b5012a265] by creating a method {{StorageService.isAvailable(InetAddress endpoint, boolean isPending) = \\{FailureDetector.instance.isAlive(endpoint) && (isPendingEndpoint || Gossiper.instance.isNormalStatus(endpoint))\\}}} that should be used instead of {{FailureDetector.isAlive(endpoint)}}, but I'm not very comfortable with this approach for the following reasons:\n* It's pretty fragile, because that check must be used instead of {{FailureDetector.isAlive()}} in every place throughout the code\n* Other developers/drivers/tools wil also need to use the same approach or a similar logic\n* We will need to modify the write path to segregate writes between natural and pending endpoints (see an example on [doPaxosCommit|https://github.com/pauloricardomg/cassandra/commit/a439edcbd8d4186ea2e3301681b7cc0b5012a265#diff-71f06c193f5b5e270cf8ac695164f43aR498])\n\nSo unless someone comes up with a magic idea, I think we're left with two options to fix this before CASSANDRA-11559:\n1. Go ahead with {{StorageService.isAvailable}} approach even with its downsides and fragility\n2. Fix this only for replacing endpoints with a different IP address, falling back to the previous behavior when the replacing node has the same IP as the old node, which should already fix this in most cloud deployments, where replacement nodes usually have different IPs\n\nI'm leaning more towards 2 since it's much simpler/safer and work on a proper general solution for 4.0 after CASSANDRA-11559. WDYT?","created":"2016-06-02T16:19:10.097+0000"},{"body":"I had a closer look at CASSANDRA-11559 and it should be possible to support that on trunk/3.x series without breaking compatibility, thus making this viable before 4.0 in a robust way. I will provide an initial patch for that soon.\n\nWe now need to decide whether we want to have a partial/fragile solution to this ticket before the ideal fix on 3.x leveraging #11559.","created":"2016-06-03T00:49:56.098+0000"},{"body":"Due to the limitations of forwarding writes to replacement nodes with the same IP, I propose initially adding this support only to replacement nodes with a different IP, since it's much simpler and we can do it in a backward-compatible way so it can probably go on 2.2+.\n\nAfter CASSANDRA-11559, we can extend this support to nodes with the same IP quite easily by setting an inactive flag on nodes being replaced and ignore these nodes on read.\n\nThe central idea is:\n{quote}\n* Add a new non-dead gossip state for replace BOOT_REPLACE\n* When receiving BOOT_REPLACE, other node adds the replacing node as bootstrapping endpoint\n* Pending ranges are calculated, and writes are sent to the replacing node during replace\n* When replacing node changes state to NORMAL, the old node is removed and the new node becomes a natural endpoint on TokenMetadata\n{quote}\n\nSince it's no longer necessary to forward hints to the replacement node when {{replace_address != broadcast_address}}, the replacement node does not need to inherit the same ID of the original node.\n\nThe replacing process remains unchanged when the replacement node has the same IP as the original node. If that's the case, I added a warn message so users know they need to run repair if the node is down for longer than {{max_hint_window_in_ms}}:\n{noformat}\nWrites will not be redirected to this node while it is performing replace because it has the same address as the node to be replaced ({}). \nIf that node has been down for longer than max_hint_window_in_ms, repair must be run after the replacement process in order to make this node consistent.\n{noformat}\n\nI adapted current dtests to test replace_address for both the old and the new path, and when {{replace_address != broadcast_address}} make sure writes are being redirected to the replacement node.\n\nInitial patch and tests below (will provide 2.2+ patches after initial review):\n||2.2||dtest||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-8523]|[branch|https://github.com/riptano/cassandra-dtest/compare/master...pauloricardomg:8523]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-dtest/lastCompletedBuild/testReport/]|","created":"2016-06-18T00:21:28.752+0000"},{"body":"Thanks! I'm happy to review.","created":"2016-06-18T00:23:09.279+0000"},{"body":"bq. Since it's no longer necessary to forward hints to the replacement node when replace_address != broadcast_address, the replacement node does not need to inherit the same ID of the original node.\n\nI don't think that's necessarily true, if a node dies at point A, and then the replacement is issued at point N, the hints between points B and M are still missed if the host id is not the same.","created":"2016-06-18T00:26:09.576+0000"},{"body":"It's not those writes that matter - the replacement will get those writes from the other nodes during streaming. The hints that you might care about are writes dropped during the replacement on the replacing node. But those should be extremely rare, and a price well worth paying for getting the vast majority of the writes vs none today.","created":"2016-06-18T00:32:08.365+0000"},{"body":"Missed that comment above - [~rlow] to review. Thanks!","created":"2016-06-20T19:08:24.843+0000"},{"body":"bq. The hints that you might care about are writes dropped during the replacement on the replacing node. \nEven these should not be a problem, because they are also forwarded to the replacement node, and if they fail they are hinted to replacement node ID.\n\nI rebased and resubmitted dtests and fixed minor typo on the original patch. I also realized it's not necessary to change original node state to {{REMOVED_TOKEN}} because it's already removed from gossip when the replacement node changes its state to {{NORMAL}}, so I removed that step.\n\nAlso added new dtests to test replace in a mixed-version environment to verify backward compatibility and to check for the warning message when replacing a node with the same address.","created":"2016-06-21T00:29:50.286+0000"},{"body":"+1 patch looks good. Really like the dtests.","created":"2016-07-29T03:50:15.961+0000"},{"body":"Thanks for taking a look! Created CASSANDRA-12344 to follow-up with support for this when the replacement node has the same address as the original node.\n\nRebased patch and dtests as well as merged up to 3.0+. All patches and CI results available below:\n\n||2.2||3.0||3.9||trunk||dtest||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-8523]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.0...pauloricardomg:3.0-8523]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.9...pauloricardomg:3.9-8523]|[branch|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:trunk-8523]|[branch|https://github.com/riptano/cassandra-dtest/compare/master...pauloricardomg:8523]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-8523-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.9-8523-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-8523-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-8523-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.9-8523-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-8523-dtest/lastCompletedBuild/testReport/]|\n\nThere were some minor merge conflicts on 3.0, and a slightly larger conflict on 3.9 due to CASSANDRA-10134, so I did some refactoring in the 3.9+ version to move most of the logic to {{prepareForReplacement}}. Can you take another look [~rlow]?\n\nDtest PR created [here|https://github.com/riptano/cassandra-dtest/pull/1155].\n\nWhile this is marked an improvement and would theoretically only go to trunk, this limitation is pretty counter-intuitive and probably hurts many users in the wild, and since the changeset is relatively small and self-contained, I think it could be interpreted as a bugfix and perhaps go on 2.2+ or maybe 3.0+. WDYT [~brandon.williams] [~jkni] ?","created":"2016-07-29T20:17:27.221+0000"},{"body":"I'll review the 3.9 version. I'm very much in favour of putting this in 2.2 and 3.0 as this hurts us badly and no doubt others suffer too.","created":"2016-07-29T20:27:04.863+0000"},{"body":"+1 on the 3.9 version too.","created":"2016-08-11T01:20:57.602+0000"},{"body":"Thanks! This will be ready to commit after [PR|https://github.com/riptano/cassandra-dtest/pull/1155] is reviewed and merged. I rebased with the updated dtests and submitted a new CI run:\n\n||2.2||3.0||trunk||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-8523]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.0...pauloricardomg:3.0-8523]|[branch|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:trunk-8523]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-8523-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-8523-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-8523-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-8523-dtest/lastCompletedBuild/testReport/]|","created":"2016-08-18T22:45:38.333+0000"},{"body":"Are you waiting for me to review the dtest PR?","created":"2016-08-18T23:07:28.643+0000"},{"body":"No, someone from the cassandra-dtest is reviewing, but since this will break existing tests it can only be commited after that is merged.","created":"2016-08-18T23:09:25.483+0000"},{"body":"Any updates about the Dtests","created":"2016-08-30T17:23:28.161+0000"},{"body":"I don't think the tests are waiting on anything at this point, but [~pauloricardomg] can confirm","created":"2016-08-30T17:26:21.745+0000"},{"body":"Waiting final +1 and then will mark as ready to commit.","created":"2016-08-30T17:28:54.613+0000"},{"body":"waiting final +1 from whom?","created":"2016-08-30T18:27:56.882+0000"},{"body":"from [PR#1155|https://github.com/riptano/cassandra-dtest/pull/1155], it's done. :)","created":"2016-08-30T18:37:12.457+0000"},{"body":"Committed as [b39d984f7bd682c7638415d65dcc4ac9bcb74e5f|https://github.com/apache/cassandra/commit/b39d984f7bd682c7638415d65dcc4ac9bcb74e5f] to 2.2, merged with 3.0 and trunk. Thanks.","created":"2016-08-31T19:31:28.991+0000"}],"conversations":[{"body":"In our operations, we make heavy use of replace_address (or replace_address_first_boot) in order to replace broken nodes. We now realize that writes are not sent to the replacement nodes while they are in hibernate state and streaming in data. This runs counter to what our expectations were, especially since we know that writes ARE sent to nodes when they are bootstrapped into the ring.\n\nIt seems like cassandra should arrange to send writes to a node that is in the process of replacing another node, just like it does for a nodes that are bootstraping. I hesitate to phrase this as \"we should send writes to a node in hibernate\" because the concept of hibernate may be useful in other contexts, as per CASSANDRA-8336. Maybe a new state is needed here?\n\nAmong other things, the fact that we don't get writes during this period makes subsequent repairs more expensive, proportional to the number of writes that we miss (and depending on the amount of data that needs to be streamed during replacement and the time it may take to rebuild secondary indexes, we could miss many many hours worth of writes). It also leaves us more exposed to consistency violations.\n","from":"reporter","subject":"Writes should be sent to a replacement node which has a new IP while it is streaming in data"},{"body":"I've checked with [~brandon.williams], and this is intended behavior, so I'm marking this as an Improvement, not a bug. Do note that the replacement node will receive all writes while it was streaming that were captured as hints, so if the stream takes less than the hint window, you should not see too much discrepancy.","from":"developer"},{"body":"I completely agree this is an improvement, but it's going to be pretty tricky, especially since we can't use the FD to determine if the node has died, at least not in its current form since that would mark the node as UP.","from":"developer"},{"body":"Marking related to CASSANDRA-8494, but may want to resolve as duplicate.","from":"developer"},{"body":"I'm out of ideas here. Signalling that we want to send writes to the replacing node is easy enough, but knowing when to stop if it fails is tricky, since we can't use the FD and it's a singleton so we can't just grab a new instance. [~jkni], [~Stefania] do you have any ideas?","from":"developer"},{"body":"I would consider a new, non dead, state for replacing a node rather than using hibernate. This should allow us to use the Failure Detector right?\n\nAlternatively, would a new Gossip property that indicates a \"shadow\" of an existing node be easier to implement? Could this property be made to expire unless it gets renewed by the shadow node? \n","from":"developer"},{"body":"bq. I would consider a new, non dead, state for replacing a node rather than using hibernate. This should allow us to use the Failure Detector right?\n\nIt looks like this would work now, because the gossiper is the only thing that registers with the FD. In the past this wasn't always true, though.\n\nbq. Alternatively, would a new Gossip property that indicates a \"shadow\" of an existing node be easier to implement?\n\nHmm, can you can explain further?","from":"developer"},{"body":"bq. Hmm, can you can explain further?\n\nWell, it would be a bit convoluted so a new Gossip state would probably be cleaner if we can add it. But the idea is to link a host to one or more hosts that act as 'shadows', that is they receive copies of all writes or hints but are not queried for reads. ","from":"developer"},{"body":"There are two scenarios we should consider when replacing a node:\n1) The replacing node has the same IP as the previous node\n2) The replacing node has a different IP as the previous node\n\nOn CASSANDRA-9244 I have gotten pretty far in an [implementation|https://issues.apache.org/jira/browse/CASSANDRA-9244?focusedCommentId=15211202&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15211202] that adds a new non-dead gossip state {{BOOT_REPLACE}} and considers the replacing endpoint as a bootstrapping pending endpoint, solving case 2 transparently.\n\nCase 1 is trickier because when the replacing node enters gossip with a non-dead state, other nodes will think the previous node is back up and send reads to him (since he is a natural endpoint).\n\nA simple way to solve this is to special-case the read path and ignore nodes in \"NON-NORMAL\" state when sending reads to natural endpoints. While this will probably solve the problem, there are quite a few different paths we need to hack to make sure this is enforced correctly (paxos, read, hints, etc), so I'm not totally comfortable with that.\n\nA more transparent but a bit costlier approach to solve case 1 would be to change the {{TokenMetadata}} to keep nodes as {{(InetAddress, UUID)}} pairs, and create a new interface to the {{FailureDetector}} keyed by {{UUID}}. This way we could keep {{(IP=127.0.0.1,UUID=1)}} in {{TokenMetadata}} as a natural endpoint, and add a replacement node {{(IP=127.0.0.1,UUID=2)}} as a pending endpoint. So, during reads, {{FD.isAlive(UUID=1)}} would return false, and natural reads would not be sent to {{(IP=127.0.0.1,UUID=1)}}, while pending writes would be sent to {{(IP=127.0.0.1,UUID=2)}} because {{FD.isAlive(UUID=2)}} would return true.\n\nI'd be happy to continue working on this, so feedback on any of the above or alternative approaches would be greatly appreciated.","from":"developer"},{"body":"I don't have a good idea for a low-hanging solution here - I think the cleanest is the combination of a new gossip state for replacing nodes + enhancement of the failure detector/token metadata. Since that's a bigger change, let's also make sure that it doesn't conflict with plans for strongly consistent membership. I'd really rather not couple gossip status to the internals of the read/write paths.\n\nI also agree that we should make sure that we kill two birds with one stone and handle this and [CASSANDRA-9244] at the same time.\n\nLet me know if I can help in design and/or review.","from":"developer"},{"body":"I agree with [~jkni] - a new gossip state and updating FD/TMD is the best way to go. I'll follow up on #9244 to get better grasp of what's going on there, as well.\n\nwrt strongly consistent membership, that effort focuses primarily on correct and linearizable changes to the cluster state machine, and not any behaviors that should be triggered due to those changes when they arrive at cluster nodes. So, I don't *think* we'd run into problems between these two efforts, but good to keep the same players involved :)","from":"developer"},{"body":"[~pauloricardomg] assigning to you, please do continue working on this.","from":"developer"},{"body":"Initial route I will take is to create a new dead state for replace, that adds the node as replacing endpoint to TokenMetadata. When a replacement endpoint is added to TokenMetadata, if there's an existing down node with the same IP, then we remove it from natural endpoints so the replacement node no longer receive reads when it becomes alive in the FD and include the replacement node as pending joining endpoint so writes are forwarded to it. After pending ranges for replace are calculated the final step is to set the node as alive in the FD and change the dead state logic to not mark \"alive\" dead state nodes as DOWN automatically (we will probably rename this nomenclature to a better name, like invisible or whatever), so the FD will work as usual and remove the replacing endpoint from TokenMetadata (and restore it as a natural endpoint) if it becomes down. The current streaming logic will probably be unaffected by this, and the node will change its state to NORMAL after stream is finished completing the replace procedure.\n\nThe only downside I can think of this approach is that we will lose hints during a failed replace, but this is not a big deal as hints are an optimization, and replace will probably take longer than max_hint_window anyway.\n\nI will start going this route, feel free to give any feedback or let me know if I'm missing something on this high level flow.","from":"developer"},{"body":"Without understanding the FD details, this sounds good. Losing hints isn't an issue, as you say.","from":"developer"},{"body":"Solving this when the replacing node has a different ip address is quite simple:\n* Add a new non-dead gossip state for replace BOOT_REPLACE\n* When receiving BOOT_REPLACE, other node adds the replacing node as bootstrapping AND replacing endpoint\n* Pending ranges are calculated, and writes are sent to the replacing node during replace\n* When replacing node changes state to NORMAL, the old node is removed and the new node becomes a natural endpoint on {{TokenMetadata}}\n\nI have a WIP patch with this approach [here|https://github.com/pauloricardomg/cassandra/commit/a4be7f7a0f298b1d52821f67ff2e927a1d1f1c5e] that is functional. The tricky part is when the replacing node has the same IP as the old node, because if this node ever joins gossip with a non-dead state, {{FailureDetector.isAlive(InetAddress)}} return true and reads are sent to the replacing node, since the original node is a natural endpoint.\n\nMy initial idea of removing the old node as a natural endpoint from {{TokenMetadata}} obviously does not work because this changes range assignments making reads go the following node in the ring, which has no data for that range. I also considered replacing the node IP in TokenMetadata with a special marker, but this will also not play along well with snitches and topologies.\n\nA clean fix to this is to enhance the node representation on {{TokenMetadata}}, which is currently a plain {{InetAddresss}}, to include the information if an endpoint is available for reads (CASSANDRA-11559), but this cannot be done before cassandra 4.0 because it will change {{TokenMetadata}} and {{AbstractReplicationStrategy}} public interfaces and other related classes, so it's a major refactoring in the codebase. One step further is to change the {{FailureDetector}} interface to query nodes by {{UUIDs}} instead of {{InetAddress}}, so a replacing endpoint will not be confused with the original node since they will have different ids.\n\nOne workaround before that would be to send a read to a normal endpoint if both {{FailureDetector.isAlive(endpoint)}} is true and the endpoint is on {{NORMAL}} state, so we avoid sending reads to a replacing endpoint with the same IP address. A big downside is that this check should also be peformed for other operations such as bootstrap, repairs, hints, batches, etc. I experimented with that route [in this commit|https://github.com/pauloricardomg/cassandra/commit/a439edcbd8d4186ea2e3301681b7cc0b5012a265] by creating a method {{StorageService.isAvailable(InetAddress endpoint, boolean isPending) = \\{FailureDetector.instance.isAlive(endpoint) && (isPendingEndpoint || Gossiper.instance.isNormalStatus(endpoint))\\}}} that should be used instead of {{FailureDetector.isAlive(endpoint)}}, but I'm not very comfortable with this approach for the following reasons:\n* It's pretty fragile, because that check must be used instead of {{FailureDetector.isAlive()}} in every place throughout the code\n* Other developers/drivers/tools wil also need to use the same approach or a similar logic\n* We will need to modify the write path to segregate writes between natural and pending endpoints (see an example on [doPaxosCommit|https://github.com/pauloricardomg/cassandra/commit/a439edcbd8d4186ea2e3301681b7cc0b5012a265#diff-71f06c193f5b5e270cf8ac695164f43aR498])\n\nSo unless someone comes up with a magic idea, I think we're left with two options to fix this before CASSANDRA-11559:\n1. Go ahead with {{StorageService.isAvailable}} approach even with its downsides and fragility\n2. Fix this only for replacing endpoints with a different IP address, falling back to the previous behavior when the replacing node has the same IP as the old node, which should already fix this in most cloud deployments, where replacement nodes usually have different IPs\n\nI'm leaning more towards 2 since it's much simpler/safer and work on a proper general solution for 4.0 after CASSANDRA-11559. WDYT?","from":"developer"},{"body":"I had a closer look at CASSANDRA-11559 and it should be possible to support that on trunk/3.x series without breaking compatibility, thus making this viable before 4.0 in a robust way. I will provide an initial patch for that soon.\n\nWe now need to decide whether we want to have a partial/fragile solution to this ticket before the ideal fix on 3.x leveraging #11559.","from":"developer"},{"body":"Due to the limitations of forwarding writes to replacement nodes with the same IP, I propose initially adding this support only to replacement nodes with a different IP, since it's much simpler and we can do it in a backward-compatible way so it can probably go on 2.2+.\n\nAfter CASSANDRA-11559, we can extend this support to nodes with the same IP quite easily by setting an inactive flag on nodes being replaced and ignore these nodes on read.\n\nThe central idea is:\n{quote}\n* Add a new non-dead gossip state for replace BOOT_REPLACE\n* When receiving BOOT_REPLACE, other node adds the replacing node as bootstrapping endpoint\n* Pending ranges are calculated, and writes are sent to the replacing node during replace\n* When replacing node changes state to NORMAL, the old node is removed and the new node becomes a natural endpoint on TokenMetadata\n{quote}\n\nSince it's no longer necessary to forward hints to the replacement node when {{replace_address != broadcast_address}}, the replacement node does not need to inherit the same ID of the original node.\n\nThe replacing process remains unchanged when the replacement node has the same IP as the original node. If that's the case, I added a warn message so users know they need to run repair if the node is down for longer than {{max_hint_window_in_ms}}:\n{noformat}\nWrites will not be redirected to this node while it is performing replace because it has the same address as the node to be replaced ({}). \nIf that node has been down for longer than max_hint_window_in_ms, repair must be run after the replacement process in order to make this node consistent.\n{noformat}\n\nI adapted current dtests to test replace_address for both the old and the new path, and when {{replace_address != broadcast_address}} make sure writes are being redirected to the replacement node.\n\nInitial patch and tests below (will provide 2.2+ patches after initial review):\n||2.2||dtest||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-8523]|[branch|https://github.com/riptano/cassandra-dtest/compare/master...pauloricardomg:8523]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-dtest/lastCompletedBuild/testReport/]|","from":"developer"},{"body":"Thanks! I'm happy to review.","from":"developer"},{"body":"bq. Since it's no longer necessary to forward hints to the replacement node when replace_address != broadcast_address, the replacement node does not need to inherit the same ID of the original node.\n\nI don't think that's necessarily true, if a node dies at point A, and then the replacement is issued at point N, the hints between points B and M are still missed if the host id is not the same.","from":"developer"},{"body":"It's not those writes that matter - the replacement will get those writes from the other nodes during streaming. The hints that you might care about are writes dropped during the replacement on the replacing node. But those should be extremely rare, and a price well worth paying for getting the vast majority of the writes vs none today.","from":"developer"},{"body":"Missed that comment above - [~rlow] to review. Thanks!","from":"developer"},{"body":"bq. The hints that you might care about are writes dropped during the replacement on the replacing node. \nEven these should not be a problem, because they are also forwarded to the replacement node, and if they fail they are hinted to replacement node ID.\n\nI rebased and resubmitted dtests and fixed minor typo on the original patch. I also realized it's not necessary to change original node state to {{REMOVED_TOKEN}} because it's already removed from gossip when the replacement node changes its state to {{NORMAL}}, so I removed that step.\n\nAlso added new dtests to test replace in a mixed-version environment to verify backward compatibility and to check for the warning message when replacing a node with the same address.","from":"developer"},{"body":"+1 patch looks good. Really like the dtests.","from":"developer"},{"body":"Thanks for taking a look! Created CASSANDRA-12344 to follow-up with support for this when the replacement node has the same address as the original node.\n\nRebased patch and dtests as well as merged up to 3.0+. All patches and CI results available below:\n\n||2.2||3.0||3.9||trunk||dtest||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-8523]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.0...pauloricardomg:3.0-8523]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.9...pauloricardomg:3.9-8523]|[branch|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:trunk-8523]|[branch|https://github.com/riptano/cassandra-dtest/compare/master...pauloricardomg:8523]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-8523-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.9-8523-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-8523-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-8523-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.9-8523-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-8523-dtest/lastCompletedBuild/testReport/]|\n\nThere were some minor merge conflicts on 3.0, and a slightly larger conflict on 3.9 due to CASSANDRA-10134, so I did some refactoring in the 3.9+ version to move most of the logic to {{prepareForReplacement}}. Can you take another look [~rlow]?\n\nDtest PR created [here|https://github.com/riptano/cassandra-dtest/pull/1155].\n\nWhile this is marked an improvement and would theoretically only go to trunk, this limitation is pretty counter-intuitive and probably hurts many users in the wild, and since the changeset is relatively small and self-contained, I think it could be interpreted as a bugfix and perhaps go on 2.2+ or maybe 3.0+. WDYT [~brandon.williams] [~jkni] ?","from":"developer"},{"body":"I'll review the 3.9 version. I'm very much in favour of putting this in 2.2 and 3.0 as this hurts us badly and no doubt others suffer too.","from":"developer"},{"body":"+1 on the 3.9 version too.","from":"developer"},{"body":"Thanks! This will be ready to commit after [PR|https://github.com/riptano/cassandra-dtest/pull/1155] is reviewed and merged. I rebased with the updated dtests and submitted a new CI run:\n\n||2.2||3.0||trunk||\n|[branch|https://github.com/apache/cassandra/compare/cassandra-2.2...pauloricardomg:2.2-8523]|[branch|https://github.com/apache/cassandra/compare/cassandra-3.0...pauloricardomg:3.0-8523]|[branch|https://github.com/apache/cassandra/compare/trunk...pauloricardomg:trunk-8523]|\n|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-8523-testall/lastCompletedBuild/testReport/]|[testall|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-8523-testall/lastCompletedBuild/testReport/]|\n|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-2.2-8523-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-3.0-8523-dtest/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/paulomotta/job/pauloricardomg-trunk-8523-dtest/lastCompletedBuild/testReport/]|","from":"developer"},{"body":"Are you waiting for me to review the dtest PR?","from":"developer"},{"body":"No, someone from the cassandra-dtest is reviewing, but since this will break existing tests it can only be commited after that is merged.","from":"developer"},{"body":"Any updates about the Dtests","from":"developer"},{"body":"I don't think the tests are waiting on anything at this point, but [~pauloricardomg] can confirm","from":"developer"},{"body":"Waiting final +1 and then will mark as ready to commit.","from":"developer"},{"body":"waiting final +1 from whom?","from":"developer"},{"body":"from [PR#1155|https://github.com/riptano/cassandra-dtest/pull/1155], it's done. :)","from":"developer"},{"body":"Committed as [b39d984f7bd682c7638415d65dcc4ac9bcb74e5f|https://github.com/apache/cassandra/commit/b39d984f7bd682c7638415d65dcc4ac9bcb74e5f] to 2.2, merged with 3.0 and trunk. Thanks.","from":"developer"}],"created":"2014-12-19T19:59:53.000+0000","description":"In our operations, we make heavy use of replace_address (or replace_address_first_boot) in order to replace broken nodes. We now realize that writes are not sent to the replacement nodes while they are in hibernate state and streaming in data. This runs counter to what our expectations were, especially since we know that writes ARE sent to nodes when they are bootstrapped into the ring.\n\nIt seems like cassandra should arrange to send writes to a node that is in the process of replacing another node, just like it does for a nodes that are bootstraping. I hesitate to phrase this as \"we should send writes to a node in hibernate\" because the concept of hibernate may be useful in other contexts, as per CASSANDRA-8336. Maybe a new state is needed here?\n\nAmong other things, the fact that we don't get writes during this period makes subsequent repairs more expensive, proportional to the number of writes that we miss (and depending on the amount of data that needs to be streamed during replacement and the time it may take to rebuild secondary indexes, we could miss many many hours worth of writes). It also leaves us more exposed to consistency violations.\n","issue_id":"12762978","key":"CASSANDRA-8523","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-08-31T19:32:06.000+0000","role":"fixed_distractor","summary":"Writes should be sent to a replacement node which has a new IP while it is streaming in data"} {"case_id":"12764043","cluster":"DISTRACTOR-CASSANDRA-8544","comments":[{"body":"Some questions:\n\nbq. It happens sometimes after restarts caused by undeletable files under Windows.\nHow frequently is \"sometimes\"?\nWhat files are undeletable?\nHow are you resolving / working around this? Does it work after another attempt to restart?\nDo you have steps to reproduce this?\n\nAttaching a patch that should print out some more details on the NPE if you can reproduce it consistently.","created":"2014-12-30T17:32:20.715+0000"},{"body":"I've seen it two times locally, one time it was reported by the customer of JetBrains Upsource.\n\nIt always with compactions_in_progress files, see exception below:\n{quote}\n14:16:38.232 [MemtableFlushWriter:4] ERROR o.a.c.service.CassandraDaemon - Exception in thread Thread[MemtableFlushWriter:4,5,main]\norg.apache.cassandra.io.FSWriteError: java.nio.file.FileSystemException: C:\\Tools\\Upsource\\data\\cassandra\\data\\system\\compactions_in_progress-55080a\nb05d9c388690a4acb25fe1f77b\\system-compactions_in_progress-tmp-ka-13253-Statistics.db: Der Prozess kann nicht auf die Datei zugreifen, da sie von ein\nem anderen Prozess verwendet wird.\n\n at org.apache.cassandra.io.util.FileUtils.deleteWithConfirm(FileUtils.java:135) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.io.util.FileUtils.deleteWithConfirm(FileUtils.java:121) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.io.sstable.SSTable.delete(SSTable.java:113) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.io.sstable.SSTableWriter.abort(SSTableWriter.java:355) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.io.sstable.SSTableWriter.abort(SSTableWriter.java:333) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.db.Memtable$FlushRunnable.writeSortedContents(Memtable.java:381) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.db.Memtable$FlushRunnable.runWith(Memtable.java:312) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) ~[cassandra-all-2.1.1.jar:2.1.1]\n at com.google.common.util.concurrent.MoreExecutors$SameThreadExecutorService.execute(MoreExecutors.java:297) ~[guava-16.0.jar:na]\n at org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1053) ~[cassandra-all-2.1.1.jar:2.1.1]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_71]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) ~[na:1.7.0_71]\n at java.lang.Thread.run(Thread.java:745) ~[na:1.7.0_71]\n{quote}\n\nThis particular exception seems to be fixed by turning off AV and Windows Search, but it's no good to fail on the next start.\n\n> How are you resolving / working around this? Does it work after another attempt to restart?\nNo workarounds attempted. It doesn't work after another restart.\n\n> Attaching a patch that should print out some more details on the NPE if you can reproduce it consistently.\nUnfortunately I can't reproduce it consistently, so no luck here.\n","created":"2015-01-05T17:35:00.350+0000"},{"body":"If turning off AV and Windows Search consistently fixes this I consider this an external/environment problem and not a problem with our code-base w/regards to stopping. Pre 3.0 (and w/memory-mapped I/O when we go back down that route), file sharing violations like this from outside programs are, by default, a problem we're going to stop the database for rather than allowing failed deletions to potentially fill a drive on a node. If you'd prefer to allow the system to continue you can reference the following: [one|http://www.datastax.com/documentation/cassandra/2.0/cassandra/configuration/configCassandra_yaml_r.html] - [two|http://www.datastax.com/dev/blog/handling-disk-failures-in-cassandra-1-2].\n\nAt the very least, you need to exclude your cassandra data and logs folder from AV scanning and also [remove it from indexing|http://windows.microsoft.com/en-us/windows/improve-windows-searches-using-index-faq#1TC=windows-7].\n\nAttaching a v1 that wraps the NPE and instead throws an FSReadError and points to log file for more information, since our original NPE error isn't terribly useful.","created":"2015-01-07T20:36:15.156+0000"},{"body":"+1","created":"2015-03-06T05:46:03.506+0000"},{"body":"Committed.","created":"2015-03-09T17:36:35.756+0000"}],"conversations":[{"body":"It happens sometimes after restarts caused by undeletable files under Windows.\n\n{quote}\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.db.ColumnFamilyStore.removeUnfinishedCompactionLeftovers(ColumnFamilyStore.java:579)\n at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:232)\n at org.apache.cassandra.service.CassandraDaemon.init(CassandraDaemon.java:377)\n at com.jetbrains.cassandra.service.CassandraServiceMain.start(CassandraServiceMain.java:81)\n ... 6 more\n{quote}","from":"reporter","subject":"Cassandra could not start with NPE in ColumnFamilyStore.removeUnfinishedCompactionLeftovers"},{"body":"Some questions:\n\nbq. It happens sometimes after restarts caused by undeletable files under Windows.\nHow frequently is \"sometimes\"?\nWhat files are undeletable?\nHow are you resolving / working around this? Does it work after another attempt to restart?\nDo you have steps to reproduce this?\n\nAttaching a patch that should print out some more details on the NPE if you can reproduce it consistently.","from":"developer"},{"body":"I've seen it two times locally, one time it was reported by the customer of JetBrains Upsource.\n\nIt always with compactions_in_progress files, see exception below:\n{quote}\n14:16:38.232 [MemtableFlushWriter:4] ERROR o.a.c.service.CassandraDaemon - Exception in thread Thread[MemtableFlushWriter:4,5,main]\norg.apache.cassandra.io.FSWriteError: java.nio.file.FileSystemException: C:\\Tools\\Upsource\\data\\cassandra\\data\\system\\compactions_in_progress-55080a\nb05d9c388690a4acb25fe1f77b\\system-compactions_in_progress-tmp-ka-13253-Statistics.db: Der Prozess kann nicht auf die Datei zugreifen, da sie von ein\nem anderen Prozess verwendet wird.\n\n at org.apache.cassandra.io.util.FileUtils.deleteWithConfirm(FileUtils.java:135) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.io.util.FileUtils.deleteWithConfirm(FileUtils.java:121) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.io.sstable.SSTable.delete(SSTable.java:113) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.io.sstable.SSTableWriter.abort(SSTableWriter.java:355) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.io.sstable.SSTableWriter.abort(SSTableWriter.java:333) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.db.Memtable$FlushRunnable.writeSortedContents(Memtable.java:381) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.db.Memtable$FlushRunnable.runWith(Memtable.java:312) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.io.util.DiskAwareRunnable.runMayThrow(DiskAwareRunnable.java:48) ~[cassandra-all-2.1.1.jar:2.1.1]\n at org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28) ~[cassandra-all-2.1.1.jar:2.1.1]\n at com.google.common.util.concurrent.MoreExecutors$SameThreadExecutorService.execute(MoreExecutors.java:297) ~[guava-16.0.jar:na]\n at org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1053) ~[cassandra-all-2.1.1.jar:2.1.1]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_71]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) ~[na:1.7.0_71]\n at java.lang.Thread.run(Thread.java:745) ~[na:1.7.0_71]\n{quote}\n\nThis particular exception seems to be fixed by turning off AV and Windows Search, but it's no good to fail on the next start.\n\n> How are you resolving / working around this? Does it work after another attempt to restart?\nNo workarounds attempted. It doesn't work after another restart.\n\n> Attaching a patch that should print out some more details on the NPE if you can reproduce it consistently.\nUnfortunately I can't reproduce it consistently, so no luck here.\n","from":"developer"},{"body":"If turning off AV and Windows Search consistently fixes this I consider this an external/environment problem and not a problem with our code-base w/regards to stopping. Pre 3.0 (and w/memory-mapped I/O when we go back down that route), file sharing violations like this from outside programs are, by default, a problem we're going to stop the database for rather than allowing failed deletions to potentially fill a drive on a node. If you'd prefer to allow the system to continue you can reference the following: [one|http://www.datastax.com/documentation/cassandra/2.0/cassandra/configuration/configCassandra_yaml_r.html] - [two|http://www.datastax.com/dev/blog/handling-disk-failures-in-cassandra-1-2].\n\nAt the very least, you need to exclude your cassandra data and logs folder from AV scanning and also [remove it from indexing|http://windows.microsoft.com/en-us/windows/improve-windows-searches-using-index-faq#1TC=windows-7].\n\nAttaching a v1 that wraps the NPE and instead throws an FSReadError and points to log file for more information, since our original NPE error isn't terribly useful.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed.","from":"developer"}],"created":"2014-12-29T16:29:27.000+0000","description":"It happens sometimes after restarts caused by undeletable files under Windows.\n\n{quote}\nCaused by: java.lang.NullPointerException\n at org.apache.cassandra.db.ColumnFamilyStore.removeUnfinishedCompactionLeftovers(ColumnFamilyStore.java:579)\n at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:232)\n at org.apache.cassandra.service.CassandraDaemon.init(CassandraDaemon.java:377)\n at com.jetbrains.cassandra.service.CassandraServiceMain.start(CassandraServiceMain.java:81)\n ... 6 more\n{quote}","issue_id":"12764043","key":"CASSANDRA-8544","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-03-09T17:36:35.000+0000","role":"fixed_distractor","summary":"Cassandra could not start with NPE in ColumnFamilyStore.removeUnfinishedCompactionLeftovers"} {"case_id":"12767307","cluster":"DISTRACTOR-CASSANDRA-8616","comments":[{"body":"[~philipthompson] can you try to reproduce this? It may be helpful to run stress for a while, then shut down the node, then run sstable2json on one of the stress tables.","created":"2015-01-13T23:31:37.978+0000"},{"body":"This indeed seems to be the case with 2.0.10, and it's even happening if you invoke sstable2json incorrectly.\n\nTo repro I ran cassandra-stress, and killed the cassandra process while stress was still running. Then I ran sstable2json (with a bad sstable filename, and with a valid one) and new commitlog files are created.\n\n{noformat}\nroot@ce986ee973c1:/cassandra# ls -ltr /var/lib/cassandra/commitlog | wc -l\n11\nroot@ce986ee973c1:/cassandra# bin/sstable2json thisfiledoesntexist.db > /dev/null\nException in thread \"main\" java.util.NoSuchElementException\n at java.util.StringTokenizer.nextToken(StringTokenizer.java:349)\n at org.apache.cassandra.io.sstable.Descriptor.fromFilename(Descriptor.java:235)\n at org.apache.cassandra.io.sstable.Descriptor.fromFilename(Descriptor.java:204)\n at org.apache.cassandra.tools.SSTableExport.main(SSTableExport.java:452)\n^C\nroot@ce986ee973c1:/cassandra# ls -ltr /var/lib/cassandra/commitlog | wc -l\n12\n{noformat}","created":"2015-01-14T00:32:26.327+0000"},{"body":"Same goes for 2.1\n\n{noformat}\nroot@e69321a30198:/cassandra# ls -ltr data/commitlog/ \ntotal 262144\n-rw-r--r-- 1 root root 33554432 Jan 14 00:37 CommitLog-4-1421195804516.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804525.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804530.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804529.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804528.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804531.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804526.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804527.log\nroot@e69321a30198:/cassandra# ls -ltr data/commitlog/ | wc -l\n9\nroot@e69321a30198:/cassandra# tools/bin/sstable2json data/data/keyspace1/standard1-85148e309b8511e4b94691de6f1c3a1d/keyspace1-standard1-ka-1-Data.db > /dev/null\nroot@e69321a30198:/cassandra# ls -ltr data/commitlog/\ntotal 262148\n-rw-r--r-- 1 root root 33554432 Jan 14 00:37 CommitLog-4-1421195804516.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804525.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804530.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804529.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804528.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804531.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804526.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804527.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:42 CommitLog-4-1421196133229.log\nroot@e69321a30198:/cassandra# ls -ltr data/commitlog/ | wc -l\n10\n{noformat}","created":"2015-01-14T00:44:15.765+0000"},{"body":"Thanks, [~rhatch].\n\n[~yukim] do you want to take this one?","created":"2015-01-14T17:13:37.404+0000"},{"body":"yup","created":"2015-01-14T17:52:40.854+0000"},{"body":"This is due to the fact that loading schema writes (updates) schema version whenever it happens.\n(I believe the behavior hasn't changed for a while, so issue here has been around way before 2.0.10.)\nAlso judging from the code, standalone scrub/upgradesstbale/sstablesplit behave the same.\n\nQuick work around is to add the way to load schema without updating schema version.","created":"2015-01-21T00:52:16.791+0000"},{"body":"+1","created":"2015-01-21T23:10:13.800+0000"},{"body":"Committed, thanks!","created":"2015-01-23T15:28:38.044+0000"},{"body":"Seems to be working fine in 2.0 latest, but latest 2.1 (c468c8b436) is still writing commitlog files when sstable2json is called.\n\nSimilar to before, this happens whether the db file argument is valid or not, and each time 1 new commitlog file is written.","created":"2015-01-23T23:19:47.841+0000"},{"body":"You are right.\nIn 2.1 and trunk, accessing schema creates new Memtable instance for schema_keyspace, and [Memtable touches CommitLog singleton|https://github.com/apache/cassandra/blob/cassandra-2.1.2/src/java/org/apache/cassandra/db/Memtable.java#L66] which [creates one commit log file when it is initialized|https://github.com/apache/cassandra/blob/cassandra-2.1.2/src/java/org/apache/cassandra/db/commitlog/CommitLog.java#L70].\n\nLooks like more work is needed to be done...","created":"2015-01-24T02:21:20.682+0000"},{"body":"I am seeing what appears to be the same issue in 1.2.16. Using sstable2json is trying to open the commitlog files for writing:\n\n{noformat}\n$ sstable2json \nException in thread \"COMMIT-LOG-ALLOCATOR\" FSWriteError in /var/lib/cassandra/commitlog/CommitLog-2-1423079861481.log ]\n\tat org.apache.cassandra.db.commitlog.CommitLogSegment.(CommitLogSegment.java:132)\n\tat org.apache.cassandra.db.commitlog.CommitLogSegment.freshSegment(CommitLogSegment.java:81)\n\tat org.apache.cassandra.db.commitlog.CommitLogAllocator.createFreshSegment(CommitLogAllocator.java:251)\n\tat org.apache.cassandra.db.commitlog.CommitLogAllocator.access$500(CommitLogAllocator.java:49)\n\tat org.apache.cassandra.db.commitlog.CommitLogAllocator$1.runMayThrow(CommitLogAllocator.java:105)\n\tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n\tat java.lang.Thread.run(Thread.java:695)\nCaused by: java.io.FileNotFoundException: /var/lib/cassandra/commitlog/CommitLog-2-1423079861481.log (Permission denied)\n\tat java.io.RandomAccessFile.open(Native Method)\n\tat java.io.RandomAccessFile.(RandomAccessFile.java:216)\n\tat org.apache.cassandra.db.commitlog.CommitLogSegment.(CommitLogSegment.java:116)\n\t... 6 more\n{noformat}","created":"2015-02-04T20:00:56.043+0000"},{"body":"Something to keep in mind for our next gen sstable2json [~slebresne]","created":"2015-07-13T04:39:51.219+0000"},{"body":"Agreed. The goal mid-term is to get something CASSANDRA-9587 so such tool really act as external tools, i.e. they don't silently load a lot of the same stuff than the server so this kind of thing cannot happen.","created":"2015-07-13T06:48:43.018+0000"},{"body":"Bumping to critical since this can prevent startup when sstable2json creates a new segment as root. ","created":"2015-10-05T18:16:10.021+0000"},{"body":"+1 to critical priority, I encountered this in production and, even worse than preventing startup, it started the node up in a mode where it was listening to clients but was unable to speak to other nodes.","created":"2016-02-09T21:24:52.870+0000"},{"body":"[~yukim] Since you did the first part of this issue, mind looking at the followup?","created":"2016-02-11T16:17:15.617+0000"},{"body":"I wrote dirty workaround patch: https://github.com/yukim/cassandra/tree/8616-2.1\n\nBasically added flag to disable CommitLogAllocator so directories/files are not created.\nWe cannot use Config.setClientMode for that because it causes another problems(basically, we cannot use ColumnFaimlyStore in client mode).\n\nAt first, I wrote offline schema loader that constructs schema from schema_* SSTables. But, offline tools just use ColumnFamilyStore everywhere, I needed to rewrite everything. So I end up above hack.\nI can save it for later, probably at the time CASSANDRA-7464 is in.\n\nIf the change looks good, I will work for other branches.","created":"2016-02-12T01:39:49.836+0000"},{"body":"We definitively want to clean how offline use internal code as it's a mess right now but for now I agree a quick and dirty (and importantly simple) fix is good enough. It would be nice to write a regression dtest for this though before committing.\n\n[~thobbs] are you good finishing review on this since you're still marked reviewer?","created":"2016-02-12T11:12:26.506+0000"},{"body":"bq. Tyler Hobbs are you good finishing review on this since you're still marked reviewer?\n\nSure, I'll review.","created":"2016-02-12T23:03:06.169+0000"},{"body":"The fix looks good to me as a temporary solution. However, I think you missed a few tools that also need the fix:\n* {{SSTableLevelResetter}}\n* {{SSTableExpiredBlockers}}\n* {{BulkLoader}}\n\nAnd in later versions, we also need to fix:\n* {{SSTableRepairedAtSetter}}\n* {{StandaloneVerifier}}\n* {{StandaloneSSTableUtil}}\n\nAs Sylvain mentions, it would be good to add regression dtests for these.","created":"2016-02-13T00:27:21.736+0000"},{"body":"Bad news. My approach didn't work.\n\nThere is one point that tries to access commit log: offline tools like sstablescrub tries to delete data from system.sstable_activity after scrubbing SSTable. This causes access to commit log and since we are trying to disable, the tool hangs at [here|https://github.com/apache/cassandra/blob/cassandra-2.1.13/src/java/org/apache/cassandra/db/commitlog/CommitLogSegmentManager.java#L271].\nPreviously I stated that many offline tools cannot run with {{Client.setClientMode(true)}}, so reading from SSTableReader tracks sstable activity even we are ofline.\n\nProbably we should consider rewriting tools so that they never use ColumnFamilyStore.","created":"2016-02-19T17:24:26.172+0000"},{"body":"bq. Probably we should consider rewriting tools so that they never use ColumnFamilyStore.\n\nThat is sounding like a better option to me, too. Without doing that, I worry that we will miss edge cases that touch the commitlog in the future.","created":"2016-02-19T18:17:30.887+0000"},{"body":"bq. That is sounding like a better option to me, too.\n\nRight, but I'll also note that some people seems to have run into this with pretty nasty consequences. So I'm all for a clean solution here, but unless \"rewriting tools so that they never use ColumnFamilyStore\" is a lot more trivial than it sounds to me, we might need a quick-n-dirty fix that is suitable for 2.1 here. And what I mean by that is that we might want to add some \"offline\" mode (different from \"client mode\" since you say we can't use that) that skips what shouldn't be done by offline tools (like deleting data from {{system.sstable_activity}}, which sounds to me like something we probably can just skip when doing an offline scrub).","created":"2016-02-22T09:55:24.179+0000"},{"body":"Progress update: There are several places other than {{system.sstable_activity}} that offline tools access(vary by version). Trying to patch everything, otherwise tools can hang (dtest is catching these).\n\n* In cassandra-2.2, offline tools (sstablesplit and posssibly other few) can update {{system.compaction_in_progress}} and {{system.compaction_history}}.\n* In cassandra-3.0+, sstablescrub against secondary index sstables can update {{system.IndexInfo}}.","created":"2016-03-30T22:55:10.068+0000"},{"body":"[~yukim] CASSANDRA-9054 is now committed. Can you check whether this is still an issue?","created":"2016-08-18T01:46:17.038+0000"},{"body":"[~snazy] the problem comes from opening schema in the tools. They need to access schema by opening {{system_schema}} keyspace that creates commit log. I tried to patch commit log part, but various \"online\" features described above came up after that. CASSANDRA-9587 can be useful here.","created":"2016-08-19T23:44:00.767+0000"},{"body":"I took another approach in the new patches.\nSince the source of accessing commit log is {{Memtable}}, I changed {{Tracker}} to have {{Memtable}} optional when only on online. (This change may also be useful for future offline tools change.)\n\nThere are several places that try to update system tables (sstable_activity, secondary index, compaction) so I manually had to disabled them by checking {{DatabaseDescriptor.isDaemonInitialized}} (in trunk, added similar method to 2.2 and 3.0), or offline tool can hang.\n\n||branch||testall||dtest||\n|[8616-2.2|https://github.com/yukim/cassandra/tree/8616-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-2.2-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-2.2-dtest/lastCompletedBuild/testReport/]|\n|[8616-3.0|https://github.com/yukim/cassandra/tree/8616-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.0-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.0-dtest/lastCompletedBuild/testReport/]|\n|[8616-trunk|https://github.com/yukim/cassandra/tree/8616-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-trunk-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-trunk-dtest/lastCompletedBuild/testReport/]|\n\n(test is still running for some)","created":"2016-09-12T21:44:20.677+0000"},{"body":"Unit test failures in trunk revealed that I need to patch where we query Memtable and SSTables since current code assumes there is at least one iterator there. Will update the patch.","created":"2016-09-13T17:19:49.951+0000"},{"body":"Patch updated and tests are running.","created":"2016-09-14T03:04:30.930+0000"},{"body":"The patches seem fine to me, but there are some problems in the tests. It seems like the 2.2 testall failures around SSTableRewriter could be related to this, since they don't show up in normal 2.2 test failures. The 3.0 and 3.3 dtests are erroring out when trying to copy test results at the end, but if I remember correctly, this can be caused by some tests not terminating, so we should investigate that. The trunk RemoveTest failure in testall could also be related to this patch.","created":"2016-09-15T16:52:37.624+0000"},{"body":"There were indeed a problem that causes dtest failure in previous patches.\nUpdated and rebased all the branches:\n\n||branch||testall||dtest||\n|[8616-2.2|https://github.com/yukim/cassandra/tree/8616-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-2.2-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-2.2-dtest/lastCompletedBuild/testReport/]|\n|[8616-3.0|https://github.com/yukim/cassandra/tree/8616-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.0-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.0-dtest/lastCompletedBuild/testReport/]|\n|[8616-3.X|https://github.com/yukim/cassandra/tree/8616-3.X]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.X-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.X-dtest/lastCompletedBuild/testReport/]|\n|[8616-trunk|https://github.com/yukim/cassandra/tree/8616-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-trunk-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-trunk-dtest/lastCompletedBuild/testReport/]|\n","created":"2016-10-10T13:55:52.598+0000"},{"body":"I totally forgot about this, sorry about that. The test results do look much better now. If you can rebase again and do one final test run, I'm +1 on committing this if the test results still look good.","created":"2016-11-29T18:05:06.171+0000"},{"body":"ping [~yukim]","created":"2016-12-07T18:38:00.182+0000"},{"body":"Rebase pushed, running tests now.","created":"2016-12-07T19:03:49.443+0000"},{"body":"I've restarted the 3.0 tests, because they didn't run for some reason. It looks like the only potential problems are in the trunk dtests, where {{offline_tools_test.TestOfflineTools.sstableupgrade_test}} and {{upgrade_internal_auth_test.TestAuthUpgrade.test_upgrade_legacy_table}} are failing.","created":"2016-12-08T23:28:10.960+0000"},{"body":"Right, I already pushed the fix and running tests again. Will update once they are done.","created":"2016-12-08T23:55:39.822+0000"},{"body":"Tests completed now.","created":"2016-12-09T03:57:52.339+0000"},{"body":"+1","created":"2016-12-13T17:08:32.653+0000"},{"body":"Thanks, committed as {{66f1aaf88d3cde5c52b13d71d3326da5eda16fb1}}.","created":"2016-12-14T00:47:00.952+0000"}],"conversations":[{"body":"There was a report of sstable2json causing commitlog segments to be written out when run. I haven't attempted to reproduce this yet, so that's all I know for now. Since sstable2json loads the conf and schema, I'm thinking that it may inadvertently be triggering the commitlog code.\n\nsstablescrub, sstableverify, and other sstable tools have the same issue.","from":"reporter","subject":"sstable tools may result in commit log segments be written"},{"body":"[~philipthompson] can you try to reproduce this? It may be helpful to run stress for a while, then shut down the node, then run sstable2json on one of the stress tables.","from":"developer"},{"body":"This indeed seems to be the case with 2.0.10, and it's even happening if you invoke sstable2json incorrectly.\n\nTo repro I ran cassandra-stress, and killed the cassandra process while stress was still running. Then I ran sstable2json (with a bad sstable filename, and with a valid one) and new commitlog files are created.\n\n{noformat}\nroot@ce986ee973c1:/cassandra# ls -ltr /var/lib/cassandra/commitlog | wc -l\n11\nroot@ce986ee973c1:/cassandra# bin/sstable2json thisfiledoesntexist.db > /dev/null\nException in thread \"main\" java.util.NoSuchElementException\n at java.util.StringTokenizer.nextToken(StringTokenizer.java:349)\n at org.apache.cassandra.io.sstable.Descriptor.fromFilename(Descriptor.java:235)\n at org.apache.cassandra.io.sstable.Descriptor.fromFilename(Descriptor.java:204)\n at org.apache.cassandra.tools.SSTableExport.main(SSTableExport.java:452)\n^C\nroot@ce986ee973c1:/cassandra# ls -ltr /var/lib/cassandra/commitlog | wc -l\n12\n{noformat}","from":"developer"},{"body":"Same goes for 2.1\n\n{noformat}\nroot@e69321a30198:/cassandra# ls -ltr data/commitlog/ \ntotal 262144\n-rw-r--r-- 1 root root 33554432 Jan 14 00:37 CommitLog-4-1421195804516.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804525.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804530.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804529.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804528.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804531.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804526.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804527.log\nroot@e69321a30198:/cassandra# ls -ltr data/commitlog/ | wc -l\n9\nroot@e69321a30198:/cassandra# tools/bin/sstable2json data/data/keyspace1/standard1-85148e309b8511e4b94691de6f1c3a1d/keyspace1-standard1-ka-1-Data.db > /dev/null\nroot@e69321a30198:/cassandra# ls -ltr data/commitlog/\ntotal 262148\n-rw-r--r-- 1 root root 33554432 Jan 14 00:37 CommitLog-4-1421195804516.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804525.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804530.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804529.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804528.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804531.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804526.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:38 CommitLog-4-1421195804527.log\n-rw-r--r-- 1 root root 33554432 Jan 14 00:42 CommitLog-4-1421196133229.log\nroot@e69321a30198:/cassandra# ls -ltr data/commitlog/ | wc -l\n10\n{noformat}","from":"developer"},{"body":"Thanks, [~rhatch].\n\n[~yukim] do you want to take this one?","from":"developer"},{"body":"yup","from":"developer"},{"body":"This is due to the fact that loading schema writes (updates) schema version whenever it happens.\n(I believe the behavior hasn't changed for a while, so issue here has been around way before 2.0.10.)\nAlso judging from the code, standalone scrub/upgradesstbale/sstablesplit behave the same.\n\nQuick work around is to add the way to load schema without updating schema version.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed, thanks!","from":"developer"},{"body":"Seems to be working fine in 2.0 latest, but latest 2.1 (c468c8b436) is still writing commitlog files when sstable2json is called.\n\nSimilar to before, this happens whether the db file argument is valid or not, and each time 1 new commitlog file is written.","from":"developer"},{"body":"You are right.\nIn 2.1 and trunk, accessing schema creates new Memtable instance for schema_keyspace, and [Memtable touches CommitLog singleton|https://github.com/apache/cassandra/blob/cassandra-2.1.2/src/java/org/apache/cassandra/db/Memtable.java#L66] which [creates one commit log file when it is initialized|https://github.com/apache/cassandra/blob/cassandra-2.1.2/src/java/org/apache/cassandra/db/commitlog/CommitLog.java#L70].\n\nLooks like more work is needed to be done...","from":"developer"},{"body":"I am seeing what appears to be the same issue in 1.2.16. Using sstable2json is trying to open the commitlog files for writing:\n\n{noformat}\n$ sstable2json \nException in thread \"COMMIT-LOG-ALLOCATOR\" FSWriteError in /var/lib/cassandra/commitlog/CommitLog-2-1423079861481.log ]\n\tat org.apache.cassandra.db.commitlog.CommitLogSegment.(CommitLogSegment.java:132)\n\tat org.apache.cassandra.db.commitlog.CommitLogSegment.freshSegment(CommitLogSegment.java:81)\n\tat org.apache.cassandra.db.commitlog.CommitLogAllocator.createFreshSegment(CommitLogAllocator.java:251)\n\tat org.apache.cassandra.db.commitlog.CommitLogAllocator.access$500(CommitLogAllocator.java:49)\n\tat org.apache.cassandra.db.commitlog.CommitLogAllocator$1.runMayThrow(CommitLogAllocator.java:105)\n\tat org.apache.cassandra.utils.WrappedRunnable.run(WrappedRunnable.java:28)\n\tat java.lang.Thread.run(Thread.java:695)\nCaused by: java.io.FileNotFoundException: /var/lib/cassandra/commitlog/CommitLog-2-1423079861481.log (Permission denied)\n\tat java.io.RandomAccessFile.open(Native Method)\n\tat java.io.RandomAccessFile.(RandomAccessFile.java:216)\n\tat org.apache.cassandra.db.commitlog.CommitLogSegment.(CommitLogSegment.java:116)\n\t... 6 more\n{noformat}","from":"developer"},{"body":"Something to keep in mind for our next gen sstable2json [~slebresne]","from":"developer"},{"body":"Agreed. The goal mid-term is to get something CASSANDRA-9587 so such tool really act as external tools, i.e. they don't silently load a lot of the same stuff than the server so this kind of thing cannot happen.","from":"developer"},{"body":"Bumping to critical since this can prevent startup when sstable2json creates a new segment as root. ","from":"developer"},{"body":"+1 to critical priority, I encountered this in production and, even worse than preventing startup, it started the node up in a mode where it was listening to clients but was unable to speak to other nodes.","from":"developer"},{"body":"[~yukim] Since you did the first part of this issue, mind looking at the followup?","from":"developer"},{"body":"I wrote dirty workaround patch: https://github.com/yukim/cassandra/tree/8616-2.1\n\nBasically added flag to disable CommitLogAllocator so directories/files are not created.\nWe cannot use Config.setClientMode for that because it causes another problems(basically, we cannot use ColumnFaimlyStore in client mode).\n\nAt first, I wrote offline schema loader that constructs schema from schema_* SSTables. But, offline tools just use ColumnFamilyStore everywhere, I needed to rewrite everything. So I end up above hack.\nI can save it for later, probably at the time CASSANDRA-7464 is in.\n\nIf the change looks good, I will work for other branches.","from":"developer"},{"body":"We definitively want to clean how offline use internal code as it's a mess right now but for now I agree a quick and dirty (and importantly simple) fix is good enough. It would be nice to write a regression dtest for this though before committing.\n\n[~thobbs] are you good finishing review on this since you're still marked reviewer?","from":"developer"},{"body":"bq. Tyler Hobbs are you good finishing review on this since you're still marked reviewer?\n\nSure, I'll review.","from":"developer"},{"body":"The fix looks good to me as a temporary solution. However, I think you missed a few tools that also need the fix:\n* {{SSTableLevelResetter}}\n* {{SSTableExpiredBlockers}}\n* {{BulkLoader}}\n\nAnd in later versions, we also need to fix:\n* {{SSTableRepairedAtSetter}}\n* {{StandaloneVerifier}}\n* {{StandaloneSSTableUtil}}\n\nAs Sylvain mentions, it would be good to add regression dtests for these.","from":"developer"},{"body":"Bad news. My approach didn't work.\n\nThere is one point that tries to access commit log: offline tools like sstablescrub tries to delete data from system.sstable_activity after scrubbing SSTable. This causes access to commit log and since we are trying to disable, the tool hangs at [here|https://github.com/apache/cassandra/blob/cassandra-2.1.13/src/java/org/apache/cassandra/db/commitlog/CommitLogSegmentManager.java#L271].\nPreviously I stated that many offline tools cannot run with {{Client.setClientMode(true)}}, so reading from SSTableReader tracks sstable activity even we are ofline.\n\nProbably we should consider rewriting tools so that they never use ColumnFamilyStore.","from":"developer"},{"body":"bq. Probably we should consider rewriting tools so that they never use ColumnFamilyStore.\n\nThat is sounding like a better option to me, too. Without doing that, I worry that we will miss edge cases that touch the commitlog in the future.","from":"developer"},{"body":"bq. That is sounding like a better option to me, too.\n\nRight, but I'll also note that some people seems to have run into this with pretty nasty consequences. So I'm all for a clean solution here, but unless \"rewriting tools so that they never use ColumnFamilyStore\" is a lot more trivial than it sounds to me, we might need a quick-n-dirty fix that is suitable for 2.1 here. And what I mean by that is that we might want to add some \"offline\" mode (different from \"client mode\" since you say we can't use that) that skips what shouldn't be done by offline tools (like deleting data from {{system.sstable_activity}}, which sounds to me like something we probably can just skip when doing an offline scrub).","from":"developer"},{"body":"Progress update: There are several places other than {{system.sstable_activity}} that offline tools access(vary by version). Trying to patch everything, otherwise tools can hang (dtest is catching these).\n\n* In cassandra-2.2, offline tools (sstablesplit and posssibly other few) can update {{system.compaction_in_progress}} and {{system.compaction_history}}.\n* In cassandra-3.0+, sstablescrub against secondary index sstables can update {{system.IndexInfo}}.","from":"developer"},{"body":"[~yukim] CASSANDRA-9054 is now committed. Can you check whether this is still an issue?","from":"developer"},{"body":"[~snazy] the problem comes from opening schema in the tools. They need to access schema by opening {{system_schema}} keyspace that creates commit log. I tried to patch commit log part, but various \"online\" features described above came up after that. CASSANDRA-9587 can be useful here.","from":"developer"},{"body":"I took another approach in the new patches.\nSince the source of accessing commit log is {{Memtable}}, I changed {{Tracker}} to have {{Memtable}} optional when only on online. (This change may also be useful for future offline tools change.)\n\nThere are several places that try to update system tables (sstable_activity, secondary index, compaction) so I manually had to disabled them by checking {{DatabaseDescriptor.isDaemonInitialized}} (in trunk, added similar method to 2.2 and 3.0), or offline tool can hang.\n\n||branch||testall||dtest||\n|[8616-2.2|https://github.com/yukim/cassandra/tree/8616-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-2.2-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-2.2-dtest/lastCompletedBuild/testReport/]|\n|[8616-3.0|https://github.com/yukim/cassandra/tree/8616-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.0-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.0-dtest/lastCompletedBuild/testReport/]|\n|[8616-trunk|https://github.com/yukim/cassandra/tree/8616-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-trunk-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-trunk-dtest/lastCompletedBuild/testReport/]|\n\n(test is still running for some)","from":"developer"},{"body":"Unit test failures in trunk revealed that I need to patch where we query Memtable and SSTables since current code assumes there is at least one iterator there. Will update the patch.","from":"developer"},{"body":"Patch updated and tests are running.","from":"developer"},{"body":"The patches seem fine to me, but there are some problems in the tests. It seems like the 2.2 testall failures around SSTableRewriter could be related to this, since they don't show up in normal 2.2 test failures. The 3.0 and 3.3 dtests are erroring out when trying to copy test results at the end, but if I remember correctly, this can be caused by some tests not terminating, so we should investigate that. The trunk RemoveTest failure in testall could also be related to this patch.","from":"developer"},{"body":"There were indeed a problem that causes dtest failure in previous patches.\nUpdated and rebased all the branches:\n\n||branch||testall||dtest||\n|[8616-2.2|https://github.com/yukim/cassandra/tree/8616-2.2]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-2.2-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-2.2-dtest/lastCompletedBuild/testReport/]|\n|[8616-3.0|https://github.com/yukim/cassandra/tree/8616-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.0-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.0-dtest/lastCompletedBuild/testReport/]|\n|[8616-3.X|https://github.com/yukim/cassandra/tree/8616-3.X]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.X-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-3.X-dtest/lastCompletedBuild/testReport/]|\n|[8616-trunk|https://github.com/yukim/cassandra/tree/8616-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-trunk-testall/lastCompletedBuild/testReport/]|[dtest|http://cassci.datastax.com/view/Dev/view/yukim/job/yukim-8616-trunk-dtest/lastCompletedBuild/testReport/]|\n","from":"developer"},{"body":"I totally forgot about this, sorry about that. The test results do look much better now. If you can rebase again and do one final test run, I'm +1 on committing this if the test results still look good.","from":"developer"},{"body":"ping [~yukim]","from":"developer"},{"body":"Rebase pushed, running tests now.","from":"developer"},{"body":"I've restarted the 3.0 tests, because they didn't run for some reason. It looks like the only potential problems are in the trunk dtests, where {{offline_tools_test.TestOfflineTools.sstableupgrade_test}} and {{upgrade_internal_auth_test.TestAuthUpgrade.test_upgrade_legacy_table}} are failing.","from":"developer"},{"body":"Right, I already pushed the fix and running tests again. Will update once they are done.","from":"developer"},{"body":"Tests completed now.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Thanks, committed as {{66f1aaf88d3cde5c52b13d71d3326da5eda16fb1}}.","from":"developer"}],"created":"2015-01-13T23:27:35.000+0000","description":"There was a report of sstable2json causing commitlog segments to be written out when run. I haven't attempted to reproduce this yet, so that's all I know for now. Since sstable2json loads the conf and schema, I'm thinking that it may inadvertently be triggering the commitlog code.\n\nsstablescrub, sstableverify, and other sstable tools have the same issue.","issue_id":"12767307","key":"CASSANDRA-8616","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-12-14T00:46:59.000+0000","role":"fixed_distractor","summary":"sstable tools may result in commit log segments be written"} {"case_id":"12769480","cluster":"DISTRACTOR-CASSANDRA-8670","comments":[{"body":"What's the path forward on this?","created":"2015-03-07T18:57:12.369+0000"},{"body":"I am getting to this now. Should be fixed in 3.0. Once I have it fixed for 3.0 we can decide about back porting to 2.1.","created":"2015-03-12T16:02:16.321+0000"},{"body":"I have a dtest that can reproduce the issue. I added to JMX gcstats the current amount of in flight direct bytebuffer memory (this includes buffers that haven't been GCed). This is the value from java.nio.Bits.\n\nWhen I took a heap dump the issue was with Netty pooling memory. Netty pooled 600 megabytes of memory after I serially (and single threaded) wrote/read 5 rows with a single 35 megabyte column each. Each pooled bit of memory was 16 megabytes. I don't know yet where Netty steady states.\n\nThis doesn't match what I recall from the original user report where the memory was being pooled as part of intracluster networking. There may be another factor like the setting for intracluster compression that influences it. It may even be that enabling intracluster compression is a work around.","created":"2015-03-16T22:21:30.872+0000"},{"body":"I can confirm that compression determines whether you see an issue with memory used by intracluster messaging.\n\nThe memory used by Netty shouldn't scale with cluster size the same way. \n\nStill figuring out how to test for the issue when Netty is dominating direct bytebuffer usage. I don't see a way to get Netty to report on it's memory usage nor a way to get the allocations done by NIO tracked.\n\nMonkey patching where art thou?","created":"2015-03-17T21:21:51.165+0000"},{"body":"Took way longer than I would have liked, but here is an alternative implementation of DataInputStream and DataOutputStreamPlus that wraps a WritableByteChannel and does any necessary buffering. \n\n[Implementation available on github.|https://github.com/apache/cassandra/compare/trunk...aweisberg:C-8670?expand=1]\n\nI am also attaching a dtest validates that almost no direct byte buffer memory is allocated even when using large columns. To check how much is allocated I used reflection on java.nio.Bits and have GCInspector supply it along with the other metrics it supplies.\n\nTo make it easy to test for I added a -D flag for testing that has Netty not pool memory and prefer non-direct byte buffers.\n\nThe only other place where I think we might run into this issue is streaming. That operates on the input/output streams from sockets. With streaming you don't connect to as many nodes, and if the thread that is used for streaming is released once streaming completes is shouldn't be a problem.\n","created":"2015-03-23T22:30:16.909+0000"},{"body":"Could we get the patch rebased so that all of the commits are adjacent, and not interspersed with merges in from 2.1/trunk? It makes it difficult to follow exactly what's been changed (I've been guilty of this approach in the past, but I think it helps clean review to always ensure every commit for a patch occurs at the end of the git log)\n\nIt's also worth discussing a potential simpler approach to this: couldn't we wrap DataInputStream, and proxy read(byte[]) to a loop over read(byte[], int, int)? For DataOutputPlus we can just change the behaviour of our DataOutputStreamPlus for byte[], which would fall through to DataOutputStreamAndChannel which could use the code you have for write(ByteBuffer) only to duplicate the behaviour here. We could (and probably should, when compression is disabled) use that in OTC to remove the indirection when filling a ByteBuffer. I'm not saying for sure this is better, but since it is much simpler it seems we should refute this approach before attempting something more involved?","created":"2015-03-24T22:35:25.789+0000"},{"body":"Do you want to see it commit by commit or should I squash it?\n\nbq. It's also worth discussing a potential simpler approach to this: couldn't we wrap DataInputStream, and proxy read(byte[]) to a loop over read(byte[], int, int)? For DataOutputPlus we can just change the behaviour of our DataOutputStreamPlus for byte[], which would fall through to DataOutputStreamAndChannel which could use the code you have for write(ByteBuffer) only to duplicate the behaviour here. We could (and probably should, when compression is disabled) use that in OTC to remove the indirection when filling a ByteBuffer. I'm not saying for sure this is better, but since it is much simpler it seems we should refute this approach before attempting something more involved?\n\nI was generally trying improve the situation by reading/writing to direct buffers and in the read case not adding another wrapper class and indirection. I have no idea what kind of output the JVM has for Streams that wrap other streams.\n\nDataOutputStreamAndChannel doesn't work with BufferedOutputStream. It doesn't flush the output stream before writing to the channel. We could always change that though.\n\nI am in favor of doing whatever doesn't have me benchmarking.","created":"2015-03-24T23:18:13.071+0000"},{"body":"bq. I am in favor of doing whatever doesn't have me benchmarking.\n\nLet's not get bogged down in microbenchmarking this stuff. IMO a little analysis of the options is sufficient, since either is likely an improvement. I don't have a preconceived answer to the question I raised, I just think it's worth assessing both options in contrast to each other. It's a bit late for me to assess that myself this evening, so I'll aim to collect my thoughts on that tomorrow. Feel free to fill in yours if you have time.","created":"2015-03-24T23:36:37.230+0000"},{"body":"bq. Do you want to see it commit by commit or should I squash it?\n\nI don't mind, so long as it's easy to squash myself (it looks to be interleaved with commits from elsewhere right now, is the difficulty)","created":"2015-03-25T10:17:16.386+0000"},{"body":"What tool are you using to review? github does a good job handling merge commits and not show them as part of the diff and it doesn't show them as individual commits.\n\nYou can also convert any github comparison to a single diff by adding .diff to the URL. That's what I used to create a [new branch|https://github.com/apache/cassandra/compare/trunk...aweisberg:C-8670-2?expand=1]","created":"2015-03-25T13:59:05.588+0000"},{"body":"bq. What tool are you using to review?\n\nI like to navigate in IntelliJ, and on the command line, so having a clean run of commits helps a lot.\n\nAfter a bit of consideration, I think there's a good justification for introducing a whole new class if we intend to fully replace DataStreamOutputAndChannel, largely because the two write paths are not at all clear, and appear to be different (the old versions of the write paths being hard to actually pin down the location of in the VM source). So having a solid handle on how it behaves, and ensuring fewer code paths are executed, seems a good thing. As such, I think this patch should replace DSOaC entirely, and remove it from the codebase. I also think this is a good opportunity to share its code with DataOutputByteBuffer, and in doing hopefully make that faster, potentially improving performance of CL append (it doesn't need to extend AbstractDataOutput, and would share most of its implementation with NIODataOutputStream if it did not).\n\nA few comments in NIODataInputStream:\n\n* readNext() should assert it is never shuffling more than 7 bytes; in fact ideally this would be done by readMinimum() to make it clearer\n* readNext() should IMO never shuffle unless it's at the end of its capacity; if it hasRemaining() and limit() != capacity() it should read on from its current limit (readMinimum can ensure there is room to fully meet its requirements)\n* readUnsignedShort() could simply be: {{ return readShort() & 0xFFFF;}} \n* available() should return the bytes in the buffer at least\n* ensureMinimum() isn't clearly named, since it is more intrinsically linked to primitive reads than it suggests, consuming the bytes and throwing EOF if it cannot read. Something like preparePrimitiveRead() (no fixed idea myself, just think it is more than \"ensureMinimum\")\n\nA few comments in NIODataOutputStreamPlus:\n* close() should flush\n* close() should clean the buffer\n* why the use of hollowBuffer? For clarity in case of restoring the cursor position during exceptions? Would be helpful to clarify with a comment. It seems like perhaps this should only be used for the first branch, though, since the second should have no risk of throwing an exception, so we can safely restore the position. It seems like it might be best to make hollowBuffer default to null, and instantiate it only if it is larger than our buffer size, otherwise first flushing our internal buffer if we haven't got enough room. This way we should rarely need the hollowBuffer.\n* We should either extend our AbstractDataOutput, or make our writeUTF method public static, so we can share it\n\nFinally, it would be nice if we didn't need to stash the OutputStream version separately. Perhaps we can reorganise the class hierarchy, so that DataOutputStreamPlus doesn't wrap an internal OutputStream, it just is a light abstract class merge of the types OutputStream and DataOutputPlus. We can introduce a WrappedDataOutputStreamPlus in its place, and AbstractDataOutput could extend our new DataOutputStreamPlus instead of the other way around (with Wrapped... extending _it_). Then we can just stash a DataOutputStreamPlus in all cases. Sound reasonable?","created":"2015-03-26T11:20:19.541+0000"},{"body":"NIODataInputStream\nbq. readNext() should assert it is never shuffling more than 7 bytes; in fact ideally this would be done by readMinimum() to make it clearer\nBy assert you mean an assert that compiles out or a precondition?\n\nbq. readNext() should IMO never shuffle unless it's at the end of its capacity; if it hasRemaining() and limit() != capacity() it should read on from its current limit (readMinimum can ensure there is room to fully meet its requirements)\nI guess I don't get when this optimization will help. I could see it hurting. You could stream through the buffer not returning to the beginning on a regular basis and end up issuing smaller then desired reads.\n\nUsers of buffered input stream get this behavior and I didn't want to change it. DataInput and company pull bytes out one at a time even for multi-byte types.\n\nNIODataOutputStreamPlus\nbq. available() should return the bytes in the buffer at least\nI duplicated the JDK behavior for NIO. DataInputStream for a socket returns 0, for a file it returns the bytes remaining to read from the file. I think it makes sense for the API when you don't have a real answer.\n\nbq. why the use of hollowBuffer? For clarity in case of restoring the cursor position during exceptions? Would be helpful to clarify with a comment. It seems like perhaps this should only be used for the first branch, though, since the second should have no risk of throwing an exception, so we can safely restore the position. It seems like it might be best to make hollowBuffer default to null, and instantiate it only if it is larger than our buffer size, otherwise first flushing our internal buffer if we haven't got enough room. This way we should rarely need the hollowBuffer.\nThe contract of the API requires that the incoming buffer not be modified. For thread safety reasons I don't modify the original buffer's position and then reset it in a finally block.\n\nI am not sure what you mean by hollow buffer larger than our buffer. It's hollow so it has no size. We also use it copy things into our buffer while preserving the original position.\n\nThe rest is reasonable.\n\n\n\n","created":"2015-03-26T19:50:34.397+0000"},{"body":"bq. By assert you mean an assert that compiles out or a precondition?\n\nI don't mind, really. It's just for clarity\n\nbq. I guess I don't get when this optimization will help. I could see it hurting. You could stream through the buffer not returning to the beginning on a regular basis and end up issuing smaller then desired reads.\n\nIt was more a suggestion for clarity - at least, in my opinion. Assuming we typically fill the buffer, it isn't really a problem, and if we don't we usually have room to fill after it (although if we were to try to fill an almost full buffer it would be a problem; but so is repeatedly shuffling a buffer that is regularly very underfilled). But perhaps some more comments explaining the behaviour of (and reasoning behind) each branch is a better solution, along with the assertions to make clear this is not a costly or common operation.\n\nbq. DataInputStream for a socket returns 0\n\nDataInputStream isn't buffered. BufferedInputStream, and the API spec in InputStream#available suggest we should return number of bytes we have buffered\n\nbq. We also use it copy things into our buffer while preserving the original position.\n\nIn this scenario it is likely cheaper to simply restore the position once done, and this approach also means we can likely typically avoid ever allocating a hollow buffer. There is also no (typical) risk of exception, so no reason to use the hollow buffer, since we can guarantee we will be able to restore its position.\n\nbq. I am not sure what you mean by hollow buffer larger then our buffer\n\nI meant parameter provided buffer\n\nThere are also some formatting issues I forgot to mention (braces on the wrong line, and lots of extra linebreaks between methods)","created":"2015-03-26T20:52:52.931+0000"},{"body":"bq. In this scenario it is likely cheaper to simply restore the position once done, and this approach also means we can likely typically avoid ever allocating a hollow buffer. There is also no (typical) risk of exception, so no reason to use the hollow buffer, since we can guarantee we will be able to restore its position.\nI've been bitten by very hard to find threading bugs from having multiple threads concurrently reading from the same ByteBuffer using relative methods. It's a small amount of code to not have to worry about it ever. It's common that you will send a message referencing an \"immutable\" object graph across multiple connections and it ends up not being so immutable because serialization for some object uses a ByteBuffer without duplicating first.\n\nI don't think allocating an extra object per stream should enter into the decision making. It's tiny in the big picture.\n\nbq. but so is repeatedly shuffling a buffer that is regularly very underfilled\nTrue, but you only have to shuffle up to 7 bytes, and if the buffer is empty the cost of filling it is going to dominate. JNI calls, multiple allocations, context switch to the kernel. \n\nIn the case of socket IO we are hoping to buffer entire message, potentially all available ones so the common case is that the buffer will be empty at the end and it doesn't matter much.\n\nFor file IO doing a sequential read the common case will be that there are handful of bytes left because we are reading a multi-byte value. If we don't shuffle in that case we will do all that work to read a few bytes and then go back to fill the rest of the buffer a second time.\n\nI'll update available() to return the buffered data.","created":"2015-03-26T21:26:42.460+0000"},{"body":"bq. to not have to worry about it ever\n\nI'm not sure you can ever not worry about it, since the default behaviour is that methods _do_ modify the position of a BB, so you have to assume when it's a risk that you need to ensure each thread has its own version via duplicate(). But if that's your rationale, just comment it to explain and I'm cool with it.\n\nbq. For file IO\n\nthis should always be fully populated, but like I said: comments to explain (and assertions) will solve everything :)","created":"2015-03-26T21:40:27.710+0000"},{"body":"I think I covered what we talked about. I followed quite a few things and this is where it lead me. I don't feel like I made a dent in terms of having less code in wide use.\n\nAbstractDataOutputStreamAndChannelPlus (formerly AbstractDataOutput) is still pretty firmly entrenched.\n","created":"2015-03-27T20:29:13.235+0000"},{"body":"I've pushed some suggestions for further refactoring [here|https://github.com/belliottsmith/cassandra/tree/8670-suggestions]. I've only looked at the overall class hierarchy, I haven't focused yet on reviewing the method implementation changes.\n\nMostly these changes flatten the class hierarchy; it's gotten deep enough I don't think there's a good reason to maintain the distinction between DataStreamOutputPlus and DataStreamOutputPlusAndChannel, especially since we often just mock up a Channel based off the OutputStream. I've also flattened NIODataOutputStream and DataOutputStreamByteBufferPlus into BufferedDataOutputStreamPlus, since we only write to the buffer if we don't exceed its size. At the same time, since we are now refactoring this whole hierarchy, I made DataOutputBuffer extend BufferedDataOutputStreamPlus, and just ensures the buffer grows as necessary, and have removed FastByteArrayOutputStream since we no longer need it.\n\nI've also stopped SequentialWriter implementing WritableByteChannel, and now pass in its internal Channel, since that's the only way the operations will benefit. As a follow up ticket, we should probably move SequentialWriter to utilising BufferedDataOutputStreamPlus directly, so that it can benefit from faster encoding of primitives\n\nLet me know what you think of the changes to the hierarchy, and once we've ironed that out we can move on to the home stretch and confirm the code changes. One other thing we could consider is dropping the \"Plus\" from everything except the interface, since it seems superfluous, and it's all fairly verbose.\n","created":"2015-03-29T23:14:03.552+0000"},{"body":"I've just pushed another update, which undoes part of the SequentialWriter changes, since all writes (esp. for compression) should go through the SW itself. It also makes one further minor change to stop using Channels.newChannel() everywhere, since that introduces an extra layer of byte shuffling unnecessarily, when DataOutputPlus implements a compatible method that can be called directly.","created":"2015-03-30T15:19:32.851+0000"},{"body":"I've pushed one more round of changes [here|https://github.com/belliottsmith/cassandra/tree/8670-2], after your follow up round (which I mention for posterity). I've made the following changes; let me know your thoughts on them:\n\n* Merged writeUTF into one method, with a fast path _only_ for ASCII characters, since this is likely to benefit most from unrolling, and the instruction cache pollution effect is small. The two separate but near identical and very large methods look almost certain to be worse due to icache misses than a single branch that is mostly predicted correctly, especially when we had multiple branches inside the loop, which were each more likely to be mispredicted. As a follow-up commit, in case you're worried by this, I've introduced a no-conditional version of sizeOfChar (which we may be able to optimise further), but I haven't performed any benchmarks to measure the difference in effect.\n* Reverted the new hollowBuffer approach for array backed buffers - I couldn't see a reason for not just directly invoking the write(byte[]) methods?\n* Based SafeMemoryWriter on DataOutputBuffer\n* Shared the UBDOSP.utfBytes and DOSP.WBC.buf in the same ThreadLocal \n* Preferred bb.hasArray() to bb.isDirect(), since it is a concrete method, so can be inlined\n* Moved writeUTFLegacy into the test case, since it's only for test purposes now\n* Fixed formatting in UnbufferedDataOutputStreamPlus (seems a good opportunity to standardise it)","created":"2015-03-31T13:26:44.965+0000"},{"body":"Unit tets pass and I am +1 on your changes.","created":"2015-03-31T16:08:45.453+0000"},{"body":"Committed","created":"2015-03-31T16:29:17.762+0000"},{"body":"A little niggle was bugging me, and I decided to check it out, and I think it warrants further consideration:\n\nThe [windows|http://hg.openjdk.java.net/jdk7/jdk7/jdk/file/00cd9dc3c2b5/src/windows/native/java/net/SocketOutputStream.c] and [solaris|http://hg.openjdk.java.net/jdk6/jdk6/jdk/annotate/1a53516ce032/src/solaris/native/java/net/SocketOutputStream.c] implementations of SocketOutputStream allocate _temporary_ (not cached) heap (not Java heap) memory of the total size of the data being written to them, if they're above the size of the stack memory they allocate by default (according to the source, this ranges from 8K to 64K on [Solaris|http://hg.openjdk.java.net/jdk7/jdk7/jdk/file/4dbd83eb0250/src/solaris/native/java/net/net_util_md.h#l123], and 2K on [Windows|http://hg.openjdk.java.net/jdk7/jdk7/jdk/file/0c27202d66c1/src/windows/native/java/net/net_util_md.h#l236], for the stack buffer, with the heap buffer being a little larger)\n\nThis raises a few questions: \n# how much of our thread stack frame ends up being used by these methods? Seems we have some headroom with even the minimum stack frame, so maybe not worth worrying about; but larger writes will incur some non-java-heap allocations without much benefit\n# should we simply copy our byte[] data into our direct buffer, and call write on this instead? this has the added advantage of avoiding the OutputStream entirely in situations where we have access to the Channel, which removes an entire codepath from the common execution path. There seems to be no downside, especially if our buffer is large enough, since the copying is exactly what the native code does anyway.\n# what is behaviour on Linux? I can't for the life of me find the linux implementation; best I can assume is it is very similar to Solaris\n","created":"2015-04-01T10:41:45.339+0000"},{"body":"Your assumption is correct {{jdk/src/solaris/native/java/net/SocketOutputStream.c}} is the Linux source.\nWhat I don't understand is why {{Java_java_net_SocketOutputStream_socketWrite0}} not just uses the given {{byte[]}} but copies to a stack/heap buffer.","created":"2015-04-01T11:14:41.174+0000"},{"body":"bq. What I don't understand is why Java_java_net_SocketOutputStream_socketWrite0 not just uses the given byte[] but copies to a stack/heap buffer.\n\nTime to safe point. Using it directly would require pausing GC until the socket write had completed, which could be lengthy (since it's blocking, it's unbounded, in fact)","created":"2015-04-01T11:18:32.813+0000"},{"body":"In what scenarios do we end up actually writing to the native methods right now? In most cases we opt to use the channel when available. The only time we should be using WrappedDataOutputStreamPlus is when we aren't actually writing to the channel of a file or socket.\n\nWhere we could really do better is when using compression, we could have the compression wrap a direct buffer which wraps the channel and avoid relying on the built in mechanisms for getting data off heap.\n\nI had some other refrigerator moments overnight as well.\n\nBufferedDataOutputStreamPlus.close() will clean the buffer even if it might not own the buffer such as when it is provided to the constructor. It also doesn't check if the channel is null.\n\nWhen used from the commit log without a channel it will throw an NPE if it exceeds the capacity of the buffer when it goes to flush. I suppose one runtime exception is as good as another. There is also the extra bounds check in ensureRemaining() which always seemed a little useless to me and the fact that it will not drop through and do efficient copies for direct byte buffers. It almost seems like there is a case for a version just for wrapping fixed size bytebuffers. ","created":"2015-04-01T13:28:16.104+0000"},{"body":"I updated the microbenchmark to do what I think is the right think WRT to dead code elimination and constant folding. Results don't change, but I it looks more like I know what we I am doing.","created":"2015-04-01T13:37:52.228+0000"},{"body":"bq. In most cases we opt to use the channel when available. ... Where we could really do better is when using compression\n\nYou're right. I thought we had more paths leftover, but it is just compression really. That's a much cleaner state of affairs!\n\nIt looks like compression will typically be under the stack buffer size, so we don't have such a problem, but it would still be nice to move to a DirectByteBuffer compressed output stream. It should be quite viable since there are now compression methods for working over these directly, but that should probably be a follow up ticket. Do you want to file, or shall I?\n\nbq. BufferedDataOutputStreamPlus.close() will clean the buffer even if it might not own the buffer such as when it is provided to the constructor. It also doesn't check if the channel is null.\n\nWe don't call close in situations where this is a problem\n\nbq. When used from the commit log without a channel it will throw an NPE if it exceeds the capacity of the buffer when it goes to flush\n\nWe presize the buffer correctly\n\nStill, on both these counts we could offer better safety if we wanted, without much cost. If we used a DataOutputBuffer it would solve these problems, and we could assert that the final buffer is == the provided buffer...\n\nbq. It almost seems like there is a case for a version just for wrapping fixed size bytebuffers.\n\nPossibly. I'm on the fence, since it's extra code cache and class hierarchy pollution, with limited positive impact. We would need to duplicate the whole of BufferedDataOutputStreamPlus. You could benchmark to see if there's an appreciable difference? If the difference is small (which I expect it is, given branch prediction), I would prefer to avoid the pollution.","created":"2015-04-01T13:45:54.038+0000"},{"body":"CASSANDRA-8887 would be part of it. Maybe an umbrella ticket and do that first to figure out how hook the compression libraries up. Can you file?\n\nAlso on trying to understand code cache pollution. How much do you really know about how many instructions are emitted when the JVM can inline and duplicate stuff out the yin yang?","created":"2015-04-01T13:52:28.911+0000"},{"body":"On more thinking I don't think having close clean the buffer if we allow people to supply one is great. It could be double free or use after free if someone makes a mistake.\n\n-I think we should had least have FileUtil.clean null out the pointer to the buffer before/after cleaning (whichever is possible) so we get as immediate a failure as possible.- -(This is done by the cleaner)- Except not all buffers we create will have cleaners... Duplicates or slices of the buffer won't pick up that the pointer was nulled, but it's better then nothing. We should also assert that the buffer wasn't provided in the constructor and throw an exception if someone did that and then called close.","created":"2015-04-01T14:10:53.224+0000"},{"body":"bq. Also on trying to understand code cache pollution. How much do you really know about how many instructions are emitted when the JVM can inline and duplicate stuff out the yin yang?\n\nWell, there are a lot of heuristics to apply, that are admittedly limited and imperfect. But in general: hotspot won't inline megamorphic call sites, and even bimorphic callsites are unlikely to be (i think probably never) inlined, only given a static despatch fast path. These heuristics are enough to guide decisions around this sufficiently well in my experience. If there are multiple implementations viable at any moment, then the callsite will not be inlined, at most its location will be, and even if it _is_ inlined, this doesn't necessarily pollute the code cache, since inlined methods are small (as are most methods) and adjacent occupancy of the cache is essentially free, unless the overall method size exceeds the cache line boundary.\n\nIn general my view, the simplest heuristic is: if the benefit is small, and it increases the number of active call sites, then let's not. This works from a code management as well as a cache pollution perspective at once.","created":"2015-04-01T14:18:28.577+0000"},{"body":"bq. We should also assert that the buffer wasn't provided in the constructor and throw an exception if someone did that and then called close.\n\nIf we forbid this entirely, and only expose it via DataOutputBuffer, then the close method is harmless (as it's a no-op)","created":"2015-04-01T14:34:58.176+0000"}],"conversations":[{"body":"If you provide a large byte array to NIO and ask it to populate the byte array from a socket it will allocate a thread local byte buffer that is the size of the requested read no matter how large it is. Old IO wraps new IO for sockets (but not files) so old IO is effected as well.\n\nEven If you are using Buffered{Input | Output}Stream you can end up passing a large byte array to NIO. The byte array read method will pass the array to NIO directly if it is larger than the internal buffer. \n\nPassing large cells between nodes as part of intra-cluster messaging can cause the NIO pooled buffers to quickly reach a high watermark and stay there. This ends up costing 2x the largest cell size because there is a buffer for input and output since they are different threads. This is further multiplied by the number of nodes in the cluster - 1 since each has a dedicated thread pair with separate thread locals.\n\nAnecdotally it appears that the cost is doubled beyond that although it isn't clear why. Possibly the control connections or possibly there is some way in which multiple \n\nNeed a workload in CI that tests the advertised limits of cells on a cluster. It would be reasonable to ratchet down the max direct memory for the test to trigger failures if a memory pooling issue is introduced. I don't think we need to test concurrently pulling in a lot of them, but it should at least work serially.\n\nThe obvious fix to address this issue would be to read in smaller chunks when dealing with large values. I think small should still be relatively large (4 megabytes) so that code that is reading from a disk can amortize the cost of a seek. It can be hard to tell what the underlying thing being read from is going to be in some of the contexts where we might choose to implement switching to reading chunks.","from":"reporter","subject":"Large columns + NIO memory pooling causes excessive direct memory usage"},{"body":"What's the path forward on this?","from":"developer"},{"body":"I am getting to this now. Should be fixed in 3.0. Once I have it fixed for 3.0 we can decide about back porting to 2.1.","from":"developer"},{"body":"I have a dtest that can reproduce the issue. I added to JMX gcstats the current amount of in flight direct bytebuffer memory (this includes buffers that haven't been GCed). This is the value from java.nio.Bits.\n\nWhen I took a heap dump the issue was with Netty pooling memory. Netty pooled 600 megabytes of memory after I serially (and single threaded) wrote/read 5 rows with a single 35 megabyte column each. Each pooled bit of memory was 16 megabytes. I don't know yet where Netty steady states.\n\nThis doesn't match what I recall from the original user report where the memory was being pooled as part of intracluster networking. There may be another factor like the setting for intracluster compression that influences it. It may even be that enabling intracluster compression is a work around.","from":"developer"},{"body":"I can confirm that compression determines whether you see an issue with memory used by intracluster messaging.\n\nThe memory used by Netty shouldn't scale with cluster size the same way. \n\nStill figuring out how to test for the issue when Netty is dominating direct bytebuffer usage. I don't see a way to get Netty to report on it's memory usage nor a way to get the allocations done by NIO tracked.\n\nMonkey patching where art thou?","from":"developer"},{"body":"Took way longer than I would have liked, but here is an alternative implementation of DataInputStream and DataOutputStreamPlus that wraps a WritableByteChannel and does any necessary buffering. \n\n[Implementation available on github.|https://github.com/apache/cassandra/compare/trunk...aweisberg:C-8670?expand=1]\n\nI am also attaching a dtest validates that almost no direct byte buffer memory is allocated even when using large columns. To check how much is allocated I used reflection on java.nio.Bits and have GCInspector supply it along with the other metrics it supplies.\n\nTo make it easy to test for I added a -D flag for testing that has Netty not pool memory and prefer non-direct byte buffers.\n\nThe only other place where I think we might run into this issue is streaming. That operates on the input/output streams from sockets. With streaming you don't connect to as many nodes, and if the thread that is used for streaming is released once streaming completes is shouldn't be a problem.\n","from":"developer"},{"body":"Could we get the patch rebased so that all of the commits are adjacent, and not interspersed with merges in from 2.1/trunk? It makes it difficult to follow exactly what's been changed (I've been guilty of this approach in the past, but I think it helps clean review to always ensure every commit for a patch occurs at the end of the git log)\n\nIt's also worth discussing a potential simpler approach to this: couldn't we wrap DataInputStream, and proxy read(byte[]) to a loop over read(byte[], int, int)? For DataOutputPlus we can just change the behaviour of our DataOutputStreamPlus for byte[], which would fall through to DataOutputStreamAndChannel which could use the code you have for write(ByteBuffer) only to duplicate the behaviour here. We could (and probably should, when compression is disabled) use that in OTC to remove the indirection when filling a ByteBuffer. I'm not saying for sure this is better, but since it is much simpler it seems we should refute this approach before attempting something more involved?","from":"developer"},{"body":"Do you want to see it commit by commit or should I squash it?\n\nbq. It's also worth discussing a potential simpler approach to this: couldn't we wrap DataInputStream, and proxy read(byte[]) to a loop over read(byte[], int, int)? For DataOutputPlus we can just change the behaviour of our DataOutputStreamPlus for byte[], which would fall through to DataOutputStreamAndChannel which could use the code you have for write(ByteBuffer) only to duplicate the behaviour here. We could (and probably should, when compression is disabled) use that in OTC to remove the indirection when filling a ByteBuffer. I'm not saying for sure this is better, but since it is much simpler it seems we should refute this approach before attempting something more involved?\n\nI was generally trying improve the situation by reading/writing to direct buffers and in the read case not adding another wrapper class and indirection. I have no idea what kind of output the JVM has for Streams that wrap other streams.\n\nDataOutputStreamAndChannel doesn't work with BufferedOutputStream. It doesn't flush the output stream before writing to the channel. We could always change that though.\n\nI am in favor of doing whatever doesn't have me benchmarking.","from":"developer"},{"body":"bq. I am in favor of doing whatever doesn't have me benchmarking.\n\nLet's not get bogged down in microbenchmarking this stuff. IMO a little analysis of the options is sufficient, since either is likely an improvement. I don't have a preconceived answer to the question I raised, I just think it's worth assessing both options in contrast to each other. It's a bit late for me to assess that myself this evening, so I'll aim to collect my thoughts on that tomorrow. Feel free to fill in yours if you have time.","from":"developer"},{"body":"bq. Do you want to see it commit by commit or should I squash it?\n\nI don't mind, so long as it's easy to squash myself (it looks to be interleaved with commits from elsewhere right now, is the difficulty)","from":"developer"},{"body":"What tool are you using to review? github does a good job handling merge commits and not show them as part of the diff and it doesn't show them as individual commits.\n\nYou can also convert any github comparison to a single diff by adding .diff to the URL. That's what I used to create a [new branch|https://github.com/apache/cassandra/compare/trunk...aweisberg:C-8670-2?expand=1]","from":"developer"},{"body":"bq. What tool are you using to review?\n\nI like to navigate in IntelliJ, and on the command line, so having a clean run of commits helps a lot.\n\nAfter a bit of consideration, I think there's a good justification for introducing a whole new class if we intend to fully replace DataStreamOutputAndChannel, largely because the two write paths are not at all clear, and appear to be different (the old versions of the write paths being hard to actually pin down the location of in the VM source). So having a solid handle on how it behaves, and ensuring fewer code paths are executed, seems a good thing. As such, I think this patch should replace DSOaC entirely, and remove it from the codebase. I also think this is a good opportunity to share its code with DataOutputByteBuffer, and in doing hopefully make that faster, potentially improving performance of CL append (it doesn't need to extend AbstractDataOutput, and would share most of its implementation with NIODataOutputStream if it did not).\n\nA few comments in NIODataInputStream:\n\n* readNext() should assert it is never shuffling more than 7 bytes; in fact ideally this would be done by readMinimum() to make it clearer\n* readNext() should IMO never shuffle unless it's at the end of its capacity; if it hasRemaining() and limit() != capacity() it should read on from its current limit (readMinimum can ensure there is room to fully meet its requirements)\n* readUnsignedShort() could simply be: {{ return readShort() & 0xFFFF;}} \n* available() should return the bytes in the buffer at least\n* ensureMinimum() isn't clearly named, since it is more intrinsically linked to primitive reads than it suggests, consuming the bytes and throwing EOF if it cannot read. Something like preparePrimitiveRead() (no fixed idea myself, just think it is more than \"ensureMinimum\")\n\nA few comments in NIODataOutputStreamPlus:\n* close() should flush\n* close() should clean the buffer\n* why the use of hollowBuffer? For clarity in case of restoring the cursor position during exceptions? Would be helpful to clarify with a comment. It seems like perhaps this should only be used for the first branch, though, since the second should have no risk of throwing an exception, so we can safely restore the position. It seems like it might be best to make hollowBuffer default to null, and instantiate it only if it is larger than our buffer size, otherwise first flushing our internal buffer if we haven't got enough room. This way we should rarely need the hollowBuffer.\n* We should either extend our AbstractDataOutput, or make our writeUTF method public static, so we can share it\n\nFinally, it would be nice if we didn't need to stash the OutputStream version separately. Perhaps we can reorganise the class hierarchy, so that DataOutputStreamPlus doesn't wrap an internal OutputStream, it just is a light abstract class merge of the types OutputStream and DataOutputPlus. We can introduce a WrappedDataOutputStreamPlus in its place, and AbstractDataOutput could extend our new DataOutputStreamPlus instead of the other way around (with Wrapped... extending _it_). Then we can just stash a DataOutputStreamPlus in all cases. Sound reasonable?","from":"developer"},{"body":"NIODataInputStream\nbq. readNext() should assert it is never shuffling more than 7 bytes; in fact ideally this would be done by readMinimum() to make it clearer\nBy assert you mean an assert that compiles out or a precondition?\n\nbq. readNext() should IMO never shuffle unless it's at the end of its capacity; if it hasRemaining() and limit() != capacity() it should read on from its current limit (readMinimum can ensure there is room to fully meet its requirements)\nI guess I don't get when this optimization will help. I could see it hurting. You could stream through the buffer not returning to the beginning on a regular basis and end up issuing smaller then desired reads.\n\nUsers of buffered input stream get this behavior and I didn't want to change it. DataInput and company pull bytes out one at a time even for multi-byte types.\n\nNIODataOutputStreamPlus\nbq. available() should return the bytes in the buffer at least\nI duplicated the JDK behavior for NIO. DataInputStream for a socket returns 0, for a file it returns the bytes remaining to read from the file. I think it makes sense for the API when you don't have a real answer.\n\nbq. why the use of hollowBuffer? For clarity in case of restoring the cursor position during exceptions? Would be helpful to clarify with a comment. It seems like perhaps this should only be used for the first branch, though, since the second should have no risk of throwing an exception, so we can safely restore the position. It seems like it might be best to make hollowBuffer default to null, and instantiate it only if it is larger than our buffer size, otherwise first flushing our internal buffer if we haven't got enough room. This way we should rarely need the hollowBuffer.\nThe contract of the API requires that the incoming buffer not be modified. For thread safety reasons I don't modify the original buffer's position and then reset it in a finally block.\n\nI am not sure what you mean by hollow buffer larger than our buffer. It's hollow so it has no size. We also use it copy things into our buffer while preserving the original position.\n\nThe rest is reasonable.\n\n\n\n","from":"developer"},{"body":"bq. By assert you mean an assert that compiles out or a precondition?\n\nI don't mind, really. It's just for clarity\n\nbq. I guess I don't get when this optimization will help. I could see it hurting. You could stream through the buffer not returning to the beginning on a regular basis and end up issuing smaller then desired reads.\n\nIt was more a suggestion for clarity - at least, in my opinion. Assuming we typically fill the buffer, it isn't really a problem, and if we don't we usually have room to fill after it (although if we were to try to fill an almost full buffer it would be a problem; but so is repeatedly shuffling a buffer that is regularly very underfilled). But perhaps some more comments explaining the behaviour of (and reasoning behind) each branch is a better solution, along with the assertions to make clear this is not a costly or common operation.\n\nbq. DataInputStream for a socket returns 0\n\nDataInputStream isn't buffered. BufferedInputStream, and the API spec in InputStream#available suggest we should return number of bytes we have buffered\n\nbq. We also use it copy things into our buffer while preserving the original position.\n\nIn this scenario it is likely cheaper to simply restore the position once done, and this approach also means we can likely typically avoid ever allocating a hollow buffer. There is also no (typical) risk of exception, so no reason to use the hollow buffer, since we can guarantee we will be able to restore its position.\n\nbq. I am not sure what you mean by hollow buffer larger then our buffer\n\nI meant parameter provided buffer\n\nThere are also some formatting issues I forgot to mention (braces on the wrong line, and lots of extra linebreaks between methods)","from":"developer"},{"body":"bq. In this scenario it is likely cheaper to simply restore the position once done, and this approach also means we can likely typically avoid ever allocating a hollow buffer. There is also no (typical) risk of exception, so no reason to use the hollow buffer, since we can guarantee we will be able to restore its position.\nI've been bitten by very hard to find threading bugs from having multiple threads concurrently reading from the same ByteBuffer using relative methods. It's a small amount of code to not have to worry about it ever. It's common that you will send a message referencing an \"immutable\" object graph across multiple connections and it ends up not being so immutable because serialization for some object uses a ByteBuffer without duplicating first.\n\nI don't think allocating an extra object per stream should enter into the decision making. It's tiny in the big picture.\n\nbq. but so is repeatedly shuffling a buffer that is regularly very underfilled\nTrue, but you only have to shuffle up to 7 bytes, and if the buffer is empty the cost of filling it is going to dominate. JNI calls, multiple allocations, context switch to the kernel. \n\nIn the case of socket IO we are hoping to buffer entire message, potentially all available ones so the common case is that the buffer will be empty at the end and it doesn't matter much.\n\nFor file IO doing a sequential read the common case will be that there are handful of bytes left because we are reading a multi-byte value. If we don't shuffle in that case we will do all that work to read a few bytes and then go back to fill the rest of the buffer a second time.\n\nI'll update available() to return the buffered data.","from":"developer"},{"body":"bq. to not have to worry about it ever\n\nI'm not sure you can ever not worry about it, since the default behaviour is that methods _do_ modify the position of a BB, so you have to assume when it's a risk that you need to ensure each thread has its own version via duplicate(). But if that's your rationale, just comment it to explain and I'm cool with it.\n\nbq. For file IO\n\nthis should always be fully populated, but like I said: comments to explain (and assertions) will solve everything :)","from":"developer"},{"body":"I think I covered what we talked about. I followed quite a few things and this is where it lead me. I don't feel like I made a dent in terms of having less code in wide use.\n\nAbstractDataOutputStreamAndChannelPlus (formerly AbstractDataOutput) is still pretty firmly entrenched.\n","from":"developer"},{"body":"I've pushed some suggestions for further refactoring [here|https://github.com/belliottsmith/cassandra/tree/8670-suggestions]. I've only looked at the overall class hierarchy, I haven't focused yet on reviewing the method implementation changes.\n\nMostly these changes flatten the class hierarchy; it's gotten deep enough I don't think there's a good reason to maintain the distinction between DataStreamOutputPlus and DataStreamOutputPlusAndChannel, especially since we often just mock up a Channel based off the OutputStream. I've also flattened NIODataOutputStream and DataOutputStreamByteBufferPlus into BufferedDataOutputStreamPlus, since we only write to the buffer if we don't exceed its size. At the same time, since we are now refactoring this whole hierarchy, I made DataOutputBuffer extend BufferedDataOutputStreamPlus, and just ensures the buffer grows as necessary, and have removed FastByteArrayOutputStream since we no longer need it.\n\nI've also stopped SequentialWriter implementing WritableByteChannel, and now pass in its internal Channel, since that's the only way the operations will benefit. As a follow up ticket, we should probably move SequentialWriter to utilising BufferedDataOutputStreamPlus directly, so that it can benefit from faster encoding of primitives\n\nLet me know what you think of the changes to the hierarchy, and once we've ironed that out we can move on to the home stretch and confirm the code changes. One other thing we could consider is dropping the \"Plus\" from everything except the interface, since it seems superfluous, and it's all fairly verbose.\n","from":"developer"},{"body":"I've just pushed another update, which undoes part of the SequentialWriter changes, since all writes (esp. for compression) should go through the SW itself. It also makes one further minor change to stop using Channels.newChannel() everywhere, since that introduces an extra layer of byte shuffling unnecessarily, when DataOutputPlus implements a compatible method that can be called directly.","from":"developer"},{"body":"I've pushed one more round of changes [here|https://github.com/belliottsmith/cassandra/tree/8670-2], after your follow up round (which I mention for posterity). I've made the following changes; let me know your thoughts on them:\n\n* Merged writeUTF into one method, with a fast path _only_ for ASCII characters, since this is likely to benefit most from unrolling, and the instruction cache pollution effect is small. The two separate but near identical and very large methods look almost certain to be worse due to icache misses than a single branch that is mostly predicted correctly, especially when we had multiple branches inside the loop, which were each more likely to be mispredicted. As a follow-up commit, in case you're worried by this, I've introduced a no-conditional version of sizeOfChar (which we may be able to optimise further), but I haven't performed any benchmarks to measure the difference in effect.\n* Reverted the new hollowBuffer approach for array backed buffers - I couldn't see a reason for not just directly invoking the write(byte[]) methods?\n* Based SafeMemoryWriter on DataOutputBuffer\n* Shared the UBDOSP.utfBytes and DOSP.WBC.buf in the same ThreadLocal \n* Preferred bb.hasArray() to bb.isDirect(), since it is a concrete method, so can be inlined\n* Moved writeUTFLegacy into the test case, since it's only for test purposes now\n* Fixed formatting in UnbufferedDataOutputStreamPlus (seems a good opportunity to standardise it)","from":"developer"},{"body":"Unit tets pass and I am +1 on your changes.","from":"developer"},{"body":"Committed","from":"developer"},{"body":"A little niggle was bugging me, and I decided to check it out, and I think it warrants further consideration:\n\nThe [windows|http://hg.openjdk.java.net/jdk7/jdk7/jdk/file/00cd9dc3c2b5/src/windows/native/java/net/SocketOutputStream.c] and [solaris|http://hg.openjdk.java.net/jdk6/jdk6/jdk/annotate/1a53516ce032/src/solaris/native/java/net/SocketOutputStream.c] implementations of SocketOutputStream allocate _temporary_ (not cached) heap (not Java heap) memory of the total size of the data being written to them, if they're above the size of the stack memory they allocate by default (according to the source, this ranges from 8K to 64K on [Solaris|http://hg.openjdk.java.net/jdk7/jdk7/jdk/file/4dbd83eb0250/src/solaris/native/java/net/net_util_md.h#l123], and 2K on [Windows|http://hg.openjdk.java.net/jdk7/jdk7/jdk/file/0c27202d66c1/src/windows/native/java/net/net_util_md.h#l236], for the stack buffer, with the heap buffer being a little larger)\n\nThis raises a few questions: \n# how much of our thread stack frame ends up being used by these methods? Seems we have some headroom with even the minimum stack frame, so maybe not worth worrying about; but larger writes will incur some non-java-heap allocations without much benefit\n# should we simply copy our byte[] data into our direct buffer, and call write on this instead? this has the added advantage of avoiding the OutputStream entirely in situations where we have access to the Channel, which removes an entire codepath from the common execution path. There seems to be no downside, especially if our buffer is large enough, since the copying is exactly what the native code does anyway.\n# what is behaviour on Linux? I can't for the life of me find the linux implementation; best I can assume is it is very similar to Solaris\n","from":"developer"},{"body":"Your assumption is correct {{jdk/src/solaris/native/java/net/SocketOutputStream.c}} is the Linux source.\nWhat I don't understand is why {{Java_java_net_SocketOutputStream_socketWrite0}} not just uses the given {{byte[]}} but copies to a stack/heap buffer.","from":"developer"},{"body":"bq. What I don't understand is why Java_java_net_SocketOutputStream_socketWrite0 not just uses the given byte[] but copies to a stack/heap buffer.\n\nTime to safe point. Using it directly would require pausing GC until the socket write had completed, which could be lengthy (since it's blocking, it's unbounded, in fact)","from":"developer"},{"body":"In what scenarios do we end up actually writing to the native methods right now? In most cases we opt to use the channel when available. The only time we should be using WrappedDataOutputStreamPlus is when we aren't actually writing to the channel of a file or socket.\n\nWhere we could really do better is when using compression, we could have the compression wrap a direct buffer which wraps the channel and avoid relying on the built in mechanisms for getting data off heap.\n\nI had some other refrigerator moments overnight as well.\n\nBufferedDataOutputStreamPlus.close() will clean the buffer even if it might not own the buffer such as when it is provided to the constructor. It also doesn't check if the channel is null.\n\nWhen used from the commit log without a channel it will throw an NPE if it exceeds the capacity of the buffer when it goes to flush. I suppose one runtime exception is as good as another. There is also the extra bounds check in ensureRemaining() which always seemed a little useless to me and the fact that it will not drop through and do efficient copies for direct byte buffers. It almost seems like there is a case for a version just for wrapping fixed size bytebuffers. ","from":"developer"},{"body":"I updated the microbenchmark to do what I think is the right think WRT to dead code elimination and constant folding. Results don't change, but I it looks more like I know what we I am doing.","from":"developer"},{"body":"bq. In most cases we opt to use the channel when available. ... Where we could really do better is when using compression\n\nYou're right. I thought we had more paths leftover, but it is just compression really. That's a much cleaner state of affairs!\n\nIt looks like compression will typically be under the stack buffer size, so we don't have such a problem, but it would still be nice to move to a DirectByteBuffer compressed output stream. It should be quite viable since there are now compression methods for working over these directly, but that should probably be a follow up ticket. Do you want to file, or shall I?\n\nbq. BufferedDataOutputStreamPlus.close() will clean the buffer even if it might not own the buffer such as when it is provided to the constructor. It also doesn't check if the channel is null.\n\nWe don't call close in situations where this is a problem\n\nbq. When used from the commit log without a channel it will throw an NPE if it exceeds the capacity of the buffer when it goes to flush\n\nWe presize the buffer correctly\n\nStill, on both these counts we could offer better safety if we wanted, without much cost. If we used a DataOutputBuffer it would solve these problems, and we could assert that the final buffer is == the provided buffer...\n\nbq. It almost seems like there is a case for a version just for wrapping fixed size bytebuffers.\n\nPossibly. I'm on the fence, since it's extra code cache and class hierarchy pollution, with limited positive impact. We would need to duplicate the whole of BufferedDataOutputStreamPlus. You could benchmark to see if there's an appreciable difference? If the difference is small (which I expect it is, given branch prediction), I would prefer to avoid the pollution.","from":"developer"},{"body":"CASSANDRA-8887 would be part of it. Maybe an umbrella ticket and do that first to figure out how hook the compression libraries up. Can you file?\n\nAlso on trying to understand code cache pollution. How much do you really know about how many instructions are emitted when the JVM can inline and duplicate stuff out the yin yang?","from":"developer"},{"body":"On more thinking I don't think having close clean the buffer if we allow people to supply one is great. It could be double free or use after free if someone makes a mistake.\n\n-I think we should had least have FileUtil.clean null out the pointer to the buffer before/after cleaning (whichever is possible) so we get as immediate a failure as possible.- -(This is done by the cleaner)- Except not all buffers we create will have cleaners... Duplicates or slices of the buffer won't pick up that the pointer was nulled, but it's better then nothing. We should also assert that the buffer wasn't provided in the constructor and throw an exception if someone did that and then called close.","from":"developer"},{"body":"bq. Also on trying to understand code cache pollution. How much do you really know about how many instructions are emitted when the JVM can inline and duplicate stuff out the yin yang?\n\nWell, there are a lot of heuristics to apply, that are admittedly limited and imperfect. But in general: hotspot won't inline megamorphic call sites, and even bimorphic callsites are unlikely to be (i think probably never) inlined, only given a static despatch fast path. These heuristics are enough to guide decisions around this sufficiently well in my experience. If there are multiple implementations viable at any moment, then the callsite will not be inlined, at most its location will be, and even if it _is_ inlined, this doesn't necessarily pollute the code cache, since inlined methods are small (as are most methods) and adjacent occupancy of the cache is essentially free, unless the overall method size exceeds the cache line boundary.\n\nIn general my view, the simplest heuristic is: if the benefit is small, and it increases the number of active call sites, then let's not. This works from a code management as well as a cache pollution perspective at once.","from":"developer"},{"body":"bq. We should also assert that the buffer wasn't provided in the constructor and throw an exception if someone did that and then called close.\n\nIf we forbid this entirely, and only expose it via DataOutputBuffer, then the close method is harmless (as it's a no-op)","from":"developer"}],"created":"2015-01-22T23:36:37.000+0000","description":"If you provide a large byte array to NIO and ask it to populate the byte array from a socket it will allocate a thread local byte buffer that is the size of the requested read no matter how large it is. Old IO wraps new IO for sockets (but not files) so old IO is effected as well.\n\nEven If you are using Buffered{Input | Output}Stream you can end up passing a large byte array to NIO. The byte array read method will pass the array to NIO directly if it is larger than the internal buffer. \n\nPassing large cells between nodes as part of intra-cluster messaging can cause the NIO pooled buffers to quickly reach a high watermark and stay there. This ends up costing 2x the largest cell size because there is a buffer for input and output since they are different threads. This is further multiplied by the number of nodes in the cluster - 1 since each has a dedicated thread pair with separate thread locals.\n\nAnecdotally it appears that the cost is doubled beyond that although it isn't clear why. Possibly the control connections or possibly there is some way in which multiple \n\nNeed a workload in CI that tests the advertised limits of cells on a cluster. It would be reasonable to ratchet down the max direct memory for the test to trigger failures if a memory pooling issue is introduced. I don't think we need to test concurrently pulling in a lot of them, but it should at least work serially.\n\nThe obvious fix to address this issue would be to read in smaller chunks when dealing with large values. I think small should still be relatively large (4 megabytes) so that code that is reading from a disk can amortize the cost of a seek. It can be hard to tell what the underlying thing being read from is going to be in some of the contexts where we might choose to implement switching to reading chunks.","issue_id":"12769480","key":"CASSANDRA-8670","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-04-03T21:44:47.000+0000","role":"fixed_distractor","summary":"Large columns + NIO memory pooling causes excessive direct memory usage"} {"case_id":"12839532","cluster":"DISTRACTOR-CASSANDRA-9630","comments":[{"body":"This is likely the lack of MessageService not setting SO_LINGER to 0 on the sockets.","created":"2016-05-11T15:57:30.883+0000"},{"body":"I am experiencing a similar problem. \n\nCassandra version: 3.0.8\nEnvironment: Amazon EC2\n\nError Case:\nWhen I restart Cassandra service on a node, after the node comes up it sees some or all of other nodes as DN even though other nodes see this node as UN. \n\nHere is the output of netstat and nodetool status for this error case:\n\n1. right after stopping cassandra service on node 10.4.68.222:\n{code}\n--------------------------------------\nip-10-4-54-176\ntcp 0 0 10.4.54.176:51268 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.54.176:56135 10.4.68.222:7000 TIME_WAIT \ntcp 1 0 10.4.54.176:43697 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.54.176:52372 10.4.68.222:7000 TIME_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-54-177\ntcp 0 0 10.4.54.177:56960 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.54.177:54539 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.54.177:32823 10.4.68.222:7000 TIME_WAIT \ntcp 1 0 10.4.54.177:48985 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-68-222\ntcp 0 0 10.4.68.222:7000 10.4.54.176:43697 FIN_WAIT2 \ntcp 0 0 10.4.68.222:7000 10.4.54.177:48985 FIN_WAIT2 \ntcp 0 0 10.4.68.222:7000 10.4.68.222:54419 TIME_WAIT \ntcp 0 0 10.4.68.222:7000 10.4.43.65:43197 FIN_WAIT2 \ntcp 0 0 10.4.68.222:7000 10.4.68.221:44149 FIN_WAIT2 \ntcp 0 0 10.4.68.222:7000 10.4.68.222:41302 TIME_WAIT \ntcp 0 0 10.4.68.222:7000 10.4.43.66:54321 FIN_WAIT2 \n--------------------------------------\n--------------------------------------\nip-10-4-68-221\ntcp 0 0 10.4.68.221:49599 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.68.221:55033 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.68.221:51628 10.4.68.222:7000 TIME_WAIT \ntcp 1 0 10.4.68.221:44149 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-43-66\ntcp 0 0 10.4.43.66:55930 10.4.68.222:7000 TIME_WAIT \ntcp 1 0 10.4.43.66:54321 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.43.66:60968 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.43.66:49087 10.4.68.222:7000 TIME_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-43-65\ntcp 1 0 10.4.43.65:43197 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.43.65:36467 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.43.65:53317 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.43.65:54897 10.4.68.222:7000 TIME_WAIT \n--------------------------------------\n{code}\n\n2. a bit after stopping cassandra service on node 10.4.68.222:\n{code}\n--------------------------------------\nip-10-4-54-176\ntcp 1 0 10.4.54.176:43697 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-54-177\n--------------------------------------\n--------------------------------------\nip-10-4-68-222\n--------------------------------------\n--------------------------------------\nip-10-4-68-221\ntcp 1 0 10.4.68.221:44149 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-43-66\ntcp 1 0 10.4.43.66:54321 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-43-65\ntcp 1 0 10.4.43.65:43197 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n{code}\n\n3. after starting cassandra service on node 10.4.68.222: \n{code}\n--------------------------------------\nip-10-4-54-176\ntcp 0 0 10.4.54.176:42460 10.4.68.222:7000 ESTABLISHED \ntcp 1 303403 10.4.54.176:43697 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.54.176:42109 10.4.68.222:7000 ESTABLISHED \n--------------------------------------\n--------------------------------------\nip-10-4-54-177\ntcp 0 0 10.4.54.177:43687 10.4.68.222:7000 ESTABLISHED \ntcp 0 0 10.4.54.177:56107 10.4.68.222:7000 ESTABLISHED \ntcp 0 0 10.4.54.177:39426 10.4.68.222:7000 ESTABLISHED \n--------------------------------------\n--------------------------------------\nip-10-4-68-222\ntcp 0 0 10.4.68.222:7000 0.0.0.0:* LISTEN \ntcp 0 0 10.4.68.222:7000 10.4.54.176:42109 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.54.177:43687 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.54.176:42460 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.43.66:55168 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.43.65:60239 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.54.177:39426 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.43.65:43480 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.68.221:54490 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.68.221:59771 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.54.177:56107 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.43.66:55581 ESTABLISHED \n--------------------------------------\n--------------------------------------\nip-10-4-68-221\ntcp 0 0 10.4.68.221:54490 10.4.68.222:7000 ESTABLISHED \ntcp 0 0 10.4.68.221:59771 10.4.68.222:7000 ESTABLISHED \ntcp 1 304316 10.4.68.221:44149 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-43-66\ntcp 1 322344 10.4.43.66:54321 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.43.66:55581 10.4.68.222:7000 ESTABLISHED \ntcp 0 0 10.4.43.66:55168 10.4.68.222:7000 ESTABLISHED \n--------------------------------------\n--------------------------------------\nip-10-4-43-65\ntcp 1 376331 10.4.43.65:43197 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.43.65:43480 10.4.68.222:7000 ESTABLISHED \ntcp 0 0 10.4.43.65:60239 10.4.68.222:7000 ESTABLISHED \n--------------------------------------\n{code}\n\n4. nodetool status on all nodes after starting cassandra service on node 10.4.68.222:\n{code}\nip-10-4-54-176\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.4.54.176 127.67 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nUN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nUN 10.4.43.66 141.94 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nUN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n--------------------------------------\n--------------------------------------\nip-10-4-54-177\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.4.54.176 127.63 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nUN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nUN 10.4.43.66 141.94 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nUN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n------------------------\n--------------------------------------\nip-10-4-68-222\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nDN 10.4.54.176 127.63 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nDN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nDN 10.4.43.66 141.94 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nDN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n--------------------------------------\n--------------------------------------\nip-10-4-68-221\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.4.54.176 127.63 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nUN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nUN 10.4.43.66 141.94 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nUN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n--------------------------------------\n--------------------------------------\nip-10-4-43-66\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.4.54.176 127.67 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nUN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nUN 10.4.43.66 141.95 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nUN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n--------------------------------------\n--------------------------------------\nip-10-4-43-65\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.4.54.176 127.67 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nUN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nUN 10.4.43.66 141.94 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nUN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n--------------------------------------\n{code}\n","created":"2016-07-28T01:13:37.556+0000"},{"body":"I noticed we're not closing the socket [if there is an exception while connecting|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/net/OutboundTcpConnection.java#L496] on {{OutboundTcpConnection}}, so a race during a node's shutdown might cause a failed connection attempt to that node remain in the {{CLOSE_WAIT}} state until the next GC, which could potentially cause this.\n\n[~farzad.panahi] Are you willing to try out [this patch|https://github.com/pauloricardomg/cassandra/commit/3f46d414b06afb607b6a97152661b10c53c103e6] to see if it fixes it? You need to replace your {{lib/apache-cassandra-3.0.8.jar}} with [apache-cassandra-3.0.8-SNAPSHOT.jar|https://issues.apache.org/jira/secure/attachment/12820814/apache-cassandra-3.0.8-SNAPSHOT.jar] and perform a rolling restart on some of the nodes and check if this will fix the issue in these nodes (if you prefer you can generate your own jar by cloning [this branch|https://github.com/pauloricardomg/cassandra/tree/3.0.6-9630] and running {{ant clean jar}}).\n\nIf this does not solve it, it would be nice if you could set the logging level of the {{org.apache.cassandra.net}} package to {{TRACE}}, either via {{nodetool setlogginglevel org.apache.cassandra.net TRACE}} or by adding {{}} to the end of your {{conf/logback.xml}}. After this, please attach the relevant information in the logs of affected nodes to this ticket for further analysis.","created":"2016-07-28T22:33:56.702+0000"},{"body":"Is there a plan to fix this issue, we are observing the same behavior in 3.9 version. We have 3 node Cassandra cluster running in 3 VMs. When we reboot one of the node (say VM1), we noticed socket connections on the other nodes (say VM2, VM3) are still in CLOSE_WAIT state, hence when we start Cassandra on the rebooted node it is not able to join the cluster. We observed nodetool status returning \"UN\" for itself and \"DN\" for other 2 nodes, however after 5-20 minutes we notice \"Connection Timeout\" exception in debug.log on the other 2 nodes (VM2 and VM3) and new socket connection being established and they are able to join the cluster.\r\n","created":"2018-01-03T01:16:44.753+0000"},{"body":"Even though we didn't hear back from someone who tested the patch, I'm quite confident this will fix the hanging sockets problem, so I will set this to patch available.\r\n\r\nWould you mind having a look [~snazy]? Patch [here|https://github.com/pauloricardomg/cassandra/tree/3.0-9630]. Submitted CI, will update after results.","created":"2018-01-10T18:00:56.185+0000"},{"body":"+1 it should fix the issue.","created":"2018-01-12T13:20:54.762+0000"},{"body":"Thanks for the review! Committed as \\{{51bf51813c4a7a9f9ad3adfe8ddac171b398816b}} to cassandra-3.0 and merged up to trunk.","created":"2018-01-15T12:42:09.074+0000"}],"conversations":[{"body":"After upgrading from Cassandra from 2.0.12 to 2.0.15, whenever we killed a cassandra process (with SIGTERM), some other nodes maintained a connection with the killed node in the CLOSE_WAIT state on port 7000 for about 5-20 minutes.\n\nSo, when we started the killed node again, other nodes could not establish a handshake because of the connections on the CLOSE_WAIT state, so they remained on the DOWN state to each other until the initial connection expired.\n\nThe problem did not happen if I ran a nodetool disablegossip before killing the node.\n\nI was able to fix this issue by reverting the CASSANDRA-8336 commits (including CASSANDRA-9238). After reverting this, cassandra now closes connection correctly when killed with -TERM, but leaves connections on CLOSE_WAIT state if I run nodetool disablethrift before killing the nodes.\n\nI did not try to reproduce the problem in a clean environment.","from":"reporter","subject":"Killing cassandra process results in unclosed connections"},{"body":"This is likely the lack of MessageService not setting SO_LINGER to 0 on the sockets.","from":"developer"},{"body":"I am experiencing a similar problem. \n\nCassandra version: 3.0.8\nEnvironment: Amazon EC2\n\nError Case:\nWhen I restart Cassandra service on a node, after the node comes up it sees some or all of other nodes as DN even though other nodes see this node as UN. \n\nHere is the output of netstat and nodetool status for this error case:\n\n1. right after stopping cassandra service on node 10.4.68.222:\n{code}\n--------------------------------------\nip-10-4-54-176\ntcp 0 0 10.4.54.176:51268 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.54.176:56135 10.4.68.222:7000 TIME_WAIT \ntcp 1 0 10.4.54.176:43697 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.54.176:52372 10.4.68.222:7000 TIME_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-54-177\ntcp 0 0 10.4.54.177:56960 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.54.177:54539 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.54.177:32823 10.4.68.222:7000 TIME_WAIT \ntcp 1 0 10.4.54.177:48985 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-68-222\ntcp 0 0 10.4.68.222:7000 10.4.54.176:43697 FIN_WAIT2 \ntcp 0 0 10.4.68.222:7000 10.4.54.177:48985 FIN_WAIT2 \ntcp 0 0 10.4.68.222:7000 10.4.68.222:54419 TIME_WAIT \ntcp 0 0 10.4.68.222:7000 10.4.43.65:43197 FIN_WAIT2 \ntcp 0 0 10.4.68.222:7000 10.4.68.221:44149 FIN_WAIT2 \ntcp 0 0 10.4.68.222:7000 10.4.68.222:41302 TIME_WAIT \ntcp 0 0 10.4.68.222:7000 10.4.43.66:54321 FIN_WAIT2 \n--------------------------------------\n--------------------------------------\nip-10-4-68-221\ntcp 0 0 10.4.68.221:49599 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.68.221:55033 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.68.221:51628 10.4.68.222:7000 TIME_WAIT \ntcp 1 0 10.4.68.221:44149 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-43-66\ntcp 0 0 10.4.43.66:55930 10.4.68.222:7000 TIME_WAIT \ntcp 1 0 10.4.43.66:54321 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.43.66:60968 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.43.66:49087 10.4.68.222:7000 TIME_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-43-65\ntcp 1 0 10.4.43.65:43197 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.43.65:36467 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.43.65:53317 10.4.68.222:7000 TIME_WAIT \ntcp 0 0 10.4.43.65:54897 10.4.68.222:7000 TIME_WAIT \n--------------------------------------\n{code}\n\n2. a bit after stopping cassandra service on node 10.4.68.222:\n{code}\n--------------------------------------\nip-10-4-54-176\ntcp 1 0 10.4.54.176:43697 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-54-177\n--------------------------------------\n--------------------------------------\nip-10-4-68-222\n--------------------------------------\n--------------------------------------\nip-10-4-68-221\ntcp 1 0 10.4.68.221:44149 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-43-66\ntcp 1 0 10.4.43.66:54321 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-43-65\ntcp 1 0 10.4.43.65:43197 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n{code}\n\n3. after starting cassandra service on node 10.4.68.222: \n{code}\n--------------------------------------\nip-10-4-54-176\ntcp 0 0 10.4.54.176:42460 10.4.68.222:7000 ESTABLISHED \ntcp 1 303403 10.4.54.176:43697 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.54.176:42109 10.4.68.222:7000 ESTABLISHED \n--------------------------------------\n--------------------------------------\nip-10-4-54-177\ntcp 0 0 10.4.54.177:43687 10.4.68.222:7000 ESTABLISHED \ntcp 0 0 10.4.54.177:56107 10.4.68.222:7000 ESTABLISHED \ntcp 0 0 10.4.54.177:39426 10.4.68.222:7000 ESTABLISHED \n--------------------------------------\n--------------------------------------\nip-10-4-68-222\ntcp 0 0 10.4.68.222:7000 0.0.0.0:* LISTEN \ntcp 0 0 10.4.68.222:7000 10.4.54.176:42109 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.54.177:43687 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.54.176:42460 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.43.66:55168 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.43.65:60239 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.54.177:39426 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.43.65:43480 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.68.221:54490 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.68.221:59771 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.54.177:56107 ESTABLISHED \ntcp 0 0 10.4.68.222:7000 10.4.43.66:55581 ESTABLISHED \n--------------------------------------\n--------------------------------------\nip-10-4-68-221\ntcp 0 0 10.4.68.221:54490 10.4.68.222:7000 ESTABLISHED \ntcp 0 0 10.4.68.221:59771 10.4.68.222:7000 ESTABLISHED \ntcp 1 304316 10.4.68.221:44149 10.4.68.222:7000 CLOSE_WAIT \n--------------------------------------\n--------------------------------------\nip-10-4-43-66\ntcp 1 322344 10.4.43.66:54321 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.43.66:55581 10.4.68.222:7000 ESTABLISHED \ntcp 0 0 10.4.43.66:55168 10.4.68.222:7000 ESTABLISHED \n--------------------------------------\n--------------------------------------\nip-10-4-43-65\ntcp 1 376331 10.4.43.65:43197 10.4.68.222:7000 CLOSE_WAIT \ntcp 0 0 10.4.43.65:43480 10.4.68.222:7000 ESTABLISHED \ntcp 0 0 10.4.43.65:60239 10.4.68.222:7000 ESTABLISHED \n--------------------------------------\n{code}\n\n4. nodetool status on all nodes after starting cassandra service on node 10.4.68.222:\n{code}\nip-10-4-54-176\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.4.54.176 127.67 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nUN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nUN 10.4.43.66 141.94 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nUN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n--------------------------------------\n--------------------------------------\nip-10-4-54-177\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.4.54.176 127.63 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nUN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nUN 10.4.43.66 141.94 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nUN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n------------------------\n--------------------------------------\nip-10-4-68-222\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nDN 10.4.54.176 127.63 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nDN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nDN 10.4.43.66 141.94 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nDN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n--------------------------------------\n--------------------------------------\nip-10-4-68-221\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.4.54.176 127.63 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nUN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nUN 10.4.43.66 141.94 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nUN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n--------------------------------------\n--------------------------------------\nip-10-4-43-66\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.4.54.176 127.67 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nUN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nUN 10.4.43.66 141.95 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nUN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n--------------------------------------\n--------------------------------------\nip-10-4-43-65\nDatacenter: us-east\n===================\nStatus=Up/Down\n|/ State=Normal/Leaving/Joining/Moving\n-- Address Load Tokens Owns (effective) Host ID Rack\nUN 10.4.54.176 127.67 GB 256 47.5% 7163bf77-2fef-4e33-81c1-0e61038dece1 1b\nUN 10.4.43.65 124.19 GB 256 46.2% 80265afb-8beb-4887-a696-fc9b75956894 1a\nUN 10.4.54.177 136.06 GB 256 50.7% b9010e24-4e92-4212-8a17-65892ea9ff66 1b\nUN 10.4.43.66 141.94 GB 256 52.3% b00fdf10-1075-4953-8a96-caf375221684 1a\nUN 10.4.68.221 137.12 GB 256 50.7% 37479ec3-7b6d-4537-975c-f9d95e92ee1d 1d\nUN 10.4.68.222 141.89 GB 256 52.7% 8df87657-c39b-405a-ba54-d60b577c1429 1d\n--------------------------------------\n{code}\n","from":"developer"},{"body":"I noticed we're not closing the socket [if there is an exception while connecting|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/net/OutboundTcpConnection.java#L496] on {{OutboundTcpConnection}}, so a race during a node's shutdown might cause a failed connection attempt to that node remain in the {{CLOSE_WAIT}} state until the next GC, which could potentially cause this.\n\n[~farzad.panahi] Are you willing to try out [this patch|https://github.com/pauloricardomg/cassandra/commit/3f46d414b06afb607b6a97152661b10c53c103e6] to see if it fixes it? You need to replace your {{lib/apache-cassandra-3.0.8.jar}} with [apache-cassandra-3.0.8-SNAPSHOT.jar|https://issues.apache.org/jira/secure/attachment/12820814/apache-cassandra-3.0.8-SNAPSHOT.jar] and perform a rolling restart on some of the nodes and check if this will fix the issue in these nodes (if you prefer you can generate your own jar by cloning [this branch|https://github.com/pauloricardomg/cassandra/tree/3.0.6-9630] and running {{ant clean jar}}).\n\nIf this does not solve it, it would be nice if you could set the logging level of the {{org.apache.cassandra.net}} package to {{TRACE}}, either via {{nodetool setlogginglevel org.apache.cassandra.net TRACE}} or by adding {{}} to the end of your {{conf/logback.xml}}. After this, please attach the relevant information in the logs of affected nodes to this ticket for further analysis.","from":"developer"},{"body":"Is there a plan to fix this issue, we are observing the same behavior in 3.9 version. We have 3 node Cassandra cluster running in 3 VMs. When we reboot one of the node (say VM1), we noticed socket connections on the other nodes (say VM2, VM3) are still in CLOSE_WAIT state, hence when we start Cassandra on the rebooted node it is not able to join the cluster. We observed nodetool status returning \"UN\" for itself and \"DN\" for other 2 nodes, however after 5-20 minutes we notice \"Connection Timeout\" exception in debug.log on the other 2 nodes (VM2 and VM3) and new socket connection being established and they are able to join the cluster.\r\n","from":"developer"},{"body":"Even though we didn't hear back from someone who tested the patch, I'm quite confident this will fix the hanging sockets problem, so I will set this to patch available.\r\n\r\nWould you mind having a look [~snazy]? Patch [here|https://github.com/pauloricardomg/cassandra/tree/3.0-9630]. Submitted CI, will update after results.","from":"developer"},{"body":"+1 it should fix the issue.","from":"developer"},{"body":"Thanks for the review! Committed as \\{{51bf51813c4a7a9f9ad3adfe8ddac171b398816b}} to cassandra-3.0 and merged up to trunk.","from":"developer"}],"created":"2015-06-22T12:29:28.000+0000","description":"After upgrading from Cassandra from 2.0.12 to 2.0.15, whenever we killed a cassandra process (with SIGTERM), some other nodes maintained a connection with the killed node in the CLOSE_WAIT state on port 7000 for about 5-20 minutes.\n\nSo, when we started the killed node again, other nodes could not establish a handshake because of the connections on the CLOSE_WAIT state, so they remained on the DOWN state to each other until the initial connection expired.\n\nThe problem did not happen if I ran a nodetool disablegossip before killing the node.\n\nI was able to fix this issue by reverting the CASSANDRA-8336 commits (including CASSANDRA-9238). After reverting this, cassandra now closes connection correctly when killed with -TERM, but leaves connections on CLOSE_WAIT state if I run nodetool disablethrift before killing the nodes.\n\nI did not try to reproduce the problem in a clean environment.","issue_id":"12839532","key":"CASSANDRA-9630","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2018-01-15T12:42:09.000+0000","role":"fixed_distractor","summary":"Killing cassandra process results in unclosed connections"} {"case_id":"12682950","cluster":"DISTRACTOR-HADOOP-10146","comments":[{"body":"This patch uses a hack to workaround the bug. Synch'ing on the streams before closing dovetails with the synch'ed {{ProcessPipeInputStream.drainInputStream}}. The hack is a safe no-op on JDK6 because it does not drain the streams.\n\nY! has been using this patch in production for 8 months. The problem was immediately reported to Oracle but a fix will not be available until around mid-year so we're providing this workaround to the community.","created":"2013-12-05T17:43:01.449+0000"},{"body":"No tests are included due to the difficulty of reproducing the bug on demand, and it's not hadoop's responsibility to have unit tests for JDK bugs.","created":"2013-12-05T17:43:59.678+0000"},{"body":"Anyone want to review? After moving to JDK7 in production, we had many NMs under load going OOM and crashing due to this bug. Task retries masked that the cluster was slowly shrinking. As noted above, we've been running production clusters for 8 months with this patch.","created":"2013-12-11T17:28:37.739+0000"},{"body":"bq. ProcessPipeInputStream.drainInputStream's will OOM allocating an array if in.available() returns a huge number, or may wreak havoc by incorrectly draining the fd.\nYou've seen OOMs in practice? Or the closing of wrong streams?\n\nI think I understand the race. Patch changes look good. Great that it's a no-op in JDK 6.\n\nMay be link to the corresponding source in openJDK here for posterity?\n\nAlso, may be you should be put the exact JVM version in the comment - to ease reasoning in the future.","created":"2013-12-11T22:22:49.799+0000"},{"body":"JDK bug has been verified as such by the devs and patched in 7u60: \nhttps://bugs.openjdk.java.net/browse/JDK-8024521","created":"2013-12-17T09:50:45.780+0000"},{"body":"Yes, we were losing ~10s of NMs/day because of OOMs caused by this bug. After the patch, no OOMs.\n\nThe referenced openjdk bug is indeed the same problem.\n\nDo I have a +1 to commit?","created":"2013-12-18T21:16:23.113+0000"},{"body":"Can you add a reference to this JIRA as well as a link to the corresponding (buggy) JVM version in the java comment for posterity? We can revisit and revert this when the time is apt.","created":"2013-12-19T02:26:03.369+0000"},{"body":"Yeah, +1 barring that point about code comment.","created":"2013-12-19T02:26:29.958+0000"},{"body":"Added comments based on info provided by Steve. Will commit when pre-commit passes unless there's objections to the comments.\n\nNote: the original patch filenames had the wrong jira...","created":"2014-01-15T19:17:49.500+0000"},{"body":"The comment is informative enough, the patch looks good to me. +1.","created":"2014-01-15T20:25:41.619+0000"},{"body":"+1","created":"2014-01-15T20:42:57.109+0000"},{"body":"Committed to trunk, branch-2, and branch-0.23.","created":"2014-01-16T18:58:28.657+0000"}],"conversations":[{"body":"JDK7's {{Process}} output streams have an async fd-close race bug. This manifests as commands run via o.a.h.u.Shell causing threads to hang, OOM, or cause other bizarre behavior. The NM is likely to encounter the bug under heavy load.\n\nSpecifically, {{ProcessBuilder}}'s {{UNIXProcess}} starts a thread to reap the process and drain stdout/stderr to avoid a lingering zombie process. A race occurs if the thread using the stream closes it, the underlying fd is recycled/reopened, while the reaper is draining it. {{ProcessPipeInputStream.drainInputStream}}'s will OOM allocating an array if {{in.available()}} returns a huge number, or may wreak havoc by incorrectly draining the fd.\n\n","from":"reporter","subject":"Workaround JDK7 Process fd close bug"},{"body":"This patch uses a hack to workaround the bug. Synch'ing on the streams before closing dovetails with the synch'ed {{ProcessPipeInputStream.drainInputStream}}. The hack is a safe no-op on JDK6 because it does not drain the streams.\n\nY! has been using this patch in production for 8 months. The problem was immediately reported to Oracle but a fix will not be available until around mid-year so we're providing this workaround to the community.","from":"developer"},{"body":"No tests are included due to the difficulty of reproducing the bug on demand, and it's not hadoop's responsibility to have unit tests for JDK bugs.","from":"developer"},{"body":"Anyone want to review? After moving to JDK7 in production, we had many NMs under load going OOM and crashing due to this bug. Task retries masked that the cluster was slowly shrinking. As noted above, we've been running production clusters for 8 months with this patch.","from":"developer"},{"body":"bq. ProcessPipeInputStream.drainInputStream's will OOM allocating an array if in.available() returns a huge number, or may wreak havoc by incorrectly draining the fd.\nYou've seen OOMs in practice? Or the closing of wrong streams?\n\nI think I understand the race. Patch changes look good. Great that it's a no-op in JDK 6.\n\nMay be link to the corresponding source in openJDK here for posterity?\n\nAlso, may be you should be put the exact JVM version in the comment - to ease reasoning in the future.","from":"developer"},{"body":"JDK bug has been verified as such by the devs and patched in 7u60: \nhttps://bugs.openjdk.java.net/browse/JDK-8024521","from":"developer"},{"body":"Yes, we were losing ~10s of NMs/day because of OOMs caused by this bug. After the patch, no OOMs.\n\nThe referenced openjdk bug is indeed the same problem.\n\nDo I have a +1 to commit?","from":"developer"},{"body":"Can you add a reference to this JIRA as well as a link to the corresponding (buggy) JVM version in the java comment for posterity? We can revisit and revert this when the time is apt.","from":"developer"},{"body":"Yeah, +1 barring that point about code comment.","from":"developer"},{"body":"Added comments based on info provided by Steve. Will commit when pre-commit passes unless there's objections to the comments.\n\nNote: the original patch filenames had the wrong jira...","from":"developer"},{"body":"The comment is informative enough, the patch looks good to me. +1.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed to trunk, branch-2, and branch-0.23.","from":"developer"}],"created":"2013-12-05T17:25:30.000+0000","description":"JDK7's {{Process}} output streams have an async fd-close race bug. This manifests as commands run via o.a.h.u.Shell causing threads to hang, OOM, or cause other bizarre behavior. The NM is likely to encounter the bug under heavy load.\n\nSpecifically, {{ProcessBuilder}}'s {{UNIXProcess}} starts a thread to reap the process and drain stdout/stderr to avoid a lingering zombie process. A race occurs if the thread using the stream closes it, the underlying fd is recycled/reopened, while the reaper is draining it. {{ProcessPipeInputStream.drainInputStream}}'s will OOM allocating an array if {{in.available()}} returns a huge number, or may wreak havoc by incorrectly draining the fd.\n\n","issue_id":"12682950","key":"HADOOP-10146","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2014-01-16T18:58:28.000+0000","role":"fixed_distractor","summary":"Workaround JDK7 Process fd close bug"} {"case_id":"12686316","cluster":"DISTRACTOR-HADOOP-10181","comments":[{"body":"+1 for this one.. We have 350 nodes so they span two subnets/VLANs, but we forward the Ganglia multicast between them. We just upgraded to CDH5 (Hadoop 2.5) and now we've lost our Ganglia metrics for half of the cluster (because the TTL is set to 1) depending on which half of the cluster you ask (as they will not span the two subnets) \n\n(Given we now have a tonne of missing metrics I would say this is more than Minor, as the alternative is use Unicast Ganglia and that is not ideal for other reasons) ","created":"2015-01-16T18:35:51.156+0000"},{"body":"We've been running the attached patch in our production cluster for a few days now. It has fixed the issue.","created":"2015-01-20T16:20:58.914+0000"},{"body":"I can fix the findbugs warning, but I'm not sure what's up with that test. It passed locally for me.","created":"2015-01-20T17:28:16.251+0000"},{"body":"This fixes the findbugs warning from the previous patch.","created":"2015-01-20T18:10:19.107+0000"},{"body":"I can get the same test timeout locally, but it happens both with and without my patch. I don't think it's related.","created":"2015-01-20T19:02:02.617+0000"},{"body":"Hi, [~ajsquared]. Thank you for providing this patch. This mostly looks good to me. Here are a few minor comments:\n\n# It appears the existing strategy for validation of configuration input in this class is to use a default if the property is unspecified, but let it throw {{NumberFormatException}} if there is a value specified, and it is non-numeric. (For example, see {{getDmax}}.) While perhaps not ideal, it's probably best to keep the validation handling consistent for the new property.\n# It's odd that the existing exception handling code used {{printStackTrace}} even though a logger is available. Even though it's not directly related to your patch, let's take this opportunity to change it to use the {{LOG}}.\n# Thank you for updating package.html too.\n\nThe test failure reported by Jenkins is unrelated. I confirmed that the test passes locally for me.\n\nI also reviewed {{GangliaContext31}} (unchanged in this patch) and it appears the subclass will work just fine with your changes in the base class.\n\nI'll be +1 after points 1 and 2 above are addressed. However, this is my first time looking at this part of the codebase. I'd prefer if we could get one more review by someone who has experience with this code before I commit it. [~tucu00] and [~tlipcon], I noticed your names in the change log. Are you interested in reviewing?","created":"2015-01-26T23:52:29.018+0000"},{"body":"Thanks for the feedback, [~cnauroth]! I've submitted a new patch to address your comments.","created":"2015-01-27T13:03:40.370+0000"},{"body":"+1 for patch v003. Andrew, thank you for addressing the feedback.\n\nI plan to wait until Monday, 2/2, to commit this, in case any committer who has prior experience with this code also wants to review.","created":"2015-01-27T18:07:35.286+0000"},{"body":"I committed this to trunk and branch-2. Andrew, thank you for contributing this patch.","created":"2015-02-02T19:25:20.858+0000"},{"body":"Great, thanks!","created":"2015-02-02T21:53:28.302+0000"},{"body":"Sorry for the late reply, I was on vacation the last few weeks. The patch looks fine to me. Thanks for reviewing and committing, Chris.","created":"2015-02-11T07:31:58.715+0000"},{"body":"[~tlipcon], no worries. Thanks for following up.","created":"2015-02-11T17:41:33.908+0000"}],"conversations":[{"body":"The GangliaContext class which is used to send Hadoop metrics to Ganglia uses a DatagramSocket to send these metrics. This works fine for Ganglia multicast setups that are all on the same VLAN. However, when working with multiple VLANs, a packet sent via DatagramSocket to a multicast address will end up with a TTL of 1. Multicast TTL indicates the number of network hops for which a particular multicast packet is valid. The packets sent by GangliaContext do not make it to ganglia aggregrators on the same multicast group, but in different VLANs.\n\nTo fix, we'd need a configuration property that specifies that multicast is to be used, and another that allows setting of the multicast packet TTL. With these set, we could then use MulticastSocket setTimeToLive() instead of just plain ol' DatagramSocket.\n","from":"reporter","subject":"GangliaContext does not work with multicast ganglia setup"},{"body":"+1 for this one.. We have 350 nodes so they span two subnets/VLANs, but we forward the Ganglia multicast between them. We just upgraded to CDH5 (Hadoop 2.5) and now we've lost our Ganglia metrics for half of the cluster (because the TTL is set to 1) depending on which half of the cluster you ask (as they will not span the two subnets) \n\n(Given we now have a tonne of missing metrics I would say this is more than Minor, as the alternative is use Unicast Ganglia and that is not ideal for other reasons) ","from":"developer"},{"body":"We've been running the attached patch in our production cluster for a few days now. It has fixed the issue.","from":"developer"},{"body":"I can fix the findbugs warning, but I'm not sure what's up with that test. It passed locally for me.","from":"developer"},{"body":"This fixes the findbugs warning from the previous patch.","from":"developer"},{"body":"I can get the same test timeout locally, but it happens both with and without my patch. I don't think it's related.","from":"developer"},{"body":"Hi, [~ajsquared]. Thank you for providing this patch. This mostly looks good to me. Here are a few minor comments:\n\n# It appears the existing strategy for validation of configuration input in this class is to use a default if the property is unspecified, but let it throw {{NumberFormatException}} if there is a value specified, and it is non-numeric. (For example, see {{getDmax}}.) While perhaps not ideal, it's probably best to keep the validation handling consistent for the new property.\n# It's odd that the existing exception handling code used {{printStackTrace}} even though a logger is available. Even though it's not directly related to your patch, let's take this opportunity to change it to use the {{LOG}}.\n# Thank you for updating package.html too.\n\nThe test failure reported by Jenkins is unrelated. I confirmed that the test passes locally for me.\n\nI also reviewed {{GangliaContext31}} (unchanged in this patch) and it appears the subclass will work just fine with your changes in the base class.\n\nI'll be +1 after points 1 and 2 above are addressed. However, this is my first time looking at this part of the codebase. I'd prefer if we could get one more review by someone who has experience with this code before I commit it. [~tucu00] and [~tlipcon], I noticed your names in the change log. Are you interested in reviewing?","from":"developer"},{"body":"Thanks for the feedback, [~cnauroth]! I've submitted a new patch to address your comments.","from":"developer"},{"body":"+1 for patch v003. Andrew, thank you for addressing the feedback.\n\nI plan to wait until Monday, 2/2, to commit this, in case any committer who has prior experience with this code also wants to review.","from":"developer"},{"body":"I committed this to trunk and branch-2. Andrew, thank you for contributing this patch.","from":"developer"},{"body":"Great, thanks!","from":"developer"},{"body":"Sorry for the late reply, I was on vacation the last few weeks. The patch looks fine to me. Thanks for reviewing and committing, Chris.","from":"developer"},{"body":"[~tlipcon], no worries. Thanks for following up.","from":"developer"}],"created":"2013-12-24T18:43:48.000+0000","description":"The GangliaContext class which is used to send Hadoop metrics to Ganglia uses a DatagramSocket to send these metrics. This works fine for Ganglia multicast setups that are all on the same VLAN. However, when working with multiple VLANs, a packet sent via DatagramSocket to a multicast address will end up with a TTL of 1. Multicast TTL indicates the number of network hops for which a particular multicast packet is valid. The packets sent by GangliaContext do not make it to ganglia aggregrators on the same multicast group, but in different VLANs.\n\nTo fix, we'd need a configuration property that specifies that multicast is to be used, and another that allows setting of the multicast packet TTL. With these set, we could then use MulticastSocket setTimeToLive() instead of just plain ol' DatagramSocket.\n","issue_id":"12686316","key":"HADOOP-10181","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2015-02-02T19:25:20.000+0000","role":"fixed_distractor","summary":"GangliaContext does not work with multicast ganglia setup"} {"case_id":"12603075","cluster":"DISTRACTOR-HADOOP-10326","comments":[{"body":"I also met with this issue today. Need to check with the patch any way. [~mdeferran] you want create and attach patch?","created":"2013-05-27T19:06:09.975+0000"},{"body":"With the above mentioned change I am able to disctp data to S3 when security is enabled also.","created":"2013-05-28T09:43:33.591+0000"},{"body":"I encountered this bug while using DistCp to S3. The patch fixed the problem. Thanks!!!","created":"2013-06-03T13:34:22.028+0000"},{"body":"[~mdeferran] - want to take a stab at this, post a patch?","created":"2014-02-03T16:46:02.498+0000"},{"body":"The fix would work, but silently ignores the fact that we won't be using security. May be, we should add a config that users should enable to allow this, just so the users know what they are doing? ","created":"2014-02-04T19:58:46.612+0000"},{"body":"Patch from Manuel:\n\nThis patch seems to fix it.\n{code}\nIndex: core/org/apache/hadoop/security/SecurityUtil.java\n===================================================================\n--- core/org/apache/hadoop/security/SecurityUtil.java (révision 1305278)\n+++ core/org/apache/hadoop/security/SecurityUtil.java (copie de travail)\n@@ -313,6 +313,9 @@\n if (authority == null || authority.isEmpty()) {\n return null;\n }\n+ if (uri.getScheme().equals(\"s3n\") || uri.getScheme().equals(\"s3\")) {\n+ return null;\n+ }\n InetSocketAddress addr = NetUtils.createSocketAddr(authority, defPort);\n return buildTokenService(addr).toString();\n }\n{code}","created":"2014-02-05T18:49:02.468+0000"},{"body":"Spoke to [~tucu00] and [~atm] offline. Both believe, YARN/MR and SecurityUtil should be agnostic to how the underlying filesystem handles tokens. The S3 client should handle ignoring these tokens. Moved to a Common JIRA to address that.","created":"2014-02-05T18:50:25.599+0000"},{"body":"Adding patch \"0001-HADOOP-10326.-s3-s3n-does-not-support-tokens.patch\". Tested on secure cluster, running distcp and wordcount against s3 data.","created":"2014-02-08T18:05:33.305+0000"},{"body":"Patch looks good to me. +1 pending Jenkins.\n\nFor some reason I can't seem to add bc as a Hadoop contributor at the moment. [~tucu00] - would you mind taking care of that in the JIRA admin console?","created":"2014-02-11T00:43:12.991+0000"},{"body":"JIRA console funny, getting \"The JIRA server could not be contacted. This may be a temporary glitch or the server may be down.\" pop up when trying to add him","created":"2014-02-11T00:48:10.454+0000"},{"body":"Yea, me too. Curious, but not a big deal. Can fix that up later.","created":"2014-02-11T00:49:23.871+0000"},{"body":"I've just committed this to trunk and branch-2.\n\nThanks a lot for the contribution, bc.","created":"2014-02-11T02:49:34.414+0000"}],"conversations":[{"body":"With Kerberos enabled, any job that is taking as input or output s3 files fails.\n\nIt can be easily reproduced with wordcount shipped in hadoop-examples.jar and a public S3 file:\n{code}\n/opt/hadoop/bin/hadoop --config /opt/hadoop/conf/ jar /opt/hadoop/hadoop-examples-1.0.0.jar wordcount s3n://ubikodpublic/test out01\n{code}\n\nreturns:\n{code}\n12/08/10 12:40:19 INFO hdfs.DFSClient: Created HDFS_DELEGATION_TOKEN token 192 for hadoop on 10.85.151.233:9000\n12/08/10 12:40:19 INFO security.TokenCache: Got dt for hdfs://aws04.machine.com:9000/mapred/staging/hadoop/.staging/job_201208101229_0004;uri=10.85.151.233:9000;t.service=10.85.151.233:9000\n12/08/10 12:40:19 INFO mapred.JobClient: Cleaning up the staging area hdfs://aws04.machine.com:9000/mapred/staging/hadoop/.staging/job_201208101229_0004\njava.lang.IllegalArgumentException: java.net.UnknownHostException: ubikodpublic\n at org.apache.hadoop.security.SecurityUtil.buildTokenService(SecurityUtil.java:293)\n at org.apache.hadoop.security.SecurityUtil.buildDTServiceName(SecurityUtil.java:317)\n at org.apache.hadoop.fs.FileSystem.getCanonicalServiceName(FileSystem.java:189)\n at org.apache.hadoop.mapreduce.security.TokenCache.obtainTokensForNamenodesInternal(TokenCache.java:92)\n at org.apache.hadoop.mapreduce.security.TokenCache.obtainTokensForNamenodes(TokenCache.java:79)\n at org.apache.hadoop.mapreduce.lib.input.FileInputFormat.listStatus(FileInputFormat.java:197)\n at org.apache.hadoop.mapreduce.lib.input.FileInputFormat.getSplits(FileInputFormat.java:252)\n\n{code}\n\n","from":"reporter","subject":"M/R jobs can not access S3 if Kerberos is enabled"},{"body":"I also met with this issue today. Need to check with the patch any way. [~mdeferran] you want create and attach patch?","from":"developer"},{"body":"With the above mentioned change I am able to disctp data to S3 when security is enabled also.","from":"developer"},{"body":"I encountered this bug while using DistCp to S3. The patch fixed the problem. Thanks!!!","from":"developer"},{"body":"[~mdeferran] - want to take a stab at this, post a patch?","from":"developer"},{"body":"The fix would work, but silently ignores the fact that we won't be using security. May be, we should add a config that users should enable to allow this, just so the users know what they are doing? ","from":"developer"},{"body":"Patch from Manuel:\n\nThis patch seems to fix it.\n{code}\nIndex: core/org/apache/hadoop/security/SecurityUtil.java\n===================================================================\n--- core/org/apache/hadoop/security/SecurityUtil.java (révision 1305278)\n+++ core/org/apache/hadoop/security/SecurityUtil.java (copie de travail)\n@@ -313,6 +313,9 @@\n if (authority == null || authority.isEmpty()) {\n return null;\n }\n+ if (uri.getScheme().equals(\"s3n\") || uri.getScheme().equals(\"s3\")) {\n+ return null;\n+ }\n InetSocketAddress addr = NetUtils.createSocketAddr(authority, defPort);\n return buildTokenService(addr).toString();\n }\n{code}","from":"developer"},{"body":"Spoke to [~tucu00] and [~atm] offline. Both believe, YARN/MR and SecurityUtil should be agnostic to how the underlying filesystem handles tokens. The S3 client should handle ignoring these tokens. Moved to a Common JIRA to address that.","from":"developer"},{"body":"Adding patch \"0001-HADOOP-10326.-s3-s3n-does-not-support-tokens.patch\". Tested on secure cluster, running distcp and wordcount against s3 data.","from":"developer"},{"body":"Patch looks good to me. +1 pending Jenkins.\n\nFor some reason I can't seem to add bc as a Hadoop contributor at the moment. [~tucu00] - would you mind taking care of that in the JIRA admin console?","from":"developer"},{"body":"JIRA console funny, getting \"The JIRA server could not be contacted. This may be a temporary glitch or the server may be down.\" pop up when trying to add him","from":"developer"},{"body":"Yea, me too. Curious, but not a big deal. Can fix that up later.","from":"developer"},{"body":"I've just committed this to trunk and branch-2.\n\nThanks a lot for the contribution, bc.","from":"developer"}],"created":"2012-08-10T13:49:57.000+0000","description":"With Kerberos enabled, any job that is taking as input or output s3 files fails.\n\nIt can be easily reproduced with wordcount shipped in hadoop-examples.jar and a public S3 file:\n{code}\n/opt/hadoop/bin/hadoop --config /opt/hadoop/conf/ jar /opt/hadoop/hadoop-examples-1.0.0.jar wordcount s3n://ubikodpublic/test out01\n{code}\n\nreturns:\n{code}\n12/08/10 12:40:19 INFO hdfs.DFSClient: Created HDFS_DELEGATION_TOKEN token 192 for hadoop on 10.85.151.233:9000\n12/08/10 12:40:19 INFO security.TokenCache: Got dt for hdfs://aws04.machine.com:9000/mapred/staging/hadoop/.staging/job_201208101229_0004;uri=10.85.151.233:9000;t.service=10.85.151.233:9000\n12/08/10 12:40:19 INFO mapred.JobClient: Cleaning up the staging area hdfs://aws04.machine.com:9000/mapred/staging/hadoop/.staging/job_201208101229_0004\njava.lang.IllegalArgumentException: java.net.UnknownHostException: ubikodpublic\n at org.apache.hadoop.security.SecurityUtil.buildTokenService(SecurityUtil.java:293)\n at org.apache.hadoop.security.SecurityUtil.buildDTServiceName(SecurityUtil.java:317)\n at org.apache.hadoop.fs.FileSystem.getCanonicalServiceName(FileSystem.java:189)\n at org.apache.hadoop.mapreduce.security.TokenCache.obtainTokensForNamenodesInternal(TokenCache.java:92)\n at org.apache.hadoop.mapreduce.security.TokenCache.obtainTokensForNamenodes(TokenCache.java:79)\n at org.apache.hadoop.mapreduce.lib.input.FileInputFormat.listStatus(FileInputFormat.java:197)\n at org.apache.hadoop.mapreduce.lib.input.FileInputFormat.getSplits(FileInputFormat.java:252)\n\n{code}\n\n","issue_id":"12603075","key":"HADOOP-10326","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2014-02-11T02:49:34.000+0000","role":"fixed_distractor","summary":"M/R jobs can not access S3 if Kerberos is enabled"} {"case_id":"12716922","cluster":"DISTRACTOR-HADOOP-10630","comments":[{"body":"A possible fix is to refresh \"currentProxy\" no matter if the failover is performed by the current thread or not. Since the process is protected by proxyProvider's lock, all the threads should be able to get the current value of currentProxy.","created":"2014-05-27T22:16:25.289+0000"},{"body":"The patch looks reasonable. Did you have a chance to verify that it fixes the issue?","created":"2014-05-28T17:33:42.253+0000"},{"body":"Not yet. Actually the issue cannot easily be reproduced since by default the client will retry/failover 10 times. I will decrease the retry number and rerun the test with/without the patch these days.","created":"2014-05-28T17:41:38.951+0000"},{"body":"+1 for the patch, once the failover tests pass.","created":"2014-05-28T17:57:41.696+0000"},{"body":"Looks like our failover tests run well with the patch during the weekend. I will commit the patch shortly.","created":"2014-06-02T21:29:05.900+0000"},{"body":"I've committed this to trunk and branch-2. Thanks for the review, [~kihwal] and [~sureshms]! And thanks [~arpitgupta] for reporting the issue.","created":"2014-06-02T21:45:56.531+0000"}],"conversations":[{"body":"In one of our system tests with NameNode HA setup, we ran 300 threads in LoadGenerator. While one of the NameNodes was already in the active state and started to serve, we still saw one of the client thread failed all the retries in a 20 seconds window. In the meanwhile, we saw a lot of following warning msg in the log:\n{noformat}\nWARN retry.RetryInvocationHandler: A failover has occurred since the start of this method invocation attempt.\n{noformat}\n\nAfter checking the code, we see the following code in RetryInvocationHandler:\n{code}\n while (true) {\n // The number of times this invocation handler has ever been failed over,\n // before this method invocation attempt. Used to prevent concurrent\n // failed method invocations from triggering multiple failover attempts.\n long invocationAttemptFailoverCount;\n synchronized (proxyProvider) {\n invocationAttemptFailoverCount = proxyProviderFailoverCount;\n }\n ......\n if (action.action == RetryAction.RetryDecision.FAILOVER_AND_RETRY) {\n // Make sure that concurrent failed method invocations only cause a\n // single actual fail over.\n synchronized (proxyProvider) {\n if (invocationAttemptFailoverCount == proxyProviderFailoverCount) {\n proxyProvider.performFailover(currentProxy.proxy);\n proxyProviderFailoverCount++;\n currentProxy = proxyProvider.getProxy();\n } else {\n LOG.warn(\"A failover has occurred since the start of this method\"\n + \" invocation attempt.\");\n }\n }\n invocationFailoverCount++;\n }\n ......\n{code}\n\nWe can see we refresh the value of currentProxy only when the thread performs the failover (while holding the monitor of the proxyProvider). Because \"currentProxy\" is not volatile, a thread that does not perform the failover (in which case it will log the warning msg) may fail to get the new value of currentProxy.","from":"reporter","subject":"Possible race condition in RetryInvocationHandler"},{"body":"A possible fix is to refresh \"currentProxy\" no matter if the failover is performed by the current thread or not. Since the process is protected by proxyProvider's lock, all the threads should be able to get the current value of currentProxy.","from":"developer"},{"body":"The patch looks reasonable. Did you have a chance to verify that it fixes the issue?","from":"developer"},{"body":"Not yet. Actually the issue cannot easily be reproduced since by default the client will retry/failover 10 times. I will decrease the retry number and rerun the test with/without the patch these days.","from":"developer"},{"body":"+1 for the patch, once the failover tests pass.","from":"developer"},{"body":"Looks like our failover tests run well with the patch during the weekend. I will commit the patch shortly.","from":"developer"},{"body":"I've committed this to trunk and branch-2. Thanks for the review, [~kihwal] and [~sureshms]! And thanks [~arpitgupta] for reporting the issue.","from":"developer"}],"created":"2014-05-27T22:11:05.000+0000","description":"In one of our system tests with NameNode HA setup, we ran 300 threads in LoadGenerator. While one of the NameNodes was already in the active state and started to serve, we still saw one of the client thread failed all the retries in a 20 seconds window. In the meanwhile, we saw a lot of following warning msg in the log:\n{noformat}\nWARN retry.RetryInvocationHandler: A failover has occurred since the start of this method invocation attempt.\n{noformat}\n\nAfter checking the code, we see the following code in RetryInvocationHandler:\n{code}\n while (true) {\n // The number of times this invocation handler has ever been failed over,\n // before this method invocation attempt. Used to prevent concurrent\n // failed method invocations from triggering multiple failover attempts.\n long invocationAttemptFailoverCount;\n synchronized (proxyProvider) {\n invocationAttemptFailoverCount = proxyProviderFailoverCount;\n }\n ......\n if (action.action == RetryAction.RetryDecision.FAILOVER_AND_RETRY) {\n // Make sure that concurrent failed method invocations only cause a\n // single actual fail over.\n synchronized (proxyProvider) {\n if (invocationAttemptFailoverCount == proxyProviderFailoverCount) {\n proxyProvider.performFailover(currentProxy.proxy);\n proxyProviderFailoverCount++;\n currentProxy = proxyProvider.getProxy();\n } else {\n LOG.warn(\"A failover has occurred since the start of this method\"\n + \" invocation attempt.\");\n }\n }\n invocationFailoverCount++;\n }\n ......\n{code}\n\nWe can see we refresh the value of currentProxy only when the thread performs the failover (while holding the monitor of the proxyProvider). Because \"currentProxy\" is not volatile, a thread that does not perform the failover (in which case it will log the warning msg) may fail to get the new value of currentProxy.","issue_id":"12716922","key":"HADOOP-10630","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2014-06-02T21:45:56.000+0000","role":"fixed_distractor","summary":"Possible race condition in RetryInvocationHandler"} {"case_id":"12720158","cluster":"DISTRACTOR-HADOOP-10673","comments":[{"body":"The patch updates the existing rpc metrics in the case of exception. HDFS unit tests will be updated to cover this.","created":"2014-06-09T18:08:30.443+0000"},{"body":"[~mingma], I think your suggestion of maintaining dedicated metrics for exceptions is useful to measure \"wasted\" calls.","created":"2014-06-10T05:53:53.659+0000"},{"body":"Thanks, Gera. Here is the updated patch. For RpcMetrics, it doesn't distinguish good call from waste call. For RpcDetailedMetrics, if the call throws exception, it will use $MethodName_$ExceptionClassSimpleName as the metrics name.","created":"2014-06-13T00:54:15.901+0000"},{"body":"Updates with unit test.","created":"2014-06-13T21:38:30.344+0000"},{"body":"[~mingma], thank you for the suggestion. I think this improvement is useful, too. One minor comment: both {{rpcMetrics}} and {{rpcDetailedMetrics}} are not final field, so how about adding null check before accessing them?\n\n{code}\n+ server.rpcMetrics.addRpcQueueTime(qTime);\n+ server.rpcMetrics.addRpcProcessingTime(processingTime);\n+ server.rpcDetailedMetrics.addProcessingTime(detailedMetricsName,\n+ processingTime);\n{code}\n\n{code}\n protected RpcMetrics rpcMetrics;\n protected RpcDetailedMetrics rpcDetailedMetrics;\n{code}","created":"2014-06-13T22:56:36.224+0000"},{"body":"Thanks, Tsuyoshi.\n\nHere is the updated patch that makes the variables final.","created":"2014-06-14T02:09:33.025+0000"},{"body":"Thanks for the updating, Ming. It looks better approach to me. One additional point: if we can make the variables final, we can remove null check in {{Server#stop()}}.\n\n{code}\n if (this.rpcMetrics != null) {\n this.rpcMetrics.shutdown();\n }\n if (this.rpcDetailedMetrics != null) {\n this.rpcDetailedMetrics.shutdown();\n }\n{code}","created":"2014-06-14T23:11:22.828+0000"},{"body":"Thanks, Tsuyoshi. [~lohit] asked a good question how the increase of the number of metrics will impact metrics subsystem. The most common one might be StandbyException, which can be thrown regardless of methods. Other exceptions are likely to be specific to certain methods. Based on the current number of metrics, JMX size; number of methods, a very rough estimate can put the increase to around 10-15% for NN.\n\nSo to be safe, we can aggregate by exception class name instead; that will still provide useful information (from the exception name, we can find the possible methods).\n\nHere is the updated patch.","created":"2014-06-16T23:24:28.470+0000"},{"body":"The patch looks good to me. +1.","created":"2014-07-15T22:59:04.387+0000"},{"body":"I've committed this to trunk and branch-2. Thanks for the contribution, [~mingma]! Thanks for the review, [~ozawa]!","created":"2014-07-15T23:15:40.487+0000"}],"conversations":[{"body":"Currently RPC metrics isn't updated when the call throws an exception. We can either update the existing metrics or have a new set of metrics in the case of exception.","from":"reporter","subject":"Update rpc metrics when the call throws an exception"},{"body":"The patch updates the existing rpc metrics in the case of exception. HDFS unit tests will be updated to cover this.","from":"developer"},{"body":"[~mingma], I think your suggestion of maintaining dedicated metrics for exceptions is useful to measure \"wasted\" calls.","from":"developer"},{"body":"Thanks, Gera. Here is the updated patch. For RpcMetrics, it doesn't distinguish good call from waste call. For RpcDetailedMetrics, if the call throws exception, it will use $MethodName_$ExceptionClassSimpleName as the metrics name.","from":"developer"},{"body":"Updates with unit test.","from":"developer"},{"body":"[~mingma], thank you for the suggestion. I think this improvement is useful, too. One minor comment: both {{rpcMetrics}} and {{rpcDetailedMetrics}} are not final field, so how about adding null check before accessing them?\n\n{code}\n+ server.rpcMetrics.addRpcQueueTime(qTime);\n+ server.rpcMetrics.addRpcProcessingTime(processingTime);\n+ server.rpcDetailedMetrics.addProcessingTime(detailedMetricsName,\n+ processingTime);\n{code}\n\n{code}\n protected RpcMetrics rpcMetrics;\n protected RpcDetailedMetrics rpcDetailedMetrics;\n{code}","from":"developer"},{"body":"Thanks, Tsuyoshi.\n\nHere is the updated patch that makes the variables final.","from":"developer"},{"body":"Thanks for the updating, Ming. It looks better approach to me. One additional point: if we can make the variables final, we can remove null check in {{Server#stop()}}.\n\n{code}\n if (this.rpcMetrics != null) {\n this.rpcMetrics.shutdown();\n }\n if (this.rpcDetailedMetrics != null) {\n this.rpcDetailedMetrics.shutdown();\n }\n{code}","from":"developer"},{"body":"Thanks, Tsuyoshi. [~lohit] asked a good question how the increase of the number of metrics will impact metrics subsystem. The most common one might be StandbyException, which can be thrown regardless of methods. Other exceptions are likely to be specific to certain methods. Based on the current number of metrics, JMX size; number of methods, a very rough estimate can put the increase to around 10-15% for NN.\n\nSo to be safe, we can aggregate by exception class name instead; that will still provide useful information (from the exception name, we can find the possible methods).\n\nHere is the updated patch.","from":"developer"},{"body":"The patch looks good to me. +1.","from":"developer"},{"body":"I've committed this to trunk and branch-2. Thanks for the contribution, [~mingma]! Thanks for the review, [~ozawa]!","from":"developer"}],"created":"2014-06-09T18:01:13.000+0000","description":"Currently RPC metrics isn't updated when the call throws an exception. We can either update the existing metrics or have a new set of metrics in the case of exception.","issue_id":"12720158","key":"HADOOP-10673","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2014-07-15T23:15:40.000+0000","role":"fixed_distractor","summary":"Update rpc metrics when the call throws an exception"} {"case_id":"12725544","cluster":"DISTRACTOR-HADOOP-10786","comments":[{"body":"Not sure why build is failing, there are no compile errors in the log. Can't reproduce build failure locally - build is failing (flaky?) even on trunk.","created":"2014-07-07T16:56:28.491+0000"},{"body":"[~Tobi], thanks for sharing this patch. We encountered the same problem on a Hadoop 2.3.0 cluster and were able to resolve it with this change.","created":"2014-09-13T04:01:16.062+0000"},{"body":"# what does this do on Java 6?\n# It may actually be possible to write a test for this with MiniKDC; this is clearly something that's not being tested for today. Adding that test would would ensure we don't regress again","created":"2014-09-13T11:14:19.866+0000"},{"body":"keytab is tagged as {{ * @ since 1.7 }}, so can't go in while Hadoop is still built against Java 6.","created":"2014-09-13T11:18:51.586+0000"},{"body":"Shouldn't this be a higher priority than 'Minor'? The end of public updates to Java 7 will be April 2015. A silent failure to re-login from keytab after TGT expiry dooms any long running process that wants to use secure RPC. Anyone who cares about security and about running the best performing supported Java runtime shortly will be forced to locally patch their core libraries. ","created":"2014-09-17T15:54:56.176+0000"},{"body":"We can use reflection to make this fix and still allow JDK6 to build and run.\n\nI've attached a patch to do this, as well as added a unit test that will catch regressions. The unit test uses the MiniKDC and verifies login from keytab and relogin from keytab in addition to simply checking that isKeytab = true when it should be.\n\n[~Tobi], thanks a lot for working on this. Let me know what you think about my suggestion and test. If you are too busy, I can also take this JIRA up.","created":"2014-11-06T01:42:21.853+0000"},{"body":"Note that the test reproduces the issue when run against JDK8 on a build without the fix.\n\nI then built and ran successfully with the fix for JDK 6, 7, and 8.","created":"2014-11-06T01:43:56.611+0000"},{"body":"That's one clean patch.","created":"2014-11-06T04:28:45.155+0000"},{"body":"Thank you [~schu], that's a great patch. What else would need to be done to land this?","created":"2014-11-06T22:30:07.117+0000"},{"body":"Thanks, Tobi. We just need to wait for reviews and make any iterations that come from the reviews. When a committer +1's, then the approved patch can be committed.","created":"2014-11-06T22:37:23.104+0000"},{"body":"{code}\n+ try {\n+ // In JDK6 and JDK7, if useKeyTab and storeKey are specified in the\n+ // Krb5LoginModule, then some number of KerberosKey objects are added\n+ // to the Subject's private credentials. However, in JDK8, a KeyTab\n+ // object is added instead. More details in HADOOP-10786.\n+ keytabClass = Class.forName(\"javax.security.auth.kerberos.KeyTab\");\n+ } catch (ClassNotFoundException cnfe) {\n+ // Ignore. javax.security.auth.kerberos.KeyTab does not exist in JDK6.\n+ }\n+ if (keytabClass != null) {\n+ this.isKeytab = !subject.getPrivateCredentials(keytabClass).isEmpty();\n+ } else {\n+ this.isKeytab = !subject.getPrivateCredentials(KerberosKey.class).isEmpty();\n+ }\n this.isKrbTkt = !subject.getPrivateCredentials(KerberosTicket.class).isEmpty();\n{code}\n\n{{forName}} is fairly slow. Since the patch is targeting 2.7 which only supports JDK7, the code should be able to use the class in compile time, though we'll need to wait until jenkins to be switched to Java 7 before this patch can land.\n\n{code}\n+ @VisibleForTesting\n+ static void setShouldRenewImmediatelyForTests(boolean immediate) {\n+ shouldRenewImmediatelyForTests = immediate;\n+ }\n{code}\n\nInstead of adding this method, it might make more sense to extract the logic of login into a separate function and call the function directly in the tests.\n","created":"2014-11-07T00:15:06.012+0000"},{"body":"bq. forName is fairly slow. Since the patch is targeting 2.7 which only supports JDK7, the code should be able to use the class in compile time, though we'll need to wait until jenkins to be switched to Java 7 before this patch can land.\n\nI'd like to commit this for 2.6, though, which will still be targeting Java 6. How about we create a constant {{KEY_TAB_CLASS}} and then do the reflection in a static initialization block? That way we only pay the lookup penalty once per JVM and the patch still works with both version of Java.","created":"2014-11-07T01:31:37.897+0000"},{"body":"Performing the reflection in a static init block sounds like a good idea.\n\nI can see how it'd be useful to extract the logic of login into a separate function and just call it directly. I'd like to make sure to exercise as much of the reloginFromKeytab logic as possible (aside from waiting for a renew window), though.\n\nThe test verifies isKeytab == true, which is good. However, if for some reason the way isKeytab changes in reloginFromKeytab (or something else changes before actual login), it'd be good to exercise this.\n\nAttaching a patch that moves the reflection to a static block.\n\nAlso, I made some additional fixes:\n\n* Fix the conditional logic when using shouldRenewImmediatelyForTests by moving the check for null TGT ahead.\n* Remove //return","created":"2014-11-07T02:33:55.574+0000"},{"body":"Resubmitting HADOOP-10786.3.patch to hopefully retrigger Hadoop QA jenkins.","created":"2014-11-07T19:24:14.282+0000"},{"body":"TestMetricsSystemImpl is unrelated. I ran the test locally and successfully with the patch. Outstanding JIRA for it at HADOOP-10062.","created":"2014-11-07T20:21:51.548+0000"},{"body":"Posting a new patch to fix a mistake. Thanks to ATM for catching it.\n\nIf the following conditional is true, we should return:\n{code}\nif (tgt != null && !shouldRenewImmediatelyForTests &&\n now < getRefreshTime(tgt)) {\n{code}\n\nI had forgotten to clean up some debugging stuff.","created":"2014-11-07T20:29:32.877+0000"},{"body":"Oops, wrong patch. Cancelling and will resubmit.","created":"2014-11-07T20:30:14.429+0000"},{"body":"Note that the correct patch is HADOOP-10786.4.patch.","created":"2014-11-07T20:31:24.346+0000"},{"body":"Thanks a lot, Stephen. The latest patch looks good to me. I'll be +1 on it pending Jenkins.\n\nHaohui - just checking, has this latest patch addressed your performance concern?\n\nThanks folks.","created":"2014-11-07T20:49:49.894+0000"},{"body":"Looks pretty good to me. Some nits:\n\n{code}\n+ private static Class KEY_TAB_CLASS = KerberosKey.class;\n{code}\n\ncan be \n\n{code}\n+ private static final Class KEY_TAB_CLASS = KerberosKey.class;\n{code}\n\n{code}\n+ public void createTestDir() {\n+ workDir = new File(System.getProperty(\"test.dir\", \"target\"));\n+ }\n{code}\n\nYou can use {{TemporaryFolder}} here so that the files can be properly cleaned up after the tests.\n\n\n{code}\n+ public File getWorkDir() {\n+ return workDir;\n+ }\n+ public MiniKdc getKdc() {\n+ return kdc;\n+ }\n+\n{code}\n\nLooks like the test can simply inline these getters to make the patch even smaller.","created":"2014-11-07T20:54:59.422+0000"},{"body":"Thanks for the comments, Haohui. Attaching patch to address your comments.\n\nKEY_TAB_CLASS will be reassigned to Class.forName(\"javax.security.auth.kerberos.KeyTab\") on JDK7/8, so I couldn't declare it final.\n\nThe new patch now uses TemporaryFolder to ensure cleanup of the dir after test run.\n\nAlso removed the unnecessary getters.","created":"2014-11-07T21:30:00.949+0000"},{"body":"Oh I see. The patch looks good to me. +1 pending jenkins.","created":"2014-11-07T21:36:55.253+0000"},{"body":"I've committed the patch to trunk, branch-2 and branch-2.6. Thanks [~schu] for the contribution.","created":"2014-11-10T01:49:18.513+0000"},{"body":"I created HADOOP-10287 to track the effort of simplifying the code in 2.7.","created":"2014-11-10T01:52:12.415+0000"},{"body":"I think it's too late/risky to put this into 2.6; let's get this into 2.7. Thanks.","created":"2014-11-10T02:06:17.756+0000"},{"body":"Moved to hadoop-2.7.0.","created":"2014-11-10T02:20:38.177+0000"},{"body":"[~schu] and [~Tobi], thank you for providing this patch.\n\nI just wanted to share with everyone that even though this bug was reported against a JDK 8 code change, it appears the same change has entered the JDK 7 code line. I am seeing the same problem in the most recent OpenJDK build. With JDK 1.7.0_79, I could not repro the problem. After upgrading to JDK 1.7.0_85, I could repro the problem. I don't know the exact minor version number within JDK 7 that first introduced the change, but it's somewhere in that range.\n\nI also confirmed that this patch fixes the problem for JDK 1.7.0_85 too. To verify, I ran the test without the corresponding fixes in {{UserGroupInformation}}. I observed that the test failed on the assertion for {{ugi.isFromKeytab()}}. Then, I applied the {{UserGroupInformation}} part of the patch, and the test passed.\n\nBottom line: If you want to run a secured Hadoop cluster on JDK 1.7.0_85 or later, then you must run Apache Hadoop 2.7.0 or later.","created":"2015-08-12T21:14:49.845+0000"},{"body":"Cherry-picked to 2.6.1","created":"2015-08-14T06:56:21.920+0000"},{"body":"Thank you, [~vinayrpet].","created":"2015-08-14T15:41:03.153+0000"},{"body":"[~vinayrpet], please stop committing to 2.6.1 (branch-2.6 or branch-2.6.1), we have a fairly elaborate parallel release-process going on for 2.6.1 and these cherry-picks are disrupting our progress. If you want something added to 2.6.1, please post it in the mailing lists. Thanks. ","created":"2015-08-14T19:07:07.328+0000"},{"body":"bq. we have a fairly elaborate parallel release-process going on for 2.6.1 and these cherry-picks are disrupting our progress.\nThanks [~vinodkv]. I am actually cherry-picking only those which are listed for 2.6.1 here https://wiki.apache.org/hadoop/Release-2.6.1-Working-Notes.\nAnyway I will stop further merges until required.\n-Thanks","created":"2015-08-17T05:43:58.188+0000"},{"body":"This wasn't originally in 2.6.1, must have been committed to 2.6, which was already 2.6.2. I just committed this to 2.6.1. Ran compilation and TestUGILoginFromKeytab before the push.","created":"2015-08-28T01:26:02.425+0000"},{"body":"Looks like the Krb5LoginModule change was backported to [JDK 1.7.0_80|http://bugs.java.com/view_bug.do?bug_id=8043932] as well, in case anyone else hits it.\n\n","created":"2015-11-21T00:41:35.192+0000"},{"body":"[~taoluo]\nThanks for your reminding~ We hit the problem in JDK 1.7.0_80. :(","created":"2016-01-14T02:48:29.957+0000"},{"body":"Hey guys,\r\n\r\nI don't see this patch (HADOOP-10786.5.patch), especially this piece of code\r\n{code:java}\r\n+ private static Class KEY_TAB_CLASS = KerberosKey.class;\r\n+ static {\r\n+ try {\r\n+ // We use KEY_TAB_CLASS to determine if the UGI is logged in from\r\n+ // keytab. In JDK6 and JDK7, if useKeyTab and storeKey are specified\r\n+ // in the Krb5LoginModule, then some number of KerberosKey objects\r\n+ // are added to the Subject's private credentials. However, in JDK8,\r\n+ // a KeyTab object is added instead. More details in HADOOP-10786.\r\n+ KEY_TAB_CLASS = Class.forName(\"javax.security.auth.kerberos.KeyTab\");\r\n+ } catch (ClassNotFoundException cnfe) {\r\n+ // Ignore. javax.security.auth.kerberos.KeyTab does not exist in JDK6.\r\n+ }\r\n+ }\r\n+\r\n\r\n- this.isKeytab = !subject.getPrivateCredentials(KerberosKey.class).isEmpty();\r\n+ this.isKeytab = !subject.getPrivateCredentials(KEY_TAB_CLASS).isEmpty();\r\n\r\n{code}\r\napplied in UserGroupInformation class of Apache Hadoop 2.7.0 or 2.7.1 or in any of the subsequent versions till 3.0.0. I'm not sure if above piece of code has been changed/moved somewhere in further/future commits.\r\n\r\n I've a Kerberized Hadoop cluster that uses 2.7.3 and a client that uses 2.7.0, my client interacts with the cluster only using HDFS FileSystem API calls and I expect Kerberos ticket renewable to be automatically handled by the client's runtime dependency. Initially login from keytab is always successful but once the TGT expires, client is unable to login again, it fails with this \r\n\r\n\r\n{noformat}\r\njava.io.IOException: Failed on local exception: java.io.IOException: javax.security.sasl.SaslException: GSS initiate failed [Caused by GSSException: No valid credentials provided (Mechanism level: Failed to find any Kerberos tgt)]; Host Details : local host is: \"hc4t03283/16.202.4.11\"; destination host is: \"hc4t02044.itcs.nircorp.net\":8020; {noformat}\r\nHow can I ascertain or at least rule out that this issue isn't caused because of HADOOP-10786? This is the very reason I looked at the UserGroupInformation class' code and I didn't find this patch applied. I'm using JDK 1.8.0_151. \r\n\r\nPlease let me know if I'm missing something obvious here. ","created":"2018-03-06T16:02:21.280+0000"},{"body":"Looks like it's in branch-2.7\r\n{code}\r\n> git log --grep HADOOP-10786 branch-2.7\r\ncommit 8f4a09b6076de9fbd6cd8ccaddf72ba9c94429ff\r\nAuthor: Vinayakumar B \r\nDate: Fri Aug 14 12:23:51 2015 +0530\r\n\r\n HADOOP-10786. Fix UGI#reloginFromKeytab on Java 8. Contributed by Stephen Chu.\r\n Moved CHANGES.txt entry to 2.6.1\r\n \r\n (cherry picked from commit e7aa81394dce61cc96d480e21204263a5f2ed153)\r\n{code}\r\n\r\nThe code has moved on a lot since that patch went in, which is why there's no match. ","created":"2018-03-07T14:33:27.228+0000"},{"body":"[~Tobi] [~schu] is this also an issue for OpenJDK 1.8?","created":"2019-01-31T02:39:10.029+0000"}],"conversations":[{"body":"Krb5LoginModule changed subtly in java 8: in particular, if useKeyTab and storeKey are specified, then only a KeyTab object is added to the Subject's private credentials, whereas in java <= 7 both a KeyTab and some number of KerberosKey objects were added.\n\nThe UGI constructor checks whether or not a keytab was used to login by looking if there are any KerberosKey objects in the Subject's private credentials. If there are, then isKeyTab is set to true, and otherwise it's set to false.\n\nThus, in java 8 isKeyTab is always false given the current UGI implementation, which makes UGI#reloginFromKeytab fail silently.\n\nAttached patch will check for a KeyTab object on the Subject, instead of a KerberosKey object. This fixes relogins from kerberos keytabs on Oracle java 8, and works on Oracle java 7 as well.","from":"reporter","subject":"Fix UGI#reloginFromKeytab on Java 8"},{"body":"Not sure why build is failing, there are no compile errors in the log. Can't reproduce build failure locally - build is failing (flaky?) even on trunk.","from":"developer"},{"body":"[~Tobi], thanks for sharing this patch. We encountered the same problem on a Hadoop 2.3.0 cluster and were able to resolve it with this change.","from":"developer"},{"body":"# what does this do on Java 6?\n# It may actually be possible to write a test for this with MiniKDC; this is clearly something that's not being tested for today. Adding that test would would ensure we don't regress again","from":"developer"},{"body":"keytab is tagged as {{ * @ since 1.7 }}, so can't go in while Hadoop is still built against Java 6.","from":"developer"},{"body":"Shouldn't this be a higher priority than 'Minor'? The end of public updates to Java 7 will be April 2015. A silent failure to re-login from keytab after TGT expiry dooms any long running process that wants to use secure RPC. Anyone who cares about security and about running the best performing supported Java runtime shortly will be forced to locally patch their core libraries. ","from":"developer"},{"body":"We can use reflection to make this fix and still allow JDK6 to build and run.\n\nI've attached a patch to do this, as well as added a unit test that will catch regressions. The unit test uses the MiniKDC and verifies login from keytab and relogin from keytab in addition to simply checking that isKeytab = true when it should be.\n\n[~Tobi], thanks a lot for working on this. Let me know what you think about my suggestion and test. If you are too busy, I can also take this JIRA up.","from":"developer"},{"body":"Note that the test reproduces the issue when run against JDK8 on a build without the fix.\n\nI then built and ran successfully with the fix for JDK 6, 7, and 8.","from":"developer"},{"body":"That's one clean patch.","from":"developer"},{"body":"Thank you [~schu], that's a great patch. What else would need to be done to land this?","from":"developer"},{"body":"Thanks, Tobi. We just need to wait for reviews and make any iterations that come from the reviews. When a committer +1's, then the approved patch can be committed.","from":"developer"},{"body":"{code}\n+ try {\n+ // In JDK6 and JDK7, if useKeyTab and storeKey are specified in the\n+ // Krb5LoginModule, then some number of KerberosKey objects are added\n+ // to the Subject's private credentials. However, in JDK8, a KeyTab\n+ // object is added instead. More details in HADOOP-10786.\n+ keytabClass = Class.forName(\"javax.security.auth.kerberos.KeyTab\");\n+ } catch (ClassNotFoundException cnfe) {\n+ // Ignore. javax.security.auth.kerberos.KeyTab does not exist in JDK6.\n+ }\n+ if (keytabClass != null) {\n+ this.isKeytab = !subject.getPrivateCredentials(keytabClass).isEmpty();\n+ } else {\n+ this.isKeytab = !subject.getPrivateCredentials(KerberosKey.class).isEmpty();\n+ }\n this.isKrbTkt = !subject.getPrivateCredentials(KerberosTicket.class).isEmpty();\n{code}\n\n{{forName}} is fairly slow. Since the patch is targeting 2.7 which only supports JDK7, the code should be able to use the class in compile time, though we'll need to wait until jenkins to be switched to Java 7 before this patch can land.\n\n{code}\n+ @VisibleForTesting\n+ static void setShouldRenewImmediatelyForTests(boolean immediate) {\n+ shouldRenewImmediatelyForTests = immediate;\n+ }\n{code}\n\nInstead of adding this method, it might make more sense to extract the logic of login into a separate function and call the function directly in the tests.\n","from":"developer"},{"body":"bq. forName is fairly slow. Since the patch is targeting 2.7 which only supports JDK7, the code should be able to use the class in compile time, though we'll need to wait until jenkins to be switched to Java 7 before this patch can land.\n\nI'd like to commit this for 2.6, though, which will still be targeting Java 6. How about we create a constant {{KEY_TAB_CLASS}} and then do the reflection in a static initialization block? That way we only pay the lookup penalty once per JVM and the patch still works with both version of Java.","from":"developer"},{"body":"Performing the reflection in a static init block sounds like a good idea.\n\nI can see how it'd be useful to extract the logic of login into a separate function and just call it directly. I'd like to make sure to exercise as much of the reloginFromKeytab logic as possible (aside from waiting for a renew window), though.\n\nThe test verifies isKeytab == true, which is good. However, if for some reason the way isKeytab changes in reloginFromKeytab (or something else changes before actual login), it'd be good to exercise this.\n\nAttaching a patch that moves the reflection to a static block.\n\nAlso, I made some additional fixes:\n\n* Fix the conditional logic when using shouldRenewImmediatelyForTests by moving the check for null TGT ahead.\n* Remove //return","from":"developer"},{"body":"Resubmitting HADOOP-10786.3.patch to hopefully retrigger Hadoop QA jenkins.","from":"developer"},{"body":"TestMetricsSystemImpl is unrelated. I ran the test locally and successfully with the patch. Outstanding JIRA for it at HADOOP-10062.","from":"developer"},{"body":"Posting a new patch to fix a mistake. Thanks to ATM for catching it.\n\nIf the following conditional is true, we should return:\n{code}\nif (tgt != null && !shouldRenewImmediatelyForTests &&\n now < getRefreshTime(tgt)) {\n{code}\n\nI had forgotten to clean up some debugging stuff.","from":"developer"},{"body":"Oops, wrong patch. Cancelling and will resubmit.","from":"developer"},{"body":"Note that the correct patch is HADOOP-10786.4.patch.","from":"developer"},{"body":"Thanks a lot, Stephen. The latest patch looks good to me. I'll be +1 on it pending Jenkins.\n\nHaohui - just checking, has this latest patch addressed your performance concern?\n\nThanks folks.","from":"developer"},{"body":"Looks pretty good to me. Some nits:\n\n{code}\n+ private static Class KEY_TAB_CLASS = KerberosKey.class;\n{code}\n\ncan be \n\n{code}\n+ private static final Class KEY_TAB_CLASS = KerberosKey.class;\n{code}\n\n{code}\n+ public void createTestDir() {\n+ workDir = new File(System.getProperty(\"test.dir\", \"target\"));\n+ }\n{code}\n\nYou can use {{TemporaryFolder}} here so that the files can be properly cleaned up after the tests.\n\n\n{code}\n+ public File getWorkDir() {\n+ return workDir;\n+ }\n+ public MiniKdc getKdc() {\n+ return kdc;\n+ }\n+\n{code}\n\nLooks like the test can simply inline these getters to make the patch even smaller.","from":"developer"},{"body":"Thanks for the comments, Haohui. Attaching patch to address your comments.\n\nKEY_TAB_CLASS will be reassigned to Class.forName(\"javax.security.auth.kerberos.KeyTab\") on JDK7/8, so I couldn't declare it final.\n\nThe new patch now uses TemporaryFolder to ensure cleanup of the dir after test run.\n\nAlso removed the unnecessary getters.","from":"developer"},{"body":"Oh I see. The patch looks good to me. +1 pending jenkins.","from":"developer"},{"body":"I've committed the patch to trunk, branch-2 and branch-2.6. Thanks [~schu] for the contribution.","from":"developer"},{"body":"I created HADOOP-10287 to track the effort of simplifying the code in 2.7.","from":"developer"},{"body":"I think it's too late/risky to put this into 2.6; let's get this into 2.7. Thanks.","from":"developer"},{"body":"Moved to hadoop-2.7.0.","from":"developer"},{"body":"[~schu] and [~Tobi], thank you for providing this patch.\n\nI just wanted to share with everyone that even though this bug was reported against a JDK 8 code change, it appears the same change has entered the JDK 7 code line. I am seeing the same problem in the most recent OpenJDK build. With JDK 1.7.0_79, I could not repro the problem. After upgrading to JDK 1.7.0_85, I could repro the problem. I don't know the exact minor version number within JDK 7 that first introduced the change, but it's somewhere in that range.\n\nI also confirmed that this patch fixes the problem for JDK 1.7.0_85 too. To verify, I ran the test without the corresponding fixes in {{UserGroupInformation}}. I observed that the test failed on the assertion for {{ugi.isFromKeytab()}}. Then, I applied the {{UserGroupInformation}} part of the patch, and the test passed.\n\nBottom line: If you want to run a secured Hadoop cluster on JDK 1.7.0_85 or later, then you must run Apache Hadoop 2.7.0 or later.","from":"developer"},{"body":"Cherry-picked to 2.6.1","from":"developer"},{"body":"Thank you, [~vinayrpet].","from":"developer"},{"body":"[~vinayrpet], please stop committing to 2.6.1 (branch-2.6 or branch-2.6.1), we have a fairly elaborate parallel release-process going on for 2.6.1 and these cherry-picks are disrupting our progress. If you want something added to 2.6.1, please post it in the mailing lists. Thanks. ","from":"developer"},{"body":"bq. we have a fairly elaborate parallel release-process going on for 2.6.1 and these cherry-picks are disrupting our progress.\nThanks [~vinodkv]. I am actually cherry-picking only those which are listed for 2.6.1 here https://wiki.apache.org/hadoop/Release-2.6.1-Working-Notes.\nAnyway I will stop further merges until required.\n-Thanks","from":"developer"},{"body":"This wasn't originally in 2.6.1, must have been committed to 2.6, which was already 2.6.2. I just committed this to 2.6.1. Ran compilation and TestUGILoginFromKeytab before the push.","from":"developer"},{"body":"Looks like the Krb5LoginModule change was backported to [JDK 1.7.0_80|http://bugs.java.com/view_bug.do?bug_id=8043932] as well, in case anyone else hits it.\n\n","from":"developer"},{"body":"[~taoluo]\nThanks for your reminding~ We hit the problem in JDK 1.7.0_80. :(","from":"developer"},{"body":"Hey guys,\r\n\r\nI don't see this patch (HADOOP-10786.5.patch), especially this piece of code\r\n{code:java}\r\n+ private static Class KEY_TAB_CLASS = KerberosKey.class;\r\n+ static {\r\n+ try {\r\n+ // We use KEY_TAB_CLASS to determine if the UGI is logged in from\r\n+ // keytab. In JDK6 and JDK7, if useKeyTab and storeKey are specified\r\n+ // in the Krb5LoginModule, then some number of KerberosKey objects\r\n+ // are added to the Subject's private credentials. However, in JDK8,\r\n+ // a KeyTab object is added instead. More details in HADOOP-10786.\r\n+ KEY_TAB_CLASS = Class.forName(\"javax.security.auth.kerberos.KeyTab\");\r\n+ } catch (ClassNotFoundException cnfe) {\r\n+ // Ignore. javax.security.auth.kerberos.KeyTab does not exist in JDK6.\r\n+ }\r\n+ }\r\n+\r\n\r\n- this.isKeytab = !subject.getPrivateCredentials(KerberosKey.class).isEmpty();\r\n+ this.isKeytab = !subject.getPrivateCredentials(KEY_TAB_CLASS).isEmpty();\r\n\r\n{code}\r\napplied in UserGroupInformation class of Apache Hadoop 2.7.0 or 2.7.1 or in any of the subsequent versions till 3.0.0. I'm not sure if above piece of code has been changed/moved somewhere in further/future commits.\r\n\r\n I've a Kerberized Hadoop cluster that uses 2.7.3 and a client that uses 2.7.0, my client interacts with the cluster only using HDFS FileSystem API calls and I expect Kerberos ticket renewable to be automatically handled by the client's runtime dependency. Initially login from keytab is always successful but once the TGT expires, client is unable to login again, it fails with this \r\n\r\n\r\n{noformat}\r\njava.io.IOException: Failed on local exception: java.io.IOException: javax.security.sasl.SaslException: GSS initiate failed [Caused by GSSException: No valid credentials provided (Mechanism level: Failed to find any Kerberos tgt)]; Host Details : local host is: \"hc4t03283/16.202.4.11\"; destination host is: \"hc4t02044.itcs.nircorp.net\":8020; {noformat}\r\nHow can I ascertain or at least rule out that this issue isn't caused because of HADOOP-10786? This is the very reason I looked at the UserGroupInformation class' code and I didn't find this patch applied. I'm using JDK 1.8.0_151. \r\n\r\nPlease let me know if I'm missing something obvious here. ","from":"developer"},{"body":"Looks like it's in branch-2.7\r\n{code}\r\n> git log --grep HADOOP-10786 branch-2.7\r\ncommit 8f4a09b6076de9fbd6cd8ccaddf72ba9c94429ff\r\nAuthor: Vinayakumar B \r\nDate: Fri Aug 14 12:23:51 2015 +0530\r\n\r\n HADOOP-10786. Fix UGI#reloginFromKeytab on Java 8. Contributed by Stephen Chu.\r\n Moved CHANGES.txt entry to 2.6.1\r\n \r\n (cherry picked from commit e7aa81394dce61cc96d480e21204263a5f2ed153)\r\n{code}\r\n\r\nThe code has moved on a lot since that patch went in, which is why there's no match. ","from":"developer"},{"body":"[~Tobi] [~schu] is this also an issue for OpenJDK 1.8?","from":"developer"}],"created":"2014-07-05T05:02:51.000+0000","description":"Krb5LoginModule changed subtly in java 8: in particular, if useKeyTab and storeKey are specified, then only a KeyTab object is added to the Subject's private credentials, whereas in java <= 7 both a KeyTab and some number of KerberosKey objects were added.\n\nThe UGI constructor checks whether or not a keytab was used to login by looking if there are any KerberosKey objects in the Subject's private credentials. If there are, then isKeyTab is set to true, and otherwise it's set to false.\n\nThus, in java 8 isKeyTab is always false given the current UGI implementation, which makes UGI#reloginFromKeytab fail silently.\n\nAttached patch will check for a KeyTab object on the Subject, instead of a KerberosKey object. This fixes relogins from kerberos keytabs on Oracle java 8, and works on Oracle java 7 as well.","issue_id":"12725544","key":"HADOOP-10786","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2014-11-10T01:49:18.000+0000","role":"fixed_distractor","summary":"Fix UGI#reloginFromKeytab on Java 8"} {"case_id":"12726069","cluster":"DISTRACTOR-HADOOP-10798","comments":[{"body":"Which filesystem? Ie. hdfs or local? GlobStatus should return a sorted list because it calls listStatus which returns a sorted list. If the glob results are not sorted, then listStatus is likely the culprit.","created":"2014-07-08T13:30:22.500+0000"},{"body":"Tested on local filesystem. works on a macosx, but not on linux.","created":"2014-07-08T16:09:20.765+0000"},{"body":"[~daryn]: smart. I was wondering why it worked on HDFS but not on localFS.\n\nI don't know how I feel about this JIRA. The obvious solution is to have localFS sort its output on Linux, but this will cripple performance by forcing us to buffer the whole list of files before we return anything (in addition to the cost of the sort itself, of course).\n\nIt would be a bit easier to do in globStatus, since we always have an array there (unlike in listStatus where we might just have an iterator), but the same issues crop up. We're slowing stuff down, for a feature most don't need.\n\nDo users really depend on this behavior, or can we just drop this from the spec? I guess the shell probably wants sorted output, to provide a consistent display. But it can sort it itself, of course. Thoughts?","created":"2014-07-14T18:08:05.675+0000"},{"body":"In my oppinion, dropping this is totally ok. I agree, that this sorting feature may not be used often. \n","created":"2014-07-21T07:51:28.271+0000"},{"body":"I thought about this some more, and I think we should just sort the array returned from globStatus. Otherwise, applications will get inconsistent behavior. Applications can always use the raw listStatus if they don't want to pay the cost of a sort. And the cost of a sort is not that high when you already have everything in an array (like we do with globStatus).","created":"2014-07-21T18:17:52.302+0000"},{"body":"I would like to drop the requirement from the API. Since no one is reporting their application is breaking, no one appears to be depending on the output being sorted. It's better to let the client pay the cost of sorting, if they need the output sorted, rather than than impose the cost on all callers. This is particularly true for FileSystem implementations that may not by default return sorted listStatus results.","created":"2015-05-19T19:17:46.652+0000"},{"body":"Hi guys, does parallel sorting rely on the returned files being sorted? If the sorted results are in multiple files, we need to read the files in specific order to maintain the global sort right?","created":"2015-05-20T06:31:50.686+0000"},{"body":"I don't feel strongly about this either way. If you want to implement the no-sort option, please add sorting to the relevant parts of the shell, though, and post a patch modifying the javadoc.","created":"2015-05-26T19:39:47.023+0000"},{"body":"I'd like to go ahead and remove the sorting language from the API. There's no need to do the sort on shell-side since it's not being done already and hence there should be no behavior change. Sorting, particularly for large or already sorted entries can be costly. If the client really needs it sorted either in the code or from the shell, it's easy for them to do so and pay the cost. \n\nThis is a minor change to the API documentation, so it'd be a good newbie task for my colleague, [~srevanka] and I'll assign it to her if we're cool with this plan.\n","created":"2015-06-17T22:20:06.776+0000"},{"body":"bq. I'd like to go ahead and remove the sorting language from the API.\n\nOK\n\nbq. There's no need to do the sort on shell-side since it's not being done already and hence there should be no behavior change.\n\nDisagree. \"ls\" on UNIX has always returned entries in sorted order. So does Hadoop's ls, except in the special case where you are using a non-HDFS filesystem. This is not a normal case, so the discrepancy got overlooked. But I think we should fix it now.\n\nWe don't need to be ultra-fast in the shell, so there seems to be no reason why we shouldn't just add a sort. At least that's my thinking right now. What do you think?","created":"2015-06-18T00:36:35.993+0000"},{"body":"Right now this JIRA is limited to correcting an incorrect javadoc. If we add in changing the output of a reasonably distant command (specifically the shell of liststatus, not just the API of FileSystem.listStatus), that's a pretty big change relative to the original JIRA. I'd be fine with having that discussion (to sort the shell output of ls) in another JIRA. Since this additional change would potentially affect current behavior for non-HDFS FileSystems, as you say, it should probably be considered an incompatible change, which again is beyond the scope of the current JIRA.\n\nCan we just restrict this one to the API correction and open another to discuss adding sorting?","created":"2015-06-18T04:21:21.959+0000"},{"body":"OK, I did a little more research into this. The {{globStatus}} code back in 2.0.3-alpha does sort the entries it returned. Around Hadoop 2.3, the sort was lost during the globber rewrite. This was a bug, but it was hidden by the fact that HDFS sorts its listStatus entries (this behavior is undocumented, but 100% consistent).\n\nSince the API documentation says that sorted entries are returned, and since this is the case for the vast majority of use-cases (i.e. when using Hadoop with HDFS), I think changing this behavior in {{globStatus}} would be an incompatible change. Any user code relying on the old documented behavior would break. Let's commit the original patch I posted to fix this situation. If we want to have a discussion about changing the API contract we can have that discussion for Hadoop 3.0 only.\n\nalso I feel that the facts that:\n1. globStatus has historically had a sort in it\n2. users who want to optimize by avoiding a sort can use listStatus\n\nstrongly suggest that changing this behavior is not a good idea, even in 3.x.","created":"2015-06-18T18:53:24.344+0000"},{"body":"OK, if this is a regression rather than an incorrect Javadoc, I'm OK with the sorting. +1 on the patch, assuming javadoc and checkstyle warnings are spurious. I'd like it to go back to 2.6 and 2.7 as well, and will backport it if you don't have the time.","created":"2015-06-18T22:51:44.406+0000"},{"body":"Thanks, [~jghoman]. It's good to have this finally fixed. Committed to 2.8","created":"2015-06-30T23:42:08.094+0000"}],"conversations":[{"body":"(FileSystem) globStatus() does not return a sorted file list anymore.\n\nBut the API says: \" ... Results are sorted by their names.\"\n\nSeems to be lost, when the Globber Object was introduced. Can't find a sort in actual code.\n\ncode to check this behavior:\n{code}\n Configuration conf = new Configuration();\n FileSystem fs = FileSystem.get(conf);\n Path path = new Path(\"/tmp/\" + System.currentTimeMillis());\n fs.mkdirs(path);\n fs.deleteOnExit(path);\n fs.createNewFile(new Path(path, \"2\"));\n fs.createNewFile(new Path(path, \"3\"));\n fs.createNewFile(new Path(path, \"1\"));\n\n FileStatus[] status = fs.globStatus(new Path(path, \"*\"));\n Collection list = new ArrayList();\n for (FileStatus f: status) {\n list.add(f.getPath().toString());\n //System.out.println(f.getPath().toString());\n }\n boolean sorted = Ordering.natural().isOrdered(list);\n Assert.assertTrue(sorted);\n{code}","from":"reporter","subject":"globStatus() should always return a sorted list of files"},{"body":"Which filesystem? Ie. hdfs or local? GlobStatus should return a sorted list because it calls listStatus which returns a sorted list. If the glob results are not sorted, then listStatus is likely the culprit.","from":"developer"},{"body":"Tested on local filesystem. works on a macosx, but not on linux.","from":"developer"},{"body":"[~daryn]: smart. I was wondering why it worked on HDFS but not on localFS.\n\nI don't know how I feel about this JIRA. The obvious solution is to have localFS sort its output on Linux, but this will cripple performance by forcing us to buffer the whole list of files before we return anything (in addition to the cost of the sort itself, of course).\n\nIt would be a bit easier to do in globStatus, since we always have an array there (unlike in listStatus where we might just have an iterator), but the same issues crop up. We're slowing stuff down, for a feature most don't need.\n\nDo users really depend on this behavior, or can we just drop this from the spec? I guess the shell probably wants sorted output, to provide a consistent display. But it can sort it itself, of course. Thoughts?","from":"developer"},{"body":"In my oppinion, dropping this is totally ok. I agree, that this sorting feature may not be used often. \n","from":"developer"},{"body":"I thought about this some more, and I think we should just sort the array returned from globStatus. Otherwise, applications will get inconsistent behavior. Applications can always use the raw listStatus if they don't want to pay the cost of a sort. And the cost of a sort is not that high when you already have everything in an array (like we do with globStatus).","from":"developer"},{"body":"I would like to drop the requirement from the API. Since no one is reporting their application is breaking, no one appears to be depending on the output being sorted. It's better to let the client pay the cost of sorting, if they need the output sorted, rather than than impose the cost on all callers. This is particularly true for FileSystem implementations that may not by default return sorted listStatus results.","from":"developer"},{"body":"Hi guys, does parallel sorting rely on the returned files being sorted? If the sorted results are in multiple files, we need to read the files in specific order to maintain the global sort right?","from":"developer"},{"body":"I don't feel strongly about this either way. If you want to implement the no-sort option, please add sorting to the relevant parts of the shell, though, and post a patch modifying the javadoc.","from":"developer"},{"body":"I'd like to go ahead and remove the sorting language from the API. There's no need to do the sort on shell-side since it's not being done already and hence there should be no behavior change. Sorting, particularly for large or already sorted entries can be costly. If the client really needs it sorted either in the code or from the shell, it's easy for them to do so and pay the cost. \n\nThis is a minor change to the API documentation, so it'd be a good newbie task for my colleague, [~srevanka] and I'll assign it to her if we're cool with this plan.\n","from":"developer"},{"body":"bq. I'd like to go ahead and remove the sorting language from the API.\n\nOK\n\nbq. There's no need to do the sort on shell-side since it's not being done already and hence there should be no behavior change.\n\nDisagree. \"ls\" on UNIX has always returned entries in sorted order. So does Hadoop's ls, except in the special case where you are using a non-HDFS filesystem. This is not a normal case, so the discrepancy got overlooked. But I think we should fix it now.\n\nWe don't need to be ultra-fast in the shell, so there seems to be no reason why we shouldn't just add a sort. At least that's my thinking right now. What do you think?","from":"developer"},{"body":"Right now this JIRA is limited to correcting an incorrect javadoc. If we add in changing the output of a reasonably distant command (specifically the shell of liststatus, not just the API of FileSystem.listStatus), that's a pretty big change relative to the original JIRA. I'd be fine with having that discussion (to sort the shell output of ls) in another JIRA. Since this additional change would potentially affect current behavior for non-HDFS FileSystems, as you say, it should probably be considered an incompatible change, which again is beyond the scope of the current JIRA.\n\nCan we just restrict this one to the API correction and open another to discuss adding sorting?","from":"developer"},{"body":"OK, I did a little more research into this. The {{globStatus}} code back in 2.0.3-alpha does sort the entries it returned. Around Hadoop 2.3, the sort was lost during the globber rewrite. This was a bug, but it was hidden by the fact that HDFS sorts its listStatus entries (this behavior is undocumented, but 100% consistent).\n\nSince the API documentation says that sorted entries are returned, and since this is the case for the vast majority of use-cases (i.e. when using Hadoop with HDFS), I think changing this behavior in {{globStatus}} would be an incompatible change. Any user code relying on the old documented behavior would break. Let's commit the original patch I posted to fix this situation. If we want to have a discussion about changing the API contract we can have that discussion for Hadoop 3.0 only.\n\nalso I feel that the facts that:\n1. globStatus has historically had a sort in it\n2. users who want to optimize by avoiding a sort can use listStatus\n\nstrongly suggest that changing this behavior is not a good idea, even in 3.x.","from":"developer"},{"body":"OK, if this is a regression rather than an incorrect Javadoc, I'm OK with the sorting. +1 on the patch, assuming javadoc and checkstyle warnings are spurious. I'd like it to go back to 2.6 and 2.7 as well, and will backport it if you don't have the time.","from":"developer"},{"body":"Thanks, [~jghoman]. It's good to have this finally fixed. Committed to 2.8","from":"developer"}],"created":"2014-07-08T13:22:01.000+0000","description":"(FileSystem) globStatus() does not return a sorted file list anymore.\n\nBut the API says: \" ... Results are sorted by their names.\"\n\nSeems to be lost, when the Globber Object was introduced. Can't find a sort in actual code.\n\ncode to check this behavior:\n{code}\n Configuration conf = new Configuration();\n FileSystem fs = FileSystem.get(conf);\n Path path = new Path(\"/tmp/\" + System.currentTimeMillis());\n fs.mkdirs(path);\n fs.deleteOnExit(path);\n fs.createNewFile(new Path(path, \"2\"));\n fs.createNewFile(new Path(path, \"3\"));\n fs.createNewFile(new Path(path, \"1\"));\n\n FileStatus[] status = fs.globStatus(new Path(path, \"*\"));\n Collection list = new ArrayList();\n for (FileStatus f: status) {\n list.add(f.getPath().toString());\n //System.out.println(f.getPath().toString());\n }\n boolean sorted = Ordering.natural().isOrdered(list);\n Assert.assertTrue(sorted);\n{code}","issue_id":"12726069","key":"HADOOP-10798","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2015-06-30T23:43:15.000+0000","role":"fixed_distractor","summary":"globStatus() should always return a sorted list of files"} {"case_id":"12727751","cluster":"DISTRACTOR-HADOOP-10852","comments":[{"body":"The _NetgroupCache_ is fixed by keeping state in only one _ConcurrentHashMap_ - {{userToNetgroupsMap}}.\n{{add}} updates {{userToNetgroupsMap}} directly.\n\nThere is a slight performance loss when invoking {{isCached}} and {{{userToNetgroupsMap}}\nBut these are invoked only during {{refresh}} and adding a new group. These invocations should be rare compared to {{getNetgroups}} invocation.","created":"2014-07-16T19:25:16.524+0000"},{"body":"updating the patch to fix the test failures.","created":"2014-08-20T00:26:52.526+0000"},{"body":"Attaching a newer patch which keeps single collection.","created":"2014-12-13T16:50:13.275+0000"},{"body":"+1 for the patch except for some formatting. Feel free to commit after fixing it.\n\n# Space between {{//}} and comment text. e.g. _//ConcurrentHashMap does not allow null values;_\n# Spaces around tokens in for e.g. _String user:users_\n\nThe patch looks good otherwise and existing unit tests look sufficient.","created":"2014-12-14T09:16:50.569+0000"},{"body":"Thanks for the review [~arpitagarwal]. Attaching the newer patch. Will commit once jenkins build is done with expected results.","created":"2014-12-15T18:45:51.725+0000"},{"body":"committed to trunk and branch-2","created":"2014-12-15T22:06:25.110+0000"}],"conversations":[{"body":"_NetgroupCache_ internally uses two ConcurrentHashMaps and a boolean variable to signal updates on one of the ConcurrentHashMap\nNone of the functions are synchronized and hence is possible to have unexpected results due to race condition between different threads.\n\nAs an example, consider the following sequence:\n\nThread1 :\n{{add}} a group\n{{netgroupToUsersMap}} is updated.\n{{netgroupToUsersMapUpdated}} is set to true.\nThread 2:\ncalls {{getNetgroups}} for a user\nDue to re-ordering, {{netgroupToUsersMapUpdated=true}} is visible, but updates in {{netgroupToUsersMap}} is not visible.\nDoes a wrong update with older {{netgroupToUsersMap}} values. ","from":"reporter","subject":"NetgroupCache is not thread-safe"},{"body":"The _NetgroupCache_ is fixed by keeping state in only one _ConcurrentHashMap_ - {{userToNetgroupsMap}}.\n{{add}} updates {{userToNetgroupsMap}} directly.\n\nThere is a slight performance loss when invoking {{isCached}} and {{{userToNetgroupsMap}}\nBut these are invoked only during {{refresh}} and adding a new group. These invocations should be rare compared to {{getNetgroups}} invocation.","from":"developer"},{"body":"updating the patch to fix the test failures.","from":"developer"},{"body":"Attaching a newer patch which keeps single collection.","from":"developer"},{"body":"+1 for the patch except for some formatting. Feel free to commit after fixing it.\n\n# Space between {{//}} and comment text. e.g. _//ConcurrentHashMap does not allow null values;_\n# Spaces around tokens in for e.g. _String user:users_\n\nThe patch looks good otherwise and existing unit tests look sufficient.","from":"developer"},{"body":"Thanks for the review [~arpitagarwal]. Attaching the newer patch. Will commit once jenkins build is done with expected results.","from":"developer"},{"body":"committed to trunk and branch-2","from":"developer"}],"created":"2014-07-16T19:20:09.000+0000","description":"_NetgroupCache_ internally uses two ConcurrentHashMaps and a boolean variable to signal updates on one of the ConcurrentHashMap\nNone of the functions are synchronized and hence is possible to have unexpected results due to race condition between different threads.\n\nAs an example, consider the following sequence:\n\nThread1 :\n{{add}} a group\n{{netgroupToUsersMap}} is updated.\n{{netgroupToUsersMapUpdated}} is set to true.\nThread 2:\ncalls {{getNetgroups}} for a user\nDue to re-ordering, {{netgroupToUsersMapUpdated=true}} is visible, but updates in {{netgroupToUsersMap}} is not visible.\nDoes a wrong update with older {{netgroupToUsersMap}} values. ","issue_id":"12727751","key":"HADOOP-10852","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2014-12-15T22:06:25.000+0000","role":"fixed_distractor","summary":"NetgroupCache is not thread-safe"} {"case_id":"12749638","cluster":"DISTRACTOR-HADOOP-11218","comments":[{"body":"I propose a more aggressive approach, that is, to remove this configuration and pick reasonable ciphers / protocols. The reason is that it requires some domain knowledges to properly configure it correctly, and misconfiguring it can lead to security holes. It would be nice to not having a configuration that can shoot the users' foot.\n\nNote that the configuration is added in 2.6 which still supports Java 6. The configuration allows users that run 2.6 on Java 7 to disable the flawed protocols / ciphers. As 2.7 is going to support Java 7 only, I think that the original motivation is fulfilled thus the configuration can be removed.","created":"2014-10-29T17:48:43.846+0000"},{"body":"We want to disable SSLv3 in Java 7 as well, and would just want to add protocols to the whitelist. I don't see how we can revert the changes altogether. Am I missing something here? ","created":"2014-10-29T17:53:52.947+0000"},{"body":"Let me try to rephrase a little bit to see whether it makes sense: I propose to not exposing the configuration to the user. Instead, the code should pick reasonable settings.\n\nWhat I'm proposing is that the code just disable SSLv3 and not exposing it as a configuration, as configuring it properly requires domain knowledges and misconfiguration can lead to security issues. Make sense?","created":"2014-10-29T18:02:00.157+0000"},{"body":"I see merit to having configurations with reasonable defaults. Configs allow users to workaround any security vulnerabilities at least until an official release with fixes comes out.","created":"2014-10-29T18:16:00.642+0000"},{"body":"I'm with [~kkambatl] on this one. I'd really love the ability to disable broken implementations until a fix comes out. We tend to have this view that as soon as we release code it gets immediately installed everywhere. The reality is that it can take weeks to deploy a new version. Having the ability to disable a broken codepath is always a nicety.\n\n","created":"2014-10-29T19:18:25.072+0000"},{"body":"Thanks for the explanation. It's fine with me as long as we have a reasonable configuration.","created":"2014-10-29T20:18:36.475+0000"},{"body":"I posted an approach for enabling TLSv1.1 and TLSv1.2 for HttpFS service in duplicate ticket. The reason for our customers to go for TLS1.2 is that current RHEL7 and Ubuntu based HDFS client gateways when used with curl can enforce which TLS level to use. The security teams wants application using curl to enforce TLSv1.2; however, in absence of server support its not feasible. Regardless, once we allow TLSv1, TLSv1.1, TLSv1.2 options as part of server config,server can choose highest level of support for TLS available and may or may not honor client request. But, atleast client application can downgrade or choose not to use TLSv1. Since we support JDK7 I propose that we add support for TLSv1.1 and TLSv1.2 for KMS and HttpFS services atleast using SSLFactory.\nPlease find the code snippet for implemented changes.\n{code:xml}\n \n{code}\n\nChanges include addition of TLSv1.1,TLSv1.2 to SSLenabledProtocols xml attribute on line 73 of file hadoop/hadoop-hdfs-project/hadoop-hdfs-httpfs/src/main/tomcat/ssl-server.xml.conf","created":"2015-10-02T05:29:57.529+0000"},{"body":"Please find the result of tests carried out.\n{noformat}\n[root@vjs-1 ~]# diff /opt/myclient/hadoop-httpfs/tomcat-conf.https/conf/server.xml /opt/myclient/hadoop-httpfs/tomcat-conf.https/conf/server_tls1.xml \n73c73\n<                clientAuth=\"false\" sslEnabledProtocols=“TLSv1,TLSv1.1,TLSv1.2,SSLv2Hello\"\n---\n>                clientAuth=\"false\" sslEnabledProtocols=\"TLSv1,SSLv2Hello\"\n\n[root@vjkc ~]# openssl s_client -connect vjs-1.vpc.myclient.com:14000  -tls1 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-1.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n\n[root@vjkc ~]# openssl s_client -connect vjs-1.vpc.myclient.com:14000  -tls1_1 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep -i Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-1.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n\n[root@vjkc ~]# openssl s_client -connect vjs-1.vpc.myclient.com:14000  -tls1_2 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep -i Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-1.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n{noformat}\n","created":"2015-10-02T05:34:01.674+0000"},{"body":"The code snippted changes for kms will be required on line 73 of file /hadoop-common-project/hadoop-kms/src/main/tomcat/ssl-server.xml.conf. The code changes are as follows:\n{code:xml}\n\n{code:xml}\n\nPlease see the excerpts from test log.\n{noformat}\n[root@vjs-kms ~]# diff /opt/myclient/hadoop-kms/tomcat-conf.https/conf/server.xml /opt/myclient/hadoop-kms/tomcat-conf.https/conf/server_tls1.xml \n73c73\n<                clientAuth=\"false\" sslEnabledProtocols=“TLSv1,TLSv1.1,TLSv1.2,SSLv2Hello\"\n---\n>                clientAuth=\"false\" sslEnabledProtocols=\"TLSv1,SSLv2Hello\"\n\n[root@vjkc ~]# openssl s_client -connect vjs-kms.vpc.myclient.com:16000  -tls1 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-kms.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n\n[root@vjkc ~]# openssl s_client -connect vjs-kms.vpc.myclient.com:16000  -tls1_1 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep -i Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-kms.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n\n[root@vjkc ~]# openssl s_client -connect vjs-kms.vpc.myclient.com:16000  -tls1_2 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep -i Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-kms.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n{noformat}\n\nPlease review my proposed changes and suggest any feedback. I will work on the patch for submission in the meantime.\n","created":"2015-10-02T05:50:38.536+0000"},{"body":"Please find attached the proposed patch for enabling TLSv1.1 and TLSv1.2 for HttpFS and KMS service. Please review and suggest any changes. This is my first contribution any advise is appreciated.","created":"2015-10-02T06:01:44.785+0000"},{"body":"The code changes have been attached already. Formally submitting the patch","created":"2015-10-02T21:12:47.622+0000"},{"body":"The two -1 are encountered across all patches. Please suggest if anything is required from my side.","created":"2015-10-05T01:51:31.927+0000"},{"body":"+1. Will commit tomorrow if there are no more comments.","created":"2015-10-13T19:34:26.519+0000"},{"body":"[~cmccabe], can this please be committed as you've reviewed it already? Considering this for a 2.7.2 RC this weekend.","created":"2015-10-30T23:50:50.566+0000"},{"body":"[~cmccabe], gentle bump again. Will be great if this can get committed in a day or two, as I am chasing a few other tickets besides this. Thanks.","created":"2015-11-02T21:26:40.077+0000"},{"body":"Moving this out into 2.7.3 in the interest of 2.7.2's progress.","created":"2015-11-06T00:21:38.160+0000"},{"body":"+1. Committing shortly","created":"2015-11-22T23:59:31.718+0000"},{"body":"I've committed the patch to trunk and branch-2. Thanks [~SINGHVJD] for the contribution.","created":"2015-11-23T00:00:58.794+0000"},{"body":"This looks like a candidate for branch-2.7 (along with HADOOP-12817).","created":"2016-07-15T18:42:48.040+0000"}],"conversations":[{"body":"HADOOP-11217 required us to specifically list the versions of TLS that KMS supports. With Hadoop 2.7 dropping support for Java 6 and Java 7 supporting TLSv1.1 and TLSv1.2, we should add them to the list.","from":"reporter","subject":"Add TLSv1.1,TLSv1.2 to KMS, HttpFS, SSLFactory"},{"body":"I propose a more aggressive approach, that is, to remove this configuration and pick reasonable ciphers / protocols. The reason is that it requires some domain knowledges to properly configure it correctly, and misconfiguring it can lead to security holes. It would be nice to not having a configuration that can shoot the users' foot.\n\nNote that the configuration is added in 2.6 which still supports Java 6. The configuration allows users that run 2.6 on Java 7 to disable the flawed protocols / ciphers. As 2.7 is going to support Java 7 only, I think that the original motivation is fulfilled thus the configuration can be removed.","from":"developer"},{"body":"We want to disable SSLv3 in Java 7 as well, and would just want to add protocols to the whitelist. I don't see how we can revert the changes altogether. Am I missing something here? ","from":"developer"},{"body":"Let me try to rephrase a little bit to see whether it makes sense: I propose to not exposing the configuration to the user. Instead, the code should pick reasonable settings.\n\nWhat I'm proposing is that the code just disable SSLv3 and not exposing it as a configuration, as configuring it properly requires domain knowledges and misconfiguration can lead to security issues. Make sense?","from":"developer"},{"body":"I see merit to having configurations with reasonable defaults. Configs allow users to workaround any security vulnerabilities at least until an official release with fixes comes out.","from":"developer"},{"body":"I'm with [~kkambatl] on this one. I'd really love the ability to disable broken implementations until a fix comes out. We tend to have this view that as soon as we release code it gets immediately installed everywhere. The reality is that it can take weeks to deploy a new version. Having the ability to disable a broken codepath is always a nicety.\n\n","from":"developer"},{"body":"Thanks for the explanation. It's fine with me as long as we have a reasonable configuration.","from":"developer"},{"body":"I posted an approach for enabling TLSv1.1 and TLSv1.2 for HttpFS service in duplicate ticket. The reason for our customers to go for TLS1.2 is that current RHEL7 and Ubuntu based HDFS client gateways when used with curl can enforce which TLS level to use. The security teams wants application using curl to enforce TLSv1.2; however, in absence of server support its not feasible. Regardless, once we allow TLSv1, TLSv1.1, TLSv1.2 options as part of server config,server can choose highest level of support for TLS available and may or may not honor client request. But, atleast client application can downgrade or choose not to use TLSv1. Since we support JDK7 I propose that we add support for TLSv1.1 and TLSv1.2 for KMS and HttpFS services atleast using SSLFactory.\nPlease find the code snippet for implemented changes.\n{code:xml}\n \n{code}\n\nChanges include addition of TLSv1.1,TLSv1.2 to SSLenabledProtocols xml attribute on line 73 of file hadoop/hadoop-hdfs-project/hadoop-hdfs-httpfs/src/main/tomcat/ssl-server.xml.conf","from":"developer"},{"body":"Please find the result of tests carried out.\n{noformat}\n[root@vjs-1 ~]# diff /opt/myclient/hadoop-httpfs/tomcat-conf.https/conf/server.xml /opt/myclient/hadoop-httpfs/tomcat-conf.https/conf/server_tls1.xml \n73c73\n<                clientAuth=\"false\" sslEnabledProtocols=“TLSv1,TLSv1.1,TLSv1.2,SSLv2Hello\"\n---\n>                clientAuth=\"false\" sslEnabledProtocols=\"TLSv1,SSLv2Hello\"\n\n[root@vjkc ~]# openssl s_client -connect vjs-1.vpc.myclient.com:14000  -tls1 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-1.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n\n[root@vjkc ~]# openssl s_client -connect vjs-1.vpc.myclient.com:14000  -tls1_1 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep -i Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-1.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n\n[root@vjkc ~]# openssl s_client -connect vjs-1.vpc.myclient.com:14000  -tls1_2 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep -i Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-1.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n{noformat}\n","from":"developer"},{"body":"The code snippted changes for kms will be required on line 73 of file /hadoop-common-project/hadoop-kms/src/main/tomcat/ssl-server.xml.conf. The code changes are as follows:\n{code:xml}\n\n{code:xml}\n\nPlease see the excerpts from test log.\n{noformat}\n[root@vjs-kms ~]# diff /opt/myclient/hadoop-kms/tomcat-conf.https/conf/server.xml /opt/myclient/hadoop-kms/tomcat-conf.https/conf/server_tls1.xml \n73c73\n<                clientAuth=\"false\" sslEnabledProtocols=“TLSv1,TLSv1.1,TLSv1.2,SSLv2Hello\"\n---\n>                clientAuth=\"false\" sslEnabledProtocols=\"TLSv1,SSLv2Hello\"\n\n[root@vjkc ~]# openssl s_client -connect vjs-kms.vpc.myclient.com:16000  -tls1 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-kms.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n\n[root@vjkc ~]# openssl s_client -connect vjs-kms.vpc.myclient.com:16000  -tls1_1 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep -i Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-kms.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n\n[root@vjkc ~]# openssl s_client -connect vjs-kms.vpc.myclient.com:16000  -tls1_2 -CAfile /opt/myclient/security/setup/ca-certs/VIJAY-WIN-HEN9IV5CAGA-CA.pem | grep -i Renegotiation\ndepth=1 DC = FCE, DC = SINGH, DC = VIJAY, CN = VIJAY-WIN-HEN9IV5CAGA-CA\nverify return:1\ndepth=0 C = US, ST = Illinois, L = Chicago, O = myclient, OU = EDHCLUSTER, CN = vjs-kms.vpc.myclient.com\nverify return:1\n\nSecure Renegotiation IS supported\n{noformat}\n\nPlease review my proposed changes and suggest any feedback. I will work on the patch for submission in the meantime.\n","from":"developer"},{"body":"Please find attached the proposed patch for enabling TLSv1.1 and TLSv1.2 for HttpFS and KMS service. Please review and suggest any changes. This is my first contribution any advise is appreciated.","from":"developer"},{"body":"The code changes have been attached already. Formally submitting the patch","from":"developer"},{"body":"The two -1 are encountered across all patches. Please suggest if anything is required from my side.","from":"developer"},{"body":"+1. Will commit tomorrow if there are no more comments.","from":"developer"},{"body":"[~cmccabe], can this please be committed as you've reviewed it already? Considering this for a 2.7.2 RC this weekend.","from":"developer"},{"body":"[~cmccabe], gentle bump again. Will be great if this can get committed in a day or two, as I am chasing a few other tickets besides this. Thanks.","from":"developer"},{"body":"Moving this out into 2.7.3 in the interest of 2.7.2's progress.","from":"developer"},{"body":"+1. Committing shortly","from":"developer"},{"body":"I've committed the patch to trunk and branch-2. Thanks [~SINGHVJD] for the contribution.","from":"developer"},{"body":"This looks like a candidate for branch-2.7 (along with HADOOP-12817).","from":"developer"}],"created":"2014-10-21T22:14:17.000+0000","description":"HADOOP-11217 required us to specifically list the versions of TLS that KMS supports. With Hadoop 2.7 dropping support for Java 6 and Java 7 supporting TLSv1.1 and TLSv1.2, we should add them to the list.","issue_id":"12749638","key":"HADOOP-11218","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2015-11-23T00:00:58.000+0000","role":"fixed_distractor","summary":"Add TLSv1.1,TLSv1.2 to KMS, HttpFS, SSLFactory"} {"case_id":"12740623","cluster":"DISTRACTOR-HADOOP-11286","comments":[{"body":"The changes in HDFS-6134 and HADOOP-10150 appear to have introduced this additional breakage for downstream users.","created":"2014-09-10T22:03:01.926+0000"},{"body":"Attached patch to CryptoUtils, [^0001-MAPREDUCE-6083-Avoid-client-use-of-deprecated-LimitI.patch] for the 2.6.0-SNAPSHOT branch. (Bumps Guava dependency to version 14.0.1, which is the last version with both LimitInputStream and the alternative, in order to minimize impact with maximal benefit.)","created":"2014-09-10T22:16:08.029+0000"},{"body":"Does this patch also remove the use of LimitInputStream in other parts of Hadoop? For example, in MiniDFSCluster?","created":"2014-10-04T11:53:31.932+0000"},{"body":"version 2.5.1:\n\njava.lang.NoClassDefFoundError: com/google/common/io/LimitInputStream\n\tat java.net.URLClassLoader$1.run(URLClassLoader.java:366)\n\tat java.net.URLClassLoader$1.run(URLClassLoader.java:355)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat java.net.URLClassLoader.findClass(URLClassLoader.java:354)\n\tat java.lang.ClassLoader.loadClass(ClassLoader.java:425)\n\tat sun.misc.Launcher$AppClassLoader.loadClass(Launcher.java:308)\n\tat java.lang.ClassLoader.loadClass(ClassLoader.java:358)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImageFormat$LoaderDelegator.load(FSImageFormat.java:223)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImage.loadFSImage(FSImage.java:913)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImage.loadFSImage(FSImage.java:899)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImage.loadFSImageFile(FSImage.java:722)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImage.loadFSImage(FSImage.java:660)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImage.recoverTransitionRead(FSImage.java:279)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.loadFSImage(FSNamesystem.java:955)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.loadFromDisk(FSNamesystem.java:700)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNode.loadNamesystem(NameNode.java:529)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNode.initialize(NameNode.java:585)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNode.(NameNode.java:751)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNode.(NameNode.java:735)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNode.createNameNode(NameNode.java:1407)\n\tat org.apache.hadoop.hdfs.MiniDFSCluster.createNameNode(MiniDFSCluster.java:998)\n\tat org.apache.hadoop.hdfs.MiniDFSCluster.createNameNodesAndSetConf(MiniDFSCluster.java:869)\n\tat org.apache.hadoop.hdfs.MiniDFSCluster.initMiniDFSCluster(MiniDFSCluster.java:704)\n\tat org.apache.hadoop.hdfs.MiniDFSCluster.(MiniDFSCluster.java:376)\n\tat org.apache.hadoop.hdfs.MiniDFSCluster$Builder.build(MiniDFSCluster.java:357)\n\tat de.mstier.hadoop.WordCountTest.setUp(WordCountTest.java:65)\n","created":"2014-10-04T11:55:35.989+0000"},{"body":"bq. Does this patch also remove the use of LimitInputStream in other parts of Hadoop? For example, in MiniDFSCluster?\n\nNo. See the referenced HDFS-7040. This only mitigates the new problem introduced in 2.6.0. HDFS-7040 addresses the issue introduced in 2.4.0, which is limited to MiniDFSCluster. This problem is worse, though, because it directly impacts many more users than the MiniDFSCluster one.","created":"2014-10-04T17:20:19.803+0000"},{"body":"Would this be more likely to be accepted for 2.6.0 if it were provided as a copied/re-implemented version of LimitInputStream instead of a dependency version change?","created":"2014-10-15T20:39:25.929+0000"},{"body":"[~ctubbsii] - Apologies for the late response. Unfortunately, we can't change guava versions in 2.6 since it would be incompatible. ","created":"2014-11-05T13:28:32.759+0000"},{"body":"[~acmurthy]: I understand. What about my previous question, about whether a fix would be accepted if implemented as a copied/re-implemented version of LimitInputStream instead of a guava version change?","created":"2014-11-06T01:27:58.583+0000"},{"body":"Yes, I'm cool with that. Tx!","created":"2014-11-06T03:01:46.020+0000"},{"body":"Uploading a new patch ([^0001-MAPREDUCE-6083-and-HDFS-7040-Avoid-Guava-s-LimitInpu.patch]) which copies the HBase solution for the same issue. It also incidentally adds HDFS-7040, which is the other places where LimitInputStream is used (but only for version 2.6.0 and later).","created":"2014-11-07T21:59:09.891+0000"},{"body":"I just committed this. Thanks [~ctubbsii]!","created":"2014-11-09T13:27:19.373+0000"}],"conversations":[{"body":"See HDFS-7040 for more background/details.\n\nIn recent 2.6.0-SNAPSHOTs, the use of LimitInputStream was added to CryptoUtils. This is part of the API components of Hadoop, which severely impacts users who were utilizing newer versions of Guava, where the @Beta and @Deprecated class, LimitInputStream, has been removed (removed in version 15 and later), beyond the impact already experienced in 2.4.0 as identified in HDFS-7040.","from":"reporter","subject":"Map/Reduce dangerously adds Guava @Beta class to CryptoUtils"},{"body":"The changes in HDFS-6134 and HADOOP-10150 appear to have introduced this additional breakage for downstream users.","from":"developer"},{"body":"Attached patch to CryptoUtils, [^0001-MAPREDUCE-6083-Avoid-client-use-of-deprecated-LimitI.patch] for the 2.6.0-SNAPSHOT branch. (Bumps Guava dependency to version 14.0.1, which is the last version with both LimitInputStream and the alternative, in order to minimize impact with maximal benefit.)","from":"developer"},{"body":"Does this patch also remove the use of LimitInputStream in other parts of Hadoop? For example, in MiniDFSCluster?","from":"developer"},{"body":"version 2.5.1:\n\njava.lang.NoClassDefFoundError: com/google/common/io/LimitInputStream\n\tat java.net.URLClassLoader$1.run(URLClassLoader.java:366)\n\tat java.net.URLClassLoader$1.run(URLClassLoader.java:355)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat java.net.URLClassLoader.findClass(URLClassLoader.java:354)\n\tat java.lang.ClassLoader.loadClass(ClassLoader.java:425)\n\tat sun.misc.Launcher$AppClassLoader.loadClass(Launcher.java:308)\n\tat java.lang.ClassLoader.loadClass(ClassLoader.java:358)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImageFormat$LoaderDelegator.load(FSImageFormat.java:223)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImage.loadFSImage(FSImage.java:913)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImage.loadFSImage(FSImage.java:899)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImage.loadFSImageFile(FSImage.java:722)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImage.loadFSImage(FSImage.java:660)\n\tat org.apache.hadoop.hdfs.server.namenode.FSImage.recoverTransitionRead(FSImage.java:279)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.loadFSImage(FSNamesystem.java:955)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.loadFromDisk(FSNamesystem.java:700)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNode.loadNamesystem(NameNode.java:529)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNode.initialize(NameNode.java:585)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNode.(NameNode.java:751)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNode.(NameNode.java:735)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNode.createNameNode(NameNode.java:1407)\n\tat org.apache.hadoop.hdfs.MiniDFSCluster.createNameNode(MiniDFSCluster.java:998)\n\tat org.apache.hadoop.hdfs.MiniDFSCluster.createNameNodesAndSetConf(MiniDFSCluster.java:869)\n\tat org.apache.hadoop.hdfs.MiniDFSCluster.initMiniDFSCluster(MiniDFSCluster.java:704)\n\tat org.apache.hadoop.hdfs.MiniDFSCluster.(MiniDFSCluster.java:376)\n\tat org.apache.hadoop.hdfs.MiniDFSCluster$Builder.build(MiniDFSCluster.java:357)\n\tat de.mstier.hadoop.WordCountTest.setUp(WordCountTest.java:65)\n","from":"developer"},{"body":"bq. Does this patch also remove the use of LimitInputStream in other parts of Hadoop? For example, in MiniDFSCluster?\n\nNo. See the referenced HDFS-7040. This only mitigates the new problem introduced in 2.6.0. HDFS-7040 addresses the issue introduced in 2.4.0, which is limited to MiniDFSCluster. This problem is worse, though, because it directly impacts many more users than the MiniDFSCluster one.","from":"developer"},{"body":"Would this be more likely to be accepted for 2.6.0 if it were provided as a copied/re-implemented version of LimitInputStream instead of a dependency version change?","from":"developer"},{"body":"[~ctubbsii] - Apologies for the late response. Unfortunately, we can't change guava versions in 2.6 since it would be incompatible. ","from":"developer"},{"body":"[~acmurthy]: I understand. What about my previous question, about whether a fix would be accepted if implemented as a copied/re-implemented version of LimitInputStream instead of a guava version change?","from":"developer"},{"body":"Yes, I'm cool with that. Tx!","from":"developer"},{"body":"Uploading a new patch ([^0001-MAPREDUCE-6083-and-HDFS-7040-Avoid-Guava-s-LimitInpu.patch]) which copies the HBase solution for the same issue. It also incidentally adds HDFS-7040, which is the other places where LimitInputStream is used (but only for version 2.6.0 and later).","from":"developer"},{"body":"I just committed this. Thanks [~ctubbsii]!","from":"developer"}],"created":"2014-09-10T22:00:31.000+0000","description":"See HDFS-7040 for more background/details.\n\nIn recent 2.6.0-SNAPSHOTs, the use of LimitInputStream was added to CryptoUtils. This is part of the API components of Hadoop, which severely impacts users who were utilizing newer versions of Guava, where the @Beta and @Deprecated class, LimitInputStream, has been removed (removed in version 15 and later), beyond the impact already experienced in 2.4.0 as identified in HDFS-7040.","issue_id":"12740623","key":"HADOOP-11286","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2014-11-09T13:27:19.000+0000","role":"fixed_distractor","summary":"Map/Reduce dangerously adds Guava @Beta class to CryptoUtils"} {"case_id":"12757103","cluster":"DISTRACTOR-HADOOP-11327","comments":[{"body":"Consider bumping the priority of this ticket - the bug exposes an edge case where a BloomFilter returns an incorrect result.","created":"2014-11-22T00:10:18.748+0000"},{"body":"Hi [~tim.luo]. I'm interested in seeing this issue resolved. Please let me know if you plan on working on it any time soon. Otherwise, would it be okay if I took it over?","created":"2015-01-13T21:51:07.228+0000"},{"body":"Hi [~epayne], you can definitely take it over.\nI've had some issues getting my local copy to build (even pre-changes) so I've been held up.","created":"2015-01-13T22:18:32.201+0000"},{"body":"Thanks [~tim.luo]. Here is patch, version 1.","created":"2015-01-13T22:49:24.598+0000"},{"body":"Thanks for picking this up, Eric!\n\nPatch looks good, but I have one nit on the unit test. Rather than unit testing not() just for the behavior of the lone bit, I think it would be better to test the other bits as well to ensure the not() method really behaves as expected. If there had been a proper not() unit test previously then it would have caught this error and other types of errors that could occur in an implementation of not().","created":"2015-01-15T21:43:16.584+0000"},{"body":"Thanks, [~jlowe], for the review and comments.\n\nI have updated the test case in version 2 of the patch.","created":"2015-01-16T17:13:14.827+0000"},{"body":"+1 lgtm. Committing this.","created":"2015-01-21T19:05:11.837+0000"},{"body":"Thanks, Eric! I committed this to trunk and branch-2.","created":"2015-01-21T19:08:26.510+0000"}],"conversations":[{"body":"There's an off-by-one error in {{BloomFilter#not()}}:\n\n{{BloomFilter#not}} calls {{BitSet#flip(0, vectorSize - 1)}}, but according to the javadoc for that method, {{toIndex}} is end-_exclusive_:\n{noformat}\n* @param toIndex index after the last bit to flip\n{noformat}\n\nThis means that the last bit in the bit array is not flipped.\nSpecifically, this was discovered in the following scenario:\n1. A new/empty {{BloomFilter}} was created with vectorSize=7.\n2. Invoke {{bloomFilter.not()}}; now expecting a bloom filter with all 7 bits (0 through 6) flipped to 1 and membershipTest(...) to always return true.\n3. However, membershipTest(...) was found to often not return true, and upon inspection, the BitSet only had bits 0 through 5 flipped.\n\nThe fix should be simple: remove the \"- 1\" from the call to {{BitSet#flip}}.","from":"reporter","subject":"BloomFilter#not() omits the last bit, resulting in an incorrect filter"},{"body":"Consider bumping the priority of this ticket - the bug exposes an edge case where a BloomFilter returns an incorrect result.","from":"developer"},{"body":"Hi [~tim.luo]. I'm interested in seeing this issue resolved. Please let me know if you plan on working on it any time soon. Otherwise, would it be okay if I took it over?","from":"developer"},{"body":"Hi [~epayne], you can definitely take it over.\nI've had some issues getting my local copy to build (even pre-changes) so I've been held up.","from":"developer"},{"body":"Thanks [~tim.luo]. Here is patch, version 1.","from":"developer"},{"body":"Thanks for picking this up, Eric!\n\nPatch looks good, but I have one nit on the unit test. Rather than unit testing not() just for the behavior of the lone bit, I think it would be better to test the other bits as well to ensure the not() method really behaves as expected. If there had been a proper not() unit test previously then it would have caught this error and other types of errors that could occur in an implementation of not().","from":"developer"},{"body":"Thanks, [~jlowe], for the review and comments.\n\nI have updated the test case in version 2 of the patch.","from":"developer"},{"body":"+1 lgtm. Committing this.","from":"developer"},{"body":"Thanks, Eric! I committed this to trunk and branch-2.","from":"developer"}],"created":"2014-11-21T22:41:37.000+0000","description":"There's an off-by-one error in {{BloomFilter#not()}}:\n\n{{BloomFilter#not}} calls {{BitSet#flip(0, vectorSize - 1)}}, but according to the javadoc for that method, {{toIndex}} is end-_exclusive_:\n{noformat}\n* @param toIndex index after the last bit to flip\n{noformat}\n\nThis means that the last bit in the bit array is not flipped.\nSpecifically, this was discovered in the following scenario:\n1. A new/empty {{BloomFilter}} was created with vectorSize=7.\n2. Invoke {{bloomFilter.not()}}; now expecting a bloom filter with all 7 bits (0 through 6) flipped to 1 and membershipTest(...) to always return true.\n3. However, membershipTest(...) was found to often not return true, and upon inspection, the BitSet only had bits 0 through 5 flipped.\n\nThe fix should be simple: remove the \"- 1\" from the call to {{BitSet#flip}}.","issue_id":"12757103","key":"HADOOP-11327","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2015-01-21T19:08:26.000+0000","role":"fixed_distractor","summary":"BloomFilter#not() omits the last bit, resulting in an incorrect filter"} {"case_id":"12757993","cluster":"DISTRACTOR-HADOOP-11340","comments":[{"body":"[~kasha] this something we can fix?","created":"2014-12-04T01:19:17.713+0000"},{"body":"My bad. Thanks [~fs111] for reporting this. Let me follow up on this. ","created":"2014-12-04T01:27:50.480+0000"},{"body":"Sorry for the delay on this - subversion was down at the time and I forgot about it after.\n\nJust updated the mds files to refer to the right tarballs. Should go through the mirrors in 24 hours. ","created":"2014-12-28T19:04:12.701+0000"}],"conversations":[{"body":"The mds file containing the checksums is referring to RC1 instead of the final release and that breaks some of our processes, since we do name matching and expect the file-name in the mds file to match the actual tarball name.\n\nSee for instance:\n\n{code}\n$ curl http://mirror.23media.de/apache/hadoop/core/hadoop-2.5.2/hadoop-2.5.2-src.tar.gz.mds\nhadoop-2.5.2-RC1-src.tar.gz: MD5 = AA 48 7A 0E 9C 9D DE 70 33 8F B4 82 E9 26\n C5 9A\nhadoop-2.5.2-RC1-src.tar.gz: SHA1 = 82FE 9B9A 580E AB73 249D 9EF7 3B8B 380C\n 43BD 7EAF\nhadoop-2.5.2-RC1-src.tar.gz: RMD160 = 54CA E137 19F8 5C39 7F92 A53E 4DA4 C70E\n 4908 CAD3\nhadoop-2.5.2-RC1-src.tar.gz: SHA224 = C9B68308 47CEC04C 9CA23491 3F33EC39\n CBF421D8 3738ACDA E40AD2C6\nhadoop-2.5.2-RC1-src.tar.gz: SHA256 = 139EF872 09C5637E 6FACA381 6CFD82CE\n 653DF2F6 1ADD240D 01C9D62E F41AD8A4\nhadoop-2.5.2-RC1-src.tar.gz: SHA384 = B2C5EA35 782BDDD9 EBFF42D7 A608F496\n 466744C3 A0D1C3A5 F212FE7B 6014D159\n 0B45B61E 930909A4 D538AC05 30336E26\nhadoop-2.5.2-RC1-src.tar.gz: SHA512 = 68CB93CC 051F8B91 ABA22913 ABD484AD\n C6D1715A EAEB2D08 829942AC 8D7D765D\n 700771D4 BEB44965 931DA771 078C3714\n 30B5C4AC CC1F29F2 114E9070 FED294B9\n{code}\n","from":"reporter","subject":"mds file of hadoop 2.5.2 is of RC1, not of final release"},{"body":"[~kasha] this something we can fix?","from":"developer"},{"body":"My bad. Thanks [~fs111] for reporting this. Let me follow up on this. ","from":"developer"},{"body":"Sorry for the delay on this - subversion was down at the time and I forgot about it after.\n\nJust updated the mds files to refer to the right tarballs. Should go through the mirrors in 24 hours. ","from":"developer"}],"created":"2014-11-26T15:58:29.000+0000","description":"The mds file containing the checksums is referring to RC1 instead of the final release and that breaks some of our processes, since we do name matching and expect the file-name in the mds file to match the actual tarball name.\n\nSee for instance:\n\n{code}\n$ curl http://mirror.23media.de/apache/hadoop/core/hadoop-2.5.2/hadoop-2.5.2-src.tar.gz.mds\nhadoop-2.5.2-RC1-src.tar.gz: MD5 = AA 48 7A 0E 9C 9D DE 70 33 8F B4 82 E9 26\n C5 9A\nhadoop-2.5.2-RC1-src.tar.gz: SHA1 = 82FE 9B9A 580E AB73 249D 9EF7 3B8B 380C\n 43BD 7EAF\nhadoop-2.5.2-RC1-src.tar.gz: RMD160 = 54CA E137 19F8 5C39 7F92 A53E 4DA4 C70E\n 4908 CAD3\nhadoop-2.5.2-RC1-src.tar.gz: SHA224 = C9B68308 47CEC04C 9CA23491 3F33EC39\n CBF421D8 3738ACDA E40AD2C6\nhadoop-2.5.2-RC1-src.tar.gz: SHA256 = 139EF872 09C5637E 6FACA381 6CFD82CE\n 653DF2F6 1ADD240D 01C9D62E F41AD8A4\nhadoop-2.5.2-RC1-src.tar.gz: SHA384 = B2C5EA35 782BDDD9 EBFF42D7 A608F496\n 466744C3 A0D1C3A5 F212FE7B 6014D159\n 0B45B61E 930909A4 D538AC05 30336E26\nhadoop-2.5.2-RC1-src.tar.gz: SHA512 = 68CB93CC 051F8B91 ABA22913 ABD484AD\n C6D1715A EAEB2D08 829942AC 8D7D765D\n 700771D4 BEB44965 931DA771 078C3714\n 30B5C4AC CC1F29F2 114E9070 FED294B9\n{code}\n","issue_id":"12757993","key":"HADOOP-11340","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2014-12-28T19:04:12.000+0000","role":"fixed_distractor","summary":"mds file of hadoop 2.5.2 is of RC1, not of final release"} {"case_id":"12770356","cluster":"DISTRACTOR-HADOOP-11512","comments":[{"body":"Working on it.","created":"2015-01-27T14:15:27.270+0000"},{"body":"[~qwertymaniac], test may need some work adjustments. ","created":"2015-01-29T21:11:00.086+0000"},{"body":"Thanks [~Ryan P] for working on this! Can you address the changes for the test as below?\n\n1. The Configuration should be set to a specific class value, like say the WritableSerialization one, but with a prefixed or suffixed space.\n2. The SerializationFactory(conf) initialization is graceful, i.e. it will not fail but just log a WARN. So the ideal test case would need to then use the factory to fetch the serializer for a class such as the Text or LongWritable class, and ensure the return is not null.","created":"2015-01-30T02:33:28.717+0000"},{"body":"Bit confused on part 2, forgive my ignorance, but testSerializationKeyIsTrimmed() will be run by the two testGetSerializer() and testGetDeserializer() methods below. \n\nAt any rate I have added a serialization class padded with whitespace on both ends. If any further adjustments are needed please let me know. ","created":"2015-02-01T18:33:57.104+0000"},{"body":"Thanks again - the (1) change looks good, but the test still will not fail without the fix, making it incomplete. We need (2) done too to achieve it.\n\nIn literal terms, from the SerializationFactory object constructed at the end of your test, invoke its {{getSerializer(Text.class)}} or so, and assert that the returned object is not 'null'. This will exercise the serialization lookup and complete the test.\n\nThe current behaviour of constructing SerializationFactory with a bad config is that it still passes omitting the bad classes it finds in the configs (choosing to log them as WARNs than to throw an exception).\n\nLet me know if this makes sense!","created":"2015-02-02T02:32:23.961+0000"},{"body":"[~qwertymaniac], much appreciated. I'll get the hang of this eventually ","created":"2015-02-03T12:31:28.899+0000"},{"body":"Alright, so hopefully I got it this time. I apologize for all the testing mishaps.\n","created":"2015-02-08T13:42:55.895+0000"},{"body":"Thanks Ryan. I've attached a cleaned up version that removes the extra changes to indentation made on some unrelated lines.\n\n+1, will commit after Jenkins runs through this.\n\nI also attached a tested branch-2 variant as the back-port was not straight-forward.","created":"2015-02-09T05:39:16.437+0000"},{"body":"Eclipse failure appears unrelated. Checking locally before pushing anyhow. Thanks again Ryan!","created":"2015-02-10T07:07:19.543+0000"},{"body":"Local mvn eclipse:eclipse passes just fine with patch. I've pushed the patch to trunk and branch-2.\n\nResolving.","created":"2015-02-10T07:23:32.827+0000"}],"conversations":[{"body":"In the file {{hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/io/serializer/SerializationFactory.java}}, we grab the IO_SERIALIZATIONS_KEY config as Configuration#getStrings(…) which does not trim the input. This could cause confusing user issues if someone manually overrides the key in the XML files/Configuration object without using the dynamic approach.\n\nThe call should instead use Configuration#getTrimmedStrings(…), so the whitespace is trimmed before the class names are searched on the classpath.","from":"reporter","subject":"Use getTrimmedStrings when reading serialization keys"},{"body":"Working on it.","from":"developer"},{"body":"[~qwertymaniac], test may need some work adjustments. ","from":"developer"},{"body":"Thanks [~Ryan P] for working on this! Can you address the changes for the test as below?\n\n1. The Configuration should be set to a specific class value, like say the WritableSerialization one, but with a prefixed or suffixed space.\n2. The SerializationFactory(conf) initialization is graceful, i.e. it will not fail but just log a WARN. So the ideal test case would need to then use the factory to fetch the serializer for a class such as the Text or LongWritable class, and ensure the return is not null.","from":"developer"},{"body":"Bit confused on part 2, forgive my ignorance, but testSerializationKeyIsTrimmed() will be run by the two testGetSerializer() and testGetDeserializer() methods below. \n\nAt any rate I have added a serialization class padded with whitespace on both ends. If any further adjustments are needed please let me know. ","from":"developer"},{"body":"Thanks again - the (1) change looks good, but the test still will not fail without the fix, making it incomplete. We need (2) done too to achieve it.\n\nIn literal terms, from the SerializationFactory object constructed at the end of your test, invoke its {{getSerializer(Text.class)}} or so, and assert that the returned object is not 'null'. This will exercise the serialization lookup and complete the test.\n\nThe current behaviour of constructing SerializationFactory with a bad config is that it still passes omitting the bad classes it finds in the configs (choosing to log them as WARNs than to throw an exception).\n\nLet me know if this makes sense!","from":"developer"},{"body":"[~qwertymaniac], much appreciated. I'll get the hang of this eventually ","from":"developer"},{"body":"Alright, so hopefully I got it this time. I apologize for all the testing mishaps.\n","from":"developer"},{"body":"Thanks Ryan. I've attached a cleaned up version that removes the extra changes to indentation made on some unrelated lines.\n\n+1, will commit after Jenkins runs through this.\n\nI also attached a tested branch-2 variant as the back-port was not straight-forward.","from":"developer"},{"body":"Eclipse failure appears unrelated. Checking locally before pushing anyhow. Thanks again Ryan!","from":"developer"},{"body":"Local mvn eclipse:eclipse passes just fine with patch. I've pushed the patch to trunk and branch-2.\n\nResolving.","from":"developer"}],"created":"2015-01-27T14:07:12.000+0000","description":"In the file {{hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/io/serializer/SerializationFactory.java}}, we grab the IO_SERIALIZATIONS_KEY config as Configuration#getStrings(…) which does not trim the input. This could cause confusing user issues if someone manually overrides the key in the XML files/Configuration object without using the dynamic approach.\n\nThe call should instead use Configuration#getTrimmedStrings(…), so the whitespace is trimmed before the class names are searched on the classpath.","issue_id":"12770356","key":"HADOOP-11512","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2015-02-10T07:23:32.000+0000","role":"fixed_distractor","summary":"Use getTrimmedStrings when reading serialization keys"} {"case_id":"12779927","cluster":"DISTRACTOR-HADOOP-11693","comments":[{"body":"Good find Duo, this explains some throttling I've seen already!\n\n| Copy blob is very costly in Azure storage and during Azure storage gc, it will be highly likely throttled.\n\nWhy can't we just \"move\" (CopyFromBlob) it to a different container? If I'm not mistaken it is a pretty cheap linkage operation, O(1) so to say.","created":"2015-03-05T23:26:42.948+0000"},{"body":"[~thomas.jungblut]\n\nWe discussed with Azure storage team.\n\nThe problem here is during copying blob, temp tables will be created. Azure storage has cleaner threads to clean these temp tables. When the number of temp tables reaches a certain number, Azure storage will throtte the copy blob operation so that no more temp tables will be created.\n\nHowever, during Azure storage gc, cleaner is blocked, which restricts the number of copy blob operations meanwhile.\n\nIt seems not simply a linkage operation. I do not fully know the details though.","created":"2015-03-05T23:36:53.381+0000"},{"body":"[~cnauroth] Please take a look.","created":"2015-03-06T02:15:14.185+0000"},{"body":"[~onpduo] you are right- I was mistaking it with deletes of containers vs. blobs. \n\nCan you please shed some light on why we archive old WALs? I would assume we can just queue them up for deletion and delete them at a rate that doesn't cause throttling while it was splitting logs. \n\nI know this might be a very \"localized\" solution to HBase, do you see any better fix than just changing the retry backoffs?","created":"2015-03-06T17:13:18.753+0000"},{"body":"[~thomas.jungblut]\n\nHere's some words from [~enis],\n{code}\n\"There is currently two services which may keep the files in the archive dir. First is a TTL process, which ensures that the WAL files are kept at least for 10 min. This is mainly for debugging. The other one is replication. If you have replication setup, the replication processes will hang on to the WAL files until they are replicated.\n\nThere was some related discussion about directly deleting those files, but that was not implemented AFAIK. HBase assumes that rename() is a cheap operation, and uses rename for not only WAL files but for data files as well.\"\n{code}\n\nHowever, in cloud, especially in Azure storage, rename() is not a cheap operation. Currently Azure storage gc only happens on page blobs and not as frequently as you imagined. So changing retry backoffs on copyblob operation seems the only short term fix here.","created":"2015-03-06T22:53:37.731+0000"},{"body":"[~apurtell], although the issue title mentions HBase, the root cause for this problem actually resides in Hadoop Common, specifically the Azure Storage {{FileSystem}} implementation. I have updated the issue title to try to make this clearer. It looks like I don't have access to move HBASE issues, so would you mind moving this back to HADOOP? Thanks!","created":"2015-03-09T15:56:29.931+0000"},{"body":"Sure, moving it back! Thanks [~cnauroth]\n","created":"2015-03-09T16:38:52.350+0000"},{"body":"Patch looks good to me. Is copyBlob the only operation that may get ServerBusy ? ","created":"2015-03-09T21:22:48.545+0000"},{"body":"[~enis]\n\nCurrently yes. Since our WASB driver is slow due to the synchronized hsync() method when writing WALs to Azure storage, we have not seen other operations being throttled.","created":"2015-03-09T21:28:33.307+0000"},{"body":"[~cnauroth] Could you take a look?","created":"2015-03-10T19:08:15.630+0000"},{"body":"Hi, [~onpduo]. I have just one question. Right now, the patch waits to see if it encounters a {{SERVER_BUSY}} error, and then restarts the operation with the redefined retry policy. Is there any reason not to just use this retry policy right from the beginning on the initial call to {{startCopyFromBlob}}?\n\nThe patch will need to be reformatted to fit Hadoop coding conventions. We indent by 2 spaces, and we wrap lines that exceed 80 characters.\n\nThanks!","created":"2015-03-10T21:14:19.042+0000"},{"body":"[~cnauroth]\nThere is a default retry policy for all the Azure storage calls, and it is initialized with the Azure storage client.\n{code}\n private static final int DEFAULT_MIN_BACKOFF_INTERVAL = 1 * 1000; // 1s\n private static final int DEFAULT_MAX_BACKOFF_INTERVAL = 30 * 1000; // 30s\n private static final int DEFAULT_BACKOFF_INTERVAL = 1 * 1000; // 1s\n private static final int DEFAULT_MAX_RETRY_ATTEMPTS = 15;\n{code}\n\nThe backoff interval is 1s. Azure storage throttling issue caused by Azure storage GC happens very rarely. Until now only one customer met this issue and only last 10-15 mins every day. \n\nBelow is the retry policy for copyblob,\n{code}\n private static final int DEFAULT_COPYBLOB_MIN_BACKOFF_INTERVAL = 3 * 1000; // 3s\n private static final int DEFAULT_COPYBLOB_MAX_BACKOFF_INTERVAL = 90 * 1000; // 90s\n private static final int DEFAULT_COPYBLOB_BACKOFF_INTERVAL = 30 * 1000; // 30s\n private static final int DEFAULT_COPYBLOB_MAX_RETRY_ATTEMPTS = 15; \n{code}\n\nThe backoff is longer, 15s. We set these values in order to let it retry as much as 15min. We only apply this to the rare Azure storage GC case so we do not lose performance in the normal cases.\n\nI have changed the format.","created":"2015-03-11T01:06:08.013+0000"},{"body":"Thanks for the explanation about normal case vs. the new backoff policy. That makes sense.\n\nThe eclipse:eclipse failure looks unrelated. I couldn't reproduce it locally.\n\nSorry to nitpick, but there are still some lines in {{AzureNativeFileSystemStore}} that exceed the 80 character limit. I know there are some existing lines in this file that already break the rule. Don't worry about cleaning up all of the existing code, but please make sure all lines touched in the patch adhere to the 80 character limit.\n\nThe findbugs warning is legitimate. I'm not sure why it's triggering now with this patch, as it appears the problem existed before the patch. We can fix this by changing the {{catch (Exception e)}} so that there are 2 separate catch clauses for {{catch (StorageException e)}} and {{catch (URISyntaxException e)}}. Each one can be rethrown wrapped as an {{AzureException}}.\n\nWe're almost there. Thanks, Duo!","created":"2015-03-11T18:21:56.373+0000"},{"body":"Please disregard the mention of an eclipse:eclipse failure in my last comment. That comment was meant for a different patch. Sorry about that.","created":"2015-03-11T18:30:31.717+0000"},{"body":"[~cnauroth]\n\nI have wrapped those lines execeeding 80 characters, and indent by 2 characters. ","created":"2015-03-11T20:19:03.225+0000"},{"body":"I have committed this to trunk, branch-2 and branch-2.7. Duo, thank you for contributing the patch and incorporating the feedback. Thomas and Enis, thank you for helping with code review.","created":"2015-03-11T21:45:55.823+0000"},{"body":"Hi [~cnauroth]\n\nI have submitted a new patch, which does the retries in WASB rather than rely on Azure Storage SDK. As I looked into the source code this week, Azure Storage SDK regards storage exception as non-retryable, so when throttling happens, the current code might still not work.\n\nCould you reopen this JIRA and review it ASAP?\n\nThanks.","created":"2015-06-05T00:38:55.089+0000"},{"body":"Hello [~onpduo]. The HADOOP-11693 patch already shipped in Apache Hadoop 2.7.0. At this point, please create a new jira to track the new change instead of attaching new patches here.\n\nBTW, I noticed that this new patch appears to be in a multi-byte character encoding (UTF-16 with BOM?). When you create the new jira, please attach an ASCII patch file.","created":"2015-06-05T02:44:27.229+0000"}],"conversations":[{"body":"One of our customers' production HBase clusters was periodically throttled by Azure storage, when HBase was archiving old WALs. HMaster aborted the region server and tried to restart it.\n\nHowever, since the cluster was still being throttled by Azure storage, the upcoming distributed log splitting also failed. Sometimes hbase:meta table was on this region server and finally showed offline, which cause the whole cluster in bad state.\n\n{code}\n2015-03-01 18:36:45,623 ERROR org.apache.hadoop.hbase.master.HMaster: Region server workernode4.hbaseproddb4001.f5.internal.cloudapp.net,60020,1424845421044 reported a fatal error:\nABORTING region server workernode4.hbaseproddb4001.f5.internal.cloudapp.net,60020,1424845421044: IOE in log roller\nCause:\norg.apache.hadoop.fs.azure.AzureException: com.microsoft.windowsazure.storage.StorageException: The server is busy.\n\tat org.apache.hadoop.fs.azurenative.AzureNativeFileSystemStore.rename(AzureNativeFileSystemStore.java:2446)\n\tat org.apache.hadoop.fs.azurenative.AzureNativeFileSystemStore.rename(AzureNativeFileSystemStore.java:2367)\n\tat org.apache.hadoop.fs.azurenative.NativeAzureFileSystem.rename(NativeAzureFileSystem.java:1960)\n\tat org.apache.hadoop.hbase.util.FSUtils.renameAndSetModifyTime(FSUtils.java:1719)\n\tat org.apache.hadoop.hbase.regionserver.wal.FSHLog.archiveLogFile(FSHLog.java:798)\n\tat org.apache.hadoop.hbase.regionserver.wal.FSHLog.cleanOldLogs(FSHLog.java:656)\n\tat org.apache.hadoop.hbase.regionserver.wal.FSHLog.rollWriter(FSHLog.java:593)\n\tat org.apache.hadoop.hbase.regionserver.LogRoller.run(LogRoller.java:97)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: com.microsoft.windowsazure.storage.StorageException: The server is busy.\n\tat com.microsoft.windowsazure.storage.StorageException.translateException(StorageException.java:163)\n\tat com.microsoft.windowsazure.storage.core.StorageRequest.materializeException(StorageRequest.java:306)\n\tat com.microsoft.windowsazure.storage.core.ExecutionEngine.executeWithRetry(ExecutionEngine.java:229)\n\tat com.microsoft.windowsazure.storage.blob.CloudBlob.startCopyFromBlob(CloudBlob.java:762)\n\tat org.apache.hadoop.fs.azurenative.StorageInterfaceImpl$CloudBlobWrapperImpl.startCopyFromBlob(StorageInterfaceImpl.java:350)\n\tat org.apache.hadoop.fs.azurenative.AzureNativeFileSystemStore.rename(AzureNativeFileSystemStore.java:2439)\n\t... 8 more\n\n2015-03-01 18:43:29,072 ERROR org.apache.hadoop.hbase.executor.EventHandler: Caught throwable while processing event M_META_SERVER_SHUTDOWN\njava.io.IOException: failed log splitting for workernode13.hbaseproddb4001.f5.internal.cloudapp.net,60020,1424845307901, will retry\n\tat org.apache.hadoop.hbase.master.handler.MetaServerShutdownHandler.process(MetaServerShutdownHandler.java:71)\n\tat org.apache.hadoop.hbase.executor.EventHandler.run(EventHandler.java:128)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.hadoop.fs.azure.AzureException: com.microsoft.windowsazure.storage.StorageException: The server is busy.\n\tat org.apache.hadoop.fs.azurenative.AzureNativeFileSystemStore.rename(AzureNativeFileSystemStore.java:2446)\n\tat org.apache.hadoop.fs.azurenative.NativeAzureFileSystem$FolderRenamePending.execute(NativeAzureFileSystem.java:393)\n\tat org.apache.hadoop.fs.azurenative.NativeAzureFileSystem.rename(NativeAzureFileSystem.java:1973)\n\tat org.apache.hadoop.hbase.master.MasterFileSystem.getLogDirs(MasterFileSystem.java:319)\n\tat org.apache.hadoop.hbase.master.MasterFileSystem.splitLog(MasterFileSystem.java:406)\n\tat org.apache.hadoop.hbase.master.MasterFileSystem.splitMetaLog(MasterFileSystem.java:302)\n\tat org.apache.hadoop.hbase.master.MasterFileSystem.splitMetaLog(MasterFileSystem.java:293)\n\tat org.apache.hadoop.hbase.master.handler.MetaServerShutdownHandler.process(MetaServerShutdownHandler.java:64)\n\t... 4 more\nCaused by: com.microsoft.windowsazure.storage.StorageException: The server is busy.\n\tat com.microsoft.windowsazure.storage.StorageException.translateException(StorageException.java:163)\n\tat com.microsoft.windowsazure.storage.core.StorageRequest.materializeException(StorageRequest.java:306)\n\tat com.microsoft.windowsazure.storage.core.ExecutionEngine.executeWithRetry(ExecutionEngine.java:229)\n\tat com.microsoft.windowsazure.storage.blob.CloudBlob.startCopyFromBlob(CloudBlob.java:762)\n\tat org.apache.hadoop.fs.azurenative.StorageInterfaceImpl$CloudBlobWrapperImpl.startCopyFromBlob(StorageInterfaceImpl.java:350)\n\tat org.apache.hadoop.fs.azurenative.AzureNativeFileSystemStore.rename(AzureNativeFileSystemStore.java:2439)\n\t... 11 more\n\nSun Mar 01 18:59:51 GMT 2015, org.apache.hadoop.hbase.client.RpcRetryingCaller@aa93ac7, org.apache.hadoop.hbase.NotServingRegionException: org.apache.hadoop.hbase.NotServingRegionException: Region hbase:meta,,1 is not online on workernode13.hbaseproddb4001.f5.internal.cloudapp.net,60020,1425235081338\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.getRegionByEncodedName(HRegionServer.java:2676)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.getRegion(HRegionServer.java:4095)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.scan(HRegionServer.java:3076)\n\tat org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:28861)\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2008)\n\tat org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:92)\n\tat org.apache.hadoop.hbase.ipc.SimpleRpcScheduler.consumerLoop(SimpleRpcScheduler.java:160)\n\tat org.apache.hadoop.hbase.ipc.SimpleRpcScheduler.access$000(SimpleRpcScheduler.java:38)\n\tat org.apache.hadoop.hbase.ipc.SimpleRpcScheduler$1.run(SimpleRpcScheduler.java:110)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}\n\nWhen archiving old WALs, WASB will do rename operation by copying src blob to destination blob and deleting the src blob. Copy blob is very costly in Azure storage and during Azure storage gc, it will be highly likely throttled. The throttling by Azure storage usually ends within 15mins. Current WASB retry policy is exponential retry, but only last at most for 2min. Short term fix will be adding a more intensive exponential retry when copy blob is throttled.","from":"reporter","subject":"Azure Storage FileSystem rename operations are throttled too aggressively to complete HBase WAL archiving."},{"body":"Good find Duo, this explains some throttling I've seen already!\n\n| Copy blob is very costly in Azure storage and during Azure storage gc, it will be highly likely throttled.\n\nWhy can't we just \"move\" (CopyFromBlob) it to a different container? If I'm not mistaken it is a pretty cheap linkage operation, O(1) so to say.","from":"developer"},{"body":"[~thomas.jungblut]\n\nWe discussed with Azure storage team.\n\nThe problem here is during copying blob, temp tables will be created. Azure storage has cleaner threads to clean these temp tables. When the number of temp tables reaches a certain number, Azure storage will throtte the copy blob operation so that no more temp tables will be created.\n\nHowever, during Azure storage gc, cleaner is blocked, which restricts the number of copy blob operations meanwhile.\n\nIt seems not simply a linkage operation. I do not fully know the details though.","from":"developer"},{"body":"[~cnauroth] Please take a look.","from":"developer"},{"body":"[~onpduo] you are right- I was mistaking it with deletes of containers vs. blobs. \n\nCan you please shed some light on why we archive old WALs? I would assume we can just queue them up for deletion and delete them at a rate that doesn't cause throttling while it was splitting logs. \n\nI know this might be a very \"localized\" solution to HBase, do you see any better fix than just changing the retry backoffs?","from":"developer"},{"body":"[~thomas.jungblut]\n\nHere's some words from [~enis],\n{code}\n\"There is currently two services which may keep the files in the archive dir. First is a TTL process, which ensures that the WAL files are kept at least for 10 min. This is mainly for debugging. The other one is replication. If you have replication setup, the replication processes will hang on to the WAL files until they are replicated.\n\nThere was some related discussion about directly deleting those files, but that was not implemented AFAIK. HBase assumes that rename() is a cheap operation, and uses rename for not only WAL files but for data files as well.\"\n{code}\n\nHowever, in cloud, especially in Azure storage, rename() is not a cheap operation. Currently Azure storage gc only happens on page blobs and not as frequently as you imagined. So changing retry backoffs on copyblob operation seems the only short term fix here.","from":"developer"},{"body":"[~apurtell], although the issue title mentions HBase, the root cause for this problem actually resides in Hadoop Common, specifically the Azure Storage {{FileSystem}} implementation. I have updated the issue title to try to make this clearer. It looks like I don't have access to move HBASE issues, so would you mind moving this back to HADOOP? Thanks!","from":"developer"},{"body":"Sure, moving it back! Thanks [~cnauroth]\n","from":"developer"},{"body":"Patch looks good to me. Is copyBlob the only operation that may get ServerBusy ? ","from":"developer"},{"body":"[~enis]\n\nCurrently yes. Since our WASB driver is slow due to the synchronized hsync() method when writing WALs to Azure storage, we have not seen other operations being throttled.","from":"developer"},{"body":"[~cnauroth] Could you take a look?","from":"developer"},{"body":"Hi, [~onpduo]. I have just one question. Right now, the patch waits to see if it encounters a {{SERVER_BUSY}} error, and then restarts the operation with the redefined retry policy. Is there any reason not to just use this retry policy right from the beginning on the initial call to {{startCopyFromBlob}}?\n\nThe patch will need to be reformatted to fit Hadoop coding conventions. We indent by 2 spaces, and we wrap lines that exceed 80 characters.\n\nThanks!","from":"developer"},{"body":"[~cnauroth]\nThere is a default retry policy for all the Azure storage calls, and it is initialized with the Azure storage client.\n{code}\n private static final int DEFAULT_MIN_BACKOFF_INTERVAL = 1 * 1000; // 1s\n private static final int DEFAULT_MAX_BACKOFF_INTERVAL = 30 * 1000; // 30s\n private static final int DEFAULT_BACKOFF_INTERVAL = 1 * 1000; // 1s\n private static final int DEFAULT_MAX_RETRY_ATTEMPTS = 15;\n{code}\n\nThe backoff interval is 1s. Azure storage throttling issue caused by Azure storage GC happens very rarely. Until now only one customer met this issue and only last 10-15 mins every day. \n\nBelow is the retry policy for copyblob,\n{code}\n private static final int DEFAULT_COPYBLOB_MIN_BACKOFF_INTERVAL = 3 * 1000; // 3s\n private static final int DEFAULT_COPYBLOB_MAX_BACKOFF_INTERVAL = 90 * 1000; // 90s\n private static final int DEFAULT_COPYBLOB_BACKOFF_INTERVAL = 30 * 1000; // 30s\n private static final int DEFAULT_COPYBLOB_MAX_RETRY_ATTEMPTS = 15; \n{code}\n\nThe backoff is longer, 15s. We set these values in order to let it retry as much as 15min. We only apply this to the rare Azure storage GC case so we do not lose performance in the normal cases.\n\nI have changed the format.","from":"developer"},{"body":"Thanks for the explanation about normal case vs. the new backoff policy. That makes sense.\n\nThe eclipse:eclipse failure looks unrelated. I couldn't reproduce it locally.\n\nSorry to nitpick, but there are still some lines in {{AzureNativeFileSystemStore}} that exceed the 80 character limit. I know there are some existing lines in this file that already break the rule. Don't worry about cleaning up all of the existing code, but please make sure all lines touched in the patch adhere to the 80 character limit.\n\nThe findbugs warning is legitimate. I'm not sure why it's triggering now with this patch, as it appears the problem existed before the patch. We can fix this by changing the {{catch (Exception e)}} so that there are 2 separate catch clauses for {{catch (StorageException e)}} and {{catch (URISyntaxException e)}}. Each one can be rethrown wrapped as an {{AzureException}}.\n\nWe're almost there. Thanks, Duo!","from":"developer"},{"body":"Please disregard the mention of an eclipse:eclipse failure in my last comment. That comment was meant for a different patch. Sorry about that.","from":"developer"},{"body":"[~cnauroth]\n\nI have wrapped those lines execeeding 80 characters, and indent by 2 characters. ","from":"developer"},{"body":"I have committed this to trunk, branch-2 and branch-2.7. Duo, thank you for contributing the patch and incorporating the feedback. Thomas and Enis, thank you for helping with code review.","from":"developer"},{"body":"Hi [~cnauroth]\n\nI have submitted a new patch, which does the retries in WASB rather than rely on Azure Storage SDK. As I looked into the source code this week, Azure Storage SDK regards storage exception as non-retryable, so when throttling happens, the current code might still not work.\n\nCould you reopen this JIRA and review it ASAP?\n\nThanks.","from":"developer"},{"body":"Hello [~onpduo]. The HADOOP-11693 patch already shipped in Apache Hadoop 2.7.0. At this point, please create a new jira to track the new change instead of attaching new patches here.\n\nBTW, I noticed that this new patch appears to be in a multi-byte character encoding (UTF-16 with BOM?). When you create the new jira, please attach an ASCII patch file.","from":"developer"}],"created":"2015-03-05T23:19:13.000+0000","description":"One of our customers' production HBase clusters was periodically throttled by Azure storage, when HBase was archiving old WALs. HMaster aborted the region server and tried to restart it.\n\nHowever, since the cluster was still being throttled by Azure storage, the upcoming distributed log splitting also failed. Sometimes hbase:meta table was on this region server and finally showed offline, which cause the whole cluster in bad state.\n\n{code}\n2015-03-01 18:36:45,623 ERROR org.apache.hadoop.hbase.master.HMaster: Region server workernode4.hbaseproddb4001.f5.internal.cloudapp.net,60020,1424845421044 reported a fatal error:\nABORTING region server workernode4.hbaseproddb4001.f5.internal.cloudapp.net,60020,1424845421044: IOE in log roller\nCause:\norg.apache.hadoop.fs.azure.AzureException: com.microsoft.windowsazure.storage.StorageException: The server is busy.\n\tat org.apache.hadoop.fs.azurenative.AzureNativeFileSystemStore.rename(AzureNativeFileSystemStore.java:2446)\n\tat org.apache.hadoop.fs.azurenative.AzureNativeFileSystemStore.rename(AzureNativeFileSystemStore.java:2367)\n\tat org.apache.hadoop.fs.azurenative.NativeAzureFileSystem.rename(NativeAzureFileSystem.java:1960)\n\tat org.apache.hadoop.hbase.util.FSUtils.renameAndSetModifyTime(FSUtils.java:1719)\n\tat org.apache.hadoop.hbase.regionserver.wal.FSHLog.archiveLogFile(FSHLog.java:798)\n\tat org.apache.hadoop.hbase.regionserver.wal.FSHLog.cleanOldLogs(FSHLog.java:656)\n\tat org.apache.hadoop.hbase.regionserver.wal.FSHLog.rollWriter(FSHLog.java:593)\n\tat org.apache.hadoop.hbase.regionserver.LogRoller.run(LogRoller.java:97)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: com.microsoft.windowsazure.storage.StorageException: The server is busy.\n\tat com.microsoft.windowsazure.storage.StorageException.translateException(StorageException.java:163)\n\tat com.microsoft.windowsazure.storage.core.StorageRequest.materializeException(StorageRequest.java:306)\n\tat com.microsoft.windowsazure.storage.core.ExecutionEngine.executeWithRetry(ExecutionEngine.java:229)\n\tat com.microsoft.windowsazure.storage.blob.CloudBlob.startCopyFromBlob(CloudBlob.java:762)\n\tat org.apache.hadoop.fs.azurenative.StorageInterfaceImpl$CloudBlobWrapperImpl.startCopyFromBlob(StorageInterfaceImpl.java:350)\n\tat org.apache.hadoop.fs.azurenative.AzureNativeFileSystemStore.rename(AzureNativeFileSystemStore.java:2439)\n\t... 8 more\n\n2015-03-01 18:43:29,072 ERROR org.apache.hadoop.hbase.executor.EventHandler: Caught throwable while processing event M_META_SERVER_SHUTDOWN\njava.io.IOException: failed log splitting for workernode13.hbaseproddb4001.f5.internal.cloudapp.net,60020,1424845307901, will retry\n\tat org.apache.hadoop.hbase.master.handler.MetaServerShutdownHandler.process(MetaServerShutdownHandler.java:71)\n\tat org.apache.hadoop.hbase.executor.EventHandler.run(EventHandler.java:128)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.hadoop.fs.azure.AzureException: com.microsoft.windowsazure.storage.StorageException: The server is busy.\n\tat org.apache.hadoop.fs.azurenative.AzureNativeFileSystemStore.rename(AzureNativeFileSystemStore.java:2446)\n\tat org.apache.hadoop.fs.azurenative.NativeAzureFileSystem$FolderRenamePending.execute(NativeAzureFileSystem.java:393)\n\tat org.apache.hadoop.fs.azurenative.NativeAzureFileSystem.rename(NativeAzureFileSystem.java:1973)\n\tat org.apache.hadoop.hbase.master.MasterFileSystem.getLogDirs(MasterFileSystem.java:319)\n\tat org.apache.hadoop.hbase.master.MasterFileSystem.splitLog(MasterFileSystem.java:406)\n\tat org.apache.hadoop.hbase.master.MasterFileSystem.splitMetaLog(MasterFileSystem.java:302)\n\tat org.apache.hadoop.hbase.master.MasterFileSystem.splitMetaLog(MasterFileSystem.java:293)\n\tat org.apache.hadoop.hbase.master.handler.MetaServerShutdownHandler.process(MetaServerShutdownHandler.java:64)\n\t... 4 more\nCaused by: com.microsoft.windowsazure.storage.StorageException: The server is busy.\n\tat com.microsoft.windowsazure.storage.StorageException.translateException(StorageException.java:163)\n\tat com.microsoft.windowsazure.storage.core.StorageRequest.materializeException(StorageRequest.java:306)\n\tat com.microsoft.windowsazure.storage.core.ExecutionEngine.executeWithRetry(ExecutionEngine.java:229)\n\tat com.microsoft.windowsazure.storage.blob.CloudBlob.startCopyFromBlob(CloudBlob.java:762)\n\tat org.apache.hadoop.fs.azurenative.StorageInterfaceImpl$CloudBlobWrapperImpl.startCopyFromBlob(StorageInterfaceImpl.java:350)\n\tat org.apache.hadoop.fs.azurenative.AzureNativeFileSystemStore.rename(AzureNativeFileSystemStore.java:2439)\n\t... 11 more\n\nSun Mar 01 18:59:51 GMT 2015, org.apache.hadoop.hbase.client.RpcRetryingCaller@aa93ac7, org.apache.hadoop.hbase.NotServingRegionException: org.apache.hadoop.hbase.NotServingRegionException: Region hbase:meta,,1 is not online on workernode13.hbaseproddb4001.f5.internal.cloudapp.net,60020,1425235081338\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.getRegionByEncodedName(HRegionServer.java:2676)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.getRegion(HRegionServer.java:4095)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.scan(HRegionServer.java:3076)\n\tat org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:28861)\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2008)\n\tat org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:92)\n\tat org.apache.hadoop.hbase.ipc.SimpleRpcScheduler.consumerLoop(SimpleRpcScheduler.java:160)\n\tat org.apache.hadoop.hbase.ipc.SimpleRpcScheduler.access$000(SimpleRpcScheduler.java:38)\n\tat org.apache.hadoop.hbase.ipc.SimpleRpcScheduler$1.run(SimpleRpcScheduler.java:110)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}\n\nWhen archiving old WALs, WASB will do rename operation by copying src blob to destination blob and deleting the src blob. Copy blob is very costly in Azure storage and during Azure storage gc, it will be highly likely throttled. The throttling by Azure storage usually ends within 15mins. Current WASB retry policy is exponential retry, but only last at most for 2min. Short term fix will be adding a more intensive exponential retry when copy blob is throttled.","issue_id":"12779927","key":"HADOOP-11693","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2015-03-11T21:45:55.000+0000","role":"fixed_distractor","summary":"Azure Storage FileSystem rename operations are throttled too aggressively to complete HBase WAL archiving."} {"case_id":"12331625","cluster":"DISTRACTOR-HADOOP-120","comments":[{"body":"This patch appears to fix the problem. I don't know whether you will object to the exception modification on ClassNotFoundException .\n\n-dk\n","created":"2006-04-05T07:58:24.000+0000"},{"body":"After thinking about it for a bit, the problem with this patch is that this is going to encode the typename in each and every record. So if your value type is ArrayWriter, you are going to spend an extra 2+strlen(\"org.apache.hadoop.io.UTF8\") bytes per a record. That's a fair amount of overhead.\n\nWe also have to be a careful with the serialization of ArrayWritable because it is used in the DFS name node logs.\n\nI'm not sure what the right solution is. Probably for right now, I would derive a subclass of ArrayWritable that is specific for your type. It isn't pretty, but it is guaranteed to be safe.\n\npublic class UTF8Array extends ArrayWritable {\n public UTF8Array() {\n super(UTF8.class);\n }\n}","created":"2006-04-05T13:05:09.000+0000"},{"body":"I realized that, of course, but I'm not expecting a whole lot of small ArrayWritable's .\n\nOwen O'Malley's solution has the problem that ArrayWritables can themselves contain ArrayWritables. We might, for example, want to have a two-component ArrayWritable where there are some integers and some UTF8s, so we need a class\n\n public class ArrayOfArrayOfSomeIntsAndSomeUTF8s extends ArrayWritable\n {\n ... filling this in gives me a headache\n }\n\n-dk\n","created":"2006-04-05T14:45:21.000+0000"},{"body":"Regarding the class name encoding: you could use the same trick that we use in Nutch, org.apache.nutch.crawl.MapWritable, which uses a sort of dictionary encoding, and as long as you use only the \"standard\" types the overhead is just 1 byte. For non-standard types, the type name is put into a dictionary (once), and henceforth only 1 byte is used, too.","created":"2006-04-05T15:13:12.000+0000"},{"body":"The big deal is probably not so much the space it takes but the TIME it takes to look up a class.\n\nOwen, I'll stop by at 11 to talk about this. I think we need to generalize the notion of an input ot output class to an input or output _descriptor_ .\n\nGotta run now -- literally.\n\n-dk\n","created":"2006-04-05T23:28:30.000+0000"},{"body":"Regarding the dictionary approach ... remember the life of a dictionary is the lifetime of a DFS, not merely a particular run or a DFS file. Also it does violate the apparent convention that the callER knows the type of the object. Now maybe that should be rethought, but we should do so explicitly, not back into it.\n\n-dk\n","created":"2006-04-05T23:31:07.000+0000"},{"body":"Dictionary is stored together inside the file, so I don't think it's a problem - you can always read it later and restore the data.\n\nCaller may be unaware of exact types stored in a Map-like structure, but it should not prevent him from knowing (and using) portions of data that he knows about, so I think this doesn't violate the convention... ok, maybe bends it a little ;)","created":"2006-04-06T02:25:40.000+0000"},{"body":"Hello all,\n\nwhat is current thinking about this bug and the right way to fix it?\n\nWhat about adding a simple javadoc comment in this class to warn users that it must be derived to be used as a Reducer value? ","created":"2006-07-18T09:42:53.000+0000"},{"body":"Hello,\nThis bug is still not fixed in the latest versions of Hadoop. I wrote a fix similar to the one submitted. But it hasnt been updated in the respository.\nIs there some issue with this fix ?","created":"2007-05-11T15:08:59.002+0000"},{"body":"Could anyone commit this patch ?","created":"2007-05-11T15:30:20.612+0000"},{"body":"This patch no longer applies to trunk. It's also not formatted according to Hadoop's conventions. Finally, this is not an acceptable approach. Rather, one could use the workaround suggested by Owen, or we could add a new Writable class that writes the name of the class in each instance, but it would not be back-compatible to make this change to ArrayWritable.","created":"2007-05-11T18:47:06.296+0000"},{"body":"I ran into this issue this morning. I'm new to Hadoop so forgive me if I failed to understand the big picture in submitting this patch.\n\nTo address, I added some documentation to ArrayWritable indicating that it needs to be subclassed to be used as input to Reducers. I also added a check for valueClass being undefined when calling readFields() - as a relatively new Hadoop user, it took me a while to figure out what the NPE I got in ReflectionUtils was and what to do with it. \n\nIt looks like some DFS datastructures depend on ArrayWritable being the way that it is, so it seems that changing the serialization format would break stuff. It might make sense to create a different ArrayWritable that serialized type information for Hadoop users specifically for MapReduce applications, but this class should probably document its contract with the outside world and throw a specific exception when it's violated. I've also added a unit test to verify the exceptions thrown or not thrown are as expected.","created":"2007-09-07T19:42:39.145+0000"},{"body":"Patch to document ArrayWritable and throw a specific exception when readFields is undefine.","created":"2007-09-07T19:43:30.051+0000"},{"body":"I think we can do this more simply by prohibiting the creation of ArrayWritable's whose valueClass is null. I'll attach a patch.","created":"2007-09-13T20:04:58.235+0000"},{"body":"Here's the new version that prohibits creation of ArrayWritable's with a null valueClass.","created":"2007-09-13T20:10:44.788+0000"},{"body":"Yes, that's definitely a better fix and if anyone is using ArrayWritable's default constructor, it should be easy to change.","created":"2007-09-13T20:24:22.055+0000"},{"body":"I just committed this. Thanks, Cameron.","created":"2007-09-18T18:26:19.889+0000"}],"conversations":[{"body":"If you have a Reducer whose value type is an ArrayWriter it gets enstreamed alright but at reconstruction type when ArrayWriter::readFields(DataInput in) runs on a DataInput that has a nonempty ArrayWriter , newInstance fails trying to instantiate the null class.","from":"reporter","subject":"Reading an ArrayWriter does not work because valueClass does not get initialized"},{"body":"This patch appears to fix the problem. I don't know whether you will object to the exception modification on ClassNotFoundException .\n\n-dk\n","from":"developer"},{"body":"After thinking about it for a bit, the problem with this patch is that this is going to encode the typename in each and every record. So if your value type is ArrayWriter, you are going to spend an extra 2+strlen(\"org.apache.hadoop.io.UTF8\") bytes per a record. That's a fair amount of overhead.\n\nWe also have to be a careful with the serialization of ArrayWritable because it is used in the DFS name node logs.\n\nI'm not sure what the right solution is. Probably for right now, I would derive a subclass of ArrayWritable that is specific for your type. It isn't pretty, but it is guaranteed to be safe.\n\npublic class UTF8Array extends ArrayWritable {\n public UTF8Array() {\n super(UTF8.class);\n }\n}","from":"developer"},{"body":"I realized that, of course, but I'm not expecting a whole lot of small ArrayWritable's .\n\nOwen O'Malley's solution has the problem that ArrayWritables can themselves contain ArrayWritables. We might, for example, want to have a two-component ArrayWritable where there are some integers and some UTF8s, so we need a class\n\n public class ArrayOfArrayOfSomeIntsAndSomeUTF8s extends ArrayWritable\n {\n ... filling this in gives me a headache\n }\n\n-dk\n","from":"developer"},{"body":"Regarding the class name encoding: you could use the same trick that we use in Nutch, org.apache.nutch.crawl.MapWritable, which uses a sort of dictionary encoding, and as long as you use only the \"standard\" types the overhead is just 1 byte. For non-standard types, the type name is put into a dictionary (once), and henceforth only 1 byte is used, too.","from":"developer"},{"body":"The big deal is probably not so much the space it takes but the TIME it takes to look up a class.\n\nOwen, I'll stop by at 11 to talk about this. I think we need to generalize the notion of an input ot output class to an input or output _descriptor_ .\n\nGotta run now -- literally.\n\n-dk\n","from":"developer"},{"body":"Regarding the dictionary approach ... remember the life of a dictionary is the lifetime of a DFS, not merely a particular run or a DFS file. Also it does violate the apparent convention that the callER knows the type of the object. Now maybe that should be rethought, but we should do so explicitly, not back into it.\n\n-dk\n","from":"developer"},{"body":"Dictionary is stored together inside the file, so I don't think it's a problem - you can always read it later and restore the data.\n\nCaller may be unaware of exact types stored in a Map-like structure, but it should not prevent him from knowing (and using) portions of data that he knows about, so I think this doesn't violate the convention... ok, maybe bends it a little ;)","from":"developer"},{"body":"Hello all,\n\nwhat is current thinking about this bug and the right way to fix it?\n\nWhat about adding a simple javadoc comment in this class to warn users that it must be derived to be used as a Reducer value? ","from":"developer"},{"body":"Hello,\nThis bug is still not fixed in the latest versions of Hadoop. I wrote a fix similar to the one submitted. But it hasnt been updated in the respository.\nIs there some issue with this fix ?","from":"developer"},{"body":"Could anyone commit this patch ?","from":"developer"},{"body":"This patch no longer applies to trunk. It's also not formatted according to Hadoop's conventions. Finally, this is not an acceptable approach. Rather, one could use the workaround suggested by Owen, or we could add a new Writable class that writes the name of the class in each instance, but it would not be back-compatible to make this change to ArrayWritable.","from":"developer"},{"body":"I ran into this issue this morning. I'm new to Hadoop so forgive me if I failed to understand the big picture in submitting this patch.\n\nTo address, I added some documentation to ArrayWritable indicating that it needs to be subclassed to be used as input to Reducers. I also added a check for valueClass being undefined when calling readFields() - as a relatively new Hadoop user, it took me a while to figure out what the NPE I got in ReflectionUtils was and what to do with it. \n\nIt looks like some DFS datastructures depend on ArrayWritable being the way that it is, so it seems that changing the serialization format would break stuff. It might make sense to create a different ArrayWritable that serialized type information for Hadoop users specifically for MapReduce applications, but this class should probably document its contract with the outside world and throw a specific exception when it's violated. I've also added a unit test to verify the exceptions thrown or not thrown are as expected.","from":"developer"},{"body":"Patch to document ArrayWritable and throw a specific exception when readFields is undefine.","from":"developer"},{"body":"I think we can do this more simply by prohibiting the creation of ArrayWritable's whose valueClass is null. I'll attach a patch.","from":"developer"},{"body":"Here's the new version that prohibits creation of ArrayWritable's with a null valueClass.","from":"developer"},{"body":"Yes, that's definitely a better fix and if anyone is using ArrayWritable's default constructor, it should be easy to change.","from":"developer"},{"body":"I just committed this. Thanks, Cameron.","from":"developer"}],"created":"2006-04-05T07:54:37.000+0000","description":"If you have a Reducer whose value type is an ArrayWriter it gets enstreamed alright but at reconstruction type when ArrayWriter::readFields(DataInput in) runs on a DataInput that has a nonempty ArrayWriter , newInstance fails trying to instantiate the null class.","issue_id":"12331625","key":"HADOOP-120","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2007-09-18T18:26:19.000+0000","role":"fixed_distractor","summary":"Reading an ArrayWriter does not work because valueClass does not get initialized"} {"case_id":"12863593","cluster":"DISTRACTOR-HADOOP-12406","comments":[{"body":"Hey Everyone: \n\nThis is related to a problem we are seeing in NUTCH, NUTCH-1084. Any chance someone could review. It looks like the test failures are actually orthogonal to this patch.","created":"2015-09-12T13:55:18.394+0000"},{"body":"[~ndouba], please use the \"Target Version\" field to express your intention. Fix-version is exclusively used by committers when a patch gets committed.\n\nIAC, 2.7.1 is done, targeting this for 2.7.2/2.8.0.","created":"2015-09-28T18:39:22.171+0000"},{"body":"[~ndouba], is this happening inside a MapReduce job running on top of YARN? MapReduce does have job.jar in the system classpath, so that's not explainable.\n\nThe only way this may happen is if you are using the MapReduce JobClassLoader for this - please let us know the values for mapreduce.job.classloader in your job.","created":"2015-10-27T00:00:43.020+0000"},{"body":"Moving this out of 2.7.2 as there's been no update in a while.","created":"2015-11-02T21:20:30.194+0000"},{"body":"Hi Vinod,\n\nI was running this on top of YARN with the default configuration enabled on top of Java 1.8 from Oracle. I think the new isolation strategy used in the newer versions of the JVM for class loaders prevents the system class loader from loading a job jar class. I didn't do anything fancy with configuration options or anything else. The only thing I added were the remote debugging options to allow me to attach to the map and reduce threads remotely and debug them step by step. I can tell you that the job jar definitely isn't in the system's class loader but in the current thread's context class loader. I verified this observation by manually inspecting the class loader object from the system and the one that comes with the thread context.\n\nCheers,\n\nNadeem","created":"2015-11-02T23:45:36.178+0000"},{"body":"Hi [~ndouba],\n\nI'm about to do a 2.7.3 Apache Hadoop release and finally got around to this again.\n\nh4. Analysis\nTo make progress, I had to read up a bit on nutch and about how to run this so that I can reproduce the bug in order to rationalize your patch. I finally succeeded in doing so! Tested this with 2.7.2 release and nutch 1.11 and using the URL feed [given at NUTCH-1084|https://issues.apache.org/jira/browse/NUTCH-1084?focusedCommentId=13882771&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-13882771]\n{code}\n~/tmp/common/hadoop-common-2.7.2/bin/hadoop jar apache-nutch-1.11.job org.apache.nutch.crawl.CrawlDbReader file:///tmp/nutch/apache-nutch-1.11/runtime/local/crawl/crawldb/ -url http://bappenas.go.id/\n{code}\n\nI can reproduce all the problems listed at NUTCH-1084 - with readdb, MR local-job-runner based job for crawling etc.\n\nThe real issue is that Nutch's readdb is client-only and *not* running a MapReduce job which was my question before. For regular MR jobs, the job-jar *is* on the system class-loader. For the client-only invocations using \"hadoop jar\" and local-job-runner, the job-jar is actually *not* on the system-classpath - that is why you are running into the issue.\n\nh4. Summary\nYour patch looks good to me. Clearly, the thread context-loader falls back to system class-loader where it is not overridden - so we are fine for all the ways of loading the classes in readFields.\n\nI'll resubmit your patch with minor commenting related changes to Jenkins and commit if Mr.Jenkins is also fine.","created":"2016-04-10T20:36:41.850+0000"},{"body":"Updated patch fixing the code comment..","created":"2016-04-10T20:41:44.061+0000"},{"body":"Hi Vinod and Nadeem, thanks for taking care of this. At Apache Nutch, we're looking forward to 2.7.3!","created":"2016-04-11T09:05:33.271+0000"},{"body":"TestReloadingX509TrustManager failure is not related, I'll see if there is an existing JIRA.\n\nChecking this in now..","created":"2016-04-11T18:59:06.538+0000"},{"body":"Committed this to trunk, branch-2 and branch-2.7. Thanks [~ndouba]!\n\nForgot to mention that I've tested that the nutch code fails without the patch and passes with.","created":"2016-04-11T19:06:48.369+0000"},{"body":"Closing the JIRA as part of 2.7.3 release.","created":"2016-08-25T22:48:16.995+0000"},{"body":"Merge the fix to branch-2.8.","created":"2017-01-05T21:25:25.169+0000"}],"conversations":[{"body":"Note: I am not an expert at JAVA, Class loaders, or Hadoop. I am just a hacker. My solution might be entirely wrong.\n\nAbstractMapWritable.readFields throws a ClassNotFoundException when reading custom writables. Debugging the job using remote debugging in IntelliJ revealed that the class loader being used in Class.forName() is different than that used by the Thread's current context (Thread.currentThread().getContextClassLoader()). The class path for the system class loader does not include the libraries of the job jar. However, the class path for the context class loader does. The proposed patch changes the class loading mechanism in readFields to use the Thread's context class loader instead of the system's default class loader.","from":"reporter","subject":"AbstractMapWritable.readFields throws ClassNotFoundException with custom writables"},{"body":"Hey Everyone: \n\nThis is related to a problem we are seeing in NUTCH, NUTCH-1084. Any chance someone could review. It looks like the test failures are actually orthogonal to this patch.","from":"developer"},{"body":"[~ndouba], please use the \"Target Version\" field to express your intention. Fix-version is exclusively used by committers when a patch gets committed.\n\nIAC, 2.7.1 is done, targeting this for 2.7.2/2.8.0.","from":"developer"},{"body":"[~ndouba], is this happening inside a MapReduce job running on top of YARN? MapReduce does have job.jar in the system classpath, so that's not explainable.\n\nThe only way this may happen is if you are using the MapReduce JobClassLoader for this - please let us know the values for mapreduce.job.classloader in your job.","from":"developer"},{"body":"Moving this out of 2.7.2 as there's been no update in a while.","from":"developer"},{"body":"Hi Vinod,\n\nI was running this on top of YARN with the default configuration enabled on top of Java 1.8 from Oracle. I think the new isolation strategy used in the newer versions of the JVM for class loaders prevents the system class loader from loading a job jar class. I didn't do anything fancy with configuration options or anything else. The only thing I added were the remote debugging options to allow me to attach to the map and reduce threads remotely and debug them step by step. I can tell you that the job jar definitely isn't in the system's class loader but in the current thread's context class loader. I verified this observation by manually inspecting the class loader object from the system and the one that comes with the thread context.\n\nCheers,\n\nNadeem","from":"developer"},{"body":"Hi [~ndouba],\n\nI'm about to do a 2.7.3 Apache Hadoop release and finally got around to this again.\n\nh4. Analysis\nTo make progress, I had to read up a bit on nutch and about how to run this so that I can reproduce the bug in order to rationalize your patch. I finally succeeded in doing so! Tested this with 2.7.2 release and nutch 1.11 and using the URL feed [given at NUTCH-1084|https://issues.apache.org/jira/browse/NUTCH-1084?focusedCommentId=13882771&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-13882771]\n{code}\n~/tmp/common/hadoop-common-2.7.2/bin/hadoop jar apache-nutch-1.11.job org.apache.nutch.crawl.CrawlDbReader file:///tmp/nutch/apache-nutch-1.11/runtime/local/crawl/crawldb/ -url http://bappenas.go.id/\n{code}\n\nI can reproduce all the problems listed at NUTCH-1084 - with readdb, MR local-job-runner based job for crawling etc.\n\nThe real issue is that Nutch's readdb is client-only and *not* running a MapReduce job which was my question before. For regular MR jobs, the job-jar *is* on the system class-loader. For the client-only invocations using \"hadoop jar\" and local-job-runner, the job-jar is actually *not* on the system-classpath - that is why you are running into the issue.\n\nh4. Summary\nYour patch looks good to me. Clearly, the thread context-loader falls back to system class-loader where it is not overridden - so we are fine for all the ways of loading the classes in readFields.\n\nI'll resubmit your patch with minor commenting related changes to Jenkins and commit if Mr.Jenkins is also fine.","from":"developer"},{"body":"Updated patch fixing the code comment..","from":"developer"},{"body":"Hi Vinod and Nadeem, thanks for taking care of this. At Apache Nutch, we're looking forward to 2.7.3!","from":"developer"},{"body":"TestReloadingX509TrustManager failure is not related, I'll see if there is an existing JIRA.\n\nChecking this in now..","from":"developer"},{"body":"Committed this to trunk, branch-2 and branch-2.7. Thanks [~ndouba]!\n\nForgot to mention that I've tested that the nutch code fails without the patch and passes with.","from":"developer"},{"body":"Closing the JIRA as part of 2.7.3 release.","from":"developer"},{"body":"Merge the fix to branch-2.8.","from":"developer"}],"created":"2015-09-12T07:09:41.000+0000","description":"Note: I am not an expert at JAVA, Class loaders, or Hadoop. I am just a hacker. My solution might be entirely wrong.\n\nAbstractMapWritable.readFields throws a ClassNotFoundException when reading custom writables. Debugging the job using remote debugging in IntelliJ revealed that the class loader being used in Class.forName() is different than that used by the Thread's current context (Thread.currentThread().getContextClassLoader()). The class path for the system class loader does not include the libraries of the job jar. However, the class path for the context class loader does. The proposed patch changes the class loading mechanism in readFields to use the Thread's context class loader instead of the system's default class loader.","issue_id":"12863593","key":"HADOOP-12406","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2016-04-11T19:06:48.000+0000","role":"fixed_distractor","summary":"AbstractMapWritable.readFields throws ClassNotFoundException with custom writables"} {"case_id":"12903596","cluster":"DISTRACTOR-HADOOP-12469","comments":[{"body":"Hi [~jira.shegalov], thanks for reporting this.\n\nI'm working on a patch to wrap the {{CopyReadException}} only once. However, I found it's hard to write a unit test which will fail. The current unit tests are not able to detect this bug because the {{map}} method throws an {{CopyReadException}} when calling {{getFileStatus}}, before it calls the {{copyFileWithRetry}}. This makes sense as the file is deleted before the copy operation. This way the unit tests pass with the following condition true (in case of ignoreFailures):\n{code}\nif (ignoreFailures && exception.getCause() instanceof\n RetriableFileCopyCommand.CopyReadException\n{code}\n\nShould we address the orthogonal {{ignoreFailures}} mutually exclusive with the atomic option separately?","created":"2015-10-09T22:06:09.480+0000"},{"body":"The release audit warning is unrelated.","created":"2015-10-10T04:47:26.219+0000"},{"body":"I've committed the patch to trunk and branch-2. Thanks [~liuml07] for the contribution.","created":"2015-10-10T08:32:48.617+0000"},{"body":"Reopening the jira.\n\nSorry for the noise -- I realized that I haven't posted my +1 on the patch before committing. Taking a closer look the patch, it does not seem to really address the problem. I'm going to revert it now.","created":"2015-10-10T10:24:47.736+0000"},{"body":"Hi [~liuml07], thanks for working on the patch. Regarding the unit test you should consider a full distcp test where you create a file and make it unusable by setting 000 permission. DistCp should be run doAs non-cluster-root user.\n\nbq. Should we address the orthogonal ignoreFailures mutually exclusive with the atomic option separately?\nSure, will you file it?\n\n","created":"2015-10-11T22:35:52.609+0000"},{"body":"[~wheat9] even if you had posted your +1, I would have appreciated a chance to review the patch before committing it. ","created":"2015-10-11T22:39:34.735+0000"},{"body":"Thanks [~wheat9] for reviewing this patch. I should make it clear that it was not to address the whole issue when I upload the patch. Sorry for the confusion.\n\nAs [~jira.shegalov] suggested, I'll update the patch with tests enabled. I filed a new issue [HADOOP-12473] about the atomic option for ignoring failures. ","created":"2015-10-12T17:40:00.363+0000"},{"body":"As [~jira.shegalov] suggested, the v1 patch changes unit test helper method {{doTestIgnoreFailures}} by creating a file and making it unusable by setting 000 permission. The unit test {{testIgnoreFailures}} can not pass in {{trunk}} code without this patch, as it should wrap the {{CopyReadException}} only once.","created":"2015-10-31T22:05:50.426+0000"},{"body":"Ah, I see what you did there. +1.","created":"2015-12-21T22:27:27.904+0000"},{"body":"The v2 patch rebases from {{trunk}} branch to resolve trivial conflicts.","created":"2016-01-21T22:22:39.087+0000"},{"body":"Hi [~liuml07], can you move doTestIgnoreFailures back underneath testDirToFile so it's easier to see what you changed?","created":"2016-01-27T18:16:26.074+0000"},{"body":"Thanks [~jira.shegalov] for your review and comments. The v4 patch moves the {{doTestIgnoreFailures()}} method back to its old position.\n\nThe patch mainly changes helper method {{doTestIgnoreFailures}} by creating a file and making it unusable by setting 000 permission. The unit test {{testIgnoreFailures}} can not pass in trunk code without this patch, as it should wrap the CopyReadException only once.","created":"2016-01-27T22:15:57.346+0000"},{"body":"Hi [~jira.shegalov], any further comments? Thanks.","created":"2016-02-23T21:50:15.559+0000"},{"body":"Can anyone commit this reviewed patch? Thanks.","created":"2016-04-27T02:29:40.326+0000"},{"body":"Per-offline discussion with [~jingzhao], the v5 patch relaxes the need of wrapping a {{CopyReadException}} exception, which happens in multiple nested call paths, and checks if any {{Throwable}} in the exception chain matches {{CopyReadException}} exception type.","created":"2016-04-28T02:12:49.172+0000"},{"body":"Re-uploading the same patch v5 to trigger Jenkins.","created":"2016-04-29T00:06:28.097+0000"},{"body":"Instead of replacing the old unit test (which deletes some source files to generate failures), maybe we can add a new unit test. +1 after addressing the comment.","created":"2016-05-03T23:18:21.282+0000"},{"body":"Thanks [~jingzhao] for the comments. The v6 patch is to address this.","created":"2016-05-04T00:54:10.106+0000"},{"body":"+1 on the latest patch. I've committed this to trunk and branch-2. Thanks [~liuml07] for the contribution!","created":"2016-05-04T17:25:31.704+0000"},{"body":"Thanks [~jingzhao] for your helpful discussion, review and commit!","created":"2016-05-04T17:50:41.975+0000"}],"conversations":[{"body":"{{RetriableFileCopyCommand.CopyReadException}} is double-wrapped via\n\n# via {{RetriableCommand::execute}}\n# via {{CopyMapper#copyFileWithRetry}}\n\nbefore {{CopyMapper::handleFailure}} tests \n{code}\nif (ignoreFailures && exception.getCause() instanceof\n RetriableFileCopyCommand.CopyReadException\n{code}\nwhich is always false.\n\nOrthogonally, ignoring failures should be mutually exclusive with the atomic option otherwise an incomplete dir is eligible for commit defeating the purpose.\n ","from":"reporter","subject":"distcp should not ignore the ignoreFailures option"},{"body":"Hi [~jira.shegalov], thanks for reporting this.\n\nI'm working on a patch to wrap the {{CopyReadException}} only once. However, I found it's hard to write a unit test which will fail. The current unit tests are not able to detect this bug because the {{map}} method throws an {{CopyReadException}} when calling {{getFileStatus}}, before it calls the {{copyFileWithRetry}}. This makes sense as the file is deleted before the copy operation. This way the unit tests pass with the following condition true (in case of ignoreFailures):\n{code}\nif (ignoreFailures && exception.getCause() instanceof\n RetriableFileCopyCommand.CopyReadException\n{code}\n\nShould we address the orthogonal {{ignoreFailures}} mutually exclusive with the atomic option separately?","from":"developer"},{"body":"The release audit warning is unrelated.","from":"developer"},{"body":"I've committed the patch to trunk and branch-2. Thanks [~liuml07] for the contribution.","from":"developer"},{"body":"Reopening the jira.\n\nSorry for the noise -- I realized that I haven't posted my +1 on the patch before committing. Taking a closer look the patch, it does not seem to really address the problem. I'm going to revert it now.","from":"developer"},{"body":"Hi [~liuml07], thanks for working on the patch. Regarding the unit test you should consider a full distcp test where you create a file and make it unusable by setting 000 permission. DistCp should be run doAs non-cluster-root user.\n\nbq. Should we address the orthogonal ignoreFailures mutually exclusive with the atomic option separately?\nSure, will you file it?\n\n","from":"developer"},{"body":"[~wheat9] even if you had posted your +1, I would have appreciated a chance to review the patch before committing it. ","from":"developer"},{"body":"Thanks [~wheat9] for reviewing this patch. I should make it clear that it was not to address the whole issue when I upload the patch. Sorry for the confusion.\n\nAs [~jira.shegalov] suggested, I'll update the patch with tests enabled. I filed a new issue [HADOOP-12473] about the atomic option for ignoring failures. ","from":"developer"},{"body":"As [~jira.shegalov] suggested, the v1 patch changes unit test helper method {{doTestIgnoreFailures}} by creating a file and making it unusable by setting 000 permission. The unit test {{testIgnoreFailures}} can not pass in {{trunk}} code without this patch, as it should wrap the {{CopyReadException}} only once.","from":"developer"},{"body":"Ah, I see what you did there. +1.","from":"developer"},{"body":"The v2 patch rebases from {{trunk}} branch to resolve trivial conflicts.","from":"developer"},{"body":"Hi [~liuml07], can you move doTestIgnoreFailures back underneath testDirToFile so it's easier to see what you changed?","from":"developer"},{"body":"Thanks [~jira.shegalov] for your review and comments. The v4 patch moves the {{doTestIgnoreFailures()}} method back to its old position.\n\nThe patch mainly changes helper method {{doTestIgnoreFailures}} by creating a file and making it unusable by setting 000 permission. The unit test {{testIgnoreFailures}} can not pass in trunk code without this patch, as it should wrap the CopyReadException only once.","from":"developer"},{"body":"Hi [~jira.shegalov], any further comments? Thanks.","from":"developer"},{"body":"Can anyone commit this reviewed patch? Thanks.","from":"developer"},{"body":"Per-offline discussion with [~jingzhao], the v5 patch relaxes the need of wrapping a {{CopyReadException}} exception, which happens in multiple nested call paths, and checks if any {{Throwable}} in the exception chain matches {{CopyReadException}} exception type.","from":"developer"},{"body":"Re-uploading the same patch v5 to trigger Jenkins.","from":"developer"},{"body":"Instead of replacing the old unit test (which deletes some source files to generate failures), maybe we can add a new unit test. +1 after addressing the comment.","from":"developer"},{"body":"Thanks [~jingzhao] for the comments. The v6 patch is to address this.","from":"developer"},{"body":"+1 on the latest patch. I've committed this to trunk and branch-2. Thanks [~liuml07] for the contribution!","from":"developer"},{"body":"Thanks [~jingzhao] for your helpful discussion, review and commit!","from":"developer"}],"created":"2015-10-09T01:30:16.000+0000","description":"{{RetriableFileCopyCommand.CopyReadException}} is double-wrapped via\n\n# via {{RetriableCommand::execute}}\n# via {{CopyMapper#copyFileWithRetry}}\n\nbefore {{CopyMapper::handleFailure}} tests \n{code}\nif (ignoreFailures && exception.getCause() instanceof\n RetriableFileCopyCommand.CopyReadException\n{code}\nwhich is always false.\n\nOrthogonally, ignoring failures should be mutually exclusive with the atomic option otherwise an incomplete dir is eligible for commit defeating the purpose.\n ","issue_id":"12903596","key":"HADOOP-12469","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2016-05-04T17:25:31.000+0000","role":"fixed_distractor","summary":"distcp should not ignore the ignoreFailures option"} {"case_id":"12952099","cluster":"DISTRACTOR-HADOOP-12948","comments":[{"body":"I think that instead of downloading/launching a standalone Apache DS server, it should use the embedded MiniKdc, which is based on Apache DS. MiniKdc implemented in HADOOP-9848 can be useful for this purpose.","created":"2016-03-21T17:57:43.376+0000"},{"body":"I agree with converting these tests to use mini-KDC. I think the startKdc profile has been broken for a long time. In practice, I suspect this means no one really runs and pays attention to those tests.","created":"2016-03-21T18:31:00.910+0000"},{"body":"+! for the move, but note that miniKDC has its own problems; it doesn't really like >1 principal\n\nThis is something which kerby could help with","created":"2016-03-22T09:59:42.360+0000"},{"body":"Thanks [~steve_l] for the comment. Yes looks like Kerby is the way to go. I wasn't paying attention to the latest community trend. :) HADOOP-12911 is the related jira to replace miniKDC with Kerby.","created":"2016-03-22T11:03:57.303+0000"},{"body":"After studying the code, I think {{TestUGIWithSecurityOn}} should be removed. Its {{testLogin}} duplicates the tests in {{TestKMS}}, and its {{testGetUGIFromKerberosSubject}} duplicates {{TestMiniKdc#testKerberosLogin}}.\n\n","created":"2016-03-24T21:21:01.211+0000"},{"body":"Rev01: Removed the files added in HADOOP-8078.","created":"2016-03-24T21:21:49.965+0000"},{"body":"Updated pom.xml to remove startKdc profile, and removed users.ldif which was added in HADOOP-8078.","created":"2016-03-24T21:30:28.004+0000"},{"body":"Filed a corresponding jira HDFS-10210 to remove HDFS side of the code.","created":"2016-03-24T21:32:49.350+0000"},{"body":"The last patch did not apply successfully. I suspect it was due to the binary keytab files removed in the patch. Uploaded rev03 which does not touch those keytab files.","created":"2016-03-25T14:46:22.605+0000"},{"body":"If you want to work with binary files in a patch, use 'git format-patch' to generate it.","created":"2016-03-25T15:32:25.443+0000"},{"body":"Thanks for the suggestion, Allen.\nLet's see if Yetus is clever enough to apply this patch.","created":"2016-03-25T17:09:31.765+0000"},{"body":"You could always test yourself:\n\n{code}\ndhcp-208:hadoop aw$ dev-support/bin/smart-apply-patch --plugins=all HADOOP-12948\nProcessing: HADOOP-12948\nHADOOP-12948 patch is being downloaded at Fri Mar 25 10:27:49 PDT 2016 from\n https://issues.apache.org/jira/secure/attachment/12795432/0001-HADOOP-12948-Maven-profile-startKdc-is-broken.patch -> Downloaded\nApplying the patch:\nFri Mar 25 10:27:49 PDT 2016\ncd /Users/aw/Src/drd/hadoop\ngit apply --binary -v --stat --apply -p0 /tmp/yetus-27292.24189/patch\nApplied patch hadoop-common-project/hadoop-common/pom.xml cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/security/TestUGIWithSecurityOn.java cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/dn1.keytab cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/nn1.keytab cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/user1.keytab cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/killKdc.sh cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/ldif/users.ldif cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/server.xml cleanly.\n hadoop-common-project/hadoop-common/pom.xml | 85 -------\n .../hadoop/security/TestUGIWithSecurityOn.java | 117 ---------\n .../src/test/resources/kdc/keytabs/dn1.keytab | Bin\n .../src/test/resources/kdc/keytabs/nn1.keytab | Bin\n .../src/test/resources/kdc/keytabs/user1.keytab | Bin\n .../src/test/resources/kdc/killKdc.sh | 19 -\n .../src/test/resources/kdc/ldif/users.ldif | 78 ------\n .../src/test/resources/kdc/server.xml | 258 --------------------\n 8 files changed, 557 deletions(-)\ndhcp-208:hadoop aw$ git status\nOn branch test\nChanges not staged for commit:\n (use \"git add/rm ...\" to update what will be committed)\n (use \"git checkout -- ...\" to discard changes in working directory)\n\n\tmodified: hadoop-common-project/hadoop-common/pom.xml\n\tdeleted: hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/security/TestUGIWithSecurityOn.java\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/dn1.keytab\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/nn1.keytab\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/user1.keytab\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/killKdc.sh\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/ldif/users.ldif\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/server.xml\n\nno changes added to commit (use \"git add\" and/or \"git commit -a\")\ndhcp-208:hadoop aw$ \n{code}","created":"2016-03-25T17:28:32.345+0000"},{"body":"Not sure why TestNativeLibraryChecker failed. This patch does not touch native code. Test Results does not show the failure.","created":"2016-03-25T20:03:27.491+0000"},{"body":"Dont worry about it. It's being fixed in HADOOP-12955.","created":"2016-03-26T01:00:00.342+0000"},{"body":"LGTM, +1 pending Jenkins","created":"2019-05-22T04:41:59.351+0000"},{"body":"Hi [~weichiu], would you remove startKdc profile from hadoop-hdfs module as well?","created":"2019-05-22T05:06:01.046+0000"},{"body":"That's taken care of by HDFS-10210.","created":"2019-05-22T13:27:40.986+0000"},{"body":"+1, committing this.","created":"2019-05-23T05:01:30.479+0000"},{"body":"Committed this to trunk. Thank you, [~weichiu]!","created":"2019-05-23T05:02:49.922+0000"}],"conversations":[{"body":"{noformat}\nmvn install -Dtest=TestUGIWithSecurityOn -DstartKdc=true\n\nmain:\n [exec] xargs: illegal option -- -\n [exec] usage: xargs [-0opt] [-E eofstr] [-I replstr [-R replacements]] [-J replstr]\n [exec] [-L number] [-n number [-x]] [-P maxprocs] [-s size]\n [exec] [utility [argument ...]]\n [exec] Result: 1\n [get] Getting: http://newverhost.com/pub//directory/apacheds/unstable/1.5/1.5.7/apacheds-1.5.7.tar.gz\n [get] To: /Users/weichiu/sandbox/hadoop/hadoop-common-project/hadoop-common/target/test-classes/kdc/downloads/apacheds-1.5.7.tar.gz\n [get] Error getting http://newverhost.com/pub//directory/apacheds/unstable/1.5/1.5.7/apacheds-1.5.7.tar.gz to /Users/weichiu/sandbox/hadoop/hadoop-common-project/hadoop-common/target/test-classes/kdc/downloads/apacheds-1.5.7.tar.gz\n[INFO] ------------------------------------------------------------------------\n[INFO] BUILD FAILURE\n[INFO] ------------------------------------------------------------------------\n[INFO] Total time: 8.448 s\n[INFO] Finished at: 2016-03-21T10:00:56-07:00\n[INFO] Final Memory: 31M/439M\n[INFO] ------------------------------------------------------------------------\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-antrun-plugin:1.7:run (kdc) on project hadoop-common: An Ant BuildException has occured: java.net.UnknownHostException: newverhost.com\n[ERROR] around Ant part ...... @ 7:244 in /Users/weichiu/sandbox/hadoop/hadoop-common-project/hadoop-common/target/antrun/build-main.xml\n[ERROR] -> [Help 1]\n[ERROR]\n[ERROR] To see the full stack trace of the errors, re-run Maven with the -e switch.\n[ERROR] Re-run Maven using the -X switch to enable full debug logging.\n[ERROR]\n[ERROR] For more information about the errors and possible solutions, please read the following articles:\n[ERROR] [Help 1] http://cwiki.apache.org/confluence/display/MAVEN/MojoExecutionException\n{noformat}\nI'm using Mac so part of the reason might be my operating system (even though the pom.xml stated it supported Mac), but the major problem is that it attempted to download apacheds from newverhost.com, which does not seem exist any more.\n\nThese tests were implemented in HADOOP-8078, and must have -DstartKdc=true in order to run them.","from":"reporter","subject":"Remove the defunct startKdc profile from hadoop-common"},{"body":"I think that instead of downloading/launching a standalone Apache DS server, it should use the embedded MiniKdc, which is based on Apache DS. MiniKdc implemented in HADOOP-9848 can be useful for this purpose.","from":"developer"},{"body":"I agree with converting these tests to use mini-KDC. I think the startKdc profile has been broken for a long time. In practice, I suspect this means no one really runs and pays attention to those tests.","from":"developer"},{"body":"+! for the move, but note that miniKDC has its own problems; it doesn't really like >1 principal\n\nThis is something which kerby could help with","from":"developer"},{"body":"Thanks [~steve_l] for the comment. Yes looks like Kerby is the way to go. I wasn't paying attention to the latest community trend. :) HADOOP-12911 is the related jira to replace miniKDC with Kerby.","from":"developer"},{"body":"After studying the code, I think {{TestUGIWithSecurityOn}} should be removed. Its {{testLogin}} duplicates the tests in {{TestKMS}}, and its {{testGetUGIFromKerberosSubject}} duplicates {{TestMiniKdc#testKerberosLogin}}.\n\n","from":"developer"},{"body":"Rev01: Removed the files added in HADOOP-8078.","from":"developer"},{"body":"Updated pom.xml to remove startKdc profile, and removed users.ldif which was added in HADOOP-8078.","from":"developer"},{"body":"Filed a corresponding jira HDFS-10210 to remove HDFS side of the code.","from":"developer"},{"body":"The last patch did not apply successfully. I suspect it was due to the binary keytab files removed in the patch. Uploaded rev03 which does not touch those keytab files.","from":"developer"},{"body":"If you want to work with binary files in a patch, use 'git format-patch' to generate it.","from":"developer"},{"body":"Thanks for the suggestion, Allen.\nLet's see if Yetus is clever enough to apply this patch.","from":"developer"},{"body":"You could always test yourself:\n\n{code}\ndhcp-208:hadoop aw$ dev-support/bin/smart-apply-patch --plugins=all HADOOP-12948\nProcessing: HADOOP-12948\nHADOOP-12948 patch is being downloaded at Fri Mar 25 10:27:49 PDT 2016 from\n https://issues.apache.org/jira/secure/attachment/12795432/0001-HADOOP-12948-Maven-profile-startKdc-is-broken.patch -> Downloaded\nApplying the patch:\nFri Mar 25 10:27:49 PDT 2016\ncd /Users/aw/Src/drd/hadoop\ngit apply --binary -v --stat --apply -p0 /tmp/yetus-27292.24189/patch\nApplied patch hadoop-common-project/hadoop-common/pom.xml cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/security/TestUGIWithSecurityOn.java cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/dn1.keytab cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/nn1.keytab cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/user1.keytab cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/killKdc.sh cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/ldif/users.ldif cleanly.\nApplied patch hadoop-common-project/hadoop-common/src/test/resources/kdc/server.xml cleanly.\n hadoop-common-project/hadoop-common/pom.xml | 85 -------\n .../hadoop/security/TestUGIWithSecurityOn.java | 117 ---------\n .../src/test/resources/kdc/keytabs/dn1.keytab | Bin\n .../src/test/resources/kdc/keytabs/nn1.keytab | Bin\n .../src/test/resources/kdc/keytabs/user1.keytab | Bin\n .../src/test/resources/kdc/killKdc.sh | 19 -\n .../src/test/resources/kdc/ldif/users.ldif | 78 ------\n .../src/test/resources/kdc/server.xml | 258 --------------------\n 8 files changed, 557 deletions(-)\ndhcp-208:hadoop aw$ git status\nOn branch test\nChanges not staged for commit:\n (use \"git add/rm ...\" to update what will be committed)\n (use \"git checkout -- ...\" to discard changes in working directory)\n\n\tmodified: hadoop-common-project/hadoop-common/pom.xml\n\tdeleted: hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/security/TestUGIWithSecurityOn.java\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/dn1.keytab\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/nn1.keytab\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/keytabs/user1.keytab\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/killKdc.sh\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/ldif/users.ldif\n\tdeleted: hadoop-common-project/hadoop-common/src/test/resources/kdc/server.xml\n\nno changes added to commit (use \"git add\" and/or \"git commit -a\")\ndhcp-208:hadoop aw$ \n{code}","from":"developer"},{"body":"Not sure why TestNativeLibraryChecker failed. This patch does not touch native code. Test Results does not show the failure.","from":"developer"},{"body":"Dont worry about it. It's being fixed in HADOOP-12955.","from":"developer"},{"body":"LGTM, +1 pending Jenkins","from":"developer"},{"body":"Hi [~weichiu], would you remove startKdc profile from hadoop-hdfs module as well?","from":"developer"},{"body":"That's taken care of by HDFS-10210.","from":"developer"},{"body":"+1, committing this.","from":"developer"},{"body":"Committed this to trunk. Thank you, [~weichiu]!","from":"developer"}],"created":"2016-03-21T17:05:34.000+0000","description":"{noformat}\nmvn install -Dtest=TestUGIWithSecurityOn -DstartKdc=true\n\nmain:\n [exec] xargs: illegal option -- -\n [exec] usage: xargs [-0opt] [-E eofstr] [-I replstr [-R replacements]] [-J replstr]\n [exec] [-L number] [-n number [-x]] [-P maxprocs] [-s size]\n [exec] [utility [argument ...]]\n [exec] Result: 1\n [get] Getting: http://newverhost.com/pub//directory/apacheds/unstable/1.5/1.5.7/apacheds-1.5.7.tar.gz\n [get] To: /Users/weichiu/sandbox/hadoop/hadoop-common-project/hadoop-common/target/test-classes/kdc/downloads/apacheds-1.5.7.tar.gz\n [get] Error getting http://newverhost.com/pub//directory/apacheds/unstable/1.5/1.5.7/apacheds-1.5.7.tar.gz to /Users/weichiu/sandbox/hadoop/hadoop-common-project/hadoop-common/target/test-classes/kdc/downloads/apacheds-1.5.7.tar.gz\n[INFO] ------------------------------------------------------------------------\n[INFO] BUILD FAILURE\n[INFO] ------------------------------------------------------------------------\n[INFO] Total time: 8.448 s\n[INFO] Finished at: 2016-03-21T10:00:56-07:00\n[INFO] Final Memory: 31M/439M\n[INFO] ------------------------------------------------------------------------\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-antrun-plugin:1.7:run (kdc) on project hadoop-common: An Ant BuildException has occured: java.net.UnknownHostException: newverhost.com\n[ERROR] around Ant part ...... @ 7:244 in /Users/weichiu/sandbox/hadoop/hadoop-common-project/hadoop-common/target/antrun/build-main.xml\n[ERROR] -> [Help 1]\n[ERROR]\n[ERROR] To see the full stack trace of the errors, re-run Maven with the -e switch.\n[ERROR] Re-run Maven using the -X switch to enable full debug logging.\n[ERROR]\n[ERROR] For more information about the errors and possible solutions, please read the following articles:\n[ERROR] [Help 1] http://cwiki.apache.org/confluence/display/MAVEN/MojoExecutionException\n{noformat}\nI'm using Mac so part of the reason might be my operating system (even though the pom.xml stated it supported Mac), but the major problem is that it attempted to download apacheds from newverhost.com, which does not seem exist any more.\n\nThese tests were implemented in HADOOP-8078, and must have -DstartKdc=true in order to run them.","issue_id":"12952099","key":"HADOOP-12948","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2019-05-23T05:02:49.000+0000","role":"fixed_distractor","summary":"Remove the defunct startKdc profile from hadoop-common"} {"case_id":"12954650","cluster":"DISTRACTOR-HADOOP-12979","comments":[{"body":"full stack\n{code}\n\nDriver stacktrace:\n at org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1457)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1445)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1444)\n at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n at org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1444)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:809)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:809)\n at scala.Option.foreach(Option.scala:257)\n at org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:809)\n ...\n Cause: java.io.IOException: ${hadoop.tmp.dir}/s3a not configured\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.confChanged(LocalDirAllocator.java:269)\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.getLocalPathForWrite(LocalDirAllocator.java:349)\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.createTmpFileForWrite(LocalDirAllocator.java:421)\n at org.apache.hadoop.fs.LocalDirAllocator.createTmpFileForWrite(LocalDirAllocator.java:198)\n at org.apache.hadoop.fs.s3a.S3AOutputStream.(S3AOutputStream.java:91)\n at org.apache.hadoop.fs.s3a.S3AFileSystem.create(S3AFileSystem.java:488)\n at org.apache.hadoop.fs.FileSystem.create(FileSystem.java:921)\n at org.apache.hadoop.fs.FileSystem.create(FileSystem.java:814)\n at org.apache.hadoop.mapred.TextOutputFormat.getRecordWriter(TextOutputFormat.java:123)\n at org.apache.spark.SparkHadoopWriter.open(SparkHadoopWriter.scala:90)\n{code}","created":"2016-03-30T14:44:31.084+0000"},{"body":"The problem is the branch take if there's no s3 buffer dir defined. the else clause is broken; it's looking for a config option which isn't there.\n{code}\n if (conf.get(BUFFER_DIR, null) != null) {\n lDirAlloc = new LocalDirAllocator(BUFFER_DIR);\n } else {\n lDirAlloc = new LocalDirAllocator(\"${hadoop.tmp.dir}/s3a\"); // HERE\n }\n{code}\n\nThe fix should be to set {{BUFFER_DIR}} to the full path desired, create the {{LocalDirAllocator(BUFFER_DIR)}} from the (possibly enhanced) config","created":"2016-03-30T14:47:08.577+0000"}],"conversations":[{"body":"Running some spark s3a tests trigger an NPE in Hadoop <=2/7l IOE in 2.8 saying \n{code}\n${hadoop.tmp.dir}/s3a not configured.\n{code}\nThat's correct: there is no configuration option on the conf called \n{code}\n${hadoop.tmp.dir}/s3a\n{code}\nThere may be one called {{hadoop.tmp.dir}}, however.\n\nEssentially s3a is sending the wrong config option down, if it can't find {{fs.s3a.buffer.dir}}","from":"reporter","subject":"IOE in S3a: ${hadoop.tmp.dir}/s3a not configured"},{"body":"full stack\n{code}\n\nDriver stacktrace:\n at org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1457)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1445)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1444)\n at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n at org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1444)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:809)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:809)\n at scala.Option.foreach(Option.scala:257)\n at org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:809)\n ...\n Cause: java.io.IOException: ${hadoop.tmp.dir}/s3a not configured\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.confChanged(LocalDirAllocator.java:269)\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.getLocalPathForWrite(LocalDirAllocator.java:349)\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.createTmpFileForWrite(LocalDirAllocator.java:421)\n at org.apache.hadoop.fs.LocalDirAllocator.createTmpFileForWrite(LocalDirAllocator.java:198)\n at org.apache.hadoop.fs.s3a.S3AOutputStream.(S3AOutputStream.java:91)\n at org.apache.hadoop.fs.s3a.S3AFileSystem.create(S3AFileSystem.java:488)\n at org.apache.hadoop.fs.FileSystem.create(FileSystem.java:921)\n at org.apache.hadoop.fs.FileSystem.create(FileSystem.java:814)\n at org.apache.hadoop.mapred.TextOutputFormat.getRecordWriter(TextOutputFormat.java:123)\n at org.apache.spark.SparkHadoopWriter.open(SparkHadoopWriter.scala:90)\n{code}","from":"developer"},{"body":"The problem is the branch take if there's no s3 buffer dir defined. the else clause is broken; it's looking for a config option which isn't there.\n{code}\n if (conf.get(BUFFER_DIR, null) != null) {\n lDirAlloc = new LocalDirAllocator(BUFFER_DIR);\n } else {\n lDirAlloc = new LocalDirAllocator(\"${hadoop.tmp.dir}/s3a\"); // HERE\n }\n{code}\n\nThe fix should be to set {{BUFFER_DIR}} to the full path desired, create the {{LocalDirAllocator(BUFFER_DIR)}} from the (possibly enhanced) config","from":"developer"}],"created":"2016-03-30T14:40:11.000+0000","description":"Running some spark s3a tests trigger an NPE in Hadoop <=2/7l IOE in 2.8 saying \n{code}\n${hadoop.tmp.dir}/s3a not configured.\n{code}\nThat's correct: there is no configuration option on the conf called \n{code}\n${hadoop.tmp.dir}/s3a\n{code}\nThere may be one called {{hadoop.tmp.dir}}, however.\n\nEssentially s3a is sending the wrong config option down, if it can't find {{fs.s3a.buffer.dir}}","issue_id":"12954650","key":"HADOOP-12979","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2017-03-03T18:55:00.000+0000","role":"fixed_distractor","summary":"IOE in S3a: ${hadoop.tmp.dir}/s3a not configured"} {"case_id":"12969760","cluster":"DISTRACTOR-HADOOP-13147","comments":[{"body":"Issue which introduced the code","created":"2016-05-14T11:59:29.264+0000"},{"body":"BTW, why is the following final?\n\nfinal public void update(int b)\n\nNone of the other methods are.","created":"2016-05-14T12:01:05.460+0000"},{"body":"Bulk update: moved all 3.2.0 non-blocker issues, please move back if it is a blocker.","created":"2018-11-23T11:56:19.056+0000"},{"body":"Bulk update: moved all 3.3.0 non-blocker issues, please move back if it is a blocker.","created":"2020-04-09T18:51:53.782+0000"},{"body":"Bulk update: moved all 3.4.0 non-blocker issues, please move back if it is a blocker. Retarget 3.5.0.","created":"2024-01-04T11:36:16.028+0000"},{"body":"What is the point of saying \"*please move back if it is a blocker.* \" if you don't intend to allow it?","created":"2024-01-04T12:58:50.589+0000"},{"body":"[~sebb] We plan to release hadoop-3.4.0 based on the trunk branch of hadoop. According to the release process, some non-blocking jira target versions need to be updated. I saw that when 3.3.0 was released in April 2020, this JIRA was not a blocking issue, so I updated the status of this JIRA. We can change the target to 3.5.0, which will not affect your PR submission.","created":"2024-01-04T13:09:40.235+0000"},{"body":"Committed to trunk.\r\nThanx [~sebb] for the contribution!!!\r\n\r\nAdded [~sebb] as Hadoop-Common Contributor to assign the ticket.\r\n\r\nWelcome to Hadoop :-) ","created":"2024-05-20T18:39:11.440+0000"}],"conversations":[{"body":"Constructors must not call overrideable methods.\n\nAn object is not guaranteed fully constructed until the constructor exits, so the subclass override may not see the fully created parent object.\n\nThis applies to:\n\nPureJavaCrc32\n","from":"reporter","subject":"Constructors must not call overrideable methods in PureJavaCrc32C"},{"body":"Issue which introduced the code","from":"developer"},{"body":"BTW, why is the following final?\n\nfinal public void update(int b)\n\nNone of the other methods are.","from":"developer"},{"body":"Bulk update: moved all 3.2.0 non-blocker issues, please move back if it is a blocker.","from":"developer"},{"body":"Bulk update: moved all 3.3.0 non-blocker issues, please move back if it is a blocker.","from":"developer"},{"body":"Bulk update: moved all 3.4.0 non-blocker issues, please move back if it is a blocker. Retarget 3.5.0.","from":"developer"},{"body":"What is the point of saying \"*please move back if it is a blocker.* \" if you don't intend to allow it?","from":"developer"},{"body":"[~sebb] We plan to release hadoop-3.4.0 based on the trunk branch of hadoop. According to the release process, some non-blocking jira target versions need to be updated. I saw that when 3.3.0 was released in April 2020, this JIRA was not a blocking issue, so I updated the status of this JIRA. We can change the target to 3.5.0, which will not affect your PR submission.","from":"developer"},{"body":"Committed to trunk.\r\nThanx [~sebb] for the contribution!!!\r\n\r\nAdded [~sebb] as Hadoop-Common Contributor to assign the ticket.\r\n\r\nWelcome to Hadoop :-) ","from":"developer"}],"created":"2016-05-14T11:58:07.000+0000","description":"Constructors must not call overrideable methods.\n\nAn object is not guaranteed fully constructed until the constructor exits, so the subclass override may not see the fully created parent object.\n\nThis applies to:\n\nPureJavaCrc32\n","issue_id":"12969760","key":"HADOOP-13147","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-05-20T18:39:23.000+0000","role":"fixed_distractor","summary":"Constructors must not call overrideable methods in PureJavaCrc32C"} {"case_id":"12993096","cluster":"DISTRACTOR-HADOOP-13434","comments":[{"body":"This patch adds quoting to the Shell class on each command that invokes bash with string parameters.","created":"2016-07-28T18:28:24.255+0000"},{"body":"This iteration removes the use of bash from the cases where it wasn't necessary.","created":"2016-07-28T22:12:03.167+0000"},{"body":"LGTM - removes bash where there are alternatives and quotes the rest.\n+1 - pending jenkins green light.\n","created":"2016-07-29T02:21:26.874+0000"},{"body":"I made some tweaks to make check style happy.","created":"2016-07-29T22:06:52.032+0000"},{"body":"{code}\n+ : new String[] { \"/bin/bash\", bashQuote(absolutePath) };\n{code}\n\nWhile we're here, let's remove the path to /bin/bash since that won't work on every platform.","created":"2016-07-30T22:29:24.240+0000"},{"body":"[~aw] - that seems reasonable and correct to me.\n[~owen.omalley] - Have you checked the unit test failure from the jenkins run - it strikes me as unrelated.","created":"2016-08-01T15:13:06.710+0000"},{"body":"The test failure looks unrelated (won't repro for me).\n\nI'd like to commit this tomorrow. Allen, I can post a separate patch for the fix you suggested after I commit this, assuming you don't object. ","created":"2016-08-01T18:03:43.505+0000"},{"body":"Pushed this through to branch-2.8. Thank you for fixing this [~owen.omalley] and [~lmccay] for the code review. Filed HADOOP-13457 for the bug-fix suggested by Allen.\n\nCherry-picking to branch-2.7 threw up many conflicts. I suspect they will be straightforward. I'll post a 2.7.x backport patch later today.","created":"2016-08-02T20:53:45.165+0000"},{"body":"Reopening to attach branch-2.7 patch.","created":"2016-08-02T21:56:37.764+0000"},{"body":"Backport of Owen's patch. The conflicts were easy to resolve.","created":"2016-08-02T21:58:00.566+0000"},{"body":"Pushed to branch-2.7 and branch-2.7.3 also based (thanks [~lmccay] for taking a look at the 2.7 patch).","created":"2016-08-03T21:36:16.431+0000"},{"body":"Don't know what happened to my +1 for the 2.7 patch - so here it is again:\n\n+1","created":"2016-08-04T15:07:28.123+0000"},{"body":"Closing the JIRA as part of 2.7.3 release.","created":"2016-08-25T22:48:35.808+0000"},{"body":"Cherry-picked it to 2.6.5 (trivial).","created":"2016-09-15T05:09:57.347+0000"}],"conversations":[{"body":"The Shell class makes assumptions that the parameters won't have spaces or other special characters, even when it invokes bash.","from":"reporter","subject":"CVE-2016-5393: Add quoting to Shell class"},{"body":"This patch adds quoting to the Shell class on each command that invokes bash with string parameters.","from":"developer"},{"body":"This iteration removes the use of bash from the cases where it wasn't necessary.","from":"developer"},{"body":"LGTM - removes bash where there are alternatives and quotes the rest.\n+1 - pending jenkins green light.\n","from":"developer"},{"body":"I made some tweaks to make check style happy.","from":"developer"},{"body":"{code}\n+ : new String[] { \"/bin/bash\", bashQuote(absolutePath) };\n{code}\n\nWhile we're here, let's remove the path to /bin/bash since that won't work on every platform.","from":"developer"},{"body":"[~aw] - that seems reasonable and correct to me.\n[~owen.omalley] - Have you checked the unit test failure from the jenkins run - it strikes me as unrelated.","from":"developer"},{"body":"The test failure looks unrelated (won't repro for me).\n\nI'd like to commit this tomorrow. Allen, I can post a separate patch for the fix you suggested after I commit this, assuming you don't object. ","from":"developer"},{"body":"Pushed this through to branch-2.8. Thank you for fixing this [~owen.omalley] and [~lmccay] for the code review. Filed HADOOP-13457 for the bug-fix suggested by Allen.\n\nCherry-picking to branch-2.7 threw up many conflicts. I suspect they will be straightforward. I'll post a 2.7.x backport patch later today.","from":"developer"},{"body":"Reopening to attach branch-2.7 patch.","from":"developer"},{"body":"Backport of Owen's patch. The conflicts were easy to resolve.","from":"developer"},{"body":"Pushed to branch-2.7 and branch-2.7.3 also based (thanks [~lmccay] for taking a look at the 2.7 patch).","from":"developer"},{"body":"Don't know what happened to my +1 for the 2.7 patch - so here it is again:\n\n+1","from":"developer"},{"body":"Closing the JIRA as part of 2.7.3 release.","from":"developer"},{"body":"Cherry-picked it to 2.6.5 (trivial).","from":"developer"}],"created":"2016-07-27T23:05:02.000+0000","description":"The Shell class makes assumptions that the parameters won't have spaces or other special characters, even when it invokes bash.","issue_id":"12993096","key":"HADOOP-13434","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2016-08-03T21:36:16.000+0000","role":"fixed_distractor","summary":"CVE-2016-5393: Add quoting to Shell class"} {"case_id":"13007752","cluster":"DISTRACTOR-HADOOP-13663","comments":[{"body":"Checking if the index is a positive number. I tried to put all the error message together in the code.","created":"2016-09-27T16:02:58.212+0000"},{"body":"Could you add a test for this into {{TestSysInfoWindows}}?","created":"2016-09-27T17:09:12.593+0000"},{"body":"Modified unit test.","created":"2016-09-27T17:14:38.723+0000"},{"body":"My cleanest Hadoop QA report in a long time :)\n[~stevel@apache.org] are you OK with this patch?","created":"2016-09-28T17:42:54.513+0000"},{"body":"+1\ncommitted to 2.8+\n\nif you need this in 2.7.4, you'll need to open a new JIRA with a patch named for a build & test against that branch. e.g {{HADOOP-13663-branch-2.7-001.patch}}","created":"2016-09-29T11:36:06.473+0000"},{"body":"Thank you [~stevel@apache.org], 2.8+ is good. No need for 2.7.4 for us.\nIf anyone wants it pushed there, I can go ahead.","created":"2016-09-29T17:20:42.056+0000"},{"body":"Re-resolving so the status is \"Fixed\" rather than \"Resolved\".","created":"2016-11-22T01:57:18.157+0000"}],"conversations":[{"body":"Sometimes, the {{NodeResourceMonitor}} tries to read the system utilization from winutils.exe and this return empty values. This triggers the following exception:\njava.lang.StringIndexOutOfBoundsException: String index out of range: -1\n\tat java.lang.String.substring(String.java:1911)\n\tat org.apache.hadoop.util.SysInfoWindows.refreshIfNeeded(SysInfoWindows.java:158)\n\tat org.apache.hadoop.util.SysInfoWindows.getPhysicalMemorySize(SysInfoWindows.java:247)\n\tat org.apache.hadoop.yarn.util.ResourceCalculatorPlugin.getPhysicalMemorySize(ResourceCalculatorPlugin.java:63)\n\tat org.apache.hadoop.yarn.server.nodemanager.NodeResourceMonitorImpl$MonitoringThread.run(NodeResourceMonitorImpl.java:139) ","from":"reporter","subject":"Index out of range in SysInfoWindows"},{"body":"Checking if the index is a positive number. I tried to put all the error message together in the code.","from":"developer"},{"body":"Could you add a test for this into {{TestSysInfoWindows}}?","from":"developer"},{"body":"Modified unit test.","from":"developer"},{"body":"My cleanest Hadoop QA report in a long time :)\n[~stevel@apache.org] are you OK with this patch?","from":"developer"},{"body":"+1\ncommitted to 2.8+\n\nif you need this in 2.7.4, you'll need to open a new JIRA with a patch named for a build & test against that branch. e.g {{HADOOP-13663-branch-2.7-001.patch}}","from":"developer"},{"body":"Thank you [~stevel@apache.org], 2.8+ is good. No need for 2.7.4 for us.\nIf anyone wants it pushed there, I can go ahead.","from":"developer"},{"body":"Re-resolving so the status is \"Fixed\" rather than \"Resolved\".","from":"developer"}],"created":"2016-09-26T23:06:22.000+0000","description":"Sometimes, the {{NodeResourceMonitor}} tries to read the system utilization from winutils.exe and this return empty values. This triggers the following exception:\njava.lang.StringIndexOutOfBoundsException: String index out of range: -1\n\tat java.lang.String.substring(String.java:1911)\n\tat org.apache.hadoop.util.SysInfoWindows.refreshIfNeeded(SysInfoWindows.java:158)\n\tat org.apache.hadoop.util.SysInfoWindows.getPhysicalMemorySize(SysInfoWindows.java:247)\n\tat org.apache.hadoop.yarn.util.ResourceCalculatorPlugin.getPhysicalMemorySize(ResourceCalculatorPlugin.java:63)\n\tat org.apache.hadoop.yarn.server.nodemanager.NodeResourceMonitorImpl$MonitoringThread.run(NodeResourceMonitorImpl.java:139) ","issue_id":"13007752","key":"HADOOP-13663","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2016-11-22T01:57:17.000+0000","role":"fixed_distractor","summary":"Index out of range in SysInfoWindows"} {"case_id":"13027245","cluster":"DISTRACTOR-HADOOP-13890","comments":[{"body":"Saw them in pre-commit unit tests: https://builds.apache.org/job/PreCommit-HDFS-Build/17824/artifact/patchprocess/patch-unit-hadoop-common-project_hadoop-common.txt and https://builds.apache.org/job/PreCommit-HDFS-Build/17824/artifact/patchprocess/patch-unit-hadoop-hdfs-project_hadoop-hdfs.txt.\n","created":"2016-12-12T01:41:38.484+0000"},{"body":"Attach a patch that fixed the unit tests.","created":"2016-12-12T07:15:52.363+0000"},{"body":"Sorry I was not able to catch this at the commit time because Jenkins only run hadoop-auth tests but the failed tests are from hadoop-common. \n\n","created":"2016-12-12T07:18:36.103+0000"},{"body":"[~xyao] thanks for updating the patch..{{TestTrashWithSecureEncryptionZones}} and {{TestSecureEncryptionZoneWithKMS}} also related.can you fix them also..?\n\n *Jenkins URL* \nhttps://builds.apache.org/job/PreCommit-HDFS-Build/17836/testReport/","created":"2016-12-12T13:26:38.138+0000"},{"body":"Thanks [~brahmareddy]. I've updated patch to cover both hadoop-common and haddop-hdfs tests. Also update the title and description.","created":"2016-12-12T16:02:56.467+0000"},{"body":"The failed unit test is not related to this change. It is tracked by https://issues.apache.org/jira/browse/HDFS-11131. \n\n","created":"2016-12-12T19:49:00.258+0000"},{"body":"+1 for the patch. Verified it fixes the unit tests. [~brahmareddy], do you have any additional comments?","created":"2016-12-12T21:39:29.704+0000"},{"body":"I plan to commit the patch by EOD today to fix the Jenkins issues unless [~brahmareddy] or other folks on the watchlist have additional comments. ","created":"2016-12-12T23:08:51.706+0000"},{"body":"This appears to be an attempt to hide an incompatibility introduced by HADOOP-13565. Generally, when a legitimate test breaks, the solution isn't to alter the test to conform to new incompatible behavior.\n","created":"2016-12-12T23:24:50.398+0000"},{"body":"Thanks [~daryn] for the comments. Yes, HADOOP-13565 enforces the check of SPNEGO SPNs to have three parts HTTP/host/realm. This will be incompatible which previously allows principal like HTTP/host before HADOOP-13565, assuming the default realm at authentication time.\n\nSince HTTP/host is a legitimate use case as you commented on HADOOP-13891, we can loosen the check added by HADOOP-13565 to allow it from KerberosAuthenticationHandler without modify the unit tests.","created":"2016-12-12T23:50:17.955+0000"},{"body":"Post a patch that fix the KerberosName parsing and release the checking to require a realm from KerberosAuthenticationHandler without modifying the failed unit tests. This way, we won't break compatibility for use case that use SPN in the form of HTTP/host assuming local realm like the failed unit tests. \n\nI've tested the patch locally against the failed tests and all of them passed. Please review, thanks!\n","created":"2016-12-13T01:29:49.248+0000"},{"body":"I guess these changes (hadoop-auth) are unlikely to trigger hadoop-common and hadoop-hdfs tests by Jenkins.","created":"2016-12-13T02:00:06.373+0000"},{"body":"In the patch v2, we remove the check for realm but keep the check for host as required based on [RFC-4559|https://tools.ietf.org/html/rfc4559]. \n{code}\n When the Kerberos Version 5 GSSAPI mechanism [RFC4121] is being used,\n the HTTP server will be using a principal name of the form of\n \"HTTP/hostname\".\n{code}\n\nIn other words, some valid UPN (User Principal Name) without hostname like HTTP@EXAMPLE.COM will be invalid for HTTP SPNEGO SPN (Service Principal Name). The RFC does not mention any requirement on realm. But based on many articles on multi-realm deployment, it is recommend to have HTTP/FQDN@Realm configured to avoid ambiguity and authentication problem in multi-realm use cases.","created":"2016-12-13T02:37:56.851+0000"},{"body":"+1 LGTM (non-binding). All hadoop-kms and hadoop-httpfs tests passed.","created":"2016-12-13T07:52:29.772+0000"},{"body":"[~xyao] Thanks for working on this JIRA.\nAfter applying your v3 patch to my local trunk branch, the \"Invalid SPNEGO sequence\" exception still exist in {{TestWebDelegationToken}}.\nHave I missed something?","created":"2016-12-14T02:37:05.358+0000"},{"body":"[~yuanbo], thanks for trying the patch. Can you post or attach the test logs of the failed test? Here is the result on my local machine after v3 patch, which has all passed.\n{code}\n-------------------------------------------------------\n T E S T S\n-------------------------------------------------------\nRunning org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken\nTests run: 12, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 4.559 sec - in org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken\n\nResults :\n\nTests run: 12, Failures: 0, Errors: 0, Skipped: 0\n\n{code}\n","created":"2016-12-14T03:25:00.787+0000"},{"body":"[~xyao] Thanks for your response.\nI've attached my test failure information and this is my git status info:\n{code}\n# Changes not staged for commit:\n# (use \"git add ...\" to update what will be committed)\n# (use \"git checkout -- ...\" to discard changes in working directory)\n#\n#\tmodified: hadoop-common-project/hadoop-auth/src/main/java/org/apache/hadoop/security/authentication/server/KerberosAuthenticationHandler.java\n#\tmodified: hadoop-common-project/hadoop-auth/src/main/java/org/apache/hadoop/security/authentication/util/KerberosName.java\n#\tmodified: hadoop-common-project/hadoop-auth/src/test/java/org/apache/hadoop/security/authentication/util/TestKerberosName.java\n#\nno changes added to commit (use \"git add\" and/or \"git commit -a\")\n{code}\n","created":"2016-12-14T03:54:08.727+0000"},{"body":"[~yuanbo], do you use Oracle JDK or IBM JDK? What's your JDK version? It passed on my machine with Oracle JDK 1.8.\n\nCan you enable trace log for KerberosAuthenticatonHandler during the test run by\n1. adding the following to the end of TestWebDelegationToken#setUp() \n{code}\n GenericTestUtils.setLogLevel(KerberosAuthenticationHandler.LOG, Level.TRACE);\n{code}\n2. change KerberosAuthenticationHandler.LOG to public in KerberosAuthenticator.java.\nand reattach the log with additional trace enabled? This will reveal the cause of SPNEGO failure from server side. Thanks in advance!","created":"2016-12-14T05:18:09.118+0000"},{"body":"Attach patch v04 to enable KerberosAuthenticationHandler trace log for tests in TestWebDelegationToken.java.\n","created":"2016-12-14T05:23:24.450+0000"},{"body":"[~yuanbo], the issue you hit is IBM JDK specific. TestWebDelegationToken was failing even without HADOOP-13565.\nBased on that, I think the failure is not related to either HADOOP-13565 or HADOOP-13890 we are trying to solve here. \nIf you want, we could fix IBM JDK issue for TestWebDelegationToken in a separate ticket later.\n\nIBM JDK 8\n{code}\n[root@c6404 hadoop]# /opt/ibm/java-x86_64-80/bin/java -version\njava version \"1.8.0\"\nJava(TM) SE Runtime Environment (build pxa6480sr3fp12-20160919_01(SR3 FP12))\nIBM J9 VM (build 2.8, JRE 1.8.0 Linux amd64-64 Compressed References 20160915_318796 (JIT enabled, AOT enabled)\nJ9VM - R28_Java8_SR3_20160915_0912_B318796\nJIT - tr.r14.java.green_20160818_122998\nGC - R28_Java8_SR3_20160915_0912_B318796_CMPRSS\nJ9CL - 20160915_318796)\nJCL - 20160914_01 based on Oracle jdk8u101-b13\n{code}\n\nGit info after revert HADOOP-13565.\n{code}\ncommit b5a719486112fa1ca60bee5eec81f41a0828b928\nAuthor: root \nDate: Wed Dec 14 07:10:36 2016 +0000\n Revert \"HADOOP-13565. KerberosAuthenticationHandler#authenticate should not rebuild SPN based on client request. Contributed by Xiaoyu Yao.\"\n \n This reverts commit 4c38f11cec0664b70e52f9563052dca8fb17c33f.\n{code} \n\nThe test still failed even after revert HADOOP-13565. \n{code}\nTests run: 12, Failures: 0, Errors: 2, Skipped: 0, Time elapsed: 10.748 sec <<< FAILURE! - in org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken\ntestKerberosDelegationTokenAuthenticator(org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken) Time elapsed: 1.571 sec <<< ERROR!\njavax.security.auth.login.LoginException: Bad JAAS configuration: unrecognized option: isInitiator\n at com.ibm.security.jgss.i18n.I18NException.throwLoginException(I18NException.java:23)\n at com.ibm.security.auth.module.Krb5LoginModule.d(Krb5LoginModule.java:57)\n at com.ibm.security.auth.module.Krb5LoginModule.a(Krb5LoginModule.java:686)\n at com.ibm.security.auth.module.Krb5LoginModule.login(Krb5LoginModule.java:214)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:95)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:55)\n at java.lang.reflect.Method.invoke(Method.java:508)\n at javax.security.auth.login.LoginContext.invoke(LoginContext.java:788)\n at javax.security.auth.login.LoginContext.access$000(LoginContext.java:196)\n at javax.security.auth.login.LoginContext$5.run(LoginContext.java:721)\n at javax.security.auth.login.LoginContext$5.run(LoginContext.java:719)\n at java.security.AccessController.doPrivileged(AccessController.java:686)\n at javax.security.auth.login.LoginContext.invokeCreatorPriv(LoginContext.java:719)\n at javax.security.auth.login.LoginContext.login(LoginContext.java:593)\n at org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken.doAsKerberosUser(TestWebDelegationToken.java:710)\n at org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken.testKerberosDelegationTokenAuthenticator(TestWebDelegationToken.java:778)\n at org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken.testKerberosDelegationTokenAuthenticator(TestWebDelegationToken.java:729)\n{code}","created":"2016-12-14T07:22:06.846+0000"},{"body":"[~xyao] Thanks for your new patch and explanation.\nMy java version is oracle-1.8, the details are here:\n{code}\njava version \"1.8.0_101\"\nJava(TM) SE Runtime Environment (build 1.8.0_101-b13)\nJava HotSpot(TM) 64-Bit Server VM (build 25.101-b13, mixed mode)\n\nll |grep java_sdk_1.8.0\nlrwxrwxrwx. 1 root root 58 Sep 6 16:06 java_sdk_1.8.0 -> /usr/lib/jvm/java-1.8.0-oracle-1.8.0.101-1jpp.1.el7.x86_64\n{code}\nThe exception from IBM JDK seems not to be \"Invalid SPNEGO sequence\" exception.\nFYI, I have attached a new log file: test_failure_1.txt.\nSince [~jzhuge], you and Apache Jenkins have reported that the test failures are passed, I'm strongly doubting these failures are related to my laptop environment. Please go ahead if you're suspended by my comment.","created":"2016-12-14T08:06:48.874+0000"},{"body":"[~yuanbo], here is what happened in your case.\n\n1. hostname {{localhost}} is mapped to principal {{HTTP/localhost}} during KerberosAuthenticationHandler.java:init.\n\n{code}\n2016-12-14 15:48:34,459 TRACE server.KerberosAuthenticationHandler (KerberosAuthenticationHandler.java:init(279)) - Map server: localhost to principal: HTTP/localhost\n{code}\n\n2. authenticate request comes in\n{code}\n2016-12-14 15:48:34,482 TRACE server.KerberosAuthenticationHandler (KerberosAuthenticationHandler.java:authenticate(400)) - SPNEGO starting for url: http://localhost:39910/foo/bar\n{code}\n\n3. The localhost to principal lookup somehow failed with an empty principal as shown below, which failed the test.\n{code}\n2016-12-14 15:48:34,495 TRACE server.KerberosAuthenticationHandler (KerberosAuthenticationHandler.java:run(421)) - SPNEGO with principals: []\n{code}\n\nThe only difference is in all the pass cases the HashMap lookup successfully find the right principal. I can't see obvious reason why the single principle is not being added into the HashMap during init(). I attach a new patch with additional tracing. [~yuanbo], can you try it out and post the result?\n\n{code}\n2016-12-13 21:12:43,918 TRACE server.KerberosAuthenticationHandler (KerberosAuthenticationHandler.java:run(421)) - SPNEGO with principals: [HTTP/localhost]\n{code}\n","created":"2016-12-14T16:36:03.513+0000"},{"body":"bq. Since John Zhuge, you and Apache Jenkins have reported that the test failures are passed, I'm strongly doubting these failures are related to my laptop environment. Please go ahead if you're suspended by my comment\n\nAgree. We can investigate your case as a separate issue. \nAppreciate if folks on the watch list can review this fix to unblock Jenkins. Thanks in advance!","created":"2016-12-14T18:21:31.970+0000"},{"body":"+1 for the v5 patch.","created":"2016-12-14T18:52:31.785+0000"},{"body":"cc: [~brahmareddy] and [~daryn], I plan to commit HADOOP-13890 shortly to unblock Jenkins and QE/test pipelines. \nIf you have additional comments regarding HADOOP-13565, we can discuss and address them with followup JIRAs.","created":"2016-12-14T18:57:05.249+0000"},{"body":"My late +1, [~xyao] thanks for taking care..Now it's addressed the incompatible which is introduced in HADOOP-13565 , [~daryn] do you think same..?\n\nStill tests can still fail in windows,may we can investigate further in follow up jira's..\n\nI feel,,,,It will be always good to run all the tests when we change common part ( which we feel,it can impact other projects).May be we change one line each project,Let jenkins run in all the projects which could have avoid this jira..","created":"2016-12-14T19:28:46.683+0000"},{"body":"[~brahmareddy], thanks for the review and tried the fix on windows. \nThe windows failure is not related to HADOOP-13565 or HADOOP-13890. I tried the TestWebDelegationToken with both changes reverted on Windows. It failed with the same error due to an issue with KerberosUtil#getDefaultRealm(). I filed a follow up ticket HADOOP-13907 to investigate and fix KerberosUtil#getDefaultRealm() on Windows.\n\n{code}\njava.lang.IllegalArgumentException: Can't get Kerberos realm\n\n\tat org.apache.hadoop.security.HadoopKerberosName.setConfiguration(HadoopKerberosName.java:65)\n\tat org.apache.hadoop.security.UserGroupInformation.initialize(UserGroupInformation.java:305)\n\tat org.apache.hadoop.security.UserGroupInformation.setConfiguration(UserGroupInformation.java:351)\n\tat org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken.testKerberosDelegationTokenAuthenticator(TestWebDelegationToken.java:746)\n\tat org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken.testKerberosDelegationTokenAuthenticator(TestWebDelegationToken.java:729)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:47)\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:44)\n\tat org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:26)\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:27)\n\tat org.junit.runners.ParentRunner.runLeaf(ParentRunner.java:271)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:70)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:50)\n\tat org.junit.runners.ParentRunner$3.run(ParentRunner.java:238)\n\tat org.junit.runners.ParentRunner$1.schedule(ParentRunner.java:63)\n\tat org.junit.runners.ParentRunner.runChildren(ParentRunner.java:236)\n\tat org.junit.runners.ParentRunner.access$000(ParentRunner.java:53)\n\tat org.junit.runners.ParentRunner$2.evaluate(ParentRunner.java:229)\n\tat org.junit.runners.ParentRunner.run(ParentRunner.java:309)\n\tat org.junit.runner.JUnitCore.run(JUnitCore.java:160)\n\tat com.intellij.junit4.JUnit4IdeaTestRunner.startRunnerWithArgs(JUnit4IdeaTestRunner.java:117)\n\tat com.intellij.junit4.JUnit4IdeaTestRunner.startRunnerWithArgs(JUnit4IdeaTestRunner.java:42)\n\tat com.intellij.rt.execution.junit.JUnitStarter.prepareStreamsAndStart(JUnitStarter.java:262)\n\tat com.intellij.rt.execution.junit.JUnitStarter.main(JUnitStarter.java:84)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat com.intellij.rt.execution.application.AppMain.main(AppMain.java:147)\nCaused by: java.lang.reflect.InvocationTargetException\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat org.apache.hadoop.security.authentication.util.KerberosUtil.getDefaultRealm(KerberosUtil.java:88)\n\tat org.apache.hadoop.security.HadoopKerberosName.setConfiguration(HadoopKerberosName.java:63)\n\t... 33 more\nCaused by: KrbException: Cannot locate default realm\n\tat sun.security.krb5.Config.getDefaultRealm(Config.java:1029)\n\t... 39 more\n{code}\n","created":"2016-12-14T21:25:11.066+0000"},{"body":"Thanks all for the reviews and discussions. I've commit the patch to trunk, branch-2 and branch-2.8\n\nI also updated the title/description of the incompatible issue introduced by HADOOP-13565 and fixed with this ticket.","created":"2016-12-14T22:04:11.554+0000"}],"conversations":[{"body":"*strong text*HADOOP-13565 introduced an incompatible check that disallowed principal like HTTP/host from being used as SPNEGO SPN. \r\nThis breaks the following test in trunk: TestWebDelegationToken, TestKMS , TestTrashWithSecureEncryptionZones and TestSecureEncryptionZoneWithKMS because they used HTTP/localhost as SPNEGO SPN assuming the default realm. This ticket is opened to bring back the support of HTTP/host as valid SPNEGO SPN. \r\n\r\nKerberosName parsing bug was discovered, fixed and included as a necessary part of this ticket along with additional unit test to cover parsing different form of principals. \r\n\r\n *Jenkins URL* \r\nhttps://builds.apache.org/job/hadoop-qbt-trunk-java8-linux-x86/251/testReport/\r\nhttps://builds.apache.org/job/PreCommit-HADOOP-Build/11240/testReport/","from":"reporter","subject":"Maintain HTTP/host as SPNEGO SPN support and fix KerberosName parsing "},{"body":"Saw them in pre-commit unit tests: https://builds.apache.org/job/PreCommit-HDFS-Build/17824/artifact/patchprocess/patch-unit-hadoop-common-project_hadoop-common.txt and https://builds.apache.org/job/PreCommit-HDFS-Build/17824/artifact/patchprocess/patch-unit-hadoop-hdfs-project_hadoop-hdfs.txt.\n","from":"developer"},{"body":"Attach a patch that fixed the unit tests.","from":"developer"},{"body":"Sorry I was not able to catch this at the commit time because Jenkins only run hadoop-auth tests but the failed tests are from hadoop-common. \n\n","from":"developer"},{"body":"[~xyao] thanks for updating the patch..{{TestTrashWithSecureEncryptionZones}} and {{TestSecureEncryptionZoneWithKMS}} also related.can you fix them also..?\n\n *Jenkins URL* \nhttps://builds.apache.org/job/PreCommit-HDFS-Build/17836/testReport/","from":"developer"},{"body":"Thanks [~brahmareddy]. I've updated patch to cover both hadoop-common and haddop-hdfs tests. Also update the title and description.","from":"developer"},{"body":"The failed unit test is not related to this change. It is tracked by https://issues.apache.org/jira/browse/HDFS-11131. \n\n","from":"developer"},{"body":"+1 for the patch. Verified it fixes the unit tests. [~brahmareddy], do you have any additional comments?","from":"developer"},{"body":"I plan to commit the patch by EOD today to fix the Jenkins issues unless [~brahmareddy] or other folks on the watchlist have additional comments. ","from":"developer"},{"body":"This appears to be an attempt to hide an incompatibility introduced by HADOOP-13565. Generally, when a legitimate test breaks, the solution isn't to alter the test to conform to new incompatible behavior.\n","from":"developer"},{"body":"Thanks [~daryn] for the comments. Yes, HADOOP-13565 enforces the check of SPNEGO SPNs to have three parts HTTP/host/realm. This will be incompatible which previously allows principal like HTTP/host before HADOOP-13565, assuming the default realm at authentication time.\n\nSince HTTP/host is a legitimate use case as you commented on HADOOP-13891, we can loosen the check added by HADOOP-13565 to allow it from KerberosAuthenticationHandler without modify the unit tests.","from":"developer"},{"body":"Post a patch that fix the KerberosName parsing and release the checking to require a realm from KerberosAuthenticationHandler without modifying the failed unit tests. This way, we won't break compatibility for use case that use SPN in the form of HTTP/host assuming local realm like the failed unit tests. \n\nI've tested the patch locally against the failed tests and all of them passed. Please review, thanks!\n","from":"developer"},{"body":"I guess these changes (hadoop-auth) are unlikely to trigger hadoop-common and hadoop-hdfs tests by Jenkins.","from":"developer"},{"body":"In the patch v2, we remove the check for realm but keep the check for host as required based on [RFC-4559|https://tools.ietf.org/html/rfc4559]. \n{code}\n When the Kerberos Version 5 GSSAPI mechanism [RFC4121] is being used,\n the HTTP server will be using a principal name of the form of\n \"HTTP/hostname\".\n{code}\n\nIn other words, some valid UPN (User Principal Name) without hostname like HTTP@EXAMPLE.COM will be invalid for HTTP SPNEGO SPN (Service Principal Name). The RFC does not mention any requirement on realm. But based on many articles on multi-realm deployment, it is recommend to have HTTP/FQDN@Realm configured to avoid ambiguity and authentication problem in multi-realm use cases.","from":"developer"},{"body":"+1 LGTM (non-binding). All hadoop-kms and hadoop-httpfs tests passed.","from":"developer"},{"body":"[~xyao] Thanks for working on this JIRA.\nAfter applying your v3 patch to my local trunk branch, the \"Invalid SPNEGO sequence\" exception still exist in {{TestWebDelegationToken}}.\nHave I missed something?","from":"developer"},{"body":"[~yuanbo], thanks for trying the patch. Can you post or attach the test logs of the failed test? Here is the result on my local machine after v3 patch, which has all passed.\n{code}\n-------------------------------------------------------\n T E S T S\n-------------------------------------------------------\nRunning org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken\nTests run: 12, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 4.559 sec - in org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken\n\nResults :\n\nTests run: 12, Failures: 0, Errors: 0, Skipped: 0\n\n{code}\n","from":"developer"},{"body":"[~xyao] Thanks for your response.\nI've attached my test failure information and this is my git status info:\n{code}\n# Changes not staged for commit:\n# (use \"git add ...\" to update what will be committed)\n# (use \"git checkout -- ...\" to discard changes in working directory)\n#\n#\tmodified: hadoop-common-project/hadoop-auth/src/main/java/org/apache/hadoop/security/authentication/server/KerberosAuthenticationHandler.java\n#\tmodified: hadoop-common-project/hadoop-auth/src/main/java/org/apache/hadoop/security/authentication/util/KerberosName.java\n#\tmodified: hadoop-common-project/hadoop-auth/src/test/java/org/apache/hadoop/security/authentication/util/TestKerberosName.java\n#\nno changes added to commit (use \"git add\" and/or \"git commit -a\")\n{code}\n","from":"developer"},{"body":"[~yuanbo], do you use Oracle JDK or IBM JDK? What's your JDK version? It passed on my machine with Oracle JDK 1.8.\n\nCan you enable trace log for KerberosAuthenticatonHandler during the test run by\n1. adding the following to the end of TestWebDelegationToken#setUp() \n{code}\n GenericTestUtils.setLogLevel(KerberosAuthenticationHandler.LOG, Level.TRACE);\n{code}\n2. change KerberosAuthenticationHandler.LOG to public in KerberosAuthenticator.java.\nand reattach the log with additional trace enabled? This will reveal the cause of SPNEGO failure from server side. Thanks in advance!","from":"developer"},{"body":"Attach patch v04 to enable KerberosAuthenticationHandler trace log for tests in TestWebDelegationToken.java.\n","from":"developer"},{"body":"[~yuanbo], the issue you hit is IBM JDK specific. TestWebDelegationToken was failing even without HADOOP-13565.\nBased on that, I think the failure is not related to either HADOOP-13565 or HADOOP-13890 we are trying to solve here. \nIf you want, we could fix IBM JDK issue for TestWebDelegationToken in a separate ticket later.\n\nIBM JDK 8\n{code}\n[root@c6404 hadoop]# /opt/ibm/java-x86_64-80/bin/java -version\njava version \"1.8.0\"\nJava(TM) SE Runtime Environment (build pxa6480sr3fp12-20160919_01(SR3 FP12))\nIBM J9 VM (build 2.8, JRE 1.8.0 Linux amd64-64 Compressed References 20160915_318796 (JIT enabled, AOT enabled)\nJ9VM - R28_Java8_SR3_20160915_0912_B318796\nJIT - tr.r14.java.green_20160818_122998\nGC - R28_Java8_SR3_20160915_0912_B318796_CMPRSS\nJ9CL - 20160915_318796)\nJCL - 20160914_01 based on Oracle jdk8u101-b13\n{code}\n\nGit info after revert HADOOP-13565.\n{code}\ncommit b5a719486112fa1ca60bee5eec81f41a0828b928\nAuthor: root \nDate: Wed Dec 14 07:10:36 2016 +0000\n Revert \"HADOOP-13565. KerberosAuthenticationHandler#authenticate should not rebuild SPN based on client request. Contributed by Xiaoyu Yao.\"\n \n This reverts commit 4c38f11cec0664b70e52f9563052dca8fb17c33f.\n{code} \n\nThe test still failed even after revert HADOOP-13565. \n{code}\nTests run: 12, Failures: 0, Errors: 2, Skipped: 0, Time elapsed: 10.748 sec <<< FAILURE! - in org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken\ntestKerberosDelegationTokenAuthenticator(org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken) Time elapsed: 1.571 sec <<< ERROR!\njavax.security.auth.login.LoginException: Bad JAAS configuration: unrecognized option: isInitiator\n at com.ibm.security.jgss.i18n.I18NException.throwLoginException(I18NException.java:23)\n at com.ibm.security.auth.module.Krb5LoginModule.d(Krb5LoginModule.java:57)\n at com.ibm.security.auth.module.Krb5LoginModule.a(Krb5LoginModule.java:686)\n at com.ibm.security.auth.module.Krb5LoginModule.login(Krb5LoginModule.java:214)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:95)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:55)\n at java.lang.reflect.Method.invoke(Method.java:508)\n at javax.security.auth.login.LoginContext.invoke(LoginContext.java:788)\n at javax.security.auth.login.LoginContext.access$000(LoginContext.java:196)\n at javax.security.auth.login.LoginContext$5.run(LoginContext.java:721)\n at javax.security.auth.login.LoginContext$5.run(LoginContext.java:719)\n at java.security.AccessController.doPrivileged(AccessController.java:686)\n at javax.security.auth.login.LoginContext.invokeCreatorPriv(LoginContext.java:719)\n at javax.security.auth.login.LoginContext.login(LoginContext.java:593)\n at org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken.doAsKerberosUser(TestWebDelegationToken.java:710)\n at org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken.testKerberosDelegationTokenAuthenticator(TestWebDelegationToken.java:778)\n at org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken.testKerberosDelegationTokenAuthenticator(TestWebDelegationToken.java:729)\n{code}","from":"developer"},{"body":"[~xyao] Thanks for your new patch and explanation.\nMy java version is oracle-1.8, the details are here:\n{code}\njava version \"1.8.0_101\"\nJava(TM) SE Runtime Environment (build 1.8.0_101-b13)\nJava HotSpot(TM) 64-Bit Server VM (build 25.101-b13, mixed mode)\n\nll |grep java_sdk_1.8.0\nlrwxrwxrwx. 1 root root 58 Sep 6 16:06 java_sdk_1.8.0 -> /usr/lib/jvm/java-1.8.0-oracle-1.8.0.101-1jpp.1.el7.x86_64\n{code}\nThe exception from IBM JDK seems not to be \"Invalid SPNEGO sequence\" exception.\nFYI, I have attached a new log file: test_failure_1.txt.\nSince [~jzhuge], you and Apache Jenkins have reported that the test failures are passed, I'm strongly doubting these failures are related to my laptop environment. Please go ahead if you're suspended by my comment.","from":"developer"},{"body":"[~yuanbo], here is what happened in your case.\n\n1. hostname {{localhost}} is mapped to principal {{HTTP/localhost}} during KerberosAuthenticationHandler.java:init.\n\n{code}\n2016-12-14 15:48:34,459 TRACE server.KerberosAuthenticationHandler (KerberosAuthenticationHandler.java:init(279)) - Map server: localhost to principal: HTTP/localhost\n{code}\n\n2. authenticate request comes in\n{code}\n2016-12-14 15:48:34,482 TRACE server.KerberosAuthenticationHandler (KerberosAuthenticationHandler.java:authenticate(400)) - SPNEGO starting for url: http://localhost:39910/foo/bar\n{code}\n\n3. The localhost to principal lookup somehow failed with an empty principal as shown below, which failed the test.\n{code}\n2016-12-14 15:48:34,495 TRACE server.KerberosAuthenticationHandler (KerberosAuthenticationHandler.java:run(421)) - SPNEGO with principals: []\n{code}\n\nThe only difference is in all the pass cases the HashMap lookup successfully find the right principal. I can't see obvious reason why the single principle is not being added into the HashMap during init(). I attach a new patch with additional tracing. [~yuanbo], can you try it out and post the result?\n\n{code}\n2016-12-13 21:12:43,918 TRACE server.KerberosAuthenticationHandler (KerberosAuthenticationHandler.java:run(421)) - SPNEGO with principals: [HTTP/localhost]\n{code}\n","from":"developer"},{"body":"bq. Since John Zhuge, you and Apache Jenkins have reported that the test failures are passed, I'm strongly doubting these failures are related to my laptop environment. Please go ahead if you're suspended by my comment\n\nAgree. We can investigate your case as a separate issue. \nAppreciate if folks on the watch list can review this fix to unblock Jenkins. Thanks in advance!","from":"developer"},{"body":"+1 for the v5 patch.","from":"developer"},{"body":"cc: [~brahmareddy] and [~daryn], I plan to commit HADOOP-13890 shortly to unblock Jenkins and QE/test pipelines. \nIf you have additional comments regarding HADOOP-13565, we can discuss and address them with followup JIRAs.","from":"developer"},{"body":"My late +1, [~xyao] thanks for taking care..Now it's addressed the incompatible which is introduced in HADOOP-13565 , [~daryn] do you think same..?\n\nStill tests can still fail in windows,may we can investigate further in follow up jira's..\n\nI feel,,,,It will be always good to run all the tests when we change common part ( which we feel,it can impact other projects).May be we change one line each project,Let jenkins run in all the projects which could have avoid this jira..","from":"developer"},{"body":"[~brahmareddy], thanks for the review and tried the fix on windows. \nThe windows failure is not related to HADOOP-13565 or HADOOP-13890. I tried the TestWebDelegationToken with both changes reverted on Windows. It failed with the same error due to an issue with KerberosUtil#getDefaultRealm(). I filed a follow up ticket HADOOP-13907 to investigate and fix KerberosUtil#getDefaultRealm() on Windows.\n\n{code}\njava.lang.IllegalArgumentException: Can't get Kerberos realm\n\n\tat org.apache.hadoop.security.HadoopKerberosName.setConfiguration(HadoopKerberosName.java:65)\n\tat org.apache.hadoop.security.UserGroupInformation.initialize(UserGroupInformation.java:305)\n\tat org.apache.hadoop.security.UserGroupInformation.setConfiguration(UserGroupInformation.java:351)\n\tat org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken.testKerberosDelegationTokenAuthenticator(TestWebDelegationToken.java:746)\n\tat org.apache.hadoop.security.token.delegation.web.TestWebDelegationToken.testKerberosDelegationTokenAuthenticator(TestWebDelegationToken.java:729)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:47)\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:44)\n\tat org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:26)\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:27)\n\tat org.junit.runners.ParentRunner.runLeaf(ParentRunner.java:271)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:70)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:50)\n\tat org.junit.runners.ParentRunner$3.run(ParentRunner.java:238)\n\tat org.junit.runners.ParentRunner$1.schedule(ParentRunner.java:63)\n\tat org.junit.runners.ParentRunner.runChildren(ParentRunner.java:236)\n\tat org.junit.runners.ParentRunner.access$000(ParentRunner.java:53)\n\tat org.junit.runners.ParentRunner$2.evaluate(ParentRunner.java:229)\n\tat org.junit.runners.ParentRunner.run(ParentRunner.java:309)\n\tat org.junit.runner.JUnitCore.run(JUnitCore.java:160)\n\tat com.intellij.junit4.JUnit4IdeaTestRunner.startRunnerWithArgs(JUnit4IdeaTestRunner.java:117)\n\tat com.intellij.junit4.JUnit4IdeaTestRunner.startRunnerWithArgs(JUnit4IdeaTestRunner.java:42)\n\tat com.intellij.rt.execution.junit.JUnitStarter.prepareStreamsAndStart(JUnitStarter.java:262)\n\tat com.intellij.rt.execution.junit.JUnitStarter.main(JUnitStarter.java:84)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat com.intellij.rt.execution.application.AppMain.main(AppMain.java:147)\nCaused by: java.lang.reflect.InvocationTargetException\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat org.apache.hadoop.security.authentication.util.KerberosUtil.getDefaultRealm(KerberosUtil.java:88)\n\tat org.apache.hadoop.security.HadoopKerberosName.setConfiguration(HadoopKerberosName.java:63)\n\t... 33 more\nCaused by: KrbException: Cannot locate default realm\n\tat sun.security.krb5.Config.getDefaultRealm(Config.java:1029)\n\t... 39 more\n{code}\n","from":"developer"},{"body":"Thanks all for the reviews and discussions. I've commit the patch to trunk, branch-2 and branch-2.8\n\nI also updated the title/description of the incompatible issue introduced by HADOOP-13565 and fixed with this ticket.","from":"developer"}],"created":"2016-12-11T07:42:41.000+0000","description":"*strong text*HADOOP-13565 introduced an incompatible check that disallowed principal like HTTP/host from being used as SPNEGO SPN. \r\nThis breaks the following test in trunk: TestWebDelegationToken, TestKMS , TestTrashWithSecureEncryptionZones and TestSecureEncryptionZoneWithKMS because they used HTTP/localhost as SPNEGO SPN assuming the default realm. This ticket is opened to bring back the support of HTTP/host as valid SPNEGO SPN. \r\n\r\nKerberosName parsing bug was discovered, fixed and included as a necessary part of this ticket along with additional unit test to cover parsing different form of principals. \r\n\r\n *Jenkins URL* \r\nhttps://builds.apache.org/job/hadoop-qbt-trunk-java8-linux-x86/251/testReport/\r\nhttps://builds.apache.org/job/PreCommit-HADOOP-Build/11240/testReport/","issue_id":"13027245","key":"HADOOP-13890","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2016-12-14T22:04:11.000+0000","role":"fixed_distractor","summary":"Maintain HTTP/host as SPNEGO SPN support and fix KerberosName parsing "} {"case_id":"13027862","cluster":"DISTRACTOR-HADOOP-13901","comments":[{"body":"The files are generated by other plugins, so we can ignore them. Attaching a simple patch to ignore them.","created":"2016-12-13T18:18:44.747+0000"},{"body":"Thanks [~ajisakaa] for fixing the issue, quite annoying. I will try the patch when I run dev-support/bin/test-patch locally.","created":"2016-12-13T18:44:25.186+0000"},{"body":"Thanks for reporting and providing a patch.\n\n{code}\n326\t .settings/org.eclipse.jdt.core.prefs\n{code}\n\nLooks like we can exclude the {{.settings}} directory?","created":"2016-12-13T19:23:38.163+0000"},{"body":"Still got 2 ASF license issues for dev-support/bin/test-patch after applying the patch:\n{noformat}\nLines that start with ????? in the ASF License report indicate files that do not have an Apache license header:\n !????? /home/jzhuge/hadoop/hadoop-tools/hadoop-ant/target/antrun/build-main.xml\n !????? /home/jzhuge/hadoop/hadoop-tools/hadoop-ant/target/.plxarc\n{noformat}","created":"2016-12-13T19:29:25.790+0000"},{"body":"the entries for {{.classpath}}, {{.project}}, and {{.settings}} only get flagged if a module is removed from the project (so like in precommit when something is moved, or if someone was switching between branches without cleaning up). I'm not sure we need to exclude them.\n\n[~jzhuge] is there a {{hadoop-tools/hadoop-ant}} module in your checkout that corresponds to those target directories? the RAT plugin should already be ignoring the {{target}} directory for anything maven knows about.","created":"2016-12-13T19:32:37.053+0000"},{"body":"My changes removed {{maven-antrun-plugin}} sections from {{hadoop-kms/pom.xml}} and {{hadoop-hdfs-httpfs/pom.xml}}.","created":"2016-12-13T22:57:54.491+0000"},{"body":"bq. Looks like we can exclude the .settings directory?\nAgreed. I'll update the patch.\n\nbq. Still got 2 ASF license issues for dev-support/bin/test-patch after applying the patch:\nHi [~jzhuge], hadoop-ant module has been removed in trunk. Would you run {{git clean -d -f}} and retry my patch?","created":"2016-12-14T02:54:55.713+0000"},{"body":"[~ajisakaa] No more ASF license issue after git clean. Sorry for the noise.","created":"2016-12-14T04:40:13.308+0000"},{"body":"No problem!","created":"2016-12-14T07:23:23.141+0000"},{"body":"02 patch: excluded the .settings directory.","created":"2016-12-14T07:23:34.336+0000"},{"body":"+1 LGTM (non-binding)","created":"2016-12-14T07:57:18.466+0000"},{"body":"Does this patch still apply? \n\nIf so: +1","created":"2017-04-28T11:41:59.921+0000"},{"body":"Rebased.","created":"2017-04-29T15:03:41.656+0000"},{"body":"+1\n","created":"2017-04-29T17:03:40.841+0000"},{"body":"Committed this to trunk. Thanks [~stevel@apache.org] for the review.","created":"2017-05-01T06:33:21.535+0000"}],"conversations":[{"body":"Hadoop-side of YETUS-473.\nhttps://builds.apache.org/job/PreCommit-HADOOP-Build/11239/artifact/patchprocess/patch-asflicense-problems.txt\n{noformat}\nLines that start with ????? in the ASF License report indicate files that do not have an Apache license header:\n !????? /testptch/hadoop/hadoop-build-tools/maven-eclipse.xml\n !????? /testptch/hadoop/hadoop-build-tools/.externalToolBuilders/Maven_Ant_Builder.launch\n !????? hadoop-client/.classpath\n !????? hadoop-client/.project\n !????? hadoop-client/.settings/org.eclipse.jdt.core.prefs\n{noformat}","from":"reporter","subject":"Fix ASF License warnings"},{"body":"The files are generated by other plugins, so we can ignore them. Attaching a simple patch to ignore them.","from":"developer"},{"body":"Thanks [~ajisakaa] for fixing the issue, quite annoying. I will try the patch when I run dev-support/bin/test-patch locally.","from":"developer"},{"body":"Thanks for reporting and providing a patch.\n\n{code}\n326\t .settings/org.eclipse.jdt.core.prefs\n{code}\n\nLooks like we can exclude the {{.settings}} directory?","from":"developer"},{"body":"Still got 2 ASF license issues for dev-support/bin/test-patch after applying the patch:\n{noformat}\nLines that start with ????? in the ASF License report indicate files that do not have an Apache license header:\n !????? /home/jzhuge/hadoop/hadoop-tools/hadoop-ant/target/antrun/build-main.xml\n !????? /home/jzhuge/hadoop/hadoop-tools/hadoop-ant/target/.plxarc\n{noformat}","from":"developer"},{"body":"the entries for {{.classpath}}, {{.project}}, and {{.settings}} only get flagged if a module is removed from the project (so like in precommit when something is moved, or if someone was switching between branches without cleaning up). I'm not sure we need to exclude them.\n\n[~jzhuge] is there a {{hadoop-tools/hadoop-ant}} module in your checkout that corresponds to those target directories? the RAT plugin should already be ignoring the {{target}} directory for anything maven knows about.","from":"developer"},{"body":"My changes removed {{maven-antrun-plugin}} sections from {{hadoop-kms/pom.xml}} and {{hadoop-hdfs-httpfs/pom.xml}}.","from":"developer"},{"body":"bq. Looks like we can exclude the .settings directory?\nAgreed. I'll update the patch.\n\nbq. Still got 2 ASF license issues for dev-support/bin/test-patch after applying the patch:\nHi [~jzhuge], hadoop-ant module has been removed in trunk. Would you run {{git clean -d -f}} and retry my patch?","from":"developer"},{"body":"[~ajisakaa] No more ASF license issue after git clean. Sorry for the noise.","from":"developer"},{"body":"No problem!","from":"developer"},{"body":"02 patch: excluded the .settings directory.","from":"developer"},{"body":"+1 LGTM (non-binding)","from":"developer"},{"body":"Does this patch still apply? \n\nIf so: +1","from":"developer"},{"body":"Rebased.","from":"developer"},{"body":"+1\n","from":"developer"},{"body":"Committed this to trunk. Thanks [~stevel@apache.org] for the review.","from":"developer"}],"created":"2016-12-13T18:16:22.000+0000","description":"Hadoop-side of YETUS-473.\nhttps://builds.apache.org/job/PreCommit-HADOOP-Build/11239/artifact/patchprocess/patch-asflicense-problems.txt\n{noformat}\nLines that start with ????? in the ASF License report indicate files that do not have an Apache license header:\n !????? /testptch/hadoop/hadoop-build-tools/maven-eclipse.xml\n !????? /testptch/hadoop/hadoop-build-tools/.externalToolBuilders/Maven_Ant_Builder.launch\n !????? hadoop-client/.classpath\n !????? hadoop-client/.project\n !????? hadoop-client/.settings/org.eclipse.jdt.core.prefs\n{noformat}","issue_id":"13027862","key":"HADOOP-13901","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2017-05-01T06:33:21.000+0000","role":"fixed_distractor","summary":"Fix ASF License warnings"} {"case_id":"13041143","cluster":"DISTRACTOR-HADOOP-14067","comments":[{"body":"There are already people who have created workarounds for this by copying the VersionInfo class - https://github.com/timveil/hive-jdbc-uber-jar#note-about-kerberos-and-the-workaround. That obviously is not a maintainable solution for this problem.\n\nWe are seeing this with Hive's jdbc jar when used with tools such as SQuirreL SQL or dbVisualizer","created":"2017-02-07T21:52:11.994+0000"},{"body":"[~thejas]\r\n\r\nThe patch is having checkstyle issues and javadoc errors.\r\n\r\nOther than that +1 from me.\r\n\r\n \r\n\r\n ","created":"2018-01-26T22:04:43.496+0000"},{"body":"+1\r\n\r\n[~thejas], please review the javadoc, checkstyle issues in the patch.","created":"2018-03-19T22:00:22.228+0000"},{"body":"Attaching 02.patch with checkstyle, javadoc fixes.\r\n","created":"2018-03-20T02:22:19.960+0000"},{"body":"03.patch - fix remaining checkstyle issue (verified locally)","created":"2018-03-20T19:14:32.702+0000"},{"body":"Uploading 03.patch for real this time!\r\n","created":"2018-03-22T00:42:00.892+0000"},{"body":"+1 (now checkstyle issues fixed)","created":"2018-03-22T06:31:50.581+0000"},{"body":"+1, Thanks for addressing style/javadoc issues. I will commit shortly.","created":"2018-03-22T15:49:58.517+0000"},{"body":"I have committed this to trunk. Thanks [~thejas]!","created":"2018-03-22T21:15:26.490+0000"},{"body":"Cherry-picked to branch-3.1","created":"2018-03-27T17:25:10.791+0000"},{"body":"Doing 3.1.0 RC1 now, moved all 3.1.1 (branch-3.1) fixes to 3.1.0 (branch-3.1.0)","created":"2018-03-29T16:57:36.413+0000"}],"conversations":[{"body":"org.apache.hadoop.util.VersionInfo loads the version-info.properties file via the current thread classloader.\nHowever, in case of applications that are using hadoop classes dynamically (eg jdbc based tools such as SQuirreL SQL) the current thread might not be the one that loaded the hadoop classes including VersionInfo, and it would fail to fine the properties file.\nThe right place to look for the properties file is in the classloader of VersionInfo class, as right version is the one that is associated with rest of the loaded hadoop classes, and not necessarily the one in current thread classloader.\n\nCreated a related jira - HADOOP-14066 to make methods to get version via VersionInfo a public api.","from":"reporter","subject":"VersionInfo should load version-info.properties from its own classloader"},{"body":"There are already people who have created workarounds for this by copying the VersionInfo class - https://github.com/timveil/hive-jdbc-uber-jar#note-about-kerberos-and-the-workaround. That obviously is not a maintainable solution for this problem.\n\nWe are seeing this with Hive's jdbc jar when used with tools such as SQuirreL SQL or dbVisualizer","from":"developer"},{"body":"[~thejas]\r\n\r\nThe patch is having checkstyle issues and javadoc errors.\r\n\r\nOther than that +1 from me.\r\n\r\n \r\n\r\n ","from":"developer"},{"body":"+1\r\n\r\n[~thejas], please review the javadoc, checkstyle issues in the patch.","from":"developer"},{"body":"Attaching 02.patch with checkstyle, javadoc fixes.\r\n","from":"developer"},{"body":"03.patch - fix remaining checkstyle issue (verified locally)","from":"developer"},{"body":"Uploading 03.patch for real this time!\r\n","from":"developer"},{"body":"+1 (now checkstyle issues fixed)","from":"developer"},{"body":"+1, Thanks for addressing style/javadoc issues. I will commit shortly.","from":"developer"},{"body":"I have committed this to trunk. Thanks [~thejas]!","from":"developer"},{"body":"Cherry-picked to branch-3.1","from":"developer"},{"body":"Doing 3.1.0 RC1 now, moved all 3.1.1 (branch-3.1) fixes to 3.1.0 (branch-3.1.0)","from":"developer"}],"created":"2017-02-07T21:49:13.000+0000","description":"org.apache.hadoop.util.VersionInfo loads the version-info.properties file via the current thread classloader.\nHowever, in case of applications that are using hadoop classes dynamically (eg jdbc based tools such as SQuirreL SQL) the current thread might not be the one that loaded the hadoop classes including VersionInfo, and it would fail to fine the properties file.\nThe right place to look for the properties file is in the classloader of VersionInfo class, as right version is the one that is associated with rest of the loaded hadoop classes, and not necessarily the one in current thread classloader.\n\nCreated a related jira - HADOOP-14066 to make methods to get version via VersionInfo a public api.","issue_id":"13041143","key":"HADOOP-14067","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2018-03-22T21:15:26.000+0000","role":"fixed_distractor","summary":"VersionInfo should load version-info.properties from its own classloader"} {"case_id":"13074392","cluster":"DISTRACTOR-HADOOP-14451","comments":[{"body":"Below are bits from stacktrace:\n\nThread1\n{code}\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.fs.RawLocalFileSystem.setPermission(RawLocalFileSystem.java:739)\n\tat org.apache.hadoop.fs.RawLocalFileSystem$LocalFSFileOutputStream.(RawLocalFileSystem.java:224)\n\tat org.apache.hadoop.fs.RawLocalFileSystem$LocalFSFileOutputStream.(RawLocalFileSystem.java:208)\n{code}\n\t\nThread2\n{code}\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.mapred.FadvisedFileRegion.transferSuccessful(FadvisedFileRegion.java:160)\n\tat org.apache.hadoop.mapred.ShuffleHandler$Shuffle$1.operationComplete(ShuffleHandler.java:1166)\n\tat org.jboss.netty.channel.DefaultChannelFuture.notifyListener(DefaultChannelFuture.java:427)\n\tat org.jboss.netty.channel.DefaultChannelFuture.notifyListeners(DefaultChannelFuture.java:413)\n{code}\n\n\t\nThe above threads looks to be blocked by below two stacks\n\t\nStack1:\n{code}\n\"New I/O worker #1\" #135 prio=5 os_prio=0 tid=0x00007f1f60817800 nid=0x697d in Object.wait() [0x00007f1f4429a000]\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.io.nativeio.NativeIO$POSIX.(NativeIO.java:184)\n\tat org.apache.hadoop.mapred.FadvisedFileRegion.transferSuccessful(FadvisedFileRegion.java:160)\n\tat org.apache.hadoop.mapred.ShuffleHandler$Shuffle$1.operationComplete(ShuffleHandler.java:1166)\n{code}\n\t\nStack2:\n{code}\n\"ContainersLauncher #16\" #365 prio=5 os_prio=0 tid=0x00007f1f49c8a800 nid=0x7cd0 in Object.wait() [0x00007f1f32891000]\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.io.nativeio.NativeIO.initNative(Native Method)\n\tat org.apache.hadoop.io.nativeio.NativeIO.(NativeIO.java:645)\n\tat org.apache.hadoop.fs.RawLocalFileSystem.setPermission(RawLocalFileSystem.java:739)\n{code}\n\t\n*Stack1* is blocked by *Stack2* as *Stack1* thread needs *NativeIO* class initialization to finish, so the problematic stack looks to be Stack 2\t\nNext in Stack 2 @ NativeIO.java:645 initNative in a native call, and it tries to initialize the native-hadoop library\n\nso in Stack 2: it try to do this in NativeIO.c via *Java_org_apache_hadoop_io_nativeio_NativeIO_initNative*\n{code}static void consts_init(JNIEnv *env) {\n jclass clazz = (*env)->FindClass(env, NATIVE_IO_POSIX_CLASS);\n #define NATIVE_IO_POSIX_CLASS \"org/apache/hadoop/io/nativeio/NativeIO$POSIX\"\n i.e create class org.apache.hadoop.io.nativeio.NativeIO$POSIX{code}\n \nbut *Stack1* is already in {{org.apache.hadoop.io.nativeio.NativeIO$POSIX.}}\nso it deadlock and all threads hang","created":"2017-05-24T03:45:21.813+0000"},{"body":"A flavor of https://bugs.openjdk.java.net/browse/JDK-8037567 but not same, rather caused by application code","created":"2017-05-24T03:49:53.987+0000"},{"body":"attaching affected nodemanager jstack\nFor the background of the issue, test NM seem to have all containers timeout after restart, and hence from jstack, arrived at above conclusion with deadlock in NativeIO","created":"2017-05-24T13:34:22.091+0000"},{"body":"Some references\nhttps://docs.oracle.com/javase/8/docs/technotes/guides/jni/spec/functions.html#FindClass \nhttps://docs.oracle.com/javase/specs/jls/se8/html/jls-12.html#jls-12.4.2 ","created":"2017-05-24T13:51:05.032+0000"},{"body":"Thank you [~ajithshetty] .Analysis makes sense to me. The same could cause dead lock.\nTried a sample program to simulate.\n{code}\n\t\tThread one=new Thread(){\n\t\t\t@Override\n\t\t\tpublic void run() {\n\t\t\t\tNativeIO.isAvailable();\n\t\t\t}\n\t\t};\t\t\n\t\tThread two=new Thread(){\n\t\t\t@Override\n\t\t\tpublic void run() {\n\t\t\t\tNativeIO.POSIX.isAvailable();\n\t\t}\n\t\t};\n\t\ttwo.start();\t\n\t\tone.start();\n\t\tone.join();\n\t\ttwo.join();\n{code}\n*stack trace*\n{code}\n\"Thread-1\" #9 prio=5 os_prio=0 tid=0x00007f1b90c5f000 nid=0x6843 in Object.wait() [0x00007f1b80ba3000]\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.io.nativeio.NativeIO.initNative(Native Method)\n\tat org.apache.hadoop.io.nativeio.NativeIO.(NativeIO.java:644)\n\tat org.apache.hadoop.io.nativeio.TestNative$1.run(TestNative.java:8)\n\n\"Thread-2\" #10 prio=5 os_prio=0 tid=0x00007f1b90c58000 nid=0x6842 in Object.wait() [0x00007f1b80ca4000]\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.io.nativeio.NativeIO$POSIX.(NativeIO.java:184)\n\tat org.apache.hadoop.io.nativeio.TestNative$2.run(TestNative.java:15)\n{code}\n\nAffected versions are *2.8.0 and hadoop3*\n\n\ncc/ [~cmccabe_impala_fa3f] ","created":"2017-05-24T18:57:16.453+0000"},{"body":"Hi [~vinayrpet] , [~rakesh_r] , [~martinw] any thoughtss??","created":"2017-05-25T09:25:41.691+0000"},{"body":"I am not actively working on this, so my thoughts may be wide of the mark. Does this issue occur with\n\n{code}\nif (NativeIO.isAvailable()) {\n NativeIO.POSIX.getCacheManipulator().posixFadviseIfPossible(identifier, fd, getPosition(), getCount(), POSIX_FADV_DONTNEED);\n}\n{code}\n\n\n\n\n","created":"2017-05-25T10:18:45.207+0000"},{"body":"[~martinw]\nWe tried the same the approach could solve the problem. But in all {{NativeIO.POSIX}} usage location we have to check.\n[~ajithshetty] Any other solution in your mind?","created":"2017-05-25T10:37:54.801+0000"},{"body":"Only other suggestion I can think of is to break down {{initNative()}} into smaller functions, e.g. {{initNativePosix()}}, {{initNativeStat()}}... Each of which is initialised by a static block within its inner class. ","created":"2017-05-25T12:01:47.660+0000"},{"body":"Attached a simple approach of avoiding deadlock.","created":"2017-05-25T12:17:46.871+0000"},{"body":"i think people here are more interested in uploading a patch rather than discussing with the submitter if analysis and approach is right. If i have added well amount of analysis, i would probably have a solution too. Anyways, feel free to assign to self if that's the case","created":"2017-05-25T12:55:19.527+0000"},{"body":"Looks like the 'simple' approach is not working for this after all for this Native IO.\n [~ajithshetty], Please continue your analysis.","created":"2017-05-25T14:47:45.637+0000"},{"body":"Uploaded the patch, as per [~martinw]'s suggestion.\r\n\r\n1. Split into separate code blocks in NativeIo.c\r\n\r\n{{initNative()}}, {{initNativePosix()}} and {{initNativeWindows()}}\r\n\r\n2. Refactored the workaround for threadsafe workaround for some platforms for {{getpwduid}} calls.\r\n Instead of looking back from Native code for static variable \r\n {{workaroundNonThreadSafePasswdCalls}}, passing itself directly via native method itself as parameter. Currently its only required for POSIX, using this variable for {{initNativePosix()}} only (same as before).\r\n\r\n3. Tests added separate classes for Windows and POSIX to load static blocks separately in forked JVMs.","created":"2018-01-22T06:02:46.165+0000"},{"body":"Verified the tests manually. \r\n\r\n{{TestNativeIoInitPosix}} timesout without src change as expected.\r\n\r\n{{TestNativeIoInitWindows}} passes with/without src change as this is only future proof test.\r\n\r\nPlease review.","created":"2018-01-22T06:08:15.787+0000"},{"body":"Fixed checkstyles.","created":"2018-01-22T09:19:48.346+0000"},{"body":"updated the patch. \r\nUnified the tests into single class.","created":"2018-02-14T09:22:12.998+0000"},{"body":"[~vinayakumarb] can u rebase?","created":"2020-05-28T13:24:41.815+0000"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Logfile || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 0m 0s{color} | {color:blue}{color} | {color:blue} Docker mode activated. {color} |\r\n| {color:red}-1{color} | {color:red} patch {color} | {color:red} 0m 13s{color} | {color:red}{color} | {color:red} HADOOP-14451 does not apply to trunk. Rebase required? Wrong Branch? See https://wiki.apache.org/hadoop/HowToContribute for help. {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| JIRA Issue | HADOOP-14451 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12910549/HADOOP-14451-04.patch |\r\n| Console output | https://ci-hadoop.apache.org/job/PreCommit-HADOOP-Build/257/console |\r\n| versions | git=2.17.1 |\r\n| Powered by | Apache Yetus 0.13.0-SNAPSHOT https://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","created":"2022-02-06T20:59:00.361+0000"}],"conversations":[{"body":"* Scenario:\r\n 1. One thread calls a static method of NativeIO, which loads static block of NativeIo.\r\n 2. Second thread calls a static method of NativeIo.POSIX, which loads a static block of NativeIO.POSIX class\r\n\r\nBoth try to lock on same object inside native code gets into deadlock.","from":"reporter","subject":"Deadlock in NativeIO"},{"body":"Below are bits from stacktrace:\n\nThread1\n{code}\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.fs.RawLocalFileSystem.setPermission(RawLocalFileSystem.java:739)\n\tat org.apache.hadoop.fs.RawLocalFileSystem$LocalFSFileOutputStream.(RawLocalFileSystem.java:224)\n\tat org.apache.hadoop.fs.RawLocalFileSystem$LocalFSFileOutputStream.(RawLocalFileSystem.java:208)\n{code}\n\t\nThread2\n{code}\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.mapred.FadvisedFileRegion.transferSuccessful(FadvisedFileRegion.java:160)\n\tat org.apache.hadoop.mapred.ShuffleHandler$Shuffle$1.operationComplete(ShuffleHandler.java:1166)\n\tat org.jboss.netty.channel.DefaultChannelFuture.notifyListener(DefaultChannelFuture.java:427)\n\tat org.jboss.netty.channel.DefaultChannelFuture.notifyListeners(DefaultChannelFuture.java:413)\n{code}\n\n\t\nThe above threads looks to be blocked by below two stacks\n\t\nStack1:\n{code}\n\"New I/O worker #1\" #135 prio=5 os_prio=0 tid=0x00007f1f60817800 nid=0x697d in Object.wait() [0x00007f1f4429a000]\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.io.nativeio.NativeIO$POSIX.(NativeIO.java:184)\n\tat org.apache.hadoop.mapred.FadvisedFileRegion.transferSuccessful(FadvisedFileRegion.java:160)\n\tat org.apache.hadoop.mapred.ShuffleHandler$Shuffle$1.operationComplete(ShuffleHandler.java:1166)\n{code}\n\t\nStack2:\n{code}\n\"ContainersLauncher #16\" #365 prio=5 os_prio=0 tid=0x00007f1f49c8a800 nid=0x7cd0 in Object.wait() [0x00007f1f32891000]\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.io.nativeio.NativeIO.initNative(Native Method)\n\tat org.apache.hadoop.io.nativeio.NativeIO.(NativeIO.java:645)\n\tat org.apache.hadoop.fs.RawLocalFileSystem.setPermission(RawLocalFileSystem.java:739)\n{code}\n\t\n*Stack1* is blocked by *Stack2* as *Stack1* thread needs *NativeIO* class initialization to finish, so the problematic stack looks to be Stack 2\t\nNext in Stack 2 @ NativeIO.java:645 initNative in a native call, and it tries to initialize the native-hadoop library\n\nso in Stack 2: it try to do this in NativeIO.c via *Java_org_apache_hadoop_io_nativeio_NativeIO_initNative*\n{code}static void consts_init(JNIEnv *env) {\n jclass clazz = (*env)->FindClass(env, NATIVE_IO_POSIX_CLASS);\n #define NATIVE_IO_POSIX_CLASS \"org/apache/hadoop/io/nativeio/NativeIO$POSIX\"\n i.e create class org.apache.hadoop.io.nativeio.NativeIO$POSIX{code}\n \nbut *Stack1* is already in {{org.apache.hadoop.io.nativeio.NativeIO$POSIX.}}\nso it deadlock and all threads hang","from":"developer"},{"body":"A flavor of https://bugs.openjdk.java.net/browse/JDK-8037567 but not same, rather caused by application code","from":"developer"},{"body":"attaching affected nodemanager jstack\nFor the background of the issue, test NM seem to have all containers timeout after restart, and hence from jstack, arrived at above conclusion with deadlock in NativeIO","from":"developer"},{"body":"Some references\nhttps://docs.oracle.com/javase/8/docs/technotes/guides/jni/spec/functions.html#FindClass \nhttps://docs.oracle.com/javase/specs/jls/se8/html/jls-12.html#jls-12.4.2 ","from":"developer"},{"body":"Thank you [~ajithshetty] .Analysis makes sense to me. The same could cause dead lock.\nTried a sample program to simulate.\n{code}\n\t\tThread one=new Thread(){\n\t\t\t@Override\n\t\t\tpublic void run() {\n\t\t\t\tNativeIO.isAvailable();\n\t\t\t}\n\t\t};\t\t\n\t\tThread two=new Thread(){\n\t\t\t@Override\n\t\t\tpublic void run() {\n\t\t\t\tNativeIO.POSIX.isAvailable();\n\t\t}\n\t\t};\n\t\ttwo.start();\t\n\t\tone.start();\n\t\tone.join();\n\t\ttwo.join();\n{code}\n*stack trace*\n{code}\n\"Thread-1\" #9 prio=5 os_prio=0 tid=0x00007f1b90c5f000 nid=0x6843 in Object.wait() [0x00007f1b80ba3000]\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.io.nativeio.NativeIO.initNative(Native Method)\n\tat org.apache.hadoop.io.nativeio.NativeIO.(NativeIO.java:644)\n\tat org.apache.hadoop.io.nativeio.TestNative$1.run(TestNative.java:8)\n\n\"Thread-2\" #10 prio=5 os_prio=0 tid=0x00007f1b90c58000 nid=0x6842 in Object.wait() [0x00007f1b80ca4000]\n java.lang.Thread.State: RUNNABLE\n\tat org.apache.hadoop.io.nativeio.NativeIO$POSIX.(NativeIO.java:184)\n\tat org.apache.hadoop.io.nativeio.TestNative$2.run(TestNative.java:15)\n{code}\n\nAffected versions are *2.8.0 and hadoop3*\n\n\ncc/ [~cmccabe_impala_fa3f] ","from":"developer"},{"body":"Hi [~vinayrpet] , [~rakesh_r] , [~martinw] any thoughtss??","from":"developer"},{"body":"I am not actively working on this, so my thoughts may be wide of the mark. Does this issue occur with\n\n{code}\nif (NativeIO.isAvailable()) {\n NativeIO.POSIX.getCacheManipulator().posixFadviseIfPossible(identifier, fd, getPosition(), getCount(), POSIX_FADV_DONTNEED);\n}\n{code}\n\n\n\n\n","from":"developer"},{"body":"[~martinw]\nWe tried the same the approach could solve the problem. But in all {{NativeIO.POSIX}} usage location we have to check.\n[~ajithshetty] Any other solution in your mind?","from":"developer"},{"body":"Only other suggestion I can think of is to break down {{initNative()}} into smaller functions, e.g. {{initNativePosix()}}, {{initNativeStat()}}... Each of which is initialised by a static block within its inner class. ","from":"developer"},{"body":"Attached a simple approach of avoiding deadlock.","from":"developer"},{"body":"i think people here are more interested in uploading a patch rather than discussing with the submitter if analysis and approach is right. If i have added well amount of analysis, i would probably have a solution too. Anyways, feel free to assign to self if that's the case","from":"developer"},{"body":"Looks like the 'simple' approach is not working for this after all for this Native IO.\n [~ajithshetty], Please continue your analysis.","from":"developer"},{"body":"Uploaded the patch, as per [~martinw]'s suggestion.\r\n\r\n1. Split into separate code blocks in NativeIo.c\r\n\r\n{{initNative()}}, {{initNativePosix()}} and {{initNativeWindows()}}\r\n\r\n2. Refactored the workaround for threadsafe workaround for some platforms for {{getpwduid}} calls.\r\n Instead of looking back from Native code for static variable \r\n {{workaroundNonThreadSafePasswdCalls}}, passing itself directly via native method itself as parameter. Currently its only required for POSIX, using this variable for {{initNativePosix()}} only (same as before).\r\n\r\n3. Tests added separate classes for Windows and POSIX to load static blocks separately in forked JVMs.","from":"developer"},{"body":"Verified the tests manually. \r\n\r\n{{TestNativeIoInitPosix}} timesout without src change as expected.\r\n\r\n{{TestNativeIoInitWindows}} passes with/without src change as this is only future proof test.\r\n\r\nPlease review.","from":"developer"},{"body":"Fixed checkstyles.","from":"developer"},{"body":"updated the patch. \r\nUnified the tests into single class.","from":"developer"},{"body":"[~vinayakumarb] can u rebase?","from":"developer"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Logfile || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 0m 0s{color} | {color:blue}{color} | {color:blue} Docker mode activated. {color} |\r\n| {color:red}-1{color} | {color:red} patch {color} | {color:red} 0m 13s{color} | {color:red}{color} | {color:red} HADOOP-14451 does not apply to trunk. Rebase required? Wrong Branch? See https://wiki.apache.org/hadoop/HowToContribute for help. {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| JIRA Issue | HADOOP-14451 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12910549/HADOOP-14451-04.patch |\r\n| Console output | https://ci-hadoop.apache.org/job/PreCommit-HADOOP-Build/257/console |\r\n| versions | git=2.17.1 |\r\n| Powered by | Apache Yetus 0.13.0-SNAPSHOT https://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","from":"developer"}],"created":"2017-05-24T03:40:22.000+0000","description":"* Scenario:\r\n 1. One thread calls a static method of NativeIO, which loads static block of NativeIo.\r\n 2. Second thread calls a static method of NativeIo.POSIX, which loads a static block of NativeIO.POSIX class\r\n\r\nBoth try to lock on same object inside native code gets into deadlock.","issue_id":"13074392","key":"HADOOP-14451","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-06-24T04:58:41.000+0000","role":"fixed_distractor","summary":"Deadlock in NativeIO"} {"case_id":"12371190","cluster":"DISTRACTOR-HADOOP-1475","comments":[{"body":"This patch clears the cache before the task tracker reinitializes itself. This prevents the file cache from thinking it has files that it in fact doesn't have, because the backing files have been deleted in the reinitialization.","created":"2007-06-18T05:57:26.176+0000"},{"body":"I just committed this. Thanks, Owen!\n\nhttp://svn.apache.org/viewvc?view=rev&rev=549200","created":"2007-06-20T19:25:23.080+0000"}],"conversations":[{"body":"All our jobs on a 600 node cluster fail. Symptom is that the local filecache disappears.\n\nIt might have to do with the fact that lost task trackers get re-initialized when they send a heartbeat again, and purge the local directory completely without updating the filecache.\n\nSide issue is;\nwhy do we get so many lost tasktrackers which then resume the heartbeat (a kind of 'bogus' lost tasktracker)?. We lost tasktrackers:\n13 in the 1st hour of the job\n18 in the 2nd hour\n33 in the 3rd hour\nThen the job failed.\n\nE.g. all the tasktrackers lost in the first 2 hours of the job got logged sometime later with a 'Status from unknown Tracker' in the jobtracker log and got reinitialized.\n\nI attach some jobracker log messages showing how the heartbeat of the lost tasktrackers come in late, sometimes less than 1 minute late, sometimes up to 16 minutes. What could be the reason? Do the heartbeats get lost? \n\n\n\n2007-06-07 13:09:08,518 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_070\n2007-06-07 13:09:48,919 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_070\n\n2007-06-07 13:39:08,740 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_075\n2007-06-07 13:41:50,810 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_075\n\n2007-06-07 14:32:29,093 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_082\n2007-06-07 14:35:34,217 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_082\n\n2007-06-07 14:15:48,856 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_085\n2007-06-07 14:20:21,337 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_085\n\n2007-06-07 15:25:49,524 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_098\n2007-06-07 15:33:56,732 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_098\n\n2007-06-07 14:49:09,203 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_106\n2007-06-07 14:54:25,538 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_106\n\n2007-06-07 15:02:29,337 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_108\n2007-06-07 15:02:57,558 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_108\n\n2007-06-07 14:19:09,022 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_112\n2007-06-07 14:19:15,273 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_112\n\n2007-06-07 14:19:08,881 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_114\n2007-06-07 14:30:03,354 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_114\n\n2007-06-07 15:42:29,579 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_116\n2007-06-07 15:43:06,422 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_116\n\n2007-06-07 14:55:49,280 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_117\n2007-06-07 14:56:38,452 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_117\n\n2007-06-07 15:15:49,461 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_120\n2007-06-07 15:31:37,028 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_120\n\n2007-06-07 15:09:09,435 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_174\n2007-06-07 15:18:31,254 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_174\n\n\n","from":"reporter","subject":"local filecache disappears"},{"body":"This patch clears the cache before the task tracker reinitializes itself. This prevents the file cache from thinking it has files that it in fact doesn't have, because the backing files have been deleted in the reinitialization.","from":"developer"},{"body":"I just committed this. Thanks, Owen!\n\nhttp://svn.apache.org/viewvc?view=rev&rev=549200","from":"developer"}],"created":"2007-06-07T23:39:10.000+0000","description":"All our jobs on a 600 node cluster fail. Symptom is that the local filecache disappears.\n\nIt might have to do with the fact that lost task trackers get re-initialized when they send a heartbeat again, and purge the local directory completely without updating the filecache.\n\nSide issue is;\nwhy do we get so many lost tasktrackers which then resume the heartbeat (a kind of 'bogus' lost tasktracker)?. We lost tasktrackers:\n13 in the 1st hour of the job\n18 in the 2nd hour\n33 in the 3rd hour\nThen the job failed.\n\nE.g. all the tasktrackers lost in the first 2 hours of the job got logged sometime later with a 'Status from unknown Tracker' in the jobtracker log and got reinitialized.\n\nI attach some jobracker log messages showing how the heartbeat of the lost tasktrackers come in late, sometimes less than 1 minute late, sometimes up to 16 minutes. What could be the reason? Do the heartbeats get lost? \n\n\n\n2007-06-07 13:09:08,518 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_070\n2007-06-07 13:09:48,919 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_070\n\n2007-06-07 13:39:08,740 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_075\n2007-06-07 13:41:50,810 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_075\n\n2007-06-07 14:32:29,093 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_082\n2007-06-07 14:35:34,217 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_082\n\n2007-06-07 14:15:48,856 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_085\n2007-06-07 14:20:21,337 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_085\n\n2007-06-07 15:25:49,524 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_098\n2007-06-07 15:33:56,732 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_098\n\n2007-06-07 14:49:09,203 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_106\n2007-06-07 14:54:25,538 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_106\n\n2007-06-07 15:02:29,337 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_108\n2007-06-07 15:02:57,558 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_108\n\n2007-06-07 14:19:09,022 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_112\n2007-06-07 14:19:15,273 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_112\n\n2007-06-07 14:19:08,881 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_114\n2007-06-07 14:30:03,354 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_114\n\n2007-06-07 15:42:29,579 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_116\n2007-06-07 15:43:06,422 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_116\n\n2007-06-07 14:55:49,280 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_117\n2007-06-07 14:56:38,452 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_117\n\n2007-06-07 15:15:49,461 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_120\n2007-06-07 15:31:37,028 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_120\n\n2007-06-07 15:09:09,435 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker tracker_174\n2007-06-07 15:18:31,254 WARN org.apache.hadoop.mapred.JobTracker: Status_from_unknown_Tracker : tracker_174\n\n\n","issue_id":"12371190","key":"HADOOP-1475","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2007-06-20T19:25:23.000+0000","role":"fixed_distractor","summary":"local filecache disappears"} {"case_id":"13125761","cluster":"DISTRACTOR-HADOOP-15129","comments":[{"body":"Re-resolve unresolved addresses for up to ipc client max retries.","created":"2017-12-19T00:40:23.493+0000"},{"body":"the IPC is very sensitive code, as DN/NN comms, so this is going to need review by people who care about that, maybe [~arpitagarwal].\r\n\r\nto me, production code looks fine, but we need reviewers from HDFS too.\r\n\r\nTest wise, \r\n* use LambdaTestUtils.intercept() & java-8 lambda expressions for those tests which throw exceptions; if your lambda returns any object, toString() will be caused on it & the text used in the exception...this can be useful\r\n* and as Client is autocloseable, you can use try-with-resources to manage its lifecycle\r\n\r\nTest failures unrelated & covered elsewhere","created":"2017-12-19T14:27:15.765+0000"},{"body":"Isn't this dupe of HDFS-8068 ?\r\nAlso it is related to HADOOP-12125.","created":"2017-12-19T14:39:34.873+0000"},{"body":"Thanks for the heads up Steve.\r\n\r\n[~Karthik Palaniappan], do you have the callstack to go with this error message?\r\n{code}\r\n2017-12-15 00:13:40,712 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n{code}\r\nIt's not logged by default but perhaps you captured it while debugging. Also I added you as a contributor.\r\n","created":"2017-12-19T18:48:34.415+0000"},{"body":"[~arpitagarwal]: Here are the notes from our investigation.\r\n\r\nstartDataNode() calls refreshNamenodes(): https://github.com/apache/hadoop/blob/trunk/hadoop-hdfs-project/hadoop-hdfs/src/main/java/org/apache/hadoop/hdfs/server/datanode/DataNode.java#L1425\r\n\r\nrefreshNamenodes has unresolved InetSocketAddress(es): https://github.com/apache/hadoop/blob/trunk/hadoop-hdfs-project/hadoop-hdfs/src/main/java/org/apache/hadoop/hdfs/server/datanode/BlockPoolManager.java#L152, and the unresolved address(es) get passed to a BPOfferService. Note that a BPOfferService is only created once for a NameServiceId.\r\n\r\nBPOS caches that ISA: https://github.com/apache/hadoop/blob/trunk/hadoop-hdfs-project/hadoop-hdfs/src/main/java/org/apache/hadoop/hdfs/server/datanode/BPOfferService.java#L134\r\n\r\nAnd creates a BPServiceActor with that ISA: https://github.com/apache/hadoop/blob/trunk/hadoop-hdfs-project/hadoop-hdfs/src/main/java/org/apache/hadoop/hdfs/server/datanode/BPOfferService.java#L138\r\n\r\nWhen the BPServiceActor tries to get NN info, it hits that exception (and logs “Problem connecting to server”): https://github.com/apache/hadoop/blob/trunk/hadoop-hdfs-project/hadoop-hdfs/src/main/java/org/apache/hadoop/hdfs/server/datanode/BPServiceActor.java#L235\r\n\r\nHere is the stacktrace of the exception:\r\n\r\n{code:java}\r\njava.net.UnknownHostException: Invalid host name: local host is: (unknown); destination host is: \"non-existent\":8020; java.net.UnknownHostException; For more details see: http://wiki.apache.org/hadoop/UnknownHost\r\n at org.apache.hadoop.hdfs.server.datanode.BPServiceActor.retrieveNamespaceInfo(BPServiceActor.java:221)\r\n at org.apache.hadoop.hdfs.server.datanode.BPServiceActor.connectToNNAndHandshake(BPServiceActor.java:263)\r\n at org.apache.hadoop.hdfs.server.datanode.BPServiceActor.run(BPServiceActor.java:752)\r\n at java.lang.Thread.run(Thread.java:748)\r\nCaused by: java.net.UnknownHostException: Invalid host name: local host is: (unknown); destination host is: \"non-existent\":8020; java.net.UnknownHostException; For more details see: http://wiki.apache.org/hadoop/UnknownHost\r\n at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\r\n at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\r\n at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\r\n at java.lang.reflect.Constructor.newInstance(Constructor.java:423)\r\n at org.apache.hadoop.net.NetUtils.wrapWithMessage(NetUtils.java:801)\r\n at org.apache.hadoop.net.NetUtils.wrapException(NetUtils.java:744)\r\n at org.apache.hadoop.ipc.Client$Connection.(Client.java:446)\r\n at org.apache.hadoop.ipc.Client.getConnection(Client.java:1524)\r\n at org.apache.hadoop.ipc.Client.call(Client.java:1375)\r\n at org.apache.hadoop.ipc.Client.call(Client.java:1339)\r\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:227)\r\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:116)\r\n at com.sun.proxy.$Proxy15.versionRequest(Unknown Source)\r\n at org.apache.hadoop.hdfs.protocolPB.DatanodeProtocolClientSideTranslatorPB.versionRequest(DatanodeProtocolClientSideTranslatorPB.java:274)\r\n at org.apache.hadoop.hdfs.server.datanode.BPServiceActor.retrieveNamespaceInfo(BPServiceActor.java:216)\r\n ... 3 more\r\n{code}\r\n\r\nBPServiceActor uses a while(true) loop that runs every 5 seconds. I don’t think it ever stops (which is a little weird).\r\n","created":"2017-12-19T20:27:41.881+0000"},{"body":"[~shahrs87]: It's related to both -- I'll add links to them.\r\n\r\n1) This is not related to proxies like 8068 and 12125\r\n\r\n2) This is a different proposed fix than 8068 -- instead of adding retries inside of HDFS, it adds retries down at the IPC layer. This is more in line with what [~jlowe] suggested in 12125.","created":"2017-12-19T20:31:49.784+0000"},{"body":"[~stevel@apache.org] I was working on incorporating your feedback, but it looks like Hadoop 2.8 still supports Java 7, so I can't use lambdas or try-with-resources.","created":"2018-01-08T21:17:59.150+0000"},{"body":"afraid not with try-with-resource. you can use lambdas with anonymous classes, as it lines you up for moving to Java 8 with ease in future; the IDE can take a class like that and switch to a Lambda at a button press.","created":"2018-01-10T11:47:25.024+0000"},{"body":"Fair enough, patch 002 uses a Java 7-style anonymous class with LambdaTestUtils.","created":"2018-01-10T19:08:59.116+0000"},{"body":"I haven't looked at the test cases yet but the change looks fine to me. Will review the new tests.\r\n\r\n[~kihwal], do you have any thoughts since you added the original re-resolution logic?","created":"2018-01-29T20:13:42.438+0000"},{"body":"{quote}local host is: (unknown); {quote}\r\ncan we pass \"localhost\" for {{NetUtils.wrapException}} instead of null. Above message in logs is little misleading.","created":"2018-01-29T21:17:46.880+0000"},{"body":"The test case {{testIpcFlakyHostResolution}} also lgtm. The other new test {{testIpcHostResolutionTimeout}} looks unrelated to this change. Do we need it [~Karthik Palaniappan]?\r\n\r\nbq. can we pass \"localhost\" for NetUtils.wrapException instead of null. Above message in logs is little misleading.\r\nGood point. The local IP address for the connection is indeterminate until the target hostname is resolved e.g. multi-homed setups. So I am okay with leaving this as null for now. [~ajayydv], let me know if that sounds reasonable.","created":"2018-02-12T19:53:00.460+0000"},{"body":"{quote} The local IP address for the connection is indeterminate until the target hostname is resolved e.g. multi-homed setups. So I am okay with leaving this as null for now.{quote}\r\n[~arpitagarwal], I see your point. I was thinking that localhost \"UNKOWN\" may mislead someone in wrong diagnosis about the problem. Wondering if we can we improve the messaging. But this can be discussed separately and doesn't need to be part of this jira.","created":"2018-02-12T21:54:00.707+0000"},{"body":"[~arpitagarwal] yes both tests are necessary. The HostResolutionTimeout test proves that if the host is always unresolvable, the code still does throw an exception. TheFlakyHostResolution test shows that if the host is only temporarily resolvable, ipc client can recover. Other tests of course prove the final case that if the host is immediately resolvable it's fine.","created":"2018-02-16T02:47:38.162+0000"},{"body":"Could always add a specific failure call and wiki link here","created":"2018-02-16T11:49:19.974+0000"},{"body":"I forgot about this for a while – is there anything blocking merging the patch?","created":"2018-04-19T01:19:08.975+0000"},{"body":"Hi Karthik! Thanks for your contribution. Could you please rebase the patch to the latest trunk? I usually apply patches using\r\n{code:java}\r\n$ git apply {code}\r\nA few suggestions:\r\n # Could you please use short descriptions in JIRA? I was told a long time ago. :)\r\n # When using JIRA numbers, could you please write HDFS-8068 (instead of just 8068) because issues often cut across several different projects, and this way JIRA creates nice links for viewers to click on?\r\n\r\nPatches are usually committed to trunk *first* and then a (possibly) different version of the patch may be committed to earlier branches like branch-2. So technically you could have used neat Lambdas in the trunk patch. ;) Its a nit though.\r\n\r\nI'm trying to find the wikipage that tried to explain certain errors. I'm afraid I rarely found them useful (its probably because we didn't really expand on those wiki pages ever), so I'm fine with a more helpful error in the logs.\r\n\r\nCould you please also comment on whether you have been running with this patch in production for any amount of time and seen / not seen any issues with it?\r\n\r\nI concur that this is extremely important code, so it behooves us to tread very carefully. ","created":"2018-12-12T11:09:06.328+0000"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Logfile || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 0m 0s{color} | {color:blue}{color} | {color:blue} Docker mode activated. {color} |\r\n| {color:red}-1{color} | {color:red} patch {color} | {color:red} 0m 12s{color} | {color:red}{color} | {color:red} HADOOP-15129 does not apply to trunk. Rebase required? Wrong Branch? See https://wiki.apache.org/hadoop/HowToContribute for help. {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| JIRA Issue | HADOOP-15129 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12905521/HADOOP-15129.002.patch |\r\n| Console output | https://ci-hadoop.apache.org/job/PreCommit-HADOOP-Build/229/console |\r\n| versions | git=2.17.1 |\r\n| Powered by | Apache Yetus 0.13.0-SNAPSHOT https://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","created":"2021-08-28T03:58:47.405+0000"},{"body":"This remains a problem for cloud infrastructure deployments, so I'd like to pick it up and see if we can get it completed. I've sent a pull request with the following changes compared to the prior revision:\r\n\r\n* Remove older code for a throw of {{UnknownHostException}}. This lies outside the retry loop, so even though the earlier patch did the right thing by placing the throw inside the retry loop, this remaining code perpetuated the problem of an infinite unresolved host.\r\n* Make minor formatting changes in the test to resolve Checkstyle issues flagged in the last Yetus run.\r\n\r\nAdditionally, I've confirmed testing of the patch in moderate-sized (200-node) Dataproc cluster deployments.\r\n\r\n[~stevel@apache.org], [~arp], [~raviprak], [~ajayydv], and [~shahrs87], can we please work on getting this reviewed and committed? I'm interested in merging this down to branch-3.3, branch-3.2, branch-2.10 and branch-2.9. The patch as-is won't apply cleanly to 2.x. If you approve, then I'll prepare separate pull requests for those branches.\r\n\r\nAlso, BTW, hello everyone. :-)","created":"2021-08-28T04:06:28.028+0000"},{"body":"Let's please also maintain credit for Karthik. (I submitted my patch with a Co-authored-by tag.)","created":"2021-08-28T04:08:02.117+0000"},{"body":"I forgot to mention one other change I made from the prior patch. I incorporated {{NetUtils.getHostname()}} into the exception message, but now I see Arpit raised a point in an earlier comment about multi-homed configurations. If it's preferred, I'm happy to switch that back to {{null}} and leave it undetermined in the exception message.","created":"2021-08-28T16:44:16.891+0000"},{"body":"[~ywskycn], thank you for the +1. I'm planning to commit this but will give it a little more time in case the earlier reviewers still want to add feedback.","created":"2021-09-01T16:19:01.776+0000"},{"body":"I have committed this to trunk, branch-3.3 and branch-3.2. I didn't end up merging down to the 2.x line like I said I would, because I retested on 2.x, and the bug isn't present there.\r\n\r\n[~Karthik Palaniappan], thank you for providing the original patch. Thank you to all of the reviewers and [~ywskycn] for the final review.","created":"2021-09-03T19:47:21.530+0000"}],"conversations":[{"body":"On startup, the Datanode creates an InetSocketAddress to register with each namenode. Though there are retries on connection failure throughout the stack, the same InetSocketAddress is reused.\r\n\r\nInetSocketAddress is an interesting class, because it resolves DNS names to IP addresses on construction, and it is never refreshed. Hadoop re-creates an InetSocketAddress in some cases just in case the remote IP has changed for a particular DNS name: https://issues.apache.org/jira/browse/HADOOP-7472.\r\n\r\nAnyway, on startup, you cna see the Datanode log: \"Namenode...remains unresolved\" -- referring to the fact that DNS lookup failed.\r\n\r\n{code:java}\r\n2017-11-02 16:01:55,115 INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Refresh request received for nameservices: null\r\n2017-11-02 16:01:55,153 WARN org.apache.hadoop.hdfs.DFSUtilClient: Namenode for null remains unresolved for ID null. Check your hdfs-site.xml file to ensure namenodes are configured properly.\r\n2017-11-02 16:01:55,156 INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Starting BPOfferServices for nameservices: \r\n2017-11-02 16:01:55,169 INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Block pool (Datanode Uuid unassigned) service to cluster-32f5-m:8020 starting to offer service\r\n{code}\r\n\r\nThe Datanode then proceeds to use this unresolved address, as it may work if the DN is configured to use a proxy. Since I'm not using a proxy, it forever prints out this message:\r\n\r\n\r\n{code:java}\r\n2017-12-15 00:13:40,712 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n2017-12-15 00:13:45,712 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n2017-12-15 00:13:50,712 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n2017-12-15 00:13:55,713 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n2017-12-15 00:14:00,713 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n\r\n{code}\r\n\r\nUnfortunately, the log doesn't contain the exception that triggered it, but the culprit is actually in IPC Client: https://github.com/apache/hadoop/blob/trunk/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/ipc/Client.java#L444.\r\n\r\nThis line was introduced in https://issues.apache.org/jira/browse/HADOOP-487 to give a clear error message when somebody mispells an address.\r\n\r\nHowever, the fix in HADOOP-7472 doesn't apply here, because that code happens in Client#getConnection after the Connection is constructed.\r\n\r\nMy proposed fix (will attach a patch) is to move this exception out of the constructor and into a place that will trigger HADOOP-7472's logic to re-resolve addresses. If the DNS failure was temporary, this will allow the connection to succeed. If not, the connection will fail after ipc client retries (default 10 seconds worth of retries).\r\n\r\nI want to fix this in ipc client rather than just in Datanode startup, as this fixes temporary DNS issues for all of Hadoop.","from":"reporter","subject":"Datanode caches namenode DNS lookup failure and cannot startup"},{"body":"Re-resolve unresolved addresses for up to ipc client max retries.","from":"developer"},{"body":"the IPC is very sensitive code, as DN/NN comms, so this is going to need review by people who care about that, maybe [~arpitagarwal].\r\n\r\nto me, production code looks fine, but we need reviewers from HDFS too.\r\n\r\nTest wise, \r\n* use LambdaTestUtils.intercept() & java-8 lambda expressions for those tests which throw exceptions; if your lambda returns any object, toString() will be caused on it & the text used in the exception...this can be useful\r\n* and as Client is autocloseable, you can use try-with-resources to manage its lifecycle\r\n\r\nTest failures unrelated & covered elsewhere","from":"developer"},{"body":"Isn't this dupe of HDFS-8068 ?\r\nAlso it is related to HADOOP-12125.","from":"developer"},{"body":"Thanks for the heads up Steve.\r\n\r\n[~Karthik Palaniappan], do you have the callstack to go with this error message?\r\n{code}\r\n2017-12-15 00:13:40,712 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n{code}\r\nIt's not logged by default but perhaps you captured it while debugging. Also I added you as a contributor.\r\n","from":"developer"},{"body":"[~arpitagarwal]: Here are the notes from our investigation.\r\n\r\nstartDataNode() calls refreshNamenodes(): https://github.com/apache/hadoop/blob/trunk/hadoop-hdfs-project/hadoop-hdfs/src/main/java/org/apache/hadoop/hdfs/server/datanode/DataNode.java#L1425\r\n\r\nrefreshNamenodes has unresolved InetSocketAddress(es): https://github.com/apache/hadoop/blob/trunk/hadoop-hdfs-project/hadoop-hdfs/src/main/java/org/apache/hadoop/hdfs/server/datanode/BlockPoolManager.java#L152, and the unresolved address(es) get passed to a BPOfferService. Note that a BPOfferService is only created once for a NameServiceId.\r\n\r\nBPOS caches that ISA: https://github.com/apache/hadoop/blob/trunk/hadoop-hdfs-project/hadoop-hdfs/src/main/java/org/apache/hadoop/hdfs/server/datanode/BPOfferService.java#L134\r\n\r\nAnd creates a BPServiceActor with that ISA: https://github.com/apache/hadoop/blob/trunk/hadoop-hdfs-project/hadoop-hdfs/src/main/java/org/apache/hadoop/hdfs/server/datanode/BPOfferService.java#L138\r\n\r\nWhen the BPServiceActor tries to get NN info, it hits that exception (and logs “Problem connecting to server”): https://github.com/apache/hadoop/blob/trunk/hadoop-hdfs-project/hadoop-hdfs/src/main/java/org/apache/hadoop/hdfs/server/datanode/BPServiceActor.java#L235\r\n\r\nHere is the stacktrace of the exception:\r\n\r\n{code:java}\r\njava.net.UnknownHostException: Invalid host name: local host is: (unknown); destination host is: \"non-existent\":8020; java.net.UnknownHostException; For more details see: http://wiki.apache.org/hadoop/UnknownHost\r\n at org.apache.hadoop.hdfs.server.datanode.BPServiceActor.retrieveNamespaceInfo(BPServiceActor.java:221)\r\n at org.apache.hadoop.hdfs.server.datanode.BPServiceActor.connectToNNAndHandshake(BPServiceActor.java:263)\r\n at org.apache.hadoop.hdfs.server.datanode.BPServiceActor.run(BPServiceActor.java:752)\r\n at java.lang.Thread.run(Thread.java:748)\r\nCaused by: java.net.UnknownHostException: Invalid host name: local host is: (unknown); destination host is: \"non-existent\":8020; java.net.UnknownHostException; For more details see: http://wiki.apache.org/hadoop/UnknownHost\r\n at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\r\n at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\r\n at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\r\n at java.lang.reflect.Constructor.newInstance(Constructor.java:423)\r\n at org.apache.hadoop.net.NetUtils.wrapWithMessage(NetUtils.java:801)\r\n at org.apache.hadoop.net.NetUtils.wrapException(NetUtils.java:744)\r\n at org.apache.hadoop.ipc.Client$Connection.(Client.java:446)\r\n at org.apache.hadoop.ipc.Client.getConnection(Client.java:1524)\r\n at org.apache.hadoop.ipc.Client.call(Client.java:1375)\r\n at org.apache.hadoop.ipc.Client.call(Client.java:1339)\r\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:227)\r\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:116)\r\n at com.sun.proxy.$Proxy15.versionRequest(Unknown Source)\r\n at org.apache.hadoop.hdfs.protocolPB.DatanodeProtocolClientSideTranslatorPB.versionRequest(DatanodeProtocolClientSideTranslatorPB.java:274)\r\n at org.apache.hadoop.hdfs.server.datanode.BPServiceActor.retrieveNamespaceInfo(BPServiceActor.java:216)\r\n ... 3 more\r\n{code}\r\n\r\nBPServiceActor uses a while(true) loop that runs every 5 seconds. I don’t think it ever stops (which is a little weird).\r\n","from":"developer"},{"body":"[~shahrs87]: It's related to both -- I'll add links to them.\r\n\r\n1) This is not related to proxies like 8068 and 12125\r\n\r\n2) This is a different proposed fix than 8068 -- instead of adding retries inside of HDFS, it adds retries down at the IPC layer. This is more in line with what [~jlowe] suggested in 12125.","from":"developer"},{"body":"[~stevel@apache.org] I was working on incorporating your feedback, but it looks like Hadoop 2.8 still supports Java 7, so I can't use lambdas or try-with-resources.","from":"developer"},{"body":"afraid not with try-with-resource. you can use lambdas with anonymous classes, as it lines you up for moving to Java 8 with ease in future; the IDE can take a class like that and switch to a Lambda at a button press.","from":"developer"},{"body":"Fair enough, patch 002 uses a Java 7-style anonymous class with LambdaTestUtils.","from":"developer"},{"body":"I haven't looked at the test cases yet but the change looks fine to me. Will review the new tests.\r\n\r\n[~kihwal], do you have any thoughts since you added the original re-resolution logic?","from":"developer"},{"body":"{quote}local host is: (unknown); {quote}\r\ncan we pass \"localhost\" for {{NetUtils.wrapException}} instead of null. Above message in logs is little misleading.","from":"developer"},{"body":"The test case {{testIpcFlakyHostResolution}} also lgtm. The other new test {{testIpcHostResolutionTimeout}} looks unrelated to this change. Do we need it [~Karthik Palaniappan]?\r\n\r\nbq. can we pass \"localhost\" for NetUtils.wrapException instead of null. Above message in logs is little misleading.\r\nGood point. The local IP address for the connection is indeterminate until the target hostname is resolved e.g. multi-homed setups. So I am okay with leaving this as null for now. [~ajayydv], let me know if that sounds reasonable.","from":"developer"},{"body":"{quote} The local IP address for the connection is indeterminate until the target hostname is resolved e.g. multi-homed setups. So I am okay with leaving this as null for now.{quote}\r\n[~arpitagarwal], I see your point. I was thinking that localhost \"UNKOWN\" may mislead someone in wrong diagnosis about the problem. Wondering if we can we improve the messaging. But this can be discussed separately and doesn't need to be part of this jira.","from":"developer"},{"body":"[~arpitagarwal] yes both tests are necessary. The HostResolutionTimeout test proves that if the host is always unresolvable, the code still does throw an exception. TheFlakyHostResolution test shows that if the host is only temporarily resolvable, ipc client can recover. Other tests of course prove the final case that if the host is immediately resolvable it's fine.","from":"developer"},{"body":"Could always add a specific failure call and wiki link here","from":"developer"},{"body":"I forgot about this for a while – is there anything blocking merging the patch?","from":"developer"},{"body":"Hi Karthik! Thanks for your contribution. Could you please rebase the patch to the latest trunk? I usually apply patches using\r\n{code:java}\r\n$ git apply {code}\r\nA few suggestions:\r\n # Could you please use short descriptions in JIRA? I was told a long time ago. :)\r\n # When using JIRA numbers, could you please write HDFS-8068 (instead of just 8068) because issues often cut across several different projects, and this way JIRA creates nice links for viewers to click on?\r\n\r\nPatches are usually committed to trunk *first* and then a (possibly) different version of the patch may be committed to earlier branches like branch-2. So technically you could have used neat Lambdas in the trunk patch. ;) Its a nit though.\r\n\r\nI'm trying to find the wikipage that tried to explain certain errors. I'm afraid I rarely found them useful (its probably because we didn't really expand on those wiki pages ever), so I'm fine with a more helpful error in the logs.\r\n\r\nCould you please also comment on whether you have been running with this patch in production for any amount of time and seen / not seen any issues with it?\r\n\r\nI concur that this is extremely important code, so it behooves us to tread very carefully. ","from":"developer"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Logfile || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 0m 0s{color} | {color:blue}{color} | {color:blue} Docker mode activated. {color} |\r\n| {color:red}-1{color} | {color:red} patch {color} | {color:red} 0m 12s{color} | {color:red}{color} | {color:red} HADOOP-15129 does not apply to trunk. Rebase required? Wrong Branch? See https://wiki.apache.org/hadoop/HowToContribute for help. {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| JIRA Issue | HADOOP-15129 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12905521/HADOOP-15129.002.patch |\r\n| Console output | https://ci-hadoop.apache.org/job/PreCommit-HADOOP-Build/229/console |\r\n| versions | git=2.17.1 |\r\n| Powered by | Apache Yetus 0.13.0-SNAPSHOT https://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","from":"developer"},{"body":"This remains a problem for cloud infrastructure deployments, so I'd like to pick it up and see if we can get it completed. I've sent a pull request with the following changes compared to the prior revision:\r\n\r\n* Remove older code for a throw of {{UnknownHostException}}. This lies outside the retry loop, so even though the earlier patch did the right thing by placing the throw inside the retry loop, this remaining code perpetuated the problem of an infinite unresolved host.\r\n* Make minor formatting changes in the test to resolve Checkstyle issues flagged in the last Yetus run.\r\n\r\nAdditionally, I've confirmed testing of the patch in moderate-sized (200-node) Dataproc cluster deployments.\r\n\r\n[~stevel@apache.org], [~arp], [~raviprak], [~ajayydv], and [~shahrs87], can we please work on getting this reviewed and committed? I'm interested in merging this down to branch-3.3, branch-3.2, branch-2.10 and branch-2.9. The patch as-is won't apply cleanly to 2.x. If you approve, then I'll prepare separate pull requests for those branches.\r\n\r\nAlso, BTW, hello everyone. :-)","from":"developer"},{"body":"Let's please also maintain credit for Karthik. (I submitted my patch with a Co-authored-by tag.)","from":"developer"},{"body":"I forgot to mention one other change I made from the prior patch. I incorporated {{NetUtils.getHostname()}} into the exception message, but now I see Arpit raised a point in an earlier comment about multi-homed configurations. If it's preferred, I'm happy to switch that back to {{null}} and leave it undetermined in the exception message.","from":"developer"},{"body":"[~ywskycn], thank you for the +1. I'm planning to commit this but will give it a little more time in case the earlier reviewers still want to add feedback.","from":"developer"},{"body":"I have committed this to trunk, branch-3.3 and branch-3.2. I didn't end up merging down to the 2.x line like I said I would, because I retested on 2.x, and the bug isn't present there.\r\n\r\n[~Karthik Palaniappan], thank you for providing the original patch. Thank you to all of the reviewers and [~ywskycn] for the final review.","from":"developer"}],"created":"2017-12-19T00:34:56.000+0000","description":"On startup, the Datanode creates an InetSocketAddress to register with each namenode. Though there are retries on connection failure throughout the stack, the same InetSocketAddress is reused.\r\n\r\nInetSocketAddress is an interesting class, because it resolves DNS names to IP addresses on construction, and it is never refreshed. Hadoop re-creates an InetSocketAddress in some cases just in case the remote IP has changed for a particular DNS name: https://issues.apache.org/jira/browse/HADOOP-7472.\r\n\r\nAnyway, on startup, you cna see the Datanode log: \"Namenode...remains unresolved\" -- referring to the fact that DNS lookup failed.\r\n\r\n{code:java}\r\n2017-11-02 16:01:55,115 INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Refresh request received for nameservices: null\r\n2017-11-02 16:01:55,153 WARN org.apache.hadoop.hdfs.DFSUtilClient: Namenode for null remains unresolved for ID null. Check your hdfs-site.xml file to ensure namenodes are configured properly.\r\n2017-11-02 16:01:55,156 INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Starting BPOfferServices for nameservices: \r\n2017-11-02 16:01:55,169 INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Block pool (Datanode Uuid unassigned) service to cluster-32f5-m:8020 starting to offer service\r\n{code}\r\n\r\nThe Datanode then proceeds to use this unresolved address, as it may work if the DN is configured to use a proxy. Since I'm not using a proxy, it forever prints out this message:\r\n\r\n\r\n{code:java}\r\n2017-12-15 00:13:40,712 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n2017-12-15 00:13:45,712 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n2017-12-15 00:13:50,712 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n2017-12-15 00:13:55,713 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n2017-12-15 00:14:00,713 WARN org.apache.hadoop.hdfs.server.datanode.DataNode: Problem connecting to server: cluster-32f5-m:8020\r\n\r\n{code}\r\n\r\nUnfortunately, the log doesn't contain the exception that triggered it, but the culprit is actually in IPC Client: https://github.com/apache/hadoop/blob/trunk/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/ipc/Client.java#L444.\r\n\r\nThis line was introduced in https://issues.apache.org/jira/browse/HADOOP-487 to give a clear error message when somebody mispells an address.\r\n\r\nHowever, the fix in HADOOP-7472 doesn't apply here, because that code happens in Client#getConnection after the Connection is constructed.\r\n\r\nMy proposed fix (will attach a patch) is to move this exception out of the constructor and into a place that will trigger HADOOP-7472's logic to re-resolve addresses. If the DNS failure was temporary, this will allow the connection to succeed. If not, the connection will fail after ipc client retries (default 10 seconds worth of retries).\r\n\r\nI want to fix this in ipc client rather than just in Datanode startup, as this fixes temporary DNS issues for all of Hadoop.","issue_id":"13125761","key":"HADOOP-15129","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2021-09-03T19:47:21.000+0000","role":"fixed_distractor","summary":"Datanode caches namenode DNS lookup failure and cannot startup"} {"case_id":"13130473","cluster":"DISTRACTOR-HADOOP-15169","comments":[{"body":"Uploaded the {{branch-2}} patch.Kindly review.","created":"2018-01-12T10:45:39.067+0000"},{"body":"Uploading the trunk patch.\r\n\r\n{{branch-2}} compile and {{javadoc}} errors are unrelated.","created":"2018-01-12T12:09:47.225+0000"},{"body":"I am not an expert in HttpServer area so please let me know if some of my comments doesn't make sense.\r\n\r\nOverall the change looks good.\r\nHave few comments.\r\n+SslSelectChannelConnectorSecure.java+\r\n* The list of enabled protocols from SSLEngine are {{SSLv2Hello, TLSv1, TLSv1.1, TLSv1.2}}.\r\nAfter the change, we will filter out {{SSLv2Hello}}.\r\nI think it should be ok.\r\n\r\n* {quote}\r\n + private boolean contains(String p, String[] c) {\r\n{quote}\r\nWould be nice to use human readable variables instead of p and c.\r\n\r\nAlso it would be nice to have some supporting test cases.\r\n\r\n","created":"2018-01-12T20:11:59.136+0000"},{"body":"In Apache Hadoop 3.x, the Jetty version is greater than 9.3.12 and it only accepts TLS 1.2 by default. I don't want to add a setting to accept TLS 1.1 or older protocols to create a security hole for now. When we have migrated to Java 11 and Jetty 9.4.x to use TLS 1.3, then we can add the setting for Jetty server.\r\n\r\nOn the other hand, in Apache Hadoop 2.x, adding the setting for HttpServer2 makes sense to me. That way we can avoid using SSLv2Hello, TLSv1, or TLSv1.1 in HttpServer2.","created":"2018-12-12T11:17:00.580+0000"},{"body":"This one looks important. How do we want to do about it?","created":"2019-08-17T11:44:41.031+0000"},{"body":"[~aajisaka] I understand your concern. However, this is merely to achieve consistency with other Hadoop components. We've got customers with legacy tools that can only support SSLv2Hello, and they aren't able to use it after upgrading to Hadoop 3.\r\n\r\n \r\n\r\n[~brahmareddy] thanks for the patch. have you tested it? Looking at Jetty's SslContextFactory implementation (SslContextFactory#selectProtocols()), after included protocols are added, it removes excluded protocols, which contains \r\n\r\n\"SSL\", \"SSLv2\", \"SSLv2Hello\", \"SSLv3\". I suspect we should reset excluded protocols before adding included protocols.","created":"2019-09-09T02:28:55.425+0000"},{"body":"{quote}We've got customers with legacy tools that can only support SSLv2Hello, and they aren't able to use it after upgrading to Hadoop 3.\r\n{quote}\r\nThanks [~weichiu]. Agreed.","created":"2019-09-10T09:46:51.588+0000"},{"body":"[~brahmareddy] After Jetty 9.2.4.v20141103 (which is the case after Hadoop 3.0), SSLv2Hello is excluded by default, and we have to explicitly reset excluded protocols otherwise your change won't take effect.\r\n\r\nI am attaching an updated patch with a test, which correctly add supported protocols specified. Please take a look.\r\n\r\nNote also that an optional include list of protocols can potentially disable vulnerable protocols in the future as well.","created":"2019-10-10T00:41:06.632+0000"},{"body":"Thanks [~brahmareddy] and [~weichiu] for the patch. It looks good to me overall.\r\n\r\nI just have one suggestion w.r.t. the handling of the excluded protocols. By default SslContextFactory will set the following (\"SSL\", \"SSLv2\", \"SSLv2Hello\", \"SSLv3\") to the excluded protocol. \r\n\r\nInstead of always reset the excluded protocol to empty, we should remove only those contained in the enabledProtocols from the excluded protocol. This way, we don't allow weak protocols not in the enable list.\r\n\r\nPlease also add a test case to ensure if use add SSLv2Hello to included protocol, SSL/SSLv2/SSLv3 should not be allowed.\r\n\r\n \r\n\r\n ","created":"2019-10-11T18:16:20.398+0000"},{"body":"Thanks [~xyao]!\r\n\r\nGood idea. I'll update the patch accordingly.\r\n\r\nHowever, Jetty won't let me enable SSLv2Hello only. I must also enable another protocol otherwise it fails right away. I think that's by design. So this doesn't really need a test case.","created":"2019-10-11T22:49:39.540+0000"},{"body":"Thanks [~weichiu] for v3 patch. It looks good to me. One minor comments: In line 551, we use equals to compare the protocol strings, do we need to handle the case when the orders are different but the protocols are same?\r\n|551|if (!enabledProtocols.equals(SSLFactory.SSL_ENABLED_PROTOCOLS_DEFAULT)) {|","created":"2019-10-14T23:15:39.182+0000"},{"body":"It is fine in the current Jetty version and Hadoop 3.3.0 where the default excluded protocols in Jetty and default enabled protocols in Hadoop don't overlap. The effect is the same where the order is reversed or not.\r\n\r\nIf one day we update to a future Jetty version that excludes more protocols and some of which are permitted by Hadoop by default, I would like them not to be enabled (Jetty's exclude list takes precedence over include list) by default, unless user consciously update the hadoop.ssl.enabled.protocols configuration.","created":"2019-10-15T00:18:07.117+0000"},{"body":"Agree, thanks [~weichiu]. +1.\r\n\r\nThere is one checkstyle issue that you can fix at commit.","created":"2019-10-15T18:52:19.939+0000"},{"body":"Pushed to trunk, thanks [~brahmareddy] for the initial patch, and [~aajisaka] and [~xyao] for the review and comments!","created":"2019-10-15T20:58:02.805+0000"}],"conversations":[{"body":"As of now *hadoop.ssl.enabled.protocols\"* will not take effect for all the http servers( only Datanodehttp server will use this config).","from":"reporter","subject":"\"hadoop.ssl.enabled.protocols\" should be considered in httpserver2"},{"body":"Uploaded the {{branch-2}} patch.Kindly review.","from":"developer"},{"body":"Uploading the trunk patch.\r\n\r\n{{branch-2}} compile and {{javadoc}} errors are unrelated.","from":"developer"},{"body":"I am not an expert in HttpServer area so please let me know if some of my comments doesn't make sense.\r\n\r\nOverall the change looks good.\r\nHave few comments.\r\n+SslSelectChannelConnectorSecure.java+\r\n* The list of enabled protocols from SSLEngine are {{SSLv2Hello, TLSv1, TLSv1.1, TLSv1.2}}.\r\nAfter the change, we will filter out {{SSLv2Hello}}.\r\nI think it should be ok.\r\n\r\n* {quote}\r\n + private boolean contains(String p, String[] c) {\r\n{quote}\r\nWould be nice to use human readable variables instead of p and c.\r\n\r\nAlso it would be nice to have some supporting test cases.\r\n\r\n","from":"developer"},{"body":"In Apache Hadoop 3.x, the Jetty version is greater than 9.3.12 and it only accepts TLS 1.2 by default. I don't want to add a setting to accept TLS 1.1 or older protocols to create a security hole for now. When we have migrated to Java 11 and Jetty 9.4.x to use TLS 1.3, then we can add the setting for Jetty server.\r\n\r\nOn the other hand, in Apache Hadoop 2.x, adding the setting for HttpServer2 makes sense to me. That way we can avoid using SSLv2Hello, TLSv1, or TLSv1.1 in HttpServer2.","from":"developer"},{"body":"This one looks important. How do we want to do about it?","from":"developer"},{"body":"[~aajisaka] I understand your concern. However, this is merely to achieve consistency with other Hadoop components. We've got customers with legacy tools that can only support SSLv2Hello, and they aren't able to use it after upgrading to Hadoop 3.\r\n\r\n \r\n\r\n[~brahmareddy] thanks for the patch. have you tested it? Looking at Jetty's SslContextFactory implementation (SslContextFactory#selectProtocols()), after included protocols are added, it removes excluded protocols, which contains \r\n\r\n\"SSL\", \"SSLv2\", \"SSLv2Hello\", \"SSLv3\". I suspect we should reset excluded protocols before adding included protocols.","from":"developer"},{"body":"{quote}We've got customers with legacy tools that can only support SSLv2Hello, and they aren't able to use it after upgrading to Hadoop 3.\r\n{quote}\r\nThanks [~weichiu]. Agreed.","from":"developer"},{"body":"[~brahmareddy] After Jetty 9.2.4.v20141103 (which is the case after Hadoop 3.0), SSLv2Hello is excluded by default, and we have to explicitly reset excluded protocols otherwise your change won't take effect.\r\n\r\nI am attaching an updated patch with a test, which correctly add supported protocols specified. Please take a look.\r\n\r\nNote also that an optional include list of protocols can potentially disable vulnerable protocols in the future as well.","from":"developer"},{"body":"Thanks [~brahmareddy] and [~weichiu] for the patch. It looks good to me overall.\r\n\r\nI just have one suggestion w.r.t. the handling of the excluded protocols. By default SslContextFactory will set the following (\"SSL\", \"SSLv2\", \"SSLv2Hello\", \"SSLv3\") to the excluded protocol. \r\n\r\nInstead of always reset the excluded protocol to empty, we should remove only those contained in the enabledProtocols from the excluded protocol. This way, we don't allow weak protocols not in the enable list.\r\n\r\nPlease also add a test case to ensure if use add SSLv2Hello to included protocol, SSL/SSLv2/SSLv3 should not be allowed.\r\n\r\n \r\n\r\n ","from":"developer"},{"body":"Thanks [~xyao]!\r\n\r\nGood idea. I'll update the patch accordingly.\r\n\r\nHowever, Jetty won't let me enable SSLv2Hello only. I must also enable another protocol otherwise it fails right away. I think that's by design. So this doesn't really need a test case.","from":"developer"},{"body":"Thanks [~weichiu] for v3 patch. It looks good to me. One minor comments: In line 551, we use equals to compare the protocol strings, do we need to handle the case when the orders are different but the protocols are same?\r\n|551|if (!enabledProtocols.equals(SSLFactory.SSL_ENABLED_PROTOCOLS_DEFAULT)) {|","from":"developer"},{"body":"It is fine in the current Jetty version and Hadoop 3.3.0 where the default excluded protocols in Jetty and default enabled protocols in Hadoop don't overlap. The effect is the same where the order is reversed or not.\r\n\r\nIf one day we update to a future Jetty version that excludes more protocols and some of which are permitted by Hadoop by default, I would like them not to be enabled (Jetty's exclude list takes precedence over include list) by default, unless user consciously update the hadoop.ssl.enabled.protocols configuration.","from":"developer"},{"body":"Agree, thanks [~weichiu]. +1.\r\n\r\nThere is one checkstyle issue that you can fix at commit.","from":"developer"},{"body":"Pushed to trunk, thanks [~brahmareddy] for the initial patch, and [~aajisaka] and [~xyao] for the review and comments!","from":"developer"}],"created":"2018-01-12T09:36:54.000+0000","description":"As of now *hadoop.ssl.enabled.protocols\"* will not take effect for all the http servers( only Datanodehttp server will use this config).","issue_id":"13130473","key":"HADOOP-15169","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2019-10-15T20:58:02.000+0000","role":"fixed_distractor","summary":"\"hadoop.ssl.enabled.protocols\" should be considered in httpserver2"} {"case_id":"13145877","cluster":"DISTRACTOR-HADOOP-15320","comments":[{"body":"Interesting. In HADOOP-14943 I'd proposed pulling up the azure one to hadoop common for shared use, spec a bit tighter what it did and then wire up S3A to it too.\r\n\r\nNow you are saying for multiTB files we don't need this code at all? well, that's good news.\r\nI see your arguments, but do think it will need be bounced past the various tools, including: hive, spark, pig to see that it all goes OK. But given S3A is using that default with no adverse consequences, I think you'll be right.\r\n\r\nAs usual: which endpoints did you run the entire hadoop-azure and hadoop-azuredatalake test suites?","created":"2018-03-17T18:08:14.907+0000"},{"body":"What testing has been done with this, already?\r\n\r\nbq. do think it will need be bounced past the various tools, including: hive, spark, pig to see that it all goes OK. But given S3A is using that default with no adverse consequences, I think you'll be right.\r\nWouldn't one expect the same results, if the pattern worked for S3A? One would expect to find framework code that is unnecessarily serial after this change. What tests did S3A run that should be repeated?\r\n\r\nbq. which endpoints did you run the entire hadoop-azure and hadoop-azuredatalake test suites?\r\nRunning these integration tests is a good idea. It's why they're there, after all.","created":"2018-03-19T23:26:44.088+0000"},{"body":"I've run a few Spark jobs on very large input file (hundreds of TB) and the getSplits() on this file took a few seconds, vs. 1.5 hours without the change.\r\n\r\nI'm in the middle of running hive tpch tests.\r\n\r\nAnything else we should run?\r\n\r\nAs [~chris.douglas] mentioned, since S3A is running file, we should be good to go for this patch.","created":"2018-03-20T01:37:46.544+0000"},{"body":"bq. Anything else we should run?\r\nAs [~stevel@apache.org] suggested, the hadoop-azure and hadoop-azuredatalake test suites and contract tests should pass.","created":"2018-03-22T20:45:36.461+0000"},{"body":"I know s3 \"appears\" to work, but I'm not actually confident that everything is getting the splits right there.\r\n\r\nThe one I want you look at is: Spark, CSV, multiGB: SPARK-22240 . That's what's been niggling at me for a while. ","created":"2018-03-23T16:59:01.983+0000"},{"body":"bq. I know s3 \"appears\" to work, but I'm not actually confident that everything is getting the splits right there.\r\nMe either, but 1.5h to generate synthetic splits is definitely wrong. If we develop a new, best practice for object stores, then we can apply that across the stores we support. The {{Blocklocations[]}} return type is pretty restrictive, but we could probably do better.\r\n\r\nbq. The one I want you look at is: Spark, CSV, multiGB: SPARK-22240 . That's what's been niggling at me for a while.\r\nMaybe I'm missing the bug. Block locations are hints for locality, not format partitioning. In that JIRA: gzip is not splittable, so a single reader is correct, absent some other preparation (saving the dictionary at offsets, writing zero-length gzip files as split markers, etc.). In general, framework parallelism should not rely exclusively on block locations...","created":"2018-03-23T18:17:17.900+0000"},{"body":"OK, maybe I'm the confused one. Given it's working for S3a, if you're happy, so am I","created":"2018-03-23T18:22:40.777+0000"},{"body":"I also run the following manual tests successfully:\r\n\r\n1) Hive TPCH test with my change on WASB and it passed with correct number of splits.\r\n\r\n2) Spark application to convert a huge CSV file to parquet.","created":"2018-03-27T21:36:17.154+0000"},{"body":"Thanks [~shanyu]. Running this through Jenkins.\r\n\r\nWe could add a unit test to signal that a change to the default behavior could affect these FS implementations, but that should be implied.","created":"2018-03-27T23:01:11.790+0000"},{"body":"Fixed checkstyle warnings. +1 from me. [~stevel@apache.org], lgty?","created":"2018-03-28T00:34:58.180+0000"},{"body":"If you are happy, I am\r\n\r\n+1","created":"2018-03-28T10:24:54.646+0000"},{"body":"I committed this. Thanks [~shanyu]","created":"2018-03-28T19:05:40.917+0000"}],"conversations":[{"body":"hadoop-azure and hadoop-azure-datalake have its own implementation of getFileBlockLocations(), which faked a list of artificial blocks based on the hard-coded block size. And each block has one host with name \"localhost\". Take a look at this code:\r\n\r\n[https://github.com/apache/hadoop/blob/release-2.9.0-RC3/hadoop-tools/hadoop-azure/src/main/java/org/apache/hadoop/fs/azure/NativeAzureFileSystem.java#L3485]\r\n\r\nThis is a unnecessary mock up for a \"remote\" file system to mimic HDFS. And the problem with this mock is that for large (~TB) files we generates lots of artificial blocks, and FileInputFormat.getSplits() is slow in calculating splits based on these blocks.\r\n\r\nWe can safely remove this customized getFileBlockLocations() implementation, fall back to the default FileSystem.getFileBlockLocations() implementation, which is to return 1 block for any file with 1 host \"localhost\". Note that this doesn't mean we will create much less splits, because the number of splits is still limited by the blockSize in FileInputFormat.computeSplitSize():\r\n{code:java}\r\nreturn Math.max(minSize, Math.min(goalSize, blockSize));{code}","from":"reporter","subject":"Remove customized getFileBlockLocations for hadoop-azure and hadoop-azure-datalake"},{"body":"Interesting. In HADOOP-14943 I'd proposed pulling up the azure one to hadoop common for shared use, spec a bit tighter what it did and then wire up S3A to it too.\r\n\r\nNow you are saying for multiTB files we don't need this code at all? well, that's good news.\r\nI see your arguments, but do think it will need be bounced past the various tools, including: hive, spark, pig to see that it all goes OK. But given S3A is using that default with no adverse consequences, I think you'll be right.\r\n\r\nAs usual: which endpoints did you run the entire hadoop-azure and hadoop-azuredatalake test suites?","from":"developer"},{"body":"What testing has been done with this, already?\r\n\r\nbq. do think it will need be bounced past the various tools, including: hive, spark, pig to see that it all goes OK. But given S3A is using that default with no adverse consequences, I think you'll be right.\r\nWouldn't one expect the same results, if the pattern worked for S3A? One would expect to find framework code that is unnecessarily serial after this change. What tests did S3A run that should be repeated?\r\n\r\nbq. which endpoints did you run the entire hadoop-azure and hadoop-azuredatalake test suites?\r\nRunning these integration tests is a good idea. It's why they're there, after all.","from":"developer"},{"body":"I've run a few Spark jobs on very large input file (hundreds of TB) and the getSplits() on this file took a few seconds, vs. 1.5 hours without the change.\r\n\r\nI'm in the middle of running hive tpch tests.\r\n\r\nAnything else we should run?\r\n\r\nAs [~chris.douglas] mentioned, since S3A is running file, we should be good to go for this patch.","from":"developer"},{"body":"bq. Anything else we should run?\r\nAs [~stevel@apache.org] suggested, the hadoop-azure and hadoop-azuredatalake test suites and contract tests should pass.","from":"developer"},{"body":"I know s3 \"appears\" to work, but I'm not actually confident that everything is getting the splits right there.\r\n\r\nThe one I want you look at is: Spark, CSV, multiGB: SPARK-22240 . That's what's been niggling at me for a while. ","from":"developer"},{"body":"bq. I know s3 \"appears\" to work, but I'm not actually confident that everything is getting the splits right there.\r\nMe either, but 1.5h to generate synthetic splits is definitely wrong. If we develop a new, best practice for object stores, then we can apply that across the stores we support. The {{Blocklocations[]}} return type is pretty restrictive, but we could probably do better.\r\n\r\nbq. The one I want you look at is: Spark, CSV, multiGB: SPARK-22240 . That's what's been niggling at me for a while.\r\nMaybe I'm missing the bug. Block locations are hints for locality, not format partitioning. In that JIRA: gzip is not splittable, so a single reader is correct, absent some other preparation (saving the dictionary at offsets, writing zero-length gzip files as split markers, etc.). In general, framework parallelism should not rely exclusively on block locations...","from":"developer"},{"body":"OK, maybe I'm the confused one. Given it's working for S3a, if you're happy, so am I","from":"developer"},{"body":"I also run the following manual tests successfully:\r\n\r\n1) Hive TPCH test with my change on WASB and it passed with correct number of splits.\r\n\r\n2) Spark application to convert a huge CSV file to parquet.","from":"developer"},{"body":"Thanks [~shanyu]. Running this through Jenkins.\r\n\r\nWe could add a unit test to signal that a change to the default behavior could affect these FS implementations, but that should be implied.","from":"developer"},{"body":"Fixed checkstyle warnings. +1 from me. [~stevel@apache.org], lgty?","from":"developer"},{"body":"If you are happy, I am\r\n\r\n+1","from":"developer"},{"body":"I committed this. Thanks [~shanyu]","from":"developer"}],"created":"2018-03-16T22:20:35.000+0000","description":"hadoop-azure and hadoop-azure-datalake have its own implementation of getFileBlockLocations(), which faked a list of artificial blocks based on the hard-coded block size. And each block has one host with name \"localhost\". Take a look at this code:\r\n\r\n[https://github.com/apache/hadoop/blob/release-2.9.0-RC3/hadoop-tools/hadoop-azure/src/main/java/org/apache/hadoop/fs/azure/NativeAzureFileSystem.java#L3485]\r\n\r\nThis is a unnecessary mock up for a \"remote\" file system to mimic HDFS. And the problem with this mock is that for large (~TB) files we generates lots of artificial blocks, and FileInputFormat.getSplits() is slow in calculating splits based on these blocks.\r\n\r\nWe can safely remove this customized getFileBlockLocations() implementation, fall back to the default FileSystem.getFileBlockLocations() implementation, which is to return 1 block for any file with 1 host \"localhost\". Note that this doesn't mean we will create much less splits, because the number of splits is still limited by the blockSize in FileInputFormat.computeSplitSize():\r\n{code:java}\r\nreturn Math.max(minSize, Math.min(goalSize, blockSize));{code}","issue_id":"13145877","key":"HADOOP-15320","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2018-03-28T19:05:40.000+0000","role":"fixed_distractor","summary":"Remove customized getFileBlockLocations for hadoop-azure and hadoop-azure-datalake"} {"case_id":"13149737","cluster":"DISTRACTOR-HADOOP-15358","comments":[{"body":"Fixed SFTPConnectionPool connections leakage","created":"2018-08-22T09:07:49.999+0000"},{"body":"+1. It looks good to me.\r\n\r\nVery precise problem definition and clean implementation with unit test. It's backward compatible as the old methods are still there (which opens the new connections).\r\n\r\nUnit test is passed + I tested it with 'dfs ls' and it worked well. I also noticed the recursive behaviour with java debug.\r\n\r\nWill commit it to the trunk soon... ","created":"2018-11-23T08:47:14.469+0000"},{"body":"Committed to the trunk. Thanks the contribution and sorry for the delayed review.\r\n\r\n(I fixed the unused import + line length checkstyle issues during the commit)","created":"2018-11-23T08:58:57.385+0000"}],"conversations":[{"body":"Methods of SFTPFileSystem operate on poolable ChannelSftp instances, thus some methods of SFTPFileSystem are chained together resulting in establishing multiple connections to the SFTP server to accomplish one compound action, those methods are listed below:\r\n # mkdirs method\r\nthe public mkdirs method acquires a new ChannelSftp from the pool [1]\r\nand then recursively creates directories, checking for the directory existence beforehand by calling the method exists[2] which delegates to the getFileStatus(ChannelSftp channel, Path file) method [3] and so on until it ends up in returning the FilesStatus instance [4]. The resource leakage occurs in the method getWorkingDirectory which calls the getHomeDirectory method [5] which in turn establishes a new connection to the sftp server instead of using an already created connection. As the mkdirs method is recursive this results in creating a huge number of connections.\r\n # open method [6]. This method returns an instance of FSDataInputStream which consumes SFTPInputStream instance which doesn't return an acquired ChannelSftp instance back to the pool but instead it closes it[7]. This leads to establishing another connection to an SFTP server when the next method is called on the FileSystem instance.\r\n\r\n\r\n[1] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L658\r\n\r\n[2] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L321\r\n\r\n[3] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L202\r\n\r\n[4] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L290\r\n\r\n[5] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L640\r\n\r\n[6] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L504\r\n\r\n[7] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPInputStream.java#L123","from":"reporter","subject":"SFTPConnectionPool connections leakage"},{"body":"Fixed SFTPConnectionPool connections leakage","from":"developer"},{"body":"+1. It looks good to me.\r\n\r\nVery precise problem definition and clean implementation with unit test. It's backward compatible as the old methods are still there (which opens the new connections).\r\n\r\nUnit test is passed + I tested it with 'dfs ls' and it worked well. I also noticed the recursive behaviour with java debug.\r\n\r\nWill commit it to the trunk soon... ","from":"developer"},{"body":"Committed to the trunk. Thanks the contribution and sorry for the delayed review.\r\n\r\n(I fixed the unused import + line length checkstyle issues during the commit)","from":"developer"}],"created":"2018-04-03T13:11:41.000+0000","description":"Methods of SFTPFileSystem operate on poolable ChannelSftp instances, thus some methods of SFTPFileSystem are chained together resulting in establishing multiple connections to the SFTP server to accomplish one compound action, those methods are listed below:\r\n # mkdirs method\r\nthe public mkdirs method acquires a new ChannelSftp from the pool [1]\r\nand then recursively creates directories, checking for the directory existence beforehand by calling the method exists[2] which delegates to the getFileStatus(ChannelSftp channel, Path file) method [3] and so on until it ends up in returning the FilesStatus instance [4]. The resource leakage occurs in the method getWorkingDirectory which calls the getHomeDirectory method [5] which in turn establishes a new connection to the sftp server instead of using an already created connection. As the mkdirs method is recursive this results in creating a huge number of connections.\r\n # open method [6]. This method returns an instance of FSDataInputStream which consumes SFTPInputStream instance which doesn't return an acquired ChannelSftp instance back to the pool but instead it closes it[7]. This leads to establishing another connection to an SFTP server when the next method is called on the FileSystem instance.\r\n\r\n\r\n[1] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L658\r\n\r\n[2] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L321\r\n\r\n[3] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L202\r\n\r\n[4] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L290\r\n\r\n[5] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L640\r\n\r\n[6] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPFileSystem.java#L504\r\n\r\n[7] https://github.com/apache/hadoop/blob/736ceab2f58fb9ab5907c5b5110bd44384038e6b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/sftp/SFTPInputStream.java#L123","issue_id":"13149737","key":"HADOOP-15358","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2018-11-23T08:58:57.000+0000","role":"fixed_distractor","summary":"SFTPConnectionPool connections leakage"} {"case_id":"13165009","cluster":"DISTRACTOR-HADOOP-15524","comments":[{"body":"Thanks [~jesmith3] for reporting this.\r\nYes we should handle this case in {{BytesWritable#setSize}} in a similar way as it's handled in {{ArrayList}}.\r\n\r\nPlease let me know if you need any help in creating a patch for this issue.","created":"2018-06-13T13:54:59.966+0000"},{"body":"Thanks, [~nandakumar131]. I put up a PR as seen above.  If a patch is still required let me know and I will include it as well.","created":"2018-06-13T15:46:56.159+0000"},{"body":"Thanks [~jesmith3] for the PR, the change looks good to me.\r\ncc: [~arpitagarwal], can you please add [~jesmith3] as a contributor and assign this jira to him.","created":"2018-06-14T04:48:39.046+0000"},{"body":"Hi [~nanda], [~arpaga] , Can we commit this Jira ? Please let me know if there are any concerns on this fix to be worried or safe to commit ? We have got a customer request on this and might need a backport .","created":"2020-05-08T07:24:17.119+0000"},{"body":"Hey [~nanda] can you take a quick look and confirm you are still +1 on this change? I probably missed the notification when you tagged me last time.","created":"2020-05-11T14:45:48.101+0000"},{"body":"Thanks for the update [~arp].\r\nI'm +1 on the change. Just retriggered Jenkins, will merge it after the build.\r\n\r\nhttps://builds.apache.org/job/hadoop-multibranch/job/PR-393/10/","created":"2020-05-11T18:14:15.278+0000"},{"body":"Thanks [~jesmith3] for the contribution. Merged the changes to trunk.","created":"2020-05-12T18:51:39.449+0000"}],"conversations":[{"body":"BytesWritable.setSize uses Integer.MAX_VALUE to initialize the internal array.  On my environment, this causes an OOME\r\n{code:java}\r\nException in thread \"main\" java.lang.OutOfMemoryError: Requested array size exceeds VM limit\r\n{code}\r\nbyte[Integer.MAX_VALUE-2] must be used to prevent this error.\r\n\r\nTested on OSX and CentOS 7 using Java version 1.8.0_131.\r\n\r\nI noticed that java.util.ArrayList contains the following\r\n{code:java}\r\n/**\r\n * The maximum size of array to allocate.\r\n * Some VMs reserve some header words in an array.\r\n * Attempts to allocate larger arrays may result in\r\n * OutOfMemoryError: Requested array size exceeds VM limit\r\n */\r\nprivate static final int MAX_ARRAY_SIZE = Integer.MAX_VALUE - 8;\r\n{code}\r\n \r\n\r\nBytesWritable.setSize should use something similar to prevent an OOME from occurring.\r\n\r\n ","from":"reporter","subject":"BytesWritable causes OOME when array size reaches Integer.MAX_VALUE"},{"body":"Thanks [~jesmith3] for reporting this.\r\nYes we should handle this case in {{BytesWritable#setSize}} in a similar way as it's handled in {{ArrayList}}.\r\n\r\nPlease let me know if you need any help in creating a patch for this issue.","from":"developer"},{"body":"Thanks, [~nandakumar131]. I put up a PR as seen above.  If a patch is still required let me know and I will include it as well.","from":"developer"},{"body":"Thanks [~jesmith3] for the PR, the change looks good to me.\r\ncc: [~arpitagarwal], can you please add [~jesmith3] as a contributor and assign this jira to him.","from":"developer"},{"body":"Hi [~nanda], [~arpaga] , Can we commit this Jira ? Please let me know if there are any concerns on this fix to be worried or safe to commit ? We have got a customer request on this and might need a backport .","from":"developer"},{"body":"Hey [~nanda] can you take a quick look and confirm you are still +1 on this change? I probably missed the notification when you tagged me last time.","from":"developer"},{"body":"Thanks for the update [~arp].\r\nI'm +1 on the change. Just retriggered Jenkins, will merge it after the build.\r\n\r\nhttps://builds.apache.org/job/hadoop-multibranch/job/PR-393/10/","from":"developer"},{"body":"Thanks [~jesmith3] for the contribution. Merged the changes to trunk.","from":"developer"}],"created":"2018-06-08T19:33:48.000+0000","description":"BytesWritable.setSize uses Integer.MAX_VALUE to initialize the internal array.  On my environment, this causes an OOME\r\n{code:java}\r\nException in thread \"main\" java.lang.OutOfMemoryError: Requested array size exceeds VM limit\r\n{code}\r\nbyte[Integer.MAX_VALUE-2] must be used to prevent this error.\r\n\r\nTested on OSX and CentOS 7 using Java version 1.8.0_131.\r\n\r\nI noticed that java.util.ArrayList contains the following\r\n{code:java}\r\n/**\r\n * The maximum size of array to allocate.\r\n * Some VMs reserve some header words in an array.\r\n * Attempts to allocate larger arrays may result in\r\n * OutOfMemoryError: Requested array size exceeds VM limit\r\n */\r\nprivate static final int MAX_ARRAY_SIZE = Integer.MAX_VALUE - 8;\r\n{code}\r\n \r\n\r\nBytesWritable.setSize should use something similar to prevent an OOME from occurring.\r\n\r\n ","issue_id":"13165009","key":"HADOOP-15524","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2020-05-12T18:51:39.000+0000","role":"fixed_distractor","summary":"BytesWritable causes OOME when array size reaches Integer.MAX_VALUE"} {"case_id":"13182310","cluster":"DISTRACTOR-HADOOP-15708","comments":[{"body":"The problem here is that the deprecation registration declares that mapping from old to new. We cannot expect a property registered with the deprecated key to be retrievable with the new key until then. \r\n\r\nIt should still be retrievable with the old key, just like any other key.\r\n\r\nAre you seeing something different?\r\n\r\nOr is the issue here that the act of registration of a deprecated key changing what/whether you can get stuff back?\r\n\r\n\r\npatch wise: nice to see a test, especially with the assertEquals() parameters in the right order. But...SLF4J must be the log API, not stdout. thanks.","created":"2018-08-31T09:35:10.611+0000"},{"body":"{quote} We cannot expect a property registered with the deprecated key to be retrievable with the new key until then. {quote}\r\nYes, I understand this. What I tried to describe is not this case.\r\nThe problem here is if you get some value from the conf then register the deprecation mappings, you are unable to read the conf value with the old key, which is not a desired behavior.\r\nIf you check the code + logs, I try to get the values with the old and new keys, but the output says:\r\n\r\n{code:java}\r\nLooked up property value with name hadoop.zk.address: null\r\nLooked up property value with name hadoop.zk.address: null\r\n{code}\r\n\r\nSo even if you use the old key, it tries to fetch values with the new key and it's null.\r\n\r\nI just meant this patch to be a proof of concept, so please disregard the stdouts, as the final patch will be more production ready.\r\nThanks!","created":"2018-08-31T10:20:50.867+0000"},{"body":"{{So to reproduce the problem:}}\r\n{{Works:}}\r\n\r\n{{1. load properties from xml}}\r\n{{2. add deprecation mapping}}\r\n{{3. read either new or old key {color:#8eb021}works{color}}}\r\n\r\n{{Does now work:}}\r\n{{1. load properties from xml}}\r\n{{{color:#59afe1}1,5. Get any completely not related property{color} }}\r\n{{2. add deprecation mapping}}\r\n{{3. read either new or old key {color:#d04437}fails{color}}}","created":"2018-08-31T10:26:41.967+0000"},{"body":"Hi [~zsiegl]!\r\nThanks for your patch!\r\nLGTM in overall, I would just rename the {{names}} array to {{newKeys}} and the corresponding loop variable {{n}} to {{newKey}}.\r\n\r\n{noformat}\r\nprivate void updatePropertiesWIthDeprecatedKeys(\r\n DeprecationContext deprecations, String[] names) {\r\n for (String n : names) {\r\n String deprecatedKey = deprecations.getReverseDeprecatedKeyMap().get(n);\r\n if (deprecatedKey != null && !getProps().containsKey(n)) {\r\n String deprecatedValue = getProps().getProperty(deprecatedKey);\r\n if (deprecatedValue != null) {\r\n getProps().setProperty(n, deprecatedValue);\r\n }\r\n }\r\n }\r\n }\r\n{noformat}\r\n\r\nThanks!","created":"2018-08-31T10:37:09.338+0000"},{"body":"[~snemeth] thanks for the review, items fixed.\r\n\r\n[~stevel@apache.org] could you review the current patch? Please let me know if you would need any further help understanding the context of the issue.","created":"2018-08-31T11:41:40.890+0000"},{"body":"Thanks [~zsiegl] for the updated patch!\r\nLGTM, +1 (non-binding)","created":"2018-08-31T12:27:48.621+0000"},{"body":"I'm going to wait to see what others say, which given its a long weekend in the US, unlikely until next week\r\n\r\nConfiguration is such as ubiquitous class that its on the list of things to be very careful about changing, even when the changes seem low risk, which is why we need that extra oversight\r\n\r\nlooking @ code; \r\n\r\n* I'd like to see a patch which doesn't add IDE-helped changes, as that complicates merging and cherry-picking, so fixes to unused imports, line endings, spacing need to be left out. Sorry.\r\n* in TestConfigurationDeprecation, use try-with-resources for writing the file.\r\n\r\n","created":"2018-08-31T17:13:32.999+0000"},{"body":"[~stevel@apache.org] Thank you for the review. Changes done, new patch on the way. Try-with-resources in the test would require a larger refactor of the test case, I suggest a different commit for that.","created":"2018-08-31T19:29:48.912+0000"},{"body":"[~stevel@apache.org] do you know of anyone who would be interested in this issue?\r\n[~xiaochen] Would you be interested to have a look at this problem? \r\n\r\nThanks :)","created":"2018-09-10T15:58:54.164+0000"},{"body":"[~jlowe] [~daryn] could someone have a look at this issue please?","created":"2018-09-24T19:48:47.324+0000"},{"body":"There's definitely something wrong here.  I was playing with the {{Configuration}} class like this:\r\n{code:java}\r\n// OLD_KEY is set to \"foo\" in the XML\r\nconf.get(OLD_KEY); // returns \"foo\"\r\nConfiguration.addDeprecations(OLD_KEY, NEW_KEY);\r\nconf.get(OLD_KEY); // returns null\r\n{code}\r\nThat doesn't seem right. Adding a deprecation looks like it breaks getting the value of the deprecated key. On top of that, if you ommit the 2nd line, it works correctly:\r\n{code:java}\r\n// OLD_KEY is set to \"foo\" in the XML\r\nConfiguration.addDeprecations(OLD_KEY, NEW_KEY);\r\nconf.get(OLD_KEY); // returns \"foo\"\r\n{code}\r\nThis means that getting the {{OLD_KEY}} and then deprecating it changes the value, which isn't the right behavior. \r\n\r\nAs for the right way to fix this, and make sure we're not breaking anything depending on this odd behavior, I'm not sure. It would be good to get some more opinions.","created":"2018-10-03T00:26:08.426+0000"},{"body":"[~rkanter] did you experience this behaviour before or after applying the patch?\r\nThis is what the patch should fix. Also before the patch: \r\n\r\n{code:java}\r\n// OLD_KEY is set to \"foo\" in the XML\r\nConfiguration.addDeprecations(OLD_KEY, NEW_KEY);\r\nconf.get(OLD_KEY); // returns \"foo\"\r\n{code}\r\n\r\nWhile:\r\n{code:java}\r\n// OLD_KEY is set to \"foo\" in the XML\r\nconf.get(SOMETHING_ELSE);\r\nConfiguration.addDeprecations(OLD_KEY, NEW_KEY);\r\nconf.get(OLD_KEY); // returns null\r\n{code}\r\n\r\n","created":"2018-10-03T01:29:05.068+0000"},{"body":"Sorry, yeah; that was before the patch.  I was just investigating what the current behavior is.  As [~stevel@apache.org] said, we have to be very careful about changing this class.  The code seems okay to me though, but let's see if anyone else has thoughts about the consequences of changing this behavior.","created":"2018-10-04T00:47:12.587+0000"},{"body":"So to make things a bit clear about the change:\r\nBefore the patch:\r\n{code:java}\r\nprivate String[] handleDeprecation(DeprecationContext deprecations,\r\n String name) {\r\n if (null != name) {\r\n name = name.trim();\r\n }\r\n // Initialize the return value with requested name\r\n String[] names = new String[]{name};\r\n // Deprecated keys are logged once and an updated names are returned\r\n DeprecatedKeyInfo keyInfo = deprecations.getDeprecatedKeyMap().get(name);\r\n if (keyInfo != null) {\r\n if (!keyInfo.getAndSetAccessed()) {\r\n logDeprecation(keyInfo.getWarningMessage(name));\r\n }\r\n // Override return value for deprecated keys\r\n names = keyInfo.newKeys;\r\n }\r\n\r\n // **** WE RETURN EARLY HERE IF THERE ARE NO OVERLAYS ***\r\n\r\n // If there are no overlay values we can return early\r\n Properties overlayProperties = getOverlay();\r\n if (overlayProperties.isEmpty()) {\r\n return names;\r\n }\r\n // Update properties and overlays with reverse lookup values\r\n for (String n : names) {\r\n String deprecatedKey = deprecations.getReverseDeprecatedKeyMap().get(n);\r\n if (deprecatedKey != null && !overlayProperties.containsKey(n)) {\r\n String deprecatedValue = overlayProperties.getProperty(deprecatedKey);\r\n if (deprecatedValue != null) {\r\n // **** STILL AT THIS LINE WE ARE UPDATING PROPS *NOT* OVERLAY ***\r\n getProps().setProperty(n, deprecatedValue);\r\n overlayProperties.setProperty(n, deprecatedValue);\r\n }\r\n }\r\n }\r\n return names;\r\n }\r\n{code}\r\n\r\nSo I have extracted the latter loop to run for all props not just overlays as it should work.\r\n\r\nTo address concerns of [~jlowe] as we talked offline:\r\n- You said that if you set deprecated and new keys as well, this patch would change the return value. To address this fear I have created a test case for this:\r\n\r\n{code:java}\r\n@Test\r\n public void testBothPropertiesAreSet() throws Exception {\r\n // SETUP\r\n final String oldZkAddressKey = \"yarn.resourcemanager.zk-address\";\r\n final String newZkAddressKey = CommonConfigurationKeys.ZK_ADDRESS;\r\n final String oldZkAddressValue = \"oldZkAddress\";\r\n final String newZkAddressValue = \"newZkAddress\";\r\n\r\n try{\r\n out = new BufferedWriter(new FileWriter(CONFIG4));\r\n startConfig();\r\n appendProperty(oldZkAddressKey, oldZkAddressValue);\r\n appendProperty(newZkAddressKey, newZkAddressValue);\r\n endConfig();\r\n\r\n Path fileResource = new Path(CONFIG4);\r\n conf.addResource(fileResource);\r\n } finally {\r\n out.close();\r\n }\r\n\r\n // ACT\r\n conf.get(oldZkAddressKey);\r\n Configuration.addDeprecations(new Configuration.DeprecationDelta[] {\r\n new Configuration.DeprecationDelta(oldZkAddressKey, newZkAddressKey)});\r\n\r\n // ASSERT\r\n assertEquals(\"New property should overwrite deprecated one\",\r\n newZkAddressValue, conf.get(oldZkAddressKey));\r\n assertEquals(\"Property should be accessible through new key\",\r\n newZkAddressValue, conf.get(newZkAddressKey));\r\n }\r\n{code}\r\n\r\nThis tests setting both properties and returning the same values before and after the patch as well.\r\nI did not add this to the patch, as this has been already covered in the previous test cases, it is juts a bit more readable this way, as a separate test. \r\n\r\nTo address fears of [~rkanter]:\r\n- You said that some may depend on the bugous functionality. That is normally a good idea to stay backward compatible, however why would you depend on getting NULL instead of the config value you needed? This is a bug. Don'd depend on bugs.\r\n\r\nIs there anything else I can do for this to be accepted? I'd be happy to.\r\n","created":"2018-10-04T19:06:22.075+0000"},{"body":"{quote}Don't depend on bugs.\r\n{quote}\r\nI agree, though we've had issues in the past.\r\n\r\n \r\n\r\nAnyway, I think we we should move forward on this. It seems pretty clear to me that the current behavior doesn't make sense and we should fix it. I'll wait about a week to give anyone else a chance to speak up about this, and then I'll commit the patch. Otherwise, I think we'll be waiting forever while everyone hopes someone else will do the commit.\r\n\r\nOne nitpick with the patch: {{updatePropertiesWIthDeprecatedKeys}} has an extra capitalization in {{WIth}}. It should be {{updatePropertiesWithDeprecatedKeys}}.","created":"2018-10-04T22:36:56.702+0000"},{"body":"[~rkanter] thanks for the review, and all the care about this. Fixed the typo. New patch uploaded.","created":"2018-10-04T22:45:44.261+0000"},{"body":"Thanks [~zsiegl] and everyone else for reviews.  Committed to trunk!","created":"2018-10-11T01:52:13.063+0000"},{"body":"Looks like we missed a trivial problem where one of the updated tests leaves behind a file, which causes test runs to complain about a wrong license.  I've filed HADOOP-15853 to fix that.","created":"2018-10-15T19:52:35.387+0000"}],"conversations":[{"body":"Hadoop Common contains a widely used Configuration class.\r\n This class can handle deprecations of properties, e.g. if property 'A' gets deprecated with an alternative property key 'B', users can access property values with keys 'A' and 'B'.\r\n Unfortunately, this does not work in one case.\r\n When a config file is specified (for instance, XML) and a property is read with the config.get() method, the config is loaded from the file at this time. \r\n If the deprecation mapping is not yet specified by the time any config value is retrieved and the XML config refers to a deprecated key, then the deprecation mapping specified, the config value cannot be retrieved neither with the deprecated nor with the new key.\r\n The attached patch contains a testcase that reproduces this wrong behavior.\r\n\r\nHere are the steps outlined what the testcase does:\r\n 1. Creates an XML config file with a deprecated property\r\n 2. Adds the config to the Configuration object\r\n 3. Retrieves the config with its deprecated key (it does not really matter which property the user gets, could be any)\r\n 4. Specifies the deprecation rules including the one defined in the config\r\n 5. Prints and asserts the property retrieved from the config with both the deprecated and the new property keys.\r\n\r\nFor reference, here is the log of one execution that actually shows what the issue is:\r\n{noformat}\r\nLoaded items: 1\r\nLooked up property value with name hadoop.zk.address: null\r\nLooked up property value with name yarn.resourcemanager.zk-address: dummyZkAddress\r\nContents of config file: [, , yarn.resourcemanager.zk-addressdummyZkAddress, ]\r\nLooked up property value with name hadoop.zk.address: null\r\n2018-08-31 10:10:06,484 INFO Configuration.deprecation (Configuration.java:logDeprecation(1397)) - yarn.resourcemanager.zk-address is deprecated. Instead, use hadoop.zk.address\r\nLooked up property value with name hadoop.zk.address: null\r\nLooked up property value with name hadoop.zk.address: null\r\n\r\njava.lang.AssertionError: \r\nExpected :dummyZkAddress\r\nActual :null\r\n{noformat}\r\n*As it's visible from the output and the code, the issue is really that if the config is retrieved either with the deprecated or the new value, Configuration both wants to serve the value with the new key.*\r\n *If the mapping is not specified before any retrieval happened, the value is only stored under the deprecated key but not the new key.*","from":"reporter","subject":"Reading values from Configuration before adding deprecations make it impossible to read value with deprecated key"},{"body":"The problem here is that the deprecation registration declares that mapping from old to new. We cannot expect a property registered with the deprecated key to be retrievable with the new key until then. \r\n\r\nIt should still be retrievable with the old key, just like any other key.\r\n\r\nAre you seeing something different?\r\n\r\nOr is the issue here that the act of registration of a deprecated key changing what/whether you can get stuff back?\r\n\r\n\r\npatch wise: nice to see a test, especially with the assertEquals() parameters in the right order. But...SLF4J must be the log API, not stdout. thanks.","from":"developer"},{"body":"{quote} We cannot expect a property registered with the deprecated key to be retrievable with the new key until then. {quote}\r\nYes, I understand this. What I tried to describe is not this case.\r\nThe problem here is if you get some value from the conf then register the deprecation mappings, you are unable to read the conf value with the old key, which is not a desired behavior.\r\nIf you check the code + logs, I try to get the values with the old and new keys, but the output says:\r\n\r\n{code:java}\r\nLooked up property value with name hadoop.zk.address: null\r\nLooked up property value with name hadoop.zk.address: null\r\n{code}\r\n\r\nSo even if you use the old key, it tries to fetch values with the new key and it's null.\r\n\r\nI just meant this patch to be a proof of concept, so please disregard the stdouts, as the final patch will be more production ready.\r\nThanks!","from":"developer"},{"body":"{{So to reproduce the problem:}}\r\n{{Works:}}\r\n\r\n{{1. load properties from xml}}\r\n{{2. add deprecation mapping}}\r\n{{3. read either new or old key {color:#8eb021}works{color}}}\r\n\r\n{{Does now work:}}\r\n{{1. load properties from xml}}\r\n{{{color:#59afe1}1,5. Get any completely not related property{color} }}\r\n{{2. add deprecation mapping}}\r\n{{3. read either new or old key {color:#d04437}fails{color}}}","from":"developer"},{"body":"Hi [~zsiegl]!\r\nThanks for your patch!\r\nLGTM in overall, I would just rename the {{names}} array to {{newKeys}} and the corresponding loop variable {{n}} to {{newKey}}.\r\n\r\n{noformat}\r\nprivate void updatePropertiesWIthDeprecatedKeys(\r\n DeprecationContext deprecations, String[] names) {\r\n for (String n : names) {\r\n String deprecatedKey = deprecations.getReverseDeprecatedKeyMap().get(n);\r\n if (deprecatedKey != null && !getProps().containsKey(n)) {\r\n String deprecatedValue = getProps().getProperty(deprecatedKey);\r\n if (deprecatedValue != null) {\r\n getProps().setProperty(n, deprecatedValue);\r\n }\r\n }\r\n }\r\n }\r\n{noformat}\r\n\r\nThanks!","from":"developer"},{"body":"[~snemeth] thanks for the review, items fixed.\r\n\r\n[~stevel@apache.org] could you review the current patch? Please let me know if you would need any further help understanding the context of the issue.","from":"developer"},{"body":"Thanks [~zsiegl] for the updated patch!\r\nLGTM, +1 (non-binding)","from":"developer"},{"body":"I'm going to wait to see what others say, which given its a long weekend in the US, unlikely until next week\r\n\r\nConfiguration is such as ubiquitous class that its on the list of things to be very careful about changing, even when the changes seem low risk, which is why we need that extra oversight\r\n\r\nlooking @ code; \r\n\r\n* I'd like to see a patch which doesn't add IDE-helped changes, as that complicates merging and cherry-picking, so fixes to unused imports, line endings, spacing need to be left out. Sorry.\r\n* in TestConfigurationDeprecation, use try-with-resources for writing the file.\r\n\r\n","from":"developer"},{"body":"[~stevel@apache.org] Thank you for the review. Changes done, new patch on the way. Try-with-resources in the test would require a larger refactor of the test case, I suggest a different commit for that.","from":"developer"},{"body":"[~stevel@apache.org] do you know of anyone who would be interested in this issue?\r\n[~xiaochen] Would you be interested to have a look at this problem? \r\n\r\nThanks :)","from":"developer"},{"body":"[~jlowe] [~daryn] could someone have a look at this issue please?","from":"developer"},{"body":"There's definitely something wrong here.  I was playing with the {{Configuration}} class like this:\r\n{code:java}\r\n// OLD_KEY is set to \"foo\" in the XML\r\nconf.get(OLD_KEY); // returns \"foo\"\r\nConfiguration.addDeprecations(OLD_KEY, NEW_KEY);\r\nconf.get(OLD_KEY); // returns null\r\n{code}\r\nThat doesn't seem right. Adding a deprecation looks like it breaks getting the value of the deprecated key. On top of that, if you ommit the 2nd line, it works correctly:\r\n{code:java}\r\n// OLD_KEY is set to \"foo\" in the XML\r\nConfiguration.addDeprecations(OLD_KEY, NEW_KEY);\r\nconf.get(OLD_KEY); // returns \"foo\"\r\n{code}\r\nThis means that getting the {{OLD_KEY}} and then deprecating it changes the value, which isn't the right behavior. \r\n\r\nAs for the right way to fix this, and make sure we're not breaking anything depending on this odd behavior, I'm not sure. It would be good to get some more opinions.","from":"developer"},{"body":"[~rkanter] did you experience this behaviour before or after applying the patch?\r\nThis is what the patch should fix. Also before the patch: \r\n\r\n{code:java}\r\n// OLD_KEY is set to \"foo\" in the XML\r\nConfiguration.addDeprecations(OLD_KEY, NEW_KEY);\r\nconf.get(OLD_KEY); // returns \"foo\"\r\n{code}\r\n\r\nWhile:\r\n{code:java}\r\n// OLD_KEY is set to \"foo\" in the XML\r\nconf.get(SOMETHING_ELSE);\r\nConfiguration.addDeprecations(OLD_KEY, NEW_KEY);\r\nconf.get(OLD_KEY); // returns null\r\n{code}\r\n\r\n","from":"developer"},{"body":"Sorry, yeah; that was before the patch.  I was just investigating what the current behavior is.  As [~stevel@apache.org] said, we have to be very careful about changing this class.  The code seems okay to me though, but let's see if anyone else has thoughts about the consequences of changing this behavior.","from":"developer"},{"body":"So to make things a bit clear about the change:\r\nBefore the patch:\r\n{code:java}\r\nprivate String[] handleDeprecation(DeprecationContext deprecations,\r\n String name) {\r\n if (null != name) {\r\n name = name.trim();\r\n }\r\n // Initialize the return value with requested name\r\n String[] names = new String[]{name};\r\n // Deprecated keys are logged once and an updated names are returned\r\n DeprecatedKeyInfo keyInfo = deprecations.getDeprecatedKeyMap().get(name);\r\n if (keyInfo != null) {\r\n if (!keyInfo.getAndSetAccessed()) {\r\n logDeprecation(keyInfo.getWarningMessage(name));\r\n }\r\n // Override return value for deprecated keys\r\n names = keyInfo.newKeys;\r\n }\r\n\r\n // **** WE RETURN EARLY HERE IF THERE ARE NO OVERLAYS ***\r\n\r\n // If there are no overlay values we can return early\r\n Properties overlayProperties = getOverlay();\r\n if (overlayProperties.isEmpty()) {\r\n return names;\r\n }\r\n // Update properties and overlays with reverse lookup values\r\n for (String n : names) {\r\n String deprecatedKey = deprecations.getReverseDeprecatedKeyMap().get(n);\r\n if (deprecatedKey != null && !overlayProperties.containsKey(n)) {\r\n String deprecatedValue = overlayProperties.getProperty(deprecatedKey);\r\n if (deprecatedValue != null) {\r\n // **** STILL AT THIS LINE WE ARE UPDATING PROPS *NOT* OVERLAY ***\r\n getProps().setProperty(n, deprecatedValue);\r\n overlayProperties.setProperty(n, deprecatedValue);\r\n }\r\n }\r\n }\r\n return names;\r\n }\r\n{code}\r\n\r\nSo I have extracted the latter loop to run for all props not just overlays as it should work.\r\n\r\nTo address concerns of [~jlowe] as we talked offline:\r\n- You said that if you set deprecated and new keys as well, this patch would change the return value. To address this fear I have created a test case for this:\r\n\r\n{code:java}\r\n@Test\r\n public void testBothPropertiesAreSet() throws Exception {\r\n // SETUP\r\n final String oldZkAddressKey = \"yarn.resourcemanager.zk-address\";\r\n final String newZkAddressKey = CommonConfigurationKeys.ZK_ADDRESS;\r\n final String oldZkAddressValue = \"oldZkAddress\";\r\n final String newZkAddressValue = \"newZkAddress\";\r\n\r\n try{\r\n out = new BufferedWriter(new FileWriter(CONFIG4));\r\n startConfig();\r\n appendProperty(oldZkAddressKey, oldZkAddressValue);\r\n appendProperty(newZkAddressKey, newZkAddressValue);\r\n endConfig();\r\n\r\n Path fileResource = new Path(CONFIG4);\r\n conf.addResource(fileResource);\r\n } finally {\r\n out.close();\r\n }\r\n\r\n // ACT\r\n conf.get(oldZkAddressKey);\r\n Configuration.addDeprecations(new Configuration.DeprecationDelta[] {\r\n new Configuration.DeprecationDelta(oldZkAddressKey, newZkAddressKey)});\r\n\r\n // ASSERT\r\n assertEquals(\"New property should overwrite deprecated one\",\r\n newZkAddressValue, conf.get(oldZkAddressKey));\r\n assertEquals(\"Property should be accessible through new key\",\r\n newZkAddressValue, conf.get(newZkAddressKey));\r\n }\r\n{code}\r\n\r\nThis tests setting both properties and returning the same values before and after the patch as well.\r\nI did not add this to the patch, as this has been already covered in the previous test cases, it is juts a bit more readable this way, as a separate test. \r\n\r\nTo address fears of [~rkanter]:\r\n- You said that some may depend on the bugous functionality. That is normally a good idea to stay backward compatible, however why would you depend on getting NULL instead of the config value you needed? This is a bug. Don'd depend on bugs.\r\n\r\nIs there anything else I can do for this to be accepted? I'd be happy to.\r\n","from":"developer"},{"body":"{quote}Don't depend on bugs.\r\n{quote}\r\nI agree, though we've had issues in the past.\r\n\r\n \r\n\r\nAnyway, I think we we should move forward on this. It seems pretty clear to me that the current behavior doesn't make sense and we should fix it. I'll wait about a week to give anyone else a chance to speak up about this, and then I'll commit the patch. Otherwise, I think we'll be waiting forever while everyone hopes someone else will do the commit.\r\n\r\nOne nitpick with the patch: {{updatePropertiesWIthDeprecatedKeys}} has an extra capitalization in {{WIth}}. It should be {{updatePropertiesWithDeprecatedKeys}}.","from":"developer"},{"body":"[~rkanter] thanks for the review, and all the care about this. Fixed the typo. New patch uploaded.","from":"developer"},{"body":"Thanks [~zsiegl] and everyone else for reviews.  Committed to trunk!","from":"developer"},{"body":"Looks like we missed a trivial problem where one of the updated tests leaves behind a file, which causes test runs to complain about a wrong license.  I've filed HADOOP-15853 to fix that.","from":"developer"}],"created":"2018-08-31T08:40:37.000+0000","description":"Hadoop Common contains a widely used Configuration class.\r\n This class can handle deprecations of properties, e.g. if property 'A' gets deprecated with an alternative property key 'B', users can access property values with keys 'A' and 'B'.\r\n Unfortunately, this does not work in one case.\r\n When a config file is specified (for instance, XML) and a property is read with the config.get() method, the config is loaded from the file at this time. \r\n If the deprecation mapping is not yet specified by the time any config value is retrieved and the XML config refers to a deprecated key, then the deprecation mapping specified, the config value cannot be retrieved neither with the deprecated nor with the new key.\r\n The attached patch contains a testcase that reproduces this wrong behavior.\r\n\r\nHere are the steps outlined what the testcase does:\r\n 1. Creates an XML config file with a deprecated property\r\n 2. Adds the config to the Configuration object\r\n 3. Retrieves the config with its deprecated key (it does not really matter which property the user gets, could be any)\r\n 4. Specifies the deprecation rules including the one defined in the config\r\n 5. Prints and asserts the property retrieved from the config with both the deprecated and the new property keys.\r\n\r\nFor reference, here is the log of one execution that actually shows what the issue is:\r\n{noformat}\r\nLoaded items: 1\r\nLooked up property value with name hadoop.zk.address: null\r\nLooked up property value with name yarn.resourcemanager.zk-address: dummyZkAddress\r\nContents of config file: [, , yarn.resourcemanager.zk-addressdummyZkAddress, ]\r\nLooked up property value with name hadoop.zk.address: null\r\n2018-08-31 10:10:06,484 INFO Configuration.deprecation (Configuration.java:logDeprecation(1397)) - yarn.resourcemanager.zk-address is deprecated. Instead, use hadoop.zk.address\r\nLooked up property value with name hadoop.zk.address: null\r\nLooked up property value with name hadoop.zk.address: null\r\n\r\njava.lang.AssertionError: \r\nExpected :dummyZkAddress\r\nActual :null\r\n{noformat}\r\n*As it's visible from the output and the code, the issue is really that if the config is retrieved either with the deprecated or the new value, Configuration both wants to serve the value with the new key.*\r\n *If the mapping is not specified before any retrieval happened, the value is only stored under the deprecated key but not the new key.*","issue_id":"13182310","key":"HADOOP-15708","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2018-10-11T01:52:13.000+0000","role":"fixed_distractor","summary":"Reading values from Configuration before adding deprecations make it impossible to read value with deprecated key"} {"case_id":"12373557","cluster":"DISTRACTOR-HADOOP-1596","comments":[{"body":"Hudson builds will fail due to this, so folks should refrain from marking thinngs \"Patch Available\" until this is resolved.","created":"2007-07-11T22:45:17.523+0000"},{"body":"Perhaps this is related to HADOOP-1587","created":"2007-07-11T23:08:17.611+0000"},{"body":"This is the result of changing job id \t HADOOP-1473\nThe initial exception in TestSymLink is hard to find but it actually is this:\n\njava.lang.NumberFormatException: For input string: \"m\"\n\tat java.lang.NumberFormatException.forInputString(NumberFormatException.java:48)\n\tat java.lang.Integer.parseInt(Integer.java:447)\n\tat java.lang.Integer.parseInt(Integer.java:497)\n\tat org.apache.hadoop.streaming.StreamUtil.getTaskInfo(StreamUtil.java:451)\n\tat org.apache.hadoop.streaming.PipeMapRed.setStreamJobDetails(PipeMapRed.java:190)\n\tat org.apache.hadoop.streaming.PipeMapRed.configure(PipeMapRed.java:132)\n\tat org.apache.hadoop.streaming.PipeMapper.configure(PipeMapper.java:61)\n\tat org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:58)\n\tat org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:82)\n\tat org.apache.hadoop.mapred.MapRunner.configure(MapRunner.java:32)\n\tat org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:58)\n\tat org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:82)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:185)\n\tat org.apache.hadoop.mapred.TaskTracker$Child.main(TaskTracker.java:1763)\n\nAfter that the task remains not configured. And we see other exceptions NPE or \"File does not exist\".\nSo the problem is that streaming was not changed to correctly parse new task ids.\n\nA more general problem with streaming is that it just dumps exceptions into stderr and proceeds like nothing happened.\n","created":"2007-07-12T01:57:30.028+0000"},{"body":"Argh. I didn't see that Konstatin had found it until I had and went to upload the patch. Anyways, here is the patch. I not only fixed parsing the jobids, I also made it throw out of configure if anything went wrong.","created":"2007-07-12T07:43:45.891+0000"},{"body":"+1","created":"2007-07-12T12:57:25.253+0000"},{"body":"I just committed this.","created":"2007-07-12T14:49:26.598+0000"}],"conversations":[{"body":"TestSymLink started failing sometime today.","from":"reporter","subject":"TestSymLink is failing"},{"body":"Hudson builds will fail due to this, so folks should refrain from marking thinngs \"Patch Available\" until this is resolved.","from":"developer"},{"body":"Perhaps this is related to HADOOP-1587","from":"developer"},{"body":"This is the result of changing job id \t HADOOP-1473\nThe initial exception in TestSymLink is hard to find but it actually is this:\n\njava.lang.NumberFormatException: For input string: \"m\"\n\tat java.lang.NumberFormatException.forInputString(NumberFormatException.java:48)\n\tat java.lang.Integer.parseInt(Integer.java:447)\n\tat java.lang.Integer.parseInt(Integer.java:497)\n\tat org.apache.hadoop.streaming.StreamUtil.getTaskInfo(StreamUtil.java:451)\n\tat org.apache.hadoop.streaming.PipeMapRed.setStreamJobDetails(PipeMapRed.java:190)\n\tat org.apache.hadoop.streaming.PipeMapRed.configure(PipeMapRed.java:132)\n\tat org.apache.hadoop.streaming.PipeMapper.configure(PipeMapper.java:61)\n\tat org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:58)\n\tat org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:82)\n\tat org.apache.hadoop.mapred.MapRunner.configure(MapRunner.java:32)\n\tat org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:58)\n\tat org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:82)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:185)\n\tat org.apache.hadoop.mapred.TaskTracker$Child.main(TaskTracker.java:1763)\n\nAfter that the task remains not configured. And we see other exceptions NPE or \"File does not exist\".\nSo the problem is that streaming was not changed to correctly parse new task ids.\n\nA more general problem with streaming is that it just dumps exceptions into stderr and proceeds like nothing happened.\n","from":"developer"},{"body":"Argh. I didn't see that Konstatin had found it until I had and went to upload the patch. Anyways, here is the patch. I not only fixed parsing the jobids, I also made it throw out of configure if anything went wrong.","from":"developer"},{"body":"+1","from":"developer"},{"body":"I just committed this.","from":"developer"}],"created":"2007-07-11T22:40:04.000+0000","description":"TestSymLink started failing sometime today.","issue_id":"12373557","key":"HADOOP-1596","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2007-07-12T14:49:26.000+0000","role":"fixed_distractor","summary":"TestSymLink is failing"} {"case_id":"12328376","cluster":"DISTRACTOR-HADOOP-16","comments":[{"body":"I'm seeing a similar problem with the latest code. I have a task that takes a lot of input files (currently ~1400, each several gigabytes each). The amount of time spent hitting hasTaskWithHit() for each task is exhorbitant, easily causing a timeout. This code does all of the work up front - if you had, say, 1400 tasks, and 2 taskrunners, you could easily dispatch the first 2xsimultaneous running tasks first, then go back and spend time calculating the rest of the matchups. Certainly, a lot of the steps in the job setup after listFiles() could be potentially slow for certain problem sizes.","created":"2006-02-10T17:19:51.000+0000"},{"body":"Everyone should probably be made aware of the strange behavior we see during indexing, at least for a relatively large number of large segments (topN=500K, depth=20) with a relatively large crawldb (50M URLs). Note that this was all performed with ipc.client.timeout set to 30 minutes.\n\nAfter launching the indexing job, the web UI shows all of the TaskTrackers, but the numbers in the \"Secs since heartbeat\" column just keep increasing. This goes on for about 10 minutes until the JobTracker finally loses all of them (and the tasks they were working on), as is shown in its log:\n\n060210 224115 parsing file:/home/crawler/nutch/conf/nutch-site.xml\n060210 225151 Lost tracker 'tracker_37064'\n060210 225151 Task 'task_m_4ftk58' has been lost.\n060210 225151 Task 'task_m_6ww2ri' has been lost.\n\n...(snip)...\n\n060210 225151 Task 'task_r_y6d190' has been lost.\n060210 225151 Lost tracker 'tracker_92921'\n060210 225151 Task 'task_m_9p24at' has been lost.\n\n...(etc)...\n\nAt this point, the web UI is still up, the job shows 0% complete, and the TaskTrackers table is empty. It goes on for an hour or so like this, during which any rational person would probably want to kill the job and start over.\n\nDon't do this! Keep the faith!!!\n\nAbout an hour later, the JobTracker magically reestablishes its connection to the TaskTrackers (which now have new names), as is shown in its log:\n\n060210 225151 Task 'task_r_yj3y3o' has been lost.\n060210 235403 Adding task 'task_m_k9u9a8' to set for tracker 'tracker_85874'\n060210 235404 Adding task 'task_m_pijt4q' to set for tracker 'tracker_61888'\n\n...(etc)...\n\nThe web UI also shows that the TaskTrackers are back (with their new names).\n\nThere's nothing in the TaskTracker logs during the initial 10 minutes, then a bunch of exiting and closing messages, until finally the TaskTrackers start \"Reinitializing local state\":\n\n060210 225403 Stopping server on 50050\n060210 230102 Server handler 4 on 50050: exiting\n\n...(snip)...\n\n060210 230105 Server handler 7 on 50050: exiting\n060210 232024 Server listener on port 50050: exiting\n060210 232403 Stopping server on 50040\n060210 234902 Server listener on port 50040: exiting\n060210 234925 Server connection on port 50040 from 192.168.1.5: exiting\n\n...(snip)...\n\n060210 235009 Server connection on port 50040 from 192.168.1.10: exiting\n060210 235013 Client connection to 192.168.1.4:50040: closing\n060210 235014 Client connection to 192.168.1.7:50040: closing\n060210 235015 Server connection on port 50040 from 192.168.1.7: exiting\n060210 235016 Server handler 0 on 50040: exiting\n\n...(snip)...\n\n060210 235024 Server handler 2 on 50040: exiting\n060210 235403 Reinitializing local state\n060210 235403 Server listener on port 50050: starting\n060210 235403 Server handler 0 on 50050: starting\n\n...(etc)...\n\nDuring the time that the TaskTrackers are lost, neither the master nor the slave machines seem to be using much of the CPU or RAM, and the DataNode logs are quiet. I suppose that it's probably I/O bound on the master machine, but even that seems mysterious to me. It would seem particularly inappropriate for the JobTracker to punt the TaskTrackers because the master was too busy to listen for their heartbeats.\n\nAt any rate, once the TaskTrackers go through the \"Reinitializing local state\" thing, the indexing job seems to proceed normally, and it eventually completes with no errors.\n","created":"2006-02-12T04:20:17.000+0000"},{"body":"Looking further into my concern about hasTaskWithCacheHit overhead (called from obtainNewMapTask in JobInProgress).... it looks like obtainNewMapTask actually calls hasTaskWithCacheHit once for each map job, but always chooses the first choice - why even call the relatively expensive hasTaskWithCacheHit if you've already chosen a value for cacheTarget? Likewise, once you've filled in both cacheTarget and hasTarget, why not break out of that for loop entirely?\n\nAm I following the code flow incorrectly, or is obtainNewMapTask essentially making maps^2 RPC calls in response to a single incoming RPC call? Even if the jobtracker and namenode are running on the same host, that could easily explain the sensitivity to timeouts in the current code.\n","created":"2006-02-13T16:22:46.000+0000"},{"body":"Bryan: I agree, this looks like a serious bug. The TaskTracker should minimize the calls it makes to the NameNode. Ideally it should only make a single call per job on each input split.","created":"2006-02-14T04:17:00.000+0000"},{"body":"\n Here's a patch for the problem. I changed two things.\n\n 1) The comment is right, there's no need to iterate through all choices in JobTracker\nafter we find a TaskInProgress to take the task.\n\n 2) Within a TaskInProgress (TIP), which tracks an individual split within a Job, we cache\nthe results of the getHints() call to the Distributed File System.\n","created":"2006-02-14T18:53:04.000+0000"},{"body":"\n Sorry, my blurb above was a little unclear.\n\n I should have said:\n\n 1) Bryan's comment is right, we don't need to iterate through the whole\nlist in JobTracker's obtainNewMapTask call. We now just do it until we\nfind a good cacheTarget or stdTarget value.\n\n 2) A TIP object tracks each individual split in the Job. We cache\nthe data at each TIP. This will be handy in case the TIP has to\nbe re-executed due to machine failure. \n\n I don't mind caching the hints aggressively, because it's \njust task-placement we're after. If the hint is wrong (which only happens \nin case of machine failure), we might send the task to a suboptimal \nmachine. No big deal.\n\n","created":"2006-02-14T18:59:18.000+0000"},{"body":"\n Patch comitted.","created":"2006-02-15T04:08:59.000+0000"},{"body":"I'm not sure this patch does quite what's desired.\n\nIt looks like before it would try as hard as possible to find a task with a cache hit, then fail to running either just the first executable task. Now, it will search for the first task that's either a cache hit *or* executable..... In principle, wouldn't you only want to break out of the loop for finding a cache hit? This, of course, still causes timeouts - is there any way to actually precompute the list of cache hits, so that it's just a matter of picking a task from a priority queue?","created":"2006-02-17T03:49:54.000+0000"},{"body":"Of course, given the existing code, maybe it's worth trying a little harder - iterate through the list as before, but keep track of the time that's been taken, and give up if it gets to 1/2 of the RPC timeout, or something of the like.","created":"2006-02-17T03:50:44.000+0000"},{"body":"But, you're right, Bryan, I think this is still not optimal. It should certainly check to see if more map tasks have local input data for the calling node before it gives up. Ideally it should not be iterating through all map tasks either, but rather create a mapping of node -> mapTask* for tasks that have local data for a node. We could keep a queue of tasks whose entries in this table have not yet been computed. When we fail to find a task with local data for a node, then we can pop a few (10?) entries off the queue and enter them in the map, and if there are still no matches, just give the node a task with remote input data. This way we'd avoid ever doing too much work in a single call, and ever iterating over all tasks.","created":"2006-02-17T05:59:32.000+0000"},{"body":"I have another idea.\n\nFirst, switch the delegation of responsiblity for job assignment. Right now, it happens in the JobTracker instance, in response to an obtainNewMapTask call. This scales very poorly. In particular, it causes the RPC-timeouts if you do any sort of serious work in obtainNewMapTask. There's another bug, just reported as HADOOP-43, which occurs if you spend too long in a RPC call, and also described in the second comment, above.\n\nSo, instead of having the JobTracker do this reactively, either:\n1) Precompute - probably most scalably done by starting a mini-job, which just computes the list of who has precached data from a given FileSplit.\n2) Compute on demand - as a TaskTracker job. This could work by some protocol of offering the TaskTracker a set of possible jobs it could do work on, letting it pick the ones it thinks are best, and return the remainders for assignment. This, of course, would only work well for instances where tasks >> tasktracker instances.\n\nI looked at implementing something like 1, but decided I think 2 is a much better option. 1 would put a lot more instantaneous demand on the namenode. Plus, once you've finished precomputing the best nodes, if nodes come or go you don't really have a solution. 2 seems to distribute both the work and some of the demand, and it makes it possible for the cluster to grow or shrink dramatically without failing to take advantage of the local storage available at each node. Unfortunately, without any pre-work, it's possible that, doing option 2, you'd pick bad subsets of work to distribute to each node, and get no local I/O improvement at all.\n\nI'd really like to see something done, perferably soon. With dozens of nodes, and hundreds of gbs of data in my current problem set, it's very nearly impossible to get the current code to make progress, without killing tasktrackers (some with lots of work units already completed). I can do some of the coding, if there's agreement for what direction to push.","created":"2006-02-21T15:01:05.000+0000"},{"body":"I'm re-opening this & assigning it to Mike.\n\nI think we can fix this by changing the TaskTracker to incrementally calculate things, as I alluded above. The TaskTracker should avoid ever doing anything that might take very long. RPC timeouts to the TaskTracker are very bad and must be avoided.\n\nWe can maintain a mapping of taskTracker->split, initially empty, and a queue of splits, initially filled with all of the splits. (If basic split-generation is too expensive, since it calls file-length on each input file,, then we can eventually change the split API into a generator, which incrementally enumerates splits and use that in place of this queue.) When a request for a task arrives from a tasktracker we can first examine the table. If any splits are present for the calling tracker, then we return one and remove it from the table. Otherwise we pop a constant number of splits (10?) off of the queue, enumerate the tasktrackers that host each split by calling the namenode, and add entries to the table. Then we consult the table again. In most cases (many more splits than tasktrackers) we should identify suitable splits. If we fail then we assign a randomly selected, non-local split to the tasktracker. Make sense?","created":"2006-02-23T05:26:33.000+0000"},{"body":"Eric Baldeschwieler wrote:\n> [...] why not just dedicate a thread \n> to planning and then load a complete plan? That can produce more \n> optimal placement and a simpler to understand initialization sequence.\n\nA separate thread in the JobTracker? That could be a good approach. We'd have a queue of submitted but as-yet unplanned jobs. The thread can then pop a job off the queue, compute its splits, then start populating a tasktracker->split table. When tasktrackers poll for work they can consult this table, potentially while the thread is still populating it.\n\nI'm hesitant to move this out of the JobTracker into the TaskTracker, since that introduces complexity. But a single thread in the JobTracker should be simple to add and should mostly solve this. +1","created":"2006-02-23T06:33:39.000+0000"},{"body":"My only concern is that this doesn't scale as well for really huge jobs, but, if implemented as you suggest (pre-filling, hand out non-matching jobs when you don't have matched jobs) seems like a reasonable trade off. \n\nI'd still think it'd be nice to better handle the case of newly-discovered tasktrackers, but that's probably much more minor than other scaling/responsiveness issues. Maybe a future enhancement for the JobTracker thread that does this work.","created":"2006-02-23T07:03:31.000+0000"},{"body":"Having a separate thread in JobTracker can even pre-compute a data locality map that will enable optimal task assignment to task trackers.\nHere is the basic idea. As the thread computes the splits for a job, it can determine which nodes have local data for each split. The thread can populate \na map mapping a node to a list of the splits whose data are on the node. When a tasktracker polls for a task, the JobTracker can look up the map by \nthe node name of the task tracker to see whether there are any un-assigned splits in the corresponding list. If yes, assign one of them to the task tracker. \nOtherwise, randomlly choose a split for the task tracker. ","created":"2006-02-24T03:12:06.000+0000"},{"body":"\n I created a separate thread on the JobTracker that handles file-splitting and\ngathering info about split-caches. The job is placed into the PREP state\nuntil it is processed by this JobTracker thread. After processing, it moves\ninto the RUNNING state. \n\n This should allow the hard split-work to continue but still allow the JobTracker\nto process heartbeats. For now, the thread lives at the JobTracker but perhaps\nsomeday we'll move it to a TaskTracker thread.\n\n I also fixed the alg for choosing a task-allocation given a tasktracker.\n(Whether cached, non-cached, or speculative.)\n\n Let me know if this fixes some of the problems you've been seeing. \n\n","created":"2006-02-28T08:46:46.000+0000"},{"body":"I applied Mike's patch, with a few small changes, fixing an NPE and ArrayOutOfBounds in dfs.\n\nThere's still a bug where, with mapred.tasktracker.tasks.maximum=2, only a single map and a single reduce are running on each node. I believe the intent is that there should be up to two of each related to a single job, so that the reduce tasks can be copying data while the maps are still running.","created":"2006-03-03T08:09:22.000+0000"},{"body":"This is great, and finally fixes my issues with a large job that would never start.\n\nHowever, the way things are in this patch (and the current code), the job doesn't get started until the background thread finishes computing all of the cache hints. This takes far too long - it took 15 minutes on a recent run. During that time, of course, no other work was getting done. How about moving the cachedHints-filling-loop to the end of initTasks(), and go ahead and set the job to RUNNING and \"tasksInited=true\" in the meantime?\n\nDoing this locally lets work commence immediately, while the cache hints continue to get filled in for future task allocations.","created":"2006-03-08T08:49:48.000+0000"}],"conversations":[{"body":"We've been using Nutch 0.8 (MapReduce) to perform some internet crawling. Things seemed to be going well until...\n\n060129 222409 Lost tracker 'tracker_56288'\n060129 222409 Task 'task_m_10gs5f' has been lost.\n060129 222409 Task 'task_m_10qhzr' has been lost.\n ........\n ........\n060129 222409 Task 'task_r_zggbwu' has been lost.\n060129 222409 Task 'task_r_zh8dao' has been lost.\n060129 222455 Server handler 8 on 8010 caught: java.net.SocketException: Socket closed\njava.net.SocketException: Socket closed\n at java.net.SocketOutputStream.socketWrite(SocketOutputStream.java:99)\n at java.net.SocketOutputStream.write(SocketOutputStream.java:136)\n at java.io.BufferedOutputStream.flushBuffer(BufferedOutputStream.java:65)\n at java.io.BufferedOutputStream.flush(BufferedOutputStream.java:123)\n at java.io.DataOutputStream.flush(DataOutputStream.java:106)\n at org.apache.nutch.ipc.Server$Handler.run(Server.java:216)\n060129 222455 Adding task 'task_m_cia5po' to set for tracker 'tracker_56288'\n060129 223711 Adding task 'task_m_ffv59i' to set for tracker 'tracker_25647'\n\nI'm hoping that someone could explain why task_m_cia5po got added to tracker_56288 after this tracker was lost.\n\nThe Crawl .main process died with the following output:\n\n060129 221129 Indexer: adding segment: /user/crawler/crawl-20060129091444/segments/20060129200246\nException in thread \"main\" java.io.IOException: timed out waiting for response\n at org.apache.nutch.ipc.Client.call(Client.java:296)\n at org.apache.nutch.ipc.RPC$Invoker.invoke(RPC.java:127)\n at $Proxy1.submitJob(Unknown Source)\n at org.apache.nutch.mapred.JobClient.submitJob(JobClient.java:259)\n at org.apache.nutch.mapred.JobClient.runJob(JobClient.java:288)\n at org.apache.nutch.indexer.Indexer.index(Indexer.java:263)\n at org.apache.nutch.crawl.Crawl.main(Crawl.java:127)\n\nHowever, it definitely seems as if the JobTracker is still waiting for the job to finish (no failed jobs).\n\nDoug Cutting's response:\nThe bug here is that the RPC call times out while the map task is computing splits. The fix is that the job tracker should not compute splits until after it has returned from the submitJob RPC. Please submit a bug in Jira to help remind us to fix this.\n","from":"reporter","subject":"RPC call times out while indexing map task is computing splits"},{"body":"I'm seeing a similar problem with the latest code. I have a task that takes a lot of input files (currently ~1400, each several gigabytes each). The amount of time spent hitting hasTaskWithHit() for each task is exhorbitant, easily causing a timeout. This code does all of the work up front - if you had, say, 1400 tasks, and 2 taskrunners, you could easily dispatch the first 2xsimultaneous running tasks first, then go back and spend time calculating the rest of the matchups. Certainly, a lot of the steps in the job setup after listFiles() could be potentially slow for certain problem sizes.","from":"developer"},{"body":"Everyone should probably be made aware of the strange behavior we see during indexing, at least for a relatively large number of large segments (topN=500K, depth=20) with a relatively large crawldb (50M URLs). Note that this was all performed with ipc.client.timeout set to 30 minutes.\n\nAfter launching the indexing job, the web UI shows all of the TaskTrackers, but the numbers in the \"Secs since heartbeat\" column just keep increasing. This goes on for about 10 minutes until the JobTracker finally loses all of them (and the tasks they were working on), as is shown in its log:\n\n060210 224115 parsing file:/home/crawler/nutch/conf/nutch-site.xml\n060210 225151 Lost tracker 'tracker_37064'\n060210 225151 Task 'task_m_4ftk58' has been lost.\n060210 225151 Task 'task_m_6ww2ri' has been lost.\n\n...(snip)...\n\n060210 225151 Task 'task_r_y6d190' has been lost.\n060210 225151 Lost tracker 'tracker_92921'\n060210 225151 Task 'task_m_9p24at' has been lost.\n\n...(etc)...\n\nAt this point, the web UI is still up, the job shows 0% complete, and the TaskTrackers table is empty. It goes on for an hour or so like this, during which any rational person would probably want to kill the job and start over.\n\nDon't do this! Keep the faith!!!\n\nAbout an hour later, the JobTracker magically reestablishes its connection to the TaskTrackers (which now have new names), as is shown in its log:\n\n060210 225151 Task 'task_r_yj3y3o' has been lost.\n060210 235403 Adding task 'task_m_k9u9a8' to set for tracker 'tracker_85874'\n060210 235404 Adding task 'task_m_pijt4q' to set for tracker 'tracker_61888'\n\n...(etc)...\n\nThe web UI also shows that the TaskTrackers are back (with their new names).\n\nThere's nothing in the TaskTracker logs during the initial 10 minutes, then a bunch of exiting and closing messages, until finally the TaskTrackers start \"Reinitializing local state\":\n\n060210 225403 Stopping server on 50050\n060210 230102 Server handler 4 on 50050: exiting\n\n...(snip)...\n\n060210 230105 Server handler 7 on 50050: exiting\n060210 232024 Server listener on port 50050: exiting\n060210 232403 Stopping server on 50040\n060210 234902 Server listener on port 50040: exiting\n060210 234925 Server connection on port 50040 from 192.168.1.5: exiting\n\n...(snip)...\n\n060210 235009 Server connection on port 50040 from 192.168.1.10: exiting\n060210 235013 Client connection to 192.168.1.4:50040: closing\n060210 235014 Client connection to 192.168.1.7:50040: closing\n060210 235015 Server connection on port 50040 from 192.168.1.7: exiting\n060210 235016 Server handler 0 on 50040: exiting\n\n...(snip)...\n\n060210 235024 Server handler 2 on 50040: exiting\n060210 235403 Reinitializing local state\n060210 235403 Server listener on port 50050: starting\n060210 235403 Server handler 0 on 50050: starting\n\n...(etc)...\n\nDuring the time that the TaskTrackers are lost, neither the master nor the slave machines seem to be using much of the CPU or RAM, and the DataNode logs are quiet. I suppose that it's probably I/O bound on the master machine, but even that seems mysterious to me. It would seem particularly inappropriate for the JobTracker to punt the TaskTrackers because the master was too busy to listen for their heartbeats.\n\nAt any rate, once the TaskTrackers go through the \"Reinitializing local state\" thing, the indexing job seems to proceed normally, and it eventually completes with no errors.\n","from":"developer"},{"body":"Looking further into my concern about hasTaskWithCacheHit overhead (called from obtainNewMapTask in JobInProgress).... it looks like obtainNewMapTask actually calls hasTaskWithCacheHit once for each map job, but always chooses the first choice - why even call the relatively expensive hasTaskWithCacheHit if you've already chosen a value for cacheTarget? Likewise, once you've filled in both cacheTarget and hasTarget, why not break out of that for loop entirely?\n\nAm I following the code flow incorrectly, or is obtainNewMapTask essentially making maps^2 RPC calls in response to a single incoming RPC call? Even if the jobtracker and namenode are running on the same host, that could easily explain the sensitivity to timeouts in the current code.\n","from":"developer"},{"body":"Bryan: I agree, this looks like a serious bug. The TaskTracker should minimize the calls it makes to the NameNode. Ideally it should only make a single call per job on each input split.","from":"developer"},{"body":"\n Here's a patch for the problem. I changed two things.\n\n 1) The comment is right, there's no need to iterate through all choices in JobTracker\nafter we find a TaskInProgress to take the task.\n\n 2) Within a TaskInProgress (TIP), which tracks an individual split within a Job, we cache\nthe results of the getHints() call to the Distributed File System.\n","from":"developer"},{"body":"\n Sorry, my blurb above was a little unclear.\n\n I should have said:\n\n 1) Bryan's comment is right, we don't need to iterate through the whole\nlist in JobTracker's obtainNewMapTask call. We now just do it until we\nfind a good cacheTarget or stdTarget value.\n\n 2) A TIP object tracks each individual split in the Job. We cache\nthe data at each TIP. This will be handy in case the TIP has to\nbe re-executed due to machine failure. \n\n I don't mind caching the hints aggressively, because it's \njust task-placement we're after. If the hint is wrong (which only happens \nin case of machine failure), we might send the task to a suboptimal \nmachine. No big deal.\n\n","from":"developer"},{"body":"\n Patch comitted.","from":"developer"},{"body":"I'm not sure this patch does quite what's desired.\n\nIt looks like before it would try as hard as possible to find a task with a cache hit, then fail to running either just the first executable task. Now, it will search for the first task that's either a cache hit *or* executable..... In principle, wouldn't you only want to break out of the loop for finding a cache hit? This, of course, still causes timeouts - is there any way to actually precompute the list of cache hits, so that it's just a matter of picking a task from a priority queue?","from":"developer"},{"body":"Of course, given the existing code, maybe it's worth trying a little harder - iterate through the list as before, but keep track of the time that's been taken, and give up if it gets to 1/2 of the RPC timeout, or something of the like.","from":"developer"},{"body":"But, you're right, Bryan, I think this is still not optimal. It should certainly check to see if more map tasks have local input data for the calling node before it gives up. Ideally it should not be iterating through all map tasks either, but rather create a mapping of node -> mapTask* for tasks that have local data for a node. We could keep a queue of tasks whose entries in this table have not yet been computed. When we fail to find a task with local data for a node, then we can pop a few (10?) entries off the queue and enter them in the map, and if there are still no matches, just give the node a task with remote input data. This way we'd avoid ever doing too much work in a single call, and ever iterating over all tasks.","from":"developer"},{"body":"I have another idea.\n\nFirst, switch the delegation of responsiblity for job assignment. Right now, it happens in the JobTracker instance, in response to an obtainNewMapTask call. This scales very poorly. In particular, it causes the RPC-timeouts if you do any sort of serious work in obtainNewMapTask. There's another bug, just reported as HADOOP-43, which occurs if you spend too long in a RPC call, and also described in the second comment, above.\n\nSo, instead of having the JobTracker do this reactively, either:\n1) Precompute - probably most scalably done by starting a mini-job, which just computes the list of who has precached data from a given FileSplit.\n2) Compute on demand - as a TaskTracker job. This could work by some protocol of offering the TaskTracker a set of possible jobs it could do work on, letting it pick the ones it thinks are best, and return the remainders for assignment. This, of course, would only work well for instances where tasks >> tasktracker instances.\n\nI looked at implementing something like 1, but decided I think 2 is a much better option. 1 would put a lot more instantaneous demand on the namenode. Plus, once you've finished precomputing the best nodes, if nodes come or go you don't really have a solution. 2 seems to distribute both the work and some of the demand, and it makes it possible for the cluster to grow or shrink dramatically without failing to take advantage of the local storage available at each node. Unfortunately, without any pre-work, it's possible that, doing option 2, you'd pick bad subsets of work to distribute to each node, and get no local I/O improvement at all.\n\nI'd really like to see something done, perferably soon. With dozens of nodes, and hundreds of gbs of data in my current problem set, it's very nearly impossible to get the current code to make progress, without killing tasktrackers (some with lots of work units already completed). I can do some of the coding, if there's agreement for what direction to push.","from":"developer"},{"body":"I'm re-opening this & assigning it to Mike.\n\nI think we can fix this by changing the TaskTracker to incrementally calculate things, as I alluded above. The TaskTracker should avoid ever doing anything that might take very long. RPC timeouts to the TaskTracker are very bad and must be avoided.\n\nWe can maintain a mapping of taskTracker->split, initially empty, and a queue of splits, initially filled with all of the splits. (If basic split-generation is too expensive, since it calls file-length on each input file,, then we can eventually change the split API into a generator, which incrementally enumerates splits and use that in place of this queue.) When a request for a task arrives from a tasktracker we can first examine the table. If any splits are present for the calling tracker, then we return one and remove it from the table. Otherwise we pop a constant number of splits (10?) off of the queue, enumerate the tasktrackers that host each split by calling the namenode, and add entries to the table. Then we consult the table again. In most cases (many more splits than tasktrackers) we should identify suitable splits. If we fail then we assign a randomly selected, non-local split to the tasktracker. Make sense?","from":"developer"},{"body":"Eric Baldeschwieler wrote:\n> [...] why not just dedicate a thread \n> to planning and then load a complete plan? That can produce more \n> optimal placement and a simpler to understand initialization sequence.\n\nA separate thread in the JobTracker? That could be a good approach. We'd have a queue of submitted but as-yet unplanned jobs. The thread can then pop a job off the queue, compute its splits, then start populating a tasktracker->split table. When tasktrackers poll for work they can consult this table, potentially while the thread is still populating it.\n\nI'm hesitant to move this out of the JobTracker into the TaskTracker, since that introduces complexity. But a single thread in the JobTracker should be simple to add and should mostly solve this. +1","from":"developer"},{"body":"My only concern is that this doesn't scale as well for really huge jobs, but, if implemented as you suggest (pre-filling, hand out non-matching jobs when you don't have matched jobs) seems like a reasonable trade off. \n\nI'd still think it'd be nice to better handle the case of newly-discovered tasktrackers, but that's probably much more minor than other scaling/responsiveness issues. Maybe a future enhancement for the JobTracker thread that does this work.","from":"developer"},{"body":"Having a separate thread in JobTracker can even pre-compute a data locality map that will enable optimal task assignment to task trackers.\nHere is the basic idea. As the thread computes the splits for a job, it can determine which nodes have local data for each split. The thread can populate \na map mapping a node to a list of the splits whose data are on the node. When a tasktracker polls for a task, the JobTracker can look up the map by \nthe node name of the task tracker to see whether there are any un-assigned splits in the corresponding list. If yes, assign one of them to the task tracker. \nOtherwise, randomlly choose a split for the task tracker. ","from":"developer"},{"body":"\n I created a separate thread on the JobTracker that handles file-splitting and\ngathering info about split-caches. The job is placed into the PREP state\nuntil it is processed by this JobTracker thread. After processing, it moves\ninto the RUNNING state. \n\n This should allow the hard split-work to continue but still allow the JobTracker\nto process heartbeats. For now, the thread lives at the JobTracker but perhaps\nsomeday we'll move it to a TaskTracker thread.\n\n I also fixed the alg for choosing a task-allocation given a tasktracker.\n(Whether cached, non-cached, or speculative.)\n\n Let me know if this fixes some of the problems you've been seeing. \n\n","from":"developer"},{"body":"I applied Mike's patch, with a few small changes, fixing an NPE and ArrayOutOfBounds in dfs.\n\nThere's still a bug where, with mapred.tasktracker.tasks.maximum=2, only a single map and a single reduce are running on each node. I believe the intent is that there should be up to two of each related to a single job, so that the reduce tasks can be copying data while the maps are still running.","from":"developer"},{"body":"This is great, and finally fixes my issues with a large job that would never start.\n\nHowever, the way things are in this patch (and the current code), the job doesn't get started until the background thread finishes computing all of the cache hints. This takes far too long - it took 15 minutes on a recent run. During that time, of course, no other work was getting done. How about moving the cachedHints-filling-loop to the end of initTasks(), and go ahead and set the job to RUNNING and \"tasksInited=true\" in the meantime?\n\nDoing this locally lets work commence immediately, while the cache hints continue to get filled in for future task allocations.","from":"developer"}],"created":"2006-02-01T08:25:38.000+0000","description":"We've been using Nutch 0.8 (MapReduce) to perform some internet crawling. Things seemed to be going well until...\n\n060129 222409 Lost tracker 'tracker_56288'\n060129 222409 Task 'task_m_10gs5f' has been lost.\n060129 222409 Task 'task_m_10qhzr' has been lost.\n ........\n ........\n060129 222409 Task 'task_r_zggbwu' has been lost.\n060129 222409 Task 'task_r_zh8dao' has been lost.\n060129 222455 Server handler 8 on 8010 caught: java.net.SocketException: Socket closed\njava.net.SocketException: Socket closed\n at java.net.SocketOutputStream.socketWrite(SocketOutputStream.java:99)\n at java.net.SocketOutputStream.write(SocketOutputStream.java:136)\n at java.io.BufferedOutputStream.flushBuffer(BufferedOutputStream.java:65)\n at java.io.BufferedOutputStream.flush(BufferedOutputStream.java:123)\n at java.io.DataOutputStream.flush(DataOutputStream.java:106)\n at org.apache.nutch.ipc.Server$Handler.run(Server.java:216)\n060129 222455 Adding task 'task_m_cia5po' to set for tracker 'tracker_56288'\n060129 223711 Adding task 'task_m_ffv59i' to set for tracker 'tracker_25647'\n\nI'm hoping that someone could explain why task_m_cia5po got added to tracker_56288 after this tracker was lost.\n\nThe Crawl .main process died with the following output:\n\n060129 221129 Indexer: adding segment: /user/crawler/crawl-20060129091444/segments/20060129200246\nException in thread \"main\" java.io.IOException: timed out waiting for response\n at org.apache.nutch.ipc.Client.call(Client.java:296)\n at org.apache.nutch.ipc.RPC$Invoker.invoke(RPC.java:127)\n at $Proxy1.submitJob(Unknown Source)\n at org.apache.nutch.mapred.JobClient.submitJob(JobClient.java:259)\n at org.apache.nutch.mapred.JobClient.runJob(JobClient.java:288)\n at org.apache.nutch.indexer.Indexer.index(Indexer.java:263)\n at org.apache.nutch.crawl.Crawl.main(Crawl.java:127)\n\nHowever, it definitely seems as if the JobTracker is still waiting for the job to finish (no failed jobs).\n\nDoug Cutting's response:\nThe bug here is that the RPC call times out while the map task is computing splits. The fix is that the job tracker should not compute splits until after it has returned from the submitJob RPC. Please submit a bug in Jira to help remind us to fix this.\n","issue_id":"12328376","key":"HADOOP-16","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-03-03T08:09:22.000+0000","role":"fixed_distractor","summary":"RPC call times out while indexing map task is computing splits"} {"case_id":"13212344","cluster":"DISTRACTOR-HADOOP-16080","comments":[{"body":"This gist contains the pom file that I used to created a relocated version of hadoop-aws.\r\n\r\nhttps://gist.github.com/keith-turner/f6dcbd33342732e42695d66509239983","created":"2019-01-28T22:27:44.035+0000"},{"body":"Seems like ideally the hadoop-aws module would avoid using any hadoop APIs with non-hadoop/relocated types.","created":"2019-01-28T22:59:18.835+0000"},{"body":"keith. I appreciate your concerns. What we really want is an object store dependency which contains all the store connectors with shaded dependencies, HADOOP-15387. \r\n\r\nNobody has volunteered to do this -yet. \r\n\r\nYou've actually started this, which makes this a more viable proposition than HADOOP-15387. \r\n\r\nDo you want to take on the challenge of a hadoop-cloud-storage-shaded artifact? I think we'll all have to help with the testing (summary: it'll be hard) but it will be appreciated by all those downstream projects.","created":"2019-01-29T12:18:49.120+0000"},{"body":"> Do you want to take on the challenge of a hadoop-cloud-storage-shaded artifact?\r\n\r\nNo, I am quite busy and HADOOP-15387 feels like the wrong direction. The envisioned hadoop-cloudstorage artifact seems misaligned with the communities and dependencies. Seems a better structure would be that hadoop-aws is an independent artifact that only uses public/stable hadoop APIs. I took a look at SemaphoredDelegatingExecutor and noticed that is marked InterfaceAudience.Private, so it seems like hadoop-aws should just not use it. However, maybe its not feasible for hadoop-aws to only use public/stable APIs. If I magically had the time I would explore making hadoop-aws more independent instead of more dependent. \r\n\r\n","created":"2019-01-29T15:51:40.525+0000"},{"body":"bq. No, I am quite busy \r\n\r\nThat is the problem we all have, I'm afraid.\r\n\r\nbq. The envisioned hadoop-cloudstorage artifact seems misaligned with the communities and dependencies. \r\n\r\nWhy so? \r\n\r\nSpark has a declared dependency on the unshaded hadoop-cloud-storage JAR: https://github.com/apache/spark/blob/master/hadoop-cloud/pom.xml#L208; so does Tez, and some other projects. Having a shaded offering would only need a change in those declarations and cover all the stores.\r\n\r\nbq. Seems a better structure would be that hadoop-aws is an independent artifact that only uses public/stable hadoop APIs. I took a look at SemaphoredDelegatingExecutor and noticed that is marked InterfaceAudience.Private, so it seems like hadoop-aws should just not use it\r\n\r\nSemaphoredDelegatingExecutor actually arrived in hadoop-aws first, HADOOP-13560; pulled up into hadoop-common by HADOOP-15309 so that it could be shared by the other object stores. It's private *within Hadoop itself*. By tagging as such, we retain the option of making incompatible changes. Similarly, we keep a lot of implementation stuff in hadoop-common, and share test suites of FS behaviours in hadoop-common-tests. That keeps maintenance costs down (do I really have to have a copy and paste of SemaphoredDelegatingExecutor? What about EtagChecksum? or all the new fs.impl stuff I'm adding in HADOOP-15229 for async IO?\r\n\r\nbq. If I magically had the time I would explore making hadoop-aws more independent instead of more dependent.\r\n\r\nThe other aspect of a shaded cloud moduleis that it would also be able to hide transitive dependencies. \r\nYou've avoided seeing that problem because you already had SLF4J, commons-*, etc on the CP, of compatible versions, and as we've switched to the shaded AWS SDK, so you don't have to worry about the jackson and httpclient problems which are complex enough that we are going to have to stop making Hadoop 2.7.x releases. But hadoop-azure does pass on its unshaded dependencies, as do some others -and I do get to deal with those problems. If we can produce a single JAR \"depend on this and you won't have classpath problems\", people will be happy. It that which tends to be the most traumatic.","created":"2019-01-29T18:15:59.730+0000"},{"body":"Hi, All.\r\nIs there any update on this JIRA? Otherwise, can we increase the `Priority` from `Major` to `Blocker`?","created":"2020-11-25T20:33:18.225+0000"},{"body":"the way to address the issue would be to do a shaded-cloud-connectors module, not the proposed \"lets only use the hadoop-client API\" policy.\r\n\r\nI've not had a chance to do that, if you want to begin instead","created":"2020-11-26T21:11:32.536+0000"},{"body":"[~stevel@apache.org] One other issue here is that, even though the APIs are private within Hadoop itself, should we avoid using Guava in the APIs themselves? this will allow Hadoop modules fit better with each other.","created":"2020-11-30T19:57:56.670+0000"},{"body":"Thanks [~csun] and [~stevel@apache.org] for your works.\r\nCommitted to branch-3.2.2, backport, compile and check at local passed.","created":"2020-12-04T06:15:00.003+0000"}],"conversations":[{"body":"I attempted to use Accumulo and S3a with the following jars on the classpath.\r\n\r\n * hadoop-client-api-3.1.1.jar\r\n * hadoop-client-runtime-3.1.1.jar\r\n * hadoop-aws-3.1.1.jar\r\n\r\nThis failed with the following exception.\r\n\r\n{noformat}\r\nException in thread \"init\" java.lang.NoSuchMethodError: org.apache.hadoop.util.SemaphoredDelegatingExecutor.(Lcom/google/common/util/concurrent/ListeningExecutorService;IZ)V\r\n at org.apache.hadoop.fs.s3a.S3AFileSystem.create(S3AFileSystem.java:769)\r\n at org.apache.hadoop.fs.FileSystem.create(FileSystem.java:1169)\r\n at org.apache.hadoop.fs.FileSystem.create(FileSystem.java:1149)\r\n at org.apache.hadoop.fs.FileSystem.create(FileSystem.java:1108)\r\n at org.apache.hadoop.fs.FileSystem.createNewFile(FileSystem.java:1413)\r\n at org.apache.accumulo.server.fs.VolumeManagerImpl.createNewFile(VolumeManagerImpl.java:184)\r\n at org.apache.accumulo.server.init.Initialize.initDirs(Initialize.java:479)\r\n at org.apache.accumulo.server.init.Initialize.initFileSystem(Initialize.java:487)\r\n at org.apache.accumulo.server.init.Initialize.initialize(Initialize.java:370)\r\n at org.apache.accumulo.server.init.Initialize.doInit(Initialize.java:348)\r\n at org.apache.accumulo.server.init.Initialize.execute(Initialize.java:967)\r\n at org.apache.accumulo.start.Main.lambda$execKeyword$0(Main.java:129)\r\n at java.lang.Thread.run(Thread.java:748)\r\n{noformat}\r\n\r\nThe problem is that {{S3AFileSystem.create()}} looks for {{SemaphoredDelegatingExecutor(com.google.common.util.concurrent.ListeningExecutorService)}} which does not exist in hadoop-client-api-3.1.1.jar. What does exist is {{SemaphoredDelegatingExecutor(org.apache.hadoop.shaded.com.google.common.util.concurrent.ListeningExecutorService)}}.\r\n\r\nTo work around this issue I created a version of hadoop-aws-3.1.1.jar that relocated references to Guava.\r\n","from":"reporter","subject":"hadoop-aws does not work with hadoop-client-api"},{"body":"This gist contains the pom file that I used to created a relocated version of hadoop-aws.\r\n\r\nhttps://gist.github.com/keith-turner/f6dcbd33342732e42695d66509239983","from":"developer"},{"body":"Seems like ideally the hadoop-aws module would avoid using any hadoop APIs with non-hadoop/relocated types.","from":"developer"},{"body":"keith. I appreciate your concerns. What we really want is an object store dependency which contains all the store connectors with shaded dependencies, HADOOP-15387. \r\n\r\nNobody has volunteered to do this -yet. \r\n\r\nYou've actually started this, which makes this a more viable proposition than HADOOP-15387. \r\n\r\nDo you want to take on the challenge of a hadoop-cloud-storage-shaded artifact? I think we'll all have to help with the testing (summary: it'll be hard) but it will be appreciated by all those downstream projects.","from":"developer"},{"body":"> Do you want to take on the challenge of a hadoop-cloud-storage-shaded artifact?\r\n\r\nNo, I am quite busy and HADOOP-15387 feels like the wrong direction. The envisioned hadoop-cloudstorage artifact seems misaligned with the communities and dependencies. Seems a better structure would be that hadoop-aws is an independent artifact that only uses public/stable hadoop APIs. I took a look at SemaphoredDelegatingExecutor and noticed that is marked InterfaceAudience.Private, so it seems like hadoop-aws should just not use it. However, maybe its not feasible for hadoop-aws to only use public/stable APIs. If I magically had the time I would explore making hadoop-aws more independent instead of more dependent. \r\n\r\n","from":"developer"},{"body":"bq. No, I am quite busy \r\n\r\nThat is the problem we all have, I'm afraid.\r\n\r\nbq. The envisioned hadoop-cloudstorage artifact seems misaligned with the communities and dependencies. \r\n\r\nWhy so? \r\n\r\nSpark has a declared dependency on the unshaded hadoop-cloud-storage JAR: https://github.com/apache/spark/blob/master/hadoop-cloud/pom.xml#L208; so does Tez, and some other projects. Having a shaded offering would only need a change in those declarations and cover all the stores.\r\n\r\nbq. Seems a better structure would be that hadoop-aws is an independent artifact that only uses public/stable hadoop APIs. I took a look at SemaphoredDelegatingExecutor and noticed that is marked InterfaceAudience.Private, so it seems like hadoop-aws should just not use it\r\n\r\nSemaphoredDelegatingExecutor actually arrived in hadoop-aws first, HADOOP-13560; pulled up into hadoop-common by HADOOP-15309 so that it could be shared by the other object stores. It's private *within Hadoop itself*. By tagging as such, we retain the option of making incompatible changes. Similarly, we keep a lot of implementation stuff in hadoop-common, and share test suites of FS behaviours in hadoop-common-tests. That keeps maintenance costs down (do I really have to have a copy and paste of SemaphoredDelegatingExecutor? What about EtagChecksum? or all the new fs.impl stuff I'm adding in HADOOP-15229 for async IO?\r\n\r\nbq. If I magically had the time I would explore making hadoop-aws more independent instead of more dependent.\r\n\r\nThe other aspect of a shaded cloud moduleis that it would also be able to hide transitive dependencies. \r\nYou've avoided seeing that problem because you already had SLF4J, commons-*, etc on the CP, of compatible versions, and as we've switched to the shaded AWS SDK, so you don't have to worry about the jackson and httpclient problems which are complex enough that we are going to have to stop making Hadoop 2.7.x releases. But hadoop-azure does pass on its unshaded dependencies, as do some others -and I do get to deal with those problems. If we can produce a single JAR \"depend on this and you won't have classpath problems\", people will be happy. It that which tends to be the most traumatic.","from":"developer"},{"body":"Hi, All.\r\nIs there any update on this JIRA? Otherwise, can we increase the `Priority` from `Major` to `Blocker`?","from":"developer"},{"body":"the way to address the issue would be to do a shaded-cloud-connectors module, not the proposed \"lets only use the hadoop-client API\" policy.\r\n\r\nI've not had a chance to do that, if you want to begin instead","from":"developer"},{"body":"[~stevel@apache.org] One other issue here is that, even though the APIs are private within Hadoop itself, should we avoid using Guava in the APIs themselves? this will allow Hadoop modules fit better with each other.","from":"developer"},{"body":"Thanks [~csun] and [~stevel@apache.org] for your works.\r\nCommitted to branch-3.2.2, backport, compile and check at local passed.","from":"developer"}],"created":"2019-01-28T22:05:14.000+0000","description":"I attempted to use Accumulo and S3a with the following jars on the classpath.\r\n\r\n * hadoop-client-api-3.1.1.jar\r\n * hadoop-client-runtime-3.1.1.jar\r\n * hadoop-aws-3.1.1.jar\r\n\r\nThis failed with the following exception.\r\n\r\n{noformat}\r\nException in thread \"init\" java.lang.NoSuchMethodError: org.apache.hadoop.util.SemaphoredDelegatingExecutor.(Lcom/google/common/util/concurrent/ListeningExecutorService;IZ)V\r\n at org.apache.hadoop.fs.s3a.S3AFileSystem.create(S3AFileSystem.java:769)\r\n at org.apache.hadoop.fs.FileSystem.create(FileSystem.java:1169)\r\n at org.apache.hadoop.fs.FileSystem.create(FileSystem.java:1149)\r\n at org.apache.hadoop.fs.FileSystem.create(FileSystem.java:1108)\r\n at org.apache.hadoop.fs.FileSystem.createNewFile(FileSystem.java:1413)\r\n at org.apache.accumulo.server.fs.VolumeManagerImpl.createNewFile(VolumeManagerImpl.java:184)\r\n at org.apache.accumulo.server.init.Initialize.initDirs(Initialize.java:479)\r\n at org.apache.accumulo.server.init.Initialize.initFileSystem(Initialize.java:487)\r\n at org.apache.accumulo.server.init.Initialize.initialize(Initialize.java:370)\r\n at org.apache.accumulo.server.init.Initialize.doInit(Initialize.java:348)\r\n at org.apache.accumulo.server.init.Initialize.execute(Initialize.java:967)\r\n at org.apache.accumulo.start.Main.lambda$execKeyword$0(Main.java:129)\r\n at java.lang.Thread.run(Thread.java:748)\r\n{noformat}\r\n\r\nThe problem is that {{S3AFileSystem.create()}} looks for {{SemaphoredDelegatingExecutor(com.google.common.util.concurrent.ListeningExecutorService)}} which does not exist in hadoop-client-api-3.1.1.jar. What does exist is {{SemaphoredDelegatingExecutor(org.apache.hadoop.shaded.com.google.common.util.concurrent.ListeningExecutorService)}}.\r\n\r\nTo work around this issue I created a version of hadoop-aws-3.1.1.jar that relocated references to Guava.\r\n","issue_id":"13212344","key":"HADOOP-16080","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2020-12-04T00:14:23.000+0000","role":"fixed_distractor","summary":"hadoop-aws does not work with hadoop-client-api"} {"case_id":"13227363","cluster":"DISTRACTOR-HADOOP-16247","comments":[{"body":"#URI.normalize returns same URI if relativePath, will not be normalized because it is Opaque. So \"path\" object will be null.\r\n\r\nI will post a new patch. It looks like If we don't decode msg prior path construction, then we may end up having more change to handle it. ","created":"2019-04-11T01:14:04.599+0000"},{"body":"fyi - [~xiaochen] [~zvenczel] [~gabor.bota]","created":"2019-04-11T01:16:14.499+0000"},{"body":"Hi [~kpalanisamy] thanks for reporting the issue!\r\nDo you have a test case that reproduces the bug? The tests pass even if I remove the fix.\r\n\r\nTo reproduce, I think you'll need to patch the code a little bit:\r\n\r\n{code}\r\n public static int readStream(String path) throws Exception{\r\n URL url = new URL(path);\r\n- URI uri = url.toURI();\r\n- FileSystem fs = FileSystem.get(uri, CONFIGURATION);\r\n- FSDataInputStream fsdis = fs.open(new Path(uri));\r\n- return fsdis.available();\r\n+ InputStream is = url.openStream();\r\n+ return is.available();\r\n+\r\n }\r\n{code}","created":"2019-04-16T11:59:16.034+0000"},{"body":"Thank you [~weichiu] for review and repro step. Posted new patch with the repro change. ","created":"2019-04-16T18:22:18.427+0000"},{"body":"I don't think this patch will work for anything but file: URIs.  Specifically, {{getSchemeSpecificPart()}} is going to return //host/path for hdfs://host/path.","created":"2019-04-16T21:42:16.742+0000"},{"body":"Thank you [~daryn] looking into it. \r\n\r\nIf users would like to use \"hdfs\" schema, then they may need to pass in URI like,\r\n\r\nURI(String scheme, String host, String path, String fragment)\r\n\r\nURI(String scheme, String userInfo, String host, int port, String path, String query, String fragment)\r\n\r\n \r\n\r\nI don't think anyone used it with \"hdfs\" schema earlier or they might use URI like above. \r\n\r\nThis simple fix is given because it is working in Hadoop 2.x but not working in Hadoop 3.x.","created":"2019-04-16T23:15:22.682+0000"},{"body":"[~daryn] [~weichiu]\r\n\r\nWould you please ack? This is critical because we can't requst users to change application code to make it work.","created":"2019-04-19T22:23:49.736+0000"},{"body":"Yes [~daryn], old change will not work for HDFS fs url.  I have a new code change now. \r\n\r\nThe problem initailly by #URL.getPath(), this will not work if URL file path have any space. So HADOOP-15217 fix is to change URL.toURI() which handle space issue. Now, there is another problem with URI. If URI contains relative/opaque file path, then path will be null in URI object. \r\n\r\nI think only change that work for both file and hdfs schema is,\r\n{code:java}\r\nFsUrlConnection.java\r\n\r\nif(uri.isOpaque() && uri.getScheme().equals(\"file\")) {\r\n is = fs.open(new Path(uri.getSchemeSpecificPart()));\r\n } else {\r\n is = fs.open(new Path(uri));\r\n }\r\n\r\n{code}\r\n \r\n\r\nThanks [~bharatviswa] for offline discussion. \r\n\r\n[~daryn] [~weichiu] Would you please suggest if any?","created":"2019-05-01T00:16:28.679+0000"},{"body":"Patch over all LGTM. Could you add some explanation for the change, it will be easy when someone is reading the code for the first time. Once this is done, I am +1 with the change.\r\n\r\nThank You [~kpalanisamy] for the fix. (I think now with this patch, it will fix the file scheme with a relative path, and for rest of the URI's it uses the old code, so it does not break anything.\r\n\r\n \r\n\r\n \r\n\r\n ","created":"2019-05-09T22:18:41.489+0000"},{"body":"Thank you [~bharatviswa] for reviewing this patch. I updated a new patch with some info. ","created":"2019-05-10T01:24:54.198+0000"},{"body":"+1 LGTM, [~kpalanisamy] can you fix checkstyle issues reported by jenkins, it is strange it has given +1, but it has checkstyle errors.\r\n\r\nI Will wait for a couple of days if no more comments I will commit this patch, as [~daryn] and [~weichiu] have already reviewed and had some comments.","created":"2019-05-10T04:38:35.161+0000"},{"body":"Thank you [~bharatviswa].","created":"2019-05-10T08:42:42.430+0000"},{"body":"As no further comments, I will commit this shortly.\r\n\r\n ","created":"2019-05-16T00:39:00.480+0000"},{"body":"Thank You [~kpalanisamy] for the contribution, [~weichiu]  and [~daryn] for the review.\r\n\r\nI have committed this to the trunk, branch-3.1, and branch-3.2.\r\n\r\n ","created":"2019-05-16T00:49:30.473+0000"}],"conversations":[{"body":"FsUrlConnection doesn't handle relativePath correctly after the change [HADOOP-15217|https://issues.apache.org/jira/browse/HADOOP-15217]\r\n\r\n{code}\r\n\r\nException in thread \"main\" java.lang.NullPointerException\r\n at org.apache.hadoop.fs.Path.isUriPathAbsolute(Path.java:385)\r\n at org.apache.hadoop.fs.Path.isAbsolute(Path.java:395)\r\n at org.apache.hadoop.fs.RawLocalFileSystem.pathToFile(RawLocalFileSystem.java:87)\r\n at org.apache.hadoop.fs.RawLocalFileSystem.deprecatedGetFileStatus(RawLocalFileSystem.java:636)\r\n at org.apache.hadoop.fs.RawLocalFileSystem.getFileLinkStatusInternal(RawLocalFileSystem.java:930)\r\n at org.apache.hadoop.fs.RawLocalFileSystem.getFileStatus(RawLocalFileSystem.java:631)\r\n at org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:454)\r\n at org.apache.hadoop.fs.ChecksumFileSystem$ChecksumFSInputChecker.(ChecksumFileSystem.java:146)\r\n at org.apache.hadoop.fs.ChecksumFileSystem.open(ChecksumFileSystem.java:347)\r\n at org.apache.hadoop.fs.FileSystem.open(FileSystem.java:899)\r\n at org.apache.hadoop.fs.FsUrlConnection.connect(FsUrlConnection.java:62)\r\n at org.apache.hadoop.fs.FsUrlConnection.getInputStream(FsUrlConnection.java:71)\r\n at java.net.URL.openStream(URL.java:1045)\r\n at UrlProblem.testRelativePath(UrlProblem.java:33)\r\n at UrlProblem.main(UrlProblem.java:19)\r\n\r\n{code}","from":"reporter","subject":"NPE in FsUrlConnection"},{"body":"#URI.normalize returns same URI if relativePath, will not be normalized because it is Opaque. So \"path\" object will be null.\r\n\r\nI will post a new patch. It looks like If we don't decode msg prior path construction, then we may end up having more change to handle it. ","from":"developer"},{"body":"fyi - [~xiaochen] [~zvenczel] [~gabor.bota]","from":"developer"},{"body":"Hi [~kpalanisamy] thanks for reporting the issue!\r\nDo you have a test case that reproduces the bug? The tests pass even if I remove the fix.\r\n\r\nTo reproduce, I think you'll need to patch the code a little bit:\r\n\r\n{code}\r\n public static int readStream(String path) throws Exception{\r\n URL url = new URL(path);\r\n- URI uri = url.toURI();\r\n- FileSystem fs = FileSystem.get(uri, CONFIGURATION);\r\n- FSDataInputStream fsdis = fs.open(new Path(uri));\r\n- return fsdis.available();\r\n+ InputStream is = url.openStream();\r\n+ return is.available();\r\n+\r\n }\r\n{code}","from":"developer"},{"body":"Thank you [~weichiu] for review and repro step. Posted new patch with the repro change. ","from":"developer"},{"body":"I don't think this patch will work for anything but file: URIs.  Specifically, {{getSchemeSpecificPart()}} is going to return //host/path for hdfs://host/path.","from":"developer"},{"body":"Thank you [~daryn] looking into it. \r\n\r\nIf users would like to use \"hdfs\" schema, then they may need to pass in URI like,\r\n\r\nURI(String scheme, String host, String path, String fragment)\r\n\r\nURI(String scheme, String userInfo, String host, int port, String path, String query, String fragment)\r\n\r\n \r\n\r\nI don't think anyone used it with \"hdfs\" schema earlier or they might use URI like above. \r\n\r\nThis simple fix is given because it is working in Hadoop 2.x but not working in Hadoop 3.x.","from":"developer"},{"body":"[~daryn] [~weichiu]\r\n\r\nWould you please ack? This is critical because we can't requst users to change application code to make it work.","from":"developer"},{"body":"Yes [~daryn], old change will not work for HDFS fs url.  I have a new code change now. \r\n\r\nThe problem initailly by #URL.getPath(), this will not work if URL file path have any space. So HADOOP-15217 fix is to change URL.toURI() which handle space issue. Now, there is another problem with URI. If URI contains relative/opaque file path, then path will be null in URI object. \r\n\r\nI think only change that work for both file and hdfs schema is,\r\n{code:java}\r\nFsUrlConnection.java\r\n\r\nif(uri.isOpaque() && uri.getScheme().equals(\"file\")) {\r\n is = fs.open(new Path(uri.getSchemeSpecificPart()));\r\n } else {\r\n is = fs.open(new Path(uri));\r\n }\r\n\r\n{code}\r\n \r\n\r\nThanks [~bharatviswa] for offline discussion. \r\n\r\n[~daryn] [~weichiu] Would you please suggest if any?","from":"developer"},{"body":"Patch over all LGTM. Could you add some explanation for the change, it will be easy when someone is reading the code for the first time. Once this is done, I am +1 with the change.\r\n\r\nThank You [~kpalanisamy] for the fix. (I think now with this patch, it will fix the file scheme with a relative path, and for rest of the URI's it uses the old code, so it does not break anything.\r\n\r\n \r\n\r\n \r\n\r\n ","from":"developer"},{"body":"Thank you [~bharatviswa] for reviewing this patch. I updated a new patch with some info. ","from":"developer"},{"body":"+1 LGTM, [~kpalanisamy] can you fix checkstyle issues reported by jenkins, it is strange it has given +1, but it has checkstyle errors.\r\n\r\nI Will wait for a couple of days if no more comments I will commit this patch, as [~daryn] and [~weichiu] have already reviewed and had some comments.","from":"developer"},{"body":"Thank you [~bharatviswa].","from":"developer"},{"body":"As no further comments, I will commit this shortly.\r\n\r\n ","from":"developer"},{"body":"Thank You [~kpalanisamy] for the contribution, [~weichiu]  and [~daryn] for the review.\r\n\r\nI have committed this to the trunk, branch-3.1, and branch-3.2.\r\n\r\n ","from":"developer"}],"created":"2019-04-11T00:52:19.000+0000","description":"FsUrlConnection doesn't handle relativePath correctly after the change [HADOOP-15217|https://issues.apache.org/jira/browse/HADOOP-15217]\r\n\r\n{code}\r\n\r\nException in thread \"main\" java.lang.NullPointerException\r\n at org.apache.hadoop.fs.Path.isUriPathAbsolute(Path.java:385)\r\n at org.apache.hadoop.fs.Path.isAbsolute(Path.java:395)\r\n at org.apache.hadoop.fs.RawLocalFileSystem.pathToFile(RawLocalFileSystem.java:87)\r\n at org.apache.hadoop.fs.RawLocalFileSystem.deprecatedGetFileStatus(RawLocalFileSystem.java:636)\r\n at org.apache.hadoop.fs.RawLocalFileSystem.getFileLinkStatusInternal(RawLocalFileSystem.java:930)\r\n at org.apache.hadoop.fs.RawLocalFileSystem.getFileStatus(RawLocalFileSystem.java:631)\r\n at org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:454)\r\n at org.apache.hadoop.fs.ChecksumFileSystem$ChecksumFSInputChecker.(ChecksumFileSystem.java:146)\r\n at org.apache.hadoop.fs.ChecksumFileSystem.open(ChecksumFileSystem.java:347)\r\n at org.apache.hadoop.fs.FileSystem.open(FileSystem.java:899)\r\n at org.apache.hadoop.fs.FsUrlConnection.connect(FsUrlConnection.java:62)\r\n at org.apache.hadoop.fs.FsUrlConnection.getInputStream(FsUrlConnection.java:71)\r\n at java.net.URL.openStream(URL.java:1045)\r\n at UrlProblem.testRelativePath(UrlProblem.java:33)\r\n at UrlProblem.main(UrlProblem.java:19)\r\n\r\n{code}","issue_id":"13227363","key":"HADOOP-16247","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2019-05-16T00:49:30.000+0000","role":"fixed_distractor","summary":"NPE in FsUrlConnection"} {"case_id":"13231915","cluster":"DISTRACTOR-HADOOP-16299","comments":[{"body":"I thought it is no use compiling Apache Hadoop with Java 11 and setting the target version to 1.8. However, now I'm thinking it is useful for testing Apache Hadoop (compiled with Java 8) with Java 11.","created":"2019-05-07T03:54:08.105+0000"},{"body":"001 patch\r\nCreated a new profile that runs only when \"-Djavac.version=11\". The disadvantage is that when using Java 12 or upper and set \"-Djavac.version\" to 12 or upper, this profile does not work.","created":"2019-05-07T05:40:25.211+0000"},{"body":"patch JGTM, though maybe a profile \"java11\" would be clearer. And yes, we will need a java 12 version in future, at least while java 8 is supported","created":"2019-05-07T09:26:33.447+0000"},{"body":"Thanks [~stevel@apache.org] for the review.\r\n\"java11\" seems misleading because there is already a profile \"jdk11\" to be activated with Java 11+. Maybe \"add-exports\", \"jigsaw\", \"module\", or something?","created":"2019-05-07T09:38:16.298+0000"},{"body":"002 patch: I found the way to discard the --add-exports settings by removing the direct usages of com.sun.jndi.ldap package. This patch is based on https://issues.apache.org/jira/secure/attachment/12951455/HADOOP-15941.1.patch. Thanks [~tasanuma0829] for the initial work.","created":"2019-05-07T12:59:41.635+0000"},{"body":"I could successfully run the LdapGroupsMapping* unit tests in the following 3 environments.\r\n* OpenJDK 8\r\n* OpenJDK 11.0.3\r\n* OpenJDK 11.0.3 + {{-Djavac.version=11}} option","created":"2019-05-08T05:28:26.299+0000"},{"body":"+1. Will commit it later.","created":"2019-05-09T05:43:15.524+0000"},{"body":"Committed to trunk. Thanks for your contribution, [~ajisakaa], and thanks for your review, [~stevel@apache.org].","created":"2019-05-09T05:52:52.990+0000"},{"body":"Thank you, [~tasanuma0829]!","created":"2019-05-09T05:56:46.312+0000"}],"conversations":[{"body":"{{mvn install -DskipTests}} fails on Java 11 without specifying {{-Djavac.version=11}}.\r\n{noformat}\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.1:compile (default-compile) on project hadoop-annotations: Fatal error compiling: error: option --add-exports not allowed with target 1.8 -> [Help 1]\r\n{noformat}\r\nHADOOP-15941 added {{--add-exports}} option when the java version is 11 but the option is not allowed when the javac target version is 1.8.","from":"reporter","subject":"[JDK 11] Build fails without specifying -Djavac.version=11"},{"body":"I thought it is no use compiling Apache Hadoop with Java 11 and setting the target version to 1.8. However, now I'm thinking it is useful for testing Apache Hadoop (compiled with Java 8) with Java 11.","from":"developer"},{"body":"001 patch\r\nCreated a new profile that runs only when \"-Djavac.version=11\". The disadvantage is that when using Java 12 or upper and set \"-Djavac.version\" to 12 or upper, this profile does not work.","from":"developer"},{"body":"patch JGTM, though maybe a profile \"java11\" would be clearer. And yes, we will need a java 12 version in future, at least while java 8 is supported","from":"developer"},{"body":"Thanks [~stevel@apache.org] for the review.\r\n\"java11\" seems misleading because there is already a profile \"jdk11\" to be activated with Java 11+. Maybe \"add-exports\", \"jigsaw\", \"module\", or something?","from":"developer"},{"body":"002 patch: I found the way to discard the --add-exports settings by removing the direct usages of com.sun.jndi.ldap package. This patch is based on https://issues.apache.org/jira/secure/attachment/12951455/HADOOP-15941.1.patch. Thanks [~tasanuma0829] for the initial work.","from":"developer"},{"body":"I could successfully run the LdapGroupsMapping* unit tests in the following 3 environments.\r\n* OpenJDK 8\r\n* OpenJDK 11.0.3\r\n* OpenJDK 11.0.3 + {{-Djavac.version=11}} option","from":"developer"},{"body":"+1. Will commit it later.","from":"developer"},{"body":"Committed to trunk. Thanks for your contribution, [~ajisakaa], and thanks for your review, [~stevel@apache.org].","from":"developer"},{"body":"Thank you, [~tasanuma0829]!","from":"developer"}],"created":"2019-05-07T03:42:10.000+0000","description":"{{mvn install -DskipTests}} fails on Java 11 without specifying {{-Djavac.version=11}}.\r\n{noformat}\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.1:compile (default-compile) on project hadoop-annotations: Fatal error compiling: error: option --add-exports not allowed with target 1.8 -> [Help 1]\r\n{noformat}\r\nHADOOP-15941 added {{--add-exports}} option when the java version is 11 but the option is not allowed when the javac target version is 1.8.","issue_id":"13231915","key":"HADOOP-16299","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2019-05-09T05:52:52.000+0000","role":"fixed_distractor","summary":"[JDK 11] Build fails without specifying -Djavac.version=11"} {"case_id":"12332617","cluster":"DISTRACTOR-HADOOP-163","comments":[{"body":"Alternately, the datanode daemon should simply exit if it cannot write to its configured data directories.","created":"2006-04-25T04:56:30.000+0000"},{"body":"\nExiting is an option. However, the datanode may still be able to read, thus to serve the existing blocks.\n","created":"2006-04-25T05:19:21.000+0000"},{"body":"Good point. So perhaps a read-only node should report itself at 100% of capacity? Then the namenode should never allocate blocks to it.","created":"2006-04-25T05:52:15.000+0000"},{"body":"I envision 24x7 systems where the datanode is automatically restarted upon failure by init or another HA component. When a partition/FS fails, it will likely remain in a failed state after restart, so reporting up and stopping to serve would be better than simply exitting, which would lead to thrashing. At the end of the day, outside intervention would be required, so the most important part is diagnosing the error and reporting it as such. reporting 100% full would not generate the same kind of attention by a correction system/person.","created":"2006-04-25T23:51:04.000+0000"},{"body":"But should a datanode with no disk or a read-only disk still send heartbeats to the namenode? I think not. So it should exit the datanode daemon loop. I think more than that is hard to specify at this point. You're talking about what we should do when we have a system that automatically restarts, and a system that monitors, etc. We don't have those systems in Hadoop today, so they're hard to code to! In the meantime, do you think it would be better to enter some zombie state, not sending heartbeats or otherwise participating in namenode network protocols, but looping sending out SOS over channels TBD?","created":"2006-04-26T05:50:16.000+0000"},{"body":"what I'm suggesting is close to your suggestion:\nif you're read-only, behave as though you're 100% full, serve only read requests, but don't mislead: report you're read-only, not that you're 100% full. Namenode will avoid new block allocations to the node, but its log will contain an error that could trigger external corrective action.\n","created":"2006-04-26T07:50:17.000+0000"},{"body":"I plan to take a simple approach to this problem. If a data node finds out that it can not write to its disk. It reports to the name node and aborts. The name node logs the error and alerts it on the http UI.\n\nI will use the existing data node protocol \"errorReport\" for the error reporting but with a minor change. In addition to the parameter error message, the rpc will also send an error code so that the name node does not need to parse the error message to figure out which action to take.","created":"2006-05-20T00:56:31.000+0000"},{"body":"Sounds good - except, what \"aborts\"? The idea of the datanode staying operating, but reporting error and not accepting further blocks is probably better, but maybe you meant \"abort the block write\". The node's blocks should probably not be counted by the namenode, but still available as a source for replication. Also, staying up means that there are fewer timeouts - it used to be that, when writing large volumes into DFS, if one or more of your nodes was full, your writer would hit a periodic timeout as connections to the (full and constantly restarting) datanode were refused. Hitting a timeout because some fraction of all resources is overused is, of course, much *much* slower than continuing to stream. Further - if the datanode periodically re-tests if the error condition has lifted, it can more immediately begin contributing to the cluster productivity again.","created":"2006-05-20T01:04:46.000+0000"},{"body":"We are trying to deal with the case that the node is misconfigured / broken. Trying to operate in these situations is hard. Simpler to fail fast, IMO. This leverages the designed strengths of HDFS. Our goal is to get the information to the operator so they can diagnose and fix the problem and seal the problem off from the cluster.\n\nThis is distinct from the case that the node is simply full. That would not trigger this condition.","created":"2006-05-20T01:34:22.000+0000"},{"body":"In this patch, if a data node finds that its data directory becomes not readable or writable, it logs the error and reports the problem to its namen ode and shut down itself. When the name node receives the error report, it lots the error and removes the data node info.\n\nA data node detects disk problem at startup time, when it receives a r/w request, after it receives a command from its name node, and before it sends out a block report. A data node will not start up if its data dir is not readable or writable. ","created":"2006-05-27T05:22:39.000+0000"},{"body":"This looks great! I just committed it. Thanks, Hairong!","created":"2006-05-27T05:42:26.000+0000"}],"conversations":[{"body":"I observed that sometime, if a file of a data node is not mounted properly, it may not be writable. In this case, any data writes will fail. The name node should stop assigning new blocks to that data node. The webpage should show that node is in an abnormal state.\n\n","from":"reporter","subject":"If a DFS datanode cannot write onto its file system. it should tell the name node not to assign new blocks to it."},{"body":"Alternately, the datanode daemon should simply exit if it cannot write to its configured data directories.","from":"developer"},{"body":"\nExiting is an option. However, the datanode may still be able to read, thus to serve the existing blocks.\n","from":"developer"},{"body":"Good point. So perhaps a read-only node should report itself at 100% of capacity? Then the namenode should never allocate blocks to it.","from":"developer"},{"body":"I envision 24x7 systems where the datanode is automatically restarted upon failure by init or another HA component. When a partition/FS fails, it will likely remain in a failed state after restart, so reporting up and stopping to serve would be better than simply exitting, which would lead to thrashing. At the end of the day, outside intervention would be required, so the most important part is diagnosing the error and reporting it as such. reporting 100% full would not generate the same kind of attention by a correction system/person.","from":"developer"},{"body":"But should a datanode with no disk or a read-only disk still send heartbeats to the namenode? I think not. So it should exit the datanode daemon loop. I think more than that is hard to specify at this point. You're talking about what we should do when we have a system that automatically restarts, and a system that monitors, etc. We don't have those systems in Hadoop today, so they're hard to code to! In the meantime, do you think it would be better to enter some zombie state, not sending heartbeats or otherwise participating in namenode network protocols, but looping sending out SOS over channels TBD?","from":"developer"},{"body":"what I'm suggesting is close to your suggestion:\nif you're read-only, behave as though you're 100% full, serve only read requests, but don't mislead: report you're read-only, not that you're 100% full. Namenode will avoid new block allocations to the node, but its log will contain an error that could trigger external corrective action.\n","from":"developer"},{"body":"I plan to take a simple approach to this problem. If a data node finds out that it can not write to its disk. It reports to the name node and aborts. The name node logs the error and alerts it on the http UI.\n\nI will use the existing data node protocol \"errorReport\" for the error reporting but with a minor change. In addition to the parameter error message, the rpc will also send an error code so that the name node does not need to parse the error message to figure out which action to take.","from":"developer"},{"body":"Sounds good - except, what \"aborts\"? The idea of the datanode staying operating, but reporting error and not accepting further blocks is probably better, but maybe you meant \"abort the block write\". The node's blocks should probably not be counted by the namenode, but still available as a source for replication. Also, staying up means that there are fewer timeouts - it used to be that, when writing large volumes into DFS, if one or more of your nodes was full, your writer would hit a periodic timeout as connections to the (full and constantly restarting) datanode were refused. Hitting a timeout because some fraction of all resources is overused is, of course, much *much* slower than continuing to stream. Further - if the datanode periodically re-tests if the error condition has lifted, it can more immediately begin contributing to the cluster productivity again.","from":"developer"},{"body":"We are trying to deal with the case that the node is misconfigured / broken. Trying to operate in these situations is hard. Simpler to fail fast, IMO. This leverages the designed strengths of HDFS. Our goal is to get the information to the operator so they can diagnose and fix the problem and seal the problem off from the cluster.\n\nThis is distinct from the case that the node is simply full. That would not trigger this condition.","from":"developer"},{"body":"In this patch, if a data node finds that its data directory becomes not readable or writable, it logs the error and reports the problem to its namen ode and shut down itself. When the name node receives the error report, it lots the error and removes the data node info.\n\nA data node detects disk problem at startup time, when it receives a r/w request, after it receives a command from its name node, and before it sends out a block report. A data node will not start up if its data dir is not readable or writable. ","from":"developer"},{"body":"This looks great! I just committed it. Thanks, Hairong!","from":"developer"}],"created":"2006-04-25T04:52:02.000+0000","description":"I observed that sometime, if a file of a data node is not mounted properly, it may not be writable. In this case, any data writes will fail. The name node should stop assigning new blocks to that data node. The webpage should show that node is in an abnormal state.\n\n","issue_id":"12332617","key":"HADOOP-163","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-05-27T05:42:26.000+0000","role":"fixed_distractor","summary":"If a DFS datanode cannot write onto its file system. it should tell the name node not to assign new blocks to it."} {"case_id":"13232730","cluster":"DISTRACTOR-HADOOP-16308","comments":[{"body":"{code}\r\n019-05-10 17:43:28,907 [Time-limited test] ERROR contract.ContractTestUtils (ContractTestUtils.java:cleanup(380)) - Error deleting in TEARDOWN - /fork-0005/test: java.io.IOException: s3a://hwdev-steve-ireland-new: FileSystem is closed!\r\njava.io.IOException: s3a://hwdev-steve-ireland-new: FileSystem is closed!\r\n\tat org.apache.hadoop.fs.s3a.S3AFileSystem.checkNotClosed(S3AFileSystem.java:2806)\r\n\tat org.apache.hadoop.fs.s3a.S3AFileSystem.entryPoint(S3AFileSystem.java:1385)\r\n\tat org.apache.hadoop.fs.s3a.S3AFileSystem.exists(S3AFileSystem.java:3344)\r\n\tat org.apache.hadoop.fs.contract.ContractTestUtils.rm(ContractTestUtils.java:401)\r\n\tat org.apache.hadoop.fs.contract.ContractTestUtils.cleanup(ContractTestUtils.java:378)\r\n\tat org.apache.hadoop.fs.contract.AbstractFSContractTestBase.deleteTestDirInTeardown(AbstractFSContractTestBase.java:213)\r\n\tat org.apache.hadoop.fs.contract.AbstractFSContractTestBase.teardown(AbstractFSContractTestBase.java:204)\r\n\tat org.apache.hadoop.fs.contract.AbstractContractSeekTest.teardown(AbstractContractSeekTest.java:79)\r\n\tat org.apache.hadoop.fs.contract.s3a.ITestS3AContractSeek.teardown(ITestS3AContractSeek.java:131)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n\tat java.lang.reflect.Method.invoke(Method.java:498)\r\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:50)\r\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\r\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:47)\r\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:33)\r\n\tat org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:55)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:298)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:292)\r\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n\tat java.lang.Thread.run(Thread.java:748)\r\n\r\n{code}","created":"2019-05-10T19:28:44.720+0000"},{"body":"found when looking through the test logs during the HADOOP-16117 process","created":"2019-05-10T19:30:40.781+0000"}],"conversations":[{"body":"the cleanup for the ITestS3AContractSeek now adds a stack trace to the logs warning that the getFileSystem() Fs has been closed. Looks like it came from the parquet seek fix patch","from":"reporter","subject":"ITestS3AContractSeek teardown closes test FS before superclass can do its cleanup"},{"body":"{code}\r\n019-05-10 17:43:28,907 [Time-limited test] ERROR contract.ContractTestUtils (ContractTestUtils.java:cleanup(380)) - Error deleting in TEARDOWN - /fork-0005/test: java.io.IOException: s3a://hwdev-steve-ireland-new: FileSystem is closed!\r\njava.io.IOException: s3a://hwdev-steve-ireland-new: FileSystem is closed!\r\n\tat org.apache.hadoop.fs.s3a.S3AFileSystem.checkNotClosed(S3AFileSystem.java:2806)\r\n\tat org.apache.hadoop.fs.s3a.S3AFileSystem.entryPoint(S3AFileSystem.java:1385)\r\n\tat org.apache.hadoop.fs.s3a.S3AFileSystem.exists(S3AFileSystem.java:3344)\r\n\tat org.apache.hadoop.fs.contract.ContractTestUtils.rm(ContractTestUtils.java:401)\r\n\tat org.apache.hadoop.fs.contract.ContractTestUtils.cleanup(ContractTestUtils.java:378)\r\n\tat org.apache.hadoop.fs.contract.AbstractFSContractTestBase.deleteTestDirInTeardown(AbstractFSContractTestBase.java:213)\r\n\tat org.apache.hadoop.fs.contract.AbstractFSContractTestBase.teardown(AbstractFSContractTestBase.java:204)\r\n\tat org.apache.hadoop.fs.contract.AbstractContractSeekTest.teardown(AbstractContractSeekTest.java:79)\r\n\tat org.apache.hadoop.fs.contract.s3a.ITestS3AContractSeek.teardown(ITestS3AContractSeek.java:131)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n\tat java.lang.reflect.Method.invoke(Method.java:498)\r\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:50)\r\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\r\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:47)\r\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:33)\r\n\tat org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:55)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:298)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:292)\r\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n\tat java.lang.Thread.run(Thread.java:748)\r\n\r\n{code}","from":"developer"},{"body":"found when looking through the test logs during the HADOOP-16117 process","from":"developer"}],"created":"2019-05-10T19:28:19.000+0000","description":"the cleanup for the ITestS3AContractSeek now adds a stack trace to the logs warning that the getFileSystem() Fs has been closed. Looks like it came from the parquet seek fix patch","issue_id":"13232730","key":"HADOOP-16308","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2026-02-05T17:03:01.000+0000","role":"fixed_distractor","summary":"ITestS3AContractSeek teardown closes test FS before superclass can do its cleanup"} {"case_id":"13236993","cluster":"DISTRACTOR-HADOOP-16341","comments":[{"body":"Build failed for \r\n\r\n{code}\r\n[ERROR] bower qunit#1.19.0 ECMDERR Failed to execute \"git clone https://github.com/jquery/qunit.git -b 1.19.0 --progress . --depth 1\", exit code of #128 Cloning into '.'... fatal: unable to access 'https://github.com/jquery/qunit.git/': GnuTLS recv error (-54): Error in the pull function.\r\n[ERROR] \r\n[ERROR] Additional error details:\r\n[ERROR] Cloning into '.'...\r\n[ERROR] fatal: unable to access 'https://github.com/jquery/qunit.git/': GnuTLS recv error (-54): Error in the pull function.\r\n[INFO] ------------------------------------------------------------------------\r\n{code}\r\n","created":"2019-06-01T04:25:21.864+0000"},{"body":"I see: its the parsing of the default values which is causing the delay. Well spotted.\r\n\r\nWe'll take github PRs now; that's probably a better way of submitting things","created":"2019-06-03T08:41:30.576+0000"},{"body":"Thanks, opened - https://github.com/apache/hadoop/pull/896","created":"2019-06-03T18:44:15.885+0000"},{"body":"[~gopalv] [~stevel@apache.org] This issue is open and i want to know why the patch was committed to trunk and other branches.\r\nDid you discuss it in other places?","created":"2019-08-21T09:59:34.426+0000"},{"body":"updated the fix version as 3.1.3 released","created":"2019-09-25T12:53:30.902+0000"},{"body":"The hadoop release 3.1.4 code freeze is today (https://cwiki.apache.org/confluence/display/HADOOP/Roadmap).\r\nPlease close this issue today or move it to a different fix version.\r\nThank you!","created":"2020-04-22T08:07:37.992+0000"},{"body":"I'm updating the fix versions and closing this based on the git log.","created":"2020-04-23T02:48:19.406+0000"}],"conversations":[{"body":" !shutdown-hook-removal.png! ","from":"reporter","subject":"ShutDownHookManager: Regressed performance on Hook removals after HADOOP-15679"},{"body":"Build failed for \r\n\r\n{code}\r\n[ERROR] bower qunit#1.19.0 ECMDERR Failed to execute \"git clone https://github.com/jquery/qunit.git -b 1.19.0 --progress . --depth 1\", exit code of #128 Cloning into '.'... fatal: unable to access 'https://github.com/jquery/qunit.git/': GnuTLS recv error (-54): Error in the pull function.\r\n[ERROR] \r\n[ERROR] Additional error details:\r\n[ERROR] Cloning into '.'...\r\n[ERROR] fatal: unable to access 'https://github.com/jquery/qunit.git/': GnuTLS recv error (-54): Error in the pull function.\r\n[INFO] ------------------------------------------------------------------------\r\n{code}\r\n","from":"developer"},{"body":"I see: its the parsing of the default values which is causing the delay. Well spotted.\r\n\r\nWe'll take github PRs now; that's probably a better way of submitting things","from":"developer"},{"body":"Thanks, opened - https://github.com/apache/hadoop/pull/896","from":"developer"},{"body":"[~gopalv] [~stevel@apache.org] This issue is open and i want to know why the patch was committed to trunk and other branches.\r\nDid you discuss it in other places?","from":"developer"},{"body":"updated the fix version as 3.1.3 released","from":"developer"},{"body":"The hadoop release 3.1.4 code freeze is today (https://cwiki.apache.org/confluence/display/HADOOP/Roadmap).\r\nPlease close this issue today or move it to a different fix version.\r\nThank you!","from":"developer"},{"body":"I'm updating the fix versions and closing this based on the git log.","from":"developer"}],"created":"2019-05-31T23:16:38.000+0000","description":" !shutdown-hook-removal.png! ","issue_id":"13236993","key":"HADOOP-16341","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2020-04-23T02:48:19.000+0000","role":"fixed_distractor","summary":"ShutDownHookManager: Regressed performance on Hook removals after HADOOP-15679"} {"case_id":"12374268","cluster":"DISTRACTOR-HADOOP-1642","comments":[{"body":"I have seen unit tests fail with this message on recent trunk versions. I don't think the bug is specific to JobControl, but is rather with LocalJobRunner.","created":"2007-07-20T17:22:27.017+0000"},{"body":"This should fix this. Can you please try this, Johan?","created":"2007-10-18T20:01:47.219+0000"},{"body":"I applied the patch to the 0.15 branch and reran my program, this time it got a bit further. The above error didn't show up.\nInstead i got this:\n\njava.io.IOException: Target /tmp/hadoop-johan/mapred/local/map_0000/file.out already exists\n at org.apache.hadoop.fs.FileUtil.checkDest(FileUtil.java:246)\n at org.apache.hadoop.fs.FileUtil.copy(FileUtil.java:125)\n at org.apache.hadoop.fs.FileUtil.copy(FileUtil.java:116)\n at org.apache.hadoop.fs.RawLocalFileSystem.rename(RawLocalFileSystem.java:180)\n at org.apache.hadoop.fs.ChecksumFileSystem.rename(ChecksumFileSystem.java:380)\n at org.apache.hadoop.mapred.MapTask$MapOutputBuffer.mergeParts(MapTask.java:501)\n at org.apache.hadoop.mapred.MapTask$MapOutputBuffer.flush(MapTask.java:607)\n at org.apache.hadoop.mapred.MapTask.run(MapTask.java:193)\n at org.apache.hadoop.mapred.LocalJobRunner$Job.run(LocalJobRunner.java:133)\n","created":"2007-10-19T10:43:18.914+0000"},{"body":"Here's another try. Does this work for you, Johan?","created":"2007-10-19T18:34:03.023+0000"},{"body":"I just committed this. Thanks, Doug!","created":"2007-10-25T17:56:58.487+0000"},{"body":"I just checked out the 0.15 branch and I don't see the code mentioned in the patch commited there? \n\nAnyway, I applied the most recent patch, rebuilt hadoop and then put the new jar files in our project. I get the same behaviour as Johan:\n\n11-01 10:57:02] Thread-1 (LocalJobRunner.java:186) - job_local_2\njava.io.IOException: Target /tmp/hadoop-adrian/mapred/local/map_0000/file.out already exists\n\tat org.apache.hadoop.fs.FileUtil.checkDest(FileUtil.java:246)\n\tat org.apache.hadoop.fs.FileUtil.copy(FileUtil.java:125)\n\tat org.apache.hadoop.fs.FileUtil.copy(FileUtil.java:116)\n\tat org.apache.hadoop.fs.RawLocalFileSystem.rename(RawLocalFileSystem.java:180)\n\tat org.apache.hadoop.fs.ChecksumFileSystem.rename(ChecksumFileSystem.java:394)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.mergeParts(MapTask.java:501)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.flush(MapTask.java:607)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:193)\n\tat org.apache.hadoop.mapred.LocalJobRunner$Job.run(LocalJobRunner.java:132)\n\nSo I don't think this issue should be marked as resolved.","created":"2007-11-01T11:34:27.513+0000"},{"body":"Re-opening, since this does not yet appear to be fixed.","created":"2007-11-02T19:46:43.030+0000"},{"body":"Code Review please. \n\nThe actual change to the source code is very small, just a one liner modification on line 117 of LocalJobRunner.java like so:\n\n String mapId = jobId + \"_map_\" + idFormat.format(i);\n\nThe problem was that in local MR and FS mode, the mapId wasn't unique enough and caused tasks to create .out files with the same name. Prepending the jobId as shown above fixes this. Maybe there is a better way to do this that is more consistent with how it is done in the distributed Job Runner? Or using a GUID of some sort? The above worked for me but if you have a preference for another way of generating the mapId I'm keen to hear it.\n\nThe bulk of the patch consists of adding a Unit test for this. What I have done is create a TestLocalJobControl which extends HadoopTestCase and sets up a mini-mr cluster in Local MR and File mode. I based the test on TestJobControl (which is a bit of a weird unit test IMHO as it does no asserts, surely it should check the job results?) . Anyway, I factored out all the common functionality into a JobControlTestUtils class and changed both tests to use this. My test failed with the \n\njava.io.IOException: Target build/test/mapred/local/map_0000/file.out\n\nmessage until the change to LocalJobRunner was made so I think it covers this issue.","created":"2007-11-05T12:47:01.519+0000"},{"body":"Patch containing fix and unit test added (HADOOP-1642-3.patch)","created":"2007-11-05T12:48:32.688+0000"},{"body":"Missing Apache license as comment at top of new unit test classes","created":"2007-11-05T18:44:39.293+0000"},{"body":"Code review, see comments for previous patch (HADOOP-1642-3.patch), all that has changed here is that I have added Apache license in comments at top of new files in accordance with license.","created":"2007-11-05T18:45:56.842+0000"},{"body":"Can anyone tell why hudson couldn't apply the patch? I generated it according to the hadoop wiki docs and can apply it just fine to the trunk myself...","created":"2007-11-06T09:52:41.626+0000"},{"body":"bq. Can anyone tell why hudson couldn't apply the patch?\n\nHere's what I see when I attempt to apply this to trunk:\n{noformat}\n % patch -p 0 < ~/Desktop/HADOOP-1642-4.patch \npatching file src/test/org/apache/hadoop/mapred/jobcontrol/TestLocalJobControl.java\npatching file src/test/org/apache/hadoop/mapred/jobcontrol/TestJobControl.java\nHunk #1 FAILED at 18.\nHunk #2 succeeded at 47 (offset 3 lines).\nHunk #3 succeeded at 90 (offset 3 lines).\nHunk #4 succeeded at 106 (offset 3 lines).\nHunk #5 succeeded at 125 (offset 3 lines).\n1 out of 5 hunks FAILED -- saving rejects to file src/test/org/apache/hadoop/mapred/jobcontrol/TestJobControl.java.rej\npatching file src/test/org/apache/hadoop/mapred/jobcontrol/JobControlTestUtils.java\npatching file src/java/org/apache/hadoop/mapred/LocalJobRunner.java\nHunk #2 FAILED at 268.\n1 out of 2 hunks FAILED -- saving rejects to file src/java/org/apache/hadoop/mapred/LocalJobRunner.java.rej\n{noformat}","created":"2007-11-06T17:12:59.996+0000"},{"body":"OK, I patched it using Eclipse and it worked. I just tried from command line and I get the same as you. Does it have to work from command line? If so I will look into it and try figure out what the issue is....","created":"2007-11-06T17:35:07.168+0000"},{"body":"Have you recently updated your Eclipse workspace to the current trunk?\n\n> Does it have to work from command line?\n\nYes.\n","created":"2007-11-06T17:53:36.783+0000"},{"body":"Attempting patch regeneration from trunk","created":"2007-11-06T18:21:09.137+0000"},{"body":"Same patch as before, but this time generated against a fresh checkout of the trunk. Works for me from the command line. All comments above for HADOOP-1642-3.patch still apply.","created":"2007-11-06T18:22:46.430+0000"},{"body":"Should work from command line now.","created":"2007-11-06T18:45:18.278+0000"},{"body":"bq. I just checked out the 0.15 branch and I don't see the code mentioned in the patch commited there?\n\nAdrian, the original fix went into 0.16.0 (trunk), could you check if it is fixed there? If not, please open another issue and attach your current patch, else we can debate about reverting Doug's original patch from 0.16.0. Thanks!","created":"2007-11-10T06:59:56.056+0000"},{"body":"Yes, the original fix is indeed in the trunk, the patch I have submitted was also built against the trunk and requires Doug's original patch and should be applied in addition to it as it solves the second issue which arose after the original patch. Can someone please code review it and see if it is acceptable?","created":"2007-11-12T10:08:01.078+0000"},{"body":"Resolving this and taking the issue forward in HADOOP-2245.","created":"2007-11-21T07:13:36.516+0000"}],"conversations":[{"body":"If I run several jobs at the same time using JobControl and the LocalJobRunner i get:\njava.io.IOException: Target /tmp/hadoop-johan/mapred/local/localRunner/job_local_1.xml already exists.\n\nIt seems like the JobControl class tries to run multiple jobs with the same jobid, causing the exception.\n\n\n","from":"reporter","subject":"Jobs using LocalJobRunner + JobControl fails"},{"body":"I have seen unit tests fail with this message on recent trunk versions. I don't think the bug is specific to JobControl, but is rather with LocalJobRunner.","from":"developer"},{"body":"This should fix this. Can you please try this, Johan?","from":"developer"},{"body":"I applied the patch to the 0.15 branch and reran my program, this time it got a bit further. The above error didn't show up.\nInstead i got this:\n\njava.io.IOException: Target /tmp/hadoop-johan/mapred/local/map_0000/file.out already exists\n at org.apache.hadoop.fs.FileUtil.checkDest(FileUtil.java:246)\n at org.apache.hadoop.fs.FileUtil.copy(FileUtil.java:125)\n at org.apache.hadoop.fs.FileUtil.copy(FileUtil.java:116)\n at org.apache.hadoop.fs.RawLocalFileSystem.rename(RawLocalFileSystem.java:180)\n at org.apache.hadoop.fs.ChecksumFileSystem.rename(ChecksumFileSystem.java:380)\n at org.apache.hadoop.mapred.MapTask$MapOutputBuffer.mergeParts(MapTask.java:501)\n at org.apache.hadoop.mapred.MapTask$MapOutputBuffer.flush(MapTask.java:607)\n at org.apache.hadoop.mapred.MapTask.run(MapTask.java:193)\n at org.apache.hadoop.mapred.LocalJobRunner$Job.run(LocalJobRunner.java:133)\n","from":"developer"},{"body":"Here's another try. Does this work for you, Johan?","from":"developer"},{"body":"I just committed this. Thanks, Doug!","from":"developer"},{"body":"I just checked out the 0.15 branch and I don't see the code mentioned in the patch commited there? \n\nAnyway, I applied the most recent patch, rebuilt hadoop and then put the new jar files in our project. I get the same behaviour as Johan:\n\n11-01 10:57:02] Thread-1 (LocalJobRunner.java:186) - job_local_2\njava.io.IOException: Target /tmp/hadoop-adrian/mapred/local/map_0000/file.out already exists\n\tat org.apache.hadoop.fs.FileUtil.checkDest(FileUtil.java:246)\n\tat org.apache.hadoop.fs.FileUtil.copy(FileUtil.java:125)\n\tat org.apache.hadoop.fs.FileUtil.copy(FileUtil.java:116)\n\tat org.apache.hadoop.fs.RawLocalFileSystem.rename(RawLocalFileSystem.java:180)\n\tat org.apache.hadoop.fs.ChecksumFileSystem.rename(ChecksumFileSystem.java:394)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.mergeParts(MapTask.java:501)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.flush(MapTask.java:607)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:193)\n\tat org.apache.hadoop.mapred.LocalJobRunner$Job.run(LocalJobRunner.java:132)\n\nSo I don't think this issue should be marked as resolved.","from":"developer"},{"body":"Re-opening, since this does not yet appear to be fixed.","from":"developer"},{"body":"Code Review please. \n\nThe actual change to the source code is very small, just a one liner modification on line 117 of LocalJobRunner.java like so:\n\n String mapId = jobId + \"_map_\" + idFormat.format(i);\n\nThe problem was that in local MR and FS mode, the mapId wasn't unique enough and caused tasks to create .out files with the same name. Prepending the jobId as shown above fixes this. Maybe there is a better way to do this that is more consistent with how it is done in the distributed Job Runner? Or using a GUID of some sort? The above worked for me but if you have a preference for another way of generating the mapId I'm keen to hear it.\n\nThe bulk of the patch consists of adding a Unit test for this. What I have done is create a TestLocalJobControl which extends HadoopTestCase and sets up a mini-mr cluster in Local MR and File mode. I based the test on TestJobControl (which is a bit of a weird unit test IMHO as it does no asserts, surely it should check the job results?) . Anyway, I factored out all the common functionality into a JobControlTestUtils class and changed both tests to use this. My test failed with the \n\njava.io.IOException: Target build/test/mapred/local/map_0000/file.out\n\nmessage until the change to LocalJobRunner was made so I think it covers this issue.","from":"developer"},{"body":"Patch containing fix and unit test added (HADOOP-1642-3.patch)","from":"developer"},{"body":"Missing Apache license as comment at top of new unit test classes","from":"developer"},{"body":"Code review, see comments for previous patch (HADOOP-1642-3.patch), all that has changed here is that I have added Apache license in comments at top of new files in accordance with license.","from":"developer"},{"body":"Can anyone tell why hudson couldn't apply the patch? I generated it according to the hadoop wiki docs and can apply it just fine to the trunk myself...","from":"developer"},{"body":"bq. Can anyone tell why hudson couldn't apply the patch?\n\nHere's what I see when I attempt to apply this to trunk:\n{noformat}\n % patch -p 0 < ~/Desktop/HADOOP-1642-4.patch \npatching file src/test/org/apache/hadoop/mapred/jobcontrol/TestLocalJobControl.java\npatching file src/test/org/apache/hadoop/mapred/jobcontrol/TestJobControl.java\nHunk #1 FAILED at 18.\nHunk #2 succeeded at 47 (offset 3 lines).\nHunk #3 succeeded at 90 (offset 3 lines).\nHunk #4 succeeded at 106 (offset 3 lines).\nHunk #5 succeeded at 125 (offset 3 lines).\n1 out of 5 hunks FAILED -- saving rejects to file src/test/org/apache/hadoop/mapred/jobcontrol/TestJobControl.java.rej\npatching file src/test/org/apache/hadoop/mapred/jobcontrol/JobControlTestUtils.java\npatching file src/java/org/apache/hadoop/mapred/LocalJobRunner.java\nHunk #2 FAILED at 268.\n1 out of 2 hunks FAILED -- saving rejects to file src/java/org/apache/hadoop/mapred/LocalJobRunner.java.rej\n{noformat}","from":"developer"},{"body":"OK, I patched it using Eclipse and it worked. I just tried from command line and I get the same as you. Does it have to work from command line? If so I will look into it and try figure out what the issue is....","from":"developer"},{"body":"Have you recently updated your Eclipse workspace to the current trunk?\n\n> Does it have to work from command line?\n\nYes.\n","from":"developer"},{"body":"Attempting patch regeneration from trunk","from":"developer"},{"body":"Same patch as before, but this time generated against a fresh checkout of the trunk. Works for me from the command line. All comments above for HADOOP-1642-3.patch still apply.","from":"developer"},{"body":"Should work from command line now.","from":"developer"},{"body":"bq. I just checked out the 0.15 branch and I don't see the code mentioned in the patch commited there?\n\nAdrian, the original fix went into 0.16.0 (trunk), could you check if it is fixed there? If not, please open another issue and attach your current patch, else we can debate about reverting Doug's original patch from 0.16.0. Thanks!","from":"developer"},{"body":"Yes, the original fix is indeed in the trunk, the patch I have submitted was also built against the trunk and requires Doug's original patch and should be applied in addition to it as it solves the second issue which arose after the original patch. Can someone please code review it and see if it is acceptable?","from":"developer"},{"body":"Resolving this and taking the issue forward in HADOOP-2245.","from":"developer"}],"created":"2007-07-20T15:57:01.000+0000","description":"If I run several jobs at the same time using JobControl and the LocalJobRunner i get:\njava.io.IOException: Target /tmp/hadoop-johan/mapred/local/localRunner/job_local_1.xml already exists.\n\nIt seems like the JobControl class tries to run multiple jobs with the same jobid, causing the exception.\n\n\n","issue_id":"12374268","key":"HADOOP-1642","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2007-11-21T07:13:36.000+0000","role":"fixed_distractor","summary":"Jobs using LocalJobRunner + JobControl fails"} {"case_id":"13249240","cluster":"DISTRACTOR-HADOOP-16494","comments":[{"body":"As [~jeagles] commented, Hadoop project release checker is available (https://checker.apache.org/projs/hadoop.html) and this page should be checked after each release is made.","created":"2019-08-07T02:31:59.891+0000"},{"body":"Committed this to trunk, branch-3.2, branch-3.1, branch-2, branch-2.9, and branch-2.8.","created":"2019-08-22T02:07:31.097+0000"},{"body":"[~aajisaka] I created release artifacts for 3.2.1 and trying to verify sha512. I see below error \r\n{noformat}\r\nrohithsharmaks@ip-172-31-38-26:~/branch-3.2.1/target/artifacts$ sha512sum -c CHANGELOG.md.sha512\r\nsha512sum: /build/source/target/artifacts/CHANGELOG.md: No such file or directory\r\n/build/source/target/artifacts/CHANGELOG.md: FAILED open or read\r\nsha512sum: WARNING: 1 listed file could not be read\r\n{noformat}\r\n\r\nWhen sha512 is created, it is taking entire folder path into account because *\"${i}\"* is entire folder patch. \r\n{noformat}\r\n for i in ${ARTIFACTS_DIR}/*; do\r\n ${GPG} --use-agent --armor --output \"${i}.asc\" --detach-sig \"${i}\"\r\n sha512sum --tag \"${i}\" > \"${i}.sha512\"\r\n done\r\n{noformat}\r\n\r\nIs this expected? or need fix to consider only file name. How was earlier md5 was running?","created":"2019-09-11T04:21:44.804+0000"},{"body":"Just checked previous releases where same as above folder path is in md5 files. Looks it is fine.","created":"2019-09-11T04:25:43.307+0000"},{"body":"[~rohithsharma] While that isn't a new thing, it is probably still going to be a source of frustration, since it is very unlikely that whoever is verifying the file is going to be doing so in a path that beings with {{/build/source/target}}, and therefore, it might be a good idea to address this in a new issue. Do you think?","created":"2019-09-11T05:11:00.310+0000"},{"body":"bq. it might be a good idea to address this in a new issue. Do you think?\r\nI am +1 for addressing this.","created":"2019-09-11T12:29:18.914+0000"}],"conversations":[{"body":"Originally reported by [~ctubbsii]: https://lists.apache.org/thread.html/db2f5d5d8600c405293ebfb3bfc415e200e59f72605c5a920a461c09@%3Cgeneral.hadoop.apache.org%3E\r\n\r\nbq. None of the artifacts seem to have valid detached checksum files that are in compliance with https://www.apache.org/dev/release-distribution There should be some \".shaXXX\" files in there, and not just the (optional) \".mds\" files.","from":"reporter","subject":"Add SHA-256 or SHA-512 checksum to release artifacts to comply with the release distribution policy"},{"body":"As [~jeagles] commented, Hadoop project release checker is available (https://checker.apache.org/projs/hadoop.html) and this page should be checked after each release is made.","from":"developer"},{"body":"Committed this to trunk, branch-3.2, branch-3.1, branch-2, branch-2.9, and branch-2.8.","from":"developer"},{"body":"[~aajisaka] I created release artifacts for 3.2.1 and trying to verify sha512. I see below error \r\n{noformat}\r\nrohithsharmaks@ip-172-31-38-26:~/branch-3.2.1/target/artifacts$ sha512sum -c CHANGELOG.md.sha512\r\nsha512sum: /build/source/target/artifacts/CHANGELOG.md: No such file or directory\r\n/build/source/target/artifacts/CHANGELOG.md: FAILED open or read\r\nsha512sum: WARNING: 1 listed file could not be read\r\n{noformat}\r\n\r\nWhen sha512 is created, it is taking entire folder path into account because *\"${i}\"* is entire folder patch. \r\n{noformat}\r\n for i in ${ARTIFACTS_DIR}/*; do\r\n ${GPG} --use-agent --armor --output \"${i}.asc\" --detach-sig \"${i}\"\r\n sha512sum --tag \"${i}\" > \"${i}.sha512\"\r\n done\r\n{noformat}\r\n\r\nIs this expected? or need fix to consider only file name. How was earlier md5 was running?","from":"developer"},{"body":"Just checked previous releases where same as above folder path is in md5 files. Looks it is fine.","from":"developer"},{"body":"[~rohithsharma] While that isn't a new thing, it is probably still going to be a source of frustration, since it is very unlikely that whoever is verifying the file is going to be doing so in a path that beings with {{/build/source/target}}, and therefore, it might be a good idea to address this in a new issue. Do you think?","from":"developer"},{"body":"bq. it might be a good idea to address this in a new issue. Do you think?\r\nI am +1 for addressing this.","from":"developer"}],"created":"2019-08-07T02:25:27.000+0000","description":"Originally reported by [~ctubbsii]: https://lists.apache.org/thread.html/db2f5d5d8600c405293ebfb3bfc415e200e59f72605c5a920a461c09@%3Cgeneral.hadoop.apache.org%3E\r\n\r\nbq. None of the artifacts seem to have valid detached checksum files that are in compliance with https://www.apache.org/dev/release-distribution There should be some \".shaXXX\" files in there, and not just the (optional) \".mds\" files.","issue_id":"13249240","key":"HADOOP-16494","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2019-08-22T02:07:31.000+0000","role":"fixed_distractor","summary":"Add SHA-256 or SHA-512 checksum to release artifacts to comply with the release distribution policy"} {"case_id":"13259467","cluster":"DISTRACTOR-HADOOP-16614","comments":[{"body":"The patch looks good to me. I will commit if no objections.","created":"2019-10-23T15:14:46.075+0000"},{"body":"Thank you [~seanlau] for the patch.\r\n+1 merged to trunk.\r\n","created":"2019-10-24T15:51:59.966+0000"},{"body":"Do we plan to backport this to 2.7 and 3.2? Spark has the same issue on arm platform, it depends on *org.fusesource.leveldbjni* not only itself, but also it depends on some hadoop 2.7(or 3.2) jar packages which using *org.fusesource.leveldbjni*,  so it will be good that we backport this to 2.7 and 3.2 :)","created":"2019-10-25T01:34:33.581+0000"},{"body":"A workaround of this is adding JVM option \"-Dlibrary.leveldbjni.path=\" to specify the dir path where libleveldbjni.so locates.\r\n\r\nThe libleveldbjni.so built on ARM platform can be extracted from the jar:\r\n[https://mvnrepository.com/artifact/org.fusesource.leveldbjni/leveldbjni-all/1.8-hw-20191105]","created":"2023-04-03T00:08:37.953+0000"}],"conversations":[{"body":"Currently, Hadoop denpend on the *leveldbjni-all:1.8* package of *org.fusesource.leveldbjni* group, but it cannot support ARM platform.\r\n\r\nsee: [https://search.maven.org/search?q=g:org.fusesource.leveldbjni]\r\n\r\nBecause the leveldbjni community is inactivity and the  code ([https://github.com/fusesource/leveldbjni]) didn't updated a long time.I will build the leveldbjni package of aarch64 platform, and upload it with other platform packages of *org.fusesource.leveldbjni* to a new *org.openlabtesting.leveldbjni* maven repo. In hadoop code, I will add a new profile aarch64 for for automatically select the *org.openlabtesting.leveldbjni* artifact group and using the aarch64 package of leveldbjni when running on ARM server, this approach has no effect on current code.","from":"reporter","subject":"Missing leveldbjni package of aarch64 platform"},{"body":"The patch looks good to me. I will commit if no objections.","from":"developer"},{"body":"Thank you [~seanlau] for the patch.\r\n+1 merged to trunk.\r\n","from":"developer"},{"body":"Do we plan to backport this to 2.7 and 3.2? Spark has the same issue on arm platform, it depends on *org.fusesource.leveldbjni* not only itself, but also it depends on some hadoop 2.7(or 3.2) jar packages which using *org.fusesource.leveldbjni*,  so it will be good that we backport this to 2.7 and 3.2 :)","from":"developer"},{"body":"A workaround of this is adding JVM option \"-Dlibrary.leveldbjni.path=\" to specify the dir path where libleveldbjni.so locates.\r\n\r\nThe libleveldbjni.so built on ARM platform can be extracted from the jar:\r\n[https://mvnrepository.com/artifact/org.fusesource.leveldbjni/leveldbjni-all/1.8-hw-20191105]","from":"developer"}],"created":"2019-09-29T02:53:13.000+0000","description":"Currently, Hadoop denpend on the *leveldbjni-all:1.8* package of *org.fusesource.leveldbjni* group, but it cannot support ARM platform.\r\n\r\nsee: [https://search.maven.org/search?q=g:org.fusesource.leveldbjni]\r\n\r\nBecause the leveldbjni community is inactivity and the  code ([https://github.com/fusesource/leveldbjni]) didn't updated a long time.I will build the leveldbjni package of aarch64 platform, and upload it with other platform packages of *org.fusesource.leveldbjni* to a new *org.openlabtesting.leveldbjni* maven repo. In hadoop code, I will add a new profile aarch64 for for automatically select the *org.openlabtesting.leveldbjni* artifact group and using the aarch64 package of leveldbjni when running on ARM server, this approach has no effect on current code.","issue_id":"13259467","key":"HADOOP-16614","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2019-10-24T15:51:59.000+0000","role":"fixed_distractor","summary":"Missing leveldbjni package of aarch64 platform"} {"case_id":"13302792","cluster":"DISTRACTOR-HADOOP-17028","comments":[{"body":"this would be good. All the object stores can be slow.\r\n\r\nFWIW, if you set fs.s3a.bucket.probe = 0 we skip that check for a bucket, but s3guard still kicks off its conversation with DynamoDB","created":"2020-05-05T13:10:58.638+0000"},{"body":"The filesystems for both the non-leaf nodes (InternalDirOfViewFs) and leaf nodes (ChRootedFileSystem) gets constructed during the viewfs initialize phase. During this the fs object gets created though FsGetter. Here is my rough idea about the implementation.\r\n * Make sure the FsGetter.getNewInstance() doesn't initialize the FileSystem object when asked for.\r\n * For InternalDirOfViewFs, the initialize is called at the constructor but this call is calling base initailize method.\r\n * For ChRootedFileSystem, if the constructor gets invoked through ChRootedFileSystem(final URI uri, Configuration conf) then it gets the fs object from FileSystem.get(uri, conf) which will initalize the fs object but this constructor is not used except in tests, so we are good. The other constructor gets the fs object as argument, so we have to make sure caller wont initialize the fs object before invoking this constructor.\r\n * When a FileSystem api gets invoked for ChRootedFileSystem, before calling the actual implementation of the underlying fs object through FilterFileSystem, it can check whether the fs object has been initialized, so that the fs object gets initialized only once. We can tap at ChRootedFileSystem.fullPath(path) to check the initialization (through a class level variable)\r\n\r\n \r\n\r\n[~umamaheswararao] let me know your thoughts about the approach. I can start working on this.","created":"2020-06-07T00:30:08.943+0000"},{"body":"PR for this: [https://github.com/apache/hadoop/pull/2260]\r\n\r\n[~umamaheswararao] can you please review the change when you get a chance. Thanks in advance","created":"2020-08-29T01:28:11.222+0000"},{"body":"Thank you [~abhishekd] for the patch. I have this in my list, I will provide my feedback soon. Thank you.","created":"2020-09-01T07:02:52.208+0000"},{"body":"Left minor comments on the PR. The approach looks reasonable to me.\r\nThe important thing is to get a Jenkins build. Looks like it could not run. May be you fork got stale. You should probably rebase on current trunk.","created":"2021-06-03T00:49:53.142+0000"},{"body":"I'll re-post here my comment from the linked PR for visibility.\r\nIt looks like [~stevel@apache.org] deprecated this class later after his comment. Don't know what prompted the refactoring, but\r\n{{org.apache.hadoop.fs.impl.FunctionsRaisingIOE.FunctionRaisingIOE}}\r\nshould be now\r\n{{org.apache.hadoop.util.functional.FunctionRaisingIOE}}\r\n\r\nMoving interfaces, which are public by default, from one package to another is considered an incompatible change, especially since the previous variant had been released. Besides, it has been committed only to branch-3.3 contributing to further divergence of the supported branches and making backporting yet harder.\r\nSo [~abhishekd], I prefer to avoid using {{FunctionRaisingIOE}} in this patch if possible. This will simplify the backport and avoid using unstable APIs.","created":"2021-06-04T19:19:42.171+0000"},{"body":"bq. It looks like Steve Loughran deprecated this class later after his comment. Don't know what prompted the refactoring, but\r\n{{org.apache.hadoop.fs.impl.FunctionsRaisingIOE.FunctionRaisingIOE}} should be now {{org.apache.hadoop.util.functional.FunctionRaisingIOE}}\r\n\r\nI moved it because it turns out that wrapping/unwrapping IOEs is critical to using this in applications using the FS API, and since the relevant methods/interfaces were not public, the only way to do that is to have them public. \r\n\r\nAccordingly I\r\n# replicated the functional interfaces in the public/unstable package {{org.apache.hadoop.util.functional}} where I'm trying to make it possible to use IOE-raising stuff (including RemoteIterator) in apps.\r\n# tagged the old ones, as Deprecated, so new code will use it.\r\n\r\n\r\nbq. Moving interfaces, which are public by default, from one package to another is considered an incompatible change, especially since the previous variant had been released\r\n\r\nTwo points to note\r\n# I tagged the original package \"Fs.impl\" as not public then, isn't it?\r\n# left the old interface alone, on the basis that if filesystems outside the hadoop codebase (gcs?) were using them, all would be good.\r\n\r\n{code}\r\n@InterfaceAudience.LimitedPrivate(\"Filesystems\")\r\n@InterfaceStability.Unstable\r\n{code}\r\n\r\ntherefore, I do not consider this to be an incompatible change since\r\n# it wasn't public, outside filesystems \r\n# it hasn't been removed, just deprecated.\r\n\r\nbq. I prefer to avoid using FunctionRaisingIOE in this patch if possible\r\n\r\nUse the o.a.h.fs.impl for compatibility across Hadoop 3.3+\r\n\r\nFor older releases, no, it's not there. Not sure what to do there.\r\n","created":"2021-06-07T13:22:14.871+0000"},{"body":"Steve, this looks like incompatible change as it replaced the parameter to a different class:\r\n{code}\r\n public static CompletableFuture eval(\r\n- FunctionsRaisingIOE.CallableRaisingIOE callable) {\r\n+ CallableRaisingIOE callable) {\r\n CompletableFuture result = new CompletableFuture<>();\r\n{code}\r\nAlso introduction of dead code should have been avoided. Looking at the history, you introduced {{org.apache.hadoop.fs.impl.FunctionsRaisingIOE.FunctionRaisingIOE}} as a part of HADOOP-15183. But it wasn't used anywhere. Then HADOOP-17450 deprecated it. One good thing about it that you can remove the deprecated classes as they have never been used.\r\nSupporting dead code is just a waste of energy it also makes back porting really hard. You don't seem to care about older versions, but many people do.","created":"2021-06-27T21:19:06.989+0000"},{"body":"bq. Steve, this looks like incompatible change as it replaced the parameter to a different class...\r\n\r\nIn a world of @Functional params, it is compatible. But I suspect that you are right about link-time matching; the JVM probably isn't going to blindly treat them as equivalent.\r\n\r\nbq. One good thing about it that you can remove the deprecated classes as they have never been used.\r\n\r\nnot used in hadoop-*. My fear was that they had been picked up/used in google GCS. I've been building that connector and I don't see it in use. \r\n\r\nGiven that I'll target it for removal in 3.3.2\r\n\r\nbq. You don't seem to care about older versions, but many people do.\r\n\r\nI really do. I am the one currently staring at the ABFS Code and a branch-2 build. At the same time, being more isolated (and, let's be ruthless: lower risk), the object storage code has been able to evolve faster than other bits of the codebase. \r\n\r\nFWIW the main troublespot in backporting cloud storage changes, esp in hadoop-azure, is actually mockito versions. something which works on mockito 2 is still likely to compile on mockito 1.x but then fail with some \"impossible\" stack trace, and I'm left trying to distinguish between \"mockito lib version issues, fix test or cut\", \"different codepath breaks mockito test\" and \"test has actually found a regression\". It's a key reason why I don't like mockito-based testing. ","created":"2021-06-28T11:00:23.129+0000"},{"body":"I committed this to trunk and branches 3.1, 3.2, 3.3. Thanks [~abhishekd] for working on it.\r\nWill need a separate patch for branch-2.10. Too many conflicts.","created":"2021-07-14T01:32:20.661+0000"},{"body":"Committed PR #3218 to branch-2.10.\r\nThank you [~abhishekd]","created":"2021-07-21T01:28:34.698+0000"}],"conversations":[{"body":"Currently viewFS initialize all configured target filesystems when viewfs#init itself.\r\n\r\nSome target file system initialization involve creating heavy objects and proxy connections. Ex: DistributedFileSystem#initialize will create DFSClient object which will create proxy connections to NN etc.\r\nFor example: if ViewFS configured with 10 target fs with hdfs uri and 2 targets with s3a.\r\n\r\nIf one of the client only work with s3a target, But ViewFS will initialize all targets irrespective of what clients interested to work with. That means, here client will create 10 DFS initializations and 2 s3a initializations. Its unnecessary to have DFS initialization here. So, it will be a good idea to initialize the target fs only when first time usage call come to particular target fs scheme. ","from":"reporter","subject":"ViewFS should initialize target filesystems lazily"},{"body":"this would be good. All the object stores can be slow.\r\n\r\nFWIW, if you set fs.s3a.bucket.probe = 0 we skip that check for a bucket, but s3guard still kicks off its conversation with DynamoDB","from":"developer"},{"body":"The filesystems for both the non-leaf nodes (InternalDirOfViewFs) and leaf nodes (ChRootedFileSystem) gets constructed during the viewfs initialize phase. During this the fs object gets created though FsGetter. Here is my rough idea about the implementation.\r\n * Make sure the FsGetter.getNewInstance() doesn't initialize the FileSystem object when asked for.\r\n * For InternalDirOfViewFs, the initialize is called at the constructor but this call is calling base initailize method.\r\n * For ChRootedFileSystem, if the constructor gets invoked through ChRootedFileSystem(final URI uri, Configuration conf) then it gets the fs object from FileSystem.get(uri, conf) which will initalize the fs object but this constructor is not used except in tests, so we are good. The other constructor gets the fs object as argument, so we have to make sure caller wont initialize the fs object before invoking this constructor.\r\n * When a FileSystem api gets invoked for ChRootedFileSystem, before calling the actual implementation of the underlying fs object through FilterFileSystem, it can check whether the fs object has been initialized, so that the fs object gets initialized only once. We can tap at ChRootedFileSystem.fullPath(path) to check the initialization (through a class level variable)\r\n\r\n \r\n\r\n[~umamaheswararao] let me know your thoughts about the approach. I can start working on this.","from":"developer"},{"body":"PR for this: [https://github.com/apache/hadoop/pull/2260]\r\n\r\n[~umamaheswararao] can you please review the change when you get a chance. Thanks in advance","from":"developer"},{"body":"Thank you [~abhishekd] for the patch. I have this in my list, I will provide my feedback soon. Thank you.","from":"developer"},{"body":"Left minor comments on the PR. The approach looks reasonable to me.\r\nThe important thing is to get a Jenkins build. Looks like it could not run. May be you fork got stale. You should probably rebase on current trunk.","from":"developer"},{"body":"I'll re-post here my comment from the linked PR for visibility.\r\nIt looks like [~stevel@apache.org] deprecated this class later after his comment. Don't know what prompted the refactoring, but\r\n{{org.apache.hadoop.fs.impl.FunctionsRaisingIOE.FunctionRaisingIOE}}\r\nshould be now\r\n{{org.apache.hadoop.util.functional.FunctionRaisingIOE}}\r\n\r\nMoving interfaces, which are public by default, from one package to another is considered an incompatible change, especially since the previous variant had been released. Besides, it has been committed only to branch-3.3 contributing to further divergence of the supported branches and making backporting yet harder.\r\nSo [~abhishekd], I prefer to avoid using {{FunctionRaisingIOE}} in this patch if possible. This will simplify the backport and avoid using unstable APIs.","from":"developer"},{"body":"bq. It looks like Steve Loughran deprecated this class later after his comment. Don't know what prompted the refactoring, but\r\n{{org.apache.hadoop.fs.impl.FunctionsRaisingIOE.FunctionRaisingIOE}} should be now {{org.apache.hadoop.util.functional.FunctionRaisingIOE}}\r\n\r\nI moved it because it turns out that wrapping/unwrapping IOEs is critical to using this in applications using the FS API, and since the relevant methods/interfaces were not public, the only way to do that is to have them public. \r\n\r\nAccordingly I\r\n# replicated the functional interfaces in the public/unstable package {{org.apache.hadoop.util.functional}} where I'm trying to make it possible to use IOE-raising stuff (including RemoteIterator) in apps.\r\n# tagged the old ones, as Deprecated, so new code will use it.\r\n\r\n\r\nbq. Moving interfaces, which are public by default, from one package to another is considered an incompatible change, especially since the previous variant had been released\r\n\r\nTwo points to note\r\n# I tagged the original package \"Fs.impl\" as not public then, isn't it?\r\n# left the old interface alone, on the basis that if filesystems outside the hadoop codebase (gcs?) were using them, all would be good.\r\n\r\n{code}\r\n@InterfaceAudience.LimitedPrivate(\"Filesystems\")\r\n@InterfaceStability.Unstable\r\n{code}\r\n\r\ntherefore, I do not consider this to be an incompatible change since\r\n# it wasn't public, outside filesystems \r\n# it hasn't been removed, just deprecated.\r\n\r\nbq. I prefer to avoid using FunctionRaisingIOE in this patch if possible\r\n\r\nUse the o.a.h.fs.impl for compatibility across Hadoop 3.3+\r\n\r\nFor older releases, no, it's not there. Not sure what to do there.\r\n","from":"developer"},{"body":"Steve, this looks like incompatible change as it replaced the parameter to a different class:\r\n{code}\r\n public static CompletableFuture eval(\r\n- FunctionsRaisingIOE.CallableRaisingIOE callable) {\r\n+ CallableRaisingIOE callable) {\r\n CompletableFuture result = new CompletableFuture<>();\r\n{code}\r\nAlso introduction of dead code should have been avoided. Looking at the history, you introduced {{org.apache.hadoop.fs.impl.FunctionsRaisingIOE.FunctionRaisingIOE}} as a part of HADOOP-15183. But it wasn't used anywhere. Then HADOOP-17450 deprecated it. One good thing about it that you can remove the deprecated classes as they have never been used.\r\nSupporting dead code is just a waste of energy it also makes back porting really hard. You don't seem to care about older versions, but many people do.","from":"developer"},{"body":"bq. Steve, this looks like incompatible change as it replaced the parameter to a different class...\r\n\r\nIn a world of @Functional params, it is compatible. But I suspect that you are right about link-time matching; the JVM probably isn't going to blindly treat them as equivalent.\r\n\r\nbq. One good thing about it that you can remove the deprecated classes as they have never been used.\r\n\r\nnot used in hadoop-*. My fear was that they had been picked up/used in google GCS. I've been building that connector and I don't see it in use. \r\n\r\nGiven that I'll target it for removal in 3.3.2\r\n\r\nbq. You don't seem to care about older versions, but many people do.\r\n\r\nI really do. I am the one currently staring at the ABFS Code and a branch-2 build. At the same time, being more isolated (and, let's be ruthless: lower risk), the object storage code has been able to evolve faster than other bits of the codebase. \r\n\r\nFWIW the main troublespot in backporting cloud storage changes, esp in hadoop-azure, is actually mockito versions. something which works on mockito 2 is still likely to compile on mockito 1.x but then fail with some \"impossible\" stack trace, and I'm left trying to distinguish between \"mockito lib version issues, fix test or cut\", \"different codepath breaks mockito test\" and \"test has actually found a regression\". It's a key reason why I don't like mockito-based testing. ","from":"developer"},{"body":"I committed this to trunk and branches 3.1, 3.2, 3.3. Thanks [~abhishekd] for working on it.\r\nWill need a separate patch for branch-2.10. Too many conflicts.","from":"developer"},{"body":"Committed PR #3218 to branch-2.10.\r\nThank you [~abhishekd]","from":"developer"}],"created":"2020-05-05T06:09:02.000+0000","description":"Currently viewFS initialize all configured target filesystems when viewfs#init itself.\r\n\r\nSome target file system initialization involve creating heavy objects and proxy connections. Ex: DistributedFileSystem#initialize will create DFSClient object which will create proxy connections to NN etc.\r\nFor example: if ViewFS configured with 10 target fs with hdfs uri and 2 targets with s3a.\r\n\r\nIf one of the client only work with s3a target, But ViewFS will initialize all targets irrespective of what clients interested to work with. That means, here client will create 10 DFS initializations and 2 s3a initializations. Its unnecessary to have DFS initialization here. So, it will be a good idea to initialize the target fs only when first time usage call come to particular target fs scheme. ","issue_id":"13302792","key":"HADOOP-17028","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2021-07-21T01:30:31.000+0000","role":"fixed_distractor","summary":"ViewFS should initialize target filesystems lazily"} {"case_id":"13370616","cluster":"DISTRACTOR-HADOOP-17629","comments":[{"body":"I believe the module error is a red herring, and the real error is just that `testAuthorization` doesn't expect a SocketException to be thrown here. `testErrorMsgForInsecureClient` will also fail [while trying to cast to RemoteException|https://github.com/apache/hadoop/blob/trunk/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/ipc/TestRPC.java#L826] because of the same error (although this doesn't fail as consistently as testAuthorization)\r\n\r\nMore interestingly however is that [earlier in the insecure client test case|https://github.com/apache/hadoop/blob/trunk/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/ipc/TestRPC.java#L803], it does the exact same cast to RemoteException, but clearly does so successfully. This clearly shows that the cast itself is not the error. This does however raise the issue of why a SocketException is being thrown in the first place","created":"2025-06-18T23:39:37.705+0000"},{"body":"Currently, it fails with below IOException. It was using WritableRpcEngine to run TestProtobufRpcProto.\r\n{code}\r\njava.io.IOException: protocolClass interface org.apache.hadoop.ipc.TestRpcBase$TestRpcService is not implemented by protocolImpl which is of class class org.apache.hadoop.ipc.protobuf.TestRpcServiceProtos$TestProtobufRpcProto$2\r\n\r\n\tat org.apache.hadoop.ipc.WritableRpcEngine$Server.(WritableRpcEngine.java:523)\r\n\tat org.apache.hadoop.ipc.WritableRpcEngine.getServer(WritableRpcEngine.java:379)\r\n\tat org.apache.hadoop.ipc.RPC$Builder.build(RPC.java:986)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:124)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:119)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:102)\r\n\tat org.apache.hadoop.ipc.TestRPC.doRPCs(TestRPC.java:646)\r\n\tat org.apache.hadoop.ipc.TestRPC.testAuthorization(TestRPC.java:705)\r\n\t...\r\n{code}","created":"2025-09-29T19:29:38.989+0000"},{"body":"Changing it to use ProtobufRpcEngine2 can fix it.","created":"2025-09-29T19:31:47.086+0000"},{"body":"The failure happens in both JDK 8 and JDK 17.  So it is not specific to JDK 17.","created":"2025-09-29T20:00:54.700+0000"},{"body":"The pull request is now merged.","created":"2025-09-30T15:55:29.122+0000"}],"conversations":[{"body":"Currently, it fails with below IOException. It was using WritableRpcEngine to run TestProtobufRpcProto.\r\n{code}\r\njava.io.IOException: protocolClass interface org.apache.hadoop.ipc.TestRpcBase$TestRpcService is not implemented by protocolImpl which is of class class org.apache.hadoop.ipc.protobuf.TestRpcServiceProtos$TestProtobufRpcProto$2\r\n\r\n\tat org.apache.hadoop.ipc.WritableRpcEngine$Server.(WritableRpcEngine.java:523)\r\n\tat org.apache.hadoop.ipc.WritableRpcEngine.getServer(WritableRpcEngine.java:379)\r\n\tat org.apache.hadoop.ipc.RPC$Builder.build(RPC.java:986)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:124)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:119)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:102)\r\n\tat org.apache.hadoop.ipc.TestRPC.doRPCs(TestRPC.java:646)\r\n\tat org.apache.hadoop.ipc.TestRPC.testAuthorization(TestRPC.java:705)\r\n\t...\r\n{code}\r\n-----\r\nh3. Original description (trimmed stack trace)\r\n{code}\r\n[ERROR] testAuthorization(org.apache.hadoop.ipc.TestRPC) Time elapsed: 1.066 s <<< ERROR!\r\njava.lang.ClassCastException: class java.net.SocketException cannot be cast to class org.apache.hadoop.ipc.RemoteException (java.net.SocketException is in module java.base of loader 'bootstrap'; org.apache.hadoop.ipc.RemoteException is in unnamed module of loader 'app')\r\n\tat org.apache.hadoop.ipc.TestRPC.doRPCs(TestRPC.java:591)\r\n\tat org.apache.hadoop.ipc.TestRPC.testAuthorization(TestRPC.java:639)\r\n\t...\r\n{code}\r\n","from":"reporter","subject":"TestRPC#testAuthorization fails"},{"body":"I believe the module error is a red herring, and the real error is just that `testAuthorization` doesn't expect a SocketException to be thrown here. `testErrorMsgForInsecureClient` will also fail [while trying to cast to RemoteException|https://github.com/apache/hadoop/blob/trunk/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/ipc/TestRPC.java#L826] because of the same error (although this doesn't fail as consistently as testAuthorization)\r\n\r\nMore interestingly however is that [earlier in the insecure client test case|https://github.com/apache/hadoop/blob/trunk/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/ipc/TestRPC.java#L803], it does the exact same cast to RemoteException, but clearly does so successfully. This clearly shows that the cast itself is not the error. This does however raise the issue of why a SocketException is being thrown in the first place","from":"developer"},{"body":"Currently, it fails with below IOException. It was using WritableRpcEngine to run TestProtobufRpcProto.\r\n{code}\r\njava.io.IOException: protocolClass interface org.apache.hadoop.ipc.TestRpcBase$TestRpcService is not implemented by protocolImpl which is of class class org.apache.hadoop.ipc.protobuf.TestRpcServiceProtos$TestProtobufRpcProto$2\r\n\r\n\tat org.apache.hadoop.ipc.WritableRpcEngine$Server.(WritableRpcEngine.java:523)\r\n\tat org.apache.hadoop.ipc.WritableRpcEngine.getServer(WritableRpcEngine.java:379)\r\n\tat org.apache.hadoop.ipc.RPC$Builder.build(RPC.java:986)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:124)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:119)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:102)\r\n\tat org.apache.hadoop.ipc.TestRPC.doRPCs(TestRPC.java:646)\r\n\tat org.apache.hadoop.ipc.TestRPC.testAuthorization(TestRPC.java:705)\r\n\t...\r\n{code}","from":"developer"},{"body":"Changing it to use ProtobufRpcEngine2 can fix it.","from":"developer"},{"body":"The failure happens in both JDK 8 and JDK 17.  So it is not specific to JDK 17.","from":"developer"},{"body":"The pull request is now merged.","from":"developer"}],"created":"2021-04-09T09:54:42.000+0000","description":"Currently, it fails with below IOException. It was using WritableRpcEngine to run TestProtobufRpcProto.\r\n{code}\r\njava.io.IOException: protocolClass interface org.apache.hadoop.ipc.TestRpcBase$TestRpcService is not implemented by protocolImpl which is of class class org.apache.hadoop.ipc.protobuf.TestRpcServiceProtos$TestProtobufRpcProto$2\r\n\r\n\tat org.apache.hadoop.ipc.WritableRpcEngine$Server.(WritableRpcEngine.java:523)\r\n\tat org.apache.hadoop.ipc.WritableRpcEngine.getServer(WritableRpcEngine.java:379)\r\n\tat org.apache.hadoop.ipc.RPC$Builder.build(RPC.java:986)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:124)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:119)\r\n\tat org.apache.hadoop.ipc.TestRpcBase.setupTestServer(TestRpcBase.java:102)\r\n\tat org.apache.hadoop.ipc.TestRPC.doRPCs(TestRPC.java:646)\r\n\tat org.apache.hadoop.ipc.TestRPC.testAuthorization(TestRPC.java:705)\r\n\t...\r\n{code}\r\n-----\r\nh3. Original description (trimmed stack trace)\r\n{code}\r\n[ERROR] testAuthorization(org.apache.hadoop.ipc.TestRPC) Time elapsed: 1.066 s <<< ERROR!\r\njava.lang.ClassCastException: class java.net.SocketException cannot be cast to class org.apache.hadoop.ipc.RemoteException (java.net.SocketException is in module java.base of loader 'bootstrap'; org.apache.hadoop.ipc.RemoteException is in unnamed module of loader 'app')\r\n\tat org.apache.hadoop.ipc.TestRPC.doRPCs(TestRPC.java:591)\r\n\tat org.apache.hadoop.ipc.TestRPC.testAuthorization(TestRPC.java:639)\r\n\t...\r\n{code}\r\n","issue_id":"13370616","key":"HADOOP-17629","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-09-30T15:55:29.000+0000","role":"fixed_distractor","summary":"TestRPC#testAuthorization fails"} {"case_id":"13379711","cluster":"DISTRACTOR-HADOOP-17723","comments":[{"body":"Hmm. This patch doesn't fix the \"header not found\" ({{Python.h: No such file}}) issue for me. It is possible that a previous version of {{python2.7}} package for {{ubuntu:bionic}} contains the header but now it is gone.\r\n\r\n{code:title=docker build -t apache/hadoop:3 -f Dockerfile_aarch64 .}\r\n#13 9.705 In file included from ast27/Parser/acceler.c:13:0:\r\n#13 9.705 ast27/Parser/../Include/pgenheaders.h:8:10: fatal error: Python.h: No such file or directory\r\n#13 9.705 #include \"Python.h\"\r\n#13 9.705 ^~~~~~~~~~\r\n#13 9.705 compilation terminated.\r\n#13 9.705 error: command 'aarch64-linux-gnu-gcc' failed with exit status 1\r\n{code}\r\n\r\nI am able to work around the issue by installing the {{python3-dev}} package (which should have the Python.h header file):\r\n\r\n{code:bash|title=Fix}\r\ndiff --git a/dev-support/docker/Dockerfile_aarch64 b/dev-support/docker/Dockerfile_aarch64\r\nindex 46818a6e234..63de24146a4 100644\r\n--- a/dev-support/docker/Dockerfile_aarch64\r\n+++ b/dev-support/docker/Dockerfile_aarch64\r\n@@ -74,6 +74,7 @@ RUN apt-get -q update \\\r\n pkg-config \\\r\n python2.7 \\\r\n python3 \\\r\n+ python3-dev \\\r\n python3-pip \\\r\n python3-pkg-resources \\\r\n python3-setuptools \\\r\n{code}\r\n\r\nNot sure why {{python3-dev}} is mentioned in the jira description but somehow omitted in the PR?\r\n\r\nWith the diff above, {{cd dev-support/docker/ && docker build -t apache/hadoop:3 -f Dockerfile_aarch64 .}} works for me. Same deal for branch-3.3.2 (need to add python3-dev). trunk (3.4.0) works without any changes.","created":"2021-12-02T10:55:24.424+0000"},{"body":"can you post a PR?\r\ni think the problem is intermittent for me.","created":"2021-12-06T03:27:54.161+0000"},{"body":"[~weichiu] Yup, will post the one-line change for branch-3.3 later.","created":"2021-12-06T19:11:25.323+0000"},{"body":"[~weichiu] Posted HADOOP-18048. The plan is to merge to branch-3.3 first then backport it to branch-3.3.2","created":"2021-12-14T23:04:09.245+0000"}],"conversations":[{"body":"Running the create-release script for Hadoop 3.3.1 on an ARM machine, docker image fails to build:\r\n\r\n{noformat}\r\n aarch64-linux-gnu-gcc -pthread -DNDEBUG -g -fwrapv -O2 -Wall -g -fstack-protector-strong -Wformat -Werror=format-security -Wdate-time -D_FORTIFY_SOURCE=2 -fPIC -Iast27/Include -I/usr/include/python3.6m -c ast27/Parser/acceler.c -o build/temp.linux-aarch64-3.6/ast27/Parser/acceler.o In file included from ast27/Parser/acceler.c:13:0: ast27/Parser/../Include/pgenheaders.h:8:10: fatal error: Python.h: No such file or directory #include \"Python.h\" ^~~~~~~~~~ compilation terminated. error: command 'aarch64-linux-gnu-gcc' failed with exit status 1\r\n\r\n\r\n{noformat}\r\n\r\nThe missing Python3.h requires python3-dev package: https://stackoverflow.com/questions/21530577/fatal-error-python-h-no-such-file-or-directory\r\n\r\nThe PhantomJS binary was built for Xenial, doesn't run after the Dockerfile migrated to Bionic/Focal. Fortunately Bionic/Focal has official PhantomJS packages.","from":"reporter","subject":"[build] fix the Dockerfile for ARM"},{"body":"Hmm. This patch doesn't fix the \"header not found\" ({{Python.h: No such file}}) issue for me. It is possible that a previous version of {{python2.7}} package for {{ubuntu:bionic}} contains the header but now it is gone.\r\n\r\n{code:title=docker build -t apache/hadoop:3 -f Dockerfile_aarch64 .}\r\n#13 9.705 In file included from ast27/Parser/acceler.c:13:0:\r\n#13 9.705 ast27/Parser/../Include/pgenheaders.h:8:10: fatal error: Python.h: No such file or directory\r\n#13 9.705 #include \"Python.h\"\r\n#13 9.705 ^~~~~~~~~~\r\n#13 9.705 compilation terminated.\r\n#13 9.705 error: command 'aarch64-linux-gnu-gcc' failed with exit status 1\r\n{code}\r\n\r\nI am able to work around the issue by installing the {{python3-dev}} package (which should have the Python.h header file):\r\n\r\n{code:bash|title=Fix}\r\ndiff --git a/dev-support/docker/Dockerfile_aarch64 b/dev-support/docker/Dockerfile_aarch64\r\nindex 46818a6e234..63de24146a4 100644\r\n--- a/dev-support/docker/Dockerfile_aarch64\r\n+++ b/dev-support/docker/Dockerfile_aarch64\r\n@@ -74,6 +74,7 @@ RUN apt-get -q update \\\r\n pkg-config \\\r\n python2.7 \\\r\n python3 \\\r\n+ python3-dev \\\r\n python3-pip \\\r\n python3-pkg-resources \\\r\n python3-setuptools \\\r\n{code}\r\n\r\nNot sure why {{python3-dev}} is mentioned in the jira description but somehow omitted in the PR?\r\n\r\nWith the diff above, {{cd dev-support/docker/ && docker build -t apache/hadoop:3 -f Dockerfile_aarch64 .}} works for me. Same deal for branch-3.3.2 (need to add python3-dev). trunk (3.4.0) works without any changes.","from":"developer"},{"body":"can you post a PR?\r\ni think the problem is intermittent for me.","from":"developer"},{"body":"[~weichiu] Yup, will post the one-line change for branch-3.3 later.","from":"developer"},{"body":"[~weichiu] Posted HADOOP-18048. The plan is to merge to branch-3.3 first then backport it to branch-3.3.2","from":"developer"}],"created":"2021-05-21T09:12:15.000+0000","description":"Running the create-release script for Hadoop 3.3.1 on an ARM machine, docker image fails to build:\r\n\r\n{noformat}\r\n aarch64-linux-gnu-gcc -pthread -DNDEBUG -g -fwrapv -O2 -Wall -g -fstack-protector-strong -Wformat -Werror=format-security -Wdate-time -D_FORTIFY_SOURCE=2 -fPIC -Iast27/Include -I/usr/include/python3.6m -c ast27/Parser/acceler.c -o build/temp.linux-aarch64-3.6/ast27/Parser/acceler.o In file included from ast27/Parser/acceler.c:13:0: ast27/Parser/../Include/pgenheaders.h:8:10: fatal error: Python.h: No such file or directory #include \"Python.h\" ^~~~~~~~~~ compilation terminated. error: command 'aarch64-linux-gnu-gcc' failed with exit status 1\r\n\r\n\r\n{noformat}\r\n\r\nThe missing Python3.h requires python3-dev package: https://stackoverflow.com/questions/21530577/fatal-error-python-h-no-such-file-or-directory\r\n\r\nThe PhantomJS binary was built for Xenial, doesn't run after the Dockerfile migrated to Bionic/Focal. Fortunately Bionic/Focal has official PhantomJS packages.","issue_id":"13379711","key":"HADOOP-17723","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2021-06-01T01:58:06.000+0000","role":"fixed_distractor","summary":"[build] fix the Dockerfile for ARM"} {"case_id":"13384984","cluster":"DISTRACTOR-HADOOP-17769","comments":[{"body":"I have created two Pull requests for both banch-2.10 and trunk.\r\nI believe this fix is needed for other 3.x too.\r\n\r\nI verified Junit-4.13.1 had the bug and would cause the exceptions described in HDFS-16072.\r\n\r\n[~ayushtkn] can you please take a look at the upgrade and commit it to both branches?\r\n","created":"2021-06-21T18:25:21.649+0000"},{"body":"Committed to trunk, branch-3.3,3.2 and 2.10\r\nThanx [~ahussein] for the contribution!!!","created":"2021-06-25T17:22:24.058+0000"},{"body":"yes -thank you!","created":"2021-06-28T11:00:58.843+0000"}],"conversations":[{"body":"JUnit 4.13.1 has a bug that is reported in Junit [issue-1652|https://github.com/junit-team/junit4/issues/1652] _Timeout ThreadGroups should not be destroyed_\r\n\r\nAfter upgrading Junit to 4.13.1 in HADOOP-17602, {{TestBlockRecovery}} started to fail regularly in branch-3.x and branch-2.10.\r\nWhile investigating the failure in branch-2.10 HDFS-16072, I found out that the bug is the main reason {{TestBlockRecovery}} started to fail because the timeout of the Junit would try to close a ThreadGroup that has been already closed which throws the {{java.lang.IllegalThreadStateException}}.\r\n\r\nThe bug has been fixed in Junit-4.13.2\r\n\r\nFor branch-3.x, HDFS-15940 did not address the root cause of the problem. Eventually, Splitting the {{TestBlockRecovery}} hid the bug, but the upgrade needs to be done so that the problem does not show up in another unit test.","from":"reporter","subject":"Upgrade JUnit to 4.13.2"},{"body":"I have created two Pull requests for both banch-2.10 and trunk.\r\nI believe this fix is needed for other 3.x too.\r\n\r\nI verified Junit-4.13.1 had the bug and would cause the exceptions described in HDFS-16072.\r\n\r\n[~ayushtkn] can you please take a look at the upgrade and commit it to both branches?\r\n","from":"developer"},{"body":"Committed to trunk, branch-3.3,3.2 and 2.10\r\nThanx [~ahussein] for the contribution!!!","from":"developer"},{"body":"yes -thank you!","from":"developer"}],"created":"2021-06-21T17:24:36.000+0000","description":"JUnit 4.13.1 has a bug that is reported in Junit [issue-1652|https://github.com/junit-team/junit4/issues/1652] _Timeout ThreadGroups should not be destroyed_\r\n\r\nAfter upgrading Junit to 4.13.1 in HADOOP-17602, {{TestBlockRecovery}} started to fail regularly in branch-3.x and branch-2.10.\r\nWhile investigating the failure in branch-2.10 HDFS-16072, I found out that the bug is the main reason {{TestBlockRecovery}} started to fail because the timeout of the Junit would try to close a ThreadGroup that has been already closed which throws the {{java.lang.IllegalThreadStateException}}.\r\n\r\nThe bug has been fixed in Junit-4.13.2\r\n\r\nFor branch-3.x, HDFS-15940 did not address the root cause of the problem. Eventually, Splitting the {{TestBlockRecovery}} hid the bug, but the upgrade needs to be done so that the problem does not show up in another unit test.","issue_id":"13384984","key":"HADOOP-17769","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2021-06-25T17:21:48.000+0000","role":"fixed_distractor","summary":"Upgrade JUnit to 4.13.2"} {"case_id":"12376985","cluster":"DISTRACTOR-HADOOP-1792","comments":[{"body":"Is anyone else seeing this? There were changes to DF.java in 0.14, but none that should cause this. Are you still running under a Cygwin shell? We depend on Cygwin binaries being on $PATH, including at least 'df' and 'bash'.","created":"2007-08-28T17:37:02.168+0000"},{"body":"It use to work without cygwin in the path var.\n\nI added it and it work yes.\n\nBut I don't understand why it use to work before, that's weird ?!","created":"2007-08-29T07:33:08.890+0000"},{"body":"I got this problem too: It appears DF is being used to compute available disk space before creating files. While Java does not appears to have support for this (see [bug 4057701|http://bugs.sun.com/bugdatabase/view_bug.do?bug_id=4057701]), it does appear that another Apache project (sort of) does:\n\n[org.apache.commons.FileSystemUtils.freeSpaceKb|http://commons.apache.org/io/api-release/org/apache/commons/io/FileSystemUtils.html#freeSpaceKb(java.lang.String)]\n\n...at least it wouldn't require that Hadoop users on Windows install Cygwin.","created":"2007-11-30T20:49:19.272+0000"},{"body":"looks like this has been fixed and works with cygwin as intended. resolving for now.","created":"2008-02-04T23:16:30.313+0000"},{"body":"I had this problem and the comments here were not very clear. I solved it by trying various things...so here is how I got it to work:\n\nMy environment:\n- Windows XP\n- Eclipse running under Windows\n- Eclipse includes the IBM plugin for mapreduce/hadoop\n- I was trying to run the mapreduce job as a java application via Eclipse\n\nSolution:\n- Installed the latest stable version of Cygwin\n- Added \"c:\\Cygwin\\bin \" (or equivalent) to Path on windows\n- Started the Cygwin shell\n- Ran the mapreduce job as a java application while the Cygwin shell was running.\n\nHope this helps clarify things for other newbies.","created":"2008-02-28T01:32:41.356+0000"},{"body":"Hello, I've done all this, but when I run my job in eclipse, I get the Visual Studio just-in-time debugger for df.exe, as if it is finding it, but it's not running correctly. \nHas anyone else had this problem?\nIt seems to occur when the map job completes...\n\n-Steve","created":"2009-11-06T18:40:29.266+0000"}],"conversations":[{"body":"My code use to work with previous version of hadoop, I upgraded to 0.14 and now:\njava.io.IOException: CreateProcess: df -k \"C:\\Documents and Settings\\Benjamin\\Local Settings\\Temp\\test14906test\\mapredLocal\" error=2\n\tat java.lang.ProcessImpl.create(Native Method)\n\tat java.lang.ProcessImpl.(Unknown Source)\n\tat java.lang.ProcessImpl.start(Unknown Source)\n\tat java.lang.ProcessBuilder.start(Unknown Source)\n\tat java.lang.Runtime.exec(Unknown Source)\n\tat java.lang.Runtime.exec(Unknown Source)\n\tat org.apache.hadoop.fs.DF.doDF(DF.java:60)\n\tat org.apache.hadoop.fs.DF.(DF.java:53)\n\tat org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.confChanged(LocalDirAllocator.java:198)\n\tat org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.getLocalPathForWrite(LocalDirAllocator.java:235)\n\tat org.apache.hadoop.fs.LocalDirAllocator.getLocalPathForWrite(LocalDirAllocator.java:124)\n\tat org.apache.hadoop.mapred.MapOutputFile.getSpillFileForWrite(MapOutputFile.java:88)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.sortAndSpillToDisk(MapTask.java:373)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.flush(MapTask.java:593)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:190)\n\tat org.apache.hadoop.mapred.LocalJobRunner$Job.run(LocalJobRunner.java:137)\n\tat org.apache.hadoop.mapred.LocalJobRunner.submitJob(LocalJobRunner.java:283)\n\tat org.apache.hadoop.mapred.JobClient.submitJob(JobClient.java:397)\n...","from":"reporter","subject":"df command doesn't exist under windows"},{"body":"Is anyone else seeing this? There were changes to DF.java in 0.14, but none that should cause this. Are you still running under a Cygwin shell? We depend on Cygwin binaries being on $PATH, including at least 'df' and 'bash'.","from":"developer"},{"body":"It use to work without cygwin in the path var.\n\nI added it and it work yes.\n\nBut I don't understand why it use to work before, that's weird ?!","from":"developer"},{"body":"I got this problem too: It appears DF is being used to compute available disk space before creating files. While Java does not appears to have support for this (see [bug 4057701|http://bugs.sun.com/bugdatabase/view_bug.do?bug_id=4057701]), it does appear that another Apache project (sort of) does:\n\n[org.apache.commons.FileSystemUtils.freeSpaceKb|http://commons.apache.org/io/api-release/org/apache/commons/io/FileSystemUtils.html#freeSpaceKb(java.lang.String)]\n\n...at least it wouldn't require that Hadoop users on Windows install Cygwin.","from":"developer"},{"body":"looks like this has been fixed and works with cygwin as intended. resolving for now.","from":"developer"},{"body":"I had this problem and the comments here were not very clear. I solved it by trying various things...so here is how I got it to work:\n\nMy environment:\n- Windows XP\n- Eclipse running under Windows\n- Eclipse includes the IBM plugin for mapreduce/hadoop\n- I was trying to run the mapreduce job as a java application via Eclipse\n\nSolution:\n- Installed the latest stable version of Cygwin\n- Added \"c:\\Cygwin\\bin \" (or equivalent) to Path on windows\n- Started the Cygwin shell\n- Ran the mapreduce job as a java application while the Cygwin shell was running.\n\nHope this helps clarify things for other newbies.","from":"developer"},{"body":"Hello, I've done all this, but when I run my job in eclipse, I get the Visual Studio just-in-time debugger for df.exe, as if it is finding it, but it's not running correctly. \nHas anyone else had this problem?\nIt seems to occur when the map job completes...\n\n-Steve","from":"developer"}],"created":"2007-08-28T10:36:51.000+0000","description":"My code use to work with previous version of hadoop, I upgraded to 0.14 and now:\njava.io.IOException: CreateProcess: df -k \"C:\\Documents and Settings\\Benjamin\\Local Settings\\Temp\\test14906test\\mapredLocal\" error=2\n\tat java.lang.ProcessImpl.create(Native Method)\n\tat java.lang.ProcessImpl.(Unknown Source)\n\tat java.lang.ProcessImpl.start(Unknown Source)\n\tat java.lang.ProcessBuilder.start(Unknown Source)\n\tat java.lang.Runtime.exec(Unknown Source)\n\tat java.lang.Runtime.exec(Unknown Source)\n\tat org.apache.hadoop.fs.DF.doDF(DF.java:60)\n\tat org.apache.hadoop.fs.DF.(DF.java:53)\n\tat org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.confChanged(LocalDirAllocator.java:198)\n\tat org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.getLocalPathForWrite(LocalDirAllocator.java:235)\n\tat org.apache.hadoop.fs.LocalDirAllocator.getLocalPathForWrite(LocalDirAllocator.java:124)\n\tat org.apache.hadoop.mapred.MapOutputFile.getSpillFileForWrite(MapOutputFile.java:88)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.sortAndSpillToDisk(MapTask.java:373)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.flush(MapTask.java:593)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:190)\n\tat org.apache.hadoop.mapred.LocalJobRunner$Job.run(LocalJobRunner.java:137)\n\tat org.apache.hadoop.mapred.LocalJobRunner.submitJob(LocalJobRunner.java:283)\n\tat org.apache.hadoop.mapred.JobClient.submitJob(JobClient.java:397)\n...","issue_id":"12376985","key":"HADOOP-1792","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-02-04T23:16:46.000+0000","role":"fixed_distractor","summary":"df command doesn't exist under windows"} {"case_id":"13424131","cluster":"DISTRACTOR-HADOOP-18090","comments":[{"body":"[~stevel@apache.org] [~busbey] Can this be accepted as a bug?","created":"2022-01-25T07:11:21.231+0000"},{"body":"I'm not sure what kind of validation you're looking for. Generally, just about anyone can submit things as a bug against hadoop (as you have done here). I believe you have filed it against the correct component.\r\n\r\nIf you mean \"can someone fix this\", that's all done essentially by folks volunteering. If no one picks up this issue you might try discussing it on the dev list. It will get much more attention if you attempt to fix things.\r\n\r\nTo me, this looks like a mismatch in expectations around SFTPFileSystem and our classpath isolated client libraries. The two questions that I would work out are\r\n\r\n# _should_ SFTPFileSystem be included in the client libraries or is it misplaced? This feels akin to the issue we had with s3a in HADOOP-16080 and ideally solved akin to HADOOP-15387\r\n# Presuming this should stay where it is, fixing it means changing the set of included relocated classes. Why isn't this class already present if it's reachable from a class we include? How extensive a change is correcting that.","created":"2022-01-25T14:23:53.573+0000"},{"body":"[~busbey] For the second question why the class isn't reachable, it seems hadoop-client-api has a compile time dependency on hadoop-common which gets shaded because of pattern.\r\nHowever there are many transitive dependencies of hadoop-common which are excluded. The list of exlusions comprise {_}com.jcraft:jsch{_}.\r\nHence in SFTPFileSystem, the pattern matching rewrites the imports to relocated path while at the same time the classes which were supposed to be relocated were excluded from POM. \r\nSheer text replacement without any class reloaction.\r\n\r\nRegarding first point, SFTPFileSystem is a part of hadoop-common. Now it would be good to have it land in its own module unlike hadoop-aws.\r\n\r\nFor fixes, I am new in this space that's why was just checking if this is indeed a bug. I know in OSS community solves the issues. Not asking you to fix :).","created":"2022-01-25T14:57:07.785+0000"}],"conversations":[{"body":"Spark 3.2.0 transitively introduces hadoop-client-api and hadoop-client-runtime dependencies.\r\n\r\nWhen we create a SFTPFileSystem instance (org.apache.hadoop.fs.sftp.SFTPFileSystem) it tries to load the relocated classes from _com.jcraft.jsch_ package.\r\n\r\nThe filesystem instance creation fails with error:\r\n{code:java}\r\njava.lang.ClassNotFoundException: org.apache.hadoop.shaded.com.jcraft.jsch.SftpException\r\n    at java.net.URLClassLoader.findClass(URLClassLoader.java:382)\r\n    at java.lang.ClassLoader.loadClass(ClassLoader.java:424)\r\n    at sun.misc.Launcher$AppClassLoader.loadClass(Launcher.java:349)\r\n    at java.lang.ClassLoader.loadClass(ClassLoader.java:357) {code}\r\n\r\nExcluding client from transitive load of spark and directly using hadoop-common/hadoop-client is the way its working for us.","from":"reporter","subject":"Exclude com/jcraft/jsch classes from being shaded/relocated"},{"body":"[~stevel@apache.org] [~busbey] Can this be accepted as a bug?","from":"developer"},{"body":"I'm not sure what kind of validation you're looking for. Generally, just about anyone can submit things as a bug against hadoop (as you have done here). I believe you have filed it against the correct component.\r\n\r\nIf you mean \"can someone fix this\", that's all done essentially by folks volunteering. If no one picks up this issue you might try discussing it on the dev list. It will get much more attention if you attempt to fix things.\r\n\r\nTo me, this looks like a mismatch in expectations around SFTPFileSystem and our classpath isolated client libraries. The two questions that I would work out are\r\n\r\n# _should_ SFTPFileSystem be included in the client libraries or is it misplaced? This feels akin to the issue we had with s3a in HADOOP-16080 and ideally solved akin to HADOOP-15387\r\n# Presuming this should stay where it is, fixing it means changing the set of included relocated classes. Why isn't this class already present if it's reachable from a class we include? How extensive a change is correcting that.","from":"developer"},{"body":"[~busbey] For the second question why the class isn't reachable, it seems hadoop-client-api has a compile time dependency on hadoop-common which gets shaded because of pattern.\r\nHowever there are many transitive dependencies of hadoop-common which are excluded. The list of exlusions comprise {_}com.jcraft:jsch{_}.\r\nHence in SFTPFileSystem, the pattern matching rewrites the imports to relocated path while at the same time the classes which were supposed to be relocated were excluded from POM. \r\nSheer text replacement without any class reloaction.\r\n\r\nRegarding first point, SFTPFileSystem is a part of hadoop-common. Now it would be good to have it land in its own module unlike hadoop-aws.\r\n\r\nFor fixes, I am new in this space that's why was just checking if this is indeed a bug. I know in OSS community solves the issues. Not asking you to fix :).","from":"developer"}],"created":"2022-01-22T09:17:58.000+0000","description":"Spark 3.2.0 transitively introduces hadoop-client-api and hadoop-client-runtime dependencies.\r\n\r\nWhen we create a SFTPFileSystem instance (org.apache.hadoop.fs.sftp.SFTPFileSystem) it tries to load the relocated classes from _com.jcraft.jsch_ package.\r\n\r\nThe filesystem instance creation fails with error:\r\n{code:java}\r\njava.lang.ClassNotFoundException: org.apache.hadoop.shaded.com.jcraft.jsch.SftpException\r\n    at java.net.URLClassLoader.findClass(URLClassLoader.java:382)\r\n    at java.lang.ClassLoader.loadClass(ClassLoader.java:424)\r\n    at sun.misc.Launcher$AppClassLoader.loadClass(Launcher.java:349)\r\n    at java.lang.ClassLoader.loadClass(ClassLoader.java:357) {code}\r\n\r\nExcluding client from transitive load of spark and directly using hadoop-common/hadoop-client is the way its working for us.","issue_id":"13424131","key":"HADOOP-18090","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-11-06T12:33:25.000+0000","role":"fixed_distractor","summary":"Exclude com/jcraft/jsch classes from being shaded/relocated"} {"case_id":"13496586","cluster":"DISTRACTOR-HADOOP-18521","comments":[{"body":"I can't provide a test for this as you need a multi GB CSV file and a build of spark configured to use your hadoop dist.\r\n\r\nThe latest build of cloudstore (https://github.com/steveloughran/cloudstore) has a command {{mkcsv}} which can create the file; the man page includes the spark binding info: https://github.com/steveloughran/cloudstore/blob/trunk/src/main/site/mkcsv.md\r\n\r\nalong with the fix, i am going to include a stream capability which the fs and stream can be probed for to declare that the fix is in. this allows for programmatic verification of the safety of releases, including with the cloudstore pathcapabilities command\r\n\r\nupdate: no, the proposed fix is insufficent","created":"2022-11-05T15:08:50.080+0000"},{"body":"We have a nice and minimal fix which can be easily backported anywhere that is needed.\r\n\r\nThat larger patch of mine was intended to\r\n# avoid adding the prefetches of closed streams to the completed list (impact: buffers will be retained until timeout)\r\n# add validation about accidental buffer reuse\r\n# handle failures in the async read caused by the close()\r\n# iostatistics for production code and testing. used in asserts and will allow us to assess value of prefetching in production code\r\n# pulling out of the methods invoked on abfs input stream into their own interface, again for testing.\r\n\r\nwe don't need this as much any more; it's something where I would like most of the features (1, 3, 4, 5) in, but we could look at that in terms of a broader review of the readbuffer feature\r\n\r\nI'm going to\r\nrebase my pr\r\ncreate a new jira for \"extend readbuffer for testing/statistics\" and move the pr to that","created":"2022-12-09T14:01:32.873+0000"},{"body":"This is fixed in HADOOP-18546; the followups were just tuning.\r\n\r\nclosing as such. my larger bit of work has some advantages (better testability, iostats of use) but that makes it too complex to put in 3.3.5 and means more work remaining to fix.\r\n\r\nwhen we do a rework of the read buffer manager some aspects of it can be applied. \r\n* iostats\r\n* tryEvict prioritising eviction of completed fetches with buffers belonging to closed windows\r\n* AbfsInputStream calls to go an interface, with unit tests\r\n\r\nIt'd also be good to include split start/end and read policy from stream to manager\r\n* don't prefetch past end of split (or at most, one block)\r\n* on random IO, use optimised policy (no prefetch? one block max)\r\n* on vectored IO: no prefetching\r\nj","created":"2022-12-19T11:11:44.981+0000"}],"conversations":[{"body":"{{AbfsInputStream.close()}} can trigger the return of buffers used for active prefetch GET requests into the ReadBufferManager free buffer pool.\r\n\r\nA subsequent prefetch by a different stream in the same process may acquire this same buffer. This can lead to risk of corruption of its own prefetched data, data which may then be returned to that other thread.\r\n\r\nThe full analysis in in the document attached to this JIRA.\r\n\r\nThe issue is fixed in Hadoop 3.3.5 \r\n\r\nh2. Emergency fix through site configuration\r\n\r\nOn releases without the fix for this (3.3.2-3.3.4), the bug can be avoided by disabling all prefetching\r\n{code:java}\r\nfs.azure.readaheadqueue.depth = 0\r\n{code}\r\n\r\nh2. Automated probes for risk of exposure\r\n\r\nThe [cloudstore|https://github.com/steveloughran/cloudstore] diagnostics JAR has a command [safeprefetch|https://github.com/steveloughran/cloudstore/blob/trunk/src/main/site/safeprefetch.md] which probes an abfs client for being vulnerable. It does this through {{PathCapabilities.hasPathCapability()}} probes. It can be invoked on the command line to validate the version/configuration\r\n\r\nConsult [the source|https://github.com/steveloughran/cloudstore/blob/trunk/src/main/java/org/apache/hadoop/fs/store/abfs/SafePrefetch.java#L96] to see how to do this programmatically.\r\n\r\nNote also that the tool's [mkcsv|https://github.com/steveloughran/cloudstore/blob/trunk/src/main/site/mkcsv.md] command can be used to generate the multi-GB CSV files needed to trigger the condition and so verify that the issue exists.\r\n\r\n\r\nh2. Microsoft Announcement\r\n\r\n{code}\r\n\r\nFrom: Sneha Vijayarajan\r\nSubject: RE: Alert ! ABFS Driver - Possible data corruption on read path\r\n\r\nHi,\r\n\r\nOne of the contributions made to ABFS Driver has a potential to cause data corruption on read\r\npath.\r\n\r\nPlease check if the below change is part of any of your releases:\r\n\r\nHADOOP-17156. Purging the buffers associated with input streams during close() by mukund-thakur\r\n· Pull Request #3285 · apache/hadoop (github.com)\r\n\r\nRCA: Scenario that can lead to data corruption:\r\n\r\nDriver allocates a bunch of prefetch buffers at init and are shared by different instances of\r\nInputStreams created within that process. These prefetch buffers could be in 3 stages –\r\n\r\n* In ReadAheadQueue : request for prefetch logged\r\n* In ProgressList : Work has begun to talk to backend store to get the requested data\r\n* In CompletedList: Prefetch data is now available for consumption.\r\n\r\nWhen multiple InputStreams have prefetch buffers across these states and close is triggered on\r\nany InputStream/s, the commit above will remove buffers allotted to respective stream from all\r\nthe 3 lists and also declare that the buffers are available for new prefetches to happen, but\r\nno action to cancel/prevent buffer from being updated with ongoing network request is done.\r\nData corruption can happen if one such freed up buffer from InProgressList is allotted to a new\r\nprefetch request and then the buffer got filled up with the previous stream’s network request.\r\n\r\nMitigation: If this change is present in any release, kindly help communicate to your customers\r\nto immediately set below config to 0 in their clusters. This will disable prefetches which can\r\nhave an impact on perf but will prevent the possibility of data corruption.\r\n\r\nfs.azure.readaheadqueue.depth: Sets the readahead queue depth in AbfsInputStream. In case the\r\nset value is negative the read ahead queue depth will be set as\r\nRuntime.getRuntime().availableProcessors(). By default the value will be 2. To disable\r\nreadaheads, set this value to 0. If your workload is doing only random reads (non-sequential)\r\nor you are seeing throttling, you may try setting this value to 0.\r\n\r\nNext steps: We are getting help to post the notifications for this in Apache groups. Work on\r\nHotFix is also ongoing. Will update this thread once the change is checked in.\r\n\r\nPlease reach out for any queries or clarifications.\r\n\r\nThanks,\r\nSneha Vijayarajan\r\n\r\n{code}\r\n ","from":"reporter","subject":"ABFS ReadBufferManager buffer sharing across concurrent HTTP requests"},{"body":"I can't provide a test for this as you need a multi GB CSV file and a build of spark configured to use your hadoop dist.\r\n\r\nThe latest build of cloudstore (https://github.com/steveloughran/cloudstore) has a command {{mkcsv}} which can create the file; the man page includes the spark binding info: https://github.com/steveloughran/cloudstore/blob/trunk/src/main/site/mkcsv.md\r\n\r\nalong with the fix, i am going to include a stream capability which the fs and stream can be probed for to declare that the fix is in. this allows for programmatic verification of the safety of releases, including with the cloudstore pathcapabilities command\r\n\r\nupdate: no, the proposed fix is insufficent","from":"developer"},{"body":"We have a nice and minimal fix which can be easily backported anywhere that is needed.\r\n\r\nThat larger patch of mine was intended to\r\n# avoid adding the prefetches of closed streams to the completed list (impact: buffers will be retained until timeout)\r\n# add validation about accidental buffer reuse\r\n# handle failures in the async read caused by the close()\r\n# iostatistics for production code and testing. used in asserts and will allow us to assess value of prefetching in production code\r\n# pulling out of the methods invoked on abfs input stream into their own interface, again for testing.\r\n\r\nwe don't need this as much any more; it's something where I would like most of the features (1, 3, 4, 5) in, but we could look at that in terms of a broader review of the readbuffer feature\r\n\r\nI'm going to\r\nrebase my pr\r\ncreate a new jira for \"extend readbuffer for testing/statistics\" and move the pr to that","from":"developer"},{"body":"This is fixed in HADOOP-18546; the followups were just tuning.\r\n\r\nclosing as such. my larger bit of work has some advantages (better testability, iostats of use) but that makes it too complex to put in 3.3.5 and means more work remaining to fix.\r\n\r\nwhen we do a rework of the read buffer manager some aspects of it can be applied. \r\n* iostats\r\n* tryEvict prioritising eviction of completed fetches with buffers belonging to closed windows\r\n* AbfsInputStream calls to go an interface, with unit tests\r\n\r\nIt'd also be good to include split start/end and read policy from stream to manager\r\n* don't prefetch past end of split (or at most, one block)\r\n* on random IO, use optimised policy (no prefetch? one block max)\r\n* on vectored IO: no prefetching\r\nj","from":"developer"}],"created":"2022-11-05T15:00:26.000+0000","description":"{{AbfsInputStream.close()}} can trigger the return of buffers used for active prefetch GET requests into the ReadBufferManager free buffer pool.\r\n\r\nA subsequent prefetch by a different stream in the same process may acquire this same buffer. This can lead to risk of corruption of its own prefetched data, data which may then be returned to that other thread.\r\n\r\nThe full analysis in in the document attached to this JIRA.\r\n\r\nThe issue is fixed in Hadoop 3.3.5 \r\n\r\nh2. Emergency fix through site configuration\r\n\r\nOn releases without the fix for this (3.3.2-3.3.4), the bug can be avoided by disabling all prefetching\r\n{code:java}\r\nfs.azure.readaheadqueue.depth = 0\r\n{code}\r\n\r\nh2. Automated probes for risk of exposure\r\n\r\nThe [cloudstore|https://github.com/steveloughran/cloudstore] diagnostics JAR has a command [safeprefetch|https://github.com/steveloughran/cloudstore/blob/trunk/src/main/site/safeprefetch.md] which probes an abfs client for being vulnerable. It does this through {{PathCapabilities.hasPathCapability()}} probes. It can be invoked on the command line to validate the version/configuration\r\n\r\nConsult [the source|https://github.com/steveloughran/cloudstore/blob/trunk/src/main/java/org/apache/hadoop/fs/store/abfs/SafePrefetch.java#L96] to see how to do this programmatically.\r\n\r\nNote also that the tool's [mkcsv|https://github.com/steveloughran/cloudstore/blob/trunk/src/main/site/mkcsv.md] command can be used to generate the multi-GB CSV files needed to trigger the condition and so verify that the issue exists.\r\n\r\n\r\nh2. Microsoft Announcement\r\n\r\n{code}\r\n\r\nFrom: Sneha Vijayarajan\r\nSubject: RE: Alert ! ABFS Driver - Possible data corruption on read path\r\n\r\nHi,\r\n\r\nOne of the contributions made to ABFS Driver has a potential to cause data corruption on read\r\npath.\r\n\r\nPlease check if the below change is part of any of your releases:\r\n\r\nHADOOP-17156. Purging the buffers associated with input streams during close() by mukund-thakur\r\n· Pull Request #3285 · apache/hadoop (github.com)\r\n\r\nRCA: Scenario that can lead to data corruption:\r\n\r\nDriver allocates a bunch of prefetch buffers at init and are shared by different instances of\r\nInputStreams created within that process. These prefetch buffers could be in 3 stages –\r\n\r\n* In ReadAheadQueue : request for prefetch logged\r\n* In ProgressList : Work has begun to talk to backend store to get the requested data\r\n* In CompletedList: Prefetch data is now available for consumption.\r\n\r\nWhen multiple InputStreams have prefetch buffers across these states and close is triggered on\r\nany InputStream/s, the commit above will remove buffers allotted to respective stream from all\r\nthe 3 lists and also declare that the buffers are available for new prefetches to happen, but\r\nno action to cancel/prevent buffer from being updated with ongoing network request is done.\r\nData corruption can happen if one such freed up buffer from InProgressList is allotted to a new\r\nprefetch request and then the buffer got filled up with the previous stream’s network request.\r\n\r\nMitigation: If this change is present in any release, kindly help communicate to your customers\r\nto immediately set below config to 0 in their clusters. This will disable prefetches which can\r\nhave an impact on perf but will prevent the possibility of data corruption.\r\n\r\nfs.azure.readaheadqueue.depth: Sets the readahead queue depth in AbfsInputStream. In case the\r\nset value is negative the read ahead queue depth will be set as\r\nRuntime.getRuntime().availableProcessors(). By default the value will be 2. To disable\r\nreadaheads, set this value to 0. If your workload is doing only random reads (non-sequential)\r\nor you are seeing throttling, you may try setting this value to 0.\r\n\r\nNext steps: We are getting help to post the notifications for this in Apache groups. Work on\r\nHotFix is also ongoing. Will update this thread once the change is checked in.\r\n\r\nPlease reach out for any queries or clarifications.\r\n\r\nThanks,\r\nSneha Vijayarajan\r\n\r\n{code}\r\n ","issue_id":"13496586","key":"HADOOP-18521","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2022-12-19T11:11:44.000+0000","role":"fixed_distractor","summary":"ABFS ReadBufferManager buffer sharing across concurrent HTTP requests"} {"case_id":"13514848","cluster":"DISTRACTOR-HADOOP-18581","comments":[{"body":"[~surendralilhore]  thanks for reporting and working on this. Are you planning to backport to other branches also..?\r\n{quote}When JN client try to re-login and it fails, it will destroy server service ticket also and NameNode not able to server client request. We can see the below error logs in NameNode log file.\r\n{quote}\r\nAny insights on when this can happen.. I checked your testcase which logout on other thread, but will be case..?\r\n\r\nAre you using kerboes-1.15 client or kerboes-1.13 client.?","created":"2023-01-08T19:40:04.883+0000"},{"body":"{quote}Are you planning to backport to other branches also..?\r\n{quote}\r\nYes [~brahmareddy], I will backport this.\r\n\r\n \r\n{quote}Any insights on when this can happen.\r\n{quote}\r\nYes, this issue happened in many prod cluster. Mostly this issue happened when one KDC is doing backup and it is not available for login request. When client trying to do the re-login but login failed because client is not able to failover to other available KDC server(failover failed because of wrong error code from first server).\r\n\r\n \r\n{quote}I checked your testcase which logout on other thread, but will be case..?\r\n{quote}\r\nYes, This is very common in NameNode and journalnode case. QJM in NameNode is client for JournalNode and it will do re-login as client, but when this re-login fail it will impact the NameNode also because [UGI#unprotectedRelogin()|https://github.com/apache/hadoop/blob/trunk/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/security/UserGroupInformation.java#L1361] destroy the NameNode ticket.\r\n\r\n ","created":"2023-01-11T03:54:36.184+0000"}],"conversations":[{"body":"Handle re-login in Server when client, server running in same JVM and client trying to re-login, but it fails.\r\n\r\nFor example, NameNode is server but in same JVM journal node client also running to push to edit logs. When JN client try to re-login and it fails, it will destroy server service ticket also and NameNode not able to server client request. We can see the below error logs in NameNode log file.\r\n\r\n \r\n{noformat}\r\nAuth failed for x.x.x.x:42199:null (GSS initiate failed) with true cause: (GSS initiate failed)\r\nAuth failed for x.x.x.x:42199:null (GSS initiate failed) with true cause: (GSS initiate failed)\r\nAuth failed for x.x.x.x:42199:null (GSS initiate failed) with true cause: (GSS initiate failed){noformat}\r\nSame discussion happened in HADOOP-17996.","from":"reporter","subject":"Handle Server KDC re-login when Server and Client run in same JVM."},{"body":"[~surendralilhore]  thanks for reporting and working on this. Are you planning to backport to other branches also..?\r\n{quote}When JN client try to re-login and it fails, it will destroy server service ticket also and NameNode not able to server client request. We can see the below error logs in NameNode log file.\r\n{quote}\r\nAny insights on when this can happen.. I checked your testcase which logout on other thread, but will be case..?\r\n\r\nAre you using kerboes-1.15 client or kerboes-1.13 client.?","from":"developer"},{"body":"{quote}Are you planning to backport to other branches also..?\r\n{quote}\r\nYes [~brahmareddy], I will backport this.\r\n\r\n \r\n{quote}Any insights on when this can happen.\r\n{quote}\r\nYes, this issue happened in many prod cluster. Mostly this issue happened when one KDC is doing backup and it is not available for login request. When client trying to do the re-login but login failed because client is not able to failover to other available KDC server(failover failed because of wrong error code from first server).\r\n\r\n \r\n{quote}I checked your testcase which logout on other thread, but will be case..?\r\n{quote}\r\nYes, This is very common in NameNode and journalnode case. QJM in NameNode is client for JournalNode and it will do re-login as client, but when this re-login fail it will impact the NameNode also because [UGI#unprotectedRelogin()|https://github.com/apache/hadoop/blob/trunk/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/security/UserGroupInformation.java#L1361] destroy the NameNode ticket.\r\n\r\n ","from":"developer"}],"created":"2022-12-20T10:01:37.000+0000","description":"Handle re-login in Server when client, server running in same JVM and client trying to re-login, but it fails.\r\n\r\nFor example, NameNode is server but in same JVM journal node client also running to push to edit logs. When JN client try to re-login and it fails, it will destroy server service ticket also and NameNode not able to server client request. We can see the below error logs in NameNode log file.\r\n\r\n \r\n{noformat}\r\nAuth failed for x.x.x.x:42199:null (GSS initiate failed) with true cause: (GSS initiate failed)\r\nAuth failed for x.x.x.x:42199:null (GSS initiate failed) with true cause: (GSS initiate failed)\r\nAuth failed for x.x.x.x:42199:null (GSS initiate failed) with true cause: (GSS initiate failed){noformat}\r\nSame discussion happened in HADOOP-17996.","issue_id":"13514848","key":"HADOOP-18581","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2023-01-08T18:26:52.000+0000","role":"fixed_distractor","summary":"Handle Server KDC re-login when Server and Client run in same JVM."} {"case_id":"13525194","cluster":"DISTRACTOR-HADOOP-18636","comments":[{"body":"{quote}when you ask actually for a file, it will look for the parent dir, but it calls {{{}mkdir(){}}}, rather than {{mkdirs()}}\r\n{quote}\r\nThis is the suspect line in question? [https://github.com/apache/hadoop/blob/7e19bc31b65f86be91451a0ec7590023b6b57a12/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/util/DiskChecker.java#L190]\r\n\r\nSneaky one.","created":"2023-02-17T13:12:55.237+0000"},{"body":"[~dannycjones] actually that method works...its recursive and does the right thing. \r\n\r\nwhat fails is the scan before that for available space -which turns out to exclude all directories which don't exist as they report 0 bytes free\r\nhttps://github.com/apache/hadoop/blob/7e19bc31b65f86be91451a0ec7590023b6b57a12/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/LocalDirAllocator.java#L417\r\n\r\nwe need to add the mkdirs() call above that line","created":"2023-02-17T19:27:09.827+0000"},{"body":"-3.3.6 release has been fixed, fix version removed 3.4.0.-","created":"2024-01-21T03:13:00.953+0000"}],"conversations":[{"body":"The s3a and abfs clients use LocalDirAllocator for allocating files in local (temporary) storage for buffering blocks to write, and, for the s3a staging committer, files being staged. \r\nWhen initialized (or when the configuration key value is updated) LocalDirAllocator enumerates all directories in the list and calls {{mkdirs()}} to create them.\r\n\r\nwhen you ask actually for a file, it will look for the parent dir, and will again call {{mkdirs()}}. \r\n\r\nBut before it does that, it looks to see if the dir has any space...if not it is excluded from the list of directories with room for data.\r\n\r\nAnd guess what: directories which don't exist report as having no space. So they get excluded -the recreation code doesn't get a chance to run.\r\n\r\n\r\n","from":"reporter","subject":"LocalDirAllocator cannot recover from directory tree deletion during the life of a filesystem client"},{"body":"{quote}when you ask actually for a file, it will look for the parent dir, but it calls {{{}mkdir(){}}}, rather than {{mkdirs()}}\r\n{quote}\r\nThis is the suspect line in question? [https://github.com/apache/hadoop/blob/7e19bc31b65f86be91451a0ec7590023b6b57a12/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/util/DiskChecker.java#L190]\r\n\r\nSneaky one.","from":"developer"},{"body":"[~dannycjones] actually that method works...its recursive and does the right thing. \r\n\r\nwhat fails is the scan before that for available space -which turns out to exclude all directories which don't exist as they report 0 bytes free\r\nhttps://github.com/apache/hadoop/blob/7e19bc31b65f86be91451a0ec7590023b6b57a12/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/LocalDirAllocator.java#L417\r\n\r\nwe need to add the mkdirs() call above that line","from":"developer"},{"body":"-3.3.6 release has been fixed, fix version removed 3.4.0.-","from":"developer"}],"created":"2023-02-17T11:40:35.000+0000","description":"The s3a and abfs clients use LocalDirAllocator for allocating files in local (temporary) storage for buffering blocks to write, and, for the s3a staging committer, files being staged. \r\nWhen initialized (or when the configuration key value is updated) LocalDirAllocator enumerates all directories in the list and calls {{mkdirs()}} to create them.\r\n\r\nwhen you ask actually for a file, it will look for the parent dir, and will again call {{mkdirs()}}. \r\n\r\nBut before it does that, it looks to see if the dir has any space...if not it is excluded from the list of directories with room for data.\r\n\r\nAnd guess what: directories which don't exist report as having no space. So they get excluded -the recreation code doesn't get a chance to run.\r\n\r\n\r\n","issue_id":"13525194","key":"HADOOP-18636","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2023-03-01T13:30:30.000+0000","role":"fixed_distractor","summary":"LocalDirAllocator cannot recover from directory tree deletion during the life of a filesystem client"} {"case_id":"13542678","cluster":"DISTRACTOR-HADOOP-18793","comments":[{"body":"I'm guessing ${UUID} directory is preserved on purpose as the commit job ID provided by Spark was not guaranteed to be unique historically, as described in the document.\r\nDeleting one's own staging directory might impact other ongoing commit jobs sharing the same ID (timestamp).\r\n\r\nhttps://hadoop.apache.org/docs/stable/hadoop-aws/tools/hadoop-aws/committers.html#Job_commit_fails_java.io.FileNotFoundException_.E2.80.9CFile_hdfs:.2F.2F....2Fstaging-uploads.2F_temporary.2F0_does_not_exist.E2.80.9D\r\n{quote}\r\nSpark generates job IDs for its committers using the current timestamp, and if two jobs/stages are started in the same second, they will have the same job ID.\r\n{quote}\r\n\r\nBut now I think it would be safe to delete it on a certain condition: fs.s3a.committer.require.uuid is enabled and there is almost no risk of ${UUID} collision.","created":"2023-07-06T10:49:00.170+0000"},{"body":"should be done in cleanupStagingDirs()\r\n\r\ncan you do a run at debug and make sure that \"Cleaning up work path\" is printed? \r\n\r\n\r\nwe should always be able to delete the dir as if the uuids are not unique then the jobs are corrupted.","created":"2023-07-06T15:18:13.888+0000"},{"body":"Hi [~stevel@apache.org], as I checked, workPath here seems to be pointing to a local fs path under fs.s3a.buffer.dir.\r\n\r\n{code:none}\r\n23/07/06 15:54:56 DEBUG AbstractS3ACommitter: Setting work path to file:/ssd/yarn/nm/usercache/${USER}/appcache/application_1687403148782_2731269/s3a/0e9bd28b-296c-4d6e-bff1-3f7049bd7ecd/_temporary/0/_temporary/attempt_202307061554568951051086712533391_0000_m_000000_0\r\n\r\n23/07/06 15:54:56 INFO AbstractS3ACommitterFactory: Using Commmitter StagingCommitter{AbstractS3ACommitter{role=Task committer attempt_202307061554568951051086712533391_0000_m_000000_0, name=directory, outputPath=s3a://${USER}/checkpoints/user/${USER}/7663e766-9630-434c-aa7d-a8f429b98dff/e67f2bbf-3850-4f63-9b4b-b722a5fbd81f, workPath=file:/ssd/yarn/nm/usercache/${USER}/appcache/application_1687403148782_2731269/s3a/0e9bd28b-296c-4d6e-bff1-3f7049bd7ecd/_temporary/0/_temporary/attempt_202307061554568951051086712533391_0000_m_000000_0}, conflictResolution=REPLACE, wrappedCommitter=FileOutputCommitter{PathOutputCommitter{context=TaskAttemptContextImpl{JobContextImpl{jobId=job_202307061554568951051086712533391_0000}; taskId=attempt_202307061554568951051086712533391_0000_m_000000_0, status=''}; org.apache.hadoop.mapreduce.lib.output.FileOutputCommitter@11ac570b}; outputPath=hdfs://nameservice1/user/${USER}/tmp/staging/${USER}/0e9bd28b-296c-4d6e-bff1-3f7049bd7ecd/staging-uploads, workPath=null, algorithmVersion=1, skipCleanup=false, ignoreCleanupFailures=true}} for s3a://${USER}/checkpoints/user/${USER}/7663e766-9630-434c-aa7d-a8f429b98dff/e67f2bbf-3850-4f63-9b4b-b722a5fbd81f\r\n\r\n23/07/06 15:54:56 DEBUG StagingCommitter: Cleaning up work path file:/ssd/yarn/nm/usercache/${USER}/appcache/application_1687403148782_2731269/s3a/0e9bd28b-296c-4d6e-bff1-3f7049bd7ecd/_temporary/0/_temporary/attempt_202307061554568951051086712533391_0000_m_000000_0\r\n23/07/06 15:54:56 INFO AbstractS3ACommitter: Task committer attempt_202307061554568951051086712533391_0000_m_000000_0: commitJob((no job ID)): duration 0:00.430s\r\n{code}","created":"2023-07-06T15:53:46.653+0000"},{"body":"oh, I see. looking at the staging committer, it does call wrappedCommitter.cleanupJob() which will delete the _temporary subdir under that staging dir which is where the .pendingset files go, but the parent dir is left alone. so although all files are cleaned up, dirs get leaked.\r\n\r\nhow about you provide a patch for this, with a new test case in ITestStagingCommitProtocol? \r\n","created":"2023-07-06T16:20:12.396+0000"},{"body":"Sure let me work on this.\r\n\r\nbq. we should always be able to delete the dir as if the uuids are not unique then the jobs are corrupted.\r\n\r\nI agree with you on this. I think we can always simply delete it.","created":"2023-07-06T16:43:04.152+0000"},{"body":"Hi [~stevel@apache.org] I have a similar issue related to this method: *cleanupStagingDirs()* using {*}magic committers{*}.\r\n\r\nWe have two spark jobs writing to the same s3a directory. We have the property *spark.hadoop.fs.s3a.committer.abort.pending.uploads=false*\r\nSo we see in logs this line: \r\n\r\nDEBUG [main] o.a.h.fs.s3a.commit.AbstractS3ACommitter (819): Not cleanup up pending uploads to s3a ...\r\n\r\n(from [https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/commit/AbstractS3ACommitter.java#L952])\r\n\r\nBut we also see in logs that when the first job finalize {*}the __magic directory is deleted{*}:\r\nINFO  [main] o.a.h.fs.s3a.commit.magic.MagicS3GuardCommitter (98): Deleting magic directory s3a://my-bucket/my-table/__magic: duration 0:00.560s\r\n\r\n(from  \r\n[https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/commit/magic/MagicS3GuardCommitter.java#L137])\r\n \r\nI'm not sure but I think that this is affecting to the second job that is still running.\r\nThe fact that __magic *is deleted recursively* (including all subdirectories like: __magic/job-1 , __magic/job-2 ...) , *could be a problem?*","created":"2023-07-07T20:51:02.567+0000"},{"body":"[~emanuelvelzi] that is HADOOP-18568; if someone can provide a patch with tests I will review it. ","created":"2023-07-08T11:44:02.383+0000"},{"body":"[~stevel@apache.org], I believe this is not exactly the same case.\r\n\r\nI want to perform a clean-up operation, but I want to ensure that I only clean the prefix associated with the job that is currently finalizing.\r\n\r\nI think something like this could be a good solution:\r\n{code:java}\r\n/**\r\n * Delete the magic directory.\r\n */\r\npublic void cleanupStagingDirs() {\r\n//Path path = magicSubdir(getOutputPath());\r\n  Path path = new Path(magicSubdir(getOutputPath()), formatAppAttemptDir(getUUID()));\r\n try(DurationInfo ignored = new DurationInfo(LOG, true,\r\n \"Deleting magic directory %s\", path)) {\r\n Invoker.ignoreIOExceptions(LOG, \"cleanup magic directory\", path.toString(),\r\n () -> deleteWithWarning(getDestFS(), path, true));\r\n }\r\n} {code}\r\n \r\nWhat do you think of this approach?","created":"2023-07-08T21:15:20.339+0000"},{"body":"[~stevel@apache.org] I have created a related Jira ticket, HADOOP-18797, to address my problem.\r\n\r\n(I apologize for interrupting this ticket with my issue.)","created":"2023-07-10T13:36:08.971+0000"}],"conversations":[{"body":"When setting up StagingCommitter and its internal FileOutputCommitter, a temporary directory that holds MPU information will be created on the default FS, which by default is to be /user/${USER}/tmp/staging/${USER}/${UUID}/staging-uploads.\r\n\r\nOn a successful job commit, its child directory (_temporary) will be [cleaned up|https://github.com/apache/hadoop/blob/a36d8adfd18e88f2752f4387ac4497aadd3a74e7/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/commit/staging/StagingCommitter.java#L516] properly, but ${UUID}/staging-uploads will remain.\r\n\r\nThis will result in having too many empty ${UUID}/staging-uploads directories under /user/${USER}/tmp/staging/${USER}, and will eventually cause an issue in an environment where the max number of items in a directory is capped (e.g. by dfs.namenode.fs-limits.max-directory-items in HDFS).\r\n\r\n{noformat}\r\nThe directory item limit of /user/${USER}/tmp/staging/${USER} is exceeded: limit=1048576 items=1048576\r\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.verifyMaxDirItems(FSDirectory.java:1205)\r\n{noformat}","from":"reporter","subject":"S3A StagingCommitter does not clean up staging-uploads directory"},{"body":"I'm guessing ${UUID} directory is preserved on purpose as the commit job ID provided by Spark was not guaranteed to be unique historically, as described in the document.\r\nDeleting one's own staging directory might impact other ongoing commit jobs sharing the same ID (timestamp).\r\n\r\nhttps://hadoop.apache.org/docs/stable/hadoop-aws/tools/hadoop-aws/committers.html#Job_commit_fails_java.io.FileNotFoundException_.E2.80.9CFile_hdfs:.2F.2F....2Fstaging-uploads.2F_temporary.2F0_does_not_exist.E2.80.9D\r\n{quote}\r\nSpark generates job IDs for its committers using the current timestamp, and if two jobs/stages are started in the same second, they will have the same job ID.\r\n{quote}\r\n\r\nBut now I think it would be safe to delete it on a certain condition: fs.s3a.committer.require.uuid is enabled and there is almost no risk of ${UUID} collision.","from":"developer"},{"body":"should be done in cleanupStagingDirs()\r\n\r\ncan you do a run at debug and make sure that \"Cleaning up work path\" is printed? \r\n\r\n\r\nwe should always be able to delete the dir as if the uuids are not unique then the jobs are corrupted.","from":"developer"},{"body":"Hi [~stevel@apache.org], as I checked, workPath here seems to be pointing to a local fs path under fs.s3a.buffer.dir.\r\n\r\n{code:none}\r\n23/07/06 15:54:56 DEBUG AbstractS3ACommitter: Setting work path to file:/ssd/yarn/nm/usercache/${USER}/appcache/application_1687403148782_2731269/s3a/0e9bd28b-296c-4d6e-bff1-3f7049bd7ecd/_temporary/0/_temporary/attempt_202307061554568951051086712533391_0000_m_000000_0\r\n\r\n23/07/06 15:54:56 INFO AbstractS3ACommitterFactory: Using Commmitter StagingCommitter{AbstractS3ACommitter{role=Task committer attempt_202307061554568951051086712533391_0000_m_000000_0, name=directory, outputPath=s3a://${USER}/checkpoints/user/${USER}/7663e766-9630-434c-aa7d-a8f429b98dff/e67f2bbf-3850-4f63-9b4b-b722a5fbd81f, workPath=file:/ssd/yarn/nm/usercache/${USER}/appcache/application_1687403148782_2731269/s3a/0e9bd28b-296c-4d6e-bff1-3f7049bd7ecd/_temporary/0/_temporary/attempt_202307061554568951051086712533391_0000_m_000000_0}, conflictResolution=REPLACE, wrappedCommitter=FileOutputCommitter{PathOutputCommitter{context=TaskAttemptContextImpl{JobContextImpl{jobId=job_202307061554568951051086712533391_0000}; taskId=attempt_202307061554568951051086712533391_0000_m_000000_0, status=''}; org.apache.hadoop.mapreduce.lib.output.FileOutputCommitter@11ac570b}; outputPath=hdfs://nameservice1/user/${USER}/tmp/staging/${USER}/0e9bd28b-296c-4d6e-bff1-3f7049bd7ecd/staging-uploads, workPath=null, algorithmVersion=1, skipCleanup=false, ignoreCleanupFailures=true}} for s3a://${USER}/checkpoints/user/${USER}/7663e766-9630-434c-aa7d-a8f429b98dff/e67f2bbf-3850-4f63-9b4b-b722a5fbd81f\r\n\r\n23/07/06 15:54:56 DEBUG StagingCommitter: Cleaning up work path file:/ssd/yarn/nm/usercache/${USER}/appcache/application_1687403148782_2731269/s3a/0e9bd28b-296c-4d6e-bff1-3f7049bd7ecd/_temporary/0/_temporary/attempt_202307061554568951051086712533391_0000_m_000000_0\r\n23/07/06 15:54:56 INFO AbstractS3ACommitter: Task committer attempt_202307061554568951051086712533391_0000_m_000000_0: commitJob((no job ID)): duration 0:00.430s\r\n{code}","from":"developer"},{"body":"oh, I see. looking at the staging committer, it does call wrappedCommitter.cleanupJob() which will delete the _temporary subdir under that staging dir which is where the .pendingset files go, but the parent dir is left alone. so although all files are cleaned up, dirs get leaked.\r\n\r\nhow about you provide a patch for this, with a new test case in ITestStagingCommitProtocol? \r\n","from":"developer"},{"body":"Sure let me work on this.\r\n\r\nbq. we should always be able to delete the dir as if the uuids are not unique then the jobs are corrupted.\r\n\r\nI agree with you on this. I think we can always simply delete it.","from":"developer"},{"body":"Hi [~stevel@apache.org] I have a similar issue related to this method: *cleanupStagingDirs()* using {*}magic committers{*}.\r\n\r\nWe have two spark jobs writing to the same s3a directory. We have the property *spark.hadoop.fs.s3a.committer.abort.pending.uploads=false*\r\nSo we see in logs this line: \r\n\r\nDEBUG [main] o.a.h.fs.s3a.commit.AbstractS3ACommitter (819): Not cleanup up pending uploads to s3a ...\r\n\r\n(from [https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/commit/AbstractS3ACommitter.java#L952])\r\n\r\nBut we also see in logs that when the first job finalize {*}the __magic directory is deleted{*}:\r\nINFO  [main] o.a.h.fs.s3a.commit.magic.MagicS3GuardCommitter (98): Deleting magic directory s3a://my-bucket/my-table/__magic: duration 0:00.560s\r\n\r\n(from  \r\n[https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/commit/magic/MagicS3GuardCommitter.java#L137])\r\n \r\nI'm not sure but I think that this is affecting to the second job that is still running.\r\nThe fact that __magic *is deleted recursively* (including all subdirectories like: __magic/job-1 , __magic/job-2 ...) , *could be a problem?*","from":"developer"},{"body":"[~emanuelvelzi] that is HADOOP-18568; if someone can provide a patch with tests I will review it. ","from":"developer"},{"body":"[~stevel@apache.org], I believe this is not exactly the same case.\r\n\r\nI want to perform a clean-up operation, but I want to ensure that I only clean the prefix associated with the job that is currently finalizing.\r\n\r\nI think something like this could be a good solution:\r\n{code:java}\r\n/**\r\n * Delete the magic directory.\r\n */\r\npublic void cleanupStagingDirs() {\r\n//Path path = magicSubdir(getOutputPath());\r\n  Path path = new Path(magicSubdir(getOutputPath()), formatAppAttemptDir(getUUID()));\r\n try(DurationInfo ignored = new DurationInfo(LOG, true,\r\n \"Deleting magic directory %s\", path)) {\r\n Invoker.ignoreIOExceptions(LOG, \"cleanup magic directory\", path.toString(),\r\n () -> deleteWithWarning(getDestFS(), path, true));\r\n }\r\n} {code}\r\n \r\nWhat do you think of this approach?","from":"developer"},{"body":"[~stevel@apache.org] I have created a related Jira ticket, HADOOP-18797, to address my problem.\r\n\r\n(I apologize for interrupting this ticket with my issue.)","from":"developer"}],"created":"2023-07-06T10:39:44.000+0000","description":"When setting up StagingCommitter and its internal FileOutputCommitter, a temporary directory that holds MPU information will be created on the default FS, which by default is to be /user/${USER}/tmp/staging/${USER}/${UUID}/staging-uploads.\r\n\r\nOn a successful job commit, its child directory (_temporary) will be [cleaned up|https://github.com/apache/hadoop/blob/a36d8adfd18e88f2752f4387ac4497aadd3a74e7/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/commit/staging/StagingCommitter.java#L516] properly, but ${UUID}/staging-uploads will remain.\r\n\r\nThis will result in having too many empty ${UUID}/staging-uploads directories under /user/${USER}/tmp/staging/${USER}, and will eventually cause an issue in an environment where the max number of items in a directory is capped (e.g. by dfs.namenode.fs-limits.max-directory-items in HDFS).\r\n\r\n{noformat}\r\nThe directory item limit of /user/${USER}/tmp/staging/${USER} is exceeded: limit=1048576 items=1048576\r\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.verifyMaxDirItems(FSDirectory.java:1205)\r\n{noformat}","issue_id":"13542678","key":"HADOOP-18793","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2023-07-08T13:04:02.000+0000","role":"fixed_distractor","summary":"S3A StagingCommitter does not clean up staging-uploads directory"} {"case_id":"13544857","cluster":"DISTRACTOR-HADOOP-18826","comments":[{"body":"based on your analysis, HADOOP-16612 is the cause","created":"2023-07-26T13:23:17.691+0000"},{"body":"but looking into getRelativePath() is a likely change HADOOP-16916\r\n\r\nnow: why doesn't anyone else see this?","created":"2023-07-26T13:25:34.263+0000"},{"body":"PR for fix: [Hadoop 18826: [ABFS] Fix for Empty Relative Path Issue Leading to GetFileStatus(\"/\") failure. by anujmodi2021 · Pull Request #5909 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/5909]","created":"2023-07-31T10:59:11.536+0000"},{"body":"[~sergey.shabalov]: this bug has been fixed and the next hadoop 3.3.x release should include it.\r\n\r\nif you need it yourself before then, you are going to have to do your own build of a suitable branch with the patch.","created":"2023-08-09T15:36:36.139+0000"}],"conversations":[{"body":"I am using hadoop-azure-3.3.0.jar and have written code:\r\n{code:java}\r\nstatic final String ROOT_DIR = \"abfs://ssh-test-fs@sshadlsgen2.dfs.core.windows.net\",\r\nConfiguration config = new Configuration();        config.set(\"fs.defaultFS\",ROOT_DIR);        config.set(\"fs.adl.oauth2.access.token.provider.type\",\"ClientCredential\");        config.set(\"fs.adl.oauth2.client.id\",\"\");        config.set(\"fs.adl.oauth2.credential\",\"\");        config.set(\"fs.adl.oauth2.refresh.url\",\"\");        config.set(\"fs.azure.account.key.sshadlsgen2.dfs.core.windows.net\",ACCESS_TOKEN);        config.set(\"fs.azure.skipUserGroupMetadataDuringInitialization\",\"true\");\r\n\tFileSystem fs = FileSystem.get(config);\r\n\tSystem.out.println( \"\\nfs:'\"+fs.toString()+\"'\");\r\n\tFileStatus status = fs.getFileStatus(new Path(ROOT_DIR)); // !!! Exception in 3.3.1-*\r\n\tSystem.out.println( \"\\nstatus:'\"+status.toString()+\"'\");\r\n {code}\r\nIt did work properly till 3.3.1. \r\n\r\nBut in 3.3.1 it fails with exception:\r\n{code:java}\r\nCaused by: Operation failed: \"Value for one of the query parameters specified in the request URI is invalid.\", 400, HEAD, https://sshadlsgen2.dfs.core.windows.net/ssh-test-fs?upn=false&action=getAccessControl&timeout=90 at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.completeExecute(AbfsRestOperation.java:218) at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.lambda$execute$0(AbfsRestOperation.java:181) at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.measureDurationOfInvocation(IOStatisticsBinding.java:494) at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.trackDurationOfInvocation(IOStatisticsBinding.java:465) at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.execute(AbfsRestOperation.java:179) at org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:942) at org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:924) at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.getFileStatus(AzureBlobFileSystemStore.java:846) at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem.getFileStatus(AzureBlobFileSystem.java:507) {code}\r\nI performed some research and found:\r\n\r\nIn hadoop-azure-3.3.0.jar we see:\r\n{code:java}\r\norg.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore{\r\n\t...\r\n\tpublic FileStatus getFileStatus(final Path path) throws IOException {\r\n\t...\r\nLine 604:\t\top = client.getAclStatus(AbfsHttpConstants.FORWARD_SLASH + AbfsHttpConstants.ROOT_PATH);\r\n\t...\r\n\t}\r\n\t...\r\n} {code}\r\nand this code produces REST request:\r\n{code:java}\r\nhttps://sshadlsgen2.dfs.core.windows.net/ssh-test-fs//?upn=false&action=getAccessControl&timeout=90\r\n  {code}\r\nThere is finalizes slash in path part \"...ssh-test-fs{*}{color:#de350b}//{color}{*}?upn=false...\" This request does work properly.\r\n\r\nBut since hadoop-azure-3.3.1.jar till latest hadoop-azure-3.3.6.jar we see:\r\n{code:java}\r\norg.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore {\r\n\t...\r\n\tpublic FileStatus getFileStatus(final Path path) throws IOException {\r\n\t\t...\r\n\t\t\t\tperfInfo.registerCallee(\"getAclStatus\");\r\nLine 846: op = client.getAclStatus(getRelativePath(path));\r\n\t\t...\r\n\t}\r\n\t...\r\n}\r\nLine 1492:\r\nprivate String getRelativePath(final Path path) {\r\n\t...\r\n\treturn path.toUri().getPath();\r\n} {code}\r\nand this code prduces REST request:\r\n{code:java}\r\nhttps://sshadlsgen2.dfs.core.windows.net/ssh-test-fs?upn=false&action=getAccessControl&timeout=90 {code}\r\nThere is not finalizes slash in path part \"...ssh-test-fs?upn=false...\" It happens because the new code \"path.toUri().getPath();\" produces empty string.\r\n\r\nThis request fails with message:\r\n{code:java}\r\nCaused by: Operation failed: \"Value for one of the query parameters specified in the request URI is invalid.\", 400, HEAD, https://sshadlsgen2.dfs.core.windows.net/ssh-test-fs?upn=false&action=getAccessControl&timeout=90\r\n\tat org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.completeExecute(AbfsRestOperation.java:218)\r\n\tat org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.lambda$execute$0(AbfsRestOperation.java:181)\r\n\tat org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.measureDurationOfInvocation(IOStatisticsBinding.java:494)\r\n\tat org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.trackDurationOfInvocation(IOStatisticsBinding.java:465)\r\n\tat org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.execute(AbfsRestOperation.java:179)\r\n\tat org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:942)\r\n\tat org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:924)\r\n\tat org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.getFileStatus(AzureBlobFileSystemStore.java:846)\r\n\tat org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem.getFileStatus(AzureBlobFileSystem.java:507) {code}\r\nSuch us it is for all hadoop-azure-3.3.*.jar versions which does use log4j 2.* not 1.2.17 we can't update using version\r\n\r\n \r\n\r\nI attach a sample of Maven project to try: test_hadoop-azure-3_3_1-FileSystem_getFileStatus - Copy.zip","from":"reporter","subject":"abfs getFileStatus(/) fails with \"Value for one of the query parameters specified in the request URI is invalid.\", 400"},{"body":"based on your analysis, HADOOP-16612 is the cause","from":"developer"},{"body":"but looking into getRelativePath() is a likely change HADOOP-16916\r\n\r\nnow: why doesn't anyone else see this?","from":"developer"},{"body":"PR for fix: [Hadoop 18826: [ABFS] Fix for Empty Relative Path Issue Leading to GetFileStatus(\"/\") failure. by anujmodi2021 · Pull Request #5909 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/5909]","from":"developer"},{"body":"[~sergey.shabalov]: this bug has been fixed and the next hadoop 3.3.x release should include it.\r\n\r\nif you need it yourself before then, you are going to have to do your own build of a suitable branch with the patch.","from":"developer"}],"created":"2023-07-25T20:35:06.000+0000","description":"I am using hadoop-azure-3.3.0.jar and have written code:\r\n{code:java}\r\nstatic final String ROOT_DIR = \"abfs://ssh-test-fs@sshadlsgen2.dfs.core.windows.net\",\r\nConfiguration config = new Configuration();        config.set(\"fs.defaultFS\",ROOT_DIR);        config.set(\"fs.adl.oauth2.access.token.provider.type\",\"ClientCredential\");        config.set(\"fs.adl.oauth2.client.id\",\"\");        config.set(\"fs.adl.oauth2.credential\",\"\");        config.set(\"fs.adl.oauth2.refresh.url\",\"\");        config.set(\"fs.azure.account.key.sshadlsgen2.dfs.core.windows.net\",ACCESS_TOKEN);        config.set(\"fs.azure.skipUserGroupMetadataDuringInitialization\",\"true\");\r\n\tFileSystem fs = FileSystem.get(config);\r\n\tSystem.out.println( \"\\nfs:'\"+fs.toString()+\"'\");\r\n\tFileStatus status = fs.getFileStatus(new Path(ROOT_DIR)); // !!! Exception in 3.3.1-*\r\n\tSystem.out.println( \"\\nstatus:'\"+status.toString()+\"'\");\r\n {code}\r\nIt did work properly till 3.3.1. \r\n\r\nBut in 3.3.1 it fails with exception:\r\n{code:java}\r\nCaused by: Operation failed: \"Value for one of the query parameters specified in the request URI is invalid.\", 400, HEAD, https://sshadlsgen2.dfs.core.windows.net/ssh-test-fs?upn=false&action=getAccessControl&timeout=90 at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.completeExecute(AbfsRestOperation.java:218) at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.lambda$execute$0(AbfsRestOperation.java:181) at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.measureDurationOfInvocation(IOStatisticsBinding.java:494) at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.trackDurationOfInvocation(IOStatisticsBinding.java:465) at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.execute(AbfsRestOperation.java:179) at org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:942) at org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:924) at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.getFileStatus(AzureBlobFileSystemStore.java:846) at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem.getFileStatus(AzureBlobFileSystem.java:507) {code}\r\nI performed some research and found:\r\n\r\nIn hadoop-azure-3.3.0.jar we see:\r\n{code:java}\r\norg.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore{\r\n\t...\r\n\tpublic FileStatus getFileStatus(final Path path) throws IOException {\r\n\t...\r\nLine 604:\t\top = client.getAclStatus(AbfsHttpConstants.FORWARD_SLASH + AbfsHttpConstants.ROOT_PATH);\r\n\t...\r\n\t}\r\n\t...\r\n} {code}\r\nand this code produces REST request:\r\n{code:java}\r\nhttps://sshadlsgen2.dfs.core.windows.net/ssh-test-fs//?upn=false&action=getAccessControl&timeout=90\r\n  {code}\r\nThere is finalizes slash in path part \"...ssh-test-fs{*}{color:#de350b}//{color}{*}?upn=false...\" This request does work properly.\r\n\r\nBut since hadoop-azure-3.3.1.jar till latest hadoop-azure-3.3.6.jar we see:\r\n{code:java}\r\norg.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore {\r\n\t...\r\n\tpublic FileStatus getFileStatus(final Path path) throws IOException {\r\n\t\t...\r\n\t\t\t\tperfInfo.registerCallee(\"getAclStatus\");\r\nLine 846: op = client.getAclStatus(getRelativePath(path));\r\n\t\t...\r\n\t}\r\n\t...\r\n}\r\nLine 1492:\r\nprivate String getRelativePath(final Path path) {\r\n\t...\r\n\treturn path.toUri().getPath();\r\n} {code}\r\nand this code prduces REST request:\r\n{code:java}\r\nhttps://sshadlsgen2.dfs.core.windows.net/ssh-test-fs?upn=false&action=getAccessControl&timeout=90 {code}\r\nThere is not finalizes slash in path part \"...ssh-test-fs?upn=false...\" It happens because the new code \"path.toUri().getPath();\" produces empty string.\r\n\r\nThis request fails with message:\r\n{code:java}\r\nCaused by: Operation failed: \"Value for one of the query parameters specified in the request URI is invalid.\", 400, HEAD, https://sshadlsgen2.dfs.core.windows.net/ssh-test-fs?upn=false&action=getAccessControl&timeout=90\r\n\tat org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.completeExecute(AbfsRestOperation.java:218)\r\n\tat org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.lambda$execute$0(AbfsRestOperation.java:181)\r\n\tat org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.measureDurationOfInvocation(IOStatisticsBinding.java:494)\r\n\tat org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.trackDurationOfInvocation(IOStatisticsBinding.java:465)\r\n\tat org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.execute(AbfsRestOperation.java:179)\r\n\tat org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:942)\r\n\tat org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:924)\r\n\tat org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.getFileStatus(AzureBlobFileSystemStore.java:846)\r\n\tat org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem.getFileStatus(AzureBlobFileSystem.java:507) {code}\r\nSuch us it is for all hadoop-azure-3.3.*.jar versions which does use log4j 2.* not 1.2.17 we can't update using version\r\n\r\n \r\n\r\nI attach a sample of Maven project to try: test_hadoop-azure-3_3_1-FileSystem_getFileStatus - Copy.zip","issue_id":"13544857","key":"HADOOP-18826","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2023-08-08T18:07:22.000+0000","role":"fixed_distractor","summary":"abfs getFileStatus(/) fails with \"Value for one of the query parameters specified in the request URI is invalid.\", 400"} {"case_id":"13549002","cluster":"DISTRACTOR-HADOOP-18870","comments":[{"body":"[~bender]\r\n\r\n{quote}\r\nProposing routing the 4-parameter method to a 5-parameter method, which instantiates the ZKConfiguration as the 5th parameter. This is a non-breaking change, as the ZKConfiguration is currently instantiated within the method.\r\n{quote}\r\n\r\nAm I missing something or you meant `ZKClientConfig` as the 5th parameter, right?\r\nChecking the linked Curator PR (https://github.com/apache/curator/pull/391/files#diff-687a4ed1252bfb4f56b3aeeb28bee4413b7df9bec4b969b72215587158ac875dR59) shows me ZKClientConfig as the 5th parameter there.\r\nCan you fix the description of the jira ?","created":"2023-09-07T01:25:55.254+0000"},{"body":"[~snemeth] you are correct. Fixed the description of the Jira. Thank you for spotting it and for the CR/merge.","created":"2023-09-07T09:01:46.384+0000"}],"conversations":[{"body":"[Curator PR#391 |https://github.com/apache/curator/pull/391/files#diff-687a4ed1252bfb4f56b3aeeb28bee4413b7df9bec4b969b72215587158ac875dR59] introduced a default method in the ZooKeeperFactory interface, hence the override of the 4-parameter NewZookeeper method in the HadoopZookeeperFactory class is not taking effect due to this. \r\n\r\nProposing routing the 4-parameter method to a 5-parameter method, which instantiates the ZKClientConfig as the 5th parameter. This is a non-breaking change, as the ZKClientConfig is currently instantiated within the method.","from":"reporter","subject":"CURATOR-599 change broke functionality introduced in HADOOP-18139 and HADOOP-18709"},{"body":"[~bender]\r\n\r\n{quote}\r\nProposing routing the 4-parameter method to a 5-parameter method, which instantiates the ZKConfiguration as the 5th parameter. This is a non-breaking change, as the ZKConfiguration is currently instantiated within the method.\r\n{quote}\r\n\r\nAm I missing something or you meant `ZKClientConfig` as the 5th parameter, right?\r\nChecking the linked Curator PR (https://github.com/apache/curator/pull/391/files#diff-687a4ed1252bfb4f56b3aeeb28bee4413b7df9bec4b969b72215587158ac875dR59) shows me ZKClientConfig as the 5th parameter there.\r\nCan you fix the description of the jira ?","from":"developer"},{"body":"[~snemeth] you are correct. Fixed the description of the Jira. Thank you for spotting it and for the CR/merge.","from":"developer"}],"created":"2023-08-29T15:04:45.000+0000","description":"[Curator PR#391 |https://github.com/apache/curator/pull/391/files#diff-687a4ed1252bfb4f56b3aeeb28bee4413b7df9bec4b969b72215587158ac875dR59] introduced a default method in the ZooKeeperFactory interface, hence the override of the 4-parameter NewZookeeper method in the HadoopZookeeperFactory class is not taking effect due to this. \r\n\r\nProposing routing the 4-parameter method to a 5-parameter method, which instantiates the ZKClientConfig as the 5th parameter. This is a non-breaking change, as the ZKClientConfig is currently instantiated within the method.","issue_id":"13549002","key":"HADOOP-18870","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2023-09-07T01:37:44.000+0000","role":"fixed_distractor","summary":"CURATOR-599 change broke functionality introduced in HADOOP-18139 and HADOOP-18709"} {"case_id":"13553554","cluster":"DISTRACTOR-HADOOP-18929","comments":[{"body":"Hi [~mthakur] \r\n\r\nI think it is due to HADOOP-18895\r\nThe build failed in the branch-3.3 PR here: [https://github.com/apache/hadoop/pull/6073#issuecomment-1721005715] it had -1 for mvn install & shadedclient\r\n\r\nCan you revert that & try, it works for me.\r\n\r\ncc. [~stevel@apache.org]","created":"2023-10-10T16:22:32.639+0000"},{"body":"Thanks, [~ayushtkn]  for checking quickly. Let me revert and try. ","created":"2023-10-10T16:48:50.144+0000"},{"body":"It looks like commons-compress 1.24.0 is the 1st commons-compress jar to have module-info.class in it.\r\n\r\nIf you are amenable, I can do a PR that excludes the commons-compress 1.24.0 module-info.class from hadoop-client-minicluster and hadoop-client-runtime jars.","created":"2023-10-10T20:02:48.264+0000"},{"body":"Oh okay. A quick followup PR will do. Thanks","created":"2023-10-10T20:04:51.901+0000"},{"body":"https://github.com/apache/hadoop/pull/6169","created":"2023-10-10T20:13:02.746+0000"},{"body":"There are quite few third party jar fixes merged in branch 3.3  around last 1 year after 3.3.6 in Jun 2023\r\n\r\nIs there any 3.3.x  next release planned soon ?","created":"2024-06-06T07:48:21.302+0000"}],"conversations":[{"body":"{noformat}\r\n[ESC[1;34mINFOESC[m] ESC[1m-------< ESC[0;36morg.apache.hadoop:hadoop-client-check-test-invariantsESC[0;1m >--------ESC[m\r\n[ESC[1;34mINFOESC[m] ESC[1mBuilding Apache Hadoop Client Packaging Invariants for Test 3.3.9-SNAPSHOT [105/111]ESC[m\r\n[ESC[1;34mINFOESC[m] ESC[1m--------------------------------[ pom ]---------------------------------ESC[m\r\n[ESC[1;34mINFOESC[m] \r\n[ESC[1;34mINFOESC[m] ESC[1m--- ESC[0;32mmaven-enforcer-plugin:3.0.0-M1:enforceESC[m ESC[1m(enforce-banned-dependencies)ESC[m @ ESC[36mhadoop-client-check-test-invariantsESC[0;1m ---ESC[m\r\n[ESC[1;34mINFOESC[m] Adding ignorable dependency: org.apache.hadoop:hadoop-annotations:null\r\n[ESC[1;34mINFOESC[m]   Adding ignore: *\r\n[ESC[1;33mWARNINGESC[m] Rule 1: org.apache.maven.plugins.enforcer.BanDuplicateClasses failed with message:\r\nDuplicate classes found:\r\n\r\n\r\n  Found in:\r\n    org.apache.hadoop:hadoop-client-minicluster:jar:3.3.9-SNAPSHOT:compile\r\n    org.apache.hadoop:hadoop-client-runtime:jar:3.3.9-SNAPSHOT:compile\r\n  Duplicate classes:\r\n    META-INF/versions/9/module-info.class\r\n\r\n{noformat}\r\nCC [~stevel@apache.org]  [~weichu] ","from":"reporter","subject":"Build failure while trying to create apache 3.3.7 release locally."},{"body":"Hi [~mthakur] \r\n\r\nI think it is due to HADOOP-18895\r\nThe build failed in the branch-3.3 PR here: [https://github.com/apache/hadoop/pull/6073#issuecomment-1721005715] it had -1 for mvn install & shadedclient\r\n\r\nCan you revert that & try, it works for me.\r\n\r\ncc. [~stevel@apache.org]","from":"developer"},{"body":"Thanks, [~ayushtkn]  for checking quickly. Let me revert and try. ","from":"developer"},{"body":"It looks like commons-compress 1.24.0 is the 1st commons-compress jar to have module-info.class in it.\r\n\r\nIf you are amenable, I can do a PR that excludes the commons-compress 1.24.0 module-info.class from hadoop-client-minicluster and hadoop-client-runtime jars.","from":"developer"},{"body":"Oh okay. A quick followup PR will do. Thanks","from":"developer"},{"body":"https://github.com/apache/hadoop/pull/6169","from":"developer"},{"body":"There are quite few third party jar fixes merged in branch 3.3  around last 1 year after 3.3.6 in Jun 2023\r\n\r\nIs there any 3.3.x  next release planned soon ?","from":"developer"}],"created":"2023-10-10T15:43:08.000+0000","description":"{noformat}\r\n[ESC[1;34mINFOESC[m] ESC[1m-------< ESC[0;36morg.apache.hadoop:hadoop-client-check-test-invariantsESC[0;1m >--------ESC[m\r\n[ESC[1;34mINFOESC[m] ESC[1mBuilding Apache Hadoop Client Packaging Invariants for Test 3.3.9-SNAPSHOT [105/111]ESC[m\r\n[ESC[1;34mINFOESC[m] ESC[1m--------------------------------[ pom ]---------------------------------ESC[m\r\n[ESC[1;34mINFOESC[m] \r\n[ESC[1;34mINFOESC[m] ESC[1m--- ESC[0;32mmaven-enforcer-plugin:3.0.0-M1:enforceESC[m ESC[1m(enforce-banned-dependencies)ESC[m @ ESC[36mhadoop-client-check-test-invariantsESC[0;1m ---ESC[m\r\n[ESC[1;34mINFOESC[m] Adding ignorable dependency: org.apache.hadoop:hadoop-annotations:null\r\n[ESC[1;34mINFOESC[m]   Adding ignore: *\r\n[ESC[1;33mWARNINGESC[m] Rule 1: org.apache.maven.plugins.enforcer.BanDuplicateClasses failed with message:\r\nDuplicate classes found:\r\n\r\n\r\n  Found in:\r\n    org.apache.hadoop:hadoop-client-minicluster:jar:3.3.9-SNAPSHOT:compile\r\n    org.apache.hadoop:hadoop-client-runtime:jar:3.3.9-SNAPSHOT:compile\r\n  Duplicate classes:\r\n    META-INF/versions/9/module-info.class\r\n\r\n{noformat}\r\nCC [~stevel@apache.org]  [~weichu] ","issue_id":"13553554","key":"HADOOP-18929","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2023-10-13T20:18:43.000+0000","role":"fixed_distractor","summary":"Build failure while trying to create apache 3.3.7 release locally."} {"case_id":"13571527","cluster":"DISTRACTOR-HADOOP-19106","comments":[{"body":"CC [~snvijaya]  [~pranavs] ","created":"2024-03-11T19:09:15.100+0000"},{"body":"Thanks for reporting this [~mthakur] \r\nHowever, I do not see these tests failing.\r\nThese tests require auth type to be set as \"SharedKey\" and works with only HNS accounts therefore also need \"fs.azure.test.namespace.enabled\" to be set to true. \r\n\r\nWith these two configs present, tests are working fine for me.\r\n\r\nPlease share the exact configuration you are using to better investigate the issue.","created":"2024-03-12T03:47:08.507+0000"},{"body":"It does fail for me with the same config mentioned in  [https://github.com/apache/hadoop/pull/6069#issuecomment-1965105331] + fs.azure.test.namespace.enabled=true. \r\n\r\n ","created":"2024-03-12T21:28:07.489+0000"},{"body":"It fails because [https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-azure/src/test/java/org/apache/hadoop/fs/azurebfs/ITestAzureBlobFileSystemAuthorization.java#L360]  returns null. \r\n\r\nand this only gets initialized when authType is SAS [https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-azure/src/main/java/org/apache/hadoop/fs/azurebfs/AzureBlobFileSystemStore.java#L1733] \r\n\r\n \r\n\r\n ","created":"2024-03-12T22:21:38.276+0000"},{"body":"Working on it...\r\nThanks for update","created":"2024-03-14T04:28:54.608+0000"},{"body":"[~mthakur] \r\nRaised a PR with fix: [HADOOP-19129: [ABFS] Test Fixes and Test Script Bug Fixes by anujmodi2021 · Pull Request #6676 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/6676]\r\n\r\nPlease review","created":"2024-04-01T09:35:51.699+0000"},{"body":"[HADOOP-19129: [ABFS] Test Fixes and Test Script Bug Fixes by anujmodi2021 · Pull Request #6676 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/6676]","created":"2024-04-13T12:24:16.508+0000"},{"body":"[~anujmodi2021]  this doesn't seem to be fixed yet? \r\n\r\n ","created":"2024-09-16T18:32:14.599+0000"},{"body":"apart from this, I see following failures in trunk. Not sure if these are my config related. \r\n{code:java}\r\n[ERROR]   ITestAzureBlobFileSystemAuthorization.testSASTokenProviderEmptySASToken:99 Expected an exception of type class org.apache.hadoop.fs.azurebfs.contracts.exceptions.SASTokenProviderException\r\n[ERROR]   ITestAzureBlobFileSystemAuthorization.testSASTokenProviderInitializeException:82 Expected an exception of type class org.apache.hadoop.fs.azurebfs.contracts.exceptions.SASTokenProviderException\r\n[ERROR]   ITestAzureBlobFileSystemAuthorization.testSASTokenProviderNullSASToken:114 Expected an exception of type class org.apache.hadoop.fs.azurebfs.contracts.exceptions.SASTokenProviderException\r\n[ERROR]   ITestAzureBlobFileSystemChooseSAS.testBothProviderFixedTokenUnset:171 Expected an exception of type class org.apache.hadoop.fs.azurebfs.contracts.exceptions.SASTokenProviderException\r\n[ERROR]   ITestFileSystemInitialization.testFileSystemCapabilities:94 [path capability fs.capability.etags.available in AzureBlobFileSystem{uri=abfs://abfs-testcontainer-ce2755d3-ff36-4196-9653-816309b8f440@mthakurdata.dfs.core.windows.net, user='mthakur', primaryUserGroup='staff'[fs.azure.capability.readahead.safe]}] expected:<[tru]e> but was:<[fals]e>\r\n[ERROR] Errors: \r\n[ERROR]   ITestAbfsReadFooterMetrics.testMetricWithIdlePeriod:331 » InvalidUri Invalid U...\r\n[ERROR]   ITestAbfsReadFooterMetrics.testReadFooterMetrics:114 » InvalidUri Invalid URI ...\r\n[ERROR]   ITestAbfsReadFooterMetrics.testReadFooterMetricsWithParquetAndNonParquet:60->testReadWriteAndSeek:217 » InvalidUri\r\n[ERROR]   ITestAbfsRestOperationException.testAuthFailException:189->AbstractAbfsIntegrationTest.getFileSystem:318 » InvalidConfigurationValue\r\n[ERROR]   ITestAbfsRestOperationException.testCustomTokenFetchRetryCount:136->testWithDifferentCustomTokenFetchRetry:155 » InvalidConfigurationValue \r\n[ERROR]   ITestAzureBlobFileSystemAuthorization.testSetPermissionUnauthorized:218->runTest:272 NullPointer\r\n[ERROR]   ITestAzureBlobFileSystemChooseSAS.testBothProviderFixedTokenConfigured:107 » SASTokenProvider\r\n[ERROR]   ITestAzureBlobFileSystemChooseSAS.testOnlyFixedTokenConfigured:132->testOnlyFixedTokenConfiguredInternal:146 » SASTokenProvider\r\n[ERROR]   ITestAzureBlobFileSystemE2E.testHttpReadTimeout »  Unexpected exception, expec...\r\n[ERROR]   ITestAzureBlobFileSystemInitAndCreate.testFileSystemInitFailsWithBlobEndpoitUrl:121->lambda$testFileSystemInitFailsWithBlobEndpoitUrl$0:123 » KeyProvider\r\n[ERROR]   ITestAbfsHttpClientRequestExecutor.testConnectionReadRecords:330->lambda$null$11:308->lambda$mockHttpOperationBehavior$6:223 ClassCast\r\n[ERROR]   ITestAbfsHttpClientRequestExecutor.testExpect100ContinueHandling:190->lambda$testExpect100ContinueHandling$3:198 » IO\r\n[ERROR]   ITestExponentialRetryPolicy.testThrottlingIntercept:106 » KeyProvider Failure ...\r\n[INFO] \r\n[ERROR] Failures: \r\n[ERROR]   ITestAbfsRenameStageFailure>TestRenameStageFailure.testDeleteTargetPaths:262 [Etag of destination file abfs://abfs-test-new@mthakurdata.dfs.core.windows.net/fork-0002/test/testDeleteTargetPaths/source.txt] expected: but was:<\"0x8DCD68B103B3FF4\">\r\n[ERROR]   ITestAbfsFileSystemContractEtag>AbstractContractEtagTest.testEtagConsistencyAcrossListAndHead:60 [path capability fs.capability.etags.available of abfs://abfs-test-new@mthakurdata.dfs.core.windows.net/fork-0002/test/testEtagConsistencyAcrossListAndHead] expected:<[tru]e> but was:<[fals]e>{code}\r\n\r\nCould you and [~pranavs]  please check these and see what all needs fixing. Thanks \r\n\r\n ","created":"2024-09-16T20:33:09.148+0000"},{"body":"Thanks for reporting [~mthakur]\r\nCan you please share what all configs you have added when facing these failures...\r\nI will be able to debug better\r\n\r\nThanks.","created":"2024-09-17T05:16:30.910+0000"},{"body":"It's the same config mentioned in [https://github.com/apache/hadoop/pull/6069#issuecomment-1965105331] \r\n\r\nalong with \r\n\r\n\r\n\r\n    fs.azure.test.namespace.enabled\r\n\r\n    true\r\n\r\n\r\n\r\n ","created":"2024-09-17T16:48:09.119+0000"},{"body":"Hi [~mthakur] \r\nSorry for the delay. I finally got some time to work on these issues.\r\nHave created a PR for the fixes: [HADOOP-18960: [ABFS] Making Contract tests run in sequential and Other Test Fixes by anujmodi2021 · Pull Request #7104 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/7104]","created":"2024-10-09T09:17:28.081+0000"},{"body":"Ran a test suite with the configs that you have [~mthakur] \r\n\r\nHere are the results: Metric related failures are known and fixed in https://github.com/apache/hadoop/pull/6847\r\n\r\n------------------------------\r\n:::: AGGREGATED TEST RESULT ::::\r\n\r\n============================================================\r\nHNS-SharedKey\r\n============================================================\r\n\r\n[ERROR] testBackoffRetryMetrics(org.apache.hadoop.fs.azurebfs.services.TestAbfsRestOperation)  Time elapsed: 3.822 s  <<< ERROR!\r\n[ERROR] testReadFooterMetrics(org.apache.hadoop.fs.azurebfs.ITestAbfsReadFooterMetrics)  Time elapsed: 1.135 s  <<< ERROR!\r\n[ERROR] testMetricWithIdlePeriod(org.apache.hadoop.fs.azurebfs.ITestAbfsReadFooterMetrics)  Time elapsed: 1.187 s  <<< ERROR!\r\n[ERROR] testReadFooterMetricsWithParquetAndNonParquet(org.apache.hadoop.fs.azurebfs.ITestAbfsReadFooterMetrics)  Time elapsed: 1.174 s  <<< ERROR!\r\n\r\n[ERROR] Tests run: 157, Failures: 0, Errors: 1, Skipped: 2\r\n[ERROR] Tests run: 652, Failures: 0, Errors: 3, Skipped: 98\r\n[WARNING] Tests run: 171, Failures: 0, Errors: 0, Skipped: 25\r\n[WARNING] Tests run: 262, Failures: 0, Errors: 0, Skipped: 10","created":"2024-10-09T12:04:24.775+0000"},{"body":"This has been merged to trunk.","created":"2024-11-12T11:23:21.633+0000"}],"conversations":[{"body":"When below config set to true all of the tests fails else it skips.\r\n\r\n\r\n\r\n    fs.azure.test.namespace.enabled\r\n\r\n    true\r\n\r\n\r\n\r\n \r\n\r\n[*ERROR*] testOpenFileAuthorized(org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemAuthorization)  Time elapsed: 0.064 s  <<< ERROR!\r\n\r\njava.lang.NullPointerException\r\n\r\n at org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemAuthorization.runTest(ITestAzureBlobFileSystemAuthorization.java:273)\r\n\r\n at org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemAuthorization.testOpenFileAuthorized(ITestAzureBlobFileSystemAuthorization.java:132)\r\n\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n\r\n at org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:59)\r\n\r\n at org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\r\n\r\n at org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:56)\r\n\r\n at org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\r\n\r\n at org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:26)\r\n\r\n at org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:27)","from":"reporter","subject":"[ABFS] All tests of. ITestAzureBlobFileSystemAuthorization fails with NPE"},{"body":"CC [~snvijaya]  [~pranavs] ","from":"developer"},{"body":"Thanks for reporting this [~mthakur] \r\nHowever, I do not see these tests failing.\r\nThese tests require auth type to be set as \"SharedKey\" and works with only HNS accounts therefore also need \"fs.azure.test.namespace.enabled\" to be set to true. \r\n\r\nWith these two configs present, tests are working fine for me.\r\n\r\nPlease share the exact configuration you are using to better investigate the issue.","from":"developer"},{"body":"It does fail for me with the same config mentioned in  [https://github.com/apache/hadoop/pull/6069#issuecomment-1965105331] + fs.azure.test.namespace.enabled=true. \r\n\r\n ","from":"developer"},{"body":"It fails because [https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-azure/src/test/java/org/apache/hadoop/fs/azurebfs/ITestAzureBlobFileSystemAuthorization.java#L360]  returns null. \r\n\r\nand this only gets initialized when authType is SAS [https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-azure/src/main/java/org/apache/hadoop/fs/azurebfs/AzureBlobFileSystemStore.java#L1733] \r\n\r\n \r\n\r\n ","from":"developer"},{"body":"Working on it...\r\nThanks for update","from":"developer"},{"body":"[~mthakur] \r\nRaised a PR with fix: [HADOOP-19129: [ABFS] Test Fixes and Test Script Bug Fixes by anujmodi2021 · Pull Request #6676 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/6676]\r\n\r\nPlease review","from":"developer"},{"body":"[HADOOP-19129: [ABFS] Test Fixes and Test Script Bug Fixes by anujmodi2021 · Pull Request #6676 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/6676]","from":"developer"},{"body":"[~anujmodi2021]  this doesn't seem to be fixed yet? \r\n\r\n ","from":"developer"},{"body":"apart from this, I see following failures in trunk. Not sure if these are my config related. \r\n{code:java}\r\n[ERROR]   ITestAzureBlobFileSystemAuthorization.testSASTokenProviderEmptySASToken:99 Expected an exception of type class org.apache.hadoop.fs.azurebfs.contracts.exceptions.SASTokenProviderException\r\n[ERROR]   ITestAzureBlobFileSystemAuthorization.testSASTokenProviderInitializeException:82 Expected an exception of type class org.apache.hadoop.fs.azurebfs.contracts.exceptions.SASTokenProviderException\r\n[ERROR]   ITestAzureBlobFileSystemAuthorization.testSASTokenProviderNullSASToken:114 Expected an exception of type class org.apache.hadoop.fs.azurebfs.contracts.exceptions.SASTokenProviderException\r\n[ERROR]   ITestAzureBlobFileSystemChooseSAS.testBothProviderFixedTokenUnset:171 Expected an exception of type class org.apache.hadoop.fs.azurebfs.contracts.exceptions.SASTokenProviderException\r\n[ERROR]   ITestFileSystemInitialization.testFileSystemCapabilities:94 [path capability fs.capability.etags.available in AzureBlobFileSystem{uri=abfs://abfs-testcontainer-ce2755d3-ff36-4196-9653-816309b8f440@mthakurdata.dfs.core.windows.net, user='mthakur', primaryUserGroup='staff'[fs.azure.capability.readahead.safe]}] expected:<[tru]e> but was:<[fals]e>\r\n[ERROR] Errors: \r\n[ERROR]   ITestAbfsReadFooterMetrics.testMetricWithIdlePeriod:331 » InvalidUri Invalid U...\r\n[ERROR]   ITestAbfsReadFooterMetrics.testReadFooterMetrics:114 » InvalidUri Invalid URI ...\r\n[ERROR]   ITestAbfsReadFooterMetrics.testReadFooterMetricsWithParquetAndNonParquet:60->testReadWriteAndSeek:217 » InvalidUri\r\n[ERROR]   ITestAbfsRestOperationException.testAuthFailException:189->AbstractAbfsIntegrationTest.getFileSystem:318 » InvalidConfigurationValue\r\n[ERROR]   ITestAbfsRestOperationException.testCustomTokenFetchRetryCount:136->testWithDifferentCustomTokenFetchRetry:155 » InvalidConfigurationValue \r\n[ERROR]   ITestAzureBlobFileSystemAuthorization.testSetPermissionUnauthorized:218->runTest:272 NullPointer\r\n[ERROR]   ITestAzureBlobFileSystemChooseSAS.testBothProviderFixedTokenConfigured:107 » SASTokenProvider\r\n[ERROR]   ITestAzureBlobFileSystemChooseSAS.testOnlyFixedTokenConfigured:132->testOnlyFixedTokenConfiguredInternal:146 » SASTokenProvider\r\n[ERROR]   ITestAzureBlobFileSystemE2E.testHttpReadTimeout »  Unexpected exception, expec...\r\n[ERROR]   ITestAzureBlobFileSystemInitAndCreate.testFileSystemInitFailsWithBlobEndpoitUrl:121->lambda$testFileSystemInitFailsWithBlobEndpoitUrl$0:123 » KeyProvider\r\n[ERROR]   ITestAbfsHttpClientRequestExecutor.testConnectionReadRecords:330->lambda$null$11:308->lambda$mockHttpOperationBehavior$6:223 ClassCast\r\n[ERROR]   ITestAbfsHttpClientRequestExecutor.testExpect100ContinueHandling:190->lambda$testExpect100ContinueHandling$3:198 » IO\r\n[ERROR]   ITestExponentialRetryPolicy.testThrottlingIntercept:106 » KeyProvider Failure ...\r\n[INFO] \r\n[ERROR] Failures: \r\n[ERROR]   ITestAbfsRenameStageFailure>TestRenameStageFailure.testDeleteTargetPaths:262 [Etag of destination file abfs://abfs-test-new@mthakurdata.dfs.core.windows.net/fork-0002/test/testDeleteTargetPaths/source.txt] expected: but was:<\"0x8DCD68B103B3FF4\">\r\n[ERROR]   ITestAbfsFileSystemContractEtag>AbstractContractEtagTest.testEtagConsistencyAcrossListAndHead:60 [path capability fs.capability.etags.available of abfs://abfs-test-new@mthakurdata.dfs.core.windows.net/fork-0002/test/testEtagConsistencyAcrossListAndHead] expected:<[tru]e> but was:<[fals]e>{code}\r\n\r\nCould you and [~pranavs]  please check these and see what all needs fixing. Thanks \r\n\r\n ","from":"developer"},{"body":"Thanks for reporting [~mthakur]\r\nCan you please share what all configs you have added when facing these failures...\r\nI will be able to debug better\r\n\r\nThanks.","from":"developer"},{"body":"It's the same config mentioned in [https://github.com/apache/hadoop/pull/6069#issuecomment-1965105331] \r\n\r\nalong with \r\n\r\n\r\n\r\n    fs.azure.test.namespace.enabled\r\n\r\n    true\r\n\r\n\r\n\r\n ","from":"developer"},{"body":"Hi [~mthakur] \r\nSorry for the delay. I finally got some time to work on these issues.\r\nHave created a PR for the fixes: [HADOOP-18960: [ABFS] Making Contract tests run in sequential and Other Test Fixes by anujmodi2021 · Pull Request #7104 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/7104]","from":"developer"},{"body":"Ran a test suite with the configs that you have [~mthakur] \r\n\r\nHere are the results: Metric related failures are known and fixed in https://github.com/apache/hadoop/pull/6847\r\n\r\n------------------------------\r\n:::: AGGREGATED TEST RESULT ::::\r\n\r\n============================================================\r\nHNS-SharedKey\r\n============================================================\r\n\r\n[ERROR] testBackoffRetryMetrics(org.apache.hadoop.fs.azurebfs.services.TestAbfsRestOperation)  Time elapsed: 3.822 s  <<< ERROR!\r\n[ERROR] testReadFooterMetrics(org.apache.hadoop.fs.azurebfs.ITestAbfsReadFooterMetrics)  Time elapsed: 1.135 s  <<< ERROR!\r\n[ERROR] testMetricWithIdlePeriod(org.apache.hadoop.fs.azurebfs.ITestAbfsReadFooterMetrics)  Time elapsed: 1.187 s  <<< ERROR!\r\n[ERROR] testReadFooterMetricsWithParquetAndNonParquet(org.apache.hadoop.fs.azurebfs.ITestAbfsReadFooterMetrics)  Time elapsed: 1.174 s  <<< ERROR!\r\n\r\n[ERROR] Tests run: 157, Failures: 0, Errors: 1, Skipped: 2\r\n[ERROR] Tests run: 652, Failures: 0, Errors: 3, Skipped: 98\r\n[WARNING] Tests run: 171, Failures: 0, Errors: 0, Skipped: 25\r\n[WARNING] Tests run: 262, Failures: 0, Errors: 0, Skipped: 10","from":"developer"},{"body":"This has been merged to trunk.","from":"developer"}],"created":"2024-03-11T19:08:43.000+0000","description":"When below config set to true all of the tests fails else it skips.\r\n\r\n\r\n\r\n    fs.azure.test.namespace.enabled\r\n\r\n    true\r\n\r\n\r\n\r\n \r\n\r\n[*ERROR*] testOpenFileAuthorized(org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemAuthorization)  Time elapsed: 0.064 s  <<< ERROR!\r\n\r\njava.lang.NullPointerException\r\n\r\n at org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemAuthorization.runTest(ITestAzureBlobFileSystemAuthorization.java:273)\r\n\r\n at org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemAuthorization.testOpenFileAuthorized(ITestAzureBlobFileSystemAuthorization.java:132)\r\n\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n\r\n at org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:59)\r\n\r\n at org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\r\n\r\n at org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:56)\r\n\r\n at org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\r\n\r\n at org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:26)\r\n\r\n at org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:27)","issue_id":"13571527","key":"HADOOP-19106","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-11-12T11:23:58.000+0000","role":"fixed_distractor","summary":"[ABFS] All tests of. ITestAzureBlobFileSystemAuthorization fails with NPE"} {"case_id":"13571787","cluster":"DISTRACTOR-HADOOP-19110","comments":[{"body":"Thanks for reporting this...\r\nIt has already been to our attention. I am working on a PR to fix all such failures...","created":"2024-03-14T04:30:09.476+0000"},{"body":"[~mthakur] \r\nRaised a PR with fix: [HADOOP-19129: [ABFS] Test Fixes and Test Script Bug Fixes by anujmodi2021 · Pull Request #6676 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/6676]\r\n\r\nPlease review","created":"2024-04-01T09:08:21.801+0000"},{"body":"[HADOOP-19129: [ABFS] Test Fixes and Test Script Bug Fixes by anujmodi2021 · Pull Request #6676 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/6676]","created":"2024-04-13T12:23:48.629+0000"},{"body":"[HADOOP-19129: [ABFS] Test Fixes and Test Script Bug Fixes by anujmodi2021 · Pull Request #6676 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/6676]","created":"2024-04-13T12:25:38.250+0000"}],"conversations":[{"body":"{code:java}\r\n[ERROR] Tests run: 6, Failures: 0, Errors: 1, Skipped: 2, Time elapsed: 91.416 s <<< FAILURE! - in org.apache.hadoop.fs.azurebfs.services.ITestExponentialRetryPolicy\r\n[ERROR] testThrottlingIntercept(org.apache.hadoop.fs.azurebfs.services.ITestExponentialRetryPolicy)  Time elapsed: 0.622 s  <<< ERROR!\r\nFailure to initialize configuration for dummy.dfs.core.windows.net key =\"null\": Invalid configuration value detected for fs.azure.account.key\r\n\tat org.apache.hadoop.fs.azurebfs.services.SimpleKeyProvider.getStorageAccountKey(SimpleKeyProvider.java:53)\r\n\tat org.apache.hadoop.fs.azurebfs.AbfsConfiguration.getStorageAccountKey(AbfsConfiguration.java:646)\r\n\tat org.apache.hadoop.fs.azurebfs.services.ITestAbfsClient.createTestClientFromCurrentContext(ITestAbfsClient.java:339)\r\n\tat org.apache.hadoop.fs.azurebfs.services.ITestExponentialRetryPolicy.testThrottlingIntercept(ITestExponentialRetryPolicy.java:106)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n\tat java.lang.reflect.Method.invoke(Method.java:498)\r\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:59) {code}","from":"reporter","subject":"ITestExponentialRetryPolicy failing in branch-3.4"},{"body":"Thanks for reporting this...\r\nIt has already been to our attention. I am working on a PR to fix all such failures...","from":"developer"},{"body":"[~mthakur] \r\nRaised a PR with fix: [HADOOP-19129: [ABFS] Test Fixes and Test Script Bug Fixes by anujmodi2021 · Pull Request #6676 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/6676]\r\n\r\nPlease review","from":"developer"},{"body":"[HADOOP-19129: [ABFS] Test Fixes and Test Script Bug Fixes by anujmodi2021 · Pull Request #6676 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/6676]","from":"developer"},{"body":"[HADOOP-19129: [ABFS] Test Fixes and Test Script Bug Fixes by anujmodi2021 · Pull Request #6676 · apache/hadoop (github.com)|https://github.com/apache/hadoop/pull/6676]","from":"developer"}],"created":"2024-03-13T18:15:32.000+0000","description":"{code:java}\r\n[ERROR] Tests run: 6, Failures: 0, Errors: 1, Skipped: 2, Time elapsed: 91.416 s <<< FAILURE! - in org.apache.hadoop.fs.azurebfs.services.ITestExponentialRetryPolicy\r\n[ERROR] testThrottlingIntercept(org.apache.hadoop.fs.azurebfs.services.ITestExponentialRetryPolicy)  Time elapsed: 0.622 s  <<< ERROR!\r\nFailure to initialize configuration for dummy.dfs.core.windows.net key =\"null\": Invalid configuration value detected for fs.azure.account.key\r\n\tat org.apache.hadoop.fs.azurebfs.services.SimpleKeyProvider.getStorageAccountKey(SimpleKeyProvider.java:53)\r\n\tat org.apache.hadoop.fs.azurebfs.AbfsConfiguration.getStorageAccountKey(AbfsConfiguration.java:646)\r\n\tat org.apache.hadoop.fs.azurebfs.services.ITestAbfsClient.createTestClientFromCurrentContext(ITestAbfsClient.java:339)\r\n\tat org.apache.hadoop.fs.azurebfs.services.ITestExponentialRetryPolicy.testThrottlingIntercept(ITestExponentialRetryPolicy.java:106)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n\tat java.lang.reflect.Method.invoke(Method.java:498)\r\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:59) {code}","issue_id":"13571787","key":"HADOOP-19110","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-04-13T12:22:34.000+0000","role":"fixed_distractor","summary":"ITestExponentialRetryPolicy failing in branch-3.4"} {"case_id":"13578246","cluster":"DISTRACTOR-HADOOP-19167","comments":[{"body":"[~skyskyhu] Sorry, I have no permission to set you as a contributor. So I cannot assignee this ticket to you. I just close it.\r\n\r\n \r\n\r\n[~slfan1989] Master, please help to do it when you are available. Thanks ","created":"2024-05-17T02:35:52.863+0000"}],"conversations":[{"body":"In one of my projects, I need to dynamically adjust compression level for different files. \r\nHowever, I found that in most cases the new compression level does not take effect as expected, the old compression level continues to be used.\r\n\r\nHere is the relevant code snippet:\r\nZStandardCodec zStandardCodec = new ZStandardCodec();\r\nzStandardCodec.setConf(conf);\r\nconf.set(\"io.compression.codec.zstd.level\", \"5\"); // level may change dynamically\r\nconf.set(\"io.compression.codec.zstd\", zStandardCodec.getClass().getName());\r\nwriter = SequenceFile.createWriter(conf, SequenceFile.Writer.file(sequenceFilePath),\r\n                                SequenceFile.Writer.keyClass(LongWritable.class),\r\n                                SequenceFile.Writer.valueClass(BytesWritable.class),\r\n                                SequenceFile.Writer.compression(CompressionType.BLOCK));\r\n\r\nThe reason is SequenceFile.Writer.init() method will call CodecPool.getCompressor(codec, null) to get a compressor. \r\nIf the compressor is a reused instance, the conf is not applied because it is passed as null:\r\npublic static Compressor getCompressor(CompressionCodec codec, Configuration conf) {\r\nCompressor compressor = borrow(compressorPool, codec.getCompressorType());\r\nif (compressor == null)\r\n\r\n{ compressor = codec.createCompressor(); LOG.info(\"Got brand-new compressor [\"+codec.getDefaultExtension()+\"]\"); }\r\n\r\nelse {\r\ncompressor.reinit(conf);   //conf is null here\r\n......\r\n\r\n \r\n\r\nPlease also refer to my unit test to reproduce the bug. \r\nTo address this bug, I modified the code to ensure that the configuration is read back from the codec when a compressor is reused.","from":"reporter","subject":"Change of Codec configuration does not work"},{"body":"[~skyskyhu] Sorry, I have no permission to set you as a contributor. So I cannot assignee this ticket to you. I just close it.\r\n\r\n \r\n\r\n[~slfan1989] Master, please help to do it when you are available. Thanks ","from":"developer"}],"created":"2024-05-06T08:50:38.000+0000","description":"In one of my projects, I need to dynamically adjust compression level for different files. \r\nHowever, I found that in most cases the new compression level does not take effect as expected, the old compression level continues to be used.\r\n\r\nHere is the relevant code snippet:\r\nZStandardCodec zStandardCodec = new ZStandardCodec();\r\nzStandardCodec.setConf(conf);\r\nconf.set(\"io.compression.codec.zstd.level\", \"5\"); // level may change dynamically\r\nconf.set(\"io.compression.codec.zstd\", zStandardCodec.getClass().getName());\r\nwriter = SequenceFile.createWriter(conf, SequenceFile.Writer.file(sequenceFilePath),\r\n                                SequenceFile.Writer.keyClass(LongWritable.class),\r\n                                SequenceFile.Writer.valueClass(BytesWritable.class),\r\n                                SequenceFile.Writer.compression(CompressionType.BLOCK));\r\n\r\nThe reason is SequenceFile.Writer.init() method will call CodecPool.getCompressor(codec, null) to get a compressor. \r\nIf the compressor is a reused instance, the conf is not applied because it is passed as null:\r\npublic static Compressor getCompressor(CompressionCodec codec, Configuration conf) {\r\nCompressor compressor = borrow(compressorPool, codec.getCompressorType());\r\nif (compressor == null)\r\n\r\n{ compressor = codec.createCompressor(); LOG.info(\"Got brand-new compressor [\"+codec.getDefaultExtension()+\"]\"); }\r\n\r\nelse {\r\ncompressor.reinit(conf);   //conf is null here\r\n......\r\n\r\n \r\n\r\nPlease also refer to my unit test to reproduce the bug. \r\nTo address this bug, I modified the code to ensure that the configuration is read back from the codec when a compressor is reused.","issue_id":"13578246","key":"HADOOP-19167","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-05-17T02:36:07.000+0000","role":"fixed_distractor","summary":"Change of Codec configuration does not work"} {"case_id":"13591442","cluster":"DISTRACTOR-HADOOP-19270","comments":[{"body":"Hi Kim,\r\n\r\nI am a newbie to this platform and this community.\r\nI would like to be guided by you folks to step into an issue and resolving it.\r\n\r\nI added a small code change by adding a sequence number that you suggested.\r\nNow do I need to create a PR for the same to get verified the changes right ? ","created":"2024-09-10T12:03:25.493+0000"},{"body":"[~nish4_nth]\r\nHello Nishanth!\r\n\r\nThank you for getting interested in this issue!\r\nI already made commit for this issue, and will make pr soon.\r\nSo instead of making pr, I would like to ask you to review my pr that will be made soon.\r\n\r\nIf you would like to review, I would mention to you in comment when I make pr!","created":"2024-09-10T14:40:25.816+0000"},{"body":"sure kim\r\nplease share your pr here for review once created.\r\nI think i can learn something from this.\r\n\r\nThanks for your words","created":"2024-09-10T15:01:06.559+0000"},{"body":"[~nish4_nth] \r\n\r\nI finally wrote pr!","created":"2024-10-21T04:58:53.328+0000"},{"body":"wow. that's great [~geatrigger]. \r\neven though I am not in a place to review/judge, I'll surely take a look at it as it could help me learn something from this.\r\n\r\nps : Is there any way that we could connect outside this maybe Instagram/X ? . I have some questions that could be helped by you if you are ok with it.\r\n\r\nThanks and Good Luck.(*)","created":"2024-10-21T05:34:34.086+0000"},{"body":"[~nish4_nth] Well, I have a mail instead.\r\n\r\nIf you have any questions about code or whatever, send mail to [geatrigger@gmail.com|mailto:geatrigger@gmail.com]\r\n\r\nThank you.","created":"2024-10-24T06:13:56.688+0000"},{"body":"Merged to trunk.","created":"2025-03-04T22:45:51.139+0000"},{"body":"[~xkrogen] Thanks for the review! [~geatrigger] Thanks for the contribution! Welcome to hadoop! add hadoop-common contributor.","created":"2025-03-05T01:25:55.969+0000"}],"conversations":[{"body":"h2. Purpose\r\n - To remove possibility of wrong-ordered log simulation\r\n\r\nh2. Why this happens?\r\n - private DelayQueue commandQueue is actually PriorityQueue that use unstable sort.\r\n - commandQueue can have order that is not same to original audit log order.\r\n - In real production, there is the commands that occur same time and should be fixed order.\r\n{code:bash}\r\n# getfileinfo before open\r\n2024-07-01 19:27:12,886 INFO FSNamesystem.audit: allowed=true ugi=xx-xx (auth:TOKEN) via hive/hadoop.example.com@EXAMPLE.PROD (auth:TOKEN) ip=/10.xx.xxx.xxx cmd=getfileinfo src=/user/hive/warehouse/a.db/b/date_id=2024-06-16/part-xxxx.gz.parquet dst=null perm=null proto=rpc\r\n2024-07-01 19:27:12,886 INFO FSNamesystem.audit: allowed=true ugi=xx-xx (auth:TOKEN) via hive/hadoop.example.com@EXAMPLE.PROD (auth:TOKEN) ip=/10.xx.xxx.xxx cmd=open src=/user/hive/warehouse/a.db/b/date_id=2024-06-16/part-xxxx.gz.parquet dst=null perm=null proto=rpc\r\n\r\n# create before setPermission\r\n# this examples have not exactly same time, but could be same when rate factor is high enough\r\n2024-07-01 17:25:30,867 INFO FSNamesystem.audit: allowed=true ugi=yy-yy@EXAMPLE.PROD (auth:KERBEROS) ip=/10.xxx.xx.xxx cmd=create src=/user/yy-yy/.staging/job_1716867484406_290658/job.xml dst=null perm=yy-yy:zzz:rw-rw-r-- proto=rpc\r\n2024-07-01 17:25:30,871 INFO FSNamesystem.audit: allowed=true ugi=yy-yy@EXAMPLE.PROD (auth:KERBEROS) ip=/10.xxx.xx.xxx cmd=setPermission src=/user/yy-yy/.staging/job_1716867484406_290658/job.xml dst=null perm=yy-yy:zzz:rw-r--r-- proto=rpc\r\n{code}\r\n\r\nh2. How much improve test accuracy when use stable sort?\r\n - Using stable sort, wrong ordered simulation could not occur.\r\n -- I fixed code to use line number of audit log in sorting criteria.\r\n -- Because it is not simple to change DelayQueue data structure to use stable sort\r\n - Multi threading or client-ip-based-partitioning could be occur in real production and affect log order, but not critical.\r\n -- Client-ip-based-partitioning is even similar to real production choas log order\r\n - This is the graph that\r\n -- use real production hdfs audit log\r\n -- compare stable sort and unstable sort with different rate(1~4)\r\n -- use 5 minutes simulation(in rate 1) ip-based-partitioned-audit-log\r\n -- shows total valid command, total read latency, total write latency\r\n!image-2024-09-19-16-40-34-947.png|width=615,height=740!\r\n - Conclusion\r\n -- Stable sort ensure almost similar valid command number.\r\n -- Unstable sort sometimes extremely high latency because of wrong ordered log simulation.","from":"reporter","subject":"Use stable sort in commandQueue"},{"body":"Hi Kim,\r\n\r\nI am a newbie to this platform and this community.\r\nI would like to be guided by you folks to step into an issue and resolving it.\r\n\r\nI added a small code change by adding a sequence number that you suggested.\r\nNow do I need to create a PR for the same to get verified the changes right ? ","from":"developer"},{"body":"[~nish4_nth]\r\nHello Nishanth!\r\n\r\nThank you for getting interested in this issue!\r\nI already made commit for this issue, and will make pr soon.\r\nSo instead of making pr, I would like to ask you to review my pr that will be made soon.\r\n\r\nIf you would like to review, I would mention to you in comment when I make pr!","from":"developer"},{"body":"sure kim\r\nplease share your pr here for review once created.\r\nI think i can learn something from this.\r\n\r\nThanks for your words","from":"developer"},{"body":"[~nish4_nth] \r\n\r\nI finally wrote pr!","from":"developer"},{"body":"wow. that's great [~geatrigger]. \r\neven though I am not in a place to review/judge, I'll surely take a look at it as it could help me learn something from this.\r\n\r\nps : Is there any way that we could connect outside this maybe Instagram/X ? . I have some questions that could be helped by you if you are ok with it.\r\n\r\nThanks and Good Luck.(*)","from":"developer"},{"body":"[~nish4_nth] Well, I have a mail instead.\r\n\r\nIf you have any questions about code or whatever, send mail to [geatrigger@gmail.com|mailto:geatrigger@gmail.com]\r\n\r\nThank you.","from":"developer"},{"body":"Merged to trunk.","from":"developer"},{"body":"[~xkrogen] Thanks for the review! [~geatrigger] Thanks for the contribution! Welcome to hadoop! add hadoop-common contributor.","from":"developer"}],"created":"2024-09-09T06:16:52.000+0000","description":"h2. Purpose\r\n - To remove possibility of wrong-ordered log simulation\r\n\r\nh2. Why this happens?\r\n - private DelayQueue commandQueue is actually PriorityQueue that use unstable sort.\r\n - commandQueue can have order that is not same to original audit log order.\r\n - In real production, there is the commands that occur same time and should be fixed order.\r\n{code:bash}\r\n# getfileinfo before open\r\n2024-07-01 19:27:12,886 INFO FSNamesystem.audit: allowed=true ugi=xx-xx (auth:TOKEN) via hive/hadoop.example.com@EXAMPLE.PROD (auth:TOKEN) ip=/10.xx.xxx.xxx cmd=getfileinfo src=/user/hive/warehouse/a.db/b/date_id=2024-06-16/part-xxxx.gz.parquet dst=null perm=null proto=rpc\r\n2024-07-01 19:27:12,886 INFO FSNamesystem.audit: allowed=true ugi=xx-xx (auth:TOKEN) via hive/hadoop.example.com@EXAMPLE.PROD (auth:TOKEN) ip=/10.xx.xxx.xxx cmd=open src=/user/hive/warehouse/a.db/b/date_id=2024-06-16/part-xxxx.gz.parquet dst=null perm=null proto=rpc\r\n\r\n# create before setPermission\r\n# this examples have not exactly same time, but could be same when rate factor is high enough\r\n2024-07-01 17:25:30,867 INFO FSNamesystem.audit: allowed=true ugi=yy-yy@EXAMPLE.PROD (auth:KERBEROS) ip=/10.xxx.xx.xxx cmd=create src=/user/yy-yy/.staging/job_1716867484406_290658/job.xml dst=null perm=yy-yy:zzz:rw-rw-r-- proto=rpc\r\n2024-07-01 17:25:30,871 INFO FSNamesystem.audit: allowed=true ugi=yy-yy@EXAMPLE.PROD (auth:KERBEROS) ip=/10.xxx.xx.xxx cmd=setPermission src=/user/yy-yy/.staging/job_1716867484406_290658/job.xml dst=null perm=yy-yy:zzz:rw-r--r-- proto=rpc\r\n{code}\r\n\r\nh2. How much improve test accuracy when use stable sort?\r\n - Using stable sort, wrong ordered simulation could not occur.\r\n -- I fixed code to use line number of audit log in sorting criteria.\r\n -- Because it is not simple to change DelayQueue data structure to use stable sort\r\n - Multi threading or client-ip-based-partitioning could be occur in real production and affect log order, but not critical.\r\n -- Client-ip-based-partitioning is even similar to real production choas log order\r\n - This is the graph that\r\n -- use real production hdfs audit log\r\n -- compare stable sort and unstable sort with different rate(1~4)\r\n -- use 5 minutes simulation(in rate 1) ip-based-partitioned-audit-log\r\n -- shows total valid command, total read latency, total write latency\r\n!image-2024-09-19-16-40-34-947.png|width=615,height=740!\r\n - Conclusion\r\n -- Stable sort ensure almost similar valid command number.\r\n -- Unstable sort sometimes extremely high latency because of wrong ordered log simulation.","issue_id":"13591442","key":"HADOOP-19270","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-03-04T22:45:51.000+0000","role":"fixed_distractor","summary":"Use stable sort in commandQueue"} {"case_id":"13591677","cluster":"DISTRACTOR-HADOOP-19271","comments":[{"body":"{code}\r\nSLF4J: Failed toString() invocation on an object of type [org.apache.hadoop.fs.azurebfs.services.AbfsManagedApacheHttpConnection]\r\nReported exception:\r\njava.lang.NullPointerException\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsManagedApacheHttpConnection.toString(AbfsManagedApacheHttpConnection.java:233)\r\n at org.slf4j.helpers.MessageFormatter.safeObjectAppend(MessageFormatter.java:277)\r\n at org.slf4j.helpers.MessageFormatter.deeplyAppendParameter(MessageFormatter.java:249)\r\n at org.slf4j.helpers.MessageFormatter.arrayFormat(MessageFormatter.java:211)\r\n at org.slf4j.helpers.MessageFormatter.arrayFormat(MessageFormatter.java:161)\r\n at org.slf4j.impl.Reload4jLoggerAdapter.debug(Reload4jLoggerAdapter.java:251)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsConnectionManager.logDebug(AbfsConnectionManager.java:204)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsConnectionManager.connect(AbfsConnectionManager.java:155)\r\n at org.apache.http.impl.execchain.MainClientExec.establishRoute(MainClientExec.java:393)\r\n at org.apache.http.impl.execchain.MainClientExec.execute(MainClientExec.java:236)\r\n at org.apache.http.impl.execchain.ProtocolExec.execute(ProtocolExec.java:186)\r\n at org.apache.http.impl.client.InternalHttpClient.doExecute(InternalHttpClient.java:185)\r\n at org.apache.http.impl.client.CloseableHttpClient.execute(CloseableHttpClient.java:83)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsApacheHttpClient.execute(AbfsApacheHttpClient.java:122)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsAHCHttpOperation.executeRequest(AbfsAHCHttpOperation.java:266)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsAHCHttpOperation.processResponse(AbfsAHCHttpOperation.java:199)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.executeHttpOperation(AbfsRestOperation.java:392)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.completeExecute(AbfsRestOperation.java:292)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.lambda$execute$0(AbfsRestOperation.java:258)\r\n at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.measureDurationOfInvocation(IOStatisticsBinding.java:494)\r\n at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.trackDurationOfInvocation(IOStatisticsBinding.java:465)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.execute(AbfsRestOperation.java:256)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:1403)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:1386)\r\n at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.getAclStatus(AzureBlobFileSystemStore.java:1673)\r\n at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem.getAclStatus(AzureBlobFileSystem.java:1302)\r\n at org.apache.hadoop.fs.store.diag.AbfsDiagnosticsInfo.validateFilesystem(AbfsDiagnosticsInfo.java:626)\r\n at org.apache.hadoop.fs.store.diag.StoreDiag.executeFileSystemOperations(StoreDiag.java:774)\r\n at org.apache.hadoop.fs.store.diag.StoreDiag.run(StoreDiag.java:230)\r\n at org.apache.hadoop.fs.store.diag.StoreDiag.run(StoreDiag.java:172)\r\n at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:82)\r\n at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:97)\r\n at org.apache.hadoop.fs.store.diag.StoreDiag.exec(StoreDiag.java:1189)\r\n at org.apache.hadoop.fs.store.diag.StoreDiag.main(StoreDiag.java:1198)\r\n at storediag.main(storediag.java:25)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at org.apache.hadoop.util.RunJar.run(RunJar.java:330)\r\n at org.apache.hadoop.util.RunJar.main(RunJar.java:245)\r\n2024-09-10 17:28:25,557 [main] DEBUG services.AbfsConnectionManager (AbfsConnectionManager.java:logDebug(204)) - Connecting [FAILED toString()] to https://stevelwales.dfs.core.windows.net:443\r\n2024-09-10 17:28:25,672 [main] DEBUG services.AbfsConnectionManager (AbfsConnectionManager.java:logDebug(204)) - Connection established: stevelwales.dfs.core.windows.net:443:-1212034775\r\n2024-09-10 17:28:25,713 [main] DEBUG services.AbfsConnectionManager (AbfsConnectionManager.java:logDebug(204)) - Connection cached: .....dfs.core.windows.net\r\n{code}\r\n","created":"2024-09-10T16:54:04.543+0000"}],"conversations":[{"body":"if {{AbfsManagedApacheHttpConnection.toString()}} is invoked and httpClientConnection is null, you get a stack trace.","from":"reporter","subject":"[ABFS]: NPE in AbfsManagedApacheHttpConnection.toString() when not connected"},{"body":"{code}\r\nSLF4J: Failed toString() invocation on an object of type [org.apache.hadoop.fs.azurebfs.services.AbfsManagedApacheHttpConnection]\r\nReported exception:\r\njava.lang.NullPointerException\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsManagedApacheHttpConnection.toString(AbfsManagedApacheHttpConnection.java:233)\r\n at org.slf4j.helpers.MessageFormatter.safeObjectAppend(MessageFormatter.java:277)\r\n at org.slf4j.helpers.MessageFormatter.deeplyAppendParameter(MessageFormatter.java:249)\r\n at org.slf4j.helpers.MessageFormatter.arrayFormat(MessageFormatter.java:211)\r\n at org.slf4j.helpers.MessageFormatter.arrayFormat(MessageFormatter.java:161)\r\n at org.slf4j.impl.Reload4jLoggerAdapter.debug(Reload4jLoggerAdapter.java:251)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsConnectionManager.logDebug(AbfsConnectionManager.java:204)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsConnectionManager.connect(AbfsConnectionManager.java:155)\r\n at org.apache.http.impl.execchain.MainClientExec.establishRoute(MainClientExec.java:393)\r\n at org.apache.http.impl.execchain.MainClientExec.execute(MainClientExec.java:236)\r\n at org.apache.http.impl.execchain.ProtocolExec.execute(ProtocolExec.java:186)\r\n at org.apache.http.impl.client.InternalHttpClient.doExecute(InternalHttpClient.java:185)\r\n at org.apache.http.impl.client.CloseableHttpClient.execute(CloseableHttpClient.java:83)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsApacheHttpClient.execute(AbfsApacheHttpClient.java:122)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsAHCHttpOperation.executeRequest(AbfsAHCHttpOperation.java:266)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsAHCHttpOperation.processResponse(AbfsAHCHttpOperation.java:199)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.executeHttpOperation(AbfsRestOperation.java:392)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.completeExecute(AbfsRestOperation.java:292)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.lambda$execute$0(AbfsRestOperation.java:258)\r\n at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.measureDurationOfInvocation(IOStatisticsBinding.java:494)\r\n at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.trackDurationOfInvocation(IOStatisticsBinding.java:465)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsRestOperation.execute(AbfsRestOperation.java:256)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:1403)\r\n at org.apache.hadoop.fs.azurebfs.services.AbfsClient.getAclStatus(AbfsClient.java:1386)\r\n at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.getAclStatus(AzureBlobFileSystemStore.java:1673)\r\n at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem.getAclStatus(AzureBlobFileSystem.java:1302)\r\n at org.apache.hadoop.fs.store.diag.AbfsDiagnosticsInfo.validateFilesystem(AbfsDiagnosticsInfo.java:626)\r\n at org.apache.hadoop.fs.store.diag.StoreDiag.executeFileSystemOperations(StoreDiag.java:774)\r\n at org.apache.hadoop.fs.store.diag.StoreDiag.run(StoreDiag.java:230)\r\n at org.apache.hadoop.fs.store.diag.StoreDiag.run(StoreDiag.java:172)\r\n at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:82)\r\n at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:97)\r\n at org.apache.hadoop.fs.store.diag.StoreDiag.exec(StoreDiag.java:1189)\r\n at org.apache.hadoop.fs.store.diag.StoreDiag.main(StoreDiag.java:1198)\r\n at storediag.main(storediag.java:25)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at org.apache.hadoop.util.RunJar.run(RunJar.java:330)\r\n at org.apache.hadoop.util.RunJar.main(RunJar.java:245)\r\n2024-09-10 17:28:25,557 [main] DEBUG services.AbfsConnectionManager (AbfsConnectionManager.java:logDebug(204)) - Connecting [FAILED toString()] to https://stevelwales.dfs.core.windows.net:443\r\n2024-09-10 17:28:25,672 [main] DEBUG services.AbfsConnectionManager (AbfsConnectionManager.java:logDebug(204)) - Connection established: stevelwales.dfs.core.windows.net:443:-1212034775\r\n2024-09-10 17:28:25,713 [main] DEBUG services.AbfsConnectionManager (AbfsConnectionManager.java:logDebug(204)) - Connection cached: .....dfs.core.windows.net\r\n{code}\r\n","from":"developer"}],"created":"2024-09-10T16:52:00.000+0000","description":"if {{AbfsManagedApacheHttpConnection.toString()}} is invoked and httpClientConnection is null, you get a stack trace.","issue_id":"13591677","key":"HADOOP-19271","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-09-16T20:34:09.000+0000","role":"fixed_distractor","summary":"[ABFS]: NPE in AbfsManagedApacheHttpConnection.toString() when not connected"} {"case_id":"13592935","cluster":"DISTRACTOR-HADOOP-19285","comments":[{"body":"FYI [~anujmodi2021]","created":"2024-09-23T10:53:35.015+0000"},{"body":"just going to highlight that this does break a test and we didn't notice. FWIW my azure storage account disappeared in some housekeeping purge a few weeks back and I only just restored it -I'm surprised nobody else did though. not enough CI tests are running off trunk.\r\n\r\n{code}\r\n[ERROR] testFileSystemCapabilities(org.apache.hadoop.fs.azurebfs.ITestFileSystemInitialization) Time elapsed: 0.23 s <<< FAILURE!\r\norg.junit.ComparisonFailure: [path capability fs.capability.etags.available in AzureBlobFileSystem{uri=abfs://abfs-testcontainer-2e77c557-180b-4156-b72f-25c995c6753a@stevelwales.dfs.core.windows.net, user='stevel', primaryUserGroup='staff'[fs.azure.capability.readahead.safe]}] expected:<[tru]e> but was:<[fals]e>\r\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\r\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\r\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\r\n\tat org.apache.hadoop.fs.azurebfs.ITestFileSystemInitialization.testFileSystemCapabilities(ITestFileSystemInitialization.java:94)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n\tat java.lang.reflect.Method.invoke(Method.java:498)\r\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:59)\r\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\r\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:56)\r\n\tat org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\r\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:26)\r\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:27)\r\n\tat org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:61)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:299)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:293)\r\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n\tat java.lang.Thread.run(Thread.java:750)\r\n\r\n[INFO]\r\n[INFO] Results:\r\n[INFO]\r\n[ERROR] Failures:\r\n[ERROR] ITestFileSystemInitialization.testFileSystemCapabilities:94 [path capability fs.capability.etags.available in AzureBlobFileSystem{uri=abfs://abfs-testcontainer-2e77c557-180b-4156-b72f-25c995c6753a@stevelwales.dfs.core.windows.net, user='stevel', primaryUserGroup='staff'[fs.azure.capability.readahead.safe]}] expected:<[tru]e> but was:<[fals]e>\r\n[INFO]\r\n{code}\r\n","created":"2024-09-23T11:17:45.600+0000"},{"body":"I believe this is also failing on trunk due to same reason:\r\nITestAbfsRenameStageFailure.testDeleteTargetPaths()","created":"2024-09-23T11:24:15.940+0000"}],"conversations":[{"body":"HADOOP-19131 accidentally deleted {{CommonPathCapabilities.ETAGS_AVAILABLE}} from the patch capabilities of abfs. \r\n\r\nrestore","from":"reporter","subject":"[ABFS] Restore ETAGS_AVAILABLE to abfs path capabilities"},{"body":"FYI [~anujmodi2021]","from":"developer"},{"body":"just going to highlight that this does break a test and we didn't notice. FWIW my azure storage account disappeared in some housekeeping purge a few weeks back and I only just restored it -I'm surprised nobody else did though. not enough CI tests are running off trunk.\r\n\r\n{code}\r\n[ERROR] testFileSystemCapabilities(org.apache.hadoop.fs.azurebfs.ITestFileSystemInitialization) Time elapsed: 0.23 s <<< FAILURE!\r\norg.junit.ComparisonFailure: [path capability fs.capability.etags.available in AzureBlobFileSystem{uri=abfs://abfs-testcontainer-2e77c557-180b-4156-b72f-25c995c6753a@stevelwales.dfs.core.windows.net, user='stevel', primaryUserGroup='staff'[fs.azure.capability.readahead.safe]}] expected:<[tru]e> but was:<[fals]e>\r\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\r\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\r\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\r\n\tat org.apache.hadoop.fs.azurebfs.ITestFileSystemInitialization.testFileSystemCapabilities(ITestFileSystemInitialization.java:94)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n\tat java.lang.reflect.Method.invoke(Method.java:498)\r\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:59)\r\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\r\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:56)\r\n\tat org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\r\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:26)\r\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:27)\r\n\tat org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:61)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:299)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:293)\r\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n\tat java.lang.Thread.run(Thread.java:750)\r\n\r\n[INFO]\r\n[INFO] Results:\r\n[INFO]\r\n[ERROR] Failures:\r\n[ERROR] ITestFileSystemInitialization.testFileSystemCapabilities:94 [path capability fs.capability.etags.available in AzureBlobFileSystem{uri=abfs://abfs-testcontainer-2e77c557-180b-4156-b72f-25c995c6753a@stevelwales.dfs.core.windows.net, user='stevel', primaryUserGroup='staff'[fs.azure.capability.readahead.safe]}] expected:<[tru]e> but was:<[fals]e>\r\n[INFO]\r\n{code}\r\n","from":"developer"},{"body":"I believe this is also failing on trunk due to same reason:\r\nITestAbfsRenameStageFailure.testDeleteTargetPaths()","from":"developer"}],"created":"2024-09-23T10:52:54.000+0000","description":"HADOOP-19131 accidentally deleted {{CommonPathCapabilities.ETAGS_AVAILABLE}} from the patch capabilities of abfs. \r\n\r\nrestore","issue_id":"13592935","key":"HADOOP-19285","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-09-23T18:24:31.000+0000","role":"fixed_distractor","summary":"[ABFS] Restore ETAGS_AVAILABLE to abfs path capabilities"} {"case_id":"13593479","cluster":"DISTRACTOR-HADOOP-19290","comments":[{"body":"This code part\r\n\r\n{noformat}\r\n public Path getChecksumFile(Path file) {\r\n return new Path(file.getParent(), \".\" + file.getName() + \".crc\");\r\n{noformat}\r\n\r\nif Path is {{/}} the {{file.getParent()}} would be {{null}} and the {{Path}} constructor will lead to NPE\r\n\r\nI plan add is not root check when we try to fetch the checksum file\r\n\r\n{noformat}\r\n boolean status = apply(p);\r\n if (status && !p.isRoot()) {\r\n Path checkFile = getChecksumFile(p);\r\n{noformat}\r\n\r\nThere won't be no checkFile for root, so we can add {{ !p.isRoot()}} and avoid this code path, actually I think we can avoid it for all directories, maybe the reason for not checking whether the path is a directory or not was to avoid an additional {{getFileStatus}} call to figure out whether the path is directory or not. but the root check is no cost so should be safe from perf side as well.\r\n","created":"2024-09-27T08:31:15.670+0000"},{"body":"Committed to trunk.\r\n\r\nThanx Everyone for the reviews!!!","created":"2024-09-28T14:05:52.939+0000"},{"body":"you going to backport to 3.4.0?","created":"2024-10-01T17:01:25.012+0000"},{"body":"c-picked to 3.4","created":"2024-10-01T17:19:35.380+0000"}],"conversations":[{"body":"Operating on / on ChecksumFileSystem throws NPE\r\n\r\n{noformat}\r\njava.lang.NullPointerException\r\n\tat org.apache.hadoop.fs.Path.(Path.java:151)\r\n\tat org.apache.hadoop.fs.Path.(Path.java:130)\r\n\tat org.apache.hadoop.fs.ChecksumFileSystem.getChecksumFile(ChecksumFileSystem.java:121)\r\n\tat org.apache.hadoop.fs.ChecksumFileSystem$FsOperation.run(ChecksumFileSystem.java:774)\r\n\tat org.apache.hadoop.fs.ChecksumFileSystem.setReplication(ChecksumFileSystem.java:884)\r\n{noformat}\r\n\r\nInternally I observed it for SetPermission but on my Mac LocalFs doesn't let me setPermission on \"/\", so I reproduced it via SetReplication which goes through the same code path","from":"reporter","subject":"Operating on / in ChecksumFileSystem throws NPE"},{"body":"This code part\r\n\r\n{noformat}\r\n public Path getChecksumFile(Path file) {\r\n return new Path(file.getParent(), \".\" + file.getName() + \".crc\");\r\n{noformat}\r\n\r\nif Path is {{/}} the {{file.getParent()}} would be {{null}} and the {{Path}} constructor will lead to NPE\r\n\r\nI plan add is not root check when we try to fetch the checksum file\r\n\r\n{noformat}\r\n boolean status = apply(p);\r\n if (status && !p.isRoot()) {\r\n Path checkFile = getChecksumFile(p);\r\n{noformat}\r\n\r\nThere won't be no checkFile for root, so we can add {{ !p.isRoot()}} and avoid this code path, actually I think we can avoid it for all directories, maybe the reason for not checking whether the path is a directory or not was to avoid an additional {{getFileStatus}} call to figure out whether the path is directory or not. but the root check is no cost so should be safe from perf side as well.\r\n","from":"developer"},{"body":"Committed to trunk.\r\n\r\nThanx Everyone for the reviews!!!","from":"developer"},{"body":"you going to backport to 3.4.0?","from":"developer"},{"body":"c-picked to 3.4","from":"developer"}],"created":"2024-09-27T08:25:17.000+0000","description":"Operating on / on ChecksumFileSystem throws NPE\r\n\r\n{noformat}\r\njava.lang.NullPointerException\r\n\tat org.apache.hadoop.fs.Path.(Path.java:151)\r\n\tat org.apache.hadoop.fs.Path.(Path.java:130)\r\n\tat org.apache.hadoop.fs.ChecksumFileSystem.getChecksumFile(ChecksumFileSystem.java:121)\r\n\tat org.apache.hadoop.fs.ChecksumFileSystem$FsOperation.run(ChecksumFileSystem.java:774)\r\n\tat org.apache.hadoop.fs.ChecksumFileSystem.setReplication(ChecksumFileSystem.java:884)\r\n{noformat}\r\n\r\nInternally I observed it for SetPermission but on my Mac LocalFs doesn't let me setPermission on \"/\", so I reproduced it via SetReplication which goes through the same code path","issue_id":"13593479","key":"HADOOP-19290","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-09-28T14:06:03.000+0000","role":"fixed_distractor","summary":"Operating on / in ChecksumFileSystem throws NPE"} {"case_id":"13594090","cluster":"DISTRACTOR-HADOOP-19299","comments":[{"body":"{code}\r\n.s3a.ITestS3AContractVectoredRead\r\n[ERROR] testSomeRangesMergedSomeUnmerged[Buffer type : array](org.apache.hadoop.fs.contract.s3a.ITestS3AContractVectoredRead) Time elapsed: 0.905 s <<< ERROR!\r\njava.util.ConcurrentModificationException\r\n at java.util.HashMap$EntrySpliterator.forEachRemaining(HashMap.java:1728)\r\n at java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:482)\r\n at java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:472)\r\n at java.util.stream.ReduceOps$ReduceOp.evaluateSequential(ReduceOps.java:708)\r\n at java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:234)\r\n at java.util.stream.ReferencePipeline.collect(ReferencePipeline.java:566)\r\n at org.apache.hadoop.fs.store.audit.HttpReferrerAuditHeader.buildHttpReferrer(HttpReferrerAuditHeader.java:182)\r\n at org.apache.hadoop.fs.s3a.audit.impl.LoggingAuditor$LoggingAuditSpan.modifyHttpRequest(LoggingAuditor.java:388)\r\n at org.apache.hadoop.fs.s3a.audit.impl.ActiveAuditManagerS3A$WrappingAuditSpan.modifyHttpRequest(ActiveAuditManagerS3A.java:871)\r\n at org.apache.hadoop.fs.s3a.audit.impl.ActiveAuditManagerS3A.modifyHttpRequest(ActiveAuditManagerS3A.java:612)\r\n at software.amazon.awssdk.core.interceptor.ExecutionInterceptorChain.modifyHttpRequestAndHttpContent(ExecutionInterceptorChain.java:89)\r\n at software.amazon.awssdk.core.internal.handler.BaseClientHandler.runModifyHttpRequestAndHttpContentInterceptors(BaseClientHandler.java:157)\r\n at software.amazon.awssdk.core.internal.handler.BaseClientHandler.finalizeSdkHttpFullRequest(BaseClientHandler.java:83)\r\n at software.amazon.awssdk.core.internal.handler.BaseSyncClientHandler.doExecute(BaseSyncClientHandler.java:151)\r\n at software.amazon.awssdk.core.internal.handler.BaseSyncClientHandler.lambda$execute$0(BaseSyncClientHandler.java:66)\r\n at software.amazon.awssdk.core.internal.handler.BaseSyncClientHandler.measureApiCallSuccess(BaseSyncClientHandler.java:182)\r\n at software.amazon.awssdk.core.internal.handler.BaseSyncClientHandler.execute(BaseSyncClientHandler.java:60)\r\n at software.amazon.awssdk.core.client.handler.SdkSyncClientHandler.execute(SdkSyncClientHandler.java:52)\r\n at software.amazon.awssdk.awscore.client.handler.AwsSyncClientHandler.execute(AwsSyncClientHandler.java:60)\r\n at software.amazon.awssdk.services.s3.DefaultS3Client.getObject(DefaultS3Client.java:5174)\r\n at software.amazon.awssdk.services.s3.S3Client.getObject(S3Client.java:9005)\r\n at org.apache.hadoop.fs.s3a.S3AFileSystem$InputStreamCallbacksImpl.getObject(S3AFileSystem.java:1934)\r\n at org.apache.hadoop.fs.s3a.S3AInputStream.lambda$getS3Object$7(S3AInputStream.java:1223)\r\n at org.apache.hadoop.fs.s3a.Invoker.once(Invoker.java:122)\r\n at org.apache.hadoop.fs.s3a.Invoker.lambda$retry$4(Invoker.java:376)\r\n at org.apache.hadoop.fs.s3a.Invoker.retryUntranslated(Invoker.java:468)\r\n at org.apache.hadoop.fs.s3a.Invoker.retry(Invoker.java:372)\r\n at org.apache.hadoop.fs.s3a.Invoker.retry(Invoker.java:347)\r\n at org.apache.hadoop.fs.s3a.S3AInputStream.getS3Object(S3AInputStream.java:1220)\r\n at org.apache.hadoop.fs.s3a.S3AInputStream.getS3ObjectInputStream(S3AInputStream.java:1117)\r\n at org.apache.hadoop.fs.s3a.S3AInputStream.readCombinedRangeAndUpdateChildren(S3AInputStream.java:963)\r\n at org.apache.hadoop.fs.s3a.S3AInputStream.lambda$readVectored$5(S3AInputStream.java:945)\r\n at org.apache.hadoop.util.SemaphoredDelegatingExecutor$RunnableWithPermitRelease.run(SemaphoredDelegatingExecutor.java:225)\r\n at org.apache.hadoop.util.SemaphoredDelegatingExecutor$RunnableWithPermitRelease.run(SemaphoredDelegatingExecutor.java:225)\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\r\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n at java.lang.Thread.run(Thread.java:750)\r\n{code}\r\n","created":"2024-10-02T17:08:38.910+0000"},{"body":"Hi [~stevel@apache.org] , can you give some repro steps for this and maybe I can pick it up? ","created":"2024-10-03T08:16:45.743+0000"},{"body":"no need, pr is up","created":"2024-10-07T12:49:11.667+0000"}],"conversations":[{"body":"Surfaced during a test run doing vector iO, where multiple parallel GETs were being issued within the same audit span, just when the header is built by enumerating the attributes.\r\n\r\n{code}\r\n queries = attributes.entrySet().stream()\r\n .filter(e -> !filter.contains(e.getKey()))\r\n .map(e -> e.getKey() + \"=\" + e.getValue())\r\n .collect(Collectors.joining(\"&\"));\r\n{code}\r\n\r\nHypothesis: multiple GET requests are conflicting in updating/reading the header.","from":"reporter","subject":"ConcurrentModificationException in HttpReferrerAuditHeader"},{"body":"{code}\r\n.s3a.ITestS3AContractVectoredRead\r\n[ERROR] testSomeRangesMergedSomeUnmerged[Buffer type : array](org.apache.hadoop.fs.contract.s3a.ITestS3AContractVectoredRead) Time elapsed: 0.905 s <<< ERROR!\r\njava.util.ConcurrentModificationException\r\n at java.util.HashMap$EntrySpliterator.forEachRemaining(HashMap.java:1728)\r\n at java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:482)\r\n at java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:472)\r\n at java.util.stream.ReduceOps$ReduceOp.evaluateSequential(ReduceOps.java:708)\r\n at java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:234)\r\n at java.util.stream.ReferencePipeline.collect(ReferencePipeline.java:566)\r\n at org.apache.hadoop.fs.store.audit.HttpReferrerAuditHeader.buildHttpReferrer(HttpReferrerAuditHeader.java:182)\r\n at org.apache.hadoop.fs.s3a.audit.impl.LoggingAuditor$LoggingAuditSpan.modifyHttpRequest(LoggingAuditor.java:388)\r\n at org.apache.hadoop.fs.s3a.audit.impl.ActiveAuditManagerS3A$WrappingAuditSpan.modifyHttpRequest(ActiveAuditManagerS3A.java:871)\r\n at org.apache.hadoop.fs.s3a.audit.impl.ActiveAuditManagerS3A.modifyHttpRequest(ActiveAuditManagerS3A.java:612)\r\n at software.amazon.awssdk.core.interceptor.ExecutionInterceptorChain.modifyHttpRequestAndHttpContent(ExecutionInterceptorChain.java:89)\r\n at software.amazon.awssdk.core.internal.handler.BaseClientHandler.runModifyHttpRequestAndHttpContentInterceptors(BaseClientHandler.java:157)\r\n at software.amazon.awssdk.core.internal.handler.BaseClientHandler.finalizeSdkHttpFullRequest(BaseClientHandler.java:83)\r\n at software.amazon.awssdk.core.internal.handler.BaseSyncClientHandler.doExecute(BaseSyncClientHandler.java:151)\r\n at software.amazon.awssdk.core.internal.handler.BaseSyncClientHandler.lambda$execute$0(BaseSyncClientHandler.java:66)\r\n at software.amazon.awssdk.core.internal.handler.BaseSyncClientHandler.measureApiCallSuccess(BaseSyncClientHandler.java:182)\r\n at software.amazon.awssdk.core.internal.handler.BaseSyncClientHandler.execute(BaseSyncClientHandler.java:60)\r\n at software.amazon.awssdk.core.client.handler.SdkSyncClientHandler.execute(SdkSyncClientHandler.java:52)\r\n at software.amazon.awssdk.awscore.client.handler.AwsSyncClientHandler.execute(AwsSyncClientHandler.java:60)\r\n at software.amazon.awssdk.services.s3.DefaultS3Client.getObject(DefaultS3Client.java:5174)\r\n at software.amazon.awssdk.services.s3.S3Client.getObject(S3Client.java:9005)\r\n at org.apache.hadoop.fs.s3a.S3AFileSystem$InputStreamCallbacksImpl.getObject(S3AFileSystem.java:1934)\r\n at org.apache.hadoop.fs.s3a.S3AInputStream.lambda$getS3Object$7(S3AInputStream.java:1223)\r\n at org.apache.hadoop.fs.s3a.Invoker.once(Invoker.java:122)\r\n at org.apache.hadoop.fs.s3a.Invoker.lambda$retry$4(Invoker.java:376)\r\n at org.apache.hadoop.fs.s3a.Invoker.retryUntranslated(Invoker.java:468)\r\n at org.apache.hadoop.fs.s3a.Invoker.retry(Invoker.java:372)\r\n at org.apache.hadoop.fs.s3a.Invoker.retry(Invoker.java:347)\r\n at org.apache.hadoop.fs.s3a.S3AInputStream.getS3Object(S3AInputStream.java:1220)\r\n at org.apache.hadoop.fs.s3a.S3AInputStream.getS3ObjectInputStream(S3AInputStream.java:1117)\r\n at org.apache.hadoop.fs.s3a.S3AInputStream.readCombinedRangeAndUpdateChildren(S3AInputStream.java:963)\r\n at org.apache.hadoop.fs.s3a.S3AInputStream.lambda$readVectored$5(S3AInputStream.java:945)\r\n at org.apache.hadoop.util.SemaphoredDelegatingExecutor$RunnableWithPermitRelease.run(SemaphoredDelegatingExecutor.java:225)\r\n at org.apache.hadoop.util.SemaphoredDelegatingExecutor$RunnableWithPermitRelease.run(SemaphoredDelegatingExecutor.java:225)\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\r\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n at java.lang.Thread.run(Thread.java:750)\r\n{code}\r\n","from":"developer"},{"body":"Hi [~stevel@apache.org] , can you give some repro steps for this and maybe I can pick it up? ","from":"developer"},{"body":"no need, pr is up","from":"developer"}],"created":"2024-10-02T17:08:20.000+0000","description":"Surfaced during a test run doing vector iO, where multiple parallel GETs were being issued within the same audit span, just when the header is built by enumerating the attributes.\r\n\r\n{code}\r\n queries = attributes.entrySet().stream()\r\n .filter(e -> !filter.contains(e.getKey()))\r\n .map(e -> e.getKey() + \"=\" + e.getValue())\r\n .collect(Collectors.joining(\"&\"));\r\n{code}\r\n\r\nHypothesis: multiple GET requests are conflicting in updating/reading the header.","issue_id":"13594090","key":"HADOOP-19299","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-10-07T16:47:14.000+0000","role":"fixed_distractor","summary":"ConcurrentModificationException in HttpReferrerAuditHeader"} {"case_id":"13595077","cluster":"DISTRACTOR-HADOOP-19309","comments":[{"body":"FWIW we've turned off the optimised option, due to some auth problem. transfer manager is hard to debug","created":"2024-10-11T16:31:00.224+0000"},{"body":"[~stevel@apache.org]  - I still see the config is enabled by default : https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/Constants.java#L1617","created":"2024-10-11T17:52:12.275+0000"},{"body":"[~stevel@apache.org]  - Additionally i see that option to enable/disable fast s3a copyFromLocal was added as part of [https://github.com/apache/hadoop/commit/ef7fb64764a9b9c77b0dbda48285f186cfcf5c4f] , I could still see the default is set to true.","created":"2024-10-14T05:35:41.288+0000"},{"body":"we turned it off in cloudera deployments as it was causing problems; v2sdk async signing related and not expected to be seen externally","created":"2024-10-15T10:58:27.683+0000"},{"body":"[~stevel@apache.org]  Sure, I had verified the hadoop-3.4.1 RC-3 for ARM64 based machines and have send my response over the email thread.","created":"2024-10-15T11:03:45.388+0000"}],"conversations":[{"body":"When the sourcePath does not contain any file scheme information, S3A CopyFromLocalFile operation fails with the following exception stack trace.\r\n{code:java}\r\n    at org.apache.hadoop.fs.s3a.impl.CopyFromLocalOperation.getFinalPath(CopyFromLocalOperation.java:360)\r\n    at org.apache.hadoop.fs.s3a.impl.CopyFromLocalOperation.uploadSourceFromFS(CopyFromLocalOperation.java:222)\r\n    at org.apache.hadoop.fs.s3a.impl.CopyFromLocalOperation.execute(CopyFromLocalOperation.java:169)\r\n    at org.apache.hadoop.fs.s3a.S3AFileSystem.lambda$copyFromLocalFile$23(S3AFileSystem.java:4217)\r\n    at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.invokeTrackingDuration(IOStatisticsBinding.java:547)\r\n    at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.lambda$trackDurationOfOperation$5(IOStatisticsBinding.java:528)\r\n    at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.trackDuration(IOStatisticsBinding.java:449)\r\n    at org.apache.hadoop.fs.s3a.S3AFileSystem.trackDurationAndSpan(S3AFileSystem.java:2871)\r\n    at org.apache.hadoop.fs.s3a.S3AFileSystem.trackDurationAndSpan(S3AFileSystem.java:2890)\r\n    at org.apache.hadoop.fs.s3a.S3AFileSystem.copyFromLocalFile(S3AFileSystem.java:4209)\r\n {code}\r\nAdditionally the failure is seen only when \r\n{color:#172b4d}*fs.s3a.optimized.copy.from.local.enabled* is enabled (which is by default). This happens only when the local source file is given without any file scheme for example : /tmp/file.txt instead of file:///tmp/file.txt.\r\n{color}\r\n \r\n{color:#172b4d}The proposal here is to add file scheme to the source if the source path does not contain the same.{color}","from":"reporter","subject":"S3A CopyFromLocalFile operation fails when the source file does not contain file scheme."},{"body":"FWIW we've turned off the optimised option, due to some auth problem. transfer manager is hard to debug","from":"developer"},{"body":"[~stevel@apache.org]  - I still see the config is enabled by default : https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/Constants.java#L1617","from":"developer"},{"body":"[~stevel@apache.org]  - Additionally i see that option to enable/disable fast s3a copyFromLocal was added as part of [https://github.com/apache/hadoop/commit/ef7fb64764a9b9c77b0dbda48285f186cfcf5c4f] , I could still see the default is set to true.","from":"developer"},{"body":"we turned it off in cloudera deployments as it was causing problems; v2sdk async signing related and not expected to be seen externally","from":"developer"},{"body":"[~stevel@apache.org]  Sure, I had verified the hadoop-3.4.1 RC-3 for ARM64 based machines and have send my response over the email thread.","from":"developer"}],"created":"2024-10-11T12:09:53.000+0000","description":"When the sourcePath does not contain any file scheme information, S3A CopyFromLocalFile operation fails with the following exception stack trace.\r\n{code:java}\r\n    at org.apache.hadoop.fs.s3a.impl.CopyFromLocalOperation.getFinalPath(CopyFromLocalOperation.java:360)\r\n    at org.apache.hadoop.fs.s3a.impl.CopyFromLocalOperation.uploadSourceFromFS(CopyFromLocalOperation.java:222)\r\n    at org.apache.hadoop.fs.s3a.impl.CopyFromLocalOperation.execute(CopyFromLocalOperation.java:169)\r\n    at org.apache.hadoop.fs.s3a.S3AFileSystem.lambda$copyFromLocalFile$23(S3AFileSystem.java:4217)\r\n    at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.invokeTrackingDuration(IOStatisticsBinding.java:547)\r\n    at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.lambda$trackDurationOfOperation$5(IOStatisticsBinding.java:528)\r\n    at org.apache.hadoop.fs.statistics.impl.IOStatisticsBinding.trackDuration(IOStatisticsBinding.java:449)\r\n    at org.apache.hadoop.fs.s3a.S3AFileSystem.trackDurationAndSpan(S3AFileSystem.java:2871)\r\n    at org.apache.hadoop.fs.s3a.S3AFileSystem.trackDurationAndSpan(S3AFileSystem.java:2890)\r\n    at org.apache.hadoop.fs.s3a.S3AFileSystem.copyFromLocalFile(S3AFileSystem.java:4209)\r\n {code}\r\nAdditionally the failure is seen only when \r\n{color:#172b4d}*fs.s3a.optimized.copy.from.local.enabled* is enabled (which is by default). This happens only when the local source file is given without any file scheme for example : /tmp/file.txt instead of file:///tmp/file.txt.\r\n{color}\r\n \r\n{color:#172b4d}The proposal here is to add file scheme to the source if the source path does not contain the same.{color}","issue_id":"13595077","key":"HADOOP-19309","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-10-25T10:12:24.000+0000","role":"fixed_distractor","summary":"S3A CopyFromLocalFile operation fails when the source file does not contain file scheme."} {"case_id":"13599167","cluster":"DISTRACTOR-HADOOP-19342","comments":[{"body":"The pull request is merged.","created":"2024-11-21T16:15:49.981+0000"}],"conversations":[{"body":"After HADOOP-19306, the following INFO messages are printed in client side. Thanks As [~stevel@apache.org] for reporting it in [this comment|https://github.com/apache/hadoop/pull/7140#issuecomment-2483881941].\r\n{code}\r\n2024-11-18 18:53:37,645 [setup] INFO security.SaslRpcServer (SaslRpcServer.java:(239)) - AuthMethod SIMPLE: code=80, mechanism=\"\"\r\n2024-11-18 18:53:37,645 [setup] INFO security.SaslRpcServer (SaslRpcServer.java:(239)) - AuthMethod KERBEROS: code=81, mechanism=\"GSSAPI\"\r\n2024-11-18 18:53:37,646 [setup] INFO security.SaslRpcServer (SaslRpcServer.java:(239)) - AuthMethod DIGEST: code=82, mechanism=\"DIGEST-MD5\"\r\n2024-11-18 18:53:37,646 [setup] INFO security.SaslRpcServer (SaslRpcServer.java:(239)) - AuthMethod TOKEN: code=82, mechanism=\"DIGEST-MD5\"\r\n2024-11-18 18:53:37,646 [setup] INFO security.SaslRpcServer (SaslRpcServer.java:(239)) - AuthMethod PLAIN: code=83, mechanism=\"PLAIN\"\r\n{code}","from":"reporter","subject":"SaslRpcServer.AuthMethod print INFO messages in client side"},{"body":"The pull request is merged.","from":"developer"}],"created":"2024-11-18T19:18:08.000+0000","description":"After HADOOP-19306, the following INFO messages are printed in client side. Thanks As [~stevel@apache.org] for reporting it in [this comment|https://github.com/apache/hadoop/pull/7140#issuecomment-2483881941].\r\n{code}\r\n2024-11-18 18:53:37,645 [setup] INFO security.SaslRpcServer (SaslRpcServer.java:(239)) - AuthMethod SIMPLE: code=80, mechanism=\"\"\r\n2024-11-18 18:53:37,645 [setup] INFO security.SaslRpcServer (SaslRpcServer.java:(239)) - AuthMethod KERBEROS: code=81, mechanism=\"GSSAPI\"\r\n2024-11-18 18:53:37,646 [setup] INFO security.SaslRpcServer (SaslRpcServer.java:(239)) - AuthMethod DIGEST: code=82, mechanism=\"DIGEST-MD5\"\r\n2024-11-18 18:53:37,646 [setup] INFO security.SaslRpcServer (SaslRpcServer.java:(239)) - AuthMethod TOKEN: code=82, mechanism=\"DIGEST-MD5\"\r\n2024-11-18 18:53:37,646 [setup] INFO security.SaslRpcServer (SaslRpcServer.java:(239)) - AuthMethod PLAIN: code=83, mechanism=\"PLAIN\"\r\n{code}","issue_id":"13599167","key":"HADOOP-19342","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2024-11-21T16:15:49.000+0000","role":"fixed_distractor","summary":"SaslRpcServer.AuthMethod print INFO messages in client side"} {"case_id":"13603599","cluster":"DISTRACTOR-HADOOP-19382","comments":[{"body":"{code}\r\n[ERROR] testFileSystemInitFailsWithBlobEndpoitUrl(org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemInitAndCreate) Time elapsed: 0.015 s <<< ERROR!\r\nFailure to initialize configuration for NAME.blob.core.windows.net key =\"null\": Invalid configuration value detected for fs.azure.account.key\r\n at org.apache.hadoop.fs.azurebfs.services.SimpleKeyProvider.getStorageAccountKey(SimpleKeyProvider.java:53)\r\n at org.apache.hadoop.fs.azurebfs.AbfsConfiguration.getStorageAccountKey(AbfsConfiguration.java:821)\r\n at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.initializeClient(AzureBlobFileSystemStore.java:1781)\r\n at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.(AzureBlobFileSystemStore.java:267)\r\n at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem.initialize(AzureBlobFileSystem.java:209)\r\n at org.apache.hadoop.fs.FileSystem.createFileSystem(FileSystem.java:3615)\r\n at org.apache.hadoop.fs.FileSystem.access$300(FileSystem.java:172)\r\n at org.apache.hadoop.fs.FileSystem$Cache.getInternal(FileSystem.java:3716)\r\n at org.apache.hadoop.fs.FileSystem$Cache.getUnique(FileSystem.java:3673)\r\n at org.apache.hadoop.fs.FileSystem.newInstance(FileSystem.java:610)\r\n at org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemInitAndCreate.lambda$testFileSystemInitFailsWithBlobEndpoitUrl$0(ITestAzureBlobFileSystemInitAndCreate.java:124)\r\n at org.apache.hadoop.test.LambdaTestUtils.intercept(LambdaTestUtils.java:500)\r\n at org.apache.hadoop.test.LambdaTestUtils.intercept(LambdaTestUtils.java:386)\r\n at org.apache.hadoop.test.LambdaTestUtils.intercept(LambdaTestUtils.java:455)\r\n at org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemInitAndCreate.testFileSystemInitFailsWithBlobEndpoitUrl(ITestAzureBlobFileSystemInitAndCreate.java:122)\r\n {code}\r\n","created":"2025-01-02T18:14:27.577+0000"},{"body":"This is a recent test from HADOOP-19187.\r\n\r\n[~anujmodi] this is one of yours, I'm afraid","created":"2025-01-02T18:15:58.109+0000"},{"body":"Thanks for reporting this [~stevel@apache.org] \r\nWill look into this on priority.","created":"2025-01-03T12:56:49.191+0000"}],"conversations":[{"body":"test failure in {{ITestAzureBlobFileSystemInitAndCreate.testFileSystemInitFailsWithBlobEndpoitUrl}} failing because the test store has a .dfs credential but not a blob one\r\n{code}\r\n\r\n[ERROR] testFileSystemInitFailsWithBlobEndpoitUrl(org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemInitAndCreate) Time elapsed: 0.015 s <<< ERROR!\r\nFailure to initialize configuration for NAME.blob.core.windows.net key =\"null\": Invalid configuration value detected for fs.azure.account.key\r\n{code}\r\n\r\n\r\nThis should downgrade to a skip.\r\n","from":"reporter","subject":"[ABFS] ITestAzureBlobFileSystemInitAndCreate failure"},{"body":"{code}\r\n[ERROR] testFileSystemInitFailsWithBlobEndpoitUrl(org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemInitAndCreate) Time elapsed: 0.015 s <<< ERROR!\r\nFailure to initialize configuration for NAME.blob.core.windows.net key =\"null\": Invalid configuration value detected for fs.azure.account.key\r\n at org.apache.hadoop.fs.azurebfs.services.SimpleKeyProvider.getStorageAccountKey(SimpleKeyProvider.java:53)\r\n at org.apache.hadoop.fs.azurebfs.AbfsConfiguration.getStorageAccountKey(AbfsConfiguration.java:821)\r\n at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.initializeClient(AzureBlobFileSystemStore.java:1781)\r\n at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.(AzureBlobFileSystemStore.java:267)\r\n at org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem.initialize(AzureBlobFileSystem.java:209)\r\n at org.apache.hadoop.fs.FileSystem.createFileSystem(FileSystem.java:3615)\r\n at org.apache.hadoop.fs.FileSystem.access$300(FileSystem.java:172)\r\n at org.apache.hadoop.fs.FileSystem$Cache.getInternal(FileSystem.java:3716)\r\n at org.apache.hadoop.fs.FileSystem$Cache.getUnique(FileSystem.java:3673)\r\n at org.apache.hadoop.fs.FileSystem.newInstance(FileSystem.java:610)\r\n at org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemInitAndCreate.lambda$testFileSystemInitFailsWithBlobEndpoitUrl$0(ITestAzureBlobFileSystemInitAndCreate.java:124)\r\n at org.apache.hadoop.test.LambdaTestUtils.intercept(LambdaTestUtils.java:500)\r\n at org.apache.hadoop.test.LambdaTestUtils.intercept(LambdaTestUtils.java:386)\r\n at org.apache.hadoop.test.LambdaTestUtils.intercept(LambdaTestUtils.java:455)\r\n at org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemInitAndCreate.testFileSystemInitFailsWithBlobEndpoitUrl(ITestAzureBlobFileSystemInitAndCreate.java:122)\r\n {code}\r\n","from":"developer"},{"body":"This is a recent test from HADOOP-19187.\r\n\r\n[~anujmodi] this is one of yours, I'm afraid","from":"developer"},{"body":"Thanks for reporting this [~stevel@apache.org] \r\nWill look into this on priority.","from":"developer"}],"created":"2025-01-02T18:13:54.000+0000","description":"test failure in {{ITestAzureBlobFileSystemInitAndCreate.testFileSystemInitFailsWithBlobEndpoitUrl}} failing because the test store has a .dfs credential but not a blob one\r\n{code}\r\n\r\n[ERROR] testFileSystemInitFailsWithBlobEndpoitUrl(org.apache.hadoop.fs.azurebfs.ITestAzureBlobFileSystemInitAndCreate) Time elapsed: 0.015 s <<< ERROR!\r\nFailure to initialize configuration for NAME.blob.core.windows.net key =\"null\": Invalid configuration value detected for fs.azure.account.key\r\n{code}\r\n\r\n\r\nThis should downgrade to a skip.\r\n","issue_id":"13603599","key":"HADOOP-19382","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-01-10T06:56:06.000+0000","role":"fixed_distractor","summary":"[ABFS] ITestAzureBlobFileSystemInitAndCreate failure"} {"case_id":"13605723","cluster":"DISTRACTOR-HADOOP-19392","comments":[{"body":"[~kunda] Welcome to Hadoop! Added as a contributor to hadoop-common.","created":"2025-01-22T22:29:12.978+0000"}],"conversations":[{"body":"Running {{mvn clean install -DskipTests}} within the dev docker container fails with the following error:\r\n{code:java}\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.10.1:testCompile (default-testCompile) on project hadoop-common: Compilation failure: Compilation failure:\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[28,35] package org.apache.ftpserver.ftplet does not exist\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[29,35] package org.apache.ftpserver.ftplet does not exist\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[30,35] package org.apache.ftpserver.ftplet does not exist\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[45,11] cannot find symbol\r\n[ERROR]   symbol:   class UserManager\r\n[ERROR]   location: class org.apache.hadoop.fs.ftp.FtpTestServer\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[80,7] cannot find symbol\r\n[ERROR]   symbol:   class Authority\r\n[ERROR]   location: class org.apache.hadoop.fs.ftp.FtpTestServer\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[80,53] cannot find symbol\r\n[ERROR]   symbol:   class FtpException\r\n[ERROR]   location: class org.apache.hadoop.fs.ftp.FtpTestServer\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/TestFTPFileSystem.java:[91,22] cannot access org.apache.ftpserver.ftplet.User\r\n[ERROR]   class file for org.apache.ftpserver.ftplet.User not found {code}\r\nThe missing classes belong to the {{org.apache.ftpserver:ftplet-api}} dependency.\r\n\r\nThe {{hadoop-common}} module's {{pom.xml}} file declares a test dependency on \r\norg.apache.ftpserver:ftpserver-core, which transitively uses {{{}ftplet-api{}}}.\r\nRunning maven with {{-X}} sheds some light:\r\n{code:java}\r\n[WARNING] The POM for org.apache.ftpserver:ftpserver-core:jar:1.0.0 is invalid, transitive dependencies (if any) will not be available: 12 problems were encountered while building the effective model for org.apache.ftpserver:ftpserver-core:1.0.0\r\n[ERROR] 'dependencies.dependency.version' for org.apache.ftpserver:ftplet-api:jar is missing. @ \r\n...{code}\r\nMy thoughts are that the specific maven version used in the dev env (3.6.3) has issues with the pom file of the specific ftpserver-core version used.\r\nUpdating the `org.apache.ftpserver` artifact versions from 1.0.0 to 1.2.0 (the latest) resolves the issue.\r\n ","from":"reporter","subject":"class org.apache.hadoop.fs.ftp.FtpTestServer does not compile"},{"body":"[~kunda] Welcome to Hadoop! Added as a contributor to hadoop-common.","from":"developer"}],"created":"2025-01-21T11:33:25.000+0000","description":"Running {{mvn clean install -DskipTests}} within the dev docker container fails with the following error:\r\n{code:java}\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.10.1:testCompile (default-testCompile) on project hadoop-common: Compilation failure: Compilation failure:\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[28,35] package org.apache.ftpserver.ftplet does not exist\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[29,35] package org.apache.ftpserver.ftplet does not exist\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[30,35] package org.apache.ftpserver.ftplet does not exist\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[45,11] cannot find symbol\r\n[ERROR]   symbol:   class UserManager\r\n[ERROR]   location: class org.apache.hadoop.fs.ftp.FtpTestServer\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[80,7] cannot find symbol\r\n[ERROR]   symbol:   class Authority\r\n[ERROR]   location: class org.apache.hadoop.fs.ftp.FtpTestServer\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/FtpTestServer.java:[80,53] cannot find symbol\r\n[ERROR]   symbol:   class FtpException\r\n[ERROR]   location: class org.apache.hadoop.fs.ftp.FtpTestServer\r\n[ERROR] /home/ykunda/hadoop/hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/fs/ftp/TestFTPFileSystem.java:[91,22] cannot access org.apache.ftpserver.ftplet.User\r\n[ERROR]   class file for org.apache.ftpserver.ftplet.User not found {code}\r\nThe missing classes belong to the {{org.apache.ftpserver:ftplet-api}} dependency.\r\n\r\nThe {{hadoop-common}} module's {{pom.xml}} file declares a test dependency on \r\norg.apache.ftpserver:ftpserver-core, which transitively uses {{{}ftplet-api{}}}.\r\nRunning maven with {{-X}} sheds some light:\r\n{code:java}\r\n[WARNING] The POM for org.apache.ftpserver:ftpserver-core:jar:1.0.0 is invalid, transitive dependencies (if any) will not be available: 12 problems were encountered while building the effective model for org.apache.ftpserver:ftpserver-core:1.0.0\r\n[ERROR] 'dependencies.dependency.version' for org.apache.ftpserver:ftplet-api:jar is missing. @ \r\n...{code}\r\nMy thoughts are that the specific maven version used in the dev env (3.6.3) has issues with the pom file of the specific ftpserver-core version used.\r\nUpdating the `org.apache.ftpserver` artifact versions from 1.0.0 to 1.2.0 (the latest) resolves the issue.\r\n ","issue_id":"13605723","key":"HADOOP-19392","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-01-22T22:27:56.000+0000","role":"fixed_distractor","summary":"class org.apache.hadoop.fs.ftp.FtpTestServer does not compile"} {"case_id":"13606497","cluster":"DISTRACTOR-HADOOP-19405","comments":[{"body":"[~stevel@apache.org] I will submit a PR to upgrade {{hadoop-tools/hadoop-aws}} from JUnit 4 to JUnit 5. Hopefully, this will resolve the issue.","created":"2025-01-28T14:50:49.137+0000"},{"body":"I'm also seeing the same issue, currently unable to run any tests. ","created":"2025-01-28T14:59:58.837+0000"},{"body":"[~ahmar] How did we verify that the unit tests are not running?","created":"2025-01-28T15:16:15.072+0000"},{"body":"run `mvn clean test` to just run unit tests\r\n\r\n \r\n\r\nrun `mvn -Dparallel-tests -DtestsThreadCount=8 clean verify` to run both ITests and unit tests.\r\n\r\n \r\n\r\nBoth commands currently running 0 tests, with the output Steve shared in the ticket.","created":"2025-01-28T15:27:34.725+0000"},{"body":"azure and distcp seem to be the same. Looking online it imples that we need that vintage junit thing on the test classpath. \r\n\r\noh, and \r\n{code}\r\nmvn test\r\n{code}\r\n\r\n is sufficient. It just isn't finding any tests","created":"2025-01-28T16:23:15.005+0000"},{"body":"[~ahmar] my s3a stream PR is rebased onto the commit before it. \r\n\r\nFWIW I'm reasonably relieved it want' just me hitting it...I've been doing java17 stuff and assumed that I'd actually got something wrong by mixing things. I've managed to get IntelliJ into that state & it doesn't let me debug unit tests any more. So having maven fail seemed related","created":"2025-01-28T16:34:49.183+0000"},{"body":"[~stevel@apache.org] \r\n\r\nI tested that if we add the following dependency, the unit tests can run normally, although the unit tests throw an error locally:\r\n{code:java}\r\n\r\n  org.junit.vintage\r\n  junit-vintage-engine\r\n  ${junit.vintage.version}\r\n {code}\r\nThe maven-surefire-plugin.version must be upgraded to 3.0.0-M4.\r\n{code:java}\r\n[ERROR] ITestS3AContractBulkDelete.validatePageSize  Time elapsed: 0.002 s  <<< ERROR!\r\njava.lang.NullPointerException: No test bucket\r\n        at org.apache.hadoop.util.Preconditions.checkNotNull(Preconditions.java:88)\r\n        at org.apache.hadoop.fs.s3a.S3ATestUtils.getTestBucketName(S3ATestUtils.java:859)\r\n        at org.apache.hadoop.fs.contract.s3a.ITestS3AContractBulkDelete.createConfiguration(ITestS3AContractBulkDelete.java:88)\r\n        at org.apache.hadoop.fs.contract.AbstractFSContractTestBase.setup(AbstractFSContractTestBase.java:188)\r\n        at org.apache.hadoop.fs.contract.AbstractContractBulkDeleteTest.setup(AbstractContractBulkDeleteTest.java:79)\r\n        at sun.reflect.GeneratedMethodAccessor1.invoke(Unknown Source)\r\n\r\n[ERROR] Tests run: 22, Failures: 0, Errors: 22, Skipped: 0 {code}\r\n ","created":"2025-01-28T16:52:42.781+0000"},{"body":"Thanks for raising thi [~stevel@apache.org] \r\nWe are also facing same issue for azure tests.","created":"2025-01-29T05:22:32.480+0000"},{"body":"[~anujmodi]  I am preparing a new PR. In a module that has not been upgraded to JUnit 5, I am adding a JUnit 4 {{junit-vintage-engine}} dependency, which can help run unit tests.\r\n\r\nTest report can be viewed at the following link:\r\n\r\nhttps://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-7335/2/testReport/","created":"2025-01-29T05:44:49.653+0000"},{"body":"[~anujmodi] rebase your PRs onto the trunk commit before the jersey update and all works, you can then rebase onto trunk once it is in.\r\n\r\nThat jersey2 update is a big piece of work, and critical for java17 -and this is just test execution problems. A new PR to restore test execution is what I'd prefer -and is what is in progress","created":"2025-01-29T11:18:39.295+0000"},{"body":"[~slfan1989] all tests with an ITest prefix are integration tests -they need to be set up with credentials in src/test/resource\r\n\r\nThat stack trace is to be expected","created":"2025-01-29T11:20:36.437+0000"}],"conversations":[{"body":"Hadoop-aws tests no longer run.\r\n\r\n{code}\r\n[INFO] -------------------------------------------------------\r\n[INFO] T E S T S\r\n[INFO] -------------------------------------------------------\r\n[INFO]\r\n[INFO] Results:\r\n[INFO]\r\n[INFO] Tests run: 0, Failures: 0, Errors: 0, Skipped: 0\r\n[INFO]\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD SUCCESS\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 2.967 s (Wall Clock)\r\n[INFO] Finished at: 2025-01-28T13:59:52Z\r\n[INFO] ------------------------------------------------------------------------\r\n{code}\r\n\r\nIf I check out to the commit before the jersey v2 update (f38d7072) all is good again.\r\n\r\n","from":"reporter","subject":"hadoop-aws and hadoop-azure tests have stopped running"},{"body":"[~stevel@apache.org] I will submit a PR to upgrade {{hadoop-tools/hadoop-aws}} from JUnit 4 to JUnit 5. Hopefully, this will resolve the issue.","from":"developer"},{"body":"I'm also seeing the same issue, currently unable to run any tests. ","from":"developer"},{"body":"[~ahmar] How did we verify that the unit tests are not running?","from":"developer"},{"body":"run `mvn clean test` to just run unit tests\r\n\r\n \r\n\r\nrun `mvn -Dparallel-tests -DtestsThreadCount=8 clean verify` to run both ITests and unit tests.\r\n\r\n \r\n\r\nBoth commands currently running 0 tests, with the output Steve shared in the ticket.","from":"developer"},{"body":"azure and distcp seem to be the same. Looking online it imples that we need that vintage junit thing on the test classpath. \r\n\r\noh, and \r\n{code}\r\nmvn test\r\n{code}\r\n\r\n is sufficient. It just isn't finding any tests","from":"developer"},{"body":"[~ahmar] my s3a stream PR is rebased onto the commit before it. \r\n\r\nFWIW I'm reasonably relieved it want' just me hitting it...I've been doing java17 stuff and assumed that I'd actually got something wrong by mixing things. I've managed to get IntelliJ into that state & it doesn't let me debug unit tests any more. So having maven fail seemed related","from":"developer"},{"body":"[~stevel@apache.org] \r\n\r\nI tested that if we add the following dependency, the unit tests can run normally, although the unit tests throw an error locally:\r\n{code:java}\r\n\r\n  org.junit.vintage\r\n  junit-vintage-engine\r\n  ${junit.vintage.version}\r\n {code}\r\nThe maven-surefire-plugin.version must be upgraded to 3.0.0-M4.\r\n{code:java}\r\n[ERROR] ITestS3AContractBulkDelete.validatePageSize  Time elapsed: 0.002 s  <<< ERROR!\r\njava.lang.NullPointerException: No test bucket\r\n        at org.apache.hadoop.util.Preconditions.checkNotNull(Preconditions.java:88)\r\n        at org.apache.hadoop.fs.s3a.S3ATestUtils.getTestBucketName(S3ATestUtils.java:859)\r\n        at org.apache.hadoop.fs.contract.s3a.ITestS3AContractBulkDelete.createConfiguration(ITestS3AContractBulkDelete.java:88)\r\n        at org.apache.hadoop.fs.contract.AbstractFSContractTestBase.setup(AbstractFSContractTestBase.java:188)\r\n        at org.apache.hadoop.fs.contract.AbstractContractBulkDeleteTest.setup(AbstractContractBulkDeleteTest.java:79)\r\n        at sun.reflect.GeneratedMethodAccessor1.invoke(Unknown Source)\r\n\r\n[ERROR] Tests run: 22, Failures: 0, Errors: 22, Skipped: 0 {code}\r\n ","from":"developer"},{"body":"Thanks for raising thi [~stevel@apache.org] \r\nWe are also facing same issue for azure tests.","from":"developer"},{"body":"[~anujmodi]  I am preparing a new PR. In a module that has not been upgraded to JUnit 5, I am adding a JUnit 4 {{junit-vintage-engine}} dependency, which can help run unit tests.\r\n\r\nTest report can be viewed at the following link:\r\n\r\nhttps://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-7335/2/testReport/","from":"developer"},{"body":"[~anujmodi] rebase your PRs onto the trunk commit before the jersey update and all works, you can then rebase onto trunk once it is in.\r\n\r\nThat jersey2 update is a big piece of work, and critical for java17 -and this is just test execution problems. A new PR to restore test execution is what I'd prefer -and is what is in progress","from":"developer"},{"body":"[~slfan1989] all tests with an ITest prefix are integration tests -they need to be set up with credentials in src/test/resource\r\n\r\nThat stack trace is to be expected","from":"developer"}],"created":"2025-01-28T14:05:46.000+0000","description":"Hadoop-aws tests no longer run.\r\n\r\n{code}\r\n[INFO] -------------------------------------------------------\r\n[INFO] T E S T S\r\n[INFO] -------------------------------------------------------\r\n[INFO]\r\n[INFO] Results:\r\n[INFO]\r\n[INFO] Tests run: 0, Failures: 0, Errors: 0, Skipped: 0\r\n[INFO]\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD SUCCESS\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 2.967 s (Wall Clock)\r\n[INFO] Finished at: 2025-01-28T13:59:52Z\r\n[INFO] ------------------------------------------------------------------------\r\n{code}\r\n\r\nIf I check out to the commit before the jersey v2 update (f38d7072) all is good again.\r\n\r\n","issue_id":"13606497","key":"HADOOP-19405","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-02-04T02:15:43.000+0000","role":"fixed_distractor","summary":"hadoop-aws and hadoop-azure tests have stopped running"} {"case_id":"13610330","cluster":"DISTRACTOR-HADOOP-19476","comments":[{"body":"Merged PR to trunk - https://github.com/apache/hadoop/pull/7452.","created":"2025-03-04T10:56:51.656+0000"}],"conversations":[{"body":"The mvnsite compilation step needs python3. Although we're installing python3 in the build environment, the python3 executable is missing.\r\nThus, we need to create a symbolic link python3 pointing to python.exe needed for mvnsite.\r\n\r\nFailure -\r\n{code}\r\n[INFO] ------------------< org.apache.hadoop:hadoop-common >-------------------\r\n[INFO] Building Apache Hadoop Common 3.5.0-SNAPSHOT [11/115]\r\n[INFO] from hadoop-common-project\\hadoop-common\\pom.xml\r\n[INFO] --------------------------------[ jar ]---------------------------------\r\n[INFO] \r\n[INFO] --- maven-clean-plugin:3.1.0:clean (default-clean) @ hadoop-common ---\r\n[INFO] Deleting C:\\hadoop\\hadoop-common-project\\hadoop-common\\target\r\n[INFO] Deleting C:\\hadoop\\hadoop-common-project\\hadoop-common\\src\\site\\markdown (includes = [UnixShellAPI.md], excludes = [])\r\n[INFO] Deleting C:\\hadoop\\hadoop-common-project\\hadoop-common\\src\\site\\resources (includes = [configuration.xsl, core-default.xml], excludes = [])\r\n[INFO] \r\n[INFO] --- exec-maven-plugin:1.3.1:exec (shelldocs) @ hadoop-common ---\r\n/usr/bin/env: 'python3': No such file or directory\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Reactor Summary for Apache Hadoop Main 3.5.0-SNAPSHOT:\r\n[INFO] \r\n[INFO] Apache Hadoop Main ................................. SUCCESS [21:20 min]\r\n[INFO] Apache Hadoop Build Tools .......................... SUCCESS [ 0.017 s]\r\n[INFO] Apache Hadoop Project POM .......................... SUCCESS [ 2.107 s]\r\n[INFO] Apache Hadoop Annotations .......................... SUCCESS [ 2.018 s]\r\n[INFO] Apache Hadoop Assemblies ........................... SUCCESS [ 0.411 s]\r\n[INFO] Apache Hadoop Project Dist POM ..................... SUCCESS [ 0.326 s]\r\n[INFO] Apache Hadoop Maven Plugins ........................ SUCCESS [ 5.055 s]\r\n[INFO] Apache Hadoop MiniKDC .............................. SUCCESS [ 2.848 s]\r\n[INFO] Apache Hadoop Auth ................................. SUCCESS [ 9.377 s]\r\n[INFO] Apache Hadoop Auth Examples ........................ SUCCESS [ 2.012 s]\r\n[INFO] Apache Hadoop Common ............................... FAILURE [ 6.722 s]\r\n...\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD FAILURE\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 21:54 min\r\n[INFO] Finished at: 2025-02-28T07:36:22Z\r\n[INFO] ------------------------------------------------------------------------\r\n[ERROR] Failed to execute goal org.codehaus.mojo:exec-maven-plugin:1.3.1:exec (shelldocs) on project hadoop-common: Command execution failed. Process exited with an error: 127 (Exit value: 127) -> [Help 1]\r\n[ERROR] \r\n{code}","from":"reporter","subject":"Create python3 symlink needed for mvnsite"},{"body":"Merged PR to trunk - https://github.com/apache/hadoop/pull/7452.","from":"developer"}],"created":"2025-03-02T05:18:49.000+0000","description":"The mvnsite compilation step needs python3. Although we're installing python3 in the build environment, the python3 executable is missing.\r\nThus, we need to create a symbolic link python3 pointing to python.exe needed for mvnsite.\r\n\r\nFailure -\r\n{code}\r\n[INFO] ------------------< org.apache.hadoop:hadoop-common >-------------------\r\n[INFO] Building Apache Hadoop Common 3.5.0-SNAPSHOT [11/115]\r\n[INFO] from hadoop-common-project\\hadoop-common\\pom.xml\r\n[INFO] --------------------------------[ jar ]---------------------------------\r\n[INFO] \r\n[INFO] --- maven-clean-plugin:3.1.0:clean (default-clean) @ hadoop-common ---\r\n[INFO] Deleting C:\\hadoop\\hadoop-common-project\\hadoop-common\\target\r\n[INFO] Deleting C:\\hadoop\\hadoop-common-project\\hadoop-common\\src\\site\\markdown (includes = [UnixShellAPI.md], excludes = [])\r\n[INFO] Deleting C:\\hadoop\\hadoop-common-project\\hadoop-common\\src\\site\\resources (includes = [configuration.xsl, core-default.xml], excludes = [])\r\n[INFO] \r\n[INFO] --- exec-maven-plugin:1.3.1:exec (shelldocs) @ hadoop-common ---\r\n/usr/bin/env: 'python3': No such file or directory\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Reactor Summary for Apache Hadoop Main 3.5.0-SNAPSHOT:\r\n[INFO] \r\n[INFO] Apache Hadoop Main ................................. SUCCESS [21:20 min]\r\n[INFO] Apache Hadoop Build Tools .......................... SUCCESS [ 0.017 s]\r\n[INFO] Apache Hadoop Project POM .......................... SUCCESS [ 2.107 s]\r\n[INFO] Apache Hadoop Annotations .......................... SUCCESS [ 2.018 s]\r\n[INFO] Apache Hadoop Assemblies ........................... SUCCESS [ 0.411 s]\r\n[INFO] Apache Hadoop Project Dist POM ..................... SUCCESS [ 0.326 s]\r\n[INFO] Apache Hadoop Maven Plugins ........................ SUCCESS [ 5.055 s]\r\n[INFO] Apache Hadoop MiniKDC .............................. SUCCESS [ 2.848 s]\r\n[INFO] Apache Hadoop Auth ................................. SUCCESS [ 9.377 s]\r\n[INFO] Apache Hadoop Auth Examples ........................ SUCCESS [ 2.012 s]\r\n[INFO] Apache Hadoop Common ............................... FAILURE [ 6.722 s]\r\n...\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD FAILURE\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 21:54 min\r\n[INFO] Finished at: 2025-02-28T07:36:22Z\r\n[INFO] ------------------------------------------------------------------------\r\n[ERROR] Failed to execute goal org.codehaus.mojo:exec-maven-plugin:1.3.1:exec (shelldocs) on project hadoop-common: Command execution failed. Process exited with an error: 127 (Exit value: 127) -> [Help 1]\r\n[ERROR] \r\n{code}","issue_id":"13610330","key":"HADOOP-19476","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-03-04T10:56:51.000+0000","role":"fixed_distractor","summary":"Create python3 symlink needed for mvnsite"} {"case_id":"13611187","cluster":"DISTRACTOR-HADOOP-19488","comments":[{"body":"We changed how we create the temp directory in HADOOP-19031, and this seems to have broken Windows. This issue does not exist on Linux or macOS.","created":"2025-03-09T16:00:27.432+0000"},{"body":"[~hexiaoqiao] could you please review this and comment? Thanks.","created":"2025-03-10T15:17:43.762+0000"},{"body":"Maven test result with the patch (based on 3.4.1) [^windows-successful-test-run.txt]","created":"2025-03-16T15:48:16.218+0000"},{"body":"Maven test result before the patch (based on 3.4.1) [^windows-test-failures-run.txt]","created":"2025-03-16T22:17:15.733+0000"},{"body":"Merged PR to trunk - https://github.com/apache/hadoop/pull/7511.","created":"2025-03-22T06:48:51.071+0000"},{"body":"I posted the backport patch.","created":"2025-03-22T17:49:31.909+0000"},{"body":"Thanks, [~gaurava].","created":"2025-03-23T23:11:36.689+0000"}],"conversations":[{"body":"On Windows, run {{hadoop jar}} (with any jar). One immediately gets the following exception:\r\n\r\n{code}\r\nException in thread \"main\" java.lang.UnsupportedOperationException: 'posix:permissions' not supported as initial attribute\r\n    at java.base/sun.nio.fs.WindowsSecurityDescriptor.fromAttribute(WindowsSecurityDescriptor.java:358)\r\n    at java.base/sun.nio.fs.WindowsFileSystemProvider.createDirectory(WindowsFileSystemProvider.java:497)\r\n    at java.base/java.nio.file.Files.createDirectory(Files.java:690)\r\n    at java.base/java.nio.file.TempFileHelper.create(TempFileHelper.java:135)\r\n    at java.base/java.nio.file.TempFileHelper.createTempDirectory(TempFileHelper.java:172)\r\n    at java.base/java.nio.file.Files.createTempDirectory(Files.java:966)\r\n    at org.apache.hadoop.util.RunJar.run(RunJar.java:296)\r\n    at org.apache.hadoop.util.RunJar.main(RunJar.java:245)\r\n{code}\r\n\r\nI'm running Windows 11 with OpenJDK 11.0.26.\r\n\r\nThis bug does not exist in 3.3.6.","from":"reporter","subject":"RunJar throws UnsupportedOperationException on Windows"},{"body":"We changed how we create the temp directory in HADOOP-19031, and this seems to have broken Windows. This issue does not exist on Linux or macOS.","from":"developer"},{"body":"[~hexiaoqiao] could you please review this and comment? Thanks.","from":"developer"},{"body":"Maven test result with the patch (based on 3.4.1) [^windows-successful-test-run.txt]","from":"developer"},{"body":"Maven test result before the patch (based on 3.4.1) [^windows-test-failures-run.txt]","from":"developer"},{"body":"Merged PR to trunk - https://github.com/apache/hadoop/pull/7511.","from":"developer"},{"body":"I posted the backport patch.","from":"developer"},{"body":"Thanks, [~gaurava].","from":"developer"}],"created":"2025-03-09T15:58:25.000+0000","description":"On Windows, run {{hadoop jar}} (with any jar). One immediately gets the following exception:\r\n\r\n{code}\r\nException in thread \"main\" java.lang.UnsupportedOperationException: 'posix:permissions' not supported as initial attribute\r\n    at java.base/sun.nio.fs.WindowsSecurityDescriptor.fromAttribute(WindowsSecurityDescriptor.java:358)\r\n    at java.base/sun.nio.fs.WindowsFileSystemProvider.createDirectory(WindowsFileSystemProvider.java:497)\r\n    at java.base/java.nio.file.Files.createDirectory(Files.java:690)\r\n    at java.base/java.nio.file.TempFileHelper.create(TempFileHelper.java:135)\r\n    at java.base/java.nio.file.TempFileHelper.createTempDirectory(TempFileHelper.java:172)\r\n    at java.base/java.nio.file.Files.createTempDirectory(Files.java:966)\r\n    at org.apache.hadoop.util.RunJar.run(RunJar.java:296)\r\n    at org.apache.hadoop.util.RunJar.main(RunJar.java:245)\r\n{code}\r\n\r\nI'm running Windows 11 with OpenJDK 11.0.26.\r\n\r\nThis bug does not exist in 3.3.6.","issue_id":"13611187","key":"HADOOP-19488","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-03-22T06:48:51.000+0000","role":"fixed_distractor","summary":"RunJar throws UnsupportedOperationException on Windows"} {"case_id":"13614533","cluster":"DISTRACTOR-HADOOP-19532","comments":[{"body":"Interestingly, this doesn't happen with later Hadoop versions, probably the metrics code does not use the same code path and avoids this incompatibility.","created":"2025-04-08T09:13:36.066+0000"},{"body":"This has been committed to trunk and branch-3.4.\r\n\r\nThanks for the reviews and committing [~slfan1989] [~stevel@apache.org] [~zeekling].","created":"2025-05-05T06:20:44.730+0000"}],"conversations":[{"body":"The commons-text version used by Hadoop is incompatible with commons-lang3 3.12.\r\n\r\nThis is actually with an older Hadoop, but with the same {_}commons-lang3{_}, _commons-configuration2_ and _commons-text_ versions as trunk:\r\n{noformat}\r\njava.lang.NoSuchMethodError: 'org.apache.commons.lang3.Range org.apache.commons.lang3.Range.of(java.lang.Comparable, java.lang.Comparable)'\r\n    at org.apache.commons.text.translate.NumericEntityEscaper.(NumericEntityEscaper.java:97)\r\n    at org.apache.commons.text.translate.NumericEntityEscaper.between(NumericEntityEscaper.java:59)\r\n    at org.apache.commons.text.StringEscapeUtils.(StringEscapeUtils.java:271)\r\n    at org.apache.commons.configuration2.PropertiesConfiguration$PropertiesReader.unescapePropertyName(PropertiesConfiguration.java:690)\r\n    at org.apache.commons.configuration2.PropertiesConfiguration$PropertiesReader.initPropertyName(PropertiesConfiguration.java:583)\r\n    at org.apache.commons.configuration2.PropertiesConfiguration$PropertiesReader.parseProperty(PropertiesConfiguration.java:640)\r\n    at org.apache.commons.configuration2.PropertiesConfiguration$PropertiesReader.nextProperty(PropertiesConfiguration.java:626)\r\n    at org.apache.commons.configuration2.PropertiesConfigurationLayout.load(PropertiesConfigurationLayout.java:443)\r\n    at org.apache.commons.configuration2.PropertiesConfiguration.read(PropertiesConfiguration.java:1500)\r\n    at org.apache.commons.configuration2.io.FileHandler.loadFromReader(FileHandler.java:712)\r\n    at org.apache.commons.configuration2.io.FileHandler.loadFromTransformedStream(FileHandler.java:782)\r\n    at org.apache.commons.configuration2.io.FileHandler.loadFromStream(FileHandler.java:738)\r\n    at org.apache.commons.configuration2.io.FileHandler.load(FileHandler.java:693)\r\n    at org.apache.commons.configuration2.io.FileHandler.load(FileHandler.java:596)\r\n    at org.apache.commons.configuration2.io.FileHandler.load(FileHandler.java:569)\r\n    at org.apache.hadoop.metrics2.impl.MetricsConfig.loadFirst(MetricsConfig.java:116)\r\n    at org.apache.hadoop.metrics2.impl.MetricsConfig.create(MetricsConfig.java:96)\r\n    at org.apache.hadoop.metrics2.impl.MetricsSystemImpl.configure(MetricsSystemImpl.java:478)\r\n    at org.apache.hadoop.metrics2.impl.MetricsSystemImpl.start(MetricsSystemImpl.java:188)\r\n    at org.apache.hadoop.metrics2.impl.MetricsSystemImpl.init(MetricsSystemImpl.java:163)\r\n    at org.apache.hadoop.metrics2.lib.DefaultMetricsSystem.init(DefaultMetricsSystem.java:62)\r\n    at org.apache.hadoop.metrics2.lib.DefaultMetricsSystem.initialize(DefaultMetricsSystem.java:58)\r\n    at org.apache.phoenix.monitoring.GlobalMetricRegistriesAdapter.(GlobalMetricRegistriesAdapter.java:63)\r\n    at org.apache.phoenix.monitoring.GlobalMetricRegistriesAdapter.(GlobalMetricRegistriesAdapter.java:53)\r\n    at org.apache.phoenix.monitoring.GlobalClientMetrics.(GlobalClientMetrics.java:138)\r\n    at org.apache.phoenix.iterate.SpoolingResultIterator.(SpoolingResultIterator.java:123)\r\n    at org.apache.phoenix.iterate.SpoolingResultIteratorTest.testSpooling(SpoolingResultIteratorTest.java:59)\r\n    at org.apache.phoenix.iterate.SpoolingResultIteratorTest.testOnDiskSpooling(SpoolingResultIteratorTest.java:72)\r\n{noformat}","from":"reporter","subject":"Update commons-lang3 to 3.17.0"},{"body":"Interestingly, this doesn't happen with later Hadoop versions, probably the metrics code does not use the same code path and avoids this incompatibility.","from":"developer"},{"body":"This has been committed to trunk and branch-3.4.\r\n\r\nThanks for the reviews and committing [~slfan1989] [~stevel@apache.org] [~zeekling].","from":"developer"}],"created":"2025-04-08T09:10:39.000+0000","description":"The commons-text version used by Hadoop is incompatible with commons-lang3 3.12.\r\n\r\nThis is actually with an older Hadoop, but with the same {_}commons-lang3{_}, _commons-configuration2_ and _commons-text_ versions as trunk:\r\n{noformat}\r\njava.lang.NoSuchMethodError: 'org.apache.commons.lang3.Range org.apache.commons.lang3.Range.of(java.lang.Comparable, java.lang.Comparable)'\r\n    at org.apache.commons.text.translate.NumericEntityEscaper.(NumericEntityEscaper.java:97)\r\n    at org.apache.commons.text.translate.NumericEntityEscaper.between(NumericEntityEscaper.java:59)\r\n    at org.apache.commons.text.StringEscapeUtils.(StringEscapeUtils.java:271)\r\n    at org.apache.commons.configuration2.PropertiesConfiguration$PropertiesReader.unescapePropertyName(PropertiesConfiguration.java:690)\r\n    at org.apache.commons.configuration2.PropertiesConfiguration$PropertiesReader.initPropertyName(PropertiesConfiguration.java:583)\r\n    at org.apache.commons.configuration2.PropertiesConfiguration$PropertiesReader.parseProperty(PropertiesConfiguration.java:640)\r\n    at org.apache.commons.configuration2.PropertiesConfiguration$PropertiesReader.nextProperty(PropertiesConfiguration.java:626)\r\n    at org.apache.commons.configuration2.PropertiesConfigurationLayout.load(PropertiesConfigurationLayout.java:443)\r\n    at org.apache.commons.configuration2.PropertiesConfiguration.read(PropertiesConfiguration.java:1500)\r\n    at org.apache.commons.configuration2.io.FileHandler.loadFromReader(FileHandler.java:712)\r\n    at org.apache.commons.configuration2.io.FileHandler.loadFromTransformedStream(FileHandler.java:782)\r\n    at org.apache.commons.configuration2.io.FileHandler.loadFromStream(FileHandler.java:738)\r\n    at org.apache.commons.configuration2.io.FileHandler.load(FileHandler.java:693)\r\n    at org.apache.commons.configuration2.io.FileHandler.load(FileHandler.java:596)\r\n    at org.apache.commons.configuration2.io.FileHandler.load(FileHandler.java:569)\r\n    at org.apache.hadoop.metrics2.impl.MetricsConfig.loadFirst(MetricsConfig.java:116)\r\n    at org.apache.hadoop.metrics2.impl.MetricsConfig.create(MetricsConfig.java:96)\r\n    at org.apache.hadoop.metrics2.impl.MetricsSystemImpl.configure(MetricsSystemImpl.java:478)\r\n    at org.apache.hadoop.metrics2.impl.MetricsSystemImpl.start(MetricsSystemImpl.java:188)\r\n    at org.apache.hadoop.metrics2.impl.MetricsSystemImpl.init(MetricsSystemImpl.java:163)\r\n    at org.apache.hadoop.metrics2.lib.DefaultMetricsSystem.init(DefaultMetricsSystem.java:62)\r\n    at org.apache.hadoop.metrics2.lib.DefaultMetricsSystem.initialize(DefaultMetricsSystem.java:58)\r\n    at org.apache.phoenix.monitoring.GlobalMetricRegistriesAdapter.(GlobalMetricRegistriesAdapter.java:63)\r\n    at org.apache.phoenix.monitoring.GlobalMetricRegistriesAdapter.(GlobalMetricRegistriesAdapter.java:53)\r\n    at org.apache.phoenix.monitoring.GlobalClientMetrics.(GlobalClientMetrics.java:138)\r\n    at org.apache.phoenix.iterate.SpoolingResultIterator.(SpoolingResultIterator.java:123)\r\n    at org.apache.phoenix.iterate.SpoolingResultIteratorTest.testSpooling(SpoolingResultIteratorTest.java:59)\r\n    at org.apache.phoenix.iterate.SpoolingResultIteratorTest.testOnDiskSpooling(SpoolingResultIteratorTest.java:72)\r\n{noformat}","issue_id":"13614533","key":"HADOOP-19532","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-05-05T06:20:44.000+0000","role":"fixed_distractor","summary":"Update commons-lang3 to 3.17.0"} {"case_id":"13616035","cluster":"DISTRACTOR-HADOOP-19551","comments":[{"body":"{noformat}\r\n[WARNING] make[2]: *** [CMakeFiles/hadoop_static.dir/build.make:286: CMakeFiles/hadoop_static.dir/main/native/src/org/apache/hadoop/crypto/random/OpensslSecureRandom.c.o] Error 1\r\n[WARNING] make[2]: *** Waiting for unfinished jobs....\r\n[WARNING] /hadoop/hadoop-common-project/hadoop-common/src/main/native/src/org/apache/hadoop/crypto/random/OpensslSecureRandom.c: In function ‘locks_setup’:\r\n[WARNING] /hadoop/hadoop-common-project/hadoop-common/src/main/native/src/org/apache/hadoop/crypto/random/OpensslSecureRandom.c:250:33: error: implicit declaration of function ‘dlsym_CRYPTO_num_locks’; did you mean ‘dlsym_CRYPTO_malloc’? [-Wimplicit-function-declaration]\r\n[WARNING] 250 | lock_cs = dlsym_CRYPTO_malloc(dlsym_CRYPTO_num_locks() * \\\r\n[WARNING] | ^~~~~~~~~~~~~~~~~~~~~~\r\n[WARNING] | dlsym_CRYPTO_malloc\r\n[WARNING] /hadoop/hadoop-common-project/hadoop-common/src/main/native/src/org/apache/hadoop/crypto/random/OpensslSecureRandom.c:257:3: error: implicit declaration of function ‘dlsym_CRYPTO_set_id_callback’; did you mean ‘CRYPTO_set_id_callback’? [-Wimplicit-function-declaration]\r\n[WARNING] 257 | dlsym_CRYPTO_set_id_callback((unsigned long (*)())pthreads_thread_id);\r\n[WARNING] | ^~~~~~~~~~~~~~~~~~~~~~~~~~~~\r\n[WARNING] | CRYPTO_set_id_callback\r\n[WARNING] /hadoop/hadoop-common-project/hadoop-common/src/main/native/src/org/apache/hadoop/crypto/random/OpensslSecureRandom.c:258:3: error: implicit declaration of function ‘dlsym_CRYPTO_set_locking_callback’; did you mean ‘CRYPTO_set_locking_callback’? [-Wimplicit-function-declaration]\r\n[WARNING] 258 | dlsym_CRYPTO_set_locking_callback((void (*)())pthreads_locking_callback);\r\n[WARNING] | ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\r\n[WARNING] | CRYPTO_set_locking_callback\r\n{noformat}\r\n","created":"2025-04-22T14:08:21.819+0000"},{"body":"{noformat}\r\n[WARNING] /hadoop/hadoop-hdfs-project/hadoop-hdfs-native-client/src/main/native/libhdfspp/include/hdfspp/uri.h:60:3: error: ‘uint16_t’ does not name a type\r\n[WARNING] 60 | uint16_t get_port() const;\r\n[WARNING] | ^~~~~~~~\r\n[WARNING] /hadoop/hadoop-hdfs-project/hadoop-hdfs-native-client/src/main/native/libhdfspp/include/hdfspp/uri.h:25:1: note: ‘uint16_t’ is defined in header ‘’; this is probably fixable by adding ‘#include ’\r\n[WARNING] 24 | #include \r\n[WARNING] +++ |+#include \r\n[WARNING] 25 | #include \r\n...\r\n{noformat}\r\n","created":"2025-04-22T14:08:40.424+0000"},{"body":"I found that a part of this JIRA is covered by HDFS-17226. I will merge it first and update [the PR#7644|https://github.com/apache/hadoop/pull/7644] here.","created":"2025-04-22T14:55:06.548+0000"}],"conversations":[{"body":"Building with {{-Pnative}} by using GCC 14 on Fedora 40 failed due to incompatible changes of GCC.","from":"reporter","subject":"Fix compilation error of native libraries on newer GCC"},{"body":"{noformat}\r\n[WARNING] make[2]: *** [CMakeFiles/hadoop_static.dir/build.make:286: CMakeFiles/hadoop_static.dir/main/native/src/org/apache/hadoop/crypto/random/OpensslSecureRandom.c.o] Error 1\r\n[WARNING] make[2]: *** Waiting for unfinished jobs....\r\n[WARNING] /hadoop/hadoop-common-project/hadoop-common/src/main/native/src/org/apache/hadoop/crypto/random/OpensslSecureRandom.c: In function ‘locks_setup’:\r\n[WARNING] /hadoop/hadoop-common-project/hadoop-common/src/main/native/src/org/apache/hadoop/crypto/random/OpensslSecureRandom.c:250:33: error: implicit declaration of function ‘dlsym_CRYPTO_num_locks’; did you mean ‘dlsym_CRYPTO_malloc’? [-Wimplicit-function-declaration]\r\n[WARNING] 250 | lock_cs = dlsym_CRYPTO_malloc(dlsym_CRYPTO_num_locks() * \\\r\n[WARNING] | ^~~~~~~~~~~~~~~~~~~~~~\r\n[WARNING] | dlsym_CRYPTO_malloc\r\n[WARNING] /hadoop/hadoop-common-project/hadoop-common/src/main/native/src/org/apache/hadoop/crypto/random/OpensslSecureRandom.c:257:3: error: implicit declaration of function ‘dlsym_CRYPTO_set_id_callback’; did you mean ‘CRYPTO_set_id_callback’? [-Wimplicit-function-declaration]\r\n[WARNING] 257 | dlsym_CRYPTO_set_id_callback((unsigned long (*)())pthreads_thread_id);\r\n[WARNING] | ^~~~~~~~~~~~~~~~~~~~~~~~~~~~\r\n[WARNING] | CRYPTO_set_id_callback\r\n[WARNING] /hadoop/hadoop-common-project/hadoop-common/src/main/native/src/org/apache/hadoop/crypto/random/OpensslSecureRandom.c:258:3: error: implicit declaration of function ‘dlsym_CRYPTO_set_locking_callback’; did you mean ‘CRYPTO_set_locking_callback’? [-Wimplicit-function-declaration]\r\n[WARNING] 258 | dlsym_CRYPTO_set_locking_callback((void (*)())pthreads_locking_callback);\r\n[WARNING] | ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\r\n[WARNING] | CRYPTO_set_locking_callback\r\n{noformat}\r\n","from":"developer"},{"body":"{noformat}\r\n[WARNING] /hadoop/hadoop-hdfs-project/hadoop-hdfs-native-client/src/main/native/libhdfspp/include/hdfspp/uri.h:60:3: error: ‘uint16_t’ does not name a type\r\n[WARNING] 60 | uint16_t get_port() const;\r\n[WARNING] | ^~~~~~~~\r\n[WARNING] /hadoop/hadoop-hdfs-project/hadoop-hdfs-native-client/src/main/native/libhdfspp/include/hdfspp/uri.h:25:1: note: ‘uint16_t’ is defined in header ‘’; this is probably fixable by adding ‘#include ’\r\n[WARNING] 24 | #include \r\n[WARNING] +++ |+#include \r\n[WARNING] 25 | #include \r\n...\r\n{noformat}\r\n","from":"developer"},{"body":"I found that a part of this JIRA is covered by HDFS-17226. I will merge it first and update [the PR#7644|https://github.com/apache/hadoop/pull/7644] here.","from":"developer"}],"created":"2025-04-22T14:08:07.000+0000","description":"Building with {{-Pnative}} by using GCC 14 on Fedora 40 failed due to incompatible changes of GCC.","issue_id":"13616035","key":"HADOOP-19551","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-05-07T06:03:39.000+0000","role":"fixed_distractor","summary":"Fix compilation error of native libraries on newer GCC"} {"case_id":"13616410","cluster":"DISTRACTOR-HADOOP-19554","comments":[{"body":"seeing a stacktrace in `ITestS3AConfiguration` which looks related; will investigate and do a followup\r\n\r\n{code}\r\n[ERROR] org.apache.hadoop.fs.s3a.ITestS3AConfiguration.testDirectoryAllocatorDefval Time elapsed: 0.029 s <<< ERROR!\r\njava.io.IOException: fs.s3a.buffer.dir not configured\r\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.confChanged(LocalDirAllocator.java:319)\r\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.getLocalPathForWrite(LocalDirAllocator.java:425)\r\n at org.apache.hadoop.fs.LocalDirAllocator.getLocalPathForWrite(LocalDirAllocator.java:171)\r\n at org.apache.hadoop.fs.LocalDirAllocator.getLocalPathForWrite(LocalDirAllocator.java:152)\r\n at org.apache.hadoop.fs.s3a.impl.S3AStoreImpl.createTemporaryFileForWriting(S3AStoreImpl.java:933)\r\n at org.apache.hadoop.fs.s3a.ITestS3AConfiguration.createTemporaryFileForWriting(ITestS3AConfiguration.java:502)\r\n at org.apache.hadoop.fs.s3a.ITestS3AConfiguration.testDirectoryAllocatorDefval(ITestS3AConfiguration.java:491)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:59)\r\n at org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\r\n at org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:56)\r\n at org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\r\n at org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:26)\r\n at org.junit.rules.ExternalResource$1.evaluate(ExternalResource.java:54)\r\n at org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:299)\r\n at org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:293)\r\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n at java.lang.Thread.run(Thread.java:750)\r\n{code}\r\n","created":"2025-05-20T11:14:21.897+0000"}],"conversations":[{"body":"In HADOOP-18636, LocalDirAllocator was modified to recreate missing dirs. But it appears that there are still codepaths which don't do that.\r\n\r\nwhen charrypicking this -please follow up with HADOOP-19573, to ensure an associated test never fails\r\n","from":"reporter","subject":"LocalDirAllocator still doesn't always recover from directory tree deletion"},{"body":"seeing a stacktrace in `ITestS3AConfiguration` which looks related; will investigate and do a followup\r\n\r\n{code}\r\n[ERROR] org.apache.hadoop.fs.s3a.ITestS3AConfiguration.testDirectoryAllocatorDefval Time elapsed: 0.029 s <<< ERROR!\r\njava.io.IOException: fs.s3a.buffer.dir not configured\r\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.confChanged(LocalDirAllocator.java:319)\r\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.getLocalPathForWrite(LocalDirAllocator.java:425)\r\n at org.apache.hadoop.fs.LocalDirAllocator.getLocalPathForWrite(LocalDirAllocator.java:171)\r\n at org.apache.hadoop.fs.LocalDirAllocator.getLocalPathForWrite(LocalDirAllocator.java:152)\r\n at org.apache.hadoop.fs.s3a.impl.S3AStoreImpl.createTemporaryFileForWriting(S3AStoreImpl.java:933)\r\n at org.apache.hadoop.fs.s3a.ITestS3AConfiguration.createTemporaryFileForWriting(ITestS3AConfiguration.java:502)\r\n at org.apache.hadoop.fs.s3a.ITestS3AConfiguration.testDirectoryAllocatorDefval(ITestS3AConfiguration.java:491)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:59)\r\n at org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\r\n at org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:56)\r\n at org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\r\n at org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:26)\r\n at org.junit.rules.ExternalResource$1.evaluate(ExternalResource.java:54)\r\n at org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:299)\r\n at org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:293)\r\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n at java.lang.Thread.run(Thread.java:750)\r\n{code}\r\n","from":"developer"}],"created":"2025-04-25T16:31:10.000+0000","description":"In HADOOP-18636, LocalDirAllocator was modified to recreate missing dirs. But it appears that there are still codepaths which don't do that.\r\n\r\nwhen charrypicking this -please follow up with HADOOP-19573, to ensure an associated test never fails\r\n","issue_id":"13616410","key":"HADOOP-19554","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-05-19T14:26:39.000+0000","role":"fixed_distractor","summary":"LocalDirAllocator still doesn't always recover from directory tree deletion"} {"case_id":"13618677","cluster":"DISTRACTOR-HADOOP-19573","comments":[{"body":"{code}\r\n[ERROR] org.apache.hadoop.fs.s3a.ITestS3AConfiguration.testDirectoryAllocatorDefval Time elapsed: 0.029 s <<< ERROR!\r\njava.io.IOException: fs.s3a.buffer.dir not configured\r\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.confChanged(LocalDirAllocator.java:319)\r\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.getLocalPathForWrite(LocalDirAllocator.java:425)\r\n at org.apache.hadoop.fs.LocalDirAllocator.getLocalPathForWrite(LocalDirAllocator.java:171)\r\n at org.apache.hadoop.fs.LocalDirAllocator.getLocalPathForWrite(LocalDirAllocator.java:152)\r\n at org.apache.hadoop.fs.s3a.impl.S3AStoreImpl.createTemporaryFileForWriting(S3AStoreImpl.java:933)\r\n at org.apache.hadoop.fs.s3a.ITestS3AConfiguration.createTemporaryFileForWriting(ITestS3AConfiguration.java:502)\r\n at org.apache.hadoop.fs.s3a.ITestS3AConfiguration.testDirectoryAllocatorDefval(ITestS3AConfiguration.java:491)\r\n{code}\r\n","created":"2025-05-20T14:03:48.856+0000"}],"conversations":[{"body":"while working on HADOOP-19554 I added a per-bucket setting for fs.s3a.buffer.dir\r\n\r\nafter this, ITestS3AConfiguration.testDirectoryAllocatorDefval() would fail in a test run of the entire class, but not if run alone.\r\n\r\nCauses\r\n* dir allocator map of config key to allocator is static; previous uses tainted outcome\r\n* per-bucket settings were't being overridden. This is complicated by the fact that \"unset\" isn't a setting, therefore can't be forced in. Instead some whitespace needs to be set.","from":"reporter","subject":"S3A: ITestS3AConfiguration.testDirectoryAllocatorDefval() failing"},{"body":"{code}\r\n[ERROR] org.apache.hadoop.fs.s3a.ITestS3AConfiguration.testDirectoryAllocatorDefval Time elapsed: 0.029 s <<< ERROR!\r\njava.io.IOException: fs.s3a.buffer.dir not configured\r\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.confChanged(LocalDirAllocator.java:319)\r\n at org.apache.hadoop.fs.LocalDirAllocator$AllocatorPerContext.getLocalPathForWrite(LocalDirAllocator.java:425)\r\n at org.apache.hadoop.fs.LocalDirAllocator.getLocalPathForWrite(LocalDirAllocator.java:171)\r\n at org.apache.hadoop.fs.LocalDirAllocator.getLocalPathForWrite(LocalDirAllocator.java:152)\r\n at org.apache.hadoop.fs.s3a.impl.S3AStoreImpl.createTemporaryFileForWriting(S3AStoreImpl.java:933)\r\n at org.apache.hadoop.fs.s3a.ITestS3AConfiguration.createTemporaryFileForWriting(ITestS3AConfiguration.java:502)\r\n at org.apache.hadoop.fs.s3a.ITestS3AConfiguration.testDirectoryAllocatorDefval(ITestS3AConfiguration.java:491)\r\n{code}\r\n","from":"developer"}],"created":"2025-05-20T13:59:50.000+0000","description":"while working on HADOOP-19554 I added a per-bucket setting for fs.s3a.buffer.dir\r\n\r\nafter this, ITestS3AConfiguration.testDirectoryAllocatorDefval() would fail in a test run of the entire class, but not if run alone.\r\n\r\nCauses\r\n* dir allocator map of config key to allocator is static; previous uses tainted outcome\r\n* per-bucket settings were't being overridden. This is complicated by the fact that \"unset\" isn't a setting, therefore can't be forced in. Instead some whitespace needs to be set.","issue_id":"13618677","key":"HADOOP-19573","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-07-08T16:57:30.000+0000","role":"fixed_distractor","summary":"S3A: ITestS3AConfiguration.testDirectoryAllocatorDefval() failing"} {"case_id":"13618886","cluster":"DISTRACTOR-HADOOP-19576","comments":[{"body":"[~stevel@apache.org]  - I understand this was done because to delete a directory in S3 express directory, All the pending upload needs to be purged. But this causes problem in insert overwrite type of job with MagicCommitter.\r\n\r\n \r\n\r\nI think we  should make  directory purge operation  \"false\" by default for all types of buckets - Any thoughts on this ?","created":"2025-05-22T10:21:22.464+0000"},{"body":"people trying to do INSERT OVERWRITE? hmm. \r\n\r\nok, set to default, mark as incompatible, and update the docs with \r\n* a section on this. \r\n* what the error means\r\n\r\npeople should always have auto cleanup enabled anyway, shouldn't they? always always always. in fact, aws cost analyzer should warn if you don't (does it do this?). and given that, shouldn't really matter much. we have CLI tools in terms of \"hadoop s3guard uploads\" to enum/purge it too. \r\n\r\nFWIW one of our qe buckets didn't do the cleanup, I used it for a scale test of MPU cleanup. We could actually have an ILoadTest for this now that createFile() lets you create a zero byte MPU","created":"2025-05-28T12:47:35.973+0000"},{"body":"Yes - it is always recommended to have S3 bucket life cycle policy to clean up dangling MPUs after a certain threshold period of time.","created":"2025-06-02T05:59:55.475+0000"}],"conversations":[{"body":"Query engines which uses Magic Committer to overwrite a directory would ideally upload the MPUs (not complete) and then delete the contents of the directory before committing the MPU.\r\n\r\n \r\n\r\nFor S3 express storage, The directory purge operation is enabled by default. Refer [here|https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/S3AFileSystem.java#L688] for code pointers.\r\n\r\n \r\n\r\nDue to this, the pending MPU uploads are purged and query fails with \r\n\r\n{{NoSuchUpload: The specified multipart upload does not exist. The upload ID might be invalid, or the multipart upload might have been aborted or completed. }}","from":"reporter","subject":"Insert Overwrite Jobs With MagicCommitter Fails On S3 Express Storage"},{"body":"[~stevel@apache.org]  - I understand this was done because to delete a directory in S3 express directory, All the pending upload needs to be purged. But this causes problem in insert overwrite type of job with MagicCommitter.\r\n\r\n \r\n\r\nI think we  should make  directory purge operation  \"false\" by default for all types of buckets - Any thoughts on this ?","from":"developer"},{"body":"people trying to do INSERT OVERWRITE? hmm. \r\n\r\nok, set to default, mark as incompatible, and update the docs with \r\n* a section on this. \r\n* what the error means\r\n\r\npeople should always have auto cleanup enabled anyway, shouldn't they? always always always. in fact, aws cost analyzer should warn if you don't (does it do this?). and given that, shouldn't really matter much. we have CLI tools in terms of \"hadoop s3guard uploads\" to enum/purge it too. \r\n\r\nFWIW one of our qe buckets didn't do the cleanup, I used it for a scale test of MPU cleanup. We could actually have an ILoadTest for this now that createFile() lets you create a zero byte MPU","from":"developer"},{"body":"Yes - it is always recommended to have S3 bucket life cycle policy to clean up dangling MPUs after a certain threshold period of time.","from":"developer"}],"created":"2025-05-22T10:18:50.000+0000","description":"Query engines which uses Magic Committer to overwrite a directory would ideally upload the MPUs (not complete) and then delete the contents of the directory before committing the MPU.\r\n\r\n \r\n\r\nFor S3 express storage, The directory purge operation is enabled by default. Refer [here|https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/S3AFileSystem.java#L688] for code pointers.\r\n\r\n \r\n\r\nDue to this, the pending MPU uploads are purged and query fails with \r\n\r\n{{NoSuchUpload: The specified multipart upload does not exist. The upload ID might be invalid, or the multipart upload might have been aborted or completed. }}","issue_id":"13618886","key":"HADOOP-19576","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-07-08T14:42:39.000+0000","role":"fixed_distractor","summary":"Insert Overwrite Jobs With MagicCommitter Fails On S3 Express Storage"} {"case_id":"12333193","cluster":"DISTRACTOR-HADOOP-196","comments":[{"body":"The constructor: \n\n\tpublic Configuration(Configuration other)\n\t\nsets the defaultResources, and the finalResources but does not use them \nbecause it then sets the properties to a clone of the other.properties.\n\nIf someone then calls addFinalResource (as does JobConf) then the\n'other' properties are just lost. This becomes a problem if you attempt to use: \n\tpublic JobConf(Class exampleClass) \n\nMy fix (there are several possible ways to fix this) is to just remember\nthe other.properties and then overlay then whenever getProps() loads the resources.\n\n","created":"2006-05-05T04:58:52.000+0000"},{"body":"Some mistakes in this submission (my first one, and I can't seem to edit them away).\n\nLet me restate the bug.\n\nThis constructor\n\n public JobConf(Configuration conf) {\n super(conf);\n initialize();\n }\n\ndoes not work as it should. The conf that is passed in gets lost.\n\nThe patch fixes this bug.\n\n- alan wootton, shopping.com","created":"2006-05-12T01:51:57.000+0000"},{"body":"Is there a reason why the patch was not submitted? I have encountered the same problem and patched it myself in a very similar way before finding out it had already been reported. Currently Configuration(conf) does not behave as expected.\n\nLorenzo Thione \nPowerset, Inc.","created":"2006-07-12T10:03:20.000+0000"},{"body":"> Is there a reason why the patch was not submitted?\n\nNot a good one. I think it just fell off my radar.\n\nIt would be good to add a unit test for this. I think the use case is roughly:\n\nConfiguration c1 = new Configuration();\nc1.set(\"foo\", \"bar\");\nConfiguratoin c2 = new Configuration(c1);\nassertEquals(c2.get(\"foo\"), \"bar\");\n\nIs that right?\n\nThe provided patch would fix this, but by setting all the values except \"foo\" twice: once when reloaded from resources and once when reloaded from overrides. Perhaps it would be better to put all resource-loaded properties in a nested Properties instance that's inherited via 'new Properties(Properties)' and only store values set directly in a top-level Properties instance. Or perhaps we should just set most of the values twice and not worry about it. Thoughts, anyone?","created":"2006-07-12T16:50:44.000+0000"},{"body":"Attached patch fixes problem with current implementation, dynamically set properties are now preserved. Extension to testcase is also provided.","created":"2006-08-31T16:43:55.000+0000"},{"body":"I just committed this. Thanks, Sami!","created":"2006-08-31T20:58:38.000+0000"}],"conversations":[{"body":"The constructor \npublic Configuration(Configuration other) ","from":"reporter","subject":"Fix buggy uselessness of Configuration( Configuration other) constructor"},{"body":"The constructor: \n\n\tpublic Configuration(Configuration other)\n\t\nsets the defaultResources, and the finalResources but does not use them \nbecause it then sets the properties to a clone of the other.properties.\n\nIf someone then calls addFinalResource (as does JobConf) then the\n'other' properties are just lost. This becomes a problem if you attempt to use: \n\tpublic JobConf(Class exampleClass) \n\nMy fix (there are several possible ways to fix this) is to just remember\nthe other.properties and then overlay then whenever getProps() loads the resources.\n\n","from":"developer"},{"body":"Some mistakes in this submission (my first one, and I can't seem to edit them away).\n\nLet me restate the bug.\n\nThis constructor\n\n public JobConf(Configuration conf) {\n super(conf);\n initialize();\n }\n\ndoes not work as it should. The conf that is passed in gets lost.\n\nThe patch fixes this bug.\n\n- alan wootton, shopping.com","from":"developer"},{"body":"Is there a reason why the patch was not submitted? I have encountered the same problem and patched it myself in a very similar way before finding out it had already been reported. Currently Configuration(conf) does not behave as expected.\n\nLorenzo Thione \nPowerset, Inc.","from":"developer"},{"body":"> Is there a reason why the patch was not submitted?\n\nNot a good one. I think it just fell off my radar.\n\nIt would be good to add a unit test for this. I think the use case is roughly:\n\nConfiguration c1 = new Configuration();\nc1.set(\"foo\", \"bar\");\nConfiguratoin c2 = new Configuration(c1);\nassertEquals(c2.get(\"foo\"), \"bar\");\n\nIs that right?\n\nThe provided patch would fix this, but by setting all the values except \"foo\" twice: once when reloaded from resources and once when reloaded from overrides. Perhaps it would be better to put all resource-loaded properties in a nested Properties instance that's inherited via 'new Properties(Properties)' and only store values set directly in a top-level Properties instance. Or perhaps we should just set most of the values twice and not worry about it. Thoughts, anyone?","from":"developer"},{"body":"Attached patch fixes problem with current implementation, dynamically set properties are now preserved. Extension to testcase is also provided.","from":"developer"},{"body":"I just committed this. Thanks, Sami!","from":"developer"}],"created":"2006-05-05T04:49:48.000+0000","description":"The constructor \npublic Configuration(Configuration other) ","issue_id":"12333193","key":"HADOOP-196","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-08-31T20:58:38.000+0000","role":"fixed_distractor","summary":"Fix buggy uselessness of Configuration( Configuration other) constructor"} {"case_id":"13622060","cluster":"DISTRACTOR-HADOOP-19600","comments":[{"body":"Merged PR to trunk - https://github.com/apache/hadoop/pull/7768.","created":"2025-06-29T06:57:24.854+0000"}],"conversations":[{"body":"The Jenkins CI for Windows Nightly build is failing since it's unable to download Apache Maven -\r\n\r\n{code}\r\n00:32:18 Step 14/78 : RUN powershell Invoke-WebRequest -URI https://downloads.apache.org/maven/maven-3/3.8.8/binaries/apache-maven-3.8.8-bin.zip -OutFile $Env:TEMP\\apache-maven-3.8.8-bin.zip\r\n00:32:18 ---> Running in a47cc5638ee5\r\n00:32:22 \u001b[91mInvoke-WebRequest : \r\n00:32:22 \u001b[0m\u001b[91m\r\n00:32:22 \u001b[0m\u001b[91m404 Not Found\r\n00:32:22 \u001b[0m\u001b[91m\r\n00:32:22 \u001b[0m\u001b[91m

Not Found

\r\n00:32:22 \u001b[0m\u001b[91m

The requested URL was not found on this server.

\r\n00:32:22 \u001b[0m\u001b[91m\r\n00:32:22 \u001b[0m\u001b[91mAt line:1 char:1\r\n00:32:22 + Invoke-WebRequest -URI https://downloads.apache.org/maven/maven-3/3.8 ...\r\n00:32:22 \u001b[0m\u001b[91m+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\r\n00:32:22 \u001b[0m\u001b[91m + CategoryInfo : InvalidOperation: (System.Net.HttpWebRequest:Htt \r\n00:32:22 \u001b[0m\u001b[91m pWebRequest) [Invoke-WebRequest], WebException\r\n00:32:22 \u001b[0m\u001b[91m + FullyQualifiedErrorId : WebCmdletWebResponseException,Microsoft.PowerShe \r\n00:32:22 ll.Commands.InvokeWebRequestCommand\r\n00:32:38 \u001b[0mThe command 'cmd /S /C powershell Invoke-WebRequest -URI https://downloads.apache.org/maven/maven-3/3.8.8/binaries/apache-maven-3.8.8-bin.zip -OutFile $Env:TEMP\\apache-maven-3.8.8-bin.zip' returned a non-zero code: 1\r\n[Pipeline] }\r\n[Pipeline] // withCredentials\r\n{code}\r\n\r\nWe need to update the URL to https://archive.apache.org/dist/maven/maven-3/3.8.8/binaries/apache-maven-3.8.8-bin.zip","from":"reporter","subject":"Fix Maven download link"},{"body":"Merged PR to trunk - https://github.com/apache/hadoop/pull/7768.","from":"developer"}],"created":"2025-06-28T18:05:06.000+0000","description":"The Jenkins CI for Windows Nightly build is failing since it's unable to download Apache Maven -\r\n\r\n{code}\r\n00:32:18 Step 14/78 : RUN powershell Invoke-WebRequest -URI https://downloads.apache.org/maven/maven-3/3.8.8/binaries/apache-maven-3.8.8-bin.zip -OutFile $Env:TEMP\\apache-maven-3.8.8-bin.zip\r\n00:32:18 ---> Running in a47cc5638ee5\r\n00:32:22 \u001b[91mInvoke-WebRequest : \r\n00:32:22 \u001b[0m\u001b[91m\r\n00:32:22 \u001b[0m\u001b[91m404 Not Found\r\n00:32:22 \u001b[0m\u001b[91m\r\n00:32:22 \u001b[0m\u001b[91m

Not Found

\r\n00:32:22 \u001b[0m\u001b[91m

The requested URL was not found on this server.

\r\n00:32:22 \u001b[0m\u001b[91m\r\n00:32:22 \u001b[0m\u001b[91mAt line:1 char:1\r\n00:32:22 + Invoke-WebRequest -URI https://downloads.apache.org/maven/maven-3/3.8 ...\r\n00:32:22 \u001b[0m\u001b[91m+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\r\n00:32:22 \u001b[0m\u001b[91m + CategoryInfo : InvalidOperation: (System.Net.HttpWebRequest:Htt \r\n00:32:22 \u001b[0m\u001b[91m pWebRequest) [Invoke-WebRequest], WebException\r\n00:32:22 \u001b[0m\u001b[91m + FullyQualifiedErrorId : WebCmdletWebResponseException,Microsoft.PowerShe \r\n00:32:22 ll.Commands.InvokeWebRequestCommand\r\n00:32:38 \u001b[0mThe command 'cmd /S /C powershell Invoke-WebRequest -URI https://downloads.apache.org/maven/maven-3/3.8.8/binaries/apache-maven-3.8.8-bin.zip -OutFile $Env:TEMP\\apache-maven-3.8.8-bin.zip' returned a non-zero code: 1\r\n[Pipeline] }\r\n[Pipeline] // withCredentials\r\n{code}\r\n\r\nWe need to update the URL to https://archive.apache.org/dist/maven/maven-3/3.8.8/binaries/apache-maven-3.8.8-bin.zip","issue_id":"13622060","key":"HADOOP-19600","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-06-29T06:57:24.000+0000","role":"fixed_distractor","summary":"Fix Maven download link"} {"case_id":"13626490","cluster":"DISTRACTOR-HADOOP-19651","comments":[{"body":"Merged PR to trunk - https://github.com/apache/hadoop/pull/7875.","created":"2025-08-16T19:37:06.702+0000"}],"conversations":[{"body":"The currently used libopenssl-3.1.4-1 is no longer available on the msys repo - https://repo.msys2.org/msys/x86_64/libopenssl-3.1.4-1-x86_64.pkg.tar.zst\r\n\r\nThe Jenkins CI thus fails to download the same and thus it fails to build the docker image for Windows.\r\n\r\n{code}\r\n00:25:33 SUCCESS: Specified value was saved.\r\n00:25:46 Removing intermediate container 5ce7355571a1\r\n00:25:46 ---> a13a4bc69545\r\n00:25:46 Step 21/78 : RUN powershell Invoke-WebRequest -Uri https://repo.msys2.org/msys/x86_64/libopenssl-3.1.4-1-x86_64.pkg.tar.zst -OutFile $Env:TEMP\\libopenssl-3.1.4-1-x86_64.pkg.tar.zst\r\n00:25:46 ---> Running in d2dafad446f9\r\n00:25:54 \u001b[91mInvoke-WebRequest : The remote server returned an error: (404) Not Found.\r\n00:25:54 \u001b[0m\u001b[91mAt line:1 char:1\r\n00:25:54 \u001b[0m\u001b[91m+ Invoke-WebRequest -Uri https://repo.msys2.org/msys/x86_64/libopenssl- ...\r\n00:25:54 + ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\r\n00:25:54 \u001b[0m\u001b[91m + CategoryInfo : InvalidOperation: (System.Net.HttpWebRequest:Htt\r\n{code}\r\n\r\nThus, we need to upgrade to the latest version to address this.","from":"reporter","subject":"Upgrade libopenssl to 3.5.2-1 needed for rsync"},{"body":"Merged PR to trunk - https://github.com/apache/hadoop/pull/7875.","from":"developer"}],"created":"2025-08-15T19:19:04.000+0000","description":"The currently used libopenssl-3.1.4-1 is no longer available on the msys repo - https://repo.msys2.org/msys/x86_64/libopenssl-3.1.4-1-x86_64.pkg.tar.zst\r\n\r\nThe Jenkins CI thus fails to download the same and thus it fails to build the docker image for Windows.\r\n\r\n{code}\r\n00:25:33 SUCCESS: Specified value was saved.\r\n00:25:46 Removing intermediate container 5ce7355571a1\r\n00:25:46 ---> a13a4bc69545\r\n00:25:46 Step 21/78 : RUN powershell Invoke-WebRequest -Uri https://repo.msys2.org/msys/x86_64/libopenssl-3.1.4-1-x86_64.pkg.tar.zst -OutFile $Env:TEMP\\libopenssl-3.1.4-1-x86_64.pkg.tar.zst\r\n00:25:46 ---> Running in d2dafad446f9\r\n00:25:54 \u001b[91mInvoke-WebRequest : The remote server returned an error: (404) Not Found.\r\n00:25:54 \u001b[0m\u001b[91mAt line:1 char:1\r\n00:25:54 \u001b[0m\u001b[91m+ Invoke-WebRequest -Uri https://repo.msys2.org/msys/x86_64/libopenssl- ...\r\n00:25:54 + ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\r\n00:25:54 \u001b[0m\u001b[91m + CategoryInfo : InvalidOperation: (System.Net.HttpWebRequest:Htt\r\n{code}\r\n\r\nThus, we need to upgrade to the latest version to address this.","issue_id":"13626490","key":"HADOOP-19651","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-08-16T19:37:06.000+0000","role":"fixed_distractor","summary":"Upgrade libopenssl to 3.5.2-1 needed for rsync"} {"case_id":"13629211","cluster":"DISTRACTOR-HADOOP-19697","comments":[{"body":"Trying to list local root \r\n\r\nIf there are dependencies needed in the HADOOP-19696 let's make sure they get into common/lib, but this registration process mustn't fail this way, so let's just have a fs.gs.impl declararation in core-default.xml\r\n\r\n{code}\r\n bin/hadoop fs -ls file:///\r\n2025-09-17 15:03:47,688 [main] WARN fs.FileSystem (FileSystem.java:loadFileSystems(3539)) - Cannot load filesystem\r\njava.util.ServiceConfigurationError: org.apache.hadoop.fs.FileSystem: org.apache.hadoop.fs.gs.GoogleHadoopFileSystem Unable to get public no-arg constructor\r\n at java.base/java.util.ServiceLoader.fail(ServiceLoader.java:586)\r\n at java.base/java.util.ServiceLoader.getConstructor(ServiceLoader.java:679)\r\n at java.base/java.util.ServiceLoader$LazyClassPathLookupIterator.hasNextService(ServiceLoader.java:1240)\r\n at java.base/java.util.ServiceLoader$LazyClassPathLookupIterator.hasNext(ServiceLoader.java:1273)\r\n at java.base/java.util.ServiceLoader$2.hasNext(ServiceLoader.java:1309)\r\n at java.base/java.util.ServiceLoader$3.hasNext(ServiceLoader.java:1393)\r\n at org.apache.hadoop.fs.FileSystem.loadFileSystems(FileSystem.java:3522)\r\n at org.apache.hadoop.fs.FileSystem.getFileSystemClass(FileSystem.java:3562)\r\n at org.apache.hadoop.fs.FileSystem.createFileSystem(FileSystem.java:3612)\r\n at org.apache.hadoop.fs.FileSystem$Cache.getInternal(FileSystem.java:3716)\r\n at org.apache.hadoop.fs.FileSystem$Cache.get(FileSystem.java:3667)\r\n at org.apache.hadoop.fs.FileSystem.get(FileSystem.java:557)\r\n at org.apache.hadoop.fs.Path.getFileSystem(Path.java:373)\r\n at org.apache.hadoop.fs.shell.PathData.expandAsGlob(PathData.java:347)\r\n at org.apache.hadoop.fs.shell.Command.expandArgument(Command.java:265)\r\n at org.apache.hadoop.fs.shell.Command.expandArguments(Command.java:248)\r\n at org.apache.hadoop.fs.shell.FsCommand.processRawArguments(FsCommand.java:105)\r\n at org.apache.hadoop.fs.shell.Command.run(Command.java:192)\r\n at org.apache.hadoop.fs.FsShell.run(FsShell.java:327)\r\n at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:82)\r\n at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:97)\r\n at org.apache.hadoop.fs.FsShell.main(FsShell.java:390)\r\nCaused by: java.lang.NoClassDefFoundError: com/google/auth/Credentials\r\n at java.base/java.lang.Class.getDeclaredConstructors0(Native Method)\r\n at java.base/java.lang.Class.privateGetDeclaredConstructors(Class.java:3373)\r\n at java.base/java.lang.Class.getConstructor0(Class.java:3578)\r\n at java.base/java.lang.Class.getConstructor(Class.java:2271)\r\n at java.base/java.util.ServiceLoader$1.run(ServiceLoader.java:666)\r\n at java.base/java.util.ServiceLoader$1.run(ServiceLoader.java:663)\r\n at java.base/java.security.AccessController.doPrivileged(AccessController.java:569)\r\n at java.base/java.util.ServiceLoader.getConstructor(ServiceLoader.java:674)\r\n ... 20 more\r\nCaused by: java.lang.ClassNotFoundException: com.google.auth.Credentials\r\n at java.base/jdk.internal.loader.BuiltinClassLoader.loadClass(BuiltinClassLoader.java:641)\r\n at java.base/jdk.internal.loader.ClassLoaders$AppClassLoader.loadClass(ClassLoaders.java:188)\r\n at java.base/java.lang.ClassLoader.loadClass(ClassLoader.java:525)\r\n ... 28 more\r\nFound 20 items\r\n---------- 1 root admin 0 2025-08-16 19:44 file:///.file\r\n...\r\n{code}\r\n","created":"2025-09-17T14:06:18.904+0000"},{"body":"Hi [~stevel@apache.org]. Not fully caught up on this one, but there is a {{fs.gs.impl}} entry in core-default.xml now. Was that all we needed?\r\n\r\nhttps://github.com/apache/hadoop/blob/trunk/hadoop-common-project/hadoop-common/src/main/resources/core-default.xml#L4508","created":"2025-10-08T04:44:25.761+0000"},{"body":"Hi [~stevel@apache.org]. Is this bug resolved now that you removed the hadoop-gcp ServiceLoader declaration in this PR?\r\n\r\nhttps://github.com/apache/hadoop/pull/7980","created":"2025-12-17T21:53:36.640+0000"}],"conversations":[{"body":"Surfaced during HADOOP-19696 and work with all the cloud connectors on the classpath.\r\n\r\nThere's a missing dependency causing the gcs connector to fail to register *when the first filesystem is instantiated*, because the service registration process loads the gcs connector class, instantiates one and asks for its schema.\r\n\r\nAs well as a sign of a problem, it's better to just add an entry in core-default.xml as this saves all classloading overhead. It's what the other asf bundled ones do.","from":"reporter","subject":"google gs connector registration failing"},{"body":"Trying to list local root \r\n\r\nIf there are dependencies needed in the HADOOP-19696 let's make sure they get into common/lib, but this registration process mustn't fail this way, so let's just have a fs.gs.impl declararation in core-default.xml\r\n\r\n{code}\r\n bin/hadoop fs -ls file:///\r\n2025-09-17 15:03:47,688 [main] WARN fs.FileSystem (FileSystem.java:loadFileSystems(3539)) - Cannot load filesystem\r\njava.util.ServiceConfigurationError: org.apache.hadoop.fs.FileSystem: org.apache.hadoop.fs.gs.GoogleHadoopFileSystem Unable to get public no-arg constructor\r\n at java.base/java.util.ServiceLoader.fail(ServiceLoader.java:586)\r\n at java.base/java.util.ServiceLoader.getConstructor(ServiceLoader.java:679)\r\n at java.base/java.util.ServiceLoader$LazyClassPathLookupIterator.hasNextService(ServiceLoader.java:1240)\r\n at java.base/java.util.ServiceLoader$LazyClassPathLookupIterator.hasNext(ServiceLoader.java:1273)\r\n at java.base/java.util.ServiceLoader$2.hasNext(ServiceLoader.java:1309)\r\n at java.base/java.util.ServiceLoader$3.hasNext(ServiceLoader.java:1393)\r\n at org.apache.hadoop.fs.FileSystem.loadFileSystems(FileSystem.java:3522)\r\n at org.apache.hadoop.fs.FileSystem.getFileSystemClass(FileSystem.java:3562)\r\n at org.apache.hadoop.fs.FileSystem.createFileSystem(FileSystem.java:3612)\r\n at org.apache.hadoop.fs.FileSystem$Cache.getInternal(FileSystem.java:3716)\r\n at org.apache.hadoop.fs.FileSystem$Cache.get(FileSystem.java:3667)\r\n at org.apache.hadoop.fs.FileSystem.get(FileSystem.java:557)\r\n at org.apache.hadoop.fs.Path.getFileSystem(Path.java:373)\r\n at org.apache.hadoop.fs.shell.PathData.expandAsGlob(PathData.java:347)\r\n at org.apache.hadoop.fs.shell.Command.expandArgument(Command.java:265)\r\n at org.apache.hadoop.fs.shell.Command.expandArguments(Command.java:248)\r\n at org.apache.hadoop.fs.shell.FsCommand.processRawArguments(FsCommand.java:105)\r\n at org.apache.hadoop.fs.shell.Command.run(Command.java:192)\r\n at org.apache.hadoop.fs.FsShell.run(FsShell.java:327)\r\n at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:82)\r\n at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:97)\r\n at org.apache.hadoop.fs.FsShell.main(FsShell.java:390)\r\nCaused by: java.lang.NoClassDefFoundError: com/google/auth/Credentials\r\n at java.base/java.lang.Class.getDeclaredConstructors0(Native Method)\r\n at java.base/java.lang.Class.privateGetDeclaredConstructors(Class.java:3373)\r\n at java.base/java.lang.Class.getConstructor0(Class.java:3578)\r\n at java.base/java.lang.Class.getConstructor(Class.java:2271)\r\n at java.base/java.util.ServiceLoader$1.run(ServiceLoader.java:666)\r\n at java.base/java.util.ServiceLoader$1.run(ServiceLoader.java:663)\r\n at java.base/java.security.AccessController.doPrivileged(AccessController.java:569)\r\n at java.base/java.util.ServiceLoader.getConstructor(ServiceLoader.java:674)\r\n ... 20 more\r\nCaused by: java.lang.ClassNotFoundException: com.google.auth.Credentials\r\n at java.base/jdk.internal.loader.BuiltinClassLoader.loadClass(BuiltinClassLoader.java:641)\r\n at java.base/jdk.internal.loader.ClassLoaders$AppClassLoader.loadClass(ClassLoaders.java:188)\r\n at java.base/java.lang.ClassLoader.loadClass(ClassLoader.java:525)\r\n ... 28 more\r\nFound 20 items\r\n---------- 1 root admin 0 2025-08-16 19:44 file:///.file\r\n...\r\n{code}\r\n","from":"developer"},{"body":"Hi [~stevel@apache.org]. Not fully caught up on this one, but there is a {{fs.gs.impl}} entry in core-default.xml now. Was that all we needed?\r\n\r\nhttps://github.com/apache/hadoop/blob/trunk/hadoop-common-project/hadoop-common/src/main/resources/core-default.xml#L4508","from":"developer"},{"body":"Hi [~stevel@apache.org]. Is this bug resolved now that you removed the hadoop-gcp ServiceLoader declaration in this PR?\r\n\r\nhttps://github.com/apache/hadoop/pull/7980","from":"developer"}],"created":"2025-09-17T14:04:27.000+0000","description":"Surfaced during HADOOP-19696 and work with all the cloud connectors on the classpath.\r\n\r\nThere's a missing dependency causing the gcs connector to fail to register *when the first filesystem is instantiated*, because the service registration process loads the gcs connector class, instantiates one and asks for its schema.\r\n\r\nAs well as a sign of a problem, it's better to just add an entry in core-default.xml as this saves all classloading overhead. It's what the other asf bundled ones do.","issue_id":"13629211","key":"HADOOP-19697","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2026-02-12T19:38:56.000+0000","role":"fixed_distractor","summary":"google gs connector registration failing"} {"case_id":"13630839","cluster":"DISTRACTOR-HADOOP-19717","comments":[{"body":"this is complicating the new thirdparty release FWIW. this should all be using the unshaded javax. Nullable/nonnull. And the hadoop-thirdparty release needs to address this stuff getting left out so 1.5.0 can be a drop-in replacement for 1.4.0\r\n{code}\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.10.1:compile (default-compile) on project hadoop-tos: Compilation failure: Compilation failure: \r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-cloud-storage-project/hadoop-tos/src/main/java/org/apache/hadoop/fs/tosfs/util/Iterables.java:[22,79] package org.apache.hadoop.thirdparty.org.checkerframework.checker.nullness.qual does not exist\r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-cloud-storage-project/hadoop-tos/src/main/java/org/apache/hadoop/fs/tosfs/util/Iterables.java:[89,29] cannot find symbol\r\n[ERROR] symbol: class Nullable\r\n[ERROR] location: class org.apache.hadoop.fs.tosfs.util.Iterables\r\n[ERROR] -> [Help 1]\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.10.1:compile (default-compile) on project hadoop-azure: Compilation failure: Compilation failure: \r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-tools/hadoop-azure/src/main/java/org/apache/hadoop/fs/azurebfs/services/AbfsLease.java:[38,79] package org.apache.hadoop.thirdparty.org.checkerframework.checker.nullness.qual does not exist\r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-tools/hadoop-azure/src/main/java/org/apache/hadoop/fs/azurebfs/services/AbfsLease.java:[183,30] cannot find symbol\r\n[ERROR] symbol: class Nullable\r\n[ERROR] -> [Help 1]\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.10.1:compile (default-compile) on project hadoop-hdfs-rbf: Compilation failure: Compilation failure: \r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-hdfs-project/hadoop-hdfs-rbf/src/main/java/org/apache/hadoop/hdfs/server/federation/router/RouterRpcServer.java:[102,79] package org.apache.hadoop.thirdparty.org.checkerframework.checker.nullness.qual does not exist\r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-hdfs-project/hadoop-hdfs-rbf/src/main/java/org/apache/hadoop/hdfs/server/federation/router/RouterRpcServer.java:[2509,30] cannot find symbol\r\n[ERROR] symbol: class NonNull\r\n[ERROR] location: class org.apache.hadoop.hdfs.server.federation.router.RouterRpcServer.AsyncThreadFactory\r\n[ERROR] -> [Help 1]\r\n[ERROR] \r\n{code}\r\n","created":"2025-10-20T16:32:20.848+0000"}],"conversations":[{"body":"In the recent build, we encountered the following issue: *org.checkerframework.checker.nullness.qual.NonNull* could not be recognized, and the following error was observed.\r\n{code:java}\r\n[ERROR] /home/jenkins/jenkins-agent/workspace/hadoop-multibranch_PR-8011/ubuntu-focal/src/hadoop-hdfs-project/hadoop-hdfs-rbf/src/main/java/org/apache/hadoop/hdfs/server/federation/router/RouterRpcServer.java:[216,50] package org.checkerframework.checker.nullness.qual does not exist\r\n[ERROR] /home/jenkins/jenkins-agent/workspace/hadoop-multibranch_PR-8011/ubuntu-focal/src/hadoop-hdfs-project/hadoop-hdfs-rbf/src/main/java/org/apache/hadoop/hdfs/server/federation/router/RouterRpcServer.java:[2509,30] cannot find symbol\r\n[ERROR] symbol: class NonNull\r\n[ERROR] location: class org.apache.hadoop.hdfs.server.federation.router.RouterRpcServer.AsyncThreadFactory\r\n {code}\r\nI checked the usage in the related modules, and we should use *org.apache.hadoop.thirdparty.org.checkerframework.checker.nullness.qual.NonNull* instead of directly using {*}org.checkerframework.checker.nullness.qual.NonNull{*}.\r\n\r\n ","from":"reporter","subject":"Resolve build error caused by missing Checker Framework (NonNull not recognized)"},{"body":"this is complicating the new thirdparty release FWIW. this should all be using the unshaded javax. Nullable/nonnull. And the hadoop-thirdparty release needs to address this stuff getting left out so 1.5.0 can be a drop-in replacement for 1.4.0\r\n{code}\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.10.1:compile (default-compile) on project hadoop-tos: Compilation failure: Compilation failure: \r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-cloud-storage-project/hadoop-tos/src/main/java/org/apache/hadoop/fs/tosfs/util/Iterables.java:[22,79] package org.apache.hadoop.thirdparty.org.checkerframework.checker.nullness.qual does not exist\r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-cloud-storage-project/hadoop-tos/src/main/java/org/apache/hadoop/fs/tosfs/util/Iterables.java:[89,29] cannot find symbol\r\n[ERROR] symbol: class Nullable\r\n[ERROR] location: class org.apache.hadoop.fs.tosfs.util.Iterables\r\n[ERROR] -> [Help 1]\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.10.1:compile (default-compile) on project hadoop-azure: Compilation failure: Compilation failure: \r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-tools/hadoop-azure/src/main/java/org/apache/hadoop/fs/azurebfs/services/AbfsLease.java:[38,79] package org.apache.hadoop.thirdparty.org.checkerframework.checker.nullness.qual does not exist\r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-tools/hadoop-azure/src/main/java/org/apache/hadoop/fs/azurebfs/services/AbfsLease.java:[183,30] cannot find symbol\r\n[ERROR] symbol: class Nullable\r\n[ERROR] -> [Help 1]\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.10.1:compile (default-compile) on project hadoop-hdfs-rbf: Compilation failure: Compilation failure: \r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-hdfs-project/hadoop-hdfs-rbf/src/main/java/org/apache/hadoop/hdfs/server/federation/router/RouterRpcServer.java:[102,79] package org.apache.hadoop.thirdparty.org.checkerframework.checker.nullness.qual does not exist\r\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-hdfs-project/hadoop-hdfs-rbf/src/main/java/org/apache/hadoop/hdfs/server/federation/router/RouterRpcServer.java:[2509,30] cannot find symbol\r\n[ERROR] symbol: class NonNull\r\n[ERROR] location: class org.apache.hadoop.hdfs.server.federation.router.RouterRpcServer.AsyncThreadFactory\r\n[ERROR] -> [Help 1]\r\n[ERROR] \r\n{code}\r\n","from":"developer"}],"created":"2025-10-07T03:49:55.000+0000","description":"In the recent build, we encountered the following issue: *org.checkerframework.checker.nullness.qual.NonNull* could not be recognized, and the following error was observed.\r\n{code:java}\r\n[ERROR] /home/jenkins/jenkins-agent/workspace/hadoop-multibranch_PR-8011/ubuntu-focal/src/hadoop-hdfs-project/hadoop-hdfs-rbf/src/main/java/org/apache/hadoop/hdfs/server/federation/router/RouterRpcServer.java:[216,50] package org.checkerframework.checker.nullness.qual does not exist\r\n[ERROR] /home/jenkins/jenkins-agent/workspace/hadoop-multibranch_PR-8011/ubuntu-focal/src/hadoop-hdfs-project/hadoop-hdfs-rbf/src/main/java/org/apache/hadoop/hdfs/server/federation/router/RouterRpcServer.java:[2509,30] cannot find symbol\r\n[ERROR] symbol: class NonNull\r\n[ERROR] location: class org.apache.hadoop.hdfs.server.federation.router.RouterRpcServer.AsyncThreadFactory\r\n {code}\r\nI checked the usage in the related modules, and we should use *org.apache.hadoop.thirdparty.org.checkerframework.checker.nullness.qual.NonNull* instead of directly using {*}org.checkerframework.checker.nullness.qual.NonNull{*}.\r\n\r\n ","issue_id":"13630839","key":"HADOOP-19717","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-10-07T15:59:30.000+0000","role":"fixed_distractor","summary":"Resolve build error caused by missing Checker Framework (NonNull not recognized)"} {"case_id":"13634520","cluster":"DISTRACTOR-HADOOP-19744","comments":[{"body":"I think it's better to just assume that SecurityManager is disabled, and get rid of the extra logic for JDK 22-23.\r\nThere is a small chance that we lose the optimization, but 22-23 is not a version we expect anyone to use for production anway.","created":"2025-11-18T08:12:32.862+0000"},{"body":"I've got code in cloudstore to turn jdk logging up to max for debugging.\r\n\r\n{code}\r\n protected void enableJvmLogging() {\r\n println(\"Enabling JVM logging\");\r\n ConsoleHandler handler = new ConsoleHandler();\r\n handler.setLevel(ALL);\r\n java.util.logging.Logger log = LogManager.getLogManager().getLogger(\"\");\r\n log.addHandler(handler);\r\n log.setLevel(ALL);\r\n }\r\n{code}\r\nMaybe we just do something here for the specific log, within the try/finally block?\r\n\r\nthat way: no need to do variants for different releases, and we have the JRE silent everywhere\r\n","created":"2025-11-18T10:48:07.817+0000"},{"body":"That could be an option, but I really don't think we need to care that much about a JDK22/23 specific issue.\r\n\r\nThe warning happens when we detect what mode Java 22/23 is in so that we can get handle the Java22-23 + enabled SecurityManager case, and is absolutely not needed for other (17-21, 24- ) JVM versions.\r\n\r\nIn the unlikely case that someone both explcitly enables the SecurityManager and is running Hadoop JDK22/23, we will be making some avoidable calls on new thread creation, which we always must do on Java24+ anyway.\r\n","created":"2025-11-18T11:03:01.527+0000"}],"conversations":[{"body":"While the code works fine, it causes Hadoop to emit warnings \r\n\r\n{noformat}\r\nWARNING: A terminally deprecated method in java.lang.System has been called\r\nWARNING: System::setSecurityManager has been called by org.apache.hadoop.security.authentication.util.SubjectUtil (file:/Users/stevel/.m2/repository/org/apache/hadoop/hadoop-auth/3.5.0-SNAPSHOT/hadoop-auth-3.5.0-SNAPSHOT.jar)\r\nWARNING: Please consider reporting this to the maintainers of org.apache.hadoop.security.authentication.util.SubjectUtil\r\n{noformat}\r\n\r\nThe SecurityManager check is only really needed for Java 22-23. \r\n\r\nWe can shortcut the check by returning true for JDK21 and earlier and false for JDK24 and later, and only check and emit warnings for 22-23, which are EOL anyway.\r\n\r\nWe can also try to duplicate the logic the JVM uses to determine whether SecurityManager is enabled without explicitly calling it, but that would be less robust than the current check.\r\n","from":"reporter","subject":"[JDK24] Do not use SecurityManager in SubjectUtil.checkThreadInheritsSubject"},{"body":"I think it's better to just assume that SecurityManager is disabled, and get rid of the extra logic for JDK 22-23.\r\nThere is a small chance that we lose the optimization, but 22-23 is not a version we expect anyone to use for production anway.","from":"developer"},{"body":"I've got code in cloudstore to turn jdk logging up to max for debugging.\r\n\r\n{code}\r\n protected void enableJvmLogging() {\r\n println(\"Enabling JVM logging\");\r\n ConsoleHandler handler = new ConsoleHandler();\r\n handler.setLevel(ALL);\r\n java.util.logging.Logger log = LogManager.getLogManager().getLogger(\"\");\r\n log.addHandler(handler);\r\n log.setLevel(ALL);\r\n }\r\n{code}\r\nMaybe we just do something here for the specific log, within the try/finally block?\r\n\r\nthat way: no need to do variants for different releases, and we have the JRE silent everywhere\r\n","from":"developer"},{"body":"That could be an option, but I really don't think we need to care that much about a JDK22/23 specific issue.\r\n\r\nThe warning happens when we detect what mode Java 22/23 is in so that we can get handle the Java22-23 + enabled SecurityManager case, and is absolutely not needed for other (17-21, 24- ) JVM versions.\r\n\r\nIn the unlikely case that someone both explcitly enables the SecurityManager and is running Hadoop JDK22/23, we will be making some avoidable calls on new thread creation, which we always must do on Java24+ anyway.\r\n","from":"developer"}],"created":"2025-11-18T06:07:59.000+0000","description":"While the code works fine, it causes Hadoop to emit warnings \r\n\r\n{noformat}\r\nWARNING: A terminally deprecated method in java.lang.System has been called\r\nWARNING: System::setSecurityManager has been called by org.apache.hadoop.security.authentication.util.SubjectUtil (file:/Users/stevel/.m2/repository/org/apache/hadoop/hadoop-auth/3.5.0-SNAPSHOT/hadoop-auth-3.5.0-SNAPSHOT.jar)\r\nWARNING: Please consider reporting this to the maintainers of org.apache.hadoop.security.authentication.util.SubjectUtil\r\n{noformat}\r\n\r\nThe SecurityManager check is only really needed for Java 22-23. \r\n\r\nWe can shortcut the check by returning true for JDK21 and earlier and false for JDK24 and later, and only check and emit warnings for 22-23, which are EOL anyway.\r\n\r\nWe can also try to duplicate the logic the JVM uses to determine whether SecurityManager is enabled without explicitly calling it, but that would be less robust than the current check.\r\n","issue_id":"13634520","key":"HADOOP-19744","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-12-15T13:26:43.000+0000","role":"fixed_distractor","summary":"[JDK24] Do not use SecurityManager in SubjectUtil.checkThreadInheritsSubject"} {"case_id":"13636761","cluster":"DISTRACTOR-HADOOP-19755","comments":[{"body":"The build that uses start-build-env.sh appears to source and target JDK 8. I am not 100% certain if this is an environment issue or not.","created":"2025-12-14T00:32:30.393+0000"},{"body":"Thanks for reporting this.\r\nIdeally this should not be the case as all the commits are verified for build before merge.\r\n\r\nStill we are looking into this and update here soon.","created":"2025-12-15T03:17:24.799+0000"},{"body":"Hi [~sjlee0] \r\nIt turned out that a recent metric related change used an API which was not available in jdk 8 hence leading to build failure in jdk 8.\r\nHere is the PR fixing the issue: [https://github.com/apache/hadoop/pull/8135] along with other relevant metric changes\r\n\r\nWe have verified this is building fine in jdk 8 as well now.","created":"2025-12-16T14:20:20.542+0000"},{"body":"Hi [~slfan1989] \r\nI know we have upgraded trunk branch to use JDK17.\r\nSeems like yetus has also started doing validations using JDK 17.\r\nDoes that mean we won't be supporting future hadoop releases on JDK 8?","created":"2025-12-18T06:02:25.608+0000"},{"body":"Then is it the build script and/or the POM that needs to be updated?","created":"2025-12-18T15:26:30.620+0000"}],"conversations":[{"body":"The hadoop build on trunk is failing in hadoop-azure.\r\n\r\nSteps to reproduce:\r\n{code}\r\n./start-build-env.sh\r\n(on container shell) mvn clean install -DskipTests -DskipShade\r\n{code}\r\n\r\nError:\r\n{code}\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.10.1:compile (default-compile) on project hadoop-azure: Compilation failure\r\n[ERROR] /home/sjlee0_gmail_com/hadoop/hadoop-tools/hadoop-azure/src/main/java/org/apache/hadoop/fs/azurebfs/utils/ResourceUtilizationUtils.java:[152,12] cannot find symbol\r\n[ERROR]   symbol:   variable ProcessHandle\r\n[ERROR]   location: class org.apache.hadoop.fs.azurebfs.utils.ResourceUtilizationUtils\r\n{code}","from":"reporter","subject":"hadoop-azure build fails"},{"body":"The build that uses start-build-env.sh appears to source and target JDK 8. I am not 100% certain if this is an environment issue or not.","from":"developer"},{"body":"Thanks for reporting this.\r\nIdeally this should not be the case as all the commits are verified for build before merge.\r\n\r\nStill we are looking into this and update here soon.","from":"developer"},{"body":"Hi [~sjlee0] \r\nIt turned out that a recent metric related change used an API which was not available in jdk 8 hence leading to build failure in jdk 8.\r\nHere is the PR fixing the issue: [https://github.com/apache/hadoop/pull/8135] along with other relevant metric changes\r\n\r\nWe have verified this is building fine in jdk 8 as well now.","from":"developer"},{"body":"Hi [~slfan1989] \r\nI know we have upgraded trunk branch to use JDK17.\r\nSeems like yetus has also started doing validations using JDK 17.\r\nDoes that mean we won't be supporting future hadoop releases on JDK 8?","from":"developer"},{"body":"Then is it the build script and/or the POM that needs to be updated?","from":"developer"}],"created":"2025-12-14T00:28:49.000+0000","description":"The hadoop build on trunk is failing in hadoop-azure.\r\n\r\nSteps to reproduce:\r\n{code}\r\n./start-build-env.sh\r\n(on container shell) mvn clean install -DskipTests -DskipShade\r\n{code}\r\n\r\nError:\r\n{code}\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.10.1:compile (default-compile) on project hadoop-azure: Compilation failure\r\n[ERROR] /home/sjlee0_gmail_com/hadoop/hadoop-tools/hadoop-azure/src/main/java/org/apache/hadoop/fs/azurebfs/utils/ResourceUtilizationUtils.java:[152,12] cannot find symbol\r\n[ERROR]   symbol:   variable ProcessHandle\r\n[ERROR]   location: class org.apache.hadoop.fs.azurebfs.utils.ResourceUtilizationUtils\r\n{code}","issue_id":"13636761","key":"HADOOP-19755","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2025-12-30T14:47:27.000+0000","role":"fixed_distractor","summary":"hadoop-azure build fails"} {"case_id":"13639805","cluster":"DISTRACTOR-HADOOP-19790","comments":[{"body":"Similar issue has been reported by other developers as well.\r\nWe are looking into the issue, If someone has any leads that will be helpful.\r\n\r\n \r\n\r\nCC: [~slfan1989] [~cnauroth] [~stevel@apache.org] ","created":"2026-01-23T07:19:32.227+0000"},{"body":"[~bhattmanish98] observed that fetching protobuf-maven-plugin from io.github.ascopes package is causing issue.\r\nIf fetchind from org.xolstice.maven.plugins as in all other places in hadoop repo fixes this issue.","created":"2026-01-23T13:26:52.322+0000"},{"body":"Hello [~anujmodi]. I'm unable to reproduce this problem.\r\n\r\nRegarding protobuf-maven-plugin from io.github.ascopes, there were these recent changes from YARN-11921:\r\n\r\nhttps://github.com/apache/hadoop/pull/8156\r\nhttps://github.com/apache/hadoop/pull/8192\r\n\r\nIs there anything earlier in the output to suggest a root cause, like maybe a failure to fully download the plugin in your local environment?\r\n\r\nCC: [~appodictic]","created":"2026-01-23T20:16:45.643+0000"},{"body":"My gut says it is s subtle maven issue.\r\n\r\n\r\n{code:java}\r\nedward@fedora:~/edgy-ansible/imaging/hadoop$ mvn -version\r\nApache Maven 3.9.9 (Red Hat 3.9.9-14)\r\nMaven home: /usr/share/maven\r\nJava version: 21.0.9, vendor: Red Hat, Inc., runtime: /usr/lib/jvm/java-21-openjdk\r\nDefault locale: en_US, platform encoding: UTF-8\r\nOS name: \"linux\", version: \"6.17.12-200.fc42.x86_64\", arch: \"amd64\", family: \"unix\"\r\n {code}","created":"2026-01-23T23:52:33.523+0000"},{"body":"{code:java}\r\nApache Maven 3.9.9 (8e8579a9e76f7d015ee5ec7bfcdc97d260186937)\r\nMaven home: /usr/share/java/maven-3\r\nJava version: 21.0.9, vendor: Alpine, runtime: /usr/lib/jvm/java-21-openjdk\r\nDefault locale: en_US, platform encoding: UTF-8\r\nOS name: \"linux\", version: \"6.17.12-200.fc42.x86_64\", arch: \"amd64\", family: \"unix\"\r\n {code}\r\nBoth of these environments work for me.","created":"2026-01-23T23:53:51.536+0000"},{"body":"[~cnauroth] \r\n\r\nhttps://github.com/ascopes/protobuf-maven-plugin/issues/853","created":"2026-01-23T23:56:05.482+0000"},{"body":"Thanks [~appodictic] [~cnauroth] \r\n\r\nSeems like it was indeed a maven issue. Upgrading to maven 3.9.12 LTS helped.\r\nBut does that mean we need to update Building.txt to refelect this requirement?","created":"2026-01-24T06:31:18.314+0000"},{"body":"there's a way to declare the minimum maven version in the project pom, currently its (3.3.0, ). If we fixed to something high (I'm on 3.9.11 FWIW) then it'd force upgrades but this is probably good all round\r\n\r\n(maven and other build tools are things I don't let homebrew control, FWIW. It's not a real package manager)","created":"2026-01-24T12:53:57.233+0000"},{"body":"[~anujmodi] I wasn't aware the plugin wasn't portable till unlder maven until this thread. I read the discussion in the ticket I reference plugin versions + maven versions. I was not clear on that. Right now we are not specifing a version of the plugin, however you can experiment with different numbers:\r\n[https://github.com/ascopes/protobuf-maven-plugin/tags]\r\n\r\nIt wouldn't be bad if we found a versions that works with older maven.\r\n\r\nIn any case, sorry to mess up your workflow and cause you to have to update. My motivation for doing this is that the xolstice plugin is EOL and the entire build doesn't work on alpine (alpine doesnt ship protobuf 2.5.0 since it is vulnerable), so my goal was to make the build more reliable in more environments. But your the broken egg in the omelet it seems. ","created":"2026-01-24T19:56:39.766+0000"},{"body":"bq. It wouldn't be bad if we found a versions that works with older maven.\r\n\r\nor are ruthless and set a recent mvn version as people need to upgrade for the build to work.","created":"2026-01-27T20:34:37.229+0000"},{"body":"Personally I feel with many upgrades going in 3.5.0 including junit, jdk. We can very well upgrade the required mvn version.","created":"2026-01-28T04:53:57.947+0000"},{"body":"bq. with many upgrades going in 3.5.0 including junit, jdk. We can very well upgrade the required mvn version.\r\n\r\n+1. \r\n\r\nsubmit a pr changing the enforced.maven.version property in hadoop-project pom","created":"2026-01-28T13:52:02.908+0000"},{"body":"[~stevel@apache.org] \r\n{code:java}\r\nsubmit a pr changing the enforced.maven.version property in hadoop-project pom {code}\r\nI have a PR ready for review for this.","created":"2026-01-29T18:03:28.795+0000"},{"body":"I think I found the root cause of one of the failing test classes on trunk (TestRMWebServicesReservation)\r\n\r\nopening discussion and a pursuing a patch in https://issues.apache.org/jira/browse/YARN-11926 ","created":"2026-02-05T21:47:04.374+0000"},{"body":"Reproduced in CI [here|https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8225/3/artifact/out/branch-mvninstall-root.txt]. I'm also seeing this in my local hadoop-build container.","created":"2026-02-18T02:32:38.260+0000"}],"conversations":[{"body":"While trying to build at hadoop level, build is failing on trunk.\r\n\r\nCommit: [https://github.com/apache/hadoop/commit/b3ce394e5aba3cdaf326b87e7cfb6be898ddd48d]\r\n\r\nError:\r\n[ERROR] Failed to execute goal io.github.ascopes:protobuf-maven-plugin:1.2.0:generate (default) on project hadoop-yarn-csi: Execution default of goal io.github.ascopes:protobuf-maven-plugin:1.2.0:generate failed: Unable to load the mojo 'generate' (or one of its required components) from the plugin 'io.github.ascopes:protobuf-maven-plugin:1.2.0': com.google.inject.ProvisionException: Unable to provision, see the following errors:\r\n[ERROR] \r\n[ERROR] 1) No implementation for io.github.ascopes.protobufmavenplugin.generate.SourceCodeGenerator was bound.\r\n[ERROR]   while locating io.github.ascopes.protobufmavenplugin.MainGenerateMojo\r\n[ERROR]   at ClassRealm[plugin>io.github.ascopes:protobuf-maven-plugin:1.2.0, parent: jdk.internal.loader.ClassLoaders$AppClassLoader@30946e09] (via modules: org.eclipse.sisu.wire.WireModule -> org.eclipse.sisu.plexus.PlexusBindingModule)\r\n[ERROR]   while locating org.apache.maven.plugin.Mojo annotated with @com.google.inject.name.Named(value=\"io.github.ascopes:protobuf-maven-plugin:1.2.0:generate\")\r\n[ERROR] \r\n[ERROR] 1 error\r\n[ERROR]       role: org.apache.maven.plugin.Mojo\r\n[ERROR]   roleHint: io.github.ascopes:protobuf-maven-plugin:1.2.0:generate","from":"reporter","subject":"Set Minimum Required Maven Version to 3.9.11"},{"body":"Similar issue has been reported by other developers as well.\r\nWe are looking into the issue, If someone has any leads that will be helpful.\r\n\r\n \r\n\r\nCC: [~slfan1989] [~cnauroth] [~stevel@apache.org] ","from":"developer"},{"body":"[~bhattmanish98] observed that fetching protobuf-maven-plugin from io.github.ascopes package is causing issue.\r\nIf fetchind from org.xolstice.maven.plugins as in all other places in hadoop repo fixes this issue.","from":"developer"},{"body":"Hello [~anujmodi]. I'm unable to reproduce this problem.\r\n\r\nRegarding protobuf-maven-plugin from io.github.ascopes, there were these recent changes from YARN-11921:\r\n\r\nhttps://github.com/apache/hadoop/pull/8156\r\nhttps://github.com/apache/hadoop/pull/8192\r\n\r\nIs there anything earlier in the output to suggest a root cause, like maybe a failure to fully download the plugin in your local environment?\r\n\r\nCC: [~appodictic]","from":"developer"},{"body":"My gut says it is s subtle maven issue.\r\n\r\n\r\n{code:java}\r\nedward@fedora:~/edgy-ansible/imaging/hadoop$ mvn -version\r\nApache Maven 3.9.9 (Red Hat 3.9.9-14)\r\nMaven home: /usr/share/maven\r\nJava version: 21.0.9, vendor: Red Hat, Inc., runtime: /usr/lib/jvm/java-21-openjdk\r\nDefault locale: en_US, platform encoding: UTF-8\r\nOS name: \"linux\", version: \"6.17.12-200.fc42.x86_64\", arch: \"amd64\", family: \"unix\"\r\n {code}","from":"developer"},{"body":"{code:java}\r\nApache Maven 3.9.9 (8e8579a9e76f7d015ee5ec7bfcdc97d260186937)\r\nMaven home: /usr/share/java/maven-3\r\nJava version: 21.0.9, vendor: Alpine, runtime: /usr/lib/jvm/java-21-openjdk\r\nDefault locale: en_US, platform encoding: UTF-8\r\nOS name: \"linux\", version: \"6.17.12-200.fc42.x86_64\", arch: \"amd64\", family: \"unix\"\r\n {code}\r\nBoth of these environments work for me.","from":"developer"},{"body":"[~cnauroth] \r\n\r\nhttps://github.com/ascopes/protobuf-maven-plugin/issues/853","from":"developer"},{"body":"Thanks [~appodictic] [~cnauroth] \r\n\r\nSeems like it was indeed a maven issue. Upgrading to maven 3.9.12 LTS helped.\r\nBut does that mean we need to update Building.txt to refelect this requirement?","from":"developer"},{"body":"there's a way to declare the minimum maven version in the project pom, currently its (3.3.0, ). If we fixed to something high (I'm on 3.9.11 FWIW) then it'd force upgrades but this is probably good all round\r\n\r\n(maven and other build tools are things I don't let homebrew control, FWIW. It's not a real package manager)","from":"developer"},{"body":"[~anujmodi] I wasn't aware the plugin wasn't portable till unlder maven until this thread. I read the discussion in the ticket I reference plugin versions + maven versions. I was not clear on that. Right now we are not specifing a version of the plugin, however you can experiment with different numbers:\r\n[https://github.com/ascopes/protobuf-maven-plugin/tags]\r\n\r\nIt wouldn't be bad if we found a versions that works with older maven.\r\n\r\nIn any case, sorry to mess up your workflow and cause you to have to update. My motivation for doing this is that the xolstice plugin is EOL and the entire build doesn't work on alpine (alpine doesnt ship protobuf 2.5.0 since it is vulnerable), so my goal was to make the build more reliable in more environments. But your the broken egg in the omelet it seems. ","from":"developer"},{"body":"bq. It wouldn't be bad if we found a versions that works with older maven.\r\n\r\nor are ruthless and set a recent mvn version as people need to upgrade for the build to work.","from":"developer"},{"body":"Personally I feel with many upgrades going in 3.5.0 including junit, jdk. We can very well upgrade the required mvn version.","from":"developer"},{"body":"bq. with many upgrades going in 3.5.0 including junit, jdk. We can very well upgrade the required mvn version.\r\n\r\n+1. \r\n\r\nsubmit a pr changing the enforced.maven.version property in hadoop-project pom","from":"developer"},{"body":"[~stevel@apache.org] \r\n{code:java}\r\nsubmit a pr changing the enforced.maven.version property in hadoop-project pom {code}\r\nI have a PR ready for review for this.","from":"developer"},{"body":"I think I found the root cause of one of the failing test classes on trunk (TestRMWebServicesReservation)\r\n\r\nopening discussion and a pursuing a patch in https://issues.apache.org/jira/browse/YARN-11926 ","from":"developer"},{"body":"Reproduced in CI [here|https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8225/3/artifact/out/branch-mvninstall-root.txt]. I'm also seeing this in my local hadoop-build container.","from":"developer"}],"created":"2026-01-23T07:07:44.000+0000","description":"While trying to build at hadoop level, build is failing on trunk.\r\n\r\nCommit: [https://github.com/apache/hadoop/commit/b3ce394e5aba3cdaf326b87e7cfb6be898ddd48d]\r\n\r\nError:\r\n[ERROR] Failed to execute goal io.github.ascopes:protobuf-maven-plugin:1.2.0:generate (default) on project hadoop-yarn-csi: Execution default of goal io.github.ascopes:protobuf-maven-plugin:1.2.0:generate failed: Unable to load the mojo 'generate' (or one of its required components) from the plugin 'io.github.ascopes:protobuf-maven-plugin:1.2.0': com.google.inject.ProvisionException: Unable to provision, see the following errors:\r\n[ERROR] \r\n[ERROR] 1) No implementation for io.github.ascopes.protobufmavenplugin.generate.SourceCodeGenerator was bound.\r\n[ERROR]   while locating io.github.ascopes.protobufmavenplugin.MainGenerateMojo\r\n[ERROR]   at ClassRealm[plugin>io.github.ascopes:protobuf-maven-plugin:1.2.0, parent: jdk.internal.loader.ClassLoaders$AppClassLoader@30946e09] (via modules: org.eclipse.sisu.wire.WireModule -> org.eclipse.sisu.plexus.PlexusBindingModule)\r\n[ERROR]   while locating org.apache.maven.plugin.Mojo annotated with @com.google.inject.name.Named(value=\"io.github.ascopes:protobuf-maven-plugin:1.2.0:generate\")\r\n[ERROR] \r\n[ERROR] 1 error\r\n[ERROR]       role: org.apache.maven.plugin.Mojo\r\n[ERROR]   roleHint: io.github.ascopes:protobuf-maven-plugin:1.2.0:generate","issue_id":"13639805","key":"HADOOP-19790","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2026-03-23T10:07:39.000+0000","role":"fixed_distractor","summary":"Set Minimum Required Maven Version to 3.9.11"} {"case_id":"12379455","cluster":"DISTRACTOR-HADOOP-1980","comments":[{"body":"These two are sort of related. When HDFS is started up, it should be possible to select the mode that HDFS comes up in.","created":"2007-11-19T23:34:30.369+0000"},{"body":"Suggested patch for trunk.\n\nIf admin puts NN in safemode during startup, NN will stay in safemode even after block ratios are satisfied. It also prints block ratios for convenience (just like the default case where admin does not enter safemode manually).\n","created":"2008-11-19T00:18:58.315+0000"},{"body":"The requirement here as I understand it is to make it possible to extend the safe mode indefinitely.\nFirst thing to do is just to start the name-node with a large extension or a >1 threshold. I guess this does not work if the name-node has already been started but the administrator needs to keep it in safe mode longer than the extension.\nThen why don't we just set really long extension in this case.\nOr even better provide an explicit option to extend safe mode from the admin command.\n{code}\nhadoop dfsadmin -safemode extend\n{code}\nThis will be much simpler patch, which would satisfy all the requirements and will have clear user api.\nOtherwise it seems rather strange from a user point of view: in order to get infinite safe mode one should enter safe mode again although it is already on.","created":"2008-12-16T02:27:58.493+0000"},{"body":"This does not preclude from setting a very large extension. \n\nThis patch only brings consistent behavior to '-safemode enter'. Once you enter safemode manually, it makes good sense for it to require a manual command to leave. It should not matter weather it is already in safemode or not.\n","created":"2008-12-16T02:44:30.968+0000"},{"body":"Updated patch with a unit test.\n\nIn some sense this is more of a bugfix rather than a new feature. It mainly aims to correct the the meaning of '-safemode enter'. Additionally normal information about % of blocks reported by datanodes is displayed to help administrators.","created":"2008-12-16T07:06:16.284+0000"},{"body":"Does not look like there is a big enthusiasm to make changes to the shell api.\nI simplified a bit Raghu's implementation of the feature.\nRenamed and tweaked the test so that it shutdowned the cluster in the final section if anything fails.","created":"2008-12-19T03:34:52.002+0000"},{"body":"+1\npatch looks good\n","created":"2008-12-19T23:54:05.243+0000"},{"body":"This is patch for 0.18 branch.","created":"2008-12-20T02:21:37.273+0000"},{"body":"I just committed this. Thank you Raghu.","created":"2008-12-20T03:13:14.806+0000"}],"conversations":[{"body":"When debugging, I'd like to be able to intentionally keep the FS in a safemode. (For example, when looking at HADOOP-1978).\nAlso, it'll be nice if the namenode can still update the webUI/report when it hits the dfs.safemode.threshold.pct.","from":"reporter","subject":"'dfsadmin -safemode enter' should prevent the namenode from leaving safemode automatically after startup"},{"body":"These two are sort of related. When HDFS is started up, it should be possible to select the mode that HDFS comes up in.","from":"developer"},{"body":"Suggested patch for trunk.\n\nIf admin puts NN in safemode during startup, NN will stay in safemode even after block ratios are satisfied. It also prints block ratios for convenience (just like the default case where admin does not enter safemode manually).\n","from":"developer"},{"body":"The requirement here as I understand it is to make it possible to extend the safe mode indefinitely.\nFirst thing to do is just to start the name-node with a large extension or a >1 threshold. I guess this does not work if the name-node has already been started but the administrator needs to keep it in safe mode longer than the extension.\nThen why don't we just set really long extension in this case.\nOr even better provide an explicit option to extend safe mode from the admin command.\n{code}\nhadoop dfsadmin -safemode extend\n{code}\nThis will be much simpler patch, which would satisfy all the requirements and will have clear user api.\nOtherwise it seems rather strange from a user point of view: in order to get infinite safe mode one should enter safe mode again although it is already on.","from":"developer"},{"body":"This does not preclude from setting a very large extension. \n\nThis patch only brings consistent behavior to '-safemode enter'. Once you enter safemode manually, it makes good sense for it to require a manual command to leave. It should not matter weather it is already in safemode or not.\n","from":"developer"},{"body":"Updated patch with a unit test.\n\nIn some sense this is more of a bugfix rather than a new feature. It mainly aims to correct the the meaning of '-safemode enter'. Additionally normal information about % of blocks reported by datanodes is displayed to help administrators.","from":"developer"},{"body":"Does not look like there is a big enthusiasm to make changes to the shell api.\nI simplified a bit Raghu's implementation of the feature.\nRenamed and tweaked the test so that it shutdowned the cluster in the final section if anything fails.","from":"developer"},{"body":"+1\npatch looks good\n","from":"developer"},{"body":"This is patch for 0.18 branch.","from":"developer"},{"body":"I just committed this. Thank you Raghu.","from":"developer"}],"created":"2007-10-02T00:59:06.000+0000","description":"When debugging, I'd like to be able to intentionally keep the FS in a safemode. (For example, when looking at HADOOP-1978).\nAlso, it'll be nice if the namenode can still update the webUI/report when it hits the dfs.safemode.threshold.pct.","issue_id":"12379455","key":"HADOOP-1980","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-12-20T03:13:14.000+0000","role":"fixed_distractor","summary":"'dfsadmin -safemode enter' should prevent the namenode from leaving safemode automatically after startup"} {"case_id":"13648295","cluster":"DISTRACTOR-HADOOP-19863","comments":[{"body":"Peter; fixed in trunk and backports in progress for 3.4 and 3.5 releases","created":"2026-05-13T18:30:32.504+0000"},{"body":"Thank you Steve.","created":"2026-05-13T18:32:30.961+0000"}],"conversations":[{"body":"As discussed in [https://github.com/apache/parquet-java/issues/2703#issuecomment-4260121705] we noticed that when vectoried IO is enabled the {{BytesRead}} metrics of Spark tasks are not correct.\r\n\r\nSpark fetches that metric via {{FileSystem.getAllStatistics}} see\r\n - [https://github.com/apache/spark/blob/5d491f62748b4b9c34bc3b5bd7390f7b5ca75053/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/FileScanRDD.scala#L98-L109] and\r\n - [https://github.com/apache/spark/blob/5d491f62748b4b9c34bc3b5bd7390f7b5ca75053/core/src/main/scala/org/apache/spark/deploy/SparkHadoopUtil.scala#L164-L170]\r\n\r\nRepro with latest Spark 4.2.0-SNAPSHOT using Hadoop 3.5.0:\r\nVectored IO is enabled by default:\r\n{code:java}\r\n➜ bin/spark-shell\r\n\r\nscala> spark.createDataFrame((0 until 5000).map(i => (i, s\"left_$i\"))).repartition(1).write.parquet(\"/tmp/t2\")\r\nscala> spark.read.parquet(\"/tmp/t2\").createOrReplaceTempView(\"t2\")\r\nscala> sql(\"SELECT * FROM t2\").collect()\r\n{code}\r\n!Screenshot 2026-04-16 at 19.02.30.png|width=85%!\r\n\r\nVectored IO is disabled explicitely:\r\n{code:java}\r\n➜ bin/spark-shell --conf spark.hadoop.parquet.hadoop.vectored.io.enabled=false\r\n\r\nscala> spark.read.parquet(\"/tmp/t2\").createOrReplaceTempView(\"t2\")\r\nscala> sql(\"SELECT * FROM t2\").collect()\r\n{code}\r\n!Screenshot 2026-04-16 at 19.03.51.png|width=85%!\r\n\r\nIn my case the generated test file size was ~45KB:\r\n{code:java}\r\n➜ ls -ll /tmp/t2\r\ntotal 88\r\n-rw-r--r--@ 1 ptoth wheel 0 Apr 16 18:57 _SUCCESS\r\n-rw-r--r--@ 1 ptoth wheel 44944 Apr 16 18:57 part-00000-cf825cf6-2fa5-46a2-b897-dbb9dc9828a7-c000.snappy.parquet{code}\r\nI believe reading the parquet footers don't go through vectored IO so the decreased 1680B probably belongs to that.\r\n\r\nThere is no data pruning in the query so the metric value should be around the file size.","from":"reporter","subject":"Incorrect Vectored IO metrics from Local Filesystem"},{"body":"Peter; fixed in trunk and backports in progress for 3.4 and 3.5 releases","from":"developer"},{"body":"Thank you Steve.","from":"developer"}],"created":"2026-04-16T17:13:51.000+0000","description":"As discussed in [https://github.com/apache/parquet-java/issues/2703#issuecomment-4260121705] we noticed that when vectoried IO is enabled the {{BytesRead}} metrics of Spark tasks are not correct.\r\n\r\nSpark fetches that metric via {{FileSystem.getAllStatistics}} see\r\n - [https://github.com/apache/spark/blob/5d491f62748b4b9c34bc3b5bd7390f7b5ca75053/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/FileScanRDD.scala#L98-L109] and\r\n - [https://github.com/apache/spark/blob/5d491f62748b4b9c34bc3b5bd7390f7b5ca75053/core/src/main/scala/org/apache/spark/deploy/SparkHadoopUtil.scala#L164-L170]\r\n\r\nRepro with latest Spark 4.2.0-SNAPSHOT using Hadoop 3.5.0:\r\nVectored IO is enabled by default:\r\n{code:java}\r\n➜ bin/spark-shell\r\n\r\nscala> spark.createDataFrame((0 until 5000).map(i => (i, s\"left_$i\"))).repartition(1).write.parquet(\"/tmp/t2\")\r\nscala> spark.read.parquet(\"/tmp/t2\").createOrReplaceTempView(\"t2\")\r\nscala> sql(\"SELECT * FROM t2\").collect()\r\n{code}\r\n!Screenshot 2026-04-16 at 19.02.30.png|width=85%!\r\n\r\nVectored IO is disabled explicitely:\r\n{code:java}\r\n➜ bin/spark-shell --conf spark.hadoop.parquet.hadoop.vectored.io.enabled=false\r\n\r\nscala> spark.read.parquet(\"/tmp/t2\").createOrReplaceTempView(\"t2\")\r\nscala> sql(\"SELECT * FROM t2\").collect()\r\n{code}\r\n!Screenshot 2026-04-16 at 19.03.51.png|width=85%!\r\n\r\nIn my case the generated test file size was ~45KB:\r\n{code:java}\r\n➜ ls -ll /tmp/t2\r\ntotal 88\r\n-rw-r--r--@ 1 ptoth wheel 0 Apr 16 18:57 _SUCCESS\r\n-rw-r--r--@ 1 ptoth wheel 44944 Apr 16 18:57 part-00000-cf825cf6-2fa5-46a2-b897-dbb9dc9828a7-c000.snappy.parquet{code}\r\nI believe reading the parquet footers don't go through vectored IO so the decreased 1680B probably belongs to that.\r\n\r\nThere is no data pruning in the query so the metric value should be around the file size.","issue_id":"13648295","key":"HADOOP-19863","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2026-05-13T18:16:40.000+0000","role":"fixed_distractor","summary":"Incorrect Vectored IO metrics from Local Filesystem"} {"case_id":"12380386","cluster":"DISTRACTOR-HADOOP-2053","comments":[{"body":"Marking this a blocker since apps that were working with 0.13 release fails with 0.14","created":"2007-10-14T04:27:05.480+0000"},{"body":"If fixing HADOOP-2043 leads us to doing a 0.14.3 release, we should include the fix for this issue in that as well.","created":"2007-10-14T04:29:24.019+0000"},{"body":"Here is a patch which frees the reference to the large {{DataOutputBuffer}} that {{BasicTypeSorterBase}} has in it's {{close}} method... this lets the GC collect away the keyValBuffer. \n\nIn absence of this patch, there is a window where both the currently active keyValBuffer and the one that should have been freed in the previous iteration are both active i.e. doubling the required amount of memory, which leads to the OutOfMemoryException.\n\nAll credit to this goes to Koji!","created":"2007-10-15T10:16:08.996+0000"},{"body":"Previously (hadoop-0.13.0) re-used the keyValBuffer by doing a *reset* on it, this led to some scenarios (HADOOP-875) where the large keyValBuffer could be kept around even after spilling, and subsequent iterations wouldn't need it, thus wasting memory.\n\nNow, we completely release it and use a new buffer in every iteration, however it means we could run into a performance-regression vis-a-vis 0.13.0, but this jira is about fixing the correctness issue.\n\nI've filed HADOOP-2054 to try and capture the creamy parts of both - lets discuss.","created":"2007-10-15T10:48:47.842+0000"},{"body":"I just committed this. Thanks, Arun!","created":"2007-10-15T17:20:06.041+0000"}],"conversations":[{"body":"In recent hadoop 0.14 we are seeing few jobs where map taskf fail with java.lang.OutOfMemoryError: Java heap space problem\nThese were the same jobs which used to work fine with 0.13\n\n\ntask_200710112103_0001_m_000015_1: java.lang.OutOfMemoryError: Java heap space\n\tat java.util.Arrays.copyOf(Arrays.java:2786)\n\tat java.io.ByteArrayOutputStream.write(ByteArrayOutputStream.java:94)\n\tat java.io.DataOutputStream.write(DataOutputStream.java:90)\n\tat org.apache.hadoop.io.Text.write(Text.java:243)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.collect(MapTask.java:340)\n","from":"reporter","subject":"OutOfMemoryError : Java heap space errors in hadoop 0.14"},{"body":"Marking this a blocker since apps that were working with 0.13 release fails with 0.14","from":"developer"},{"body":"If fixing HADOOP-2043 leads us to doing a 0.14.3 release, we should include the fix for this issue in that as well.","from":"developer"},{"body":"Here is a patch which frees the reference to the large {{DataOutputBuffer}} that {{BasicTypeSorterBase}} has in it's {{close}} method... this lets the GC collect away the keyValBuffer. \n\nIn absence of this patch, there is a window where both the currently active keyValBuffer and the one that should have been freed in the previous iteration are both active i.e. doubling the required amount of memory, which leads to the OutOfMemoryException.\n\nAll credit to this goes to Koji!","from":"developer"},{"body":"Previously (hadoop-0.13.0) re-used the keyValBuffer by doing a *reset* on it, this led to some scenarios (HADOOP-875) where the large keyValBuffer could be kept around even after spilling, and subsequent iterations wouldn't need it, thus wasting memory.\n\nNow, we completely release it and use a new buffer in every iteration, however it means we could run into a performance-regression vis-a-vis 0.13.0, but this jira is about fixing the correctness issue.\n\nI've filed HADOOP-2054 to try and capture the creamy parts of both - lets discuss.","from":"developer"},{"body":"I just committed this. Thanks, Arun!","from":"developer"}],"created":"2007-10-13T18:34:58.000+0000","description":"In recent hadoop 0.14 we are seeing few jobs where map taskf fail with java.lang.OutOfMemoryError: Java heap space problem\nThese were the same jobs which used to work fine with 0.13\n\n\ntask_200710112103_0001_m_000015_1: java.lang.OutOfMemoryError: Java heap space\n\tat java.util.Arrays.copyOf(Arrays.java:2786)\n\tat java.io.ByteArrayOutputStream.write(ByteArrayOutputStream.java:94)\n\tat java.io.DataOutputStream.write(DataOutputStream.java:90)\n\tat org.apache.hadoop.io.Text.write(Text.java:243)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.collect(MapTask.java:340)\n","issue_id":"12380386","key":"HADOOP-2053","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2007-10-15T17:20:06.000+0000","role":"fixed_distractor","summary":"OutOfMemoryError : Java heap space errors in hadoop 0.14"} {"case_id":"12381005","cluster":"DISTRACTOR-HADOOP-2095","comments":[{"body":"The problem was gone after I set the compressMapOutput attribute to false.\n","created":"2007-10-23T22:26:02.378+0000"},{"body":"Is it true that native compression was in use when you saw the problem?","created":"2007-10-25T04:26:38.371+0000"},{"body":"Yes.\n","created":"2007-10-25T04:31:05.001+0000"},{"body":"My hunch is this that we are ending up with too many files in the ramfs (the map outputs for the failing reduce are small in size). So when merge is initiated on those files, we end up creating too many codecs for decompressing the files, and since they are native, we encounter OOM. So one way to validate this is to decrease the value of mapred.inmem.merge.threshold (defaults to 1000) to something like 200 and then see whether the OOM goes away. The other alternative is to increase the heap size per task.","created":"2007-10-25T05:41:05.460+0000"},{"body":"FYI, I ran into the exactly same issue with massive failures (nearly all reduces), using a 0 value for mapred.inmem.merge.threshold (letting the framework select the threshold)\n\nusing native compression\n\n1GB of heap space, 1350 nodes\n\nmapred.inmem.merge.threshold\t0\nmapred.reduce.parallel.copies\t20\ntasktracker.http.threads\t30\nmapred.map.tasks\t13008\nmapred.reduce.tasks\t3600\nfs.inmemory.size.mb\t200\nio.seqfile.sorter.recordlimit\t1000000\nio.sort.mb\t200\nio.sort.factor\t300\nmapred.map.output.compression.type\tRECORD\nmapred.map.output.compression.codec\torg.apache.hadoop.io.compress.DefaultCodec\nmapred.compress.map.output\ttrue\n","created":"2007-10-26T20:07:57.212+0000"},{"body":"Christian, please don't use 0 for mapred.inmem.merge.threshold. It basically disables the threshold for initiating merge, and, in this case, merge on the ramfs files gets triggered only when the ramfs has collected enough bytes of data (actually 50% of fs.inmemory.size.mb). Depending on how big each map output is, we will accumulate that many files there. For each file, a native codec will be initialized, and I suspect that if we create too many codecs, we will encounter OOM (given a certain child heap size). So, I think it is worth trying the app with a reduced mapred.inmem.merge.threshold like 50/100, just to ensure that we keep a control on the number of codecs we create. Christian/Runping, could you please try this tweak and let us know? Thanks.","created":"2007-10-27T03:35:33.417+0000"},{"body":"Just to clarify, a codec-pool is maintained in the compression module. So, the number of codecs created will be in the order of mapred.inmem.merge.threshold.","created":"2007-10-27T03:48:06.859+0000"},{"body":"BTW, another thing that needs to be controlled is the io.sort.factor which would control the number of files that gets merged at once post the shuffle (the final merge of the on-disk files). The number of intermediate files opened would be as per the value of io.sort.factor and you might encounter OOM if that is large. So i would recommend that for working around this issue, the io.sort.factor value should be in the same order as the value of mapred.inmem.merge.threshold (that successfully does ramfs merges during the shuffle).","created":"2007-10-28T08:42:36.055+0000"},{"body":"Christian,\n Can you generate a set of task files for the reduce with the failure with keep.failed.task.files set to true, so that we can run it in the isolation runner and a profiler?\n","created":"2007-10-30T18:56:06.568+0000"},{"body":"Noticed that codec pool is not used in o.a.h.i.SequenceFile.CompressedBytes.writeUncompressedBytes. This needs to be fixed.","created":"2007-10-31T05:39:37.372+0000"},{"body":"Christian/Runping - Could you please re-run your jobs using this debug patch? It just logs the no. of created native codecs and the stack trace to check where they are being created from. Thanks!\n\nMeanwhile I'll plug away and see how to fix SequenceFile.CompressedBytes as Devaraj pointed out...","created":"2007-10-31T18:34:33.421+0000"},{"body":"Here is an early patch which fixed {{SequenceFile.CompressedBytes}} to use a {{CodecPool}} passed along by the {{SequenceFile.Reader}}... \n\nChristian/Runping -I'd appreciate if you could try this patch alongwith the previous debug patch while I try to test it at my end. Thanks!","created":"2007-10-31T18:46:34.239+0000"},{"body":"Sorry, previous patch had the leading path wrong...","created":"2007-10-31T18:49:25.882+0000"},{"body":"Sorry, previous patch had the leading path wrong...","created":"2007-10-31T18:50:56.717+0000"},{"body":"I still see failures after shuffling during final sort:\n\njava.lang.OutOfMemoryError: Java heap space\n\tat org.apache.hadoop.io.DataOutputBuffer$Buffer.write(DataOutputBuffer.java:52)\n\tat org.apache.hadoop.io.DataOutputBuffer.write(DataOutputBuffer.java:90)\n\tat org.apache.hadoop.io.SequenceFile$Reader.readBuffer(SequenceFile.java:1535)\n\tat org.apache.hadoop.io.SequenceFile$Reader.readBlock(SequenceFile.java:1574)\n\tat org.apache.hadoop.io.SequenceFile$Reader.nextRawKey(SequenceFile.java:1878)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$SegmentDescriptor.nextRawKey(SequenceFile.java:2894)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue.merge(SequenceFile.java:2694)\n\tat org.apache.hadoop.io.SequenceFile$Sorter.merge(SequenceFile.java:2478)\n\tat org.apache.hadoop.mapred.ReduceTask.run(ReduceTask.java:298)\n\tat org.apache.hadoop.mapred.TaskTracker$Child.main(TaskTracker.java:2049)\n\nor\n\njava.lang.OutOfMemoryError: Java heap space\n\tat org.apache.hadoop.io.compress.DecompressorStream.(DecompressorStream.java:43)\n\tat org.apache.hadoop.io.compress.DefaultCodec.createInputStream(DefaultCodec.java:71)\n\tat org.apache.hadoop.io.SequenceFile$Reader.init(SequenceFile.java:1480)\n\tat org.apache.hadoop.io.SequenceFile$Reader.(SequenceFile.java:1379)\n\tat org.apache.hadoop.io.SequenceFile$Reader.(SequenceFile.java:1302)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$SegmentDescriptor.nextRawKey(SequenceFile.java:2877)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue.merge(SequenceFile.java:2694)\n\tat org.apache.hadoop.io.SequenceFile$Sorter.merge(SequenceFile.java:2478)\n\tat org.apache.hadoop.mapred.ReduceTask.run(ReduceTask.java:298)\n\tat org.apache.hadoop.mapred.TaskTracker$Child.main(TaskTracker.java:2049)\n\nConfiguration:\n\nusing native compression\n1GB of heap space, 1350 nodes\nmapred.inmem.merge.threshold 1000\nmapred.reduce.parallel.copies 10\ntasktracker.http.threads 10\nmapred.map.tasks 2500\nmapred.reduce.tasks 2500\nfs.inmemory.size.mb 200\nio.seqfile.sorter.recordlimit 1000000\nio.sort.mb 200\nio.sort.factor 1000\nmapred.map.output.compression.type BLOCK\nmapred.map.output.compression.codec org.apache.hadoop.io.compress.DefaultCodec\nmapred.compress.map.output true\n\nI tried 2 runs:\n1) io.seqfile.compress.blocksize = 1000000 --> 1084 successful reduces, 935 failures\n2) io.seqfile.compress.blocksize = 131072 --> 2286 successful reduces, 1032 failures\n\nThe failures all seem to occur after shuffling, in the final merge-sort. Because the patch uses a pool of codecs I thought I should be able to keep a high sort.factor (to reduce the amount of multi-phasic merge-sort).","created":"2008-02-18T19:45:14.013+0000"},{"body":"Thanks for trying this out Christian! Do you have the logs of the failed tasks somewhere i.e. the syslog file?\n\nClearly this needs more work... sigh.","created":"2008-02-18T19:56:01.271+0000"},{"body":"Logs are available. I send you the pointers offline.","created":"2008-02-18T20:16:39.024+0000"},{"body":"\n1000 as merge factor seems very high.\nHow much data each reducer isexpected to process?\n","created":"2008-02-18T20:17:23.761+0000"},{"body":"With compression turned off we never had any problems with high merge factor (which reduces time spent in merging). Concerning the amount of data processed by reducers, the particular job has a few (imbalanced) reducers processing 100+ GB uncompressed. With block compression turned on we were hoping to reduce the merging even further.","created":"2008-02-18T20:36:41.414+0000"},{"body":"1000 is a much higher merge factor than is expected by the design.","created":"2008-02-19T19:47:13.864+0000"},{"body":"I'm removing this as a blocker for 0.16.1 while we continue discussions...","created":"2008-02-21T05:20:58.765+0000"},{"body":"What is the RAM overhead of an active decompressor? We can probably compute a reasonable merge size from that. If we are seeing lots of very small files, maybe we should consider just working with the uncompressed data for small inputs? That may use less total RAM and it should be easy to determine that on the fly. \n\nYou could cap the number of compressed files in RAM and when you add a new file you could choose to uncompress the smallest file you have to allow you to stay under the limit. You could also merge and recompress your small files, this obviously uses a lot more CPU.","created":"2008-02-28T07:15:52.679+0000"},{"body":"\nI'd like to elaborate a bit alone Eric's comment.\n\nThe root cause of the current problem is that when we merge too many small segments, \nwe need to allocate too many codecs, thus run out of memory.\nThe main point is to treat large segments and small segments differently.\nIt is intuitive that merging a lot of small segments incurrs too much overhead and may be inferior to quick sort.\n\nLet's say we have a fixed memory budget for in mem merge of fetched map output.\nWhen we fetch a small segment, we extract the records out of the segment and put in an area for future quick sort.\nWhen we fetch a big segment, we leave it in the in mem file.\nWhen the number of in-mem file reach certain limit (say 100), or the total memory consumption reach the budget, we sort the extracted \nrecords, and merge them along with the inmem files.\n\nThis way, we can guarantee that we will not exceed the total memory budget.\n\nThoughts?\n","created":"2008-04-24T23:24:33.975+0000"},{"body":"I would like to suggest a very simple improvement.\n\n1) Compute the maximum number of usable decompressors.\n2) Download splits until ram is full, or we have reached the limit.\n3) If ram is not full\n b) continue downloading splits, but now decompress them as they are loaded\n\n4) Now merge and dump all the splits, decompressing the first N on the fly\n\nThis is very simple and works in almost all cases. A refinement would be to decompress the smallest split each time you load a new split beyond the merge limit.\n\n---\n\nThe above seems like it would be very simple to code and would work well in the face of large splits (the merge limit is not reached) and many small splits (many are merged in the first pass). It would be ok in the face of medium splits, which seems like the worst case.\n\nA more optimal algorithm would presumably merge in ram, compressing on the fly and so on, but this is very complex and has many corner cases.","created":"2008-04-25T04:57:43.528+0000"},{"body":"+1 for Eric's proposal.\n\nA minor refinement to discuss:\n\n3) If ram is not full\nc) initiate the merge with the in-memory compressed files which fit the quota; continue downloading the compressed splits and stuff the compressed split as-is in the InMemoryFileSystem (on ram).\n\nThe next pass of the merge will pick it up and run with it. It has the advantage of saving ram and keeping the merge-code simple: each split is either compressed (compression enabled) or not (compression disabled).\n\nThoughts?","created":"2008-04-25T23:45:46.721+0000"},{"body":"I think that runs contrary to the goal of the exercise, which is to merge as many inputs at once as possible (with as little code complexity as possible). We don't want to add extra passes in the case where the inputs are very small.\n\nPossible refinements from there might include:\n- merging and compressing back into ram (might be very good in some cases)\n- Releasing ram as inputs are merged to allow interleaving of the reads without wasting ram\n\nPS We should be sure to not have the merge factor considered in this code! Merge factor if it should exist at all should be about the number of inputs from disk. Even that is pretty awkward.","created":"2008-04-28T17:05:38.542+0000"},{"body":"It is worth noting that we'd need to reserve mapred.reduce.parallel.copies (default 5, but 20 on yahoo's clusters) extra codecs for decompressing on the fly. It isn't clear to me that it would be a win having the codec owned by the copier thread rather than the merger. \n\nYou would also need to add a header that gives the uncompressed size of the map outputs, so that the reducer knows how big the uncompressed file is.\n\nOf course getting rid of the 4 compression codecs using something like HADOOP-3315, would help a *lot*.","created":"2008-04-28T23:13:42.637+0000"},{"body":"Somre refinements after further discusssions...\n\nPrologue: Use a compressed outputstream for the intermediate sequence files i.e. map-outputs, not {record|block}-compressed sequence files. This cuts down no. of decompressors required at the reducers. Add headers to ensure that the reducer can query each map to find out the exact compressed and uncompressed sizes _before_ it actually copies the data.\n\n1) Compute the maximum number of usable decompressors (this is going to get tricky with direct-buffers taking up non-heap space which is a fraction of the -Xmx of the jvm).\n2) Download map-outputs until ramfs is full; or we have reached the decompressors' limit.\n3) Trigger InMemFSMergeThread to start the merge. (Currently a new InMemFSMergeThread created for every triggered merge, I plan to fix it so that we use one and only one thread.)\n4) If ramfs is full, _suspend_ the shuffle; else keep shuffling into memory.\n\nEssentially the idea is that we pack as much into memory before we initiate the merge, this saves us trips to disk (the output of the merge) which, as Devaraj has shown (http://issues.apache.org/jira/browse/HADOOP-3297?focusedCommentId=12592816#action_12592816) leads to much better overall performance.\n\nOf course the above discussion is valid _iff_ we are dealing with small map-outputs.\n\nFor the contra-case where map-outputs are large, we need a threshold which says:\nIf a map-output > 10% of ramfs, then shuffle into memory if possible (i.e. space available in ramfs), else shuffle to disk. \nThis ensures that we do not needlessly throttle shuffle if map-outputs are too big to fit into ramfs. \n\nThoughts?\n\n----\n\nBefore I jump in and make changes I'm currently trying to simulate the above behaviour and publish some numbers... watch this space.","created":"2008-05-08T07:04:48.326+0000"},{"body":"To clarify some of the thinking above... The short term goal is not to find the optimal solution. It is to get something done that is clean, understandable and works acceptably well in all cases. We can refine from there.\n\nto expand on the above suggestions:\n\nI suggest that for objects larger than 25% of RAM, we always just send them directly to disk. A simple rule that let's us reason more easily about other case. I don't think the 10% number above can be replaced with thiis.\n\nWe need to understand how we pause the copies without doing lots of polling. Again, I suggest keeping it simple for now. What about simply setting a global flag the first time a thread starts to read an input that is < 25% of buffer RAM (not piped directly to disk) and doesn't fit in the remaining space. Other readers will then pause until this semaphore is cleared. It is ok if races happen where a few threads try at the same time. If copies fail, we will need to clear this semaphore too.\n\nWe want to be sure not to wait until RAM is totally full before starting the merge, because this might allow a single slow copy to brownout the system. I suggest a simple rule, such as wait until the semaphore discussed above is set and copies filling at least 50% of RAM have completed. Then merge.\n\nOnce all of the above is done, we can file new jiras to improve things. Ideas include:\n- Freeing storage as we merge, so fetches can be interleaved\n- decompressing small segments as we read so we can increase the number of compressed objects merged\n- ...\n","created":"2008-05-08T20:53:28.171+0000"},{"body":"bq. To clarify some of the thinking above... The short term goal is not to find the optimal solution. It is to get something done that is clean, understandable and works acceptably well in all cases. We can refine from there.\n\n+1\n\nAlong a similar tangent, here is an even simpler proposal for a first-cut. Please bear in mind that these are a result of the fact that we have noticed that the merge code is actually a juicier target to fix; on large jobs we have noticed that reduces (which started _after_ all maps were completed) were spending way more time in merge rather than in shuffle: 13mins in shuffle, 17mins in merge and 15mins in reduce. So, the idea to fix the OOM to ensure better reliability in a reasonably straight-forward, simple manner and then go after merge in HADOOP-3366.\n\nHere goes:\n1. Use a compressed stream for map-outputs, not {record|block}-compressed sequence-files. Ensure both compressed and decompressed sizes are available for the reduce _before_ it actually shuffles the bytes.\n2. Decompress map-outputs as soon as they are shuffled into ramfs if there is enough space, else suspend shuffle as described above by Eric. This ensures that we never hit the #codecs limit, we just need as many codecs as the no. of threads doing the shuffle. The idea is that the merge-factor is anyway going to be limited by the #codecs, we might as well burn-up RAM. We could try and store the compressed outputs as a further refinement in a separate issue.\n3. If RAM is more than (say) 50% full, we start merging in-memory. Also, initially we should use up as much RAM as possible, allowing for some slack. Therefore I propose we do away with the *fs.inmemory.size.mb* config knob and use 3/4 or 2/3 of the heap-size available as the RAM limit.\n4. If the split is greater than 10% or 25% of available RAM limit, and there is on RAM available we shuffle directly to disk (compressed).\n5. The output of merge is compressed and written to disk, which potentially could be merged along with (4) above.\n\nHopefully this is reasonably simple and coherent. I'll put up more thoughts on HADOOP-3366 about merge improvements, current pitfalls etc.\n\nThoughts?","created":"2008-05-13T06:51:58.182+0000"},{"body":"How would the case of multiple threads being suspended be handled? If we have a choice, we should place into ramfs the smaller files first.\nI am slightly concerned about the cpu cycles we will burn for compress/decompress the output of merges. What is the reason for compressing the output of merge?\n","created":"2008-05-13T12:49:00.491+0000"},{"body":"bq. How would the case of multiple threads being suspended be handled? If we have a choice, we should place into ramfs the smaller files first.\nAll threads wait on a single semaphore and we could do a 'notifyAll' ...\n\nbq. I am slightly concerned about the cpu cycles we will burn for compress/decompress the output of merges. What is the reason for compressing the output of merge?\nThe compression of merge-outputs is to reduce temporary disk usage on the reducer node. Also it has the nice property of keeping all data on disk compressed, since compressed splits which can't fit in-memory go straight to disk. ","created":"2008-05-13T14:27:36.202+0000"},{"body":"\nWe can expect better performance if we use lz0 as the codec for compressing the map outputs and the merge outputs.\nSome more benchmarking may be required to confirm this, though.\n ","created":"2008-05-13T16:23:03.719+0000"},{"body":"Here is a nearly complete patch (I need fix a couple of failing test-cases). \n\nHighlights:\n1. Rework sort/merge to use the new IFile rather than SequenceFiles.\n2. Compression for intermediate map-outputs now implies that the entire file is compressed, no more record/block compression. This helps the codec cost.\n3. Rework intermediate merge to ensure there are no spurious copies of keys/values.\n\nBenchmarks:\nI ran this with a single reducer job where 2500 maps produced 5MB of data each. Trunk takes 45 mins for the job to complete with the on-disk merge taking nearly 25-30mins and the final merge (as records are fed to 'reduce') taking 12-13mins. With this patch the on-disk merge takes 20-22mins and the final merge takes around 7mins, an overall improvement of nearly 30%.","created":"2008-06-04T09:55:40.474+0000"},{"body":"I looked at most of the patch except MapTask. +1 on the patch subject to unit tests passing and the sort benchmark running well. ","created":"2008-06-04T19:12:23.241+0000"},{"body":"Thanks for the review Devaraj!\n\nHere is an updated version of the patch with some minor changes...","created":"2008-06-04T20:54:14.844+0000"},{"body":"just one comment:\n\n+ LOG.info(\"Sent out \" + totalRead + \" bytes (starting from offset: \" + startOffset + \" of outputFile: \" + mapOutputFileName + \" for reduce: \" + reduce + \" from map: \" + mapId + \" given \" + partLength + \"/\" + rawPartLength);\n\nyou might want to wrap this around!! :)\n\n","created":"2008-06-04T21:10:46.171+0000"},{"body":"Fixed the log statement... thanks for the review Mahadev!","created":"2008-06-04T21:34:52.733+0000"},{"body":"Pretty-fied another log message...","created":"2008-06-04T23:56:00.079+0000"},{"body":"The javac warning was:\n{noformat}\n [javac] /zonestorage/hudson/home/hudson/hudson/jobs/Hadoop-Patch/workspace/trunk/src/java/org/apache/hadoop/mapred/JobConf.java:467: warning: [dep-ann] deprecated name isnt annotated with @Deprecated\n [javac] public void setMapOutputCompressionType(CompressionType style) {\n [javac] ^\n [javac] /zonestorage/hudson/home/hudson/hudson/jobs/Hadoop-Patch/workspace/trunk/src/java/org/apache/hadoop/mapred/JobConf.java:482: warning: [dep-ann] deprecated name isnt annotated with @Deprecated\n [javac] public CompressionType getMapOutputCompressionType() {\n [javac] ^\n{noformat}\n\nFixed now.\n\nThe test-case failure seems unrelated and works on both Linux and Mac.","created":"2008-06-05T02:49:41.975+0000"},{"body":"Why start the merge before the first fetcher thread is blocked due to \nlack of RAM? This seems like the wrong trade-off to me. Especially \nfor shuffles before the MAP is done.\n\nBe sure to account for brown-out / race conditions. IE we may have a \nfew really slow reads. Hence we should only wait for the first \nthread to block. This should also influence the size of max element \nwe will write to RAM. EG I'd suggest the MAX object stored to RAM be \nless than the (space used) / shuffle threads /2. This will insure \nyou have a full set of elements to merge when you start merging. You \nmight want to code an \"assertion\" for that.\n\n\n\n","created":"2008-06-05T04:02:46.629+0000"},{"body":"I just committed this, many thanks to Devaraj, Chris & Mahadev for review it!","created":"2008-06-05T04:07:24.966+0000"},{"body":"Eric, doesn't it make sense to have a bit of buffer space free even during the merge? On machines where merge is quick (faster CPUs) it will allow shuffle to continue unimpeded... \nCurrently a merge is triggered when the buffer is half-full, which I plan to tweak a bit. \n\nMeanwhile I plan to use HADOOP-3366 to continue the 'stalled shuffle' rework.","created":"2008-06-05T04:32:23.556+0000"},{"body":"Let's get some numbers. Producing the largest possible runs should be a long term goal. It sounds like this will be an improvement over what we had.\n\nAn optimal solution would fill RAM (with compressed data) and then release that RAM as it merged, allowing the shuffle to be interleaved.\nMaybe you should create a new JIRA to track possible future improvements?","created":"2008-06-05T05:59:25.977+0000"}],"conversations":[{"body":"One of the reducers of my job failed with the following exceptions.\nThe failure caused the whole job fail eventually.\nJava heapsize was 768MB and sort.io.mb was 140.\n\n\n2007-10-23 19:24:06,100 WARN org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Intermediate Merge of the inmemory files threw an exception: java.lang.OutOfMemoryError: Java heap space\n\tat org.apache.hadoop.io.compress.DecompressorStream.(DecompressorStream.java:43)\n\tat org.apache.hadoop.io.compress.DefaultCodec.createInputStream(DefaultCodec.java:71)\n\tat org.apache.hadoop.io.SequenceFile$Reader.init(SequenceFile.java:1345)\n\tat org.apache.hadoop.io.SequenceFile$Reader.(SequenceFile.java:1231)\n\tat org.apache.hadoop.io.SequenceFile$Reader.(SequenceFile.java:1154)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$SegmentDescriptor.nextRawKey(SequenceFile.java:2726)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue.merge(SequenceFile.java:2543)\n\tat org.apache.hadoop.io.SequenceFile$Sorter.merge(SequenceFile.java:2297)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$InMemFSMergeThread.run(ReduceTask.java:1311)\n2007-10-23 19:24:06,102 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 done copying task_200710231912_0001_m_001428_0 output .\n2007-10-23 19:24:06,185 INFO org.apache.hadoop.fs.FileSystem: Initialized InMemoryFileSystem: ramfs://mapoutput31952838/task_200710231912_0001_r_000020_2/map_1423.out-0 of size (in bytes): 209715200\n2007-10-23 19:24:06,193 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,193 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001215_0 output from xxx\n2007-10-23 19:24:06,188 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001211_0 output from xxx\n2007-10-23 19:24:06,185 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryOutputStream.close(InMemoryFileSystem.java:161)\n\tat org.apache.hadoop.fs.FSDataOutputStream$PositionCache.close(FSDataOutputStream.java:49)\n\tat org.apache.hadoop.fs.FSDataOutputStream.close(FSDataOutputStream.java:64)\n\tat org.apache.hadoop.fs.ChecksumFileSystem$ChecksumFSOutputSummer.close(ChecksumFileSystem.java:312)\n\tat org.apache.hadoop.fs.FSDataOutputStream$PositionCache.close(FSDataOutputStream.java:49)\n\tat org.apache.hadoop.fs.FSDataOutputStream.close(FSDataOutputStream.java:64)\n\tat org.apache.hadoop.mapred.MapOutputLocation.getFile(MapOutputLocation.java:253)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:713)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,199 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001247_0 output from .\n2007-10-23 19:24:06,200 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,204 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001422_0 output from .\n2007-10-23 19:24:06,207 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,209 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001278_0 output from .\n2007-10-23 19:24:06,198 WARN org.apache.hadoop.mapred.TaskTracker: Error running child\njava.io.IOException: task_200710231912_0001_r_000020_2The reduce copier failed\n\tat org.apache.hadoop.mapred.ReduceTask.run(ReduceTask.java:253)\n\tat org.apache.hadoop.mapred.TaskTracker$Child.main(TaskTracker.java:1760)\n2007-10-23 19:24:06,198 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,231 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001531_0 output from .\n2007-10-23 19:24:06,197 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,237 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001227_0 output from .\n2007-10-23 19:24:06,196 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n","from":"reporter","subject":"Reducer failed due to Out ofMemory"},{"body":"The problem was gone after I set the compressMapOutput attribute to false.\n","from":"developer"},{"body":"Is it true that native compression was in use when you saw the problem?","from":"developer"},{"body":"Yes.\n","from":"developer"},{"body":"My hunch is this that we are ending up with too many files in the ramfs (the map outputs for the failing reduce are small in size). So when merge is initiated on those files, we end up creating too many codecs for decompressing the files, and since they are native, we encounter OOM. So one way to validate this is to decrease the value of mapred.inmem.merge.threshold (defaults to 1000) to something like 200 and then see whether the OOM goes away. The other alternative is to increase the heap size per task.","from":"developer"},{"body":"FYI, I ran into the exactly same issue with massive failures (nearly all reduces), using a 0 value for mapred.inmem.merge.threshold (letting the framework select the threshold)\n\nusing native compression\n\n1GB of heap space, 1350 nodes\n\nmapred.inmem.merge.threshold\t0\nmapred.reduce.parallel.copies\t20\ntasktracker.http.threads\t30\nmapred.map.tasks\t13008\nmapred.reduce.tasks\t3600\nfs.inmemory.size.mb\t200\nio.seqfile.sorter.recordlimit\t1000000\nio.sort.mb\t200\nio.sort.factor\t300\nmapred.map.output.compression.type\tRECORD\nmapred.map.output.compression.codec\torg.apache.hadoop.io.compress.DefaultCodec\nmapred.compress.map.output\ttrue\n","from":"developer"},{"body":"Christian, please don't use 0 for mapred.inmem.merge.threshold. It basically disables the threshold for initiating merge, and, in this case, merge on the ramfs files gets triggered only when the ramfs has collected enough bytes of data (actually 50% of fs.inmemory.size.mb). Depending on how big each map output is, we will accumulate that many files there. For each file, a native codec will be initialized, and I suspect that if we create too many codecs, we will encounter OOM (given a certain child heap size). So, I think it is worth trying the app with a reduced mapred.inmem.merge.threshold like 50/100, just to ensure that we keep a control on the number of codecs we create. Christian/Runping, could you please try this tweak and let us know? Thanks.","from":"developer"},{"body":"Just to clarify, a codec-pool is maintained in the compression module. So, the number of codecs created will be in the order of mapred.inmem.merge.threshold.","from":"developer"},{"body":"BTW, another thing that needs to be controlled is the io.sort.factor which would control the number of files that gets merged at once post the shuffle (the final merge of the on-disk files). The number of intermediate files opened would be as per the value of io.sort.factor and you might encounter OOM if that is large. So i would recommend that for working around this issue, the io.sort.factor value should be in the same order as the value of mapred.inmem.merge.threshold (that successfully does ramfs merges during the shuffle).","from":"developer"},{"body":"Christian,\n Can you generate a set of task files for the reduce with the failure with keep.failed.task.files set to true, so that we can run it in the isolation runner and a profiler?\n","from":"developer"},{"body":"Noticed that codec pool is not used in o.a.h.i.SequenceFile.CompressedBytes.writeUncompressedBytes. This needs to be fixed.","from":"developer"},{"body":"Christian/Runping - Could you please re-run your jobs using this debug patch? It just logs the no. of created native codecs and the stack trace to check where they are being created from. Thanks!\n\nMeanwhile I'll plug away and see how to fix SequenceFile.CompressedBytes as Devaraj pointed out...","from":"developer"},{"body":"Here is an early patch which fixed {{SequenceFile.CompressedBytes}} to use a {{CodecPool}} passed along by the {{SequenceFile.Reader}}... \n\nChristian/Runping -I'd appreciate if you could try this patch alongwith the previous debug patch while I try to test it at my end. Thanks!","from":"developer"},{"body":"Sorry, previous patch had the leading path wrong...","from":"developer"},{"body":"Sorry, previous patch had the leading path wrong...","from":"developer"},{"body":"I still see failures after shuffling during final sort:\n\njava.lang.OutOfMemoryError: Java heap space\n\tat org.apache.hadoop.io.DataOutputBuffer$Buffer.write(DataOutputBuffer.java:52)\n\tat org.apache.hadoop.io.DataOutputBuffer.write(DataOutputBuffer.java:90)\n\tat org.apache.hadoop.io.SequenceFile$Reader.readBuffer(SequenceFile.java:1535)\n\tat org.apache.hadoop.io.SequenceFile$Reader.readBlock(SequenceFile.java:1574)\n\tat org.apache.hadoop.io.SequenceFile$Reader.nextRawKey(SequenceFile.java:1878)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$SegmentDescriptor.nextRawKey(SequenceFile.java:2894)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue.merge(SequenceFile.java:2694)\n\tat org.apache.hadoop.io.SequenceFile$Sorter.merge(SequenceFile.java:2478)\n\tat org.apache.hadoop.mapred.ReduceTask.run(ReduceTask.java:298)\n\tat org.apache.hadoop.mapred.TaskTracker$Child.main(TaskTracker.java:2049)\n\nor\n\njava.lang.OutOfMemoryError: Java heap space\n\tat org.apache.hadoop.io.compress.DecompressorStream.(DecompressorStream.java:43)\n\tat org.apache.hadoop.io.compress.DefaultCodec.createInputStream(DefaultCodec.java:71)\n\tat org.apache.hadoop.io.SequenceFile$Reader.init(SequenceFile.java:1480)\n\tat org.apache.hadoop.io.SequenceFile$Reader.(SequenceFile.java:1379)\n\tat org.apache.hadoop.io.SequenceFile$Reader.(SequenceFile.java:1302)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$SegmentDescriptor.nextRawKey(SequenceFile.java:2877)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue.merge(SequenceFile.java:2694)\n\tat org.apache.hadoop.io.SequenceFile$Sorter.merge(SequenceFile.java:2478)\n\tat org.apache.hadoop.mapred.ReduceTask.run(ReduceTask.java:298)\n\tat org.apache.hadoop.mapred.TaskTracker$Child.main(TaskTracker.java:2049)\n\nConfiguration:\n\nusing native compression\n1GB of heap space, 1350 nodes\nmapred.inmem.merge.threshold 1000\nmapred.reduce.parallel.copies 10\ntasktracker.http.threads 10\nmapred.map.tasks 2500\nmapred.reduce.tasks 2500\nfs.inmemory.size.mb 200\nio.seqfile.sorter.recordlimit 1000000\nio.sort.mb 200\nio.sort.factor 1000\nmapred.map.output.compression.type BLOCK\nmapred.map.output.compression.codec org.apache.hadoop.io.compress.DefaultCodec\nmapred.compress.map.output true\n\nI tried 2 runs:\n1) io.seqfile.compress.blocksize = 1000000 --> 1084 successful reduces, 935 failures\n2) io.seqfile.compress.blocksize = 131072 --> 2286 successful reduces, 1032 failures\n\nThe failures all seem to occur after shuffling, in the final merge-sort. Because the patch uses a pool of codecs I thought I should be able to keep a high sort.factor (to reduce the amount of multi-phasic merge-sort).","from":"developer"},{"body":"Thanks for trying this out Christian! Do you have the logs of the failed tasks somewhere i.e. the syslog file?\n\nClearly this needs more work... sigh.","from":"developer"},{"body":"Logs are available. I send you the pointers offline.","from":"developer"},{"body":"\n1000 as merge factor seems very high.\nHow much data each reducer isexpected to process?\n","from":"developer"},{"body":"With compression turned off we never had any problems with high merge factor (which reduces time spent in merging). Concerning the amount of data processed by reducers, the particular job has a few (imbalanced) reducers processing 100+ GB uncompressed. With block compression turned on we were hoping to reduce the merging even further.","from":"developer"},{"body":"1000 is a much higher merge factor than is expected by the design.","from":"developer"},{"body":"I'm removing this as a blocker for 0.16.1 while we continue discussions...","from":"developer"},{"body":"What is the RAM overhead of an active decompressor? We can probably compute a reasonable merge size from that. If we are seeing lots of very small files, maybe we should consider just working with the uncompressed data for small inputs? That may use less total RAM and it should be easy to determine that on the fly. \n\nYou could cap the number of compressed files in RAM and when you add a new file you could choose to uncompress the smallest file you have to allow you to stay under the limit. You could also merge and recompress your small files, this obviously uses a lot more CPU.","from":"developer"},{"body":"\nI'd like to elaborate a bit alone Eric's comment.\n\nThe root cause of the current problem is that when we merge too many small segments, \nwe need to allocate too many codecs, thus run out of memory.\nThe main point is to treat large segments and small segments differently.\nIt is intuitive that merging a lot of small segments incurrs too much overhead and may be inferior to quick sort.\n\nLet's say we have a fixed memory budget for in mem merge of fetched map output.\nWhen we fetch a small segment, we extract the records out of the segment and put in an area for future quick sort.\nWhen we fetch a big segment, we leave it in the in mem file.\nWhen the number of in-mem file reach certain limit (say 100), or the total memory consumption reach the budget, we sort the extracted \nrecords, and merge them along with the inmem files.\n\nThis way, we can guarantee that we will not exceed the total memory budget.\n\nThoughts?\n","from":"developer"},{"body":"I would like to suggest a very simple improvement.\n\n1) Compute the maximum number of usable decompressors.\n2) Download splits until ram is full, or we have reached the limit.\n3) If ram is not full\n b) continue downloading splits, but now decompress them as they are loaded\n\n4) Now merge and dump all the splits, decompressing the first N on the fly\n\nThis is very simple and works in almost all cases. A refinement would be to decompress the smallest split each time you load a new split beyond the merge limit.\n\n---\n\nThe above seems like it would be very simple to code and would work well in the face of large splits (the merge limit is not reached) and many small splits (many are merged in the first pass). It would be ok in the face of medium splits, which seems like the worst case.\n\nA more optimal algorithm would presumably merge in ram, compressing on the fly and so on, but this is very complex and has many corner cases.","from":"developer"},{"body":"+1 for Eric's proposal.\n\nA minor refinement to discuss:\n\n3) If ram is not full\nc) initiate the merge with the in-memory compressed files which fit the quota; continue downloading the compressed splits and stuff the compressed split as-is in the InMemoryFileSystem (on ram).\n\nThe next pass of the merge will pick it up and run with it. It has the advantage of saving ram and keeping the merge-code simple: each split is either compressed (compression enabled) or not (compression disabled).\n\nThoughts?","from":"developer"},{"body":"I think that runs contrary to the goal of the exercise, which is to merge as many inputs at once as possible (with as little code complexity as possible). We don't want to add extra passes in the case where the inputs are very small.\n\nPossible refinements from there might include:\n- merging and compressing back into ram (might be very good in some cases)\n- Releasing ram as inputs are merged to allow interleaving of the reads without wasting ram\n\nPS We should be sure to not have the merge factor considered in this code! Merge factor if it should exist at all should be about the number of inputs from disk. Even that is pretty awkward.","from":"developer"},{"body":"It is worth noting that we'd need to reserve mapred.reduce.parallel.copies (default 5, but 20 on yahoo's clusters) extra codecs for decompressing on the fly. It isn't clear to me that it would be a win having the codec owned by the copier thread rather than the merger. \n\nYou would also need to add a header that gives the uncompressed size of the map outputs, so that the reducer knows how big the uncompressed file is.\n\nOf course getting rid of the 4 compression codecs using something like HADOOP-3315, would help a *lot*.","from":"developer"},{"body":"Somre refinements after further discusssions...\n\nPrologue: Use a compressed outputstream for the intermediate sequence files i.e. map-outputs, not {record|block}-compressed sequence files. This cuts down no. of decompressors required at the reducers. Add headers to ensure that the reducer can query each map to find out the exact compressed and uncompressed sizes _before_ it actually copies the data.\n\n1) Compute the maximum number of usable decompressors (this is going to get tricky with direct-buffers taking up non-heap space which is a fraction of the -Xmx of the jvm).\n2) Download map-outputs until ramfs is full; or we have reached the decompressors' limit.\n3) Trigger InMemFSMergeThread to start the merge. (Currently a new InMemFSMergeThread created for every triggered merge, I plan to fix it so that we use one and only one thread.)\n4) If ramfs is full, _suspend_ the shuffle; else keep shuffling into memory.\n\nEssentially the idea is that we pack as much into memory before we initiate the merge, this saves us trips to disk (the output of the merge) which, as Devaraj has shown (http://issues.apache.org/jira/browse/HADOOP-3297?focusedCommentId=12592816#action_12592816) leads to much better overall performance.\n\nOf course the above discussion is valid _iff_ we are dealing with small map-outputs.\n\nFor the contra-case where map-outputs are large, we need a threshold which says:\nIf a map-output > 10% of ramfs, then shuffle into memory if possible (i.e. space available in ramfs), else shuffle to disk. \nThis ensures that we do not needlessly throttle shuffle if map-outputs are too big to fit into ramfs. \n\nThoughts?\n\n----\n\nBefore I jump in and make changes I'm currently trying to simulate the above behaviour and publish some numbers... watch this space.","from":"developer"},{"body":"To clarify some of the thinking above... The short term goal is not to find the optimal solution. It is to get something done that is clean, understandable and works acceptably well in all cases. We can refine from there.\n\nto expand on the above suggestions:\n\nI suggest that for objects larger than 25% of RAM, we always just send them directly to disk. A simple rule that let's us reason more easily about other case. I don't think the 10% number above can be replaced with thiis.\n\nWe need to understand how we pause the copies without doing lots of polling. Again, I suggest keeping it simple for now. What about simply setting a global flag the first time a thread starts to read an input that is < 25% of buffer RAM (not piped directly to disk) and doesn't fit in the remaining space. Other readers will then pause until this semaphore is cleared. It is ok if races happen where a few threads try at the same time. If copies fail, we will need to clear this semaphore too.\n\nWe want to be sure not to wait until RAM is totally full before starting the merge, because this might allow a single slow copy to brownout the system. I suggest a simple rule, such as wait until the semaphore discussed above is set and copies filling at least 50% of RAM have completed. Then merge.\n\nOnce all of the above is done, we can file new jiras to improve things. Ideas include:\n- Freeing storage as we merge, so fetches can be interleaved\n- decompressing small segments as we read so we can increase the number of compressed objects merged\n- ...\n","from":"developer"},{"body":"bq. To clarify some of the thinking above... The short term goal is not to find the optimal solution. It is to get something done that is clean, understandable and works acceptably well in all cases. We can refine from there.\n\n+1\n\nAlong a similar tangent, here is an even simpler proposal for a first-cut. Please bear in mind that these are a result of the fact that we have noticed that the merge code is actually a juicier target to fix; on large jobs we have noticed that reduces (which started _after_ all maps were completed) were spending way more time in merge rather than in shuffle: 13mins in shuffle, 17mins in merge and 15mins in reduce. So, the idea to fix the OOM to ensure better reliability in a reasonably straight-forward, simple manner and then go after merge in HADOOP-3366.\n\nHere goes:\n1. Use a compressed stream for map-outputs, not {record|block}-compressed sequence-files. Ensure both compressed and decompressed sizes are available for the reduce _before_ it actually shuffles the bytes.\n2. Decompress map-outputs as soon as they are shuffled into ramfs if there is enough space, else suspend shuffle as described above by Eric. This ensures that we never hit the #codecs limit, we just need as many codecs as the no. of threads doing the shuffle. The idea is that the merge-factor is anyway going to be limited by the #codecs, we might as well burn-up RAM. We could try and store the compressed outputs as a further refinement in a separate issue.\n3. If RAM is more than (say) 50% full, we start merging in-memory. Also, initially we should use up as much RAM as possible, allowing for some slack. Therefore I propose we do away with the *fs.inmemory.size.mb* config knob and use 3/4 or 2/3 of the heap-size available as the RAM limit.\n4. If the split is greater than 10% or 25% of available RAM limit, and there is on RAM available we shuffle directly to disk (compressed).\n5. The output of merge is compressed and written to disk, which potentially could be merged along with (4) above.\n\nHopefully this is reasonably simple and coherent. I'll put up more thoughts on HADOOP-3366 about merge improvements, current pitfalls etc.\n\nThoughts?","from":"developer"},{"body":"How would the case of multiple threads being suspended be handled? If we have a choice, we should place into ramfs the smaller files first.\nI am slightly concerned about the cpu cycles we will burn for compress/decompress the output of merges. What is the reason for compressing the output of merge?\n","from":"developer"},{"body":"bq. How would the case of multiple threads being suspended be handled? If we have a choice, we should place into ramfs the smaller files first.\nAll threads wait on a single semaphore and we could do a 'notifyAll' ...\n\nbq. I am slightly concerned about the cpu cycles we will burn for compress/decompress the output of merges. What is the reason for compressing the output of merge?\nThe compression of merge-outputs is to reduce temporary disk usage on the reducer node. Also it has the nice property of keeping all data on disk compressed, since compressed splits which can't fit in-memory go straight to disk. ","from":"developer"},{"body":"\nWe can expect better performance if we use lz0 as the codec for compressing the map outputs and the merge outputs.\nSome more benchmarking may be required to confirm this, though.\n ","from":"developer"},{"body":"Here is a nearly complete patch (I need fix a couple of failing test-cases). \n\nHighlights:\n1. Rework sort/merge to use the new IFile rather than SequenceFiles.\n2. Compression for intermediate map-outputs now implies that the entire file is compressed, no more record/block compression. This helps the codec cost.\n3. Rework intermediate merge to ensure there are no spurious copies of keys/values.\n\nBenchmarks:\nI ran this with a single reducer job where 2500 maps produced 5MB of data each. Trunk takes 45 mins for the job to complete with the on-disk merge taking nearly 25-30mins and the final merge (as records are fed to 'reduce') taking 12-13mins. With this patch the on-disk merge takes 20-22mins and the final merge takes around 7mins, an overall improvement of nearly 30%.","from":"developer"},{"body":"I looked at most of the patch except MapTask. +1 on the patch subject to unit tests passing and the sort benchmark running well. ","from":"developer"},{"body":"Thanks for the review Devaraj!\n\nHere is an updated version of the patch with some minor changes...","from":"developer"},{"body":"just one comment:\n\n+ LOG.info(\"Sent out \" + totalRead + \" bytes (starting from offset: \" + startOffset + \" of outputFile: \" + mapOutputFileName + \" for reduce: \" + reduce + \" from map: \" + mapId + \" given \" + partLength + \"/\" + rawPartLength);\n\nyou might want to wrap this around!! :)\n\n","from":"developer"},{"body":"Fixed the log statement... thanks for the review Mahadev!","from":"developer"},{"body":"Pretty-fied another log message...","from":"developer"},{"body":"The javac warning was:\n{noformat}\n [javac] /zonestorage/hudson/home/hudson/hudson/jobs/Hadoop-Patch/workspace/trunk/src/java/org/apache/hadoop/mapred/JobConf.java:467: warning: [dep-ann] deprecated name isnt annotated with @Deprecated\n [javac] public void setMapOutputCompressionType(CompressionType style) {\n [javac] ^\n [javac] /zonestorage/hudson/home/hudson/hudson/jobs/Hadoop-Patch/workspace/trunk/src/java/org/apache/hadoop/mapred/JobConf.java:482: warning: [dep-ann] deprecated name isnt annotated with @Deprecated\n [javac] public CompressionType getMapOutputCompressionType() {\n [javac] ^\n{noformat}\n\nFixed now.\n\nThe test-case failure seems unrelated and works on both Linux and Mac.","from":"developer"},{"body":"Why start the merge before the first fetcher thread is blocked due to \nlack of RAM? This seems like the wrong trade-off to me. Especially \nfor shuffles before the MAP is done.\n\nBe sure to account for brown-out / race conditions. IE we may have a \nfew really slow reads. Hence we should only wait for the first \nthread to block. This should also influence the size of max element \nwe will write to RAM. EG I'd suggest the MAX object stored to RAM be \nless than the (space used) / shuffle threads /2. This will insure \nyou have a full set of elements to merge when you start merging. You \nmight want to code an \"assertion\" for that.\n\n\n\n","from":"developer"},{"body":"I just committed this, many thanks to Devaraj, Chris & Mahadev for review it!","from":"developer"},{"body":"Eric, doesn't it make sense to have a bit of buffer space free even during the merge? On machines where merge is quick (faster CPUs) it will allow shuffle to continue unimpeded... \nCurrently a merge is triggered when the buffer is half-full, which I plan to tweak a bit. \n\nMeanwhile I plan to use HADOOP-3366 to continue the 'stalled shuffle' rework.","from":"developer"},{"body":"Let's get some numbers. Producing the largest possible runs should be a long term goal. It sounds like this will be an improvement over what we had.\n\nAn optimal solution would fill RAM (with compressed data) and then release that RAM as it merged, allowing the shuffle to be interleaved.\nMaybe you should create a new JIRA to track possible future improvements?","from":"developer"}],"created":"2007-10-23T20:38:18.000+0000","description":"One of the reducers of my job failed with the following exceptions.\nThe failure caused the whole job fail eventually.\nJava heapsize was 768MB and sort.io.mb was 140.\n\n\n2007-10-23 19:24:06,100 WARN org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Intermediate Merge of the inmemory files threw an exception: java.lang.OutOfMemoryError: Java heap space\n\tat org.apache.hadoop.io.compress.DecompressorStream.(DecompressorStream.java:43)\n\tat org.apache.hadoop.io.compress.DefaultCodec.createInputStream(DefaultCodec.java:71)\n\tat org.apache.hadoop.io.SequenceFile$Reader.init(SequenceFile.java:1345)\n\tat org.apache.hadoop.io.SequenceFile$Reader.(SequenceFile.java:1231)\n\tat org.apache.hadoop.io.SequenceFile$Reader.(SequenceFile.java:1154)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$SegmentDescriptor.nextRawKey(SequenceFile.java:2726)\n\tat org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue.merge(SequenceFile.java:2543)\n\tat org.apache.hadoop.io.SequenceFile$Sorter.merge(SequenceFile.java:2297)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$InMemFSMergeThread.run(ReduceTask.java:1311)\n2007-10-23 19:24:06,102 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 done copying task_200710231912_0001_m_001428_0 output .\n2007-10-23 19:24:06,185 INFO org.apache.hadoop.fs.FileSystem: Initialized InMemoryFileSystem: ramfs://mapoutput31952838/task_200710231912_0001_r_000020_2/map_1423.out-0 of size (in bytes): 209715200\n2007-10-23 19:24:06,193 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,193 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001215_0 output from xxx\n2007-10-23 19:24:06,188 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001211_0 output from xxx\n2007-10-23 19:24:06,185 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryOutputStream.close(InMemoryFileSystem.java:161)\n\tat org.apache.hadoop.fs.FSDataOutputStream$PositionCache.close(FSDataOutputStream.java:49)\n\tat org.apache.hadoop.fs.FSDataOutputStream.close(FSDataOutputStream.java:64)\n\tat org.apache.hadoop.fs.ChecksumFileSystem$ChecksumFSOutputSummer.close(ChecksumFileSystem.java:312)\n\tat org.apache.hadoop.fs.FSDataOutputStream$PositionCache.close(FSDataOutputStream.java:49)\n\tat org.apache.hadoop.fs.FSDataOutputStream.close(FSDataOutputStream.java:64)\n\tat org.apache.hadoop.mapred.MapOutputLocation.getFile(MapOutputLocation.java:253)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:713)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,199 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001247_0 output from .\n2007-10-23 19:24:06,200 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,204 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001422_0 output from .\n2007-10-23 19:24:06,207 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,209 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001278_0 output from .\n2007-10-23 19:24:06,198 WARN org.apache.hadoop.mapred.TaskTracker: Error running child\njava.io.IOException: task_200710231912_0001_r_000020_2The reduce copier failed\n\tat org.apache.hadoop.mapred.ReduceTask.run(ReduceTask.java:253)\n\tat org.apache.hadoop.mapred.TaskTracker$Child.main(TaskTracker.java:1760)\n2007-10-23 19:24:06,198 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,231 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001531_0 output from .\n2007-10-23 19:24:06,197 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n2007-10-23 19:24:06,237 INFO org.apache.hadoop.mapred.ReduceTask: task_200710231912_0001_r_000020_2 Copying task_200710231912_0001_m_001227_0 output from .\n2007-10-23 19:24:06,196 ERROR org.apache.hadoop.mapred.ReduceTask: Map output copy failure: java.lang.NullPointerException\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$FileAttributes.access$300(InMemoryFileSystem.java:366)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem$InMemoryFileStatus.(InMemoryFileSystem.java:378)\n\tat org.apache.hadoop.fs.InMemoryFileSystem$RawInMemoryFileSystem.getFileStatus(InMemoryFileSystem.java:283)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:251)\n\tat org.apache.hadoop.fs.FileSystem.getLength(FileSystem.java:449)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.copyOutput(ReduceTask.java:738)\n\tat org.apache.hadoop.mapred.ReduceTask$ReduceCopier$MapOutputCopier.run(ReduceTask.java:665)\n\n","issue_id":"12381005","key":"HADOOP-2095","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-06-05T04:07:25.000+0000","role":"fixed_distractor","summary":"Reducer failed due to Out ofMemory"} {"case_id":"12381891","cluster":"DISTRACTOR-HADOOP-2159","comments":[{"body":"I would check for floating point precision problems. If the calculation was done carelessly, you may end up not being able to reach the desired goal of 1.0.","created":"2007-11-06T17:48:37.422+0000"},{"body":"Did you mean to say:\nthe namenode does +*NOT*+ turn off safemode automatically","created":"2007-11-06T18:04:25.550+0000"},{"body":"Yes, I missed the *not*. Thanks for pointing out.","created":"2007-11-06T18:32:03.723+0000"},{"body":" Name Node keeps track of the total number of valid block it received in safe mode. A valid block is a block that belongs to a file. The counter is called blockSafe. The name node does not leave the safe mode automatically if the ratio of blockSafe to the total number of valid blocks is less the threshold.\n\n I see a bug in maintaining this counter. Before the counter is incremented, the name node check if the block is valid. Before it does not do the check before this counter is decremented.\n\nWhen a dfs cluster is started, if an early started data node has stale blocks, the name node will ask the data node to delete the stale blocks as the reply to its first block report. If its second block report comes in when the name node is still in safe mode, those blocks will be removed from the blocks map, and the blockSafe counter will also be decremented even though those blocks are invalid. So the cluster will end up with a blockSafe counter that's smaller than the number of valid blocks in namenode. If the threshold is set to be 1, the cluster will not be able to leave the safe mode. ","created":"2008-05-21T22:33:35.975+0000"},{"body":"We also had one cluster that didn't come out of safemode.\n\nfsck showed, \n{noformat}\nStatus: HEALTHY\n Total dirs: 270001\n Total files: 1036456\n Total blocks: 1982902 (avg. block size 92619603 B)\n Minimally replicated blocks: 1982902 (100.00001 %)\n Over-replicated blocks: 0 (0.0 %)\n Under-replicated blocks: 0 (0.0 %)\n Mis-replicated blocks: 0 (0.0 %)\n Default replication factor: 3\n Average block replication: 3.0000212\n Missing replicas: 0 (0.0 %)\nThe filesystem under path '/' is HEALTHY\n{noformat}\n\nI attached btrace to obtain the state of Namenode.namesystem.safeMode, \n\n{noformat}\n-bash-3.1$ pgrep -f NameNode\n30490\n\n-bash-3.1$ btrace -cp hadoop-core.jar 30490 DFSSafeModeTrace.java\nentered org.apache.hadoop.dfs.FSNamesystem.isInSafeMode\n\norg.apache.hadoop.dfs.FSNamesystem@3c992fa5\n{threshold=1.0, extension=30000, safeReplication=1, reached=0, blockTotal=1982902, blockSafe=1982812,\nthis$0=org.apache.hadoop.dfs.FSNamesystem@3c992fa5, }\n{noformat}\n\nThis shows that blockSafe < blockTotal, which supports Hairong's comment above.\n\n(dfs.safemode.threshold.pct is set to 1.0f)","created":"2008-05-21T22:51:46.190+0000"},{"body":"Please note that the cluster did have stale blocks. As shown in the fsck result, the ratio of the minimally replicated blocks to the total number of valid blocks is greater than 100%,","created":"2008-05-21T22:55:13.375+0000"},{"body":"Here is a patch that should fix the bug.","created":"2008-05-21T22:59:27.402+0000"},{"body":"+1 the patch looks good","created":"2008-05-22T02:15:15.936+0000"},{"body":"The patch is an one-line change that was analyzed and tested on a real cluster. But a unit test is not trival, so it is not required.\n\nI just committed this.","created":"2008-05-30T23:58:12.541+0000"}],"conversations":[{"body":"Occasionally (not easy to reproduce) the namenode does turn off safemode automatically, although fsck does not report any missing or under-replicated blocks (safemode threshold set to 1.0).\n\nAt this moment I do not have any additional information which could help analyze the issue.","from":"reporter","subject":"Namenode stuck in safemode"},{"body":"I would check for floating point precision problems. If the calculation was done carelessly, you may end up not being able to reach the desired goal of 1.0.","from":"developer"},{"body":"Did you mean to say:\nthe namenode does +*NOT*+ turn off safemode automatically","from":"developer"},{"body":"Yes, I missed the *not*. Thanks for pointing out.","from":"developer"},{"body":" Name Node keeps track of the total number of valid block it received in safe mode. A valid block is a block that belongs to a file. The counter is called blockSafe. The name node does not leave the safe mode automatically if the ratio of blockSafe to the total number of valid blocks is less the threshold.\n\n I see a bug in maintaining this counter. Before the counter is incremented, the name node check if the block is valid. Before it does not do the check before this counter is decremented.\n\nWhen a dfs cluster is started, if an early started data node has stale blocks, the name node will ask the data node to delete the stale blocks as the reply to its first block report. If its second block report comes in when the name node is still in safe mode, those blocks will be removed from the blocks map, and the blockSafe counter will also be decremented even though those blocks are invalid. So the cluster will end up with a blockSafe counter that's smaller than the number of valid blocks in namenode. If the threshold is set to be 1, the cluster will not be able to leave the safe mode. ","from":"developer"},{"body":"We also had one cluster that didn't come out of safemode.\n\nfsck showed, \n{noformat}\nStatus: HEALTHY\n Total dirs: 270001\n Total files: 1036456\n Total blocks: 1982902 (avg. block size 92619603 B)\n Minimally replicated blocks: 1982902 (100.00001 %)\n Over-replicated blocks: 0 (0.0 %)\n Under-replicated blocks: 0 (0.0 %)\n Mis-replicated blocks: 0 (0.0 %)\n Default replication factor: 3\n Average block replication: 3.0000212\n Missing replicas: 0 (0.0 %)\nThe filesystem under path '/' is HEALTHY\n{noformat}\n\nI attached btrace to obtain the state of Namenode.namesystem.safeMode, \n\n{noformat}\n-bash-3.1$ pgrep -f NameNode\n30490\n\n-bash-3.1$ btrace -cp hadoop-core.jar 30490 DFSSafeModeTrace.java\nentered org.apache.hadoop.dfs.FSNamesystem.isInSafeMode\n\norg.apache.hadoop.dfs.FSNamesystem@3c992fa5\n{threshold=1.0, extension=30000, safeReplication=1, reached=0, blockTotal=1982902, blockSafe=1982812,\nthis$0=org.apache.hadoop.dfs.FSNamesystem@3c992fa5, }\n{noformat}\n\nThis shows that blockSafe < blockTotal, which supports Hairong's comment above.\n\n(dfs.safemode.threshold.pct is set to 1.0f)","from":"developer"},{"body":"Please note that the cluster did have stale blocks. As shown in the fsck result, the ratio of the minimally replicated blocks to the total number of valid blocks is greater than 100%,","from":"developer"},{"body":"Here is a patch that should fix the bug.","from":"developer"},{"body":"+1 the patch looks good","from":"developer"},{"body":"The patch is an one-line change that was analyzed and tested on a real cluster. But a unit test is not trival, so it is not required.\n\nI just committed this.","from":"developer"}],"created":"2007-11-06T06:34:53.000+0000","description":"Occasionally (not easy to reproduce) the namenode does turn off safemode automatically, although fsck does not report any missing or under-replicated blocks (safemode threshold set to 1.0).\n\nAt this moment I do not have any additional information which could help analyze the issue.","issue_id":"12381891","key":"HADOOP-2159","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-06-02T23:34:17.000+0000","role":"fixed_distractor","summary":"Namenode stuck in safemode"} {"case_id":"12382110","cluster":"DISTRACTOR-HADOOP-2172","comments":[{"body":"The claim in particular (from IRC) is that random access to a MapFile stored on the LocalFileSystem has become much slower, and that much of the time is taken seeking the CRC file. Seeks within the current buffer should not require system calls, but perhaps that optimization was lost?\n\nI hacked TestArrayFile to test this, and ArrayFile.get(0) run 100,000 times took around 1 second in 0.13 and now takes over ten.\n","created":"2007-11-08T19:09:58.964+0000"},{"body":"This patch brings back the PositionCache from 0.13.1. It appears the position cache previously did not keep track of long skip(long) or int read(). This has been added.\n\nPasses all the tests on 0.14.3 and has been tested on a 35 node cluster.","created":"2007-11-08T19:11:00.271+0000"},{"body":"This restores the 0.13 performance for my benchmark. +1\n\nHairong, can you please review this, since you removed PositionCache as a part of HADOOP-1470? Thanks!","created":"2007-11-08T19:25:46.139+0000"},{"body":"Here's another approach to fixing this. The specific problem is that getPosition() is slow on the local filesystem (making a system call). We can fix this for all filesystems, as in the prior patch, since other filesystems might also have slow implementations of getPosition(), or we can fix it specifically for the local filesystem, as in this new patch, since other filesystems should be expected to implement this efficiently.\n\nJohan, does this version work for you too?\n\nWhat do folks prefer?","created":"2007-11-08T21:01:04.425+0000"},{"body":"I was going to write the following at the same time as above comment :\n\nbq. I am not sure about the class hierarchy. getPos() is a property of FSInputStream. So when a filesystem needs to return a FSDataInputStream, LocalFileSystem returns a PositionCache... similar to how returns a BufferedFSInputStream currently.\n\n","created":"2007-11-08T21:14:25.361+0000"},{"body":"Raghu: I take it you're voting for caching in local FS?\n\nAnother place we could cache the position is in BufferedFSInputStream. That's the place where we depend on getPos() being fast. When someone seeks, in order to check whether the seek is within the buffer, we need to know where the buffer is in the file. We currently call getPos() on the underlying stream, which is slow on the local filesystem impl, since it makes a system call. So is this a better place to cache? I started to code that, but it will take more code, since BufferedFSInputStream doesn't already override all the position-changing methods. So I'm currently leaning towards pushing the cache down into the local fs impl.\n","created":"2007-11-08T21:33:02.664+0000"},{"body":"LocalFileSystem is a ChecksumFileSystem. So the file position is cached in FSInputChecker. I still do not understand why taking off PositionCache causes such a huge performance degragation.","created":"2007-11-08T21:40:58.328+0000"},{"body":"> I take it you're voting for caching in local FS?\n\npretty much. I was thinking of on the lines of having PositionCache class similar to first patch that can be used by LocalFileSystem like this :\n\n{code} return new FSDataInputStream(new BufferedFSInputStream(\n new LocalFSFileInputStream(f), bufferSize)); {code}\nwould be replaced by :\n{code} return new FSDataInputStream(new BufferedFSInputStream(\n new PositionCache(new LocalFSFileInputStream(f)), bufferSize));{code}","created":"2007-11-08T21:57:14.420+0000"},{"body":"> So the file position is cached in FSInputChecker.\n\nThe stack traces I'm seeing show it being called as follows:\n{noformat}\norg.apache.hadoop.fs.RawLocalFileSystem$LocalFSFileInputStream.getPos(RawLocalFileSystem.java:81)\norg.apache.hadoop.fs.BufferedFSInputStream.getPos(BufferedFSInputStream.java:48)\norg.apache.hadoop.fs.FSDataInputStream.getPos(FSDataInputStream.java:41)\norg.apache.hadoop.fs.ChecksumFileSystem$ChecksumFSInputChecker.readChunk(ChecksumFileSystem.java:196)\n{noformat}\nEach such call now results in a system call, where before it was cached here too.\n","created":"2007-11-08T22:25:01.586+0000"},{"body":"Another version of patch that essentially rearranges first patch and is used only by LocalFileSystem.\n\nBtw, FSDataOutputStream also has its own PositionCache similar to first patch. I didn't change that.","created":"2007-11-08T23:51:05.216+0000"},{"body":"The patch looks good. But personally I'd like to cache the file position in BufferedFSInputStream. It removes one java file and reduces one layer of input stream.","created":"2007-11-09T00:56:51.025+0000"},{"body":"> The patch looks good. But personally I'd like to cache the file position in BufferedFSInputStream. It removes one java file and reduces one layer of input stream.\n\nIn that case, Doug's patch is the smallest I guess. Only correction needed there is that return value of 0 from fis.read() is a valid return and position should be incremented.","created":"2007-11-09T01:00:27.431+0000"},{"body":"If the cache is moved to local filesystem only I assume the dfs and other filesystems have a similar cache? Or do they not need one?\n\nPerformance with either of these three patches are pretty much identical for me so my only concern is what happens with other filesystems.","created":"2007-11-09T11:09:22.852+0000"},{"body":"Yes, DFS and any ChecksumFileSystem like LocalFileSystem all already cache file poition somewhere. That's why I took out CachedPoisition in FSDataInputStream. The problem is caused by RawLocalFileSystem not caching its file position.\n\nDoug and I talked over IM and he thinks it is better to put the cached file position in the highest level as possible so it can be shared. I feel it is better cached in BufferedFSInputStream.","created":"2007-11-09T17:58:41.423+0000"},{"body":"I took a look at BufferedFSInputStream. It is quite hard to implement file position cache there because its read methods are implement in its parent class java.io.BufferedInputStream. Now I dont have any strong opinion on where we should cache the file position.","created":"2007-11-09T18:15:35.198+0000"},{"body":"There was no strong consensus here, so we'll go with the smallest patch. Re-attaching that version now, so that Hudson will grab it.\n\n> Only correction needed there is that return value of 0 from fis.read() is a valid return and position should be incremented.\n\nIncrementing by zero is a no-op, so I don't see this as a bug.","created":"2007-11-13T19:47:35.456+0000"},{"body":">> Only correction needed there is that return value of 0 from fis.read() is a valid return and position should be incremented.\n\n> Incrementing by zero is a no-op, so I don't see this as a bug.\n\nwhen {{in.read()}} returns zero, it implies it read one byte whose value is zero. So position should be incremented by one. I think we should cancel the patch.","created":"2007-11-13T21:48:09.250+0000"},{"body":"> when in.read() returns zero, it implies it read one byte whose value is zero.\n\nOops. Good point. The fact that this passed unit tests shows that we never actually call read() on this stream, since it's always buffered, but still, it shouldn't have a buggy implementation. Thanks for catching that.","created":"2007-11-13T22:01:18.096+0000"},{"body":"New version that correctly handles the case when read() returns zero.","created":"2007-11-13T22:22:33.567+0000"},{"body":"I just committed this.","created":"2007-11-14T23:36:22.466+0000"}],"conversations":[{"body":"The PositionCache in FSDataInputStream seems to have been removed in HADOOP-1470. This causes for example MapFile.get usage to be extremely slow as the file position isn't cached in memory.","from":"reporter","subject":"PositionCache was removed from FSDataInputStream, causes extremely bad MapFile performance"},{"body":"The claim in particular (from IRC) is that random access to a MapFile stored on the LocalFileSystem has become much slower, and that much of the time is taken seeking the CRC file. Seeks within the current buffer should not require system calls, but perhaps that optimization was lost?\n\nI hacked TestArrayFile to test this, and ArrayFile.get(0) run 100,000 times took around 1 second in 0.13 and now takes over ten.\n","from":"developer"},{"body":"This patch brings back the PositionCache from 0.13.1. It appears the position cache previously did not keep track of long skip(long) or int read(). This has been added.\n\nPasses all the tests on 0.14.3 and has been tested on a 35 node cluster.","from":"developer"},{"body":"This restores the 0.13 performance for my benchmark. +1\n\nHairong, can you please review this, since you removed PositionCache as a part of HADOOP-1470? Thanks!","from":"developer"},{"body":"Here's another approach to fixing this. The specific problem is that getPosition() is slow on the local filesystem (making a system call). We can fix this for all filesystems, as in the prior patch, since other filesystems might also have slow implementations of getPosition(), or we can fix it specifically for the local filesystem, as in this new patch, since other filesystems should be expected to implement this efficiently.\n\nJohan, does this version work for you too?\n\nWhat do folks prefer?","from":"developer"},{"body":"I was going to write the following at the same time as above comment :\n\nbq. I am not sure about the class hierarchy. getPos() is a property of FSInputStream. So when a filesystem needs to return a FSDataInputStream, LocalFileSystem returns a PositionCache... similar to how returns a BufferedFSInputStream currently.\n\n","from":"developer"},{"body":"Raghu: I take it you're voting for caching in local FS?\n\nAnother place we could cache the position is in BufferedFSInputStream. That's the place where we depend on getPos() being fast. When someone seeks, in order to check whether the seek is within the buffer, we need to know where the buffer is in the file. We currently call getPos() on the underlying stream, which is slow on the local filesystem impl, since it makes a system call. So is this a better place to cache? I started to code that, but it will take more code, since BufferedFSInputStream doesn't already override all the position-changing methods. So I'm currently leaning towards pushing the cache down into the local fs impl.\n","from":"developer"},{"body":"LocalFileSystem is a ChecksumFileSystem. So the file position is cached in FSInputChecker. I still do not understand why taking off PositionCache causes such a huge performance degragation.","from":"developer"},{"body":"> I take it you're voting for caching in local FS?\n\npretty much. I was thinking of on the lines of having PositionCache class similar to first patch that can be used by LocalFileSystem like this :\n\n{code} return new FSDataInputStream(new BufferedFSInputStream(\n new LocalFSFileInputStream(f), bufferSize)); {code}\nwould be replaced by :\n{code} return new FSDataInputStream(new BufferedFSInputStream(\n new PositionCache(new LocalFSFileInputStream(f)), bufferSize));{code}","from":"developer"},{"body":"> So the file position is cached in FSInputChecker.\n\nThe stack traces I'm seeing show it being called as follows:\n{noformat}\norg.apache.hadoop.fs.RawLocalFileSystem$LocalFSFileInputStream.getPos(RawLocalFileSystem.java:81)\norg.apache.hadoop.fs.BufferedFSInputStream.getPos(BufferedFSInputStream.java:48)\norg.apache.hadoop.fs.FSDataInputStream.getPos(FSDataInputStream.java:41)\norg.apache.hadoop.fs.ChecksumFileSystem$ChecksumFSInputChecker.readChunk(ChecksumFileSystem.java:196)\n{noformat}\nEach such call now results in a system call, where before it was cached here too.\n","from":"developer"},{"body":"Another version of patch that essentially rearranges first patch and is used only by LocalFileSystem.\n\nBtw, FSDataOutputStream also has its own PositionCache similar to first patch. I didn't change that.","from":"developer"},{"body":"The patch looks good. But personally I'd like to cache the file position in BufferedFSInputStream. It removes one java file and reduces one layer of input stream.","from":"developer"},{"body":"> The patch looks good. But personally I'd like to cache the file position in BufferedFSInputStream. It removes one java file and reduces one layer of input stream.\n\nIn that case, Doug's patch is the smallest I guess. Only correction needed there is that return value of 0 from fis.read() is a valid return and position should be incremented.","from":"developer"},{"body":"If the cache is moved to local filesystem only I assume the dfs and other filesystems have a similar cache? Or do they not need one?\n\nPerformance with either of these three patches are pretty much identical for me so my only concern is what happens with other filesystems.","from":"developer"},{"body":"Yes, DFS and any ChecksumFileSystem like LocalFileSystem all already cache file poition somewhere. That's why I took out CachedPoisition in FSDataInputStream. The problem is caused by RawLocalFileSystem not caching its file position.\n\nDoug and I talked over IM and he thinks it is better to put the cached file position in the highest level as possible so it can be shared. I feel it is better cached in BufferedFSInputStream.","from":"developer"},{"body":"I took a look at BufferedFSInputStream. It is quite hard to implement file position cache there because its read methods are implement in its parent class java.io.BufferedInputStream. Now I dont have any strong opinion on where we should cache the file position.","from":"developer"},{"body":"There was no strong consensus here, so we'll go with the smallest patch. Re-attaching that version now, so that Hudson will grab it.\n\n> Only correction needed there is that return value of 0 from fis.read() is a valid return and position should be incremented.\n\nIncrementing by zero is a no-op, so I don't see this as a bug.","from":"developer"},{"body":">> Only correction needed there is that return value of 0 from fis.read() is a valid return and position should be incremented.\n\n> Incrementing by zero is a no-op, so I don't see this as a bug.\n\nwhen {{in.read()}} returns zero, it implies it read one byte whose value is zero. So position should be incremented by one. I think we should cancel the patch.","from":"developer"},{"body":"> when in.read() returns zero, it implies it read one byte whose value is zero.\n\nOops. Good point. The fact that this passed unit tests shows that we never actually call read() on this stream, since it's always buffered, but still, it shouldn't have a buggy implementation. Thanks for catching that.","from":"developer"},{"body":"New version that correctly handles the case when read() returns zero.","from":"developer"},{"body":"I just committed this.","from":"developer"}],"created":"2007-11-08T17:19:55.000+0000","description":"The PositionCache in FSDataInputStream seems to have been removed in HADOOP-1470. This causes for example MapFile.get usage to be extremely slow as the file position isn't cached in memory.","issue_id":"12382110","key":"HADOOP-2172","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2007-11-14T23:36:22.000+0000","role":"fixed_distractor","summary":"PositionCache was removed from FSDataInputStream, causes extremely bad MapFile performance"} {"case_id":"12382953","cluster":"DISTRACTOR-HADOOP-2245","comments":[{"body":"Attaching Adrian's patch https://issues.apache.org/jira/secure/attachment/12369042/HADOOP-1642-5.patch against this issue.","created":"2007-11-21T09:07:55.841+0000"},{"body":"I just committed this. Thanks, Adrian!","created":"2007-11-23T08:59:27.413+0000"},{"body":"I'm reopening this after reverting the patch.\n\nThis patch can be significantly improved, especially the test-cases. Specifically it copies code from TestJobControl into JobControlTestUtils, for re-use, and yet doesn't remove the copied code from TestJobControl. It also introduced new compiler warnings, which somehow weren't caught by the patch-process. \n\nOverall, my bad, I should have been more careful while committing this.","created":"2007-11-23T11:55:35.986+0000"},{"body":"I propose we commit the more crucial part of Adrian's patch i.e. the fix to {{LocalJobRunner}} (attached here) since that helps with the existing test-cases seem to fail and Adrian can fix the newer test-cases and submit that as a new patch. Thoughts?","created":"2007-11-23T12:01:51.776+0000"},{"body":"I have already removed the duplicate code from TestJobControl and am working on the compiler warnings as we speak so you can either go ahead with the smaller patch or wait.","created":"2007-11-23T12:10:00.351+0000"},{"body":"Sounds good, if you prefer so we can wait for the better patch. ","created":"2007-11-23T12:18:28.885+0000"},{"body":"Removed unused code in TestJobControl and made types more specific in TestJobControl to get rid of compiler warnings.","created":"2007-11-23T12:32:24.099+0000"},{"body":"Same as previous patch plus the change to LocalJobRunner.java","created":"2007-11-23T12:41:31.943+0000"},{"body":"Now with all discussed changes, tests passing etc. Code review please.","created":"2007-11-23T14:28:56.326+0000"},{"body":"New patch now available incoporating Arun's feedback.","created":"2007-11-23T14:33:15.665+0000"},{"body":"No idea why contrib tests fail on this, ran them a few times here and they go just fine.","created":"2007-11-23T17:50:31.450+0000"},{"body":"The patch seems to have generated new Findbugs warning(s) but the Findbugs warning-link is not accesible from jira. I am going to resubmit the patch again. Adrian, could you please take a look at the Findbugs warnings.","created":"2007-11-27T12:33:07.989+0000"},{"body":"Attempting to get hold of the Findbugs warnings.","created":"2007-11-27T12:33:50.666+0000"},{"body":"The findbugs error is probably related to \n\nhttps://issues.apache.org/jira/browse/HADOOP-2272\n\nwhich should now be fixed. So let's see what Hudson says... (all patches submitted since towards the end of last week that generated findbugs errors should also be resubmitted).","created":"2007-11-27T12:46:03.213+0000"},{"body":"Will resubmit to see if the javadoc issue goes away (javadoc fix as per HADOOP-2289)","created":"2007-11-27T18:59:17.227+0000"},{"body":"I just committed this. Thanks, Adrian!","created":"2007-11-29T11:31:42.496+0000"},{"body":"I think you forgot to commit the newly added class\n\nJobControlTestUtils\n\nThe trunk core tests won't compile until you do this....","created":"2007-11-29T11:47:48.784+0000"},{"body":"Done. Thanks for pointing this out.","created":"2007-11-29T12:16:29.845+0000"}],"conversations":[{"body":"The tests org.apache.hadoop.mapred.lib.aggregate.TestAggregates.testAggregates and org.apache.hadoop.record.TestRecordMR.testMapred fail intermittently with the problem \"/map_0000/file.out already exists\". There is a patch on HADOOP-1642 that should address the issue.\n","from":"reporter","subject":"TestRecordMR and TestAggregates fail once in a while"},{"body":"Attaching Adrian's patch https://issues.apache.org/jira/secure/attachment/12369042/HADOOP-1642-5.patch against this issue.","from":"developer"},{"body":"I just committed this. Thanks, Adrian!","from":"developer"},{"body":"I'm reopening this after reverting the patch.\n\nThis patch can be significantly improved, especially the test-cases. Specifically it copies code from TestJobControl into JobControlTestUtils, for re-use, and yet doesn't remove the copied code from TestJobControl. It also introduced new compiler warnings, which somehow weren't caught by the patch-process. \n\nOverall, my bad, I should have been more careful while committing this.","from":"developer"},{"body":"I propose we commit the more crucial part of Adrian's patch i.e. the fix to {{LocalJobRunner}} (attached here) since that helps with the existing test-cases seem to fail and Adrian can fix the newer test-cases and submit that as a new patch. Thoughts?","from":"developer"},{"body":"I have already removed the duplicate code from TestJobControl and am working on the compiler warnings as we speak so you can either go ahead with the smaller patch or wait.","from":"developer"},{"body":"Sounds good, if you prefer so we can wait for the better patch. ","from":"developer"},{"body":"Removed unused code in TestJobControl and made types more specific in TestJobControl to get rid of compiler warnings.","from":"developer"},{"body":"Same as previous patch plus the change to LocalJobRunner.java","from":"developer"},{"body":"Now with all discussed changes, tests passing etc. Code review please.","from":"developer"},{"body":"New patch now available incoporating Arun's feedback.","from":"developer"},{"body":"No idea why contrib tests fail on this, ran them a few times here and they go just fine.","from":"developer"},{"body":"The patch seems to have generated new Findbugs warning(s) but the Findbugs warning-link is not accesible from jira. I am going to resubmit the patch again. Adrian, could you please take a look at the Findbugs warnings.","from":"developer"},{"body":"Attempting to get hold of the Findbugs warnings.","from":"developer"},{"body":"The findbugs error is probably related to \n\nhttps://issues.apache.org/jira/browse/HADOOP-2272\n\nwhich should now be fixed. So let's see what Hudson says... (all patches submitted since towards the end of last week that generated findbugs errors should also be resubmitted).","from":"developer"},{"body":"Will resubmit to see if the javadoc issue goes away (javadoc fix as per HADOOP-2289)","from":"developer"},{"body":"I just committed this. Thanks, Adrian!","from":"developer"},{"body":"I think you forgot to commit the newly added class\n\nJobControlTestUtils\n\nThe trunk core tests won't compile until you do this....","from":"developer"},{"body":"Done. Thanks for pointing this out.","from":"developer"}],"created":"2007-11-21T07:12:46.000+0000","description":"The tests org.apache.hadoop.mapred.lib.aggregate.TestAggregates.testAggregates and org.apache.hadoop.record.TestRecordMR.testMapred fail intermittently with the problem \"/map_0000/file.out already exists\". There is a patch on HADOOP-1642 that should address the issue.\n","issue_id":"12382953","key":"HADOOP-2245","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2007-11-29T11:31:42.000+0000","role":"fixed_distractor","summary":"TestRecordMR and TestAggregates fail once in a while"} {"case_id":"12383273","cluster":"DISTRACTOR-HADOOP-2285","comments":[{"body":"A better implementation, very similar to the one you might find in BufferedReader (but with bytes instead of chars), would be the best solution. The basic issue stems from the fact that read() is called on a per byte basis in the current readLine implementation rather than using the more efficient read([], o, l) call thats available and then manipulating the array.","created":"2007-11-27T05:05:46.555+0000"},{"body":"Ok, here is a patch that does:\n 1. Avoids encoding the data as a string and stores it directly in the Text object, which avoids the encode/decode cycle.\n 2. Merges the buffer and the readLine code, so that it can use direct buffer access.\n 3. Adds a new method to Text to append bytes.\n 4. Adds a new method to Text to clear back to an empty string.\n 5. Adds test cases for the new functionality.\n\nUsing the benchmark on HADOOP-2406, I see a 3x speed up using TextInputFormat. (32 seconds down to 11). This should be a big win for any jobs that scan a lot of text data.","created":"2007-12-27T00:18:28.627+0000"},{"body":"+1 Looks good; I'm seeing better than a 3x speedup.","created":"2008-01-07T23:59:56.823+0000"},{"body":"Minor nit: this patch removes a public constructor rather than add a new one:\n\n{noformat}\n- public LineRecordReader(InputStream in, long offset, long endOffset)\n+ public LineRecordReader(InputStream in, long offset, long endOffset,\n+ Configuration job)\n{noformat}","created":"2008-01-08T09:04:26.716+0000"},{"body":"Attaching a simple fix to my previous comment on Owen's behalf...","created":"2008-01-08T09:28:36.662+0000"},{"body":"+1; speed was increased.","created":"2008-01-08T09:49:13.259+0000"},{"body":"Use LineReader.DEFAULT_BUFFER_SIZE for conf-less LineRecordReader","created":"2008-01-08T20:25:35.956+0000"},{"body":"I just committed this. Thanks Owen!","created":"2008-01-08T23:30:16.196+0000"}],"conversations":[{"body":"The LineRecordReader reads from the source byte by byte, which seems to be half as fast as if the readLine method was defined on the memory buffer directly instead of as an InputStream.","from":"reporter","subject":"TextInputFormat is slow compared to reading files."},{"body":"A better implementation, very similar to the one you might find in BufferedReader (but with bytes instead of chars), would be the best solution. The basic issue stems from the fact that read() is called on a per byte basis in the current readLine implementation rather than using the more efficient read([], o, l) call thats available and then manipulating the array.","from":"developer"},{"body":"Ok, here is a patch that does:\n 1. Avoids encoding the data as a string and stores it directly in the Text object, which avoids the encode/decode cycle.\n 2. Merges the buffer and the readLine code, so that it can use direct buffer access.\n 3. Adds a new method to Text to append bytes.\n 4. Adds a new method to Text to clear back to an empty string.\n 5. Adds test cases for the new functionality.\n\nUsing the benchmark on HADOOP-2406, I see a 3x speed up using TextInputFormat. (32 seconds down to 11). This should be a big win for any jobs that scan a lot of text data.","from":"developer"},{"body":"+1 Looks good; I'm seeing better than a 3x speedup.","from":"developer"},{"body":"Minor nit: this patch removes a public constructor rather than add a new one:\n\n{noformat}\n- public LineRecordReader(InputStream in, long offset, long endOffset)\n+ public LineRecordReader(InputStream in, long offset, long endOffset,\n+ Configuration job)\n{noformat}","from":"developer"},{"body":"Attaching a simple fix to my previous comment on Owen's behalf...","from":"developer"},{"body":"+1; speed was increased.","from":"developer"},{"body":"Use LineReader.DEFAULT_BUFFER_SIZE for conf-less LineRecordReader","from":"developer"},{"body":"I just committed this. Thanks Owen!","from":"developer"}],"created":"2007-11-26T22:44:18.000+0000","description":"The LineRecordReader reads from the source byte by byte, which seems to be half as fast as if the readLine method was defined on the memory buffer directly instead of as an InputStream.","issue_id":"12383273","key":"HADOOP-2285","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-01-08T23:30:28.000+0000","role":"fixed_distractor","summary":"TextInputFormat is slow compared to reading files."} {"case_id":"12344080","cluster":"DISTRACTOR-HADOOP-286","comments":[{"body":"It looks like the following scenario leads to this exception.\nLEASE_PERIOD = 60 sec is a global constants defining for how long a lease is issued.\nDFSClient.LeaseChecker renews this client leases every 30 sec = LEASE_PERIOD/2.\nIf the renewLease() fails then the client retries to renew every second.\nOne of the most popular reasons the renewLease() fails is because it timeouts SocketTimeoutException.\nThis happens when the namenode is busy, which is not unusual since we lock it for each operation.\nThe socket timeout is defined by the config parameter \"ipc.client.timeout\", which is set to 60 sec in\nhadoop-default.xml That means that the renewLease() can last up to 60 seconds and the lease will\nexpire the next time the client tries to renew it, which could be up to 90 seconds after the lease was\ncreated or renewed last time.\nSo there are 2 simple solutions to the problem:\n1) to increase LEASE_PERIOD\n2) to decrease ipc.client.timeout\n\nA related problem is that DFSClient sends lease renew requests no matter what every 30 seconds \nor less. It looks like the DFSClient has enough information to send renew messages only if it really \nholds a lease. A simple solution would be avoid calling renewLease() when \nDFSClient.pendingCreates is empty.\nThis could substantially decrease overall net traffic for map/reduce.\n\n","created":"2006-07-19T01:21:16.000+0000"},{"body":"This is a very simple patch that renews leases only when pendingCreates is not empty.\nThis prevents the client from sending lease renewal messages when the client\nis not writing into dfs, say just reading or doing local stuff.\nThis should make the name node less busy.\n\nI tried to change the ipc.client.timeout from 60 secs to 20 secs.\nOn my 3 node cluster everything worked fine.\nOn a large cluster the timeout was changed only for the DFSClient.\nThe LeaseExpiredException does not appear anymore.\nBut we need more statistics on that, especially with slower networks.\nThe ipc timeout is global for all ipc connections, so if we make it\nsmaller there is a risk that long lasting operations like block transfers\nwill start to timeout. I haven't seen it.\nIf anybody is willing to try 20 sec ipc timeout please post the results.\nFailing early, and retrying might make things faster in general.","created":"2006-07-20T22:00:06.000+0000"},{"body":"+1 for not requesting a lease unless a write operation is required (i.e. the patch).","created":"2006-09-06T19:06:55.000+0000"},{"body":"Patch for not renewing leases when pendingCreates is empty.\nThis is a scalability issue.","created":"2006-09-06T20:32:05.000+0000"},{"body":"I just committed this. Thanks, Konstantin!","created":"2006-09-06T22:36:48.000+0000"}],"conversations":[{"body":"\nLoading local files to dfs through hadoop dfs -copyFromLocal failed due to the following exception:\n\ncopyFromLocal: org.apache.hadoop.dfs.LeaseExpiredException: No lease on output_crawled.1.txt\n at org.apache.hadoop.dfs.FSNamesystem.getAdditionalBlock(FSNamesystem.java:414)\n at org.apache.hadoop.dfs.NameNode.addBlock(NameNode.java:190)\n at sun.reflect.GeneratedMethodAccessor9.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:585)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:243)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:231)\n","from":"reporter","subject":"copyFromLocal throws LeaseExpiredException"},{"body":"It looks like the following scenario leads to this exception.\nLEASE_PERIOD = 60 sec is a global constants defining for how long a lease is issued.\nDFSClient.LeaseChecker renews this client leases every 30 sec = LEASE_PERIOD/2.\nIf the renewLease() fails then the client retries to renew every second.\nOne of the most popular reasons the renewLease() fails is because it timeouts SocketTimeoutException.\nThis happens when the namenode is busy, which is not unusual since we lock it for each operation.\nThe socket timeout is defined by the config parameter \"ipc.client.timeout\", which is set to 60 sec in\nhadoop-default.xml That means that the renewLease() can last up to 60 seconds and the lease will\nexpire the next time the client tries to renew it, which could be up to 90 seconds after the lease was\ncreated or renewed last time.\nSo there are 2 simple solutions to the problem:\n1) to increase LEASE_PERIOD\n2) to decrease ipc.client.timeout\n\nA related problem is that DFSClient sends lease renew requests no matter what every 30 seconds \nor less. It looks like the DFSClient has enough information to send renew messages only if it really \nholds a lease. A simple solution would be avoid calling renewLease() when \nDFSClient.pendingCreates is empty.\nThis could substantially decrease overall net traffic for map/reduce.\n\n","from":"developer"},{"body":"This is a very simple patch that renews leases only when pendingCreates is not empty.\nThis prevents the client from sending lease renewal messages when the client\nis not writing into dfs, say just reading or doing local stuff.\nThis should make the name node less busy.\n\nI tried to change the ipc.client.timeout from 60 secs to 20 secs.\nOn my 3 node cluster everything worked fine.\nOn a large cluster the timeout was changed only for the DFSClient.\nThe LeaseExpiredException does not appear anymore.\nBut we need more statistics on that, especially with slower networks.\nThe ipc timeout is global for all ipc connections, so if we make it\nsmaller there is a risk that long lasting operations like block transfers\nwill start to timeout. I haven't seen it.\nIf anybody is willing to try 20 sec ipc timeout please post the results.\nFailing early, and retrying might make things faster in general.","from":"developer"},{"body":"+1 for not requesting a lease unless a write operation is required (i.e. the patch).","from":"developer"},{"body":"Patch for not renewing leases when pendingCreates is empty.\nThis is a scalability issue.","from":"developer"},{"body":"I just committed this. Thanks, Konstantin!","from":"developer"}],"created":"2006-06-07T23:50:34.000+0000","description":"\nLoading local files to dfs through hadoop dfs -copyFromLocal failed due to the following exception:\n\ncopyFromLocal: org.apache.hadoop.dfs.LeaseExpiredException: No lease on output_crawled.1.txt\n at org.apache.hadoop.dfs.FSNamesystem.getAdditionalBlock(FSNamesystem.java:414)\n at org.apache.hadoop.dfs.NameNode.addBlock(NameNode.java:190)\n at sun.reflect.GeneratedMethodAccessor9.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:585)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:243)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:231)\n","issue_id":"12344080","key":"HADOOP-286","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-09-06T22:36:48.000+0000","role":"fixed_distractor","summary":"copyFromLocal throws LeaseExpiredException"} {"case_id":"12392505","cluster":"DISTRACTOR-HADOOP-3113","comments":[{"body":"I volunteer to exercise any patch posted (Smile)\n\nCan the configuration be specified on a per-file basis? If not, this approach has less value since hbase is rarely the only user of an HDFS installation.\n\nWhats the default config. for block report periodicity vs. lease timeout currently? (Is this dfs.blockreport.intervalMsec -- 1hr -- vs. which config?) This config. could not be done per file, right? It'd be a global config?\n\nThanks Dhruba","created":"2008-03-28T21:20:53.529+0000"},{"body":"I can make the configurable to be per file, but maybe it makes more sense to make it applicable to the entire system. The reason being that datanodes do not know much about the name of the HDFS file that a block belongs to. To make this configurable \"per file\" would need lots of protocol change.\n\nThe default for block report is 1 hour and the default for lease timeout is 1 hour too. This needs to be changed so that block reports are sent evenry 30 minutes.\n\nIf a client dies while writing to the last block of that file, that block is not yet part of the blocksmap in the namenode. (A block gets inserted in the blocksmap when a complete block is received by the datanode and it sends a blockReceived message to the namenode). If the lease for this file on the namenode expires before the block report from the datanode arrives, then the namenode will erroneously think that no datanodes have a copy of that block. As part of lease recovery, the namenode will delete the last block of the file because it has no entry in the blocksMap. To prevent this from occuring, the block report periodicity should be set to 30 minutes.","created":"2008-03-30T06:03:26.070+0000"},{"body":"This patch prevents the datanodes from creating temporary files for blocks that are being written to.\n\nThe configuration parameter dfs.blockreport.intervalMsec should be changed from a default value of 3600000 to 1800000.\n","created":"2008-03-30T08:01:30.435+0000"},{"body":"> dhruba borthakur - 29/Mar/08 11:03 PM\n> I can make the configurable to be per file, but maybe it makes more sense to make it applicable to the\n> entire system. The reason being that datanodes do not know much about the name of the HDFS file\n> that a block belongs to. To make this configurable \"per file\" would need lots of protocol change.\n\nI don't think it needs to be per file. Aside from our redo log, other files are written and then immediately\nclosed and re-opened for read.\n\n> If a client dies while writing to the last block of that file, that block is not yet part of the blocksmap in the\n> namenode. (A block gets inserted in the blocksmap when a complete block is received by the datanode\n> and it sends a blockReceived message to the namenode). If the lease for this file on the namenode\n> expires before the block report from the datanode arrives, then the namenode will erroneously think\n> that no datanodes have a copy of that block. As part of lease recovery, the namenode will delete the\n> last block of the file because it has no entry in the blocksMap. To prevent this from occuring, the block\n> report periodicity should be set to 30 minutes.\n\nI think this is ok, but let me give a scenario to verify that my understanding is correct.\n\nWe open our redo log and flush it either every N seconds or after M records have been written.\nIf the process writing the log crashes, we will notice much sooner than the file lease timeout.\nAt that point another process should be able to open the file for read, and all flushed data \nwill be visible, unflushed data will not. Since the amount of unflushed data should be small\nthe amount of data lost should be minimal. Once the redo log has been read and processed,\nthe file will be deleted by the process reading the file.\n\nIf this is how this patch works, +1.","created":"2008-03-30T16:49:22.086+0000"},{"body":"Hi Jim, There is one issue that you pointed out. The patch, as it currently stands, will not behave the way you want. If the process writing the log dies, the last block is not fully written to datanode(s). This means that the namenode does not (yet) know the block locations of the last block. This will be available to the namenode only at the next block report from the datanode(s). If you happen to reopen the file before the next block report arrives you will not see the last block. Let me see if I can come up with something for this one.","created":"2008-03-31T17:09:47.875+0000"},{"body":"> I can make the configurable to be per file, but maybe it makes more sense to make it applicable to the entire system. \n\nWhen the setting is not per file:\n\n1. HBase often is new-kid on the block and but one of the many users of an HDFS install. HBase installers may find it difficult convincing HDFS admins they need to make the change.\n2. If per-file, HBase can manage the configuration. Otherwise, its a two-step process. HBase installers may plain forget.\n3. Minor-point: HBase needs append only for its Write-Ahead Log; nowhere else.\n\nIf its an 'entire system' setting, will it require an HDFS restart to take effect?\n\nSounds like we should set the blocksize for our Write-Ahead Log to be a good deal smaller than default.\n\nI took a quick look at the patch.\n\nThe hadoop-default.xml entry description is all on one line. You might want to break it up. Also, is there a downside to setting the dfs.datanode.skipTmpFile flag (Reading the description, in my head I'm thinking there must be or why even bother with this configuration?)\n\nOtherwise, patch looks good to me. \n\nDo you have a suggestion for a test I might run to exercise this new functionality?\n\nThanks Dhruba\n\n","created":"2008-04-01T23:17:38.206+0000"},{"body":"Hi Stack, thanks for your comments. I am still trying to figure out a way to address Jim's earlier requirement.... if the writer dies, how to make the file available for writing without waiting for the next block report. Let me see if I can come up with something to address this issue.","created":"2008-04-07T07:57:49.686+0000"},{"body":"There isn't any short-cut for this one. We need lease recovery HADOOP-3310 as a pre-requisite.","created":"2008-04-28T21:39:19.713+0000"},{"body":"This is the first version of the patch. When the client encounters an error while writing to a block, it invokes generation-stamp recovery on the primary datanode.","created":"2008-05-24T08:21:25.620+0000"},{"body":"This patch moves all tmp files to the real data directory on a datanode restart.","created":"2008-06-03T19:28:36.025+0000"},{"body":"The latest patch keeps blocks that are being written in the tmp directory. However, when a block is finalized it moves into the real block directory. Also, at datanode restart, all blocks from the tmp directory move to the real block directory. It is prudent to keep the blocks in the tmp directory while they are being updated because some datanode-local processing (e.g. CRC validation) might need to occur when they get moved from tmp dir to real block dir (at datanode restart).","created":"2008-06-03T19:59:31.915+0000"},{"body":"Fixed a unit test TestInterDatanodeProtocol that was trying to finalize a block inspite of the fact that the file was already closed.","created":"2008-06-03T21:06:52.034+0000"},{"body":"Patch looks good. Some minor comments:\n\n- In INodeFileUnderConstruction.setTargets(...), add \"this.primaryNodeIndex = -1;\"\n\n- Add private to INodeFileUnderConstruction.targets. Then, call pendingFile.setTargets(...) in FSNamesystem.getAdditionalBlock(...).\n","created":"2008-06-03T22:16:21.832+0000"},{"body":"Incorporated Nicholas' code review comments.","created":"2008-06-04T06:27:37.453+0000"},{"body":"+1 codes look good","created":"2008-06-04T17:31:53.333+0000"},{"body":"I just committed this.","created":"2008-06-04T17:54:03.985+0000"},{"body":"> An application can invoke sync on the FSDataOutputStream to really, really persist data in HDFS!\n\nHot dog!","created":"2008-06-04T21:24:01.357+0000"}],"conversations":[{"body":"DFSOutputStream has a method called flush() that persists block locations on the namenode and sends all outstanding data to all datanodes in the pipeline. However, this data goes to the tmp file on the datanode(s). When the block is closed, the tmp files is renamed to be the real block file. If the datanode(s) dies before the block is compete, then entire block is lost. This behaviour wil be fixed in HADOOP-1700.\n\nHowever, in the short term, a configuration paramater can be used to allow datanodes to write to the real block file directly, thereby avoiding writing to the tmp file. This means that data that is flushed successfully by a client does not get lost even if the datanode(s) or client dies.\n\nThe Namenode already has code to pick the largest replica (if multiple datanodes have different sizes of this block). Also, the namenode has code to not trigger replication request if the file is still being written to.\n\nThe only caveat that I can think of is that the block report periodicity should be much much smaller that the lease timeout period. A block report adds the being-written-to blocks to the blocksMap thereby avoiding any cleanup that a lease expiry processing might have otherwise done.\n\nNot all requirements specified by HADOOP-1700 are supported by this approach, but it could still be helpful (in the short term) for a wide range of applications.\n\n\n\n","from":"reporter","subject":"DFSOututStream.flush() should flush data to real block file on DataNode."},{"body":"I volunteer to exercise any patch posted (Smile)\n\nCan the configuration be specified on a per-file basis? If not, this approach has less value since hbase is rarely the only user of an HDFS installation.\n\nWhats the default config. for block report periodicity vs. lease timeout currently? (Is this dfs.blockreport.intervalMsec -- 1hr -- vs. which config?) This config. could not be done per file, right? It'd be a global config?\n\nThanks Dhruba","from":"developer"},{"body":"I can make the configurable to be per file, but maybe it makes more sense to make it applicable to the entire system. The reason being that datanodes do not know much about the name of the HDFS file that a block belongs to. To make this configurable \"per file\" would need lots of protocol change.\n\nThe default for block report is 1 hour and the default for lease timeout is 1 hour too. This needs to be changed so that block reports are sent evenry 30 minutes.\n\nIf a client dies while writing to the last block of that file, that block is not yet part of the blocksmap in the namenode. (A block gets inserted in the blocksmap when a complete block is received by the datanode and it sends a blockReceived message to the namenode). If the lease for this file on the namenode expires before the block report from the datanode arrives, then the namenode will erroneously think that no datanodes have a copy of that block. As part of lease recovery, the namenode will delete the last block of the file because it has no entry in the blocksMap. To prevent this from occuring, the block report periodicity should be set to 30 minutes.","from":"developer"},{"body":"This patch prevents the datanodes from creating temporary files for blocks that are being written to.\n\nThe configuration parameter dfs.blockreport.intervalMsec should be changed from a default value of 3600000 to 1800000.\n","from":"developer"},{"body":"> dhruba borthakur - 29/Mar/08 11:03 PM\n> I can make the configurable to be per file, but maybe it makes more sense to make it applicable to the\n> entire system. The reason being that datanodes do not know much about the name of the HDFS file\n> that a block belongs to. To make this configurable \"per file\" would need lots of protocol change.\n\nI don't think it needs to be per file. Aside from our redo log, other files are written and then immediately\nclosed and re-opened for read.\n\n> If a client dies while writing to the last block of that file, that block is not yet part of the blocksmap in the\n> namenode. (A block gets inserted in the blocksmap when a complete block is received by the datanode\n> and it sends a blockReceived message to the namenode). If the lease for this file on the namenode\n> expires before the block report from the datanode arrives, then the namenode will erroneously think\n> that no datanodes have a copy of that block. As part of lease recovery, the namenode will delete the\n> last block of the file because it has no entry in the blocksMap. To prevent this from occuring, the block\n> report periodicity should be set to 30 minutes.\n\nI think this is ok, but let me give a scenario to verify that my understanding is correct.\n\nWe open our redo log and flush it either every N seconds or after M records have been written.\nIf the process writing the log crashes, we will notice much sooner than the file lease timeout.\nAt that point another process should be able to open the file for read, and all flushed data \nwill be visible, unflushed data will not. Since the amount of unflushed data should be small\nthe amount of data lost should be minimal. Once the redo log has been read and processed,\nthe file will be deleted by the process reading the file.\n\nIf this is how this patch works, +1.","from":"developer"},{"body":"Hi Jim, There is one issue that you pointed out. The patch, as it currently stands, will not behave the way you want. If the process writing the log dies, the last block is not fully written to datanode(s). This means that the namenode does not (yet) know the block locations of the last block. This will be available to the namenode only at the next block report from the datanode(s). If you happen to reopen the file before the next block report arrives you will not see the last block. Let me see if I can come up with something for this one.","from":"developer"},{"body":"> I can make the configurable to be per file, but maybe it makes more sense to make it applicable to the entire system. \n\nWhen the setting is not per file:\n\n1. HBase often is new-kid on the block and but one of the many users of an HDFS install. HBase installers may find it difficult convincing HDFS admins they need to make the change.\n2. If per-file, HBase can manage the configuration. Otherwise, its a two-step process. HBase installers may plain forget.\n3. Minor-point: HBase needs append only for its Write-Ahead Log; nowhere else.\n\nIf its an 'entire system' setting, will it require an HDFS restart to take effect?\n\nSounds like we should set the blocksize for our Write-Ahead Log to be a good deal smaller than default.\n\nI took a quick look at the patch.\n\nThe hadoop-default.xml entry description is all on one line. You might want to break it up. Also, is there a downside to setting the dfs.datanode.skipTmpFile flag (Reading the description, in my head I'm thinking there must be or why even bother with this configuration?)\n\nOtherwise, patch looks good to me. \n\nDo you have a suggestion for a test I might run to exercise this new functionality?\n\nThanks Dhruba\n\n","from":"developer"},{"body":"Hi Stack, thanks for your comments. I am still trying to figure out a way to address Jim's earlier requirement.... if the writer dies, how to make the file available for writing without waiting for the next block report. Let me see if I can come up with something to address this issue.","from":"developer"},{"body":"There isn't any short-cut for this one. We need lease recovery HADOOP-3310 as a pre-requisite.","from":"developer"},{"body":"This is the first version of the patch. When the client encounters an error while writing to a block, it invokes generation-stamp recovery on the primary datanode.","from":"developer"},{"body":"This patch moves all tmp files to the real data directory on a datanode restart.","from":"developer"},{"body":"The latest patch keeps blocks that are being written in the tmp directory. However, when a block is finalized it moves into the real block directory. Also, at datanode restart, all blocks from the tmp directory move to the real block directory. It is prudent to keep the blocks in the tmp directory while they are being updated because some datanode-local processing (e.g. CRC validation) might need to occur when they get moved from tmp dir to real block dir (at datanode restart).","from":"developer"},{"body":"Fixed a unit test TestInterDatanodeProtocol that was trying to finalize a block inspite of the fact that the file was already closed.","from":"developer"},{"body":"Patch looks good. Some minor comments:\n\n- In INodeFileUnderConstruction.setTargets(...), add \"this.primaryNodeIndex = -1;\"\n\n- Add private to INodeFileUnderConstruction.targets. Then, call pendingFile.setTargets(...) in FSNamesystem.getAdditionalBlock(...).\n","from":"developer"},{"body":"Incorporated Nicholas' code review comments.","from":"developer"},{"body":"+1 codes look good","from":"developer"},{"body":"I just committed this.","from":"developer"},{"body":"> An application can invoke sync on the FSDataOutputStream to really, really persist data in HDFS!\n\nHot dog!","from":"developer"}],"created":"2008-03-27T22:32:32.000+0000","description":"DFSOutputStream has a method called flush() that persists block locations on the namenode and sends all outstanding data to all datanodes in the pipeline. However, this data goes to the tmp file on the datanode(s). When the block is closed, the tmp files is renamed to be the real block file. If the datanode(s) dies before the block is compete, then entire block is lost. This behaviour wil be fixed in HADOOP-1700.\n\nHowever, in the short term, a configuration paramater can be used to allow datanodes to write to the real block file directly, thereby avoiding writing to the tmp file. This means that data that is flushed successfully by a client does not get lost even if the datanode(s) or client dies.\n\nThe Namenode already has code to pick the largest replica (if multiple datanodes have different sizes of this block). Also, the namenode has code to not trigger replication request if the file is still being written to.\n\nThe only caveat that I can think of is that the block report periodicity should be much much smaller that the lease timeout period. A block report adds the being-written-to blocks to the blocksMap thereby avoiding any cleanup that a lease expiry processing might have otherwise done.\n\nNot all requirements specified by HADOOP-1700 are supported by this approach, but it could still be helpful (in the short term) for a wide range of applications.\n\n\n\n","issue_id":"12392505","key":"HADOOP-3113","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-06-04T17:54:04.000+0000","role":"fixed_distractor","summary":"DFSOututStream.flush() should flush data to real block file on DataNode."} {"case_id":"12392618","cluster":"DISTRACTOR-HADOOP-3131","comments":[{"body":"screenshot of broken progress counters","created":"2008-03-29T01:22:32.950+0000"},{"body":"The problem was that SequenceFile.Sorter.MergeQueue calculates progress as (total size of keys and values read) / (total size of files to be merged on disk). When a file is compressed, the file size is much smaller than the combined sizes of the keys and values. In fact, there is also a problem when compression is turned off - the code returns progress less than 100% because it does not count bytes in the file that are not part of keys and values, such as the header and length fields. This patch changes MergeQueue to use the position in the input stream to calculate number of bytes read from disk and divide that by the total amount of data to be merged.","created":"2008-06-10T23:32:34.358+0000"},{"body":"Patch that changes progress computation to use position in input stream.","created":"2008-06-10T23:35:15.146+0000"},{"body":"Matei, unfortunately this patch won't work for hadoop-0.18.* since HADOOP-2095 changed Map-Reduce to use a new file format called IFile for intermediate sort/merge...","created":"2008-06-11T00:02:06.019+0000"},{"body":"interesting. can a patch be submitted for 17.* as well? (generally seems useful and 17 has a few minor releases to come ..)","created":"2008-06-11T00:45:38.619+0000"},{"body":"Looking at the patch submitted for HADOOP-2095, it seems that it has the same problem (by doing totalBytesProcessed += (key.getLength()-key.getPosition()) + \n (value.getLength()-value.getPosition())). I can submit a separate patch against 18 to fix that, but it would also be good to place this in 17 because 18 is not getting released for a while.\n","created":"2008-06-11T18:09:59.346+0000"},{"body":"Is anyone still interested in me writing a patch for 0.18, or have you fixed this issue in IFile already?","created":"2008-06-24T22:45:03.078+0000"},{"body":"Here's a patch that fixes IFile too.","created":"2008-07-09T17:59:41.276+0000"},{"body":"New patch against trunk to fix the previous merge conflict.","created":"2008-07-14T19:00:28.034+0000"},{"body":"New patch fixing findbugs problem.","created":"2008-07-19T00:40:24.360+0000"},{"body":"Matei, am I missing something here or should IFile.Reader.getPosition return in.getPosition and not 'bytesRead' which is decompressed bytes' count?","created":"2008-07-21T22:11:59.673+0000"},{"body":"Thanks for pointing that out.. here's a new patch that fixes it and includes a test with compression for IFile.","created":"2008-07-21T22:50:03.227+0000"},{"body":"Fixes compile error.","created":"2008-07-22T17:20:20.073+0000"},{"body":"Matei, a minor nit: looking through your patch I realised that 'progress reporting' based on rawIn.getPosition might actually off due to the buffering done by IFile.Reader (see IFile.Reader.readData). Should we fix it too? (Of course, it's only a temporary glitch i.e. till the buffered data is consumed and won't be too bad...)","created":"2008-07-22T18:26:55.252+0000"},{"body":"I'm not sure it's worth fixing that, because we don't need perfect progress reporting, just a rough guide to tell whether a task is doing something, and what rate it's working at. With compression enabled, it would also be difficult to figure out which spot in the buffer corresponds to which byte in the compressed file, if we were to use the position in the buffer to figure out progress.","created":"2008-07-22T18:31:56.905+0000"},{"body":"I do not think that the test failure has anything to do with this patch, but will resubmit the patch to make sure that this is true.","created":"2008-07-23T13:56:36.495+0000"},{"body":"Resubmitting.","created":"2008-07-23T13:56:48.799+0000"},{"body":"I don't think the TestCLI failure has anything to do with this: HADOOP-3809 should address it.","created":"2008-07-23T15:31:02.035+0000"},{"body":"Matei, sorry I missed this piece the first time around:\n\n{noformat}\n+ for (Segment s: segmentsToMerge) {\n+ totalBytesProcessed += s.getPosition(); // Count initial bytes read\n+ }\n+ if (totalBytes != 0) {\n+ mergeProgress.set(totalBytesProcessed * progPerByte);\n+ } else {\n+ mergeProgress.set(1.0f);\n+ }\n{noformat}\n\nAt best it reports progress slightly early (i.e. before the final merge begins) and at worst it provides completely wrong progress value during the merging of intermediate map-outputs since all output for all reduces is in a single file. Hence {{s.getPosition}} is hopelessly off as a measure of merge progress... I vote we just do away with that block.","created":"2008-07-24T08:43:36.766+0000"},{"body":"New patch fixing issues that Arun pointed out.","created":"2008-07-24T18:30:22.576+0000"},{"body":"Hi arun, would you like to code-review this patch one final time? Thanks.","created":"2008-07-30T23:35:07.494+0000"},{"body":"I just committed this. Thanks, Matei! (It's been a long-drawn affair, appreciate your patience!)","created":"2008-07-31T00:06:07.681+0000"},{"body":"Awesome, thanks!","created":"2008-07-31T20:34:30.224+0000"}],"conversations":[{"body":"Enabling map output compression and setting the compression type to BLOCK causes the progress counters during the reduce to go crazy and report progress counts over 100%.\n\nThis is problematic for speculative execution because it thinks the tasks are doing fine.","from":"reporter","subject":"enabling BLOCK compression for map outputs breaks the reduce progress counters"},{"body":"screenshot of broken progress counters","from":"developer"},{"body":"The problem was that SequenceFile.Sorter.MergeQueue calculates progress as (total size of keys and values read) / (total size of files to be merged on disk). When a file is compressed, the file size is much smaller than the combined sizes of the keys and values. In fact, there is also a problem when compression is turned off - the code returns progress less than 100% because it does not count bytes in the file that are not part of keys and values, such as the header and length fields. This patch changes MergeQueue to use the position in the input stream to calculate number of bytes read from disk and divide that by the total amount of data to be merged.","from":"developer"},{"body":"Patch that changes progress computation to use position in input stream.","from":"developer"},{"body":"Matei, unfortunately this patch won't work for hadoop-0.18.* since HADOOP-2095 changed Map-Reduce to use a new file format called IFile for intermediate sort/merge...","from":"developer"},{"body":"interesting. can a patch be submitted for 17.* as well? (generally seems useful and 17 has a few minor releases to come ..)","from":"developer"},{"body":"Looking at the patch submitted for HADOOP-2095, it seems that it has the same problem (by doing totalBytesProcessed += (key.getLength()-key.getPosition()) + \n (value.getLength()-value.getPosition())). I can submit a separate patch against 18 to fix that, but it would also be good to place this in 17 because 18 is not getting released for a while.\n","from":"developer"},{"body":"Is anyone still interested in me writing a patch for 0.18, or have you fixed this issue in IFile already?","from":"developer"},{"body":"Here's a patch that fixes IFile too.","from":"developer"},{"body":"New patch against trunk to fix the previous merge conflict.","from":"developer"},{"body":"New patch fixing findbugs problem.","from":"developer"},{"body":"Matei, am I missing something here or should IFile.Reader.getPosition return in.getPosition and not 'bytesRead' which is decompressed bytes' count?","from":"developer"},{"body":"Thanks for pointing that out.. here's a new patch that fixes it and includes a test with compression for IFile.","from":"developer"},{"body":"Fixes compile error.","from":"developer"},{"body":"Matei, a minor nit: looking through your patch I realised that 'progress reporting' based on rawIn.getPosition might actually off due to the buffering done by IFile.Reader (see IFile.Reader.readData). Should we fix it too? (Of course, it's only a temporary glitch i.e. till the buffered data is consumed and won't be too bad...)","from":"developer"},{"body":"I'm not sure it's worth fixing that, because we don't need perfect progress reporting, just a rough guide to tell whether a task is doing something, and what rate it's working at. With compression enabled, it would also be difficult to figure out which spot in the buffer corresponds to which byte in the compressed file, if we were to use the position in the buffer to figure out progress.","from":"developer"},{"body":"I do not think that the test failure has anything to do with this patch, but will resubmit the patch to make sure that this is true.","from":"developer"},{"body":"Resubmitting.","from":"developer"},{"body":"I don't think the TestCLI failure has anything to do with this: HADOOP-3809 should address it.","from":"developer"},{"body":"Matei, sorry I missed this piece the first time around:\n\n{noformat}\n+ for (Segment s: segmentsToMerge) {\n+ totalBytesProcessed += s.getPosition(); // Count initial bytes read\n+ }\n+ if (totalBytes != 0) {\n+ mergeProgress.set(totalBytesProcessed * progPerByte);\n+ } else {\n+ mergeProgress.set(1.0f);\n+ }\n{noformat}\n\nAt best it reports progress slightly early (i.e. before the final merge begins) and at worst it provides completely wrong progress value during the merging of intermediate map-outputs since all output for all reduces is in a single file. Hence {{s.getPosition}} is hopelessly off as a measure of merge progress... I vote we just do away with that block.","from":"developer"},{"body":"New patch fixing issues that Arun pointed out.","from":"developer"},{"body":"Hi arun, would you like to code-review this patch one final time? Thanks.","from":"developer"},{"body":"I just committed this. Thanks, Matei! (It's been a long-drawn affair, appreciate your patience!)","from":"developer"},{"body":"Awesome, thanks!","from":"developer"}],"created":"2008-03-29T01:21:46.000+0000","description":"Enabling map output compression and setting the compression type to BLOCK causes the progress counters during the reduce to go crazy and report progress counts over 100%.\n\nThis is problematic for speculative execution because it thinks the tasks are doing fine.","issue_id":"12392618","key":"HADOOP-3131","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-07-31T00:06:07.000+0000","role":"fixed_distractor","summary":"enabling BLOCK compression for map outputs breaks the reduce progress counters"} {"case_id":"12392962","cluster":"DISTRACTOR-HADOOP-3155","comments":[{"body":"In the job, were there any map failures or lost trackers leading to killed maps? If so, could you please check the JobTracker logs and see whether the failed/lost tasks got reexecuted.\nWere many reducers stuck? ","created":"2008-04-02T20:30:07.419+0000"},{"body":"\nI did not notice map failures\nAll maps completed successfully.\nFour reducers got stuck.\nA few reducers got re-executed because of too many failures in fetching map output.\n\n","created":"2008-04-02T23:46:55.947+0000"},{"body":"\n\nI think the problem was due to the slow network link of a machine.\n\nThis issue can be ignored for now.\n","created":"2008-04-03T03:25:47.959+0000"},{"body":"We have been running into the same issue, very unlikely to be a network problem:\n\n#nodes 1350\n#maps: 100,000 (running about 5 minutes each)\n#reduces: 2,500\n\nThe maps were finished after about 4 hours, the data shuffling took about 8 hours, the amount of data very small.\n\nTypical pattern of reduces logs:\n\n2008-05-05 19:32:43,215 INFO org.apache.hadoop.mapred.ReduceTask: task_200805050640_0001_r_001079_0 Need 6 map output(s)\n2008-05-05 19:32:43,216 INFO org.apache.hadoop.mapred.ReduceTask: task_200805050640_0001_r_001079_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-05-05 19:32:43,216 INFO org.apache.hadoop.mapred.ReduceTask: task_200805050640_0001_r_001079_0 Got 5 known map output location(s); scheduling...\n2008-05-05 19:32:43,216 INFO org.apache.hadoop.mapred.ReduceTask: task_200805050640_0001_r_001079_0 Scheduled 0 of 5 known outputs (0 slow hosts and 5 dup hosts)\n\nAlso saw log messages with '(xxx slow hosts and 0 dup hosts)\n","created":"2008-05-05T22:18:00.970+0000"},{"body":"I just encountered this using hadoop-0.17.0 on a 7-node cluster. The performance of the affected reduce tasks has slowed dramatically - \"24.43%\n reduce > copy (195 of 266 at 0.00 MB/s)\". However the progress percentage IS gradually increasing at the rate of approximately 0.10% every 20 minutes.\n\nIs this related to this penalty box error that is showing up in my logs?\n\n2008-06-02 18:37:43,437 INFO org.apache.hadoop.mapred.ReduceTask: Task task_200806021621_0002_r_000013_0: Failed fetch #6 from task_200806021621_0002_m_000244_0\n2008-06-02 18:37:43,437 WARN org.apache.hadoop.mapred.ReduceTask: task_200806021621_0002_r_000013_0 adding host blah.corp.XXX.XXX to penalty box, next contact in 128 seconds\n\nIt seems that hadoop is autodetecting the wrong hostname for some of the cluster nodes, leading to issues when one node tries to contact another to retrieve information. I have no idea if this is related to the issue.\n\nRandom observations:\n\n2 of 28 reduce tasks are running extremely slowly.\nAll slow reduce tasks have precisely the same progress percentage.","created":"2008-06-03T15:34:38.822+0000"},{"body":"This most likely is because of the current backoff & retrial strategies in the shuffle. There are some jiras open to improve them - HADOOP-3478 & HADOOP-3327.","created":"2008-06-08T13:40:19.128+0000"},{"body":"I noticed something similar occuring to the original issue description (\n2008-04-02 17:17:44,640 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts))\n) repeated over and over again\nIt seems that this is triggered if different nodes are configured to use a different dfs name, including the following\n\non the TaskTracker and JobTracker instances\n\nfs.default.name = \"foo.example.com:50001\"\n\non the DataNode instances\n\nfs.default.name = \"hdfs://foo.example.com:50001/\"\n\nThis seems to cause the reducers to be unable to find the mapper output","created":"2008-07-11T18:20:55.413+0000"},{"body":"I'm also getting this behaviour with 0.17.1, on an EC2 hosted cluster, 14 worker (m1.small) and 1 control (NameNode/JobTracker) VMs.\n\nThe job was rather small (40 map tasks, with minimal data to each, they fetch their input internally from S3), 20 reducers. streaming.jar.\n\nAs far as I have been able to notice:\n\n-) no progress on reducers (I had cases where jobs have been hanging for multiples of 12 hours)\n-) the configuration is exactly the same on all nodes, so the DFS name idea from the previous comment does not apply.\n-) killing the reduce tasks restarts them on the same node. There they go to the same state (with the same number of needed outputs) very quickly, the job being not that big. Then they hang at the same place.\n-) the only remedy that I have found if one call this so is restarting the cluster. So this hang is not 100% input dependant.\n\nAndreas","created":"2008-07-24T08:30:57.907+0000"},{"body":"Andreas, the next time you notice this could you please run:\n\n$ bin/hadoop job -events 0 1000000\n\nand provide us the output? Thanks!","created":"2008-07-24T08:49:33.202+0000"},{"body":"We're also running into this on 0.16.2. The job has 50 mappers, 895 reducers, and all mappers finished exactly once (no speculative execution).\n\nWhen the failures occur, they seem to effect all the reduce tasks assigned to a tasktracker; they all get stuck waiting for the same number of map outputs. The number does vary from tasktracker to tasktracker (and many tasktrackers complete the reducers successfully). The task's syslog reports:\n\n{code}\n2008-08-02 06:10:02,686 INFO org.apache.hadoop.mapred.ReduceTask: task_200807311630_0007_r_000006_0 Ignoring obsolete copy result for Map Task: task_200807311630_0007_m_000001_0 from host: HOST-1131\n{code}\n\nThe number of those lines will be equal to the number of map outputs they (later) get stuck waiting for. The result that is claimed to be obsolete has already been copied, though:\n\n{code}\n2008-08-02 06:09:21,461 INFO org.apache.hadoop.mapred.ReduceTask: task_200807311630_0007_r_000006_0 Copying task_200807311630_0007_m_000001_0 output from HOST-1131\n....\n2008-08-02 06:10:02,683 INFO org.apache.hadoop.mapred.ReduceTask: task_200807311630_0007_r_000006_0 done copying task_200807311630_0007_m_000001_0 output from HOST-1131\n{code}\n\nAs a work-around, manually killing the task until it gets assigned to a working tasktracker allows the job to complete.\n\nAttachment {{task_200807311630_0007_r_000006_0.syslog.gz}} is the full syslog from one such failed task, I'll upload the output of {{./bin/apollo-hadoop job -events job_200807311630_0007 0 10000000}} next.\n","created":"2008-08-02T21:21:16.045+0000"},{"body":"Thanks for the logs, Rick.\nAre there lost task trackers during the job?\nCan you attach JobTracker log and TaskTracker log on which task_200807311630_0007_r_000006_0 ran?\nAnd if you upload log for another stuck reducer, that will be very useful.","created":"2008-08-04T10:44:09.001+0000"},{"body":"Here are debug patches logging task completion events at the tasktracker for branches 0.16 and 0.17.\nCould you guys please apply the respective patches and run applications.\n\nAnd if this problem occurs again, please upload the tasktracker logs for a _good_ task tracker as well as the logs for a tasktracker on which reducers get stuck.\n\nAnd also upload the result of _kill -3 _ for the problematic tasktracker.","created":"2008-08-05T12:22:07.655+0000"},{"body":"Took me awhile to assemble the logs, but here they are. Included are:\n\n- all reducer task syslog files from two failed tasktrackers and one good tasktracker\n- the tasktracker logs from the same hosts\n- the jobtracker log entries for this job\n- stackdumps (kill -QUIT output) from the tasktracker and task VM created on a hung tasktracker while the job was running.","created":"2008-08-05T18:51:01.848+0000"},{"body":"Thanks Rick for the effort.\n\nOne more question: What sort of failures did you observe on bad tasktrackers? I dont see any failures in tasktracker logs or the job tracker logs.","created":"2008-08-06T04:53:00.113+0000"},{"body":"Thanks Amareshwari. These tasks don't fail --- they just don't make any progress. In the case of job 200807311630_0007, the stuck task attempts were eventually manually killed until they were assigned to a non-stuck tasktracker.\n\nWe've patched the tasktrackers to log the extra debugging, though since restarting them to deploy the patch, this problem hasn't recurred (we saw failures like this for two jobs in a row before the update, so it seems related to some kind of state accumulated in either the tasktrackers or jobtracker).","created":"2008-08-06T07:13:04.679+0000"},{"body":"{code}\n\"Map-events fetcher for all reduce tasks on tracker_HOST-6026:localhost.localdomain/127.0.0.1:53932\" daemon prio=10 tid=0x7ff5d800 nid=0x510b in Object.wait() [0x81e0a000..0x81e0a1b0]\n java.lang.Thread.State: WAITING (on object monitor)\n\tat java.lang.Object.wait(Native Method)\n\t- waiting on <0x87981ee0> (a java.util.TreeMap)\n\tat java.lang.Object.wait(Object.java:485)\n\tat org.apache.hadoop.mapred.TaskTracker$MapEventsFetcherThread.run(TaskTracker.java:511)\n\t- locked <0x87981ee0> (a java.util.TreeMap)\n{code}\nThe above _kill -QUIT_ output of the tasktracker shows that MapEventsFetcherThread is in WAITING state, which is abnormal with the scenario.\n\nIt says that the Thread is waiting at line number 511:\n{code}\n509 while (((fList = reducesInShuffle()).size()) == 0) {\n510 try {\n511 runningJobs.wait();\n512 } catch (InterruptedException e) {\n513 LOG.info(\"Shutting down: \" + getName());\n514 return;\n515 }\n516 }\n{code}\nMapEventsFetcherThread will wake up from here if there are any reducesInShuffle. \nOne scenario is the thread saw that there are no reducesInShuffle and is waiting, and the notifications are missed. But, \nSince the reducers have fetched some map outputs, that says that the thread woke up and fetched some events. And since the reduces are still in shuffle, the fetcherThread should not go to wait again.\n\nIf I'm not missing something, there could be a bad data structure that is causing the issue.\n","created":"2008-08-06T08:42:50.767+0000"},{"body":"Facebook job hang from 10/08","created":"2008-10-08T16:44:59.567+0000"},{"body":"hey folks - i just uploaded the output from bin/hadoop job -events 0 1000000\n\nwe are beginning to reproduce this problem once in a while and it's hugely disruptive. the cluster is in this state right now - we will try to not restart it if any information can be extracted that can help debugging this .. Please let us know.","created":"2008-10-08T16:46:59.806+0000"},{"body":"btw - one observation - failing the reduce task does not fix the problem. the restarted reducer ends up hanging at precisely the same point (waiting for same N number of map-outputs - although i can't tell if it's the exact same map-outputs - but it's the exact same 'N' all the time). That makes me suspect that the issue is not a race condition (which was one of the possibilities discussed in hadoop-4360).","created":"2008-10-08T16:55:49.360+0000"},{"body":"Some more information from facebook:\n\nCompletion Events from TASKTRACKER: (NOTE: line 1-3 and line 4-6 are duplicates)\nNextId: 57\nID\tEvent ID\tTask ID\tID within Job\tType\tStatus\n0\t0\ttask_200810061724_5371_m_000019_0\t19\tMap\tSUCCEEDED\n1\t0\ttask_200810061724_5371_m_000023_0\t23\tMap\tSUCCEEDED\n2\t0\ttask_200810061724_5371_m_000008_0\t8\tMap\tSUCCEEDED\n3\t0\ttask_200810061724_5371_m_000019_0\t19\tMap\tSUCCEEDED\n4\t0\ttask_200810061724_5371_m_000023_0\t23\tMap\tSUCCEEDED\n5\t0\ttask_200810061724_5371_m_000008_0\t8\tMap\tSUCCEEDED\n6\t0\ttask_200810061724_5371_m_000012_0\t12\tMap\tSUCCEEDED\n7\t0\ttask_200810061724_5371_m_000017_0\t17\tMap\tSUCCEEDED\n8\t0\ttask_200810061724_5371_m_000021_0\t21\tMap\tSUCCEEDED\n9\t0\ttask_200810061724_5371_m_000011_0\t11\tMap\tSUCCEEDED\n10\t0\ttask_200810061724_5371_m_000015_0\t15\tMap\tSUCCEEDED\n11\t0\ttask_200810061724_5371_m_000020_0\t20\tMap\tSUCCEEDED\n12\t0\ttask_200810061724_5371_m_000003_0\t3\tMap\tSUCCEEDED\n13\t0\ttask_200810061724_5371_m_000009_0\t9\tMap\tSUCCEEDED\n14\t0\ttask_200810061724_5371_m_000004_0\t4\tMap\tSUCCEEDED\n15\t0\ttask_200810061724_5371_m_000005_0\t5\tMap\tSUCCEEDED\n16\t0\ttask_200810061724_5371_m_000013_0\t13\tMap\tSUCCEEDED\n17\t0\ttask_200810061724_5371_m_000014_0\t14\tMap\tSUCCEEDED\n18\t0\ttask_200810061724_5371_m_000006_0\t6\tMap\tSUCCEEDED\n19\t0\ttask_200810061724_5371_m_000002_0\t2\tMap\tSUCCEEDED\n20\t0\ttask_200810061724_5371_m_000007_0\t7\tMap\tSUCCEEDED\n21\t0\ttask_200810061724_5371_m_000010_0\t10\tMap\tSUCCEEDED\n22\t0\ttask_200810061724_5371_m_000001_0\t1\tMap\tSUCCEEDED\n23\t0\ttask_200810061724_5371_m_000016_0\t16\tMap\tSUCCEEDED\n24\t0\ttask_200810061724_5371_m_000018_0\t18\tMap\tSUCCEEDED\n25\t0\ttask_200810061724_5371_m_000018_1\t18\tMap\tKILLED\n\n\nCompletion Events from JOBTRACKER:\nID\tEvent ID\tTask ID\tID within Job\tType\tStatus\n0\t0\ttask_200810061724_5371_m_000019_0\t19\tMap\tSUCCEEDED\n1\t1\ttask_200810061724_5371_m_000023_0\t23\tMap\tSUCCEEDED\n2\t2\ttask_200810061724_5371_m_000008_0\t8\tMap\tSUCCEEDED\n3\t3\ttask_200810061724_5371_m_000024_0\t24\tMap\tSUCCEEDED\n4\t4\ttask_200810061724_5371_m_000022_0\t22\tMap\tSUCCEEDED\n5\t5\ttask_200810061724_5371_m_000000_0\t0\tMap\tSUCCEEDED\n6\t6\ttask_200810061724_5371_m_000012_0\t12\tMap\tSUCCEEDED\n7\t7\ttask_200810061724_5371_m_000017_0\t17\tMap\tSUCCEEDED\n8\t8\ttask_200810061724_5371_m_000021_0\t21\tMap\tSUCCEEDED\n9\t9\ttask_200810061724_5371_m_000011_0\t11\tMap\tSUCCEEDED\n10\t10\ttask_200810061724_5371_m_000015_0\t15\tMap\tSUCCEEDED\n11\t11\ttask_200810061724_5371_m_000020_0\t20\tMap\tSUCCEEDED\n12\t12\ttask_200810061724_5371_m_000003_0\t3\tMap\tSUCCEEDED\n13\t13\ttask_200810061724_5371_m_000009_0\t9\tMap\tSUCCEEDED\n14\t14\ttask_200810061724_5371_m_000004_0\t4\tMap\tSUCCEEDED\n15\t15\ttask_200810061724_5371_m_000005_0\t5\tMap\tSUCCEEDED\n16\t16\ttask_200810061724_5371_m_000013_0\t13\tMap\tSUCCEEDED\n17\t17\ttask_200810061724_5371_m_000014_0\t14\tMap\tSUCCEEDED\n18\t18\ttask_200810061724_5371_m_000006_0\t6\tMap\tSUCCEEDED\n19\t19\ttask_200810061724_5371_m_000002_0\t2\tMap\tSUCCEEDED\n20\t20\ttask_200810061724_5371_m_000007_0\t7\tMap\tSUCCEEDED\n21\t21\ttask_200810061724_5371_m_000010_0\t10\tMap\tSUCCEEDED\n22\t22\ttask_200810061724_5371_m_000001_0\t1\tMap\tSUCCEEDED\n23\t23\ttask_200810061724_5371_m_000016_0\t16\tMap\tSUCCEEDED\n24\t24\ttask_200810061724_5371_m_000018_0\t18\tMap\tSUCCEEDED\n25\t25\ttask_200810061724_5371_r_000000_0\t0\tReduce\tSUCCEEDED\n26\t26\ttask_200810061724_5371_r_000005_0\t5\tReduce\tSUCCEEDED\n27\t27\ttask_200810061724_5371_r_000009_0\t9\tReduce\tSUCCEEDED\n28\t28\ttask_200810061724_5371_r_000007_0\t7\tReduce\tSUCCEEDED\n29\t29\ttask_200810061724_5371_r_000004_0\t4\tReduce\tSUCCEEDED\n30\t30\ttask_200810061724_5371_r_000030_0\t30\tReduce\tSUCCEEDED\n31\t31\ttask_200810061724_5371_r_000026_0\t26\tReduce\tSUCCEEDED\n32\t32\ttask_200810061724_5371_r_000024_0\t24\tReduce\tSUCCEEDED\n33\t33\ttask_200810061724_5371_r_000018_0\t18\tReduce\tSUCCEEDED\n34\t34\ttask_200810061724_5371_r_000021_0\t21\tReduce\tSUCCEEDED\n35\t35\ttask_200810061724_5371_r_000015_0\t15\tReduce\tSUCCEEDED\n36\t36\ttask_200810061724_5371_r_000019_0\t19\tReduce\tSUCCEEDED\n37\t37\ttask_200810061724_5371_r_000022_0\t22\tReduce\tSUCCEEDED\n38\t38\ttask_200810061724_5371_r_000008_0\t8\tReduce\tSUCCEEDED\n39\t39\ttask_200810061724_5371_r_000028_0\t28\tReduce\tSUCCEEDED\n40\t40\ttask_200810061724_5371_r_000001_0\t1\tReduce\tSUCCEEDED\n41\t41\ttask_200810061724_5371_r_000014_0\t14\tReduce\tSUCCEEDED\n42\t42\ttask_200810061724_5371_r_000025_0\t25\tReduce\tSUCCEEDED\n43\t43\ttask_200810061724_5371_r_000003_0\t3\tReduce\tSUCCEEDED\n44\t44\ttask_200810061724_5371_r_000010_0\t10\tReduce\tSUCCEEDED\n45\t45\ttask_200810061724_5371_r_000017_0\t17\tReduce\tSUCCEEDED\n46\t46\ttask_200810061724_5371_r_000016_0\t16\tReduce\tSUCCEEDED\n47\t47\ttask_200810061724_5371_r_000002_0\t2\tReduce\tSUCCEEDED\n48\t48\ttask_200810061724_5371_r_000011_0\t11\tReduce\tSUCCEEDED\n49\t49\ttask_200810061724_5371_r_000023_0\t23\tReduce\tSUCCEEDED\n50\t50\ttask_200810061724_5371_r_000029_0\t29\tReduce\tSUCCEEDED\n51\t51\ttask_200810061724_5371_r_000020_0\t20\tReduce\tSUCCEEDED\n52\t52\ttask_200810061724_5371_r_000027_0\t27\tReduce\tSUCCEEDED\n53\t53\ttask_200810061724_5371_m_000018_1\t18\tMap\tKILLED\n54\t54\ttask_200810061724_5371_r_000012_0\t12\tReduce\tSUCCEEDED\n55\t55\ttask_200810061724_5371_r_000006_0\t6\tReduce\tKILLED\n56\t56\ttask_200810061724_5371_r_000006_1\t6\tReduce\tSUCCEEDED\n\n\nBasically TaskTracker get some duplicates for some mysterious reasons. Dhruba also finds that at the time of these duplicate we see RPC time out on the task tracker.\n","created":"2008-10-08T20:39:14.747+0000"},{"body":"The problem is that when the TaskTacker re-initializes its own data structures, it interrupts the fetcher thread but does not wait for the fetcher thread to exit. This means that multiple instances of the fetcher thread can be active at the same time, thus inserting duplicate entries in the fetched list.","created":"2008-10-08T21:57:01.074+0000"},{"body":"Review comments welcome.","created":"2008-10-08T21:57:30.228+0000"},{"body":"Catch Interrupts while waiting in the getTaskCompletionEvents RPC.","created":"2008-10-08T22:46:38.406+0000"},{"body":"+1\n","created":"2008-10-09T00:50:28.344+0000"},{"body":"Good catch Dhruba! \n\nCan you please verify (via extra logging if necessary) that this is indeed the cause? Just to be safe! *smile*\n\nOTOH, shouldn't we also go ahead and purge _all_ state in RunningJobs etc. when in TaskTracker.close? i.e. clear TaskTracker.runningJobs (thus the previous 'FetchStatus' etc.) That probably is a related bug worth fixing too?","created":"2008-10-09T01:06:07.348+0000"},{"body":"Dhruba - i just realized based on the previous comments internally that this patch will not work.\n\nthe issue is that interruptedexceptions are ignored by ipc/Client.java. So they are not a reliable way of shutting down the mapeventsfetcher. once consumed by the wait() code inside Client.java - the interrupt is gone and the thread will never terminate.\n\nSo we need to set some flag as well and check that at the top of the loop. If the code is blocked in RPC layer - it will eventually time out. If it is blocked in the wait() call inside the fetcher thread - then the interrupt will work. In either case - it should check for a shutdown flag at the top of the loop and use that to trigger termination.","created":"2008-10-09T01:50:12.935+0000"},{"body":"it would seem that the problem is not so much a race between the shutdown and the initialize of the mapeventsfetcher - as perhaps that mapeventsfetcher is not shutdown at all (because of the interrupts being swallowed).","created":"2008-10-09T02:10:21.797+0000"},{"body":"@Arun: I deduced this scenario from looking at the tasktracker logs. So, my theory is indeed validated. In out cluster, it typically ocurs during periods of high load. The TaskTracker.close() already clears most data structures. In this case, after the data strctures are cleared, the first instance of the fercher thread re-populated the data structure and then exited. The second instance of the fetcher thread continued using the data structure. Hence the duplicate values. Does this make sense?\n\n@Joydeep: I did see that the code in the ipc layer ignored InterruptedExceptions. The question is why does the ipc code catch and ignore InterruptedExceptions? Shouldn's it just throw InterruptedException?","created":"2008-10-09T06:15:05.799+0000"},{"body":"I am tentatively marking this issue for 0.19 because my belief is that it causes most of the tasktrackers to freeze up during times of heavy load on JT. I am guessing this might have occured with Christian Kunz's cluster earlier.","created":"2008-10-09T06:17:06.698+0000"},{"body":"@Joydeep: do you mean the following piece of code in waitForProxy function?\n{code}\n try {\n Thread.sleep(1000);\n } catch (InterruptedException ie) {\n // IGNORE\n }\n{code}\n\nwaitForProxy is called by the main thread (that calls TaskTracker.initialize()). It's not called by the FetchThread. So FetchThread will never ignore InterruptedException, correct?\n","created":"2008-10-09T06:36:29.390+0000"},{"body":"i don't know why the ipc layer ignores interruptedexceptions. but the practical matter is that it does - and my naiive suspicion is that changing that would be non trivial. also - since the ipc layer always times out anyway - interrupts are not required to break out of RPC calls.\n\nas i mentioned - it's no longer a race condition. the assumption that the thread.interrupt() works to kill the mapeventsfetcher is inherently wrong. my speculation is that both the threads run forever - which is why new jobs/tasks also hit the same problem when they run on this TT (i am assuming this since we had many many jobs that hit the same problem - and they started hours apart).\n\nwe need a shutdown flag in addition to the interrupt and need to check that flag in the eventsfetcher loop. otherwise thread.join will hang. (also - if u look at general documentation on the web about using interrupts in java - everyone advises using a flag in addition as general practice).\n\ni would be curious if there are other parts of the hadoop code base that suffer from the same assumptions (about thread.interrupt being caught and thrown).","created":"2008-10-09T06:54:24.400+0000"},{"body":"@Joydeep: I see what you mean. It's the Client.java. \n\nThere are 3 methods named \"call\" in Client.java. We are calling the one that actually throws InterruptedException.\n\n{code}\n public Writable call(Writable param, InetSocketAddress address)\n throws InterruptedException, IOException {\n\n public Writable call(Writable param, InetSocketAddress addr, \n UserGroupInformation ticket) \n throws InterruptedException, IOException {\n\n public Writable[] call(Writable[] params, InetSocketAddress[] addresses)\n throws IOException {\n{code}\n\nThe swallowing of InterruptedException definitely needs to be fixed, but it seems it does not affect this patch?\n","created":"2008-10-09T07:01:09.763+0000"},{"body":"@Zheng: The ipc code in Client.java swallows InterruptedExceptions. So, if the fetcher thread is blocked inside the RPC code, and then it gets interrupted, it will swallow the InterruptedException. If this occurs, the fetcher thread will not exit.","created":"2008-10-09T07:02:34.850+0000"},{"body":"@Zheng - my assumption (from looking at the code a long time back) is that the Proxy code (invoked by calls like getTaskcompletion) winds it's ways to calls to the ipc/Client.java routines. If u thumb through ipc/Client.java - the code ignores Interrupts. This is safe since it always has timeouts - so it doesn't need interrupts as a way of unblocking. but it also means that one cannot use Interrupted exception as a way of detecting shutdown request.\n\nuse interrupts to unblock. use flags to signal and check for shutdown.","created":"2008-10-09T07:25:54.654+0000"},{"body":"@Joydeep:\n\nI just saw the line \"Connection connection = getConnection(addr, ticket);\" inside this \"call\" method that is called by RPC.Invoker.invoke. \nInside the getConnection method there is indeed a catch block ignoring InterruptedException.\n\nSo it's true that we can be ignoring InterruptedException during the calling of getTaskcompletion -> RPC.Invoker.invoke -> call(Writable param, InetSocketAddress addr, UserGroupInformation ticket) -> getConnection.\n\n{code}\n public Writable call(Writable param, InetSocketAddress addr, \n UserGroupInformation ticket) \n throws InterruptedException, IOException {\n Connection connection = getConnection(addr, ticket);\n Call call = new Call(param);\n synchronized (call) {\n connection.sendParam(call); // send the parameter\n long wait = timeout;\n do {\n call.wait(wait); // wait for the result\n wait = timeout - (System.currentTimeMillis() - call.lastActivity);\n } while (!call.done && wait > 0);\n\n if (call.error != null) {\n throw new RemoteException(call.errorClass, call.error);\n } else if (!call.done) {\n throw new SocketTimeoutException(\"timed out waiting for rpc response\");\n } else {\n return call.value;\n }\n }\n }\n{code}\n","created":"2008-10-09T07:37:50.940+0000"},{"body":"GetConnection does swallow InterruptedException. To verify if this is the culprit of the problem, check the TaskTracker log if you see a message starting with \"Retrying connect to server: \".","created":"2008-10-09T18:57:24.526+0000"},{"body":"Before rethrowing InterruptedException, make sure to remove the connection out of the connections queue.","created":"2008-10-09T19:17:37.247+0000"},{"body":"A modified patch that exits the fetchLoop as soon as the fetcher thread is interrupted.\n\nCan somebody please provide me review comments?","created":"2008-10-10T00:06:23.931+0000"},{"body":"bq. I deduced this scenario from looking at the tasktracker logs. So, my theory is indeed validated. In out cluster, it typically ocurs during periods of high load. The TaskTracker.close() already clears most data structures. In this case, after the data strctures are cleared, the first instance of the fercher thread re-populated the data structure and then exited. The second instance of the fetcher thread continued using the data structure. Hence the duplicate values. Does this make sense?\nDhruba, I don't quite see how this could happen (maybe I am missing something and I hope so *smile*). Even if you have two threads live at some point, both of them would be sharing the same object (FetchStatus) and hence should be seeing the same fromEventId, no?\nSo if the first thread repopulated the datastructure, the second thread should be seeing the same fromEventId object (and hence would be seeing the same value), no? Also, note that two concurrent fetching of map completion events for the same job cannot happen since the fetch is synchronized on the fromEventId object.\n\nI do agree that we should not have two concurrent fetcher threads live simultaneously.","created":"2008-10-10T04:24:59.932+0000"},{"body":"Another thing I forgot to ask - did you \"Lost tracker \" lines in the JobTracker log for those trackers where this problem started showing up. During normal execution, TaskTracker.close will be invoked only when a tracker is lost. ","created":"2008-10-10T04:28:36.853+0000"},{"body":"1. say two threads start with same fromEventId.\n2. Both get same set of events\n3. now each of them increments fromEventId - so we get duplicate events and miss out on getting the real ones.\n\nin 17.1 - there is no synchronized clause on fromEventId. it's possible that this problem does not exist in 19 (or since whenever the synchronized block was added)","created":"2008-10-10T06:09:10.691+0000"},{"body":"bq. in 17.1 - there is no synchronized clause on fromEventId. it's possible that this problem does not exist in 19 (or since whenever the synchronized block was added)\nOh that's true! \n\n+1 for the patch","created":"2008-10-10T06:19:13.470+0000"},{"body":"Maybe we should make this patch work for 0.18.3 as well. Thoughts? \nTrue that this problem might not show up with 0.19 due to the synchronization around fromEventId, but it makes sense to join on the thread before starting a new one.","created":"2008-10-10T09:51:42.078+0000"},{"body":"Have to merge with latest trunk","created":"2008-10-14T17:47:03.391+0000"},{"body":"Merged with latest trunk.\n\n@Devaraj: thanks for reviewing this patch. I am currently targeting this for trunk only (not for 0.18 branch). If it needs to go into 0l18, please let me know.","created":"2008-10-14T17:48:49.537+0000"},{"body":"I think we should get a patch for 17/18 as well.. In 0.19, this problem may not be there at all due to the synchronization introduced around the fromEventId object. It makes sense to fix the bug where it surely shows up...","created":"2008-10-14T19:43:12.048+0000"},{"body":"Fix findbugs error.","created":"2008-10-15T18:18:45.346+0000"},{"body":"Fix findbugs issue.","created":"2008-10-15T18:19:15.255+0000"},{"body":"I am marking this as a blocker for 0.19.","created":"2008-10-15T18:20:46.725+0000"},{"body":"I just committed this. Thanks, Dhruba!","created":"2008-10-16T22:57:10.998+0000"},{"body":"Experienced this issue on 0.18.3 as well. I've attached a patch against this branch.","created":"2009-05-14T20:51:34.555+0000"}],"conversations":[{"body":"This happened with hadoop-0.16.2:\n\nIn relatively small job (a few hundreds of mappers and reducers), reducers were stuck at shuffling.\nI saw the lines like the following repeated hundreds of thousands of times over a few hours:\n\n2008-04-02 17:17:44,640 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:46,643 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:46,643 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:46,643 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:46,643 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:48,645 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:48,645 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:48,645 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:48,645 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:50,647 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:50,647 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:50,647 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:50,647 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:52,649 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:52,650 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:52,650 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:52,650 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:54,651 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:54,652 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:54,652 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:54,652 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:56,654 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:56,654 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:56,654 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:56,654 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:58,656 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:58,656 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:58,656 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:58,656 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:00,658 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:00,658 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:00,658 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:00,658 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:02,660 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:02,661 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:02,661 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:02,661 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:04,662 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:04,663 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:04,663 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:04,663 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:06,664 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:06,665 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:06,665 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:06,665 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:08,667 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:08,667 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:08,667 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:08,667 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:10,669 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:10,669 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:10,669 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:10,669 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:12,671 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:12,671 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:12,671 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:12,671 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:14,673 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:14,674 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:14,674 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:14,674 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:16,675 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:16,676 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:16,676 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:16,676 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:18,678 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:18,678 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:18,678 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:18,678 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:20,680 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:20,680 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:20,680 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:20,680 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:22,682 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:22,682 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:22,682 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n","from":"reporter","subject":"reducers stuck at shuffling "},{"body":"In the job, were there any map failures or lost trackers leading to killed maps? If so, could you please check the JobTracker logs and see whether the failed/lost tasks got reexecuted.\nWere many reducers stuck? ","from":"developer"},{"body":"\nI did not notice map failures\nAll maps completed successfully.\nFour reducers got stuck.\nA few reducers got re-executed because of too many failures in fetching map output.\n\n","from":"developer"},{"body":"\n\nI think the problem was due to the slow network link of a machine.\n\nThis issue can be ignored for now.\n","from":"developer"},{"body":"We have been running into the same issue, very unlikely to be a network problem:\n\n#nodes 1350\n#maps: 100,000 (running about 5 minutes each)\n#reduces: 2,500\n\nThe maps were finished after about 4 hours, the data shuffling took about 8 hours, the amount of data very small.\n\nTypical pattern of reduces logs:\n\n2008-05-05 19:32:43,215 INFO org.apache.hadoop.mapred.ReduceTask: task_200805050640_0001_r_001079_0 Need 6 map output(s)\n2008-05-05 19:32:43,216 INFO org.apache.hadoop.mapred.ReduceTask: task_200805050640_0001_r_001079_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-05-05 19:32:43,216 INFO org.apache.hadoop.mapred.ReduceTask: task_200805050640_0001_r_001079_0 Got 5 known map output location(s); scheduling...\n2008-05-05 19:32:43,216 INFO org.apache.hadoop.mapred.ReduceTask: task_200805050640_0001_r_001079_0 Scheduled 0 of 5 known outputs (0 slow hosts and 5 dup hosts)\n\nAlso saw log messages with '(xxx slow hosts and 0 dup hosts)\n","from":"developer"},{"body":"I just encountered this using hadoop-0.17.0 on a 7-node cluster. The performance of the affected reduce tasks has slowed dramatically - \"24.43%\n reduce > copy (195 of 266 at 0.00 MB/s)\". However the progress percentage IS gradually increasing at the rate of approximately 0.10% every 20 minutes.\n\nIs this related to this penalty box error that is showing up in my logs?\n\n2008-06-02 18:37:43,437 INFO org.apache.hadoop.mapred.ReduceTask: Task task_200806021621_0002_r_000013_0: Failed fetch #6 from task_200806021621_0002_m_000244_0\n2008-06-02 18:37:43,437 WARN org.apache.hadoop.mapred.ReduceTask: task_200806021621_0002_r_000013_0 adding host blah.corp.XXX.XXX to penalty box, next contact in 128 seconds\n\nIt seems that hadoop is autodetecting the wrong hostname for some of the cluster nodes, leading to issues when one node tries to contact another to retrieve information. I have no idea if this is related to the issue.\n\nRandom observations:\n\n2 of 28 reduce tasks are running extremely slowly.\nAll slow reduce tasks have precisely the same progress percentage.","from":"developer"},{"body":"This most likely is because of the current backoff & retrial strategies in the shuffle. There are some jiras open to improve them - HADOOP-3478 & HADOOP-3327.","from":"developer"},{"body":"I noticed something similar occuring to the original issue description (\n2008-04-02 17:17:44,640 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts))\n) repeated over and over again\nIt seems that this is triggered if different nodes are configured to use a different dfs name, including the following\n\non the TaskTracker and JobTracker instances\n\nfs.default.name = \"foo.example.com:50001\"\n\non the DataNode instances\n\nfs.default.name = \"hdfs://foo.example.com:50001/\"\n\nThis seems to cause the reducers to be unable to find the mapper output","from":"developer"},{"body":"I'm also getting this behaviour with 0.17.1, on an EC2 hosted cluster, 14 worker (m1.small) and 1 control (NameNode/JobTracker) VMs.\n\nThe job was rather small (40 map tasks, with minimal data to each, they fetch their input internally from S3), 20 reducers. streaming.jar.\n\nAs far as I have been able to notice:\n\n-) no progress on reducers (I had cases where jobs have been hanging for multiples of 12 hours)\n-) the configuration is exactly the same on all nodes, so the DFS name idea from the previous comment does not apply.\n-) killing the reduce tasks restarts them on the same node. There they go to the same state (with the same number of needed outputs) very quickly, the job being not that big. Then they hang at the same place.\n-) the only remedy that I have found if one call this so is restarting the cluster. So this hang is not 100% input dependant.\n\nAndreas","from":"developer"},{"body":"Andreas, the next time you notice this could you please run:\n\n$ bin/hadoop job -events 0 1000000\n\nand provide us the output? Thanks!","from":"developer"},{"body":"We're also running into this on 0.16.2. The job has 50 mappers, 895 reducers, and all mappers finished exactly once (no speculative execution).\n\nWhen the failures occur, they seem to effect all the reduce tasks assigned to a tasktracker; they all get stuck waiting for the same number of map outputs. The number does vary from tasktracker to tasktracker (and many tasktrackers complete the reducers successfully). The task's syslog reports:\n\n{code}\n2008-08-02 06:10:02,686 INFO org.apache.hadoop.mapred.ReduceTask: task_200807311630_0007_r_000006_0 Ignoring obsolete copy result for Map Task: task_200807311630_0007_m_000001_0 from host: HOST-1131\n{code}\n\nThe number of those lines will be equal to the number of map outputs they (later) get stuck waiting for. The result that is claimed to be obsolete has already been copied, though:\n\n{code}\n2008-08-02 06:09:21,461 INFO org.apache.hadoop.mapred.ReduceTask: task_200807311630_0007_r_000006_0 Copying task_200807311630_0007_m_000001_0 output from HOST-1131\n....\n2008-08-02 06:10:02,683 INFO org.apache.hadoop.mapred.ReduceTask: task_200807311630_0007_r_000006_0 done copying task_200807311630_0007_m_000001_0 output from HOST-1131\n{code}\n\nAs a work-around, manually killing the task until it gets assigned to a working tasktracker allows the job to complete.\n\nAttachment {{task_200807311630_0007_r_000006_0.syslog.gz}} is the full syslog from one such failed task, I'll upload the output of {{./bin/apollo-hadoop job -events job_200807311630_0007 0 10000000}} next.\n","from":"developer"},{"body":"Thanks for the logs, Rick.\nAre there lost task trackers during the job?\nCan you attach JobTracker log and TaskTracker log on which task_200807311630_0007_r_000006_0 ran?\nAnd if you upload log for another stuck reducer, that will be very useful.","from":"developer"},{"body":"Here are debug patches logging task completion events at the tasktracker for branches 0.16 and 0.17.\nCould you guys please apply the respective patches and run applications.\n\nAnd if this problem occurs again, please upload the tasktracker logs for a _good_ task tracker as well as the logs for a tasktracker on which reducers get stuck.\n\nAnd also upload the result of _kill -3 _ for the problematic tasktracker.","from":"developer"},{"body":"Took me awhile to assemble the logs, but here they are. Included are:\n\n- all reducer task syslog files from two failed tasktrackers and one good tasktracker\n- the tasktracker logs from the same hosts\n- the jobtracker log entries for this job\n- stackdumps (kill -QUIT output) from the tasktracker and task VM created on a hung tasktracker while the job was running.","from":"developer"},{"body":"Thanks Rick for the effort.\n\nOne more question: What sort of failures did you observe on bad tasktrackers? I dont see any failures in tasktracker logs or the job tracker logs.","from":"developer"},{"body":"Thanks Amareshwari. These tasks don't fail --- they just don't make any progress. In the case of job 200807311630_0007, the stuck task attempts were eventually manually killed until they were assigned to a non-stuck tasktracker.\n\nWe've patched the tasktrackers to log the extra debugging, though since restarting them to deploy the patch, this problem hasn't recurred (we saw failures like this for two jobs in a row before the update, so it seems related to some kind of state accumulated in either the tasktrackers or jobtracker).","from":"developer"},{"body":"{code}\n\"Map-events fetcher for all reduce tasks on tracker_HOST-6026:localhost.localdomain/127.0.0.1:53932\" daemon prio=10 tid=0x7ff5d800 nid=0x510b in Object.wait() [0x81e0a000..0x81e0a1b0]\n java.lang.Thread.State: WAITING (on object monitor)\n\tat java.lang.Object.wait(Native Method)\n\t- waiting on <0x87981ee0> (a java.util.TreeMap)\n\tat java.lang.Object.wait(Object.java:485)\n\tat org.apache.hadoop.mapred.TaskTracker$MapEventsFetcherThread.run(TaskTracker.java:511)\n\t- locked <0x87981ee0> (a java.util.TreeMap)\n{code}\nThe above _kill -QUIT_ output of the tasktracker shows that MapEventsFetcherThread is in WAITING state, which is abnormal with the scenario.\n\nIt says that the Thread is waiting at line number 511:\n{code}\n509 while (((fList = reducesInShuffle()).size()) == 0) {\n510 try {\n511 runningJobs.wait();\n512 } catch (InterruptedException e) {\n513 LOG.info(\"Shutting down: \" + getName());\n514 return;\n515 }\n516 }\n{code}\nMapEventsFetcherThread will wake up from here if there are any reducesInShuffle. \nOne scenario is the thread saw that there are no reducesInShuffle and is waiting, and the notifications are missed. But, \nSince the reducers have fetched some map outputs, that says that the thread woke up and fetched some events. And since the reduces are still in shuffle, the fetcherThread should not go to wait again.\n\nIf I'm not missing something, there could be a bad data structure that is causing the issue.\n","from":"developer"},{"body":"Facebook job hang from 10/08","from":"developer"},{"body":"hey folks - i just uploaded the output from bin/hadoop job -events 0 1000000\n\nwe are beginning to reproduce this problem once in a while and it's hugely disruptive. the cluster is in this state right now - we will try to not restart it if any information can be extracted that can help debugging this .. Please let us know.","from":"developer"},{"body":"btw - one observation - failing the reduce task does not fix the problem. the restarted reducer ends up hanging at precisely the same point (waiting for same N number of map-outputs - although i can't tell if it's the exact same map-outputs - but it's the exact same 'N' all the time). That makes me suspect that the issue is not a race condition (which was one of the possibilities discussed in hadoop-4360).","from":"developer"},{"body":"Some more information from facebook:\n\nCompletion Events from TASKTRACKER: (NOTE: line 1-3 and line 4-6 are duplicates)\nNextId: 57\nID\tEvent ID\tTask ID\tID within Job\tType\tStatus\n0\t0\ttask_200810061724_5371_m_000019_0\t19\tMap\tSUCCEEDED\n1\t0\ttask_200810061724_5371_m_000023_0\t23\tMap\tSUCCEEDED\n2\t0\ttask_200810061724_5371_m_000008_0\t8\tMap\tSUCCEEDED\n3\t0\ttask_200810061724_5371_m_000019_0\t19\tMap\tSUCCEEDED\n4\t0\ttask_200810061724_5371_m_000023_0\t23\tMap\tSUCCEEDED\n5\t0\ttask_200810061724_5371_m_000008_0\t8\tMap\tSUCCEEDED\n6\t0\ttask_200810061724_5371_m_000012_0\t12\tMap\tSUCCEEDED\n7\t0\ttask_200810061724_5371_m_000017_0\t17\tMap\tSUCCEEDED\n8\t0\ttask_200810061724_5371_m_000021_0\t21\tMap\tSUCCEEDED\n9\t0\ttask_200810061724_5371_m_000011_0\t11\tMap\tSUCCEEDED\n10\t0\ttask_200810061724_5371_m_000015_0\t15\tMap\tSUCCEEDED\n11\t0\ttask_200810061724_5371_m_000020_0\t20\tMap\tSUCCEEDED\n12\t0\ttask_200810061724_5371_m_000003_0\t3\tMap\tSUCCEEDED\n13\t0\ttask_200810061724_5371_m_000009_0\t9\tMap\tSUCCEEDED\n14\t0\ttask_200810061724_5371_m_000004_0\t4\tMap\tSUCCEEDED\n15\t0\ttask_200810061724_5371_m_000005_0\t5\tMap\tSUCCEEDED\n16\t0\ttask_200810061724_5371_m_000013_0\t13\tMap\tSUCCEEDED\n17\t0\ttask_200810061724_5371_m_000014_0\t14\tMap\tSUCCEEDED\n18\t0\ttask_200810061724_5371_m_000006_0\t6\tMap\tSUCCEEDED\n19\t0\ttask_200810061724_5371_m_000002_0\t2\tMap\tSUCCEEDED\n20\t0\ttask_200810061724_5371_m_000007_0\t7\tMap\tSUCCEEDED\n21\t0\ttask_200810061724_5371_m_000010_0\t10\tMap\tSUCCEEDED\n22\t0\ttask_200810061724_5371_m_000001_0\t1\tMap\tSUCCEEDED\n23\t0\ttask_200810061724_5371_m_000016_0\t16\tMap\tSUCCEEDED\n24\t0\ttask_200810061724_5371_m_000018_0\t18\tMap\tSUCCEEDED\n25\t0\ttask_200810061724_5371_m_000018_1\t18\tMap\tKILLED\n\n\nCompletion Events from JOBTRACKER:\nID\tEvent ID\tTask ID\tID within Job\tType\tStatus\n0\t0\ttask_200810061724_5371_m_000019_0\t19\tMap\tSUCCEEDED\n1\t1\ttask_200810061724_5371_m_000023_0\t23\tMap\tSUCCEEDED\n2\t2\ttask_200810061724_5371_m_000008_0\t8\tMap\tSUCCEEDED\n3\t3\ttask_200810061724_5371_m_000024_0\t24\tMap\tSUCCEEDED\n4\t4\ttask_200810061724_5371_m_000022_0\t22\tMap\tSUCCEEDED\n5\t5\ttask_200810061724_5371_m_000000_0\t0\tMap\tSUCCEEDED\n6\t6\ttask_200810061724_5371_m_000012_0\t12\tMap\tSUCCEEDED\n7\t7\ttask_200810061724_5371_m_000017_0\t17\tMap\tSUCCEEDED\n8\t8\ttask_200810061724_5371_m_000021_0\t21\tMap\tSUCCEEDED\n9\t9\ttask_200810061724_5371_m_000011_0\t11\tMap\tSUCCEEDED\n10\t10\ttask_200810061724_5371_m_000015_0\t15\tMap\tSUCCEEDED\n11\t11\ttask_200810061724_5371_m_000020_0\t20\tMap\tSUCCEEDED\n12\t12\ttask_200810061724_5371_m_000003_0\t3\tMap\tSUCCEEDED\n13\t13\ttask_200810061724_5371_m_000009_0\t9\tMap\tSUCCEEDED\n14\t14\ttask_200810061724_5371_m_000004_0\t4\tMap\tSUCCEEDED\n15\t15\ttask_200810061724_5371_m_000005_0\t5\tMap\tSUCCEEDED\n16\t16\ttask_200810061724_5371_m_000013_0\t13\tMap\tSUCCEEDED\n17\t17\ttask_200810061724_5371_m_000014_0\t14\tMap\tSUCCEEDED\n18\t18\ttask_200810061724_5371_m_000006_0\t6\tMap\tSUCCEEDED\n19\t19\ttask_200810061724_5371_m_000002_0\t2\tMap\tSUCCEEDED\n20\t20\ttask_200810061724_5371_m_000007_0\t7\tMap\tSUCCEEDED\n21\t21\ttask_200810061724_5371_m_000010_0\t10\tMap\tSUCCEEDED\n22\t22\ttask_200810061724_5371_m_000001_0\t1\tMap\tSUCCEEDED\n23\t23\ttask_200810061724_5371_m_000016_0\t16\tMap\tSUCCEEDED\n24\t24\ttask_200810061724_5371_m_000018_0\t18\tMap\tSUCCEEDED\n25\t25\ttask_200810061724_5371_r_000000_0\t0\tReduce\tSUCCEEDED\n26\t26\ttask_200810061724_5371_r_000005_0\t5\tReduce\tSUCCEEDED\n27\t27\ttask_200810061724_5371_r_000009_0\t9\tReduce\tSUCCEEDED\n28\t28\ttask_200810061724_5371_r_000007_0\t7\tReduce\tSUCCEEDED\n29\t29\ttask_200810061724_5371_r_000004_0\t4\tReduce\tSUCCEEDED\n30\t30\ttask_200810061724_5371_r_000030_0\t30\tReduce\tSUCCEEDED\n31\t31\ttask_200810061724_5371_r_000026_0\t26\tReduce\tSUCCEEDED\n32\t32\ttask_200810061724_5371_r_000024_0\t24\tReduce\tSUCCEEDED\n33\t33\ttask_200810061724_5371_r_000018_0\t18\tReduce\tSUCCEEDED\n34\t34\ttask_200810061724_5371_r_000021_0\t21\tReduce\tSUCCEEDED\n35\t35\ttask_200810061724_5371_r_000015_0\t15\tReduce\tSUCCEEDED\n36\t36\ttask_200810061724_5371_r_000019_0\t19\tReduce\tSUCCEEDED\n37\t37\ttask_200810061724_5371_r_000022_0\t22\tReduce\tSUCCEEDED\n38\t38\ttask_200810061724_5371_r_000008_0\t8\tReduce\tSUCCEEDED\n39\t39\ttask_200810061724_5371_r_000028_0\t28\tReduce\tSUCCEEDED\n40\t40\ttask_200810061724_5371_r_000001_0\t1\tReduce\tSUCCEEDED\n41\t41\ttask_200810061724_5371_r_000014_0\t14\tReduce\tSUCCEEDED\n42\t42\ttask_200810061724_5371_r_000025_0\t25\tReduce\tSUCCEEDED\n43\t43\ttask_200810061724_5371_r_000003_0\t3\tReduce\tSUCCEEDED\n44\t44\ttask_200810061724_5371_r_000010_0\t10\tReduce\tSUCCEEDED\n45\t45\ttask_200810061724_5371_r_000017_0\t17\tReduce\tSUCCEEDED\n46\t46\ttask_200810061724_5371_r_000016_0\t16\tReduce\tSUCCEEDED\n47\t47\ttask_200810061724_5371_r_000002_0\t2\tReduce\tSUCCEEDED\n48\t48\ttask_200810061724_5371_r_000011_0\t11\tReduce\tSUCCEEDED\n49\t49\ttask_200810061724_5371_r_000023_0\t23\tReduce\tSUCCEEDED\n50\t50\ttask_200810061724_5371_r_000029_0\t29\tReduce\tSUCCEEDED\n51\t51\ttask_200810061724_5371_r_000020_0\t20\tReduce\tSUCCEEDED\n52\t52\ttask_200810061724_5371_r_000027_0\t27\tReduce\tSUCCEEDED\n53\t53\ttask_200810061724_5371_m_000018_1\t18\tMap\tKILLED\n54\t54\ttask_200810061724_5371_r_000012_0\t12\tReduce\tSUCCEEDED\n55\t55\ttask_200810061724_5371_r_000006_0\t6\tReduce\tKILLED\n56\t56\ttask_200810061724_5371_r_000006_1\t6\tReduce\tSUCCEEDED\n\n\nBasically TaskTracker get some duplicates for some mysterious reasons. Dhruba also finds that at the time of these duplicate we see RPC time out on the task tracker.\n","from":"developer"},{"body":"The problem is that when the TaskTacker re-initializes its own data structures, it interrupts the fetcher thread but does not wait for the fetcher thread to exit. This means that multiple instances of the fetcher thread can be active at the same time, thus inserting duplicate entries in the fetched list.","from":"developer"},{"body":"Review comments welcome.","from":"developer"},{"body":"Catch Interrupts while waiting in the getTaskCompletionEvents RPC.","from":"developer"},{"body":"+1\n","from":"developer"},{"body":"Good catch Dhruba! \n\nCan you please verify (via extra logging if necessary) that this is indeed the cause? Just to be safe! *smile*\n\nOTOH, shouldn't we also go ahead and purge _all_ state in RunningJobs etc. when in TaskTracker.close? i.e. clear TaskTracker.runningJobs (thus the previous 'FetchStatus' etc.) That probably is a related bug worth fixing too?","from":"developer"},{"body":"Dhruba - i just realized based on the previous comments internally that this patch will not work.\n\nthe issue is that interruptedexceptions are ignored by ipc/Client.java. So they are not a reliable way of shutting down the mapeventsfetcher. once consumed by the wait() code inside Client.java - the interrupt is gone and the thread will never terminate.\n\nSo we need to set some flag as well and check that at the top of the loop. If the code is blocked in RPC layer - it will eventually time out. If it is blocked in the wait() call inside the fetcher thread - then the interrupt will work. In either case - it should check for a shutdown flag at the top of the loop and use that to trigger termination.","from":"developer"},{"body":"it would seem that the problem is not so much a race between the shutdown and the initialize of the mapeventsfetcher - as perhaps that mapeventsfetcher is not shutdown at all (because of the interrupts being swallowed).","from":"developer"},{"body":"@Arun: I deduced this scenario from looking at the tasktracker logs. So, my theory is indeed validated. In out cluster, it typically ocurs during periods of high load. The TaskTracker.close() already clears most data structures. In this case, after the data strctures are cleared, the first instance of the fercher thread re-populated the data structure and then exited. The second instance of the fetcher thread continued using the data structure. Hence the duplicate values. Does this make sense?\n\n@Joydeep: I did see that the code in the ipc layer ignored InterruptedExceptions. The question is why does the ipc code catch and ignore InterruptedExceptions? Shouldn's it just throw InterruptedException?","from":"developer"},{"body":"I am tentatively marking this issue for 0.19 because my belief is that it causes most of the tasktrackers to freeze up during times of heavy load on JT. I am guessing this might have occured with Christian Kunz's cluster earlier.","from":"developer"},{"body":"@Joydeep: do you mean the following piece of code in waitForProxy function?\n{code}\n try {\n Thread.sleep(1000);\n } catch (InterruptedException ie) {\n // IGNORE\n }\n{code}\n\nwaitForProxy is called by the main thread (that calls TaskTracker.initialize()). It's not called by the FetchThread. So FetchThread will never ignore InterruptedException, correct?\n","from":"developer"},{"body":"i don't know why the ipc layer ignores interruptedexceptions. but the practical matter is that it does - and my naiive suspicion is that changing that would be non trivial. also - since the ipc layer always times out anyway - interrupts are not required to break out of RPC calls.\n\nas i mentioned - it's no longer a race condition. the assumption that the thread.interrupt() works to kill the mapeventsfetcher is inherently wrong. my speculation is that both the threads run forever - which is why new jobs/tasks also hit the same problem when they run on this TT (i am assuming this since we had many many jobs that hit the same problem - and they started hours apart).\n\nwe need a shutdown flag in addition to the interrupt and need to check that flag in the eventsfetcher loop. otherwise thread.join will hang. (also - if u look at general documentation on the web about using interrupts in java - everyone advises using a flag in addition as general practice).\n\ni would be curious if there are other parts of the hadoop code base that suffer from the same assumptions (about thread.interrupt being caught and thrown).","from":"developer"},{"body":"@Joydeep: I see what you mean. It's the Client.java. \n\nThere are 3 methods named \"call\" in Client.java. We are calling the one that actually throws InterruptedException.\n\n{code}\n public Writable call(Writable param, InetSocketAddress address)\n throws InterruptedException, IOException {\n\n public Writable call(Writable param, InetSocketAddress addr, \n UserGroupInformation ticket) \n throws InterruptedException, IOException {\n\n public Writable[] call(Writable[] params, InetSocketAddress[] addresses)\n throws IOException {\n{code}\n\nThe swallowing of InterruptedException definitely needs to be fixed, but it seems it does not affect this patch?\n","from":"developer"},{"body":"@Zheng: The ipc code in Client.java swallows InterruptedExceptions. So, if the fetcher thread is blocked inside the RPC code, and then it gets interrupted, it will swallow the InterruptedException. If this occurs, the fetcher thread will not exit.","from":"developer"},{"body":"@Zheng - my assumption (from looking at the code a long time back) is that the Proxy code (invoked by calls like getTaskcompletion) winds it's ways to calls to the ipc/Client.java routines. If u thumb through ipc/Client.java - the code ignores Interrupts. This is safe since it always has timeouts - so it doesn't need interrupts as a way of unblocking. but it also means that one cannot use Interrupted exception as a way of detecting shutdown request.\n\nuse interrupts to unblock. use flags to signal and check for shutdown.","from":"developer"},{"body":"@Joydeep:\n\nI just saw the line \"Connection connection = getConnection(addr, ticket);\" inside this \"call\" method that is called by RPC.Invoker.invoke. \nInside the getConnection method there is indeed a catch block ignoring InterruptedException.\n\nSo it's true that we can be ignoring InterruptedException during the calling of getTaskcompletion -> RPC.Invoker.invoke -> call(Writable param, InetSocketAddress addr, UserGroupInformation ticket) -> getConnection.\n\n{code}\n public Writable call(Writable param, InetSocketAddress addr, \n UserGroupInformation ticket) \n throws InterruptedException, IOException {\n Connection connection = getConnection(addr, ticket);\n Call call = new Call(param);\n synchronized (call) {\n connection.sendParam(call); // send the parameter\n long wait = timeout;\n do {\n call.wait(wait); // wait for the result\n wait = timeout - (System.currentTimeMillis() - call.lastActivity);\n } while (!call.done && wait > 0);\n\n if (call.error != null) {\n throw new RemoteException(call.errorClass, call.error);\n } else if (!call.done) {\n throw new SocketTimeoutException(\"timed out waiting for rpc response\");\n } else {\n return call.value;\n }\n }\n }\n{code}\n","from":"developer"},{"body":"GetConnection does swallow InterruptedException. To verify if this is the culprit of the problem, check the TaskTracker log if you see a message starting with \"Retrying connect to server: \".","from":"developer"},{"body":"Before rethrowing InterruptedException, make sure to remove the connection out of the connections queue.","from":"developer"},{"body":"A modified patch that exits the fetchLoop as soon as the fetcher thread is interrupted.\n\nCan somebody please provide me review comments?","from":"developer"},{"body":"bq. I deduced this scenario from looking at the tasktracker logs. So, my theory is indeed validated. In out cluster, it typically ocurs during periods of high load. The TaskTracker.close() already clears most data structures. In this case, after the data strctures are cleared, the first instance of the fercher thread re-populated the data structure and then exited. The second instance of the fetcher thread continued using the data structure. Hence the duplicate values. Does this make sense?\nDhruba, I don't quite see how this could happen (maybe I am missing something and I hope so *smile*). Even if you have two threads live at some point, both of them would be sharing the same object (FetchStatus) and hence should be seeing the same fromEventId, no?\nSo if the first thread repopulated the datastructure, the second thread should be seeing the same fromEventId object (and hence would be seeing the same value), no? Also, note that two concurrent fetching of map completion events for the same job cannot happen since the fetch is synchronized on the fromEventId object.\n\nI do agree that we should not have two concurrent fetcher threads live simultaneously.","from":"developer"},{"body":"Another thing I forgot to ask - did you \"Lost tracker \" lines in the JobTracker log for those trackers where this problem started showing up. During normal execution, TaskTracker.close will be invoked only when a tracker is lost. ","from":"developer"},{"body":"1. say two threads start with same fromEventId.\n2. Both get same set of events\n3. now each of them increments fromEventId - so we get duplicate events and miss out on getting the real ones.\n\nin 17.1 - there is no synchronized clause on fromEventId. it's possible that this problem does not exist in 19 (or since whenever the synchronized block was added)","from":"developer"},{"body":"bq. in 17.1 - there is no synchronized clause on fromEventId. it's possible that this problem does not exist in 19 (or since whenever the synchronized block was added)\nOh that's true! \n\n+1 for the patch","from":"developer"},{"body":"Maybe we should make this patch work for 0.18.3 as well. Thoughts? \nTrue that this problem might not show up with 0.19 due to the synchronization around fromEventId, but it makes sense to join on the thread before starting a new one.","from":"developer"},{"body":"Have to merge with latest trunk","from":"developer"},{"body":"Merged with latest trunk.\n\n@Devaraj: thanks for reviewing this patch. I am currently targeting this for trunk only (not for 0.18 branch). If it needs to go into 0l18, please let me know.","from":"developer"},{"body":"I think we should get a patch for 17/18 as well.. In 0.19, this problem may not be there at all due to the synchronization introduced around the fromEventId object. It makes sense to fix the bug where it surely shows up...","from":"developer"},{"body":"Fix findbugs error.","from":"developer"},{"body":"Fix findbugs issue.","from":"developer"},{"body":"I am marking this as a blocker for 0.19.","from":"developer"},{"body":"I just committed this. Thanks, Dhruba!","from":"developer"},{"body":"Experienced this issue on 0.18.3 as well. I've attached a patch against this branch.","from":"developer"}],"created":"2008-04-02T18:27:51.000+0000","description":"This happened with hadoop-0.16.2:\n\nIn relatively small job (a few hundreds of mappers and reducers), reducers were stuck at shuffling.\nI saw the lines like the following repeated hundreds of thousands of times over a few hours:\n\n2008-04-02 17:17:44,640 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:44,641 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:46,643 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:46,643 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:46,643 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:46,643 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:48,645 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:48,645 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:48,645 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:48,645 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:50,647 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:50,647 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:50,647 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:50,647 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:52,649 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:52,650 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:52,650 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:52,650 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:54,651 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:54,652 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:54,652 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:54,652 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:56,654 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:56,654 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:56,654 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:56,654 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:17:58,656 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:17:58,656 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:17:58,656 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:17:58,656 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:00,658 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:00,658 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:00,658 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:00,658 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:02,660 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:02,661 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:02,661 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:02,661 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:04,662 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:04,663 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:04,663 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:04,663 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:06,664 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:06,665 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:06,665 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:06,665 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:08,667 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:08,667 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:08,667 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:08,667 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:10,669 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:10,669 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:10,669 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:10,669 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:12,671 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:12,671 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:12,671 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:12,671 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:14,673 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:14,674 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:14,674 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:14,674 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:16,675 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:16,676 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:16,676 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:16,676 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:18,678 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:18,678 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:18,678 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:18,678 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:20,680 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:20,680 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:20,680 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n2008-04-02 17:18:20,680 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Scheduled 0 of 0 known outputs (0 slow hosts and 0 dup hosts)\n2008-04-02 17:18:22,682 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Need 2 map output(s)\n2008-04-02 17:18:22,682 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0: Got 0 new map-outputs & 0 obsolete map-outputs from tasktracker and 0 map-outputs from previous failures\n2008-04-02 17:18:22,682 INFO org.apache.hadoop.mapred.ReduceTask: task_200804021200_0337_r_000008_0 Got 0 known map output location(s); scheduling...\n","issue_id":"12392962","key":"HADOOP-3155","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-10-16T22:57:11.000+0000","role":"fixed_distractor","summary":"reducers stuck at shuffling "} {"case_id":"12393665","cluster":"DISTRACTOR-HADOOP-3232","comments":[{"body":"Example client exception:\n\n08/04/10 10:09:27 INFO dfs.DFSClient: Exception in createBlockOutputStream java.io.IOException: Bad connect ack with firstBadLink 10.0.5.76:50010\n08/04/10 10:09:27 INFO dfs.DFSClient: Abandoning block blk_-5192866954303337577\n08/04/10 10:09:27 INFO dfs.DFSClient: Waiting to find target node: 10.0.5.70:50010\n\nDatanode log output from around the same time:\n\n2008-04-10 10:09:15,758 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 got response for connect ack from downstream datanode with firstbadlink as 10.0.5.76:50010\n2008-04-10 10:09:15,758 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 forwarding connect ack to upstream firstbadlink is 10.0.5.76:50010\n2008-04-10 10:09:15,758 INFO org.apache.hadoop.dfs.DataNode: PacketResponder blk_-9149681676832791404 2 Exception java.io.EOFException\n at java.io.DataInputStream.readFully(DataInputStream.java:180)\n at java.io.DataInputStream.readLong(DataInputStream.java:399)\n at org.apache.hadoop.dfs.DataNode$PacketResponder.run(DataNode.java:1822)\n at java.lang.Thread.run(Thread.java:619)\n\n2008-04-10 10:09:15,758 INFO org.apache.hadoop.dfs.DataNode: PacketResponder 2 for block blk_-9149681676832791404 terminating\n2008-04-10 10:09:16,793 INFO org.apache.hadoop.dfs.DataNode: writeBlock blk_1511572447827516117 received exception java.net.SocketTimeoutException: Read timed out\n2008-04-10 10:09:16,793 ERROR org.apache.hadoop.dfs.DataNode: 10.0.5.70:50010:DataXceiver: java.net.SocketTimeoutException: Read timed out\n at java.net.SocketInputStream.socketRead0(Native Method)\n at java.net.SocketInputStream.read(SocketInputStream.java:129)\n at java.net.SocketInputStream.read(SocketInputStream.java:182)\n at java.io.DataInputStream.readByte(DataInputStream.java:248)\n at org.apache.hadoop.io.WritableUtils.readVLong(WritableUtils.java:324)\n at org.apache.hadoop.io.WritableUtils.readVInt(WritableUtils.java:346)\n at org.apache.hadoop.io.Text.readString(Text.java:413)\n at org.apache.hadoop.dfs.DataNode$DataXceiver.writeBlock(DataNode.java:1117)\n at org.apache.hadoop.dfs.DataNode$DataXceiver.run(DataNode.java:938)\n at java.lang.Thread.run(Thread.java:619)\n\n2008-04-10 10:09:23,895 INFO org.apache.hadoop.dfs.DataNode: Receiving block blk_-7442302015902809712 src: /10.0.5.74:55546 dest: /10.0.5.74:50010\n2008-04-10 10:09:23,896 INFO org.apache.hadoop.dfs.DataNode: Datanode 0 forwarding connect ack to upstream firstbadlink is\n2008-04-10 10:09:23,927 INFO org.apache.hadoop.dfs.DataNode: Received block blk_-7442302015902809712 of size 902 from /10.0.5.74\n2008-04-10 10:09:23,928 INFO org.apache.hadoop.dfs.DataNode: PacketResponder 0 for block blk_-7442302015902809712 terminating\n2008-04-10 10:09:23,937 INFO org.apache.hadoop.dfs.DataNode: Receiving block blk_4161972554165500020 src: /10.0.11.7:41256 dest: /10.0.11.7:50010\n2008-04-10 10:09:23,968 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 got response for connect ack from downstream datanode with firstbadlink as\n2008-04-10 10:09:23,969 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 forwarding connect ack to upstream firstbadlink is\n2008-04-10 10:09:24,034 INFO org.apache.hadoop.dfs.DataNode: Received block blk_4161972554165500020 of size 95 from /10.0.11.7\n2008-04-10 10:09:24,034 INFO org.apache.hadoop.dfs.DataNode: PacketResponder 2 for block blk_4161972554165500020 terminating\n2008-04-10 10:09:27,664 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 got response for connect ack from downstream datanode with firstbadlink as 10.0.5.76:50010\n2008-04-10 10:09:27,664 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 forwarding connect ack to upstream firstbadlink is 10.0.5.76:50010\n2008-04-10 10:09:27,665 INFO org.apache.hadoop.dfs.DataNode: PacketResponder blk_-5192866954303337577 2 Exception java.io.EOFException\n at java.io.DataInputStream.readFully(DataInputStream.java:180)\n at java.io.DataInputStream.readLong(DataInputStream.java:399)\n at org.apache.hadoop.dfs.DataNode$PacketResponder.run(DataNode.java:1822)\n at java.lang.Thread.run(Thread.java:619)\n\n","created":"2008-04-10T15:02:33.296+0000"},{"body":"Stacktrace of a datanode that lost contact with the namenode. Not sure it helps but doesn't hurt.","created":"2008-04-10T15:06:46.968+0000"},{"body":"Namenode stacktrace","created":"2008-04-10T15:10:24.466+0000"},{"body":"Found 68 of these when doing a stackdump in a datanode that had lost contact with the namenode.\n\n\"org.apache.hadoop.dfs.DataNode$DataXceiver@5b02a6\" daemon prio=10 tid=0x7192bc00 nid=0x5bc6 waiting for monitor entry [0x70cee000..0x70cef0c0]\n java.lang.Thread.State: BLOCKED (on object monitor)\n at org.apache.hadoop.dfs.FSDataset.getFile(FSDataset.java:867)\n - waiting to lock <0x77ce9360> (a org.apache.hadoop.dfs.FSDataset)\n at org.apache.hadoop.dfs.FSDataset.isValidBlock(FSDataset.java:795)\n at org.apache.hadoop.dfs.FSDataset.writeToBlock(FSDataset.java:614)\n at org.apache.hadoop.dfs.DataNode$BlockReceiver.(DataNode.java:1995)\n at org.apache.hadoop.dfs.DataNode$DataXceiver.writeBlock(DataNode.java:1074)\n at org.apache.hadoop.dfs.DataNode$DataXceiver.run(DataNode.java:938)\n at java.lang.Thread.run(Thread.java:619)","created":"2008-04-10T17:04:26.119+0000"},{"body":"> Found 68 of these when doing a stackdump in a datanode that had lost contact with the namenode.\n\nCould you attach the full stack trace of this datanode?","created":"2008-04-10T17:33:14.977+0000"},{"body":"As requested","created":"2008-04-10T17:50:07.487+0000"},{"body":"\nWhile you are at it, could you attach log (.log file) for this datanode as well? The log file show what activity is going on now. Also list approximate time the stack trace was taken.\n\nMy observation is that you are writing a lot of blocks.. and the datanode that looks blocked is blocked while listing all the blocks on the native filesystem. It does this every hour when it sends block reports. Till now nothing suspicious other than heavy write traffic and slow disks. Check iostat on the machine. What is the hardware like?\n\nThe two main threads from the DataNode are :\n\n# locks 0x782f1348 : {noformat}\n\"DataNode: [/var/storage/1/dfs/data,/var/storage/2/dfs/data,/var/storage/3/dfs/data,/var/storage/4/dfs/data]\" daemon prio=10 tid=0x72409000\n nid=0x44f6 runnable [0x71d8a000..0x71d8aec0]\n java.lang.Thread.State: RUNNABLE\n at java.io.UnixFileSystem.list(Native Method)\n at java.io.File.list(File.java:973)\n at java.io.File.listFiles(File.java:1051)\n at org.apache.hadoop.dfs.FSDataset$FSDir.getBlockInfo(FSDataset.java:153)\n at org.apache.hadoop.dfs.FSDataset$FSDir.getBlockInfo(FSDataset.java:149)\n at org.apache.hadoop.dfs.FSDataset$FSDir.getBlockInfo(FSDataset.java:149)\n at org.apache.hadoop.dfs.FSDataset$FSVolume.getBlockInfo(FSDataset.java:368)\n at org.apache.hadoop.dfs.FSDataset$FSVolumeSet.getBlockInfo(FSDataset.java:434)\n - locked <0x782f1348> (a org.apache.hadoop.dfs.FSDataset$FSVolumeSet)\n at org.apache.hadoop.dfs.FSDataset.getBlockReport(FSDataset.java:781)\n at org.apache.hadoop.dfs.DataNode.offerService(DataNode.java:642)\n at org.apache.hadoop.dfs.DataNode.run(DataNode.java:2431)\n at java.lang.Thread.run(Thread.java:619)\n{noformat}\n# locked 0x77ce9360 and waiting on 0x782f1348 {noformat}\n\"org.apache.hadoop.dfs.DataNode$DataXceiver@101f287\" daemon prio=10 tid=0x71906400 nid=0x5a93 waiting for monitor entry [0x712de000..0x712defc0]\n java.lang.Thread.State: BLOCKED (on object monitor)\n\tat org.apache.hadoop.dfs.FSDataset.writeToBlock(FSDataset.java:665)\n\t- waiting to lock <0x782f1348> (a org.apache.hadoop.dfs.FSDataset$FSVolumeSet)\n\t- locked <0x77ce9360> (a org.apache.hadoop.dfs.FSDataset)\n\tat org.apache.hadoop.dfs.DataNode$BlockReceiver.(DataNode.java:1995)\n\tat org.apache.hadoop.dfs.DataNode$DataXceiver.writeBlock(DataNode.java:1074)\n\tat org.apache.hadoop.dfs.DataNode$DataXceiver.run(DataNode.java:938)\n\tat java.lang.Thread.run(Thread.java:619)\n{noformat}\n# most other threads are waiting on 0x77ce9360\n\n\n\n","created":"2008-04-10T18:07:55.478+0000"},{"body":"Is is there any change in load on the cluster compared to when it was running 0.15?","created":"2008-04-10T18:10:27.936+0000"},{"body":"Here's the log file for that machine. \nHardware: 8 cores @ 1.86ghz, 8gb ram, 4x750gb 7200 prm disks with gigabit ethernet connections.\n\nExample iostat from a node that just lost contact with the namenode.\n\navg-cpu: %user %nice %sys %iowait %idle\n 18.88 0.00 1.60 0.99 78.53\n\nDevice: rrqm/s wrqm/s r/s w/s rsec/s wsec/s rkB/s wkB/s avgrq-sz avgqu-sz await svctm %util\nsda 0.55 0.72 4.77 4.42 837.81 1831.38 418.91 915.69 290.43 0.26 28.17 7.28 6.69\nsdb 0.46 0.87 4.42 3.92 771.67 1443.65 385.84 721.83 265.93 0.05 5.43 7.87 6.56\nsdc 0.54 0.59 4.59 3.72 823.05 1676.57 411.52 838.28 300.87 0.11 13.06 7.92 6.58\nsdd 0.45 0.63 4.32 3.51 756.46 1433.20 378.23 716.60 279.92 0.29 36.78 7.95 6.22","created":"2008-04-10T18:17:39.089+0000"},{"body":"Johan, you can enclose the formatted text inside {_noformat_} ... {_noformat_} so that it is easier to read.","created":"2008-04-10T20:31:37.610+0000"},{"body":"What is the exact iostat command you ran? Is it an average over last 5 seconds or so or overall average?\n\nFrom the data\n{noformat}\n2008-04-10 09:23:08,667 INFO org.apache.hadoop.dfs.DataNode: BlockReport of 497572 blocks got processed in 381086 msecs\n{noformat}\n\nA few things about this: \n- You have around 500K blocks.. mostly of them very small (verification time is very short). This is an order of magnitude larger than our datanodes have here at Yahoo.\n- The above log message says it too 6.5 min for block report. Most of this time I would think is for listing the files in the local directory.\n- Such a large number of blocks should cause similar problem with 0.15 also. Anything you think is different?\n\nThough DataNode should be able to handle larger number of blocks better, I don't think this is a blocker for 0.16.3 release (unless iostat shows something else). Do you agree?\n\nBtw, are you planning to have large number of small blocks? Its going to limit NameNode scalability.\n\n","created":"2008-04-10T21:44:39.830+0000"},{"body":"I am currently marking this for 0.18. 0.16.3 is close to be released and looks like any fix for this would not be trivial. Also it does not seem like a regression yet. But if it turns out to be different, we can make this a candidate for 0.16.4 or 0.17.\n","created":"2008-04-10T21:56:41.656+0000"},{"body":"Here's the second output from iostat -x 30 from a datanode that just lost contact.\n\n{noformat}\navg-cpu: %user %nice %sys %iowait %idle\n 94.26 0.00 0.83 0.00 4.91\n\nDevice: rrqm/s wrqm/s r/s w/s rsec/s wsec/s rkB/s wkB/s avgrq-sz avgqu-sz await svctm %util\nsda 0.40 0.20 137.45 5.10 2031.86 2680.44 1015.93 1340.22 33.06 1.93 13.53 6.72 95.73\nsdb 0.23 0.03 0.33 1.20 145.02 552.22 72.51 276.11 454.87 0.02 10.22 4.78 0.73\nsdc 0.50 0.03 0.70 8.76 307.10 7137.72 153.55 3568.86 786.69 4.70 496.94 6.90 6.53\nsdd 0.47 0.07 0.83 0.53 315.63 4.83 157.81 2.42 234.56 0.06 44.39 10.73 1.47\n{noformat}\n\nI'm aware that we have quite a lot of small files, it's an issue we're working on. I guess we'll have to ramp up the priority.\nWhat you're saying that the block reports are causing this makes sense. I'll merge a few of these directories of small files and see if it improves.\n\nPerhaps the difference between 0.15 and 0.16 was just that we did hit it quite hard after the upgrade as data had queued up,\nI'm going to see how it behaves today.","created":"2008-04-11T10:35:24.164+0000"},{"body":"We've reduced the number of blocks per node to about 150 000, but we're still seeing the same problem as before.","created":"2008-04-17T15:09:50.496+0000"},{"body":"can you check messages similar to \"[...]BlockReport of 497572 blocks got processed in 381086 msecs\" on such data nodes?","created":"2008-04-17T19:22:15.109+0000"},{"body":"One thing I am not sure yet is why your sda is so much busier than other disks.. somehow all the native filesystem metadata falls on one disk? Another possibility is swapping but size of each is very small (~512 bytes).","created":"2008-04-17T19:25:25.395+0000"},{"body":"> \"DataNode: [/var/storage/1/dfs/data,/var/storage/2/dfs/data,/var/storage/3/dfs/data,/var/storage/4/dfs/data]\" \nis each of /var/storage/[1-4] mounted on different disk or all are on sda?\n","created":"2008-04-17T19:28:17.534+0000"},{"body":"All those directories are on separate disks.\n\nOne thing I didn't think about was that even though we merged a lot of files to reduce the number of blocks there would still be tons of files in dfs/data/previous, since I had not run -finalizeUpgrade.\nSo if I'm not mistaken the du command would still go through all those, taking quite some time. I've now finalized the upgrade and we're seeing better performance: BlockReport of 147083 blocks got processed in 25432 msecs\n\nAlthough having the datanodes lose contact with the namenode because it's checking disk usage seems like quite a serious bug to me.\n\nnot sure why sda is more busy, although that is where the logs are located","created":"2008-04-18T09:25:02.220+0000"},{"body":"Yes, looks like DU goes through previous also... So instead of both block report and DU causing problems now it is DU..\n\nCan you clarify if \"datanodes lose contact\" means NameNode actually marks them \"dead\"?\n\n> Although having the datanodes lose contact with the namenode because it's checking disk usage seems like quite a serious bug to me.\n\nI agree. Doing these in the background without blocking normal DataNode functions takes a little bit of restructuring. We should keep this jira open.\n\n> not sure why sda is more busy, although that is where the logs are located\nThis might help your situation. If you find more info, please inform us.\n","created":"2008-04-18T18:21:20.352+0000"},{"body":"Created a simple patch that makes DU run the shell command in a new thread and never block on getUsage(). It does change the behavior a bit, but in most cases it shouldn't be a problem.\nI've not had a chance to test this on a real cluster yet, works in my local testing though.\n\nUnfortunately it turns out this only solves part of the problem. Block reports are still an issue.\nThat's harder to solve though, not sure how to restructure that. Ideas?","created":"2008-05-01T14:34:49.542+0000"},{"body":"Regd block reports, one option is not to do the fs scan at all.","created":"2008-05-01T14:53:14.838+0000"},{"body":"> Block reports are still an issue.\n\nWill incremental block reports (HADOOP-1079) resolve this?\n","created":"2008-05-01T17:09:27.212+0000"},{"body":"> Will incremental block reports (HADOOP-1079) resolve this?\nIf it avoids scanning the file system, yes. From the above jira it is not clear if it will avoid comparing in-memory map and on-disk blocks. \n\nEven as simple a change as commenting out fs scan is enough for this problem, I think.","created":"2008-05-01T17:39:21.010+0000"},{"body":"If I understand this correctly the purpose of a block report is to let the namenode know if a block mysteriously goes missing on the data node.\n\nWould it make sense to have this as part of the block verification mechanism described in HADOOP-2012? Is it already?\nThat way the block report could always be sent from memory and doesn't have to hit the disk.\n\nOf course just commenting out the fs scan as Raghu suggest would also work, how big is the issue of blocks going missing?","created":"2008-05-02T13:09:36.914+0000"},{"body":"Shall we fix the DU issue first, it seems easier. Could someone review my patch, please?\nThen we can create another ticket for the block report problem.","created":"2008-05-09T10:00:06.653+0000"},{"body":"regd the patch,\n> It does change the behavior a bit, but in most cases it shouldn't be a problem.\n\nI haven't looked at it properly yet, could you describe what the change in behavior is? Also not sure why it needs to change Shell stuff. Could the desired behavor for DF be implemented in DF class?\n","created":"2008-05-09T15:44:46.447+0000"},{"body":"> Also not sure why it needs to change Shell stuff. [ ... ]\n\nI think this was because Johan wanted to make DF implement Runnable, and there was a conflict, since Shell already has a method named 'run'. But this changes public APIs incompatibly, and is thus not a good approach.\n\nJohan, perhaps instead we could define a nested class in DF.java that extends Thread and overrides run() there, or implements Runnable, if you prefer.\n\nAlso the default interval is dfs.blockreport.intervalMsec, which seems rather long. It should really be related to the heartbeat interval, no? Moreover, we shouldn't use a DFS parameter in a generic FS class. So default interval should be something safe, perhaps hardwired to 10 minutes or somesuch, and DFS should override that when it constructs a DF, if it needs. Does that make sense?\n\nAnd perhaps the thread should run 'df' first, then sleep, so that values are available to clients sooner?\n\nFinally, and most imporant, what evidence do you have that DU is in fact causing problems? It is run very infrequently, not strictly synchronized with block reports but rather triggered by heartbeats. A slow DU would thus result in a delayed heartbeat. None of the stack traces above indicate that DU is blocking other activities of the datanode.","created":"2008-05-09T16:37:20.914+0000"},{"body":"Cleaned up the patch a bit, now only changes DU.java + the test","created":"2008-05-09T16:37:50.302+0000"},{"body":"Doug: You're right about the Runnable/run() bit, just as you wrote that I adapted the patch as you suggested.\n\nI agree about the interval, I'll change it.\nThis patch is for DU, the DF returns so quickly that it shouldn't cause an issue.\n\nIn the DU constructor the command is run once so that we get values straight away, I thought this would be better since then we know for sure there's correct values in there once the object is created.\n\nI'll try to recreate the situation to produce good evidence, but off the top of my head DU is used to decide what volume to write to in writeToBlock in FSDataset, so it causes problems with writing blocks if it takes too long. We've seen quite a lot of this.\nAs you say it doesn't run that often, but often enough to cause us problems.","created":"2008-05-09T16:56:48.691+0000"},{"body":"This addresses the first of my above concerns, but not the other three.\n\nAlso, does that test in fact succeed? It seems to me that the thread will not have yet had a chance to update its usage count before the assertEquals(), no?\n","created":"2008-05-09T16:59:12.611+0000"},{"body":"Oops. I posted my previous comment before I saw your last comment. You're right, I was confusing DU and DF. And I missed the refresh in the ctor. Sorry! That resolves most of my concerns. Since this is called much more frequently than I was thinking, it makes sense for it not to be synchronous. I think my only remaining concern is the default interval, which you've said you'd address. Thanks!","created":"2008-05-09T17:07:11.468+0000"},{"body":"\nIt was my mistake to say 'DF' where I meant 'DU'.\n\nPatch looks good. Couple of comments :\n\n- it need not disallow interval of zero. In that case, you could just not start the thread and invoke run() as before. Since DU is a utilitiy it is used (or could be used) outside DataNode.\n\n- persistent thread : in normal case, since the thread works only once in a while, it could be created only when it needs to run. This would be an improvement, I don't mean it as a hard requirement for this patch. This is more inline with the prev behaviour since if getUsed() is not called, then there is no penalty. \n\nAre you using this (or prev) patch in your environment?\n\nThis certainly improves DN stability with large number of blocks. We still need to keep in mind that DU has very noticeable penalty. Say it takes around 10min (as in your case)... then it implies 15% of the time DN will be extremely I/O starved. This will have very noticeable affect on I/O intensive applications. ","created":"2008-05-09T19:29:13.324+0000"},{"body":"Updated patch with the suggestions from Doug and Raghu.\nInterval defaults to 10min. If the incoming interval is 0 the previous behavior is used.\n\nFindbugs doesn't like that I start a thread in the constructor, but afaik it's the only way without adding a start method to the class and I assume you don't want to change the interface.\nThis passes all the tests+checkstyle on my local machine. Also added a bunch of javadoc.\n\nRaghu: I am using a previous patch on our cluster yes, no problems so far.\nI'm not sure what you mean by not having a permanent thread. How would we update the value without blocking on getUsed in that case? You say it's not a hard requirement, I hope you can accept the patch anyway.\n\n","created":"2008-05-13T13:40:37.224+0000"},{"body":"> I'm not sure what you mean by not having a permanent thread. How would we update the value without blocking on getUsed in that case?\n\nThis thread will stay idle pretty much most of the time. So we could start a thread inside getUsed() (and possibly in other accessor methods if interval has passed) and make the thread exit after running du. This is no less accurate than current implementation. Or you could schedule a periodic thread using Java 'Executor'. Even if you do keep the persistent thread, could you add comment if you agree that it need not be persistent.. we might implement that later.\n\nRegd the patch: \n# 'lock' is not required. You can synchronize on DU.this. \n# Also DURefreshThread should either be static class or not keep a ref to du. ","created":"2008-05-14T08:04:04.463+0000"},{"body":"1. Changed\n2. Removed the reference to DU and call DU.this.run() instead.\n\nScheduling it with an Executor would still have a running thread that keeps track of the scheduling, no? ScheduledThreadPoolExecutor for example.\nThe other example was starting a thread inside getUsed() then returning the old value while the DU runs? Wouldn't that cause problems where getUsed is called very infrequently? Perhaps this is never an issue?\n\nAnyway, I added the comment about improving this with a non permanent thread anyway, in the most common case I guess that would work fine.\n\nAs mentioned this patch will cause one new findbugs error, starting a thread in the constructor, hard to avoid without having a new method to start it, breaking the public interface.","created":"2008-05-14T11:20:31.781+0000"},{"body":"\n> Scheduling it with an Executor would still have a running thread that keeps track of the scheduling, no? \nThat is implementation dependent. JavaDoc does not say. There might be just one thread that handles many such tasks. We at least let the implementation of Executors to optimize it.\n\nDU looks like simple utility to users. If they do create these multiple times, each when ever they need, they will be surprised to see threads hanging around. In that sense, it might be better to add a \"start()\" method so that it is explicit to them that a thread might be started and it needs to be shutdown(). \n\n+1 over all.","created":"2008-05-14T15:31:22.218+0000"},{"body":"Actually requiring start() is not so bad. Also DataNodes could invoke shutdown() when in side its shutdown. It will also remove the findbugs warning.\n","created":"2008-05-14T16:41:48.965+0000"},{"body":"Patch failed, my bad, ran all my tests with java6 so didn't catch the the @Overrides that eclipse likes to put in. It doesn't work well with java5.\n\nRaghu: Sure, I can add a start method if you guys want one.","created":"2008-05-14T17:06:50.806+0000"},{"body":"Updated patch with start and shutdown methods. If start isn't invoked the previous behavior of running on demand will be executed.\n\nI'm having issues running the unit tests on my machine on a clean trunk, TestDatanodeBlockScanner fails, so I hope this one passes all the tests.","created":"2008-05-20T10:23:11.634+0000"},{"body":"> TestDatanodeBlockScanner fails,.. \n\nPlease update your trunk, a fix for this was committed yesterday. If it still fails on your machine, let us know. ","created":"2008-05-20T15:14:06.033+0000"},{"body":"+1. Thanks for multiple iterations. If a user does not call start(), the interval given will be silently ignored.. in that sense, interval could be argument to start(). But I will commit v6.path for now.\n","created":"2008-05-22T18:02:28.593+0000"},{"body":"I just committed this. Thanks Johan!","created":"2008-05-22T18:22:59.820+0000"},{"body":"I propose an alternate solution for this.\nIf the block information was managed by having a inotify task (in linux/solaris), and the windows equivalent which I forget, the datanode could be informed each time a file in the dfs tree is created, updated, or deleted.\n\nWith this information being delivered, it can maintain an accurate block map with only 1 full scan of the datanode blocks, at start time.\n\nWith this algorithm the data nodes will be able to scale to a much larger number of blocks.\n\nThe other thing is the way the sync blocks on the FSDataset.FSVolumeSet are held totally aggravates this bug in 0.18.1.\n\nThe jason@attributor.com address will be going away shortly, I will be switching to jason.hadoop@gmail.com in the next little bit.","created":"2009-01-10T04:56:54.891+0000"},{"body":"Sounds very interesting. I suggest opening a new ticket describing the solution in more detail so the discussion can continue there instead of in this closed ticket.","created":"2009-01-12T10:29:35.633+0000"}],"conversations":[{"body":"I recently upgraded to 0.16.2 from 0.15.2 on our 10 node cluster.\nUnfortunately we're seeing datanode timeout issues. In previous versions we've often seen in the nn webui that one or two datanodes \"last contact\" goes from the usual 0-3 sec to ~200-300 before it drops down to 0 again.\n\nThis causes mild discomfort but the big problems appear when all nodes do this at once, as happened a few times after the upgrade.\nIt was suggested that this could be due to namenode garbage collection, but looking at the gc log output it doesn't seem to be the case.","from":"reporter","subject":"Datanodes time out"},{"body":"Example client exception:\n\n08/04/10 10:09:27 INFO dfs.DFSClient: Exception in createBlockOutputStream java.io.IOException: Bad connect ack with firstBadLink 10.0.5.76:50010\n08/04/10 10:09:27 INFO dfs.DFSClient: Abandoning block blk_-5192866954303337577\n08/04/10 10:09:27 INFO dfs.DFSClient: Waiting to find target node: 10.0.5.70:50010\n\nDatanode log output from around the same time:\n\n2008-04-10 10:09:15,758 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 got response for connect ack from downstream datanode with firstbadlink as 10.0.5.76:50010\n2008-04-10 10:09:15,758 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 forwarding connect ack to upstream firstbadlink is 10.0.5.76:50010\n2008-04-10 10:09:15,758 INFO org.apache.hadoop.dfs.DataNode: PacketResponder blk_-9149681676832791404 2 Exception java.io.EOFException\n at java.io.DataInputStream.readFully(DataInputStream.java:180)\n at java.io.DataInputStream.readLong(DataInputStream.java:399)\n at org.apache.hadoop.dfs.DataNode$PacketResponder.run(DataNode.java:1822)\n at java.lang.Thread.run(Thread.java:619)\n\n2008-04-10 10:09:15,758 INFO org.apache.hadoop.dfs.DataNode: PacketResponder 2 for block blk_-9149681676832791404 terminating\n2008-04-10 10:09:16,793 INFO org.apache.hadoop.dfs.DataNode: writeBlock blk_1511572447827516117 received exception java.net.SocketTimeoutException: Read timed out\n2008-04-10 10:09:16,793 ERROR org.apache.hadoop.dfs.DataNode: 10.0.5.70:50010:DataXceiver: java.net.SocketTimeoutException: Read timed out\n at java.net.SocketInputStream.socketRead0(Native Method)\n at java.net.SocketInputStream.read(SocketInputStream.java:129)\n at java.net.SocketInputStream.read(SocketInputStream.java:182)\n at java.io.DataInputStream.readByte(DataInputStream.java:248)\n at org.apache.hadoop.io.WritableUtils.readVLong(WritableUtils.java:324)\n at org.apache.hadoop.io.WritableUtils.readVInt(WritableUtils.java:346)\n at org.apache.hadoop.io.Text.readString(Text.java:413)\n at org.apache.hadoop.dfs.DataNode$DataXceiver.writeBlock(DataNode.java:1117)\n at org.apache.hadoop.dfs.DataNode$DataXceiver.run(DataNode.java:938)\n at java.lang.Thread.run(Thread.java:619)\n\n2008-04-10 10:09:23,895 INFO org.apache.hadoop.dfs.DataNode: Receiving block blk_-7442302015902809712 src: /10.0.5.74:55546 dest: /10.0.5.74:50010\n2008-04-10 10:09:23,896 INFO org.apache.hadoop.dfs.DataNode: Datanode 0 forwarding connect ack to upstream firstbadlink is\n2008-04-10 10:09:23,927 INFO org.apache.hadoop.dfs.DataNode: Received block blk_-7442302015902809712 of size 902 from /10.0.5.74\n2008-04-10 10:09:23,928 INFO org.apache.hadoop.dfs.DataNode: PacketResponder 0 for block blk_-7442302015902809712 terminating\n2008-04-10 10:09:23,937 INFO org.apache.hadoop.dfs.DataNode: Receiving block blk_4161972554165500020 src: /10.0.11.7:41256 dest: /10.0.11.7:50010\n2008-04-10 10:09:23,968 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 got response for connect ack from downstream datanode with firstbadlink as\n2008-04-10 10:09:23,969 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 forwarding connect ack to upstream firstbadlink is\n2008-04-10 10:09:24,034 INFO org.apache.hadoop.dfs.DataNode: Received block blk_4161972554165500020 of size 95 from /10.0.11.7\n2008-04-10 10:09:24,034 INFO org.apache.hadoop.dfs.DataNode: PacketResponder 2 for block blk_4161972554165500020 terminating\n2008-04-10 10:09:27,664 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 got response for connect ack from downstream datanode with firstbadlink as 10.0.5.76:50010\n2008-04-10 10:09:27,664 INFO org.apache.hadoop.dfs.DataNode: Datanode 2 forwarding connect ack to upstream firstbadlink is 10.0.5.76:50010\n2008-04-10 10:09:27,665 INFO org.apache.hadoop.dfs.DataNode: PacketResponder blk_-5192866954303337577 2 Exception java.io.EOFException\n at java.io.DataInputStream.readFully(DataInputStream.java:180)\n at java.io.DataInputStream.readLong(DataInputStream.java:399)\n at org.apache.hadoop.dfs.DataNode$PacketResponder.run(DataNode.java:1822)\n at java.lang.Thread.run(Thread.java:619)\n\n","from":"developer"},{"body":"Stacktrace of a datanode that lost contact with the namenode. Not sure it helps but doesn't hurt.","from":"developer"},{"body":"Namenode stacktrace","from":"developer"},{"body":"Found 68 of these when doing a stackdump in a datanode that had lost contact with the namenode.\n\n\"org.apache.hadoop.dfs.DataNode$DataXceiver@5b02a6\" daemon prio=10 tid=0x7192bc00 nid=0x5bc6 waiting for monitor entry [0x70cee000..0x70cef0c0]\n java.lang.Thread.State: BLOCKED (on object monitor)\n at org.apache.hadoop.dfs.FSDataset.getFile(FSDataset.java:867)\n - waiting to lock <0x77ce9360> (a org.apache.hadoop.dfs.FSDataset)\n at org.apache.hadoop.dfs.FSDataset.isValidBlock(FSDataset.java:795)\n at org.apache.hadoop.dfs.FSDataset.writeToBlock(FSDataset.java:614)\n at org.apache.hadoop.dfs.DataNode$BlockReceiver.(DataNode.java:1995)\n at org.apache.hadoop.dfs.DataNode$DataXceiver.writeBlock(DataNode.java:1074)\n at org.apache.hadoop.dfs.DataNode$DataXceiver.run(DataNode.java:938)\n at java.lang.Thread.run(Thread.java:619)","from":"developer"},{"body":"> Found 68 of these when doing a stackdump in a datanode that had lost contact with the namenode.\n\nCould you attach the full stack trace of this datanode?","from":"developer"},{"body":"As requested","from":"developer"},{"body":"\nWhile you are at it, could you attach log (.log file) for this datanode as well? The log file show what activity is going on now. Also list approximate time the stack trace was taken.\n\nMy observation is that you are writing a lot of blocks.. and the datanode that looks blocked is blocked while listing all the blocks on the native filesystem. It does this every hour when it sends block reports. Till now nothing suspicious other than heavy write traffic and slow disks. Check iostat on the machine. What is the hardware like?\n\nThe two main threads from the DataNode are :\n\n# locks 0x782f1348 : {noformat}\n\"DataNode: [/var/storage/1/dfs/data,/var/storage/2/dfs/data,/var/storage/3/dfs/data,/var/storage/4/dfs/data]\" daemon prio=10 tid=0x72409000\n nid=0x44f6 runnable [0x71d8a000..0x71d8aec0]\n java.lang.Thread.State: RUNNABLE\n at java.io.UnixFileSystem.list(Native Method)\n at java.io.File.list(File.java:973)\n at java.io.File.listFiles(File.java:1051)\n at org.apache.hadoop.dfs.FSDataset$FSDir.getBlockInfo(FSDataset.java:153)\n at org.apache.hadoop.dfs.FSDataset$FSDir.getBlockInfo(FSDataset.java:149)\n at org.apache.hadoop.dfs.FSDataset$FSDir.getBlockInfo(FSDataset.java:149)\n at org.apache.hadoop.dfs.FSDataset$FSVolume.getBlockInfo(FSDataset.java:368)\n at org.apache.hadoop.dfs.FSDataset$FSVolumeSet.getBlockInfo(FSDataset.java:434)\n - locked <0x782f1348> (a org.apache.hadoop.dfs.FSDataset$FSVolumeSet)\n at org.apache.hadoop.dfs.FSDataset.getBlockReport(FSDataset.java:781)\n at org.apache.hadoop.dfs.DataNode.offerService(DataNode.java:642)\n at org.apache.hadoop.dfs.DataNode.run(DataNode.java:2431)\n at java.lang.Thread.run(Thread.java:619)\n{noformat}\n# locked 0x77ce9360 and waiting on 0x782f1348 {noformat}\n\"org.apache.hadoop.dfs.DataNode$DataXceiver@101f287\" daemon prio=10 tid=0x71906400 nid=0x5a93 waiting for monitor entry [0x712de000..0x712defc0]\n java.lang.Thread.State: BLOCKED (on object monitor)\n\tat org.apache.hadoop.dfs.FSDataset.writeToBlock(FSDataset.java:665)\n\t- waiting to lock <0x782f1348> (a org.apache.hadoop.dfs.FSDataset$FSVolumeSet)\n\t- locked <0x77ce9360> (a org.apache.hadoop.dfs.FSDataset)\n\tat org.apache.hadoop.dfs.DataNode$BlockReceiver.(DataNode.java:1995)\n\tat org.apache.hadoop.dfs.DataNode$DataXceiver.writeBlock(DataNode.java:1074)\n\tat org.apache.hadoop.dfs.DataNode$DataXceiver.run(DataNode.java:938)\n\tat java.lang.Thread.run(Thread.java:619)\n{noformat}\n# most other threads are waiting on 0x77ce9360\n\n\n\n","from":"developer"},{"body":"Is is there any change in load on the cluster compared to when it was running 0.15?","from":"developer"},{"body":"Here's the log file for that machine. \nHardware: 8 cores @ 1.86ghz, 8gb ram, 4x750gb 7200 prm disks with gigabit ethernet connections.\n\nExample iostat from a node that just lost contact with the namenode.\n\navg-cpu: %user %nice %sys %iowait %idle\n 18.88 0.00 1.60 0.99 78.53\n\nDevice: rrqm/s wrqm/s r/s w/s rsec/s wsec/s rkB/s wkB/s avgrq-sz avgqu-sz await svctm %util\nsda 0.55 0.72 4.77 4.42 837.81 1831.38 418.91 915.69 290.43 0.26 28.17 7.28 6.69\nsdb 0.46 0.87 4.42 3.92 771.67 1443.65 385.84 721.83 265.93 0.05 5.43 7.87 6.56\nsdc 0.54 0.59 4.59 3.72 823.05 1676.57 411.52 838.28 300.87 0.11 13.06 7.92 6.58\nsdd 0.45 0.63 4.32 3.51 756.46 1433.20 378.23 716.60 279.92 0.29 36.78 7.95 6.22","from":"developer"},{"body":"Johan, you can enclose the formatted text inside {_noformat_} ... {_noformat_} so that it is easier to read.","from":"developer"},{"body":"What is the exact iostat command you ran? Is it an average over last 5 seconds or so or overall average?\n\nFrom the data\n{noformat}\n2008-04-10 09:23:08,667 INFO org.apache.hadoop.dfs.DataNode: BlockReport of 497572 blocks got processed in 381086 msecs\n{noformat}\n\nA few things about this: \n- You have around 500K blocks.. mostly of them very small (verification time is very short). This is an order of magnitude larger than our datanodes have here at Yahoo.\n- The above log message says it too 6.5 min for block report. Most of this time I would think is for listing the files in the local directory.\n- Such a large number of blocks should cause similar problem with 0.15 also. Anything you think is different?\n\nThough DataNode should be able to handle larger number of blocks better, I don't think this is a blocker for 0.16.3 release (unless iostat shows something else). Do you agree?\n\nBtw, are you planning to have large number of small blocks? Its going to limit NameNode scalability.\n\n","from":"developer"},{"body":"I am currently marking this for 0.18. 0.16.3 is close to be released and looks like any fix for this would not be trivial. Also it does not seem like a regression yet. But if it turns out to be different, we can make this a candidate for 0.16.4 or 0.17.\n","from":"developer"},{"body":"Here's the second output from iostat -x 30 from a datanode that just lost contact.\n\n{noformat}\navg-cpu: %user %nice %sys %iowait %idle\n 94.26 0.00 0.83 0.00 4.91\n\nDevice: rrqm/s wrqm/s r/s w/s rsec/s wsec/s rkB/s wkB/s avgrq-sz avgqu-sz await svctm %util\nsda 0.40 0.20 137.45 5.10 2031.86 2680.44 1015.93 1340.22 33.06 1.93 13.53 6.72 95.73\nsdb 0.23 0.03 0.33 1.20 145.02 552.22 72.51 276.11 454.87 0.02 10.22 4.78 0.73\nsdc 0.50 0.03 0.70 8.76 307.10 7137.72 153.55 3568.86 786.69 4.70 496.94 6.90 6.53\nsdd 0.47 0.07 0.83 0.53 315.63 4.83 157.81 2.42 234.56 0.06 44.39 10.73 1.47\n{noformat}\n\nI'm aware that we have quite a lot of small files, it's an issue we're working on. I guess we'll have to ramp up the priority.\nWhat you're saying that the block reports are causing this makes sense. I'll merge a few of these directories of small files and see if it improves.\n\nPerhaps the difference between 0.15 and 0.16 was just that we did hit it quite hard after the upgrade as data had queued up,\nI'm going to see how it behaves today.","from":"developer"},{"body":"We've reduced the number of blocks per node to about 150 000, but we're still seeing the same problem as before.","from":"developer"},{"body":"can you check messages similar to \"[...]BlockReport of 497572 blocks got processed in 381086 msecs\" on such data nodes?","from":"developer"},{"body":"One thing I am not sure yet is why your sda is so much busier than other disks.. somehow all the native filesystem metadata falls on one disk? Another possibility is swapping but size of each is very small (~512 bytes).","from":"developer"},{"body":"> \"DataNode: [/var/storage/1/dfs/data,/var/storage/2/dfs/data,/var/storage/3/dfs/data,/var/storage/4/dfs/data]\" \nis each of /var/storage/[1-4] mounted on different disk or all are on sda?\n","from":"developer"},{"body":"All those directories are on separate disks.\n\nOne thing I didn't think about was that even though we merged a lot of files to reduce the number of blocks there would still be tons of files in dfs/data/previous, since I had not run -finalizeUpgrade.\nSo if I'm not mistaken the du command would still go through all those, taking quite some time. I've now finalized the upgrade and we're seeing better performance: BlockReport of 147083 blocks got processed in 25432 msecs\n\nAlthough having the datanodes lose contact with the namenode because it's checking disk usage seems like quite a serious bug to me.\n\nnot sure why sda is more busy, although that is where the logs are located","from":"developer"},{"body":"Yes, looks like DU goes through previous also... So instead of both block report and DU causing problems now it is DU..\n\nCan you clarify if \"datanodes lose contact\" means NameNode actually marks them \"dead\"?\n\n> Although having the datanodes lose contact with the namenode because it's checking disk usage seems like quite a serious bug to me.\n\nI agree. Doing these in the background without blocking normal DataNode functions takes a little bit of restructuring. We should keep this jira open.\n\n> not sure why sda is more busy, although that is where the logs are located\nThis might help your situation. If you find more info, please inform us.\n","from":"developer"},{"body":"Created a simple patch that makes DU run the shell command in a new thread and never block on getUsage(). It does change the behavior a bit, but in most cases it shouldn't be a problem.\nI've not had a chance to test this on a real cluster yet, works in my local testing though.\n\nUnfortunately it turns out this only solves part of the problem. Block reports are still an issue.\nThat's harder to solve though, not sure how to restructure that. Ideas?","from":"developer"},{"body":"Regd block reports, one option is not to do the fs scan at all.","from":"developer"},{"body":"> Block reports are still an issue.\n\nWill incremental block reports (HADOOP-1079) resolve this?\n","from":"developer"},{"body":"> Will incremental block reports (HADOOP-1079) resolve this?\nIf it avoids scanning the file system, yes. From the above jira it is not clear if it will avoid comparing in-memory map and on-disk blocks. \n\nEven as simple a change as commenting out fs scan is enough for this problem, I think.","from":"developer"},{"body":"If I understand this correctly the purpose of a block report is to let the namenode know if a block mysteriously goes missing on the data node.\n\nWould it make sense to have this as part of the block verification mechanism described in HADOOP-2012? Is it already?\nThat way the block report could always be sent from memory and doesn't have to hit the disk.\n\nOf course just commenting out the fs scan as Raghu suggest would also work, how big is the issue of blocks going missing?","from":"developer"},{"body":"Shall we fix the DU issue first, it seems easier. Could someone review my patch, please?\nThen we can create another ticket for the block report problem.","from":"developer"},{"body":"regd the patch,\n> It does change the behavior a bit, but in most cases it shouldn't be a problem.\n\nI haven't looked at it properly yet, could you describe what the change in behavior is? Also not sure why it needs to change Shell stuff. Could the desired behavor for DF be implemented in DF class?\n","from":"developer"},{"body":"> Also not sure why it needs to change Shell stuff. [ ... ]\n\nI think this was because Johan wanted to make DF implement Runnable, and there was a conflict, since Shell already has a method named 'run'. But this changes public APIs incompatibly, and is thus not a good approach.\n\nJohan, perhaps instead we could define a nested class in DF.java that extends Thread and overrides run() there, or implements Runnable, if you prefer.\n\nAlso the default interval is dfs.blockreport.intervalMsec, which seems rather long. It should really be related to the heartbeat interval, no? Moreover, we shouldn't use a DFS parameter in a generic FS class. So default interval should be something safe, perhaps hardwired to 10 minutes or somesuch, and DFS should override that when it constructs a DF, if it needs. Does that make sense?\n\nAnd perhaps the thread should run 'df' first, then sleep, so that values are available to clients sooner?\n\nFinally, and most imporant, what evidence do you have that DU is in fact causing problems? It is run very infrequently, not strictly synchronized with block reports but rather triggered by heartbeats. A slow DU would thus result in a delayed heartbeat. None of the stack traces above indicate that DU is blocking other activities of the datanode.","from":"developer"},{"body":"Cleaned up the patch a bit, now only changes DU.java + the test","from":"developer"},{"body":"Doug: You're right about the Runnable/run() bit, just as you wrote that I adapted the patch as you suggested.\n\nI agree about the interval, I'll change it.\nThis patch is for DU, the DF returns so quickly that it shouldn't cause an issue.\n\nIn the DU constructor the command is run once so that we get values straight away, I thought this would be better since then we know for sure there's correct values in there once the object is created.\n\nI'll try to recreate the situation to produce good evidence, but off the top of my head DU is used to decide what volume to write to in writeToBlock in FSDataset, so it causes problems with writing blocks if it takes too long. We've seen quite a lot of this.\nAs you say it doesn't run that often, but often enough to cause us problems.","from":"developer"},{"body":"This addresses the first of my above concerns, but not the other three.\n\nAlso, does that test in fact succeed? It seems to me that the thread will not have yet had a chance to update its usage count before the assertEquals(), no?\n","from":"developer"},{"body":"Oops. I posted my previous comment before I saw your last comment. You're right, I was confusing DU and DF. And I missed the refresh in the ctor. Sorry! That resolves most of my concerns. Since this is called much more frequently than I was thinking, it makes sense for it not to be synchronous. I think my only remaining concern is the default interval, which you've said you'd address. Thanks!","from":"developer"},{"body":"\nIt was my mistake to say 'DF' where I meant 'DU'.\n\nPatch looks good. Couple of comments :\n\n- it need not disallow interval of zero. In that case, you could just not start the thread and invoke run() as before. Since DU is a utilitiy it is used (or could be used) outside DataNode.\n\n- persistent thread : in normal case, since the thread works only once in a while, it could be created only when it needs to run. This would be an improvement, I don't mean it as a hard requirement for this patch. This is more inline with the prev behaviour since if getUsed() is not called, then there is no penalty. \n\nAre you using this (or prev) patch in your environment?\n\nThis certainly improves DN stability with large number of blocks. We still need to keep in mind that DU has very noticeable penalty. Say it takes around 10min (as in your case)... then it implies 15% of the time DN will be extremely I/O starved. This will have very noticeable affect on I/O intensive applications. ","from":"developer"},{"body":"Updated patch with the suggestions from Doug and Raghu.\nInterval defaults to 10min. If the incoming interval is 0 the previous behavior is used.\n\nFindbugs doesn't like that I start a thread in the constructor, but afaik it's the only way without adding a start method to the class and I assume you don't want to change the interface.\nThis passes all the tests+checkstyle on my local machine. Also added a bunch of javadoc.\n\nRaghu: I am using a previous patch on our cluster yes, no problems so far.\nI'm not sure what you mean by not having a permanent thread. How would we update the value without blocking on getUsed in that case? You say it's not a hard requirement, I hope you can accept the patch anyway.\n\n","from":"developer"},{"body":"> I'm not sure what you mean by not having a permanent thread. How would we update the value without blocking on getUsed in that case?\n\nThis thread will stay idle pretty much most of the time. So we could start a thread inside getUsed() (and possibly in other accessor methods if interval has passed) and make the thread exit after running du. This is no less accurate than current implementation. Or you could schedule a periodic thread using Java 'Executor'. Even if you do keep the persistent thread, could you add comment if you agree that it need not be persistent.. we might implement that later.\n\nRegd the patch: \n# 'lock' is not required. You can synchronize on DU.this. \n# Also DURefreshThread should either be static class or not keep a ref to du. ","from":"developer"},{"body":"1. Changed\n2. Removed the reference to DU and call DU.this.run() instead.\n\nScheduling it with an Executor would still have a running thread that keeps track of the scheduling, no? ScheduledThreadPoolExecutor for example.\nThe other example was starting a thread inside getUsed() then returning the old value while the DU runs? Wouldn't that cause problems where getUsed is called very infrequently? Perhaps this is never an issue?\n\nAnyway, I added the comment about improving this with a non permanent thread anyway, in the most common case I guess that would work fine.\n\nAs mentioned this patch will cause one new findbugs error, starting a thread in the constructor, hard to avoid without having a new method to start it, breaking the public interface.","from":"developer"},{"body":"\n> Scheduling it with an Executor would still have a running thread that keeps track of the scheduling, no? \nThat is implementation dependent. JavaDoc does not say. There might be just one thread that handles many such tasks. We at least let the implementation of Executors to optimize it.\n\nDU looks like simple utility to users. If they do create these multiple times, each when ever they need, they will be surprised to see threads hanging around. In that sense, it might be better to add a \"start()\" method so that it is explicit to them that a thread might be started and it needs to be shutdown(). \n\n+1 over all.","from":"developer"},{"body":"Actually requiring start() is not so bad. Also DataNodes could invoke shutdown() when in side its shutdown. It will also remove the findbugs warning.\n","from":"developer"},{"body":"Patch failed, my bad, ran all my tests with java6 so didn't catch the the @Overrides that eclipse likes to put in. It doesn't work well with java5.\n\nRaghu: Sure, I can add a start method if you guys want one.","from":"developer"},{"body":"Updated patch with start and shutdown methods. If start isn't invoked the previous behavior of running on demand will be executed.\n\nI'm having issues running the unit tests on my machine on a clean trunk, TestDatanodeBlockScanner fails, so I hope this one passes all the tests.","from":"developer"},{"body":"> TestDatanodeBlockScanner fails,.. \n\nPlease update your trunk, a fix for this was committed yesterday. If it still fails on your machine, let us know. ","from":"developer"},{"body":"+1. Thanks for multiple iterations. If a user does not call start(), the interval given will be silently ignored.. in that sense, interval could be argument to start(). But I will commit v6.path for now.\n","from":"developer"},{"body":"I just committed this. Thanks Johan!","from":"developer"},{"body":"I propose an alternate solution for this.\nIf the block information was managed by having a inotify task (in linux/solaris), and the windows equivalent which I forget, the datanode could be informed each time a file in the dfs tree is created, updated, or deleted.\n\nWith this information being delivered, it can maintain an accurate block map with only 1 full scan of the datanode blocks, at start time.\n\nWith this algorithm the data nodes will be able to scale to a much larger number of blocks.\n\nThe other thing is the way the sync blocks on the FSDataset.FSVolumeSet are held totally aggravates this bug in 0.18.1.\n\nThe jason@attributor.com address will be going away shortly, I will be switching to jason.hadoop@gmail.com in the next little bit.","from":"developer"},{"body":"Sounds very interesting. I suggest opening a new ticket describing the solution in more detail so the discussion can continue there instead of in this closed ticket.","from":"developer"}],"created":"2008-04-10T14:57:31.000+0000","description":"I recently upgraded to 0.16.2 from 0.15.2 on our 10 node cluster.\nUnfortunately we're seeing datanode timeout issues. In previous versions we've often seen in the nn webui that one or two datanodes \"last contact\" goes from the usual 0-3 sec to ~200-300 before it drops down to 0 again.\n\nThis causes mild discomfort but the big problems appear when all nodes do this at once, as happened a few times after the upgrade.\nIt was suggested that this could be due to namenode garbage collection, but looking at the gc log output it doesn't seem to be the case.","issue_id":"12393665","key":"HADOOP-3232","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-05-22T18:22:59.000+0000","role":"fixed_distractor","summary":"Datanodes time out"} {"case_id":"12396742","cluster":"DISTRACTOR-HADOOP-3442","comments":[{"body":"\nWhen run gridmix with the latest 0.17 code, some job failed with the mappers throwing the following exception:\n\njava.lang.StackOverflowError\n\tat org.apache.hadoop.util.QuickSort.fix(QuickSort.java:29)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:58)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n","created":"2008-05-23T21:51:40.820+0000"},{"body":"II I understand the code correctly, it seems that if the input data is already sorted, the quick sort will behave badly -- recurrsion depth will be equal to the data size.\n","created":"2008-05-23T22:22:30.686+0000"},{"body":"I haven't verified this yet, but if what Runping is saying is true, does it warrant a 0.17.1 release?","created":"2008-05-26T03:46:27.582+0000"},{"body":"The recursion will always be on the smaller of the two partitions first, so even a degenerate case should not overflow the stack. I ran the testcase with 10 * 2^17 elements (default 5% of the default 100MB allocated to record storage in MapTask), 10x that, sorted, reversed, all equal, etc. and could not reproduce this. However, if the comparator were flawed, that could cause the behavior you're seeing.","created":"2008-05-28T18:44:50.556+0000"},{"body":"It's not sorted data (which should produce relatively even partitions per lines 58-60), but horribly unlucky pivot selections that could produce this. Since the JVM doesn't appear to recognize the tail recursion, it's possible that a series- meaning hundreds- of sequential median values sampled from the left, right, and midpoint could be so skewed that the recursion could top out (a bad comparator might hurt pivot selection). This should be _fantastically_ unlikely, though. Does this happen with HADOOP-3308 applied?","created":"2008-05-28T22:21:02.297+0000"},{"body":"Runping, are you able to reproduce the problem?\n\nFor already sorted list, I think the codes will pick the median (i.e. the best choice). I suspect that the bug, if there is any, is outside util.QuickSort.","created":"2008-05-29T00:59:10.060+0000"},{"body":"I've attached a patch to remove the recursion. On my machine, this actually made things faster too. Not sure if that's real or just a quirk with my test case.","created":"2008-05-31T00:33:21.181+0000"},{"body":"I agree with the comments that the problem was probably outside of the quicksort. You could also check your java stack size to be sure it's not somehow set _really_ low.","created":"2008-05-31T00:44:25.158+0000"},{"body":"This may be useful to add: http://en.wikipedia.org/wiki/Introsort (essentially, switch from quicksort to heapsort once for subsorts beyond a certain depth).","created":"2008-06-01T00:41:37.084+0000"},{"body":"Alternatively, we could just effect the tail recursion with a loop (see attached; credit to Nicholas for the idea). I'm not wild about creating new java.util.Stack instances, but it's not likely to be a performance issue either way. As for switching to heapsort from insertion sort for < K or log\\(n) values, though I'd be surprised to learn that was a bottleneck, it sounds like a cool idea to try.\n\nThe recursion itself doesn't seem to be the problem, though; that there's a case where it becomes unbounded is. This is showing up often enough that there is likely a framework problem, but nobody has explained how they reproduce it, yet.","created":"2008-06-02T21:10:11.115+0000"},{"body":"Can you try testing my version in your environment to see if you get a performance change? I'm thinking that there may be some JIT thing going on that makes the no-recursion version better. I did get a significant (like, 20%+ less time), repeatable, performance gain when using my version.\n\nIntrosort still uses insertionSort instead of quicksort for low n. It uses heapsort when quicksort has recursed deeply; unlike quicksort, heapsort has guaranteed worst-case performance of O(n log n), avoiding both the huge recursion (it's not recursive) and the O(n*n) runtime that such a recursion produces. But it's generally slower, so introsort uses it only when quicksort has hit a bad case.\n\nThanks!","created":"2008-06-02T21:22:47.938+0000"},{"body":"bq. Can you try testing my version in your environment to see if you get a performance change?\n\nI tried running the test case with 10M elements and saw only very slight performance improvements over the original, excluding the sequential test case, where it degraded. I'm running 1.5.0_13 on a Mac, so YMMV.\n\nbq. Introsort still uses insertionSort instead of quicksort for low n. It uses heapsort when quicksort has recursed deeply\n\nAh, I understand. Still, the issue here is not performance, but correctness. Until we verify that the implementation of QuickSort is correct but pushed into a degenerate case, I think we should assume that some aspect of the framework is incorrect and try to determine why we're topping out the stack. Once we isolate the issue, then we can look into new sorting strategies if they solve the problem, but I'm not sure we can say without fear of refutation that the recursion alone is causing this, yet. Clearly the two are related, but we need a reproducible test case before we can start proposing remedies, no? If there's a problem with the Quicksort implementation but we switch to heapsort after it hits the maximum recursion depth, then we mask a bug with the optimization.","created":"2008-06-02T22:34:29.665+0000"},{"body":"Okay, thanks for running that test.\n\nWhat you write about correctness makes sense to me. Thanks....","created":"2008-06-03T01:29:04.398+0000"},{"body":"This bug has hit a lot of our code. One thing I've found is that it depends on the amount of data per reducer. I.e. tripling the number of reducers causes the stack overflow to not occur.","created":"2008-06-04T19:38:29.805+0000"},{"body":"Again, repeating what Chris asked for earlier - can we have a testcase to reproduce the issue. That would be of huge help.","created":"2008-06-04T19:46:59.316+0000"},{"body":"bq. One thing I've found is that it depends on the amount of data per reducer. I.e. tripling the number of reducers causes the stack overflow to not occur.\n\nThat makes sense; the spill is sorted by partition, then by key. Increasing the number of reducers will avoid bad patterns in the data.\n\nI'm attaching a patch that will spill some internal data structures on a StackOverflowError and a verification tool. It's a server-side patch, so it must be deployed with the TaskTrackers. It should write to the output directory for the job. Once the spills are pulled to local disk, running the tool *should* reproduce the error. If someone who can reproduce this would compress and post the buffer/index data here, it would be hugely helpful.\n\nWhat's written to disk are the entries in the MapOutputBuffer (keys and values) for a single spill and a table recording partitions and key/value lengths.","created":"2008-06-04T19:50:48.996+0000"},{"body":"The jobs it's failing on are running on data I can't release. \n\nI am generating a synthetic data set and a stripped down mapper to try to reproduce the problem.","created":"2008-06-04T20:00:48.076+0000"},{"body":"bq. The jobs it's failing on are running on data I can't release.\n\nWould it be possible to apply the patch, validate the spills with the checker, and post some debugging information? In particular, patching QuickSort to catch StackOverflowError and print out its left and right indices as it unwinds would be the first thing I'd try with a reproducible test. If you can't include the kvbuffer data, the kvindices data would be modestly helpful and would reveal only lengths and partitions. Also, even if the checker doesn't reproduce the error, the \"STACK OVERFLOW\" log message from the failed mapper should include the buffer and accounting indices. Those would also be valuable and reveal nothing about your data.","created":"2008-06-04T20:35:47.792+0000"},{"body":"Attached is \n* a python script that generates a data file that causes the bug in the attached job class (generate_synth_data.py)\n* test driver/mapper (Test.java)\n* utility class (LongPair.java)\n\nThis has caused the StackOverflow on a single node test cluster (namenode/jobtracker/slaves on one box)","created":"2008-06-05T00:03:45.086+0000"},{"body":"\nThe patch looks good.\n\nI think it is better to add randomization to the patched code.\n\nTwo possibilities:\n1. instead to choose (p+r)/2, choose random(p, r) as the pivotal candidate.\n2. Randomize the input data before calling quicksort (O(n)) cost.\n","created":"2008-06-05T02:48:13.455+0000"},{"body":"Is this a new bug? Has the sort code been change recently or something that feeds it?\n\nDon't the classic texts have chapters on pivot selection?\n\nHere are some references (and support for runping's randomization suggestion)\n- http://en.wikipedia.org/wiki/Quicksort\n","created":"2008-06-05T06:25:53.357+0000"},{"body":"One possible bug causing this problem is integer overflow: for example, \n{code}\nfix(s, (p+r) >>> 1, p);\n{code}\nif p+r > Integer.MAX_VALUE, then the result won't be correct.\n\nThis leads to a question: do we need long for the index?","created":"2008-06-05T17:40:46.320+0000"},{"body":"I forgot >>> is unsigned right shift, so (p+r)>>>1 won't be overflow. However, if p+r < 0, it is a problem.","created":"2008-06-05T18:00:56.243+0000"},{"body":"bq. if p+r < 0, it is a problem.\n\nUntil the first compare, when the IndexedSortable (i.e. MapOutputBuffer) will throw an ArrayIndexOutOfBoundsException.\n\nbq. Is this a new bug? Has the sort code been change recently or something that feeds it?\n\nThe QuickSort is new in 0.17; we used to use MergeSort.\n\nbq. Don't the classic texts have chapters on pivot selection?\n\nThe default HashPartitioner- which almost everybody uses- should produce fairly random distributions, save equal keys (which is why increasing the number of reducers usually makes this go away). It was verified that the \"median-of-three\" partitioning would be sufficient to handle the sorted, single-reducer case, which is the canonical worst-case for quicksort. In the literature reviewed, cases where \"median-of-three\" partitioning fails are spoken of as being deliberately crafted, i.e. as possible DoS vectors, not as something common to real-world data. Someone was able to get a sharable spill and indices, so we'll do some analysis to figure out why it's more common than we thought it would be.\n\nbq. Randomize the input data before calling quicksort (O\\(n)) cost\n\nRandomizing the indices should be pretty cheap, so this is probably a good idea. Since we're confident that this really is a problem with the recursion, we should also investigate Chad's suggestion of Introsort. This is what the STL uses: http://www.sgi.com/tech/stl/sort.html (see [2])","created":"2008-06-05T19:15:53.938+0000"},{"body":"Analysis of the data (thanks to everyone who provided their test cases) led us to consider the following degenerate case:\n\nConsider a partition:\n{noformat}\na_n, a_1, a_2, ... , a_n-2, a_n-1\n{noformat}\n\nWhere {{a_1 ... a_n-1}} are sorted. The median of three partitioning will consider {{a_n}}, {{a_n/2}}, and {{a_n-1}} and select {{a_n-1}} as the pivot. While the sort runs:\n{noformat}\na_n-1, a_1, a_2, ... , a_n-2, a_n\n{noformat}\n\nThe left index will run all the way to {{a_n}} and swap the pivot into place, yielding the following:\n{noformat}\na_n-2, a_1, a_2, ... , a_n-3, a_n-1, a_n\n{noformat}\n\nSo the next partition will get:\n{noformat}\na_n-2, a_1, a_2, ... , a_n-4, a_n-3\n{noformat}\nSo while sorted data will yield a series of optimal partitions, nearly sorted data like this can cause the sort to fall into a degenerate case. Among the suggestions to ameliorate this:\n# Consider the median and two random offsets for the median-of-three partitioning (or three random offsets, etc.)\n# Always pick a random pivot\n# After swapping the pivot into place, swap what it replaced into a random position in the left partition\n\nRandomizing the input data makes this case far less common and Introsort regards it as an inevitable, degenerate case; both are also sound additions.","created":"2008-06-05T22:37:57.838+0000"},{"body":"1. QuickSort is n*logn only on average. In the worst case it is n-square. So whatever strategy you choose there will be an input sequence (randomized or not) that will lead to the degenerate behavior. So why QuickSort if there are hundreds of other sorting algorithms?\n2. Recursion is imperative and good for describing algorithms. In practice it should be eliminated. Standard techniques are applicable in this case.\n","created":"2008-06-05T23:06:06.540+0000"},{"body":"Oops! This should not have been unassigned from 0.18.0","created":"2008-06-07T01:31:23.352+0000"},{"body":"Proposed fix for 0.18:\n* Resurrect map.sort.class property to control which sort implementation to use, add HeapSort & BatcherSort as valid targets; keep QuickSort as the default\n* Change QuickSort to recur only on the smaller partition (limits recursion depth to log\\(n))\n* Change QuickSort to use HeapSort once it reaches an unreasonable recursion depth (2*ceil(log\\(n))) (Introspection Sort)\n\nFix for 0.17.1:\n* Only fix the recursion. The degenerate cases for QuickSort with median-of-three partitioning really are uncommon.","created":"2008-06-09T02:21:07.756+0000"},{"body":"Some quick questions - what's the complexity in time/space of BatcherSort. For HeapSort, there is a o.a.h.u.PriorityQueue. Could we use that?","created":"2008-06-09T03:53:22.246+0000"},{"body":"bq. what's the complexity in time/space of BatcherSort\n\nIt takes O(n*log\\(n)^2) time; without executing any of its stages in parallel, it's probably not worth using in practice before the alternatives, but it's as fast as QuickSort for small data sets.\n\nbq. For HeapSort, there is a o.a.h.u.PriorityQueue. Could we use that?\n\nI don't think it would be an improvement. HeapSort effects the sort in-place, while using PriorityQueue would require additional allocations and several unnecessary abstractions. HeapSort isn't a full heap implementation, either, so I'd argue that it's not adding redundant code.","created":"2008-06-09T04:18:17.456+0000"},{"body":"That's fair...","created":"2008-06-09T04:54:32.516+0000"},{"body":"- QuickSort.getMaxDepth can be more aggressive: since the index are int, we have n <= 2^32 and lg n <= 32. So, we could safely let QuickSort.getMaxDepth be something around (lg n)^3 or it should depend on the maximum size of system stack.\n\n- HeapSort is a singleton. Define a static variable HeapSort.INSTANCE. Then, we don't have to new HeapSort() in QuickSort.\n\n- public methods HeapSort.sort(IndexedSortable s, int p, int r) and QuickSort.sort(...) need javadoc.\n\n- There are quite a few sorting classes/interfaces. Consider create a new package for them.\n\n- Remove BatcherSort in this patch. Create a new issue if it is useful.\n","created":"2008-06-09T21:24:35.732+0000"},{"body":"bq. QuickSort.getMaxDepth can be more aggressive: since the index are int, we have n <= 2^32 and lg n <= 32. So, we could safely let QuickSort.getMaxDepth be something around (lg n)^3 or it should depend on the maximum size of system stack.\n\nThe 2*log\\(n) heuristic was intended to bail out of a worst case, not to protect against the StackOverflowError. The recursion on the system stack is already limited to log\\(n), but quicksort will always have O(n^2) cases that this needs to handle. I'm happy to discuss an appropriate value for k, but k*log\\(n) seems appropriate to this purpose.\n\nbq. HeapSort is a singleton. Define a static variable HeapSort.INSTANCE. Then, we don't have to new HeapSort() in QuickSort.\nbq. There are quite a few sorting classes/interfaces. Consider create a new package for them.\n\nOf course, you're right; all the algorithms should be singletons, and further, having sort algorithms as objects at all is not the best design. Still, given that MapTask remains the only user, adding factories, moving these into a separate package, etc. seems like overkill at this point. Would it be sufficient to add a singleton instance of HeapSort to QuickSort?\n\nbq. public methods HeapSort.sort(IndexedSortable s, int p, int r) and QuickSort.sort(...) need javadoc.\n\nOK.\n\nbq. Remove BatcherSort in this patch. Create a new issue if it is useful.\n\nFair enough. BatcherSort is quick for small sets of data (particularly when compared to QuickSort, which uses an insertion sort for small partitions), but small datasets aren't exactly a typical use case. :)","created":"2008-06-10T07:36:46.506+0000"},{"body":"bq. The 2*log(n) heuristic was intended to bail out of a worst case, not to protect against the StackOverflowError. The recursion on the system stack is already limited to log(n), but quicksort will always have O(n^2) cases that this needs to handle. I'm happy to discuss an appropriate value for k, but k*log(n) seems appropriate to this purpose.\n\nI guess it is safe to choose a larger k. For example, with k = 8, the stack depth is bounded by 250 in practice, but a lot fewer cases of input data patterns will need to invoke the heapsort.\nEven with all these fixes, I still believe it is good to add randomization to pivotal selection (or randomize the input data upfront).","created":"2008-06-10T10:29:08.000+0000"},{"body":"> Even with all these fixes, I still believe it is good to add randomization to pivotal selection (or randomize the input data upfront).\n\nI also like randomized quick sort. However, we better fix the stack overflow problem first. Then, we could think about how to improve the sorting algorithm.\n\n> The 2*log\\(n) heuristic was intended to bail out of a worst case, not to protect against the StackOverflowError\n\nI think the 2*log n heuristic not only bail out of worst cases but also good cases since it is overly strict.\n\nI agree that we should force on fixing the problem in this issue. 3442-3.patch is already very good. +1\n","created":"2008-06-10T17:52:04.229+0000"},{"body":"bq. I guess it is safe to choose a larger k. For example, with k = 8, the stack depth is bounded by 250 in practice, but a lot fewer cases of input data patterns will need to invoke the heapsort.\nbq. I think the 2*log n heuristic not only bail out of worst cases but also good cases since it is overly strict.\n\nThe IntroSort [paper|http://citeseer.ist.psu.edu/musser97introspective.html] suggested that 2 * floor(log\\(n)) produced good results. Given that the most degenerate cases cascade, deferring the change to heapsort may effect longer running times. Still, it does seem low. I ran some informal benchmarks with a random set of elements in nearly sorted, sawtooth, and pipe organ patterns and using k = 2 switched to heapsort pretty aggressively, so I increased k to 4. The average running times were all within a few hundred milliseconds, though; we probably don't want to get too carried away optimizing and tweaking an operation typically accounting for no more than a few seconds out of every spill.","created":"2008-06-11T00:25:50.434+0000"},{"body":"I just committed this.","created":"2008-06-11T00:57:36.988+0000"}],"conversations":[{"body":"","from":"reporter","subject":"QuickSort may get into unbounded recursion"},{"body":"\nWhen run gridmix with the latest 0.17 code, some job failed with the mappers throwing the following exception:\n\njava.lang.StackOverflowError\n\tat org.apache.hadoop.util.QuickSort.fix(QuickSort.java:29)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:58)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n\tat org.apache.hadoop.util.QuickSort.sort(QuickSort.java:82)\n","from":"developer"},{"body":"II I understand the code correctly, it seems that if the input data is already sorted, the quick sort will behave badly -- recurrsion depth will be equal to the data size.\n","from":"developer"},{"body":"I haven't verified this yet, but if what Runping is saying is true, does it warrant a 0.17.1 release?","from":"developer"},{"body":"The recursion will always be on the smaller of the two partitions first, so even a degenerate case should not overflow the stack. I ran the testcase with 10 * 2^17 elements (default 5% of the default 100MB allocated to record storage in MapTask), 10x that, sorted, reversed, all equal, etc. and could not reproduce this. However, if the comparator were flawed, that could cause the behavior you're seeing.","from":"developer"},{"body":"It's not sorted data (which should produce relatively even partitions per lines 58-60), but horribly unlucky pivot selections that could produce this. Since the JVM doesn't appear to recognize the tail recursion, it's possible that a series- meaning hundreds- of sequential median values sampled from the left, right, and midpoint could be so skewed that the recursion could top out (a bad comparator might hurt pivot selection). This should be _fantastically_ unlikely, though. Does this happen with HADOOP-3308 applied?","from":"developer"},{"body":"Runping, are you able to reproduce the problem?\n\nFor already sorted list, I think the codes will pick the median (i.e. the best choice). I suspect that the bug, if there is any, is outside util.QuickSort.","from":"developer"},{"body":"I've attached a patch to remove the recursion. On my machine, this actually made things faster too. Not sure if that's real or just a quirk with my test case.","from":"developer"},{"body":"I agree with the comments that the problem was probably outside of the quicksort. You could also check your java stack size to be sure it's not somehow set _really_ low.","from":"developer"},{"body":"This may be useful to add: http://en.wikipedia.org/wiki/Introsort (essentially, switch from quicksort to heapsort once for subsorts beyond a certain depth).","from":"developer"},{"body":"Alternatively, we could just effect the tail recursion with a loop (see attached; credit to Nicholas for the idea). I'm not wild about creating new java.util.Stack instances, but it's not likely to be a performance issue either way. As for switching to heapsort from insertion sort for < K or log\\(n) values, though I'd be surprised to learn that was a bottleneck, it sounds like a cool idea to try.\n\nThe recursion itself doesn't seem to be the problem, though; that there's a case where it becomes unbounded is. This is showing up often enough that there is likely a framework problem, but nobody has explained how they reproduce it, yet.","from":"developer"},{"body":"Can you try testing my version in your environment to see if you get a performance change? I'm thinking that there may be some JIT thing going on that makes the no-recursion version better. I did get a significant (like, 20%+ less time), repeatable, performance gain when using my version.\n\nIntrosort still uses insertionSort instead of quicksort for low n. It uses heapsort when quicksort has recursed deeply; unlike quicksort, heapsort has guaranteed worst-case performance of O(n log n), avoiding both the huge recursion (it's not recursive) and the O(n*n) runtime that such a recursion produces. But it's generally slower, so introsort uses it only when quicksort has hit a bad case.\n\nThanks!","from":"developer"},{"body":"bq. Can you try testing my version in your environment to see if you get a performance change?\n\nI tried running the test case with 10M elements and saw only very slight performance improvements over the original, excluding the sequential test case, where it degraded. I'm running 1.5.0_13 on a Mac, so YMMV.\n\nbq. Introsort still uses insertionSort instead of quicksort for low n. It uses heapsort when quicksort has recursed deeply\n\nAh, I understand. Still, the issue here is not performance, but correctness. Until we verify that the implementation of QuickSort is correct but pushed into a degenerate case, I think we should assume that some aspect of the framework is incorrect and try to determine why we're topping out the stack. Once we isolate the issue, then we can look into new sorting strategies if they solve the problem, but I'm not sure we can say without fear of refutation that the recursion alone is causing this, yet. Clearly the two are related, but we need a reproducible test case before we can start proposing remedies, no? If there's a problem with the Quicksort implementation but we switch to heapsort after it hits the maximum recursion depth, then we mask a bug with the optimization.","from":"developer"},{"body":"Okay, thanks for running that test.\n\nWhat you write about correctness makes sense to me. Thanks....","from":"developer"},{"body":"This bug has hit a lot of our code. One thing I've found is that it depends on the amount of data per reducer. I.e. tripling the number of reducers causes the stack overflow to not occur.","from":"developer"},{"body":"Again, repeating what Chris asked for earlier - can we have a testcase to reproduce the issue. That would be of huge help.","from":"developer"},{"body":"bq. One thing I've found is that it depends on the amount of data per reducer. I.e. tripling the number of reducers causes the stack overflow to not occur.\n\nThat makes sense; the spill is sorted by partition, then by key. Increasing the number of reducers will avoid bad patterns in the data.\n\nI'm attaching a patch that will spill some internal data structures on a StackOverflowError and a verification tool. It's a server-side patch, so it must be deployed with the TaskTrackers. It should write to the output directory for the job. Once the spills are pulled to local disk, running the tool *should* reproduce the error. If someone who can reproduce this would compress and post the buffer/index data here, it would be hugely helpful.\n\nWhat's written to disk are the entries in the MapOutputBuffer (keys and values) for a single spill and a table recording partitions and key/value lengths.","from":"developer"},{"body":"The jobs it's failing on are running on data I can't release. \n\nI am generating a synthetic data set and a stripped down mapper to try to reproduce the problem.","from":"developer"},{"body":"bq. The jobs it's failing on are running on data I can't release.\n\nWould it be possible to apply the patch, validate the spills with the checker, and post some debugging information? In particular, patching QuickSort to catch StackOverflowError and print out its left and right indices as it unwinds would be the first thing I'd try with a reproducible test. If you can't include the kvbuffer data, the kvindices data would be modestly helpful and would reveal only lengths and partitions. Also, even if the checker doesn't reproduce the error, the \"STACK OVERFLOW\" log message from the failed mapper should include the buffer and accounting indices. Those would also be valuable and reveal nothing about your data.","from":"developer"},{"body":"Attached is \n* a python script that generates a data file that causes the bug in the attached job class (generate_synth_data.py)\n* test driver/mapper (Test.java)\n* utility class (LongPair.java)\n\nThis has caused the StackOverflow on a single node test cluster (namenode/jobtracker/slaves on one box)","from":"developer"},{"body":"\nThe patch looks good.\n\nI think it is better to add randomization to the patched code.\n\nTwo possibilities:\n1. instead to choose (p+r)/2, choose random(p, r) as the pivotal candidate.\n2. Randomize the input data before calling quicksort (O(n)) cost.\n","from":"developer"},{"body":"Is this a new bug? Has the sort code been change recently or something that feeds it?\n\nDon't the classic texts have chapters on pivot selection?\n\nHere are some references (and support for runping's randomization suggestion)\n- http://en.wikipedia.org/wiki/Quicksort\n","from":"developer"},{"body":"One possible bug causing this problem is integer overflow: for example, \n{code}\nfix(s, (p+r) >>> 1, p);\n{code}\nif p+r > Integer.MAX_VALUE, then the result won't be correct.\n\nThis leads to a question: do we need long for the index?","from":"developer"},{"body":"I forgot >>> is unsigned right shift, so (p+r)>>>1 won't be overflow. However, if p+r < 0, it is a problem.","from":"developer"},{"body":"bq. if p+r < 0, it is a problem.\n\nUntil the first compare, when the IndexedSortable (i.e. MapOutputBuffer) will throw an ArrayIndexOutOfBoundsException.\n\nbq. Is this a new bug? Has the sort code been change recently or something that feeds it?\n\nThe QuickSort is new in 0.17; we used to use MergeSort.\n\nbq. Don't the classic texts have chapters on pivot selection?\n\nThe default HashPartitioner- which almost everybody uses- should produce fairly random distributions, save equal keys (which is why increasing the number of reducers usually makes this go away). It was verified that the \"median-of-three\" partitioning would be sufficient to handle the sorted, single-reducer case, which is the canonical worst-case for quicksort. In the literature reviewed, cases where \"median-of-three\" partitioning fails are spoken of as being deliberately crafted, i.e. as possible DoS vectors, not as something common to real-world data. Someone was able to get a sharable spill and indices, so we'll do some analysis to figure out why it's more common than we thought it would be.\n\nbq. Randomize the input data before calling quicksort (O\\(n)) cost\n\nRandomizing the indices should be pretty cheap, so this is probably a good idea. Since we're confident that this really is a problem with the recursion, we should also investigate Chad's suggestion of Introsort. This is what the STL uses: http://www.sgi.com/tech/stl/sort.html (see [2])","from":"developer"},{"body":"Analysis of the data (thanks to everyone who provided their test cases) led us to consider the following degenerate case:\n\nConsider a partition:\n{noformat}\na_n, a_1, a_2, ... , a_n-2, a_n-1\n{noformat}\n\nWhere {{a_1 ... a_n-1}} are sorted. The median of three partitioning will consider {{a_n}}, {{a_n/2}}, and {{a_n-1}} and select {{a_n-1}} as the pivot. While the sort runs:\n{noformat}\na_n-1, a_1, a_2, ... , a_n-2, a_n\n{noformat}\n\nThe left index will run all the way to {{a_n}} and swap the pivot into place, yielding the following:\n{noformat}\na_n-2, a_1, a_2, ... , a_n-3, a_n-1, a_n\n{noformat}\n\nSo the next partition will get:\n{noformat}\na_n-2, a_1, a_2, ... , a_n-4, a_n-3\n{noformat}\nSo while sorted data will yield a series of optimal partitions, nearly sorted data like this can cause the sort to fall into a degenerate case. Among the suggestions to ameliorate this:\n# Consider the median and two random offsets for the median-of-three partitioning (or three random offsets, etc.)\n# Always pick a random pivot\n# After swapping the pivot into place, swap what it replaced into a random position in the left partition\n\nRandomizing the input data makes this case far less common and Introsort regards it as an inevitable, degenerate case; both are also sound additions.","from":"developer"},{"body":"1. QuickSort is n*logn only on average. In the worst case it is n-square. So whatever strategy you choose there will be an input sequence (randomized or not) that will lead to the degenerate behavior. So why QuickSort if there are hundreds of other sorting algorithms?\n2. Recursion is imperative and good for describing algorithms. In practice it should be eliminated. Standard techniques are applicable in this case.\n","from":"developer"},{"body":"Oops! This should not have been unassigned from 0.18.0","from":"developer"},{"body":"Proposed fix for 0.18:\n* Resurrect map.sort.class property to control which sort implementation to use, add HeapSort & BatcherSort as valid targets; keep QuickSort as the default\n* Change QuickSort to recur only on the smaller partition (limits recursion depth to log\\(n))\n* Change QuickSort to use HeapSort once it reaches an unreasonable recursion depth (2*ceil(log\\(n))) (Introspection Sort)\n\nFix for 0.17.1:\n* Only fix the recursion. The degenerate cases for QuickSort with median-of-three partitioning really are uncommon.","from":"developer"},{"body":"Some quick questions - what's the complexity in time/space of BatcherSort. For HeapSort, there is a o.a.h.u.PriorityQueue. Could we use that?","from":"developer"},{"body":"bq. what's the complexity in time/space of BatcherSort\n\nIt takes O(n*log\\(n)^2) time; without executing any of its stages in parallel, it's probably not worth using in practice before the alternatives, but it's as fast as QuickSort for small data sets.\n\nbq. For HeapSort, there is a o.a.h.u.PriorityQueue. Could we use that?\n\nI don't think it would be an improvement. HeapSort effects the sort in-place, while using PriorityQueue would require additional allocations and several unnecessary abstractions. HeapSort isn't a full heap implementation, either, so I'd argue that it's not adding redundant code.","from":"developer"},{"body":"That's fair...","from":"developer"},{"body":"- QuickSort.getMaxDepth can be more aggressive: since the index are int, we have n <= 2^32 and lg n <= 32. So, we could safely let QuickSort.getMaxDepth be something around (lg n)^3 or it should depend on the maximum size of system stack.\n\n- HeapSort is a singleton. Define a static variable HeapSort.INSTANCE. Then, we don't have to new HeapSort() in QuickSort.\n\n- public methods HeapSort.sort(IndexedSortable s, int p, int r) and QuickSort.sort(...) need javadoc.\n\n- There are quite a few sorting classes/interfaces. Consider create a new package for them.\n\n- Remove BatcherSort in this patch. Create a new issue if it is useful.\n","from":"developer"},{"body":"bq. QuickSort.getMaxDepth can be more aggressive: since the index are int, we have n <= 2^32 and lg n <= 32. So, we could safely let QuickSort.getMaxDepth be something around (lg n)^3 or it should depend on the maximum size of system stack.\n\nThe 2*log\\(n) heuristic was intended to bail out of a worst case, not to protect against the StackOverflowError. The recursion on the system stack is already limited to log\\(n), but quicksort will always have O(n^2) cases that this needs to handle. I'm happy to discuss an appropriate value for k, but k*log\\(n) seems appropriate to this purpose.\n\nbq. HeapSort is a singleton. Define a static variable HeapSort.INSTANCE. Then, we don't have to new HeapSort() in QuickSort.\nbq. There are quite a few sorting classes/interfaces. Consider create a new package for them.\n\nOf course, you're right; all the algorithms should be singletons, and further, having sort algorithms as objects at all is not the best design. Still, given that MapTask remains the only user, adding factories, moving these into a separate package, etc. seems like overkill at this point. Would it be sufficient to add a singleton instance of HeapSort to QuickSort?\n\nbq. public methods HeapSort.sort(IndexedSortable s, int p, int r) and QuickSort.sort(...) need javadoc.\n\nOK.\n\nbq. Remove BatcherSort in this patch. Create a new issue if it is useful.\n\nFair enough. BatcherSort is quick for small sets of data (particularly when compared to QuickSort, which uses an insertion sort for small partitions), but small datasets aren't exactly a typical use case. :)","from":"developer"},{"body":"bq. The 2*log(n) heuristic was intended to bail out of a worst case, not to protect against the StackOverflowError. The recursion on the system stack is already limited to log(n), but quicksort will always have O(n^2) cases that this needs to handle. I'm happy to discuss an appropriate value for k, but k*log(n) seems appropriate to this purpose.\n\nI guess it is safe to choose a larger k. For example, with k = 8, the stack depth is bounded by 250 in practice, but a lot fewer cases of input data patterns will need to invoke the heapsort.\nEven with all these fixes, I still believe it is good to add randomization to pivotal selection (or randomize the input data upfront).","from":"developer"},{"body":"> Even with all these fixes, I still believe it is good to add randomization to pivotal selection (or randomize the input data upfront).\n\nI also like randomized quick sort. However, we better fix the stack overflow problem first. Then, we could think about how to improve the sorting algorithm.\n\n> The 2*log\\(n) heuristic was intended to bail out of a worst case, not to protect against the StackOverflowError\n\nI think the 2*log n heuristic not only bail out of worst cases but also good cases since it is overly strict.\n\nI agree that we should force on fixing the problem in this issue. 3442-3.patch is already very good. +1\n","from":"developer"},{"body":"bq. I guess it is safe to choose a larger k. For example, with k = 8, the stack depth is bounded by 250 in practice, but a lot fewer cases of input data patterns will need to invoke the heapsort.\nbq. I think the 2*log n heuristic not only bail out of worst cases but also good cases since it is overly strict.\n\nThe IntroSort [paper|http://citeseer.ist.psu.edu/musser97introspective.html] suggested that 2 * floor(log\\(n)) produced good results. Given that the most degenerate cases cascade, deferring the change to heapsort may effect longer running times. Still, it does seem low. I ran some informal benchmarks with a random set of elements in nearly sorted, sawtooth, and pipe organ patterns and using k = 2 switched to heapsort pretty aggressively, so I increased k to 4. The average running times were all within a few hundred milliseconds, though; we probably don't want to get too carried away optimizing and tweaking an operation typically accounting for no more than a few seconds out of every spill.","from":"developer"},{"body":"I just committed this.","from":"developer"}],"created":"2008-05-23T21:50:08.000+0000","description":"","issue_id":"12396742","key":"HADOOP-3442","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-06-11T00:57:37.000+0000","role":"fixed_distractor","summary":"QuickSort may get into unbounded recursion"} {"case_id":"12347270","cluster":"DISTRACTOR-HADOOP-423","comments":[{"body":"To fix this, we'll detect and normalize relative paths in the Path class. For backwards compatibility, we'll allow people to rename or remove existing files/dirs that contain '.' or '..' . Or do people think that's not necessary because it's unlikely anyone has been using '.' and '..' in their paths intentionally? (Though I suppose if they did unintentionally, it might be nice to let them remove them.)\n ","created":"2006-09-07T00:54:47.000+0000"},{"body":"The patch also includes changes in Path.java as well as added tests in TestPath.java.","created":"2006-09-20T22:16:00.000+0000"},{"body":"I think, as a safeguard, we should probably also disallow \".\" and \"..\" as file or directory names at the namenode, throwing an exception for attempts to create directories or files with these names.\n\nWhile we're at it, should we also prohibit \"/\" and \":\" in names? I think so. Currently it is impossible to create a Path with a slash in a name, although colons are permitted.","created":"2006-09-21T22:17:56.000+0000"},{"body":"I can add a check to disallow \".\" and \"..\" (and prohibit \":\") as file or directory names. Should this be done in DFSClient or the namenode? We can save an rpc by putting it in the client.","created":"2006-09-25T23:26:03.000+0000"},{"body":"I think we want it as a sanity check in the namenode, so that we're sure they never enter the filesystem. I don't think it's worth trying to optimize this.","created":"2006-09-25T23:46:23.000+0000"},{"body":"This patch also contains the check in the namenode. ","created":"2006-09-27T22:58:31.000+0000"},{"body":"I think the check in the NameNode isn't quite right. We should never permit \".\" or \"..\" at any point in a path here, since paths should always be fully qualifited and normalized by the time they reach the namenode, right?\n\nAlso, it seems a shame to have to call Path.toString(). Perhaps we should add a Path method like:\n\n public String element(int i) { return elements[i]; }\n\nThen the path checker could be implemented with something like:\n\nif (!path.isAbsolute()) {\n throw ... \n}\nfor (int i = 0; i < path.depth(); i++) {\n String element = path.element(i);\n if (element.equals(\"..\") || element.equals(\".\") || element.indexOf(\"/\") >= 0) {\n throw ...\n }\n}\n","created":"2006-09-28T17:47:41.000+0000"},{"body":"You're right, I've changed the logic in the namenode check.\n\nI don't understand where you don't want to call Path.toString(), in DFSClient? The namenode only gets a String, which I parse for the offending sequences. Do you mean we should convert it to a Path there?","created":"2006-09-28T23:26:50.000+0000"},{"body":"> I don't understand where you don't want to call Path.toString(), in DFSClient?\n\nYou're right. I mistakenly assumed that src.toString() was used to convert a Path to a String, when it's really used to convert a UTF8 to a String. Sorry!","created":"2006-09-28T23:32:31.000+0000"},{"body":"I just committed this. Thanks, Wendy!","created":"2006-09-29T20:15:27.000+0000"}],"conversations":[{"body":"Paths containing '.' or '..' are not normalized -- the corresponding directories actually seem to be created. This is inconsistent and unexpected behavior, and is particularly painful when running from a working directory, using relative paths that start with ...","from":"reporter","subject":"file paths are not normalized"},{"body":"To fix this, we'll detect and normalize relative paths in the Path class. For backwards compatibility, we'll allow people to rename or remove existing files/dirs that contain '.' or '..' . Or do people think that's not necessary because it's unlikely anyone has been using '.' and '..' in their paths intentionally? (Though I suppose if they did unintentionally, it might be nice to let them remove them.)\n ","from":"developer"},{"body":"The patch also includes changes in Path.java as well as added tests in TestPath.java.","from":"developer"},{"body":"I think, as a safeguard, we should probably also disallow \".\" and \"..\" as file or directory names at the namenode, throwing an exception for attempts to create directories or files with these names.\n\nWhile we're at it, should we also prohibit \"/\" and \":\" in names? I think so. Currently it is impossible to create a Path with a slash in a name, although colons are permitted.","from":"developer"},{"body":"I can add a check to disallow \".\" and \"..\" (and prohibit \":\") as file or directory names. Should this be done in DFSClient or the namenode? We can save an rpc by putting it in the client.","from":"developer"},{"body":"I think we want it as a sanity check in the namenode, so that we're sure they never enter the filesystem. I don't think it's worth trying to optimize this.","from":"developer"},{"body":"This patch also contains the check in the namenode. ","from":"developer"},{"body":"I think the check in the NameNode isn't quite right. We should never permit \".\" or \"..\" at any point in a path here, since paths should always be fully qualifited and normalized by the time they reach the namenode, right?\n\nAlso, it seems a shame to have to call Path.toString(). Perhaps we should add a Path method like:\n\n public String element(int i) { return elements[i]; }\n\nThen the path checker could be implemented with something like:\n\nif (!path.isAbsolute()) {\n throw ... \n}\nfor (int i = 0; i < path.depth(); i++) {\n String element = path.element(i);\n if (element.equals(\"..\") || element.equals(\".\") || element.indexOf(\"/\") >= 0) {\n throw ...\n }\n}\n","from":"developer"},{"body":"You're right, I've changed the logic in the namenode check.\n\nI don't understand where you don't want to call Path.toString(), in DFSClient? The namenode only gets a String, which I parse for the offending sequences. Do you mean we should convert it to a Path there?","from":"developer"},{"body":"> I don't understand where you don't want to call Path.toString(), in DFSClient?\n\nYou're right. I mistakenly assumed that src.toString() was used to convert a Path to a String, when it's really used to convert a UTF8 to a String. Sorry!","from":"developer"},{"body":"I just committed this. Thanks, Wendy!","from":"developer"}],"created":"2006-08-03T20:25:53.000+0000","description":"Paths containing '.' or '..' are not normalized -- the corresponding directories actually seem to be created. This is inconsistent and unexpected behavior, and is particularly painful when running from a working directory, using relative paths that start with ...","issue_id":"12347270","key":"HADOOP-423","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-09-29T20:15:27.000+0000","role":"fixed_distractor","summary":"file paths are not normalized"} {"case_id":"12347638","cluster":"DISTRACTOR-HADOOP-438","comments":[{"body":"To fix this we're going to enforce a pathname limit in the Path class constructor. . (8K in length and 1K in depth.) The constructor will throw an exception which will be passed back to the client. ","created":"2006-09-07T01:25:07.000+0000"},{"body":"We've changed our plan. Instead of enforcing the length in the Path constructor, we are going to enforce it in mkdirs, addFile, and renameTo. The reason for this is that the limitation is not actually in Path, but in the filesystem, so we should not enforce it in Path. An exception will still be passed back to the client. \n\nIf you object to this plan or prefer the original, please comment. \n\n\n","created":"2006-09-11T21:59:11.000+0000"},{"body":"I just wanted to mention that there are file systems out there without path restrictions.\n\nhttp://en.wikipedia.org/wiki/Comparison_of_file_systems#Limits\n","created":"2006-09-11T23:18:52.000+0000"},{"body":"I just committed this. Thanks, Wendy!","created":"2006-09-14T20:13:10.000+0000"}],"conversations":[{"body":"I was trying to create a deep hierarchy of directories using DFS mkdirs().\nWhen the path to the leaf directory became long (~20000) DFS was still able to create\ndirectories with these names, but UTF8 started truncating long strings resulting in\nincorrect logging of namespace edits. That later crashed the namenode during restart,\nwhen it was trying to reproduce file creation logged in the edits file with truncated names.\nUTF8 is deprecated now so we will have to replace it with Text.\nWith UTF8 we should enforce a pathname limit of 0xffff/3 = 21845\nWith Text it is going to be larger. Not sure what the exact number is.\n","from":"reporter","subject":"DFS pathname limitation."},{"body":"To fix this we're going to enforce a pathname limit in the Path class constructor. . (8K in length and 1K in depth.) The constructor will throw an exception which will be passed back to the client. ","from":"developer"},{"body":"We've changed our plan. Instead of enforcing the length in the Path constructor, we are going to enforce it in mkdirs, addFile, and renameTo. The reason for this is that the limitation is not actually in Path, but in the filesystem, so we should not enforce it in Path. An exception will still be passed back to the client. \n\nIf you object to this plan or prefer the original, please comment. \n\n\n","from":"developer"},{"body":"I just wanted to mention that there are file systems out there without path restrictions.\n\nhttp://en.wikipedia.org/wiki/Comparison_of_file_systems#Limits\n","from":"developer"},{"body":"I just committed this. Thanks, Wendy!","from":"developer"}],"created":"2006-08-10T01:08:43.000+0000","description":"I was trying to create a deep hierarchy of directories using DFS mkdirs().\nWhen the path to the leaf directory became long (~20000) DFS was still able to create\ndirectories with these names, but UTF8 started truncating long strings resulting in\nincorrect logging of namespace edits. That later crashed the namenode during restart,\nwhen it was trying to reproduce file creation logged in the edits file with truncated names.\nUTF8 is deprecated now so we will have to replace it with Text.\nWith UTF8 we should enforce a pathname limit of 0xffff/3 = 21845\nWith Text it is going to be larger. Not sure what the exact number is.\n","issue_id":"12347638","key":"HADOOP-438","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-09-14T20:13:10.000+0000","role":"fixed_distractor","summary":"DFS pathname limitation."} {"case_id":"12406534","cluster":"DISTRACTOR-HADOOP-4422","comments":[{"body":"Simple patch that removes bucket creation.","created":"2008-10-15T21:44:26.266+0000"},{"body":"This patch needs to be re-generated without the a/ b/ stuff for Hudson to be able to apply it. It must apply with 'patch -p 0 < foo.patch' when connected to trunk.","created":"2008-10-16T17:29:13.328+0000"},{"body":"Patch applies with -p0 now (used git diff --no-prefix).","created":"2008-10-16T19:57:51.040+0000"},{"body":"Marked as an incompatible change, since existing code that relies on bucket creation will need to be changed.","created":"2008-10-16T20:38:32.441+0000"},{"body":"For consistency, we should make the same change to Jets3tFileSystemStore.\n\nAlso, regarding the tests, there are two unit tests: Jets3tS3FileSystemContractTest and Jets3tNativeS3FileSystemContractTest which can be run manually to test the S3 integration. The only difference with this patch is that the buckets they run against must already exist - so I don't think any change to the tests are needed.","created":"2008-10-21T20:10:28.371+0000"},{"body":"Cancelling patch pending change to Jets3tFileSystemStore.","created":"2008-11-12T00:57:23.445+0000"},{"body":"Bucket creation also removed from Jets3tFileSystemStore.","created":"2008-11-21T22:25:43.278+0000"},{"body":"I ran the tests as follows after setting the correct test buckets and keys in src/test/hadoop-site.xml:\n\nant -Dtestcase=Jets3tS3FileSystemContractTest test\nant -Dtestcase=Jets3tNativeS3FileSystemContractTest test\n\nThey seem to pass:\n\nTestsuite: org.apache.hadoop.fs.s3.Jets3tS3FileSystemContractTest\nTests run: 25, Failures: 0, Errors: 0, Time elapsed: 131.575 sec\nTestsuite: org.apache.hadoop.fs.s3native.Jets3tNativeS3FileSystemContractTest\nTests run: 26, Failures: 0, Errors: 0, Time elapsed: 52.694 sec\n\nHowever, they both produce hundreds of warnings:\n\n(s3) 2008-11-21 15:51:28,800 WARN httpclient.RestS3Service (RestS3Service.java:performRequest(317)) - Response '/%2Ftest' - Unexpected response code 404, expected 200\n(s3n) 2008-11-21 15:34:55,646 WARN httpclient.RestS3Service (RestS3Service.java:performRequest(317)) - Response '/test' - Unexpected response code 404, expected 200\n\nAny ideas?","created":"2008-11-21T22:31:55.876+0000"},{"body":"Never mind about the warnings from Jets3t during testing. They are expected.","created":"2008-11-22T00:25:01.564+0000"},{"body":"I've just committed this. Thanks David!","created":"2008-11-25T12:03:23.229+0000"},{"body":"Edit release note for publication.","created":"2009-03-03T01:52:10.085+0000"}],"conversations":[{"body":"Both S3 file systems (s3 and s3n) try to create the bucket at every initialization. This is bad because\n\n* Every S3 operation costs money. These unnecessary calls are an unnecessary expense.\n* These calls can fail when called concurrently. This makes the file system unusable in large jobs.\n* Any operation, such as a \"fs -ls\", creates a bucket. This is counter-intuitive and undesirable.\n\nThe initialization code should assume the bucket exists:\n\n* Creating a bucket is a very rare operation. Accounts are limited to 100 buckets.\n* Any check at initialization for bucket existence is a waste of money.\n\nPer Amazon: \"Because bucket operations work against a centralized, global resource space, it is not appropriate to make bucket create or delete calls on the high availability code path of your application. It is better to create or delete buckets in a separate initialization or setup routine that you run less often.\"\n","from":"reporter","subject":"S3 file systems should not create bucket"},{"body":"Simple patch that removes bucket creation.","from":"developer"},{"body":"This patch needs to be re-generated without the a/ b/ stuff for Hudson to be able to apply it. It must apply with 'patch -p 0 < foo.patch' when connected to trunk.","from":"developer"},{"body":"Patch applies with -p0 now (used git diff --no-prefix).","from":"developer"},{"body":"Marked as an incompatible change, since existing code that relies on bucket creation will need to be changed.","from":"developer"},{"body":"For consistency, we should make the same change to Jets3tFileSystemStore.\n\nAlso, regarding the tests, there are two unit tests: Jets3tS3FileSystemContractTest and Jets3tNativeS3FileSystemContractTest which can be run manually to test the S3 integration. The only difference with this patch is that the buckets they run against must already exist - so I don't think any change to the tests are needed.","from":"developer"},{"body":"Cancelling patch pending change to Jets3tFileSystemStore.","from":"developer"},{"body":"Bucket creation also removed from Jets3tFileSystemStore.","from":"developer"},{"body":"I ran the tests as follows after setting the correct test buckets and keys in src/test/hadoop-site.xml:\n\nant -Dtestcase=Jets3tS3FileSystemContractTest test\nant -Dtestcase=Jets3tNativeS3FileSystemContractTest test\n\nThey seem to pass:\n\nTestsuite: org.apache.hadoop.fs.s3.Jets3tS3FileSystemContractTest\nTests run: 25, Failures: 0, Errors: 0, Time elapsed: 131.575 sec\nTestsuite: org.apache.hadoop.fs.s3native.Jets3tNativeS3FileSystemContractTest\nTests run: 26, Failures: 0, Errors: 0, Time elapsed: 52.694 sec\n\nHowever, they both produce hundreds of warnings:\n\n(s3) 2008-11-21 15:51:28,800 WARN httpclient.RestS3Service (RestS3Service.java:performRequest(317)) - Response '/%2Ftest' - Unexpected response code 404, expected 200\n(s3n) 2008-11-21 15:34:55,646 WARN httpclient.RestS3Service (RestS3Service.java:performRequest(317)) - Response '/test' - Unexpected response code 404, expected 200\n\nAny ideas?","from":"developer"},{"body":"Never mind about the warnings from Jets3t during testing. They are expected.","from":"developer"},{"body":"I've just committed this. Thanks David!","from":"developer"},{"body":"Edit release note for publication.","from":"developer"}],"created":"2008-10-15T21:43:51.000+0000","description":"Both S3 file systems (s3 and s3n) try to create the bucket at every initialization. This is bad because\n\n* Every S3 operation costs money. These unnecessary calls are an unnecessary expense.\n* These calls can fail when called concurrently. This makes the file system unusable in large jobs.\n* Any operation, such as a \"fs -ls\", creates a bucket. This is counter-intuitive and undesirable.\n\nThe initialization code should assume the bucket exists:\n\n* Creating a bucket is a very rare operation. Accounts are limited to 100 buckets.\n* Any check at initialization for bucket existence is a waste of money.\n\nPer Amazon: \"Because bucket operations work against a centralized, global resource space, it is not appropriate to make bucket create or delete calls on the high availability code path of your application. It is better to create or delete buckets in a separate initialization or setup routine that you run less often.\"\n","issue_id":"12406534","key":"HADOOP-4422","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-11-25T12:03:23.000+0000","role":"fixed_distractor","summary":"S3 file systems should not create bucket"} {"case_id":"12407043","cluster":"DISTRACTOR-HADOOP-4498","comments":[{"body":"changed component affected","created":"2008-10-22T23:29:49.301+0000"},{"body":"changed title","created":"2008-10-22T23:30:56.833+0000"},{"body":"Stuffs jobName between \\Q and \\E so it is treated as a literal value.","created":"2008-10-22T23:35:16.787+0000"},{"body":"All but org.apache.hadoop.fs.TestLocalDirAllocator tests pass. \n\nDid not include a unit test since most of these methods are private to JobHistory..\n\n [exec] -1 overall. \n [exec] \n [exec] +1 @author. The patch does not contain any @author tags.\n [exec] \n [exec] -1 tests included. The patch doesn't appear to include any new or modified tests.\n [exec] Please justify why no tests are needed for this patch.\n [exec] \n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec] \n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec] \n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.","created":"2008-10-23T16:39:59.480+0000"},{"body":"If I'm not missing something, the following String replaceAll is not useful.\n{noformat} \n+ private static String escapeRegexChars( String string ) {\n+ return \"\\\\Q\"+string.replaceAll(\"\\\\\\\\E\", \"\\\\\\\\E\")+\"\\\\E\";\n+ }\n+\n{noformat}\n\nIf possible can you add unit test with jobNames having metacharacters and \\E etc. ","created":"2008-10-24T03:47:50.157+0000"},{"body":"Relates to HADOOP-4017","created":"2008-10-24T03:50:07.157+0000"},{"body":"Chris,\nCan you help us reproduce this bug? What should be the jobname which can lead to this issue? Plz include a testcase which\n{code}\n1) Submits a job with the *special* job name\n2) Check if the job submission succeeds \n{code}\nThis test should fail without the fix.","created":"2008-10-24T04:44:56.846+0000"},{"body":"I'll add a unit test for this. ","created":"2008-10-24T05:20:21.322+0000"},{"body":"good catch on the regex, this works much better\n\n{noformat}\n private static String escapeRegexChars( String string ) {\n return \"\\\\Q\"+string.replaceAll(\"\\\\\\\\E\", \"\\\\\\\\E\\\\\\\\\\\\\\\\E\\\\\\\\Q\")+\"\\\\E\";\n }\n{noformat}\n\nvia a println of the pattern..\n\n{noformat}\njobName pattern\n------------------------------------------------\n[[name] value] \\Q[[name] value]\\E\nname \\Evalue] \\Qname \\E\\\\E\\Qvalue]\\E\n{noformat}\n\nlet me know if this doesn't seem right. i'll run the full test suite in the morning.","created":"2008-10-24T06:51:03.275+0000"},{"body":"escapes \\E properly. includes two unit tests as suggested","created":"2008-10-24T18:26:22.305+0000"},{"body":"all tests pass except TestLocalDirAllocator, which always fails for me (even with a clean tree).\n\n[exec] +1 overall. \n [exec] \n [exec] +1 @author. The patch does not contain any @author tags.\n [exec] \n [exec] +1 tests included. The patch appears to include 3 new or modified tests.\n [exec] \n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec] \n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec] \n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.","created":"2008-10-24T18:35:41.069+0000"},{"body":"I submitted the job with name : [[name] value], without the patch and the job was successful. \nI was able to parse the history of the jobs with complex names too, both from the web ui and history viewer.\nAlso the Testcase: testComplexName passes without the patch, though the Testcase: testComplexNameWithRegex failed\n\nCan you tell us how to reproduce the issue?\n\n","created":"2008-10-27T10:36:36.165+0000"},{"body":"both tests now properly exercise the bug","created":"2008-10-27T15:03:05.682+0000"},{"body":"Thank for being patient with this issues.\n\nthe first test now reflects the issue in the same manner it presented. ","created":"2008-10-27T15:04:42.240+0000"},{"body":"+1\nLatest patch looks good.","created":"2008-10-29T10:29:06.091+0000"},{"body":"anyone willing to commit this so we can get into the 0.19.x release?","created":"2008-10-30T00:24:16.563+0000"},{"body":"The following seems not working well.\n{noformat}\njobName pattern\n------------------------------------------------\nabc\\\\Exyz \\Qabc\\\\E\\\\E\\Qxyz\\E\n{noformat}","created":"2008-10-30T20:37:31.664+0000"},{"body":"> The following seems not working well...\nOops, I was wrong about that. The replacement is perfect.","created":"2008-10-30T20:48:05.045+0000"},{"body":"great news its working for your test case. thanks for testing this.","created":"2008-10-30T20:58:45.099+0000"},{"body":"I just committed this. Thanks, Chris!","created":"2008-10-31T00:11:32.626+0000"}],"conversations":[{"body":"JobHistory#getJobHistoryFileName throws a PatternSyntaxException if the jobName includes characters that could be regex metacharacters or metasequences and are subsequently considered malformed.\n\nTrace:\nException in thread \"initJobs\" java.util.regex.PatternSyntaxException: Unclosed character class near index 97\n....\n\tat java.util.regex.Pattern.error(Pattern.java:1713)\n\tat java.util.regex.Pattern.clazz(Pattern.java:2254)\n\tat java.util.regex.Pattern.clazz(Pattern.java:2210)\n\tat java.util.regex.Pattern.sequence(Pattern.java:1818)\n\tat java.util.regex.Pattern.expr(Pattern.java:1752)\n\tat java.util.regex.Pattern.compile(Pattern.java:1460)\n\tat java.util.regex.Pattern.(Pattern.java:1133)\n\tat java.util.regex.Pattern.compile(Pattern.java:823)\n\tat org.apache.hadoop.mapred.JobHistory$JobInfo.getJobHistoryFileName(JobHistory.java:632)\n\tat org.apache.hadoop.mapred.JobHistory$JobInfo.finalizeRecovery(JobHistory.java:740)\n\tat org.apache.hadoop.mapred.JobTracker.finalizeJob(JobTracker.java:1532)\n\tat org.apache.hadoop.mapred.JobInProgress.garbageCollect(JobInProgress.java:2232)\n\tat org.apache.hadoop.mapred.JobInProgress.terminateJob(JobInProgress.java:1938)\n\tat org.apache.hadoop.mapred.JobInProgress.terminate(JobInProgress.java:1953)\n\tat org.apache.hadoop.mapred.JobInProgress.fail(JobInProgress.java:2012)\n\tat org.apache.hadoop.mapred.EagerTaskInitializationListener$JobInitThread.run(EagerTaskInitializationListener.java:62)\n\tat java.lang.Thread.run(Thread.java:637)","from":"reporter","subject":"JobHistory does not escape literal jobName when used in a regex pattern"},{"body":"changed component affected","from":"developer"},{"body":"changed title","from":"developer"},{"body":"Stuffs jobName between \\Q and \\E so it is treated as a literal value.","from":"developer"},{"body":"All but org.apache.hadoop.fs.TestLocalDirAllocator tests pass. \n\nDid not include a unit test since most of these methods are private to JobHistory..\n\n [exec] -1 overall. \n [exec] \n [exec] +1 @author. The patch does not contain any @author tags.\n [exec] \n [exec] -1 tests included. The patch doesn't appear to include any new or modified tests.\n [exec] Please justify why no tests are needed for this patch.\n [exec] \n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec] \n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec] \n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.","from":"developer"},{"body":"If I'm not missing something, the following String replaceAll is not useful.\n{noformat} \n+ private static String escapeRegexChars( String string ) {\n+ return \"\\\\Q\"+string.replaceAll(\"\\\\\\\\E\", \"\\\\\\\\E\")+\"\\\\E\";\n+ }\n+\n{noformat}\n\nIf possible can you add unit test with jobNames having metacharacters and \\E etc. ","from":"developer"},{"body":"Relates to HADOOP-4017","from":"developer"},{"body":"Chris,\nCan you help us reproduce this bug? What should be the jobname which can lead to this issue? Plz include a testcase which\n{code}\n1) Submits a job with the *special* job name\n2) Check if the job submission succeeds \n{code}\nThis test should fail without the fix.","from":"developer"},{"body":"I'll add a unit test for this. ","from":"developer"},{"body":"good catch on the regex, this works much better\n\n{noformat}\n private static String escapeRegexChars( String string ) {\n return \"\\\\Q\"+string.replaceAll(\"\\\\\\\\E\", \"\\\\\\\\E\\\\\\\\\\\\\\\\E\\\\\\\\Q\")+\"\\\\E\";\n }\n{noformat}\n\nvia a println of the pattern..\n\n{noformat}\njobName pattern\n------------------------------------------------\n[[name] value] \\Q[[name] value]\\E\nname \\Evalue] \\Qname \\E\\\\E\\Qvalue]\\E\n{noformat}\n\nlet me know if this doesn't seem right. i'll run the full test suite in the morning.","from":"developer"},{"body":"escapes \\E properly. includes two unit tests as suggested","from":"developer"},{"body":"all tests pass except TestLocalDirAllocator, which always fails for me (even with a clean tree).\n\n[exec] +1 overall. \n [exec] \n [exec] +1 @author. The patch does not contain any @author tags.\n [exec] \n [exec] +1 tests included. The patch appears to include 3 new or modified tests.\n [exec] \n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec] \n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec] \n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.","from":"developer"},{"body":"I submitted the job with name : [[name] value], without the patch and the job was successful. \nI was able to parse the history of the jobs with complex names too, both from the web ui and history viewer.\nAlso the Testcase: testComplexName passes without the patch, though the Testcase: testComplexNameWithRegex failed\n\nCan you tell us how to reproduce the issue?\n\n","from":"developer"},{"body":"both tests now properly exercise the bug","from":"developer"},{"body":"Thank for being patient with this issues.\n\nthe first test now reflects the issue in the same manner it presented. ","from":"developer"},{"body":"+1\nLatest patch looks good.","from":"developer"},{"body":"anyone willing to commit this so we can get into the 0.19.x release?","from":"developer"},{"body":"The following seems not working well.\n{noformat}\njobName pattern\n------------------------------------------------\nabc\\\\Exyz \\Qabc\\\\E\\\\E\\Qxyz\\E\n{noformat}","from":"developer"},{"body":"> The following seems not working well...\nOops, I was wrong about that. The replacement is perfect.","from":"developer"},{"body":"great news its working for your test case. thanks for testing this.","from":"developer"},{"body":"I just committed this. Thanks, Chris!","from":"developer"}],"created":"2008-10-22T22:40:02.000+0000","description":"JobHistory#getJobHistoryFileName throws a PatternSyntaxException if the jobName includes characters that could be regex metacharacters or metasequences and are subsequently considered malformed.\n\nTrace:\nException in thread \"initJobs\" java.util.regex.PatternSyntaxException: Unclosed character class near index 97\n....\n\tat java.util.regex.Pattern.error(Pattern.java:1713)\n\tat java.util.regex.Pattern.clazz(Pattern.java:2254)\n\tat java.util.regex.Pattern.clazz(Pattern.java:2210)\n\tat java.util.regex.Pattern.sequence(Pattern.java:1818)\n\tat java.util.regex.Pattern.expr(Pattern.java:1752)\n\tat java.util.regex.Pattern.compile(Pattern.java:1460)\n\tat java.util.regex.Pattern.(Pattern.java:1133)\n\tat java.util.regex.Pattern.compile(Pattern.java:823)\n\tat org.apache.hadoop.mapred.JobHistory$JobInfo.getJobHistoryFileName(JobHistory.java:632)\n\tat org.apache.hadoop.mapred.JobHistory$JobInfo.finalizeRecovery(JobHistory.java:740)\n\tat org.apache.hadoop.mapred.JobTracker.finalizeJob(JobTracker.java:1532)\n\tat org.apache.hadoop.mapred.JobInProgress.garbageCollect(JobInProgress.java:2232)\n\tat org.apache.hadoop.mapred.JobInProgress.terminateJob(JobInProgress.java:1938)\n\tat org.apache.hadoop.mapred.JobInProgress.terminate(JobInProgress.java:1953)\n\tat org.apache.hadoop.mapred.JobInProgress.fail(JobInProgress.java:2012)\n\tat org.apache.hadoop.mapred.EagerTaskInitializationListener$JobInitThread.run(EagerTaskInitializationListener.java:62)\n\tat java.lang.Thread.run(Thread.java:637)","issue_id":"12407043","key":"HADOOP-4498","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-10-31T00:11:32.000+0000","role":"fixed_distractor","summary":"JobHistory does not escape literal jobName when used in a regex pattern"} {"case_id":"12407133","cluster":"DISTRACTOR-HADOOP-4510","comments":[{"body":"The simple solution is to make the method public.\n\nThe alternative of not having RecordWriters write to this magic directory seems more daunting considering this bit of Hadoop changes every release. ","created":"2008-10-24T00:33:00.584+0000"},{"body":"changes protected to public","created":"2008-10-24T00:34:21.630+0000"},{"body":"Users can get the task's temporary output path from the api FileOutputFormat.getWorkOutputPath(), where he can create side files and etc. \nThe method FileOutputFormat.getTaskOutputPath() is used internally by the RecordWriters. And it is protected sothat outputFormats extending FileOutputFormat can use this in RecordWriters. I dont see a reason why this should be made public.","created":"2008-10-24T03:38:13.844+0000"},{"body":"We prefer it public because we write through the FileOutputFormat class via a RecordWriter, which internally (magically) inserts the temp path and task id path at the end of the intended path. \n\nThis is done so that speculative execution will succeed. And we would like to benefit from this behavior, so aren't really asking that it change.\n\nThe side effect is that we have no way of finding the actual location of the data written and then moving it to where it was intended to be written. \n\nSince we don't have multiple (named) output collectors, we must emulate the behavior through our own api.","created":"2008-10-24T05:16:31.385+0000"},{"body":"I just committed this. Thanks, Chris!","created":"2008-10-26T04:17:32.130+0000"},{"body":"I see this is marked as fixed for 0.20. Any way it can leak into the 0.19.x release?","created":"2008-10-30T00:22:55.336+0000"},{"body":"Now this is pushed to the 0.19 branch, Jira should reflect the change.","created":"2008-10-31T00:30:15.085+0000"},{"body":"Great news, thanks!","created":"2008-10-31T00:35:04.915+0000"}],"conversations":[{"body":"o.a.h.m.FileOutputFormat#getTaskOutputPath() is protected. \n\nHaving access to a task output directory as used internally by RecordWriters is quite handy. This is especially true if the user is attempting to serialize out data in a similar fashion as the output collector.","from":"reporter","subject":"FileOutputFormat protects getTaskOutputPath"},{"body":"The simple solution is to make the method public.\n\nThe alternative of not having RecordWriters write to this magic directory seems more daunting considering this bit of Hadoop changes every release. ","from":"developer"},{"body":"changes protected to public","from":"developer"},{"body":"Users can get the task's temporary output path from the api FileOutputFormat.getWorkOutputPath(), where he can create side files and etc. \nThe method FileOutputFormat.getTaskOutputPath() is used internally by the RecordWriters. And it is protected sothat outputFormats extending FileOutputFormat can use this in RecordWriters. I dont see a reason why this should be made public.","from":"developer"},{"body":"We prefer it public because we write through the FileOutputFormat class via a RecordWriter, which internally (magically) inserts the temp path and task id path at the end of the intended path. \n\nThis is done so that speculative execution will succeed. And we would like to benefit from this behavior, so aren't really asking that it change.\n\nThe side effect is that we have no way of finding the actual location of the data written and then moving it to where it was intended to be written. \n\nSince we don't have multiple (named) output collectors, we must emulate the behavior through our own api.","from":"developer"},{"body":"I just committed this. Thanks, Chris!","from":"developer"},{"body":"I see this is marked as fixed for 0.20. Any way it can leak into the 0.19.x release?","from":"developer"},{"body":"Now this is pushed to the 0.19 branch, Jira should reflect the change.","from":"developer"},{"body":"Great news, thanks!","from":"developer"}],"created":"2008-10-24T00:30:46.000+0000","description":"o.a.h.m.FileOutputFormat#getTaskOutputPath() is protected. \n\nHaving access to a task output directory as used internally by RecordWriters is quite handy. This is especially true if the user is attempting to serialize out data in a similar fashion as the output collector.","issue_id":"12407133","key":"HADOOP-4510","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-10-26T04:17:32.000+0000","role":"fixed_distractor","summary":"FileOutputFormat protects getTaskOutputPath"} {"case_id":"12408175","cluster":"DISTRACTOR-HADOOP-4626","comments":[{"body":"These should be replaced in the xml source with relative links to api/org/apache/hadoop/...\n","created":"2008-11-10T19:46:30.426+0000"},{"body":"The setSafeMode() link indeed is a missing link, instead of a broken link, since the hdfs api docs are not available in the public web site. Similarly, links to HDFS Java API do not exist.","created":"2009-02-20T20:07:24.443+0000"},{"body":"4626_20090526.patch: fix API version.\n\nThe setSafeMode() link will be fixed in HADOOP-5918.","created":"2009-05-26T22:16:58.485+0000"},{"body":"{noformat}\n [exec] -1 overall. \n [exec] \n [exec] +1 @author. The patch does not contain any @author tags.\n [exec] \n [exec] -1 tests included. The patch doesn't appear to include any new or modified tests.\n [exec] Please justify why no tests are needed for this patch.\n [exec] \n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec] \n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec] \n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec] \n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n [exec] \n [exec] +1 release audit. The applied patch does not increase the total number of release audit warnings.\n{noformat}\nThis is a documentation change. No sure why \"ant test-patch\" cannot detect it but -1 on tests included.","created":"2009-05-27T14:43:22.156+0000"},{"body":"I checked the generated doc manually. The links are correct.","created":"2009-05-27T14:48:21.816+0000"},{"body":"+1 \nGenerated doc manually and verified for link correctness. Patch looks good.\nThanks Nicholas. ","created":"2009-05-27T21:57:59.554+0000"},{"body":"I have committed this to 0.20 and above.","created":"2009-05-27T22:07:29.336+0000"}],"conversations":[{"body":"For example, in http://hadoop.apache.org/core/docs/r0.17.2/hdfs_user_guide.html, the setSafeMode() link is pointing to http ://hadoop.apache.org/core/docs/current/..., which is not an 0.17 doc.","from":"reporter","subject":"API link in forrest doc should point to the same version of hadoop."},{"body":"These should be replaced in the xml source with relative links to api/org/apache/hadoop/...\n","from":"developer"},{"body":"The setSafeMode() link indeed is a missing link, instead of a broken link, since the hdfs api docs are not available in the public web site. Similarly, links to HDFS Java API do not exist.","from":"developer"},{"body":"4626_20090526.patch: fix API version.\n\nThe setSafeMode() link will be fixed in HADOOP-5918.","from":"developer"},{"body":"{noformat}\n [exec] -1 overall. \n [exec] \n [exec] +1 @author. The patch does not contain any @author tags.\n [exec] \n [exec] -1 tests included. The patch doesn't appear to include any new or modified tests.\n [exec] Please justify why no tests are needed for this patch.\n [exec] \n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec] \n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec] \n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec] \n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n [exec] \n [exec] +1 release audit. The applied patch does not increase the total number of release audit warnings.\n{noformat}\nThis is a documentation change. No sure why \"ant test-patch\" cannot detect it but -1 on tests included.","from":"developer"},{"body":"I checked the generated doc manually. The links are correct.","from":"developer"},{"body":"+1 \nGenerated doc manually and verified for link correctness. Patch looks good.\nThanks Nicholas. ","from":"developer"},{"body":"I have committed this to 0.20 and above.","from":"developer"}],"created":"2008-11-10T19:14:28.000+0000","description":"For example, in http://hadoop.apache.org/core/docs/r0.17.2/hdfs_user_guide.html, the setSafeMode() link is pointing to http ://hadoop.apache.org/core/docs/current/..., which is not an 0.17 doc.","issue_id":"12408175","key":"HADOOP-4626","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-05-27T22:07:29.000+0000","role":"fixed_distractor","summary":"API link in forrest doc should point to the same version of hadoop."} {"case_id":"12408866","cluster":"DISTRACTOR-HADOOP-4692","comments":[{"body":"The block file of blk_B on DN1 shows that the on-disk block size is 134205440. So the only replica of this block is truncated and therefore corrupted but reading this block does not cause ChecksumException.","created":"2008-11-20T00:28:41.112+0000"},{"body":"Currently NameNode does not detect that a replica is truncated and therefore corrupted. One way to solve this is to let block report handling check each block's\nlength and mark those truncated blocks as corrupted. Also when NN receives a new block that's truncated, NN should mark it as corrupted instead of adding it to recent invalidates directly. Once NameNode finds out all replicas are corrupted, it will stop replicating/deleting a block.\n\n","created":"2008-11-20T00:34:39.889+0000"},{"body":"This is a duplicate of HADOOP-3314, but this really writes up the problem better... can we close the older ticket?\n\nWe've run into this issue locally, and it's rather debilitating because it can result in \"silent corruptions\": these truncations can accumulate for a long time without anything noticing. If you are running with 2 replicas (hey, not all of us can afford all that raw disk space...) and lose a data node, then this can result in a nasty surprise if the second copy had this truncation problem.\n\nThis in fact has caused corruption for about 500 files locally.","created":"2008-11-25T04:17:12.749+0000"},{"body":"When an inconsistently-sized file is found, we need to trigger a block scan of all the possible sources. This depends on our ability to trigger manual scans, which is precisely what HADOOP-4865 is for.","created":"2008-12-14T20:28:30.604+0000"},{"body":"Attached a file which is my first whack of a patch that builds on HADOOP-4865.","created":"2008-12-14T20:31:18.921+0000"},{"body":"I'll be able to work on an updated patch tomorrow -- for now, this approach appears to be working.","created":"2008-12-16T17:31:13.674+0000"},{"body":"Bah - the approach works to trigger verification, but the verification doesn't catch the fact that there's too little data (the metadata is computed for the truncated block. In fact, the block does verify just fine!","created":"2008-12-16T18:12:58.988+0000"},{"body":"Another idea is to pass the NN recorded block length when replicating a block. (currently -1 is passed). When sender sees that its on-disk length is less than the asked leghth, report corrupt block to NN and stop replication.","created":"2008-12-16T18:41:15.550+0000"},{"body":"A patch is attached for review.","created":"2009-01-10T00:52:13.418+0000"},{"body":"+1 for reporting a block as corrupt to NN.\nRegd implementation :\n\nThe patch makes BlockSender to report the corruption (implicitly assuming that null client implies a transfer). I think this approach mixes higher level policy with lower level implementaion. \n\nMy suggestion would be to make BlockSender throw an excection (it throws IOException now, it could throw TruncatedBlockException in stead). Then make the block transfer thread (in DataNode.java) to catch it and report the corrupt block to NN.\n\n","created":"2009-01-13T00:39:25.790+0000"},{"body":"Thanks Raghu for the comment. Yes, I like your suggestion. But still with your approach, BiockSender needs to know if the block reading is for block transfer or not by checking if the client name before throwing TruncateBlockException. Would this be OK?\n\nAnother question is what should BlockSender do if the on-disk block length is longer than the NN recorded length? Currently block replication only copies the number of bytes recorded by NN. Is this a good idea?","created":"2009-01-13T00:47:27.883+0000"},{"body":"> BiockSender needs to know if the block reading is for block transfer or not by checking if the client name before throwing TruncateBlockException. Would this be OK? \n\nI don't think so. Right now it always throws IOException. We just needs to change the exception so that higher levels can distinguish. \n\n> Another question is what should BlockSender do if the on-disk block length is longer than the NN recorded length? Currently block replication only copies the number of bytes recorded by NN. Is this a good idea?\n\nCopying only the bytes requested by NN is ok (as far as NN is concerned). Similar to previous comment, I don't think BlockSender should worry about it, but some higher level in DataNode... I am +0 on fixing \"extra data\" issue. But if we want to, DataTransfer thread could check for the right size before even creating a BlockSender. ","created":"2009-01-13T01:12:44.322+0000"},{"body":"> Copying only the bytes requested by NN is ok (as far as NN is concerned).\n\nI am still not sure of this. If the block is being written to, a longer block is also a corrupt block. If the block is being written to, then copying partial data is useless.\n\nHi Dhruba, could you please clarify if it is possible that ReplicationMonitor may replicate a block that's being written to after the introduction of sync & append?","created":"2009-01-14T21:43:48.325+0000"},{"body":"f the NN and the DN have the same generation stamp, then the file is either not-open or the file is marked as \"under construction\" at the namenode.\n4:47 so, the NN will not start any new replication requests for these blocks (via HADOOP-5027)","created":"2009-01-15T00:50:18.450+0000"},{"body":"> so, the NN will not start any new replication requests for these blocks (via HADOOP-5027) \n\nNN won't start new requests but what if there are scheduled requests?","created":"2009-01-15T19:31:19.936+0000"},{"body":"Previously scheduled replication requests willl complete and the new destination datanode will send a blockReceived message to the NN. In the meantime, if the file has been opened for \"append\", then the generation stamp on the namenode should have been bumped. If the blockReceived arrived at the NN after the generation stamp has been bumped, then the blockReceived will not be able to find this block in the blocksMap.\n\ni think this should not cause any issues. Any race condition that I might have missed?","created":"2009-01-15T22:21:11.726+0000"},{"body":"Ok if HADOOP-5027 makes sure that blocks under construction do not add to blocksMap, I will treat on-disk blocks whose length is inconsistent with NN recorded length as corrupt and datanodes stop replicating them.","created":"2009-01-16T22:38:47.418+0000"},{"body":"My understanding is that the NN will send the block length ( as recorded in NN metadata) to the source datanode of the replication request. The source datanode will verify that this length matches the length of the block file on disk. If it does not match, then the source datanode will not replicate the block. Is my understanding correct?","created":"2009-01-17T00:51:57.303+0000"},{"body":"In the current trunk, the source datanode ignores the block length that NN sent and uses the on-disk block length to transfer the block.\n\nWhat I plan to do is that when receiving a block replication request, datanode first checks if this block is under construction or not by looking at the ongoingCreates list. If yes, stop replicating the block. Otherwise check if the on-disk block length is the same as the block length sent by NN. If no, report NN corrupt blocks and stop replicating. Otherwise, start replicated the block.","created":"2009-01-22T21:25:30.857+0000"},{"body":"With this patch, this issue does not depend on HADOOP-5027 any more.","created":"2009-01-23T01:19:02.815+0000"},{"body":" [exec]\n [exec] +1 @author. The patch does not contain any @author tags.\n [exec]\n [exec] +1 tests included. The patch appears to include 9 new or modified tests.\n [exec]\n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec]\n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec]\n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec]\n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n [exec]\n\nAnt test-core had the following known failures:\n [junit] Tests run: 1, Failures: 1, Errors: 0, Time elapsed: 1.35 sec\n [junit] Test org.apache.hadoop.http.TestGlobalFilter FAILED\n [junit] Running org.apache.hadoop.mapreduce.TestMapReduceLocal\n [junit] Tests run: 1, Failures: 0, Errors: 1, Time elapsed: 29.481 sec","created":"2009-01-26T19:59:13.442+0000"},{"body":"[See related comment here|https://issues.apache.org/jira/browse/HADOOP-5027#action_12668136]","created":"2009-01-28T20:14:09.508+0000"},{"body":"Compared to the last patch, this patch has two changes:\n1. Remove the check isUnderConstruction because isValid return false if a block is still under construction.\n2. If a block's on-disk block size is bigger than the NN recorded length, do not mark it as corrupt. Instead copy the number of bytes that NN asks for.\n\nChange 2 is being cautious. When I work on HADOOP-5133, I realize that in the current trunk, a block's length may not be finalized even when a file is closed. So marking a block to be corrupt using NN recorded length is too dangerous.","created":"2009-02-18T19:32:47.772+0000"},{"body":"+1. The patch looks good makes good sense to me. \n\nRegd correctness and how it fits in the larger context of related fixes like HADOOP-5133, HADOO-5027, I haven't looked into much. That area of HDFS is under lot of flux.\n\nbtw, the 'links' for the jira says HADOOP-3314 is a duplicate of this, is it still true? Mostly HADOOP-3314 still needs to be fixed.","created":"2009-02-18T20:01:45.703+0000"},{"body":"I've just committed this.\n\nRegarding Raghu's concern, this patch is based on the assumption that the length of a block (identified by its block id & generation stamp) recorded on the NN's side can only grow but never shrink. I will keep an eye that HADOOP-5133 and HADOOP-5027 observe this assumption.","created":"2009-02-19T21:58:16.252+0000"}],"conversations":[{"body":"Our cluster has an under-replicated block with only one replica, assuming its block id is B. NameNode log shows that NameNode is in an infinite loop replicating/deleting the block.\n\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* ask DN1 to replicate blk_B to datanode(s) DN2, DN3\nWARN org.apache.hadoop.fs.FSNamesystem: Inconsistent size for block blk_B reported from DN2 current size is 134217728 reported size is 134205440\nWARN org.apache.hadoop.fs.FSNamesystem: Deleting block blk_B from DN2\nINFO org.apache.hadoop.dfs.StateChange: DIR* NameSystem.invalidateBlock: blk_B on DN2\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* NameSystem.delete: blk_B is added to invalidSet of DN2\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* NameSystem.addStoredBlock: blockMap updated: DN2 is added to blk_B size 134217728\nWARN org.apache.hadoop.fs.FSNamesystem: Inconsistent size for block blk_-B reported from DN3 current size is 134217728 reported size is 134205440\nWARN org.apache.hadoop.fs.FSNamesystem: Deleting block blk_B from DN3\nINFO org.apache.hadoop.dfs.StateChange: DIR* NameSystem.invalidateBlock: blk_B on DN3\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* NameSystem.delete: blk_B is added to invalidSet of DN3\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* NameSystem.addStoredBlock: blockMap updated: DN3 is added to blk_B size 134217728\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* ask DN1 to replicate blk_B to datanode(s) DN4, DN5\n...\n","from":"reporter","subject":" Namenode in infinite loop for replicating/deleting corrupted block"},{"body":"The block file of blk_B on DN1 shows that the on-disk block size is 134205440. So the only replica of this block is truncated and therefore corrupted but reading this block does not cause ChecksumException.","from":"developer"},{"body":"Currently NameNode does not detect that a replica is truncated and therefore corrupted. One way to solve this is to let block report handling check each block's\nlength and mark those truncated blocks as corrupted. Also when NN receives a new block that's truncated, NN should mark it as corrupted instead of adding it to recent invalidates directly. Once NameNode finds out all replicas are corrupted, it will stop replicating/deleting a block.\n\n","from":"developer"},{"body":"This is a duplicate of HADOOP-3314, but this really writes up the problem better... can we close the older ticket?\n\nWe've run into this issue locally, and it's rather debilitating because it can result in \"silent corruptions\": these truncations can accumulate for a long time without anything noticing. If you are running with 2 replicas (hey, not all of us can afford all that raw disk space...) and lose a data node, then this can result in a nasty surprise if the second copy had this truncation problem.\n\nThis in fact has caused corruption for about 500 files locally.","from":"developer"},{"body":"When an inconsistently-sized file is found, we need to trigger a block scan of all the possible sources. This depends on our ability to trigger manual scans, which is precisely what HADOOP-4865 is for.","from":"developer"},{"body":"Attached a file which is my first whack of a patch that builds on HADOOP-4865.","from":"developer"},{"body":"I'll be able to work on an updated patch tomorrow -- for now, this approach appears to be working.","from":"developer"},{"body":"Bah - the approach works to trigger verification, but the verification doesn't catch the fact that there's too little data (the metadata is computed for the truncated block. In fact, the block does verify just fine!","from":"developer"},{"body":"Another idea is to pass the NN recorded block length when replicating a block. (currently -1 is passed). When sender sees that its on-disk length is less than the asked leghth, report corrupt block to NN and stop replication.","from":"developer"},{"body":"A patch is attached for review.","from":"developer"},{"body":"+1 for reporting a block as corrupt to NN.\nRegd implementation :\n\nThe patch makes BlockSender to report the corruption (implicitly assuming that null client implies a transfer). I think this approach mixes higher level policy with lower level implementaion. \n\nMy suggestion would be to make BlockSender throw an excection (it throws IOException now, it could throw TruncatedBlockException in stead). Then make the block transfer thread (in DataNode.java) to catch it and report the corrupt block to NN.\n\n","from":"developer"},{"body":"Thanks Raghu for the comment. Yes, I like your suggestion. But still with your approach, BiockSender needs to know if the block reading is for block transfer or not by checking if the client name before throwing TruncateBlockException. Would this be OK?\n\nAnother question is what should BlockSender do if the on-disk block length is longer than the NN recorded length? Currently block replication only copies the number of bytes recorded by NN. Is this a good idea?","from":"developer"},{"body":"> BiockSender needs to know if the block reading is for block transfer or not by checking if the client name before throwing TruncateBlockException. Would this be OK? \n\nI don't think so. Right now it always throws IOException. We just needs to change the exception so that higher levels can distinguish. \n\n> Another question is what should BlockSender do if the on-disk block length is longer than the NN recorded length? Currently block replication only copies the number of bytes recorded by NN. Is this a good idea?\n\nCopying only the bytes requested by NN is ok (as far as NN is concerned). Similar to previous comment, I don't think BlockSender should worry about it, but some higher level in DataNode... I am +0 on fixing \"extra data\" issue. But if we want to, DataTransfer thread could check for the right size before even creating a BlockSender. ","from":"developer"},{"body":"> Copying only the bytes requested by NN is ok (as far as NN is concerned).\n\nI am still not sure of this. If the block is being written to, a longer block is also a corrupt block. If the block is being written to, then copying partial data is useless.\n\nHi Dhruba, could you please clarify if it is possible that ReplicationMonitor may replicate a block that's being written to after the introduction of sync & append?","from":"developer"},{"body":"f the NN and the DN have the same generation stamp, then the file is either not-open or the file is marked as \"under construction\" at the namenode.\n4:47 so, the NN will not start any new replication requests for these blocks (via HADOOP-5027)","from":"developer"},{"body":"> so, the NN will not start any new replication requests for these blocks (via HADOOP-5027) \n\nNN won't start new requests but what if there are scheduled requests?","from":"developer"},{"body":"Previously scheduled replication requests willl complete and the new destination datanode will send a blockReceived message to the NN. In the meantime, if the file has been opened for \"append\", then the generation stamp on the namenode should have been bumped. If the blockReceived arrived at the NN after the generation stamp has been bumped, then the blockReceived will not be able to find this block in the blocksMap.\n\ni think this should not cause any issues. Any race condition that I might have missed?","from":"developer"},{"body":"Ok if HADOOP-5027 makes sure that blocks under construction do not add to blocksMap, I will treat on-disk blocks whose length is inconsistent with NN recorded length as corrupt and datanodes stop replicating them.","from":"developer"},{"body":"My understanding is that the NN will send the block length ( as recorded in NN metadata) to the source datanode of the replication request. The source datanode will verify that this length matches the length of the block file on disk. If it does not match, then the source datanode will not replicate the block. Is my understanding correct?","from":"developer"},{"body":"In the current trunk, the source datanode ignores the block length that NN sent and uses the on-disk block length to transfer the block.\n\nWhat I plan to do is that when receiving a block replication request, datanode first checks if this block is under construction or not by looking at the ongoingCreates list. If yes, stop replicating the block. Otherwise check if the on-disk block length is the same as the block length sent by NN. If no, report NN corrupt blocks and stop replicating. Otherwise, start replicated the block.","from":"developer"},{"body":"With this patch, this issue does not depend on HADOOP-5027 any more.","from":"developer"},{"body":" [exec]\n [exec] +1 @author. The patch does not contain any @author tags.\n [exec]\n [exec] +1 tests included. The patch appears to include 9 new or modified tests.\n [exec]\n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec]\n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec]\n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec]\n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n [exec]\n\nAnt test-core had the following known failures:\n [junit] Tests run: 1, Failures: 1, Errors: 0, Time elapsed: 1.35 sec\n [junit] Test org.apache.hadoop.http.TestGlobalFilter FAILED\n [junit] Running org.apache.hadoop.mapreduce.TestMapReduceLocal\n [junit] Tests run: 1, Failures: 0, Errors: 1, Time elapsed: 29.481 sec","from":"developer"},{"body":"[See related comment here|https://issues.apache.org/jira/browse/HADOOP-5027#action_12668136]","from":"developer"},{"body":"Compared to the last patch, this patch has two changes:\n1. Remove the check isUnderConstruction because isValid return false if a block is still under construction.\n2. If a block's on-disk block size is bigger than the NN recorded length, do not mark it as corrupt. Instead copy the number of bytes that NN asks for.\n\nChange 2 is being cautious. When I work on HADOOP-5133, I realize that in the current trunk, a block's length may not be finalized even when a file is closed. So marking a block to be corrupt using NN recorded length is too dangerous.","from":"developer"},{"body":"+1. The patch looks good makes good sense to me. \n\nRegd correctness and how it fits in the larger context of related fixes like HADOOP-5133, HADOO-5027, I haven't looked into much. That area of HDFS is under lot of flux.\n\nbtw, the 'links' for the jira says HADOOP-3314 is a duplicate of this, is it still true? Mostly HADOOP-3314 still needs to be fixed.","from":"developer"},{"body":"I've just committed this.\n\nRegarding Raghu's concern, this patch is based on the assumption that the length of a block (identified by its block id & generation stamp) recorded on the NN's side can only grow but never shrink. I will keep an eye that HADOOP-5133 and HADOOP-5027 observe this assumption.","from":"developer"}],"created":"2008-11-20T00:24:08.000+0000","description":"Our cluster has an under-replicated block with only one replica, assuming its block id is B. NameNode log shows that NameNode is in an infinite loop replicating/deleting the block.\n\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* ask DN1 to replicate blk_B to datanode(s) DN2, DN3\nWARN org.apache.hadoop.fs.FSNamesystem: Inconsistent size for block blk_B reported from DN2 current size is 134217728 reported size is 134205440\nWARN org.apache.hadoop.fs.FSNamesystem: Deleting block blk_B from DN2\nINFO org.apache.hadoop.dfs.StateChange: DIR* NameSystem.invalidateBlock: blk_B on DN2\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* NameSystem.delete: blk_B is added to invalidSet of DN2\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* NameSystem.addStoredBlock: blockMap updated: DN2 is added to blk_B size 134217728\nWARN org.apache.hadoop.fs.FSNamesystem: Inconsistent size for block blk_-B reported from DN3 current size is 134217728 reported size is 134205440\nWARN org.apache.hadoop.fs.FSNamesystem: Deleting block blk_B from DN3\nINFO org.apache.hadoop.dfs.StateChange: DIR* NameSystem.invalidateBlock: blk_B on DN3\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* NameSystem.delete: blk_B is added to invalidSet of DN3\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* NameSystem.addStoredBlock: blockMap updated: DN3 is added to blk_B size 134217728\nINFO org.apache.hadoop.dfs.StateChange: BLOCK* ask DN1 to replicate blk_B to datanode(s) DN4, DN5\n...\n","issue_id":"12408866","key":"HADOOP-4692","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-02-19T21:58:16.000+0000","role":"fixed_distractor","summary":" Namenode in infinite loop for replicating/deleting corrupted block"} {"case_id":"12409114","cluster":"DISTRACTOR-HADOOP-4716","comments":[{"body":"Seems it gets stuck trying to copy output files. These messages loop until it times out.\n\n [exec] [junit] 2008-11-23 08:19:01,944 INFO mapred.TaskTracker (TaskTracker.java:reportProgress(1898)) - attempt_200811230815_0001_r_000000_0 0.10666667% reduce > copy (16 of 50 at 0.00 MB/s) >\n [exec] [junit] 2008-11-23 08:19:03,104 INFO mapred.TaskTracker (TaskTracker.java:tryToGetOutputSize(1514)) - org.apache.hadoop.util.DiskChecker$DiskErrorException: Could not find taskTracker/jobcache/job_200811230815_0001/attempt_200811230815_0001_r_000000_0/output/file.out in any of the configured local directories\n","created":"2008-11-24T12:51:47.028+0000"},{"body":"Johan, it might look like the line\n{code}\n[exec] [junit] 2008-11-23 08:19:01,944 INFO mapred.TaskTracker (TaskTracker.java:reportProgress(1898)) - \n attempt_200811230815_0001_r_000000_0 0.10666667% reduce > copy (16 of 50 at 0.00 MB/s) >\n{code}\nis in a loop but if you observe the part \n{code}\ncopy (16 of 50 at 0.00 MB/s)\n ^^\n{code}\nyou will see that the number (marked with ^^) keeps on changing. The job runs 50 map tasks and hence the reducer copies 50 maps outputs. Let me know if this is not the case. I ran this test on my box and it passed. I will run it in a loop to see if I can reproduce this. Can you attach the logs for this test?","created":"2008-11-26T06:09:32.626+0000"},{"body":"Here's an example of a similar behavior. As you can see it gets stuck when trying to copy the map output, in this case for roughly 10 minutes before the test times out.","created":"2008-11-26T15:24:08.976+0000"},{"body":"The JobTracker upon restart rebuilds the _task-completion-event_ list. Here there are events from the tracker which was lost upon restart. When the task-tracker (re)connects it re-sizes its own _task-completion-event_ list. Hence the tracker retains the missing map's events. After some time the jobtracker finds out that the tracker is lost and kills all the maps that were run on the lost tracker and re-executes them. The tracker will have the _task-completion-event_ list like \n{code}\n1. SUC m1-t1\n2. SUC m2-t2\n3. SUC m3-t1\n4. SUC m4-t2\n5. KIL m1-t1\n6. KIL m3-t1\n7. SUC m1-t2\n8. SUC m3-t2\n{code}\nThe reducer takes _m1-t1_ and starts pulling map output from _t1_. Note that when the reducer fails on _m1_ it checks that _m1_ is _OBSOLETE_ and then ignores it. The test case times out because it takes fair amount of time (~3mins) to fail once. So this doesnt look like a bug but a limitation. The reason this issue is not commonly seen is because the reducer actually starts late and hence the tracker has the latest updates which prevents the reducer to take up maps from the lost tracker. I could easily reproduce this problem when the reducer was scheduled early. \n----\nOne thing that can be done here is to make _num-reducers=0_ as the test case doesnt actually require reducers. But actually its better to have reducers as it makes the testcase strict and hence better. So if we decide to keep reducers then there should be some way to control the timeout (~3min --> ~5 secs). Thoughts?","created":"2008-11-27T13:54:14.923+0000"},{"body":"Also, the reducers build up a list of known map outputs from the map completion events obtained from the task-tracker. Upon restart (HADOOP-3245), the reducer doesn't clear this list. The problem arises (on a very small cluster e.g. test cases) when a map completes and this completion event is not flushed to the history file (in buffer) and the tracker on which it ran gets lost. In such a case the (new) JobTracker has no idea about the map and since the reducer's list is stale, it takes some time to figure out that the map location is bad and use the other (re-executed by the _new_ JobTracker) one. Note than on a large cluster, this should not be an issue as, upon failure, new maps from different host will be tried and after sometime the new location for such (dangling) maps will be passed by the JobTracker. We have 2 choices\n- stale data : This can help in cases where a tracker has not yet joined but the data (map's output) is still valid/available. The drawback being the case where a tracker is lost and data becomes unavailable. In such a case the location will be retried again and again until the (newly re-executed) map's output is pulled from some other tracker. Here the time will be wasted in pulling map output from a lost tracker/node and waiting for the (dangling) maps (from that node) to be re-executed.\n- fresh data : This can help in cases where few trackers go down. The drawback being that trackers that are up and ready to serve the map output will be ignored since they are yet to join. Here the time will be wasted in waiting for the tracker to _formally_ join back.","created":"2008-12-02T04:28:58.789+0000"},{"body":"Also the following tweaks causes the test _TestJobTrackerRestartWithLostTracker_ to pass consistently.\n- _timeout_ : Set DEFAULT_READ_TIMEOUT = 3 sec (default = 3min) and DEFAULT_CONNECT_TIMEOUT = 100 msecs (default = 3 secs) in {{ReduceTask.java}}.\n- _freshness_ : clearing the _known output_ list in {{ReduceTask.java}} upon a restart.","created":"2008-12-02T04:36:50.328+0000"},{"body":"I think we should not expose such low-level timeouts to the configuration (at least not until it is really required to do so).. \nAlso, in this testcase, it seems reasonable to not have the reducer at all.","created":"2008-12-10T06:09:19.949+0000"},{"body":"Attaching a patch that \n- reduced max backoff so that the retries happen frequently\n- cleared the _known-outputs_ structure so that the reducer doesnt consider events not known to jobtracker\n \nTested on my box and works fine. Not tested it throughly though. Somethings that would also help : \n- decrease the num-maps from 50 to 30. This will reduce the test's runt time.","created":"2008-12-10T10:03:40.825+0000"},{"body":"Marking as a blocker, as multiple patches are failing on Hudson on this.","created":"2008-12-15T03:42:45.305+0000"},{"body":"Looks like HADOOP-4683 solves this issue for now. The way HADOOP-4683 solves this issue is that it fetches all the events in one go and latest events are fetched faster. So the fact that a map location is faulty/missing is immediately passed onto the reducer which adds the map output location to the _ignore list_. I couldn't reproduce this bug with trunk. Also HADOOP-4220 tries to reduce the run-time of this testcase. ","created":"2008-12-15T11:44:59.435+0000"},{"body":"Looks like {{TestJobTrackerRestartWithLostTracker}} can still fail. Consider a map that is \n- hosted by a tracker that will be lost \n- the map completion event is not logged to the job history (i.e its in the buffer)\n\nCall it a _hanging-map_. Once the reducer reaches the _hanging-map_, it will be stuck there forever as the map location is not known to the jobtracker and hence wont be added to the _ignore-list_. The previous fix solves the issue. What it does is :\n- clears the old/stale mapping of map output locations\n- reduces the backoff time so that the reducer doesnt back off for long\n\nUpdates the patch to trunk.\n","created":"2008-12-16T09:47:48.720+0000"},{"body":"You missed one check for null for the returned list from mapLocations. That should be added..","created":"2008-12-19T04:59:07.344+0000"},{"body":"Incorporated Devaraj's comments. Result of _test-patch_ \n{noformat}\n[exec] +1 overall. \n [exec] \n [exec] +1 @author. The patch does not contain any @author tags.\n [exec] \n [exec] +1 tests included. The patch appears to include 3 new or modified tests.\n [exec] \n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec] \n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec] \n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec] \n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n{noformat}","created":"2008-12-19T06:28:22.314+0000"},{"body":"if knownOutputsByLoc is null, we should break out of the loop instead of continuing","created":"2008-12-19T08:49:58.799+0000"},{"body":"In this patch we make no implicit assumption. Let it run in the loop. It should not be a hit as retry-fetches and num-hosts are not expected to be much. Added comments.","created":"2008-12-19T10:41:51.997+0000"},{"body":"I just committed this. Thanks, Amar!","created":"2008-12-22T04:42:18.019+0000"}],"conversations":[{"body":"This test frequently times out: org.apache.hadoop.mapred.TestJobTrackerRestartWithLostTracker.testRestartWithLostTracker\nExample: http://hudson.zones.apache.org/hudson/job/Hadoop-Patch/3637/testReport/org.apache.hadoop.mapred/TestJobTrackerRestartWithLostTracker/testRestartWithLostTracker/","from":"reporter","subject":"testRestartWithLostTracker frequently times out"},{"body":"Seems it gets stuck trying to copy output files. These messages loop until it times out.\n\n [exec] [junit] 2008-11-23 08:19:01,944 INFO mapred.TaskTracker (TaskTracker.java:reportProgress(1898)) - attempt_200811230815_0001_r_000000_0 0.10666667% reduce > copy (16 of 50 at 0.00 MB/s) >\n [exec] [junit] 2008-11-23 08:19:03,104 INFO mapred.TaskTracker (TaskTracker.java:tryToGetOutputSize(1514)) - org.apache.hadoop.util.DiskChecker$DiskErrorException: Could not find taskTracker/jobcache/job_200811230815_0001/attempt_200811230815_0001_r_000000_0/output/file.out in any of the configured local directories\n","from":"developer"},{"body":"Johan, it might look like the line\n{code}\n[exec] [junit] 2008-11-23 08:19:01,944 INFO mapred.TaskTracker (TaskTracker.java:reportProgress(1898)) - \n attempt_200811230815_0001_r_000000_0 0.10666667% reduce > copy (16 of 50 at 0.00 MB/s) >\n{code}\nis in a loop but if you observe the part \n{code}\ncopy (16 of 50 at 0.00 MB/s)\n ^^\n{code}\nyou will see that the number (marked with ^^) keeps on changing. The job runs 50 map tasks and hence the reducer copies 50 maps outputs. Let me know if this is not the case. I ran this test on my box and it passed. I will run it in a loop to see if I can reproduce this. Can you attach the logs for this test?","from":"developer"},{"body":"Here's an example of a similar behavior. As you can see it gets stuck when trying to copy the map output, in this case for roughly 10 minutes before the test times out.","from":"developer"},{"body":"The JobTracker upon restart rebuilds the _task-completion-event_ list. Here there are events from the tracker which was lost upon restart. When the task-tracker (re)connects it re-sizes its own _task-completion-event_ list. Hence the tracker retains the missing map's events. After some time the jobtracker finds out that the tracker is lost and kills all the maps that were run on the lost tracker and re-executes them. The tracker will have the _task-completion-event_ list like \n{code}\n1. SUC m1-t1\n2. SUC m2-t2\n3. SUC m3-t1\n4. SUC m4-t2\n5. KIL m1-t1\n6. KIL m3-t1\n7. SUC m1-t2\n8. SUC m3-t2\n{code}\nThe reducer takes _m1-t1_ and starts pulling map output from _t1_. Note that when the reducer fails on _m1_ it checks that _m1_ is _OBSOLETE_ and then ignores it. The test case times out because it takes fair amount of time (~3mins) to fail once. So this doesnt look like a bug but a limitation. The reason this issue is not commonly seen is because the reducer actually starts late and hence the tracker has the latest updates which prevents the reducer to take up maps from the lost tracker. I could easily reproduce this problem when the reducer was scheduled early. \n----\nOne thing that can be done here is to make _num-reducers=0_ as the test case doesnt actually require reducers. But actually its better to have reducers as it makes the testcase strict and hence better. So if we decide to keep reducers then there should be some way to control the timeout (~3min --> ~5 secs). Thoughts?","from":"developer"},{"body":"Also, the reducers build up a list of known map outputs from the map completion events obtained from the task-tracker. Upon restart (HADOOP-3245), the reducer doesn't clear this list. The problem arises (on a very small cluster e.g. test cases) when a map completes and this completion event is not flushed to the history file (in buffer) and the tracker on which it ran gets lost. In such a case the (new) JobTracker has no idea about the map and since the reducer's list is stale, it takes some time to figure out that the map location is bad and use the other (re-executed by the _new_ JobTracker) one. Note than on a large cluster, this should not be an issue as, upon failure, new maps from different host will be tried and after sometime the new location for such (dangling) maps will be passed by the JobTracker. We have 2 choices\n- stale data : This can help in cases where a tracker has not yet joined but the data (map's output) is still valid/available. The drawback being the case where a tracker is lost and data becomes unavailable. In such a case the location will be retried again and again until the (newly re-executed) map's output is pulled from some other tracker. Here the time will be wasted in pulling map output from a lost tracker/node and waiting for the (dangling) maps (from that node) to be re-executed.\n- fresh data : This can help in cases where few trackers go down. The drawback being that trackers that are up and ready to serve the map output will be ignored since they are yet to join. Here the time will be wasted in waiting for the tracker to _formally_ join back.","from":"developer"},{"body":"Also the following tweaks causes the test _TestJobTrackerRestartWithLostTracker_ to pass consistently.\n- _timeout_ : Set DEFAULT_READ_TIMEOUT = 3 sec (default = 3min) and DEFAULT_CONNECT_TIMEOUT = 100 msecs (default = 3 secs) in {{ReduceTask.java}}.\n- _freshness_ : clearing the _known output_ list in {{ReduceTask.java}} upon a restart.","from":"developer"},{"body":"I think we should not expose such low-level timeouts to the configuration (at least not until it is really required to do so).. \nAlso, in this testcase, it seems reasonable to not have the reducer at all.","from":"developer"},{"body":"Attaching a patch that \n- reduced max backoff so that the retries happen frequently\n- cleared the _known-outputs_ structure so that the reducer doesnt consider events not known to jobtracker\n \nTested on my box and works fine. Not tested it throughly though. Somethings that would also help : \n- decrease the num-maps from 50 to 30. This will reduce the test's runt time.","from":"developer"},{"body":"Marking as a blocker, as multiple patches are failing on Hudson on this.","from":"developer"},{"body":"Looks like HADOOP-4683 solves this issue for now. The way HADOOP-4683 solves this issue is that it fetches all the events in one go and latest events are fetched faster. So the fact that a map location is faulty/missing is immediately passed onto the reducer which adds the map output location to the _ignore list_. I couldn't reproduce this bug with trunk. Also HADOOP-4220 tries to reduce the run-time of this testcase. ","from":"developer"},{"body":"Looks like {{TestJobTrackerRestartWithLostTracker}} can still fail. Consider a map that is \n- hosted by a tracker that will be lost \n- the map completion event is not logged to the job history (i.e its in the buffer)\n\nCall it a _hanging-map_. Once the reducer reaches the _hanging-map_, it will be stuck there forever as the map location is not known to the jobtracker and hence wont be added to the _ignore-list_. The previous fix solves the issue. What it does is :\n- clears the old/stale mapping of map output locations\n- reduces the backoff time so that the reducer doesnt back off for long\n\nUpdates the patch to trunk.\n","from":"developer"},{"body":"You missed one check for null for the returned list from mapLocations. That should be added..","from":"developer"},{"body":"Incorporated Devaraj's comments. Result of _test-patch_ \n{noformat}\n[exec] +1 overall. \n [exec] \n [exec] +1 @author. The patch does not contain any @author tags.\n [exec] \n [exec] +1 tests included. The patch appears to include 3 new or modified tests.\n [exec] \n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec] \n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec] \n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec] \n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n{noformat}","from":"developer"},{"body":"if knownOutputsByLoc is null, we should break out of the loop instead of continuing","from":"developer"},{"body":"In this patch we make no implicit assumption. Let it run in the loop. It should not be a hit as retry-fetches and num-hosts are not expected to be much. Added comments.","from":"developer"},{"body":"I just committed this. Thanks, Amar!","from":"developer"}],"created":"2008-11-24T10:36:58.000+0000","description":"This test frequently times out: org.apache.hadoop.mapred.TestJobTrackerRestartWithLostTracker.testRestartWithLostTracker\nExample: http://hudson.zones.apache.org/hudson/job/Hadoop-Patch/3637/testReport/org.apache.hadoop.mapred/TestJobTrackerRestartWithLostTracker/testRestartWithLostTracker/","issue_id":"12409114","key":"HADOOP-4716","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-12-22T04:42:18.000+0000","role":"fixed_distractor","summary":"testRestartWithLostTracker frequently times out"} {"case_id":"12409472","cluster":"DISTRACTOR-HADOOP-4742","comments":[{"body":"Yes, I think this is indeed a problem. The proposed solution should be able to fix the problem.","created":"2008-12-01T19:12:15.613+0000"},{"body":"Thanks Wang for your contribution. I redid the patch against the trunk.","created":"2008-12-03T23:54:13.528+0000"},{"body":"This is the patch for branch 0.18.","created":"2008-12-03T23:58:13.452+0000"},{"body":"ant test-core succeded:\nBUILD SUCCESSFUL\nTotal time: 115 minutes 14 seconds\n\nant test-patch result:\n [exec] -1 overall.\n\n [exec] +1 @author. The patch does not contain any @author tags.\n\n [exec] -1 tests included. The patch doesn't appear to include any new or modified tes\nts.\n [exec] Please justify why no tests are needed for this patch.\n\n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n\n [exec] +1 javac. The applied patch does not increase the total number of javac compil\ner warnings.\n\n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n\n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n","created":"2008-12-05T23:16:37.722+0000"},{"body":"I've just committed this. Thanks you, Wang!","created":"2008-12-05T23:43:01.160+0000"},{"body":"Thanks Hairong! I learned much the issue progress of hadoop here. And I think I will do more next time. :)","created":"2008-12-05T23:48:57.651+0000"}],"conversations":[{"body":"We recently deployed a 0.18.1 cluster and did some test. And we found\nif we corrupt a block, the namenode will find it and replicate it as soon as\na client read that block. However, the namenode will delete a health block\n(the source of the above replication operation) at the same time, (I think this\nissue may affect all 0.18 tree.)\n\nHaving did some trace, I find in FSNamesystem.addStoredBlock(), it will\ncheck the number of replications after add the block to blocksMap:\n\n | NumberReplicas num = countNodes(storedBlock);\n | int numLiveReplicas = num.liveReplicas();\n | int numCurrentReplica = numLiveReplicas\n | + pendingReplications.getNumReplicas(block);\n\nwhich means all the live replicas and pending replications will be\ncounted. But in the end of FSNamesystem.blockReceived(), which\ncalls the addStoredBlock(), it will call addStoredBlock() first, then\nreduce the pendingReplications count.\n\n | //\n | // Modify the blocks->datanode map and node's map.\n | //\n | addStoredBlock(block, node, delHintNode );\n | pendingReplications.remove(block);\n\nHence, the newly replicated replica will be counted twice, and then\nwill be marked as excess and lead to a mistake deletion.\n\nI think change the counting lines in blockReceived(), may solve this\nissue:\n\n--- FSNamesystem.java-orig 2008-11-28 13:34:40.000000000 +0800\n+++ FSNamesystem.java 2008-11-28 13:54:12.000000000 +0800\n@@ -3152,8 +3152,8 @@\n //\n // Modify the blocks->datanode map and node's map.\n //\n- addStoredBlock(block, node, delHintNode );\n pendingReplications.remove(block);\n+ addStoredBlock(block, node, delHintNode );\n }\n\n long[] getStats() throws IOException {\n\nThe following is the logs for the mistake deletion, with additional\nlogging info inserted by me.\n\n2008-11-28 11:22:08,866 INFO org.apache.hadoop.dfs.StateChange: *DIR*\nNameNode.reportBadBlocks\n2008-11-28 11:22:08,866 INFO org.apache.hadoop.dfs.StateChange: BLOCK\nNameSystem.addToCorruptReplicasMap: blk_3828935579548953768 added as\ncorrupt on 192.168.33.51:50010 by /192.168.33.51\n2008-11-28 11:22:10,179 INFO org.apache.hadoop.dfs.StateChange: BLOCK*\nask 192.168.33.50:50010 to replicate blk_3828935579548953768_1184 to\ndatanode(s) 192.168.33.45:50010\n2008-11-28 11:22:12,629 INFO org.apache.hadoop.dfs.StateChange: BLOCK*\nNameSystem.addStoredBlock: blockMap updated: 192.168.33.45:50010 is\nadded to blk_3828935579548953768_1184 size 67108864\n2008-11-28 11:22:12,629 INFO org.apache.hadoop.dfs.StateChange: Wang\nXu* NameSystem.addStoredBlock: current replicas 4 in which has 1\npendings\n2008-11-28 11:22:12,630 INFO org.apache.hadoop.dfs.StateChange: DIR*\nNameSystem.invalidateBlock: blk_3828935579548953768_1184 on\n192.168.33.51:50010\n2008-11-28 11:22:12,630 INFO org.apache.hadoop.dfs.StateChange: BLOCK*\nNameSystem.delete: blk_3828935579548953768 is added to invalidSet of\n192.168.33.51:50010\n2008-11-28 11:22:13,180 INFO org.apache.hadoop.dfs.StateChange: BLOCK*\nask 192.168.33.44:50010 to delete blk_3828935579548953768_1184\n2008-11-28 11:22:13,181 INFO org.apache.hadoop.dfs.StateChange: BLOCK*\nask 192.168.33.51:50010 to delete blk_3828935579548953768_1184\n\n","from":"reporter","subject":"Mistake delete replica in hadoop 0.18.1"},{"body":"Yes, I think this is indeed a problem. The proposed solution should be able to fix the problem.","from":"developer"},{"body":"Thanks Wang for your contribution. I redid the patch against the trunk.","from":"developer"},{"body":"This is the patch for branch 0.18.","from":"developer"},{"body":"ant test-core succeded:\nBUILD SUCCESSFUL\nTotal time: 115 minutes 14 seconds\n\nant test-patch result:\n [exec] -1 overall.\n\n [exec] +1 @author. The patch does not contain any @author tags.\n\n [exec] -1 tests included. The patch doesn't appear to include any new or modified tes\nts.\n [exec] Please justify why no tests are needed for this patch.\n\n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n\n [exec] +1 javac. The applied patch does not increase the total number of javac compil\ner warnings.\n\n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n\n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n","from":"developer"},{"body":"I've just committed this. Thanks you, Wang!","from":"developer"},{"body":"Thanks Hairong! I learned much the issue progress of hadoop here. And I think I will do more next time. :)","from":"developer"}],"created":"2008-11-29T07:49:38.000+0000","description":"We recently deployed a 0.18.1 cluster and did some test. And we found\nif we corrupt a block, the namenode will find it and replicate it as soon as\na client read that block. However, the namenode will delete a health block\n(the source of the above replication operation) at the same time, (I think this\nissue may affect all 0.18 tree.)\n\nHaving did some trace, I find in FSNamesystem.addStoredBlock(), it will\ncheck the number of replications after add the block to blocksMap:\n\n | NumberReplicas num = countNodes(storedBlock);\n | int numLiveReplicas = num.liveReplicas();\n | int numCurrentReplica = numLiveReplicas\n | + pendingReplications.getNumReplicas(block);\n\nwhich means all the live replicas and pending replications will be\ncounted. But in the end of FSNamesystem.blockReceived(), which\ncalls the addStoredBlock(), it will call addStoredBlock() first, then\nreduce the pendingReplications count.\n\n | //\n | // Modify the blocks->datanode map and node's map.\n | //\n | addStoredBlock(block, node, delHintNode );\n | pendingReplications.remove(block);\n\nHence, the newly replicated replica will be counted twice, and then\nwill be marked as excess and lead to a mistake deletion.\n\nI think change the counting lines in blockReceived(), may solve this\nissue:\n\n--- FSNamesystem.java-orig 2008-11-28 13:34:40.000000000 +0800\n+++ FSNamesystem.java 2008-11-28 13:54:12.000000000 +0800\n@@ -3152,8 +3152,8 @@\n //\n // Modify the blocks->datanode map and node's map.\n //\n- addStoredBlock(block, node, delHintNode );\n pendingReplications.remove(block);\n+ addStoredBlock(block, node, delHintNode );\n }\n\n long[] getStats() throws IOException {\n\nThe following is the logs for the mistake deletion, with additional\nlogging info inserted by me.\n\n2008-11-28 11:22:08,866 INFO org.apache.hadoop.dfs.StateChange: *DIR*\nNameNode.reportBadBlocks\n2008-11-28 11:22:08,866 INFO org.apache.hadoop.dfs.StateChange: BLOCK\nNameSystem.addToCorruptReplicasMap: blk_3828935579548953768 added as\ncorrupt on 192.168.33.51:50010 by /192.168.33.51\n2008-11-28 11:22:10,179 INFO org.apache.hadoop.dfs.StateChange: BLOCK*\nask 192.168.33.50:50010 to replicate blk_3828935579548953768_1184 to\ndatanode(s) 192.168.33.45:50010\n2008-11-28 11:22:12,629 INFO org.apache.hadoop.dfs.StateChange: BLOCK*\nNameSystem.addStoredBlock: blockMap updated: 192.168.33.45:50010 is\nadded to blk_3828935579548953768_1184 size 67108864\n2008-11-28 11:22:12,629 INFO org.apache.hadoop.dfs.StateChange: Wang\nXu* NameSystem.addStoredBlock: current replicas 4 in which has 1\npendings\n2008-11-28 11:22:12,630 INFO org.apache.hadoop.dfs.StateChange: DIR*\nNameSystem.invalidateBlock: blk_3828935579548953768_1184 on\n192.168.33.51:50010\n2008-11-28 11:22:12,630 INFO org.apache.hadoop.dfs.StateChange: BLOCK*\nNameSystem.delete: blk_3828935579548953768 is added to invalidSet of\n192.168.33.51:50010\n2008-11-28 11:22:13,180 INFO org.apache.hadoop.dfs.StateChange: BLOCK*\nask 192.168.33.44:50010 to delete blk_3828935579548953768_1184\n2008-11-28 11:22:13,181 INFO org.apache.hadoop.dfs.StateChange: BLOCK*\nask 192.168.33.51:50010 to delete blk_3828935579548953768_1184\n\n","issue_id":"12409472","key":"HADOOP-4742","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2008-12-05T23:43:01.000+0000","role":"fixed_distractor","summary":"Mistake delete replica in hadoop 0.18.1"} {"case_id":"12410921","cluster":"DISTRACTOR-HADOOP-4906","comments":[{"body":"A user complained about this on core-user@ supplying the following trace:\n{noformat}\njava.lang.OutOfMemoryError: Java heap space \nat java.util.Arrays.copyOf(Arrays.java:2882) \nat java.lang.AbstractStringBuilder.expandCapacity(AbstractStringBuilder.java:100)\nat java.lang.AbstractStringBuilder.append(AbstractStringBuilder.java:390) \nat java.lang.StringBuffer.append(StringBuffer.java:224) \nat com.sun.org.apache.xerces.internal.dom.DeferredDocumentImpl.getNodeValueString(DeferredDocumentImpl.java:1167)\nat com.sun.org.apache.xerces.internal.dom.DeferredDocumentImpl.getNodeValueString(DeferredDocumentImpl.java:1120)\nat com.sun.org.apache.xerces.internal.dom.DeferredTextImpl.synchronizeData(DeferredTextImpl.java:93)\nat com.sun.org.apache.xerces.internal.dom.CharacterDataImpl.getData(CharacterDataImpl.java:160)\nat org.apache.hadoop.conf.Configuration.loadResource(Configuration.java:928)\nat org.apache.hadoop.conf.Configuration.loadResources(Configuration.java:851)\nat org.apache.hadoop.conf.Configuration.getProps(Configuration.java:819) \nat org.apache.hadoop.conf.Configuration.get(Configuration.java:278) \nat org.apache.hadoop.conf.Configuration.getBoolean(Configuration.java:446) \nat org.apache.hadoop.mapred.JobConf.getKeepFailedTaskFiles(JobConf.java:308) \nat org.apache.hadoop.mapred.TaskTracker$TaskInProgress.setJobConf(TaskTracker.java:1506)\nat org.apache.hadoop.mapred.TaskTracker.launchTaskForJob(TaskTracker.java:727)\nat org.apache.hadoop.mapred.TaskTracker.localizeJob(TaskTracker.java:721) \nat org.apache.hadoop.mapred.TaskTracker.startNewTask(TaskTracker.java:1306) \nat org.apache.hadoop.mapred.TaskTracker.offerService(TaskTracker.java:946) \nat org.apache.hadoop.mapred.TaskTracker.run(TaskTracker.java:1343) \nat org.apache.hadoop.mapred.TaskTracker.main(TaskTracker.java:2354)\n{noformat}","created":"2008-12-17T19:17:48.739+0000"},{"body":"I have seen this happen too. In one case the input path entry in our JobConf used up quite a bit of of ram since we were trying to merge a lot of small files into bigger ones. The JobConf for each task was kept in ram until the Job completed, but since we were running a lot of tasks the TaskTracker ran out of ram before the job could finish.","created":"2008-12-18T10:18:31.319+0000"},{"body":"Yes this is a bug. References to completed/failed TaskInProgress objects from a datastructure (TaskTracker.tasks Map datastructure) are never cleared. The TaskInProgress object also has a JobConf reference... ","created":"2008-12-18T17:29:44.806+0000"},{"body":"TaskInProgress object references are held by Map *tasks* and Map *runningJobs*. These are not cleared until the job is complete.\nThese are not really required to be hanging around till the job completion. As soon as the task finishes, these can be cleared.","created":"2008-12-23T06:54:08.267+0000"},{"body":"By reducing the TaskTracker heap size and running SleepJob with high number of map tasks I could reproduce the OOM.\nWith the attached patch, TaskInProgress objects are being garbage collected and I am not seeing the OOM with the same settings.","created":"2008-12-23T13:01:49.270+0000"},{"body":"bq. These are not really required to be hanging around till the job completion. As soon as the task finishes, these can be cleared. \n\nUnfortunately this isn't true. We need the TaskInProgress object in TaskTracker.tasks at least for failing the TaskAttempt of maps whose map-outputs are lost... see TaskTracker.mapOutputLost. I guess we need to just set defaultJobConf/localJobConf to null in TaskInProgress.cleanup for now? Sigh!\n\n----\n\nUnrelated rant: My head still hurts from having to track this down in the mess that the TaskTracker has evolved into... I guess we need to seriously start thinking of cleaning up the TaskTracker, moving TaskTracker.TaskInProgress (very bad name too!) to it's own file, simplifying the interaction between the child TaskAttempt and the TaskTracker etc. Thoughts? ","created":"2008-12-23T16:04:14.213+0000"},{"body":"bq. I guess we need to just set defaultJobConf/localJobConf to null in TaskInProgress.cleanup for now\n+1. I would say that we also nullify the TIP.diagnosticInfo since this is user settable too. Since a TIP has no other user settable data, and all TIPs of a particular job are removed when the latter completes, it should be okay to keep the TIPs in memory..","created":"2008-12-24T03:42:59.889+0000"},{"body":"sigh! just setting TaskInProgress.localJobConf to null doesn't seem to make JobConf object getting garbage collected. TaskInProgress.localJobConf references are being indirectly held by other classes -> MapOutputFile, MapTask and MapTaskRunner. so we need to set jobConf in these as well to null. ","created":"2008-12-24T11:04:49.363+0000"},{"body":"this patch sets the TaskInProgress#localJobConf object references to null.","created":"2008-12-29T14:06:31.373+0000"},{"body":"Since this problem affects the 0.19 branch, can we get this fix into the 0.19 branch as well? Thanks.","created":"2008-12-29T16:11:08.572+0000"},{"body":"Setting the JobConf references to null in TaskInProgress#cleanup had problem since the TaskRunner thread may need the ref later. Had an offline discussion with Devaraj and decided that TaskRunner#run would be the better place to clean the references.\nThis patch makes the JobConf references to null in the finally block of TaskRunner#run. Also had to keep some data in member variables of MapOutputFile and TaskInProgress so that JobConf reference is not required when task finishes.","created":"2009-01-05T11:26:30.839+0000"},{"body":"It struck to me that there might be a better approach to the one we have taken, since the present one is hard to maintain. \nCurrently there is a new JobConf object created for each task via new JobConf(rjob.jobFile) in TaskTracker#localizeJob. Since this object is created from a file, new references are created for all properties. Instead of that what if we create JobConf as below:\n\n{code}\n//keep a reference to JobConf in RunningJob\nrjob.localized = true;\nrjob.jobConf = localJobConf;\n{code}\n\n{code}\n//create the task level JobConf\nlaunchTaskForJob(tip, new JobConf(rjob.jobConf)); \n{code}\n\nThis way the properties for task's conf will be shallow cloned resulting in only reference copy instead of full new object creation. Also there would be a side benefit as there is no need to parse xml file and create JobConf for each task.\nThoughts ?","created":"2009-01-07T12:21:39.162+0000"},{"body":"Being little clear on my last comment. The code could look like this in TaskTracker.java:\n\n{code}\nprivate void localizeJob(TaskInProgress tip) throws IOException {\n.......\n.......\nsynchronized (rjob) {\n if (!rjob.localized) {\n ........\n ........\n //keep a reference to JobConf in RunningJob\n rjob.localized = true;\n rjob.jobConf = localJobConf;\n }\n}\n //create the task level JobConf \n launchTaskForJob(tip, new JobConf(rjob.jobConf)); \n}\n{code}","created":"2009-01-07T12:40:28.191+0000"},{"body":"This seems like a good interim solution. +1\n\nI'd be happy to commit this if we could run a large sleep-job and confirm the fix. Thanks!","created":"2009-01-13T22:30:58.039+0000"},{"body":"Ran a SleepJob by putting a large dummy property value (20MB) in JobConf. \n* Without the patch, the OOM came after running 44 map tasks. This is expected as the default TT heap size is 1GB so the OOM should come within 50 tasks (20MB *50) executions.\n* With the patch, the no of tasks ran successfully past 1000, then I stopped it. During the run, monitored the heap size which remains almost steady.","created":"2009-01-13T23:03:03.593+0000"},{"body":"does this bug exist in 0.19 as well? If so, can we get it into 0.19 branch too?","created":"2009-01-13T23:59:04.235+0000"},{"body":"bq. does this bug exist in 0.19 as well? If so, can we get it into 0.19 branch too? \nyes. Will upload the patch for 0.19 shortly.","created":"2009-01-14T00:23:00.920+0000"},{"body":"all tests passed. ant test-patch :\n-1 overall.\n [exec]\n [exec] +1 @author. The patch does not contain any @author tags.\n [exec]\n [exec] -1 tests included. The patch doesn't appear to include any new or modified tests.\n [exec] Please justify why no tests are needed for this patch.\n [exec]\n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec]\n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec]\n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec]\n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\nIt is not easy to write a test case for this.\n","created":"2009-01-14T20:55:13.537+0000"},{"body":"same patch applies to 0.19 branch as well.","created":"2009-01-14T22:21:15.549+0000"},{"body":"Minor nit - we should remove TaskTracker.RunningJob.jobFile since it isn't used anymore...","created":"2009-01-15T00:47:55.765+0000"},{"body":"incorporated Arun's comment","created":"2009-01-15T01:32:38.492+0000"},{"body":"I just committed this. Thanks, Sharad!","created":"2009-01-16T00:35:13.081+0000"}],"conversations":[{"body":"Looks like the TaskTracker isn't cleaning up correctly after completed/failed tasks, I suspect that the JobConfs aren't being deallocated. Eventually the TaskTracker runs out of memory after running several tasks.","from":"reporter","subject":"TaskTracker running out of memory after running several tasks"},{"body":"A user complained about this on core-user@ supplying the following trace:\n{noformat}\njava.lang.OutOfMemoryError: Java heap space \nat java.util.Arrays.copyOf(Arrays.java:2882) \nat java.lang.AbstractStringBuilder.expandCapacity(AbstractStringBuilder.java:100)\nat java.lang.AbstractStringBuilder.append(AbstractStringBuilder.java:390) \nat java.lang.StringBuffer.append(StringBuffer.java:224) \nat com.sun.org.apache.xerces.internal.dom.DeferredDocumentImpl.getNodeValueString(DeferredDocumentImpl.java:1167)\nat com.sun.org.apache.xerces.internal.dom.DeferredDocumentImpl.getNodeValueString(DeferredDocumentImpl.java:1120)\nat com.sun.org.apache.xerces.internal.dom.DeferredTextImpl.synchronizeData(DeferredTextImpl.java:93)\nat com.sun.org.apache.xerces.internal.dom.CharacterDataImpl.getData(CharacterDataImpl.java:160)\nat org.apache.hadoop.conf.Configuration.loadResource(Configuration.java:928)\nat org.apache.hadoop.conf.Configuration.loadResources(Configuration.java:851)\nat org.apache.hadoop.conf.Configuration.getProps(Configuration.java:819) \nat org.apache.hadoop.conf.Configuration.get(Configuration.java:278) \nat org.apache.hadoop.conf.Configuration.getBoolean(Configuration.java:446) \nat org.apache.hadoop.mapred.JobConf.getKeepFailedTaskFiles(JobConf.java:308) \nat org.apache.hadoop.mapred.TaskTracker$TaskInProgress.setJobConf(TaskTracker.java:1506)\nat org.apache.hadoop.mapred.TaskTracker.launchTaskForJob(TaskTracker.java:727)\nat org.apache.hadoop.mapred.TaskTracker.localizeJob(TaskTracker.java:721) \nat org.apache.hadoop.mapred.TaskTracker.startNewTask(TaskTracker.java:1306) \nat org.apache.hadoop.mapred.TaskTracker.offerService(TaskTracker.java:946) \nat org.apache.hadoop.mapred.TaskTracker.run(TaskTracker.java:1343) \nat org.apache.hadoop.mapred.TaskTracker.main(TaskTracker.java:2354)\n{noformat}","from":"developer"},{"body":"I have seen this happen too. In one case the input path entry in our JobConf used up quite a bit of of ram since we were trying to merge a lot of small files into bigger ones. The JobConf for each task was kept in ram until the Job completed, but since we were running a lot of tasks the TaskTracker ran out of ram before the job could finish.","from":"developer"},{"body":"Yes this is a bug. References to completed/failed TaskInProgress objects from a datastructure (TaskTracker.tasks Map datastructure) are never cleared. The TaskInProgress object also has a JobConf reference... ","from":"developer"},{"body":"TaskInProgress object references are held by Map *tasks* and Map *runningJobs*. These are not cleared until the job is complete.\nThese are not really required to be hanging around till the job completion. As soon as the task finishes, these can be cleared.","from":"developer"},{"body":"By reducing the TaskTracker heap size and running SleepJob with high number of map tasks I could reproduce the OOM.\nWith the attached patch, TaskInProgress objects are being garbage collected and I am not seeing the OOM with the same settings.","from":"developer"},{"body":"bq. These are not really required to be hanging around till the job completion. As soon as the task finishes, these can be cleared. \n\nUnfortunately this isn't true. We need the TaskInProgress object in TaskTracker.tasks at least for failing the TaskAttempt of maps whose map-outputs are lost... see TaskTracker.mapOutputLost. I guess we need to just set defaultJobConf/localJobConf to null in TaskInProgress.cleanup for now? Sigh!\n\n----\n\nUnrelated rant: My head still hurts from having to track this down in the mess that the TaskTracker has evolved into... I guess we need to seriously start thinking of cleaning up the TaskTracker, moving TaskTracker.TaskInProgress (very bad name too!) to it's own file, simplifying the interaction between the child TaskAttempt and the TaskTracker etc. Thoughts? ","from":"developer"},{"body":"bq. I guess we need to just set defaultJobConf/localJobConf to null in TaskInProgress.cleanup for now\n+1. I would say that we also nullify the TIP.diagnosticInfo since this is user settable too. Since a TIP has no other user settable data, and all TIPs of a particular job are removed when the latter completes, it should be okay to keep the TIPs in memory..","from":"developer"},{"body":"sigh! just setting TaskInProgress.localJobConf to null doesn't seem to make JobConf object getting garbage collected. TaskInProgress.localJobConf references are being indirectly held by other classes -> MapOutputFile, MapTask and MapTaskRunner. so we need to set jobConf in these as well to null. ","from":"developer"},{"body":"this patch sets the TaskInProgress#localJobConf object references to null.","from":"developer"},{"body":"Since this problem affects the 0.19 branch, can we get this fix into the 0.19 branch as well? Thanks.","from":"developer"},{"body":"Setting the JobConf references to null in TaskInProgress#cleanup had problem since the TaskRunner thread may need the ref later. Had an offline discussion with Devaraj and decided that TaskRunner#run would be the better place to clean the references.\nThis patch makes the JobConf references to null in the finally block of TaskRunner#run. Also had to keep some data in member variables of MapOutputFile and TaskInProgress so that JobConf reference is not required when task finishes.","from":"developer"},{"body":"It struck to me that there might be a better approach to the one we have taken, since the present one is hard to maintain. \nCurrently there is a new JobConf object created for each task via new JobConf(rjob.jobFile) in TaskTracker#localizeJob. Since this object is created from a file, new references are created for all properties. Instead of that what if we create JobConf as below:\n\n{code}\n//keep a reference to JobConf in RunningJob\nrjob.localized = true;\nrjob.jobConf = localJobConf;\n{code}\n\n{code}\n//create the task level JobConf\nlaunchTaskForJob(tip, new JobConf(rjob.jobConf)); \n{code}\n\nThis way the properties for task's conf will be shallow cloned resulting in only reference copy instead of full new object creation. Also there would be a side benefit as there is no need to parse xml file and create JobConf for each task.\nThoughts ?","from":"developer"},{"body":"Being little clear on my last comment. The code could look like this in TaskTracker.java:\n\n{code}\nprivate void localizeJob(TaskInProgress tip) throws IOException {\n.......\n.......\nsynchronized (rjob) {\n if (!rjob.localized) {\n ........\n ........\n //keep a reference to JobConf in RunningJob\n rjob.localized = true;\n rjob.jobConf = localJobConf;\n }\n}\n //create the task level JobConf \n launchTaskForJob(tip, new JobConf(rjob.jobConf)); \n}\n{code}","from":"developer"},{"body":"This seems like a good interim solution. +1\n\nI'd be happy to commit this if we could run a large sleep-job and confirm the fix. Thanks!","from":"developer"},{"body":"Ran a SleepJob by putting a large dummy property value (20MB) in JobConf. \n* Without the patch, the OOM came after running 44 map tasks. This is expected as the default TT heap size is 1GB so the OOM should come within 50 tasks (20MB *50) executions.\n* With the patch, the no of tasks ran successfully past 1000, then I stopped it. During the run, monitored the heap size which remains almost steady.","from":"developer"},{"body":"does this bug exist in 0.19 as well? If so, can we get it into 0.19 branch too?","from":"developer"},{"body":"bq. does this bug exist in 0.19 as well? If so, can we get it into 0.19 branch too? \nyes. Will upload the patch for 0.19 shortly.","from":"developer"},{"body":"all tests passed. ant test-patch :\n-1 overall.\n [exec]\n [exec] +1 @author. The patch does not contain any @author tags.\n [exec]\n [exec] -1 tests included. The patch doesn't appear to include any new or modified tests.\n [exec] Please justify why no tests are needed for this patch.\n [exec]\n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec]\n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec]\n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec]\n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\nIt is not easy to write a test case for this.\n","from":"developer"},{"body":"same patch applies to 0.19 branch as well.","from":"developer"},{"body":"Minor nit - we should remove TaskTracker.RunningJob.jobFile since it isn't used anymore...","from":"developer"},{"body":"incorporated Arun's comment","from":"developer"},{"body":"I just committed this. Thanks, Sharad!","from":"developer"}],"created":"2008-12-17T19:15:37.000+0000","description":"Looks like the TaskTracker isn't cleaning up correctly after completed/failed tasks, I suspect that the JobConfs aren't being deallocated. Eventually the TaskTracker runs out of memory after running several tasks.","issue_id":"12410921","key":"HADOOP-4906","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-01-16T00:35:13.000+0000","role":"fixed_distractor","summary":"TaskTracker running out of memory after running several tasks"} {"case_id":"12348901","cluster":"DISTRACTOR-HADOOP-493","comments":[{"body":"\nClusters don't seem to show this problem anymore. DFS still reports wrong value for 'used bytes' (see HADOOP-599), but it is not weird value like above. It always shows % used is zero. Will close this for now and will reopen if HADOOP-599 does not fix this.\n\n\n","created":"2006-10-23T20:35:31.000+0000"},{"body":"\neither it is already fixed or HADOOP-599 might fix it.","created":"2006-10-23T20:36:44.000+0000"}],"conversations":[{"body":"The information presented by hadoop dfs -report is incorrect.\nIn the following example, the total raw bytes are correct, the used raw bytes are incorrect and should actually be on the order of 138TB.\nThe total effective bytes is correct, and the actual replication factor is really a healthy 3.\n\nyarnon@g1010:build:942>hadoop dfs 2> /dev/null -report | head\nTotal raw bytes: 240761579888640 (224226.69 Gb)\nUsed raw bytes: 35934229005360 (33466.35 Gb)\n% used: 14.92%\n\nTotal effective bytes: 47093545108964 (43859.28 Gb)\nEffective replication multiplier: 0.7630393703047027\n-------------------------------------------------\n\n\nIn the following example again the total raw bytes and total effective bytes are correct, but the used raw bytes, % used and effective replication are a tad off the mark.\n\nhadoop dfs 2> /dev/null -report | head\nTotal raw bytes: 236476179599360 (220235.60 Gb)\nUsed raw bytes: -22028236104431 (-2.15 k)\n% used: -9.31%\n\nTotal effective bytes: 55609049637201 (51789.96 Gb)\nEffective replication multiplier: -0.39612682194976206\n-------------------------------------------------\n","from":"reporter","subject":"hadoop dfs -report information is incorrect"},{"body":"\nClusters don't seem to show this problem anymore. DFS still reports wrong value for 'used bytes' (see HADOOP-599), but it is not weird value like above. It always shows % used is zero. Will close this for now and will reopen if HADOOP-599 does not fix this.\n\n\n","from":"developer"},{"body":"\neither it is already fixed or HADOOP-599 might fix it.","from":"developer"}],"created":"2006-08-30T06:16:49.000+0000","description":"The information presented by hadoop dfs -report is incorrect.\nIn the following example, the total raw bytes are correct, the used raw bytes are incorrect and should actually be on the order of 138TB.\nThe total effective bytes is correct, and the actual replication factor is really a healthy 3.\n\nyarnon@g1010:build:942>hadoop dfs 2> /dev/null -report | head\nTotal raw bytes: 240761579888640 (224226.69 Gb)\nUsed raw bytes: 35934229005360 (33466.35 Gb)\n% used: 14.92%\n\nTotal effective bytes: 47093545108964 (43859.28 Gb)\nEffective replication multiplier: 0.7630393703047027\n-------------------------------------------------\n\n\nIn the following example again the total raw bytes and total effective bytes are correct, but the used raw bytes, % used and effective replication are a tad off the mark.\n\nhadoop dfs 2> /dev/null -report | head\nTotal raw bytes: 236476179599360 (220235.60 Gb)\nUsed raw bytes: -22028236104431 (-2.15 k)\n% used: -9.31%\n\nTotal effective bytes: 55609049637201 (51789.96 Gb)\nEffective replication multiplier: -0.39612682194976206\n-------------------------------------------------\n","issue_id":"12348901","key":"HADOOP-493","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-10-23T20:36:44.000+0000","role":"fixed_distractor","summary":"hadoop dfs -report information is incorrect"} {"case_id":"12329291","cluster":"DISTRACTOR-HADOOP-50","comments":[{"body":"I think this is a valid concern. Most filesystems work poorly with thousands of files in a single directory. My recent tests on ext3 show that listing the data directory with 50,000 blocks takes several seconds.\n\nFSDataset:80 contains a commented out section, which seems to address this issue. Anyone knows why it's not used?","created":"2006-03-26T04:04:07.000+0000"},{"body":"Hi Andrzej,\n\nI wrote this code and got it 90% working some time ago, but then had to abandon\nit for a more important bug. It is not ready to go in its current state, but shouldn't\nbe too hard. I can bring this code back to life..\n\n--Mike","created":"2006-03-29T02:04:45.000+0000"},{"body":"That would be very useful. I've seen similar solutions in many places (e.g. squid, or Mozilla cache dir).\n\nCurrently, each time a block report is sent we need to list this huge dir. That's still ok, it's infrequent enough. However, each time we need to access a block, a correct file needs to be open. Inside the native code JVM uses an open(2) call, which causes the OS to perform a name-to-inode lookup. Even though OS is caching partial results of this lookup (in Linux this is known as dcache/dentries), still depending on the size of this LRU cache and the FS implementation details, doing real lookups for e.g. new blocks or newly requested blocks may take a long time.\n\nHaving said that, I'm not sure what would be the real performance benefit of this change, perhaps you could come up with a simpler test first...?","created":"2006-03-29T18:21:45.000+0000"},{"body":"\nThis fixes the multiple-directory storage problem. It\nlazily creates a single level of 512 subdirectories, into which\nthe blocks are allocated according to the lower 9 bits of the\nblock id. If mankind ever needs more blocks than this, it is easy\nto add an additional subdir layer and select on the lowest-but-9\nbits of the blockid.\n\nThis change is backwards-compatible with the previous block\nlayout. Old blocks in the single-layer dir will always be kept in\nthat format; we don't migrate them. New blocks will always be added \nto the new hierarchy.\n\nIf both versions of the storage system are present, we always test\nthe new one first. If that fails, we test the old one. (The new test\nshould be faster, so we do it first.)\n\nPlease let me know if this patch works for you.","created":"2006-04-23T12:17:55.000+0000"},{"body":"+1\nI didn't look at the patch, but I carefully read the commented out code. I vote yes. (is it already in? - close this issue).","created":"2006-05-26T08:36:29.000+0000"},{"body":"This was done as part of HADOOP-64","created":"2006-09-07T08:03:09.000+0000"}],"conversations":[{"body":"The datanode currently stores all file blocks in a single directory. With 32MB blocks and terabyte filesystems, this will create too many files in a single directory for many filesystems. Thus blocks should be stored in multiple directories, perhaps even a shallow hierarchy.","from":"reporter","subject":"dfs datanode should store blocks in multiple directories"},{"body":"I think this is a valid concern. Most filesystems work poorly with thousands of files in a single directory. My recent tests on ext3 show that listing the data directory with 50,000 blocks takes several seconds.\n\nFSDataset:80 contains a commented out section, which seems to address this issue. Anyone knows why it's not used?","from":"developer"},{"body":"Hi Andrzej,\n\nI wrote this code and got it 90% working some time ago, but then had to abandon\nit for a more important bug. It is not ready to go in its current state, but shouldn't\nbe too hard. I can bring this code back to life..\n\n--Mike","from":"developer"},{"body":"That would be very useful. I've seen similar solutions in many places (e.g. squid, or Mozilla cache dir).\n\nCurrently, each time a block report is sent we need to list this huge dir. That's still ok, it's infrequent enough. However, each time we need to access a block, a correct file needs to be open. Inside the native code JVM uses an open(2) call, which causes the OS to perform a name-to-inode lookup. Even though OS is caching partial results of this lookup (in Linux this is known as dcache/dentries), still depending on the size of this LRU cache and the FS implementation details, doing real lookups for e.g. new blocks or newly requested blocks may take a long time.\n\nHaving said that, I'm not sure what would be the real performance benefit of this change, perhaps you could come up with a simpler test first...?","from":"developer"},{"body":"\nThis fixes the multiple-directory storage problem. It\nlazily creates a single level of 512 subdirectories, into which\nthe blocks are allocated according to the lower 9 bits of the\nblock id. If mankind ever needs more blocks than this, it is easy\nto add an additional subdir layer and select on the lowest-but-9\nbits of the blockid.\n\nThis change is backwards-compatible with the previous block\nlayout. Old blocks in the single-layer dir will always be kept in\nthat format; we don't migrate them. New blocks will always be added \nto the new hierarchy.\n\nIf both versions of the storage system are present, we always test\nthe new one first. If that fails, we test the old one. (The new test\nshould be faster, so we do it first.)\n\nPlease let me know if this patch works for you.","from":"developer"},{"body":"+1\nI didn't look at the patch, but I carefully read the commented out code. I vote yes. (is it already in? - close this issue).","from":"developer"},{"body":"This was done as part of HADOOP-64","from":"developer"}],"created":"2006-02-22T07:25:57.000+0000","description":"The datanode currently stores all file blocks in a single directory. With 32MB blocks and terabyte filesystems, this will create too many files in a single directory for many filesystems. Thus blocks should be stored in multiple directories, perhaps even a shallow hierarchy.","issue_id":"12329291","key":"HADOOP-50","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-09-07T08:03:46.000+0000","role":"fixed_distractor","summary":"dfs datanode should store blocks in multiple directories"} {"case_id":"12414453","cluster":"DISTRACTOR-HADOOP-5210","comments":[{"body":"Screen shot showing progress > 100%","created":"2009-02-10T07:37:34.335+0000"},{"body":"This could be because of the way we compute mergeProgress during merges in the reduce. The mergeProgress is a function of the totalBytesProcessed and the totalBytesProcessed is incremented for every segment considered during merge. So if we have multi-level merges, we would run into a case where we report more progress per byte since many bytes would make hit the disk but they would be again considered for the next level merge and so on.. ","created":"2009-02-18T06:11:52.736+0000"},{"body":"I've seen this problem, too. Am happy to send my local config info if that's useful.","created":"2009-03-11T17:43:48.976+0000"},{"body":"As Devaraj mentioned, the problem is in the calculation of mergeProgress when multi-level merges happen.\nAttaching patch that fixes the issue. Please review and provide your comments.","created":"2009-03-12T11:16:20.759+0000"},{"body":"Carrying over bytes from intermediate merges to the final reduce phase is not correct. ","created":"2009-03-13T11:40:51.050+0000"},{"body":"Yes Jothi. When the intermediate merges complete, we can say that the sortPhase is completed and if we reset the variable totalBytesProcessed before the final merge, we can use that for calculating the progress of reducePhase(the 3rd phase of reduce task). Patch of HADOOP-3131 removed this resetting of totalBytesProcessed.\n\nMatei, Would you please check if your patch(of JIRA 3131) removed this reset intentionally and if I am missing out something ?\n\nAttaching patch which resets the bytes-processed to zero before final merge.\nPlease review and provide your comments.","created":"2009-03-16T09:58:09.857+0000"},{"body":"Jothi offline suggested to remove some unnecessary code from merge().\n\nAttaching new patch with that change.","created":"2009-03-19T04:06:17.081+0000"},{"body":"TestReduceTask was failing with earlier patch because of ignoring the starting bytes read from segments in the final merge.\n\nAttaching the patch that resets totalBytesProcessed to the number of bytes read in this final merge(instead of 0).\n\nPlease review and provide your comments.","created":"2009-03-20T04:28:37.548+0000"},{"body":"+1. Patch looks good. ","created":"2009-03-23T09:37:01.291+0000"},{"body":"I just committed this. Thanks, Ravi!","created":"2009-03-25T09:00:15.874+0000"},{"body":"It would be nice if this patch can be committed to branch 0.20 also.\nThe same patch applies to branch 0.20 also.\nDevaraj, Would you please commit this to 0.20 ?","created":"2009-05-15T06:15:48.435+0000"},{"body":"I committed this to the 0.20 branch.","created":"2009-05-20T09:54:24.175+0000"}],"conversations":[{"body":"When the total map outputs size (reduce input size) is high, the reported progress is greater than 100%.","from":"reporter","subject":"Reduce Task Progress shows > 100% when the total size of map outputs (for a single reducer) is high "},{"body":"Screen shot showing progress > 100%","from":"developer"},{"body":"This could be because of the way we compute mergeProgress during merges in the reduce. The mergeProgress is a function of the totalBytesProcessed and the totalBytesProcessed is incremented for every segment considered during merge. So if we have multi-level merges, we would run into a case where we report more progress per byte since many bytes would make hit the disk but they would be again considered for the next level merge and so on.. ","from":"developer"},{"body":"I've seen this problem, too. Am happy to send my local config info if that's useful.","from":"developer"},{"body":"As Devaraj mentioned, the problem is in the calculation of mergeProgress when multi-level merges happen.\nAttaching patch that fixes the issue. Please review and provide your comments.","from":"developer"},{"body":"Carrying over bytes from intermediate merges to the final reduce phase is not correct. ","from":"developer"},{"body":"Yes Jothi. When the intermediate merges complete, we can say that the sortPhase is completed and if we reset the variable totalBytesProcessed before the final merge, we can use that for calculating the progress of reducePhase(the 3rd phase of reduce task). Patch of HADOOP-3131 removed this resetting of totalBytesProcessed.\n\nMatei, Would you please check if your patch(of JIRA 3131) removed this reset intentionally and if I am missing out something ?\n\nAttaching patch which resets the bytes-processed to zero before final merge.\nPlease review and provide your comments.","from":"developer"},{"body":"Jothi offline suggested to remove some unnecessary code from merge().\n\nAttaching new patch with that change.","from":"developer"},{"body":"TestReduceTask was failing with earlier patch because of ignoring the starting bytes read from segments in the final merge.\n\nAttaching the patch that resets totalBytesProcessed to the number of bytes read in this final merge(instead of 0).\n\nPlease review and provide your comments.","from":"developer"},{"body":"+1. Patch looks good. ","from":"developer"},{"body":"I just committed this. Thanks, Ravi!","from":"developer"},{"body":"It would be nice if this patch can be committed to branch 0.20 also.\nThe same patch applies to branch 0.20 also.\nDevaraj, Would you please commit this to 0.20 ?","from":"developer"},{"body":"I committed this to the 0.20 branch.","from":"developer"}],"created":"2009-02-10T07:36:29.000+0000","description":"When the total map outputs size (reduce input size) is high, the reported progress is greater than 100%.","issue_id":"12414453","key":"HADOOP-5210","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-03-25T09:00:15.000+0000","role":"fixed_distractor","summary":"Reduce Task Progress shows > 100% when the total size of map outputs (for a single reducer) is high "} {"case_id":"12415851","cluster":"DISTRACTOR-HADOOP-5367","comments":[{"body":"This could be mostly because of HADOOP-5269 and HADOOP-5235. There are committed to branch 0.19.2. Can you please try out 0.19.2 by checking out branch 0.19 ?","created":"2009-03-02T03:56:55.091+0000"},{"body":"I also meet this such issue.\n\nAfter a long time running of MapReduce (about 200 jobs have completed). The MapReduce job is huaguped forever.\n(1) The JobTracker always logs:\n2009-03-17 16:29:39,997 INFO org.apache.hadoop.mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200903171247_0387_m_000015_1\n\n(2) The Job cannot complete and stopp at 79% forever.\n\n(3) All TaskTrackers may hungup at the sametime, since the logs of each TaskTracer stop at that time.\n\n\nAnd nefore the hangup. I can also find odds and ends such logs of JobTracker, such as.\n2009-03-17 16:29:21,767 INFO org.apache.hadoop.mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200903171247_0387_m_000015_1\n\n\nAnd another experience is:\nOne time, I found the task slot cannot reach the capability, maybe some slot is also hungup.","created":"2009-03-17T15:27:59.838+0000"},{"body":"I am using branch-0.19, it seems fine.","created":"2009-03-27T05:06:16.166+0000"},{"body":"Got fixed in 0.19.2","created":"2009-03-30T06:32:10.982+0000"}],"conversations":[{"body":"Hi,\n\nAfter I while, my cluster will only run the reduce tasks sequentially (each reducer running on the same node), the other nodes stay empty. The map phase however will run the jobs on all the nodes, also after such a \"long\" reduce phase has completed. But the reduce phase will then be again executed sequentially. This happens in my cluster after about 160 successfully completed jobs. (Some jobs have reducer set to 0!). \nAs possible solution I have to restart the mapreduce service.\n\nI didn't notice this behaviour in version 0.19.0. I can't use version 0.19.0 because of the multipleoutput bug when setting reducers to 0.\n\nAnoter site node which might be related. I also tried running the jobs with speculative execution set to on. My cluster would always hold back one reducer and only run it (in multiple instances) after the first of the other 6 reducers had finished, instead of launching all of them at the same time.\n\n\nBelow is a short extract from related logfile. It's full of these kind of entries.\n\n09/02/28 12:48:07 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0051_r_000006_1\n09/02/28 12:48:08 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0041_r_000002_1\n09/02/28 12:48:08 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0083_r_000006_1\n09/02/28 12:48:08 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0041_r_000005_1\n09/02/28 12:48:10 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0105_r_000006_1\n09/02/28 12:48:10 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0102_r_000006_1\n09/02/28 12:48:12 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0051_r_000006_1\n09/02/28 12:48:13 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0041_r_000002_1\n09/02/28 12:48:13 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0083_r_000006_1\n09/02/28 12:48:13 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0041_r_000005_1\n","from":"reporter","subject":"After some jobs have finished, Reducer will run new job's reduce tasks sequentially and not in parallel (mapred.JobTracker: Serious problem. While updating status, cannot find taskid...)"},{"body":"This could be mostly because of HADOOP-5269 and HADOOP-5235. There are committed to branch 0.19.2. Can you please try out 0.19.2 by checking out branch 0.19 ?","from":"developer"},{"body":"I also meet this such issue.\n\nAfter a long time running of MapReduce (about 200 jobs have completed). The MapReduce job is huaguped forever.\n(1) The JobTracker always logs:\n2009-03-17 16:29:39,997 INFO org.apache.hadoop.mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200903171247_0387_m_000015_1\n\n(2) The Job cannot complete and stopp at 79% forever.\n\n(3) All TaskTrackers may hungup at the sametime, since the logs of each TaskTracer stop at that time.\n\n\nAnd nefore the hangup. I can also find odds and ends such logs of JobTracker, such as.\n2009-03-17 16:29:21,767 INFO org.apache.hadoop.mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200903171247_0387_m_000015_1\n\n\nAnd another experience is:\nOne time, I found the task slot cannot reach the capability, maybe some slot is also hungup.","from":"developer"},{"body":"I am using branch-0.19, it seems fine.","from":"developer"},{"body":"Got fixed in 0.19.2","from":"developer"}],"created":"2009-02-28T13:10:01.000+0000","description":"Hi,\n\nAfter I while, my cluster will only run the reduce tasks sequentially (each reducer running on the same node), the other nodes stay empty. The map phase however will run the jobs on all the nodes, also after such a \"long\" reduce phase has completed. But the reduce phase will then be again executed sequentially. This happens in my cluster after about 160 successfully completed jobs. (Some jobs have reducer set to 0!). \nAs possible solution I have to restart the mapreduce service.\n\nI didn't notice this behaviour in version 0.19.0. I can't use version 0.19.0 because of the multipleoutput bug when setting reducers to 0.\n\nAnoter site node which might be related. I also tried running the jobs with speculative execution set to on. My cluster would always hold back one reducer and only run it (in multiple instances) after the first of the other 6 reducers had finished, instead of launching all of them at the same time.\n\n\nBelow is a short extract from related logfile. It's full of these kind of entries.\n\n09/02/28 12:48:07 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0051_r_000006_1\n09/02/28 12:48:08 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0041_r_000002_1\n09/02/28 12:48:08 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0083_r_000006_1\n09/02/28 12:48:08 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0041_r_000005_1\n09/02/28 12:48:10 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0105_r_000006_1\n09/02/28 12:48:10 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0102_r_000006_1\n09/02/28 12:48:12 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0051_r_000006_1\n09/02/28 12:48:13 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0041_r_000002_1\n09/02/28 12:48:13 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0083_r_000006_1\n09/02/28 12:48:13 INFO mapred.JobTracker: Serious problem. While updating status, cannot find taskid attempt_200902271700_0041_r_000005_1\n","issue_id":"12415851","key":"HADOOP-5367","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-03-30T06:32:11.000+0000","role":"fixed_distractor","summary":"After some jobs have finished, Reducer will run new job's reduce tasks sequentially and not in parallel (mapred.JobTracker: Serious problem. While updating status, cannot find taskid...)"} {"case_id":"12417354","cluster":"DISTRACTOR-HADOOP-5539","comments":[{"body":"I'd like to see this fixed as well since one reason I've enabled map output compression is to reduce disk space usage by the mapreduce framework. It appears that currently the map outputs are simply decompressed as soon as they have been downloaded by the reducer.","created":"2009-03-20T18:16:31.853+0000"},{"body":"I verified that it is doing the same on map task no intermediate.x file from o.a.h.mapred.Merger are getting compressed.","created":"2009-03-22T04:06:27.548+0000"},{"body":"this should fix the problem I had to make a few new constructors. I left the old constructors that these files where using because not sure if any other tasks using these. this patch will apply to 0.19-branch I have not worked any on trunk so might need to try dry-run before applying to trunk. tested on my end and working correctly now with this patch.\n","created":"2009-03-22T08:58:35.055+0000"},{"body":"added back to read map","created":"2009-03-24T03:08:38.341+0000"},{"body":"Someone can use my patch as a starting point.\n\nThe ReduceTask.java call that is the problem is line 2145\nThe maptask.java call that is the problem is line 1269\n\nI use streaming without a combiner so that should be looked at also to see if it uses o.a.h.mapred.Merger\n\nthe basic problem is the codec is not passed from these function to the merger so its always null the call to \no.a.h.mapred.Merger should include codec somehow if compression is not used then codec is null \nin both ReduceTask and Maptask.\n\nI thank this is a major bug that effects all MR jobs with disk bandwidth that uses compression.\n","created":"2009-03-24T03:12:20.412+0000"},{"body":"bq. added back to read map \n\nIt's a blocker; it will be resolved and backported to 0.20 at least. The road map isn't; the PA queue defines the set of patches that can be committed. The fix version is usually set when it's actually resolved, so where it was committed is documented.","created":"2009-03-24T04:06:10.008+0000"},{"body":"Oh, I see; the patch is for 0.19. My mistake. ","created":"2009-03-24T05:16:57.957+0000"},{"body":"The patch looks good. A few minor points: \n\n# The new MergeQueue constructor could call the existing constructor and then set the codec later.\n{code}\npublic MergeQueue(Configuration conf, FileSystem fs, \n List> segments, RawComparator comparator,\n Progressable reporter, boolean sortSegments, CompressionCodec codec) {\n this(conf, fs, segments, comparator, reporter, sortSegments);\n this.codec = codec;\n }\n{code}\n\n# For the new merge methods, should we place the Codec argument after the valueClass argument (instead of being the last argument) to maintain consistency with the other method that does take the codec argument?\n\nWould you be able to provide patches for trunk and 20-branch as well?","created":"2009-05-12T08:48:03.244+0000"},{"body":"I got to many thing going on right now to make a new patch fill free to mod my patch to work the way you want and use it to build a patch for trunk I would like to see this fixed in 0.20.1 if at all possible. this will be the one thing holding me up from upgrading to hbase 0.20 when it becomes ready.","created":"2009-05-12T19:01:09.653+0000"},{"body":"Updated the patch to trunk","created":"2009-05-15T07:18:10.296+0000"},{"body":"Could somebody review this patch? Thanks.","created":"2009-05-18T11:14:11.786+0000"},{"body":"Patch looks good.\n\nThis patch clashes with HADOOP-5572. Need to update this patch once HADOOP-5572 gets committed.\n\nHADOOP-5572 changes mergeParts() to call merge() with boolean sortSegments ----- this avoids one new signature of merge() from your patch.\n","created":"2009-05-19T18:09:55.284+0000"},{"body":"Patch updated to trunk","created":"2009-05-21T03:30:27.615+0000"},{"body":"Patch looks good.\n+1","created":"2009-05-21T11:18:35.164+0000"},{"body":"Patch for the 20 branch","created":"2009-05-27T09:39:31.989+0000"},{"body":"I just committed this. Thanks Jothi and Billy!","created":"2009-05-28T11:37:50.983+0000"},{"body":"Why no unit test? Why no javadoc for new methods?\n\nIf you tested this manually, what steps did you perform?","created":"2009-05-28T15:11:39.230+0000"},{"body":"no commit for 0.19 branch?","created":"2009-05-28T16:15:41.542+0000"},{"body":"bq. Why no unit test? If you tested this manually, what steps did you perform?\n\nIt is pretty difficult to write a unit test for this patch as this patch just enables compression during intermediate merges. The files that are created during the intermediate merges are consumed soon after they are created and the final merged file was compressed even without this patch. I did the same test as Billy had done -- add print statements in the framework code (Merger.java) to verify if compression was turned on during intermediate merges.\n\nbq. Why no javadoc for new methods?\n\nThe newly added methods are in Merger, which is a mapred package private class\n\nbq. no commit for 0.19 branch?\nBilly, from this comment https://issues.apache.org/jira/browse/HADOOP-5539?focusedCommentId=12708570&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#action_12708570, we thought you needed this only for 0.20. If you need it for 0.19 branch as well, I can generate a patch for that too.","created":"2009-05-29T03:29:32.207+0000"},{"body":"No I do not need it my version patch with my original patch for 0.19 but other might sense there is still a lot of older version in production that will update to 0.19 branch now that it has a few minor releases on it.\n","created":"2009-05-29T04:37:58.477+0000"}],"conversations":[{"body":"hadoop-site.xml :\nmapred.compress.map.output = true\n\nmap output files are compressed but when the in memory merger closes \non the reduce the on disk merger runs to reduce input files to <= io.sort.factor if needed. \n\nwhen this happens it outputs files called intermediate.x files these \ndo not maintain compression setting the writer (o.a.h.mapred.Merger.class line 432)\npasses the codec but I added some logging and its always null map output compression set true or false.\n\nThis causes task to fail if they can not hold the uncompressed size of the data of the reduce its holding\nI thank this is just and oversight of the codec not getting set correctly for the on disk merges.\n\n{code}\n2009-03-20 01:30:30,005 INFO org.apache.hadoop.mapred.Merger: Merging 30 intermediate segments out of a total of 3000\n2009-03-20 01:30:30,005 INFO org.apache.hadoop.mapred.Merger: intermediate.1 used codec: null\n{code}\n\nI added \n{code}\n // added my me\n\t if (codec != null){\n\t LOG.info(\"intermediate.\" + passNo + \" used codec: \" + codec.toString());\n\t } else {\n\t LOG.info(\"intermediate.\" + passNo + \" used codec: Null\");\n\t }\n\t // end added by me\n{code}\nJust before the creation of the writer o.a.h.mapred.Merger.class line 432\nand it outputs the second line above.\n\nI have confirmed this with the logging and I have looked at the files on the disk of the tasktracker. I can read the data in \nthe intermediate files clearly telling me that there not compressed but I can not read the map.out files direct from the map output\ntelling me the compression is working on the map end but not on the on disk merge that produces the intermediate.\n\nI can see no benefit for these not maintaining the compression setting and as it looks they where intended to maintain it.\n","from":"reporter","subject":"o.a.h.mapred.Merger not maintaining map out compression on intermediate files"},{"body":"I'd like to see this fixed as well since one reason I've enabled map output compression is to reduce disk space usage by the mapreduce framework. It appears that currently the map outputs are simply decompressed as soon as they have been downloaded by the reducer.","from":"developer"},{"body":"I verified that it is doing the same on map task no intermediate.x file from o.a.h.mapred.Merger are getting compressed.","from":"developer"},{"body":"this should fix the problem I had to make a few new constructors. I left the old constructors that these files where using because not sure if any other tasks using these. this patch will apply to 0.19-branch I have not worked any on trunk so might need to try dry-run before applying to trunk. tested on my end and working correctly now with this patch.\n","from":"developer"},{"body":"added back to read map","from":"developer"},{"body":"Someone can use my patch as a starting point.\n\nThe ReduceTask.java call that is the problem is line 2145\nThe maptask.java call that is the problem is line 1269\n\nI use streaming without a combiner so that should be looked at also to see if it uses o.a.h.mapred.Merger\n\nthe basic problem is the codec is not passed from these function to the merger so its always null the call to \no.a.h.mapred.Merger should include codec somehow if compression is not used then codec is null \nin both ReduceTask and Maptask.\n\nI thank this is a major bug that effects all MR jobs with disk bandwidth that uses compression.\n","from":"developer"},{"body":"bq. added back to read map \n\nIt's a blocker; it will be resolved and backported to 0.20 at least. The road map isn't; the PA queue defines the set of patches that can be committed. The fix version is usually set when it's actually resolved, so where it was committed is documented.","from":"developer"},{"body":"Oh, I see; the patch is for 0.19. My mistake. ","from":"developer"},{"body":"The patch looks good. A few minor points: \n\n# The new MergeQueue constructor could call the existing constructor and then set the codec later.\n{code}\npublic MergeQueue(Configuration conf, FileSystem fs, \n List> segments, RawComparator comparator,\n Progressable reporter, boolean sortSegments, CompressionCodec codec) {\n this(conf, fs, segments, comparator, reporter, sortSegments);\n this.codec = codec;\n }\n{code}\n\n# For the new merge methods, should we place the Codec argument after the valueClass argument (instead of being the last argument) to maintain consistency with the other method that does take the codec argument?\n\nWould you be able to provide patches for trunk and 20-branch as well?","from":"developer"},{"body":"I got to many thing going on right now to make a new patch fill free to mod my patch to work the way you want and use it to build a patch for trunk I would like to see this fixed in 0.20.1 if at all possible. this will be the one thing holding me up from upgrading to hbase 0.20 when it becomes ready.","from":"developer"},{"body":"Updated the patch to trunk","from":"developer"},{"body":"Could somebody review this patch? Thanks.","from":"developer"},{"body":"Patch looks good.\n\nThis patch clashes with HADOOP-5572. Need to update this patch once HADOOP-5572 gets committed.\n\nHADOOP-5572 changes mergeParts() to call merge() with boolean sortSegments ----- this avoids one new signature of merge() from your patch.\n","from":"developer"},{"body":"Patch updated to trunk","from":"developer"},{"body":"Patch looks good.\n+1","from":"developer"},{"body":"Patch for the 20 branch","from":"developer"},{"body":"I just committed this. Thanks Jothi and Billy!","from":"developer"},{"body":"Why no unit test? Why no javadoc for new methods?\n\nIf you tested this manually, what steps did you perform?","from":"developer"},{"body":"no commit for 0.19 branch?","from":"developer"},{"body":"bq. Why no unit test? If you tested this manually, what steps did you perform?\n\nIt is pretty difficult to write a unit test for this patch as this patch just enables compression during intermediate merges. The files that are created during the intermediate merges are consumed soon after they are created and the final merged file was compressed even without this patch. I did the same test as Billy had done -- add print statements in the framework code (Merger.java) to verify if compression was turned on during intermediate merges.\n\nbq. Why no javadoc for new methods?\n\nThe newly added methods are in Merger, which is a mapred package private class\n\nbq. no commit for 0.19 branch?\nBilly, from this comment https://issues.apache.org/jira/browse/HADOOP-5539?focusedCommentId=12708570&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#action_12708570, we thought you needed this only for 0.20. If you need it for 0.19 branch as well, I can generate a patch for that too.","from":"developer"},{"body":"No I do not need it my version patch with my original patch for 0.19 but other might sense there is still a lot of older version in production that will update to 0.19 branch now that it has a few minor releases on it.\n","from":"developer"}],"created":"2009-03-20T06:51:59.000+0000","description":"hadoop-site.xml :\nmapred.compress.map.output = true\n\nmap output files are compressed but when the in memory merger closes \non the reduce the on disk merger runs to reduce input files to <= io.sort.factor if needed. \n\nwhen this happens it outputs files called intermediate.x files these \ndo not maintain compression setting the writer (o.a.h.mapred.Merger.class line 432)\npasses the codec but I added some logging and its always null map output compression set true or false.\n\nThis causes task to fail if they can not hold the uncompressed size of the data of the reduce its holding\nI thank this is just and oversight of the codec not getting set correctly for the on disk merges.\n\n{code}\n2009-03-20 01:30:30,005 INFO org.apache.hadoop.mapred.Merger: Merging 30 intermediate segments out of a total of 3000\n2009-03-20 01:30:30,005 INFO org.apache.hadoop.mapred.Merger: intermediate.1 used codec: null\n{code}\n\nI added \n{code}\n // added my me\n\t if (codec != null){\n\t LOG.info(\"intermediate.\" + passNo + \" used codec: \" + codec.toString());\n\t } else {\n\t LOG.info(\"intermediate.\" + passNo + \" used codec: Null\");\n\t }\n\t // end added by me\n{code}\nJust before the creation of the writer o.a.h.mapred.Merger.class line 432\nand it outputs the second line above.\n\nI have confirmed this with the logging and I have looked at the files on the disk of the tasktracker. I can read the data in \nthe intermediate files clearly telling me that there not compressed but I can not read the map.out files direct from the map output\ntelling me the compression is working on the map end but not on the on disk merge that produces the intermediate.\n\nI can see no benefit for these not maintaining the compression setting and as it looks they where intended to maintain it.\n","issue_id":"12417354","key":"HADOOP-5539","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-05-28T11:37:51.000+0000","role":"fixed_distractor","summary":"o.a.h.mapred.Merger not maintaining map out compression on intermediate files"} {"case_id":"12424365","cluster":"DISTRACTOR-HADOOP-5762","comments":[{"body":"The problem happens if any of the source paths is an empty directory. It's because source directories are not \"visited\" in the inner loop that checks for files and directories to be copied and, thus, are not added to _distcp_src_files.","created":"2009-05-07T22:06:24.055+0000"},{"body":"Hi Rodrigo, which version is your patch for? Does this problem still exist on trunk?","created":"2009-05-07T22:56:59.315+0000"},{"body":"The problem still exists on trunk.\n\nThe patch was generated based on trunk.","created":"2009-05-08T01:16:26.685+0000"},{"body":"Just realized I had uploaded the wrong patch. I'm correcting it now.","created":"2009-05-09T23:42:59.396+0000"},{"body":"Correct patch.","created":"2009-05-10T00:16:08.455+0000"},{"body":"The root directory was not being included in the list of files/directories to be copied.\n\nBesides, copying was aborted if no data was being copied (only directories or empty files). The correct test should be on the number of files being copied.","created":"2009-05-10T00:18:43.136+0000"},{"body":"{code}\n+ return srcCount > 0;\n{code}\nChecking srcCount may not work well since srcCount counts all src paths. It will be >0 even there is nothing to copy.\n\nWe probably need a new varible dirCount and then return fileCount > 0 || dirCount > 0.","created":"2009-05-11T20:18:19.860+0000"},{"body":"You are right, Nicholas. What about using fileCount for that and just add empty directories to it?","created":"2009-05-11T20:31:04.650+0000"},{"body":"> What about using fileCount for that and just add empty directories to it?\nfileCount is also used for the -filelimit option which does not include dirs.\n\nFor creating directories, we don't really need a job. Why don't we make these directories during setup? i.e. mkdir instead of appending them to src_writer/dst_writer. What do you think?","created":"2009-05-11T20:42:02.416+0000"},{"body":"It could be done. Actually, my first attempt to fix this bug did exactly that. Then I noticed that method copy() was already programmed to copy empty directories and, except for the root source, directories were being added to src_writer on method setup. I think it's a good design option to let copy be the only method that actually writes something to the destination. This makes the code simpler and more elegant (thus, more maintainable).\n\nI'm don't like the idea of creating only the root src directory (in case it's empty) on setup(). It just doesn't look good. And creating all empty src subdirectories on setup might force us to do a lot of extra checking on setup to cope with failures (checks that already exist on copy()). I don't think many people will be using distcp to copy empty directories and it doesn't look like the performance gain will compensate the loss in code simplicity.","created":"2009-05-11T21:41:02.001+0000"},{"body":"Uploading a new patch.","created":"2009-05-12T23:34:05.369+0000"},{"body":"+1 The latest patch looks good.","created":"2009-05-27T14:49:46.446+0000"},{"body":"The failed tests do not seem related. I will commit this soon.","created":"2009-06-03T22:05:09.237+0000"},{"body":"I have committed this. Thanks, Rodrigo!","created":"2009-06-03T22:11:36.390+0000"}],"conversations":[{"body":"If I have an empty directory /testdir1 and then I run the command bin/hadoop distcp /testdir1 /testdir2, the command completes successfully, but does not create the empty directory /testdir2.\n\n","from":"reporter","subject":"distcp does not copy empty directories"},{"body":"The problem happens if any of the source paths is an empty directory. It's because source directories are not \"visited\" in the inner loop that checks for files and directories to be copied and, thus, are not added to _distcp_src_files.","from":"developer"},{"body":"Hi Rodrigo, which version is your patch for? Does this problem still exist on trunk?","from":"developer"},{"body":"The problem still exists on trunk.\n\nThe patch was generated based on trunk.","from":"developer"},{"body":"Just realized I had uploaded the wrong patch. I'm correcting it now.","from":"developer"},{"body":"Correct patch.","from":"developer"},{"body":"The root directory was not being included in the list of files/directories to be copied.\n\nBesides, copying was aborted if no data was being copied (only directories or empty files). The correct test should be on the number of files being copied.","from":"developer"},{"body":"{code}\n+ return srcCount > 0;\n{code}\nChecking srcCount may not work well since srcCount counts all src paths. It will be >0 even there is nothing to copy.\n\nWe probably need a new varible dirCount and then return fileCount > 0 || dirCount > 0.","from":"developer"},{"body":"You are right, Nicholas. What about using fileCount for that and just add empty directories to it?","from":"developer"},{"body":"> What about using fileCount for that and just add empty directories to it?\nfileCount is also used for the -filelimit option which does not include dirs.\n\nFor creating directories, we don't really need a job. Why don't we make these directories during setup? i.e. mkdir instead of appending them to src_writer/dst_writer. What do you think?","from":"developer"},{"body":"It could be done. Actually, my first attempt to fix this bug did exactly that. Then I noticed that method copy() was already programmed to copy empty directories and, except for the root source, directories were being added to src_writer on method setup. I think it's a good design option to let copy be the only method that actually writes something to the destination. This makes the code simpler and more elegant (thus, more maintainable).\n\nI'm don't like the idea of creating only the root src directory (in case it's empty) on setup(). It just doesn't look good. And creating all empty src subdirectories on setup might force us to do a lot of extra checking on setup to cope with failures (checks that already exist on copy()). I don't think many people will be using distcp to copy empty directories and it doesn't look like the performance gain will compensate the loss in code simplicity.","from":"developer"},{"body":"Uploading a new patch.","from":"developer"},{"body":"+1 The latest patch looks good.","from":"developer"},{"body":"The failed tests do not seem related. I will commit this soon.","from":"developer"},{"body":"I have committed this. Thanks, Rodrigo!","from":"developer"}],"created":"2009-05-01T01:14:30.000+0000","description":"If I have an empty directory /testdir1 and then I run the command bin/hadoop distcp /testdir1 /testdir2, the command completes successfully, but does not create the empty directory /testdir2.\n\n","issue_id":"12424365","key":"HADOOP-5762","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-06-03T22:11:36.000+0000","role":"fixed_distractor","summary":"distcp does not copy empty directories"} {"case_id":"12425606","cluster":"DISTRACTOR-HADOOP-5850","comments":[{"body":"Attaching patch to fix this issue.\n - With this patch, jobs will 0 maps or no input still run JobSetUp, any number of reduces( which do nothing), and the JobCleanUp task.\n - Removed the 0-splits check in JobInProgress.initTasks() and added checks so that cleanup task doesn't launch before setup tasks when number of splits is zero.\n - Renamed TestEmptyJobWithDFS to TestEmptyJob, removed HDFS dependence to quicken the test, added checks for verifying the number of map and reduce tasks run for an empty-job.","created":"2009-05-19T06:50:45.849+0000"},{"body":"Changes to JobInProgress looks fine.\nCould you add an assertion to the test case to verify presence of empty output directories as well? That should fail without the patch and should pass with the patch.","created":"2009-05-19T09:37:37.610+0000"},{"body":"Attaching new patch with the suggested changes. The test now fails without the core changes and succeeds with.","created":"2009-05-19T10:38:14.399+0000"},{"body":"ant test-patch results:\n{code}\n [exec] +1 overall.\n [exec]\n [exec] +1 @author. The patch does not contain any @author tags.\n [exec]\n [exec] +1 tests included. The patch appears to include 6 new or modified tests.\n [exec]\n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec]\n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec]\n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec]\n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n [exec]\n [exec] +1 release audit. The applied patch does not increase the total number of release audit warnings.\n [exec]\n [exec]\n [exec]\n [exec]\n [exec] ======================================================================\n [exec] ======================================================================\n [exec] Finished build.\n [exec] ======================================================================\n [exec] ======================================================================\n{code}","created":"2009-05-19T11:03:48.008+0000"},{"body":"Should the output directories be empty or should they produce zero-length part files ?","created":"2009-05-19T22:43:38.218+0000"},{"body":"bq. Should the output directories be empty or should they produce zero-length part files ? \nThe job will run N number of reduces that the user intends and they should produce zero-length part files. The patch does the same.","created":"2009-05-20T07:35:57.767+0000"},{"body":"You should be able to use LazyOutputFormat if you wanted no part files to be produced.","created":"2009-05-20T09:00:23.074+0000"},{"body":"Attaching final patch. The previous patch had some problems because of which TestSetupAndCleanupFailure failed.\n\nThis patch fixes those problems. It passed ant test-patch, core and contrib tests except the following which failed/timeout even without this patch and are unrelated to the changes in this patch.\n\nFailed:\n - org.apache.hadoop.streaming.TestMultipleCachefiles\n - org.apache.hadoop.streaming.TestStreamingBadRecords\n - org.apache.hadoop.streaming.TestSymLink\n\nTimedout:\n - org.apache.hadoop.mapred.TestJobInProgressListener FAILED (timeout)\n - org.apache.hadoop.mapred.TestQueueCapacities FAILED (timeout)","created":"2009-05-20T20:14:30.478+0000"},{"body":"Forgot to summarize. With this patch,\n - Jobs with zero maps will still run job setup and cleanup tasks.\n - If number of reduces is non-zero, reduces are run leaving behind the corresponding number of empty part-files in the output directory.\n - If the number of reduces is also zero, an empty output directory is left behind.\n - The map progress(and reduce progress if number of reduces is zero) is set to 1.0 once the job cleanup task finishes.","created":"2009-05-20T20:18:02.395+0000"},{"body":"Patch for branch 0.20.","created":"2009-05-21T03:27:34.424+0000"},{"body":"Setting Map Progress and reduce progress to 1.0f should not be done in updateTaskStatus() method, it can be done when setup completes in completedTask().\nCatching NullPointerException in testcase doesn't seem correct, just throw it out if there is any.\n","created":"2009-05-21T08:57:25.185+0000"},{"body":"The patch seems to introduce a new error. The taskdetails.jsp page of job setup and cleanup tasks throws a NullPointerException. This is the case with all the jobs' setup and cleanup tasks. Below is the Exception seen on the UI:\njava.lang.NullPointerException\nat org.apache.hadoop.mapred.TaskInProgress.getSplitNodes(TaskInProgress.java:1034)\nat org.apache.hadoop.mapred.taskdetails_jsp._jspService(taskdetails_jsp.java:288)\nat org.apache.jasper.runtime.HttpJspBase.service(HttpJspBase.java:97)\nat javax.servlet.http.HttpServlet.service(HttpServlet.java:820)\nat org.mortbay.jetty.servlet.ServletHolder.handle(ServletHolder.java:502)\nat org.mortbay.jetty.servlet.ServletHandler.handle(ServletHandler.java:363)\nat org.mortbay.jetty.security.SecurityHandler.handle(SecurityHandler.java:216)\nat org.mortbay.jetty.servlet.SessionHandler.handle(SessionHandler.java:181)\nat org.mortbay.jetty.handler.ContextHandler.handle(ContextHandler.java:766)\nat org.mortbay.jetty.webapp.WebAppContext.handle(WebAppContext.java:4","created":"2009-05-21T10:57:47.992+0000"},{"body":"Had offline discussion with Nigel. Attaching a screenshot for the Exception thrown in taskdetails.jsp page.","created":"2009-05-22T03:52:26.615+0000"},{"body":"Thank you Ramya for finding out a problem with the patch! I am working on fixing this and uploading a new patch.","created":"2009-05-22T04:11:26.468+0000"},{"body":"Attaching patch incorporating the review comments and fixing the issues.","created":"2009-05-22T11:27:46.363+0000"},{"body":"changes look fine to me","created":"2009-05-22T11:40:20.059+0000"},{"body":"ant test-patch and run-test-mapred targets passed with the patch.","created":"2009-05-22T14:50:14.706+0000"},{"body":"Patch for branch-20.","created":"2009-05-22T15:09:06.123+0000"},{"body":"I just committed this. Thanks, Vinod!","created":"2009-05-22T15:29:37.345+0000"},{"body":"With the above fix, when a job (writing to DFS) with 0 maps and >0 reduces is submitted, the cluster hangs completely.\nThe TT logs show \"INFO org.apache.hadoop.mapred.TaskTracker: Resending 'status' to ' with reponseId 'ID\" infinitely and the JT throws java.io.IOException: java.lang.ArithmeticException forever. \nBelow is the stacktrace:\n{noformat} \n2009-05-25 08:13:00,124 INFO org.apache.hadoop.ipc.Server: IPC Server handler 37 on , call heartbeat(org.apache.hadoop.mapred.TaskTrackerStatus@14d128c, false, false, true, 3231) from : error: java.io.IOException:\n java.lang.ArithmeticException: / by zero\njava.io.IOException: java.lang.ArithmeticException: / by zero\n at org.apache.hadoop.mapred.ResourceEstimator.getEstimatedMapOutputSize(ResourceEstimator.java:85)\n at org.apache.hadoop.mapred.JobInProgress.findNewMapTask(JobInProgress.java:1729)\n at org.apache.hadoop.mapred.JobInProgress.obtainNewMapTask(JobInProgress.java:978)\n at org.apache.hadoop.mapred.CapacityTaskScheduler$MapSchedulingMgr.obtainNewTask(CapacityTaskScheduler.java:572)\n at org.apache.hadoop.mapred.CapacityTaskScheduler$TaskSchedulingMgr.getTaskFromQueue(CapacityTaskScheduler.java:418)\n at org.apache.hadoop.mapred.CapacityTaskScheduler$TaskSchedulingMgr.assignTasks(CapacityTaskScheduler.java:498)\n at org.apache.hadoop.mapred.CapacityTaskScheduler$TaskSchedulingMgr.access$500(CapacityTaskScheduler.java:277)\n at org.apache.hadoop.mapred.CapacityTaskScheduler.assignTasks(CapacityTaskScheduler.java:977)\n at org.apache.hadoop.mapred.JobTracker.heartbeat(JobTracker.java:2605)\n at sun.reflect.GeneratedMethodAccessor7.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:508)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:959)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:955)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:396)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:953)\n{noformat} \nIn such a case, all the jobs hang infinitely without progressing and the cluster is completely down.\nThis problem is solved only when the no-map job is killed. Once the job is killed the cluster is back running and the other jobs proceed smoothly.\n","created":"2009-05-25T10:43:19.517+0000"},{"body":"This situation occurs if scheduler is invoked and it calls job.obtainNewMapTask() while the job-clean-up task of this job is still running. Discussed with this Devaraj who concurs that obtainNewMapTask()/obtainNewReduceTask() should return immediately, doing nothing, when job-cleanup is running. Will address this in a new JIRA.","created":"2009-05-25T11:02:47.497+0000"},{"body":"bq. [...] Will address this in a new JIRA.\nHADOOP-5908\n","created":"2009-05-25T12:05:00.201+0000"}],"conversations":[{"body":"Currently, the framework ignores jobs that have 0 maps. This is incorrect. Many pipelines need the job to run (if nothing else, to create the output directory!) so that subsequent jobs don't fail. Effectively, there will be no map tasks and the reduce tasks should immediately set up the Reducer and RecordWriter and then call close on both since there are no inputs to the reduce. I believe it should just work if we remove the check...","from":"reporter","subject":"map/reduce doesn't run jobs with 0 maps"},{"body":"Attaching patch to fix this issue.\n - With this patch, jobs will 0 maps or no input still run JobSetUp, any number of reduces( which do nothing), and the JobCleanUp task.\n - Removed the 0-splits check in JobInProgress.initTasks() and added checks so that cleanup task doesn't launch before setup tasks when number of splits is zero.\n - Renamed TestEmptyJobWithDFS to TestEmptyJob, removed HDFS dependence to quicken the test, added checks for verifying the number of map and reduce tasks run for an empty-job.","from":"developer"},{"body":"Changes to JobInProgress looks fine.\nCould you add an assertion to the test case to verify presence of empty output directories as well? That should fail without the patch and should pass with the patch.","from":"developer"},{"body":"Attaching new patch with the suggested changes. The test now fails without the core changes and succeeds with.","from":"developer"},{"body":"ant test-patch results:\n{code}\n [exec] +1 overall.\n [exec]\n [exec] +1 @author. The patch does not contain any @author tags.\n [exec]\n [exec] +1 tests included. The patch appears to include 6 new or modified tests.\n [exec]\n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec]\n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec]\n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec]\n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n [exec]\n [exec] +1 release audit. The applied patch does not increase the total number of release audit warnings.\n [exec]\n [exec]\n [exec]\n [exec]\n [exec] ======================================================================\n [exec] ======================================================================\n [exec] Finished build.\n [exec] ======================================================================\n [exec] ======================================================================\n{code}","from":"developer"},{"body":"Should the output directories be empty or should they produce zero-length part files ?","from":"developer"},{"body":"bq. Should the output directories be empty or should they produce zero-length part files ? \nThe job will run N number of reduces that the user intends and they should produce zero-length part files. The patch does the same.","from":"developer"},{"body":"You should be able to use LazyOutputFormat if you wanted no part files to be produced.","from":"developer"},{"body":"Attaching final patch. The previous patch had some problems because of which TestSetupAndCleanupFailure failed.\n\nThis patch fixes those problems. It passed ant test-patch, core and contrib tests except the following which failed/timeout even without this patch and are unrelated to the changes in this patch.\n\nFailed:\n - org.apache.hadoop.streaming.TestMultipleCachefiles\n - org.apache.hadoop.streaming.TestStreamingBadRecords\n - org.apache.hadoop.streaming.TestSymLink\n\nTimedout:\n - org.apache.hadoop.mapred.TestJobInProgressListener FAILED (timeout)\n - org.apache.hadoop.mapred.TestQueueCapacities FAILED (timeout)","from":"developer"},{"body":"Forgot to summarize. With this patch,\n - Jobs with zero maps will still run job setup and cleanup tasks.\n - If number of reduces is non-zero, reduces are run leaving behind the corresponding number of empty part-files in the output directory.\n - If the number of reduces is also zero, an empty output directory is left behind.\n - The map progress(and reduce progress if number of reduces is zero) is set to 1.0 once the job cleanup task finishes.","from":"developer"},{"body":"Patch for branch 0.20.","from":"developer"},{"body":"Setting Map Progress and reduce progress to 1.0f should not be done in updateTaskStatus() method, it can be done when setup completes in completedTask().\nCatching NullPointerException in testcase doesn't seem correct, just throw it out if there is any.\n","from":"developer"},{"body":"The patch seems to introduce a new error. The taskdetails.jsp page of job setup and cleanup tasks throws a NullPointerException. This is the case with all the jobs' setup and cleanup tasks. Below is the Exception seen on the UI:\njava.lang.NullPointerException\nat org.apache.hadoop.mapred.TaskInProgress.getSplitNodes(TaskInProgress.java:1034)\nat org.apache.hadoop.mapred.taskdetails_jsp._jspService(taskdetails_jsp.java:288)\nat org.apache.jasper.runtime.HttpJspBase.service(HttpJspBase.java:97)\nat javax.servlet.http.HttpServlet.service(HttpServlet.java:820)\nat org.mortbay.jetty.servlet.ServletHolder.handle(ServletHolder.java:502)\nat org.mortbay.jetty.servlet.ServletHandler.handle(ServletHandler.java:363)\nat org.mortbay.jetty.security.SecurityHandler.handle(SecurityHandler.java:216)\nat org.mortbay.jetty.servlet.SessionHandler.handle(SessionHandler.java:181)\nat org.mortbay.jetty.handler.ContextHandler.handle(ContextHandler.java:766)\nat org.mortbay.jetty.webapp.WebAppContext.handle(WebAppContext.java:4","from":"developer"},{"body":"Had offline discussion with Nigel. Attaching a screenshot for the Exception thrown in taskdetails.jsp page.","from":"developer"},{"body":"Thank you Ramya for finding out a problem with the patch! I am working on fixing this and uploading a new patch.","from":"developer"},{"body":"Attaching patch incorporating the review comments and fixing the issues.","from":"developer"},{"body":"changes look fine to me","from":"developer"},{"body":"ant test-patch and run-test-mapred targets passed with the patch.","from":"developer"},{"body":"Patch for branch-20.","from":"developer"},{"body":"I just committed this. Thanks, Vinod!","from":"developer"},{"body":"With the above fix, when a job (writing to DFS) with 0 maps and >0 reduces is submitted, the cluster hangs completely.\nThe TT logs show \"INFO org.apache.hadoop.mapred.TaskTracker: Resending 'status' to ' with reponseId 'ID\" infinitely and the JT throws java.io.IOException: java.lang.ArithmeticException forever. \nBelow is the stacktrace:\n{noformat} \n2009-05-25 08:13:00,124 INFO org.apache.hadoop.ipc.Server: IPC Server handler 37 on , call heartbeat(org.apache.hadoop.mapred.TaskTrackerStatus@14d128c, false, false, true, 3231) from : error: java.io.IOException:\n java.lang.ArithmeticException: / by zero\njava.io.IOException: java.lang.ArithmeticException: / by zero\n at org.apache.hadoop.mapred.ResourceEstimator.getEstimatedMapOutputSize(ResourceEstimator.java:85)\n at org.apache.hadoop.mapred.JobInProgress.findNewMapTask(JobInProgress.java:1729)\n at org.apache.hadoop.mapred.JobInProgress.obtainNewMapTask(JobInProgress.java:978)\n at org.apache.hadoop.mapred.CapacityTaskScheduler$MapSchedulingMgr.obtainNewTask(CapacityTaskScheduler.java:572)\n at org.apache.hadoop.mapred.CapacityTaskScheduler$TaskSchedulingMgr.getTaskFromQueue(CapacityTaskScheduler.java:418)\n at org.apache.hadoop.mapred.CapacityTaskScheduler$TaskSchedulingMgr.assignTasks(CapacityTaskScheduler.java:498)\n at org.apache.hadoop.mapred.CapacityTaskScheduler$TaskSchedulingMgr.access$500(CapacityTaskScheduler.java:277)\n at org.apache.hadoop.mapred.CapacityTaskScheduler.assignTasks(CapacityTaskScheduler.java:977)\n at org.apache.hadoop.mapred.JobTracker.heartbeat(JobTracker.java:2605)\n at sun.reflect.GeneratedMethodAccessor7.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:508)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:959)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:955)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:396)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:953)\n{noformat} \nIn such a case, all the jobs hang infinitely without progressing and the cluster is completely down.\nThis problem is solved only when the no-map job is killed. Once the job is killed the cluster is back running and the other jobs proceed smoothly.\n","from":"developer"},{"body":"This situation occurs if scheduler is invoked and it calls job.obtainNewMapTask() while the job-clean-up task of this job is still running. Discussed with this Devaraj who concurs that obtainNewMapTask()/obtainNewReduceTask() should return immediately, doing nothing, when job-cleanup is running. Will address this in a new JIRA.","from":"developer"},{"body":"bq. [...] Will address this in a new JIRA.\nHADOOP-5908\n","from":"developer"}],"created":"2009-05-15T16:59:13.000+0000","description":"Currently, the framework ignores jobs that have 0 maps. This is incorrect. Many pipelines need the job to run (if nothing else, to create the output directory!) so that subsequent jobs don't fail. Effectively, there will be no map tasks and the reduce tasks should immediately set up the Reducer and RecordWriter and then call close on both since there are no inputs to the reduce. I believe it should just work if we remove the check...","issue_id":"12425606","key":"HADOOP-5850","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-05-22T15:29:37.000+0000","role":"fixed_distractor","summary":"map/reduce doesn't run jobs with 0 maps"} {"case_id":"12353314","cluster":"DISTRACTOR-HADOOP-604","comments":[{"body":"I was investigating HADOOP-643. The problem is that the FSDataSet is not protected by locks. In this case, data-receive thread is receiving some data from another datanaode and is traversing the dataset to validate that the directory is valid. At the same time, the main datanode thread (offerService()) could be modifying the FSDataset. Protecting the DataXceiveServer thread from the offerservice thread is all this is needed.\n\nIt appears that the DataTransfer does not need any protection.","created":"2006-10-26T21:41:54.000+0000"},{"body":"\nThanks Dhruba. Yes, lot of accesses from offerServer() and other threads is not protected. I just started reading the surounding code to see exactly how this exception occurs what we should protect.\n","created":"2006-10-27T18:16:47.000+0000"},{"body":"\nWhat should this thread do if it gets a run time exception like this? Should that be part of the patch for this?\n","created":"2006-10-28T00:12:05.000+0000"},{"body":"it's best effort:\nif you can recover, by restarting the thread or otherwise, then recover.\nOtherwise, take down the entire process, so it's clear that you're down and require manual intervention.","created":"2006-10-28T00:29:07.000+0000"},{"body":"\nThis patch fixes the NPE and synchronizes datanode's directories and maps.\n","created":"2006-11-08T01:42:20.000+0000"},{"body":"I just committed this. Thanks, Raghu!","created":"2006-11-08T23:00:22.000+0000"}],"conversations":[{"body":"I wonder whether handling of 'children' in FSDataset is thread-safe.\n\nException in thread \"org.apache.hadoop.dfs.DataNode$DataXceiveServer@19b04e2\" java.lang.NullPointerException\n at org.apache.hadoop.dfs.FSDataset$FSDir.checkDirTree(FSDataset.java:162)\n at org.apache.hadoop.dfs.FSDataset$FSDir.checkDirTree(FSDataset.java:162)\n at org.apache.hadoop.dfs.FSDataset$FSDir.checkDirTree(FSDataset.java:162)\n at org.apache.hadoop.dfs.FSDataset$FSDir.checkDirTree(FSDataset.java:162)\n at org.apache.hadoop.dfs.FSDataset$FSVolume.checkDirs(FSDataset.java:238)\n at org.apache.hadoop.dfs.FSDataset$FSVolumeSet.checkDirs(FSDataset.java:326)\n at org.apache.hadoop.dfs.FSDataset.checkDataDir(FSDataset.java:522)\n at org.apache.hadoop.dfs.DataNode$DataXceiveServer.run(DataNode.java:472)\n at java.lang.Thread.run(Thread.java:595)\n","from":"reporter","subject":"DataNodes get NullPointerException and become unresponsive on 50010"},{"body":"I was investigating HADOOP-643. The problem is that the FSDataSet is not protected by locks. In this case, data-receive thread is receiving some data from another datanaode and is traversing the dataset to validate that the directory is valid. At the same time, the main datanode thread (offerService()) could be modifying the FSDataset. Protecting the DataXceiveServer thread from the offerservice thread is all this is needed.\n\nIt appears that the DataTransfer does not need any protection.","from":"developer"},{"body":"\nThanks Dhruba. Yes, lot of accesses from offerServer() and other threads is not protected. I just started reading the surounding code to see exactly how this exception occurs what we should protect.\n","from":"developer"},{"body":"\nWhat should this thread do if it gets a run time exception like this? Should that be part of the patch for this?\n","from":"developer"},{"body":"it's best effort:\nif you can recover, by restarting the thread or otherwise, then recover.\nOtherwise, take down the entire process, so it's clear that you're down and require manual intervention.","from":"developer"},{"body":"\nThis patch fixes the NPE and synchronizes datanode's directories and maps.\n","from":"developer"},{"body":"I just committed this. Thanks, Raghu!","from":"developer"}],"created":"2006-10-16T17:49:56.000+0000","description":"I wonder whether handling of 'children' in FSDataset is thread-safe.\n\nException in thread \"org.apache.hadoop.dfs.DataNode$DataXceiveServer@19b04e2\" java.lang.NullPointerException\n at org.apache.hadoop.dfs.FSDataset$FSDir.checkDirTree(FSDataset.java:162)\n at org.apache.hadoop.dfs.FSDataset$FSDir.checkDirTree(FSDataset.java:162)\n at org.apache.hadoop.dfs.FSDataset$FSDir.checkDirTree(FSDataset.java:162)\n at org.apache.hadoop.dfs.FSDataset$FSDir.checkDirTree(FSDataset.java:162)\n at org.apache.hadoop.dfs.FSDataset$FSVolume.checkDirs(FSDataset.java:238)\n at org.apache.hadoop.dfs.FSDataset$FSVolumeSet.checkDirs(FSDataset.java:326)\n at org.apache.hadoop.dfs.FSDataset.checkDataDir(FSDataset.java:522)\n at org.apache.hadoop.dfs.DataNode$DataXceiveServer.run(DataNode.java:472)\n at java.lang.Thread.run(Thread.java:595)\n","issue_id":"12353314","key":"HADOOP-604","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-11-08T23:00:22.000+0000","role":"fixed_distractor","summary":"DataNodes get NullPointerException and become unresponsive on 50010"} {"case_id":"12430515","cluster":"DISTRACTOR-HADOOP-6151","comments":[{"body":"I believe the transforms should be:\n 1. & -> &amp;\n 2. < -> &lt;\n 3. > -> &gt;\n 4. ' -> &apos;\n 5. \"-> &quot;\n\nAs long as we do those transforms, any html that the user includes in their data will just be treated as literal text rather than html commands.","created":"2009-07-15T22:55:16.710+0000"},{"body":"This patch introduces an input filter for all of the servlets and jsp pages that quotes all of the html active characters in the parameters. This means that all of the cross site scripting attacks based on bad urls should be fixed.\n\nI'll file a follow up jira to fix the vector where the values in the job need to be quoted.","created":"2009-09-17T22:00:49.939+0000"},{"body":"I forgot the --no-prefix..","created":"2009-09-18T02:26:14.462+0000"},{"body":"{noformat}\n [exec] -1 overall. \n [exec] \n [exec] +1 @author. The patch does not contain any @author tags.\n [exec] \n [exec] +1 tests included. The patch appears to include 2 new or modified tests.\n [exec] \n [exec] -1 javadoc. The javadoc tool appears to have generated 1 warning messages.\n [exec] \n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec] \n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec] \n [exec] +1 release audit. The applied patch does not increase the total number of release audit warnings.\n\n[javadoc] /snip/common/src/.../HtmlQuoting.java:145: warning - @return tag has no arguments.\n[javadoc] /snip/common/src/.../HtmlQuoting.java:73: warning - @param argument \"buffer\" is not a parameter name.\n[javadoc] /snip/common/src/.../HtmlQuoting.java:73: warning - @param argument \"add\" is not a parameter name.\n{noformat}\n\n* The unit test should use JUnit4 test annotations instead of JUnit3 TestCase\n* HttpServer::printRequest looks useful for debugging, but should probably be left out\n* The static \\*Bytes fields should be final\n* The @return docs for \"needsQuoting\" could be more explicit","created":"2009-09-18T04:31:08.299+0000"},{"body":"Messed up the JavaDoc. Now fixed.","created":"2009-09-18T04:49:17.696+0000"},{"body":"This patch addresses Chris' comments.","created":"2009-09-18T06:01:11.871+0000"},{"body":"+1","created":"2009-09-18T06:14:53.321+0000"},{"body":"I just committed this.","created":"2009-09-18T16:33:30.565+0000"},{"body":"This patch is for 0.20. (not to be committed)","created":"2009-12-05T01:35:52.553+0000"}],"conversations":[{"body":"We need to quote html characters that come from user generated data. Otherwise, all of the web ui's have cross site scripting attack, etc.","from":"reporter","subject":"The servlets should quote html characters"},{"body":"I believe the transforms should be:\n 1. & -> &amp;\n 2. < -> &lt;\n 3. > -> &gt;\n 4. ' -> &apos;\n 5. \"-> &quot;\n\nAs long as we do those transforms, any html that the user includes in their data will just be treated as literal text rather than html commands.","from":"developer"},{"body":"This patch introduces an input filter for all of the servlets and jsp pages that quotes all of the html active characters in the parameters. This means that all of the cross site scripting attacks based on bad urls should be fixed.\n\nI'll file a follow up jira to fix the vector where the values in the job need to be quoted.","from":"developer"},{"body":"I forgot the --no-prefix..","from":"developer"},{"body":"{noformat}\n [exec] -1 overall. \n [exec] \n [exec] +1 @author. The patch does not contain any @author tags.\n [exec] \n [exec] +1 tests included. The patch appears to include 2 new or modified tests.\n [exec] \n [exec] -1 javadoc. The javadoc tool appears to have generated 1 warning messages.\n [exec] \n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec] \n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec] \n [exec] +1 release audit. The applied patch does not increase the total number of release audit warnings.\n\n[javadoc] /snip/common/src/.../HtmlQuoting.java:145: warning - @return tag has no arguments.\n[javadoc] /snip/common/src/.../HtmlQuoting.java:73: warning - @param argument \"buffer\" is not a parameter name.\n[javadoc] /snip/common/src/.../HtmlQuoting.java:73: warning - @param argument \"add\" is not a parameter name.\n{noformat}\n\n* The unit test should use JUnit4 test annotations instead of JUnit3 TestCase\n* HttpServer::printRequest looks useful for debugging, but should probably be left out\n* The static \\*Bytes fields should be final\n* The @return docs for \"needsQuoting\" could be more explicit","from":"developer"},{"body":"Messed up the JavaDoc. Now fixed.","from":"developer"},{"body":"This patch addresses Chris' comments.","from":"developer"},{"body":"+1","from":"developer"},{"body":"I just committed this.","from":"developer"},{"body":"This patch is for 0.20. (not to be committed)","from":"developer"}],"created":"2009-07-15T16:24:17.000+0000","description":"We need to quote html characters that come from user generated data. Otherwise, all of the web ui's have cross site scripting attack, etc.","issue_id":"12430515","key":"HADOOP-6151","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-09-18T16:33:30.000+0000","role":"fixed_distractor","summary":"The servlets should quote html characters"} {"case_id":"12433708","cluster":"DISTRACTOR-HADOOP-6207","comments":[{"body":"(Note: I have patches, but I can't figure out where libhdfs has gone after the project split!)\n\nHere are the leaks I found so far:\nhdfsJniHelper.c: constructNewArrayString -> reference to newly created string is not released after it has been added to array. This leads to leaks when the array is destroyed, as all strings have reference count of at least one.\nhdfs.c: hdfsOpen -> reference to jAttrString is never destroyed, resulting in leak\nfuse-dfs.c: When a file is closed, the reference to the file system object is not destroyed.\n\nThe first one is the most critical - in FUSE, a string for each group the user is in is created per file open. This results in many hundreds of bytes leaking per file open.\n\nI was able to detect these by running FUSE with a heap size of 4MB, then opening and closing files until I got OOM exceptions.","created":"2009-08-21T17:41:23.894+0000"},{"body":"bq. hdfsJniHelper.c: constructNewArrayString -> reference to newly created string is not released after it has been added to array. This leads to leaks when the array is destroyed, as all strings have reference count of at least one.\n\nconstructNewArrayString got deleted because it was buggy and not actually used by anyone.\n\nbq. hdfs.c: hdfsOpen -> reference to jAttrString is never destroyed, resulting in leak\n\nI can't find jAttrString in the new code; it seems to have been removed by an earlier change.\n\nbq. fuse-dfs.c: When a file is closed, the reference to the file system object is not destroyed.\n\nFixed by HDFS-3608, which added a background thread which cleans up unused connections.\n\nIt seems like we should close this bug because these issues have been resolved.","created":"2012-08-16T00:29:38.235+0000"},{"body":"There is a high probability that this has been fixed and/or is stale. So closing it as such. If not, please file a new jira.","created":"2014-07-23T21:50:24.812+0000"}],"conversations":[{"body":"libhdfs leaks many objects during normal operation. This becomes exacerbated by long-running processes (such as FUSE-DFS).","from":"reporter","subject":"libhdfs leaks object references"},{"body":"(Note: I have patches, but I can't figure out where libhdfs has gone after the project split!)\n\nHere are the leaks I found so far:\nhdfsJniHelper.c: constructNewArrayString -> reference to newly created string is not released after it has been added to array. This leads to leaks when the array is destroyed, as all strings have reference count of at least one.\nhdfs.c: hdfsOpen -> reference to jAttrString is never destroyed, resulting in leak\nfuse-dfs.c: When a file is closed, the reference to the file system object is not destroyed.\n\nThe first one is the most critical - in FUSE, a string for each group the user is in is created per file open. This results in many hundreds of bytes leaking per file open.\n\nI was able to detect these by running FUSE with a heap size of 4MB, then opening and closing files until I got OOM exceptions.","from":"developer"},{"body":"bq. hdfsJniHelper.c: constructNewArrayString -> reference to newly created string is not released after it has been added to array. This leads to leaks when the array is destroyed, as all strings have reference count of at least one.\n\nconstructNewArrayString got deleted because it was buggy and not actually used by anyone.\n\nbq. hdfs.c: hdfsOpen -> reference to jAttrString is never destroyed, resulting in leak\n\nI can't find jAttrString in the new code; it seems to have been removed by an earlier change.\n\nbq. fuse-dfs.c: When a file is closed, the reference to the file system object is not destroyed.\n\nFixed by HDFS-3608, which added a background thread which cleans up unused connections.\n\nIt seems like we should close this bug because these issues have been resolved.","from":"developer"},{"body":"There is a high probability that this has been fixed and/or is stale. So closing it as such. If not, please file a new jira.","from":"developer"}],"created":"2009-08-21T17:36:40.000+0000","description":"libhdfs leaks many objects during normal operation. This becomes exacerbated by long-running processes (such as FUSE-DFS).","issue_id":"12433708","key":"HADOOP-6207","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2014-07-23T21:50:24.000+0000","role":"fixed_distractor","summary":"libhdfs leaks object references"} {"case_id":"12434590","cluster":"DISTRACTOR-HADOOP-6231","comments":[{"body":"A patch that implements this in the way suggested at https://issues.apache.org/jira/browse/HADOOP-6097?focusedCommentId=12745027&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#action_12745027.","created":"2009-09-02T07:00:08.870+0000"},{"body":"The unit test should use JUnit4 conventions rather than JUnit3, but aside from that the fix looks good","created":"2009-09-02T07:58:58.506+0000"},{"body":"Thanks for the suggestion, Chris. Here's a modified patch.","created":"2009-09-03T01:06:34.357+0000"},{"body":"I've just committed this.","created":"2009-09-04T23:38:18.058+0000"},{"body":"Patch against the 0.20 branch.","created":"2009-09-15T14:42:36.892+0000"},{"body":"This fixes a bug affecting all releases from 0.18 (when Hadoop archives were introduced). Without this patch, only one Hadoop archive can be opened at a time:\n\n{noformat}\n$ hadoop dfs -ls har:///user/knoguchi/test.har har:///user/knoguchi/test2.har\nFound 1 items\ndrw-r--r-- - knoguchi users 0 2009-08-18 18:52 /user/knoguchi/test.har/user\nls: Invalid file name: /user/knoguchi/test2.har in har:///user/knoguchi/test.har\n\n$ hadoop dfs -ls har:///user/knoguchi/test2.har har:///user/knoguchi/test.har\nFound 1 items\ndrw------- - knoguchi users 0 2009-08-17 19:15 /user/knoguchi/test2.har/user\nls: Invalid file name: /user/knoguchi/test.har in har:///user/knoguchi/test2.har\n{noformat}","created":"2009-09-15T14:49:18.796+0000"},{"body":" [exec]\n [exec] +1 overall.\n [exec]\n [exec] +1 @author. The patch does not contain any @author tags.\n [exec]\n [exec] +1 tests included. The patch appears to include 3 new or modified tests.\n [exec]\n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec]\n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec]\n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec]\n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n [exec]\n\noutput of ant test-patch.... am still running ant test and will post the results.... ","created":"2009-09-17T22:39:32.118+0000"},{"body":"I just committed this to 0.20 branch. thanks tom and ben!","created":"2009-09-18T18:47:23.380+0000"},{"body":"Isn't this really an improvement (not a bug). Can you write a release note for this?","created":"2009-09-18T20:11:42.318+0000"},{"body":"nigel,\n This is really a bug, given that multiple archives in a single wont work without this patch. Please look at HADOOP-6097 for more details.\n I will add a release note on this.\n\nthanks","created":"2009-09-18T21:06:34.829+0000"},{"body":"sorry pressed enter too soon,\n I meant \"multiple archives in a single jvm wont work\"\n\n","created":"2009-09-18T21:07:27.952+0000"}],"conversations":[{"body":"HAR filesystem instances should not be cached, so this JIRA seeks to provide a general mechanism for disabling the cache on a per-filesystem basis. (Carried over from HADOOP-6097.)","from":"reporter","subject":"Allow caching of filesystem instances to be disabled on a per-instance basis"},{"body":"A patch that implements this in the way suggested at https://issues.apache.org/jira/browse/HADOOP-6097?focusedCommentId=12745027&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#action_12745027.","from":"developer"},{"body":"The unit test should use JUnit4 conventions rather than JUnit3, but aside from that the fix looks good","from":"developer"},{"body":"Thanks for the suggestion, Chris. Here's a modified patch.","from":"developer"},{"body":"I've just committed this.","from":"developer"},{"body":"Patch against the 0.20 branch.","from":"developer"},{"body":"This fixes a bug affecting all releases from 0.18 (when Hadoop archives were introduced). Without this patch, only one Hadoop archive can be opened at a time:\n\n{noformat}\n$ hadoop dfs -ls har:///user/knoguchi/test.har har:///user/knoguchi/test2.har\nFound 1 items\ndrw-r--r-- - knoguchi users 0 2009-08-18 18:52 /user/knoguchi/test.har/user\nls: Invalid file name: /user/knoguchi/test2.har in har:///user/knoguchi/test.har\n\n$ hadoop dfs -ls har:///user/knoguchi/test2.har har:///user/knoguchi/test.har\nFound 1 items\ndrw------- - knoguchi users 0 2009-08-17 19:15 /user/knoguchi/test2.har/user\nls: Invalid file name: /user/knoguchi/test.har in har:///user/knoguchi/test2.har\n{noformat}","from":"developer"},{"body":" [exec]\n [exec] +1 overall.\n [exec]\n [exec] +1 @author. The patch does not contain any @author tags.\n [exec]\n [exec] +1 tests included. The patch appears to include 3 new or modified tests.\n [exec]\n [exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n [exec]\n [exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n [exec]\n [exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n [exec]\n [exec] +1 Eclipse classpath. The patch retains Eclipse classpath integrity.\n [exec]\n\noutput of ant test-patch.... am still running ant test and will post the results.... ","from":"developer"},{"body":"I just committed this to 0.20 branch. thanks tom and ben!","from":"developer"},{"body":"Isn't this really an improvement (not a bug). Can you write a release note for this?","from":"developer"},{"body":"nigel,\n This is really a bug, given that multiple archives in a single wont work without this patch. Please look at HADOOP-6097 for more details.\n I will add a release note on this.\n\nthanks","from":"developer"},{"body":"sorry pressed enter too soon,\n I meant \"multiple archives in a single jvm wont work\"\n\n","from":"developer"}],"created":"2009-09-02T06:50:36.000+0000","description":"HAR filesystem instances should not be cached, so this JIRA seeks to provide a general mechanism for disabling the cache on a per-filesystem basis. (Carried over from HADOOP-6097.)","issue_id":"12434590","key":"HADOOP-6231","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-09-18T18:47:23.000+0000","role":"fixed_distractor","summary":"Allow caching of filesystem instances to be disabled on a per-instance basis"} {"case_id":"12401189","cluster":"DISTRACTOR-HADOOP-6234","comments":[{"body":"No, it isn't acceptable to break backwards compatability.\n\nSymbolic is ok, since it isn't ambiguous. To use straight octal we either need a flag like a leading o or use a new attribute and handle the old attribute in decimal.\n\nMy other concern in this area is against supporting octal via a leading 0 on all integer attributes. That will lead to massive confusion, in my opinion.\n","created":"2009-08-18T15:57:55.262+0000"},{"body":"I'd guess the massive confusion at this point would be that we're using decimal rather than octal, but yeah, changing at this point cold-turkey would probably just confuse a our established users rather than our new users. My concern over using a lead 0 is that li/unix only uses octal for chmod, whether or not a leading 0 is specified. I'd like to end up with something that matches that behavior.\nIn that case, the best thing is probably to deprecate the current key and introduce a new one with octal/symbolic semantics. That's the approach I'll plan on using. Too bad HADOOP-6105 isn't up and runnning yet.","created":"2009-08-18T19:37:50.922+0000"},{"body":"this is one of those things where i think it is acceptable to break existing users, given proper release notes and enough warning. really. if we are willing to break api's then breaking something as relatively minor as permissions in the conf file does not seem like that big of deal.\n\nperhaps this will be the jira that finally provides 'upgrade' notes instead of just 'release' notes, like most real world operating systems.","created":"2009-08-20T03:41:25.868+0000"},{"body":"Done with most of the work. I very much would like to use HADOOP-6105, but it's not yet committed. I'll be out of town until next Wednesday, so attaching prelim patch for review and hopefully 6105 will be ready to go when I get back.\n\nPatch:\n * Creates a new config option, dfs.umaskmode, which will take either an octal or symbolic umask value.\n * Separates out the permission parsing from the shell code. This allows us to use the exact same code for parsing in both places.\n * Does not change external view of any of the permission processing\n * Creates a new PermissionParser and a package-private UmaskParser to handle the subtle differences between what regular permissions and umask permissions allow\n * Creates new unit tests to verify we're getting the right value from the conf file. Most of the parsing code testing is free, since we're now using common code in all the parsing.\n\nStill need to:\n* Use 6105 to handle old key\n* Test against hdfs","created":"2009-08-28T02:37:01.559+0000"},{"body":"Moved to Common, where the affected files actually reside","created":"2009-09-02T18:02:38.113+0000"},{"body":"Finished patch. Rather than wait for 6105, handcoded check for prior value. When 6105 is committed, will update to use its functionality. \n\nCompared to previous patch, provided backwards compatibility, updated documentation and split one test into three. \n\nOf note, HDFS needs to be updated due to this patch and will open a new JIRA for that now.\n\nPasses all unit tests.\n\nTest-patch:\n{noformat}\n[exec] +1 overall. \n[exec] \n[exec] +1 @author. The patch does not contain any @author tags.\n[exec] \n[exec] +1 tests included. The patch appears to include 3 new or modified tests.\n[exec] \n[exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n[exec] \n[exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n[exec] \n[exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n[exec] \n[exec] +1 release audit. The applied patch does not increase the total number of release audit warnings.\n{noformat}","created":"2009-09-04T02:09:46.537+0000"},{"body":"submitting patch, although already ran unit tests and test-patch.","created":"2009-09-04T02:11:58.219+0000"},{"body":"+1 for the patch.","created":"2009-09-04T21:53:06.506+0000"},{"body":"Committed the change. Thanks Jakob.","created":"2009-09-08T19:09:27.749+0000"},{"body":"Editorial pass over all release notes prior to publication of 0.21.","created":"2009-10-09T03:51:38.157+0000"},{"body":"attaching a patch for branch 0.20. This patch has some differences from the 0.21 patch. The changes related to sticky bits had to be removed since 0.20 does not have sticky bit feature.","created":"2009-11-17T23:11:34.021+0000"},{"body":"Patch for 20 looks good. Essentially it is octal/symbolic umask settings in a universe where sticky bits don't exist.\n\nComments:\n* Line 379 - remove reference to t in comment\n* Line 526 - Add back reference to X being prohibited\n* Line 552 - Extraneous import of FSNamesystem\n\nThese are minor and otherwise +1.","created":"2009-11-18T00:26:21.530+0000"},{"body":"For the 0.20 patch, we should test whether it has the problem described in HDFS-760.","created":"2009-11-18T23:39:49.617+0000"},{"body":"New patch that addresses Jakob's comments.","created":"2009-11-20T17:38:35.800+0000"},{"body":"For 0.20 patch, the problem described in HDFS-760 is not relevant. On trunk we use a different backward compatibility mechanism to map deprecated configuration key to new configuration key, using DeprecatedKeyInfo. That patch has its own mechanism to provide backward compatibility.","created":"2009-11-20T17:43:55.121+0000"},{"body":"updated patch looks good. +1.","created":"2009-11-20T19:41:08.899+0000"}],"conversations":[{"body":"Currently, the settings for the default umask in Hadoop configuration files require the input format be in decimal. Considering that every admin in the world thinks of permissions in octal and/or symbolic format, the config files should really use those two formats and drop decimal.\n\n[... and, yes, I'm aware this breaks backwards compatibility. But in this case, I think that is perfectly acceptable.]","from":"reporter","subject":"Permission configuration files should use octal and symbolic"},{"body":"No, it isn't acceptable to break backwards compatability.\n\nSymbolic is ok, since it isn't ambiguous. To use straight octal we either need a flag like a leading o or use a new attribute and handle the old attribute in decimal.\n\nMy other concern in this area is against supporting octal via a leading 0 on all integer attributes. That will lead to massive confusion, in my opinion.\n","from":"developer"},{"body":"I'd guess the massive confusion at this point would be that we're using decimal rather than octal, but yeah, changing at this point cold-turkey would probably just confuse a our established users rather than our new users. My concern over using a lead 0 is that li/unix only uses octal for chmod, whether or not a leading 0 is specified. I'd like to end up with something that matches that behavior.\nIn that case, the best thing is probably to deprecate the current key and introduce a new one with octal/symbolic semantics. That's the approach I'll plan on using. Too bad HADOOP-6105 isn't up and runnning yet.","from":"developer"},{"body":"this is one of those things where i think it is acceptable to break existing users, given proper release notes and enough warning. really. if we are willing to break api's then breaking something as relatively minor as permissions in the conf file does not seem like that big of deal.\n\nperhaps this will be the jira that finally provides 'upgrade' notes instead of just 'release' notes, like most real world operating systems.","from":"developer"},{"body":"Done with most of the work. I very much would like to use HADOOP-6105, but it's not yet committed. I'll be out of town until next Wednesday, so attaching prelim patch for review and hopefully 6105 will be ready to go when I get back.\n\nPatch:\n * Creates a new config option, dfs.umaskmode, which will take either an octal or symbolic umask value.\n * Separates out the permission parsing from the shell code. This allows us to use the exact same code for parsing in both places.\n * Does not change external view of any of the permission processing\n * Creates a new PermissionParser and a package-private UmaskParser to handle the subtle differences between what regular permissions and umask permissions allow\n * Creates new unit tests to verify we're getting the right value from the conf file. Most of the parsing code testing is free, since we're now using common code in all the parsing.\n\nStill need to:\n* Use 6105 to handle old key\n* Test against hdfs","from":"developer"},{"body":"Moved to Common, where the affected files actually reside","from":"developer"},{"body":"Finished patch. Rather than wait for 6105, handcoded check for prior value. When 6105 is committed, will update to use its functionality. \n\nCompared to previous patch, provided backwards compatibility, updated documentation and split one test into three. \n\nOf note, HDFS needs to be updated due to this patch and will open a new JIRA for that now.\n\nPasses all unit tests.\n\nTest-patch:\n{noformat}\n[exec] +1 overall. \n[exec] \n[exec] +1 @author. The patch does not contain any @author tags.\n[exec] \n[exec] +1 tests included. The patch appears to include 3 new or modified tests.\n[exec] \n[exec] +1 javadoc. The javadoc tool did not generate any warning messages.\n[exec] \n[exec] +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n[exec] \n[exec] +1 findbugs. The patch does not introduce any new Findbugs warnings.\n[exec] \n[exec] +1 release audit. The applied patch does not increase the total number of release audit warnings.\n{noformat}","from":"developer"},{"body":"submitting patch, although already ran unit tests and test-patch.","from":"developer"},{"body":"+1 for the patch.","from":"developer"},{"body":"Committed the change. Thanks Jakob.","from":"developer"},{"body":"Editorial pass over all release notes prior to publication of 0.21.","from":"developer"},{"body":"attaching a patch for branch 0.20. This patch has some differences from the 0.21 patch. The changes related to sticky bits had to be removed since 0.20 does not have sticky bit feature.","from":"developer"},{"body":"Patch for 20 looks good. Essentially it is octal/symbolic umask settings in a universe where sticky bits don't exist.\n\nComments:\n* Line 379 - remove reference to t in comment\n* Line 526 - Add back reference to X being prohibited\n* Line 552 - Extraneous import of FSNamesystem\n\nThese are minor and otherwise +1.","from":"developer"},{"body":"For the 0.20 patch, we should test whether it has the problem described in HDFS-760.","from":"developer"},{"body":"New patch that addresses Jakob's comments.","from":"developer"},{"body":"For 0.20 patch, the problem described in HDFS-760 is not relevant. On trunk we use a different backward compatibility mechanism to map deprecated configuration key to new configuration key, using DeprecatedKeyInfo. That patch has its own mechanism to provide backward compatibility.","from":"developer"},{"body":"updated patch looks good. +1.","from":"developer"}],"created":"2008-07-28T21:05:27.000+0000","description":"Currently, the settings for the default umask in Hadoop configuration files require the input format be in decimal. Considering that every admin in the world thinks of permissions in octal and/or symbolic format, the config files should really use those two formats and drop decimal.\n\n[... and, yes, I'm aware this breaks backwards compatibility. But in this case, I think that is perfectly acceptable.]","issue_id":"12401189","key":"HADOOP-6234","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-09-08T19:09:27.000+0000","role":"fixed_distractor","summary":"Permission configuration files should use octal and symbolic"} {"case_id":"12436547","cluster":"DISTRACTOR-HADOOP-6285","comments":[{"body":"This patch fixes the signature of getParameterMap and adds testcases of HttpServer.","created":"2009-09-25T17:27:17.619+0000"},{"body":"The new test needs to create build/webapps/test. *sigh*","created":"2009-09-25T18:21:22.248+0000"},{"body":"+1 patch looks good.","created":"2009-09-25T20:21:45.352+0000"},{"body":"I just committed this.","created":"2009-09-28T21:26:45.970+0000"},{"body":"This patch applies to 0.20. (not to be committed)","created":"2009-12-05T01:38:43.714+0000"}],"conversations":[{"body":"In the HDFS tests, we see:\n{noformat}\njava.lang.ClassCastException: [Ljava.lang.String; cannot be cast to java.lang.String\n\tat org.apache.hadoop.http.HttpServer$QuotingInputFilter$RequestQuoter.getParameterMap(HttpServer.java:591)\n\tat org.apache.hadoop.hdfs.server.namenode.FsckServlet.doGet(FsckServlet.java:44)\n\tat javax.servlet.http.HttpServlet.service(HttpServlet.java:707)\n\tat javax.servlet.http.HttpServlet.service(HttpServlet.java:820)\n\tat org.mortbay.jetty.servlet.ServletHolder.handle(ServletHolder.java:502)\n\tat org.mortbay.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1124)\n\tat org.apache.hadoop.http.HttpServer$QuotingInputFilter.doFilter(HttpServer.java:613)\n\tat org.mortbay.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1115)\n\tat org.mortbay.jetty.servlet.ServletHandler.handle(ServletHandler.java:361)\n\tat org.mortbay.jetty.security.SecurityHandler.handle(SecurityHandler.java:216)\n\tat org.mortbay.jetty.servlet.SessionHandler.handle(SessionHandler.java:181)\n\tat org.mortbay.jetty.handler.ContextHandler.handle(ContextHandler.java:766)\n\tat org.mortbay.jetty.webapp.WebAppContext.handle(WebAppContext.java:417)\n\tat org.mortbay.jetty.handler.ContextHandlerCollection.handle(ContextHandlerCollection.java:230)\n\tat org.mortbay.jetty.handler.HandlerWrapper.handle(HandlerWrapper.java:152)\n\tat org.mortbay.jetty.Server.handle(Server.java:324)\n\tat org.mortbay.jetty.HttpConnection.handleRequest(HttpConnection.java:534)\n\tat org.mortbay.jetty.HttpConnection$RequestHandler.headerComplete(HttpConnection.java:864)\n\tat org.mortbay.jetty.HttpParser.parseNext(HttpParser.java:533)\n\tat org.mortbay.jetty.HttpParser.parseAvailable(HttpParser.java:207)\n\tat org.mortbay.jetty.HttpConnection.handle(HttpConnection.java:403)\n\tat org.mortbay.io.nio.SelectChannelEndPoint.run(SelectChannelEndPoint.java:409)\n\tat org.mortbay.thread.QueuedThreadPool$PoolThread.run(QueuedThreadPool.java:522)\n{noformat}\n\n","from":"reporter","subject":"HttpServer.QuotingInputFilter has the wrong signature for getParameterMap"},{"body":"This patch fixes the signature of getParameterMap and adds testcases of HttpServer.","from":"developer"},{"body":"The new test needs to create build/webapps/test. *sigh*","from":"developer"},{"body":"+1 patch looks good.","from":"developer"},{"body":"I just committed this.","from":"developer"},{"body":"This patch applies to 0.20. (not to be committed)","from":"developer"}],"created":"2009-09-24T23:00:03.000+0000","description":"In the HDFS tests, we see:\n{noformat}\njava.lang.ClassCastException: [Ljava.lang.String; cannot be cast to java.lang.String\n\tat org.apache.hadoop.http.HttpServer$QuotingInputFilter$RequestQuoter.getParameterMap(HttpServer.java:591)\n\tat org.apache.hadoop.hdfs.server.namenode.FsckServlet.doGet(FsckServlet.java:44)\n\tat javax.servlet.http.HttpServlet.service(HttpServlet.java:707)\n\tat javax.servlet.http.HttpServlet.service(HttpServlet.java:820)\n\tat org.mortbay.jetty.servlet.ServletHolder.handle(ServletHolder.java:502)\n\tat org.mortbay.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1124)\n\tat org.apache.hadoop.http.HttpServer$QuotingInputFilter.doFilter(HttpServer.java:613)\n\tat org.mortbay.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1115)\n\tat org.mortbay.jetty.servlet.ServletHandler.handle(ServletHandler.java:361)\n\tat org.mortbay.jetty.security.SecurityHandler.handle(SecurityHandler.java:216)\n\tat org.mortbay.jetty.servlet.SessionHandler.handle(SessionHandler.java:181)\n\tat org.mortbay.jetty.handler.ContextHandler.handle(ContextHandler.java:766)\n\tat org.mortbay.jetty.webapp.WebAppContext.handle(WebAppContext.java:417)\n\tat org.mortbay.jetty.handler.ContextHandlerCollection.handle(ContextHandlerCollection.java:230)\n\tat org.mortbay.jetty.handler.HandlerWrapper.handle(HandlerWrapper.java:152)\n\tat org.mortbay.jetty.Server.handle(Server.java:324)\n\tat org.mortbay.jetty.HttpConnection.handleRequest(HttpConnection.java:534)\n\tat org.mortbay.jetty.HttpConnection$RequestHandler.headerComplete(HttpConnection.java:864)\n\tat org.mortbay.jetty.HttpParser.parseNext(HttpParser.java:533)\n\tat org.mortbay.jetty.HttpParser.parseAvailable(HttpParser.java:207)\n\tat org.mortbay.jetty.HttpConnection.handle(HttpConnection.java:403)\n\tat org.mortbay.io.nio.SelectChannelEndPoint.run(SelectChannelEndPoint.java:409)\n\tat org.mortbay.thread.QueuedThreadPool$PoolThread.run(QueuedThreadPool.java:522)\n{noformat}\n\n","issue_id":"12436547","key":"HADOOP-6285","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2009-09-28T21:26:46.000+0000","role":"fixed_distractor","summary":"HttpServer.QuotingInputFilter has the wrong signature for getParameterMap"} {"case_id":"12446495","cluster":"DISTRACTOR-HADOOP-6508","comments":[{"body":"Some simple experiments we have done:\n\n1. We brought up 500 node cluster with CompositeContext containing Contexts C1 and C2. The metric for number of trackers is correct in C1( same number as in web UI) and is wrong in C2.\n2. We brought up 500 node cluster with CompositeContext containing Contexts C2 and C1 (Interchanged the order of contexts from the earlier). Then, the metric for number of trackers is correct in C2( same number as in web UI) and is wrong in C1.\n\nHere, the code path, \"in which the metric for number of trackers is incremented\", is JobTracker.addNewTracker() whenever a new tracker is added. Since first context's value matches with the one on web UI. There is no bug in JobTracker updating code. Also there is no bug in individual context implementations, because it is always second context showing wrong values.\nThus, this leaves there is bug in metrics framework i.e. CompositeContext.\n\n","created":"2010-01-27T04:31:35.097+0000"},{"body":"Analyzing the following JobTrackerMetricsInst code with CompositeContext as the MetricsContext:\n{code}\n MetricsContext context = MetricsUtil.getContext(\"mapred\");\n metricsRecord = MetricsUtil.createRecord(context, \"jobtracker\");\n metricsRecord.setTag(\"sessionId\", sessionId);\n context.registerUpdater(this);\n{code}\n\nDetails on each line of code:\n{code}\n MetricsContext context = MetricsUtil.getContext(\"mapred\");\n{code}\nThis code creates a CompositeContext(CC), which creates all its sub-contexts and calls startsMonitoring on all the\nsubcontext. Thus there are as many threads(monitoring) as the number of sub-contexts. Here, each thread calls\ndoUpdates() followed by emitRecords() in the configured periods.\n\n{code}\n metricsRecord = MetricsUtil.createRecord(context, \"jobtracker\");\n{code}\nThis code creates a MetricsRecord for CompositeContext, which is a Proxy which has a delegator for all the sub-records. This record invokes\nevery method call on all its sub-records.\n\n{code}\n context.registerUpdater(this);\n{code}\nThis code registers JobTracker as the updater for all the sub-contexts.\n\nPutting above relation pictorially (see the attached png): JobTracker (JT) has CompositeContext(CC) and ProxyMetricsRecord (PR). CC starts\nContext1(C1) and Context2(C2) 's timer threads. C1 has MetricsRecord (R1) and C2 has MetricsRecord(R2). PR delagates\nall method calls on it to R1 and R2. C1and C2 register JobTracker as the updater.\n\nBoth C1 and C2 call JT.doUpdates at specified periods. The code flow for doUpdates from C1 or C2:\n 1. Set/Incr methods on ProxyRecord are delegated to both R1 and R2. On the JobTracker code these calls are\nsynchronized.\n 2. MetricsRecord.update() boils down to R1.update() and R2.update() irrespective of whether it is from C1 or C2. This\ncall is not synchronized on JT.\n\nMoreover, MetricsRecord javadoc clearly says: \nDifferent threads should *not* use the same MetricsRecord instance at the same time. . So, the problem here is that the ProxyRecord is shared between two threads without synchronization. There is possibility for race if the above updates are not synchronized. I think this could be the most likely cause for seeing incorrect values with CompositeContext.\n","created":"2010-01-27T05:30:41.849+0000"},{"body":"One solution we thought of is to make CompositeContext a true middle man. CompositeContext registers JobTracker as the updater. Its sub-contexts register CompositeContext as the updater. CompositeContext's MetricRecord is any other record which is updated by JobTracker, and it updates its sub-contexts on the call to doUpdates(). Here, CompositeContext takes care of synchronization for different threads accessing its record. \n\nSee the attached png for the updated relation between JobTracker and CompositeContext. JobTracker (JT) has CompositeContext(CC) and CompositeRecord (CR). CC has a monitoring thread which calls doUpdates periodically, updates CR. CC also starts Context1(C1) and Context2(C2) 's timer threads. C1 has MetricsRecord (R1) and C2 has MetricsRecord(R2). CC updates R1 or R2 from the value of CR, if the doUpdates call is from C1 or C2, respectively. Here, CC registers JT as the updater, C1 and C2 register CC as the updater.\n\nThoughts?","created":"2010-01-27T06:08:57.145+0000"},{"body":"The metrics2 design ensures all the backends see the same metrics for a given snapshot. We tested and deployed the parallel backends with metrics2 in qa and production clusters for more than 8 months. No related issues are found so far.","created":"2011-05-23T19:00:44.889+0000"},{"body":"Should this really be resolved/fixed or is this JIRA in some other state?","created":"2011-06-23T19:07:28.924+0000"}],"conversations":[{"body":"In our clusters, when we use CompositeContext with two contexts, second context gets wrong values.\nThis problem is consistent on 500 (and above) node cluster.","from":"reporter","subject":"Incorrect values for metrics with CompositeContext"},{"body":"Some simple experiments we have done:\n\n1. We brought up 500 node cluster with CompositeContext containing Contexts C1 and C2. The metric for number of trackers is correct in C1( same number as in web UI) and is wrong in C2.\n2. We brought up 500 node cluster with CompositeContext containing Contexts C2 and C1 (Interchanged the order of contexts from the earlier). Then, the metric for number of trackers is correct in C2( same number as in web UI) and is wrong in C1.\n\nHere, the code path, \"in which the metric for number of trackers is incremented\", is JobTracker.addNewTracker() whenever a new tracker is added. Since first context's value matches with the one on web UI. There is no bug in JobTracker updating code. Also there is no bug in individual context implementations, because it is always second context showing wrong values.\nThus, this leaves there is bug in metrics framework i.e. CompositeContext.\n\n","from":"developer"},{"body":"Analyzing the following JobTrackerMetricsInst code with CompositeContext as the MetricsContext:\n{code}\n MetricsContext context = MetricsUtil.getContext(\"mapred\");\n metricsRecord = MetricsUtil.createRecord(context, \"jobtracker\");\n metricsRecord.setTag(\"sessionId\", sessionId);\n context.registerUpdater(this);\n{code}\n\nDetails on each line of code:\n{code}\n MetricsContext context = MetricsUtil.getContext(\"mapred\");\n{code}\nThis code creates a CompositeContext(CC), which creates all its sub-contexts and calls startsMonitoring on all the\nsubcontext. Thus there are as many threads(monitoring) as the number of sub-contexts. Here, each thread calls\ndoUpdates() followed by emitRecords() in the configured periods.\n\n{code}\n metricsRecord = MetricsUtil.createRecord(context, \"jobtracker\");\n{code}\nThis code creates a MetricsRecord for CompositeContext, which is a Proxy which has a delegator for all the sub-records. This record invokes\nevery method call on all its sub-records.\n\n{code}\n context.registerUpdater(this);\n{code}\nThis code registers JobTracker as the updater for all the sub-contexts.\n\nPutting above relation pictorially (see the attached png): JobTracker (JT) has CompositeContext(CC) and ProxyMetricsRecord (PR). CC starts\nContext1(C1) and Context2(C2) 's timer threads. C1 has MetricsRecord (R1) and C2 has MetricsRecord(R2). PR delagates\nall method calls on it to R1 and R2. C1and C2 register JobTracker as the updater.\n\nBoth C1 and C2 call JT.doUpdates at specified periods. The code flow for doUpdates from C1 or C2:\n 1. Set/Incr methods on ProxyRecord are delegated to both R1 and R2. On the JobTracker code these calls are\nsynchronized.\n 2. MetricsRecord.update() boils down to R1.update() and R2.update() irrespective of whether it is from C1 or C2. This\ncall is not synchronized on JT.\n\nMoreover, MetricsRecord javadoc clearly says: \nDifferent threads should *not* use the same MetricsRecord instance at the same time. . So, the problem here is that the ProxyRecord is shared between two threads without synchronization. There is possibility for race if the above updates are not synchronized. I think this could be the most likely cause for seeing incorrect values with CompositeContext.\n","from":"developer"},{"body":"One solution we thought of is to make CompositeContext a true middle man. CompositeContext registers JobTracker as the updater. Its sub-contexts register CompositeContext as the updater. CompositeContext's MetricRecord is any other record which is updated by JobTracker, and it updates its sub-contexts on the call to doUpdates(). Here, CompositeContext takes care of synchronization for different threads accessing its record. \n\nSee the attached png for the updated relation between JobTracker and CompositeContext. JobTracker (JT) has CompositeContext(CC) and CompositeRecord (CR). CC has a monitoring thread which calls doUpdates periodically, updates CR. CC also starts Context1(C1) and Context2(C2) 's timer threads. C1 has MetricsRecord (R1) and C2 has MetricsRecord(R2). CC updates R1 or R2 from the value of CR, if the doUpdates call is from C1 or C2, respectively. Here, CC registers JT as the updater, C1 and C2 register CC as the updater.\n\nThoughts?","from":"developer"},{"body":"The metrics2 design ensures all the backends see the same metrics for a given snapshot. We tested and deployed the parallel backends with metrics2 in qa and production clusters for more than 8 months. No related issues are found so far.","from":"developer"},{"body":"Should this really be resolved/fixed or is this JIRA in some other state?","from":"developer"}],"created":"2010-01-25T03:59:39.000+0000","description":"In our clusters, when we use CompositeContext with two contexts, second context gets wrong values.\nThis problem is consistent on 500 (and above) node cluster.","issue_id":"12446495","key":"HADOOP-6508","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2011-05-23T19:01:25.000+0000","role":"fixed_distractor","summary":"Incorrect values for metrics with CompositeContext"} {"case_id":"12354163","cluster":"DISTRACTOR-HADOOP-652","comments":[{"body":"\nThis will touch addBlock() in FSDataset.java. As part of this I would like the fix the following as well:\n\na) currently subdirectories are created only in the last sub directory (e.g. subdir63). \nb) remove siblings array I think it only increase s recursion in addBlock().\n\n","created":"2006-11-02T22:02:42.000+0000"},{"body":"I create and delete a lot of files in DFS, so I see that speed of DataNode extremely go down!\n\n> a) currently subdirectories are created only in the last sub directory (e.g. subdir63).\n\nNot only subdir63. After the DataNode restarts, subdirectories will be created in other subdirXY, because \"File[] files = dir.listFiles();\" in FSDir constructor lists subdirectories and files in arbitrary order and last subdirectory will be other. Two branches: subdir63 and other one. It is not bug. Other code process this type of tree properly.\n\n> b) remove siblings array I think it only increase s recursion in addBlock().\n\nRecursion is not good idea, because it very slow when DataNode stores a lot of blocks. I think this algorithm should be changed in the future.\n\n\n\nHere is my tested solution of this bug:\n\nNew method clearPath() in FSDir:\n\n> void clearPath( File f ) {\n> if ( dir.compareTo( f ) == 0 ) numBlocks--;\n> else {\n> if ( ( siblings != null ) && ( myIdx != ( siblings.length - 1 ) ) )\n> siblings[ myIdx + 1 ].clearPath( f );\n> else if ( children != null )\n> children[ 0 ].clearPath( f );\n> }\n> }\n\nNew method clearPath() in FSVolume:\n\n> void clearPath( File f ) {\n> dataDir.clearPath( f );\n> }\n\nChanges in invalidate() method in FSDataset:\n\n< blockMap.remove(invalidBlks[i]);\n\n> synchronized ( ongoingCreates ) {\n> blockMap.remove( invalidBlks[ i ] );\n> FSVolume v = volumeMap.get( invalidBlks[ i ] );\n> volumeMap.remove( invalidBlks[ i ] );\n> v.clearPath( f.getParentFile() );\n> }\n\nAnd changes in getFile() method in FSDataset:\n\n< return blockMap.get(b);\n\n> synchronized ( ongoingCreates ) {\n> return blockMap.get( b );\n> }\n\nNow I will try to create patch file properly.\n\nP.S. Also I set dfs.blockreport.intervalMsec = 10000 ( 10 - 30 sec ) in order to prevent lowering NameNode's speed. Because NameNode holds deleted blocks in its datastructures across block reports.\n","created":"2006-11-10T12:17:37.000+0000"},{"body":"diff FSDataset.java.org FSDataset.java.my > FSDataset.java.patch","created":"2006-11-10T13:23:15.000+0000"},{"body":"\nYou should update your code. FSDataset.java has changed recently. Also do a 'diff -u' for submitting a patch.\n\nThis fixes only updating numBlock and does not fix recursion. I am thinking of fixing both (a) and (b) in comment #1 above.\n","created":"2006-11-10T16:56:30.000+0000"},{"body":"Yes. Synchronization in FSDataset already changed. But I think synchronized block in invalidate() method should be smaller, because getFile() call is already synchronized and f.delete() call may be blocked in OS in some situations. So I submit the patch with these changes and with \"-u\" option, as you instruct me.","created":"2006-11-13T10:58:38.000+0000"},{"body":"\nThanks Vladimir.\n\nI have couple of suggestions: \n\nregd locking: synchronizing getFile() and removing from maps seperately is not currect (we will have situation where a block exists in the map but file does not exist). I would suggest moving f.delete() to outside synchronize(). So we delete after removing from maps. \n\nclearPath(): Currently its cost (in compareTo() and in recursion) and proportional to number of directories instead of depth of dir tree. One option to go to the currect child based on the path. On second thoght, you can leave it as it is we will fix it when we fix addBlock() (to remove siblings variable).\n\nFor now you just move f.delete() to outside synchrnoized block and submit the patch.","created":"2006-11-14T23:58:48.000+0000"},{"body":"OK. I just moved f.delete() to outside synchronized block and left recursion as it is. Thanks Raghu.","created":"2006-11-15T13:13:45.000+0000"},{"body":"I just committed this. Thanks Vladimir & Raghu!","created":"2006-11-20T23:12:38.000+0000"}],"conversations":[{"body":"\nCurrently when a block is deleted, DataNode just deletes the physical file and updates its map. We need to update more things. For e.g. numBlocks in FSDir is not decremented.. effect of this would be that we will create more subdirectories than necessary. It might not show up badly yet since numBlocks gets correct value when the dataNode restarts. I have to see what else needs to be updated.\n\n","from":"reporter","subject":"Not all Datastructures are updated when a block is deleted"},{"body":"\nThis will touch addBlock() in FSDataset.java. As part of this I would like the fix the following as well:\n\na) currently subdirectories are created only in the last sub directory (e.g. subdir63). \nb) remove siblings array I think it only increase s recursion in addBlock().\n\n","from":"developer"},{"body":"I create and delete a lot of files in DFS, so I see that speed of DataNode extremely go down!\n\n> a) currently subdirectories are created only in the last sub directory (e.g. subdir63).\n\nNot only subdir63. After the DataNode restarts, subdirectories will be created in other subdirXY, because \"File[] files = dir.listFiles();\" in FSDir constructor lists subdirectories and files in arbitrary order and last subdirectory will be other. Two branches: subdir63 and other one. It is not bug. Other code process this type of tree properly.\n\n> b) remove siblings array I think it only increase s recursion in addBlock().\n\nRecursion is not good idea, because it very slow when DataNode stores a lot of blocks. I think this algorithm should be changed in the future.\n\n\n\nHere is my tested solution of this bug:\n\nNew method clearPath() in FSDir:\n\n> void clearPath( File f ) {\n> if ( dir.compareTo( f ) == 0 ) numBlocks--;\n> else {\n> if ( ( siblings != null ) && ( myIdx != ( siblings.length - 1 ) ) )\n> siblings[ myIdx + 1 ].clearPath( f );\n> else if ( children != null )\n> children[ 0 ].clearPath( f );\n> }\n> }\n\nNew method clearPath() in FSVolume:\n\n> void clearPath( File f ) {\n> dataDir.clearPath( f );\n> }\n\nChanges in invalidate() method in FSDataset:\n\n< blockMap.remove(invalidBlks[i]);\n\n> synchronized ( ongoingCreates ) {\n> blockMap.remove( invalidBlks[ i ] );\n> FSVolume v = volumeMap.get( invalidBlks[ i ] );\n> volumeMap.remove( invalidBlks[ i ] );\n> v.clearPath( f.getParentFile() );\n> }\n\nAnd changes in getFile() method in FSDataset:\n\n< return blockMap.get(b);\n\n> synchronized ( ongoingCreates ) {\n> return blockMap.get( b );\n> }\n\nNow I will try to create patch file properly.\n\nP.S. Also I set dfs.blockreport.intervalMsec = 10000 ( 10 - 30 sec ) in order to prevent lowering NameNode's speed. Because NameNode holds deleted blocks in its datastructures across block reports.\n","from":"developer"},{"body":"diff FSDataset.java.org FSDataset.java.my > FSDataset.java.patch","from":"developer"},{"body":"\nYou should update your code. FSDataset.java has changed recently. Also do a 'diff -u' for submitting a patch.\n\nThis fixes only updating numBlock and does not fix recursion. I am thinking of fixing both (a) and (b) in comment #1 above.\n","from":"developer"},{"body":"Yes. Synchronization in FSDataset already changed. But I think synchronized block in invalidate() method should be smaller, because getFile() call is already synchronized and f.delete() call may be blocked in OS in some situations. So I submit the patch with these changes and with \"-u\" option, as you instruct me.","from":"developer"},{"body":"\nThanks Vladimir.\n\nI have couple of suggestions: \n\nregd locking: synchronizing getFile() and removing from maps seperately is not currect (we will have situation where a block exists in the map but file does not exist). I would suggest moving f.delete() to outside synchronize(). So we delete after removing from maps. \n\nclearPath(): Currently its cost (in compareTo() and in recursion) and proportional to number of directories instead of depth of dir tree. One option to go to the currect child based on the path. On second thoght, you can leave it as it is we will fix it when we fix addBlock() (to remove siblings variable).\n\nFor now you just move f.delete() to outside synchrnoized block and submit the patch.","from":"developer"},{"body":"OK. I just moved f.delete() to outside synchronized block and left recursion as it is. Thanks Raghu.","from":"developer"},{"body":"I just committed this. Thanks Vladimir & Raghu!","from":"developer"}],"created":"2006-10-27T18:49:44.000+0000","description":"\nCurrently when a block is deleted, DataNode just deletes the physical file and updates its map. We need to update more things. For e.g. numBlocks in FSDir is not decremented.. effect of this would be that we will create more subdirectories than necessary. It might not show up badly yet since numBlocks gets correct value when the dataNode restarts. I have to see what else needs to be updated.\n\n","issue_id":"12354163","key":"HADOOP-652","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-11-20T23:12:38.000+0000","role":"fixed_distractor","summary":"Not all Datastructures are updated when a block is deleted"} {"case_id":"12473595","cluster":"DISTRACTOR-HADOOP-6945","comments":[{"body":"+1. And this is a big enough annoyance to back-port to 0.20. It clearly has very small scope of breaking things (HADOOP-6184). I currently remove the jackson jar from the hadoop lib dir on all my nodes (CDH3 0.20).\n\nAs long as so many hadoop jars come first in a task's path, any feature that creates a new jar dependency should be _heavily_ scrutinized before inclusion. ","created":"2010-09-08T16:31:06.444+0000"},{"body":"+1 It's a pain in the butt to have to strip this out of all the nodes in our cluster. Please either remove it or bump the version to a much more recent one. ","created":"2011-05-04T05:38:14.908+0000"},{"body":"given that Jackson is only used to serialize the config, and also that generating valid JSON is fairly straightforward, why not get rid of the jackson dependency at all and just have a json serializer class? jackson (or something like json-lib or gson) could just be used to validate the content in the unit tests.","created":"2011-05-04T14:11:49.271+0000"},{"body":"Which is the status of that ticket? We are having the same trouble. For example, Avro depends on a higher Jackson version, so there are troubles when using it.","created":"2012-03-05T13:10:34.251+0000"},{"body":"This was fixed as part of HADOOP-7606 and HADOOP-7470. ","created":"2012-03-05T15:01:43.917+0000"},{"body":"Both CDH3 and 1.0 break if using Avro 1.6.2. Avro 1.6.3 (currently a release candidate) avoids using the Jackson bits that cause the issue when a Jackson library between 1.3.x and 1.5.x is on the classpath (as with CDH3).\n\nIn our cluster we removed Jackson from all task tracker nodes (yet again, it was OK for a while). One day Hadoop won't put a bunch of unnecessary crap on M/R app classpaths and clearly distinguish libraries that are needed in this context from things that are needed only for the framework. Jackson is one of several libraries with similar issues.","created":"2012-03-05T16:50:52.915+0000"}],"conversations":[{"body":"HADOOP-6184 added the ability to serialize the Configuration to JSON. However its inclusion of jackson-1.0.1 means that any map-reduce task that depends (directly or indirectly) on Jackson, can only use Jackson 1.0.1 APIs. The 100% fix is to give the task's classpath priority over the core/common classpath (MAPREDUCE-1700/MAPREDUCE-1938) so that a job can use newer revisions of the library.\n\nFor a nearer term fix, is it possible to upgrade the included Jackson library to a recent version (1.5.x+)? The APIs are backwards compatible. alernatively, its possible to eliminate the use of Jackson for this minor feature as the Configuration object could be serialized to JSON without a 3rd party library.\n\nAn ancilliary issue is that the jackson.version referenced in ivy.xml is never specified in the source tree (ivy/libraries.properties).","from":"reporter","subject":"Inclusion of Old Jackson-JSON Breaks tasks using Avro (or any task depending on Jackson JSON)"},{"body":"+1. And this is a big enough annoyance to back-port to 0.20. It clearly has very small scope of breaking things (HADOOP-6184). I currently remove the jackson jar from the hadoop lib dir on all my nodes (CDH3 0.20).\n\nAs long as so many hadoop jars come first in a task's path, any feature that creates a new jar dependency should be _heavily_ scrutinized before inclusion. ","from":"developer"},{"body":"+1 It's a pain in the butt to have to strip this out of all the nodes in our cluster. Please either remove it or bump the version to a much more recent one. ","from":"developer"},{"body":"given that Jackson is only used to serialize the config, and also that generating valid JSON is fairly straightforward, why not get rid of the jackson dependency at all and just have a json serializer class? jackson (or something like json-lib or gson) could just be used to validate the content in the unit tests.","from":"developer"},{"body":"Which is the status of that ticket? We are having the same trouble. For example, Avro depends on a higher Jackson version, so there are troubles when using it.","from":"developer"},{"body":"This was fixed as part of HADOOP-7606 and HADOOP-7470. ","from":"developer"},{"body":"Both CDH3 and 1.0 break if using Avro 1.6.2. Avro 1.6.3 (currently a release candidate) avoids using the Jackson bits that cause the issue when a Jackson library between 1.3.x and 1.5.x is on the classpath (as with CDH3).\n\nIn our cluster we removed Jackson from all task tracker nodes (yet again, it was OK for a while). One day Hadoop won't put a bunch of unnecessary crap on M/R app classpaths and clearly distinguish libraries that are needed in this context from things that are needed only for the framework. Jackson is one of several libraries with similar issues.","from":"developer"}],"created":"2010-09-08T15:41:20.000+0000","description":"HADOOP-6184 added the ability to serialize the Configuration to JSON. However its inclusion of jackson-1.0.1 means that any map-reduce task that depends (directly or indirectly) on Jackson, can only use Jackson 1.0.1 APIs. The 100% fix is to give the task's classpath priority over the core/common classpath (MAPREDUCE-1700/MAPREDUCE-1938) so that a job can use newer revisions of the library.\n\nFor a nearer term fix, is it possible to upgrade the included Jackson library to a recent version (1.5.x+)? The APIs are backwards compatible. alernatively, its possible to eliminate the use of Jackson for this minor feature as the Configuration object could be serialized to JSON without a 3rd party library.\n\nAn ancilliary issue is that the jackson.version referenced in ivy.xml is never specified in the source tree (ivy/libraries.properties).","issue_id":"12473595","key":"HADOOP-6945","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2012-03-05T15:03:23.000+0000","role":"fixed_distractor","summary":"Inclusion of Old Jackson-JSON Breaks tasks using Avro (or any task depending on Jackson JSON)"} {"case_id":"12327927","cluster":"DISTRACTOR-HADOOP-7","comments":[{"body":"\n Say! Here's a patch that implements the JobTracker rewrite. Please take a look and let me know what you think....","created":"2006-01-21T05:42:08.000+0000"},{"body":"As Mr Burns would say \"eggcelent\" I'll give this a try. BTW, is it possible to implement functionality that would start jobs that are lagging on nodes that have completed tasks like google does? for example if your 90% done and the last 10 jobs are hung because of bad hardware, slow response or failure and have the ability to redo the long running jobs in parallel on alternate nodes and complete the first one that finishes? this way if you have a huge crawl and certain nodes slow or fail those jobs can be alternated on completed nodes to try and wrap up and terminate any dead jobs when done?\n\nhope that makes sense..","created":"2006-01-21T08:09:08.000+0000"},{"body":"Byron, that's exactly what Mike means by \"speculative execution\".","created":"2006-01-22T12:37:31.000+0000"},{"body":"Guys,\n\nGreg Barish and the folks who worked on the Theseus planning system for information agents at USC did a lot of work on this subject. Theseus is implemented in java, and as I recall, includes the capability to perform speculative execution...\n\nhttp://www.isi.edu/~barish/","created":"2006-01-22T13:56:38.000+0000"},{"body":"so thats what they call it :) thanks","created":"2006-01-22T20:26:15.000+0000"},{"body":"\n One more thing:\n\n You'll see in this patch that ReduceTasks contain a 2D array of map tasks. They used to\ncontain a 1D array, one map task for each map-split of the data.\n\n This 2D business is less than perfect, so let me explain...\n\n Each Reduce task needs to know when its map task predecessors have completed.\nIn the old days, this was easy. There was one map id for each split of the data, so for\nk splits, there were k map ids to know about.\n\n But in a world with speculative execution and early-start of reduce tasks, there could\nbe multiple possible map tasks that work on the same split of data. So for each of the\nk splits, there could be up to M tasks working on it. Thus, the reduce task knows about\nk * M map ids. When any one of the M for each split has completed, the reduce task knows\nit can move on.\n\n But this is a little silly. The JobTracker has a \"TaskInProgress\" abstraction that represents\nthe idea of a \"split's worth of work\". A single TIP contains M map task ids. Instead of the\nreduce task looking all over for map task ids, it should just deal with TIP ids. That way,\nwe're back to a 1D array, and the reduce task code is easier to understand.\n\n Anyway, it will still work as is. I'll improve the code in a future patch. For the moment,\nI'll let this patch stand. I just wanted to let people know...\n","created":"2006-01-23T13:58:06.000+0000"},{"body":"I tested this patch and jobs seem to run into deadlocks when one node crashes while others are loading map output data from that node. Here some lines tasktracker on node B that tries to copy map data from node A which has crashed:\n\n060124 181752 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 181753 task_r_7jjqag 0.17820947% reduce > copy >\n060124 181753 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 181754 task_r_7jjqag 0.17820947% reduce > copy >\n060124 181754 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 181755 task_r_7jjqag 0.17820947% reduce > copy >\n060124 181755 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 181756 task_r_7jjqag 0.17820947% reduce > copy >\n[...]\n060124 223510 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 223511 task_r_7jjqag 0.17820947% reduce > copy >\n060124 223511 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 223512 task_r_7jjqag 0.17820947% reduce > copy >\n060124 223512 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 223513 task_r_7jjqag 0.17820947% reduce > copy >\n060124 223513 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n\nNode A was removed from the jobtracker's node list but it seems like not all tasks depending on that node have been killed.","created":"2006-01-25T06:47:33.000+0000"},{"body":"Mike committed this last week.","created":"2006-02-07T03:38:43.000+0000"},{"body":"This is likely me being dumb, but I don't think this issue is fixed.\n\nWhen I run any of the provided example programs wordcount/grep (also pi with specualtive excecution enabled) reduce tasks does not start before all map tasks have completed.\n\nMy cluster contains three nodes and I am running Hadoop 0.4.0.","created":"2006-07-18T11:43:28.000+0000"}],"conversations":[{"body":"The MapReduce JobTracker is not great at allocating tasks to TaskTracker worker nodes.\n\nHere are the problems:\n1) There is no speculative execution of tasks\n2) Reduce tasks must wait until all map tasks are completed before doing any work\n3) TaskTrackers don't distinguish between Map and Reduce jobs. Also, the number of\ntasks at a single node is limited to some constant. That means you can get weird deadlock\nproblems upon machine failure. The reduces take up all the available execution slots, but they\ndon't do productive work, because they're waiting for a map task to complete. Of course, that\nmap task won't even be started until the reduce tasks finish, so you can see the problem...\n4) The JobTracker is so complicated that it's hard to fix any of these.\n\n\nThe right solution is a rewrite of the JobTracker to be a lot more flexible in task handling.\nIt has to be a lot simpler. One way to make it simpler is to add an abstraction I'll call\n\"TaskInProgress\". Jobs are broken into chunks called TasksInProgress. All the TaskInProgress\nobjects must be complete, somehow, before the Job is complete.\n\nA single TaskInProgress can be executed by one or more Tasks. TaskTrackers are assigned Tasks.\nIf a Task fails, we report it back to the JobTracker, where the TaskInProgress lives. The TIP can then\ndecide whether to launch additional Tasks or not.\n\nSpeculative execution is handled within the TIP. It simply launches multiple Tasks in parallel. The\nTaskTrackers have no idea that these Tasks are actually doing the same chunk of work. The TIP\nis complete when any one of its Tasks are complete.\n\n","from":"reporter","subject":"MapReduce has a series of problems concerning task-allocation to worker nodes"},{"body":"\n Say! Here's a patch that implements the JobTracker rewrite. Please take a look and let me know what you think....","from":"developer"},{"body":"As Mr Burns would say \"eggcelent\" I'll give this a try. BTW, is it possible to implement functionality that would start jobs that are lagging on nodes that have completed tasks like google does? for example if your 90% done and the last 10 jobs are hung because of bad hardware, slow response or failure and have the ability to redo the long running jobs in parallel on alternate nodes and complete the first one that finishes? this way if you have a huge crawl and certain nodes slow or fail those jobs can be alternated on completed nodes to try and wrap up and terminate any dead jobs when done?\n\nhope that makes sense..","from":"developer"},{"body":"Byron, that's exactly what Mike means by \"speculative execution\".","from":"developer"},{"body":"Guys,\n\nGreg Barish and the folks who worked on the Theseus planning system for information agents at USC did a lot of work on this subject. Theseus is implemented in java, and as I recall, includes the capability to perform speculative execution...\n\nhttp://www.isi.edu/~barish/","from":"developer"},{"body":"so thats what they call it :) thanks","from":"developer"},{"body":"\n One more thing:\n\n You'll see in this patch that ReduceTasks contain a 2D array of map tasks. They used to\ncontain a 1D array, one map task for each map-split of the data.\n\n This 2D business is less than perfect, so let me explain...\n\n Each Reduce task needs to know when its map task predecessors have completed.\nIn the old days, this was easy. There was one map id for each split of the data, so for\nk splits, there were k map ids to know about.\n\n But in a world with speculative execution and early-start of reduce tasks, there could\nbe multiple possible map tasks that work on the same split of data. So for each of the\nk splits, there could be up to M tasks working on it. Thus, the reduce task knows about\nk * M map ids. When any one of the M for each split has completed, the reduce task knows\nit can move on.\n\n But this is a little silly. The JobTracker has a \"TaskInProgress\" abstraction that represents\nthe idea of a \"split's worth of work\". A single TIP contains M map task ids. Instead of the\nreduce task looking all over for map task ids, it should just deal with TIP ids. That way,\nwe're back to a 1D array, and the reduce task code is easier to understand.\n\n Anyway, it will still work as is. I'll improve the code in a future patch. For the moment,\nI'll let this patch stand. I just wanted to let people know...\n","from":"developer"},{"body":"I tested this patch and jobs seem to run into deadlocks when one node crashes while others are loading map output data from that node. Here some lines tasktracker on node B that tries to copy map data from node A which has crashed:\n\n060124 181752 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 181753 task_r_7jjqag 0.17820947% reduce > copy >\n060124 181753 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 181754 task_r_7jjqag 0.17820947% reduce > copy >\n060124 181754 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 181755 task_r_7jjqag 0.17820947% reduce > copy >\n060124 181755 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 181756 task_r_7jjqag 0.17820947% reduce > copy >\n[...]\n060124 223510 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 223511 task_r_7jjqag 0.17820947% reduce > copy >\n060124 223511 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 223512 task_r_7jjqag 0.17820947% reduce > copy >\n060124 223512 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n060124 223513 task_r_7jjqag 0.17820947% reduce > copy >\n060124 223513 task_r_27r56x 0.2212838% reduce > copy > task_m_1ujusx@nodeA:50040\n\nNode A was removed from the jobtracker's node list but it seems like not all tasks depending on that node have been killed.","from":"developer"},{"body":"Mike committed this last week.","from":"developer"},{"body":"This is likely me being dumb, but I don't think this issue is fixed.\n\nWhen I run any of the provided example programs wordcount/grep (also pi with specualtive excecution enabled) reduce tasks does not start before all map tasks have completed.\n\nMy cluster contains three nodes and I am running Hadoop 0.4.0.","from":"developer"}],"created":"2006-01-21T05:40:53.000+0000","description":"The MapReduce JobTracker is not great at allocating tasks to TaskTracker worker nodes.\n\nHere are the problems:\n1) There is no speculative execution of tasks\n2) Reduce tasks must wait until all map tasks are completed before doing any work\n3) TaskTrackers don't distinguish between Map and Reduce jobs. Also, the number of\ntasks at a single node is limited to some constant. That means you can get weird deadlock\nproblems upon machine failure. The reduces take up all the available execution slots, but they\ndon't do productive work, because they're waiting for a map task to complete. Of course, that\nmap task won't even be started until the reduce tasks finish, so you can see the problem...\n4) The JobTracker is so complicated that it's hard to fix any of these.\n\n\nThe right solution is a rewrite of the JobTracker to be a lot more flexible in task handling.\nIt has to be a lot simpler. One way to make it simpler is to add an abstraction I'll call\n\"TaskInProgress\". Jobs are broken into chunks called TasksInProgress. All the TaskInProgress\nobjects must be complete, somehow, before the Job is complete.\n\nA single TaskInProgress can be executed by one or more Tasks. TaskTrackers are assigned Tasks.\nIf a Task fails, we report it back to the JobTracker, where the TaskInProgress lives. The TIP can then\ndecide whether to launch additional Tasks or not.\n\nSpeculative execution is handled within the TIP. It simply launches multiple Tasks in parallel. The\nTaskTrackers have no idea that these Tasks are actually doing the same chunk of work. The TIP\nis complete when any one of its Tasks are complete.\n\n","issue_id":"12327927","key":"HADOOP-7","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-02-07T03:38:43.000+0000","role":"fixed_distractor","summary":"MapReduce has a series of problems concerning task-allocation to worker nodes"} {"case_id":"12478209","cluster":"DISTRACTOR-HADOOP-7006","comments":[{"body":"It appears that this was introduced with changeset 949658. Here is a partial diff of org.apache.hadoop.fs.FileUtil:\n\n{noformat}\n@@ -258,7 +258,7 @@\n Configuration conf, String addString) throws IOException {\n dstFile = checkDest(srcDir.getName(), dstFS, dstFile, false);\n \n- if (!srcFS.getFileStatus(srcDir).isDir())\n+ if (srcFS.getFileStatus(srcDir).isDirectory())\n return false;\n{noformat}\n\nNotice that in addition to switching from isDir() to isDirectory(), this change also dropped the negation on the front of the condition. I'll attach a simple one-line patch to restore functionality. I've also added a unit test to cover the FileUtil.copyMerge API. I'd like to volunteer for HADOOP-6387 next, and I want this unit test in place as a regression test while I work on that.\n","created":"2010-10-25T06:48:46.189+0000"},{"body":"I've attached the patch. Could I please have a code review? Thank you.","created":"2010-10-25T06:51:47.268+0000"},{"body":"I've reviewed the patch, and it looks good to me.","created":"2010-10-25T16:39:34.991+0000"},{"body":"Tests failed for me until I also changed the file merge to sort the files. This is probably a good idea anyway. Here's a new version of the patch with that change.","created":"2010-10-25T19:43:56.286+0000"},{"body":"Thanks, Aaron and Doug.\n\nI'm following the directions from http://wiki.apache.org/hadoop/HowToContribute . I don't see the \"Submit Patch\" link to trigger Hudson though. Is there something else that I need to do?\n","created":"2010-10-26T17:11:09.053+0000"},{"body":"Doug already marked the ticket patch available.\n\nNot quite sure why Hudson hasn't picked it up yet - might just be taking some time to get around to it (the HDFS tests take quite a while to run.)","created":"2010-10-26T17:21:25.025+0000"},{"body":"Ran test-patch by hand.\n\n{code}\n-1 overall. \n\n +1 @author. The patch does not contain any @author tags.\n\n +1 tests included. The patch appears to include 3 new or modified tests.\n\n -1 javadoc. The javadoc tool appears to have generated 1 warning messages.\n\n +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n\n +1 findbugs. The patch does not introduce any new Findbugs warnings.\n\n +1 release audit. The applied patch does not increase the total number of release audit warnings.\n\n +1 system tests framework. The patch passed system tests framework compile.\n{code}\n\nThe javadoc test seems broken, as it fails even when I run test-patch with an empty patch file.","created":"2010-10-26T19:29:31.626+0000"},{"body":"I just committed this. Thanks, Chris!","created":"2010-10-26T21:16:34.715+0000"}],"conversations":[{"body":"Running the codebase from trunk, the hadoop fs -getmerge command does not work. As implemented in prior versions (i.e. 0.20.2), I could run hadoop fs -getmerge pointed at a directory containing multiple files. It would merge all files into a single file on the local file system. Running the same command using the codebase from trunk, it looks like nothing happens.\n","from":"reporter","subject":"hadoop fs -getmerge does not work using codebase from trunk."},{"body":"It appears that this was introduced with changeset 949658. Here is a partial diff of org.apache.hadoop.fs.FileUtil:\n\n{noformat}\n@@ -258,7 +258,7 @@\n Configuration conf, String addString) throws IOException {\n dstFile = checkDest(srcDir.getName(), dstFS, dstFile, false);\n \n- if (!srcFS.getFileStatus(srcDir).isDir())\n+ if (srcFS.getFileStatus(srcDir).isDirectory())\n return false;\n{noformat}\n\nNotice that in addition to switching from isDir() to isDirectory(), this change also dropped the negation on the front of the condition. I'll attach a simple one-line patch to restore functionality. I've also added a unit test to cover the FileUtil.copyMerge API. I'd like to volunteer for HADOOP-6387 next, and I want this unit test in place as a regression test while I work on that.\n","from":"developer"},{"body":"I've attached the patch. Could I please have a code review? Thank you.","from":"developer"},{"body":"I've reviewed the patch, and it looks good to me.","from":"developer"},{"body":"Tests failed for me until I also changed the file merge to sort the files. This is probably a good idea anyway. Here's a new version of the patch with that change.","from":"developer"},{"body":"Thanks, Aaron and Doug.\n\nI'm following the directions from http://wiki.apache.org/hadoop/HowToContribute . I don't see the \"Submit Patch\" link to trigger Hudson though. Is there something else that I need to do?\n","from":"developer"},{"body":"Doug already marked the ticket patch available.\n\nNot quite sure why Hudson hasn't picked it up yet - might just be taking some time to get around to it (the HDFS tests take quite a while to run.)","from":"developer"},{"body":"Ran test-patch by hand.\n\n{code}\n-1 overall. \n\n +1 @author. The patch does not contain any @author tags.\n\n +1 tests included. The patch appears to include 3 new or modified tests.\n\n -1 javadoc. The javadoc tool appears to have generated 1 warning messages.\n\n +1 javac. The applied patch does not increase the total number of javac compiler warnings.\n\n +1 findbugs. The patch does not introduce any new Findbugs warnings.\n\n +1 release audit. The applied patch does not increase the total number of release audit warnings.\n\n +1 system tests framework. The patch passed system tests framework compile.\n{code}\n\nThe javadoc test seems broken, as it fails even when I run test-patch with an empty patch file.","from":"developer"},{"body":"I just committed this. Thanks, Chris!","from":"developer"}],"created":"2010-10-25T06:40:02.000+0000","description":"Running the codebase from trunk, the hadoop fs -getmerge command does not work. As implemented in prior versions (i.e. 0.20.2), I could run hadoop fs -getmerge pointed at a directory containing multiple files. It would merge all files into a single file on the local file system. Running the same command using the codebase from trunk, it looks like nothing happens.\n","issue_id":"12478209","key":"HADOOP-7006","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2010-10-26T21:16:34.000+0000","role":"fixed_distractor","summary":"hadoop fs -getmerge does not work using codebase from trunk."} {"case_id":"12330009","cluster":"DISTRACTOR-HADOOP-71","comments":[{"body":"Here is a patch that adds the FileSystem as an argument to the constructor.","created":"2006-03-09T08:19:01.000+0000"},{"body":"Doug,\n This one seems to have dropped off your radar. The fix isn't mandatory any more, because we created a custom config for the junit tests. However, it still seems cleaner to use the provided FileSystem rather than fetching the default one.","created":"2006-03-23T01:05:55.000+0000"},{"body":"This looks good, although it makes an incompatible changes to a public APIs used by Nutch. So we should probably continue to support the old method signatures too, or else make a co-ordinated change to Nutch.","created":"2006-03-23T03:35:09.000+0000"},{"body":"This patch also fixes the TextInputFormat and the framework's handling of the output path too. I will file a follow-up but to get deprecate the current FileSystem.get(Configuration) method and replace it like FileSystem.getDefault(Configuration). There are many place still left in the code where the default file system is gotten. However, with this patch, I was able to run map/reduce jobs (including a unit test) with the input and output going to non-default file systems. ","created":"2007-06-06T05:54:10.124+0000"},{"body":"Owen, this will break InputFormats and OutputFormats that use the FileSystem passed to getRecordReader() and getRecordWriter(), since null is now passed. Is that intended? Is it now an error to touch that parameter? If so, we should at least provide some warning in the release notes, no? It'd be nicer to have a new API and then deprecate this one...","created":"2007-06-14T22:00:22.392+0000"},{"body":"*sigh*\n\nI guess it is a bit too aggressive. The point of doing that is that any OutputFormats who use that FileSystem are broken in exactly the way fixed by this patch. Furthermore, the parameter has been named 'ignored' for many releases. (Only OutputFormats were given nulls, because InputFormats don't take a FileSystem. I think I removed the InputFormat FileSystem parameter a long time ago.)","created":"2007-06-18T06:13:02.027+0000"},{"body":"This patch no longer applies to trunk, and I still think we need to make it more back-compatible.","created":"2007-06-20T18:57:10.210+0000"},{"body":"Updated to trunk and continued passing in the default fs.","created":"2007-08-09T22:51:23.345+0000"},{"body":"Realized that one of the changes wasn't needed.","created":"2007-08-09T23:01:28.145+0000"},{"body":"Sorry for the lateness on this one. It had fallen off my radar as being bounced. We need this one.","created":"2007-08-09T23:02:17.682+0000"},{"body":"+1","created":"2007-08-10T18:24:20.968+0000"},{"body":"I just committed this.","created":"2007-08-10T20:26:25.226+0000"}],"conversations":[{"body":"The mapred.TestSequenceFileInputFormat test was failing when run with a conf directory that pointed to a dfs cluster. The reason was that SequenceFileRecordReader was using the default FileSystem from the config, while the test program was assuming the \"local\" file system was being used.","from":"reporter","subject":"The SequenceFileRecordReader uses the default FileSystem rather than the supplied one"},{"body":"Here is a patch that adds the FileSystem as an argument to the constructor.","from":"developer"},{"body":"Doug,\n This one seems to have dropped off your radar. The fix isn't mandatory any more, because we created a custom config for the junit tests. However, it still seems cleaner to use the provided FileSystem rather than fetching the default one.","from":"developer"},{"body":"This looks good, although it makes an incompatible changes to a public APIs used by Nutch. So we should probably continue to support the old method signatures too, or else make a co-ordinated change to Nutch.","from":"developer"},{"body":"This patch also fixes the TextInputFormat and the framework's handling of the output path too. I will file a follow-up but to get deprecate the current FileSystem.get(Configuration) method and replace it like FileSystem.getDefault(Configuration). There are many place still left in the code where the default file system is gotten. However, with this patch, I was able to run map/reduce jobs (including a unit test) with the input and output going to non-default file systems. ","from":"developer"},{"body":"Owen, this will break InputFormats and OutputFormats that use the FileSystem passed to getRecordReader() and getRecordWriter(), since null is now passed. Is that intended? Is it now an error to touch that parameter? If so, we should at least provide some warning in the release notes, no? It'd be nicer to have a new API and then deprecate this one...","from":"developer"},{"body":"*sigh*\n\nI guess it is a bit too aggressive. The point of doing that is that any OutputFormats who use that FileSystem are broken in exactly the way fixed by this patch. Furthermore, the parameter has been named 'ignored' for many releases. (Only OutputFormats were given nulls, because InputFormats don't take a FileSystem. I think I removed the InputFormat FileSystem parameter a long time ago.)","from":"developer"},{"body":"This patch no longer applies to trunk, and I still think we need to make it more back-compatible.","from":"developer"},{"body":"Updated to trunk and continued passing in the default fs.","from":"developer"},{"body":"Realized that one of the changes wasn't needed.","from":"developer"},{"body":"Sorry for the lateness on this one. It had fallen off my radar as being bounced. We need this one.","from":"developer"},{"body":"+1","from":"developer"},{"body":"I just committed this.","from":"developer"}],"created":"2006-03-09T08:18:06.000+0000","description":"The mapred.TestSequenceFileInputFormat test was failing when run with a conf directory that pointed to a dfs cluster. The reason was that SequenceFileRecordReader was using the default FileSystem from the config, while the test program was assuming the \"local\" file system was being used.","issue_id":"12330009","key":"HADOOP-71","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2007-08-10T20:26:25.000+0000","role":"fixed_distractor","summary":"The SequenceFileRecordReader uses the default FileSystem rather than the supplied one"} {"case_id":"12514139","cluster":"DISTRACTOR-HADOOP-7461","comments":[{"body":"I've upgraded this to Major in the hope that someone will pay attention. This bug still affects 0.20.205.0, four months after the bug was filed. It causes total failure when running in non-distributed mode. The fix is trivial for whoever manages the POM -- just add the missing dependency! It's a bad sign that a bug like this has persisted this long.\n","created":"2011-11-03T03:06:37.071+0000"},{"body":"updated the core pom template to include the jackson dependency.","created":"2011-12-14T17:43:39.943+0000"},{"body":"Two comments:\nFirst, to request/propose that a bug be fixed in a given release, please add that release to the \"Target Version\" field in the bug report.\nSecond, if the fix is trivial, please construct a patch for the fix, test it, and upload it to the Jira, rather than complaining that no one (else) has fixed it in four months.","created":"2011-12-14T17:47:00.243+0000"},{"body":"Giri, two typos:\n+ jackson-mapper-asl\nshould be \n+ jackson-mapper-asl\n\nand the closing \n+ \nshould be\n+ \n","created":"2011-12-14T17:51:59.944+0000"},{"body":"mybad , fixed the typos.\n\ntested the patch by doing mvn-install and verified if the pom includes the jackson dependency.\n\nsteps:\nant -Dcompile.native=true -Dcompile.c++=true mvn-install\ncheck the mvn cache \n~/.m2/repository/org/apache/hadoop/hadoop-core/1.0.1-SNAPSHOT/hadoop-core-1.0.1-SNAPSHOT.pom","created":"2011-12-14T18:52:03.496+0000"},{"body":"+1. Please commit to branch-1.0 and branch-1. Thanks.","created":"2011-12-14T18:56:11.894+0000"},{"body":"committed to both 1 and 1.0 branches, thanks Matt","created":"2011-12-14T19:18:34.016+0000"},{"body":"Closed upon release of version 1.0.0.","created":"2011-12-28T10:03:32.220+0000"},{"body":"This was indeed fixed in 1.0.0, so no need to note it as fixed in 1.1.0.","created":"2012-08-24T00:05:26.939+0000"}],"conversations":[{"body":"(COMMENT: This bug still affects 0.20.205.0, four months after the bug was filed. This causes total failure, and the fix is trivial for whoever manages the POM -- just add the missing dependency! --ben)\n\nThis issue was identified and the fix & workaround was documented at \n\nhttps://issues.cloudera.org/browse/DISTRO-44\n\nThe issue affects use of Hadoop 0.20.203.0 from the Maven central repo. I built a job using that maven repo and ran it, resulting in this failure:\n\nException in thread \"main\" java.lang.NoClassDefFoundError: org/codehaus/jackson/map/JsonMappingException\n\tat thinkbig.hadoop.inputformat.TestXmlInputFormat.run(TestXmlInputFormat.java:18)\n\tat thinkbig.hadoop.inputformat.TestXmlInputFormat.main(TestXmlInputFormat.java:23)\nCaused by: java.lang.ClassNotFoundException: org.codehaus.jackson.map.JsonMappingException\n\n\n\n","from":"reporter","subject":"Jackson Dependency Not Declared in Hadoop POM"},{"body":"I've upgraded this to Major in the hope that someone will pay attention. This bug still affects 0.20.205.0, four months after the bug was filed. It causes total failure when running in non-distributed mode. The fix is trivial for whoever manages the POM -- just add the missing dependency! It's a bad sign that a bug like this has persisted this long.\n","from":"developer"},{"body":"updated the core pom template to include the jackson dependency.","from":"developer"},{"body":"Two comments:\nFirst, to request/propose that a bug be fixed in a given release, please add that release to the \"Target Version\" field in the bug report.\nSecond, if the fix is trivial, please construct a patch for the fix, test it, and upload it to the Jira, rather than complaining that no one (else) has fixed it in four months.","from":"developer"},{"body":"Giri, two typos:\n+ jackson-mapper-asl\nshould be \n+ jackson-mapper-asl\n\nand the closing \n+ \nshould be\n+ \n","from":"developer"},{"body":"mybad , fixed the typos.\n\ntested the patch by doing mvn-install and verified if the pom includes the jackson dependency.\n\nsteps:\nant -Dcompile.native=true -Dcompile.c++=true mvn-install\ncheck the mvn cache \n~/.m2/repository/org/apache/hadoop/hadoop-core/1.0.1-SNAPSHOT/hadoop-core-1.0.1-SNAPSHOT.pom","from":"developer"},{"body":"+1. Please commit to branch-1.0 and branch-1. Thanks.","from":"developer"},{"body":"committed to both 1 and 1.0 branches, thanks Matt","from":"developer"},{"body":"Closed upon release of version 1.0.0.","from":"developer"},{"body":"This was indeed fixed in 1.0.0, so no need to note it as fixed in 1.1.0.","from":"developer"}],"created":"2011-07-15T00:12:34.000+0000","description":"(COMMENT: This bug still affects 0.20.205.0, four months after the bug was filed. This causes total failure, and the fix is trivial for whoever manages the POM -- just add the missing dependency! --ben)\n\nThis issue was identified and the fix & workaround was documented at \n\nhttps://issues.cloudera.org/browse/DISTRO-44\n\nThe issue affects use of Hadoop 0.20.203.0 from the Maven central repo. I built a job using that maven repo and ran it, resulting in this failure:\n\nException in thread \"main\" java.lang.NoClassDefFoundError: org/codehaus/jackson/map/JsonMappingException\n\tat thinkbig.hadoop.inputformat.TestXmlInputFormat.run(TestXmlInputFormat.java:18)\n\tat thinkbig.hadoop.inputformat.TestXmlInputFormat.main(TestXmlInputFormat.java:23)\nCaused by: java.lang.ClassNotFoundException: org.codehaus.jackson.map.JsonMappingException\n\n\n\n","issue_id":"12514139","key":"HADOOP-7461","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2011-12-14T19:18:33.000+0000","role":"fixed_distractor","summary":"Jackson Dependency Not Declared in Hadoop POM"} {"case_id":"12531158","cluster":"DISTRACTOR-HADOOP-7817","comments":[{"body":"patch forthcoming","created":"2011-11-11T17:49:00.449+0000"},{"body":"Yes, the position will come as zero, because it directly opens the FileOutPutstream in append mode. This will ensures that the data to go at the end of the file.\nfixing this would be simple, we can set the position on fileChannel position to the length of the file. But i am not sure this is really a bug or not, because the author intention is to leave as it is equal behavior with FileOutPutStream when opening in append mode....?","created":"2011-11-11T18:40:20.443+0000"},{"body":"\nFWIW: \nDFSClient.append does:\n\n return new FSDataOutputStream(out, statistics, out.getInitialLen());\n\nWas planning on submitting a patch on RawLocalFileSystem w/\n\n public FSDataOutputStream append(Path f, int bufferSize,\n Progressable progress) throws IOException {\n if (!exists(f)) {\n throw new FileNotFoundException(\"File \" + f + \" not found\");\n }\n FileStatus status = getFileStatus(f);\n if (status.isDirectory()) {\n throw new IOException(\"Cannot append to a diretory (=\" + f + \" )\");\n }\n return new FSDataOutputStream(new BufferedOutputStream(\n new LocalFSFileOutputStream(f, true), bufferSize), statistics,status.getLen());\n }","created":"2011-11-11T19:15:51.908+0000"},{"body":"I think the -1 javadoc mentioned above is not related to my patch.","created":"2011-12-07T00:47:56.150+0000"},{"body":"Is anyone looking at this? I have another patch that depends on this and would like to get some feedback.\n\nThanks!","created":"2012-01-03T21:37:48.729+0000"},{"body":"Cancelling patch, as it no longer applies to trunk.","created":"2015-03-11T23:39:18.781+0000"},{"body":"Assigning to myself to update the patch. [~ktomasette] , please feel free to assign back if you would like to continue work on this JIRA","created":"2015-05-28T06:02:19.338+0000"},{"body":"Updated patch on the latest trunk code base.","created":"2015-06-03T10:09:33.670+0000"},{"body":"Thanks for updating the patch [~kanaka]. \nSource changes looks good.\n\nHave some nits in test.\n1. {{fs}} is not closed after test.\n2. {{in}} also not closed.\n3. Use default value 'build/test/data' for {{System.getProperty(\"test.build.data\", \".\")}}","created":"2015-06-09T09:47:23.262+0000"},{"body":"4. {{assertEquals(pos, out.getPos());}}, better this be asserted against hardcoded value 4. \n5. {code}+ StringBuffer buff = new StringBuffer();\n+ FSDataInputStream in = fs.open(path1);\n+ int c;\n+ while ((c = in.read()) != -1) {\n+ buff.append((char) c);\n+ }\n+\n+ //Verify the content\n+ assertEquals(\"text1 text2\", buff.toString());{code}\nAbove one could be also written as \n{code} FSDataInputStream in = fs.open(path1);\n byte[] buf = new byte[in.available()];\n in.read(buf);\n // Verify the content\n assertEquals(\"text1 text2\", new String(buf));{code}","created":"2015-06-09T09:58:25.842+0000"},{"body":"I think, test can be added in existing {{TestLocalFileSystem}} class itself instead of creating new file.","created":"2015-06-09T10:00:22.223+0000"},{"body":"Thanks for the review [~vinayrpet]. I have moved the test code to {{TestLocalFileSystem}} and handled the comments except the first one to close the fs. \n\nWe can close {{FileSystem}} by making {{fileSys}} static in a method with annotation {{@AfterClass}} but I felt it may be un related change for this JIRA context as it's an old file and some other test classes for example {{TestCryptoStreamsForLocalFS}} are also not closing localFS. \nSo, we can file a separate JIRA to handle this in all such files. What is your opinion?\n\n","created":"2015-06-09T13:53:27.549+0000"},{"body":"bq. We can close FileSystem by making fileSys static in a method with annotation @AfterClass but I felt it may be un related change for this JIRA context as it's an old file and some other test classes for example TestCryptoStreamsForLocalFS are also not closing localFS. \nSo, we can file a separate JIRA to handle this in all such files. What is your opinion?\nSince the whole test is moved to {{TestLocalFileSystem}}, this change is not needed for local fs.\nAnyway, my main concern was the closing of open streams, thats handled in latest patch. So I am fine with that.\n\nLatest patch looks good to me. +1.\nWill commit soon","created":"2015-06-10T05:25:26.857+0000"},{"body":"Committed to trunk and branch-2.\nThanks for the contribution [~ktomasette] and [~kanaka].","created":"2015-06-10T05:41:00.708+0000"},{"body":"Cherry-picked this to branch-2.7 and branch-2.6 to Fix {{TestSequenceFileAppend}} from HADOOP-7139\n\nThere were minor conflicts. Attached committed patches for reference.","created":"2016-04-06T02:57:00.466+0000"},{"body":"Closing the JIRA as part of 2.7.3 release.","created":"2016-08-25T22:48:40.742+0000"}],"conversations":[{"body":"When RawLocalFileSyste.append() is called it returns an FSDataOutputStream whose .getPos() returns 0.\ngetPos() should return position in the file where appends will start writing.\n\n","from":"reporter","subject":"RawLocalFileSystem.append() should give FSDataOutputStream with accurate .getPos()"},{"body":"patch forthcoming","from":"developer"},{"body":"Yes, the position will come as zero, because it directly opens the FileOutPutstream in append mode. This will ensures that the data to go at the end of the file.\nfixing this would be simple, we can set the position on fileChannel position to the length of the file. But i am not sure this is really a bug or not, because the author intention is to leave as it is equal behavior with FileOutPutStream when opening in append mode....?","from":"developer"},{"body":"\nFWIW: \nDFSClient.append does:\n\n return new FSDataOutputStream(out, statistics, out.getInitialLen());\n\nWas planning on submitting a patch on RawLocalFileSystem w/\n\n public FSDataOutputStream append(Path f, int bufferSize,\n Progressable progress) throws IOException {\n if (!exists(f)) {\n throw new FileNotFoundException(\"File \" + f + \" not found\");\n }\n FileStatus status = getFileStatus(f);\n if (status.isDirectory()) {\n throw new IOException(\"Cannot append to a diretory (=\" + f + \" )\");\n }\n return new FSDataOutputStream(new BufferedOutputStream(\n new LocalFSFileOutputStream(f, true), bufferSize), statistics,status.getLen());\n }","from":"developer"},{"body":"I think the -1 javadoc mentioned above is not related to my patch.","from":"developer"},{"body":"Is anyone looking at this? I have another patch that depends on this and would like to get some feedback.\n\nThanks!","from":"developer"},{"body":"Cancelling patch, as it no longer applies to trunk.","from":"developer"},{"body":"Assigning to myself to update the patch. [~ktomasette] , please feel free to assign back if you would like to continue work on this JIRA","from":"developer"},{"body":"Updated patch on the latest trunk code base.","from":"developer"},{"body":"Thanks for updating the patch [~kanaka]. \nSource changes looks good.\n\nHave some nits in test.\n1. {{fs}} is not closed after test.\n2. {{in}} also not closed.\n3. Use default value 'build/test/data' for {{System.getProperty(\"test.build.data\", \".\")}}","from":"developer"},{"body":"4. {{assertEquals(pos, out.getPos());}}, better this be asserted against hardcoded value 4. \n5. {code}+ StringBuffer buff = new StringBuffer();\n+ FSDataInputStream in = fs.open(path1);\n+ int c;\n+ while ((c = in.read()) != -1) {\n+ buff.append((char) c);\n+ }\n+\n+ //Verify the content\n+ assertEquals(\"text1 text2\", buff.toString());{code}\nAbove one could be also written as \n{code} FSDataInputStream in = fs.open(path1);\n byte[] buf = new byte[in.available()];\n in.read(buf);\n // Verify the content\n assertEquals(\"text1 text2\", new String(buf));{code}","from":"developer"},{"body":"I think, test can be added in existing {{TestLocalFileSystem}} class itself instead of creating new file.","from":"developer"},{"body":"Thanks for the review [~vinayrpet]. I have moved the test code to {{TestLocalFileSystem}} and handled the comments except the first one to close the fs. \n\nWe can close {{FileSystem}} by making {{fileSys}} static in a method with annotation {{@AfterClass}} but I felt it may be un related change for this JIRA context as it's an old file and some other test classes for example {{TestCryptoStreamsForLocalFS}} are also not closing localFS. \nSo, we can file a separate JIRA to handle this in all such files. What is your opinion?\n\n","from":"developer"},{"body":"bq. We can close FileSystem by making fileSys static in a method with annotation @AfterClass but I felt it may be un related change for this JIRA context as it's an old file and some other test classes for example TestCryptoStreamsForLocalFS are also not closing localFS. \nSo, we can file a separate JIRA to handle this in all such files. What is your opinion?\nSince the whole test is moved to {{TestLocalFileSystem}}, this change is not needed for local fs.\nAnyway, my main concern was the closing of open streams, thats handled in latest patch. So I am fine with that.\n\nLatest patch looks good to me. +1.\nWill commit soon","from":"developer"},{"body":"Committed to trunk and branch-2.\nThanks for the contribution [~ktomasette] and [~kanaka].","from":"developer"},{"body":"Cherry-picked this to branch-2.7 and branch-2.6 to Fix {{TestSequenceFileAppend}} from HADOOP-7139\n\nThere were minor conflicts. Attached committed patches for reference.","from":"developer"},{"body":"Closing the JIRA as part of 2.7.3 release.","from":"developer"}],"created":"2011-11-11T17:48:39.000+0000","description":"When RawLocalFileSyste.append() is called it returns an FSDataOutputStream whose .getPos() returns 0.\ngetPos() should return position in the file where appends will start writing.\n\n","issue_id":"12531158","key":"HADOOP-7817","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2015-06-10T05:41:00.000+0000","role":"fixed_distractor","summary":"RawLocalFileSystem.append() should give FSDataOutputStream with accurate .getPos()"} {"case_id":"12534976","cluster":"DISTRACTOR-HADOOP-7917","comments":[{"body":"changing script that compiles proto files to handle cywin and windows version of protoc which does not handle UNIX style paths.\n\nNote that hadoop-yarn-common is failing in windows but this is not related to the proto compilation but to the saveVersion.sh script, MAPREDUCE-3540 .","created":"2011-12-13T16:37:10.508+0000"},{"body":"patch tested on win7, osx, ubuntu","created":"2011-12-13T16:38:31.490+0000"},{"body":"Just verified the patch, it works fine for me except mapred compilation failure due to saveVersion.sh problem.\nmvn eclipse:eclipse also works fine for common and hdfs. Also imported into eclipse. \n\nThanks Alejandro, for the patch. Looks good to me.\n\n+1 from my side.\n\nWhy patch command failed in Jenkins? do you have idea? ( is it because of cross projects? )","created":"2011-12-13T17:43:30.539+0000"},{"body":"+1 the patch looks good.","created":"2011-12-15T07:31:56.118+0000"},{"body":"Thanks for the patch Alejandro, it works. ","created":"2011-12-15T10:28:55.101+0000"},{"body":"committed to trunk and branch-0.23","created":"2011-12-15T14:47:50.365+0000"},{"body":"Hi Alejandro,\n\nlooks CHANGES.txt has typo mistake.\n\n{quote}HADOOP_7917. compilation of protobuf files fails in windows/cygwin. (tucu){quote}\n\nissue id should have \"-\" instead of \"_\".\n\nThanks\nUma","created":"2011-12-19T16:29:06.572+0000"},{"body":"Uma,\n\nThanks for pointing this out, I'll piggyback on the next commit I do to fix the typo.\n\nAlejandro","created":"2011-12-19T17:46:29.954+0000"},{"body":"Hello, I downloaded 0.24.0 and tried to build under Windows/Cygwin.\nProto file compilation still seems to fail. I'm running mvn from within Cygwin though, not windows command prompt.\nAnd if I switch around the true and false logic for IS_WIN variable, it works fine.\n\nIs this by design?? Your clarification would be much helpful!! :)","created":"2012-03-16T06:00:23.457+0000"}],"conversations":[{"body":"HADOOP-7899 & HDFS-2511 introduced compilation of proto files as part of the build.\n\nSuch compilation is failing in windows/cygwin","from":"reporter","subject":"compilation of protobuf files fails in windows/cygwin"},{"body":"changing script that compiles proto files to handle cywin and windows version of protoc which does not handle UNIX style paths.\n\nNote that hadoop-yarn-common is failing in windows but this is not related to the proto compilation but to the saveVersion.sh script, MAPREDUCE-3540 .","from":"developer"},{"body":"patch tested on win7, osx, ubuntu","from":"developer"},{"body":"Just verified the patch, it works fine for me except mapred compilation failure due to saveVersion.sh problem.\nmvn eclipse:eclipse also works fine for common and hdfs. Also imported into eclipse. \n\nThanks Alejandro, for the patch. Looks good to me.\n\n+1 from my side.\n\nWhy patch command failed in Jenkins? do you have idea? ( is it because of cross projects? )","from":"developer"},{"body":"+1 the patch looks good.","from":"developer"},{"body":"Thanks for the patch Alejandro, it works. ","from":"developer"},{"body":"committed to trunk and branch-0.23","from":"developer"},{"body":"Hi Alejandro,\n\nlooks CHANGES.txt has typo mistake.\n\n{quote}HADOOP_7917. compilation of protobuf files fails in windows/cygwin. (tucu){quote}\n\nissue id should have \"-\" instead of \"_\".\n\nThanks\nUma","from":"developer"},{"body":"Uma,\n\nThanks for pointing this out, I'll piggyback on the next commit I do to fix the typo.\n\nAlejandro","from":"developer"},{"body":"Hello, I downloaded 0.24.0 and tried to build under Windows/Cygwin.\nProto file compilation still seems to fail. I'm running mvn from within Cygwin though, not windows command prompt.\nAnd if I switch around the true and false logic for IS_WIN variable, it works fine.\n\nIs this by design?? Your clarification would be much helpful!! :)","from":"developer"}],"created":"2011-12-13T16:31:59.000+0000","description":"HADOOP-7899 & HDFS-2511 introduced compilation of proto files as part of the build.\n\nSuch compilation is failing in windows/cygwin","issue_id":"12534976","key":"HADOOP-7917","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2011-12-15T14:47:50.000+0000","role":"fixed_distractor","summary":"compilation of protobuf files fails in windows/cygwin"} {"case_id":"12536346","cluster":"DISTRACTOR-HADOOP-7940","comments":[{"body":"Hey Aaron,\n\nWould you be interested in contributing a fix and a test case for this as well?\n\nThanks,\nHarsh","created":"2012-01-15T06:56:43.755+0000"},{"body":"Index: hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/io/Text.java\n===================================================================\n--- hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/io/Text.java\t(revision 1294407)\n+++ hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/io/Text.java\t(working copy)\n@@ -239,6 +239,7 @@\n */\n public void clear() {\n length = 0;\n+ bytes = EMPTY_BYTES;\n }\n \n /*\nIndex: hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/io/TestText.java\n===================================================================\n--- hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/io/TestText.java\t(revision 1294407)\n+++ hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/io/TestText.java\t(working copy)\n@@ -192,6 +192,16 @@\n assertTrue(text.find(\"\\u20ac\", 5)==11);\n }\n \n+ public void testClear() {\n+\tText text = new Text();\n+\tassertEquals(\"\", text.toString());\n+\tassertEquals(0, text.getBytes().length);\n+\ttext = new Text(\"abcd\\u20acbdcd\\u20ac\");\n+\ttext.clear();\n+\tassertEquals(\"\", text.toString());\n+\tassertEquals(0, text.getBytes().length);\n+ }\n+\n public void testFindAfterUpdatingContents() throws Exception {\n Text text = new Text(\"abcd\");\n text.set(\"a\".getBytes());\n","created":"2012-02-28T00:12:24.948+0000"},{"body":"Hi Csaba,\n\nThanks for contributing! Mind attaching a patch against trunk? (attach files and be sure to check \"Grant license to ASF for inclusion in ASF works\"). This is necessary to incorporate your change into the code. More info here: http://wiki.apache.org/hadoop/HowToContribute\n\nThanks,\nEli","created":"2012-02-28T16:34:25.006+0000"},{"body":"Hi Eli!\n\nI've read HowToContribute page. This patch for the trunk. I've also attached the patch to the issue.\n\nBest regards,\nCsaba","created":"2012-02-28T17:29:41.603+0000"},{"body":"+1, TestViewFsTrash failure is unrelated to this patch (There's already a separate bug report available for its failing though). Test fails without fix.\n\nCommitting to trunk and branch-0.23 shortly.","created":"2012-02-29T06:39:20.861+0000"},{"body":"Committed to branch-0.23 and trunk. Thanks a lot for your contribution Csaba! We're hoping to see many more :)","created":"2012-02-29T11:08:27.551+0000"},{"body":"toString, getLength and other Text methods behave correctly after clear(). This will hurt the performance of the class when calling append often and then clear because you wont be able to use the byte array size that you had previously scaled up to.\n\nI dont think there was any expectation that getBytes would return a 0 length array, it says right in the javadoc: \"Returns the raw bytes; however, only data up to {@link #getLength()} is valid.\" which in this case is valid, 0 of the bytes are valid","created":"2012-04-25T03:37:45.215+0000"},{"body":"This will likely hurt the performance problem described in HADOOP-6109.\n\nI was looking at the javadoc in Hadoop 1.0.2. The more recent javadoc adds: \"Please use {@link #copyBytes()} if you need the returned array to be precisely the length of the data.\" which makes it even clearer than you must pay attention to the length if you use getBytes()","created":"2012-04-25T04:02:38.948+0000"},{"body":"[~jdonofrio] - You're right there. I've filed HADOOP-8323 to revert this.","created":"2012-04-27T07:11:00.251+0000"},{"body":"Jim - I'm actually not convinced this hurts at all. I've decided not to have this reverted in second thought, and I've explained the reasoning at HADOOP-8323.\n\nClear call ought to clear memory and thats what this change actually intends to do (though the test case may lead you astray - filed HADOOP-8324).\n\nThe current javadocs indeed cover the getBytes usage behavior as you've pointed out. So if you'd like to keep the size in its increased state, why clear() it?","created":"2012-04-27T09:20:29.538+0000"},{"body":"*Note:* This was reverted by HADOOP-8323. Please see that ticket for more information.","created":"2012-05-06T11:37:07.086+0000"}],"conversations":[{"body":"LineReader reader = new LineReader(in, 4096);\n...\n\nText text = new Text();\nwhile((reader.readLine(text)) > 0) {\n ...\n text.clear();\n}\n}\n\nEven the clear() method is called each time, some bytes are still not filled as zero.\nSo, when reader.readLine(text) is called in a loop, some bytes are dirty which was from last call.","from":"reporter","subject":"method clear() in org.apache.hadoop.io.Text does not work"},{"body":"Hey Aaron,\n\nWould you be interested in contributing a fix and a test case for this as well?\n\nThanks,\nHarsh","from":"developer"},{"body":"Index: hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/io/Text.java\n===================================================================\n--- hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/io/Text.java\t(revision 1294407)\n+++ hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/io/Text.java\t(working copy)\n@@ -239,6 +239,7 @@\n */\n public void clear() {\n length = 0;\n+ bytes = EMPTY_BYTES;\n }\n \n /*\nIndex: hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/io/TestText.java\n===================================================================\n--- hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/io/TestText.java\t(revision 1294407)\n+++ hadoop-common-project/hadoop-common/src/test/java/org/apache/hadoop/io/TestText.java\t(working copy)\n@@ -192,6 +192,16 @@\n assertTrue(text.find(\"\\u20ac\", 5)==11);\n }\n \n+ public void testClear() {\n+\tText text = new Text();\n+\tassertEquals(\"\", text.toString());\n+\tassertEquals(0, text.getBytes().length);\n+\ttext = new Text(\"abcd\\u20acbdcd\\u20ac\");\n+\ttext.clear();\n+\tassertEquals(\"\", text.toString());\n+\tassertEquals(0, text.getBytes().length);\n+ }\n+\n public void testFindAfterUpdatingContents() throws Exception {\n Text text = new Text(\"abcd\");\n text.set(\"a\".getBytes());\n","from":"developer"},{"body":"Hi Csaba,\n\nThanks for contributing! Mind attaching a patch against trunk? (attach files and be sure to check \"Grant license to ASF for inclusion in ASF works\"). This is necessary to incorporate your change into the code. More info here: http://wiki.apache.org/hadoop/HowToContribute\n\nThanks,\nEli","from":"developer"},{"body":"Hi Eli!\n\nI've read HowToContribute page. This patch for the trunk. I've also attached the patch to the issue.\n\nBest regards,\nCsaba","from":"developer"},{"body":"+1, TestViewFsTrash failure is unrelated to this patch (There's already a separate bug report available for its failing though). Test fails without fix.\n\nCommitting to trunk and branch-0.23 shortly.","from":"developer"},{"body":"Committed to branch-0.23 and trunk. Thanks a lot for your contribution Csaba! We're hoping to see many more :)","from":"developer"},{"body":"toString, getLength and other Text methods behave correctly after clear(). This will hurt the performance of the class when calling append often and then clear because you wont be able to use the byte array size that you had previously scaled up to.\n\nI dont think there was any expectation that getBytes would return a 0 length array, it says right in the javadoc: \"Returns the raw bytes; however, only data up to {@link #getLength()} is valid.\" which in this case is valid, 0 of the bytes are valid","from":"developer"},{"body":"This will likely hurt the performance problem described in HADOOP-6109.\n\nI was looking at the javadoc in Hadoop 1.0.2. The more recent javadoc adds: \"Please use {@link #copyBytes()} if you need the returned array to be precisely the length of the data.\" which makes it even clearer than you must pay attention to the length if you use getBytes()","from":"developer"},{"body":"[~jdonofrio] - You're right there. I've filed HADOOP-8323 to revert this.","from":"developer"},{"body":"Jim - I'm actually not convinced this hurts at all. I've decided not to have this reverted in second thought, and I've explained the reasoning at HADOOP-8323.\n\nClear call ought to clear memory and thats what this change actually intends to do (though the test case may lead you astray - filed HADOOP-8324).\n\nThe current javadocs indeed cover the getBytes usage behavior as you've pointed out. So if you'd like to keep the size in its increased state, why clear() it?","from":"developer"},{"body":"*Note:* This was reverted by HADOOP-8323. Please see that ticket for more information.","from":"developer"}],"created":"2011-12-24T14:50:11.000+0000","description":"LineReader reader = new LineReader(in, 4096);\n...\n\nText text = new Text();\nwhile((reader.readLine(text)) > 0) {\n ...\n text.clear();\n}\n}\n\nEven the clear() method is called each time, some bytes are still not filled as zero.\nSo, when reader.readLine(text) is called in a loop, some bytes are dirty which was from last call.","issue_id":"12536346","key":"HADOOP-7940","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2012-02-29T11:08:27.000+0000","role":"fixed_distractor","summary":"method clear() in org.apache.hadoop.io.Text does not work"} {"case_id":"12538377","cluster":"DISTRACTOR-HADOOP-7974","comments":[{"body":"Patch that uses a get-parent call instead of hacking with strings.","created":"2012-01-15T05:06:02.616+0000"},{"body":"(Also fixed some diff-surrounding whitesp. and indentation issues)","created":"2012-01-15T05:07:37.479+0000"},{"body":"+1 javadoc failures are unrelated","created":"2012-01-15T08:51:12.857+0000"},{"body":"I've committed this and merged to 23. Thanks Harsh!","created":"2012-01-15T08:54:40.610+0000"},{"body":"It's been fixed for a while.","created":"2012-02-08T02:09:28.874+0000"},{"body":"TestViewFsTrash occasionally fails. Filed HADOOP-8110. I am not sure if this issue causes it. It would be great if you can take a look.","created":"2012-02-24T18:53:52.281+0000"}],"conversations":[{"body":"HADOOP-7284 added a test called TestViewFsTrash which contains the following code to determine the user's home directory. It only works if the user's directory is one level deep, and breaks if the home directory is more than one level deep (eg user hudson, who's home dir might be /usr/lib/hudson instead of /home/hudson).\n\n{code}\n // create a link for home directory so that trash path works\n // set up viewfs's home dir root to point to home dir root on target\n // But home dir is different on linux, mac etc.\n // Figure it out by calling home dir on target\n \n String homeDir = fsTarget.getHomeDirectory().toUri().getPath();\n int indexOf2ndSlash = homeDir.indexOf('/', 1);\n String homeDirRoot = homeDir.substring(0, indexOf2ndSlash);\n ConfigUtil.addLink(conf, homeDirRoot,\n fsTarget.makeQualified(new Path(homeDirRoot)).toUri()); \n ConfigUtil.setHomeDirConf(conf, homeDirRoot);\n Log.info(\"Home dir base \" + homeDirRoot);\n{code}\n\nSeems like we should instead search from the end of the path for the last slash and use that as the base, ie ask the home directory for its parent.\n\n\n\n\n\n\n","from":"reporter","subject":"TestViewFsTrash incorrectly determines the user's home directory"},{"body":"Patch that uses a get-parent call instead of hacking with strings.","from":"developer"},{"body":"(Also fixed some diff-surrounding whitesp. and indentation issues)","from":"developer"},{"body":"+1 javadoc failures are unrelated","from":"developer"},{"body":"I've committed this and merged to 23. Thanks Harsh!","from":"developer"},{"body":"It's been fixed for a while.","from":"developer"},{"body":"TestViewFsTrash occasionally fails. Filed HADOOP-8110. I am not sure if this issue causes it. It would be great if you can take a look.","from":"developer"}],"created":"2012-01-13T23:50:17.000+0000","description":"HADOOP-7284 added a test called TestViewFsTrash which contains the following code to determine the user's home directory. It only works if the user's directory is one level deep, and breaks if the home directory is more than one level deep (eg user hudson, who's home dir might be /usr/lib/hudson instead of /home/hudson).\n\n{code}\n // create a link for home directory so that trash path works\n // set up viewfs's home dir root to point to home dir root on target\n // But home dir is different on linux, mac etc.\n // Figure it out by calling home dir on target\n \n String homeDir = fsTarget.getHomeDirectory().toUri().getPath();\n int indexOf2ndSlash = homeDir.indexOf('/', 1);\n String homeDirRoot = homeDir.substring(0, indexOf2ndSlash);\n ConfigUtil.addLink(conf, homeDirRoot,\n fsTarget.makeQualified(new Path(homeDirRoot)).toUri()); \n ConfigUtil.setHomeDirConf(conf, homeDirRoot);\n Log.info(\"Home dir base \" + homeDirRoot);\n{code}\n\nSeems like we should instead search from the end of the path for the last slash and use that as the base, ie ask the home directory for its parent.\n\n\n\n\n\n\n","issue_id":"12538377","key":"HADOOP-7974","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2012-02-08T02:09:28.000+0000","role":"fixed_distractor","summary":"TestViewFsTrash incorrectly determines the user's home directory"} {"case_id":"12541687","cluster":"DISTRACTOR-HADOOP-8031","comments":[{"body":"Simple patch","created":"2012-03-06T21:35:37.311+0000"},{"body":"Thanks for contributing Elias. Can you update TestConfiguration with a case that will fail w/o your patch?\n\nI've rebased your patch on trunk.","created":"2012-05-26T01:52:58.891+0000"},{"body":"Patch attached.","created":"2012-05-26T01:53:33.925+0000"},{"body":"Eli,\n\nThanks for your response. I would like to reproduce the problem but I'd have to somehow embed the .xml file inside a .jar and adjust the test classpath to match. I'd likely have to isolate the test from the rest of the existing classpath as well. Maybe you could guide me through this?\n\nThanks.","created":"2012-05-26T03:39:41.873+0000"},{"body":"Thanks Elias, \n\nI have noticed that the patch no longer applies due to the recent changes to Configuration.java. I have made the required changes so it can apply cleanly on trunk. \n\nAlso changed:\n\n{code}\n+ doc = builder.parse((InputStream)name);\n{code}\n\nto:\n\n{code}\n+ doc = parse(builder, (InputStream) resource);\n{code}\n\nto insure the InputStream will be closed.\n\nRegarding testing, I am not sure if there is an easy way to reproduce this issue in a unit test.","created":"2012-08-23T08:19:17.857+0000"},{"body":"+1","created":"2012-08-23T23:22:01.442+0000"},{"body":"Thanks Elias. Committed to trunk and branch-2.","created":"2012-08-23T23:30:10.814+0000"},{"body":"Hey Tucu. I think this commit broke the way in which relative xincludes are handled in Configuration. I have some development confs which use xinclude with non-absolute paths, and it used to successfully pick up the included files from my conf directory. Now, it seems to be looking in the current working directory instead.\n\nIs it possible to fix the code so that the relative paths are resolved the same as before? I think xinclude is relatively common for deployments.","created":"2012-08-30T01:00:00.858+0000"},{"body":"I confirmed that reverting this patch locally restored the old behavior.\n\nIf we can't maintain the old behavior, we should at least mark this as an incompatible change. But I bet it's doable to both fix it and have relative xincludes.","created":"2012-08-30T01:13:20.255+0000"},{"body":"Hi Todd, It is weird that this patch caused this behavior change. The patch didn't modify the builder or the docBuilderFactory, and it is still docBuilderFactory.setXIncludeAware(true). In essence, the patch is basically using DocumentBuilder#parse(InputStream uri.openStream()) instead of DocumentBuilder#parse(String uri.toString()). Seems there is a change in implementation of both parse methods, which seems weird.","created":"2012-08-30T01:37:01.362+0000"},{"body":"Yea, I don't know much about the underlying API, but it definitely changed the behavior. It's still trying to do the xinclude, it's just looking in cwd instead of my conf dir.","created":"2012-08-30T01:43:27.626+0000"},{"body":"I looked into the implementation of javax.xml.parsers.DocumentBuilder and org.xml.sax.InputSource and there is a difference when the DocumentBuilder parse(String) method is used versus parse(InputStream). Basically we need to use parse(InputStream is, String systemId) which provides a base for resolving relative URIs. Here is a new patch that fixes this issue. It needs to be applied on top of the previously committed patch. I am not sure if we need to create a new ticket since this one is already committed.","created":"2012-08-30T04:08:56.879+0000"},{"body":"I'd either file a new jira explaining the bug with a new patch with fix/test, or revert this patch and post a new one that includes the fix.","created":"2012-08-30T18:35:51.845+0000"},{"body":"Thanks Eli! I have created HADOOP-8749.","created":"2012-08-30T18:59:37.616+0000"},{"body":"Re-resolving this since Ahmed is addressing my issue in HADOOP-8749","created":"2012-08-30T20:16:47.375+0000"}],"conversations":[{"body":"While running a hadoop client within RHQ (monitoring software) using its classloader, I see this:\n\n2012-02-07 09:15:25,313 INFO [ResourceContainer.invoker.daemon-2] (org.apache.hadoop.conf.Configuration)- parsing jar:file:/usr/local/rhq-agent/data/tmp/rhq-hadoop-plugin-4.3.0-SNAPSHOT.jar6856622641102893436.classloader/hadoop-core-0.20.2+737+1.jar7204287718482036191.tmp!/core-default.xml\n2012-02-07 09:15:25,318 ERROR [InventoryManager.discovery-1] (rhq.core.pc.inventory.InventoryManager)- Failed to start component for Resource[id=16290, type=NameNode, key=NameNode:/usr/lib/hadoop-0.20, name=NameNode, parent=vg61l01ad-hadoop002.apple.com] from synchronized merge.\norg.rhq.core.clientapi.agent.PluginContainerException: Failed to start component for resource Resource[id=16290, type=NameNode, key=NameNode:/usr/lib/hadoop-0.20, name=NameNode, parent=vg61l01ad-hadoop002.apple.com].\nCaused by: java.lang.RuntimeException: core-site.xml not found\n\tat org.apache.hadoop.conf.Configuration.loadResource(Configuration.java:1308)\n\tat org.apache.hadoop.conf.Configuration.loadResources(Configuration.java:1228)\n\tat org.apache.hadoop.conf.Configuration.getProps(Configuration.java:1169)\n\tat org.apache.hadoop.conf.Configuration.set(Configuration.java:438)\n\nThis is because the URL\n\njar:file:/usr/local/rhq-agent/data/tmp/rhq-hadoop-plugin-4.3.0-SNAPSHOT.jar6856622641102893436.classloader/hadoop-core-0.20.2+737+1.jar7204287718482036191.tmp!/core-default.xml\n\ncannot be found by DocumentBuilder (doesn't understand it). (Note: the logs are for an old version of Configuration class, but the new version has the same code.)\n\nThe solution is to obtain the resource stream directly from the URL object itself.\n\nThat is to say:\n\n{code}\n URL url = getResource((String)name);\n- if (url != null) {\n- if (!quiet) {\n- LOG.info(\"parsing \" + url);\n- }\n- doc = builder.parse(url.toString());\n- }\n+ doc = builder.parse(url.openStream());\n{code}\n\nNote: I have a full patch pending approval at Apple for this change, including some cleanup.","from":"reporter","subject":"Configuration class fails to find embedded .jar resources; should use URL.openStream()"},{"body":"Simple patch","from":"developer"},{"body":"Thanks for contributing Elias. Can you update TestConfiguration with a case that will fail w/o your patch?\n\nI've rebased your patch on trunk.","from":"developer"},{"body":"Patch attached.","from":"developer"},{"body":"Eli,\n\nThanks for your response. I would like to reproduce the problem but I'd have to somehow embed the .xml file inside a .jar and adjust the test classpath to match. I'd likely have to isolate the test from the rest of the existing classpath as well. Maybe you could guide me through this?\n\nThanks.","from":"developer"},{"body":"Thanks Elias, \n\nI have noticed that the patch no longer applies due to the recent changes to Configuration.java. I have made the required changes so it can apply cleanly on trunk. \n\nAlso changed:\n\n{code}\n+ doc = builder.parse((InputStream)name);\n{code}\n\nto:\n\n{code}\n+ doc = parse(builder, (InputStream) resource);\n{code}\n\nto insure the InputStream will be closed.\n\nRegarding testing, I am not sure if there is an easy way to reproduce this issue in a unit test.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Thanks Elias. Committed to trunk and branch-2.","from":"developer"},{"body":"Hey Tucu. I think this commit broke the way in which relative xincludes are handled in Configuration. I have some development confs which use xinclude with non-absolute paths, and it used to successfully pick up the included files from my conf directory. Now, it seems to be looking in the current working directory instead.\n\nIs it possible to fix the code so that the relative paths are resolved the same as before? I think xinclude is relatively common for deployments.","from":"developer"},{"body":"I confirmed that reverting this patch locally restored the old behavior.\n\nIf we can't maintain the old behavior, we should at least mark this as an incompatible change. But I bet it's doable to both fix it and have relative xincludes.","from":"developer"},{"body":"Hi Todd, It is weird that this patch caused this behavior change. The patch didn't modify the builder or the docBuilderFactory, and it is still docBuilderFactory.setXIncludeAware(true). In essence, the patch is basically using DocumentBuilder#parse(InputStream uri.openStream()) instead of DocumentBuilder#parse(String uri.toString()). Seems there is a change in implementation of both parse methods, which seems weird.","from":"developer"},{"body":"Yea, I don't know much about the underlying API, but it definitely changed the behavior. It's still trying to do the xinclude, it's just looking in cwd instead of my conf dir.","from":"developer"},{"body":"I looked into the implementation of javax.xml.parsers.DocumentBuilder and org.xml.sax.InputSource and there is a difference when the DocumentBuilder parse(String) method is used versus parse(InputStream). Basically we need to use parse(InputStream is, String systemId) which provides a base for resolving relative URIs. Here is a new patch that fixes this issue. It needs to be applied on top of the previously committed patch. I am not sure if we need to create a new ticket since this one is already committed.","from":"developer"},{"body":"I'd either file a new jira explaining the bug with a new patch with fix/test, or revert this patch and post a new one that includes the fix.","from":"developer"},{"body":"Thanks Eli! I have created HADOOP-8749.","from":"developer"},{"body":"Re-resolving this since Ahmed is addressing my issue in HADOOP-8749","from":"developer"}],"created":"2012-02-07T20:22:24.000+0000","description":"While running a hadoop client within RHQ (monitoring software) using its classloader, I see this:\n\n2012-02-07 09:15:25,313 INFO [ResourceContainer.invoker.daemon-2] (org.apache.hadoop.conf.Configuration)- parsing jar:file:/usr/local/rhq-agent/data/tmp/rhq-hadoop-plugin-4.3.0-SNAPSHOT.jar6856622641102893436.classloader/hadoop-core-0.20.2+737+1.jar7204287718482036191.tmp!/core-default.xml\n2012-02-07 09:15:25,318 ERROR [InventoryManager.discovery-1] (rhq.core.pc.inventory.InventoryManager)- Failed to start component for Resource[id=16290, type=NameNode, key=NameNode:/usr/lib/hadoop-0.20, name=NameNode, parent=vg61l01ad-hadoop002.apple.com] from synchronized merge.\norg.rhq.core.clientapi.agent.PluginContainerException: Failed to start component for resource Resource[id=16290, type=NameNode, key=NameNode:/usr/lib/hadoop-0.20, name=NameNode, parent=vg61l01ad-hadoop002.apple.com].\nCaused by: java.lang.RuntimeException: core-site.xml not found\n\tat org.apache.hadoop.conf.Configuration.loadResource(Configuration.java:1308)\n\tat org.apache.hadoop.conf.Configuration.loadResources(Configuration.java:1228)\n\tat org.apache.hadoop.conf.Configuration.getProps(Configuration.java:1169)\n\tat org.apache.hadoop.conf.Configuration.set(Configuration.java:438)\n\nThis is because the URL\n\njar:file:/usr/local/rhq-agent/data/tmp/rhq-hadoop-plugin-4.3.0-SNAPSHOT.jar6856622641102893436.classloader/hadoop-core-0.20.2+737+1.jar7204287718482036191.tmp!/core-default.xml\n\ncannot be found by DocumentBuilder (doesn't understand it). (Note: the logs are for an old version of Configuration class, but the new version has the same code.)\n\nThe solution is to obtain the resource stream directly from the URL object itself.\n\nThat is to say:\n\n{code}\n URL url = getResource((String)name);\n- if (url != null) {\n- if (!quiet) {\n- LOG.info(\"parsing \" + url);\n- }\n- doc = builder.parse(url.toString());\n- }\n+ doc = builder.parse(url.openStream());\n{code}\n\nNote: I have a full patch pending approval at Apple for this change, including some cleanup.","issue_id":"12541687","key":"HADOOP-8031","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2012-08-30T20:16:47.000+0000","role":"fixed_distractor","summary":"Configuration class fails to find embedded .jar resources; should use URL.openStream()"} {"case_id":"12547550","cluster":"DISTRACTOR-HADOOP-8197","comments":[{"body":"verified visually that log WARNs are written once","created":"2012-03-22T15:14:18.303+0000"},{"body":"+1, patch looks good.\n\nOne optional nitpick though: Would {{warnOnceIfDeprecated}} be a better method name?","created":"2012-03-22T16:23:22.870+0000"},{"body":"same patch renaming method per Harsh's suggestion","created":"2012-03-22T16:46:09.757+0000"},{"body":"Committed to trunk and branch-0.23","created":"2012-03-22T16:49:51.332+0000"},{"body":"Pulled this into branch-0.23","created":"2012-07-30T14:22:23.181+0000"}],"conversations":[{"body":"The logic to do print a warning only once per deprecated key does not work:\n\n{code}\n2012-03-21 22:32:58,121 WARN Configuration:661 - user.name is deprecated. Instead, use mapreduce.job.user.name\n....\n2012-03-21 22:32:58,123 WARN Configuration:661 - fs.default.name is deprecated. Instead, use fs.defaultFS\n...\n2012-03-21 22:32:58,130 WARN Configuration:661 - mapred.job.tracker is deprecated. Instead, use mapreduce.jobtracker.address\n2012-03-21 22:32:58,351 WARN Configuration:345 - fs.default.name is deprecated. Instead, use fs.defaultFS\n...\n2012-03-21 22:32:58,843 WARN Configuration:661 - user.name is deprecated. Instead, use mapreduce.job.user.name\n2012-03-21 22:32:58,844 WARN Configuration:661 - mapred.job.tracker is deprecated. Instead, use mapreduce.jobtracker.address\n2012-03-21 22:32:58,844 WARN Configuration:661 - fs.default.name is deprecated. Instead, use fs.defaultFS\n{code}","from":"reporter","subject":"Configuration logs WARNs on every use of a deprecated key"},{"body":"verified visually that log WARNs are written once","from":"developer"},{"body":"+1, patch looks good.\n\nOne optional nitpick though: Would {{warnOnceIfDeprecated}} be a better method name?","from":"developer"},{"body":"same patch renaming method per Harsh's suggestion","from":"developer"},{"body":"Committed to trunk and branch-0.23","from":"developer"},{"body":"Pulled this into branch-0.23","from":"developer"}],"created":"2012-03-22T05:48:44.000+0000","description":"The logic to do print a warning only once per deprecated key does not work:\n\n{code}\n2012-03-21 22:32:58,121 WARN Configuration:661 - user.name is deprecated. Instead, use mapreduce.job.user.name\n....\n2012-03-21 22:32:58,123 WARN Configuration:661 - fs.default.name is deprecated. Instead, use fs.defaultFS\n...\n2012-03-21 22:32:58,130 WARN Configuration:661 - mapred.job.tracker is deprecated. Instead, use mapreduce.jobtracker.address\n2012-03-21 22:32:58,351 WARN Configuration:345 - fs.default.name is deprecated. Instead, use fs.defaultFS\n...\n2012-03-21 22:32:58,843 WARN Configuration:661 - user.name is deprecated. Instead, use mapreduce.job.user.name\n2012-03-21 22:32:58,844 WARN Configuration:661 - mapred.job.tracker is deprecated. Instead, use mapreduce.jobtracker.address\n2012-03-21 22:32:58,844 WARN Configuration:661 - fs.default.name is deprecated. Instead, use fs.defaultFS\n{code}","issue_id":"12547550","key":"HADOOP-8197","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2012-03-22T16:49:51.000+0000","role":"fixed_distractor","summary":"Configuration logs WARNs on every use of a deprecated key"} {"case_id":"12550497","cluster":"DISTRACTOR-HADOOP-8268","comments":[{"body":"I didnt get why this patch didnt apply. it works fine here against trunk. It might be CR/LF issue?","created":"2012-04-11T13:52:11.860+0000"},{"body":"It doesn't apply because test-patch doesn't work on files that are under hadoop-project-dist.\n\nYou should probably just say what testing you did in a comment here. A reviewer should probably apply it themself before committing it.","created":"2012-04-12T17:58:38.224+0000"},{"body":"I did testing with:\n\nmvn package -Pdist -Dtar -DskipTests","created":"2012-04-12T19:55:55.446+0000"},{"body":"I really need to get this committed. This change is trivial (just XML markup) and it prevents me from using company build infrastructure for my contributions to hadoop.","created":"2012-04-27T11:24:58.838+0000"},{"body":"re-uploaded patch to see if new hadoop builds checking system will kick in","created":"2012-05-03T06:55:56.268+0000"},{"body":"The change looks fine. Before we commit it in though, did you run a build on Windows+Cygwin as well, since those are what you're changing the commands for?\n\nAlso, why do you see a failure only on hadoop-project-dist, when the same issue exists in so many other pom.xml files?\n\n{code}\n./hadoop-common-project/hadoop-common/pom.xml: which cygpath 2> /dev/null\n./hadoop-common-project/hadoop-common/pom.xml: mkdir -p $JAVA_DIR 2> /dev/null\n./hadoop-common-project/hadoop-common/pom.xml: for PROTO_FILE in `ls $PROTO_DIR/*.proto 2> /dev/null`\n./hadoop-common-project/hadoop-common/pom.xml: which cygpath 2> /dev/null\n./hadoop-common-project/hadoop-common/pom.xml: mkdir -p $JAVA_DIR 2> /dev/null\n./hadoop-common-project/hadoop-common/pom.xml: for PROTO_FILE in `ls $PROTO_DIR/*.proto 2> /dev/null`\n./hadoop-hdfs-project/hadoop-hdfs/pom.xml: which cygpath 2> /dev/null\n./hadoop-hdfs-project/hadoop-hdfs/pom.xml: mkdir -p $JAVA_DIR 2> /dev/null\n./hadoop-hdfs-project/hadoop-hdfs/pom.xml: for PROTO_FILE in `ls $PROTO_DIR/*.proto 2> /dev/null`\n./hadoop-hdfs-project/hadoop-hdfs-httpfs/pom.xml: which cygpath 2> /dev/null\n./hadoop-hdfs-project/hadoop-hdfs-httpfs/pom.xml: which cygpath 2> /dev/null\n./hadoop-mapreduce-project/pom.xml: which cygpath 2> /dev/null\n{code}\n\nWouldn't you like them all fixed?\n\nbq. -1 javadoc. The javadoc tool appears to have generated 2 warning messages.\n\nThis is cause of HADOOP-8359.\n\n","created":"2012-05-05T17:48:43.377+0000"},{"body":"(There may be other > usage as well, I just searched for \"2>\" there above.)","created":"2012-05-05T17:49:57.314+0000"},{"body":"bq. The change looks fine. Before we commit it in though, did you run a build on Windows+Cygwin as well, since those are what you're changing the commands for.\n\nNevermind this comment. I manually verified we don't break anything:\n\n{code}\n➜ hadoop-project-dist git:(trunk) ✗ cat ./target/dist-copynativelibs.sh\nwhich cygpath 2> /dev/null\n{code}","created":"2012-05-05T17:56:09.224+0000"},{"body":"I have not tested this on cygwin because i do not have it installed.\n\nIf some thing is in other POM files then it needs to be fixed as well, i just submitted patch for first POM on which our project validation system aborted build.","created":"2012-05-06T05:48:31.159+0000"},{"body":"Commit all pom files with \"2>\" you have discovered. If there will be more problems then i will open new ticket.","created":"2012-05-07T12:59:13.807+0000"},{"body":"I believe this patch fixes all pom.xml occurrences of \"2>\" to {{2>}}.\n\nRunning a grep against all pom.xml files for \" >\" to pick up stdout redirector issues turned up no results. Hence I'm thinking that this patch suffices to fix the redirectors-caused validation issues properly at least.\n\nPatch was tested by running Radim's command and inspecting the generated files, to ensure they carried \">\" and not {{>}} directly.\n\nPlease review.\n\n@Radim - Would still be good if you can post us a list of all bad xml instances and we fix all that in this JIRA. I have none setup with me right now, but lemme know if I need to do that by self.","created":"2012-05-07T13:15:36.430+0000"},{"body":"Lets take another approach. Standard unix command for validating POMs is part of libxml2 package:\n\nfind . -name \\*.xml -exec xmllint --valid --noout {} \\;\n\nto make xmllint validation work, you need to add proper schema to every pom which is good practice anyway.\n\n\n\nhadoop has 43 pom.xml files.","created":"2012-05-14T16:12:58.268+0000"},{"body":"command was wrong. libxml2 does not work well with XSD schemas. C version of xerces 3 parser works fine:\n\nfind . -name *.xml -exec PParse -n -s -f -v=always {} \\;","created":"2012-05-15T15:44:28.207+0000"},{"body":"added xml schema to maven POMs","created":"2012-05-15T15:49:34.930+0000"},{"body":"guys, seriously commit this already. It takes you months to commit such trivial change like this.","created":"2012-05-22T13:08:25.604+0000"},{"body":"+1 pending jenkins. Sorry for the delay Radim.","created":"2012-05-22T14:08:25.455+0000"},{"body":"Re-upping on Radim's behalf to trigger QA Bot.","created":"2012-05-22T14:09:16.764+0000"},{"body":"merged trunk","created":"2012-05-24T14:24:42.372+0000"},{"body":"Reupload patch. JIRA was down, QA bot was unable to post results.","created":"2012-05-25T04:34:33.067+0000"},{"body":"Patch had a few issues: New lines were Windows-style (trailing CRs in it), and wasn't -p0 applicable (although not an issue these days with test-patch, as it tries upto three levels).\n\nHere's a more compatible svn patch. Lets hit QA Bot with this.\n\nLocally mvn 3.x passes an install run (without tests) for me. So am +1. Will commit this in by monday regardless of QA (would be good to have it run though) unless someone else objects. The same kinda lines is already used in HBase, and thats another basis for my +1 here.\n\nThanks Radim.\n\n(Here it goes…)","created":"2012-05-25T13:27:19.737+0000"},{"body":"Fixed filename of patch.","created":"2012-05-25T13:28:20.988+0000"},{"body":"reupload for bot","created":"2012-05-26T07:17:07.964+0000"},{"body":"Here is bot output. It will most likely fail to post results to JIRA.\n\nhttps://builds.apache.org/job/PreCommit-HADOOP-Build/1040/","created":"2012-05-26T07:24:01.706+0000"},{"body":"Thanks Radim!\n\n+1 based on a local {{mvn clean install -Dtest=FooBar}} as well. Committing to branch-2 and trunk shortly.","created":"2012-05-28T14:46:18.704+0000"},{"body":"- Committed revision 1343272 to trunk.\n- svn merge -c 1343272 to branch-2 committed as revision 1343275.\n\nThanks for this and also your continuing contributions Radim! :)","created":"2012-05-28T14:59:12.298+0000"}],"conversations":[{"body":"In a few pom files there are embedded ant commands which contains '>' - redirection. This makes XML file invalid and this POM file can not be deployed into validating Maven repository managers such as Artifactory.","from":"reporter","subject":"A few pom.xml across Hadoop project may fail XML validation"},{"body":"I didnt get why this patch didnt apply. it works fine here against trunk. It might be CR/LF issue?","from":"developer"},{"body":"It doesn't apply because test-patch doesn't work on files that are under hadoop-project-dist.\n\nYou should probably just say what testing you did in a comment here. A reviewer should probably apply it themself before committing it.","from":"developer"},{"body":"I did testing with:\n\nmvn package -Pdist -Dtar -DskipTests","from":"developer"},{"body":"I really need to get this committed. This change is trivial (just XML markup) and it prevents me from using company build infrastructure for my contributions to hadoop.","from":"developer"},{"body":"re-uploaded patch to see if new hadoop builds checking system will kick in","from":"developer"},{"body":"The change looks fine. Before we commit it in though, did you run a build on Windows+Cygwin as well, since those are what you're changing the commands for?\n\nAlso, why do you see a failure only on hadoop-project-dist, when the same issue exists in so many other pom.xml files?\n\n{code}\n./hadoop-common-project/hadoop-common/pom.xml: which cygpath 2> /dev/null\n./hadoop-common-project/hadoop-common/pom.xml: mkdir -p $JAVA_DIR 2> /dev/null\n./hadoop-common-project/hadoop-common/pom.xml: for PROTO_FILE in `ls $PROTO_DIR/*.proto 2> /dev/null`\n./hadoop-common-project/hadoop-common/pom.xml: which cygpath 2> /dev/null\n./hadoop-common-project/hadoop-common/pom.xml: mkdir -p $JAVA_DIR 2> /dev/null\n./hadoop-common-project/hadoop-common/pom.xml: for PROTO_FILE in `ls $PROTO_DIR/*.proto 2> /dev/null`\n./hadoop-hdfs-project/hadoop-hdfs/pom.xml: which cygpath 2> /dev/null\n./hadoop-hdfs-project/hadoop-hdfs/pom.xml: mkdir -p $JAVA_DIR 2> /dev/null\n./hadoop-hdfs-project/hadoop-hdfs/pom.xml: for PROTO_FILE in `ls $PROTO_DIR/*.proto 2> /dev/null`\n./hadoop-hdfs-project/hadoop-hdfs-httpfs/pom.xml: which cygpath 2> /dev/null\n./hadoop-hdfs-project/hadoop-hdfs-httpfs/pom.xml: which cygpath 2> /dev/null\n./hadoop-mapreduce-project/pom.xml: which cygpath 2> /dev/null\n{code}\n\nWouldn't you like them all fixed?\n\nbq. -1 javadoc. The javadoc tool appears to have generated 2 warning messages.\n\nThis is cause of HADOOP-8359.\n\n","from":"developer"},{"body":"(There may be other > usage as well, I just searched for \"2>\" there above.)","from":"developer"},{"body":"bq. The change looks fine. Before we commit it in though, did you run a build on Windows+Cygwin as well, since those are what you're changing the commands for.\n\nNevermind this comment. I manually verified we don't break anything:\n\n{code}\n➜ hadoop-project-dist git:(trunk) ✗ cat ./target/dist-copynativelibs.sh\nwhich cygpath 2> /dev/null\n{code}","from":"developer"},{"body":"I have not tested this on cygwin because i do not have it installed.\n\nIf some thing is in other POM files then it needs to be fixed as well, i just submitted patch for first POM on which our project validation system aborted build.","from":"developer"},{"body":"Commit all pom files with \"2>\" you have discovered. If there will be more problems then i will open new ticket.","from":"developer"},{"body":"I believe this patch fixes all pom.xml occurrences of \"2>\" to {{2>}}.\n\nRunning a grep against all pom.xml files for \" >\" to pick up stdout redirector issues turned up no results. Hence I'm thinking that this patch suffices to fix the redirectors-caused validation issues properly at least.\n\nPatch was tested by running Radim's command and inspecting the generated files, to ensure they carried \">\" and not {{>}} directly.\n\nPlease review.\n\n@Radim - Would still be good if you can post us a list of all bad xml instances and we fix all that in this JIRA. I have none setup with me right now, but lemme know if I need to do that by self.","from":"developer"},{"body":"Lets take another approach. Standard unix command for validating POMs is part of libxml2 package:\n\nfind . -name \\*.xml -exec xmllint --valid --noout {} \\;\n\nto make xmllint validation work, you need to add proper schema to every pom which is good practice anyway.\n\n\n\nhadoop has 43 pom.xml files.","from":"developer"},{"body":"command was wrong. libxml2 does not work well with XSD schemas. C version of xerces 3 parser works fine:\n\nfind . -name *.xml -exec PParse -n -s -f -v=always {} \\;","from":"developer"},{"body":"added xml schema to maven POMs","from":"developer"},{"body":"guys, seriously commit this already. It takes you months to commit such trivial change like this.","from":"developer"},{"body":"+1 pending jenkins. Sorry for the delay Radim.","from":"developer"},{"body":"Re-upping on Radim's behalf to trigger QA Bot.","from":"developer"},{"body":"merged trunk","from":"developer"},{"body":"Reupload patch. JIRA was down, QA bot was unable to post results.","from":"developer"},{"body":"Patch had a few issues: New lines were Windows-style (trailing CRs in it), and wasn't -p0 applicable (although not an issue these days with test-patch, as it tries upto three levels).\n\nHere's a more compatible svn patch. Lets hit QA Bot with this.\n\nLocally mvn 3.x passes an install run (without tests) for me. So am +1. Will commit this in by monday regardless of QA (would be good to have it run though) unless someone else objects. The same kinda lines is already used in HBase, and thats another basis for my +1 here.\n\nThanks Radim.\n\n(Here it goes…)","from":"developer"},{"body":"Fixed filename of patch.","from":"developer"},{"body":"reupload for bot","from":"developer"},{"body":"Here is bot output. It will most likely fail to post results to JIRA.\n\nhttps://builds.apache.org/job/PreCommit-HADOOP-Build/1040/","from":"developer"},{"body":"Thanks Radim!\n\n+1 based on a local {{mvn clean install -Dtest=FooBar}} as well. Committing to branch-2 and trunk shortly.","from":"developer"},{"body":"- Committed revision 1343272 to trunk.\n- svn merge -c 1343272 to branch-2 committed as revision 1343275.\n\nThanks for this and also your continuing contributions Radim! :)","from":"developer"}],"created":"2012-04-11T11:12:45.000+0000","description":"In a few pom files there are embedded ant commands which contains '>' - redirection. This makes XML file invalid and this POM file can not be deployed into validating Maven repository managers such as Artifactory.","issue_id":"12550497","key":"HADOOP-8268","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2012-05-28T14:59:12.000+0000","role":"fixed_distractor","summary":"A few pom.xml across Hadoop project may fail XML validation"} {"case_id":"12556916","cluster":"DISTRACTOR-HADOOP-8424","comments":[{"body":"The problem was that hadoop jars were included before the webapps in the classpath for dev environments. In normal deployments scenarios, the directory structure is different and the order is maintained.","created":"2012-05-22T23:34:06.503+0000"},{"body":"Looks good overall. Three minor comments:\n1. It seems that you misplaced the comment \"for developers, add Hadoop classes to CLASSPATH\", it should be left on its original location\n2. Can you please remove the \"if exist %HADOOP_CORE_HOME%\\...\" before \"for\" loops as it is not needed\n3. Can you also please break the two if's and place: \n{code}\nfor %%i in (%HADOOP_CORE_HOME%\\build\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n)\n{code}\nunder \"for releases, add core hadoop jar & webapps to CLASSPATH\", and place:\n{code}\nfor %%i in (%HADOOP_CORE_HOME%\\build\\ivy\\lib\\Hadoop\\common\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n)\n{code}\nunder: \"add libs to CLASSPATH\" right after\n{code}for %%i in (%HADOOP_CORE_HOME%\\lib\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n){code}\nThis way, we will be more consistent with the script layout from \"bin\\hadoop\" used for non-Windows platforms.\n\nSeparately, what do you think about having a tracking Jira on making .sh and .cmd scripts fully consistent, as they aren't at the moment?\n","created":"2012-05-29T00:21:36.546+0000"},{"body":"1) Fixed\n2) Not sure thats ok because the same script is used to start stuff from a release directory structure that may not have those build directories\n3) that would mix up the dev and release class path additions. I prefer it the way it is with the 2 scenarios separated out.\n\nI would take your argument further and suggest replacing the scripts with a platform independent script.","created":"2012-06-04T22:44:21.061+0000"},{"body":"Thanks Bikas, looks good.\n\nFor #2, you would have the same result without {{if}}, as the number of iterations in the {{for}} loop would be zero.","created":"2012-06-07T00:51:16.947+0000"},{"body":"which for loop? I have a feeling we are looking at different parts of the code :)\nsince this is a trivial change that unblocks the web ui for devs, would it be ok to commit this and you can make the fixes you are alluding to? \n","created":"2012-06-07T18:27:01.120+0000"},{"body":"+1, Oh, sorry, didn't get used to commenting \"+1\" :)\n\nWhat I wanted to say is that the following:\n{code}\nif exist %HADOOP_CORE_HOME%\\build (\n for %%i in (%HADOOP_CORE_HOME%\\build\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n )\n if exist %HADOOP_CORE_HOME%\\build\\ivy\\lib\\Hadoop\\common (\n for %%i in (%HADOOP_CORE_HOME%\\build\\ivy\\lib\\Hadoop\\common\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n )\n )\n)\n{code}\n\nis equivalent to:\n\n{code}\nfor %%i in (%HADOOP_CORE_HOME%\\build\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n)\n\nfor %%i in (%HADOOP_CORE_HOME%\\build\\ivy\\lib\\Hadoop\\common\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n)\n{code}\n\nThis does not have any impact on the functionality, so it is fine to commit.","created":"2012-06-07T18:40:00.507+0000"},{"body":"Yes, we were looking at different parts!\nAttached new patch.","created":"2012-06-07T19:20:14.220+0000"},{"body":"Looks good, thanks!","created":"2012-06-07T21:55:38.412+0000"},{"body":"I just committed this. Thanks Bikas!","created":"2012-06-11T14:06:19.921+0000"}],"conversations":[{"body":"The classpath is setup to include the hadoop jars before the build webapps directory and that upsets jetty when it is trying to resolve the webapp classes.","from":"reporter","subject":"Web UI broken on Windows because classpath not setup correctly"},{"body":"The problem was that hadoop jars were included before the webapps in the classpath for dev environments. In normal deployments scenarios, the directory structure is different and the order is maintained.","from":"developer"},{"body":"Looks good overall. Three minor comments:\n1. It seems that you misplaced the comment \"for developers, add Hadoop classes to CLASSPATH\", it should be left on its original location\n2. Can you please remove the \"if exist %HADOOP_CORE_HOME%\\...\" before \"for\" loops as it is not needed\n3. Can you also please break the two if's and place: \n{code}\nfor %%i in (%HADOOP_CORE_HOME%\\build\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n)\n{code}\nunder \"for releases, add core hadoop jar & webapps to CLASSPATH\", and place:\n{code}\nfor %%i in (%HADOOP_CORE_HOME%\\build\\ivy\\lib\\Hadoop\\common\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n)\n{code}\nunder: \"add libs to CLASSPATH\" right after\n{code}for %%i in (%HADOOP_CORE_HOME%\\lib\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n){code}\nThis way, we will be more consistent with the script layout from \"bin\\hadoop\" used for non-Windows platforms.\n\nSeparately, what do you think about having a tracking Jira on making .sh and .cmd scripts fully consistent, as they aren't at the moment?\n","from":"developer"},{"body":"1) Fixed\n2) Not sure thats ok because the same script is used to start stuff from a release directory structure that may not have those build directories\n3) that would mix up the dev and release class path additions. I prefer it the way it is with the 2 scenarios separated out.\n\nI would take your argument further and suggest replacing the scripts with a platform independent script.","from":"developer"},{"body":"Thanks Bikas, looks good.\n\nFor #2, you would have the same result without {{if}}, as the number of iterations in the {{for}} loop would be zero.","from":"developer"},{"body":"which for loop? I have a feeling we are looking at different parts of the code :)\nsince this is a trivial change that unblocks the web ui for devs, would it be ok to commit this and you can make the fixes you are alluding to? \n","from":"developer"},{"body":"+1, Oh, sorry, didn't get used to commenting \"+1\" :)\n\nWhat I wanted to say is that the following:\n{code}\nif exist %HADOOP_CORE_HOME%\\build (\n for %%i in (%HADOOP_CORE_HOME%\\build\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n )\n if exist %HADOOP_CORE_HOME%\\build\\ivy\\lib\\Hadoop\\common (\n for %%i in (%HADOOP_CORE_HOME%\\build\\ivy\\lib\\Hadoop\\common\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n )\n )\n)\n{code}\n\nis equivalent to:\n\n{code}\nfor %%i in (%HADOOP_CORE_HOME%\\build\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n)\n\nfor %%i in (%HADOOP_CORE_HOME%\\build\\ivy\\lib\\Hadoop\\common\\*.jar) do (\n set CLASSPATH=!CLASSPATH!;%%i\n)\n{code}\n\nThis does not have any impact on the functionality, so it is fine to commit.","from":"developer"},{"body":"Yes, we were looking at different parts!\nAttached new patch.","from":"developer"},{"body":"Looks good, thanks!","from":"developer"},{"body":"I just committed this. Thanks Bikas!","from":"developer"}],"created":"2012-05-22T20:27:02.000+0000","description":"The classpath is setup to include the hadoop jars before the build webapps directory and that upsets jetty when it is trying to resolve the webapp classes.","issue_id":"12556916","key":"HADOOP-8424","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2012-06-11T14:06:19.000+0000","role":"fixed_distractor","summary":"Web UI broken on Windows because classpath not setup correctly"} {"case_id":"12601554","cluster":"DISTRACTOR-HADOOP-8655","comments":[{"body":"I have found a similar Bug And a fix, MAPREDUCE-4512. Please reffer the patch, and kindly encorporate the same.\nWhile fixing I too have encounted such a senario, I think this occur at the end of the buffer which would capture 4096 Charactors.\nMy understanding is the ending and begining of next buffer can and the delimiter indexses are not properly handled.\nThis is resulting in some or the other bugs.\n\nTried solving , but the fix resulted in some new bugs. The once all the senario is caught we can ensure a posible fix.","created":"2012-08-06T11:13:32.627+0000"},{"body":"A few lines of change in LineReader, also incorporaed the MAPREDUCE-4512 patch","created":"2012-08-06T11:24:34.235+0000"},{"body":"A few lines of change in LineReader, also incorporaed the MAPREDUCE-4512 patch","created":"2012-08-06T11:26:02.886+0000"},{"body":"As with MAPREDUCE-4512, I moved this to project Hadoop Common since that's where the patch needs to be applied.\n\nIn the future, please don't set the Reviewed flag unless the patch has been reviewed and approved by someone in the community. I see no record of that occurring, so I've cleared that flag. Also the Fix versions field is intended to mark where the patch has been integrated, so please don't set this field. If you'd like to indicate what versions you'd like to have the patch committed to, use the Target Versions field instead.\n","created":"2012-08-06T15:24:38.237+0000"},{"body":"Thank you Meria for including my patch as well.\nThats Jason, for have a closer look, and merging it.\nPlease guide about howt to categorize , the bug\nAs this was a issue faced in MAP REDUCE\nAnd was supposed to be raised in HADOOP","created":"2012-08-06T15:39:01.550+0000"},{"body":"Finding the right JIRA project is a straightforward mapping from the top-level projects in the code base:\n\n* Anything under hadoop-common-project maps to Hadoop Common\n* Anything under hadoop-hdfs-project maps to Hadoop HDFS\n* Anything under hadoop-mapreduce-project maps to Hadoop Map/Reduce","created":"2012-08-06T15:53:58.704+0000"},{"body":"The issue occurs when the buffer that reads the input file content, at a particular instance, ends with a character or character sequence that matches the head of the record delimiter.\n\nFor example, in the above case, while reading the file, the buffer's end bytes at an instance might be as follows,\n\n........3.\n\nThe default buffer size is 4096 bytes.Hence the input should be more than 4096 bytes and the last bytes of the buffer should match the head of the delimiter...Please guide how to create test case for the patch..\n\n\n\n \n\n","created":"2012-08-08T04:48:16.941+0000"},{"body":"Patch with test case. This patch holds good with the test case of HADOOP-8654.","created":"2012-08-16T14:33:40.658+0000"},{"body":"Comments from a quick perusal of the patch:\n\n* Please post patches with appropriate names, this is HADOOP-8655 but the patch name implies it's for HADOOP-8654.\n* Patch needs to be updated to trunk. HADOOP-8654 has been committed since this patch was posted.\n* Patch contains tabs, please convert to spaces.\n* Why were the {{InterfaceAudience}} and {{InterfaceStability}} decorators removed?\n","created":"2012-08-16T14:56:45.819+0000"},{"body":"Patch is commited aganist 04ba22681a494bf718dff7926e783c75bf64c2c7 taking care of HADOOP-8654,","created":"2012-08-16T15:43:24.173+0000"},{"body":"The code looks good, but it looks like you put TestLineReader.java under the main directory when it should be under the test directory. It will not compile under main. I also haven't had a chance to look at it in depth.","created":"2012-08-20T18:04:25.919+0000"},{"body":"Revised the patch as per Robert Joseph Evans comments","created":"2012-08-21T07:16:15.577+0000"},{"body":"Could any body clarify about \norg.apache.hadoop.ha.TestZKFailoverController Unit Test\n\n","created":"2012-08-21T09:01:40.435+0000"},{"body":"The TestZKFailoverController failure is unrelated, see HADOOP-8591.","created":"2012-08-21T13:22:10.387+0000"},{"body":"Thanks Robert Joseph Evans & Jason Lowe , for providing the info,\nIf I am not wrong, ZKFailoverController itself has a problem , and that is being reflected here.\nIf so, I hope this could be closed,\nLets listen from Arun AK, as well,\nHope his data sets would respond positevely.","created":"2012-08-21T13:58:59.600+0000"},{"body":"Gelesh,\n\nThe new patch looks better, but I still have a few comments.\n\n # Please make sure you follow the style guide. It should follow [Sun's code conventions|http://java.sun.com/docs/codeconv/] except indentation is 2 spaces, not 4. There are still tabs everywhere throughout the code and there are many lines that go over 80 characters in length. Comments are included in the 80 character limit.\n # In the test getTestData method is only called once, and is very specific to the single test method. I would prefer to see it inlined in testCustomDeliminator.\n # I appreciate that you want to explain what is happening in your code, but I don't think you need quite so many comments. For example you don't need to reference HADOOP-8654. There should be test cases added with HADOOP-8654 to validate that there were no regression. ","created":"2012-08-21T15:15:10.918+0000"},{"body":"Thank you Robert Joseph Evans,\nThis patch is updated as per your comments","created":"2012-08-22T16:51:20.937+0000"},{"body":"Since Hadoop_QA automated testing has not acted upon the previous patch, re uploading the same","created":"2012-08-23T02:28:00.337+0000"},{"body":"Gelesh :I had tried out the patch that you have posted herein. That really solves my problem. Thanks a lot for the patch. Is the patch that you re uploaded same as before? Do I need to apply this new patch?","created":"2012-08-23T16:16:27.846+0000"},{"body":"Thanks for the patch Gelesh +1.\n\nI checked this into trunk and branch-2.","created":"2012-08-23T17:00:34.581+0000"}],"conversations":[{"body":"Set textinputformat.record.delimiter as \"\"\n\nSuppose the input is a text file with the following content\n1User12User23User34User45User5\n\nMapper was expected to get value as \n\nValue 1 - 1User1\nValue 2 - 2User2\nValue 3 - 3User3\nValue 4 - 4User4\nValue 5 - 5User5\n\nAccording to this bug Mapper gets value\n\nValue 1 - entity>1User1\nValue 2 - id>2User2\nValue 3 - 3id>User3\nValue 4 - 4User4name>\nValue 5 - 5User5\n\nThe pattern shown above need not occur for value 1,2,3 necessarily. The bug occurs at some random positions in the map input.\n ","from":"reporter","subject":"In TextInputFormat, while specifying textinputformat.record.delimiter the character/character sequences in data file similar to starting character/starting character sequence in delimiter were found missing in certain cases in the Map Output"},{"body":"I have found a similar Bug And a fix, MAPREDUCE-4512. Please reffer the patch, and kindly encorporate the same.\nWhile fixing I too have encounted such a senario, I think this occur at the end of the buffer which would capture 4096 Charactors.\nMy understanding is the ending and begining of next buffer can and the delimiter indexses are not properly handled.\nThis is resulting in some or the other bugs.\n\nTried solving , but the fix resulted in some new bugs. The once all the senario is caught we can ensure a posible fix.","from":"developer"},{"body":"A few lines of change in LineReader, also incorporaed the MAPREDUCE-4512 patch","from":"developer"},{"body":"A few lines of change in LineReader, also incorporaed the MAPREDUCE-4512 patch","from":"developer"},{"body":"As with MAPREDUCE-4512, I moved this to project Hadoop Common since that's where the patch needs to be applied.\n\nIn the future, please don't set the Reviewed flag unless the patch has been reviewed and approved by someone in the community. I see no record of that occurring, so I've cleared that flag. Also the Fix versions field is intended to mark where the patch has been integrated, so please don't set this field. If you'd like to indicate what versions you'd like to have the patch committed to, use the Target Versions field instead.\n","from":"developer"},{"body":"Thank you Meria for including my patch as well.\nThats Jason, for have a closer look, and merging it.\nPlease guide about howt to categorize , the bug\nAs this was a issue faced in MAP REDUCE\nAnd was supposed to be raised in HADOOP","from":"developer"},{"body":"Finding the right JIRA project is a straightforward mapping from the top-level projects in the code base:\n\n* Anything under hadoop-common-project maps to Hadoop Common\n* Anything under hadoop-hdfs-project maps to Hadoop HDFS\n* Anything under hadoop-mapreduce-project maps to Hadoop Map/Reduce","from":"developer"},{"body":"The issue occurs when the buffer that reads the input file content, at a particular instance, ends with a character or character sequence that matches the head of the record delimiter.\n\nFor example, in the above case, while reading the file, the buffer's end bytes at an instance might be as follows,\n\n........3.\n\nThe default buffer size is 4096 bytes.Hence the input should be more than 4096 bytes and the last bytes of the buffer should match the head of the delimiter...Please guide how to create test case for the patch..\n\n\n\n \n\n","from":"developer"},{"body":"Patch with test case. This patch holds good with the test case of HADOOP-8654.","from":"developer"},{"body":"Comments from a quick perusal of the patch:\n\n* Please post patches with appropriate names, this is HADOOP-8655 but the patch name implies it's for HADOOP-8654.\n* Patch needs to be updated to trunk. HADOOP-8654 has been committed since this patch was posted.\n* Patch contains tabs, please convert to spaces.\n* Why were the {{InterfaceAudience}} and {{InterfaceStability}} decorators removed?\n","from":"developer"},{"body":"Patch is commited aganist 04ba22681a494bf718dff7926e783c75bf64c2c7 taking care of HADOOP-8654,","from":"developer"},{"body":"The code looks good, but it looks like you put TestLineReader.java under the main directory when it should be under the test directory. It will not compile under main. I also haven't had a chance to look at it in depth.","from":"developer"},{"body":"Revised the patch as per Robert Joseph Evans comments","from":"developer"},{"body":"Could any body clarify about \norg.apache.hadoop.ha.TestZKFailoverController Unit Test\n\n","from":"developer"},{"body":"The TestZKFailoverController failure is unrelated, see HADOOP-8591.","from":"developer"},{"body":"Thanks Robert Joseph Evans & Jason Lowe , for providing the info,\nIf I am not wrong, ZKFailoverController itself has a problem , and that is being reflected here.\nIf so, I hope this could be closed,\nLets listen from Arun AK, as well,\nHope his data sets would respond positevely.","from":"developer"},{"body":"Gelesh,\n\nThe new patch looks better, but I still have a few comments.\n\n # Please make sure you follow the style guide. It should follow [Sun's code conventions|http://java.sun.com/docs/codeconv/] except indentation is 2 spaces, not 4. There are still tabs everywhere throughout the code and there are many lines that go over 80 characters in length. Comments are included in the 80 character limit.\n # In the test getTestData method is only called once, and is very specific to the single test method. I would prefer to see it inlined in testCustomDeliminator.\n # I appreciate that you want to explain what is happening in your code, but I don't think you need quite so many comments. For example you don't need to reference HADOOP-8654. There should be test cases added with HADOOP-8654 to validate that there were no regression. ","from":"developer"},{"body":"Thank you Robert Joseph Evans,\nThis patch is updated as per your comments","from":"developer"},{"body":"Since Hadoop_QA automated testing has not acted upon the previous patch, re uploading the same","from":"developer"},{"body":"Gelesh :I had tried out the patch that you have posted herein. That really solves my problem. Thanks a lot for the patch. Is the patch that you re uploaded same as before? Do I need to apply this new patch?","from":"developer"},{"body":"Thanks for the patch Gelesh +1.\n\nI checked this into trunk and branch-2.","from":"developer"}],"created":"2012-08-06T11:01:33.000+0000","description":"Set textinputformat.record.delimiter as \"\"\n\nSuppose the input is a text file with the following content\n1User12User23User34User45User5\n\nMapper was expected to get value as \n\nValue 1 - 1User1\nValue 2 - 2User2\nValue 3 - 3User3\nValue 4 - 4User4\nValue 5 - 5User5\n\nAccording to this bug Mapper gets value\n\nValue 1 - entity>1User1\nValue 2 - id>2User2\nValue 3 - 3id>User3\nValue 4 - 4User4name>\nValue 5 - 5User5\n\nThe pattern shown above need not occur for value 1,2,3 necessarily. The bug occurs at some random positions in the map input.\n ","issue_id":"12601554","key":"HADOOP-8655","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2012-08-23T17:00:34.000+0000","role":"fixed_distractor","summary":"In TextInputFormat, while specifying textinputformat.record.delimiter the character/character sequences in data file similar to starting character/starting character sequence in delimiter were found missing in certain cases in the Map Output"} {"case_id":"12330363","cluster":"DISTRACTOR-HADOOP-89","comments":[{"body":"This patch makes a file visible in the file system as soon as it is created by an application. However, the data blocks are associated with a file when the file gets closed.\n\nIf a DFS client A has created a file and is writing data to it, another DFS client will see the file size as zero until A closes the file.","created":"2007-08-06T19:56:45.123+0000"},{"body":"I see 2 problems with this patch:\n# It does not close any of the 3 issues raised in the description above.\nYou can see files in the ls, but you cannot see file progress and you still loose all accumulated data if the file is not closed.\n# Implementation-wise I would expect that when files are made visible we will get rid of the pending create structure, \nbut it is still there, so now you will have to synchronize pending create data with the corresponding inode.\n\nI'd propose to make files both visible and readable while they are still being created. Namely, we should let\nclients read portions of the file that have been written and replicated on data-nodes.","created":"2007-08-09T21:59:10.869+0000"},{"body":"This patch does not change APIs and disk formats. The idea is that files that are created by clients appear immediately in the namespace. Other clients can see the file and get its attributes. This is a big change in semantics and I would rather do this in the beginning of a release than towards the end of the release cycle. This is a step in implementing full-fledged \"appends\" for HDFS files (HADOOP-1700).\n\nRegarding \"removing pending creates\" structures, I agree that they should move into a new kind of inode. I can do it after you introduce the concept of class-hierarchy of inodes (HADOOP-1687).\n\nPlease let me know if the above sounds good to you and makes you agree that this patch is commit-able.\n\n\n","created":"2007-08-10T00:31:38.187+0000"},{"body":"1. Contents of new blocks that are appended to a file is visible to clients as soon as the datanode reports the block to the namenode. This means that data is visible to clients even when the block metadata is not yet persisted on disk. This approach lets us avoid a fs-transaction into the edit log for every new block allocation.\n\n2. The block allocation for a file is persisted in the edit log when the file is closed.\n\n3. A new API FSDataOutputStream.sync() allows an application to make data persistent on disk even before the file is closed. The invocation of this API causes a transaction to be logged into the edits log to record the blocks that are currently allocated to the file. An application that is recording data to a log file will periodically invoke this API to ensure that the contents of the log file persist even if the application dies before closing the file.\n\n4. The FsShell utility has a new command that is invoked as \"bin/hadoop dfs -tail [-f] \". When the \"-f\" option is used, the FsShell utility will periodically poll for changes to the filesize. When a filesize change is detected, it will re-open the file and will display the new contents that were added to the file.\n","created":"2007-08-28T06:59:30.682+0000"},{"body":"This patch implements 1, 2 and 4 of the above.","created":"2007-09-04T19:40:07.792+0000"},{"body":"merged patch with latest trunk.","created":"2007-09-06T18:22:23.777+0000"},{"body":"This patch does not apply anymore.\n\nWith HADOOP-1700 on the horizon does it make sense to introduce tail -f here?\nTail -f works until the client that is writing to the file is alive. If it dies before closing the file all information\nreported by tail -f does not exist in the system anymore. So in a way tail -f reports an illusive data that is not \nguaranteed to be present in the future.\nThis patch has a lot of important internal changes, like removing pendingCreates etc.\nBut introducing tail -f at this point may cause confusion among potential users.\n","created":"2007-09-12T02:19:40.359+0000"},{"body":"Merged patch with latest trunk. Konstantin's view was that user's might get confused if they can view data in a file that gets thrown out if the namenode restarts. Keeping with his view, i have removed the \"tail\" command from this patch.\n\nA follow-on patch will introduce a new API \"flush\" that will persist blocks before file is closed. ","created":"2007-09-12T21:27:06.551+0000"},{"body":"+1\nMy only concern is that in order to add or remove blocks you reallocate the array of blocks\nso that it's length always equals the number of blocks in the file.\nI tested this patch with small files (1 or 2 blocks) and did not see any degradation in performance.\nSince our average file size is 1.5 blocks this is good. \nIMO we should still test it for bigger files to make sure there is no slow-down.","created":"2007-09-18T01:52:09.269+0000"},{"body":"1. merged patch with latest trunk.\n2. Re-introduced the \"tail\" command from dfs shell.","created":"2007-09-18T18:07:33.427+0000"},{"body":"Marking it Patch Available to trigger automatic tests,","created":"2007-09-18T18:09:48.365+0000"},{"body":"> 2. Re-introduced the \"tail\" command from dfs shell.\n\nWhat is the point of reviews and discussions then?","created":"2007-09-19T01:20:56.757+0000"},{"body":"The contrib test failure si nto related to this patch, being tracked by HADOOP-1924","created":"2007-09-19T21:30:11.295+0000"},{"body":"Dhruba convinced me that it is worth committing the patch with the tail functionality in it.\nThere are 2 types of failures that lead to a loss of entire file data. \n# The name-node failure, and \n# the client failure.\n\nIf the name-node dies the length of an incomplete file will be set to 0, which correspond to the current behavior, when we just loose the entire file.\nIf the client dies the name-node automatically closes all files created by the client as long as it detects the client lease expiration.\nThe last one is the most common case of failure, and the code provides protection for from loosing data in the case.","created":"2007-09-19T22:01:02.073+0000"},{"body":"I just committed this.","created":"2007-09-19T22:14:18.973+0000"},{"body":"This feature alters user-facing behavior. Until now, a file's existence was an indication that the task producing the file is complete. Please mark it such in release notes. Thanks.","created":"2007-09-25T21:15:38.449+0000"},{"body":"The CHANGES.txt file lists this issue under the section titled \"NEW FEATURES\". So, this should figure prominently in the release notes.","created":"2007-09-25T21:50:28.350+0000"},{"body":"It is actually HADOOP-1708 that introduced the feature that \"files appear in namespace as soon as they are created\". I changed the release notes (CHANGES.txt) to ensure that this change appears in the INCOMPATIBLE section rather than in NEW FEATURES section.","created":"2007-10-01T07:09:15.363+0000"},{"body":"That was what was decided in the last group meeting between you, me, Owen,\nSameer, etc. That's why I introduced the changes to FsShell.java.\n\n\n\n\n","created":"2007-11-03T05:42:13.636+0000"}],"conversations":[{"body":"the current behaviour, whereby a file is not visible until it is closed has several flaws,including:\n1. no practical way to know if a file/job is progressing\n2. no way to implement files that never close, such as log files\n3. failure to close a file results in loss of the file\n\nThe part of the file that's written should be visible.","from":"reporter","subject":"files are not visible until they are closed"},{"body":"This patch makes a file visible in the file system as soon as it is created by an application. However, the data blocks are associated with a file when the file gets closed.\n\nIf a DFS client A has created a file and is writing data to it, another DFS client will see the file size as zero until A closes the file.","from":"developer"},{"body":"I see 2 problems with this patch:\n# It does not close any of the 3 issues raised in the description above.\nYou can see files in the ls, but you cannot see file progress and you still loose all accumulated data if the file is not closed.\n# Implementation-wise I would expect that when files are made visible we will get rid of the pending create structure, \nbut it is still there, so now you will have to synchronize pending create data with the corresponding inode.\n\nI'd propose to make files both visible and readable while they are still being created. Namely, we should let\nclients read portions of the file that have been written and replicated on data-nodes.","from":"developer"},{"body":"This patch does not change APIs and disk formats. The idea is that files that are created by clients appear immediately in the namespace. Other clients can see the file and get its attributes. This is a big change in semantics and I would rather do this in the beginning of a release than towards the end of the release cycle. This is a step in implementing full-fledged \"appends\" for HDFS files (HADOOP-1700).\n\nRegarding \"removing pending creates\" structures, I agree that they should move into a new kind of inode. I can do it after you introduce the concept of class-hierarchy of inodes (HADOOP-1687).\n\nPlease let me know if the above sounds good to you and makes you agree that this patch is commit-able.\n\n\n","from":"developer"},{"body":"1. Contents of new blocks that are appended to a file is visible to clients as soon as the datanode reports the block to the namenode. This means that data is visible to clients even when the block metadata is not yet persisted on disk. This approach lets us avoid a fs-transaction into the edit log for every new block allocation.\n\n2. The block allocation for a file is persisted in the edit log when the file is closed.\n\n3. A new API FSDataOutputStream.sync() allows an application to make data persistent on disk even before the file is closed. The invocation of this API causes a transaction to be logged into the edits log to record the blocks that are currently allocated to the file. An application that is recording data to a log file will periodically invoke this API to ensure that the contents of the log file persist even if the application dies before closing the file.\n\n4. The FsShell utility has a new command that is invoked as \"bin/hadoop dfs -tail [-f] \". When the \"-f\" option is used, the FsShell utility will periodically poll for changes to the filesize. When a filesize change is detected, it will re-open the file and will display the new contents that were added to the file.\n","from":"developer"},{"body":"This patch implements 1, 2 and 4 of the above.","from":"developer"},{"body":"merged patch with latest trunk.","from":"developer"},{"body":"This patch does not apply anymore.\n\nWith HADOOP-1700 on the horizon does it make sense to introduce tail -f here?\nTail -f works until the client that is writing to the file is alive. If it dies before closing the file all information\nreported by tail -f does not exist in the system anymore. So in a way tail -f reports an illusive data that is not \nguaranteed to be present in the future.\nThis patch has a lot of important internal changes, like removing pendingCreates etc.\nBut introducing tail -f at this point may cause confusion among potential users.\n","from":"developer"},{"body":"Merged patch with latest trunk. Konstantin's view was that user's might get confused if they can view data in a file that gets thrown out if the namenode restarts. Keeping with his view, i have removed the \"tail\" command from this patch.\n\nA follow-on patch will introduce a new API \"flush\" that will persist blocks before file is closed. ","from":"developer"},{"body":"+1\nMy only concern is that in order to add or remove blocks you reallocate the array of blocks\nso that it's length always equals the number of blocks in the file.\nI tested this patch with small files (1 or 2 blocks) and did not see any degradation in performance.\nSince our average file size is 1.5 blocks this is good. \nIMO we should still test it for bigger files to make sure there is no slow-down.","from":"developer"},{"body":"1. merged patch with latest trunk.\n2. Re-introduced the \"tail\" command from dfs shell.","from":"developer"},{"body":"Marking it Patch Available to trigger automatic tests,","from":"developer"},{"body":"> 2. Re-introduced the \"tail\" command from dfs shell.\n\nWhat is the point of reviews and discussions then?","from":"developer"},{"body":"The contrib test failure si nto related to this patch, being tracked by HADOOP-1924","from":"developer"},{"body":"Dhruba convinced me that it is worth committing the patch with the tail functionality in it.\nThere are 2 types of failures that lead to a loss of entire file data. \n# The name-node failure, and \n# the client failure.\n\nIf the name-node dies the length of an incomplete file will be set to 0, which correspond to the current behavior, when we just loose the entire file.\nIf the client dies the name-node automatically closes all files created by the client as long as it detects the client lease expiration.\nThe last one is the most common case of failure, and the code provides protection for from loosing data in the case.","from":"developer"},{"body":"I just committed this.","from":"developer"},{"body":"This feature alters user-facing behavior. Until now, a file's existence was an indication that the task producing the file is complete. Please mark it such in release notes. Thanks.","from":"developer"},{"body":"The CHANGES.txt file lists this issue under the section titled \"NEW FEATURES\". So, this should figure prominently in the release notes.","from":"developer"},{"body":"It is actually HADOOP-1708 that introduced the feature that \"files appear in namespace as soon as they are created\". I changed the release notes (CHANGES.txt) to ensure that this change appears in the INCOMPATIBLE section rather than in NEW FEATURES section.","from":"developer"},{"body":"That was what was decided in the last group meeting between you, me, Owen,\nSameer, etc. That's why I introduced the changes to FsShell.java.\n\n\n\n\n","from":"developer"}],"created":"2006-03-17T10:16:35.000+0000","description":"the current behaviour, whereby a file is not visible until it is closed has several flaws,including:\n1. no practical way to know if a file/job is progressing\n2. no way to implement files that never close, such as log files\n3. failure to close a file results in loss of the file\n\nThe part of the file that's written should be visible.","issue_id":"12330363","key":"HADOOP-89","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2007-09-19T22:14:20.000+0000","role":"fixed_distractor","summary":"files are not visible until they are closed"} {"case_id":"12361016","cluster":"DISTRACTOR-HADOOP-917","comments":[{"body":"Attached patch reproduces the NPE in trunk version of hadoop:\n\njava.lang.NullPointerException\n at org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue.merge(SequenceFile.java:2392)\n at org.apache.hadoop.io.SequenceFile$Sorter.merge(SequenceFile.java:2087)\n at org.apache.hadoop.mapred.MapTask$MapOutputBuffer.mergeParts(MapTask.java:498)\n at org.apache.hadoop.mapred.MapTask.run(MapTask.java:191)\n at org.apache.hadoop.mapred.LocalJobRunner$Job.run(LocalJobRunner.java:109)\n","created":"2007-01-29T19:57:51.477+0000"},{"body":"It looks like if a map output sorter has to do a multi-way merge, it crashes with a null pointer trying to create the temporary file name.","created":"2007-02-06T21:49:58.413+0000"},{"body":"The problem was that the merge code was assuming that outputFile was set and it wasn't in that context. I've changed the API to the merge code so that the methods that don't have an output file pass in the tmpDirectory where files should be created.","created":"2007-02-06T23:11:21.755+0000"},{"body":"I just committed this. Thanks, Owen!","created":"2007-02-07T00:03:04.961+0000"},{"body":"Thanks Owen! patched fixed the problem. ","created":"2007-02-07T00:25:17.378+0000"}],"conversations":[{"body":"After nutch started using hadoop 0.10.1 the following Exception started to appear:\n\njava.lang.NullPointerException\n\tat org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue.merge(SequenceFile.java:2158)\n\tat org.apache.hadoop.io.SequenceFile$Sorter.merge(SequenceFile.java:1892)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.mergeParts(MapTask.java:498)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:191)\n\tat org.apache.hadoop.mapred.TaskTracker$Child.main(TaskTracker.java:1367)\n\nAnyone know the cure?","from":"reporter","subject":"NPE in org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue"},{"body":"Attached patch reproduces the NPE in trunk version of hadoop:\n\njava.lang.NullPointerException\n at org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue.merge(SequenceFile.java:2392)\n at org.apache.hadoop.io.SequenceFile$Sorter.merge(SequenceFile.java:2087)\n at org.apache.hadoop.mapred.MapTask$MapOutputBuffer.mergeParts(MapTask.java:498)\n at org.apache.hadoop.mapred.MapTask.run(MapTask.java:191)\n at org.apache.hadoop.mapred.LocalJobRunner$Job.run(LocalJobRunner.java:109)\n","from":"developer"},{"body":"It looks like if a map output sorter has to do a multi-way merge, it crashes with a null pointer trying to create the temporary file name.","from":"developer"},{"body":"The problem was that the merge code was assuming that outputFile was set and it wasn't in that context. I've changed the API to the merge code so that the methods that don't have an output file pass in the tmpDirectory where files should be created.","from":"developer"},{"body":"I just committed this. Thanks, Owen!","from":"developer"},{"body":"Thanks Owen! patched fixed the problem. ","from":"developer"}],"created":"2007-01-22T19:10:05.000+0000","description":"After nutch started using hadoop 0.10.1 the following Exception started to appear:\n\njava.lang.NullPointerException\n\tat org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue.merge(SequenceFile.java:2158)\n\tat org.apache.hadoop.io.SequenceFile$Sorter.merge(SequenceFile.java:1892)\n\tat org.apache.hadoop.mapred.MapTask$MapOutputBuffer.mergeParts(MapTask.java:498)\n\tat org.apache.hadoop.mapred.MapTask.run(MapTask.java:191)\n\tat org.apache.hadoop.mapred.TaskTracker$Child.main(TaskTracker.java:1367)\n\nAnyone know the cure?","issue_id":"12361016","key":"HADOOP-917","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2007-02-07T00:03:04.000+0000","role":"fixed_distractor","summary":"NPE in org.apache.hadoop.io.SequenceFile$Sorter$MergeQueue"} {"case_id":"12330368","cluster":"DISTRACTOR-HADOOP-92","comments":[{"body":"specifically, I'm hoping that we can create a well known file in the output directory with a log of all interesting job output.\n\nPerhaps we should put it in a subdirectory /INFO/ of the output directory, so that when the directory is used for future input, it is not confused with data? This will also allow us to add other things later.","created":"2006-03-17T13:12:16.000+0000"},{"body":"It could be done with a logs directory in output_dir. So, the output_dir/logs/tasks/ would contain all the information per task (a file per task). This file would contain information like -- machine name, start time, end time, result status, error messages. Also, there should be a file per job output_dir/logs/job_id.log. This file would contains job specific data -- no of mapreduce tasks that were running at some time, number of machines it was running on, and the start time and end time of the job. This log could also contain information on which of the task failed, so that the user can take a look at the respective task log to know why it failed.","created":"2006-03-22T05:05:42.000+0000"},{"body":"Another approach to this would be to expose this through the JobClient API. Job-related events can be reported to the job client. Events can be queued in the job tracker and the JobClient can retrieve them as it polls for job status. Then the JobClient can decide where to log them. By default they can be logged to standard error.\n\nThe events I think one might care about are:\n - task start (task_id, type & host)\n - task completion (task_id)\n - task failure (task_id, error message)\n\nThe jobtracker already tracks most of this, so I don't think this places a huge new burden on the jobtracker.\n\nI don't like polluting the job's output directory with log data since it would require changes to the InputFormat implementations and other code to make them skip this specially named sub-directory (unless the name begins with a dot, which the fs code already ignores).\n\nIn any case, we could add an option to JobClient to log to the job's output fs.\n","created":"2006-03-23T04:02:36.000+0000"},{"body":"here is patch that reports the machine on which a specific task failed. Clicking on the jobid takes you to a page where the information is displayed as before (with a little more parsing). But each of the task's is made clickable to show all the attempts that were made to execute this task. On clikcing on this task, you can see the machine on which a particular task attempt was executed and what the error was.","created":"2006-03-30T07:37:43.000+0000"},{"body":"I am including a patch that improves upon the job tracker interface. The tipid is made clickable so that it shows all the attempts for a particular task with the machine name. I found it useful, so am uplodaing it again.","created":"2006-04-13T04:24:30.000+0000"},{"body":"I took this patch for a spin. +1 on commit.","created":"2006-04-13T05:01:40.000+0000"},{"body":"I've been using it today too. +1 on commit.","created":"2006-04-13T05:20:00.000+0000"},{"body":"This mostly looks good. I committed it with the following changes:\n\n1. TaskStatus is not a public class, so a public method in another public class should not return it. For now, I just made the method package-private, since the jsp pages are compiled in the same package. Longer-term we should probably copy this information into the TaskReport, the public version of a TaskStatus, or something. This information should be available to other applications through a public API, and through an RPC in the JobSubmissionProtocol.\n\n2. I renamed JobTracker.getallTaskStatus to be getTaskStatuses, like the JobInProgress method, and also using correct camel-case.\n\n3. Rather than add a new TaskStatus contructor, leaving an old one that is no longer called, I removed the old one. We don't need dead code around. This is package-private, so we can be sure that no code outside of this package uses the old constructor.\n\n4. The patch removed some inter-method whitespace. I restored it.\n","created":"2006-04-13T07:04:45.000+0000"}],"conversations":[{"body":"Currently Mapreduce does not tell you which machine failed to execute the task. Also, it would be nice to have features wherein there is a log report with each job, saying the number of tasks it ran (reporting which one failed and on which machine, listing any error information it can) with the start/end/execute time of each task. ","from":"reporter","subject":"Error Reporting/logging in MapReduce"},{"body":"specifically, I'm hoping that we can create a well known file in the output directory with a log of all interesting job output.\n\nPerhaps we should put it in a subdirectory /INFO/ of the output directory, so that when the directory is used for future input, it is not confused with data? This will also allow us to add other things later.","from":"developer"},{"body":"It could be done with a logs directory in output_dir. So, the output_dir/logs/tasks/ would contain all the information per task (a file per task). This file would contain information like -- machine name, start time, end time, result status, error messages. Also, there should be a file per job output_dir/logs/job_id.log. This file would contains job specific data -- no of mapreduce tasks that were running at some time, number of machines it was running on, and the start time and end time of the job. This log could also contain information on which of the task failed, so that the user can take a look at the respective task log to know why it failed.","from":"developer"},{"body":"Another approach to this would be to expose this through the JobClient API. Job-related events can be reported to the job client. Events can be queued in the job tracker and the JobClient can retrieve them as it polls for job status. Then the JobClient can decide where to log them. By default they can be logged to standard error.\n\nThe events I think one might care about are:\n - task start (task_id, type & host)\n - task completion (task_id)\n - task failure (task_id, error message)\n\nThe jobtracker already tracks most of this, so I don't think this places a huge new burden on the jobtracker.\n\nI don't like polluting the job's output directory with log data since it would require changes to the InputFormat implementations and other code to make them skip this specially named sub-directory (unless the name begins with a dot, which the fs code already ignores).\n\nIn any case, we could add an option to JobClient to log to the job's output fs.\n","from":"developer"},{"body":"here is patch that reports the machine on which a specific task failed. Clicking on the jobid takes you to a page where the information is displayed as before (with a little more parsing). But each of the task's is made clickable to show all the attempts that were made to execute this task. On clikcing on this task, you can see the machine on which a particular task attempt was executed and what the error was.","from":"developer"},{"body":"I am including a patch that improves upon the job tracker interface. The tipid is made clickable so that it shows all the attempts for a particular task with the machine name. I found it useful, so am uplodaing it again.","from":"developer"},{"body":"I took this patch for a spin. +1 on commit.","from":"developer"},{"body":"I've been using it today too. +1 on commit.","from":"developer"},{"body":"This mostly looks good. I committed it with the following changes:\n\n1. TaskStatus is not a public class, so a public method in another public class should not return it. For now, I just made the method package-private, since the jsp pages are compiled in the same package. Longer-term we should probably copy this information into the TaskReport, the public version of a TaskStatus, or something. This information should be available to other applications through a public API, and through an RPC in the JobSubmissionProtocol.\n\n2. I renamed JobTracker.getallTaskStatus to be getTaskStatuses, like the JobInProgress method, and also using correct camel-case.\n\n3. Rather than add a new TaskStatus contructor, leaving an old one that is no longer called, I removed the old one. We don't need dead code around. This is package-private, so we can be sure that no code outside of this package uses the old constructor.\n\n4. The patch removed some inter-method whitespace. I restored it.\n","from":"developer"}],"created":"2006-03-17T12:48:15.000+0000","description":"Currently Mapreduce does not tell you which machine failed to execute the task. Also, it would be nice to have features wherein there is a log report with each job, saying the number of tasks it ran (reporting which one failed and on which machine, listing any error information it can) with the start/end/execute time of each task. ","issue_id":"12330368","key":"HADOOP-92","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2006-04-13T07:04:45.000+0000","role":"fixed_distractor","summary":"Error Reporting/logging in MapReduce"} {"case_id":"12643456","cluster":"DISTRACTOR-HADOOP-9485","comments":[{"body":"Let's also add a unit test, to make sure this doesn't happen again.","created":"2013-04-18T23:19:51.814+0000"},{"body":"I thought about this a little more, and I think changing the default socket factory class is probably a bit too risky.\n\nIt's better to set compatible defaults in the code. So we simply have the code default {{dfs.client.use.legacy.blockreader}} to true, which matches any socket factory the user may have.\n\nIn {{hdfs-default.xml}}, in contrast, we set {{dfs.client.use.legacy.blockreader}} to false, reflecting the fact that we also set \n{{hadoop.rpc.socket.factory.class.default}} to {{org.apache.hadoop.net.StandardSocketFactory}} in core-default.xml.\n\nThis fixes the problem, and has the least chance of breaking any existing users.","created":"2013-04-29T21:48:17.490+0000"},{"body":"don't change hadoop.rpc.socket.factory.class.default","created":"2013-04-30T21:27:58.002+0000"},{"body":"this had to be updated due to HDFS-4305 getting committed in the meantime.","created":"2013-05-01T00:42:55.444+0000"},{"body":"I disagree with the approach of having differing defaults between the code and the hdfs-default.xml. Why is it risky to change the default in the code to match the default from the XML file? 99.9% of non-buggy users should be using hdfs-default.xml (which is shipped in our jars) or manually overriding. If someone is relying on the behavior with the XML file not present, I don't think we need to maintain that case.","created":"2013-05-02T06:10:58.943+0000"},{"body":"I think the issue here is that running without XML defaults works in branch-1, and when upgrading to branch-2 or a distribution based on that, people view this kind of regression as a stumbling block.\n\nbq. Why is it risky to change the default in the code to match the default from the XML file? \n\nCurrently, you can set {{hadoop.rpc.socket.factory.class.default = null}} and we'll use {{SocketFactory.getDefault()}}. If we change this behavior we may break users who are relying on it. It certainly breaks {{TestSaslRPC}}. It would be an incompatible change that would probably cause problems for users.\n\nI suppose we could make the code default to the same thing as the XML file, but not change the meaning of setting the key to null. That might work.","created":"2013-05-02T18:06:44.443+0000"},{"body":"Yea, I _think_ we can differentiate between \"explicitly set to null\" and \"not explicitly set\". Either way, I think I'd rather break one or two folks who are setting to null to get that behavior than have a default against the legacy blockreader which we'd like to eventually remove.","created":"2013-05-03T15:56:59.475+0000"},{"body":"This laest patch (which Jenkins hasn't run yet) should do exactly that. Only by explicitly setting it to null or empty will you get the null behavior.","created":"2013-05-03T17:12:26.233+0000"},{"body":"+1, the latest patch looks good to me. I'm going to commit this momentarily.","created":"2013-05-10T21:40:06.167+0000"},{"body":"I've just committed this to trunk and branch-2.\n\nThanks a lot for the contribution, Colin.","created":"2013-05-10T21:50:08.862+0000"}],"conversations":[{"body":"In {{core-default.xml}}, {{hadoop.rpc.socket.factory.class.default}} defaults to {{org.apache.hadoop.net.StandardSocketFactory}}. However, in {{CommonConfigurationKeysPublic.java}}, there is no default for this key. This is inconsistent (defaults in the code versus defaults in the XML files should match.) It also leads to problems with {{RemoteBlockReader2}}, since the default {{SocketFactory}} creates a {{Socket}} without an associated channel. {{RemoteBlockReader2}} cannot use such a {{Socket}}.\n\nThis bug only really becomes apparent when you create a {{Configuration}} using the {{Configuration(loadDefaults=true)}} constructor. Thanks to AB Srinivasan for his help in discovering this bug.","from":"reporter","subject":"No default value in the code for hadoop.rpc.socket.factory.class.default"},{"body":"Let's also add a unit test, to make sure this doesn't happen again.","from":"developer"},{"body":"I thought about this a little more, and I think changing the default socket factory class is probably a bit too risky.\n\nIt's better to set compatible defaults in the code. So we simply have the code default {{dfs.client.use.legacy.blockreader}} to true, which matches any socket factory the user may have.\n\nIn {{hdfs-default.xml}}, in contrast, we set {{dfs.client.use.legacy.blockreader}} to false, reflecting the fact that we also set \n{{hadoop.rpc.socket.factory.class.default}} to {{org.apache.hadoop.net.StandardSocketFactory}} in core-default.xml.\n\nThis fixes the problem, and has the least chance of breaking any existing users.","from":"developer"},{"body":"don't change hadoop.rpc.socket.factory.class.default","from":"developer"},{"body":"this had to be updated due to HDFS-4305 getting committed in the meantime.","from":"developer"},{"body":"I disagree with the approach of having differing defaults between the code and the hdfs-default.xml. Why is it risky to change the default in the code to match the default from the XML file? 99.9% of non-buggy users should be using hdfs-default.xml (which is shipped in our jars) or manually overriding. If someone is relying on the behavior with the XML file not present, I don't think we need to maintain that case.","from":"developer"},{"body":"I think the issue here is that running without XML defaults works in branch-1, and when upgrading to branch-2 or a distribution based on that, people view this kind of regression as a stumbling block.\n\nbq. Why is it risky to change the default in the code to match the default from the XML file? \n\nCurrently, you can set {{hadoop.rpc.socket.factory.class.default = null}} and we'll use {{SocketFactory.getDefault()}}. If we change this behavior we may break users who are relying on it. It certainly breaks {{TestSaslRPC}}. It would be an incompatible change that would probably cause problems for users.\n\nI suppose we could make the code default to the same thing as the XML file, but not change the meaning of setting the key to null. That might work.","from":"developer"},{"body":"Yea, I _think_ we can differentiate between \"explicitly set to null\" and \"not explicitly set\". Either way, I think I'd rather break one or two folks who are setting to null to get that behavior than have a default against the legacy blockreader which we'd like to eventually remove.","from":"developer"},{"body":"This laest patch (which Jenkins hasn't run yet) should do exactly that. Only by explicitly setting it to null or empty will you get the null behavior.","from":"developer"},{"body":"+1, the latest patch looks good to me. I'm going to commit this momentarily.","from":"developer"},{"body":"I've just committed this to trunk and branch-2.\n\nThanks a lot for the contribution, Colin.","from":"developer"}],"created":"2013-04-18T22:56:51.000+0000","description":"In {{core-default.xml}}, {{hadoop.rpc.socket.factory.class.default}} defaults to {{org.apache.hadoop.net.StandardSocketFactory}}. However, in {{CommonConfigurationKeysPublic.java}}, there is no default for this key. This is inconsistent (defaults in the code versus defaults in the XML files should match.) It also leads to problems with {{RemoteBlockReader2}}, since the default {{SocketFactory}} creates a {{Socket}} without an associated channel. {{RemoteBlockReader2}} cannot use such a {{Socket}}.\n\nThis bug only really becomes apparent when you create a {{Configuration}} using the {{Configuration(loadDefaults=true)}} constructor. Thanks to AB Srinivasan for his help in discovering this bug.","issue_id":"12643456","key":"HADOOP-9485","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2013-05-10T21:50:08.000+0000","role":"fixed_distractor","summary":"No default value in the code for hadoop.rpc.socket.factory.class.default"} {"case_id":"12662634","cluster":"DISTRACTOR-HADOOP-9851","comments":[{"body":"This is big a problem if you use winbind along with hadoop security setup.","created":"2016-06-02T21:00:45.233+0000"},{"body":"Patch 01:\r\n\"-\" sign has to remain at 1st place otherwise it will be compiled as range","created":"2017-10-20T14:43:06.300+0000"},{"body":"Bulk update: moved all 3.2.0 non-blocker issues, please move back if it is a blocker.","created":"2018-11-23T11:55:33.943+0000"},{"body":"Bulk update: moved all 3.3.0 non-blocker issues, please move back if it is a blocker.","created":"2020-04-10T18:15:38.311+0000"},{"body":"Thanx [~boky01] for the patch.\r\n Well I think adding support for + shouldn't be a problem, but the pattern is there since quite long. Not sure whether they didn't add support for \"+\" deliberately or not.. \r\n Since it isn't going to specific to HDFS, would be good if we get some more opinions\r\n[~stevel@apache.org] [~aajisaka] any thoughts on this?","created":"2020-06-05T15:51:26.790+0000"},{"body":"[~boky01] seems you allowed + for windows as well, I think windows doesn't support + in user name.\r\nCan you check once, I checked in Linux and it allows + in user name. if not, we can just allow for linux and this should be good to go then.","created":"2020-06-10T05:54:23.654+0000"},{"body":"[~ayushtkn],\r\nWindows remained unchanged only Linux will allow + sign.","created":"2020-06-15T11:37:08.600+0000"},{"body":"Can you give a check to the checkstyle complains\nTest failures seems unrelated.\n+1 once checkstyle issue is fixed\n","created":"2020-06-15T22:12:44.800+0000"},{"body":"[~ayushtkn],\r\nThe checkstyle is not caused by my patch. The indentation was wrong even before my patch.\r\nI did not fix it because I did not want bigger patch than needed and indentation fixes decrease the readability in diff tools. But I am not sure what is the best practice here.","created":"2020-06-16T10:28:57.696+0000"},{"body":"Should be fine then, we can keep the existing indentation to match the rest of the file","created":"2020-06-16T10:49:30.514+0000"},{"body":"Committed to trunk\r\n\r\nThanx [~boky01]  for the fix and [~h0tbird]  for the report!!!","created":"2020-06-17T08:34:38.378+0000"}],"conversations":[{"body":"I intend to set user and group:\n\n*User:* _MYCOMPANY+marc.villacorta_\n*Group:* hadoop\n\nwhere _'+'_ is what we use as a winbind separator.\n\nAnd this is what I get:\n{code:none}\nsudo -u hdfs hadoop fs -touchz /tmp/test.txt\nsudo -u hdfs hadoop fs -chown MYCOMPANY+marc.villacorta:hadoop /tmp/test.txt\n-chown: 'MYCOMPANY+marc.villacorta:hadoop' does not match expected pattern for [owner][:group].\nUsage: hadoop fs [generic options] -chown [-R] [OWNER][:[GROUP]] PATH...\n{code}\n\nI am using version: 2.0.0-cdh4.3.0\n\nQuote [source|http://h30097.www3.hp.com/docs/iass/OSIS_62/MAN/MAN8/0044____.HTM]:\n\n{quote}\nwinbind separator\n The winbind separator option allows you to specify how NT domain names\n and user names are combined into unix user names when presented to\n users. By default, winbindd will use the traditional '\\' separator so\n that the unix user names look like DOMAIN\\username. In some cases this\n separator character may cause problems as the '\\' character has\n special meaning in unix shells. In that case you can use the winbind\n separator option to specify an alternative separator character. Good\n alternatives may be '/' (although that conflicts with the unix\n directory separator) or a '+ 'character. The '+' character appears to\n be the best choice for 100% compatibility with existing unix\n utilities, but may be an aesthetically bad choice depending on your\n taste.\n\n Default: winbind separator = \\\n\n Example: winbind separator = +\n{quote}","from":"reporter","subject":"dfs -chown does not like \"+\" plus sign in user name"},{"body":"This is big a problem if you use winbind along with hadoop security setup.","from":"developer"},{"body":"Patch 01:\r\n\"-\" sign has to remain at 1st place otherwise it will be compiled as range","from":"developer"},{"body":"Bulk update: moved all 3.2.0 non-blocker issues, please move back if it is a blocker.","from":"developer"},{"body":"Bulk update: moved all 3.3.0 non-blocker issues, please move back if it is a blocker.","from":"developer"},{"body":"Thanx [~boky01] for the patch.\r\n Well I think adding support for + shouldn't be a problem, but the pattern is there since quite long. Not sure whether they didn't add support for \"+\" deliberately or not.. \r\n Since it isn't going to specific to HDFS, would be good if we get some more opinions\r\n[~stevel@apache.org] [~aajisaka] any thoughts on this?","from":"developer"},{"body":"[~boky01] seems you allowed + for windows as well, I think windows doesn't support + in user name.\r\nCan you check once, I checked in Linux and it allows + in user name. if not, we can just allow for linux and this should be good to go then.","from":"developer"},{"body":"[~ayushtkn],\r\nWindows remained unchanged only Linux will allow + sign.","from":"developer"},{"body":"Can you give a check to the checkstyle complains\nTest failures seems unrelated.\n+1 once checkstyle issue is fixed\n","from":"developer"},{"body":"[~ayushtkn],\r\nThe checkstyle is not caused by my patch. The indentation was wrong even before my patch.\r\nI did not fix it because I did not want bigger patch than needed and indentation fixes decrease the readability in diff tools. But I am not sure what is the best practice here.","from":"developer"},{"body":"Should be fine then, we can keep the existing indentation to match the rest of the file","from":"developer"},{"body":"Committed to trunk\r\n\r\nThanx [~boky01]  for the fix and [~h0tbird]  for the report!!!","from":"developer"}],"created":"2013-08-08T14:55:18.000+0000","description":"I intend to set user and group:\n\n*User:* _MYCOMPANY+marc.villacorta_\n*Group:* hadoop\n\nwhere _'+'_ is what we use as a winbind separator.\n\nAnd this is what I get:\n{code:none}\nsudo -u hdfs hadoop fs -touchz /tmp/test.txt\nsudo -u hdfs hadoop fs -chown MYCOMPANY+marc.villacorta:hadoop /tmp/test.txt\n-chown: 'MYCOMPANY+marc.villacorta:hadoop' does not match expected pattern for [owner][:group].\nUsage: hadoop fs [generic options] -chown [-R] [OWNER][:[GROUP]] PATH...\n{code}\n\nI am using version: 2.0.0-cdh4.3.0\n\nQuote [source|http://h30097.www3.hp.com/docs/iass/OSIS_62/MAN/MAN8/0044____.HTM]:\n\n{quote}\nwinbind separator\n The winbind separator option allows you to specify how NT domain names\n and user names are combined into unix user names when presented to\n users. By default, winbindd will use the traditional '\\' separator so\n that the unix user names look like DOMAIN\\username. In some cases this\n separator character may cause problems as the '\\' character has\n special meaning in unix shells. In that case you can use the winbind\n separator option to specify an alternative separator character. Good\n alternatives may be '/' (although that conflicts with the unix\n directory separator) or a '+ 'character. The '+' character appears to\n be the best choice for 100% compatibility with existing unix\n utilities, but may be an aesthetically bad choice depending on your\n taste.\n\n Default: winbind separator = \\\n\n Example: winbind separator = +\n{quote}","issue_id":"12662634","key":"HADOOP-9851","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2020-06-17T08:34:38.000+0000","role":"fixed_distractor","summary":"dfs -chown does not like \"+\" plus sign in user name"} {"case_id":"12680362","cluster":"DISTRACTOR-HBASE-10015","comments":[{"body":"Sample 0.94 patch.","created":"2013-11-20T20:12:19.393+0000"},{"body":"The 3x improvement I see for tall tables. For wider tables the improvement is less pronounced (for 5 columns I see a 20% or so improvement).\n\nAlso a note as to why I think this should be correct: Even with the current synchronized, a next/peek/reseek that has already started, will finish with the old reader. That has not changed.\n\nNeed to look closer at races between close() and next/peek/reseek, as close could be called from another thread (I think), when the lease expired.","created":"2013-11-20T20:24:06.631+0000"},{"body":"Actually the main difference is between ScanWildcardColumnTracker and ExplicitColumnTracker. With ExplicitColumnTracker there are still many reseeks that outweigh the performance improvement seen here.","created":"2013-11-20T20:30:50.614+0000"},{"body":"It is interesting. Java synchronization w/o thread contention cost is close to zero. You would see the difference only when you run multiple threads accessing the same StoreScanner.","created":"2013-11-20T20:39:29.016+0000"},{"body":"That is true as far as actual thread synchronization goes.\n\nEvery synchronize still places a read and write memory fence, though, and I think that is the effect we're seeing. I was a bit surprised about the magnitude of this myself.\n\nI can try switching one of my cores off and measure this again, if the issues is memory fencing we should not see any improvement with one core only.\n","created":"2013-11-20T20:56:04.156+0000"},{"body":"I just ran mys tests and found no difference at all, but they were single-thread StoreScanner. ","created":"2013-11-20T21:20:17.478+0000"},{"body":"Nope. I see the same effect with just one core.\nWill on some other machines and different versions of the JDK.","created":"2013-11-20T21:20:21.630+0000"},{"body":"Close is fine, since it is only triggered via RegionScannerImpl, which is synchronized.\nThe overhead I measured *could* explain the numbers I've seen here:\nhttps://issues.apache.org/jira/browse/HBASE-9440?focusedCommentId=13767047&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-13767047\n\n(Compared to the raw scan speed in an HFile I found medium overhead per column and a lot of overhead per rows, and StoreScanner.peek()/next(), etc, are mostly called per row).","created":"2013-11-20T23:23:32.915+0000"},{"body":"bq. I just ran mys tests and found no difference at all, but they were single-thread StoreScanner. \n\nHmm... All my tests are single threaded. Different JVM versions?\nAre you testing with tall tables?","created":"2013-11-20T23:30:42.922+0000"},{"body":"The patch doesn't look wrong, but I'm surprised at the improvement. If not too onerous, perhaps you could attach a test case that reproduces it?","created":"2013-11-20T23:31:12.858+0000"},{"body":"It'll be hard to disentangle the test. Basically I am testing with a table with a single column family, and filtering all data at the server with a ValueFilter. Lemme have an extra look at the code.","created":"2013-11-20T23:33:46.895+0000"},{"body":"Could the test be isolated in a variant of HFilePerfEval? I've had luck isolating the BlockCache in a similar variant (HBASE-9806).","created":"2013-11-20T23:41:32.318+0000"},{"body":"{quote}\nDifferent JVM versions?\n{quote}\n1.6_56 Mac OSX 10.7.5. I concur with myself one more time: the cost of synchronized is very low when there is no thread contention.","created":"2013-11-21T00:30:50.367+0000"},{"body":"Strange... On my desktop machine (12 cores, JDK6) I also measure no difference.\nMy Laptop has 2 cores and OpenJDK7. I wonder whether I am seeing an OpenJDK bug.\n","created":"2013-11-21T00:42:16.324+0000"},{"body":"Confirmed on my laptop with OpenJDK, 2.5x improvement over 10 runs very low standard deviation.\nGoing to do some microbenchmarks.","created":"2013-11-21T01:06:03.701+0000"},{"body":"Wait... No, I do see the same improvement in the JDK6, 12 core box (I forgot to switch the server over).\nI'll attach a quick and dirty benchmark.","created":"2013-11-21T01:12:15.595+0000"},{"body":"Test code. This is some crap extracted from various tests, don't judge me on that code.. I will deny that I've ever written that.\nUse it with TestLoad . It will create and seed the table if needed.\nThen it runs the scan tests 5 times (after a prime round) and calculates mean and standard deviation.\n\nDid I mention that this is hack? :)\n\nFor the record, I see a 2-3x scan speed improvement on both OpenJDK7 on a old'ish 2 core machine as well as with Oracle JDK6 on a 12 core machine. In both cases the scan is from a single client only.\n","created":"2013-11-21T01:24:25.345+0000"},{"body":"bq. I concur with myself one more time: the cost of synchronized is very low when there is no thread contention.\n\nWell, you're wrong twice then :)\n\nJust try it... Call a synchronized method in a loop a few 100 million times. Then remove the synchronized. Make sure the method returns something, such as a reference to a member, so it is not optimized immediately.\n\nOn my test machines (JDK6 and JDK7) the latter is at least 40x faster on some machines it's 63x faster. All just a single thread.\n\nAs I said before, synchronized does more than exclusion.\n# it barres JVM from reordering instructions\n# it places memory fences (both read and write), which \\-depending on exact hw- disallows instruction reordering of the CPU\n# it may flush cache lines\n","created":"2013-11-21T02:26:24.869+0000"},{"body":"If somebody else could run the test I attached that'd be great. I ran against a local single node HBase on top of a single node HDFS cluster, with all data in the blockcache. Maybe a unittest against a mini cluster would be better.","created":"2013-11-21T02:32:29.125+0000"},{"body":"May be I am wrong (empty synchronized method call cost on my laptop is 25 ns) but my own tests on StoreScanner show 0 improvement. \n\nCode is simple:\n\ncreate region, populate with data (make sure data is in a cache) , then\n\n{code}\n LOG.info(\"Test store scanner\");\n Scan scan = new Scan();\n scan.setStartRow(region.getStartKey());\n scan.setStopRow(region.getEndKey());\n Store store = region.getStore(CF);\n StoreScanner scanner = new StoreScanner(store, store.getScanInfo(), scan, null);\n long start = System.currentTimeMillis();\n int total = 0;\n List result = new ArrayList();\n while(scanner.next(result)){\n total++; result.clear();\n }\n \n LOG.info(\"Test store scanner finished. Found \"+total +\" in \"+(System.currentTimeMillis() - start)+\"ms\");\n{code}\n\nThis test shows exact the same time for both: default StoreScanner and *unsynchronized* StoreScanner. The scan is not very fast: 1-1.5M rows per sec (rows are relatively small: 1 CF + 5 CQ, ~ 120 bytes )\n\n ","created":"2013-11-21T03:02:44.084+0000"},{"body":"For all microbenchmarks, please add the following command -line args:\n\n{code}\n-XX:+UseBiasedLocking -XX:BiasedLockingStartupDelay=0\n{code}\n\nOracle JVM does not enable biased locking (single thread lock optimization) first 4s of program execution.","created":"2013-11-21T03:52:08.313+0000"},{"body":"Can you try at the RegionScanner level, which maintains a heap of StoreScanners and calls peek() quite frequently. In my sampling sampling profiler I saw StoreScanner.peek() come us as some of the top method where the time is spent (and does nothing but doing a compare on a local member then calls peek on its heap).\n\nAlso 25ns are more than 100 cycles on modern HW, in line what I would expect from memory stalls caused by the fences.\n","created":"2013-11-21T04:09:04.676+0000"},{"body":"Will add those options. All the tests run at least 15s, though.\n\nOh, also rereading your comment above... I saw the most significant improvement per row (i.e. with 5 CQ you'd see less of an improvement). This very much looks like it is an interaction between KeyValueHeap and StoreScanner, where the cost seems to be mostly per row, rather than KeyValue.","created":"2013-11-21T04:13:09.249+0000"},{"body":"{quote}\nAlso 25ns are more than 100 cycles on modern HW, in line what I would expect from memory stalls caused by the fences.\n{quote}\nWith biased locking on, its ~ 2ns. I will try RegionScanner.","created":"2013-11-21T04:15:30.769+0000"},{"body":"Added a quick and dirty perf test. The test will fail at the end and failure message is the runtime (I find this the most convenient).\n\nI get:\n10 runs mean:2315.9 sigma:121.37417352962696\nand\n10 runs mean:4721.9 sigma:153.24911092727422\nwith the changes to StoreScanner reverted.\n","created":"2013-11-21T06:22:47.369+0000"},{"body":"One last note before I call it quits today.\nThe win appears to reduce the per row overhead by 50-60%. With 5 CQs this shrinks proportionally to 10-15% or so.\n","created":"2013-11-21T06:47:44.367+0000"},{"body":"If somebody could run the attached unit test that would be greatly appreciated.","created":"2013-11-21T17:33:45.656+0000"},{"body":"I tried to run the unit test but got:\n{code}\ntestScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance) Time elapsed: 0.007 sec <<< ERROR!\norg.apache.hadoop.ipc.RemoteException: java.io.IOException: File /user/tyu/hbase/hbase.version could only be replicated to 0 nodes, instead of 1\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getAdditionalBlock(FSNamesystem.java:1558)\n at org.apache.hadoop.hdfs.server.namenode.NameNode.addBlock(NameNode.java:696)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n{code}","created":"2013-11-21T18:45:57.334+0000"},{"body":"That would be related to your setup I think. Seems like it can't start the mini cluster.","created":"2013-11-21T18:54:14.834+0000"},{"body":"Without: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:2224.3 sigma:24.178709642989634\n\nWith: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:1373.3 sigma:14.920120642943875\n\nThis is on an EC2 c3.2xlarge, which uses 16 of 20 available hardware threads of a two socket board with Xeon E5-2680s, backed by SSDs. ","created":"2013-11-21T19:02:24.721+0000"},{"body":"[~tedyu@apache.org] that's probably HBASE-5711, do 'umask 022' first.","created":"2013-11-21T19:06:14.428+0000"},{"body":"@Andy:\nThanks for the hint.\nI ran 'umask 022' first but got the same result on Mac.","created":"2013-11-21T19:10:25.165+0000"},{"body":"200 iterations of TestStoreScanner, TestAtomicOperation and TestAcidGuarantees passed based on 0.94 patch.","created":"2013-11-21T19:46:04.188+0000"},{"body":"I'll make a trunk patch later today.","created":"2013-11-21T20:22:26.017+0000"},{"body":"I ran RegionScanner and see ~10% improvement on on a very narrow rows only (8 store files in a region). This is 0.94.6. By the way the performance difference between StoreScanner and RegionScanner is huge (almost 3x times).","created":"2013-11-21T21:21:41.376+0000"},{"body":"Thanks [~vrodionov], we'll look at RegionScanner next. :) My evil plans is to eventually get rid of all synchronization during scanning and provide exclusion by containment instead.\n\nThe overall improvement is real. Validated on various different machines. The effect of memory stalls is probably more pronounced in the running server.\n\nIn any case, I think we agree that this patch can't make things worse (provide it is correct, of course).\nI also tried with count\\(*) queries in Phoenix on tall tables (to rule out some anomalies with Filters). I see a 90% improvement there - again only on tall tables.\n","created":"2013-11-21T22:14:48.947+0000"},{"body":"Trunk patch.","created":"2013-11-21T22:15:36.986+0000"},{"body":"Some more results:\n10 runs mean:1851.0 sigma:127.71374240855992\nvs (without StoreScanner patch)\n10 runs mean:2795.8 sigma:57.52182194611015\n","created":"2013-11-21T22:31:48.557+0000"},{"body":"{code}\n0: jdbc:phoenix:localhost> select count(*) from \"tableTall\";\n+----------+\n| COUNT(1) |\n+----------+\n| 20000000 |\n+----------+\n1 row selected (3.946 seconds)\n{code}\nWith patch:\n{code}\n0: jdbc:phoenix:localhost> select count(*) from \"tableTall\";\n+----------+\n| COUNT(1) |\n+----------+\n| 20000000 |\n+----------+\n1 row selected (2.915 seconds)\n{code}\n\nPhoenix is using FAST_DIFF encoding and the ExplicitColumnTracker, so the relative gain is only 25% there (one CQ per row)","created":"2013-11-21T22:41:12.713+0000"},{"body":"The failures are all NPEs. Two of them could be related, the last is during DFS cluster shutdown.","created":"2013-11-22T01:27:47.078+0000"},{"body":"Got together with some of our hardware exports. We measured some of a perf counters and found that with the patch we see 50% less branch-misses, which is significant.","created":"2013-11-22T01:32:45.288+0000"},{"body":"Arrghh... No idea how the first patch could actually work *at all*. If you look closely you'll see that the updateReader condition was never unset. So on each an every call to next() it would force a reset of the scanner stack.\n\nHere's an update, which resets the condition atomically.","created":"2013-11-22T01:55:13.406+0000"},{"body":"And a trunk version. Note with v2 I see the same improvement.","created":"2013-11-22T01:58:04.093+0000"},{"body":"10 runs mean:2770.5 sigma:85.37827592543668, without\n10 runs mean:1791.2 sigma:50.81495842761264, with","created":"2013-11-22T02:02:32.084+0000"},{"body":"Still getting the NPE in trunk in TestHRegion with v2, though (0.94 is fine with v2)... Looking.\n","created":"2013-11-22T02:33:15.499+0000"},{"body":"The numbers are great. Patch looks good given the premise. Was going to suggest volatile instead of AtomicBoolean but looks like the getAndSet is fundamental so withdraw this nit. Is test failure because of this patch?\n\n","created":"2013-11-22T04:44:50.072+0000"},{"body":"Unfortunately I am no longer sure it actually works. I identified the problem:\nIn HStore.completeCompaction we call notifyChangedReadersObservers, which calls updateReaders on all StoreScanners. Before this patch, this method would block if there was a StoreScanner in the middle of a next/seek/reseek/etc. So before that patch one would guarantee that after notifyChangedReadersObservers returns it is safe to remove the compacted files. That is no longer true.\n\nI'll see if I can come up with something. It is a shame that we have to synchronize and get 10's of millions of branch misses per second, just so we can compact a few times a day.\n","created":"2013-11-22T05:02:00.215+0000"},{"body":"A possible solution is keeping updateReaders() all methods that actually call checkReseek synchronized.\nSo peek() could still go without synchronization. Will test the performance with that.","created":"2013-11-22T05:37:56.862+0000"},{"body":"That seems to be correct. updateReaders now waits until all operations that are effected by the reader changes to finish. peek() does not need to be synchronized as long as we do not change scanner stack from under its feet, which we avoid by deferring to the scanner thread to do that.\n\nThe unittest now yields this:\n10 runs mean:4449.2 sigma:24.395901295094635, without patch\n10 runs mean:2609.2 sigma:67.94085663280968, with patch\n\nSo, still a significant win, albeit to quite what it was before.\n","created":"2013-11-22T05:44:38.760+0000"},{"body":"New v3 patch for 0.94","created":"2013-11-22T05:45:18.368+0000"},{"body":"And trunk.\nThis patch should be correct for all scenarios.","created":"2013-11-22T05:45:54.080+0000"},{"body":"Can we have it so no changing of readers during a scan?","created":"2013-11-22T05:53:58.962+0000"},{"body":"Oops. Attached wrong trunk patch. This is right one. All good :)","created":"2013-11-22T05:57:03.134+0000"},{"body":"I think it would be hard to\n# push reader switching to the scanning thread\n# guarantee that the readers will eventually be switched to the compaction can proceed and finished\n\nWhat if a scanner isn't closed? Or nobody ever calls next() on it? Then we have to wait for the lease to expire before the compaction can finish.\n\nnext()/seek()/reseek() are not as critical as they are not called remotely as often as peek. peek() is also a very small method and (according to our perf expert here) small synchronized methods have the greatest potential to throw the branch prediction off.\n","created":"2013-11-22T06:03:28.591+0000"},{"body":"Phoenix with v3 (different machines than above, so do not compare absolute numbers).\nWithout patch:\n{code}\n0: jdbc:phoenix:localhost> select count(*) from \"myNew1\";\n+----------+\n| COUNT(1) |\n+----------+\n| 30000000 |\n+----------+\n1 row selected (18.417 seconds)\n{code}\n\nWith patch:\n{code}\n0: jdbc:phoenix:localhost> select count(*) from \"myNew1\";\n+----------+\n| COUNT(1) |\n+----------+\n| 30000000 |\n+----------+\n1 row selected (13.319 seconds)\n{code}\n","created":"2013-11-22T06:20:14.430+0000"},{"body":"bq. Then we have to wait for the lease to expire before the compaction can finish.\n\nI was thinking this might not be the end of the world.\n\nWhat if we closed and reopened the current scanner when readers are changed. It is a rare event.","created":"2013-11-22T06:23:09.820+0000"},{"body":"Actually now it does not even need the AtomicBoolean anymore.","created":"2013-11-22T06:26:38.045+0000"},{"body":"The closing and reopening is what needs the synchronization :)\nA thread might still be using the scanner. If we can make it so that we do not have to synchronize all methods in StoreScanner, and still do it safely I'd be happy... But frankly ATM I do not see how without other expensive synchronization.\n","created":"2013-11-22T06:30:31.973+0000"},{"body":"Removed AtomicBoolean","created":"2013-11-22T06:37:54.110+0000"},{"body":"Same for trunk.\n(This give a few % better performance)","created":"2013-11-22T06:38:29.292+0000"},{"body":"bq. The closing and reopening is what needs the synchronization \n\nI'm not helping. \n\nThis looks like an issue:\n\n+ if (shouldUpdateReaders) {\n+ shouldUpdateReaders = false;\n\nIt is happening outside a sync block and shouldUpdateReaders is not volatile.\n\n","created":"2013-11-22T06:58:50.768+0000"},{"body":"checkUpdateReaders is called from checkReseek, which is only called from synchronized methods. :)\n\nYou are helping. Looking at this from different angles is good.","created":"2013-11-22T07:02:25.345+0000"},{"body":"OK. Would suggest running hadoopqa a few times. Tests around changing readers are usually pretty good at finding issues if any.","created":"2013-11-22T07:17:11.778+0000"},{"body":"Yep. Found issues with the first two versions of this.","created":"2013-11-22T07:26:42.752+0000"},{"body":"Get another run in.","created":"2013-11-22T17:31:49.426+0000"},{"body":"Thanks Stack.","created":"2013-11-22T18:04:55.855+0000"},{"body":"Need to look at the find bugs thing","created":"2013-11-22T21:03:27.939+0000"},{"body":"+1 on commit.","created":"2013-11-22T21:40:00.218+0000"},{"body":"There's more stuff to do. I got a crash course on CPU design from Hussam Mousa, one of our performance experts.\nWe saw a lot of CPU frontend and backend stalls during scanning, so there is potential for a lot more improvements.\n\nI'm still looking for ways to remove all synchronization from StoreScanner and RegionScannerImpl.\n","created":"2013-11-23T00:45:45.597+0000"},{"body":"I'll look at the findbugs issue, do some more tests, and then commit.\n\nAn interesting metric we have to start to pay more attention is the scan cost per KV.\nFor example scanning through 1m rows with one CQ is *much* slower than scanning through 100k rows with 10 CQs, even though it touches the same number of KVs. This patch helps a bit to even that out.\n\nAs [~stack] and I said in the comments here, it should be possible to remove all synchronization form StoreScanner and RegionScannerImpl. It would require some refactoring.","created":"2013-11-23T00:50:12.049+0000"},{"body":"bq. It would require some refactoring.\n\nLets make a plan.","created":"2013-11-23T01:07:50.627+0000"},{"body":"The findbugs warning is a dud, it complains about StoreScanner.heap being locked 79% of the cases.","created":"2013-11-23T01:28:18.951+0000"},{"body":"{quote}\n so there is potential for a lot more improvements\n{quote}\n\nKV creation (new) is one of the serious bottlenecks, but I have no idea how to not create new instances on *next*. I have done some HBase internal hacks to get Maximum possible performance from scan. It is the multi-threaded application (in my case - 8HT threads) and scans on StoreFileScanner directly (data is cached 100%). The table was tall and narrow.\n\nStock HBase was able to reach 50M KV per sec\nStock with KV reuse (hack) - 90M KV per sec.\n\n\n","created":"2013-11-23T01:30:23.280+0000"},{"body":"Here's another idea:\n* we already lock the RegionScannerImpl.next/seek/reseek, etc.\n* what if we pass the RegionScannerImpl instance as a lock object to StoreScanner\n* in updateReaders() we'd then lock on that RegionScannerImpl instance, or we'd call a special synchronized method in RegionScannerImpl and have it update the readers.\n\nThe lock scope would be broader, but it should be correct, and it would allow us to remove all locking from StoreScanner.\nI'll experiment with that.\n","created":"2013-11-23T01:38:24.862+0000"},{"body":"[~vrodionov], yeah, need to work on that too. Is that with block encoding? When using block encoding we need to copy the underlying byte[]. When no block encoding is used making a new KV is just a few dozen bytes.\n","created":"2013-11-23T01:40:50.584+0000"},{"body":"bq. pass the RegionScannerImpl instance as a lock object to StoreScanner\n\nAlas, that works fine for scanning, but not for compactions/flushes (which also use StoreScanner, and can actually override the StoreScanner via a coprocessor hook).\n","created":"2013-11-23T05:47:39.902+0000"},{"body":"bq. but not for compactions/flushes ....\n\nTell us more? Am interested.","created":"2013-11-23T05:58:49.928+0000"},{"body":"I thought, since we lock all operations at the RegioScannerImpl level anyway, we could exclude a Store's updateReaders at that level too. That works fine if we have a RegionScanner, but in the case of compactions and flushes we don't one, so I have no way to protect the StoreScanner reading on behalf of a flush or a compaction.\n","created":"2013-11-23T06:19:55.192+0000"},{"body":"And it would be silly creating a regionscannerimpl to do a storescan. We need to do update readers differently. Compactions and flushes are rare and background tasks. They should defer to the ongoing scans.","created":"2013-11-23T06:45:13.403+0000"},{"body":"Yeah, I'm coming around to that. :)\nWe have to figure out how to do that without excessive synchronization. At the minimum we have to guarantee that nobody is currently using the readers in question before we retire them.\nAlternatively we wait for all running scanners to finish. How do we do that?\n","created":"2013-11-23T07:06:26.981+0000"},{"body":"I can come up with various schemes for that, but nothing that is cheaper than just synchronizing all the methods.","created":"2013-11-23T07:17:08.339+0000"},{"body":"+ Scans run free for N seconds or nanoseconds since checking this should be cheap(?) and then they go to a checkpoint where they look to see if they should reset. ChangedReaders blocks until checkpoint has been cleared.\n+ Scans check for closing being set on each op (I suppose this check of a volatile would be just as bad as a synchronization).\n+ We refcount outstanding scanners and only delete compacted files when refcount goes to zero. Not sure how we'd switch in flushes unless we prevent the flush happening while ongoing scan (eek).\n\n","created":"2013-11-23T07:37:36.715+0000"},{"body":"Something like this. \n\nOn order for the compaction to make progress we could assign scanners to epochs, where everytime we change the readers for a store we go to a new epoch. If all scanners for an epoch are either done or have switched to the new epoch, we can retire the readers of that epoch.\n\nIn any case, just commit this change and keep working on it? As Vladimir points out there are other issues with more serious performance implications.","created":"2013-11-23T17:49:27.390+0000"},{"body":"bq. On order for the compaction to make progress we could assign scanners to epochs, where everytime we change the readers for a store we go to a new epoch. If all scanners for an epoch are either done or have switched to the new epoch, we can retire the readers of that epoch.\n\nI haven't thought about this nearly as much as you have recently but that sounds promising.","created":"2013-11-23T18:35:43.363+0000"},{"body":"+1 on committing what is done already.\n\nepoch sounds like refcounting? Yeah, no hurry deleting the old stuff as long as it is done eventually. Accounting would be easier if we could move/rename files under the scanner as long is it does not disrupt (maybe I can try this).\n\nI like your idea of lockless scanning. Would be good to put it up as a goal even if hard to attain, if only to orientate which way progress lies.","created":"2013-11-23T18:42:27.792+0000"},{"body":"Epoch would be like reference counting per distinct sweet of readers. With just a reference count of scanners I'd be worried that compaction would never make any progress.","created":"2013-11-23T22:35:42.637+0000"},{"body":"Actually this is not quite right in 0.94, because it does not have HBASE-6499. (seek does not call checkReseek). I'll make that change as well, also checkUpdatedReaders can be folded into checkReseek for better readability.\n","created":"2013-11-24T04:08:13.071+0000"},{"body":"No longer sure that patch is good. The problem was that StoreScanner.peek() is synchronized.\n\nWith the patch peek() will use the old scanner stack, looking at MemstoreScanner and StoreFileScanner it *is* correct (both scanners just return the current KV and StoreFileScanner does not touch its reader during peek), but it is fragile, it also requires StoreScanner to hang on to all StoreFileScanners and the MemstoreScanner, until checkReseek is called or the scanner is closed. This in turn means that we keep references to the readers open, etc.\n\nI keep doing that... Filing jiras with patches and then finding that the fix is a bad idea.\nWell, at least I learned about modern CPU and that synchronized in tight loops is very expensive.\n\nI will keep thinking about this. Unscheduling for now.","created":"2013-11-24T05:02:18.253+0000"},{"body":"So you are thinking this cannot be completed until after we add delayed clean up of no-longer-used files? Only then can we safely remove synchronizations?\n\nHow many threads we talking anyways? It should be uncontended. The only thread is the current handler asking to return scan results -- this changes as different handlers come in on each bulk next invocation -- and then an incidental update readers request..and that is it?","created":"2013-11-24T18:20:58.518+0000"},{"body":"I think so (to your first point).\n\nThis is (almost) never contented, but StoreScanner.peek() is called *very* frequently (including the compares in KeyValueHeap) and the memory fences enforced by synchronized cause a slowdown.\n\nI did notice that during flushed and compactions we do *not* register any listeners for changed readers, so my earlier idea of just synchronizing on the RegionScannerImpl should work after all.\n","created":"2013-11-24T23:47:54.040+0000"},{"body":"Here's yet another sample patch (for 0.94). Adds a setter for a syncObject to KeyValueScanner. (if I wanted to change the StoreScanner constructor then we need to change the coprocessor region observer to pass this in as well, so I opted for a setter instead). When a StoreScanner is created by a RegionScanner it passes a reference to this as the syncObject.\nSo now all locking is - when needed - done through the RegionScannerImpl and I see the same performance improvement.\nAs said above StoreScanner for flushes and compactions are never invalidated anyway.\n\nWe have coarsened the lock from StoreScanner.next/peek/seek/etc to RegionScannerImpl.next/peek/seek/etc.\n\nNow, this is not really pretty. Is this worth the performance gain for tall table scans? Up to 2x for really tall tables with small KVs, proportionally less for wider tables and larger KVs.\n","created":"2013-11-25T03:32:06.781+0000"},{"body":"Was just looking at making a trunk patch.\nIn trunk StoreScanners for compaction *do* register a change observer? But not in 0.94?!\n\nSo - sigh - the patch I just made for 0.94 will not work in trunk. And maybe more importantly: Is there a lingering bug in 0.94?","created":"2013-11-25T04:02:41.334+0000"},{"body":"Continuing my monologue here... :)\n\nTo my surprise I get a 25-30% perf improvement when I just replace intrinsic locking (synchronized) with ReentrantLock - again for very tall tables only.\nsynchronized with biased locking should outperform the ReentrantLock in uncontended cases, but it does not (all my tests are against a real RegionServer, so the initial delay for BiasedLocking has not effect here). \n\nSomething is going on.\n\nThis is on JDK7 on old'ish 2 core machine. Will try on some other machines as well. If this bears out on other machines as well it would be safe and quick win.\n","created":"2013-11-25T05:53:56.641+0000"},{"body":"Parking the ReentrantLock patch for 0.94 here.","created":"2013-11-25T05:54:33.921+0000"},{"body":"Verified on different machine architecture with JDK6 (using the attached unittest).\n4400ms vs 3200ms (JDK7, 2 core machine)\n2600mx vs 2000ms (JDK6, 12 core machine)\n\nAgain, if somebody is still reading and could verify the latest patch with the earlier unittest on their hardware that would be greatly appreciated.\n","created":"2013-11-25T06:20:22.413+0000"},{"body":"On mac I see this w/ jvm1.8:\n\nFailed tests: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:3032.9 sigma:60.501983438561744\nFailed tests: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:2322.4 sigma:119.92430946226041\n\nLet me try another box. How can I be sure there is no regression for wide tables?","created":"2013-11-25T18:31:47.020+0000"},{"body":"JDK 7, 16 vcores, SSD, xfs\n\nWithout: 5 runs, Mean: 14.4136 Sigma: 0.17013712116995514\n\nWith the 0.94 -lock patch: 5 runs, Mean: 10.614 Sigma: 0.18904179432072687\n\nEdit: This is with the 'TestLoad' utility","created":"2013-11-25T18:33:10.429+0000"},{"body":"This is the same system with TestScanFilterPerformance:\n\nWithout: 10 runs mean:2248.7 sigma:36.59795076230362\n\nWith the 0.94 -lock patch: 10 runs mean:1860.3 sigma:29.397448868906977","created":"2013-11-25T18:39:31.544+0000"},{"body":"On Linux with jdk 1.7 :\n\nwith patch:\nFailed tests: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:2297.5 sigma:23.29484921608208\n\nwithout:\nFailed tests: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:3242.4 sigma:732.0478399667605","created":"2013-11-25T18:41:19.227+0000"},{"body":"So it seems this is universal.\nOur performance dude has actually confirmed that this is an expected outcome. I'll gather some more CPU metrics as I find time today.\n\n[~stack], these locks are taken in StoreScanner, which is \"row\" based (unlike StoreFileScanner and MemstoreScanner). So the frequent locks/unlocks there we'd mostly see per row. Wider rows won't be slower, but the speedup effect would be proportionally less. To be sure I'll validate with wider tables (maybe 20 CQs). I will also double check that biased locking is in effect.\n","created":"2013-11-25T19:29:52.693+0000"},{"body":"10 columns, 100 byte values, 1st and 4th selected.\n5 runs, Mean: 13.756 Sigma: 0.04210938137755054, with lock patch\n5 runs, Mean: 14.2028 Sigma: 0.08452313292821084, without\n\n10 columns, all selected\n5 runs, Mean: 8.577 Sigma: 0.09060463564299566\n5 runs, Mean: 9.9846 Sigma: 0.1023749969474969, without\n\nPer KV cost now is dominant. In no scenario have I observed this to be slower.\n","created":"2013-11-25T23:00:16.323+0000"},{"body":"+1","created":"2013-11-25T23:02:01.790+0000"},{"body":"I force enabled biased locking. No improvement with intrinsic locking; according to the docs it is enabled in JDK6+ anyway, was just making sure.","created":"2013-11-26T00:12:08.868+0000"},{"body":"If we do ref counting, can't we still finish the compaction, but not archive the files as long as there are scanners against them. We can do it by decoupling the store file archiving from compaction. At the worst case on RS crash, we might end up already compacted files which won't affect semantics. \nRegardless, +1 on -lock patch just for the gains. ","created":"2013-11-26T01:01:30.310+0000"},{"body":"Yeah, that's the idea. We'd delay archiving any HFile until no scanners are referring to it any more.","created":"2013-11-26T01:21:07.528+0000"},{"body":"oh, ok. I missed that somehow from the above discussion. Is there a jira yet? I can help with this.","created":"2013-11-26T01:53:39.387+0000"},{"body":"The trick will be to do all this without the need to synchronize anything in StoreScanner.\nMaybe there is a way to bring my idea of lock coarsening further: Above I suggest to just have updateReaders lock on the RegionScannerImpl, because that is locked anyway. The problem was that - at least in trunk - we also want to be notified during flushes and compactions and in that case we do not have a RegionScannerImpl. So idea is: Lock the StoreScanner itself from the compaction/flush code and pass itself as the lock object. For flushes we do it in a single loop, for compaction we could do that in chunks of 10000 or so.\n\nNow in this issue I already attached a bunch of different patches. Lemme commit the -lock patch here and close this. We can discuss further on a new jira.\n\nThanks [~apurtell], [~stack], and [~tedyu@apache.org] for running the perf unittests.\n","created":"2013-11-26T05:03:05.808+0000"},{"body":"Trunk patch.","created":"2013-11-26T18:17:04.539+0000"},{"body":"Will commit if it comes through clean.\nChecked again in the sampling profiler, StoreScanner.peek() moved from 1st place to 15th or so.","created":"2013-11-26T18:57:46.573+0000"},{"body":"+1 to 0.96 and trunk.","created":"2013-11-26T19:18:22.177+0000"},{"body":"Will check the various new warnings.","created":"2013-11-26T20:57:08.231+0000"},{"body":"The javac warnings are these:\n{code}\n[INFO] Compiling 150 source files to /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/trunk/hbase-common/target/classes\n[WARNING] /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/trunk/hbase-common/src/main/java/org/apache/hadoop/hbase/util/Bytes.java:[51,15] sun.misc.Unsafe is Sun proprietary API and may be removed in a future release\n[WARNING] /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/trunk/hbase-common/src/main/java/org/apache/hadoop/hbase/util/Bytes.java:[1110,19] sun.misc.Unsafe is Sun proprietary API and may be removed in a future release\n[WARNING] /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/trunk/hbase-common/src/main/java/org/apache/hadoop/hbase/util/Bytes.java:[1116,21] sun.misc.Unsafe is Sun proprietary API and may be removed in a future release\n[WARNING] /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/trunk/hbase-common/src/main/java/org/apache/hadoop/hbase/util/Bytes.java:[1121,28] sun.misc.Unsafe is Sun proprietary API and may be removed in a future release\n{code}\nThese are not new.\n\nSame with the Javadoc warnings (no new ones from this patch).\nSame for findbugs.\n\nGoing to commit.","created":"2013-11-26T21:01:55.816+0000"},{"body":"Committed to all branches.\n\nInterestingly now I find that for tall tables region.getCoprocessorHost().postScannerFilterRow takes the new top spot.\n(with a sampling profiler)","created":"2013-11-26T21:11:24.159+0000"},{"body":"Released in 0.96.1. Issue closed.","created":"2013-12-16T18:46:44.532+0000"}],"conversations":[{"body":"Did some more profiling (this time with a sampling profiler) and StoreScanner.peek() showed up a lot in the samples. At first that was surprising, but peek is synchronized, so it seems a lot of the sync'ing cost is eaten there.\nIt seems the only reason we have to synchronize all these methods is because a concurrent flush or compaction can change the scanner stack, other than that only a single thread should access a StoreScanner at any given time.\nSo replaced updateReaders() with some code that just indicates to the scanner that the readers should be updated and then make it the using thread's responsibility to do the work.\nThe perf improvement from this is staggering. I am seeing somewhere around 3x scan performance improvement across all scenarios.\n\nNow, the hard part is to reason about whether this is 100% correct. I ran TestAtomicOperation and TestAcidGuarantees a few times in a loop, all still pass.\n\nWill attach a sample patch.","from":"reporter","subject":"Replace intrinsic locking with explicit locks in StoreScanner"},{"body":"Sample 0.94 patch.","from":"developer"},{"body":"The 3x improvement I see for tall tables. For wider tables the improvement is less pronounced (for 5 columns I see a 20% or so improvement).\n\nAlso a note as to why I think this should be correct: Even with the current synchronized, a next/peek/reseek that has already started, will finish with the old reader. That has not changed.\n\nNeed to look closer at races between close() and next/peek/reseek, as close could be called from another thread (I think), when the lease expired.","from":"developer"},{"body":"Actually the main difference is between ScanWildcardColumnTracker and ExplicitColumnTracker. With ExplicitColumnTracker there are still many reseeks that outweigh the performance improvement seen here.","from":"developer"},{"body":"It is interesting. Java synchronization w/o thread contention cost is close to zero. You would see the difference only when you run multiple threads accessing the same StoreScanner.","from":"developer"},{"body":"That is true as far as actual thread synchronization goes.\n\nEvery synchronize still places a read and write memory fence, though, and I think that is the effect we're seeing. I was a bit surprised about the magnitude of this myself.\n\nI can try switching one of my cores off and measure this again, if the issues is memory fencing we should not see any improvement with one core only.\n","from":"developer"},{"body":"I just ran mys tests and found no difference at all, but they were single-thread StoreScanner. ","from":"developer"},{"body":"Nope. I see the same effect with just one core.\nWill on some other machines and different versions of the JDK.","from":"developer"},{"body":"Close is fine, since it is only triggered via RegionScannerImpl, which is synchronized.\nThe overhead I measured *could* explain the numbers I've seen here:\nhttps://issues.apache.org/jira/browse/HBASE-9440?focusedCommentId=13767047&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-13767047\n\n(Compared to the raw scan speed in an HFile I found medium overhead per column and a lot of overhead per rows, and StoreScanner.peek()/next(), etc, are mostly called per row).","from":"developer"},{"body":"bq. I just ran mys tests and found no difference at all, but they were single-thread StoreScanner. \n\nHmm... All my tests are single threaded. Different JVM versions?\nAre you testing with tall tables?","from":"developer"},{"body":"The patch doesn't look wrong, but I'm surprised at the improvement. If not too onerous, perhaps you could attach a test case that reproduces it?","from":"developer"},{"body":"It'll be hard to disentangle the test. Basically I am testing with a table with a single column family, and filtering all data at the server with a ValueFilter. Lemme have an extra look at the code.","from":"developer"},{"body":"Could the test be isolated in a variant of HFilePerfEval? I've had luck isolating the BlockCache in a similar variant (HBASE-9806).","from":"developer"},{"body":"{quote}\nDifferent JVM versions?\n{quote}\n1.6_56 Mac OSX 10.7.5. I concur with myself one more time: the cost of synchronized is very low when there is no thread contention.","from":"developer"},{"body":"Strange... On my desktop machine (12 cores, JDK6) I also measure no difference.\nMy Laptop has 2 cores and OpenJDK7. I wonder whether I am seeing an OpenJDK bug.\n","from":"developer"},{"body":"Confirmed on my laptop with OpenJDK, 2.5x improvement over 10 runs very low standard deviation.\nGoing to do some microbenchmarks.","from":"developer"},{"body":"Wait... No, I do see the same improvement in the JDK6, 12 core box (I forgot to switch the server over).\nI'll attach a quick and dirty benchmark.","from":"developer"},{"body":"Test code. This is some crap extracted from various tests, don't judge me on that code.. I will deny that I've ever written that.\nUse it with TestLoad . It will create and seed the table if needed.\nThen it runs the scan tests 5 times (after a prime round) and calculates mean and standard deviation.\n\nDid I mention that this is hack? :)\n\nFor the record, I see a 2-3x scan speed improvement on both OpenJDK7 on a old'ish 2 core machine as well as with Oracle JDK6 on a 12 core machine. In both cases the scan is from a single client only.\n","from":"developer"},{"body":"bq. I concur with myself one more time: the cost of synchronized is very low when there is no thread contention.\n\nWell, you're wrong twice then :)\n\nJust try it... Call a synchronized method in a loop a few 100 million times. Then remove the synchronized. Make sure the method returns something, such as a reference to a member, so it is not optimized immediately.\n\nOn my test machines (JDK6 and JDK7) the latter is at least 40x faster on some machines it's 63x faster. All just a single thread.\n\nAs I said before, synchronized does more than exclusion.\n# it barres JVM from reordering instructions\n# it places memory fences (both read and write), which \\-depending on exact hw- disallows instruction reordering of the CPU\n# it may flush cache lines\n","from":"developer"},{"body":"If somebody else could run the test I attached that'd be great. I ran against a local single node HBase on top of a single node HDFS cluster, with all data in the blockcache. Maybe a unittest against a mini cluster would be better.","from":"developer"},{"body":"May be I am wrong (empty synchronized method call cost on my laptop is 25 ns) but my own tests on StoreScanner show 0 improvement. \n\nCode is simple:\n\ncreate region, populate with data (make sure data is in a cache) , then\n\n{code}\n LOG.info(\"Test store scanner\");\n Scan scan = new Scan();\n scan.setStartRow(region.getStartKey());\n scan.setStopRow(region.getEndKey());\n Store store = region.getStore(CF);\n StoreScanner scanner = new StoreScanner(store, store.getScanInfo(), scan, null);\n long start = System.currentTimeMillis();\n int total = 0;\n List result = new ArrayList();\n while(scanner.next(result)){\n total++; result.clear();\n }\n \n LOG.info(\"Test store scanner finished. Found \"+total +\" in \"+(System.currentTimeMillis() - start)+\"ms\");\n{code}\n\nThis test shows exact the same time for both: default StoreScanner and *unsynchronized* StoreScanner. The scan is not very fast: 1-1.5M rows per sec (rows are relatively small: 1 CF + 5 CQ, ~ 120 bytes )\n\n ","from":"developer"},{"body":"For all microbenchmarks, please add the following command -line args:\n\n{code}\n-XX:+UseBiasedLocking -XX:BiasedLockingStartupDelay=0\n{code}\n\nOracle JVM does not enable biased locking (single thread lock optimization) first 4s of program execution.","from":"developer"},{"body":"Can you try at the RegionScanner level, which maintains a heap of StoreScanners and calls peek() quite frequently. In my sampling sampling profiler I saw StoreScanner.peek() come us as some of the top method where the time is spent (and does nothing but doing a compare on a local member then calls peek on its heap).\n\nAlso 25ns are more than 100 cycles on modern HW, in line what I would expect from memory stalls caused by the fences.\n","from":"developer"},{"body":"Will add those options. All the tests run at least 15s, though.\n\nOh, also rereading your comment above... I saw the most significant improvement per row (i.e. with 5 CQ you'd see less of an improvement). This very much looks like it is an interaction between KeyValueHeap and StoreScanner, where the cost seems to be mostly per row, rather than KeyValue.","from":"developer"},{"body":"{quote}\nAlso 25ns are more than 100 cycles on modern HW, in line what I would expect from memory stalls caused by the fences.\n{quote}\nWith biased locking on, its ~ 2ns. I will try RegionScanner.","from":"developer"},{"body":"Added a quick and dirty perf test. The test will fail at the end and failure message is the runtime (I find this the most convenient).\n\nI get:\n10 runs mean:2315.9 sigma:121.37417352962696\nand\n10 runs mean:4721.9 sigma:153.24911092727422\nwith the changes to StoreScanner reverted.\n","from":"developer"},{"body":"One last note before I call it quits today.\nThe win appears to reduce the per row overhead by 50-60%. With 5 CQs this shrinks proportionally to 10-15% or so.\n","from":"developer"},{"body":"If somebody could run the attached unit test that would be greatly appreciated.","from":"developer"},{"body":"I tried to run the unit test but got:\n{code}\ntestScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance) Time elapsed: 0.007 sec <<< ERROR!\norg.apache.hadoop.ipc.RemoteException: java.io.IOException: File /user/tyu/hbase/hbase.version could only be replicated to 0 nodes, instead of 1\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getAdditionalBlock(FSNamesystem.java:1558)\n at org.apache.hadoop.hdfs.server.namenode.NameNode.addBlock(NameNode.java:696)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n{code}","from":"developer"},{"body":"That would be related to your setup I think. Seems like it can't start the mini cluster.","from":"developer"},{"body":"Without: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:2224.3 sigma:24.178709642989634\n\nWith: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:1373.3 sigma:14.920120642943875\n\nThis is on an EC2 c3.2xlarge, which uses 16 of 20 available hardware threads of a two socket board with Xeon E5-2680s, backed by SSDs. ","from":"developer"},{"body":"[~tedyu@apache.org] that's probably HBASE-5711, do 'umask 022' first.","from":"developer"},{"body":"@Andy:\nThanks for the hint.\nI ran 'umask 022' first but got the same result on Mac.","from":"developer"},{"body":"200 iterations of TestStoreScanner, TestAtomicOperation and TestAcidGuarantees passed based on 0.94 patch.","from":"developer"},{"body":"I'll make a trunk patch later today.","from":"developer"},{"body":"I ran RegionScanner and see ~10% improvement on on a very narrow rows only (8 store files in a region). This is 0.94.6. By the way the performance difference between StoreScanner and RegionScanner is huge (almost 3x times).","from":"developer"},{"body":"Thanks [~vrodionov], we'll look at RegionScanner next. :) My evil plans is to eventually get rid of all synchronization during scanning and provide exclusion by containment instead.\n\nThe overall improvement is real. Validated on various different machines. The effect of memory stalls is probably more pronounced in the running server.\n\nIn any case, I think we agree that this patch can't make things worse (provide it is correct, of course).\nI also tried with count\\(*) queries in Phoenix on tall tables (to rule out some anomalies with Filters). I see a 90% improvement there - again only on tall tables.\n","from":"developer"},{"body":"Trunk patch.","from":"developer"},{"body":"Some more results:\n10 runs mean:1851.0 sigma:127.71374240855992\nvs (without StoreScanner patch)\n10 runs mean:2795.8 sigma:57.52182194611015\n","from":"developer"},{"body":"{code}\n0: jdbc:phoenix:localhost> select count(*) from \"tableTall\";\n+----------+\n| COUNT(1) |\n+----------+\n| 20000000 |\n+----------+\n1 row selected (3.946 seconds)\n{code}\nWith patch:\n{code}\n0: jdbc:phoenix:localhost> select count(*) from \"tableTall\";\n+----------+\n| COUNT(1) |\n+----------+\n| 20000000 |\n+----------+\n1 row selected (2.915 seconds)\n{code}\n\nPhoenix is using FAST_DIFF encoding and the ExplicitColumnTracker, so the relative gain is only 25% there (one CQ per row)","from":"developer"},{"body":"The failures are all NPEs. Two of them could be related, the last is during DFS cluster shutdown.","from":"developer"},{"body":"Got together with some of our hardware exports. We measured some of a perf counters and found that with the patch we see 50% less branch-misses, which is significant.","from":"developer"},{"body":"Arrghh... No idea how the first patch could actually work *at all*. If you look closely you'll see that the updateReader condition was never unset. So on each an every call to next() it would force a reset of the scanner stack.\n\nHere's an update, which resets the condition atomically.","from":"developer"},{"body":"And a trunk version. Note with v2 I see the same improvement.","from":"developer"},{"body":"10 runs mean:2770.5 sigma:85.37827592543668, without\n10 runs mean:1791.2 sigma:50.81495842761264, with","from":"developer"},{"body":"Still getting the NPE in trunk in TestHRegion with v2, though (0.94 is fine with v2)... Looking.\n","from":"developer"},{"body":"The numbers are great. Patch looks good given the premise. Was going to suggest volatile instead of AtomicBoolean but looks like the getAndSet is fundamental so withdraw this nit. Is test failure because of this patch?\n\n","from":"developer"},{"body":"Unfortunately I am no longer sure it actually works. I identified the problem:\nIn HStore.completeCompaction we call notifyChangedReadersObservers, which calls updateReaders on all StoreScanners. Before this patch, this method would block if there was a StoreScanner in the middle of a next/seek/reseek/etc. So before that patch one would guarantee that after notifyChangedReadersObservers returns it is safe to remove the compacted files. That is no longer true.\n\nI'll see if I can come up with something. It is a shame that we have to synchronize and get 10's of millions of branch misses per second, just so we can compact a few times a day.\n","from":"developer"},{"body":"A possible solution is keeping updateReaders() all methods that actually call checkReseek synchronized.\nSo peek() could still go without synchronization. Will test the performance with that.","from":"developer"},{"body":"That seems to be correct. updateReaders now waits until all operations that are effected by the reader changes to finish. peek() does not need to be synchronized as long as we do not change scanner stack from under its feet, which we avoid by deferring to the scanner thread to do that.\n\nThe unittest now yields this:\n10 runs mean:4449.2 sigma:24.395901295094635, without patch\n10 runs mean:2609.2 sigma:67.94085663280968, with patch\n\nSo, still a significant win, albeit to quite what it was before.\n","from":"developer"},{"body":"New v3 patch for 0.94","from":"developer"},{"body":"And trunk.\nThis patch should be correct for all scenarios.","from":"developer"},{"body":"Can we have it so no changing of readers during a scan?","from":"developer"},{"body":"Oops. Attached wrong trunk patch. This is right one. All good :)","from":"developer"},{"body":"I think it would be hard to\n# push reader switching to the scanning thread\n# guarantee that the readers will eventually be switched to the compaction can proceed and finished\n\nWhat if a scanner isn't closed? Or nobody ever calls next() on it? Then we have to wait for the lease to expire before the compaction can finish.\n\nnext()/seek()/reseek() are not as critical as they are not called remotely as often as peek. peek() is also a very small method and (according to our perf expert here) small synchronized methods have the greatest potential to throw the branch prediction off.\n","from":"developer"},{"body":"Phoenix with v3 (different machines than above, so do not compare absolute numbers).\nWithout patch:\n{code}\n0: jdbc:phoenix:localhost> select count(*) from \"myNew1\";\n+----------+\n| COUNT(1) |\n+----------+\n| 30000000 |\n+----------+\n1 row selected (18.417 seconds)\n{code}\n\nWith patch:\n{code}\n0: jdbc:phoenix:localhost> select count(*) from \"myNew1\";\n+----------+\n| COUNT(1) |\n+----------+\n| 30000000 |\n+----------+\n1 row selected (13.319 seconds)\n{code}\n","from":"developer"},{"body":"bq. Then we have to wait for the lease to expire before the compaction can finish.\n\nI was thinking this might not be the end of the world.\n\nWhat if we closed and reopened the current scanner when readers are changed. It is a rare event.","from":"developer"},{"body":"Actually now it does not even need the AtomicBoolean anymore.","from":"developer"},{"body":"The closing and reopening is what needs the synchronization :)\nA thread might still be using the scanner. If we can make it so that we do not have to synchronize all methods in StoreScanner, and still do it safely I'd be happy... But frankly ATM I do not see how without other expensive synchronization.\n","from":"developer"},{"body":"Removed AtomicBoolean","from":"developer"},{"body":"Same for trunk.\n(This give a few % better performance)","from":"developer"},{"body":"bq. The closing and reopening is what needs the synchronization \n\nI'm not helping. \n\nThis looks like an issue:\n\n+ if (shouldUpdateReaders) {\n+ shouldUpdateReaders = false;\n\nIt is happening outside a sync block and shouldUpdateReaders is not volatile.\n\n","from":"developer"},{"body":"checkUpdateReaders is called from checkReseek, which is only called from synchronized methods. :)\n\nYou are helping. Looking at this from different angles is good.","from":"developer"},{"body":"OK. Would suggest running hadoopqa a few times. Tests around changing readers are usually pretty good at finding issues if any.","from":"developer"},{"body":"Yep. Found issues with the first two versions of this.","from":"developer"},{"body":"Get another run in.","from":"developer"},{"body":"Thanks Stack.","from":"developer"},{"body":"Need to look at the find bugs thing","from":"developer"},{"body":"+1 on commit.","from":"developer"},{"body":"There's more stuff to do. I got a crash course on CPU design from Hussam Mousa, one of our performance experts.\nWe saw a lot of CPU frontend and backend stalls during scanning, so there is potential for a lot more improvements.\n\nI'm still looking for ways to remove all synchronization from StoreScanner and RegionScannerImpl.\n","from":"developer"},{"body":"I'll look at the findbugs issue, do some more tests, and then commit.\n\nAn interesting metric we have to start to pay more attention is the scan cost per KV.\nFor example scanning through 1m rows with one CQ is *much* slower than scanning through 100k rows with 10 CQs, even though it touches the same number of KVs. This patch helps a bit to even that out.\n\nAs [~stack] and I said in the comments here, it should be possible to remove all synchronization form StoreScanner and RegionScannerImpl. It would require some refactoring.","from":"developer"},{"body":"bq. It would require some refactoring.\n\nLets make a plan.","from":"developer"},{"body":"The findbugs warning is a dud, it complains about StoreScanner.heap being locked 79% of the cases.","from":"developer"},{"body":"{quote}\n so there is potential for a lot more improvements\n{quote}\n\nKV creation (new) is one of the serious bottlenecks, but I have no idea how to not create new instances on *next*. I have done some HBase internal hacks to get Maximum possible performance from scan. It is the multi-threaded application (in my case - 8HT threads) and scans on StoreFileScanner directly (data is cached 100%). The table was tall and narrow.\n\nStock HBase was able to reach 50M KV per sec\nStock with KV reuse (hack) - 90M KV per sec.\n\n\n","from":"developer"},{"body":"Here's another idea:\n* we already lock the RegionScannerImpl.next/seek/reseek, etc.\n* what if we pass the RegionScannerImpl instance as a lock object to StoreScanner\n* in updateReaders() we'd then lock on that RegionScannerImpl instance, or we'd call a special synchronized method in RegionScannerImpl and have it update the readers.\n\nThe lock scope would be broader, but it should be correct, and it would allow us to remove all locking from StoreScanner.\nI'll experiment with that.\n","from":"developer"},{"body":"[~vrodionov], yeah, need to work on that too. Is that with block encoding? When using block encoding we need to copy the underlying byte[]. When no block encoding is used making a new KV is just a few dozen bytes.\n","from":"developer"},{"body":"bq. pass the RegionScannerImpl instance as a lock object to StoreScanner\n\nAlas, that works fine for scanning, but not for compactions/flushes (which also use StoreScanner, and can actually override the StoreScanner via a coprocessor hook).\n","from":"developer"},{"body":"bq. but not for compactions/flushes ....\n\nTell us more? Am interested.","from":"developer"},{"body":"I thought, since we lock all operations at the RegioScannerImpl level anyway, we could exclude a Store's updateReaders at that level too. That works fine if we have a RegionScanner, but in the case of compactions and flushes we don't one, so I have no way to protect the StoreScanner reading on behalf of a flush or a compaction.\n","from":"developer"},{"body":"And it would be silly creating a regionscannerimpl to do a storescan. We need to do update readers differently. Compactions and flushes are rare and background tasks. They should defer to the ongoing scans.","from":"developer"},{"body":"Yeah, I'm coming around to that. :)\nWe have to figure out how to do that without excessive synchronization. At the minimum we have to guarantee that nobody is currently using the readers in question before we retire them.\nAlternatively we wait for all running scanners to finish. How do we do that?\n","from":"developer"},{"body":"I can come up with various schemes for that, but nothing that is cheaper than just synchronizing all the methods.","from":"developer"},{"body":"+ Scans run free for N seconds or nanoseconds since checking this should be cheap(?) and then they go to a checkpoint where they look to see if they should reset. ChangedReaders blocks until checkpoint has been cleared.\n+ Scans check for closing being set on each op (I suppose this check of a volatile would be just as bad as a synchronization).\n+ We refcount outstanding scanners and only delete compacted files when refcount goes to zero. Not sure how we'd switch in flushes unless we prevent the flush happening while ongoing scan (eek).\n\n","from":"developer"},{"body":"Something like this. \n\nOn order for the compaction to make progress we could assign scanners to epochs, where everytime we change the readers for a store we go to a new epoch. If all scanners for an epoch are either done or have switched to the new epoch, we can retire the readers of that epoch.\n\nIn any case, just commit this change and keep working on it? As Vladimir points out there are other issues with more serious performance implications.","from":"developer"},{"body":"bq. On order for the compaction to make progress we could assign scanners to epochs, where everytime we change the readers for a store we go to a new epoch. If all scanners for an epoch are either done or have switched to the new epoch, we can retire the readers of that epoch.\n\nI haven't thought about this nearly as much as you have recently but that sounds promising.","from":"developer"},{"body":"+1 on committing what is done already.\n\nepoch sounds like refcounting? Yeah, no hurry deleting the old stuff as long as it is done eventually. Accounting would be easier if we could move/rename files under the scanner as long is it does not disrupt (maybe I can try this).\n\nI like your idea of lockless scanning. Would be good to put it up as a goal even if hard to attain, if only to orientate which way progress lies.","from":"developer"},{"body":"Epoch would be like reference counting per distinct sweet of readers. With just a reference count of scanners I'd be worried that compaction would never make any progress.","from":"developer"},{"body":"Actually this is not quite right in 0.94, because it does not have HBASE-6499. (seek does not call checkReseek). I'll make that change as well, also checkUpdatedReaders can be folded into checkReseek for better readability.\n","from":"developer"},{"body":"No longer sure that patch is good. The problem was that StoreScanner.peek() is synchronized.\n\nWith the patch peek() will use the old scanner stack, looking at MemstoreScanner and StoreFileScanner it *is* correct (both scanners just return the current KV and StoreFileScanner does not touch its reader during peek), but it is fragile, it also requires StoreScanner to hang on to all StoreFileScanners and the MemstoreScanner, until checkReseek is called or the scanner is closed. This in turn means that we keep references to the readers open, etc.\n\nI keep doing that... Filing jiras with patches and then finding that the fix is a bad idea.\nWell, at least I learned about modern CPU and that synchronized in tight loops is very expensive.\n\nI will keep thinking about this. Unscheduling for now.","from":"developer"},{"body":"So you are thinking this cannot be completed until after we add delayed clean up of no-longer-used files? Only then can we safely remove synchronizations?\n\nHow many threads we talking anyways? It should be uncontended. The only thread is the current handler asking to return scan results -- this changes as different handlers come in on each bulk next invocation -- and then an incidental update readers request..and that is it?","from":"developer"},{"body":"I think so (to your first point).\n\nThis is (almost) never contented, but StoreScanner.peek() is called *very* frequently (including the compares in KeyValueHeap) and the memory fences enforced by synchronized cause a slowdown.\n\nI did notice that during flushed and compactions we do *not* register any listeners for changed readers, so my earlier idea of just synchronizing on the RegionScannerImpl should work after all.\n","from":"developer"},{"body":"Here's yet another sample patch (for 0.94). Adds a setter for a syncObject to KeyValueScanner. (if I wanted to change the StoreScanner constructor then we need to change the coprocessor region observer to pass this in as well, so I opted for a setter instead). When a StoreScanner is created by a RegionScanner it passes a reference to this as the syncObject.\nSo now all locking is - when needed - done through the RegionScannerImpl and I see the same performance improvement.\nAs said above StoreScanner for flushes and compactions are never invalidated anyway.\n\nWe have coarsened the lock from StoreScanner.next/peek/seek/etc to RegionScannerImpl.next/peek/seek/etc.\n\nNow, this is not really pretty. Is this worth the performance gain for tall table scans? Up to 2x for really tall tables with small KVs, proportionally less for wider tables and larger KVs.\n","from":"developer"},{"body":"Was just looking at making a trunk patch.\nIn trunk StoreScanners for compaction *do* register a change observer? But not in 0.94?!\n\nSo - sigh - the patch I just made for 0.94 will not work in trunk. And maybe more importantly: Is there a lingering bug in 0.94?","from":"developer"},{"body":"Continuing my monologue here... :)\n\nTo my surprise I get a 25-30% perf improvement when I just replace intrinsic locking (synchronized) with ReentrantLock - again for very tall tables only.\nsynchronized with biased locking should outperform the ReentrantLock in uncontended cases, but it does not (all my tests are against a real RegionServer, so the initial delay for BiasedLocking has not effect here). \n\nSomething is going on.\n\nThis is on JDK7 on old'ish 2 core machine. Will try on some other machines as well. If this bears out on other machines as well it would be safe and quick win.\n","from":"developer"},{"body":"Parking the ReentrantLock patch for 0.94 here.","from":"developer"},{"body":"Verified on different machine architecture with JDK6 (using the attached unittest).\n4400ms vs 3200ms (JDK7, 2 core machine)\n2600mx vs 2000ms (JDK6, 12 core machine)\n\nAgain, if somebody is still reading and could verify the latest patch with the earlier unittest on their hardware that would be greatly appreciated.\n","from":"developer"},{"body":"On mac I see this w/ jvm1.8:\n\nFailed tests: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:3032.9 sigma:60.501983438561744\nFailed tests: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:2322.4 sigma:119.92430946226041\n\nLet me try another box. How can I be sure there is no regression for wide tables?","from":"developer"},{"body":"JDK 7, 16 vcores, SSD, xfs\n\nWithout: 5 runs, Mean: 14.4136 Sigma: 0.17013712116995514\n\nWith the 0.94 -lock patch: 5 runs, Mean: 10.614 Sigma: 0.18904179432072687\n\nEdit: This is with the 'TestLoad' utility","from":"developer"},{"body":"This is the same system with TestScanFilterPerformance:\n\nWithout: 10 runs mean:2248.7 sigma:36.59795076230362\n\nWith the 0.94 -lock patch: 10 runs mean:1860.3 sigma:29.397448868906977","from":"developer"},{"body":"On Linux with jdk 1.7 :\n\nwith patch:\nFailed tests: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:2297.5 sigma:23.29484921608208\n\nwithout:\nFailed tests: testScanFilterPerformance(org.apache.hadoop.hbase.regionserver.TestScanFilterPerformance): 10 runs mean:3242.4 sigma:732.0478399667605","from":"developer"},{"body":"So it seems this is universal.\nOur performance dude has actually confirmed that this is an expected outcome. I'll gather some more CPU metrics as I find time today.\n\n[~stack], these locks are taken in StoreScanner, which is \"row\" based (unlike StoreFileScanner and MemstoreScanner). So the frequent locks/unlocks there we'd mostly see per row. Wider rows won't be slower, but the speedup effect would be proportionally less. To be sure I'll validate with wider tables (maybe 20 CQs). I will also double check that biased locking is in effect.\n","from":"developer"},{"body":"10 columns, 100 byte values, 1st and 4th selected.\n5 runs, Mean: 13.756 Sigma: 0.04210938137755054, with lock patch\n5 runs, Mean: 14.2028 Sigma: 0.08452313292821084, without\n\n10 columns, all selected\n5 runs, Mean: 8.577 Sigma: 0.09060463564299566\n5 runs, Mean: 9.9846 Sigma: 0.1023749969474969, without\n\nPer KV cost now is dominant. In no scenario have I observed this to be slower.\n","from":"developer"},{"body":"+1","from":"developer"},{"body":"I force enabled biased locking. No improvement with intrinsic locking; according to the docs it is enabled in JDK6+ anyway, was just making sure.","from":"developer"},{"body":"If we do ref counting, can't we still finish the compaction, but not archive the files as long as there are scanners against them. We can do it by decoupling the store file archiving from compaction. At the worst case on RS crash, we might end up already compacted files which won't affect semantics. \nRegardless, +1 on -lock patch just for the gains. ","from":"developer"},{"body":"Yeah, that's the idea. We'd delay archiving any HFile until no scanners are referring to it any more.","from":"developer"},{"body":"oh, ok. I missed that somehow from the above discussion. Is there a jira yet? I can help with this.","from":"developer"},{"body":"The trick will be to do all this without the need to synchronize anything in StoreScanner.\nMaybe there is a way to bring my idea of lock coarsening further: Above I suggest to just have updateReaders lock on the RegionScannerImpl, because that is locked anyway. The problem was that - at least in trunk - we also want to be notified during flushes and compactions and in that case we do not have a RegionScannerImpl. So idea is: Lock the StoreScanner itself from the compaction/flush code and pass itself as the lock object. For flushes we do it in a single loop, for compaction we could do that in chunks of 10000 or so.\n\nNow in this issue I already attached a bunch of different patches. Lemme commit the -lock patch here and close this. We can discuss further on a new jira.\n\nThanks [~apurtell], [~stack], and [~tedyu@apache.org] for running the perf unittests.\n","from":"developer"},{"body":"Trunk patch.","from":"developer"},{"body":"Will commit if it comes through clean.\nChecked again in the sampling profiler, StoreScanner.peek() moved from 1st place to 15th or so.","from":"developer"},{"body":"+1 to 0.96 and trunk.","from":"developer"},{"body":"Will check the various new warnings.","from":"developer"},{"body":"The javac warnings are these:\n{code}\n[INFO] Compiling 150 source files to /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/trunk/hbase-common/target/classes\n[WARNING] /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/trunk/hbase-common/src/main/java/org/apache/hadoop/hbase/util/Bytes.java:[51,15] sun.misc.Unsafe is Sun proprietary API and may be removed in a future release\n[WARNING] /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/trunk/hbase-common/src/main/java/org/apache/hadoop/hbase/util/Bytes.java:[1110,19] sun.misc.Unsafe is Sun proprietary API and may be removed in a future release\n[WARNING] /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/trunk/hbase-common/src/main/java/org/apache/hadoop/hbase/util/Bytes.java:[1116,21] sun.misc.Unsafe is Sun proprietary API and may be removed in a future release\n[WARNING] /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/trunk/hbase-common/src/main/java/org/apache/hadoop/hbase/util/Bytes.java:[1121,28] sun.misc.Unsafe is Sun proprietary API and may be removed in a future release\n{code}\nThese are not new.\n\nSame with the Javadoc warnings (no new ones from this patch).\nSame for findbugs.\n\nGoing to commit.","from":"developer"},{"body":"Committed to all branches.\n\nInterestingly now I find that for tall tables region.getCoprocessorHost().postScannerFilterRow takes the new top spot.\n(with a sampling profiler)","from":"developer"},{"body":"Released in 0.96.1. Issue closed.","from":"developer"}],"created":"2013-11-20T20:10:58.000+0000","description":"Did some more profiling (this time with a sampling profiler) and StoreScanner.peek() showed up a lot in the samples. At first that was surprising, but peek is synchronized, so it seems a lot of the sync'ing cost is eaten there.\nIt seems the only reason we have to synchronize all these methods is because a concurrent flush or compaction can change the scanner stack, other than that only a single thread should access a StoreScanner at any given time.\nSo replaced updateReaders() with some code that just indicates to the scanner that the readers should be updated and then make it the using thread's responsibility to do the work.\nThe perf improvement from this is staggering. I am seeing somewhere around 3x scan performance improvement across all scenarios.\n\nNow, the hard part is to reason about whether this is 100% correct. I ran TestAtomicOperation and TestAcidGuarantees a few times in a loop, all still pass.\n\nWill attach a sample patch.","issue_id":"12680362","key":"HBASE-10015","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2013-11-26T21:11:24.000+0000","role":"fixed_distractor","summary":"Replace intrinsic locking with explicit locks in StoreScanner"} {"case_id":"12688557","cluster":"DISTRACTOR-HBASE-10330","comments":[{"body":"Is that the case in 0.94 as well? (I'll check)","created":"2014-01-13T17:03:09.242+0000"},{"body":"Seems like it would have a straightforward fix. ","created":"2014-01-13T18:00:02.464+0000"},{"body":"Raising the priority won't get this issue fixed faster [~bijucd], but providing a patch might.","created":"2014-06-27T18:23:50.400+0000"},{"body":"TableOutputFormat doesn't have this problem:\n{code}\n public void close(TaskAttemptContext context)\n throws IOException {\n table.close();\n }\n{code}","created":"2014-06-27T19:42:50.050+0000"},{"body":"Running test suite now","created":"2014-06-27T19:47:19.562+0000"},{"body":"+1, pending buildbot's opinion.","created":"2014-06-27T20:44:16.308+0000"},{"body":"QA bot is blocked by compilation issue.\nTest suite passed on Linux:\n{code}\nTests run: 4, Failures: 0, Errors: 0, Skipped: 0\n\n[INFO]\n[INFO] ------------------------------------------------------------------------\n[INFO] Building HBase - Assembly 0.99.0-SNAPSHOT\n[INFO] ------------------------------------------------------------------------\n[INFO]\n[INFO] --- maven-remote-resources-plugin:1.4:process (default) @ hbase-assembly ---\n[INFO]\n[INFO] --- maven-dependency-plugin:2.8:build-classpath (create-hbase-generated-classpath) @ hbase-assembly ---\n[INFO] Wrote classpath file '/homes/hortonzy/trunk/target/cached_classpath.txt'.\n[INFO] ------------------------------------------------------------------------\n[INFO] Reactor Summary:\n[INFO]\n[INFO] HBase ............................................. SUCCESS [1.789s]\n[INFO] HBase - Common .................................... SUCCESS [30.709s]\n[INFO] HBase - Protocol .................................. SUCCESS [0.272s]\n[INFO] HBase - Client .................................... SUCCESS [47.358s]\n[INFO] HBase - Hadoop Compatibility ...................... SUCCESS [5.062s]\n[INFO] HBase - Hadoop Two Compatibility .................. SUCCESS [1.462s]\n[INFO] HBase - Prefix Tree ............................... SUCCESS [2.550s]\n[INFO] HBase - Server .................................... SUCCESS [1:22:00.851s]\n[INFO] HBase - Testing Util .............................. SUCCESS [1.310s]\n[INFO] HBase - Thrift .................................... SUCCESS [2:03.932s]\n[INFO] HBase - Shell ..................................... SUCCESS [1:45.499s]\n[INFO] HBase - Integration Tests ......................... SUCCESS [0.877s]\n[INFO] HBase - Examples .................................. SUCCESS [5.575s]\n[INFO] HBase - Assembly .................................. SUCCESS [0.913s]\n[INFO] ------------------------------------------------------------------------\n[INFO] BUILD SUCCESS\n[INFO] ------------------------------------------------------------------------\n[INFO] Total time: 1:27:28.965s\n[INFO] Finished at: Fri Jun 27 21:39:30 UTC 2014\n[INFO] Final Memory: 49M/432M\n{code}","created":"2014-06-27T21:40:30.368+0000"},{"body":"Planning to integrate later today if there is no more review comment.","created":"2014-06-27T22:42:15.594+0000"},{"body":"Closing this issue after 0.99.0 release. ","created":"2015-02-21T23:29:24.103+0000"}],"conversations":[{"body":"As far as I can tell, TableInputFormat creates an instance of HTable which is used by TableRecordReaderImpl. However TableRecordReaderImpl.close() only closes the scanner, not the table. In turn the HTable's HConnection's reference count is never decreased which leads to leaking HConnections.\n\nTableOutputFormat might have a similar bug.","from":"reporter","subject":"TableInputFormat/TableRecordReaderImpl leaks HTable"},{"body":"Is that the case in 0.94 as well? (I'll check)","from":"developer"},{"body":"Seems like it would have a straightforward fix. ","from":"developer"},{"body":"Raising the priority won't get this issue fixed faster [~bijucd], but providing a patch might.","from":"developer"},{"body":"TableOutputFormat doesn't have this problem:\n{code}\n public void close(TaskAttemptContext context)\n throws IOException {\n table.close();\n }\n{code}","from":"developer"},{"body":"Running test suite now","from":"developer"},{"body":"+1, pending buildbot's opinion.","from":"developer"},{"body":"QA bot is blocked by compilation issue.\nTest suite passed on Linux:\n{code}\nTests run: 4, Failures: 0, Errors: 0, Skipped: 0\n\n[INFO]\n[INFO] ------------------------------------------------------------------------\n[INFO] Building HBase - Assembly 0.99.0-SNAPSHOT\n[INFO] ------------------------------------------------------------------------\n[INFO]\n[INFO] --- maven-remote-resources-plugin:1.4:process (default) @ hbase-assembly ---\n[INFO]\n[INFO] --- maven-dependency-plugin:2.8:build-classpath (create-hbase-generated-classpath) @ hbase-assembly ---\n[INFO] Wrote classpath file '/homes/hortonzy/trunk/target/cached_classpath.txt'.\n[INFO] ------------------------------------------------------------------------\n[INFO] Reactor Summary:\n[INFO]\n[INFO] HBase ............................................. SUCCESS [1.789s]\n[INFO] HBase - Common .................................... SUCCESS [30.709s]\n[INFO] HBase - Protocol .................................. SUCCESS [0.272s]\n[INFO] HBase - Client .................................... SUCCESS [47.358s]\n[INFO] HBase - Hadoop Compatibility ...................... SUCCESS [5.062s]\n[INFO] HBase - Hadoop Two Compatibility .................. SUCCESS [1.462s]\n[INFO] HBase - Prefix Tree ............................... SUCCESS [2.550s]\n[INFO] HBase - Server .................................... SUCCESS [1:22:00.851s]\n[INFO] HBase - Testing Util .............................. SUCCESS [1.310s]\n[INFO] HBase - Thrift .................................... SUCCESS [2:03.932s]\n[INFO] HBase - Shell ..................................... SUCCESS [1:45.499s]\n[INFO] HBase - Integration Tests ......................... SUCCESS [0.877s]\n[INFO] HBase - Examples .................................. SUCCESS [5.575s]\n[INFO] HBase - Assembly .................................. SUCCESS [0.913s]\n[INFO] ------------------------------------------------------------------------\n[INFO] BUILD SUCCESS\n[INFO] ------------------------------------------------------------------------\n[INFO] Total time: 1:27:28.965s\n[INFO] Finished at: Fri Jun 27 21:39:30 UTC 2014\n[INFO] Final Memory: 49M/432M\n{code}","from":"developer"},{"body":"Planning to integrate later today if there is no more review comment.","from":"developer"},{"body":"Closing this issue after 0.99.0 release. ","from":"developer"}],"created":"2014-01-13T15:51:25.000+0000","description":"As far as I can tell, TableInputFormat creates an instance of HTable which is used by TableRecordReaderImpl. However TableRecordReaderImpl.close() only closes the scanner, not the table. In turn the HTable's HConnection's reference count is never decreased which leads to leaking HConnections.\n\nTableOutputFormat might have a similar bug.","issue_id":"12688557","key":"HBASE-10330","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2014-06-28T02:21:21.000+0000","role":"fixed_distractor","summary":"TableInputFormat/TableRecordReaderImpl leaks HTable"} {"case_id":"12707760","cluster":"DISTRACTOR-HBASE-10958","comments":[{"body":"One workaround we found is to completely disable compactions, then when you need to run them you have to force flush the regions that have bulk loaded file first and ensure that bulk loads aren't coming in at the same time.\n\nWorkloads that are strictly doing incremental bulk loads aren't affected, you need a mix of bulk loaded files and normal Puts.\n\nA hacky solution could be to force flush when bulk loading with seqids and grab the next sequence id that comes after the memstore flush to go to the bulk loaded file. This means that bulk loading needs to initiate a flush, get the sequence id under the region write lock, then do the bulk load. We don't need to wait for the flush to happen... unless the possibility for the bulk loaded file to be compacted before the flush is done is high enough.","created":"2014-04-10T18:17:19.017+0000"},{"body":"Seems fine to force a flush before bulk loading.","created":"2014-04-10T18:43:51.709+0000"},{"body":"Here's a quick hack I put together to show one solution. In this case I inline a flush with the {{bulkLoadHFiles}} call and I modified the {{internalFlushcache}} code to be able to get more state back and also return a sequential ID that ends up being in between two memstores.\n\nI tested that it works following the steps listed in this jira's description.\n\nI don't really like having to wait for the flush to happen since it could take a lot of time to finish making bulk loading way slower. On the other hand, it's safer than just requesting a flush asynchronously and then grabbing a new sequence ID from HLog. Maybe it would still be fine since you need to also have a compaction to trigger the bug...","created":"2014-04-10T23:45:32.585+0000"},{"body":"Here's a less intrusive hack that shows what it looks like to request an asynchronous flush from {{bulkLoadHFiles}}. I also tested that it works like the previous patch.\n\nThe benefits are that the bulk loader doesn't have to wait and it's a lot less code, but the blindspot is still there and it widens as the region server gets more load; for example, if a lot of flushing is happening, which is likely with this patch, then the flush requests are queued and a compacting can kick in and at that moment if the region server dies then data is lost.","created":"2014-04-11T00:02:09.546+0000"},{"body":"I'll throw my 2 canadien cents and say that I prefer the \"removing-the-second-lane\" solution (the one where we do a synchronous flush prior to bulk load). My rationale is that the other approach still leaves us with a blindspot, even if greatly reduced and we would effectively still have the bug (potentially). ","created":"2014-04-11T04:30:08.941+0000"},{"body":"I'm w/ [~alexandre.normand].\n\nFlushState should be FlushResult? Does it have to be public?\n\nShould below be HBaseIOE (you know who is watching) or some specialization on HBaseIOE, FlushFailedIOE(FlushState)?\n\nthrow new IOException(\n\nWhat is happening here:\n\n+ seqId = fs.flushSequenceId;\n\nThis is the 'next' seqid after the flush or the seqid that was written into the hfile that was flushed?\n\nDo we need to up the seqid if the flush one is being used for a bulk load so the next edit in memstore has a different number? Or that is done already elsewhere?\n\nGood stuff JD. A unit test would be too hard to conjure? Maybe describe then instead how you replicate.","created":"2014-04-11T05:02:15.520+0000"},{"body":"bq. FlushState should be FlushResult? Does it have to be public?\n\nSome external classes to that package call {{flushcache()}} directly. Also, that method is public so whatever it returns needs to be as visible.\n\nbq. Should below be HBaseIOE (you know who is watching) or some specialization on HBaseIOE, FlushFailedIOE(FlushState)?\n\nIt's inline with the rest of that method, which the bulk loading client seems to process correctly. Not that it shouldn't be considered, but maybe in a different jira?\n\n{quote}\nThis is the 'next' seqid after the flush or the seqid that was written into the hfile that was flushed?\nDo we need to up the seqid if the flush one is being used for a bulk load so the next edit in memstore has a different number? Or that is done already elsewhere?\n{quote}\n\nYeah that part is hard to follow, needs at least some documentation (hey it is a hack!). So it's set here:\n\n+ return new FlushState(flushSeqId, compactionRequested);\n\nThat {{flushSeqId}} comes from here:\n\n{code}\n Long startSeqId = wal.startCacheFlush(this.getRegionInfo().getEncodedNameAsBytes());\n if (startSeqId == null) {\n status.setStatus(\"Flush will not be started for [\" + this.getRegionInfo().getEncodedName()\n + \"] - WAL is going away\");\n return false;\n }\n flushSeqId = startSeqId.longValue();\n{code}\n\nSo it's a sequence id that's the same one as the one used to signal the flush. I thought about creating a new one just after, but I'm not sure if it's necessary since that {{startSeqId}} will be after all the MemStore edits.\n\nbq. A unit test would be too hard to conjure? Maybe describe then instead how you replicate.\n\nDoesn't seem too hard to write... at least the recreating what I'm doing manually shouldn't take too long to write, but testing the corner cases seems much harder.\n\nI'm still wondering if there's a more elegant solution.","created":"2014-04-11T16:27:31.718+0000"},{"body":"Patch that addresses most of Stack's comments and based on trunk. Still no unit test, spent too much time today trying to understand why {{TestHRegion}} fails so much (see HBASE-10312).","created":"2014-04-12T00:12:25.292+0000"},{"body":"New patch with a unit test inside {{TestWALReplay}}. I've done a bit of refactoring and I'm tempted to go a step further and somehow extract the common bits from testRegionMadeOfBulkLoadedFilesOnly and testCompactedBulkLoadedFiles but it seems a bit messy.","created":"2014-04-15T20:20:44.606+0000"},{"body":"Oh and I found that testRegionMadeOfBulkLoadedFilesOnly wasn't really testing WAL replay with a bulk loaded file, the KV was being added at LATEST_TIMESTAMP so it wasn't visible being so far in the future. I made it so we check that we got all the rows back.","created":"2014-04-15T20:21:55.933+0000"},{"body":"Patch looks great [~jdcryans]. Minor nit is that you do\n\n+ if (fs.flushSucceeded()) {\n+ seqId = fs.flushSequenceId;\n+ } else if (fs.result == FlushResult.Result.CANNOT_FLUSH_MEMSTORE_EMPTY) {\n\nflushSucceeded method (should it be isFlushSucceeded) and then you go get the result by accessing the data member directly. Minor inconsistency.\n\n+1 on commit if hadoopqa ok. Can address above if you want on commit.","created":"2014-04-15T20:35:04.888+0000"},{"body":"That's a real test failure, the user that bulk loads doesn't have the permission to flush. Looking.","created":"2014-04-15T22:14:08.566+0000"},{"body":"Oh; that's an unforeseen problem. Does it make sense for a user to be able to bulkload but not to flush? (I suppose it does as bulk loading is not in principle different from inserting data via Put/Delete).\n","created":"2014-04-15T22:50:23.025+0000"},{"body":"Yeah, and bulk loading itself is kind of like a flush since you end up with an HFile.","created":"2014-04-15T22:53:42.166+0000"},{"body":"Right now bulk loading is WRITE and flush is ADMIN. Problematic!","created":"2014-04-15T22:58:00.403+0000"},{"body":"Solutions on top of my head:\n\n- Do doAs inside the method. The rationale being that it's the region server that's trying to do that here, not the user. I'm not sure if this is doable, and [~mbertozzi] is laughing at me.\n\n- Matteo thinks that we should just make bulk load an ADMIN method. I agree but I'm not fond of breaking secure setups that use bulk loads. [~apurtell]?","created":"2014-04-15T23:06:10.501+0000"},{"body":"Wait. Are you saying the region server itself cannot issue a flush as part of the bulkLoadHFiles RPC?\nThe region server is flushing all the time on its own behalf (when the memstore is full, etc).\n","created":"2014-04-15T23:51:55.484+0000"},{"body":"bq. Wait. Are you saying the region server itself cannot issue a flush as part of the bulkLoadHFiles RPC?\n\nExact. That's why the User.runAs (and then run as the region server itself) solution seems to make sense to me.\n\nI read some more code, and it seems bulk load actually requires CREATE even though the call itself requires WRITE. From TestAccessController:\n\n{code}\n // User performing bulk loads must have privilege to read table metadata\n // (ADMIN or CREATE)\n verifyAllowed(bulkLoadAction, SUPERUSER, USER_ADMIN, USER_OWNER, USER_CREATE);\n verifyDenied(bulkLoadAction, USER_RW, USER_NONE, USER_RO);\n{code}\n\nSo another options is to set flush and compact as CREATE actions (in line with create table, alter, disable, enable, delete).","created":"2014-04-15T23:58:25.436+0000"},{"body":"I guess I did not realize that the authorization extends to any action issued from an RPC, rather than just authorizing the RPC call itself.\n","created":"2014-04-16T00:10:51.886+0000"},{"body":"We expect READ and WRITE perms granted in a fine grained way to constraint who can do individual ops that only collectively add up to cluster impacting events like compactions, splits, and flushes. For actions that can have a global cluster impact, we'd like ADMIN to be granted sparingly to admins or delegates. IIRC enable and disable are ADMIN actions also, since disabling or enabling a 10000 region table has consequences. CREATE is kind of a middle ground for schema reads and updates, but in terms of schema update that's splitting hairs I suppose since a schema update of said large table would also have consequences of the same scale.\n\nBulk load is a special snowflake because it's a series of puts (so, WRITE) yet obviously more than that as mentioned, we need to flush, and moving files in place will probably kick off compaction. Making bulk load an ADMIN action, or CREATE, makes sense to me also.","created":"2014-04-16T04:17:26.256+0000"},{"body":"And the bit about needing CREATE or ADMIN to read schema metadata, this is to protect potentially sensitive information in the metadata that an ordinary user granted only READ or READ+WRITE access to the table has no need to see, that was HBASE-8692","created":"2014-04-16T04:22:00.345+0000"},{"body":"Very nice explanation and insights on these permissions!\nLooking previously at this page: https://hbase.apache.org/book/hbase.accesscontrol.configuration.html\nwe only have pretty vague info over there.\nAlso, the 'write' permission for 'flush' and 'compact' in 'Table 8.1'. Are they even correct?","created":"2014-04-16T06:11:33.512+0000"},{"body":"bq. we'd like ADMIN to be granted sparingly to admins or delegates\n\n+1\n\nbq. IIRC enable and disable are ADMIN actions also, since disabling or enabling a 10000 region table has consequences.\n\nThis is what I see in the code in trunk:\n\n{code}\nrequirePermission(\"preBulkLoadHFile\", ....getTableDesc().getTableName(), el.getFirst(), null, Permission.Action.WRITE);\nrequirePermission(\"enableTable\", tableName, null, null, Action.ADMIN, Action.CREATE);\nrequirePermission(\"disableTable\", tableName, null, null, Action.ADMIN, Action.CREATE);\nrequirePermission(\"compact\", getTableName(e.getEnvironment()), null, null, Action.ADMIN);\nrequirePermission(\"flush\", getTableName(e.getEnvironment()), null, null, Action.ADMIN);\n{code}\n\nIMO flush should have lower or same perms as disableTable.\n\nSo here's a list of changes I believe are needed:\n\n - preBulkLoadHFile goes from WRITE to CREATE (seems more in line with what's really needed to bulk load given the code I posted yesterday)\n - compact/flush go from ADMIN to ADMIN or CREATE\n\nThis should not have an impact on the current users. If we can agree on the changes, I'll open a new jira that's going to be blocking this one.","created":"2014-04-16T15:59:21.499+0000"},{"body":"Maybe \"CREATE\" no longer expresses what it now implies...?\n\nI can see that folks would not want to grant users CREATE (or ADMIN) so that they cannot create/drop/enable/disable tables, but still do allow them to load data via bulk load. That would now no longer possible.\n\nAlso it would seem more sensible to me that if a user bulk loads some data and then for *technical* reasons the region server decides to flush there should be no additional right needed; just a user does not need permission to flush only because a Put happens to cause a flush.\nHave we dismissed this option?","created":"2014-04-16T17:58:53.436+0000"},{"body":"bq. I can see that folks would not want to grant users CREATE (or ADMIN) so that they cannot create/drop/enable/disable tables, but still do allow them to load data via bulk load. That would now no longer possible.\n\nBulk load already needs CREATE as shown above in TestAccessController.\n\nbq. Have we dismissed this option?\n\nI don't think so, but no one has been pushing for it. It creates a precedent, nowhere else in the code do we use User.runAs to override the current user that came in via a RPC (I'm not even sure if it works, but that's because I don't know that code very well).","created":"2014-04-16T18:07:41.831+0000"},{"body":"Missed the test comment. In that case let's do what you suggest.\n(Since this in an existing issue I might sill want release 0.94.19 before we fix it depending on whether this needs more discussion)","created":"2014-04-16T18:18:29.798+0000"},{"body":"Just did a quick test. The requirement on 'CREATE' for bulk load seems to come from here. Is this even intended?\n{code}\nException in thread \"main\" org.apache.hadoop.hbase.security.AccessDeniedException: org.apache.hadoop.hbase.security.AccessDeniedException: Insufficient permissions (user=user1@IBM.COM, scope=TestTable, family=, action=CREATE)\n at org.apache.hadoop.hbase.security.access.AccessController.requirePermission(AccessController.java:356)\n at org.apache.hadoop.hbase.security.access.AccessController.preGetTableDescriptors(AccessController.java:1513)\n at org.apache.hadoop.hbase.master.MasterCoprocessorHost.preGetTableDescriptors(MasterCoprocessorHost.java:1260)\n at org.apache.hadoop.hbase.master.HMaster.getTableDescriptors(HMaster.java:2569)\n at org.apache.hadoop.hbase.protobuf.generated.MasterProtos$MasterService$2.callBlockingMethod(MasterProtos.java:40438)\n at org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2150)\n ...\n at org.apache.hadoop.hbase.protobuf.ProtobufUtil.getRemoteException(ProtobufUtil.java:235)\n at org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.getHTableDescriptor(HConnectionManager.java:2632)\n at org.apache.hadoop.hbase.client.HTable.getTableDescriptor(HTable.java:548)\n at org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles.doBulkLoad(LoadIncrementalHFiles.java:233)\n at org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles.run(LoadIncrementalHFiles.java:820)\n{code}","created":"2014-04-16T18:29:20.308+0000"},{"body":"bq. The requirement on 'CREATE' for bulk load seems to come from here. Is this even intended?\n\nThanks for doing this. There's also another place where this is called, splitStoreFile(), in order to get an HCD. I don't think it's necessary to call it twice, but it seems necessary to call it at least once else you can't:\n\n- verify that the families exist\n- get the schema for each family so that we create the HFiles with the correct configurations when splitting","created":"2014-04-16T21:56:04.393+0000"},{"body":"Forgot to reply to [~stack]'s comment:\n\nbq. flushSucceeded method (should it be isFlushSucceeded) and then you go get the result by accessing the data member directly. Minor inconsistency.\n\nflushSucceeded() isn't just looking up a field though, it's checking two things. ","created":"2014-04-16T23:51:44.715+0000"},{"body":"bq. IMO flush should have lower or same perms as disableTable.\n\nThat seems fine. \n\nbq. Maybe \"CREATE\" no longer expresses what it now implies...?\n\nAt the least we have an imperfect idea of when a user should be able to create tables and administer them, just \"not administer them too much\"","created":"2014-04-17T03:20:08.801+0000"},{"body":"bq. The requirement on 'CREATE' for bulk load seems to come from [ getTableDescriptors ]. Is this even intended?\n\nYes.\n\nCREATE is overloaded to mean \"RESTRICTED ADMIN\", with ADMIN-ish privilege required because table schema is considered potentially sensitive. \n\nOn another issue Francis Liu and I discussed the notion of creating a new permission 'SCHEMA' which would grant permission to read schema metadata. Now as then it seems maybe not quite needed (yet). CREATE and ADMIN would have such SCHEMA permission implicitly, so how useful would SCHEMA be, and there would still need a grant beyond WRITE for bulk loading.","created":"2014-04-17T03:37:50.130+0000"},{"body":"Attaching 2 patches.\n\nOne is a backport for 0.94. While doing the backport I saw that a TestSnapshotFromMaster was failing and Matteo was able to see that it was an error in my patch, flushing always returned that it needed compaction. I need to rerun all the tests now but it was the only one that failed (haven't tried with security either). I also added a test in TestHRegion for that.\n\nThe second patch is for trunk, in which I ported the same test to TestHRegion. Interestingly, it didn't work. I found that in 0.94 we compact if num_files > compactionThreshold, but in trunk it's >=, so it seems that we compact more often now. This patch also has the fixes from [~stack]'s comments.","created":"2014-04-18T00:40:17.861+0000"},{"body":"FYI rather than do this:\n\n+ String method = \"testFlushResult\";\n\nYou can do this:\n\n method = name.getMethodName();\n\n... because in TestHRegion it does this:\n\n @Rule public TestName name = new TestName();\n\nSee here http://stackoverflow.com/questions/473401/get-name-of-currently-executing-test-in-junit-4\n\nOn the sometime an accessor, sometime not... not going to argue. nit.\n\nPatch LGTM (where G==Great)","created":"2014-04-18T04:32:15.362+0000"},{"body":"bq. method = name.getMethodName();\n\nWill fix (looks like I copied that from one of the few methods in that class that doesn't do it).","created":"2014-04-18T04:53:28.916+0000"},{"body":"bq. Will fix (looks like I copied that from one of the few methods in that class that doesn't do it).\n\nit is a nit. fix on commit?","created":"2014-04-18T05:03:23.377+0000"},{"body":"Nice. +1","created":"2014-04-18T21:40:04.551+0000"},{"body":"Moving to 0.94.20.","created":"2014-04-21T18:35:28.741+0000"},{"body":"Committed to 0.96 and up. Like for HBASE-11008, I'm waiting to commit to 0.94 or I can open a backport jira.","created":"2014-04-25T21:06:08.353+0000"},{"body":"Are you waiting for me to commit [~jdcryans]? Just making sure we're not mutually waiting.","created":"2014-04-30T20:39:58.648+0000"},{"body":"Sorry, went to do something else, lemme get that in (with HBASE-11008 first).","created":"2014-04-30T20:42:27.270+0000"},{"body":"Now committed to 0.94 (took some time because I wanted to double check the tests I modified were all green). Thanks everyone.","created":"2014-04-30T22:09:09.855+0000"},{"body":"You 'd man.","created":"2014-04-30T23:40:04.345+0000"}],"conversations":[{"body":"We found an issue with bulk loads causing data loss when assigning sequence ids (HBASE-6630) that is triggered when replaying recovered edits. We're nicknaming this issue *Blindspot*.\n\nThe problem is that the sequence id given to a bulk loaded file is higher than those of the edits in the region's memstore. When replaying recovered edits, the rule to skip some of them is that they have to be _lower than the highest sequence id_. In other words, the edits that have a sequence id lower than the highest one in the store files *should* have also been flushed. This is not the case with bulk loaded files since we now have an HFile with a sequence id higher than unflushed edits.\n\nThe log recovery code takes this into account by simply skipping the bulk loaded files, but this \"bulk loaded status\" is *lost* on compaction. The edits in the logs that have a sequence id lower than the bulk loaded file that got compacted are put in a blind spot and are skipped during replay.\n\nHere's the easiest way to recreate this issue:\n - Create an empty table\n - Put one row in it (let's say it gets seqid 1)\n - Bulk load one file (it gets seqid 2). I used ImporTsv and set hbase.mapreduce.bulkload.assign.sequenceNumbers.\n - Bulk load a second file the same way (it gets seqid 3).\n - Major compact the table (the new file has seqid 3 and isn't considered bulk loaded).\n - Kill the region server that holds the table's region.\n - Scan the table once the region is made available again. The first row, at seqid 1, will be missing since the HFile with seqid 3 makes us believe that everything that came before it was flushed.","from":"reporter","subject":"[dataloss] Bulk loading with seqids can prevent some log entries from being replayed"},{"body":"One workaround we found is to completely disable compactions, then when you need to run them you have to force flush the regions that have bulk loaded file first and ensure that bulk loads aren't coming in at the same time.\n\nWorkloads that are strictly doing incremental bulk loads aren't affected, you need a mix of bulk loaded files and normal Puts.\n\nA hacky solution could be to force flush when bulk loading with seqids and grab the next sequence id that comes after the memstore flush to go to the bulk loaded file. This means that bulk loading needs to initiate a flush, get the sequence id under the region write lock, then do the bulk load. We don't need to wait for the flush to happen... unless the possibility for the bulk loaded file to be compacted before the flush is done is high enough.","from":"developer"},{"body":"Seems fine to force a flush before bulk loading.","from":"developer"},{"body":"Here's a quick hack I put together to show one solution. In this case I inline a flush with the {{bulkLoadHFiles}} call and I modified the {{internalFlushcache}} code to be able to get more state back and also return a sequential ID that ends up being in between two memstores.\n\nI tested that it works following the steps listed in this jira's description.\n\nI don't really like having to wait for the flush to happen since it could take a lot of time to finish making bulk loading way slower. On the other hand, it's safer than just requesting a flush asynchronously and then grabbing a new sequence ID from HLog. Maybe it would still be fine since you need to also have a compaction to trigger the bug...","from":"developer"},{"body":"Here's a less intrusive hack that shows what it looks like to request an asynchronous flush from {{bulkLoadHFiles}}. I also tested that it works like the previous patch.\n\nThe benefits are that the bulk loader doesn't have to wait and it's a lot less code, but the blindspot is still there and it widens as the region server gets more load; for example, if a lot of flushing is happening, which is likely with this patch, then the flush requests are queued and a compacting can kick in and at that moment if the region server dies then data is lost.","from":"developer"},{"body":"I'll throw my 2 canadien cents and say that I prefer the \"removing-the-second-lane\" solution (the one where we do a synchronous flush prior to bulk load). My rationale is that the other approach still leaves us with a blindspot, even if greatly reduced and we would effectively still have the bug (potentially). ","from":"developer"},{"body":"I'm w/ [~alexandre.normand].\n\nFlushState should be FlushResult? Does it have to be public?\n\nShould below be HBaseIOE (you know who is watching) or some specialization on HBaseIOE, FlushFailedIOE(FlushState)?\n\nthrow new IOException(\n\nWhat is happening here:\n\n+ seqId = fs.flushSequenceId;\n\nThis is the 'next' seqid after the flush or the seqid that was written into the hfile that was flushed?\n\nDo we need to up the seqid if the flush one is being used for a bulk load so the next edit in memstore has a different number? Or that is done already elsewhere?\n\nGood stuff JD. A unit test would be too hard to conjure? Maybe describe then instead how you replicate.","from":"developer"},{"body":"bq. FlushState should be FlushResult? Does it have to be public?\n\nSome external classes to that package call {{flushcache()}} directly. Also, that method is public so whatever it returns needs to be as visible.\n\nbq. Should below be HBaseIOE (you know who is watching) or some specialization on HBaseIOE, FlushFailedIOE(FlushState)?\n\nIt's inline with the rest of that method, which the bulk loading client seems to process correctly. Not that it shouldn't be considered, but maybe in a different jira?\n\n{quote}\nThis is the 'next' seqid after the flush or the seqid that was written into the hfile that was flushed?\nDo we need to up the seqid if the flush one is being used for a bulk load so the next edit in memstore has a different number? Or that is done already elsewhere?\n{quote}\n\nYeah that part is hard to follow, needs at least some documentation (hey it is a hack!). So it's set here:\n\n+ return new FlushState(flushSeqId, compactionRequested);\n\nThat {{flushSeqId}} comes from here:\n\n{code}\n Long startSeqId = wal.startCacheFlush(this.getRegionInfo().getEncodedNameAsBytes());\n if (startSeqId == null) {\n status.setStatus(\"Flush will not be started for [\" + this.getRegionInfo().getEncodedName()\n + \"] - WAL is going away\");\n return false;\n }\n flushSeqId = startSeqId.longValue();\n{code}\n\nSo it's a sequence id that's the same one as the one used to signal the flush. I thought about creating a new one just after, but I'm not sure if it's necessary since that {{startSeqId}} will be after all the MemStore edits.\n\nbq. A unit test would be too hard to conjure? Maybe describe then instead how you replicate.\n\nDoesn't seem too hard to write... at least the recreating what I'm doing manually shouldn't take too long to write, but testing the corner cases seems much harder.\n\nI'm still wondering if there's a more elegant solution.","from":"developer"},{"body":"Patch that addresses most of Stack's comments and based on trunk. Still no unit test, spent too much time today trying to understand why {{TestHRegion}} fails so much (see HBASE-10312).","from":"developer"},{"body":"New patch with a unit test inside {{TestWALReplay}}. I've done a bit of refactoring and I'm tempted to go a step further and somehow extract the common bits from testRegionMadeOfBulkLoadedFilesOnly and testCompactedBulkLoadedFiles but it seems a bit messy.","from":"developer"},{"body":"Oh and I found that testRegionMadeOfBulkLoadedFilesOnly wasn't really testing WAL replay with a bulk loaded file, the KV was being added at LATEST_TIMESTAMP so it wasn't visible being so far in the future. I made it so we check that we got all the rows back.","from":"developer"},{"body":"Patch looks great [~jdcryans]. Minor nit is that you do\n\n+ if (fs.flushSucceeded()) {\n+ seqId = fs.flushSequenceId;\n+ } else if (fs.result == FlushResult.Result.CANNOT_FLUSH_MEMSTORE_EMPTY) {\n\nflushSucceeded method (should it be isFlushSucceeded) and then you go get the result by accessing the data member directly. Minor inconsistency.\n\n+1 on commit if hadoopqa ok. Can address above if you want on commit.","from":"developer"},{"body":"That's a real test failure, the user that bulk loads doesn't have the permission to flush. Looking.","from":"developer"},{"body":"Oh; that's an unforeseen problem. Does it make sense for a user to be able to bulkload but not to flush? (I suppose it does as bulk loading is not in principle different from inserting data via Put/Delete).\n","from":"developer"},{"body":"Yeah, and bulk loading itself is kind of like a flush since you end up with an HFile.","from":"developer"},{"body":"Right now bulk loading is WRITE and flush is ADMIN. Problematic!","from":"developer"},{"body":"Solutions on top of my head:\n\n- Do doAs inside the method. The rationale being that it's the region server that's trying to do that here, not the user. I'm not sure if this is doable, and [~mbertozzi] is laughing at me.\n\n- Matteo thinks that we should just make bulk load an ADMIN method. I agree but I'm not fond of breaking secure setups that use bulk loads. [~apurtell]?","from":"developer"},{"body":"Wait. Are you saying the region server itself cannot issue a flush as part of the bulkLoadHFiles RPC?\nThe region server is flushing all the time on its own behalf (when the memstore is full, etc).\n","from":"developer"},{"body":"bq. Wait. Are you saying the region server itself cannot issue a flush as part of the bulkLoadHFiles RPC?\n\nExact. That's why the User.runAs (and then run as the region server itself) solution seems to make sense to me.\n\nI read some more code, and it seems bulk load actually requires CREATE even though the call itself requires WRITE. From TestAccessController:\n\n{code}\n // User performing bulk loads must have privilege to read table metadata\n // (ADMIN or CREATE)\n verifyAllowed(bulkLoadAction, SUPERUSER, USER_ADMIN, USER_OWNER, USER_CREATE);\n verifyDenied(bulkLoadAction, USER_RW, USER_NONE, USER_RO);\n{code}\n\nSo another options is to set flush and compact as CREATE actions (in line with create table, alter, disable, enable, delete).","from":"developer"},{"body":"I guess I did not realize that the authorization extends to any action issued from an RPC, rather than just authorizing the RPC call itself.\n","from":"developer"},{"body":"We expect READ and WRITE perms granted in a fine grained way to constraint who can do individual ops that only collectively add up to cluster impacting events like compactions, splits, and flushes. For actions that can have a global cluster impact, we'd like ADMIN to be granted sparingly to admins or delegates. IIRC enable and disable are ADMIN actions also, since disabling or enabling a 10000 region table has consequences. CREATE is kind of a middle ground for schema reads and updates, but in terms of schema update that's splitting hairs I suppose since a schema update of said large table would also have consequences of the same scale.\n\nBulk load is a special snowflake because it's a series of puts (so, WRITE) yet obviously more than that as mentioned, we need to flush, and moving files in place will probably kick off compaction. Making bulk load an ADMIN action, or CREATE, makes sense to me also.","from":"developer"},{"body":"And the bit about needing CREATE or ADMIN to read schema metadata, this is to protect potentially sensitive information in the metadata that an ordinary user granted only READ or READ+WRITE access to the table has no need to see, that was HBASE-8692","from":"developer"},{"body":"Very nice explanation and insights on these permissions!\nLooking previously at this page: https://hbase.apache.org/book/hbase.accesscontrol.configuration.html\nwe only have pretty vague info over there.\nAlso, the 'write' permission for 'flush' and 'compact' in 'Table 8.1'. Are they even correct?","from":"developer"},{"body":"bq. we'd like ADMIN to be granted sparingly to admins or delegates\n\n+1\n\nbq. IIRC enable and disable are ADMIN actions also, since disabling or enabling a 10000 region table has consequences.\n\nThis is what I see in the code in trunk:\n\n{code}\nrequirePermission(\"preBulkLoadHFile\", ....getTableDesc().getTableName(), el.getFirst(), null, Permission.Action.WRITE);\nrequirePermission(\"enableTable\", tableName, null, null, Action.ADMIN, Action.CREATE);\nrequirePermission(\"disableTable\", tableName, null, null, Action.ADMIN, Action.CREATE);\nrequirePermission(\"compact\", getTableName(e.getEnvironment()), null, null, Action.ADMIN);\nrequirePermission(\"flush\", getTableName(e.getEnvironment()), null, null, Action.ADMIN);\n{code}\n\nIMO flush should have lower or same perms as disableTable.\n\nSo here's a list of changes I believe are needed:\n\n - preBulkLoadHFile goes from WRITE to CREATE (seems more in line with what's really needed to bulk load given the code I posted yesterday)\n - compact/flush go from ADMIN to ADMIN or CREATE\n\nThis should not have an impact on the current users. If we can agree on the changes, I'll open a new jira that's going to be blocking this one.","from":"developer"},{"body":"Maybe \"CREATE\" no longer expresses what it now implies...?\n\nI can see that folks would not want to grant users CREATE (or ADMIN) so that they cannot create/drop/enable/disable tables, but still do allow them to load data via bulk load. That would now no longer possible.\n\nAlso it would seem more sensible to me that if a user bulk loads some data and then for *technical* reasons the region server decides to flush there should be no additional right needed; just a user does not need permission to flush only because a Put happens to cause a flush.\nHave we dismissed this option?","from":"developer"},{"body":"bq. I can see that folks would not want to grant users CREATE (or ADMIN) so that they cannot create/drop/enable/disable tables, but still do allow them to load data via bulk load. That would now no longer possible.\n\nBulk load already needs CREATE as shown above in TestAccessController.\n\nbq. Have we dismissed this option?\n\nI don't think so, but no one has been pushing for it. It creates a precedent, nowhere else in the code do we use User.runAs to override the current user that came in via a RPC (I'm not even sure if it works, but that's because I don't know that code very well).","from":"developer"},{"body":"Missed the test comment. In that case let's do what you suggest.\n(Since this in an existing issue I might sill want release 0.94.19 before we fix it depending on whether this needs more discussion)","from":"developer"},{"body":"Just did a quick test. The requirement on 'CREATE' for bulk load seems to come from here. Is this even intended?\n{code}\nException in thread \"main\" org.apache.hadoop.hbase.security.AccessDeniedException: org.apache.hadoop.hbase.security.AccessDeniedException: Insufficient permissions (user=user1@IBM.COM, scope=TestTable, family=, action=CREATE)\n at org.apache.hadoop.hbase.security.access.AccessController.requirePermission(AccessController.java:356)\n at org.apache.hadoop.hbase.security.access.AccessController.preGetTableDescriptors(AccessController.java:1513)\n at org.apache.hadoop.hbase.master.MasterCoprocessorHost.preGetTableDescriptors(MasterCoprocessorHost.java:1260)\n at org.apache.hadoop.hbase.master.HMaster.getTableDescriptors(HMaster.java:2569)\n at org.apache.hadoop.hbase.protobuf.generated.MasterProtos$MasterService$2.callBlockingMethod(MasterProtos.java:40438)\n at org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2150)\n ...\n at org.apache.hadoop.hbase.protobuf.ProtobufUtil.getRemoteException(ProtobufUtil.java:235)\n at org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.getHTableDescriptor(HConnectionManager.java:2632)\n at org.apache.hadoop.hbase.client.HTable.getTableDescriptor(HTable.java:548)\n at org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles.doBulkLoad(LoadIncrementalHFiles.java:233)\n at org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles.run(LoadIncrementalHFiles.java:820)\n{code}","from":"developer"},{"body":"bq. The requirement on 'CREATE' for bulk load seems to come from here. Is this even intended?\n\nThanks for doing this. There's also another place where this is called, splitStoreFile(), in order to get an HCD. I don't think it's necessary to call it twice, but it seems necessary to call it at least once else you can't:\n\n- verify that the families exist\n- get the schema for each family so that we create the HFiles with the correct configurations when splitting","from":"developer"},{"body":"Forgot to reply to [~stack]'s comment:\n\nbq. flushSucceeded method (should it be isFlushSucceeded) and then you go get the result by accessing the data member directly. Minor inconsistency.\n\nflushSucceeded() isn't just looking up a field though, it's checking two things. ","from":"developer"},{"body":"bq. IMO flush should have lower or same perms as disableTable.\n\nThat seems fine. \n\nbq. Maybe \"CREATE\" no longer expresses what it now implies...?\n\nAt the least we have an imperfect idea of when a user should be able to create tables and administer them, just \"not administer them too much\"","from":"developer"},{"body":"bq. The requirement on 'CREATE' for bulk load seems to come from [ getTableDescriptors ]. Is this even intended?\n\nYes.\n\nCREATE is overloaded to mean \"RESTRICTED ADMIN\", with ADMIN-ish privilege required because table schema is considered potentially sensitive. \n\nOn another issue Francis Liu and I discussed the notion of creating a new permission 'SCHEMA' which would grant permission to read schema metadata. Now as then it seems maybe not quite needed (yet). CREATE and ADMIN would have such SCHEMA permission implicitly, so how useful would SCHEMA be, and there would still need a grant beyond WRITE for bulk loading.","from":"developer"},{"body":"Attaching 2 patches.\n\nOne is a backport for 0.94. While doing the backport I saw that a TestSnapshotFromMaster was failing and Matteo was able to see that it was an error in my patch, flushing always returned that it needed compaction. I need to rerun all the tests now but it was the only one that failed (haven't tried with security either). I also added a test in TestHRegion for that.\n\nThe second patch is for trunk, in which I ported the same test to TestHRegion. Interestingly, it didn't work. I found that in 0.94 we compact if num_files > compactionThreshold, but in trunk it's >=, so it seems that we compact more often now. This patch also has the fixes from [~stack]'s comments.","from":"developer"},{"body":"FYI rather than do this:\n\n+ String method = \"testFlushResult\";\n\nYou can do this:\n\n method = name.getMethodName();\n\n... because in TestHRegion it does this:\n\n @Rule public TestName name = new TestName();\n\nSee here http://stackoverflow.com/questions/473401/get-name-of-currently-executing-test-in-junit-4\n\nOn the sometime an accessor, sometime not... not going to argue. nit.\n\nPatch LGTM (where G==Great)","from":"developer"},{"body":"bq. method = name.getMethodName();\n\nWill fix (looks like I copied that from one of the few methods in that class that doesn't do it).","from":"developer"},{"body":"bq. Will fix (looks like I copied that from one of the few methods in that class that doesn't do it).\n\nit is a nit. fix on commit?","from":"developer"},{"body":"Nice. +1","from":"developer"},{"body":"Moving to 0.94.20.","from":"developer"},{"body":"Committed to 0.96 and up. Like for HBASE-11008, I'm waiting to commit to 0.94 or I can open a backport jira.","from":"developer"},{"body":"Are you waiting for me to commit [~jdcryans]? Just making sure we're not mutually waiting.","from":"developer"},{"body":"Sorry, went to do something else, lemme get that in (with HBASE-11008 first).","from":"developer"},{"body":"Now committed to 0.94 (took some time because I wanted to double check the tests I modified were all green). Thanks everyone.","from":"developer"},{"body":"You 'd man.","from":"developer"}],"created":"2014-04-10T18:08:02.000+0000","description":"We found an issue with bulk loads causing data loss when assigning sequence ids (HBASE-6630) that is triggered when replaying recovered edits. We're nicknaming this issue *Blindspot*.\n\nThe problem is that the sequence id given to a bulk loaded file is higher than those of the edits in the region's memstore. When replaying recovered edits, the rule to skip some of them is that they have to be _lower than the highest sequence id_. In other words, the edits that have a sequence id lower than the highest one in the store files *should* have also been flushed. This is not the case with bulk loaded files since we now have an HFile with a sequence id higher than unflushed edits.\n\nThe log recovery code takes this into account by simply skipping the bulk loaded files, but this \"bulk loaded status\" is *lost* on compaction. The edits in the logs that have a sequence id lower than the bulk loaded file that got compacted are put in a blind spot and are skipped during replay.\n\nHere's the easiest way to recreate this issue:\n - Create an empty table\n - Put one row in it (let's say it gets seqid 1)\n - Bulk load one file (it gets seqid 2). I used ImporTsv and set hbase.mapreduce.bulkload.assign.sequenceNumbers.\n - Bulk load a second file the same way (it gets seqid 3).\n - Major compact the table (the new file has seqid 3 and isn't considered bulk loaded).\n - Kill the region server that holds the table's region.\n - Scan the table once the region is made available again. The first row, at seqid 1, will be missing since the HFile with seqid 3 makes us believe that everything that came before it was flushed.","issue_id":"12707760","key":"HBASE-10958","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2014-04-30T22:09:09.000+0000","role":"fixed_distractor","summary":"[dataloss] Bulk loading with seqids can prevent some log entries from being replayed"} {"case_id":"12411558","cluster":"DISTRACTOR-HBASE-1107","comments":[{"body":"Right above line 322 in HStoreScanner.java is this comment:\n\\\\\n{code}\n// I think it is safe getting key from mem at this stage -- it shouldn't have\n// been flushed yet\n{code}\n\nSo I think this is no longer true.","created":"2009-01-02T18:30:59.658+0000"},{"body":"Attached patch 1107-1 addresses the symptom. Is it enough?","created":"2009-01-02T18:45:11.842+0000"},{"body":"I think this NPE is a symptom of our making the list of changed readers observers sloppier using CopyOnWrite. I think the scanner has been closed. The first thing done on close is that we remove ourselves from the list of changed readers observers but my guess is that there is an outstanding iteration going on. It still has a reference and invoked the observer code AFTER the clearing of the memcache scanner.\n\nI'm not sure your attached patch safe enough Andrew. My worry is that between your new check on #320, this.keys[MEMS_INDEX] could be cleared before its use on #327 by another thread. What you think?\n\nThe observer code -- updateReaders method -- shouldn't run if close has been invoked methinks. I was going to add in a 'close' flag. Seems like overkill but ain't sure what else to use as 'close' indicator:\n\n{code}\nIndex: src/java/org/apache/hadoop/hbase/regionserver/HStoreScanner.java\n===================================================================\n--- src/java/org/apache/hadoop/hbase/regionserver/HStoreScanner.java (revision 731185)\n+++ src/java/org/apache/hadoop/hbase/regionserver/HStoreScanner.java (working copy)\n@@ -26,6 +26,7 @@\n import java.util.Set;\n import java.util.SortedMap;\n import java.util.TreeMap;\n+import java.util.concurrent.atomic.AtomicBoolean;\n import java.util.concurrent.locks.ReentrantReadWriteLock;\n \n import org.apache.commons.logging.Log;\n@@ -60,6 +61,8 @@\n // Used around transition from no storefile to the first.\n private final ReentrantReadWriteLock lock = new ReentrantReadWriteLock();\n \n+ private final AtomicBoolean closing = new AtomicBoolean(false);\n+ \n /** Create an Scanner with a handle on the memcache and HStore files. */\n @SuppressWarnings(\"unchecked\")\n HStoreScanner(HStore store, byte [][] targetCols, byte [] firstRow,\n@@ -294,6 +297,7 @@\n }\n \n public void close() {\n+ this.closing.set(true);\n this.store.deleteChangedReaderObserver(this);\n doClose();\n }\n@@ -309,6 +313,9 @@\n // Implementation of ChangedReadersObserver\n \n public void updateReaders() throws IOException {\n+ if (this.closing.get()) {\n+ return;\n+ }\n this.lock.writeLock().lock();\n try {\n MapFile.Reader [] readers = this.store.getReaders();\n{code}","created":"2009-01-05T20:23:36.805+0000"},{"body":"Concur that the scanner is closed when the observer gets it.\n\n+1\n\nI'll commit. ","created":"2009-01-05T20:34:40.684+0000"},{"body":"Committed. Passes all local tests. ","created":"2009-01-05T21:28:46.370+0000"},{"body":"Seen again, on 0.19.0 release:\n\n2009-01-25 03:56:39,324 FATAL org.apache.hadoop.hbase.regionserver.MemcacheFlusher: Replay of hlog required. Forcing server shutdown org.apache.hadoop.hbase.DroppedSnapshotException: region: urls,http|img3.megavideo.com|80|a|8|8b3035357a222073d30d09f2c4ae85.jpg,1232408450783\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:896)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:789)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushRegion(MemcacheFlusher.java:227)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushSomeRegions(MemcacheFlusher.java:291)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.reclaimMemcacheMemory(MemcacheFlusher.java:261)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.batchUpdates(HRegionServer.java:1614)\n at sun.reflect.GeneratedMethodAccessor11.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:632)\n at org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:895)\nCaused by: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HStoreScanner.updateReaders(HStoreScanner.java:330)\n at org.apache.hadoop.hbase.regionserver.HStore.notifyChangedReadersObservers(HStore.java:737)\n at org.apache.hadoop.hbase.regionserver.HStore.updateReaders(HStore.java:725)\n at org.apache.hadoop.hbase.regionserver.HStore.internalFlushCache(HStore.java:694)\n at org.apache.hadoop.hbase.regionserver.HStore.flushCache(HStore.java:630)\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:881)\n ... 10 more\n","created":"2009-01-25T04:08:19.075+0000"},{"body":"Thats crazy! I must be digging the hole in the wrong spot?","created":"2009-01-25T06:54:38.699+0000"},{"body":"Got another of these today.\n\n2009-02-09 00:16:55,081 FATAL org.apache.hadoop.hbase.regionserver.MemcacheFlusher: Replay of hlog required. Forcing server shutdown\norg.apache.hadoop.hbase.DroppedSnapshotException: region: content,26ac2c3fd24ac7ba1032e7b68973b9fa,1233995252171\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:896)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:789)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushRegion(MemcacheFlusher.java:227)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushSomeRegions(MemcacheFlusher.java:291)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.reclaimMemcacheMemory(MemcacheFlusher.java:261)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.batchUpdates(HRegionServer.java:1619)\n at sun.reflect.GeneratedMethodAccessor12.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:632)\n at org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:895)\nCaused by: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HStoreScanner.updateReaders(HStoreScanner.java:330)\n at org.apache.hadoop.hbase.regionserver.HStore.notifyChangedReadersObservers(HStore.java:737)\n at org.apache.hadoop.hbase.regionserver.HStore.updateReaders(HStore.java:725)\n at org.apache.hadoop.hbase.regionserver.HStore.internalFlushCache(HStore.java:694)\n at org.apache.hadoop.hbase.regionserver.HStore.flushCache(HStore.java:630)\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:881)\n ... 10 more\n","created":"2009-02-09T00:40:11.548+0000"},{"body":"This problem has been taking down regionservers regularly for me as well. Running hbase 0.19.0 on hadoop 0.19.0 with 4 slaves. It's becoming a major stability issue for us.\n\n2009-03-13 16:30:47,728 FATAL org.apache.hadoop.hbase.regionserver.MemcacheFlusher: Replay of hlog required. Forcing server shutdown\norg.apache.hadoop.hbase.DroppedSnapshotException: region: sessions,^@^@^@^@^@^@^Cd^@^@^A^_�x�q^@.IPHONEc8532c72874b66d1165c2f62ed91ef353eb9,1235185327583\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:896)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:789)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushRegion(MemcacheFlusher.java:227)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.run(MemcacheFlusher.java:137)\nCaused by: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HStoreScanner.updateReaders(HStoreScanner.java:330)\n at org.apache.hadoop.hbase.regionserver.HStore.notifyChangedReadersObservers(HStore.java:737)\n at org.apache.hadoop.hbase.regionserver.HStore.updateReaders(HStore.java:725)\n at org.apache.hadoop.hbase.regionserver.HStore.internalFlushCache(HStore.java:694)\n at org.apache.hadoop.hbase.regionserver.HStore.flushCache(HStore.java:630)\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:881)\n ... 3 more","created":"2009-03-13T23:43:50.916+0000"},{"body":"The code in updateReaders was changed significantly for 0.19.1 (We don't open readers we already had open).\n\nThe NPE above as Andrew has already pointed out comes from here:\n\n{code}\n // I think its safe getting key from mem at this stage -- it shouldn't have\n // been flushed yet\n this.scanners[HSFS_INDEX] = new StoreFileScanner(this.store,\n this.timestamp, this. targetCols, this.keys[MEMS_INDEX].getRow());\n{code}\n\nIt looks like its the this.keys[MEMS_INDEX] is null.\n\nIn testing while doing the 0.19.1 changes I was able to make this NPE. It happened when updateReaders was called after the both memory and the file scanners had been exhausted but the scanner had not yet been closed down fully. The attempt to get from the memory key was failing because it had been set to null after its last next.\n\nNow, we keep the last next key and use that instead of trying to go to memcache (thinking on it, memcache exhaustion can always happen ahead of store file scanner exhaustion).","created":"2009-03-14T18:31:45.694+0000"},{"body":"So head of 0.19 branch includes a resolution for this?","created":"2009-03-19T22:24:59.902+0000"},{"body":"Yeah, should be fixed by 0.19.1 by my reckoning.","created":"2009-03-19T23:02:45.361+0000"},{"body":"Marking as resolved in 0.19.1","created":"2009-03-20T20:33:45.792+0000"}],"conversations":[{"body":"2009-01-01 23:55:41,629 FATAL org.apache.hadoop.hbase.regionserver.MemcacheFlusher: Replay of hlog required. Forcing server shutdown\norg.apache.hadoop.hbase.DroppedSnapshotException: region: content,cff13605e2ea6ce0b221ac864687bf08,1230777531253\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:880)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:773)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushRegion(MemcacheFlusher.java:227)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.run(MemcacheFlusher.java:137)\nCaused by: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HStoreScanner.updateReaders(HStoreScanner.java:322)\n at org.apache.hadoop.hbase.regionserver.HStore.notifyChangedReadersObservers(HStore.java:737)\n at org.apache.hadoop.hbase.regionserver.HStore.updateReaders(HStore.java:725)\n at org.apache.hadoop.hbase.regionserver.HStore.internalFlushCache(HStore.java:694)\n at org.apache.hadoop.hbase.regionserver.HStore.flushCache(HStore.java:630)\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:865)\n ... 3 more\n","from":"reporter","subject":"NPE in HStoreScanner.updateReaders"},{"body":"Right above line 322 in HStoreScanner.java is this comment:\n\\\\\n{code}\n// I think it is safe getting key from mem at this stage -- it shouldn't have\n// been flushed yet\n{code}\n\nSo I think this is no longer true.","from":"developer"},{"body":"Attached patch 1107-1 addresses the symptom. Is it enough?","from":"developer"},{"body":"I think this NPE is a symptom of our making the list of changed readers observers sloppier using CopyOnWrite. I think the scanner has been closed. The first thing done on close is that we remove ourselves from the list of changed readers observers but my guess is that there is an outstanding iteration going on. It still has a reference and invoked the observer code AFTER the clearing of the memcache scanner.\n\nI'm not sure your attached patch safe enough Andrew. My worry is that between your new check on #320, this.keys[MEMS_INDEX] could be cleared before its use on #327 by another thread. What you think?\n\nThe observer code -- updateReaders method -- shouldn't run if close has been invoked methinks. I was going to add in a 'close' flag. Seems like overkill but ain't sure what else to use as 'close' indicator:\n\n{code}\nIndex: src/java/org/apache/hadoop/hbase/regionserver/HStoreScanner.java\n===================================================================\n--- src/java/org/apache/hadoop/hbase/regionserver/HStoreScanner.java (revision 731185)\n+++ src/java/org/apache/hadoop/hbase/regionserver/HStoreScanner.java (working copy)\n@@ -26,6 +26,7 @@\n import java.util.Set;\n import java.util.SortedMap;\n import java.util.TreeMap;\n+import java.util.concurrent.atomic.AtomicBoolean;\n import java.util.concurrent.locks.ReentrantReadWriteLock;\n \n import org.apache.commons.logging.Log;\n@@ -60,6 +61,8 @@\n // Used around transition from no storefile to the first.\n private final ReentrantReadWriteLock lock = new ReentrantReadWriteLock();\n \n+ private final AtomicBoolean closing = new AtomicBoolean(false);\n+ \n /** Create an Scanner with a handle on the memcache and HStore files. */\n @SuppressWarnings(\"unchecked\")\n HStoreScanner(HStore store, byte [][] targetCols, byte [] firstRow,\n@@ -294,6 +297,7 @@\n }\n \n public void close() {\n+ this.closing.set(true);\n this.store.deleteChangedReaderObserver(this);\n doClose();\n }\n@@ -309,6 +313,9 @@\n // Implementation of ChangedReadersObserver\n \n public void updateReaders() throws IOException {\n+ if (this.closing.get()) {\n+ return;\n+ }\n this.lock.writeLock().lock();\n try {\n MapFile.Reader [] readers = this.store.getReaders();\n{code}","from":"developer"},{"body":"Concur that the scanner is closed when the observer gets it.\n\n+1\n\nI'll commit. ","from":"developer"},{"body":"Committed. Passes all local tests. ","from":"developer"},{"body":"Seen again, on 0.19.0 release:\n\n2009-01-25 03:56:39,324 FATAL org.apache.hadoop.hbase.regionserver.MemcacheFlusher: Replay of hlog required. Forcing server shutdown org.apache.hadoop.hbase.DroppedSnapshotException: region: urls,http|img3.megavideo.com|80|a|8|8b3035357a222073d30d09f2c4ae85.jpg,1232408450783\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:896)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:789)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushRegion(MemcacheFlusher.java:227)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushSomeRegions(MemcacheFlusher.java:291)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.reclaimMemcacheMemory(MemcacheFlusher.java:261)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.batchUpdates(HRegionServer.java:1614)\n at sun.reflect.GeneratedMethodAccessor11.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:632)\n at org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:895)\nCaused by: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HStoreScanner.updateReaders(HStoreScanner.java:330)\n at org.apache.hadoop.hbase.regionserver.HStore.notifyChangedReadersObservers(HStore.java:737)\n at org.apache.hadoop.hbase.regionserver.HStore.updateReaders(HStore.java:725)\n at org.apache.hadoop.hbase.regionserver.HStore.internalFlushCache(HStore.java:694)\n at org.apache.hadoop.hbase.regionserver.HStore.flushCache(HStore.java:630)\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:881)\n ... 10 more\n","from":"developer"},{"body":"Thats crazy! I must be digging the hole in the wrong spot?","from":"developer"},{"body":"Got another of these today.\n\n2009-02-09 00:16:55,081 FATAL org.apache.hadoop.hbase.regionserver.MemcacheFlusher: Replay of hlog required. Forcing server shutdown\norg.apache.hadoop.hbase.DroppedSnapshotException: region: content,26ac2c3fd24ac7ba1032e7b68973b9fa,1233995252171\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:896)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:789)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushRegion(MemcacheFlusher.java:227)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushSomeRegions(MemcacheFlusher.java:291)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.reclaimMemcacheMemory(MemcacheFlusher.java:261)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.batchUpdates(HRegionServer.java:1619)\n at sun.reflect.GeneratedMethodAccessor12.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:632)\n at org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:895)\nCaused by: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HStoreScanner.updateReaders(HStoreScanner.java:330)\n at org.apache.hadoop.hbase.regionserver.HStore.notifyChangedReadersObservers(HStore.java:737)\n at org.apache.hadoop.hbase.regionserver.HStore.updateReaders(HStore.java:725)\n at org.apache.hadoop.hbase.regionserver.HStore.internalFlushCache(HStore.java:694)\n at org.apache.hadoop.hbase.regionserver.HStore.flushCache(HStore.java:630)\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:881)\n ... 10 more\n","from":"developer"},{"body":"This problem has been taking down regionservers regularly for me as well. Running hbase 0.19.0 on hadoop 0.19.0 with 4 slaves. It's becoming a major stability issue for us.\n\n2009-03-13 16:30:47,728 FATAL org.apache.hadoop.hbase.regionserver.MemcacheFlusher: Replay of hlog required. Forcing server shutdown\norg.apache.hadoop.hbase.DroppedSnapshotException: region: sessions,^@^@^@^@^@^@^Cd^@^@^A^_�x�q^@.IPHONEc8532c72874b66d1165c2f62ed91ef353eb9,1235185327583\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:896)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:789)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushRegion(MemcacheFlusher.java:227)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.run(MemcacheFlusher.java:137)\nCaused by: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HStoreScanner.updateReaders(HStoreScanner.java:330)\n at org.apache.hadoop.hbase.regionserver.HStore.notifyChangedReadersObservers(HStore.java:737)\n at org.apache.hadoop.hbase.regionserver.HStore.updateReaders(HStore.java:725)\n at org.apache.hadoop.hbase.regionserver.HStore.internalFlushCache(HStore.java:694)\n at org.apache.hadoop.hbase.regionserver.HStore.flushCache(HStore.java:630)\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:881)\n ... 3 more","from":"developer"},{"body":"The code in updateReaders was changed significantly for 0.19.1 (We don't open readers we already had open).\n\nThe NPE above as Andrew has already pointed out comes from here:\n\n{code}\n // I think its safe getting key from mem at this stage -- it shouldn't have\n // been flushed yet\n this.scanners[HSFS_INDEX] = new StoreFileScanner(this.store,\n this.timestamp, this. targetCols, this.keys[MEMS_INDEX].getRow());\n{code}\n\nIt looks like its the this.keys[MEMS_INDEX] is null.\n\nIn testing while doing the 0.19.1 changes I was able to make this NPE. It happened when updateReaders was called after the both memory and the file scanners had been exhausted but the scanner had not yet been closed down fully. The attempt to get from the memory key was failing because it had been set to null after its last next.\n\nNow, we keep the last next key and use that instead of trying to go to memcache (thinking on it, memcache exhaustion can always happen ahead of store file scanner exhaustion).","from":"developer"},{"body":"So head of 0.19 branch includes a resolution for this?","from":"developer"},{"body":"Yeah, should be fixed by 0.19.1 by my reckoning.","from":"developer"},{"body":"Marking as resolved in 0.19.1","from":"developer"}],"created":"2009-01-02T01:37:45.000+0000","description":"2009-01-01 23:55:41,629 FATAL org.apache.hadoop.hbase.regionserver.MemcacheFlusher: Replay of hlog required. Forcing server shutdown\norg.apache.hadoop.hbase.DroppedSnapshotException: region: content,cff13605e2ea6ce0b221ac864687bf08,1230777531253\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:880)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:773)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.flushRegion(MemcacheFlusher.java:227)\n at org.apache.hadoop.hbase.regionserver.MemcacheFlusher.run(MemcacheFlusher.java:137)\nCaused by: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HStoreScanner.updateReaders(HStoreScanner.java:322)\n at org.apache.hadoop.hbase.regionserver.HStore.notifyChangedReadersObservers(HStore.java:737)\n at org.apache.hadoop.hbase.regionserver.HStore.updateReaders(HStore.java:725)\n at org.apache.hadoop.hbase.regionserver.HStore.internalFlushCache(HStore.java:694)\n at org.apache.hadoop.hbase.regionserver.HStore.flushCache(HStore.java:630)\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:865)\n ... 3 more\n","issue_id":"12411558","key":"HBASE-1107","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2009-03-20T20:33:45.000+0000","role":"fixed_distractor","summary":"NPE in HStoreScanner.updateReaders"} {"case_id":"12716122","cluster":"DISTRACTOR-HBASE-11236","comments":[{"body":"Don't see it any more. Close it.","created":"2014-05-30T17:37:21.394+0000"},{"body":"I've seen this on 1.2 a decent ammount","created":"2015-08-31T22:52:50.632+0000"},{"body":"I also observed same log in product environment and regions were not opened.\n{noformat}\n2015-08-31 19:04:25,448 | WARN | PriorityRpcServer.handler=13,queue=1,port=21300 | RegionServer *.*.*.*,21302,1441017551749 indicates a last flushed sequence id (53128) that is less than the previous last flushed sequence id (53131) for region hbase:meta,,1 Ignoring. | org.apache.hadoop.hbase.master.ServerManager.updateLastFlushedSequenceIds(ServerManager.java:299)\n{noformat}","created":"2015-09-01T01:43:36.522+0000"},{"body":"I think this could happen without HBASE-13811 where [~stack] fix a issue that we may report a wrong last flushed sequence id if a flush is aborted.\n\n[~eclark] Any more informations? What's happened to the region before this log(reassignment or flush?)? Maybe there are other issues.\n\nThanks.","created":"2015-09-01T02:51:56.275+0000"},{"body":"[~Apache9] Looks like root cause is different, we have HBASE-13811 in our version. I am also trying to figure out the relevant info from logs.","created":"2015-09-01T09:15:14.143+0000"},{"body":"[~pankaj2461], could you provide more information on this? I just wonder whether you suffer data loss on this? ","created":"2016-09-17T18:33:39.618+0000"},{"body":"Sorry [~syuanjiang], couldn't find any useful information that time. Later I didn't see this log again.","created":"2016-09-19T03:53:06.788+0000"},{"body":"ignore flushed sequence id can cause data loss, please ref to https://issues.apache.org/jira/browse/HBASE-16649?focusedCommentId=15502490&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15502490","created":"2016-09-20T01:29:42.240+0000"},{"body":"we are having the same issue with hbase-1.2.0+cdh5.8.0+160-1:\n\n2016-09-22 02:34:21,805 WARN [B.defaultRpcServer.handler=43,queue=3,port=60000] master.ServerManager: RegionServer dr2,60020,1470414843847 indicates a last flushed sequence id (6178551) that is less than the previous last flushed sequence id (9091225) for region user_entry_tags,?,1472495644435.f5cbb00e42bcfd353e2041818b69fd6f. Ignoring.\n","created":"2016-09-22T09:43:32.558+0000"},{"body":"I am on version 1.1.2, and a data loss bug landed me here. \n\nWhat happened to me is that a data loss has been identified on region cea9c145a49489e2ebbc683c9bc0b545. Then grepping on the regionId, I found nothing but below:\n\nhbase-hbase-master-hqhd02nm01.pclc0.merkle.local.log.20:2017-06-30 05:47:20,237 WARN [B.priority.fifo.QRpcServer.handler=2,queue=0,port=16000] master.ServerManager: RegionServer hqhd02dt031.pclc0.merkle.local,16020,1498275002632 indicates a last flushed sequence id (20717979) that is less than the previous last flushed sequence id (24918589) for region CR_CDI_CHHEQPROD_HASH_RECORD,f8886,1492710130131.cea9c145a49489e2ebbc683c9bc0b545. Ignoring.\nhbase-hbase-master-hqhd02nm01.pclc0.merkle.local.log.20:2017-06-30 05:47:23,255 WARN [B.priority.fifo.QRpcServer.handler=19,queue=1,port=16000] master.ServerManager: RegionServer hqhd02dt031.pclc0.merkle.local,16020,1498275002632 indicates a last flushed sequence id (20717979) that is less than the previous last flushed sequence id (24918589) for region CR_CDI_CHHEQPROD_HASH_RECORD,f8886,1492710130131.cea9c145a49489e2ebbc683c9bc0b545. Ignoring.\nhbase-hbase-master-hqhd02nm01.pclc0.merkle.local.log.20:2017-06-30 05:47:26,273 WARN [B.priority.fifo.QRpcServer.handler=0,queue=0,port=16000] master.ServerManager: RegionServer hqhd02dt031.pclc0.merkle.local,16020,1498275002632 indicates a last flushed sequence id (20717979) that is less than the previous last flushed sequence id (24918589) for region CR_CDI_CHHEQPROD_HASH_RECORD,f8886,1492710130131.cea9c145a49489e2ebbc683c9bc0b545. Ignoring.\nhbase-hbase-master-hqhd02nm01.pclc0.merkle.local.log.20:2017-06-30 05:47:29,290 WARN [B.priority.fifo.QRpcServer.handler=16,queue=0,port=16000] master.ServerManager: RegionServer hqhd02dt031.pclc0.merkle.local,16020,1498275002632 indicates a last flushed sequence id (20717979) that is less than the previous last flushed sequence id (24918589) for region CR_CDI_CHHEQPROD_HASH_RECORD,f8886,1492710130131.cea9c145a49489e2ebbc683c9bc0b545. Ignoring.\nh","created":"2017-06-30T16:14:46.371+0000"},{"body":"HBASE-16721 may address this problem. According to the above comments, this issue had happened in the 1.1.2, 1.2, and hbase-1.2.0+cdh5.8.0.\nHBASE-16721 had be merged into 1.1.8+ and 1.2.4+, and [cdh5.9.2|https://archive.cloudera.com/cdh5/cdh/5/hbase-1.2.0-cdh5.9.2.CHANGES.txt]. We should close this jira If there is no more victims.","created":"2017-07-17T16:00:56.271+0000"},{"body":"Resolving as [~chia7712] suggests after research. Assigning issue to Chia-Ping since he did the work.","created":"2018-03-02T04:18:36.390+0000"}],"conversations":[{"body":"I got lots of error messages like this:\n\n{quote}\n2014-05-22 08:58:59,793 DEBUG [RpcServer.handler=1,port=20020] master.ServerManager: RegionServer a2428.halxg.cloudera.com,20020,1400742071109 indicates a last flushed sequence id (numberOfStores=9, numberOfStorefiles=2, storefileUncompressedSizeMB=517, storefileSizeMB=517, compressionRatio=1.0000, memstoreSizeMB=0, storefileIndexSizeMB=0, readRequestsCount=0, writeRequestsCount=0, rootIndexSizeKB=34, totalStaticIndexSizeKB=381, totalStaticBloomSizeKB=0, totalCompactingKVs=0, currentCompactedKVs=0, compactionProgressPct=NaN) that is less than the previous last flushed sequence id (605446) for region IntegrationTestBigLinkedList, �A��*t�^FU�2��0,1400740489477.a44d3e309b5a7e29355f6faa0d3a4095. Ignoring.\n{quote}\n\nRegionLoad.toString doesn't print out the last flushed sequence id passed in. Why is it less than the previous one?","from":"reporter","subject":"Last flushed sequence id is ignored by ServerManager"},{"body":"Don't see it any more. Close it.","from":"developer"},{"body":"I've seen this on 1.2 a decent ammount","from":"developer"},{"body":"I also observed same log in product environment and regions were not opened.\n{noformat}\n2015-08-31 19:04:25,448 | WARN | PriorityRpcServer.handler=13,queue=1,port=21300 | RegionServer *.*.*.*,21302,1441017551749 indicates a last flushed sequence id (53128) that is less than the previous last flushed sequence id (53131) for region hbase:meta,,1 Ignoring. | org.apache.hadoop.hbase.master.ServerManager.updateLastFlushedSequenceIds(ServerManager.java:299)\n{noformat}","from":"developer"},{"body":"I think this could happen without HBASE-13811 where [~stack] fix a issue that we may report a wrong last flushed sequence id if a flush is aborted.\n\n[~eclark] Any more informations? What's happened to the region before this log(reassignment or flush?)? Maybe there are other issues.\n\nThanks.","from":"developer"},{"body":"[~Apache9] Looks like root cause is different, we have HBASE-13811 in our version. I am also trying to figure out the relevant info from logs.","from":"developer"},{"body":"[~pankaj2461], could you provide more information on this? I just wonder whether you suffer data loss on this? ","from":"developer"},{"body":"Sorry [~syuanjiang], couldn't find any useful information that time. Later I didn't see this log again.","from":"developer"},{"body":"ignore flushed sequence id can cause data loss, please ref to https://issues.apache.org/jira/browse/HBASE-16649?focusedCommentId=15502490&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15502490","from":"developer"},{"body":"we are having the same issue with hbase-1.2.0+cdh5.8.0+160-1:\n\n2016-09-22 02:34:21,805 WARN [B.defaultRpcServer.handler=43,queue=3,port=60000] master.ServerManager: RegionServer dr2,60020,1470414843847 indicates a last flushed sequence id (6178551) that is less than the previous last flushed sequence id (9091225) for region user_entry_tags,?,1472495644435.f5cbb00e42bcfd353e2041818b69fd6f. Ignoring.\n","from":"developer"},{"body":"I am on version 1.1.2, and a data loss bug landed me here. \n\nWhat happened to me is that a data loss has been identified on region cea9c145a49489e2ebbc683c9bc0b545. Then grepping on the regionId, I found nothing but below:\n\nhbase-hbase-master-hqhd02nm01.pclc0.merkle.local.log.20:2017-06-30 05:47:20,237 WARN [B.priority.fifo.QRpcServer.handler=2,queue=0,port=16000] master.ServerManager: RegionServer hqhd02dt031.pclc0.merkle.local,16020,1498275002632 indicates a last flushed sequence id (20717979) that is less than the previous last flushed sequence id (24918589) for region CR_CDI_CHHEQPROD_HASH_RECORD,f8886,1492710130131.cea9c145a49489e2ebbc683c9bc0b545. Ignoring.\nhbase-hbase-master-hqhd02nm01.pclc0.merkle.local.log.20:2017-06-30 05:47:23,255 WARN [B.priority.fifo.QRpcServer.handler=19,queue=1,port=16000] master.ServerManager: RegionServer hqhd02dt031.pclc0.merkle.local,16020,1498275002632 indicates a last flushed sequence id (20717979) that is less than the previous last flushed sequence id (24918589) for region CR_CDI_CHHEQPROD_HASH_RECORD,f8886,1492710130131.cea9c145a49489e2ebbc683c9bc0b545. Ignoring.\nhbase-hbase-master-hqhd02nm01.pclc0.merkle.local.log.20:2017-06-30 05:47:26,273 WARN [B.priority.fifo.QRpcServer.handler=0,queue=0,port=16000] master.ServerManager: RegionServer hqhd02dt031.pclc0.merkle.local,16020,1498275002632 indicates a last flushed sequence id (20717979) that is less than the previous last flushed sequence id (24918589) for region CR_CDI_CHHEQPROD_HASH_RECORD,f8886,1492710130131.cea9c145a49489e2ebbc683c9bc0b545. Ignoring.\nhbase-hbase-master-hqhd02nm01.pclc0.merkle.local.log.20:2017-06-30 05:47:29,290 WARN [B.priority.fifo.QRpcServer.handler=16,queue=0,port=16000] master.ServerManager: RegionServer hqhd02dt031.pclc0.merkle.local,16020,1498275002632 indicates a last flushed sequence id (20717979) that is less than the previous last flushed sequence id (24918589) for region CR_CDI_CHHEQPROD_HASH_RECORD,f8886,1492710130131.cea9c145a49489e2ebbc683c9bc0b545. Ignoring.\nh","from":"developer"},{"body":"HBASE-16721 may address this problem. According to the above comments, this issue had happened in the 1.1.2, 1.2, and hbase-1.2.0+cdh5.8.0.\nHBASE-16721 had be merged into 1.1.8+ and 1.2.4+, and [cdh5.9.2|https://archive.cloudera.com/cdh5/cdh/5/hbase-1.2.0-cdh5.9.2.CHANGES.txt]. We should close this jira If there is no more victims.","from":"developer"},{"body":"Resolving as [~chia7712] suggests after research. Assigning issue to Chia-Ping since he did the work.","from":"developer"}],"created":"2014-05-22T16:31:37.000+0000","description":"I got lots of error messages like this:\n\n{quote}\n2014-05-22 08:58:59,793 DEBUG [RpcServer.handler=1,port=20020] master.ServerManager: RegionServer a2428.halxg.cloudera.com,20020,1400742071109 indicates a last flushed sequence id (numberOfStores=9, numberOfStorefiles=2, storefileUncompressedSizeMB=517, storefileSizeMB=517, compressionRatio=1.0000, memstoreSizeMB=0, storefileIndexSizeMB=0, readRequestsCount=0, writeRequestsCount=0, rootIndexSizeKB=34, totalStaticIndexSizeKB=381, totalStaticBloomSizeKB=0, totalCompactingKVs=0, currentCompactedKVs=0, compactionProgressPct=NaN) that is less than the previous last flushed sequence id (605446) for region IntegrationTestBigLinkedList, �A��*t�^FU�2��0,1400740489477.a44d3e309b5a7e29355f6faa0d3a4095. Ignoring.\n{quote}\n\nRegionLoad.toString doesn't print out the last flushed sequence id passed in. Why is it less than the previous one?","issue_id":"12716122","key":"HBASE-11236","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2018-03-02T04:18:36.000+0000","role":"fixed_distractor","summary":"Last flushed sequence id is ignored by ServerManager"} {"case_id":"12723282","cluster":"DISTRACTOR-HBASE-11405","comments":[{"body":"Related question: Should be disable parallel hbcks permanently? A quick look at the code suggests this might cause inconsistencies as well. ","created":"2014-06-24T06:14:36.645+0000"},{"body":"bq.Related question: Should be disable parallel hbcks permanently?\n+1 for that.","created":"2014-06-24T07:22:36.870+0000"},{"body":"I agree. hbck in parallel might create some mess... We might want to block that.","created":"2014-06-24T18:44:41.726+0000"},{"body":"+1, would be a better fix to test in HBCK startup if another is running and error out if so","created":"2014-06-24T23:56:32.400+0000"},{"body":"- Thanks for the +1s. I was thinking of writing a patch which creates an ephemeral znode /hbase/hbck-in-progress while starting hbck and deletes it while exiting. If a new hbck instance sees already existing znode, it aborts. However [~mbertozzi] says it is not a good approach to directly deal with ZK connections as the internal implementations might change breaking this feature and also since we are trying to remove ZK dependency (via other JIRAs) on HBase, implementations like this might make that transition difficult. So, In future when we have a lock() based system, each hbck should probably request a lock from master and the other hbcks should wait on it.\n\n- Feel free to suggest any other approaches you have in mind.","created":"2014-06-25T18:39:11.314+0000"},{"body":"Use a 0 length file in HDFS as the lock. Present an interactive option for the user to ignore the lock file and proceed.","created":"2014-06-25T22:35:52.066+0000"},{"body":"This patch creates a file in / and maintains a lease on it till hbck exits. Parallel instances of hbck check if such a file exists and bail out if necessary. Added a shutdownhook() to this class to clean up the lock file incase user kills hbck with a SIGTERM. I make sure that the cleanup is done only once using a flag.","created":"2014-07-02T13:52:42.550+0000"},{"body":"{code}\n+ HBCK_LOCK_PATH = new Path(new Path(FSUtils.getRootDir(getConf()), \n+ HConstants.HBASE_TEMP_DIRECTORY), BALANCER_LOCK_FILE);\n{code}\nThe lock file is used by hbck. Should the name of lock file reflect this ?\n{code}\n+ // Make sure tmp dir exists, if not create it\n+ fs.mkdirs(new Path(FSUtils.getRootDir(getConf()), HConstants.HBASE_TEMP_DIRECTORY));\n{code}\nCheck the return value from mkdirs().\n{code}\n+ setRetCode(-1);\n+ LOG.info(\"Another balancer is running - Exiting this instance\");\n{code}\nlog level should be error above.\n\n","created":"2014-07-02T16:03:55.098+0000"},{"body":"bq. The lock file is used by hbck. Should the name of lock file reflect this ?\n\nAgreed. Typo, will make this change.\n\nbq. Check the return value from mkdirs().\n\nAny particular reason why we should check that? It silently fails if the directory already exists. May be I can put a check like if(!fs.exists(dir)) {fs.mkdirs(dir);} ?\n\nbq. log level should be error above.\n\nAgreed, will make this change too.\n","created":"2014-07-02T16:30:02.509+0000"},{"body":"If the return value from mkdirs() is false, directory isn't created successfully.\nSee javadoc of mkdirs():\n{code}\n * @return true if the directory creation succeeds; false otherwise\n * @throws IOException\n */\n public static boolean mkdirs(FileSystem fs, Path dir, FsPermission permission)\n{code}","created":"2014-07-02T16:35:55.378+0000"},{"body":"If directory creation had problem, FSUtils.create() is likely to fail.\nI think no extra action is needed for the mkdirs() call.","created":"2014-07-02T17:41:41.132+0000"},{"body":"v1 addresses comments from Ted. ","created":"2014-07-03T04:01:52.354+0000"},{"body":"{noformat}\nbusbey2-MBA:hbase busbey$ git status\nOn branch master\nYour branch is up-to-date with 'origin/master'.\n\nnothing to commit, working directory clean\nbusbey2-MBA:hbase busbey$ git apply --check ~/Downloads/HBASE-11405-trunk.patch.1 \nerror: patch failed: hbase-server/src/main/java/org/apache/hadoop/hbase/util/HBaseFsck.java:105\nerror: hbase-server/src/main/java/org/apache/hadoop/hbase/util/HBaseFsck.java: patch does not apply\nerror: patch failed: hbase-server/src/test/java/org/apache/hadoop/hbase/util/TestHBaseFsck.java:36\nerror: hbase-server/src/test/java/org/apache/hadoop/hbase/util/TestHBaseFsck.java: patch does not apply\n{noformat}\n\nPatch no longer applies to master. [~bharathv] can you rebase?\n\nCould you then also upload to ReviewBoard so it's easier to give review feedback?","created":"2014-08-12T15:29:50.206+0000"},{"body":"Rebased the patch to trunk.","created":"2014-08-22T09:16:15.746+0000"},{"body":"The unit test or the structure of the code likely needs to be modified -- I get this when I try to run the test and this is likely due to the shutdown hook.\n\n\n\n{code}\njon@swoop:~/proj/hbase-trunk$ mvn clean test -Dtest=TestHBaseFsck\n....\n-------------------------------------------------------\n T E S T S\n-------------------------------------------------------\nRunning org.apache.hadoop.hbase.util.TestHBaseFsck\n\nResults :\n\nTests run: 0, Failures: 0, Errors: 0, Skipped: 0\n...\n[INFO] ------------------------------------------------------------------------\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-surefire-plugin:2.17:test (default-test) on project hbase-server: ExecutionException: java.lang.RuntimeException: The forked VM terminated without properly saying goodbye. VM crash or System.exit called?\n[ERROR] Command was /bin/sh -c cd /home/jon/proj/hbase-trunk/hbase-server && /opt/jdk1.7.0_25/jre/bin/java -enableassertions -XX:MaxDirectMemorySize=1G -Xmx1900m -XX:MaxPermSize=256m -Djava.security.egd=file:/dev/./urandom -Djava.net.preferIPv4Stack=true -Djava.awt.headless=true -jar /home/jon/proj/hbase-trunk/hbase-server/target/surefire/surefirebooter6944807203322450992.jar /home/jon/proj/hbase-trunk/hbase-server/target/surefire/surefire1768607263272388944tmp /home/jon/proj/hbase-trunk/hbase-server/target/surefire/surefire_084228475421972179tmp\n[ERROR] -> [Help 1]\n...\n{code}","created":"2014-09-11T19:02:45.364+0000"},{"body":"attached updated patch with minor imports fixes.","created":"2014-09-11T19:04:03.116+0000"},{"body":"\nLet's tell the user where the lock file is in case it was left by an hbck instance terminated with kill -9 that left the tmp file behind. Also say to delete if you are sure the other hbck is dead.\n{code}\n+ // Check if another instance of balancer is running\n+ hbckOutFd = checkAndMarkRunningHbck(); \n+ if (hbckOutFd == null) {\n+ LOG.error(\"Another balancer is running - Exiting this instance\");\n+ setRetCode(-1);\n+ Runtime.getRuntime().exit(-1);\n+ }\n{code}","created":"2014-09-11T19:06:17.126+0000"},{"body":"[~bharathv] are you still working on this? would you mind if I took it over?","created":"2014-09-11T20:49:40.133+0000"},{"body":"I ran TestHBaseFsck#testParallelHbck and was able to reproduce the error Jon mentioned above.\nI think it was caused by the following statement:\n{code}\n+ Runtime.getRuntime().exit(-1);\n{code}\n","created":"2014-09-11T20:56:38.238+0000"},{"body":"Patch v3 removes the Runtime.getRuntime().exit() call.\nIOE is thrown instead.\n\nTest case is adjusted accordingly.","created":"2014-09-11T20:58:27.967+0000"},{"body":"Thanks [~jon@cloudera.com] and [~tedyu@apache.org]. The tests fail because we abruptly kill the jvm. Ted's version fixes that. I just added a few log lines as per Jon's suggestion. [~busbey] Mind taking a look at this version? https://reviews.apache.org/r/25587","created":"2014-09-12T17:52:33.147+0000"},{"body":"[~bharathv]:\nMind attaching patch here ?","created":"2014-09-12T21:46:14.590+0000"},{"body":"[~tedyu@apache.org] Sorry, attached it now. Made some minor changes [Comments on review board]\n[~busbey] Since we are throwing an IOEx , doFsck won't return an HBaseFsck object and hence it doesn't matter what return value we set. Since we are doing\n\n{code:title=|borderStyle=solid}\n if (e.getMessage().contains(\"Duplicate hbck\")) {\n fail = false;\n }\n }\n // If we reach here, then an exception was caught\n if (fail) fail();\n return null;\n }\n{code}\n\nIt actually confirms that the issue is due to multiple hbcks. This approach looks good to me. What do you think?","created":"2014-09-15T06:52:09.067+0000"},{"body":"I believe the issue with failing tests is probably because surefire is running tests in parallel (which is stopping parallel hbcks since they are all using the same cluster object). If I run the failed tests individually, they pass. There is a mvn profile \"nonParallelTests\" but it doesn't work for me either. Any ideas?","created":"2014-09-15T15:27:53.268+0000"},{"body":"v6 summary\n\n- Moved unlockHbck() from exec() to onlineHbck() method, since some of the tests directly call onlineHbck() rather than going via exec. This makes sure we cleanup the lock in tests\n- Small bug in compareAndSet(), misread the doc, fixed it, also setting it to (false) inside the block to make sure tests using the same HBck object doesn't fail.","created":"2014-09-15T15:29:25.628+0000"},{"body":"thanks bharath, and thanks sean and ted for reviews. committed to branch-1 and master.","created":"2014-09-17T19:54:29.994+0000"},{"body":"Sorry, reverted and reopened since sean had more comments on review board.","created":"2014-09-17T20:00:45.785+0000"},{"body":"v7\n- Addressed Sean's comments on the RB.\n- Few minor changes","created":"2014-09-18T21:31:36.802+0000"},{"body":"[~ busbey]:\nHave all of your comments been addressed ?","created":"2014-09-18T23:04:43.354+0000"},{"body":"Yep. I'm +1 on the latest RB version, which I believe is the same as v7.","created":"2014-09-18T23:11:07.494+0000"},{"body":"Integrated to branch-1 and master.\n\nThanks for the patch, bharath.","created":"2014-09-19T01:02:22.722+0000"},{"body":"The pull back to branch-1 broke compilation. Looks like hte imports are different in the test class there.","created":"2014-09-19T02:00:43.496+0000"},{"body":"I've reverted this from branch-1 for the time being. \n\nReminder, please test before committing to branches!","created":"2014-09-19T02:09:43.417+0000"},{"body":"Will pay attention next time.","created":"2014-09-19T02:14:20.223+0000"},{"body":"With rebased patch for branch-1:\n{code}\nRunning org.apache.hadoop.hbase.util.TestHBaseFsck\nTests run: 44, Failures: 0, Errors: 0, Skipped: 1, Time elapsed: 334.241 sec\n{code}","created":"2014-09-19T16:34:41.078+0000"},{"body":"Integrated to branch-1 again.","created":"2014-09-19T16:41:00.066+0000"},{"body":"Is this not an issue in 0.94 and 0.98?","created":"2014-09-19T17:34:10.510+0000"},{"body":"definitely present in both 0.94 and 0.98. Do you prefer a backport or known issues text? I can pull together either.","created":"2014-09-19T17:37:28.539+0000"},{"body":"11405-1.0.txt applies cleanly on 0.98\n[~apurtell]:\nDo you want this ?","created":"2014-09-19T17:46:26.360+0000"},{"body":"We can attach backport patches here since neither 0.99.1 nor 2.0.0 are released.","created":"2014-09-19T18:09:54.377+0000"},{"body":"+1 for 0.98\nWhy was this branch left out of consideration initially (along with 0.94)? ","created":"2014-09-20T00:08:38.464+0000"},{"body":"Reopened and updated fix versions","created":"2014-09-20T00:09:25.061+0000"},{"body":"Integrated to 0.98 as well.\n\nI can prepare patch for 0.94 if Sean is busy.","created":"2014-09-20T02:53:29.690+0000"},{"body":"I'm chasing down a failure on the hadoop-1 profile and working on 0.94.\nShould be later tonight.\n\nIf someone else can get it done earlier, that'd be great.\n\n-- \nSean\n\n","created":"2014-09-20T03:12:40.693+0000"},{"body":"You mean this one ?\n{code}\ntestParallelHbck(org.apache.hadoop.hbase.util.TestHBaseFsck) Time elapsed: 60.152 sec <<< FAILURE!\njava.lang.AssertionError\n at org.apache.hadoop.hbase.util.TestHBaseFsck.testParallelHbck(TestHBaseFsck.java:552)\n{code}\nSigh - let me revert from 0.98 for now.","created":"2014-09-20T03:16:04.203+0000"},{"body":"okay, the test failure is caused because on [hadoop 1 dfs client transparently retries the file create for up to 5 minutes|https://github.com/apache/hadoop/blob/release-1.2.1/src/hdfs/org/apache/hadoop/hdfs/DFSClient.java#L163].\n\nLooking at the logs of the hadoop 1 run, I can see the second hbck instance fail to create the file. Then later after the first one cleans up, the second picks up again and keeps going. Thus both succeed and the assert fails. Same failure on 0.94.\n\nI'd rather not insert a 5 minute sleep in the hbck to ensure the second instance fails. Anyone have some idea of a different way we can test for the branches that support hadoop 1?","created":"2014-09-20T11:25:57.937+0000"},{"body":"Just a shot in the dark: Use mockito to swap in a different retry policy for creates? ","created":"2014-09-20T17:07:31.883+0000"},{"body":"Here's an amended version specific to each of 0.98 and 0.94 that works under profiles for hadoop 1.0, 1.1, and 2.x. Due to the amount of change, I added an Amending-Author to the commit message.\n\nI've run through {{-Dtest=TestHBaseFsck}} for the above profiles and {{-PrunSmallTests}} for whatever the default profile is on each branch. NB running TestHBaseFsck on 0.94 required {{-PlocalTests}}.\n\nFixing this involved exposing a previously private method for HBaseTestingUtility under 0.94. Since nothing else needed it in 0.98 I marked it deprecated with a note that it would disappear with Hadoop 1.x support. Since the method isn't in branch-1 or later, that should work out well.\n\nOne worry: some of the TestHBaseFsck tests get a bit flaky on 0.94 with this change on my laptop (which admittedly is underpowered). However, even leaving in a single retry period causes the parallel test to fail because the first invocation finished in < 1 sec.","created":"2014-09-24T22:09:51.809+0000"},{"body":"The 0.98 patch lgtm","created":"2014-09-24T23:05:14.312+0000"},{"body":"Rebased patch for 0.98","created":"2014-09-25T16:17:59.937+0000"},{"body":"Integrated to 0.98\n\nRunning test for 0.94 now.","created":"2014-09-25T16:28:51.338+0000"},{"body":"Integrated to 0.94 as well.\n\nThanks Sean and bharath","created":"2014-09-25T16:48:42.207+0000"}],"conversations":[{"body":"This is because of the following piece of code in hbck\n\n{code:borderStyle=solid}\n boolean oldBalancer = admin.setBalancerRunning(false, true);\n try {\n onlineConsistencyRepair();\n }\n finally {\n admin.setBalancerRunning(oldBalancer, false);\n }\n{code}\n\nNewer invocations set oldBalancer to false as it was disabled by previous invocations and this disables balancer permanently unless its manually turned on by the user. Easy to reproduce, just run hbck 100 times in a loop in 2 different sessions and you can see that balancer is set to false in the HMaster logs.","from":"reporter","subject":"Multiple invocations of hbck in parallel disables balancer permanently "},{"body":"Related question: Should be disable parallel hbcks permanently? A quick look at the code suggests this might cause inconsistencies as well. ","from":"developer"},{"body":"bq.Related question: Should be disable parallel hbcks permanently?\n+1 for that.","from":"developer"},{"body":"I agree. hbck in parallel might create some mess... We might want to block that.","from":"developer"},{"body":"+1, would be a better fix to test in HBCK startup if another is running and error out if so","from":"developer"},{"body":"- Thanks for the +1s. I was thinking of writing a patch which creates an ephemeral znode /hbase/hbck-in-progress while starting hbck and deletes it while exiting. If a new hbck instance sees already existing znode, it aborts. However [~mbertozzi] says it is not a good approach to directly deal with ZK connections as the internal implementations might change breaking this feature and also since we are trying to remove ZK dependency (via other JIRAs) on HBase, implementations like this might make that transition difficult. So, In future when we have a lock() based system, each hbck should probably request a lock from master and the other hbcks should wait on it.\n\n- Feel free to suggest any other approaches you have in mind.","from":"developer"},{"body":"Use a 0 length file in HDFS as the lock. Present an interactive option for the user to ignore the lock file and proceed.","from":"developer"},{"body":"This patch creates a file in / and maintains a lease on it till hbck exits. Parallel instances of hbck check if such a file exists and bail out if necessary. Added a shutdownhook() to this class to clean up the lock file incase user kills hbck with a SIGTERM. I make sure that the cleanup is done only once using a flag.","from":"developer"},{"body":"{code}\n+ HBCK_LOCK_PATH = new Path(new Path(FSUtils.getRootDir(getConf()), \n+ HConstants.HBASE_TEMP_DIRECTORY), BALANCER_LOCK_FILE);\n{code}\nThe lock file is used by hbck. Should the name of lock file reflect this ?\n{code}\n+ // Make sure tmp dir exists, if not create it\n+ fs.mkdirs(new Path(FSUtils.getRootDir(getConf()), HConstants.HBASE_TEMP_DIRECTORY));\n{code}\nCheck the return value from mkdirs().\n{code}\n+ setRetCode(-1);\n+ LOG.info(\"Another balancer is running - Exiting this instance\");\n{code}\nlog level should be error above.\n\n","from":"developer"},{"body":"bq. The lock file is used by hbck. Should the name of lock file reflect this ?\n\nAgreed. Typo, will make this change.\n\nbq. Check the return value from mkdirs().\n\nAny particular reason why we should check that? It silently fails if the directory already exists. May be I can put a check like if(!fs.exists(dir)) {fs.mkdirs(dir);} ?\n\nbq. log level should be error above.\n\nAgreed, will make this change too.\n","from":"developer"},{"body":"If the return value from mkdirs() is false, directory isn't created successfully.\nSee javadoc of mkdirs():\n{code}\n * @return true if the directory creation succeeds; false otherwise\n * @throws IOException\n */\n public static boolean mkdirs(FileSystem fs, Path dir, FsPermission permission)\n{code}","from":"developer"},{"body":"If directory creation had problem, FSUtils.create() is likely to fail.\nI think no extra action is needed for the mkdirs() call.","from":"developer"},{"body":"v1 addresses comments from Ted. ","from":"developer"},{"body":"{noformat}\nbusbey2-MBA:hbase busbey$ git status\nOn branch master\nYour branch is up-to-date with 'origin/master'.\n\nnothing to commit, working directory clean\nbusbey2-MBA:hbase busbey$ git apply --check ~/Downloads/HBASE-11405-trunk.patch.1 \nerror: patch failed: hbase-server/src/main/java/org/apache/hadoop/hbase/util/HBaseFsck.java:105\nerror: hbase-server/src/main/java/org/apache/hadoop/hbase/util/HBaseFsck.java: patch does not apply\nerror: patch failed: hbase-server/src/test/java/org/apache/hadoop/hbase/util/TestHBaseFsck.java:36\nerror: hbase-server/src/test/java/org/apache/hadoop/hbase/util/TestHBaseFsck.java: patch does not apply\n{noformat}\n\nPatch no longer applies to master. [~bharathv] can you rebase?\n\nCould you then also upload to ReviewBoard so it's easier to give review feedback?","from":"developer"},{"body":"Rebased the patch to trunk.","from":"developer"},{"body":"The unit test or the structure of the code likely needs to be modified -- I get this when I try to run the test and this is likely due to the shutdown hook.\n\n\n\n{code}\njon@swoop:~/proj/hbase-trunk$ mvn clean test -Dtest=TestHBaseFsck\n....\n-------------------------------------------------------\n T E S T S\n-------------------------------------------------------\nRunning org.apache.hadoop.hbase.util.TestHBaseFsck\n\nResults :\n\nTests run: 0, Failures: 0, Errors: 0, Skipped: 0\n...\n[INFO] ------------------------------------------------------------------------\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-surefire-plugin:2.17:test (default-test) on project hbase-server: ExecutionException: java.lang.RuntimeException: The forked VM terminated without properly saying goodbye. VM crash or System.exit called?\n[ERROR] Command was /bin/sh -c cd /home/jon/proj/hbase-trunk/hbase-server && /opt/jdk1.7.0_25/jre/bin/java -enableassertions -XX:MaxDirectMemorySize=1G -Xmx1900m -XX:MaxPermSize=256m -Djava.security.egd=file:/dev/./urandom -Djava.net.preferIPv4Stack=true -Djava.awt.headless=true -jar /home/jon/proj/hbase-trunk/hbase-server/target/surefire/surefirebooter6944807203322450992.jar /home/jon/proj/hbase-trunk/hbase-server/target/surefire/surefire1768607263272388944tmp /home/jon/proj/hbase-trunk/hbase-server/target/surefire/surefire_084228475421972179tmp\n[ERROR] -> [Help 1]\n...\n{code}","from":"developer"},{"body":"attached updated patch with minor imports fixes.","from":"developer"},{"body":"\nLet's tell the user where the lock file is in case it was left by an hbck instance terminated with kill -9 that left the tmp file behind. Also say to delete if you are sure the other hbck is dead.\n{code}\n+ // Check if another instance of balancer is running\n+ hbckOutFd = checkAndMarkRunningHbck(); \n+ if (hbckOutFd == null) {\n+ LOG.error(\"Another balancer is running - Exiting this instance\");\n+ setRetCode(-1);\n+ Runtime.getRuntime().exit(-1);\n+ }\n{code}","from":"developer"},{"body":"[~bharathv] are you still working on this? would you mind if I took it over?","from":"developer"},{"body":"I ran TestHBaseFsck#testParallelHbck and was able to reproduce the error Jon mentioned above.\nI think it was caused by the following statement:\n{code}\n+ Runtime.getRuntime().exit(-1);\n{code}\n","from":"developer"},{"body":"Patch v3 removes the Runtime.getRuntime().exit() call.\nIOE is thrown instead.\n\nTest case is adjusted accordingly.","from":"developer"},{"body":"Thanks [~jon@cloudera.com] and [~tedyu@apache.org]. The tests fail because we abruptly kill the jvm. Ted's version fixes that. I just added a few log lines as per Jon's suggestion. [~busbey] Mind taking a look at this version? https://reviews.apache.org/r/25587","from":"developer"},{"body":"[~bharathv]:\nMind attaching patch here ?","from":"developer"},{"body":"[~tedyu@apache.org] Sorry, attached it now. Made some minor changes [Comments on review board]\n[~busbey] Since we are throwing an IOEx , doFsck won't return an HBaseFsck object and hence it doesn't matter what return value we set. Since we are doing\n\n{code:title=|borderStyle=solid}\n if (e.getMessage().contains(\"Duplicate hbck\")) {\n fail = false;\n }\n }\n // If we reach here, then an exception was caught\n if (fail) fail();\n return null;\n }\n{code}\n\nIt actually confirms that the issue is due to multiple hbcks. This approach looks good to me. What do you think?","from":"developer"},{"body":"I believe the issue with failing tests is probably because surefire is running tests in parallel (which is stopping parallel hbcks since they are all using the same cluster object). If I run the failed tests individually, they pass. There is a mvn profile \"nonParallelTests\" but it doesn't work for me either. Any ideas?","from":"developer"},{"body":"v6 summary\n\n- Moved unlockHbck() from exec() to onlineHbck() method, since some of the tests directly call onlineHbck() rather than going via exec. This makes sure we cleanup the lock in tests\n- Small bug in compareAndSet(), misread the doc, fixed it, also setting it to (false) inside the block to make sure tests using the same HBck object doesn't fail.","from":"developer"},{"body":"thanks bharath, and thanks sean and ted for reviews. committed to branch-1 and master.","from":"developer"},{"body":"Sorry, reverted and reopened since sean had more comments on review board.","from":"developer"},{"body":"v7\n- Addressed Sean's comments on the RB.\n- Few minor changes","from":"developer"},{"body":"[~ busbey]:\nHave all of your comments been addressed ?","from":"developer"},{"body":"Yep. I'm +1 on the latest RB version, which I believe is the same as v7.","from":"developer"},{"body":"Integrated to branch-1 and master.\n\nThanks for the patch, bharath.","from":"developer"},{"body":"The pull back to branch-1 broke compilation. Looks like hte imports are different in the test class there.","from":"developer"},{"body":"I've reverted this from branch-1 for the time being. \n\nReminder, please test before committing to branches!","from":"developer"},{"body":"Will pay attention next time.","from":"developer"},{"body":"With rebased patch for branch-1:\n{code}\nRunning org.apache.hadoop.hbase.util.TestHBaseFsck\nTests run: 44, Failures: 0, Errors: 0, Skipped: 1, Time elapsed: 334.241 sec\n{code}","from":"developer"},{"body":"Integrated to branch-1 again.","from":"developer"},{"body":"Is this not an issue in 0.94 and 0.98?","from":"developer"},{"body":"definitely present in both 0.94 and 0.98. Do you prefer a backport or known issues text? I can pull together either.","from":"developer"},{"body":"11405-1.0.txt applies cleanly on 0.98\n[~apurtell]:\nDo you want this ?","from":"developer"},{"body":"We can attach backport patches here since neither 0.99.1 nor 2.0.0 are released.","from":"developer"},{"body":"+1 for 0.98\nWhy was this branch left out of consideration initially (along with 0.94)? ","from":"developer"},{"body":"Reopened and updated fix versions","from":"developer"},{"body":"Integrated to 0.98 as well.\n\nI can prepare patch for 0.94 if Sean is busy.","from":"developer"},{"body":"I'm chasing down a failure on the hadoop-1 profile and working on 0.94.\nShould be later tonight.\n\nIf someone else can get it done earlier, that'd be great.\n\n-- \nSean\n\n","from":"developer"},{"body":"You mean this one ?\n{code}\ntestParallelHbck(org.apache.hadoop.hbase.util.TestHBaseFsck) Time elapsed: 60.152 sec <<< FAILURE!\njava.lang.AssertionError\n at org.apache.hadoop.hbase.util.TestHBaseFsck.testParallelHbck(TestHBaseFsck.java:552)\n{code}\nSigh - let me revert from 0.98 for now.","from":"developer"},{"body":"okay, the test failure is caused because on [hadoop 1 dfs client transparently retries the file create for up to 5 minutes|https://github.com/apache/hadoop/blob/release-1.2.1/src/hdfs/org/apache/hadoop/hdfs/DFSClient.java#L163].\n\nLooking at the logs of the hadoop 1 run, I can see the second hbck instance fail to create the file. Then later after the first one cleans up, the second picks up again and keeps going. Thus both succeed and the assert fails. Same failure on 0.94.\n\nI'd rather not insert a 5 minute sleep in the hbck to ensure the second instance fails. Anyone have some idea of a different way we can test for the branches that support hadoop 1?","from":"developer"},{"body":"Just a shot in the dark: Use mockito to swap in a different retry policy for creates? ","from":"developer"},{"body":"Here's an amended version specific to each of 0.98 and 0.94 that works under profiles for hadoop 1.0, 1.1, and 2.x. Due to the amount of change, I added an Amending-Author to the commit message.\n\nI've run through {{-Dtest=TestHBaseFsck}} for the above profiles and {{-PrunSmallTests}} for whatever the default profile is on each branch. NB running TestHBaseFsck on 0.94 required {{-PlocalTests}}.\n\nFixing this involved exposing a previously private method for HBaseTestingUtility under 0.94. Since nothing else needed it in 0.98 I marked it deprecated with a note that it would disappear with Hadoop 1.x support. Since the method isn't in branch-1 or later, that should work out well.\n\nOne worry: some of the TestHBaseFsck tests get a bit flaky on 0.94 with this change on my laptop (which admittedly is underpowered). However, even leaving in a single retry period causes the parallel test to fail because the first invocation finished in < 1 sec.","from":"developer"},{"body":"The 0.98 patch lgtm","from":"developer"},{"body":"Rebased patch for 0.98","from":"developer"},{"body":"Integrated to 0.98\n\nRunning test for 0.94 now.","from":"developer"},{"body":"Integrated to 0.94 as well.\n\nThanks Sean and bharath","from":"developer"}],"created":"2014-06-24T05:39:43.000+0000","description":"This is because of the following piece of code in hbck\n\n{code:borderStyle=solid}\n boolean oldBalancer = admin.setBalancerRunning(false, true);\n try {\n onlineConsistencyRepair();\n }\n finally {\n admin.setBalancerRunning(oldBalancer, false);\n }\n{code}\n\nNewer invocations set oldBalancer to false as it was disabled by previous invocations and this disables balancer permanently unless its manually turned on by the user. Easy to reproduce, just run hbck 100 times in a loop in 2 different sessions and you can see that balancer is set to false in the HMaster logs.","issue_id":"12723282","key":"HBASE-11405","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2014-09-25T16:48:42.000+0000","role":"fixed_distractor","summary":"Multiple invocations of hbck in parallel disables balancer permanently "} {"case_id":"12736274","cluster":"DISTRACTOR-HBASE-11813","comments":[{"body":"Ping [~stack], this came in on HBASE-7899","created":"2014-08-23T19:46:01.789+0000"},{"body":"Great, we are eagerly looking forward to the patch :)","created":"2014-08-23T20:53:30.675+0000"},{"body":"The code has been in hbase a good while now. The issue I think is this.cellScanner = this.iterator.next().cellScanner(); where the iterator never finishes. I cannot repro it locally. Its some particularly combo of cell count and lists of cellscanners that is triggering it.","created":"2014-08-23T23:03:20.189+0000"},{"body":"Not that a scanner that never returns is better but can we do this without recursion? ","created":"2014-08-23T23:42:55.052+0000"},{"body":"I'd suspect this one:\n\n{code}\n /**\n * Flatten the map of cells out under the CellScanner\n * @param map Map of Cell Lists; for example, the map of families to Cells that is used\n * inside Put, etc., keeping Cells organized by family.\n * @return CellScanner interface over cellIterable\n */\n public static CellScanner createCellScanner(final NavigableMap> map) {\n return new CellScanner() {\n private final Iterator>> entries =\n map.entrySet().iterator();\n private Iterator currentIterator = null;\n private Cell currentCell;\n\n @Override\n public Cell current() {\n return this.currentCell;\n }\n\n @Override\n public boolean advance() {\n if (this.currentIterator == null) {\n if (!this.entries.hasNext()) return false;\n this.currentIterator = this.entries.next().getValue().iterator();\n }\n if (this.currentIterator.hasNext()) {\n this.currentCell = this.currentIterator.next();\n return true;\n }\n this.currentCell = null;\n this.currentIterator = null;\n return advance();\n }\n };\n }\n{code}\nlooks the one Andrew mentioned would not trigger advance method in server side...while the other one is widely used in server side code paths..coprocessor or end point related..","created":"2014-08-24T04:00:47.880+0000"},{"body":"Patch for 0.98.\n\nWas able to repro a stackoverflow by making up a CellScannerable made of 10k+ CellScanners. Hopefully that is the same as what is going on here.\n\nUndoes recursion and adds tests that do CellScannerables of 100k and a few edge cases.\n\nChanges names of default handlers go give small clue as to what is running.\n\nLet me post master patch. Its a little different.","created":"2014-08-24T06:52:38.177+0000"},{"body":"Rebase","created":"2014-08-24T07:18:08.011+0000"},{"body":"Master patch.","created":"2014-08-24T07:19:24.387+0000"},{"body":"[~Schabby] Suggest you enable DEBUG. This patch below should catch the overflow error, dump some detail on the particular invocation, and allow you keep going:\n\ndiff --git a/hbase-server/src/main/java/org/apache/hadoop/hbase/ipc/CallRunner.java b/hbase-server/src/main/java/org/apache/hadoop/hbase/ipc/CallRunner.java\nindex 31484bb..da2afe0 100644\n--- a/hbase-server/src/main/java/org/apache/hadoop/hbase/ipc/CallRunner.java\n+++ b/hbase-server/src/main/java/org/apache/hadoop/hbase/ipc/CallRunner.java\n@@ -136,9 +136,9 @@ public class CallRunner {\n \"this means that the server was processing a \" +\n \"request but the client went away. The error message was: \" +\n cce.getMessage());\n- } catch (Exception e) {\n+ } catch (Throwable e) {\n RpcServer.LOG.warn(Thread.currentThread().getName()\n- + \": caught: \" + StringUtils.stringifyException(e));\n+ + \": caught: \" + StringUtils.stringifyException(e) + \" call=\" + getCall());\n }\n }\n\nNo guarantees! I tried it and works when no problems.","created":"2014-08-24T07:55:32.983+0000"},{"body":"\noops..it already points to line 210(got fever,brain is not so clear)\nThanks Stack","created":"2014-08-24T09:19:47.203+0000"},{"body":"Sorry for my asking, but where do I get the patch from exactly? git://git.apache.org/hbase.git has the last commits about 15 hours ago.\n\nThanks, Johannes","created":"2014-08-24T09:45:47.420+0000"},{"body":"Ah, nevermind. I believe I just have to apply the patched attached to this bug ticket.","created":"2014-08-24T10:14:16.405+0000"},{"body":"The patch is live in our production cluster for about 2 hours now. So far no RS crash...","created":"2014-08-24T15:16:01.463+0000"},{"body":"Thanks for reporting back [~Schabby]. Please let us know if this is still looking good after a day or so. ","created":"2014-08-24T16:29:00.969+0000"},{"body":"Retry patch. Above failure seems unrelated.\n\nChatting w/ [~Schabby] offline, we may be over the overflow issue and on to new pastures. Corruption? TBD.","created":"2014-08-25T01:51:07.924+0000"},{"body":"bq.-1 core tests. The patch failed these unit tests:\norg.apache.hadoop.hbase.TestCellUtil\nSeems the test case is related to the patch.","created":"2014-08-25T05:25:47.135+0000"},{"body":"Quick update from me. \n\nWe have Stacks patch running for a day now. The StackOverflowExceptions did not occur again, all RS are operational and the cluster did not hang. We keep our home-compiled HBase running now until the patch makes it to an official release.\n\nOur client code still queries very large batches at times. Random-access of 100k records in one query is likely for us. Large batches were the original cause of this issue. With the recursion issue resolved, we now observed two non-dramatic cases where the client timed out and a ChannelClosedException was thrown on the server side without killing the RS. Stack and I suspect that a large query is taking to long to process/transmit, but we havent figured out the root cause yet (region is consistent). We will adjust our logging a bit and keep an eye on it. Besides these two cases, nothing happend so far.\n\nThank you all for the quick fix and the responsiveness!","created":"2014-08-25T12:22:01.622+0000"},{"body":"Previous failure was because I attached test only and not the fix by mistake. Here is fix and test.","created":"2014-08-25T15:09:59.517+0000"},{"body":"Review?","created":"2014-08-25T18:00:11.708+0000"},{"body":"javadoc is unrelated. I can fix on commit though: [WARNING] /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/wal/WALCellCodec.java:96: warning - Tag @link: reference not found: cellCodecClsName","created":"2014-08-25T18:01:27.755+0000"},{"body":"+1 v3 trunk and 0.98 patches. (Reviewed the CellUtil changes, they are identical.)","created":"2014-08-25T21:45:36.418+0000"},{"body":"Thanks Andrew.\n\nCommitted to 0.98+","created":"2014-08-25T21:56:41.329+0000"},{"body":"Closing this issue after 0.99.0 release. ","created":"2015-02-21T23:34:02.531+0000"}],"conversations":[{"body":"On user@hbase, johannes.schaback@visual-meta.com reported:\n{quote}\nwe face a serious issue with our HBase production cluster for two days now. Every couple minutes, a random RegionServer gets stuck and does not process any requests. In addition this causes the other RegionServers to freeze within a minute which brings down the entire cluster. Stopping the affected RegionServer unblocks the cluster and everything comes back to normal.\n{quote}\n\nSubsequent troubleshooting reveals that RPC is getting stuck because we are losing RPC handlers. In the .out files we have this:\n{noformat}\nException in thread \"defaultRpcServer.handler=5,queue=2,port=60020\"\njava.lang.StackOverflowError\n at org.apache.hadoop.hbase.CellUtil$1.advance(CellUtil.java:210)\n at org.apache.hadoop.hbase.CellUtil$1.advance(CellUtil.java:210)\n at org.apache.hadoop.hbase.CellUtil$1.advance(CellUtil.java:210)\n at org.apache.hadoop.hbase.CellUtil$1.advance(CellUtil.java:210)\n[...]\nException in thread \"defaultRpcServer.handler=5,queue=2,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=18,queue=0,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=23,queue=2,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=24,queue=0,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=2,queue=2,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=11,queue=2,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=25,queue=1,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=20,queue=2,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=19,queue=1,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=15,queue=0,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=1,queue=1,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=7,queue=1,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=4,queue=1,port=60020\"\njava.lang.StackOverflowError​\n{noformat}\n\nThat is the anonymous CellScanner instance we create from CellUtil#createCellScanner:\n{code}\n​ return new CellScanner() {\n private final Iterator iterator = cellScannerables.iterator();\n private CellScanner cellScanner = null;\n\n @Override\n public Cell current() {\n return this.cellScanner != null? this.cellScanner.current(): null;\n }\n\n @Override\n public boolean advance() throws IOException {\n if (this.cellScanner == null) {\n if (!this.iterator.hasNext()) return false;\n this.cellScanner = this.iterator.next().cellScanner();\n }\n if (this.cellScanner.advance()) return true;\n this.cellScanner = null;\n---> return advance();\n }\n };\n{code}\n\nThat final return statement is the immediate problem.\n\nWe should also fix this so the RegionServer aborts if it loses a handler to an Error. \n","from":"reporter","subject":"CellScanner#advance may overflow stack"},{"body":"Ping [~stack], this came in on HBASE-7899","from":"developer"},{"body":"Great, we are eagerly looking forward to the patch :)","from":"developer"},{"body":"The code has been in hbase a good while now. The issue I think is this.cellScanner = this.iterator.next().cellScanner(); where the iterator never finishes. I cannot repro it locally. Its some particularly combo of cell count and lists of cellscanners that is triggering it.","from":"developer"},{"body":"Not that a scanner that never returns is better but can we do this without recursion? ","from":"developer"},{"body":"I'd suspect this one:\n\n{code}\n /**\n * Flatten the map of cells out under the CellScanner\n * @param map Map of Cell Lists; for example, the map of families to Cells that is used\n * inside Put, etc., keeping Cells organized by family.\n * @return CellScanner interface over cellIterable\n */\n public static CellScanner createCellScanner(final NavigableMap> map) {\n return new CellScanner() {\n private final Iterator>> entries =\n map.entrySet().iterator();\n private Iterator currentIterator = null;\n private Cell currentCell;\n\n @Override\n public Cell current() {\n return this.currentCell;\n }\n\n @Override\n public boolean advance() {\n if (this.currentIterator == null) {\n if (!this.entries.hasNext()) return false;\n this.currentIterator = this.entries.next().getValue().iterator();\n }\n if (this.currentIterator.hasNext()) {\n this.currentCell = this.currentIterator.next();\n return true;\n }\n this.currentCell = null;\n this.currentIterator = null;\n return advance();\n }\n };\n }\n{code}\nlooks the one Andrew mentioned would not trigger advance method in server side...while the other one is widely used in server side code paths..coprocessor or end point related..","from":"developer"},{"body":"Patch for 0.98.\n\nWas able to repro a stackoverflow by making up a CellScannerable made of 10k+ CellScanners. Hopefully that is the same as what is going on here.\n\nUndoes recursion and adds tests that do CellScannerables of 100k and a few edge cases.\n\nChanges names of default handlers go give small clue as to what is running.\n\nLet me post master patch. Its a little different.","from":"developer"},{"body":"Rebase","from":"developer"},{"body":"Master patch.","from":"developer"},{"body":"[~Schabby] Suggest you enable DEBUG. This patch below should catch the overflow error, dump some detail on the particular invocation, and allow you keep going:\n\ndiff --git a/hbase-server/src/main/java/org/apache/hadoop/hbase/ipc/CallRunner.java b/hbase-server/src/main/java/org/apache/hadoop/hbase/ipc/CallRunner.java\nindex 31484bb..da2afe0 100644\n--- a/hbase-server/src/main/java/org/apache/hadoop/hbase/ipc/CallRunner.java\n+++ b/hbase-server/src/main/java/org/apache/hadoop/hbase/ipc/CallRunner.java\n@@ -136,9 +136,9 @@ public class CallRunner {\n \"this means that the server was processing a \" +\n \"request but the client went away. The error message was: \" +\n cce.getMessage());\n- } catch (Exception e) {\n+ } catch (Throwable e) {\n RpcServer.LOG.warn(Thread.currentThread().getName()\n- + \": caught: \" + StringUtils.stringifyException(e));\n+ + \": caught: \" + StringUtils.stringifyException(e) + \" call=\" + getCall());\n }\n }\n\nNo guarantees! I tried it and works when no problems.","from":"developer"},{"body":"\noops..it already points to line 210(got fever,brain is not so clear)\nThanks Stack","from":"developer"},{"body":"Sorry for my asking, but where do I get the patch from exactly? git://git.apache.org/hbase.git has the last commits about 15 hours ago.\n\nThanks, Johannes","from":"developer"},{"body":"Ah, nevermind. I believe I just have to apply the patched attached to this bug ticket.","from":"developer"},{"body":"The patch is live in our production cluster for about 2 hours now. So far no RS crash...","from":"developer"},{"body":"Thanks for reporting back [~Schabby]. Please let us know if this is still looking good after a day or so. ","from":"developer"},{"body":"Retry patch. Above failure seems unrelated.\n\nChatting w/ [~Schabby] offline, we may be over the overflow issue and on to new pastures. Corruption? TBD.","from":"developer"},{"body":"bq.-1 core tests. The patch failed these unit tests:\norg.apache.hadoop.hbase.TestCellUtil\nSeems the test case is related to the patch.","from":"developer"},{"body":"Quick update from me. \n\nWe have Stacks patch running for a day now. The StackOverflowExceptions did not occur again, all RS are operational and the cluster did not hang. We keep our home-compiled HBase running now until the patch makes it to an official release.\n\nOur client code still queries very large batches at times. Random-access of 100k records in one query is likely for us. Large batches were the original cause of this issue. With the recursion issue resolved, we now observed two non-dramatic cases where the client timed out and a ChannelClosedException was thrown on the server side without killing the RS. Stack and I suspect that a large query is taking to long to process/transmit, but we havent figured out the root cause yet (region is consistent). We will adjust our logging a bit and keep an eye on it. Besides these two cases, nothing happend so far.\n\nThank you all for the quick fix and the responsiveness!","from":"developer"},{"body":"Previous failure was because I attached test only and not the fix by mistake. Here is fix and test.","from":"developer"},{"body":"Review?","from":"developer"},{"body":"javadoc is unrelated. I can fix on commit though: [WARNING] /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/wal/WALCellCodec.java:96: warning - Tag @link: reference not found: cellCodecClsName","from":"developer"},{"body":"+1 v3 trunk and 0.98 patches. (Reviewed the CellUtil changes, they are identical.)","from":"developer"},{"body":"Thanks Andrew.\n\nCommitted to 0.98+","from":"developer"},{"body":"Closing this issue after 0.99.0 release. ","from":"developer"}],"created":"2014-08-23T19:42:59.000+0000","description":"On user@hbase, johannes.schaback@visual-meta.com reported:\n{quote}\nwe face a serious issue with our HBase production cluster for two days now. Every couple minutes, a random RegionServer gets stuck and does not process any requests. In addition this causes the other RegionServers to freeze within a minute which brings down the entire cluster. Stopping the affected RegionServer unblocks the cluster and everything comes back to normal.\n{quote}\n\nSubsequent troubleshooting reveals that RPC is getting stuck because we are losing RPC handlers. In the .out files we have this:\n{noformat}\nException in thread \"defaultRpcServer.handler=5,queue=2,port=60020\"\njava.lang.StackOverflowError\n at org.apache.hadoop.hbase.CellUtil$1.advance(CellUtil.java:210)\n at org.apache.hadoop.hbase.CellUtil$1.advance(CellUtil.java:210)\n at org.apache.hadoop.hbase.CellUtil$1.advance(CellUtil.java:210)\n at org.apache.hadoop.hbase.CellUtil$1.advance(CellUtil.java:210)\n[...]\nException in thread \"defaultRpcServer.handler=5,queue=2,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=18,queue=0,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=23,queue=2,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=24,queue=0,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=2,queue=2,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=11,queue=2,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=25,queue=1,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=20,queue=2,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=19,queue=1,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=15,queue=0,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=1,queue=1,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=7,queue=1,port=60020\"\njava.lang.StackOverflowError\nException in thread \"defaultRpcServer.handler=4,queue=1,port=60020\"\njava.lang.StackOverflowError​\n{noformat}\n\nThat is the anonymous CellScanner instance we create from CellUtil#createCellScanner:\n{code}\n​ return new CellScanner() {\n private final Iterator iterator = cellScannerables.iterator();\n private CellScanner cellScanner = null;\n\n @Override\n public Cell current() {\n return this.cellScanner != null? this.cellScanner.current(): null;\n }\n\n @Override\n public boolean advance() throws IOException {\n if (this.cellScanner == null) {\n if (!this.iterator.hasNext()) return false;\n this.cellScanner = this.iterator.next().cellScanner();\n }\n if (this.cellScanner.advance()) return true;\n this.cellScanner = null;\n---> return advance();\n }\n };\n{code}\n\nThat final return statement is the immediate problem.\n\nWe should also fix this so the RegionServer aborts if it loses a handler to an Error. \n","issue_id":"12736274","key":"HBASE-11813","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2014-08-25T21:56:41.000+0000","role":"fixed_distractor","summary":"CellScanner#advance may overflow stack"} {"case_id":"12738344","cluster":"DISTRACTOR-HBASE-11876","comments":[{"body":"Let's discuss. The one problem I see that a client (coprocessor) could call nextRaw(...) cause lots of load on the RS, that would not register as a scan metric on the RS unless the coprocessor is a good citizen. (of course a called can already wreak havoc by failing to start/stop the region operation or to synchronize on the scanner object)\n\nIf no objections I'll make a patch.","created":"2014-09-02T04:52:55.399+0000"},{"body":"Sounds good to me. Want to get a patch into .6? I will be tagging the next RC tonight.","created":"2014-09-02T16:38:01.675+0000"},{"body":"Sounds good as long as broadcast wide that this is how nextRaw works.","created":"2014-09-02T19:02:56.339+0000"},{"body":"Given the time constraints I'll take this","created":"2014-09-02T21:45:40.390+0000"},{"body":"Attached patch removes metrics update from HRegion.RegionScannerImpl#nextRaw, keeps track of the equivalent metric in RSRpcServices#scan, and updates the metric once per batch. When limiting the result set by size we use KeyValue#heapSize, which also accounts for object heap overheads, but when calculating scan metrics we use KeyValue#getLength, which is the on wire size of the key value(s). I've maintained this distinction.","created":"2014-09-02T22:54:54.159+0000"},{"body":"+1 (non-binding). But just a small question about this block\n- for (; i < rows\n- && currentScanResultSize < maxResultSize; ) {\n+ for (; i < rows; i++) {\n+ // Stop collecting results if maxScannerResultSize is set and we have exceeded it\n+ if ((maxScannerResultSize < Long.MAX_VALUE) &&\n+ (currentScanResultSize >= maxResultSize)) {\n+ break;\n+ }\n // Collect values to be returned here\n boolean moreRows = scanner.nextRaw(values);\n if (!values.isEmpty()) {\n- if (maxScannerResultSize < Long.MAX_VALUE){\n- for (Cell kv : values) {\n- currentScanResultSize += KeyValueUtil.ensureKeyValue(kv).heapSize();\n- }\n+ for (Cell cell : values) {\n+ KeyValue kv = KeyValueUtil.ensureKeyValue(cell);\n+ currentScanResultSize += kv.heapSize();\n+ totalKvSize += kv.getLength();\n }\n results.add(Result.create(values, null, stale));\n- i++;\n }\n{code}\nThe increment for variable to i moved out of the if (!values.isEmpty()) block. Does this mean that values is never going to be empty? If so, then can we do away without it?\n","created":"2014-09-02T23:12:00.077+0000"},{"body":"Sry messed with the code block. This is the one I'm referring to\n{code}\n synchronized(scanner) {\n boolean stale = (region.getRegionInfo().getReplicaId() != 0);\n- for (; i < rows\n- && currentScanResultSize < maxResultSize; ) {\n+ for (; i < rows; i++) {\n+ // Stop collecting results if maxScannerResultSize is set and we have exceeded it\n+ if ((maxScannerResultSize < Long.MAX_VALUE) &&\n+ (currentScanResultSize >= maxResultSize)) {\n+ break;\n+ }\n // Collect values to be returned here\n boolean moreRows = scanner.nextRaw(values);\n if (!values.isEmpty()) {\n- if (maxScannerResultSize < Long.MAX_VALUE){\n- for (Cell kv : values) {\n- currentScanResultSize += KeyValueUtil.ensureKeyValue(kv).heapSize();\n- }\n+ for (Cell cell : values) {\n+ KeyValue kv = KeyValueUtil.ensureKeyValue(cell);\n+ currentScanResultSize += kv.heapSize();\n+ totalKvSize += kv.getLength();\n }\n results.add(Result.create(values, null, stale));\n- i++;\n }\n{code}","created":"2014-09-02T23:15:01.388+0000"},{"body":"bq. The increment for variable to i moved out of the if (!values.isEmpty()) block.\n\nYeah you're right I should not have done that. Updated patch.","created":"2014-09-02T23:19:11.034+0000"},{"body":"Any chance for a quick pass over this [~lhofhansl] or [~stack]? I'd like to get this in for .6, tonight.","created":"2014-09-03T00:05:38.706+0000"},{"body":"Patch for 0.98","created":"2014-09-03T00:40:16.433+0000"},{"body":"+1","created":"2014-09-03T00:52:30.790+0000"},{"body":"Thanks [~eclark], will commit shortly unless objection.","created":"2014-09-03T01:03:23.742+0000"},{"body":"Committed to 0.98+","created":"2014-09-03T01:29:33.884+0000"},{"body":"Sorry was traveling today. Was just finding some time to make a quick patch, and it's already done. :)\n\nBelated +1 on patch then, that's almost exactly how I would have coded it up. Should be a nice improvement.\n","created":"2014-09-03T06:27:44.598+0000"},{"body":"The trunk change passed precommit but the 0.98 version requires an addendum to fix a failing test, see https://builds.apache.org/job/HBase-0.98/493/testReport/org.apache.hadoop.hbase.regionserver/TestRegionServerMetrics/testScanNext/","created":"2014-09-03T13:44:32.531+0000"},{"body":"There are two code paths in 0.98 which call nextRaw from scanning code. One isn't found by Eclipse reference search but grep did the trick. Will revert previous commit and push the fixed version as soon as local tests check out.","created":"2014-09-03T15:34:07.197+0000"},{"body":"Pushed to 0.98. All hbase-server tests pass locally.","created":"2014-09-03T16:43:14.985+0000"},{"body":"Closing this issue after 0.99.0 release. ","created":"2015-02-21T23:34:57.703+0000"}],"conversations":[{"body":"I added the RegionScanner.nextRaw(...) to allow \"smart\" client to avoid some of the default work that HBase is doing, such as {start|stop}RegionOperation and synchronized(scanner) for each row.\n\nMetrics should follow the same approach. Collecting them per row is expensive and a caller should have the option to collect those later or to avoid collecting them completely.\n\nWe can also save some cycles in RSRcpServices.scan(...) if we updated the metric only once/batch instead of each row.\n","from":"reporter","subject":"RegionScanner.nextRaw(...) should not update metrics"},{"body":"Let's discuss. The one problem I see that a client (coprocessor) could call nextRaw(...) cause lots of load on the RS, that would not register as a scan metric on the RS unless the coprocessor is a good citizen. (of course a called can already wreak havoc by failing to start/stop the region operation or to synchronize on the scanner object)\n\nIf no objections I'll make a patch.","from":"developer"},{"body":"Sounds good to me. Want to get a patch into .6? I will be tagging the next RC tonight.","from":"developer"},{"body":"Sounds good as long as broadcast wide that this is how nextRaw works.","from":"developer"},{"body":"Given the time constraints I'll take this","from":"developer"},{"body":"Attached patch removes metrics update from HRegion.RegionScannerImpl#nextRaw, keeps track of the equivalent metric in RSRpcServices#scan, and updates the metric once per batch. When limiting the result set by size we use KeyValue#heapSize, which also accounts for object heap overheads, but when calculating scan metrics we use KeyValue#getLength, which is the on wire size of the key value(s). I've maintained this distinction.","from":"developer"},{"body":"+1 (non-binding). But just a small question about this block\n- for (; i < rows\n- && currentScanResultSize < maxResultSize; ) {\n+ for (; i < rows; i++) {\n+ // Stop collecting results if maxScannerResultSize is set and we have exceeded it\n+ if ((maxScannerResultSize < Long.MAX_VALUE) &&\n+ (currentScanResultSize >= maxResultSize)) {\n+ break;\n+ }\n // Collect values to be returned here\n boolean moreRows = scanner.nextRaw(values);\n if (!values.isEmpty()) {\n- if (maxScannerResultSize < Long.MAX_VALUE){\n- for (Cell kv : values) {\n- currentScanResultSize += KeyValueUtil.ensureKeyValue(kv).heapSize();\n- }\n+ for (Cell cell : values) {\n+ KeyValue kv = KeyValueUtil.ensureKeyValue(cell);\n+ currentScanResultSize += kv.heapSize();\n+ totalKvSize += kv.getLength();\n }\n results.add(Result.create(values, null, stale));\n- i++;\n }\n{code}\nThe increment for variable to i moved out of the if (!values.isEmpty()) block. Does this mean that values is never going to be empty? If so, then can we do away without it?\n","from":"developer"},{"body":"Sry messed with the code block. This is the one I'm referring to\n{code}\n synchronized(scanner) {\n boolean stale = (region.getRegionInfo().getReplicaId() != 0);\n- for (; i < rows\n- && currentScanResultSize < maxResultSize; ) {\n+ for (; i < rows; i++) {\n+ // Stop collecting results if maxScannerResultSize is set and we have exceeded it\n+ if ((maxScannerResultSize < Long.MAX_VALUE) &&\n+ (currentScanResultSize >= maxResultSize)) {\n+ break;\n+ }\n // Collect values to be returned here\n boolean moreRows = scanner.nextRaw(values);\n if (!values.isEmpty()) {\n- if (maxScannerResultSize < Long.MAX_VALUE){\n- for (Cell kv : values) {\n- currentScanResultSize += KeyValueUtil.ensureKeyValue(kv).heapSize();\n- }\n+ for (Cell cell : values) {\n+ KeyValue kv = KeyValueUtil.ensureKeyValue(cell);\n+ currentScanResultSize += kv.heapSize();\n+ totalKvSize += kv.getLength();\n }\n results.add(Result.create(values, null, stale));\n- i++;\n }\n{code}","from":"developer"},{"body":"bq. The increment for variable to i moved out of the if (!values.isEmpty()) block.\n\nYeah you're right I should not have done that. Updated patch.","from":"developer"},{"body":"Any chance for a quick pass over this [~lhofhansl] or [~stack]? I'd like to get this in for .6, tonight.","from":"developer"},{"body":"Patch for 0.98","from":"developer"},{"body":"+1","from":"developer"},{"body":"Thanks [~eclark], will commit shortly unless objection.","from":"developer"},{"body":"Committed to 0.98+","from":"developer"},{"body":"Sorry was traveling today. Was just finding some time to make a quick patch, and it's already done. :)\n\nBelated +1 on patch then, that's almost exactly how I would have coded it up. Should be a nice improvement.\n","from":"developer"},{"body":"The trunk change passed precommit but the 0.98 version requires an addendum to fix a failing test, see https://builds.apache.org/job/HBase-0.98/493/testReport/org.apache.hadoop.hbase.regionserver/TestRegionServerMetrics/testScanNext/","from":"developer"},{"body":"There are two code paths in 0.98 which call nextRaw from scanning code. One isn't found by Eclipse reference search but grep did the trick. Will revert previous commit and push the fixed version as soon as local tests check out.","from":"developer"},{"body":"Pushed to 0.98. All hbase-server tests pass locally.","from":"developer"},{"body":"Closing this issue after 0.99.0 release. ","from":"developer"}],"created":"2014-09-02T04:50:45.000+0000","description":"I added the RegionScanner.nextRaw(...) to allow \"smart\" client to avoid some of the default work that HBase is doing, such as {start|stop}RegionOperation and synchronized(scanner) for each row.\n\nMetrics should follow the same approach. Collecting them per row is expensive and a caller should have the option to collect those later or to avoid collecting them completely.\n\nWe can also save some cycles in RSRcpServices.scan(...) if we updated the metric only once/batch instead of each row.\n","issue_id":"12738344","key":"HBASE-11876","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2014-09-03T16:43:14.000+0000","role":"fixed_distractor","summary":"RegionScanner.nextRaw(...) should not update metrics"} {"case_id":"12415413","cluster":"DISTRACTOR-HBASE-1212","comments":[{"body":"As I noted on HBASE-1274, we can try using the file names to do the sequencing:\n\nbq. What about using something that sorts lexicographically and sorting the dir.listFiles we use when grabbing the store files? Then when we move around two files with the same sequence ids to do a merge we can rename the actual files.","created":"2009-03-24T03:43:39.637+0000"},{"body":"Moving to 0.20.0","created":"2009-03-25T09:55:58.338+0000"},{"body":"I like the jgray idea that we use modtime instead of an edit number.","created":"2009-04-22T19:00:32.110+0000"},{"body":"Thinking on it, this event should be extremely rare. Sequence ids are monotonically increasing in a running regionserver. Across a cluster, two files of the same family would have to end with same sequenceid. Then whats the likelihood that of all regions on cluster these are the two to merge (Merge is a little-used tool to date).\n\nTo fix, would need to look at the content of the two files and make a judgement as to which should come before the other -- which has the most recent edits. Maybe we could do something basic like let the file with the largest size prevail over the smaller. Once we'd figure which file to bring to the fore, we need to rewrite the hfile so we can change the sequence id. Since we're rewriting one of the files at least, might as well compact them.\n\nWe could move to modification times. That should simplify this sequenceid story. It wouldn't remove this issue. We'd still have to figure which store file to favor if two happened to have same mod time.\n\nIn bigtable, chubby owns the storefiles/sstables. Maybe thats where we should go so we don't have sequenceids anymore?\n\nMoving out of 0.20.0 because this issue rare and amount of work to address is large.\n\n","created":"2009-05-11T17:16:04.089+0000"},{"body":"Hi there,\nWhen using util.merge on a table (created under hbase 0.20.6, but now running under 0.90.1) I get the following exception very often (which I think is the result of this jira):\n{code} \nFATAL util.Merge: Merge failed\njava.io.IOException: Files have same sequenceid: 1283640296746\n\tat org.apache.hadoop.hbase.regionserver.HRegion.merge(HRegion.java:2785)\n\tat org.apache.hadoop.hbase.util.Merge.merge(Merge.java:287)\n\tat org.apache.hadoop.hbase.util.Merge.mergeTwoRegions(Merge.java:238)\n\tat org.apache.hadoop.hbase.util.Merge.run(Merge.java:110)\n\tat org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:65)\n\tat org.apache.hadoop.hbase.util.Merge.main(Merge.java:379)\n{code}\n\nUnfortunately at this point some families of the region were already moved to the new merged region which leaves the table in an inconsistent state.\nIs there any way to detect this failure case before running the merge tool (except for opening the store files in advance and calling getMaxSequenceId)?\n\nSame issue is also mentioned in the [comments|https://issues.apache.org/jira/browse/HBASE-1621?focusedCommentId=12993477#comment-12993477] of HBASE-1621","created":"2011-03-02T16:13:12.378+0000"},{"body":"This is worse now that bulk upload is popular; all get low seqid so likelihood of clashing more likely.\n\nMade it so I had to do manual merge doing fixup of user cluster.","created":"2011-03-04T00:37:46.778+0000"},{"body":"now that we dont key files by their sequence id in Store.java, this is simpler: just allow duplicate sequence ids.","created":"2011-03-04T00:50:42.301+0000"},{"body":"Can someone comment, if this is still an issue?","created":"2012-06-06T08:42:36.092+0000"},{"body":"FYI, I'm facing this issue today while trying to merge all the regions of a table. I have created a 54 regions table, and while trying to merge all of them back into a single one I got this error. I'm able to reproduce it.","created":"2012-12-05T18:44:44.746+0000"},{"body":"Attached logs from a merge failure.","created":"2012-12-05T21:27:58.199+0000"},{"body":"Based on Ryan's comment on March 3rd 2011 I did some more testing.\n\nHere is how I proceed:\nCreation of a test table.\nTruncate the test table if already there and not empty.\nInsert 100000 records\nhbck => ok.\nSplit into 256 regions using HTML interface\nBalancer + major_compact\nTry to merge all the regions 2 by 2 until only 1 region remaining or until error.\nIf no error, restart from split. (Usually need to do it twice)\n\nNow, at this stage, I have 2 regions with the same sequence ID that I'm not able to merge with the current un-patched HRegion code.\n\nDisplayed:starting merge of regions: testtable,?\\xEC\\x80C^\\xB2Q\\xE8,1354799289826.eb1b2213a6b09a9d989c39abe3d6d18b. and testtable,?\\xEC\\xF9\\x88\\xBE{Sb,1354799292172.2d81e98f11bdda12dbf06c11eefefb39. into new region {NAME => 'testtable,?\\xEC\\x80C^\\xB2Q\\xE8,1354799434888.e4e403c2bf3b3b2561f26bfb70ef48d1.', STARTKEY => '?\\xEC\\x80C^\\xB2Q\\xE8', ENDKEY => '?\\xEDw\\xF5\\x16\\xC4\\x81\\xB7', ENCODED => e4e403c2bf3b3b2561f26bfb70ef48d1,} with start key and end key \n\n\nStart-hbase, hbck => One issue on the newly created region (HBASE-7287 e4e403c2bf3b3b2561f26bfb70ef48d)\nrowcount => 12/12/06 08:15:14 INFO mapred.JobClient: ROWS=100000\n\nApply HBASE-1212 patch, stop HBase and restart the merge.\nMerge went well up to the end (all regions merged 2 by 2 until only 1 region remaining)\nStart-hbase, hbck => One issue on the newly created region (HBASE-7287 e4e403c2bf3b3b2561f26bfb70ef48d) Same as before, so no additionnal issue because of HBASE-1212.\nrowcount => 12/12/06 08:21:54 INFO mapred.JobClient: ROWS=100000\n\n\nThe issue created by the merge when not patched (HBASE-7287 or HBASE-1212) is just an empty directory created which can be manually removed.\n\nIf we apply HBASE-1212, then there is no more needs to apply HBASE-7287. If we want to make more tests for HBASE-1212 then it might be good to apply HBASE-7287 to avoid system inconsistencies.\n\nBased on my tests I can confirm that we can remove the sequenceID check and the merge will still be working fine even if 2 files got the same sequenceID.","created":"2012-12-06T13:29:42.965+0000"},{"body":"Patch to remove the sequenceID check.","created":"2012-12-06T13:30:17.700+0000"},{"body":"Can you attach patch for trunk ?\n\nThanks","created":"2012-12-06T15:26:34.013+0000"},{"body":"Hi Ted,\n\nI have already attached the patch. I did it on my eclipse which is configured with the trunk. So it should be correct? Or have I done anything wrong?\n\nThanks.","created":"2012-12-06T16:07:58.175+0000"},{"body":"From your patch:\n{code}\nIndex: src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n{code}\nIn trunk, the path would be hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\nCan you use 'svn diff' command to produce next patch under the root of your workspace ?","created":"2012-12-06T17:17:24.937+0000"},{"body":"Oh, sorry about that! Did not figured I missed that.\n\nDebian come with VSN < 1.7 so this command line was not working for me. I installed 1.7, run it and got the patch for all the modifications I made. So attached is the extract for HBASE-1212. Can I mark HBASE-7287 as resolved by this one?","created":"2012-12-06T18:29:27.194+0000"},{"body":"Change looks good to me. Taking a cursory look at the code we're indeed not using the sequenceId solely to distinguish HFiles.\nCan somebody reconfirm that.\n","created":"2012-12-06T19:00:20.197+0000"},{"body":"Re-submitting the refreshed patch.","created":"2013-01-26T02:41:19.010+0000"},{"body":"We will use sequenceId to sort HFiles, like in compaction.\nWhen merging, there is no overlap between these two files from different regions, I think it's ok even if two files have the same sequence ID since we don't use it to distingush the files in Store any more \n\nPatch looks good for me","created":"2013-01-28T06:02:35.569+0000"},{"body":"+1 from me.","created":"2013-01-28T16:03:08.107+0000"},{"body":"Should I upload one for 0.94?","created":"2013-01-28T16:08:59.167+0000"},{"body":"Integrated to trunk.\n\nThanks for the patch, Jean-Marc.\n\nThanks for the review, Chunhui.","created":"2013-01-28T16:55:32.642+0000"},{"body":"Was integrated a while back. Resolving.","created":"2013-04-23T06:33:31.360+0000"}],"conversations":[{"body":"Currently merging two regions, the merge tool will compare their sequence ids. If same, it will decrement one. It needs to do this because on region open, files are keyed by their sequenceid; if two the same, one will erase the other.\n\nWell, with the move to the aggregating hfile format, the sequenceid is written when the file is created and its no longer written into an aside file but as metadata on to the end of the file. Changing the sequenceid is no longer an option.\n\nThis issue is about figuring a solution for the rare case where two store files have same sequence id AND we want to merge the two regions.","from":"reporter","subject":"merge tool expects regions all have different sequence ids"},{"body":"As I noted on HBASE-1274, we can try using the file names to do the sequencing:\n\nbq. What about using something that sorts lexicographically and sorting the dir.listFiles we use when grabbing the store files? Then when we move around two files with the same sequence ids to do a merge we can rename the actual files.","from":"developer"},{"body":"Moving to 0.20.0","from":"developer"},{"body":"I like the jgray idea that we use modtime instead of an edit number.","from":"developer"},{"body":"Thinking on it, this event should be extremely rare. Sequence ids are monotonically increasing in a running regionserver. Across a cluster, two files of the same family would have to end with same sequenceid. Then whats the likelihood that of all regions on cluster these are the two to merge (Merge is a little-used tool to date).\n\nTo fix, would need to look at the content of the two files and make a judgement as to which should come before the other -- which has the most recent edits. Maybe we could do something basic like let the file with the largest size prevail over the smaller. Once we'd figure which file to bring to the fore, we need to rewrite the hfile so we can change the sequence id. Since we're rewriting one of the files at least, might as well compact them.\n\nWe could move to modification times. That should simplify this sequenceid story. It wouldn't remove this issue. We'd still have to figure which store file to favor if two happened to have same mod time.\n\nIn bigtable, chubby owns the storefiles/sstables. Maybe thats where we should go so we don't have sequenceids anymore?\n\nMoving out of 0.20.0 because this issue rare and amount of work to address is large.\n\n","from":"developer"},{"body":"Hi there,\nWhen using util.merge on a table (created under hbase 0.20.6, but now running under 0.90.1) I get the following exception very often (which I think is the result of this jira):\n{code} \nFATAL util.Merge: Merge failed\njava.io.IOException: Files have same sequenceid: 1283640296746\n\tat org.apache.hadoop.hbase.regionserver.HRegion.merge(HRegion.java:2785)\n\tat org.apache.hadoop.hbase.util.Merge.merge(Merge.java:287)\n\tat org.apache.hadoop.hbase.util.Merge.mergeTwoRegions(Merge.java:238)\n\tat org.apache.hadoop.hbase.util.Merge.run(Merge.java:110)\n\tat org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:65)\n\tat org.apache.hadoop.hbase.util.Merge.main(Merge.java:379)\n{code}\n\nUnfortunately at this point some families of the region were already moved to the new merged region which leaves the table in an inconsistent state.\nIs there any way to detect this failure case before running the merge tool (except for opening the store files in advance and calling getMaxSequenceId)?\n\nSame issue is also mentioned in the [comments|https://issues.apache.org/jira/browse/HBASE-1621?focusedCommentId=12993477#comment-12993477] of HBASE-1621","from":"developer"},{"body":"This is worse now that bulk upload is popular; all get low seqid so likelihood of clashing more likely.\n\nMade it so I had to do manual merge doing fixup of user cluster.","from":"developer"},{"body":"now that we dont key files by their sequence id in Store.java, this is simpler: just allow duplicate sequence ids.","from":"developer"},{"body":"Can someone comment, if this is still an issue?","from":"developer"},{"body":"FYI, I'm facing this issue today while trying to merge all the regions of a table. I have created a 54 regions table, and while trying to merge all of them back into a single one I got this error. I'm able to reproduce it.","from":"developer"},{"body":"Attached logs from a merge failure.","from":"developer"},{"body":"Based on Ryan's comment on March 3rd 2011 I did some more testing.\n\nHere is how I proceed:\nCreation of a test table.\nTruncate the test table if already there and not empty.\nInsert 100000 records\nhbck => ok.\nSplit into 256 regions using HTML interface\nBalancer + major_compact\nTry to merge all the regions 2 by 2 until only 1 region remaining or until error.\nIf no error, restart from split. (Usually need to do it twice)\n\nNow, at this stage, I have 2 regions with the same sequence ID that I'm not able to merge with the current un-patched HRegion code.\n\nDisplayed:starting merge of regions: testtable,?\\xEC\\x80C^\\xB2Q\\xE8,1354799289826.eb1b2213a6b09a9d989c39abe3d6d18b. and testtable,?\\xEC\\xF9\\x88\\xBE{Sb,1354799292172.2d81e98f11bdda12dbf06c11eefefb39. into new region {NAME => 'testtable,?\\xEC\\x80C^\\xB2Q\\xE8,1354799434888.e4e403c2bf3b3b2561f26bfb70ef48d1.', STARTKEY => '?\\xEC\\x80C^\\xB2Q\\xE8', ENDKEY => '?\\xEDw\\xF5\\x16\\xC4\\x81\\xB7', ENCODED => e4e403c2bf3b3b2561f26bfb70ef48d1,} with start key and end key \n\n\nStart-hbase, hbck => One issue on the newly created region (HBASE-7287 e4e403c2bf3b3b2561f26bfb70ef48d)\nrowcount => 12/12/06 08:15:14 INFO mapred.JobClient: ROWS=100000\n\nApply HBASE-1212 patch, stop HBase and restart the merge.\nMerge went well up to the end (all regions merged 2 by 2 until only 1 region remaining)\nStart-hbase, hbck => One issue on the newly created region (HBASE-7287 e4e403c2bf3b3b2561f26bfb70ef48d) Same as before, so no additionnal issue because of HBASE-1212.\nrowcount => 12/12/06 08:21:54 INFO mapred.JobClient: ROWS=100000\n\n\nThe issue created by the merge when not patched (HBASE-7287 or HBASE-1212) is just an empty directory created which can be manually removed.\n\nIf we apply HBASE-1212, then there is no more needs to apply HBASE-7287. If we want to make more tests for HBASE-1212 then it might be good to apply HBASE-7287 to avoid system inconsistencies.\n\nBased on my tests I can confirm that we can remove the sequenceID check and the merge will still be working fine even if 2 files got the same sequenceID.","from":"developer"},{"body":"Patch to remove the sequenceID check.","from":"developer"},{"body":"Can you attach patch for trunk ?\n\nThanks","from":"developer"},{"body":"Hi Ted,\n\nI have already attached the patch. I did it on my eclipse which is configured with the trunk. So it should be correct? Or have I done anything wrong?\n\nThanks.","from":"developer"},{"body":"From your patch:\n{code}\nIndex: src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n{code}\nIn trunk, the path would be hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\nCan you use 'svn diff' command to produce next patch under the root of your workspace ?","from":"developer"},{"body":"Oh, sorry about that! Did not figured I missed that.\n\nDebian come with VSN < 1.7 so this command line was not working for me. I installed 1.7, run it and got the patch for all the modifications I made. So attached is the extract for HBASE-1212. Can I mark HBASE-7287 as resolved by this one?","from":"developer"},{"body":"Change looks good to me. Taking a cursory look at the code we're indeed not using the sequenceId solely to distinguish HFiles.\nCan somebody reconfirm that.\n","from":"developer"},{"body":"Re-submitting the refreshed patch.","from":"developer"},{"body":"We will use sequenceId to sort HFiles, like in compaction.\nWhen merging, there is no overlap between these two files from different regions, I think it's ok even if two files have the same sequence ID since we don't use it to distingush the files in Store any more \n\nPatch looks good for me","from":"developer"},{"body":"+1 from me.","from":"developer"},{"body":"Should I upload one for 0.94?","from":"developer"},{"body":"Integrated to trunk.\n\nThanks for the patch, Jean-Marc.\n\nThanks for the review, Chunhui.","from":"developer"},{"body":"Was integrated a while back. Resolving.","from":"developer"}],"created":"2009-02-23T22:30:11.000+0000","description":"Currently merging two regions, the merge tool will compare their sequence ids. If same, it will decrement one. It needs to do this because on region open, files are keyed by their sequenceid; if two the same, one will erase the other.\n\nWell, with the move to the aggregating hfile format, the sequenceid is written when the file is created and its no longer written into an aside file but as metadata on to the end of the file. Changing the sequenceid is no longer an option.\n\nThis issue is about figuring a solution for the rare case where two store files have same sequence id AND we want to merge the two regions.","issue_id":"12415413","key":"HBASE-1212","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2013-04-23T06:33:31.000+0000","role":"fixed_distractor","summary":"merge tool expects regions all have different sequence ids"} {"case_id":"12747101","cluster":"DISTRACTOR-HBASE-12219","comments":[{"body":"With the TableStateManager from HBASE-7767 this shouldn't be a problem, since the table list of tables can be accessed directly from memory.","created":"2014-10-09T19:45:57.977+0000"},{"body":"bq. With the TableStateManager from HBASE-7767 this shouldn't be a problem,\n\nThat won't be a solution for 0.98. \n\nI don't think a TTL based cache will be the right solution if it might serve stale descriptors that have been changed on disk but haven't expired yet. What about maintaining a cache that retains entries indefinitely and is also updated in place by the other master functions that deal with schema changes?","created":"2014-10-10T17:21:05.405+0000"},{"body":"Thats a good idea [~apurtell]. I'm not really concerned about the staleness if the TTL is as short as 1 second which still make a huge difference in the throughput. For now I'm going to modify my patch to address your concerns and let me know what you think.","created":"2014-10-10T18:31:40.146+0000"},{"body":"bq. I'm not really concerned about the staleness if the TTL is as short as 1 second \n\nThen definitely the suggested alternative could make a big difference in the number of NN ops totally, if descriptors for a table rarely or never change.","created":"2014-10-10T18:36:25.027+0000"},{"body":"This patch addresses [~apurtell] suggestions.\n","created":"2014-10-17T16:48:05.679+0000"},{"body":"Changed summary to reflect this no longer requires to configure a TTL to refresh the cached entries. Also aded a new property to load the HTDs at startup of the master (enabled by default) {{hbase.master.preload.tabledescriptors}}\n","created":"2014-10-17T16:51:22.202+0000"},{"body":"https://reviews.apache.org/r/26878","created":"2014-10-17T17:06:50.956+0000"},{"body":"Y-axis requests normalized, X-axis time in milliseconds as measured from the client. Test done in a MBP @2.3GHz in pseudo mode with HBase defaults, 2k tables and tables added continuously during the test.","created":"2014-10-17T17:10:59.753+0000"},{"body":"Can you update the patch to be based on master? It looks like the same NN reading code is in the cache there.","created":"2014-10-17T17:21:28.135+0000"},{"body":"Tests should pass once HBASE-12380 is committed.\n","created":"2014-10-30T20:03:25.445+0000"},{"body":"Reapply so get another hadoopqa run ","created":"2014-10-30T21:26:37.352+0000"},{"body":"new patch, fixed some nits.","created":"2014-10-30T21:46:44.485+0000"},{"body":"v2 failed after turning off the cache in the the FSTableDescriptors constructor. I think it should be fine to use v1 instead, will upload new patch.","created":"2014-10-31T00:13:25.072+0000"},{"body":"Test failure not related, see HBASE-11819.","created":"2014-10-31T03:32:37.828+0000"},{"body":"Applied to master. You want to make a 0.99 patch [~esteban] (it didn't cherry-pick nicely... lots 'off')","created":"2014-10-31T04:14:12.676+0000"},{"body":"Attaching patch for 0.99. Is that good for you [~enis]?","created":"2014-10-31T08:56:02.176+0000"},{"body":"Cancelled 0.99 patch for now, it was consistent but had some formatting issues.","created":"2014-10-31T09:52:21.774+0000"},{"body":"Pushed the branch-1 patch (Pushed the master patch yesterday). Thanks [~esteban]","created":"2014-10-31T17:53:39.709+0000"},{"body":"Missing change.","created":"2014-10-31T18:23:14.690+0000"},{"body":"Applied addendum. Thanks [~esteban]","created":"2014-10-31T18:34:01.690+0000"},{"body":"Attached patch for 0.98 cc:[~apurtell] ","created":"2014-10-31T18:34:57.257+0000"},{"body":"0.98 patch doesn't apply cleanly:\n{noformat}\npatching file hbase-server/src/main/java/org/apache/hadoop/hbase/TableDescriptors.java\npatching file hbase-server/src/main/java/org/apache/hadoop/hbase/master/HMaster.java\nHunk #1 succeeded at 362 (offset 3 lines).\nHunk #2 succeeded at 489 (offset 3 lines).\nHunk #3 succeeded at 811 with fuzz 1 (offset 3 lines).\npatching file hbase-server/src/main/java/org/apache/hadoop/hbase/master/handler/CreateTableHandler.java\npatching file hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java\nHunk #4 FAILED at 1309.\n1 out of 10 hunks FAILED -- saving rejects to file hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java.rej\npatching file hbase-server/src/main/java/org/apache/hadoop/hbase/util/FSTableDescriptors.java\nHunk #3 FAILED at 88.\nHunk #4 FAILED at 120.\nHunk #5 FAILED at 130.\nHunk #6 FAILED at 147.\nHunk #7 succeeded at 197 (offset 7 lines).\nHunk #8 succeeded at 270 (offset 7 lines).\nHunk #9 succeeded at 289 (offset 7 lines).\nHunk #10 succeeded at 307 (offset 7 lines).\nHunk #11 succeeded at 321 (offset 7 lines).\nHunk #12 succeeded at 337 (offset 7 lines).\nHunk #13 succeeded at 356 (offset 7 lines).\nHunk #14 succeeded at 396 (offset 7 lines).\nHunk #15 succeeded at 421 (offset 7 lines).\nHunk #16 succeeded at 465 (offset 7 lines).\nHunk #17 succeeded at 473 (offset 7 lines).\nHunk #18 FAILED at 481.\nHunk #19 succeeded at 541 (offset 7 lines).\nHunk #20 succeeded at 581 (offset 7 lines).\nHunk #21 succeeded at 596 (offset 7 lines).\nHunk #22 succeeded at 622 (offset 7 lines).\nHunk #23 succeeded at 648 (offset 7 lines).\nHunk #24 succeeded at 686 (offset 7 lines).\nHunk #25 succeeded at 712 (offset 7 lines).\nHunk #26 succeeded at 720 (offset 7 lines).\nHunk #27 succeeded at 752 (offset 7 lines).\n5 out of 27 hunks FAILED -- saving rejects to file hbase-server/src/main/java/org/apache/hadoop/hbase/util/FSTableDescriptors.java.rej\npatching file hbase-server/src/test/java/org/apache/hadoop/hbase/master/TestCatalogJanitor.java\npatching file hbase-server/src/test/java/org/apache/hadoop/hbase/util/TestFSTableDescriptors.java\nHunk #3 succeeded at 395 (offset 1 line).\n{noformat}\n","created":"2014-10-31T20:22:08.002+0000"},{"body":"I see, I made the patch in a branch that was 29 commits behind. Let me fix that for you.","created":"2014-10-31T22:10:04.200+0000"},{"body":"Try this one [~apurtell] my apologies for this mishap .","created":"2014-10-31T23:18:19.982+0000"},{"body":"Thanks [~esteban], that's better. Pushed to 0.98 ","created":"2014-10-31T23:43:30.900+0000"},{"body":"Reverted branch-1 patch and addendum. Builds are unstable starting w/ this patch going in. I'm reverting till build is back to stable again then will put stuff back.","created":"2014-11-01T23:04:49.762+0000"},{"body":"Builds on branch-1 are blue again after backing this out. I think this the zombie maker. Leaving open till we figure why.","created":"2014-11-02T03:32:38.519+0000"},{"body":"Using [~manukranthk]'s awesome findHangingTests script, it looks like the set of runs that were red all had org.apache.hadoop.hbase.client.TestAdmin hang, which caused the Surefire-forked process to time out after 15 minutes and fail the Maven build.","created":"2014-11-02T04:55:53.469+0000"},{"body":"I disabled forked process timeouts and reran TestAdmin before [~stack]'s commit, which revealed this:\n{code}\ntestTruncateTablePreservingSplits(org.apache.hadoop.hbase.client.TestAdmin) Time elapsed: 242.311 sec <<< ERROR!\norg.apache.hadoop.hbase.client.RetriesExhaustedException: Failed after attempts=7, exceptions:\nSat Nov 01 22:54:33 PDT 2014, null, java.net.SocketTimeoutException: callTimeout=60000, callDuration=291372: row '' on table 'testTruncateTablePreservingSplits\n\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.throwEnrichedException(RpcRetryingCallerWithReadReplicas.java:261)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas.call(ScannerCallableWithReplicas.java:199)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas.call(ScannerCallableWithReplicas.java:56)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithoutRetries(RpcRetryingCaller.java:196)\n\tat org.apache.hadoop.hbase.client.ClientScanner.call(ClientScanner.java:287)\n\tat org.apache.hadoop.hbase.client.ClientScanner.nextScanner(ClientScanner.java:267)\n\tat org.apache.hadoop.hbase.client.ClientScanner.initializeScannerInConstruction(ClientScanner.java:139)\n\tat org.apache.hadoop.hbase.client.ClientScanner.(ClientScanner.java:134)\n\tat org.apache.hadoop.hbase.client.HTable.getScanner(HTable.java:789)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.countRows(HBaseTestingUtility.java:1894)\n\tat org.apache.hadoop.hbase.client.TestAdmin.testTruncateTable(TestAdmin.java:393)\n\tat org.apache.hadoop.hbase.client.TestAdmin.testTruncateTablePreservingSplits(TestAdmin.java:369)\nCaused by: java.net.SocketTimeoutException: callTimeout=60000, callDuration=291372: row '' on table 'testTruncateTablePreservingSplits\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:155)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.hadoop.hbase.client.RetriesExhaustedException: Can't get the location\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.getRegionLocations(RpcRetryingCallerWithReadReplicas.java:299)\n\tat org.apache.hadoop.hbase.client.ScannerCallable.prepare(ScannerCallable.java:136)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:121)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.hadoop.hbase.client.NoServerForRegionException: No server address listed in hbase:meta for region testTruncateTablePreservingSplits,,1414907431375.c73f9cd53b251842620244ea108213d5. containing row \n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.locateRegionInMeta(ConnectionManager.java:1227)\n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.locateRegion(ConnectionManager.java:1093)\n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.relocateRegion(ConnectionManager.java:1064)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.getRegionLocations(RpcRetryingCallerWithReadReplicas.java:288)\n\tat org.apache.hadoop.hbase.client.ScannerCallable.prepare(ScannerCallable.java:136)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:121)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n\ntestTruncateTable(org.apache.hadoop.hbase.client.TestAdmin) Time elapsed: 242.6 sec <<< ERROR!\norg.apache.hadoop.hbase.client.RetriesExhaustedException: Failed after attempts=7, exceptions:\nSat Nov 01 23:02:52 PDT 2014, null, java.net.SocketTimeoutException: callTimeout=60000, callDuration=292098: row '' on table 'testTruncateTable\n\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.throwEnrichedException(RpcRetryingCallerWithReadReplicas.java:261)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas.call(ScannerCallableWithReplicas.java:199)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas.call(ScannerCallableWithReplicas.java:56)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithoutRetries(RpcRetryingCaller.java:196)\n\tat org.apache.hadoop.hbase.client.ClientScanner.call(ClientScanner.java:287)\n\tat org.apache.hadoop.hbase.client.ClientScanner.nextScanner(ClientScanner.java:267)\n\tat org.apache.hadoop.hbase.client.ClientScanner.initializeScannerInConstruction(ClientScanner.java:139)\n\tat org.apache.hadoop.hbase.client.ClientScanner.(ClientScanner.java:134)\n\tat org.apache.hadoop.hbase.client.HTable.getScanner(HTable.java:789)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.countRows(HBaseTestingUtility.java:1894)\n\tat org.apache.hadoop.hbase.client.TestAdmin.testTruncateTable(TestAdmin.java:393)\n\tat org.apache.hadoop.hbase.client.TestAdmin.testTruncateTable(TestAdmin.java:364)\nCaused by: java.net.SocketTimeoutException: callTimeout=60000, callDuration=292098: row '' on table 'testTruncateTable\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:155)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.hadoop.hbase.client.RetriesExhaustedException: Can't get the location\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.getRegionLocations(RpcRetryingCallerWithReadReplicas.java:299)\n\tat org.apache.hadoop.hbase.client.ScannerCallable.prepare(ScannerCallable.java:136)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:121)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.hadoop.hbase.client.NoServerForRegionException: No server address listed in hbase:meta for region testTruncateTable,,1414907930959.4e6ea2189874de8fc4b408137fa5d38f. containing row \n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.locateRegionInMeta(ConnectionManager.java:1227)\n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.locateRegion(ConnectionManager.java:1093)\n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.relocateRegion(ConnectionManager.java:1064)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.getRegionLocations(RpcRetryingCallerWithReadReplicas.java:288)\n\tat org.apache.hadoop.hbase.client.ScannerCallable.prepare(ScannerCallable.java:136)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:121)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}\n\nI also ran {{jstack -F}} on the process; let me know if you want me to put that up on pastebin, [~esteban].","created":"2014-11-02T06:07:08.501+0000"},{"body":"Good digging [~dimaspivak] How you disable? Setting in pom? You ran locally or up on Apache. Should we have a config which disables forking when we have a zombie amok?","created":"2014-11-02T23:18:44.568+0000"},{"body":"I ran on my internal rig (don't have an Apache account) and added {{-Dsurefire.timeout=0}} to the mvn command which set the process timeout to unlimited. Can definitely have a Jenkins parameter in the builds job that does that when we notice that the build has gone red because of zombies.","created":"2014-11-02T23:28:36.091+0000"},{"body":"The timeout setting above is a nice tip. When hunting zombies locally you can also set the first part and second part fork modes to \"always\" so each test runs in its own VM, then loop the unit test suite, watch for stragglers, then jstack. Works well because you can be sure every stack in the dump is relevant for the hung test. ","created":"2014-11-02T23:41:20.428+0000"},{"body":"Looks like truncation isn't working now in 0.98. I was complaining on the wrong issue before, see https://issues.apache.org/jira/browse/HBASE-12142?focusedCommentId=14195604&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-14195604 for log detail and steps to reproduce.","created":"2014-11-04T03:34:39.461+0000"},{"body":"The issue was a mismatch how truncate table should work after HBASE-7767. Both HBASE-8332 and HBASE-12142 use {{tempdir}} instead of {{tempTableDir}}:\n\n{code}\n Path tempTableDir = FSUtils.getTableDir(tempdir, this.tableName);\n new FSTableDescriptors(server.getConfiguration())\n .createTableDescriptorForTableDirectory(tempTableDir, getTableDescriptor(), false);\n{code}\n\nWhich is the correct behavior from HBASE-7767 instead of FSTD. createTableDescriptorForTableDirectory(tempdir).\n\nThanks for [~mbertozzi] for the brainstorming to understand where this issue came from.\n","created":"2014-11-04T11:06:05.606+0000"},{"body":"I'm planning to roll the 0.98.8 RC0 this Friday. We'll need a fix for truncate issues on 0.98 branch before then to avoid a revert of this change. Thanks!","created":"2014-11-04T20:19:59.610+0000"},{"body":"[~apurtell] the addendum for 0.98 should solve the original problem, can you give it a try?","created":"2014-11-04T23:00:50.817+0000"},{"body":"I applied HBASE-12219-0.99.v1.patch Lets see if branch-1 stays stable. If so, will apply 0.98.","created":"2014-11-04T23:11:05.716+0000"},{"body":"I applied the addendum patch to 0.98. A quick test with the minicluster and shell truncate command looks ok. TestAdmin truncate tests pass. Please apply the addendum to 0.98 whenever ready [~stack]","created":"2014-11-04T23:25:03.961+0000"},{"body":"branch-1 build looks good.\n\nI applied the 0.98 addendum presuming branch-1 good.\n\nThanks for the patches and the perseverence [~esteban]","created":"2014-11-05T00:17:58.824+0000"},{"body":"[~larsh] tried this on 0.94 and seems good, do you want to give it a try too?","created":"2014-11-07T00:15:26.425+0000"},{"body":"Closing this issue after 0.99.2 release.","created":"2015-02-21T23:45:41.180+0000"},{"body":"Sorry I missed this JIRA at the time, but I have a couple of concerns if I'm understanding this change correctly. I'd like to check to see if I am. It looks like the effect is to back out the directory modtime caching and instead have the master just maintain a persistent in memory cache.\n\nPrevious to this change it was safe to have any processes on the cluster use FSTableDescriptors to read or write table descriptors. Updates would be atomic and consistent, immediately available to all readers. However, the cost of that was a HDFS NN operation on every table descriptor read to prove it had not changed. Since in practice it appears that only the master process ever updates an existing table descriptor, it should be safe to have the master skip the directory modtime checks and proactively update its cached copy. (hbck can also create table descriptors for orphaned tables but hopefully those don't happen to tables the master has already cached).\n\nThis change makes the master descriptor reads faster but imposes the constraint that only the active master should update table descriptors - any other writers would cause the master cache to become stale. It also means that no other processes should use the cache the same way as the master could change the data and cause stale caches. Assuming this is the case, I think we'd be better served by reflecting that in the FSTableDescriptors API and javadoc. For example, currently most constructors and usages now default to keeping a persistent cache as well as allowing updates which sets a bad example for new uses. There are also no warnings in the javadoc about the new contract. Possibly better would be to make the default constructor be read only and have persistent caching disabled. Then another constructor for the master allowing both writes and persistent caching.\n\nThis change also seems to remove all table descriptor caching from the region servers (the old directory modtime caching is gone and the new caching is disabled for region servers). Thanks to HBASE-8778 reloading from the FS each time is cheaper than it used to be, but this change still increases the cost from 1 NN operation (check directory modtime) to 2 NN + 3 DN operations (find current file, get its block locations, open block, read close block). This slows things down a bit again for mass assignments/balances on huge tables. It seems better for the region servers to retain the directory modtime caching, but simply skip the modtime check when running inside the master.\n\nDoes that understanding of this change sound correct - or did I botch it? Sorry I missed it at the time. If that sounds right, a follow up JIRA may be good, and if I see our table assignments slower from this and no one else gets to it I can try to put up the changes.","created":"2015-07-31T06:21:53.771+0000"}],"conversations":[{"body":"Currently table descriptors and tables are cached once they are accessed for the first time. Next calls to the master only require a trip to HDFS to lookup the modified time in order to reload the table descriptors if modified. However in clusters with a large number of tables or concurrent clients and this can be too aggressive to HDFS and the master causing contention to process other requests. A simple solution is to have a TTL based cached for FSTableDescriptors#getAll() and FSTableDescriptors#TableDescriptorAndModtime() that can allow the master to process those calls faster without causing contention without having to perform a trip to HDFS for every call. to listtables() or getTableDescriptor()","from":"reporter","subject":"Cache more efficiently getAll() and get() in FSTableDescriptors"},{"body":"With the TableStateManager from HBASE-7767 this shouldn't be a problem, since the table list of tables can be accessed directly from memory.","from":"developer"},{"body":"bq. With the TableStateManager from HBASE-7767 this shouldn't be a problem,\n\nThat won't be a solution for 0.98. \n\nI don't think a TTL based cache will be the right solution if it might serve stale descriptors that have been changed on disk but haven't expired yet. What about maintaining a cache that retains entries indefinitely and is also updated in place by the other master functions that deal with schema changes?","from":"developer"},{"body":"Thats a good idea [~apurtell]. I'm not really concerned about the staleness if the TTL is as short as 1 second which still make a huge difference in the throughput. For now I'm going to modify my patch to address your concerns and let me know what you think.","from":"developer"},{"body":"bq. I'm not really concerned about the staleness if the TTL is as short as 1 second \n\nThen definitely the suggested alternative could make a big difference in the number of NN ops totally, if descriptors for a table rarely or never change.","from":"developer"},{"body":"This patch addresses [~apurtell] suggestions.\n","from":"developer"},{"body":"Changed summary to reflect this no longer requires to configure a TTL to refresh the cached entries. Also aded a new property to load the HTDs at startup of the master (enabled by default) {{hbase.master.preload.tabledescriptors}}\n","from":"developer"},{"body":"https://reviews.apache.org/r/26878","from":"developer"},{"body":"Y-axis requests normalized, X-axis time in milliseconds as measured from the client. Test done in a MBP @2.3GHz in pseudo mode with HBase defaults, 2k tables and tables added continuously during the test.","from":"developer"},{"body":"Can you update the patch to be based on master? It looks like the same NN reading code is in the cache there.","from":"developer"},{"body":"Tests should pass once HBASE-12380 is committed.\n","from":"developer"},{"body":"Reapply so get another hadoopqa run ","from":"developer"},{"body":"new patch, fixed some nits.","from":"developer"},{"body":"v2 failed after turning off the cache in the the FSTableDescriptors constructor. I think it should be fine to use v1 instead, will upload new patch.","from":"developer"},{"body":"Test failure not related, see HBASE-11819.","from":"developer"},{"body":"Applied to master. You want to make a 0.99 patch [~esteban] (it didn't cherry-pick nicely... lots 'off')","from":"developer"},{"body":"Attaching patch for 0.99. Is that good for you [~enis]?","from":"developer"},{"body":"Cancelled 0.99 patch for now, it was consistent but had some formatting issues.","from":"developer"},{"body":"Pushed the branch-1 patch (Pushed the master patch yesterday). Thanks [~esteban]","from":"developer"},{"body":"Missing change.","from":"developer"},{"body":"Applied addendum. Thanks [~esteban]","from":"developer"},{"body":"Attached patch for 0.98 cc:[~apurtell] ","from":"developer"},{"body":"0.98 patch doesn't apply cleanly:\n{noformat}\npatching file hbase-server/src/main/java/org/apache/hadoop/hbase/TableDescriptors.java\npatching file hbase-server/src/main/java/org/apache/hadoop/hbase/master/HMaster.java\nHunk #1 succeeded at 362 (offset 3 lines).\nHunk #2 succeeded at 489 (offset 3 lines).\nHunk #3 succeeded at 811 with fuzz 1 (offset 3 lines).\npatching file hbase-server/src/main/java/org/apache/hadoop/hbase/master/handler/CreateTableHandler.java\npatching file hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java\nHunk #4 FAILED at 1309.\n1 out of 10 hunks FAILED -- saving rejects to file hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java.rej\npatching file hbase-server/src/main/java/org/apache/hadoop/hbase/util/FSTableDescriptors.java\nHunk #3 FAILED at 88.\nHunk #4 FAILED at 120.\nHunk #5 FAILED at 130.\nHunk #6 FAILED at 147.\nHunk #7 succeeded at 197 (offset 7 lines).\nHunk #8 succeeded at 270 (offset 7 lines).\nHunk #9 succeeded at 289 (offset 7 lines).\nHunk #10 succeeded at 307 (offset 7 lines).\nHunk #11 succeeded at 321 (offset 7 lines).\nHunk #12 succeeded at 337 (offset 7 lines).\nHunk #13 succeeded at 356 (offset 7 lines).\nHunk #14 succeeded at 396 (offset 7 lines).\nHunk #15 succeeded at 421 (offset 7 lines).\nHunk #16 succeeded at 465 (offset 7 lines).\nHunk #17 succeeded at 473 (offset 7 lines).\nHunk #18 FAILED at 481.\nHunk #19 succeeded at 541 (offset 7 lines).\nHunk #20 succeeded at 581 (offset 7 lines).\nHunk #21 succeeded at 596 (offset 7 lines).\nHunk #22 succeeded at 622 (offset 7 lines).\nHunk #23 succeeded at 648 (offset 7 lines).\nHunk #24 succeeded at 686 (offset 7 lines).\nHunk #25 succeeded at 712 (offset 7 lines).\nHunk #26 succeeded at 720 (offset 7 lines).\nHunk #27 succeeded at 752 (offset 7 lines).\n5 out of 27 hunks FAILED -- saving rejects to file hbase-server/src/main/java/org/apache/hadoop/hbase/util/FSTableDescriptors.java.rej\npatching file hbase-server/src/test/java/org/apache/hadoop/hbase/master/TestCatalogJanitor.java\npatching file hbase-server/src/test/java/org/apache/hadoop/hbase/util/TestFSTableDescriptors.java\nHunk #3 succeeded at 395 (offset 1 line).\n{noformat}\n","from":"developer"},{"body":"I see, I made the patch in a branch that was 29 commits behind. Let me fix that for you.","from":"developer"},{"body":"Try this one [~apurtell] my apologies for this mishap .","from":"developer"},{"body":"Thanks [~esteban], that's better. Pushed to 0.98 ","from":"developer"},{"body":"Reverted branch-1 patch and addendum. Builds are unstable starting w/ this patch going in. I'm reverting till build is back to stable again then will put stuff back.","from":"developer"},{"body":"Builds on branch-1 are blue again after backing this out. I think this the zombie maker. Leaving open till we figure why.","from":"developer"},{"body":"Using [~manukranthk]'s awesome findHangingTests script, it looks like the set of runs that were red all had org.apache.hadoop.hbase.client.TestAdmin hang, which caused the Surefire-forked process to time out after 15 minutes and fail the Maven build.","from":"developer"},{"body":"I disabled forked process timeouts and reran TestAdmin before [~stack]'s commit, which revealed this:\n{code}\ntestTruncateTablePreservingSplits(org.apache.hadoop.hbase.client.TestAdmin) Time elapsed: 242.311 sec <<< ERROR!\norg.apache.hadoop.hbase.client.RetriesExhaustedException: Failed after attempts=7, exceptions:\nSat Nov 01 22:54:33 PDT 2014, null, java.net.SocketTimeoutException: callTimeout=60000, callDuration=291372: row '' on table 'testTruncateTablePreservingSplits\n\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.throwEnrichedException(RpcRetryingCallerWithReadReplicas.java:261)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas.call(ScannerCallableWithReplicas.java:199)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas.call(ScannerCallableWithReplicas.java:56)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithoutRetries(RpcRetryingCaller.java:196)\n\tat org.apache.hadoop.hbase.client.ClientScanner.call(ClientScanner.java:287)\n\tat org.apache.hadoop.hbase.client.ClientScanner.nextScanner(ClientScanner.java:267)\n\tat org.apache.hadoop.hbase.client.ClientScanner.initializeScannerInConstruction(ClientScanner.java:139)\n\tat org.apache.hadoop.hbase.client.ClientScanner.(ClientScanner.java:134)\n\tat org.apache.hadoop.hbase.client.HTable.getScanner(HTable.java:789)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.countRows(HBaseTestingUtility.java:1894)\n\tat org.apache.hadoop.hbase.client.TestAdmin.testTruncateTable(TestAdmin.java:393)\n\tat org.apache.hadoop.hbase.client.TestAdmin.testTruncateTablePreservingSplits(TestAdmin.java:369)\nCaused by: java.net.SocketTimeoutException: callTimeout=60000, callDuration=291372: row '' on table 'testTruncateTablePreservingSplits\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:155)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.hadoop.hbase.client.RetriesExhaustedException: Can't get the location\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.getRegionLocations(RpcRetryingCallerWithReadReplicas.java:299)\n\tat org.apache.hadoop.hbase.client.ScannerCallable.prepare(ScannerCallable.java:136)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:121)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.hadoop.hbase.client.NoServerForRegionException: No server address listed in hbase:meta for region testTruncateTablePreservingSplits,,1414907431375.c73f9cd53b251842620244ea108213d5. containing row \n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.locateRegionInMeta(ConnectionManager.java:1227)\n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.locateRegion(ConnectionManager.java:1093)\n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.relocateRegion(ConnectionManager.java:1064)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.getRegionLocations(RpcRetryingCallerWithReadReplicas.java:288)\n\tat org.apache.hadoop.hbase.client.ScannerCallable.prepare(ScannerCallable.java:136)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:121)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n\ntestTruncateTable(org.apache.hadoop.hbase.client.TestAdmin) Time elapsed: 242.6 sec <<< ERROR!\norg.apache.hadoop.hbase.client.RetriesExhaustedException: Failed after attempts=7, exceptions:\nSat Nov 01 23:02:52 PDT 2014, null, java.net.SocketTimeoutException: callTimeout=60000, callDuration=292098: row '' on table 'testTruncateTable\n\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.throwEnrichedException(RpcRetryingCallerWithReadReplicas.java:261)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas.call(ScannerCallableWithReplicas.java:199)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas.call(ScannerCallableWithReplicas.java:56)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithoutRetries(RpcRetryingCaller.java:196)\n\tat org.apache.hadoop.hbase.client.ClientScanner.call(ClientScanner.java:287)\n\tat org.apache.hadoop.hbase.client.ClientScanner.nextScanner(ClientScanner.java:267)\n\tat org.apache.hadoop.hbase.client.ClientScanner.initializeScannerInConstruction(ClientScanner.java:139)\n\tat org.apache.hadoop.hbase.client.ClientScanner.(ClientScanner.java:134)\n\tat org.apache.hadoop.hbase.client.HTable.getScanner(HTable.java:789)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.countRows(HBaseTestingUtility.java:1894)\n\tat org.apache.hadoop.hbase.client.TestAdmin.testTruncateTable(TestAdmin.java:393)\n\tat org.apache.hadoop.hbase.client.TestAdmin.testTruncateTable(TestAdmin.java:364)\nCaused by: java.net.SocketTimeoutException: callTimeout=60000, callDuration=292098: row '' on table 'testTruncateTable\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:155)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.hadoop.hbase.client.RetriesExhaustedException: Can't get the location\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.getRegionLocations(RpcRetryingCallerWithReadReplicas.java:299)\n\tat org.apache.hadoop.hbase.client.ScannerCallable.prepare(ScannerCallable.java:136)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:121)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.hadoop.hbase.client.NoServerForRegionException: No server address listed in hbase:meta for region testTruncateTable,,1414907930959.4e6ea2189874de8fc4b408137fa5d38f. containing row \n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.locateRegionInMeta(ConnectionManager.java:1227)\n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.locateRegion(ConnectionManager.java:1093)\n\tat org.apache.hadoop.hbase.client.ConnectionManager$HConnectionImplementation.relocateRegion(ConnectionManager.java:1064)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCallerWithReadReplicas.getRegionLocations(RpcRetryingCallerWithReadReplicas.java:288)\n\tat org.apache.hadoop.hbase.client.ScannerCallable.prepare(ScannerCallable.java:136)\n\tat org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:121)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:294)\n\tat org.apache.hadoop.hbase.client.ScannerCallableWithReplicas$RetryingRPC.call(ScannerCallableWithReplicas.java:275)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}\n\nI also ran {{jstack -F}} on the process; let me know if you want me to put that up on pastebin, [~esteban].","from":"developer"},{"body":"Good digging [~dimaspivak] How you disable? Setting in pom? You ran locally or up on Apache. Should we have a config which disables forking when we have a zombie amok?","from":"developer"},{"body":"I ran on my internal rig (don't have an Apache account) and added {{-Dsurefire.timeout=0}} to the mvn command which set the process timeout to unlimited. Can definitely have a Jenkins parameter in the builds job that does that when we notice that the build has gone red because of zombies.","from":"developer"},{"body":"The timeout setting above is a nice tip. When hunting zombies locally you can also set the first part and second part fork modes to \"always\" so each test runs in its own VM, then loop the unit test suite, watch for stragglers, then jstack. Works well because you can be sure every stack in the dump is relevant for the hung test. ","from":"developer"},{"body":"Looks like truncation isn't working now in 0.98. I was complaining on the wrong issue before, see https://issues.apache.org/jira/browse/HBASE-12142?focusedCommentId=14195604&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-14195604 for log detail and steps to reproduce.","from":"developer"},{"body":"The issue was a mismatch how truncate table should work after HBASE-7767. Both HBASE-8332 and HBASE-12142 use {{tempdir}} instead of {{tempTableDir}}:\n\n{code}\n Path tempTableDir = FSUtils.getTableDir(tempdir, this.tableName);\n new FSTableDescriptors(server.getConfiguration())\n .createTableDescriptorForTableDirectory(tempTableDir, getTableDescriptor(), false);\n{code}\n\nWhich is the correct behavior from HBASE-7767 instead of FSTD. createTableDescriptorForTableDirectory(tempdir).\n\nThanks for [~mbertozzi] for the brainstorming to understand where this issue came from.\n","from":"developer"},{"body":"I'm planning to roll the 0.98.8 RC0 this Friday. We'll need a fix for truncate issues on 0.98 branch before then to avoid a revert of this change. Thanks!","from":"developer"},{"body":"[~apurtell] the addendum for 0.98 should solve the original problem, can you give it a try?","from":"developer"},{"body":"I applied HBASE-12219-0.99.v1.patch Lets see if branch-1 stays stable. If so, will apply 0.98.","from":"developer"},{"body":"I applied the addendum patch to 0.98. A quick test with the minicluster and shell truncate command looks ok. TestAdmin truncate tests pass. Please apply the addendum to 0.98 whenever ready [~stack]","from":"developer"},{"body":"branch-1 build looks good.\n\nI applied the 0.98 addendum presuming branch-1 good.\n\nThanks for the patches and the perseverence [~esteban]","from":"developer"},{"body":"[~larsh] tried this on 0.94 and seems good, do you want to give it a try too?","from":"developer"},{"body":"Closing this issue after 0.99.2 release.","from":"developer"},{"body":"Sorry I missed this JIRA at the time, but I have a couple of concerns if I'm understanding this change correctly. I'd like to check to see if I am. It looks like the effect is to back out the directory modtime caching and instead have the master just maintain a persistent in memory cache.\n\nPrevious to this change it was safe to have any processes on the cluster use FSTableDescriptors to read or write table descriptors. Updates would be atomic and consistent, immediately available to all readers. However, the cost of that was a HDFS NN operation on every table descriptor read to prove it had not changed. Since in practice it appears that only the master process ever updates an existing table descriptor, it should be safe to have the master skip the directory modtime checks and proactively update its cached copy. (hbck can also create table descriptors for orphaned tables but hopefully those don't happen to tables the master has already cached).\n\nThis change makes the master descriptor reads faster but imposes the constraint that only the active master should update table descriptors - any other writers would cause the master cache to become stale. It also means that no other processes should use the cache the same way as the master could change the data and cause stale caches. Assuming this is the case, I think we'd be better served by reflecting that in the FSTableDescriptors API and javadoc. For example, currently most constructors and usages now default to keeping a persistent cache as well as allowing updates which sets a bad example for new uses. There are also no warnings in the javadoc about the new contract. Possibly better would be to make the default constructor be read only and have persistent caching disabled. Then another constructor for the master allowing both writes and persistent caching.\n\nThis change also seems to remove all table descriptor caching from the region servers (the old directory modtime caching is gone and the new caching is disabled for region servers). Thanks to HBASE-8778 reloading from the FS each time is cheaper than it used to be, but this change still increases the cost from 1 NN operation (check directory modtime) to 2 NN + 3 DN operations (find current file, get its block locations, open block, read close block). This slows things down a bit again for mass assignments/balances on huge tables. It seems better for the region servers to retain the directory modtime caching, but simply skip the modtime check when running inside the master.\n\nDoes that understanding of this change sound correct - or did I botch it? Sorry I missed it at the time. If that sounds right, a follow up JIRA may be good, and if I see our table assignments slower from this and no one else gets to it I can try to put up the changes.","from":"developer"}],"created":"2014-10-09T19:31:29.000+0000","description":"Currently table descriptors and tables are cached once they are accessed for the first time. Next calls to the master only require a trip to HDFS to lookup the modified time in order to reload the table descriptors if modified. However in clusters with a large number of tables or concurrent clients and this can be too aggressive to HDFS and the master causing contention to process other requests. A simple solution is to have a TTL based cached for FSTableDescriptors#getAll() and FSTableDescriptors#TableDescriptorAndModtime() that can allow the master to process those calls faster without causing contention without having to perform a trip to HDFS for every call. to listtables() or getTableDescriptor()","issue_id":"12747101","key":"HBASE-12219","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2014-11-05T00:17:58.000+0000","role":"fixed_distractor","summary":"Cache more efficiently getAll() and get() in FSTableDescriptors"} {"case_id":"12751777","cluster":"DISTRACTOR-HBASE-12386","comments":[{"body":"Looking at the code it seems that once the remote zk peers lookup fails, the refresh ts is updated and the return list of RS peers is empty.\n\nNext time org.apache.hadoop.hbase.replication.regionserver.ReplicationSinkManager does not retry the lookup on the next polling as the following condition is not met:\n{code:java}\nif (endpoint.getLastRegionServerUpdate() > this.lastUpdateToPeers) {\n LOG.info(\"Current list of sinks is out of date, updating\");\n chooseSinks();\n}\n{code}\n\nA fix would be to force a refresh when the list of peers is empty:\n{code:java}\nif (replicationPeers.getTimestampOfLastChangeToPeer(peerClusterId) > this.lastUpdateToPeers\n || sinks.isEmpty()) {\n LOG.info(\"Current list of sinks is out of date or empty, updating\");\n chooseSinks();\n}\n{code}\n\nNote that this is not reproducing in 0.94 where it seems the refresh is happening in this case.\n","created":"2014-10-30T20:27:00.646+0000"},{"body":"{code}\n+ if (endpoint.getLastRegionServerUpdate() > this.lastUpdateToPeers || sinks.isEmpty()) {\n+ LOG.info(\"Current list of sinks is out of date or empty, updating\");\n{code}\nIt would helpful if the condition (list out of date or empty) is stated clearly in the log message.","created":"2014-10-30T22:05:08.379+0000"},{"body":"{{\"Current list of sinks is out of date or empty, updating\"}} seems clear enough to me.\n\n+1 on patch.\n\nOne thing we have to think through is what happens when the slave cluster is down for a bit. We'd chose sinks again on each call. I think that's OK especially since we dialed down the retry interval to 5mins recently after a bit.\n\nAlso, we can still be a bad situation where RegionServers die and restart at the slave cluster, we could go down to a single RS at the peers before we try to choose sinks again. That's for another issue.","created":"2014-10-30T22:33:56.795+0000"},{"body":"Pushed to 0.98+ \nThanks for the patch [~amuraru]!","created":"2014-11-01T01:04:28.954+0000"},{"body":"Closing this issue after 0.99.2 release.","created":"2015-02-21T23:45:01.022+0000"},{"body":"Committed w/o a JIRA ID\r\n\r\ncommit 0505072c5182841ad1a28d798527c69bcc3348f0\r\nAuthor: Adrian Muraru \r\nDate: Thu Oct 30 23:50:02 2014 +0200\r\n\r\n Replication gets stuck following a transient zookeeper error to remote peer cluster\r\n\r\n Signed-off-by: Andrew Purtell \r\n\r\n","created":"2018-04-04T17:34:07.006+0000"},{"body":"Sorry about that.","created":"2018-04-04T18:56:56.915+0000"},{"body":"[~apurtell] No worries sir. I have done 10 for any one by anyone else. Just noting these facts in issue as I try to align JIRA and git for branch-2.","created":"2018-04-04T19:06:58.024+0000"}],"conversations":[{"body":"Following a transient ZK error replication gets stuck and remote peers are never updated.\n\nSource region servers are reporting continuously the following error in logs:\n\"No replication sinks are available\"\n\n","from":"reporter","subject":"Replication gets stuck following a transient zookeeper error to remote peer cluster"},{"body":"Looking at the code it seems that once the remote zk peers lookup fails, the refresh ts is updated and the return list of RS peers is empty.\n\nNext time org.apache.hadoop.hbase.replication.regionserver.ReplicationSinkManager does not retry the lookup on the next polling as the following condition is not met:\n{code:java}\nif (endpoint.getLastRegionServerUpdate() > this.lastUpdateToPeers) {\n LOG.info(\"Current list of sinks is out of date, updating\");\n chooseSinks();\n}\n{code}\n\nA fix would be to force a refresh when the list of peers is empty:\n{code:java}\nif (replicationPeers.getTimestampOfLastChangeToPeer(peerClusterId) > this.lastUpdateToPeers\n || sinks.isEmpty()) {\n LOG.info(\"Current list of sinks is out of date or empty, updating\");\n chooseSinks();\n}\n{code}\n\nNote that this is not reproducing in 0.94 where it seems the refresh is happening in this case.\n","from":"developer"},{"body":"{code}\n+ if (endpoint.getLastRegionServerUpdate() > this.lastUpdateToPeers || sinks.isEmpty()) {\n+ LOG.info(\"Current list of sinks is out of date or empty, updating\");\n{code}\nIt would helpful if the condition (list out of date or empty) is stated clearly in the log message.","from":"developer"},{"body":"{{\"Current list of sinks is out of date or empty, updating\"}} seems clear enough to me.\n\n+1 on patch.\n\nOne thing we have to think through is what happens when the slave cluster is down for a bit. We'd chose sinks again on each call. I think that's OK especially since we dialed down the retry interval to 5mins recently after a bit.\n\nAlso, we can still be a bad situation where RegionServers die and restart at the slave cluster, we could go down to a single RS at the peers before we try to choose sinks again. That's for another issue.","from":"developer"},{"body":"Pushed to 0.98+ \nThanks for the patch [~amuraru]!","from":"developer"},{"body":"Closing this issue after 0.99.2 release.","from":"developer"},{"body":"Committed w/o a JIRA ID\r\n\r\ncommit 0505072c5182841ad1a28d798527c69bcc3348f0\r\nAuthor: Adrian Muraru \r\nDate: Thu Oct 30 23:50:02 2014 +0200\r\n\r\n Replication gets stuck following a transient zookeeper error to remote peer cluster\r\n\r\n Signed-off-by: Andrew Purtell \r\n\r\n","from":"developer"},{"body":"Sorry about that.","from":"developer"},{"body":"[~apurtell] No worries sir. I have done 10 for any one by anyone else. Just noting these facts in issue as I try to align JIRA and git for branch-2.","from":"developer"}],"created":"2014-10-30T20:18:30.000+0000","description":"Following a transient ZK error replication gets stuck and remote peers are never updated.\n\nSource region servers are reporting continuously the following error in logs:\n\"No replication sinks are available\"\n\n","issue_id":"12751777","key":"HBASE-12386","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2014-11-01T01:04:28.000+0000","role":"fixed_distractor","summary":"Replication gets stuck following a transient zookeeper error to remote peer cluster"} {"case_id":"12754873","cluster":"DISTRACTOR-HBASE-12465","comments":[{"body":"This issue might be a user uses hbase tmp folder as Import tool temporary output folder while HBase will try to recreate(delete and then create) tmp folder during starts. Therefore it cause HMaster can't start.\n\n[~saint.ack@gmail.com] Do u think any error from checkTempDir inside HMaster#createInitialFileSystemLayout is fatal? If it's fatal then we don't need do any thing for the JIRA otherwise we catch the error/log it and move on.","created":"2014-12-10T02:42:53.590+0000"},{"body":"Ping [~saint.ack@gmail.com] any thoughts on this? Thanks.","created":"2014-12-12T01:12:55.955+0000"},{"body":"I can write up a little test [~jeffreyz]? I'm not sure whether an exception out of createInitialFileSystemLayout is fatal...","created":"2014-12-12T01:19:42.966+0000"},{"body":"[~saint.ack@gmail.com] [~jeffreyz] If an exception out of createInitialFileSystemLayout is not fatal, can we catch the exception and make sure the tmp directory exist and continue?","created":"2014-12-14T20:54:01.764+0000"},{"body":"Jeffrey: We ran into this issue on one of our clusters last week. Looking at your JIRA updates, I couldn't make out if you were able to figure out a fix. Would sharing our logs be of any help here? Thanks!","created":"2015-06-07T21:55:07.497+0000"},{"body":"If of use, Jeffrey, Alicia, et al. the errors we (Biju and Sudarshan) saw were:\n\nFatal to the running HMaster as well as any HMaster starting:\n{code}\n2015-06-06 02:08:52,374 ERROR org.apache.hadoop.hbase.backup.HFileArchiver: Failed to archive class org.apache.hadoop.hbase.backup.HFi\nleArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e/f2806ebb34ab493ebe4b623fac585776_SeqId_385900210_\n2015-06-06 02:08:52,374 WARN org.apache.hadoop.hbase.backup.HFileArchiver: Couldn't archive class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e/f2806ebb34ab493ebe4b623fac585776_SeqId_385900210_ into backup directory: hdfs://cluster1/hbase/archive/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e\n2015-06-06 02:08:52,374 WARN org.apache.hadoop.hbase.backup.HFileArchiver: Failed to complete archive of: [class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e/f2806ebb34ab493ebe4b623fac585776_SeqId_385900210_]. Those files are still in the original location, and they may slow down reads.\n2015-06-06 02:08:52,374 FATAL org.apache.hadoop.hbase.master.HMaster: Unhandled exception. Starting shutdown.\njava.io.IOException: Received error when attempting to archive files ([class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e]), cannot delete region directory.\n at org.apache.hadoop.hbase.backup.HFileArchiver.archiveRegion(HFileArchiver.java:148)\n at org.apache.hadoop.hbase.master.MasterFileSystem.checkTempDir(MasterFileSystem.java:503)\n at org.apache.hadoop.hbase.master.MasterFileSystem.createInitialFileSystemLayout(MasterFileSystem.java:149)\n at org.apache.hadoop.hbase.master.MasterFileSystem.(MasterFileSystem.java:127)\n at org.apache.hadoop.hbase.master.HMaster.finishInitialization(HMaster.java:789) \n at org.apache.hadoop.hbase.master.HMaster.run(HMaster.java:606)\n at java.lang.Thread.run(Thread.java:745) \n2015-06-06 02:08:52,375 INFO org.apache.hadoop.hbase.master.HMaster: Aborting\n2015-06-06 02:08:52,375 DEBUG org.apache.hadoop.hbase.master.HMaster: Stopping service threads\n[... more shutting down here ...]\n2015-06-06 02:28:15,583 DEBUG org.apache.hadoop.hbase.master.ActiveMasterManager: A master is now available\n2015-06-06 02:28:15,583 INFO org.apache.hadoop.hbase.master.ActiveMasterManager: Registered Active Master=cluster1-bcpc-r2n7.example.com,60000,1433572093679\n2015-06-06 02:28:15,588 INFO org.apache.hadoop.conf.Configuration.deprecation: fs.default.name is deprecated. Instead, use fs.defaultFS\n2015-06-06 02:28:15,968 INFO org.apache.hadoop.conf.Configuration.deprecation: hadoop.native.lib is deprecated. Instead, use io.native.lib.available\n2015-06-06 02:28:16,164 DEBUG org.apache.hadoop.hbase.util.FSTableDescriptors: Current tableInfoPath = hdfs://cluster1/hbase/data/hbase/meta/.tabledesc/.tableinfo.0000000001\n2015-06-06 02:28:16,195 DEBUG org.apache.hadoop.hbase.util.FSTableDescriptors: TableInfo already exists.. Skipping creation\n2015-06-06 02:28:17,850 DEBUG org.apache.hadoop.hbase.backup.HFileArchiver: ARCHIVING hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf\n2015-06-06 02:28:17,860 DEBUG org.apache.hadoop.hbase.backup.HFileArchiver: Archiving [class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e]\n2015-06-06 02:28:17,885 WARN org.apache.hadoop.hbase.backup.HFileArchiver: Failed to archive class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e/f2806ebb34ab493ebe4b623fac585776_SeqId_385900210_ on try #0\norg.apache.hadoop.security.AccessControlException: Permission denied: user=hbase, access=WRITE, inode=\"/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e/f2806ebb34ab493ebe4b623fac585776_SeqId_385900210_\":userName:supergroup:-rwxr-xr-x\n at org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.check(FSPermissionChecker.java:234)\n at org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:164)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkPermission(FSNamesystem.java:5202)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkPermission(FSNamesystem.java:5184)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkPathAccess(FSNamesystem.java:5146)\n[... more file permission failures ...]\n{code}","created":"2015-06-08T04:10:37.801+0000"},{"body":"[~clayb] Which version of Hbase are you using? Is it a secure or unsecure cluster? Thanks.","created":"2015-06-09T18:22:59.943+0000"},{"body":"This one was as unsecure cluster running 0.96.","created":"2015-06-14T23:21:16.819+0000"},{"body":"0.96 is EOL. Can you upgrade to a recent 0.98 release and retest? Thanks.\n","created":"2015-06-15T18:10:37.293+0000"},{"body":"I am using Hbase version HBase 0.98.6 and have the same issue.","created":"2015-07-20T13:51:21.251+0000"},{"body":"The issue should have been fixed (see PHOENIX-976). The solution is to always use secure BulkLoad as default and to configure staging directory different from Hbase root temp directory. ","created":"2015-08-27T19:07:40.242+0000"},{"body":"Add \"org.apache.hadoop.hbase.security.access.SecureBulkLoadEndpoint\" into config \"\"hbase.coprocessor.region.classes\", all bulk load will use SecureBulkLoadEndpoint. Check HBASE-12052 for details. There is also a related issue HBASE-12533 on HBase bulk load endpoint.","created":"2015-09-02T18:18:37.807+0000"}],"conversations":[{"body":"- Start of HBase master fails due to the following error found in the log.\n\n2014-11-11 20:25:58,860 WARN org.apache.hadoop.hbase.backup.HFileArchiver: \nFailed to archive class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePa\nth,file:hdfs://YYYY/hbase/.tmp/data/default/tbl/00820520f5cb7839395e83f40c8d97c2/e/52bf9eee7a27460c8d9e2a26fa43c918_SeqId_282271246_ on try #1\norg.apache.hadoop.security.AccessControlException: Permission denied: \nuser=hbase,access=WRITE,inode=\"/hbase/.tmp/data/default/tbl/00820520f5cb7839395e83f40c8d97c2/e/52bf9eee7a27460c8d9e2a26fa43c918_SeqId_282271246_\":devuser:supergroup:-rwxr-xr-x\n- All the files that hbase master was complaining about are created under an users user-id instead on \"hbase\" user resulting in incorrect access permission for the master to act on.\n- Looks like this was due to bulk load done using LoadIncrementalHFiles program.\n- HBASE-12052 is another scenario similar to this one. ","from":"reporter","subject":"HBase master start fails due to incorrect file creations"},{"body":"This issue might be a user uses hbase tmp folder as Import tool temporary output folder while HBase will try to recreate(delete and then create) tmp folder during starts. Therefore it cause HMaster can't start.\n\n[~saint.ack@gmail.com] Do u think any error from checkTempDir inside HMaster#createInitialFileSystemLayout is fatal? If it's fatal then we don't need do any thing for the JIRA otherwise we catch the error/log it and move on.","from":"developer"},{"body":"Ping [~saint.ack@gmail.com] any thoughts on this? Thanks.","from":"developer"},{"body":"I can write up a little test [~jeffreyz]? I'm not sure whether an exception out of createInitialFileSystemLayout is fatal...","from":"developer"},{"body":"[~saint.ack@gmail.com] [~jeffreyz] If an exception out of createInitialFileSystemLayout is not fatal, can we catch the exception and make sure the tmp directory exist and continue?","from":"developer"},{"body":"Jeffrey: We ran into this issue on one of our clusters last week. Looking at your JIRA updates, I couldn't make out if you were able to figure out a fix. Would sharing our logs be of any help here? Thanks!","from":"developer"},{"body":"If of use, Jeffrey, Alicia, et al. the errors we (Biju and Sudarshan) saw were:\n\nFatal to the running HMaster as well as any HMaster starting:\n{code}\n2015-06-06 02:08:52,374 ERROR org.apache.hadoop.hbase.backup.HFileArchiver: Failed to archive class org.apache.hadoop.hbase.backup.HFi\nleArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e/f2806ebb34ab493ebe4b623fac585776_SeqId_385900210_\n2015-06-06 02:08:52,374 WARN org.apache.hadoop.hbase.backup.HFileArchiver: Couldn't archive class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e/f2806ebb34ab493ebe4b623fac585776_SeqId_385900210_ into backup directory: hdfs://cluster1/hbase/archive/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e\n2015-06-06 02:08:52,374 WARN org.apache.hadoop.hbase.backup.HFileArchiver: Failed to complete archive of: [class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e/f2806ebb34ab493ebe4b623fac585776_SeqId_385900210_]. Those files are still in the original location, and they may slow down reads.\n2015-06-06 02:08:52,374 FATAL org.apache.hadoop.hbase.master.HMaster: Unhandled exception. Starting shutdown.\njava.io.IOException: Received error when attempting to archive files ([class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e]), cannot delete region directory.\n at org.apache.hadoop.hbase.backup.HFileArchiver.archiveRegion(HFileArchiver.java:148)\n at org.apache.hadoop.hbase.master.MasterFileSystem.checkTempDir(MasterFileSystem.java:503)\n at org.apache.hadoop.hbase.master.MasterFileSystem.createInitialFileSystemLayout(MasterFileSystem.java:149)\n at org.apache.hadoop.hbase.master.MasterFileSystem.(MasterFileSystem.java:127)\n at org.apache.hadoop.hbase.master.HMaster.finishInitialization(HMaster.java:789) \n at org.apache.hadoop.hbase.master.HMaster.run(HMaster.java:606)\n at java.lang.Thread.run(Thread.java:745) \n2015-06-06 02:08:52,375 INFO org.apache.hadoop.hbase.master.HMaster: Aborting\n2015-06-06 02:08:52,375 DEBUG org.apache.hadoop.hbase.master.HMaster: Stopping service threads\n[... more shutting down here ...]\n2015-06-06 02:28:15,583 DEBUG org.apache.hadoop.hbase.master.ActiveMasterManager: A master is now available\n2015-06-06 02:28:15,583 INFO org.apache.hadoop.hbase.master.ActiveMasterManager: Registered Active Master=cluster1-bcpc-r2n7.example.com,60000,1433572093679\n2015-06-06 02:28:15,588 INFO org.apache.hadoop.conf.Configuration.deprecation: fs.default.name is deprecated. Instead, use fs.defaultFS\n2015-06-06 02:28:15,968 INFO org.apache.hadoop.conf.Configuration.deprecation: hadoop.native.lib is deprecated. Instead, use io.native.lib.available\n2015-06-06 02:28:16,164 DEBUG org.apache.hadoop.hbase.util.FSTableDescriptors: Current tableInfoPath = hdfs://cluster1/hbase/data/hbase/meta/.tabledesc/.tableinfo.0000000001\n2015-06-06 02:28:16,195 DEBUG org.apache.hadoop.hbase.util.FSTableDescriptors: TableInfo already exists.. Skipping creation\n2015-06-06 02:28:17,850 DEBUG org.apache.hadoop.hbase.backup.HFileArchiver: ARCHIVING hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf\n2015-06-06 02:28:17,860 DEBUG org.apache.hadoop.hbase.backup.HFileArchiver: Archiving [class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e]\n2015-06-06 02:28:17,885 WARN org.apache.hadoop.hbase.backup.HFileArchiver: Failed to archive class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePath, file:hdfs://cluster1/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e/f2806ebb34ab493ebe4b623fac585776_SeqId_385900210_ on try #0\norg.apache.hadoop.security.AccessControlException: Permission denied: user=hbase, access=WRITE, inode=\"/hbase/.tmp/data/default/xyxyTablenamE12345/002cd7fbf10def3bb3149ed85707fabf/e/f2806ebb34ab493ebe4b623fac585776_SeqId_385900210_\":userName:supergroup:-rwxr-xr-x\n at org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.check(FSPermissionChecker.java:234)\n at org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:164)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkPermission(FSNamesystem.java:5202)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkPermission(FSNamesystem.java:5184)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkPathAccess(FSNamesystem.java:5146)\n[... more file permission failures ...]\n{code}","from":"developer"},{"body":"[~clayb] Which version of Hbase are you using? Is it a secure or unsecure cluster? Thanks.","from":"developer"},{"body":"This one was as unsecure cluster running 0.96.","from":"developer"},{"body":"0.96 is EOL. Can you upgrade to a recent 0.98 release and retest? Thanks.\n","from":"developer"},{"body":"I am using Hbase version HBase 0.98.6 and have the same issue.","from":"developer"},{"body":"The issue should have been fixed (see PHOENIX-976). The solution is to always use secure BulkLoad as default and to configure staging directory different from Hbase root temp directory. ","from":"developer"},{"body":"Add \"org.apache.hadoop.hbase.security.access.SecureBulkLoadEndpoint\" into config \"\"hbase.coprocessor.region.classes\", all bulk load will use SecureBulkLoadEndpoint. Check HBASE-12052 for details. There is also a related issue HBASE-12533 on HBase bulk load endpoint.","from":"developer"}],"created":"2014-11-12T20:13:28.000+0000","description":"- Start of HBase master fails due to the following error found in the log.\n\n2014-11-11 20:25:58,860 WARN org.apache.hadoop.hbase.backup.HFileArchiver: \nFailed to archive class org.apache.hadoop.hbase.backup.HFileArchiver$FileablePa\nth,file:hdfs://YYYY/hbase/.tmp/data/default/tbl/00820520f5cb7839395e83f40c8d97c2/e/52bf9eee7a27460c8d9e2a26fa43c918_SeqId_282271246_ on try #1\norg.apache.hadoop.security.AccessControlException: Permission denied: \nuser=hbase,access=WRITE,inode=\"/hbase/.tmp/data/default/tbl/00820520f5cb7839395e83f40c8d97c2/e/52bf9eee7a27460c8d9e2a26fa43c918_SeqId_282271246_\":devuser:supergroup:-rwxr-xr-x\n- All the files that hbase master was complaining about are created under an users user-id instead on \"hbase\" user resulting in incorrect access permission for the master to act on.\n- Looks like this was due to bulk load done using LoadIncrementalHFiles program.\n- HBASE-12052 is another scenario similar to this one. ","issue_id":"12754873","key":"HBASE-12465","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2015-09-02T18:18:49.000+0000","role":"fixed_distractor","summary":"HBase master start fails due to incorrect file creations"} {"case_id":"12757523","cluster":"DISTRACTOR-HBASE-12565","comments":[{"body":"Is it really incorrect, though? doMiniBatchMutation makes no guarantees about whether it commits all the data or not. The only part guaranteed is that no partial rows are visible (via MVCC) that guarantee is still in tact.\n\nNot saying this is good, but none of the HBase guarantees are broken.\nMaybe the problem is that a client gets an exception and assumes that no data was written?\n","created":"2014-11-25T04:36:13.294+0000"},{"body":"That is correct, it does not have to commit all the data or not. It should however accurately relay which rows were written and which ones were not with status codes similar to its 0.94 counterpart. It just looks like it needs a bit of attention with regards to error semantics.","created":"2014-11-25T04:56:44.017+0000"},{"body":"I see. Yeah, that's broken.","created":"2014-11-25T05:07:16.955+0000"},{"body":"I am currently working on a patch/fix.","created":"2014-11-25T17:18:17.960+0000"},{"body":"Applies the two part recommended solution from the description.\n\nDivided the existing batchMutate() test into two tests, one that performs\nbatchMutate while no row locks are held, one that performs batch mutate \nwhile multiple row locks are held and concurrently another thread calls \nregion.close().\n","created":"2014-12-04T19:16:51.889+0000"},{"body":"Fixes line length warnings and hanging unit test.","created":"2014-12-04T23:45:20.556+0000"},{"body":"Looks good to me. We might delay a requested close a bit. Flush and compact do that as well, so that seems OK.\n\nCould somebody else have a look too? [~stack], [~apurtell]?\n","created":"2014-12-05T00:06:29.119+0000"},{"body":"Patch looks good to me. Seems right that close, etc., should hold until the batch goes in. The 'internal' row lock refactor is good making it so we only use the guts of it down in mini batch.\n\nI'd like to get another hadoopqa run in to make sure TestAtomicOperation passes. hadoopqa is broke at mo. Hopefully will be back soon.","created":"2014-12-05T01:22:03.974+0000"},{"body":"Retry.Does the failed test pass locally?","created":"2014-12-05T06:40:34.994+0000"},{"body":"Pushed to branch-1+ Thanks for the patch Keith.","created":"2014-12-05T18:23:53.815+0000"},{"body":"Closing this issue after 1.0.0 release.","created":"2015-02-21T23:48:52.294+0000"}],"conversations":[{"body":"The following sequence of events is possible to occur in HRegion's batchMutate() call:\n\n1. caller attempts to call HRegion.batchMutate() with a batch of N>1 records\n2. batchMutate acquires region lock in startRegionOperation, then calls doMiniBatchMutation()\n3. doMiniBatchMutation acquires one row lock\n4. Region closes\n5. doMiniBatchMutation attempts to acquire second row lock.\n\nWhen this happens, the lock acquisition will also attempt to acquire the region lock, which fails (because the region is closing). At this stage, doMiniBatchMutation will stop writing further, BUT it WILL write data for the rows whose locks have already been acquired, and advance the index in MiniBatchOperationInProgress. Then, after it terminates successfully, batchMutate() will loop around a second time, and attempt AGAIN to acquire the region closing lock. When that happens, a NotServingRegionException is thrown back to the caller.\n\nThus, we have a race condition where partial data can be written when a region server is closing.\n\nThe main problem stems from the location of startRegionOperation() calls in batchMutate and doMiniBatchMutation():\n\n1. batchMutate() reacquires the region lock with each iteration of the loop, which can cause some successful writes to occur, but then fail on others\n2. getRowLock() attempts to acquire the region lock once for each row, which allows doMiniBatchMutation to terminate early; this forces batchMutate() to use multiple iterations and results in condition 1 being hit.\n\nThere appears to be two parts to the solution as well:\n\n1. open an internal path so that doMiniBatchMutation() can acquire row locks without checking for region closure. This will have the added benefit of a significant performance improvement during large batch mutations.\n2. move the startRegionOperation() out of the loop in batchMutate() so that multiple iterations of doMiniBatchMutation will not cause the operation to fail.","from":"reporter","subject":"Race condition in HRegion.batchMutate() causes partial data to be written when region closes"},{"body":"Is it really incorrect, though? doMiniBatchMutation makes no guarantees about whether it commits all the data or not. The only part guaranteed is that no partial rows are visible (via MVCC) that guarantee is still in tact.\n\nNot saying this is good, but none of the HBase guarantees are broken.\nMaybe the problem is that a client gets an exception and assumes that no data was written?\n","from":"developer"},{"body":"That is correct, it does not have to commit all the data or not. It should however accurately relay which rows were written and which ones were not with status codes similar to its 0.94 counterpart. It just looks like it needs a bit of attention with regards to error semantics.","from":"developer"},{"body":"I see. Yeah, that's broken.","from":"developer"},{"body":"I am currently working on a patch/fix.","from":"developer"},{"body":"Applies the two part recommended solution from the description.\n\nDivided the existing batchMutate() test into two tests, one that performs\nbatchMutate while no row locks are held, one that performs batch mutate \nwhile multiple row locks are held and concurrently another thread calls \nregion.close().\n","from":"developer"},{"body":"Fixes line length warnings and hanging unit test.","from":"developer"},{"body":"Looks good to me. We might delay a requested close a bit. Flush and compact do that as well, so that seems OK.\n\nCould somebody else have a look too? [~stack], [~apurtell]?\n","from":"developer"},{"body":"Patch looks good to me. Seems right that close, etc., should hold until the batch goes in. The 'internal' row lock refactor is good making it so we only use the guts of it down in mini batch.\n\nI'd like to get another hadoopqa run in to make sure TestAtomicOperation passes. hadoopqa is broke at mo. Hopefully will be back soon.","from":"developer"},{"body":"Retry.Does the failed test pass locally?","from":"developer"},{"body":"Pushed to branch-1+ Thanks for the patch Keith.","from":"developer"},{"body":"Closing this issue after 1.0.0 release.","from":"developer"}],"created":"2014-11-24T21:55:48.000+0000","description":"The following sequence of events is possible to occur in HRegion's batchMutate() call:\n\n1. caller attempts to call HRegion.batchMutate() with a batch of N>1 records\n2. batchMutate acquires region lock in startRegionOperation, then calls doMiniBatchMutation()\n3. doMiniBatchMutation acquires one row lock\n4. Region closes\n5. doMiniBatchMutation attempts to acquire second row lock.\n\nWhen this happens, the lock acquisition will also attempt to acquire the region lock, which fails (because the region is closing). At this stage, doMiniBatchMutation will stop writing further, BUT it WILL write data for the rows whose locks have already been acquired, and advance the index in MiniBatchOperationInProgress. Then, after it terminates successfully, batchMutate() will loop around a second time, and attempt AGAIN to acquire the region closing lock. When that happens, a NotServingRegionException is thrown back to the caller.\n\nThus, we have a race condition where partial data can be written when a region server is closing.\n\nThe main problem stems from the location of startRegionOperation() calls in batchMutate and doMiniBatchMutation():\n\n1. batchMutate() reacquires the region lock with each iteration of the loop, which can cause some successful writes to occur, but then fail on others\n2. getRowLock() attempts to acquire the region lock once for each row, which allows doMiniBatchMutation to terminate early; this forces batchMutate() to use multiple iterations and results in condition 1 being hit.\n\nThere appears to be two parts to the solution as well:\n\n1. open an internal path so that doMiniBatchMutation() can acquire row locks without checking for region closure. This will have the added benefit of a significant performance improvement during large batch mutations.\n2. move the startRegionOperation() out of the loop in batchMutate() so that multiple iterations of doMiniBatchMutation will not cause the operation to fail.","issue_id":"12757523","key":"HBASE-12565","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2014-12-05T18:23:53.000+0000","role":"fixed_distractor","summary":"Race condition in HRegion.batchMutate() causes partial data to be written when region closes"} {"case_id":"12772446","cluster":"DISTRACTOR-HBASE-12971","comments":[{"body":"What about adding a configurable upper bound for sleep time? ","created":"2015-02-05T00:16:22.565+0000"},{"body":"Yeah, this is bad!\nWhat's the rationale for squaring the retry multiplier for socket timeouts? I can see maybe doubling it.\n\nDoes anybody disagree with that? If not, I'll change the logic to double max retries for socket timeouts instead of squaring.\n","created":"2015-02-10T04:34:54.178+0000"},{"body":"I guess it just worked nicely this way with the value being 10.\nI'd like to change this to 2x in all branches.","created":"2015-02-10T04:36:31.637+0000"},{"body":"[~lhofhansl] I wonder why socket timeout sleep time should be calculated by max retry count? ","created":"2015-02-10T05:28:58.470+0000"},{"body":"I do not think we should invent yet another config option.\nWe can already configure replication.source.socketTimeoutMultiplier, it's just about a good default.\n\nIn fact with that in mind maybe the socketTimeoutMultiplier should just be maxRetriesMultiplier (we declared maxRetriesMultiplier to be a good maximum since we configured it that way, on a socket timeout it seems good to wait for that maximum immediately).\n\nEverybody good with that (socketTimeoutMultiplier = maxRetriesMultiplier)?\n","created":"2015-02-10T06:58:32.614+0000"},{"body":"Just explain my thought here.\nI think maxRetriesMultiplier carries two things here, one is max retry count, and the other is max sleep time. I want to retry 300 times does not mean I want to sleep 300s when last retry.\nSo I suggest to have two config here, maxRetryCount and maxRetrySleepTime. If socketTimeout then just use maxRetrySleepTime.\n\nOf course your solution can work for this issue [~lhofhansl], but I really think we should use another model to calculate sleep time(base * retry is too simple :) )","created":"2015-02-10T08:08:32.992+0000"},{"body":"+1 on two configuration parameters: {{maxRetryCount and maxRetrySleepTime}}","created":"2015-02-10T08:43:56.104+0000"},{"body":"We already have a ton of config parameters. \n\nWill this be good enough?:\n{quote}\nWe can already configure replication.source.socketTimeoutMultiplier, it's just about a good default.\nIn fact with that in mind maybe the socketTimeoutMultiplier should just be maxRetriesMultiplier (we declared maxRetriesMultiplier to be a good maximum since we configured it that way, on a socket timeout it seems good to wait for that maximum immediately).\nEverybody good with that (socketTimeoutMultiplier = maxRetriesMultiplier)?\n{quote}\nBecause if so let's make the change, add a release note, and we are done here.","created":"2015-02-10T16:31:44.378+0000"},{"body":"bq. +1 on two configuration parameters: maxRetryCount and maxRetrySleepTime\n\nI disagree. We have retry interval and retry count now.\nLet's not as another thing. If interval+count dues not work let's change the whole thing.","created":"2015-02-11T02:07:27.348+0000"},{"body":"{quote}\nLet's not as another thing. If interval+count dues not work let's change the whole thing.\n{quote}\nFine. Can do it in another issue if current interval+count does not work.","created":"2015-02-11T02:41:27.833+0000"},{"body":"So here's a patch of breathtaking complexity.\n\nI suggest we only apply to the branches where we already increase the maxretries to 300. (1.0.1, 1.1, and 2.0).","created":"2015-02-12T19:13:28.942+0000"},{"body":"bq. So here's a patch of breathtaking complexity.\n\nlol\n\n+1","created":"2015-02-12T19:22:30.287+0000"},{"body":"Separately.\nYou don't want to make this change in 0.98 Lars? \nEarlier you said:\n{quote}\nWe can already configure replication.source.socketTimeoutMultiplier, it's just about a good default.\nIn fact with that in mind maybe the socketTimeoutMultiplier should just be maxRetriesMultiplier (we declared maxRetriesMultiplier to be a good maximum since we configured it that way, on a socket timeout it seems good to wait for that maximum immediately).\n{quote}\nDoes that logic not hold? I realize there will be a behavioral change if we stop squaring the max retries multiplier here, but following the above it borders on a bug.\n\nIf we don't apply this to 0.98 I think we are going to be bit by this in production at some point. ","created":"2015-02-12T19:25:23.067+0000"},{"body":"[~lhofhansl] Makes sense!\nWould it make sense to also sync default value for {{this.maxRetriesMultiplier = this.conf.getInt(\"replication.source.maxretriesmultiplier\", 10);}} ?\nIn ReplicationSource.java#L169 and HBaseInterClusterReplicationEndpoint.java#L79 ?\n{{HBaseInterClusterReplicationEndpoint}} duplicates a lot of code from {{ReplicationSource.java}} so we should also factor it out - but probably in a different patch\n","created":"2015-02-12T19:29:02.887+0000"},{"body":"[~amuraru], agreed.\n\n[~apurtell], the default is max retries is 10 (i.e. 10s with the default sleep interval). For socket timeouts that is too small (I think).\nIn fact I see the following comment as to why socket timeouts are handled differently:\n{code}\n // This exception means we waited for more than 60s and nothing\n // happened, the cluster is alive and calling it right away\n // even for a test just makes things worse.\n{code}\n\nI do not see us setting a socket timeout anywhere, so the 60s much be an assumption/default.\nMaybe the retry after a socket timeout should at least wait for 60s. (i.e. max(60s, sleep interval * max retries).\n","created":"2015-02-12T19:40:21.543+0000"},{"body":"At the same time I do not want to complicate this unduly.","created":"2015-02-12T19:54:04.934+0000"},{"body":"bq. Maybe the retry after a socket timeout should at least wait for 60s. (i.e. max(60s, sleep interval * max retries).\n\nMaybe, ok, we can do that somewhere else to follow up","created":"2015-02-12T19:55:11.142+0000"},{"body":"We can drive this even further. We probably want distinguish between a connect and read timeout.\n\nOr another question is: Why handle a socket timeout exception differently (from other problem) in the first place? We can get a TO exception, why not just increase the sleep timer and sleep?\n\nAnyway, let's just fix the immediate problem(s) that we introduced in HBASE-11367 here.","created":"2015-02-12T20:08:17.873+0000"},{"body":"Fixes the default max retries as well.\nWill commit commit to 1.0.1+ unless I hear objections.\n","created":"2015-02-12T20:10:14.950+0000"},{"body":"I was wrong. The timeouts default to this:\n{code}\n public final static int DEFAULT_SOCKET_TIMEOUT_CONNECT = 10000; // 10 seconds\n public final static int DEFAULT_SOCKET_TIMEOUT_READ = 20000; // 20 seconds\n public final static int DEFAULT_SOCKET_TIMEOUT_WRITE = 60000; // 60 seconds\n{code}\n(from RpcClient)","created":"2015-02-13T06:58:54.058+0000"},{"body":"Pushed to 1.0.1, 1.1, and 2.0","created":"2015-02-14T06:19:11.841+0000"},{"body":"Closing this issue after 1.0.0 release.","created":"2015-02-21T23:49:25.893+0000"}],"conversations":[{"body":"We are setting in hbase-site the default value of 300 for {{replication.source.maxretriesmultiplier}} introduced in HBASE-11964.\n\nWhile this value works fine to recover for transient errors with remote ZK quorum from the peer Hbase cluster - it proved to have side effects in the code introduced in HBASE-11367 Pluggable replication endpoint, where the default is much lower (10).\nSee:\n1. https://github.com/apache/hbase/blob/master/hbase-server/src/main/java/org/apache/hadoop/hbase/replication/regionserver/ReplicationSource.java#L169\n2. https://github.com/apache/hbase/blob/master/hbase-server/src/main/java/org/apache/hadoop/hbase/replication/regionserver/HBaseInterClusterReplicationEndpoint.java#L79\n\nThe the two default values are definitely conflicting - when {{replication.source.maxretriesmultiplier}} is set in the hbase-site to 300 this will lead to a sleep time of 300*300 (25h!) when a sockettimeout exception is thrown.\n\n","from":"reporter","subject":"Replication stuck due to large default value for replication.source.maxretriesmultiplier"},{"body":"What about adding a configurable upper bound for sleep time? ","from":"developer"},{"body":"Yeah, this is bad!\nWhat's the rationale for squaring the retry multiplier for socket timeouts? I can see maybe doubling it.\n\nDoes anybody disagree with that? If not, I'll change the logic to double max retries for socket timeouts instead of squaring.\n","from":"developer"},{"body":"I guess it just worked nicely this way with the value being 10.\nI'd like to change this to 2x in all branches.","from":"developer"},{"body":"[~lhofhansl] I wonder why socket timeout sleep time should be calculated by max retry count? ","from":"developer"},{"body":"I do not think we should invent yet another config option.\nWe can already configure replication.source.socketTimeoutMultiplier, it's just about a good default.\n\nIn fact with that in mind maybe the socketTimeoutMultiplier should just be maxRetriesMultiplier (we declared maxRetriesMultiplier to be a good maximum since we configured it that way, on a socket timeout it seems good to wait for that maximum immediately).\n\nEverybody good with that (socketTimeoutMultiplier = maxRetriesMultiplier)?\n","from":"developer"},{"body":"Just explain my thought here.\nI think maxRetriesMultiplier carries two things here, one is max retry count, and the other is max sleep time. I want to retry 300 times does not mean I want to sleep 300s when last retry.\nSo I suggest to have two config here, maxRetryCount and maxRetrySleepTime. If socketTimeout then just use maxRetrySleepTime.\n\nOf course your solution can work for this issue [~lhofhansl], but I really think we should use another model to calculate sleep time(base * retry is too simple :) )","from":"developer"},{"body":"+1 on two configuration parameters: {{maxRetryCount and maxRetrySleepTime}}","from":"developer"},{"body":"We already have a ton of config parameters. \n\nWill this be good enough?:\n{quote}\nWe can already configure replication.source.socketTimeoutMultiplier, it's just about a good default.\nIn fact with that in mind maybe the socketTimeoutMultiplier should just be maxRetriesMultiplier (we declared maxRetriesMultiplier to be a good maximum since we configured it that way, on a socket timeout it seems good to wait for that maximum immediately).\nEverybody good with that (socketTimeoutMultiplier = maxRetriesMultiplier)?\n{quote}\nBecause if so let's make the change, add a release note, and we are done here.","from":"developer"},{"body":"bq. +1 on two configuration parameters: maxRetryCount and maxRetrySleepTime\n\nI disagree. We have retry interval and retry count now.\nLet's not as another thing. If interval+count dues not work let's change the whole thing.","from":"developer"},{"body":"{quote}\nLet's not as another thing. If interval+count dues not work let's change the whole thing.\n{quote}\nFine. Can do it in another issue if current interval+count does not work.","from":"developer"},{"body":"So here's a patch of breathtaking complexity.\n\nI suggest we only apply to the branches where we already increase the maxretries to 300. (1.0.1, 1.1, and 2.0).","from":"developer"},{"body":"bq. So here's a patch of breathtaking complexity.\n\nlol\n\n+1","from":"developer"},{"body":"Separately.\nYou don't want to make this change in 0.98 Lars? \nEarlier you said:\n{quote}\nWe can already configure replication.source.socketTimeoutMultiplier, it's just about a good default.\nIn fact with that in mind maybe the socketTimeoutMultiplier should just be maxRetriesMultiplier (we declared maxRetriesMultiplier to be a good maximum since we configured it that way, on a socket timeout it seems good to wait for that maximum immediately).\n{quote}\nDoes that logic not hold? I realize there will be a behavioral change if we stop squaring the max retries multiplier here, but following the above it borders on a bug.\n\nIf we don't apply this to 0.98 I think we are going to be bit by this in production at some point. ","from":"developer"},{"body":"[~lhofhansl] Makes sense!\nWould it make sense to also sync default value for {{this.maxRetriesMultiplier = this.conf.getInt(\"replication.source.maxretriesmultiplier\", 10);}} ?\nIn ReplicationSource.java#L169 and HBaseInterClusterReplicationEndpoint.java#L79 ?\n{{HBaseInterClusterReplicationEndpoint}} duplicates a lot of code from {{ReplicationSource.java}} so we should also factor it out - but probably in a different patch\n","from":"developer"},{"body":"[~amuraru], agreed.\n\n[~apurtell], the default is max retries is 10 (i.e. 10s with the default sleep interval). For socket timeouts that is too small (I think).\nIn fact I see the following comment as to why socket timeouts are handled differently:\n{code}\n // This exception means we waited for more than 60s and nothing\n // happened, the cluster is alive and calling it right away\n // even for a test just makes things worse.\n{code}\n\nI do not see us setting a socket timeout anywhere, so the 60s much be an assumption/default.\nMaybe the retry after a socket timeout should at least wait for 60s. (i.e. max(60s, sleep interval * max retries).\n","from":"developer"},{"body":"At the same time I do not want to complicate this unduly.","from":"developer"},{"body":"bq. Maybe the retry after a socket timeout should at least wait for 60s. (i.e. max(60s, sleep interval * max retries).\n\nMaybe, ok, we can do that somewhere else to follow up","from":"developer"},{"body":"We can drive this even further. We probably want distinguish between a connect and read timeout.\n\nOr another question is: Why handle a socket timeout exception differently (from other problem) in the first place? We can get a TO exception, why not just increase the sleep timer and sleep?\n\nAnyway, let's just fix the immediate problem(s) that we introduced in HBASE-11367 here.","from":"developer"},{"body":"Fixes the default max retries as well.\nWill commit commit to 1.0.1+ unless I hear objections.\n","from":"developer"},{"body":"I was wrong. The timeouts default to this:\n{code}\n public final static int DEFAULT_SOCKET_TIMEOUT_CONNECT = 10000; // 10 seconds\n public final static int DEFAULT_SOCKET_TIMEOUT_READ = 20000; // 20 seconds\n public final static int DEFAULT_SOCKET_TIMEOUT_WRITE = 60000; // 60 seconds\n{code}\n(from RpcClient)","from":"developer"},{"body":"Pushed to 1.0.1, 1.1, and 2.0","from":"developer"},{"body":"Closing this issue after 1.0.0 release.","from":"developer"}],"created":"2015-02-04T18:21:10.000+0000","description":"We are setting in hbase-site the default value of 300 for {{replication.source.maxretriesmultiplier}} introduced in HBASE-11964.\n\nWhile this value works fine to recover for transient errors with remote ZK quorum from the peer Hbase cluster - it proved to have side effects in the code introduced in HBASE-11367 Pluggable replication endpoint, where the default is much lower (10).\nSee:\n1. https://github.com/apache/hbase/blob/master/hbase-server/src/main/java/org/apache/hadoop/hbase/replication/regionserver/ReplicationSource.java#L169\n2. https://github.com/apache/hbase/blob/master/hbase-server/src/main/java/org/apache/hadoop/hbase/replication/regionserver/HBaseInterClusterReplicationEndpoint.java#L79\n\nThe the two default values are definitely conflicting - when {{replication.source.maxretriesmultiplier}} is set in the hbase-site to 300 this will lead to a sleep time of 300*300 (25h!) when a sockettimeout exception is thrown.\n\n","issue_id":"12772446","key":"HBASE-12971","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2015-02-14T06:19:11.000+0000","role":"fixed_distractor","summary":"Replication stuck due to large default value for replication.source.maxretriesmultiplier"} {"case_id":"12828919","cluster":"DISTRACTOR-HBASE-13664","comments":[{"body":"This patch works for both branch-1 and branch-1.1","created":"2015-05-11T21:24:01.924+0000"},{"body":"Please name your patch with 'branch-1' in its name so that QA bot applies it on correct branch.","created":"2015-05-11T21:53:20.523+0000"},{"body":"Thanks. Added a copy of the original file with branch-1 in it, and removed the old copy.","created":"2015-05-11T22:22:40.185+0000"},{"body":"where's the \"-1 core tests. The patch failed these unit tests:\" coming from? I also couldn't see any javadoc warnings related to this change.\n\nWhat's the next step from here?","created":"2015-05-12T17:30:29.780+0000"},{"body":"dev-support/findHangingTests.py reported:\n{code}\nFetching the console output from the URL\nPrinting hanging tests\nHanging test : org.apache.hadoop.hbase.mapreduce.TestTableInputFormatScan1\nHanging test : org.apache.hadoop.hbase.mapreduce.TestCellCounter\nHanging test : org.apache.hadoop.hbase.mapreduce.TestHFileOutputFormat2\nHanging test : org.apache.hadoop.hbase.mapreduce.TestMultiTableInputFormat\nPrinting Failing tests\nFailing test : org.apache.hadoop.hbase.client.TestSnapshotCloneIndependence\nFailing test : org.apache.hadoop.hbase.master.handler.TestEnableTableHandler\n{code}\nMaybe attach patch one more time ?","created":"2015-05-12T17:33:49.201+0000"},{"body":"re-attaching the same file to see if the test clear up.","created":"2015-05-12T17:37:51.615+0000"},{"body":"Thanks for the patch, Solomon.","created":"2015-05-12T20:32:13.896+0000"},{"body":"Thanks for the push. branch-1 and branch-1.0 were updated. branch-1.1 and branch-1.1.0 should probably be updated as well. How can I help make that happen?","created":"2015-05-12T21:22:18.138+0000"},{"body":"Pushed to branch-1.1 as well.\n\nbranch-1.1.0 is for 1.1.0 release.","created":"2015-05-12T21:29:52.578+0000"},{"body":"gracias","created":"2015-05-12T21:33:54.213+0000"},{"body":"Compilation failed in hbase-thrift/src/main/java/org/apache/hadoop/hbase/thrift2/ThriftHBaseServiceHandler.java in branch-1.0","created":"2015-05-12T22:13:16.308+0000"},{"body":"Testing an addendum based on https://issues.apache.org/jira/secure/attachment/12706955/HBASE-13201-2-branch-1.patch which didn't get patched into branch-1.0","created":"2015-05-12T22:52:26.737+0000"},{"body":"This is a patch to fix branch-1.0's problems with the patch applied to branch-1 and branch-1.1. branch-1.0 needed an additional change that was previously submitted with HBASE-13201 and applied to the -1 and -1.1 branches, but not -1.0.","created":"2015-05-12T23:15:16.823+0000"},{"body":"[~ndimiduk] was kind enough to get the right changes from HBASE-13201 into branch-1.0.\n\n[~tedyu] Do I need to attach the same HBASE-13664 patch again to set off the testing cycle?","created":"2015-05-13T18:55:10.437+0000"},{"body":"Re-integrated into branch-1.0","created":"2015-05-13T18:59:46.608+0000"},{"body":"The reported OOMing and other errors in TestDistributedLogSplitting shouldn't have anything to do with ConnectionCache or the thrift server changes. ConnectionCache is only used in the REST and thrift server.","created":"2015-05-13T20:25:24.592+0000"},{"body":"Thanks for double-checking Solomon.","created":"2015-05-13T20:30:57.778+0000"},{"body":"Closing this issue after 1.0.2 release.","created":"2015-08-31T22:39:29.667+0000"}],"conversations":[{"body":"ConnectionCache uses the new HBase interfaces in master. Update branch-1* to use them as well.","from":"reporter","subject":"Use HBase 1.0 interfaces in ConnectionCache"},{"body":"This patch works for both branch-1 and branch-1.1","from":"developer"},{"body":"Please name your patch with 'branch-1' in its name so that QA bot applies it on correct branch.","from":"developer"},{"body":"Thanks. Added a copy of the original file with branch-1 in it, and removed the old copy.","from":"developer"},{"body":"where's the \"-1 core tests. The patch failed these unit tests:\" coming from? I also couldn't see any javadoc warnings related to this change.\n\nWhat's the next step from here?","from":"developer"},{"body":"dev-support/findHangingTests.py reported:\n{code}\nFetching the console output from the URL\nPrinting hanging tests\nHanging test : org.apache.hadoop.hbase.mapreduce.TestTableInputFormatScan1\nHanging test : org.apache.hadoop.hbase.mapreduce.TestCellCounter\nHanging test : org.apache.hadoop.hbase.mapreduce.TestHFileOutputFormat2\nHanging test : org.apache.hadoop.hbase.mapreduce.TestMultiTableInputFormat\nPrinting Failing tests\nFailing test : org.apache.hadoop.hbase.client.TestSnapshotCloneIndependence\nFailing test : org.apache.hadoop.hbase.master.handler.TestEnableTableHandler\n{code}\nMaybe attach patch one more time ?","from":"developer"},{"body":"re-attaching the same file to see if the test clear up.","from":"developer"},{"body":"Thanks for the patch, Solomon.","from":"developer"},{"body":"Thanks for the push. branch-1 and branch-1.0 were updated. branch-1.1 and branch-1.1.0 should probably be updated as well. How can I help make that happen?","from":"developer"},{"body":"Pushed to branch-1.1 as well.\n\nbranch-1.1.0 is for 1.1.0 release.","from":"developer"},{"body":"gracias","from":"developer"},{"body":"Compilation failed in hbase-thrift/src/main/java/org/apache/hadoop/hbase/thrift2/ThriftHBaseServiceHandler.java in branch-1.0","from":"developer"},{"body":"Testing an addendum based on https://issues.apache.org/jira/secure/attachment/12706955/HBASE-13201-2-branch-1.patch which didn't get patched into branch-1.0","from":"developer"},{"body":"This is a patch to fix branch-1.0's problems with the patch applied to branch-1 and branch-1.1. branch-1.0 needed an additional change that was previously submitted with HBASE-13201 and applied to the -1 and -1.1 branches, but not -1.0.","from":"developer"},{"body":"[~ndimiduk] was kind enough to get the right changes from HBASE-13201 into branch-1.0.\n\n[~tedyu] Do I need to attach the same HBASE-13664 patch again to set off the testing cycle?","from":"developer"},{"body":"Re-integrated into branch-1.0","from":"developer"},{"body":"The reported OOMing and other errors in TestDistributedLogSplitting shouldn't have anything to do with ConnectionCache or the thrift server changes. ConnectionCache is only used in the REST and thrift server.","from":"developer"},{"body":"Thanks for double-checking Solomon.","from":"developer"},{"body":"Closing this issue after 1.0.2 release.","from":"developer"}],"created":"2015-05-11T19:09:04.000+0000","description":"ConnectionCache uses the new HBase interfaces in master. Update branch-1* to use them as well.","issue_id":"12828919","key":"HBASE-13664","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2015-05-13T19:00:00.000+0000","role":"fixed_distractor","summary":"Use HBase 1.0 interfaces in ConnectionCache"} {"case_id":"12834590","cluster":"DISTRACTOR-HBASE-13825","comments":[{"body":"One option is to use CodedInputStream#setSizeLimit in the client to effectively disable this check by setting it to Integer.MAX.","created":"2015-06-03T05:07:45.695+0000"},{"body":"Thanks for the suggestion [~apurtell] , this is what the stack trace suggests but please can you help with a code snippet? When you say \"change it in the client\" do you mean the Hbase client or the application client calling the \"get\". I am only able/permitted to use pre-built Hbase jars from maven so cannot change Hbase code in any way.\n\nCodedInputStream.setSizeLimit() suggests using a static method which does not exist. Furthermore I have no instances of CodedInputStream in by application client so where should I set this size limit?\n\nIs it work adding a HBase parameter for this? ","created":"2015-06-03T08:27:54.575+0000"},{"body":"bq. When you say \"change it in the client\" do you mean the Hbase client or...\n\nThe HBase client","created":"2015-06-04T03:04:12.995+0000"},{"body":"This probably needs to be configurable as a Hbase option as we cannot change client code.","created":"2015-06-06T21:35:47.081+0000"},{"body":"Wondering if you have regionserver logs for that event by chance? Also curious if you have hbase.table.max.rowsize property set in the config.","created":"2015-06-17T10:39:26.503+0000"},{"body":"Hi [~mantonov], we don't have hbase.table.max.rowsize set but we also don't see any RowTooBigExceptions being thrown either in the region server logs (which I cannot send out unfortunately)","created":"2015-06-17T12:34:18.495+0000"},{"body":"Last week on the mailing list someone else wrote in with a problem similar to this, where the common issue is hitting the static CodedInputStream limit in the client. The new case was deserializing HBase PB types in a MapReduce worker, something that can't be helped with server side response limit options. I am going to pick up this issue next week and plan to address it with a site configuration option for adjusting the static CodedInputStream limit. ","created":"2015-07-27T00:03:30.927+0000"},{"body":"HBASE-14076 related.","created":"2015-07-27T13:54:20.969+0000"},{"body":"I followed the reference to HBASE-14076 over to HBASE-13230. The solution there is to use the static helper ProtobufUtil#mergeDelimitedFrom wherever we've written a delimited message and would use mergeDelmitedFrom to read it back in, since the delimited message format begins with the total message size encoded in vint32. We use the encoded size to adjust the CodedInputStream limit as needed. \n\nPatches here also address relevant uses of #mergeFrom. We use Integer.MAX_VALUE as the size limit for CodedInputStream where it is not known. In some places it's unlikely a message processed there will exceed 64 MB, but I made a change anyway. It is harmless and consistent to use ProtobufUtil#mergeFrom.\n\nbranch-1 and 0.98 patches also incorporate HBASE-14076.\n\nReviewboard: https://reviews.apache.org/r/37062/\n\n/cc [~stack] Touched a lot of your code here.","created":"2015-08-04T03:02:04.683+0000"},{"body":"Dang forgot to set patch available, let's do that now...","created":"2015-08-04T16:15:16.925+0000"},{"body":"If you have a sec [~esteban], the branch-1 and 0.98 patches here incorporate your work on HBASE-14076, what do you think?","created":"2015-08-04T16:16:38.008+0000"},{"body":"The precommit test was bad because someone killed our test JVM externally:\n{noformat}\nxecutionException: java.lang.RuntimeException: The forked VM terminated without properly saying goodbye. VM crash or System.exit called?\n[ERROR] Command was /bin/sh -c cd /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/hbase-server && /home/jenkins/jenkins-slave/tools/hudson.model.JDK/jdk-1.7u51/jre/bin/java -enableassertions -XX:MaxDirectMemorySize=1G -Xmx2800m -XX:MaxPermSize=256m -Djava.security.egd=file:/dev/./urandom -Djava.net.preferIPv4Stack=true -Djava.awt.headless=true -jar /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/hbase-server/target/surefire/surefirebooter8543005017696418773.jar /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/hbase-server/target/surefire/surefire2508603723119457542tmp /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/hbase-server/target/surefire/surefire_9171447714369068258025tmp\n{noformat}\n\nThe \"zombie\" was AmbariManagementControllerTest, that's not us. \n\nTests pass for me locally. \n\nLet me check on that checkstyle thing","created":"2015-08-04T19:12:25.821+0000"},{"body":"Valid checkstyle issues:\n- ClusterID: Unused import - com.google.protobuf.InvalidProtocolBufferException . Missed that one.\n- HColumnDescriptor: Unused import - com.google.protobuf.InvalidProtocolBufferException . Also missed this one.\n\nNothing else jumps out as relevant or related. New patches coming up.\n","created":"2015-08-04T19:31:15.792+0000"},{"body":"New patches for branch-1 and 0.98 fix unused imports.","created":"2015-08-04T19:35:03.873+0000"},{"body":"+1 [~apurtell] also I think you addressed some of the comments from [~anoopsamjohn] from HBASE-14076. I'm going to open a JIRA to port the changes to master as well. Thanks!","created":"2015-08-04T21:27:17.164+0000"},{"body":"Thanks [~esteban]. Or, if you like, I can add what you've identified as missing to the patch for master here. If so, what would that be","created":"2015-08-04T22:03:40.923+0000"},{"body":"Ok, I have a +1, going to commit shortly","created":"2015-08-07T00:43:25.808+0000"},{"body":"Committed to 0.98 and up","created":"2015-08-07T16:14:02.157+0000"},{"body":"Closing this issue after 1.0.2 release.","created":"2015-08-31T22:39:50.663+0000"}],"conversations":[{"body":"When performing a get operation on a column family with more than 64MB of data, the operation fails with:\n\nCaused by: Portable(java.io.IOException): Call to host:port failed on local exception: com.google.protobuf.InvalidProtocolBufferException: Protocol message was too large. May be malicious. Use CodedInputStream.setSizeLimit() to increase the size limit.\n at org.apache.hadoop.hbase.ipc.RpcClient.wrapException(RpcClient.java:1481)\n at org.apache.hadoop.hbase.ipc.RpcClient.call(RpcClient.java:1453)\n at org.apache.hadoop.hbase.ipc.RpcClient.callBlockingMethod(RpcClient.java:1653)\n at org.apache.hadoop.hbase.ipc.RpcClient$BlockingRpcChannelImplementation.callBlockingMethod(RpcClient.java:1711)\n at org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$BlockingStub.get(ClientProtos.java:27308)\n at org.apache.hadoop.hbase.protobuf.ProtobufUtil.get(ProtobufUtil.java:1381)\n at org.apache.hadoop.hbase.client.HTable$3.call(HTable.java:753)\n at org.apache.hadoop.hbase.client.HTable$3.call(HTable.java:751)\n at org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:120)\n at org.apache.hadoop.hbase.client.HTable.get(HTable.java:756)\n at org.apache.hadoop.hbase.client.HTable.get(HTable.java:765)\n at org.apache.hadoop.hbase.client.HTablePool$PooledHTable.get(HTablePool.java:395)\n\nThis may be related to https://issues.apache.org/jira/browse/HBASE-11747 but that issue is related to cluster status. \n\nScan and put operations on the same data work fine\n\nTested on a 1.0.0 cluster with both 1.0.1 and 1.0.0 clients.\n\n\n","from":"reporter","subject":"Use ProtobufUtil#mergeFrom and ProtobufUtil#mergeDelimitedFrom in place of builder methods of same name"},{"body":"One option is to use CodedInputStream#setSizeLimit in the client to effectively disable this check by setting it to Integer.MAX.","from":"developer"},{"body":"Thanks for the suggestion [~apurtell] , this is what the stack trace suggests but please can you help with a code snippet? When you say \"change it in the client\" do you mean the Hbase client or the application client calling the \"get\". I am only able/permitted to use pre-built Hbase jars from maven so cannot change Hbase code in any way.\n\nCodedInputStream.setSizeLimit() suggests using a static method which does not exist. Furthermore I have no instances of CodedInputStream in by application client so where should I set this size limit?\n\nIs it work adding a HBase parameter for this? ","from":"developer"},{"body":"bq. When you say \"change it in the client\" do you mean the Hbase client or...\n\nThe HBase client","from":"developer"},{"body":"This probably needs to be configurable as a Hbase option as we cannot change client code.","from":"developer"},{"body":"Wondering if you have regionserver logs for that event by chance? Also curious if you have hbase.table.max.rowsize property set in the config.","from":"developer"},{"body":"Hi [~mantonov], we don't have hbase.table.max.rowsize set but we also don't see any RowTooBigExceptions being thrown either in the region server logs (which I cannot send out unfortunately)","from":"developer"},{"body":"Last week on the mailing list someone else wrote in with a problem similar to this, where the common issue is hitting the static CodedInputStream limit in the client. The new case was deserializing HBase PB types in a MapReduce worker, something that can't be helped with server side response limit options. I am going to pick up this issue next week and plan to address it with a site configuration option for adjusting the static CodedInputStream limit. ","from":"developer"},{"body":"HBASE-14076 related.","from":"developer"},{"body":"I followed the reference to HBASE-14076 over to HBASE-13230. The solution there is to use the static helper ProtobufUtil#mergeDelimitedFrom wherever we've written a delimited message and would use mergeDelmitedFrom to read it back in, since the delimited message format begins with the total message size encoded in vint32. We use the encoded size to adjust the CodedInputStream limit as needed. \n\nPatches here also address relevant uses of #mergeFrom. We use Integer.MAX_VALUE as the size limit for CodedInputStream where it is not known. In some places it's unlikely a message processed there will exceed 64 MB, but I made a change anyway. It is harmless and consistent to use ProtobufUtil#mergeFrom.\n\nbranch-1 and 0.98 patches also incorporate HBASE-14076.\n\nReviewboard: https://reviews.apache.org/r/37062/\n\n/cc [~stack] Touched a lot of your code here.","from":"developer"},{"body":"Dang forgot to set patch available, let's do that now...","from":"developer"},{"body":"If you have a sec [~esteban], the branch-1 and 0.98 patches here incorporate your work on HBASE-14076, what do you think?","from":"developer"},{"body":"The precommit test was bad because someone killed our test JVM externally:\n{noformat}\nxecutionException: java.lang.RuntimeException: The forked VM terminated without properly saying goodbye. VM crash or System.exit called?\n[ERROR] Command was /bin/sh -c cd /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/hbase-server && /home/jenkins/jenkins-slave/tools/hudson.model.JDK/jdk-1.7u51/jre/bin/java -enableassertions -XX:MaxDirectMemorySize=1G -Xmx2800m -XX:MaxPermSize=256m -Djava.security.egd=file:/dev/./urandom -Djava.net.preferIPv4Stack=true -Djava.awt.headless=true -jar /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/hbase-server/target/surefire/surefirebooter8543005017696418773.jar /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/hbase-server/target/surefire/surefire2508603723119457542tmp /home/jenkins/jenkins-slave/workspace/PreCommit-HBASE-Build/hbase-server/target/surefire/surefire_9171447714369068258025tmp\n{noformat}\n\nThe \"zombie\" was AmbariManagementControllerTest, that's not us. \n\nTests pass for me locally. \n\nLet me check on that checkstyle thing","from":"developer"},{"body":"Valid checkstyle issues:\n- ClusterID: Unused import - com.google.protobuf.InvalidProtocolBufferException . Missed that one.\n- HColumnDescriptor: Unused import - com.google.protobuf.InvalidProtocolBufferException . Also missed this one.\n\nNothing else jumps out as relevant or related. New patches coming up.\n","from":"developer"},{"body":"New patches for branch-1 and 0.98 fix unused imports.","from":"developer"},{"body":"+1 [~apurtell] also I think you addressed some of the comments from [~anoopsamjohn] from HBASE-14076. I'm going to open a JIRA to port the changes to master as well. Thanks!","from":"developer"},{"body":"Thanks [~esteban]. Or, if you like, I can add what you've identified as missing to the patch for master here. If so, what would that be","from":"developer"},{"body":"Ok, I have a +1, going to commit shortly","from":"developer"},{"body":"Committed to 0.98 and up","from":"developer"},{"body":"Closing this issue after 1.0.2 release.","from":"developer"}],"created":"2015-06-02T13:48:26.000+0000","description":"When performing a get operation on a column family with more than 64MB of data, the operation fails with:\n\nCaused by: Portable(java.io.IOException): Call to host:port failed on local exception: com.google.protobuf.InvalidProtocolBufferException: Protocol message was too large. May be malicious. Use CodedInputStream.setSizeLimit() to increase the size limit.\n at org.apache.hadoop.hbase.ipc.RpcClient.wrapException(RpcClient.java:1481)\n at org.apache.hadoop.hbase.ipc.RpcClient.call(RpcClient.java:1453)\n at org.apache.hadoop.hbase.ipc.RpcClient.callBlockingMethod(RpcClient.java:1653)\n at org.apache.hadoop.hbase.ipc.RpcClient$BlockingRpcChannelImplementation.callBlockingMethod(RpcClient.java:1711)\n at org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$BlockingStub.get(ClientProtos.java:27308)\n at org.apache.hadoop.hbase.protobuf.ProtobufUtil.get(ProtobufUtil.java:1381)\n at org.apache.hadoop.hbase.client.HTable$3.call(HTable.java:753)\n at org.apache.hadoop.hbase.client.HTable$3.call(HTable.java:751)\n at org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithRetries(RpcRetryingCaller.java:120)\n at org.apache.hadoop.hbase.client.HTable.get(HTable.java:756)\n at org.apache.hadoop.hbase.client.HTable.get(HTable.java:765)\n at org.apache.hadoop.hbase.client.HTablePool$PooledHTable.get(HTablePool.java:395)\n\nThis may be related to https://issues.apache.org/jira/browse/HBASE-11747 but that issue is related to cluster status. \n\nScan and put operations on the same data work fine\n\nTested on a 1.0.0 cluster with both 1.0.1 and 1.0.0 clients.\n\n\n","issue_id":"12834590","key":"HBASE-13825","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2015-08-07T16:14:02.000+0000","role":"fixed_distractor","summary":"Use ProtobufUtil#mergeFrom and ProtobufUtil#mergeDelimitedFrom in place of builder methods of same name"} {"case_id":"12845520","cluster":"DISTRACTOR-HBASE-14098","comments":[{"body":"Related to HBASE-10052?","created":"2015-07-16T08:58:52.719+0000"},{"body":"Pretty simple patch to try and see how this works.","created":"2015-07-16T09:12:34.723+0000"},{"body":"Yeah related but on the other side.","created":"2015-07-16T09:22:10.646+0000"},{"body":"Oh... This is for the output files of a compaction... Aren't we likely going to read the compacted data again (especially since we just dropped it all from the block cache)?\n","created":"2015-07-16T09:47:05.352+0000"},{"body":"Yes it's possible that we will read the file again. It's also likely that on large compactions we will just be blowing the fs cache out of the water. Compacting 32gb on machine with 32gb memory free means that nothing else can be in the fs cache. No /bin/bash, no inodes, nothing.\n\nMy plan is likely to set this only for large compactions; the thought being that large compactions are much more likely to have stale data in them. I'm going to test the current patch out on a cluster that's doing really large compactions right now. If I see any positive changes then we can make this smarter.","created":"2015-07-16T16:16:25.730+0000"},{"body":"That makes sense. Thanks [~eclark]. A cheap approximation might be to set it for major compactions. Agree that ideally we'd decided based on the size of the output file(s) compared to the available RAM for the OS buffer cache.\n","created":"2015-07-17T09:29:55.544+0000"},{"body":"I tested this out yesterday on a cluster with 1.1.\nWhen compacting 100+ GB storfiles I can reliably make the linux kernel pause trying to thrash and find pages to clear out.\nWith just dropping the write's out of cache things are better and there are fewer stalls.\nWith dropping the caches behind writes and reads from a compaction there are no noticeable stalls ( ones that atop sees that take > 1 second ).\n\nSo I think that we should provide the ability to drop caches behind reads and writes on large compactions ( those that would run on the large compaction threads ). However because this could adversely affect people with less data per machine I think that we should have this turned off by default.\n\nThoughts?","created":"2015-07-17T17:43:45.015+0000"},{"body":"Awesome! \n\nCache flush stalls can also be avoided by configuring the Linux disk cache. Could you do a {{sysctl -a | grep dirty}} on one of the boxes and paste vm.dirty_background_ratio, vm.dirty_background_bytes, and vm.dirty_ratio here?\n\nJust curious :) Your previous comment made sense anyway. Caching anything that is usually read sequentially and larger than the available RAM never makes sense.\n","created":"2015-07-17T18:04:25.479+0000"},{"body":"{code}\nvm.dirty_background_ratio = 3\nvm.dirty_ratio = 20\n{code}\n\nThe interesting part is that these stalls aren't really caused by dirty pages. There's only 400mb of dirty pages. These stalls seem to be caused by thrashing the cache. Too much reading/writing actually causes reads to happen on the root disk to re-page in executables.","created":"2015-07-17T18:30:13.578+0000"},{"body":"Patch with test fixes for stripe compaction.\n\nI put it up for review here: https://reviews.facebook.net/D42681","created":"2015-07-20T20:52:42.704+0000"},{"body":"Style fixes","created":"2015-07-20T22:20:47.377+0000"},{"body":"TestIOFencing works on my box 3 times in a row. Checking master.","created":"2015-07-21T18:47:09.525+0000"},{"body":"Yeah that test looks good. I think it's just a little flakey.","created":"2015-07-22T04:54:02.880+0000"},{"body":"Ping? This one solved kernel stalls for us.","created":"2015-07-29T16:58:35.376+0000"},{"body":"bq.if (this.conf.getBoolean(\"hbase.regionserver.compaction.private.readers\", true)) \nSo with this change compactions will always use new readers so that the OS pages of those files are not cached based on the new setting that says dropBehind on compaction?","created":"2015-07-30T06:39:22.298+0000"},{"body":"Was just going to +1... Missed the private readers bit.\nI added the private.readers features in order to decouple compactions readers from \"operational\" readers (user scans, gets, etc).\nDefensively I defaulted this to false, but I think we can safely enable that always... But if we wanted to stay defensive we could default to the value of \"drop caches behind compactions\".\n\nActually, in either case +1 from me.","created":"2015-07-30T17:27:02.051+0000"},{"body":"The patch went stale while I was on vacation. Here's a rebase.\n\nYes this does turn on private readers by default. We've been running it for a while and I haven't seen any down sides, so I feel pretty sure that it's not too big a risk.","created":"2015-08-11T14:50:50.941+0000"},{"body":"Whoops missed the mob file stuff being there.","created":"2015-08-11T15:50:40.644+0000"},{"body":"Committing soon unless there are any other comments.","created":"2015-08-12T19:04:03.668+0000"},{"body":"The patch didn't apply cleanly to branch-1. Here's the backport.","created":"2015-08-12T21:34:00.079+0000"},{"body":"Pushed to master, branch-1, and branch-1.2","created":"2015-08-13T02:48:41.670+0000"},{"body":"bq. Yes this does turn on private readers by default. We've been running it for a while and I haven't seen any down sides, so I feel pretty sure that it's not too big a risk.\n\nCool. I put it in a while back, but haven't seen anybody but us using it. Did you notice improvements in scan performance when you enabled it? (sorry off topic)\n\nCool to have this committed.","created":"2015-08-13T04:11:44.086+0000"},{"body":"This change drops ugly thread dumps in our logs:\n\n{code}\n9034 2015-08-26 11:18:38,175 DEBUG [Time-limited test] hfile.HFile$WriterFactory(308): Unable to set drop behind on /Users/stack/checkouts/hbase.git.commit/hbase-server/target/test-data/ed67f436-1d46-43f#\n9035 java.lang.UnsupportedOperationException: the wrapped stream does not support setting the drop-behind caching setting.\n9036 › at org.apache.hadoop.fs.FSDataOutputStream.setDropBehind(FSDataOutputStream.java:150)\n9037 › at org.apache.hadoop.hbase.io.hfile.HFile$WriterFactory.create(HFile.java:306)\n9038 › at org.apache.hadoop.hbase.regionserver.StoreFile$Writer.(StoreFile.java:787)\n9039 › at org.apache.hadoop.hbase.regionserver.StoreFile$Writer.(StoreFile.java:742)\n9040 › at org.apache.hadoop.hbase.regionserver.StoreFile$WriterBuilder.build(StoreFile.java:682)\n9041 › at org.apache.hadoop.hbase.regionserver.HStore.createWriterInTmp(HStore.java:1029)\n9042 › at org.apache.hadoop.hbase.regionserver.DefaultStoreFlusher.flushSnapshot(DefaultStoreFlusher.java:66)\n9043 › at org.apache.hadoop.hbase.regionserver.HStore.flushCache(HStore.java:932)\n9044 › at org.apache.hadoop.hbase.regionserver.HStore$StoreFlusherImpl.flushCache(HStore.java:2069)\n9045 › at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushCacheAndCommit(HRegion.java:2312)\n9046 › at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2048)\n9047 › at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2010)\n9048 › at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:1901)\n9049 › at org.apache.hadoop.hbase.regionserver.HRegion.flush(HRegion.java:1827)\n9050 › at org.apache.hadoop.hbase.regionserver.TestPerColumnFamilyFlush.testSelectiveFlushWhenNotEnabled(TestPerColumnFamilyFlush.java:303)\n9051 › at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n9052 › at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n9053 › at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n9054 › at java.lang.reflect.Method.invoke(Method.java:606)\n9055 › at org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:50)\n9056 › at org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\n9057 › at org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:47)\n9058 › at org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\n9059 › at org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:298)\n9060 › at org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:292)\n9061 › at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n9062 › at java.lang.Thread.run(Thread.java:744)\n{code}\n\nIt looks like an issue and only after looking in code do I see it informative. I should change this to be TRACE level with DEBUG emitting just that the feature is not available? (This is in testing master with its version of hadoop)","created":"2015-08-26T18:27:54.812+0000"}],"conversations":[{"body":"","from":"reporter","subject":"Allow dropping caches behind compactions"},{"body":"Related to HBASE-10052?","from":"developer"},{"body":"Pretty simple patch to try and see how this works.","from":"developer"},{"body":"Yeah related but on the other side.","from":"developer"},{"body":"Oh... This is for the output files of a compaction... Aren't we likely going to read the compacted data again (especially since we just dropped it all from the block cache)?\n","from":"developer"},{"body":"Yes it's possible that we will read the file again. It's also likely that on large compactions we will just be blowing the fs cache out of the water. Compacting 32gb on machine with 32gb memory free means that nothing else can be in the fs cache. No /bin/bash, no inodes, nothing.\n\nMy plan is likely to set this only for large compactions; the thought being that large compactions are much more likely to have stale data in them. I'm going to test the current patch out on a cluster that's doing really large compactions right now. If I see any positive changes then we can make this smarter.","from":"developer"},{"body":"That makes sense. Thanks [~eclark]. A cheap approximation might be to set it for major compactions. Agree that ideally we'd decided based on the size of the output file(s) compared to the available RAM for the OS buffer cache.\n","from":"developer"},{"body":"I tested this out yesterday on a cluster with 1.1.\nWhen compacting 100+ GB storfiles I can reliably make the linux kernel pause trying to thrash and find pages to clear out.\nWith just dropping the write's out of cache things are better and there are fewer stalls.\nWith dropping the caches behind writes and reads from a compaction there are no noticeable stalls ( ones that atop sees that take > 1 second ).\n\nSo I think that we should provide the ability to drop caches behind reads and writes on large compactions ( those that would run on the large compaction threads ). However because this could adversely affect people with less data per machine I think that we should have this turned off by default.\n\nThoughts?","from":"developer"},{"body":"Awesome! \n\nCache flush stalls can also be avoided by configuring the Linux disk cache. Could you do a {{sysctl -a | grep dirty}} on one of the boxes and paste vm.dirty_background_ratio, vm.dirty_background_bytes, and vm.dirty_ratio here?\n\nJust curious :) Your previous comment made sense anyway. Caching anything that is usually read sequentially and larger than the available RAM never makes sense.\n","from":"developer"},{"body":"{code}\nvm.dirty_background_ratio = 3\nvm.dirty_ratio = 20\n{code}\n\nThe interesting part is that these stalls aren't really caused by dirty pages. There's only 400mb of dirty pages. These stalls seem to be caused by thrashing the cache. Too much reading/writing actually causes reads to happen on the root disk to re-page in executables.","from":"developer"},{"body":"Patch with test fixes for stripe compaction.\n\nI put it up for review here: https://reviews.facebook.net/D42681","from":"developer"},{"body":"Style fixes","from":"developer"},{"body":"TestIOFencing works on my box 3 times in a row. Checking master.","from":"developer"},{"body":"Yeah that test looks good. I think it's just a little flakey.","from":"developer"},{"body":"Ping? This one solved kernel stalls for us.","from":"developer"},{"body":"bq.if (this.conf.getBoolean(\"hbase.regionserver.compaction.private.readers\", true)) \nSo with this change compactions will always use new readers so that the OS pages of those files are not cached based on the new setting that says dropBehind on compaction?","from":"developer"},{"body":"Was just going to +1... Missed the private readers bit.\nI added the private.readers features in order to decouple compactions readers from \"operational\" readers (user scans, gets, etc).\nDefensively I defaulted this to false, but I think we can safely enable that always... But if we wanted to stay defensive we could default to the value of \"drop caches behind compactions\".\n\nActually, in either case +1 from me.","from":"developer"},{"body":"The patch went stale while I was on vacation. Here's a rebase.\n\nYes this does turn on private readers by default. We've been running it for a while and I haven't seen any down sides, so I feel pretty sure that it's not too big a risk.","from":"developer"},{"body":"Whoops missed the mob file stuff being there.","from":"developer"},{"body":"Committing soon unless there are any other comments.","from":"developer"},{"body":"The patch didn't apply cleanly to branch-1. Here's the backport.","from":"developer"},{"body":"Pushed to master, branch-1, and branch-1.2","from":"developer"},{"body":"bq. Yes this does turn on private readers by default. We've been running it for a while and I haven't seen any down sides, so I feel pretty sure that it's not too big a risk.\n\nCool. I put it in a while back, but haven't seen anybody but us using it. Did you notice improvements in scan performance when you enabled it? (sorry off topic)\n\nCool to have this committed.","from":"developer"},{"body":"This change drops ugly thread dumps in our logs:\n\n{code}\n9034 2015-08-26 11:18:38,175 DEBUG [Time-limited test] hfile.HFile$WriterFactory(308): Unable to set drop behind on /Users/stack/checkouts/hbase.git.commit/hbase-server/target/test-data/ed67f436-1d46-43f#\n9035 java.lang.UnsupportedOperationException: the wrapped stream does not support setting the drop-behind caching setting.\n9036 › at org.apache.hadoop.fs.FSDataOutputStream.setDropBehind(FSDataOutputStream.java:150)\n9037 › at org.apache.hadoop.hbase.io.hfile.HFile$WriterFactory.create(HFile.java:306)\n9038 › at org.apache.hadoop.hbase.regionserver.StoreFile$Writer.(StoreFile.java:787)\n9039 › at org.apache.hadoop.hbase.regionserver.StoreFile$Writer.(StoreFile.java:742)\n9040 › at org.apache.hadoop.hbase.regionserver.StoreFile$WriterBuilder.build(StoreFile.java:682)\n9041 › at org.apache.hadoop.hbase.regionserver.HStore.createWriterInTmp(HStore.java:1029)\n9042 › at org.apache.hadoop.hbase.regionserver.DefaultStoreFlusher.flushSnapshot(DefaultStoreFlusher.java:66)\n9043 › at org.apache.hadoop.hbase.regionserver.HStore.flushCache(HStore.java:932)\n9044 › at org.apache.hadoop.hbase.regionserver.HStore$StoreFlusherImpl.flushCache(HStore.java:2069)\n9045 › at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushCacheAndCommit(HRegion.java:2312)\n9046 › at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2048)\n9047 › at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2010)\n9048 › at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:1901)\n9049 › at org.apache.hadoop.hbase.regionserver.HRegion.flush(HRegion.java:1827)\n9050 › at org.apache.hadoop.hbase.regionserver.TestPerColumnFamilyFlush.testSelectiveFlushWhenNotEnabled(TestPerColumnFamilyFlush.java:303)\n9051 › at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n9052 › at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n9053 › at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n9054 › at java.lang.reflect.Method.invoke(Method.java:606)\n9055 › at org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:50)\n9056 › at org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\n9057 › at org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:47)\n9058 › at org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\n9059 › at org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:298)\n9060 › at org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:292)\n9061 › at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n9062 › at java.lang.Thread.run(Thread.java:744)\n{code}\n\nIt looks like an issue and only after looking in code do I see it informative. I should change this to be TRACE level with DEBUG emitting just that the feature is not available? (This is in testing master with its version of hadoop)","from":"developer"}],"created":"2015-07-16T08:44:02.000+0000","description":"","issue_id":"12845520","key":"HBASE-14098","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2015-08-13T02:48:41.000+0000","role":"fixed_distractor","summary":"Allow dropping caches behind compactions"} {"case_id":"12857835","cluster":"DISTRACTOR-HBASE-14280","comments":[{"body":"{code}\n884\t 2.6.0\n{code}\nPlease don't upgrade hadoop dependency.\n\nCan you use reflection ?\n\nFormatting is off as well: indentation should be 4 spaces.","created":"2015-08-21T18:20:11.821+0000"},{"body":"New patch compatible with all version of hadoop","created":"2015-09-08T20:09:04.750+0000"},{"body":"{code}\n74\t }catch(NoSuchMethodError e){\n{code}\nPlease leave a space between right curly and catch, catch and '(', ')' and right curly.\n\n'else' should follow the right curly (on same line).","created":"2015-09-08T20:21:40.327+0000"},{"body":"Hi [~tedyu],\n\nPatch Updated with the review comments.\n\nRegards,\nAnkit Singhal","created":"2015-09-09T07:00:09.836+0000"},{"body":"lgtm\n\nWaiting for green QA run.","created":"2015-09-09T10:22:22.348+0000"},{"body":"Were checkstyle warnings related to the patch ?","created":"2015-09-16T14:01:07.354+0000"},{"body":"Checkstyle warnings are from the patch:\n{code}\n\n\n\n{code}","created":"2015-09-18T20:50:07.193+0000"},{"body":"Patch v4 addresses checkstyle warnings.","created":"2015-09-19T15:06:59.316+0000"},{"body":"Thanks for the patch, Ankit","created":"2015-09-19T17:57:38.378+0000"},{"body":"TestFSHDFSUtils fails in branch-1.1","created":"2015-09-21T16:49:16.975+0000"},{"body":"Fixed the test case failure on Branch 1.1","created":"2015-09-22T11:07:16.470+0000"},{"body":"Bulk closing 1.1.3 issues.","created":"2016-01-27T15:29:02.331+0000"}],"conversations":[{"body":"Caused by: org.apache.hadoop.hbase.ipc.RemoteWithExtrasException(java.io.IOException): java.io.IOException: Wrong FS: hdfs://ha-aggregation-nameservice1/hbase_upload/82c89692-6e78-46ef-bbea-c9e825318bfe/A/1aaaa31358d641c69d6c34b803c187b0, expected: hdfs://ha-hbase-nameservice1\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2113)\n\tat org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:108)\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor.consumerLoop(RpcExecutor.java:114)\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$1.run(RpcExecutor.java:94)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: java.lang.IllegalArgumentException: Wrong FS: hdfs://ha-aggregation-nameservice1/hbase_upload/82c89692-6e78-46ef-bbea-c9e825318bfe/A/1aaaa31358d641c69d6c34b803c187b0, expected: hdfs://ha-hbase-nameservice1\n\tat org.apache.hadoop.fs.FileSystem.checkPath(FileSystem.java:645)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem.getPathName(DistributedFileSystem.java:193)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem.access$000(DistributedFileSystem.java:105)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem$19.doCall(DistributedFileSystem.java:1136)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem$19.doCall(DistributedFileSystem.java:1132)\n\tat org.apache.hadoop.fs.FileSystemLinkResolver.resolve(FileSystemLinkResolver.java:81)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem.getFileStatus(DistributedFileSystem.java:1132)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:414)\n\tat org.apache.hadoop.fs.FileSystem.exists(FileSystem.java:1423)\n\tat org.apache.hadoop.hbase.regionserver.HRegionFileSystem.commitStoreFile(HRegionFileSystem.java:372)\n\tat org.apache.hadoop.hbase.regionserver.HRegionFileSystem.bulkLoadStoreFile(HRegionFileSystem.java:451)\n\tat org.apache.hadoop.hbase.regionserver.HStore.bulkLoadHFile(HStore.java:750)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.bulkLoadHFiles(HRegion.java:4894)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.bulkLoadHFiles(HRegion.java:4799)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.bulkLoadHFile(HRegionServer.java:3377)\n\tat org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:29996)\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2078)\n\t... 4 more\n\n\tat org.apache.hadoop.hbase.ipc.RpcClient.call(RpcClient.java:1498)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.callBlockingMethod(RpcClient.java:1684)\n\tat org.apache.hadoop.hbase.ipc.RpcClient$BlockingRpcChannelImplementation.callBlockingMethod(RpcClient.java:1737)\n\tat org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$BlockingStub.bulkLoadHFile(ClientProtos.java:29276)\n\tat org.apache.hadoop.hbase.protobuf.ProtobufUtil.bulkLoadHFile(ProtobufUtil.java:1548)\n\t... 11 more\n","from":"reporter","subject":"Bulk Upload from HA cluster to remote HA hbase cluster fails"},{"body":"{code}\n884\t 2.6.0\n{code}\nPlease don't upgrade hadoop dependency.\n\nCan you use reflection ?\n\nFormatting is off as well: indentation should be 4 spaces.","from":"developer"},{"body":"New patch compatible with all version of hadoop","from":"developer"},{"body":"{code}\n74\t }catch(NoSuchMethodError e){\n{code}\nPlease leave a space between right curly and catch, catch and '(', ')' and right curly.\n\n'else' should follow the right curly (on same line).","from":"developer"},{"body":"Hi [~tedyu],\n\nPatch Updated with the review comments.\n\nRegards,\nAnkit Singhal","from":"developer"},{"body":"lgtm\n\nWaiting for green QA run.","from":"developer"},{"body":"Were checkstyle warnings related to the patch ?","from":"developer"},{"body":"Checkstyle warnings are from the patch:\n{code}\n\n\n\n{code}","from":"developer"},{"body":"Patch v4 addresses checkstyle warnings.","from":"developer"},{"body":"Thanks for the patch, Ankit","from":"developer"},{"body":"TestFSHDFSUtils fails in branch-1.1","from":"developer"},{"body":"Fixed the test case failure on Branch 1.1","from":"developer"},{"body":"Bulk closing 1.1.3 issues.","from":"developer"}],"created":"2015-08-21T12:06:18.000+0000","description":"Caused by: org.apache.hadoop.hbase.ipc.RemoteWithExtrasException(java.io.IOException): java.io.IOException: Wrong FS: hdfs://ha-aggregation-nameservice1/hbase_upload/82c89692-6e78-46ef-bbea-c9e825318bfe/A/1aaaa31358d641c69d6c34b803c187b0, expected: hdfs://ha-hbase-nameservice1\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2113)\n\tat org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:108)\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor.consumerLoop(RpcExecutor.java:114)\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$1.run(RpcExecutor.java:94)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: java.lang.IllegalArgumentException: Wrong FS: hdfs://ha-aggregation-nameservice1/hbase_upload/82c89692-6e78-46ef-bbea-c9e825318bfe/A/1aaaa31358d641c69d6c34b803c187b0, expected: hdfs://ha-hbase-nameservice1\n\tat org.apache.hadoop.fs.FileSystem.checkPath(FileSystem.java:645)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem.getPathName(DistributedFileSystem.java:193)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem.access$000(DistributedFileSystem.java:105)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem$19.doCall(DistributedFileSystem.java:1136)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem$19.doCall(DistributedFileSystem.java:1132)\n\tat org.apache.hadoop.fs.FileSystemLinkResolver.resolve(FileSystemLinkResolver.java:81)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem.getFileStatus(DistributedFileSystem.java:1132)\n\tat org.apache.hadoop.fs.FilterFileSystem.getFileStatus(FilterFileSystem.java:414)\n\tat org.apache.hadoop.fs.FileSystem.exists(FileSystem.java:1423)\n\tat org.apache.hadoop.hbase.regionserver.HRegionFileSystem.commitStoreFile(HRegionFileSystem.java:372)\n\tat org.apache.hadoop.hbase.regionserver.HRegionFileSystem.bulkLoadStoreFile(HRegionFileSystem.java:451)\n\tat org.apache.hadoop.hbase.regionserver.HStore.bulkLoadHFile(HStore.java:750)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.bulkLoadHFiles(HRegion.java:4894)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.bulkLoadHFiles(HRegion.java:4799)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.bulkLoadHFile(HRegionServer.java:3377)\n\tat org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:29996)\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2078)\n\t... 4 more\n\n\tat org.apache.hadoop.hbase.ipc.RpcClient.call(RpcClient.java:1498)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.callBlockingMethod(RpcClient.java:1684)\n\tat org.apache.hadoop.hbase.ipc.RpcClient$BlockingRpcChannelImplementation.callBlockingMethod(RpcClient.java:1737)\n\tat org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$BlockingStub.bulkLoadHFile(ClientProtos.java:29276)\n\tat org.apache.hadoop.hbase.protobuf.ProtobufUtil.bulkLoadHFile(ProtobufUtil.java:1548)\n\t... 11 more\n","issue_id":"12857835","key":"HBASE-14280","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2015-09-22T16:23:42.000+0000","role":"fixed_distractor","summary":"Bulk Upload from HA cluster to remote HA hbase cluster fails"} {"case_id":"12857913","cluster":"DISTRACTOR-HBASE-14283","comments":[{"body":"We suspect the person in HBASE-13830 also ran into this bug, based on his similar exception, but can’t know for sure without more information from him.","created":"2015-08-21T17:31:08.823+0000"},{"body":"In the patch I also added an extra unit test to TestFromClientside, that also tests reverse scan, but is unrelated to this bug, just as an additional test to strengthen the test suite. ","created":"2015-08-21T17:34:27.973+0000"},{"body":"Anyone have any comments on this bug and the fix? Perhaps [~zjushch] (from HBASE-4811 which implemented reverse scan) or [~mikhail] (from HBASE-3857 which implemented HFileV2) or [~liyintang] (from HBASE-4532 which added the bloom filter method(s) I have a question about)? Or maybe one of them can tell me the person(s) who would be best suited to examine the bug and proposed fixes?","created":"2015-08-25T04:49:45.779+0000"},{"body":"Attaching the patch for QA. \n[~benlau]\nSuggest to rename the patch based on the JIRA id. Will take a look at the patch ASAP. Thanks for the patch. Lets see what the QA says. ","created":"2015-08-25T05:37:11.046+0000"},{"body":"Thanks [~ramkrishna.s.vasudevan@gmail.com] for showing me the conventions for the patch process. The tests look like they passed, with 1 minor style comment. I can fix that but want to address the bloom filter blocks issue first. Who would be best for me to ask about it, how about you? Is it reasonable to change the HFile.Reader interface so that the HFile reader (instead of the higher level StoreFile reader) is in charge of deserializing and holding the bloom filter data structures? Can't see why not but maybe I'm missing something.","created":"2015-08-25T18:44:41.988+0000"},{"body":"We use below method to get the previous block\npublic HFileBlock readBlock(long dataBlockOffset, long onDiskBlockSize,\n final boolean cacheBlock, boolean pread, final boolean isCompaction,\n boolean updateCacheMetrics, BlockType expectedBlockType,\n DataBlockEncoding expectedDataBlockEncoding)\n\nSo there no BlockType check and looping? May be it will read a block and see that block is not the expected one and go to next block and check for type. In seek before case instead of going fwd we should be going backward in case the expected block type is matching with the cur block type. That way of solution will work?","created":"2015-08-26T14:55:52.349+0000"},{"body":"Hi Anoop. The problem isn't that we read a previous block and see that the block is not the expected type. prevBlockOffset guarantees that we can seek to the previous block of the same type as the current one. See the comments on HFileBlock.getPrevBlockOffset(). We are always seeking to the previous data block, we are simply not calculating how much to read correctly once we have seeked to that previous data block because our prev data block size calculation can include other blocks because of the layout of scannable section in HFileV2+. We need a way of knowing apriori what the size of the previous data block is. The method you describe is used in HFileReaderImpl.readNextDataBlock(). Note that the reason this method works is because this method can use the method curBlock.getNextBlockOnDiskSizeWithHeader(). We need something similar to that when seeking backwards in order to achieve optimal performance. Let me know if I misunderstood what you meant. ","created":"2015-08-26T16:58:25.626+0000"},{"body":"[~benlau]\nLet me take a look at this tomorrow morning my time. ","created":"2015-08-26T17:17:12.622+0000"},{"body":"bq. Is it reasonable to change the HFile.Reader interface so that the HFile reader (instead of the higher level StoreFile reader)\nI think you can try this if you think that will help in the bloom area. Also it is an private interface so it is fine to change. ","created":"2015-08-27T16:26:24.627+0000"},{"body":"Thanks ramkrishna. I will take a crack at fixing the calculation in the presence of bloom filters later and see how an interface update looks. ","created":"2015-08-27T17:20:05.263+0000"},{"body":"Here's a V2 of the patch that handles bloom filter blocks. It requires some interface changes that blur the line a bit between the StoreFile reader and the HFile reader which is not ideal but there isn't really any other way to fix this currently in a performant way for bloom filters. Let me know what you guys think. I have attached the patch to the ticket for review/feedback. ","created":"2015-08-29T00:36:47.167+0000"},{"body":"{code}\n+ throw new IllegalArgumentException(\"Data block offset was provided that is actually \"\n+ + \"an index or bloom block offset: \" + dataBlockOffset1);\n{code}\nI understand the meaning of above message. But strictly speaking, offset is just a numeric value. Itself wouldn't indicate whether the block is a data block or index / bloom block.\n\nProviding number of index blocks found in the exception message would be helpful:\n{code}\n+ if (indexBlocksFound > 1) {\n+ throw new IllegalStateException(\"Found more than 1 block of this type between 2 \"\n+ + \"consecutive data blocks: \" + type);\n{code}\n{code}\n+ public BloomFilter getGeneralBloomFilter(BloomType bloomFilterType) throws IOException {\n+ if (generalBloomFilterLoaded) {\n+ return generalBloomFilter;\n{code}\nCan generalBloomFilterLoaded be replaced with checking generalBloomFilter not being null ? Similar comment applies to deleteBloomFilterLoaded\n{code}\n+ if (bloomFilterType == BloomType.NONE) {\n+ throw new IOException(\"Valid bloom filter type not found in FileInfo\");\n{code}\nCan information from FileInfo be included in the exception message to facilitate debugging ?\n\n\n","created":"2015-08-29T20:47:51.306+0000"},{"body":"Correct me if I'm wrong but I think I can't do a generalBloomFilter == null check because there are 3 states not 2: (1) Tried to load filter and it exists (2) Tried to load filter and it does not exist (3) Have not tried to load filter yet. If we rely on the generalBloomFilter == null check we can't distinguish between (2) and (3) which means we would end up trying to reload the filter unnecessarily.\n\n{quote}\nCan information from FileInfo be included in the exception message to facilitate debugging ?\n{quote}\n\nWhat information from FileInfo should be provided? The point of the message (maybe it needs to be revised for clarity) is that we found an unexpected BloomFilter in the HFile-- unexpected because the HFile FileInfo metadata claims there is no bloom filter (type = NONE). I will clarify the msg a bit. \n\nI'll fix the other issues, thanks.","created":"2015-08-31T21:06:34.066+0000"},{"body":"Can a singleton of type BloomFilter be created so that we can use it to denote case 2 ?\nMeaning generalBloomFilter is initially null. When we find no bloom filter exists, assign generalBloomFilter the singleton.\n\nThanks","created":"2015-08-31T21:13:34.263+0000"},{"body":"Putting patch on reviewboard would make review easier.\n\nThere're some typo's in current patch.","created":"2015-08-31T21:22:40.716+0000"},{"body":"The other parts of the code consider BloomFilter to be a nullable type, not just here. I don't know if it makes sense to change that in this patch. It is a bit overkill to use a null object here and seems to increase complexity more than it eliminates currently (new class to eliminate a load flag), since unlike in some common use cases of having a null object, we can't avoid checking for the null object here.","created":"2015-08-31T21:46:16.925+0000"},{"body":"Alright I'll put it on reviewboard, thanks.","created":"2015-08-31T21:46:59.193+0000"},{"body":"Alright, you can keep the checking as it is now.","created":"2015-08-31T21:50:10.998+0000"},{"body":"Reviewboard link with updated version of patch (v3): https://reviews.apache.org/r/37971/","created":"2015-08-31T23:05:44.090+0000"},{"body":"Going through the details of this issue and the patch. \nWe have the optimization of reading next block's header along with every block to avoid 2 reads (1st header and get data size and then do second read). This works well for forward read. In case of backward read, where we have to read the prev block, we tried to solve the case by calculating the prev blocks size from offset of 2 blocks (considering no other blocks in btw these 2 data blocks).. As per the bug, this assumption is not true. Good find. If we have the prev block data size also along with prev block offset, we were good. But we dont have that in current HFiles. When we dont know the data size of a block already, we are passing -1 so that it will read block by 2 reads.\nSeeing the patch it is very complex. And there are additional overhead also. Also it is not complete fix as it can not handle more than 2 level index block. So IMHO it will be better to avoid this kind of complex fix. \nWe can do the fix in 2 steps.\n1. Bump the minor version of the HFile and add the prev block size also along with offset. When this info is there in HFile we can safely use that and do read of prev block in one read.\n2. For reading old files where this meta data is NOT available, just pass -1 and let it read in 2 steps. It is ok.. Any way once the patch is applied and over the run the older files will get compacted to new file and this will have the additional meta info.","created":"2015-09-09T02:54:07.372+0000"},{"body":"Yep, we talked with Anoop and agree that the patch adds a lot of complexity for a fix that doesn't fix the issue 100%. The portion of the patch that is required to fix the bug for bloom filters is especially long. We thought to aim for a longer term fix later, but based on our discussion with Anoop it sounds like a backwards compatible, complete fix that adds the necessary metadata to HFile should not be too complicated/much work (eg does not involve creating new HFileReader implementation or other infrastructure). We will submit a new patch later with a final fix. We will keep the unit tests from the 1st patch since they are still applicable.","created":"2015-09-09T03:33:25.529+0000"},{"body":"Just to clarify, the prev block data size is what is missing here. So that is going to be added per block and this information is added to the hfile's metadata? So there is going to be a change in the HFileblock's SerDe format?","created":"2015-09-09T11:02:25.407+0000"},{"body":"Yes we would be changing the serialization for HFileBlock header. It would have a new field for the previous data block size, for block of the same type (same semantics as prevBlockOffset now). Any objections?","created":"2015-09-09T16:46:51.325+0000"},{"body":"Am fine. The HFile's metadata has to be used while reading the HFileblock. Can look at the patch once posted.","created":"2015-09-09T17:07:50.865+0000"},{"body":"Hey guys, I started looking into updating the HFile serialization to support reverse scans per previous comments. One thing that immediately struck me as being a possible problem is that the header sizes appear to be hardcoded into HConstants.java (HConstants.HFILEBLOCK_HEADER_SIZE), rather than being read from the HFile block header or HFile metadata itself. \n\nThis seems to imply that if I add more fields to the header and then do a rolling restart to update all region servers to have my code, any old region server that hasn't updated yet and is processing the new HFiles will not realize the header is bigger now and that there is stuff they need to skip / ignore. This might necessitate a 2-step restart process with 2 rolling restarts. \n\nRestart 1 to update all RS to have the appropriate new reading code. Restart 2 will enable writes by setting an HBase config option (false by default) to start writing the new HFiles. Am I missing something and this 2-step rolling restart is not necessary for some reason? It seems unlikely people would find this process palatable but is there a better alternative? \n\nAlternatively I can turn this into a non-backwards compatible major version update instead of a minor version update and require a full cluster restart but that is kind of harsh in its own way. Opinions/thoughts?","created":"2015-09-15T05:13:09.453+0000"},{"body":"After talking to some committers in HBase, it seems that unless there is a very strong case / no viable alternative, all new patches to HBase should not require a full cluster restart. Hence, we will be going with the 2-rolling-restart approach as described above. It requires the cluster operator to do 2 rolling restarts and set a new config but that should not be too burdensome for a major upgrade. This rolling-restart-compatible approach is a bit more messy/complicated code-wise so let us look a bit into the best way to do this.","created":"2015-09-15T23:24:33.853+0000"},{"body":"Hi guys, I have posted a new patch on review board. See https://reviews.apache.org/r/38720/. The patch adds support for HFileV4 and uses it to fix/optimize reverse scan. The patch is designed to be rolling-restartable in 2 phases, as discussed in an above comment on Sep 14.\n\nAs currently posted, I think the patch has to go into HBase 2.0 since it changes HConstants which is marked @Stable. It turns out that assumptions about the header size and contents are hardcoded in several places, primarily HConstants.java.\n\nLet me know what you guys think of the patch. More eyes would be better because the changes are somewhat farther reaching than they sounded initially and the HFile format has a long history that I’m not as familiar with as most people around here.\n\nI also need to reach out to Facebook later for feedback since this change seems like it will affect their external Memcached block cache. I have marked TODO in the patch in several places where I will need to talk with Facebook.\n\nAlso, since we are adding a new HFileV4, now would be a good time to fix anything else that is broken or add additional metadata for future optimizations that we missed out on in HFileV3. Someone could add more metadata after the patch stands on its own as a complete fix for reverse scans. ","created":"2015-09-24T17:33:43.616+0000"},{"body":"HFileV4 would be a pretty drastic change. I suggest you raise a thread on dev to discuss what other things we might want to add/change with a v4.","created":"2015-09-24T17:36:31.814+0000"},{"body":"Hi Nick, that sounds like a good idea, I will do that.","created":"2015-09-24T17:50:09.329+0000"},{"body":"We could do a v4 for 2.0. Should run this idea by [~mbertozzi] ","created":"2015-09-26T01:23:05.798+0000"},{"body":"v4 probably is a bad name.. \nthis is more a v2/3.x since (the original idea) was just a fix in the format to solve a bug.\nnot a complete format redesign. everything should be the same aside adding one field on the level block. \nBen was also suggesting a double rolling upgrade to migrate to that, more or less like we did for 3.\nI'm ok to have it only in 2.0, but to me this can also be done in a 1.3. just switch the name from v4 to 3.x :)\n\nif we want to extend the scope of redesigning the file-format, that will be another topic.\nI know that [~lhofhansl], during hbasecon, was suggesting format changes to group by qualifiers and similar.\nbut that will not be just related to fixing this bug and it will probably be in 2.0 only I guess.","created":"2015-09-26T01:58:05.068+0000"},{"body":"As Matteo mentions, we discussed this issue briefly with him and Stack about a week ago or so. [~mbertozzi] I thought we had asked if updating the major version would be fine and the answer was in the affirmative but I might've misunderstood or misunderstood the degree of affirmation. (IIRC the reason was something along the lines of cluster operators updating HFile.FORMAT_VERSION_KEY in their config being a desirable property of the 2nd rolling upgrade or something.)\n\nIn any case though the reason we initially were going to do a 3.x but moved to a 4.0 was because (1) major versions (but not in HFiles?) generally denote level of backward compatibility and the new HFiles produced in this patch cannot be read by an HFileV3.X reader (2) the patch requires enough changes to assumptions in the serialization code (eg regarding header size or block cache) that it doesn't seem appropriate as a minor version change and (3) if we are following the rules we've set for ourselves, the changes in HConstants alone (annotated as Stable) mean this patch should be going into HBase 2.0 (admittedly rules can always be bent).\n\nUpdating only the minor version because the format change currently fixes only 1 bug (as opposed to 10 bugs or adding a new feature) seems to be the wrong way to think of versioning, IMO. If our concern is that we would like a fix for this bug in 1.3 and not wait until 2.0, we could also commit a shorter term fix for 1.3 that just always reads the block header or does an optimistic read and falls back on reading the block header if the read fails from block size expectations (configurable, optimistic off by default). In combination with an expected size correction for index blocks (perhaps not for bloom filter blocks since that fix is a messy addition and also violates some API abstraction layers in StoreFile) it might be fine in most scenarios, especially if the cluster operator is allowed to change the inline block chunking config for the cluster that needs to do reverse scans. Within Yahoo internally we will probably go with a bandaid fix like this for now so that users can use reverse scans and still get 'ok' even if not max performance. (Also sub-100% performance is better than getting exceptions about block sizes ;) )\n\nIf people would be okay with dividing it up this way-- short term fix (no HFile changes) for 1.3 and a longer term fix (HFileV4) for HBase 2.0 I could create a separate ticket with the newest patch as a starting point for HFileV4 and submit another patch for this ticket that implements a configurable 'choose your tradeoff' fix as described in the previous paragraph.","created":"2015-09-26T03:50:40.624+0000"},{"body":"I was making the point for a 2/3.x just because, people was already scared by the name v4 for the format. even if there are no real changes to the format itself aside few things. also we had already an \"incompatibile\" change in a minor version of hfile (the reason the minor version was introduced) and it was more or less the same thing as this one. In that case was adding a field to the header for checksum HBASE-5074.\n\nin any case i'm pretty sure we can get the changes in a 1.x too. off by default and people can decide if they care about that or not. more or less the same thing happened for the hfile v2.1 (checksum), v2.2 (protobufs) and v3 (tags). \nbut that's not my call, the next RM for 1.3 will make the call. but you have a +1 for me to change the hfile format in 2.0\n","created":"2015-09-26T04:22:01.572+0000"},{"body":"bq. I know that Lars Hofhansl, during hbasecon, was suggesting format changes to group by qualifiers and similar.\n\nI've suggested many things over the years :)\nI think this was about storing keys and values separately in HFile is indexing each with the start position. So if now scan with a filter, we can slog through the filters pretty quickly and only reassemble those Cells that match. Or we can do fast aggregates over the value... Both need significant other plumbing as well.","created":"2015-09-26T05:12:04.471+0000"},{"body":"When we were designing tags we accepted some limitations of HFile that later were problematic, specifically, we couldn't vary cell encoding on a block by block basis. Even if no cells use tags in a file, we'd bloat each cell with a short. Later we introduced a whole file optimization for this issue but clearly we'd have more opportunities to employ it if we could vary encoding strategy on a block-by-block basis. I thought about introducing an extensible pbufed block header. It didn't make sense for the tag serialization issue - we would trade one type of bloat for another, those additional header bytes will end up in the block cache - but if there are multiple use cases for it lined up, a new extensible pbufed 'block header' could be worthwhile. Would make future block level changes less likely to be incompatible changes too. \n\nDoes the introduction of something like that require a major version bump? I think so. I'd like to see us be more like semver with HFile versioning, and if we're on the same page about that, this is a major version bump because earlier versioned readers won't be able to handle the change. \n\nAlso opens the door to interesting things like using different block encoding strategies on a block by block basis according to characteristics of the cells to be encoded within.","created":"2015-09-26T17:28:42.755+0000"},{"body":"We can fix this bug now (in all applicable versions) by reading the header and then data. This 2 reads have perf impact but better no bug. Can we get that in first? For the 2.0 version (branch-1 also?) we can decide on how to add the new data to header (with major version bump or minor bump).. Fixing this bug as such we can give priority and get it done soon?","created":"2015-10-05T11:13:01.343+0000"},{"body":"Hey guys, sorry, I should be able to get back to this soon. Finishing up an unrelated project right now. I didn't know that minor versions in HFiles were also non-backwards compatible. That's one less reason then to make this a major version bump. If anyone has a strong preference for this fix to go into a V3.X I can change the patch to use minor version (eg for header size calculation) when I have time to do it. If not I'll leave it as V4 since it's a little simpler in the code as a major version bump. My original intention btw if it wasn't clear was that this wouldn't be the only change in a V4, just the first change that would go into a V4, whose format/contents is not yet meant to be final even when this patch is committed, i.e. V4 would be essentially a WIP with more changes suggested and implemented in other tickets and eventually released in HBase 2.0.\n\n[~anoop.hbase] I'm down for committing a short-term read-the-header-always fix for now and then discussing the longer term solution second. Which branches do you want the patch for?","created":"2015-10-07T04:07:03.273+0000"},{"body":"I would say commit for all applicable branches including Trunk. The better fix wherever we need (trunk , branch-1) pls add a TODO in the patch with Jira# which will fix that.","created":"2015-10-07T10:59:03.931+0000"},{"body":"Alright I'll work on that. I created HBASE-14576 for the longer term fix, so this ticket will just be to implement the short term fix we described.","created":"2015-10-08T00:30:41.685+0000"},{"body":"Short term patches for this bug per discussion.","created":"2015-10-08T04:17:16.018+0000"},{"body":"Hmmm, can't seem to edit Jira comments. Anyways, I attached short term patches to the ticket per discussion for all versions of HBase from 0.98 to master. The patches are mostly the same other than 0.98. There were a couple of utility methods that were missing/private in earlier versions of HBase 1.X that I backported or changed to public to mirror the master version of the patch. (Let me know if that isn't kosher for some reason.) I probably won't have time to update the patches this week but I'll look at feedback and make appropriate changes when I get the chance next week.","created":"2015-10-08T04:23:16.289+0000"},{"body":"lgtm\n\nQA report picked up hadoop tests.","created":"2015-10-08T07:52:21.949+0000"},{"body":"+1","created":"2015-10-09T10:39:43.949+0000"},{"body":"So anything I should change in the patches? How many +1's are needed? Does someone else need to +1? ","created":"2015-10-14T04:23:28.692+0000"},{"body":"Attach the same patch once more for a clean QA run.. Then we can commit it.. Ted can you pls commit it after QA? ","created":"2015-10-14T04:34:24.299+0000"},{"body":"Am I understanding correctly that we always incur two reads now, even when we're scanning forward? If so, that seems unfortunate.","created":"2015-10-14T04:45:48.539+0000"},{"body":"No it is for backward scan only. [~benlau] confirm once if am wrong.","created":"2015-10-14T04:58:53.644+0000"},{"body":"[~lhofhansl] and [~anoop.hbase], I probably should’ve called it out explicitly but yes technically, it would affect forward scans (among other things), but I don’t think in any noticeable way. The patched code is called by HalfStoreFileReader.getLastKey(). \n\nThis getLastKey() method is needed to figure out the split point of a region which is eg used in splitting a region. Extra IO op there but given that region splits don’t happen frequently nor at a very high rate when they do I think it is fine. The getLastKey() is also used by StoreFile.Reader for passesKeyRangeFilter() check to determine whether a store is applicable to a scan, both forward/reverse. It affects the initial scan creation as well as later next() RPC calls I think. Although this is too bad, it doesn’t really matter I think in practice, because the block of the last key will get cached in the BlockCache the first time we need to know the last key for that halfstore. So repeated calls later in the region server will not incur any overhead. Let me know if this addresses your concerns or not.\n\nIncidentally I think that caching is one of the primary reasons why there seems to be a decent # of people using reverse scan but almost no one reporting this bug— because caching often hides it, either completely or nondeterministically (requests succeeding on later retries). We had some problems initially reproducing this bug reliably because if we ran forward scans in a table concurrently with the reverse scans, it would cause the blocks to become cached and so certain key ranges that previously would’ve caused reverse scan to fail suddenly started working just fine.\n\nRe: Attaching the same patch, let me know if I'm doing this wrong but it sounds like all I should have to do is just upload the same patch file for master but with a different name and the QA tests will pick it up. I will do that.","created":"2015-10-14T09:24:34.993+0000"},{"body":"Attached a new patch for master, same as the previous patch but with 'reupload' in the name..","created":"2015-10-14T09:26:37.333+0000"},{"body":"bq. Although this is too bad, it doesn’t really matter I think in practice, because the block of the last key will get cached in the BlockCache the first time we need to know the last key for that halfstore. So repeated calls later in the region server will not incur any overhead. Let me know if this addresses your concerns or not.\n\nWell... There are scanning workloads where the data being scanned won't fit into the blockcache and so the user would be calling Scan#setCacheBlocks(boolean) with 'false'.\n\nIf we are only doing twice as many reads as before per block only for HalfStoreFileReader, then I think this is ok, because HalfStoreFileReader is only used between the time when a split takes place and when compaction on the daughters is complete, references are removed, and the parent is deleted. Please add comments indicating where we are taking on suboptimal behavior. \n\nIf we will always issue two reads for a block in all cases, then that is really unfortunate and I'd ask that be revisited.","created":"2015-10-16T00:32:26.460+0000"},{"body":"[~andrew.purtell@gmail.com] Let me know if I'm missing something, but I think there is more than 1 scan in play here. I think you're talking about an external hbase client scan. I'm talking about an internal hbase scan opened up by the regionserver and which we know for a fact caches the block. See the implementation of HalfStoreFileReader.getLastKey(), it is creating an internal scanner that does cache. Furthermore, the results of the method aren't cached in the Reader class (eg as a variable) and since the method is called repeatedly in the codebase it seems likely that the author's expectation was that the block cache would work correctly and make an internal cache for the file reader redundant. So not only does this scenario happen only for newly split regions but it only happens for the first time. I can add more comments to the patches if it is really necessary but there is already a comment in the code indicating that this fix is not performant and is meant to be updated by a later ticket whose jira # is listed.","created":"2015-10-16T01:29:11.837+0000"},{"body":"[~andrew.purtell@gmail.com] Is the above explanation agreeable?","created":"2015-10-20T21:20:43.390+0000"},{"body":"Still there [~andrew.purtell@gmail.com]? Anyone else have comments/questions?","created":"2015-10-26T17:25:15.760+0000"},{"body":"Ok, agreeable.","created":"2015-10-26T21:04:03.881+0000"},{"body":"Thanks Andrew. Can we merge these patches or is there still something that needs to be done or reviewed?","created":"2015-10-26T21:18:25.456+0000"},{"body":"[~benlau] I don't see more outstanding feedback and there's already a +1 here from Anoop. Let me check on some things locally and then commit if it looks good.","created":"2015-10-26T21:21:43.293+0000"},{"body":"Pushed to 0.98 and up.","created":"2015-10-26T23:45:17.519+0000"},{"body":"That's odd, something went wrong with the patch for 1.1 branch. That compiled fine for me but perhaps I overlooked something, or the patch became stale. I will take a look later tonight.","created":"2015-10-27T01:30:37.259+0000"},{"body":"Not a problem. This compiled for me too so I thought but there's a small missing static helper function in CellUtil. I'll commit it as an addendum now.","created":"2015-10-27T01:38:49.033+0000"},{"body":"No worries. Just pushed this small addendum to branch-1.1. ","created":"2015-10-27T01:40:00.340+0000"},{"body":"Sorry, I was trying to commit HBASE-14682, and pushed the addendum in the process without checking here. ","created":"2015-10-27T01:41:38.418+0000"},{"body":"I'll back mine out because the build is broken again (two CellUtil#matchingTimestamps) \n","created":"2015-10-27T01:44:29.827+0000"},{"body":"[~andrew.purtell@gmail.com] Was the wrong patch applied to 1.1 or am I misunderstanding something? So I looked at https://github.com/apache/hbase/commits/branch-1.1 specifically https://github.com/apache/hbase/commit/0db04a1705e5e8cc04cc9c010ddfc5612f60cfec and it is missing the CellUtil method that I had in my HBASE-14283-branch-1.1.patch attached. It's fine if this is fixed as an addendum but there's nothing else amiss right (i.e. the other branches have the right patches applied?)","created":"2015-10-27T02:40:52.449+0000"},{"body":"No worries Ben. Not sure how I can help. The other branches got the right patches. ","created":"2015-10-27T02:54:00.891+0000"},{"body":"Ok thanks just wanted to confirm the other branches were fine as far as we know, eg we didnt swap 1.1 patch with 0.98 patch and need to fix/update 0.98 branch. Thanks. It looks like there are some test failures but I think they are not related to the patch. I will rerun the tests that fail tonight locally after they appear in this ticket.","created":"2015-10-27T03:00:25.925+0000"},{"body":"Thanks for the perseverance [~benlau].. Hope u will now work on the other Jira to handle it in better way wrt perf impact. Thanks...","created":"2015-10-27T05:20:49.001+0000"},{"body":"Reran the failures that looked relevant on the various branches, seems the tests are just unstable. When I get the time I'll try to restart the discussion about updating the HFile serialization for more efficient reverse scans in HBASE-14576. ","created":"2015-10-27T08:44:27.440+0000"},{"body":"Bulk closing 1.1.3 issues.","created":"2016-01-27T15:28:27.737+0000"}],"conversations":[{"body":"Reverse scans do not work if an HFile contains inline bloom blocks or leaf level index blocks. The reason is because the seekBefore() call calculates the previous data block’s size by assuming data blocks are contiguous which is not the case in HFile V2 and beyond.\n\nAttached is a first cut patch (targeting bcef28eefaf192b0ad48c8011f98b8e944340da5 on trunk) which includes:\n(1) a unit test which exposes the bug and demonstrates failures for both inline bloom blocks and inline index blocks\n(2) a proposed fix for inline index blocks that does not require a new HFile version change, but is only performant for 1 and 2-level indexes and not 3+. 3+ requires an HFile format update for optimal performance. \n\nThis patch does not fix the bloom filter blocks bug. But the fix should be similar to the case of inline index blocks. The reason I haven’t made the change yet is I want to confirm that you guys would be fine with me revising the HFile.Reader interface.\n\nSpecifically, these 2 functions (getGeneralBloomFilterMetadata and getDeleteBloomFilterMetadata) need to return the BloomFilter. Right now the HFileReader class doesn’t have a reference to the bloom filters (and hence their indices) and only constructs the IO streams and hence has no way to know where the bloom blocks are in the HFile. It seems that the HFile.Reader bloom method comments state that they “know nothing about how that metadata is structured” but I do not know if that is a requirement of the abstraction (why?) or just an incidental current property. \n\nWe would like to do 3 things with community approval:\n(1) Update the HFile.Reader interface and implementation to contain and return BloomFilters directly rather than unstructured IO streams\n(2) Merge the fixes for index blocks and bloom blocks into open source\n(3) Create a new Jira ticket for open source HBase to add a ‘prevBlockSize’ field in the block header in the next HFile version, so that seekBefore() calls can not only be correct but performant in all cases.","from":"reporter","subject":"Reverse scan doesn’t work with HFile inline index/bloom blocks"},{"body":"We suspect the person in HBASE-13830 also ran into this bug, based on his similar exception, but can’t know for sure without more information from him.","from":"developer"},{"body":"In the patch I also added an extra unit test to TestFromClientside, that also tests reverse scan, but is unrelated to this bug, just as an additional test to strengthen the test suite. ","from":"developer"},{"body":"Anyone have any comments on this bug and the fix? Perhaps [~zjushch] (from HBASE-4811 which implemented reverse scan) or [~mikhail] (from HBASE-3857 which implemented HFileV2) or [~liyintang] (from HBASE-4532 which added the bloom filter method(s) I have a question about)? Or maybe one of them can tell me the person(s) who would be best suited to examine the bug and proposed fixes?","from":"developer"},{"body":"Attaching the patch for QA. \n[~benlau]\nSuggest to rename the patch based on the JIRA id. Will take a look at the patch ASAP. Thanks for the patch. Lets see what the QA says. ","from":"developer"},{"body":"Thanks [~ramkrishna.s.vasudevan@gmail.com] for showing me the conventions for the patch process. The tests look like they passed, with 1 minor style comment. I can fix that but want to address the bloom filter blocks issue first. Who would be best for me to ask about it, how about you? Is it reasonable to change the HFile.Reader interface so that the HFile reader (instead of the higher level StoreFile reader) is in charge of deserializing and holding the bloom filter data structures? Can't see why not but maybe I'm missing something.","from":"developer"},{"body":"We use below method to get the previous block\npublic HFileBlock readBlock(long dataBlockOffset, long onDiskBlockSize,\n final boolean cacheBlock, boolean pread, final boolean isCompaction,\n boolean updateCacheMetrics, BlockType expectedBlockType,\n DataBlockEncoding expectedDataBlockEncoding)\n\nSo there no BlockType check and looping? May be it will read a block and see that block is not the expected one and go to next block and check for type. In seek before case instead of going fwd we should be going backward in case the expected block type is matching with the cur block type. That way of solution will work?","from":"developer"},{"body":"Hi Anoop. The problem isn't that we read a previous block and see that the block is not the expected type. prevBlockOffset guarantees that we can seek to the previous block of the same type as the current one. See the comments on HFileBlock.getPrevBlockOffset(). We are always seeking to the previous data block, we are simply not calculating how much to read correctly once we have seeked to that previous data block because our prev data block size calculation can include other blocks because of the layout of scannable section in HFileV2+. We need a way of knowing apriori what the size of the previous data block is. The method you describe is used in HFileReaderImpl.readNextDataBlock(). Note that the reason this method works is because this method can use the method curBlock.getNextBlockOnDiskSizeWithHeader(). We need something similar to that when seeking backwards in order to achieve optimal performance. Let me know if I misunderstood what you meant. ","from":"developer"},{"body":"[~benlau]\nLet me take a look at this tomorrow morning my time. ","from":"developer"},{"body":"bq. Is it reasonable to change the HFile.Reader interface so that the HFile reader (instead of the higher level StoreFile reader)\nI think you can try this if you think that will help in the bloom area. Also it is an private interface so it is fine to change. ","from":"developer"},{"body":"Thanks ramkrishna. I will take a crack at fixing the calculation in the presence of bloom filters later and see how an interface update looks. ","from":"developer"},{"body":"Here's a V2 of the patch that handles bloom filter blocks. It requires some interface changes that blur the line a bit between the StoreFile reader and the HFile reader which is not ideal but there isn't really any other way to fix this currently in a performant way for bloom filters. Let me know what you guys think. I have attached the patch to the ticket for review/feedback. ","from":"developer"},{"body":"{code}\n+ throw new IllegalArgumentException(\"Data block offset was provided that is actually \"\n+ + \"an index or bloom block offset: \" + dataBlockOffset1);\n{code}\nI understand the meaning of above message. But strictly speaking, offset is just a numeric value. Itself wouldn't indicate whether the block is a data block or index / bloom block.\n\nProviding number of index blocks found in the exception message would be helpful:\n{code}\n+ if (indexBlocksFound > 1) {\n+ throw new IllegalStateException(\"Found more than 1 block of this type between 2 \"\n+ + \"consecutive data blocks: \" + type);\n{code}\n{code}\n+ public BloomFilter getGeneralBloomFilter(BloomType bloomFilterType) throws IOException {\n+ if (generalBloomFilterLoaded) {\n+ return generalBloomFilter;\n{code}\nCan generalBloomFilterLoaded be replaced with checking generalBloomFilter not being null ? Similar comment applies to deleteBloomFilterLoaded\n{code}\n+ if (bloomFilterType == BloomType.NONE) {\n+ throw new IOException(\"Valid bloom filter type not found in FileInfo\");\n{code}\nCan information from FileInfo be included in the exception message to facilitate debugging ?\n\n\n","from":"developer"},{"body":"Correct me if I'm wrong but I think I can't do a generalBloomFilter == null check because there are 3 states not 2: (1) Tried to load filter and it exists (2) Tried to load filter and it does not exist (3) Have not tried to load filter yet. If we rely on the generalBloomFilter == null check we can't distinguish between (2) and (3) which means we would end up trying to reload the filter unnecessarily.\n\n{quote}\nCan information from FileInfo be included in the exception message to facilitate debugging ?\n{quote}\n\nWhat information from FileInfo should be provided? The point of the message (maybe it needs to be revised for clarity) is that we found an unexpected BloomFilter in the HFile-- unexpected because the HFile FileInfo metadata claims there is no bloom filter (type = NONE). I will clarify the msg a bit. \n\nI'll fix the other issues, thanks.","from":"developer"},{"body":"Can a singleton of type BloomFilter be created so that we can use it to denote case 2 ?\nMeaning generalBloomFilter is initially null. When we find no bloom filter exists, assign generalBloomFilter the singleton.\n\nThanks","from":"developer"},{"body":"Putting patch on reviewboard would make review easier.\n\nThere're some typo's in current patch.","from":"developer"},{"body":"The other parts of the code consider BloomFilter to be a nullable type, not just here. I don't know if it makes sense to change that in this patch. It is a bit overkill to use a null object here and seems to increase complexity more than it eliminates currently (new class to eliminate a load flag), since unlike in some common use cases of having a null object, we can't avoid checking for the null object here.","from":"developer"},{"body":"Alright I'll put it on reviewboard, thanks.","from":"developer"},{"body":"Alright, you can keep the checking as it is now.","from":"developer"},{"body":"Reviewboard link with updated version of patch (v3): https://reviews.apache.org/r/37971/","from":"developer"},{"body":"Going through the details of this issue and the patch. \nWe have the optimization of reading next block's header along with every block to avoid 2 reads (1st header and get data size and then do second read). This works well for forward read. In case of backward read, where we have to read the prev block, we tried to solve the case by calculating the prev blocks size from offset of 2 blocks (considering no other blocks in btw these 2 data blocks).. As per the bug, this assumption is not true. Good find. If we have the prev block data size also along with prev block offset, we were good. But we dont have that in current HFiles. When we dont know the data size of a block already, we are passing -1 so that it will read block by 2 reads.\nSeeing the patch it is very complex. And there are additional overhead also. Also it is not complete fix as it can not handle more than 2 level index block. So IMHO it will be better to avoid this kind of complex fix. \nWe can do the fix in 2 steps.\n1. Bump the minor version of the HFile and add the prev block size also along with offset. When this info is there in HFile we can safely use that and do read of prev block in one read.\n2. For reading old files where this meta data is NOT available, just pass -1 and let it read in 2 steps. It is ok.. Any way once the patch is applied and over the run the older files will get compacted to new file and this will have the additional meta info.","from":"developer"},{"body":"Yep, we talked with Anoop and agree that the patch adds a lot of complexity for a fix that doesn't fix the issue 100%. The portion of the patch that is required to fix the bug for bloom filters is especially long. We thought to aim for a longer term fix later, but based on our discussion with Anoop it sounds like a backwards compatible, complete fix that adds the necessary metadata to HFile should not be too complicated/much work (eg does not involve creating new HFileReader implementation or other infrastructure). We will submit a new patch later with a final fix. We will keep the unit tests from the 1st patch since they are still applicable.","from":"developer"},{"body":"Just to clarify, the prev block data size is what is missing here. So that is going to be added per block and this information is added to the hfile's metadata? So there is going to be a change in the HFileblock's SerDe format?","from":"developer"},{"body":"Yes we would be changing the serialization for HFileBlock header. It would have a new field for the previous data block size, for block of the same type (same semantics as prevBlockOffset now). Any objections?","from":"developer"},{"body":"Am fine. The HFile's metadata has to be used while reading the HFileblock. Can look at the patch once posted.","from":"developer"},{"body":"Hey guys, I started looking into updating the HFile serialization to support reverse scans per previous comments. One thing that immediately struck me as being a possible problem is that the header sizes appear to be hardcoded into HConstants.java (HConstants.HFILEBLOCK_HEADER_SIZE), rather than being read from the HFile block header or HFile metadata itself. \n\nThis seems to imply that if I add more fields to the header and then do a rolling restart to update all region servers to have my code, any old region server that hasn't updated yet and is processing the new HFiles will not realize the header is bigger now and that there is stuff they need to skip / ignore. This might necessitate a 2-step restart process with 2 rolling restarts. \n\nRestart 1 to update all RS to have the appropriate new reading code. Restart 2 will enable writes by setting an HBase config option (false by default) to start writing the new HFiles. Am I missing something and this 2-step rolling restart is not necessary for some reason? It seems unlikely people would find this process palatable but is there a better alternative? \n\nAlternatively I can turn this into a non-backwards compatible major version update instead of a minor version update and require a full cluster restart but that is kind of harsh in its own way. Opinions/thoughts?","from":"developer"},{"body":"After talking to some committers in HBase, it seems that unless there is a very strong case / no viable alternative, all new patches to HBase should not require a full cluster restart. Hence, we will be going with the 2-rolling-restart approach as described above. It requires the cluster operator to do 2 rolling restarts and set a new config but that should not be too burdensome for a major upgrade. This rolling-restart-compatible approach is a bit more messy/complicated code-wise so let us look a bit into the best way to do this.","from":"developer"},{"body":"Hi guys, I have posted a new patch on review board. See https://reviews.apache.org/r/38720/. The patch adds support for HFileV4 and uses it to fix/optimize reverse scan. The patch is designed to be rolling-restartable in 2 phases, as discussed in an above comment on Sep 14.\n\nAs currently posted, I think the patch has to go into HBase 2.0 since it changes HConstants which is marked @Stable. It turns out that assumptions about the header size and contents are hardcoded in several places, primarily HConstants.java.\n\nLet me know what you guys think of the patch. More eyes would be better because the changes are somewhat farther reaching than they sounded initially and the HFile format has a long history that I’m not as familiar with as most people around here.\n\nI also need to reach out to Facebook later for feedback since this change seems like it will affect their external Memcached block cache. I have marked TODO in the patch in several places where I will need to talk with Facebook.\n\nAlso, since we are adding a new HFileV4, now would be a good time to fix anything else that is broken or add additional metadata for future optimizations that we missed out on in HFileV3. Someone could add more metadata after the patch stands on its own as a complete fix for reverse scans. ","from":"developer"},{"body":"HFileV4 would be a pretty drastic change. I suggest you raise a thread on dev to discuss what other things we might want to add/change with a v4.","from":"developer"},{"body":"Hi Nick, that sounds like a good idea, I will do that.","from":"developer"},{"body":"We could do a v4 for 2.0. Should run this idea by [~mbertozzi] ","from":"developer"},{"body":"v4 probably is a bad name.. \nthis is more a v2/3.x since (the original idea) was just a fix in the format to solve a bug.\nnot a complete format redesign. everything should be the same aside adding one field on the level block. \nBen was also suggesting a double rolling upgrade to migrate to that, more or less like we did for 3.\nI'm ok to have it only in 2.0, but to me this can also be done in a 1.3. just switch the name from v4 to 3.x :)\n\nif we want to extend the scope of redesigning the file-format, that will be another topic.\nI know that [~lhofhansl], during hbasecon, was suggesting format changes to group by qualifiers and similar.\nbut that will not be just related to fixing this bug and it will probably be in 2.0 only I guess.","from":"developer"},{"body":"As Matteo mentions, we discussed this issue briefly with him and Stack about a week ago or so. [~mbertozzi] I thought we had asked if updating the major version would be fine and the answer was in the affirmative but I might've misunderstood or misunderstood the degree of affirmation. (IIRC the reason was something along the lines of cluster operators updating HFile.FORMAT_VERSION_KEY in their config being a desirable property of the 2nd rolling upgrade or something.)\n\nIn any case though the reason we initially were going to do a 3.x but moved to a 4.0 was because (1) major versions (but not in HFiles?) generally denote level of backward compatibility and the new HFiles produced in this patch cannot be read by an HFileV3.X reader (2) the patch requires enough changes to assumptions in the serialization code (eg regarding header size or block cache) that it doesn't seem appropriate as a minor version change and (3) if we are following the rules we've set for ourselves, the changes in HConstants alone (annotated as Stable) mean this patch should be going into HBase 2.0 (admittedly rules can always be bent).\n\nUpdating only the minor version because the format change currently fixes only 1 bug (as opposed to 10 bugs or adding a new feature) seems to be the wrong way to think of versioning, IMO. If our concern is that we would like a fix for this bug in 1.3 and not wait until 2.0, we could also commit a shorter term fix for 1.3 that just always reads the block header or does an optimistic read and falls back on reading the block header if the read fails from block size expectations (configurable, optimistic off by default). In combination with an expected size correction for index blocks (perhaps not for bloom filter blocks since that fix is a messy addition and also violates some API abstraction layers in StoreFile) it might be fine in most scenarios, especially if the cluster operator is allowed to change the inline block chunking config for the cluster that needs to do reverse scans. Within Yahoo internally we will probably go with a bandaid fix like this for now so that users can use reverse scans and still get 'ok' even if not max performance. (Also sub-100% performance is better than getting exceptions about block sizes ;) )\n\nIf people would be okay with dividing it up this way-- short term fix (no HFile changes) for 1.3 and a longer term fix (HFileV4) for HBase 2.0 I could create a separate ticket with the newest patch as a starting point for HFileV4 and submit another patch for this ticket that implements a configurable 'choose your tradeoff' fix as described in the previous paragraph.","from":"developer"},{"body":"I was making the point for a 2/3.x just because, people was already scared by the name v4 for the format. even if there are no real changes to the format itself aside few things. also we had already an \"incompatibile\" change in a minor version of hfile (the reason the minor version was introduced) and it was more or less the same thing as this one. In that case was adding a field to the header for checksum HBASE-5074.\n\nin any case i'm pretty sure we can get the changes in a 1.x too. off by default and people can decide if they care about that or not. more or less the same thing happened for the hfile v2.1 (checksum), v2.2 (protobufs) and v3 (tags). \nbut that's not my call, the next RM for 1.3 will make the call. but you have a +1 for me to change the hfile format in 2.0\n","from":"developer"},{"body":"bq. I know that Lars Hofhansl, during hbasecon, was suggesting format changes to group by qualifiers and similar.\n\nI've suggested many things over the years :)\nI think this was about storing keys and values separately in HFile is indexing each with the start position. So if now scan with a filter, we can slog through the filters pretty quickly and only reassemble those Cells that match. Or we can do fast aggregates over the value... Both need significant other plumbing as well.","from":"developer"},{"body":"When we were designing tags we accepted some limitations of HFile that later were problematic, specifically, we couldn't vary cell encoding on a block by block basis. Even if no cells use tags in a file, we'd bloat each cell with a short. Later we introduced a whole file optimization for this issue but clearly we'd have more opportunities to employ it if we could vary encoding strategy on a block-by-block basis. I thought about introducing an extensible pbufed block header. It didn't make sense for the tag serialization issue - we would trade one type of bloat for another, those additional header bytes will end up in the block cache - but if there are multiple use cases for it lined up, a new extensible pbufed 'block header' could be worthwhile. Would make future block level changes less likely to be incompatible changes too. \n\nDoes the introduction of something like that require a major version bump? I think so. I'd like to see us be more like semver with HFile versioning, and if we're on the same page about that, this is a major version bump because earlier versioned readers won't be able to handle the change. \n\nAlso opens the door to interesting things like using different block encoding strategies on a block by block basis according to characteristics of the cells to be encoded within.","from":"developer"},{"body":"We can fix this bug now (in all applicable versions) by reading the header and then data. This 2 reads have perf impact but better no bug. Can we get that in first? For the 2.0 version (branch-1 also?) we can decide on how to add the new data to header (with major version bump or minor bump).. Fixing this bug as such we can give priority and get it done soon?","from":"developer"},{"body":"Hey guys, sorry, I should be able to get back to this soon. Finishing up an unrelated project right now. I didn't know that minor versions in HFiles were also non-backwards compatible. That's one less reason then to make this a major version bump. If anyone has a strong preference for this fix to go into a V3.X I can change the patch to use minor version (eg for header size calculation) when I have time to do it. If not I'll leave it as V4 since it's a little simpler in the code as a major version bump. My original intention btw if it wasn't clear was that this wouldn't be the only change in a V4, just the first change that would go into a V4, whose format/contents is not yet meant to be final even when this patch is committed, i.e. V4 would be essentially a WIP with more changes suggested and implemented in other tickets and eventually released in HBase 2.0.\n\n[~anoop.hbase] I'm down for committing a short-term read-the-header-always fix for now and then discussing the longer term solution second. Which branches do you want the patch for?","from":"developer"},{"body":"I would say commit for all applicable branches including Trunk. The better fix wherever we need (trunk , branch-1) pls add a TODO in the patch with Jira# which will fix that.","from":"developer"},{"body":"Alright I'll work on that. I created HBASE-14576 for the longer term fix, so this ticket will just be to implement the short term fix we described.","from":"developer"},{"body":"Short term patches for this bug per discussion.","from":"developer"},{"body":"Hmmm, can't seem to edit Jira comments. Anyways, I attached short term patches to the ticket per discussion for all versions of HBase from 0.98 to master. The patches are mostly the same other than 0.98. There were a couple of utility methods that were missing/private in earlier versions of HBase 1.X that I backported or changed to public to mirror the master version of the patch. (Let me know if that isn't kosher for some reason.) I probably won't have time to update the patches this week but I'll look at feedback and make appropriate changes when I get the chance next week.","from":"developer"},{"body":"lgtm\n\nQA report picked up hadoop tests.","from":"developer"},{"body":"+1","from":"developer"},{"body":"So anything I should change in the patches? How many +1's are needed? Does someone else need to +1? ","from":"developer"},{"body":"Attach the same patch once more for a clean QA run.. Then we can commit it.. Ted can you pls commit it after QA? ","from":"developer"},{"body":"Am I understanding correctly that we always incur two reads now, even when we're scanning forward? If so, that seems unfortunate.","from":"developer"},{"body":"No it is for backward scan only. [~benlau] confirm once if am wrong.","from":"developer"},{"body":"[~lhofhansl] and [~anoop.hbase], I probably should’ve called it out explicitly but yes technically, it would affect forward scans (among other things), but I don’t think in any noticeable way. The patched code is called by HalfStoreFileReader.getLastKey(). \n\nThis getLastKey() method is needed to figure out the split point of a region which is eg used in splitting a region. Extra IO op there but given that region splits don’t happen frequently nor at a very high rate when they do I think it is fine. The getLastKey() is also used by StoreFile.Reader for passesKeyRangeFilter() check to determine whether a store is applicable to a scan, both forward/reverse. It affects the initial scan creation as well as later next() RPC calls I think. Although this is too bad, it doesn’t really matter I think in practice, because the block of the last key will get cached in the BlockCache the first time we need to know the last key for that halfstore. So repeated calls later in the region server will not incur any overhead. Let me know if this addresses your concerns or not.\n\nIncidentally I think that caching is one of the primary reasons why there seems to be a decent # of people using reverse scan but almost no one reporting this bug— because caching often hides it, either completely or nondeterministically (requests succeeding on later retries). We had some problems initially reproducing this bug reliably because if we ran forward scans in a table concurrently with the reverse scans, it would cause the blocks to become cached and so certain key ranges that previously would’ve caused reverse scan to fail suddenly started working just fine.\n\nRe: Attaching the same patch, let me know if I'm doing this wrong but it sounds like all I should have to do is just upload the same patch file for master but with a different name and the QA tests will pick it up. I will do that.","from":"developer"},{"body":"Attached a new patch for master, same as the previous patch but with 'reupload' in the name..","from":"developer"},{"body":"bq. Although this is too bad, it doesn’t really matter I think in practice, because the block of the last key will get cached in the BlockCache the first time we need to know the last key for that halfstore. So repeated calls later in the region server will not incur any overhead. Let me know if this addresses your concerns or not.\n\nWell... There are scanning workloads where the data being scanned won't fit into the blockcache and so the user would be calling Scan#setCacheBlocks(boolean) with 'false'.\n\nIf we are only doing twice as many reads as before per block only for HalfStoreFileReader, then I think this is ok, because HalfStoreFileReader is only used between the time when a split takes place and when compaction on the daughters is complete, references are removed, and the parent is deleted. Please add comments indicating where we are taking on suboptimal behavior. \n\nIf we will always issue two reads for a block in all cases, then that is really unfortunate and I'd ask that be revisited.","from":"developer"},{"body":"[~andrew.purtell@gmail.com] Let me know if I'm missing something, but I think there is more than 1 scan in play here. I think you're talking about an external hbase client scan. I'm talking about an internal hbase scan opened up by the regionserver and which we know for a fact caches the block. See the implementation of HalfStoreFileReader.getLastKey(), it is creating an internal scanner that does cache. Furthermore, the results of the method aren't cached in the Reader class (eg as a variable) and since the method is called repeatedly in the codebase it seems likely that the author's expectation was that the block cache would work correctly and make an internal cache for the file reader redundant. So not only does this scenario happen only for newly split regions but it only happens for the first time. I can add more comments to the patches if it is really necessary but there is already a comment in the code indicating that this fix is not performant and is meant to be updated by a later ticket whose jira # is listed.","from":"developer"},{"body":"[~andrew.purtell@gmail.com] Is the above explanation agreeable?","from":"developer"},{"body":"Still there [~andrew.purtell@gmail.com]? Anyone else have comments/questions?","from":"developer"},{"body":"Ok, agreeable.","from":"developer"},{"body":"Thanks Andrew. Can we merge these patches or is there still something that needs to be done or reviewed?","from":"developer"},{"body":"[~benlau] I don't see more outstanding feedback and there's already a +1 here from Anoop. Let me check on some things locally and then commit if it looks good.","from":"developer"},{"body":"Pushed to 0.98 and up.","from":"developer"},{"body":"That's odd, something went wrong with the patch for 1.1 branch. That compiled fine for me but perhaps I overlooked something, or the patch became stale. I will take a look later tonight.","from":"developer"},{"body":"Not a problem. This compiled for me too so I thought but there's a small missing static helper function in CellUtil. I'll commit it as an addendum now.","from":"developer"},{"body":"No worries. Just pushed this small addendum to branch-1.1. ","from":"developer"},{"body":"Sorry, I was trying to commit HBASE-14682, and pushed the addendum in the process without checking here. ","from":"developer"},{"body":"I'll back mine out because the build is broken again (two CellUtil#matchingTimestamps) \n","from":"developer"},{"body":"[~andrew.purtell@gmail.com] Was the wrong patch applied to 1.1 or am I misunderstanding something? So I looked at https://github.com/apache/hbase/commits/branch-1.1 specifically https://github.com/apache/hbase/commit/0db04a1705e5e8cc04cc9c010ddfc5612f60cfec and it is missing the CellUtil method that I had in my HBASE-14283-branch-1.1.patch attached. It's fine if this is fixed as an addendum but there's nothing else amiss right (i.e. the other branches have the right patches applied?)","from":"developer"},{"body":"No worries Ben. Not sure how I can help. The other branches got the right patches. ","from":"developer"},{"body":"Ok thanks just wanted to confirm the other branches were fine as far as we know, eg we didnt swap 1.1 patch with 0.98 patch and need to fix/update 0.98 branch. Thanks. It looks like there are some test failures but I think they are not related to the patch. I will rerun the tests that fail tonight locally after they appear in this ticket.","from":"developer"},{"body":"Thanks for the perseverance [~benlau].. Hope u will now work on the other Jira to handle it in better way wrt perf impact. Thanks...","from":"developer"},{"body":"Reran the failures that looked relevant on the various branches, seems the tests are just unstable. When I get the time I'll try to restart the discussion about updating the HFile serialization for more efficient reverse scans in HBASE-14576. ","from":"developer"},{"body":"Bulk closing 1.1.3 issues.","from":"developer"}],"created":"2015-08-21T17:30:55.000+0000","description":"Reverse scans do not work if an HFile contains inline bloom blocks or leaf level index blocks. The reason is because the seekBefore() call calculates the previous data block’s size by assuming data blocks are contiguous which is not the case in HFile V2 and beyond.\n\nAttached is a first cut patch (targeting bcef28eefaf192b0ad48c8011f98b8e944340da5 on trunk) which includes:\n(1) a unit test which exposes the bug and demonstrates failures for both inline bloom blocks and inline index blocks\n(2) a proposed fix for inline index blocks that does not require a new HFile version change, but is only performant for 1 and 2-level indexes and not 3+. 3+ requires an HFile format update for optimal performance. \n\nThis patch does not fix the bloom filter blocks bug. But the fix should be similar to the case of inline index blocks. The reason I haven’t made the change yet is I want to confirm that you guys would be fine with me revising the HFile.Reader interface.\n\nSpecifically, these 2 functions (getGeneralBloomFilterMetadata and getDeleteBloomFilterMetadata) need to return the BloomFilter. Right now the HFileReader class doesn’t have a reference to the bloom filters (and hence their indices) and only constructs the IO streams and hence has no way to know where the bloom blocks are in the HFile. It seems that the HFile.Reader bloom method comments state that they “know nothing about how that metadata is structured” but I do not know if that is a requirement of the abstraction (why?) or just an incidental current property. \n\nWe would like to do 3 things with community approval:\n(1) Update the HFile.Reader interface and implementation to contain and return BloomFilters directly rather than unstructured IO streams\n(2) Merge the fixes for index blocks and bloom blocks into open source\n(3) Create a new Jira ticket for open source HBase to add a ‘prevBlockSize’ field in the block header in the next HFile version, so that seekBefore() calls can not only be correct but performant in all cases.","issue_id":"12857913","key":"HBASE-14283","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2015-10-26T23:45:17.000+0000","role":"fixed_distractor","summary":"Reverse scan doesn’t work with HFile inline index/bloom blocks"} {"case_id":"12863285","cluster":"DISTRACTOR-HBASE-14406","comments":[{"body":"Following condition will result in the same filter. It will have data loss with the current filter construction.\ncol1 > 4 && col2 < 3\ncol1 > 4 || col2 < 3","created":"2015-09-11T05:05:12.592+0000"},{"body":"Ted Malaska\n[~malaskat]'s response\nOn one you 100% right. We're good if we have one RowKey filter and one Column Filter or multiple column filters that are and joined. But an or on two different columns will cause rows to be wrongly filtered.\n\nIt is quite common in sql query there are filters on multiple columns.","created":"2015-09-11T05:09:01.315+0000"},{"body":"[~zhazhan]:\nI copied your 1st comment to description.\nIt is good practice to give people good summary.","created":"2015-09-11T05:10:31.497+0000"},{"body":"Thank again for finding this. Sorry for missing this in the code.\n\nLet me know if you want to make the patch, or if someone else would like to make the patch. If not I have time early next week.\n\nThanks again.","created":"2015-09-11T12:11:11.083+0000"},{"body":"Let me know if anyone would like to do this. If not I plan to take it on monday.\n\nI think I have a solution in my head. Should take a day or so to code up.","created":"2015-09-13T00:37:01.298+0000"},{"body":"Cool going to started tonight.\n\nThanks [~aafri] ","created":"2015-09-14T18:36:58.361+0000"},{"body":"Made really good progress today. There is about 4 more hours of work left.\n\nSo hopefully in the next day or so.","created":"2015-09-15T20:59:12.700+0000"},{"body":"First draft. \n\n1. Added full support for dynamic logic to be pushed down in the scan's filter\n2. Added the ability to turn on or off the push down filter functionality\n3. Added a bunch of tests\n\n","created":"2015-09-20T17:29:51.570+0000"},{"body":"Sorry for the delay. It was a long week.\n\nThis patch has all of Ted Yu's connects addressed as best that I could. Please review thanks.\n\nAnd thank you Ted Yu for your Review.","created":"2015-09-25T23:42:33.469+0000"},{"body":"Clarified things and adding a more asserts to test the RowKeyFilter outcomes for all queries that use it.","created":"2015-09-27T02:09:56.064+0000"},{"body":"@Ted:\nAre you adding more tests ?\n\nI will go over the review board one more time.","created":"2015-09-30T14:38:05.370+0000"},{"body":"Sorry Hadoop World took me out of commition for awhile, but I'm back now.\n\nI looked at the review board. I don't see anything else to add. What else has to happen to close this jira out. I have the whole week to focus on this jira.","created":"2015-10-11T02:32:27.299+0000"},{"body":"@Ted:\nI will review one more time.\nHave your tests covered the following clause ?\n\nwhere rowkey < 1 and col > 2\n\nwhere rowkey < 1 or col > 3","created":"2015-10-12T16:43:52.808+0000"},{"body":"I think the bug from last time was the following two\n\n( rowkey < 1 or col > 2 )\n\nand\n\n( colA < 1 or colB > 2 )\n\nThe functionality of (rowkey < 1 and col > 2) worked in the last patch\n\nBut here are some related tests that should cover both cases\ntest(\"Test SQL point and range combo\") \ntest(\"Test OR logic with a one RowKey and One column\")\ntest(\"Test two complete range non merge rowKey query\")\ntest(\"Test OR logic with a two columns\")\n","created":"2015-10-12T16:58:14.586+0000"},{"body":"[~malaskat]\n\nCan you add a test case as \n val results = sqlContext.sql(\"SELECT KEY_FIELD, B_FIELD, A_FIELD FROM hbaseTable1 \" +\n \"WHERE \" +\n \"( KEY_FIELD >= 'get4' or A_FIELD <= 'foo2') ).take(10)","created":"2015-10-12T17:31:23.058+0000"},{"body":"[~zhanzhang] np. Lets me add it now. It will take hopefully less then an hour.","created":"2015-10-12T17:44:57.855+0000"},{"body":"Applied worked for Zhan Zhang and Ted Yu","created":"2015-10-12T19:02:56.109+0000"},{"body":"Then to Than\n","created":"2015-10-12T21:01:31.820+0000"},{"body":"Running test against patch v4:\n{code}\nDiscovery starting.\nDiscovery completed in 1 second, 98 milliseconds.\nRun starting. Expected test count is: 43\nHBaseDStreamFunctionsSuite:\nFormatting using clusterid: testClusterID\n- bulkput to test HBase client *** FAILED ***\n java.lang.NullPointerException:\n at org.apache.hadoop.hbase.CellUtil.cloneValue(CellUtil.java:98)\n at org.apache.hadoop.hbase.spark.HBaseDStreamFunctionsSuite$$anonfun$1.apply$mcV$sp(HBaseDStreamFunctionsSuite.scala:116)\n at org.apache.hadoop.hbase.spark.HBaseDStreamFunctionsSuite$$anonfun$1.apply(HBaseDStreamFunctionsSuite.scala:62)\n at org.apache.hadoop.hbase.spark.HBaseDStreamFunctionsSuite$$anonfun$1.apply(HBaseDStreamFunctionsSuite.scala:62)\n at org.scalatest.Transformer$$anonfun$apply$1.apply$mcV$sp(Transformer.scala:22)\n at org.scalatest.OutcomeOf$class.outcomeOf(OutcomeOf.scala:85)\n at org.scalatest.OutcomeOf$.outcomeOf(OutcomeOf.scala:104)\n at org.scalatest.Transformer.apply(Transformer.scala:22)\n at org.scalatest.Transformer.apply(Transformer.scala:20)\n at org.scalatest.FunSuiteLike$$anon$1.apply(FunSuiteLike.scala:166)\n{code}","created":"2015-10-12T21:06:18.095+0000"},{"body":"Hmm. Most likely a race condition. Let me look into it. I didn't change that code and all the tests work on my side.\n\nLet me see what I can do.","created":"2015-10-12T21:14:31.821+0000"},{"body":"Yup race condition changing awaitTerminationOrTimeout to awaitTermination patch coming once everything passes tests and build.","created":"2015-10-12T21:18:02.217+0000"},{"body":"Hmm let me research how it is done in the unit tests in Spark core and I will try to do the same.","created":"2015-10-12T21:28:57.731+0000"},{"body":"Fixed DStream race condition","created":"2015-10-12T21:53:03.479+0000"},{"body":"Based on patch v6:\n{code}\nHBaseContextSuite:\n*** RUN ABORTED ***\n java.io.IOException: Unable to determine a plan to assign region(s)\n at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)\n at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n at java.lang.reflect.Constructor.newInstance(Constructor.java:526)\n at org.apache.hadoop.ipc.RemoteException.instantiateException(RemoteException.java:106)\n at org.apache.hadoop.ipc.RemoteException.unwrapRemoteException(RemoteException.java:95)\n at org.apache.hadoop.hbase.util.ForeignExceptionUtil.toIOException(ForeignExceptionUtil.java:45)\n at org.apache.hadoop.hbase.client.HBaseAdmin$ProcedureFuture.convertResult(HBaseAdmin.java:4513)\n at org.apache.hadoop.hbase.client.HBaseAdmin$ProcedureFuture.waitProcedureResult(HBaseAdmin.java:4471)\n at org.apache.hadoop.hbase.client.HBaseAdmin$ProcedureFuture.get(HBaseAdmin.java:4405)\n ...\n Cause: org.apache.hadoop.ipc.RemoteException: Unable to determine a plan to assign region(s)\n at org.apache.hadoop.hbase.master.AssignmentManager.assign(AssignmentManager.java:1497)\n at org.apache.hadoop.hbase.util.ModifyRegionUtils.assignRegions(ModifyRegionUtils.java:254)\n at org.apache.hadoop.hbase.master.procedure.CreateTableProcedure.assignRegions(CreateTableProcedure.java:430)\n at org.apache.hadoop.hbase.master.procedure.CreateTableProcedure.executeFromState(CreateTableProcedure.java:126)\n at org.apache.hadoop.hbase.master.procedure.CreateTableProcedure.executeFromState(CreateTableProcedure.java:57)\n at org.apache.hadoop.hbase.procedure2.StateMachineProcedure.execute(StateMachineProcedure.java:119)\n at org.apache.hadoop.hbase.procedure2.Procedure.doExecute(Procedure.java:442)\n at org.apache.hadoop.hbase.procedure2.ProcedureExecutor.execProcedure(ProcedureExecutor.java:1002)\n at org.apache.hadoop.hbase.procedure2.ProcedureExecutor.execLoop(ProcedureExecutor.java:793)\n at org.apache.hadoop.hbase.procedure2.ProcedureExecutor.execLoop(ProcedureExecutor.java:746)\n{code}","created":"2015-10-12T22:08:31.395+0000"},{"body":"That last issue looks to be outside of the Spark code. Looks to be something wrong with the HBaseTestingUtility in the build environment. I've ran these test a dozen times today through the following command and have not seen this issue.\n\ncd hbase-spark\nmvn -Dtest=NoUnitTests clean verify\n\nWhat are the options to resolve this issue?","created":"2015-10-12T22:22:24.180+0000"},{"body":"I agree the cause might be in hbase core.\n\nHowever under hbase-spark/target/surefire-reports/ , I only found:\n{code}\nls ~/trunk/hbase-spark/target/surefire-reports/\nTEST-org.apache.hadoop.hbase.spark.DynamicLogicExpressionSuite.xml\tTestSuite.txt\nTEST-org.apache.hadoop.hbase.spark.HBaseDStreamFunctionsSuite.xml\n{code}\nDo you know where I can find the log corresponding to the period when HBaseContextSuite ran ?","created":"2015-10-12T22:26:58.374+0000"},{"body":"HBaseTestingUtility is used in all the following test suites:\n\n* BulkLoadSuite\n* DefaultSourceSuite\n* HBaseContextSuite\n* HBaseDStreamFunctionsSuite\n* HBaseRDDFunctionsSuite\n* TestJavaHBaseContext\n\nAs for the logs When I run \"mvn -Dtest=NoUnitTests clean verify\" I get the following files in my hbase-spark/surefire-reports/ folder.\n\nExample I see \"TEST-org.apache.hadoop.hbase.spark.HBaseContextSuite.xml\"\n\nI will upload a compressed bundle of what I see in that director.\n\n\n\n\n","created":"2015-10-12T22:39:12.331+0000"},{"body":"Attachment example of what I see after I run mvn -Dtest=NoUnitTests clean verify\n","created":"2015-10-12T22:40:01.877+0000"},{"body":"I was able to find target/surefire-reports/TEST-org.apache.hadoop.hbase.spark.HBaseContextSuite.xml after a good run.\nHowever, it didn't contain any log related to hbase core.","created":"2015-10-12T22:46:31.353+0000"},{"body":"Whats going on with the build, it doesn't look like it even got to the Spark Tests.\n\nIs there anything I need to change?\n\nLet me know I'm open to working on this all tomorrow if need be.","created":"2015-10-13T00:20:03.792+0000"},{"body":"Yeah it passed. :) \n\nSo [~ted_yu] and [~zhanzhang] let me know what else I need to do and thank you super much for being so engaged today and helping me through this jira.\n\nTed Malaska","created":"2015-10-13T02:18:52.408+0000"},{"body":"Looking at hbase-protocol/src/main/protobuf/Filter.proto , SQLPredicatePushDownCellToColumnMapping and SQLPredicatePushDownFilter are used by hbase-spark module only.\nIs it possible to separate these two from Filter.proto and put into .proto file in hbase-spark module ?","created":"2015-10-13T22:42:26.519+0000"},{"body":"Well that does make sense. Let me look into that tomorrow.","created":"2015-10-13T22:45:43.130+0000"},{"body":"Moved ProtoBufs to hbase-spark and out of hbase-protoco","created":"2015-10-14T12:51:05.811+0000"},{"body":"What went wrong with the build?","created":"2015-10-14T18:06:34.674+0000"},{"body":"{code}\npatching file hbase-spark/pom.xml\nHunk #2 FAILED at 595.\n1 out of 2 hunks FAILED -- saving rejects to file hbase-spark/pom.xml.rej\n{code}\nHere is the tip of pom.xml.rej :\n{code}\n***************\n*** 587,592 ****\n \n \n \n- \n\n --- 595,652 ----\n \n \n \n\n+ \n+ \n+ \n+ \n+ skip-rpc-tests\n+ \n+ \n+ skip-rpc-tests\n+ \n+ \n+ \n+ true\n+ \n+ \n+ \n{code}","created":"2015-10-14T18:27:05.478+0000"},{"body":"OK I will make this change in the next hour or so.\n\nThanks Ted Yu","created":"2015-10-15T18:03:26.519+0000"},{"body":"rebasing pom.\n\nOn a side note something funky happen to hbase when I rebased. On my computer all the unit tests are broken in the master branch unless I do a \"reverse look-up-able IP\" to make it work. This issue is unrelated to my patch it was something else that change recently.\n\nI get this error \njava.io.IOException: java.lang.RuntimeException: Could not resolve Kerberos principal name: java.net.UnknownHostException: tmalaska-MBP-2.home: tmalaska-MBP-2.home: nodename nor servname provided, or not known\n\nOn a HBaseTestingUtility startMiniCluster\n\nI have tested this on more then one computer with friends. \n\nSo repeat the patch should be good, but something is not good with HBase in the latest master with respeck to HBaseTestingUtility startMiniCluster\n\n","created":"2015-10-15T21:03:06.742+0000"},{"body":"[~tedyu] just brought this to my attention WRT HBASE-14594 which was recently applied.\n\nCan you provide a command to run the tests that you see fail, [~malaskat]?\n\nbq. java.io.IOException: java.lang.RuntimeException: Could not resolve Kerberos principal name: java.net.UnknownHostException: tmalaska-MBP-2.home: tmalaska-MBP-2.home: nodename nor servname provided, or not known\n\nGiven that DNS is a central component to the kerberos security model, I'm surprise to hear you say that this ever worked for you before...","created":"2015-10-15T21:18:08.467+0000"},{"body":"I just did \"mvn -Dtest=NoUnitTests clean verify\" in the hbase-spark folder\n\nIt totally worked before I rebased then I rebased and it didn't work.\n\nI also got a fresh copy of master (so with out this patch) and I tried it on my box and two other people's boxes and all three failed.\n\nThe change in the host file was successful on one of the boxes to fix the problem. I will try to repeat the fix when I get home.","created":"2015-10-15T21:21:46.580+0000"},{"body":"bq. just brought this to my attention WRT HBASE-14594 which was recently applied.\n\nOh, it's also not readily apparent to me how the changes from HBASE-14594 would have broken anything as it just added a layer of indirection in what got invoked.","created":"2015-10-15T21:21:52.107+0000"},{"body":"Also this is my version\n\ntmalaska-MBP-2:hbase-spark ted.malaska$ git log | head -n 1\ncommit d5ed46bc9f9285f75d2d906ec9c120cb408827df","created":"2015-10-15T21:22:19.131+0000"},{"body":"The code that break is just\n\n var TEST_UTIL: HBaseTestingUtility = new HBaseTestingUtility\n TEST_UTIL.startMiniCluster() //BOOM\n\nThere is nothing else that runs no Spark stuff no nothing. Just HBaseTestingUtility","created":"2015-10-15T21:23:41.460+0000"},{"body":"BTW here is the full stack trace when the host file is not updated to do a reverse look up.\n\n","created":"2015-10-15T21:59:19.493+0000"},{"body":"I tried this build and it doesn't have the hbase unit testing problem\n\ntmalaska-MBP-2:hbase-spark ted.malaska$ git log | head -n 1\ncommit 8f95318f6252c1c0b7a073619525eae6d991f47b","created":"2015-10-15T22:39:48.470+0000"},{"body":"OK I rebuilt everything and restarted my computer. And everything is fine.\n\nI'm not sure what caused the problem originally but unit tests in patch 9 work on my local.\n\nSorry for the false alarm\n","created":"2015-10-15T22:54:51.480+0000"},{"body":"Just double checking and uploading the newest version","created":"2015-10-15T22:57:05.911+0000"},{"body":"There should be FilterProtos.java in hbase-spark module.\n\nDid you forget to include it in the patch ?","created":"2015-10-15T23:13:46.484+0000"},{"body":"yup it is in there\n\nhttps://reviews.apache.org/r/38536/diff/9#4\n\nLet me know if you don't see it\n","created":"2015-10-15T23:22:58.614+0000"},{"body":"I was talking about patch v10 - I don't see this file there.","created":"2015-10-15T23:24:04.156+0000"},{"body":"I just looked at \n\nhttps://issues.apache.org/jira/secure/attachment/12766912/HBASE-14406.10.patch\n\nand search for \n\ndiff --git a/hbase-spark/src/main/protobuf/Filter.proto b/hbase-spark/src/main/protobuf/Filter.proto\n\nIt's there. Let me know if I missed something","created":"2015-10-15T23:26:41.649+0000"},{"body":"Also the diff number of the review board if off by one. Which is my fault I skipped version 8. It never made it to up loaded :)\n\nSo version 9 on reviewBoard is version 10 on jira","created":"2015-10-15T23:28:37.395+0000"},{"body":"grr the build system didn't generate the proto classes\n\nLet me do some research","created":"2015-10-15T23:42:30.867+0000"},{"body":"Ohh [~ted_yu] so I need to add the generated file into the patch. Now I understand what you are saying.\n\nSorry I was reading to fast.\n\nWill make new patch now","created":"2015-10-15T23:47:59.664+0000"},{"body":"Added FilterProtos.java to git","created":"2015-10-15T23:55:15.003+0000"},{"body":"Thanks Zhan for reviewing.\n\nThanks Ted for the patch.","created":"2015-10-16T18:26:13.886+0000"},{"body":"OMG I own you guys a beer. That was a long patch. Thank you both.","created":"2015-10-16T18:27:53.922+0000"},{"body":"own -> owe","created":"2015-10-16T18:29:01.893+0000"},{"body":"Reverted from branch-2/2.0.0 by HBASE-18817","created":"2018-04-05T23:37:29.827+0000"},{"body":"hbase-spark module reverted from branch-2/2.0.0 by HBASE-18817","created":"2018-04-06T04:05:29.326+0000"}],"conversations":[{"body":"Following condition will result in the same filter. It will have data loss with the current filter construction.\ncol1 > 4 && col2 < 3\ncol1 > 4 || col2 < 3","from":"reporter","subject":"The dataframe datasource filter is wrong, and will result in data loss or unexpected behavior"},{"body":"Following condition will result in the same filter. It will have data loss with the current filter construction.\ncol1 > 4 && col2 < 3\ncol1 > 4 || col2 < 3","from":"developer"},{"body":"Ted Malaska\n[~malaskat]'s response\nOn one you 100% right. We're good if we have one RowKey filter and one Column Filter or multiple column filters that are and joined. But an or on two different columns will cause rows to be wrongly filtered.\n\nIt is quite common in sql query there are filters on multiple columns.","from":"developer"},{"body":"[~zhazhan]:\nI copied your 1st comment to description.\nIt is good practice to give people good summary.","from":"developer"},{"body":"Thank again for finding this. Sorry for missing this in the code.\n\nLet me know if you want to make the patch, or if someone else would like to make the patch. If not I have time early next week.\n\nThanks again.","from":"developer"},{"body":"Let me know if anyone would like to do this. If not I plan to take it on monday.\n\nI think I have a solution in my head. Should take a day or so to code up.","from":"developer"},{"body":"Cool going to started tonight.\n\nThanks [~aafri] ","from":"developer"},{"body":"Made really good progress today. There is about 4 more hours of work left.\n\nSo hopefully in the next day or so.","from":"developer"},{"body":"First draft. \n\n1. Added full support for dynamic logic to be pushed down in the scan's filter\n2. Added the ability to turn on or off the push down filter functionality\n3. Added a bunch of tests\n\n","from":"developer"},{"body":"Sorry for the delay. It was a long week.\n\nThis patch has all of Ted Yu's connects addressed as best that I could. Please review thanks.\n\nAnd thank you Ted Yu for your Review.","from":"developer"},{"body":"Clarified things and adding a more asserts to test the RowKeyFilter outcomes for all queries that use it.","from":"developer"},{"body":"@Ted:\nAre you adding more tests ?\n\nI will go over the review board one more time.","from":"developer"},{"body":"Sorry Hadoop World took me out of commition for awhile, but I'm back now.\n\nI looked at the review board. I don't see anything else to add. What else has to happen to close this jira out. I have the whole week to focus on this jira.","from":"developer"},{"body":"@Ted:\nI will review one more time.\nHave your tests covered the following clause ?\n\nwhere rowkey < 1 and col > 2\n\nwhere rowkey < 1 or col > 3","from":"developer"},{"body":"I think the bug from last time was the following two\n\n( rowkey < 1 or col > 2 )\n\nand\n\n( colA < 1 or colB > 2 )\n\nThe functionality of (rowkey < 1 and col > 2) worked in the last patch\n\nBut here are some related tests that should cover both cases\ntest(\"Test SQL point and range combo\") \ntest(\"Test OR logic with a one RowKey and One column\")\ntest(\"Test two complete range non merge rowKey query\")\ntest(\"Test OR logic with a two columns\")\n","from":"developer"},{"body":"[~malaskat]\n\nCan you add a test case as \n val results = sqlContext.sql(\"SELECT KEY_FIELD, B_FIELD, A_FIELD FROM hbaseTable1 \" +\n \"WHERE \" +\n \"( KEY_FIELD >= 'get4' or A_FIELD <= 'foo2') ).take(10)","from":"developer"},{"body":"[~zhanzhang] np. Lets me add it now. It will take hopefully less then an hour.","from":"developer"},{"body":"Applied worked for Zhan Zhang and Ted Yu","from":"developer"},{"body":"Then to Than\n","from":"developer"},{"body":"Running test against patch v4:\n{code}\nDiscovery starting.\nDiscovery completed in 1 second, 98 milliseconds.\nRun starting. Expected test count is: 43\nHBaseDStreamFunctionsSuite:\nFormatting using clusterid: testClusterID\n- bulkput to test HBase client *** FAILED ***\n java.lang.NullPointerException:\n at org.apache.hadoop.hbase.CellUtil.cloneValue(CellUtil.java:98)\n at org.apache.hadoop.hbase.spark.HBaseDStreamFunctionsSuite$$anonfun$1.apply$mcV$sp(HBaseDStreamFunctionsSuite.scala:116)\n at org.apache.hadoop.hbase.spark.HBaseDStreamFunctionsSuite$$anonfun$1.apply(HBaseDStreamFunctionsSuite.scala:62)\n at org.apache.hadoop.hbase.spark.HBaseDStreamFunctionsSuite$$anonfun$1.apply(HBaseDStreamFunctionsSuite.scala:62)\n at org.scalatest.Transformer$$anonfun$apply$1.apply$mcV$sp(Transformer.scala:22)\n at org.scalatest.OutcomeOf$class.outcomeOf(OutcomeOf.scala:85)\n at org.scalatest.OutcomeOf$.outcomeOf(OutcomeOf.scala:104)\n at org.scalatest.Transformer.apply(Transformer.scala:22)\n at org.scalatest.Transformer.apply(Transformer.scala:20)\n at org.scalatest.FunSuiteLike$$anon$1.apply(FunSuiteLike.scala:166)\n{code}","from":"developer"},{"body":"Hmm. Most likely a race condition. Let me look into it. I didn't change that code and all the tests work on my side.\n\nLet me see what I can do.","from":"developer"},{"body":"Yup race condition changing awaitTerminationOrTimeout to awaitTermination patch coming once everything passes tests and build.","from":"developer"},{"body":"Hmm let me research how it is done in the unit tests in Spark core and I will try to do the same.","from":"developer"},{"body":"Fixed DStream race condition","from":"developer"},{"body":"Based on patch v6:\n{code}\nHBaseContextSuite:\n*** RUN ABORTED ***\n java.io.IOException: Unable to determine a plan to assign region(s)\n at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)\n at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n at java.lang.reflect.Constructor.newInstance(Constructor.java:526)\n at org.apache.hadoop.ipc.RemoteException.instantiateException(RemoteException.java:106)\n at org.apache.hadoop.ipc.RemoteException.unwrapRemoteException(RemoteException.java:95)\n at org.apache.hadoop.hbase.util.ForeignExceptionUtil.toIOException(ForeignExceptionUtil.java:45)\n at org.apache.hadoop.hbase.client.HBaseAdmin$ProcedureFuture.convertResult(HBaseAdmin.java:4513)\n at org.apache.hadoop.hbase.client.HBaseAdmin$ProcedureFuture.waitProcedureResult(HBaseAdmin.java:4471)\n at org.apache.hadoop.hbase.client.HBaseAdmin$ProcedureFuture.get(HBaseAdmin.java:4405)\n ...\n Cause: org.apache.hadoop.ipc.RemoteException: Unable to determine a plan to assign region(s)\n at org.apache.hadoop.hbase.master.AssignmentManager.assign(AssignmentManager.java:1497)\n at org.apache.hadoop.hbase.util.ModifyRegionUtils.assignRegions(ModifyRegionUtils.java:254)\n at org.apache.hadoop.hbase.master.procedure.CreateTableProcedure.assignRegions(CreateTableProcedure.java:430)\n at org.apache.hadoop.hbase.master.procedure.CreateTableProcedure.executeFromState(CreateTableProcedure.java:126)\n at org.apache.hadoop.hbase.master.procedure.CreateTableProcedure.executeFromState(CreateTableProcedure.java:57)\n at org.apache.hadoop.hbase.procedure2.StateMachineProcedure.execute(StateMachineProcedure.java:119)\n at org.apache.hadoop.hbase.procedure2.Procedure.doExecute(Procedure.java:442)\n at org.apache.hadoop.hbase.procedure2.ProcedureExecutor.execProcedure(ProcedureExecutor.java:1002)\n at org.apache.hadoop.hbase.procedure2.ProcedureExecutor.execLoop(ProcedureExecutor.java:793)\n at org.apache.hadoop.hbase.procedure2.ProcedureExecutor.execLoop(ProcedureExecutor.java:746)\n{code}","from":"developer"},{"body":"That last issue looks to be outside of the Spark code. Looks to be something wrong with the HBaseTestingUtility in the build environment. I've ran these test a dozen times today through the following command and have not seen this issue.\n\ncd hbase-spark\nmvn -Dtest=NoUnitTests clean verify\n\nWhat are the options to resolve this issue?","from":"developer"},{"body":"I agree the cause might be in hbase core.\n\nHowever under hbase-spark/target/surefire-reports/ , I only found:\n{code}\nls ~/trunk/hbase-spark/target/surefire-reports/\nTEST-org.apache.hadoop.hbase.spark.DynamicLogicExpressionSuite.xml\tTestSuite.txt\nTEST-org.apache.hadoop.hbase.spark.HBaseDStreamFunctionsSuite.xml\n{code}\nDo you know where I can find the log corresponding to the period when HBaseContextSuite ran ?","from":"developer"},{"body":"HBaseTestingUtility is used in all the following test suites:\n\n* BulkLoadSuite\n* DefaultSourceSuite\n* HBaseContextSuite\n* HBaseDStreamFunctionsSuite\n* HBaseRDDFunctionsSuite\n* TestJavaHBaseContext\n\nAs for the logs When I run \"mvn -Dtest=NoUnitTests clean verify\" I get the following files in my hbase-spark/surefire-reports/ folder.\n\nExample I see \"TEST-org.apache.hadoop.hbase.spark.HBaseContextSuite.xml\"\n\nI will upload a compressed bundle of what I see in that director.\n\n\n\n\n","from":"developer"},{"body":"Attachment example of what I see after I run mvn -Dtest=NoUnitTests clean verify\n","from":"developer"},{"body":"I was able to find target/surefire-reports/TEST-org.apache.hadoop.hbase.spark.HBaseContextSuite.xml after a good run.\nHowever, it didn't contain any log related to hbase core.","from":"developer"},{"body":"Whats going on with the build, it doesn't look like it even got to the Spark Tests.\n\nIs there anything I need to change?\n\nLet me know I'm open to working on this all tomorrow if need be.","from":"developer"},{"body":"Yeah it passed. :) \n\nSo [~ted_yu] and [~zhanzhang] let me know what else I need to do and thank you super much for being so engaged today and helping me through this jira.\n\nTed Malaska","from":"developer"},{"body":"Looking at hbase-protocol/src/main/protobuf/Filter.proto , SQLPredicatePushDownCellToColumnMapping and SQLPredicatePushDownFilter are used by hbase-spark module only.\nIs it possible to separate these two from Filter.proto and put into .proto file in hbase-spark module ?","from":"developer"},{"body":"Well that does make sense. Let me look into that tomorrow.","from":"developer"},{"body":"Moved ProtoBufs to hbase-spark and out of hbase-protoco","from":"developer"},{"body":"What went wrong with the build?","from":"developer"},{"body":"{code}\npatching file hbase-spark/pom.xml\nHunk #2 FAILED at 595.\n1 out of 2 hunks FAILED -- saving rejects to file hbase-spark/pom.xml.rej\n{code}\nHere is the tip of pom.xml.rej :\n{code}\n***************\n*** 587,592 ****\n \n \n \n- \n\n --- 595,652 ----\n \n \n \n\n+ \n+ \n+ \n+ \n+ skip-rpc-tests\n+ \n+ \n+ skip-rpc-tests\n+ \n+ \n+ \n+ true\n+ \n+ \n+ \n{code}","from":"developer"},{"body":"OK I will make this change in the next hour or so.\n\nThanks Ted Yu","from":"developer"},{"body":"rebasing pom.\n\nOn a side note something funky happen to hbase when I rebased. On my computer all the unit tests are broken in the master branch unless I do a \"reverse look-up-able IP\" to make it work. This issue is unrelated to my patch it was something else that change recently.\n\nI get this error \njava.io.IOException: java.lang.RuntimeException: Could not resolve Kerberos principal name: java.net.UnknownHostException: tmalaska-MBP-2.home: tmalaska-MBP-2.home: nodename nor servname provided, or not known\n\nOn a HBaseTestingUtility startMiniCluster\n\nI have tested this on more then one computer with friends. \n\nSo repeat the patch should be good, but something is not good with HBase in the latest master with respeck to HBaseTestingUtility startMiniCluster\n\n","from":"developer"},{"body":"[~tedyu] just brought this to my attention WRT HBASE-14594 which was recently applied.\n\nCan you provide a command to run the tests that you see fail, [~malaskat]?\n\nbq. java.io.IOException: java.lang.RuntimeException: Could not resolve Kerberos principal name: java.net.UnknownHostException: tmalaska-MBP-2.home: tmalaska-MBP-2.home: nodename nor servname provided, or not known\n\nGiven that DNS is a central component to the kerberos security model, I'm surprise to hear you say that this ever worked for you before...","from":"developer"},{"body":"I just did \"mvn -Dtest=NoUnitTests clean verify\" in the hbase-spark folder\n\nIt totally worked before I rebased then I rebased and it didn't work.\n\nI also got a fresh copy of master (so with out this patch) and I tried it on my box and two other people's boxes and all three failed.\n\nThe change in the host file was successful on one of the boxes to fix the problem. I will try to repeat the fix when I get home.","from":"developer"},{"body":"bq. just brought this to my attention WRT HBASE-14594 which was recently applied.\n\nOh, it's also not readily apparent to me how the changes from HBASE-14594 would have broken anything as it just added a layer of indirection in what got invoked.","from":"developer"},{"body":"Also this is my version\n\ntmalaska-MBP-2:hbase-spark ted.malaska$ git log | head -n 1\ncommit d5ed46bc9f9285f75d2d906ec9c120cb408827df","from":"developer"},{"body":"The code that break is just\n\n var TEST_UTIL: HBaseTestingUtility = new HBaseTestingUtility\n TEST_UTIL.startMiniCluster() //BOOM\n\nThere is nothing else that runs no Spark stuff no nothing. Just HBaseTestingUtility","from":"developer"},{"body":"BTW here is the full stack trace when the host file is not updated to do a reverse look up.\n\n","from":"developer"},{"body":"I tried this build and it doesn't have the hbase unit testing problem\n\ntmalaska-MBP-2:hbase-spark ted.malaska$ git log | head -n 1\ncommit 8f95318f6252c1c0b7a073619525eae6d991f47b","from":"developer"},{"body":"OK I rebuilt everything and restarted my computer. And everything is fine.\n\nI'm not sure what caused the problem originally but unit tests in patch 9 work on my local.\n\nSorry for the false alarm\n","from":"developer"},{"body":"Just double checking and uploading the newest version","from":"developer"},{"body":"There should be FilterProtos.java in hbase-spark module.\n\nDid you forget to include it in the patch ?","from":"developer"},{"body":"yup it is in there\n\nhttps://reviews.apache.org/r/38536/diff/9#4\n\nLet me know if you don't see it\n","from":"developer"},{"body":"I was talking about patch v10 - I don't see this file there.","from":"developer"},{"body":"I just looked at \n\nhttps://issues.apache.org/jira/secure/attachment/12766912/HBASE-14406.10.patch\n\nand search for \n\ndiff --git a/hbase-spark/src/main/protobuf/Filter.proto b/hbase-spark/src/main/protobuf/Filter.proto\n\nIt's there. Let me know if I missed something","from":"developer"},{"body":"Also the diff number of the review board if off by one. Which is my fault I skipped version 8. It never made it to up loaded :)\n\nSo version 9 on reviewBoard is version 10 on jira","from":"developer"},{"body":"grr the build system didn't generate the proto classes\n\nLet me do some research","from":"developer"},{"body":"Ohh [~ted_yu] so I need to add the generated file into the patch. Now I understand what you are saying.\n\nSorry I was reading to fast.\n\nWill make new patch now","from":"developer"},{"body":"Added FilterProtos.java to git","from":"developer"},{"body":"Thanks Zhan for reviewing.\n\nThanks Ted for the patch.","from":"developer"},{"body":"OMG I own you guys a beer. That was a long patch. Thank you both.","from":"developer"},{"body":"own -> owe","from":"developer"},{"body":"Reverted from branch-2/2.0.0 by HBASE-18817","from":"developer"},{"body":"hbase-spark module reverted from branch-2/2.0.0 by HBASE-18817","from":"developer"}],"created":"2015-09-11T04:59:55.000+0000","description":"Following condition will result in the same filter. It will have data loss with the current filter construction.\ncol1 > 4 && col2 < 3\ncol1 > 4 || col2 < 3","issue_id":"12863285","key":"HBASE-14406","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2015-10-16T18:26:13.000+0000","role":"fixed_distractor","summary":"The dataframe datasource filter is wrong, and will result in data loss or unexpected behavior"} {"case_id":"12427237","cluster":"DISTRACTOR-HBASE-1485","comments":[{"body":"Just to say Gets and Scanners should work the same, whatever we decide.","created":"2009-06-05T18:33:12.705+0000"},{"body":"I've had at least three people with a use case for this.\n\nMight create a couple sub-tasks here so we can at least head in the right direction.\n\nFirst, we need to make scanners ignore duplicate versions of the same column. The trickiest part is, how do we determine which to keep? We want to always come from the latest storefile, but I believe their IDs are still random and not timestamps? We might need to make that change to fix this. Would also then require a modification to the KVHeap to take this into account, all other things considered equal.\n\nOnce we have scanners working, that will mean the proper thing is enforced on major (and if we want, minor) compactions.\n\nGets will only work once we re-implement Gets as an optimized scan (taking advantage of bloom filters, mostly).\n\n\nI remember why I punted this to 0.20.1, the tricky part at the beginning is pretty tough and touches a good bit of core read-path code.\n\nRevisiting now, we'll see. Anyone else interested in this / want to work on it?","created":"2009-07-09T18:40:29.655+0000"},{"body":"Yeah, as ryan suggested, we should exploit sequenceid -- maybe name file for sequenceid?","created":"2009-07-10T05:29:11.611+0000"},{"body":"Bumped to 0.21","created":"2009-08-12T22:50:50.927+0000"},{"body":"Related to hbase-997","created":"2009-08-21T22:42:44.950+0000"},{"body":"From the list:\n\n{code}\nOn Tue, Jan 26, 2010 at 9:36 AM, Rod Cope wrote:\n> Hi,\n>\n> I¹m seeing behavior on 0.20.2 and 0.20.3 that doesn¹t seem quite right and\n> would like to know if this is by design, a bug, or something I¹m doing\n> wrong.\n>\n> Background:\n>\n> When I do a put that includes a timestamp like this (conceptually ­ I know\n> this is not the actual API), it works just fine.\n>  put ³table², ³family², ³column², ³bbb², 12345\n>\n> Then, if I do another put in the same client code using the same timestamp\n> like this...\n>  put ³table², ³family², ³column², ³aaa², 12345\n>\n> ...and I create a scanner, grab a Result, and iterate over all values using\n> list(), I get this...\n>  ³table², ³family², ³column², ³aaa², 12345\n>\n> So far, so good.  Now, if I truncate the table from the shell and run a new\n> program that does a flush() on the table between the two put¹s, but does it\n> in the same client program back-to-back, I also get the same results from\n> list().\n>\n> -----\n>\n> Problem:\n>\n> Here¹s where the trouble starts.  I truncate the table and run a new program\n> that puts ³bbb², flushes the table, and quits.  Here¹s what I get from\n> list():\n>  ³table², ³family², ³column², ³bbb², 12345\n>\n> Then I run another program that puts ³aaa², flushes, and quits.  Here¹s what\n> I get from list():\n>  ³table², ³family², ³column², ³aaa², 12345\n>  ³table², ³family², ³column², ³bbb², 12345\n>\n> And if I then run a third program that puts ³ccc², flushes, and quits, I get\n> this from list():\n>  ³table², ³family², ³column², ³ccc², 12345\n>  ³table², ³family², ³column², ³bbb², 12345\n>  ³table², ³family², ³column², ³aaa², 12345\n>\n> I¹m getting three different values for identical\n> table/family/qualifier/timestamp tuples.  Does this seem right?  There also\n> doesn¹t seem to be a defined sort order, probably because the timestamps are\n> identical.\n>\n> Also, if instead of using list(), I use getMap(), then I always only get a\n> single result.  The single result is always the last item in the lists above\n> (i.e., ³bbb² then ³bbb² then ³aaa²).  I get identical results from using\n> getNoVersionMap().\n>\n> I suspect that this same behavior could occur when HBase decides to flush on\n> its own, but I could be wrong.  As you can imagine, this can cause problems\n> because clients can¹t know from the results of calling list() which value is\n> ³right² or ³newest².  They also can¹t rely on getMap() or getNoVersionMap()\n> because the single result that gets returned is not necessarily ³right² or\n> ³newest².\n>\n> I¹ve reproduced everything above in a stand-alone installation and also with\n> a 7 regionserver cluster with the final 0.20.3.  I started down this\n> debugging path originally because I ran into this problem on the 7\n> regionserver cluster with one table of 100+ regions.  I was flushing\n> programmatically at the end of some large imports because I'm doing\n> setWriteToWAL(false) for load performance.\n>\n> Am I doing something wrong?  Did I miss an HBase assumption about flushing\n> and/or identical timestamps?\n>\n> Any help would be much appreciated.\n{code}","created":"2010-01-27T00:48:42.429+0000"},{"body":"Right now Get = 1 row Scans, thus making this issue at least twice as easy :-)","created":"2010-06-10T21:54:20.229+0000"},{"body":"Hurray!","created":"2010-09-03T02:29:32.053+0000"},{"body":"Patch for this issue posted at https://review.cloudera.org/r/780/","created":"2010-09-06T22:55:44.006+0000"},{"body":"We've tried the patch posted at https://review.cloudera.org/r/780/ here at Outerthought.\nThe attached file is a unit test showing that the patch works at first sight.\nHowever, performing updates on existing timestamps, in combination with triggering major compactions things don't work as expected.\n\nI've used negative-assertions in order to make the tests succeed, and added a comment where we would expect the result to be otherwise.\n\nI've also added a test with the example where a row is deleted and then an update on an older timestamp afterwards remains hidden by the delete.","created":"2010-09-07T15:44:49.636+0000"},{"body":"Great stuff Evert. This is really helpful.\n\nIn the delete case, this is still the expected behavior. This patch doesn't change anything with regards to insertions past deletions.","created":"2010-09-07T15:53:05.205+0000"},{"body":"The attached test TestCellUpdates.java also intializes a new HBaseTestingUtility for each test method.\nSetting this up only once for the whole test class causes issues when triggering the major compaction which I haven't been able to pinpoint yet.","created":"2010-09-07T15:57:08.284+0000"},{"body":"Thanks Evert, for point this out. As Jonathan mentioned, the behavior remains the same for the delete case. I will look more into the flushing case and keep you updated.","created":"2010-09-07T15:58:51.785+0000"},{"body":"I understand that the delete case isn't related to this jira. It would have been better to put it in a separate unit test.","created":"2010-09-07T16:15:30.749+0000"},{"body":"This is resolved. The fixed patch and explanation for this has been posted at https://review.cloudera.org/r/780/. Let me know if there are further questions. \n\nThanks again Evert, for making us aware of this case!","created":"2010-09-07T19:30:12.488+0000"},{"body":"Final patch posted by Pranav on reviewboard. Has been reviewed and approved for commit by Ryan and myself. Tests passing besides what already fails on trunk.","created":"2010-09-08T16:55:53.900+0000"},{"body":"Committed to trunk. Thanks Pranav! Great work!","created":"2010-09-08T17:23:06.125+0000"},{"body":"Thanks all!","created":"2010-09-08T19:13:27.937+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T13:01:15.858+0000"}],"conversations":[{"body":"As of now, both gets and scanners will end up returning all duplicate versions of a column. The ordering of them is indeterminate.\n\nWe need to decide what the desired/expected behavior should be and make it happen.\n\nNote: It's nearly impossible for this to work with Gets as they are now implemented in 1304 so this is really a Scanner issue. To implement this correctly with Gets, we would have to undo basically all the optimizations that Gets do and making them far slower than a Scanner.","from":"reporter","subject":"Wrong or indeterminate behavior when there are duplicate versions of a column"},{"body":"Just to say Gets and Scanners should work the same, whatever we decide.","from":"developer"},{"body":"I've had at least three people with a use case for this.\n\nMight create a couple sub-tasks here so we can at least head in the right direction.\n\nFirst, we need to make scanners ignore duplicate versions of the same column. The trickiest part is, how do we determine which to keep? We want to always come from the latest storefile, but I believe their IDs are still random and not timestamps? We might need to make that change to fix this. Would also then require a modification to the KVHeap to take this into account, all other things considered equal.\n\nOnce we have scanners working, that will mean the proper thing is enforced on major (and if we want, minor) compactions.\n\nGets will only work once we re-implement Gets as an optimized scan (taking advantage of bloom filters, mostly).\n\n\nI remember why I punted this to 0.20.1, the tricky part at the beginning is pretty tough and touches a good bit of core read-path code.\n\nRevisiting now, we'll see. Anyone else interested in this / want to work on it?","from":"developer"},{"body":"Yeah, as ryan suggested, we should exploit sequenceid -- maybe name file for sequenceid?","from":"developer"},{"body":"Bumped to 0.21","from":"developer"},{"body":"Related to hbase-997","from":"developer"},{"body":"From the list:\n\n{code}\nOn Tue, Jan 26, 2010 at 9:36 AM, Rod Cope wrote:\n> Hi,\n>\n> I¹m seeing behavior on 0.20.2 and 0.20.3 that doesn¹t seem quite right and\n> would like to know if this is by design, a bug, or something I¹m doing\n> wrong.\n>\n> Background:\n>\n> When I do a put that includes a timestamp like this (conceptually ­ I know\n> this is not the actual API), it works just fine.\n>  put ³table², ³family², ³column², ³bbb², 12345\n>\n> Then, if I do another put in the same client code using the same timestamp\n> like this...\n>  put ³table², ³family², ³column², ³aaa², 12345\n>\n> ...and I create a scanner, grab a Result, and iterate over all values using\n> list(), I get this...\n>  ³table², ³family², ³column², ³aaa², 12345\n>\n> So far, so good.  Now, if I truncate the table from the shell and run a new\n> program that does a flush() on the table between the two put¹s, but does it\n> in the same client program back-to-back, I also get the same results from\n> list().\n>\n> -----\n>\n> Problem:\n>\n> Here¹s where the trouble starts.  I truncate the table and run a new program\n> that puts ³bbb², flushes the table, and quits.  Here¹s what I get from\n> list():\n>  ³table², ³family², ³column², ³bbb², 12345\n>\n> Then I run another program that puts ³aaa², flushes, and quits.  Here¹s what\n> I get from list():\n>  ³table², ³family², ³column², ³aaa², 12345\n>  ³table², ³family², ³column², ³bbb², 12345\n>\n> And if I then run a third program that puts ³ccc², flushes, and quits, I get\n> this from list():\n>  ³table², ³family², ³column², ³ccc², 12345\n>  ³table², ³family², ³column², ³bbb², 12345\n>  ³table², ³family², ³column², ³aaa², 12345\n>\n> I¹m getting three different values for identical\n> table/family/qualifier/timestamp tuples.  Does this seem right?  There also\n> doesn¹t seem to be a defined sort order, probably because the timestamps are\n> identical.\n>\n> Also, if instead of using list(), I use getMap(), then I always only get a\n> single result.  The single result is always the last item in the lists above\n> (i.e., ³bbb² then ³bbb² then ³aaa²).  I get identical results from using\n> getNoVersionMap().\n>\n> I suspect that this same behavior could occur when HBase decides to flush on\n> its own, but I could be wrong.  As you can imagine, this can cause problems\n> because clients can¹t know from the results of calling list() which value is\n> ³right² or ³newest².  They also can¹t rely on getMap() or getNoVersionMap()\n> because the single result that gets returned is not necessarily ³right² or\n> ³newest².\n>\n> I¹ve reproduced everything above in a stand-alone installation and also with\n> a 7 regionserver cluster with the final 0.20.3.  I started down this\n> debugging path originally because I ran into this problem on the 7\n> regionserver cluster with one table of 100+ regions.  I was flushing\n> programmatically at the end of some large imports because I'm doing\n> setWriteToWAL(false) for load performance.\n>\n> Am I doing something wrong?  Did I miss an HBase assumption about flushing\n> and/or identical timestamps?\n>\n> Any help would be much appreciated.\n{code}","from":"developer"},{"body":"Right now Get = 1 row Scans, thus making this issue at least twice as easy :-)","from":"developer"},{"body":"Hurray!","from":"developer"},{"body":"Patch for this issue posted at https://review.cloudera.org/r/780/","from":"developer"},{"body":"We've tried the patch posted at https://review.cloudera.org/r/780/ here at Outerthought.\nThe attached file is a unit test showing that the patch works at first sight.\nHowever, performing updates on existing timestamps, in combination with triggering major compactions things don't work as expected.\n\nI've used negative-assertions in order to make the tests succeed, and added a comment where we would expect the result to be otherwise.\n\nI've also added a test with the example where a row is deleted and then an update on an older timestamp afterwards remains hidden by the delete.","from":"developer"},{"body":"Great stuff Evert. This is really helpful.\n\nIn the delete case, this is still the expected behavior. This patch doesn't change anything with regards to insertions past deletions.","from":"developer"},{"body":"The attached test TestCellUpdates.java also intializes a new HBaseTestingUtility for each test method.\nSetting this up only once for the whole test class causes issues when triggering the major compaction which I haven't been able to pinpoint yet.","from":"developer"},{"body":"Thanks Evert, for point this out. As Jonathan mentioned, the behavior remains the same for the delete case. I will look more into the flushing case and keep you updated.","from":"developer"},{"body":"I understand that the delete case isn't related to this jira. It would have been better to put it in a separate unit test.","from":"developer"},{"body":"This is resolved. The fixed patch and explanation for this has been posted at https://review.cloudera.org/r/780/. Let me know if there are further questions. \n\nThanks again Evert, for making us aware of this case!","from":"developer"},{"body":"Final patch posted by Pranav on reviewboard. Has been reviewed and approved for commit by Ryan and myself. Tests passing besides what already fails on trunk.","from":"developer"},{"body":"Committed to trunk. Thanks Pranav! Great work!","from":"developer"},{"body":"Thanks all!","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2009-06-05T18:08:05.000+0000","description":"As of now, both gets and scanners will end up returning all duplicate versions of a column. The ordering of them is indeterminate.\n\nWe need to decide what the desired/expected behavior should be and make it happen.\n\nNote: It's nearly impossible for this to work with Gets as they are now implemented in 1304 so this is really a Scanner issue. To implement this correctly with Gets, we would have to undo basically all the optimizations that Gets do and making them far slower than a Scanner.","issue_id":"12427237","key":"HBASE-1485","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2010-09-08T17:23:06.000+0000","role":"fixed_distractor","summary":"Wrong or indeterminate behavior when there are duplicate versions of a column"} {"case_id":"12428692","cluster":"DISTRACTOR-HBASE-1574","comments":[{"body":"How are you thinking this Ryan, that they should be buffered too, or just batched and sent over?","created":"2009-06-24T17:56:56.880+0000"},{"body":"Moving to 0.21","created":"2009-06-30T23:29:39.976+0000"},{"body":"Pulling into 0.20.1. We need it (I won't up the RPC version when I add the client/server methods since its a pure addition).","created":"2009-09-14T01:04:51.484+0000"},{"body":"Any chance of a review J-D?\n\nThis is not a pretty addition. In particular, the code in HConnectionManager is effectively duplicated. I tried to make a class to hold common code but it just made stuff worse; state is held in too many variables across the full breadth of the method. I factored out what I could. It would help if Put and Delete had an ancestor or even if it was just an interface such as Row with a getRow method in it or ComparableRow.\n\n","created":"2009-09-14T23:48:26.838+0000"},{"body":"Hang on.. had an idea.... back in a minute.","created":"2009-09-14T23:48:41.967+0000"},{"body":"I'm not against having a common base class or interface, like update or so. The reason that we didn't do that in the first place was the different way that we were using them, for example with buffering and so on, but this might change now when we are doing this multi delete stuff. ","created":"2009-09-15T00:22:01.857+0000"},{"body":"This is a bit better. The \"common\" interface is Row.. things that have a row can implement Row. Row has a getRow method. Thats it. I had RowComparable as the interface but thought it better to leave the two concepts distinct in case we want Row without Comparable.\n\nIn this patch, the duplication of code is gone. We instantiate a Batch private utility class to run the batching.\n\nJ-D (or Holstad), review please.\n\nI'm not upping the RPC version number because I want this to go into 0.20.1 (upping RPC version will make 0.20.0 incompatible with 0.20.1). I figure its ok because this is straight addtion of new method.","created":"2009-09-15T00:55:52.590+0000"},{"body":"In HRS, why do you copied the code from the delete method instead of calling it in a for()?\n\nIn HTable, I think there's something missing in the comment. \"If exception, list will have \"\n\nI like the refactoring of HCM and that it handles more than Put, that's the way to go. I think processBatchOfRows should have a new name, my goal when I wrote it was that it be the single method to handle any batches but now it's different.","created":"2009-09-15T15:05:35.689+0000"},{"body":"The batch delete is different from delete, no? Would be hard to refactor so they shared commonage.\n\nYeah, will fix the HTable comment on commit?\n\nYeah, processBatchOfRows is not a good name but I don't want to change it for this commit to 0.20 branch. I could open new issue for 0.21 where we look at these batch operations -- including batch get -- and figure out better method namings deprecating the old?\n\nOne other thing, I was going to change Row to be package private (Any reaction to this new interface? Do you think Row a good name? Should it be HasRow or something?).\n\nThanks J-D.\n\nSt.Ack","created":"2009-09-15T15:21:58.718+0000"},{"body":"I think Row is good, the rest is ok. +1 that you fix on commit.","created":"2009-09-15T15:29:40.056+0000"},{"body":"Patch looks good. \n\nI agree that we might have to rethink the structure of that code as it looks today if we want to support all these batch actions, to make it in a proper way.\nI would be good to know how people are using the batch methods to see what code we share between the different calls. ","created":"2009-09-15T17:43:04.324+0000"},{"body":"Commited branch and trunk (Thanks for review J-D and Holstad). I made Row package private and fixed javadoc on commit.","created":"2009-09-15T20:30:26.312+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T13:02:09.319+0000"}],"conversations":[{"body":"in 880 there is no way to do a batch delete (anymore?). We should add one back in.","from":"reporter","subject":"Client and server APIs to do batch deletes."},{"body":"How are you thinking this Ryan, that they should be buffered too, or just batched and sent over?","from":"developer"},{"body":"Moving to 0.21","from":"developer"},{"body":"Pulling into 0.20.1. We need it (I won't up the RPC version when I add the client/server methods since its a pure addition).","from":"developer"},{"body":"Any chance of a review J-D?\n\nThis is not a pretty addition. In particular, the code in HConnectionManager is effectively duplicated. I tried to make a class to hold common code but it just made stuff worse; state is held in too many variables across the full breadth of the method. I factored out what I could. It would help if Put and Delete had an ancestor or even if it was just an interface such as Row with a getRow method in it or ComparableRow.\n\n","from":"developer"},{"body":"Hang on.. had an idea.... back in a minute.","from":"developer"},{"body":"I'm not against having a common base class or interface, like update or so. The reason that we didn't do that in the first place was the different way that we were using them, for example with buffering and so on, but this might change now when we are doing this multi delete stuff. ","from":"developer"},{"body":"This is a bit better. The \"common\" interface is Row.. things that have a row can implement Row. Row has a getRow method. Thats it. I had RowComparable as the interface but thought it better to leave the two concepts distinct in case we want Row without Comparable.\n\nIn this patch, the duplication of code is gone. We instantiate a Batch private utility class to run the batching.\n\nJ-D (or Holstad), review please.\n\nI'm not upping the RPC version number because I want this to go into 0.20.1 (upping RPC version will make 0.20.0 incompatible with 0.20.1). I figure its ok because this is straight addtion of new method.","from":"developer"},{"body":"In HRS, why do you copied the code from the delete method instead of calling it in a for()?\n\nIn HTable, I think there's something missing in the comment. \"If exception, list will have \"\n\nI like the refactoring of HCM and that it handles more than Put, that's the way to go. I think processBatchOfRows should have a new name, my goal when I wrote it was that it be the single method to handle any batches but now it's different.","from":"developer"},{"body":"The batch delete is different from delete, no? Would be hard to refactor so they shared commonage.\n\nYeah, will fix the HTable comment on commit?\n\nYeah, processBatchOfRows is not a good name but I don't want to change it for this commit to 0.20 branch. I could open new issue for 0.21 where we look at these batch operations -- including batch get -- and figure out better method namings deprecating the old?\n\nOne other thing, I was going to change Row to be package private (Any reaction to this new interface? Do you think Row a good name? Should it be HasRow or something?).\n\nThanks J-D.\n\nSt.Ack","from":"developer"},{"body":"I think Row is good, the rest is ok. +1 that you fix on commit.","from":"developer"},{"body":"Patch looks good. \n\nI agree that we might have to rethink the structure of that code as it looks today if we want to support all these batch actions, to make it in a proper way.\nI would be good to know how people are using the batch methods to see what code we share between the different calls. ","from":"developer"},{"body":"Commited branch and trunk (Thanks for review J-D and Holstad). I made Row package private and fixed javadoc on commit.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2009-06-23T18:51:37.000+0000","description":"in 880 there is no way to do a batch delete (anymore?). We should add one back in.","issue_id":"12428692","key":"HBASE-1574","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2009-09-15T20:30:26.000+0000","role":"fixed_distractor","summary":"Client and server APIs to do batch deletes."} {"case_id":"12428823","cluster":"DISTRACTOR-HBASE-1583","comments":[{"body":"some thoughts:\n\n- get rid of compactions during startup/shutdown. There is no reason to run compactions at those times.\n- dont worry about compacting reference regions as quickly, we can wait a few minutes if necessary.\n- make shut down flush memstore as fast as possible, then get the heck outta there.\n\nwe need to do more work on bringing up regions as fast as possible in 0.21. The master needs to blast region assignments out there in a threaded/performant manner, rather than waiting for 3-second checkins to do their magic.\n\nFinally I have decided to up the size of a region on my own set-up. 256MB seems a little too small, and with even twice as large regions, i end up with less than half the region count!","created":"2009-06-25T01:17:07.769+0000"},{"body":"@stack:\n\n> startup is a mess with our assigning out regions an rebalancing at same time. By time that the\n> compactions on open run, it can be near an hour before whole thing settles down and becomes\n> useable\n\nsafe mode was to prevent rebalancing during startup. Are we not using safe mode anymore?\n\n@ryan\n\n> The master needs to blast region assignments out there in a threaded/performant manner\n\nWhat I was planning to do for 0.21 when I put region assignments in ZK was to make a znode\nwhose children are unassigned regions. A region server can then decide if it is too busy or\nnot, and if not, remove the unassigned region from the list and add it to its list of regions being\nserved once the region is available.\n\nThis removes the master from region assignment. It would then only need to detect unassigned\nregions.","created":"2009-06-25T01:35:37.107+0000"},{"body":"Safe mode is still there. Thats just period during which all machines report in and during which we hand out catalog regions. After safe mode elapses, then the mayhem breaks out as master tries to hand out 6k regions ten or so at a time balancing at same time.\n\nRegion assignment needs to part of larger scale rewrite of master function. Hows does a master figure a region unassigned? It reads .META. table to figure current state. We need to be careful how we bridge scan of .META. and read of zk.","created":"2009-06-25T03:34:23.758+0000"},{"body":"It would also be neato to have more optimal global assignment of regions. Any given table should be spread out as much as possible.\n\n","created":"2009-06-25T03:40:29.495+0000"},{"body":"I like your idea, ryan. Next generation of balancing needs to take into account actual usage that comprises a complex notion of \"load\" including memcache, er memstore, usage, block cache usage, and as you say, an additional factor could be how many other regions from the same table are already assigned to that regionserver.","created":"2009-06-25T22:31:29.233+0000"},{"body":"My understanding for 0.21 is the region assignment process is going to be largely unmediated by the master, except for the case where the master finds an unassigned region in META and puts up a node into the \"to be assigned\" queue out in ZK. My opinion this is the way to go, but how a regionserver is to judge its load relative to others, or even learn about the load of others, is an unanswered question. Furthermore, there must be a mechanism in place such that some regionserver will take on a new region if all others have passed on it. Is there an issue up for this type of stuff yet? ","created":"2009-06-26T00:11:32.089+0000"},{"body":"We should open an issue Andrew. What'll we call it? Master redesign?","created":"2009-06-26T03:39:45.925+0000"},{"body":"+1 on this issue\n\nAlso btw the Master redesign is HBASE-1110.","created":"2009-06-26T15:54:27.514+0000"},{"body":"I suggested that we do not do come out of safe mode until all regions have been assigned when we added safe mode and make the regions not run compactions while in safe mode I thank that would be an easy fix for this problem\nI have seen the same thing when you have region that are behind on compactions after a shutdown on start up compaction tie up reassignments.\n\nBilly\n","created":"2009-06-29T15:33:43.320+0000"},{"body":"+1 on the kill compaction on shutdown why not just kill the thread and not let it finish it will get done on restart so we will not lose anything.\n","created":"2009-06-29T15:35:56.360+0000"},{"body":"Small jgray suggested optimization; don't compact on open unless has References","created":"2009-07-03T21:36:33.775+0000"},{"body":"Here is a patch to disable compaction on close of a region (helps with disable of table too).","created":"2009-07-16T18:36:42.217+0000"},{"body":"+1 on 1583-nocompactonclose.patch","created":"2009-07-16T19:20:20.645+0000"},{"body":"This patch adds to the previous. It adds no compaction on open unless region has references.","created":"2009-07-16T19:25:41.700+0000"},{"body":"Looks good to me.","created":"2009-07-16T19:58:32.732+0000"},{"body":"With this patch applied and with HBASE-1058 reverted, up and down as well as enable/disable is much smoother.\n\nLet me now take a look at 'safe mode'. Talking with JK, looks like its not working properly. All regions were supposed to be assigned while in safe mode but not working.","created":"2009-07-16T20:58:06.324+0000"},{"body":"Ack on safe mode having oddities. I brought down a cluster with 133 regions cleanly and restarted it just now. Right away ~128 regions were assigned out. The rest were assigned out a few minutes later. Invoking 'enable' on a incompletely assigned table prodded the master into some action but did not bring things up all the way as this kludge has done in the past. ","created":"2009-07-16T21:21:25.836+0000"},{"body":"I took a look at this safe mode stuff. Its broken. Will open an issue. Whats happening is that we exit safe mode near immediately after startup because initial MetaScanner scan does nothing except set that initial scan has completed (though it did nothing -- original idea was that initialScan would do first scan of the newly deployed .META.). So, we exit safe mode near immediately after startup.\n\nFixing metascanner so initial scan doesn't happen till we've scanned actual deploy so safe mode stays in place while deploy is going on kills our assignment rate. It crawls. I gave up trying to debug more since these above patches undoing compactions on close and open seem to be enough to close this issue at least for 0.20.0 release.","created":"2009-07-16T23:16:24.732+0000"},{"body":"Committed to branch and trunk. Resolving this issue. There is more we can do but this I think is enough for 0.20.0.","created":"2009-07-16T23:30:52.631+0000"}],"conversations":[{"body":"Starting and stopping a loaded large cluster is way too flakey and takes too long. This is 0.19.x but same issues apply to TRUNK I'd say.\n\nAt pset with our > 100 nodes carrying 6k regions:\n\n+ shutdown takes way too long.... maybe ten minutes or so. We compact regions inline with shutdown. We should just go down. It doesn't seem like all regionservers go down everytime either.\n+ startup is a mess with our assigning out regions an rebalancing at same time. By time that the compactions on open run, it can be near an hour before whole thing settles down and becomes useable","from":"reporter","subject":"Start/Stop of large cluster untenable"},{"body":"some thoughts:\n\n- get rid of compactions during startup/shutdown. There is no reason to run compactions at those times.\n- dont worry about compacting reference regions as quickly, we can wait a few minutes if necessary.\n- make shut down flush memstore as fast as possible, then get the heck outta there.\n\nwe need to do more work on bringing up regions as fast as possible in 0.21. The master needs to blast region assignments out there in a threaded/performant manner, rather than waiting for 3-second checkins to do their magic.\n\nFinally I have decided to up the size of a region on my own set-up. 256MB seems a little too small, and with even twice as large regions, i end up with less than half the region count!","from":"developer"},{"body":"@stack:\n\n> startup is a mess with our assigning out regions an rebalancing at same time. By time that the\n> compactions on open run, it can be near an hour before whole thing settles down and becomes\n> useable\n\nsafe mode was to prevent rebalancing during startup. Are we not using safe mode anymore?\n\n@ryan\n\n> The master needs to blast region assignments out there in a threaded/performant manner\n\nWhat I was planning to do for 0.21 when I put region assignments in ZK was to make a znode\nwhose children are unassigned regions. A region server can then decide if it is too busy or\nnot, and if not, remove the unassigned region from the list and add it to its list of regions being\nserved once the region is available.\n\nThis removes the master from region assignment. It would then only need to detect unassigned\nregions.","from":"developer"},{"body":"Safe mode is still there. Thats just period during which all machines report in and during which we hand out catalog regions. After safe mode elapses, then the mayhem breaks out as master tries to hand out 6k regions ten or so at a time balancing at same time.\n\nRegion assignment needs to part of larger scale rewrite of master function. Hows does a master figure a region unassigned? It reads .META. table to figure current state. We need to be careful how we bridge scan of .META. and read of zk.","from":"developer"},{"body":"It would also be neato to have more optimal global assignment of regions. Any given table should be spread out as much as possible.\n\n","from":"developer"},{"body":"I like your idea, ryan. Next generation of balancing needs to take into account actual usage that comprises a complex notion of \"load\" including memcache, er memstore, usage, block cache usage, and as you say, an additional factor could be how many other regions from the same table are already assigned to that regionserver.","from":"developer"},{"body":"My understanding for 0.21 is the region assignment process is going to be largely unmediated by the master, except for the case where the master finds an unassigned region in META and puts up a node into the \"to be assigned\" queue out in ZK. My opinion this is the way to go, but how a regionserver is to judge its load relative to others, or even learn about the load of others, is an unanswered question. Furthermore, there must be a mechanism in place such that some regionserver will take on a new region if all others have passed on it. Is there an issue up for this type of stuff yet? ","from":"developer"},{"body":"We should open an issue Andrew. What'll we call it? Master redesign?","from":"developer"},{"body":"+1 on this issue\n\nAlso btw the Master redesign is HBASE-1110.","from":"developer"},{"body":"I suggested that we do not do come out of safe mode until all regions have been assigned when we added safe mode and make the regions not run compactions while in safe mode I thank that would be an easy fix for this problem\nI have seen the same thing when you have region that are behind on compactions after a shutdown on start up compaction tie up reassignments.\n\nBilly\n","from":"developer"},{"body":"+1 on the kill compaction on shutdown why not just kill the thread and not let it finish it will get done on restart so we will not lose anything.\n","from":"developer"},{"body":"Small jgray suggested optimization; don't compact on open unless has References","from":"developer"},{"body":"Here is a patch to disable compaction on close of a region (helps with disable of table too).","from":"developer"},{"body":"+1 on 1583-nocompactonclose.patch","from":"developer"},{"body":"This patch adds to the previous. It adds no compaction on open unless region has references.","from":"developer"},{"body":"Looks good to me.","from":"developer"},{"body":"With this patch applied and with HBASE-1058 reverted, up and down as well as enable/disable is much smoother.\n\nLet me now take a look at 'safe mode'. Talking with JK, looks like its not working properly. All regions were supposed to be assigned while in safe mode but not working.","from":"developer"},{"body":"Ack on safe mode having oddities. I brought down a cluster with 133 regions cleanly and restarted it just now. Right away ~128 regions were assigned out. The rest were assigned out a few minutes later. Invoking 'enable' on a incompletely assigned table prodded the master into some action but did not bring things up all the way as this kludge has done in the past. ","from":"developer"},{"body":"I took a look at this safe mode stuff. Its broken. Will open an issue. Whats happening is that we exit safe mode near immediately after startup because initial MetaScanner scan does nothing except set that initial scan has completed (though it did nothing -- original idea was that initialScan would do first scan of the newly deployed .META.). So, we exit safe mode near immediately after startup.\n\nFixing metascanner so initial scan doesn't happen till we've scanned actual deploy so safe mode stays in place while deploy is going on kills our assignment rate. It crawls. I gave up trying to debug more since these above patches undoing compactions on close and open seem to be enough to close this issue at least for 0.20.0 release.","from":"developer"},{"body":"Committed to branch and trunk. Resolving this issue. There is more we can do but this I think is enough for 0.20.0.","from":"developer"}],"created":"2009-06-25T01:05:08.000+0000","description":"Starting and stopping a loaded large cluster is way too flakey and takes too long. This is 0.19.x but same issues apply to TRUNK I'd say.\n\nAt pset with our > 100 nodes carrying 6k regions:\n\n+ shutdown takes way too long.... maybe ten minutes or so. We compact regions inline with shutdown. We should just go down. It doesn't seem like all regionservers go down everytime either.\n+ startup is a mess with our assigning out regions an rebalancing at same time. By time that the compactions on open run, it can be near an hour before whole thing settles down and becomes useable","issue_id":"12428823","key":"HBASE-1583","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2009-07-16T23:30:52.000+0000","role":"fixed_distractor","summary":"Start/Stop of large cluster untenable"} {"case_id":"12974271","cluster":"DISTRACTOR-HBASE-15925","comments":[{"body":"FYI [~mantonov], [~busbey]","created":"2016-05-31T16:15:52.475+0000"},{"body":"-01\n\n - add properties to the top level pom that match those in the hadoop-2.0 profile","created":"2016-06-09T18:32:19.385+0000"},{"body":"Skimmed, looks good to me, except typo on the last line (nit).","created":"2016-06-09T19:21:33.183+0000"},{"body":"you fine with fixing that nit on commit, provided precommit doesn't explode in an unexpected way?","created":"2016-06-09T19:42:19.402+0000"},{"body":"totally, +1","created":"2016-06-09T20:29:09.599+0000"},{"body":"thanks for fixing the issue.\nwhen will 1.2.2 be published in maven central?\nhttp://mvnrepository.com/artifact/org.apache.hbase/hbase-testing-util\n","created":"2016-06-12T08:56:17.381+0000"},{"body":"bq. when will 1.2.2 be published in maven central?\n\nWe'll have to hold a release vote, but we've been waiting to do that on the 1.2.z line for about 2 months. I'm hoping to get an RC started in time to have it close by the end of the week.","created":"2016-06-13T13:55:42.876+0000"},{"body":"has 1.2.2 been published?\notherwise, is there something I can do to help make it happen?","created":"2016-07-07T13:40:15.692+0000"},{"body":"Nice timing [~dportabella]. The latest RC for 1.2.2 has been posted, and includes a staged maven repository -- this is what will be promoted to maven central assuming this release passes. Take a look: https://repository.apache.org/content/repositories/orgapachehbase-1142/\n\nFrom what I can tell, this is not fixed in the RC. https://repository.apache.org/content/repositories/orgapachehbase-1142/org/apache/hbase/hbase-testing-util/1.2.2/hbase-testing-util-1.2.2.pom\n\nFYI [~busbey]","created":"2016-07-08T01:29:34.291+0000"},{"body":"Reopening issue. Sorry fellas.","created":"2016-07-08T01:30:07.347+0000"},{"body":"Comparing the pom in the repo on r2 vs that in central for 1.2.1, I don't understand why the slf4j.version property a couple lines down is apparently parsed correctly while the compat module variable is not.","created":"2016-07-08T01:35:48.606+0000"},{"body":"I'm no longer convinced this can be confirmed with visual inspection. [~dportabella] can you plug this repository and version into your project and verify? Maybe you can contribute a patch to test this to our [hbase-downstreamer|https://github.com/saintstack/hbase-downstreamer] project.","created":"2016-07-08T01:38:48.467+0000"},{"body":"I no longer understand.\nI've tried with 1.2.2 and it works, but now it also works on 1.2.1 from maven central.\n\nI've tried as follows:\n\n{code}\n$ mvn archetype:generate -DgroupId=com.example -DartifactId=App -DarchetypeArtifactId=maven-archetype-quickstart -DinteractiveMode=true\n{code}\n\nI've added the hbase-testing-util 1.2.2 dependency and the https://repository.apache.org/content/groups/staging/ repository.\n\npom.xml\n{code}\n\n 4.0.0\n com.example\n App\n jar\n 1.0-SNAPSHOT\n App\n http://maven.apache.org\n\n \n \n junit\n junit\n 3.8.1\n test\n \n \n org.apache.hbase\n hbase-testing-util\n 1.2.2\n \n \n\n \n \n apache_stage\n Apache Stage\n https://repository.apache.org/content/groups/staging/\n \n \n\n{code}\n\nAnd run:\n{code}\n$ rm -Rf /david/.m2/repository/org/apache/hbase/\n$ mvn clean\n$ mvn package\n$ mvn exec:java -Dexec.mainClass=\"com.example.App\"\n{code}\n\nand it works.\nhowever, removing the repository and using 1.2.1 works also now.\n\nI am investigating.\n","created":"2016-07-08T14:01:17.311+0000"},{"body":"I first reported the problem when running a scala code that depends on hbase version \"0.98.7-hadoop2\".\n\nUpdating this code to use hbase version 1.2.1 (or anything from 1.x to 1.21) fails with:\n{code}\nUnresolved dependencies path: org.apache.hbase:${compat.module}:1.2.1\n{code}\n\nit works now using hbase version 1.2.2.\n\nif you want to try:\n{code}\n$ git clone https://github.com/dportabella/spark-examples.git\n# edit build.sbt:\n replace: val hbaseVersion = \"1.2.2\"\n add: resolvers += \"Apache Staging\" at \"https://repository.apache.org/content/groups/staging/\"\n\n$ rm -Rf /david/.ivy2//cache/org.apache.hbase\n$ rm -Rf /david/.m2//repository/org/apache/hbase/\n$ sbt run\n{code}\n\nSo, I don't know why maven does not complain on 1.2.1 (see my post above),\nbut sbt complains on 1.2.1 and now it works on 1.2.2.\nSo, I guess the problem is solved. Thanks!\n","created":"2016-07-08T14:58:41.232+0000"},{"body":"{quote}\nI'm no longer convinced this can be confirmed with visual inspection. \n{quote}\n\nNope, can't be verified by looking because it's a matter of some tools treating Maven pom's incorrectly (most notably sbt). I'm not aware of any web tool that e.g. uses sbt to make a effective pom we could check.\n\n{quote}\nSo, I don't know why maven does not complain on 1.2.1 (see my post above),\nbut sbt complains on 1.2.1 and now it works on 1.2.2.\nSo, I guess the problem is solved. Thanks!\n{quote}\n\nGlad to hear sbt now works! Maven should work correctly both before and after this patch (before only so long as it's a version that correctly handles default profiles, which should be all of the maven 3s).\n\nI'll resolve this again now that we have things confirmed. Thanks for rechecking!","created":"2016-07-08T15:41:11.514+0000"},{"body":"{quote}\nhas 1.2.2 been published?\notherwise, is there something I can do to help make it happen?\n{quote}\n\nTo answer this question, it hasn't been published yet. We're on our 3rd release candidate, which so far looks good. The vote is scheduled to close on Tuesday 12 July. If you'd like more details, check out the thread on dev@ '[VOTE] Third release candidate for HBase 1.2.2 (RC2)'.\n\nAll votes on the release candidate help. You can check out some of the existing votes to see what kinds of things folks check.","created":"2016-07-08T15:56:57.850+0000"},{"body":"{quote}\n> So, I don't know why maven does not complain on 1.2.1 (see my post above),\n> but sbt complains on 1.2.1 and now it works on 1.2.2.\n> So, I guess the problem is solved. Thanks!\nGlad to hear sbt now works! Maven should work correctly both before and after this patch (before only so long as it's a version that correctly handles default profiles, which should be all of the maven 3s).\n{quote}\nI am not sure that you understood what I mean. I will rephrase it:\n\nhbase-testing-util version 1.2.1 has a problem, that has a dependency with a variable not evaluated (it's the point of this ticket actually)\nhttp://mvnrepository.com/artifact/org.apache.hbase/hbase-testing-util/1.2.1\n\nso, a maven pom.xml which declares a dependency on hbase-testing-util version 1.2.1 should fail, as it cannot download all of its dependencies.\n\nhowever, maven does not fail (and it should) as I mentioned in a this post above. so, maybe there is a bug on maven.\nhttps://issues.apache.org/jira/browse/HBASE-15925?focusedCommentId=15367707&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15367707\n\non the other hand, sbt complains if a project depends on hbase-testing-util version 1.2.1, as it should.\nand, sbt succeeds if a project depends on hbase-testing-util version 1.2.2, meaning that this ticket has fixed the problem on hbase.\n\nso, in this sense, there is possibly a bug on maven, not on sbt.\n","created":"2016-07-08T16:30:12.101+0000"},{"body":"{quote}\nhbase-testing-util version 1.2.1 has a problem, that has a dependency with a variable not evaluated (it's the point of this ticket actually)\nhttp://mvnrepository.com/artifact/org.apache.hbase/hbase-testing-util/1.2.1\n{quote}\n\nvariables don't get evaluated until runtime. They're allowed in all parts of dependency coordinates.\n\n{quote}\nso, a maven pom.xml which declares a dependency on hbase-testing-util version 1.2.1 should fail, as it cannot download all of its dependencies.\nhowever, maven does not fail (and it should) as I mentioned in a this post above. so, maybe there is a bug on maven.\nhttps://issues.apache.org/jira/browse/HBASE-15925?focusedCommentId=15367707&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15367707\n{quote}\n\nThis is correct behavior (evaluating the variable at runtime, then getting said dependency), which is why Maven works.\n\n{quote}\non the other hand, sbt complains if a project depends on hbase-testing-util version 1.2.1, as it should.\nand, sbt succeeds if a project depends on hbase-testing-util version 1.2.2, meaning that this ticket has fixed the problem on hbase.\nso, in this sense, there is possibly a bug on maven, not on sbt.\n{quote}\n\nWell, I'm glad we agree that the proximal issue is solved. :)\n\nThe underlying bug is in sbt. It isn't properly activating profiles. In particular, we rely on [\"property isn't defined at invocation\"|http://books.sonatype.com/mvnref-book/reference/profiles-sect-activation.html#profiles-sect-activation-by-absence] activation to set a bunch of defaults around what version of hadoop gets used. ([this one here in 1.2.2RC2|https://github.com/apache/hbase/blob/1.2.2RC2/pom.xml#L1935]).\n\nThis JIRA fixed the issue by ensuring we have defined defaults for all the needed properties that match that default profile, so that when tools fail to activate the profiles we ask for they get something that works-at-all.\n","created":"2016-07-08T16:48:51.830+0000"},{"body":"[~dportabella] can you verify 1.2.2rc2 specifically with it's repository: https://repository.apache.org/content/repositories/orgapachehbase-1142/\n\nI don't know exactly which artifacts are published into https://repository.apache.org/content/groups/staging/ whereas the content of orgapachehbase-1142 are exactly what will become 1.2.2 when the vote passes.","created":"2016-07-08T16:55:18.201+0000"},{"body":"oh, i understand now.\nthanks for the explanation. :)\n\n","created":"2016-07-08T16:59:02.818+0000"},{"body":"{quote}\nDavid Portabella can you verify 1.2.2rc2 specifically with it's repository: https://repository.apache.org/content/repositories/orgapachehbase-1142/\n{quote}\n\nI try as follows:\n{code}\n$ git clone https://github.com/dportabella/spark-examples.git\n# edit build.sbt:\n replace: val hbaseVersion = \"1.2.2rc2\"\n add: resolvers += \"Apache Staging\" at \"https://repository.apache.org/content/repositories/orgapachehbase-1142/\"\n\n$ rm -Rf /david/.ivy2/cache/org.apache.hbase\n$ rm -Rf /david/.m2/repository/org/apache/hbase/\n$ sbt run\n{code}\n\nresult:\n{code}\n[info] Resolving org.apache.hbase#hbase-testing-util;1.2.2rc2 ...\n[warn] \tmodule not found: org.apache.hbase#hbase-testing-util;1.2.2rc2\n[warn] https://repository.apache.org/content/repositories/orgapachehbase-1142/org/apache/hbase/hbase-testing-util/1.2.2rc2/hbase-testing-util-1.2.2rc2.pom\n{code}\n\nit does not find 1.2.2rc2. am i doing sthy wrong?\n\n","created":"2016-07-08T19:58:53.142+0000"},{"body":"Yeah, with that repository URL, the version is just 1.2.2 -- no rc part.","created":"2016-07-08T20:07:32.286+0000"},{"body":"Ok, yes, it works :)\n","created":"2016-07-08T22:12:03.071+0000"},{"body":"Sean, I've filled a bug report at sbt based on your answer. could you please check it?\nhttps://github.com/sbt/sbt/issues/2666","created":"2016-07-11T17:16:13.329+0000"},{"body":"I see that two new versions are published in maven central (1.2.2 and 1.2.3), that's great!\n\none question:\nthis version 1.2.1 (and all previous versions) had this direct dependency's artifactId on ${compat.module}, and we could see this here:\nhttp://mvnrepository.com/artifact/org.apache.hbase/hbase-testing-util/1.2.1\n\nhowever, I don't see this anymore. why?\n\nDoes this mean that you have fixed and replaced the 1.2.1 version (and all the previous ones)?\nIf so, isn't that incorrect? a release version should stay immutable forever.\nCompiling a code which depends on a release dependency should always produce the same result.\n\nOr maybe mvnrepository.com now resolves these variables, and it was not the case when this ticket was open?\n","created":"2016-09-21T13:55:20.270+0000"}],"conversations":[{"body":"Looks like we've regressed on HBASE-8488. Have a look at the dependency artifacts list on http://mvnrepository.com/artifact/org.apache.hbase/hbase-testing-util/1.2.1. Notice the direct dependency's artifactId is {{$\\{compat.module\\}}}.","from":"reporter","subject":"compat-module maven variable not evaluated"},{"body":"FYI [~mantonov], [~busbey]","from":"developer"},{"body":"-01\n\n - add properties to the top level pom that match those in the hadoop-2.0 profile","from":"developer"},{"body":"Skimmed, looks good to me, except typo on the last line (nit).","from":"developer"},{"body":"you fine with fixing that nit on commit, provided precommit doesn't explode in an unexpected way?","from":"developer"},{"body":"totally, +1","from":"developer"},{"body":"thanks for fixing the issue.\nwhen will 1.2.2 be published in maven central?\nhttp://mvnrepository.com/artifact/org.apache.hbase/hbase-testing-util\n","from":"developer"},{"body":"bq. when will 1.2.2 be published in maven central?\n\nWe'll have to hold a release vote, but we've been waiting to do that on the 1.2.z line for about 2 months. I'm hoping to get an RC started in time to have it close by the end of the week.","from":"developer"},{"body":"has 1.2.2 been published?\notherwise, is there something I can do to help make it happen?","from":"developer"},{"body":"Nice timing [~dportabella]. The latest RC for 1.2.2 has been posted, and includes a staged maven repository -- this is what will be promoted to maven central assuming this release passes. Take a look: https://repository.apache.org/content/repositories/orgapachehbase-1142/\n\nFrom what I can tell, this is not fixed in the RC. https://repository.apache.org/content/repositories/orgapachehbase-1142/org/apache/hbase/hbase-testing-util/1.2.2/hbase-testing-util-1.2.2.pom\n\nFYI [~busbey]","from":"developer"},{"body":"Reopening issue. Sorry fellas.","from":"developer"},{"body":"Comparing the pom in the repo on r2 vs that in central for 1.2.1, I don't understand why the slf4j.version property a couple lines down is apparently parsed correctly while the compat module variable is not.","from":"developer"},{"body":"I'm no longer convinced this can be confirmed with visual inspection. [~dportabella] can you plug this repository and version into your project and verify? Maybe you can contribute a patch to test this to our [hbase-downstreamer|https://github.com/saintstack/hbase-downstreamer] project.","from":"developer"},{"body":"I no longer understand.\nI've tried with 1.2.2 and it works, but now it also works on 1.2.1 from maven central.\n\nI've tried as follows:\n\n{code}\n$ mvn archetype:generate -DgroupId=com.example -DartifactId=App -DarchetypeArtifactId=maven-archetype-quickstart -DinteractiveMode=true\n{code}\n\nI've added the hbase-testing-util 1.2.2 dependency and the https://repository.apache.org/content/groups/staging/ repository.\n\npom.xml\n{code}\n\n 4.0.0\n com.example\n App\n jar\n 1.0-SNAPSHOT\n App\n http://maven.apache.org\n\n \n \n junit\n junit\n 3.8.1\n test\n \n \n org.apache.hbase\n hbase-testing-util\n 1.2.2\n \n \n\n \n \n apache_stage\n Apache Stage\n https://repository.apache.org/content/groups/staging/\n \n \n\n{code}\n\nAnd run:\n{code}\n$ rm -Rf /david/.m2/repository/org/apache/hbase/\n$ mvn clean\n$ mvn package\n$ mvn exec:java -Dexec.mainClass=\"com.example.App\"\n{code}\n\nand it works.\nhowever, removing the repository and using 1.2.1 works also now.\n\nI am investigating.\n","from":"developer"},{"body":"I first reported the problem when running a scala code that depends on hbase version \"0.98.7-hadoop2\".\n\nUpdating this code to use hbase version 1.2.1 (or anything from 1.x to 1.21) fails with:\n{code}\nUnresolved dependencies path: org.apache.hbase:${compat.module}:1.2.1\n{code}\n\nit works now using hbase version 1.2.2.\n\nif you want to try:\n{code}\n$ git clone https://github.com/dportabella/spark-examples.git\n# edit build.sbt:\n replace: val hbaseVersion = \"1.2.2\"\n add: resolvers += \"Apache Staging\" at \"https://repository.apache.org/content/groups/staging/\"\n\n$ rm -Rf /david/.ivy2//cache/org.apache.hbase\n$ rm -Rf /david/.m2//repository/org/apache/hbase/\n$ sbt run\n{code}\n\nSo, I don't know why maven does not complain on 1.2.1 (see my post above),\nbut sbt complains on 1.2.1 and now it works on 1.2.2.\nSo, I guess the problem is solved. Thanks!\n","from":"developer"},{"body":"{quote}\nI'm no longer convinced this can be confirmed with visual inspection. \n{quote}\n\nNope, can't be verified by looking because it's a matter of some tools treating Maven pom's incorrectly (most notably sbt). I'm not aware of any web tool that e.g. uses sbt to make a effective pom we could check.\n\n{quote}\nSo, I don't know why maven does not complain on 1.2.1 (see my post above),\nbut sbt complains on 1.2.1 and now it works on 1.2.2.\nSo, I guess the problem is solved. Thanks!\n{quote}\n\nGlad to hear sbt now works! Maven should work correctly both before and after this patch (before only so long as it's a version that correctly handles default profiles, which should be all of the maven 3s).\n\nI'll resolve this again now that we have things confirmed. Thanks for rechecking!","from":"developer"},{"body":"{quote}\nhas 1.2.2 been published?\notherwise, is there something I can do to help make it happen?\n{quote}\n\nTo answer this question, it hasn't been published yet. We're on our 3rd release candidate, which so far looks good. The vote is scheduled to close on Tuesday 12 July. If you'd like more details, check out the thread on dev@ '[VOTE] Third release candidate for HBase 1.2.2 (RC2)'.\n\nAll votes on the release candidate help. You can check out some of the existing votes to see what kinds of things folks check.","from":"developer"},{"body":"{quote}\n> So, I don't know why maven does not complain on 1.2.1 (see my post above),\n> but sbt complains on 1.2.1 and now it works on 1.2.2.\n> So, I guess the problem is solved. Thanks!\nGlad to hear sbt now works! Maven should work correctly both before and after this patch (before only so long as it's a version that correctly handles default profiles, which should be all of the maven 3s).\n{quote}\nI am not sure that you understood what I mean. I will rephrase it:\n\nhbase-testing-util version 1.2.1 has a problem, that has a dependency with a variable not evaluated (it's the point of this ticket actually)\nhttp://mvnrepository.com/artifact/org.apache.hbase/hbase-testing-util/1.2.1\n\nso, a maven pom.xml which declares a dependency on hbase-testing-util version 1.2.1 should fail, as it cannot download all of its dependencies.\n\nhowever, maven does not fail (and it should) as I mentioned in a this post above. so, maybe there is a bug on maven.\nhttps://issues.apache.org/jira/browse/HBASE-15925?focusedCommentId=15367707&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15367707\n\non the other hand, sbt complains if a project depends on hbase-testing-util version 1.2.1, as it should.\nand, sbt succeeds if a project depends on hbase-testing-util version 1.2.2, meaning that this ticket has fixed the problem on hbase.\n\nso, in this sense, there is possibly a bug on maven, not on sbt.\n","from":"developer"},{"body":"{quote}\nhbase-testing-util version 1.2.1 has a problem, that has a dependency with a variable not evaluated (it's the point of this ticket actually)\nhttp://mvnrepository.com/artifact/org.apache.hbase/hbase-testing-util/1.2.1\n{quote}\n\nvariables don't get evaluated until runtime. They're allowed in all parts of dependency coordinates.\n\n{quote}\nso, a maven pom.xml which declares a dependency on hbase-testing-util version 1.2.1 should fail, as it cannot download all of its dependencies.\nhowever, maven does not fail (and it should) as I mentioned in a this post above. so, maybe there is a bug on maven.\nhttps://issues.apache.org/jira/browse/HBASE-15925?focusedCommentId=15367707&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15367707\n{quote}\n\nThis is correct behavior (evaluating the variable at runtime, then getting said dependency), which is why Maven works.\n\n{quote}\non the other hand, sbt complains if a project depends on hbase-testing-util version 1.2.1, as it should.\nand, sbt succeeds if a project depends on hbase-testing-util version 1.2.2, meaning that this ticket has fixed the problem on hbase.\nso, in this sense, there is possibly a bug on maven, not on sbt.\n{quote}\n\nWell, I'm glad we agree that the proximal issue is solved. :)\n\nThe underlying bug is in sbt. It isn't properly activating profiles. In particular, we rely on [\"property isn't defined at invocation\"|http://books.sonatype.com/mvnref-book/reference/profiles-sect-activation.html#profiles-sect-activation-by-absence] activation to set a bunch of defaults around what version of hadoop gets used. ([this one here in 1.2.2RC2|https://github.com/apache/hbase/blob/1.2.2RC2/pom.xml#L1935]).\n\nThis JIRA fixed the issue by ensuring we have defined defaults for all the needed properties that match that default profile, so that when tools fail to activate the profiles we ask for they get something that works-at-all.\n","from":"developer"},{"body":"[~dportabella] can you verify 1.2.2rc2 specifically with it's repository: https://repository.apache.org/content/repositories/orgapachehbase-1142/\n\nI don't know exactly which artifacts are published into https://repository.apache.org/content/groups/staging/ whereas the content of orgapachehbase-1142 are exactly what will become 1.2.2 when the vote passes.","from":"developer"},{"body":"oh, i understand now.\nthanks for the explanation. :)\n\n","from":"developer"},{"body":"{quote}\nDavid Portabella can you verify 1.2.2rc2 specifically with it's repository: https://repository.apache.org/content/repositories/orgapachehbase-1142/\n{quote}\n\nI try as follows:\n{code}\n$ git clone https://github.com/dportabella/spark-examples.git\n# edit build.sbt:\n replace: val hbaseVersion = \"1.2.2rc2\"\n add: resolvers += \"Apache Staging\" at \"https://repository.apache.org/content/repositories/orgapachehbase-1142/\"\n\n$ rm -Rf /david/.ivy2/cache/org.apache.hbase\n$ rm -Rf /david/.m2/repository/org/apache/hbase/\n$ sbt run\n{code}\n\nresult:\n{code}\n[info] Resolving org.apache.hbase#hbase-testing-util;1.2.2rc2 ...\n[warn] \tmodule not found: org.apache.hbase#hbase-testing-util;1.2.2rc2\n[warn] https://repository.apache.org/content/repositories/orgapachehbase-1142/org/apache/hbase/hbase-testing-util/1.2.2rc2/hbase-testing-util-1.2.2rc2.pom\n{code}\n\nit does not find 1.2.2rc2. am i doing sthy wrong?\n\n","from":"developer"},{"body":"Yeah, with that repository URL, the version is just 1.2.2 -- no rc part.","from":"developer"},{"body":"Ok, yes, it works :)\n","from":"developer"},{"body":"Sean, I've filled a bug report at sbt based on your answer. could you please check it?\nhttps://github.com/sbt/sbt/issues/2666","from":"developer"},{"body":"I see that two new versions are published in maven central (1.2.2 and 1.2.3), that's great!\n\none question:\nthis version 1.2.1 (and all previous versions) had this direct dependency's artifactId on ${compat.module}, and we could see this here:\nhttp://mvnrepository.com/artifact/org.apache.hbase/hbase-testing-util/1.2.1\n\nhowever, I don't see this anymore. why?\n\nDoes this mean that you have fixed and replaced the 1.2.1 version (and all the previous ones)?\nIf so, isn't that incorrect? a release version should stay immutable forever.\nCompiling a code which depends on a release dependency should always produce the same result.\n\nOr maybe mvnrepository.com now resolves these variables, and it was not the case when this ticket was open?\n","from":"developer"}],"created":"2016-05-31T16:13:57.000+0000","description":"Looks like we've regressed on HBASE-8488. Have a look at the dependency artifacts list on http://mvnrepository.com/artifact/org.apache.hbase/hbase-testing-util/1.2.1. Notice the direct dependency's artifactId is {{$\\{compat.module\\}}}.","issue_id":"12974271","key":"HBASE-15925","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2016-07-08T15:41:11.000+0000","role":"fixed_distractor","summary":"compat-module maven variable not evaluated"} {"case_id":"13012329","cluster":"DISTRACTOR-HBASE-16841","comments":[{"body":"Upload the first patch for review.\n1. Take the snapshot for mob files at last in {{EnabledTableSnapshotHandler}}.\n2. Refine the code when taking a snapshot for a disabled table.","created":"2016-10-14T17:30:21.236+0000"},{"body":"Is it possible to add a test ?","created":"2016-10-14T17:35:34.452+0000"},{"body":"Thanks [~tedyu@apache.org], sure, let me try.","created":"2016-10-14T17:44:56.955+0000"},{"body":"Upload a new patch to fix the failures in tests about restoring snapshots.\nHi [~tedyu], it is a little difficult to add a test for this case, it needs some delayed flush in some regions during the snapshot which is hard to mimic. Mob doesn't allow a configurable flusher which is designed by purpose to reduce the configurations when using mob.\nI think this is the issue that caused the failures sometimes in unit tests. Hope this patch can fix them all.\nHi [~mbertozzi], do you want to look at this patch? Thanks!","created":"2016-10-25T08:20:23.941+0000"},{"body":"Upload a new patch V3 to remove unused logger from tests.","created":"2016-10-25T08:28:09.925+0000"},{"body":"The findbugs and test failures should not be related with this patch.","created":"2016-10-25T11:02:51.312+0000"},{"body":"Patch v3 is good by me.","created":"2016-10-25T14:10:25.878+0000"},{"body":"Patch LGTM but good to get [~mbertozzi]'s opinion [~jingcheng.du@intel.com] He is out the next few days but will be back sir.","created":"2016-10-25T16:24:40.566+0000"},{"body":"Thanks a lot [~stack]. I'll wait for his response.","created":"2016-10-26T01:58:47.781+0000"},{"body":"Hi [~jingcheng.du@intel.com], about Unit test case, adding a coprocessor for the table, sleep in preFlush() will delay the flush of the memstore with mob cells. will this be easy to reproduce the issue? Thanks.","created":"2016-10-28T20:02:49.662+0000"},{"body":"Thanks a lot [~huaxiang], I'll try to add the tests in this way.","created":"2016-10-31T06:57:34.176+0000"},{"body":"Upload a new patch V4 to add a test to cover this case.\nHi [~mbertozzi], do you want to take a look at this patch? Thanks.","created":"2016-11-10T09:42:10.280+0000"},{"body":"Looks good overall.\nNew test passes.\n{code}\n+ } catch (InterruptedException e1) {\n+ throw new IOException(e1);\n{code}\nConsider throwing InterruptedIOException","created":"2016-11-10T17:48:49.844+0000"},{"body":"Thanks [~tedyu]!\nUpload a new patch V5 according Ted's comments and fix the check style issues.","created":"2016-11-11T02:00:30.875+0000"},{"body":"Hi [~mbertozzi], would you mind taking a look if the patch is good to go? Thanks.","created":"2016-11-22T10:28:54.956+0000"},{"body":"sorry, missed the ping. \nthe snapshot code moved, so the patch does not apply to current master, need a little change.\npatch looks good. only thing is that we can replace the two for loop that find if hcd.isMobEnabled() with MobUtil.hasMobColumns(htd).","created":"2016-12-01T23:23:19.247+0000"},{"body":"Thanks [~mbertozzi].\nUpload a new patch V6 according to the comments. Is this one good to go?","created":"2016-12-02T07:07:51.323+0000"},{"body":"+1","created":"2016-12-02T15:11:09.165+0000"},{"body":"Pushed to the master branch. Thanks [~stack], [~tedyu], [~mbertozzi] and [~huaxiang] for the review.","created":"2016-12-05T08:27:08.113+0000"}],"conversations":[{"body":"Running the following steps will probably lose MOB data when working with snapshots.\n1. Create a mob-enabled table by running create 't1', {NAME => 'f1', IS_MOB => true, MOB_THRESHOLD => 0}.\n2. Put millions of data.\n3. Run {{snapshot 't1','t1_snapshot'}} to take a snapshot for this table t1.\n4. Run {{clone_snapshot 't1_snapshot','t1_cloned'}} to clone this snapshot.\n5. Run {{delete_snapshot 't1_snapshot'}} to delete this snapshot.\n6. Run {{disable 't1'}} and {{delete 't1'}} to delete the table.\n7. Now go to the archive directory of t1, the number of .link directories is different from the number of hfiles which means some data will be lost after the hfile cleaner runs.\n\nThis is because, when taking a snapshot on a enabled mob table, each region flushes itself and takes a snapshot, and the mob snapshot is taken only if the current region is first region of the table. At that time, the flushing of some regions might not be finished, and some mob files are not flushed to disk yet. Eventually some mob files are not recorded in the snapshot manifest.\nTo solve this, we need to take the mob snapshot at last after the snapshots on all the online and offline regions are finished in {{EnabledTableSnapshotHandler}}.","from":"reporter","subject":"Data loss in MOB files after cloning a snapshot and deleting that snapshot"},{"body":"Upload the first patch for review.\n1. Take the snapshot for mob files at last in {{EnabledTableSnapshotHandler}}.\n2. Refine the code when taking a snapshot for a disabled table.","from":"developer"},{"body":"Is it possible to add a test ?","from":"developer"},{"body":"Thanks [~tedyu@apache.org], sure, let me try.","from":"developer"},{"body":"Upload a new patch to fix the failures in tests about restoring snapshots.\nHi [~tedyu], it is a little difficult to add a test for this case, it needs some delayed flush in some regions during the snapshot which is hard to mimic. Mob doesn't allow a configurable flusher which is designed by purpose to reduce the configurations when using mob.\nI think this is the issue that caused the failures sometimes in unit tests. Hope this patch can fix them all.\nHi [~mbertozzi], do you want to look at this patch? Thanks!","from":"developer"},{"body":"Upload a new patch V3 to remove unused logger from tests.","from":"developer"},{"body":"The findbugs and test failures should not be related with this patch.","from":"developer"},{"body":"Patch v3 is good by me.","from":"developer"},{"body":"Patch LGTM but good to get [~mbertozzi]'s opinion [~jingcheng.du@intel.com] He is out the next few days but will be back sir.","from":"developer"},{"body":"Thanks a lot [~stack]. I'll wait for his response.","from":"developer"},{"body":"Hi [~jingcheng.du@intel.com], about Unit test case, adding a coprocessor for the table, sleep in preFlush() will delay the flush of the memstore with mob cells. will this be easy to reproduce the issue? Thanks.","from":"developer"},{"body":"Thanks a lot [~huaxiang], I'll try to add the tests in this way.","from":"developer"},{"body":"Upload a new patch V4 to add a test to cover this case.\nHi [~mbertozzi], do you want to take a look at this patch? Thanks.","from":"developer"},{"body":"Looks good overall.\nNew test passes.\n{code}\n+ } catch (InterruptedException e1) {\n+ throw new IOException(e1);\n{code}\nConsider throwing InterruptedIOException","from":"developer"},{"body":"Thanks [~tedyu]!\nUpload a new patch V5 according Ted's comments and fix the check style issues.","from":"developer"},{"body":"Hi [~mbertozzi], would you mind taking a look if the patch is good to go? Thanks.","from":"developer"},{"body":"sorry, missed the ping. \nthe snapshot code moved, so the patch does not apply to current master, need a little change.\npatch looks good. only thing is that we can replace the two for loop that find if hcd.isMobEnabled() with MobUtil.hasMobColumns(htd).","from":"developer"},{"body":"Thanks [~mbertozzi].\nUpload a new patch V6 according to the comments. Is this one good to go?","from":"developer"},{"body":"+1","from":"developer"},{"body":"Pushed to the master branch. Thanks [~stack], [~tedyu], [~mbertozzi] and [~huaxiang] for the review.","from":"developer"}],"created":"2016-10-14T11:32:00.000+0000","description":"Running the following steps will probably lose MOB data when working with snapshots.\n1. Create a mob-enabled table by running create 't1', {NAME => 'f1', IS_MOB => true, MOB_THRESHOLD => 0}.\n2. Put millions of data.\n3. Run {{snapshot 't1','t1_snapshot'}} to take a snapshot for this table t1.\n4. Run {{clone_snapshot 't1_snapshot','t1_cloned'}} to clone this snapshot.\n5. Run {{delete_snapshot 't1_snapshot'}} to delete this snapshot.\n6. Run {{disable 't1'}} and {{delete 't1'}} to delete the table.\n7. Now go to the archive directory of t1, the number of .link directories is different from the number of hfiles which means some data will be lost after the hfile cleaner runs.\n\nThis is because, when taking a snapshot on a enabled mob table, each region flushes itself and takes a snapshot, and the mob snapshot is taken only if the current region is first region of the table. At that time, the flushing of some regions might not be finished, and some mob files are not flushed to disk yet. Eventually some mob files are not recorded in the snapshot manifest.\nTo solve this, we need to take the mob snapshot at last after the snapshots on all the online and offline regions are finished in {{EnabledTableSnapshotHandler}}.","issue_id":"13012329","key":"HBASE-16841","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2016-12-05T08:27:31.000+0000","role":"fixed_distractor","summary":"Data loss in MOB files after cloning a snapshot and deleting that snapshot"} {"case_id":"13047229","cluster":"DISTRACTOR-HBASE-17713","comments":[{"body":"If I understand the JSON specification correctly, a String literal as above is valid JSON. It does not need to be specified in object notation. Have a look at the following references\n\n* http://stackoverflow.com/questions/13318420/is-a-single-string-value-considered-valid-json\n* http://www.freeformatter.com/json-validator.html\n\nThe appropriate test in the codebase of HBase (TestVersionResource#doTestGetStorageClusterVersionJSON) shows that JSON is returned as the content type.","created":"2017-03-03T15:03:52.328+0000"},{"body":"[~janh], thanks. You are right, the response is valid JSON. \nHowever, this still looks like a strange return value.","created":"2017-03-13T02:50:32.570+0000"},{"body":"I also found this problem when using rest api.I can fix it. thanks :)","created":"2017-09-05T13:06:08.737+0000"},{"body":"After this patch\n\nXML:\n{code:xml}\n\n\n{code}\n\nJSON:\n{code:java}\n{\"Version\": \"2.0.0-alpha3-SNAPSHOT\"}\n{code}\n\n","created":"2017-09-05T13:16:20.303+0000"},{"body":"lgtm, pending QA","created":"2017-09-05T15:47:10.504+0000"},{"body":"Please fix test failure:\n{code}\ntestParsingClusterVersion(org.apache.hadoop.hbase.rest.client.TestXmlParsing) Time elapsed: 0.023 sec <<< FAILURE!\njava.lang.AssertionError: expected:<2.0.0> but was:\n\tat org.apache.hadoop.hbase.rest.client.TestXmlParsing.testParsingClusterVersion(TestXmlParsing.java:52)\n{code}","created":"2017-09-06T03:07:30.761+0000"},{"body":"The v1 patch fixes the UT. pending QA.Thanks","created":"2017-09-06T04:05:15.821+0000"},{"body":"v1 patch looks good.Upload patch for branch-1. :)","created":"2017-09-06T04:57:59.329+0000"},{"body":"Thanks for the patch, Guangxu","created":"2017-09-06T14:37:13.200+0000"},{"body":"Thanks [~yuzhihong@gmail.com] for review.:)","created":"2017-09-06T16:03:21.254+0000"}],"conversations":[{"body":"Hbase REST API, this interface `get 'version/cluster'`, when I use the header `Accept: application/json`, the response is not JSON but plain text.\n{code:none}\n curl -X GET \\\n -H \"Accept: application/json\" \\\n \"http://localhost:8888/version/cluster\"\n\n \"1.2.2\"\n{code}\n\nBut when I use `Accept: text/xml`, the response is correct XML.\n\n{code:none}\n curl -X GET \\\n -H \"Accept: text/xml\" \\\n \"http://localhost:8888/version/cluster\"\n\n 1.2.2\n{code}","from":"reporter","subject":"the interface '/version/cluster' with header 'Accept: application/json' return is not JSON but plain text"},{"body":"If I understand the JSON specification correctly, a String literal as above is valid JSON. It does not need to be specified in object notation. Have a look at the following references\n\n* http://stackoverflow.com/questions/13318420/is-a-single-string-value-considered-valid-json\n* http://www.freeformatter.com/json-validator.html\n\nThe appropriate test in the codebase of HBase (TestVersionResource#doTestGetStorageClusterVersionJSON) shows that JSON is returned as the content type.","from":"developer"},{"body":"[~janh], thanks. You are right, the response is valid JSON. \nHowever, this still looks like a strange return value.","from":"developer"},{"body":"I also found this problem when using rest api.I can fix it. thanks :)","from":"developer"},{"body":"After this patch\n\nXML:\n{code:xml}\n\n\n{code}\n\nJSON:\n{code:java}\n{\"Version\": \"2.0.0-alpha3-SNAPSHOT\"}\n{code}\n\n","from":"developer"},{"body":"lgtm, pending QA","from":"developer"},{"body":"Please fix test failure:\n{code}\ntestParsingClusterVersion(org.apache.hadoop.hbase.rest.client.TestXmlParsing) Time elapsed: 0.023 sec <<< FAILURE!\njava.lang.AssertionError: expected:<2.0.0> but was:\n\tat org.apache.hadoop.hbase.rest.client.TestXmlParsing.testParsingClusterVersion(TestXmlParsing.java:52)\n{code}","from":"developer"},{"body":"The v1 patch fixes the UT. pending QA.Thanks","from":"developer"},{"body":"v1 patch looks good.Upload patch for branch-1. :)","from":"developer"},{"body":"Thanks for the patch, Guangxu","from":"developer"},{"body":"Thanks [~yuzhihong@gmail.com] for review.:)","from":"developer"}],"created":"2017-03-01T08:49:27.000+0000","description":"Hbase REST API, this interface `get 'version/cluster'`, when I use the header `Accept: application/json`, the response is not JSON but plain text.\n{code:none}\n curl -X GET \\\n -H \"Accept: application/json\" \\\n \"http://localhost:8888/version/cluster\"\n\n \"1.2.2\"\n{code}\n\nBut when I use `Accept: text/xml`, the response is correct XML.\n\n{code:none}\n curl -X GET \\\n -H \"Accept: text/xml\" \\\n \"http://localhost:8888/version/cluster\"\n\n 1.2.2\n{code}","issue_id":"13047229","key":"HBASE-17713","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2017-09-06T14:37:13.000+0000","role":"fixed_distractor","summary":"the interface '/version/cluster' with header 'Accept: application/json' return is not JSON but plain text"} {"case_id":"13071045","cluster":"DISTRACTOR-HBASE-18030","comments":[{"body":"Ya some corruption has happened. So are you able to consistently reproduce this case when you do the similar increments on your table? I think this is a critical bug. any tests possible?","created":"2017-05-11T04:57:23.594+0000"},{"body":"I copied that table to another table and it worked for few days. But yesterday in the new table, again facing the same issue.\nLet me know what kind of tests you are asking, I will do that.","created":"2017-05-11T05:26:37.122+0000"},{"body":"Test in the sense if you have identified certain pattern and that can be reproduced with a test case. ","created":"2017-05-11T06:07:40.165+0000"},{"body":"My table is having two column family (weekly, daily) and using TTL.\nFor me, these are very basic features for tables with increments. The weekly column family has small files and I don't see them compacting (total 5 HFiles). This is not the case with daily column family, which has just 1 HFile.\n\nhbase(main):016:0> get 'table-name', 'rowkey', 'daily:a'\n0 row(s) in 0.0100 seconds\n\nhbase(main):017:0> get 'table-name', 'rowkey', 'weekly:a'\nERROR: java.io.IOException: Invalid currTagsLen -32723. Block offset: 0, block length: 69222, position: 3447 (without header).\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2172)\n\nA cron runs at my application service (3 of them) which flushes all the increments it accumulated, one by one to hbase.\n\nNothing suspicious about the whole flow or any patterns I have observed. \nI am trying to remove TTL and moving to just 1 column family to see if it is fixing the issue.","created":"2017-05-11T06:48:07.568+0000"},{"body":"CAn you try major_compacting the files in the weekly column family (that has 5 Hfiles). If major compaction itself fails then again it is the same issue due to corrupted block.","created":"2017-05-11T07:08:17.684+0000"},{"body":"Reading your comments again, when you say you have TTL you mean the Column family TTL or do you have per Cell TTL? Seems like you have per Cell TTL in your increments?","created":"2017-05-11T07:22:22.106+0000"},{"body":"hbase(main):022:0> major_compact 'tablename'\n0 row(s) in 0.1280 seconds\n\nCompaction is done.\n\nafter every error line (Invalid currTagsLen -32712) this follows:\n2017-05-11 07:31:00,745 INFO [hdps01.labs.ops.use1a.i.riva.co,16020,1493962926376_ChoreService_2] regionserver.HRegionServer: hdps01.labs.ops.use1a.i.riva.co,16020,1493962926376-MemstoreFlusherChore requesting flush for region ,,1493782145821.db8a11a78836a33cdd045e5c6fcf4b6c. after a delay of 10433\n\nI can still see 5 HFiles for weekly column family","created":"2017-05-11T07:34:46.104+0000"},{"body":"I have per cell TTL (i.e. I am setting TTL per Increment)","created":"2017-05-11T09:42:26.018+0000"},{"body":"In that case you are using Tag feature of hbase. Internally per cell TTL is implemented as Tags. Ok will now look into the code to understand what could be possibly wrong and come back here.\nIf possible try to replicate your use case with a simple test case with per Cell TTL. ","created":"2017-05-11T09:47:54.152+0000"},{"body":"BTW, which version of Hbase are you using? Is it 1.1.2 only as specified in 'Affected Versions'","created":"2017-05-11T09:58:31.785+0000"},{"body":"I am using HBase 1.1.2","created":"2017-05-11T10:17:13.978+0000"},{"body":"Checked the code. Nothing obvious in this area. But I suspect some other issues when every increment has a TTL. But that should not throw a negative tag offset. \nbq.2017-05-11 07:31:00,745 INFO [hdps01.labs.ops.use1a.i.riva.co,16020,1493962926376_ChoreService_2] regionserver.HRegionServer: hdps01.labs.ops.use1a.i.riva.co,16020,1493962926376-MemstoreFlusherChore requesting flush for region ,,1493782145821.db8a11a78836a33cdd045e5c6fcf4b6c. after a delay of 10433\nThat is because you still have data in your memstore may be. You can do a manual flush on the table to see if this goes off. \n","created":"2017-05-12T05:12:45.188+0000"},{"body":"I believe I know what is the issue. Will try giving a patch soon.","created":"2017-05-12T08:28:47.841+0000"},{"body":"Pls try once. I will have a detailed look and see any other missing area. Will add a UT also. This is just an initial patch","created":"2017-05-12T08:36:35.842+0000"},{"body":"bq.But I suspect some other issues when every increment has a TTL.\nThis is what I suspected. Every increment with a TTL will have multiple TTL tags in it. So we should remove older ones and only add the new one for that increment. \nBut that should not cause the tag offset negative thing if the overall tags len is anyway inside the KV..","created":"2017-05-12T09:26:15.884+0000"},{"body":"Yes. The op performed here is Increments. Each of the increment Mutation will have a TTL associated with it. The per cell TTL info will be written to Cell as tag. So say for the 1st Cell we have one TTL tag in it. WHen we increment this, we will write a new Cell and in that we will copy all tags from existing old cell and as the new Mutation having a TTL, we will add a new Tag also to it and finally the new Cell will have 2 tags. (Duplicated TTL tags).. This will grow with every incr op. Within a Cell , the total tags length is written in 2 bytes. So once this duplicated tags increase to a level and total tags bytes exceed Short.MAX_Value, we will write the tags len as -ve. ","created":"2017-05-12T09:52:50.417+0000"},{"body":"Duplicated tags I had found the problem already for sure. But did not expect the tags to grow in length because in KV construction we already had a check to see if tags is greater than Tag.MAXLEN. So here it is Short overflow, so even that check would have been bypassed.\nAlso we need fix in Storefile etc also if this negative overflow happens. \n","created":"2017-05-12T10:04:56.886+0000"},{"body":"Fixing the problem with duplicated TTL tags. \nThe visibility and ACL tags already handle this possible duplicated tags issue (With increments/Append)","created":"2017-06-05T17:15:07.657+0000"},{"body":"lgtm","created":"2017-06-05T17:26:12.860+0000"},{"body":"bq.The visibility and ACL tags already handle this possible duplicated tags issue (With increments/Append)\nIf already handled it is fine. \nJust asking - Should we unify the way the tag removal happens for all the cases (and future cases that are possible)?\n+1.","created":"2017-06-06T04:02:22.369+0000"},{"body":"You mean remove duplicate Tag types? May be not. As such we dont have any restriction like there should not be duplicated tags in a Cell. So it depends on the Tag type.. I tried to make it like unique Tag types in a Cell way. But then gave up with above thinking. As of now this way is ok. When we allow custom tags also, lets think.\nWill commit this.\nThis has to be committed to branch-1, branch-1.3, branch-1.2 at least.","created":"2017-06-06T04:38:47.845+0000"}],"conversations":[{"body":"2017-04-29 14:24:14,135 ERROR [B.fifo.QRpcServer.handler=49,queue=1,port=16020] ipc.RpcServer: Unexpected throwable object java.lang.IllegalStateException: Invalid currTagsLen -32712. Block offset: 3707853, block length: 72841, position: 0 (without header). at org.apache.hadoop.hbase.io.hfile.HFileReaderV3$ScannerV3.checkTagsLen(HFileReaderV3.java:226)\n\nI am not not using any hbase tags feature.\nThe Increment operation from the application side is triggering this error. The same is happening when scanner is run on this table. It feels that one or more particular HFile block is corrupt (with negative tagLength).\n\nhbase(main):007:0> scan 'table-name', {LIMIT=>1,STARTROW=>'ad:event_count:a'}\nReturning the result\n\n\nhbase(main):008:0> scan 'table-name', {LIMIT=>1,STARTROW=>'ad:event_count:b'}\nROW COLUMN+CELL \nERROR: java.io.IOException: java.lang.IllegalStateException: Invalid currTagsLen -32701. Block offset: 272031, block length: 72441, position: 32487 (without header).\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl.handleException(HRegion.java:5607)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl.(HRegion.java:5579)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.instantiateRegionScanner(HRegion.java:2627)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.getScanner(HRegion.java:2613)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.getScanner(HRegion.java:2595)\n\tat org.apache.hadoop.hbase.regionserver.RSRpcServices.scan(RSRpcServices.java:2282)\n\tat org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:32295)","from":"reporter","subject":"Per Cell TTL tags may get duplicated with increments/Append causing tags length overflow"},{"body":"Ya some corruption has happened. So are you able to consistently reproduce this case when you do the similar increments on your table? I think this is a critical bug. any tests possible?","from":"developer"},{"body":"I copied that table to another table and it worked for few days. But yesterday in the new table, again facing the same issue.\nLet me know what kind of tests you are asking, I will do that.","from":"developer"},{"body":"Test in the sense if you have identified certain pattern and that can be reproduced with a test case. ","from":"developer"},{"body":"My table is having two column family (weekly, daily) and using TTL.\nFor me, these are very basic features for tables with increments. The weekly column family has small files and I don't see them compacting (total 5 HFiles). This is not the case with daily column family, which has just 1 HFile.\n\nhbase(main):016:0> get 'table-name', 'rowkey', 'daily:a'\n0 row(s) in 0.0100 seconds\n\nhbase(main):017:0> get 'table-name', 'rowkey', 'weekly:a'\nERROR: java.io.IOException: Invalid currTagsLen -32723. Block offset: 0, block length: 69222, position: 3447 (without header).\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2172)\n\nA cron runs at my application service (3 of them) which flushes all the increments it accumulated, one by one to hbase.\n\nNothing suspicious about the whole flow or any patterns I have observed. \nI am trying to remove TTL and moving to just 1 column family to see if it is fixing the issue.","from":"developer"},{"body":"CAn you try major_compacting the files in the weekly column family (that has 5 Hfiles). If major compaction itself fails then again it is the same issue due to corrupted block.","from":"developer"},{"body":"Reading your comments again, when you say you have TTL you mean the Column family TTL or do you have per Cell TTL? Seems like you have per Cell TTL in your increments?","from":"developer"},{"body":"hbase(main):022:0> major_compact 'tablename'\n0 row(s) in 0.1280 seconds\n\nCompaction is done.\n\nafter every error line (Invalid currTagsLen -32712) this follows:\n2017-05-11 07:31:00,745 INFO [hdps01.labs.ops.use1a.i.riva.co,16020,1493962926376_ChoreService_2] regionserver.HRegionServer: hdps01.labs.ops.use1a.i.riva.co,16020,1493962926376-MemstoreFlusherChore requesting flush for region ,,1493782145821.db8a11a78836a33cdd045e5c6fcf4b6c. after a delay of 10433\n\nI can still see 5 HFiles for weekly column family","from":"developer"},{"body":"I have per cell TTL (i.e. I am setting TTL per Increment)","from":"developer"},{"body":"In that case you are using Tag feature of hbase. Internally per cell TTL is implemented as Tags. Ok will now look into the code to understand what could be possibly wrong and come back here.\nIf possible try to replicate your use case with a simple test case with per Cell TTL. ","from":"developer"},{"body":"BTW, which version of Hbase are you using? Is it 1.1.2 only as specified in 'Affected Versions'","from":"developer"},{"body":"I am using HBase 1.1.2","from":"developer"},{"body":"Checked the code. Nothing obvious in this area. But I suspect some other issues when every increment has a TTL. But that should not throw a negative tag offset. \nbq.2017-05-11 07:31:00,745 INFO [hdps01.labs.ops.use1a.i.riva.co,16020,1493962926376_ChoreService_2] regionserver.HRegionServer: hdps01.labs.ops.use1a.i.riva.co,16020,1493962926376-MemstoreFlusherChore requesting flush for region ,,1493782145821.db8a11a78836a33cdd045e5c6fcf4b6c. after a delay of 10433\nThat is because you still have data in your memstore may be. You can do a manual flush on the table to see if this goes off. \n","from":"developer"},{"body":"I believe I know what is the issue. Will try giving a patch soon.","from":"developer"},{"body":"Pls try once. I will have a detailed look and see any other missing area. Will add a UT also. This is just an initial patch","from":"developer"},{"body":"bq.But I suspect some other issues when every increment has a TTL.\nThis is what I suspected. Every increment with a TTL will have multiple TTL tags in it. So we should remove older ones and only add the new one for that increment. \nBut that should not cause the tag offset negative thing if the overall tags len is anyway inside the KV..","from":"developer"},{"body":"Yes. The op performed here is Increments. Each of the increment Mutation will have a TTL associated with it. The per cell TTL info will be written to Cell as tag. So say for the 1st Cell we have one TTL tag in it. WHen we increment this, we will write a new Cell and in that we will copy all tags from existing old cell and as the new Mutation having a TTL, we will add a new Tag also to it and finally the new Cell will have 2 tags. (Duplicated TTL tags).. This will grow with every incr op. Within a Cell , the total tags length is written in 2 bytes. So once this duplicated tags increase to a level and total tags bytes exceed Short.MAX_Value, we will write the tags len as -ve. ","from":"developer"},{"body":"Duplicated tags I had found the problem already for sure. But did not expect the tags to grow in length because in KV construction we already had a check to see if tags is greater than Tag.MAXLEN. So here it is Short overflow, so even that check would have been bypassed.\nAlso we need fix in Storefile etc also if this negative overflow happens. \n","from":"developer"},{"body":"Fixing the problem with duplicated TTL tags. \nThe visibility and ACL tags already handle this possible duplicated tags issue (With increments/Append)","from":"developer"},{"body":"lgtm","from":"developer"},{"body":"bq.The visibility and ACL tags already handle this possible duplicated tags issue (With increments/Append)\nIf already handled it is fine. \nJust asking - Should we unify the way the tag removal happens for all the cases (and future cases that are possible)?\n+1.","from":"developer"},{"body":"You mean remove duplicate Tag types? May be not. As such we dont have any restriction like there should not be duplicated tags in a Cell. So it depends on the Tag type.. I tried to make it like unique Tag types in a Cell way. But then gave up with above thinking. As of now this way is ok. When we allow custom tags also, lets think.\nWill commit this.\nThis has to be committed to branch-1, branch-1.3, branch-1.2 at least.","from":"developer"}],"created":"2017-05-11T04:31:01.000+0000","description":"2017-04-29 14:24:14,135 ERROR [B.fifo.QRpcServer.handler=49,queue=1,port=16020] ipc.RpcServer: Unexpected throwable object java.lang.IllegalStateException: Invalid currTagsLen -32712. Block offset: 3707853, block length: 72841, position: 0 (without header). at org.apache.hadoop.hbase.io.hfile.HFileReaderV3$ScannerV3.checkTagsLen(HFileReaderV3.java:226)\n\nI am not not using any hbase tags feature.\nThe Increment operation from the application side is triggering this error. The same is happening when scanner is run on this table. It feels that one or more particular HFile block is corrupt (with negative tagLength).\n\nhbase(main):007:0> scan 'table-name', {LIMIT=>1,STARTROW=>'ad:event_count:a'}\nReturning the result\n\n\nhbase(main):008:0> scan 'table-name', {LIMIT=>1,STARTROW=>'ad:event_count:b'}\nROW COLUMN+CELL \nERROR: java.io.IOException: java.lang.IllegalStateException: Invalid currTagsLen -32701. Block offset: 272031, block length: 72441, position: 32487 (without header).\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl.handleException(HRegion.java:5607)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl.(HRegion.java:5579)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.instantiateRegionScanner(HRegion.java:2627)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.getScanner(HRegion.java:2613)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.getScanner(HRegion.java:2595)\n\tat org.apache.hadoop.hbase.regionserver.RSRpcServices.scan(RSRpcServices.java:2282)\n\tat org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:32295)","issue_id":"13071045","key":"HBASE-18030","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2017-06-06T06:59:12.000+0000","role":"fixed_distractor","summary":"Per Cell TTL tags may get duplicated with increments/Append causing tags length overflow"} {"case_id":"12435494","cluster":"DISTRACTOR-HBASE-1831","comments":[{"body":"Filters should not be run on the client-side at all.\n\nServer needs to be able to tell whether Scan should continue to the next region or not.","created":"2009-09-23T04:35:45.597+0000"},{"body":"Chatted with Jon.\n\n1. Need to flag client to STOP. HTable#next internally uses the batch version of next (so can do prefetching of rows). An empty list of Results is always sent -- never null. We'll add passing null as the flag that filter is done; do not move to next region (Client has code to handle null list so if an old hbase version connects, it won't break; it'll just not do the STOP properly).\n2. There is a non-batch next in the ipc interface. I was thinking of deprecating it and moving internals to use the batch interface only, but these internal uses of scanners do not carry filters so will just leave them for 0.20.1.\n3. Filters carry state. How do we get the state across region transitions? Again chatting with Jon, will do the following. If a Scanner has a filter, and we got back a non-empty list, its time to move to the next region. Just before we move to the next region, we'll make another call to the old server -- Scanner.getFilter -- whose result is the deserialized filter. The deserialized filter will be passed then to the next region. In this manner filters will be able to carry their state forward. Downside is extra RPC call IF scanning with a filter.","created":"2009-09-23T05:04:45.208+0000"},{"body":"On 3., above, carrying stateful filters across regions, it can't be done easily in 0.20.x because we can get a NotServingRegionException at any time. Also, if nothing goes wrong and we just exhaust a scanner on a particular region, the last thing done over in RegionServer is cleanup of scanner so a subsequent getFilter call would have nothing to pick up on once it'd arrived at the RegionServer. Statefulness has to be done client-side; filters need to allow specifying a client-component.","created":"2009-09-23T19:32:52.525+0000"},{"body":"@jgray You say, \"Filters should not be run on the client-side at all.\"\n\nI don't know how else we can do stateful scanners that ride over regions.","created":"2009-09-23T19:42:45.485+0000"},{"body":"So, chatting w/ Ryan on this topic, he suggests that support for stateful filters is going to be a bear -- there's splitting and then what if regionserver crashes, etc. At the moment we figure a filter wants to stop the scan by testing the last regions endrow against the filter run client-side. Would it be enough my removing this and just adding the flag above which allows filters server-side to stop the scan (by passing back null result out of batch next)? There'd be no getFilter to pick up a filters state and pass it from one region to the next.\n\nThere is going to be a problem though when fellas want to do filters that return 20 row results only or that want to have a filter skip 1000 rows; to do this, we'd something to run client-side.","created":"2009-09-23T19:57:24.634+0000"},{"body":"Here is a bit of a start. It adds javadoc to the ipc interfaces about new meaning of null. It then adds a new flag to the client-side nextScanner method. It renames the method that checks for the end row in Scan.... Now to work on server-side.","created":"2009-09-23T20:44:15.982+0000"},{"body":"I agree that this is not going to be easy. Still think that filters should not be run client-side (even if they are, it will require additional/new information to be sent from client to server).\n\nRunning an offset of 1000 rows client-side completely negates the value in using filters in the first place... to not have to send back all that data. That's the reason client-side filters don't really work, they require you to send back otherwise filtered-out data.\n\nWhat I would like to see as a long-term solution to this:\n\n1st: Add ability for server to say STOP and not go to next region\n2nd: Correctness for stateful scanners under non-split, non-failure scenarios (less correctness / fail-fast if encountering issues)\n3rd: Correctness/robustness for stateful scanners under splits and failures","created":"2009-09-24T00:06:00.383+0000"},{"body":"v2... complete but for test. Working on that now.","created":"2009-09-25T00:32:04.099+0000"},{"body":"Still not done","created":"2009-09-27T22:54:02.022+0000"},{"body":"Adds two ugly tests. One under filter that puts up three regions and then checks at the Region level that filters are doing right thing (Why can't i instantiate an HRegionServer and test from its interface -- its currently way too hard to put one of these up... requires there be a master.. .it shouldn't). Other test if ugly from client side. Splits table then makes sure RowFilter is returning right results around the row boundary. I can assert counts but I can't assert that only a subset of regions are being accessed with asserts. To do the latter, I added logging and it required eyeballing but you can see in the logs that yes we do not go to next region if filter says we're done.\n\nA few tests seem to be failing.... Looking into it.","created":"2009-10-02T04:22:09.804+0000"},{"body":"This version passes all tests.","created":"2009-10-02T05:07:25.096+0000"},{"body":"Needs review.","created":"2009-10-03T05:18:15.954+0000"},{"body":"Testing this now. There were two rejects against latest SVN 0.20, one in HTable, one in HConnectionManager. I fixed them up and attached the result as -v6","created":"2009-10-05T23:36:06.559+0000"},{"body":"+1 all tests pass here","created":"2009-10-06T00:30:23.945+0000"},{"body":"Thanks for review Andrew. Committed branch and trunk.","created":"2009-10-06T03:26:26.189+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T13:01:47.911+0000"}],"conversations":[{"body":"Right now, a client replays part of the Filter locally by calling filterRowKey() and filterAllRemaining() to determine whether it should continue to the next region.\n\nA number of new filters rely on filterKeyValue() and other calls to alter state. It's also a false assumption that all rows/keys affecting a filter returning true for FAR will be seen client-side (what about those that failed the filter).\n\nThis issue is about dealing with Filters properly from the client-side.","from":"reporter","subject":"Scanning API must be reworked to allow for fully functional Filters client-side"},{"body":"Filters should not be run on the client-side at all.\n\nServer needs to be able to tell whether Scan should continue to the next region or not.","from":"developer"},{"body":"Chatted with Jon.\n\n1. Need to flag client to STOP. HTable#next internally uses the batch version of next (so can do prefetching of rows). An empty list of Results is always sent -- never null. We'll add passing null as the flag that filter is done; do not move to next region (Client has code to handle null list so if an old hbase version connects, it won't break; it'll just not do the STOP properly).\n2. There is a non-batch next in the ipc interface. I was thinking of deprecating it and moving internals to use the batch interface only, but these internal uses of scanners do not carry filters so will just leave them for 0.20.1.\n3. Filters carry state. How do we get the state across region transitions? Again chatting with Jon, will do the following. If a Scanner has a filter, and we got back a non-empty list, its time to move to the next region. Just before we move to the next region, we'll make another call to the old server -- Scanner.getFilter -- whose result is the deserialized filter. The deserialized filter will be passed then to the next region. In this manner filters will be able to carry their state forward. Downside is extra RPC call IF scanning with a filter.","from":"developer"},{"body":"On 3., above, carrying stateful filters across regions, it can't be done easily in 0.20.x because we can get a NotServingRegionException at any time. Also, if nothing goes wrong and we just exhaust a scanner on a particular region, the last thing done over in RegionServer is cleanup of scanner so a subsequent getFilter call would have nothing to pick up on once it'd arrived at the RegionServer. Statefulness has to be done client-side; filters need to allow specifying a client-component.","from":"developer"},{"body":"@jgray You say, \"Filters should not be run on the client-side at all.\"\n\nI don't know how else we can do stateful scanners that ride over regions.","from":"developer"},{"body":"So, chatting w/ Ryan on this topic, he suggests that support for stateful filters is going to be a bear -- there's splitting and then what if regionserver crashes, etc. At the moment we figure a filter wants to stop the scan by testing the last regions endrow against the filter run client-side. Would it be enough my removing this and just adding the flag above which allows filters server-side to stop the scan (by passing back null result out of batch next)? There'd be no getFilter to pick up a filters state and pass it from one region to the next.\n\nThere is going to be a problem though when fellas want to do filters that return 20 row results only or that want to have a filter skip 1000 rows; to do this, we'd something to run client-side.","from":"developer"},{"body":"Here is a bit of a start. It adds javadoc to the ipc interfaces about new meaning of null. It then adds a new flag to the client-side nextScanner method. It renames the method that checks for the end row in Scan.... Now to work on server-side.","from":"developer"},{"body":"I agree that this is not going to be easy. Still think that filters should not be run client-side (even if they are, it will require additional/new information to be sent from client to server).\n\nRunning an offset of 1000 rows client-side completely negates the value in using filters in the first place... to not have to send back all that data. That's the reason client-side filters don't really work, they require you to send back otherwise filtered-out data.\n\nWhat I would like to see as a long-term solution to this:\n\n1st: Add ability for server to say STOP and not go to next region\n2nd: Correctness for stateful scanners under non-split, non-failure scenarios (less correctness / fail-fast if encountering issues)\n3rd: Correctness/robustness for stateful scanners under splits and failures","from":"developer"},{"body":"v2... complete but for test. Working on that now.","from":"developer"},{"body":"Still not done","from":"developer"},{"body":"Adds two ugly tests. One under filter that puts up three regions and then checks at the Region level that filters are doing right thing (Why can't i instantiate an HRegionServer and test from its interface -- its currently way too hard to put one of these up... requires there be a master.. .it shouldn't). Other test if ugly from client side. Splits table then makes sure RowFilter is returning right results around the row boundary. I can assert counts but I can't assert that only a subset of regions are being accessed with asserts. To do the latter, I added logging and it required eyeballing but you can see in the logs that yes we do not go to next region if filter says we're done.\n\nA few tests seem to be failing.... Looking into it.","from":"developer"},{"body":"This version passes all tests.","from":"developer"},{"body":"Needs review.","from":"developer"},{"body":"Testing this now. There were two rejects against latest SVN 0.20, one in HTable, one in HConnectionManager. I fixed them up and attached the result as -v6","from":"developer"},{"body":"+1 all tests pass here","from":"developer"},{"body":"Thanks for review Andrew. Committed branch and trunk.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2009-09-12T01:06:29.000+0000","description":"Right now, a client replays part of the Filter locally by calling filterRowKey() and filterAllRemaining() to determine whether it should continue to the next region.\n\nA number of new filters rely on filterKeyValue() and other calls to alter state. It's also a false assumption that all rows/keys affecting a filter returning true for FAR will be seen client-side (what about those that failed the filter).\n\nThis issue is about dealing with Filters properly from the client-side.","issue_id":"12435494","key":"HBASE-1831","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2009-10-06T03:26:26.000+0000","role":"fixed_distractor","summary":"Scanning API must be reworked to allow for fully functional Filters client-side"} {"case_id":"13090752","cluster":"DISTRACTOR-HBASE-18471","comments":[{"body":"Another difference in {{scanner}} behavior, after the test code is run\n{code}\nhbase(main):023:0> scan 'test_dml'\nROW COLUMN+CELL \n myRow column=myFamily:, timestamp=4, value=myValue \n myRow column=myFamily:myQualifier, timestamp=1, value=myValue \n1 row(s) in 0.0330 seconds\n\nhbase(main):024:0> scan 'test_dml', {VERSIONS => 10}\nROW COLUMN+CELL \n myRow column=myFamily:, timestamp=4, value=myValue \n myRow column=myFamily:, timestamp=3, value=myValue \n1 row(s) in 0.0290 seconds\n{code}","created":"2017-08-01T01:09:41.754+0000"},{"body":"We add a cell without qualifier for deleting all columns of the specified family.\n{code}\n public Delete addFamily(final byte [] family, final long timestamp) {\n if (timestamp < 0) {\n throw new IllegalArgumentException(\"Timestamp cannot be negative. ts=\" + timestamp);\n }\n List list = familyMap.get(family);\n if(list == null) {\n list = new ArrayList<>(1);\n } else if(!list.isEmpty()) {\n list.clear();\n }\n KeyValue kv = new KeyValue(row, family, null, timestamp, KeyValue.Type.DeleteFamily);\n list.add(kv);\n familyMap.put(family, list);\n return this;\n }\n{code}\n\nThe *hint cell* created by CellUtil.createLastOnRowCol force the KVScanner to skip the remaining cells which have the same family/qualifier with *hint cell*\n{code}\n public Cell getKeyForNextColumn(Cell cell) {\n ColumnCount nextColumn = columns.getColumnHint();\n if (nextColumn == null) {\n return CellUtil.createLastOnRowCol(cell);\n } else {\n return CellUtil.createFirstOnRowCol(cell, nextColumn.getBuffer(), nextColumn.getOffset(),\n nextColumn.getLength());\n }\n }\n{code}\n\n","created":"2017-08-02T13:15:02.149+0000"},{"body":"Ping for reviews. This bug makes our hbase lose data. We should resolve it before next release.","created":"2017-08-14T06:26:39.828+0000"},{"body":"lgtm\n\nCan you post patch for branch-1.2 ?","created":"2017-08-14T07:20:29.706+0000"},{"body":"Will post patch for branch-1 and branch-1.2\n\n{quote}\n+ /**\n+ * @param cell current cell\n+ * @return Null if we have no idea about the \"hint\" of next key\n+ */\n{quote}\nThe above comment is wrong...I will remove it.","created":"2017-08-14T08:16:20.365+0000"},{"body":"Nice debug. Ya this is a serious issue but did not unearth till now !\n{code}\n if (type != Type.Minimum.getCode()) {\n3148\t type = KeyValue.Type.values()[KeyValue.Type.codeToType(type).ordinal() - 1].getCode();\n3149\t } else if (ts != HConstants.OLDEST_TIMESTAMP) {\n3150\t ts = ts - 1;\n3151\t type = Type.Maximum.getCode();\n3152\t }\n{code}\nSo as in the test, we will call with a Put type cell. So the type is not min and 1st if is through. We will use the type as one below the PUT type which is again MINIMUM. But then we will NOT change the ts. That is the diff what we make (?)\nIn our compare logic, there is special treatment for min and max I guess. We wont even see the TS then (?) I may be wrong. If so am not sure how we will solve issue here.","created":"2017-08-14T10:58:28.995+0000"},{"body":"bq. In our compare logic, there is special treatment for min and max I guess. We wont even see the TS\nDo you mean this special treatment?\n{code:title=CellComparator#compareWithoutRow}\n if (lFamLength + lQualLength == 0\n && left.getTypeByte() == Type.Minimum.getCode()) {\n // left is \"bigger\", i.e. it appears later in the sorted order\n return 1;\n }\n if (rFamLength + rQualLength == 0\n && right.getTypeByte() == Type.Minimum.getCode()) {\n return -1;\n }\n{code}\nThere is no special treatment for max. Also i don' find any case that the cell has empty column and PUT type. (i mean this patch won't generate the cell which has empty column and MINIMUM type.)","created":"2017-08-14T13:56:52.253+0000"},{"body":"v1 patch\n# remove incorrect comment\n# add more tests (Ted's suggestion)","created":"2017-08-14T14:17:21.523+0000"},{"body":"Will check this tomorrow. ","created":"2017-08-14T18:02:50.056+0000"},{"body":"Reviewing this. So reading the patch and the description. After a DeleteFamily here a Put with empty qualifier is being added. What happens if a put with a qualifier which is lexographically smaller than the existing QUALIFIER happens after the DeleteFamily?","created":"2017-08-16T09:53:36.187+0000"},{"body":"bq. After a DeleteFamily here a Put\nbq. What happens if a put with a qualifier which is lexographically smaller than the existing QUALIFIER happens after the DeleteFamily?\nDo you mean all puts happen *after* the DeleteFamily? If yes, the DeleteFamily cell doesn't delete any PUT cells. Please correct me if I was wrong.","created":"2017-08-16T12:13:38.234+0000"},{"body":"My doubt is correct.\nSay assume we have qual1 and qual0.\nWe first do a put for qual1 and then we add a deleteFamily.\nSay in the same test case after DeleteFamily is added, if we do puts for qual0 (instead of empty qual as done now) every thing works fine. I think the simple reason is because just after adding \nPut (qual1), Delete family, put(qual0, val0), put(qual0, val1) - the Deletefamily always sorts out first becuase it knows that qual0 is lesser than qual1 and so while scanning DeleteFamily always peeks out as the first Cell when we do StoreScanner#next().\nBut when an empty qualifier is added the sorting takes a different pattern.\nIdeally a cell with qualifier qual0 and a cell without qualifier the cell without qualifier should sort first and then the cell with qualifier.\nI think in CellComparator#compareColums if we can handle this then we are able to solve this issue?\n{code}\n if(lclength != 0 && rclength == 0) {\n // means the right hand side should be sorted lower.\n return 1;\n }\n if(lclength == 0 && rclength != 0) {\n // means the right hand side should be sorted higher.\n return -1;\n }\n{code}\nWhat do you think [~chia7712]?","created":"2017-08-16T18:29:38.067+0000"},{"body":"The bug is due to the key returned from getKeyForNextColumn(). The key say that *we should skip all cells which have the same qualifier with current cell*, so the DeleteFamily is skipped.\nex1: Put (qual1), Delete family, put(qual0, val0), put(qual0, val1)\nThe *DeleteFamily* always peeks out as the first Cell, so every thing works fine.\n\nex2: Put (qual1), Delete family, put(empty, val0), put(empty, val1)\nThe *put(empty, val1)* always peeks out as the first Cell. After put(empty, val0) peeks out, the matcher say \"SeekNextColumn\". StoreScanner use the key from matcher#getKeyForNextColumn() to seek next column. The layout of the key is shown below:\n|row|fam|empty|HConstants.OLDEST_TIMESTAMP|Type.Minimum| \nHence, the DeleteFamily is skipped because its timestamp is larger than the key.\n\nbq. Ideally a cell with qualifier qual0 and a cell without qualifier the cell without qualifier should sort first and then the cell with qualifier.\nIt seems to me that the CellComparator works normally about this.\n\nbq. I think in CellComparator#compareColums if we can handle this then we are able to solve this issue?\nPardon me, I fail to catch what you mean here.","created":"2017-08-17T02:30:47.002+0000"},{"body":"bq.The key say that we should skip all cells which have the same qualifier with current cell, so the DeleteFamily is skipped.\nSo my main doubt was that with the first cell that is peeked why is that a difference. I expected qual0 and empty qualifier both to be lesser than the qual1. Will check the patch.","created":"2017-08-17T04:19:03.204+0000"},{"body":"Ya CellComparator works fine. My bad.","created":"2017-08-17T04:34:09.230+0000"},{"body":"[~chia7712], I have a question, cell with type 'DeleteFamily' or 'DeleteColumn' should always sort before any kvs in one row, right? So why a normal kv with empty column sorted before delete mark and then the delete marker is skipped by 'SeekNextColumn' ?","created":"2017-08-17T05:56:27.254+0000"},{"body":"bq. why a normal kv with empty column sorted before delete mark\n{code:title=CellComparator#compareWithoutRow}\n...\n diff = compareTimestamps(left, right);\n if (diff != 0) return diff;\n\n // Compare types. Let the delete types sort ahead of puts; i.e. types\n // of higher numbers sort before those of lesser numbers. Maximum (255)\n // appears ahead of everything, and minimum (0) appears after\n // everything.\n return (0xff & right.getTypeByte()) - (0xff & left.getTypeByte());\n{code}\nWe compare the TS first, so normal kv with empty column can be sorted before delete mark if the kv's ts is larger(newer) than delete mark.","created":"2017-08-17T06:04:11.258+0000"},{"body":"IMHO the first peek() that happens after a StoreScanner is created does not involve the getKeyForNextColumn(). It is just what ever comes from the Memstore CSLM. ( when the cells are in memory)\nWhen put(qual1, val) and DeleteFamily() is followed by put(empty, val0) - the comparator compares the latest put and the deleteFamily sees that the qualifier is same and it goes to the timestamp. In the timestamp the put sorts out first because that is the latest. In ts we purposefully consider the latest to appear first.\n\nBut in the case put(qual1, val) and DeleteFamily() is followed by put(qual0, val0) - in this case the put is considered to be larger than the delete family and hence put sorts out next. \nThat is why on the first peek itself things change. And then yes with getKeyForNextColumn() we tend to get the next value. So this patch tries to avoid that. I really don't know if there is a better way to work on this currently. ","created":"2017-08-17T06:12:30.634+0000"},{"body":"[~chia7712] - Am just seeing your comment after I was debugging the flow. You seem to say the same point. ","created":"2017-08-17T06:14:00.868+0000"},{"body":"{quote}\nWe compare the TS first, so normal kv with empty column can be sorted before delete mark if the kv's ts is larger(newer) than delete mark.\n{quote}\nI think this is not right, the delete marker should always be read before any normal kv. So that the DeleteTracker can collect all delete markers before working.","created":"2017-08-17T06:17:59.262+0000"},{"body":"bq. I really don't know if there is a better way to work on this currently.\nThe cost of current solution is lower performance if there are many cells which have empty qualifier. We can't seek to next column directly. That say, we need to parse all cells one by one. But so far i don't find any better way.","created":"2017-08-17T06:30:50.955+0000"},{"body":"bq. the delete marker should always be read before any normal kv.\nWe may scan many useless delete marks if they are related to older PUT cells.","created":"2017-08-17T06:32:59.980+0000"},{"body":"IMHO in an ideal case there is always a qualifier and puts with no qualifiers are rare, unless we see this type of issues with cell level ACLs where an user has tried issuing a delete family but some other malicious user tries to retrieve the older version of the cells by doing a Put with empty qualifier.","created":"2017-08-17T06:35:15.893+0000"},{"body":"[~ram_krish] Thanks for the great case. Any more suggestions?","created":"2017-08-17T11:08:54.929+0000"},{"body":"So you make a cell by just changing the type to minimum and so the seek tries to seek to a cell lesser than the current Cell (the cell with empty qualifier). So we find the deleteFamily. \nI hope all other delete varieties does not have this problem because I think they will be having a qualifier associated with it.\nI don have any suggestion currently to make the comparator more intelligent in this case. +1.","created":"2017-08-18T07:14:27.790+0000"},{"body":"BTW this is great debugging :)","created":"2017-08-18T07:16:47.829+0000"},{"body":"Thanks for the comment. [~ram_krish]\nNot that I hated current solution less, but that I hated data loss more. TODO will be left in getKeyForNextColumn() for reminding someone to give a better way.\n\nWill commit it tomorrow if no objections.\n\n","created":"2017-08-18T07:29:13.805+0000"},{"body":"Thanks for all reviews.","created":"2017-08-18T18:18:43.856+0000"}],"conversations":[{"body":"The qualifier of a deleted row (with keep deleted cells true) re-appears after re-inserting the same row multiple times (with different timestamp) with an empty qualifier.\n\nScenario:\n# Put row with family and qualifier (timestamp 1).\n# Delete entire row (timestamp 2).\n# Put same row again with family without qualifier (timestamp 3).\nA scan (latest version) returns the row with family without qualifier, version 3 (which is correct).\n# Put the same row again with family without qualifier (timestamp 4).\nA scan (latest version) returns multiple rows:\n* the row with family without qualifier, version 4 (which is correct).\n* the row with family with qualifier, version 1 (which is wrong).\n\nThere is a test scenario attached.\noutput:\n 13:42:53,952 [main] client.HBaseAdmin - Started disable of test_dml\n 13:42:55,801 [main] client.HBaseAdmin - Disabled test_dml\n 13:42:57,256 [main] client.HBaseAdmin - Deleted test_dml\n 13:42:58,592 [main] client.HBaseAdmin - Created test_dml\nPut row: 'myRow' with family: 'myFamily' with qualifier: 'myQualifier' with timestamp: '1'\nScan printout =>\n Row: 'myRow', Timestamp: '1', Family: 'myFamily', Qualifier: 'myQualifier', Value: 'myValue'\nDelete row: 'myRow'\nScan printout =>\nPut row: 'myRow' with family: 'myFamily' with qualifier: 'null' with timestamp: '3'\nScan printout =>\n Row: 'myRow', Timestamp: '3', Family: 'myFamily', Qualifier: '', Value: 'myValue'\nPut row: 'myRow' with family: 'myFamily' with qualifier: 'null' with timestamp: '4'\nScan printout =>\n Row: 'myRow', Timestamp: '4', Family: 'myFamily', Qualifier: '', Value: 'myValue'\n {color:red}Row: 'myRow', Timestamp: '1', Family: 'myFamily', Qualifier: 'myQualifier', Value: 'myValue'{color}\n","from":"reporter","subject":"The DeleteFamily cell is skipped when StoreScanner seeks to next column"},{"body":"Another difference in {{scanner}} behavior, after the test code is run\n{code}\nhbase(main):023:0> scan 'test_dml'\nROW COLUMN+CELL \n myRow column=myFamily:, timestamp=4, value=myValue \n myRow column=myFamily:myQualifier, timestamp=1, value=myValue \n1 row(s) in 0.0330 seconds\n\nhbase(main):024:0> scan 'test_dml', {VERSIONS => 10}\nROW COLUMN+CELL \n myRow column=myFamily:, timestamp=4, value=myValue \n myRow column=myFamily:, timestamp=3, value=myValue \n1 row(s) in 0.0290 seconds\n{code}","from":"developer"},{"body":"We add a cell without qualifier for deleting all columns of the specified family.\n{code}\n public Delete addFamily(final byte [] family, final long timestamp) {\n if (timestamp < 0) {\n throw new IllegalArgumentException(\"Timestamp cannot be negative. ts=\" + timestamp);\n }\n List list = familyMap.get(family);\n if(list == null) {\n list = new ArrayList<>(1);\n } else if(!list.isEmpty()) {\n list.clear();\n }\n KeyValue kv = new KeyValue(row, family, null, timestamp, KeyValue.Type.DeleteFamily);\n list.add(kv);\n familyMap.put(family, list);\n return this;\n }\n{code}\n\nThe *hint cell* created by CellUtil.createLastOnRowCol force the KVScanner to skip the remaining cells which have the same family/qualifier with *hint cell*\n{code}\n public Cell getKeyForNextColumn(Cell cell) {\n ColumnCount nextColumn = columns.getColumnHint();\n if (nextColumn == null) {\n return CellUtil.createLastOnRowCol(cell);\n } else {\n return CellUtil.createFirstOnRowCol(cell, nextColumn.getBuffer(), nextColumn.getOffset(),\n nextColumn.getLength());\n }\n }\n{code}\n\n","from":"developer"},{"body":"Ping for reviews. This bug makes our hbase lose data. We should resolve it before next release.","from":"developer"},{"body":"lgtm\n\nCan you post patch for branch-1.2 ?","from":"developer"},{"body":"Will post patch for branch-1 and branch-1.2\n\n{quote}\n+ /**\n+ * @param cell current cell\n+ * @return Null if we have no idea about the \"hint\" of next key\n+ */\n{quote}\nThe above comment is wrong...I will remove it.","from":"developer"},{"body":"Nice debug. Ya this is a serious issue but did not unearth till now !\n{code}\n if (type != Type.Minimum.getCode()) {\n3148\t type = KeyValue.Type.values()[KeyValue.Type.codeToType(type).ordinal() - 1].getCode();\n3149\t } else if (ts != HConstants.OLDEST_TIMESTAMP) {\n3150\t ts = ts - 1;\n3151\t type = Type.Maximum.getCode();\n3152\t }\n{code}\nSo as in the test, we will call with a Put type cell. So the type is not min and 1st if is through. We will use the type as one below the PUT type which is again MINIMUM. But then we will NOT change the ts. That is the diff what we make (?)\nIn our compare logic, there is special treatment for min and max I guess. We wont even see the TS then (?) I may be wrong. If so am not sure how we will solve issue here.","from":"developer"},{"body":"bq. In our compare logic, there is special treatment for min and max I guess. We wont even see the TS\nDo you mean this special treatment?\n{code:title=CellComparator#compareWithoutRow}\n if (lFamLength + lQualLength == 0\n && left.getTypeByte() == Type.Minimum.getCode()) {\n // left is \"bigger\", i.e. it appears later in the sorted order\n return 1;\n }\n if (rFamLength + rQualLength == 0\n && right.getTypeByte() == Type.Minimum.getCode()) {\n return -1;\n }\n{code}\nThere is no special treatment for max. Also i don' find any case that the cell has empty column and PUT type. (i mean this patch won't generate the cell which has empty column and MINIMUM type.)","from":"developer"},{"body":"v1 patch\n# remove incorrect comment\n# add more tests (Ted's suggestion)","from":"developer"},{"body":"Will check this tomorrow. ","from":"developer"},{"body":"Reviewing this. So reading the patch and the description. After a DeleteFamily here a Put with empty qualifier is being added. What happens if a put with a qualifier which is lexographically smaller than the existing QUALIFIER happens after the DeleteFamily?","from":"developer"},{"body":"bq. After a DeleteFamily here a Put\nbq. What happens if a put with a qualifier which is lexographically smaller than the existing QUALIFIER happens after the DeleteFamily?\nDo you mean all puts happen *after* the DeleteFamily? If yes, the DeleteFamily cell doesn't delete any PUT cells. Please correct me if I was wrong.","from":"developer"},{"body":"My doubt is correct.\nSay assume we have qual1 and qual0.\nWe first do a put for qual1 and then we add a deleteFamily.\nSay in the same test case after DeleteFamily is added, if we do puts for qual0 (instead of empty qual as done now) every thing works fine. I think the simple reason is because just after adding \nPut (qual1), Delete family, put(qual0, val0), put(qual0, val1) - the Deletefamily always sorts out first becuase it knows that qual0 is lesser than qual1 and so while scanning DeleteFamily always peeks out as the first Cell when we do StoreScanner#next().\nBut when an empty qualifier is added the sorting takes a different pattern.\nIdeally a cell with qualifier qual0 and a cell without qualifier the cell without qualifier should sort first and then the cell with qualifier.\nI think in CellComparator#compareColums if we can handle this then we are able to solve this issue?\n{code}\n if(lclength != 0 && rclength == 0) {\n // means the right hand side should be sorted lower.\n return 1;\n }\n if(lclength == 0 && rclength != 0) {\n // means the right hand side should be sorted higher.\n return -1;\n }\n{code}\nWhat do you think [~chia7712]?","from":"developer"},{"body":"The bug is due to the key returned from getKeyForNextColumn(). The key say that *we should skip all cells which have the same qualifier with current cell*, so the DeleteFamily is skipped.\nex1: Put (qual1), Delete family, put(qual0, val0), put(qual0, val1)\nThe *DeleteFamily* always peeks out as the first Cell, so every thing works fine.\n\nex2: Put (qual1), Delete family, put(empty, val0), put(empty, val1)\nThe *put(empty, val1)* always peeks out as the first Cell. After put(empty, val0) peeks out, the matcher say \"SeekNextColumn\". StoreScanner use the key from matcher#getKeyForNextColumn() to seek next column. The layout of the key is shown below:\n|row|fam|empty|HConstants.OLDEST_TIMESTAMP|Type.Minimum| \nHence, the DeleteFamily is skipped because its timestamp is larger than the key.\n\nbq. Ideally a cell with qualifier qual0 and a cell without qualifier the cell without qualifier should sort first and then the cell with qualifier.\nIt seems to me that the CellComparator works normally about this.\n\nbq. I think in CellComparator#compareColums if we can handle this then we are able to solve this issue?\nPardon me, I fail to catch what you mean here.","from":"developer"},{"body":"bq.The key say that we should skip all cells which have the same qualifier with current cell, so the DeleteFamily is skipped.\nSo my main doubt was that with the first cell that is peeked why is that a difference. I expected qual0 and empty qualifier both to be lesser than the qual1. Will check the patch.","from":"developer"},{"body":"Ya CellComparator works fine. My bad.","from":"developer"},{"body":"[~chia7712], I have a question, cell with type 'DeleteFamily' or 'DeleteColumn' should always sort before any kvs in one row, right? So why a normal kv with empty column sorted before delete mark and then the delete marker is skipped by 'SeekNextColumn' ?","from":"developer"},{"body":"bq. why a normal kv with empty column sorted before delete mark\n{code:title=CellComparator#compareWithoutRow}\n...\n diff = compareTimestamps(left, right);\n if (diff != 0) return diff;\n\n // Compare types. Let the delete types sort ahead of puts; i.e. types\n // of higher numbers sort before those of lesser numbers. Maximum (255)\n // appears ahead of everything, and minimum (0) appears after\n // everything.\n return (0xff & right.getTypeByte()) - (0xff & left.getTypeByte());\n{code}\nWe compare the TS first, so normal kv with empty column can be sorted before delete mark if the kv's ts is larger(newer) than delete mark.","from":"developer"},{"body":"IMHO the first peek() that happens after a StoreScanner is created does not involve the getKeyForNextColumn(). It is just what ever comes from the Memstore CSLM. ( when the cells are in memory)\nWhen put(qual1, val) and DeleteFamily() is followed by put(empty, val0) - the comparator compares the latest put and the deleteFamily sees that the qualifier is same and it goes to the timestamp. In the timestamp the put sorts out first because that is the latest. In ts we purposefully consider the latest to appear first.\n\nBut in the case put(qual1, val) and DeleteFamily() is followed by put(qual0, val0) - in this case the put is considered to be larger than the delete family and hence put sorts out next. \nThat is why on the first peek itself things change. And then yes with getKeyForNextColumn() we tend to get the next value. So this patch tries to avoid that. I really don't know if there is a better way to work on this currently. ","from":"developer"},{"body":"[~chia7712] - Am just seeing your comment after I was debugging the flow. You seem to say the same point. ","from":"developer"},{"body":"{quote}\nWe compare the TS first, so normal kv with empty column can be sorted before delete mark if the kv's ts is larger(newer) than delete mark.\n{quote}\nI think this is not right, the delete marker should always be read before any normal kv. So that the DeleteTracker can collect all delete markers before working.","from":"developer"},{"body":"bq. I really don't know if there is a better way to work on this currently.\nThe cost of current solution is lower performance if there are many cells which have empty qualifier. We can't seek to next column directly. That say, we need to parse all cells one by one. But so far i don't find any better way.","from":"developer"},{"body":"bq. the delete marker should always be read before any normal kv.\nWe may scan many useless delete marks if they are related to older PUT cells.","from":"developer"},{"body":"IMHO in an ideal case there is always a qualifier and puts with no qualifiers are rare, unless we see this type of issues with cell level ACLs where an user has tried issuing a delete family but some other malicious user tries to retrieve the older version of the cells by doing a Put with empty qualifier.","from":"developer"},{"body":"[~ram_krish] Thanks for the great case. Any more suggestions?","from":"developer"},{"body":"So you make a cell by just changing the type to minimum and so the seek tries to seek to a cell lesser than the current Cell (the cell with empty qualifier). So we find the deleteFamily. \nI hope all other delete varieties does not have this problem because I think they will be having a qualifier associated with it.\nI don have any suggestion currently to make the comparator more intelligent in this case. +1.","from":"developer"},{"body":"BTW this is great debugging :)","from":"developer"},{"body":"Thanks for the comment. [~ram_krish]\nNot that I hated current solution less, but that I hated data loss more. TODO will be left in getKeyForNextColumn() for reminding someone to give a better way.\n\nWill commit it tomorrow if no objections.\n\n","from":"developer"},{"body":"Thanks for all reviews.","from":"developer"}],"created":"2017-07-28T11:38:13.000+0000","description":"The qualifier of a deleted row (with keep deleted cells true) re-appears after re-inserting the same row multiple times (with different timestamp) with an empty qualifier.\n\nScenario:\n# Put row with family and qualifier (timestamp 1).\n# Delete entire row (timestamp 2).\n# Put same row again with family without qualifier (timestamp 3).\nA scan (latest version) returns the row with family without qualifier, version 3 (which is correct).\n# Put the same row again with family without qualifier (timestamp 4).\nA scan (latest version) returns multiple rows:\n* the row with family without qualifier, version 4 (which is correct).\n* the row with family with qualifier, version 1 (which is wrong).\n\nThere is a test scenario attached.\noutput:\n 13:42:53,952 [main] client.HBaseAdmin - Started disable of test_dml\n 13:42:55,801 [main] client.HBaseAdmin - Disabled test_dml\n 13:42:57,256 [main] client.HBaseAdmin - Deleted test_dml\n 13:42:58,592 [main] client.HBaseAdmin - Created test_dml\nPut row: 'myRow' with family: 'myFamily' with qualifier: 'myQualifier' with timestamp: '1'\nScan printout =>\n Row: 'myRow', Timestamp: '1', Family: 'myFamily', Qualifier: 'myQualifier', Value: 'myValue'\nDelete row: 'myRow'\nScan printout =>\nPut row: 'myRow' with family: 'myFamily' with qualifier: 'null' with timestamp: '3'\nScan printout =>\n Row: 'myRow', Timestamp: '3', Family: 'myFamily', Qualifier: '', Value: 'myValue'\nPut row: 'myRow' with family: 'myFamily' with qualifier: 'null' with timestamp: '4'\nScan printout =>\n Row: 'myRow', Timestamp: '4', Family: 'myFamily', Qualifier: '', Value: 'myValue'\n {color:red}Row: 'myRow', Timestamp: '1', Family: 'myFamily', Qualifier: 'myQualifier', Value: 'myValue'{color}\n","issue_id":"13090752","key":"HBASE-18471","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2017-08-18T18:18:43.000+0000","role":"fixed_distractor","summary":"The DeleteFamily cell is skipped when StoreScanner seeks to next column"} {"case_id":"12436937","cluster":"DISTRACTOR-HBASE-1876","comments":[{"body":"We should make this a recommended patch for hadoop installs running hbase. Let me add it to 'Getting Started' list.","created":"2009-09-30T23:33:36.762+0000"},{"body":"Should we roll the HDFS-630 patch into the patched Hadoop jar included in the HBase distrib, alongside the patch for HDFS-127? ","created":"2009-10-05T22:22:19.863+0000"},{"body":"Its not client-side only like the hdfs-127 patch. It needs the namenode patched. I'm adding a note to our getting started recommending adding this patch to your hadoop install. I think that enough for hbase 0.20.1.","created":"2009-10-05T22:27:35.231+0000"},{"body":"I added to our 'Getting Started' a recommendation that users apply hdfs-630 to their hadoop cluster on branch and trunk.\n\nI think need for hdfs-630 is going to become more apparent as we test new sync/append, especially on small clusters.\n\nMoving out of 0.20.1 now....","created":"2009-10-05T22:38:50.880+0000"},{"body":"I adapted the previous patch for HDFS-630 to 0.21.x branch \nhttps://issues.apache.org/jira/secure/attachment/12422242/0001-Fix-HDFS-630-for-0.21.patch","created":"2009-10-15T16:00:43.190+0000"},{"body":"I did more testing of hdfs-630. For sure it helps with the above situation. To underline how necessary we think this patch is, especially when cluster is small, I've add the patch to the hadoop-hdfs.jar bundled with hbase.","created":"2009-11-15T00:29:16.614+0000"},{"body":"stack: the patched DFSClient is not compatible with unpatched NameNode, so if we're going to include the patch in the hadoop-hdfs.jar we need to explain that it must be used with a patched NameNode as well. ","created":"2009-11-15T09:59:40.002+0000"},{"body":"@Cosmin: Thats bad. Thanks. Let me undo.","created":"2009-11-15T16:34:20.482+0000"},{"body":"I undid bundling an hadoop-hdfs patched with hdfs-630 being part of hbase deploy.","created":"2009-12-12T23:20:58.586+0000"},{"body":"Linking to HDFS-630","created":"2009-12-14T09:50:24.466+0000"},{"body":"hdfs-630 is in 0.21 hadoop and hadoop trunk. I just suggested that it get added to 0.20-append branch. If it goes in, we can resolve this issue against hbase 0.21.","created":"2010-06-03T17:45:24.184+0000"},{"body":"Stack:\nbq. hdfs-630 is in 0.21 hadoop and hadoop trunk. I just suggested that it get added to 0.20-append branch. If it goes in, we can resolve this issue against hbase 0.21.\n\n+1\n\nWe have been using a Hadoop patched with 630 internally so HBase is stable on small-ish DFS clusters. \n\n\n\n \n\n","created":"2010-06-04T18:26:57.446+0000"},{"body":"Dhruba committed a hdfs-630 but then I saw that Cosmin commented suggesting that Dhruba use a patch Todd made for 0.20 branch. I wonder why? Let me ask him.\n\n","created":"2010-06-04T18:39:08.552+0000"},{"body":"oh, nm... it hasn't been committed yet. I misread the issue.","created":"2010-06-04T18:40:27.173+0000"},{"body":"@Cosmin Can we close this? branch-0.20-append, what we have checked into hbase and what we expect to run on now has hdfs-630.","created":"2010-07-16T23:44:34.537+0000"},{"body":"We're currently running CDH3b2 and it looks good. I think it's safe to close it now with HDFS-630 committed to 0.20-append.","created":"2010-07-17T07:51:57.197+0000"},{"body":"Closing w/ Cosmin's blessing.","created":"2010-07-17T15:38:05.537+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T13:02:00.063+0000"}],"conversations":[{"body":"A dead datanode in the cluster can lead to multiple HRegionServer failures and corrupted data. The HRegionServer failures can be reproduced consistently on a 7 machines cluster with approx 2000 regions.\n\nSteps to reproduce\n\nThe easiest and safest way is to reproduce it for the .META. table, however it will work with any table. \n\nLocate a datanode that stores the .META. files and kill -9 it. \nIn order to get multiple writes to the .META. table bring up or shut down a region server this will eventually cause a flush on the memstore\n\n2009-09-25 09:26:17,775 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Flush requested on .META.,demo__assets,asset_283132172,1252898166036,1253265069920\n2009-09-25 09:26:17,775 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Started memstore flush for region .META.,demo__assets,asset_283132172,1252898166036,1253265069920. Current region memstore si\nze 16.3k\n2009-09-25 09:26:17,791 INFO org.apache.hadoop.hdfs.DFSClient: Exception in createBlockOutputStream java.io.IOException: Bad connect ack with firstBadLink 10.72.79.108:50010\n2009-09-25 09:26:17,791 INFO org.apache.hadoop.hdfs.DFSClient: Abandoning block blk_-8767099282771605606_176852\n\n\n\nThe DFSClient will retry for 3 times, but there's a high chance it will try on the same failed datanode (it takes around 10 minutes for dead datanode to be removed from cluster)\n\n\n\n\n2009-09-25 09:26:41,810 WARN org.apache.hadoop.hdfs.DFSClient: DataStreamer Exception: java.io.IOException: Unable to create new block.\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream.nextBlockOutputStream(DFSClient.java:2814)\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream.access$2000(DFSClient.java:2078)\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream$DataStreamer.run(DFSClient.java:2264)\n\n2009-09-25 09:26:41,810 WARN org.apache.hadoop.hdfs.DFSClient: Error Recovery for block blk_5317304716016587434_176852 bad datanode[2] nodes == null\n2009-09-25 09:26:41,810 WARN org.apache.hadoop.hdfs.DFSClient: Could not get block locations. Source file \"/hbase/.META./225980069/info/5573114819456511457\" - Aborting...\n2009-09-25 09:26:41,810 FATAL org.apache.hadoop.hbase.regionserver.MemStoreFlusher: Replay of hlog required. Forcing server shutdown\norg.apache.hadoop.hbase.DroppedSnapshotException: region: .META.,demo__assets,asset_283132172,1252898166036,1253265069920\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:942)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:835)\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:241)\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.run(MemStoreFlusher.java:149)\nCaused by: java.io.IOException: Bad connect ack with firstBadLink 10.72.79.108:50010\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream.createBlockOutputStream(DFSClient.java:2872)\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream.nextBlockOutputStream(DFSClient.java:2795)\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream.access$2000(DFSClient.java:2078)\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream$DataStreamer.run(DFSClient.java:2264)\n\n\n\nAfter the HRegionServer shuts down itself the regions will be reassigned however you might hit this \n\n\n\n2009-09-26 08:04:23,646 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_OPEN: .META.,demo__assets,asset_283132172,1252898166036,1253265069920\n2009-09-26 08:04:23,684 WARN org.apache.hadoop.hbase.regionserver.Store: Skipping hdfs://b0:9000/hbase/.META./225980069/historian/1432202951743803786 because its empty. HBASE-646 DATA LOSS?\n...\n2009-09-26 08:04:23,776 INFO org.apache.hadoop.hbase.regionserver.HRegion: region .META.,demo__assets,asset_283132172,1252898166036,1253265069920/225980069 available; sequence id is 1331458484\n\nWe ended up with corrupted data in .META. \"info:server\" after master got confirmation that it was updated from the HRegionServer that got DroppedSnapshotException\n\nSince after a cluster restart server:info will be correct, .META. is safer to test with. Also to detect data corruption you can just scan .META. get the start key for each region and attempt to retrieve it from the corresponding table. If .META. is corrupted you get a NotServingRegionException. \n\nThis issue is related to https://issues.apache.org/jira/browse/HDFS-630 \n\nI attached a patch for HDFS-630 https://issues.apache.org/jira/secure/attachment/12420919/HDFS-630.patch that fixes this problem. \n","from":"reporter","subject":"DroppedSnapshotException when flushing memstore after a datanode dies"},{"body":"We should make this a recommended patch for hadoop installs running hbase. Let me add it to 'Getting Started' list.","from":"developer"},{"body":"Should we roll the HDFS-630 patch into the patched Hadoop jar included in the HBase distrib, alongside the patch for HDFS-127? ","from":"developer"},{"body":"Its not client-side only like the hdfs-127 patch. It needs the namenode patched. I'm adding a note to our getting started recommending adding this patch to your hadoop install. I think that enough for hbase 0.20.1.","from":"developer"},{"body":"I added to our 'Getting Started' a recommendation that users apply hdfs-630 to their hadoop cluster on branch and trunk.\n\nI think need for hdfs-630 is going to become more apparent as we test new sync/append, especially on small clusters.\n\nMoving out of 0.20.1 now....","from":"developer"},{"body":"I adapted the previous patch for HDFS-630 to 0.21.x branch \nhttps://issues.apache.org/jira/secure/attachment/12422242/0001-Fix-HDFS-630-for-0.21.patch","from":"developer"},{"body":"I did more testing of hdfs-630. For sure it helps with the above situation. To underline how necessary we think this patch is, especially when cluster is small, I've add the patch to the hadoop-hdfs.jar bundled with hbase.","from":"developer"},{"body":"stack: the patched DFSClient is not compatible with unpatched NameNode, so if we're going to include the patch in the hadoop-hdfs.jar we need to explain that it must be used with a patched NameNode as well. ","from":"developer"},{"body":"@Cosmin: Thats bad. Thanks. Let me undo.","from":"developer"},{"body":"I undid bundling an hadoop-hdfs patched with hdfs-630 being part of hbase deploy.","from":"developer"},{"body":"Linking to HDFS-630","from":"developer"},{"body":"hdfs-630 is in 0.21 hadoop and hadoop trunk. I just suggested that it get added to 0.20-append branch. If it goes in, we can resolve this issue against hbase 0.21.","from":"developer"},{"body":"Stack:\nbq. hdfs-630 is in 0.21 hadoop and hadoop trunk. I just suggested that it get added to 0.20-append branch. If it goes in, we can resolve this issue against hbase 0.21.\n\n+1\n\nWe have been using a Hadoop patched with 630 internally so HBase is stable on small-ish DFS clusters. \n\n\n\n \n\n","from":"developer"},{"body":"Dhruba committed a hdfs-630 but then I saw that Cosmin commented suggesting that Dhruba use a patch Todd made for 0.20 branch. I wonder why? Let me ask him.\n\n","from":"developer"},{"body":"oh, nm... it hasn't been committed yet. I misread the issue.","from":"developer"},{"body":"@Cosmin Can we close this? branch-0.20-append, what we have checked into hbase and what we expect to run on now has hdfs-630.","from":"developer"},{"body":"We're currently running CDH3b2 and it looks good. I think it's safe to close it now with HDFS-630 committed to 0.20-append.","from":"developer"},{"body":"Closing w/ Cosmin's blessing.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2009-09-30T16:48:43.000+0000","description":"A dead datanode in the cluster can lead to multiple HRegionServer failures and corrupted data. The HRegionServer failures can be reproduced consistently on a 7 machines cluster with approx 2000 regions.\n\nSteps to reproduce\n\nThe easiest and safest way is to reproduce it for the .META. table, however it will work with any table. \n\nLocate a datanode that stores the .META. files and kill -9 it. \nIn order to get multiple writes to the .META. table bring up or shut down a region server this will eventually cause a flush on the memstore\n\n2009-09-25 09:26:17,775 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Flush requested on .META.,demo__assets,asset_283132172,1252898166036,1253265069920\n2009-09-25 09:26:17,775 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Started memstore flush for region .META.,demo__assets,asset_283132172,1252898166036,1253265069920. Current region memstore si\nze 16.3k\n2009-09-25 09:26:17,791 INFO org.apache.hadoop.hdfs.DFSClient: Exception in createBlockOutputStream java.io.IOException: Bad connect ack with firstBadLink 10.72.79.108:50010\n2009-09-25 09:26:17,791 INFO org.apache.hadoop.hdfs.DFSClient: Abandoning block blk_-8767099282771605606_176852\n\n\n\nThe DFSClient will retry for 3 times, but there's a high chance it will try on the same failed datanode (it takes around 10 minutes for dead datanode to be removed from cluster)\n\n\n\n\n2009-09-25 09:26:41,810 WARN org.apache.hadoop.hdfs.DFSClient: DataStreamer Exception: java.io.IOException: Unable to create new block.\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream.nextBlockOutputStream(DFSClient.java:2814)\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream.access$2000(DFSClient.java:2078)\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream$DataStreamer.run(DFSClient.java:2264)\n\n2009-09-25 09:26:41,810 WARN org.apache.hadoop.hdfs.DFSClient: Error Recovery for block blk_5317304716016587434_176852 bad datanode[2] nodes == null\n2009-09-25 09:26:41,810 WARN org.apache.hadoop.hdfs.DFSClient: Could not get block locations. Source file \"/hbase/.META./225980069/info/5573114819456511457\" - Aborting...\n2009-09-25 09:26:41,810 FATAL org.apache.hadoop.hbase.regionserver.MemStoreFlusher: Replay of hlog required. Forcing server shutdown\norg.apache.hadoop.hbase.DroppedSnapshotException: region: .META.,demo__assets,asset_283132172,1252898166036,1253265069920\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:942)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:835)\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:241)\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.run(MemStoreFlusher.java:149)\nCaused by: java.io.IOException: Bad connect ack with firstBadLink 10.72.79.108:50010\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream.createBlockOutputStream(DFSClient.java:2872)\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream.nextBlockOutputStream(DFSClient.java:2795)\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream.access$2000(DFSClient.java:2078)\n at org.apache.hadoop.hdfs.DFSClient$DFSOutputStream$DataStreamer.run(DFSClient.java:2264)\n\n\n\nAfter the HRegionServer shuts down itself the regions will be reassigned however you might hit this \n\n\n\n2009-09-26 08:04:23,646 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_OPEN: .META.,demo__assets,asset_283132172,1252898166036,1253265069920\n2009-09-26 08:04:23,684 WARN org.apache.hadoop.hbase.regionserver.Store: Skipping hdfs://b0:9000/hbase/.META./225980069/historian/1432202951743803786 because its empty. HBASE-646 DATA LOSS?\n...\n2009-09-26 08:04:23,776 INFO org.apache.hadoop.hbase.regionserver.HRegion: region .META.,demo__assets,asset_283132172,1252898166036,1253265069920/225980069 available; sequence id is 1331458484\n\nWe ended up with corrupted data in .META. \"info:server\" after master got confirmation that it was updated from the HRegionServer that got DroppedSnapshotException\n\nSince after a cluster restart server:info will be correct, .META. is safer to test with. Also to detect data corruption you can just scan .META. get the start key for each region and attempt to retrieve it from the corresponding table. If .META. is corrupted you get a NotServingRegionException. \n\nThis issue is related to https://issues.apache.org/jira/browse/HDFS-630 \n\nI attached a patch for HDFS-630 https://issues.apache.org/jira/secure/attachment/12420919/HDFS-630.patch that fixes this problem. \n","issue_id":"12436937","key":"HBASE-1876","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2010-07-17T15:38:05.000+0000","role":"fixed_distractor","summary":"DroppedSnapshotException when flushing memstore after a datanode dies"} {"case_id":"13120125","cluster":"DISTRACTOR-HBASE-19321","comments":[{"body":"Looking at the stack trace in hung state:\r\n{code}\r\n\"Time-limited test\" #814 daemon prio=5 os_prio=31 tid=0x00007fa39d5d7000 nid=0x47a13 in Object.wait() [0x0000700022aa9000]\r\n java.lang.Thread.State: WAITING (on object monitor)\r\n at java.lang.Object.wait(Native Method)\r\n at java.lang.Object.wait(Object.java:502)\r\n at org.apache.curator.framework.state.ConnectionStateManager.blockUntilConnected(ConnectionStateManager.java:224)\r\n - locked <0x0000000791798dd8> (a org.apache.curator.framework.state.ConnectionStateManager)\r\n at org.apache.curator.framework.imps.CuratorFrameworkImpl.blockUntilConnected(CuratorFrameworkImpl.java:266)\r\n at org.apache.curator.framework.imps.CuratorFrameworkImpl.blockUntilConnected(CuratorFrameworkImpl.java:272)\r\n at org.apache.hadoop.hbase.client.ZKAsyncRegistry.(ZKAsyncRegistry.java:84)\r\n{code}\r\nIt seems that blockUntilConnected() never finished.\r\n","created":"2017-11-22T02:55:31.407+0000"},{"body":"[~Apache9]:\r\nCan you take a look ?","created":"2017-11-22T02:55:55.580+0000"},{"body":"Let me take a look. If no cluster then no doubt we can not connect to zk... It does not make sense to create a zk registry.","created":"2017-11-22T03:00:26.631+0000"},{"body":"OK, the test is used to test we will fail if no cluster.\r\n\r\nMark it as Ignored I'd say. Will be fixed automatically after we remove the blockUntilConnected.\r\n\r\nThanks.","created":"2017-11-22T03:04:14.295+0000"},{"body":"Can variant of blockUntilConnected be used / added which takes into account zookeeper connection issue ?","created":"2017-11-22T03:14:38.238+0000"},{"body":"curator provides: blockUntilConnected(int, TimeUnit)","created":"2017-11-22T03:18:31.196+0000"},{"body":"Yeah there is method\r\n\r\n{code}\r\npublic boolean blockUntilConnected(int maxWaitTime, TimeUnit units) throws InterruptedException;\r\n{code}\r\n\r\nI think 1 or 2 seconds is enough? You can provide a patch, and also add the Thread.interrupt you want.\r\n\r\nThanks.","created":"2017-11-22T03:19:27.787+0000"},{"body":"You made all the previous changes.\r\n\r\nIt would be better you fix it.","created":"2017-11-22T03:23:09.774+0000"},{"body":"Do not have much time. Thanks.","created":"2017-11-22T03:27:55.962+0000"},{"body":"If you too do not have much time, I could find another person to fix this. What do you think? Thanks.","created":"2017-11-22T03:30:50.055+0000"},{"body":"Sounds like you're a manager now.\r\n\r\nThis is high priority issue - without the fix, HBASE-19200 should be reverted.","created":"2017-11-22T03:33:42.415+0000"},{"body":"I always say that, you can do what you think is right, if you want to revert then just revert, a test broken is enough to revert a commit. But I also have my own work plan, and also my judgement on which issue is more important. If you think this one is high priority then just do it.\r\n\r\nThanks.","created":"2017-11-22T03:44:27.182+0000"},{"body":"I have fixed this issue, but I can not assign this issue to myself, if someone can assign it to me?","created":"2017-11-22T05:52:00.078+0000"},{"body":"Assigned to you [~wuguoquan]. Please upload the patch. Thanks.","created":"2017-11-22T06:00:23.460+0000"},{"body":"+1. Let's wait for the pre commit result.","created":"2017-11-22T06:16:01.377+0000"},{"body":"ok, thanks","created":"2017-11-22T06:19:05.236+0000"},{"body":"Where is the Thread.interrupt() call ?\r\n\r\nHow is 2 second interval determined ? Isn't this short for real scenario ?","created":"2017-11-22T06:19:18.073+0000"},{"body":"{quote}\r\nA temporary workaround is to call blockUntilConnected when constructing ZKAsyncRegistry but this is not a production level code. \r\n{quote}\r\nI already said this is not a production level code, why do you keep asking 'what's the implication in production scenario', and 'Isn't this short for real scenario'? I can not get your point.\r\n\r\nThanks.","created":"2017-11-22T06:34:46.417+0000"},{"body":"Pushed to master and branch-2.\r\n\r\nThanks [~wuguoquan] for the contributing.","created":"2017-11-22T07:52:27.067+0000"},{"body":"same to you ","created":"2017-11-22T09:44:42.348+0000"},{"body":"Guoquan:\r\nTestAcidGuarantees started to timeout (again).\r\nCheck flaky test dashbaord:\r\nhttps://builds.apache.org/job/HBASE-Flaky-Tests/23484/","created":"2017-11-22T14:19:49.618+0000"},{"body":"Where is the dashboard? TestAcidGuarantees is not failed in your url?\n\nAnd now TestAcidGuarantees has been promoted to LargeTests, so if it still timeout then there must be other problems. It has been marked as flakey tests long long ago.\n\nThanks.","created":"2017-11-22T14:32:09.334+0000"},{"body":"Guoquan:\r\nDid you run TestAcidGuarantees locally (since patch was attached) ?\r\n\r\nbq. It has been marked as flakey tests long long ago\r\n\r\nI disagree. It got stuck after HBASE-19200 went in.\r\n\r\nFlaky test dashboard is here:\r\nhttps://builds.apache.org/job/HBASE-Find-Flaky-Tests/lastSuccessfulBuild/artifact/dashboard.html\r\n\r\nYou can see clearly the last 5 blue bars which indicate timeout.","created":"2017-11-22T14:38:09.230+0000"},{"body":"bq. I already said this is not a production level code\r\n\r\nhbase-2 is getting into beta phase where dev / QA at various companies have started testing the daily build(s).\r\n\r\nIt would be weird when they find out that what worked yesterday no longer works for the new daily build.\r\n\r\nIn this case, TestAcidGuarantees can be thought of as mini environment which can tell us how robust the change is.\r\nTestAcidGuarantees is a basic test. When developers make change to locking in HRegion, in-memory compaction, etc, they need to know that they are working with non-flaky base.","created":"2017-11-22T14:53:30.455+0000"},{"body":"Then please help finding out the root cause instead of commanding others and repeating useless word again and again. You are not a manager in the community, no one need to accept your command. And even HBASE-19200 is not found by you. I will not reply your comments about this topic any more, it is just wasting of time. You can revert everything you dislike since you are a commiter. I will revert back when I have time to digging in here.","created":"2017-11-22T14:57:24.586+0000"},{"body":"I tried changing the duration passed to blockUntilConnected from 2 to 30.\r\nThere was no improvement in the runtime of TestAcidGuarantees.\r\n\r\nNeed to dig deeper when I have time.","created":"2017-11-22T16:20:40.162+0000"},{"body":"{code}\r\n2017-11-22 13:13:28,358 ERROR [MemStoreFlusher.1] regionserver.HRegion(1210): Asked to modify this region's (TestAcidGuarantees,,1511356384562.67f5ede11bc9d37151794455003c8ee9.) memstoreSize to a negative value which is incorrect. Current memstoreSize=251900, delta=-281670\r\njava.lang.Exception\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.checkNegativeMemStoreDataSize(HRegion.java:1210)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.decrMemStoreSize(HRegion.java:1203)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.internalFlushCacheAndCommit(HRegion.java:2639)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2355)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2327)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:2218)\r\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:498)\r\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:467)\r\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.access$900(MemStoreFlusher.java:69)\r\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher$FlushHandler.run(MemStoreFlusher.java:252)\r\n\tat java.lang.Thread.run(Thread.java:748)\r\n{code}\r\nIt seems the eager policy cause the negative size of memstore, so it breaks the close of table (If the memstore size is negative, the rs will be aborted when closing the region.). We don't close the table before HBASE-19311 hence the issue doesn't hurt the {{TestAcidGuarantees}} previously. As [~Apache9] optimized the {{TestAcidGuarantees}} by creating the mini cluster once and creating/deleting the table for each test case, the issue hurts {{TestAcidGuarantees}} now. I revert the HBASE-19311 locally, and then run the {{TestAcidGuarantees}} with eager policy. The error about the negative size sill happens. We should file a issue to disable the eager policy for {{TestAcidGuarantees}} before we have time to dig in. [~tedyu] [~Apache9] WDYT?","created":"2017-11-22T16:30:55.983+0000"},{"body":"bq. We should file a issue to disable the eager policy for TestAcidGuarantees\r\n\r\nSince you are already working on HBASE-19266, this can be handled there - replacing EAGER with ADAPTIVE policy.","created":"2017-11-22T16:40:46.717+0000"},{"body":"bq. Since you are already working on HBASE-19266, this can be handled there - replacing EAGER with ADAPTIVE policy.\r\nWe still create the {{TestAcidGuarantees}} for EAGER but it will be denoted with {{Ignore}}. Let us discuss them in HBASE-19266. What about closing this issue? [~ted_yu]","created":"2017-11-22T16:46:18.116+0000"}],"conversations":[{"body":"From https://builds.apache.org/job/HBASE-Flaky-Tests/23477/testReport/junit/org.apache.hadoop.hbase.client/TestAdmin2/testCheckHBaseAvailableWithoutCluster/ :\r\n{code}\r\norg.junit.runners.model.TestTimedOutException: test timed out after 300000 milliseconds\r\n\tat org.apache.hadoop.hbase.client.TestAdmin2.testCheckHBaseAvailableWithoutCluster(TestAdmin2.java:573)\r\n{code}\r\nIt seems this started hanging after HBASE-19313","from":"reporter","subject":"ZKAsyncRegistry ctor would hang when zookeeper cluster is not available"},{"body":"Looking at the stack trace in hung state:\r\n{code}\r\n\"Time-limited test\" #814 daemon prio=5 os_prio=31 tid=0x00007fa39d5d7000 nid=0x47a13 in Object.wait() [0x0000700022aa9000]\r\n java.lang.Thread.State: WAITING (on object monitor)\r\n at java.lang.Object.wait(Native Method)\r\n at java.lang.Object.wait(Object.java:502)\r\n at org.apache.curator.framework.state.ConnectionStateManager.blockUntilConnected(ConnectionStateManager.java:224)\r\n - locked <0x0000000791798dd8> (a org.apache.curator.framework.state.ConnectionStateManager)\r\n at org.apache.curator.framework.imps.CuratorFrameworkImpl.blockUntilConnected(CuratorFrameworkImpl.java:266)\r\n at org.apache.curator.framework.imps.CuratorFrameworkImpl.blockUntilConnected(CuratorFrameworkImpl.java:272)\r\n at org.apache.hadoop.hbase.client.ZKAsyncRegistry.(ZKAsyncRegistry.java:84)\r\n{code}\r\nIt seems that blockUntilConnected() never finished.\r\n","from":"developer"},{"body":"[~Apache9]:\r\nCan you take a look ?","from":"developer"},{"body":"Let me take a look. If no cluster then no doubt we can not connect to zk... It does not make sense to create a zk registry.","from":"developer"},{"body":"OK, the test is used to test we will fail if no cluster.\r\n\r\nMark it as Ignored I'd say. Will be fixed automatically after we remove the blockUntilConnected.\r\n\r\nThanks.","from":"developer"},{"body":"Can variant of blockUntilConnected be used / added which takes into account zookeeper connection issue ?","from":"developer"},{"body":"curator provides: blockUntilConnected(int, TimeUnit)","from":"developer"},{"body":"Yeah there is method\r\n\r\n{code}\r\npublic boolean blockUntilConnected(int maxWaitTime, TimeUnit units) throws InterruptedException;\r\n{code}\r\n\r\nI think 1 or 2 seconds is enough? You can provide a patch, and also add the Thread.interrupt you want.\r\n\r\nThanks.","from":"developer"},{"body":"You made all the previous changes.\r\n\r\nIt would be better you fix it.","from":"developer"},{"body":"Do not have much time. Thanks.","from":"developer"},{"body":"If you too do not have much time, I could find another person to fix this. What do you think? Thanks.","from":"developer"},{"body":"Sounds like you're a manager now.\r\n\r\nThis is high priority issue - without the fix, HBASE-19200 should be reverted.","from":"developer"},{"body":"I always say that, you can do what you think is right, if you want to revert then just revert, a test broken is enough to revert a commit. But I also have my own work plan, and also my judgement on which issue is more important. If you think this one is high priority then just do it.\r\n\r\nThanks.","from":"developer"},{"body":"I have fixed this issue, but I can not assign this issue to myself, if someone can assign it to me?","from":"developer"},{"body":"Assigned to you [~wuguoquan]. Please upload the patch. Thanks.","from":"developer"},{"body":"+1. Let's wait for the pre commit result.","from":"developer"},{"body":"ok, thanks","from":"developer"},{"body":"Where is the Thread.interrupt() call ?\r\n\r\nHow is 2 second interval determined ? Isn't this short for real scenario ?","from":"developer"},{"body":"{quote}\r\nA temporary workaround is to call blockUntilConnected when constructing ZKAsyncRegistry but this is not a production level code. \r\n{quote}\r\nI already said this is not a production level code, why do you keep asking 'what's the implication in production scenario', and 'Isn't this short for real scenario'? I can not get your point.\r\n\r\nThanks.","from":"developer"},{"body":"Pushed to master and branch-2.\r\n\r\nThanks [~wuguoquan] for the contributing.","from":"developer"},{"body":"same to you ","from":"developer"},{"body":"Guoquan:\r\nTestAcidGuarantees started to timeout (again).\r\nCheck flaky test dashbaord:\r\nhttps://builds.apache.org/job/HBASE-Flaky-Tests/23484/","from":"developer"},{"body":"Where is the dashboard? TestAcidGuarantees is not failed in your url?\n\nAnd now TestAcidGuarantees has been promoted to LargeTests, so if it still timeout then there must be other problems. It has been marked as flakey tests long long ago.\n\nThanks.","from":"developer"},{"body":"Guoquan:\r\nDid you run TestAcidGuarantees locally (since patch was attached) ?\r\n\r\nbq. It has been marked as flakey tests long long ago\r\n\r\nI disagree. It got stuck after HBASE-19200 went in.\r\n\r\nFlaky test dashboard is here:\r\nhttps://builds.apache.org/job/HBASE-Find-Flaky-Tests/lastSuccessfulBuild/artifact/dashboard.html\r\n\r\nYou can see clearly the last 5 blue bars which indicate timeout.","from":"developer"},{"body":"bq. I already said this is not a production level code\r\n\r\nhbase-2 is getting into beta phase where dev / QA at various companies have started testing the daily build(s).\r\n\r\nIt would be weird when they find out that what worked yesterday no longer works for the new daily build.\r\n\r\nIn this case, TestAcidGuarantees can be thought of as mini environment which can tell us how robust the change is.\r\nTestAcidGuarantees is a basic test. When developers make change to locking in HRegion, in-memory compaction, etc, they need to know that they are working with non-flaky base.","from":"developer"},{"body":"Then please help finding out the root cause instead of commanding others and repeating useless word again and again. You are not a manager in the community, no one need to accept your command. And even HBASE-19200 is not found by you. I will not reply your comments about this topic any more, it is just wasting of time. You can revert everything you dislike since you are a commiter. I will revert back when I have time to digging in here.","from":"developer"},{"body":"I tried changing the duration passed to blockUntilConnected from 2 to 30.\r\nThere was no improvement in the runtime of TestAcidGuarantees.\r\n\r\nNeed to dig deeper when I have time.","from":"developer"},{"body":"{code}\r\n2017-11-22 13:13:28,358 ERROR [MemStoreFlusher.1] regionserver.HRegion(1210): Asked to modify this region's (TestAcidGuarantees,,1511356384562.67f5ede11bc9d37151794455003c8ee9.) memstoreSize to a negative value which is incorrect. Current memstoreSize=251900, delta=-281670\r\njava.lang.Exception\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.checkNegativeMemStoreDataSize(HRegion.java:1210)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.decrMemStoreSize(HRegion.java:1203)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.internalFlushCacheAndCommit(HRegion.java:2639)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2355)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2327)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:2218)\r\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:498)\r\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:467)\r\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.access$900(MemStoreFlusher.java:69)\r\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher$FlushHandler.run(MemStoreFlusher.java:252)\r\n\tat java.lang.Thread.run(Thread.java:748)\r\n{code}\r\nIt seems the eager policy cause the negative size of memstore, so it breaks the close of table (If the memstore size is negative, the rs will be aborted when closing the region.). We don't close the table before HBASE-19311 hence the issue doesn't hurt the {{TestAcidGuarantees}} previously. As [~Apache9] optimized the {{TestAcidGuarantees}} by creating the mini cluster once and creating/deleting the table for each test case, the issue hurts {{TestAcidGuarantees}} now. I revert the HBASE-19311 locally, and then run the {{TestAcidGuarantees}} with eager policy. The error about the negative size sill happens. We should file a issue to disable the eager policy for {{TestAcidGuarantees}} before we have time to dig in. [~tedyu] [~Apache9] WDYT?","from":"developer"},{"body":"bq. We should file a issue to disable the eager policy for TestAcidGuarantees\r\n\r\nSince you are already working on HBASE-19266, this can be handled there - replacing EAGER with ADAPTIVE policy.","from":"developer"},{"body":"bq. Since you are already working on HBASE-19266, this can be handled there - replacing EAGER with ADAPTIVE policy.\r\nWe still create the {{TestAcidGuarantees}} for EAGER but it will be denoted with {{Ignore}}. Let us discuss them in HBASE-19266. What about closing this issue? [~ted_yu]","from":"developer"}],"created":"2017-11-22T02:42:04.000+0000","description":"From https://builds.apache.org/job/HBASE-Flaky-Tests/23477/testReport/junit/org.apache.hadoop.hbase.client/TestAdmin2/testCheckHBaseAvailableWithoutCluster/ :\r\n{code}\r\norg.junit.runners.model.TestTimedOutException: test timed out after 300000 milliseconds\r\n\tat org.apache.hadoop.hbase.client.TestAdmin2.testCheckHBaseAvailableWithoutCluster(TestAdmin2.java:573)\r\n{code}\r\nIt seems this started hanging after HBASE-19313","issue_id":"13120125","key":"HBASE-19321","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2017-11-22T16:47:58.000+0000","role":"fixed_distractor","summary":"ZKAsyncRegistry ctor would hang when zookeeper cluster is not available"} {"case_id":"12439012","cluster":"DISTRACTOR-HBASE-1933","comments":[{"body":"Can you write the pom and counsel on where we should put them when we release? Thanks.","created":"2009-10-24T20:07:38.864+0000"},{"body":"Yes of course\n\nHere is a sample pom.xml for hbase 0.20.0\n\n\n\n 4.0.0\n org.apache.hadoop\n hbase\n 0.20.0\n\n\nWhen you release hbase you could upload the jar and this pom to the public maven repository. For instance, the hadoop avro project seems to be uploaded on the public repo. You could contact the hadoop avro team to know how do they upload their last release (avro 1.2.0)","created":"2009-10-24T20:59:42.785+0000"},{"body":"How to upload a jar to the Apache directory ?\n\nThe public answer :\nhttp://maven.apache.org/guides/mini/guide-central-repository-upload.html\n\nThe apache internal answer :\nhttp://www.apache.org/dev/repository-faq.html\n\n","created":"2009-10-24T21:40:58.231+0000"},{"body":"Here is a better pom, but you have to add the dependencies tag too:\n\n\n 4.0.0\n org.apache.hadoop\n hbase\n jar\n Hadoop - HBase\n 0.20.0\n HBase is the Hadoop database\n http://hadoop.apache.org/hbase\n \n \n The Apache Software License, Version 2.0\n http://www.apache.org/licenses/LICENSE-2.0.txt\n repo\n \n \n \n http://svn.apache.org/viewcvs.cgi/hadoop/hbase\n \n\n\nDo you plan to migrate from ant to maven ?","created":"2009-10-25T19:52:32.544+0000"},{"body":"There are no plans to migrate to maven, it doesn't generally solve any problem we are having and is complex. If you can create a patch that would add the appropriate ant task to upload to a maven repository, that would be great! Thanks!","created":"2009-10-25T19:56:00.864+0000"},{"body":"+1 on closing this issue as invalid. ","created":"2009-10-25T21:04:42.140+0000"},{"body":"I could help to upload your release to central maven repository.\n\nI could provide a repository (on googlecode for sample) with the hadoop project and then ask for a sync from the central repo and this repo.\n\n\n\nAndrew: If this issue is invalid, feel free to explain here why. thanks.","created":"2009-10-25T22:50:23.218+0000"},{"body":"You can upload artifacts to the m2 repository without using Maven -indeed, Ivy happily pulls them down for your ant builds. \n\nAccordingly, I don' t think Invalid is the right way to close this bugrep.\n\nWhat is important is that the ASF don't let any apache release artifacts depend on any snapshot releases of other apache artifacts, so its only recently (Hadoop 0.21+) that its been possible, with HADOOP-3305. Without a copy of hadoop JARS, there's been little gain from having HBase in there. \n\nAccordingly I would recommend re-opening and marking as dependent on HADOOP-3302.","created":"2009-10-26T12:22:36.297+0000"},{"body":"Sorry, my misunderstanding. The tail of this issue looked like a request to switch the build system from ant to maven. ","created":"2009-10-26T22:10:14.100+0000"},{"body":"This relates to http://jira.codehaus.org/browse/MAVENUPLOAD-2642 (and HADOOP-3302). From Steve:\n\n\"Historically it's not been possible to upload anything because Hadoop\ncore depended on an unreleased snapshot of commons-cli, but as it's\nbeen rolled back to one that has shipped, it should be possible to\nresolve everything. All we need for HBase is to put together a POM\nwith the core dependencies and then put up the artifacts that run\nagainst the not-quite-released Hadoop 0.21 version, which will be the\nfirst one with fully resolvable dependencies.\"","created":"2009-11-24T15:26:22.557+0000"},{"body":"Preferably you would use the repository.apache.org system to get the artifacts synced to central, since they are, well apache artifacts ;-)","created":"2009-11-25T03:38:39.201+0000"},{"body":"@Brian Thanks.","created":"2009-11-25T04:12:10.357+0000"},{"body":"Hadoop is pushing jars to apache repository on successful build. In hadoop build.xml I see a mvn-deploy target that http puts jars to the apache snapshot repo.","created":"2010-01-04T20:13:42.709+0000"},{"body":"Pulling into 0.21 and assigning myself.","created":"2010-01-04T20:14:37.598+0000"},{"body":"Would not this sort of depend on HBASE-1433 to complete the chain. Also - given that thrift and zookeeper do not have a maven repository yet - this might be tricky ? \n","created":"2010-01-04T20:22:17.007+0000"},{"body":"This would be very useful for third-party integration . +1 on the idea. Please make sure that the releases of hbase do not depend on the snapshots of other dependent projects ( that makes debugging / retrieving source / installation a nightmare ). ","created":"2010-01-05T23:49:08.181+0000"},{"body":".bq Please make sure that the releases of hbase do not depend on the snapshots of other dependent projects\n\nHow would you suggest w/ specify the dependency in the pom then?","created":"2010-01-06T05:09:56.856+0000"},{"body":"{quote}\n.bq Please make sure that the releases of hbase do not depend on the snapshots of other dependent projects\n\nHow would you suggest w/ specify the dependency in the pom then?\n{quote}\n\nI meant - \"preferably have releases of hbase dependent on *releases* of other dependent projects \" , so that it easy to keep track of the corresponding source code + bugs if any. \n\n","created":"2010-01-06T05:15:03.763+0000"},{"body":"This issue kind of , depends on the maven artifacts of the other 2 projects (thrift / zookeeper) . So people interested in this - can vote / contribute to the other ones. Zookeeper is going to be available eventually soon (3.3.0). Thrift needs help though. \n\nAssuming that is the case - we should be all set for the publishing of .pom / ivy.xml for hbase . ","created":"2010-01-08T02:01:17.494+0000"},{"body":"I haven't checked yet but maybe we can push our thrift and zk to maven ourselves to the apache maven repo. We can ask the respective projects if its ok w/ them but we could do a manual put w/ a coarse pom?","created":"2010-01-08T04:46:15.287+0000"},{"body":"| I haven't checked yet but maybe we can push our thrift and zk to maven ourselves to the apache maven repo. We can ask the respective projects if its ok w/ them but we could do a manual put w/ a coarse pom?\n\nAs long as the project maintainers are ok with the same - that should work. Caveats:\n\n* If we are going to push it ourselves - then the snapshot release becomes kind of orphaned. With zk - I foresee less issues since they already have their maven integration in place and it is just a question of them integrating with their build process to publish to apache snapshot repository ( like - core, hdfs, mapreduce etc. ). \n\n\n* Thrift is tricky because there is no build process available yet. Even if we push the latest code to snapshot (we can't retroactively push version to maven repository , I guess) , that would mean that we need to look back into hbase about upgrading thrift dependency. It is definitely doable but just to be aware, though. \n","created":"2010-01-08T05:43:25.127+0000"},{"body":"THRIFT-363 is coming along pretty and should be close to getting published soon. That would leave the dependency on zk only and then we should be ready to publish hbase nightly snapshots from build. ","created":"2010-01-17T19:21:53.009+0000"},{"body":"The zk lads are talking about zk published to maven as part of the zk 3.3.0 release. It survived the edit of issues prepping for 3.3.0. Here is their issue on the matter: https://issues.apache.org/jira/browse/ZOOKEEPER-224.","created":"2010-01-17T20:00:40.344+0000"},{"body":"As mentioned in person -THRIFT-363 actually publishes mvn artifacts that can be used with both mvn and ivy. So for now - we can help check to publish mvn repository to a location of our choice and see if we would be able to integrate with thrift. ","created":"2010-01-28T07:18:25.007+0000"},{"body":"Making this a blocker as reminder to self to do this soon, at least thrift.","created":"2010-02-05T07:05:37.640+0000"},{"body":"For artifact publishing of hadoop","created":"2010-03-05T18:40:34.305+0000"},{"body":"'Maven'ization helps this process. ","created":"2010-03-07T00:49:11.956+0000"},{"body":"zk 3.3.0 is going to be released soon, with mvn artifacts., as seen from the latest update to zk ticket. \n\nHBASE-2255 discusses the artifact for the patched version of hadoop (hdfs sync + others) as needed by hbase , that I believe, would eventually be available in a public repository, say - groupId - org.apache.hbase and artifactId - \"hadoop-0.20.2-patched\" , say.. \n\nAs before - THRIFT-363 is pending still. \n\n","created":"2010-03-21T00:44:35.741+0000"},{"body":"zk 3.3.0 now published to mvn repositories. HBASE-2392 takes care of the upgradation. \n\nTHRIFT-363 is checked in now too. So , eventually thrift team will publish their artifacts as well. \n\nSo - all decks cleared for hbase mvn repository publishing here , when it we do a 0.21.0 release. ","created":"2010-03-30T21:31:37.681+0000"},{"body":"@Kay Kay You a have a patch so we can start publishing snapshots now?","created":"2010-04-02T18:55:14.502+0000"},{"body":"@stack: Will look into the patch over the weekend. Given that we are already in maven , I believe it should not be very hard from now. ","created":"2010-04-02T20:06:56.895+0000"},{"body":"In HBASE-2394 we are currently working on using the parent Apache pom which already has the necessary repositories defined as far as I know so it should be easier to publish jars after that has been finished.\n\nJust a heads up to avoid duplicate work.","created":"2010-04-03T17:40:20.310+0000"},{"body":"Thanks Lars for the heads -up. That would definitely address some of the presence of xml tags as needed by the repo publication process. Any idea - how far are we in that process. Let us know.\n\n Otherwise- I was planning to modify the .. with 1 entry for HBase publishing and a writable-user name, that any of you committers can override in ~/.m2/settings.xml to start publishing for snapshots for the moment. \n\nLet me know your thoughts on the same. ","created":"2010-04-09T01:37:44.157+0000"},{"body":".bq Otherwise- I was planning to modify the .. with 1 entry for HBase publishing and a writable-user name, that any of you committers can override in ~/.m2/settings.xml to start publishing for snapshots for the moment.\n\nAbove sounds good to me Kay Kay.","created":"2010-04-09T04:24:25.413+0000"},{"body":"1. Create a settings.xml as outlined in the patch. \n\n2. Place it by default in ~/.m2/settings.xml (preferred) or in a custom location and specify the same as input using the -s option \n\n3. Invoke the deploy target \n\n$ mvn -s settings.xml deploy \n\n\nTarget url-s copied from apache parent pom ( version 7) \n\nPrivate distributionManagement url-s removed, since I believe they were needed for zk / thrift publishing and not needed anymore. \n","created":"2010-04-09T06:00:04.137+0000"},{"body":"| I was planning to modify the .. \n\nOops - meant to say . ","created":"2010-04-09T06:01:02.746+0000"},{"body":"I committed the above patch. Its doing something, publishing something here: https://repository.apache.org/content/repositories/snapshots/org/apache/hbase/\n\nBut something is up still... trying to figure it (w/ Kay Kay's help). Added 'deploy' to hudson build targets. Hopefully hudson user can deploy to mvn repo for us. #2010 build will tell us.","created":"2010-04-12T23:52:23.618+0000"},{"body":"The artifacts published to the apache snapshot repositories looks ok. (hbase-core , hbase-contrib-stargate-core , hbase-contrib-transactional would be needed by most users). \n\nThe patch contains 'apache standard' server ids ( as in the pom.xml of apache ) so hoping that appropriate server ids are set up in the hudson environment, that will 'deploy' . As stack mentioned above - keeping fingers crossed until then. \n\n\n","created":"2010-04-13T00:08:20.657+0000"},{"body":"An username with write access to 'org.apache.hbase' groupId is required for deploy to succeed. INFRA-2609 , raised with nexus/hudson infrastructure to seek best practices and use appropriately. \n\n\n","created":"2010-04-13T00:20:02.737+0000"},{"body":"As per update on INFRA-2609 , the machines should have been configured to publish to all groupId under org.apache.* . So - may be - we are good here for an automatic deploy. \n","created":"2010-04-13T00:41:50.191+0000"},{"body":"Yes. It (hudson deploy) works. \n\n1211 ( http://hudson.zones.apache.org/hudson/view/HBase/job/HBase-Patch/1211/console ) uploads the artifacts successfully after running the tests. Great.\n","created":"2010-04-13T06:22:04.017+0000"},{"body":"Closing. Our successful hudson builds are now dumped into the apache m2 snapshot repository. Thanks to those that made this possible.","created":"2010-04-13T14:55:11.592+0000"},{"body":"0.20.4 HBASE jars missed on http://repo2.maven.org/maven2/org/apache/hadoop/","created":"2010-05-07T09:52:55.157+0000"},{"body":"Sergey, our TRUNK is maven build, not the branch 0.20.4 was made from. We don't have a pom for 0.20 branch. We could hand-publish it I suppose if someone spent some time on a pom.","created":"2010-05-07T14:57:50.617+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T13:01:50.098+0000"}],"conversations":[{"body":"There are many cool release of hadoop hbase and this project is an apache project, as the maven project.\n\nBut the released jars must be download manually and then deploy to a private repository before they can be used by developer using maven2.\n\nPlease could you upload the hbase jars on the public maven2 repository ?\n\nOf course, we can help to deploy those artifact if necessary.\n\n","from":"reporter","subject":"Upload Hbase jars to a public maven repository"},{"body":"Can you write the pom and counsel on where we should put them when we release? Thanks.","from":"developer"},{"body":"Yes of course\n\nHere is a sample pom.xml for hbase 0.20.0\n\n\n\n 4.0.0\n org.apache.hadoop\n hbase\n 0.20.0\n\n\nWhen you release hbase you could upload the jar and this pom to the public maven repository. For instance, the hadoop avro project seems to be uploaded on the public repo. You could contact the hadoop avro team to know how do they upload their last release (avro 1.2.0)","from":"developer"},{"body":"How to upload a jar to the Apache directory ?\n\nThe public answer :\nhttp://maven.apache.org/guides/mini/guide-central-repository-upload.html\n\nThe apache internal answer :\nhttp://www.apache.org/dev/repository-faq.html\n\n","from":"developer"},{"body":"Here is a better pom, but you have to add the dependencies tag too:\n\n\n 4.0.0\n org.apache.hadoop\n hbase\n jar\n Hadoop - HBase\n 0.20.0\n HBase is the Hadoop database\n http://hadoop.apache.org/hbase\n \n \n The Apache Software License, Version 2.0\n http://www.apache.org/licenses/LICENSE-2.0.txt\n repo\n \n \n \n http://svn.apache.org/viewcvs.cgi/hadoop/hbase\n \n\n\nDo you plan to migrate from ant to maven ?","from":"developer"},{"body":"There are no plans to migrate to maven, it doesn't generally solve any problem we are having and is complex. If you can create a patch that would add the appropriate ant task to upload to a maven repository, that would be great! Thanks!","from":"developer"},{"body":"+1 on closing this issue as invalid. ","from":"developer"},{"body":"I could help to upload your release to central maven repository.\n\nI could provide a repository (on googlecode for sample) with the hadoop project and then ask for a sync from the central repo and this repo.\n\n\n\nAndrew: If this issue is invalid, feel free to explain here why. thanks.","from":"developer"},{"body":"You can upload artifacts to the m2 repository without using Maven -indeed, Ivy happily pulls them down for your ant builds. \n\nAccordingly, I don' t think Invalid is the right way to close this bugrep.\n\nWhat is important is that the ASF don't let any apache release artifacts depend on any snapshot releases of other apache artifacts, so its only recently (Hadoop 0.21+) that its been possible, with HADOOP-3305. Without a copy of hadoop JARS, there's been little gain from having HBase in there. \n\nAccordingly I would recommend re-opening and marking as dependent on HADOOP-3302.","from":"developer"},{"body":"Sorry, my misunderstanding. The tail of this issue looked like a request to switch the build system from ant to maven. ","from":"developer"},{"body":"This relates to http://jira.codehaus.org/browse/MAVENUPLOAD-2642 (and HADOOP-3302). From Steve:\n\n\"Historically it's not been possible to upload anything because Hadoop\ncore depended on an unreleased snapshot of commons-cli, but as it's\nbeen rolled back to one that has shipped, it should be possible to\nresolve everything. All we need for HBase is to put together a POM\nwith the core dependencies and then put up the artifacts that run\nagainst the not-quite-released Hadoop 0.21 version, which will be the\nfirst one with fully resolvable dependencies.\"","from":"developer"},{"body":"Preferably you would use the repository.apache.org system to get the artifacts synced to central, since they are, well apache artifacts ;-)","from":"developer"},{"body":"@Brian Thanks.","from":"developer"},{"body":"Hadoop is pushing jars to apache repository on successful build. In hadoop build.xml I see a mvn-deploy target that http puts jars to the apache snapshot repo.","from":"developer"},{"body":"Pulling into 0.21 and assigning myself.","from":"developer"},{"body":"Would not this sort of depend on HBASE-1433 to complete the chain. Also - given that thrift and zookeeper do not have a maven repository yet - this might be tricky ? \n","from":"developer"},{"body":"This would be very useful for third-party integration . +1 on the idea. Please make sure that the releases of hbase do not depend on the snapshots of other dependent projects ( that makes debugging / retrieving source / installation a nightmare ). ","from":"developer"},{"body":".bq Please make sure that the releases of hbase do not depend on the snapshots of other dependent projects\n\nHow would you suggest w/ specify the dependency in the pom then?","from":"developer"},{"body":"{quote}\n.bq Please make sure that the releases of hbase do not depend on the snapshots of other dependent projects\n\nHow would you suggest w/ specify the dependency in the pom then?\n{quote}\n\nI meant - \"preferably have releases of hbase dependent on *releases* of other dependent projects \" , so that it easy to keep track of the corresponding source code + bugs if any. \n\n","from":"developer"},{"body":"This issue kind of , depends on the maven artifacts of the other 2 projects (thrift / zookeeper) . So people interested in this - can vote / contribute to the other ones. Zookeeper is going to be available eventually soon (3.3.0). Thrift needs help though. \n\nAssuming that is the case - we should be all set for the publishing of .pom / ivy.xml for hbase . ","from":"developer"},{"body":"I haven't checked yet but maybe we can push our thrift and zk to maven ourselves to the apache maven repo. We can ask the respective projects if its ok w/ them but we could do a manual put w/ a coarse pom?","from":"developer"},{"body":"| I haven't checked yet but maybe we can push our thrift and zk to maven ourselves to the apache maven repo. We can ask the respective projects if its ok w/ them but we could do a manual put w/ a coarse pom?\n\nAs long as the project maintainers are ok with the same - that should work. Caveats:\n\n* If we are going to push it ourselves - then the snapshot release becomes kind of orphaned. With zk - I foresee less issues since they already have their maven integration in place and it is just a question of them integrating with their build process to publish to apache snapshot repository ( like - core, hdfs, mapreduce etc. ). \n\n\n* Thrift is tricky because there is no build process available yet. Even if we push the latest code to snapshot (we can't retroactively push version to maven repository , I guess) , that would mean that we need to look back into hbase about upgrading thrift dependency. It is definitely doable but just to be aware, though. \n","from":"developer"},{"body":"THRIFT-363 is coming along pretty and should be close to getting published soon. That would leave the dependency on zk only and then we should be ready to publish hbase nightly snapshots from build. ","from":"developer"},{"body":"The zk lads are talking about zk published to maven as part of the zk 3.3.0 release. It survived the edit of issues prepping for 3.3.0. Here is their issue on the matter: https://issues.apache.org/jira/browse/ZOOKEEPER-224.","from":"developer"},{"body":"As mentioned in person -THRIFT-363 actually publishes mvn artifacts that can be used with both mvn and ivy. So for now - we can help check to publish mvn repository to a location of our choice and see if we would be able to integrate with thrift. ","from":"developer"},{"body":"Making this a blocker as reminder to self to do this soon, at least thrift.","from":"developer"},{"body":"For artifact publishing of hadoop","from":"developer"},{"body":"'Maven'ization helps this process. ","from":"developer"},{"body":"zk 3.3.0 is going to be released soon, with mvn artifacts., as seen from the latest update to zk ticket. \n\nHBASE-2255 discusses the artifact for the patched version of hadoop (hdfs sync + others) as needed by hbase , that I believe, would eventually be available in a public repository, say - groupId - org.apache.hbase and artifactId - \"hadoop-0.20.2-patched\" , say.. \n\nAs before - THRIFT-363 is pending still. \n\n","from":"developer"},{"body":"zk 3.3.0 now published to mvn repositories. HBASE-2392 takes care of the upgradation. \n\nTHRIFT-363 is checked in now too. So , eventually thrift team will publish their artifacts as well. \n\nSo - all decks cleared for hbase mvn repository publishing here , when it we do a 0.21.0 release. ","from":"developer"},{"body":"@Kay Kay You a have a patch so we can start publishing snapshots now?","from":"developer"},{"body":"@stack: Will look into the patch over the weekend. Given that we are already in maven , I believe it should not be very hard from now. ","from":"developer"},{"body":"In HBASE-2394 we are currently working on using the parent Apache pom which already has the necessary repositories defined as far as I know so it should be easier to publish jars after that has been finished.\n\nJust a heads up to avoid duplicate work.","from":"developer"},{"body":"Thanks Lars for the heads -up. That would definitely address some of the presence of xml tags as needed by the repo publication process. Any idea - how far are we in that process. Let us know.\n\n Otherwise- I was planning to modify the .. with 1 entry for HBase publishing and a writable-user name, that any of you committers can override in ~/.m2/settings.xml to start publishing for snapshots for the moment. \n\nLet me know your thoughts on the same. ","from":"developer"},{"body":".bq Otherwise- I was planning to modify the .. with 1 entry for HBase publishing and a writable-user name, that any of you committers can override in ~/.m2/settings.xml to start publishing for snapshots for the moment.\n\nAbove sounds good to me Kay Kay.","from":"developer"},{"body":"1. Create a settings.xml as outlined in the patch. \n\n2. Place it by default in ~/.m2/settings.xml (preferred) or in a custom location and specify the same as input using the -s option \n\n3. Invoke the deploy target \n\n$ mvn -s settings.xml deploy \n\n\nTarget url-s copied from apache parent pom ( version 7) \n\nPrivate distributionManagement url-s removed, since I believe they were needed for zk / thrift publishing and not needed anymore. \n","from":"developer"},{"body":"| I was planning to modify the .. \n\nOops - meant to say . ","from":"developer"},{"body":"I committed the above patch. Its doing something, publishing something here: https://repository.apache.org/content/repositories/snapshots/org/apache/hbase/\n\nBut something is up still... trying to figure it (w/ Kay Kay's help). Added 'deploy' to hudson build targets. Hopefully hudson user can deploy to mvn repo for us. #2010 build will tell us.","from":"developer"},{"body":"The artifacts published to the apache snapshot repositories looks ok. (hbase-core , hbase-contrib-stargate-core , hbase-contrib-transactional would be needed by most users). \n\nThe patch contains 'apache standard' server ids ( as in the pom.xml of apache ) so hoping that appropriate server ids are set up in the hudson environment, that will 'deploy' . As stack mentioned above - keeping fingers crossed until then. \n\n\n","from":"developer"},{"body":"An username with write access to 'org.apache.hbase' groupId is required for deploy to succeed. INFRA-2609 , raised with nexus/hudson infrastructure to seek best practices and use appropriately. \n\n\n","from":"developer"},{"body":"As per update on INFRA-2609 , the machines should have been configured to publish to all groupId under org.apache.* . So - may be - we are good here for an automatic deploy. \n","from":"developer"},{"body":"Yes. It (hudson deploy) works. \n\n1211 ( http://hudson.zones.apache.org/hudson/view/HBase/job/HBase-Patch/1211/console ) uploads the artifacts successfully after running the tests. Great.\n","from":"developer"},{"body":"Closing. Our successful hudson builds are now dumped into the apache m2 snapshot repository. Thanks to those that made this possible.","from":"developer"},{"body":"0.20.4 HBASE jars missed on http://repo2.maven.org/maven2/org/apache/hadoop/","from":"developer"},{"body":"Sergey, our TRUNK is maven build, not the branch 0.20.4 was made from. We don't have a pom for 0.20 branch. We could hand-publish it I suppose if someone spent some time on a pom.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2009-10-24T20:01:22.000+0000","description":"There are many cool release of hadoop hbase and this project is an apache project, as the maven project.\n\nBut the released jars must be download manually and then deploy to a private repository before they can be used by developer using maven2.\n\nPlease could you upload the hbase jars on the public maven2 repository ?\n\nOf course, we can help to deploy those artifact if necessary.\n\n","issue_id":"12439012","key":"HBASE-1933","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2010-04-13T14:55:11.000+0000","role":"fixed_distractor","summary":"Upload Hbase jars to a public maven repository"} {"case_id":"13120841","cluster":"DISTRACTOR-HBASE-19349","comments":[{"body":"{code}\r\n \r\n{code}\r\nI saw this comment in pom.xml. This comment seems outdated and can be removed too? [~stack]","created":"2017-11-28T07:04:27.655+0000"},{"body":"Yeah..nothing in dependency tree....that's weird!","created":"2017-11-28T23:10:09.740+0000"},{"body":"I checed the new module hbase-zookeeper and didn't find anything...","created":"2017-11-28T23:28:48.492+0000"},{"body":"My first guess would be, might be coming from some changed hadoop dep where it's not excluded.","created":"2017-11-28T23:36:03.713+0000"},{"body":"Do you take a look about HBASE-19089. It may be related.","created":"2017-11-28T23:40:52.762+0000"},{"body":"It's just fixing missing deps in assembly.\r\nEven if the jar doesn't get included before that change (haven't tested), that change itself won't be the cause, it would be just surfacing the problem present elsewhere. ","created":"2017-11-28T23:48:42.738+0000"},{"body":"bq. that change itself won't be the cause, it would be just surfacing the problem present elsewhere.\r\nOk. I found server-api-2.5.jar in the comment https://issues.apache.org/jira/browse/HBASE-19089?focusedCommentId=16221504&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-16221504 . This can help us find the root cause.","created":"2017-11-28T23:59:02.712+0000"},{"body":"Cool, let me revert that change and see if it's not there. If so, that change can indeed help us find root cause. ","created":"2017-11-29T00:16:52.014+0000"},{"body":"Yup, not present before the change. So it's coming from one of the modules being added in this change.","created":"2017-11-29T00:37:26.695+0000"},{"body":"This seems to do the 'right' thing but it is awful. Running mvn w/ -X it seems like the dependency comes in via hadoop-hdfs test jar. Let me see if I can do better.","created":"2017-12-07T22:54:14.769+0000"},{"body":".002 \r\n \r\n HBASE-19349 Introduce wrong version depencency of servlet-api jar\r\n\r\n Move the hadoop-hdfs guava exclude in modules up to the top pom.\r\n Looks like an exclude in a module is not additive but rather exclusive\r\n blanking out the top level set of exclusions.\r\n\r\n Tested by looking in lib dir of the built tarball.","created":"2017-12-07T23:22:25.397+0000"},{"body":"Assigning [~zghaobac] because he did the hard work figuring which jar was the problem.","created":"2017-12-07T23:23:16.751+0000"},{"body":"bq. Looks like an exclude in a module is not additive but rather exclusive blanking out the top level set of exclusions.\r\n (╯°□°)╯","created":"2017-12-07T23:25:44.484+0000"},{"body":"Some of them had htrace exclusion, what about that?\r\n","created":"2017-12-07T23:26:34.635+0000"},{"body":"This says otherwise: https://stackoverflow.com/questions/10734565/effect-of-overriding-exclusions-in-maven-dependency\r\n\r\n","created":"2017-12-07T23:31:03.017+0000"},{"body":"Thanks' [~appy] Sounds like I'm wrong then (╯°□°)╯ The patch does the right thing though for me (w/o it I see what [~zghaobac] was seeing). On htrace, they were already excluded up in the top-level pom.","created":"2017-12-08T01:07:45.558+0000"},{"body":"Just throwing some perplexing questions in case you found the answers to them:\r\n- Why doesn't it show up in dependency tree?\r\n- It's not from javax.servlet:servlet-api, right? We already exclude that one in top poms. Would be nuts if it's that one.\r\n\r\nThough it feels wrong that we know the absolute reason, patch seems reasonable so +1 on that.","created":"2017-12-08T01:31:41.065+0000"},{"body":"bq. Why doesn't it show up in dependency tree?\r\n\r\nMaven dependency tool seems to have holes and doesn't do a complete job in my experience (dependency:analysis/dependency:tree -- the cleaning of the poms was using these tools but then had to fine-tune by running against jenkins because I couldn't repo its failures locally).\r\n\r\nbq. It's not from javax.servlet:servlet-api, right? We already exclude that one in top poms. Would be nuts if it's that one.\r\n\r\nI don't thing so. I made the association with hadoop-hdfs by adding -X running assembly step.\r\n\r\nThanks for review and questions [~appy] Let me try this patch on clean checkout and if it works, will commit.\r\n\r\n\r\n","created":"2017-12-08T01:41:34.750+0000"},{"body":"Checked it again. Works. Pushed to master and branch-2. Thanks for review [~appy] Thanks for the detective work up front [~zghaobac]","created":"2017-12-08T02:09:38.089+0000"},{"body":"Thanks [~stack] for the patch. Double check by build a tarball based the latest code. servlet-api-2.5.jar is gone. :-)","created":"2017-12-08T13:19:09.150+0000"},{"body":"[~zghaobac] You did the hard part. Thank you.","created":"2017-12-08T17:18:26.017+0000"}],"conversations":[{"body":"Build a tarball.\r\n{code}\r\nmvn -DskipTests clean install && mvn -DskipTests package assembly:single\r\ntar zxvf hbase-2.0.0-beta-1-SNAPSHOT-bin.tar.gz\r\n{code}\r\nThen I found there is a servlet-api-2.5.jar in the lib directory. The right depencency should be javax.servlet-api-3.1.0.jar.\r\n\r\nStart a distributed cluster with this tarball. And got exception when access Master/RS info jsp.\r\n{code}\r\n2017-11-27,10:02:05,066 WARN org.eclipse.jetty.server.HttpChannel: /\r\njava.lang.NoSuchMethodError: javax.servlet.http.HttpServletRequest.isAsyncSupported()Z\r\n at org.eclipse.jetty.server.ResourceService.sendData(ResourceService.java:689)\r\n at org.eclipse.jetty.server.ResourceService.doGet(ResourceService.java:294)\r\n at org.eclipse.jetty.servlet.DefaultServlet.doGet(DefaultServlet.java:458)\r\n at javax.servlet.http.HttpServlet.service(HttpServlet.java:707)\r\n at javax.servlet.http.HttpServlet.service(HttpServlet.java:820)\r\n at org.eclipse.jetty.servlet.ServletHolder.handle(ServletHolder.java:841)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1650)\r\n at org.apache.hadoop.hbase.http.lib.StaticUserWebFilter$StaticUserFilter.doFilter(StaticUserWebFilter.java:113)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1637)\r\n at org.apache.hadoop.hbase.http.ClickjackingPreventionFilter.doFilter(ClickjackingPreventionFilter.java:48)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1637)\r\n at org.apache.hadoop.hbase.http.HttpServer$QuotingInputFilter.doFilter(HttpServer.java:1374)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1637)\r\n at org.apache.hadoop.hbase.http.NoCacheFilter.doFilter(NoCacheFilter.java:49)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1637)\r\n at org.apache.hadoop.hbase.http.NoCacheFilter.doFilter(NoCacheFilter.java:49)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1637)\r\n at org.eclipse.jetty.servlet.ServletHandler.doHandle(ServletHandler.java:533)\r\n{code}\r\n\r\nTry mvn depencency:tree but didn't find why servlet-api-2.5.jar was introduced.\r\n\r\nI download hbase-2.0.0-alpha4-bin.tar.gz and didn't find servlet-api-2.5.jar. And build a tar from hbase-2.0.0-alpha4-src.tar.gz and didn't find servlet-api-2.5.jar, too. So this may be introduced by recently commits. And should fix this when release 2.0.0-beta1.\r\n","from":"reporter","subject":"Introduce wrong version depencency of servlet-api jar"},{"body":"{code}\r\n \r\n{code}\r\nI saw this comment in pom.xml. This comment seems outdated and can be removed too? [~stack]","from":"developer"},{"body":"Yeah..nothing in dependency tree....that's weird!","from":"developer"},{"body":"I checed the new module hbase-zookeeper and didn't find anything...","from":"developer"},{"body":"My first guess would be, might be coming from some changed hadoop dep where it's not excluded.","from":"developer"},{"body":"Do you take a look about HBASE-19089. It may be related.","from":"developer"},{"body":"It's just fixing missing deps in assembly.\r\nEven if the jar doesn't get included before that change (haven't tested), that change itself won't be the cause, it would be just surfacing the problem present elsewhere. ","from":"developer"},{"body":"bq. that change itself won't be the cause, it would be just surfacing the problem present elsewhere.\r\nOk. I found server-api-2.5.jar in the comment https://issues.apache.org/jira/browse/HBASE-19089?focusedCommentId=16221504&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-16221504 . This can help us find the root cause.","from":"developer"},{"body":"Cool, let me revert that change and see if it's not there. If so, that change can indeed help us find root cause. ","from":"developer"},{"body":"Yup, not present before the change. So it's coming from one of the modules being added in this change.","from":"developer"},{"body":"This seems to do the 'right' thing but it is awful. Running mvn w/ -X it seems like the dependency comes in via hadoop-hdfs test jar. Let me see if I can do better.","from":"developer"},{"body":".002 \r\n \r\n HBASE-19349 Introduce wrong version depencency of servlet-api jar\r\n\r\n Move the hadoop-hdfs guava exclude in modules up to the top pom.\r\n Looks like an exclude in a module is not additive but rather exclusive\r\n blanking out the top level set of exclusions.\r\n\r\n Tested by looking in lib dir of the built tarball.","from":"developer"},{"body":"Assigning [~zghaobac] because he did the hard work figuring which jar was the problem.","from":"developer"},{"body":"bq. Looks like an exclude in a module is not additive but rather exclusive blanking out the top level set of exclusions.\r\n (╯°□°)╯","from":"developer"},{"body":"Some of them had htrace exclusion, what about that?\r\n","from":"developer"},{"body":"This says otherwise: https://stackoverflow.com/questions/10734565/effect-of-overriding-exclusions-in-maven-dependency\r\n\r\n","from":"developer"},{"body":"Thanks' [~appy] Sounds like I'm wrong then (╯°□°)╯ The patch does the right thing though for me (w/o it I see what [~zghaobac] was seeing). On htrace, they were already excluded up in the top-level pom.","from":"developer"},{"body":"Just throwing some perplexing questions in case you found the answers to them:\r\n- Why doesn't it show up in dependency tree?\r\n- It's not from javax.servlet:servlet-api, right? We already exclude that one in top poms. Would be nuts if it's that one.\r\n\r\nThough it feels wrong that we know the absolute reason, patch seems reasonable so +1 on that.","from":"developer"},{"body":"bq. Why doesn't it show up in dependency tree?\r\n\r\nMaven dependency tool seems to have holes and doesn't do a complete job in my experience (dependency:analysis/dependency:tree -- the cleaning of the poms was using these tools but then had to fine-tune by running against jenkins because I couldn't repo its failures locally).\r\n\r\nbq. It's not from javax.servlet:servlet-api, right? We already exclude that one in top poms. Would be nuts if it's that one.\r\n\r\nI don't thing so. I made the association with hadoop-hdfs by adding -X running assembly step.\r\n\r\nThanks for review and questions [~appy] Let me try this patch on clean checkout and if it works, will commit.\r\n\r\n\r\n","from":"developer"},{"body":"Checked it again. Works. Pushed to master and branch-2. Thanks for review [~appy] Thanks for the detective work up front [~zghaobac]","from":"developer"},{"body":"Thanks [~stack] for the patch. Double check by build a tarball based the latest code. servlet-api-2.5.jar is gone. :-)","from":"developer"},{"body":"[~zghaobac] You did the hard part. Thank you.","from":"developer"}],"created":"2017-11-27T06:40:28.000+0000","description":"Build a tarball.\r\n{code}\r\nmvn -DskipTests clean install && mvn -DskipTests package assembly:single\r\ntar zxvf hbase-2.0.0-beta-1-SNAPSHOT-bin.tar.gz\r\n{code}\r\nThen I found there is a servlet-api-2.5.jar in the lib directory. The right depencency should be javax.servlet-api-3.1.0.jar.\r\n\r\nStart a distributed cluster with this tarball. And got exception when access Master/RS info jsp.\r\n{code}\r\n2017-11-27,10:02:05,066 WARN org.eclipse.jetty.server.HttpChannel: /\r\njava.lang.NoSuchMethodError: javax.servlet.http.HttpServletRequest.isAsyncSupported()Z\r\n at org.eclipse.jetty.server.ResourceService.sendData(ResourceService.java:689)\r\n at org.eclipse.jetty.server.ResourceService.doGet(ResourceService.java:294)\r\n at org.eclipse.jetty.servlet.DefaultServlet.doGet(DefaultServlet.java:458)\r\n at javax.servlet.http.HttpServlet.service(HttpServlet.java:707)\r\n at javax.servlet.http.HttpServlet.service(HttpServlet.java:820)\r\n at org.eclipse.jetty.servlet.ServletHolder.handle(ServletHolder.java:841)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1650)\r\n at org.apache.hadoop.hbase.http.lib.StaticUserWebFilter$StaticUserFilter.doFilter(StaticUserWebFilter.java:113)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1637)\r\n at org.apache.hadoop.hbase.http.ClickjackingPreventionFilter.doFilter(ClickjackingPreventionFilter.java:48)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1637)\r\n at org.apache.hadoop.hbase.http.HttpServer$QuotingInputFilter.doFilter(HttpServer.java:1374)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1637)\r\n at org.apache.hadoop.hbase.http.NoCacheFilter.doFilter(NoCacheFilter.java:49)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1637)\r\n at org.apache.hadoop.hbase.http.NoCacheFilter.doFilter(NoCacheFilter.java:49)\r\n at org.eclipse.jetty.servlet.ServletHandler$CachedChain.doFilter(ServletHandler.java:1637)\r\n at org.eclipse.jetty.servlet.ServletHandler.doHandle(ServletHandler.java:533)\r\n{code}\r\n\r\nTry mvn depencency:tree but didn't find why servlet-api-2.5.jar was introduced.\r\n\r\nI download hbase-2.0.0-alpha4-bin.tar.gz and didn't find servlet-api-2.5.jar. And build a tar from hbase-2.0.0-alpha4-src.tar.gz and didn't find servlet-api-2.5.jar, too. So this may be introduced by recently commits. And should fix this when release 2.0.0-beta1.\r\n","issue_id":"13120841","key":"HBASE-19349","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2017-12-08T02:09:38.000+0000","role":"fixed_distractor","summary":"Introduce wrong version depencency of servlet-api jar"} {"case_id":"12441098","cluster":"DISTRACTOR-HBASE-1989","comments":[{"body":"I'll take this one. Redoing HCD any ways to go against ZK instead of as a Writable so can change it then.","created":"2009-11-19T04:58:30.814+0000"},{"body":"I did a small part of this before I found this issue. I only changed HBaseAdmin - this patch deprecates the old methods and changes their usage.\n\nBut as I said: This is only a small part of the renaming process but perhaps it is useful anyway","created":"2010-02-12T15:48:07.433+0000"},{"body":"This is a start. Will I commit it as a start in on this issue, as a part 1? I'm working on the other part over in master rewrite where HColumnDescriptor becomes instead ColumnFamilyDefinition, etc., throughout the code base.","created":"2010-02-12T19:56:32.855+0000"},{"body":"Marking as patch available. Has work started by LarsF.","created":"2011-03-31T15:07:07.014+0000"},{"body":"Stale","created":"2012-05-30T22:36:43.789+0000"},{"body":"Wow, this is from the way-back machine (2009). But I still think it's a good idea from a terminology standpoint to clear this up in the API","created":"2012-05-30T23:12:16.003+0000"},{"body":"Stale issue. Reopen if still relevant (if there's new activity)","created":"2014-06-08T21:54:40.038+0000"},{"body":"I'd like to work on this again because the naming is still confusing. I'll reopen and attach a patch that'll deprecate the \"old\" versions in 2.0.0 so they can be removed in 3.0.0.","created":"2015-04-30T08:21:28.390+0000"},{"body":"Attaching a patch that fixes all methods in {{Admin}} and {{HBaseAdmin}}.\n\nThis does not change {{HColumnDescriptor}} as that'd take too long for now. If I have time I'll do that in another issue.","created":"2015-05-05T20:22:37.724+0000"},{"body":"bq.void deleteColumnFamily(final TableName tableName, final byte[] columnName) throws IOException\nCan you make the param name as 'columnFamily '\n\nHBaseAdmin.java\n{quote}\naddColumnFamily(final byte[] tableName, HColumnDescriptor columnFamily)\naddColumnFamily(final String tableName, HColumnDescriptor columnFamily)\n{quote}\nMay be no need to add newly. Just deprecate its counterparts with replacement as \naddColumnFamily(final TableName tableName, final HColumnDescriptor columnFamily)\n(?)\nSame case applicable for delete/modify","created":"2015-05-06T08:59:32.616+0000"},{"body":"Thanks for taking a look and the comments.\n\nI've attached v1 of the patch that should fix the Checkstyle warning and addresses your comments. I've renamed all parameters I could find in those two classes to {{columnFamily}} and I've removed those extra methods I added.","created":"2015-05-06T12:57:02.211+0000"},{"body":"Latest patch LGTM","created":"2015-05-06T14:55:20.216+0000"},{"body":"[~anoop.hbase] Thanks for taking a look. Jenkins agrees with you. I think this is good to commit if there are no further comments.","created":"2015-05-09T08:23:50.124+0000"},{"body":"Pushed to master. Thanks [~lars_francke] Nice.","created":"2015-05-11T03:48:34.284+0000"},{"body":"this got pushed without a proper commit message (missing jira reference). any objection to revert + reapply with commit?","created":"2015-05-11T05:38:22.486+0000"},{"body":"Thanks stack!\n\n[~busbey] I do not object, thank you for catching this.","created":"2015-05-11T06:12:22.917+0000"},{"body":"Reverted because of bad commit message (Thanks again [~busbey]) and then reapplied.","created":"2015-05-11T16:47:36.148+0000"}],"conversations":[{"body":"Consider the classes Admin and HColumnDescriptor.\n\nHColumnDescriptor is really referring to a \"column family\" and not a \"column\" (i.e., family:qualifer).\n\nLikewise, in Admin there is a method called \"addColumn\" that takes an HColumnDescriptor instance.\n\nI labeled this a bug in the sense that it produces conceptual confusion because there is a big difference between a column and column-family in HBase and these terms should be used consistently. The code works, though.\n \n\n","from":"reporter","subject":"Admin (et al.) not accurate with Column vs. Column-Family usage"},{"body":"I'll take this one. Redoing HCD any ways to go against ZK instead of as a Writable so can change it then.","from":"developer"},{"body":"I did a small part of this before I found this issue. I only changed HBaseAdmin - this patch deprecates the old methods and changes their usage.\n\nBut as I said: This is only a small part of the renaming process but perhaps it is useful anyway","from":"developer"},{"body":"This is a start. Will I commit it as a start in on this issue, as a part 1? I'm working on the other part over in master rewrite where HColumnDescriptor becomes instead ColumnFamilyDefinition, etc., throughout the code base.","from":"developer"},{"body":"Marking as patch available. Has work started by LarsF.","from":"developer"},{"body":"Stale","from":"developer"},{"body":"Wow, this is from the way-back machine (2009). But I still think it's a good idea from a terminology standpoint to clear this up in the API","from":"developer"},{"body":"Stale issue. Reopen if still relevant (if there's new activity)","from":"developer"},{"body":"I'd like to work on this again because the naming is still confusing. I'll reopen and attach a patch that'll deprecate the \"old\" versions in 2.0.0 so they can be removed in 3.0.0.","from":"developer"},{"body":"Attaching a patch that fixes all methods in {{Admin}} and {{HBaseAdmin}}.\n\nThis does not change {{HColumnDescriptor}} as that'd take too long for now. If I have time I'll do that in another issue.","from":"developer"},{"body":"bq.void deleteColumnFamily(final TableName tableName, final byte[] columnName) throws IOException\nCan you make the param name as 'columnFamily '\n\nHBaseAdmin.java\n{quote}\naddColumnFamily(final byte[] tableName, HColumnDescriptor columnFamily)\naddColumnFamily(final String tableName, HColumnDescriptor columnFamily)\n{quote}\nMay be no need to add newly. Just deprecate its counterparts with replacement as \naddColumnFamily(final TableName tableName, final HColumnDescriptor columnFamily)\n(?)\nSame case applicable for delete/modify","from":"developer"},{"body":"Thanks for taking a look and the comments.\n\nI've attached v1 of the patch that should fix the Checkstyle warning and addresses your comments. I've renamed all parameters I could find in those two classes to {{columnFamily}} and I've removed those extra methods I added.","from":"developer"},{"body":"Latest patch LGTM","from":"developer"},{"body":"[~anoop.hbase] Thanks for taking a look. Jenkins agrees with you. I think this is good to commit if there are no further comments.","from":"developer"},{"body":"Pushed to master. Thanks [~lars_francke] Nice.","from":"developer"},{"body":"this got pushed without a proper commit message (missing jira reference). any objection to revert + reapply with commit?","from":"developer"},{"body":"Thanks stack!\n\n[~busbey] I do not object, thank you for catching this.","from":"developer"},{"body":"Reverted because of bad commit message (Thanks again [~busbey]) and then reapplied.","from":"developer"}],"created":"2009-11-18T21:27:49.000+0000","description":"Consider the classes Admin and HColumnDescriptor.\n\nHColumnDescriptor is really referring to a \"column family\" and not a \"column\" (i.e., family:qualifer).\n\nLikewise, in Admin there is a method called \"addColumn\" that takes an HColumnDescriptor instance.\n\nI labeled this a bug in the sense that it produces conceptual confusion because there is a big difference between a column and column-family in HBase and these terms should be used consistently. The code works, though.\n \n\n","issue_id":"12441098","key":"HBASE-1989","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2015-05-11T16:47:36.000+0000","role":"fixed_distractor","summary":"Admin (et al.) not accurate with Column vs. Column-Family usage"} {"case_id":"13159462","cluster":"DISTRACTOR-HBASE-20587","comments":[{"body":"Argh, missed that hbase-client still has a reference to Jackson's JsonMapper.","created":"2018-05-15T16:36:35.132+0000"},{"body":"Looks like we can't rip out Jackson from the hbase-client as all operations (e.g. Get, Put) use it for implementing {{toString()}}.\r\n\r\nCan still pull it out of hbase-common and put that server-side in hbase-http.","created":"2018-05-15T16:41:57.134+0000"},{"body":"You mean our JsonMapper has a reference to Jackson's ObjectMapper.","created":"2018-05-15T16:42:28.181+0000"},{"body":"should we shade jackson in hbase-thirdparty?","created":"2018-05-15T16:44:44.251+0000"},{"body":"{quote}You mean our JsonMapper has a reference to Jackson's ObjectMapper.\r\n{quote}\r\nYes, thank you :)\r\n{quote}should we shade jackson in hbase-thirdparty?\r\n{quote}\r\nJust had that same thought over on HBASE-20852. I'm not sure if there's a reason we haven't yet shaded it, but that would help.","created":"2018-05-15T16:46:54.230+0000"},{"body":"Let me use this issue to pull Jackson completely out of the client path and replace it with the gson we have in hbase-thirdparty (big thanks to [~Apache9] for that pointer).","created":"2018-05-16T18:45:52.696+0000"},{"body":"Any updates here?","created":"2018-06-19T13:56:44.816+0000"},{"body":"Nah, haven't gotten around to this yet, Duo. Bumped to 2.2.0. Thanks for checking.","created":"2018-06-20T04:16:46.575+0000"},{"body":"[~elserj] Any updates here, sir? If not, will move this to 2.3.0. Thanks.","created":"2019-02-13T03:19:45.959+0000"},{"body":"Moved it out. Need to come back :)","created":"2019-02-13T14:10:06.180+0000"},{"body":"I can work on this if you do not mind [~elserj].","created":"2019-02-20T03:40:02.199+0000"},{"body":"Our rest framework also uses jackson and the replace work is not straight-forward so do not include it. Can do it in another issue I think.","created":"2019-02-20T10:53:59.842+0000"},{"body":"No qualms here, thanks for picking it up!\r\n\r\nYour first patch looks OK, but seems like might have caused some of those test failures. I'm sure you're on that already.\r\n\r\nI see you were fixing some other style/static-analysis things while you were in some files. Bloats the patch a little, but probably better long term. Thanks for having an eye out for them.","created":"2019-02-20T15:15:12.944+0000"},{"body":"Let's land HBASE-21939 first, and then in this issue we can completely remove the jackson dependencies, not only for client.","created":"2019-02-21T00:38:25.573+0000"},{"body":"Give up on HBASE-21939 for now as we also use jackson for xml...\r\n\r\nLet's resolve this first...","created":"2019-02-21T02:40:28.806+0000"},{"body":"Review board link:\r\n\r\nhttps://reviews.apache.org/r/70029/","created":"2019-02-22T02:09:28.303+0000"},{"body":"OK seems the SuppressWarnings can not be added to a local var... Let me just remove it since it is not a problem. Can purge these warnings together in another issue I think.","created":"2019-02-22T08:15:45.696+0000"},{"body":"Pushed to branch-2.2+.\r\n\r\nThanks all for reviewing.","created":"2019-02-22T09:06:32.790+0000"}],"conversations":[{"body":"HBASE-20582 got me looking at how we use Jackson. It appears that we moved some JSON code from hbase-server into hbase-common via HBASE-19053. But, there seems to be no good reason why this code should live there and not in hbase-http instead. Keeping Jackson off the user's classpath is a nice goal.\r\n\r\nFYI [~appy], [~mdrob]","from":"reporter","subject":"Replace Jackson with shaded thirdparty gson"},{"body":"Argh, missed that hbase-client still has a reference to Jackson's JsonMapper.","from":"developer"},{"body":"Looks like we can't rip out Jackson from the hbase-client as all operations (e.g. Get, Put) use it for implementing {{toString()}}.\r\n\r\nCan still pull it out of hbase-common and put that server-side in hbase-http.","from":"developer"},{"body":"You mean our JsonMapper has a reference to Jackson's ObjectMapper.","from":"developer"},{"body":"should we shade jackson in hbase-thirdparty?","from":"developer"},{"body":"{quote}You mean our JsonMapper has a reference to Jackson's ObjectMapper.\r\n{quote}\r\nYes, thank you :)\r\n{quote}should we shade jackson in hbase-thirdparty?\r\n{quote}\r\nJust had that same thought over on HBASE-20852. I'm not sure if there's a reason we haven't yet shaded it, but that would help.","from":"developer"},{"body":"Let me use this issue to pull Jackson completely out of the client path and replace it with the gson we have in hbase-thirdparty (big thanks to [~Apache9] for that pointer).","from":"developer"},{"body":"Any updates here?","from":"developer"},{"body":"Nah, haven't gotten around to this yet, Duo. Bumped to 2.2.0. Thanks for checking.","from":"developer"},{"body":"[~elserj] Any updates here, sir? If not, will move this to 2.3.0. Thanks.","from":"developer"},{"body":"Moved it out. Need to come back :)","from":"developer"},{"body":"I can work on this if you do not mind [~elserj].","from":"developer"},{"body":"Our rest framework also uses jackson and the replace work is not straight-forward so do not include it. Can do it in another issue I think.","from":"developer"},{"body":"No qualms here, thanks for picking it up!\r\n\r\nYour first patch looks OK, but seems like might have caused some of those test failures. I'm sure you're on that already.\r\n\r\nI see you were fixing some other style/static-analysis things while you were in some files. Bloats the patch a little, but probably better long term. Thanks for having an eye out for them.","from":"developer"},{"body":"Let's land HBASE-21939 first, and then in this issue we can completely remove the jackson dependencies, not only for client.","from":"developer"},{"body":"Give up on HBASE-21939 for now as we also use jackson for xml...\r\n\r\nLet's resolve this first...","from":"developer"},{"body":"Review board link:\r\n\r\nhttps://reviews.apache.org/r/70029/","from":"developer"},{"body":"OK seems the SuppressWarnings can not be added to a local var... Let me just remove it since it is not a problem. Can purge these warnings together in another issue I think.","from":"developer"},{"body":"Pushed to branch-2.2+.\r\n\r\nThanks all for reviewing.","from":"developer"}],"created":"2018-05-15T16:34:04.000+0000","description":"HBASE-20582 got me looking at how we use Jackson. It appears that we moved some JSON code from hbase-server into hbase-common via HBASE-19053. But, there seems to be no good reason why this code should live there and not in hbase-http instead. Keeping Jackson off the user's classpath is a nice goal.\r\n\r\nFYI [~appy], [~mdrob]","issue_id":"13159462","key":"HBASE-20587","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2019-02-22T09:06:32.000+0000","role":"fixed_distractor","summary":"Replace Jackson with shaded thirdparty gson"} {"case_id":"12444359","cluster":"DISTRACTOR-HBASE-2077","comments":[{"body":"jdcryans helped narrow it down to an issue in Leases where the lease would polled from the queue at the same time it was being renewed. since there is no sychronization protection there is a race condition. I am attaching a patch that should fix the problem though it is very difficult to reproduce.\n","created":"2009-12-30T07:11:19.595+0000"},{"body":"Patch to fix race condition","created":"2009-12-30T07:17:05.840+0000"},{"body":"I ran the client test and it passes. Committed to branch and trunk. Made Sam a new contributor, thanks for the patch!","created":"2009-12-30T07:40:13.266+0000"},{"body":"Patch looks good to me.","created":"2009-12-30T17:47:15.710+0000"},{"body":"This is not fixed in 0.20.3, even tho the patch got in we still get the error:\n\n{code}\n2010-02-22 19:18:06,638 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Scanner -4405675591371793964 lease expired\n2010-02-22 19:18:06,639 ERROR org.apache.hadoop.hbase.regionserver.HRegionServer: \njava.lang.NullPointerException\n\tat org.apache.hadoop.hbase.KeyValue$KVComparator.compare(KeyValue.java:1310)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:136)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:127)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:117)\n\tat java.util.PriorityQueue.siftDownUsingComparator(PriorityQueue.java:641)\n\tat java.util.PriorityQueue.siftDown(PriorityQueue.java:612)\n\tat java.util.PriorityQueue.poll(PriorityQueue.java:523)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap.next(KeyValueHeap.java:113)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.nextInternal(HRegion.java:1807)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.next(HRegion.java:1771)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.next(HRegionServer.java:1894)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:657)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:915)\n{code}","created":"2010-02-23T01:38:42.255+0000"},{"body":"The way we are using the DelayQueue looks broken. We wait for leaseCheckFrequency (or the delay time, whichever happens first) in Leases and then we just expire the scanner. I think there's times when we really expire scanners that aren't done yet. This patch adds a check in run() to verify the element is really expired.","created":"2010-02-23T02:46:44.211+0000"},{"body":"Moving in 0.20.4 and marking critical.","created":"2010-02-23T02:47:24.183+0000"},{"body":"Not sure whether this is related, but we've seen similar behaviour:\n\nWe've seen several of these exact stack traces (on 0.20.4-dev) which were accompanied by an immediate region server shutdown. At start it seemed that this error causes the shutdown but in every case, after a thorough examination of the logs, we also found a zookeeper session timeout. Eventually we discovered it was a faulty disk causing large delays at unpredictable times.\n\nIn our cases we also had messages like:\n2010-02-05 16:44:20,821 WARN org.apache.hadoop.hbase.util.Sleeper: We slept 63152ms, ten times longer than scheduled: 1000\nand\n2010-02-05 16:44:26,448 ERROR org.apache.hadoop.hbase.regionserver.HRegionServer: ZooKeeper session expired\n\nI can provide more details if this is relevant\n","created":"2010-02-23T11:54:58.358+0000"},{"body":"@Yoram\n\nThe NPEs are enough to kill a region servers, what you pasted is a GC pause that took more than 1 minute. But, I see how that could happen... If for some reason the user scans more than 1 row (scanner caching), and that during the scan the thread is paused for some reason, there's a good possibility that on the next iteration in HRS.next() the lease has already been expired. \n\nI would love to see a log of such an event.","created":"2010-02-23T16:32:12.312+0000"},{"body":"I've applied the patch to 0.20.3, and things look like they are working fine for me now. The same map/reduce job which crashed with the above stack trace now completes without error.\nIn my case, I have an expensive reduce phase, which is running on 1701037 rows input. It looks like this was taking so long to complete that the scanner timeout problem was happening.","created":"2010-02-23T19:24:14.514+0000"},{"body":"@Ian, that's good news.\n\nAlso I got Yoram's logs off-list, I'll take a look tomorrow and see if there's anything that's in the scope of this jira.","created":"2010-02-24T03:39:45.760+0000"},{"body":"WRT Yoram's logs, it went like I thought it was eg a scan begins before the GC pause and when it ends the lease is expired and around the same time a scanner tries to finish. Does it make sense to handle such a case? From a client perspective, a pause of 1 minute probably already timed it out.\n\nI will go forward and commit this patch as it fixed Ian's problem.","created":"2010-02-25T00:37:46.297+0000"},{"body":"Second patch committed to branch and trunk.","created":"2010-02-25T01:26:02.533+0000"},{"body":"Further testing shows that my patch in fact introduced a bug that was keeping the leases opened for a long time. Reopening to see if we can dig deeper in Ian's issue.","created":"2010-02-26T00:56:45.187+0000"},{"body":"So I'm thinking about taking a new approach to this bug. Since the major problem here is that we need knowledge of any client using a scanner (or anything else lease-related), I think we should add a new AtomicInteger inside Leases.Lease and increment it every time a user renews that lease. When you are down with the lease-related action, you decrease the AtomicInteger. This protects us from any GC pause happening while, for example, a scanner is next'ing 100 rows and gets a 60 secs pause right in the middle.","created":"2010-03-31T18:23:46.160+0000"},{"body":"Patch that implements my latest idea. The user has the choice of keeping track of the lease usage or not. Passes the tests that currently pass.","created":"2010-03-31T20:43:40.738+0000"},{"body":"What if you renewed the lease on entry and on the way out as insurance\nagainst long-running 'next' invocation? Would that be 'cleaner'?\n\n\nOn Wed, Mar 31, 2010 at 1:45 PM, Jean-Daniel Cryans (JIRA)\n","created":"2010-04-02T18:18:29.760+0000"},{"body":"How would it be cleaner?","created":"2010-04-05T20:08:58.544+0000"},{"body":"No new long that you increment/decrement and keep an account of.\n\nThe method name stays as renewLease rather than talk about increment/decrement (\"Why do I have to do increment/decrement on a lease when all I'm interested in is lease renewal\"). ","created":"2010-04-05T20:50:44.035+0000"},{"body":"bq. The method name stays as renewLease rather than talk about increment/decrement (\"Why do I have to do increment/decrement on a lease when all I'm interested in is lease renewal\").\n\nThe bulk of the issue is about not timing out a lease that someone currently uses, whether there's a GC or not. To be certain we have acquired the lease during the whole operation, unless we renew the lease after each line, I don't see how we can insure the same level of safety that my patch offers.\n\nThis patch also allows multiple users to share the lease if it's needed (hence incrementing/decrementing).","created":"2010-04-06T04:53:16.402+0000"},{"body":".bq This patch also allows multiple users to share the lease if it's needed (hence incrementing/decrementing).\n\nThis seems perverse to me. When would such a usecase make sense?\n\n.bq To be certain we have acquired the lease during the whole operation, unless we renew the lease after each line, I don't see how we can insure the same level of safety that my patch offers.\n\nOk. Makes sense that while the scanner is inside the server, then the lease moves to a different 'state'. My suggestion doesn't cover case of our timing out because of GC while scanner is a server-side resident. \n\nWhy not remove the lease on entry and then renew it on the way out? (IMO, the increment/decrement semantic is confusing).\n","created":"2010-04-06T05:22:58.390+0000"},{"body":"Let's punt this to 0.20.5 then","created":"2010-04-08T18:21:54.198+0000"},{"body":"Marking these as fixed against 0.21.0 rather than against 0.20.5.","created":"2010-05-12T23:52:25.384+0000"},{"body":"What's the status of this patch in 0.20 branch? It seems it was committed then reverted, but the revert didn't actually do a full revert (left a new getExpirationTime() method in there). We're seeing the issue on 0.20.4. Does anyone have thoughts on a good fix?","created":"2010-05-29T02:53:31.526+0000"},{"body":"The first patch was committed, the second reverted (maybe I forgot something in there tho). HBASE-2503 plays around the same part of the code, and I'm pretty sure it fixes the NPE but I can't tell for sure since to trip on it you need a heavily GCing region server.\n\nSo it should be fixed in 0.20.5, it would be awesome if you can confirm.","created":"2010-05-29T17:03:27.107+0000"},{"body":"core issue: no concurrency control between next() calls and lease timeouts. While nice to have, this is a corner case and can't hold up 0.90","created":"2010-10-05T21:35:53.967+0000"},{"body":"Here is a suggestion where we remove lease from leases while we are processing a request then on the way out in a finally we renew lease.","created":"2011-05-20T20:57:34.480+0000"},{"body":"Ahemm.. this is a version that actually works (TestFromClientSide is a good test for this change).","created":"2011-05-21T04:27:41.063+0000"},{"body":"+1 on latest patch, I like it.","created":"2011-06-21T20:40:18.456+0000"},{"body":"Applied 2077-v4.txt to branch and trunk as a 'part2' on this issue. I'm now closing this since its gone all over the place. We've not seen Sam's original issue in a while and it'll look different in current codebase; lets open new issue then.","created":"2011-06-21T20:58:12.642+0000"},{"body":"Yeah the NPEs are gone for me but I was still able to easily get myself into a situation that triggered expired leases by just scanning a table. This patch will solve that nicely.","created":"2011-06-21T21:02:32.776+0000"},{"body":"This is long since committed, but just a request:\n\nIn the future could we open separate JIRAs rather than doing a \"part 2\" when the commits are more than a day apart? It's very difficult to figure out what went on in the history of this JIRA, since it was committed for 0.20 in Dec '09, briefly amended in Feb '10, amendation partially reverted the next day, and then another change in Jun '11 for 0.90.4 to solve an entirely different bug than the description indicates. This makes it very difficult to support past branches or maintain distributions, since it appears this was fixed long ago but in fact 0.90.3 lacks a major part of the JIRA.","created":"2011-08-09T20:07:39.098+0000"},{"body":"Sorry Todd. Will be better going forward.","created":"2011-08-09T20:35:16.929+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T13:01:44.615+0000"}],"conversations":[{"body":"2009-12-29 18:05:55,432 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Scanner -4250070597157694417 lease expired\n2009-12-29 18:05:55,443 ERROR org.apache.hadoop.hbase.regionserver.HRegionServer: \njava.lang.NullPointerException\n\tat org.apache.hadoop.hbase.KeyValue$KVComparator.compare(KeyValue.java:1310)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:136)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:127)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:117)\n\tat java.util.PriorityQueue.siftDownUsingComparator(PriorityQueue.java:641)\n\tat java.util.PriorityQueue.siftDown(PriorityQueue.java:612)\n\tat java.util.PriorityQueue.poll(PriorityQueue.java:523)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap.next(KeyValueHeap.java:113)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.nextInternal(HRegion.java:1776)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.next(HRegion.java:1719)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.next(HRegionServer.java:1944)\n\tat sun.reflect.GeneratedMethodAccessor13.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:648)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:915)\n2009-12-29 18:05:55,446 INFO org.apache.hadoop.ipc.HBaseServer: IPC Server handler 7 on 55260, call next(-4250070597157694417, 10000) from 192.168.1.90:54011: error: java.io.IOException: java.lang.NullPointerException\njava.io.IOException: java.lang.NullPointerException\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.convertThrowableToIOE(HRegionServer.java:869)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.convertThrowableToIOE(HRegionServer.java:859)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.next(HRegionServer.java:1965)\n\tat sun.reflect.GeneratedMethodAccessor13.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:648)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:915)\nCaused by: java.lang.NullPointerException\n\tat org.apache.hadoop.hbase.KeyValue$KVComparator.compare(KeyValue.java:1310)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:136)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:127)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:117)\n\tat java.util.PriorityQueue.siftDownUsingComparator(PriorityQueue.java:641)\n\tat java.util.PriorityQueue.siftDown(PriorityQueue.java:612)\n\tat java.util.PriorityQueue.poll(PriorityQueue.java:523)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap.next(KeyValueHeap.java:113)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.nextInternal(HRegion.java:1776)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.next(HRegion.java:1719)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.next(HRegionServer.java:1944)\n\t... 5 more\n2009-12-29 18:05:55,447 WARN org.apache.hadoop.ipc.HBaseServer: IPC Server Responder, call next(-4250070597157694417, 10000) from 192.168.1.90:54011: output error\n2009-12-29 18:05:55,448 INFO org.apache.hadoop.ipc.HBaseServer: IPC Server handler 7 on 55260 caught: java.nio.channels.ClosedChannelException\n\tat sun.nio.ch.SocketChannelImpl.ensureWriteOpen(SocketChannelImpl.java:126)\n\tat sun.nio.ch.SocketChannelImpl.write(SocketChannelImpl.java:324)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer.channelWrite(HBaseServer.java:1125)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Responder.processResponse(HBaseServer.java:615)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Responder.doRespond(HBaseServer.java:679)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:943)\n\n2009-12-29 18:05:56,322 INFO org.apache.hadoop.ipc.HBaseServer: Stopping server on 55260\n2009-12-29 18:05:56,322 INFO org.apache.hadoop.ipc.HBaseServer: Stopping IPC Server listener on 55260\n","from":"reporter","subject":"NullPointerException with an open scanner that expired causing an immediate region server shutdown"},{"body":"jdcryans helped narrow it down to an issue in Leases where the lease would polled from the queue at the same time it was being renewed. since there is no sychronization protection there is a race condition. I am attaching a patch that should fix the problem though it is very difficult to reproduce.\n","from":"developer"},{"body":"Patch to fix race condition","from":"developer"},{"body":"I ran the client test and it passes. Committed to branch and trunk. Made Sam a new contributor, thanks for the patch!","from":"developer"},{"body":"Patch looks good to me.","from":"developer"},{"body":"This is not fixed in 0.20.3, even tho the patch got in we still get the error:\n\n{code}\n2010-02-22 19:18:06,638 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Scanner -4405675591371793964 lease expired\n2010-02-22 19:18:06,639 ERROR org.apache.hadoop.hbase.regionserver.HRegionServer: \njava.lang.NullPointerException\n\tat org.apache.hadoop.hbase.KeyValue$KVComparator.compare(KeyValue.java:1310)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:136)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:127)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:117)\n\tat java.util.PriorityQueue.siftDownUsingComparator(PriorityQueue.java:641)\n\tat java.util.PriorityQueue.siftDown(PriorityQueue.java:612)\n\tat java.util.PriorityQueue.poll(PriorityQueue.java:523)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap.next(KeyValueHeap.java:113)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.nextInternal(HRegion.java:1807)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.next(HRegion.java:1771)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.next(HRegionServer.java:1894)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:657)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:915)\n{code}","from":"developer"},{"body":"The way we are using the DelayQueue looks broken. We wait for leaseCheckFrequency (or the delay time, whichever happens first) in Leases and then we just expire the scanner. I think there's times when we really expire scanners that aren't done yet. This patch adds a check in run() to verify the element is really expired.","from":"developer"},{"body":"Moving in 0.20.4 and marking critical.","from":"developer"},{"body":"Not sure whether this is related, but we've seen similar behaviour:\n\nWe've seen several of these exact stack traces (on 0.20.4-dev) which were accompanied by an immediate region server shutdown. At start it seemed that this error causes the shutdown but in every case, after a thorough examination of the logs, we also found a zookeeper session timeout. Eventually we discovered it was a faulty disk causing large delays at unpredictable times.\n\nIn our cases we also had messages like:\n2010-02-05 16:44:20,821 WARN org.apache.hadoop.hbase.util.Sleeper: We slept 63152ms, ten times longer than scheduled: 1000\nand\n2010-02-05 16:44:26,448 ERROR org.apache.hadoop.hbase.regionserver.HRegionServer: ZooKeeper session expired\n\nI can provide more details if this is relevant\n","from":"developer"},{"body":"@Yoram\n\nThe NPEs are enough to kill a region servers, what you pasted is a GC pause that took more than 1 minute. But, I see how that could happen... If for some reason the user scans more than 1 row (scanner caching), and that during the scan the thread is paused for some reason, there's a good possibility that on the next iteration in HRS.next() the lease has already been expired. \n\nI would love to see a log of such an event.","from":"developer"},{"body":"I've applied the patch to 0.20.3, and things look like they are working fine for me now. The same map/reduce job which crashed with the above stack trace now completes without error.\nIn my case, I have an expensive reduce phase, which is running on 1701037 rows input. It looks like this was taking so long to complete that the scanner timeout problem was happening.","from":"developer"},{"body":"@Ian, that's good news.\n\nAlso I got Yoram's logs off-list, I'll take a look tomorrow and see if there's anything that's in the scope of this jira.","from":"developer"},{"body":"WRT Yoram's logs, it went like I thought it was eg a scan begins before the GC pause and when it ends the lease is expired and around the same time a scanner tries to finish. Does it make sense to handle such a case? From a client perspective, a pause of 1 minute probably already timed it out.\n\nI will go forward and commit this patch as it fixed Ian's problem.","from":"developer"},{"body":"Second patch committed to branch and trunk.","from":"developer"},{"body":"Further testing shows that my patch in fact introduced a bug that was keeping the leases opened for a long time. Reopening to see if we can dig deeper in Ian's issue.","from":"developer"},{"body":"So I'm thinking about taking a new approach to this bug. Since the major problem here is that we need knowledge of any client using a scanner (or anything else lease-related), I think we should add a new AtomicInteger inside Leases.Lease and increment it every time a user renews that lease. When you are down with the lease-related action, you decrease the AtomicInteger. This protects us from any GC pause happening while, for example, a scanner is next'ing 100 rows and gets a 60 secs pause right in the middle.","from":"developer"},{"body":"Patch that implements my latest idea. The user has the choice of keeping track of the lease usage or not. Passes the tests that currently pass.","from":"developer"},{"body":"What if you renewed the lease on entry and on the way out as insurance\nagainst long-running 'next' invocation? Would that be 'cleaner'?\n\n\nOn Wed, Mar 31, 2010 at 1:45 PM, Jean-Daniel Cryans (JIRA)\n","from":"developer"},{"body":"How would it be cleaner?","from":"developer"},{"body":"No new long that you increment/decrement and keep an account of.\n\nThe method name stays as renewLease rather than talk about increment/decrement (\"Why do I have to do increment/decrement on a lease when all I'm interested in is lease renewal\"). ","from":"developer"},{"body":"bq. The method name stays as renewLease rather than talk about increment/decrement (\"Why do I have to do increment/decrement on a lease when all I'm interested in is lease renewal\").\n\nThe bulk of the issue is about not timing out a lease that someone currently uses, whether there's a GC or not. To be certain we have acquired the lease during the whole operation, unless we renew the lease after each line, I don't see how we can insure the same level of safety that my patch offers.\n\nThis patch also allows multiple users to share the lease if it's needed (hence incrementing/decrementing).","from":"developer"},{"body":".bq This patch also allows multiple users to share the lease if it's needed (hence incrementing/decrementing).\n\nThis seems perverse to me. When would such a usecase make sense?\n\n.bq To be certain we have acquired the lease during the whole operation, unless we renew the lease after each line, I don't see how we can insure the same level of safety that my patch offers.\n\nOk. Makes sense that while the scanner is inside the server, then the lease moves to a different 'state'. My suggestion doesn't cover case of our timing out because of GC while scanner is a server-side resident. \n\nWhy not remove the lease on entry and then renew it on the way out? (IMO, the increment/decrement semantic is confusing).\n","from":"developer"},{"body":"Let's punt this to 0.20.5 then","from":"developer"},{"body":"Marking these as fixed against 0.21.0 rather than against 0.20.5.","from":"developer"},{"body":"What's the status of this patch in 0.20 branch? It seems it was committed then reverted, but the revert didn't actually do a full revert (left a new getExpirationTime() method in there). We're seeing the issue on 0.20.4. Does anyone have thoughts on a good fix?","from":"developer"},{"body":"The first patch was committed, the second reverted (maybe I forgot something in there tho). HBASE-2503 plays around the same part of the code, and I'm pretty sure it fixes the NPE but I can't tell for sure since to trip on it you need a heavily GCing region server.\n\nSo it should be fixed in 0.20.5, it would be awesome if you can confirm.","from":"developer"},{"body":"core issue: no concurrency control between next() calls and lease timeouts. While nice to have, this is a corner case and can't hold up 0.90","from":"developer"},{"body":"Here is a suggestion where we remove lease from leases while we are processing a request then on the way out in a finally we renew lease.","from":"developer"},{"body":"Ahemm.. this is a version that actually works (TestFromClientSide is a good test for this change).","from":"developer"},{"body":"+1 on latest patch, I like it.","from":"developer"},{"body":"Applied 2077-v4.txt to branch and trunk as a 'part2' on this issue. I'm now closing this since its gone all over the place. We've not seen Sam's original issue in a while and it'll look different in current codebase; lets open new issue then.","from":"developer"},{"body":"Yeah the NPEs are gone for me but I was still able to easily get myself into a situation that triggered expired leases by just scanning a table. This patch will solve that nicely.","from":"developer"},{"body":"This is long since committed, but just a request:\n\nIn the future could we open separate JIRAs rather than doing a \"part 2\" when the commits are more than a day apart? It's very difficult to figure out what went on in the history of this JIRA, since it was committed for 0.20 in Dec '09, briefly amended in Feb '10, amendation partially reverted the next day, and then another change in Jun '11 for 0.90.4 to solve an entirely different bug than the description indicates. This makes it very difficult to support past branches or maintain distributions, since it appears this was fixed long ago but in fact 0.90.3 lacks a major part of the JIRA.","from":"developer"},{"body":"Sorry Todd. Will be better going forward.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2009-12-30T07:08:45.000+0000","description":"2009-12-29 18:05:55,432 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Scanner -4250070597157694417 lease expired\n2009-12-29 18:05:55,443 ERROR org.apache.hadoop.hbase.regionserver.HRegionServer: \njava.lang.NullPointerException\n\tat org.apache.hadoop.hbase.KeyValue$KVComparator.compare(KeyValue.java:1310)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:136)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:127)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:117)\n\tat java.util.PriorityQueue.siftDownUsingComparator(PriorityQueue.java:641)\n\tat java.util.PriorityQueue.siftDown(PriorityQueue.java:612)\n\tat java.util.PriorityQueue.poll(PriorityQueue.java:523)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap.next(KeyValueHeap.java:113)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.nextInternal(HRegion.java:1776)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.next(HRegion.java:1719)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.next(HRegionServer.java:1944)\n\tat sun.reflect.GeneratedMethodAccessor13.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:648)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:915)\n2009-12-29 18:05:55,446 INFO org.apache.hadoop.ipc.HBaseServer: IPC Server handler 7 on 55260, call next(-4250070597157694417, 10000) from 192.168.1.90:54011: error: java.io.IOException: java.lang.NullPointerException\njava.io.IOException: java.lang.NullPointerException\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.convertThrowableToIOE(HRegionServer.java:869)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.convertThrowableToIOE(HRegionServer.java:859)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.next(HRegionServer.java:1965)\n\tat sun.reflect.GeneratedMethodAccessor13.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:648)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:915)\nCaused by: java.lang.NullPointerException\n\tat org.apache.hadoop.hbase.KeyValue$KVComparator.compare(KeyValue.java:1310)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:136)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:127)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap$KVScannerComparator.compare(KeyValueHeap.java:117)\n\tat java.util.PriorityQueue.siftDownUsingComparator(PriorityQueue.java:641)\n\tat java.util.PriorityQueue.siftDown(PriorityQueue.java:612)\n\tat java.util.PriorityQueue.poll(PriorityQueue.java:523)\n\tat org.apache.hadoop.hbase.regionserver.KeyValueHeap.next(KeyValueHeap.java:113)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.nextInternal(HRegion.java:1776)\n\tat org.apache.hadoop.hbase.regionserver.HRegion$RegionScanner.next(HRegion.java:1719)\n\tat org.apache.hadoop.hbase.regionserver.HRegionServer.next(HRegionServer.java:1944)\n\t... 5 more\n2009-12-29 18:05:55,447 WARN org.apache.hadoop.ipc.HBaseServer: IPC Server Responder, call next(-4250070597157694417, 10000) from 192.168.1.90:54011: output error\n2009-12-29 18:05:55,448 INFO org.apache.hadoop.ipc.HBaseServer: IPC Server handler 7 on 55260 caught: java.nio.channels.ClosedChannelException\n\tat sun.nio.ch.SocketChannelImpl.ensureWriteOpen(SocketChannelImpl.java:126)\n\tat sun.nio.ch.SocketChannelImpl.write(SocketChannelImpl.java:324)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer.channelWrite(HBaseServer.java:1125)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Responder.processResponse(HBaseServer.java:615)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Responder.doRespond(HBaseServer.java:679)\n\tat org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:943)\n\n2009-12-29 18:05:56,322 INFO org.apache.hadoop.ipc.HBaseServer: Stopping server on 55260\n2009-12-29 18:05:56,322 INFO org.apache.hadoop.ipc.HBaseServer: Stopping IPC Server listener on 55260\n","issue_id":"12444359","key":"HBASE-2077","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2011-06-21T20:58:12.000+0000","role":"fixed_distractor","summary":"NullPointerException with an open scanner that expired causing an immediate region server shutdown"} {"case_id":"13178096","cluster":"DISTRACTOR-HBASE-21032","comments":[{"body":"Thanks for this report Andrey [~timoha], so, do you expect the partial results could be more?","created":"2018-08-10T03:39:09.675+0000"},{"body":"Yes, it should always try to fit MaxResultSize of Cells into a partial ScanResponse. You can try out the difference by running the code I provided on hbase 2.0.0 (will return only ~2-3 results) and hbase 2.1.0 returns ~260 ScanResponses (2X that if you account for heartbeats).","created":"2018-08-10T16:26:39.307+0000"},{"body":"[~timoha]\r\n\r\nthis is caused by HBASE-20457. This patch set read type to *STREAM* when a limit is reached. (patch applied to 2.1.0 but 2.0.0)\r\n\r\n#checkTimeLimit will always return true after this since \"returnImmediately\" is set to true.\r\n\r\n ","created":"2018-08-12T11:13:50.057+0000"},{"body":"[~timoha]\r\n\r\nFound a solution for you.\r\n\r\nJust need to add this line:\r\n\r\n*scan.setReadType(ReadType.PREAD);*\r\n\r\nthen the behavior of hbase-2.1 will be the same as 2.0. Basically setting read type to PREAD explicitly can prevent read type from switching to STREAM automatically when conditions met. ","created":"2018-08-12T22:38:51.493+0000"},{"body":"ping [~Apache9], would you mind taking a look at this phenomenon.","created":"2018-08-13T02:40:45.191+0000"},{"body":"Will take a look.","created":"2018-08-13T02:42:11.979+0000"},{"body":"The ScannerContext is per rpc call, so theoretically the returnImmediately will only take effect once as it should lead to a immediately return of the rpc call, and next time the ScannerContext will be recreated and returnImmediately will be false.\r\n\r\nLet me see what's the problem here.","created":"2018-08-13T03:02:13.771+0000"},{"body":"OK, I think the problem is that you use the same startRow and stopRow. Let me check what's broken here.","created":"2018-08-13T03:52:46.206+0000"},{"body":"The fix is simple, we should use the stored readType field instead of get it from the Scan object. The UT is more important, so give the credit to [~timoha].","created":"2018-08-20T09:29:20.039+0000"},{"body":"+1 from me. Reattaching for a rerun while [~Apache9] is offline to see if failure related or flakey.","created":"2018-08-20T19:55:37.471+0000"},{"body":"Timeout failure, not related i think.\r\n*{color:green}+1{color}*\r\n\r\n[~timoha], do you want to try applying the patch for confirmation?\r\n","created":"2018-08-21T03:00:19.901+0000"},{"body":"Sure, giving it a try by applying to 5a40eae63e290c8a12b1e7d4dd01fc98ba09573d of branch-2.1 and deploying onto our dev.","created":"2018-08-21T16:58:16.621+0000"},{"body":"I'm hitting another issue I've had before with 2.1 when upgrading dev cluster. HBase master is stuck initializing because it can't get to hbase:namespace region:\r\n\r\n{{18/08/21 18:52:28 INFO client.RpcRetryingCallerImpl: Call exception, tries=32, retries=46, started=471534 ms ago, cancelled=false, msg=org.apache.hadoop.hbase.NotServingRegionException: hbase:namespace,,1508805323559.2ad4d95f7b9d9ba0d746b8da50a7f9a7. is not online on regionserver-0,16020,1534877035908}}\r\n{{    at org.apache.hadoop.hbase.regionserver.HRegionServer.getRegionByEncodedName(HRegionServer.java:3287)}}\r\n{{    at org.apache.hadoop.hbase.regionserver.HRegionServer.getRegion(HRegionServer.java:3264)}}\r\n{{    at org.apache.hadoop.hbase.regionserver.RSRpcServices.getRegion(RSRpcServices.java:1428)}}\r\n{{    at org.apache.hadoop.hbase.regionserver.RSRpcServices.get(RSRpcServices.java:2443)}}\r\n{{    at org.apache.hadoop.hbase.shaded.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:41998)}}\r\n{{    at org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:409)}}\r\n{{    at org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:130)}}\r\n{{    at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:324)}}\r\n{{    at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:304)}}\r\n{{, details=row 'default' on table 'hbase:namespace' at region=hbase:namespace,,1508805323559.2ad4d95f7b9d9ba0d746b8da50a7f9a7., hostname=regionserver-0,16020,1534538897512, seqNum=172}}\r\n\r\n \r\n\r\nAnd there's nothing about hbase:namespace in regionserver-0's logs. Don't really know how to work around this.","created":"2018-08-21T18:57:10.750+0000"},{"body":"Is your issue what is on the tail of HBASE-20671 [~timoha]? Thanks.","created":"2018-08-21T19:01:50.204+0000"},{"body":"Even though we don't use read replicas, it seems like it's the same issue. The last row loaded from hbase:meta was:\r\n\r\n18/08/21 19:46:21 INFO assignment.RegionStateStore: Load hbase:meta entry region=2ad4d95f7b9d9ba0d746b8da50a7f9a7, regionState=OPEN, lastHost=regionserver-0,16020,1534538897512, regionLocation=regionserver-3,16020,1533765879856, openSeqNum=172\r\n\r\nWhich, based on region encoded name, is hbase:namespace. However, the regionserver is different.\r\n\r\nI think I'm going to halt deploying the patched version for now. Instead, I've run App.java that I've attached and it looks like it behaves as expected.","created":"2018-08-21T20:24:56.931+0000"},{"body":"Ok. Will commit this. There'll be a separate issue on the namespace issue. Thanks [~timoha]","created":"2018-08-21T20:28:10.419+0000"},{"body":"Pushed to branch-2.1+. Thanks for patch [~timoha] and [~Apache9]","created":"2018-08-21T20:32:42.601+0000"}],"conversations":[{"body":"I have a long row with a bunch of columns that I'm scanning with setAllowPartialResults(true). In the response I'm getting the first partial ScanResponse being around 2MB with multiple cells while all of the consequent ones being 1 cell per ScanResponse. After digging more, I found that each of those single cell ScanResponse partials are preceded by a heartbeat (zero cells). This results in two requests per cell to a regionserver.\r\n\r\nI've attached code to reproduce it on hbase version 2.1.0 (it works as expected on 2.0.0 and 2.0.1).\r\n\r\n[^App.java]\r\n\r\nI'm fairly certain it's a serverside issue as [gohbase|https://github.com/tsuna/gohbase] client is having the same issue. I have not tried to reproduce this with multi-row scan.","from":"reporter","subject":"ScanResponses contain only one cell each"},{"body":"Thanks for this report Andrey [~timoha], so, do you expect the partial results could be more?","from":"developer"},{"body":"Yes, it should always try to fit MaxResultSize of Cells into a partial ScanResponse. You can try out the difference by running the code I provided on hbase 2.0.0 (will return only ~2-3 results) and hbase 2.1.0 returns ~260 ScanResponses (2X that if you account for heartbeats).","from":"developer"},{"body":"[~timoha]\r\n\r\nthis is caused by HBASE-20457. This patch set read type to *STREAM* when a limit is reached. (patch applied to 2.1.0 but 2.0.0)\r\n\r\n#checkTimeLimit will always return true after this since \"returnImmediately\" is set to true.\r\n\r\n ","from":"developer"},{"body":"[~timoha]\r\n\r\nFound a solution for you.\r\n\r\nJust need to add this line:\r\n\r\n*scan.setReadType(ReadType.PREAD);*\r\n\r\nthen the behavior of hbase-2.1 will be the same as 2.0. Basically setting read type to PREAD explicitly can prevent read type from switching to STREAM automatically when conditions met. ","from":"developer"},{"body":"ping [~Apache9], would you mind taking a look at this phenomenon.","from":"developer"},{"body":"Will take a look.","from":"developer"},{"body":"The ScannerContext is per rpc call, so theoretically the returnImmediately will only take effect once as it should lead to a immediately return of the rpc call, and next time the ScannerContext will be recreated and returnImmediately will be false.\r\n\r\nLet me see what's the problem here.","from":"developer"},{"body":"OK, I think the problem is that you use the same startRow and stopRow. Let me check what's broken here.","from":"developer"},{"body":"The fix is simple, we should use the stored readType field instead of get it from the Scan object. The UT is more important, so give the credit to [~timoha].","from":"developer"},{"body":"+1 from me. Reattaching for a rerun while [~Apache9] is offline to see if failure related or flakey.","from":"developer"},{"body":"Timeout failure, not related i think.\r\n*{color:green}+1{color}*\r\n\r\n[~timoha], do you want to try applying the patch for confirmation?\r\n","from":"developer"},{"body":"Sure, giving it a try by applying to 5a40eae63e290c8a12b1e7d4dd01fc98ba09573d of branch-2.1 and deploying onto our dev.","from":"developer"},{"body":"I'm hitting another issue I've had before with 2.1 when upgrading dev cluster. HBase master is stuck initializing because it can't get to hbase:namespace region:\r\n\r\n{{18/08/21 18:52:28 INFO client.RpcRetryingCallerImpl: Call exception, tries=32, retries=46, started=471534 ms ago, cancelled=false, msg=org.apache.hadoop.hbase.NotServingRegionException: hbase:namespace,,1508805323559.2ad4d95f7b9d9ba0d746b8da50a7f9a7. is not online on regionserver-0,16020,1534877035908}}\r\n{{    at org.apache.hadoop.hbase.regionserver.HRegionServer.getRegionByEncodedName(HRegionServer.java:3287)}}\r\n{{    at org.apache.hadoop.hbase.regionserver.HRegionServer.getRegion(HRegionServer.java:3264)}}\r\n{{    at org.apache.hadoop.hbase.regionserver.RSRpcServices.getRegion(RSRpcServices.java:1428)}}\r\n{{    at org.apache.hadoop.hbase.regionserver.RSRpcServices.get(RSRpcServices.java:2443)}}\r\n{{    at org.apache.hadoop.hbase.shaded.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:41998)}}\r\n{{    at org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:409)}}\r\n{{    at org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:130)}}\r\n{{    at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:324)}}\r\n{{    at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:304)}}\r\n{{, details=row 'default' on table 'hbase:namespace' at region=hbase:namespace,,1508805323559.2ad4d95f7b9d9ba0d746b8da50a7f9a7., hostname=regionserver-0,16020,1534538897512, seqNum=172}}\r\n\r\n \r\n\r\nAnd there's nothing about hbase:namespace in regionserver-0's logs. Don't really know how to work around this.","from":"developer"},{"body":"Is your issue what is on the tail of HBASE-20671 [~timoha]? Thanks.","from":"developer"},{"body":"Even though we don't use read replicas, it seems like it's the same issue. The last row loaded from hbase:meta was:\r\n\r\n18/08/21 19:46:21 INFO assignment.RegionStateStore: Load hbase:meta entry region=2ad4d95f7b9d9ba0d746b8da50a7f9a7, regionState=OPEN, lastHost=regionserver-0,16020,1534538897512, regionLocation=regionserver-3,16020,1533765879856, openSeqNum=172\r\n\r\nWhich, based on region encoded name, is hbase:namespace. However, the regionserver is different.\r\n\r\nI think I'm going to halt deploying the patched version for now. Instead, I've run App.java that I've attached and it looks like it behaves as expected.","from":"developer"},{"body":"Ok. Will commit this. There'll be a separate issue on the namespace issue. Thanks [~timoha]","from":"developer"},{"body":"Pushed to branch-2.1+. Thanks for patch [~timoha] and [~Apache9]","from":"developer"}],"created":"2018-08-09T18:33:29.000+0000","description":"I have a long row with a bunch of columns that I'm scanning with setAllowPartialResults(true). In the response I'm getting the first partial ScanResponse being around 2MB with multiple cells while all of the consequent ones being 1 cell per ScanResponse. After digging more, I found that each of those single cell ScanResponse partials are preceded by a heartbeat (zero cells). This results in two requests per cell to a regionserver.\r\n\r\nI've attached code to reproduce it on hbase version 2.1.0 (it works as expected on 2.0.0 and 2.0.1).\r\n\r\n[^App.java]\r\n\r\nI'm fairly certain it's a serverside issue as [gohbase|https://github.com/tsuna/gohbase] client is having the same issue. I have not tried to reproduce this with multi-row scan.","issue_id":"13178096","key":"HBASE-21032","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2018-08-21T20:32:42.000+0000","role":"fixed_distractor","summary":"ScanResponses contain only one cell each"} {"case_id":"13182328","cluster":"DISTRACTOR-HBASE-21135","comments":[{"body":"Initially I tried to do replacements in beanshell condition itself. But it seems it is too late to do so as the enforcer fails to evaluate the beanshell syntax evben then. Attached an initial patch [^HBASE-21135.master.001.patch] which replaces all backslashes (if any, say in windows path) with forward slash via a plugin. ","created":"2018-08-31T10:46:44.283+0000"},{"body":"With patch on Windows:\r\n{noformat}\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Reactor Summary:\r\n[INFO]\r\n[INFO] Apache HBase ....................................... SUCCESS [ 3.441 s]\r\n[INFO] Apache HBase - Checkstyle .......................... SUCCESS [ 1.043 s]\r\n[INFO] Apache HBase - Build Support ....................... SUCCESS [ 0.057 s]\r\n[INFO] Apache HBase - Error Prone Rules ................... SUCCESS [ 1.385 s]\r\n[INFO] Apache HBase - Annotations ......................... SUCCESS [ 0.614 s]\r\n[INFO] Apache HBase - Build Configuration ................. SUCCESS [ 0.128 s]\r\n[INFO] Apache HBase - Shaded Protocol ..................... SUCCESS [ 34.409 s]\r\n[INFO] Apache HBase - Common .............................. SUCCESS [ 24.600 s]\r\n[INFO] Apache HBase - Metrics API ......................... SUCCESS [ 2.099 s]\r\n[INFO] Apache HBase - Hadoop Compatibility ................ SUCCESS [ 2.799 s]\r\n[INFO] Apache HBase - Metrics Implementation .............. SUCCESS [ 2.513 s]\r\n[INFO] Apache HBase - Hadoop Two Compatibility ............ SUCCESS [ 3.540 s]\r\n[INFO] Apache HBase - Protocol ............................ SUCCESS [ 11.765 s]\r\n[INFO] Apache HBase - Client .............................. SUCCESS [ 12.878 s]\r\n[INFO] Apache HBase - Zookeeper ........................... SUCCESS [ 3.142 s]\r\n[INFO] Apache HBase - Replication ......................... SUCCESS [ 3.655 s]\r\n[INFO] Apache HBase - Resource Bundle ..................... SUCCESS [ 0.328 s]\r\n[INFO] Apache HBase - HTTP ................................ SUCCESS [ 6.054 s]\r\n[INFO] Apache HBase - Procedure ........................... SUCCESS [ 3.555 s]\r\n[INFO] Apache HBase - Server .............................. SUCCESS [ 41.261 s]\r\n[INFO] Apache HBase - MapReduce ........................... SUCCESS [ 10.958 s]\r\n[INFO] Apache HBase - Testing Util ........................ SUCCESS [ 4.142 s]\r\n[INFO] Apache HBase - Thrift .............................. SUCCESS [ 12.591 s]\r\n[INFO] Apache HBase - RSGroup ............................. SUCCESS [ 7.119 s]\r\n[INFO] Apache HBase - Shell ............................... SUCCESS [ 4.232 s]\r\n[INFO] Apache HBase - Coprocessor Endpoint ................ SUCCESS [ 7.678 s]\r\n[INFO] Apache HBase - Integration Tests ................... SUCCESS [ 6.985 s]\r\n[INFO] Apache HBase - Rest ................................ SUCCESS [ 9.629 s]\r\n[INFO] Apache HBase - Examples ............................ SUCCESS [ 6.832 s]\r\n[INFO] Apache HBase - Shaded .............................. SUCCESS [ 0.272 s]\r\n[INFO] Apache HBase - Shaded - Client (with Hadoop bundled) SUCCESS [ 1.451 s]\r\n[INFO] Apache HBase - Shaded - Client ..................... SUCCESS [ 1.466 s]\r\n[INFO] Apache HBase - Shaded - MapReduce .................. SUCCESS [ 4.534 s]\r\n[INFO] Apache HBase - External Block Cache ................ SUCCESS [ 2.963 s]\r\n[INFO] Apache HBase - Assembly ............................ SUCCESS [ 13.044 s]\r\n[INFO] Apache HBase Shaded Packaging Invariants ........... SUCCESS [ 14.888 s]\r\n[INFO] Apache HBase Shaded Packaging Invariants (with Hadoop bundled) SUCCESS [ 12.077 s]\r\n[INFO] Apache HBase - Archetypes .......................... SUCCESS [ 0.121 s]\r\n[INFO] Apache HBase - Exemplar for hbase-client archetype . SUCCESS [ 3.492 s]\r\n[INFO] Apache HBase - Exemplar for hbase-shaded-client archetype SUCCESS [ 1.708 s]\r\n[INFO] Apache HBase - Archetype builder ................... SUCCESS [ 0.665 s]\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD SUCCESS\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 04:48 min\r\n[INFO] Finished at: 2018-08-31T16:11:25+05:30\r\n[INFO] Final Memory: 229M/1474M\r\n[INFO] ------------------------------------------------------------------------\r\n{noformat}\r\n\r\n","created":"2018-08-31T10:47:33.449+0000"},{"body":"With patch on Ubuntu:\r\n\r\n{code:java}\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Reactor Summary:\r\n[INFO] \r\n[INFO] Apache HBase ....................................... SUCCESS [ 2.909 s]\r\n[INFO] Apache HBase - Checkstyle .......................... SUCCESS [ 0.744 s]\r\n[INFO] Apache HBase - Build Support ....................... SUCCESS [ 0.059 s]\r\n[INFO] Apache HBase - Error Prone Rules ................... SUCCESS [ 0.772 s]\r\n[INFO] Apache HBase - Annotations ......................... SUCCESS [ 0.463 s]\r\n[INFO] Apache HBase - Build Configuration ................. SUCCESS [ 0.104 s]\r\n[INFO] Apache HBase - Shaded Protocol ..................... SUCCESS [ 29.708 s]\r\n[INFO] Apache HBase - Common .............................. SUCCESS [ 6.303 s]\r\n[INFO] Apache HBase - Metrics API ......................... SUCCESS [ 1.378 s]\r\n[INFO] Apache HBase - Hadoop Compatibility ................ SUCCESS [ 1.703 s]\r\n[INFO] Apache HBase - Metrics Implementation .............. SUCCESS [ 1.256 s]\r\n[INFO] Apache HBase - Hadoop Two Compatibility ............ SUCCESS [ 1.935 s]\r\n[INFO] Apache HBase - Protocol ............................ SUCCESS [ 7.504 s]\r\n[INFO] Apache HBase - Client .............................. SUCCESS [ 7.923 s]\r\n[INFO] Apache HBase - Zookeeper ........................... SUCCESS [ 1.697 s]\r\n[INFO] Apache HBase - Replication ......................... SUCCESS [ 1.949 s]\r\n[INFO] Apache HBase - Resource Bundle ..................... SUCCESS [ 0.192 s]\r\n[INFO] Apache HBase - HTTP ................................ SUCCESS [ 4.039 s]\r\n[INFO] Apache HBase - Procedure ........................... SUCCESS [ 2.151 s]\r\n[INFO] Apache HBase - Server .............................. SUCCESS [ 26.699 s]\r\n[INFO] Apache HBase - MapReduce ........................... SUCCESS [ 5.589 s]\r\n[INFO] Apache HBase - Testing Util ........................ SUCCESS [ 2.789 s]\r\n[INFO] Apache HBase - Thrift .............................. SUCCESS [ 5.217 s]\r\n[INFO] Apache HBase - RSGroup ............................. SUCCESS [ 2.605 s]\r\n[INFO] Apache HBase - Shell ............................... SUCCESS [ 1.540 s]\r\n[INFO] Apache HBase - Coprocessor Endpoint ................ SUCCESS [ 2.338 s]\r\n[INFO] Apache HBase - Backup .............................. SUCCESS [ 1.895 s]\r\n[INFO] Apache HBase - Integration Tests ................... SUCCESS [ 2.285 s]\r\n[INFO] Apache HBase - Rest ................................ SUCCESS [ 2.777 s]\r\n[INFO] Apache HBase - Examples ............................ SUCCESS [ 2.009 s]\r\n[INFO] Apache HBase - Shaded .............................. SUCCESS [ 0.177 s]\r\n[INFO] Apache HBase - Shaded - Client (with Hadoop bundled) SUCCESS [ 0.853 s]\r\n[INFO] Apache HBase - Shaded - Client ..................... SUCCESS [ 0.835 s]\r\n[INFO] Apache HBase - Shaded - MapReduce .................. SUCCESS [ 2.661 s]\r\n[INFO] Apache HBase - External Block Cache ................ SUCCESS [ 1.530 s]\r\n[INFO] Apache HBase - Spark ............................... SUCCESS [ 34.849 s]\r\n[INFO] Apache HBase - Spark Integration Tests ............. SUCCESS [ 1.839 s]\r\n[INFO] Apache HBase - Assembly ............................ SUCCESS [ 3.908 s]\r\n[INFO] Apache HBase Shaded Packaging Invariants ........... SUCCESS [ 0.841 s]\r\n[INFO] Apache HBase Shaded Packaging Invariants (with Hadoop bundled) SUCCESS [ 0.563 s]\r\n[INFO] Apache HBase - Archetypes .......................... SUCCESS [ 0.037 s]\r\n[INFO] Apache HBase - Exemplar for hbase-client archetype . SUCCESS [ 1.093 s]\r\n[INFO] Apache HBase - Exemplar for hbase-shaded-client archetype SUCCESS [ 0.825 s]\r\n[INFO] Apache HBase - Archetype builder ................... SUCCESS [ 0.368 s]\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD SUCCESS\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 03:01 min\r\n[INFO] Finished at: 2018-08-31T16:05:58+05:30\r\n[INFO] Final Memory: 262M/1033M\r\n[INFO] ------------------------------------------------------------------------\r\n{code}\r\n","created":"2018-08-31T10:48:11.235+0000"},{"body":"Ping [~busbey], [~mdrob@cloudera.com] as you are more familiar with HBASE-16351.","created":"2018-08-31T10:49:47.496+0000"},{"body":"Patch LGTM and enforcer passes on my Mac with it, but I don't have a Windows machine to test on available right now. Pretty sure the test failures are unrelated too. I think [~busbey] has a Windows machine, so I'll wait for him to report back.\r\n \r\n","created":"2018-08-31T14:21:19.166+0000"},{"body":"I do in fact have a windows build machine. Let me go dig it up. Might be late before I finish with a build, I don't think it's been set up to build the hbase project before.","created":"2018-08-31T15:15:58.708+0000"},{"body":"the compile and javac error looks concerning though unrelated. ","created":"2018-08-31T15:16:35.775+0000"},{"body":"I can't build master on my windows box before or after this patch. Maven fails in both cases with an error in hbase-common about finding bash.\r\n\r\nIf things still build fine on unix-y platforms I'm not opposed to the patch. But I'd need to know more about what building on windows is supposed to look like in order to weigh in on wether or not things get better there.","created":"2018-08-31T18:50:29.043+0000"},{"body":"{quote}I can't build master on my windows box before or after this patch. Maven fails in both cases with an error in hbase-common about finding bash.\r\n{quote}\r\n[~busbey] Do you have [cygwin|https://hbase.apache.org/cygwin.html] installed? I compile hbase using cygwin.","created":"2018-09-14T09:57:20.612+0000"},{"body":"No, I do not have cygwin since the ref guide didn't mention anything about it.\r\n\r\nI'm +1 on including this to help out short term. But we should have a DISCUSS thread on dev@ about the future of building on windows. There's too many gaps right now e.g. nothing in the ref guide about set up and no nightly tests that building on windows works. If we're not going to make those efforts we should expressly state that it isn't expected to work.","created":"2018-09-14T15:53:08.871+0000"},{"body":"{quote}But we should have a DISCUSS thread on dev@ about the future of building on windows.\r\n{quote}\r\nThat would be great. I think we should at least be able to compile hbase correctly on windows. I use windows for development at my workplace; I need to compile the code before pushing any changes to my work-repo. Sometimes, I also need to run unit tests via eclipse. I believe many other developers might be using windows as their development environment. Although I doubt very few would be using windows as test environment. Ensuring that hbase can be compiled on windows would save a lot of trouble windows developers may face initially while setting up hbase for the first time. In fact today I found we cannot compile on windows with \"{{-Prelease\"}} profile. I was thinking about fixing it and submitting a patch for the same; given that hbase community is eager to support compilation on windows.","created":"2018-09-14T18:29:22.413+0000"},{"body":"[~nihaljain.cs] ,  [^HBASE-21135.master.001.patch] is good in branch-2.  I tested it  in master, it solved HBase-Shaded Module, but there are still problems in Apache HBase - Shaded - MapReduce on  the Windows OS. Have you met?\r\n{code:java}\r\nINFO] Apache HBase - Shaded .............................. SUCCESS [ 1.115 s]\r\n[INFO] Apache HBase - Shaded - Client (with Hadoop bundled) SUCCESS [ 30.302 s]\r\n[INFO] Apache HBase - Shaded - Client ..................... SUCCESS [ 28.375 s]\r\n[INFO] Apache HBase - Shaded - MapReduce .................. FAILURE [ 6.677 s]\r\n[INFO] Apache HBase - External Block Cache ................ SKIPPED\r\n[INFO] Apache HBase - Assembly ............................ SKIPPED\r\n[INFO] Apache HBase Shaded Packaging Invariants ........... SKIPPED\r\n[INFO] Apache HBase Shaded Packaging Invariants (with Hadoop bundled) SKIPPED\r\n[INFO] Apache HBase - Archetypes .......................... SKIPPED\r\n[INFO] Apache HBase - Exemplar for hbase-client archetype . SKIPPED\r\n[INFO] Apache HBase - Exemplar for hbase-shaded-client archetype SKIPPED\r\n[INFO] Apache HBase - Archetype builder ................... SKIPPED\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD FAILURE\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 08:59 min\r\n[INFO] Finished at: 2018-12-29T11:25:09+08:00\r\n[INFO] Final Memory: 217M/967M\r\n[INFO] ------------------------------------------------------------------------\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce (check-aggregate-license) on project hbase-shaded-mapreduce: Some Enforcer rules have failed. Look above for specific messages explaining why the rule failed. -> [Help 1]\r\n[ERROR] \r\n[ERROR] To see the full stack trace of the errors, re-run Maven with the -e switch.\r\n[ERROR] Re-run Maven using the -X switch to enable full debug logging.\r\n{code}","created":"2018-12-29T11:46:14.692+0000"},{"body":"[~xu qinya] Thanks for trying the patch. I tried compiling the latest master build with patch on windows. I didnot face any error in hbase-mapreduce-shaded module. Do you mind sharing the detailed error by adding -X to your mvn compile command. For eg:\r\n{quote}{{mvn clean install -DskipTests -X}}\r\n{quote}","created":"2018-12-31T08:55:33.509+0000"},{"body":"[~nihaljain.cs]  Sorry, It's a problem with my test environment. I changed it. Compiling the latest master build with patch on windows is succeed. ","created":"2019-01-01T02:55:19.837+0000"},{"body":"{quote}Compiling the latest master build with patch on windows is succeed.\r\n{quote}\r\nAwesome. Thanks [~xu qinya] for testing the patch.","created":"2019-01-01T20:05:07.166+0000"},{"body":"this problem is facing in the branch-1 also.","created":"2019-01-03T13:16:54.444+0000"},{"body":"[~sreenivasulureddy] does this patch fix the issue?","created":"2019-01-03T13:49:22.905+0000"},{"body":"[~nihaljain.cs]\r\nYes this patch is fixing the issue.","created":"2019-01-03T14:24:16.711+0000"},{"body":"Thanks [~sreenivasulureddy].\r\n\r\nSo the windows build would fail in all branches having HBASE-16351.","created":"2019-01-03T18:03:04.652+0000"},{"body":"ping [~busbey] ","created":"2019-01-05T08:13:14.886+0000"},{"body":"thanks for the ping! in my queue for this coming week. my only windows machine was a few states away during the holidays. :)\r\n\r\nCan someone summarize the prerequisites and steps I should expect here?\r\n\r\nIs it:\r\n\r\n* install cygwin\r\n* run maven goals as usual from within cygwin\r\n\r\nif I have normal windows installs of java and maven do I need to do something special about using them under cygwin?","created":"2019-01-05T18:51:56.607+0000"},{"body":"[~busbey]\r\n\r\nyep,install *maven,java,cygwin* in the Windows OS,then running mvn cmd to test will be ok. \r\nThese two patches are very,very useful for the windows users(poor guys) to learn and contribute to the HBase.","created":"2019-01-07T10:13:59.198+0000"},{"body":"I cloned the latest master branch, but i got the same problem with you.\r\n\r\nI see the pom.xml inside, it didn't change to your patch,why?\r\n\r\nMy branch is at 2019/1/20","created":"2019-01-20T12:56:49.483+0000"},{"body":"{quote}I see the pom.xml inside, it didn't change to your patch,why?\r\n{quote}\r\nHi [~pingsutw], this patch has not been merged yet.\r\n","created":"2019-01-21T06:54:38.153+0000"},{"body":"ping [~busbey]\r\n\r\nIt couldn't be better to backport these two patches to branch-1.x!","created":"2019-01-21T11:51:36.693+0000"},{"body":"Ping [~apurtell] FYI","created":"2019-01-25T13:05:23.887+0000"},{"body":"It's fine to commit this to branch-1 etc as long as it doesn't break the build for everyone else. We don't support Cygwin build environments. The discussion happened a while back. I have a vague memory of when we ripped out some Cygwin specific build stuff a long time ago. That said, its fine to accept patches that help there, as long as they don't break the build on any UNIX derived OS.","created":"2019-01-25T19:56:46.373+0000"},{"body":"{quote}It's fine to commit this to branch-1 etc as long as it doesn't break the build for everyone else. \r\n{quote}\r\nHave tested license check on linux. Works as expected. Earlier, [~mdrob] had confirmed that it passed on mac too (See [comment|https://issues.apache.org/jira/browse/HBASE-21135?focusedCommentId=16598766&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-16598766])\r\n\r\n ","created":"2019-02-06T06:58:28.747+0000"},{"body":"[~busbey] If the patch looks fine, could you commit this to affected branches?","created":"2019-02-09T11:31:43.396+0000"},{"body":"progress! I have installed cygwin on my windows build box and successfully built {{mvn -DskipTests package}} using this patch.\r\n\r\nI just need to confirm again not breaking things on linux, which I'm doing now.\r\n\r\n(I didn't have HBASE-21656 in place, so I'll need to follow up there on why things didn't fail)","created":"2019-03-08T17:43:41.315+0000"},{"body":"confirmed things work locally for me after patch. that included adding a dependency with a banned license to make sure enforcer would find the right file. Pushed to master so far; I'll push to all active branches as I manage to similarly confirm build integrity.","created":"2019-03-08T20:08:31.529+0000"},{"body":"This change has introduced a new Maven warning on branch-1\r\n\r\n[WARNING] Some problems were encountered while building the effective model for org.apache.hbase:hbase-resource-bundle:jar:1.5.0\r\n[WARNING] 'build.plugins.plugin.(groupId:artifactId)' must be unique but found duplicate declaration of plugin org.codehaus.mojo:build-helper-maven-plugin @ org.apache.hbase:hbase:1.5.0, hbase/pom.xml, line 837, column 15\r\n\r\n ","created":"2019-03-28T16:58:01.574+0000"},{"body":"There is a trivial fix. I am committing it everywhere. WIll post the addendum patch here","created":"2019-03-28T17:03:00.660+0000"}],"conversations":[{"body":"License check via enforce plugin throws following error during build on windows:\r\n{code:java}\r\nSourced file: inline evaluation of: ``File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-ar . . . '' Token Parsing Error: Lexical error at line 1, column 29.  Encountered: \"D\" (68), after : \"\\\"D:\\\\\": {code}\r\nComplete stacktrace with command\r\n{code:java}\r\nmvn clean install -DskipTests -X\r\n{code}\r\nis as follows:\r\n{noformat}\r\n[INFO] --- maven-enforcer-plugin:3.0.0-M1:enforce (check-aggregate-license) @ hbase-shaded ---\r\n[DEBUG] Configuring mojo org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce from plugin realm ClassRealm[plugin>org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1, parent: sun.misc.Launcher$AppClassLoader@55f96302]\r\n[DEBUG] Configuring mojo 'org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce' with basic configurator -->\r\n[DEBUG] (s) fail = true\r\n[DEBUG] (s) failFast = false\r\n[DEBUG] (f) ignoreCache = false\r\n[DEBUG] (f) mojoExecution = org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce {execution: check-aggregate-license}\r\n[DEBUG] (s) project = MavenProject: org.apache.hbase:hbase-shaded:2.1.1-SNAPSHOT @ D:\\DS\\HBase_2\\hbase\\hbase-shaded\\pom.xml\r\n[DEBUG] (s) condition = File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\");\r\n\r\n // Beanshell does not support try-with-resources,\r\n // so we must close this scanner manually\r\n Scanner scanner = new Scanner(license);\r\n\r\n while (scanner.hasNextLine()) {\r\n if (scanner.nextLine().startsWith(\"ERROR:\")) {\r\n scanner.close();\r\n return false;\r\n }\r\n }\r\n scanner.close();\r\n return true;\r\n[DEBUG] (s) message = License errors detected, for more detail find ERROR in\r\n D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\r\n[DEBUG] (s) rules = [org.apache.maven.plugins.enforcer.EvaluateBeanshell@7e307087]\r\n[DEBUG] (s) session = org.apache.maven.execution.MavenSession@5e1218b4\r\n[DEBUG] (s) skip = false\r\n[DEBUG] -- end configuration --\r\n[DEBUG] Executing rule: org.apache.maven.plugins.enforcer.EvaluateBeanshell\r\n[DEBUG] Echo condition : File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\");\r\n\r\n // Beanshell does not support try-with-resources,\r\n // so we must close this scanner manually\r\n Scanner scanner = new Scanner(license);\r\n\r\n while (scanner.hasNextLine()) {\r\n if (scanner.nextLine().startsWith(\"ERROR:\")) {\r\n scanner.close();\r\n return false;\r\n }\r\n }\r\n scanner.close();\r\n return true;\r\n[DEBUG] Echo script : File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\");\r\n\r\n // Beanshell does not support try-with-resources,\r\n // so we must close this scanner manually\r\n Scanner scanner = new Scanner(license);\r\n\r\n while (scanner.hasNextLine()) {\r\n if (scanner.nextLine().startsWith(\"ERROR:\")) {\r\n scanner.close();\r\n return false;\r\n }\r\n }\r\n scanner.close();\r\n return true;\r\n[DEBUG] Adding failure due to exception\r\norg.apache.maven.enforcer.rule.api.EnforcerRuleException: Couldn't evaluate condition: File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\");\r\n\r\n // Beanshell does not support try-with-resources,\r\n // so we must close this scanner manually\r\n Scanner scanner = new Scanner(license);\r\n\r\n while (scanner.hasNextLine()) {\r\n if (scanner.nextLine().startsWith(\"ERROR:\")) {\r\n scanner.close();\r\n return false;\r\n }\r\n }\r\n scanner.close();\r\n return true;\r\n at org.apache.maven.plugins.enforcer.EvaluateBeanshell.evaluateCondition(EvaluateBeanshell.java:107)\r\n at org.apache.maven.plugins.enforcer.EvaluateBeanshell.execute(EvaluateBeanshell.java:72)\r\n at org.apache.maven.plugins.enforcer.EnforceMojo.execute(EnforceMojo.java:202)\r\n at org.apache.maven.plugin.DefaultBuildPluginManager.executeMojo(DefaultBuildPluginManager.java:134)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:208)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:154)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:146)\r\n at org.apache.maven.lifecycle.internal.LifecycleModuleBuilder.buildProject(LifecycleModuleBuilder.java:117)\r\n at org.apache.maven.lifecycle.internal.LifecycleModuleBuilder.buildProject(LifecycleModuleBuilder.java:81)\r\n at org.apache.maven.lifecycle.internal.builder.singlethreaded.SingleThreadedBuilder.build(SingleThreadedBuilder.java:51)\r\n at org.apache.maven.lifecycle.internal.LifecycleStarter.execute(LifecycleStarter.java:128)\r\n at org.apache.maven.DefaultMaven.doExecute(DefaultMaven.java:309)\r\n at org.apache.maven.DefaultMaven.doExecute(DefaultMaven.java:194)\r\n at org.apache.maven.DefaultMaven.execute(DefaultMaven.java:107)\r\n at org.apache.maven.cli.MavenCli.execute(MavenCli.java:993)\r\n at org.apache.maven.cli.MavenCli.doMain(MavenCli.java:345)\r\n at org.apache.maven.cli.MavenCli.main(MavenCli.java:191)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.launchEnhanced(Launcher.java:289)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.launch(Launcher.java:229)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.mainWithExitCode(Launcher.java:415)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.main(Launcher.java:356)\r\nCaused by: Sourced file: inline evaluation of: ``File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-ar . . . '' Token Parsing Error: Lexical error at line 1, column 29. Encountered: \"D\" (68), after : \"\\\"D:\\\\\": \r\n\r\n at bsh.Interpreter.eval(Unknown Source)\r\n at bsh.Interpreter.eval(Unknown Source)\r\n at bsh.Interpreter.eval(Unknown Source)\r\n at org.apache.maven.plugins.enforcer.EvaluateBeanshell.evaluateCondition(EvaluateBeanshell.java:102)\r\n ... 24 more\r\n[WARNING] Rule 0: org.apache.maven.plugins.enforcer.EvaluateBeanshell failed with message:\r\nCouldn't evaluate condition: File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\");\r\n\r\n // Beanshell does not support try-with-resources,\r\n // so we must close this scanner manually\r\n Scanner scanner = new Scanner(license);\r\n\r\n while (scanner.hasNextLine()) {\r\n if (scanner.nextLine().startsWith(\"ERROR:\")) {\r\n scanner.close();\r\n return false;\r\n }\r\n }\r\n scanner.close();\r\n return true;\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Reactor Summary:\r\n[INFO]\r\n[INFO] Apache HBase - Shaded .............................. FAILURE [ 2.388 s]\r\n[INFO] Apache HBase - Shaded - Client (with Hadoop bundled) SKIPPED\r\n[INFO] Apache HBase - Shaded - Client ..................... SKIPPED\r\n[INFO] Apache HBase - Shaded - MapReduce .................. SKIPPED\r\n[INFO] Apache HBase - External Block Cache ................ SKIPPED\r\n[INFO] Apache HBase - Assembly ............................ SKIPPED\r\n[INFO] Apache HBase Shaded Packaging Invariants ........... SKIPPED\r\n[INFO] Apache HBase Shaded Packaging Invariants (with Hadoop bundled) SKIPPED\r\n[INFO] Apache HBase - Archetypes .......................... SKIPPED\r\n[INFO] Apache HBase - Exemplar for hbase-client archetype . SKIPPED\r\n[INFO] Apache HBase - Exemplar for hbase-shaded-client archetype SKIPPED\r\n[INFO] Apache HBase - Archetype builder ................... SKIPPED\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD FAILURE\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 4.822 s\r\n[INFO] Finished at: 2018-08-31T15:39:00+05:30\r\n[INFO] Final Memory: 41M/421M\r\n[INFO] ------------------------------------------------------------------------\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce (check-aggregate-license) on project hbase-shaded: Some Enforcer rules have failed. Look above for specific messages explaining why the rule failed. -> [Help 1]\r\norg.apache.maven.lifecycle.LifecycleExecutionException: Failed to execute goal org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce (check-aggregate-license) on project hbase-shaded: Some Enforcer rules have failed. Look above for specific messages explaining why the rule failed.\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:213)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:154)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:146)\r\n at org.apache.maven.lifecycle.internal.LifecycleModuleBuilder.buildProject(LifecycleModuleBuilder.java:117)\r\n at org.apache.maven.lifecycle.internal.LifecycleModuleBuilder.buildProject(LifecycleModuleBuilder.java:81)\r\n at org.apache.maven.lifecycle.internal.builder.singlethreaded.SingleThreadedBuilder.build(SingleThreadedBuilder.java:51)\r\n at org.apache.maven.lifecycle.internal.LifecycleStarter.execute(LifecycleStarter.java:128)\r\n at org.apache.maven.DefaultMaven.doExecute(DefaultMaven.java:309)\r\n at org.apache.maven.DefaultMaven.doExecute(DefaultMaven.java:194)\r\n at org.apache.maven.DefaultMaven.execute(DefaultMaven.java:107)\r\n at org.apache.maven.cli.MavenCli.execute(MavenCli.java:993)\r\n at org.apache.maven.cli.MavenCli.doMain(MavenCli.java:345)\r\n at org.apache.maven.cli.MavenCli.main(MavenCli.java:191)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.launchEnhanced(Launcher.java:289)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.launch(Launcher.java:229)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.mainWithExitCode(Launcher.java:415)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.main(Launcher.java:356)\r\nCaused by: org.apache.maven.plugin.MojoExecutionException: Some Enforcer rules have failed. Look above for specific messages explaining why the rule failed.\r\n at org.apache.maven.plugins.enforcer.EnforceMojo.execute(EnforceMojo.java:243)\r\n at org.apache.maven.plugin.DefaultBuildPluginManager.executeMojo(DefaultBuildPluginManager.java:134)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:208)\r\n ... 20 more\r\n[ERROR]\r\n[ERROR]\r\n[ERROR] For more information about the errors and possible solutions, please read the following articles:\r\n[ERROR] [Help 1] http://cwiki.apache.org/confluence/display/MAVEN/MojoExecutionException\r\n{noformat}\r\n ","from":"reporter","subject":"Build fails on windows as it fails to parse windows path during license check"},{"body":"Initially I tried to do replacements in beanshell condition itself. But it seems it is too late to do so as the enforcer fails to evaluate the beanshell syntax evben then. Attached an initial patch [^HBASE-21135.master.001.patch] which replaces all backslashes (if any, say in windows path) with forward slash via a plugin. ","from":"developer"},{"body":"With patch on Windows:\r\n{noformat}\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Reactor Summary:\r\n[INFO]\r\n[INFO] Apache HBase ....................................... SUCCESS [ 3.441 s]\r\n[INFO] Apache HBase - Checkstyle .......................... SUCCESS [ 1.043 s]\r\n[INFO] Apache HBase - Build Support ....................... SUCCESS [ 0.057 s]\r\n[INFO] Apache HBase - Error Prone Rules ................... SUCCESS [ 1.385 s]\r\n[INFO] Apache HBase - Annotations ......................... SUCCESS [ 0.614 s]\r\n[INFO] Apache HBase - Build Configuration ................. SUCCESS [ 0.128 s]\r\n[INFO] Apache HBase - Shaded Protocol ..................... SUCCESS [ 34.409 s]\r\n[INFO] Apache HBase - Common .............................. SUCCESS [ 24.600 s]\r\n[INFO] Apache HBase - Metrics API ......................... SUCCESS [ 2.099 s]\r\n[INFO] Apache HBase - Hadoop Compatibility ................ SUCCESS [ 2.799 s]\r\n[INFO] Apache HBase - Metrics Implementation .............. SUCCESS [ 2.513 s]\r\n[INFO] Apache HBase - Hadoop Two Compatibility ............ SUCCESS [ 3.540 s]\r\n[INFO] Apache HBase - Protocol ............................ SUCCESS [ 11.765 s]\r\n[INFO] Apache HBase - Client .............................. SUCCESS [ 12.878 s]\r\n[INFO] Apache HBase - Zookeeper ........................... SUCCESS [ 3.142 s]\r\n[INFO] Apache HBase - Replication ......................... SUCCESS [ 3.655 s]\r\n[INFO] Apache HBase - Resource Bundle ..................... SUCCESS [ 0.328 s]\r\n[INFO] Apache HBase - HTTP ................................ SUCCESS [ 6.054 s]\r\n[INFO] Apache HBase - Procedure ........................... SUCCESS [ 3.555 s]\r\n[INFO] Apache HBase - Server .............................. SUCCESS [ 41.261 s]\r\n[INFO] Apache HBase - MapReduce ........................... SUCCESS [ 10.958 s]\r\n[INFO] Apache HBase - Testing Util ........................ SUCCESS [ 4.142 s]\r\n[INFO] Apache HBase - Thrift .............................. SUCCESS [ 12.591 s]\r\n[INFO] Apache HBase - RSGroup ............................. SUCCESS [ 7.119 s]\r\n[INFO] Apache HBase - Shell ............................... SUCCESS [ 4.232 s]\r\n[INFO] Apache HBase - Coprocessor Endpoint ................ SUCCESS [ 7.678 s]\r\n[INFO] Apache HBase - Integration Tests ................... SUCCESS [ 6.985 s]\r\n[INFO] Apache HBase - Rest ................................ SUCCESS [ 9.629 s]\r\n[INFO] Apache HBase - Examples ............................ SUCCESS [ 6.832 s]\r\n[INFO] Apache HBase - Shaded .............................. SUCCESS [ 0.272 s]\r\n[INFO] Apache HBase - Shaded - Client (with Hadoop bundled) SUCCESS [ 1.451 s]\r\n[INFO] Apache HBase - Shaded - Client ..................... SUCCESS [ 1.466 s]\r\n[INFO] Apache HBase - Shaded - MapReduce .................. SUCCESS [ 4.534 s]\r\n[INFO] Apache HBase - External Block Cache ................ SUCCESS [ 2.963 s]\r\n[INFO] Apache HBase - Assembly ............................ SUCCESS [ 13.044 s]\r\n[INFO] Apache HBase Shaded Packaging Invariants ........... SUCCESS [ 14.888 s]\r\n[INFO] Apache HBase Shaded Packaging Invariants (with Hadoop bundled) SUCCESS [ 12.077 s]\r\n[INFO] Apache HBase - Archetypes .......................... SUCCESS [ 0.121 s]\r\n[INFO] Apache HBase - Exemplar for hbase-client archetype . SUCCESS [ 3.492 s]\r\n[INFO] Apache HBase - Exemplar for hbase-shaded-client archetype SUCCESS [ 1.708 s]\r\n[INFO] Apache HBase - Archetype builder ................... SUCCESS [ 0.665 s]\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD SUCCESS\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 04:48 min\r\n[INFO] Finished at: 2018-08-31T16:11:25+05:30\r\n[INFO] Final Memory: 229M/1474M\r\n[INFO] ------------------------------------------------------------------------\r\n{noformat}\r\n\r\n","from":"developer"},{"body":"With patch on Ubuntu:\r\n\r\n{code:java}\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Reactor Summary:\r\n[INFO] \r\n[INFO] Apache HBase ....................................... SUCCESS [ 2.909 s]\r\n[INFO] Apache HBase - Checkstyle .......................... SUCCESS [ 0.744 s]\r\n[INFO] Apache HBase - Build Support ....................... SUCCESS [ 0.059 s]\r\n[INFO] Apache HBase - Error Prone Rules ................... SUCCESS [ 0.772 s]\r\n[INFO] Apache HBase - Annotations ......................... SUCCESS [ 0.463 s]\r\n[INFO] Apache HBase - Build Configuration ................. SUCCESS [ 0.104 s]\r\n[INFO] Apache HBase - Shaded Protocol ..................... SUCCESS [ 29.708 s]\r\n[INFO] Apache HBase - Common .............................. SUCCESS [ 6.303 s]\r\n[INFO] Apache HBase - Metrics API ......................... SUCCESS [ 1.378 s]\r\n[INFO] Apache HBase - Hadoop Compatibility ................ SUCCESS [ 1.703 s]\r\n[INFO] Apache HBase - Metrics Implementation .............. SUCCESS [ 1.256 s]\r\n[INFO] Apache HBase - Hadoop Two Compatibility ............ SUCCESS [ 1.935 s]\r\n[INFO] Apache HBase - Protocol ............................ SUCCESS [ 7.504 s]\r\n[INFO] Apache HBase - Client .............................. SUCCESS [ 7.923 s]\r\n[INFO] Apache HBase - Zookeeper ........................... SUCCESS [ 1.697 s]\r\n[INFO] Apache HBase - Replication ......................... SUCCESS [ 1.949 s]\r\n[INFO] Apache HBase - Resource Bundle ..................... SUCCESS [ 0.192 s]\r\n[INFO] Apache HBase - HTTP ................................ SUCCESS [ 4.039 s]\r\n[INFO] Apache HBase - Procedure ........................... SUCCESS [ 2.151 s]\r\n[INFO] Apache HBase - Server .............................. SUCCESS [ 26.699 s]\r\n[INFO] Apache HBase - MapReduce ........................... SUCCESS [ 5.589 s]\r\n[INFO] Apache HBase - Testing Util ........................ SUCCESS [ 2.789 s]\r\n[INFO] Apache HBase - Thrift .............................. SUCCESS [ 5.217 s]\r\n[INFO] Apache HBase - RSGroup ............................. SUCCESS [ 2.605 s]\r\n[INFO] Apache HBase - Shell ............................... SUCCESS [ 1.540 s]\r\n[INFO] Apache HBase - Coprocessor Endpoint ................ SUCCESS [ 2.338 s]\r\n[INFO] Apache HBase - Backup .............................. SUCCESS [ 1.895 s]\r\n[INFO] Apache HBase - Integration Tests ................... SUCCESS [ 2.285 s]\r\n[INFO] Apache HBase - Rest ................................ SUCCESS [ 2.777 s]\r\n[INFO] Apache HBase - Examples ............................ SUCCESS [ 2.009 s]\r\n[INFO] Apache HBase - Shaded .............................. SUCCESS [ 0.177 s]\r\n[INFO] Apache HBase - Shaded - Client (with Hadoop bundled) SUCCESS [ 0.853 s]\r\n[INFO] Apache HBase - Shaded - Client ..................... SUCCESS [ 0.835 s]\r\n[INFO] Apache HBase - Shaded - MapReduce .................. SUCCESS [ 2.661 s]\r\n[INFO] Apache HBase - External Block Cache ................ SUCCESS [ 1.530 s]\r\n[INFO] Apache HBase - Spark ............................... SUCCESS [ 34.849 s]\r\n[INFO] Apache HBase - Spark Integration Tests ............. SUCCESS [ 1.839 s]\r\n[INFO] Apache HBase - Assembly ............................ SUCCESS [ 3.908 s]\r\n[INFO] Apache HBase Shaded Packaging Invariants ........... SUCCESS [ 0.841 s]\r\n[INFO] Apache HBase Shaded Packaging Invariants (with Hadoop bundled) SUCCESS [ 0.563 s]\r\n[INFO] Apache HBase - Archetypes .......................... SUCCESS [ 0.037 s]\r\n[INFO] Apache HBase - Exemplar for hbase-client archetype . SUCCESS [ 1.093 s]\r\n[INFO] Apache HBase - Exemplar for hbase-shaded-client archetype SUCCESS [ 0.825 s]\r\n[INFO] Apache HBase - Archetype builder ................... SUCCESS [ 0.368 s]\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD SUCCESS\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 03:01 min\r\n[INFO] Finished at: 2018-08-31T16:05:58+05:30\r\n[INFO] Final Memory: 262M/1033M\r\n[INFO] ------------------------------------------------------------------------\r\n{code}\r\n","from":"developer"},{"body":"Ping [~busbey], [~mdrob@cloudera.com] as you are more familiar with HBASE-16351.","from":"developer"},{"body":"Patch LGTM and enforcer passes on my Mac with it, but I don't have a Windows machine to test on available right now. Pretty sure the test failures are unrelated too. I think [~busbey] has a Windows machine, so I'll wait for him to report back.\r\n \r\n","from":"developer"},{"body":"I do in fact have a windows build machine. Let me go dig it up. Might be late before I finish with a build, I don't think it's been set up to build the hbase project before.","from":"developer"},{"body":"the compile and javac error looks concerning though unrelated. ","from":"developer"},{"body":"I can't build master on my windows box before or after this patch. Maven fails in both cases with an error in hbase-common about finding bash.\r\n\r\nIf things still build fine on unix-y platforms I'm not opposed to the patch. But I'd need to know more about what building on windows is supposed to look like in order to weigh in on wether or not things get better there.","from":"developer"},{"body":"{quote}I can't build master on my windows box before or after this patch. Maven fails in both cases with an error in hbase-common about finding bash.\r\n{quote}\r\n[~busbey] Do you have [cygwin|https://hbase.apache.org/cygwin.html] installed? I compile hbase using cygwin.","from":"developer"},{"body":"No, I do not have cygwin since the ref guide didn't mention anything about it.\r\n\r\nI'm +1 on including this to help out short term. But we should have a DISCUSS thread on dev@ about the future of building on windows. There's too many gaps right now e.g. nothing in the ref guide about set up and no nightly tests that building on windows works. If we're not going to make those efforts we should expressly state that it isn't expected to work.","from":"developer"},{"body":"{quote}But we should have a DISCUSS thread on dev@ about the future of building on windows.\r\n{quote}\r\nThat would be great. I think we should at least be able to compile hbase correctly on windows. I use windows for development at my workplace; I need to compile the code before pushing any changes to my work-repo. Sometimes, I also need to run unit tests via eclipse. I believe many other developers might be using windows as their development environment. Although I doubt very few would be using windows as test environment. Ensuring that hbase can be compiled on windows would save a lot of trouble windows developers may face initially while setting up hbase for the first time. In fact today I found we cannot compile on windows with \"{{-Prelease\"}} profile. I was thinking about fixing it and submitting a patch for the same; given that hbase community is eager to support compilation on windows.","from":"developer"},{"body":"[~nihaljain.cs] ,  [^HBASE-21135.master.001.patch] is good in branch-2.  I tested it  in master, it solved HBase-Shaded Module, but there are still problems in Apache HBase - Shaded - MapReduce on  the Windows OS. Have you met?\r\n{code:java}\r\nINFO] Apache HBase - Shaded .............................. SUCCESS [ 1.115 s]\r\n[INFO] Apache HBase - Shaded - Client (with Hadoop bundled) SUCCESS [ 30.302 s]\r\n[INFO] Apache HBase - Shaded - Client ..................... SUCCESS [ 28.375 s]\r\n[INFO] Apache HBase - Shaded - MapReduce .................. FAILURE [ 6.677 s]\r\n[INFO] Apache HBase - External Block Cache ................ SKIPPED\r\n[INFO] Apache HBase - Assembly ............................ SKIPPED\r\n[INFO] Apache HBase Shaded Packaging Invariants ........... SKIPPED\r\n[INFO] Apache HBase Shaded Packaging Invariants (with Hadoop bundled) SKIPPED\r\n[INFO] Apache HBase - Archetypes .......................... SKIPPED\r\n[INFO] Apache HBase - Exemplar for hbase-client archetype . SKIPPED\r\n[INFO] Apache HBase - Exemplar for hbase-shaded-client archetype SKIPPED\r\n[INFO] Apache HBase - Archetype builder ................... SKIPPED\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD FAILURE\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 08:59 min\r\n[INFO] Finished at: 2018-12-29T11:25:09+08:00\r\n[INFO] Final Memory: 217M/967M\r\n[INFO] ------------------------------------------------------------------------\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce (check-aggregate-license) on project hbase-shaded-mapreduce: Some Enforcer rules have failed. Look above for specific messages explaining why the rule failed. -> [Help 1]\r\n[ERROR] \r\n[ERROR] To see the full stack trace of the errors, re-run Maven with the -e switch.\r\n[ERROR] Re-run Maven using the -X switch to enable full debug logging.\r\n{code}","from":"developer"},{"body":"[~xu qinya] Thanks for trying the patch. I tried compiling the latest master build with patch on windows. I didnot face any error in hbase-mapreduce-shaded module. Do you mind sharing the detailed error by adding -X to your mvn compile command. For eg:\r\n{quote}{{mvn clean install -DskipTests -X}}\r\n{quote}","from":"developer"},{"body":"[~nihaljain.cs]  Sorry, It's a problem with my test environment. I changed it. Compiling the latest master build with patch on windows is succeed. ","from":"developer"},{"body":"{quote}Compiling the latest master build with patch on windows is succeed.\r\n{quote}\r\nAwesome. Thanks [~xu qinya] for testing the patch.","from":"developer"},{"body":"this problem is facing in the branch-1 also.","from":"developer"},{"body":"[~sreenivasulureddy] does this patch fix the issue?","from":"developer"},{"body":"[~nihaljain.cs]\r\nYes this patch is fixing the issue.","from":"developer"},{"body":"Thanks [~sreenivasulureddy].\r\n\r\nSo the windows build would fail in all branches having HBASE-16351.","from":"developer"},{"body":"ping [~busbey] ","from":"developer"},{"body":"thanks for the ping! in my queue for this coming week. my only windows machine was a few states away during the holidays. :)\r\n\r\nCan someone summarize the prerequisites and steps I should expect here?\r\n\r\nIs it:\r\n\r\n* install cygwin\r\n* run maven goals as usual from within cygwin\r\n\r\nif I have normal windows installs of java and maven do I need to do something special about using them under cygwin?","from":"developer"},{"body":"[~busbey]\r\n\r\nyep,install *maven,java,cygwin* in the Windows OS,then running mvn cmd to test will be ok. \r\nThese two patches are very,very useful for the windows users(poor guys) to learn and contribute to the HBase.","from":"developer"},{"body":"I cloned the latest master branch, but i got the same problem with you.\r\n\r\nI see the pom.xml inside, it didn't change to your patch,why?\r\n\r\nMy branch is at 2019/1/20","from":"developer"},{"body":"{quote}I see the pom.xml inside, it didn't change to your patch,why?\r\n{quote}\r\nHi [~pingsutw], this patch has not been merged yet.\r\n","from":"developer"},{"body":"ping [~busbey]\r\n\r\nIt couldn't be better to backport these two patches to branch-1.x!","from":"developer"},{"body":"Ping [~apurtell] FYI","from":"developer"},{"body":"It's fine to commit this to branch-1 etc as long as it doesn't break the build for everyone else. We don't support Cygwin build environments. The discussion happened a while back. I have a vague memory of when we ripped out some Cygwin specific build stuff a long time ago. That said, its fine to accept patches that help there, as long as they don't break the build on any UNIX derived OS.","from":"developer"},{"body":"{quote}It's fine to commit this to branch-1 etc as long as it doesn't break the build for everyone else. \r\n{quote}\r\nHave tested license check on linux. Works as expected. Earlier, [~mdrob] had confirmed that it passed on mac too (See [comment|https://issues.apache.org/jira/browse/HBASE-21135?focusedCommentId=16598766&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-16598766])\r\n\r\n ","from":"developer"},{"body":"[~busbey] If the patch looks fine, could you commit this to affected branches?","from":"developer"},{"body":"progress! I have installed cygwin on my windows build box and successfully built {{mvn -DskipTests package}} using this patch.\r\n\r\nI just need to confirm again not breaking things on linux, which I'm doing now.\r\n\r\n(I didn't have HBASE-21656 in place, so I'll need to follow up there on why things didn't fail)","from":"developer"},{"body":"confirmed things work locally for me after patch. that included adding a dependency with a banned license to make sure enforcer would find the right file. Pushed to master so far; I'll push to all active branches as I manage to similarly confirm build integrity.","from":"developer"},{"body":"This change has introduced a new Maven warning on branch-1\r\n\r\n[WARNING] Some problems were encountered while building the effective model for org.apache.hbase:hbase-resource-bundle:jar:1.5.0\r\n[WARNING] 'build.plugins.plugin.(groupId:artifactId)' must be unique but found duplicate declaration of plugin org.codehaus.mojo:build-helper-maven-plugin @ org.apache.hbase:hbase:1.5.0, hbase/pom.xml, line 837, column 15\r\n\r\n ","from":"developer"},{"body":"There is a trivial fix. I am committing it everywhere. WIll post the addendum patch here","from":"developer"}],"created":"2018-08-31T10:15:31.000+0000","description":"License check via enforce plugin throws following error during build on windows:\r\n{code:java}\r\nSourced file: inline evaluation of: ``File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-ar . . . '' Token Parsing Error: Lexical error at line 1, column 29.  Encountered: \"D\" (68), after : \"\\\"D:\\\\\": {code}\r\nComplete stacktrace with command\r\n{code:java}\r\nmvn clean install -DskipTests -X\r\n{code}\r\nis as follows:\r\n{noformat}\r\n[INFO] --- maven-enforcer-plugin:3.0.0-M1:enforce (check-aggregate-license) @ hbase-shaded ---\r\n[DEBUG] Configuring mojo org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce from plugin realm ClassRealm[plugin>org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1, parent: sun.misc.Launcher$AppClassLoader@55f96302]\r\n[DEBUG] Configuring mojo 'org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce' with basic configurator -->\r\n[DEBUG] (s) fail = true\r\n[DEBUG] (s) failFast = false\r\n[DEBUG] (f) ignoreCache = false\r\n[DEBUG] (f) mojoExecution = org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce {execution: check-aggregate-license}\r\n[DEBUG] (s) project = MavenProject: org.apache.hbase:hbase-shaded:2.1.1-SNAPSHOT @ D:\\DS\\HBase_2\\hbase\\hbase-shaded\\pom.xml\r\n[DEBUG] (s) condition = File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\");\r\n\r\n // Beanshell does not support try-with-resources,\r\n // so we must close this scanner manually\r\n Scanner scanner = new Scanner(license);\r\n\r\n while (scanner.hasNextLine()) {\r\n if (scanner.nextLine().startsWith(\"ERROR:\")) {\r\n scanner.close();\r\n return false;\r\n }\r\n }\r\n scanner.close();\r\n return true;\r\n[DEBUG] (s) message = License errors detected, for more detail find ERROR in\r\n D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\r\n[DEBUG] (s) rules = [org.apache.maven.plugins.enforcer.EvaluateBeanshell@7e307087]\r\n[DEBUG] (s) session = org.apache.maven.execution.MavenSession@5e1218b4\r\n[DEBUG] (s) skip = false\r\n[DEBUG] -- end configuration --\r\n[DEBUG] Executing rule: org.apache.maven.plugins.enforcer.EvaluateBeanshell\r\n[DEBUG] Echo condition : File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\");\r\n\r\n // Beanshell does not support try-with-resources,\r\n // so we must close this scanner manually\r\n Scanner scanner = new Scanner(license);\r\n\r\n while (scanner.hasNextLine()) {\r\n if (scanner.nextLine().startsWith(\"ERROR:\")) {\r\n scanner.close();\r\n return false;\r\n }\r\n }\r\n scanner.close();\r\n return true;\r\n[DEBUG] Echo script : File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\");\r\n\r\n // Beanshell does not support try-with-resources,\r\n // so we must close this scanner manually\r\n Scanner scanner = new Scanner(license);\r\n\r\n while (scanner.hasNextLine()) {\r\n if (scanner.nextLine().startsWith(\"ERROR:\")) {\r\n scanner.close();\r\n return false;\r\n }\r\n }\r\n scanner.close();\r\n return true;\r\n[DEBUG] Adding failure due to exception\r\norg.apache.maven.enforcer.rule.api.EnforcerRuleException: Couldn't evaluate condition: File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\");\r\n\r\n // Beanshell does not support try-with-resources,\r\n // so we must close this scanner manually\r\n Scanner scanner = new Scanner(license);\r\n\r\n while (scanner.hasNextLine()) {\r\n if (scanner.nextLine().startsWith(\"ERROR:\")) {\r\n scanner.close();\r\n return false;\r\n }\r\n }\r\n scanner.close();\r\n return true;\r\n at org.apache.maven.plugins.enforcer.EvaluateBeanshell.evaluateCondition(EvaluateBeanshell.java:107)\r\n at org.apache.maven.plugins.enforcer.EvaluateBeanshell.execute(EvaluateBeanshell.java:72)\r\n at org.apache.maven.plugins.enforcer.EnforceMojo.execute(EnforceMojo.java:202)\r\n at org.apache.maven.plugin.DefaultBuildPluginManager.executeMojo(DefaultBuildPluginManager.java:134)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:208)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:154)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:146)\r\n at org.apache.maven.lifecycle.internal.LifecycleModuleBuilder.buildProject(LifecycleModuleBuilder.java:117)\r\n at org.apache.maven.lifecycle.internal.LifecycleModuleBuilder.buildProject(LifecycleModuleBuilder.java:81)\r\n at org.apache.maven.lifecycle.internal.builder.singlethreaded.SingleThreadedBuilder.build(SingleThreadedBuilder.java:51)\r\n at org.apache.maven.lifecycle.internal.LifecycleStarter.execute(LifecycleStarter.java:128)\r\n at org.apache.maven.DefaultMaven.doExecute(DefaultMaven.java:309)\r\n at org.apache.maven.DefaultMaven.doExecute(DefaultMaven.java:194)\r\n at org.apache.maven.DefaultMaven.execute(DefaultMaven.java:107)\r\n at org.apache.maven.cli.MavenCli.execute(MavenCli.java:993)\r\n at org.apache.maven.cli.MavenCli.doMain(MavenCli.java:345)\r\n at org.apache.maven.cli.MavenCli.main(MavenCli.java:191)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.launchEnhanced(Launcher.java:289)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.launch(Launcher.java:229)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.mainWithExitCode(Launcher.java:415)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.main(Launcher.java:356)\r\nCaused by: Sourced file: inline evaluation of: ``File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-ar . . . '' Token Parsing Error: Lexical error at line 1, column 29. Encountered: \"D\" (68), after : \"\\\"D:\\\\\": \r\n\r\n at bsh.Interpreter.eval(Unknown Source)\r\n at bsh.Interpreter.eval(Unknown Source)\r\n at bsh.Interpreter.eval(Unknown Source)\r\n at org.apache.maven.plugins.enforcer.EvaluateBeanshell.evaluateCondition(EvaluateBeanshell.java:102)\r\n ... 24 more\r\n[WARNING] Rule 0: org.apache.maven.plugins.enforcer.EvaluateBeanshell failed with message:\r\nCouldn't evaluate condition: File license = new File(\"D:\\DS\\HBase_2\\hbase\\hbase-shaded\\target/maven-shared-archive-resources/META-INF/LICENSE\");\r\n\r\n // Beanshell does not support try-with-resources,\r\n // so we must close this scanner manually\r\n Scanner scanner = new Scanner(license);\r\n\r\n while (scanner.hasNextLine()) {\r\n if (scanner.nextLine().startsWith(\"ERROR:\")) {\r\n scanner.close();\r\n return false;\r\n }\r\n }\r\n scanner.close();\r\n return true;\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Reactor Summary:\r\n[INFO]\r\n[INFO] Apache HBase - Shaded .............................. FAILURE [ 2.388 s]\r\n[INFO] Apache HBase - Shaded - Client (with Hadoop bundled) SKIPPED\r\n[INFO] Apache HBase - Shaded - Client ..................... SKIPPED\r\n[INFO] Apache HBase - Shaded - MapReduce .................. SKIPPED\r\n[INFO] Apache HBase - External Block Cache ................ SKIPPED\r\n[INFO] Apache HBase - Assembly ............................ SKIPPED\r\n[INFO] Apache HBase Shaded Packaging Invariants ........... SKIPPED\r\n[INFO] Apache HBase Shaded Packaging Invariants (with Hadoop bundled) SKIPPED\r\n[INFO] Apache HBase - Archetypes .......................... SKIPPED\r\n[INFO] Apache HBase - Exemplar for hbase-client archetype . SKIPPED\r\n[INFO] Apache HBase - Exemplar for hbase-shaded-client archetype SKIPPED\r\n[INFO] Apache HBase - Archetype builder ................... SKIPPED\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] BUILD FAILURE\r\n[INFO] ------------------------------------------------------------------------\r\n[INFO] Total time: 4.822 s\r\n[INFO] Finished at: 2018-08-31T15:39:00+05:30\r\n[INFO] Final Memory: 41M/421M\r\n[INFO] ------------------------------------------------------------------------\r\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce (check-aggregate-license) on project hbase-shaded: Some Enforcer rules have failed. Look above for specific messages explaining why the rule failed. -> [Help 1]\r\norg.apache.maven.lifecycle.LifecycleExecutionException: Failed to execute goal org.apache.maven.plugins:maven-enforcer-plugin:3.0.0-M1:enforce (check-aggregate-license) on project hbase-shaded: Some Enforcer rules have failed. Look above for specific messages explaining why the rule failed.\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:213)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:154)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:146)\r\n at org.apache.maven.lifecycle.internal.LifecycleModuleBuilder.buildProject(LifecycleModuleBuilder.java:117)\r\n at org.apache.maven.lifecycle.internal.LifecycleModuleBuilder.buildProject(LifecycleModuleBuilder.java:81)\r\n at org.apache.maven.lifecycle.internal.builder.singlethreaded.SingleThreadedBuilder.build(SingleThreadedBuilder.java:51)\r\n at org.apache.maven.lifecycle.internal.LifecycleStarter.execute(LifecycleStarter.java:128)\r\n at org.apache.maven.DefaultMaven.doExecute(DefaultMaven.java:309)\r\n at org.apache.maven.DefaultMaven.doExecute(DefaultMaven.java:194)\r\n at org.apache.maven.DefaultMaven.execute(DefaultMaven.java:107)\r\n at org.apache.maven.cli.MavenCli.execute(MavenCli.java:993)\r\n at org.apache.maven.cli.MavenCli.doMain(MavenCli.java:345)\r\n at org.apache.maven.cli.MavenCli.main(MavenCli.java:191)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.launchEnhanced(Launcher.java:289)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.launch(Launcher.java:229)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.mainWithExitCode(Launcher.java:415)\r\n at org.codehaus.plexus.classworlds.launcher.Launcher.main(Launcher.java:356)\r\nCaused by: org.apache.maven.plugin.MojoExecutionException: Some Enforcer rules have failed. Look above for specific messages explaining why the rule failed.\r\n at org.apache.maven.plugins.enforcer.EnforceMojo.execute(EnforceMojo.java:243)\r\n at org.apache.maven.plugin.DefaultBuildPluginManager.executeMojo(DefaultBuildPluginManager.java:134)\r\n at org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:208)\r\n ... 20 more\r\n[ERROR]\r\n[ERROR]\r\n[ERROR] For more information about the errors and possible solutions, please read the following articles:\r\n[ERROR] [Help 1] http://cwiki.apache.org/confluence/display/MAVEN/MojoExecutionException\r\n{noformat}\r\n ","issue_id":"13182328","key":"HBASE-21135","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2019-03-28T17:14:10.000+0000","role":"fixed_distractor","summary":"Build fails on windows as it fails to parse windows path during license check"} {"case_id":"13198882","cluster":"DISTRACTOR-HBASE-21487","comments":[{"body":"Agree that the two procedures should aggregate rather than be an either/or. You have a suggested fix [~arshiya9414]? Thanks.","created":"2018-11-16T16:16:47.407+0000"},{"body":"Issue is that new descriptor is being built before submitting the ModifyTableProcedure in HMaster class.In order to maintain the consistency ,the new descriptor can be built inside ModifyTableProcedure class but the class expects new descriptor and also the coprocessor hooks are maintained in HMaster class to which the solution cannot be met or should we make the design as like earlier with many procedures(AddColumn/ModifyColumn/DeleteColumn) which will solve the problem?","created":"2018-11-23T14:38:38.429+0000"},{"body":"Yes, it is a problem, in 2.x, we only have ModifyTableProcedure, and the Tabledescriptor is passed in with no lock. I think maybe we should acquire the table exclusive lock before constructing the new table descriptor to ensure we are using the newest version. Are you preparing a patch, [~arshiya9414]?","created":"2018-11-23T16:21:52.323+0000"},{"body":"No, I'm not preparing any patch as of now,I was just thinking about the solutions and did not get any clear approach.\r\nwhen do we release the table lock if we grab the lock before constructing the new table descriptor? As per my understanding, when we release the lock before submitting the procedure, the issue wont be solved as the descriptor will override and it will lead to same unexpected results.If we release after submitting the procedure,the ModifyTableProcedure won't start until the table lock is released.Is this right?","created":"2018-11-26T14:00:52.014+0000"},{"body":"I think we could record the seqenceID of the current tabledescriptor on FS(and pass to ModifyTableProcedure), then in prepare state of ModifyTableProcedure, we can check whether if the sequenceID of the current table descriptor from FS is different from the ID we recorded. If they are not equal, that means someone has already changed the descriptor, we can abort the modify procedure if so.","created":"2018-11-26T15:55:33.212+0000"},{"body":"From the above solution only one operation will be succeeded because we will abort one of the procedure if ID is not equal.Is this acceptable behavior? Or else after checking the sequence ID of descriptor in prepare state,if its not equal can we just throw exception and retry again by building new descriptor and submitting procedure again?","created":"2018-11-27T11:37:15.320+0000"},{"body":"[~allan163] Do you have any suggestions on the issue?","created":"2018-12-03T06:29:53.313+0000"},{"body":"{quote}\r\nFrom the above solution only one operation will be succeeded because we will abort one of the procedure if ID is not equal.Is this acceptable behavior?\r\n{quote}\r\nYes, I think it is acceptable\r\n{quote}\r\nOr else after checking the sequence ID of descriptor in prepare state,if its not equal can we just throw exception and retry again by building new descriptor and submitting procedure again?\r\n{quote}\r\nSince there is only one kind of procedure(ModifyTableProcedure) here, when it is begin to execute, it doesn't know how to build the new descriptor, unless we record some info into the procedure.","created":"2018-12-03T06:37:56.469+0000"},{"body":"{quote}record the seqenceID of the current tabledescriptor on FS(and pass to ModifyTableProcedure),\r\n{quote}\r\nI think , SequenceID concept is specific to FSTableDescriptors implementation which is one of the impl of TableDescriptors interface. So IMO we should not depend on sequenceID here for comparison. \r\n\r\nMay be we can pass the old_table_descriptor also in ModifyTableProcedure. In MODIFY_TABLE_PREPARE step, we can compare the old_table_descriptor with current_table_descriptor, if they are not same then we can throw exception ... \r\n\r\n\r\n","created":"2018-12-04T08:48:46.132+0000"},{"body":"[~allan163], any thoughts on the issue?","created":"2018-12-09T12:16:16.493+0000"},{"body":"{quote}\r\nMay be we can pass the old_table_descriptor also in ModifyTableProcedure. In MODIFY_TABLE_PREPARE step, we can compare the old_table_descriptor with current_table_descriptor, if they are not same then we can throw exception ...\r\n{quote}\r\nI think you are right, it is better to compare table descriptors directly.","created":"2018-12-09T13:38:07.815+0000"},{"body":"Attached patch on behalf of [~arshiya9414].\r\n\r\n[~allan163], can you please add [~arshiya9414] as a HBase contributer and assign this Jira to her.\r\n\r\n ","created":"2019-01-17T13:38:30.668+0000"},{"body":"Ping [~allan163], kindly review the patch.","created":"2019-01-17T13:39:28.247+0000"},{"body":"I missed to add newly added file.Addressed the same with v2 patch.","created":"2019-01-18T04:33:08.557+0000"},{"body":"{code}\r\n+ public ModifyTableProcedure(final MasterProcedureEnv env,\r\n+ final TableDescriptor newTableDescriptor, final ProcedurePrepareLatch latch,\r\n+ final TableDescriptor oldTableDescriptor, final boolean shouldCheckDescriptor)\r\n+ throws HBaseIOException {\r\n+ this(env, newTableDescriptor, latch);\r\n+ this.unmodifiedTableDescriptor = oldTableDescriptor;\r\n+ this.shouldCheckDescriptor = shouldCheckDescriptor;\r\n+ }\r\n+\r\n{code}\r\nConstructor with less arguments should call constructors with more arguments,and extra arguments use default values, not the way around like in the patch. Except this, the patch looks great.","created":"2019-02-22T06:46:37.077+0000"},{"body":"Thanks for reviewing [~allan163], addressed the above comment in the .05 patch.","created":"2019-02-22T11:28:20.484+0000"},{"body":"{code:java}\r\nif (shouldCheckDescriptor) {\r\n2653\tsubmitProcedure(new ModifyTableProcedure(procedureExecutor.getEnvironment(),\r\n2654\tnewDescriptor, latch, oldDescriptor, shouldCheckDescriptor));\r\n2655\t} else {\r\n2656\tsubmitProcedure(\r\n2657\tnew ModifyTableProcedure(procedureExecutor.getEnvironment(), newDescriptor, latch));\r\n2658\t}\r\n{code}\r\nsubmitProcedure( new ModifyTableProcedure(procedureExecutor.getEnvironment(), newDescriptor, latch, shouldCheckDescriptor)); directly? No need to check shouldCheckDescriptor.","created":"2019-02-26T08:50:15.943+0000"},{"body":" \r\n{code:java}\r\npreflightChecks(env, null/* No table checks; if changing peers, table can be online */);\r\nthis.modifiedTableDescriptor = newTableDescriptor;\r\n{code}\r\nThe preflightChecks should after this.modifiedTableDescriptor = newTableDescriptor; As the check will call getTableName and it need modifiedTableDescriptor.\r\n\r\n \r\n{code:java}\r\nprivate void initilize() {code}\r\nAdd two parameter unmodifiedTableDescriptor and shouldCheckDescriptor for it? And no need to assign them again.\r\n\r\n ","created":"2019-02-26T09:03:47.478+0000"},{"body":"Addressed the comments in the 06 patch.Thanks [~zghaobac]","created":"2019-02-28T10:23:56.435+0000"},{"body":"+1.","created":"2019-02-28T10:34:29.876+0000"},{"body":"Pushed to branch-2.2+. Thanks [~arshiya9414] for contributing.","created":"2019-03-02T05:26:49.616+0000"}],"conversations":[{"body":"Concurrent modifyTable or add/delete/modify columnFamily leads to incorrect result. After HBASE-18893, The behavior of add/delete/modify column family during concurrent operation is changed compare to branch-1.When one client is adding cf2 and another one cf3 .. In branch-1 final result will be cf1,cf2,cf3 but now either cf1,cf2 OR cf1,cf3 will be the outcome depending on which ModifyTableProcedure executed finally.Its because new table descriptor is constructed before submitting the ModifyTableProcedure in HMaster class and its not guarded by any lock.\r\n\r\n*Steps to reproduce*\r\n\r\n1.Create table 't' with column family 'f1'\r\n2.Client-1 and Client-2 requests to add column family 'f2' and 'f3' on table 't' concurrently.\r\n\r\n*Expected Result*\r\nTable should have three column families(f1,f2,f3)\r\n\r\n*Actual Result*\r\nTable 't' will have column family either (f1,f2) or (f1,f3)\r\n","from":"reporter","subject":"Concurrent modify table ops can lead to unexpected results"},{"body":"Agree that the two procedures should aggregate rather than be an either/or. You have a suggested fix [~arshiya9414]? Thanks.","from":"developer"},{"body":"Issue is that new descriptor is being built before submitting the ModifyTableProcedure in HMaster class.In order to maintain the consistency ,the new descriptor can be built inside ModifyTableProcedure class but the class expects new descriptor and also the coprocessor hooks are maintained in HMaster class to which the solution cannot be met or should we make the design as like earlier with many procedures(AddColumn/ModifyColumn/DeleteColumn) which will solve the problem?","from":"developer"},{"body":"Yes, it is a problem, in 2.x, we only have ModifyTableProcedure, and the Tabledescriptor is passed in with no lock. I think maybe we should acquire the table exclusive lock before constructing the new table descriptor to ensure we are using the newest version. Are you preparing a patch, [~arshiya9414]?","from":"developer"},{"body":"No, I'm not preparing any patch as of now,I was just thinking about the solutions and did not get any clear approach.\r\nwhen do we release the table lock if we grab the lock before constructing the new table descriptor? As per my understanding, when we release the lock before submitting the procedure, the issue wont be solved as the descriptor will override and it will lead to same unexpected results.If we release after submitting the procedure,the ModifyTableProcedure won't start until the table lock is released.Is this right?","from":"developer"},{"body":"I think we could record the seqenceID of the current tabledescriptor on FS(and pass to ModifyTableProcedure), then in prepare state of ModifyTableProcedure, we can check whether if the sequenceID of the current table descriptor from FS is different from the ID we recorded. If they are not equal, that means someone has already changed the descriptor, we can abort the modify procedure if so.","from":"developer"},{"body":"From the above solution only one operation will be succeeded because we will abort one of the procedure if ID is not equal.Is this acceptable behavior? Or else after checking the sequence ID of descriptor in prepare state,if its not equal can we just throw exception and retry again by building new descriptor and submitting procedure again?","from":"developer"},{"body":"[~allan163] Do you have any suggestions on the issue?","from":"developer"},{"body":"{quote}\r\nFrom the above solution only one operation will be succeeded because we will abort one of the procedure if ID is not equal.Is this acceptable behavior?\r\n{quote}\r\nYes, I think it is acceptable\r\n{quote}\r\nOr else after checking the sequence ID of descriptor in prepare state,if its not equal can we just throw exception and retry again by building new descriptor and submitting procedure again?\r\n{quote}\r\nSince there is only one kind of procedure(ModifyTableProcedure) here, when it is begin to execute, it doesn't know how to build the new descriptor, unless we record some info into the procedure.","from":"developer"},{"body":"{quote}record the seqenceID of the current tabledescriptor on FS(and pass to ModifyTableProcedure),\r\n{quote}\r\nI think , SequenceID concept is specific to FSTableDescriptors implementation which is one of the impl of TableDescriptors interface. So IMO we should not depend on sequenceID here for comparison. \r\n\r\nMay be we can pass the old_table_descriptor also in ModifyTableProcedure. In MODIFY_TABLE_PREPARE step, we can compare the old_table_descriptor with current_table_descriptor, if they are not same then we can throw exception ... \r\n\r\n\r\n","from":"developer"},{"body":"[~allan163], any thoughts on the issue?","from":"developer"},{"body":"{quote}\r\nMay be we can pass the old_table_descriptor also in ModifyTableProcedure. In MODIFY_TABLE_PREPARE step, we can compare the old_table_descriptor with current_table_descriptor, if they are not same then we can throw exception ...\r\n{quote}\r\nI think you are right, it is better to compare table descriptors directly.","from":"developer"},{"body":"Attached patch on behalf of [~arshiya9414].\r\n\r\n[~allan163], can you please add [~arshiya9414] as a HBase contributer and assign this Jira to her.\r\n\r\n ","from":"developer"},{"body":"Ping [~allan163], kindly review the patch.","from":"developer"},{"body":"I missed to add newly added file.Addressed the same with v2 patch.","from":"developer"},{"body":"{code}\r\n+ public ModifyTableProcedure(final MasterProcedureEnv env,\r\n+ final TableDescriptor newTableDescriptor, final ProcedurePrepareLatch latch,\r\n+ final TableDescriptor oldTableDescriptor, final boolean shouldCheckDescriptor)\r\n+ throws HBaseIOException {\r\n+ this(env, newTableDescriptor, latch);\r\n+ this.unmodifiedTableDescriptor = oldTableDescriptor;\r\n+ this.shouldCheckDescriptor = shouldCheckDescriptor;\r\n+ }\r\n+\r\n{code}\r\nConstructor with less arguments should call constructors with more arguments,and extra arguments use default values, not the way around like in the patch. Except this, the patch looks great.","from":"developer"},{"body":"Thanks for reviewing [~allan163], addressed the above comment in the .05 patch.","from":"developer"},{"body":"{code:java}\r\nif (shouldCheckDescriptor) {\r\n2653\tsubmitProcedure(new ModifyTableProcedure(procedureExecutor.getEnvironment(),\r\n2654\tnewDescriptor, latch, oldDescriptor, shouldCheckDescriptor));\r\n2655\t} else {\r\n2656\tsubmitProcedure(\r\n2657\tnew ModifyTableProcedure(procedureExecutor.getEnvironment(), newDescriptor, latch));\r\n2658\t}\r\n{code}\r\nsubmitProcedure( new ModifyTableProcedure(procedureExecutor.getEnvironment(), newDescriptor, latch, shouldCheckDescriptor)); directly? No need to check shouldCheckDescriptor.","from":"developer"},{"body":" \r\n{code:java}\r\npreflightChecks(env, null/* No table checks; if changing peers, table can be online */);\r\nthis.modifiedTableDescriptor = newTableDescriptor;\r\n{code}\r\nThe preflightChecks should after this.modifiedTableDescriptor = newTableDescriptor; As the check will call getTableName and it need modifiedTableDescriptor.\r\n\r\n \r\n{code:java}\r\nprivate void initilize() {code}\r\nAdd two parameter unmodifiedTableDescriptor and shouldCheckDescriptor for it? And no need to assign them again.\r\n\r\n ","from":"developer"},{"body":"Addressed the comments in the 06 patch.Thanks [~zghaobac]","from":"developer"},{"body":"+1.","from":"developer"},{"body":"Pushed to branch-2.2+. Thanks [~arshiya9414] for contributing.","from":"developer"}],"created":"2018-11-16T11:28:16.000+0000","description":"Concurrent modifyTable or add/delete/modify columnFamily leads to incorrect result. After HBASE-18893, The behavior of add/delete/modify column family during concurrent operation is changed compare to branch-1.When one client is adding cf2 and another one cf3 .. In branch-1 final result will be cf1,cf2,cf3 but now either cf1,cf2 OR cf1,cf3 will be the outcome depending on which ModifyTableProcedure executed finally.Its because new table descriptor is constructed before submitting the ModifyTableProcedure in HMaster class and its not guarded by any lock.\r\n\r\n*Steps to reproduce*\r\n\r\n1.Create table 't' with column family 'f1'\r\n2.Client-1 and Client-2 requests to add column family 'f2' and 'f3' on table 't' concurrently.\r\n\r\n*Expected Result*\r\nTable should have three column families(f1,f2,f3)\r\n\r\n*Actual Result*\r\nTable 't' will have column family either (f1,f2) or (f1,f3)\r\n","issue_id":"13198882","key":"HBASE-21487","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2019-03-02T05:26:49.000+0000","role":"fixed_distractor","summary":"Concurrent modify table ops can lead to unexpected results"} {"case_id":"13205431","cluster":"DISTRACTOR-HBASE-21620","comments":[{"body":"[~mohamed.meeran], I written a UT for it, Yeah , you are right, it's a bug.... will provide a patch for this. ","created":"2018-12-20T03:41:07.017+0000"},{"body":"Catch the stack in regionserver : \r\n{code}\r\n\"RpcServer.default.FPBQ.Fifo.handler=4,queue=0,port=39128\" #154 daemon prio=5 os_prio=0 tid=0x00007fb875839000 nid=0x202d runnable [0x00007fb7535f3000]\r\n java.lang.Thread.State: RUNNABLE\r\n at org.apache.hadoop.hbase.filter.FilterListBase.compareCell(FilterListBase.java:86)\r\n at org.apache.hadoop.hbase.filter.FilterListWithOR.getNextCellHint(FilterListWithOR.java:371)\r\n at org.apache.hadoop.hbase.filter.FilterList.getNextCellHint(FilterList.java:265)\r\n at org.apache.hadoop.hbase.regionserver.querymatcher.UserScanQueryMatcher.getNextKeyHint(UserScanQueryMatcher.java:96)\r\n at org.apache.hadoop.hbase.regionserver.StoreScanner.next(StoreScanner.java:686)\r\n at org.apache.hadoop.hbase.regionserver.KeyValueHeap.next(KeyValueHeap.java:152)\r\n at org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl.populateResult(HRegion.java:6292)\r\n at org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl.nextInternal(HRegion.java:6452)\r\n at org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl.nextRaw(HRegion.java:6224)\r\n at org.apache.hadoop.hbase.regionserver.RSRpcServices.scan(RSRpcServices.java:2882)\r\n - locked <0x00000006cc21a338> (a org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl)\r\n at org.apache.hadoop.hbase.regionserver.RSRpcServices.scan(RSRpcServices.java:3131)\r\n at org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:36613)\r\n at org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2380)\r\n at org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:124)\r\n at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:297)\r\n at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:277)\r\n{code}","created":"2018-12-20T03:49:47.753+0000"},{"body":"[~mohamed.meeran], I think HBASE-201620.branch-1.patch can help you. Let's see what Hadoop QA says. ","created":"2018-12-20T08:50:37.364+0000"},{"body":"Thanks [~openinx]. It is working in our environment with the given patch.. ","created":"2018-12-20T10:22:22.112+0000"},{"body":"Discussed with Guanghao, we can use the preivous return code and cell to calculate the return code in some case, then we can push the filterlist's return code in a further step. So upload the patch.v1.","created":"2018-12-20T14:15:14.878+0000"},{"body":"This looks like a great fix but it looks risky putting it into branch-2.0 or branch-2.1? All tests did pass which is a strong argument in favor and there are two of you discussing this change. What you think lads? Should it go in? The stuff that troubles me is changes like this:\r\n\r\n if (filter.filterAllRemaining() || !shouldPassCurrentCellToFilter(prevCell, c, prevCode)) {\r\n\r\nwhere we remove the shouldPassCurrentCellToFilter bit in the patch.\r\n\r\nThanks.\r\n","created":"2018-12-21T01:56:22.864+0000"},{"body":"It's a critial bug. when I run the UT without the fix, my RegionServer was caught in an infinite loop. to be clear, it was in StoreScanner#next(...) . the filterList return a hint cell which was even smaller than the current cell in heap.peak(), so we got stuck here (the cpu cost too much and high load )\r\n{code}\r\n case SEEK_NEXT_USING_HINT:\r\n Cell nextKV = matcher.getNextKeyHint(cell);\r\n if (nextKV != null) {\r\n seekAsDirection(nextKV);\r\n NextState stateAfterSeekByHint = needToReturn(outResult);\r\n if (stateAfterSeekByHint != null) {\r\n return scannerContext.setScannerState(stateAfterSeekByHint).hasMoreValues();\r\n }\r\n } else {\r\n heap.next();\r\n }\r\n{code}\r\nbq. What you think lads? Should it go in? \r\nSo I think we should let this go in branch-2.0 and branch2.1. \r\nbq. The stuff that troubles me is changes like this: ... where we remove the shouldPassCurrentCellToFilter bit in the patch.\r\nIn fact, we did not remove the shouldPassCurrentCellToFilter. I refactor the shouldPassCurrentCellToFilter to calculateReturnCodeByPrevCellAndRC, which means will calculate the return code when testing shouldPassCurrentCellToFilter based on the previous return code and previous cell. you can see here: \r\n{code}\r\n- ReturnCode localRC = filter.filterCell(c);\r\n+ ReturnCode localRC = calculateReturnCodeByPrevCellAndRC(filter, c, prevCell, prevCode);\r\n+ if (localRC == null) {\r\n+ // Can not get return code based on previous cell and previous return code. In other words,\r\n+ // we should pass the current cell to this sub-filter to get the return code, and it won't\r\n+ // impact the sub-filter's internal state.\r\n+ localRC = filter.filterCell(c);\r\n+ }\r\n{code}\r\nThanks.","created":"2018-12-21T02:11:56.939+0000"},{"body":"The TestRegionServerAbortTimeout is unrelated to this patch, and passed under my local. [~zghaobac] any other concerns ? If no, the checkstyle will be fixed when committing.","created":"2018-12-21T06:42:52.413+0000"},{"body":"[~openinx] Can you please tell for which Version of HBase you provided the v1 and v2 patch? We tried applying the v2.patch in HBase-1.4.8 but it results in Compilation error. The Filter instance doesn't have filterCell() method.\r\n\r\nThough the filterCell() method is found in 2.0.0. Can you provide the patch for 1.4.8? ","created":"2018-12-21T07:07:29.945+0000"},{"body":"bq. Can you please tell for which Version of HBase you provided the v1 and v2 patch? \r\n[~mohamed.meeran], The patch is still be rewiewing, so please wait a moment. patch for branch-1 will upload soon too. Thanks. ","created":"2018-12-21T09:38:33.112+0000"},{"body":"It looks like a ugly bug, +1 for branch-2.0 and branch-2.1 ","created":"2018-12-21T09:55:18.574+0000"},{"body":"+1","created":"2018-12-21T10:37:38.084+0000"},{"body":"+1\r\n\r\nWill retry the hadoopqa.","created":"2018-12-21T17:53:25.112+0000"},{"body":"+1 for branch-2.0+ given can make for infinite loop.","created":"2018-12-21T17:54:22.937+0000"},{"body":"Resolving after fixing checkstyle on commit to branch-2.0+ (so can roll RC). Opened sub-issue for backport to branch-1.","created":"2018-12-21T23:24:27.532+0000"},{"body":"[~openinx] We upgraded hbase from 1.2.5 to 1.4.8 with the patch you provided, in our production servers. We identified a performance degrade in Scanning a long row(around half a million columns) with multiple Column Prefix Filters(around 10 filters). The scan now takes around 1.2 seconds, while the same scan in 1.2.5 took only 0.3 seconds max. We've included the column keys in the attached file(columnkey.txt). Is there a way to increase the performance?\r\n\r\n \r\n\r\n*Table description:*\r\n\r\n{NAME => 'MCF', DATA_BLOCK_ENCODING => 'NONE', BLOOMFILTER => 'ROWCOL', COMPRESSION => 'SNAPPY', VERSIONS => '1', TTL => '2592000 SECONDS (30 DAYS)', MIN_VERSIONS => '1', KEEP_DELETED_CELLS => 'TTL', BLOCKSIZE => '65536', IN_MEMORY => 'false', BLOCKCACHE => 'true'}   \r\n\r\n \r\n\r\n*Scan query :*\r\n\r\nscan 'namespace:tablename', \\{ STARTROW => 'row', ENDROW => 'row', FILTER => \"ColumnPrefixFilter('1545212621603120001_') OR ColumnPrefixFilter('1546841752667120001_') OR ColumnPrefixFilter('1545387980301120001_') OR ColumnPrefixFilter('1544677866436120001_') OR ColumnPrefixFilter('1546252381017120001_') OR ColumnPrefixFilter('1546247122010120001_') OR ColumnPrefixFilter('2866449000003420001_w') OR ColumnPrefixFilter('1545221612425120001_') OR ColumnPrefixFilter('1546582197395120001_') OR ColumnPrefixFilter('1545798618753120001_')\"} \r\n\r\n ","created":"2019-01-08T13:00:38.650+0000"},{"body":"I guess the bottleneck is here, Let me have a test..\r\n{code}\r\ndiff --git a/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java b/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\r\nindex d51fdf0701..f0fbf3e8a1 100644\r\n--- a/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\r\n+++ b/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\r\n@@ -684,7 +684,7 @@ public class StoreScanner extends NonReversedNonLazyKeyValueScanner\r\n \r\n case SEEK_NEXT_USING_HINT:\r\n Cell nextKV = matcher.getNextKeyHint(cell);\r\n- if (nextKV != null) {\r\n+ if (nextKV != null && comparator.compare(nextKV, cell) > 0) {\r\n seekAsDirection(nextKV);\r\n NextState stateAfterSeekByHint = needToReturn(outResult);\r\n if (stateAfterSeekByHint != null) {\r\n{code}","created":"2019-01-08T13:43:03.444+0000"},{"body":"[~openinx] any update on this?","created":"2019-01-10T07:26:19.846+0000"},{"body":"Not yet, you can try to have a simple test after remove the && comparator.compare(nextKV, cell) > 0 , and compare the performance. ","created":"2019-01-10T08:53:07.033+0000"},{"body":"[~openinx] we removed the *(comparator.compare(nextKV, cell) > 0)* condition and tested, but still the scan is slow.\r\n\r\n ","created":"2019-01-11T07:24:17.309+0000"},{"body":"Any improvement ? another reason maybe: the filter list concated by OR will choose the minimal forward step among all sub-filters. in this patch, we have stricter restrictions on all sub filters, include those sub-filter whose has non-null RC returned in calculateReturnCodeByPrevCellAndRC (previously, we will skip to merge this sub-filter's rc, but it's wrong in some case), and merge all of the sub-filter's RC, this is also some time cost.","created":"2019-01-12T09:54:55.507+0000"},{"body":"[~openinx] There was no improvement even after removing the compare condition. Is there any other way to optimize the scan?","created":"2019-01-14T11:42:11.238+0000"},{"body":"We tried the same scan using HBase-2.1.2. The scan is even more slower than 1.4.8 (It takes around 3-4 seconds now). Because of this issue we are unable to update our production clusters from 1.2.5 to a higher version. [~openinx] Can you please provide a fix for this?\r\n\r\n ","created":"2019-01-17T07:53:55.688+0000"},{"body":"[~KarthickRam] , How can I import your columnkey.txt to my hbase cluster to reproduce your case ? ","created":"2019-01-17T12:07:38.238+0000"},{"body":"[~openinx] please use this (HBaseFileImport.java) file to import the columns in your cluster. Please replace the ZK Quorum, topnode and namespace:tablename. ","created":"2019-01-17T12:32:01.539+0000"},{"body":"Thanks [~KarthickRam], I found some points to optimize for your case, and filed an issue HBASE-21734 for this. let's discuss there. ","created":"2019-01-17T13:38:44.367+0000"},{"body":"[~openinx] While scanning a row (around 10 lakhs columns) with  100 column prefixes , it takes around  4 secs  in hbase-1.2.5 and when same query is executed in hbase-1.4.9 it takes around 50 secs .\r\n\r\nAnyway to optimise this??\r\n\r\nP.S we have applied the patch provided in HBASE-21620 and  HBASE-21734 . Attached *qualifiers*.*txt* file , use the HBaseFileImport.java file provided to populate in your table and use *scanquery.txt* to query ","created":"2019-05-09T08:31:31.631+0000"},{"body":"How could I read your input files and verify the problem ? ","created":"2019-05-09T08:48:17.869+0000"},{"body":"[~openinx]  please use the attached *HBaseFileImport.java* file and change the filename from _\"columnkey.txt\"_ to _\"qualifiers.txt\"_ to import the columns in your table. You also have to replace the ZK quorum, topnode and namespace:tablename. After importing the columns you can use scan query given in *scanquery.txt* in hbase shell.","created":"2019-05-09T11:28:27.579+0000"},{"body":"[~openinx] Any update on this ?","created":"2019-05-12T05:31:27.261+0000"},{"body":"[~openinx] We are facing this issue in production. Any update on this?","created":"2019-05-21T09:23:11.888+0000"},{"body":"OK. Mind to open an seperate issue for this ? I will take some time later, still some working on hand :-) ","created":"2019-05-21T10:05:22.803+0000"},{"body":"[~openinx] I've raised a separate issue here [https://issues.apache.org/jira/browse/HBASE-22448].\r\n\r\n ","created":"2019-05-21T10:34:00.971+0000"}],"conversations":[{"body":"In some cases, unable to get the scan results when using more than one column prefix filter.\r\n\r\nAttached a java file to import the data which we used and a text file containing the values..\r\n\r\nWhile executing the following query (hbase shell as well as java program) it is waiting indefinitely and after RPC timeout we got the following error.. Also we noticed high cpu, high load average and very frequent young gc  in the region server containing this row...\r\n\r\nscan 'namespace:tablename',\\{STARTROW => 'test',ENDROW => 'test', FILTER => \"ColumnPrefixFilter('1544770422942010001_') OR ColumnPrefixFilter('1544769883529010001_')\"}\r\n\r\nROW                                                  COLUMN+CELL                                                                   ERROR: Call id=18, waitTime=60005, rpcTimetout=60000\r\n\r\n \r\n\r\nNote: Table scan operation and scan with a single column prefix filter works fine in this case.\r\n\r\nWhen we check the same query in hbase-1.2.5 it is working fine.\r\n\r\nCan you please help me on this..","from":"reporter","subject":"Problem in scan query when using more than one column prefix filter in some cases."},{"body":"[~mohamed.meeran], I written a UT for it, Yeah , you are right, it's a bug.... will provide a patch for this. ","from":"developer"},{"body":"Catch the stack in regionserver : \r\n{code}\r\n\"RpcServer.default.FPBQ.Fifo.handler=4,queue=0,port=39128\" #154 daemon prio=5 os_prio=0 tid=0x00007fb875839000 nid=0x202d runnable [0x00007fb7535f3000]\r\n java.lang.Thread.State: RUNNABLE\r\n at org.apache.hadoop.hbase.filter.FilterListBase.compareCell(FilterListBase.java:86)\r\n at org.apache.hadoop.hbase.filter.FilterListWithOR.getNextCellHint(FilterListWithOR.java:371)\r\n at org.apache.hadoop.hbase.filter.FilterList.getNextCellHint(FilterList.java:265)\r\n at org.apache.hadoop.hbase.regionserver.querymatcher.UserScanQueryMatcher.getNextKeyHint(UserScanQueryMatcher.java:96)\r\n at org.apache.hadoop.hbase.regionserver.StoreScanner.next(StoreScanner.java:686)\r\n at org.apache.hadoop.hbase.regionserver.KeyValueHeap.next(KeyValueHeap.java:152)\r\n at org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl.populateResult(HRegion.java:6292)\r\n at org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl.nextInternal(HRegion.java:6452)\r\n at org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl.nextRaw(HRegion.java:6224)\r\n at org.apache.hadoop.hbase.regionserver.RSRpcServices.scan(RSRpcServices.java:2882)\r\n - locked <0x00000006cc21a338> (a org.apache.hadoop.hbase.regionserver.HRegion$RegionScannerImpl)\r\n at org.apache.hadoop.hbase.regionserver.RSRpcServices.scan(RSRpcServices.java:3131)\r\n at org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:36613)\r\n at org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:2380)\r\n at org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:124)\r\n at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:297)\r\n at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:277)\r\n{code}","from":"developer"},{"body":"[~mohamed.meeran], I think HBASE-201620.branch-1.patch can help you. Let's see what Hadoop QA says. ","from":"developer"},{"body":"Thanks [~openinx]. It is working in our environment with the given patch.. ","from":"developer"},{"body":"Discussed with Guanghao, we can use the preivous return code and cell to calculate the return code in some case, then we can push the filterlist's return code in a further step. So upload the patch.v1.","from":"developer"},{"body":"This looks like a great fix but it looks risky putting it into branch-2.0 or branch-2.1? All tests did pass which is a strong argument in favor and there are two of you discussing this change. What you think lads? Should it go in? The stuff that troubles me is changes like this:\r\n\r\n if (filter.filterAllRemaining() || !shouldPassCurrentCellToFilter(prevCell, c, prevCode)) {\r\n\r\nwhere we remove the shouldPassCurrentCellToFilter bit in the patch.\r\n\r\nThanks.\r\n","from":"developer"},{"body":"It's a critial bug. when I run the UT without the fix, my RegionServer was caught in an infinite loop. to be clear, it was in StoreScanner#next(...) . the filterList return a hint cell which was even smaller than the current cell in heap.peak(), so we got stuck here (the cpu cost too much and high load )\r\n{code}\r\n case SEEK_NEXT_USING_HINT:\r\n Cell nextKV = matcher.getNextKeyHint(cell);\r\n if (nextKV != null) {\r\n seekAsDirection(nextKV);\r\n NextState stateAfterSeekByHint = needToReturn(outResult);\r\n if (stateAfterSeekByHint != null) {\r\n return scannerContext.setScannerState(stateAfterSeekByHint).hasMoreValues();\r\n }\r\n } else {\r\n heap.next();\r\n }\r\n{code}\r\nbq. What you think lads? Should it go in? \r\nSo I think we should let this go in branch-2.0 and branch2.1. \r\nbq. The stuff that troubles me is changes like this: ... where we remove the shouldPassCurrentCellToFilter bit in the patch.\r\nIn fact, we did not remove the shouldPassCurrentCellToFilter. I refactor the shouldPassCurrentCellToFilter to calculateReturnCodeByPrevCellAndRC, which means will calculate the return code when testing shouldPassCurrentCellToFilter based on the previous return code and previous cell. you can see here: \r\n{code}\r\n- ReturnCode localRC = filter.filterCell(c);\r\n+ ReturnCode localRC = calculateReturnCodeByPrevCellAndRC(filter, c, prevCell, prevCode);\r\n+ if (localRC == null) {\r\n+ // Can not get return code based on previous cell and previous return code. In other words,\r\n+ // we should pass the current cell to this sub-filter to get the return code, and it won't\r\n+ // impact the sub-filter's internal state.\r\n+ localRC = filter.filterCell(c);\r\n+ }\r\n{code}\r\nThanks.","from":"developer"},{"body":"The TestRegionServerAbortTimeout is unrelated to this patch, and passed under my local. [~zghaobac] any other concerns ? If no, the checkstyle will be fixed when committing.","from":"developer"},{"body":"[~openinx] Can you please tell for which Version of HBase you provided the v1 and v2 patch? We tried applying the v2.patch in HBase-1.4.8 but it results in Compilation error. The Filter instance doesn't have filterCell() method.\r\n\r\nThough the filterCell() method is found in 2.0.0. Can you provide the patch for 1.4.8? ","from":"developer"},{"body":"bq. Can you please tell for which Version of HBase you provided the v1 and v2 patch? \r\n[~mohamed.meeran], The patch is still be rewiewing, so please wait a moment. patch for branch-1 will upload soon too. Thanks. ","from":"developer"},{"body":"It looks like a ugly bug, +1 for branch-2.0 and branch-2.1 ","from":"developer"},{"body":"+1","from":"developer"},{"body":"+1\r\n\r\nWill retry the hadoopqa.","from":"developer"},{"body":"+1 for branch-2.0+ given can make for infinite loop.","from":"developer"},{"body":"Resolving after fixing checkstyle on commit to branch-2.0+ (so can roll RC). Opened sub-issue for backport to branch-1.","from":"developer"},{"body":"[~openinx] We upgraded hbase from 1.2.5 to 1.4.8 with the patch you provided, in our production servers. We identified a performance degrade in Scanning a long row(around half a million columns) with multiple Column Prefix Filters(around 10 filters). The scan now takes around 1.2 seconds, while the same scan in 1.2.5 took only 0.3 seconds max. We've included the column keys in the attached file(columnkey.txt). Is there a way to increase the performance?\r\n\r\n \r\n\r\n*Table description:*\r\n\r\n{NAME => 'MCF', DATA_BLOCK_ENCODING => 'NONE', BLOOMFILTER => 'ROWCOL', COMPRESSION => 'SNAPPY', VERSIONS => '1', TTL => '2592000 SECONDS (30 DAYS)', MIN_VERSIONS => '1', KEEP_DELETED_CELLS => 'TTL', BLOCKSIZE => '65536', IN_MEMORY => 'false', BLOCKCACHE => 'true'}   \r\n\r\n \r\n\r\n*Scan query :*\r\n\r\nscan 'namespace:tablename', \\{ STARTROW => 'row', ENDROW => 'row', FILTER => \"ColumnPrefixFilter('1545212621603120001_') OR ColumnPrefixFilter('1546841752667120001_') OR ColumnPrefixFilter('1545387980301120001_') OR ColumnPrefixFilter('1544677866436120001_') OR ColumnPrefixFilter('1546252381017120001_') OR ColumnPrefixFilter('1546247122010120001_') OR ColumnPrefixFilter('2866449000003420001_w') OR ColumnPrefixFilter('1545221612425120001_') OR ColumnPrefixFilter('1546582197395120001_') OR ColumnPrefixFilter('1545798618753120001_')\"} \r\n\r\n ","from":"developer"},{"body":"I guess the bottleneck is here, Let me have a test..\r\n{code}\r\ndiff --git a/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java b/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\r\nindex d51fdf0701..f0fbf3e8a1 100644\r\n--- a/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\r\n+++ b/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\r\n@@ -684,7 +684,7 @@ public class StoreScanner extends NonReversedNonLazyKeyValueScanner\r\n \r\n case SEEK_NEXT_USING_HINT:\r\n Cell nextKV = matcher.getNextKeyHint(cell);\r\n- if (nextKV != null) {\r\n+ if (nextKV != null && comparator.compare(nextKV, cell) > 0) {\r\n seekAsDirection(nextKV);\r\n NextState stateAfterSeekByHint = needToReturn(outResult);\r\n if (stateAfterSeekByHint != null) {\r\n{code}","from":"developer"},{"body":"[~openinx] any update on this?","from":"developer"},{"body":"Not yet, you can try to have a simple test after remove the && comparator.compare(nextKV, cell) > 0 , and compare the performance. ","from":"developer"},{"body":"[~openinx] we removed the *(comparator.compare(nextKV, cell) > 0)* condition and tested, but still the scan is slow.\r\n\r\n ","from":"developer"},{"body":"Any improvement ? another reason maybe: the filter list concated by OR will choose the minimal forward step among all sub-filters. in this patch, we have stricter restrictions on all sub filters, include those sub-filter whose has non-null RC returned in calculateReturnCodeByPrevCellAndRC (previously, we will skip to merge this sub-filter's rc, but it's wrong in some case), and merge all of the sub-filter's RC, this is also some time cost.","from":"developer"},{"body":"[~openinx] There was no improvement even after removing the compare condition. Is there any other way to optimize the scan?","from":"developer"},{"body":"We tried the same scan using HBase-2.1.2. The scan is even more slower than 1.4.8 (It takes around 3-4 seconds now). Because of this issue we are unable to update our production clusters from 1.2.5 to a higher version. [~openinx] Can you please provide a fix for this?\r\n\r\n ","from":"developer"},{"body":"[~KarthickRam] , How can I import your columnkey.txt to my hbase cluster to reproduce your case ? ","from":"developer"},{"body":"[~openinx] please use this (HBaseFileImport.java) file to import the columns in your cluster. Please replace the ZK Quorum, topnode and namespace:tablename. ","from":"developer"},{"body":"Thanks [~KarthickRam], I found some points to optimize for your case, and filed an issue HBASE-21734 for this. let's discuss there. ","from":"developer"},{"body":"[~openinx] While scanning a row (around 10 lakhs columns) with  100 column prefixes , it takes around  4 secs  in hbase-1.2.5 and when same query is executed in hbase-1.4.9 it takes around 50 secs .\r\n\r\nAnyway to optimise this??\r\n\r\nP.S we have applied the patch provided in HBASE-21620 and  HBASE-21734 . Attached *qualifiers*.*txt* file , use the HBaseFileImport.java file provided to populate in your table and use *scanquery.txt* to query ","from":"developer"},{"body":"How could I read your input files and verify the problem ? ","from":"developer"},{"body":"[~openinx]  please use the attached *HBaseFileImport.java* file and change the filename from _\"columnkey.txt\"_ to _\"qualifiers.txt\"_ to import the columns in your table. You also have to replace the ZK quorum, topnode and namespace:tablename. After importing the columns you can use scan query given in *scanquery.txt* in hbase shell.","from":"developer"},{"body":"[~openinx] Any update on this ?","from":"developer"},{"body":"[~openinx] We are facing this issue in production. Any update on this?","from":"developer"},{"body":"OK. Mind to open an seperate issue for this ? I will take some time later, still some working on hand :-) ","from":"developer"},{"body":"[~openinx] I've raised a separate issue here [https://issues.apache.org/jira/browse/HBASE-22448].\r\n\r\n ","from":"developer"}],"created":"2018-12-19T16:05:47.000+0000","description":"In some cases, unable to get the scan results when using more than one column prefix filter.\r\n\r\nAttached a java file to import the data which we used and a text file containing the values..\r\n\r\nWhile executing the following query (hbase shell as well as java program) it is waiting indefinitely and after RPC timeout we got the following error.. Also we noticed high cpu, high load average and very frequent young gc  in the region server containing this row...\r\n\r\nscan 'namespace:tablename',\\{STARTROW => 'test',ENDROW => 'test', FILTER => \"ColumnPrefixFilter('1544770422942010001_') OR ColumnPrefixFilter('1544769883529010001_')\"}\r\n\r\nROW                                                  COLUMN+CELL                                                                   ERROR: Call id=18, waitTime=60005, rpcTimetout=60000\r\n\r\n \r\n\r\nNote: Table scan operation and scan with a single column prefix filter works fine in this case.\r\n\r\nWhen we check the same query in hbase-1.2.5 it is working fine.\r\n\r\nCan you please help me on this..","issue_id":"13205431","key":"HBASE-21620","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2018-12-21T23:24:27.000+0000","role":"fixed_distractor","summary":"Problem in scan query when using more than one column prefix filter in some cases."} {"case_id":"12455120","cluster":"DISTRACTOR-HBASE-2180","comments":[{"body":"Using pread -- its present in the code already, its just commented out on the line following the above cited by Ryan -- I see doubled throughput when 16 clients concurrently random reading out of a single regionserver so it helps. i'll try and get some more numbers in here (I see 'wa' in top running at about the same for both cases but the regionserver is definetly working harder for the pread case using about double the CPU).\n\nNumbers are not that good though -- about 50ms latency doing a random read when 16 concurrent clients. This is a RS carrying 16M rows on 92 regions where there is 1 storefile only in the family and 4DNs under it.\n\nWay back when we were looking at pread, it improved the random read latency by some small percentage IIRC, about 11%, but then scan speed slowed some... but these would have been for the case of low numbers of concurrent clients.\n\nIts scanning 27k rows/second before the pread change using single client. And 21k/second after.\n\nLet me get some more numbers... up the concurrent client count and get some other points on how pread changes throughput.\n","created":"2010-02-03T23:58:33.150+0000"},{"body":"Link to issue that suggests we pread when doing random read and read when scanning.","created":"2010-02-04T02:31:22.411+0000"},{"body":"We ran some testing that show improved performance by commenting in the change in BoundedFileInputStream, by replacing the synchronized statement with the one commented out, that uses the PositionalReader interface\n\n{noformat}\n //synchronized (in) {\n // in.seek(pos);\n // ret = in.read(b, off, n);\n //}\n ret = in.read(pos, b, off, n);\n{noformat}","created":"2010-02-04T21:40:38.363+0000"},{"body":"This patch has gets do preads fetching blocks and uses the old seek+read for scans.\n\nPatch removes the old HFile.Reader.getScanner methods and replaces both with a getScanner that takes two arguments -- whether to cache blocks read and whether to use pread or not pulling in the block. I got rid of the old getScanners to force all getScanners to be explicit about what they want regards caching and pread.\n\nThis patch does not include tests. Its hard to test for this performance change.\n\nA further improvement would recognize short scans -- i.e. scans that are < an hfile block size. In this case, we'd want to pread rather than seek+scan (especially so when scan one row replaces get)\n\n","created":"2010-02-05T00:29:49.358+0000"},{"body":"This patch includes fixes for tests making them use new getScanner method and includes small PE fix when --rows is small (We would NPE). I might need a v3. A test is failing (TestGetDeleteTracker). Need to investigate.\n\nIn testing on something that tries to resemble the yahoo papers testing -- ~20M rows per server, 116 regions on a RS and only one replica -- this patch seems to double the throughput if ~20 concurrent clients on a RS. I tested scans and scan speeds are what they were w/ this patch in place. They have not deterioated.\n\nOne thing I noticed was that scanning when the data is not local -- i.e. the data is in a DN on another machine -- there is added latency for sure.... taking maybe 25% as long again for the test to complete. I need to see if same is true of random reads. Cosmin suggested that the yahoo test with its single replica only might be doing lots of remote accessing and could be incurring the extra latency.","created":"2010-02-05T07:42:33.795+0000"},{"body":"+1 thanks for doing this!","created":"2010-02-06T22:59:29.262+0000"},{"body":"Committed branch and trunk.","created":"2010-02-07T05:29:52.524+0000"},{"body":"Really commit to TRUNK.","created":"2010-02-08T23:40:12.648+0000"},{"body":"Committed a while back. Resolving.","created":"2010-02-19T00:18:32.975+0000"},{"body":"After applying this patch to 0.20.3 I got the following errors in my regionserver logs when doing high loads of gets and puts:\n\n2010-02-25 11:44:08,243 INFO org.apache.hadoop.hbase.regionserver.HRegion: compaction completed on region inrdb_ticket,\\x07R\\x00\\x00\\x00\\x00\\x80\\xFF\\xFF\\xFF\\x7F\\x00\\x00\\x00\\x01,1267094341820 i\nn 6sec\n1177:java.net.BindException: Cannot assign requested address\n at sun.nio.ch.Net.connect(Native Method)\n at sun.nio.ch.SocketChannelImpl.connect(SocketChannelImpl.java:507)\n at org.apache.hadoop.net.SocketIOWithTimeout.connect(SocketIOWithTimeout.java:192)\n at org.apache.hadoop.net.NetUtils.connect(NetUtils.java:404)\n at org.apache.hadoop.hdfs.DFSClient$DFSInputStream.fetchBlockByteRange(DFSClient.java:1825)\n at org.apache.hadoop.hdfs.DFSClient$DFSInputStream.read(DFSClient.java:1898)\n at org.apache.hadoop.fs.FSDataInputStream.read(FSDataInputStream.java:46)\n at org.apache.hadoop.hbase.io.hfile.BoundedRangeFileInputStream.read(BoundedRangeFileInputStream.java:101)\n at org.apache.hadoop.hbase.io.hfile.BoundedRangeFileInputStream.read(BoundedRangeFileInputStream.java:88)\n at org.apache.hadoop.hbase.io.hfile.BoundedRangeFileInputStream.read(BoundedRangeFileInputStream.java:81)\n at org.apache.hadoop.io.compress.BlockDecompressorStream.rawReadInt(BlockDecompressorStream.java:121)\n at org.apache.hadoop.io.compress.BlockDecompressorStream.getCompressedData(BlockDecompressorStream.java:96)\n at org.apache.hadoop.io.compress.BlockDecompressorStream.decompress(BlockDecompressorStream.java:82)\n at org.apache.hadoop.io.compress.DecompressorStream.read(DecompressorStream.java:74)\n at java.io.BufferedInputStream.read1(BufferedInputStream.java:256)\n at java.io.BufferedInputStream.read(BufferedInputStream.java:317)\n at org.apache.hadoop.io.IOUtils.readFully(IOUtils.java:100)\n at org.apache.hadoop.hbase.io.hfile.HFile$Reader.decompress(HFile.java:1018)\n at org.apache.hadoop.hbase.io.hfile.HFile$Reader.readBlock(HFile.java:966)\n at org.apache.hadoop.hbase.io.hfile.HFile$Reader$Scanner.next(HFile.java:1159)\n at org.apache.hadoop.hbase.regionserver.StoreFileGetScan.getStoreFile(StoreFileGetScan.java:108)\n at org.apache.hadoop.hbase.regionserver.StoreFileGetScan.get(StoreFileGetScan.java:65)\n at org.apache.hadoop.hbase.regionserver.Store.get(Store.java:1463)\n at org.apache.hadoop.hbase.regionserver.HRegion.get(HRegion.java:2396)\n at org.apache.hadoop.hbase.regionserver.HRegion.get(HRegion.java:2385)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.get(HRegionServer.java:1731)\n at sun.reflect.GeneratedMethodAccessor7.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:657)\n at org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:915)\n\nThe DataNode logs are fine (no maximum xcievers exceeded errors). Turns out that the OS was running out of port numbers. netstat showed more than 20,000 connections in TIME_WAIT state. Reverting to the original hbase-0.20.3 jar solved the problem. Only very few (<10) TIME_WAIT connections even after running gets/puts for a while.\n\nSo it looks like this patch causes some network connection issues. Any ideas if that could be the case?\n\nPS Running only gets seems to be fine, but I've mostly run tests with reads from the block cache.","created":"2010-02-25T11:11:07.310+0000"},{"body":"Reopening to take a look.\n\nI have a vague recollection of stuff not being closed down if not all is read out of the socket. Thanks for reporting this Erik.","created":"2010-02-25T12:03:55.067+0000"},{"body":"I was thinking of a very old issue, HADOOP-2341, but that was about CLOSE_WAIT, not TIME_WAIT. Erik I presume the TIME_WAIT are on the datanode side? I suppose there could be an issue here if many random reads in a short amount of time and the minimum segment lifetime (MSL) time is long in your tcp/ip implementation. Do you know what it is? 2minutes seems default reading up on the internets so could be in TIME_WAIT for 4 minutes. This what you are seeing you think Erik? They go away after a while? Whats the OS? This would seem to be a new issue then. We need pread that does keep-alive reusing sockets (Todd!).","created":"2010-02-25T16:22:22.242+0000"},{"body":"I saw both CLOSE_WAIT and TIME_WAIT. Maybe CLOSE_WAIT was in the majority. Connections were mostly to the data node.\n\n$ uname -a\nLinux inrdb-worker1.ripe.net 2.6.18-164.11.1.el5 #1 SMP Wed Jan 20 07:32:21 EST 2010 x86_64 x86_64 x86_64 GNU/Linux\n\nThey do go after a while, since after a few of the \"Cannot assign requested address\" exceptions the server starts working again.\n\nUnfortunately I'll be away for the weekend and won't be able to investigate further. I wonder why so many connections are being opened so quickly that the server runs out of ports within a few minutes of starting the gets/puts?","created":"2010-02-25T16:32:56.552+0000"},{"body":".bq I wonder why so many connections are being opened so quickly that the server runs out of ports within a few minutes of starting the gets/puts?\n\nGets used hdfs pread. pread opens a socket per access. My guess is that high rate of gets soon overwhelms the time each socket takes to clean up after close. What kinda rates are we talking here Erik?","created":"2010-02-27T13:08:26.865+0000"},{"body":"In the absence of reusing sockets, I think the TIME_WAIT issue could be dealt with on the system level by toggling /proc/sys/net/ipv4/tcp_tw_recycle","created":"2010-03-03T21:23:30.242+0000"},{"body":"Resolving against 0.20.4. I opened hbase-2492 to cover underlying new socket per pread.","created":"2010-04-26T16:17:16.325+0000"}],"conversations":[{"body":"deep in the HFile read path, there is this code:\n\n synchronized (in) {\n in.seek(pos);\n ret = in.read(b, off, n);\n }\n\n\nthis makes it so that only 1 read per file per thread is active. this prevents the OS and hardware from being able to do IO scheduling by optimizing lots of concurrent reads. \n\nWe need to either use a reentrant API (pread may be partially reentrant according to Todd) or use multiple stream objects, 1 per scanner/thread.","from":"reporter","subject":"Bad random read performance from synchronizing hfile.fddatainputstream"},{"body":"Using pread -- its present in the code already, its just commented out on the line following the above cited by Ryan -- I see doubled throughput when 16 clients concurrently random reading out of a single regionserver so it helps. i'll try and get some more numbers in here (I see 'wa' in top running at about the same for both cases but the regionserver is definetly working harder for the pread case using about double the CPU).\n\nNumbers are not that good though -- about 50ms latency doing a random read when 16 concurrent clients. This is a RS carrying 16M rows on 92 regions where there is 1 storefile only in the family and 4DNs under it.\n\nWay back when we were looking at pread, it improved the random read latency by some small percentage IIRC, about 11%, but then scan speed slowed some... but these would have been for the case of low numbers of concurrent clients.\n\nIts scanning 27k rows/second before the pread change using single client. And 21k/second after.\n\nLet me get some more numbers... up the concurrent client count and get some other points on how pread changes throughput.\n","from":"developer"},{"body":"Link to issue that suggests we pread when doing random read and read when scanning.","from":"developer"},{"body":"We ran some testing that show improved performance by commenting in the change in BoundedFileInputStream, by replacing the synchronized statement with the one commented out, that uses the PositionalReader interface\n\n{noformat}\n //synchronized (in) {\n // in.seek(pos);\n // ret = in.read(b, off, n);\n //}\n ret = in.read(pos, b, off, n);\n{noformat}","from":"developer"},{"body":"This patch has gets do preads fetching blocks and uses the old seek+read for scans.\n\nPatch removes the old HFile.Reader.getScanner methods and replaces both with a getScanner that takes two arguments -- whether to cache blocks read and whether to use pread or not pulling in the block. I got rid of the old getScanners to force all getScanners to be explicit about what they want regards caching and pread.\n\nThis patch does not include tests. Its hard to test for this performance change.\n\nA further improvement would recognize short scans -- i.e. scans that are < an hfile block size. In this case, we'd want to pread rather than seek+scan (especially so when scan one row replaces get)\n\n","from":"developer"},{"body":"This patch includes fixes for tests making them use new getScanner method and includes small PE fix when --rows is small (We would NPE). I might need a v3. A test is failing (TestGetDeleteTracker). Need to investigate.\n\nIn testing on something that tries to resemble the yahoo papers testing -- ~20M rows per server, 116 regions on a RS and only one replica -- this patch seems to double the throughput if ~20 concurrent clients on a RS. I tested scans and scan speeds are what they were w/ this patch in place. They have not deterioated.\n\nOne thing I noticed was that scanning when the data is not local -- i.e. the data is in a DN on another machine -- there is added latency for sure.... taking maybe 25% as long again for the test to complete. I need to see if same is true of random reads. Cosmin suggested that the yahoo test with its single replica only might be doing lots of remote accessing and could be incurring the extra latency.","from":"developer"},{"body":"+1 thanks for doing this!","from":"developer"},{"body":"Committed branch and trunk.","from":"developer"},{"body":"Really commit to TRUNK.","from":"developer"},{"body":"Committed a while back. Resolving.","from":"developer"},{"body":"After applying this patch to 0.20.3 I got the following errors in my regionserver logs when doing high loads of gets and puts:\n\n2010-02-25 11:44:08,243 INFO org.apache.hadoop.hbase.regionserver.HRegion: compaction completed on region inrdb_ticket,\\x07R\\x00\\x00\\x00\\x00\\x80\\xFF\\xFF\\xFF\\x7F\\x00\\x00\\x00\\x01,1267094341820 i\nn 6sec\n1177:java.net.BindException: Cannot assign requested address\n at sun.nio.ch.Net.connect(Native Method)\n at sun.nio.ch.SocketChannelImpl.connect(SocketChannelImpl.java:507)\n at org.apache.hadoop.net.SocketIOWithTimeout.connect(SocketIOWithTimeout.java:192)\n at org.apache.hadoop.net.NetUtils.connect(NetUtils.java:404)\n at org.apache.hadoop.hdfs.DFSClient$DFSInputStream.fetchBlockByteRange(DFSClient.java:1825)\n at org.apache.hadoop.hdfs.DFSClient$DFSInputStream.read(DFSClient.java:1898)\n at org.apache.hadoop.fs.FSDataInputStream.read(FSDataInputStream.java:46)\n at org.apache.hadoop.hbase.io.hfile.BoundedRangeFileInputStream.read(BoundedRangeFileInputStream.java:101)\n at org.apache.hadoop.hbase.io.hfile.BoundedRangeFileInputStream.read(BoundedRangeFileInputStream.java:88)\n at org.apache.hadoop.hbase.io.hfile.BoundedRangeFileInputStream.read(BoundedRangeFileInputStream.java:81)\n at org.apache.hadoop.io.compress.BlockDecompressorStream.rawReadInt(BlockDecompressorStream.java:121)\n at org.apache.hadoop.io.compress.BlockDecompressorStream.getCompressedData(BlockDecompressorStream.java:96)\n at org.apache.hadoop.io.compress.BlockDecompressorStream.decompress(BlockDecompressorStream.java:82)\n at org.apache.hadoop.io.compress.DecompressorStream.read(DecompressorStream.java:74)\n at java.io.BufferedInputStream.read1(BufferedInputStream.java:256)\n at java.io.BufferedInputStream.read(BufferedInputStream.java:317)\n at org.apache.hadoop.io.IOUtils.readFully(IOUtils.java:100)\n at org.apache.hadoop.hbase.io.hfile.HFile$Reader.decompress(HFile.java:1018)\n at org.apache.hadoop.hbase.io.hfile.HFile$Reader.readBlock(HFile.java:966)\n at org.apache.hadoop.hbase.io.hfile.HFile$Reader$Scanner.next(HFile.java:1159)\n at org.apache.hadoop.hbase.regionserver.StoreFileGetScan.getStoreFile(StoreFileGetScan.java:108)\n at org.apache.hadoop.hbase.regionserver.StoreFileGetScan.get(StoreFileGetScan.java:65)\n at org.apache.hadoop.hbase.regionserver.Store.get(Store.java:1463)\n at org.apache.hadoop.hbase.regionserver.HRegion.get(HRegion.java:2396)\n at org.apache.hadoop.hbase.regionserver.HRegion.get(HRegion.java:2385)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.get(HRegionServer.java:1731)\n at sun.reflect.GeneratedMethodAccessor7.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:657)\n at org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:915)\n\nThe DataNode logs are fine (no maximum xcievers exceeded errors). Turns out that the OS was running out of port numbers. netstat showed more than 20,000 connections in TIME_WAIT state. Reverting to the original hbase-0.20.3 jar solved the problem. Only very few (<10) TIME_WAIT connections even after running gets/puts for a while.\n\nSo it looks like this patch causes some network connection issues. Any ideas if that could be the case?\n\nPS Running only gets seems to be fine, but I've mostly run tests with reads from the block cache.","from":"developer"},{"body":"Reopening to take a look.\n\nI have a vague recollection of stuff not being closed down if not all is read out of the socket. Thanks for reporting this Erik.","from":"developer"},{"body":"I was thinking of a very old issue, HADOOP-2341, but that was about CLOSE_WAIT, not TIME_WAIT. Erik I presume the TIME_WAIT are on the datanode side? I suppose there could be an issue here if many random reads in a short amount of time and the minimum segment lifetime (MSL) time is long in your tcp/ip implementation. Do you know what it is? 2minutes seems default reading up on the internets so could be in TIME_WAIT for 4 minutes. This what you are seeing you think Erik? They go away after a while? Whats the OS? This would seem to be a new issue then. We need pread that does keep-alive reusing sockets (Todd!).","from":"developer"},{"body":"I saw both CLOSE_WAIT and TIME_WAIT. Maybe CLOSE_WAIT was in the majority. Connections were mostly to the data node.\n\n$ uname -a\nLinux inrdb-worker1.ripe.net 2.6.18-164.11.1.el5 #1 SMP Wed Jan 20 07:32:21 EST 2010 x86_64 x86_64 x86_64 GNU/Linux\n\nThey do go after a while, since after a few of the \"Cannot assign requested address\" exceptions the server starts working again.\n\nUnfortunately I'll be away for the weekend and won't be able to investigate further. I wonder why so many connections are being opened so quickly that the server runs out of ports within a few minutes of starting the gets/puts?","from":"developer"},{"body":".bq I wonder why so many connections are being opened so quickly that the server runs out of ports within a few minutes of starting the gets/puts?\n\nGets used hdfs pread. pread opens a socket per access. My guess is that high rate of gets soon overwhelms the time each socket takes to clean up after close. What kinda rates are we talking here Erik?","from":"developer"},{"body":"In the absence of reusing sockets, I think the TIME_WAIT issue could be dealt with on the system level by toggling /proc/sys/net/ipv4/tcp_tw_recycle","from":"developer"},{"body":"Resolving against 0.20.4. I opened hbase-2492 to cover underlying new socket per pread.","from":"developer"}],"created":"2010-02-02T22:11:31.000+0000","description":"deep in the HFile read path, there is this code:\n\n synchronized (in) {\n in.seek(pos);\n ret = in.read(b, off, n);\n }\n\n\nthis makes it so that only 1 read per file per thread is active. this prevents the OS and hardware from being able to do IO scheduling by optimizing lots of concurrent reads. \n\nWe need to either use a reentrant API (pread may be partially reentrant according to Todd) or use multiple stream objects, 1 per scanner/thread.","issue_id":"12455120","key":"HBASE-2180","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2010-04-26T16:17:16.000+0000","role":"fixed_distractor","summary":"Bad random read performance from synchronizing hfile.fddatainputstream"} {"case_id":"13231190","cluster":"DISTRACTOR-HBASE-22349","comments":[{"body":"This is a very good observation. One of my co-worker observed and debugged the similar issue in our environment.\r\n\r\nIf we don't want RS holds 0 regions, maybe besides tweaking 'minCostNeedBalance', we can introduce a rule that when RS holds 0 region, it triggers balancing regardless. \r\n\r\n Or, we can adjust cost() for this class :\r\n\r\nstatic class PrimaryRegionCountSkewCostFunction\r\n\r\nto make this factor impacting more than others?","created":"2019-06-28T22:55:00.163+0000"},{"body":"The scenario as originally described is fixed by HBASE-24139. However, I would like to propose using this to track other cases where we should execute the balancer, like one server with much fewer regions, or much more regions, than the average server in the cluster. (Take the original scenario, and instead of having 0 regions on the server, have only 1 region on the server.)\r\n\r\nI can think of a few options:\r\n# Use some hot/cold threshold like 50%. Compute the average regions per server. If a server has a region count which is >150% or <50% of this average, allow the balancer to run (short-circuit in {{needsBalance}})\r\n# Find outliers using some type of standard deviation, and short-circuit run in {{needsBalance}} if one is found.\r\n# Introduce a \"force run\" of the balancer on some timed interval.\r\n\r\nI'm inclined to try option 1. Option 3 sounds appealing to me, because it is a backstop to catch all of the cases which are ignored by {{minCostNeedBalance}}. However, other operators may find it too interrupting, if they need little to no region movement in the cluster.\r\n\r\nFor reference, one scenario where we find ourselves in this undesirable state is by running {{region_mover}} at the same time as the load balancer. As stated in the {{region_mover}} comments, those two operations will conflict. The result can be one regionserver which has double the regions of any other server in the cluster. And if {{minCostNeedBalance}} is not exceeded, which is not difficult in a sizable cluster, one regionserver will run with double the load indefinitely.","created":"2022-04-14T23:29:48.419+0000"}],"conversations":[{"body":"HBASE-24139 allows the load balancer to run when one server has 0 regions and another server has more than 1 region. This is a special case of a more generic problem, where one server has far too few or far too many regions. The StochasticLoadBalancer defaults may decide the cluster is \"balanced enough\" according to {{hbase.master.balancer.stochastic.minCostNeedBalance}}, even though one server may have a far higher or lower number of regions compared to the rest of the cluster.\r\n\r\nOne specific example of this we have seen is when we use {{RegionMover}} to move regions back to a restarted RegionServer, if the {{StochasticLoadBalancer}} happens to be running. The load balancer sees a newly restarted RegionServer with 0 regions, and after HBASE-24139, it will balance regions to this server. Simultaneously, {{RegionMover}} moves back regions. The end result is that the newly restarted RegionServer has twice the load of any other server in the cluster. Future iterations of the load balancer do nothing, as the cluster cost does not exceed {{minCostNeedBalance}}.\r\n\r\nAnother example is if the load balancer makes very slow progress on a cluster, it may not move the average cluster load to a newly restarted regionserver in one iteration. But after the first iteration, the balancer may again not run due to cluster cost not exceeding {{minCostNeedBalance}}.\r\n\r\nWe can propose a solution where we reuse the {{slop}} concept in {{SimpleLoadBalancer}} and use this to extend the HBASE-24139 logic for deciding to run the balancer as long as there is a \"sloppy\" server in the cluster.\r\n\r\n+*Previous Description Notes Below, which are relevant, but as stated, were already fixed by HBASE-24139*+\r\n\r\nIn EMR cluster, whenever I replace one of the nodes, the regions never get rebalanced.\r\n\r\nThe default minCostNeedBalance set to 0.05 is too high.\r\n\r\nThe region count on the servers were: 21, 21, 20, 20, 20, 20, 21, 20, 20, 20 = 203\r\n\r\nOnce a node(region server) got replaced with a new node (terminated and EMR recreated a node), the region count on the servers became: 23, 0, 23, 22, 22, 22, 22, 23, 23, 23 = 203\r\n\r\nFrom hbase-master-logs, I can see the below WARN which indicates that the default minCostNeedBalance does not hold good for these scenarios.\r\n\r\n##\r\n\r\n2019-04-29 09:31:37,027 WARN  [ip-172-31-35-122.ec2.internal,16000,1556524892897_ChoreService_1] cleaner.CleanerChore: WALs outstanding under hdfs://ip-172-31-35-122.ec2.internal:8020/user/hbase/oldWALs2019-04-29 09:31:42,920 INFO  [ip-172-31-35-122.ec2.internal,16000,1556524892897_ChoreService_1] balancer.StochasticLoadBalancer: Skipping load balancing because balanced cluster; total cost is 52.041826194833405, sum multiplier is 1102.0 min cost which need balance is 0.05\r\n\r\n##\r\n\r\nTo mitigate this, I had to modify the default minCostNeedBalance to lower value like 0.01f and restart Region Servers and Hbase Master. After modifying this value to 0.01f I could see the regions getting re-balanced.\r\n\r\nThis has led me to the following questions which I would like to get it answered from the HBase experts.\r\n\r\n1)What are the factors that affect the value of total cost and sum multiplier? How could we determine the right minCostNeedBalance value for any cluster?\r\n\r\n2)How did Hbase arrive at setting the default value to 0.05f? Is it optimal value? If yes, then what is the recommended way to mitigate this scenario? \r\n\r\nAttached: Steps to reproduce\r\n\r\n \r\n\r\nNote: HBase-17565 patch is already applied.","from":"reporter","subject":"Stochastic Load Balancer skips balancing when node is replaced in cluster"},{"body":"This is a very good observation. One of my co-worker observed and debugged the similar issue in our environment.\r\n\r\nIf we don't want RS holds 0 regions, maybe besides tweaking 'minCostNeedBalance', we can introduce a rule that when RS holds 0 region, it triggers balancing regardless. \r\n\r\n Or, we can adjust cost() for this class :\r\n\r\nstatic class PrimaryRegionCountSkewCostFunction\r\n\r\nto make this factor impacting more than others?","from":"developer"},{"body":"The scenario as originally described is fixed by HBASE-24139. However, I would like to propose using this to track other cases where we should execute the balancer, like one server with much fewer regions, or much more regions, than the average server in the cluster. (Take the original scenario, and instead of having 0 regions on the server, have only 1 region on the server.)\r\n\r\nI can think of a few options:\r\n# Use some hot/cold threshold like 50%. Compute the average regions per server. If a server has a region count which is >150% or <50% of this average, allow the balancer to run (short-circuit in {{needsBalance}})\r\n# Find outliers using some type of standard deviation, and short-circuit run in {{needsBalance}} if one is found.\r\n# Introduce a \"force run\" of the balancer on some timed interval.\r\n\r\nI'm inclined to try option 1. Option 3 sounds appealing to me, because it is a backstop to catch all of the cases which are ignored by {{minCostNeedBalance}}. However, other operators may find it too interrupting, if they need little to no region movement in the cluster.\r\n\r\nFor reference, one scenario where we find ourselves in this undesirable state is by running {{region_mover}} at the same time as the load balancer. As stated in the {{region_mover}} comments, those two operations will conflict. The result can be one regionserver which has double the regions of any other server in the cluster. And if {{minCostNeedBalance}} is not exceeded, which is not difficult in a sizable cluster, one regionserver will run with double the load indefinitely.","from":"developer"}],"created":"2019-05-02T05:34:27.000+0000","description":"HBASE-24139 allows the load balancer to run when one server has 0 regions and another server has more than 1 region. This is a special case of a more generic problem, where one server has far too few or far too many regions. The StochasticLoadBalancer defaults may decide the cluster is \"balanced enough\" according to {{hbase.master.balancer.stochastic.minCostNeedBalance}}, even though one server may have a far higher or lower number of regions compared to the rest of the cluster.\r\n\r\nOne specific example of this we have seen is when we use {{RegionMover}} to move regions back to a restarted RegionServer, if the {{StochasticLoadBalancer}} happens to be running. The load balancer sees a newly restarted RegionServer with 0 regions, and after HBASE-24139, it will balance regions to this server. Simultaneously, {{RegionMover}} moves back regions. The end result is that the newly restarted RegionServer has twice the load of any other server in the cluster. Future iterations of the load balancer do nothing, as the cluster cost does not exceed {{minCostNeedBalance}}.\r\n\r\nAnother example is if the load balancer makes very slow progress on a cluster, it may not move the average cluster load to a newly restarted regionserver in one iteration. But after the first iteration, the balancer may again not run due to cluster cost not exceeding {{minCostNeedBalance}}.\r\n\r\nWe can propose a solution where we reuse the {{slop}} concept in {{SimpleLoadBalancer}} and use this to extend the HBASE-24139 logic for deciding to run the balancer as long as there is a \"sloppy\" server in the cluster.\r\n\r\n+*Previous Description Notes Below, which are relevant, but as stated, were already fixed by HBASE-24139*+\r\n\r\nIn EMR cluster, whenever I replace one of the nodes, the regions never get rebalanced.\r\n\r\nThe default minCostNeedBalance set to 0.05 is too high.\r\n\r\nThe region count on the servers were: 21, 21, 20, 20, 20, 20, 21, 20, 20, 20 = 203\r\n\r\nOnce a node(region server) got replaced with a new node (terminated and EMR recreated a node), the region count on the servers became: 23, 0, 23, 22, 22, 22, 22, 23, 23, 23 = 203\r\n\r\nFrom hbase-master-logs, I can see the below WARN which indicates that the default minCostNeedBalance does not hold good for these scenarios.\r\n\r\n##\r\n\r\n2019-04-29 09:31:37,027 WARN  [ip-172-31-35-122.ec2.internal,16000,1556524892897_ChoreService_1] cleaner.CleanerChore: WALs outstanding under hdfs://ip-172-31-35-122.ec2.internal:8020/user/hbase/oldWALs2019-04-29 09:31:42,920 INFO  [ip-172-31-35-122.ec2.internal,16000,1556524892897_ChoreService_1] balancer.StochasticLoadBalancer: Skipping load balancing because balanced cluster; total cost is 52.041826194833405, sum multiplier is 1102.0 min cost which need balance is 0.05\r\n\r\n##\r\n\r\nTo mitigate this, I had to modify the default minCostNeedBalance to lower value like 0.01f and restart Region Servers and Hbase Master. After modifying this value to 0.01f I could see the regions getting re-balanced.\r\n\r\nThis has led me to the following questions which I would like to get it answered from the HBase experts.\r\n\r\n1)What are the factors that affect the value of total cost and sum multiplier? How could we determine the right minCostNeedBalance value for any cluster?\r\n\r\n2)How did Hbase arrive at setting the default value to 0.05f? Is it optimal value? If yes, then what is the recommended way to mitigate this scenario? \r\n\r\nAttached: Steps to reproduce\r\n\r\n \r\n\r\nNote: HBase-17565 patch is already applied.","issue_id":"13231190","key":"HBASE-22349","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2022-04-28T21:56:04.000+0000","role":"fixed_distractor","summary":"Stochastic Load Balancer skips balancing when node is replaced in cluster"} {"case_id":"13232694","cluster":"DISTRACTOR-HBASE-22393","comments":[{"body":"+1, lgtm.","created":"2019-05-10T17:35:15.983+0000"},{"body":"* Why aren't you creating a dependency reduced pom? As is this will include curator classes but still list curator as a dependency\r\n* Why not shade all dependencies? It'll make it easier to use hboss with a variety of hbase versions.\r\n* for the shading pattern please use something that won't conflict with either of the hbase-thirdparty artifacts or the hbase-shaded stuff.","created":"2019-05-10T17:41:20.791+0000"},{"body":"bq. or the shading pattern please use something that won't conflict with either of the hbase-thirdparty artifacts or the hbase-shaded stuff.\r\n\r\nsorry, left off: but also please include a package name that makes it clear these are relocated. e.g. {{org.apache.hadoop.hbase.oss.thirdparty}}","created":"2019-05-10T17:43:03.243+0000"},{"body":"{quote}Why aren't you creating a dependency reduced pom?{quote}\r\n\r\nThat's ... a very good question. Now shading all the dependencies except for things in test & provided scope with the prefix you suggested. Instead of re-shading all of hbase-thirdparty, the only thing I'm using is Guava's VisibleForTesting so I just shade that directly and dropped the old dependency.","created":"2019-05-10T19:47:26.642+0000"},{"body":"Tested latest patch. It's shading *org.apache.log4j*, but it's missing *log4j* in the _includes_ section, so am getting class not found while trying to start RSes:\r\n{noformat}\r\n2019-05-14 03:47:13,301 ERROR org.apache.hadoop.hbase.regionserver.HRegionServer: Failed construction RegionServer\r\njava.lang.NoClassDefFoundError: org/apache/hadoop/hbase/oss/thirdparty/org/apache/log4j/Level\r\n at org.apache.hadoop.hbase.oss.thirdparty.org.slf4j.LoggerFactory.bind(LoggerFactory.java:150)\r\n at org.apache.hadoop.hbase.oss.thirdparty.org.slf4j.LoggerFactory.performInitialization(LoggerFactory.java:124)\r\n at org.apache.hadoop.hbase.oss.thirdparty.org.slf4j.LoggerFactory.getILoggerFactory(LoggerFactory.java:412)\r\n at org.apache.hadoop.hbase.oss.thirdparty.org.slf4j.LoggerFactory.getLogger(LoggerFactory.java:357)\r\n at org.apache.hadoop.hbase.oss.thirdparty.org.slf4j.LoggerFactory.getLogger(LoggerFactory.java:383)\r\n at org.apache.hadoop.hbase.oss.HBaseObjectStoreSemantics.(HBaseObjectStoreSemantics.java:98)\r\n{noformat}\r\nAlso faced similar issues with *javax* related classes. I don't think we really need to shade it, since it's platform provided, we wouldn't face version conflicts here, right?\r\n\r\nI had managed to have it working by adding *log4j* in the _includes_ section and removing _relocation_ definition for *javax*.","created":"2019-05-14T11:18:23.035+0000"},{"body":"yep, no shading logging libraries or things that don't play well with shading. here's the set we use in the main project:\r\n\r\nhttps://github.com/apache/hbase/blob/68f14c19ff79e36b17e99c7e848c19ce5e0164d5/hbase-shaded/pom.xml#L134\r\n","created":"2019-05-14T14:56:14.797+0000"},{"body":"The javax stuff turned out to be from jsr305 not being excluded everywhere - it's actually a banned dependency in HBase proper.\r\n\r\nI've removed the slf4j-log4j and log4j JARs from any involvement in the shading after discussing the specifics with [~busbey]. slf4j-api is still included.\r\n\r\nLooks good to me, but I can't conveniently deploy this onto a real cluster, right now. I trust [~wchevreuil] can kindly give this a good test much quicker than I can :)","created":"2019-05-14T23:34:20.086+0000"},{"body":"Latest patch is relocating references to any class in *javax* domain, so it causes *NoClassDefFoundError* on process loading the generated hboss jar:\r\n\r\n{noformat}\r\nException in thread \"main\" java.lang.NoClassDefFoundError: org/apache/hadoop/hbase/oss/thirdparty/javax/security/sasl/SaslException\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.zookeeper.ClientCnxn.(ClientCnxn.java:400)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.zookeeper.ClientCnxn.(ClientCnxn.java:359)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.zookeeper.ZooKeeper.(ZooKeeper.java:447)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.utils.DefaultZookeeperFactory.newZooKeeper(DefaultZookeeperFactory.java:29)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.framework.imps.CuratorFrameworkImpl$2.newZooKeeper(CuratorFrameworkImpl.java:191)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.HandleHolder$1.getZooKeeper(HandleHolder.java:101)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.HandleHolder.getZooKeeper(HandleHolder.java:57)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.ConnectionState.reset(ConnectionState.java:201)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.ConnectionState.start(ConnectionState.java:111)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.CuratorZookeeperClient.start(CuratorZookeeperClient.java:214)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.framework.imps.CuratorFrameworkImpl.start(CuratorFrameworkImpl.java:308)\r\n\tat org.apache.hadoop.hbase.oss.sync.ZKTreeLockManager.initialize(ZKTreeLockManager.java:90)\r\n{noformat}\r\n\r\nSince the offending dependency (jsr305) is already marked for exclusion, I believe we can safely remove this *javax relocation rule* from shade section in the pom. Attaching a patch example that worked on my tests (notice I don't have *javax* on the relocations portion).","created":"2019-05-16T19:01:19.913+0000"},{"body":"[~busbey], any comments on the latest patch and previous comment above? If that looks ok now, can we get it committed?","created":"2019-05-21T12:13:50.948+0000"},{"body":"[^0001-HBASE-22393.patch] looks OK to me, but would probably be good to wait for Busbey to come back :)","created":"2019-05-28T17:47:20.382+0000"},{"body":"I'm pulling down the latest and reviewing today/tomorrow.","created":"2019-05-28T20:37:21.738+0000"},{"body":"Updated patch on [PR#2|https://github.com/apache/hbase-filesystem/pull/2]. [~mackrorysd] and/or [~wchevreuil] can you take a look. I think this one is ready to go.","created":"2019-05-30T07:07:03.978+0000"},{"body":"lgtm.","created":"2019-05-31T11:39:45.668+0000"},{"body":"thanks folks!","created":"2019-05-31T15:44:01.478+0000"},{"body":"Belated +1 from me too, although a couple of the other maven features that have come into play here are a bit foreign to me still.","created":"2019-05-31T17:31:05.947+0000"}],"conversations":[{"body":"Hadoop uses a very old version of Curator, and if it ever ends up in the classpath with HBOSS you can get this:\r\n\r\n{code}Exception in thread \"main\" java.lang.IllegalAccessError: tried to access method org.apache.curator.framework.recipes.locks.InterProcessMutex.isOwnedByCurrentThread()Z from class org.apache.hadoop.hbase.oss.sync.ZKTreeLockManager{code}\r\n\r\nI think the simplest solution is to just shade Curator.","from":"reporter","subject":"HBOSS: Shaded external dependencies to avoid conflicts with Hadoop and HBase"},{"body":"+1, lgtm.","from":"developer"},{"body":"* Why aren't you creating a dependency reduced pom? As is this will include curator classes but still list curator as a dependency\r\n* Why not shade all dependencies? It'll make it easier to use hboss with a variety of hbase versions.\r\n* for the shading pattern please use something that won't conflict with either of the hbase-thirdparty artifacts or the hbase-shaded stuff.","from":"developer"},{"body":"bq. or the shading pattern please use something that won't conflict with either of the hbase-thirdparty artifacts or the hbase-shaded stuff.\r\n\r\nsorry, left off: but also please include a package name that makes it clear these are relocated. e.g. {{org.apache.hadoop.hbase.oss.thirdparty}}","from":"developer"},{"body":"{quote}Why aren't you creating a dependency reduced pom?{quote}\r\n\r\nThat's ... a very good question. Now shading all the dependencies except for things in test & provided scope with the prefix you suggested. Instead of re-shading all of hbase-thirdparty, the only thing I'm using is Guava's VisibleForTesting so I just shade that directly and dropped the old dependency.","from":"developer"},{"body":"Tested latest patch. It's shading *org.apache.log4j*, but it's missing *log4j* in the _includes_ section, so am getting class not found while trying to start RSes:\r\n{noformat}\r\n2019-05-14 03:47:13,301 ERROR org.apache.hadoop.hbase.regionserver.HRegionServer: Failed construction RegionServer\r\njava.lang.NoClassDefFoundError: org/apache/hadoop/hbase/oss/thirdparty/org/apache/log4j/Level\r\n at org.apache.hadoop.hbase.oss.thirdparty.org.slf4j.LoggerFactory.bind(LoggerFactory.java:150)\r\n at org.apache.hadoop.hbase.oss.thirdparty.org.slf4j.LoggerFactory.performInitialization(LoggerFactory.java:124)\r\n at org.apache.hadoop.hbase.oss.thirdparty.org.slf4j.LoggerFactory.getILoggerFactory(LoggerFactory.java:412)\r\n at org.apache.hadoop.hbase.oss.thirdparty.org.slf4j.LoggerFactory.getLogger(LoggerFactory.java:357)\r\n at org.apache.hadoop.hbase.oss.thirdparty.org.slf4j.LoggerFactory.getLogger(LoggerFactory.java:383)\r\n at org.apache.hadoop.hbase.oss.HBaseObjectStoreSemantics.(HBaseObjectStoreSemantics.java:98)\r\n{noformat}\r\nAlso faced similar issues with *javax* related classes. I don't think we really need to shade it, since it's platform provided, we wouldn't face version conflicts here, right?\r\n\r\nI had managed to have it working by adding *log4j* in the _includes_ section and removing _relocation_ definition for *javax*.","from":"developer"},{"body":"yep, no shading logging libraries or things that don't play well with shading. here's the set we use in the main project:\r\n\r\nhttps://github.com/apache/hbase/blob/68f14c19ff79e36b17e99c7e848c19ce5e0164d5/hbase-shaded/pom.xml#L134\r\n","from":"developer"},{"body":"The javax stuff turned out to be from jsr305 not being excluded everywhere - it's actually a banned dependency in HBase proper.\r\n\r\nI've removed the slf4j-log4j and log4j JARs from any involvement in the shading after discussing the specifics with [~busbey]. slf4j-api is still included.\r\n\r\nLooks good to me, but I can't conveniently deploy this onto a real cluster, right now. I trust [~wchevreuil] can kindly give this a good test much quicker than I can :)","from":"developer"},{"body":"Latest patch is relocating references to any class in *javax* domain, so it causes *NoClassDefFoundError* on process loading the generated hboss jar:\r\n\r\n{noformat}\r\nException in thread \"main\" java.lang.NoClassDefFoundError: org/apache/hadoop/hbase/oss/thirdparty/javax/security/sasl/SaslException\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.zookeeper.ClientCnxn.(ClientCnxn.java:400)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.zookeeper.ClientCnxn.(ClientCnxn.java:359)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.zookeeper.ZooKeeper.(ZooKeeper.java:447)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.utils.DefaultZookeeperFactory.newZooKeeper(DefaultZookeeperFactory.java:29)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.framework.imps.CuratorFrameworkImpl$2.newZooKeeper(CuratorFrameworkImpl.java:191)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.HandleHolder$1.getZooKeeper(HandleHolder.java:101)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.HandleHolder.getZooKeeper(HandleHolder.java:57)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.ConnectionState.reset(ConnectionState.java:201)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.ConnectionState.start(ConnectionState.java:111)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.CuratorZookeeperClient.start(CuratorZookeeperClient.java:214)\r\n\tat org.apache.hadoop.hbase.oss.thirdparty.org.apache.curator.framework.imps.CuratorFrameworkImpl.start(CuratorFrameworkImpl.java:308)\r\n\tat org.apache.hadoop.hbase.oss.sync.ZKTreeLockManager.initialize(ZKTreeLockManager.java:90)\r\n{noformat}\r\n\r\nSince the offending dependency (jsr305) is already marked for exclusion, I believe we can safely remove this *javax relocation rule* from shade section in the pom. Attaching a patch example that worked on my tests (notice I don't have *javax* on the relocations portion).","from":"developer"},{"body":"[~busbey], any comments on the latest patch and previous comment above? If that looks ok now, can we get it committed?","from":"developer"},{"body":"[^0001-HBASE-22393.patch] looks OK to me, but would probably be good to wait for Busbey to come back :)","from":"developer"},{"body":"I'm pulling down the latest and reviewing today/tomorrow.","from":"developer"},{"body":"Updated patch on [PR#2|https://github.com/apache/hbase-filesystem/pull/2]. [~mackrorysd] and/or [~wchevreuil] can you take a look. I think this one is ready to go.","from":"developer"},{"body":"lgtm.","from":"developer"},{"body":"thanks folks!","from":"developer"},{"body":"Belated +1 from me too, although a couple of the other maven features that have come into play here are a bit foreign to me still.","from":"developer"}],"created":"2019-05-10T15:53:00.000+0000","description":"Hadoop uses a very old version of Curator, and if it ever ends up in the classpath with HBOSS you can get this:\r\n\r\n{code}Exception in thread \"main\" java.lang.IllegalAccessError: tried to access method org.apache.curator.framework.recipes.locks.InterProcessMutex.isOwnedByCurrentThread()Z from class org.apache.hadoop.hbase.oss.sync.ZKTreeLockManager{code}\r\n\r\nI think the simplest solution is to just shade Curator.","issue_id":"13232694","key":"HBASE-22393","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2019-05-31T15:48:36.000+0000","role":"fixed_distractor","summary":"HBOSS: Shaded external dependencies to avoid conflicts with Hadoop and HBase"} {"case_id":"13236041","cluster":"DISTRACTOR-HBASE-22487","comments":[{"body":"This is a trivial patch which simply removes the uncalled code.","created":"2019-05-28T17:45:30.088+0000"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 3m 22s{color} | {color:blue} Docker mode activated. {color} |\r\n|| || || || {color:brown} Prechecks {color} ||\r\n| {color:green}+1{color} | {color:green} hbaseanti {color} | {color:green} 0m 0s{color} | {color:green} Patch does not have any anti-patterns. {color} |\r\n| {color:green}+1{color} | {color:green} @author {color} | {color:green} 0m 0s{color} | {color:green} The patch does not contain any @author tags. {color} |\r\n| {color:orange}-0{color} | {color:orange} test4tests {color} | {color:orange} 0m 0s{color} | {color:orange} The patch doesn't appear to include any new or modified tests. Please justify why no new tests are needed for this patch. Also please list what manual steps were performed to verify this patch. {color} |\r\n|| || || || {color:brown} master Compile Tests {color} ||\r\n| {color:green}+1{color} | {color:green} mvninstall {color} | {color:green} 5m 45s{color} | {color:green} master passed {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 1m 7s{color} | {color:green} master passed {color} |\r\n| {color:green}+1{color} | {color:green} checkstyle {color} | {color:green} 1m 30s{color} | {color:green} master passed {color} |\r\n| {color:green}+1{color} | {color:green} shadedjars {color} | {color:green} 5m 49s{color} | {color:green} branch has no errors when building our shaded downstream artifacts. {color} |\r\n| {color:green}+1{color} | {color:green} findbugs {color} | {color:green} 4m 24s{color} | {color:green} master passed {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 43s{color} | {color:green} master passed {color} |\r\n|| || || || {color:brown} Patch Compile Tests {color} ||\r\n| {color:green}+1{color} | {color:green} mvninstall {color} | {color:green} 5m 18s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 1m 5s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} javac {color} | {color:green} 1m 5s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} checkstyle {color} | {color:green} 1m 29s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} whitespace {color} | {color:green} 0m 0s{color} | {color:green} The patch has no whitespace issues. {color} |\r\n| {color:green}+1{color} | {color:green} shadedjars {color} | {color:green} 5m 48s{color} | {color:green} patch has no errors when building our shaded downstream artifacts. {color} |\r\n| {color:green}+1{color} | {color:green} hadoopcheck {color} | {color:green} 22m 28s{color} | {color:green} Patch does not cause any errors with Hadoop 2.8.5 2.9.2 or 3.0.3 3.1.2. {color} |\r\n| {color:green}+1{color} | {color:green} findbugs {color} | {color:green} 4m 42s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 44s{color} | {color:green} the patch passed {color} |\r\n|| || || || {color:brown} Other Tests {color} ||\r\n| {color:red}-1{color} | {color:red} unit {color} | {color:red}263m 46s{color} | {color:red} hbase-server in the patch failed. {color} |\r\n| {color:green}+1{color} | {color:green} asflicense {color} | {color:green} 1m 31s{color} | {color:green} The patch does not generate ASF License warnings. {color} |\r\n| {color:black}{color} | {color:black} {color} | {color:black}335m 29s{color} | {color:black} {color} |\r\n\\\\\r\n\\\\\r\n|| Reason || Tests ||\r\n| Failed junit tests | hadoop.hbase.replication.multiwal.TestReplicationSyncUpToolWithMultipleWAL |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| Docker | Client=17.05.0-ce Server=17.05.0-ce base: https://builds.apache.org/job/PreCommit-HBASE-Build/439/artifact/patchprocess/Dockerfile |\r\n| JIRA Issue | HBASE-22487 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12970051/0001-HBASE-22487-getMostLoadedRegions-is-unused.patch |\r\n| Optional Tests | dupname asflicense javac javadoc unit findbugs shadedjars hadoopcheck hbaseanti checkstyle compile |\r\n| uname | Linux b72a46cb1dda 4.4.0-143-generic #169-Ubuntu SMP Thu Feb 7 07:56:38 UTC 2019 x86_64 GNU/Linux |\r\n| Build tool | maven |\r\n| Personality | dev-support/hbase-personality.sh |\r\n| git revision | master / 6ddf893a34 |\r\n| maven | version: Apache Maven 3.5.4 (1edded0938998edf8bf061f1ceb3cfdeccf443fe; 2018-06-17T18:33:14Z) |\r\n| Default Java | 1.8.0_181 |\r\n| findbugs | v3.1.11 |\r\n| unit | https://builds.apache.org/job/PreCommit-HBASE-Build/439/artifact/patchprocess/patch-unit-hbase-server.txt |\r\n| Test Results | https://builds.apache.org/job/PreCommit-HBASE-Build/439/testReport/ |\r\n| Max. process+thread count | 5232 (vs. ulimit of 10000) |\r\n| modules | C: hbase-server U: hbase-server |\r\n| Console output | https://builds.apache.org/job/PreCommit-HBASE-Build/439/console |\r\n| Powered by | Apache Yetus 0.9.0 http://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","created":"2019-05-28T23:44:49.174+0000"},{"body":"lgtm +1 (non binding). Failed UT seems unrelated, ran this patch locally and same test passed.","created":"2019-05-29T12:20:12.357+0000"},{"body":"+1, I'm committing this patch now.","created":"2019-05-30T13:57:13.460+0000"},{"body":"Thanks for the patch [~clayb]. Pushed to all branches.","created":"2019-05-30T16:48:14.638+0000"}],"conversations":[{"body":"Reading {{HRegionServer.java}}, I noticed the function {{getMostLoadedRegions}} this seems to replicate functionality found now in the {{StochasticLoadBalancer.}} Further, it seems the only consumer of this function was removed in HBASE-805.","from":"reporter","subject":"getMostLoadedRegions is unused"},{"body":"This is a trivial patch which simply removes the uncalled code.","from":"developer"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 3m 22s{color} | {color:blue} Docker mode activated. {color} |\r\n|| || || || {color:brown} Prechecks {color} ||\r\n| {color:green}+1{color} | {color:green} hbaseanti {color} | {color:green} 0m 0s{color} | {color:green} Patch does not have any anti-patterns. {color} |\r\n| {color:green}+1{color} | {color:green} @author {color} | {color:green} 0m 0s{color} | {color:green} The patch does not contain any @author tags. {color} |\r\n| {color:orange}-0{color} | {color:orange} test4tests {color} | {color:orange} 0m 0s{color} | {color:orange} The patch doesn't appear to include any new or modified tests. Please justify why no new tests are needed for this patch. Also please list what manual steps were performed to verify this patch. {color} |\r\n|| || || || {color:brown} master Compile Tests {color} ||\r\n| {color:green}+1{color} | {color:green} mvninstall {color} | {color:green} 5m 45s{color} | {color:green} master passed {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 1m 7s{color} | {color:green} master passed {color} |\r\n| {color:green}+1{color} | {color:green} checkstyle {color} | {color:green} 1m 30s{color} | {color:green} master passed {color} |\r\n| {color:green}+1{color} | {color:green} shadedjars {color} | {color:green} 5m 49s{color} | {color:green} branch has no errors when building our shaded downstream artifacts. {color} |\r\n| {color:green}+1{color} | {color:green} findbugs {color} | {color:green} 4m 24s{color} | {color:green} master passed {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 43s{color} | {color:green} master passed {color} |\r\n|| || || || {color:brown} Patch Compile Tests {color} ||\r\n| {color:green}+1{color} | {color:green} mvninstall {color} | {color:green} 5m 18s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 1m 5s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} javac {color} | {color:green} 1m 5s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} checkstyle {color} | {color:green} 1m 29s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} whitespace {color} | {color:green} 0m 0s{color} | {color:green} The patch has no whitespace issues. {color} |\r\n| {color:green}+1{color} | {color:green} shadedjars {color} | {color:green} 5m 48s{color} | {color:green} patch has no errors when building our shaded downstream artifacts. {color} |\r\n| {color:green}+1{color} | {color:green} hadoopcheck {color} | {color:green} 22m 28s{color} | {color:green} Patch does not cause any errors with Hadoop 2.8.5 2.9.2 or 3.0.3 3.1.2. {color} |\r\n| {color:green}+1{color} | {color:green} findbugs {color} | {color:green} 4m 42s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 44s{color} | {color:green} the patch passed {color} |\r\n|| || || || {color:brown} Other Tests {color} ||\r\n| {color:red}-1{color} | {color:red} unit {color} | {color:red}263m 46s{color} | {color:red} hbase-server in the patch failed. {color} |\r\n| {color:green}+1{color} | {color:green} asflicense {color} | {color:green} 1m 31s{color} | {color:green} The patch does not generate ASF License warnings. {color} |\r\n| {color:black}{color} | {color:black} {color} | {color:black}335m 29s{color} | {color:black} {color} |\r\n\\\\\r\n\\\\\r\n|| Reason || Tests ||\r\n| Failed junit tests | hadoop.hbase.replication.multiwal.TestReplicationSyncUpToolWithMultipleWAL |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| Docker | Client=17.05.0-ce Server=17.05.0-ce base: https://builds.apache.org/job/PreCommit-HBASE-Build/439/artifact/patchprocess/Dockerfile |\r\n| JIRA Issue | HBASE-22487 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12970051/0001-HBASE-22487-getMostLoadedRegions-is-unused.patch |\r\n| Optional Tests | dupname asflicense javac javadoc unit findbugs shadedjars hadoopcheck hbaseanti checkstyle compile |\r\n| uname | Linux b72a46cb1dda 4.4.0-143-generic #169-Ubuntu SMP Thu Feb 7 07:56:38 UTC 2019 x86_64 GNU/Linux |\r\n| Build tool | maven |\r\n| Personality | dev-support/hbase-personality.sh |\r\n| git revision | master / 6ddf893a34 |\r\n| maven | version: Apache Maven 3.5.4 (1edded0938998edf8bf061f1ceb3cfdeccf443fe; 2018-06-17T18:33:14Z) |\r\n| Default Java | 1.8.0_181 |\r\n| findbugs | v3.1.11 |\r\n| unit | https://builds.apache.org/job/PreCommit-HBASE-Build/439/artifact/patchprocess/patch-unit-hbase-server.txt |\r\n| Test Results | https://builds.apache.org/job/PreCommit-HBASE-Build/439/testReport/ |\r\n| Max. process+thread count | 5232 (vs. ulimit of 10000) |\r\n| modules | C: hbase-server U: hbase-server |\r\n| Console output | https://builds.apache.org/job/PreCommit-HBASE-Build/439/console |\r\n| Powered by | Apache Yetus 0.9.0 http://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","from":"developer"},{"body":"lgtm +1 (non binding). Failed UT seems unrelated, ran this patch locally and same test passed.","from":"developer"},{"body":"+1, I'm committing this patch now.","from":"developer"},{"body":"Thanks for the patch [~clayb]. Pushed to all branches.","from":"developer"}],"created":"2019-05-28T16:38:39.000+0000","description":"Reading {{HRegionServer.java}}, I noticed the function {{getMostLoadedRegions}} this seems to replicate functionality found now in the {{StochasticLoadBalancer.}} Further, it seems the only consumer of this function was removed in HBASE-805.","issue_id":"13236041","key":"HBASE-22487","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2019-05-30T16:48:14.000+0000","role":"fixed_distractor","summary":"getMostLoadedRegions is unused"} {"case_id":"13237444","cluster":"DISTRACTOR-HBASE-22538","comments":[{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 0m 23s{color} | {color:blue} Docker mode activated. {color} |\r\n|| || || || {color:brown} Prechecks {color} ||\r\n| {color:green}+1{color} | {color:green} @author {color} | {color:green} 0m 0s{color} | {color:green} The patch does not contain any @author tags. {color} |\r\n|| || || || {color:brown} branch-1.4 Compile Tests {color} ||\r\n|| || || || {color:brown} Patch Compile Tests {color} ||\r\n| {color:red}-1{color} | {color:red} rubocop {color} | {color:red} 0m 6s{color} | {color:red} The patch generated 4 new + 353 unchanged - 8 fixed = 357 total (was 361) {color} |\r\n| {color:green}+1{color} | {color:green} ruby-lint {color} | {color:green} 0m 2s{color} | {color:green} There were no new ruby-lint issues. {color} |\r\n| {color:green}+1{color} | {color:green} whitespace {color} | {color:green} 0m 0s{color} | {color:green} The patch has no whitespace issues. {color} |\r\n|| || || || {color:brown} Other Tests {color} ||\r\n| {color:green}+1{color} | {color:green} asflicense {color} | {color:green} 0m 43s{color} | {color:green} The patch does not generate ASF License warnings. {color} |\r\n| {color:black}{color} | {color:black} {color} | {color:black} 2m 9s{color} | {color:black} {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| Docker | Client=17.05.0-ce Server=17.05.0-ce base: https://builds.apache.org/job/PreCommit-HBASE-Build/490/artifact/patchprocess/Dockerfile |\r\n| JIRA Issue | HBASE-22538 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12970813/HBASE-22538.branch-1.4.001.patch |\r\n| Optional Tests | dupname asflicense rubocop ruby_lint |\r\n| uname | Linux a097a654c21d 4.4.0-143-generic #169~14.04.2-Ubuntu SMP Wed Feb 13 15:00:41 UTC 2019 x86_64 x86_64 x86_64 GNU/Linux |\r\n| Build tool | maven |\r\n| Personality | dev-support/hbase-personality.sh |\r\n| git revision | branch-1.4 / 81efa17 |\r\n| maven | version: Apache Maven 3.0.5 |\r\n| rubocop | v0.71.0 |\r\n| rubocop | https://builds.apache.org/job/PreCommit-HBASE-Build/490/artifact/patchprocess/diff-patch-rubocop.txt |\r\n| ruby-lint | v2.3.1 |\r\n| Max. process+thread count | 35 (vs. ulimit of 10000) |\r\n| modules | C: . U: . |\r\n| Console output | https://builds.apache.org/job/PreCommit-HBASE-Build/490/console |\r\n| Powered by | Apache Yetus 0.9.0 http://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","created":"2019-06-04T10:33:53.140+0000"},{"body":"[~Jeongdae Kim] Thanks for the detailed explanation which helped understand the issue and the fix :-).","created":"2019-06-04T10:45:22.964+0000"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 18m 50s{color} | {color:blue} Docker mode activated. {color} |\r\n|| || || || {color:brown} Prechecks {color} ||\r\n| {color:green}+1{color} | {color:green} @author {color} | {color:green} 0m 0s{color} | {color:green} The patch does not contain any @author tags. {color} |\r\n|| || || || {color:brown} branch-1.4 Compile Tests {color} ||\r\n|| || || || {color:brown} Patch Compile Tests {color} ||\r\n| {color:red}-1{color} | {color:red} rubocop {color} | {color:red} 0m 5s{color} | {color:red} The patch generated 2 new + 353 unchanged - 8 fixed = 355 total (was 361) {color} |\r\n| {color:green}+1{color} | {color:green} ruby-lint {color} | {color:green} 0m 2s{color} | {color:green} There were no new ruby-lint issues. {color} |\r\n| {color:green}+1{color} | {color:green} whitespace {color} | {color:green} 0m 0s{color} | {color:green} The patch has no whitespace issues. {color} |\r\n|| || || || {color:brown} Other Tests {color} ||\r\n| {color:green}+1{color} | {color:green} asflicense {color} | {color:green} 0m 47s{color} | {color:green} The patch does not generate ASF License warnings. {color} |\r\n| {color:black}{color} | {color:black} {color} | {color:black} 20m 42s{color} | {color:black} {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| Docker | Client=17.05.0-ce Server=17.05.0-ce base: https://builds.apache.org/job/PreCommit-HBASE-Build/512/artifact/patchprocess/Dockerfile |\r\n| JIRA Issue | HBASE-22538 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12971397/HBASE-22538.branch-1.4.002.patch |\r\n| Optional Tests | dupname asflicense rubocop ruby_lint |\r\n| uname | Linux ec865fcb68e6 4.4.0-138-generic #164-Ubuntu SMP Tue Oct 2 17:16:02 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux |\r\n| Build tool | maven |\r\n| Personality | dev-support/hbase-personality.sh |\r\n| git revision | branch-1.4 / e10bfa0 |\r\n| maven | version: Apache Maven 3.0.5 |\r\n| rubocop | v0.71.0 |\r\n| rubocop | https://builds.apache.org/job/PreCommit-HBASE-Build/512/artifact/patchprocess/diff-patch-rubocop.txt |\r\n| ruby-lint | v2.3.1 |\r\n| Max. process+thread count | 40 (vs. ulimit of 10000) |\r\n| modules | C: . U: . |\r\n| Console output | https://builds.apache.org/job/PreCommit-HBASE-Build/512/console |\r\n| Powered by | Apache Yetus 0.9.0 http://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","created":"2019-06-11T01:32:44.344+0000"},{"body":"I handled 2 related warnings of 4 from rubocop results, and remaining 2 warnings doesn't come from this patch.\r\n{noformat}\r\n/testptch/hbase/bin/region_mover.rb:90:51: C: Style/StringLiterals: Prefer single-quoted strings when you don't need string interpolation or special symbols. \r\n/testptch/hbase/bin/region_mover.rb:91:81: C: Metrics/LineLength: Line is too long. [81/80]{noformat}\r\n{noformat}\r\n/testptch/hbase/bin/region_mover.rb:64:1: C: Metrics/AbcSize: Assignment Branch Condition size for getServerNameForRegion is too high. [41.79/15] \r\n/testptch/hbase/bin/region_mover.rb:64:1: C: Metrics/MethodLength: Method has too many lines. [29/10] {noformat}\r\n \r\n\r\nCould someone review this issue, please?","created":"2019-06-11T02:24:49.225+0000"},{"body":"+1\r\nBetter than the original code, certainly.\r\nLet me commit","created":"2019-06-20T22:20:33.921+0000"}],"conversations":[{"body":"We can stop or restart region servers gracefully using graceful_stop.sh command\r\nThis command should guarantee that all regions are moved out before shutting down a region server.\r\n\r\nHowever, sometimes i saw many requests failed while restarting a region server with this command in our production clusters(v1.2.5)\r\naffected clients got many RegionServerStoppedExceptions and exhausted retry count.\r\n\r\nI found it took 0.03 sec to move a region, it’s too fast. and, moving(unloading) regions in the region server wasn’t finished, even didn’t closed yet when region server got shutdown signal.\r\nBecause a region server serving regions (didn't be closed) were stopped, clients got many exception (RegionServerStoppedException)\r\n\r\nBut, region_mover should wait until a region is served by other region server(meta changed)\r\nhttps://github.com/apache/hbase/blob/branch-1.2/bin/region_mover.rb#L153\r\n\r\nI figured out why this early shutdown happened. \r\na) our clusters use upper case hostname\r\nb) region server makes ServerName with lowercase hostname, and it will be sent to the master\r\nhttps://github.com/apache/hbase/blob/branch-1.2/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java#L542\r\nc) when updating meta, server name will keep its own case\r\nhttps://github.com/apache/hbase/blob/branch-1.2/hbase-client/src/main/java/org/apache/hadoop/hbase/MetaTableAccessor.java#L1527\r\nd) region_mover.rb just compare b) and c), so it is always false\r\nhttps://github.com/apache/hbase/blob/branch-1.2/bin/region_mover.rb#L91\r\nhttps://github.com/apache/hbase/blob/branch-1.2/bin/region_mover.rb#L52\r\n\r\nI think region_mover should compare server name between master and meta with the same case(lower)\r\n\r\nWith patch, I confirmed region_mover waited until finishing moving all regions, then triggered shutting down region sever. (also observed only RegionMovedException before shutdown log, and no exception after starting shutdown)\r\n","from":"reporter","subject":"Prevent graceful_stop.sh from shutting down RS too early before finishing unloading regions"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 0m 23s{color} | {color:blue} Docker mode activated. {color} |\r\n|| || || || {color:brown} Prechecks {color} ||\r\n| {color:green}+1{color} | {color:green} @author {color} | {color:green} 0m 0s{color} | {color:green} The patch does not contain any @author tags. {color} |\r\n|| || || || {color:brown} branch-1.4 Compile Tests {color} ||\r\n|| || || || {color:brown} Patch Compile Tests {color} ||\r\n| {color:red}-1{color} | {color:red} rubocop {color} | {color:red} 0m 6s{color} | {color:red} The patch generated 4 new + 353 unchanged - 8 fixed = 357 total (was 361) {color} |\r\n| {color:green}+1{color} | {color:green} ruby-lint {color} | {color:green} 0m 2s{color} | {color:green} There were no new ruby-lint issues. {color} |\r\n| {color:green}+1{color} | {color:green} whitespace {color} | {color:green} 0m 0s{color} | {color:green} The patch has no whitespace issues. {color} |\r\n|| || || || {color:brown} Other Tests {color} ||\r\n| {color:green}+1{color} | {color:green} asflicense {color} | {color:green} 0m 43s{color} | {color:green} The patch does not generate ASF License warnings. {color} |\r\n| {color:black}{color} | {color:black} {color} | {color:black} 2m 9s{color} | {color:black} {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| Docker | Client=17.05.0-ce Server=17.05.0-ce base: https://builds.apache.org/job/PreCommit-HBASE-Build/490/artifact/patchprocess/Dockerfile |\r\n| JIRA Issue | HBASE-22538 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12970813/HBASE-22538.branch-1.4.001.patch |\r\n| Optional Tests | dupname asflicense rubocop ruby_lint |\r\n| uname | Linux a097a654c21d 4.4.0-143-generic #169~14.04.2-Ubuntu SMP Wed Feb 13 15:00:41 UTC 2019 x86_64 x86_64 x86_64 GNU/Linux |\r\n| Build tool | maven |\r\n| Personality | dev-support/hbase-personality.sh |\r\n| git revision | branch-1.4 / 81efa17 |\r\n| maven | version: Apache Maven 3.0.5 |\r\n| rubocop | v0.71.0 |\r\n| rubocop | https://builds.apache.org/job/PreCommit-HBASE-Build/490/artifact/patchprocess/diff-patch-rubocop.txt |\r\n| ruby-lint | v2.3.1 |\r\n| Max. process+thread count | 35 (vs. ulimit of 10000) |\r\n| modules | C: . U: . |\r\n| Console output | https://builds.apache.org/job/PreCommit-HBASE-Build/490/console |\r\n| Powered by | Apache Yetus 0.9.0 http://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","from":"developer"},{"body":"[~Jeongdae Kim] Thanks for the detailed explanation which helped understand the issue and the fix :-).","from":"developer"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 18m 50s{color} | {color:blue} Docker mode activated. {color} |\r\n|| || || || {color:brown} Prechecks {color} ||\r\n| {color:green}+1{color} | {color:green} @author {color} | {color:green} 0m 0s{color} | {color:green} The patch does not contain any @author tags. {color} |\r\n|| || || || {color:brown} branch-1.4 Compile Tests {color} ||\r\n|| || || || {color:brown} Patch Compile Tests {color} ||\r\n| {color:red}-1{color} | {color:red} rubocop {color} | {color:red} 0m 5s{color} | {color:red} The patch generated 2 new + 353 unchanged - 8 fixed = 355 total (was 361) {color} |\r\n| {color:green}+1{color} | {color:green} ruby-lint {color} | {color:green} 0m 2s{color} | {color:green} There were no new ruby-lint issues. {color} |\r\n| {color:green}+1{color} | {color:green} whitespace {color} | {color:green} 0m 0s{color} | {color:green} The patch has no whitespace issues. {color} |\r\n|| || || || {color:brown} Other Tests {color} ||\r\n| {color:green}+1{color} | {color:green} asflicense {color} | {color:green} 0m 47s{color} | {color:green} The patch does not generate ASF License warnings. {color} |\r\n| {color:black}{color} | {color:black} {color} | {color:black} 20m 42s{color} | {color:black} {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| Docker | Client=17.05.0-ce Server=17.05.0-ce base: https://builds.apache.org/job/PreCommit-HBASE-Build/512/artifact/patchprocess/Dockerfile |\r\n| JIRA Issue | HBASE-22538 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12971397/HBASE-22538.branch-1.4.002.patch |\r\n| Optional Tests | dupname asflicense rubocop ruby_lint |\r\n| uname | Linux ec865fcb68e6 4.4.0-138-generic #164-Ubuntu SMP Tue Oct 2 17:16:02 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux |\r\n| Build tool | maven |\r\n| Personality | dev-support/hbase-personality.sh |\r\n| git revision | branch-1.4 / e10bfa0 |\r\n| maven | version: Apache Maven 3.0.5 |\r\n| rubocop | v0.71.0 |\r\n| rubocop | https://builds.apache.org/job/PreCommit-HBASE-Build/512/artifact/patchprocess/diff-patch-rubocop.txt |\r\n| ruby-lint | v2.3.1 |\r\n| Max. process+thread count | 40 (vs. ulimit of 10000) |\r\n| modules | C: . U: . |\r\n| Console output | https://builds.apache.org/job/PreCommit-HBASE-Build/512/console |\r\n| Powered by | Apache Yetus 0.9.0 http://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","from":"developer"},{"body":"I handled 2 related warnings of 4 from rubocop results, and remaining 2 warnings doesn't come from this patch.\r\n{noformat}\r\n/testptch/hbase/bin/region_mover.rb:90:51: C: Style/StringLiterals: Prefer single-quoted strings when you don't need string interpolation or special symbols. \r\n/testptch/hbase/bin/region_mover.rb:91:81: C: Metrics/LineLength: Line is too long. [81/80]{noformat}\r\n{noformat}\r\n/testptch/hbase/bin/region_mover.rb:64:1: C: Metrics/AbcSize: Assignment Branch Condition size for getServerNameForRegion is too high. [41.79/15] \r\n/testptch/hbase/bin/region_mover.rb:64:1: C: Metrics/MethodLength: Method has too many lines. [29/10] {noformat}\r\n \r\n\r\nCould someone review this issue, please?","from":"developer"},{"body":"+1\r\nBetter than the original code, certainly.\r\nLet me commit","from":"developer"}],"created":"2019-06-04T10:10:35.000+0000","description":"We can stop or restart region servers gracefully using graceful_stop.sh command\r\nThis command should guarantee that all regions are moved out before shutting down a region server.\r\n\r\nHowever, sometimes i saw many requests failed while restarting a region server with this command in our production clusters(v1.2.5)\r\naffected clients got many RegionServerStoppedExceptions and exhausted retry count.\r\n\r\nI found it took 0.03 sec to move a region, it’s too fast. and, moving(unloading) regions in the region server wasn’t finished, even didn’t closed yet when region server got shutdown signal.\r\nBecause a region server serving regions (didn't be closed) were stopped, clients got many exception (RegionServerStoppedException)\r\n\r\nBut, region_mover should wait until a region is served by other region server(meta changed)\r\nhttps://github.com/apache/hbase/blob/branch-1.2/bin/region_mover.rb#L153\r\n\r\nI figured out why this early shutdown happened. \r\na) our clusters use upper case hostname\r\nb) region server makes ServerName with lowercase hostname, and it will be sent to the master\r\nhttps://github.com/apache/hbase/blob/branch-1.2/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java#L542\r\nc) when updating meta, server name will keep its own case\r\nhttps://github.com/apache/hbase/blob/branch-1.2/hbase-client/src/main/java/org/apache/hadoop/hbase/MetaTableAccessor.java#L1527\r\nd) region_mover.rb just compare b) and c), so it is always false\r\nhttps://github.com/apache/hbase/blob/branch-1.2/bin/region_mover.rb#L91\r\nhttps://github.com/apache/hbase/blob/branch-1.2/bin/region_mover.rb#L52\r\n\r\nI think region_mover should compare server name between master and meta with the same case(lower)\r\n\r\nWith patch, I confirmed region_mover waited until finishing moving all regions, then triggered shutting down region sever. (also observed only RegionMovedException before shutdown log, and no exception after starting shutdown)\r\n","issue_id":"13237444","key":"HBASE-22538","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2019-06-20T22:25:36.000+0000","role":"fixed_distractor","summary":"Prevent graceful_stop.sh from shutting down RS too early before finishing unloading regions"} {"case_id":"12457251","cluster":"DISTRACTOR-HBASE-2256","comments":[{"body":"isnt this due to the millisecond timestamp nature of our puts and delete markers? If the next put has the same millisecond TS as an existing delete record you will be masked by the previous delete.","created":"2010-02-24T03:05:29.752+0000"},{"body":"Yes, I'm sure that the millisecond nature of timestamps comes in to play here. However, I'm not setting any timestamps, and was under the impression that hbase would always reflect the state of the last operation done. Is this not a valid assumption?\n\nA related question. If I do two puts (w/latest timestamp), am I guaranteed to see the last one? I'm sure many users operate under this assumption.\n\nSo, for a given row, I'm doing a delete of an entire row, then a put of two cells in different families. Then I do a get.\n\nMost times I'll see all of the latest put. Sometimes I see nothing for that row. Sometimes I see just one family/cell from the previous put. It seems with 0.20.2 I would either get all the row, or none. However with 0.20.3 I will see all or just one family. Returning just one family seems even more wrong.\n\nSo suppose I wish to wipe existing cells of a row, then write some new cells. Is the only way to reliably do this to put a 1 ms pause between the delete and the put? This would hurt my throughput...\n\n","created":"2010-02-24T18:09:16.150+0000"},{"body":"bq. Yes, I'm sure that the millisecond nature of timestamps comes in to play here. However, I'm not setting any timestamps, and was under the impression that hbase would always reflect the state of the last operation done. Is this not a valid assumption?\n\nDo you see this delete problem even on a fully distributed setup? In my experience, they only happen in unit tests where all components are in the same JVM whereas when network is involved some milliseconds will separate two consecutive operations. \n\nbq. A related question. If I do two puts (w/latest timestamp), am I guaranteed to see the last one? I'm sure many users operate under this assumption.\n\nIf they have the same timestamp, there's no guarantee.\n\nbq. So, for a given row, I'm doing a delete of an entire row, then a put of two cells in different families. Then I do a get.\n\nSee my first comment, it's ok when not done in unit tests. Like the bigtable paper says, if you need a finer granularity than millisecond you may need to redefine the timestamps (using something like microseconds).","created":"2010-02-24T18:15:57.002+0000"},{"body":"I have only noticed this in unit tests and a non-distributed setup. However, this Delete,Put happens in ITHbase's IndexRegion which means that even in a distributed setup the client for the put/delete and the regionserver handling them could be in the same JVM.\n\nFor the put, put case, it seems to me this could be a real issue. Even in distributed setup sequential puts could happen in the same ms no? However, I did a similar test for Put after Put and it seems to always work. If it did not, I'm sure users would have complained loudly by now.\n\nFrom my point of view, it would be nice to have this behavior for Put after Delete as well. \n\nI'm not saying I need finer granularity tham ms. Just that when I'm never explicitly messing with timestamps, I always see a \"correct\" view that reflects my last operation. I just skimmed over the bigtable paper, and could not find an explication about what they do in this case..\n\n\n","created":"2010-02-24T18:49:26.621+0000"},{"body":"there is a millisecond resolution, and it might be difficult to get better without changing the storage format so we can get nanos in there.\n\nbut still, for most people, doing a put - delete - put all within 1 millisecond is a not common. maybe it might be possible to change something so we dont have to run up against this issue?","created":"2010-02-26T08:17:19.386+0000"},{"body":"What if I could do something like this:\n\nPut put1 = ...\nHTable.put(put1);\nDelete delete = new Delete(...).guaranteeAfter(put1);\nHTable.delete(delete);\nPut put2 = new Put(...).guaranteeAfter(delete);\nHTable.put(put2)\n\nIt seems like the distributed case isn't a problem since it's so unlikely, but the delete then put seems more plausible. We could set the timestamp to 1ms after the delete if needed. The occasional write will get a timestamp a few ms in the future, which doesn't seem that bad. I think this solves Clint's requirement of seeing a correct view without explicitly messing with timestamps.","created":"2010-04-22T18:07:52.176+0000"},{"body":"We ran into this problem recently in our production code. A single hbase client needed to first clean up of several columns by deleting them and then put a subset those columns back in with new values. Frequently the delete and put call would happen in the same millisecond thus masking the put. For now we have implemented a fix on our side but would be nice to see a real fix for this, where the region servers handle this more gracefully.\n\nMaybe we could log some warnings when this occurs for possible easier debugging? This was an extremely difficult problem to find.\n\nOr maybe someone has a clever solution? \n ","created":"2011-03-22T23:00:00.528+0000"},{"body":"@Nathaniel This should be fixed as by-product of hbase-2856.","created":"2011-03-24T23:09:22.980+0000"},{"body":"Currently timestamp for Put and Delete is in milliseconds.\nIf we can use System.nanoTime(), the chance of this issue happening would be very low.","created":"2011-04-07T20:24:26.974+0000"},{"body":"Probably can't use nano time, it wraps around too frequently.\n","created":"2011-04-07T20:30:07.296+0000"},{"body":"From http://download.oracle.com/javase/1.5.0/docs/api/java/lang/System.html#nanoTime%28%29 :\nDifferences in successive calls that span greater than approximately 292 years (263 nanoseconds) will not accurately compute elapsed time due to numerical overflow. ","created":"2011-04-07T20:41:43.253+0000"},{"body":"more importantly, nanotime is elapsed time since some arbitrary system-local reference point, and has no meaning on an absolute scale.","created":"2011-04-07T20:44:46.120+0000"},{"body":"{code}\n\t long l = System.nanoTime();\n\t long l2 = System.currentTimeMillis();\n{code}\nLooking at the values of l (1302209826865074000) and l2 (1302209826865), nanoTime is aligned with time in millis.\nAssuming nano and milli timestamps correlate, we can devise (correction) mechanism in master and region servers such that (corrected) nano timestamp reflects the actual millisecond timestamp.","created":"2011-04-07T21:02:08.136+0000"},{"body":"I think this would be a hacky non-solution, regardless of whether it's epoch nanos or not.","created":"2011-04-07T21:09:15.364+0000"},{"body":"We(XiaoMi) fixed this issue with introducing a ScanDeleteTrackerWithMVCC.\nMy workmate [~fenghh] will upload a patch soon.","created":"2013-05-27T12:54:20.194+0000"},{"body":"Please refer to https://issues.apache.org/jira/browse/HBASE-8721 for our fix","created":"2013-06-09T10:08:55.643+0000"},{"body":"Also encountered similar problem. What about this solution?\n{code}\npublic class IncrementingWallTimeEnvironmentEdge implements EnvironmentEdge {\n private long clock = -1;\n\n public IncrementingWallTimeEnvironmentEdge() {\n }\n\n @Override\n public long currentTimeMillis() {\n long wallTime = System.currentTimeMillis() << 10; // ~us, or any arbitrary scaling factor\n\n synchronized (this) {\n if (clock < wallTime) {\n clock = wallTime;\n }\n return clock++;\n }\n }\n}\n{code}\nThis would solve this problem and guarantee the timestamp aligned with the wall time clock in milliseconds as long as we set the scaling factor to a larger enough number (i.e. make sure the speed of logical clock is slower than System.currentTimeMillis()). Shift factor from 10 ~20 (1M-1G qps) is proper value for current server configuration which also will not introduce wrapping around concern (584M ~ 0.57M year).","created":"2013-11-18T08:08:44.922+0000"},{"body":"[~heliangliang] I like this notion (It is related a bit to HBASE-8927). The call to currentTimeMillis is done frequently. I think we'd combine this change with an attempt at removing as many calls to currentTimeMillis as possible. We might replace the currentTimeMillis calls that are for tiiming with nanotime calls instead and use this new class for Cell version).","created":"2013-11-22T08:02:28.223+0000"},{"body":"It happened in HBase 1.2.3. My code is \n{code}\nDelete del = new Delete(row.getBytes());\ntable.delete(del);\n\nList puts = myPuts();\ntable.put(puts);\n{code}\n\nMost of time worked fine. But sometimes the data in HBase went crazy. So I don't trust the data stored.","created":"2016-11-11T09:56:25.917+0000"},{"body":"Please open a new issue [~gfeng] This one is a long time closed. World has changed a bunch since this issue too so probably different cause. Thanks.","created":"2016-11-11T17:34:00.590+0000"},{"body":"HBASE-15968 has been resolved. It adds option to enable change in comparator which should fix this issue.","created":"2017-07-30T12:02:28.878+0000"}],"conversations":[{"body":"Doing a Delete of a whole row, followed immediately by a put to that row will sometimes miss a cell. Attached is a test to provoke the issue.","from":"reporter","subject":"Delete row, followed quickly to put of the same row will sometimes fail."},{"body":"isnt this due to the millisecond timestamp nature of our puts and delete markers? If the next put has the same millisecond TS as an existing delete record you will be masked by the previous delete.","from":"developer"},{"body":"Yes, I'm sure that the millisecond nature of timestamps comes in to play here. However, I'm not setting any timestamps, and was under the impression that hbase would always reflect the state of the last operation done. Is this not a valid assumption?\n\nA related question. If I do two puts (w/latest timestamp), am I guaranteed to see the last one? I'm sure many users operate under this assumption.\n\nSo, for a given row, I'm doing a delete of an entire row, then a put of two cells in different families. Then I do a get.\n\nMost times I'll see all of the latest put. Sometimes I see nothing for that row. Sometimes I see just one family/cell from the previous put. It seems with 0.20.2 I would either get all the row, or none. However with 0.20.3 I will see all or just one family. Returning just one family seems even more wrong.\n\nSo suppose I wish to wipe existing cells of a row, then write some new cells. Is the only way to reliably do this to put a 1 ms pause between the delete and the put? This would hurt my throughput...\n\n","from":"developer"},{"body":"bq. Yes, I'm sure that the millisecond nature of timestamps comes in to play here. However, I'm not setting any timestamps, and was under the impression that hbase would always reflect the state of the last operation done. Is this not a valid assumption?\n\nDo you see this delete problem even on a fully distributed setup? In my experience, they only happen in unit tests where all components are in the same JVM whereas when network is involved some milliseconds will separate two consecutive operations. \n\nbq. A related question. If I do two puts (w/latest timestamp), am I guaranteed to see the last one? I'm sure many users operate under this assumption.\n\nIf they have the same timestamp, there's no guarantee.\n\nbq. So, for a given row, I'm doing a delete of an entire row, then a put of two cells in different families. Then I do a get.\n\nSee my first comment, it's ok when not done in unit tests. Like the bigtable paper says, if you need a finer granularity than millisecond you may need to redefine the timestamps (using something like microseconds).","from":"developer"},{"body":"I have only noticed this in unit tests and a non-distributed setup. However, this Delete,Put happens in ITHbase's IndexRegion which means that even in a distributed setup the client for the put/delete and the regionserver handling them could be in the same JVM.\n\nFor the put, put case, it seems to me this could be a real issue. Even in distributed setup sequential puts could happen in the same ms no? However, I did a similar test for Put after Put and it seems to always work. If it did not, I'm sure users would have complained loudly by now.\n\nFrom my point of view, it would be nice to have this behavior for Put after Delete as well. \n\nI'm not saying I need finer granularity tham ms. Just that when I'm never explicitly messing with timestamps, I always see a \"correct\" view that reflects my last operation. I just skimmed over the bigtable paper, and could not find an explication about what they do in this case..\n\n\n","from":"developer"},{"body":"there is a millisecond resolution, and it might be difficult to get better without changing the storage format so we can get nanos in there.\n\nbut still, for most people, doing a put - delete - put all within 1 millisecond is a not common. maybe it might be possible to change something so we dont have to run up against this issue?","from":"developer"},{"body":"What if I could do something like this:\n\nPut put1 = ...\nHTable.put(put1);\nDelete delete = new Delete(...).guaranteeAfter(put1);\nHTable.delete(delete);\nPut put2 = new Put(...).guaranteeAfter(delete);\nHTable.put(put2)\n\nIt seems like the distributed case isn't a problem since it's so unlikely, but the delete then put seems more plausible. We could set the timestamp to 1ms after the delete if needed. The occasional write will get a timestamp a few ms in the future, which doesn't seem that bad. I think this solves Clint's requirement of seeing a correct view without explicitly messing with timestamps.","from":"developer"},{"body":"We ran into this problem recently in our production code. A single hbase client needed to first clean up of several columns by deleting them and then put a subset those columns back in with new values. Frequently the delete and put call would happen in the same millisecond thus masking the put. For now we have implemented a fix on our side but would be nice to see a real fix for this, where the region servers handle this more gracefully.\n\nMaybe we could log some warnings when this occurs for possible easier debugging? This was an extremely difficult problem to find.\n\nOr maybe someone has a clever solution? \n ","from":"developer"},{"body":"@Nathaniel This should be fixed as by-product of hbase-2856.","from":"developer"},{"body":"Currently timestamp for Put and Delete is in milliseconds.\nIf we can use System.nanoTime(), the chance of this issue happening would be very low.","from":"developer"},{"body":"Probably can't use nano time, it wraps around too frequently.\n","from":"developer"},{"body":"From http://download.oracle.com/javase/1.5.0/docs/api/java/lang/System.html#nanoTime%28%29 :\nDifferences in successive calls that span greater than approximately 292 years (263 nanoseconds) will not accurately compute elapsed time due to numerical overflow. ","from":"developer"},{"body":"more importantly, nanotime is elapsed time since some arbitrary system-local reference point, and has no meaning on an absolute scale.","from":"developer"},{"body":"{code}\n\t long l = System.nanoTime();\n\t long l2 = System.currentTimeMillis();\n{code}\nLooking at the values of l (1302209826865074000) and l2 (1302209826865), nanoTime is aligned with time in millis.\nAssuming nano and milli timestamps correlate, we can devise (correction) mechanism in master and region servers such that (corrected) nano timestamp reflects the actual millisecond timestamp.","from":"developer"},{"body":"I think this would be a hacky non-solution, regardless of whether it's epoch nanos or not.","from":"developer"},{"body":"We(XiaoMi) fixed this issue with introducing a ScanDeleteTrackerWithMVCC.\nMy workmate [~fenghh] will upload a patch soon.","from":"developer"},{"body":"Please refer to https://issues.apache.org/jira/browse/HBASE-8721 for our fix","from":"developer"},{"body":"Also encountered similar problem. What about this solution?\n{code}\npublic class IncrementingWallTimeEnvironmentEdge implements EnvironmentEdge {\n private long clock = -1;\n\n public IncrementingWallTimeEnvironmentEdge() {\n }\n\n @Override\n public long currentTimeMillis() {\n long wallTime = System.currentTimeMillis() << 10; // ~us, or any arbitrary scaling factor\n\n synchronized (this) {\n if (clock < wallTime) {\n clock = wallTime;\n }\n return clock++;\n }\n }\n}\n{code}\nThis would solve this problem and guarantee the timestamp aligned with the wall time clock in milliseconds as long as we set the scaling factor to a larger enough number (i.e. make sure the speed of logical clock is slower than System.currentTimeMillis()). Shift factor from 10 ~20 (1M-1G qps) is proper value for current server configuration which also will not introduce wrapping around concern (584M ~ 0.57M year).","from":"developer"},{"body":"[~heliangliang] I like this notion (It is related a bit to HBASE-8927). The call to currentTimeMillis is done frequently. I think we'd combine this change with an attempt at removing as many calls to currentTimeMillis as possible. We might replace the currentTimeMillis calls that are for tiiming with nanotime calls instead and use this new class for Cell version).","from":"developer"},{"body":"It happened in HBase 1.2.3. My code is \n{code}\nDelete del = new Delete(row.getBytes());\ntable.delete(del);\n\nList puts = myPuts();\ntable.put(puts);\n{code}\n\nMost of time worked fine. But sometimes the data in HBase went crazy. So I don't trust the data stored.","from":"developer"},{"body":"Please open a new issue [~gfeng] This one is a long time closed. World has changed a bunch since this issue too so probably different cause. Thanks.","from":"developer"},{"body":"HBASE-15968 has been resolved. It adds option to enable change in comparator which should fix this issue.","from":"developer"}],"created":"2010-02-24T03:00:41.000+0000","description":"Doing a Delete of a whole row, followed immediately by a put to that row will sometimes miss a cell. Attached is a test to provoke the issue.","issue_id":"12457251","key":"HBASE-2256","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2022-06-11T23:07:55.000+0000","role":"fixed_distractor","summary":"Delete row, followed quickly to put of the same row will sometimes fail."} {"case_id":"13249233","cluster":"DISTRACTOR-HBASE-22806","comments":[{"body":"bq. Expected: no cells (all cells marked as deleted) in CF. Actual: C1, C2 shows up automatically in CF\r\nIt makes sense to me. Please try imagine your CF have billions of cells in production env, it's an unacceptable to tomb every cells just after calling Admin#deleteColumnFamily. First it's time consuming, second it's a whole CF full scan which would press you RS. If you take a look at the source code, you will understand why C1 & C2 show up again.\r\n\r\nbq. if doing Admin.modifyColumnFamily() without actually changing anything, then after step 4 the cells won't come back.\r\nHow do you build the {{ColumnFamilyDescriptor}}? Calling modifyColumnFamily will overwrite your table's meta data, C1 & C2 are still there and don't get tombed but just you can't see it because of meta mismatch, if i understand correctly.\r\n\r\n","created":"2019-08-07T03:29:21.816+0000"},{"body":"Here's a complete code example:\r\n\r\n{code:java}\r\nTableName tableName = TableName.valueOf(\"table\");\r\nbyte[] family1 = Bytes.toBytes(\"family1\");\r\nbyte[] family2 = Bytes.toBytes(\"family2\");\r\nbyte[] row = Bytes.toBytes(\"row\");\r\nbyte[] qualifier = Bytes.toBytes(\"qualifier\");\r\nbyte[] value = Bytes.toBytes(\"value\");\r\n\r\nColumnFamilyDescriptor familyDesc1 =\r\n ColumnFamilyDescriptorBuilder\r\n .newBuilder(family1)\r\n .build();\r\nColumnFamilyDescriptor familyDesc2 =\r\n ColumnFamilyDescriptorBuilder\r\n .newBuilder(family2)\r\n .build();\r\n\r\n\r\nTableDescriptor tableDesc = TableDescriptorBuilder\r\n .newBuilder(tableName)\r\n .setColumnFamily(familyDesc1)\r\n .setColumnFamily(familyDesc2)\r\n .build();\r\n\r\ntry {\r\n Connection conn = ConnectionFactory.createConnection();\r\n Admin admin = conn.getAdmin();\r\n\r\n // step 1\r\n admin.createTable(tableDesc);\r\n\r\n // step 2\r\n Table table = conn.getTable(tableName);\r\n table.put(new Put(row).addColumn(family2, qualifier, value));\r\n TableDescriptor currentDesc = table.getDescriptor();\r\n table.close();\r\n\r\n // step 3 (addition): uncommenting this no-op modification makes this work\r\n // admin.modifyTable(currentDesc);\r\n\r\n // read the data before deleting and recreating (it should exist)\r\n table = conn.getTable(tableName);\r\n Result resultGood = table.get(new Get(row));\r\n System.out.println(\"should be 'value': \" + Bytes.toStringBinary(resultGood.value()));\r\n table.close();\r\n\r\n // step 3\r\n admin.deleteColumnFamily(tableName, family2);\r\n\r\n // step 4\r\n admin.addColumnFamily(tableName, familyDesc2);\r\n\r\n /*\r\n // also occurs with two modify tables\r\n TableDescriptor removedDesc =\r\n TableDescriptorBuilder.newBuilder(tableDesc).removeColumnFamily(family2).build();\r\n // step 3\r\n admin.modifyTable(removedDesc);\r\n // step 4\r\n admin.modifyTable(tableDesc);\r\n */\r\n\r\n // read the data after deleting and recreating (it should exist)\r\n table = conn.getTable(tableName);\r\n Result resultBad = table.get(new Get(row));\r\n System.out.println(\"should be null: \" + Bytes.toStringBinary(resultBad.value()));\r\n table.close();\r\n\r\n // cleanup\r\n admin.disableTable(tableName);\r\n admin.deleteTable(tableName);\r\n} catch (IOException e) {\r\n System.err.println(e);\r\n}\r\n{code}\r\n\r\nThe output looks like:\r\n\r\n{code:none}\r\nshould be 'value': value\r\nshould be null: value\r\n{code}\r\n\r\nIf the no-op {{modifyTable}} is uncommented, the output is:\r\n\r\n{code:none}\r\nshould be 'value': value\r\nshould be null: null\r\n{code}\r\n\r\nI've (tried to...?) attached a project that should compile and run against a {{localhost}} HBase. (I was running it in IntelliJ.)","created":"2019-08-07T04:15:39.894+0000"},{"body":"Maybe related to HBASE-21596.","created":"2019-08-07T08:27:36.026+0000"},{"body":"{quote}Please try imagine your CF have billions of cells in production env, it's an unacceptable to tomb every cells just after calling Admin#deleteColumnFamily.\r\n{quote}\r\nMy understanding is CFs are physically different files/directories on HDFS and therefore a delete column family should be a simple HDFS atomic delete or move/rename operation; tombing cells, I would think, should not be necessary - is this not the case with HBase's implementation?\r\n\r\nI will note that Google Bigtable [explicitly calls out| [https://cloud.google.com/bigtable/docs/managing-tables#deleting_column_families]] data from deleted column families are unrecoverable.\r\n\r\nDoes this have the potential for data, once protected by a column family scoped ACL, to no longer be secure?","created":"2019-08-08T02:44:08.318+0000"},{"body":"bq. is this not the case with HBase's implementation?\r\nNo, it is not.\r\n\r\nbq. Does this have the potential for data, once protected by a column family scoped ACL, to no longer be secure?\r\nTo be more specific? Can't understand the question.","created":"2019-08-08T03:10:58.735+0000"},{"body":"[~wchevreuil] hmm no, different story IMO.","created":"2019-08-08T03:11:48.868+0000"},{"body":"{{admin.modifyTable(currentDesc);}} this method may delete cf on HDFS as you said, but not {{deleteColumnFamily}} and {{modify***}} this kind of CF levels.\r\n[~tmoschou]","created":"2019-08-08T03:19:05.590+0000"},{"body":"{quote}How do you build the {{ColumnFamilyDescriptor}}?\r\n{quote}\r\n\r\nUsing {{ColumnFamilyDescriptorBuilder.newBuilder(existingCf)}} without changing any config will give the expected result.\r\n\r\n[~reidchan]","created":"2019-08-08T03:20:42.694+0000"},{"body":"Please check source code: {{HMaster#modifyColumn, HMaster#deleteColumn, HMaster#modifyTable}}\r\n\r\nif you find any unreasonable or bug, feel free to attach patches.","created":"2019-08-08T03:46:57.092+0000"},{"body":"I think it's clear that there is _something_ weird going on, because {{admin.modifyTable(currentDesc)}} seems like it should do nothing at all (it is \"modifying\" the table to have the same {{TableDescriptor}}), but:\r\n\r\n* the cells are _not_ deleted when that call doesn't run \r\n* the cells are deleted when that call does run\r\n\r\nDo you agree that this is unreasonable?\r\n\r\nPlease confirm whether this is behaviour is correct or not, before we waste time trying to fix it.\r\n","created":"2019-08-08T04:07:22.961+0000"},{"body":"bq. it should do nothing .. have the same TableDescriptor.. ,the cells are deleted when that call does run.\r\nYes, it is unreasonable.\r\n\r\nCommunity encourages volunteering, but under your context 'waste time' which sounds unwilling and being forced, I'd suggest you just set it aside and let someone have interest to take it.:)\r\n","created":"2019-08-08T04:20:42.935+0000"},{"body":"Thanks.\r\n\r\nIt would only be a waste of time if it wasn't actually a problem (for example, if we spent two days trying to fix the issue, only for our patch to be rejected because the current behaviour is correct), so I wanted to make sure we had confirmation that it was a bug.","created":"2019-08-08T04:24:54.165+0000"},{"body":"The problem happens when region memstore contain entries for the deleted CF and CF is deleted dynamically (without disabling the table). Since we delete the CF from FS first and then reopen the region, during reopen RS will flush the memstore content to FS. So deleted CF store will contain the memstore content for the deleted CF. ","created":"2019-08-20T07:00:19.258+0000"},{"body":"(y)[~pankaj2461], make sense to me, is there an upcoming patch? (expecting)","created":"2019-08-20T07:06:06.340+0000"},{"body":"So actual problem is the stale CF.","created":"2019-08-20T07:13:18.296+0000"},{"body":"Attached the UT patch.","created":"2019-08-20T07:14:41.561+0000"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 0m 0s{color} | {color:blue} Docker mode activated. {color} |\r\n| {color:red}-1{color} | {color:red} patch {color} | {color:red} 0m 8s{color} | {color:red} HBASE-22806 does not apply to master. Rebase required? Wrong Branch? See https://yetus.apache.org/documentation/in-progress/precommit-patchnames for help. {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| JIRA Issue | HBASE-22806 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12978032/HBASE-22806_UT.patch |\r\n| Console output | https://builds.apache.org/job/PreCommit-HBASE-Build/806/console |\r\n| Powered by | Apache Yetus 0.9.0 http://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","created":"2019-08-21T03:30:27.690+0000"},{"body":"ping [~pankaj2461], please check the QA warning.","created":"2019-08-21T03:34:35.377+0000"},{"body":"[~reidchan] I just uploaded the UT to reproduce this issue. I'm checking the solution, will upload the patch soon.","created":"2019-08-22T11:17:01.373+0000"},{"body":"[~reidchan] Have raised the PR, kindly review.","created":"2019-08-23T11:49:11.245+0000"},{"body":"Pushed on branch-2.1+. Probably not pertinent on branch-1. Thanks for the patch [~pankaj2461].","created":"2019-08-25T02:52:12.557+0000"},{"body":"Issue is applicable to branch-1 also but fix will be different as BulkReOpen reopens the region asynchronously unlike GeneralBulkAssigner,\r\n\r\n{code}\r\n\r\n/**\r\n * Reopen the regions asynchronously, so always returns true immediately.\r\n * @return true\r\n */\r\n @Override\r\n protected boolean waitUntilDone(long timeout) {\r\n return true;\r\n }\r\n\r\n{code}\r\n\r\nTo fix this problem in branch-1, we need to  add region reopen completion wait logic or disable-enable the table.\r\n\r\n[~apurtell]  Please provide your opinion.","created":"2019-08-25T10:20:54.628+0000"},{"body":"bq. To fix this problem in branch-1, we need to add region reopen completion wait logic or disable-enable the table.\r\n\r\nNeither option sounds great. Adding the wait-for-completion seems less disruptive. We can try that. [~pankaj2461]","created":"2019-08-27T22:39:22.252+0000"}],"conversations":[{"body":"While deleting the CF dynamically (without disabling the table), CF dirs are not cleared from FS when region memstore contain entries for that CF. \r\n\r\nSince we delete the CF from FS first and then reopen the region, during reopen RS will flush the memstore content to FS. So deleted CF store will contain the memstore content for the deleted CF. \r\n\r\nSo adding back same CF will have old entries.","from":"reporter","subject":"Deleted CF are not cleared if memstore contain entries"},{"body":"bq. Expected: no cells (all cells marked as deleted) in CF. Actual: C1, C2 shows up automatically in CF\r\nIt makes sense to me. Please try imagine your CF have billions of cells in production env, it's an unacceptable to tomb every cells just after calling Admin#deleteColumnFamily. First it's time consuming, second it's a whole CF full scan which would press you RS. If you take a look at the source code, you will understand why C1 & C2 show up again.\r\n\r\nbq. if doing Admin.modifyColumnFamily() without actually changing anything, then after step 4 the cells won't come back.\r\nHow do you build the {{ColumnFamilyDescriptor}}? Calling modifyColumnFamily will overwrite your table's meta data, C1 & C2 are still there and don't get tombed but just you can't see it because of meta mismatch, if i understand correctly.\r\n\r\n","from":"developer"},{"body":"Here's a complete code example:\r\n\r\n{code:java}\r\nTableName tableName = TableName.valueOf(\"table\");\r\nbyte[] family1 = Bytes.toBytes(\"family1\");\r\nbyte[] family2 = Bytes.toBytes(\"family2\");\r\nbyte[] row = Bytes.toBytes(\"row\");\r\nbyte[] qualifier = Bytes.toBytes(\"qualifier\");\r\nbyte[] value = Bytes.toBytes(\"value\");\r\n\r\nColumnFamilyDescriptor familyDesc1 =\r\n ColumnFamilyDescriptorBuilder\r\n .newBuilder(family1)\r\n .build();\r\nColumnFamilyDescriptor familyDesc2 =\r\n ColumnFamilyDescriptorBuilder\r\n .newBuilder(family2)\r\n .build();\r\n\r\n\r\nTableDescriptor tableDesc = TableDescriptorBuilder\r\n .newBuilder(tableName)\r\n .setColumnFamily(familyDesc1)\r\n .setColumnFamily(familyDesc2)\r\n .build();\r\n\r\ntry {\r\n Connection conn = ConnectionFactory.createConnection();\r\n Admin admin = conn.getAdmin();\r\n\r\n // step 1\r\n admin.createTable(tableDesc);\r\n\r\n // step 2\r\n Table table = conn.getTable(tableName);\r\n table.put(new Put(row).addColumn(family2, qualifier, value));\r\n TableDescriptor currentDesc = table.getDescriptor();\r\n table.close();\r\n\r\n // step 3 (addition): uncommenting this no-op modification makes this work\r\n // admin.modifyTable(currentDesc);\r\n\r\n // read the data before deleting and recreating (it should exist)\r\n table = conn.getTable(tableName);\r\n Result resultGood = table.get(new Get(row));\r\n System.out.println(\"should be 'value': \" + Bytes.toStringBinary(resultGood.value()));\r\n table.close();\r\n\r\n // step 3\r\n admin.deleteColumnFamily(tableName, family2);\r\n\r\n // step 4\r\n admin.addColumnFamily(tableName, familyDesc2);\r\n\r\n /*\r\n // also occurs with two modify tables\r\n TableDescriptor removedDesc =\r\n TableDescriptorBuilder.newBuilder(tableDesc).removeColumnFamily(family2).build();\r\n // step 3\r\n admin.modifyTable(removedDesc);\r\n // step 4\r\n admin.modifyTable(tableDesc);\r\n */\r\n\r\n // read the data after deleting and recreating (it should exist)\r\n table = conn.getTable(tableName);\r\n Result resultBad = table.get(new Get(row));\r\n System.out.println(\"should be null: \" + Bytes.toStringBinary(resultBad.value()));\r\n table.close();\r\n\r\n // cleanup\r\n admin.disableTable(tableName);\r\n admin.deleteTable(tableName);\r\n} catch (IOException e) {\r\n System.err.println(e);\r\n}\r\n{code}\r\n\r\nThe output looks like:\r\n\r\n{code:none}\r\nshould be 'value': value\r\nshould be null: value\r\n{code}\r\n\r\nIf the no-op {{modifyTable}} is uncommented, the output is:\r\n\r\n{code:none}\r\nshould be 'value': value\r\nshould be null: null\r\n{code}\r\n\r\nI've (tried to...?) attached a project that should compile and run against a {{localhost}} HBase. (I was running it in IntelliJ.)","from":"developer"},{"body":"Maybe related to HBASE-21596.","from":"developer"},{"body":"{quote}Please try imagine your CF have billions of cells in production env, it's an unacceptable to tomb every cells just after calling Admin#deleteColumnFamily.\r\n{quote}\r\nMy understanding is CFs are physically different files/directories on HDFS and therefore a delete column family should be a simple HDFS atomic delete or move/rename operation; tombing cells, I would think, should not be necessary - is this not the case with HBase's implementation?\r\n\r\nI will note that Google Bigtable [explicitly calls out| [https://cloud.google.com/bigtable/docs/managing-tables#deleting_column_families]] data from deleted column families are unrecoverable.\r\n\r\nDoes this have the potential for data, once protected by a column family scoped ACL, to no longer be secure?","from":"developer"},{"body":"bq. is this not the case with HBase's implementation?\r\nNo, it is not.\r\n\r\nbq. Does this have the potential for data, once protected by a column family scoped ACL, to no longer be secure?\r\nTo be more specific? Can't understand the question.","from":"developer"},{"body":"[~wchevreuil] hmm no, different story IMO.","from":"developer"},{"body":"{{admin.modifyTable(currentDesc);}} this method may delete cf on HDFS as you said, but not {{deleteColumnFamily}} and {{modify***}} this kind of CF levels.\r\n[~tmoschou]","from":"developer"},{"body":"{quote}How do you build the {{ColumnFamilyDescriptor}}?\r\n{quote}\r\n\r\nUsing {{ColumnFamilyDescriptorBuilder.newBuilder(existingCf)}} without changing any config will give the expected result.\r\n\r\n[~reidchan]","from":"developer"},{"body":"Please check source code: {{HMaster#modifyColumn, HMaster#deleteColumn, HMaster#modifyTable}}\r\n\r\nif you find any unreasonable or bug, feel free to attach patches.","from":"developer"},{"body":"I think it's clear that there is _something_ weird going on, because {{admin.modifyTable(currentDesc)}} seems like it should do nothing at all (it is \"modifying\" the table to have the same {{TableDescriptor}}), but:\r\n\r\n* the cells are _not_ deleted when that call doesn't run \r\n* the cells are deleted when that call does run\r\n\r\nDo you agree that this is unreasonable?\r\n\r\nPlease confirm whether this is behaviour is correct or not, before we waste time trying to fix it.\r\n","from":"developer"},{"body":"bq. it should do nothing .. have the same TableDescriptor.. ,the cells are deleted when that call does run.\r\nYes, it is unreasonable.\r\n\r\nCommunity encourages volunteering, but under your context 'waste time' which sounds unwilling and being forced, I'd suggest you just set it aside and let someone have interest to take it.:)\r\n","from":"developer"},{"body":"Thanks.\r\n\r\nIt would only be a waste of time if it wasn't actually a problem (for example, if we spent two days trying to fix the issue, only for our patch to be rejected because the current behaviour is correct), so I wanted to make sure we had confirmation that it was a bug.","from":"developer"},{"body":"The problem happens when region memstore contain entries for the deleted CF and CF is deleted dynamically (without disabling the table). Since we delete the CF from FS first and then reopen the region, during reopen RS will flush the memstore content to FS. So deleted CF store will contain the memstore content for the deleted CF. ","from":"developer"},{"body":"(y)[~pankaj2461], make sense to me, is there an upcoming patch? (expecting)","from":"developer"},{"body":"So actual problem is the stale CF.","from":"developer"},{"body":"Attached the UT patch.","from":"developer"},{"body":"| (x) *{color:red}-1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 0m 0s{color} | {color:blue} Docker mode activated. {color} |\r\n| {color:red}-1{color} | {color:red} patch {color} | {color:red} 0m 8s{color} | {color:red} HBASE-22806 does not apply to master. Rebase required? Wrong Branch? See https://yetus.apache.org/documentation/in-progress/precommit-patchnames for help. {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| JIRA Issue | HBASE-22806 |\r\n| JIRA Patch URL | https://issues.apache.org/jira/secure/attachment/12978032/HBASE-22806_UT.patch |\r\n| Console output | https://builds.apache.org/job/PreCommit-HBASE-Build/806/console |\r\n| Powered by | Apache Yetus 0.9.0 http://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","from":"developer"},{"body":"ping [~pankaj2461], please check the QA warning.","from":"developer"},{"body":"[~reidchan] I just uploaded the UT to reproduce this issue. I'm checking the solution, will upload the patch soon.","from":"developer"},{"body":"[~reidchan] Have raised the PR, kindly review.","from":"developer"},{"body":"Pushed on branch-2.1+. Probably not pertinent on branch-1. Thanks for the patch [~pankaj2461].","from":"developer"},{"body":"Issue is applicable to branch-1 also but fix will be different as BulkReOpen reopens the region asynchronously unlike GeneralBulkAssigner,\r\n\r\n{code}\r\n\r\n/**\r\n * Reopen the regions asynchronously, so always returns true immediately.\r\n * @return true\r\n */\r\n @Override\r\n protected boolean waitUntilDone(long timeout) {\r\n return true;\r\n }\r\n\r\n{code}\r\n\r\nTo fix this problem in branch-1, we need to  add region reopen completion wait logic or disable-enable the table.\r\n\r\n[~apurtell]  Please provide your opinion.","from":"developer"},{"body":"bq. To fix this problem in branch-1, we need to add region reopen completion wait logic or disable-enable the table.\r\n\r\nNeither option sounds great. Adding the wait-for-completion seems less disruptive. We can try that. [~pankaj2461]","from":"developer"}],"created":"2019-08-07T00:56:22.000+0000","description":"While deleting the CF dynamically (without disabling the table), CF dirs are not cleared from FS when region memstore contain entries for that CF. \r\n\r\nSince we delete the CF from FS first and then reopen the region, during reopen RS will flush the memstore content to FS. So deleted CF store will contain the memstore content for the deleted CF. \r\n\r\nSo adding back same CF will have old entries.","issue_id":"13249233","key":"HBASE-22806","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2019-08-25T02:52:12.000+0000","role":"fixed_distractor","summary":"Deleted CF are not cleared if memstore contain entries"} {"case_id":"13262997","cluster":"DISTRACTOR-HBASE-23185","comments":[{"body":"Nice fix [~lineyshinya]. I merged to tip of branch-1. [~apurtell] Shout if you want it to go back to other branches.","created":"2019-10-21T16:45:32.822+0000"},{"body":"[~stack] Shout! I think the 1.4 that just went out is the last but there might be more 1.3s. So if not too much trouble...","created":"2019-10-22T00:51:01.205+0000"},{"body":"Thank you for reviewing and merging! Will this be released in 1.5.1? May I create a backport PR for branch-1.3 and branch-1.4 (I think I miss 1.4.11, but I don't know branch-1.4 will be 1.4.12 or EOL after 1.4.11?)?\r\n\r\nDo I need to create a ticket for backporting?","created":"2019-10-22T03:12:21.346+0000"},{"body":"Reopening so we can backport.\r\n\r\n[~lineyshinya] mind making patches for older branches ? I tried backporting the commit on branch-1 but rejects.\r\n\r\nIt looks like tip of branch-1 is where 1.5 was cut from so whatever is next off the tip of branch-1 -- whether 1.5.1 or 1.6.0, it will get this patch.\r\n\r\nNo need to open a new issue. I reopened this one as host for the backports. Thansk.","created":"2019-10-22T05:00:07.975+0000"},{"body":"Thank you for reopening and detailed info. Yes, I'll create a patch for branch-1.3 and branch-1.4 soon.","created":"2019-10-22T06:10:59.007+0000"},{"body":"Sorry, seems my change breaks some tests.","created":"2019-10-23T09:46:36.847+0000"},{"body":"I reverted this from branch-1 because it's causing TestMasterNoCluster and TestCatalogJanitor to fail there.","created":"2019-10-26T05:05:34.392+0000"},{"body":"[~lineyshinya] You see Sean's message above? Saw it because I tripped over https://github.com/apache/hbase/pull/748/files. Thanks. Once it goes in successfully here I can do the backports. Thanks.","created":"2019-10-30T19:56:01.481+0000"},{"body":"#748 should work fine because it contains test fix: [https://github.com/apache/hbase/pull/748/commits/45c7e3588f6c503d8f95f2cf6387549ee380da0b]\r\n\r\n \r\n\r\nCould you confirm it? I run hbase-server tests and finished successfully.","created":"2019-10-31T03:44:19.183+0000"},{"body":"[https://github.com/apache/hbase/pull/748#issuecomment-547403969]\r\n\r\n \r\n\r\nThis also contains hbase-server tests passing.","created":"2019-10-31T03:45:17.243+0000"},{"body":"I merged the commit. WIll check nightly to see if the two tests Sean identified start failing again. Thanks for the fix [~lineyshinya]","created":"2019-10-31T22:27:25.071+0000"},{"body":"Thank you too all! I hope to see all green and to have a good night tonight ;)","created":"2019-11-01T07:54:26.991+0000"},{"body":"Checked build last night. Looks like both of the tests [~busbey] referenced passed:\r\n\r\nhttps://builds.apache.org/view/H-L/view/HBase/job/HBase%20Nightly/job/branch-1/1124/testReport/org.apache.hadoop.hbase.master/TestMasterNoCluster/\r\n\r\nhttps://builds.apache.org/view/H-L/view/HBase/job/HBase%20Nightly/job/branch-1/1124/testReport/org.apache.hadoop.hbase.master/TestCatalogJanitor/\r\n\r\nLet me try backports.\r\n\r\n","created":"2019-11-01T16:37:23.915+0000"},{"body":"Re-resolving. Merged branch-1 and branch-1.3 and branch-1.4 patches. Thanks for patches [~lineyshinya]","created":"2019-11-01T19:25:46.861+0000"},{"body":"Ah, you already merged #745 and #743... Sorry, but I didn't yet add the fixing the test code commit yet...\r\n\r\nSo we need revert them or add a test fix commit as soon as possible.\r\n\r\nI'll create a test fix PR for these branches today, but I need several hours because I'm now out.\r\n\r\nIf you need to fix these branches ASAP, please revert them once...","created":"2019-11-02T06:12:40.839+0000"},{"body":"Test code fix for branch-1.3 [https://github.com/apache/hbase/pull/786] [~stack]\r\n\r\nIf you prefer to revert once and retry it with test code fix let me know please and revert the backport commit. I'll update #786 like #748.\r\n\r\n \r\n\r\nI'm not sure this commit is necessary for branch-1.4, I'll wait for the CI report here","created":"2019-11-02T06:36:29.764+0000"},{"body":"Let me reopen. Let me then apply your test code changes. Thank you for being on top of this [~lineyshinya].","created":"2019-11-02T16:09:25.941+0000"},{"body":"Created fixing PR for branch-1.4: [https://github.com/apache/hbase/pull/788] because it seems branch-1.4 will failed without this.\r\n\r\nbranch-1.3: [https://github.com/apache/hbase/pull/786]\r\n\r\n \r\n\r\nThank you for helping me to fix this!","created":"2019-11-02T17:38:46.735+0000"},{"body":"Thank you [~lineyshinya]. I merged both patches. Lets leave this issue open till we see that both branch-1.3 and branch-1.4 pass for the above mentioned test failures.","created":"2019-11-02T18:27:38.178+0000"},{"body":"Took a look at recent 1.4 nightly.\r\n\r\nhttps://builds.apache.org/view/H-L/view/HBase/job/HBase%20Nightly/job/branch-1.4/1076/testReport/org.apache.hadoop.hbase.master/TestMasterNoCluster/\r\n\r\n1.3 nightly\r\n\r\nhttps://builds.apache.org/view/H-L/view/HBase/job/HBase%20Nightly/job/branch-1.3/1027/testReport/org.apache.hadoop.hbase.master/TestMasterNoCluster/\r\n\r\nThese seems good.\r\n\r\nTestCatalogJanitor might be still sick... It fails here on last 1.3 nightly https://builds.apache.org/job/HBase%20Nightly/job/branch-1.4/1075/ \r\n\r\nWill see how it does in next run. Leaving open in meantime.\r\n","created":"2019-11-04T17:13:00.606+0000"},{"body":"TestCatalogJanitor is in the flakies list still. ... it is added as a flakie when 1.3 runs ... https://builds.apache.org/job/HBase%20Nightly/job/branch-1.3/1029/consoleFull ... but no mention here https://builds.apache.org/view/H-L/view/HBase/job/HBase-Find-Flaky-Tests/job/branch-1.3/lastSuccessfulBuild/artifact/dashboard.html which is odd. The test is still in the code base. Will let it percolate another while then will dig in.\r\n\r\n","created":"2019-11-07T23:44:12.964+0000"},{"body":"If it's not in that list at all it probably passed every time. Check the Jenkins test summary for the job to confirm","created":"2019-11-07T23:48:21.143+0000"},{"body":"Yeah all passes\r\n\r\nhttps://builds.apache.org/job/HBase-Flaky-Tests/job/branch-1.3/test_results_analyzer/","created":"2019-11-07T23:49:52.109+0000"},{"body":"Thats a nice tool. Thanks for intercession [~busbey] and pointer.","created":"2019-11-07T23:53:37.403+0000"},{"body":"Resolving as done. Thanks for the patch [~lineyshinya]","created":"2019-11-07T23:54:24.312+0000"}],"conversations":[{"body":"When we analyzed the performance of our hbase application with many puts, we found that Configuration methods use many CPU resources:\r\n\r\n!Screenshot from 2019-10-18 12-38-14.png|width=460,height=205!\r\n\r\nAs you can see, getTable().put() is calling Configuration methods which cause regex or synchronization by Hashtable.\r\n\r\nThis should not happen in 0.99.2 because https://issues.apache.org/jira/browse/HBASE-12128 addressed such an issue.\r\n However, it's reproducing nowadays by bugs or leakages after many code evoluations between 0.9x and 1.x.\r\n # [https://github.com/apache/hbase/blob/dd9eadb00f9dcd071a246482a11dfc7d63845f00/hbase-client/src/main/java/org/apache/hadoop/hbase/client/HTable.java#L369-L374]\r\n ** finishSetup is called every new HTable() e.g. every con.getTable()\r\n ** So getInt is called everytime and it does regex\r\n # [https://github.com/apache/hbase/blob/dd9eadb00f9dcd071a246482a11dfc7d63845f00/hbase-client/src/main/java/org/apache/hadoop/hbase/client/BufferedMutatorImpl.java#L115]\r\n ** BufferedMutatorImpl is created every first put for HTable e.g. con.getTable().put()\r\n ** Create ConnectionConf every time in BufferedMutatorImpl constructor\r\n ** ConnectionConf gets config value in the constructor\r\n # [https://github.com/apache/hbase/blob/dd9eadb00f9dcd071a246482a11dfc7d63845f00/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncProcess.java#L326]\r\n ** AsyncProcess is created in BufferedMutatorImpl constructor, so new AsyncProcess is created by con.getTable().put()\r\n ** AsyncProcess parse many configurations\r\n\r\nSo, con.getTable().put() is heavy operation for CPU because of getting config value.\r\n\r\n \r\n\r\nWith in-house patch for this issue, we observed about 10% improvement on max-throughput (e.g. CPU usage) at client-side:\r\n\r\n!Screenshot from 2019-10-18 13-03-24.png|width=508,height=223!\r\n\r\n \r\n\r\nSeems branch-2 is not affected because client implementation has been changed dramatically.\r\n  ","from":"reporter","subject":"High cpu usage because getTable()#put() gets config value every time"},{"body":"Nice fix [~lineyshinya]. I merged to tip of branch-1. [~apurtell] Shout if you want it to go back to other branches.","from":"developer"},{"body":"[~stack] Shout! I think the 1.4 that just went out is the last but there might be more 1.3s. So if not too much trouble...","from":"developer"},{"body":"Thank you for reviewing and merging! Will this be released in 1.5.1? May I create a backport PR for branch-1.3 and branch-1.4 (I think I miss 1.4.11, but I don't know branch-1.4 will be 1.4.12 or EOL after 1.4.11?)?\r\n\r\nDo I need to create a ticket for backporting?","from":"developer"},{"body":"Reopening so we can backport.\r\n\r\n[~lineyshinya] mind making patches for older branches ? I tried backporting the commit on branch-1 but rejects.\r\n\r\nIt looks like tip of branch-1 is where 1.5 was cut from so whatever is next off the tip of branch-1 -- whether 1.5.1 or 1.6.0, it will get this patch.\r\n\r\nNo need to open a new issue. I reopened this one as host for the backports. Thansk.","from":"developer"},{"body":"Thank you for reopening and detailed info. Yes, I'll create a patch for branch-1.3 and branch-1.4 soon.","from":"developer"},{"body":"Sorry, seems my change breaks some tests.","from":"developer"},{"body":"I reverted this from branch-1 because it's causing TestMasterNoCluster and TestCatalogJanitor to fail there.","from":"developer"},{"body":"[~lineyshinya] You see Sean's message above? Saw it because I tripped over https://github.com/apache/hbase/pull/748/files. Thanks. Once it goes in successfully here I can do the backports. Thanks.","from":"developer"},{"body":"#748 should work fine because it contains test fix: [https://github.com/apache/hbase/pull/748/commits/45c7e3588f6c503d8f95f2cf6387549ee380da0b]\r\n\r\n \r\n\r\nCould you confirm it? I run hbase-server tests and finished successfully.","from":"developer"},{"body":"[https://github.com/apache/hbase/pull/748#issuecomment-547403969]\r\n\r\n \r\n\r\nThis also contains hbase-server tests passing.","from":"developer"},{"body":"I merged the commit. WIll check nightly to see if the two tests Sean identified start failing again. Thanks for the fix [~lineyshinya]","from":"developer"},{"body":"Thank you too all! I hope to see all green and to have a good night tonight ;)","from":"developer"},{"body":"Checked build last night. Looks like both of the tests [~busbey] referenced passed:\r\n\r\nhttps://builds.apache.org/view/H-L/view/HBase/job/HBase%20Nightly/job/branch-1/1124/testReport/org.apache.hadoop.hbase.master/TestMasterNoCluster/\r\n\r\nhttps://builds.apache.org/view/H-L/view/HBase/job/HBase%20Nightly/job/branch-1/1124/testReport/org.apache.hadoop.hbase.master/TestCatalogJanitor/\r\n\r\nLet me try backports.\r\n\r\n","from":"developer"},{"body":"Re-resolving. Merged branch-1 and branch-1.3 and branch-1.4 patches. Thanks for patches [~lineyshinya]","from":"developer"},{"body":"Ah, you already merged #745 and #743... Sorry, but I didn't yet add the fixing the test code commit yet...\r\n\r\nSo we need revert them or add a test fix commit as soon as possible.\r\n\r\nI'll create a test fix PR for these branches today, but I need several hours because I'm now out.\r\n\r\nIf you need to fix these branches ASAP, please revert them once...","from":"developer"},{"body":"Test code fix for branch-1.3 [https://github.com/apache/hbase/pull/786] [~stack]\r\n\r\nIf you prefer to revert once and retry it with test code fix let me know please and revert the backport commit. I'll update #786 like #748.\r\n\r\n \r\n\r\nI'm not sure this commit is necessary for branch-1.4, I'll wait for the CI report here","from":"developer"},{"body":"Let me reopen. Let me then apply your test code changes. Thank you for being on top of this [~lineyshinya].","from":"developer"},{"body":"Created fixing PR for branch-1.4: [https://github.com/apache/hbase/pull/788] because it seems branch-1.4 will failed without this.\r\n\r\nbranch-1.3: [https://github.com/apache/hbase/pull/786]\r\n\r\n \r\n\r\nThank you for helping me to fix this!","from":"developer"},{"body":"Thank you [~lineyshinya]. I merged both patches. Lets leave this issue open till we see that both branch-1.3 and branch-1.4 pass for the above mentioned test failures.","from":"developer"},{"body":"Took a look at recent 1.4 nightly.\r\n\r\nhttps://builds.apache.org/view/H-L/view/HBase/job/HBase%20Nightly/job/branch-1.4/1076/testReport/org.apache.hadoop.hbase.master/TestMasterNoCluster/\r\n\r\n1.3 nightly\r\n\r\nhttps://builds.apache.org/view/H-L/view/HBase/job/HBase%20Nightly/job/branch-1.3/1027/testReport/org.apache.hadoop.hbase.master/TestMasterNoCluster/\r\n\r\nThese seems good.\r\n\r\nTestCatalogJanitor might be still sick... It fails here on last 1.3 nightly https://builds.apache.org/job/HBase%20Nightly/job/branch-1.4/1075/ \r\n\r\nWill see how it does in next run. Leaving open in meantime.\r\n","from":"developer"},{"body":"TestCatalogJanitor is in the flakies list still. ... it is added as a flakie when 1.3 runs ... https://builds.apache.org/job/HBase%20Nightly/job/branch-1.3/1029/consoleFull ... but no mention here https://builds.apache.org/view/H-L/view/HBase/job/HBase-Find-Flaky-Tests/job/branch-1.3/lastSuccessfulBuild/artifact/dashboard.html which is odd. The test is still in the code base. Will let it percolate another while then will dig in.\r\n\r\n","from":"developer"},{"body":"If it's not in that list at all it probably passed every time. Check the Jenkins test summary for the job to confirm","from":"developer"},{"body":"Yeah all passes\r\n\r\nhttps://builds.apache.org/job/HBase-Flaky-Tests/job/branch-1.3/test_results_analyzer/","from":"developer"},{"body":"Thats a nice tool. Thanks for intercession [~busbey] and pointer.","from":"developer"},{"body":"Resolving as done. Thanks for the patch [~lineyshinya]","from":"developer"}],"created":"2019-10-18T04:04:26.000+0000","description":"When we analyzed the performance of our hbase application with many puts, we found that Configuration methods use many CPU resources:\r\n\r\n!Screenshot from 2019-10-18 12-38-14.png|width=460,height=205!\r\n\r\nAs you can see, getTable().put() is calling Configuration methods which cause regex or synchronization by Hashtable.\r\n\r\nThis should not happen in 0.99.2 because https://issues.apache.org/jira/browse/HBASE-12128 addressed such an issue.\r\n However, it's reproducing nowadays by bugs or leakages after many code evoluations between 0.9x and 1.x.\r\n # [https://github.com/apache/hbase/blob/dd9eadb00f9dcd071a246482a11dfc7d63845f00/hbase-client/src/main/java/org/apache/hadoop/hbase/client/HTable.java#L369-L374]\r\n ** finishSetup is called every new HTable() e.g. every con.getTable()\r\n ** So getInt is called everytime and it does regex\r\n # [https://github.com/apache/hbase/blob/dd9eadb00f9dcd071a246482a11dfc7d63845f00/hbase-client/src/main/java/org/apache/hadoop/hbase/client/BufferedMutatorImpl.java#L115]\r\n ** BufferedMutatorImpl is created every first put for HTable e.g. con.getTable().put()\r\n ** Create ConnectionConf every time in BufferedMutatorImpl constructor\r\n ** ConnectionConf gets config value in the constructor\r\n # [https://github.com/apache/hbase/blob/dd9eadb00f9dcd071a246482a11dfc7d63845f00/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncProcess.java#L326]\r\n ** AsyncProcess is created in BufferedMutatorImpl constructor, so new AsyncProcess is created by con.getTable().put()\r\n ** AsyncProcess parse many configurations\r\n\r\nSo, con.getTable().put() is heavy operation for CPU because of getting config value.\r\n\r\n \r\n\r\nWith in-house patch for this issue, we observed about 10% improvement on max-throughput (e.g. CPU usage) at client-side:\r\n\r\n!Screenshot from 2019-10-18 13-03-24.png|width=508,height=223!\r\n\r\n \r\n\r\nSeems branch-2 is not affected because client implementation has been changed dramatically.\r\n  ","issue_id":"13262997","key":"HBASE-23185","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2019-11-07T23:54:24.000+0000","role":"fixed_distractor","summary":"High cpu usage because getTable()#put() gets config value every time"} {"case_id":"13264003","cluster":"DISTRACTOR-HBASE-23205","comments":[{"body":"In this PR([https://github.com/apache/hbase/pull/749]), I tried to update log position only when WAL rolled, and removed log position updates in reader side to resolve concurrency issue by case 2).\r\n\r\nI added some tests to check case 1) and case 3)\r\n * TestReplicationSource.testSetLogPositionForWALCurrentlyReadingWhenLogsRolled for case 1) \r\n * TestWALEntryStream.testReplicationSourceWALReaderThreadWithFilter for case 3)\r\n\r\n \r\n\r\nI think all flaws i mentioned can be fixed by this patch.","created":"2019-10-24T09:32:55.885+0000"},{"body":"Thanks for the contribution [~Jeongdae Kim]! I had merged it into branch-1. Do you need this backported to branch-1.4 as well?","created":"2020-01-02T12:01:58.046+0000"},{"body":">  Do you need this backported to branch-1.4 as well?\r\n\r\nI for one am eagerly awaiting 1.4.13 and hoping it includes this fix. I'm seeing the aborting RS prob with the \"Failed to write replication wal position\" error when running 1.4.12.","created":"2020-01-02T20:23:57.631+0000"},{"body":"[~wchevreuil] Thank you for supporting me for a quite long time. :)\r\n\r\n \r\n{quote}Do you need this backported to branch-1.4 as well?\r\n{quote}\r\n \r\n\r\nYes, let me prepare a patch for branch-1.4. but I'm not sure that I should take care of 1.5.x, because of no more maintenance releases for 1.5.x according to this [https://lists.apache.org/thread.html/cc66ce601a0bcd4fc809139f3f79f63607ee9f5bb4dcffcb1a92b2e6@%3Cdev.hbase.apache.org%3E]\r\n\r\n(and there is no branch for 1.5)","created":"2020-01-03T07:29:08.536+0000"},{"body":"[~Jeongdae Kim], [~whitney13], let me cherry-pick this branch-1 commit into branch-1.4.\r\n\r\n \r\n{quote}Yes, let me prepare a patch for branch-1.4. but I'm not sure that I should take care of 1.5.x, because of no more maintenance releases for 1.5.x according to this\r\n{quote}\r\n There's still plans for 1.5 maintenance releases, there's no branch-1.5 because that's the current minor version, so everything that's on branch-1 is currently targeted to be present on the next 1.5.x release  (1.5.1, in this case). If folks decide to raise the minor version to \"1.6\", then a \"branch-1.5\" would be created, and changes on branch-1 would go to \"1.6.x\" releases. \r\n\r\nLet me work on the cherry-pick for \"1.4\". If it applies cleanly, no need for an additional patch, [~Jeongdae Kim] . ","created":"2020-01-03T12:13:24.937+0000"},{"body":"Cherrypick didn't apply cleanly, it resulted in several conflicts. [~Jeongdae Kim], mind put on a 1.4 patch here?","created":"2020-01-03T12:30:57.233+0000"},{"body":"Thanks for explanation about releases and branches! I already made the PR for 1.4 branch.\r\n\r\n[https://github.com/apache/hbase/pull/980]\r\n\r\nI can made a patch file if you need it.\r\n\r\nThanks!","created":"2020-01-06T08:30:03.634+0000"},{"body":"Thanks again for the work here, [~Jeongdae Kim]! I had merged PR #980 into branch-1.4, so it will be included on the next 1.4 maintenance release (1.4.13). Marking this Jira as resolved.","created":"2020-01-09T10:12:44.571+0000"}],"conversations":[{"body":"We observed a lot of old WALs were not removed from archives and their corresponding replication queues, while testing with 1.4.10.\r\n stacked old WALs are empty or have no entries to be replicated (not in replication table_cfs)\r\n  \r\n As described in HBASE-22784, if no entries to be replicated are appended to WALs, log position will never be updated. As a consequence, all WALs won’t be removed. this issue happened since HBASE-15995.\r\n  \r\n I think old WALs would not be stacked with HBASE-22784. but, it still have something to be fixed as below\r\n\r\n  case 1) Log position could be updated wrongly, when log rolled, because lastWalPath of batches might not point to WAL currently being read.\r\n * For example,  after last entry added in a batch were read from P1 position in the WAL W1\r\n and then WAL rolled, and reader read until it reaches the end of old wals and continue reading entries from new WAL W2, and then it reached batch size. current read position for W2 is P2. In this case, the batch being passed to a shipper have walPath W1 and P2, so shipper will try to update position P2 for W1. it may result in data inconsistency in recovery case or update failure to zookeeper (znode could not exist by previous log position updates, i guess this case is the same case as HBASE-23169 ?)\r\n\r\n \r\n\r\n  case 2) Log position could be not updated or updated to wrong position by pendingShipment flag introduced from HBASE-22784\r\n * In shipper thread, it would not be guaranteed to update log position always, by setting pendingShipment to false.\r\n If  reader set the flag to true, right after shipper set it to false during {color:#24292e}updateLogPosition(), shipper won’t update log position.{color}\r\n On the other hand, while reader read filtered entries, If shipper set to false reader will update log position to current read position. it may lose data in recovery case.\r\n\r\n \r\n\r\n  case 3) A lot of log position updates could be happened, when most of WAL entries are filtered by TableCfWALEntryFilter.\r\n * I think it would be better to reduce the number of log updates in that case, because\r\n ## zookeeper writes are more expensive operations than reads.(since writes involve synchronizing the state of all servers),\r\n ## even if read position was not updated, it would be harmless because all entries will be filtered out again in recovery process.\r\n * It would be enough to update log position only when wal rolled in that case. (to cleanup old wals)\r\n\r\n \r\n\r\n-In addition, During this work, i found a minor bug which is updating replication buffer size wrongly by decreasing total buffer size with the size of bulk loaded files.-\r\n -I’d like to fix it, if it’s ok.-\r\n\r\nI removed the changes above and made a separate jira : HBASE-23254","from":"reporter","subject":"Correctly update the position of WALs currently being replicated."},{"body":"In this PR([https://github.com/apache/hbase/pull/749]), I tried to update log position only when WAL rolled, and removed log position updates in reader side to resolve concurrency issue by case 2).\r\n\r\nI added some tests to check case 1) and case 3)\r\n * TestReplicationSource.testSetLogPositionForWALCurrentlyReadingWhenLogsRolled for case 1) \r\n * TestWALEntryStream.testReplicationSourceWALReaderThreadWithFilter for case 3)\r\n\r\n \r\n\r\nI think all flaws i mentioned can be fixed by this patch.","from":"developer"},{"body":"Thanks for the contribution [~Jeongdae Kim]! I had merged it into branch-1. Do you need this backported to branch-1.4 as well?","from":"developer"},{"body":">  Do you need this backported to branch-1.4 as well?\r\n\r\nI for one am eagerly awaiting 1.4.13 and hoping it includes this fix. I'm seeing the aborting RS prob with the \"Failed to write replication wal position\" error when running 1.4.12.","from":"developer"},{"body":"[~wchevreuil] Thank you for supporting me for a quite long time. :)\r\n\r\n \r\n{quote}Do you need this backported to branch-1.4 as well?\r\n{quote}\r\n \r\n\r\nYes, let me prepare a patch for branch-1.4. but I'm not sure that I should take care of 1.5.x, because of no more maintenance releases for 1.5.x according to this [https://lists.apache.org/thread.html/cc66ce601a0bcd4fc809139f3f79f63607ee9f5bb4dcffcb1a92b2e6@%3Cdev.hbase.apache.org%3E]\r\n\r\n(and there is no branch for 1.5)","from":"developer"},{"body":"[~Jeongdae Kim], [~whitney13], let me cherry-pick this branch-1 commit into branch-1.4.\r\n\r\n \r\n{quote}Yes, let me prepare a patch for branch-1.4. but I'm not sure that I should take care of 1.5.x, because of no more maintenance releases for 1.5.x according to this\r\n{quote}\r\n There's still plans for 1.5 maintenance releases, there's no branch-1.5 because that's the current minor version, so everything that's on branch-1 is currently targeted to be present on the next 1.5.x release  (1.5.1, in this case). If folks decide to raise the minor version to \"1.6\", then a \"branch-1.5\" would be created, and changes on branch-1 would go to \"1.6.x\" releases. \r\n\r\nLet me work on the cherry-pick for \"1.4\". If it applies cleanly, no need for an additional patch, [~Jeongdae Kim] . ","from":"developer"},{"body":"Cherrypick didn't apply cleanly, it resulted in several conflicts. [~Jeongdae Kim], mind put on a 1.4 patch here?","from":"developer"},{"body":"Thanks for explanation about releases and branches! I already made the PR for 1.4 branch.\r\n\r\n[https://github.com/apache/hbase/pull/980]\r\n\r\nI can made a patch file if you need it.\r\n\r\nThanks!","from":"developer"},{"body":"Thanks again for the work here, [~Jeongdae Kim]! I had merged PR #980 into branch-1.4, so it will be included on the next 1.4 maintenance release (1.4.13). Marking this Jira as resolved.","from":"developer"}],"created":"2019-10-23T11:37:44.000+0000","description":"We observed a lot of old WALs were not removed from archives and their corresponding replication queues, while testing with 1.4.10.\r\n stacked old WALs are empty or have no entries to be replicated (not in replication table_cfs)\r\n  \r\n As described in HBASE-22784, if no entries to be replicated are appended to WALs, log position will never be updated. As a consequence, all WALs won’t be removed. this issue happened since HBASE-15995.\r\n  \r\n I think old WALs would not be stacked with HBASE-22784. but, it still have something to be fixed as below\r\n\r\n  case 1) Log position could be updated wrongly, when log rolled, because lastWalPath of batches might not point to WAL currently being read.\r\n * For example,  after last entry added in a batch were read from P1 position in the WAL W1\r\n and then WAL rolled, and reader read until it reaches the end of old wals and continue reading entries from new WAL W2, and then it reached batch size. current read position for W2 is P2. In this case, the batch being passed to a shipper have walPath W1 and P2, so shipper will try to update position P2 for W1. it may result in data inconsistency in recovery case or update failure to zookeeper (znode could not exist by previous log position updates, i guess this case is the same case as HBASE-23169 ?)\r\n\r\n \r\n\r\n  case 2) Log position could be not updated or updated to wrong position by pendingShipment flag introduced from HBASE-22784\r\n * In shipper thread, it would not be guaranteed to update log position always, by setting pendingShipment to false.\r\n If  reader set the flag to true, right after shipper set it to false during {color:#24292e}updateLogPosition(), shipper won’t update log position.{color}\r\n On the other hand, while reader read filtered entries, If shipper set to false reader will update log position to current read position. it may lose data in recovery case.\r\n\r\n \r\n\r\n  case 3) A lot of log position updates could be happened, when most of WAL entries are filtered by TableCfWALEntryFilter.\r\n * I think it would be better to reduce the number of log updates in that case, because\r\n ## zookeeper writes are more expensive operations than reads.(since writes involve synchronizing the state of all servers),\r\n ## even if read position was not updated, it would be harmless because all entries will be filtered out again in recovery process.\r\n * It would be enough to update log position only when wal rolled in that case. (to cleanup old wals)\r\n\r\n \r\n\r\n-In addition, During this work, i found a minor bug which is updating replication buffer size wrongly by decreasing total buffer size with the size of bulk loaded files.-\r\n -I’d like to fix it, if it’s ok.-\r\n\r\nI removed the changes above and made a separate jira : HBASE-23254","issue_id":"13264003","key":"HBASE-23205","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-01-09T10:12:52.000+0000","role":"fixed_distractor","summary":"Correctly update the position of WALs currently being replicated."} {"case_id":"13277330","cluster":"DISTRACTOR-HBASE-23633","comments":[{"body":"What I want is an indication in Master log that a Region is not opening because hfiles are corrupt or we have dangling references. Currently it just says fail but not why (smile).","created":"2020-01-03T18:45:56.471+0000"},{"body":"I also observed this problem during test, many regions *FAILED* to open due to CorruptHFileException. \r\n{noformat}\r\n2020-01-29 07:07:13,911 | INFO | RS_OPEN_REGION-RS-IP:RS-PORT-2 | Validating hfile at hdfs://cluster/hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793 for inclusion in store family region usertable01,user35466,1580220595485.a2f0e8b46399ce55e864d4ee7311c845. | org.apache.hadoop.hbase.regionserver.HStore.assertBulkLoadHFileOk(HStore.java:730)\r\n2020-01-29 07:07:13,930 | ERROR | RS_OPEN_REGION-RS-IP:RS-PORT-2 | Failed open of region=usertable01,user35466,1580220595485.a2f0e8b46399ce55e864d4ee7311c845., starting to roll back the global memstore size. | org.apache.hadoop.hbase.regionserver.handler.OpenRegionHandler.openRegion(OpenRegionHandler.java:386)\r\norg.apache.hadoop.hbase.io.hfile.CorruptHFileException: Problem reading HFile Trailer from file hdfs://cluster/hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793\r\n\tat org.apache.hadoop.hbase.io.hfile.HFile.openReader(HFile.java:503)\r\n\tat org.apache.hadoop.hbase.io.hfile.HFile.createReader(HFile.java:562)\r\n\tat org.apache.hadoop.hbase.regionserver.HStore.assertBulkLoadHFileOk(HStore.java:732)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.loadRecoveredHFilesIfAny(HRegion.java:4905)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.initializeRegionInternals(HRegion.java:863)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.initialize(HRegion.java:824)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.openHRegion(HRegion.java:7023)\r\n{noformat}\r\n\r\n\r\nAfter digging more into the log, observed this problem occured when \"split-log-closeStream\" thread was splitting WAL into hfile and Region Server abort due to some region. So the \"split-log-closeStream\" thread was interrupted and left the recovered hfile in an intermediate state.\r\n\r\n{noformat}\r\n2020-01-28 23:01:04,962 | WARN | RS_LOG_REPLAY_OPS-8-5-179-5:RS-PORT-0 | log splitting of WALs/RS-IP,RS-PORT,1580220469213-splitting/RS-IP%2CRS-PORT%2C1580220469213.1580222580793 interrupted, resigning | org.apache.hadoop.hbase.regionserver.SplitLogWorker$1.exec(SplitLogWorker.java:111)\r\njava.io.InterruptedIOException\r\n\tat org.apache.hadoop.hbase.wal.BoundedRecoveredHFilesOutputSink.writeRemainingEntryBuffers(BoundedRecoveredHFilesOutputSink.java:186)\r\n\tat org.apache.hadoop.hbase.wal.BoundedRecoveredHFilesOutputSink.close(BoundedRecoveredHFilesOutputSink.java:155)\r\n\tat org.apache.hadoop.hbase.wal.WALSplitter.splitLogFile(WALSplitter.java:404)\r\n\tat org.apache.hadoop.hbase.wal.WALSplitter.splitLogFile(WALSplitter.java:225)\r\n\tat org.apache.hadoop.hbase.regionserver.SplitLogWorker$1.exec(SplitLogWorker.java:105)\r\n\tat org.apache.hadoop.hbase.regionserver.handler.WALSplitterHandler.process(WALSplitterHandler.java:72)\r\n\tat org.apache.hadoop.hbase.executor.EventHandler.run(EventHandler.java:129)\r\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n\tat java.lang.Thread.run(Thread.java:748)\r\nCaused by: java.lang.InterruptedException\r\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject.reportInterruptAfterWait(AbstractQueuedSynchronizer.java:2014)\r\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject.await(AbstractQueuedSynchronizer.java:2048)\r\n\tat java.util.concurrent.LinkedBlockingQueue.take(LinkedBlockingQueue.java:442)\r\n\tat java.util.concurrent.ExecutorCompletionService.take(ExecutorCompletionService.java:193)\r\n\tat org.apache.hadoop.hbase.wal.BoundedRecoveredHFilesOutputSink.writeRemainingEntryBuffers(BoundedRecoveredHFilesOutputSink.java:179)\r\n\t... 9 more\r\n\r\n{noformat}\r\n\r\nFurther I checked and confirmed from the NN audit log that file was not written completelty and RS went down,\r\n{noformat}\r\n2020-01-28 23:01:04,946 | INFO | IPC Server handler 125 on 25000 | BLOCK* allocate blk_1092127264_18392260, replicas=DN-IP1:DN-PORT, DN-IP2:DN-PORT, DN-IP3:DN-PORT for /hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793 | FSDirWriteFileOp.java:856\r\n----\r\n2020-01-29 00:01:04,956 | INFO | org.apache.hadoop.hdfs.server.namenode.LeaseManager$Monitor@862fb5 | Recovering [Lease. Holder: DFSClient_NONMAPREDUCE_-1098699935_1, pending creates: 21], src=/hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793 | FSNamesystem.java:3344\r\n2020-01-29 00:01:04,957 | WARN | org.apache.hadoop.hdfs.server.namenode.LeaseManager$Monitor@862fb5 | DIR* NameSystem.internalReleaseLease: File /hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793 has not been closed. Lease recovery is in progress. RecoveryId = 18395023 for block blk_1092127264_18392260 | FSNamesystem.java:3470\r\n2020-01-29 00:01:14,504 | INFO | IPC Server handler 0 on 25006 | commitBlockSynchronization(oldBlock=BP-2062589142-192.168.250.11-1574429102552:blk_1092127264_18392260, file=/hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793, newgenerationstamp=18395023, newlength=0, newtargets=[]) successful | FSNamesystem.java:3748\r\n{noformat}\r\n\r\nSince WAL split was interrupted, so HMaster will recover it by resubmitting the WAL split task. So there may not be data loss IMO.\r\n\r\nWe should cleanup such corrupted hfile. What do you think [~zghao] [~stack] sir ?\r\n\r\n","created":"2020-01-29T12:08:46.707+0000"},{"body":"In my test scenario all corrupted hfiles were of zero length. We should check & delete zero length file during recovered hfile bulkload, the same way how it handled while replaying edits,\r\n\r\n{code}\r\n private long loadRecoveredHFilesIfAny(Collection stores) throws IOException {\r\n Path regionDir = getWALRegionDir();\r\n long maxSeqId = -1;\r\n for (HStore store : stores) {\r\n String familyName = store.getColumnFamilyName();\r\n FileStatus[] files =\r\n WALSplitUtil.getRecoveredHFiles(fs.getFileSystem(), regionDir, familyName);\r\n if (files != null && files.length != 0) {\r\n for (FileStatus file : files) {\r\n // Check and delete the zero length file\r\n if (isZeroLengthThenDelete(fs.getFileSystem(), file.getPath())) {\r\n continue;\r\n }\r\n store.assertBulkLoadHFileOk(file.getPath());\r\n{code}","created":"2020-01-30T13:10:50.886+0000"},{"body":"[~pankajkumar] Are you working for this now?","created":"2020-02-02T02:30:31.535+0000"},{"body":"Will raise PR with UT.","created":"2020-02-03T18:49:15.095+0000"},{"body":"[~pankajkumar] As the PR not updated long time, I merged it and open a new issue to add ut.","created":"2020-03-22T08:45:00.949+0000"}],"conversations":[{"body":"Copy the comment from PR review.\r\n\r\n \r\n\r\nIf the file is a corrupt HFile, an exception will be thrown here, which will cause the region to fail to open.\r\nMaybe we can add a new parameter to control whether to skip the exception, similar to recover edits which has a parameter \"hbase.hregion.edits.replay.skip.errors\";\r\n\r\n \r\n\r\nRegions that can't be opened because of detached References or corrupt hfiles are a fact-of-life. We need work on this issue. This will be a new variant on the problem -- i.e. bad recovered hfiles.\r\n\r\nOn adding a config to ignore bad files and just open, thats a bit dangerous as per @infraio .... as it could mean silent data loss.","from":"reporter","subject":"Find a way to handle the corrupt recovered hfiles"},{"body":"What I want is an indication in Master log that a Region is not opening because hfiles are corrupt or we have dangling references. Currently it just says fail but not why (smile).","from":"developer"},{"body":"I also observed this problem during test, many regions *FAILED* to open due to CorruptHFileException. \r\n{noformat}\r\n2020-01-29 07:07:13,911 | INFO | RS_OPEN_REGION-RS-IP:RS-PORT-2 | Validating hfile at hdfs://cluster/hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793 for inclusion in store family region usertable01,user35466,1580220595485.a2f0e8b46399ce55e864d4ee7311c845. | org.apache.hadoop.hbase.regionserver.HStore.assertBulkLoadHFileOk(HStore.java:730)\r\n2020-01-29 07:07:13,930 | ERROR | RS_OPEN_REGION-RS-IP:RS-PORT-2 | Failed open of region=usertable01,user35466,1580220595485.a2f0e8b46399ce55e864d4ee7311c845., starting to roll back the global memstore size. | org.apache.hadoop.hbase.regionserver.handler.OpenRegionHandler.openRegion(OpenRegionHandler.java:386)\r\norg.apache.hadoop.hbase.io.hfile.CorruptHFileException: Problem reading HFile Trailer from file hdfs://cluster/hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793\r\n\tat org.apache.hadoop.hbase.io.hfile.HFile.openReader(HFile.java:503)\r\n\tat org.apache.hadoop.hbase.io.hfile.HFile.createReader(HFile.java:562)\r\n\tat org.apache.hadoop.hbase.regionserver.HStore.assertBulkLoadHFileOk(HStore.java:732)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.loadRecoveredHFilesIfAny(HRegion.java:4905)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.initializeRegionInternals(HRegion.java:863)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.initialize(HRegion.java:824)\r\n\tat org.apache.hadoop.hbase.regionserver.HRegion.openHRegion(HRegion.java:7023)\r\n{noformat}\r\n\r\n\r\nAfter digging more into the log, observed this problem occured when \"split-log-closeStream\" thread was splitting WAL into hfile and Region Server abort due to some region. So the \"split-log-closeStream\" thread was interrupted and left the recovered hfile in an intermediate state.\r\n\r\n{noformat}\r\n2020-01-28 23:01:04,962 | WARN | RS_LOG_REPLAY_OPS-8-5-179-5:RS-PORT-0 | log splitting of WALs/RS-IP,RS-PORT,1580220469213-splitting/RS-IP%2CRS-PORT%2C1580220469213.1580222580793 interrupted, resigning | org.apache.hadoop.hbase.regionserver.SplitLogWorker$1.exec(SplitLogWorker.java:111)\r\njava.io.InterruptedIOException\r\n\tat org.apache.hadoop.hbase.wal.BoundedRecoveredHFilesOutputSink.writeRemainingEntryBuffers(BoundedRecoveredHFilesOutputSink.java:186)\r\n\tat org.apache.hadoop.hbase.wal.BoundedRecoveredHFilesOutputSink.close(BoundedRecoveredHFilesOutputSink.java:155)\r\n\tat org.apache.hadoop.hbase.wal.WALSplitter.splitLogFile(WALSplitter.java:404)\r\n\tat org.apache.hadoop.hbase.wal.WALSplitter.splitLogFile(WALSplitter.java:225)\r\n\tat org.apache.hadoop.hbase.regionserver.SplitLogWorker$1.exec(SplitLogWorker.java:105)\r\n\tat org.apache.hadoop.hbase.regionserver.handler.WALSplitterHandler.process(WALSplitterHandler.java:72)\r\n\tat org.apache.hadoop.hbase.executor.EventHandler.run(EventHandler.java:129)\r\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n\tat java.lang.Thread.run(Thread.java:748)\r\nCaused by: java.lang.InterruptedException\r\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject.reportInterruptAfterWait(AbstractQueuedSynchronizer.java:2014)\r\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject.await(AbstractQueuedSynchronizer.java:2048)\r\n\tat java.util.concurrent.LinkedBlockingQueue.take(LinkedBlockingQueue.java:442)\r\n\tat java.util.concurrent.ExecutorCompletionService.take(ExecutorCompletionService.java:193)\r\n\tat org.apache.hadoop.hbase.wal.BoundedRecoveredHFilesOutputSink.writeRemainingEntryBuffers(BoundedRecoveredHFilesOutputSink.java:179)\r\n\t... 9 more\r\n\r\n{noformat}\r\n\r\nFurther I checked and confirmed from the NN audit log that file was not written completelty and RS went down,\r\n{noformat}\r\n2020-01-28 23:01:04,946 | INFO | IPC Server handler 125 on 25000 | BLOCK* allocate blk_1092127264_18392260, replicas=DN-IP1:DN-PORT, DN-IP2:DN-PORT, DN-IP3:DN-PORT for /hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793 | FSDirWriteFileOp.java:856\r\n----\r\n2020-01-29 00:01:04,956 | INFO | org.apache.hadoop.hdfs.server.namenode.LeaseManager$Monitor@862fb5 | Recovering [Lease. Holder: DFSClient_NONMAPREDUCE_-1098699935_1, pending creates: 21], src=/hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793 | FSNamesystem.java:3344\r\n2020-01-29 00:01:04,957 | WARN | org.apache.hadoop.hdfs.server.namenode.LeaseManager$Monitor@862fb5 | DIR* NameSystem.internalReleaseLease: File /hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793 has not been closed. Lease recovery is in progress. RecoveryId = 18395023 for block blk_1092127264_18392260 | FSNamesystem.java:3470\r\n2020-01-29 00:01:14,504 | INFO | IPC Server handler 0 on 25006 | commitBlockSynchronization(oldBlock=BP-2062589142-192.168.250.11-1574429102552:blk_1092127264_18392260, file=/hbase/data/default/usertable01/a2f0e8b46399ce55e864d4ee7311c845/family/recovered.hfiles/0000000000000000290-RS-IP%2CRS-PORT%2C1580220469213.1580222580793, newgenerationstamp=18395023, newlength=0, newtargets=[]) successful | FSNamesystem.java:3748\r\n{noformat}\r\n\r\nSince WAL split was interrupted, so HMaster will recover it by resubmitting the WAL split task. So there may not be data loss IMO.\r\n\r\nWe should cleanup such corrupted hfile. What do you think [~zghao] [~stack] sir ?\r\n\r\n","from":"developer"},{"body":"In my test scenario all corrupted hfiles were of zero length. We should check & delete zero length file during recovered hfile bulkload, the same way how it handled while replaying edits,\r\n\r\n{code}\r\n private long loadRecoveredHFilesIfAny(Collection stores) throws IOException {\r\n Path regionDir = getWALRegionDir();\r\n long maxSeqId = -1;\r\n for (HStore store : stores) {\r\n String familyName = store.getColumnFamilyName();\r\n FileStatus[] files =\r\n WALSplitUtil.getRecoveredHFiles(fs.getFileSystem(), regionDir, familyName);\r\n if (files != null && files.length != 0) {\r\n for (FileStatus file : files) {\r\n // Check and delete the zero length file\r\n if (isZeroLengthThenDelete(fs.getFileSystem(), file.getPath())) {\r\n continue;\r\n }\r\n store.assertBulkLoadHFileOk(file.getPath());\r\n{code}","from":"developer"},{"body":"[~pankajkumar] Are you working for this now?","from":"developer"},{"body":"Will raise PR with UT.","from":"developer"},{"body":"[~pankajkumar] As the PR not updated long time, I merged it and open a new issue to add ut.","from":"developer"}],"created":"2020-01-03T09:08:17.000+0000","description":"Copy the comment from PR review.\r\n\r\n \r\n\r\nIf the file is a corrupt HFile, an exception will be thrown here, which will cause the region to fail to open.\r\nMaybe we can add a new parameter to control whether to skip the exception, similar to recover edits which has a parameter \"hbase.hregion.edits.replay.skip.errors\";\r\n\r\n \r\n\r\nRegions that can't be opened because of detached References or corrupt hfiles are a fact-of-life. We need work on this issue. This will be a new variant on the problem -- i.e. bad recovered hfiles.\r\n\r\nOn adding a config to ignore bad files and just open, thats a bit dangerous as per @infraio .... as it could mean silent data loss.","issue_id":"13277330","key":"HBASE-23633","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-03-22T15:07:09.000+0000","role":"fixed_distractor","summary":"Find a way to handle the corrupt recovered hfiles"} {"case_id":"13294701","cluster":"DISTRACTOR-HBASE-24074","comments":[{"body":"Which version's hbase? Is there other running peer procedure?","created":"2020-03-28T10:33:12.602+0000"},{"body":"It's HBase-2.2.3. It looks strange because we are accessing oldsources in  a synchronized way.\r\n\r\n{code}\r\n\r\n synchronized (this.oldsources) {\r\n List previousQueueIds = new ArrayList<>();\r\n for (ReplicationSourceInterface oldSource : this.oldsources) {\r\n\r\n{code}","created":"2020-03-28T10:47:17.348+0000"},{"body":"Problem is, we are removing the list item while iterating the same list in foreach loop.\r\n\r\n \r\n\r\nhttps://github.com/apache/hbase/blob/4565f74af2e9c6f1b3e069cb36e9cad40289337c/hbase-server/src/main/java/org/apache/hadoop/hbase/replication/regionserver/ReplicationSourceManager.java#L499\r\n\r\n \r\n\r\nIterator will solve this issue, will raise PR","created":"2020-04-03T08:36:38.059+0000"},{"body":"Pushed to branch-2.2+. Let me know if you disagree [~zghao]","created":"2020-04-09T23:44:34.593+0000"}],"conversations":[{"body":"Peer update failed with ConcurrentModificationException,\r\n{noformat}\r\n2020-03-28 11:36:49,770 | WARN | RpcServer.default.FPBQ.Fifo.handler=49,queue=4,port=hm_port | Refresh peer peer1 for UPDATE_CONFIG on rs_host,rs_port,1585228936395 failed | org.apache.hadoop.hbase.master.replication.RefreshPeerProcedure.complete(RefreshPeerProcedure.java:114)\r\njava.util.ConcurrentModificationException via szvphispra08478,21302,1585228936395:java.util.ConcurrentModificationException: \r\n\tat org.apache.hadoop.hbase.procedure2.RemoteProcedureException.fromProto(RemoteProcedureException.java:124)\r\n\tat org.apache.hadoop.hbase.master.MasterRpcServices.lambda$reportProcedureDone$4(MasterRpcServices.java:2576)\r\n\tat java.util.ArrayList.forEach(ArrayList.java:1257)\r\n\tat java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1082)\r\n\tat org.apache.hadoop.hbase.master.MasterRpcServices.reportProcedureDone(MasterRpcServices.java:2571)\r\n\tat org.apache.hadoop.hbase.shaded.protobuf.generated.RegionServerStatusProtos$RegionServerStatusService$2.callBlockingMethod(RegionServerStatusProtos.java:15341)\r\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:413)\r\n\tat org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:133)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:338)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:318)\r\nCaused by: java.util.ConcurrentModificationException: \r\n\tat java.util.ArrayList$Itr.checkForComodification(ArrayList.java:909)\r\n\tat java.util.ArrayList$Itr.next(ArrayList.java:859)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSourceManager.refreshSources(ReplicationSourceManager.java:399)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.PeerProcedureHandlerImpl.updatePeerConfig(PeerProcedureHandlerImpl.java:123)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.RefreshPeerCallable.call(RefreshPeerCallable.java:68)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.RefreshPeerCallable.call(RefreshPeerCallable.java:34)\r\n\tat org.apache.hadoop.hbase.regionserver.handler.RSProcedureHandler.process(RSProcedureHandler.java:47)\r\n\tat org.apache.hadoop.hbase.executor.EventHandler.run(EventHandler.java:104)\r\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n\tat java.lang.Thread.run(Thread.java:748)\r\n\r\n{noformat}","from":"reporter","subject":"ConcurrentModificationException occurred in ReplicationSourceManager while refreshing the peer"},{"body":"Which version's hbase? Is there other running peer procedure?","from":"developer"},{"body":"It's HBase-2.2.3. It looks strange because we are accessing oldsources in  a synchronized way.\r\n\r\n{code}\r\n\r\n synchronized (this.oldsources) {\r\n List previousQueueIds = new ArrayList<>();\r\n for (ReplicationSourceInterface oldSource : this.oldsources) {\r\n\r\n{code}","from":"developer"},{"body":"Problem is, we are removing the list item while iterating the same list in foreach loop.\r\n\r\n \r\n\r\nhttps://github.com/apache/hbase/blob/4565f74af2e9c6f1b3e069cb36e9cad40289337c/hbase-server/src/main/java/org/apache/hadoop/hbase/replication/regionserver/ReplicationSourceManager.java#L499\r\n\r\n \r\n\r\nIterator will solve this issue, will raise PR","from":"developer"},{"body":"Pushed to branch-2.2+. Let me know if you disagree [~zghao]","from":"developer"}],"created":"2020-03-28T10:14:05.000+0000","description":"Peer update failed with ConcurrentModificationException,\r\n{noformat}\r\n2020-03-28 11:36:49,770 | WARN | RpcServer.default.FPBQ.Fifo.handler=49,queue=4,port=hm_port | Refresh peer peer1 for UPDATE_CONFIG on rs_host,rs_port,1585228936395 failed | org.apache.hadoop.hbase.master.replication.RefreshPeerProcedure.complete(RefreshPeerProcedure.java:114)\r\njava.util.ConcurrentModificationException via szvphispra08478,21302,1585228936395:java.util.ConcurrentModificationException: \r\n\tat org.apache.hadoop.hbase.procedure2.RemoteProcedureException.fromProto(RemoteProcedureException.java:124)\r\n\tat org.apache.hadoop.hbase.master.MasterRpcServices.lambda$reportProcedureDone$4(MasterRpcServices.java:2576)\r\n\tat java.util.ArrayList.forEach(ArrayList.java:1257)\r\n\tat java.util.Collections$UnmodifiableCollection.forEach(Collections.java:1082)\r\n\tat org.apache.hadoop.hbase.master.MasterRpcServices.reportProcedureDone(MasterRpcServices.java:2571)\r\n\tat org.apache.hadoop.hbase.shaded.protobuf.generated.RegionServerStatusProtos$RegionServerStatusService$2.callBlockingMethod(RegionServerStatusProtos.java:15341)\r\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:413)\r\n\tat org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:133)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:338)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:318)\r\nCaused by: java.util.ConcurrentModificationException: \r\n\tat java.util.ArrayList$Itr.checkForComodification(ArrayList.java:909)\r\n\tat java.util.ArrayList$Itr.next(ArrayList.java:859)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSourceManager.refreshSources(ReplicationSourceManager.java:399)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.PeerProcedureHandlerImpl.updatePeerConfig(PeerProcedureHandlerImpl.java:123)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.RefreshPeerCallable.call(RefreshPeerCallable.java:68)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.RefreshPeerCallable.call(RefreshPeerCallable.java:34)\r\n\tat org.apache.hadoop.hbase.regionserver.handler.RSProcedureHandler.process(RSProcedureHandler.java:47)\r\n\tat org.apache.hadoop.hbase.executor.EventHandler.run(EventHandler.java:104)\r\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n\tat java.lang.Thread.run(Thread.java:748)\r\n\r\n{noformat}","issue_id":"13294701","key":"HBASE-24074","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-04-09T23:44:34.000+0000","role":"fixed_distractor","summary":"ConcurrentModificationException occurred in ReplicationSourceManager while refreshing the peer"} {"case_id":"12461334","cluster":"DISTRACTOR-HBASE-2414","comments":[{"body":"Something along the following lines would help accomplish this:\n\n- In a version of the tests, the master and region server threads each have a method to receive a string representation of the task they have to carry out\n- They use reflection and execute the call only when called by the external caller. That way we can \"serialize\" the distributed execution.\n","created":"2010-04-07T00:45:01.883+0000"},{"body":"Here is an ugly part-patch. Its an attempt at doing a cluster that has no RPC. Master and RegionServer would invoke each other's methods directly rather than go via RPC. This is not the right approach -- subclassing HMaster and RS and trying to shutdown the networking. It'd take for ever to do and it'd be ugly when done.\n\nLet me think more on making it so regionservers and masters take a 'script' as K suggests above.","created":"2010-04-13T06:19:14.071+0000"},{"body":"Failed attempt 2: Starts a master with disabled scanners and then fakes it out by invoking master methods directly calling region start and region server heartbeat methods, passing back the region open, etc. messages. All works pretty good coming up but then master wants to connect to the non-running meta to add edis tot the table itself. We could blank this networking out but in at least a few of the tests we want to build updates to fail legitimately because .meta. is not available. This tack may be worth further dev but we need something better.\n\nLooking now at bringing up the minihdfscluster and then doing direct injection as per the above patch on the individual daemons to get what I want. Also looking at changing the proxy so its not over rpc but goes local instead.","created":"2010-04-22T17:57:50.928+0000"},{"body":"This approach seems more promising. It makes all the interfaces used networking configurable and then provides direct connect between master and regionservers with no network involved. The instances are readily available so can do stuff like remove meta and then send in a close message. Not done yet. Just posting what I have so far. Will keep at it.","created":"2010-04-23T08:05:59.028+0000"},{"body":"Marking blocker and pulling into 0.20.5 and 0.21","created":"2010-04-23T08:07:16.996+0000"},{"body":"This version better integrates the use of nonetwork implementations of pertinent interfaces. Not done yet. Need to work it into HBaseTestingUtility better and still need to prove that testing cluster transistion with this stuff is viable (it still looks so, just need to do the proof still). If want to take a look now, start with TestClusterTransistions.","created":"2010-04-23T19:17:52.180+0000"},{"body":"Cleaned up errors on shutdown. Now clean start/stop of nonetwork cluster (and old stuff confirmed working as it used to). Next is try to do actual unit test.","created":"2010-04-24T05:47:59.140+0000"},{"body":"This patch is not for review. Its a mess that has two testing techniques wrapped up in it still and is in need of cleanup.\n\nMy pursuit of a 'direct'/nonetwork hbase ran into the weeds; i.e the 'first' technique. All of the client logic down in HConnnectionManager#TableServers would need to be redone so its 'direct'. Subclassing HConnectionManager#TableServers helped in that I could leverage what was already there but still, a lot to be done. So, i tried doing reproduction of cluster bad-cases using old-school minihbasecluster machinations (technique 'two').\n\nThe pursuit of technique 'one' opened up our minihbasecluster code making it so I was able to write a simple test to repro what is seen here in the stack trace at the head of this issue (patch includes the code in the TestClusterTransition junit test). \n\nHere is what my test is showing...\n\n{code}\n2010-04-25 20:31:07,946 INFO [RegionServer:1] regionserver.HRegionServer(649): aborting server at: 192.168.1.106:63335\n2010-04-25 20:31:07,956 INFO [main-EventThread] master.ServerManager$ServerExpirer(831): 192.168.1.106,63335,1272252645013 znode expired\n2010-04-25 20:31:07,957 INFO [main-EventThread] master.RegionManager(795): META region removed from onlineMetaRegions\n2010-04-25 20:31:07,958 DEBUG [RegionServer:1] zookeeper.ZooKeeperWrapper(682): Closed connection with ZooKeeper\n2010-04-25 20:31:07,958 INFO [RegionServer:1] regionserver.HRegionServer(696): RegionServer:1 exiting\n2010-04-25 20:31:07,960 INFO [main] regionserver.HRegionServer(261): My address is 192.168.1.106:0\n2010-04-25 20:31:07,961 INFO [main] ipc.HBaseRpcMetrics(52): Initializing RPC Metrics with hostName=HRegionServer, port=63412\n2010-04-25 20:31:07,962 INFO [main] regionserver.MemStoreFlusher(102): globalMemStoreLimit=32.6m, globalMemStoreLimitLowMark=20.4m, maxHeap=81.4m\n2010-04-25 20:31:07,963 INFO [main] regionserver.HRegionServer$MajorCompactionChecker(984): Runs every 1000000ms\n2010-04-25 20:31:07,968 DEBUG [HMaster] master.HMaster(506): Processing todo: ProcessServerShutdown of 192.168.1.106,63335,1272252645013\n...\n# While processing server shutdown in came the new RS instance w/ same port and load balance kicks in....\n...\n2010-04-25 20:31:08,187 INFO [RegionServer:1] regionserver.HRegionServer(1202): HRegionServer started at: 192.168.1.106:63412\n2010-04-25 20:31:08,188 DEBUG [RegionServer:1] zookeeper.ZooKeeperWrapper(398): Read ZNode /hbase/root-region-server got 192.168.1.106:63333\n2010-04-25 20:31:08,218 DEBUG [pool-1-thread-1] regionserver.HLog$1(1278): Thread got 53 to process\n2010-04-25 20:31:08,222 DEBUG [IPC Server handler 4 on 60000] master.RegionManager$LoadBalancer(1447): Server is overloaded: load=15, avg=7.5, slop=0.3\n...\n# Then fell into... this while processing a close region\n\n2010-04-25 20:31:08,360 DEBUG [HMaster] master.HMaster(506): Processing todo: ProcessRegionClose of 2428,fff,1272252656267, false, reassign: true\n2010-04-25 20:31:08,362 DEBUG [HMaster] master.RetryableMetaOperation(95): Exception in RetryableMetaOperation: \njava.lang.NullPointerException\n\tat org.apache.hadoop.hbase.master.RetryableMetaOperation.doWithRetries(RetryableMetaOperation.java:65)\n\tat org.apache.hadoop.hbase.master.ProcessRegionClose.process(ProcessRegionClose.java:89)\n\tat org.apache.hadoop.hbase.master.HMaster.processToDoQueue(HMaster.java:510)\n\tat org.apache.hadoop.hbase.master.HMaster.run(HMaster.java:445)\n2010-04-25 20:31:08,367 WARN [HMaster] master.HMaster(546): Processing pending operations: ProcessRegionClose of 2428,fff,1272252656267, false, reassign: true\njava.lang.RuntimeException: java.lang.NullPointerException\n\tat org.apache.hadoop.hbase.master.RetryableMetaOperation.doWithRetries(RetryableMetaOperation.java:96)\n\tat org.apache.hadoop.hbase.master.ProcessRegionClose.process(ProcessRegionClose.java:89)\n\tat org.apache.hadoop.hbase.master.HMaster.processToDoQueue(HMaster.java:510)\n\tat org.apache.hadoop.hbase.master.HMaster.run(HMaster.java:445)\nCaused by: java.lang.NullPointerException\n\tat org.apache.hadoop.hbase.master.RetryableMetaOperation.doWithRetries(RetryableMetaOperation.java:65)\n\t... 3 more\n..\n\nand so on...\n{code}\n\nIt required some knowledge of minihbasecluster internals but its not too bad methinks. I've added a bunch of doc. so others can follow. \n\nLet me clean up and repro more of the recent cluster failings in unit test scenario using minihbasecluster.","created":"2010-04-26T03:52:36.230+0000"},{"body":"Cleaned up patch. Adds a 'kill' method to HRS. Included TestClusterTransition test will not shutdown because we are trying to do a close and can't because meta is not online nor will it ever come online (HBASE-2428).","created":"2010-04-27T07:16:42.425+0000"},{"body":"Trying to use powermock and mockito inside a minihbasecluster context I run into the following issues (after struggling past similar around commons logging):\n\n{code}\n\nTestcase: testRegionCloseWhenNoMetaHBase2428 took 1.659 sec\n Caused an ERROR\nloader constraint violation: loader (instance of org/powermock/core/classloader/MockClassLoader) previously initiated loading for a different type with name \"javax/management/MBeanServer\"\njava.lang.LinkageError: loader constraint violation: loader (instance of org/powermock/core/classloader/MockClassLoader) previously initiated loading for a different type with name \"javax/management/MBeanServer\"\n at java.lang.ClassLoader.defineClass1(Native Method)\n at java.lang.ClassLoader.defineClass(ClassLoader.java:700)\n at java.lang.ClassLoader.defineClass(ClassLoader.java:545)\n at org.powermock.core.classloader.MockClassLoader.loadUnmockedClass(MockClassLoader.java:190)\n at org.powermock.core.classloader.MockClassLoader.loadModifiedClass(MockClassLoader.java:148)\n at org.powermock.core.classloader.DeferSupportingClassLoader.loadClass(DeferSupportingClassLoader.java:63)\n at java.lang.ClassLoader.loadClass(ClassLoader.java:254)\n at java.lang.ClassLoader.loadClassInternal(ClassLoader.java:399)\n at org.apache.hadoop.metrics.util.MBeanUtil.registerMBean(MBeanUtil.java:53)\n at org.apache.hadoop.ipc.metrics.RpcActivityMBean.(RpcActivityMBean.java:70)\n at org.apache.hadoop.ipc.metrics.RpcMetrics.(RpcMetrics.java:64)\n at org.apache.hadoop.ipc.Server.(Server.java:1028)\n at org.apache.hadoop.ipc.RPC$Server.(RPC.java:488)\n at org.apache.hadoop.ipc.RPC.getServer(RPC.java:450)\n at org.apache.hadoop.hdfs.server.namenode.NameNode.initialize(NameNode.java:191)\n at org.apache.hadoop.hdfs.server.namenode.NameNode.(NameNode.java:279)\n at org.apache.hadoop.hdfs.server.namenode.NameNode.createNameNode(NameNode.java:956)\n at org.apache.hadoop.hdfs.MiniDFSCluster.(MiniDFSCluster.java:275)\n at org.apache.hadoop.hbase.HBaseTestingUtility.startMiniCluster(HBaseTestingUtility.java:200)\n at org.apache.hadoop.hbase.TestClusterTransitions.beforeAllTests(TestClusterTransitions.java:44)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl$PowerMockJUnit44MethodRunner.executeTest(PowerMockJUnit44RunnerDelegateImpl.java:309)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit47RunnerDelegateImpl$PowerMockJUnit47MethodRunner.executeTestInSuper(PowerMockJUnit47RunnerDelegateImpl.java:112)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit47RunnerDelegateImpl$PowerMockJUnit47MethodRunner.executeTest(PowerMockJUnit47RunnerDelegateImpl.java:73)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl$PowerMockJUnit44MethodRunner.runBeforesThenTestThenAfters(PowerMockJUnit44RunnerDelegateImpl.java:297)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl.invokeTestMethod(PowerMockJUnit44RunnerDelegateImpl.java:222)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl.runMethods(PowerMockJUnit44RunnerDelegateImpl.java:161)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl$1.run(PowerMockJUnit44RunnerDelegateImpl.java:135)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl.run(PowerMockJUnit44RunnerDelegateImpl.java:133)\n at org.powermock.modules.junit4.common.internal.impl.JUnit4TestSuiteChunkerImpl.run(JUnit4TestSuiteChunkerImpl.java:112)\n at org.powermock.modules.junit4.common.internal.impl.AbstractCommonPowerMockRunner.run(AbstractCommonPowerMockRunner.java:55)\n{code}\n\nThe above classloading uglyness looks like it would take a while to undo -- if its possible at all -- and when done, I'd end up with a fragile test environment is my guess. I'm going to pass on this combo unless others have input.","created":"2010-04-27T19:26:07.516+0000"},{"body":"Attached patch should have sufficient vocabulary for writing tests that can repro recent rash of issues. It adds means of adding listeners to master events -- both before and after event happens. The before lets you put off the event. it'll be requeued, added to the DelayQueue (this happens currently in Master if an issue processing a RegionServerOperation, it gets requeued to be tried again later). The patch also adds to minihbasecluster simple means of sending regionserver \"events'.\n\nPatch includes test that reproduces hbase-2428. Let me add more unit tests for other issues.","created":"2010-04-28T06:35:47.486+0000"},{"body":"I want to commit this patch. Can I get a review. Its kinda polluted in that it contains:\n\n1. Means of testing state transitions across the master. Tests can register listeners. In the listener you should be able to delay, cancel, count, etc., RegionServerOperations. It should be possible in the listener simulating uglyness seen out on live clusters.\n2. A fix for the HBASE-2428 bug.\n3. Facility added to TestHBaseUtility and MiniHBaseCluster\n4. I started refactoring of the queue of RegionServerOperations in master moving it out to a separate file making it testable but then I ran into fact that RegionServerOperations each have a reference to master AND they can put themselves back on the queue -- the circularity baffles. This has to be fixed but will do in a separate patch.\n\nHere is a commit message with some detail on the commit:\n\n{code}\nM src/test/org/apache/hadoop/hbase/HBaseTestingUtility.java\n Broke up startMiniHBaseCluster into smaller methods so can mix and\n match pieces of minihbasecluster toward other ends.\n (setupClusterBuildDir, isRunningCluster, getMiniHBaseCluster): Added.\nM src/test/org/apache/hadoop/hbase/TestInfoServers.java\nM src/test/org/apache/hadoop/hbase/TestRegionRebalancing.java\nM src/test/org/apache/hadoop/hbase/HBaseClusterTestCase.java\nM src/test/org/apache/hadoop/hbase/regionserver/TestLogRolling.java\nM src/test/org/apache/hadoop/hbase/regionserver/DisabledTestRegionServerExit.java\nM src/test/org/apache/hadoop/hbase/mapreduce/TestTableIndex.java\nM src/test/org/apache/hadoop/hbase/mapred/TestTableIndex.java\n Ripple from change of MiniHBaseCluster.getRegionThreads to\n getRegionServerThreads.\nM src/test/org/apache/hadoop/hbase/MiniHBaseCluster.java\n Added new MiniHBaseClusterMaster that is override of HMaster so\n I can piggyback messages for designated regionservers atop the\n heartbeat: close region, etc.\n (getServerWithMeta, addMessageToSendRegionServer): Added.\nA src/test/org/apache/hadoop/hbase/master/TestRegionServerOperationQueue.java\n Stubbed out test of new RegionServerOperationQueue class.\nA src/test/org/apache/hadoop/hbase/master/TestMasterTransistions.java\n Test master cluster transistions. Includes unit test of hbase-2428.\nM src/test/org/apache/hadoop/hbase/util/TestMigration.java\n Disable migration test. Nothing to migrate yet and besides it was\n trying to load a 0.19 hbase data tar.gz that has since been removed.\nM src/contrib/stargate/build.xml\n Added a copyright.\nM src/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java\n Documentation and moved some methods from down at tail of the class\n where they were in among static methods used parsing cmd-line usage\n up above the usage and master startup static methods.\n Added fix for issue broken by an hbase-1215 commit where we were\n looking at wrong address (Grep for r796326 for more).\nM src/java/org/apache/hadoop/hbase/LocalHBaseCluster.java\n Moved bulk out to new JVMClusterUtils class and made accessible.\n Added passing of HMaster.class to instantiate to facilitate\n passing of TestHMaster.class.\nM src/java/org/apache/hadoop/hbase/master/RegionServerOperationQueue.java\n The RegionServerOperations queues moved out to their own class from\n Master. Allows listeners to register and get notice before and after\n a RegionServerOperation is processed. Includes part of bug fix for\n hbase-2428. When an error processing a RegionServerOperation, we'd fall\n into the IOException catch. We'd then put the operation back on the delay\n queue for later processing only we'd not reset its expiration. It\n would therefore run again immmediately... fail again, and so on.\n Changed the return from process to be an enum rather than true/false\n so I don't have to have do things like call checkfs down in here and\n I don't need to have a master instance around.\nM src/java/org/apache/hadoop/hbase/master/ServerManager.java\n How we add RegionServerOperation instances has changed to go via\n RegionServerOperationQueue now.\nM src/java/org/apache/hadoop/hbase/master/ProcessServerShutdown.java\n (getDeadServerAddress): Added.\nM src/java/org/apache/hadoop/hbase/master/RegionServerOperationListener.java\n Listener interface to implement if interested in watching RegionServerOperations.\nM src/java/org/apache/hadoop/hbase/master/HMaster.java\n Moved the RegionServerOperation code out of here to\n RegionServerOperationQueue.\n (adornRegionServerAnswer, constructMaster): Added.\nM src/java/org/apache/hadoop/hbase/master/ProcessRegionOpen.java\n Comment.\nM src/java/org/apache/hadoop/hbase/master/ProcessRegionClose.java\n Added in *fix* for 2428 NPE. For now did what happens in\n ProcessRegionOpen for symmetry's sake but it needs to be replaced.\nM src/java/org/apache/hadoop/hbase/master/RegionServerOperation.java\n (resetExpiration): Added.\nM src/java/org/apache/hadoop/hbase/master/ProcessRegionStatusChange.java\n Javaadoc.\nM src/java/org/apache/hadoop/hbase/util/Threads.java\n (threadDumpingIsAlive): Added from LocalHBaseCluster.\n (sleep): Added.\nM src/java/org/apache/hadoop/hbase/util/JVMClusterUtil.java\n New class that has facility moved from LocalHBaseCluster with added\n javadoc and made accessible. Needed testing.\n{code}","created":"2010-04-29T02:38:52.179+0000"},{"body":"Awesome stuff Stack! Patch looks great, just a couple of small comments:\n\n- Revert src/java/org/apache/hadoop/hbase/master/ProcessRegionStatusChange.java (no change in file)\n- src/java/org/apache/hadoop/hbase/master/ProcessServerShutdown.java:60 Don't need row parameter in \"ToDoEntry(final byte [] row, final HRegionInfo info)\" anymore.\n","created":"2010-04-29T18:17:01.607+0000"},{"body":"Thanks Karthik. Will fix above on commit.","created":"2010-04-30T00:58:44.918+0000"},{"body":"Committed to branch and trunk. Thanks for review Karthik. I committed with your suggested changes. I'm resolving this issue though I'm sure there is probably a better way of testing cluster transitions. This should do for now. We can open a new JIRA for NG or improvements.","created":"2010-04-30T06:55:08.158+0000"},{"body":"Marking these as fixed against 0.21.0 rather than against 0.20.5.","created":"2010-05-12T23:52:30.003+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T12:42:00.942+0000"}],"conversations":[{"body":"We keep finding good cases that are reasonably hard to test, yet the test suite does not encode these. \nFor example: \nHBASE-2413 Master does not respect generation stamps, may result in meta getting permanently offlined\nHBASE-2312 Possible data loss when RS goes into GC pause while rolling HLog\n\nI am sure there are many more such \"scenarios\" we should put into the unit tests. \n\n","from":"reporter","subject":"Enhance test suite to be able to specify distributed scenarios"},{"body":"Something along the following lines would help accomplish this:\n\n- In a version of the tests, the master and region server threads each have a method to receive a string representation of the task they have to carry out\n- They use reflection and execute the call only when called by the external caller. That way we can \"serialize\" the distributed execution.\n","from":"developer"},{"body":"Here is an ugly part-patch. Its an attempt at doing a cluster that has no RPC. Master and RegionServer would invoke each other's methods directly rather than go via RPC. This is not the right approach -- subclassing HMaster and RS and trying to shutdown the networking. It'd take for ever to do and it'd be ugly when done.\n\nLet me think more on making it so regionservers and masters take a 'script' as K suggests above.","from":"developer"},{"body":"Failed attempt 2: Starts a master with disabled scanners and then fakes it out by invoking master methods directly calling region start and region server heartbeat methods, passing back the region open, etc. messages. All works pretty good coming up but then master wants to connect to the non-running meta to add edis tot the table itself. We could blank this networking out but in at least a few of the tests we want to build updates to fail legitimately because .meta. is not available. This tack may be worth further dev but we need something better.\n\nLooking now at bringing up the minihdfscluster and then doing direct injection as per the above patch on the individual daemons to get what I want. Also looking at changing the proxy so its not over rpc but goes local instead.","from":"developer"},{"body":"This approach seems more promising. It makes all the interfaces used networking configurable and then provides direct connect between master and regionservers with no network involved. The instances are readily available so can do stuff like remove meta and then send in a close message. Not done yet. Just posting what I have so far. Will keep at it.","from":"developer"},{"body":"Marking blocker and pulling into 0.20.5 and 0.21","from":"developer"},{"body":"This version better integrates the use of nonetwork implementations of pertinent interfaces. Not done yet. Need to work it into HBaseTestingUtility better and still need to prove that testing cluster transistion with this stuff is viable (it still looks so, just need to do the proof still). If want to take a look now, start with TestClusterTransistions.","from":"developer"},{"body":"Cleaned up errors on shutdown. Now clean start/stop of nonetwork cluster (and old stuff confirmed working as it used to). Next is try to do actual unit test.","from":"developer"},{"body":"This patch is not for review. Its a mess that has two testing techniques wrapped up in it still and is in need of cleanup.\n\nMy pursuit of a 'direct'/nonetwork hbase ran into the weeds; i.e the 'first' technique. All of the client logic down in HConnnectionManager#TableServers would need to be redone so its 'direct'. Subclassing HConnectionManager#TableServers helped in that I could leverage what was already there but still, a lot to be done. So, i tried doing reproduction of cluster bad-cases using old-school minihbasecluster machinations (technique 'two').\n\nThe pursuit of technique 'one' opened up our minihbasecluster code making it so I was able to write a simple test to repro what is seen here in the stack trace at the head of this issue (patch includes the code in the TestClusterTransition junit test). \n\nHere is what my test is showing...\n\n{code}\n2010-04-25 20:31:07,946 INFO [RegionServer:1] regionserver.HRegionServer(649): aborting server at: 192.168.1.106:63335\n2010-04-25 20:31:07,956 INFO [main-EventThread] master.ServerManager$ServerExpirer(831): 192.168.1.106,63335,1272252645013 znode expired\n2010-04-25 20:31:07,957 INFO [main-EventThread] master.RegionManager(795): META region removed from onlineMetaRegions\n2010-04-25 20:31:07,958 DEBUG [RegionServer:1] zookeeper.ZooKeeperWrapper(682): Closed connection with ZooKeeper\n2010-04-25 20:31:07,958 INFO [RegionServer:1] regionserver.HRegionServer(696): RegionServer:1 exiting\n2010-04-25 20:31:07,960 INFO [main] regionserver.HRegionServer(261): My address is 192.168.1.106:0\n2010-04-25 20:31:07,961 INFO [main] ipc.HBaseRpcMetrics(52): Initializing RPC Metrics with hostName=HRegionServer, port=63412\n2010-04-25 20:31:07,962 INFO [main] regionserver.MemStoreFlusher(102): globalMemStoreLimit=32.6m, globalMemStoreLimitLowMark=20.4m, maxHeap=81.4m\n2010-04-25 20:31:07,963 INFO [main] regionserver.HRegionServer$MajorCompactionChecker(984): Runs every 1000000ms\n2010-04-25 20:31:07,968 DEBUG [HMaster] master.HMaster(506): Processing todo: ProcessServerShutdown of 192.168.1.106,63335,1272252645013\n...\n# While processing server shutdown in came the new RS instance w/ same port and load balance kicks in....\n...\n2010-04-25 20:31:08,187 INFO [RegionServer:1] regionserver.HRegionServer(1202): HRegionServer started at: 192.168.1.106:63412\n2010-04-25 20:31:08,188 DEBUG [RegionServer:1] zookeeper.ZooKeeperWrapper(398): Read ZNode /hbase/root-region-server got 192.168.1.106:63333\n2010-04-25 20:31:08,218 DEBUG [pool-1-thread-1] regionserver.HLog$1(1278): Thread got 53 to process\n2010-04-25 20:31:08,222 DEBUG [IPC Server handler 4 on 60000] master.RegionManager$LoadBalancer(1447): Server is overloaded: load=15, avg=7.5, slop=0.3\n...\n# Then fell into... this while processing a close region\n\n2010-04-25 20:31:08,360 DEBUG [HMaster] master.HMaster(506): Processing todo: ProcessRegionClose of 2428,fff,1272252656267, false, reassign: true\n2010-04-25 20:31:08,362 DEBUG [HMaster] master.RetryableMetaOperation(95): Exception in RetryableMetaOperation: \njava.lang.NullPointerException\n\tat org.apache.hadoop.hbase.master.RetryableMetaOperation.doWithRetries(RetryableMetaOperation.java:65)\n\tat org.apache.hadoop.hbase.master.ProcessRegionClose.process(ProcessRegionClose.java:89)\n\tat org.apache.hadoop.hbase.master.HMaster.processToDoQueue(HMaster.java:510)\n\tat org.apache.hadoop.hbase.master.HMaster.run(HMaster.java:445)\n2010-04-25 20:31:08,367 WARN [HMaster] master.HMaster(546): Processing pending operations: ProcessRegionClose of 2428,fff,1272252656267, false, reassign: true\njava.lang.RuntimeException: java.lang.NullPointerException\n\tat org.apache.hadoop.hbase.master.RetryableMetaOperation.doWithRetries(RetryableMetaOperation.java:96)\n\tat org.apache.hadoop.hbase.master.ProcessRegionClose.process(ProcessRegionClose.java:89)\n\tat org.apache.hadoop.hbase.master.HMaster.processToDoQueue(HMaster.java:510)\n\tat org.apache.hadoop.hbase.master.HMaster.run(HMaster.java:445)\nCaused by: java.lang.NullPointerException\n\tat org.apache.hadoop.hbase.master.RetryableMetaOperation.doWithRetries(RetryableMetaOperation.java:65)\n\t... 3 more\n..\n\nand so on...\n{code}\n\nIt required some knowledge of minihbasecluster internals but its not too bad methinks. I've added a bunch of doc. so others can follow. \n\nLet me clean up and repro more of the recent cluster failings in unit test scenario using minihbasecluster.","from":"developer"},{"body":"Cleaned up patch. Adds a 'kill' method to HRS. Included TestClusterTransition test will not shutdown because we are trying to do a close and can't because meta is not online nor will it ever come online (HBASE-2428).","from":"developer"},{"body":"Trying to use powermock and mockito inside a minihbasecluster context I run into the following issues (after struggling past similar around commons logging):\n\n{code}\n\nTestcase: testRegionCloseWhenNoMetaHBase2428 took 1.659 sec\n Caused an ERROR\nloader constraint violation: loader (instance of org/powermock/core/classloader/MockClassLoader) previously initiated loading for a different type with name \"javax/management/MBeanServer\"\njava.lang.LinkageError: loader constraint violation: loader (instance of org/powermock/core/classloader/MockClassLoader) previously initiated loading for a different type with name \"javax/management/MBeanServer\"\n at java.lang.ClassLoader.defineClass1(Native Method)\n at java.lang.ClassLoader.defineClass(ClassLoader.java:700)\n at java.lang.ClassLoader.defineClass(ClassLoader.java:545)\n at org.powermock.core.classloader.MockClassLoader.loadUnmockedClass(MockClassLoader.java:190)\n at org.powermock.core.classloader.MockClassLoader.loadModifiedClass(MockClassLoader.java:148)\n at org.powermock.core.classloader.DeferSupportingClassLoader.loadClass(DeferSupportingClassLoader.java:63)\n at java.lang.ClassLoader.loadClass(ClassLoader.java:254)\n at java.lang.ClassLoader.loadClassInternal(ClassLoader.java:399)\n at org.apache.hadoop.metrics.util.MBeanUtil.registerMBean(MBeanUtil.java:53)\n at org.apache.hadoop.ipc.metrics.RpcActivityMBean.(RpcActivityMBean.java:70)\n at org.apache.hadoop.ipc.metrics.RpcMetrics.(RpcMetrics.java:64)\n at org.apache.hadoop.ipc.Server.(Server.java:1028)\n at org.apache.hadoop.ipc.RPC$Server.(RPC.java:488)\n at org.apache.hadoop.ipc.RPC.getServer(RPC.java:450)\n at org.apache.hadoop.hdfs.server.namenode.NameNode.initialize(NameNode.java:191)\n at org.apache.hadoop.hdfs.server.namenode.NameNode.(NameNode.java:279)\n at org.apache.hadoop.hdfs.server.namenode.NameNode.createNameNode(NameNode.java:956)\n at org.apache.hadoop.hdfs.MiniDFSCluster.(MiniDFSCluster.java:275)\n at org.apache.hadoop.hbase.HBaseTestingUtility.startMiniCluster(HBaseTestingUtility.java:200)\n at org.apache.hadoop.hbase.TestClusterTransitions.beforeAllTests(TestClusterTransitions.java:44)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl$PowerMockJUnit44MethodRunner.executeTest(PowerMockJUnit44RunnerDelegateImpl.java:309)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit47RunnerDelegateImpl$PowerMockJUnit47MethodRunner.executeTestInSuper(PowerMockJUnit47RunnerDelegateImpl.java:112)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit47RunnerDelegateImpl$PowerMockJUnit47MethodRunner.executeTest(PowerMockJUnit47RunnerDelegateImpl.java:73)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl$PowerMockJUnit44MethodRunner.runBeforesThenTestThenAfters(PowerMockJUnit44RunnerDelegateImpl.java:297)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl.invokeTestMethod(PowerMockJUnit44RunnerDelegateImpl.java:222)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl.runMethods(PowerMockJUnit44RunnerDelegateImpl.java:161)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl$1.run(PowerMockJUnit44RunnerDelegateImpl.java:135)\n at org.powermock.modules.junit4.internal.impl.PowerMockJUnit44RunnerDelegateImpl.run(PowerMockJUnit44RunnerDelegateImpl.java:133)\n at org.powermock.modules.junit4.common.internal.impl.JUnit4TestSuiteChunkerImpl.run(JUnit4TestSuiteChunkerImpl.java:112)\n at org.powermock.modules.junit4.common.internal.impl.AbstractCommonPowerMockRunner.run(AbstractCommonPowerMockRunner.java:55)\n{code}\n\nThe above classloading uglyness looks like it would take a while to undo -- if its possible at all -- and when done, I'd end up with a fragile test environment is my guess. I'm going to pass on this combo unless others have input.","from":"developer"},{"body":"Attached patch should have sufficient vocabulary for writing tests that can repro recent rash of issues. It adds means of adding listeners to master events -- both before and after event happens. The before lets you put off the event. it'll be requeued, added to the DelayQueue (this happens currently in Master if an issue processing a RegionServerOperation, it gets requeued to be tried again later). The patch also adds to minihbasecluster simple means of sending regionserver \"events'.\n\nPatch includes test that reproduces hbase-2428. Let me add more unit tests for other issues.","from":"developer"},{"body":"I want to commit this patch. Can I get a review. Its kinda polluted in that it contains:\n\n1. Means of testing state transitions across the master. Tests can register listeners. In the listener you should be able to delay, cancel, count, etc., RegionServerOperations. It should be possible in the listener simulating uglyness seen out on live clusters.\n2. A fix for the HBASE-2428 bug.\n3. Facility added to TestHBaseUtility and MiniHBaseCluster\n4. I started refactoring of the queue of RegionServerOperations in master moving it out to a separate file making it testable but then I ran into fact that RegionServerOperations each have a reference to master AND they can put themselves back on the queue -- the circularity baffles. This has to be fixed but will do in a separate patch.\n\nHere is a commit message with some detail on the commit:\n\n{code}\nM src/test/org/apache/hadoop/hbase/HBaseTestingUtility.java\n Broke up startMiniHBaseCluster into smaller methods so can mix and\n match pieces of minihbasecluster toward other ends.\n (setupClusterBuildDir, isRunningCluster, getMiniHBaseCluster): Added.\nM src/test/org/apache/hadoop/hbase/TestInfoServers.java\nM src/test/org/apache/hadoop/hbase/TestRegionRebalancing.java\nM src/test/org/apache/hadoop/hbase/HBaseClusterTestCase.java\nM src/test/org/apache/hadoop/hbase/regionserver/TestLogRolling.java\nM src/test/org/apache/hadoop/hbase/regionserver/DisabledTestRegionServerExit.java\nM src/test/org/apache/hadoop/hbase/mapreduce/TestTableIndex.java\nM src/test/org/apache/hadoop/hbase/mapred/TestTableIndex.java\n Ripple from change of MiniHBaseCluster.getRegionThreads to\n getRegionServerThreads.\nM src/test/org/apache/hadoop/hbase/MiniHBaseCluster.java\n Added new MiniHBaseClusterMaster that is override of HMaster so\n I can piggyback messages for designated regionservers atop the\n heartbeat: close region, etc.\n (getServerWithMeta, addMessageToSendRegionServer): Added.\nA src/test/org/apache/hadoop/hbase/master/TestRegionServerOperationQueue.java\n Stubbed out test of new RegionServerOperationQueue class.\nA src/test/org/apache/hadoop/hbase/master/TestMasterTransistions.java\n Test master cluster transistions. Includes unit test of hbase-2428.\nM src/test/org/apache/hadoop/hbase/util/TestMigration.java\n Disable migration test. Nothing to migrate yet and besides it was\n trying to load a 0.19 hbase data tar.gz that has since been removed.\nM src/contrib/stargate/build.xml\n Added a copyright.\nM src/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java\n Documentation and moved some methods from down at tail of the class\n where they were in among static methods used parsing cmd-line usage\n up above the usage and master startup static methods.\n Added fix for issue broken by an hbase-1215 commit where we were\n looking at wrong address (Grep for r796326 for more).\nM src/java/org/apache/hadoop/hbase/LocalHBaseCluster.java\n Moved bulk out to new JVMClusterUtils class and made accessible.\n Added passing of HMaster.class to instantiate to facilitate\n passing of TestHMaster.class.\nM src/java/org/apache/hadoop/hbase/master/RegionServerOperationQueue.java\n The RegionServerOperations queues moved out to their own class from\n Master. Allows listeners to register and get notice before and after\n a RegionServerOperation is processed. Includes part of bug fix for\n hbase-2428. When an error processing a RegionServerOperation, we'd fall\n into the IOException catch. We'd then put the operation back on the delay\n queue for later processing only we'd not reset its expiration. It\n would therefore run again immmediately... fail again, and so on.\n Changed the return from process to be an enum rather than true/false\n so I don't have to have do things like call checkfs down in here and\n I don't need to have a master instance around.\nM src/java/org/apache/hadoop/hbase/master/ServerManager.java\n How we add RegionServerOperation instances has changed to go via\n RegionServerOperationQueue now.\nM src/java/org/apache/hadoop/hbase/master/ProcessServerShutdown.java\n (getDeadServerAddress): Added.\nM src/java/org/apache/hadoop/hbase/master/RegionServerOperationListener.java\n Listener interface to implement if interested in watching RegionServerOperations.\nM src/java/org/apache/hadoop/hbase/master/HMaster.java\n Moved the RegionServerOperation code out of here to\n RegionServerOperationQueue.\n (adornRegionServerAnswer, constructMaster): Added.\nM src/java/org/apache/hadoop/hbase/master/ProcessRegionOpen.java\n Comment.\nM src/java/org/apache/hadoop/hbase/master/ProcessRegionClose.java\n Added in *fix* for 2428 NPE. For now did what happens in\n ProcessRegionOpen for symmetry's sake but it needs to be replaced.\nM src/java/org/apache/hadoop/hbase/master/RegionServerOperation.java\n (resetExpiration): Added.\nM src/java/org/apache/hadoop/hbase/master/ProcessRegionStatusChange.java\n Javaadoc.\nM src/java/org/apache/hadoop/hbase/util/Threads.java\n (threadDumpingIsAlive): Added from LocalHBaseCluster.\n (sleep): Added.\nM src/java/org/apache/hadoop/hbase/util/JVMClusterUtil.java\n New class that has facility moved from LocalHBaseCluster with added\n javadoc and made accessible. Needed testing.\n{code}","from":"developer"},{"body":"Awesome stuff Stack! Patch looks great, just a couple of small comments:\n\n- Revert src/java/org/apache/hadoop/hbase/master/ProcessRegionStatusChange.java (no change in file)\n- src/java/org/apache/hadoop/hbase/master/ProcessServerShutdown.java:60 Don't need row parameter in \"ToDoEntry(final byte [] row, final HRegionInfo info)\" anymore.\n","from":"developer"},{"body":"Thanks Karthik. Will fix above on commit.","from":"developer"},{"body":"Committed to branch and trunk. Thanks for review Karthik. I committed with your suggested changes. I'm resolving this issue though I'm sure there is probably a better way of testing cluster transitions. This should do for now. We can open a new JIRA for NG or improvements.","from":"developer"},{"body":"Marking these as fixed against 0.21.0 rather than against 0.20.5.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2010-04-07T00:44:40.000+0000","description":"We keep finding good cases that are reasonably hard to test, yet the test suite does not encode these. \nFor example: \nHBASE-2413 Master does not respect generation stamps, may result in meta getting permanently offlined\nHBASE-2312 Possible data loss when RS goes into GC pause while rolling HLog\n\nI am sure there are many more such \"scenarios\" we should put into the unit tests. \n\n","issue_id":"12461334","key":"HBASE-2414","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2010-04-30T06:55:08.000+0000","role":"fixed_distractor","summary":"Enhance test suite to be able to specify distributed scenarios"} {"case_id":"13297452","cluster":"DISTRACTOR-HBASE-24158","comments":[{"body":"Its an old test. It cam in w/ HBASE-16945. Back then it was running w/ 20 threads. Recently I downed it to 10. Will keep an eye on it.","created":"2020-04-09T18:02:07.733+0000"},{"body":"Attached dumb patch that does nothing but use a few less resources; no notion if helps.","created":"2020-04-09T18:04:31.997+0000"},{"body":"Hmm... Yeah, it failed in the last two nightly runs on branch-2. Let me push this patch on branch-2 and see if it helps.","created":"2020-04-09T18:17:46.910+0000"},{"body":"Ok. Pushed below on branch-2. Lets see if helps.\r\n{code}\r\ncommit 2d117963809929e9523e47e94b7e5faab7db8e1a (HEAD -> 2)\r\nAuthor: stack \r\nDate: Thu Apr 9 11:03:22 2020 -0700\r\n\r\n HBASE-24158 [Flakey Tests] TestAsyncTableGetMultiThreaded\r\n\r\ndiff --git a/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestAsyncTableGetMultiThreaded.java b/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestAsyncTableGetMultiThreaded.java\r\nindex 7c896957fa..ce6bc05ca0 100644\r\n--- a/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestAsyncTableGetMultiThreaded.java\r\n+++ b/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestAsyncTableGetMultiThreaded.java\r\n@@ -98,7 +98,7 @@ public class TestAsyncTableGetMultiThreaded {\r\n TEST_UTIL.getConfiguration().set(CompactingMemStore.COMPACTING_MEMSTORE_TYPE_KEY,\r\n String.valueOf(memoryCompaction));\r\n\r\n- TEST_UTIL.startMiniCluster(5);\r\n+ TEST_UTIL.startMiniCluster(3);\r\n SPLIT_KEYS = new byte[8][];\r\n for (int i = 111; i < 999; i += 111) {\r\n SPLIT_KEYS[i / 111 - 1] = Bytes.toBytes(String.format(\"%03d\", i));\r\n@@ -134,7 +134,7 @@ public class TestAsyncTableGetMultiThreaded {\r\n @Test\r\n public void test() throws Exception {\r\n LOG.info(\"====== Test started ======\");\r\n- int numThreads = 10;\r\n+ int numThreads = 7;\r\n AtomicBoolean stop = new AtomicBoolean(false);\r\n ExecutorService executor =\r\n Executors.newFixedThreadPool(numThreads, Threads.newDaemonThreadFactory(\"TestAsyncGet-\"));\r\n{code}","created":"2020-04-09T18:18:50.112+0000"},{"body":"Hmm... It just failed branch-2 nightly because it could not write the xml file. Suspicious. I ran it locally under mission control and it uses less CPU and not much memory. Keeping an eye on it.","created":"2020-04-11T01:28:13.803+0000"},{"body":"And it just failed branch-2.3 #30 nightly with unable to write xml file.\r\n\r\nLet me push this so all branches are same. Will open new follow-on to track failed finish of test after get more info on it.","created":"2020-04-13T18:50:42.626+0000"},{"body":"Doubt this a fix. Pushing though so all branches are same.","created":"2020-04-13T18:52:49.474+0000"},{"body":"Failed last night with unable to write the xml:\r\n\r\nFailed to read test report file /home/jenkins/jenkins-slave/workspace/HBase_Nightly_branch-2/output-jdk8-hadoop3/archiver/hbase-server/target/surefire-reports/TEST-org.apache.hadoop.hbase.client.TestAsyncTableGetMultiThreaded.xml\r\norg.dom4j.DocumentException: Error on line 92 of document : XML document structures must start and end within the same entity. Nested exception: XML document structures must start and end within the same entity.","created":"2020-04-15T15:41:59.892+0000"},{"body":"Got this checking local runs...\r\n{code}\r\n2020-04-15 20:27:40,882 ERROR [RPCClient-NioEventLoopGroup-6-3] util.FutureUtils(70): Unexpected error caught when processing CompletableFuture\r\n java.lang.NullPointerException\r\n at org.apache.hadoop.hbase.client.AsyncRegionLocatorHelper.canUpdateOnError(AsyncRegionLocatorHelper.java:49)\r\n at org.apache.hadoop.hbase.client.AsyncRegionLocatorHelper.updateCachedLocationOnError(AsyncRegionLocatorHelper.java:61)\r\n at org.apache.hadoop.hbase.client.AsyncNonMetaRegionLocator.updateCachedLocationOnError(AsyncNonMetaRegionLocator.java:610)\r\n...\r\n{code}\r\n\r\nI added logging to the background threads that were doing continuous gets. In the thread dump output when test times out, I saw one or two of the background threads. Looking at where they were last going back through logs, in at least one case the region had just split, the locator couldn't find the region. The handling of the null lcoation generated the above which messed up the background thread... It got stuck.\r\n\r\nTesting patch locally....","created":"2020-04-16T06:10:36.693+0000"},{"body":"Reopening to apply addendum to address NPE above.","created":"2020-04-16T15:00:13.682+0000"},{"body":"Pushed below addendum to branch-2.2+\r\n\r\n{code}\r\ncommit 50f5cf1f94a10b14fd3def2aab495682128d1816 (HEAD -> 2.2, origin/branch-2.2)\r\nAuthor: stack \r\nDate: Thu Apr 16 08:02:33 2020 -0700\r\n\r\n HBASE-24158 [Flakey Tests] TestAsyncTableGetMultiThreaded\r\n Addendum to address NPE\r\n\r\ndiff --git a/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncRegionLocatorHelper.java b/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncRegionLocatorHelper.java\r\nindex 764650afba..4c6cd5a011 100644\r\n--- a/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncRegionLocatorHelper.java\r\n+++ b/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncRegionLocatorHelper.java\r\n@@ -1,4 +1,4 @@\r\n-/**\r\n+/*\r\n * Licensed to the Apache Software Foundation (ASF) under one\r\n * or more contributor license agreements. See the NOTICE file\r\n * distributed with this work for additional information\r\n@@ -19,7 +19,6 @@ package org.apache.hadoop.hbase.client;\r\n\r\n import static org.apache.hadoop.hbase.exceptions.ClientExceptionsUtil.findException;\r\n import static org.apache.hadoop.hbase.exceptions.ClientExceptionsUtil.isMetaClearingException;\r\n-\r\n import java.util.Arrays;\r\n import java.util.function.Consumer;\r\n import java.util.function.Function;\r\n@@ -45,7 +44,13 @@ final class AsyncRegionLocatorHelper {\r\n static boolean canUpdateOnError(HRegionLocation loc, HRegionLocation oldLoc) {\r\n // Do not need to update if no such location, or the location is newer, or the location is not\r\n // the same with us\r\n- return oldLoc != null && oldLoc.getSeqNum() <= loc.getSeqNum() &&\r\n+ if (loc == null || loc.getServerName() == null) {\r\n+ return false;\r\n+ }\r\n+ if (oldLoc == null || oldLoc.getServerName() == null) {\r\n+ return false;\r\n+ }\r\n+ return oldLoc.getSeqNum() <= loc.getSeqNum() &&\r\n oldLoc.getServerName().equals(loc.getServerName());\r\n }\r\n{code}","created":"2020-04-16T15:06:34.022+0000"},{"body":"Re-resolving after pushing addendum (see attached). Pushed to branch-2.2+ -- its a but in the AsyncRegionLocatorHelper.","created":"2020-04-16T15:08:52.255+0000"}],"conversations":[{"body":"I've already cut down the number of threads used by this test but it failed in nightly last night unable to close out its xml and locally it failed too in a run overnight. I ran it under harness and it seems well-behaved. It doesn't use much memory -- 700MB -- and thread counts are usual (~450). It does use near 100% CPU which is a little unusual. Otherwise, looks fine.\r\n\r\nLet me keep an eye on it. Could down the thread count more and use less processes... this makes it use less CPU. There does seems a bunch of overlap with tests done elsewhere.","from":"reporter","subject":"[Flakey Tests] TestAsyncTableGetMultiThreaded"},{"body":"Its an old test. It cam in w/ HBASE-16945. Back then it was running w/ 20 threads. Recently I downed it to 10. Will keep an eye on it.","from":"developer"},{"body":"Attached dumb patch that does nothing but use a few less resources; no notion if helps.","from":"developer"},{"body":"Hmm... Yeah, it failed in the last two nightly runs on branch-2. Let me push this patch on branch-2 and see if it helps.","from":"developer"},{"body":"Ok. Pushed below on branch-2. Lets see if helps.\r\n{code}\r\ncommit 2d117963809929e9523e47e94b7e5faab7db8e1a (HEAD -> 2)\r\nAuthor: stack \r\nDate: Thu Apr 9 11:03:22 2020 -0700\r\n\r\n HBASE-24158 [Flakey Tests] TestAsyncTableGetMultiThreaded\r\n\r\ndiff --git a/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestAsyncTableGetMultiThreaded.java b/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestAsyncTableGetMultiThreaded.java\r\nindex 7c896957fa..ce6bc05ca0 100644\r\n--- a/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestAsyncTableGetMultiThreaded.java\r\n+++ b/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestAsyncTableGetMultiThreaded.java\r\n@@ -98,7 +98,7 @@ public class TestAsyncTableGetMultiThreaded {\r\n TEST_UTIL.getConfiguration().set(CompactingMemStore.COMPACTING_MEMSTORE_TYPE_KEY,\r\n String.valueOf(memoryCompaction));\r\n\r\n- TEST_UTIL.startMiniCluster(5);\r\n+ TEST_UTIL.startMiniCluster(3);\r\n SPLIT_KEYS = new byte[8][];\r\n for (int i = 111; i < 999; i += 111) {\r\n SPLIT_KEYS[i / 111 - 1] = Bytes.toBytes(String.format(\"%03d\", i));\r\n@@ -134,7 +134,7 @@ public class TestAsyncTableGetMultiThreaded {\r\n @Test\r\n public void test() throws Exception {\r\n LOG.info(\"====== Test started ======\");\r\n- int numThreads = 10;\r\n+ int numThreads = 7;\r\n AtomicBoolean stop = new AtomicBoolean(false);\r\n ExecutorService executor =\r\n Executors.newFixedThreadPool(numThreads, Threads.newDaemonThreadFactory(\"TestAsyncGet-\"));\r\n{code}","from":"developer"},{"body":"Hmm... It just failed branch-2 nightly because it could not write the xml file. Suspicious. I ran it locally under mission control and it uses less CPU and not much memory. Keeping an eye on it.","from":"developer"},{"body":"And it just failed branch-2.3 #30 nightly with unable to write xml file.\r\n\r\nLet me push this so all branches are same. Will open new follow-on to track failed finish of test after get more info on it.","from":"developer"},{"body":"Doubt this a fix. Pushing though so all branches are same.","from":"developer"},{"body":"Failed last night with unable to write the xml:\r\n\r\nFailed to read test report file /home/jenkins/jenkins-slave/workspace/HBase_Nightly_branch-2/output-jdk8-hadoop3/archiver/hbase-server/target/surefire-reports/TEST-org.apache.hadoop.hbase.client.TestAsyncTableGetMultiThreaded.xml\r\norg.dom4j.DocumentException: Error on line 92 of document : XML document structures must start and end within the same entity. Nested exception: XML document structures must start and end within the same entity.","from":"developer"},{"body":"Got this checking local runs...\r\n{code}\r\n2020-04-15 20:27:40,882 ERROR [RPCClient-NioEventLoopGroup-6-3] util.FutureUtils(70): Unexpected error caught when processing CompletableFuture\r\n java.lang.NullPointerException\r\n at org.apache.hadoop.hbase.client.AsyncRegionLocatorHelper.canUpdateOnError(AsyncRegionLocatorHelper.java:49)\r\n at org.apache.hadoop.hbase.client.AsyncRegionLocatorHelper.updateCachedLocationOnError(AsyncRegionLocatorHelper.java:61)\r\n at org.apache.hadoop.hbase.client.AsyncNonMetaRegionLocator.updateCachedLocationOnError(AsyncNonMetaRegionLocator.java:610)\r\n...\r\n{code}\r\n\r\nI added logging to the background threads that were doing continuous gets. In the thread dump output when test times out, I saw one or two of the background threads. Looking at where they were last going back through logs, in at least one case the region had just split, the locator couldn't find the region. The handling of the null lcoation generated the above which messed up the background thread... It got stuck.\r\n\r\nTesting patch locally....","from":"developer"},{"body":"Reopening to apply addendum to address NPE above.","from":"developer"},{"body":"Pushed below addendum to branch-2.2+\r\n\r\n{code}\r\ncommit 50f5cf1f94a10b14fd3def2aab495682128d1816 (HEAD -> 2.2, origin/branch-2.2)\r\nAuthor: stack \r\nDate: Thu Apr 16 08:02:33 2020 -0700\r\n\r\n HBASE-24158 [Flakey Tests] TestAsyncTableGetMultiThreaded\r\n Addendum to address NPE\r\n\r\ndiff --git a/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncRegionLocatorHelper.java b/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncRegionLocatorHelper.java\r\nindex 764650afba..4c6cd5a011 100644\r\n--- a/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncRegionLocatorHelper.java\r\n+++ b/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncRegionLocatorHelper.java\r\n@@ -1,4 +1,4 @@\r\n-/**\r\n+/*\r\n * Licensed to the Apache Software Foundation (ASF) under one\r\n * or more contributor license agreements. See the NOTICE file\r\n * distributed with this work for additional information\r\n@@ -19,7 +19,6 @@ package org.apache.hadoop.hbase.client;\r\n\r\n import static org.apache.hadoop.hbase.exceptions.ClientExceptionsUtil.findException;\r\n import static org.apache.hadoop.hbase.exceptions.ClientExceptionsUtil.isMetaClearingException;\r\n-\r\n import java.util.Arrays;\r\n import java.util.function.Consumer;\r\n import java.util.function.Function;\r\n@@ -45,7 +44,13 @@ final class AsyncRegionLocatorHelper {\r\n static boolean canUpdateOnError(HRegionLocation loc, HRegionLocation oldLoc) {\r\n // Do not need to update if no such location, or the location is newer, or the location is not\r\n // the same with us\r\n- return oldLoc != null && oldLoc.getSeqNum() <= loc.getSeqNum() &&\r\n+ if (loc == null || loc.getServerName() == null) {\r\n+ return false;\r\n+ }\r\n+ if (oldLoc == null || oldLoc.getServerName() == null) {\r\n+ return false;\r\n+ }\r\n+ return oldLoc.getSeqNum() <= loc.getSeqNum() &&\r\n oldLoc.getServerName().equals(loc.getServerName());\r\n }\r\n{code}","from":"developer"},{"body":"Re-resolving after pushing addendum (see attached). Pushed to branch-2.2+ -- its a but in the AsyncRegionLocatorHelper.","from":"developer"}],"created":"2020-04-09T17:56:38.000+0000","description":"I've already cut down the number of threads used by this test but it failed in nightly last night unable to close out its xml and locally it failed too in a run overnight. I ran it under harness and it seems well-behaved. It doesn't use much memory -- 700MB -- and thread counts are usual (~450). It does use near 100% CPU which is a little unusual. Otherwise, looks fine.\r\n\r\nLet me keep an eye on it. Could down the thread count more and use less processes... this makes it use less CPU. There does seems a bunch of overlap with tests done elsewhere.","issue_id":"13297452","key":"HBASE-24158","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-04-16T15:08:52.000+0000","role":"fixed_distractor","summary":"[Flakey Tests] TestAsyncTableGetMultiThreaded"} {"case_id":"13298423","cluster":"DISTRACTOR-HBASE-24190","comments":[{"body":"Pushed to master, branch-2, 2.3, 2.2, 2.1, branch-1, 1.4. Thanks [~shahrs87] for the fix.","created":"2020-05-11T06:59:02.685+0000"},{"body":"The commits applied do not conform to the project requirements for including a Jira ticket and matching between the commit title and jira summary. Responsible committer, please revert and reapply everywhere. Thanks.","created":"2020-05-13T19:26:14.204+0000"},{"body":"Fixed Jira id issue in respective commits. Marking it resolved.","created":"2020-05-14T08:39:32.296+0000"}],"conversations":[{"body":"In hbase-20586 (https://issues.apache.org/jira/browse/HBASE-20586)\r\n\r\n(commit_sha: [https://github.com/apache/hbase/commit/cd61bcc0] )\r\n\r\nThe code added ([SyncTable.java|https://github.com/apache/hbase/commit/cd61bcc0#diff-d1b79635f33483bf6226609e91fd1cc3]) for the use of *hbase.security.authentication* is case-sensitive. So users setting it to “KERBEROS” won’t take effect. \r\n\r\n \r\n{code:java}\r\n private void initCredentialsForHBase(String zookeeper, Job job) throws IOException {\r\n   Configuration peerConf = HBaseConfiguration.createClusterConf(job.getConfiguration(), zookeeper);\r\n   if(peerConf.get(\"hbase.security.authentication\").equals(\"kerberos\")){\r\n TableMapReduceUtil.initCredentialsForCluster(job, peerConf);    }\r\n }\r\n{code}\r\n \r\n\r\nHowever, in current code base, other uses of *hbase.security.authentication* are all case-insensitive. For example in *MasterFileSystem.java.* \r\n\r\n \r\n{code:java}\r\npublic MasterFileSystem(Configuration conf) throws IOException{   \r\n ...   \r\n this.isSecurityEnabled = \"kerberos\".equalsIgnoreCase(conf.get(\"hbase.security.authentication\"));  \r\n ... \r\n}\r\n{code}\r\n \r\n\r\nThe doc in GitHub repo is also misleading (Giving upper-case value).\r\n{quote}As a distributed database, HBase must be able to authenticate users and HBase services across an untrusted network. Clients and HBase services are treated equivalently in terms of authentication (and this is the only time we will draw such a distinction).\r\n\r\nThere are currently three modes of authentication which are supported by HBase today via the configuration property {{hbase.security.authentication}}\r\n\r\n{{1.SIMPLE}}\r\n\r\n{{2.KERBROS}}\r\n\r\n{{3.TOKEN}}\r\n{quote}\r\nUsers may misconfigure the parameter because of the case-senstive problem.\r\n\r\n*How To Fix*\r\n\r\nUsing *eqaulsIgnoreCase* API consistently in every place when using *hbase.security.authentication* or make it clear in Doc.","from":"reporter","subject":"Make kerberos value of hbase.security.authentication property case insensitive"},{"body":"Pushed to master, branch-2, 2.3, 2.2, 2.1, branch-1, 1.4. Thanks [~shahrs87] for the fix.","from":"developer"},{"body":"The commits applied do not conform to the project requirements for including a Jira ticket and matching between the commit title and jira summary. Responsible committer, please revert and reapply everywhere. Thanks.","from":"developer"},{"body":"Fixed Jira id issue in respective commits. Marking it resolved.","from":"developer"}],"created":"2020-04-15T00:56:23.000+0000","description":"In hbase-20586 (https://issues.apache.org/jira/browse/HBASE-20586)\r\n\r\n(commit_sha: [https://github.com/apache/hbase/commit/cd61bcc0] )\r\n\r\nThe code added ([SyncTable.java|https://github.com/apache/hbase/commit/cd61bcc0#diff-d1b79635f33483bf6226609e91fd1cc3]) for the use of *hbase.security.authentication* is case-sensitive. So users setting it to “KERBEROS” won’t take effect. \r\n\r\n \r\n{code:java}\r\n private void initCredentialsForHBase(String zookeeper, Job job) throws IOException {\r\n   Configuration peerConf = HBaseConfiguration.createClusterConf(job.getConfiguration(), zookeeper);\r\n   if(peerConf.get(\"hbase.security.authentication\").equals(\"kerberos\")){\r\n TableMapReduceUtil.initCredentialsForCluster(job, peerConf);    }\r\n }\r\n{code}\r\n \r\n\r\nHowever, in current code base, other uses of *hbase.security.authentication* are all case-insensitive. For example in *MasterFileSystem.java.* \r\n\r\n \r\n{code:java}\r\npublic MasterFileSystem(Configuration conf) throws IOException{   \r\n ...   \r\n this.isSecurityEnabled = \"kerberos\".equalsIgnoreCase(conf.get(\"hbase.security.authentication\"));  \r\n ... \r\n}\r\n{code}\r\n \r\n\r\nThe doc in GitHub repo is also misleading (Giving upper-case value).\r\n{quote}As a distributed database, HBase must be able to authenticate users and HBase services across an untrusted network. Clients and HBase services are treated equivalently in terms of authentication (and this is the only time we will draw such a distinction).\r\n\r\nThere are currently three modes of authentication which are supported by HBase today via the configuration property {{hbase.security.authentication}}\r\n\r\n{{1.SIMPLE}}\r\n\r\n{{2.KERBROS}}\r\n\r\n{{3.TOKEN}}\r\n{quote}\r\nUsers may misconfigure the parameter because of the case-senstive problem.\r\n\r\n*How To Fix*\r\n\r\nUsing *eqaulsIgnoreCase* API consistently in every place when using *hbase.security.authentication* or make it clear in Doc.","issue_id":"13298423","key":"HBASE-24190","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-05-14T08:39:32.000+0000","role":"fixed_distractor","summary":"Make kerberos value of hbase.security.authentication property case insensitive"} {"case_id":"13300656","cluster":"DISTRACTOR-HBASE-24247","comments":[{"body":"Pushed to branch-2.3+. Thanks for review [~janh]\r\n","created":"2020-04-28T20:28:42.812+0000"},{"body":"This breaks TestAdminShell.\r\n\r\n{noformat}\r\nFailure: exception expected but none was thrown.\r\ntest_merge_regions(Hbase::AdminRegionTest)\r\nsrc/test/ruby/hbase/admin_test.rb:648:in `block in test_merge_regions'\r\n 645: command(:merge_region, region1,region1,region1)\r\n 646: end\r\n 647: # 3 non-adjacent regions without forcible=true\r\n => 648: assert_raise(RuntimeError) do\r\n 649: command(:merge_region, region1,region2,region4)\r\n 650: end\r\n 651: # 2 adjacent regions\r\n{noformat}","created":"2020-04-29T10:52:21.438+0000"},{"body":"Thank you for fingering it [~zhangduo] Let me fix.","created":"2020-04-29T15:44:10.459+0000"},{"body":"Re-resolved. Pushed an addendum on branch-2.3+. Set 'force' on merge that is scheduled by hbck2 fixMeta. Previous tried to be 'smart' about it but favor keeping current semantic/expectation where any attempt at merging non-adjacent regions requires 'force'.","created":"2020-04-29T21:40:08.086+0000"}],"conversations":[{"body":"Below is a multi-merge created by FixMeta provoked by 'hbck2 fixMeta'. The merge is legitimate in that indeed all Regions overlap. The merge is cutoff off at the current max of 10 Regions-at-a-time (which is another issue). The merge fails though because two Regions in the Set of Regions to merge are not adjacent when we do our pre-flight check. We could 'force' the merge but better if the 'check' is improved.\r\n\r\n{code}\r\n2020-04-22 22:04:57,048 WARN org.apache.hadoop.hbase.master.assignment.MergeTableRegionsProcedure: Unable to merge non-adjacent or non-overlapping regions 50b9f911320f64d0ab54a7606a6cdb77, 15877a8df3987176b12a2e2c4712c95f when force=false\r\n2020-04-22 22:04:57,048 WARN org.apache.hadoop.hbase.master.MetaFixer: Failed overlap fix of [{ENCODED => 6f880442573f4ca0c2536ce2352e4883, NAME => 'X,,1567882650838.6f880442573f4ca0c2536ce2352e4883.', STARTKEY => '', ENDKEY => '\\x01\\x02\\x05\\x01\\x03\\x02\\x01\\x01\\x01\\x01\\x02201908310200\\x00\\x00\\x048.1-11B117\\x00\\x00\\x00\\x00\\x00\\x00iPad4,1\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00'}, {ENCODED => 98af7f02916e014c07ac099724c7ffaf, NAME => 'X,\\x01\\x01\\x05\\x01\\x01,1558718898305.98af7f02916e014c07ac099724c7ffaf.', STARTKEY => '\\x01\\x01\\x05\\x01\\x01', ENDKEY => '\\x01\\x01\\x05\\x01\\x02'}, {ENCODED => d1a0b8772432c1148cdf7a8fad8e770c, NAME => 'X,\\x01\\x01\\x05\\x01\\x02,1558718898305.d1a0b8772432c1148cdf7a8fad8e770c.', STARTKEY => '\\x01\\x01\\x05\\x01\\x02', ENDKEY => '\\x01\\x01\\x05\\x01\\x03'}, {ENCODED => 99738e58d057dafb861116a3efcb0285, NAME => 'X,\\x01\\x01\\x05\\x01\\x03,1558718898305.99738e58d057dafb861116a3efcb0285.', STARTKEY => '\\x01\\x01\\x05\\x01\\x03', ENDKEY => '\\x01\\x01\\x05\\x02\\x01'}, {ENCODED => 50b9f911320f64d0ab54a7606a6cdb77, NAME => 'X,\\x01\\x01\\x05\\x02\\x01,1558718898305.50b9f911320f64d0ab54a7606a6cdb77.', STARTKEY => '\\x01\\x01\\x05\\x02\\x01', ENDKEY => '\\x01\\x01\\x05\\x02\\x02'}, {ENCODED => 15877a8df3987176b12a2e2c4712c95f, NAME => 'X,\\x01\\x01\\x05\\x02\\x03,1558718898305.15877a8df3987176b12a2e2c4712c95f.', STARTKEY => '\\x01\\x01\\x05\\x02\\x03', ENDKEY => '\\x01\\x01\\x06\\x01\\x01'}, {ENCODED => d5f0929fffbaec29ca99d4d0cd90c491, NAME => 'X,\\x01\\x01\\x06\\x01\\x01,1558718898305.d5f0929fffbaec29ca99d4d0cd90c491.', STARTKEY => '\\x01\\x01\\x06\\x01\\x01', ENDKEY => '\\x01\\x01\\x06\\x01\\x02'}, {ENCODED => 8d72ed0d1d635511a323abef7026ec4f, NAME => 'X,\\x01\\x01\\x06\\x01\\x03,1558718898305.8d72ed0d1d635511a323abef7026ec4f.', STARTKEY => '\\x01\\x01\\x06\\x01\\x03', ENDKEY => '\\x01\\x01\\x06\\x02\\x01'}, {ENCODED => 977f5a0e2f77a91531000d358f9a8eba, NAME => 'X,\\x01\\x01\\x06\\x02\\x01,1558718898305.977f5a0e2f77a91531000d358f9a8eba.', STARTKEY => '\\x01\\x01\\x06\\x02\\x01', ENDKEY => '\\x01\\x01\\x06\\x02\\x02'}, {ENCODED => 21cdc09d13ae1ecefc6531786229f2ec, NAME => 'X,\\x01\\x01\\x06\\x02\\x03,1558718898305.21cdc09d13ae1ecefc6531786229f2ec.', STARTKEY => '\\x01\\x01\\x06\\x02\\x03', ENDKEY => '\\x01\\x01\\x07\\x01\\x01'}]\r\norg.apache.hadoop.hbase.exceptions.MergeRegionException: Unable to merge non-adjacent or non-overlapping regions 50b9f911320f64d0ab54a7606a6cdb77, 15877a8df3987176b12a2e2c4712c95f when force=false\r\n at org.apache.hadoop.hbase.master.assignment.MergeTableRegionsProcedure.checkRegionsToMerge(MergeTableRegionsProcedure.java:140)\r\n at org.apache.hadoop.hbase.master.assignment.MergeTableRegionsProcedure.(MergeTableRegionsProcedure.java:105)\r\n at org.apache.hadoop.hbase.master.HMaster$2.run(HMaster.java:1961)\r\n at org.apache.hadoop.hbase.master.procedure.MasterProcedureUtil.submitProcedure(MasterProcedureUtil.java:134)\r\n at org.apache.hadoop.hbase.master.HMaster.mergeRegions(HMaster.java:1955)\r\n at org.apache.hadoop.hbase.master.MetaFixer.fixOverlaps(MetaFixer.java:221)\r\n at org.apache.hadoop.hbase.master.MetaFixer.fix(MetaFixer.java:77)\r\n at org.apache.hadoop.hbase.master.MasterRpcServices.fixMeta(MasterRpcServices.java:2649)\r\n at org.apache.hadoop.hbase.shaded.protobuf.generated.MasterProtos$HbckService$2.callBlockingMethod(MasterProtos.java)\r\n at org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:388)\r\n at org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:133)\r\n at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:338)\r\n at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:318)\r\n\r\n{code}","from":"reporter","subject":"Failed multi-merge because two regions not adjacent (legitimately)."},{"body":"Pushed to branch-2.3+. Thanks for review [~janh]\r\n","from":"developer"},{"body":"This breaks TestAdminShell.\r\n\r\n{noformat}\r\nFailure: exception expected but none was thrown.\r\ntest_merge_regions(Hbase::AdminRegionTest)\r\nsrc/test/ruby/hbase/admin_test.rb:648:in `block in test_merge_regions'\r\n 645: command(:merge_region, region1,region1,region1)\r\n 646: end\r\n 647: # 3 non-adjacent regions without forcible=true\r\n => 648: assert_raise(RuntimeError) do\r\n 649: command(:merge_region, region1,region2,region4)\r\n 650: end\r\n 651: # 2 adjacent regions\r\n{noformat}","from":"developer"},{"body":"Thank you for fingering it [~zhangduo] Let me fix.","from":"developer"},{"body":"Re-resolved. Pushed an addendum on branch-2.3+. Set 'force' on merge that is scheduled by hbck2 fixMeta. Previous tried to be 'smart' about it but favor keeping current semantic/expectation where any attempt at merging non-adjacent regions requires 'force'.","from":"developer"}],"created":"2020-04-23T21:05:32.000+0000","description":"Below is a multi-merge created by FixMeta provoked by 'hbck2 fixMeta'. The merge is legitimate in that indeed all Regions overlap. The merge is cutoff off at the current max of 10 Regions-at-a-time (which is another issue). The merge fails though because two Regions in the Set of Regions to merge are not adjacent when we do our pre-flight check. We could 'force' the merge but better if the 'check' is improved.\r\n\r\n{code}\r\n2020-04-22 22:04:57,048 WARN org.apache.hadoop.hbase.master.assignment.MergeTableRegionsProcedure: Unable to merge non-adjacent or non-overlapping regions 50b9f911320f64d0ab54a7606a6cdb77, 15877a8df3987176b12a2e2c4712c95f when force=false\r\n2020-04-22 22:04:57,048 WARN org.apache.hadoop.hbase.master.MetaFixer: Failed overlap fix of [{ENCODED => 6f880442573f4ca0c2536ce2352e4883, NAME => 'X,,1567882650838.6f880442573f4ca0c2536ce2352e4883.', STARTKEY => '', ENDKEY => '\\x01\\x02\\x05\\x01\\x03\\x02\\x01\\x01\\x01\\x01\\x02201908310200\\x00\\x00\\x048.1-11B117\\x00\\x00\\x00\\x00\\x00\\x00iPad4,1\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00'}, {ENCODED => 98af7f02916e014c07ac099724c7ffaf, NAME => 'X,\\x01\\x01\\x05\\x01\\x01,1558718898305.98af7f02916e014c07ac099724c7ffaf.', STARTKEY => '\\x01\\x01\\x05\\x01\\x01', ENDKEY => '\\x01\\x01\\x05\\x01\\x02'}, {ENCODED => d1a0b8772432c1148cdf7a8fad8e770c, NAME => 'X,\\x01\\x01\\x05\\x01\\x02,1558718898305.d1a0b8772432c1148cdf7a8fad8e770c.', STARTKEY => '\\x01\\x01\\x05\\x01\\x02', ENDKEY => '\\x01\\x01\\x05\\x01\\x03'}, {ENCODED => 99738e58d057dafb861116a3efcb0285, NAME => 'X,\\x01\\x01\\x05\\x01\\x03,1558718898305.99738e58d057dafb861116a3efcb0285.', STARTKEY => '\\x01\\x01\\x05\\x01\\x03', ENDKEY => '\\x01\\x01\\x05\\x02\\x01'}, {ENCODED => 50b9f911320f64d0ab54a7606a6cdb77, NAME => 'X,\\x01\\x01\\x05\\x02\\x01,1558718898305.50b9f911320f64d0ab54a7606a6cdb77.', STARTKEY => '\\x01\\x01\\x05\\x02\\x01', ENDKEY => '\\x01\\x01\\x05\\x02\\x02'}, {ENCODED => 15877a8df3987176b12a2e2c4712c95f, NAME => 'X,\\x01\\x01\\x05\\x02\\x03,1558718898305.15877a8df3987176b12a2e2c4712c95f.', STARTKEY => '\\x01\\x01\\x05\\x02\\x03', ENDKEY => '\\x01\\x01\\x06\\x01\\x01'}, {ENCODED => d5f0929fffbaec29ca99d4d0cd90c491, NAME => 'X,\\x01\\x01\\x06\\x01\\x01,1558718898305.d5f0929fffbaec29ca99d4d0cd90c491.', STARTKEY => '\\x01\\x01\\x06\\x01\\x01', ENDKEY => '\\x01\\x01\\x06\\x01\\x02'}, {ENCODED => 8d72ed0d1d635511a323abef7026ec4f, NAME => 'X,\\x01\\x01\\x06\\x01\\x03,1558718898305.8d72ed0d1d635511a323abef7026ec4f.', STARTKEY => '\\x01\\x01\\x06\\x01\\x03', ENDKEY => '\\x01\\x01\\x06\\x02\\x01'}, {ENCODED => 977f5a0e2f77a91531000d358f9a8eba, NAME => 'X,\\x01\\x01\\x06\\x02\\x01,1558718898305.977f5a0e2f77a91531000d358f9a8eba.', STARTKEY => '\\x01\\x01\\x06\\x02\\x01', ENDKEY => '\\x01\\x01\\x06\\x02\\x02'}, {ENCODED => 21cdc09d13ae1ecefc6531786229f2ec, NAME => 'X,\\x01\\x01\\x06\\x02\\x03,1558718898305.21cdc09d13ae1ecefc6531786229f2ec.', STARTKEY => '\\x01\\x01\\x06\\x02\\x03', ENDKEY => '\\x01\\x01\\x07\\x01\\x01'}]\r\norg.apache.hadoop.hbase.exceptions.MergeRegionException: Unable to merge non-adjacent or non-overlapping regions 50b9f911320f64d0ab54a7606a6cdb77, 15877a8df3987176b12a2e2c4712c95f when force=false\r\n at org.apache.hadoop.hbase.master.assignment.MergeTableRegionsProcedure.checkRegionsToMerge(MergeTableRegionsProcedure.java:140)\r\n at org.apache.hadoop.hbase.master.assignment.MergeTableRegionsProcedure.(MergeTableRegionsProcedure.java:105)\r\n at org.apache.hadoop.hbase.master.HMaster$2.run(HMaster.java:1961)\r\n at org.apache.hadoop.hbase.master.procedure.MasterProcedureUtil.submitProcedure(MasterProcedureUtil.java:134)\r\n at org.apache.hadoop.hbase.master.HMaster.mergeRegions(HMaster.java:1955)\r\n at org.apache.hadoop.hbase.master.MetaFixer.fixOverlaps(MetaFixer.java:221)\r\n at org.apache.hadoop.hbase.master.MetaFixer.fix(MetaFixer.java:77)\r\n at org.apache.hadoop.hbase.master.MasterRpcServices.fixMeta(MasterRpcServices.java:2649)\r\n at org.apache.hadoop.hbase.shaded.protobuf.generated.MasterProtos$HbckService$2.callBlockingMethod(MasterProtos.java)\r\n at org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:388)\r\n at org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:133)\r\n at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:338)\r\n at org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:318)\r\n\r\n{code}","issue_id":"13300656","key":"HBASE-24247","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-04-29T21:40:08.000+0000","role":"fixed_distractor","summary":"Failed multi-merge because two regions not adjacent (legitimately)."} {"case_id":"13305078","cluster":"DISTRACTOR-HBASE-24376","comments":[{"body":"FYI [~ndimiduk], kind of understand the root cause, will put a patch tomorrow.","created":"2020-05-15T00:52:27.641+0000"},{"body":"Once you're settled on the fix, mind updating the Jira description to outline what the problem was? Thanks!","created":"2020-05-18T22:07:56.066+0000"},{"body":"[~ndimiduk], I just updated the root cause in the description, thanks for review.","created":"2020-05-19T16:54:59.178+0000"},{"body":"Hi [~zghao], do you want this Jira in 2.2? If so, I will do a backport to 2.2 branch, thanks.","created":"2020-05-25T23:11:39.227+0000"},{"body":"Resolve it for now.","created":"2020-05-25T23:13:44.862+0000"},{"body":"{quote}bq. do you want this Jira in 2.2? If so, I will do a backport to 2.2 branch, thanks.\r\n{quote}\r\nYes. Thanks.","created":"2020-05-26T01:29:21.615+0000"},{"body":"Will post a backport for 2.2.","created":"2020-05-26T17:23:25.169+0000"},{"body":"I faced with the issue in branch-1 (HBase 1.2.0-cdh5.16.2):\r\n\r\n2020-07-02 16:41:18,337 INFO org.apache.hadoop.hbase.master.normalizer.MergeNormalizationPlan: Executing merging normalization plan: MergeNormalizationPlan{firstRegion={ENC\r\nODED => feba09265266f1c3c090bc42cc90becc, NAME => 'tableName,Aw-BEZ0JD4M3HrvA4Yks,1593695478557.feba09265266f1c3c090bc42cc90becc.', STARTKEY => 'Aw-BEZ0JD4M3HrvA4Yks', EN\r\nDKEY => 'D0sYMT716R0tyHPGk8ii'}, secondRegion={ENCODED => 8003ecbf849c4f5e27bf5956ec0729cc, NAME => 'TableName,B_zjT044PCvwQ4I53Q5m,1593695479990.8003ecbf849c4f5e27bf5956\r\nec0729cc.', STARTKEY => 'B_zjT044PCvwQ4I53Q5m', ENDKEY => 'FCzjFZ4Vhb0hpVtn6VxP'}}\r\n\r\n \r\n\r\n ","created":"2020-07-06T14:46:13.409+0000"},{"body":"Hi [~ruslan.sabitov], last time I checked, for branch-1, it is ok. In your case, these two regions are overlapping with each other, so it is ok. It seems that there is already overlaps which means the table is not at a consistent state.","created":"2020-07-15T16:41:36.126+0000"},{"body":"[~huaxiangsun] thank you for your reply. Here is another example:\r\n\r\n2020-06-04 17:31:59,611 INFO org.apache.hadoop.hbase.master.normalizer.MergeNormalizationPlan: Executing merging normalization plan: MergeNormalizationPlan\\{firstRegion={ENCODED => 5d2d90c3697145c15bc2dc1841bbd140, NAME => 'tableName,-7-7d--TPxHgzWi3x6Nw,1591257734317.5d2d90c3697145c15bc2dc1841bbd140.', STARTKEY => '-7-7d--TPxHgzWi3x6Nw', ENDKEY => '-F--'}, secondRegion=\\{ENCODED => 5efe3811574f9f390d4cc8fc6099aa18, NAME => 'tableName,-Myro7PP4ncDkwLSQqVv,1591257751881.5efe3811574f9f390d4cc8fc6099aa18.', STARTKEY => '-Myro7PP4ncDkwLSQqVv', ENDKEY => '-V–'}}\r\n\r\n \r\n\r\nI parsed HBase Master log and collected all lines with text MergeNormalizationPlan and where r1 ENDKEY is not equal r2 STARTKEY:\r\n\r\n[https://pastebin.com/5wqr28br]\r\n\r\n \r\n\r\nI'd like to add this table becomes inconsistent each time I enable normalization for the table. Do you have any ideas how to check that the problem in the normalization?","created":"2020-07-16T06:03:17.308+0000"},{"body":"I think in hbase-1, the normalizer uses a force flag to do the merge, which means that if the two regions are not next to each other, it will still merge them.\r\n\r\nAre you sure that inconsistency is caused by normalizer? I.e, before normalizer run, table is consistent, after normalizer run, there is inconsistency. If that is the case, the issue is that normalizer merges two non-adjacent regions, which will cause overlaps. \r\n\r\nThere is one such issue with hbase-2, but I checked the code, hbase-1 seems ok. \r\n\r\nYou can go over the master log, dump out meta table, and inconsistency report from hbck to check if that is the case.","created":"2020-07-16T17:02:54.362+0000"},{"body":"[~ruslan.sabitov] if you have more concerns on this topic, can you please take them to the user@ list? We try to keep JIRA (and GitHub) focused on topics of active development. Thanks.","created":"2020-07-20T22:48:08.953+0000"}],"conversations":[{"body":"Currently, we found normalizer was merging regions which are non-adjacent, it will cause inconsistencies in the cluster.\r\n{code:java}\r\n439055 2020-05-08 17:47:09,814 INFO org.apache.hadoop.hbase.master.normalizer.MergeNormalizationPlan: Executing merging normalization plan: MergeNormalizationPlan{firstRegion={ENCODED => 47fe236a5e3649ded95cb64ad0c08492, NAME => 'TABLE,\\x03\\x01\\x05\\x01\\x04\\x02,1554838974870.47fe236a5e3649ded95cb64ad 0c08492.', STARTKEY => '\\x03\\x01\\x05\\x01\\x04\\x02', ENDKEY => '\\x03\\x01\\x05\\x01\\x04\\x02\\x01\\x02\\x02201904082200\\x00\\x00\\x03Mac\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00iMac13,1\\x00\\x00\\x00\\x00\\x00\\x049.3-14E260\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x02\\x05'}, secondRegion={ENCODED => 0c0f2aa67f4329d5c4 8ba0320f173d31, NAME => 'TABLE,\\x03\\x01\\x05\\x02\\x01\\x01,1554830735526.0c0f2aa67f4329d5c48ba0320f173d31.', STARTKEY => '\\x03\\x01\\x05\\x02\\x01\\x01', ENDKEY => '\\x03\\x01\\x05\\x02\\x01\\x02'}}\r\n439056 2020-05-08 17:47:11,438 INFO org.apache.hadoop.hbase.ScheduledChore: CatalogJanitor-*****:16000 average execution time: 1676219193 ns.\r\n439057 2020-05-08 17:47:11,730 INFO org.apache.hadoop.hbase.master.HMaster: Client=null/null merge regions [47fe236a5e3649ded95cb64ad0c08492], [0c0f2aa67f4329d5c48ba0320f173d31]\r\n {code}\r\n \r\n\r\nThe root cause is that getMergeNormalizationPlan() uses a list of regionInfo which is ordered by regionName. regionName does not necessary guarantee the order of STARTKEY (let's say 'aa1', 'aa1!', in order of regionName, it will be 'aa1!' followed by 'aa1'. This will result in normalizer merging non-adjacent regions into one and creates overlaps. This is not an issue in branch-1 as the list is already ordered by RegionInfo.COMPARATOR in normalizer.\r\n\r\n ","from":"reporter","subject":"MergeNormalizer is merging non-adjacent regions and causing region overlaps/holes."},{"body":"FYI [~ndimiduk], kind of understand the root cause, will put a patch tomorrow.","from":"developer"},{"body":"Once you're settled on the fix, mind updating the Jira description to outline what the problem was? Thanks!","from":"developer"},{"body":"[~ndimiduk], I just updated the root cause in the description, thanks for review.","from":"developer"},{"body":"Hi [~zghao], do you want this Jira in 2.2? If so, I will do a backport to 2.2 branch, thanks.","from":"developer"},{"body":"Resolve it for now.","from":"developer"},{"body":"{quote}bq. do you want this Jira in 2.2? If so, I will do a backport to 2.2 branch, thanks.\r\n{quote}\r\nYes. Thanks.","from":"developer"},{"body":"Will post a backport for 2.2.","from":"developer"},{"body":"I faced with the issue in branch-1 (HBase 1.2.0-cdh5.16.2):\r\n\r\n2020-07-02 16:41:18,337 INFO org.apache.hadoop.hbase.master.normalizer.MergeNormalizationPlan: Executing merging normalization plan: MergeNormalizationPlan{firstRegion={ENC\r\nODED => feba09265266f1c3c090bc42cc90becc, NAME => 'tableName,Aw-BEZ0JD4M3HrvA4Yks,1593695478557.feba09265266f1c3c090bc42cc90becc.', STARTKEY => 'Aw-BEZ0JD4M3HrvA4Yks', EN\r\nDKEY => 'D0sYMT716R0tyHPGk8ii'}, secondRegion={ENCODED => 8003ecbf849c4f5e27bf5956ec0729cc, NAME => 'TableName,B_zjT044PCvwQ4I53Q5m,1593695479990.8003ecbf849c4f5e27bf5956\r\nec0729cc.', STARTKEY => 'B_zjT044PCvwQ4I53Q5m', ENDKEY => 'FCzjFZ4Vhb0hpVtn6VxP'}}\r\n\r\n \r\n\r\n ","from":"developer"},{"body":"Hi [~ruslan.sabitov], last time I checked, for branch-1, it is ok. In your case, these two regions are overlapping with each other, so it is ok. It seems that there is already overlaps which means the table is not at a consistent state.","from":"developer"},{"body":"[~huaxiangsun] thank you for your reply. Here is another example:\r\n\r\n2020-06-04 17:31:59,611 INFO org.apache.hadoop.hbase.master.normalizer.MergeNormalizationPlan: Executing merging normalization plan: MergeNormalizationPlan\\{firstRegion={ENCODED => 5d2d90c3697145c15bc2dc1841bbd140, NAME => 'tableName,-7-7d--TPxHgzWi3x6Nw,1591257734317.5d2d90c3697145c15bc2dc1841bbd140.', STARTKEY => '-7-7d--TPxHgzWi3x6Nw', ENDKEY => '-F--'}, secondRegion=\\{ENCODED => 5efe3811574f9f390d4cc8fc6099aa18, NAME => 'tableName,-Myro7PP4ncDkwLSQqVv,1591257751881.5efe3811574f9f390d4cc8fc6099aa18.', STARTKEY => '-Myro7PP4ncDkwLSQqVv', ENDKEY => '-V–'}}\r\n\r\n \r\n\r\nI parsed HBase Master log and collected all lines with text MergeNormalizationPlan and where r1 ENDKEY is not equal r2 STARTKEY:\r\n\r\n[https://pastebin.com/5wqr28br]\r\n\r\n \r\n\r\nI'd like to add this table becomes inconsistent each time I enable normalization for the table. Do you have any ideas how to check that the problem in the normalization?","from":"developer"},{"body":"I think in hbase-1, the normalizer uses a force flag to do the merge, which means that if the two regions are not next to each other, it will still merge them.\r\n\r\nAre you sure that inconsistency is caused by normalizer? I.e, before normalizer run, table is consistent, after normalizer run, there is inconsistency. If that is the case, the issue is that normalizer merges two non-adjacent regions, which will cause overlaps. \r\n\r\nThere is one such issue with hbase-2, but I checked the code, hbase-1 seems ok. \r\n\r\nYou can go over the master log, dump out meta table, and inconsistency report from hbck to check if that is the case.","from":"developer"},{"body":"[~ruslan.sabitov] if you have more concerns on this topic, can you please take them to the user@ list? We try to keep JIRA (and GitHub) focused on topics of active development. Thanks.","from":"developer"}],"created":"2020-05-15T00:48:42.000+0000","description":"Currently, we found normalizer was merging regions which are non-adjacent, it will cause inconsistencies in the cluster.\r\n{code:java}\r\n439055 2020-05-08 17:47:09,814 INFO org.apache.hadoop.hbase.master.normalizer.MergeNormalizationPlan: Executing merging normalization plan: MergeNormalizationPlan{firstRegion={ENCODED => 47fe236a5e3649ded95cb64ad0c08492, NAME => 'TABLE,\\x03\\x01\\x05\\x01\\x04\\x02,1554838974870.47fe236a5e3649ded95cb64ad 0c08492.', STARTKEY => '\\x03\\x01\\x05\\x01\\x04\\x02', ENDKEY => '\\x03\\x01\\x05\\x01\\x04\\x02\\x01\\x02\\x02201904082200\\x00\\x00\\x03Mac\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00iMac13,1\\x00\\x00\\x00\\x00\\x00\\x049.3-14E260\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x02\\x05'}, secondRegion={ENCODED => 0c0f2aa67f4329d5c4 8ba0320f173d31, NAME => 'TABLE,\\x03\\x01\\x05\\x02\\x01\\x01,1554830735526.0c0f2aa67f4329d5c48ba0320f173d31.', STARTKEY => '\\x03\\x01\\x05\\x02\\x01\\x01', ENDKEY => '\\x03\\x01\\x05\\x02\\x01\\x02'}}\r\n439056 2020-05-08 17:47:11,438 INFO org.apache.hadoop.hbase.ScheduledChore: CatalogJanitor-*****:16000 average execution time: 1676219193 ns.\r\n439057 2020-05-08 17:47:11,730 INFO org.apache.hadoop.hbase.master.HMaster: Client=null/null merge regions [47fe236a5e3649ded95cb64ad0c08492], [0c0f2aa67f4329d5c48ba0320f173d31]\r\n {code}\r\n \r\n\r\nThe root cause is that getMergeNormalizationPlan() uses a list of regionInfo which is ordered by regionName. regionName does not necessary guarantee the order of STARTKEY (let's say 'aa1', 'aa1!', in order of regionName, it will be 'aa1!' followed by 'aa1'. This will result in normalizer merging non-adjacent regions into one and creates overlaps. This is not an issue in branch-1 as the list is already ordered by RegionInfo.COMPARATOR in normalizer.\r\n\r\n ","issue_id":"13305078","key":"HBASE-24376","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-05-25T23:13:59.000+0000","role":"fixed_distractor","summary":"MergeNormalizer is merging non-adjacent regions and causing region overlaps/holes."} {"case_id":"13312320","cluster":"DISTRACTOR-HBASE-24588","comments":[{"body":"{quote}the calling context should handle the resolution of those futures in a single place.\r\n{quote}\r\nHMaster should call future.get() for all tables after plans are submitted for all tables. That might be better right? And if polling throws any Exception, we log it only at one place: HMaster(caller) and we maintain a count say normalizationPlanFailureCount and if it is helpful, maybe expose it in metric? Thought?","created":"2020-06-19T08:04:41.386+0000"},{"body":"[~ndimiduk] If you already had some patch or started working on this, please let me know, I will unassign myself. If not, let me try to work on this and see if I can make a patch ready for review.","created":"2020-06-19T08:08:10.932+0000"},{"body":"Thank you [~ndimiduk] for the review.\r\n\r\nI believe we can have this in 2.3, but I am fine with 2.3.1 also, you have already taken so many such last moment Jiras and I don't want to give more headache :)\r\n\r\nSo let me keep this till branch-2 for now?","created":"2020-06-26T16:45:38.794+0000"},{"body":"This is a good fix as the existing behavior is broken. +1 for branch-2.3 and 2.3.0rc1.","created":"2020-06-26T17:11:06.596+0000"},{"body":"True, let me create two backport PRs for branch-2 and branch-2.3. After QA build results, we can merge them.","created":"2020-06-26T17:14:20.576+0000"}],"conversations":[{"body":"I left a comment on a merged [commit|https://github.com/apache/hbase/commit/5d0e0fc5fd09bddb2d766d1e24e28e472961f454#r39987289] where a little discussion has blossomed. Right now the normalizer produces two types of plans: \"split\" or \"merge\". The master receives the list of plans and executes them in a simple loop. The way it does the actual execution is delegated off to the plan implementation.\r\n\r\nThe bug I noticed is that the two implementations are subtly different. Both use async APIs to submit procedures, but \"split\" blocks on completion while \"merge\" does not. Furthermore, because \"split\" blocks, it's able to capture any exception that's thrown, while \"merge\" cannot.\r\n\r\nThese implementations should be made consistent. My thinking at the moment is this {{execute}} method should instead be named {{submit}}, creating and returning the {{Future}} that represents whatever work it submitted, and the calling context should handle the resolution of those futures in a single place.","from":"reporter","subject":"Normalizer plan execution is not consistent between plan types"},{"body":"{quote}the calling context should handle the resolution of those futures in a single place.\r\n{quote}\r\nHMaster should call future.get() for all tables after plans are submitted for all tables. That might be better right? And if polling throws any Exception, we log it only at one place: HMaster(caller) and we maintain a count say normalizationPlanFailureCount and if it is helpful, maybe expose it in metric? Thought?","from":"developer"},{"body":"[~ndimiduk] If you already had some patch or started working on this, please let me know, I will unassign myself. If not, let me try to work on this and see if I can make a patch ready for review.","from":"developer"},{"body":"Thank you [~ndimiduk] for the review.\r\n\r\nI believe we can have this in 2.3, but I am fine with 2.3.1 also, you have already taken so many such last moment Jiras and I don't want to give more headache :)\r\n\r\nSo let me keep this till branch-2 for now?","from":"developer"},{"body":"This is a good fix as the existing behavior is broken. +1 for branch-2.3 and 2.3.0rc1.","from":"developer"},{"body":"True, let me create two backport PRs for branch-2 and branch-2.3. After QA build results, we can merge them.","from":"developer"}],"created":"2020-06-18T23:50:57.000+0000","description":"I left a comment on a merged [commit|https://github.com/apache/hbase/commit/5d0e0fc5fd09bddb2d766d1e24e28e472961f454#r39987289] where a little discussion has blossomed. Right now the normalizer produces two types of plans: \"split\" or \"merge\". The master receives the list of plans and executes them in a simple loop. The way it does the actual execution is delegated off to the plan implementation.\r\n\r\nThe bug I noticed is that the two implementations are subtly different. Both use async APIs to submit procedures, but \"split\" blocks on completion while \"merge\" does not. Furthermore, because \"split\" blocks, it's able to capture any exception that's thrown, while \"merge\" cannot.\r\n\r\nThese implementations should be made consistent. My thinking at the moment is this {{execute}} method should instead be named {{submit}}, creating and returning the {{Future}} that represents whatever work it submitted, and the calling context should handle the resolution of those futures in a single place.","issue_id":"13312320","key":"HBASE-24588","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-06-27T19:46:37.000+0000","role":"fixed_distractor","summary":"Normalizer plan execution is not consistent between plan types"} {"case_id":"13312917","cluster":"DISTRACTOR-HBASE-24615","comments":[{"body":"ping [~dmanning] and [~shahrs87]","created":"2020-07-06T06:21:43.929+0000"},{"body":"[~wenfeiyi666] just fyi you dont have to create a separate PR for branch-2. You just need to have a PR for master branch and the committer will try to backport to all the other branches. If the rebase work is more then he/she will let you know to create another PR for those branch. Thank you !","created":"2020-07-06T16:48:31.611+0000"},{"body":"thanks, I go it.","created":"2020-07-07T01:54:27.966+0000"},{"body":"ping [~shahrs87]","created":"2020-07-13T06:03:33.264+0000"},{"body":"[~wenfeiyi666] Approved and requested one of the committer [~vjasani] to help review and commit the patch. Thank you [~wenfeiyi666] for the contribution.","created":"2020-07-14T15:49:46.825+0000"},{"body":"[~wenfeiyi666] can you please raise PR for branch-2 and branch-1? Although backports are directly applied to branch-2.x, but TestMutableRangeHistogram is not able to import MutableSizeHistogram due to dependency issue.\r\n\r\nThanks for the patch, it is already merged in master branch.","created":"2020-07-14T18:58:08.635+0000"},{"body":"ok, no problem","created":"2020-07-15T02:08:30.159+0000"},{"body":"branch-1 branch-1.3 branch-1.4 exists the bug, branch-1.0 branch-1.1 branch-1.2 not exists.","created":"2020-07-15T02:49:07.352+0000"},{"body":"[~zghao] are you fine with this patch landing on branch-2.2 as I see 2.2.6 RC preparation going on?","created":"2020-07-15T10:22:56.010+0000"},{"body":"For now, I have merged changes to master, branch-2, 2.3, branch-1.","created":"2020-07-15T11:13:12.205+0000"},{"body":"{code:java}\r\n long val = snapshot.getCount();\r\n if (val - cumNum > 0) {\r\n metricsRecordBuilder.addCounter(\r\n Interns.info(name + \"_\" + rangeType + \"_\" + ranges[ranges.length - 1] + \"-inf\", desc),\r\n val - cumNum);\r\n }\r\n{code}\r\nI am not fully understand the fix. I thought the last bucket was handled in previous code, too?","created":"2020-07-16T00:20:13.721+0000"},{"body":"[~zghao] The snippet you shared is for overflow bucket (for the count that falls outside of ranges[length-1])\r\nIn this jira we fixed the case where we are not updating the count between the range: ranges[length-2] and ranges[length-1]\r\nHope it makes sense. ","created":"2020-07-16T00:28:04.956+0000"},{"body":"{quote}The snippet you shared is for overflow bucket (for the count that falls outside of ranges[length-1])\r\nIn this jira we fixed the case where we are not updating the count between the range: ranges[length-2] and ranges[length-1]\r\n{quote}\r\n \r\n\r\nGot it. +1 for branch-2.2. Let me cherry-pick it.","created":"2020-07-16T00:36:55.818+0000"}],"conversations":[{"body":"We are not processing the distribution for last bucket. \r\n\r\nhttps://github.com/apache/hbase/blob/master/hbase-hadoop-compat/src/main/java/org/apache/hadoop/metrics2/lib/MutableRangeHistogram.java#L70\r\n\r\n{code:java}\r\n public void updateSnapshotRangeMetrics(MetricsRecordBuilder metricsRecordBuilder,\r\n Snapshot snapshot) {\r\n long priorRange = 0;\r\n long cumNum = 0;\r\n\r\n final long[] ranges = getRanges();\r\n final String rangeType = getRangeType();\r\n for (int i = 0; i < ranges.length - 1; i++) { -----> The bug lies here. We are not processing last bucket.\r\n long val = snapshot.getCountAtOrBelow(ranges[i]);\r\n if (val - cumNum > 0) {\r\n metricsRecordBuilder.addCounter(\r\n Interns.info(name + \"_\" + rangeType + \"_\" + priorRange + \"-\" + ranges[i], desc),\r\n val - cumNum);\r\n }\r\n priorRange = ranges[i];\r\n cumNum = val;\r\n }\r\n long val = snapshot.getCount();\r\n if (val - cumNum > 0) {\r\n metricsRecordBuilder.addCounter(\r\n Interns.info(name + \"_\" + rangeType + \"_\" + ranges[ranges.length - 1] + \"-inf\", desc),\r\n val - cumNum);\r\n }\r\n }\r\n{code}\r\n","from":"reporter","subject":"MutableRangeHistogram#updateSnapshotRangeMetrics doesn't calculate the distribution for last bucket."},{"body":"ping [~dmanning] and [~shahrs87]","from":"developer"},{"body":"[~wenfeiyi666] just fyi you dont have to create a separate PR for branch-2. You just need to have a PR for master branch and the committer will try to backport to all the other branches. If the rebase work is more then he/she will let you know to create another PR for those branch. Thank you !","from":"developer"},{"body":"thanks, I go it.","from":"developer"},{"body":"ping [~shahrs87]","from":"developer"},{"body":"[~wenfeiyi666] Approved and requested one of the committer [~vjasani] to help review and commit the patch. Thank you [~wenfeiyi666] for the contribution.","from":"developer"},{"body":"[~wenfeiyi666] can you please raise PR for branch-2 and branch-1? Although backports are directly applied to branch-2.x, but TestMutableRangeHistogram is not able to import MutableSizeHistogram due to dependency issue.\r\n\r\nThanks for the patch, it is already merged in master branch.","from":"developer"},{"body":"ok, no problem","from":"developer"},{"body":"branch-1 branch-1.3 branch-1.4 exists the bug, branch-1.0 branch-1.1 branch-1.2 not exists.","from":"developer"},{"body":"[~zghao] are you fine with this patch landing on branch-2.2 as I see 2.2.6 RC preparation going on?","from":"developer"},{"body":"For now, I have merged changes to master, branch-2, 2.3, branch-1.","from":"developer"},{"body":"{code:java}\r\n long val = snapshot.getCount();\r\n if (val - cumNum > 0) {\r\n metricsRecordBuilder.addCounter(\r\n Interns.info(name + \"_\" + rangeType + \"_\" + ranges[ranges.length - 1] + \"-inf\", desc),\r\n val - cumNum);\r\n }\r\n{code}\r\nI am not fully understand the fix. I thought the last bucket was handled in previous code, too?","from":"developer"},{"body":"[~zghao] The snippet you shared is for overflow bucket (for the count that falls outside of ranges[length-1])\r\nIn this jira we fixed the case where we are not updating the count between the range: ranges[length-2] and ranges[length-1]\r\nHope it makes sense. ","from":"developer"},{"body":"{quote}The snippet you shared is for overflow bucket (for the count that falls outside of ranges[length-1])\r\nIn this jira we fixed the case where we are not updating the count between the range: ranges[length-2] and ranges[length-1]\r\n{quote}\r\n \r\n\r\nGot it. +1 for branch-2.2. Let me cherry-pick it.","from":"developer"}],"created":"2020-06-22T22:24:16.000+0000","description":"We are not processing the distribution for last bucket. \r\n\r\nhttps://github.com/apache/hbase/blob/master/hbase-hadoop-compat/src/main/java/org/apache/hadoop/metrics2/lib/MutableRangeHistogram.java#L70\r\n\r\n{code:java}\r\n public void updateSnapshotRangeMetrics(MetricsRecordBuilder metricsRecordBuilder,\r\n Snapshot snapshot) {\r\n long priorRange = 0;\r\n long cumNum = 0;\r\n\r\n final long[] ranges = getRanges();\r\n final String rangeType = getRangeType();\r\n for (int i = 0; i < ranges.length - 1; i++) { -----> The bug lies here. We are not processing last bucket.\r\n long val = snapshot.getCountAtOrBelow(ranges[i]);\r\n if (val - cumNum > 0) {\r\n metricsRecordBuilder.addCounter(\r\n Interns.info(name + \"_\" + rangeType + \"_\" + priorRange + \"-\" + ranges[i], desc),\r\n val - cumNum);\r\n }\r\n priorRange = ranges[i];\r\n cumNum = val;\r\n }\r\n long val = snapshot.getCount();\r\n if (val - cumNum > 0) {\r\n metricsRecordBuilder.addCounter(\r\n Interns.info(name + \"_\" + rangeType + \"_\" + ranges[ranges.length - 1] + \"-inf\", desc),\r\n val - cumNum);\r\n }\r\n }\r\n{code}\r\n","issue_id":"13312917","key":"HBASE-24615","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-07-15T11:14:10.000+0000","role":"fixed_distractor","summary":"MutableRangeHistogram#updateSnapshotRangeMetrics doesn't calculate the distribution for last bucket."} {"case_id":"13314384","cluster":"DISTRACTOR-HBASE-24661","comments":[{"body":"[~stack] is aware of this issue, the addendum seems causing the issue. It passes without this addendum.\r\n{code:java}\r\ncommit bfec964aef09a6474d35b379fb4000cf3e8f86cc\r\nAuthor: stack \r\nDate:   Mon Jun 29 14:19:44 2020 -0700\r\n\r\n\r\n    HBASE-24648 Remove the legacy 'forceSplit' related code at region server side (#1990)\r\n    Addendum to fix TestHeapSize\r\n\r\n {code}","created":"2020-06-30T22:46:52.347+0000"},{"body":"Addendum by me broke this [~huaxiangsun]?\r\n\r\n\r\ncommit bfec964aef09a6474d35b379fb4000cf3e8f86cc\r\nAuthor: stack \r\nDate: Mon Jun 29 14:19:44 2020 -0700\r\n HBASE-24648 Remove the legacy 'forceSplit' related code at region server side (#1990)\r\n Addendum to fix TestHeapSize\r\n\r\n\r\n\r\n+1 on your fix.","created":"2020-06-30T22:47:05.285+0000"},{"body":"[~stack], yeah, it works when backing to the commit before this addendum. ","created":"2020-06-30T22:56:25.088+0000"},{"body":"Reverted. Thanks [~huaxiangsun]","created":"2020-06-30T23:08:05.375+0000"},{"body":"Thanks [~stack] for quick turnaround. ","created":"2020-06-30T23:18:14.563+0000"}],"conversations":[{"body":"{code:java}\r\nINFO] \r\n[INFO] --- maven-surefire-plugin:3.0.0-M4:test (default-test) @ hbase-server ---\r\n[INFO] \r\n[INFO] -------------------------------------------------------\r\n[INFO]  T E S T S\r\n[INFO] -------------------------------------------------------\r\n[INFO] Running org.apache.hadoop.hbase.io.TestHeapSize\r\n[ERROR] Tests run: 6, Failures: 1, Errors: 0, Skipped: 0, Time elapsed: 1.884 s <<< FAILURE! - in org.apache.hadoop.hbase.io.TestHeapSize\r\n[ERROR] org.apache.hadoop.hbase.io.TestHeapSize.testSizes  Time elapsed: 0.308 s  <<< FAILURE!\r\njava.lang.AssertionError: expected:<368> but was:<360>\r\n\tat org.apache.hadoop.hbase.io.TestHeapSize.testSizes(TestHeapSize.java:493)\r\n\r\n\r\n[INFO] \r\n[INFO] Results:\r\n[INFO] \r\n[ERROR] Failures: \r\n[ERROR]   TestHeapSize.testSizes:493 expected:<368> but was:<360>\r\n[INFO] \r\n[ERROR] Tests run: 6, Failures: 1, Errors: 0, Skipped: 0\r\n[INFO] \r\n{code}","from":"reporter","subject":"TestHeapSize.testSizes failure"},{"body":"[~stack] is aware of this issue, the addendum seems causing the issue. It passes without this addendum.\r\n{code:java}\r\ncommit bfec964aef09a6474d35b379fb4000cf3e8f86cc\r\nAuthor: stack \r\nDate:   Mon Jun 29 14:19:44 2020 -0700\r\n\r\n\r\n    HBASE-24648 Remove the legacy 'forceSplit' related code at region server side (#1990)\r\n    Addendum to fix TestHeapSize\r\n\r\n {code}","from":"developer"},{"body":"Addendum by me broke this [~huaxiangsun]?\r\n\r\n\r\ncommit bfec964aef09a6474d35b379fb4000cf3e8f86cc\r\nAuthor: stack \r\nDate: Mon Jun 29 14:19:44 2020 -0700\r\n HBASE-24648 Remove the legacy 'forceSplit' related code at region server side (#1990)\r\n Addendum to fix TestHeapSize\r\n\r\n\r\n\r\n+1 on your fix.","from":"developer"},{"body":"[~stack], yeah, it works when backing to the commit before this addendum. ","from":"developer"},{"body":"Reverted. Thanks [~huaxiangsun]","from":"developer"},{"body":"Thanks [~stack] for quick turnaround. ","from":"developer"}],"created":"2020-06-30T22:17:14.000+0000","description":"{code:java}\r\nINFO] \r\n[INFO] --- maven-surefire-plugin:3.0.0-M4:test (default-test) @ hbase-server ---\r\n[INFO] \r\n[INFO] -------------------------------------------------------\r\n[INFO]  T E S T S\r\n[INFO] -------------------------------------------------------\r\n[INFO] Running org.apache.hadoop.hbase.io.TestHeapSize\r\n[ERROR] Tests run: 6, Failures: 1, Errors: 0, Skipped: 0, Time elapsed: 1.884 s <<< FAILURE! - in org.apache.hadoop.hbase.io.TestHeapSize\r\n[ERROR] org.apache.hadoop.hbase.io.TestHeapSize.testSizes  Time elapsed: 0.308 s  <<< FAILURE!\r\njava.lang.AssertionError: expected:<368> but was:<360>\r\n\tat org.apache.hadoop.hbase.io.TestHeapSize.testSizes(TestHeapSize.java:493)\r\n\r\n\r\n[INFO] \r\n[INFO] Results:\r\n[INFO] \r\n[ERROR] Failures: \r\n[ERROR]   TestHeapSize.testSizes:493 expected:<368> but was:<360>\r\n[INFO] \r\n[ERROR] Tests run: 6, Failures: 1, Errors: 0, Skipped: 0\r\n[INFO] \r\n{code}","issue_id":"13314384","key":"HBASE-24661","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-06-30T23:08:05.000+0000","role":"fixed_distractor","summary":"TestHeapSize.testSizes failure"} {"case_id":"13314891","cluster":"DISTRACTOR-HBASE-24675","comments":[{"body":"Does this apply to branch-1 by any chance? If it does, can you raise a PR?\r\n\r\nThanks","created":"2020-07-17T17:29:43.759+0000"},{"body":"{quote}\r\nDoes this apply to branch-1 by any chance? If it does, can you raise a PR?\r\n{quote}\r\nYest, it is applicable to branch-1 also. Raised PR","created":"2020-07-20T09:01:50.218+0000"},{"body":"Reopening for remaining backports.","created":"2020-07-20T09:11:08.213+0000"},{"body":"| (/) *{color:green}+1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 1m 36s{color} | {color:blue} Docker mode activated. {color} |\r\n|| || || || {color:brown} Prechecks {color} ||\r\n| {color:green}+1{color} | {color:green} dupname {color} | {color:green} 0m 0s{color} | {color:green} No case conflicting files found. {color} |\r\n| {color:green}+1{color} | {color:green} hbaseanti {color} | {color:green} 0m 0s{color} | {color:green} Patch does not have any anti-patterns. {color} |\r\n| {color:green}+1{color} | {color:green} @author {color} | {color:green} 0m 0s{color} | {color:green} The patch does not contain any @author tags. {color} |\r\n| {color:orange}-0{color} | {color:orange} test4tests {color} | {color:orange} 0m 0s{color} | {color:orange} The patch doesn't appear to include any new or modified tests. Please justify why no new tests are needed for this patch. Also please list what manual steps were performed to verify this patch. {color} |\r\n|| || || || {color:brown} branch-1 Compile Tests {color} ||\r\n| {color:green}+1{color} | {color:green} mvninstall {color} | {color:green} 9m 52s{color} | {color:green} branch-1 passed {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 0m 18s{color} | {color:green} branch-1 passed with JDK v1.8.0_252 {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 0m 23s{color} | {color:green} branch-1 passed with JDK v1.7.0_272 {color} |\r\n| {color:green}+1{color} | {color:green} checkstyle {color} | {color:green} 0m 26s{color} | {color:green} branch-1 passed {color} |\r\n| {color:green}+1{color} | {color:green} shadedjars {color} | {color:green} 3m 19s{color} | {color:green} branch has no errors when building our shaded downstream artifacts. {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 30s{color} | {color:green} branch-1 passed with JDK v1.8.0_252 {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 19s{color} | {color:green} branch-1 passed with JDK v1.7.0_272 {color} |\r\n| {color:blue}0{color} | {color:blue} spotbugs {color} | {color:blue} 1m 13s{color} | {color:blue} Used deprecated FindBugs config; considering switching to SpotBugs. {color} |\r\n| {color:green}+1{color} | {color:green} findbugs {color} | {color:green} 1m 10s{color} | {color:green} branch-1 passed {color} |\r\n|| || || || {color:brown} Patch Compile Tests {color} ||\r\n| {color:green}+1{color} | {color:green} mvninstall {color} | {color:green} 2m 6s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 0m 17s{color} | {color:green} the patch passed with JDK v1.8.0_252 {color} |\r\n| {color:green}+1{color} | {color:green} javac {color} | {color:green} 0m 17s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 0m 25s{color} | {color:green} the patch passed with JDK v1.7.0_272 {color} |\r\n| {color:green}+1{color} | {color:green} javac {color} | {color:green} 0m 25s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} checkstyle {color} | {color:green} 0m 17s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} whitespace {color} | {color:green} 0m 0s{color} | {color:green} The patch has no whitespace issues. {color} |\r\n| {color:green}+1{color} | {color:green} shadedjars {color} | {color:green} 3m 8s{color} | {color:green} patch has no errors when building our shaded downstream artifacts. {color} |\r\n| {color:green}+1{color} | {color:green} hadoopcheck {color} | {color:green} 4m 59s{color} | {color:green} Patch does not cause any errors with Hadoop 2.8.5 2.9.2. {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 14s{color} | {color:green} the patch passed with JDK v1.8.0_252 {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 19s{color} | {color:green} the patch passed with JDK v1.7.0_272 {color} |\r\n| {color:green}+1{color} | {color:green} findbugs {color} | {color:green} 1m 1s{color} | {color:green} the patch passed {color} |\r\n|| || || || {color:brown} Other Tests {color} ||\r\n| {color:green}+1{color} | {color:green} unit {color} | {color:green} 14m 33s{color} | {color:green} hbase-rsgroup in the patch passed. {color} |\r\n| {color:green}+1{color} | {color:green} asflicense {color} | {color:green} 0m 21s{color} | {color:green} The patch does not generate ASF License warnings. {color} |\r\n| {color:black}{color} | {color:black} {color} | {color:black} 49m 23s{color} | {color:black} {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| Docker | Client=19.03.9 Server=19.03.9 base: https://builds.apache.org/job/HBase-PreCommit-GitHub-PR/job/PR-2102/1/artifact/out/Dockerfile |\r\n| GITHUB PR | https://github.com/apache/hbase/pull/2102 |\r\n| JIRA Issue | HBASE-24675 |\r\n| Optional Tests | dupname asflicense javac javadoc unit spotbugs findbugs shadedjars hadoopcheck hbaseanti checkstyle compile |\r\n| uname | Linux 574f2add6021 4.15.0-101-generic #102-Ubuntu SMP Mon May 11 10:07:26 UTC 2020 x86_64 x86_64 x86_64 GNU/Linux |\r\n| Build tool | maven |\r\n| Personality | /home/jenkins/jenkins-slave/workspace/Base-PreCommit-GitHub-PR_PR-2102/out/precommit/personality/provided.sh |\r\n| git revision | branch-1 / fb0fb58 |\r\n| Default Java | 1.7.0_272 |\r\n| Multi-JDK versions | /usr/lib/jvm/zulu-8-amd64:1.8.0_252 /usr/lib/jvm/zulu-7-amd64:1.7.0_272 |\r\n| Test Results | https://builds.apache.org/job/HBase-PreCommit-GitHub-PR/job/PR-2102/1/testReport/ |\r\n| Max. process+thread count | 1593 (vs. ulimit of 10000) |\r\n| modules | C: hbase-rsgroup U: hbase-rsgroup |\r\n| Console output | https://builds.apache.org/job/HBase-PreCommit-GitHub-PR/job/PR-2102/1/console |\r\n| versions | git=1.9.1 maven=3.0.5 findbugs=3.0.1 |\r\n| Powered by | Apache Yetus 0.11.1 https://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","created":"2020-07-20T10:23:09.613+0000"},{"body":"Backported to branch-1. We can wait for 2.2.6 RC to go through before backporting this to branch-2.2 ?","created":"2020-07-20T13:44:22.937+0000"}],"conversations":[{"body":"Steps to reproduce:\r\n# Install a HBase cluster with three RS(rs1,rs2 and rs3) and one Master\r\n# Create two rsgroups r1 and r2 and move rs1 to r1 and rs2 to r2\r\n{code:java}\r\nadd_rsgroup 'r1';add_rsgroup 'r2';move_servers_rsgroup 'r1',['host1:16020'];move_servers_rsgroup 'r2',['host2:16020']\r\n{code}\r\n# Restart Master\r\n# Run list_rsgroups for hbase shell, all region servers are assigned to default regroup.\r\n","from":"reporter","subject":"On Master restart all servers are assigned to default rsgroup."},{"body":"Does this apply to branch-1 by any chance? If it does, can you raise a PR?\r\n\r\nThanks","from":"developer"},{"body":"{quote}\r\nDoes this apply to branch-1 by any chance? If it does, can you raise a PR?\r\n{quote}\r\nYest, it is applicable to branch-1 also. Raised PR","from":"developer"},{"body":"Reopening for remaining backports.","from":"developer"},{"body":"| (/) *{color:green}+1 overall{color}* |\r\n\\\\\r\n\\\\\r\n|| Vote || Subsystem || Runtime || Comment ||\r\n| {color:blue}0{color} | {color:blue} reexec {color} | {color:blue} 1m 36s{color} | {color:blue} Docker mode activated. {color} |\r\n|| || || || {color:brown} Prechecks {color} ||\r\n| {color:green}+1{color} | {color:green} dupname {color} | {color:green} 0m 0s{color} | {color:green} No case conflicting files found. {color} |\r\n| {color:green}+1{color} | {color:green} hbaseanti {color} | {color:green} 0m 0s{color} | {color:green} Patch does not have any anti-patterns. {color} |\r\n| {color:green}+1{color} | {color:green} @author {color} | {color:green} 0m 0s{color} | {color:green} The patch does not contain any @author tags. {color} |\r\n| {color:orange}-0{color} | {color:orange} test4tests {color} | {color:orange} 0m 0s{color} | {color:orange} The patch doesn't appear to include any new or modified tests. Please justify why no new tests are needed for this patch. Also please list what manual steps were performed to verify this patch. {color} |\r\n|| || || || {color:brown} branch-1 Compile Tests {color} ||\r\n| {color:green}+1{color} | {color:green} mvninstall {color} | {color:green} 9m 52s{color} | {color:green} branch-1 passed {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 0m 18s{color} | {color:green} branch-1 passed with JDK v1.8.0_252 {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 0m 23s{color} | {color:green} branch-1 passed with JDK v1.7.0_272 {color} |\r\n| {color:green}+1{color} | {color:green} checkstyle {color} | {color:green} 0m 26s{color} | {color:green} branch-1 passed {color} |\r\n| {color:green}+1{color} | {color:green} shadedjars {color} | {color:green} 3m 19s{color} | {color:green} branch has no errors when building our shaded downstream artifacts. {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 30s{color} | {color:green} branch-1 passed with JDK v1.8.0_252 {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 19s{color} | {color:green} branch-1 passed with JDK v1.7.0_272 {color} |\r\n| {color:blue}0{color} | {color:blue} spotbugs {color} | {color:blue} 1m 13s{color} | {color:blue} Used deprecated FindBugs config; considering switching to SpotBugs. {color} |\r\n| {color:green}+1{color} | {color:green} findbugs {color} | {color:green} 1m 10s{color} | {color:green} branch-1 passed {color} |\r\n|| || || || {color:brown} Patch Compile Tests {color} ||\r\n| {color:green}+1{color} | {color:green} mvninstall {color} | {color:green} 2m 6s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 0m 17s{color} | {color:green} the patch passed with JDK v1.8.0_252 {color} |\r\n| {color:green}+1{color} | {color:green} javac {color} | {color:green} 0m 17s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} compile {color} | {color:green} 0m 25s{color} | {color:green} the patch passed with JDK v1.7.0_272 {color} |\r\n| {color:green}+1{color} | {color:green} javac {color} | {color:green} 0m 25s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} checkstyle {color} | {color:green} 0m 17s{color} | {color:green} the patch passed {color} |\r\n| {color:green}+1{color} | {color:green} whitespace {color} | {color:green} 0m 0s{color} | {color:green} The patch has no whitespace issues. {color} |\r\n| {color:green}+1{color} | {color:green} shadedjars {color} | {color:green} 3m 8s{color} | {color:green} patch has no errors when building our shaded downstream artifacts. {color} |\r\n| {color:green}+1{color} | {color:green} hadoopcheck {color} | {color:green} 4m 59s{color} | {color:green} Patch does not cause any errors with Hadoop 2.8.5 2.9.2. {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 14s{color} | {color:green} the patch passed with JDK v1.8.0_252 {color} |\r\n| {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 0m 19s{color} | {color:green} the patch passed with JDK v1.7.0_272 {color} |\r\n| {color:green}+1{color} | {color:green} findbugs {color} | {color:green} 1m 1s{color} | {color:green} the patch passed {color} |\r\n|| || || || {color:brown} Other Tests {color} ||\r\n| {color:green}+1{color} | {color:green} unit {color} | {color:green} 14m 33s{color} | {color:green} hbase-rsgroup in the patch passed. {color} |\r\n| {color:green}+1{color} | {color:green} asflicense {color} | {color:green} 0m 21s{color} | {color:green} The patch does not generate ASF License warnings. {color} |\r\n| {color:black}{color} | {color:black} {color} | {color:black} 49m 23s{color} | {color:black} {color} |\r\n\\\\\r\n\\\\\r\n|| Subsystem || Report/Notes ||\r\n| Docker | Client=19.03.9 Server=19.03.9 base: https://builds.apache.org/job/HBase-PreCommit-GitHub-PR/job/PR-2102/1/artifact/out/Dockerfile |\r\n| GITHUB PR | https://github.com/apache/hbase/pull/2102 |\r\n| JIRA Issue | HBASE-24675 |\r\n| Optional Tests | dupname asflicense javac javadoc unit spotbugs findbugs shadedjars hadoopcheck hbaseanti checkstyle compile |\r\n| uname | Linux 574f2add6021 4.15.0-101-generic #102-Ubuntu SMP Mon May 11 10:07:26 UTC 2020 x86_64 x86_64 x86_64 GNU/Linux |\r\n| Build tool | maven |\r\n| Personality | /home/jenkins/jenkins-slave/workspace/Base-PreCommit-GitHub-PR_PR-2102/out/precommit/personality/provided.sh |\r\n| git revision | branch-1 / fb0fb58 |\r\n| Default Java | 1.7.0_272 |\r\n| Multi-JDK versions | /usr/lib/jvm/zulu-8-amd64:1.8.0_252 /usr/lib/jvm/zulu-7-amd64:1.7.0_272 |\r\n| Test Results | https://builds.apache.org/job/HBase-PreCommit-GitHub-PR/job/PR-2102/1/testReport/ |\r\n| Max. process+thread count | 1593 (vs. ulimit of 10000) |\r\n| modules | C: hbase-rsgroup U: hbase-rsgroup |\r\n| Console output | https://builds.apache.org/job/HBase-PreCommit-GitHub-PR/job/PR-2102/1/console |\r\n| versions | git=1.9.1 maven=3.0.5 findbugs=3.0.1 |\r\n| Powered by | Apache Yetus 0.11.1 https://yetus.apache.org |\r\n\r\n\r\nThis message was automatically generated.\r\n\r\n","from":"developer"},{"body":"Backported to branch-1. We can wait for 2.2.6 RC to go through before backporting this to branch-2.2 ?","from":"developer"}],"created":"2020-07-03T13:19:44.000+0000","description":"Steps to reproduce:\r\n# Install a HBase cluster with three RS(rs1,rs2 and rs3) and one Master\r\n# Create two rsgroups r1 and r2 and move rs1 to r1 and rs2 to r2\r\n{code:java}\r\nadd_rsgroup 'r1';add_rsgroup 'r2';move_servers_rsgroup 'r1',['host1:16020'];move_servers_rsgroup 'r2',['host2:16020']\r\n{code}\r\n# Restart Master\r\n# Run list_rsgroups for hbase shell, all region servers are assigned to default regroup.\r\n","issue_id":"13314891","key":"HBASE-24675","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-07-23T17:28:40.000+0000","role":"fixed_distractor","summary":"On Master restart all servers are assigned to default rsgroup."} {"case_id":"13316641","cluster":"DISTRACTOR-HBASE-24738","comments":[{"body":"[~stack] sir schema is hard coded as \"http\", we should support \"https\" also.","created":"2020-07-14T08:31:46.787+0000"},{"body":"[~bitoffdev] FYI","created":"2020-07-14T14:45:41.962+0000"},{"body":"Hey [~pankajkumar], thanks for doing this! Sorry for the delay.\r\n\r\nMy understanding is that so far, client configurations have been pretty simple and just need either a Zookeeper-based or master-based connection registry. The issue would seem to be that there is no notion of the protocol (ie. HTTP or HTTPS).\r\n\r\nThe implementation uncommented by your PR seems like it should work, although it uses configuration keys intended for the server. +It seems like more of a \"workaround\" solution than the ideal one, but maybe that is okay+.\r\n\r\nThe main alternatives I see:\r\n # We could have the client attempt an HTTP connection and try HTTPS if the first connection gets dropped.\r\n # We could allow users to specify SSL/TLS in the processlist command itself.\r\n\r\nIt would be great to settle on a pattern for any commands that need to connect to the info port of our region servers.\r\n\r\nWorth a footnote is that we might be able to have Jetty upgrade HTTP to HTTPS automatically, but that would be server-side and perhaps a more long-term change.\r\n\r\nIt may be okay to go with your proposed changes, but I wanted to mention my thoughts first. What do you think?","created":"2020-07-23T14:40:00.834+0000"},{"body":"{quote} client configurations have been pretty simple and just need either a Zookeeper-based or master-based connection registry\r\n{quote}\r\nMake sense. \r\n\r\n \r\n{quote}We could have the client attempt an HTTP connection and try HTTPS if the first connection gets dropped.\r\n{quote}\r\nWill update the PR with above approach.","created":"2020-07-28T12:23:50.527+0000"},{"body":"{quote}Will update the PR with above approach.\r\n{quote}\r\nAwesome; thanks!","created":"2020-07-28T15:50:52.024+0000"},{"body":"Merged to branch-2 and master. Thanks for patch [~pankajkumar] . Thanks for review [~bitoffdev]","created":"2020-07-28T16:13:36.127+0000"}],"conversations":[{"body":"HBase Shell command \"processlist\" fails with ERROR: Unexpected end of file from server when HBase SSL enabled.\r\n\r\n \r\n\r\nBelow code is commented since beginning, see HBASE-4368,\r\n\r\n[https://github.com/apache/hbase/blob/8076eafb187ce32d4a78aef482b6218d85a985ac/hbase-shell/src/main/ruby/hbase/taskmonitor.rb#L85]","from":"reporter","subject":"[Shell] processlist command fails with ERROR: Unexpected end of file from server when SSL enabled"},{"body":"[~stack] sir schema is hard coded as \"http\", we should support \"https\" also.","from":"developer"},{"body":"[~bitoffdev] FYI","from":"developer"},{"body":"Hey [~pankajkumar], thanks for doing this! Sorry for the delay.\r\n\r\nMy understanding is that so far, client configurations have been pretty simple and just need either a Zookeeper-based or master-based connection registry. The issue would seem to be that there is no notion of the protocol (ie. HTTP or HTTPS).\r\n\r\nThe implementation uncommented by your PR seems like it should work, although it uses configuration keys intended for the server. +It seems like more of a \"workaround\" solution than the ideal one, but maybe that is okay+.\r\n\r\nThe main alternatives I see:\r\n # We could have the client attempt an HTTP connection and try HTTPS if the first connection gets dropped.\r\n # We could allow users to specify SSL/TLS in the processlist command itself.\r\n\r\nIt would be great to settle on a pattern for any commands that need to connect to the info port of our region servers.\r\n\r\nWorth a footnote is that we might be able to have Jetty upgrade HTTP to HTTPS automatically, but that would be server-side and perhaps a more long-term change.\r\n\r\nIt may be okay to go with your proposed changes, but I wanted to mention my thoughts first. What do you think?","from":"developer"},{"body":"{quote} client configurations have been pretty simple and just need either a Zookeeper-based or master-based connection registry\r\n{quote}\r\nMake sense. \r\n\r\n \r\n{quote}We could have the client attempt an HTTP connection and try HTTPS if the first connection gets dropped.\r\n{quote}\r\nWill update the PR with above approach.","from":"developer"},{"body":"{quote}Will update the PR with above approach.\r\n{quote}\r\nAwesome; thanks!","from":"developer"},{"body":"Merged to branch-2 and master. Thanks for patch [~pankajkumar] . Thanks for review [~bitoffdev]","from":"developer"}],"created":"2020-07-14T08:23:42.000+0000","description":"HBase Shell command \"processlist\" fails with ERROR: Unexpected end of file from server when HBase SSL enabled.\r\n\r\n \r\n\r\nBelow code is commented since beginning, see HBASE-4368,\r\n\r\n[https://github.com/apache/hbase/blob/8076eafb187ce32d4a78aef482b6218d85a985ac/hbase-shell/src/main/ruby/hbase/taskmonitor.rb#L85]","issue_id":"13316641","key":"HBASE-24738","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-07-28T16:13:36.000+0000","role":"fixed_distractor","summary":"[Shell] processlist command fails with ERROR: Unexpected end of file from server when SSL enabled"} {"case_id":"13316788","cluster":"DISTRACTOR-HBASE-24742","comments":[{"body":"There are two observations:\r\n1. We do not need to check for \"fake\" keys inserted by the ROWCOL BF logic if there are not ROWCOL BFs (or if they are not used)\r\n2. We can extend the identity-compare of the nextIndexedKey across multiple calls. It's just an optimization and not for correctness.\r\n","created":"2020-07-15T00:29:31.590+0000"},{"body":"Here's a patch.\r\nPlease have a careful look, especially at the part that turns previousIndexedKey into a member.\r\n","created":"2020-07-15T00:31:20.712+0000"},{"body":"To add more color, following is the tight loop that Lars is talking about\r\n\r\n{noformat}\r\n protected boolean trySkipToNextColumn(Cell cell) throws IOException {\r\n Cell nextCell = null;\r\n // used to guard against a changed next indexed key by doing a identity comparison\r\n // when the identity changes we need to compare the bytes again\r\n Cell previousIndexedKey = null;\r\n do {\r\n Cell nextIndexedKey = getNextIndexedKey();\r\n if (nextIndexedKey != null && nextIndexedKey != KeyValueScanner.NO_NEXT_INDEXED_KEY &&\r\n (nextIndexedKey == previousIndexedKey ||\r\n matcher.compareKeyForNextColumn(nextIndexedKey, cell) >= 0)) { <=====\r\n this.heap.next();\r\n ++kvsScanned;\r\n previousIndexedKey = nextIndexedKey;\r\n } else {\r\n return false;\r\n }\r\n } while ((nextCell = this.heap.peek()) != null && CellUtil.matchingRowColumn(cell, nextCell));\r\n // We need this check because it may happen that the new scanner that we get\r\n // during heap.next() is requiring reseek due of fake KV previously generated for\r\n // ROWCOL bloom filter optimization. See HBASE-19863 for more details\r\n if (nextCell != null && matcher.compareKeyForNextColumn(nextCell, cell) < 0) {. <===\r\n return false;\r\n }\r\n return true;\r\n }\r\n{noformat}\r\n\r\nSpecifically that was added to prevent SQM from matching the skipped rows but it turns out that it does may more compare checks than what it was before. To test our theory we've undone the loop and let the SQM match the rows and we gained almost ~30% back in scans with explicit column filters. But again as discussed in HBASE-17958, that comes at an expense of correctness that filters shouldn't see skipped rows.\r\n\r\n[~zghao] [~zhangduo] FYI since you were involved in the original jira fix and implementation.","created":"2020-07-15T00:33:17.379+0000"},{"body":"Passes the test added in HBASE-19863 and brings the runtime of a test Phoenix query from 5.8s to 4.2s.\r\n(This is for a fully compacted table and VERSIONS=1, which represents the worst case, where the two linked jiras triple the number of comparisons per Cell).\r\n\r\nI'll post a PR tomorrow.","created":"2020-07-15T00:34:47.503+0000"},{"body":"I think the discussion in HBASE-17958 is enough to show that the logic is necessary? Let’s not go back to let the filters deal with strange cells. Will take a look at the patch.\n\nAnyway, back to the problem, I agree we have done too many bytes comparison. We should find a general way to deal with it. Was imagine that, we add some methods in the ScannerContext, to record whether we have changed the row, or family, or column, or version, so for most cases we do not need to do bytes comparison again?\n\nThanks.","created":"2020-07-15T01:05:25.784+0000"},{"body":"There is a regression in SKIP hint handling in branch-2, see HBASE-24637. Is this related? \r\n\r\nIn the HBASE-24637 case the difference is not so much more comparisons on a hot path, which would be good to fix of course, no question, but a regression with wider scope with respect to reseeking (I/O) that causes, proportional to number of cells/columns, more work in the whole stack from file or blockcache to hfile reader up to SQM. \r\n\r\nI went out on vacation (and am still out) before tracking this down. I was/am planning a review of any commit that touches SQM and friends. This was a bit daunting because (I am guessing) the number of commits from circa 1.3 to 2.2 is more than a handful. Maybe not, that would be nice. Perhaps this issue finds it but I suspect more changes in this area were committed after HBASE-17958 and HBASE-19863.","created":"2020-07-15T01:18:01.241+0000"},{"body":"> I think the discussion in HBASE-17958 is enough to show that the logic is necessary\r\n\r\nYep, not suggesting that we undo that patch. Instead we should comprehensively fix the codepaths to not do extra byte compares. So agree with you.\r\n\r\n> I was/am planning a review of any commit that touches SQM and friends. This was a bit daunting because (I am guessing) the number of commits from circa 1.3 to 2.2 is more than a handful. \r\n\r\n[~apurtell] look at the flame graph I attached. I also noticed the bump in number of re-seeks. Based on my analysis I think HBASE-17958 is related, there are more cases where the skip hinting fails in the above code I pasted. Overall I think both the issues are related. We just tested a part of the fix (which is reduce the number of byte comparisons) but we need to analyze the code properly to see where the hinting fails and then re-seeks, which is essentially your jira.","created":"2020-07-15T01:50:43.269+0000"},{"body":"And on the fake cell, maybe we should not pass it to upper layer? I think the intention here, is to avoid do a real seek for a store scanner, if it just needs to know its order in the KeyValueScanner, and delay the actual seek. Now I believe the code logic is to peek the kv and compare them directly. Maybe we could introduce something like getKeyForComparison? So we actually call peek or next, we will done the real seek and get a 'real' kv. In general, I think the open too many internal things to upper layer in KeyValueScanner, we have seek, reseek, requestSeek, realSeekDone, enforceSeek... Maybe when we introduced them at the first place, they were doing well and improving performance, but later when we added more code and fixed bugs, people will misuse them and cause performance regression...\r\n\r\nLooking at the POC, #1 is good, but I'm a bit nervous for #2, as in the discussion in HBASE-17958, We claimed that the index key could be changed during the next call. Will learn more when the PR is ready.\r\n\r\nThanks.","created":"2020-07-15T02:04:54.471+0000"},{"body":"[~zhangduo] I'm with you there. The second part was harder to reason about, and I feel a bit less easy about it.\r\nIn the end it's an optimization to save a comparison, previousIndexedKey and nextIndexedKey will never accidentally be the same (as in identical), so I *think* it should OK.\r\n\r\nI'm not sure how we can avoid passing the fake keys up, since it is designed to handle things and the \"upper\" heap. At least not without a lot refactoring.\r\n\r\n[~apurtell] I'll take a look at HBASE-24637 (I'm a bit thinly spread, though)\r\n","created":"2020-07-15T05:21:54.908+0000"},{"body":"Upon second thought. Perhaps a seek (not a reseek, but a seek that could actually goes backwards) could make it so that previousIndexedKey and nextIndexedKey are accidentally the same and we *still* would have to do the compare.\r\n\r\nIn my test the majority of the improvement came from the first change.\r\n","created":"2020-07-15T05:27:02.930+0000"},{"body":"{quote}\r\nI'm not sure how we can avoid passing the fake keys up, since it is designed to handle things and the \"upper\" heap. At least not without a lot refactoring.\r\n{quote}\r\n\r\nWe need a big refactoring for this change, I believe. Will open a brainstorm issue later, as I'm a bit busy these days. Make summary for the first half and make plan for the second half...","created":"2020-07-15T13:39:03.139+0000"},{"body":"Created a PR for observation #1 above.","created":"2020-07-15T16:08:27.910+0000"},{"body":"Merged into branch-1.\r\n\r\nI'll look into master/branch-2, but my feeling is that things are quite different there.\r\n\r\n[~apurtell] (before you yell at me for not looking at branch-2/master) :)","created":"2020-07-16T18:10:17.433+0000"},{"body":"Lemme put this into branch-2 and master as well.","created":"2020-07-16T19:56:55.371+0000"},{"body":"Master (and branch-2) patch. Will just apply as they're the same as the branch-1 patch.","created":"2020-07-16T20:05:40.409+0000"},{"body":"Also pushed to branch-2 and master.","created":"2020-07-16T20:17:36.447+0000"},{"body":"Sounds good. The other issue is available for branch-2 findings. ","created":"2020-07-16T20:21:20.464+0000"},{"body":"Reopen for 2.3 backport, https://github.com/apache/hbase/pull/2081","created":"2020-07-18T00:01:05.221+0000"},{"body":"I'm still retrying toward a passing precommit run for branch-2.3. In the mean time, would it be possible to get a correctness test case to go along with this change? This seems like too fundamental of an area to be tweaking without test coverage. Thanks.","created":"2020-07-21T18:06:49.209+0000"},{"body":"Looks like HBASE-19863 added some coverage. Do we need more than that?","created":"2020-07-21T18:21:25.541+0000"},{"body":"bq. Looks like HBASE-19863 added some coverage. Do we need more than that?\r\n\r\nIt's hard to say; those tests pass with and without this change. We pushed a change to critical section of code to all maintenance branches that was not explicitly accompanied by updated test coverage. It will go out in patch releases on three release lines. I'm just asking.","created":"2020-07-21T22:54:11.870+0000"},{"body":"Re: Tests. This is purely an internal optimization of another optimization (no kidding) with no functional impact. For the SEEK optimization we have tests that assert the number of SEEKs vs SKIPs during scanning.\r\n\r\nI cannot think of any useful additional tests. Lemme perhaps check if there are SEEK vs SKIP tests with ROWCOL BFs enabled. Or [~bharathv], could you also perhaps have a look as I'm off the next few weeks.\r\n","created":"2020-08-02T20:10:33.384+0000"}],"conversations":[{"body":"In our testing of HBase 1.3 against the current tip of branch-1 we saw a 30% slowdown in scanning scenarios.\r\n\r\nWe tracked it back to HBASE-17958 and HBASE-19863.\r\nBoth add comparisons to one of the tightest HBase has.\r\n\r\n[~bharathv]","from":"reporter","subject":"Improve performance of SKIP vs SEEK logic"},{"body":"There are two observations:\r\n1. We do not need to check for \"fake\" keys inserted by the ROWCOL BF logic if there are not ROWCOL BFs (or if they are not used)\r\n2. We can extend the identity-compare of the nextIndexedKey across multiple calls. It's just an optimization and not for correctness.\r\n","from":"developer"},{"body":"Here's a patch.\r\nPlease have a careful look, especially at the part that turns previousIndexedKey into a member.\r\n","from":"developer"},{"body":"To add more color, following is the tight loop that Lars is talking about\r\n\r\n{noformat}\r\n protected boolean trySkipToNextColumn(Cell cell) throws IOException {\r\n Cell nextCell = null;\r\n // used to guard against a changed next indexed key by doing a identity comparison\r\n // when the identity changes we need to compare the bytes again\r\n Cell previousIndexedKey = null;\r\n do {\r\n Cell nextIndexedKey = getNextIndexedKey();\r\n if (nextIndexedKey != null && nextIndexedKey != KeyValueScanner.NO_NEXT_INDEXED_KEY &&\r\n (nextIndexedKey == previousIndexedKey ||\r\n matcher.compareKeyForNextColumn(nextIndexedKey, cell) >= 0)) { <=====\r\n this.heap.next();\r\n ++kvsScanned;\r\n previousIndexedKey = nextIndexedKey;\r\n } else {\r\n return false;\r\n }\r\n } while ((nextCell = this.heap.peek()) != null && CellUtil.matchingRowColumn(cell, nextCell));\r\n // We need this check because it may happen that the new scanner that we get\r\n // during heap.next() is requiring reseek due of fake KV previously generated for\r\n // ROWCOL bloom filter optimization. See HBASE-19863 for more details\r\n if (nextCell != null && matcher.compareKeyForNextColumn(nextCell, cell) < 0) {. <===\r\n return false;\r\n }\r\n return true;\r\n }\r\n{noformat}\r\n\r\nSpecifically that was added to prevent SQM from matching the skipped rows but it turns out that it does may more compare checks than what it was before. To test our theory we've undone the loop and let the SQM match the rows and we gained almost ~30% back in scans with explicit column filters. But again as discussed in HBASE-17958, that comes at an expense of correctness that filters shouldn't see skipped rows.\r\n\r\n[~zghao] [~zhangduo] FYI since you were involved in the original jira fix and implementation.","from":"developer"},{"body":"Passes the test added in HBASE-19863 and brings the runtime of a test Phoenix query from 5.8s to 4.2s.\r\n(This is for a fully compacted table and VERSIONS=1, which represents the worst case, where the two linked jiras triple the number of comparisons per Cell).\r\n\r\nI'll post a PR tomorrow.","from":"developer"},{"body":"I think the discussion in HBASE-17958 is enough to show that the logic is necessary? Let’s not go back to let the filters deal with strange cells. Will take a look at the patch.\n\nAnyway, back to the problem, I agree we have done too many bytes comparison. We should find a general way to deal with it. Was imagine that, we add some methods in the ScannerContext, to record whether we have changed the row, or family, or column, or version, so for most cases we do not need to do bytes comparison again?\n\nThanks.","from":"developer"},{"body":"There is a regression in SKIP hint handling in branch-2, see HBASE-24637. Is this related? \r\n\r\nIn the HBASE-24637 case the difference is not so much more comparisons on a hot path, which would be good to fix of course, no question, but a regression with wider scope with respect to reseeking (I/O) that causes, proportional to number of cells/columns, more work in the whole stack from file or blockcache to hfile reader up to SQM. \r\n\r\nI went out on vacation (and am still out) before tracking this down. I was/am planning a review of any commit that touches SQM and friends. This was a bit daunting because (I am guessing) the number of commits from circa 1.3 to 2.2 is more than a handful. Maybe not, that would be nice. Perhaps this issue finds it but I suspect more changes in this area were committed after HBASE-17958 and HBASE-19863.","from":"developer"},{"body":"> I think the discussion in HBASE-17958 is enough to show that the logic is necessary\r\n\r\nYep, not suggesting that we undo that patch. Instead we should comprehensively fix the codepaths to not do extra byte compares. So agree with you.\r\n\r\n> I was/am planning a review of any commit that touches SQM and friends. This was a bit daunting because (I am guessing) the number of commits from circa 1.3 to 2.2 is more than a handful. \r\n\r\n[~apurtell] look at the flame graph I attached. I also noticed the bump in number of re-seeks. Based on my analysis I think HBASE-17958 is related, there are more cases where the skip hinting fails in the above code I pasted. Overall I think both the issues are related. We just tested a part of the fix (which is reduce the number of byte comparisons) but we need to analyze the code properly to see where the hinting fails and then re-seeks, which is essentially your jira.","from":"developer"},{"body":"And on the fake cell, maybe we should not pass it to upper layer? I think the intention here, is to avoid do a real seek for a store scanner, if it just needs to know its order in the KeyValueScanner, and delay the actual seek. Now I believe the code logic is to peek the kv and compare them directly. Maybe we could introduce something like getKeyForComparison? So we actually call peek or next, we will done the real seek and get a 'real' kv. In general, I think the open too many internal things to upper layer in KeyValueScanner, we have seek, reseek, requestSeek, realSeekDone, enforceSeek... Maybe when we introduced them at the first place, they were doing well and improving performance, but later when we added more code and fixed bugs, people will misuse them and cause performance regression...\r\n\r\nLooking at the POC, #1 is good, but I'm a bit nervous for #2, as in the discussion in HBASE-17958, We claimed that the index key could be changed during the next call. Will learn more when the PR is ready.\r\n\r\nThanks.","from":"developer"},{"body":"[~zhangduo] I'm with you there. The second part was harder to reason about, and I feel a bit less easy about it.\r\nIn the end it's an optimization to save a comparison, previousIndexedKey and nextIndexedKey will never accidentally be the same (as in identical), so I *think* it should OK.\r\n\r\nI'm not sure how we can avoid passing the fake keys up, since it is designed to handle things and the \"upper\" heap. At least not without a lot refactoring.\r\n\r\n[~apurtell] I'll take a look at HBASE-24637 (I'm a bit thinly spread, though)\r\n","from":"developer"},{"body":"Upon second thought. Perhaps a seek (not a reseek, but a seek that could actually goes backwards) could make it so that previousIndexedKey and nextIndexedKey are accidentally the same and we *still* would have to do the compare.\r\n\r\nIn my test the majority of the improvement came from the first change.\r\n","from":"developer"},{"body":"{quote}\r\nI'm not sure how we can avoid passing the fake keys up, since it is designed to handle things and the \"upper\" heap. At least not without a lot refactoring.\r\n{quote}\r\n\r\nWe need a big refactoring for this change, I believe. Will open a brainstorm issue later, as I'm a bit busy these days. Make summary for the first half and make plan for the second half...","from":"developer"},{"body":"Created a PR for observation #1 above.","from":"developer"},{"body":"Merged into branch-1.\r\n\r\nI'll look into master/branch-2, but my feeling is that things are quite different there.\r\n\r\n[~apurtell] (before you yell at me for not looking at branch-2/master) :)","from":"developer"},{"body":"Lemme put this into branch-2 and master as well.","from":"developer"},{"body":"Master (and branch-2) patch. Will just apply as they're the same as the branch-1 patch.","from":"developer"},{"body":"Also pushed to branch-2 and master.","from":"developer"},{"body":"Sounds good. The other issue is available for branch-2 findings. ","from":"developer"},{"body":"Reopen for 2.3 backport, https://github.com/apache/hbase/pull/2081","from":"developer"},{"body":"I'm still retrying toward a passing precommit run for branch-2.3. In the mean time, would it be possible to get a correctness test case to go along with this change? This seems like too fundamental of an area to be tweaking without test coverage. Thanks.","from":"developer"},{"body":"Looks like HBASE-19863 added some coverage. Do we need more than that?","from":"developer"},{"body":"bq. Looks like HBASE-19863 added some coverage. Do we need more than that?\r\n\r\nIt's hard to say; those tests pass with and without this change. We pushed a change to critical section of code to all maintenance branches that was not explicitly accompanied by updated test coverage. It will go out in patch releases on three release lines. I'm just asking.","from":"developer"},{"body":"Re: Tests. This is purely an internal optimization of another optimization (no kidding) with no functional impact. For the SEEK optimization we have tests that assert the number of SEEKs vs SKIPs during scanning.\r\n\r\nI cannot think of any useful additional tests. Lemme perhaps check if there are SEEK vs SKIP tests with ROWCOL BFs enabled. Or [~bharathv], could you also perhaps have a look as I'm off the next few weeks.\r\n","from":"developer"}],"created":"2020-07-15T00:27:52.000+0000","description":"In our testing of HBase 1.3 against the current tip of branch-1 we saw a 30% slowdown in scanning scenarios.\r\n\r\nWe tracked it back to HBASE-17958 and HBASE-19863.\r\nBoth add comparisons to one of the tightest HBase has.\r\n\r\n[~bharathv]","issue_id":"13316788","key":"HBASE-24742","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-07-21T22:51:02.000+0000","role":"fixed_distractor","summary":"Improve performance of SKIP vs SEEK logic"} {"case_id":"13320017","cluster":"DISTRACTOR-HBASE-24794","comments":[{"body":"did PR linking break again?\r\n\r\nhttps://github.com/apache/hbase/pull/2174","created":"2020-07-30T06:06:14.225+0000"}],"conversations":[{"body":"had a cluster fail after upgrade from hbase 1 because all writes to meta failed.\r\n\r\nmaster started in maintenance mode looks like (RS hosting meta in non-maint would look similar starting with {{HRegion.doBatchMutate}}):\r\n\r\n{code}\r\n2020-07-28 17:52:56,553 WARN org.apache.hadoop.hbase.regionserver.HRegion: Failed getting lock, row=some_user_table\r\njava.io.IOException: Timed out waiting for lock for row: some_user_table in region 1588230740\r\n at org.apache.hadoop.hbase.regionserver.HRegion.getRowLockInternal(HRegion.java:5863)\r\n at org.apache.hadoop.hbase.regionserver.HRegion$BatchOperation.lockRowsAndBuildMiniBatch(HRegion.java:3322)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.doMiniBatchMutate(HRegion.java:4018)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:3992)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:3923)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:3914)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:3928)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.doBatchMutate(HRegion.java:4255)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.put(HRegion.java:3047)\r\n at org.apache.hadoop.hbase.regionserver.RSRpcServices.mutate(RSRpcServices.java:2827)\r\n at org.apache.hadoop.hbase.client.ClientServiceCallable.doMutate(ClientServiceCallable.java:55)\r\n at org.apache.hadoop.hbase.client.HTable$3.rpcCall(HTable.java:538)\r\n at org.apache.hadoop.hbase.client.HTable$3.rpcCall(HTable.java:533)\r\n at org.apache.hadoop.hbase.client.RegionServerCallable.call(RegionServerCallable.java:127)\r\n at org.apache.hadoop.hbase.client.RpcRetryingCallerImpl.callWithRetries(RpcRetryingCallerImpl.java:107)\r\n at org.apache.hadoop.hbase.client.HTable.put(HTable.java:542)\r\n at org.apache.hadoop.hbase.MetaTableAccessor.put(MetaTableAccessor.java:1339)\r\n at org.apache.hadoop.hbase.MetaTableAccessor.putToMetaTable(MetaTableAccessor.java:1329)\r\n at org.apache.hadoop.hbase.MetaTableAccessor.updateTableState(MetaTableAccessor.java:1672)\r\n at org.apache.hadoop.hbase.MetaTableAccessor.updateTableState(MetaTableAccessor.java:1112)\r\n at org.apache.hadoop.hbase.master.TableStateManager.fixTableStates(TableStateManager.java:296)\r\n at org.apache.hadoop.hbase.master.TableStateManager.start(TableStateManager.java:269)\r\n at org.apache.hadoop.hbase.master.HMaster.finishActiveMasterInitialization(HMaster.java:1004)\r\n at org.apache.hadoop.hbase.master.HMaster.startActiveMasterManager(HMaster.java:2274)\r\n at org.apache.hadoop.hbase.master.HMaster.lambda$run$0(HMaster.java:583)\r\n at java.lang.Thread.run(Thread.java:745)\r\n{code}\r\n\r\nlogging roughly 6k times /second.\r\n\r\nfailure was caused by a change in behavior for {{hbase.rowlock.wait.duration}} in HBASE-17210 (so 1.4.0+, 2.0.0+). Prior to that change setting the config <= 0 meant that row locks would succeed only if they were immediately available. After the change we fail the lock attempt without checking the lock at all.\r\n\r\nworkaround: set {{hbase.rowlock.wait.duration}} to a small positive number, e.g. 1, if you want row locks to fail quickly.","from":"reporter","subject":"hbase.rowlock.wait.duration should not be <= 0"},{"body":"did PR linking break again?\r\n\r\nhttps://github.com/apache/hbase/pull/2174","from":"developer"}],"created":"2020-07-29T16:50:28.000+0000","description":"had a cluster fail after upgrade from hbase 1 because all writes to meta failed.\r\n\r\nmaster started in maintenance mode looks like (RS hosting meta in non-maint would look similar starting with {{HRegion.doBatchMutate}}):\r\n\r\n{code}\r\n2020-07-28 17:52:56,553 WARN org.apache.hadoop.hbase.regionserver.HRegion: Failed getting lock, row=some_user_table\r\njava.io.IOException: Timed out waiting for lock for row: some_user_table in region 1588230740\r\n at org.apache.hadoop.hbase.regionserver.HRegion.getRowLockInternal(HRegion.java:5863)\r\n at org.apache.hadoop.hbase.regionserver.HRegion$BatchOperation.lockRowsAndBuildMiniBatch(HRegion.java:3322)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.doMiniBatchMutate(HRegion.java:4018)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:3992)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:3923)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:3914)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.batchMutate(HRegion.java:3928)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.doBatchMutate(HRegion.java:4255)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.put(HRegion.java:3047)\r\n at org.apache.hadoop.hbase.regionserver.RSRpcServices.mutate(RSRpcServices.java:2827)\r\n at org.apache.hadoop.hbase.client.ClientServiceCallable.doMutate(ClientServiceCallable.java:55)\r\n at org.apache.hadoop.hbase.client.HTable$3.rpcCall(HTable.java:538)\r\n at org.apache.hadoop.hbase.client.HTable$3.rpcCall(HTable.java:533)\r\n at org.apache.hadoop.hbase.client.RegionServerCallable.call(RegionServerCallable.java:127)\r\n at org.apache.hadoop.hbase.client.RpcRetryingCallerImpl.callWithRetries(RpcRetryingCallerImpl.java:107)\r\n at org.apache.hadoop.hbase.client.HTable.put(HTable.java:542)\r\n at org.apache.hadoop.hbase.MetaTableAccessor.put(MetaTableAccessor.java:1339)\r\n at org.apache.hadoop.hbase.MetaTableAccessor.putToMetaTable(MetaTableAccessor.java:1329)\r\n at org.apache.hadoop.hbase.MetaTableAccessor.updateTableState(MetaTableAccessor.java:1672)\r\n at org.apache.hadoop.hbase.MetaTableAccessor.updateTableState(MetaTableAccessor.java:1112)\r\n at org.apache.hadoop.hbase.master.TableStateManager.fixTableStates(TableStateManager.java:296)\r\n at org.apache.hadoop.hbase.master.TableStateManager.start(TableStateManager.java:269)\r\n at org.apache.hadoop.hbase.master.HMaster.finishActiveMasterInitialization(HMaster.java:1004)\r\n at org.apache.hadoop.hbase.master.HMaster.startActiveMasterManager(HMaster.java:2274)\r\n at org.apache.hadoop.hbase.master.HMaster.lambda$run$0(HMaster.java:583)\r\n at java.lang.Thread.run(Thread.java:745)\r\n{code}\r\n\r\nlogging roughly 6k times /second.\r\n\r\nfailure was caused by a change in behavior for {{hbase.rowlock.wait.duration}} in HBASE-17210 (so 1.4.0+, 2.0.0+). Prior to that change setting the config <= 0 meant that row locks would succeed only if they were immediately available. After the change we fail the lock attempt without checking the lock at all.\r\n\r\nworkaround: set {{hbase.rowlock.wait.duration}} to a small positive number, e.g. 1, if you want row locks to fail quickly.","issue_id":"13320017","key":"HBASE-24794","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-07-30T20:22:47.000+0000","role":"fixed_distractor","summary":"hbase.rowlock.wait.duration should not be <= 0"} {"case_id":"13320733","cluster":"DISTRACTOR-HBASE-24813","comments":[{"body":"This cause very serious problem since the ReplicationSource thread can not quit and then we flood the our log in UT with\r\n\r\n{noformat}\r\nInterrupting source thread for peer xxx without cleaning buffer usage\r\n{noformat}\r\n\r\nAnd then blow up the /tmp spaces, as one single std*deferred file(which is used to buffer stdout and stderr for surefire xml report) can be several tens of GBs.\r\n\r\nThis is a very critical problem that cause the jenkins node unavailable so let me revert it first.\r\n\r\nPlease consider fixing the problem before committing again.\r\n\r\nThanks a lot.","created":"2020-10-06T13:15:09.015+0000"},{"body":"Thanks [~zhangduo], if you managed to recover test logs/outputs, could you attach it here? ","created":"2020-10-06T14:06:47.797+0000"},{"body":"And it could even consume all the spaces for the data disk...\r\n\r\nI've created a jenkins job to execute shell command on jenkins node, and this is what I found on hbase7...\r\n\r\n{noformat}\r\n103G\t/home/jenkins/jenkins-home/workspace/HBase_HBase_Nightly_HBASE-25152@2/component/hbase-server/target/surefire-reports/org.apache.hadoop.hbase.replication.TestMasterReplication-output.txt\r\n100G\t/home/jenkins/jenkins-home/workspace/HBase_HBase_Nightly_HBASE-25152@2/component/hbase-server/target/surefire-reports/org.apache.hadoop.hbase.replication.TestReplicationStatusBothNormalAndRecoveryLagging-output.txt\r\n5.9G\t/home/jenkins/jenkins-home/workspace/HBase_HBase_Nightly_HBASE-25152@2/component/hbase-server/target/surefire-reports/TEST-org.apache.hadoop.hbase.replication.TestReplicationStatusBothNormalAndRecoveryLagging.xml\r\n5.4G\t/home/jenkins/jenkins-home/workspace/HBase_HBase_Nightly_HBASE-25152@2/component/hbase-server/target/surefire-reports/TEST-org.apache.hadoop.hbase.replication.TestMasterReplication.xml\r\n{noformat}\r\n\r\nhttps://ci-hadoop.apache.org/job/HBase/job/Test-Docker-Version/36/consoleFull","created":"2020-10-06T14:14:16.883+0000"},{"body":"{quote}\r\nThanks Duo Zhang, if you managed to recover test logs/outputs, could you attach it here?\r\n{quote}\r\n\r\nSee the attached TestReplicationSyncUpTool.log. I ran it locally and then force stopped it before it ate all my local spaces...\r\n\r\nThanks.","created":"2020-10-06T14:27:21.240+0000"},{"body":"Please take a look at HBASE-25117, that may fix this problem?\r\n\r\n ","created":"2020-10-09T02:50:21.173+0000"},{"body":"Moving out of 2.4.","created":"2020-11-23T19:12:48.250+0000"},{"body":"This should be re-resolved since the bug was addressed in HBASE-25117, yes?","created":"2021-01-04T17:31:18.157+0000"},{"body":"Yes, also got an approval for latest PR, let me commit it.","created":"2021-01-05T10:20:09.796+0000"},{"body":"Moving out to 2.3.5.","created":"2021-01-06T18:48:45.542+0000"},{"body":"Merged into master, branch-2, branch-2.4 and branch-2.2. Waiting on [~huaxiangsun] green sign to merge into branch-2.3 once he's done with 2.3.4 release.","created":"2021-01-12T10:51:59.428+0000"},{"body":"I'm resolving this for 2.4.1 RC0. Reopen and adjust fix versions for 2.3 at commit time but not before please. ","created":"2021-01-12T23:59:29.715+0000"}],"conversations":[{"body":"Following investigations on the issue described by [~elserj] on HBASE-24779, we found out that once a peer is removed, thus killing peers related *ReplicationSource* instance, it may leave *ReplicationSourceManager.totalBufferUsed* inconsistent. This can happen if *ReplicationSourceWALReader* had put some entries on its queue to be processed by *ReplicationSourceShipper,* but the peer removal killed the shipper before it could process the pending entries. When *ReplicationSourceWALReader* thread add entries to the queue, it increments *ReplicationSourceManager.totalBufferUsed* with the sum of the entries sizes. When those entries are read by *ReplicationSourceShipper,* *ReplicationSourceManager.totalBufferUsed* is then decreased. We should also decrease *ReplicationSourceManager.totalBufferUsed* when *ReplicationSource* is terminated, otherwise those unprocessed entries size would be consuming *ReplicationSourceManager.totalBufferUsed __*indefinitely, unless the RS gets restarted. This may be a problem for deployments with multiple peers, or if new peers are added.**","from":"reporter","subject":"ReplicationSource should clear buffer usage on ReplicationSourceManager upon termination"},{"body":"This cause very serious problem since the ReplicationSource thread can not quit and then we flood the our log in UT with\r\n\r\n{noformat}\r\nInterrupting source thread for peer xxx without cleaning buffer usage\r\n{noformat}\r\n\r\nAnd then blow up the /tmp spaces, as one single std*deferred file(which is used to buffer stdout and stderr for surefire xml report) can be several tens of GBs.\r\n\r\nThis is a very critical problem that cause the jenkins node unavailable so let me revert it first.\r\n\r\nPlease consider fixing the problem before committing again.\r\n\r\nThanks a lot.","from":"developer"},{"body":"Thanks [~zhangduo], if you managed to recover test logs/outputs, could you attach it here? ","from":"developer"},{"body":"And it could even consume all the spaces for the data disk...\r\n\r\nI've created a jenkins job to execute shell command on jenkins node, and this is what I found on hbase7...\r\n\r\n{noformat}\r\n103G\t/home/jenkins/jenkins-home/workspace/HBase_HBase_Nightly_HBASE-25152@2/component/hbase-server/target/surefire-reports/org.apache.hadoop.hbase.replication.TestMasterReplication-output.txt\r\n100G\t/home/jenkins/jenkins-home/workspace/HBase_HBase_Nightly_HBASE-25152@2/component/hbase-server/target/surefire-reports/org.apache.hadoop.hbase.replication.TestReplicationStatusBothNormalAndRecoveryLagging-output.txt\r\n5.9G\t/home/jenkins/jenkins-home/workspace/HBase_HBase_Nightly_HBASE-25152@2/component/hbase-server/target/surefire-reports/TEST-org.apache.hadoop.hbase.replication.TestReplicationStatusBothNormalAndRecoveryLagging.xml\r\n5.4G\t/home/jenkins/jenkins-home/workspace/HBase_HBase_Nightly_HBASE-25152@2/component/hbase-server/target/surefire-reports/TEST-org.apache.hadoop.hbase.replication.TestMasterReplication.xml\r\n{noformat}\r\n\r\nhttps://ci-hadoop.apache.org/job/HBase/job/Test-Docker-Version/36/consoleFull","from":"developer"},{"body":"{quote}\r\nThanks Duo Zhang, if you managed to recover test logs/outputs, could you attach it here?\r\n{quote}\r\n\r\nSee the attached TestReplicationSyncUpTool.log. I ran it locally and then force stopped it before it ate all my local spaces...\r\n\r\nThanks.","from":"developer"},{"body":"Please take a look at HBASE-25117, that may fix this problem?\r\n\r\n ","from":"developer"},{"body":"Moving out of 2.4.","from":"developer"},{"body":"This should be re-resolved since the bug was addressed in HBASE-25117, yes?","from":"developer"},{"body":"Yes, also got an approval for latest PR, let me commit it.","from":"developer"},{"body":"Moving out to 2.3.5.","from":"developer"},{"body":"Merged into master, branch-2, branch-2.4 and branch-2.2. Waiting on [~huaxiangsun] green sign to merge into branch-2.3 once he's done with 2.3.4 release.","from":"developer"},{"body":"I'm resolving this for 2.4.1 RC0. Reopen and adjust fix versions for 2.3 at commit time but not before please. ","from":"developer"}],"created":"2020-08-03T19:20:41.000+0000","description":"Following investigations on the issue described by [~elserj] on HBASE-24779, we found out that once a peer is removed, thus killing peers related *ReplicationSource* instance, it may leave *ReplicationSourceManager.totalBufferUsed* inconsistent. This can happen if *ReplicationSourceWALReader* had put some entries on its queue to be processed by *ReplicationSourceShipper,* but the peer removal killed the shipper before it could process the pending entries. When *ReplicationSourceWALReader* thread add entries to the queue, it increments *ReplicationSourceManager.totalBufferUsed* with the sum of the entries sizes. When those entries are read by *ReplicationSourceShipper,* *ReplicationSourceManager.totalBufferUsed* is then decreased. We should also decrease *ReplicationSourceManager.totalBufferUsed* when *ReplicationSource* is terminated, otherwise those unprocessed entries size would be consuming *ReplicationSourceManager.totalBufferUsed __*indefinitely, unless the RS gets restarted. This may be a problem for deployments with multiple peers, or if new peers are added.**","issue_id":"13320733","key":"HBASE-24813","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-01-12T23:59:57.000+0000","role":"fixed_distractor","summary":"ReplicationSource should clear buffer usage on ReplicationSourceManager upon termination"} {"case_id":"13322359","cluster":"DISTRACTOR-HBASE-24874","comments":[{"body":"We do not have UTs for this two commands? IIRC I did not see any failures on the hbase-shell UTs.","created":"2020-08-13T01:43:48.594+0000"},{"body":"[~zhangduo], we do have unit tests for both these cases in TestAdminShell.java. These shell unit tests are actually the reason that one of my open PRs ([https://github.com/apache/hbase/pull/2232]) is failing CI testing.\r\n\r\nI just took a look at the nightly build for both JDK 8 and 11. hbase-shell is passing for both of them. (/)\r\n\r\nI think the problem is that *TestAdminShell is not always run in CI*. It was definitely run for my recently opened PR. However, *TestAdminShell now appears in the list of flaky excludes* for the master branch, so it is not being run on nightly builds. This may have other effects, but I'm not familiar enough with our CI to comment further at the moment. See [https://ci-hadoop.apache.org/job/HBase/job/HBase-Find-Flaky-Tests/job/master/lastSuccessfulBuild/artifact/dashboard.html]\r\n\r\nI'll have a chance to look into this further tomorrow.","created":"2020-08-13T02:40:31.141+0000"},{"body":"OK. Thank you for taking a look of this.","created":"2020-08-13T06:10:58.120+0000"},{"body":"One of the things that fascinated me about this bug is that it only *affects the JDK 11 builds*. The JDK 8 builds work just fine. Hopefully, I'll get the chance to look into this more, but my initial guess is that this effect is a result of \"JEP 181: Nest-Based Access Control,\" added in JDK 11 (See the release notes: [https://www.oracle.com/java/technologies/javase/jdk-11-relnote.html#JDK-8010319).]\r\n\r\nAnyway, there are two changes that need to be made:\r\n\r\n1. We need a new way to accept coprocessor spec strings from the alter command. I suggested a solution on this issue relating to CoprocessorDescriptors: [HBASE-20119|https://issues.apache.org/jira/browse/HBASE-20119?focusedCommentId=17178013&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-17178013]\r\n 2. We need a new way to print table attributes in the describe command. This is easy.","created":"2020-08-14T18:58:31.534+0000"},{"body":"Thanks for the fix [~bitoffdev]","created":"2020-08-18T20:01:14.746+0000"}],"conversations":[{"body":"HBASE-20819 prepared us for HBase 3.x by removing usages of the deprecated HTableDescriptor and HColumnDescriptor classes from the shell. However, it did use two methods from the ModifiableTableDescriptor, which was only public for compatibility/migration and was marked with {{@InterfaceAudience.Private}}. When {{ModifiableTableDescriptor}} was made private last week by HBASE-24507 it broke two hbase-shell commands (*describe* and *alter* when used to set a coprocessor) that were using methods from {{ModifiableTableDescriptor}} (these methods are not present on the general {{TableDescriptor}} interface).\r\n\r\nThis story will remove the two references in hbase-shell to methods on the now-private {{ModifiableTableDescriptor}} class and will find appropriate replacements for the calls.","from":"reporter","subject":"Fix hbase-shell access to ModifiableTableDescriptor methods"},{"body":"We do not have UTs for this two commands? IIRC I did not see any failures on the hbase-shell UTs.","from":"developer"},{"body":"[~zhangduo], we do have unit tests for both these cases in TestAdminShell.java. These shell unit tests are actually the reason that one of my open PRs ([https://github.com/apache/hbase/pull/2232]) is failing CI testing.\r\n\r\nI just took a look at the nightly build for both JDK 8 and 11. hbase-shell is passing for both of them. (/)\r\n\r\nI think the problem is that *TestAdminShell is not always run in CI*. It was definitely run for my recently opened PR. However, *TestAdminShell now appears in the list of flaky excludes* for the master branch, so it is not being run on nightly builds. This may have other effects, but I'm not familiar enough with our CI to comment further at the moment. See [https://ci-hadoop.apache.org/job/HBase/job/HBase-Find-Flaky-Tests/job/master/lastSuccessfulBuild/artifact/dashboard.html]\r\n\r\nI'll have a chance to look into this further tomorrow.","from":"developer"},{"body":"OK. Thank you for taking a look of this.","from":"developer"},{"body":"One of the things that fascinated me about this bug is that it only *affects the JDK 11 builds*. The JDK 8 builds work just fine. Hopefully, I'll get the chance to look into this more, but my initial guess is that this effect is a result of \"JEP 181: Nest-Based Access Control,\" added in JDK 11 (See the release notes: [https://www.oracle.com/java/technologies/javase/jdk-11-relnote.html#JDK-8010319).]\r\n\r\nAnyway, there are two changes that need to be made:\r\n\r\n1. We need a new way to accept coprocessor spec strings from the alter command. I suggested a solution on this issue relating to CoprocessorDescriptors: [HBASE-20119|https://issues.apache.org/jira/browse/HBASE-20119?focusedCommentId=17178013&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-17178013]\r\n 2. We need a new way to print table attributes in the describe command. This is easy.","from":"developer"},{"body":"Thanks for the fix [~bitoffdev]","from":"developer"}],"created":"2020-08-12T20:18:07.000+0000","description":"HBASE-20819 prepared us for HBase 3.x by removing usages of the deprecated HTableDescriptor and HColumnDescriptor classes from the shell. However, it did use two methods from the ModifiableTableDescriptor, which was only public for compatibility/migration and was marked with {{@InterfaceAudience.Private}}. When {{ModifiableTableDescriptor}} was made private last week by HBASE-24507 it broke two hbase-shell commands (*describe* and *alter* when used to set a coprocessor) that were using methods from {{ModifiableTableDescriptor}} (these methods are not present on the general {{TableDescriptor}} interface).\r\n\r\nThis story will remove the two references in hbase-shell to methods on the now-private {{ModifiableTableDescriptor}} class and will find appropriate replacements for the calls.","issue_id":"13322359","key":"HBASE-24874","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-08-18T20:01:14.000+0000","role":"fixed_distractor","summary":"Fix hbase-shell access to ModifiableTableDescriptor methods"} {"case_id":"13322844","cluster":"DISTRACTOR-HBASE-24885","comments":[{"body":"Thanks for finding this issue [~Bo Cui].\r\n\r\nWhen we assign explicitly a region which is in already OPEN state then it may lead to Double Assignment or stuck in RIT.\r\n\r\nIf HMaster choose the same server for assignment where it is already OPEN then RS will reject the request saying already OPEN and region will stuck in RIT, otherwise double assignment problem will happen.\r\n\r\nWe need add prevalidation logic before creating TRSP for this case. Ping [~stack] [~zhangduo].","created":"2020-08-17T07:22:10.092+0000"},{"body":"Is this only a problem of HBCK2? Or we could meet the same problem when calling admin.assign?","created":"2020-08-17T07:57:38.377+0000"},{"body":"bq. Is this only a problem of HBCK2? Or we could meet the same problem when calling admin.assign?\r\n\r\nyeah, TRSP can judge current regionState in TRSP#executeFromState(REGION_STATE_TRANSITION_GET_ASSIGN_CANDIDATE/REGION_STATE_TRANSITION_CLOSE)\r\nhttps://github.com/apache/hbase/blob/7335dbc8345298c57b8da4ccba640d5432b3fde9/hbase-server/src/main/java/org/apache/hadoop/hbase/master/assignment/TransitRegionStateProcedure.java#L340","created":"2020-08-17T08:22:58.546+0000"},{"body":"{quote}Is this only a problem of HBCK2?\r\n{quote}\r\nYeah, MasterRpcServices.assigns() caller will may face this problem.\r\n\r\nIn Am.createOneAssignProcedure() we create TRSP without any precheck.\r\n\r\n[https://github.com/apache/hbase/blob/7335dbc8345298c57b8da4ccba640d5432b3fde9/hbase-server/src/main/java/org/apache/hadoop/hbase/master/MasterRpcServices.java#L2576]","created":"2020-08-17T08:29:18.918+0000"},{"body":"Want me to take this [~Bo Cui]. I had seen the second case occur dealing w/ a Replica that failed assign and had added investigation to my todo list. Thanks for digging in here.","created":"2020-08-17T14:44:16.711+0000"},{"body":"So this is only a proble for HBCK2? Then I think it is the caller's duty to make sure that you do not do something wrong... HBCK2 is for fixing assignment problems, so the method is designed to skip all the checks and force to do an assignment...\r\n\r\nThis is the comment of the method\r\n\r\n{code}\r\n /**\r\n * A 'raw' version of assign that does bulk and skirts Master state checks (assigns can be made\r\n * during Master startup). For use by Hbck2.\r\n */\r\n{code}","created":"2020-08-17T14:56:52.042+0000"},{"body":"{quote}Then I think it is the caller's duty to make sure that you do not do something wrong... HBCK2 is for fixing assignment problems, so the method is designed to skip all the checks and force to do an assignment..\r\n{quote}\r\nWe have an '--override' flag on assigns/unassigns in hbck2. When you pass this flag, I agree it should do as you describe above. When you do not pass this flag, if it can, the system should stop you shooting yourself in the foot with clear messaging on why it is not doing as you asked.\r\n\r\nMaking repairs on largish clusters with thousands of assignment problems, I make mistakes especially when running hasty bulk 'fixes'; i end up swapping one problem for a new one or as is the case here, doubling the problems that need fixing because I made the wrong call.\r\n\r\nLet me fix the comment too. It needs some color.","created":"2020-08-17T17:33:04.933+0000"},{"body":"{quote}Want me to take this [~Bo Cui]. I had seen the second case occur dealing w/ a Replica that failed assign and had added investigation to my todo list. Thanks for digging in here.{quote}\r\n[~stack] thx, i have assigned to u.","created":"2020-08-18T02:08:18.018+0000"},{"body":"I put up a PR....","created":"2020-08-19T19:33:27.492+0000"},{"body":"Merged. Thanks for review [~zhangduo]\r\n\r\n ","created":"2020-08-24T16:23:59.193+0000"}],"conversations":[{"body":"If a region has been assign to rs1 and then client assigns region again by \"hbck2 assigns\"\r\n\r\n1、if  regionPlan is region to be assign to rs2,the region will be opened on rs1 and rs2.\r\n\r\nmaster log:\r\n{quote}WARN org.apache.hadoop.hbase.master.assignment.AssignmentManager: rit=OPEN, location=rs2, table=tableName, region=reionName reported OPEN on server=rs1 but state has otherwise\r\n{quote}\r\n2、if regionPlan is region to be assign to rs1, the TransitRegionStateProcedure and OpenRegionProcedure will stuck. because rs1 is not responding to master\r\n rslog:\r\n{quote}Receiving OPEN for the region:{}, which we are already trying to OPEN - ignoring this new request for this region.\r\n{quote}\r\n ","from":"reporter","subject":"STUCK RIT by hbck2 assigns"},{"body":"Thanks for finding this issue [~Bo Cui].\r\n\r\nWhen we assign explicitly a region which is in already OPEN state then it may lead to Double Assignment or stuck in RIT.\r\n\r\nIf HMaster choose the same server for assignment where it is already OPEN then RS will reject the request saying already OPEN and region will stuck in RIT, otherwise double assignment problem will happen.\r\n\r\nWe need add prevalidation logic before creating TRSP for this case. Ping [~stack] [~zhangduo].","from":"developer"},{"body":"Is this only a problem of HBCK2? Or we could meet the same problem when calling admin.assign?","from":"developer"},{"body":"bq. Is this only a problem of HBCK2? Or we could meet the same problem when calling admin.assign?\r\n\r\nyeah, TRSP can judge current regionState in TRSP#executeFromState(REGION_STATE_TRANSITION_GET_ASSIGN_CANDIDATE/REGION_STATE_TRANSITION_CLOSE)\r\nhttps://github.com/apache/hbase/blob/7335dbc8345298c57b8da4ccba640d5432b3fde9/hbase-server/src/main/java/org/apache/hadoop/hbase/master/assignment/TransitRegionStateProcedure.java#L340","from":"developer"},{"body":"{quote}Is this only a problem of HBCK2?\r\n{quote}\r\nYeah, MasterRpcServices.assigns() caller will may face this problem.\r\n\r\nIn Am.createOneAssignProcedure() we create TRSP without any precheck.\r\n\r\n[https://github.com/apache/hbase/blob/7335dbc8345298c57b8da4ccba640d5432b3fde9/hbase-server/src/main/java/org/apache/hadoop/hbase/master/MasterRpcServices.java#L2576]","from":"developer"},{"body":"Want me to take this [~Bo Cui]. I had seen the second case occur dealing w/ a Replica that failed assign and had added investigation to my todo list. Thanks for digging in here.","from":"developer"},{"body":"So this is only a proble for HBCK2? Then I think it is the caller's duty to make sure that you do not do something wrong... HBCK2 is for fixing assignment problems, so the method is designed to skip all the checks and force to do an assignment...\r\n\r\nThis is the comment of the method\r\n\r\n{code}\r\n /**\r\n * A 'raw' version of assign that does bulk and skirts Master state checks (assigns can be made\r\n * during Master startup). For use by Hbck2.\r\n */\r\n{code}","from":"developer"},{"body":"{quote}Then I think it is the caller's duty to make sure that you do not do something wrong... HBCK2 is for fixing assignment problems, so the method is designed to skip all the checks and force to do an assignment..\r\n{quote}\r\nWe have an '--override' flag on assigns/unassigns in hbck2. When you pass this flag, I agree it should do as you describe above. When you do not pass this flag, if it can, the system should stop you shooting yourself in the foot with clear messaging on why it is not doing as you asked.\r\n\r\nMaking repairs on largish clusters with thousands of assignment problems, I make mistakes especially when running hasty bulk 'fixes'; i end up swapping one problem for a new one or as is the case here, doubling the problems that need fixing because I made the wrong call.\r\n\r\nLet me fix the comment too. It needs some color.","from":"developer"},{"body":"{quote}Want me to take this [~Bo Cui]. I had seen the second case occur dealing w/ a Replica that failed assign and had added investigation to my todo list. Thanks for digging in here.{quote}\r\n[~stack] thx, i have assigned to u.","from":"developer"},{"body":"I put up a PR....","from":"developer"},{"body":"Merged. Thanks for review [~zhangduo]\r\n\r\n ","from":"developer"}],"created":"2020-08-15T03:08:22.000+0000","description":"If a region has been assign to rs1 and then client assigns region again by \"hbck2 assigns\"\r\n\r\n1、if  regionPlan is region to be assign to rs2,the region will be opened on rs1 and rs2.\r\n\r\nmaster log:\r\n{quote}WARN org.apache.hadoop.hbase.master.assignment.AssignmentManager: rit=OPEN, location=rs2, table=tableName, region=reionName reported OPEN on server=rs1 but state has otherwise\r\n{quote}\r\n2、if regionPlan is region to be assign to rs1, the TransitRegionStateProcedure and OpenRegionProcedure will stuck. because rs1 is not responding to master\r\n rslog:\r\n{quote}Receiving OPEN for the region:{}, which we are already trying to OPEN - ignoring this new request for this region.\r\n{quote}\r\n ","issue_id":"13322844","key":"HBASE-24885","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-08-24T16:23:59.000+0000","role":"fixed_distractor","summary":"STUCK RIT by hbck2 assigns"} {"case_id":"13323810","cluster":"DISTRACTOR-HBASE-24916","comments":[{"body":"Can you give more details about which hbck2 command gives the problematic report? And are you saying two regions for different tables have the same name, in hbase meta (as per the \"r2\", in your example)?","created":"2020-08-21T08:18:23.558+0000"},{"body":"I think hole was created directly by deleting the region from meta. Hole was not created by any problematic hbck2 command.\r\n\r\nWhen a region hole is created by deleting last region of a table then hole has the region pair as \r\n\r\nBut when a region hole is created by deleting the first region then hole has the region pair as . \r\n\r\nThis seems wrong. This hole region pair should be \r\n\r\n","created":"2020-08-24T04:13:54.415+0000"},{"body":"Suppose a cluster with two tables t1 and t2. \r\nTable t1 with regions r1, r2, r3 and table t2 with regions s1,s2,s3\r\n\r\nWhen a region hole is created by deleting s3 then hole has the region pair as \r\n\r\nBut when a region hole is created by deleting s1 then hole has the region pair as . Here both regions belong to different tables\r\n\r\nThis seems wrong. This hole's region pair should be ","created":"2020-08-24T04:22:12.005+0000"},{"body":"{quote}\r\nSuppose a cluster with two tables t1 and t2.\r\nTable t1 with regions r1, r2, r3 and table t2 with regions s1,s2,s3\r\n\r\nWhen a region hole is created by deleting s3 then hole has the region pair as \r\n\r\nBut when a region hole is created by deleting s1 then hole has the region pair as . Here both regions belong to different tables\r\n\r\nThis seems wrong. This hole's region pair should be \r\n{quote}\r\n\r\nThanks for the explanation, now it's clear to me. Apparently, the meta scan is not differentiating regions for different tables.","created":"2020-08-28T15:17:53.757+0000"},{"body":"Thanks for digging in here [~arshad.mohammad] . Merged to branch-2.3+","created":"2020-08-30T17:07:58.316+0000"},{"body":"Thanks [~stack] for reviewing and merging it.\r\nThanks [~kanaka] for reviewing it. \r\nGood finding [~a00408367]. Thanks for reporting it.\r\n","created":"2020-08-31T14:22:53.135+0000"}],"conversations":[{"body":"HBCK2 Holes: Wrong region details in HBCK report ,if a region is missing in meta and start key is empty\r\n\r\nfor hole, it reports two region where hole is present.\r\n\r\nScenario where first region is missing for eg: in below scenario r1 is missing in meta\r\n\r\nstart key    end key   region\r\n\r\nempty     10        r1\r\n\r\n10           20        r2\r\n\r\n20           30        r2\r\n\r\n30          empty  r3\r\n\r\n \r\n\r\nIn our test hbck report contains two region where first region belongs to different table and second region is r2.\r\n\r\n ","from":"reporter","subject":"Region hole contains wrong regions pair when hole is created by first region deletion"},{"body":"Can you give more details about which hbck2 command gives the problematic report? And are you saying two regions for different tables have the same name, in hbase meta (as per the \"r2\", in your example)?","from":"developer"},{"body":"I think hole was created directly by deleting the region from meta. Hole was not created by any problematic hbck2 command.\r\n\r\nWhen a region hole is created by deleting last region of a table then hole has the region pair as \r\n\r\nBut when a region hole is created by deleting the first region then hole has the region pair as . \r\n\r\nThis seems wrong. This hole region pair should be \r\n\r\n","from":"developer"},{"body":"Suppose a cluster with two tables t1 and t2. \r\nTable t1 with regions r1, r2, r3 and table t2 with regions s1,s2,s3\r\n\r\nWhen a region hole is created by deleting s3 then hole has the region pair as \r\n\r\nBut when a region hole is created by deleting s1 then hole has the region pair as . Here both regions belong to different tables\r\n\r\nThis seems wrong. This hole's region pair should be ","from":"developer"},{"body":"{quote}\r\nSuppose a cluster with two tables t1 and t2.\r\nTable t1 with regions r1, r2, r3 and table t2 with regions s1,s2,s3\r\n\r\nWhen a region hole is created by deleting s3 then hole has the region pair as \r\n\r\nBut when a region hole is created by deleting s1 then hole has the region pair as . Here both regions belong to different tables\r\n\r\nThis seems wrong. This hole's region pair should be \r\n{quote}\r\n\r\nThanks for the explanation, now it's clear to me. Apparently, the meta scan is not differentiating regions for different tables.","from":"developer"},{"body":"Thanks for digging in here [~arshad.mohammad] . Merged to branch-2.3+","from":"developer"},{"body":"Thanks [~stack] for reviewing and merging it.\r\nThanks [~kanaka] for reviewing it. \r\nGood finding [~a00408367]. Thanks for reporting it.\r\n","from":"developer"}],"created":"2020-08-20T17:11:42.000+0000","description":"HBCK2 Holes: Wrong region details in HBCK report ,if a region is missing in meta and start key is empty\r\n\r\nfor hole, it reports two region where hole is present.\r\n\r\nScenario where first region is missing for eg: in below scenario r1 is missing in meta\r\n\r\nstart key    end key   region\r\n\r\nempty     10        r1\r\n\r\n10           20        r2\r\n\r\n20           30        r2\r\n\r\n30          empty  r3\r\n\r\n \r\n\r\nIn our test hbck report contains two region where first region belongs to different table and second region is r2.\r\n\r\n ","issue_id":"13323810","key":"HBASE-24916","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-08-30T17:07:58.000+0000","role":"fixed_distractor","summary":"Region hole contains wrong regions pair when hole is created by first region deletion"} {"case_id":"13329880","cluster":"DISTRACTOR-HBASE-25115","comments":[{"body":"This maybe cased by HFileBlockIndex#rootBlockContainingKey().\r\n\r\nWhen we use HFileBlockIndex#rootBlockContainingKey() with a cell created by PrivateCellUtil.createFirstOnRow(), it will return -1.","created":"2020-09-29T09:08:27.581+0000"},{"body":"a simple fix for this issue.","created":"2020-09-30T09:10:40.943+0000"},{"body":"[~zcq_rambo] Yes, your code is simplified, I will modify my PR on master branch","created":"2020-09-30T12:07:55.323+0000"},{"body":"[~niuyulin] [~zcq_rambo] I just realized that we have 2 PRs: one for master and another for branch-2.2, both from different authors.\r\n\r\nUsually, only Jira assignee creates PR and we create PR for just master branch as long as the patch is not huge and patch is not significantly different across multiple release branches. For now it's fine, will merge both PRs after some time. Also, master PR has test, so let me include that for branch-2.2 commit also.\r\n\r\nThanks to both of you for finding and fixing this nice bug.","created":"2020-10-04T10:26:37.508+0000"},{"body":"[~zcq_rambo] I have just added you to the contributers list, so going forward you can assign Jira to yourself if it is unassigned and you would like to work on it.","created":"2020-10-04T10:31:00.199+0000"}],"conversations":[{"body":"This issue can be reproduced by below steps:\r\n * make a hfile contains two rows '000' and '001';\r\n\r\n{code:java}\r\nD:\\bin>hbase hfile -p -f /hbase/data/default/test2/df76e4acab5398e70be332f6807ec3ba/f1/fda213c556d540a58d29d6bd85931dcd\r\nK: 000/f1:a/1601282789548/Put/vlen=4/seqid=4 V: aaaa\r\nK: 001/f1:a/1601282792779/Put/vlen=4/seqid=5 V: aaaa\r\nScanned kv count -> 2{code}\r\n * '001' can be seeked to;\r\n\r\n{code:java}\r\nD:\\bin>hbase hfile -e -w 001 -f /hbase/data/default/test2/df76e4acab5398e70be332f6807ec3ba/f1/fda213c556d540a58d29d6bd85931dcd\r\nK: 001/f1:a/1601282792779/Put/vlen=4/seqid=5\r\nScanned kv count -> 1{code}\r\n * but '000' can't be seeked to;\r\n\r\n{code:java}\r\nD:\\bin>hbase hfile -e -w 000 -f /hbase/data/default/test2/df76e4acab5398e70be332f6807ec3ba/f1/fda213c556d540a58d29d6bd85931dcd\r\nScanned kv count -> 0{code}\r\n In HFilePrettyPrinter we use \"scanner.seekTo(PrivateCellUtil.createFirstOnRow(this.row))\" to seek to row.But this method will retrurn -1 when the row is the first row of hfile.","from":"reporter","subject":"HFilePrettyPrinter can't seek to the row which is the first row of a hfile"},{"body":"This maybe cased by HFileBlockIndex#rootBlockContainingKey().\r\n\r\nWhen we use HFileBlockIndex#rootBlockContainingKey() with a cell created by PrivateCellUtil.createFirstOnRow(), it will return -1.","from":"developer"},{"body":"a simple fix for this issue.","from":"developer"},{"body":"[~zcq_rambo] Yes, your code is simplified, I will modify my PR on master branch","from":"developer"},{"body":"[~niuyulin] [~zcq_rambo] I just realized that we have 2 PRs: one for master and another for branch-2.2, both from different authors.\r\n\r\nUsually, only Jira assignee creates PR and we create PR for just master branch as long as the patch is not huge and patch is not significantly different across multiple release branches. For now it's fine, will merge both PRs after some time. Also, master PR has test, so let me include that for branch-2.2 commit also.\r\n\r\nThanks to both of you for finding and fixing this nice bug.","from":"developer"},{"body":"[~zcq_rambo] I have just added you to the contributers list, so going forward you can assign Jira to yourself if it is unassigned and you would like to work on it.","from":"developer"}],"created":"2020-09-29T08:02:55.000+0000","description":"This issue can be reproduced by below steps:\r\n * make a hfile contains two rows '000' and '001';\r\n\r\n{code:java}\r\nD:\\bin>hbase hfile -p -f /hbase/data/default/test2/df76e4acab5398e70be332f6807ec3ba/f1/fda213c556d540a58d29d6bd85931dcd\r\nK: 000/f1:a/1601282789548/Put/vlen=4/seqid=4 V: aaaa\r\nK: 001/f1:a/1601282792779/Put/vlen=4/seqid=5 V: aaaa\r\nScanned kv count -> 2{code}\r\n * '001' can be seeked to;\r\n\r\n{code:java}\r\nD:\\bin>hbase hfile -e -w 001 -f /hbase/data/default/test2/df76e4acab5398e70be332f6807ec3ba/f1/fda213c556d540a58d29d6bd85931dcd\r\nK: 001/f1:a/1601282792779/Put/vlen=4/seqid=5\r\nScanned kv count -> 1{code}\r\n * but '000' can't be seeked to;\r\n\r\n{code:java}\r\nD:\\bin>hbase hfile -e -w 000 -f /hbase/data/default/test2/df76e4acab5398e70be332f6807ec3ba/f1/fda213c556d540a58d29d6bd85931dcd\r\nScanned kv count -> 0{code}\r\n In HFilePrettyPrinter we use \"scanner.seekTo(PrivateCellUtil.createFirstOnRow(this.row))\" to seek to row.But this method will retrurn -1 when the row is the first row of hfile.","issue_id":"13329880","key":"HBASE-25115","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-10-04T10:55:53.000+0000","role":"fixed_distractor","summary":"HFilePrettyPrinter can't seek to the row which is the first row of a hfile"} {"case_id":"13330260","cluster":"DISTRACTOR-HBASE-25130","comments":[{"body":"FYI [~apurtell]","created":"2020-10-01T00:20:35.853+0000"},{"body":"Are you planning to submit a patch [~sandeep.guggilam]?","created":"2020-10-16T01:36:30.407+0000"},{"body":"[~apurtell] Didn't get time to go over the full workflow and design a solution of how to deal with this. I have two other JIRAs in my queue and might pick this up only after they are completed. Sorry for the delay. Let me unassign it for others to take it . WIll pick it up again if it remains unassigned once I am done with other JIRAs","created":"2020-10-16T05:46:54.547+0000"},{"body":"Thanks for the update [~sandeep.guggilam]","created":"2020-10-16T18:29:43.175+0000"},{"body":"[~vjasani] Can you please assign this to me. Thanks","created":"2020-12-14T14:29:58.618+0000"},{"body":"Done, you have HBase Jira contributor access now so you can assign Jiras to yourself going forward.","created":"2020-12-14T14:34:15.446+0000"},{"body":"Possible approaches I could think of tackling the above issue:\r\n\r\n1. Remove the the region entry from serverHolding map in case of mergeOverlap repair.\r\n * Once deleteMetaRegion() gets executed, we can call HbckRepair to remove the region from serverHoldings.\r\n * Here HbckRepair would call ServerManager and ServerManager will manage to call AssignmentManager to remove the region entry from serverHolding map via RegionStates.\r\n\r\n2. Keep cleaning out the unwanted entries from serverHoldings map via chore cleaner, if the entries for the region is not present in META. This could be an expensive operation, also would it be safe to cleanup based on the above logic ?\r\n\r\n[~apurtell] [~vjasani]  Please let me know your feedback on the above approaches or if we can handle it better. Thanks\r\n\r\n ","created":"2021-01-04T12:53:00.122+0000"},{"body":"[~rkrahul324] This is the issue:\r\n\r\nbq. the offline RPC doesn’t remove it from the serverHoldings map unless the new state is MERGED/SPLIT\r\n\r\nThe offline RPC should also remove the region from serverHoldings map if the new state is OFFLINE. ","created":"2021-01-20T18:19:40.212+0000"},{"body":"Thanks  [~apurtell]. That clear things up. ","created":"2021-01-21T07:34:33.417+0000"},{"body":"While trying to clear overlapped region entries from _serverHoldings_ map on running hbck repair. I noticed the HRI that \r\n\r\n_regionAssignments_(_TreeMap_) has an entry of does not equal to the HRI that has to be offlined. For eg \r\n\r\n*\r\n\r\n{ENCODED => fa18b66587f8f7a1de791ffefe364a48, NAME => 'test,,1611911426615.fa18b66587f8f7a1de791ffefe364a48.', STARTKEY => '', ENDKEY => ''}\r\n\r\n* is the metadata of _HRI_ in _regionAssignment_ map where as \r\n\r\n*\r\n\r\n{ENCODED => fa18b66587f8f7a1de791ffefe364a48, NAME => 'test,,1611911426615.fa18b66587f8f7a1de791ffefe364a48.', STARTKEY => '', ENDKEY => '', OFFLINE => true, SPLIT => true}\r\n\r\n* is the metadata of _HRI_ which has to go offline. \r\n So, while it tries to remove the HRI entry to be offline from _regionAssignment_ via _regionAssignment.remove(hri_), it was not able to find any and thus couldn't go ahead with _regionOffline_ operation further. \r\n\r\nI am confused here, as both _HRI_ objects i.e the one which has to go offline and the one in _regionAssignments_ map should point to same object?\r\n A random thought, do we need to update the logic of equals(on basis of _encodedRegionName_) for HRI so that both of the above considered as equals ? \r\n\r\n[~vjasani] [~apurtell] Can you please help. Thanks\r\n\r\nBtw, I reproed the overlap scenario via adding a bug in split and rollback scenario if that matters anyway.","created":"2021-01-29T11:51:44.174+0000"},{"body":"{quote}do we need to update the logic of equals(on basis of _encodedRegionName_) for HRI so that both of the above considered as equals ? \r\n{quote}\r\nI think comparing full regionName sounds better option here.","created":"2021-02-04T15:11:52.184+0000"},{"body":"hbck has been changed a lot in hbase 2, update this ticket to include branch-1 in title.","created":"2021-06-17T21:33:23.411+0000"},{"body":"https://github.com/apache/hbase/pull/3402","created":"2021-06-28T21:05:03.087+0000"}],"conversations":[{"body":"{color:#1d1c1d}Incase of repairing overlaps, hbck  essentially calls the closeRegion RPC on RS followed by offline RPC on Master to offline all the overlap regions that would be merged into a new region. {color}\r\n\r\n{color:#1d1c1d}However the offline RPC doesn’t remove it from the serverHoldings map unless the new state is MERGED/SPLIT ([https://github.com/apache/hbase/blob/branch-1/hbase-server/src/main/java/org/apache/hadoop/hbase/master/RegionStates.java#L719]) b{color}{color:#1d1c1d}ut the new state in this case is OFFLINE. {color}\r\n\r\n{color:#1d1c1d}This is actually intended to match with the META entries and would be removed later when the region is online on a different server. However, in our case , the region would never be online on a new server, hence the region info is never cleared from the map that is used by balancer and SCP for incorrect reeassignment.{color}\r\n\r\n{color:#1d1c1d}We might need to tackle this by removing the entries from the map when hbck actually deletes{color}{color:#1d1c1d} the meta entries for this region which kind of matches the in-memory map’s expectation with the META state.{color}","from":"reporter","subject":"[branch-1] Masters in-memory serverHoldings map is not cleared during hbck repair"},{"body":"FYI [~apurtell]","from":"developer"},{"body":"Are you planning to submit a patch [~sandeep.guggilam]?","from":"developer"},{"body":"[~apurtell] Didn't get time to go over the full workflow and design a solution of how to deal with this. I have two other JIRAs in my queue and might pick this up only after they are completed. Sorry for the delay. Let me unassign it for others to take it . WIll pick it up again if it remains unassigned once I am done with other JIRAs","from":"developer"},{"body":"Thanks for the update [~sandeep.guggilam]","from":"developer"},{"body":"[~vjasani] Can you please assign this to me. Thanks","from":"developer"},{"body":"Done, you have HBase Jira contributor access now so you can assign Jiras to yourself going forward.","from":"developer"},{"body":"Possible approaches I could think of tackling the above issue:\r\n\r\n1. Remove the the region entry from serverHolding map in case of mergeOverlap repair.\r\n * Once deleteMetaRegion() gets executed, we can call HbckRepair to remove the region from serverHoldings.\r\n * Here HbckRepair would call ServerManager and ServerManager will manage to call AssignmentManager to remove the region entry from serverHolding map via RegionStates.\r\n\r\n2. Keep cleaning out the unwanted entries from serverHoldings map via chore cleaner, if the entries for the region is not present in META. This could be an expensive operation, also would it be safe to cleanup based on the above logic ?\r\n\r\n[~apurtell] [~vjasani]  Please let me know your feedback on the above approaches or if we can handle it better. Thanks\r\n\r\n ","from":"developer"},{"body":"[~rkrahul324] This is the issue:\r\n\r\nbq. the offline RPC doesn’t remove it from the serverHoldings map unless the new state is MERGED/SPLIT\r\n\r\nThe offline RPC should also remove the region from serverHoldings map if the new state is OFFLINE. ","from":"developer"},{"body":"Thanks  [~apurtell]. That clear things up. ","from":"developer"},{"body":"While trying to clear overlapped region entries from _serverHoldings_ map on running hbck repair. I noticed the HRI that \r\n\r\n_regionAssignments_(_TreeMap_) has an entry of does not equal to the HRI that has to be offlined. For eg \r\n\r\n*\r\n\r\n{ENCODED => fa18b66587f8f7a1de791ffefe364a48, NAME => 'test,,1611911426615.fa18b66587f8f7a1de791ffefe364a48.', STARTKEY => '', ENDKEY => ''}\r\n\r\n* is the metadata of _HRI_ in _regionAssignment_ map where as \r\n\r\n*\r\n\r\n{ENCODED => fa18b66587f8f7a1de791ffefe364a48, NAME => 'test,,1611911426615.fa18b66587f8f7a1de791ffefe364a48.', STARTKEY => '', ENDKEY => '', OFFLINE => true, SPLIT => true}\r\n\r\n* is the metadata of _HRI_ which has to go offline. \r\n So, while it tries to remove the HRI entry to be offline from _regionAssignment_ via _regionAssignment.remove(hri_), it was not able to find any and thus couldn't go ahead with _regionOffline_ operation further. \r\n\r\nI am confused here, as both _HRI_ objects i.e the one which has to go offline and the one in _regionAssignments_ map should point to same object?\r\n A random thought, do we need to update the logic of equals(on basis of _encodedRegionName_) for HRI so that both of the above considered as equals ? \r\n\r\n[~vjasani] [~apurtell] Can you please help. Thanks\r\n\r\nBtw, I reproed the overlap scenario via adding a bug in split and rollback scenario if that matters anyway.","from":"developer"},{"body":"{quote}do we need to update the logic of equals(on basis of _encodedRegionName_) for HRI so that both of the above considered as equals ? \r\n{quote}\r\nI think comparing full regionName sounds better option here.","from":"developer"},{"body":"hbck has been changed a lot in hbase 2, update this ticket to include branch-1 in title.","from":"developer"},{"body":"https://github.com/apache/hbase/pull/3402","from":"developer"}],"created":"2020-10-01T00:20:20.000+0000","description":"{color:#1d1c1d}Incase of repairing overlaps, hbck  essentially calls the closeRegion RPC on RS followed by offline RPC on Master to offline all the overlap regions that would be merged into a new region. {color}\r\n\r\n{color:#1d1c1d}However the offline RPC doesn’t remove it from the serverHoldings map unless the new state is MERGED/SPLIT ([https://github.com/apache/hbase/blob/branch-1/hbase-server/src/main/java/org/apache/hadoop/hbase/master/RegionStates.java#L719]) b{color}{color:#1d1c1d}ut the new state in this case is OFFLINE. {color}\r\n\r\n{color:#1d1c1d}This is actually intended to match with the META entries and would be removed later when the region is online on a different server. However, in our case , the region would never be online on a new server, hence the region info is never cleared from the map that is used by balancer and SCP for incorrect reeassignment.{color}\r\n\r\n{color:#1d1c1d}We might need to tackle this by removing the entries from the map when hbck actually deletes{color}{color:#1d1c1d} the meta entries for this region which kind of matches the in-memory map’s expectation with the META state.{color}","issue_id":"13330260","key":"HBASE-25130","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-06-28T21:05:03.000+0000","role":"fixed_distractor","summary":"[branch-1] Masters in-memory serverHoldings map is not cleared during hbck repair"} {"case_id":"12463722","cluster":"DISTRACTOR-HBASE-2514","comments":[{"body":"I ran into this when upgrading to 0.20.5-rc5. I forgot to copy the hadoop-gpl-compression jar. The RS didn't shut down though, it just kept spewing those exceptions.","created":"2010-06-22T17:12:28.210+0000"},{"body":"Relating to HBASE-2681. HBASE-2681 says to fall back to NONE if lib is not available. It even has a patch. I don't think that right after thinking on it. I think refusing to open it is way to go. Ryan brings up point over in the other issue that while the CF might say LZO, the store files may have be written with something else, some other compression that we need a lib for. So, on open, its not good enough to look at the CF config. only. We need to as part of storefile open -- which happens as part of region deploy anyways -- at this time verify that we have the codecs we need to deploy region.","created":"2010-07-24T03:59:54.994+0000"},{"body":"Making this critical and bringing into 0.90.0. I think this of import given we want to talk up use of lzo.","created":"2010-07-24T04:02:48.566+0000"},{"body":"We hit that problem today on some of our prod machines at StumbleUpon. {{!cool}}.","created":"2010-08-04T08:47:46.408+0000"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/719/\n-----------------------------------------------------------\n\nReview request for hbase.\n\n\nSummary\n-------\n\nTrying to make some progress on this issue, here's an untested patch that adds testing of codecs when the region server starts, then uses that to check the schema of regions it tries to open. That's basically what this jira's scope is about.\n\nSome things I don't like the way I did it:\n - For all the users that don't do LZO, they will have a WARN every time a RS starts (in its log)\n - Does not cover cases where the files are LZOed, but the schema is NONE or GZ\n\nFinally, there's no way to let user know about the errors unless he checks the logs, but a least we won't throw ugly exceptions.\n\n\nThis addresses bug HBASE-2514.\n http://issues.apache.org/jira/browse/HBASE-2514\n\n\nDiffs\n-----\n\n /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java 989204 \n /trunk/src/main/java/org/apache/hadoop/hbase/util/CompressionTest.java 989204 \n\nDiff: http://review.cloudera.org/r/719/diff\n\n\nTesting\n-------\n\n\nThanks,\n\nJean-Daniel\n\n\n","created":"2010-08-25T19:09:27.027+0000"},{"body":"im thinking we might want to do a different approach, since checking the schema isnt comprehensive enough as JD pointed out above. Instead if we had a list of \"required codecs\" we could abort the regionserver during startup. This would allow the admin to take action to correct the deployment.\n\nI'm not sure there is a good reason to have only part of your RS being able to open LZO regions and others passing and having the master bounce regions around until they stick.","created":"2010-10-18T23:12:38.848+0000"},{"body":"Message from: \"Ryan Rawson\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/1046/\n-----------------------------------------------------------\n\n(Updated 2010-10-22 16:38:10.455573)\n\n\nReview request for hbase.\n\n\nSummary\n-------\n\nRegionServer should refuse to be assigned a region that use LZO when LZO isn't available\n\n- instead of checking regions during open, i do a check and throw during the construction of the RS that is is able to instantiate the listed codecs.\n\n\nThis addresses bug HBASE-2514.\n http://issues.apache.org/jira/browse/HBASE-2514\n\n\nDiffs (updated)\n-----\n\n trunk/CHANGES.txt 1024074 \n trunk/src/main/java/org/apache/hadoop/hbase/io/hfile/Compression.java 1024073 \n trunk/src/main/java/org/apache/hadoop/hbase/io/hfile/HFile.java 1024073 \n trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java 1024073 \n trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java 1024074 \n trunk/src/main/java/org/apache/hadoop/hbase/util/CompressionTest.java 1024073 \n trunk/src/test/java/org/apache/hadoop/hbase/util/TestCompressionTest.java PRE-CREATION \n\nDiff: http://review.cloudera.org/r/1046/diff\n\n\nTesting\n-------\n\n\nThanks,\n\nRyan\n\n\n","created":"2010-10-22T23:39:14.921+0000"},{"body":"committed, with extra feature which needs documentation: HBASE-3146","created":"2010-10-23T00:40:32.018+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T12:43:44.713+0000"}],"conversations":[{"body":"If a RegionServer is assigned a region that uses LZO but the required libraries aren't installed on that RegionServer, the server will fail unexpectedly after throwing a {{java.lang.ClassNotFoundException: com.hadoop.compression.lzo.LzoCodec}}\n\n{code}\n\n2010-05-04 16:57:27,258 FATAL org.apache.hadoop.hbase.regionserver.MemStoreFlusher: Replay of hlog required. Forcing server shutdown\norg.apache.hadoop.hbase.DroppedSnapshotException: region: tsdb,,1273011287339\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:994)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:887)\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:255)\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.run(MemStoreFlusher.java:142)\nCaused by: java.lang.RuntimeException: java.lang.ClassNotFoundException: com.hadoop.compression.lzo.LzoCodec\n at org.apache.hadoop.hbase.io.hfile.Compression$Algorithm$1.getCodec(Compression.java:91)\n at org.apache.hadoop.hbase.io.hfile.Compression$Algorithm.getCompressor(Compression.java:196)\n at org.apache.hadoop.hbase.io.hfile.HFile$Writer.getCompressingStream(HFile.java:388)\n at org.apache.hadoop.hbase.io.hfile.HFile$Writer.newBlock(HFile.java:374)\n at org.apache.hadoop.hbase.io.hfile.HFile$Writer.checkBlockBoundary(HFile.java:345)\n at org.apache.hadoop.hbase.io.hfile.HFile$Writer.append(HFile.java:517)\n at org.apache.hadoop.hbase.io.hfile.HFile$Writer.append(HFile.java:482)\n at org.apache.hadoop.hbase.regionserver.Store.internalFlushCache(Store.java:558)\n at org.apache.hadoop.hbase.regionserver.Store.flushCache(Store.java:522)\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:979)\n ... 3 more\nCaused by: java.lang.ClassNotFoundException: com.hadoop.compression.lzo.LzoCodec\n at java.net.URLClassLoader$1.run(URLClassLoader.java:200)\n at java.security.AccessController.doPrivileged(Native Method)\n at java.net.URLClassLoader.findClass(URLClassLoader.java:188)\n at java.lang.ClassLoader.loadClass(ClassLoader.java:315)\n at sun.misc.Launcher$AppClassLoader.loadClass(Launcher.java:330)\n at java.lang.ClassLoader.loadClass(ClassLoader.java:250)\n at org.apache.hadoop.hbase.io.hfile.Compression$Algorithm$1.getCodec(Compression.java:87)\n ... 12 more\n{code}","from":"reporter","subject":"RegionServer should refuse to be assigned a region that use LZO when LZO isn't available"},{"body":"I ran into this when upgrading to 0.20.5-rc5. I forgot to copy the hadoop-gpl-compression jar. The RS didn't shut down though, it just kept spewing those exceptions.","from":"developer"},{"body":"Relating to HBASE-2681. HBASE-2681 says to fall back to NONE if lib is not available. It even has a patch. I don't think that right after thinking on it. I think refusing to open it is way to go. Ryan brings up point over in the other issue that while the CF might say LZO, the store files may have be written with something else, some other compression that we need a lib for. So, on open, its not good enough to look at the CF config. only. We need to as part of storefile open -- which happens as part of region deploy anyways -- at this time verify that we have the codecs we need to deploy region.","from":"developer"},{"body":"Making this critical and bringing into 0.90.0. I think this of import given we want to talk up use of lzo.","from":"developer"},{"body":"We hit that problem today on some of our prod machines at StumbleUpon. {{!cool}}.","from":"developer"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/719/\n-----------------------------------------------------------\n\nReview request for hbase.\n\n\nSummary\n-------\n\nTrying to make some progress on this issue, here's an untested patch that adds testing of codecs when the region server starts, then uses that to check the schema of regions it tries to open. That's basically what this jira's scope is about.\n\nSome things I don't like the way I did it:\n - For all the users that don't do LZO, they will have a WARN every time a RS starts (in its log)\n - Does not cover cases where the files are LZOed, but the schema is NONE or GZ\n\nFinally, there's no way to let user know about the errors unless he checks the logs, but a least we won't throw ugly exceptions.\n\n\nThis addresses bug HBASE-2514.\n http://issues.apache.org/jira/browse/HBASE-2514\n\n\nDiffs\n-----\n\n /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java 989204 \n /trunk/src/main/java/org/apache/hadoop/hbase/util/CompressionTest.java 989204 \n\nDiff: http://review.cloudera.org/r/719/diff\n\n\nTesting\n-------\n\n\nThanks,\n\nJean-Daniel\n\n\n","from":"developer"},{"body":"im thinking we might want to do a different approach, since checking the schema isnt comprehensive enough as JD pointed out above. Instead if we had a list of \"required codecs\" we could abort the regionserver during startup. This would allow the admin to take action to correct the deployment.\n\nI'm not sure there is a good reason to have only part of your RS being able to open LZO regions and others passing and having the master bounce regions around until they stick.","from":"developer"},{"body":"Message from: \"Ryan Rawson\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/1046/\n-----------------------------------------------------------\n\n(Updated 2010-10-22 16:38:10.455573)\n\n\nReview request for hbase.\n\n\nSummary\n-------\n\nRegionServer should refuse to be assigned a region that use LZO when LZO isn't available\n\n- instead of checking regions during open, i do a check and throw during the construction of the RS that is is able to instantiate the listed codecs.\n\n\nThis addresses bug HBASE-2514.\n http://issues.apache.org/jira/browse/HBASE-2514\n\n\nDiffs (updated)\n-----\n\n trunk/CHANGES.txt 1024074 \n trunk/src/main/java/org/apache/hadoop/hbase/io/hfile/Compression.java 1024073 \n trunk/src/main/java/org/apache/hadoop/hbase/io/hfile/HFile.java 1024073 \n trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java 1024073 \n trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java 1024074 \n trunk/src/main/java/org/apache/hadoop/hbase/util/CompressionTest.java 1024073 \n trunk/src/test/java/org/apache/hadoop/hbase/util/TestCompressionTest.java PRE-CREATION \n\nDiff: http://review.cloudera.org/r/1046/diff\n\n\nTesting\n-------\n\n\nThanks,\n\nRyan\n\n\n","from":"developer"},{"body":"committed, with extra feature which needs documentation: HBASE-3146","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2010-05-05T00:06:55.000+0000","description":"If a RegionServer is assigned a region that uses LZO but the required libraries aren't installed on that RegionServer, the server will fail unexpectedly after throwing a {{java.lang.ClassNotFoundException: com.hadoop.compression.lzo.LzoCodec}}\n\n{code}\n\n2010-05-04 16:57:27,258 FATAL org.apache.hadoop.hbase.regionserver.MemStoreFlusher: Replay of hlog required. Forcing server shutdown\norg.apache.hadoop.hbase.DroppedSnapshotException: region: tsdb,,1273011287339\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:994)\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:887)\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:255)\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.run(MemStoreFlusher.java:142)\nCaused by: java.lang.RuntimeException: java.lang.ClassNotFoundException: com.hadoop.compression.lzo.LzoCodec\n at org.apache.hadoop.hbase.io.hfile.Compression$Algorithm$1.getCodec(Compression.java:91)\n at org.apache.hadoop.hbase.io.hfile.Compression$Algorithm.getCompressor(Compression.java:196)\n at org.apache.hadoop.hbase.io.hfile.HFile$Writer.getCompressingStream(HFile.java:388)\n at org.apache.hadoop.hbase.io.hfile.HFile$Writer.newBlock(HFile.java:374)\n at org.apache.hadoop.hbase.io.hfile.HFile$Writer.checkBlockBoundary(HFile.java:345)\n at org.apache.hadoop.hbase.io.hfile.HFile$Writer.append(HFile.java:517)\n at org.apache.hadoop.hbase.io.hfile.HFile$Writer.append(HFile.java:482)\n at org.apache.hadoop.hbase.regionserver.Store.internalFlushCache(Store.java:558)\n at org.apache.hadoop.hbase.regionserver.Store.flushCache(Store.java:522)\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:979)\n ... 3 more\nCaused by: java.lang.ClassNotFoundException: com.hadoop.compression.lzo.LzoCodec\n at java.net.URLClassLoader$1.run(URLClassLoader.java:200)\n at java.security.AccessController.doPrivileged(Native Method)\n at java.net.URLClassLoader.findClass(URLClassLoader.java:188)\n at java.lang.ClassLoader.loadClass(ClassLoader.java:315)\n at sun.misc.Launcher$AppClassLoader.loadClass(Launcher.java:330)\n at java.lang.ClassLoader.loadClass(ClassLoader.java:250)\n at org.apache.hadoop.hbase.io.hfile.Compression$Algorithm$1.getCodec(Compression.java:87)\n ... 12 more\n{code}","issue_id":"12463722","key":"HBASE-2514","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2010-10-23T00:40:32.000+0000","role":"fixed_distractor","summary":"RegionServer should refuse to be assigned a region that use LZO when LZO isn't available"} {"case_id":"13339359","cluster":"DISTRACTOR-HBASE-25255","comments":[{"body":"There is no xxx-output file so upload the xml here.","created":"2020-11-08T03:39:21.043+0000"},{"body":"OK, I think the problem here is because a race between loading meta and create rs group table.\r\n\r\nIn CreateTableProcedure.waitInitialized, we have this\r\n\r\n{code}\r\n @Override\r\n protected boolean waitInitialized(MasterProcedureEnv env) {\r\n if (getTableName().isSystemTable()) {\r\n // Creating system table is part of the initialization, so do not wait here.\r\n return false;\r\n }\r\n return super.waitInitialized(env);\r\n }\r\n{code}\r\n\r\nWhich means when creating rs group table, we will not wait for meta loaded, so it is possible that before meta loaded is finished, we could add the region state node for rsgroup table to AssignmentManager with OFFLINE state, and then we will call processOfflineRegions to assign the offline regions, but at the same time, the CreateTableProcedure will assign it too, thus we get this assertion error.\r\n\r\nOn master branch, meta table does not rely on any other system tables so I think we could just change the above method to let CreateTableProcedure always wait meta loaded, as we do not use CreateTableProcedure to create meta, and also carefully change the order of when to call processOfflineRegions.\r\n\r\nBut on branch-2, we still need to create namespace table... So we need to find another solution for branch-2.\r\n\r\n[~stack] FYI.","created":"2020-11-08T03:54:30.775+0000"},{"body":"{quote} it is possible that before meta loaded is finished, we could add the region state node for rsgroup table to AssignmentManager with OFFLINE state, and then we will call processOfflineRegions to assign the offline regions, but at the same time, the CreateTableProcedure will assign it too, thus we get this assertion error.\r\n{quote}\r\nOther system tables have same problem?","created":"2020-11-13T00:46:02.360+0000"},{"body":"{quote}\r\nOther system tables have same problem?\r\n{quote}\r\n\r\nIt think so. In the PR here I moved the meta loaded event trigger after processOfflineRegions, and modified CreateTableProcedure to always wait for meta loaded so we will not create other system tables before processOfflineRegions are done.\r\nOn branch-2 there is another problem that namespace table must be created before meta table, so CreateTableProcedure need to check whether it is creating namespace table, if so it should not wait for meta loaded event.","created":"2020-11-13T02:13:04.764+0000"},{"body":"Oh, the waitInitialized of MasterProedureEnv will wait until master is initialized, which is not suitable for system table I suppose. Let me push an addendum for wait for meta loaded for all the system tables.","created":"2020-11-13T07:15:15.341+0000"},{"body":"Pushed to branch-2.2+.\r\n\r\nThanks [~zghao] for reviewing.","created":"2020-11-13T08:26:17.219+0000"}],"conversations":[{"body":"Saw this when setup TestRSGroupsKillRS\r\n\r\n{noformat}\r\n2020-11-07 16:29:54,565 ERROR [master/e476f4f509a7:0:becomeActiveMaster] helpers.MarkerIgnoringBase(159): Failed to become active master\r\njava.lang.AssertionError\r\n\tat org.apache.hadoop.hbase.master.assignment.RegionStateNode.setProcedure(RegionStateNode.java:198)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.createAssignProcedure(AssignmentManager.java:647)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.lambda$null$6(AssignmentManager.java:878)\r\n\tat java.util.stream.ReferencePipeline$3$1.accept(ReferencePipeline.java:193)\r\n\tat java.util.stream.ReferencePipeline$3$1.accept(ReferencePipeline.java:193)\r\n\tat java.util.ArrayList$ArrayListSpliterator.forEachRemaining(ArrayList.java:1382)\r\n\tat java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:482)\r\n\tat java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:472)\r\n\tat java.util.stream.ForEachOps$ForEachOp.evaluateSequential(ForEachOps.java:150)\r\n\tat java.util.stream.ForEachOps$ForEachOp$OfRef.evaluateSequential(ForEachOps.java:173)\r\n\tat java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:234)\r\n\tat java.util.stream.ReferencePipeline.forEach(ReferencePipeline.java:485)\r\n\tat java.util.stream.ReferencePipeline$7$1.accept(ReferencePipeline.java:272)\r\n\tat java.util.HashMap$EntrySpliterator.forEachRemaining(HashMap.java:1699)\r\n\tat java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:482)\r\n\tat java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:472)\r\n\tat java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:546)\r\n\tat java.util.stream.AbstractPipeline.evaluateToArrayNode(AbstractPipeline.java:260)\r\n\tat java.util.stream.ReferencePipeline.toArray(ReferencePipeline.java:505)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.createAssignProcedures(AssignmentManager.java:879)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.createRoundRobinAssignProcedures(AssignmentManager.java:759)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.createRoundRobinAssignProcedures(AssignmentManager.java:775)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.processOfflineRegions(AssignmentManager.java:1513)\r\n\tat org.apache.hadoop.hbase.master.HMaster.finishActiveMasterInitialization(HMaster.java:1012)\r\n\tat org.apache.hadoop.hbase.master.HMaster.startActiveMasterManager(HMaster.java:2116)\r\n\tat org.apache.hadoop.hbase.master.HMaster.lambda$run$0(HMaster.java:515)\r\n\tat java.lang.Thread.run(Thread.java:748)\r\n{noformat}","from":"reporter","subject":"Master fails to initialize when creating rs group table"},{"body":"There is no xxx-output file so upload the xml here.","from":"developer"},{"body":"OK, I think the problem here is because a race between loading meta and create rs group table.\r\n\r\nIn CreateTableProcedure.waitInitialized, we have this\r\n\r\n{code}\r\n @Override\r\n protected boolean waitInitialized(MasterProcedureEnv env) {\r\n if (getTableName().isSystemTable()) {\r\n // Creating system table is part of the initialization, so do not wait here.\r\n return false;\r\n }\r\n return super.waitInitialized(env);\r\n }\r\n{code}\r\n\r\nWhich means when creating rs group table, we will not wait for meta loaded, so it is possible that before meta loaded is finished, we could add the region state node for rsgroup table to AssignmentManager with OFFLINE state, and then we will call processOfflineRegions to assign the offline regions, but at the same time, the CreateTableProcedure will assign it too, thus we get this assertion error.\r\n\r\nOn master branch, meta table does not rely on any other system tables so I think we could just change the above method to let CreateTableProcedure always wait meta loaded, as we do not use CreateTableProcedure to create meta, and also carefully change the order of when to call processOfflineRegions.\r\n\r\nBut on branch-2, we still need to create namespace table... So we need to find another solution for branch-2.\r\n\r\n[~stack] FYI.","from":"developer"},{"body":"{quote} it is possible that before meta loaded is finished, we could add the region state node for rsgroup table to AssignmentManager with OFFLINE state, and then we will call processOfflineRegions to assign the offline regions, but at the same time, the CreateTableProcedure will assign it too, thus we get this assertion error.\r\n{quote}\r\nOther system tables have same problem?","from":"developer"},{"body":"{quote}\r\nOther system tables have same problem?\r\n{quote}\r\n\r\nIt think so. In the PR here I moved the meta loaded event trigger after processOfflineRegions, and modified CreateTableProcedure to always wait for meta loaded so we will not create other system tables before processOfflineRegions are done.\r\nOn branch-2 there is another problem that namespace table must be created before meta table, so CreateTableProcedure need to check whether it is creating namespace table, if so it should not wait for meta loaded event.","from":"developer"},{"body":"Oh, the waitInitialized of MasterProedureEnv will wait until master is initialized, which is not suitable for system table I suppose. Let me push an addendum for wait for meta loaded for all the system tables.","from":"developer"},{"body":"Pushed to branch-2.2+.\r\n\r\nThanks [~zghao] for reviewing.","from":"developer"}],"created":"2020-11-08T03:37:57.000+0000","description":"Saw this when setup TestRSGroupsKillRS\r\n\r\n{noformat}\r\n2020-11-07 16:29:54,565 ERROR [master/e476f4f509a7:0:becomeActiveMaster] helpers.MarkerIgnoringBase(159): Failed to become active master\r\njava.lang.AssertionError\r\n\tat org.apache.hadoop.hbase.master.assignment.RegionStateNode.setProcedure(RegionStateNode.java:198)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.createAssignProcedure(AssignmentManager.java:647)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.lambda$null$6(AssignmentManager.java:878)\r\n\tat java.util.stream.ReferencePipeline$3$1.accept(ReferencePipeline.java:193)\r\n\tat java.util.stream.ReferencePipeline$3$1.accept(ReferencePipeline.java:193)\r\n\tat java.util.ArrayList$ArrayListSpliterator.forEachRemaining(ArrayList.java:1382)\r\n\tat java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:482)\r\n\tat java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:472)\r\n\tat java.util.stream.ForEachOps$ForEachOp.evaluateSequential(ForEachOps.java:150)\r\n\tat java.util.stream.ForEachOps$ForEachOp$OfRef.evaluateSequential(ForEachOps.java:173)\r\n\tat java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:234)\r\n\tat java.util.stream.ReferencePipeline.forEach(ReferencePipeline.java:485)\r\n\tat java.util.stream.ReferencePipeline$7$1.accept(ReferencePipeline.java:272)\r\n\tat java.util.HashMap$EntrySpliterator.forEachRemaining(HashMap.java:1699)\r\n\tat java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:482)\r\n\tat java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:472)\r\n\tat java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:546)\r\n\tat java.util.stream.AbstractPipeline.evaluateToArrayNode(AbstractPipeline.java:260)\r\n\tat java.util.stream.ReferencePipeline.toArray(ReferencePipeline.java:505)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.createAssignProcedures(AssignmentManager.java:879)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.createRoundRobinAssignProcedures(AssignmentManager.java:759)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.createRoundRobinAssignProcedures(AssignmentManager.java:775)\r\n\tat org.apache.hadoop.hbase.master.assignment.AssignmentManager.processOfflineRegions(AssignmentManager.java:1513)\r\n\tat org.apache.hadoop.hbase.master.HMaster.finishActiveMasterInitialization(HMaster.java:1012)\r\n\tat org.apache.hadoop.hbase.master.HMaster.startActiveMasterManager(HMaster.java:2116)\r\n\tat org.apache.hadoop.hbase.master.HMaster.lambda$run$0(HMaster.java:515)\r\n\tat java.lang.Thread.run(Thread.java:748)\r\n{noformat}","issue_id":"13339359","key":"HBASE-25255","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-11-13T08:26:17.000+0000","role":"fixed_distractor","summary":"Master fails to initialize when creating rs group table"} {"case_id":"13340890","cluster":"DISTRACTOR-HBASE-25292","comments":[{"body":"Porting changes from tested internal forks now. Should have PRs by the end of the day.","created":"2020-11-16T19:07:49.225+0000"},{"body":"PRs are up. It look longer for tests to run than anticipated.","created":"2020-11-17T23:19:18.432+0000"}],"conversations":[{"body":"We sometimes cache InetSocketAddress in data structures in an attempt to optimize away potential nameservice (DNS) lookups. This is, in general, an anti-pattern, because once an InetSocketAddress is resolved, resolution is never attempted again. The ideal pattern for connect() is ISA instantiation just before the connect() call, with no reuse of the ISA instance. For bind() we presume the local identity won't change while the process is live so usage and caching can be relaxed in that case.\r\n\r\nIf I can restate my proposal for a usage convention for InetSocketAddress, it would be this: Network identities should be bound late. This means addresses should be resolved at the last possible moment. Also, network identity mappings can change, so our code should not inappropriately cache them; otherwise we might miss a change and fail to operate normally.\r\n\r\nI have reviewed the code for InetSocketAddress usage and in my opinion sometimes we are caching ISA acceptably, and in other cases we are not.\r\n\r\nCorrect cases:\r\n * We cache ISA for RPC connections, so we don't potentially do a lookup for every Call. However, we resolve the address earlier than we need to. The code can be improved by moving resolution to just before where we connect().\r\n * Use of ISA with bind. Typical uses like bindAddress, listenerAddress, initialIsa, or localAddress.\r\n ** (There is no harm to keep direct use of ISA for bind() but these could all be replaced with Address and on-demand create of ISA just before bind().\r\n * ClusterStatusPublisher in master.\r\n * Netty integration. Netty accepts and supplies ISA, no choice there, but we want to resolve at channel create time and cache for lifetime of the channel anyway. If a remote host goes away and is replaced with a new identity, a new channel/connection will be created with a new resolution just before connect() via higher layer error handling.\r\n\r\nIncorrect cases that can be fixed:\r\n * RPC stubs. Remote clients may be recycled and replaced with new instances where the network identities (DNS name to IP address mapping) have changed--. HBASE-14544 attempts to work around DNS instability in data centers of years past in a way that, in my opinion, is the wrong thing to do in the modern era. In modern datacenters, in public cloud, and especially in kubernetes environments, DNS mappings are dynamic and subject to frequent change. It is just never the right thing to do to cache them. I intend to propose a revert of HBASE-14544. Reverting this simplifies some code a bit. That is the only reason: this is in my opinion some legacy that can be dropped, one fewer configuration variable (yay!), but if this part of the proposal is controversial it can be skipped.\r\n * RPC stubs again. When looking up the IP address of the remote host when creating a stub key we also make a key even if the resolution fails. This is the wrong thing to do. If we can't resolve the remote address, we can't contact the server. Making a stub that can't communicate is pointless. Throw an exception instead.\r\n * Favored nodes. Although the HDFS API requires InetSocketAddress, we don't have to make up a list right away and cache them forever. We can use Address to record the list of favored nodes and convert from Address to InetSocketAddress on demand (when we go to create the HFile). This will allow us to resolve datanode hostnames just before they are needed. In public cloud, kubernetes, and or some private datacenter service deployment options, datanode servers may have their network identities (DNS name -> IP address mapping) changed over time. We can and should avoid inappropriate caching that may cause us to indefinitely use an incorrect address when contacting a favored node. \r\n * Sometimes we use ISA when Address is just as good. For example, the dead servers list. If we are going to pay some attention to ISA usage discipline, let's remove the cases where we use ISA as a host and port pair but do not need to do so. Address works just as well and doesn't present an opportunity for misuse. Another example would be the RPC client concurrentCounterCache.\r\n ** We could do a lot more substitutions than what is proposed. All of the ISA uses with bind() are okay as is but could also be updated.\r\n\r\nIncorrect cases that cannot be fixed:\r\n * hbase-external-blockcache: We have to resolve all of the memcached locations up front because the memcached client constructor requires ISA instances. So we have to hope that the network identities (DNS name -> IP address mapping) does not change for any in the list. This is beyond our control.\r\n\r\nWhile in this area it is trivial to add new client connect metrics for number of potential nameservice lookups (whenever we instantiate an ISA) and number of failed nameservice lookups (if the instantiated ISA is unresolved).\r\n\r\nWhile in this area I also noticed we often directly access a field in ConnectionId where there is also a getter, so good practice is to use the getter instead.","from":"reporter","subject":"Improve InetSocketAddress usage discipline"},{"body":"Porting changes from tested internal forks now. Should have PRs by the end of the day.","from":"developer"},{"body":"PRs are up. It look longer for tests to run than anticipated.","from":"developer"}],"created":"2020-11-16T18:14:56.000+0000","description":"We sometimes cache InetSocketAddress in data structures in an attempt to optimize away potential nameservice (DNS) lookups. This is, in general, an anti-pattern, because once an InetSocketAddress is resolved, resolution is never attempted again. The ideal pattern for connect() is ISA instantiation just before the connect() call, with no reuse of the ISA instance. For bind() we presume the local identity won't change while the process is live so usage and caching can be relaxed in that case.\r\n\r\nIf I can restate my proposal for a usage convention for InetSocketAddress, it would be this: Network identities should be bound late. This means addresses should be resolved at the last possible moment. Also, network identity mappings can change, so our code should not inappropriately cache them; otherwise we might miss a change and fail to operate normally.\r\n\r\nI have reviewed the code for InetSocketAddress usage and in my opinion sometimes we are caching ISA acceptably, and in other cases we are not.\r\n\r\nCorrect cases:\r\n * We cache ISA for RPC connections, so we don't potentially do a lookup for every Call. However, we resolve the address earlier than we need to. The code can be improved by moving resolution to just before where we connect().\r\n * Use of ISA with bind. Typical uses like bindAddress, listenerAddress, initialIsa, or localAddress.\r\n ** (There is no harm to keep direct use of ISA for bind() but these could all be replaced with Address and on-demand create of ISA just before bind().\r\n * ClusterStatusPublisher in master.\r\n * Netty integration. Netty accepts and supplies ISA, no choice there, but we want to resolve at channel create time and cache for lifetime of the channel anyway. If a remote host goes away and is replaced with a new identity, a new channel/connection will be created with a new resolution just before connect() via higher layer error handling.\r\n\r\nIncorrect cases that can be fixed:\r\n * RPC stubs. Remote clients may be recycled and replaced with new instances where the network identities (DNS name to IP address mapping) have changed--. HBASE-14544 attempts to work around DNS instability in data centers of years past in a way that, in my opinion, is the wrong thing to do in the modern era. In modern datacenters, in public cloud, and especially in kubernetes environments, DNS mappings are dynamic and subject to frequent change. It is just never the right thing to do to cache them. I intend to propose a revert of HBASE-14544. Reverting this simplifies some code a bit. That is the only reason: this is in my opinion some legacy that can be dropped, one fewer configuration variable (yay!), but if this part of the proposal is controversial it can be skipped.\r\n * RPC stubs again. When looking up the IP address of the remote host when creating a stub key we also make a key even if the resolution fails. This is the wrong thing to do. If we can't resolve the remote address, we can't contact the server. Making a stub that can't communicate is pointless. Throw an exception instead.\r\n * Favored nodes. Although the HDFS API requires InetSocketAddress, we don't have to make up a list right away and cache them forever. We can use Address to record the list of favored nodes and convert from Address to InetSocketAddress on demand (when we go to create the HFile). This will allow us to resolve datanode hostnames just before they are needed. In public cloud, kubernetes, and or some private datacenter service deployment options, datanode servers may have their network identities (DNS name -> IP address mapping) changed over time. We can and should avoid inappropriate caching that may cause us to indefinitely use an incorrect address when contacting a favored node. \r\n * Sometimes we use ISA when Address is just as good. For example, the dead servers list. If we are going to pay some attention to ISA usage discipline, let's remove the cases where we use ISA as a host and port pair but do not need to do so. Address works just as well and doesn't present an opportunity for misuse. Another example would be the RPC client concurrentCounterCache.\r\n ** We could do a lot more substitutions than what is proposed. All of the ISA uses with bind() are okay as is but could also be updated.\r\n\r\nIncorrect cases that cannot be fixed:\r\n * hbase-external-blockcache: We have to resolve all of the memcached locations up front because the memcached client constructor requires ISA instances. So we have to hope that the network identities (DNS name -> IP address mapping) does not change for any in the list. This is beyond our control.\r\n\r\nWhile in this area it is trivial to add new client connect metrics for number of potential nameservice lookups (whenever we instantiate an ISA) and number of failed nameservice lookups (if the instantiated ISA is unresolved).\r\n\r\nWhile in this area I also noticed we often directly access a field in ConnectionId where there is also a getter, so good practice is to use the getter instead.","issue_id":"13340890","key":"HBASE-25292","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-12-04T18:26:12.000+0000","role":"fixed_distractor","summary":"Improve InetSocketAddress usage discipline"} {"case_id":"13341561","cluster":"DISTRACTOR-HBASE-25307","comments":[{"body":"Checked the code again, all the places where we access the connections PoolMap is under synchronzed(connections), so I do not think we need to make the PoolMap itself thread safe.\r\n\r\nAnd since it is only used in AbstractRpcClient, I suggest we just rename it to RpcConnectionMap, and move it to the ipc package and change it to package private.\r\n\r\nThoughts? Thanks.","created":"2020-11-26T06:59:09.855+0000"},{"body":"{quote}Checked the code again, all the places where we access the connections PoolMap is under synchronzed(connections), so I do not think we need to make the PoolMap itself thread safe.\r\n{quote}\r\n\r\nIt is true, but if you check the PR, our goal was to get rid of the synchronization in {{AbstractRpcClient}}. It would require some extra work, so we didn't do that, but I think we should keep {{PoolMap}} synchronized and remove synchronization from {{AbstractRpcClient}}.\r\n\r\nbq. And since it is only used in AbstractRpcClient, I suggest we just rename it to RpcConnectionMap, and move it to the ipc package and change it to package private.\r\n\r\nIt is possible. It started as a bugfix, I didn't want to do such a huge change. We also have to backport this fix to branch-2* branches, because this bug also exists there.","created":"2020-11-26T09:38:07.208+0000"},{"body":"{quote}\r\nIt is true, but if you check the PR, our goal was to get rid of the synchronization in AbstractRpcClient. It would require some extra work, so we didn't do that, but I think we should keep PoolMap synchronized and remove synchronization from AbstractRpcClient.\r\n{quote}\r\n\r\nThis is not easy, it relates to the logic for the cleanup of the connections, where we need to set and check the last touch time along with the modification to the PoolMap. If we want to do this then we need to design a lot of methods in PoolMap so we can put all the code under the lock inside PoolMap, like what you have done for getOrCreate. For me I do not think it is worth to do the abstraction if we make PoolMap package private and only use it in AbstractRpcClient.\r\n\r\nThanks.","created":"2020-11-27T06:12:58.453+0000"},{"body":"bq. This is not easy, it relates to the logic for the cleanup of the connections, where we need to set and check the last touch time along with the modification to the PoolMap.\r\n\r\nYes, we would need a functionality like {{close()}} and {{drainTo()}} which ensure no new connection can be created and every existing connection is closed properly.\r\n\r\nI am not against further refactoring, the previous implementation wasn't very good.","created":"2020-11-30T12:09:29.811+0000"}],"conversations":[{"body":"We got NPE after setting {{hbase.client.ipc.pool.type}} to {{thread-local}}:\r\n{noformat}\r\n20/11/18 01:53:04 ERROR yarn.ApplicationMaster: User class threw exception: java.lang.NullPointerException\r\njava.lang.NullPointerException\r\n at org.apache.hadoop.hbase.ipc.AbstractRpcClient.close(AbstractRpcClient.java:496)\r\n at org.apache.hadoop.hbase.client.ConnectionImplementation.close(ConnectionImplementation.java:1944)\r\n at org.apache.hadoop.hbase.mapreduce.TableInputFormatBase.close(TableInputFormatBase.java:660)\r\n{noformat}\r\nThe root cause of the issue is probably at {{PoolMap.ThreadLocalPool.values()}}:\r\n{code:java}\r\npublic Collection values() {\r\n List values = new ArrayList<>();\r\n values.add(get());\r\n return values;\r\n}\r\n{code}\r\nIt adds {{null}} into the collection if the current thread does not have any resources which leads to NPE later.\r\n\r\nI traced the usages of values() and it should return every resource, not just that one which is attached to the caller thread.","from":"reporter","subject":"ThreadLocal pooling leads to NullPointerException"},{"body":"Checked the code again, all the places where we access the connections PoolMap is under synchronzed(connections), so I do not think we need to make the PoolMap itself thread safe.\r\n\r\nAnd since it is only used in AbstractRpcClient, I suggest we just rename it to RpcConnectionMap, and move it to the ipc package and change it to package private.\r\n\r\nThoughts? Thanks.","from":"developer"},{"body":"{quote}Checked the code again, all the places where we access the connections PoolMap is under synchronzed(connections), so I do not think we need to make the PoolMap itself thread safe.\r\n{quote}\r\n\r\nIt is true, but if you check the PR, our goal was to get rid of the synchronization in {{AbstractRpcClient}}. It would require some extra work, so we didn't do that, but I think we should keep {{PoolMap}} synchronized and remove synchronization from {{AbstractRpcClient}}.\r\n\r\nbq. And since it is only used in AbstractRpcClient, I suggest we just rename it to RpcConnectionMap, and move it to the ipc package and change it to package private.\r\n\r\nIt is possible. It started as a bugfix, I didn't want to do such a huge change. We also have to backport this fix to branch-2* branches, because this bug also exists there.","from":"developer"},{"body":"{quote}\r\nIt is true, but if you check the PR, our goal was to get rid of the synchronization in AbstractRpcClient. It would require some extra work, so we didn't do that, but I think we should keep PoolMap synchronized and remove synchronization from AbstractRpcClient.\r\n{quote}\r\n\r\nThis is not easy, it relates to the logic for the cleanup of the connections, where we need to set and check the last touch time along with the modification to the PoolMap. If we want to do this then we need to design a lot of methods in PoolMap so we can put all the code under the lock inside PoolMap, like what you have done for getOrCreate. For me I do not think it is worth to do the abstraction if we make PoolMap package private and only use it in AbstractRpcClient.\r\n\r\nThanks.","from":"developer"},{"body":"bq. This is not easy, it relates to the logic for the cleanup of the connections, where we need to set and check the last touch time along with the modification to the PoolMap.\r\n\r\nYes, we would need a functionality like {{close()}} and {{drainTo()}} which ensure no new connection can be created and every existing connection is closed properly.\r\n\r\nI am not against further refactoring, the previous implementation wasn't very good.","from":"developer"}],"created":"2020-11-19T11:14:57.000+0000","description":"We got NPE after setting {{hbase.client.ipc.pool.type}} to {{thread-local}}:\r\n{noformat}\r\n20/11/18 01:53:04 ERROR yarn.ApplicationMaster: User class threw exception: java.lang.NullPointerException\r\njava.lang.NullPointerException\r\n at org.apache.hadoop.hbase.ipc.AbstractRpcClient.close(AbstractRpcClient.java:496)\r\n at org.apache.hadoop.hbase.client.ConnectionImplementation.close(ConnectionImplementation.java:1944)\r\n at org.apache.hadoop.hbase.mapreduce.TableInputFormatBase.close(TableInputFormatBase.java:660)\r\n{noformat}\r\nThe root cause of the issue is probably at {{PoolMap.ThreadLocalPool.values()}}:\r\n{code:java}\r\npublic Collection values() {\r\n List values = new ArrayList<>();\r\n values.add(get());\r\n return values;\r\n}\r\n{code}\r\nIt adds {{null}} into the collection if the current thread does not have any resources which leads to NPE later.\r\n\r\nI traced the usages of values() and it should return every resource, not just that one which is attached to the caller thread.","issue_id":"13341561","key":"HBASE-25307","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-11-30T12:04:46.000+0000","role":"fixed_distractor","summary":"ThreadLocal pooling leads to NullPointerException"} {"case_id":"13344307","cluster":"DISTRACTOR-HBASE-25361","comments":[{"body":"Pushed to branch-2.3+. Thanks for review [~huaxiang] . Hopefully this fixes the flakey.","created":"2020-12-05T16:59:22.952+0000"},{"body":"Reopen to apply addendum. Test just failed on me. Looking at my PR again, I'm missing reset of a counter so the new wait actually doesn't happen. Here is the addendum:\r\n\r\n \r\n\r\ncommit 7d0a687e5798a2f4ca3190b409169f7e17a75b34 (HEAD -> m, origin/master, origin/HEAD)\r\nAuthor: stack \r\nDate: Sat Dec 5 14:00:18 2020 -0800\r\n\r\nHBASE-25361 [Flakey Tests] branch-2 TestMetaRegionLocationCache.testStandByMetaLocations (#2736)\r\n Addendum; Reset counter so we actually wait in the new loop added by the\r\n above.\r\n\r\n{{diff --git a/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestMetaRegionLocationCache.java b/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestMetaRegionLocationCache.java}}\r\n{{index 577e15cedf..2bcddc9ea7 100644}}\r\n{{--- a/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestMetaRegionLocationCache.java}}\r\n{{+++ b/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestMetaRegionLocationCache.java}}\r\n{{@@ -99,6 +99,7 @@ public class TestMetaRegionLocationCache {}}\r\n{{ ZKWatcher zk = master.getZooKeeper();}}\r\n{{ List metaZnodes = zk.getMetaReplicaNodes();}}\r\n{{ // Wait till all replicas available.}}\r\n{{+ retries = 0;}}\r\n{{ while (master.getMetaRegionLocationCache().getMetaRegionLocations().get().size() !=}}\r\n{{ metaZnodes.size()) {}}\r\n{{ Thread.sleep(1000);}}","created":"2020-12-05T22:10:08.928+0000"},{"body":"Re-resovling. Pushed the addendum to 2.3+ (doesn't go to 2.2)","created":"2020-12-05T22:10:33.737+0000"}],"conversations":[{"body":"This just happened to me running tests locally. I see it in the current flakies list [https://ci-hadoop.apache.org/job/HBase/job/HBase-Flaky-Tests/job/branch-2/310/testReport/junit/org.apache.hadoop.hbase.client/TestMetaRegionLocationCache/testStandByMetaLocations/]\r\n\r\n \r\n\r\nIt looks like a simple timing issue. Let me put up a simple patch to see if it helps.","from":"reporter","subject":"[Flakey Tests] branch-2 TestMetaRegionLocationCache.testStandByMetaLocations"},{"body":"Pushed to branch-2.3+. Thanks for review [~huaxiang] . Hopefully this fixes the flakey.","from":"developer"},{"body":"Reopen to apply addendum. Test just failed on me. Looking at my PR again, I'm missing reset of a counter so the new wait actually doesn't happen. Here is the addendum:\r\n\r\n \r\n\r\ncommit 7d0a687e5798a2f4ca3190b409169f7e17a75b34 (HEAD -> m, origin/master, origin/HEAD)\r\nAuthor: stack \r\nDate: Sat Dec 5 14:00:18 2020 -0800\r\n\r\nHBASE-25361 [Flakey Tests] branch-2 TestMetaRegionLocationCache.testStandByMetaLocations (#2736)\r\n Addendum; Reset counter so we actually wait in the new loop added by the\r\n above.\r\n\r\n{{diff --git a/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestMetaRegionLocationCache.java b/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestMetaRegionLocationCache.java}}\r\n{{index 577e15cedf..2bcddc9ea7 100644}}\r\n{{--- a/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestMetaRegionLocationCache.java}}\r\n{{+++ b/hbase-server/src/test/java/org/apache/hadoop/hbase/client/TestMetaRegionLocationCache.java}}\r\n{{@@ -99,6 +99,7 @@ public class TestMetaRegionLocationCache {}}\r\n{{ ZKWatcher zk = master.getZooKeeper();}}\r\n{{ List metaZnodes = zk.getMetaReplicaNodes();}}\r\n{{ // Wait till all replicas available.}}\r\n{{+ retries = 0;}}\r\n{{ while (master.getMetaRegionLocationCache().getMetaRegionLocations().get().size() !=}}\r\n{{ metaZnodes.size()) {}}\r\n{{ Thread.sleep(1000);}}","from":"developer"},{"body":"Re-resovling. Pushed the addendum to 2.3+ (doesn't go to 2.2)","from":"developer"}],"created":"2020-12-04T23:41:48.000+0000","description":"This just happened to me running tests locally. I see it in the current flakies list [https://ci-hadoop.apache.org/job/HBase/job/HBase-Flaky-Tests/job/branch-2/310/testReport/junit/org.apache.hadoop.hbase.client/TestMetaRegionLocationCache/testStandByMetaLocations/]\r\n\r\n \r\n\r\nIt looks like a simple timing issue. Let me put up a simple patch to see if it helps.","issue_id":"13344307","key":"HBASE-25361","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2020-12-05T22:10:33.000+0000","role":"fixed_distractor","summary":"[Flakey Tests] branch-2 TestMetaRegionLocationCache.testStandByMetaLocations"} {"case_id":"13347244","cluster":"DISTRACTOR-HBASE-25432","comments":[{"body":"Thanks for filing this [~xiaoheipangzi]. I agree that we need security checks for setTableStateInMeta().","created":"2020-12-22T19:58:26.287+0000"},{"body":"We also find that Hbck.fixMeta also lack of security check, non-admin can also fix the meta, below is log!\r\n  \r\n 2020-12-23 06:26:20,947 INFO  [RpcServer.default.FPBQ.Fifo.handler=28,queue=1,port=16000] master.MetaFixer: Fixed hole by adding \\{ENCODED => e70948da53cc8a6ce7f7a270a53b884a, NAME => 'TestTable,00000000000000000000051557,1608704780922.e70948da53cc8a6ce7f7a270a53b884a.', STARTKEY => '00000000000000000000051557', ENDKEY => '00000000000000000000056244'}; region is NOT assigned (assign to online)\r\n\r\n \r\n\r\nit seems that one user can write region into other users' table!","created":"2020-12-23T06:29:09.407+0000"},{"body":"We should ensure rpc implementation uses access controller regardless of client, whether it is hbck or something else using it.\r\n\r\n[~xiaoheipangzi] are you planning to work on this Jira? Just want to confirm.\r\n\r\nThanks","created":"2020-12-23T07:04:35.673+0000"},{"body":"yes, i will try my best to fix it.","created":"2020-12-23T07:19:56.140+0000"},{"body":"Thank you [~xiaoheipangzi]. Really appreciate your efforts (y)","created":"2020-12-23T08:47:14.013+0000"},{"body":"[~vjasani] \r\n\r\nHi, i have submit a patch, when we try fixMeta or setTableStateInMeta as non-admin, it will generate exception like:\r\n{code:java}\r\n org.apache.hadoop.hbase.security.AccessDeniedException: Insufficient permissions for user 'user1' (global, action=ADMIN)\r\n{code}","created":"2020-12-23T09:07:03.975+0000"},{"body":"Great, Thanks [~xiaoheipangzi]. Can you please raise github pull request? HBase community no longer supports patch file based QA run (we will update documentation for the same).","created":"2020-12-23T11:55:21.948+0000"},{"body":"I pushed on branch-2.3+. [~vjasani] noticed that the branch-1 patch is missing some pieces. Leaving issue open till branch-1 PR makes it in.","created":"2020-12-28T19:02:18.674+0000"},{"body":"Thanks for the contribution [~xiaoheipangzi].\r\n\r\nThanks for the reviews and merge [~zhangduo] [~elserj] [~stack]","created":"2020-12-29T12:07:40.626+0000"},{"body":"Reopening for branch-2.2 backport.","created":"2021-01-07T13:38:16.738+0000"},{"body":"Backported to branch-2.2.","created":"2021-01-07T13:39:16.403+0000"}],"conversations":[{"body":"setTableStateInMeta and fixMeta can be accessed only through Admin rights","from":"reporter","subject":"we should add security checks for setTableStateInMeta and fixMeta"},{"body":"Thanks for filing this [~xiaoheipangzi]. I agree that we need security checks for setTableStateInMeta().","from":"developer"},{"body":"We also find that Hbck.fixMeta also lack of security check, non-admin can also fix the meta, below is log!\r\n  \r\n 2020-12-23 06:26:20,947 INFO  [RpcServer.default.FPBQ.Fifo.handler=28,queue=1,port=16000] master.MetaFixer: Fixed hole by adding \\{ENCODED => e70948da53cc8a6ce7f7a270a53b884a, NAME => 'TestTable,00000000000000000000051557,1608704780922.e70948da53cc8a6ce7f7a270a53b884a.', STARTKEY => '00000000000000000000051557', ENDKEY => '00000000000000000000056244'}; region is NOT assigned (assign to online)\r\n\r\n \r\n\r\nit seems that one user can write region into other users' table!","from":"developer"},{"body":"We should ensure rpc implementation uses access controller regardless of client, whether it is hbck or something else using it.\r\n\r\n[~xiaoheipangzi] are you planning to work on this Jira? Just want to confirm.\r\n\r\nThanks","from":"developer"},{"body":"yes, i will try my best to fix it.","from":"developer"},{"body":"Thank you [~xiaoheipangzi]. Really appreciate your efforts (y)","from":"developer"},{"body":"[~vjasani] \r\n\r\nHi, i have submit a patch, when we try fixMeta or setTableStateInMeta as non-admin, it will generate exception like:\r\n{code:java}\r\n org.apache.hadoop.hbase.security.AccessDeniedException: Insufficient permissions for user 'user1' (global, action=ADMIN)\r\n{code}","from":"developer"},{"body":"Great, Thanks [~xiaoheipangzi]. Can you please raise github pull request? HBase community no longer supports patch file based QA run (we will update documentation for the same).","from":"developer"},{"body":"I pushed on branch-2.3+. [~vjasani] noticed that the branch-1 patch is missing some pieces. Leaving issue open till branch-1 PR makes it in.","from":"developer"},{"body":"Thanks for the contribution [~xiaoheipangzi].\r\n\r\nThanks for the reviews and merge [~zhangduo] [~elserj] [~stack]","from":"developer"},{"body":"Reopening for branch-2.2 backport.","from":"developer"},{"body":"Backported to branch-2.2.","from":"developer"}],"created":"2020-12-22T03:59:04.000+0000","description":"setTableStateInMeta and fixMeta can be accessed only through Admin rights","issue_id":"13347244","key":"HBASE-25432","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-01-07T13:39:16.000+0000","role":"fixed_distractor","summary":"we should add security checks for setTableStateInMeta and fixMeta"} {"case_id":"13347835","cluster":"DISTRACTOR-HBASE-25445","comments":[{"body":"[~vjasani] Can you please assign this to me? Thanks. ","created":"2021-01-03T06:10:19.367+0000"},{"body":"Done [~dasanjan1296]. You have Jira contributor access now, you can assign Jiras yourself going forward. Thanks for picking this one.","created":"2021-01-03T06:40:59.959+0000"},{"body":"[~dasanjan1296] Have you started the fix? I'm testing the fix on my cluster and will raise PR soon.","created":"2021-01-04T03:22:25.093+0000"},{"body":"[~mokai87] Yes, I have started working on the fix and I'm testing it out. ","created":"2021-01-04T04:41:40.983+0000"},{"body":"[~vjasani] [~mokai87] I have raised a PR with the fix. Please take a look. ","created":"2021-01-05T05:13:25.678+0000"},{"body":"[~mokai87] To avoid misunderstanding, we should assign Jira to ourself when we start working on it even if it is going to take more time coming up with PR. The longer we keep Jira unassigned, more contributors will assume that no one is working on Jira as of now and hence they can assign it to themselves. Once you assign Jira to yourself, no one will try to reassign Jira to them without your formal permission on Jira :)\r\n\r\nThanks for filing this nice Jira!","created":"2021-01-05T06:07:33.954+0000"},{"body":"Got it. Thanks for the PR.","created":"2021-01-05T06:24:09.072+0000"},{"body":"This issues arises when we have WAL split NOT managed by zk ? In Zk managed this works? Trying to understand. Sorry did not see patch","created":"2021-01-06T05:11:59.475+0000"},{"body":"Yeah Anoop, this problem occured when we enabled Proc based WAL split, ZK based WAL split is working fine.","created":"2021-01-06T05:14:09.844+0000"},{"body":"It is this way from the time Proc based WAL split is introduced? Can you pls update the Affected versions accordingly? ","created":"2021-01-06T09:37:41.932+0000"},{"body":"This bug is there since HBASE-21588 implemented, also we identified a TODO for this issue during HBASE-24632 but missed later.","created":"2021-01-06T18:06:54.787+0000"},{"body":"FYI [~ndimiduk] [~huaxiangsun] Would you like to wait for some time to get this merged before spinning 2.3.4 RC? We are just waiting for a unit test and it's good to go.\r\n\r\nI just included 2.3.4 in fix versions, please update it as per your decision.","created":"2021-01-06T19:19:11.989+0000"},{"body":"I checked the diff, it is straightforward. Can we create a new Jira as a followup to add a unitest and merge this diff? \r\n\r\nThis will unblock both sides, thanks.","created":"2021-01-07T00:46:43.926+0000"},{"body":"I created HBASE-25470 as a followup for unitest case. Please merge and backport this Jira to all affected branches, thanks.","created":"2021-01-07T00:52:55.535+0000"},{"body":"We are going to spin the RC tomorrow, hopefully, it will be merged to 2.3 by tomorrow, thanks.","created":"2021-01-07T00:53:51.128+0000"},{"body":"Thanks [~huaxiangsun]. Since all reviews were happening on PR #2844, [~dasanjan1296] has checked in test already. Hopefully this should be merged very soon (at least before start of day in PST today).","created":"2021-01-07T06:43:33.513+0000"},{"body":"Thanks for the contribution [~dasanjan1296] and thanks for filing the bug [~mokai87].\r\n\r\nThanks [~huaxiangsun] for waiting one more day to get this fix in. This is unblocked now for 2.3.4 RC preparation.","created":"2021-01-07T10:34:17.871+0000"},{"body":"Thanks [~vjasani], [~dasanjan1296], and [~mokai87]!","created":"2021-01-07T17:52:37.592+0000"}],"conversations":[{"body":"If 'hbase.wal.dir' and 'hbase.rootdir' are configured to different filesystem, SplitWALRemoteProcedure archived split WAL failed since SplitWALManager using wrong fs instance. SplitWALManager should use WAL corresponding fs instance.\r\n\r\nSteps to Reproduce:\r\n * Configure 'hbase.wal.dir' and 'hbase.rootdir' so that they point to different fs instances.\r\n * Start HBase with multiple RS. \r\n * Create a couple of tables and some rows in them so that the RSs get assigned with some regions. \r\n * Take any RS with non-zero number of regions offline. \r\n * Check master logs for \"Wrong FS\" error as shown in the screenshot attached. \r\n\r\n ","from":"reporter","subject":"Old WALs archive fails in procedure based WAL split"},{"body":"[~vjasani] Can you please assign this to me? Thanks. ","from":"developer"},{"body":"Done [~dasanjan1296]. You have Jira contributor access now, you can assign Jiras yourself going forward. Thanks for picking this one.","from":"developer"},{"body":"[~dasanjan1296] Have you started the fix? I'm testing the fix on my cluster and will raise PR soon.","from":"developer"},{"body":"[~mokai87] Yes, I have started working on the fix and I'm testing it out. ","from":"developer"},{"body":"[~vjasani] [~mokai87] I have raised a PR with the fix. Please take a look. ","from":"developer"},{"body":"[~mokai87] To avoid misunderstanding, we should assign Jira to ourself when we start working on it even if it is going to take more time coming up with PR. The longer we keep Jira unassigned, more contributors will assume that no one is working on Jira as of now and hence they can assign it to themselves. Once you assign Jira to yourself, no one will try to reassign Jira to them without your formal permission on Jira :)\r\n\r\nThanks for filing this nice Jira!","from":"developer"},{"body":"Got it. Thanks for the PR.","from":"developer"},{"body":"This issues arises when we have WAL split NOT managed by zk ? In Zk managed this works? Trying to understand. Sorry did not see patch","from":"developer"},{"body":"Yeah Anoop, this problem occured when we enabled Proc based WAL split, ZK based WAL split is working fine.","from":"developer"},{"body":"It is this way from the time Proc based WAL split is introduced? Can you pls update the Affected versions accordingly? ","from":"developer"},{"body":"This bug is there since HBASE-21588 implemented, also we identified a TODO for this issue during HBASE-24632 but missed later.","from":"developer"},{"body":"FYI [~ndimiduk] [~huaxiangsun] Would you like to wait for some time to get this merged before spinning 2.3.4 RC? We are just waiting for a unit test and it's good to go.\r\n\r\nI just included 2.3.4 in fix versions, please update it as per your decision.","from":"developer"},{"body":"I checked the diff, it is straightforward. Can we create a new Jira as a followup to add a unitest and merge this diff? \r\n\r\nThis will unblock both sides, thanks.","from":"developer"},{"body":"I created HBASE-25470 as a followup for unitest case. Please merge and backport this Jira to all affected branches, thanks.","from":"developer"},{"body":"We are going to spin the RC tomorrow, hopefully, it will be merged to 2.3 by tomorrow, thanks.","from":"developer"},{"body":"Thanks [~huaxiangsun]. Since all reviews were happening on PR #2844, [~dasanjan1296] has checked in test already. Hopefully this should be merged very soon (at least before start of day in PST today).","from":"developer"},{"body":"Thanks for the contribution [~dasanjan1296] and thanks for filing the bug [~mokai87].\r\n\r\nThanks [~huaxiangsun] for waiting one more day to get this fix in. This is unblocked now for 2.3.4 RC preparation.","from":"developer"},{"body":"Thanks [~vjasani], [~dasanjan1296], and [~mokai87]!","from":"developer"}],"created":"2020-12-25T07:06:57.000+0000","description":"If 'hbase.wal.dir' and 'hbase.rootdir' are configured to different filesystem, SplitWALRemoteProcedure archived split WAL failed since SplitWALManager using wrong fs instance. SplitWALManager should use WAL corresponding fs instance.\r\n\r\nSteps to Reproduce:\r\n * Configure 'hbase.wal.dir' and 'hbase.rootdir' so that they point to different fs instances.\r\n * Start HBase with multiple RS. \r\n * Create a couple of tables and some rows in them so that the RSs get assigned with some regions. \r\n * Take any RS with non-zero number of regions offline. \r\n * Check master logs for \"Wrong FS\" error as shown in the screenshot attached. \r\n\r\n ","issue_id":"13347835","key":"HBASE-25445","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-01-07T10:34:17.000+0000","role":"fixed_distractor","summary":"Old WALs archive fails in procedure based WAL split"} {"case_id":"13350811","cluster":"DISTRACTOR-HBASE-25473","comments":[{"body":"See below for when I run:\r\n\r\n\r\n ./dev-support/checkcompatibility.py --annotation org.apache.yetus.audience.InterfaceAudience.Public rel/2.3.3 2.3.4RC0\r\n\r\n\r\n{code}\r\nINFO:root:Annotations are: ['org.apache.yetus.audience.InterfaceAudience.Public']\r\nINFO:root:Annotations path: /Users/stack/checkouts/hbase.apache.git/target/compat-check/annotations.txt\r\n...\r\njava.io.IOException: META-INF/license : could not create directory\r\n\tat sun.tools.jar.Main.extractFile(Main.java:1045)\r\n\tat sun.tools.jar.Main.extract(Main.java:981)\r\n\tat sun.tools.jar.Main.run(Main.java:311)\r\n\tat sun.tools.jar.Main.main(Main.java:1288)\r\nERROR: can't extract '/Users/stack/checkouts/hbase.apache.git/target/compat-check/src/hbase-shaded/hbase-shaded-testing-util/target/hbase-shaded-testing-util-2.3.3.jar'\r\nINFO:root:Results: {}\r\nTraceback (most recent call last):\r\n File \"./dev-support/checkcompatibility.py\", line 530, in \r\n main()\r\n File \"./dev-support/checkcompatibility.py\", line 526, in main\r\n args.compare_warnings))\r\n File \"./dev-support/checkcompatibility.py\", line 232, in compare_results\r\n if tool_results[check][issue_type] > known_count]\r\nKeyError: 'binary'\r\n{code}","created":"2021-01-07T17:54:06.585+0000"},{"body":"The hbase-shaded-testing-util jar just fails when you undo it (jar -xf *.jar) on mac os x because it bundles a LICENSE file and then has a license dir with dependency licenses in it. This causes checkcompatibility.py fail when run on mac os x (file name case insenstitivity 'feature').\r\n\r\nMy original failure was inside of a vm up on virtualbox running docker (on a mac) so should have been insulated against this fs messing.","created":"2021-01-07T22:16:03.156+0000"},{"body":"Still fails cryptic even w/ debug in it and the shaded testing util excluded:\r\n{code}\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst/hbase-protocol/target/hbase-protocol-2.3.4.jar\r\nINFO:root:Annotations are: ['org.apache.yetus.audience.InterfaceAudience.Public']\r\nINFO:root:Annotations path: /home/vagrant/hbase-rm/output/hbase/target/compat-check/annotations.txt\r\nINFO:root:Results: {}\r\nTraceback (most recent call last):\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 530, in \r\n main()\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 526, in main\r\n args.compare_warnings))\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 232, in compare_results\r\n if tool_results[check][issue_type] > known_count]\r\nKeyError: 'binary'\r\ncp: cannot stat './hbase/target/compat-check/report.html': No such file or directory\r\n2021-01-07T23:24:37Z generate_api_report stop (2258 seconds)\r\ncp: cannot stat 'api*.html': No such file or directory\r\n{code}","created":"2021-01-07T23:36:50.902+0000"},{"body":"Turns out my filter for excluding hbase-shaded-testing-util was incorrect. Fixing the filter made checkcompatibility pass again inside docker on mac. I checked on a gcp vm yesterday and it didn't have issues. Excluding hbase-shaded-testing-util should be fine for now given it just an aggregate of other jars. Let me put up a patch to fix this failing on mac and then make a subtask for reorg'ing hbase-shaded-testing-util assembly.","created":"2021-01-08T17:28:52.104+0000"},{"body":"FYI [~busbey]... Weird stuff like this is your kinda thing.","created":"2021-01-08T17:38:34.807+0000"},{"body":"Merged. Thanks for review [~ndimiduk]","created":"2021-01-08T22:48:29.955+0000"}],"conversations":[{"body":"Below is excerpt from log of RC0 2.3.4 build this afternoon when we try to do the compatibility check (inside docker create-release build). It had failed for me previously on another run of different versions. The failure is unfortunately cryptic.\r\n\r\n{code}\r\n2021-01-07T01:08:34Z generate_api_report start\r\nINFO:root:Source revision: rel/2.3.3\r\nINFO:root:Destination revision: 2.3.4RC0\r\nINFO:root:Filtering classes using 1 annotation(s):\r\nINFO:root: org.apache.yetus.audience.InterfaceAudience.Public\r\nINFO:root:Downloading Java ACC...\r\nINFO:root:Java ACC version: 2.4\r\n\r\nINFO:root:Removing scratch dir /home/vagrant/hbase-rm/output/hbase/target/compat-check\r\nINFO:root:Creating empty scratch dir /home/vagrant/hbase-rm/output/hbase/target/compat-check\r\nINFO:root:Checking out 5f7fe66efbdfc1c0b3262c0677f17e4055ba94cd in /home/vagrant/hbase-rm/output/hbase/target/compat-check/src\r\nINFO:root:Checking out ec79612b9ac6afd73fa8fc7f7f5c8d3bac277679 in /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst\r\nINFO:root:Building in /home/vagrant/hbase-rm/output/hbase/target/compat-check/src\r\nINFO:root:Building in /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst\r\nINFO:root:Will check compatibility between original jars:\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/src/hbase-shaded/hbase-shaded-client-byo-hadoop/target/original-hbase-shaded-client-byo-hadoop-2.3.3.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/src/hbase-external-blockcache/target/hbase-external-blockcache-2.3.3.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/src/hbase-shaded/hbase-shaded-client/target/original-hbase-shaded-client-2.3.3.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/src/hbase-zookeeper/target/hbase-zookeeper-2.3.3.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/src/hbase-common/target/hbase-common-2.3.3.jar\r\n\r\n....\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst/hbase-shaded/hbase-shaded-testing-util-tester/target/hbase-shaded-testing-util-tester-2.3.4.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst/hbase-archetypes/hbase-shaded-client-project/target/hbase-shaded-client-project-2.3.4.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst/hbase-protocol/target/hbase-protocol-2.3.4.jar\r\nINFO:root:Annotations are: ['org.apache.yetus.audience.InterfaceAudience.Public']\r\nINFO:root:Annotations path: /home/vagrant/hbase-rm/output/hbase/target/compat-check/annotations.txt\r\nINFO:root:Results: {}\r\nTraceback (most recent call last):\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 530, in \r\n main()\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 526, in main\r\n args.compare_warnings))\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 232, in compare_results\r\n if tool_results[check][issue_type] > known_count]\r\n{code}","from":"reporter","subject":"[create-release] checkcompatibility.py failing with \"KeyError: 'binary'\""},{"body":"See below for when I run:\r\n\r\n\r\n ./dev-support/checkcompatibility.py --annotation org.apache.yetus.audience.InterfaceAudience.Public rel/2.3.3 2.3.4RC0\r\n\r\n\r\n{code}\r\nINFO:root:Annotations are: ['org.apache.yetus.audience.InterfaceAudience.Public']\r\nINFO:root:Annotations path: /Users/stack/checkouts/hbase.apache.git/target/compat-check/annotations.txt\r\n...\r\njava.io.IOException: META-INF/license : could not create directory\r\n\tat sun.tools.jar.Main.extractFile(Main.java:1045)\r\n\tat sun.tools.jar.Main.extract(Main.java:981)\r\n\tat sun.tools.jar.Main.run(Main.java:311)\r\n\tat sun.tools.jar.Main.main(Main.java:1288)\r\nERROR: can't extract '/Users/stack/checkouts/hbase.apache.git/target/compat-check/src/hbase-shaded/hbase-shaded-testing-util/target/hbase-shaded-testing-util-2.3.3.jar'\r\nINFO:root:Results: {}\r\nTraceback (most recent call last):\r\n File \"./dev-support/checkcompatibility.py\", line 530, in \r\n main()\r\n File \"./dev-support/checkcompatibility.py\", line 526, in main\r\n args.compare_warnings))\r\n File \"./dev-support/checkcompatibility.py\", line 232, in compare_results\r\n if tool_results[check][issue_type] > known_count]\r\nKeyError: 'binary'\r\n{code}","from":"developer"},{"body":"The hbase-shaded-testing-util jar just fails when you undo it (jar -xf *.jar) on mac os x because it bundles a LICENSE file and then has a license dir with dependency licenses in it. This causes checkcompatibility.py fail when run on mac os x (file name case insenstitivity 'feature').\r\n\r\nMy original failure was inside of a vm up on virtualbox running docker (on a mac) so should have been insulated against this fs messing.","from":"developer"},{"body":"Still fails cryptic even w/ debug in it and the shaded testing util excluded:\r\n{code}\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst/hbase-protocol/target/hbase-protocol-2.3.4.jar\r\nINFO:root:Annotations are: ['org.apache.yetus.audience.InterfaceAudience.Public']\r\nINFO:root:Annotations path: /home/vagrant/hbase-rm/output/hbase/target/compat-check/annotations.txt\r\nINFO:root:Results: {}\r\nTraceback (most recent call last):\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 530, in \r\n main()\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 526, in main\r\n args.compare_warnings))\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 232, in compare_results\r\n if tool_results[check][issue_type] > known_count]\r\nKeyError: 'binary'\r\ncp: cannot stat './hbase/target/compat-check/report.html': No such file or directory\r\n2021-01-07T23:24:37Z generate_api_report stop (2258 seconds)\r\ncp: cannot stat 'api*.html': No such file or directory\r\n{code}","from":"developer"},{"body":"Turns out my filter for excluding hbase-shaded-testing-util was incorrect. Fixing the filter made checkcompatibility pass again inside docker on mac. I checked on a gcp vm yesterday and it didn't have issues. Excluding hbase-shaded-testing-util should be fine for now given it just an aggregate of other jars. Let me put up a patch to fix this failing on mac and then make a subtask for reorg'ing hbase-shaded-testing-util assembly.","from":"developer"},{"body":"FYI [~busbey]... Weird stuff like this is your kinda thing.","from":"developer"},{"body":"Merged. Thanks for review [~ndimiduk]","from":"developer"}],"created":"2021-01-07T06:46:59.000+0000","description":"Below is excerpt from log of RC0 2.3.4 build this afternoon when we try to do the compatibility check (inside docker create-release build). It had failed for me previously on another run of different versions. The failure is unfortunately cryptic.\r\n\r\n{code}\r\n2021-01-07T01:08:34Z generate_api_report start\r\nINFO:root:Source revision: rel/2.3.3\r\nINFO:root:Destination revision: 2.3.4RC0\r\nINFO:root:Filtering classes using 1 annotation(s):\r\nINFO:root: org.apache.yetus.audience.InterfaceAudience.Public\r\nINFO:root:Downloading Java ACC...\r\nINFO:root:Java ACC version: 2.4\r\n\r\nINFO:root:Removing scratch dir /home/vagrant/hbase-rm/output/hbase/target/compat-check\r\nINFO:root:Creating empty scratch dir /home/vagrant/hbase-rm/output/hbase/target/compat-check\r\nINFO:root:Checking out 5f7fe66efbdfc1c0b3262c0677f17e4055ba94cd in /home/vagrant/hbase-rm/output/hbase/target/compat-check/src\r\nINFO:root:Checking out ec79612b9ac6afd73fa8fc7f7f5c8d3bac277679 in /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst\r\nINFO:root:Building in /home/vagrant/hbase-rm/output/hbase/target/compat-check/src\r\nINFO:root:Building in /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst\r\nINFO:root:Will check compatibility between original jars:\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/src/hbase-shaded/hbase-shaded-client-byo-hadoop/target/original-hbase-shaded-client-byo-hadoop-2.3.3.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/src/hbase-external-blockcache/target/hbase-external-blockcache-2.3.3.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/src/hbase-shaded/hbase-shaded-client/target/original-hbase-shaded-client-2.3.3.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/src/hbase-zookeeper/target/hbase-zookeeper-2.3.3.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/src/hbase-common/target/hbase-common-2.3.3.jar\r\n\r\n....\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst/hbase-shaded/hbase-shaded-testing-util-tester/target/hbase-shaded-testing-util-tester-2.3.4.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst/hbase-archetypes/hbase-shaded-client-project/target/hbase-shaded-client-project-2.3.4.jar\r\n /home/vagrant/hbase-rm/output/hbase/target/compat-check/dst/hbase-protocol/target/hbase-protocol-2.3.4.jar\r\nINFO:root:Annotations are: ['org.apache.yetus.audience.InterfaceAudience.Public']\r\nINFO:root:Annotations path: /home/vagrant/hbase-rm/output/hbase/target/compat-check/annotations.txt\r\nINFO:root:Results: {}\r\nTraceback (most recent call last):\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 530, in \r\n main()\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 526, in main\r\n args.compare_warnings))\r\n File \"./hbase/dev-support/checkcompatibility.py\", line 232, in compare_results\r\n if tool_results[check][issue_type] > known_count]\r\n{code}","issue_id":"13350811","key":"HBASE-25473","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-01-08T22:48:29.000+0000","role":"fixed_distractor","summary":"[create-release] checkcompatibility.py failing with \"KeyError: 'binary'\""} {"case_id":"13351685","cluster":"DISTRACTOR-HBASE-25498","comments":[{"body":"Merged to master branch (master branch is src for doc). Thanks for the patch [~SWH12]","created":"2021-01-30T22:14:31.410+0000"}],"conversations":[{"body":"When I configured https for the HBase web UI according to the document(set {{hbase.ssl.enabled}} to {{true}} in _hbase-site.xml_), there are the following errors:\r\n{quote}2021-01-29 11:25:19,360 ERROR [main] master.HMasterCommandLine: Master exiting2021-01-29 11:25:19,360 ERROR [main] master.HMasterCommandLine: Master exitingjava.lang.RuntimeException: Failed construction of Master: class org.apache.hadoop.hbase.master.HMaster.  at org.apache.hadoop.hbase.master.HMaster.constructMaster(HMaster.java:2910) at org.apache.hadoop.hbase.master.HMasterCommandLine.startMaster(HMasterCommandLine.java:236) at org.apache.hadoop.hbase.master.HMasterCommandLine.run(HMasterCommandLine.java:140) at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:76) at org.apache.hadoop.hbase.util.ServerCommandLine.doMain(ServerCommandLine.java:149) at org.apache.hadoop.hbase.master.HMaster.main(HMaster.java:2921)Caused by: java.io.IOException: Problem starting http server at org.apache.hadoop.hbase.http.HttpServer.start(HttpServer.java:1006) at org.apache.hadoop.hbase.http.InfoServer.start(InfoServer.java:100) at org.apache.hadoop.hbase.regionserver.HRegionServer.putUpWebUI(HRegionServer.java:2015) at org.apache.hadoop.hbase.regionserver.HRegionServer.(HRegionServer.java:627) at org.apache.hadoop.hbase.master.HMaster.(HMaster.java:472) at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method) at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62) at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45) at java.lang.reflect.Constructor.newInstance(Constructor.java:423) at org.apache.hadoop.hbase.master.HMaster.constructMaster(HMaster.java:2903) ... 5 moreCaused by: java.lang.IllegalStateException: no valid keystore\r\n{quote}\r\nAccording to the following link (https://community.cloudera.com/t5/Community-Articles/Enable-HTTPS-SSL-for-HBASE-Master-UI/ta-p/247031), I found that SSL certificate and SSL configuration are required. In the configuration document of hbase, we do not need to write a detailed configuration process, but we need to be reminded: only modifying the parameters according to the configuration steps will not take effect.\r\n\r\nI think we need to add hints to the document(Please prepare SSL certificate and ssl configuration file in advance):\r\n{quote}\r\nh3. 59.1. Using Secure HTTP (HTTPS) for the Web UI\r\n\r\nA default HBase install uses insecure HTTP connections for Web UIs for the master and region servers. To enable secure HTTP (HTTPS) connections instead, set {{hbase.ssl.enabled}} to {{true}} in _hbase-site.xml(Please prepare SSL certificate and ssl configuration file in advance)_.\r\n{quote}","from":"reporter","subject":"[Documentation] Add a comment when using Secure HTTP (HTTPS) for the Web UI"},{"body":"Merged to master branch (master branch is src for doc). Thanks for the patch [~SWH12]","from":"developer"}],"created":"2021-01-12T08:39:25.000+0000","description":"When I configured https for the HBase web UI according to the document(set {{hbase.ssl.enabled}} to {{true}} in _hbase-site.xml_), there are the following errors:\r\n{quote}2021-01-29 11:25:19,360 ERROR [main] master.HMasterCommandLine: Master exiting2021-01-29 11:25:19,360 ERROR [main] master.HMasterCommandLine: Master exitingjava.lang.RuntimeException: Failed construction of Master: class org.apache.hadoop.hbase.master.HMaster.  at org.apache.hadoop.hbase.master.HMaster.constructMaster(HMaster.java:2910) at org.apache.hadoop.hbase.master.HMasterCommandLine.startMaster(HMasterCommandLine.java:236) at org.apache.hadoop.hbase.master.HMasterCommandLine.run(HMasterCommandLine.java:140) at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:76) at org.apache.hadoop.hbase.util.ServerCommandLine.doMain(ServerCommandLine.java:149) at org.apache.hadoop.hbase.master.HMaster.main(HMaster.java:2921)Caused by: java.io.IOException: Problem starting http server at org.apache.hadoop.hbase.http.HttpServer.start(HttpServer.java:1006) at org.apache.hadoop.hbase.http.InfoServer.start(InfoServer.java:100) at org.apache.hadoop.hbase.regionserver.HRegionServer.putUpWebUI(HRegionServer.java:2015) at org.apache.hadoop.hbase.regionserver.HRegionServer.(HRegionServer.java:627) at org.apache.hadoop.hbase.master.HMaster.(HMaster.java:472) at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method) at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62) at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45) at java.lang.reflect.Constructor.newInstance(Constructor.java:423) at org.apache.hadoop.hbase.master.HMaster.constructMaster(HMaster.java:2903) ... 5 moreCaused by: java.lang.IllegalStateException: no valid keystore\r\n{quote}\r\nAccording to the following link (https://community.cloudera.com/t5/Community-Articles/Enable-HTTPS-SSL-for-HBASE-Master-UI/ta-p/247031), I found that SSL certificate and SSL configuration are required. In the configuration document of hbase, we do not need to write a detailed configuration process, but we need to be reminded: only modifying the parameters according to the configuration steps will not take effect.\r\n\r\nI think we need to add hints to the document(Please prepare SSL certificate and ssl configuration file in advance):\r\n{quote}\r\nh3. 59.1. Using Secure HTTP (HTTPS) for the Web UI\r\n\r\nA default HBase install uses insecure HTTP connections for Web UIs for the master and region servers. To enable secure HTTP (HTTPS) connections instead, set {{hbase.ssl.enabled}} to {{true}} in _hbase-site.xml(Please prepare SSL certificate and ssl configuration file in advance)_.\r\n{quote}","issue_id":"13351685","key":"HBASE-25498","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-01-30T22:14:31.000+0000","role":"fixed_distractor","summary":"[Documentation] Add a comment when using Secure HTTP (HTTPS) for the Web UI"} {"case_id":"13351932","cluster":"DISTRACTOR-HBASE-25501","comments":[{"body":"Thanks for the contribution [~rda3mon].","created":"2021-01-26T06:40:50.532+0000"}],"conversations":[{"body":" \r\n{code:java}\r\n// code placeholder\r\n$ sudo /etc/init.d/yak master hbase backup create \r\nPlease make sure that backup is enabled on the cluster. To enable backup, in hbase-site.xml, set:\r\n hbase.backup.enable=true\r\nhbase.master.logcleaner.plugins=YOUR_PLUGINS,org.apache.hadoop.hbase.backup.master.BackupLogCleaner\r\nhbase.procedure.master.classes=YOUR_CLASSES,org.apache.hadoop.hbase.backup.master.LogRollMasterProcedureManager\r\nhbase.procedure.regionserver.classes=YOUR_CLASSES,org.apache.hadoop.hbase.backup.regionserver.LogRollRegionServerProcedureManager\r\nhbase.coprocessor.region.classes=YOUR_CLASSES,org.apache.hadoop.hbase.backup.BackupObserver\r\nand restart the clusterUsage: hbase backup create [options]\r\n type \"full\" to create a full backup image\r\n \"incremental\" to create an incremental backup image\r\n backup_path Full path to store the backup imageOptions:\r\n -b Bandwidth per task (MapReduce task) in MB/s\r\n -d Enable debug loggings\r\n -q Yarn queue name to run backup create command on\r\n -s Backup set to backup, mutually exclusive with -t (table list)\r\n -t Table name list, comma-separated.\r\n -w Number of parallel MapReduce tasks to execute {code}\r\nParameters -b, -q, -w are not being used when creating a export snapshot request \r\n{code:java}\r\n// code placeholder\r\n\r\nfor (TableName table : backupInfo.getTables()) {\r\n // Currently we simply set the sub copy tasks by counting the table snapshot number, we can\r\n // calculate the real files' size for the percentage in the future.\r\n // backupCopier.setSubTaskPercntgInWholeTask(1f / numOfSnapshots);\r\n int res;\r\n String[] args = new String[4];\r\n args[0] = \"-snapshot\";\r\n args[1] = backupInfo.getSnapshotName(table);\r\n args[2] = \"-copy-to\";\r\n args[3] = backupInfo.getTableBackupDir(table);\r\n\r\n String jobname = \"Full-Backup_\" + backupInfo.getBackupId() + \"_\" + table.getNameAsString();\r\n if (LOG.isDebugEnabled()) {\r\n LOG.debug(\"Setting snapshot copy job name to : \" + jobname);\r\n }\r\n conf.set(JOB_NAME_CONF_KEY, jobname);\r\n\r\n LOG.debug(\"Copy snapshot \" + args[1] + \" to \" + args[3]);\r\n res = copyService.copy(backupInfo, backupManager, conf, BackupType.FULL, args);\r\n\r\n // if one snapshot export failed, do not continue for remained snapshots\r\n if (res != 0) {\r\n LOG.error(\"Exporting Snapshot \" + args[1] + \" failed with return code: \" + res + \".\");\r\n\r\n throw new IOException(\"Failed of exporting snapshot \" + args[1] + \" to \" + args[3]\r\n + \" with reason code \" + res);\r\n }\r\n\r\n conf.unset(JOB_NAME_CONF_KEY);\r\n LOG.info(\"Snapshot copy \" + args[1] + \" finished.\");\r\n} {code}","from":"reporter","subject":"Backup not using parameters such as bandwidth, workers, etc while exporting snapshot"},{"body":"Thanks for the contribution [~rda3mon].","from":"developer"}],"created":"2021-01-13T07:00:42.000+0000","description":" \r\n{code:java}\r\n// code placeholder\r\n$ sudo /etc/init.d/yak master hbase backup create \r\nPlease make sure that backup is enabled on the cluster. To enable backup, in hbase-site.xml, set:\r\n hbase.backup.enable=true\r\nhbase.master.logcleaner.plugins=YOUR_PLUGINS,org.apache.hadoop.hbase.backup.master.BackupLogCleaner\r\nhbase.procedure.master.classes=YOUR_CLASSES,org.apache.hadoop.hbase.backup.master.LogRollMasterProcedureManager\r\nhbase.procedure.regionserver.classes=YOUR_CLASSES,org.apache.hadoop.hbase.backup.regionserver.LogRollRegionServerProcedureManager\r\nhbase.coprocessor.region.classes=YOUR_CLASSES,org.apache.hadoop.hbase.backup.BackupObserver\r\nand restart the clusterUsage: hbase backup create [options]\r\n type \"full\" to create a full backup image\r\n \"incremental\" to create an incremental backup image\r\n backup_path Full path to store the backup imageOptions:\r\n -b Bandwidth per task (MapReduce task) in MB/s\r\n -d Enable debug loggings\r\n -q Yarn queue name to run backup create command on\r\n -s Backup set to backup, mutually exclusive with -t (table list)\r\n -t Table name list, comma-separated.\r\n -w Number of parallel MapReduce tasks to execute {code}\r\nParameters -b, -q, -w are not being used when creating a export snapshot request \r\n{code:java}\r\n// code placeholder\r\n\r\nfor (TableName table : backupInfo.getTables()) {\r\n // Currently we simply set the sub copy tasks by counting the table snapshot number, we can\r\n // calculate the real files' size for the percentage in the future.\r\n // backupCopier.setSubTaskPercntgInWholeTask(1f / numOfSnapshots);\r\n int res;\r\n String[] args = new String[4];\r\n args[0] = \"-snapshot\";\r\n args[1] = backupInfo.getSnapshotName(table);\r\n args[2] = \"-copy-to\";\r\n args[3] = backupInfo.getTableBackupDir(table);\r\n\r\n String jobname = \"Full-Backup_\" + backupInfo.getBackupId() + \"_\" + table.getNameAsString();\r\n if (LOG.isDebugEnabled()) {\r\n LOG.debug(\"Setting snapshot copy job name to : \" + jobname);\r\n }\r\n conf.set(JOB_NAME_CONF_KEY, jobname);\r\n\r\n LOG.debug(\"Copy snapshot \" + args[1] + \" to \" + args[3]);\r\n res = copyService.copy(backupInfo, backupManager, conf, BackupType.FULL, args);\r\n\r\n // if one snapshot export failed, do not continue for remained snapshots\r\n if (res != 0) {\r\n LOG.error(\"Exporting Snapshot \" + args[1] + \" failed with return code: \" + res + \".\");\r\n\r\n throw new IOException(\"Failed of exporting snapshot \" + args[1] + \" to \" + args[3]\r\n + \" with reason code \" + res);\r\n }\r\n\r\n conf.unset(JOB_NAME_CONF_KEY);\r\n LOG.info(\"Snapshot copy \" + args[1] + \" finished.\");\r\n} {code}","issue_id":"13351932","key":"HBASE-25501","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-01-26T06:40:50.000+0000","role":"fixed_distractor","summary":"Backup not using parameters such as bandwidth, workers, etc while exporting snapshot"} {"case_id":"13360085","cluster":"DISTRACTOR-HBASE-25594","comments":[{"body":"Whoops. I just saw this [~akiraluca] I just pushed \r\n\r\ncommit f4e1ab7b1d87a6af91e0f18012d730c5069ded8a\r\nAuthor: Michael Stack \r\nDate: Mon Mar 15 13:24:15 2021 -0700\r\n\r\n HBASE-25663 Make graceful_stop localhostname compare match even if fqdn (#3048)\r\n\r\nIt just removed the local host stuff and if you want to do local execution, just use 'localhost'. Does this work for you or mess you up?","created":"2021-03-15T21:23:02.594+0000"},{"body":"[~stack]  I think that will perfectly work for what I was trying to achieve.\r\n\r\nThank you.\r\n\r\nHope you backport it to branch-2.","created":"2021-03-16T01:14:16.526+0000"},{"body":"[~stack] I left a comment in your PR [https://github.com/apache/hbase/pull/3048#pullrequestreview-612818495]\r\n\r\nI would appreciate if you have a look at it since I am not 100% it works as expected.","created":"2021-03-16T03:20:49.218+0000"},{"body":"Reopening to apply the fix here instead of what was over in HBASE-25663.","created":"2021-03-16T04:34:41.102+0000"},{"body":"I pushed your PR to 2.4+ [~akiraluca] (Added the doc changes from HBASE-25663 here when I pushed). Thanks for the PR. Shout if you need it to go back further.","created":"2021-03-16T04:44:41.431+0000"},{"body":"Reopening to apply addendum.","created":"2021-03-18T19:08:46.088+0000"},{"body":"I pushed the below to 2.4+\r\n{code}\r\nI pushed this to 2.3+\r\n\r\ncommit 728d4f5ab12fd2631b1ef0a7c61203e9acfb05f0 (HEAD -> 2.3, origin/branch-2.3)\r\nAuthor: Javier Akira Luca de Tena \r\nDate: Fri Mar 19 04:04:54 2021 +0900\r\n\r\n HBOPS-25594 Make easier to use graceful_stop on localhost mode (#3054)\r\n\r\n Co-authored-by: Javier \r\n\r\ndiff --git a/bin/graceful_stop.sh b/bin/graceful_stop.sh\r\nindex 89e3dd939c..e565929606 100755\r\n--- a/bin/graceful_stop.sh\r\n+++ b/bin/graceful_stop.sh\r\n@@ -32,7 +32,7 @@ moving regions\"\r\n echo \" maxthreads xx Limit the number of threads used by the region mover. Default value is 1.\"\r\n echo \" movetimeout xx Timeout for moving regions. If regions are not moved by the timeout value,\\\r\n exit with error. Default value is INT_MAX.\"\r\n- echo \" hostname Hostname of server we are to stop\"\r\n+ echo \" hostname Hostname to stop; match what HBase uses; pass 'localhost' if local to avoid ssh\"\r\n echo \" e|failfast Set -e so exit immediately if any command exits with non-zero status\"\r\n echo \" nob| nobalancer Do not manage balancer states. This is only used as optimization in \\\r\n rolling_restart.sh to avoid multiple calls to hbase shell\"\r\n@@ -100,6 +100,10 @@ localhostname=`/bin/hostname`\r\n if [ \"$localhostname\" == \"$hostname\" ]; then\r\n local=true\r\n fi\r\n+if [ \"$localhostname\" == \"$hostname\" ] || [ \"$hostname\" == \"localhost\" ]; then\r\n+ local=true\r\n+ hostname=$localhostname\r\n+fi\r\n\r\n if [ \"$nob\" == \"true\" ]; then\r\n log \"[ $0 ] skipping disabling balancer -nob argument is used\"\r\n{code}","created":"2021-03-18T19:09:03.081+0000"},{"body":"The commit has an incorrect Jira ID, let me fix it.","created":"2021-03-19T11:29:04.047+0000"},{"body":"Reverted the commit and reapplied with the correct commit message.","created":"2021-03-19T11:34:28.248+0000"},{"body":"Thanks again....[~psomogyi]","created":"2021-03-20T20:45:56.896+0000"},{"body":"Reopen to apply addendum below\r\n{code}\r\ncommit 326835e8372cc83092e0ec127650438ff153476a (HEAD -> m, origin/master, origin/HEAD)\r\nAuthor: stack \r\nDate: Sat Mar 20 13:47:18 2021 -0700\r\n\r\n HBASE-25594 Make easier to use graceful_stop on localhost mode (#3054)\r\n Addendum.\r\n\r\ndiff --git a/bin/graceful_stop.sh b/bin/graceful_stop.sh\r\nindex 05919ce72d..fc18239830 100755\r\n--- a/bin/graceful_stop.sh\r\n+++ b/bin/graceful_stop.sh\r\n@@ -105,9 +105,6 @@ filename=\"/tmp/$hostname\"\r\n local=\r\n localhostname=`/bin/hostname -f`\r\n\r\n-if [ \"$localhostname\" == \"$hostname\" ]; then\r\n- local=true\r\n-fi\r\n if [ \"$localhostname\" == \"$hostname\" ] || [ \"$hostname\" == \"localhost\" ]; then\r\n local=true\r\n hostname=$localhostname\r\n{code}","created":"2021-03-20T20:48:59.081+0000"},{"body":"Pushed addendum on branch-2.4+","created":"2021-03-20T20:49:15.363+0000"}],"conversations":[{"body":"We usually use graceful_stop.sh from the Master to restart RegionServers. However, in some scenarios we may not have privileges to restart remote RegionServers (it uses ssh).\r\n But we can still use graceful_stop.sh on the same host we want to restart.\r\n\r\nIn order to detect the execution at localhost, graceful_stop.sh uses /bin/hostname.\r\n [https://github.com/apache/hbase/blob/cfbae4d3a37e7ac4d795461c3e19406a2786838d/bin/graceful_stop.sh#L106-L110]\r\n\r\nWhen RegionMover strips the host to not include it in the list of target hosts, we filter it out by checking all RegionServer hosts in the cluster:\r\n [https://github.com/apache/hbase/blob/branch-2/hbase-server/src/main/java/org/apache/hadoop/hbase/util/RegionMover.java#L382-L384]\r\n [https://github.com/apache/hbase/blob/cfbae4d3a37e7ac4d795461c3e19406a2786838d/hbase-server/src/main/java/org/apache/hadoop/hbase/util/RegionMover.java#L692]\r\n\r\nBut the list of RegionServer hosts returned by Admin#getRegionServers are FDQN, while the hostname provided from graceful_stop.sh is not FDQN, making the comparison fail.\r\n\r\nSame happens for branch-1 region_mover.rb, which is the place I reproduced in my environment: \r\n[https://github.com/apache/hbase/blob/f9a91488b2c39320bed502619bf7adb765c79de6/bin/region_mover.rb#L305]\r\n[https://github.com/apache/hbase/blob/f9a91488b2c39320bed502619bf7adb765c79de6/bin/region_mover.rb#L175]\r\n [https://github.com/apache/hbase/blob/f9a91488b2c39320bed502619bf7adb765c79de6/bin/region_mover.rb#L186-L192]\r\n\r\n \r\n\r\nThis can be fixed just by using \"/bin/hostname -f\" in the graceful_stop.sh script.\r\n\r\nWill provide patch soon.","from":"reporter","subject":"graceful_stop.sh fails to unload regions when ran at localhost"},{"body":"Whoops. I just saw this [~akiraluca] I just pushed \r\n\r\ncommit f4e1ab7b1d87a6af91e0f18012d730c5069ded8a\r\nAuthor: Michael Stack \r\nDate: Mon Mar 15 13:24:15 2021 -0700\r\n\r\n HBASE-25663 Make graceful_stop localhostname compare match even if fqdn (#3048)\r\n\r\nIt just removed the local host stuff and if you want to do local execution, just use 'localhost'. Does this work for you or mess you up?","from":"developer"},{"body":"[~stack]  I think that will perfectly work for what I was trying to achieve.\r\n\r\nThank you.\r\n\r\nHope you backport it to branch-2.","from":"developer"},{"body":"[~stack] I left a comment in your PR [https://github.com/apache/hbase/pull/3048#pullrequestreview-612818495]\r\n\r\nI would appreciate if you have a look at it since I am not 100% it works as expected.","from":"developer"},{"body":"Reopening to apply the fix here instead of what was over in HBASE-25663.","from":"developer"},{"body":"I pushed your PR to 2.4+ [~akiraluca] (Added the doc changes from HBASE-25663 here when I pushed). Thanks for the PR. Shout if you need it to go back further.","from":"developer"},{"body":"Reopening to apply addendum.","from":"developer"},{"body":"I pushed the below to 2.4+\r\n{code}\r\nI pushed this to 2.3+\r\n\r\ncommit 728d4f5ab12fd2631b1ef0a7c61203e9acfb05f0 (HEAD -> 2.3, origin/branch-2.3)\r\nAuthor: Javier Akira Luca de Tena \r\nDate: Fri Mar 19 04:04:54 2021 +0900\r\n\r\n HBOPS-25594 Make easier to use graceful_stop on localhost mode (#3054)\r\n\r\n Co-authored-by: Javier \r\n\r\ndiff --git a/bin/graceful_stop.sh b/bin/graceful_stop.sh\r\nindex 89e3dd939c..e565929606 100755\r\n--- a/bin/graceful_stop.sh\r\n+++ b/bin/graceful_stop.sh\r\n@@ -32,7 +32,7 @@ moving regions\"\r\n echo \" maxthreads xx Limit the number of threads used by the region mover. Default value is 1.\"\r\n echo \" movetimeout xx Timeout for moving regions. If regions are not moved by the timeout value,\\\r\n exit with error. Default value is INT_MAX.\"\r\n- echo \" hostname Hostname of server we are to stop\"\r\n+ echo \" hostname Hostname to stop; match what HBase uses; pass 'localhost' if local to avoid ssh\"\r\n echo \" e|failfast Set -e so exit immediately if any command exits with non-zero status\"\r\n echo \" nob| nobalancer Do not manage balancer states. This is only used as optimization in \\\r\n rolling_restart.sh to avoid multiple calls to hbase shell\"\r\n@@ -100,6 +100,10 @@ localhostname=`/bin/hostname`\r\n if [ \"$localhostname\" == \"$hostname\" ]; then\r\n local=true\r\n fi\r\n+if [ \"$localhostname\" == \"$hostname\" ] || [ \"$hostname\" == \"localhost\" ]; then\r\n+ local=true\r\n+ hostname=$localhostname\r\n+fi\r\n\r\n if [ \"$nob\" == \"true\" ]; then\r\n log \"[ $0 ] skipping disabling balancer -nob argument is used\"\r\n{code}","from":"developer"},{"body":"The commit has an incorrect Jira ID, let me fix it.","from":"developer"},{"body":"Reverted the commit and reapplied with the correct commit message.","from":"developer"},{"body":"Thanks again....[~psomogyi]","from":"developer"},{"body":"Reopen to apply addendum below\r\n{code}\r\ncommit 326835e8372cc83092e0ec127650438ff153476a (HEAD -> m, origin/master, origin/HEAD)\r\nAuthor: stack \r\nDate: Sat Mar 20 13:47:18 2021 -0700\r\n\r\n HBASE-25594 Make easier to use graceful_stop on localhost mode (#3054)\r\n Addendum.\r\n\r\ndiff --git a/bin/graceful_stop.sh b/bin/graceful_stop.sh\r\nindex 05919ce72d..fc18239830 100755\r\n--- a/bin/graceful_stop.sh\r\n+++ b/bin/graceful_stop.sh\r\n@@ -105,9 +105,6 @@ filename=\"/tmp/$hostname\"\r\n local=\r\n localhostname=`/bin/hostname -f`\r\n\r\n-if [ \"$localhostname\" == \"$hostname\" ]; then\r\n- local=true\r\n-fi\r\n if [ \"$localhostname\" == \"$hostname\" ] || [ \"$hostname\" == \"localhost\" ]; then\r\n local=true\r\n hostname=$localhostname\r\n{code}","from":"developer"},{"body":"Pushed addendum on branch-2.4+","from":"developer"}],"created":"2021-02-22T11:48:18.000+0000","description":"We usually use graceful_stop.sh from the Master to restart RegionServers. However, in some scenarios we may not have privileges to restart remote RegionServers (it uses ssh).\r\n But we can still use graceful_stop.sh on the same host we want to restart.\r\n\r\nIn order to detect the execution at localhost, graceful_stop.sh uses /bin/hostname.\r\n [https://github.com/apache/hbase/blob/cfbae4d3a37e7ac4d795461c3e19406a2786838d/bin/graceful_stop.sh#L106-L110]\r\n\r\nWhen RegionMover strips the host to not include it in the list of target hosts, we filter it out by checking all RegionServer hosts in the cluster:\r\n [https://github.com/apache/hbase/blob/branch-2/hbase-server/src/main/java/org/apache/hadoop/hbase/util/RegionMover.java#L382-L384]\r\n [https://github.com/apache/hbase/blob/cfbae4d3a37e7ac4d795461c3e19406a2786838d/hbase-server/src/main/java/org/apache/hadoop/hbase/util/RegionMover.java#L692]\r\n\r\nBut the list of RegionServer hosts returned by Admin#getRegionServers are FDQN, while the hostname provided from graceful_stop.sh is not FDQN, making the comparison fail.\r\n\r\nSame happens for branch-1 region_mover.rb, which is the place I reproduced in my environment: \r\n[https://github.com/apache/hbase/blob/f9a91488b2c39320bed502619bf7adb765c79de6/bin/region_mover.rb#L305]\r\n[https://github.com/apache/hbase/blob/f9a91488b2c39320bed502619bf7adb765c79de6/bin/region_mover.rb#L175]\r\n [https://github.com/apache/hbase/blob/f9a91488b2c39320bed502619bf7adb765c79de6/bin/region_mover.rb#L186-L192]\r\n\r\n \r\n\r\nThis can be fixed just by using \"/bin/hostname -f\" in the graceful_stop.sh script.\r\n\r\nWill provide patch soon.","issue_id":"13360085","key":"HBASE-25594","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-03-20T20:49:15.000+0000","role":"fixed_distractor","summary":"graceful_stop.sh fails to unload regions when ran at localhost"} {"case_id":"13360278","cluster":"DISTRACTOR-HBASE-25598","comments":[{"body":"Thanks [~zhangduo] for reviewing.\r\n\r\nMerged to master and all active branch-2.x.","created":"2021-02-24T06:50:34.071+0000"}],"conversations":[{"body":"In some PRs, I got the following errors in UT results.\r\n{code:java}\r\n[ERROR] Errors: \r\n[ERROR] org.apache.hadoop.hbase.client.TestFromClientSide5.testScanMetrics[0]\r\n[ERROR] Run 1: TestFromClientSide5.testScanMetrics:1018 Did not count the result bytes expected:<60> but was:<120>\r\n[ERROR] Run 2: TestFromClientSide5.testScanMetrics:1036 Did not count the result bytes expected:<60> but was:<180>\r\n[ERROR] Run 3: TestFromClientSide5.testScanMetrics:951 » MasterRegistryFetch Exception making...\r\n[INFO] \r\n[ERROR] org.apache.hadoop.hbase.client.TestFromClientSideWithCoprocessor5.testScanMetrics[1]\r\n[ERROR] Run 1: TestFromClientSideWithCoprocessor5>TestFromClientSide5.testScanMetrics:1036 Did not count the result bytes expected:<60> but was:<120>\r\n[ERROR] Run 2: TestFromClientSideWithCoprocessor5>TestFromClientSide5.testScanMetrics:951 » IO\r\n[ERROR] Run 3: TestFromClientSideWithCoprocessor5>TestFromClientSide5.testScanMetrics:951 » IO\r\n[INFO] \r\n{code}\r\nI read the code further and found that this UT is flaky.\r\n{code:java}\r\n// check byte counters\r\nscan2 = new Scan();\r\nscan2.setScanMetricsEnabled(true);\r\nscan2.setCaching(1);\r\ntry (ResultScanner scanner = ht.getScanner(scan2)) {\r\n int numBytes = 0;\r\n for (Result result : scanner.next(1)) {\r\n for (Cell cell : result.listCells()) {\r\n numBytes += PrivateCellUtil.estimatedSerializedSizeOf(cell);\r\n }\r\n }\r\n scanner.close();\r\n ScanMetrics scanMetrics = scanner.getScanMetrics();\r\n assertEquals(\"Did not count the result bytes\", numBytes,\r\n scanMetrics.countOfBytesInResults.get());\r\n}\r\n{code}\r\nIn the code above, it is to check scanMetrics.countOfBytesInResults, but just get only ONE row by scanner.next(1) . A total of 3 rows are inserted into the table, and scanner prefetch from server in advance until maxCacheSize is exceeded, see [here|https://github.com/apache/hbase/blob/5fa15cfde3d77e77ffb1f09d60dce4db264f3831/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncTableResultScanner.java#L94].\r\n\r\nSo if scanner prefetch more than one row before closing scanner, the UT fails. we can reproduce this problem steadily by sleeping before scanner.close().","from":"reporter","subject":"TestFromClientSide5.testScanMetrics is flaky"},{"body":"Thanks [~zhangduo] for reviewing.\r\n\r\nMerged to master and all active branch-2.x.","from":"developer"}],"created":"2021-02-23T08:17:29.000+0000","description":"In some PRs, I got the following errors in UT results.\r\n{code:java}\r\n[ERROR] Errors: \r\n[ERROR] org.apache.hadoop.hbase.client.TestFromClientSide5.testScanMetrics[0]\r\n[ERROR] Run 1: TestFromClientSide5.testScanMetrics:1018 Did not count the result bytes expected:<60> but was:<120>\r\n[ERROR] Run 2: TestFromClientSide5.testScanMetrics:1036 Did not count the result bytes expected:<60> but was:<180>\r\n[ERROR] Run 3: TestFromClientSide5.testScanMetrics:951 » MasterRegistryFetch Exception making...\r\n[INFO] \r\n[ERROR] org.apache.hadoop.hbase.client.TestFromClientSideWithCoprocessor5.testScanMetrics[1]\r\n[ERROR] Run 1: TestFromClientSideWithCoprocessor5>TestFromClientSide5.testScanMetrics:1036 Did not count the result bytes expected:<60> but was:<120>\r\n[ERROR] Run 2: TestFromClientSideWithCoprocessor5>TestFromClientSide5.testScanMetrics:951 » IO\r\n[ERROR] Run 3: TestFromClientSideWithCoprocessor5>TestFromClientSide5.testScanMetrics:951 » IO\r\n[INFO] \r\n{code}\r\nI read the code further and found that this UT is flaky.\r\n{code:java}\r\n// check byte counters\r\nscan2 = new Scan();\r\nscan2.setScanMetricsEnabled(true);\r\nscan2.setCaching(1);\r\ntry (ResultScanner scanner = ht.getScanner(scan2)) {\r\n int numBytes = 0;\r\n for (Result result : scanner.next(1)) {\r\n for (Cell cell : result.listCells()) {\r\n numBytes += PrivateCellUtil.estimatedSerializedSizeOf(cell);\r\n }\r\n }\r\n scanner.close();\r\n ScanMetrics scanMetrics = scanner.getScanMetrics();\r\n assertEquals(\"Did not count the result bytes\", numBytes,\r\n scanMetrics.countOfBytesInResults.get());\r\n}\r\n{code}\r\nIn the code above, it is to check scanMetrics.countOfBytesInResults, but just get only ONE row by scanner.next(1) . A total of 3 rows are inserted into the table, and scanner prefetch from server in advance until maxCacheSize is exceeded, see [here|https://github.com/apache/hbase/blob/5fa15cfde3d77e77ffb1f09d60dce4db264f3831/hbase-client/src/main/java/org/apache/hadoop/hbase/client/AsyncTableResultScanner.java#L94].\r\n\r\nSo if scanner prefetch more than one row before closing scanner, the UT fails. we can reproduce this problem steadily by sleeping before scanner.close().","issue_id":"13360278","key":"HBASE-25598","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-02-24T06:50:34.000+0000","role":"fixed_distractor","summary":"TestFromClientSide5.testScanMetrics is flaky"} {"case_id":"13367298","cluster":"DISTRACTOR-HBASE-25692","comments":[{"body":"I've verified that all 2.x and master branches have this problem. I'll double check what's going on with 1.x when we have a fix that can be cherry-picked.","created":"2021-03-24T16:47:30.035+0000"},{"body":"Merged to 2.3+. Shout if you want it to go elsewhere [~elserj].","created":"2021-03-29T19:21:25.307+0000"}],"conversations":[{"body":"I was looking at an HBase user's cluster with [~danilocop] where they saw two otherwise identical clusters where one of them was regularly had sockets in CLOSE_WAIT going from RegionServers to a distributed storage appliance.\r\n\r\nAfter a lot of analysis, we eventually figured out that these sockets in CLOSE_WAIT were directly related to an FSDataInputStream which we forgot to close inside of the RegionServer. The subtlety was that only one of these HBase clusters was set up to do replication (to the other cluster). The HBase cluster experiencing this problem was shipping edits to a peer, and had previously been using Phoenix. At some point, the cluster had Phoenix removed from it.\r\n\r\nWhat we found was that replication still had WALs to ship which were for Phoenix tables. Phoenix, in this version, still used the custom WALCellCodec; however, this codec class was missing from the RS classpath after the owner of the cluster removed Phoenix.\r\n\r\nWhen we try to instantiate the Codec implementation via ReflectionUtils, we end up throwing an UnsupportedOperationException which wraps a NoClassDefFoundException. However, in WALFactory, we _only_ close the FSDataInputStream when we catch an IOException. \r\n\r\nThus, replication sits in a \"fast\" loop, trying to ship these edits, each time leaking a new socket because of the InputStream not being closed. There is an obvious workaround for this specific issue, but we should not leak this inside HBase.\r\n\r\nApproximate, 2.1.x stack trace which lead us to this is below.\r\n{noformat}\r\n2021-03-11 18:19:20,364 ERROR org.apache.hadoop.hbase.replication.regionserver.ReplicationSourceWALReader: Failed to read stream of replication entries\r\njava.io.IOException: Cannot get log reader\r\n\tat org.apache.hadoop.hbase.wal.WALFactory.createReader(WALFactory.java:366)\r\n\tat org.apache.hadoop.hbase.wal.WALFactory.createReader(WALFactory.java:303)\r\n\tat org.apache.hadoop.hbase.wal.WALFactory.createReader(WALFactory.java:291)\r\n\tat org.apache.hadoop.hbase.wal.WALFactory.createReader(WALFactory.java:427)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.WALEntryStream.openReader(WALEntryStream.java:354)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.WALEntryStream.openNextLog(WALEntryStream.java:302)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.WALEntryStream.checkReader(WALEntryStream.java:293)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.WALEntryStream.tryAdvanceEntry(WALEntryStream.java:174)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.WALEntryStream.hasNext(WALEntryStream.java:100)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSourceWALReader.readWALEntries(ReplicationSourceWALReader.java:192)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSourceWALReader.run(ReplicationSourceWALReader.java:138)\r\nCaused by: java.lang.UnsupportedOperationException: Unable to find org.apache.hadoop.hbase.regionserver.wal.IndexedWALEditCodec\r\n\tat org.apache.hadoop.hbase.util.ReflectionUtils.instantiateWithCustomCtor(ReflectionUtils.java:47)\r\n\tat org.apache.hadoop.hbase.regionserver.wal.WALCellCodec.create(WALCellCodec.java:106)\r\n\tat org.apache.hadoop.hbase.regionserver.wal.ProtobufLogReader.getCodec(ProtobufLogReader.java:301)\r\n\tat org.apache.hadoop.hbase.regionserver.wal.ProtobufLogReader.initAfterCompression(ProtobufLogReader.java:311)\r\n\tat org.apache.hadoop.hbase.regionserver.wal.ReaderBase.init(ReaderBase.java:81)\r\n\tat org.apache.hadoop.hbase.regionserver.wal.ProtobufLogReader.init(ProtobufLogReader.java:168)\r\n\tat org.apache.hadoop.hbase.wal.WALFactory.createReader(WALFactory.java:321)\r\n\t... 10 more\r\nCaused by: java.lang.ClassNotFoundException: org.apache.hadoop.hbase.regionserver.wal.IndexedWALEditCodec\r\n\tat java.net.URLClassLoader.findClass(URLClassLoader.java:381)\r\n\tat java.lang.ClassLoader.loadClass(ClassLoader.java:424)\r\n\tat sun.misc.Launcher$AppClassLoader.loadClass(Launcher.java:349)\r\n\tat java.lang.ClassLoader.loadClass(ClassLoader.java:357)\r\n\tat java.lang.Class.forName0(Native Method)\r\n\tat java.lang.Class.forName(Class.java:264)\r\n\tat org.apache.hadoop.hbase.util.ReflectionUtils.instantiateWithCustomCtor(ReflectionUtils.java:43)\r\n\t... 16 more\r\n{noformat}","from":"reporter","subject":"Failure to instantiate WALCellCodec leaks socket in replication"},{"body":"I've verified that all 2.x and master branches have this problem. I'll double check what's going on with 1.x when we have a fix that can be cherry-picked.","from":"developer"},{"body":"Merged to 2.3+. Shout if you want it to go elsewhere [~elserj].","from":"developer"}],"created":"2021-03-24T16:45:19.000+0000","description":"I was looking at an HBase user's cluster with [~danilocop] where they saw two otherwise identical clusters where one of them was regularly had sockets in CLOSE_WAIT going from RegionServers to a distributed storage appliance.\r\n\r\nAfter a lot of analysis, we eventually figured out that these sockets in CLOSE_WAIT were directly related to an FSDataInputStream which we forgot to close inside of the RegionServer. The subtlety was that only one of these HBase clusters was set up to do replication (to the other cluster). The HBase cluster experiencing this problem was shipping edits to a peer, and had previously been using Phoenix. At some point, the cluster had Phoenix removed from it.\r\n\r\nWhat we found was that replication still had WALs to ship which were for Phoenix tables. Phoenix, in this version, still used the custom WALCellCodec; however, this codec class was missing from the RS classpath after the owner of the cluster removed Phoenix.\r\n\r\nWhen we try to instantiate the Codec implementation via ReflectionUtils, we end up throwing an UnsupportedOperationException which wraps a NoClassDefFoundException. However, in WALFactory, we _only_ close the FSDataInputStream when we catch an IOException. \r\n\r\nThus, replication sits in a \"fast\" loop, trying to ship these edits, each time leaking a new socket because of the InputStream not being closed. There is an obvious workaround for this specific issue, but we should not leak this inside HBase.\r\n\r\nApproximate, 2.1.x stack trace which lead us to this is below.\r\n{noformat}\r\n2021-03-11 18:19:20,364 ERROR org.apache.hadoop.hbase.replication.regionserver.ReplicationSourceWALReader: Failed to read stream of replication entries\r\njava.io.IOException: Cannot get log reader\r\n\tat org.apache.hadoop.hbase.wal.WALFactory.createReader(WALFactory.java:366)\r\n\tat org.apache.hadoop.hbase.wal.WALFactory.createReader(WALFactory.java:303)\r\n\tat org.apache.hadoop.hbase.wal.WALFactory.createReader(WALFactory.java:291)\r\n\tat org.apache.hadoop.hbase.wal.WALFactory.createReader(WALFactory.java:427)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.WALEntryStream.openReader(WALEntryStream.java:354)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.WALEntryStream.openNextLog(WALEntryStream.java:302)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.WALEntryStream.checkReader(WALEntryStream.java:293)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.WALEntryStream.tryAdvanceEntry(WALEntryStream.java:174)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.WALEntryStream.hasNext(WALEntryStream.java:100)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSourceWALReader.readWALEntries(ReplicationSourceWALReader.java:192)\r\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSourceWALReader.run(ReplicationSourceWALReader.java:138)\r\nCaused by: java.lang.UnsupportedOperationException: Unable to find org.apache.hadoop.hbase.regionserver.wal.IndexedWALEditCodec\r\n\tat org.apache.hadoop.hbase.util.ReflectionUtils.instantiateWithCustomCtor(ReflectionUtils.java:47)\r\n\tat org.apache.hadoop.hbase.regionserver.wal.WALCellCodec.create(WALCellCodec.java:106)\r\n\tat org.apache.hadoop.hbase.regionserver.wal.ProtobufLogReader.getCodec(ProtobufLogReader.java:301)\r\n\tat org.apache.hadoop.hbase.regionserver.wal.ProtobufLogReader.initAfterCompression(ProtobufLogReader.java:311)\r\n\tat org.apache.hadoop.hbase.regionserver.wal.ReaderBase.init(ReaderBase.java:81)\r\n\tat org.apache.hadoop.hbase.regionserver.wal.ProtobufLogReader.init(ProtobufLogReader.java:168)\r\n\tat org.apache.hadoop.hbase.wal.WALFactory.createReader(WALFactory.java:321)\r\n\t... 10 more\r\nCaused by: java.lang.ClassNotFoundException: org.apache.hadoop.hbase.regionserver.wal.IndexedWALEditCodec\r\n\tat java.net.URLClassLoader.findClass(URLClassLoader.java:381)\r\n\tat java.lang.ClassLoader.loadClass(ClassLoader.java:424)\r\n\tat sun.misc.Launcher$AppClassLoader.loadClass(Launcher.java:349)\r\n\tat java.lang.ClassLoader.loadClass(ClassLoader.java:357)\r\n\tat java.lang.Class.forName0(Native Method)\r\n\tat java.lang.Class.forName(Class.java:264)\r\n\tat org.apache.hadoop.hbase.util.ReflectionUtils.instantiateWithCustomCtor(ReflectionUtils.java:43)\r\n\t... 16 more\r\n{noformat}","issue_id":"13367298","key":"HBASE-25692","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-03-29T19:21:25.000+0000","role":"fixed_distractor","summary":"Failure to instantiate WALCellCodec leaks socket in replication"} {"case_id":"13368362","cluster":"DISTRACTOR-HBASE-25710","comments":[{"body":"Merged to master. Thanks for PR [~shenshengli]","created":"2021-03-29T18:57:09.339+0000"}],"conversations":[{"body":"The error is shown below:\r\n\r\n19:49:24.213 [main] ERROR org.apache.hadoop.hbase.backup.RestoreDriver - Error while running restore backup\r\njava.io.IOException: Can not restore from backup directory (check Hadoop and HBase logs)\r\n at org.apache.hadoop.hbase.backup.mapreduce.MapReduceRestoreJob.run(MapReduceRestoreJob.java:110) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.util.RestoreTool.incrementalRestoreTable(RestoreTool.java:202) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.impl.RestoreTablesClient.restoreImages(RestoreTablesClient.java:178) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.impl.RestoreTablesClient.restore(RestoreTablesClient.java:221) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.impl.RestoreTablesClient.execute(RestoreTablesClient.java:258) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.impl.BackupAdminImpl.restore(BackupAdminImpl.java:520) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.RestoreDriver.parseAndRun(RestoreDriver.java:179) [hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.RestoreDriver.doWork(RestoreDriver.java:220) [hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.RestoreDriver.run(RestoreDriver.java:256) [hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:76) [hadoop-common-3.1.1.3.0.1.0-187.jar:?]\r\n at org.apache.hadoop.hbase.backup.RestoreDriver.main(RestoreDriver.java:228) [hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\nCaused by: java.io.IOException: No input paths specified in job","from":"reporter","subject":"During the recovery process, an error is thrown if there is an incremental backup of data that has not been updated"},{"body":"Merged to master. Thanks for PR [~shenshengli]","from":"developer"}],"created":"2021-03-29T11:57:09.000+0000","description":"The error is shown below:\r\n\r\n19:49:24.213 [main] ERROR org.apache.hadoop.hbase.backup.RestoreDriver - Error while running restore backup\r\njava.io.IOException: Can not restore from backup directory (check Hadoop and HBase logs)\r\n at org.apache.hadoop.hbase.backup.mapreduce.MapReduceRestoreJob.run(MapReduceRestoreJob.java:110) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.util.RestoreTool.incrementalRestoreTable(RestoreTool.java:202) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.impl.RestoreTablesClient.restoreImages(RestoreTablesClient.java:178) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.impl.RestoreTablesClient.restore(RestoreTablesClient.java:221) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.impl.RestoreTablesClient.execute(RestoreTablesClient.java:258) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.impl.BackupAdminImpl.restore(BackupAdminImpl.java:520) ~[hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.RestoreDriver.parseAndRun(RestoreDriver.java:179) [hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.RestoreDriver.doWork(RestoreDriver.java:220) [hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.hbase.backup.RestoreDriver.run(RestoreDriver.java:256) [hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\n at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:76) [hadoop-common-3.1.1.3.0.1.0-187.jar:?]\r\n at org.apache.hadoop.hbase.backup.RestoreDriver.main(RestoreDriver.java:228) [hbase-backup-3.0.0-SNAPSHOT.jar:3.0.0-SNAPSHOT]\r\nCaused by: java.io.IOException: No input paths specified in job","issue_id":"13368362","key":"HBASE-25710","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-03-29T18:57:09.000+0000","role":"fixed_distractor","summary":"During the recovery process, an error is thrown if there is an incremental backup of data that has not been updated"} {"case_id":"13372457","cluster":"DISTRACTOR-HBASE-25774","comments":[{"body":"FYI [~zhangduo], [~niuyulin], [~stack] ,many PRs are blocked by some flaky UTs... ","created":"2021-04-16T12:49:52.124+0000"},{"body":"Why we meet this IllegalStateException? We interrupt the initialization of a ReplicationSource?","created":"2021-04-16T13:08:43.131+0000"},{"body":"Yes, [~zhangduo]. There are no race conditions between ReplicationSourceManager#addSource and PeerProcedureHandlerImpl#refreshPeerState.\r\n\r\nAs a result, when RSes in UT restart, and is starting the replication service by `startReplicationService`, concurrently `transitReplicationPeerSyncReplicationState` (UT line 91) will make the initialize of a replication source terminate and throws Exception, then the RS will be aborted by this exception because it is in starting the replication service. \r\n\r\nLogs can be seen, \r\n{code:java}\r\n76d634:45149.replicationSource,1] regionserver.HRegionServer(2351): STOPPED: Unexpected exception in RS:2;ece3af76d634:45149.replicationSource,1\r\n{code}","created":"2021-04-16T16:54:26.154+0000"},{"body":"Let me check the code. Thanks for the information.","created":"2021-04-21T06:52:45.948+0000"},{"body":"I think we need to dig more on why the initialization is failed...\r\n\r\nLet me check the full log.\r\n\r\nWill report back later.","created":"2021-04-21T07:18:25.996+0000"},{"body":"OK, this is the problem...\r\n\r\n{noformat}\r\n2021-04-11T11:14:18,231 ERROR [Thread-1083] replication.TestSyncReplicationStandbyKillRS(84): Failed to kill RS\r\njava.lang.NullPointerException: null\r\n\tat org.apache.hadoop.hbase.master.HMaster.getProcedures(HMaster.java:3096) ~[classes/:?]\r\n\tat org.apache.hadoop.hbase.master.ServerManager.areDeadServersInProgress(ServerManager.java:500) ~[classes/:?]\r\n\tat org.apache.hadoop.hbase.replication.TestSyncReplicationStandbyKillRS.waitForRSShutdownToStartAndFinish(TestSyncReplicationStandbyKillRS.java:114) ~[test-classes/:?]\r\n\tat org.apache.hadoop.hbase.replication.TestSyncReplicationStandbyKillRS.lambda$testStandbyKillRegionServer$0(TestSyncReplicationStandbyKillRS.java:78) ~[test-classes/:?]\r\n{noformat}","created":"2021-04-21T07:59:58.866+0000"},{"body":"[~zhangduo] (y) \r\n\r\nDid this exception throw after this log `[Time-limited test] hbase.HBaseTestingUtility(1276): Shutting down minicluster`?\r\n\r\n ","created":"2021-04-21T09:55:47.905+0000"},{"body":"Merged #3189 to master. Will keep an eye on the recent pre commit results.","created":"2021-04-22T02:17:16.952+0000"},{"body":"https://ci-hadoop.apache.org/job/HBase/job/HBase-PreCommit-GitHub-PR/job/PR-3195/1/artifact/yetus-jdk8-hadoop3-check/output/\r\n\r\nFailed in the verification stage.\r\n\r\nLet me dig more.","created":"2021-04-23T07:19:21.876+0000"},{"body":"{noformat}\r\n2021-04-28T06:22:27,806 DEBUG [Time-limited test] replication.TestSyncReplicationStandbyKillRS(110): Going to verify the result, 1000 records expected\r\n2021-04-28T06:22:27,808 DEBUG [RPCClient-NioEventLoopGroup-7-6] client.AsyncRegionLocatorHelper(63): Try updating region=hbase:meta,,1.1588230740, hostname=6839e4a4a176,39473,1619590928211, seqNum=-1 , the old value is region=hbase:meta,,1.1588230740, hostname=6839e4a4a176,39473,1619590928211, seqNum=-1, error=java.net.ConnectException: Call to address=6839e4a4a176:39473null failed on connection exception: org.apache.hbase.thirdparty.io.netty.channel.AbstractChannel$AnnotatedConnectException: Connection refused: 6839e4a4a176/172.17.0.3:39473\r\n2021-04-28T06:22:27,809 DEBUG [RPCClient-NioEventLoopGroup-7-6] client.AsyncRegionLocatorHelper(71): The actual exception when updating region=hbase:meta,,1.1588230740, hostname=6839e4a4a176,39473,1619590928211, seqNum=-1 is java.net.ConnectException: Connection refused\r\n2021-04-28T06:22:27,809 DEBUG [RPCClient-NioEventLoopGroup-7-6] client.AsyncRegionLocatorHelper(87): Try removing region=hbase:meta,,1.1588230740, hostname=6839e4a4a176,39473,1619590928211, seqNum=-1 from cache\r\n2021-04-28T06:22:27,809 DEBUG [RPCClient-NioEventLoopGroup-7-6] ipc.FailedServers(53): Added failed server with address 6839e4a4a176:39473 to list caused by org.apache.hbase.thirdparty.io.netty.channel.AbstractChannel$AnnotatedConnectException: Connection refused: 6839e4a4a176/172.17.0.3:39473\r\n2021-04-28T06:22:27,823 DEBUG [RS:3;6839e4a4a176:42643.replicationSource.wal-reader.6839e4a4a176%2C42643%2C1619590943524,1] wal.ProtobufLogReader(422): EOF at position 381\r\n2021-04-28T06:22:27,848 DEBUG [RS:4;6839e4a4a176:44305.replicationSource.wal-reader.6839e4a4a176%2C44305%2C1619590944923,1] wal.ProtobufLogReader(422): EOF at position 83\r\n2021-04-28T06:22:27,894 DEBUG [RS:5;6839e4a4a176:45961.replicationSource.wal-reader.6839e4a4a176%2C45961%2C1619590947375,1] wal.ProtobufLogReader(422): EOF at position 83\r\n2021-04-28T06:22:27,916 DEBUG [Async-Client-Retry-Timer-pool-0] client.ConnectionUtils(548): Start fetching meta region location from registry\r\n2021-04-28T06:22:27,920 DEBUG [RPCClient-NioEventLoopGroup-7-4] client.ConnectionUtils(556): The fetched meta region location is [region=hbase:meta,,1.1588230740, hostname=6839e4a4a176,44305,1619590944923, seqNum=-1]\r\n2021-04-28T06:22:27,921 DEBUG [RPCClient-NioEventLoopGroup-7-4] ipc.RpcConnection(122): Using SIMPLE authentication for service=ClientService, sasl=false\r\n2021-04-28T06:22:27,924 INFO [RS-EventLoopGroup-14-2] ipc.ServerRpcConnection(557): Connection from 172.17.0.3:37390, version=3.0.0-SNAPSHOT, sasl=false, ugi=jenkins (auth:SIMPLE), service=ClientService\r\n2021-04-28T06:22:27,932 DEBUG [RS:3;6839e4a4a176:42643.replicationSource.wal-reader.6839e4a4a176%2C42643%2C1619590943524,1] wal.ProtobufLogReader(422): EOF at position 381\r\n2021-04-28T06:22:27,955 DEBUG [RS:4;6839e4a4a176:44305.replicationSource.wal-reader.6839e4a4a176%2C44305%2C1619590944923,1] wal.ProtobufLogReader(422): EOF at position 83\r\n2021-04-28T06:22:27,965 DEBUG [RPCClient-NioEventLoopGroup-7-10] client.AsyncNonMetaRegionLocator(355): The fetched location of 'SyncRep', row='\\x00\\x00\\x00\\x00', locateType=CURRENT is [region=SyncRep,,1619590931098.3c8d57984856848ef8526f4bc5b3fbb8., hostname=6839e4a4a176,42643,1619590943524, seqNum=9]\r\n2021-04-28T06:22:27,967 DEBUG [RPCClient-NioEventLoopGroup-7-10] ipc.RpcConnection(122): Using SIMPLE authentication for service=ClientService, sasl=false\r\n2021-04-28T06:22:27,970 INFO [RS-EventLoopGroup-13-1] ipc.ServerRpcConnection(557): Connection from 172.17.0.3:58964, version=3.0.0-SNAPSHOT, sasl=false, ugi=jenkins (auth:SIMPLE), service=ClientService\r\n2021-04-28T06:22:27,971 DEBUG [RpcServer.default.FPBQ.Fifo.handler=2,queue=0,port=42643] ipc.CallRunner(149): callId: 28 service: ClientService methodName: Get size: 107 connection: 172.17.0.3:58964 deadline: 1619591007970, exception=org.apache.hadoop.hbase.DoNotRetryIOException: SyncRep,,1619590931098.3c8d57984856848ef8526f4bc5b3fbb8. is in STANDBY state.\r\n2021-04-28T06:22:28,018 DEBUG [RS:5;6839e4a4a176:45961.replicationSource.wal-reader.6839e4a4a176%2C45961%2C1619590947375,1] wal.ProtobufLogReader(422): EOF at position 83\r\n2021-04-28T06:22:28,072 DEBUG [RS:3;6839e4a4a176:42643.replicationSource.wal-reader.6839e4a4a176%2C42643%2C1619590943524,1] wal.ProtobufLogReader(422): EOF at position 381\r\n2021-04-28T06:22:28,072 DEBUG [RS:4;6839e4a4a176:44305.replicationSource.wal-reader.6839e4a4a176%2C44305%2C1619590944923,1] wal.ProtobufLogReader(422): EOF at position 83\r\n{noformat}\r\n\r\nSeems the problem is that, we failed in the verify stage, and then we scheduled another execution and messed up the output.\r\n\r\nChecked the junit output, the errors are\r\n\r\n{noformat}\r\n[ERROR] org.apache.hadoop.hbase.replication.TestSyncReplicationStandbyKillRS.testStandbyKillRegionServer Time elapsed: 14.416 s <<< ERROR!\r\norg.apache.hadoop.hbase.DoNotRetryIOException: \r\norg.apache.hadoop.hbase.DoNotRetryIOException: SyncRep,,1619590931098.3c8d57984856848ef8526f4bc5b3fbb8. is in STANDBY state.\r\n\tat org.apache.hadoop.hbase.regionserver.RSRpcServices.rejectIfInStandByState(RSRpcServices.java:2546)\r\n\tat org.apache.hadoop.hbase.regionserver.RSRpcServices.get(RSRpcServices.java:2568)\r\n\tat org.apache.hadoop.hbase.shaded.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:45249)\r\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:395)\r\n\tat org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:135)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:338)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:318)\r\n\r\n\tat org.apache.hadoop.hbase.replication.TestSyncReplicationStandbyKillRS.testStandbyKillRegionServer(TestSyncReplicationStandbyKillRS.java:111)\r\nCaused by: org.apache.hadoop.hbase.ipc.RemoteWithExtrasException: \r\norg.apache.hadoop.hbase.DoNotRetryIOException: SyncRep,,1619590931098.3c8d57984856848ef8526f4bc5b3fbb8. is in STANDBY state.\r\n\tat org.apache.hadoop.hbase.regionserver.RSRpcServices.rejectIfInStandByState(RSRpcServices.java:2546)\r\n\tat org.apache.hadoop.hbase.regionserver.RSRpcServices.get(RSRpcServices.java:2568)\r\n\tat org.apache.hadoop.hbase.shaded.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:45249)\r\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:395)\r\n\tat org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:135)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:338)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:318)\r\n{noformat}\r\n\r\nThis means we have incosistent state between master and regionserver, as we have verified that the peer has already been transit to DA.\r\n\r\nLet me dig more.","created":"2021-04-30T04:14:27.225+0000"},{"body":"OK, I think I found a possible race here.\r\n\r\nWe will use AbstractPeerProcedure.refreshPeer to refresh the peer cache at region server side after we update the external peer storage at master side. In this method, we use ServerManager.getOnlineServersList to get the region servers which needs to be refreshed.\r\n\r\nHere comes the problem. At region server side, the initialization sequence is\r\n1. Call regionServerStartup\r\n2. Initialize the regionserver, include loading external peer storage to fill the peer cache.\r\n3. Periodically call regionServerReport.\r\n\r\nAt master side, we will only add the region server to online server list in the first regionServerReport call.\r\n\r\nSo in general, it is possible that, a region server initialized the peer cache before we update the peer storage, and after we get the region servers to fresh, it calls regionServerReport to register itself to online server list.\r\n\r\nThis does not only effect sync replication peer state refresh, but also normal peer modification.\r\n\r\nThe fix is straight forward I think, we need to store the list for region servers which have already called regionServerStartup, and also send RefreshPeerProcedure to these region servers.\r\n","created":"2021-05-06T06:11:47.860+0000"},{"body":"Wrote a UT to reproduce the problem.\r\n\r\nhttps://github.com/Apache9/hbase/commit/50c90a28444a79329d4c23ba2b84354af93403b4","created":"2021-05-06T16:32:17.122+0000"},{"body":"[~apurtell] Do you think we should include this fix in 2.4.3? Or maybe 2.4.4 since the problem has been there for a long time and all active 2.x releases are effected...","created":"2021-05-06T16:35:07.625+0000"},{"body":"Sure, 2.4.3 hasn't gone out yet and I'm going to make a new RC soon. We can get this in. If it can be done by EOB Friday. If not as you say the problem has been around for a while and can go into 2.4.4.","created":"2021-05-06T16:39:50.771+0000"},{"body":"OK, the problem is introduced by HBASE-25032, so the only effected released version is 2.3.5...","created":"2021-05-07T01:40:44.734+0000"},{"body":"So it does not only effect peer modification related procedures. For all the procedures which needs to refresh state on region server, we need to get all the region servers which have called regionServerStartup.\r\n\r\nAnd since it has not been released in 2.4.3 yet, I plan to change the priority to blocker and set fix versions to 2.4.3 and 2.3.6.\r\n\r\nShout if you have other opinions [~apurtell] [~ndimiduk].\r\n\r\nThanks.","created":"2021-05-07T01:51:40.996+0000"},{"body":"Skimmed the code, I do not think it is easy to fix as we call isServerOnline in many places, especially in RSProcedureDispatcher, we will give up if isServerOnline returns, which means we assume that we will only send procedures to online servers.\r\n\r\nSo now I prefer we just revert HBASE-25032, and use another way to not assign regions to regionservers which are not fully initialized yet.\r\n\r\nThanks.","created":"2021-05-07T02:25:23.730+0000"},{"body":"Better to be sure. Let’s revert and reopen that issue. ","created":"2021-05-07T03:08:30.069+0000"},{"body":"It has already been released in 2.3.5 so I do not think we could just revert it from all the code base.\r\n\r\nWhat I mean is provide a new patch here, which reverts the modification in HBASE-25032 and uses another approach to archive the same goal.","created":"2021-05-07T03:19:26.964+0000"},{"body":"We can revert it in 2.3.5 and make a 2.3.5.1 with just that one change, and revert it everywhere else, and not hold up 2.4.3 further. If that is acceptable I will do 2.3.5.1 and 2.4.3 at the same time. ","created":"2021-05-07T03:32:59.942+0000"},{"body":"{quote}\r\nWe can revert it in 2.3.5 and make a 2.3.5.1 with just that one change, and revert it everywhere else, and not hold up 2.4.3 further. If that is acceptable I will do 2.3.5.1 and 2.4.3 at the same time.\r\n{quote}\r\n\r\nI'm OK with this approach. [~ndimiduk] WDYT?\r\n\r\nThanks.","created":"2021-05-07T04:05:32.703+0000"},{"body":"I am reverting HBASE-25032 from master, branch-2, branch-2.3, and branch-2.4 now and will make the 2.3.5.1 and 2.4.3 releases. Voting starts Monday.","created":"2021-05-08T00:48:31.559+0000"},{"body":"Resolving via revert of HBASE-25032","created":"2021-05-08T01:30:34.421+0000"},{"body":"We didn't commit anything to branch other than master, so should we resolve this issue as fixed and set fix versions for all active branches? Not sure, just asking...","created":"2021-05-08T01:33:34.524+0000"},{"body":"Oh, good, just noticed that you committed the revert patch as HBASE-25774.","created":"2021-05-08T01:46:09.967+0000"},{"body":"Yes, I marked the revert with HBASE-25774 . I think we need to set the fix versions for it. I also updated fix versions on HBASE-25032. We are pulling the 2.3.5 release from the distribution mirrors once 2.3.5.1 is out, and 2.3.5.1 will have a correct change log, so I think we will be good. Please let me know if you'd like to see something done differently.","created":"2021-05-08T02:01:28.401+0000"},{"body":"Thanks for the addendum [~zhangduo], I noticed the issue this morning","created":"2021-05-08T19:19:46.390+0000"},{"body":"Nice find, couldn't think of this race when reviewing HBASE-25032, my bad. Agree that reverting it is best short term solution until we fix it cleanly. Coming to the fix, it seems like the issue here is the definition of what \"online\" means. I think we should split it into two states, something like INITIALIZED, REGISTERED. First state means that the RS has initialized (set during regionServerStartup()) but is waiting to be marked ready by master and the second one means that it is ready to receive requests (set in first report). Certain procedures (like refresh peer etc) that are interested in the all servers while code paths like AM are interested the REGISTERED ones. We should audit the code for usages of ServerManager carefully to make sure all code paths are addressed, WDYT? FYI [~caroliney14]","created":"2021-05-09T03:41:46.377+0000"},{"body":"I think the main point of HBASE-25032 is about region assignment.\r\n\r\nThe entry point of the final candidates is ServerManager.createDestinationServersList, I think we could just filter out the region servers which haven't done the first regionServerReport yet.\r\n\r\nOn adding states, it may confuse the developers as we have a ServerState enum, in AssignmentManager related code. Maybe just add a boolean flag in ServerMetrics, something like\r\n\r\nboolean isInitialized();\r\n\r\nAnd in regionServerStartup, we set this flag to false, and in regionServerReport, we set it to true.\r\n\r\nThanks.","created":"2021-05-09T04:16:23.651+0000"},{"body":"Looks like the discussion for how to hand 2.3.5 was split between here and HBASE-25032. I read to the end of the comments on that jira before I did this one, and so I left my opinion over there.","created":"2021-05-10T18:46:56.256+0000"},{"body":"Thanks for figuring the race [~zhangduo] (My bad too for not seeing it on review...)","created":"2021-05-10T20:23:31.745+0000"}],"conversations":[{"body":"[https://ci-hadoop.apache.org/job/HBase/job/HBase-PreCommit-GitHub-PR/job/PR-3025/9/testReport/org.apache.hadoop.hbase.replication/TestSyncReplicationStandbyKillRS/precommit_checks___yetus_jdk8_Hadoop3_checks______/]\r\n{code:java}\r\n...[truncated 391170 chars]...\r\n76d634:45149.replicationSource,1] regionserver.HRegionServer(2351): STOPPED: Unexpected exception in RS:2;ece3af76d634:45149.replicationSource,1\r\n2021-04-11T11:14:40,268 INFO [RS:2;ece3af76d634:45149] regionserver.HeapMemoryManager(218): Stopping\r\n2021-04-11T11:14:40,268 INFO [MemStoreFlusher.0] regionserver.MemStoreFlusher$FlushHandler(384): MemStoreFlusher.0 exiting\r\n2021-04-11T11:14:40,268 INFO [RS:2;ece3af76d634:45149] flush.RegionServerFlushTableProcedureManager(118): Stopping region server flush procedure manager abruptly.\r\n2021-04-11T11:14:40,270 INFO [RS:2;ece3af76d634:45149] snapshot.RegionServerSnapshotManager(136): Stopping RegionServerSnapshotManager abruptly.\r\n2021-04-11T11:14:40,270 INFO [RS:2;ece3af76d634:45149] regionserver.HRegionServer(1146): aborting server ece3af76d634,45149,1618139661734\r\n2021-04-11T11:14:40,272 ERROR [ReplicationExecutor-0.replicationSource,1-ece3af76d634,44745,1618139625245] regionserver.ReplicationSource(428): Unexpected exception in ReplicationExecutor-0.replicationSource,1-ece3af76d634,44745,1618139625245 currentPath=null\r\njava.lang.IllegalStateException: Source should be active.\r\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSource.initialize(ReplicationSource.java:547) ~[classes/:?]\r\n\tat java.lang.Thread.run(Thread.java:748) [?:1.8.0_282]\r\n2021-04-11T11:14:40,272 DEBUG [ReplicationExecutor-0.replicationSource,1-ece3af76d634,44745,1618139625245] regionserver.HRegionServer(2576): Abort already in progress. Ignoring the current request with reason: Unexpected exception in ReplicationExecutor-0.replicationSource,1-ece3af76d634,44745,1618139625245\r\n{code}\r\nMaybe it should use HBASE-24877 to avoid failure of the initialize of ReplicationSource.\r\n\r\n ","from":"reporter","subject":"ServerManager.getOnlineServer may miss some region servers when refreshing state in some procedure implementations"},{"body":"FYI [~zhangduo], [~niuyulin], [~stack] ,many PRs are blocked by some flaky UTs... ","from":"developer"},{"body":"Why we meet this IllegalStateException? We interrupt the initialization of a ReplicationSource?","from":"developer"},{"body":"Yes, [~zhangduo]. There are no race conditions between ReplicationSourceManager#addSource and PeerProcedureHandlerImpl#refreshPeerState.\r\n\r\nAs a result, when RSes in UT restart, and is starting the replication service by `startReplicationService`, concurrently `transitReplicationPeerSyncReplicationState` (UT line 91) will make the initialize of a replication source terminate and throws Exception, then the RS will be aborted by this exception because it is in starting the replication service. \r\n\r\nLogs can be seen, \r\n{code:java}\r\n76d634:45149.replicationSource,1] regionserver.HRegionServer(2351): STOPPED: Unexpected exception in RS:2;ece3af76d634:45149.replicationSource,1\r\n{code}","from":"developer"},{"body":"Let me check the code. Thanks for the information.","from":"developer"},{"body":"I think we need to dig more on why the initialization is failed...\r\n\r\nLet me check the full log.\r\n\r\nWill report back later.","from":"developer"},{"body":"OK, this is the problem...\r\n\r\n{noformat}\r\n2021-04-11T11:14:18,231 ERROR [Thread-1083] replication.TestSyncReplicationStandbyKillRS(84): Failed to kill RS\r\njava.lang.NullPointerException: null\r\n\tat org.apache.hadoop.hbase.master.HMaster.getProcedures(HMaster.java:3096) ~[classes/:?]\r\n\tat org.apache.hadoop.hbase.master.ServerManager.areDeadServersInProgress(ServerManager.java:500) ~[classes/:?]\r\n\tat org.apache.hadoop.hbase.replication.TestSyncReplicationStandbyKillRS.waitForRSShutdownToStartAndFinish(TestSyncReplicationStandbyKillRS.java:114) ~[test-classes/:?]\r\n\tat org.apache.hadoop.hbase.replication.TestSyncReplicationStandbyKillRS.lambda$testStandbyKillRegionServer$0(TestSyncReplicationStandbyKillRS.java:78) ~[test-classes/:?]\r\n{noformat}","from":"developer"},{"body":"[~zhangduo] (y) \r\n\r\nDid this exception throw after this log `[Time-limited test] hbase.HBaseTestingUtility(1276): Shutting down minicluster`?\r\n\r\n ","from":"developer"},{"body":"Merged #3189 to master. Will keep an eye on the recent pre commit results.","from":"developer"},{"body":"https://ci-hadoop.apache.org/job/HBase/job/HBase-PreCommit-GitHub-PR/job/PR-3195/1/artifact/yetus-jdk8-hadoop3-check/output/\r\n\r\nFailed in the verification stage.\r\n\r\nLet me dig more.","from":"developer"},{"body":"{noformat}\r\n2021-04-28T06:22:27,806 DEBUG [Time-limited test] replication.TestSyncReplicationStandbyKillRS(110): Going to verify the result, 1000 records expected\r\n2021-04-28T06:22:27,808 DEBUG [RPCClient-NioEventLoopGroup-7-6] client.AsyncRegionLocatorHelper(63): Try updating region=hbase:meta,,1.1588230740, hostname=6839e4a4a176,39473,1619590928211, seqNum=-1 , the old value is region=hbase:meta,,1.1588230740, hostname=6839e4a4a176,39473,1619590928211, seqNum=-1, error=java.net.ConnectException: Call to address=6839e4a4a176:39473null failed on connection exception: org.apache.hbase.thirdparty.io.netty.channel.AbstractChannel$AnnotatedConnectException: Connection refused: 6839e4a4a176/172.17.0.3:39473\r\n2021-04-28T06:22:27,809 DEBUG [RPCClient-NioEventLoopGroup-7-6] client.AsyncRegionLocatorHelper(71): The actual exception when updating region=hbase:meta,,1.1588230740, hostname=6839e4a4a176,39473,1619590928211, seqNum=-1 is java.net.ConnectException: Connection refused\r\n2021-04-28T06:22:27,809 DEBUG [RPCClient-NioEventLoopGroup-7-6] client.AsyncRegionLocatorHelper(87): Try removing region=hbase:meta,,1.1588230740, hostname=6839e4a4a176,39473,1619590928211, seqNum=-1 from cache\r\n2021-04-28T06:22:27,809 DEBUG [RPCClient-NioEventLoopGroup-7-6] ipc.FailedServers(53): Added failed server with address 6839e4a4a176:39473 to list caused by org.apache.hbase.thirdparty.io.netty.channel.AbstractChannel$AnnotatedConnectException: Connection refused: 6839e4a4a176/172.17.0.3:39473\r\n2021-04-28T06:22:27,823 DEBUG [RS:3;6839e4a4a176:42643.replicationSource.wal-reader.6839e4a4a176%2C42643%2C1619590943524,1] wal.ProtobufLogReader(422): EOF at position 381\r\n2021-04-28T06:22:27,848 DEBUG [RS:4;6839e4a4a176:44305.replicationSource.wal-reader.6839e4a4a176%2C44305%2C1619590944923,1] wal.ProtobufLogReader(422): EOF at position 83\r\n2021-04-28T06:22:27,894 DEBUG [RS:5;6839e4a4a176:45961.replicationSource.wal-reader.6839e4a4a176%2C45961%2C1619590947375,1] wal.ProtobufLogReader(422): EOF at position 83\r\n2021-04-28T06:22:27,916 DEBUG [Async-Client-Retry-Timer-pool-0] client.ConnectionUtils(548): Start fetching meta region location from registry\r\n2021-04-28T06:22:27,920 DEBUG [RPCClient-NioEventLoopGroup-7-4] client.ConnectionUtils(556): The fetched meta region location is [region=hbase:meta,,1.1588230740, hostname=6839e4a4a176,44305,1619590944923, seqNum=-1]\r\n2021-04-28T06:22:27,921 DEBUG [RPCClient-NioEventLoopGroup-7-4] ipc.RpcConnection(122): Using SIMPLE authentication for service=ClientService, sasl=false\r\n2021-04-28T06:22:27,924 INFO [RS-EventLoopGroup-14-2] ipc.ServerRpcConnection(557): Connection from 172.17.0.3:37390, version=3.0.0-SNAPSHOT, sasl=false, ugi=jenkins (auth:SIMPLE), service=ClientService\r\n2021-04-28T06:22:27,932 DEBUG [RS:3;6839e4a4a176:42643.replicationSource.wal-reader.6839e4a4a176%2C42643%2C1619590943524,1] wal.ProtobufLogReader(422): EOF at position 381\r\n2021-04-28T06:22:27,955 DEBUG [RS:4;6839e4a4a176:44305.replicationSource.wal-reader.6839e4a4a176%2C44305%2C1619590944923,1] wal.ProtobufLogReader(422): EOF at position 83\r\n2021-04-28T06:22:27,965 DEBUG [RPCClient-NioEventLoopGroup-7-10] client.AsyncNonMetaRegionLocator(355): The fetched location of 'SyncRep', row='\\x00\\x00\\x00\\x00', locateType=CURRENT is [region=SyncRep,,1619590931098.3c8d57984856848ef8526f4bc5b3fbb8., hostname=6839e4a4a176,42643,1619590943524, seqNum=9]\r\n2021-04-28T06:22:27,967 DEBUG [RPCClient-NioEventLoopGroup-7-10] ipc.RpcConnection(122): Using SIMPLE authentication for service=ClientService, sasl=false\r\n2021-04-28T06:22:27,970 INFO [RS-EventLoopGroup-13-1] ipc.ServerRpcConnection(557): Connection from 172.17.0.3:58964, version=3.0.0-SNAPSHOT, sasl=false, ugi=jenkins (auth:SIMPLE), service=ClientService\r\n2021-04-28T06:22:27,971 DEBUG [RpcServer.default.FPBQ.Fifo.handler=2,queue=0,port=42643] ipc.CallRunner(149): callId: 28 service: ClientService methodName: Get size: 107 connection: 172.17.0.3:58964 deadline: 1619591007970, exception=org.apache.hadoop.hbase.DoNotRetryIOException: SyncRep,,1619590931098.3c8d57984856848ef8526f4bc5b3fbb8. is in STANDBY state.\r\n2021-04-28T06:22:28,018 DEBUG [RS:5;6839e4a4a176:45961.replicationSource.wal-reader.6839e4a4a176%2C45961%2C1619590947375,1] wal.ProtobufLogReader(422): EOF at position 83\r\n2021-04-28T06:22:28,072 DEBUG [RS:3;6839e4a4a176:42643.replicationSource.wal-reader.6839e4a4a176%2C42643%2C1619590943524,1] wal.ProtobufLogReader(422): EOF at position 381\r\n2021-04-28T06:22:28,072 DEBUG [RS:4;6839e4a4a176:44305.replicationSource.wal-reader.6839e4a4a176%2C44305%2C1619590944923,1] wal.ProtobufLogReader(422): EOF at position 83\r\n{noformat}\r\n\r\nSeems the problem is that, we failed in the verify stage, and then we scheduled another execution and messed up the output.\r\n\r\nChecked the junit output, the errors are\r\n\r\n{noformat}\r\n[ERROR] org.apache.hadoop.hbase.replication.TestSyncReplicationStandbyKillRS.testStandbyKillRegionServer Time elapsed: 14.416 s <<< ERROR!\r\norg.apache.hadoop.hbase.DoNotRetryIOException: \r\norg.apache.hadoop.hbase.DoNotRetryIOException: SyncRep,,1619590931098.3c8d57984856848ef8526f4bc5b3fbb8. is in STANDBY state.\r\n\tat org.apache.hadoop.hbase.regionserver.RSRpcServices.rejectIfInStandByState(RSRpcServices.java:2546)\r\n\tat org.apache.hadoop.hbase.regionserver.RSRpcServices.get(RSRpcServices.java:2568)\r\n\tat org.apache.hadoop.hbase.shaded.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:45249)\r\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:395)\r\n\tat org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:135)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:338)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:318)\r\n\r\n\tat org.apache.hadoop.hbase.replication.TestSyncReplicationStandbyKillRS.testStandbyKillRegionServer(TestSyncReplicationStandbyKillRS.java:111)\r\nCaused by: org.apache.hadoop.hbase.ipc.RemoteWithExtrasException: \r\norg.apache.hadoop.hbase.DoNotRetryIOException: SyncRep,,1619590931098.3c8d57984856848ef8526f4bc5b3fbb8. is in STANDBY state.\r\n\tat org.apache.hadoop.hbase.regionserver.RSRpcServices.rejectIfInStandByState(RSRpcServices.java:2546)\r\n\tat org.apache.hadoop.hbase.regionserver.RSRpcServices.get(RSRpcServices.java:2568)\r\n\tat org.apache.hadoop.hbase.shaded.protobuf.generated.ClientProtos$ClientService$2.callBlockingMethod(ClientProtos.java:45249)\r\n\tat org.apache.hadoop.hbase.ipc.RpcServer.call(RpcServer.java:395)\r\n\tat org.apache.hadoop.hbase.ipc.CallRunner.run(CallRunner.java:135)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:338)\r\n\tat org.apache.hadoop.hbase.ipc.RpcExecutor$Handler.run(RpcExecutor.java:318)\r\n{noformat}\r\n\r\nThis means we have incosistent state between master and regionserver, as we have verified that the peer has already been transit to DA.\r\n\r\nLet me dig more.","from":"developer"},{"body":"OK, I think I found a possible race here.\r\n\r\nWe will use AbstractPeerProcedure.refreshPeer to refresh the peer cache at region server side after we update the external peer storage at master side. In this method, we use ServerManager.getOnlineServersList to get the region servers which needs to be refreshed.\r\n\r\nHere comes the problem. At region server side, the initialization sequence is\r\n1. Call regionServerStartup\r\n2. Initialize the regionserver, include loading external peer storage to fill the peer cache.\r\n3. Periodically call regionServerReport.\r\n\r\nAt master side, we will only add the region server to online server list in the first regionServerReport call.\r\n\r\nSo in general, it is possible that, a region server initialized the peer cache before we update the peer storage, and after we get the region servers to fresh, it calls regionServerReport to register itself to online server list.\r\n\r\nThis does not only effect sync replication peer state refresh, but also normal peer modification.\r\n\r\nThe fix is straight forward I think, we need to store the list for region servers which have already called regionServerStartup, and also send RefreshPeerProcedure to these region servers.\r\n","from":"developer"},{"body":"Wrote a UT to reproduce the problem.\r\n\r\nhttps://github.com/Apache9/hbase/commit/50c90a28444a79329d4c23ba2b84354af93403b4","from":"developer"},{"body":"[~apurtell] Do you think we should include this fix in 2.4.3? Or maybe 2.4.4 since the problem has been there for a long time and all active 2.x releases are effected...","from":"developer"},{"body":"Sure, 2.4.3 hasn't gone out yet and I'm going to make a new RC soon. We can get this in. If it can be done by EOB Friday. If not as you say the problem has been around for a while and can go into 2.4.4.","from":"developer"},{"body":"OK, the problem is introduced by HBASE-25032, so the only effected released version is 2.3.5...","from":"developer"},{"body":"So it does not only effect peer modification related procedures. For all the procedures which needs to refresh state on region server, we need to get all the region servers which have called regionServerStartup.\r\n\r\nAnd since it has not been released in 2.4.3 yet, I plan to change the priority to blocker and set fix versions to 2.4.3 and 2.3.6.\r\n\r\nShout if you have other opinions [~apurtell] [~ndimiduk].\r\n\r\nThanks.","from":"developer"},{"body":"Skimmed the code, I do not think it is easy to fix as we call isServerOnline in many places, especially in RSProcedureDispatcher, we will give up if isServerOnline returns, which means we assume that we will only send procedures to online servers.\r\n\r\nSo now I prefer we just revert HBASE-25032, and use another way to not assign regions to regionservers which are not fully initialized yet.\r\n\r\nThanks.","from":"developer"},{"body":"Better to be sure. Let’s revert and reopen that issue. ","from":"developer"},{"body":"It has already been released in 2.3.5 so I do not think we could just revert it from all the code base.\r\n\r\nWhat I mean is provide a new patch here, which reverts the modification in HBASE-25032 and uses another approach to archive the same goal.","from":"developer"},{"body":"We can revert it in 2.3.5 and make a 2.3.5.1 with just that one change, and revert it everywhere else, and not hold up 2.4.3 further. If that is acceptable I will do 2.3.5.1 and 2.4.3 at the same time. ","from":"developer"},{"body":"{quote}\r\nWe can revert it in 2.3.5 and make a 2.3.5.1 with just that one change, and revert it everywhere else, and not hold up 2.4.3 further. If that is acceptable I will do 2.3.5.1 and 2.4.3 at the same time.\r\n{quote}\r\n\r\nI'm OK with this approach. [~ndimiduk] WDYT?\r\n\r\nThanks.","from":"developer"},{"body":"I am reverting HBASE-25032 from master, branch-2, branch-2.3, and branch-2.4 now and will make the 2.3.5.1 and 2.4.3 releases. Voting starts Monday.","from":"developer"},{"body":"Resolving via revert of HBASE-25032","from":"developer"},{"body":"We didn't commit anything to branch other than master, so should we resolve this issue as fixed and set fix versions for all active branches? Not sure, just asking...","from":"developer"},{"body":"Oh, good, just noticed that you committed the revert patch as HBASE-25774.","from":"developer"},{"body":"Yes, I marked the revert with HBASE-25774 . I think we need to set the fix versions for it. I also updated fix versions on HBASE-25032. We are pulling the 2.3.5 release from the distribution mirrors once 2.3.5.1 is out, and 2.3.5.1 will have a correct change log, so I think we will be good. Please let me know if you'd like to see something done differently.","from":"developer"},{"body":"Thanks for the addendum [~zhangduo], I noticed the issue this morning","from":"developer"},{"body":"Nice find, couldn't think of this race when reviewing HBASE-25032, my bad. Agree that reverting it is best short term solution until we fix it cleanly. Coming to the fix, it seems like the issue here is the definition of what \"online\" means. I think we should split it into two states, something like INITIALIZED, REGISTERED. First state means that the RS has initialized (set during regionServerStartup()) but is waiting to be marked ready by master and the second one means that it is ready to receive requests (set in first report). Certain procedures (like refresh peer etc) that are interested in the all servers while code paths like AM are interested the REGISTERED ones. We should audit the code for usages of ServerManager carefully to make sure all code paths are addressed, WDYT? FYI [~caroliney14]","from":"developer"},{"body":"I think the main point of HBASE-25032 is about region assignment.\r\n\r\nThe entry point of the final candidates is ServerManager.createDestinationServersList, I think we could just filter out the region servers which haven't done the first regionServerReport yet.\r\n\r\nOn adding states, it may confuse the developers as we have a ServerState enum, in AssignmentManager related code. Maybe just add a boolean flag in ServerMetrics, something like\r\n\r\nboolean isInitialized();\r\n\r\nAnd in regionServerStartup, we set this flag to false, and in regionServerReport, we set it to true.\r\n\r\nThanks.","from":"developer"},{"body":"Looks like the discussion for how to hand 2.3.5 was split between here and HBASE-25032. I read to the end of the comments on that jira before I did this one, and so I left my opinion over there.","from":"developer"},{"body":"Thanks for figuring the race [~zhangduo] (My bad too for not seeing it on review...)","from":"developer"}],"created":"2021-04-14T16:54:29.000+0000","description":"[https://ci-hadoop.apache.org/job/HBase/job/HBase-PreCommit-GitHub-PR/job/PR-3025/9/testReport/org.apache.hadoop.hbase.replication/TestSyncReplicationStandbyKillRS/precommit_checks___yetus_jdk8_Hadoop3_checks______/]\r\n{code:java}\r\n...[truncated 391170 chars]...\r\n76d634:45149.replicationSource,1] regionserver.HRegionServer(2351): STOPPED: Unexpected exception in RS:2;ece3af76d634:45149.replicationSource,1\r\n2021-04-11T11:14:40,268 INFO [RS:2;ece3af76d634:45149] regionserver.HeapMemoryManager(218): Stopping\r\n2021-04-11T11:14:40,268 INFO [MemStoreFlusher.0] regionserver.MemStoreFlusher$FlushHandler(384): MemStoreFlusher.0 exiting\r\n2021-04-11T11:14:40,268 INFO [RS:2;ece3af76d634:45149] flush.RegionServerFlushTableProcedureManager(118): Stopping region server flush procedure manager abruptly.\r\n2021-04-11T11:14:40,270 INFO [RS:2;ece3af76d634:45149] snapshot.RegionServerSnapshotManager(136): Stopping RegionServerSnapshotManager abruptly.\r\n2021-04-11T11:14:40,270 INFO [RS:2;ece3af76d634:45149] regionserver.HRegionServer(1146): aborting server ece3af76d634,45149,1618139661734\r\n2021-04-11T11:14:40,272 ERROR [ReplicationExecutor-0.replicationSource,1-ece3af76d634,44745,1618139625245] regionserver.ReplicationSource(428): Unexpected exception in ReplicationExecutor-0.replicationSource,1-ece3af76d634,44745,1618139625245 currentPath=null\r\njava.lang.IllegalStateException: Source should be active.\r\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSource.initialize(ReplicationSource.java:547) ~[classes/:?]\r\n\tat java.lang.Thread.run(Thread.java:748) [?:1.8.0_282]\r\n2021-04-11T11:14:40,272 DEBUG [ReplicationExecutor-0.replicationSource,1-ece3af76d634,44745,1618139625245] regionserver.HRegionServer(2576): Abort already in progress. Ignoring the current request with reason: Unexpected exception in ReplicationExecutor-0.replicationSource,1-ece3af76d634,44745,1618139625245\r\n{code}\r\nMaybe it should use HBASE-24877 to avoid failure of the initialize of ReplicationSource.\r\n\r\n ","issue_id":"13372457","key":"HBASE-25774","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-05-08T01:30:34.000+0000","role":"fixed_distractor","summary":"ServerManager.getOnlineServer may miss some region servers when refreshing state in some procedure implementations"} {"case_id":"13377329","cluster":"DISTRACTOR-HBASE-25869","comments":[{"body":"*WAL Compression Results*\r\n\r\nSite configuration:\r\n{noformat}\r\n\r\n\r\n hbase.master.logcleaner.ttl\r\n 604800000\r\n\r\n\r\n\r\n\r\n hbase.regionserver.wal.enablecompression\r\n true\r\n\r\n\r\n\r\n\r\n hbase.regionserver.wal.value.enablecompression\r\n true\r\n\r\n\r\n\r\n\r\n hbase.master.logcleaner.ttl\r\n 604800000\r\n\r\n\r\n\r\n\r\n hbase.regionserver.wal.enablecompression\r\n true\r\n\r\n\r\n\r\n\r\n hbase.regionserver.wal.value.enablecompression\r\n true\r\n\r\n\r\n HERE\r\n readNextEntryAndSetPosition();\r\n if (currentEntry == null) {\r\n if (checkAllBytesParsed()) { // now we're certain we're done with this log file\r\n dequeueCurrentLog();\r\n if (openNextLog()) {\r\n readNextEntryAndSetPosition();\r\n }\r\n }\r\n }\r\n } // no other logs, we've simply hit the end of the current open log. Do nothing\r\n }\r\n }\r\n // do nothing if we don't have a WAL Reader (e.g. if there's no logs in queue)\r\n }\r\n{code}\r\n\r\nIn resetReader, we call the following methods, WALEntryStream#resetReader ----> ProtobufLogReader#reset ---> ProtobufLogReader#initInternal.\r\nIn ProtobufLogReader#initInternal, we try to create the whole reader object from scratch to see if any new data has been written.\r\nWe reset all the fields of ProtobufLogReader except for ReaderBase#fileLength.\r\nWe calculate whether trailer is present or not depending on fileLength.","from":"reporter","subject":"Seeing a spike in uncleanlyClosedWALs metric."},{"body":"Thanks for the fix [~shahrs87] !","from":"developer"},{"body":"Thank you [~apurtell] for the review and commit and [~bharathv] [~vjasani] for the reviews. ","from":"developer"},{"body":"This change causes a consistent test failure on master branch (at least). It's been showing up in recent precommit reports. \r\n\r\n{noformat}\r\n[ERROR] Failures: \r\n[ERROR] org.apache.hadoop.hbase.replication.regionserver.TestWALEntryStream.testCleanClosedWALs\r\n[ERROR] Run 1: TestWALEntryStream.testCleanClosedWALs:882 expected:<0> but was:<1>\r\n[ERROR] Run 2: TestWALEntryStream.testCleanClosedWALs:882 expected:<0> but was:<1>\r\n[ERROR] Run 3: TestWALEntryStream.testCleanClosedWALs:882 expected:<0> but was:<1>\r\n{noformat}\r\n\r\nWe need an addendum test fix, or a revert. ","from":"developer"},{"body":"[~apurtell] I created this ticket https://issues.apache.org/jira/browse/HBASE-25932 to track the fix.","from":"developer"},{"body":"The test failure is being tracked by HBASE-25932. Still leaving this open, because we shouldn't have a consistently failing test checked in for long. If it's going to take a while to resolve, better to revert this until its ready. If it can be solved in a day or two, that's fine. ","from":"developer"},{"body":"[~apurtell] let's wait until tomorrow. If  unable to find a fix then lets revert in master and branch-2.","from":"developer"},{"body":"HBASE-25932 is making progress. Seems like eventually it will go in and then this issue can be re-resolved. \r\n\r\nWe shouldn't close this until the issue is resolved one way or another, though. I've linked the JIRAs.","from":"developer"},{"body":"HBASE-25932 is now committed to master/branch-2/branch2.4.\r\n\r\n[~apurtell] Whats the general guidance on back porting to branch-2.3? That branch has diverged quite a bit and this patch doesn't apply cleanly. ","from":"developer"},{"body":"bq. Whats the general guidance on back porting to branch-2.3? \r\n\r\n[~bharathv]\r\n\r\nIt is a live branch that we are still releasing from, so should receive all relevant bug fixes -- and this issue is relevant according to that criteria -- and changes that are meaningful for cross-branch compatibility (i.e. impacting an upgrade from 1.x, or impacting an upgrade to 2.4 or later).\r\n\r\nbq. That branch has diverged quite a bit and this patch doesn't apply cleanly. \r\n\r\nIt's fine to resolve this issue without a 2.3 fix version and open a subtask or another jira for a backport to 2.3, especially if a new PR is advisable due to divergence.","from":"developer"},{"body":"Thanks, opened HBASE-25957 as a subtask. ","from":"developer"},{"body":"This commit also landed on branch-2.3 which was not planned based on the previous comments and HBASE-25957 subtask. This commit broke branch-2.3 builds so let me revert the change there.","from":"developer"},{"body":"[~psomogyi] Instead of reverting this commit, could we pick HBASE-25932 to branch-2.3 ? If yes, then I can put up a PR quickly. Cc [~bharathv] [~apurtell]","from":"developer"},{"body":"Pushed the revert commit to branch-2.3. Resolving.","from":"developer"},{"body":"[~shahrs87] I think it was committed and reverted from 2.3, still would be nice to have a working patch in that branch (if you have spare cycles) :-).","from":"developer"},{"body":"The branches have diverged a lot. It is not trivial work to backport them easily. Unfortunately don't have cycles in few days. Will get back to it later.","from":"developer"}],"created":"2021-05-26T14:41:25.000+0000","description":"Getting the following log line in all of our production clusters when WALEntryStream is dequeuing WAL file.\r\n\r\n{noformat}\r\n 2021-05-02 04:01:30,437 DEBUG [04901996] regionserver.WALEntryStream - Reached the end of WAL file hdfs://. It was not closed cleanly, so we did not parse 8 bytes of data. This is normally ok.\r\n{noformat}\r\nThe 8 bytes are usually the trailer serialized size (SIZE_OF_INT (4bytes) + \"LAWP\" (4 bytes) = 8 bytes)\r\n\r\nWhile dequeue'ing the WAL file from WALEntryStream, we reset the reader here.\r\n[WALEntryStream|https://github.com/apache/hbase/blob/branch-1/hbase-server/src/main/java/org/apache/hadoop/hbase/replication/regionserver/WALEntryStream.java#L199-L221]\r\n\r\n{code:java}\r\n private void tryAdvanceEntry() throws IOException {\r\n if (checkReader()) {\r\n readNextEntryAndSetPosition();\r\n if (currentEntry == null) { // no more entries in this log file - see if log was rolled\r\n if (logQueue.getQueue(walGroupId).size() > 1) { // log was rolled\r\n // Before dequeueing, we should always get one more attempt at reading.\r\n // This is in case more entries came in after we opened the reader,\r\n // and a new log was enqueued while we were reading. See HBASE-6758\r\n resetReader(); ---> HERE\r\n readNextEntryAndSetPosition();\r\n if (currentEntry == null) {\r\n if (checkAllBytesParsed()) { // now we're certain we're done with this log file\r\n dequeueCurrentLog();\r\n if (openNextLog()) {\r\n readNextEntryAndSetPosition();\r\n }\r\n }\r\n }\r\n } // no other logs, we've simply hit the end of the current open log. Do nothing\r\n }\r\n }\r\n // do nothing if we don't have a WAL Reader (e.g. if there's no logs in queue)\r\n }\r\n{code}\r\n\r\nIn resetReader, we call the following methods, WALEntryStream#resetReader ----> ProtobufLogReader#reset ---> ProtobufLogReader#initInternal.\r\nIn ProtobufLogReader#initInternal, we try to create the whole reader object from scratch to see if any new data has been written.\r\nWe reset all the fields of ProtobufLogReader except for ReaderBase#fileLength.\r\nWe calculate whether trailer is present or not depending on fileLength.","issue_id":"13380557","key":"HBASE-25924","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-06-10T17:29:52.000+0000","role":"fixed_distractor","summary":"Seeing a spike in uncleanlyClosedWALs metric."} {"case_id":"13380663","cluster":"DISTRACTOR-HBASE-25929","comments":[{"body":"I add a UT to reproduce the problem. And I will upload a patch later.","created":"2021-05-27T05:06:48.986+0000"},{"body":"Excellent! RC error is already a pain as it is really hard to find out the root cause.","created":"2021-05-27T06:26:31.171+0000"},{"body":"Pused to branch-2.3+. Thanks [~anoop.hbase] and [~zhangduo] for reviewing.","created":"2021-06-03T10:02:49.957+0000"}],"conversations":[{"body":"In our cluster, we found region servers may be crashed in several cases.\r\n\r\nIn hs_err_pid27712.log:\r\n{code:java}\r\nJava frames: (J=compiled Java code, j=interpreted, Vv=VM code)\r\nJ 2687 sun.misc.Unsafe.copyMemory(Ljava/lang/Object;JLjava/lang/Object;JJ)V (0 bytes) @ 0x00007f85c987eda7 [0x00007f85c987ed40+0x67]\r\nJ 5884 C1 org.apache.hadoop.hbase.util.UnsafeAccess.unsafeCopy(Ljava/lang/Object;JLjava/lang/Object;JJ)V (62 bytes) @ 0x00007f85c93fd904 [0x00007f85c93fd780+0x184]\r\nJ 4274 C1 org.apache.hadoop.hbase.util.UnsafeAccess.copy(Ljava/nio/ByteBuffer;I[BII)V (73 bytes) @ 0x00007f85c9d57a94 [0x00007f85c9d574a0+0x5f4]\r\nJ 5211 C2 org.apache.hadoop.hbase.util.ByteBufferUtils.copyFromBufferToArray([BLjava/nio/ByteBuffer;III)V (69 bytes) @ 0x00007f85ca039a34 [0x00007f85ca0399a0+0x94]\r\nJ 5985 C1 org.apache.hadoop.hbase.CellUtil.copyQualifierTo(Lorg/apache/hadoop/hbase/Cell;[BI)I (59 bytes) @ 0x00007f85c9296a34 [0x00007f85c92964c0+0x574]\r\nJ 6011 C1 org.apache.hadoop.hbase.ByteBufferKeyValue.getQualifierArray()[B (5 bytes) @ 0x00007f85c913e094 [0x00007f85c913d4c0+0xbd4]\r\nJ 6004 C1 org.apache.hadoop.hbase.CellUtil.getCellKeyAsString(Lorg/apache/hadoop/hbase/Cell;Ljava/util/function/Function;)Ljava/lang/String; (211 bytes) @ 0x00007f85c93737b4 [0x00007f85c93722e0+0x14d4]\r\nJ 6000 C1 org.apache.hadoop.hbase.CellUtil.getCellKeyAsString(Lorg/apache/hadoop/hbase/Cell;)Ljava/lang/String; (10 bytes) @ 0x00007f85c9854d14 [0x00007f85c9854ba0+0x174]\r\nj org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.getMidpoint(Lorg/apache/hadoop/hbase/CellComparator;Lorg/apache/hadoop/hbase/Cell;Lorg/apache/hadoop/hbase/Cell;Lorg/apache/hadoop/hbase/io/hfile/HFileContext;)Lorg/apache/hadoop/hbase/Cell;+132\r\nj org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.finishBlock()V+102\r\nj org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.checkBlockBoundary()V+32\r\nj org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.append(Lorg/apache/hadoop/hbase/Cell;)V+77\r\nj org.apache.hadoop.hbase.regionserver.StoreFileWriter.append(Lorg/apache/hadoop/hbase/Cell;)V+20\r\nj org.apache.hadoop.hbase.regionserver.compactions.Compactor.performCompaction(Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$FileDetails;Lorg/apache/hadoop/hbase/regionserver/InternalScanner;Lorg/apache/hadoop/hbase/regionserver/CellSink;JZLorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;ZI)Z+318\r\nj org.apache.hadoop.hbase.regionserver.compactions.Compactor.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionRequestImpl;Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$InternalScannerFactory;Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$CellSinkFactory;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+221\r\nj org.apache.hadoop.hbase.regionserver.compactions.DefaultCompactor.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionRequestImpl;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+12\r\nj org.apache.hadoop.hbase.regionserver.DefaultStoreEngine$DefaultCompactionContext.compact(Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+16\r\nj org.apache.hadoop.hbase.regionserver.HStore.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionContext;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+194\r\n{code}\r\nIn hs_err_pid28814.log:\r\n{code:java}\r\nStack: [0x00007f6d8e69b000,0x00007f6d8e6dc000], sp=0x00007f6d8e6d9e88, free space=251k\r\nNative frames: (J=compiled Java code, j=interpreted, Vv=VM code, C=native code)\r\nV [libjvm.so+0x747fa0]\r\nJ 2989 sun.misc.Unsafe.copyMemory(Ljava/lang/Object;JLjava/lang/Object;JJ)V (0 bytes) @ 0x00007f751db756e1 [0x00007f751db75600+0xe1]\r\nj org.apache.hadoop.hbase.util.UnsafeAccess.unsafeCopy(Ljava/lang/Object;JLjava/lang/Object;JJ)V+36\r\nj org.apache.hadoop.hbase.util.UnsafeAccess.copy(Ljava/nio/ByteBuffer;I[BII)V+69\r\nj org.apache.hadoop.hbase.util.ByteBufferUtils.copyFromBufferToArray([BLjava/nio/ByteBuffer;III)V+39\r\nj org.apache.hadoop.hbase.CellUtil.copyQualifierTo(Lorg/apache/hadoop/hbase/Cell;[BI)I+31\r\nJ 12082 C2 org.apache.hadoop.hbase.ByteBufferKeyValue.getQualifierArray()[B (5 bytes) @ 0x00007f751ef15fbc [0x00007f751ef15dc0+0x1fc]\r\nJ 16584 C2 org.apache.hadoop.hbase.CellUtil.getCellKeyAsString(Lorg/apache/hadoop/hbase/Cell;Ljava/util/function/Function;)Ljava/lang/String; (211 bytes) @ 0x00007f751fe320b8 [0x00007f751fe31b80+0x538]\r\nJ 17007 C2 org.apache.hadoop.hbase.regionserver.StoreFileWriter.append(Lorg/apache/hadoop/hbase/Cell;)V (31 bytes) @ 0x00007f751fc2c0f4 [0x00007f751fc2aac0+0x1634]\r\nJ 17178 C2 org.apache.hadoop.hbase.regionserver.compactions.Compactor.performCompaction(Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$FileDetails;Lorg/apache/hadoop/hbase/regionserver/InternalScanner;Lorg/apache/hadoop/hbase/regionserver/CellSink;JZLorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;ZI)Z (767 bytes) @ 0x00007f751f8e330c [0x00007f751f8e2960+0x9ac]\r\nj org.apache.hadoop.hbase.regionserver.compactions.Compactor.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionRequestImpl;Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$InternalScannerFactory;Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$CellSinkFactory;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+221\r\nj org.apache.hadoop.hbase.regionserver.compactions.DefaultCompactor.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionRequestImpl;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+12\r\nj org.apache.hadoop.hbase.regionserver.DefaultStoreEngine$DefaultCompactionContext.compact(Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+16\r\nj org.apache.hadoop.hbase.regionserver.HStore.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionContext;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+194\r\n{code}\r\nSometimes, RS is not crashed but we can see the following logs:\r\n{code:java}\r\n2021-05-27T12:53:54,465 ERROR [RpcServer.default.FPBQ.Fifo.handler=2,queue=0,port=40769-shortCompactions-2] regionserver.CompactSplit$CompactionRunner(640): Compaction failed Request=regionName=t1,user00000000000000000080,1622091224198.cf38ea5f2ea0d90163b53c2b2fd329d2., storeName=A, fileCount=2, fileSize=2.0 M (1009.7 K, 1009.7 K), priority=1, time=16220912334172021-05-27T12:53:54,465 ERROR [RpcServer.default.FPBQ.Fifo.handler=2,queue=0,port=40769-shortCompactions-2] regionserver.CompactSplit$CompactionRunner(640): Compaction failed Request=regionName=t1,user00000000000000000080,1622091224198.cf38ea5f2ea0d90163b53c2b2fd329d2., storeName=A, fileCount=2, fileSize=2.0 M (1009.7 K, 1009.7 K), priority=1, time=1622091233417java.lang.IllegalArgumentException: Left byte array sorts after right row; left=user00000000000000000080, right=user00000000000000000001 at org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.getMinimumMidpointArray(HFileWriterImpl.java:445) ~[classes/:?] at org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.getMidpoint(HFileWriterImpl.java:390) ~[classes/:?]\r\n{code}\r\nBecause byte buffers of cells have been already released, but RS still need to use these cells.","from":"reporter","subject":"RegionServer JVM crash when compaction"},{"body":"I add a UT to reproduce the problem. And I will upload a patch later.","from":"developer"},{"body":"Excellent! RC error is already a pain as it is really hard to find out the root cause.","from":"developer"},{"body":"Pused to branch-2.3+. Thanks [~anoop.hbase] and [~zhangduo] for reviewing.","from":"developer"}],"created":"2021-05-27T04:58:10.000+0000","description":"In our cluster, we found region servers may be crashed in several cases.\r\n\r\nIn hs_err_pid27712.log:\r\n{code:java}\r\nJava frames: (J=compiled Java code, j=interpreted, Vv=VM code)\r\nJ 2687 sun.misc.Unsafe.copyMemory(Ljava/lang/Object;JLjava/lang/Object;JJ)V (0 bytes) @ 0x00007f85c987eda7 [0x00007f85c987ed40+0x67]\r\nJ 5884 C1 org.apache.hadoop.hbase.util.UnsafeAccess.unsafeCopy(Ljava/lang/Object;JLjava/lang/Object;JJ)V (62 bytes) @ 0x00007f85c93fd904 [0x00007f85c93fd780+0x184]\r\nJ 4274 C1 org.apache.hadoop.hbase.util.UnsafeAccess.copy(Ljava/nio/ByteBuffer;I[BII)V (73 bytes) @ 0x00007f85c9d57a94 [0x00007f85c9d574a0+0x5f4]\r\nJ 5211 C2 org.apache.hadoop.hbase.util.ByteBufferUtils.copyFromBufferToArray([BLjava/nio/ByteBuffer;III)V (69 bytes) @ 0x00007f85ca039a34 [0x00007f85ca0399a0+0x94]\r\nJ 5985 C1 org.apache.hadoop.hbase.CellUtil.copyQualifierTo(Lorg/apache/hadoop/hbase/Cell;[BI)I (59 bytes) @ 0x00007f85c9296a34 [0x00007f85c92964c0+0x574]\r\nJ 6011 C1 org.apache.hadoop.hbase.ByteBufferKeyValue.getQualifierArray()[B (5 bytes) @ 0x00007f85c913e094 [0x00007f85c913d4c0+0xbd4]\r\nJ 6004 C1 org.apache.hadoop.hbase.CellUtil.getCellKeyAsString(Lorg/apache/hadoop/hbase/Cell;Ljava/util/function/Function;)Ljava/lang/String; (211 bytes) @ 0x00007f85c93737b4 [0x00007f85c93722e0+0x14d4]\r\nJ 6000 C1 org.apache.hadoop.hbase.CellUtil.getCellKeyAsString(Lorg/apache/hadoop/hbase/Cell;)Ljava/lang/String; (10 bytes) @ 0x00007f85c9854d14 [0x00007f85c9854ba0+0x174]\r\nj org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.getMidpoint(Lorg/apache/hadoop/hbase/CellComparator;Lorg/apache/hadoop/hbase/Cell;Lorg/apache/hadoop/hbase/Cell;Lorg/apache/hadoop/hbase/io/hfile/HFileContext;)Lorg/apache/hadoop/hbase/Cell;+132\r\nj org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.finishBlock()V+102\r\nj org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.checkBlockBoundary()V+32\r\nj org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.append(Lorg/apache/hadoop/hbase/Cell;)V+77\r\nj org.apache.hadoop.hbase.regionserver.StoreFileWriter.append(Lorg/apache/hadoop/hbase/Cell;)V+20\r\nj org.apache.hadoop.hbase.regionserver.compactions.Compactor.performCompaction(Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$FileDetails;Lorg/apache/hadoop/hbase/regionserver/InternalScanner;Lorg/apache/hadoop/hbase/regionserver/CellSink;JZLorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;ZI)Z+318\r\nj org.apache.hadoop.hbase.regionserver.compactions.Compactor.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionRequestImpl;Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$InternalScannerFactory;Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$CellSinkFactory;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+221\r\nj org.apache.hadoop.hbase.regionserver.compactions.DefaultCompactor.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionRequestImpl;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+12\r\nj org.apache.hadoop.hbase.regionserver.DefaultStoreEngine$DefaultCompactionContext.compact(Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+16\r\nj org.apache.hadoop.hbase.regionserver.HStore.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionContext;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+194\r\n{code}\r\nIn hs_err_pid28814.log:\r\n{code:java}\r\nStack: [0x00007f6d8e69b000,0x00007f6d8e6dc000], sp=0x00007f6d8e6d9e88, free space=251k\r\nNative frames: (J=compiled Java code, j=interpreted, Vv=VM code, C=native code)\r\nV [libjvm.so+0x747fa0]\r\nJ 2989 sun.misc.Unsafe.copyMemory(Ljava/lang/Object;JLjava/lang/Object;JJ)V (0 bytes) @ 0x00007f751db756e1 [0x00007f751db75600+0xe1]\r\nj org.apache.hadoop.hbase.util.UnsafeAccess.unsafeCopy(Ljava/lang/Object;JLjava/lang/Object;JJ)V+36\r\nj org.apache.hadoop.hbase.util.UnsafeAccess.copy(Ljava/nio/ByteBuffer;I[BII)V+69\r\nj org.apache.hadoop.hbase.util.ByteBufferUtils.copyFromBufferToArray([BLjava/nio/ByteBuffer;III)V+39\r\nj org.apache.hadoop.hbase.CellUtil.copyQualifierTo(Lorg/apache/hadoop/hbase/Cell;[BI)I+31\r\nJ 12082 C2 org.apache.hadoop.hbase.ByteBufferKeyValue.getQualifierArray()[B (5 bytes) @ 0x00007f751ef15fbc [0x00007f751ef15dc0+0x1fc]\r\nJ 16584 C2 org.apache.hadoop.hbase.CellUtil.getCellKeyAsString(Lorg/apache/hadoop/hbase/Cell;Ljava/util/function/Function;)Ljava/lang/String; (211 bytes) @ 0x00007f751fe320b8 [0x00007f751fe31b80+0x538]\r\nJ 17007 C2 org.apache.hadoop.hbase.regionserver.StoreFileWriter.append(Lorg/apache/hadoop/hbase/Cell;)V (31 bytes) @ 0x00007f751fc2c0f4 [0x00007f751fc2aac0+0x1634]\r\nJ 17178 C2 org.apache.hadoop.hbase.regionserver.compactions.Compactor.performCompaction(Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$FileDetails;Lorg/apache/hadoop/hbase/regionserver/InternalScanner;Lorg/apache/hadoop/hbase/regionserver/CellSink;JZLorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;ZI)Z (767 bytes) @ 0x00007f751f8e330c [0x00007f751f8e2960+0x9ac]\r\nj org.apache.hadoop.hbase.regionserver.compactions.Compactor.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionRequestImpl;Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$InternalScannerFactory;Lorg/apache/hadoop/hbase/regionserver/compactions/Compactor$CellSinkFactory;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+221\r\nj org.apache.hadoop.hbase.regionserver.compactions.DefaultCompactor.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionRequestImpl;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+12\r\nj org.apache.hadoop.hbase.regionserver.DefaultStoreEngine$DefaultCompactionContext.compact(Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+16\r\nj org.apache.hadoop.hbase.regionserver.HStore.compact(Lorg/apache/hadoop/hbase/regionserver/compactions/CompactionContext;Lorg/apache/hadoop/hbase/regionserver/throttle/ThroughputController;Lorg/apache/hadoop/hbase/security/User;)Ljava/util/List;+194\r\n{code}\r\nSometimes, RS is not crashed but we can see the following logs:\r\n{code:java}\r\n2021-05-27T12:53:54,465 ERROR [RpcServer.default.FPBQ.Fifo.handler=2,queue=0,port=40769-shortCompactions-2] regionserver.CompactSplit$CompactionRunner(640): Compaction failed Request=regionName=t1,user00000000000000000080,1622091224198.cf38ea5f2ea0d90163b53c2b2fd329d2., storeName=A, fileCount=2, fileSize=2.0 M (1009.7 K, 1009.7 K), priority=1, time=16220912334172021-05-27T12:53:54,465 ERROR [RpcServer.default.FPBQ.Fifo.handler=2,queue=0,port=40769-shortCompactions-2] regionserver.CompactSplit$CompactionRunner(640): Compaction failed Request=regionName=t1,user00000000000000000080,1622091224198.cf38ea5f2ea0d90163b53c2b2fd329d2., storeName=A, fileCount=2, fileSize=2.0 M (1009.7 K, 1009.7 K), priority=1, time=1622091233417java.lang.IllegalArgumentException: Left byte array sorts after right row; left=user00000000000000000080, right=user00000000000000000001 at org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.getMinimumMidpointArray(HFileWriterImpl.java:445) ~[classes/:?] at org.apache.hadoop.hbase.io.hfile.HFileWriterImpl.getMidpoint(HFileWriterImpl.java:390) ~[classes/:?]\r\n{code}\r\nBecause byte buffers of cells have been already released, but RS still need to use these cells.","issue_id":"13380663","key":"HBASE-25929","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-06-03T10:05:02.000+0000","role":"fixed_distractor","summary":"RegionServer JVM crash when compaction"} {"case_id":"13382158","cluster":"DISTRACTOR-HBASE-25970","comments":[{"body":"Merged to master. Thanks [~wchevreuil] and [~pankajkumar] for the reviews!","created":"2021-06-05T06:58:51.197+0000"},{"body":"Reopening for branch-2 and branch-2.5 backports.","created":"2022-08-19T10:42:06.833+0000"},{"body":"Merged to branch-2 and branch-2.5.","created":"2022-08-19T10:54:23.285+0000"}],"conversations":[{"body":"Active MOB files can be deleted by MobFileCleanerChore. The MOB_FILE_REFS are wrongly concatenated when multiple files or tables are stored. During the MOB cleanup HBase wants to keep the incorrectly concatenated file and considers the real file as deletable.\r\n\r\n{noformat}\r\n2021-06-03 15:03:09,876 TRACE org.apache.hadoop.hbase.mob.MobFileCleanerChore: Specific mob references found for store=hdfs://c1377-node2.coelab.cloudera.com:8020/hbase/data/default/IntegrationTestIngestWithMOB/1df6561244ecc434f505501e6bda9af0/test_cf/2956445a41184488b914e93199df905a : {IntegrationTestIngestWithMOB=[d41d8cd98f00b204e9800998ecf8427e20210603a6c8c0e5bf8548649b0009b156452c21_1df6561244ecc434f505501e6bda9af0d41d8cd98f00b204e9800998ecf8427e20210603497714eabbe5418c9d52382a4808e986_1df6561244ecc434f505501e6bda9af0, d41d8cd98f00b204e9800998ecf8427e202106037464cbddc86943a1bad77d52a3bd412c_1df6561244ecc434f505501e6bda9af0]} \r\n...\r\n2021-06-03 15:03:10,208 TRACE org.apache.hadoop.hbase.mob.MobFileCleanerChore: Archiving MOB file hdfs://c1377-node2.coelab.cloudera.com:8020/hbase/mobdir/data/default/IntegrationTestIngestWithMOB/e9b5d936e7f55a4f1c3246a8d5ce53c2/test_cf/d41d8cd98f00b204e9800998ecf8427e20210603497714eabbe5418c9d52382a4808e986_1df6561244ecc434f505501e6bda9af0 creation time=1622729860319\r\n2021-06-03 15:03:10,467 DEBUG org.apache.hadoop.hbase.mob.MobFileCleanerChore: MOB Cleaner is archiving: hdfs://c1377-node2.coelab.cloudera.com:8020/hbase/mobdir/data/default/IntegrationTestIngestWithMOB/e9b5d936e7f55a4f1c3246a8d5ce53c2/test_cf/d41d8cd98f00b204e9800998ecf8427e20210603497714eabbe5418c9d52382a4808e986_1df6561244ecc434f505501e6bda9af0\r\n{noformat}","from":"reporter","subject":"MOB data loss - incorrect concatenation of MOB_FILE_REFS "},{"body":"Merged to master. Thanks [~wchevreuil] and [~pankajkumar] for the reviews!","from":"developer"},{"body":"Reopening for branch-2 and branch-2.5 backports.","from":"developer"},{"body":"Merged to branch-2 and branch-2.5.","from":"developer"}],"created":"2021-06-04T14:09:27.000+0000","description":"Active MOB files can be deleted by MobFileCleanerChore. The MOB_FILE_REFS are wrongly concatenated when multiple files or tables are stored. During the MOB cleanup HBase wants to keep the incorrectly concatenated file and considers the real file as deletable.\r\n\r\n{noformat}\r\n2021-06-03 15:03:09,876 TRACE org.apache.hadoop.hbase.mob.MobFileCleanerChore: Specific mob references found for store=hdfs://c1377-node2.coelab.cloudera.com:8020/hbase/data/default/IntegrationTestIngestWithMOB/1df6561244ecc434f505501e6bda9af0/test_cf/2956445a41184488b914e93199df905a : {IntegrationTestIngestWithMOB=[d41d8cd98f00b204e9800998ecf8427e20210603a6c8c0e5bf8548649b0009b156452c21_1df6561244ecc434f505501e6bda9af0d41d8cd98f00b204e9800998ecf8427e20210603497714eabbe5418c9d52382a4808e986_1df6561244ecc434f505501e6bda9af0, d41d8cd98f00b204e9800998ecf8427e202106037464cbddc86943a1bad77d52a3bd412c_1df6561244ecc434f505501e6bda9af0]} \r\n...\r\n2021-06-03 15:03:10,208 TRACE org.apache.hadoop.hbase.mob.MobFileCleanerChore: Archiving MOB file hdfs://c1377-node2.coelab.cloudera.com:8020/hbase/mobdir/data/default/IntegrationTestIngestWithMOB/e9b5d936e7f55a4f1c3246a8d5ce53c2/test_cf/d41d8cd98f00b204e9800998ecf8427e20210603497714eabbe5418c9d52382a4808e986_1df6561244ecc434f505501e6bda9af0 creation time=1622729860319\r\n2021-06-03 15:03:10,467 DEBUG org.apache.hadoop.hbase.mob.MobFileCleanerChore: MOB Cleaner is archiving: hdfs://c1377-node2.coelab.cloudera.com:8020/hbase/mobdir/data/default/IntegrationTestIngestWithMOB/e9b5d936e7f55a4f1c3246a8d5ce53c2/test_cf/d41d8cd98f00b204e9800998ecf8427e20210603497714eabbe5418c9d52382a4808e986_1df6561244ecc434f505501e6bda9af0\r\n{noformat}","issue_id":"13382158","key":"HBASE-25970","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2022-08-19T10:54:23.000+0000","role":"fixed_distractor","summary":"MOB data loss - incorrect concatenation of MOB_FILE_REFS "} {"case_id":"13382792","cluster":"DISTRACTOR-HBASE-25984","comments":[{"body":"Obtained a heap dump and poked around quite a bit and observed that in-memory state of the {{RingBufferEventHandler}} is corrupted. After analyzing it and eyeballing the code (long story short), came up with the following theory...\r\n\r\nThis is the current implementation of SyncFuture#reset()\r\n{noformat}\r\n1 synchronized SyncFuture reset(final long sequence, Span span) {\r\n2 if (t != null && t != Thread.currentThread()) throw new IllegalStateException();\r\n3 t = Thread.currentThread();\r\n4 if (!isDone()) throw new IllegalStateException(\"\" + sequence + \" \" + Thread.currentThread());\r\n5 this.doneSequence = NOT_DONE;\r\n6 this.ringBufferSequence = sequence;\r\n7 this.span = span;\r\n8 this.throwable = null;\r\n9 return this;\r\n10 }\r\n{noformat}\r\nWe can see, there are guards against overwriting un-finished sync futures with checks on ‘current thread’ (L2) and !isDone() (L4). These are not tight checks, consider the following sequence of interleaved actions that can result in a deadlock..\r\n # RPC handler#1 → WAL sync - SyncFuture#1 (from ThreadLocalCache) → Add to RingBuffer - state: NOT_DONE\r\n # RingBufferConsumer#onEvent() → SyncFuture#1 dispatch → old state: NOT_DONE, new state: DONE (via releaseSyncFuture() in SyncRunner.run())\r\n # RPCHandler#1 (reused for another region operation) → ThreadLocal SyncFuture#1 NOT_DONE (via reset()) (*this goes through because isDone() now returns true and L2 always returns true because we reuse the ThreadLocal instance*).\r\n #  RingBufferConsumer#onEvent() - attainSafePoint() -> works on the overwritten sync future state (from step 3)\r\n #  ======== DEADLOCK ======= (as the sync future remains in the state forever)\r\n\r\nThe problem here is that once a SyncFuture is marked DONE, it is eligible for reuse. Even though ring buffer still has references to it and is checking on it to attain a safe point, in the background another handler can just overwrite it resulting in ring buffer operating on a wrong future and deadlocking the system. Very subtle bug.. I'm able to reproduce this in a carefully crafted unit test (attached). I have a fix solves this problem, will create a PR for review soon.","created":"2021-06-08T21:19:42.997+0000"},{"body":"This has been a recurring problem for us under load deadlocking a bunch of region servers and severely dropping the ingestion throughput until our probes trigger a force abort of the RS pods and the throughput recovers. I cannot share the heap dump for obvious reasons but can share any other details that reviewers might be interested in.","created":"2021-06-08T21:23:50.364+0000"},{"body":"[~bharathv] I think you can create PR. Support to run QA with uploaded patch is no longer available IIRC.","created":"2021-06-09T12:56:26.986+0000"},{"body":"After reading the AsyncFSWAL implementation, I think the overwrites of the futures are possible but it does not cause a deadlock because of the way the safe point is attained. I uploaded a draft patch that removes the usage of ThreadLocals for both the WAL implementations. It only matters for the FSHLog implementation but in general it seems risky to use ThreadLocals that are prone to overwrites. [~zhangduo] Any thoughts?","created":"2021-06-09T21:02:41.748+0000"},{"body":"FWIW I would prefer we use ThreadLocals only when there is no reasonable alternative. I don't think that bar is reached here, because as the PR demonstrates, a shared cache can work. I'm familiar with this work and some microbenchmarks done on the result (for branch-1, though) where the shared cache approach does not hurt performance and in fact produces a small performance benefit. Let me hold back on further comment until we have microbenchmarks comparing thread local vs shared cache approaches for async WAL.","created":"2021-06-09T21:07:55.030+0000"},{"body":"PR is merged. What is the status of this JIRA? Backports in progress? ","created":"2021-06-17T16:47:23.307+0000"},{"body":"Back port PRs are WIP (PR - jira auto link isn't working, it will catch up soon).\r\n\r\nhttps://github.com/apache/hbase/pull/3392\r\nhttps://github.com/apache/hbase/pull/3393\r\nhttps://github.com/apache/hbase/pull/3394","created":"2021-06-17T17:52:14.582+0000"},{"body":"git bisect flags this pr as why we have 100% fail running TestPostIncrementAndAppendBeforeWAL on branch 2.3 (Run on mac or see bottom of [https://ci-hadoop.apache.org/view/HBase/job/HBase/job/HBase-Find-Flaky-Tests/job/branch-2.3/lastSuccessfulBuild/artifact/output/dashboard.html)] Let me see if can fix.","created":"2021-07-24T22:13:55.419+0000"},{"body":"Ignore my previous comment. Bisect (or more likely the pilot) identified the wrong issue... it is not this that is cause of failed test.","created":"2021-07-26T17:42:49.212+0000"}],"conversations":[{"body":"We use FSHLog as the WAL implementation (branch-1 based) and under heavy load we noticed the WAL system gets locked up due to a subtle bug involving racy code with sync future reuse. This bug applies to all FSHLog implementations across branches.\r\n\r\nSymptoms:\r\n\r\nOn heavily loaded clusters with large write load we noticed that the region servers are hanging abruptly with filled up handler queues and stuck MVCC indicating appends/syncs not making any progress.\r\n\r\n{noformat}\r\n WARN [8,queue=9,port=60020] regionserver.MultiVersionConcurrencyControl - STUCK for : 296000 millis. MultiVersionConcurrencyControl{readPoint=172383686, writePoint=172383690, regionName=1ce4003ab60120057734ffe367667dca}\r\n WARN [6,queue=2,port=60020] regionserver.MultiVersionConcurrencyControl - STUCK for : 296000 millis. MultiVersionConcurrencyControl{readPoint=171504376, writePoint=171504381, regionName=7c441d7243f9f504194dae6bf2622631}\r\n{noformat}\r\n\r\nAll the handlers are stuck waiting for the sync futures and timing out.\r\n\r\n{noformat}\r\n java.lang.Object.wait(Native Method)\r\n org.apache.hadoop.hbase.regionserver.wal.SyncFuture.get(SyncFuture.java:183)\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog.blockOnSync(FSHLog.java:1509)\r\n .....\r\n{noformat}\r\n\r\nLog rolling is stuck because it was unable to attain a safe point\r\n\r\n{noformat}\r\n java.util.concurrent.CountDownLatch.await(CountDownLatch.java:277) \r\norg.apache.hadoop.hbase.regionserver.wal.FSHLog$SafePointZigZagLatch.waitSafePoint(FSHLog.java:1799)\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog.replaceWriter(FSHLog.java:900)\r\n{noformat}\r\n\r\nand the Ring buffer consumer thinks that there are some outstanding syncs that need to finish..\r\n\r\n{noformat}\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog$RingBufferEventHandler.attainSafePoint(FSHLog.java:2031)\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog$RingBufferEventHandler.onEvent(FSHLog.java:1999)\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog$RingBufferEventHandler.onEvent(FSHLog.java:1857)\r\n{noformat}\r\n\r\nOn the other hand, SyncRunner threads are idle and just waiting for work implying that there are no pending SyncFutures that need to be run\r\n\r\n{noformat}\r\n sun.misc.Unsafe.park(Native Method)\r\n java.util.concurrent.locks.LockSupport.park(LockSupport.java:175)\r\n java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject.await(AbstractQueuedSynchronizer.java:2039)\r\n java.util.concurrent.LinkedBlockingQueue.take(LinkedBlockingQueue.java:442)\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog$SyncRunner.run(FSHLog.java:1297)\r\n java.lang.Thread.run(Thread.java:748)\r\n{noformat}\r\n\r\nOverall the WAL system is dead locked and could make no progress until it was aborted. I got to the bottom of this issue and have a patch that can fix it (more details in the comments due to word limit in the description).","from":"reporter","subject":"FSHLog WAL lockup with sync future reuse [RS deadlock]"},{"body":"Obtained a heap dump and poked around quite a bit and observed that in-memory state of the {{RingBufferEventHandler}} is corrupted. After analyzing it and eyeballing the code (long story short), came up with the following theory...\r\n\r\nThis is the current implementation of SyncFuture#reset()\r\n{noformat}\r\n1 synchronized SyncFuture reset(final long sequence, Span span) {\r\n2 if (t != null && t != Thread.currentThread()) throw new IllegalStateException();\r\n3 t = Thread.currentThread();\r\n4 if (!isDone()) throw new IllegalStateException(\"\" + sequence + \" \" + Thread.currentThread());\r\n5 this.doneSequence = NOT_DONE;\r\n6 this.ringBufferSequence = sequence;\r\n7 this.span = span;\r\n8 this.throwable = null;\r\n9 return this;\r\n10 }\r\n{noformat}\r\nWe can see, there are guards against overwriting un-finished sync futures with checks on ‘current thread’ (L2) and !isDone() (L4). These are not tight checks, consider the following sequence of interleaved actions that can result in a deadlock..\r\n # RPC handler#1 → WAL sync - SyncFuture#1 (from ThreadLocalCache) → Add to RingBuffer - state: NOT_DONE\r\n # RingBufferConsumer#onEvent() → SyncFuture#1 dispatch → old state: NOT_DONE, new state: DONE (via releaseSyncFuture() in SyncRunner.run())\r\n # RPCHandler#1 (reused for another region operation) → ThreadLocal SyncFuture#1 NOT_DONE (via reset()) (*this goes through because isDone() now returns true and L2 always returns true because we reuse the ThreadLocal instance*).\r\n #  RingBufferConsumer#onEvent() - attainSafePoint() -> works on the overwritten sync future state (from step 3)\r\n #  ======== DEADLOCK ======= (as the sync future remains in the state forever)\r\n\r\nThe problem here is that once a SyncFuture is marked DONE, it is eligible for reuse. Even though ring buffer still has references to it and is checking on it to attain a safe point, in the background another handler can just overwrite it resulting in ring buffer operating on a wrong future and deadlocking the system. Very subtle bug.. I'm able to reproduce this in a carefully crafted unit test (attached). I have a fix solves this problem, will create a PR for review soon.","from":"developer"},{"body":"This has been a recurring problem for us under load deadlocking a bunch of region servers and severely dropping the ingestion throughput until our probes trigger a force abort of the RS pods and the throughput recovers. I cannot share the heap dump for obvious reasons but can share any other details that reviewers might be interested in.","from":"developer"},{"body":"[~bharathv] I think you can create PR. Support to run QA with uploaded patch is no longer available IIRC.","from":"developer"},{"body":"After reading the AsyncFSWAL implementation, I think the overwrites of the futures are possible but it does not cause a deadlock because of the way the safe point is attained. I uploaded a draft patch that removes the usage of ThreadLocals for both the WAL implementations. It only matters for the FSHLog implementation but in general it seems risky to use ThreadLocals that are prone to overwrites. [~zhangduo] Any thoughts?","from":"developer"},{"body":"FWIW I would prefer we use ThreadLocals only when there is no reasonable alternative. I don't think that bar is reached here, because as the PR demonstrates, a shared cache can work. I'm familiar with this work and some microbenchmarks done on the result (for branch-1, though) where the shared cache approach does not hurt performance and in fact produces a small performance benefit. Let me hold back on further comment until we have microbenchmarks comparing thread local vs shared cache approaches for async WAL.","from":"developer"},{"body":"PR is merged. What is the status of this JIRA? Backports in progress? ","from":"developer"},{"body":"Back port PRs are WIP (PR - jira auto link isn't working, it will catch up soon).\r\n\r\nhttps://github.com/apache/hbase/pull/3392\r\nhttps://github.com/apache/hbase/pull/3393\r\nhttps://github.com/apache/hbase/pull/3394","from":"developer"},{"body":"git bisect flags this pr as why we have 100% fail running TestPostIncrementAndAppendBeforeWAL on branch 2.3 (Run on mac or see bottom of [https://ci-hadoop.apache.org/view/HBase/job/HBase/job/HBase-Find-Flaky-Tests/job/branch-2.3/lastSuccessfulBuild/artifact/output/dashboard.html)] Let me see if can fix.","from":"developer"},{"body":"Ignore my previous comment. Bisect (or more likely the pilot) identified the wrong issue... it is not this that is cause of failed test.","from":"developer"}],"created":"2021-06-08T21:08:59.000+0000","description":"We use FSHLog as the WAL implementation (branch-1 based) and under heavy load we noticed the WAL system gets locked up due to a subtle bug involving racy code with sync future reuse. This bug applies to all FSHLog implementations across branches.\r\n\r\nSymptoms:\r\n\r\nOn heavily loaded clusters with large write load we noticed that the region servers are hanging abruptly with filled up handler queues and stuck MVCC indicating appends/syncs not making any progress.\r\n\r\n{noformat}\r\n WARN [8,queue=9,port=60020] regionserver.MultiVersionConcurrencyControl - STUCK for : 296000 millis. MultiVersionConcurrencyControl{readPoint=172383686, writePoint=172383690, regionName=1ce4003ab60120057734ffe367667dca}\r\n WARN [6,queue=2,port=60020] regionserver.MultiVersionConcurrencyControl - STUCK for : 296000 millis. MultiVersionConcurrencyControl{readPoint=171504376, writePoint=171504381, regionName=7c441d7243f9f504194dae6bf2622631}\r\n{noformat}\r\n\r\nAll the handlers are stuck waiting for the sync futures and timing out.\r\n\r\n{noformat}\r\n java.lang.Object.wait(Native Method)\r\n org.apache.hadoop.hbase.regionserver.wal.SyncFuture.get(SyncFuture.java:183)\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog.blockOnSync(FSHLog.java:1509)\r\n .....\r\n{noformat}\r\n\r\nLog rolling is stuck because it was unable to attain a safe point\r\n\r\n{noformat}\r\n java.util.concurrent.CountDownLatch.await(CountDownLatch.java:277) \r\norg.apache.hadoop.hbase.regionserver.wal.FSHLog$SafePointZigZagLatch.waitSafePoint(FSHLog.java:1799)\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog.replaceWriter(FSHLog.java:900)\r\n{noformat}\r\n\r\nand the Ring buffer consumer thinks that there are some outstanding syncs that need to finish..\r\n\r\n{noformat}\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog$RingBufferEventHandler.attainSafePoint(FSHLog.java:2031)\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog$RingBufferEventHandler.onEvent(FSHLog.java:1999)\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog$RingBufferEventHandler.onEvent(FSHLog.java:1857)\r\n{noformat}\r\n\r\nOn the other hand, SyncRunner threads are idle and just waiting for work implying that there are no pending SyncFutures that need to be run\r\n\r\n{noformat}\r\n sun.misc.Unsafe.park(Native Method)\r\n java.util.concurrent.locks.LockSupport.park(LockSupport.java:175)\r\n java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject.await(AbstractQueuedSynchronizer.java:2039)\r\n java.util.concurrent.LinkedBlockingQueue.take(LinkedBlockingQueue.java:442)\r\n org.apache.hadoop.hbase.regionserver.wal.FSHLog$SyncRunner.run(FSHLog.java:1297)\r\n java.lang.Thread.run(Thread.java:748)\r\n{noformat}\r\n\r\nOverall the WAL system is dead locked and could make no progress until it was aborted. I got to the bottom of this issue and have a patch that can fix it (more details in the comments due to word limit in the description).","issue_id":"13382792","key":"HBASE-25984","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-06-19T00:49:50.000+0000","role":"fixed_distractor","summary":"FSHLog WAL lockup with sync future reuse [RS deadlock]"} {"case_id":"13383708","cluster":"DISTRACTOR-HBASE-26001","comments":[{"body":"It would be better, if you [~xytss123] could create pr for other branches as well, thx!","created":"2021-06-18T03:16:42.307+0000"},{"body":"Cherry picked to branch-2.3 & branch-2.4 as well. Thanks [~xytss123] for the contributions.","created":"2021-06-21T07:37:53.911+0000"},{"body":"Reopening to revert from branch-2.3.   The  new test added here is failing 100% on branch-2.3. See bottom of [https://ci-hadoop.apache.org/view/HBase/job/HBase/job/HBase-Find-Flaky-Tests/job/branch-2.3/lastSuccessfulBuild/artifact/output/dashboard.html] Thanks.","created":"2021-07-26T17:58:47.452+0000"},{"body":"Removed 2.3.6 as fix version after revert.","created":"2021-07-26T18:34:55.967+0000"},{"body":"I run the tests with branch-2.3 locally. And got the following message :\r\n\r\n2021-07-27 15:14:46,016 DEBUG [master/192.168.1.127:0:becomeActiveMaster] asyncfs.FanOutOneBlockAsyncDFSOutputHelper(265): ClientProtocol::create wrong number of arguments, should be hadoop 3.2 or below\r\n2021-07-27 15:14:46,016 DEBUG [master/192.168.1.127:0:becomeActiveMaster] asyncfs.FanOutOneBlockAsyncDFSOutputHelper(271): ClientProtocol::create wrong number of arguments, should be hadoop 2.x\r\n2021-07-27 15:14:46,029 DEBUG [master/192.168.1.127:0:becomeActiveMaster] asyncfs.FanOutOneBlockAsyncDFSOutputHelper(280): can not find SHOULD_REPLICATE flag, should be hadoop 2.x\r\njava.lang.IllegalArgumentException: No enum constant org.apache.hadoop.fs.CreateFlag.SHOULD_REPLICATE\r\n\r\nI also rerun the tests under branch-2 and it passed. \r\nIt seems the tests failure is due to a version problem of hadoop. [~stack]","created":"2021-07-27T07:33:48.937+0000"},{"body":"I don't think the IllegalArgumentException causes the test to fail. The version of hadoop in hbase 2.3 is old but would be odd needing to update it to make a unit test pass? Seems like there is something about the 2.3 hbase context that makes the test fail. I didn't spend much time on it. [~xytss123]  Thanks.","created":"2021-07-28T05:18:27.475+0000"}],"conversations":[{"body":"AccessController postIncrementBeforeWAL() and postAppendBeforeWAL() methods will rewrite the new cell's tags by the old cell's. This will makes the other kinds of tag in new cell invisible (such as TTL tag) after this. As in Increment and Append operations, the new cell has already catch forward all tags of the old cell and TTL tag from mutation operation, here in AccessController we do not need to rewrite the tags once again. Also, the TTL tag of newCell will be invisible in the new created cell. Actually, in Increment and Append operations, the newCell has already copied all tags of the oldCell. So the oldCell is useless here.\r\n\r\n{code:java}\r\nprivate Cell createNewCellWithTags(Mutation mutation, Cell oldCell, Cell newCell) {\r\n // Collect any ACLs from the old cell\r\n List tags = Lists.newArrayList();\r\n List aclTags = Lists.newArrayList();\r\n ListMultimap perms = ArrayListMultimap.create();\r\n if (oldCell != null) {\r\n Iterator tagIterator = PrivateCellUtil.tagsIterator(oldCell);\r\n while (tagIterator.hasNext()) {\r\n Tag tag = tagIterator.next();\r\n if (tag.getType() != PermissionStorage.ACL_TAG_TYPE) {\r\n // Not an ACL tag, just carry it through\r\n if (LOG.isTraceEnabled()) {\r\n LOG.trace(\"Carrying forward tag from \" + oldCell + \": type \" + tag.getType()\r\n + \" length \" + tag.getValueLength());\r\n }\r\n tags.add(tag);\r\n } else {\r\n aclTags.add(tag);\r\n }\r\n }\r\n }\r\n\r\n // Do we have an ACL on the operation?\r\n byte[] aclBytes = mutation.getACL();\r\n if (aclBytes != null) {\r\n // Yes, use it\r\n tags.add(new ArrayBackedTag(PermissionStorage.ACL_TAG_TYPE, aclBytes));\r\n } else {\r\n // No, use what we carried forward\r\n if (perms != null) {\r\n // TODO: If we collected ACLs from more than one tag we may have a\r\n // List of size > 1, this can be collapsed into a single\r\n // Permission\r\n if (LOG.isTraceEnabled()) {\r\n LOG.trace(\"Carrying forward ACLs from \" + oldCell + \": \" + perms);\r\n }\r\n tags.addAll(aclTags);\r\n }\r\n }\r\n\r\n // If we have no tags to add, just return\r\n if (tags.isEmpty()) {\r\n return newCell;\r\n }\r\n // Here the new cell's tags will be in visible.\r\n return PrivateCellUtil.createCell(newCell, tags);\r\n }\r\n{code}\r\n","from":"reporter","subject":"When turn on access control, the cell level TTL of Increment and Append operations is invalid."},{"body":"It would be better, if you [~xytss123] could create pr for other branches as well, thx!","from":"developer"},{"body":"Cherry picked to branch-2.3 & branch-2.4 as well. Thanks [~xytss123] for the contributions.","from":"developer"},{"body":"Reopening to revert from branch-2.3.   The  new test added here is failing 100% on branch-2.3. See bottom of [https://ci-hadoop.apache.org/view/HBase/job/HBase/job/HBase-Find-Flaky-Tests/job/branch-2.3/lastSuccessfulBuild/artifact/output/dashboard.html] Thanks.","from":"developer"},{"body":"Removed 2.3.6 as fix version after revert.","from":"developer"},{"body":"I run the tests with branch-2.3 locally. And got the following message :\r\n\r\n2021-07-27 15:14:46,016 DEBUG [master/192.168.1.127:0:becomeActiveMaster] asyncfs.FanOutOneBlockAsyncDFSOutputHelper(265): ClientProtocol::create wrong number of arguments, should be hadoop 3.2 or below\r\n2021-07-27 15:14:46,016 DEBUG [master/192.168.1.127:0:becomeActiveMaster] asyncfs.FanOutOneBlockAsyncDFSOutputHelper(271): ClientProtocol::create wrong number of arguments, should be hadoop 2.x\r\n2021-07-27 15:14:46,029 DEBUG [master/192.168.1.127:0:becomeActiveMaster] asyncfs.FanOutOneBlockAsyncDFSOutputHelper(280): can not find SHOULD_REPLICATE flag, should be hadoop 2.x\r\njava.lang.IllegalArgumentException: No enum constant org.apache.hadoop.fs.CreateFlag.SHOULD_REPLICATE\r\n\r\nI also rerun the tests under branch-2 and it passed. \r\nIt seems the tests failure is due to a version problem of hadoop. [~stack]","from":"developer"},{"body":"I don't think the IllegalArgumentException causes the test to fail. The version of hadoop in hbase 2.3 is old but would be odd needing to update it to make a unit test pass? Seems like there is something about the 2.3 hbase context that makes the test fail. I didn't spend much time on it. [~xytss123]  Thanks.","from":"developer"}],"created":"2021-06-14T11:32:02.000+0000","description":"AccessController postIncrementBeforeWAL() and postAppendBeforeWAL() methods will rewrite the new cell's tags by the old cell's. This will makes the other kinds of tag in new cell invisible (such as TTL tag) after this. As in Increment and Append operations, the new cell has already catch forward all tags of the old cell and TTL tag from mutation operation, here in AccessController we do not need to rewrite the tags once again. Also, the TTL tag of newCell will be invisible in the new created cell. Actually, in Increment and Append operations, the newCell has already copied all tags of the oldCell. So the oldCell is useless here.\r\n\r\n{code:java}\r\nprivate Cell createNewCellWithTags(Mutation mutation, Cell oldCell, Cell newCell) {\r\n // Collect any ACLs from the old cell\r\n List tags = Lists.newArrayList();\r\n List aclTags = Lists.newArrayList();\r\n ListMultimap perms = ArrayListMultimap.create();\r\n if (oldCell != null) {\r\n Iterator tagIterator = PrivateCellUtil.tagsIterator(oldCell);\r\n while (tagIterator.hasNext()) {\r\n Tag tag = tagIterator.next();\r\n if (tag.getType() != PermissionStorage.ACL_TAG_TYPE) {\r\n // Not an ACL tag, just carry it through\r\n if (LOG.isTraceEnabled()) {\r\n LOG.trace(\"Carrying forward tag from \" + oldCell + \": type \" + tag.getType()\r\n + \" length \" + tag.getValueLength());\r\n }\r\n tags.add(tag);\r\n } else {\r\n aclTags.add(tag);\r\n }\r\n }\r\n }\r\n\r\n // Do we have an ACL on the operation?\r\n byte[] aclBytes = mutation.getACL();\r\n if (aclBytes != null) {\r\n // Yes, use it\r\n tags.add(new ArrayBackedTag(PermissionStorage.ACL_TAG_TYPE, aclBytes));\r\n } else {\r\n // No, use what we carried forward\r\n if (perms != null) {\r\n // TODO: If we collected ACLs from more than one tag we may have a\r\n // List of size > 1, this can be collapsed into a single\r\n // Permission\r\n if (LOG.isTraceEnabled()) {\r\n LOG.trace(\"Carrying forward ACLs from \" + oldCell + \": \" + perms);\r\n }\r\n tags.addAll(aclTags);\r\n }\r\n }\r\n\r\n // If we have no tags to add, just return\r\n if (tags.isEmpty()) {\r\n return newCell;\r\n }\r\n // Here the new cell's tags will be in visible.\r\n return PrivateCellUtil.createCell(newCell, tags);\r\n }\r\n{code}\r\n","issue_id":"13383708","key":"HBASE-26001","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-07-26T18:34:55.000+0000","role":"fixed_distractor","summary":"When turn on access control, the cell level TTL of Increment and Append operations is invalid."} {"case_id":"13386464","cluster":"DISTRACTOR-HBASE-26036","comments":[{"body":"Excellent. How did you debug and find the skipped close the leak? Was the BYTEBUFF_ALLOCATOR_CLASS addition how you debugged?","created":"2021-06-29T16:32:26.881+0000"},{"body":"Hi,[~stack],the scanner closed in get() is a little conspicuous when I looking through the codes according to the coredump logs. But the pressure in UT is not big enough  to verify the suspect by letting the program auto rewrite the released DBBs. I add a test BYTEBUFF_ALLOCATOR_CLASS to overwrite the DBBs right after released, results shows that checkAndMutate has read dirty data.","created":"2021-06-30T06:01:31.012+0000"},{"body":"[~Xiaolin Ha] may we see what your BYTEBUFF_ALLOCATOR_CLASS replacement to see what it does? And mind saying more what you mean by 'auto rewrite the released DBBs'? Thanks.","created":"2021-06-30T14:50:03.662+0000"},{"body":"Hi, [~stack], thanks for this question. The original purpose of BYTEBUFF_ALLOCATOR_CLASS is to simplify rewrite of released DBBs, without it the problem can also be reproduced, I have made another UT in [https://github.com/apache/hbase/pull/3449,] please take a look.\r\n\r\nIn the ByteBuffAllocator, it uses DBBs in the pool priority over allocating new ones. Once a DBB is released and deallocated, it will be put to the ByteBuff pool. \r\n\r\nThe test case process is that,\r\n # checkAndMutate performs getting cells, the ByteBuffAllocator allocate a new DBB for the read table block, but the HRegion.get() releases the DBB inner the method before the cell value is checked;\r\n # after get, wait a little while, another thread will deallocate the DBB it used and put back to the pool;\r\n # some other get/scan calls use the pooled DBBs to read other table blocks;\r\n # the checkAndMutate get the cell timestamp and check the cell value, but the content of the cell referred DBB was changed;\r\n\r\nI think the customized BYTEBUFF_ALLOCATOR_CLASS is very useful for debugging such problems, it can help to reduce some rewrite works of the program and simplify the test codes. Thanks.\r\n\r\n ","created":"2021-07-01T15:22:26.800+0000"},{"body":"HBASE-25187 seems not related to this issue, should be HBASE-25981?  [~Xiaolin Ha]","created":"2021-07-02T02:00:53.363+0000"},{"body":"Sweet. Thanks [~Xiaolin Ha] for the explanation.\r\n\r\nI tried [https://github.com/apache/hbase/pull/3449  |https://github.com/apache/hbase/pull/3449,]Is it supposed to fail? It doesn't for me on linux/mac hbase-2.3. I hacked [#3436|https://github.com/apache/hbase/pull/3436] pr so only the test and support for alternate BYTEBUFF_ALLOCATOR_CLASS and it fails.  Nice.","created":"2021-07-02T02:47:34.205+0000"},{"body":"Hi, [~stack], the PR #3449 has reproduced the problem at [https://ci-hadoop.apache.org/job/HBase/job/HBase-PreCommit-GitHub-PR/job/PR-3449/1/testReport/]","created":"2021-07-02T03:07:14.753+0000"},{"body":"Basically the internal calls should reach to Region level for doing ops. Like Reads in case of checkAndXXX op etc. Even when CPs do ops All has to be (And will be ideally) at Region level not at RS level which will try to serialize the results (for over wire transfer)","created":"2021-07-02T03:08:56.347+0000"},{"body":"Hi, [~filtertip] , thanks for reminding. This is not a same problem as in HBASE-25981, that is a monitor life circle problem.\r\n\r\nI referred HBASE-25187 because it used the cached row length and key length for the cell, which makes the program read dirty data, without this cache, the regionserver JVM will crash.","created":"2021-07-02T03:15:03.258+0000"},{"body":"[~Xiaolin Ha]. Ok, get it.","created":"2021-07-02T05:09:26.572+0000"},{"body":"[~anoop.hbase] Thanks for the comments, I agree with you. ","created":"2021-07-02T08:14:16.450+0000"},{"body":"{quote}\r\ncheckAndMutate performs getting cells, the ByteBuffAllocator allocate a new DBB for the read table block, but the HRegion.get() releases the DBB inner the method before the cell value is checked;\r\n{quote}\r\n\r\nBasically I think this is the problem we need to fix. This means the get method itself returns broken cells to callers.\r\nDid you find out which issues broke the HRegion.get method?\r\n\r\nThanks.","created":"2021-07-03T03:36:58.064+0000"},{"body":"Hi, [~zhangduo], we can see that inner the HReigon.getInternal(), the HRegion.get() itself closed the scanner right after scanning, it is a normal action,\r\n{code:java}\r\nScan scan = new Scan(get);\r\nif (scan.getLoadColumnFamiliesOnDemandValue() == null) {\r\n scan.setLoadColumnFamiliesOnDemand(isLoadingCfsOnDemandDefault());\r\n}\r\ntry (RegionScanner scanner = getScanner(scan, null, nonceGroup, nonce)) {\r\n scanner.next(results);\r\n}\r\n{code}\r\nbut it doesn't think about the outer methods, which won't know the DBBs of results are released.\r\n\r\nTo fix HRegion.get(), I thought there are two ways,\r\n # add the scanner above to the close callback of the RPC call, but it is a little strange for the region to get the caller, and there is already get(Get get, HRegion region, RegionScannersCloseCallBack closeCallBack, RpcCallContext context) in the RSRpcServices. What's more, for the operations like checkAndPut, the DBBs that get results used should be released as early as possibly, it need not to keep until the end of RPC call.\r\n # before return the results, copy them into heap. But this may brings larger heap pressure.\r\n\r\nLooking through the places that HRegion.get() is used, mostly of them are for test purposes, except where I changed in the PR.\r\n\r\nDo you have any advise for fixing this issue?\r\n\r\nThanks.\r\n\r\n ","created":"2021-07-05T03:53:04.278+0000"},{"body":"In case of get calls from Client layer, the RSRpcServices make use of scan APIs in Region and serialize that result cells (to CellBlock or Result PB) and then only close the scanner. So the result cells are copied here. But in case of get() API in HRegion, this is not happening and that is the issue. I am of the opinion that we should copy the result and then close the scanner. This API is exposed to CPs. Any CP can face this kind of issue. +1 to Duo's idea of fixing this.\r\n\r\nFor checkAndXXX methods, we can start using the getScanner way and use the data and close scanner then only. (As in patch) That can avoid heap copy and heap pressure. These are our internal impl method and we can change it. (Pls keep some fat notes why we dont use HRegion#get() API when such a straight forward one is available with us)","created":"2021-07-05T07:42:30.556+0000"},{"body":"[~anoop.hbase] thanks, I'll change the patch as your advice.","created":"2021-07-05T09:43:03.271+0000"},{"body":"I was going to ask how we guarantee integrity between return from HRegion#get and the copy on to the wire to send back to the client over RPC and we can't per [~anoop.hbase] 's helpful explanation. As suggested here and in PR, needs big fat WARNING.  Good stuff.","created":"2021-07-05T18:10:33.457+0000"},{"body":"Merged to branch-2 and master.\r\nThanks [~stack] for reviewing.","created":"2021-07-14T08:44:12.182+0000"},{"body":"I backported this beautiful fix to 2.4.5. It wouldn't go back to branch-2.3 cleanly, unfortunately.","created":"2021-07-14T16:41:33.811+0000"}],"conversations":[{"body":"Before HBASE-25187, we found there are regionserver JVM crashing problems on our production clusters, the coredump infos are as follows,\r\n{code:java}\r\nStack: [0x00007f621ba8d000,0x00007f621bb8e000], sp=0x00007f621bb8c0e0, free space=1020k\r\nNative frames: (J=compiled Java code, j=interpreted, Vv=VM code, C=native code)\r\nJ 10829 C2 org.apache.hadoop.hbase.ByteBufferKeyValue.getTimestamp()J (9 bytes) @ 0x00007f6a5ee11b2d [0x00007f6a5ee11ae0+0x4d]\r\nJ 22844 C2 org.apache.hadoop.hbase.regionserver.HRegion.doCheckAndRowMutate([B[B[BLorg/apache/hadoop/hbase/filter/CompareFilter$CompareOp;Lorg/apache/hadoop/hbase/filter/ByteArrayComparable;Lorg/apache/hadoop/hbase/client/RowMutations;Lorg/apache/hadoop/hbase/client/Mutation;Z)Z (540 bytes) @ 0x00007f6a60bed144 [0x00007f6a60beb320+0x1e24]\r\nJ 17972 C2 org.apache.hadoop.hbase.regionserver.RSRpcServices.checkAndRowMutate(Lorg/apache/hadoop/hbase/regionserver/Region;Ljava/util/List;Lorg/apache/hadoop/hbase/CellScanner;[B[B[BLorg/apache/hadoop/hbase/filter/CompareFilter$CompareOp;Lorg/apache/hadoop/hbase/filter/ByteArrayComparable;Lorg/apache/hadoop/hbase/shaded/protobuf/generated/ClientProtos$RegionActionResult$Builder;)Z (312 bytes) @ 0x00007f6a5f4a7ed0 [0x00007f6a5f4a6f40+0xf90]\r\nJ 26197 C2 org.apache.hadoop.hbase.regionserver.RSRpcServices.multi(Lorg/apache/hbase/thirdparty/com/google/protobuf/RpcController;Lorg/apache/hadoop/hbase/shaded/protobuf/generated/ClientProtos$MultiRequest;)Lorg/apache/hadoop/hbase/shaded/protobuf/generated/ClientProtos$MultiResponse; (644 bytes) @ 0x00007f6a61538b0c [0x00007f6a61537940+0x11cc]\r\nJ 26332 C2 org.apache.hadoop.hbase.ipc.RpcServer.call(Lorg/apache/hadoop/hbase/ipc/RpcCall;Lorg/apache/hadoop/hbase/monitoring/MonitoredRPCHandler;)Lorg/apache/hadoop/hbase/util/Pair; (566 bytes) @ 0x00007f6a615e8228 [0x00007f6a615e79c0+0x868]\r\nJ 20563 C2 org.apache.hadoop.hbase.ipc.CallRunner.run()V (1196 bytes) @ 0x00007f6a60711a4c [0x00007f6a60711000+0xa4c]\r\nJ 19656% C2 org.apache.hadoop.hbase.ipc.RpcExecutor.consumerLoop(Ljava/util/concurrent/BlockingQueue;Ljava/util/concurrent/atomic/AtomicInteger;)V (338 bytes) @ 0x00007f6a6039a414 [0x00007f6a6039a320+0xf4]\r\nj org.apache.hadoop.hbase.ipc.RpcExecutor$1.run()V+24\r\nj java.lang.Thread.run()V+11\r\nv ~StubRoutines::call_stub\r\n{code}\r\nI have made a UT to reproduce this error, it can occur 100%。\r\n\r\nAfter HBASE-25187,the check result of the checkAndMutate will be false, because it read wrong/dirty data from the released ByteBuff.","from":"reporter","subject":"DBB released too early and dirty data for some operations"},{"body":"Excellent. How did you debug and find the skipped close the leak? Was the BYTEBUFF_ALLOCATOR_CLASS addition how you debugged?","from":"developer"},{"body":"Hi,[~stack],the scanner closed in get() is a little conspicuous when I looking through the codes according to the coredump logs. But the pressure in UT is not big enough  to verify the suspect by letting the program auto rewrite the released DBBs. I add a test BYTEBUFF_ALLOCATOR_CLASS to overwrite the DBBs right after released, results shows that checkAndMutate has read dirty data.","from":"developer"},{"body":"[~Xiaolin Ha] may we see what your BYTEBUFF_ALLOCATOR_CLASS replacement to see what it does? And mind saying more what you mean by 'auto rewrite the released DBBs'? Thanks.","from":"developer"},{"body":"Hi, [~stack], thanks for this question. The original purpose of BYTEBUFF_ALLOCATOR_CLASS is to simplify rewrite of released DBBs, without it the problem can also be reproduced, I have made another UT in [https://github.com/apache/hbase/pull/3449,] please take a look.\r\n\r\nIn the ByteBuffAllocator, it uses DBBs in the pool priority over allocating new ones. Once a DBB is released and deallocated, it will be put to the ByteBuff pool. \r\n\r\nThe test case process is that,\r\n # checkAndMutate performs getting cells, the ByteBuffAllocator allocate a new DBB for the read table block, but the HRegion.get() releases the DBB inner the method before the cell value is checked;\r\n # after get, wait a little while, another thread will deallocate the DBB it used and put back to the pool;\r\n # some other get/scan calls use the pooled DBBs to read other table blocks;\r\n # the checkAndMutate get the cell timestamp and check the cell value, but the content of the cell referred DBB was changed;\r\n\r\nI think the customized BYTEBUFF_ALLOCATOR_CLASS is very useful for debugging such problems, it can help to reduce some rewrite works of the program and simplify the test codes. Thanks.\r\n\r\n ","from":"developer"},{"body":"HBASE-25187 seems not related to this issue, should be HBASE-25981?  [~Xiaolin Ha]","from":"developer"},{"body":"Sweet. Thanks [~Xiaolin Ha] for the explanation.\r\n\r\nI tried [https://github.com/apache/hbase/pull/3449  |https://github.com/apache/hbase/pull/3449,]Is it supposed to fail? It doesn't for me on linux/mac hbase-2.3. I hacked [#3436|https://github.com/apache/hbase/pull/3436] pr so only the test and support for alternate BYTEBUFF_ALLOCATOR_CLASS and it fails.  Nice.","from":"developer"},{"body":"Hi, [~stack], the PR #3449 has reproduced the problem at [https://ci-hadoop.apache.org/job/HBase/job/HBase-PreCommit-GitHub-PR/job/PR-3449/1/testReport/]","from":"developer"},{"body":"Basically the internal calls should reach to Region level for doing ops. Like Reads in case of checkAndXXX op etc. Even when CPs do ops All has to be (And will be ideally) at Region level not at RS level which will try to serialize the results (for over wire transfer)","from":"developer"},{"body":"Hi, [~filtertip] , thanks for reminding. This is not a same problem as in HBASE-25981, that is a monitor life circle problem.\r\n\r\nI referred HBASE-25187 because it used the cached row length and key length for the cell, which makes the program read dirty data, without this cache, the regionserver JVM will crash.","from":"developer"},{"body":"[~Xiaolin Ha]. Ok, get it.","from":"developer"},{"body":"[~anoop.hbase] Thanks for the comments, I agree with you. ","from":"developer"},{"body":"{quote}\r\ncheckAndMutate performs getting cells, the ByteBuffAllocator allocate a new DBB for the read table block, but the HRegion.get() releases the DBB inner the method before the cell value is checked;\r\n{quote}\r\n\r\nBasically I think this is the problem we need to fix. This means the get method itself returns broken cells to callers.\r\nDid you find out which issues broke the HRegion.get method?\r\n\r\nThanks.","from":"developer"},{"body":"Hi, [~zhangduo], we can see that inner the HReigon.getInternal(), the HRegion.get() itself closed the scanner right after scanning, it is a normal action,\r\n{code:java}\r\nScan scan = new Scan(get);\r\nif (scan.getLoadColumnFamiliesOnDemandValue() == null) {\r\n scan.setLoadColumnFamiliesOnDemand(isLoadingCfsOnDemandDefault());\r\n}\r\ntry (RegionScanner scanner = getScanner(scan, null, nonceGroup, nonce)) {\r\n scanner.next(results);\r\n}\r\n{code}\r\nbut it doesn't think about the outer methods, which won't know the DBBs of results are released.\r\n\r\nTo fix HRegion.get(), I thought there are two ways,\r\n # add the scanner above to the close callback of the RPC call, but it is a little strange for the region to get the caller, and there is already get(Get get, HRegion region, RegionScannersCloseCallBack closeCallBack, RpcCallContext context) in the RSRpcServices. What's more, for the operations like checkAndPut, the DBBs that get results used should be released as early as possibly, it need not to keep until the end of RPC call.\r\n # before return the results, copy them into heap. But this may brings larger heap pressure.\r\n\r\nLooking through the places that HRegion.get() is used, mostly of them are for test purposes, except where I changed in the PR.\r\n\r\nDo you have any advise for fixing this issue?\r\n\r\nThanks.\r\n\r\n ","from":"developer"},{"body":"In case of get calls from Client layer, the RSRpcServices make use of scan APIs in Region and serialize that result cells (to CellBlock or Result PB) and then only close the scanner. So the result cells are copied here. But in case of get() API in HRegion, this is not happening and that is the issue. I am of the opinion that we should copy the result and then close the scanner. This API is exposed to CPs. Any CP can face this kind of issue. +1 to Duo's idea of fixing this.\r\n\r\nFor checkAndXXX methods, we can start using the getScanner way and use the data and close scanner then only. (As in patch) That can avoid heap copy and heap pressure. These are our internal impl method and we can change it. (Pls keep some fat notes why we dont use HRegion#get() API when such a straight forward one is available with us)","from":"developer"},{"body":"[~anoop.hbase] thanks, I'll change the patch as your advice.","from":"developer"},{"body":"I was going to ask how we guarantee integrity between return from HRegion#get and the copy on to the wire to send back to the client over RPC and we can't per [~anoop.hbase] 's helpful explanation. As suggested here and in PR, needs big fat WARNING.  Good stuff.","from":"developer"},{"body":"Merged to branch-2 and master.\r\nThanks [~stack] for reviewing.","from":"developer"},{"body":"I backported this beautiful fix to 2.4.5. It wouldn't go back to branch-2.3 cleanly, unfortunately.","from":"developer"}],"created":"2021-06-29T08:31:14.000+0000","description":"Before HBASE-25187, we found there are regionserver JVM crashing problems on our production clusters, the coredump infos are as follows,\r\n{code:java}\r\nStack: [0x00007f621ba8d000,0x00007f621bb8e000], sp=0x00007f621bb8c0e0, free space=1020k\r\nNative frames: (J=compiled Java code, j=interpreted, Vv=VM code, C=native code)\r\nJ 10829 C2 org.apache.hadoop.hbase.ByteBufferKeyValue.getTimestamp()J (9 bytes) @ 0x00007f6a5ee11b2d [0x00007f6a5ee11ae0+0x4d]\r\nJ 22844 C2 org.apache.hadoop.hbase.regionserver.HRegion.doCheckAndRowMutate([B[B[BLorg/apache/hadoop/hbase/filter/CompareFilter$CompareOp;Lorg/apache/hadoop/hbase/filter/ByteArrayComparable;Lorg/apache/hadoop/hbase/client/RowMutations;Lorg/apache/hadoop/hbase/client/Mutation;Z)Z (540 bytes) @ 0x00007f6a60bed144 [0x00007f6a60beb320+0x1e24]\r\nJ 17972 C2 org.apache.hadoop.hbase.regionserver.RSRpcServices.checkAndRowMutate(Lorg/apache/hadoop/hbase/regionserver/Region;Ljava/util/List;Lorg/apache/hadoop/hbase/CellScanner;[B[B[BLorg/apache/hadoop/hbase/filter/CompareFilter$CompareOp;Lorg/apache/hadoop/hbase/filter/ByteArrayComparable;Lorg/apache/hadoop/hbase/shaded/protobuf/generated/ClientProtos$RegionActionResult$Builder;)Z (312 bytes) @ 0x00007f6a5f4a7ed0 [0x00007f6a5f4a6f40+0xf90]\r\nJ 26197 C2 org.apache.hadoop.hbase.regionserver.RSRpcServices.multi(Lorg/apache/hbase/thirdparty/com/google/protobuf/RpcController;Lorg/apache/hadoop/hbase/shaded/protobuf/generated/ClientProtos$MultiRequest;)Lorg/apache/hadoop/hbase/shaded/protobuf/generated/ClientProtos$MultiResponse; (644 bytes) @ 0x00007f6a61538b0c [0x00007f6a61537940+0x11cc]\r\nJ 26332 C2 org.apache.hadoop.hbase.ipc.RpcServer.call(Lorg/apache/hadoop/hbase/ipc/RpcCall;Lorg/apache/hadoop/hbase/monitoring/MonitoredRPCHandler;)Lorg/apache/hadoop/hbase/util/Pair; (566 bytes) @ 0x00007f6a615e8228 [0x00007f6a615e79c0+0x868]\r\nJ 20563 C2 org.apache.hadoop.hbase.ipc.CallRunner.run()V (1196 bytes) @ 0x00007f6a60711a4c [0x00007f6a60711000+0xa4c]\r\nJ 19656% C2 org.apache.hadoop.hbase.ipc.RpcExecutor.consumerLoop(Ljava/util/concurrent/BlockingQueue;Ljava/util/concurrent/atomic/AtomicInteger;)V (338 bytes) @ 0x00007f6a6039a414 [0x00007f6a6039a320+0xf4]\r\nj org.apache.hadoop.hbase.ipc.RpcExecutor$1.run()V+24\r\nj java.lang.Thread.run()V+11\r\nv ~StubRoutines::call_stub\r\n{code}\r\nI have made a UT to reproduce this error, it can occur 100%。\r\n\r\nAfter HBASE-25187,the check result of the checkAndMutate will be false, because it read wrong/dirty data from the released ByteBuff.","issue_id":"13386464","key":"HBASE-26036","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-07-14T08:44:12.000+0000","role":"fixed_distractor","summary":"DBB released too early and dirty data for some operations"} {"case_id":"13387489","cluster":"DISTRACTOR-HBASE-26063","comments":[{"body":"Merged to master.\r\n\r\nThanks [~niuyulin] for reviewing.","created":"2021-07-03T13:46:41.950+0000"},{"body":"bq. The other is because of opentelemetry. It packs some jar in the fat agent jar, and the name of the directory for holding the extracted jar content still ends .jar, which makes the current script think it is a jar and cause problems.\r\n\r\nRan into this issue with 2.5, backported to branch-2 and branch-2.5. Release scripts use the api check tool version committed in the branch. ","created":"2022-06-01T01:28:57.356+0000"}],"conversations":[{"body":"There are basically two problems.\r\n\r\nOne is the findbugs-maven-plugin, on rel/2.0.0, its version is 3.0.0, and it will lead the below error\r\n{noformat}\r\n[ERROR] Failed to execute goal org.codehaus.mojo:findbugs-maven-plugin:3.0.0:findbugs (default) on project hbase: Unable to parse configuration of mojo org.codehaus.mojo:findbugs-maven-plugin:3.0.0:findbugs for parameter pluginArtifacts: Cannot assign configuration entry 'pluginArtifacts' with value '${plugin.artifacts}' of type java.util.Collections.UnmodifiableRandomAccessList to property of type java.util.ArrayList -> [Help 1]\r\n{noformat}\r\n\r\nThe other is because of opentelemetry. It packs some jar in the fat agent jar, and the name of the directory for holding the extracted jar content still ends .jar, which makes the current script think it is a jar and cause problems.\r\n{noformat}\r\njava.io.FileNotFoundException: /home/zhangduo/hbase/code/target/compat-check/dst/hbase-assembly/target/dependency/opentelemetry-javaagent-1.0.1-all-jar/META-INF/licenses/okhttp-3.14.9.jar (是一个目录)\r\n at java.io.FileInputStream.open0(Native Method)\r\n at java.io.FileInputStream.open(FileInputStream.java:195)\r\n at java.io.FileInputStream.(FileInputStream.java:138)\r\n at java.io.FileInputStream.(FileInputStream.java:93)\r\n at sun.tools.jar.Main.run(Main.java:307)\r\n at sun.tools.jar.Main.main(Main.java:1288)\r\nERROR: can't extract '/home/zhangduo/hbase/code/target/compat-check/dst/hbase-assembly/target/dependency/opentelemetry-javaagent-1.0.1-all-jar/META-INF/licenses/okhttp-3.14.9.jar'\r\nINFO:root:Results: {}\r\nTraceback (most recent call last):\r\n File \"./checkcompatibility.py\", line 539, in \r\n main()\r\n File \"./checkcompatibility.py\", line 535, in main\r\n args.compare_warnings))\r\n File \"./checkcompatibility.py\", line 233, in compare_results\r\n if compare_tool_results_count(tool_results, check, issue_type, known_count)]\r\n File \"./checkcompatibility.py\", line 252, in compare_tool_results_count\r\n return tool_results[check][issue_type] > known_count\r\nKeyError: 'binary'\r\n{noformat}\r\n\r\n","from":"reporter","subject":"The current checkcompatibility.py script can not compare master and rel/2.0.0"},{"body":"Merged to master.\r\n\r\nThanks [~niuyulin] for reviewing.","from":"developer"},{"body":"bq. The other is because of opentelemetry. It packs some jar in the fat agent jar, and the name of the directory for holding the extracted jar content still ends .jar, which makes the current script think it is a jar and cause problems.\r\n\r\nRan into this issue with 2.5, backported to branch-2 and branch-2.5. Release scripts use the api check tool version committed in the branch. ","from":"developer"}],"created":"2021-07-03T10:33:35.000+0000","description":"There are basically two problems.\r\n\r\nOne is the findbugs-maven-plugin, on rel/2.0.0, its version is 3.0.0, and it will lead the below error\r\n{noformat}\r\n[ERROR] Failed to execute goal org.codehaus.mojo:findbugs-maven-plugin:3.0.0:findbugs (default) on project hbase: Unable to parse configuration of mojo org.codehaus.mojo:findbugs-maven-plugin:3.0.0:findbugs for parameter pluginArtifacts: Cannot assign configuration entry 'pluginArtifacts' with value '${plugin.artifacts}' of type java.util.Collections.UnmodifiableRandomAccessList to property of type java.util.ArrayList -> [Help 1]\r\n{noformat}\r\n\r\nThe other is because of opentelemetry. It packs some jar in the fat agent jar, and the name of the directory for holding the extracted jar content still ends .jar, which makes the current script think it is a jar and cause problems.\r\n{noformat}\r\njava.io.FileNotFoundException: /home/zhangduo/hbase/code/target/compat-check/dst/hbase-assembly/target/dependency/opentelemetry-javaagent-1.0.1-all-jar/META-INF/licenses/okhttp-3.14.9.jar (是一个目录)\r\n at java.io.FileInputStream.open0(Native Method)\r\n at java.io.FileInputStream.open(FileInputStream.java:195)\r\n at java.io.FileInputStream.(FileInputStream.java:138)\r\n at java.io.FileInputStream.(FileInputStream.java:93)\r\n at sun.tools.jar.Main.run(Main.java:307)\r\n at sun.tools.jar.Main.main(Main.java:1288)\r\nERROR: can't extract '/home/zhangduo/hbase/code/target/compat-check/dst/hbase-assembly/target/dependency/opentelemetry-javaagent-1.0.1-all-jar/META-INF/licenses/okhttp-3.14.9.jar'\r\nINFO:root:Results: {}\r\nTraceback (most recent call last):\r\n File \"./checkcompatibility.py\", line 539, in \r\n main()\r\n File \"./checkcompatibility.py\", line 535, in main\r\n args.compare_warnings))\r\n File \"./checkcompatibility.py\", line 233, in compare_results\r\n if compare_tool_results_count(tool_results, check, issue_type, known_count)]\r\n File \"./checkcompatibility.py\", line 252, in compare_tool_results_count\r\n return tool_results[check][issue_type] > known_count\r\nKeyError: 'binary'\r\n{noformat}\r\n\r\n","issue_id":"13387489","key":"HBASE-26063","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-07-03T13:46:41.000+0000","role":"fixed_distractor","summary":"The current checkcompatibility.py script can not compare master and rel/2.0.0"} {"case_id":"13392688","cluster":"DISTRACTOR-HBASE-26155","comments":[{"body":"Merged to master and branch2.3+, thanks [~stack]  [~zhangduo] for reviewing.","created":"2021-08-12T09:24:48.685+0000"}],"conversations":[{"body":"There are scanner close caused regionserver JVM coredump problems on our production clusters.\r\n{code:java}\r\nStack: [0x00007fca4b0cc000,0x00007fca4b1cd000], sp=0x00007fca4b1cb0d8, free space=1020k\r\nNative frames: (J=compiled Java code, j=interpreted, Vv=VM code, C=native code)\r\nV [libjvm.so+0x7fd314]\r\nJ 2810 sun.misc.Unsafe.copyMemory(Ljava/lang/Object;JLjava/lang/Object;JJ)V (0 bytes) @ 0x00007fdae55a9e61 [0x00007fdae55a9d80+0xe1]\r\nj org.apache.hadoop.hbase.util.UnsafeAccess.unsafeCopy(Ljava/lang/Object;JLjava/lang/Object;JJ)V+36\r\nj org.apache.hadoop.hbase.util.UnsafeAccess.copy(Ljava/nio/ByteBuffer;I[BII)V+69\r\nj org.apache.hadoop.hbase.util.ByteBufferUtils.copyFromBufferToArray([BLjava/nio/ByteBuffer;III)V+39\r\nj org.apache.hadoop.hbase.CellUtil.copyQualifierTo(Lorg/apache/hadoop/hbase/Cell;[BI)I+31\r\nj org.apache.hadoop.hbase.KeyValueUtil.appendKeyTo(Lorg/apache/hadoop/hbase/Cell;[BI)I+43\r\nJ 14724 C2 org.apache.hadoop.hbase.regionserver.StoreScanner.shipped()V (51 bytes) @ 0x00007fdae6a298d0 [0x00007fdae6a29780+0x150]\r\nJ 21387 C2 org.apache.hadoop.hbase.regionserver.RSRpcServices$RegionScannerShippedCallBack.run()V (53 bytes) @ 0x00007fdae622bab8 [0x00007fdae622acc0+0xdf8]\r\nJ 26353 C2 org.apache.hadoop.hbase.ipc.ServerCall.setResponse(Lorg/apache/hbase/thirdparty/com/google/protobuf/Message;Lorg/apache/hadoop/hbase/CellScanner;Ljava/lang/Throwable;Ljava/lang/String;)V (384 bytes) @ 0x00007fdae7f139d8 [0x00007fdae7f12980+0x1058]\r\nJ 26226 C2 org.apache.hadoop.hbase.ipc.CallRunner.run()V (1554 bytes) @ 0x00007fdae959f68c [0x00007fdae959e400+0x128c]\r\nJ 19598% C2 org.apache.hadoop.hbase.ipc.RpcExecutor.consumerLoop(Ljava/util/concurrent/BlockingQueue;Ljava/util/concurrent/atomic/AtomicInteger;)V (338 bytes) @ 0x00007fdae81c54d4 [0x00007fdae81c53e0+0xf4]\r\n{code}\r\nThere are also scan rpc errors when coredump happens at the handler,\r\n\r\n!scan-error.png|width=585,height=235!\r\n\r\nI found some clue in the logs, that some blocks may be replaced when its nextBlockOnDiskSize less than the newly one in the method \r\n\r\n \r\n{code:java}\r\npublic static boolean shouldReplaceExistingCacheBlock(BlockCache blockCache,\r\n BlockCacheKey cacheKey, Cacheable newBlock) {\r\n if (cacheKey.toString().indexOf(\".\") != -1) { // reference file\r\n LOG.warn(\"replace existing cached block, cache key is : \" + cacheKey);\r\n return true;\r\n }\r\n Cacheable existingBlock = blockCache.getBlock(cacheKey, false, false, false);\r\n if (existingBlock == null) {\r\n return true;\r\n }\r\n try {\r\n int comparison = BlockCacheUtil.validateBlockAddition(existingBlock, newBlock, cacheKey);\r\n if (comparison < 0) {\r\n LOG.warn(\"Cached block contents differ by nextBlockOnDiskSize, the new block has \"\r\n + \"nextBlockOnDiskSize set. Caching new block.\");\r\n return true;\r\n......{code}\r\n \r\n\r\nAnd the block will be replaced if it is not in the RAMCache but in the BucketCache.\r\n\r\nWhen using \r\n\r\n \r\n{code:java}\r\nprivate void putIntoBackingMap(BlockCacheKey key, BucketEntry bucketEntry) {\r\n BucketEntry previousEntry = backingMap.put(key, bucketEntry);\r\n if (previousEntry != null && previousEntry != bucketEntry) {\r\n ReentrantReadWriteLock lock = offsetLock.getLock(previousEntry.offset());\r\n lock.writeLock().lock();\r\n try {\r\n blockEvicted(key, previousEntry, false);\r\n } finally {\r\n lock.writeLock().unlock();\r\n }\r\n }\r\n}\r\n{code}\r\nto replace the old block, to avoid previous bucket entry mem leak, the previous bucket entry will be force released regardless of RPC references to it.\r\n\r\n \r\n{code:java}\r\nvoid blockEvicted(BlockCacheKey cacheKey, BucketEntry bucketEntry, boolean decrementBlockNumber) {\r\n bucketAllocator.freeBlock(bucketEntry.offset());\r\n realCacheSize.add(-1 * bucketEntry.getLength());\r\n blocksByHFile.remove(cacheKey);\r\n if (decrementBlockNumber) {\r\n this.blockNumber.decrement();\r\n }\r\n}\r\n{code}\r\nI used the check of RPC reference before replace bucket entry, and it works, no coredumps until now.\r\n\r\n \r\n\r\nThat is:\r\n{code:java}\r\npublic void cacheBlockWithWait(BlockCacheKey cacheKey, Cacheable cachedItem, boolean inMemory,\r\n boolean wait) {\r\n if (cacheEnabled) {\r\n if (backingMap.containsKey(cacheKey) || ramCache.containsKey(cacheKey)) {\r\n if (BlockCacheUtil.shouldReplaceExistingCacheBlock(this, cacheKey, cachedItem)) {\r\n BucketEntry bucketEntry = backingMap.get(cacheKey);\r\n if (bucketEntry != null && bucketEntry.isRpcRef()) {\r\n // avoid replace when there are RPC refs for the bucket entry in bucket cache\r\n return;\r\n }\r\n cacheBlockWithWaitInternal(cacheKey, cachedItem, inMemory, wait);\r\n }\r\n } else {\r\n cacheBlockWithWaitInternal(cacheKey, cachedItem, inMemory, wait);\r\n }\r\n }\r\n}\r\n{code}\r\n ","from":"reporter","subject":"JVM crash when scan"},{"body":"Merged to master and branch2.3+, thanks [~stack]  [~zhangduo] for reviewing.","from":"developer"}],"created":"2021-07-30T07:44:33.000+0000","description":"There are scanner close caused regionserver JVM coredump problems on our production clusters.\r\n{code:java}\r\nStack: [0x00007fca4b0cc000,0x00007fca4b1cd000], sp=0x00007fca4b1cb0d8, free space=1020k\r\nNative frames: (J=compiled Java code, j=interpreted, Vv=VM code, C=native code)\r\nV [libjvm.so+0x7fd314]\r\nJ 2810 sun.misc.Unsafe.copyMemory(Ljava/lang/Object;JLjava/lang/Object;JJ)V (0 bytes) @ 0x00007fdae55a9e61 [0x00007fdae55a9d80+0xe1]\r\nj org.apache.hadoop.hbase.util.UnsafeAccess.unsafeCopy(Ljava/lang/Object;JLjava/lang/Object;JJ)V+36\r\nj org.apache.hadoop.hbase.util.UnsafeAccess.copy(Ljava/nio/ByteBuffer;I[BII)V+69\r\nj org.apache.hadoop.hbase.util.ByteBufferUtils.copyFromBufferToArray([BLjava/nio/ByteBuffer;III)V+39\r\nj org.apache.hadoop.hbase.CellUtil.copyQualifierTo(Lorg/apache/hadoop/hbase/Cell;[BI)I+31\r\nj org.apache.hadoop.hbase.KeyValueUtil.appendKeyTo(Lorg/apache/hadoop/hbase/Cell;[BI)I+43\r\nJ 14724 C2 org.apache.hadoop.hbase.regionserver.StoreScanner.shipped()V (51 bytes) @ 0x00007fdae6a298d0 [0x00007fdae6a29780+0x150]\r\nJ 21387 C2 org.apache.hadoop.hbase.regionserver.RSRpcServices$RegionScannerShippedCallBack.run()V (53 bytes) @ 0x00007fdae622bab8 [0x00007fdae622acc0+0xdf8]\r\nJ 26353 C2 org.apache.hadoop.hbase.ipc.ServerCall.setResponse(Lorg/apache/hbase/thirdparty/com/google/protobuf/Message;Lorg/apache/hadoop/hbase/CellScanner;Ljava/lang/Throwable;Ljava/lang/String;)V (384 bytes) @ 0x00007fdae7f139d8 [0x00007fdae7f12980+0x1058]\r\nJ 26226 C2 org.apache.hadoop.hbase.ipc.CallRunner.run()V (1554 bytes) @ 0x00007fdae959f68c [0x00007fdae959e400+0x128c]\r\nJ 19598% C2 org.apache.hadoop.hbase.ipc.RpcExecutor.consumerLoop(Ljava/util/concurrent/BlockingQueue;Ljava/util/concurrent/atomic/AtomicInteger;)V (338 bytes) @ 0x00007fdae81c54d4 [0x00007fdae81c53e0+0xf4]\r\n{code}\r\nThere are also scan rpc errors when coredump happens at the handler,\r\n\r\n!scan-error.png|width=585,height=235!\r\n\r\nI found some clue in the logs, that some blocks may be replaced when its nextBlockOnDiskSize less than the newly one in the method \r\n\r\n \r\n{code:java}\r\npublic static boolean shouldReplaceExistingCacheBlock(BlockCache blockCache,\r\n BlockCacheKey cacheKey, Cacheable newBlock) {\r\n if (cacheKey.toString().indexOf(\".\") != -1) { // reference file\r\n LOG.warn(\"replace existing cached block, cache key is : \" + cacheKey);\r\n return true;\r\n }\r\n Cacheable existingBlock = blockCache.getBlock(cacheKey, false, false, false);\r\n if (existingBlock == null) {\r\n return true;\r\n }\r\n try {\r\n int comparison = BlockCacheUtil.validateBlockAddition(existingBlock, newBlock, cacheKey);\r\n if (comparison < 0) {\r\n LOG.warn(\"Cached block contents differ by nextBlockOnDiskSize, the new block has \"\r\n + \"nextBlockOnDiskSize set. Caching new block.\");\r\n return true;\r\n......{code}\r\n \r\n\r\nAnd the block will be replaced if it is not in the RAMCache but in the BucketCache.\r\n\r\nWhen using \r\n\r\n \r\n{code:java}\r\nprivate void putIntoBackingMap(BlockCacheKey key, BucketEntry bucketEntry) {\r\n BucketEntry previousEntry = backingMap.put(key, bucketEntry);\r\n if (previousEntry != null && previousEntry != bucketEntry) {\r\n ReentrantReadWriteLock lock = offsetLock.getLock(previousEntry.offset());\r\n lock.writeLock().lock();\r\n try {\r\n blockEvicted(key, previousEntry, false);\r\n } finally {\r\n lock.writeLock().unlock();\r\n }\r\n }\r\n}\r\n{code}\r\nto replace the old block, to avoid previous bucket entry mem leak, the previous bucket entry will be force released regardless of RPC references to it.\r\n\r\n \r\n{code:java}\r\nvoid blockEvicted(BlockCacheKey cacheKey, BucketEntry bucketEntry, boolean decrementBlockNumber) {\r\n bucketAllocator.freeBlock(bucketEntry.offset());\r\n realCacheSize.add(-1 * bucketEntry.getLength());\r\n blocksByHFile.remove(cacheKey);\r\n if (decrementBlockNumber) {\r\n this.blockNumber.decrement();\r\n }\r\n}\r\n{code}\r\nI used the check of RPC reference before replace bucket entry, and it works, no coredumps until now.\r\n\r\n \r\n\r\nThat is:\r\n{code:java}\r\npublic void cacheBlockWithWait(BlockCacheKey cacheKey, Cacheable cachedItem, boolean inMemory,\r\n boolean wait) {\r\n if (cacheEnabled) {\r\n if (backingMap.containsKey(cacheKey) || ramCache.containsKey(cacheKey)) {\r\n if (BlockCacheUtil.shouldReplaceExistingCacheBlock(this, cacheKey, cachedItem)) {\r\n BucketEntry bucketEntry = backingMap.get(cacheKey);\r\n if (bucketEntry != null && bucketEntry.isRpcRef()) {\r\n // avoid replace when there are RPC refs for the bucket entry in bucket cache\r\n return;\r\n }\r\n cacheBlockWithWaitInternal(cacheKey, cachedItem, inMemory, wait);\r\n }\r\n } else {\r\n cacheBlockWithWaitInternal(cacheKey, cachedItem, inMemory, wait);\r\n }\r\n }\r\n}\r\n{code}\r\n ","issue_id":"13392688","key":"HBASE-26155","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2021-08-12T09:24:48.000+0000","role":"fixed_distractor","summary":"JVM crash when scan"} {"case_id":"12380573","cluster":"DISTRACTOR-HBASE-269","comments":[{"body":"Generally looks good. A few comments:\n\n1. 401 is \"Unauthorized\", instead I would use 404 \"Not Found\" if a resource doesn't exist.\n2. The result of a successful PUT is 201 \"Created\".\n3. What do you think the representation of the resources will be?\n\nAnd in case you haven't seen it already, I found \"RESTful Web Services\" by by Leonard Richardson and Sam Ruby (http://www.oreilly.com/catalog/9780596529260/) really useful.","created":"2007-10-17T20:38:34.029+0000"},{"body":"Thanks for feedback Tom. \n\n401 rather than 404 was because I was watching TV at same time as writing the issue (smile). Regards 3. above, if I understand your question, there is no typing of hbase content -- as far as hbase is concerned, its all bytes -- so returned cell data should default mimetype binary/octet-stream. Later, might add switching resource format returned keyed off request headers. \n\nFor cluster, table, and column descriptors, default text/plain UTF-8 or perhaps text/xml if specified in request headers.\n\nThanks too for pointer to Leonard's book.","created":"2007-10-17T20:56:14.656+0000"},{"body":"Would the REST api support getting whole rows at a time? Perhaps off the GET /TABLE/ROW. The data could be packed into an XML document. \n\nI'd love to see this one get done, as it would get me the Ruby HBase client I've always wanted. ","created":"2007-11-08T01:03:29.103+0000"},{"body":"Have you seen how rows are dumped out in xhtml in hql.jsp? Would that work for you? You can't wrap binary data in XML so I don't think row representation as XML would be a general soln. If you explicitly request a cell, then we could pass back binary with (data length in http headers). \n\nI'm thinking that main difference between the REST implemenation and current UI HQL querying on the master would be that REST clients would go first to the master but would then be redirected to the regionserver hosting the row that is being read or updated (as opposed to all being done on master HQL'ing). I'm thinking REST clients would have to be able to deal with retry if data had been moved by the time they arrived at what had been the data server. The alternative would be something like having the regionserver redirect back to the master since it is arbiter of where everything is but I imagine that could get complicated quickly what with there often being some lag before master learns that data has moved and has had chance to orchestrate redeploy in the new location.","created":"2007-11-08T17:27:38.366+0000"},{"body":"HADOOP-2171 talks about creating a shell server which I think would be better than burdening the master with having to handle REST requests.","created":"2007-11-08T17:36:32.178+0000"},{"body":"I haven't seen the output of hql.jsp, but I'll try and track it down. You're right, you can't stick pure binary data in XML, but you can use Base64 encoded binary data. It's a little bigger, but it cuts down on the number of requests you would have to make, which I think will be the crucial bottleneck. Of course, single column requests for pure binary data should still be available.\n\nI agree with your idea of the REST requests redirecting around between the master and the region servers. That keeps the underlying architecture correct, just with a different face on it. \n\n","created":"2007-11-08T17:57:50.462+0000"},{"body":"HADOOP-2171 strikes me as a little odd. A client sends HQL to a server that parses the HQL to run HTable client operations against hbase. There are no load savings over running shell on client machine that I can see.\n\nI don't see a problem having the master handle REST requests. Master is generally lightly loaded. It will take a lot of traffic to make it break a sweat. The masters REST load would add the master fielding HTTP redirects -- a minor imposition. Should the REST load become burdensome, folks could put up an intermediary serve or take the load off the master by making their clients smarter doing HTable-like caching of data locations. ","created":"2007-11-08T18:31:16.866+0000"},{"body":"See the master UI at http://MASTER:PORT (default is localhost:60000). On the master homepage, there is a HQL link that does read-only queries against tables making best effort at outputting results in XHTML (Maybe this is good enough to get you going?).\n\nWe could Base64 the data. We could also tie a boat anchor to the server too as a means of slowing it down (smile). I suppose the RESTful way to do it is just allow clients say what they can accept in request headers (Don't know if you can stipulate xml with base64 encoded content via http request headers). If the server has support, do as the client asks.","created":"2007-11-08T18:59:15.571+0000"},{"body":"Well, either Base64 encoding the data and sending it all together takes longer than requesting each column individually, or it's the other way around. I'd be interested in seeing which way it really is. \n\nI think it is perfectly acceptable to do the column-only approach in the short term and evaluate the row oriented approach later.","created":"2007-11-08T20:55:33.560+0000"},{"body":"Bryan Duxbury added http://wiki.apache.org/lucene-hadoop/Hbase/HbaseRest.","created":"2007-11-15T06:10:46.434+0000"},{"body":"First cut at RESTful interface. Implements metainfo, gets, and scanners. Does not yet support put. Bunch of TODOs:\n\n+ Returning results multipart/related is crippled by lack of support in the container; jetty has a MultipartResponse class but can't set properly qualified Content-Type with boundary and start parameters... they get stripped. Need to figure out how to fix this (maybe jetty 6 does it better).\n+ Need to agree on timestamp format to use (ISO8601?)\n+ Need to fix HTable so it has table metadata; until then, you need to specify a column getting a scanner.\n\nHere's some samples run against a simple table named 'x' with column family 'x:' with following contents:\n\n{code}\nHbase> select * from x;\n+-------------------------+-------------------------+-------------------------+\n| Row | Column | Cell |\n+-------------------------+-------------------------+-------------------------+\n| x | x: | xyz |\n+-------------------------+-------------------------+-------------------------+\n| xyz | x:abc | abc |\n+-------------------------+-------------------------+-------------------------+\n| xyz | x:xyz | xyzxyz |\n+-------------------------+-------------------------+-------------------------+\n3 row(s) in set (0.19 sec)\n{code}\n\nIn below session I'm using curl. Doesn't have DELETE and I fake PUT with the -T option uploading a file:\n\n{code}\n$ curl http://localhost:60010/api/\n\n\n \nx\n
\n\n$ curl --header 'Accept: text/plain' http://localhost:60010/api/\nx\n\n$ curl --header 'Accept: text/plain' http://localhost:60010/api/x\nname: x, families: {x:={name: x, max versions: 3, compression: NONE, in memory: false, max length: 2147483647, bloom filter: none}}\n\n$ curl --header 'Accept: text/xml' http://localhost:60010/api/x\n\n\n \nx\n \n \n \n \nx:\n \n \nNONE\n \n \nNONE\n \n \n3\n \n \n2147483647\n \n \n \n
\n\n$ curl --header 'Accept: text/xml' http://localhost:60010/api/x/regions\n\n\n \n\n# only one region and its start key is null, the default table start key\n\n$ curl --verbose --header 'Accept: text/xml' http://localhost:60010/api/x/scanner\n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> GET /api/x/scanner HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: text/xml\n\n< HTTP/1.1 404 No+handler\n< Date: Mon, 19 Nov 2007 23:22:23 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Content-Type: text/html\n< Content-Length: 1228\n\n\nError 404 No handler\n\n\n...\n\n# Fails because currently you must specify a column name (To be fixed)\n\n$ curl --verbose --header 'Accept: text/xml' -T /tmp/diff.txt http://localhost:60010/api/x/scanner?column=x:\n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> PUT /api/x/scanner?column=x: HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: text/xml\nContent-Length: 7096\nExpect: 100-continue\n\n< HTTP/1.1 100 Continue\n< HTTP/1.1 201 Created\n< Date: Mon, 19 Nov 2007 23:23:33 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Location: //api /x/scanner/88316f77\n< Content-Length: 0\n* Connection #0 to host localhost left intact\n* Closing connection #0\n\n$ curl --verbose --header 'Accept: text/xml' -T /tmp/diff.txt http://localhost:60010/api/x/scanner/88316f77 \n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> PUT /api/x/scanner/88316f77 HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: text/xml\nContent-Length: 7096\nExpect: 100-continue\n\n< HTTP/1.1 100 Continue\n< HTTP/1.1 200 OK\n< Date: Mon, 19 Nov 2007 23:23:56 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Content-Type: text/xml;charset=UTF-8\n< Transfer-Encoding: chunked\n\n\n \nx\n \n \n1195372581842\n \n \n \nx:\n \n \neHl6\n \n \n* Connection #0 to host localhost left intact\n* Closing connection #0\n\n\n\n$ curl --verbose --header 'Accept: multipart/related' -T /' -T /tmp/diff.txt http://localhost:60010/api/x/scanner/88316f77\n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> PUT /api/x/scanner/88316f77 HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: multipart/related\nContent-Length: 7096\nExpect: 100-continue\n\n< HTTP/1.1 100 Continue\n< HTTP/1.1 200 OK\n< Date: Mon, 19 Nov 2007 23:24:26 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Content-Type: multipart/related\n< Content-Length: 814\n--org.mortbay.http.MultiPartResponse.boundary.f97mkcfm\nContent-Type: application/octet-stream\nContent-Description: row\nContent-Transfer-Encoding: binary\nContent-Length: 3\n\nxyz\n--org.mortbay.http.MultiPartResponse.boundary.f97mkcfm\nContent-Type: application/octet-stream\nContent-Description: timestamp\nContent-Transfer-Encoding: binary\nContent-Length: 13\n\n1195372609009\n--org.mortbay.http.MultiPartResponse.boundary.f97mkcfm\nContent-Type: application/octet-stream\nContent-Description: x:abc\nContent-Transfer-Encoding: binary\nContent-Length: 3\n\nabc\n--org.mortbay.http.MultiPartResponse.boundary.f97mkcfm\nContent-Type: application/octet-stream\nContent-Description: x:xyz\nContent-Transfer-Encoding: binary\nContent-Length: 6\n\nxyzxyz\n--org.mortbay.http.MultiPartResponse.boundary.f97mkcfm--\n* Connection #0 to host localhost left intact\n* Closing connection #0\n\n$ curl --verbose --header 'Accept: multipart/related' http://localhost:60010/api/x/row/x \n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> GET /api/x/row/x HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: multipart/related\n\n< HTTP/1.1 200 OK\n< Date: Mon, 19 Nov 2007 23:24:55 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Content-Type: multipart/related\n< Content-Length: 240\n--org.mortbay.http.MultiPartResponse.boundary.f97mkyj8\nContent-Type: application/octet-stream\nContent-Description: x:\nContent-Transfer-Encoding: binary\nContent-Length: 3\n\nxyz\n--org.mortbay.http.MultiPartResponse.boundary.f97mkyj8--\n* Connection #0 to host localhost left intact\n* Closing connection #0\n\n$ curl --verbose --header http://localhost:60010/api/x/row/x\ncurl: no URL specified!\ncurl: try 'curl --help' or 'curl --manual' for more information\ndurruti:~/Documents/checkouts/hadoop-trunk stack$ curl --verbose http://localhost:60010/api/x/row/x\n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> GET /api/x/row/x HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: */*\n\n< HTTP/1.1 200 OK\n< Date: Mon, 19 Nov 2007 23:25:22 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Content-Type: text/xml;charset=UTF-8\n< Transfer-Encoding: chunked\n\n\n \n \nx:\n \n \neHl6\n \n \n* Connection #0 to host localhost left intact\n* Closing connection #0\n\n{code}\n","created":"2007-11-19T23:27:10.370+0000"},{"body":"Updated patch for REST.java:\n\n-PUT/POST on a row without a timestamp works\n-DELETE on a row without a timestamp works\n","created":"2007-11-21T18:15:28.660+0000"},{"body":"Patch looks great\n\n+ 80 characters per line (unless its just silly doing return) and no tabs (Spacings should be two-char rather than usual 4-char).\n+ There is a define for COLUMN at top of the class. Use that and you might avoid the \"column\" in one place and \"columns\" elsewhere.\n+ I like the way you workaround lack of get(Text [] columns).\n+ Cast here is unnecessary: +^I^I^I^I^IText current_column = (Text)columns_retrieved[i];$\n+ For the below, perhaps just let the exception out:\n\n+^I^I} catch (javax.xml.parsers.ParserConfigurationException e) {$\n+^I^I^Iresponse.setStatus(500);$\n+^I^I^Ireturn;$\n+^I^I} catch (org.xml.sax.SAXException e){$\n+^I^I^Iresponse.setStatus(500);$\n+^I^I^Ireturn;$\n+^I^I}$\n\nI think jetty will do the 'right' thing (500 code plus stack trace).\n\n+ You should probably wrap the put after you call the startupdate in a try/finally. If an exception, call abort.\n","created":"2007-11-21T19:51:45.186+0000"},{"body":"Latest version of the REST functionality. Supports xml-formatted gets/puts, metadata requests, use of timestamps.","created":"2007-11-27T18:49:00.779+0000"},{"body":"Here's a version of patch w/o tabs, some formatting fixes, a license and fixed up javadoc comments -- in particular, the class comment has been changed because this patch addresses a few of the items mentioned in the TODO list.","created":"2007-11-27T20:24:14.762+0000"},{"body":"Refactored REST.java into several classes for easier digestion. ","created":"2007-11-27T23:43:42.763+0000"},{"body":"+ Each class needs a license. Copy one from adjacent hbase classes.\n+ You have tabs in this patch. Need to replaced with two spaces.\n+ Most of the imports per class are not pertinent (if you had eclipse working, it'd help you here: smile)\n+ You might instantiate HBaseConfiguration once and then pass it to each of the handler classes in their constructors. Same for admin and table instances? Otherwise, you'll have three copies of each?\n+ Each of your new handlers needs at least a class comment as javadoc saying what each does.\n\nOtherwise, I think they way you have broken apart the fat REST class is a big improvement. ","created":"2007-11-28T04:31:28.526+0000"},{"body":"-Added licenses\n-Handler constructors all take a HBaseConfiguration and HBaseAdmin instances now\n-Removed all the unnecessary imports from all the classes\n-Added class comments to all handlers\n\n\nAlso, I don't seem to be able to detect these tabs you're talking about. I thought it might be TextMate screwing me here, but I cat'd out the files and they look like they have two spaces to me.","created":"2007-11-28T17:48:05.595+0000"},{"body":"Latest version of the patch just has some more polish and a little better behavior in the area of scanners.","created":"2007-11-28T20:59:01.809+0000"},{"body":"Removed tabs, unused imports, and minor formatting changes. Also added override of NotSupported so could return messages that tell client why not supported (I was having trouble figuring why my curl upload wasn't working).\n\nI did some basic testing. Was able to put, get, scan. Its working good enough for a first version. +1.","created":"2007-11-29T00:04:04.111+0000"},{"body":"how is this feature going to handle calling rows with /'s in the row keys?\n\nexample above\nGET http://MASTER:PORT/TABLENAME/ROW/COLUMNNAME/:\n\nmy example call\nGET http://192.168.1.200:60010/webdata/com.example.www/:http/source/\n\nwhere \"com.example.www/:http\" is the row key and source is column \n\nthe / in the row key will kill the request","created":"2007-11-29T03:15:52.607+0000"},{"body":"Keys that have special characters should be URL encoded.\n\n\n\n","created":"2007-11-29T03:35:52.544+0000"},{"body":"I am still getting an error with urlencode looks like the : is throwing it off\n\ncalling row key\ncom.example.www/:http\n{code:xml} \n[root@PE1750-2 hbase]# curl --verbose http://192.168.1.200:60010/api/webdata/row/com.example.www%2F%3Ahttp\n* About to connect() to 192.168.1.200 port 60010\n* Trying 192.168.1.200... * connected\n* Connected to 192.168.1.200 (192.168.1.200) port 60010\n> GET /api/webdata/row/com.example.www%2F%3Ahttp HTTP/1.1\nUser-Agent: curl/7.12.1 (i686-redhat-linux-gnu) libcurl/7.12.1 OpenSSL/0.9.7a zlib/1.2.1.2 libidn/0.5.6\nHost: 192.168.1.200:60010\nPragma: no-cache\nAccept: */*\n\n< HTTP/1.1 500 For+input+string%3A+%22%3Ahttp%22\n< Date: Fri, 30 Nov 2007 00:53:31 GMT\n< Server: Jetty/5.1.4 (Linux/2.6.9-55.0.12.ELsmp i386 java/1.5.0_12\n< Content-Type: text/html\n< Content-Length: 1282\n< Connection: close\n\n\nError 500 For input string: \":http\"\n\n\n

HTTP ERROR: 500

For input string: \":http\"
\n

RequestURI=/api/webdata/row/com.example.www/:http

\n

Powered by Jetty://

\n\n\n\n* Closing connection #0\n{code} ","created":"2007-11-30T00:55:07.623+0000"},{"body":"I will check into this.\n\n\n\n","created":"2007-11-30T01:27:48.378+0000"},{"body":"Bryan: Try request.getRequestURI() instead of request.getPathInfo in getPathSegments method; the%2F is decoded as a slash when getPathInfo is used which is messing up the parse.","created":"2007-11-30T04:48:07.925+0000"},{"body":"Submitting preliminary REST implementation.","created":"2007-11-30T18:43:31.371+0000"},{"body":"Committed (Failed tests were unrelated to this patch which doesn't add any new tests and is code that doesn't run at unit test time). Resolving.","created":"2007-11-30T20:17:22.864+0000"},{"body":"Thanks for the patch Bryan.","created":"2007-11-30T20:17:49.730+0000"},{"body":"I am still getting the same error as above in my last post\nI am downloading the latest patch from here and HADOOP-2224 and applying them to trunk ver 598780 the latest one i know to build successful.\n\nI can pulll records that row key does not have /'s in them but not with /'s\nIs there something I am doing wrong?\n","created":"2007-12-02T00:46:41.432+0000"},{"body":"I am also having problems getting the with put/post on rows/columns working\n\nI keep getting a return code of 200 and the row is not inserted. The Wiki page saids it should be 201 but the example on the same page is showing a return code of 200 also can you verily the return code and that its working?\n\nI would like to see some kind of example in some programming language on how you are posting so I could post in place of saving to file and using put\n\nI am trying to write a php class around this to use for my project I am working on so sorry for all the questions. \n\nThanks\n\n\n","created":"2007-12-02T06:03:12.884+0000"},{"body":"You are right that the code and spec. are out of sync Billy. Inside in the putRowXml, there is following code on successful put around line #324 of the putRowXml:\n\n{code}\n // respond with a 200\n response.setStatus(200); \n{code}\n\nShould be 201.\n\nIts odd that its returning success but nothing is added. Looking at code, that shouldn't be possible.\n\nKeep on asking questions (and finding bugs). Thanks Billy.","created":"2007-12-03T17:25:52.234+0000"},{"body":"Now that I think about it, I don't think that put/post should return a 201, because we're not always creating a new resource, in the HTTP sense. We should just change the spec to say 200.","created":"2007-12-03T19:59:10.745+0000"},{"body":"After lots of work trying to get the php curl to work with the post option. Looks like we will need support for content type \"x-www-form-urlencoded\". I have tried many ways to get curl to encode the data and send as text/xml but using the post fields option in php curl sends the data with content type application/x-www-form-urlencoded. Sense that is not a supported content type I get back a http error\n406 Unsupported Accept Header Content: application/x-www-form-urlencoded\n\nSo is there a way we can add support for that?\n\nIf so we could still use the xml format as the data put make sure to urldecode it first.\nalso can we have a set filed to post the xml to that the app knows about like \"xmldata\" or something like that for post\n\nSo we can set the post field:\n\nExample \nxmldata=' a: YQ== ';\n\nThen the app knows where to look for the xml data.\n","created":"2007-12-03T20:02:12.347+0000"},{"body":"I got the put option working now its returning 200 and the data is in the table now so thats good.\n\nBut to use the put option you have to save the data to a file before putting the data via php curl.\nThats an extra step on inserting data in to the tables I would like to skip and use post option.\n","created":"2007-12-03T20:09:02.619+0000"},{"body":"I am unfamiliar with php's HTTP library. However, I don't think we \nshould change the content-type accepted for XML formatted data. An \nXML entity body is definitely NOT x-www-form-urlencoded. Hacking it \nto work around that essentially breaks the HTTP spec.\n\nI would suggest finding out if you can get lower-level access to the \nHTTP session than what php's library is giving you. This is not \nincredibly complicated functionality.\n\nAs a last resort, I invite you to submit a patch that can read the \nxml out of the postdata when encoded as x-www-form-urlencoded and \nwe'll find a way to work it in.\n\n\n\n","created":"2007-12-03T20:23:53.752+0000"},{"body":"not sure how this happens but here is the screen output\nI thank this is the reasion I was getting 200 html code btu could not see the results in shell.\nI rak this after restarting all hadoop and hbase so nothing should be cached in any sense\n\n{code:xml} \n[root@PE1750-1 bin]# pwd\n/hadoop/src/contrib/hbase/bin\n[root@PE1750-1 bin]# curl -v http://192.168.1.200:60010/api/webdata/row/10/\n* About to connect() to 192.168.1.200 port 60010\n* Trying 192.168.1.200... * connected\n* Connected to 192.168.1.200 (192.168.1.200) port 60010\n> GET /api/webdata/row/10/ HTTP/1.1\nUser-Agent: curl/7.12.1 (i686-redhat-linux-gnu) libcurl/7.12.1 OpenSSL/0.9.7a zlib/1.2.1.2 libidn/0.5.6\nHost: 192.168.1.200:60010\nPragma: no-cache\nAccept: */*\n\n< HTTP/1.1 200 OK\n< Date: Mon, 03 Dec 2007 20:51:30 GMT\n< Server: Jetty/5.1.4 (Linux/2.6.9-55.0.12.ELsmp i386 java/1.5.0_12\n< Content-Type: text/xml;charset=UTF-8\n< Transfer-Encoding: chunked\n\n\n \n \nstime:\n \n \nNDU2\n \n \n \n \nstime:now\n \n \nNzg5\n \n \n* Connection #0 to host 192.168.1.200 left intact\n* Closing connection #0\n[root@PE1750-1 bin]# ./hbase shell\nHbase Shell, 0.0.2 version.\nCopyright (c) 2007 by udanax, licensed to Apache Software Foundation.\nType 'help;' for usage.\n\nhql > select * from webdata;\n+-------------------------+-------------------------+-------------------------+\n| Row | Column | Cell |\n+-------------------------+-------------------------+-------------------------+\n0 row(s) in set (0.58 sec)\nhql > exit;\n[root@PE1750-1 bin]#\n{code} \n","created":"2007-12-03T20:52:24.179+0000"},{"body":"ok got a lot of work done on the php class to work with this interface\n\ngot insert,select,delete,scanner working by using socket connection (cut curl out).\n\nStill got some features to add to it like formating options where output is going to be xml or other.\nI guess when I am done I could make a new feature and upload a patch for the php class if you guys let php be added to the project.\n\nAny idea on when/if a way to get more then the latest version of max_versions returned will be added?\nsay I have a max_versions = 3 \nHow would I get the second oldest row?\n\nI have no need for it on this project but could see where it would be handy on others.\nmaybe a call like this\nGET /[table_name]/row/[row_key]?columns=x:&versions=3\n\n\nso far found row keys that have /'s in them do not get inserted return 500 error \nand I found that row keys that have a ? in them get truncate at the ? so it gets inserted but \nSay row key = \"aaa?bbb\"\nthe row would get inserted as row key \"aaa\" only","created":"2007-12-04T05:18:11.590+0000"},{"body":"Hey Billy,\n\nCould you post your PHP class? I need to use hbase from a PHP client and was wondering I could start from yours.\n\nThanks.","created":"2008-01-08T01:12:02.222+0000"},{"body":"it is linked from this issue I just submitted it.\n\nHADOOP-2546\n","created":"2008-01-08T06:18:18.722+0000"}],"conversations":[{"body":"A RESTful interface would be one means of making hbase accessible to clients that are not java. It might look something like the below:\n\n+ An HTTP GET of http://MASTER:PORT/ outputs the master's attributes: online meta regions, list of tables, etc.: i.e. what you see now when you go to http://MASTER:PORT/master.jsp.\n+ An HTTP GET of http://MASTER:PORT/TABLENAME: 200 if tables exists and HTableDescription (mimetype: text/plain or text/xml) or 401 if no such table. HTTP DELETE would drop the table. HTTP PUT would add one.\n+ An HTTP GET of http://MASTER:PORT/TABLENAME/ROW: 200 if row exists and 401 if not.\n+ An HTTP GET of http://MASTER:PORT/TABLENAME/ROW/COLUMNFAMILY: HColumnDescriptor (mimetype: text/plain or text/xml) or 401 if no such table.\n+ An HTTP GET of http://MASTER:PORT/TABLENAME/ROW/COLUMNNAME/: 200 and latest version (mimetype: binary/octet-stream) or 401 if no such cell. HTTP DELETE would delete the cell. HTTP PUT would add a new version.\n+ An HTTP GET of http://MASTER:PORT/TABLENAME/ROW/COLUMNNAME/TIMESTAMP: 200 (mimetype: binary/octet-stream) or 401 if no such cell. HTTP DELETE would remove. HTTP PUT would put this record.\n+ Browser originally goes against master but master then redirects to the hosting region server to serve, update, delete, etc. the addressed cell","from":"reporter","subject":"[hbase] RESTful interface"},{"body":"Generally looks good. A few comments:\n\n1. 401 is \"Unauthorized\", instead I would use 404 \"Not Found\" if a resource doesn't exist.\n2. The result of a successful PUT is 201 \"Created\".\n3. What do you think the representation of the resources will be?\n\nAnd in case you haven't seen it already, I found \"RESTful Web Services\" by by Leonard Richardson and Sam Ruby (http://www.oreilly.com/catalog/9780596529260/) really useful.","from":"developer"},{"body":"Thanks for feedback Tom. \n\n401 rather than 404 was because I was watching TV at same time as writing the issue (smile). Regards 3. above, if I understand your question, there is no typing of hbase content -- as far as hbase is concerned, its all bytes -- so returned cell data should default mimetype binary/octet-stream. Later, might add switching resource format returned keyed off request headers. \n\nFor cluster, table, and column descriptors, default text/plain UTF-8 or perhaps text/xml if specified in request headers.\n\nThanks too for pointer to Leonard's book.","from":"developer"},{"body":"Would the REST api support getting whole rows at a time? Perhaps off the GET /TABLE/ROW. The data could be packed into an XML document. \n\nI'd love to see this one get done, as it would get me the Ruby HBase client I've always wanted. ","from":"developer"},{"body":"Have you seen how rows are dumped out in xhtml in hql.jsp? Would that work for you? You can't wrap binary data in XML so I don't think row representation as XML would be a general soln. If you explicitly request a cell, then we could pass back binary with (data length in http headers). \n\nI'm thinking that main difference between the REST implemenation and current UI HQL querying on the master would be that REST clients would go first to the master but would then be redirected to the regionserver hosting the row that is being read or updated (as opposed to all being done on master HQL'ing). I'm thinking REST clients would have to be able to deal with retry if data had been moved by the time they arrived at what had been the data server. The alternative would be something like having the regionserver redirect back to the master since it is arbiter of where everything is but I imagine that could get complicated quickly what with there often being some lag before master learns that data has moved and has had chance to orchestrate redeploy in the new location.","from":"developer"},{"body":"HADOOP-2171 talks about creating a shell server which I think would be better than burdening the master with having to handle REST requests.","from":"developer"},{"body":"I haven't seen the output of hql.jsp, but I'll try and track it down. You're right, you can't stick pure binary data in XML, but you can use Base64 encoded binary data. It's a little bigger, but it cuts down on the number of requests you would have to make, which I think will be the crucial bottleneck. Of course, single column requests for pure binary data should still be available.\n\nI agree with your idea of the REST requests redirecting around between the master and the region servers. That keeps the underlying architecture correct, just with a different face on it. \n\n","from":"developer"},{"body":"HADOOP-2171 strikes me as a little odd. A client sends HQL to a server that parses the HQL to run HTable client operations against hbase. There are no load savings over running shell on client machine that I can see.\n\nI don't see a problem having the master handle REST requests. Master is generally lightly loaded. It will take a lot of traffic to make it break a sweat. The masters REST load would add the master fielding HTTP redirects -- a minor imposition. Should the REST load become burdensome, folks could put up an intermediary serve or take the load off the master by making their clients smarter doing HTable-like caching of data locations. ","from":"developer"},{"body":"See the master UI at http://MASTER:PORT (default is localhost:60000). On the master homepage, there is a HQL link that does read-only queries against tables making best effort at outputting results in XHTML (Maybe this is good enough to get you going?).\n\nWe could Base64 the data. We could also tie a boat anchor to the server too as a means of slowing it down (smile). I suppose the RESTful way to do it is just allow clients say what they can accept in request headers (Don't know if you can stipulate xml with base64 encoded content via http request headers). If the server has support, do as the client asks.","from":"developer"},{"body":"Well, either Base64 encoding the data and sending it all together takes longer than requesting each column individually, or it's the other way around. I'd be interested in seeing which way it really is. \n\nI think it is perfectly acceptable to do the column-only approach in the short term and evaluate the row oriented approach later.","from":"developer"},{"body":"Bryan Duxbury added http://wiki.apache.org/lucene-hadoop/Hbase/HbaseRest.","from":"developer"},{"body":"First cut at RESTful interface. Implements metainfo, gets, and scanners. Does not yet support put. Bunch of TODOs:\n\n+ Returning results multipart/related is crippled by lack of support in the container; jetty has a MultipartResponse class but can't set properly qualified Content-Type with boundary and start parameters... they get stripped. Need to figure out how to fix this (maybe jetty 6 does it better).\n+ Need to agree on timestamp format to use (ISO8601?)\n+ Need to fix HTable so it has table metadata; until then, you need to specify a column getting a scanner.\n\nHere's some samples run against a simple table named 'x' with column family 'x:' with following contents:\n\n{code}\nHbase> select * from x;\n+-------------------------+-------------------------+-------------------------+\n| Row | Column | Cell |\n+-------------------------+-------------------------+-------------------------+\n| x | x: | xyz |\n+-------------------------+-------------------------+-------------------------+\n| xyz | x:abc | abc |\n+-------------------------+-------------------------+-------------------------+\n| xyz | x:xyz | xyzxyz |\n+-------------------------+-------------------------+-------------------------+\n3 row(s) in set (0.19 sec)\n{code}\n\nIn below session I'm using curl. Doesn't have DELETE and I fake PUT with the -T option uploading a file:\n\n{code}\n$ curl http://localhost:60010/api/\n\n\n \nx\n
\n\n$ curl --header 'Accept: text/plain' http://localhost:60010/api/\nx\n\n$ curl --header 'Accept: text/plain' http://localhost:60010/api/x\nname: x, families: {x:={name: x, max versions: 3, compression: NONE, in memory: false, max length: 2147483647, bloom filter: none}}\n\n$ curl --header 'Accept: text/xml' http://localhost:60010/api/x\n\n\n \nx\n \n \n \n \nx:\n \n \nNONE\n \n \nNONE\n \n \n3\n \n \n2147483647\n \n \n \n
\n\n$ curl --header 'Accept: text/xml' http://localhost:60010/api/x/regions\n\n\n \n\n# only one region and its start key is null, the default table start key\n\n$ curl --verbose --header 'Accept: text/xml' http://localhost:60010/api/x/scanner\n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> GET /api/x/scanner HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: text/xml\n\n< HTTP/1.1 404 No+handler\n< Date: Mon, 19 Nov 2007 23:22:23 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Content-Type: text/html\n< Content-Length: 1228\n\n\nError 404 No handler\n\n\n...\n\n# Fails because currently you must specify a column name (To be fixed)\n\n$ curl --verbose --header 'Accept: text/xml' -T /tmp/diff.txt http://localhost:60010/api/x/scanner?column=x:\n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> PUT /api/x/scanner?column=x: HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: text/xml\nContent-Length: 7096\nExpect: 100-continue\n\n< HTTP/1.1 100 Continue\n< HTTP/1.1 201 Created\n< Date: Mon, 19 Nov 2007 23:23:33 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Location: //api /x/scanner/88316f77\n< Content-Length: 0\n* Connection #0 to host localhost left intact\n* Closing connection #0\n\n$ curl --verbose --header 'Accept: text/xml' -T /tmp/diff.txt http://localhost:60010/api/x/scanner/88316f77 \n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> PUT /api/x/scanner/88316f77 HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: text/xml\nContent-Length: 7096\nExpect: 100-continue\n\n< HTTP/1.1 100 Continue\n< HTTP/1.1 200 OK\n< Date: Mon, 19 Nov 2007 23:23:56 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Content-Type: text/xml;charset=UTF-8\n< Transfer-Encoding: chunked\n\n\n \nx\n \n \n1195372581842\n \n \n \nx:\n \n \neHl6\n \n \n* Connection #0 to host localhost left intact\n* Closing connection #0\n\n\n\n$ curl --verbose --header 'Accept: multipart/related' -T /' -T /tmp/diff.txt http://localhost:60010/api/x/scanner/88316f77\n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> PUT /api/x/scanner/88316f77 HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: multipart/related\nContent-Length: 7096\nExpect: 100-continue\n\n< HTTP/1.1 100 Continue\n< HTTP/1.1 200 OK\n< Date: Mon, 19 Nov 2007 23:24:26 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Content-Type: multipart/related\n< Content-Length: 814\n--org.mortbay.http.MultiPartResponse.boundary.f97mkcfm\nContent-Type: application/octet-stream\nContent-Description: row\nContent-Transfer-Encoding: binary\nContent-Length: 3\n\nxyz\n--org.mortbay.http.MultiPartResponse.boundary.f97mkcfm\nContent-Type: application/octet-stream\nContent-Description: timestamp\nContent-Transfer-Encoding: binary\nContent-Length: 13\n\n1195372609009\n--org.mortbay.http.MultiPartResponse.boundary.f97mkcfm\nContent-Type: application/octet-stream\nContent-Description: x:abc\nContent-Transfer-Encoding: binary\nContent-Length: 3\n\nabc\n--org.mortbay.http.MultiPartResponse.boundary.f97mkcfm\nContent-Type: application/octet-stream\nContent-Description: x:xyz\nContent-Transfer-Encoding: binary\nContent-Length: 6\n\nxyzxyz\n--org.mortbay.http.MultiPartResponse.boundary.f97mkcfm--\n* Connection #0 to host localhost left intact\n* Closing connection #0\n\n$ curl --verbose --header 'Accept: multipart/related' http://localhost:60010/api/x/row/x \n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> GET /api/x/row/x HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: multipart/related\n\n< HTTP/1.1 200 OK\n< Date: Mon, 19 Nov 2007 23:24:55 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Content-Type: multipart/related\n< Content-Length: 240\n--org.mortbay.http.MultiPartResponse.boundary.f97mkyj8\nContent-Type: application/octet-stream\nContent-Description: x:\nContent-Transfer-Encoding: binary\nContent-Length: 3\n\nxyz\n--org.mortbay.http.MultiPartResponse.boundary.f97mkyj8--\n* Connection #0 to host localhost left intact\n* Closing connection #0\n\n$ curl --verbose --header http://localhost:60010/api/x/row/x\ncurl: no URL specified!\ncurl: try 'curl --help' or 'curl --manual' for more information\ndurruti:~/Documents/checkouts/hadoop-trunk stack$ curl --verbose http://localhost:60010/api/x/row/x\n* About to connect() to localhost port 60010\n* Trying ::1... * connected\n* Connected to localhost (::1) port 60010\n> GET /api/x/row/x HTTP/1.1\nUser-Agent: curl/7.13.1 (powerpc-apple-darwin8.0) libcurl/7.13.1 OpenSSL/0.9.7l zlib/1.2.3\nHost: localhost:60010\nPragma: no-cache\nAccept: */*\n\n< HTTP/1.1 200 OK\n< Date: Mon, 19 Nov 2007 23:25:22 GMT\n< Server: Jetty/5.1.4 (Mac OS X/10.4.10 i386 java/1.5.0_07\n< Content-Type: text/xml;charset=UTF-8\n< Transfer-Encoding: chunked\n\n\n \n \nx:\n \n \neHl6\n \n \n* Connection #0 to host localhost left intact\n* Closing connection #0\n\n{code}\n","from":"developer"},{"body":"Updated patch for REST.java:\n\n-PUT/POST on a row without a timestamp works\n-DELETE on a row without a timestamp works\n","from":"developer"},{"body":"Patch looks great\n\n+ 80 characters per line (unless its just silly doing return) and no tabs (Spacings should be two-char rather than usual 4-char).\n+ There is a define for COLUMN at top of the class. Use that and you might avoid the \"column\" in one place and \"columns\" elsewhere.\n+ I like the way you workaround lack of get(Text [] columns).\n+ Cast here is unnecessary: +^I^I^I^I^IText current_column = (Text)columns_retrieved[i];$\n+ For the below, perhaps just let the exception out:\n\n+^I^I} catch (javax.xml.parsers.ParserConfigurationException e) {$\n+^I^I^Iresponse.setStatus(500);$\n+^I^I^Ireturn;$\n+^I^I} catch (org.xml.sax.SAXException e){$\n+^I^I^Iresponse.setStatus(500);$\n+^I^I^Ireturn;$\n+^I^I}$\n\nI think jetty will do the 'right' thing (500 code plus stack trace).\n\n+ You should probably wrap the put after you call the startupdate in a try/finally. If an exception, call abort.\n","from":"developer"},{"body":"Latest version of the REST functionality. Supports xml-formatted gets/puts, metadata requests, use of timestamps.","from":"developer"},{"body":"Here's a version of patch w/o tabs, some formatting fixes, a license and fixed up javadoc comments -- in particular, the class comment has been changed because this patch addresses a few of the items mentioned in the TODO list.","from":"developer"},{"body":"Refactored REST.java into several classes for easier digestion. ","from":"developer"},{"body":"+ Each class needs a license. Copy one from adjacent hbase classes.\n+ You have tabs in this patch. Need to replaced with two spaces.\n+ Most of the imports per class are not pertinent (if you had eclipse working, it'd help you here: smile)\n+ You might instantiate HBaseConfiguration once and then pass it to each of the handler classes in their constructors. Same for admin and table instances? Otherwise, you'll have three copies of each?\n+ Each of your new handlers needs at least a class comment as javadoc saying what each does.\n\nOtherwise, I think they way you have broken apart the fat REST class is a big improvement. ","from":"developer"},{"body":"-Added licenses\n-Handler constructors all take a HBaseConfiguration and HBaseAdmin instances now\n-Removed all the unnecessary imports from all the classes\n-Added class comments to all handlers\n\n\nAlso, I don't seem to be able to detect these tabs you're talking about. I thought it might be TextMate screwing me here, but I cat'd out the files and they look like they have two spaces to me.","from":"developer"},{"body":"Latest version of the patch just has some more polish and a little better behavior in the area of scanners.","from":"developer"},{"body":"Removed tabs, unused imports, and minor formatting changes. Also added override of NotSupported so could return messages that tell client why not supported (I was having trouble figuring why my curl upload wasn't working).\n\nI did some basic testing. Was able to put, get, scan. Its working good enough for a first version. +1.","from":"developer"},{"body":"how is this feature going to handle calling rows with /'s in the row keys?\n\nexample above\nGET http://MASTER:PORT/TABLENAME/ROW/COLUMNNAME/:\n\nmy example call\nGET http://192.168.1.200:60010/webdata/com.example.www/:http/source/\n\nwhere \"com.example.www/:http\" is the row key and source is column \n\nthe / in the row key will kill the request","from":"developer"},{"body":"Keys that have special characters should be URL encoded.\n\n\n\n","from":"developer"},{"body":"I am still getting an error with urlencode looks like the : is throwing it off\n\ncalling row key\ncom.example.www/:http\n{code:xml} \n[root@PE1750-2 hbase]# curl --verbose http://192.168.1.200:60010/api/webdata/row/com.example.www%2F%3Ahttp\n* About to connect() to 192.168.1.200 port 60010\n* Trying 192.168.1.200... * connected\n* Connected to 192.168.1.200 (192.168.1.200) port 60010\n> GET /api/webdata/row/com.example.www%2F%3Ahttp HTTP/1.1\nUser-Agent: curl/7.12.1 (i686-redhat-linux-gnu) libcurl/7.12.1 OpenSSL/0.9.7a zlib/1.2.1.2 libidn/0.5.6\nHost: 192.168.1.200:60010\nPragma: no-cache\nAccept: */*\n\n< HTTP/1.1 500 For+input+string%3A+%22%3Ahttp%22\n< Date: Fri, 30 Nov 2007 00:53:31 GMT\n< Server: Jetty/5.1.4 (Linux/2.6.9-55.0.12.ELsmp i386 java/1.5.0_12\n< Content-Type: text/html\n< Content-Length: 1282\n< Connection: close\n\n\nError 500 For input string: \":http\"\n\n\n

HTTP ERROR: 500

For input string: \":http\"
\n

RequestURI=/api/webdata/row/com.example.www/:http

\n

Powered by Jetty://

\n\n\n\n* Closing connection #0\n{code} ","from":"developer"},{"body":"I will check into this.\n\n\n\n","from":"developer"},{"body":"Bryan: Try request.getRequestURI() instead of request.getPathInfo in getPathSegments method; the%2F is decoded as a slash when getPathInfo is used which is messing up the parse.","from":"developer"},{"body":"Submitting preliminary REST implementation.","from":"developer"},{"body":"Committed (Failed tests were unrelated to this patch which doesn't add any new tests and is code that doesn't run at unit test time). Resolving.","from":"developer"},{"body":"Thanks for the patch Bryan.","from":"developer"},{"body":"I am still getting the same error as above in my last post\nI am downloading the latest patch from here and HADOOP-2224 and applying them to trunk ver 598780 the latest one i know to build successful.\n\nI can pulll records that row key does not have /'s in them but not with /'s\nIs there something I am doing wrong?\n","from":"developer"},{"body":"I am also having problems getting the with put/post on rows/columns working\n\nI keep getting a return code of 200 and the row is not inserted. The Wiki page saids it should be 201 but the example on the same page is showing a return code of 200 also can you verily the return code and that its working?\n\nI would like to see some kind of example in some programming language on how you are posting so I could post in place of saving to file and using put\n\nI am trying to write a php class around this to use for my project I am working on so sorry for all the questions. \n\nThanks\n\n\n","from":"developer"},{"body":"You are right that the code and spec. are out of sync Billy. Inside in the putRowXml, there is following code on successful put around line #324 of the putRowXml:\n\n{code}\n // respond with a 200\n response.setStatus(200); \n{code}\n\nShould be 201.\n\nIts odd that its returning success but nothing is added. Looking at code, that shouldn't be possible.\n\nKeep on asking questions (and finding bugs). Thanks Billy.","from":"developer"},{"body":"Now that I think about it, I don't think that put/post should return a 201, because we're not always creating a new resource, in the HTTP sense. We should just change the spec to say 200.","from":"developer"},{"body":"After lots of work trying to get the php curl to work with the post option. Looks like we will need support for content type \"x-www-form-urlencoded\". I have tried many ways to get curl to encode the data and send as text/xml but using the post fields option in php curl sends the data with content type application/x-www-form-urlencoded. Sense that is not a supported content type I get back a http error\n406 Unsupported Accept Header Content: application/x-www-form-urlencoded\n\nSo is there a way we can add support for that?\n\nIf so we could still use the xml format as the data put make sure to urldecode it first.\nalso can we have a set filed to post the xml to that the app knows about like \"xmldata\" or something like that for post\n\nSo we can set the post field:\n\nExample \nxmldata=' a: YQ== ';\n\nThen the app knows where to look for the xml data.\n","from":"developer"},{"body":"I got the put option working now its returning 200 and the data is in the table now so thats good.\n\nBut to use the put option you have to save the data to a file before putting the data via php curl.\nThats an extra step on inserting data in to the tables I would like to skip and use post option.\n","from":"developer"},{"body":"I am unfamiliar with php's HTTP library. However, I don't think we \nshould change the content-type accepted for XML formatted data. An \nXML entity body is definitely NOT x-www-form-urlencoded. Hacking it \nto work around that essentially breaks the HTTP spec.\n\nI would suggest finding out if you can get lower-level access to the \nHTTP session than what php's library is giving you. This is not \nincredibly complicated functionality.\n\nAs a last resort, I invite you to submit a patch that can read the \nxml out of the postdata when encoded as x-www-form-urlencoded and \nwe'll find a way to work it in.\n\n\n\n","from":"developer"},{"body":"not sure how this happens but here is the screen output\nI thank this is the reasion I was getting 200 html code btu could not see the results in shell.\nI rak this after restarting all hadoop and hbase so nothing should be cached in any sense\n\n{code:xml} \n[root@PE1750-1 bin]# pwd\n/hadoop/src/contrib/hbase/bin\n[root@PE1750-1 bin]# curl -v http://192.168.1.200:60010/api/webdata/row/10/\n* About to connect() to 192.168.1.200 port 60010\n* Trying 192.168.1.200... * connected\n* Connected to 192.168.1.200 (192.168.1.200) port 60010\n> GET /api/webdata/row/10/ HTTP/1.1\nUser-Agent: curl/7.12.1 (i686-redhat-linux-gnu) libcurl/7.12.1 OpenSSL/0.9.7a zlib/1.2.1.2 libidn/0.5.6\nHost: 192.168.1.200:60010\nPragma: no-cache\nAccept: */*\n\n< HTTP/1.1 200 OK\n< Date: Mon, 03 Dec 2007 20:51:30 GMT\n< Server: Jetty/5.1.4 (Linux/2.6.9-55.0.12.ELsmp i386 java/1.5.0_12\n< Content-Type: text/xml;charset=UTF-8\n< Transfer-Encoding: chunked\n\n\n \n \nstime:\n \n \nNDU2\n \n \n \n \nstime:now\n \n \nNzg5\n \n \n* Connection #0 to host 192.168.1.200 left intact\n* Closing connection #0\n[root@PE1750-1 bin]# ./hbase shell\nHbase Shell, 0.0.2 version.\nCopyright (c) 2007 by udanax, licensed to Apache Software Foundation.\nType 'help;' for usage.\n\nhql > select * from webdata;\n+-------------------------+-------------------------+-------------------------+\n| Row | Column | Cell |\n+-------------------------+-------------------------+-------------------------+\n0 row(s) in set (0.58 sec)\nhql > exit;\n[root@PE1750-1 bin]#\n{code} \n","from":"developer"},{"body":"ok got a lot of work done on the php class to work with this interface\n\ngot insert,select,delete,scanner working by using socket connection (cut curl out).\n\nStill got some features to add to it like formating options where output is going to be xml or other.\nI guess when I am done I could make a new feature and upload a patch for the php class if you guys let php be added to the project.\n\nAny idea on when/if a way to get more then the latest version of max_versions returned will be added?\nsay I have a max_versions = 3 \nHow would I get the second oldest row?\n\nI have no need for it on this project but could see where it would be handy on others.\nmaybe a call like this\nGET /[table_name]/row/[row_key]?columns=x:&versions=3\n\n\nso far found row keys that have /'s in them do not get inserted return 500 error \nand I found that row keys that have a ? in them get truncate at the ? so it gets inserted but \nSay row key = \"aaa?bbb\"\nthe row would get inserted as row key \"aaa\" only","from":"developer"},{"body":"Hey Billy,\n\nCould you post your PHP class? I need to use hbase from a PHP client and was wondering I could start from yours.\n\nThanks.","from":"developer"},{"body":"it is linked from this issue I just submitted it.\n\nHADOOP-2546\n","from":"developer"}],"created":"2007-10-17T05:20:06.000+0000","description":"A RESTful interface would be one means of making hbase accessible to clients that are not java. It might look something like the below:\n\n+ An HTTP GET of http://MASTER:PORT/ outputs the master's attributes: online meta regions, list of tables, etc.: i.e. what you see now when you go to http://MASTER:PORT/master.jsp.\n+ An HTTP GET of http://MASTER:PORT/TABLENAME: 200 if tables exists and HTableDescription (mimetype: text/plain or text/xml) or 401 if no such table. HTTP DELETE would drop the table. HTTP PUT would add one.\n+ An HTTP GET of http://MASTER:PORT/TABLENAME/ROW: 200 if row exists and 401 if not.\n+ An HTTP GET of http://MASTER:PORT/TABLENAME/ROW/COLUMNFAMILY: HColumnDescriptor (mimetype: text/plain or text/xml) or 401 if no such table.\n+ An HTTP GET of http://MASTER:PORT/TABLENAME/ROW/COLUMNNAME/: 200 and latest version (mimetype: binary/octet-stream) or 401 if no such cell. HTTP DELETE would delete the cell. HTTP PUT would add a new version.\n+ An HTTP GET of http://MASTER:PORT/TABLENAME/ROW/COLUMNNAME/TIMESTAMP: 200 (mimetype: binary/octet-stream) or 401 if no such cell. HTTP DELETE would remove. HTTP PUT would put this record.\n+ Browser originally goes against master but master then redirects to the hosting region server to serve, update, delete, etc. the addressed cell","issue_id":"12380573","key":"HBASE-269","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2007-11-30T20:17:22.000+0000","role":"fixed_distractor","summary":"[hbase] RESTful interface"} {"case_id":"13550561","cluster":"DISTRACTOR-HBASE-28082","comments":[{"body":"I think the problem is related to your usage of multiwal for hbase.wal.provider. It seems like that feature adds a \".regiongroup-#\" suffix to the WAL path name. I do think the bug is in the backup system, we should fix BackupUtils#parseHostFromOldLog to ignore that before trying to extract a ServerName.","created":"2023-09-13T16:39:26.444+0000"},{"body":"I have a patch that works by making less assumptions about the actual file name other than that it starts with a ServerName (host,port,...). \r\n\r\nCan't seem to attach patch files though.\r\n\r\n{code:java}\r\nFrom 40d88d9253c78e04823af49f199684bd8ac03966 Mon Sep 17 00:00:00 2001\r\nFrom: Jan Van Besien \r\nDate: Mon, 2 Oct 2023 11:07:59 +0200\r\nSubject: [PATCH] HBASE-28082 more lenient WAL hostname parsing\r\n\r\nMake the hostname parsing in BackupUtils#parseHostFromOldLog more lenient\r\nby not making any assumptions about the name of the file other than that\r\nit starts with a org.apache.hadoop.hbase.ServerName.\r\n---\r\n .../hadoop/hbase/backup/util/BackupUtils.java | 10 +++++----\r\n .../hadoop/hbase/backup/TestBackupUtils.java | 22 +++++++++++--------\r\n 2 files changed, 19 insertions(+), 13 deletions(-)\r\n\r\ndiff --git a/hbase-backup/src/main/java/org/apache/hadoop/hbase/backup/util/BackupUtils.java b/hbase-backup/src/main/java/org/apache/hadoop/hbase/backup/util/BackupUtils.java\r\nindex 5be8eed3952..a920b55bca9 100644\r\n--- a/hbase-backup/src/main/java/org/apache/hadoop/hbase/backup/util/BackupUtils.java\r\n+++ b/hbase-backup/src/main/java/org/apache/hadoop/hbase/backup/util/BackupUtils.java\r\n@@ -30,6 +30,7 @@ import java.util.Map;\r\n import java.util.Map.Entry;\r\n import java.util.TreeMap;\r\n import java.util.TreeSet;\r\n+import com.google.common.collect.Iterables;\r\n import org.apache.hadoop.conf.Configuration;\r\n import org.apache.hadoop.fs.FSDataOutputStream;\r\n import org.apache.hadoop.fs.FileStatus;\r\n@@ -365,10 +366,11 @@ public final class BackupUtils {\r\n return null;\r\n }\r\n try {\r\n- String n = p.getName();\r\n- int idx = n.lastIndexOf(LOGNAME_SEPARATOR);\r\n- String s = URLDecoder.decode(n.substring(0, idx), \"UTF8\");\r\n- return ServerName.valueOf(s).getAddress().toString();\r\n+ String urlDecodedName = URLDecoder.decode(p.getName(), \"UTF8\");\r\n+ Iterable nameSplitsOnComma = Splitter.on(\",\").split(urlDecodedName);\r\n+ String host = Iterables.get(nameSplitsOnComma, 0);\r\n+ String port = Iterables.get(nameSplitsOnComma, 1);\r\n+ return host + \":\" + port;\r\n } catch (Exception e) {\r\n LOG.warn(\"Skip log file (can't parse): {}\", p);\r\n return null;\r\ndiff --git a/hbase-backup/src/test/java/org/apache/hadoop/hbase/backup/TestBackupUtils.java b/hbase-backup/src/test/java/org/apache/hadoop/hbase/backup/TestBackupUtils.java\r\nindex 831ec309cfc..6aebe8db082 100644\r\n--- a/hbase-backup/src/test/java/org/apache/hadoop/hbase/backup/TestBackupUtils.java\r\n+++ b/hbase-backup/src/test/java/org/apache/hadoop/hbase/backup/TestBackupUtils.java\r\n@@ -87,21 +87,25 @@ public class TestBackupUtils {\r\n \r\n @Test\r\n public void testFilesystemWalHostNameParsing() throws IOException {\r\n- String host = \"localhost\";\r\n+ String host = \"a-region-server.domain.com\";\r\n int port = 60030;\r\n ServerName serverName = ServerName.valueOf(host, port, 1234);\r\n Path walRootDir = CommonFSUtils.getWALRootDir(conf);\r\n Path oldLogDir = new Path(walRootDir, HConstants.HREGION_OLDLOGDIR_NAME);\r\n \r\n- Path testWalPath = new Path(oldLogDir,\r\n- serverName.toString() + BackupUtils.LOGNAME_SEPARATOR + EnvironmentEdgeManager.currentTime());\r\n- Path testMasterWalPath =\r\n- new Path(oldLogDir, testWalPath.getName() + MasterRegionFactory.ARCHIVED_WAL_SUFFIX);\r\n+ Path testOldWalPath = new Path(oldLogDir,\r\n+ serverName + BackupUtils.LOGNAME_SEPARATOR + EnvironmentEdgeManager.currentTime());\r\n+ Assert.assertEquals(host + Addressing.HOSTNAME_PORT_SEPARATOR + port, BackupUtils.parseHostFromOldLog(testOldWalPath));\r\n \r\n- String parsedHost = BackupUtils.parseHostFromOldLog(testMasterWalPath);\r\n- Assert.assertNull(parsedHost);\r\n+ Path testMasterWalPath =\r\n+ new Path(oldLogDir, testOldWalPath.getName() + MasterRegionFactory.ARCHIVED_WAL_SUFFIX);\r\n+ Assert.assertNull(BackupUtils.parseHostFromOldLog(testMasterWalPath));\r\n \r\n- parsedHost = BackupUtils.parseHostFromOldLog(testWalPath);\r\n- Assert.assertEquals(parsedHost, host + Addressing.HOSTNAME_PORT_SEPARATOR + port);\r\n+ // org.apache.hadoop.hbase.wal.BoundedGroupingStrategy does this\r\n+ Path testOldWalWithRegionGroupingPath = new Path(oldLogDir, \r\n+ serverName + BackupUtils.LOGNAME_SEPARATOR + serverName + \r\n+ BackupUtils.LOGNAME_SEPARATOR + \"regiongroup-0\" + BackupUtils.LOGNAME_SEPARATOR + \r\n+ EnvironmentEdgeManager.currentTime());\r\n+ Assert.assertEquals(host + Addressing.HOSTNAME_PORT_SEPARATOR + port, BackupUtils.parseHostFromOldLog(testOldWalWithRegionGroupingPath));\r\n }\r\n }\r\n-- \r\n2.41.0\r\n{code}\r\n","created":"2023-10-02T09:15:13.800+0000"},{"body":"The preferred way to contribute these days is by submitting a PR at https://github.com/apache/hbase","created":"2023-10-02T11:08:36.262+0000"},{"body":"okay, here it is: https://github.com/apache/hbase/pull/5445","created":"2023-10-02T11:15:50.702+0000"},{"body":"Pushed to branch-2+.\r\n\r\nThanks [~janvanbesien] for contributing!","created":"2023-10-07T01:40:30.646+0000"},{"body":"I lost track of this, but just ran into the issue in our environment and was pleased to find it solved. Thanks [~janvanbesien]!","created":"2024-04-01T19:09:17.812+0000"}],"conversations":[{"body":"I am testing HBase backup functionality, and noticed following warning when running \"hbase backup create incremental ...\":\r\n\r\n \r\n{noformat}\r\n23/09/13 15:44:10 WARN org.apache.hadoop.hbase.backup.util.BackupUtils: Skip log file (can't parse): hdfs://hdfsns/hbase/hbase/oldWALs/hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.regiongroup-0.1694609969312{noformat}\r\nIt appears in my setup, the oldWALs are indeed given names that seem to break \"ServerName.valueOf(s)\" in \"BackupUtils#parseHostFromOldLog(Path p)\":\r\n\r\n \r\n\r\n \r\n{noformat}\r\nuser@hadoop-client-769bc9946-xqrt2:/$ hdfs dfs -ls hdfs:///hbase/hbase/oldWALs\r\nFound 42 items\r\n-rw-r--r--   1 hbase hbase     775421 2023-09-13 13:14 hdfs:///hbase/hbase/oldWALs/hbase-master-0.minikube-shared%2C16000%2C1694609954719.hbase-master-0.minikube-shared%2C16000%2C1694609954719.regiongroup-0.1694609957984$masterlocalwal$\r\n-rw-r--r--   1 hbase hbase      26059 2023-09-13 13:29 hdfs:///hbase/hbase/oldWALs/hbase-master-0.minikube-shared%2C16000%2C1694609954719.hbase-master-0.minikube-shared%2C16000%2C1694609954719.regiongroup-0.1694610867894$masterlocalwal$\r\n...\r\n-rw-r--r--   1 hbase hbase     242479 2023-09-13 14:16 hdfs:///hbase/hbase/oldWALs/hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.regiongroup-0.1694609969312\r\n-rw-r--r--   1 hbase hbase       4364 2023-09-13 14:16 hdfs:///hbase/hbase/oldWALs/hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.regiongroup-0.1694610188654\r\n...\r\n-rw-r--r--   1 hbase hbase      70802 2023-09-13 13:15 hdfs:///hbase/hbase/oldWALs/hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.meta.1694609970025.meta\r\n-rw-r--r--   1 hbase hbase         93 2023-09-13 13:04 hdfs:///hbase/hbase/oldWALs/hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.meta.1694610188627.meta\r\n...{noformat}\r\nI'd say this is not a bug in the backup system, but rather in whatever gives the oldWAL files its name. I'm however not that familiar with HBase code to find where these files are created. Any pointers are appreciated.\r\n\r\nGiven that this causes some logs to be missed during backup, I guess this can lead to data loss in a backup restore?\r\n\r\n ","from":"reporter","subject":"oldWALs naming can be incompatible with HBase backup"},{"body":"I think the problem is related to your usage of multiwal for hbase.wal.provider. It seems like that feature adds a \".regiongroup-#\" suffix to the WAL path name. I do think the bug is in the backup system, we should fix BackupUtils#parseHostFromOldLog to ignore that before trying to extract a ServerName.","from":"developer"},{"body":"I have a patch that works by making less assumptions about the actual file name other than that it starts with a ServerName (host,port,...). \r\n\r\nCan't seem to attach patch files though.\r\n\r\n{code:java}\r\nFrom 40d88d9253c78e04823af49f199684bd8ac03966 Mon Sep 17 00:00:00 2001\r\nFrom: Jan Van Besien \r\nDate: Mon, 2 Oct 2023 11:07:59 +0200\r\nSubject: [PATCH] HBASE-28082 more lenient WAL hostname parsing\r\n\r\nMake the hostname parsing in BackupUtils#parseHostFromOldLog more lenient\r\nby not making any assumptions about the name of the file other than that\r\nit starts with a org.apache.hadoop.hbase.ServerName.\r\n---\r\n .../hadoop/hbase/backup/util/BackupUtils.java | 10 +++++----\r\n .../hadoop/hbase/backup/TestBackupUtils.java | 22 +++++++++++--------\r\n 2 files changed, 19 insertions(+), 13 deletions(-)\r\n\r\ndiff --git a/hbase-backup/src/main/java/org/apache/hadoop/hbase/backup/util/BackupUtils.java b/hbase-backup/src/main/java/org/apache/hadoop/hbase/backup/util/BackupUtils.java\r\nindex 5be8eed3952..a920b55bca9 100644\r\n--- a/hbase-backup/src/main/java/org/apache/hadoop/hbase/backup/util/BackupUtils.java\r\n+++ b/hbase-backup/src/main/java/org/apache/hadoop/hbase/backup/util/BackupUtils.java\r\n@@ -30,6 +30,7 @@ import java.util.Map;\r\n import java.util.Map.Entry;\r\n import java.util.TreeMap;\r\n import java.util.TreeSet;\r\n+import com.google.common.collect.Iterables;\r\n import org.apache.hadoop.conf.Configuration;\r\n import org.apache.hadoop.fs.FSDataOutputStream;\r\n import org.apache.hadoop.fs.FileStatus;\r\n@@ -365,10 +366,11 @@ public final class BackupUtils {\r\n return null;\r\n }\r\n try {\r\n- String n = p.getName();\r\n- int idx = n.lastIndexOf(LOGNAME_SEPARATOR);\r\n- String s = URLDecoder.decode(n.substring(0, idx), \"UTF8\");\r\n- return ServerName.valueOf(s).getAddress().toString();\r\n+ String urlDecodedName = URLDecoder.decode(p.getName(), \"UTF8\");\r\n+ Iterable nameSplitsOnComma = Splitter.on(\",\").split(urlDecodedName);\r\n+ String host = Iterables.get(nameSplitsOnComma, 0);\r\n+ String port = Iterables.get(nameSplitsOnComma, 1);\r\n+ return host + \":\" + port;\r\n } catch (Exception e) {\r\n LOG.warn(\"Skip log file (can't parse): {}\", p);\r\n return null;\r\ndiff --git a/hbase-backup/src/test/java/org/apache/hadoop/hbase/backup/TestBackupUtils.java b/hbase-backup/src/test/java/org/apache/hadoop/hbase/backup/TestBackupUtils.java\r\nindex 831ec309cfc..6aebe8db082 100644\r\n--- a/hbase-backup/src/test/java/org/apache/hadoop/hbase/backup/TestBackupUtils.java\r\n+++ b/hbase-backup/src/test/java/org/apache/hadoop/hbase/backup/TestBackupUtils.java\r\n@@ -87,21 +87,25 @@ public class TestBackupUtils {\r\n \r\n @Test\r\n public void testFilesystemWalHostNameParsing() throws IOException {\r\n- String host = \"localhost\";\r\n+ String host = \"a-region-server.domain.com\";\r\n int port = 60030;\r\n ServerName serverName = ServerName.valueOf(host, port, 1234);\r\n Path walRootDir = CommonFSUtils.getWALRootDir(conf);\r\n Path oldLogDir = new Path(walRootDir, HConstants.HREGION_OLDLOGDIR_NAME);\r\n \r\n- Path testWalPath = new Path(oldLogDir,\r\n- serverName.toString() + BackupUtils.LOGNAME_SEPARATOR + EnvironmentEdgeManager.currentTime());\r\n- Path testMasterWalPath =\r\n- new Path(oldLogDir, testWalPath.getName() + MasterRegionFactory.ARCHIVED_WAL_SUFFIX);\r\n+ Path testOldWalPath = new Path(oldLogDir,\r\n+ serverName + BackupUtils.LOGNAME_SEPARATOR + EnvironmentEdgeManager.currentTime());\r\n+ Assert.assertEquals(host + Addressing.HOSTNAME_PORT_SEPARATOR + port, BackupUtils.parseHostFromOldLog(testOldWalPath));\r\n \r\n- String parsedHost = BackupUtils.parseHostFromOldLog(testMasterWalPath);\r\n- Assert.assertNull(parsedHost);\r\n+ Path testMasterWalPath =\r\n+ new Path(oldLogDir, testOldWalPath.getName() + MasterRegionFactory.ARCHIVED_WAL_SUFFIX);\r\n+ Assert.assertNull(BackupUtils.parseHostFromOldLog(testMasterWalPath));\r\n \r\n- parsedHost = BackupUtils.parseHostFromOldLog(testWalPath);\r\n- Assert.assertEquals(parsedHost, host + Addressing.HOSTNAME_PORT_SEPARATOR + port);\r\n+ // org.apache.hadoop.hbase.wal.BoundedGroupingStrategy does this\r\n+ Path testOldWalWithRegionGroupingPath = new Path(oldLogDir, \r\n+ serverName + BackupUtils.LOGNAME_SEPARATOR + serverName + \r\n+ BackupUtils.LOGNAME_SEPARATOR + \"regiongroup-0\" + BackupUtils.LOGNAME_SEPARATOR + \r\n+ EnvironmentEdgeManager.currentTime());\r\n+ Assert.assertEquals(host + Addressing.HOSTNAME_PORT_SEPARATOR + port, BackupUtils.parseHostFromOldLog(testOldWalWithRegionGroupingPath));\r\n }\r\n }\r\n-- \r\n2.41.0\r\n{code}\r\n","from":"developer"},{"body":"The preferred way to contribute these days is by submitting a PR at https://github.com/apache/hbase","from":"developer"},{"body":"okay, here it is: https://github.com/apache/hbase/pull/5445","from":"developer"},{"body":"Pushed to branch-2+.\r\n\r\nThanks [~janvanbesien] for contributing!","from":"developer"},{"body":"I lost track of this, but just ran into the issue in our environment and was pleased to find it solved. Thanks [~janvanbesien]!","from":"developer"}],"created":"2023-09-13T16:19:51.000+0000","description":"I am testing HBase backup functionality, and noticed following warning when running \"hbase backup create incremental ...\":\r\n\r\n \r\n{noformat}\r\n23/09/13 15:44:10 WARN org.apache.hadoop.hbase.backup.util.BackupUtils: Skip log file (can't parse): hdfs://hdfsns/hbase/hbase/oldWALs/hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.regiongroup-0.1694609969312{noformat}\r\nIt appears in my setup, the oldWALs are indeed given names that seem to break \"ServerName.valueOf(s)\" in \"BackupUtils#parseHostFromOldLog(Path p)\":\r\n\r\n \r\n\r\n \r\n{noformat}\r\nuser@hadoop-client-769bc9946-xqrt2:/$ hdfs dfs -ls hdfs:///hbase/hbase/oldWALs\r\nFound 42 items\r\n-rw-r--r--   1 hbase hbase     775421 2023-09-13 13:14 hdfs:///hbase/hbase/oldWALs/hbase-master-0.minikube-shared%2C16000%2C1694609954719.hbase-master-0.minikube-shared%2C16000%2C1694609954719.regiongroup-0.1694609957984$masterlocalwal$\r\n-rw-r--r--   1 hbase hbase      26059 2023-09-13 13:29 hdfs:///hbase/hbase/oldWALs/hbase-master-0.minikube-shared%2C16000%2C1694609954719.hbase-master-0.minikube-shared%2C16000%2C1694609954719.regiongroup-0.1694610867894$masterlocalwal$\r\n...\r\n-rw-r--r--   1 hbase hbase     242479 2023-09-13 14:16 hdfs:///hbase/hbase/oldWALs/hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.regiongroup-0.1694609969312\r\n-rw-r--r--   1 hbase hbase       4364 2023-09-13 14:16 hdfs:///hbase/hbase/oldWALs/hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.regiongroup-0.1694610188654\r\n...\r\n-rw-r--r--   1 hbase hbase      70802 2023-09-13 13:15 hdfs:///hbase/hbase/oldWALs/hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.meta.1694609970025.meta\r\n-rw-r--r--   1 hbase hbase         93 2023-09-13 13:04 hdfs:///hbase/hbase/oldWALs/hbase-region-0.hbase-region.minikube-shared.svc.cluster.local%2C16020%2C1694609964681.meta.1694610188627.meta\r\n...{noformat}\r\nI'd say this is not a bug in the backup system, but rather in whatever gives the oldWAL files its name. I'm however not that familiar with HBase code to find where these files are created. Any pointers are appreciated.\r\n\r\nGiven that this causes some logs to be missed during backup, I guess this can lead to data loss in a backup restore?\r\n\r\n ","issue_id":"13550561","key":"HBASE-28082","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2023-10-07T01:40:30.000+0000","role":"fixed_distractor","summary":"oldWALs naming can be incompatible with HBase backup"} {"case_id":"12468502","cluster":"DISTRACTOR-HBASE-2812","comments":[{"body":"What I see a master that thinks that some region is hosted on a region server, that that region server receives the message but does nothing about it (in the scope of that log). Bigger log snippet and thread dump would really be helpful, or even a basic unit test that shows the issue.\n\nAlso the fix version is for 0.20.6 but you wrote that you used 0.89... so can we move it only to trunk in order to release 0.20.6? Thanks Sam.","created":"2010-07-13T19:50:58.307+0000"},{"body":"Didn't see any objections for 2 days, moving to 0.90","created":"2010-07-15T18:46:40.836+0000"},{"body":"I have also run into this, using 0.20.6. I have a test environment that runs a set of automated tests with our latest code every hour. Some of these tests are set to disable, drop, and recreate a table in order to make an empty one. After succeeding for 5 to 10 executions, it seems to get into this state. I'm not able to disable and get rid of the test table until hbase is restarted, so it is definitely hampering our ability to do automated testing with hbase.\n\nThe logs I have look similar, though the master and region server are separate processes with separate logs in my environment.\n\nDoes anyone have any ideas as to the cause of this? Is there any more information that would help in diagnosing?","created":"2010-08-05T16:38:34.998+0000"},{"body":"Dave sent me the logs, and it really looks like a case of HBASE-2755:\n\n{noformat}\n2010-08-04 19:26:58,291 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_ReferralQueue,,1280975190155 to setClosing list\n2010-08-04 19:26:58,565 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scanning meta region {server: 192.168.21.66:60020, regionname: .META.,,1, startKey: <>}\n2010-08-04 19:26:58,654 INFO org.apache.hadoop.hbase.master.ServerManager: Processing MSG_REPORT_CLOSE: test_ReferralQueue,,1280975190155 from hslave3,60020,1280951333279; 1 of 1\n2010-08-04 19:26:58,654 DEBUG org.apache.hadoop.hbase.master.HMaster: Processing todo: ProcessRegionClose of test_ReferralQueue,,1280975190155, true, reassign: false\n2010-08-04 19:26:58,656 INFO org.apache.hadoop.hbase.master.ProcessRegionClose$1: region closed: test_ReferralQueue,,1280975190155\n2010-08-04 19:26:58,686 DEBUG org.apache.hadoop.hbase.master.BaseScanner: GET on test_ReferralQueue,,1280975190155 got different startcode than SCAN: sc=0, serverAddress=1280951333279\n2010-08-04 19:26:58,686 DEBUG org.apache.hadoop.hbase.master.BaseScanner: Current assignment of test_ReferralQueue,,1280975190155 is not valid; serverAddress=, startCode=0 unknown.\n{noformat}\n\nHere the base scanner saw that that region was half unassigned (atomicity problem?) since we don't see the message that \"sa\" is different. It means that the base scanner reassigns the region because it believes that:\n\n - the region isn't assigned to anyone\n - that it's not online\n - that it's not in transition\n\nThis means that we should also verify our assumption that the HRI.isOffline() is correct when we read it. I think we can augment HBASE-2755's patch to do that, since doing the GET we already read the new HRI. Making a patch.","created":"2010-08-05T17:42:50.312+0000"},{"body":"Patch generated from 0.20.6 (so that Dave can test) that incorporates HBASE-2755 and adds the checking of the newly obtained HRI.","created":"2010-08-05T18:01:53.744+0000"},{"body":"The automated tests have now run for close to 100 executions against 0.20.6 + this patch without any problems this time, so it appears to fix the issue. Is anyone up to review and commit it to the 0.20 branch?","created":"2010-08-09T15:40:24.156+0000"},{"body":"So a review would be missing (although my patch is mostly 2755), but also the confirmation that it fixes Sam's problem. If not, then we should open a new jira and commit the fix there.","created":"2010-08-09T19:00:30.061+0000"},{"body":"+1 Patch looks good to me.","created":"2010-08-10T23:23:51.699+0000"},{"body":"Since I still didn't get the confirmation from Sam that this patch fixed his issue, I opened HBASE-2927 and posted the patch there. Keeping this jira open until we can tell for sure that the original issue is fixed.","created":"2010-08-18T19:34:41.538+0000"},{"body":"Table enable/disable is different since hbase-2692 went in. Should make this a non-issue.","created":"2010-09-01T22:19:08.036+0000"},{"body":"Moving into 0.20.7. This I believe is fixed in TRUNK (0.90).","created":"2010-09-25T06:34:31.388+0000"}],"conversations":[{"body":"I see this in the client after it gives up:\n\nhbase(main):006:0> disable 'test_schema'\n\nERROR: org.apache.hadoop.hbase.RegionException: Retries exhausted, it took too long to wait for the table test_schema to be disabled.\n\nHere is some help for this command:\n Disable the named table: e.g. \"hbase> disable 't1'\"\n\nand this in the server log, a set of about 5 reports it is closing per disable call:\n\n2010-07-03 15:19:47,554 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:47,554 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:47,555 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:47,576 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:47,576 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:48,567 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:48,567 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:48,568 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:48,577 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:48,578 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:49,580 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:49,580 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:49,581 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:50,580 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:50,581 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:50,592 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:50,592 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:50,593 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:51,581 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:51,581 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:52,605 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:52,605 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:52,606 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:52,703 INFO org.apache.hadoop.hbase.master.ServerManager: 1 region servers, 0 dead, average load 3.0\n2010-07-03 15:19:52,863 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scanning meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>}\n2010-07-03 15:19:52,867 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scan of 1 row(s) of meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>} complete\n2010-07-03 15:19:53,585 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:53,585 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:53,671 DEBUG org.apache.hadoop.hbase.io.hfile.LruBlockCache: Cache Stats: Sizes: Total=1.6292953MB (1708440), Free=196.70822MB (206263512), Max=198.33751MB (207971952), Counts: Blocks=2, Access=52, Hit=50, Miss=2, Evictions=0, Evicted=0, Ratios: Hit Ratio=96.15384340286255%, Miss Ratio=3.8461539894342422%, Evicted/Run=NaN\n2010-07-03 15:19:54,617 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:54,617 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:54,618 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:54,821 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scanning meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>}\n2010-07-03 15:19:54,825 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scan of 2 row(s) of meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>} complete\n2010-07-03 15:19:54,825 INFO org.apache.hadoop.hbase.master.BaseScanner: All 1 .META. region(s) scanned\n2010-07-03 15:19:55,589 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:55,589 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:58,629 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:58,629 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:58,630 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:59,595 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:59,595 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:20:52,706 INFO org.apache.hadoop.hbase.master.ServerManager: 1 region servers, 0 dead, average load 3.0\n2010-07-03 15:20:52,866 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scanning meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>}\n2010-07-03 15:20:52,873 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scan of 1 row(s) of meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>} complete\n2010-07-03 15:20:53,671 DEBUG org.apache.hadoop.hbase.io.hfile.LruBlockCache: Cache Stats: Sizes: Total=1.6292953MB (1708440), Free=196.70822MB (206263512), Max=198.33751MB (207971952), Counts: Blocks=2, Access=54, Hit=52, Miss=2, Evictions=0, Evicted=0, Ratios: Hit Ratio=96.29629850387573%, Miss Ratio=3.7037037312984467%, Evicted/Run=NaN\n2010-07-03 15:20:54,823 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scanning meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>}\n2010-07-03 15:20:54,828 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scan of 2 row(s) of meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>} complete\n2010-07-03 15:20:54,828 INFO org.apache.hadoop.hbase.master.BaseScanner: All 1 .META. region(s) scanned\n2010-07-03 15:21:52,709 INFO org.apache.hadoop.hbase.master.ServerManager: 1 region servers, 0 dead, average load 3.0\n2010-07-03 15:21:52,869 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scanning meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>}\n2010-07-03 15:21:52,876 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scan of 1 row(s) of meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>} complete\n2010-07-03 15:21:53,672 DEBUG org.apache.hadoop.hbase.io.hfile.LruBlockCache: Cache Stats: Sizes: Total=1.6292953MB (1708440), Free=196.70822MB (206263512), Max=198.33751MB (207971952), Counts: Blocks=2, Access=56, Hit=54, Miss=2, Evictions=0, Evicted=0, Ratios: Hit Ratio=96.42857313156128%, Miss Ratio=3.57142873108387%, Evicted/Run=NaN\n2010-07-03 15:21:54,826 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scanning meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>}\n2010-07-03 15:21:54,830 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scan of 2 row(s) of meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>} complete\n2010-07-03 15:21:54,830 INFO org.apache.hadoop.hbase.master.BaseScanner: All 1 .META. region(s) scanned\n","from":"reporter","subject":"Disable 'table' fails to complete frustrating my ability to test easily"},{"body":"What I see a master that thinks that some region is hosted on a region server, that that region server receives the message but does nothing about it (in the scope of that log). Bigger log snippet and thread dump would really be helpful, or even a basic unit test that shows the issue.\n\nAlso the fix version is for 0.20.6 but you wrote that you used 0.89... so can we move it only to trunk in order to release 0.20.6? Thanks Sam.","from":"developer"},{"body":"Didn't see any objections for 2 days, moving to 0.90","from":"developer"},{"body":"I have also run into this, using 0.20.6. I have a test environment that runs a set of automated tests with our latest code every hour. Some of these tests are set to disable, drop, and recreate a table in order to make an empty one. After succeeding for 5 to 10 executions, it seems to get into this state. I'm not able to disable and get rid of the test table until hbase is restarted, so it is definitely hampering our ability to do automated testing with hbase.\n\nThe logs I have look similar, though the master and region server are separate processes with separate logs in my environment.\n\nDoes anyone have any ideas as to the cause of this? Is there any more information that would help in diagnosing?","from":"developer"},{"body":"Dave sent me the logs, and it really looks like a case of HBASE-2755:\n\n{noformat}\n2010-08-04 19:26:58,291 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_ReferralQueue,,1280975190155 to setClosing list\n2010-08-04 19:26:58,565 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scanning meta region {server: 192.168.21.66:60020, regionname: .META.,,1, startKey: <>}\n2010-08-04 19:26:58,654 INFO org.apache.hadoop.hbase.master.ServerManager: Processing MSG_REPORT_CLOSE: test_ReferralQueue,,1280975190155 from hslave3,60020,1280951333279; 1 of 1\n2010-08-04 19:26:58,654 DEBUG org.apache.hadoop.hbase.master.HMaster: Processing todo: ProcessRegionClose of test_ReferralQueue,,1280975190155, true, reassign: false\n2010-08-04 19:26:58,656 INFO org.apache.hadoop.hbase.master.ProcessRegionClose$1: region closed: test_ReferralQueue,,1280975190155\n2010-08-04 19:26:58,686 DEBUG org.apache.hadoop.hbase.master.BaseScanner: GET on test_ReferralQueue,,1280975190155 got different startcode than SCAN: sc=0, serverAddress=1280951333279\n2010-08-04 19:26:58,686 DEBUG org.apache.hadoop.hbase.master.BaseScanner: Current assignment of test_ReferralQueue,,1280975190155 is not valid; serverAddress=, startCode=0 unknown.\n{noformat}\n\nHere the base scanner saw that that region was half unassigned (atomicity problem?) since we don't see the message that \"sa\" is different. It means that the base scanner reassigns the region because it believes that:\n\n - the region isn't assigned to anyone\n - that it's not online\n - that it's not in transition\n\nThis means that we should also verify our assumption that the HRI.isOffline() is correct when we read it. I think we can augment HBASE-2755's patch to do that, since doing the GET we already read the new HRI. Making a patch.","from":"developer"},{"body":"Patch generated from 0.20.6 (so that Dave can test) that incorporates HBASE-2755 and adds the checking of the newly obtained HRI.","from":"developer"},{"body":"The automated tests have now run for close to 100 executions against 0.20.6 + this patch without any problems this time, so it appears to fix the issue. Is anyone up to review and commit it to the 0.20 branch?","from":"developer"},{"body":"So a review would be missing (although my patch is mostly 2755), but also the confirmation that it fixes Sam's problem. If not, then we should open a new jira and commit the fix there.","from":"developer"},{"body":"+1 Patch looks good to me.","from":"developer"},{"body":"Since I still didn't get the confirmation from Sam that this patch fixed his issue, I opened HBASE-2927 and posted the patch there. Keeping this jira open until we can tell for sure that the original issue is fixed.","from":"developer"},{"body":"Table enable/disable is different since hbase-2692 went in. Should make this a non-issue.","from":"developer"},{"body":"Moving into 0.20.7. This I believe is fixed in TRUNK (0.90).","from":"developer"}],"created":"2010-07-03T22:23:50.000+0000","description":"I see this in the client after it gives up:\n\nhbase(main):006:0> disable 'test_schema'\n\nERROR: org.apache.hadoop.hbase.RegionException: Retries exhausted, it took too long to wait for the table test_schema to be disabled.\n\nHere is some help for this command:\n Disable the named table: e.g. \"hbase> disable 't1'\"\n\nand this in the server log, a set of about 5 reports it is closing per disable call:\n\n2010-07-03 15:19:47,554 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:47,554 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:47,555 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:47,576 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:47,576 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:48,567 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:48,567 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:48,568 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:48,577 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:48,578 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:49,580 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:49,580 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:49,581 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:50,580 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:50,581 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:50,592 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:50,592 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:50,593 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:51,581 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:51,581 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:52,605 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:52,605 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:52,606 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:52,703 INFO org.apache.hadoop.hbase.master.ServerManager: 1 region servers, 0 dead, average load 3.0\n2010-07-03 15:19:52,863 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scanning meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>}\n2010-07-03 15:19:52,867 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scan of 1 row(s) of meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>} complete\n2010-07-03 15:19:53,585 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:53,585 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:53,671 DEBUG org.apache.hadoop.hbase.io.hfile.LruBlockCache: Cache Stats: Sizes: Total=1.6292953MB (1708440), Free=196.70822MB (206263512), Max=198.33751MB (207971952), Counts: Blocks=2, Access=52, Hit=50, Miss=2, Evictions=0, Evicted=0, Ratios: Hit Ratio=96.15384340286255%, Miss Ratio=3.8461539894342422%, Evicted/Run=NaN\n2010-07-03 15:19:54,617 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:54,617 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:54,618 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:54,821 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scanning meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>}\n2010-07-03 15:19:54,825 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scan of 2 row(s) of meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>} complete\n2010-07-03 15:19:54,825 INFO org.apache.hadoop.hbase.master.BaseScanner: All 1 .META. region(s) scanned\n2010-07-03 15:19:55,589 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:55,589 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:58,629 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing unserved regions\n2010-07-03 15:19:58,629 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Processing regions currently being served\n2010-07-03 15:19:58,630 DEBUG org.apache.hadoop.hbase.master.ChangeTableState: Adding region test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081. to setClosing list\n2010-07-03 15:19:59,595 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:19:59,595 INFO org.apache.hadoop.hbase.regionserver.HRegionServer: Worker: MSG_REGION_CLOSE: test_schema,,1278195322074.65c77aedf2f2a08d161a188dd2dd5081.\n2010-07-03 15:20:52,706 INFO org.apache.hadoop.hbase.master.ServerManager: 1 region servers, 0 dead, average load 3.0\n2010-07-03 15:20:52,866 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scanning meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>}\n2010-07-03 15:20:52,873 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scan of 1 row(s) of meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>} complete\n2010-07-03 15:20:53,671 DEBUG org.apache.hadoop.hbase.io.hfile.LruBlockCache: Cache Stats: Sizes: Total=1.6292953MB (1708440), Free=196.70822MB (206263512), Max=198.33751MB (207971952), Counts: Blocks=2, Access=54, Hit=52, Miss=2, Evictions=0, Evicted=0, Ratios: Hit Ratio=96.29629850387573%, Miss Ratio=3.7037037312984467%, Evicted/Run=NaN\n2010-07-03 15:20:54,823 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scanning meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>}\n2010-07-03 15:20:54,828 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scan of 2 row(s) of meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>} complete\n2010-07-03 15:20:54,828 INFO org.apache.hadoop.hbase.master.BaseScanner: All 1 .META. region(s) scanned\n2010-07-03 15:21:52,709 INFO org.apache.hadoop.hbase.master.ServerManager: 1 region servers, 0 dead, average load 3.0\n2010-07-03 15:21:52,869 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scanning meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>}\n2010-07-03 15:21:52,876 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.rootScanner scan of 1 row(s) of meta region {server: 192.168.2.1:54389, regionname: -ROOT-,,0.70236052, startKey: <>} complete\n2010-07-03 15:21:53,672 DEBUG org.apache.hadoop.hbase.io.hfile.LruBlockCache: Cache Stats: Sizes: Total=1.6292953MB (1708440), Free=196.70822MB (206263512), Max=198.33751MB (207971952), Counts: Blocks=2, Access=56, Hit=54, Miss=2, Evictions=0, Evicted=0, Ratios: Hit Ratio=96.42857313156128%, Miss Ratio=3.57142873108387%, Evicted/Run=NaN\n2010-07-03 15:21:54,826 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scanning meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>}\n2010-07-03 15:21:54,830 INFO org.apache.hadoop.hbase.master.BaseScanner: RegionManager.metaScanner scan of 2 row(s) of meta region {server: 192.168.2.1:54389, regionname: .META.,,1.1028785192, startKey: <>} complete\n2010-07-03 15:21:54,830 INFO org.apache.hadoop.hbase.master.BaseScanner: All 1 .META. region(s) scanned\n","issue_id":"12468502","key":"HBASE-2812","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2014-07-19T01:03:45.000+0000","role":"fixed_distractor","summary":"Disable 'table' fails to complete frustrating my ability to test easily"} {"case_id":"13570796","cluster":"DISTRACTOR-HBASE-28420","comments":[{"body":"I do not fully understand the problem...\r\n\r\nDoes the new 'active' master actually finish the active master initialization?\r\nIf so, I think region server will start to report to the new active master, as the old active master will fail to persistent to procedure store and the procedure report will fail.\r\n\r\nIf not, then the problem is why the old active master hangs there for 1 hours, without letting other active masters take the charge...","created":"2024-03-05T09:52:50.271+0000"},{"body":">>Does the new 'active' master actually finish the active master initialization?\r\n\r\nThe new 'active' master was not able to finish the initialization. It got stuck while starting AM ([code|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/master/HMaster.java#L1069]). Before that, it has started other master services. \r\n\r\n>>If so, I think region server will start to report to the new active master, as the old active master will fail to persistent to procedure store and the procedure report will fail.\r\n\r\nRegion servers always try to report old master only till they don't get an error. First, they check if master services are running([code|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/master/MasterRpcServices.java#L2381]) which was true in this case as I don't see us setting this flag to false while aborting the master. Later I didn't see any failure in reporting, maybe there is no persistence while reporting([code|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/master/procedure/ServerRemoteProcedure.java#L130]), I see it waking up an event.  \r\n\r\n>>If not, then the problem is why the old active master hangs there for 1 hour, without letting other active masters take the charge...\r\n\r\nThe Old active master didn't hang up. It finished the abortion in 30 seconds. But before finishing it acknowledged the report for remoteProcedureDone which should be acknowledged by the new active master to process further. ","created":"2024-03-05T14:13:49.537+0000"},{"body":"{quote}\r\nThe Old active master didn't hang up. It finished the abortion in 30 seconds. But before finishing it acknowledged the report for remoteProcedureDone which should be acknowledged by the new active master to process further. \r\n{quote}\r\n\r\nAs I said above, if a master has already aborted, then it can not accept the report from region servers, as it will fail when persist the state to procedure store. When becoming the active master, the new active master will call recoverLease to finish the procedure store's wal so the old master can not write to it any more, this is the fencing way here.\r\n\r\nSo if you find out that the old active master can still accept report from region servers, then it is not dead yet. There is no problem for it to accept the report. Even if it is down immediately after persisting the state, the new active master will load the persisted state and move the procedure forward.\r\n\r\nSo I still do not fully understand what is going on here, why the old active master does not quit as it is aborted? Why the new active master hang when initializing AM? Because meta not online? What is the state of the SCP for the region server which holds the meta region?","created":"2024-03-06T02:52:00.062+0000"},{"body":"OK, the SCP was scheduled by the new master, but the region server still reported to the old master? As I said above, in this scenario, the report to old master will fail because it can not persist the state to procedure store, since the new active master has already done the fencing on procedure store.\r\n\r\nYou can see the code in MasterRegion.replayWALs, it uses the same way with our normal fencing on region server crashes.\r\n\r\nSo I think the problem here is that, how could the old master finishes the remoteProcedureDone, while new active master has already done the fencing on master region.","created":"2024-03-06T03:03:25.595+0000"},{"body":"Oh, it is SplitWALRemoteProcedure procedure done, not TRSP. Maybe the problem is that, we do not persist anything when accepting SplitWALRemoteProcedure's reportProcedureDone...","created":"2024-03-06T03:05:12.159+0000"},{"body":"Confirmed, this is what we have in SplitWALRemoteProcedure.complete method\r\n\r\n{code}\r\n @Override\r\n protected void complete(MasterProcedureEnv env, Throwable error) {\r\n if (error == null) {\r\n try {\r\n env.getMasterServices().getSplitWALManager().archive(walPath);\r\n } catch (IOException e) {\r\n LOG.warn(\"Failed split of {}; ignore...\", walPath, e);\r\n }\r\n succ = true;\r\n } else {\r\n if (error instanceof DoNotRetryIOException) {\r\n LOG.warn(\"Sent {} to wrong server {}, try another\", walPath, targetServer, error);\r\n succ = true;\r\n } else {\r\n LOG.warn(\"Failed split of {}, retry...\", walPath, error);\r\n succ = false;\r\n }\r\n }\r\n }\r\n{code}\r\n\r\nWe do not persist anything into the procedure store, so even the new active master has already finished fencing, the old master could still accept the report, and cause the new active master does not know the finishing of a procedure.\r\n\r\nThis should be a common problem for remote procedures other than region assigning, we should have a unified solution to fix them all.\r\n\r\nThanks.","created":"2024-03-06T03:13:15.876+0000"},{"body":"Before proceeding further we do checkServiceStarted(). I was thinking of adding a check for aborted as well. \r\n{noformat}\r\n@Override\r\npublic ReportProcedureDoneResponse reportProcedureDone(RpcController controller,\r\n ReportProcedureDoneRequest request) throws ServiceException {\r\n // Check Masters is up and ready for duty before progressing. Remote side will keep trying.\r\n try {\r\n this.master.checkServiceStarted();\r\n } catch (ServerNotRunningYetException snrye) {\r\n throw new ServiceException(snrye);\r\n }\r\n.\r\n.\r\n.\r\n\r\n@InterfaceAudience.Private\r\nprotected void checkServiceStarted() throws ServerNotRunningYetException {\r\n if (!serviceStarted) {\r\n throw new ServerNotRunningYetException(\"Server is not running yet\");\r\n }\r\n}{noformat}\r\nlike \r\n{noformat}\r\n@InterfaceAudience.Private\r\nprotected void checkServiceStarted() throws ServerNotRunningYetException {\r\n if (!serviceStarted || isAborted() || isStopped()) {\r\n throw new ServerNotRunningYetException(\"Server is not running yet\");\r\n }\r\n}\r\n{noformat}\r\nDo you think this would be the right solution? ","created":"2024-03-06T03:36:48.295+0000"},{"body":"No, this can not fix all the problems.\r\n\r\nConsider the old active master has a long full gc, and the new active master has alredy started and scheduled a SCP, just like the scenario here. And when the region server wants to call the reportProcedureDone, the old active master recovered, and before the session expired exception triggered the abortion(it is in a background thread), the old active master receives the reportProcedureDone and finishes it. Then we will hang there for a very long time, just like here.","created":"2024-03-06T04:05:43.282+0000"},{"body":"[~umesh9414] Do you have any plans here? If no clear solution yet, I could write a simple POC PR to show how to fix the problem. WDYT?\r\n\r\nThanks.","created":"2024-03-25T14:06:16.525+0000"},{"body":"[~umesh9414] i think the issue here is that new master doesn't know that a procedure is scheduled as it is not stored in procedure store. A simple check might not fix it. A proper solution is to store this type(SplitWALRemoteProcedure) of procedure in proc store so when a new master comes up it reads the proc store for current ongoing procs and gets to know that there is a remote proc schedued and it needs to check for the progress of that. In the above case new master doesn't even know that a procedure was scheduled by old master.","created":"2024-03-26T09:20:51.233+0000"},{"body":"[~mnpoonia] in this case, the new master only scheduled all the procedures, so unknown to the master is not the case here. \r\n\r\n[~zhangduo] I am planning to solve this. I don't have any clear solution. Should we add some timed checks on the master side or simply add some persistence on the master side?","created":"2024-03-26T09:41:24.169+0000"},{"body":"[~umesh9414] Thank you for the clarification.\r\n\r\nI was wondering what happens if the new master also restarts in between and another host becomes master or is in process of becoming active master?","created":"2024-03-26T10:08:48.874+0000"},{"body":"Summarising the offfline discussion with [~umesh9414] \r\n\r\nMost procedures are stored in proc store. So if we start storing this proc also in store and when we get a report from RS we will try to update the state to success in the store. Since active master has changed, the old master would not be able to update the store. This will ensure that old master doesn't process the request and instead throws an exception.\r\n\r\n[~zhangduo] Are you suggesting a similar approach?\r\n\r\n[~umesh9414]  do you want to create a draft PR to give an idea of what we are planning?","created":"2024-03-26T11:44:22.974+0000"},{"body":"You can see the code in RegionRemoteProcedureBase, what we have done in reportTransition. We should do the same in ServerRemoteProcedure, in the remoteOperationDone method, we should update the procedure's field to store the result, and then wake up the procedure to continue the later processing.","created":"2024-03-26T11:55:39.341+0000"},{"body":"A nasty problem here is that, ServerRemoteProcedure does not has its own serializing implementation, so we need to persist some states in all the sub procedures, which means we need to touch multiple protobuf messages, and write some redundant code, but this is a must if we want to keep compatibility...","created":"2024-03-26T11:57:08.884+0000"},{"body":"I take a look at this. I checked reportTransition method of RegionRemoteProcedureBase and remoteOperationDone of ServerRemoteProcedure. \r\n\r\nI would need more understanding of protobuf as well. Working on it. ","created":"2024-03-29T20:30:47.308+0000"},{"body":"Any updates here? [~umesh9414]","created":"2024-04-09T06:30:08.017+0000"},{"body":"I've pushed the patch to master and branch-3.\r\n\r\n[~umesh9414] Please provide a PR for branch-2? The patch can not be applied to branch-2 cleanly.\r\n\r\nThanks.","created":"2024-05-31T15:51:44.600+0000"},{"body":"Pushed to all active branches.\r\n\r\nThanks [~umesh9414] for contributing!","created":"2024-06-05T14:42:26.402+0000"},{"body":"Happy to contribute :)","created":"2024-06-05T15:49:49.058+0000"}],"conversations":[{"body":"When the Active Hmaster is in the process of abortion and another HMaster is becoming Active HMaster,at the same time if any region server reports the completion of the remote procedure, it generally goes to the old active HMaster because of the cached value of rssStub -> [code|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java#L2829] ([caller method|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java#L3941]). On the Master side ([code|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/master/MasterRpcServices.java#L2381]), It did check if the service is started but that returns true if the master is in the process of abortion(I didn't see when we are setting this flag false while abortion).  \r\n\r\nThis issue becomes *critical* when *ServerCrash of meta hosting RS and master failover* happens at the same time and hbase:meta got stuck in the offline state.\r\n\r\nLogs for abortion start of HMaster \r\n{noformat}\r\n2024-02-02 07:33:11,581 ERROR [PEWorker-6] master.HMaster - ***** ABORTING master server4-1xxx,61000,1705169084562:\r\nFAILED persisting region=52d36581218e00a2668776cfea897132 state=CLOSING *****{noformat}\r\n{noformat}\r\n2024-02-02 07:33:40,999 INFO [master/server4-1xxx:61000] regionserver.HRegionServer - Exiting; \r\nstopping=hbase2b-mnds4-1-ia2.ops.sfdc.net,61000,1705169084562; zookeeper connection closed.{noformat}\r\nit took almost 30 seconds to abort the HMaster.\r\n\r\n \r\n\r\nLogs of starting SCP for meta carrying host. (This SCP is started by the new active HMaster)\r\n{noformat}\r\n2024-02-02 07:33:32,622 INFO [aster/server3-1xxx61000:becomeActiveMaster] assignment.AssignmentManager - Scheduled\r\nServerCrashProcedure pid=3305546 for server5-1xxx61020,1706857451955 (carryingMeta=true) server5-1-\r\nxxx61020,1706857451955/CRASHED/regionCount=1/lock=java.util.concurrent.locks.ReentrantReadWriteLock@1b0a5293[Write \r\nlocks = 1, Read locks = 0], oldState=ONLINE.{noformat}\r\ninitialization of remote procedure\r\n{noformat}\r\n2024-02-02 07:33:33,178 INFO [PEWorker-4] procedure2.ProcedureExecutor - Initialized subprocedures=[{pid=3305548, \r\nppid=3305547, state=RUNNABLE; SplitWALRemoteProcedure server5-1-\r\nxxxxt%2C61020%2C1706857451955.meta.1706858156058.meta, worker=server4-1-xxxx,61020,1705169180881}]{noformat}\r\nLogs of remote procedure handling on Old Active Hmaster(server4-1xxx,61000) (in the process of abortion)\r\n{noformat}\r\n2024-02-02 07:33:37,990 DEBUG [r.default.FPBQ.Fifo.handler=243,queue=9,port=61000] master.HMaster - Remote procedure \r\ndone, pid=3305548{noformat}\r\nThis should be handled by the new active HMaster so that it can wake up the suspended Procedure on the new Active Hmaster. As the new ActiveHMaster was not able to wake that up, SCP procedure got stuck thus meta stayed OFFLINE. \r\n\r\n \r\n\r\nLogs of Hmaster trying to becomeActivehmaster but stuck-\r\n{noformat}\r\n2024-02-02 07:33:43,159 WARN [aster/server3-1-ia2:61000:becomeActiveMaster] master.HMaster - hbase:meta,,1.1588230740 \r\nis NOT online; state={1588230740 state=OPEN, ts=1706859212481, server=server5-1-xxx,61020,1706857451955}; \r\nServerCrashProcedures=true. Master startup cannot progress, in holding-pattern until region onlined.{noformat}\r\nAfter this master was stuck till we did hmaster failover to come out of this situation. ","from":"reporter","subject":"Aborting Active HMaster is not rejecting remote Procedure Reports"},{"body":"I do not fully understand the problem...\r\n\r\nDoes the new 'active' master actually finish the active master initialization?\r\nIf so, I think region server will start to report to the new active master, as the old active master will fail to persistent to procedure store and the procedure report will fail.\r\n\r\nIf not, then the problem is why the old active master hangs there for 1 hours, without letting other active masters take the charge...","from":"developer"},{"body":">>Does the new 'active' master actually finish the active master initialization?\r\n\r\nThe new 'active' master was not able to finish the initialization. It got stuck while starting AM ([code|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/master/HMaster.java#L1069]). Before that, it has started other master services. \r\n\r\n>>If so, I think region server will start to report to the new active master, as the old active master will fail to persistent to procedure store and the procedure report will fail.\r\n\r\nRegion servers always try to report old master only till they don't get an error. First, they check if master services are running([code|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/master/MasterRpcServices.java#L2381]) which was true in this case as I don't see us setting this flag to false while aborting the master. Later I didn't see any failure in reporting, maybe there is no persistence while reporting([code|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/master/procedure/ServerRemoteProcedure.java#L130]), I see it waking up an event.  \r\n\r\n>>If not, then the problem is why the old active master hangs there for 1 hour, without letting other active masters take the charge...\r\n\r\nThe Old active master didn't hang up. It finished the abortion in 30 seconds. But before finishing it acknowledged the report for remoteProcedureDone which should be acknowledged by the new active master to process further. ","from":"developer"},{"body":"{quote}\r\nThe Old active master didn't hang up. It finished the abortion in 30 seconds. But before finishing it acknowledged the report for remoteProcedureDone which should be acknowledged by the new active master to process further. \r\n{quote}\r\n\r\nAs I said above, if a master has already aborted, then it can not accept the report from region servers, as it will fail when persist the state to procedure store. When becoming the active master, the new active master will call recoverLease to finish the procedure store's wal so the old master can not write to it any more, this is the fencing way here.\r\n\r\nSo if you find out that the old active master can still accept report from region servers, then it is not dead yet. There is no problem for it to accept the report. Even if it is down immediately after persisting the state, the new active master will load the persisted state and move the procedure forward.\r\n\r\nSo I still do not fully understand what is going on here, why the old active master does not quit as it is aborted? Why the new active master hang when initializing AM? Because meta not online? What is the state of the SCP for the region server which holds the meta region?","from":"developer"},{"body":"OK, the SCP was scheduled by the new master, but the region server still reported to the old master? As I said above, in this scenario, the report to old master will fail because it can not persist the state to procedure store, since the new active master has already done the fencing on procedure store.\r\n\r\nYou can see the code in MasterRegion.replayWALs, it uses the same way with our normal fencing on region server crashes.\r\n\r\nSo I think the problem here is that, how could the old master finishes the remoteProcedureDone, while new active master has already done the fencing on master region.","from":"developer"},{"body":"Oh, it is SplitWALRemoteProcedure procedure done, not TRSP. Maybe the problem is that, we do not persist anything when accepting SplitWALRemoteProcedure's reportProcedureDone...","from":"developer"},{"body":"Confirmed, this is what we have in SplitWALRemoteProcedure.complete method\r\n\r\n{code}\r\n @Override\r\n protected void complete(MasterProcedureEnv env, Throwable error) {\r\n if (error == null) {\r\n try {\r\n env.getMasterServices().getSplitWALManager().archive(walPath);\r\n } catch (IOException e) {\r\n LOG.warn(\"Failed split of {}; ignore...\", walPath, e);\r\n }\r\n succ = true;\r\n } else {\r\n if (error instanceof DoNotRetryIOException) {\r\n LOG.warn(\"Sent {} to wrong server {}, try another\", walPath, targetServer, error);\r\n succ = true;\r\n } else {\r\n LOG.warn(\"Failed split of {}, retry...\", walPath, error);\r\n succ = false;\r\n }\r\n }\r\n }\r\n{code}\r\n\r\nWe do not persist anything into the procedure store, so even the new active master has already finished fencing, the old master could still accept the report, and cause the new active master does not know the finishing of a procedure.\r\n\r\nThis should be a common problem for remote procedures other than region assigning, we should have a unified solution to fix them all.\r\n\r\nThanks.","from":"developer"},{"body":"Before proceeding further we do checkServiceStarted(). I was thinking of adding a check for aborted as well. \r\n{noformat}\r\n@Override\r\npublic ReportProcedureDoneResponse reportProcedureDone(RpcController controller,\r\n ReportProcedureDoneRequest request) throws ServiceException {\r\n // Check Masters is up and ready for duty before progressing. Remote side will keep trying.\r\n try {\r\n this.master.checkServiceStarted();\r\n } catch (ServerNotRunningYetException snrye) {\r\n throw new ServiceException(snrye);\r\n }\r\n.\r\n.\r\n.\r\n\r\n@InterfaceAudience.Private\r\nprotected void checkServiceStarted() throws ServerNotRunningYetException {\r\n if (!serviceStarted) {\r\n throw new ServerNotRunningYetException(\"Server is not running yet\");\r\n }\r\n}{noformat}\r\nlike \r\n{noformat}\r\n@InterfaceAudience.Private\r\nprotected void checkServiceStarted() throws ServerNotRunningYetException {\r\n if (!serviceStarted || isAborted() || isStopped()) {\r\n throw new ServerNotRunningYetException(\"Server is not running yet\");\r\n }\r\n}\r\n{noformat}\r\nDo you think this would be the right solution? ","from":"developer"},{"body":"No, this can not fix all the problems.\r\n\r\nConsider the old active master has a long full gc, and the new active master has alredy started and scheduled a SCP, just like the scenario here. And when the region server wants to call the reportProcedureDone, the old active master recovered, and before the session expired exception triggered the abortion(it is in a background thread), the old active master receives the reportProcedureDone and finishes it. Then we will hang there for a very long time, just like here.","from":"developer"},{"body":"[~umesh9414] Do you have any plans here? If no clear solution yet, I could write a simple POC PR to show how to fix the problem. WDYT?\r\n\r\nThanks.","from":"developer"},{"body":"[~umesh9414] i think the issue here is that new master doesn't know that a procedure is scheduled as it is not stored in procedure store. A simple check might not fix it. A proper solution is to store this type(SplitWALRemoteProcedure) of procedure in proc store so when a new master comes up it reads the proc store for current ongoing procs and gets to know that there is a remote proc schedued and it needs to check for the progress of that. In the above case new master doesn't even know that a procedure was scheduled by old master.","from":"developer"},{"body":"[~mnpoonia] in this case, the new master only scheduled all the procedures, so unknown to the master is not the case here. \r\n\r\n[~zhangduo] I am planning to solve this. I don't have any clear solution. Should we add some timed checks on the master side or simply add some persistence on the master side?","from":"developer"},{"body":"[~umesh9414] Thank you for the clarification.\r\n\r\nI was wondering what happens if the new master also restarts in between and another host becomes master or is in process of becoming active master?","from":"developer"},{"body":"Summarising the offfline discussion with [~umesh9414] \r\n\r\nMost procedures are stored in proc store. So if we start storing this proc also in store and when we get a report from RS we will try to update the state to success in the store. Since active master has changed, the old master would not be able to update the store. This will ensure that old master doesn't process the request and instead throws an exception.\r\n\r\n[~zhangduo] Are you suggesting a similar approach?\r\n\r\n[~umesh9414]  do you want to create a draft PR to give an idea of what we are planning?","from":"developer"},{"body":"You can see the code in RegionRemoteProcedureBase, what we have done in reportTransition. We should do the same in ServerRemoteProcedure, in the remoteOperationDone method, we should update the procedure's field to store the result, and then wake up the procedure to continue the later processing.","from":"developer"},{"body":"A nasty problem here is that, ServerRemoteProcedure does not has its own serializing implementation, so we need to persist some states in all the sub procedures, which means we need to touch multiple protobuf messages, and write some redundant code, but this is a must if we want to keep compatibility...","from":"developer"},{"body":"I take a look at this. I checked reportTransition method of RegionRemoteProcedureBase and remoteOperationDone of ServerRemoteProcedure. \r\n\r\nI would need more understanding of protobuf as well. Working on it. ","from":"developer"},{"body":"Any updates here? [~umesh9414]","from":"developer"},{"body":"I've pushed the patch to master and branch-3.\r\n\r\n[~umesh9414] Please provide a PR for branch-2? The patch can not be applied to branch-2 cleanly.\r\n\r\nThanks.","from":"developer"},{"body":"Pushed to all active branches.\r\n\r\nThanks [~umesh9414] for contributing!","from":"developer"},{"body":"Happy to contribute :)","from":"developer"}],"created":"2024-03-05T08:53:45.000+0000","description":"When the Active Hmaster is in the process of abortion and another HMaster is becoming Active HMaster,at the same time if any region server reports the completion of the remote procedure, it generally goes to the old active HMaster because of the cached value of rssStub -> [code|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java#L2829] ([caller method|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegionServer.java#L3941]). On the Master side ([code|https://github.com/apache/hbase/blob/branch-2.5/hbase-server/src/main/java/org/apache/hadoop/hbase/master/MasterRpcServices.java#L2381]), It did check if the service is started but that returns true if the master is in the process of abortion(I didn't see when we are setting this flag false while abortion).  \r\n\r\nThis issue becomes *critical* when *ServerCrash of meta hosting RS and master failover* happens at the same time and hbase:meta got stuck in the offline state.\r\n\r\nLogs for abortion start of HMaster \r\n{noformat}\r\n2024-02-02 07:33:11,581 ERROR [PEWorker-6] master.HMaster - ***** ABORTING master server4-1xxx,61000,1705169084562:\r\nFAILED persisting region=52d36581218e00a2668776cfea897132 state=CLOSING *****{noformat}\r\n{noformat}\r\n2024-02-02 07:33:40,999 INFO [master/server4-1xxx:61000] regionserver.HRegionServer - Exiting; \r\nstopping=hbase2b-mnds4-1-ia2.ops.sfdc.net,61000,1705169084562; zookeeper connection closed.{noformat}\r\nit took almost 30 seconds to abort the HMaster.\r\n\r\n \r\n\r\nLogs of starting SCP for meta carrying host. (This SCP is started by the new active HMaster)\r\n{noformat}\r\n2024-02-02 07:33:32,622 INFO [aster/server3-1xxx61000:becomeActiveMaster] assignment.AssignmentManager - Scheduled\r\nServerCrashProcedure pid=3305546 for server5-1xxx61020,1706857451955 (carryingMeta=true) server5-1-\r\nxxx61020,1706857451955/CRASHED/regionCount=1/lock=java.util.concurrent.locks.ReentrantReadWriteLock@1b0a5293[Write \r\nlocks = 1, Read locks = 0], oldState=ONLINE.{noformat}\r\ninitialization of remote procedure\r\n{noformat}\r\n2024-02-02 07:33:33,178 INFO [PEWorker-4] procedure2.ProcedureExecutor - Initialized subprocedures=[{pid=3305548, \r\nppid=3305547, state=RUNNABLE; SplitWALRemoteProcedure server5-1-\r\nxxxxt%2C61020%2C1706857451955.meta.1706858156058.meta, worker=server4-1-xxxx,61020,1705169180881}]{noformat}\r\nLogs of remote procedure handling on Old Active Hmaster(server4-1xxx,61000) (in the process of abortion)\r\n{noformat}\r\n2024-02-02 07:33:37,990 DEBUG [r.default.FPBQ.Fifo.handler=243,queue=9,port=61000] master.HMaster - Remote procedure \r\ndone, pid=3305548{noformat}\r\nThis should be handled by the new active HMaster so that it can wake up the suspended Procedure on the new Active Hmaster. As the new ActiveHMaster was not able to wake that up, SCP procedure got stuck thus meta stayed OFFLINE. \r\n\r\n \r\n\r\nLogs of Hmaster trying to becomeActivehmaster but stuck-\r\n{noformat}\r\n2024-02-02 07:33:43,159 WARN [aster/server3-1-ia2:61000:becomeActiveMaster] master.HMaster - hbase:meta,,1.1588230740 \r\nis NOT online; state={1588230740 state=OPEN, ts=1706859212481, server=server5-1-xxx,61020,1706857451955}; \r\nServerCrashProcedures=true. Master startup cannot progress, in holding-pattern until region onlined.{noformat}\r\nAfter this master was stuck till we did hmaster failover to come out of this situation. ","issue_id":"13570796","key":"HBASE-28420","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2024-06-05T14:42:26.000+0000","role":"fixed_distractor","summary":"Aborting Active HMaster is not rejecting remote Procedure Reports"} {"case_id":"12469634","cluster":"DISTRACTOR-HBASE-2849","comments":[{"body":"Reformatting a little bit.","created":"2010-07-19T17:02:21.489+0000"},{"body":"Client+ZK interaction just got a makeover in the master rewrite branch. It now handles master failovers properly and re-initializes zk connections but I'm not sure it will sustain a zk restart. Any chance of a reproducing unit test?","created":"2010-07-19T17:32:13.551+0000"},{"body":"http://hadoop.apache.org/zookeeper/docs/r3.3.1/api/org/apache/zookeeper/ZooKeeper.html\nbq. If for some reason, the client fails to send heart beats to the server for a prolonged period of time (exceeding the sessionTimeout value, for instance), the server will expire the session, and the session ID will become invalid. The client object will no longer be usable. To make ZooKeeper API calls, the application must create a new client object.\n\nSo apparently, a new {{ZooKeeper}} object must be created when the session becomes invalid. This sounds like a bad API, not sure why they did it this way. In HBase's source code, it seems that the only thing that creates a {{ZooKeeper}} instance is in {{ZooKeeperWrapper#reconnectToZk}}. This method, although it's public, is only called from 3 other methods in that class: the constructor, {{exists}} and {{deleteUnassignedRegion}}. The latter, {{deleteUnassignedRegion}}, is only used by the master. The former, {{exists}}, is only called from the following locations:\n* {{ZKUnassignedWatcher}}'s constructor. This is only used in the master.\n* {{RSZookeeperUpdater#startRegionCloseEvent}}. This is only used in the region server.\n* {{ZooKeeperWrapper#createOrUpdateUnassignedRegion}}. This is only used by the master's {{RegionManager}}.\n* {{ZooKeeperWrapper#createUnassignedRegion}} and {{ZooKeeperWrapper#updateUnassignedRegion}}. Those two methods, even though they're public, are only called from {{ZooKeeperWrapper#createOrUpdateUnassignedRegion}}, which itself is only used by the master's {{RegionManager}}.\n\nIn other words, for someone writing an HBase application, only a single {{ZooKeeper}} instance gets created when the {{ZooKeeperWrapper}} is instantiated. Any failure that causes the client's session to become invalid will is unrecoverable with the current code and the client has to be killed and restarted.\n\nJonathan, is the work being done for the master rewrite branch going to address this issue? Bear in mind that here I'm concerned about HBase *client* applications.","created":"2010-07-23T23:01:50.854+0000"},{"body":"It would be possible for the client to try to reconnect after being expired. It's not built-in to the way I have it now but it's possible to add it.","created":"2010-07-23T23:06:21.657+0000"},{"body":"Patch that fixes the issue. Actually there was some logic I didn't notice earlier in {{HConnectionManager}} to attempt to deal with ZK failures and reconnect when needed, but the code wasn't doing the right thing and didn't work when there was a disconnection between the HBase client and the ZK quorum. So the patch is rather simple and consists in fixing the existing logic in {{HConnectionManager.ClientZKWatcher}}.\n\nI tested this by starting a long running HBase application, killing the whole ZooKeeper ensemble and restarting it. The application experiences a hiccup while ZK is unavailable and is able to recover automatically soon after the ZK quorum is back online. Someone else is more than welcome to write a unit test that simulates this scenario if they feel like it.","created":"2010-07-24T03:58:09.421+0000"},{"body":"Committed. Thats for the 'duh' patch Benôit.","created":"2010-07-24T05:19:53.765+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T12:40:34.339+0000"}],"conversations":[{"body":"Someone made mention of this loop last week but I don't think I filed an issue. Here is another instance, again from a secret hbase admirer:\n\n\"It seems that when Zookeeper dies and restarts, all client applications need to be restarted too. I just restarted HBase in non-distributed mode (which includes a ZK) and now my application can't reconnect to ZK unless I restart it too. I'm stuck in this loop:\n\n{code}\n2010-07-19 00:13:05,725 INFO org.apache.zookeeper.server.NIOServerCnxn:\n Closed socket connection for client /127.0.0.1:55153 (no session established for client)\n2010-07-19 00:13:07,052 INFO org.apache.zookeeper.server.NIOServerCnxn:\n Accepted socket connection from /127.0.0.1:55154\n2010-07-19 00:13:07,053 INFO org.apache.zookeeper.server.NIOServerCnxn:\n Refusing session request for client /127.0.0.1:55154 as it has seen zxid 0xf5 our last zxid is 0xd7\n client must try another server\n{code}\n\"","from":"reporter","subject":"HBase clients cannot recover when their ZooKeeper session becomes invalid"},{"body":"Reformatting a little bit.","from":"developer"},{"body":"Client+ZK interaction just got a makeover in the master rewrite branch. It now handles master failovers properly and re-initializes zk connections but I'm not sure it will sustain a zk restart. Any chance of a reproducing unit test?","from":"developer"},{"body":"http://hadoop.apache.org/zookeeper/docs/r3.3.1/api/org/apache/zookeeper/ZooKeeper.html\nbq. If for some reason, the client fails to send heart beats to the server for a prolonged period of time (exceeding the sessionTimeout value, for instance), the server will expire the session, and the session ID will become invalid. The client object will no longer be usable. To make ZooKeeper API calls, the application must create a new client object.\n\nSo apparently, a new {{ZooKeeper}} object must be created when the session becomes invalid. This sounds like a bad API, not sure why they did it this way. In HBase's source code, it seems that the only thing that creates a {{ZooKeeper}} instance is in {{ZooKeeperWrapper#reconnectToZk}}. This method, although it's public, is only called from 3 other methods in that class: the constructor, {{exists}} and {{deleteUnassignedRegion}}. The latter, {{deleteUnassignedRegion}}, is only used by the master. The former, {{exists}}, is only called from the following locations:\n* {{ZKUnassignedWatcher}}'s constructor. This is only used in the master.\n* {{RSZookeeperUpdater#startRegionCloseEvent}}. This is only used in the region server.\n* {{ZooKeeperWrapper#createOrUpdateUnassignedRegion}}. This is only used by the master's {{RegionManager}}.\n* {{ZooKeeperWrapper#createUnassignedRegion}} and {{ZooKeeperWrapper#updateUnassignedRegion}}. Those two methods, even though they're public, are only called from {{ZooKeeperWrapper#createOrUpdateUnassignedRegion}}, which itself is only used by the master's {{RegionManager}}.\n\nIn other words, for someone writing an HBase application, only a single {{ZooKeeper}} instance gets created when the {{ZooKeeperWrapper}} is instantiated. Any failure that causes the client's session to become invalid will is unrecoverable with the current code and the client has to be killed and restarted.\n\nJonathan, is the work being done for the master rewrite branch going to address this issue? Bear in mind that here I'm concerned about HBase *client* applications.","from":"developer"},{"body":"It would be possible for the client to try to reconnect after being expired. It's not built-in to the way I have it now but it's possible to add it.","from":"developer"},{"body":"Patch that fixes the issue. Actually there was some logic I didn't notice earlier in {{HConnectionManager}} to attempt to deal with ZK failures and reconnect when needed, but the code wasn't doing the right thing and didn't work when there was a disconnection between the HBase client and the ZK quorum. So the patch is rather simple and consists in fixing the existing logic in {{HConnectionManager.ClientZKWatcher}}.\n\nI tested this by starting a long running HBase application, killing the whole ZooKeeper ensemble and restarting it. The application experiences a hiccup while ZK is unavailable and is able to recover automatically soon after the ZK quorum is back online. Someone else is more than welcome to write a unit test that simulates this scenario if they feel like it.","from":"developer"},{"body":"Committed. Thats for the 'duh' patch Benôit.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2010-07-19T16:52:10.000+0000","description":"Someone made mention of this loop last week but I don't think I filed an issue. Here is another instance, again from a secret hbase admirer:\n\n\"It seems that when Zookeeper dies and restarts, all client applications need to be restarted too. I just restarted HBase in non-distributed mode (which includes a ZK) and now my application can't reconnect to ZK unless I restart it too. I'm stuck in this loop:\n\n{code}\n2010-07-19 00:13:05,725 INFO org.apache.zookeeper.server.NIOServerCnxn:\n Closed socket connection for client /127.0.0.1:55153 (no session established for client)\n2010-07-19 00:13:07,052 INFO org.apache.zookeeper.server.NIOServerCnxn:\n Accepted socket connection from /127.0.0.1:55154\n2010-07-19 00:13:07,053 INFO org.apache.zookeeper.server.NIOServerCnxn:\n Refusing session request for client /127.0.0.1:55154 as it has seen zxid 0xf5 our last zxid is 0xd7\n client must try another server\n{code}\n\"","issue_id":"12469634","key":"HBASE-2849","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2010-07-24T05:19:53.000+0000","role":"fixed_distractor","summary":"HBase clients cannot recover when their ZooKeeper session becomes invalid"} {"case_id":"13578273","cluster":"DISTRACTOR-HBASE-28569","comments":[{"body":"I think the problem here is that, if the join has been interrupted, we should not consider the result as correct any more?","created":"2024-05-06T13:02:19.387+0000"},{"body":"[~zhangduo] I've added a PR to address this - https://github.com/apache/hbase/pull/6266","created":"2024-10-09T23:16:43.464+0000"},{"body":"Would it be possible to get some feedback on the PR? We're seeing an uptick of instances of this bug.","created":"2025-03-12T17:04:52.453+0000"},{"body":"Pushed to all active branches.\r\n\r\nThanks [~claireiacono] for contributing!","created":"2025-04-07T15:50:08.574+0000"}],"conversations":[{"body":"There is a race condition that can happen when a regionserver aborts initialisation while splitting a WAL from another regionserver. This race leads to writing the WAL trailer for recovered edits while the writer threads are still running, thus the trailer gets interleaved with the edits corrupting the recovered edits file (and preventing the region to be assigned).\r\nWe've seen this happening on HBase 2.4.17, but looking at the latest code it seems that the race can still happen there.\r\nThe sequence of operations that leads to this issue:\r\n * {{org.apache.hadoop.hbase.wal.WALSplitter.splitWAL}} calls {{outputSink.close()}} after adding all the entries to the buffers\r\n * The output sink is {{org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink}} and its {{close}} method calls first {{finishWriterThreads}} in a try block which in turn will call {{finish}} on every thread and then join it to make sure it's done.\r\n * However if the splitter thread gets interrupted because of RS aborting, the join will get interrupted and {{finishWriterThreads}} will rethrow without waiting for the writer threads to stop.\r\n * This is problematic because coming back to {{org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink.close}} it will call {{closeWriters}} in a finally block (so it will execute even when the join was interrupted).\r\n * {{closeWriters}} will call {{org.apache.hadoop.hbase.wal.AbstractRecoveredEditsOutputSink.closeRecoveredEditsWriter}} which will call {{close}} on {{{}editWriter.writer{}}}.\r\n * When {{editWriter.writer}} is {{{}org.apache.hadoop.hbase.regionserver.wal.ProtobufLogWriter{}}}, its {{close}} method will write the trailer before closing the file.\r\n * This trailer write will now go in parallel with writer threads writing entries causing corruption.\r\n * If there are no other errors, {{closeWriters}} will succeed renaming all temporary files to final recovered edits, causing problems next time the region is assigned.\r\n\r\nLogs evidence supporting the above flow:\r\nAbort is triggered (because it failed to open the WAL due to some ongoing infra issue):\r\n{noformat}\r\nregionserver-2 regionserver 06:22:00.384 [RS_OPEN_META-regionserver/host01:16201-0] ERROR org.apache.hadoop.hbase.regionserver.HRegionServer - ***** ABORTING region server host01,16201,1709187641249: WAL can not clean up after init failed *****{noformat}\r\nWe can see that the writer threads were still active after closing (even considering that the\r\nordering in the log might not be accurate, we see that they die because the channel is closed while still writing, not because they're stopping):\r\n{noformat}\r\nregionserver-2 regionserver 06:22:09.662 [DataStreamer for file /hbase/data/default/aeris_v2/53308260a6b22eaf6ebb8353f7df3077/recovered.edits/0000000003169600719-host02%2C16201%2C1709180140645.1709186722780.temp block BP-1645452845-192.168.2.230-1615455682886:blk_1076340939_2645368] WARN  org.apache.hadoop.hdfs.DataStreamer - Error Recovery for BP-1645452845-192.168.2.230-1615455682886:blk_1076340939_2645368 in pipeline [DatanodeInfoWithStorage[192.168.2.230:15010,DS-2aa201ab-1027-47ec-b05f-b39d795fda85,DISK], DatanodeInfoWithStorage[192.168.2.232:15010,DS-39651d5a-67d2-4126-88f0-45cdee967dab,DISK], Datanode\r\nInfoWithStorage[192.168.2.231:15010,DS-e08a1d17-f7b1-4e39-9713-9706bd762f48,DISK]]: datanode 2(DatanodeInfoWithStorage[192.168.2.231:15010,DS-e08a1d17-f7b1-4e39-9713-9706bd762f48,DISK]) is bad.\r\nregionserver-2 regionserver 06:22:09.742 [split-log-closeStream-pool-1] INFO  org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink - Closed recovered edits writer path=hdfs://mycluster/hbase/data/default/aeris_v2/53308260a6b22eaf6ebb8353f7df3077/recovered.edits/0000000003169600719-host02%2C16201%\r\n2C1709180140645.1709186722780.temp (wrote 5949 edits, skipped 0 edits in 93 ms)\r\nregionserver-2 regionserver 06:22:09.743 [RS_LOG_REPLAY_OPS-regionserver/host01:16201-1-Writer-0] ERROR org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink - Failed to write log entry aeris_v2/53308260a6b22eaf6ebb8353f7df3077/3169611655=[#edits: 8 = ] to log\r\nregionserver-2 regionserver java.nio.channels.ClosedChannelException: null\r\nregionserver-2 regionserver    at org.apache.hadoop.hdfs.ExceptionLastSeen.throwException4Close(ExceptionLastSeen.java:73) ~[hadoop-hdfs-client-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.hdfs.DFSOutputStream.checkClosed(DFSOutputStream.java:153) ~[hadoop-hdfs-client-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.fs.FSOutputSummer.write(FSOutputSummer.java:105) ~[hadoop-common-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.fs.FSDataOutputStream$PositionCache.write(FSDataOutputStream.java:57) ~[hadoop-common-3.2.4.jar:?]\r\nregionserver-2 regionserver    at java.io.DataOutputStream.write(Unknown Source) ~[?:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.io.ByteBufferWriterOutputStream.write(ByteBufferWriterOutputStream.java:94) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.KeyValue.write(KeyValue.java:2093) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.KeyValueUtil.oswrite(KeyValueUtil.java:737) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.wal.WALCellCodec$EnsureKvEncoder.write(WALCellCodec.java:365) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.wal.ProtobufLogWriter.append(ProtobufLogWriter.java:58) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.AbstractRecoveredEditsOutputSink$RecoveredEditsWriter.writeRegionEntries(AbstractRecoveredEditsOutputSink.java:237) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink.append(RecoveredEditsOutputSink.java:66) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.writeBuffer(OutputSink.java:249) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.doRun(OutputSink.java:241) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.run(OutputSink.java:211) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver 06:22:09.745 [RS_LOG_REPLAY_OPS-regionserver/host01:16201-1-Writer-0] ERROR org.apache.hadoop.hbase.wal.OutputSink - Exiting thread\r\nregionserver-2 regionserver java.nio.channels.ClosedChannelException: null\r\nregionserver-2 regionserver    at org.apache.hadoop.hdfs.ExceptionLastSeen.throwException4Close(ExceptionLastSeen.java:73) ~[hadoop-hdfs-client-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.hdfs.DFSOutputStream.checkClosed(DFSOutputStream.java:153) ~[hadoop-hdfs-client-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.fs.FSOutputSummer.write(FSOutputSummer.java:105) ~[hadoop-common-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.fs.FSDataOutputStream$PositionCache.write(FSDataOutputStream.java:57) ~[hadoop-common-3.2.4.jar:?]\r\nregionserver-2 regionserver    at java.io.DataOutputStream.write(Unknown Source) ~[?:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.io.ByteBufferWriterOutputStream.write(ByteBufferWriterOutputStream.java:94) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.KeyValue.write(KeyValue.java:2093) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.KeyValueUtil.oswrite(KeyValueUtil.java:737) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.wal.WALCellCodec$EnsureKvEncoder.write(WALCellCodec.java:365) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.wal.ProtobufLogWriter.append(ProtobufLogWriter.java:58) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.AbstractRecoveredEditsOutputSink$RecoveredEditsWriter.writeRegionEntries(AbstractRecoveredEditsOutputSink.java:237) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink.append(RecoveredEditsOutputSink.java:66) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.writeBuffer(OutputSink.java:249) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.doRun(OutputSink.java:241) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.run(OutputSink.java:211) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver 06:22:09.749 [split-log-closeStream-pool-1] INFO  org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink - Rename recovered edits hdfs://mycluster/hbase/data/default/aeris_v2/53308260a6b22eaf6ebb8353f7df3077/recovered.edits/0000000003169600719-host02%2C16201%2C1709180140645.1709186722780.temp to hdfs://mycluster/hbase/data/default/aeris_v2/53308260a6b22eaf6ebb8353f7df3077/recovered.edits/0000000003169611654{noformat}\r\nAnd we can see the {{finishWriterThreads}} was interrupted:\r\n{noformat}\r\nregionserver-2 regionserver 06:22:20.643 [RS_LOG_REPLAY_OPS-regionserver/host01:16201-1] WARN  org.apache.hadoop.hbase.regionserver.SplitLogWorker - Resigning, interrupted splitting WAL hdfs://mycluster/hbase/WALs/host02,16201,1709180140645-splitting/host02%2C16201%2C1709180140645.1709186722780\r\nregionserver-2 regionserver java.io.InterruptedIOException: null\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink.finishWriterThreads(OutputSink.java:139) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink.close(RecoveredEditsOutputSink.java:94) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.WALSplitter.splitWAL(WALSplitter.java:414) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.WALSplitter.splitLogFile(WALSplitter.java:201) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.SplitLogWorker.splitLog(SplitLogWorker.java:108) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.SplitWALCallable.call(SplitWALCallable.java:100) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.SplitWALCallable.call(SplitWALCallable.java:46) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.handler.RSProcedureHandler.process(RSProcedureHandler.java:49) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.executor.EventHandler.run(EventHandler.java:98) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [?:?]\r\nregionserver-2 regionserver    at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [?:?]\r\nregionserver-2 regionserver    at java.lang.Thread.run(Unknown Source) [?:?]\r\nregionserver-2 regionserver Caused by: java.lang.InterruptedException\r\nregionserver-2 regionserver    at java.lang.Object.wait(Native Method) ~[?:?]\r\nregionserver-2 regionserver    at java.lang.Thread.join(Unknown Source) ~[?:?]\r\nregionserver-2 regionserver    at java.lang.Thread.join(Unknown Source) ~[?:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink.finishWriterThreads(OutputSink.java:137) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    ... 11 more{noformat}\r\n \r\nThe corrupted recovered edits file contained one additional piece of evidence towards a race condition. The 0 bytes that form the 32 bit trailer length were scattered through the file while the marker {{LAWP}} was written in one piece. This is because in OpenJDK11\r\n{{writeInt}} writes each byte individually but byte arrays are written all at once.\r\n \r\nOriginal bug report from Vlad Hanciuta.","from":"reporter","subject":"Race condition during WAL splitting leading to corrupt recovered.edits"},{"body":"I think the problem here is that, if the join has been interrupted, we should not consider the result as correct any more?","from":"developer"},{"body":"[~zhangduo] I've added a PR to address this - https://github.com/apache/hbase/pull/6266","from":"developer"},{"body":"Would it be possible to get some feedback on the PR? We're seeing an uptick of instances of this bug.","from":"developer"},{"body":"Pushed to all active branches.\r\n\r\nThanks [~claireiacono] for contributing!","from":"developer"}],"created":"2024-05-06T12:48:11.000+0000","description":"There is a race condition that can happen when a regionserver aborts initialisation while splitting a WAL from another regionserver. This race leads to writing the WAL trailer for recovered edits while the writer threads are still running, thus the trailer gets interleaved with the edits corrupting the recovered edits file (and preventing the region to be assigned).\r\nWe've seen this happening on HBase 2.4.17, but looking at the latest code it seems that the race can still happen there.\r\nThe sequence of operations that leads to this issue:\r\n * {{org.apache.hadoop.hbase.wal.WALSplitter.splitWAL}} calls {{outputSink.close()}} after adding all the entries to the buffers\r\n * The output sink is {{org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink}} and its {{close}} method calls first {{finishWriterThreads}} in a try block which in turn will call {{finish}} on every thread and then join it to make sure it's done.\r\n * However if the splitter thread gets interrupted because of RS aborting, the join will get interrupted and {{finishWriterThreads}} will rethrow without waiting for the writer threads to stop.\r\n * This is problematic because coming back to {{org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink.close}} it will call {{closeWriters}} in a finally block (so it will execute even when the join was interrupted).\r\n * {{closeWriters}} will call {{org.apache.hadoop.hbase.wal.AbstractRecoveredEditsOutputSink.closeRecoveredEditsWriter}} which will call {{close}} on {{{}editWriter.writer{}}}.\r\n * When {{editWriter.writer}} is {{{}org.apache.hadoop.hbase.regionserver.wal.ProtobufLogWriter{}}}, its {{close}} method will write the trailer before closing the file.\r\n * This trailer write will now go in parallel with writer threads writing entries causing corruption.\r\n * If there are no other errors, {{closeWriters}} will succeed renaming all temporary files to final recovered edits, causing problems next time the region is assigned.\r\n\r\nLogs evidence supporting the above flow:\r\nAbort is triggered (because it failed to open the WAL due to some ongoing infra issue):\r\n{noformat}\r\nregionserver-2 regionserver 06:22:00.384 [RS_OPEN_META-regionserver/host01:16201-0] ERROR org.apache.hadoop.hbase.regionserver.HRegionServer - ***** ABORTING region server host01,16201,1709187641249: WAL can not clean up after init failed *****{noformat}\r\nWe can see that the writer threads were still active after closing (even considering that the\r\nordering in the log might not be accurate, we see that they die because the channel is closed while still writing, not because they're stopping):\r\n{noformat}\r\nregionserver-2 regionserver 06:22:09.662 [DataStreamer for file /hbase/data/default/aeris_v2/53308260a6b22eaf6ebb8353f7df3077/recovered.edits/0000000003169600719-host02%2C16201%2C1709180140645.1709186722780.temp block BP-1645452845-192.168.2.230-1615455682886:blk_1076340939_2645368] WARN  org.apache.hadoop.hdfs.DataStreamer - Error Recovery for BP-1645452845-192.168.2.230-1615455682886:blk_1076340939_2645368 in pipeline [DatanodeInfoWithStorage[192.168.2.230:15010,DS-2aa201ab-1027-47ec-b05f-b39d795fda85,DISK], DatanodeInfoWithStorage[192.168.2.232:15010,DS-39651d5a-67d2-4126-88f0-45cdee967dab,DISK], Datanode\r\nInfoWithStorage[192.168.2.231:15010,DS-e08a1d17-f7b1-4e39-9713-9706bd762f48,DISK]]: datanode 2(DatanodeInfoWithStorage[192.168.2.231:15010,DS-e08a1d17-f7b1-4e39-9713-9706bd762f48,DISK]) is bad.\r\nregionserver-2 regionserver 06:22:09.742 [split-log-closeStream-pool-1] INFO  org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink - Closed recovered edits writer path=hdfs://mycluster/hbase/data/default/aeris_v2/53308260a6b22eaf6ebb8353f7df3077/recovered.edits/0000000003169600719-host02%2C16201%\r\n2C1709180140645.1709186722780.temp (wrote 5949 edits, skipped 0 edits in 93 ms)\r\nregionserver-2 regionserver 06:22:09.743 [RS_LOG_REPLAY_OPS-regionserver/host01:16201-1-Writer-0] ERROR org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink - Failed to write log entry aeris_v2/53308260a6b22eaf6ebb8353f7df3077/3169611655=[#edits: 8 = ] to log\r\nregionserver-2 regionserver java.nio.channels.ClosedChannelException: null\r\nregionserver-2 regionserver    at org.apache.hadoop.hdfs.ExceptionLastSeen.throwException4Close(ExceptionLastSeen.java:73) ~[hadoop-hdfs-client-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.hdfs.DFSOutputStream.checkClosed(DFSOutputStream.java:153) ~[hadoop-hdfs-client-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.fs.FSOutputSummer.write(FSOutputSummer.java:105) ~[hadoop-common-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.fs.FSDataOutputStream$PositionCache.write(FSDataOutputStream.java:57) ~[hadoop-common-3.2.4.jar:?]\r\nregionserver-2 regionserver    at java.io.DataOutputStream.write(Unknown Source) ~[?:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.io.ByteBufferWriterOutputStream.write(ByteBufferWriterOutputStream.java:94) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.KeyValue.write(KeyValue.java:2093) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.KeyValueUtil.oswrite(KeyValueUtil.java:737) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.wal.WALCellCodec$EnsureKvEncoder.write(WALCellCodec.java:365) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.wal.ProtobufLogWriter.append(ProtobufLogWriter.java:58) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.AbstractRecoveredEditsOutputSink$RecoveredEditsWriter.writeRegionEntries(AbstractRecoveredEditsOutputSink.java:237) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink.append(RecoveredEditsOutputSink.java:66) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.writeBuffer(OutputSink.java:249) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.doRun(OutputSink.java:241) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.run(OutputSink.java:211) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver 06:22:09.745 [RS_LOG_REPLAY_OPS-regionserver/host01:16201-1-Writer-0] ERROR org.apache.hadoop.hbase.wal.OutputSink - Exiting thread\r\nregionserver-2 regionserver java.nio.channels.ClosedChannelException: null\r\nregionserver-2 regionserver    at org.apache.hadoop.hdfs.ExceptionLastSeen.throwException4Close(ExceptionLastSeen.java:73) ~[hadoop-hdfs-client-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.hdfs.DFSOutputStream.checkClosed(DFSOutputStream.java:153) ~[hadoop-hdfs-client-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.fs.FSOutputSummer.write(FSOutputSummer.java:105) ~[hadoop-common-3.2.4.jar:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.fs.FSDataOutputStream$PositionCache.write(FSDataOutputStream.java:57) ~[hadoop-common-3.2.4.jar:?]\r\nregionserver-2 regionserver    at java.io.DataOutputStream.write(Unknown Source) ~[?:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.io.ByteBufferWriterOutputStream.write(ByteBufferWriterOutputStream.java:94) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.KeyValue.write(KeyValue.java:2093) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.KeyValueUtil.oswrite(KeyValueUtil.java:737) ~[hbase-common-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.wal.WALCellCodec$EnsureKvEncoder.write(WALCellCodec.java:365) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.wal.ProtobufLogWriter.append(ProtobufLogWriter.java:58) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.AbstractRecoveredEditsOutputSink$RecoveredEditsWriter.writeRegionEntries(AbstractRecoveredEditsOutputSink.java:237) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink.append(RecoveredEditsOutputSink.java:66) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.writeBuffer(OutputSink.java:249) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.doRun(OutputSink.java:241) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink$WriterThread.run(OutputSink.java:211) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver 06:22:09.749 [split-log-closeStream-pool-1] INFO  org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink - Rename recovered edits hdfs://mycluster/hbase/data/default/aeris_v2/53308260a6b22eaf6ebb8353f7df3077/recovered.edits/0000000003169600719-host02%2C16201%2C1709180140645.1709186722780.temp to hdfs://mycluster/hbase/data/default/aeris_v2/53308260a6b22eaf6ebb8353f7df3077/recovered.edits/0000000003169611654{noformat}\r\nAnd we can see the {{finishWriterThreads}} was interrupted:\r\n{noformat}\r\nregionserver-2 regionserver 06:22:20.643 [RS_LOG_REPLAY_OPS-regionserver/host01:16201-1] WARN  org.apache.hadoop.hbase.regionserver.SplitLogWorker - Resigning, interrupted splitting WAL hdfs://mycluster/hbase/WALs/host02,16201,1709180140645-splitting/host02%2C16201%2C1709180140645.1709186722780\r\nregionserver-2 regionserver java.io.InterruptedIOException: null\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink.finishWriterThreads(OutputSink.java:139) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.RecoveredEditsOutputSink.close(RecoveredEditsOutputSink.java:94) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.WALSplitter.splitWAL(WALSplitter.java:414) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.WALSplitter.splitLogFile(WALSplitter.java:201) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.SplitLogWorker.splitLog(SplitLogWorker.java:108) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.SplitWALCallable.call(SplitWALCallable.java:100) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.SplitWALCallable.call(SplitWALCallable.java:46) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.regionserver.handler.RSProcedureHandler.process(RSProcedureHandler.java:49) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.executor.EventHandler.run(EventHandler.java:98) [hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [?:?]\r\nregionserver-2 regionserver    at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [?:?]\r\nregionserver-2 regionserver    at java.lang.Thread.run(Unknown Source) [?:?]\r\nregionserver-2 regionserver Caused by: java.lang.InterruptedException\r\nregionserver-2 regionserver    at java.lang.Object.wait(Native Method) ~[?:?]\r\nregionserver-2 regionserver    at java.lang.Thread.join(Unknown Source) ~[?:?]\r\nregionserver-2 regionserver    at java.lang.Thread.join(Unknown Source) ~[?:?]\r\nregionserver-2 regionserver    at org.apache.hadoop.hbase.wal.OutputSink.finishWriterThreads(OutputSink.java:137) ~[hbase-server-2.4.17.jar:2.4.17]\r\nregionserver-2 regionserver    ... 11 more{noformat}\r\n \r\nThe corrupted recovered edits file contained one additional piece of evidence towards a race condition. The 0 bytes that form the 32 bit trailer length were scattered through the file while the marker {{LAWP}} was written in one piece. This is because in OpenJDK11\r\n{{writeInt}} writes each byte individually but byte arrays are written all at once.\r\n \r\nOriginal bug report from Vlad Hanciuta.","issue_id":"13578273","key":"HBASE-28569","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2025-04-07T15:50:08.000+0000","role":"fixed_distractor","summary":"Race condition during WAL splitting leading to corrupt recovered.edits"} {"case_id":"13579511","cluster":"DISTRACTOR-HBASE-28599","comments":[{"body":"Hi,[~robiee17]. This is youngjukim and I'm \boperating HBase clusters in my company from Korea. Thank you for explaining the detailed problem definition, reproduction process with fix \bsuggestions. I'd like to do HBase contributing. Is it okay if I try to solve this issue?","created":"2024-05-18T15:20:33.829+0000"},{"body":"Hello, [~zhangduo] Could you review this PR?\r\n\r\n[https://github.com/apache/hbase/pull/5927]","created":"2024-05-19T12:53:18.823+0000"},{"body":"Pushed to all active branches.\r\n\r\nThanks [~fjvbn2003] for contributing and [~robiee17] for the great analyzing!","created":"2024-05-20T01:38:12.733+0000"}],"conversations":[{"body":"*Issue:*\r\n`RowTooBigException` is thrown when a duplicate increment RPC call is attempted.\r\n\r\n*Expected Behavior:*\r\n1. The initial RPC increment call should time out for some reason.\r\n2. The duplicate RPC call should be converted to a GET request and fetch the result that I am trying to increment.\r\n3. The result should contain only the qualifier that I am attempting to increment.\r\n\r\n*Actual Behavior:*\r\n1. The initial RPC increment call timed out, which is expected.\r\n2. The duplicate RPC call is converted to a GET request but fails to clone the qualifier into the GET request.\r\n3. Hence, the GET request attempts to retrieve all qualifiers for the given row and columnfamily, resulting in a `RowTooBigException`.\r\n\r\n*Steps to Reproduce:*\r\n1. Ensure a row with a total value size exceeding `hbase.table.max.rowsize` (default = 1073741824) exists.\r\n2. Nonce property should be enabled `hbase.client.nonces.enabled` which is actually defaulted to true.\r\n3. Attempt to increment a qualifier against the same row.\r\n4. In my case, I am using a postIncrement co-processor which may cause a delay (longer than the RPC timeout property).\r\n5. A duplicate increment call should be triggered, which tries to get the value rather than increment it.\r\n6. The GET request actually tries to retrieve all the qualifiers for the row, resulting in a `RowTooBigException`.\r\n\r\n*Insights:*\r\nUpon further debugging, I found that qualifiers are not cloned into the GET instance due to incorrect usage of [CellScanner.advance|https://github.com/apache/hbase/blob/7ebd4381261fefd78fc2acf258a95184f4147cee/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java#L3833]\r\n\r\n*Fix Suggestion:*\r\nRemoving the `!` operation from `while (!CellScanner.advance)` may resolve the issue.\r\n\r\nAttached Exception Stack Trace for reference.","from":"reporter","subject":"RowTooBigException is thrown when duplicate increment RPC call is attempted"},{"body":"Hi,[~robiee17]. This is youngjukim and I'm \boperating HBase clusters in my company from Korea. Thank you for explaining the detailed problem definition, reproduction process with fix \bsuggestions. I'd like to do HBase contributing. Is it okay if I try to solve this issue?","from":"developer"},{"body":"Hello, [~zhangduo] Could you review this PR?\r\n\r\n[https://github.com/apache/hbase/pull/5927]","from":"developer"},{"body":"Pushed to all active branches.\r\n\r\nThanks [~fjvbn2003] for contributing and [~robiee17] for the great analyzing!","from":"developer"}],"created":"2024-05-16T11:45:29.000+0000","description":"*Issue:*\r\n`RowTooBigException` is thrown when a duplicate increment RPC call is attempted.\r\n\r\n*Expected Behavior:*\r\n1. The initial RPC increment call should time out for some reason.\r\n2. The duplicate RPC call should be converted to a GET request and fetch the result that I am trying to increment.\r\n3. The result should contain only the qualifier that I am attempting to increment.\r\n\r\n*Actual Behavior:*\r\n1. The initial RPC increment call timed out, which is expected.\r\n2. The duplicate RPC call is converted to a GET request but fails to clone the qualifier into the GET request.\r\n3. Hence, the GET request attempts to retrieve all qualifiers for the given row and columnfamily, resulting in a `RowTooBigException`.\r\n\r\n*Steps to Reproduce:*\r\n1. Ensure a row with a total value size exceeding `hbase.table.max.rowsize` (default = 1073741824) exists.\r\n2. Nonce property should be enabled `hbase.client.nonces.enabled` which is actually defaulted to true.\r\n3. Attempt to increment a qualifier against the same row.\r\n4. In my case, I am using a postIncrement co-processor which may cause a delay (longer than the RPC timeout property).\r\n5. A duplicate increment call should be triggered, which tries to get the value rather than increment it.\r\n6. The GET request actually tries to retrieve all the qualifiers for the row, resulting in a `RowTooBigException`.\r\n\r\n*Insights:*\r\nUpon further debugging, I found that qualifiers are not cloned into the GET instance due to incorrect usage of [CellScanner.advance|https://github.com/apache/hbase/blob/7ebd4381261fefd78fc2acf258a95184f4147cee/hbase-server/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java#L3833]\r\n\r\n*Fix Suggestion:*\r\nRemoving the `!` operation from `while (!CellScanner.advance)` may resolve the issue.\r\n\r\nAttached Exception Stack Trace for reference.","issue_id":"13579511","key":"HBASE-28599","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2024-05-20T01:38:12.000+0000","role":"fixed_distractor","summary":"RowTooBigException is thrown when duplicate increment RPC call is attempted"} {"case_id":"12382932","cluster":"DISTRACTOR-HBASE-286","comments":[{"body":"Do you have an idea of how to update a metadata for table initializing?","created":"2007-11-21T01:45:06.001+0000"},{"body":"I am not sure how its coded right now for drop and create commands from the shell but I figured we could borrow from that\n\nSteps I was suggestion:\n\n1. Get col data for table with - > show tables;\n2. Drop table\n3. Create table from the show table data\n\nAll the code should be in the drop and create commands code to take care of the meta table.\nJust need to come up with a way to parse the show table data for the table we need.\nI have never coded in Java so this might be harder then I thank.\n\nfrom dropping tables and creating them in the past on my setup dropping takes about 10sec but create right after is done in < 1s","created":"2007-11-21T02:07:23.262+0000"},{"body":"That would be fine, Billy. \n\nIn addition, We should think about delete statement with no WHERE clause.\nAnd, I'd like to discuss about performance advantages of these functions except functional advantages.","created":"2007-11-22T02:23:14.184+0000"},{"body":"- Truncate table is used to clean all data from a table.\n\nSYNTAX: TRUNCATE TABLE table_name;\n\n{code}\nhql > insert into test (a,b) values ('aa','bb') where row='row1';\n1 row inserted successfully. (0.41 sec)\nhql > truncate table test;\n'test' is successfully truncated.\nhql > select * from test;\n0 row(s) in set. (1.16 sec)\nhql > exit;\n{code}\n\n- We need to get away from functionalities and focus on the internal algorithm issues. I'll make these issues.\n- i didn't make the test case becuase it already exists. (It just steps of creating drop repeats.)\n","created":"2007-12-01T05:22:18.206+0000"},{"body":"submitting.","created":"2007-12-01T05:23:05.766+0000"},{"body":"Should it be closed?","created":"2007-12-20T14:35:16.609+0000"},{"body":"Edward: This patch won't apply for me. Could you please regenerate? I'm at r605980 12/20/2007. FYI, this kinda thing does not belong as a class comment: \"But, We need to get away from functionalities and focus on the internal algorithm issues. - Edward.\" \n\nThanks","created":"2007-12-20T17:24:49.922+0000"},{"body":"regenerated patch","created":"2007-12-21T00:36:37.994+0000"},{"body":"submitting","created":"2007-12-21T00:37:17.569+0000"},{"body":"help output patch","created":"2007-12-21T02:09:04.895+0000"},{"body":"Patch that puts together help-string.patch and 2240.2 (I also tried it locally and it works +1).","created":"2007-12-24T17:31:43.937+0000"},{"body":"Try hudson.","created":"2007-12-24T17:32:17.835+0000"},{"body":"It's not running.","created":"2007-12-25T22:33:15.454+0000"},{"body":"Edward: There is no point rescheduling same patch. Hudsons complaint is that the patch just doesn't apply (I just tried it and indeed it doesn't).","created":"2007-12-29T06:54:39.193+0000"},{"body":"Oh. I see.","created":"2007-12-29T07:04:46.160+0000"},{"body":"not fit.\ncanceling.","created":"2008-01-03T05:42:20.633+0000"},{"body":"Here's a clean-up patch.","created":"2008-01-03T06:47:32.633+0000"},{"body":"submitting.","created":"2008-01-03T06:47:53.760+0000"},{"body":"Edward: Your new truncate command doesn't have an entry in the 'help' output. Without one, I don't think folks will realize the functionality exists.","created":"2008-01-04T04:25:53.255+0000"},{"body":"Oh, sorry. I missed it.","created":"2008-01-04T05:35:06.215+0000"},{"body":"Added help-output.","created":"2008-01-04T05:35:37.927+0000"},{"body":"submitting.","created":"2008-01-04T05:36:02.330+0000"},{"body":"Committed after confirming it works by trying it locally (Test failures are erratics unrelated to shell failures). Resolving. Thanks for the patch Edward.","created":"2008-01-06T04:46:32.843+0000"},{"body":"I can't find a 'TruncateCommand.java'\nPlease check this problem!!\n\nsrc/contrib/hbase/src/java/org/apache/hadoop/hbase/shell/TruncateCommand.java\n","created":"2008-01-06T06:02:30.653+0000"},{"body":"Thanks Edward. I forgot to add it. I just committed it.","created":"2008-01-06T06:12:35.810+0000"}],"conversations":[{"body":"Would be nice to have a way to truncate the tables from the shell. With out doing a drop and create your self. Maybe the truncate could issue the drop and create command for you based off the layout in the the table before the truncate.","from":"reporter","subject":"Truncate for hbase"},{"body":"Do you have an idea of how to update a metadata for table initializing?","from":"developer"},{"body":"I am not sure how its coded right now for drop and create commands from the shell but I figured we could borrow from that\n\nSteps I was suggestion:\n\n1. Get col data for table with - > show tables;\n2. Drop table\n3. Create table from the show table data\n\nAll the code should be in the drop and create commands code to take care of the meta table.\nJust need to come up with a way to parse the show table data for the table we need.\nI have never coded in Java so this might be harder then I thank.\n\nfrom dropping tables and creating them in the past on my setup dropping takes about 10sec but create right after is done in < 1s","from":"developer"},{"body":"That would be fine, Billy. \n\nIn addition, We should think about delete statement with no WHERE clause.\nAnd, I'd like to discuss about performance advantages of these functions except functional advantages.","from":"developer"},{"body":"- Truncate table is used to clean all data from a table.\n\nSYNTAX: TRUNCATE TABLE table_name;\n\n{code}\nhql > insert into test (a,b) values ('aa','bb') where row='row1';\n1 row inserted successfully. (0.41 sec)\nhql > truncate table test;\n'test' is successfully truncated.\nhql > select * from test;\n0 row(s) in set. (1.16 sec)\nhql > exit;\n{code}\n\n- We need to get away from functionalities and focus on the internal algorithm issues. I'll make these issues.\n- i didn't make the test case becuase it already exists. (It just steps of creating drop repeats.)\n","from":"developer"},{"body":"submitting.","from":"developer"},{"body":"Should it be closed?","from":"developer"},{"body":"Edward: This patch won't apply for me. Could you please regenerate? I'm at r605980 12/20/2007. FYI, this kinda thing does not belong as a class comment: \"But, We need to get away from functionalities and focus on the internal algorithm issues. - Edward.\" \n\nThanks","from":"developer"},{"body":"regenerated patch","from":"developer"},{"body":"submitting","from":"developer"},{"body":"help output patch","from":"developer"},{"body":"Patch that puts together help-string.patch and 2240.2 (I also tried it locally and it works +1).","from":"developer"},{"body":"Try hudson.","from":"developer"},{"body":"It's not running.","from":"developer"},{"body":"Edward: There is no point rescheduling same patch. Hudsons complaint is that the patch just doesn't apply (I just tried it and indeed it doesn't).","from":"developer"},{"body":"Oh. I see.","from":"developer"},{"body":"not fit.\ncanceling.","from":"developer"},{"body":"Here's a clean-up patch.","from":"developer"},{"body":"submitting.","from":"developer"},{"body":"Edward: Your new truncate command doesn't have an entry in the 'help' output. Without one, I don't think folks will realize the functionality exists.","from":"developer"},{"body":"Oh, sorry. I missed it.","from":"developer"},{"body":"Added help-output.","from":"developer"},{"body":"submitting.","from":"developer"},{"body":"Committed after confirming it works by trying it locally (Test failures are erratics unrelated to shell failures). Resolving. Thanks for the patch Edward.","from":"developer"},{"body":"I can't find a 'TruncateCommand.java'\nPlease check this problem!!\n\nsrc/contrib/hbase/src/java/org/apache/hadoop/hbase/shell/TruncateCommand.java\n","from":"developer"},{"body":"Thanks Edward. I forgot to add it. I just committed it.","from":"developer"}],"created":"2007-11-21T00:48:39.000+0000","description":"Would be nice to have a way to truncate the tables from the shell. With out doing a drop and create your self. Maybe the truncate could issue the drop and create command for you based off the layout in the the table before the truncate.","issue_id":"12382932","key":"HBASE-286","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2008-01-06T04:46:32.000+0000","role":"fixed_distractor","summary":"Truncate for hbase"} {"case_id":"13590958","cluster":"DISTRACTOR-HBASE-28812","comments":[{"body":"May be related to HBASE-28577? We removed KVComparator in HBASE-28577?\r\n\r\nBut it is a bit strange that we should not use this class any more, it is for 1.x...","created":"2024-09-04T07:47:02.504+0000"},{"body":"[~zhangduo] Thank you for the reply! It seems yes. Upgrading from 2.6.0 to the commit before HBASE-28577 will succeed.\r\n\r\n \r\n\r\nI tested upgrading from 2.6.0 to the following 2 commits from master branch\r\n{code:java}\r\nUpgrade crashed. (The error log looks similar. I attached the failure log: hbase--master-440ed844e077.log)\r\n\r\ncommit 419666b8eb8a881724fe6f65e8235a4220824e51 (HEAD)\r\nAuthor: lixiaobao <977734161@qq.com>\r\nDate:   Wed May 22 18:34:42 2024 +0800    HBASE-28577 Remove deprecated methods in KeyValue (#5883)\r\n    \r\n    Co-authored-by: lixiaobao \r\n    Co-authored-by: 李小保 \r\n    Signed-off-by: Duo Zhang \r\n\r\n\r\n===\r\n\r\nUpgrade succeeded.\r\n\r\ncommit 3b18ba664a6dcde344e13fe9305c272592195c03\r\nAuthor: Nick Dimiduk \r\nDate:   Wed May 22 10:05:54 2024 +0200    HBASE-28605 Add ErrorProne ban on Hadoop shaded thirdparty jars (#5918)\r\n    \r\n    This change results in this error on master at `3a3dd66e21`.\r\n    \r\n    ```\r\n    [WARNING] Rule 2: de.skuzzle.enforcer.restrictimports.rule.RestrictImports failed with message:\r\n    \r\n    Banned imports detected:\r\n    Reason: Use shaded version in hbase-thirdparty{code}\r\n ","created":"2024-09-04T17:40:12.345+0000"},{"body":"Ah, on hbase-2.x, we still want to maintain compatibility with 1.x, so we will force write 1.x comparator name to the trailer of HFile...\r\n\r\n{code}\r\n private String getHBase1CompatibleName(final String comparator) {\r\n if (\r\n comparator.equals(CellComparatorImpl.class.getName())\r\n || comparator.equals(InnerStoreCellComparator.class.getName())\r\n ) {\r\n return KeyValue.COMPARATOR.getClass().getName();\r\n }\r\n if (comparator.equals(MetaCellComparator.class.getName())) {\r\n return KeyValue.META_COMPARATOR.getClass().getName();\r\n }\r\n return comparator;\r\n }\r\n{code}\r\n\r\nSo we still need to check the comparator name for the old KVComparator.\r\n\r\nLet me open a PR.","created":"2024-09-05T11:52:58.690+0000"},{"body":"[~kehan5800] The PR is ready, could you please verify if it works after applying the PR?\r\n\r\nThanks.","created":"2024-09-09T14:47:25.982+0000"},{"body":"[Duo Zhang|https://issues.apache.org/jira/secure/ViewProfile.jspa?name=zhangduo] Thank you for the PR! I have applied the patch to a030e80998, and it's working.","created":"2024-09-09T17:02:03.308+0000"},{"body":"Pushed to master and branch-3.\r\n\r\nThanks [~kehan5800] for reporting and verifying.\r\n\r\nThanks [~ndimiduk] for reviewing!","created":"2024-09-16T14:11:25.861+0000"}],"conversations":[{"body":"I am trying to upgrade from 2.6.0 (stable release) to 3.0.0. I built 3.0.0 using the following commit (a030e8099840e640684a68b6e4a79e7c1d5a6823)\r\n{code:java}\r\ncommit a030e8099840e640684a68b6e4a79e7c1d5a6823 (HEAD -> branch-3, upstream/branch-3)\r\nAuthor: Ray Mattingly \r\nDate:   Mon Sep 2 04:38:29 2024 -0400    HBASE-28697 Don't clean bulk load system entries until backup is complete (#6089)\r\n    \r\n    Co-authored-by: Ray Mattingly \r\n{code}\r\nHowever, the HMaster would crash during the upgrade process.\r\nh1. Reproduce\r\n\r\nStep1: Start up 2.6.0 cluster (1 HDFS, 1 HM, 1 RS)\r\n\r\nStep2: Stop the entire cluster\r\n\r\nStep3: Upgrade to 3.0.0 cluster.\r\n\r\nHMaster will crash with the following error message\r\n{code:java}\r\n2024-09-04T04:29:18,917 WARN  [master/hmaster:16000:becomeActiveMaster] regionserver.HRegion: Failed initialize of region= master:store,,1.1595e783b53d99cd5eef43b6debb2682., starting to roll back memstore\r\njava.io.IOException: java.io.IOException: org.apache.hadoop.hbase.io.hfile.CorruptHFileException: Problem reading HFile Trailer from file hdfs://master:8020/hbase/MasterData/data/master/store/1595e783b53d99cd5eef43b6debb2682/info/82c6d244b6244c179cdbafcead00ed75\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.initializeStores(HRegion.java:1215) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.initializeStores(HRegion.java:1158) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.initializeRegionInternals(HRegion.java:1030) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.initialize(HRegion.java:974) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.openHRegion(HRegion.java:7794) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.openHRegionFromTableDir(HRegion.java:7749) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.region.MasterRegion.open(MasterRegion.java:277) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.region.MasterRegion.create(MasterRegion.java:432) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.region.MasterRegionFactory.create(MasterRegionFactory.java:135) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.HMaster.finishActiveMasterInitialization(HMaster.java:1003) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.HMaster.startActiveMasterManager(HMaster.java:2524) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.HMaster.lambda$run$0(HMaster.java:613) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.trace.TraceUtil.lambda$tracedRunnable$2(TraceUtil.java:155) ~[hbase-common-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at java.lang.Thread.run(Thread.java:833) ~[?:?]\r\nCaused by: java.io.IOException: org.apache.hadoop.hbase.io.hfile.CorruptHFileException: Problem reading HFile Trailer from file hdfs://master:8020/hbase/MasterData/data/master/store/1595e783b53d99cd5eef43b6debb2682/info/82c6d244b6244c179cdbafcead00ed75\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.openStoreFiles(StoreEngine.java:289) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.initialize(StoreEngine.java:339) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStore.(HStore.java:301) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.instantiateHStore(HRegion.java:6924) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion$1.call(HRegion.java:1181) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion$1.call(HRegion.java:1178) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) ~[?:?]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) ~[?:?]\r\n        ... 1 more\r\nCaused by: org.apache.hadoop.hbase.io.hfile.CorruptHFileException: Problem reading HFile Trailer from file hdfs://master:8020/hbase/MasterData/data/master/store/1595e783b53d99cd5eef43b6debb2682/info/82c6d244b6244c179cdbafcead00ed75\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.initTrailerAndContext(HFileInfo.java:359) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.(HFileInfo.java:132) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreFileInfo.initHFileInfo(StoreFileInfo.java:763) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.open(HStoreFile.java:395) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.initReader(HStoreFile.java:524) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.createStoreFileAndReader(StoreEngine.java:226) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.lambda$openStoreFiles$0(StoreEngine.java:267) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) ~[?:?]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) ~[?:?]\r\n        ... 1 more\r\nCaused by: java.io.IOException: java.lang.ClassNotFoundException: org.apache.hadoop.hbase.KeyValue$KVComparator\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.getComparatorClass(FixedFileTrailer.java:578) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.deserializeFromPB(FixedFileTrailer.java:304) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.deserialize(FixedFileTrailer.java:250) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.readFromStream(FixedFileTrailer.java:407) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.initTrailerAndContext(HFileInfo.java:349) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.(HFileInfo.java:132) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreFileInfo.initHFileInfo(StoreFileInfo.java:763) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.open(HStoreFile.java:395) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.initReader(HStoreFile.java:524) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.createStoreFileAndReader(StoreEngine.java:226) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.lambda$openStoreFiles$0(StoreEngine.java:267) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) ~[?:?]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) ~[?:?]\r\n        ... 1 more\r\nCaused by: java.lang.ClassNotFoundException: org.apache.hadoop.hbase.KeyValue$KVComparator\r\n        at jdk.internal.loader.BuiltinClassLoader.loadClass(BuiltinClassLoader.java:641) ~[?:?]\r\n        at jdk.internal.loader.ClassLoaders$AppClassLoader.loadClass(ClassLoaders.java:188) ~[?:?]\r\n        at java.lang.ClassLoader.loadClass(ClassLoader.java:520) ~[?:?]\r\n        at java.lang.Class.forName0(Native Method) ~[?:?]\r\n        at java.lang.Class.forName(Class.java:375) ~[?:?]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.getComparatorClass(FixedFileTrailer.java:576) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.deserializeFromPB(FixedFileTrailer.java:304) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.deserialize(FixedFileTrailer.java:250) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.readFromStream(FixedFileTrailer.java:407) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.initTrailerAndContext(HFileInfo.java:349) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.(HFileInfo.java:132) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreFileInfo.initHFileInfo(StoreFileInfo.java:763) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.open(HStoreFile.java:395) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.initReader(HStoreFile.java:524) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.createStoreFileAndReader(StoreEngine.java:226) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.lambda$openStoreFiles$0(StoreEngine.java:267) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) ~[?:?]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) ~[?:?]\r\n        ... 1 more {code}\r\nThis problem seems to be introduced recently, and I can still upgrade from 2.6.0 to 3.0.0 using the previous commits (E.g. commit from May 24: 516c89e8597fb6ed391f9e85e594f8b7e5b56e38). \r\n\r\nI have attached the hmaster log.","from":"reporter","subject":"Upgrade from 2.6.0 to 3.0.0 crashed"},{"body":"May be related to HBASE-28577? We removed KVComparator in HBASE-28577?\r\n\r\nBut it is a bit strange that we should not use this class any more, it is for 1.x...","from":"developer"},{"body":"[~zhangduo] Thank you for the reply! It seems yes. Upgrading from 2.6.0 to the commit before HBASE-28577 will succeed.\r\n\r\n \r\n\r\nI tested upgrading from 2.6.0 to the following 2 commits from master branch\r\n{code:java}\r\nUpgrade crashed. (The error log looks similar. I attached the failure log: hbase--master-440ed844e077.log)\r\n\r\ncommit 419666b8eb8a881724fe6f65e8235a4220824e51 (HEAD)\r\nAuthor: lixiaobao <977734161@qq.com>\r\nDate:   Wed May 22 18:34:42 2024 +0800    HBASE-28577 Remove deprecated methods in KeyValue (#5883)\r\n    \r\n    Co-authored-by: lixiaobao \r\n    Co-authored-by: 李小保 \r\n    Signed-off-by: Duo Zhang \r\n\r\n\r\n===\r\n\r\nUpgrade succeeded.\r\n\r\ncommit 3b18ba664a6dcde344e13fe9305c272592195c03\r\nAuthor: Nick Dimiduk \r\nDate:   Wed May 22 10:05:54 2024 +0200    HBASE-28605 Add ErrorProne ban on Hadoop shaded thirdparty jars (#5918)\r\n    \r\n    This change results in this error on master at `3a3dd66e21`.\r\n    \r\n    ```\r\n    [WARNING] Rule 2: de.skuzzle.enforcer.restrictimports.rule.RestrictImports failed with message:\r\n    \r\n    Banned imports detected:\r\n    Reason: Use shaded version in hbase-thirdparty{code}\r\n ","from":"developer"},{"body":"Ah, on hbase-2.x, we still want to maintain compatibility with 1.x, so we will force write 1.x comparator name to the trailer of HFile...\r\n\r\n{code}\r\n private String getHBase1CompatibleName(final String comparator) {\r\n if (\r\n comparator.equals(CellComparatorImpl.class.getName())\r\n || comparator.equals(InnerStoreCellComparator.class.getName())\r\n ) {\r\n return KeyValue.COMPARATOR.getClass().getName();\r\n }\r\n if (comparator.equals(MetaCellComparator.class.getName())) {\r\n return KeyValue.META_COMPARATOR.getClass().getName();\r\n }\r\n return comparator;\r\n }\r\n{code}\r\n\r\nSo we still need to check the comparator name for the old KVComparator.\r\n\r\nLet me open a PR.","from":"developer"},{"body":"[~kehan5800] The PR is ready, could you please verify if it works after applying the PR?\r\n\r\nThanks.","from":"developer"},{"body":"[Duo Zhang|https://issues.apache.org/jira/secure/ViewProfile.jspa?name=zhangduo] Thank you for the PR! I have applied the patch to a030e80998, and it's working.","from":"developer"},{"body":"Pushed to master and branch-3.\r\n\r\nThanks [~kehan5800] for reporting and verifying.\r\n\r\nThanks [~ndimiduk] for reviewing!","from":"developer"}],"created":"2024-09-04T05:22:44.000+0000","description":"I am trying to upgrade from 2.6.0 (stable release) to 3.0.0. I built 3.0.0 using the following commit (a030e8099840e640684a68b6e4a79e7c1d5a6823)\r\n{code:java}\r\ncommit a030e8099840e640684a68b6e4a79e7c1d5a6823 (HEAD -> branch-3, upstream/branch-3)\r\nAuthor: Ray Mattingly \r\nDate:   Mon Sep 2 04:38:29 2024 -0400    HBASE-28697 Don't clean bulk load system entries until backup is complete (#6089)\r\n    \r\n    Co-authored-by: Ray Mattingly \r\n{code}\r\nHowever, the HMaster would crash during the upgrade process.\r\nh1. Reproduce\r\n\r\nStep1: Start up 2.6.0 cluster (1 HDFS, 1 HM, 1 RS)\r\n\r\nStep2: Stop the entire cluster\r\n\r\nStep3: Upgrade to 3.0.0 cluster.\r\n\r\nHMaster will crash with the following error message\r\n{code:java}\r\n2024-09-04T04:29:18,917 WARN  [master/hmaster:16000:becomeActiveMaster] regionserver.HRegion: Failed initialize of region= master:store,,1.1595e783b53d99cd5eef43b6debb2682., starting to roll back memstore\r\njava.io.IOException: java.io.IOException: org.apache.hadoop.hbase.io.hfile.CorruptHFileException: Problem reading HFile Trailer from file hdfs://master:8020/hbase/MasterData/data/master/store/1595e783b53d99cd5eef43b6debb2682/info/82c6d244b6244c179cdbafcead00ed75\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.initializeStores(HRegion.java:1215) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.initializeStores(HRegion.java:1158) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.initializeRegionInternals(HRegion.java:1030) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.initialize(HRegion.java:974) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.openHRegion(HRegion.java:7794) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.openHRegionFromTableDir(HRegion.java:7749) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.region.MasterRegion.open(MasterRegion.java:277) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.region.MasterRegion.create(MasterRegion.java:432) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.region.MasterRegionFactory.create(MasterRegionFactory.java:135) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.HMaster.finishActiveMasterInitialization(HMaster.java:1003) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.HMaster.startActiveMasterManager(HMaster.java:2524) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.master.HMaster.lambda$run$0(HMaster.java:613) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.trace.TraceUtil.lambda$tracedRunnable$2(TraceUtil.java:155) ~[hbase-common-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at java.lang.Thread.run(Thread.java:833) ~[?:?]\r\nCaused by: java.io.IOException: org.apache.hadoop.hbase.io.hfile.CorruptHFileException: Problem reading HFile Trailer from file hdfs://master:8020/hbase/MasterData/data/master/store/1595e783b53d99cd5eef43b6debb2682/info/82c6d244b6244c179cdbafcead00ed75\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.openStoreFiles(StoreEngine.java:289) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.initialize(StoreEngine.java:339) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStore.(HStore.java:301) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion.instantiateHStore(HRegion.java:6924) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion$1.call(HRegion.java:1181) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HRegion$1.call(HRegion.java:1178) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) ~[?:?]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) ~[?:?]\r\n        ... 1 more\r\nCaused by: org.apache.hadoop.hbase.io.hfile.CorruptHFileException: Problem reading HFile Trailer from file hdfs://master:8020/hbase/MasterData/data/master/store/1595e783b53d99cd5eef43b6debb2682/info/82c6d244b6244c179cdbafcead00ed75\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.initTrailerAndContext(HFileInfo.java:359) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.(HFileInfo.java:132) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreFileInfo.initHFileInfo(StoreFileInfo.java:763) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.open(HStoreFile.java:395) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.initReader(HStoreFile.java:524) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.createStoreFileAndReader(StoreEngine.java:226) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.lambda$openStoreFiles$0(StoreEngine.java:267) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) ~[?:?]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) ~[?:?]\r\n        ... 1 more\r\nCaused by: java.io.IOException: java.lang.ClassNotFoundException: org.apache.hadoop.hbase.KeyValue$KVComparator\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.getComparatorClass(FixedFileTrailer.java:578) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.deserializeFromPB(FixedFileTrailer.java:304) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.deserialize(FixedFileTrailer.java:250) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.readFromStream(FixedFileTrailer.java:407) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.initTrailerAndContext(HFileInfo.java:349) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.(HFileInfo.java:132) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreFileInfo.initHFileInfo(StoreFileInfo.java:763) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.open(HStoreFile.java:395) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.initReader(HStoreFile.java:524) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.createStoreFileAndReader(StoreEngine.java:226) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.lambda$openStoreFiles$0(StoreEngine.java:267) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) ~[?:?]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) ~[?:?]\r\n        ... 1 more\r\nCaused by: java.lang.ClassNotFoundException: org.apache.hadoop.hbase.KeyValue$KVComparator\r\n        at jdk.internal.loader.BuiltinClassLoader.loadClass(BuiltinClassLoader.java:641) ~[?:?]\r\n        at jdk.internal.loader.ClassLoaders$AppClassLoader.loadClass(ClassLoaders.java:188) ~[?:?]\r\n        at java.lang.ClassLoader.loadClass(ClassLoader.java:520) ~[?:?]\r\n        at java.lang.Class.forName0(Native Method) ~[?:?]\r\n        at java.lang.Class.forName(Class.java:375) ~[?:?]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.getComparatorClass(FixedFileTrailer.java:576) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.deserializeFromPB(FixedFileTrailer.java:304) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.deserialize(FixedFileTrailer.java:250) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.FixedFileTrailer.readFromStream(FixedFileTrailer.java:407) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.initTrailerAndContext(HFileInfo.java:349) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.io.hfile.HFileInfo.(HFileInfo.java:132) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreFileInfo.initHFileInfo(StoreFileInfo.java:763) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.open(HStoreFile.java:395) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.HStoreFile.initReader(HStoreFile.java:524) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.createStoreFileAndReader(StoreEngine.java:226) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at org.apache.hadoop.hbase.regionserver.StoreEngine.lambda$openStoreFiles$0(StoreEngine.java:267) ~[hbase-server-3.0.0-beta-2-SNAPSHOT.jar:3.0.0-beta-2-SNAPSHOT]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) ~[?:?]\r\n        at java.util.concurrent.FutureTask.run(FutureTask.java:264) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) ~[?:?]\r\n        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) ~[?:?]\r\n        ... 1 more {code}\r\nThis problem seems to be introduced recently, and I can still upgrade from 2.6.0 to 3.0.0 using the previous commits (E.g. commit from May 24: 516c89e8597fb6ed391f9e85e594f8b7e5b56e38). \r\n\r\nI have attached the hmaster log.","issue_id":"13590958","key":"HBASE-28812","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2024-09-16T14:11:25.000+0000","role":"fixed_distractor","summary":"Upgrade from 2.6.0 to 3.0.0 crashed"} {"case_id":"13597605","cluster":"DISTRACTOR-HBASE-28955","comments":[{"body":"Unfortunately this can only go into branch-3+, as the constructor for DFSOutputStream has been changed, and on branch-2 there is no way to not create the data streamer in DFSOutputStream.\r\n\r\nThanks [~stoty] for reviewing!","created":"2024-11-06T14:29:52.691+0000"},{"body":"I'm not sure I get it, [~zhangduo] .\r\nDFSOutputStream is coming from Hadoop.\r\n\r\nCan't we use our usual reflection tricks to search for the Hadoop 3 specific constuctor when creating the dummy DFSOutputStream, and fall back to passing null if it is not available (Since we don't need to pass it it in that case) ?","created":"2024-11-06T14:51:55.673+0000"},{"body":"The problem is for complilation. There is only one protected constructor for DFSOutputStream, and the parameters are different between hadoop2 and hadoop3. There is no way for us to write a code which could compile for both hadoop3 and hadoop2...\r\n\r\nA possible way is to separate the DummyDFSOutputStream to a separated module, and when compiling with hadoop2 we do not link it, or maybe we coud discuss again whether we still want to support hadoop2 for newer 2.x releases?\r\n\r\nThanks.","created":"2024-11-07T01:51:30.165+0000"},{"body":"Thanks for the explanation, I did not think that through. Some things cannot be band-aided with reflection.\r\n\r\nI think that at least 2.6.x should support Hadoop 3.4.1, unless we plan to EOL it and release 2.7.x without Hadoop2 support. (and if we solve that, there is no reason not backport that to 2.5)\r\n\r\nIn my mind, the thing keeping ppl on 2.x is going to be HBase 1.x API support, rather than Hadoop2 support, so dropping Hadoop2 is a possibility, but I think we should properly support Hadoop 3.4.1 ASAP in active releases, and dropping Hadoop2 support won't help with that.\r\n\r\nI think that adding an hbase-hadoop3-compat module is the best course of action.","created":"2024-11-07T06:39:24.730+0000"},{"body":"Will file a backport issue for testing possible solutions.","created":"2024-11-07T06:44:50.850+0000"},{"body":"Should we revert to Hadoop 3.4.0 or 3.3.5 on branch-2 in the meantime ?","created":"2024-11-07T07:12:47.288+0000"}],"conversations":[{"body":"When working with hadoop 3.4.x, we saw this in the stdout file\r\n\r\n{noformat}\r\nException in thread \"LeaseRenewer:zhangduo@home\" java.lang.NullPointerException: Cannot invoke \"org.apache.hadoop.hdfs.DFSOutputStream.getNamespace()\" because \"outputStream\" is null\r\n at org.apache.hadoop.hdfs.DFSClient.getNamespaces(DFSClient.java:596)\r\n at org.apache.hadoop.hdfs.DFSClient.renewLease(DFSClient.java:618)\r\n at org.apache.hadoop.hdfs.client.impl.LeaseRenewer.renew(LeaseRenewer.java:425)\r\n at org.apache.hadoop.hdfs.client.impl.LeaseRenewer.run(LeaseRenewer.java:445)\r\n at org.apache.hadoop.hdfs.client.impl.LeaseRenewer.access$800(LeaseRenewer.java:77)\r\n at org.apache.hadoop.hdfs.client.impl.LeaseRenewer$1.run(LeaseRenewer.java:336)\r\n at java.base/java.lang.Thread.run(Thread.java:840)\r\n{noformat}\r\n\r\nThis is because in newer DFSClient implementation, we need to pass namespace when renewer lease so we can not just pass null as DFSOutputStream when calling DFSClient.beginFileLease. We should find a way to deal with it.","from":"reporter","subject":"Improve lease renew for FanOutOneBlockAsyncDFSOutput"},{"body":"Unfortunately this can only go into branch-3+, as the constructor for DFSOutputStream has been changed, and on branch-2 there is no way to not create the data streamer in DFSOutputStream.\r\n\r\nThanks [~stoty] for reviewing!","from":"developer"},{"body":"I'm not sure I get it, [~zhangduo] .\r\nDFSOutputStream is coming from Hadoop.\r\n\r\nCan't we use our usual reflection tricks to search for the Hadoop 3 specific constuctor when creating the dummy DFSOutputStream, and fall back to passing null if it is not available (Since we don't need to pass it it in that case) ?","from":"developer"},{"body":"The problem is for complilation. There is only one protected constructor for DFSOutputStream, and the parameters are different between hadoop2 and hadoop3. There is no way for us to write a code which could compile for both hadoop3 and hadoop2...\r\n\r\nA possible way is to separate the DummyDFSOutputStream to a separated module, and when compiling with hadoop2 we do not link it, or maybe we coud discuss again whether we still want to support hadoop2 for newer 2.x releases?\r\n\r\nThanks.","from":"developer"},{"body":"Thanks for the explanation, I did not think that through. Some things cannot be band-aided with reflection.\r\n\r\nI think that at least 2.6.x should support Hadoop 3.4.1, unless we plan to EOL it and release 2.7.x without Hadoop2 support. (and if we solve that, there is no reason not backport that to 2.5)\r\n\r\nIn my mind, the thing keeping ppl on 2.x is going to be HBase 1.x API support, rather than Hadoop2 support, so dropping Hadoop2 is a possibility, but I think we should properly support Hadoop 3.4.1 ASAP in active releases, and dropping Hadoop2 support won't help with that.\r\n\r\nI think that adding an hbase-hadoop3-compat module is the best course of action.","from":"developer"},{"body":"Will file a backport issue for testing possible solutions.","from":"developer"},{"body":"Should we revert to Hadoop 3.4.0 or 3.3.5 on branch-2 in the meantime ?","from":"developer"}],"created":"2024-11-04T11:15:52.000+0000","description":"When working with hadoop 3.4.x, we saw this in the stdout file\r\n\r\n{noformat}\r\nException in thread \"LeaseRenewer:zhangduo@home\" java.lang.NullPointerException: Cannot invoke \"org.apache.hadoop.hdfs.DFSOutputStream.getNamespace()\" because \"outputStream\" is null\r\n at org.apache.hadoop.hdfs.DFSClient.getNamespaces(DFSClient.java:596)\r\n at org.apache.hadoop.hdfs.DFSClient.renewLease(DFSClient.java:618)\r\n at org.apache.hadoop.hdfs.client.impl.LeaseRenewer.renew(LeaseRenewer.java:425)\r\n at org.apache.hadoop.hdfs.client.impl.LeaseRenewer.run(LeaseRenewer.java:445)\r\n at org.apache.hadoop.hdfs.client.impl.LeaseRenewer.access$800(LeaseRenewer.java:77)\r\n at org.apache.hadoop.hdfs.client.impl.LeaseRenewer$1.run(LeaseRenewer.java:336)\r\n at java.base/java.lang.Thread.run(Thread.java:840)\r\n{noformat}\r\n\r\nThis is because in newer DFSClient implementation, we need to pass namespace when renewer lease so we can not just pass null as DFSOutputStream when calling DFSClient.beginFileLease. We should find a way to deal with it.","issue_id":"13597605","key":"HBASE-28955","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2024-11-06T14:29:52.000+0000","role":"fixed_distractor","summary":"Improve lease renew for FanOutOneBlockAsyncDFSOutput"} {"case_id":"13600308","cluster":"DISTRACTOR-HBASE-29003","comments":[{"body":"[~ndimiduk] [~rmattingly] [~bbeaudreault] - this ticket may interest you.","created":"2024-12-17T15:06:29.424+0000"},{"body":"For scenario 2, should we require that an incremental backup is not possible after a table is deleted? This seems like an error to me. Truncating a table or deleting and recreating it should invalidate any previous backups as the basis for new incremental backups. A backup of the previous instance of the table should not be applicable to a new instance of a table with the same name.","created":"2025-01-07T15:20:27.257+0000"},{"body":"That would be an a different approach to solve this issue, I assumed incremental backups were intended to work this way. What should be the restriction for changing the column families? Adding families is OK, but removing triggers the full-backup-required scenario?\r\n\r\nImplementing it that way would require more invasive work into the backup system table though, to track when incrementals are (not) allowed.\r\n\r\nI do wonder now, for scenario 2, whether the incremental logs (ie: all data besides bulk loads) will work as expected in this scenario, or also cause similar issues. Do you see issues in the current suggested PR [~ndimiduk] , or are you just thinking out loud about the best action?\r\n\r\nA similar no-incremental-backup-allowed mechanism can be used for HBASE-28084.","created":"2025-01-07T18:28:30.250+0000"},{"body":"bq. Do you see issues in the current suggested PR Nick Dimiduk , or are you just thinking out loud about the best action?\r\n\r\nI'm mostly talking out loud. I like the way that you propose scenarios and expected behaviors. In this case, your scenario strikes me as invalid, but I'm curious what other folks think.\r\n\r\nIs there a use-case where we'd want an incremental backup that follows a truncated table? Would the incremental include the truncation action?\r\n\r\nFor the case of deleting a table, we have no concept of an \"instance id\" of a table like we do with a region server's start timestamp. It seems like we'd need a similar mechanism in order to properly reason about acceptable actions pertaining to old instances of a table of a given name.","created":"2025-01-08T11:47:33.471+0000"},{"body":"I agree having a \"table instance id\" would be ideal. It would allow the backup system to check whether it is dealing with the same table as the previous backup. In the lack of this, using an observer to register table re-creation (as I do in the PR) is possible. Only downside of using an observer is that it needs to be configured by the user (and the fact that existing users might forget this).\r\n\r\nAn advantage of using an observer which does bookkeeping in the backup system table, is that the same bookkeeping system can be used by HBASE-28084 (= it should be impossible to create an incremental backup if you delete the latest incremental backup).","created":"2025-01-09T10:04:27.510+0000"},{"body":"Thank you [~dieterdp_ng]. Please note the creation of https://issues.apache.org/jira/browse/HBASE-29271, which I believe we should tackle as a follow up to this work.","created":"2025-04-27T18:45:13.603+0000"}],"conversations":[{"body":"As part of the incremental backup mechanism, HBase tracks which files were bulk-loaded (since the last backup).\r\n\r\nThis data is stored in the backup:system_bulk table. Entries are added when a bulk load occurs through the BackupObserver co-processor. Entries are deleted when an incremental backup is completed.\r\n\r\nThere are 2 flaws in this implementation:\r\n\r\n1) Performing a full backup should clear the list. Imagine following scenario:\r\n * Create a full backup B1 of table T.\r\n * Perform a bulk load L1.\r\n * Take a full backup B2 of table T.\r\n * Take an incremental backup of table T.\r\n ** The data stored for this backup will include L1, even though that data is already present due to B2. (This is an inefficiency, not a real error.)\r\n\r\n2) Performing a table deletion should clear the list of bulk-loaded files. Imagine the following scenario:\r\n * Create a full backup of table T.\r\n * Perform a bulk-load B1 into T.\r\n * Disable, delete and recreate T.\r\n * Create an incremental backup (taking a full backup instead is similar to the previous case)\r\n\r\n ** The backup will contain B1, even though it doesn't belong there.\r\n\r\n \r\n\r\nNote that this *can also cause backup corruption* after a backup restores (which is how we encountered this issue), which makes this problem less niche than the above scenarios indicate. Backup restore effectively uses bulk loads as well, so users could run into following scenario, where they are trying to restore data corruption:\r\n * (create an environment with backup B1 (time t), backup B2 (time t2 > t).\r\n * Users notice data corruption, and restore backup B2 after clearing the table\r\n * Users notice data corruption is already present, and restore backup B1 after clearing the table.\r\n * Users find data corruption solved, and resume regular backup cycle from here on.\r\n ** Any incremental backup taken will contain the (possible corrupt) data from B2 (due to the restore operation using bulk operations). The backups will be affected until a FULL backup is taken after an incremental backup (so this could span a period of weeks assuming bi-weekly/monthly full backups).\r\n\r\nA minimal reproduction example:\r\n{code:java}\r\necho \"create 'table', 'cf'; put 'table', 'row1', 'cf:a', 'value1', 1400523142819\" | bin/hbase shell -n\r\nbin/hbase backup create full file:/tmp/backup -t table -i\r\necho \"disable 'table'; drop 'table'\" | bin/hbase shell -n\r\n# Empty\r\necho \"scan 'backup:system_bulk'\" | bin/hbase shell -n\r\n\r\nbin/hbase restore file:/tmp/backup backup_1732787972748 -t \"table\"\r\n# 1 entry\r\necho \"scan 'backup:system_bulk'\" | bin/hbase shell -n\r\n\r\necho \"disable 'table'; drop 'table'\" | bin/hbase shell -n\r\n# 1 entry\r\necho \"scan 'backup:system_bulk'\" | bin/hbase shell -necho \"create 'table', 'cf'; put 'table', 'row1', 'cf:b', 'value2', 1400523142819\" | bin/hbase shell -n\r\nbin/hbase backup create full file:/tmp/backup -t table -i\r\necho \"scan 'backup:system_bulk'\" | bin/hbase shell -n\r\n\r\necho \"put 'table', 'row1', 'cf:b', 'value3', 1400523142819\" | bin/hbase shell -n\r\nbin/hbase backup create incremental file:/tmp/backup -t table -i\r\n# Emtpy\r\necho \"scan 'backup:system_bulk'\" | bin/hbase shell -n\r\n\r\necho \"disable 'table'; drop 'table'\" | bin/hbase shell -n\r\nbin/hbase restore file:/tmp/backup backup_1732788098586 -t \"table\"\r\n\r\n# Will contain \"value1\" (unexpected) and \"value3\" (expected)\r\necho \"scan 'table'\" | bin/hbase shell -n\r\n {code}","from":"reporter","subject":"Proper bulk load tracking"},{"body":"[~ndimiduk] [~rmattingly] [~bbeaudreault] - this ticket may interest you.","from":"developer"},{"body":"For scenario 2, should we require that an incremental backup is not possible after a table is deleted? This seems like an error to me. Truncating a table or deleting and recreating it should invalidate any previous backups as the basis for new incremental backups. A backup of the previous instance of the table should not be applicable to a new instance of a table with the same name.","from":"developer"},{"body":"That would be an a different approach to solve this issue, I assumed incremental backups were intended to work this way. What should be the restriction for changing the column families? Adding families is OK, but removing triggers the full-backup-required scenario?\r\n\r\nImplementing it that way would require more invasive work into the backup system table though, to track when incrementals are (not) allowed.\r\n\r\nI do wonder now, for scenario 2, whether the incremental logs (ie: all data besides bulk loads) will work as expected in this scenario, or also cause similar issues. Do you see issues in the current suggested PR [~ndimiduk] , or are you just thinking out loud about the best action?\r\n\r\nA similar no-incremental-backup-allowed mechanism can be used for HBASE-28084.","from":"developer"},{"body":"bq. Do you see issues in the current suggested PR Nick Dimiduk , or are you just thinking out loud about the best action?\r\n\r\nI'm mostly talking out loud. I like the way that you propose scenarios and expected behaviors. In this case, your scenario strikes me as invalid, but I'm curious what other folks think.\r\n\r\nIs there a use-case where we'd want an incremental backup that follows a truncated table? Would the incremental include the truncation action?\r\n\r\nFor the case of deleting a table, we have no concept of an \"instance id\" of a table like we do with a region server's start timestamp. It seems like we'd need a similar mechanism in order to properly reason about acceptable actions pertaining to old instances of a table of a given name.","from":"developer"},{"body":"I agree having a \"table instance id\" would be ideal. It would allow the backup system to check whether it is dealing with the same table as the previous backup. In the lack of this, using an observer to register table re-creation (as I do in the PR) is possible. Only downside of using an observer is that it needs to be configured by the user (and the fact that existing users might forget this).\r\n\r\nAn advantage of using an observer which does bookkeeping in the backup system table, is that the same bookkeeping system can be used by HBASE-28084 (= it should be impossible to create an incremental backup if you delete the latest incremental backup).","from":"developer"},{"body":"Thank you [~dieterdp_ng]. Please note the creation of https://issues.apache.org/jira/browse/HBASE-29271, which I believe we should tackle as a follow up to this work.","from":"developer"}],"created":"2024-11-28T13:10:51.000+0000","description":"As part of the incremental backup mechanism, HBase tracks which files were bulk-loaded (since the last backup).\r\n\r\nThis data is stored in the backup:system_bulk table. Entries are added when a bulk load occurs through the BackupObserver co-processor. Entries are deleted when an incremental backup is completed.\r\n\r\nThere are 2 flaws in this implementation:\r\n\r\n1) Performing a full backup should clear the list. Imagine following scenario:\r\n * Create a full backup B1 of table T.\r\n * Perform a bulk load L1.\r\n * Take a full backup B2 of table T.\r\n * Take an incremental backup of table T.\r\n ** The data stored for this backup will include L1, even though that data is already present due to B2. (This is an inefficiency, not a real error.)\r\n\r\n2) Performing a table deletion should clear the list of bulk-loaded files. Imagine the following scenario:\r\n * Create a full backup of table T.\r\n * Perform a bulk-load B1 into T.\r\n * Disable, delete and recreate T.\r\n * Create an incremental backup (taking a full backup instead is similar to the previous case)\r\n\r\n ** The backup will contain B1, even though it doesn't belong there.\r\n\r\n \r\n\r\nNote that this *can also cause backup corruption* after a backup restores (which is how we encountered this issue), which makes this problem less niche than the above scenarios indicate. Backup restore effectively uses bulk loads as well, so users could run into following scenario, where they are trying to restore data corruption:\r\n * (create an environment with backup B1 (time t), backup B2 (time t2 > t).\r\n * Users notice data corruption, and restore backup B2 after clearing the table\r\n * Users notice data corruption is already present, and restore backup B1 after clearing the table.\r\n * Users find data corruption solved, and resume regular backup cycle from here on.\r\n ** Any incremental backup taken will contain the (possible corrupt) data from B2 (due to the restore operation using bulk operations). The backups will be affected until a FULL backup is taken after an incremental backup (so this could span a period of weeks assuming bi-weekly/monthly full backups).\r\n\r\nA minimal reproduction example:\r\n{code:java}\r\necho \"create 'table', 'cf'; put 'table', 'row1', 'cf:a', 'value1', 1400523142819\" | bin/hbase shell -n\r\nbin/hbase backup create full file:/tmp/backup -t table -i\r\necho \"disable 'table'; drop 'table'\" | bin/hbase shell -n\r\n# Empty\r\necho \"scan 'backup:system_bulk'\" | bin/hbase shell -n\r\n\r\nbin/hbase restore file:/tmp/backup backup_1732787972748 -t \"table\"\r\n# 1 entry\r\necho \"scan 'backup:system_bulk'\" | bin/hbase shell -n\r\n\r\necho \"disable 'table'; drop 'table'\" | bin/hbase shell -n\r\n# 1 entry\r\necho \"scan 'backup:system_bulk'\" | bin/hbase shell -necho \"create 'table', 'cf'; put 'table', 'row1', 'cf:b', 'value2', 1400523142819\" | bin/hbase shell -n\r\nbin/hbase backup create full file:/tmp/backup -t table -i\r\necho \"scan 'backup:system_bulk'\" | bin/hbase shell -n\r\n\r\necho \"put 'table', 'row1', 'cf:b', 'value3', 1400523142819\" | bin/hbase shell -n\r\nbin/hbase backup create incremental file:/tmp/backup -t table -i\r\n# Emtpy\r\necho \"scan 'backup:system_bulk'\" | bin/hbase shell -n\r\n\r\necho \"disable 'table'; drop 'table'\" | bin/hbase shell -n\r\nbin/hbase restore file:/tmp/backup backup_1732788098586 -t \"table\"\r\n\r\n# Will contain \"value1\" (unexpected) and \"value3\" (expected)\r\necho \"scan 'table'\" | bin/hbase shell -n\r\n {code}","issue_id":"13600308","key":"HBASE-29003","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2025-04-27T18:45:13.000+0000","role":"fixed_distractor","summary":"Proper bulk load tracking"} {"case_id":"12471578","cluster":"DISTRACTOR-HBASE-2915","comments":[{"body":"There is another deadlock that needs fixing in the scope of this jira. Since the split code was redone, there's a deadlock when SplitTransaction acquires the splitAndCloses writelock while a flush is running for it. It looks like:\n\n{noformat}\n\n\"regionserver60021.compactor\" daemon prio=10 tid=0x00007fc31845b800 nid=0x5f62 in Object.wait() [0x00007fc31e9e7000]\n java.lang.Thread.State: WAITING (on object monitor)\n\tat java.lang.Object.wait(Native Method)\n\tat java.lang.Object.wait(Object.java:485)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.close(HRegion.java:493)\n\t- locked <0x00007fc336877998> (a org.apache.hadoop.hbase.regionserver.HRegion$WriteState)\n\tat org.apache.hadoop.hbase.regionserver.SplitTransaction.execute(SplitTransaction.java:213)\n\tat org.apache.hadoop.hbase.regionserver.SplitTransaction.execute(SplitTransaction.java:186)\n\tat org.apache.hadoop.hbase.regionserver.CompactSplitThread.split(CompactSplitThread.java:157)\n\tat org.apache.hadoop.hbase.regionserver.CompactSplitThread.run(CompactSplitThread.java:87)\n\n\"regionserver60021.cacheFlusher\" daemon prio=10 tid=0x00007fc31845a000 nid=0x5f61 waiting on condition [0x00007fc31eae8000]\n java.lang.Thread.State: WAITING (parking)\n\tat sun.misc.Unsafe.park(Native Method)\n\t- parking to wait for <0x00007fc336561750> (a java.util.concurrent.locks.ReentrantReadWriteLock$NonfairSync)\n\tat java.util.concurrent.locks.LockSupport.park(LockSupport.java:158)\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer.parkAndCheckInterrupt(AbstractQueuedSynchronizer.java:747)\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer.doAcquireShared(AbstractQueuedSynchronizer.java:877)\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer.acquireShared(AbstractQueuedSynchronizer.java:1197)\n\tat java.util.concurrent.locks.ReentrantReadWriteLock$ReadLock.lock(ReentrantReadWriteLock.java:594)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:793)\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:249)\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:223)\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.run(MemStoreFlusher.java:146)\n{noformat}","created":"2010-08-19T17:21:19.271+0000"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/\n-----------------------------------------------------------\n\nReview request for hbase.\n\n\nSummary\n-------\n\nThis patch removes newScannerLock and renames splitAndClose lock to just \"lock\". Every operation is now required to obtain the read lock on \"lock\" before doing anything (including getting a row lock). This is done by calling openRegionTransaction inside a try statement and by calling closeRegionTransaction in finally.\n\nflushcache got refactored some more in order to do the locking in the proper order; first get the read lock, then do the writestate handling.\n\nFinally, it removes the need to have a writeLock when flushing when subclassers give atomic work do to via internalPreFlushcacheCommit. This means that this patch breaks external contribs. This is required to keep our whole locking mechanism simpler.\n\n\nThis addresses bug HBASE-2915.\n http://issues.apache.org/jira/browse/HBASE-2915\n\n\nDiffs\n-----\n\n /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java 987300 \n /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java 987300 \n /trunk/src/test/java/org/apache/hadoop/hbase/regionserver/TestSplitTransaction.java 987300 \n\nDiff: http://review.cloudera.org/r/691/diff\n\n\nTesting\n-------\n\n5 concurrent ICV threads + randomWrite 3 + scans on a single RS. I'm also in the process of deploying it on a cluster.\n\n\nThanks,\n\nJean-Daniel\n\n\n","created":"2010-08-19T22:13:58.075+0000"},{"body":"Message from: \"Ryan Rawson\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review966\n-----------------------------------------------------------\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n oh wow i cant believe this was ever here\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n i thought we agreed that the closing flag had to be set BEFORE the write lock was acquired to prevent race conditions?\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n lets just excise this and break compile time compatibility. Also remove internalPreFlushcacheCommit too\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n ditto remove this whole try/finally bit\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n maybe we shouldnt call this 'transaction' might confuse people into thinking we support real transactions... not sure what to call it at this moment tho\n\n\n- Ryan\n\n\n\n","created":"2010-08-19T22:52:46.597+0000"},{"body":"Message from: \"Ted Yu\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review967\n-----------------------------------------------------------\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n How about naming this method openRegionProlog ?\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Please add:\n It has to be called inside the corresponding finally block\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n How about naming this method closeRegionEpilog ?\n\n\n- Ted\n\n\n\n","created":"2010-08-19T22:59:36.300+0000"},{"body":"Message from: \"Ted Yu\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review969\n-----------------------------------------------------------\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Or regionOperationProlog()\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n and regionOperationEpilog()\n\n\n- Ted\n\n\n\n","created":"2010-08-19T23:33:35.074+0000"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review973\n-----------------------------------------------------------\n\n\nFurther testing shows that I completely forgot to properly lock incrementColumnValue, thus missing the point of the title of the jira :S Uploading a new patch soon that addresses that, along with comments from the reviews.\n\n- Jean-Daniel\n\n\n\n","created":"2010-08-20T22:22:55.695+0000"},{"body":"Message from: stack@duboce.net\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review975\n-----------------------------------------------------------\n\n\noh, one other thing, in discussions, we talked of no longer needing to wait on row locks to expire... I don't see this being excised from the close method. Should that be in here?\n\n- stack\n\n\n\n","created":"2010-08-20T23:19:41.564+0000"},{"body":"Message from: stack@duboce.net\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\nShip it!\n\n\n+1 Nice fix. A few comments below.\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n I suppose this order is ok if the first thing we do on entrance to HRegion is get the read lock before check of closing.\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Yeah, just remove.\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Aren't these lines unnecessary? openRegionTransaction does it?\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n So, this javadoc is good but do you think we need some more doc? Does there need to be more detail on new locking regime? Maybe there is no more to be said that what is here in this paragraph. You've done all the work unravelling our lock mess. With time your nice unravelling will rot unless its clear what the pattern is. I'm just trying to think of ways of preventing that happening.\n\n\n- stack\n\n\n\n","created":"2010-08-20T23:24:13.729+0000"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 507\nbq. > \nbq. >\nbq. > I suppose this order is ok if the first thing we do on entrance to HRegion is get the read lock before check of closing.\n\nSo I just redid that part. setClosing is first taken so that when the client threads arrive they can fast fail by looking at closing.get before trying to get the readLock.\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 712\nbq. > \nbq. >\nbq. > Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\n\nYeah the issue with compact and flush is that the callers don't expect to see NSRE, the want null values.\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 1115\nbq. > \nbq. >\nbq. > Aren't these lines unnecessary? openRegionTransaction does it?\n\nGood catch.\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 3142\nbq. > \nbq. >\nbq. > So, this javadoc is good but do you think we need some more doc? Does there need to be more detail on new locking regime? Maybe there is no more to be said that what is here in this paragraph. You've done all the work unravelling our lock mess. With time your nice unravelling will rot unless its clear what the pattern is. I'm just trying to think of ways of preventing that happening.\n\nYeah I'll included some more javadoc, maybe with code examples?\n\n\n- Jean-Daniel\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\n\n\n","created":"2010-08-20T23:32:56.968+0000"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n\nbq. On 2010-08-20 15:58:23, stack wrote:\nbq. > oh, one other thing, in discussions, we talked of no longer needing to wait on row locks to expire... I don't see this being excised from the close method. Should that be in here?\n\nYep, it needs to go.\n\n\n- Jean-Daniel\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review975\n-----------------------------------------------------------\n\n\n\n","created":"2010-08-20T23:35:36.748+0000"},{"body":"Message from: stack@duboce.net\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 507\nbq. > \nbq. >\nbq. > I suppose this order is ok if the first thing we do on entrance to HRegion is get the read lock before check of closing.\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. So I just redid that part. setClosing is first taken so that when the client threads arrive they can fast fail by looking at closing.get before trying to get the readLock.\n\nDon't you have to check again the setClosing after you get the read lock?\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 712\nbq. > \nbq. >\nbq. > Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. Yeah the issue with compact and flush is that the callers don't expect to see NSRE, the want null values.\n\nOK. Not important. This is deep internal stuff or make a version that takes a flag on whether to throw exception (default throws exception .. might get messy though... not important).\n\n\n- stack\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\n\n\n","created":"2010-08-20T23:40:58.621+0000"},{"body":"Message from: stack@duboce.net\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 507\nbq. > \nbq. >\nbq. > I suppose this order is ok if the first thing we do on entrance to HRegion is get the read lock before check of closing.\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. So I just redid that part. setClosing is first taken so that when the client threads arrive they can fast fail by looking at closing.get before trying to get the readLock.\nbq. \nbq. stack wrote:\nbq. Don't you have to check again the setClosing after you get the read lock?\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. I'm about to post a new patch, but it looks like this and it has to be called just before \"try\" instead of inside:\nbq. if (this.closing.get()) {\nbq. throw new NotServingRegionException(regionInfo.getRegionNameAsString() +\nbq. \" is closing\");\nbq. }\nbq. lock.readLock().lock();\nbq. if (this.closed.get()) {\nbq. throw new NotServingRegionException(regionInfo.getRegionNameAsString() +\nbq. \" is closed\");\nbq. }\n\nThat looks right.\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 712\nbq. > \nbq. >\nbq. > Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. Yeah the issue with compact and flush is that the callers don't expect to see NSRE, the want null values.\nbq. \nbq. stack wrote:\nbq. OK. Not important. This is deep internal stuff or make a version that takes a flag on whether to throw exception (default throws exception .. might get messy though... not important).\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. I'm afraid those little methods could be clogged fast.\n\nYeah. Not important.\n\n\n- stack\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\n\n\n","created":"2010-08-20T23:44:42.988+0000"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 507\nbq. > \nbq. >\nbq. > I suppose this order is ok if the first thing we do on entrance to HRegion is get the read lock before check of closing.\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. So I just redid that part. setClosing is first taken so that when the client threads arrive they can fast fail by looking at closing.get before trying to get the readLock.\nbq. \nbq. stack wrote:\nbq. Don't you have to check again the setClosing after you get the read lock?\n\nI'm about to post a new patch, but it looks like this and it has to be called just before \"try\" instead of inside:\n if (this.closing.get()) {\n throw new NotServingRegionException(regionInfo.getRegionNameAsString() +\n \" is closing\");\n }\n lock.readLock().lock();\n if (this.closed.get()) {\n throw new NotServingRegionException(regionInfo.getRegionNameAsString() +\n \" is closed\");\n }\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 712\nbq. > \nbq. >\nbq. > Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. Yeah the issue with compact and flush is that the callers don't expect to see NSRE, the want null values.\nbq. \nbq. stack wrote:\nbq. OK. Not important. This is deep internal stuff or make a version that takes a flag on whether to throw exception (default throws exception .. might get messy though... not important).\n\nI'm afraid those little methods could be clogged fast.\n\n\n- Jean-Daniel\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\n\n\n","created":"2010-08-20T23:50:28.797+0000"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 712\nbq. > \nbq. >\nbq. > Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. Yeah the issue with compact and flush is that the callers don't expect to see NSRE, the want null values.\nbq. \nbq. stack wrote:\nbq. OK. Not important. This is deep internal stuff or make a version that takes a flag on whether to throw exception (default throws exception .. might get messy though... not important).\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. I'm afraid those little methods could be clogged fast.\nbq. \nbq. stack wrote:\nbq. Yeah. Not important.\n\nMoreover, and this should be added to the documentation, none of the user-facing methods now directly use the lock, it has to go through those new methods. Means that flush, compact, split and close are the only ones with direct access.\n\n\n- Jean-Daniel\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\n\n\n","created":"2010-08-20T23:53:42.612+0000"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/\n-----------------------------------------------------------\n\n(Updated 2010-08-20 17:43:52.668799)\n\n\nReview request for hbase.\n\n\nChanges\n-------\n\nUpdated patch with \"closing\" handling and ICV locking, still needs more javadoc.\n\n\nSummary\n-------\n\nThis patch removes newScannerLock and renames splitAndClose lock to just \"lock\". Every operation is now required to obtain the read lock on \"lock\" before doing anything (including getting a row lock). This is done by calling openRegionTransaction inside a try statement and by calling closeRegionTransaction in finally.\n\nflushcache got refactored some more in order to do the locking in the proper order; first get the read lock, then do the writestate handling.\n\nFinally, it removes the need to have a writeLock when flushing when subclassers give atomic work do to via internalPreFlushcacheCommit. This means that this patch breaks external contribs. This is required to keep our whole locking mechanism simpler.\n\n\nThis addresses bug HBASE-2915.\n http://issues.apache.org/jira/browse/HBASE-2915\n\n\nDiffs (updated)\n-----\n\n /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java 987355 \n /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java 987355 \n /trunk/src/test/java/org/apache/hadoop/hbase/regionserver/TestSplitTransaction.java 987355 \n\nDiff: http://review.cloudera.org/r/691/diff\n\n\nTesting\n-------\n\n5 concurrent ICV threads + randomWrite 3 + scans on a single RS. I'm also in the process of deploying it on a cluster.\n\n\nThanks,\n\nJean-Daniel\n\n\n","created":"2010-08-21T00:59:56.953+0000"},{"body":"Message from: stack@duboce.net\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review988\n-----------------------------------------------------------\n\nShip it!\n\n\n+1\n\nif below are issues address on commit.\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Any reason this does not follow the pattern used elsewhere? You are not testing closing before getting the read lock?\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n ditto\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Is this supposed to be inside the try?\n\n\n- stack\n\n\n\n","created":"2010-08-21T04:10:36.739+0000"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n\nbq. On 2010-08-20 20:49:15, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 716\nbq. > \nbq. >\nbq. > Any reason this does not follow the pattern used elsewhere? You are not testing closing before getting the read lock?\n\nmmm didn't want to copy the same chunk of code over there, but yeah it will also give us a fail-fast behavior which will speed up flushing.\n\n\nbq. On 2010-08-20 20:49:15, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 785\nbq. > \nbq. >\nbq. > ditto\n\nditto\n\n\nbq. On 2010-08-20 20:49:15, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 2960\nbq. > \nbq. >\nbq. > Is this supposed to be inside the try?\n\nPer that method's javadoc, no. The reason is that if it's all in a try block, and that closing is true, then you won't hold a readLock and you can't do a isHeldByCurrent thread on that lock. So I thought we could catch IllegalStateException in closeRegionOperation, but that would be really ugly. This leaves us to the current situation where if an exception is thrown when we are on this line:\n if (this.closed.get()) {\nthat it would leave the readLock locked by the thread, although in this case the region is closing so we're getting rid of it anyways. Although, thinking about it, I should probably do this instead:\n\n if (this.closed.get()) {\n lock.readLock().unlock();\n throw new NotServingRegionException(regionInfo.getRegionNameAsString() +\n \" is closed\");\n }\n\n\n- Jean-Daniel\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review988\n-----------------------------------------------------------\n\n\n\n","created":"2010-08-23T17:00:25.162+0000"},{"body":"This patch is the same as the last one I posted on rb, with the addition of the unlocking in case the region is closed but the thread was able to get a readLock. I tested it on a 13 nodes cluster with 200 YCSB threads hitting it in different ways, didn't see any deadlock. Going to commit.","created":"2010-08-23T22:43:26.343+0000"},{"body":"Committed to trunk, thanks for the review Stack!","created":"2010-08-23T22:49:59.689+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T12:40:34.093+0000"}],"conversations":[{"body":"HRegion.ICV gets a row lock then gets a newScanner lock.\n\nHRegion.close gets a newScanner lock, slitCloseLock and finally waits for all row locks to finish.\n\nIf the ICV got the row lock and then close got the newScannerLock, both end up waiting on the other. This was introduced when Get became a Scan.\n\nStack thinks we can get rid of the newScannerLock in close since we setClosing to true.","from":"reporter","subject":"Deadlock between HRegion.ICV and HRegion.close"},{"body":"There is another deadlock that needs fixing in the scope of this jira. Since the split code was redone, there's a deadlock when SplitTransaction acquires the splitAndCloses writelock while a flush is running for it. It looks like:\n\n{noformat}\n\n\"regionserver60021.compactor\" daemon prio=10 tid=0x00007fc31845b800 nid=0x5f62 in Object.wait() [0x00007fc31e9e7000]\n java.lang.Thread.State: WAITING (on object monitor)\n\tat java.lang.Object.wait(Native Method)\n\tat java.lang.Object.wait(Object.java:485)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.close(HRegion.java:493)\n\t- locked <0x00007fc336877998> (a org.apache.hadoop.hbase.regionserver.HRegion$WriteState)\n\tat org.apache.hadoop.hbase.regionserver.SplitTransaction.execute(SplitTransaction.java:213)\n\tat org.apache.hadoop.hbase.regionserver.SplitTransaction.execute(SplitTransaction.java:186)\n\tat org.apache.hadoop.hbase.regionserver.CompactSplitThread.split(CompactSplitThread.java:157)\n\tat org.apache.hadoop.hbase.regionserver.CompactSplitThread.run(CompactSplitThread.java:87)\n\n\"regionserver60021.cacheFlusher\" daemon prio=10 tid=0x00007fc31845a000 nid=0x5f61 waiting on condition [0x00007fc31eae8000]\n java.lang.Thread.State: WAITING (parking)\n\tat sun.misc.Unsafe.park(Native Method)\n\t- parking to wait for <0x00007fc336561750> (a java.util.concurrent.locks.ReentrantReadWriteLock$NonfairSync)\n\tat java.util.concurrent.locks.LockSupport.park(LockSupport.java:158)\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer.parkAndCheckInterrupt(AbstractQueuedSynchronizer.java:747)\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer.doAcquireShared(AbstractQueuedSynchronizer.java:877)\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer.acquireShared(AbstractQueuedSynchronizer.java:1197)\n\tat java.util.concurrent.locks.ReentrantReadWriteLock$ReadLock.lock(ReentrantReadWriteLock.java:594)\n\tat org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:793)\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:249)\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:223)\n\tat org.apache.hadoop.hbase.regionserver.MemStoreFlusher.run(MemStoreFlusher.java:146)\n{noformat}","from":"developer"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/\n-----------------------------------------------------------\n\nReview request for hbase.\n\n\nSummary\n-------\n\nThis patch removes newScannerLock and renames splitAndClose lock to just \"lock\". Every operation is now required to obtain the read lock on \"lock\" before doing anything (including getting a row lock). This is done by calling openRegionTransaction inside a try statement and by calling closeRegionTransaction in finally.\n\nflushcache got refactored some more in order to do the locking in the proper order; first get the read lock, then do the writestate handling.\n\nFinally, it removes the need to have a writeLock when flushing when subclassers give atomic work do to via internalPreFlushcacheCommit. This means that this patch breaks external contribs. This is required to keep our whole locking mechanism simpler.\n\n\nThis addresses bug HBASE-2915.\n http://issues.apache.org/jira/browse/HBASE-2915\n\n\nDiffs\n-----\n\n /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java 987300 \n /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java 987300 \n /trunk/src/test/java/org/apache/hadoop/hbase/regionserver/TestSplitTransaction.java 987300 \n\nDiff: http://review.cloudera.org/r/691/diff\n\n\nTesting\n-------\n\n5 concurrent ICV threads + randomWrite 3 + scans on a single RS. I'm also in the process of deploying it on a cluster.\n\n\nThanks,\n\nJean-Daniel\n\n\n","from":"developer"},{"body":"Message from: \"Ryan Rawson\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review966\n-----------------------------------------------------------\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n oh wow i cant believe this was ever here\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n i thought we agreed that the closing flag had to be set BEFORE the write lock was acquired to prevent race conditions?\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n lets just excise this and break compile time compatibility. Also remove internalPreFlushcacheCommit too\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n ditto remove this whole try/finally bit\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n maybe we shouldnt call this 'transaction' might confuse people into thinking we support real transactions... not sure what to call it at this moment tho\n\n\n- Ryan\n\n\n\n","from":"developer"},{"body":"Message from: \"Ted Yu\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review967\n-----------------------------------------------------------\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n How about naming this method openRegionProlog ?\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Please add:\n It has to be called inside the corresponding finally block\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n How about naming this method closeRegionEpilog ?\n\n\n- Ted\n\n\n\n","from":"developer"},{"body":"Message from: \"Ted Yu\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review969\n-----------------------------------------------------------\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Or regionOperationProlog()\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n and regionOperationEpilog()\n\n\n- Ted\n\n\n\n","from":"developer"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review973\n-----------------------------------------------------------\n\n\nFurther testing shows that I completely forgot to properly lock incrementColumnValue, thus missing the point of the title of the jira :S Uploading a new patch soon that addresses that, along with comments from the reviews.\n\n- Jean-Daniel\n\n\n\n","from":"developer"},{"body":"Message from: stack@duboce.net\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review975\n-----------------------------------------------------------\n\n\noh, one other thing, in discussions, we talked of no longer needing to wait on row locks to expire... I don't see this being excised from the close method. Should that be in here?\n\n- stack\n\n\n\n","from":"developer"},{"body":"Message from: stack@duboce.net\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\nShip it!\n\n\n+1 Nice fix. A few comments below.\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n I suppose this order is ok if the first thing we do on entrance to HRegion is get the read lock before check of closing.\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Yeah, just remove.\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Aren't these lines unnecessary? openRegionTransaction does it?\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n So, this javadoc is good but do you think we need some more doc? Does there need to be more detail on new locking regime? Maybe there is no more to be said that what is here in this paragraph. You've done all the work unravelling our lock mess. With time your nice unravelling will rot unless its clear what the pattern is. I'm just trying to think of ways of preventing that happening.\n\n\n- stack\n\n\n\n","from":"developer"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 507\nbq. > \nbq. >\nbq. > I suppose this order is ok if the first thing we do on entrance to HRegion is get the read lock before check of closing.\n\nSo I just redid that part. setClosing is first taken so that when the client threads arrive they can fast fail by looking at closing.get before trying to get the readLock.\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 712\nbq. > \nbq. >\nbq. > Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\n\nYeah the issue with compact and flush is that the callers don't expect to see NSRE, the want null values.\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 1115\nbq. > \nbq. >\nbq. > Aren't these lines unnecessary? openRegionTransaction does it?\n\nGood catch.\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 3142\nbq. > \nbq. >\nbq. > So, this javadoc is good but do you think we need some more doc? Does there need to be more detail on new locking regime? Maybe there is no more to be said that what is here in this paragraph. You've done all the work unravelling our lock mess. With time your nice unravelling will rot unless its clear what the pattern is. I'm just trying to think of ways of preventing that happening.\n\nYeah I'll included some more javadoc, maybe with code examples?\n\n\n- Jean-Daniel\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\n\n\n","from":"developer"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n\nbq. On 2010-08-20 15:58:23, stack wrote:\nbq. > oh, one other thing, in discussions, we talked of no longer needing to wait on row locks to expire... I don't see this being excised from the close method. Should that be in here?\n\nYep, it needs to go.\n\n\n- Jean-Daniel\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review975\n-----------------------------------------------------------\n\n\n\n","from":"developer"},{"body":"Message from: stack@duboce.net\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 507\nbq. > \nbq. >\nbq. > I suppose this order is ok if the first thing we do on entrance to HRegion is get the read lock before check of closing.\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. So I just redid that part. setClosing is first taken so that when the client threads arrive they can fast fail by looking at closing.get before trying to get the readLock.\n\nDon't you have to check again the setClosing after you get the read lock?\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 712\nbq. > \nbq. >\nbq. > Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. Yeah the issue with compact and flush is that the callers don't expect to see NSRE, the want null values.\n\nOK. Not important. This is deep internal stuff or make a version that takes a flag on whether to throw exception (default throws exception .. might get messy though... not important).\n\n\n- stack\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\n\n\n","from":"developer"},{"body":"Message from: stack@duboce.net\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 507\nbq. > \nbq. >\nbq. > I suppose this order is ok if the first thing we do on entrance to HRegion is get the read lock before check of closing.\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. So I just redid that part. setClosing is first taken so that when the client threads arrive they can fast fail by looking at closing.get before trying to get the readLock.\nbq. \nbq. stack wrote:\nbq. Don't you have to check again the setClosing after you get the read lock?\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. I'm about to post a new patch, but it looks like this and it has to be called just before \"try\" instead of inside:\nbq. if (this.closing.get()) {\nbq. throw new NotServingRegionException(regionInfo.getRegionNameAsString() +\nbq. \" is closing\");\nbq. }\nbq. lock.readLock().lock();\nbq. if (this.closed.get()) {\nbq. throw new NotServingRegionException(regionInfo.getRegionNameAsString() +\nbq. \" is closed\");\nbq. }\n\nThat looks right.\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 712\nbq. > \nbq. >\nbq. > Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. Yeah the issue with compact and flush is that the callers don't expect to see NSRE, the want null values.\nbq. \nbq. stack wrote:\nbq. OK. Not important. This is deep internal stuff or make a version that takes a flag on whether to throw exception (default throws exception .. might get messy though... not important).\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. I'm afraid those little methods could be clogged fast.\n\nYeah. Not important.\n\n\n- stack\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\n\n\n","from":"developer"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 507\nbq. > \nbq. >\nbq. > I suppose this order is ok if the first thing we do on entrance to HRegion is get the read lock before check of closing.\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. So I just redid that part. setClosing is first taken so that when the client threads arrive they can fast fail by looking at closing.get before trying to get the readLock.\nbq. \nbq. stack wrote:\nbq. Don't you have to check again the setClosing after you get the read lock?\n\nI'm about to post a new patch, but it looks like this and it has to be called just before \"try\" instead of inside:\n if (this.closing.get()) {\n throw new NotServingRegionException(regionInfo.getRegionNameAsString() +\n \" is closing\");\n }\n lock.readLock().lock();\n if (this.closed.get()) {\n throw new NotServingRegionException(regionInfo.getRegionNameAsString() +\n \" is closed\");\n }\n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 712\nbq. > \nbq. >\nbq. > Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. Yeah the issue with compact and flush is that the callers don't expect to see NSRE, the want null values.\nbq. \nbq. stack wrote:\nbq. OK. Not important. This is deep internal stuff or make a version that takes a flag on whether to throw exception (default throws exception .. might get messy though... not important).\n\nI'm afraid those little methods could be clogged fast.\n\n\n- Jean-Daniel\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\n\n\n","from":"developer"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n\nbq. On 2010-08-20 15:57:34, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 712\nbq. > \nbq. >\nbq. > Seems like you could use your opentransaction/closetransaction methods here and in flush too to be consistent?\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. Yeah the issue with compact and flush is that the callers don't expect to see NSRE, the want null values.\nbq. \nbq. stack wrote:\nbq. OK. Not important. This is deep internal stuff or make a version that takes a flag on whether to throw exception (default throws exception .. might get messy though... not important).\nbq. \nbq. Jean-Daniel Cryans wrote:\nbq. I'm afraid those little methods could be clogged fast.\nbq. \nbq. stack wrote:\nbq. Yeah. Not important.\n\nMoreover, and this should be added to the documentation, none of the user-facing methods now directly use the lock, it has to go through those new methods. Means that flush, compact, split and close are the only ones with direct access.\n\n\n- Jean-Daniel\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review974\n-----------------------------------------------------------\n\n\n\n","from":"developer"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/\n-----------------------------------------------------------\n\n(Updated 2010-08-20 17:43:52.668799)\n\n\nReview request for hbase.\n\n\nChanges\n-------\n\nUpdated patch with \"closing\" handling and ICV locking, still needs more javadoc.\n\n\nSummary\n-------\n\nThis patch removes newScannerLock and renames splitAndClose lock to just \"lock\". Every operation is now required to obtain the read lock on \"lock\" before doing anything (including getting a row lock). This is done by calling openRegionTransaction inside a try statement and by calling closeRegionTransaction in finally.\n\nflushcache got refactored some more in order to do the locking in the proper order; first get the read lock, then do the writestate handling.\n\nFinally, it removes the need to have a writeLock when flushing when subclassers give atomic work do to via internalPreFlushcacheCommit. This means that this patch breaks external contribs. This is required to keep our whole locking mechanism simpler.\n\n\nThis addresses bug HBASE-2915.\n http://issues.apache.org/jira/browse/HBASE-2915\n\n\nDiffs (updated)\n-----\n\n /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java 987355 \n /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java 987355 \n /trunk/src/test/java/org/apache/hadoop/hbase/regionserver/TestSplitTransaction.java 987355 \n\nDiff: http://review.cloudera.org/r/691/diff\n\n\nTesting\n-------\n\n5 concurrent ICV threads + randomWrite 3 + scans on a single RS. I'm also in the process of deploying it on a cluster.\n\n\nThanks,\n\nJean-Daniel\n\n\n","from":"developer"},{"body":"Message from: stack@duboce.net\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review988\n-----------------------------------------------------------\n\nShip it!\n\n\n+1\n\nif below are issues address on commit.\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Any reason this does not follow the pattern used elsewhere? You are not testing closing before getting the read lock?\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n ditto\n\n\n\n/trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java\n\n\n Is this supposed to be inside the try?\n\n\n- stack\n\n\n\n","from":"developer"},{"body":"Message from: \"Jean-Daniel Cryans\" \n\n\nbq. On 2010-08-20 20:49:15, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 716\nbq. > \nbq. >\nbq. > Any reason this does not follow the pattern used elsewhere? You are not testing closing before getting the read lock?\n\nmmm didn't want to copy the same chunk of code over there, but yeah it will also give us a fail-fast behavior which will speed up flushing.\n\n\nbq. On 2010-08-20 20:49:15, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 785\nbq. > \nbq. >\nbq. > ditto\n\nditto\n\n\nbq. On 2010-08-20 20:49:15, stack wrote:\nbq. > /trunk/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java, line 2960\nbq. > \nbq. >\nbq. > Is this supposed to be inside the try?\n\nPer that method's javadoc, no. The reason is that if it's all in a try block, and that closing is true, then you won't hold a readLock and you can't do a isHeldByCurrent thread on that lock. So I thought we could catch IllegalStateException in closeRegionOperation, but that would be really ugly. This leaves us to the current situation where if an exception is thrown when we are on this line:\n if (this.closed.get()) {\nthat it would leave the readLock locked by the thread, although in this case the region is closing so we're getting rid of it anyways. Although, thinking about it, I should probably do this instead:\n\n if (this.closed.get()) {\n lock.readLock().unlock();\n throw new NotServingRegionException(regionInfo.getRegionNameAsString() +\n \" is closed\");\n }\n\n\n- Jean-Daniel\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/691/#review988\n-----------------------------------------------------------\n\n\n\n","from":"developer"},{"body":"This patch is the same as the last one I posted on rb, with the addition of the unlocking in case the region is closed but the thread was able to get a readLock. I tested it on a 13 nodes cluster with 200 YCSB threads hitting it in different ways, didn't see any deadlock. Going to commit.","from":"developer"},{"body":"Committed to trunk, thanks for the review Stack!","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2010-08-13T19:20:04.000+0000","description":"HRegion.ICV gets a row lock then gets a newScanner lock.\n\nHRegion.close gets a newScanner lock, slitCloseLock and finally waits for all row locks to finish.\n\nIf the ICV got the row lock and then close got the newScannerLock, both end up waiting on the other. This was introduced when Get became a Scan.\n\nStack thinks we can get rid of the newScannerLock in close since we setClosing to true.","issue_id":"12471578","key":"HBASE-2915","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2010-08-23T22:49:59.000+0000","role":"fixed_distractor","summary":"Deadlock between HRegion.ICV and HRegion.close"} {"case_id":"12471721","cluster":"DISTRACTOR-HBASE-2920","comments":[{"body":"Simple patch that checks if the value is null before checking emptiness, along with unit test that fails without this patch.","created":"2010-08-25T16:51:57.335+0000"},{"body":"+1","created":"2010-08-25T18:00:03.988+0000"},{"body":"Committed to trunk, thanks for looking at it Stack!","created":"2010-08-25T19:00:09.139+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T12:42:40.763+0000"}],"conversations":[{"body":"From John Beatty on the ML:\n\n{quote}\nThanks Ryan, but I seem to be missing something then. It NPEs for me.\nWhen running against 0.89.20100726 and providing a null expected value\nI get the below stack trace (and works like a champ when I provide a\nbyte[0]. I also don't see the transformation you're referring to in\nHTable.\n\n(for reference,\nhttp://svn.apache.org/viewvc/hbase/branches/0.89.20100726/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java?view=markup)\n\njava.io.IOException: java.io.IOException: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HRegionServer.convertThrowableToIOE(HRegionServer.java:845)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.convertThrowableToIOE(HRegionServer.java:835)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.checkAndMutate(HRegionServer.java:1754)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.checkAndPut(HRegionServer.java:1773)\n at sun.reflect.GeneratedMethodAccessor8.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:576)\n at org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:919)\nCaused by: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HRegion.checkAndMutate(HRegion.java:1616)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.checkAndMutate(HRegionServer.java:1751)\n ... 6 more\n{quote}\n\nLooking in the code, I'm not sure either where the null conversion is done, even worse is that we don't even have unit tests! It should be put intoTestFromClientSide.","from":"reporter","subject":"HTable.checkAndPut/Delete doesn't handle null values"},{"body":"Simple patch that checks if the value is null before checking emptiness, along with unit test that fails without this patch.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed to trunk, thanks for looking at it Stack!","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2010-08-16T17:27:00.000+0000","description":"From John Beatty on the ML:\n\n{quote}\nThanks Ryan, but I seem to be missing something then. It NPEs for me.\nWhen running against 0.89.20100726 and providing a null expected value\nI get the below stack trace (and works like a champ when I provide a\nbyte[0]. I also don't see the transformation you're referring to in\nHTable.\n\n(for reference,\nhttp://svn.apache.org/viewvc/hbase/branches/0.89.20100726/src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java?view=markup)\n\njava.io.IOException: java.io.IOException: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HRegionServer.convertThrowableToIOE(HRegionServer.java:845)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.convertThrowableToIOE(HRegionServer.java:835)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.checkAndMutate(HRegionServer.java:1754)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.checkAndPut(HRegionServer.java:1773)\n at sun.reflect.GeneratedMethodAccessor8.invoke(Unknown Source)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:576)\n at org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:919)\nCaused by: java.lang.NullPointerException\n at org.apache.hadoop.hbase.regionserver.HRegion.checkAndMutate(HRegion.java:1616)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.checkAndMutate(HRegionServer.java:1751)\n ... 6 more\n{quote}\n\nLooking in the code, I'm not sure either where the null conversion is done, even worse is that we don't even have unit tests! It should be put intoTestFromClientSide.","issue_id":"12471721","key":"HBASE-2920","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2010-08-25T19:00:09.000+0000","role":"fixed_distractor","summary":"HTable.checkAndPut/Delete doesn't handle null values"} {"case_id":"13614898","cluster":"DISTRACTOR-HBASE-29251","comments":[{"body":"FYI [~apurtell] ","created":"2025-04-10T21:21:33.511+0000"},{"body":"> Provide retries for the proc store persist failures\r\n\r\nIf we had added retries, would it have eventually succeeded in writing to hdfs? [~vjasani]\r\nI am trying to evaluate options or do we want both? Retry for some fixed number of times and then abort.","created":"2025-04-11T16:10:01.638+0000"},{"body":"{quote}do we want both?\r\n{quote}\r\nIn a way yes, it would be better to have both. Retries with exponential backoff and eventual abort. Although if we want to abort, I think it would be better to take it up as separate Jira because we won't be able to propagate failure from proc executors all the way to active master, we will need some sort of chore to do that.","created":"2025-04-11T16:17:23.907+0000"},{"body":"I think master should abort if procedure store fails to persist. This is not the case?","created":"2025-04-12T03:40:09.073+0000"},{"body":"It is not the case. However, the reason why we only attempt once to store the procedure state once is because we avoid making RPC call (given that the master local region is local, it is not hosted elsewhere), so we directly call Region#put(). This has no retry, we should wrap this with retries and exponential backoff first, at least that can be main focus of this Jira.","created":"2025-04-12T04:05:27.550+0000"},{"body":"No, you should not retry here, since retry does not help anything. The WAL system is designed to write entries in serial, so if a previous operation does not finish yet, latter operations will wait there forever.\r\n\r\nSo if a put operation fails because of WAL timeout, the only solution is to abort master and go with the failure recovery logic.","created":"2025-04-12T04:19:37.485+0000"},{"body":"In that case, don't we have problem with aborts too? Let's say datanodes are not healthy and WAL sync is not going to go through for a few minutes, we will keep aborting masters one by one, correct?","created":"2025-04-12T04:27:57.146+0000"},{"body":"Besides, if meta or namespace comes in RIT by any chance (due to SCP), new active master initialization might get delayed or stuck too, until we manually intervene.\r\n\r\nI am thinking how worse it can be if hdfs health gets temporarily compromised.","created":"2025-04-12T05:12:45.958+0000"},{"body":"So how do you plan to recover if HDFS in trouble? As I said above, retrying does not help here, the only valid way is to abort and go with the failure recovery logic...\r\n\r\nThe key thing here is to get HDFS back ASAP, or you may introduce another WAL system which does not rely on HDFS and has better availibility...","created":"2025-04-12T05:44:58.070+0000"},{"body":"Let's think from remote client perspective for a moment: the client initiated mutation encounters issue with WAL appends and it gets IOE (not DoNotRetryIOE IIRC, but please help me refresh my memory if I am wrong), wouldn't client go through RPC retry mechanism and re-initiate the mutation?\r\n{quote}The key thing here is to get HDFS back ASAP\r\n{quote}\r\nThis incident had no significant issue in hdfs, only temporary one. But I agree that hdfs health is significantly important to Proc V2 framework.","created":"2025-04-12T17:19:51.925+0000"},{"body":"You are free to retry but it does not help. And at region server side, if we get a WALSyncTimeoutException, we will abort the region server.\r\n\r\n{code}\r\n /**\r\n * If {@link WAL#sync} get a timeout exception, the only correct way is to abort the region\r\n * server, as the design of {@link WAL#sync}, is to succeed or die, there is no 'failure'. It\r\n * is usually not a big deal is because we set a very large default value(5 minutes) for\r\n * {@link AbstractFSWAL#WAL_SYNC_TIMEOUT_MS}, usually the WAL system will abort the region\r\n * server if it can not finish the sync within 5 minutes.\r\n */\r\n if (ioe instanceof WALSyncTimeoutIOException) {\r\n if (rsServices != null) {\r\n rsServices.abort(\"WAL sync timeout,forcing server shutdown\", ioe);\r\n }\r\n }\r\n throw ioe;\r\n{code}","created":"2025-04-13T06:06:43.874+0000"},{"body":"I was not much worried about WALSyncTimeoutIOException because that also kills master:\r\n{code:java}\r\npublic void update(UpdateMasterRegion action) throws IOException {\r\n try {\r\n action.update(region);\r\n flusherAndCompactor.onUpdate();\r\n } catch (WALSyncTimeoutIOException e) {\r\n LOG.error(HBaseMarkers.FATAL, \"WAL sync timeout. Aborting server.\");\r\n server.abort(\"WAL sync timeout\", e);\r\n throw e;\r\n }\r\n} {code}\r\nAll I am trying to determine is whether something like this makes sense? we retry minimal times e.g. 3 times with a few seconds delay and then abort the master, rather than just aborting the master with just one attempt. Someone might be just restarting datanodes and the datanode pipeline might become healthy in few seconds or a minute or two.\r\n\r\nEven though writes are sequential, giving up with just single attempt seems bit extreme, doesn't it? (unless of course if it is WALSyncTimeoutIOException)\r\n\r\nI understand that the retry might not help much for persistent failure case, however we are only trying to add a bit more wait time for the WAL writes. If the update goes through with smaller num of attempts, even the earlier WALEdit would have been successful, given the sequential nature.","created":"2025-04-13T23:21:23.396+0000"},{"body":"The exception mentioned in the description is exactly the same with WALSyncTimeoutException...\r\n\r\nThis is a potential bug in FSHLog implementation in 2.x, where we will throw the exception out, in AsyncFSWAL and FSHLog in 3.x, we will not throw this exception out but retry forever in the WAL system.\r\n\r\nIn general, for the master local region, once we hit exception, the only way is to abort, I can not recall for which exception we could retry here...","created":"2025-04-14T06:12:31.700+0000"},{"body":"The exception mentioned in the description comes from hdfs DataStreamer directly as it cannot get ack from pipelined datanodes. In a way, yes it can be said to be similar to that of WALSyncTimeoutIOException, although both have different origins.\r\n{quote}This is a potential bug in FSHLog implementation in 2.x, where we will throw the exception out, in AsyncFSWAL and FSHLog in 3.x, we will not throw this exception out but retry forever in the WAL system.\r\n{quote}\r\nI see, are you talking about HBASE-27231?\r\n\r\nIt seems that aborting the master right away might not be bad idea, we need to make progress in persisting the proc state in some way and the lower num of retries could help only few times. The goal is to not get stuck here. The only catch is, sometimes we might run into the abort loops so hdfs health recovery in short duration is important.\r\n\r\nLet me give more thoughts and create draft PR in a few days unless the consensus changes.","created":"2025-04-14T06:30:29.835+0000"},{"body":"Yes, in HBASE-27231 we changed the behavior for FSHLog too.","created":"2025-04-14T06:50:41.364+0000"},{"body":"I recently encountered a similar situation where the master encountered an IOException while writing to MasterRegion , causing the ServerCrashProcedure to get stuck. The reason for the IOException was that when MasterRegion executed a {{put}} to write to HDFS 3 replica, a bad disk on another machine was selected.\r\n\r\nI think if we add retries to the {{put}} operation, MasterRegion might be able to successfully write after rolling the WAL and continue the ServerCrashProcedure. Although this does not guarantee a solution, it seems like a faster way to recover, as it is faster than restarting the Master.\r\n\r\n\r\nYour discussion helped me a lot.","created":"2025-04-14T12:35:42.445+0000"},{"body":"I encountered an internal incident which related to MasterRegion as well.\r\n\r\nI'm feeling some of the exceptions could be retried, in our case, `RegionTooBusyException`, retry and sleep with backoff indeed helps (waiting for flush finished). “Retries with exponential backoff and eventual abort” It is worthy.\r\n\r\nWhile some are not, Master needs to be aborted.\r\n\r\nI could contribute my internal fix","created":"2025-04-18T04:24:44.249+0000"},{"body":"For information, why could a Master encounter `RegionTooBusyException`.\r\n\r\nIt is because, my production is runnning on K8s. K8s env had a issue which led to all pods under a rsgroup crashed, therefore no actual RS (could only assign null or localhost,1,1) could be assigned regions. Then TRSP was looping persisting state of `GET_CANDIDATE` and `OPEN`, very massively, leading to `RegionTooBusyException`.\r\n\r\n","created":"2025-04-18T04:32:11.143+0000"},{"body":"Good to hear from you too [~reidchan] :)\r\n\r\nI could add list of retriable errors and apply few retries too, whereas for generic IOE, abort the master. PR: [https://github.com/apache/hbase/pull/6910]\r\n\r\nThough your above scenario looks like the one which would have led to retries exhaustion anyways, right? However, any genuine cause of RegionTooBusyException could benefit from few retries with exponential backoff, followed by eventual abort. WDYT [~zhangduo]?","created":"2025-04-18T04:40:47.890+0000"},{"body":"Since the mutation is not remote call in this case, we don't need to worry about retriable errors like ServerNotRunningYetException, PleaseHoldException, CallQueueTooBigException, RegionOpeningException, RegionMovedException, NotServingRegionException etc.\r\n\r\nIt seems RegionTooBusyException would be the only eligible one?","created":"2025-04-18T04:56:05.633+0000"},{"body":"{quote} which would have led to retries exhaustion anyways, right? {quote} \r\n\r\nYes, but this leads to another issue which might not related to this ticket, I modified some persistency logic in TRSP to avoid massive flooding to MasterRegion.\r\n\r\n \r\n{quote} It seems RegionTooBusyException would be the only eligible one?{quote}\r\nYes, I only met this one so far","created":"2025-04-18T06:10:53.610+0000"},{"body":" \r\n{quote}bq. which would have led to retries exhaustion anyways, right?\r\n{quote}\r\n{quote}Yes, but this leads to another issue which might not related to this ticket, I modified some persistency logic in TRSP to avoid massive flooding to MasterRegion.\r\n{quote}\r\nThis is great! If not done already, could you please commit these changes? btw we also use rsgroup and we have also seen bogus server issue when for some reasons, all servers of one small rsgroup goes down.\r\n\r\n \r\n\r\n[~reidchan] [~zhangduo] [~apurtell] [~prathyu6] I have added retriable error handling as well, please review: [https://github.com/apache/hbase/pull/6910]","created":"2025-04-18T22:23:19.111+0000"},{"body":"Merged the PR and backported to branch-3, branch-2, 2.6 and 2.5 (2.5 needed a bit more changes as the WAL timeout error leading to server abort is not present there).\r\n\r\nThanks everyone for the discussions/reviews on Jira and PR!","created":"2025-04-23T04:40:41.213+0000"}],"conversations":[{"body":"When a given regionserver stops or aborts, the corresponding ServerCrashProcedure is initiated by the active master. We have recently come across a case where initial state of the SCP SERVER_CRASH_START could not be persisted in the local region store:\r\n{code:java}\r\n2025-04-09 19:00:23,538 ERROR [RegionServerTracker-0] region.RegionProcedureStore - Failed to update proc pid=60020, state=RUNNABLE:SERVER_CRASH_START; ServerCrashProcedure server1,60020,1731526432248, splitWal=true, meta=false\r\njava.io.InterruptedIOException: No ack received after 55s and a timeout of 55s\r\n    at org.apache.hadoop.hdfs.DataStreamer.waitForAckedSeqno(DataStreamer.java:938)\r\n    at org.apache.hadoop.hdfs.DFSOutputStream.flushOrSync(DFSOutputStream.java:692)\r\n    at org.apache.hadoop.hdfs.DFSOutputStream.hflush(DFSOutputStream.java:580)\r\n    at org.apache.hadoop.fs.FSDataOutputStream.hflush(FSDataOutputStream.java:136)\r\n    at org.apache.hadoop.hbase.regionserver.wal.ProtobufLogWriter.sync(ProtobufLogWriter.java:85)\r\n    at org.apache.hadoop.hbase.regionserver.wal.FSHLog$SyncRunner.run(FSHLog.java:666) {code}\r\n \r\n\r\nThis led to no further action on the SCP, it stayed stuck until the active master was restarted manually.\r\n\r\nAfter the manual restart, new active master was able to proceed further with SCP:\r\n{code:java}\r\n2025-04-09 20:43:07,693 DEBUG [master/hmaster-3:60000:becomeActiveMaster] procedure2.ProcedureExecutor - Stored pid=60771, state=RUNNABLE:SERVER_CRASH_START; ServerCrashProcedure server1,60020,1731526432248, splitWal=true, meta=false\r\n\r\n2025-04-09 20:44:15,312 INFO  [PEWorker-18] procedure2.ProcedureExecutor - Finished pid=60771, state=SUCCESS; ServerCrashProcedure server1,60020,1731526432248, splitWal=true, meta=false in 1 mins, 7.667 sec {code}\r\n \r\n\r\nWhile it is well known that for active master to be operate without functional issues, the file system backing the master local region should be healthy. It is however worth noting that hdfs can have issues and master should be able to recover the procedures like SCP unless hdfs issues persist for longer duration.\r\n\r\nA couple of proposals:\r\n * Provide retries for the proc store persist failures\r\n * Abort active master for new master to continue the recovery (deployment systems usually ensure that the aborted servers are auto-started e.g. k8s or ambari)","from":"reporter","subject":"Procedure gets stuck if the procedure state cannot be persisted"},{"body":"FYI [~apurtell] ","from":"developer"},{"body":"> Provide retries for the proc store persist failures\r\n\r\nIf we had added retries, would it have eventually succeeded in writing to hdfs? [~vjasani]\r\nI am trying to evaluate options or do we want both? Retry for some fixed number of times and then abort.","from":"developer"},{"body":"{quote}do we want both?\r\n{quote}\r\nIn a way yes, it would be better to have both. Retries with exponential backoff and eventual abort. Although if we want to abort, I think it would be better to take it up as separate Jira because we won't be able to propagate failure from proc executors all the way to active master, we will need some sort of chore to do that.","from":"developer"},{"body":"I think master should abort if procedure store fails to persist. This is not the case?","from":"developer"},{"body":"It is not the case. However, the reason why we only attempt once to store the procedure state once is because we avoid making RPC call (given that the master local region is local, it is not hosted elsewhere), so we directly call Region#put(). This has no retry, we should wrap this with retries and exponential backoff first, at least that can be main focus of this Jira.","from":"developer"},{"body":"No, you should not retry here, since retry does not help anything. The WAL system is designed to write entries in serial, so if a previous operation does not finish yet, latter operations will wait there forever.\r\n\r\nSo if a put operation fails because of WAL timeout, the only solution is to abort master and go with the failure recovery logic.","from":"developer"},{"body":"In that case, don't we have problem with aborts too? Let's say datanodes are not healthy and WAL sync is not going to go through for a few minutes, we will keep aborting masters one by one, correct?","from":"developer"},{"body":"Besides, if meta or namespace comes in RIT by any chance (due to SCP), new active master initialization might get delayed or stuck too, until we manually intervene.\r\n\r\nI am thinking how worse it can be if hdfs health gets temporarily compromised.","from":"developer"},{"body":"So how do you plan to recover if HDFS in trouble? As I said above, retrying does not help here, the only valid way is to abort and go with the failure recovery logic...\r\n\r\nThe key thing here is to get HDFS back ASAP, or you may introduce another WAL system which does not rely on HDFS and has better availibility...","from":"developer"},{"body":"Let's think from remote client perspective for a moment: the client initiated mutation encounters issue with WAL appends and it gets IOE (not DoNotRetryIOE IIRC, but please help me refresh my memory if I am wrong), wouldn't client go through RPC retry mechanism and re-initiate the mutation?\r\n{quote}The key thing here is to get HDFS back ASAP\r\n{quote}\r\nThis incident had no significant issue in hdfs, only temporary one. But I agree that hdfs health is significantly important to Proc V2 framework.","from":"developer"},{"body":"You are free to retry but it does not help. And at region server side, if we get a WALSyncTimeoutException, we will abort the region server.\r\n\r\n{code}\r\n /**\r\n * If {@link WAL#sync} get a timeout exception, the only correct way is to abort the region\r\n * server, as the design of {@link WAL#sync}, is to succeed or die, there is no 'failure'. It\r\n * is usually not a big deal is because we set a very large default value(5 minutes) for\r\n * {@link AbstractFSWAL#WAL_SYNC_TIMEOUT_MS}, usually the WAL system will abort the region\r\n * server if it can not finish the sync within 5 minutes.\r\n */\r\n if (ioe instanceof WALSyncTimeoutIOException) {\r\n if (rsServices != null) {\r\n rsServices.abort(\"WAL sync timeout,forcing server shutdown\", ioe);\r\n }\r\n }\r\n throw ioe;\r\n{code}","from":"developer"},{"body":"I was not much worried about WALSyncTimeoutIOException because that also kills master:\r\n{code:java}\r\npublic void update(UpdateMasterRegion action) throws IOException {\r\n try {\r\n action.update(region);\r\n flusherAndCompactor.onUpdate();\r\n } catch (WALSyncTimeoutIOException e) {\r\n LOG.error(HBaseMarkers.FATAL, \"WAL sync timeout. Aborting server.\");\r\n server.abort(\"WAL sync timeout\", e);\r\n throw e;\r\n }\r\n} {code}\r\nAll I am trying to determine is whether something like this makes sense? we retry minimal times e.g. 3 times with a few seconds delay and then abort the master, rather than just aborting the master with just one attempt. Someone might be just restarting datanodes and the datanode pipeline might become healthy in few seconds or a minute or two.\r\n\r\nEven though writes are sequential, giving up with just single attempt seems bit extreme, doesn't it? (unless of course if it is WALSyncTimeoutIOException)\r\n\r\nI understand that the retry might not help much for persistent failure case, however we are only trying to add a bit more wait time for the WAL writes. If the update goes through with smaller num of attempts, even the earlier WALEdit would have been successful, given the sequential nature.","from":"developer"},{"body":"The exception mentioned in the description is exactly the same with WALSyncTimeoutException...\r\n\r\nThis is a potential bug in FSHLog implementation in 2.x, where we will throw the exception out, in AsyncFSWAL and FSHLog in 3.x, we will not throw this exception out but retry forever in the WAL system.\r\n\r\nIn general, for the master local region, once we hit exception, the only way is to abort, I can not recall for which exception we could retry here...","from":"developer"},{"body":"The exception mentioned in the description comes from hdfs DataStreamer directly as it cannot get ack from pipelined datanodes. In a way, yes it can be said to be similar to that of WALSyncTimeoutIOException, although both have different origins.\r\n{quote}This is a potential bug in FSHLog implementation in 2.x, where we will throw the exception out, in AsyncFSWAL and FSHLog in 3.x, we will not throw this exception out but retry forever in the WAL system.\r\n{quote}\r\nI see, are you talking about HBASE-27231?\r\n\r\nIt seems that aborting the master right away might not be bad idea, we need to make progress in persisting the proc state in some way and the lower num of retries could help only few times. The goal is to not get stuck here. The only catch is, sometimes we might run into the abort loops so hdfs health recovery in short duration is important.\r\n\r\nLet me give more thoughts and create draft PR in a few days unless the consensus changes.","from":"developer"},{"body":"Yes, in HBASE-27231 we changed the behavior for FSHLog too.","from":"developer"},{"body":"I recently encountered a similar situation where the master encountered an IOException while writing to MasterRegion , causing the ServerCrashProcedure to get stuck. The reason for the IOException was that when MasterRegion executed a {{put}} to write to HDFS 3 replica, a bad disk on another machine was selected.\r\n\r\nI think if we add retries to the {{put}} operation, MasterRegion might be able to successfully write after rolling the WAL and continue the ServerCrashProcedure. Although this does not guarantee a solution, it seems like a faster way to recover, as it is faster than restarting the Master.\r\n\r\n\r\nYour discussion helped me a lot.","from":"developer"},{"body":"I encountered an internal incident which related to MasterRegion as well.\r\n\r\nI'm feeling some of the exceptions could be retried, in our case, `RegionTooBusyException`, retry and sleep with backoff indeed helps (waiting for flush finished). “Retries with exponential backoff and eventual abort” It is worthy.\r\n\r\nWhile some are not, Master needs to be aborted.\r\n\r\nI could contribute my internal fix","from":"developer"},{"body":"For information, why could a Master encounter `RegionTooBusyException`.\r\n\r\nIt is because, my production is runnning on K8s. K8s env had a issue which led to all pods under a rsgroup crashed, therefore no actual RS (could only assign null or localhost,1,1) could be assigned regions. Then TRSP was looping persisting state of `GET_CANDIDATE` and `OPEN`, very massively, leading to `RegionTooBusyException`.\r\n\r\n","from":"developer"},{"body":"Good to hear from you too [~reidchan] :)\r\n\r\nI could add list of retriable errors and apply few retries too, whereas for generic IOE, abort the master. PR: [https://github.com/apache/hbase/pull/6910]\r\n\r\nThough your above scenario looks like the one which would have led to retries exhaustion anyways, right? However, any genuine cause of RegionTooBusyException could benefit from few retries with exponential backoff, followed by eventual abort. WDYT [~zhangduo]?","from":"developer"},{"body":"Since the mutation is not remote call in this case, we don't need to worry about retriable errors like ServerNotRunningYetException, PleaseHoldException, CallQueueTooBigException, RegionOpeningException, RegionMovedException, NotServingRegionException etc.\r\n\r\nIt seems RegionTooBusyException would be the only eligible one?","from":"developer"},{"body":"{quote} which would have led to retries exhaustion anyways, right? {quote} \r\n\r\nYes, but this leads to another issue which might not related to this ticket, I modified some persistency logic in TRSP to avoid massive flooding to MasterRegion.\r\n\r\n \r\n{quote} It seems RegionTooBusyException would be the only eligible one?{quote}\r\nYes, I only met this one so far","from":"developer"},{"body":" \r\n{quote}bq. which would have led to retries exhaustion anyways, right?\r\n{quote}\r\n{quote}Yes, but this leads to another issue which might not related to this ticket, I modified some persistency logic in TRSP to avoid massive flooding to MasterRegion.\r\n{quote}\r\nThis is great! If not done already, could you please commit these changes? btw we also use rsgroup and we have also seen bogus server issue when for some reasons, all servers of one small rsgroup goes down.\r\n\r\n \r\n\r\n[~reidchan] [~zhangduo] [~apurtell] [~prathyu6] I have added retriable error handling as well, please review: [https://github.com/apache/hbase/pull/6910]","from":"developer"},{"body":"Merged the PR and backported to branch-3, branch-2, 2.6 and 2.5 (2.5 needed a bit more changes as the WAL timeout error leading to server abort is not present there).\r\n\r\nThanks everyone for the discussions/reviews on Jira and PR!","from":"developer"}],"created":"2025-04-10T20:33:40.000+0000","description":"When a given regionserver stops or aborts, the corresponding ServerCrashProcedure is initiated by the active master. We have recently come across a case where initial state of the SCP SERVER_CRASH_START could not be persisted in the local region store:\r\n{code:java}\r\n2025-04-09 19:00:23,538 ERROR [RegionServerTracker-0] region.RegionProcedureStore - Failed to update proc pid=60020, state=RUNNABLE:SERVER_CRASH_START; ServerCrashProcedure server1,60020,1731526432248, splitWal=true, meta=false\r\njava.io.InterruptedIOException: No ack received after 55s and a timeout of 55s\r\n    at org.apache.hadoop.hdfs.DataStreamer.waitForAckedSeqno(DataStreamer.java:938)\r\n    at org.apache.hadoop.hdfs.DFSOutputStream.flushOrSync(DFSOutputStream.java:692)\r\n    at org.apache.hadoop.hdfs.DFSOutputStream.hflush(DFSOutputStream.java:580)\r\n    at org.apache.hadoop.fs.FSDataOutputStream.hflush(FSDataOutputStream.java:136)\r\n    at org.apache.hadoop.hbase.regionserver.wal.ProtobufLogWriter.sync(ProtobufLogWriter.java:85)\r\n    at org.apache.hadoop.hbase.regionserver.wal.FSHLog$SyncRunner.run(FSHLog.java:666) {code}\r\n \r\n\r\nThis led to no further action on the SCP, it stayed stuck until the active master was restarted manually.\r\n\r\nAfter the manual restart, new active master was able to proceed further with SCP:\r\n{code:java}\r\n2025-04-09 20:43:07,693 DEBUG [master/hmaster-3:60000:becomeActiveMaster] procedure2.ProcedureExecutor - Stored pid=60771, state=RUNNABLE:SERVER_CRASH_START; ServerCrashProcedure server1,60020,1731526432248, splitWal=true, meta=false\r\n\r\n2025-04-09 20:44:15,312 INFO  [PEWorker-18] procedure2.ProcedureExecutor - Finished pid=60771, state=SUCCESS; ServerCrashProcedure server1,60020,1731526432248, splitWal=true, meta=false in 1 mins, 7.667 sec {code}\r\n \r\n\r\nWhile it is well known that for active master to be operate without functional issues, the file system backing the master local region should be healthy. It is however worth noting that hdfs can have issues and master should be able to recover the procedures like SCP unless hdfs issues persist for longer duration.\r\n\r\nA couple of proposals:\r\n * Provide retries for the proc store persist failures\r\n * Abort active master for new master to continue the recovery (deployment systems usually ensure that the aborted servers are auto-started e.g. k8s or ambari)","issue_id":"13614898","key":"HBASE-29251","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2025-04-23T04:40:53.000+0000","role":"fixed_distractor","summary":"Procedure gets stuck if the procedure state cannot be persisted"} {"case_id":"12473415","cluster":"DISTRACTOR-HBASE-2964","comments":[{"body":"Fixing this is a little tricky. We could short-circuit the IPC path when detecting that a region is hosted in the same process, and thus avoid going through handlers (this is what the datanode does in the block recovery code). However, you still can have a situation where two regionservers are trying to talk to each other and end up in a deadlock.\n\nAnother option is to add a timeout to these RPCs, abort the split and try again later if it fails.\n\nAnother thing that might help is to have the start of the split transaction flag the table as \"going offline\", and before taking the readlock, other accessors of the table can check for this case and immediately throw NSRE rather than blocking once the split is in progress.","created":"2010-09-07T01:18:31.747+0000"},{"body":"HBASE-2782 is another solution - if we provide QoS for the META and ROOT tables so they always have a few reserved handlers in a separate thread pool, we can avoid this issue.","created":"2010-09-07T01:22:27.008+0000"},{"body":"I agree this a blocker on 0.90.x","created":"2010-09-07T04:20:10.320+0000"},{"body":"As noted on the list, this seems to be due to HBASE-2461.\n\nPrior to 2461, when we split, we would close the region before doing any of the writes to META, and didn't hold any locks while doing the META updates. Now we keep the write lock all the way through, even after closing the region.\n\nI think simply moving the writeLock().unlock() up after the this.parent.close(false) in SplitTransaction should fix this issue. I'm testing that change on my test cluster now.","created":"2010-09-07T05:14:19.591+0000"},{"body":"I also had to move the \"new HTable\" call outside of the lock, since the HTable constructor does an RPC.\n\nThis patch seems to fix the issue for me. Running an overnight load test - if it's still going in the morning I'd say we're good :)","created":"2010-09-07T07:15:12.096+0000"},{"body":"Overnight test completed OK with that patch. I think we should rebuild the rc with this if Stack thinks it looks good.","created":"2010-09-07T14:35:59.566+0000"},{"body":"Message from: \"Todd Lipcon\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/798/\n-----------------------------------------------------------\n\nReview request for hbase and stack.\n\n\nSummary\n-------\n\nMoves all RPCs outside of the region writeLock - the writeLock is now only used long enough to set the 'closing' flag. When we drop the lock any waiters will see 'closing' upon acquiring the lock, and thus throw NSRE.\n\nIn the case that we abort the split, it will reopen the region as before. Accessors will have gotten NSRE but will just come back to the same region eventually.\n\n\nThis addresses bug HBASE-2964.\n http://issues.apache.org/jira/browse/HBASE-2964\n\n\nDiffs\n-----\n\n src/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java 3507c0d \n\nDiff: http://review.cloudera.org/r/798/diff\n\n\nTesting\n-------\n\nYCSB testing on my cluster - it used to deadlock due to this bug within an hour. I ran a 5 hour load test overnight and it worked OK.\n\n\nThanks,\n\nTodd\n\n\n","created":"2010-09-07T18:00:59.011+0000"},{"body":"Message from: stack@duboce.net\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/798/#review1110\n-----------------------------------------------------------\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java\n\n\n Let me make a version of this patch that takes care of rollback -- currently rollback expects the lock to be held on entrance; this will not be the case post close if above applied.\n\n\n- stack\n\n\n\n","created":"2010-09-07T18:30:57.340+0000"},{"body":"Hmmm... now I'm thinking instead that we punt locking up here in splittransaction completely. The core issue comes of an incorrect mapping of old splitLock on to new region 'lock'. Looking at what was done under the old splitLock, it all looks safe in the face of concurrency. Down in the region close, its already taking out the region write lock. Let me make a different kinda patch.","created":"2010-09-07T19:31:43.562+0000"},{"body":"Message from: stack@duboce.net\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/798/\n-----------------------------------------------------------\n\n(Updated 2010-09-07 13:38:39.968517)\n\n\nReview request for hbase and stack.\n\n\nChanges\n-------\n\nThis version removes from SplitTransaction the setting of the this.parent.lock completely. Its not needed. Down in the parent close, it takes out the write lock.\n\nIn the past, we had a split lock and a close lock (splitLock and splitsAndClosesLock). The split lock was held across the split while daughter regions were calculated and during close, actual split and update of .META. As part of lock pruning, an error made in hbase-2641, was using splitsAndClosesLock where splitLock was used previously -- and even expanding the scope of what splitLock used cover).\n\nLooking, splitLock looks like it could have served some purpose preventing two threads contending over splitting (splits make objects in filesystem and move stuff around), but we don't really need this in current HBase since only CompactSplitThread runs splits -- even in new master regime where client can call a splitRegion. Later when we want to run multiple concurrent split transactions, we'll need to reexamine.\n\n\nSummary\n-------\n\nMoves all RPCs outside of the region writeLock - the writeLock is now only used long enough to set the 'closing' flag. When we drop the lock any waiters will see 'closing' upon acquiring the lock, and thus throw NSRE.\n\nIn the case that we abort the split, it will reopen the region as before. Accessors will have gotten NSRE but will just come back to the same region eventually.\n\n\nThis addresses bug HBASE-2964.\n http://issues.apache.org/jira/browse/HBASE-2964\n\n\nDiffs (updated)\n-----\n\n src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java a692125 \n src/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java 3507c0d \n src/test/java/org/apache/hadoop/hbase/regionserver/TestSplitTransaction.java a245d97 \n\nDiff: http://review.cloudera.org/r/798/diff\n\n\nTesting\n-------\n\nYCSB testing on my cluster - it used to deadlock due to this bug within an hour. I ran a 5 hour load test overnight and it worked OK.\n\n\nThanks,\n\nTodd\n\n\n","created":"2010-09-07T20:48:27.770+0000"},{"body":"Message from: \"Todd Lipcon\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/798/#review1122\n-----------------------------------------------------------\n\n\nSeems to make sense. Let me try it on a cluster before I +1 it\n\n\nsrc/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java\n\n\n maybe now we can do an:\n \n assert !this.parent.lock.writeLock().isHeldByCurrentThread() : \"Unsafe to hold write lock while performing RPCs\";\n\n\n- Todd\n\n\n\n","created":"2010-09-08T01:39:15.868+0000"},{"body":"+1 to stack's patch from reviewboard. Imported about 550G over night, worked OK.","created":"2010-09-08T16:32:35.958+0000"},{"body":"Message from: stack@duboce.net\n\n\nbq. On 2010-09-07 18:33:16, Todd Lipcon wrote:\nbq. > src/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java, line 207\nbq. > \nbq. >\nbq. > maybe now we can do an:\nbq. > \nbq. > assert !this.parent.lock.writeLock().isHeldByCurrentThread() : \"Unsafe to hold write lock while performing RPCs\";\n\nI'll add in this assert\n\n\n- stack\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/798/#review1122\n-----------------------------------------------------------\n\n\n\n","created":"2010-09-08T16:48:51.211+0000"},{"body":"Thanks for review and for testing Todd (applied to TRUNK and to 0.89.20100830 branch.","created":"2010-09-08T16:58:31.651+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T12:42:29.775+0000"}],"conversations":[{"body":"In testing the 0.89.20100830 rc, I ran into a deadlock with the following situation:\n\n- All of the IPC Handler threads are blocked on the region lock, which is held by CompactSplitThread.\n- CompactSplitThread is in the process of trying to edit META to create the offline parent. META happens to be on the same server as is executing the split.\n\nTherefore, the CompactSplitThread is trying to connect back to itself, but all of the handler threads are blocked, so the IPC never happens. Thus, the entire RS gets deadlocked.","from":"reporter","subject":"Deadlock when RS tries to RPC to itself inside SplitTransaction"},{"body":"Fixing this is a little tricky. We could short-circuit the IPC path when detecting that a region is hosted in the same process, and thus avoid going through handlers (this is what the datanode does in the block recovery code). However, you still can have a situation where two regionservers are trying to talk to each other and end up in a deadlock.\n\nAnother option is to add a timeout to these RPCs, abort the split and try again later if it fails.\n\nAnother thing that might help is to have the start of the split transaction flag the table as \"going offline\", and before taking the readlock, other accessors of the table can check for this case and immediately throw NSRE rather than blocking once the split is in progress.","from":"developer"},{"body":"HBASE-2782 is another solution - if we provide QoS for the META and ROOT tables so they always have a few reserved handlers in a separate thread pool, we can avoid this issue.","from":"developer"},{"body":"I agree this a blocker on 0.90.x","from":"developer"},{"body":"As noted on the list, this seems to be due to HBASE-2461.\n\nPrior to 2461, when we split, we would close the region before doing any of the writes to META, and didn't hold any locks while doing the META updates. Now we keep the write lock all the way through, even after closing the region.\n\nI think simply moving the writeLock().unlock() up after the this.parent.close(false) in SplitTransaction should fix this issue. I'm testing that change on my test cluster now.","from":"developer"},{"body":"I also had to move the \"new HTable\" call outside of the lock, since the HTable constructor does an RPC.\n\nThis patch seems to fix the issue for me. Running an overnight load test - if it's still going in the morning I'd say we're good :)","from":"developer"},{"body":"Overnight test completed OK with that patch. I think we should rebuild the rc with this if Stack thinks it looks good.","from":"developer"},{"body":"Message from: \"Todd Lipcon\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/798/\n-----------------------------------------------------------\n\nReview request for hbase and stack.\n\n\nSummary\n-------\n\nMoves all RPCs outside of the region writeLock - the writeLock is now only used long enough to set the 'closing' flag. When we drop the lock any waiters will see 'closing' upon acquiring the lock, and thus throw NSRE.\n\nIn the case that we abort the split, it will reopen the region as before. Accessors will have gotten NSRE but will just come back to the same region eventually.\n\n\nThis addresses bug HBASE-2964.\n http://issues.apache.org/jira/browse/HBASE-2964\n\n\nDiffs\n-----\n\n src/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java 3507c0d \n\nDiff: http://review.cloudera.org/r/798/diff\n\n\nTesting\n-------\n\nYCSB testing on my cluster - it used to deadlock due to this bug within an hour. I ran a 5 hour load test overnight and it worked OK.\n\n\nThanks,\n\nTodd\n\n\n","from":"developer"},{"body":"Message from: stack@duboce.net\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/798/#review1110\n-----------------------------------------------------------\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java\n\n\n Let me make a version of this patch that takes care of rollback -- currently rollback expects the lock to be held on entrance; this will not be the case post close if above applied.\n\n\n- stack\n\n\n\n","from":"developer"},{"body":"Hmmm... now I'm thinking instead that we punt locking up here in splittransaction completely. The core issue comes of an incorrect mapping of old splitLock on to new region 'lock'. Looking at what was done under the old splitLock, it all looks safe in the face of concurrency. Down in the region close, its already taking out the region write lock. Let me make a different kinda patch.","from":"developer"},{"body":"Message from: stack@duboce.net\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/798/\n-----------------------------------------------------------\n\n(Updated 2010-09-07 13:38:39.968517)\n\n\nReview request for hbase and stack.\n\n\nChanges\n-------\n\nThis version removes from SplitTransaction the setting of the this.parent.lock completely. Its not needed. Down in the parent close, it takes out the write lock.\n\nIn the past, we had a split lock and a close lock (splitLock and splitsAndClosesLock). The split lock was held across the split while daughter regions were calculated and during close, actual split and update of .META. As part of lock pruning, an error made in hbase-2641, was using splitsAndClosesLock where splitLock was used previously -- and even expanding the scope of what splitLock used cover).\n\nLooking, splitLock looks like it could have served some purpose preventing two threads contending over splitting (splits make objects in filesystem and move stuff around), but we don't really need this in current HBase since only CompactSplitThread runs splits -- even in new master regime where client can call a splitRegion. Later when we want to run multiple concurrent split transactions, we'll need to reexamine.\n\n\nSummary\n-------\n\nMoves all RPCs outside of the region writeLock - the writeLock is now only used long enough to set the 'closing' flag. When we drop the lock any waiters will see 'closing' upon acquiring the lock, and thus throw NSRE.\n\nIn the case that we abort the split, it will reopen the region as before. Accessors will have gotten NSRE but will just come back to the same region eventually.\n\n\nThis addresses bug HBASE-2964.\n http://issues.apache.org/jira/browse/HBASE-2964\n\n\nDiffs (updated)\n-----\n\n src/main/java/org/apache/hadoop/hbase/regionserver/HRegion.java a692125 \n src/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java 3507c0d \n src/test/java/org/apache/hadoop/hbase/regionserver/TestSplitTransaction.java a245d97 \n\nDiff: http://review.cloudera.org/r/798/diff\n\n\nTesting\n-------\n\nYCSB testing on my cluster - it used to deadlock due to this bug within an hour. I ran a 5 hour load test overnight and it worked OK.\n\n\nThanks,\n\nTodd\n\n\n","from":"developer"},{"body":"Message from: \"Todd Lipcon\" \n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/798/#review1122\n-----------------------------------------------------------\n\n\nSeems to make sense. Let me try it on a cluster before I +1 it\n\n\nsrc/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java\n\n\n maybe now we can do an:\n \n assert !this.parent.lock.writeLock().isHeldByCurrentThread() : \"Unsafe to hold write lock while performing RPCs\";\n\n\n- Todd\n\n\n\n","from":"developer"},{"body":"+1 to stack's patch from reviewboard. Imported about 550G over night, worked OK.","from":"developer"},{"body":"Message from: stack@duboce.net\n\n\nbq. On 2010-09-07 18:33:16, Todd Lipcon wrote:\nbq. > src/main/java/org/apache/hadoop/hbase/regionserver/SplitTransaction.java, line 207\nbq. > \nbq. >\nbq. > maybe now we can do an:\nbq. > \nbq. > assert !this.parent.lock.writeLock().isHeldByCurrentThread() : \"Unsafe to hold write lock while performing RPCs\";\n\nI'll add in this assert\n\n\n- stack\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttp://review.cloudera.org/r/798/#review1122\n-----------------------------------------------------------\n\n\n\n","from":"developer"},{"body":"Thanks for review and for testing Todd (applied to TRUNK and to 0.89.20100830 branch.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2010-09-07T01:16:02.000+0000","description":"In testing the 0.89.20100830 rc, I ran into a deadlock with the following situation:\n\n- All of the IPC Handler threads are blocked on the region lock, which is held by CompactSplitThread.\n- CompactSplitThread is in the process of trying to edit META to create the offline parent. META happens to be on the same server as is executing the split.\n\nTherefore, the CompactSplitThread is trying to connect back to itself, but all of the handler threads are blocked, so the IPC never happens. Thus, the entire RS gets deadlocked.","issue_id":"12473415","key":"HBASE-2964","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2010-09-08T16:58:31.000+0000","role":"fixed_distractor","summary":"Deadlock when RS tries to RPC to itself inside SplitTransaction"} {"case_id":"13633001","cluster":"DISTRACTOR-HBASE-29696","comments":[{"body":"I'd like to work on this issue. I plan to use the synthetic boundaries only for calculating the intermediate split keys, then restore the original start and end rows on the first and last generated splits. This keeps the final split open-ended and prevents binary row keys such as 0xFFFF from being skipped. I will also add regression tests for both empty-to-empty and non-empty-to-empty ranges.","created":"2026-07-22T08:36:14.560+0000"},{"body":"I have prepared a fix for this issue.\r\n\r\n`createNInputSplitsUniform` still uses synthetic finite boundaries while calling `Bytes.split`, but it now restores the original `TableSplit` start and end rows before constructing the generated splits.\r\n\r\nTherefore, when the original end row is empty, the last split remains open-ended instead of using the temporary all-`0xFF` key. This prevents binary row keys such as `0xFFFF` from being excluded from the scan.\r\n\r\nI also added a regression test covering both empty-to-empty and non-empty-to-empty ranges. The test verifies preservation of the original boundaries and continuity between adjacent splits.","created":"2026-07-22T12:52:30.192+0000"},{"body":"Thank you [~mazhengxuan] for the contribution!\r\n\r\nPushed to the following branches:\r\n\r\n* master\r\n* branch-3\r\n* branch-3.0\r\n* branch-2\r\n* branch-2.6\r\n* branch-2.5\r\n\r\nThanks to [~noslowerdna] for the review!","created":"2026-07-30T06:06:00.741+0000"}],"conversations":[{"body":"*How to reproduce*\r\n\r\nCreate a table with a single region and add the row key \r\n{code:java}\r\nnew byte[] {-1, -1}{code}\r\nPerform a MapReduce job with at least this setting:\r\n{code:java}\r\n\"hbase.mapreduce.tableinput.mappers.per.region\" = 2 {code}\r\nThe row key is missing in the scan result.\r\n\r\n \r\n\r\n*Analysis*\r\nIn TableInputF{color:#172b4d}ormatBase#createNInputSplitsUniform there is this code:\r\n{color}\r\n{code:java}\r\n// For special case: startRow or endRow is empty\r\nif (startRow.length == 0 && endRow.length == 0) {\r\n startRow = new byte[1];\r\n endRow = new byte[1];\r\n startRow[0] = 0;\r\n endRow[0] = -1;\r\n}\r\nif (startRow.length == 0 && endRow.length != 0) {\r\n startRow = new byte[1];\r\n startRow[0] = 0;\r\n}\r\nif (startRow.length != 0 && endRow.length == 0) {\r\n endRow = new byte[startRow.length];\r\n for (int k = 0; k < startRow.length; k++) {\r\n endRow[k] = -1;\r\n }\r\n} {code}\r\nUnfortunately, in the first 'if', the endRow is set to \r\n{code:java}\r\nnew byte[] {-1} {code}\r\nBut what if there is a row key with\r\n{code:java}\r\nnew byte[] {-1, -1}{code}\r\nThis row key is after the endRow and will be ignored by the scan. \r\nThis is also an issue in the third 'if'. Since a row key can be of potentially of an infinite length, setting the end row in the third 'if' also prevents longer row keys (compared to the start row) to be ignored. \r\n\r\nTherefore, in both situations, the end row should stay empty. I can imagine the endRow is set for the next step:\r\n{code:java}\r\n// Split Region into n chunks evenly\r\nbyte[][] splitKeys = Bytes.split(startRow, endRow, true, n - 1); {code}\r\nIn that case, some compensation should be done later to return the correct input splits. ","from":"reporter","subject":"TableInputFormatBase with NUM_MAPPERS_PER_REGION produces incorrect last InputSplit"},{"body":"I'd like to work on this issue. I plan to use the synthetic boundaries only for calculating the intermediate split keys, then restore the original start and end rows on the first and last generated splits. This keeps the final split open-ended and prevents binary row keys such as 0xFFFF from being skipped. I will also add regression tests for both empty-to-empty and non-empty-to-empty ranges.","from":"developer"},{"body":"I have prepared a fix for this issue.\r\n\r\n`createNInputSplitsUniform` still uses synthetic finite boundaries while calling `Bytes.split`, but it now restores the original `TableSplit` start and end rows before constructing the generated splits.\r\n\r\nTherefore, when the original end row is empty, the last split remains open-ended instead of using the temporary all-`0xFF` key. This prevents binary row keys such as `0xFFFF` from being excluded from the scan.\r\n\r\nI also added a regression test covering both empty-to-empty and non-empty-to-empty ranges. The test verifies preservation of the original boundaries and continuity between adjacent splits.","from":"developer"},{"body":"Thank you [~mazhengxuan] for the contribution!\r\n\r\nPushed to the following branches:\r\n\r\n* master\r\n* branch-3\r\n* branch-3.0\r\n* branch-2\r\n* branch-2.6\r\n* branch-2.5\r\n\r\nThanks to [~noslowerdna] for the review!","from":"developer"}],"created":"2025-10-31T12:35:54.000+0000","description":"*How to reproduce*\r\n\r\nCreate a table with a single region and add the row key \r\n{code:java}\r\nnew byte[] {-1, -1}{code}\r\nPerform a MapReduce job with at least this setting:\r\n{code:java}\r\n\"hbase.mapreduce.tableinput.mappers.per.region\" = 2 {code}\r\nThe row key is missing in the scan result.\r\n\r\n \r\n\r\n*Analysis*\r\nIn TableInputF{color:#172b4d}ormatBase#createNInputSplitsUniform there is this code:\r\n{color}\r\n{code:java}\r\n// For special case: startRow or endRow is empty\r\nif (startRow.length == 0 && endRow.length == 0) {\r\n startRow = new byte[1];\r\n endRow = new byte[1];\r\n startRow[0] = 0;\r\n endRow[0] = -1;\r\n}\r\nif (startRow.length == 0 && endRow.length != 0) {\r\n startRow = new byte[1];\r\n startRow[0] = 0;\r\n}\r\nif (startRow.length != 0 && endRow.length == 0) {\r\n endRow = new byte[startRow.length];\r\n for (int k = 0; k < startRow.length; k++) {\r\n endRow[k] = -1;\r\n }\r\n} {code}\r\nUnfortunately, in the first 'if', the endRow is set to \r\n{code:java}\r\nnew byte[] {-1} {code}\r\nBut what if there is a row key with\r\n{code:java}\r\nnew byte[] {-1, -1}{code}\r\nThis row key is after the endRow and will be ignored by the scan. \r\nThis is also an issue in the third 'if'. Since a row key can be of potentially of an infinite length, setting the end row in the third 'if' also prevents longer row keys (compared to the start row) to be ignored. \r\n\r\nTherefore, in both situations, the end row should stay empty. I can imagine the endRow is set for the next step:\r\n{code:java}\r\n// Split Region into n chunks evenly\r\nbyte[][] splitKeys = Bytes.split(startRow, endRow, true, n - 1); {code}\r\nIn that case, some compensation should be done later to return the correct input splits. ","issue_id":"13633001","key":"HBASE-29696","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2026-07-30T06:06:24.000+0000","role":"fixed_distractor","summary":"TableInputFormatBase with NUM_MAPPERS_PER_REGION produces incorrect last InputSplit"} {"case_id":"13655252","cluster":"DISTRACTOR-HBASE-30248","comments":[{"body":"I do not have assign permission in JIRA yet. Could someone assign this issue to me?","created":"2026-07-08T06:31:01.465+0000"},{"body":"I tried to assign this issue to mazhengxuan, but it looks like the user is not assignable in the HBASE project yet.  Could someone with JIRA project permission please assign HBASE-30248 to mazhengxuan, or add them as an assignable user?  In the meantime, I am fine with mazhengxuan working on this issue, and I can help verify the patch with the original reproducer once a PR is available.","created":"2026-07-08T07:01:57.049+0000"},{"body":"{{I opened [https://github.com/apache/hbase/pull/8468] for this issue.}}\r\n\r\nThe patch validates hbase.snapshot.region.timeout at all RegionServer snapshot call sites, falls back to the documented default for values less than or equal to zero, and adds unit coverage.\r\n\r\n{{Hello Jiang He, could you please verify the patch with the original reproducer?}}","created":"2026-07-12T14:47:48.124+0000"},{"body":"Hello mazhengxuan, I verified the patch with the original reproducer.\r\n\r\nWith hbase.snapshot.region.timeout=-300000, the RegionServer no longer aborts during snapshot procedure initialization. The log shows that the value falls back to the default:\r\n\r\nhbase.snapshot.region.timeout must be greater than 0, but is -300000. Using default value 300000 The online-snapshot procedure was initialized and started successfully, and the master completed initialization.\r\n\r\nI also verified hbase.snapshot.region.timeout=0. It also falls back to the default value and HBase starts successfully.","created":"2026-07-13T02:53:19.087+0000"},{"body":"Thank you, Jiang He, for verifying the patch with the original reproducer and testing both negative and zero values. I'm glad to hear that the RegionServer no longer aborts and HBase starts successfully.\r\n\r\nThis confirms that the patch addresses the reported issue. I’ll continue following up on the PR review and CI process. Thanks again for your help!","created":"2026-07-13T07:20:45.829+0000"},{"body":"Pushed to all active branches.\r\n\r\nThanks [~mazhengxuan] [~velpro] ","created":"2026-07-27T15:20:41.248+0000"},{"body":"Cherry-picked to branch-3.0 too.","created":"2026-07-29T09:26:44.378+0000"}],"conversations":[{"body":" When `hbase.snapshot.region.timeout` is configured as a negative value, the RegionServer aborts during initialization.\r\n\r\n  The negative value is passed as the keep-alive time to a `ThreadPoolExecutor`, which throws `IllegalArgumentException`.\r\n\r\n  ## Reproducer\r\n\r\n  Use the following `hbase-site.xml` entries:\r\n\r\n  ```xml\r\n  \r\n    \r\n      hbase.rootdir\r\n      /tmp/hbase-tmp\r\n    \r\n    \r\n      hbase.cluster.distributed\r\n      false\r\n    \r\n    \r\n      hbase.unsafe.stream.capability.enforce\r\n      false\r\n    \r\n    \r\n      hbase.snapshot.region.timeout\r\n      -300000\r\n    \r\n  \r\n\r\n  Then start HBase.\r\n\r\n  Actual result\r\n\r\n  The RegionServer aborts:\r\n\r\n  ERROR [RS:0;...] regionserver.HRegionServer:\r\n  ***** ABORTING region server ... Initialization of RS failed. Hence aborting RS. *****\r\n  java.lang.IllegalArgumentException\r\n      at java.base/java.util.concurrent.ThreadPoolExecutor.(ThreadPoolExecutor.java:1293)\r\n      at java.base/java.util.concurrent.ThreadPoolExecutor.(ThreadPoolExecutor.java:1215)\r\n      at org.apache.hadoop.hbase.procedure.ProcedureMember.defaultPool(ProcedureMember.java:86)\r\n      at org.apache.hadoop.hbase.regionserver.snapshot.RegionServerSnapshotManager.initialize(RegionServerSnapshotManager.j\r\n  ava:391)\r\n\r\n  Root cause\r\n\r\n  hbase.snapshot.region.timeout is read in RegionServerSnapshotManager.initialize and used as the keep-alive time when\r\n  creating the snapshot procedure member thread pool.\r\n\r\n  A negative value is invalid for ThreadPoolExecutor.\r\n\r\n  Expected result\r\n\r\n  HBase should validate hbase.snapshot.region.timeout and reject negative values with a clear configuration error, or fall\r\n  back to a safe default.\r\n\r\n  The RegionServer should not abort with a raw IllegalArgumentException.\r\n\r\n  Notes\r\n\r\n  The failure was found by configuration fuzzing.","from":"reporter","subject":"Negative hbase.snapshot.region.timeout causes RegionServer abort during snapshot procedure initialization"},{"body":"I do not have assign permission in JIRA yet. Could someone assign this issue to me?","from":"developer"},{"body":"I tried to assign this issue to mazhengxuan, but it looks like the user is not assignable in the HBASE project yet.  Could someone with JIRA project permission please assign HBASE-30248 to mazhengxuan, or add them as an assignable user?  In the meantime, I am fine with mazhengxuan working on this issue, and I can help verify the patch with the original reproducer once a PR is available.","from":"developer"},{"body":"{{I opened [https://github.com/apache/hbase/pull/8468] for this issue.}}\r\n\r\nThe patch validates hbase.snapshot.region.timeout at all RegionServer snapshot call sites, falls back to the documented default for values less than or equal to zero, and adds unit coverage.\r\n\r\n{{Hello Jiang He, could you please verify the patch with the original reproducer?}}","from":"developer"},{"body":"Hello mazhengxuan, I verified the patch with the original reproducer.\r\n\r\nWith hbase.snapshot.region.timeout=-300000, the RegionServer no longer aborts during snapshot procedure initialization. The log shows that the value falls back to the default:\r\n\r\nhbase.snapshot.region.timeout must be greater than 0, but is -300000. Using default value 300000 The online-snapshot procedure was initialized and started successfully, and the master completed initialization.\r\n\r\nI also verified hbase.snapshot.region.timeout=0. It also falls back to the default value and HBase starts successfully.","from":"developer"},{"body":"Thank you, Jiang He, for verifying the patch with the original reproducer and testing both negative and zero values. I'm glad to hear that the RegionServer no longer aborts and HBase starts successfully.\r\n\r\nThis confirms that the patch addresses the reported issue. I’ll continue following up on the PR review and CI process. Thanks again for your help!","from":"developer"},{"body":"Pushed to all active branches.\r\n\r\nThanks [~mazhengxuan] [~velpro] ","from":"developer"},{"body":"Cherry-picked to branch-3.0 too.","from":"developer"}],"created":"2026-06-22T08:16:44.000+0000","description":" When `hbase.snapshot.region.timeout` is configured as a negative value, the RegionServer aborts during initialization.\r\n\r\n  The negative value is passed as the keep-alive time to a `ThreadPoolExecutor`, which throws `IllegalArgumentException`.\r\n\r\n  ## Reproducer\r\n\r\n  Use the following `hbase-site.xml` entries:\r\n\r\n  ```xml\r\n  \r\n    \r\n      hbase.rootdir\r\n      /tmp/hbase-tmp\r\n    \r\n    \r\n      hbase.cluster.distributed\r\n      false\r\n    \r\n    \r\n      hbase.unsafe.stream.capability.enforce\r\n      false\r\n    \r\n    \r\n      hbase.snapshot.region.timeout\r\n      -300000\r\n    \r\n  \r\n\r\n  Then start HBase.\r\n\r\n  Actual result\r\n\r\n  The RegionServer aborts:\r\n\r\n  ERROR [RS:0;...] regionserver.HRegionServer:\r\n  ***** ABORTING region server ... Initialization of RS failed. Hence aborting RS. *****\r\n  java.lang.IllegalArgumentException\r\n      at java.base/java.util.concurrent.ThreadPoolExecutor.(ThreadPoolExecutor.java:1293)\r\n      at java.base/java.util.concurrent.ThreadPoolExecutor.(ThreadPoolExecutor.java:1215)\r\n      at org.apache.hadoop.hbase.procedure.ProcedureMember.defaultPool(ProcedureMember.java:86)\r\n      at org.apache.hadoop.hbase.regionserver.snapshot.RegionServerSnapshotManager.initialize(RegionServerSnapshotManager.j\r\n  ava:391)\r\n\r\n  Root cause\r\n\r\n  hbase.snapshot.region.timeout is read in RegionServerSnapshotManager.initialize and used as the keep-alive time when\r\n  creating the snapshot procedure member thread pool.\r\n\r\n  A negative value is invalid for ThreadPoolExecutor.\r\n\r\n  Expected result\r\n\r\n  HBase should validate hbase.snapshot.region.timeout and reject negative values with a clear configuration error, or fall\r\n  back to a safe default.\r\n\r\n  The RegionServer should not abort with a raw IllegalArgumentException.\r\n\r\n  Notes\r\n\r\n  The failure was found by configuration fuzzing.","issue_id":"13655252","key":"HBASE-30248","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2026-07-29T09:26:44.000+0000","role":"fixed_distractor","summary":"Negative hbase.snapshot.region.timeout causes RegionServer abort during snapshot procedure initialization"} {"case_id":"12372715","cluster":"DISTRACTOR-HBASE-309","comments":[{"body":"If you delete all the values in a row, the row will become invisible until the region is compacted, at which point it will be removed.\n\nHowever, the real issues here are:\n\n- How do you (easily) delete a whole row?\n\n- How do you (easily) delete all the members of a column family?","created":"2007-07-11T22:21:59.795+0000"},{"body":"Regards deleting all members of a column family for a paritcular row, we could add a special entry into the cell at row:columnfamily. Gets and scans would check the asked for cell at row:column -- where column here is columnfamily + qualifier -- AND the value if any in the row:columnfamily position. Deleting a row would add the special value to each columnfamily for the designated row.\n\nNot pretty. ","created":"2007-08-24T22:59:59.847+0000"},{"body":"After playing in HADOOP-1784, here's some thoughts on this issue.\n\nHADOOP-1784 added deleteAll. It takes a timestamp and a column name. Internally it asks a new getKeys method for a list of all keys that match the passed column and that are older than or equal to the passed timestamp.\n\nTo delete a whole row, we could add a deleteAll that just took a timestamp and a version of getKeys that returned all entries on a row -- not just all entries in a row that match the passed column. This new deleteAll would then work like the current deleteAll adding a delete cell for every key returned by the getKeys call.\n\nDeleting all members of a column family would be a matter of removing all store files under the column family HStore (and any matching entries in memcache).","created":"2007-09-10T19:51:58.125+0000"},{"body":"HADOOP-2339 does the last item mentioned above","created":"2007-12-04T02:35:11.780+0000"},{"body":"Seems like this is pretty important for the API to be complete. Elevating to Major.","created":"2007-12-06T00:51:11.490+0000"},{"body":"Here goes.\n\nThis patch adds deleteAll(Text row) and deleteAll(Text row, long ts). It deletes all the cells for the row.","created":"2007-12-06T07:32:50.600+0000"},{"body":"There was some unnecessary garbage in HStore.java that this patch removes.","created":"2007-12-06T16:58:47.400+0000"},{"body":"Bryan Duxbury - 06/Dec/07 08:58 AM\n> There was some unnecessary garbage in HStore.java that this patch removes. \n\nI see no changes to HStore in this patch. Please explain.","created":"2007-12-06T17:30:35.216+0000"},{"body":"I meant that the original patch (1550.patch) had added some garbage \nto HStore.java which I removed in the second patch, since it wasn't \nused for anything.\n\n\n\n","created":"2007-12-06T17:43:50.381+0000"},{"body":"Tests pass locally, except for TestScanner2, which has shown to fail even without the patch.","created":"2007-12-06T17:46:38.863+0000"},{"body":"Couple of comments on patch B:\n\n+ The public HTable.deleteAll methods needs javadoc (Whats there is not javadoc). \n+ The javadoc in HRegionInterface looks like it was copy/pasted (I think your patch will fail in the javadoc validation up on hudson). Same in HRegion.","created":"2007-12-06T18:01:09.013+0000"},{"body":"This patch cleans up stack's javadoc suggestions.","created":"2007-12-06T18:14:38.954+0000"},{"body":"v3 is missing the unit test that was present in v2. Is that intentional?\n\nLooking at the v2 unit test, it adds values to column families only -- not to qualified column families. Does the deletaAll work if the you add a value to COLUMN_FAMILY[0] + \"some_qualifier\"?","created":"2007-12-06T22:45:03.608+0000"},{"body":"Stack was right. My test case was under-done. After expanding the test cases, it was clear that the patch didn't work.","created":"2007-12-07T02:07:18.870+0000"},{"body":"Phew. This patch includes the missing TestDeleteAll, which is expanded to better test the deleteAll behavior, and it passes now. The local tests suite passed all except for TestScanner2, which has a known issue at the moment.","created":"2007-12-07T20:18:03.464+0000"},{"body":"Only critical things are that teardown should pass the miniHdfs instance to StaticTestEnvironment.shutdownDfs (I know, the other tests need to do this too -- we need to fix this) and you ahve redundant imports (I'll give you $5 dollars if you come by and let me help you do your eclipse setup).\n\nOtherwise, non-blocking minor items:\n\n+ You might add test that adds a pure column family, one w/o qualifier -- just to be sure that works. You might assert that cell values actually made it into hbase before you check to see if they have been deleted.\n+ line lengths are generally <= 80 FYI.\n+ Do you need to call the 'fail' -- just asking? Any harm letting the null fall into the subsequent assertEquals?\n+ Could you do 'Test.getLength() <= 0' instead of getColumn().toString().equals(\"\") and save on a convertion to String? Same here 'origin.getColumn().equals(new Text())'.\n\nThis is interesting:\n\n- if(HLogEdit.isDeleted(readval.get())) {\n+ if(isDeleted(readkey, readval.get(), true, deletes)) {\n\nDId you find a bug in our getFull?\n\nChange this javadoc -- 'Test that the @param target ...' -- to target (drop the @param -- might confuse javadoc tool or at minimum, looks odd in produced javadoc). Same for @param origin.\n\nJust FYI, this:\n\n{code}\n+ if (target.getRow().equals(origin.getRow())) {\n+ // check the timestamp\n+ return target.getTimestamp() <= origin.getTimestamp();\n+ } else {\n+ return false;\n+ }\n{code}\ncould be written as:\n{code}\nreturn (target.getRow().equals(origin.getRow())) ? target.getTimestamp() <= origin.getTimestamp();: false;\n{code}\n\nJust FYI.\n\nOtherwise, good on you. You've (nearly) nailed a tough nugget.\n\n\n","created":"2007-12-07T21:02:43.463+0000"},{"body":"OK, fixed the javadoc issues and lengthy lines. \n\nThe null check in assertCellValuePresent is necessary because if the result is null, attempting to create a string out of it throws an exception instead of a nice assertion failure.\n\nAnd yes, I did in fact find a bug in getFull. It would only check if the cell was deleted in the HStore, not the memcache. This bug was introduced by me some time ago when I \"fixed\" getFull to not return deleted values. Now it's *really* fixed.","created":"2007-12-07T22:05:08.455+0000"},{"body":"Added running StaticTestEnvironment minihdfs shutdown. Its close of fs before shutting down the minidfs seems to be what fixed the occasionaly hudson hang. Also added for my own edification to the unit test a check that pure columnfamily cell entries get removed.\n\n+1 on patch. Ran the unit test and it passed.","created":"2007-12-07T22:29:32.037+0000"},{"body":"Sending to Hudson.","created":"2007-12-07T22:30:17.043+0000"},{"body":"Committed. Resolving. Thanks Bryan.\n\nOnly thing outstanding is deleting all cells from a particular family. Made a new issue HADOOP-2384 for that.","created":"2007-12-08T00:45:02.076+0000"}],"conversations":[{"body":"There is no support in hbase currently for deleting a row -- i.e. remove all columns and their versions keyed by a particular row id. Nor is there a means of passing in a row id and column family name having hbase delete all members of the column family (for the designated row).","from":"reporter","subject":"[hbase] No means of deleting a'row' nor all members of a column family"},{"body":"If you delete all the values in a row, the row will become invisible until the region is compacted, at which point it will be removed.\n\nHowever, the real issues here are:\n\n- How do you (easily) delete a whole row?\n\n- How do you (easily) delete all the members of a column family?","from":"developer"},{"body":"Regards deleting all members of a column family for a paritcular row, we could add a special entry into the cell at row:columnfamily. Gets and scans would check the asked for cell at row:column -- where column here is columnfamily + qualifier -- AND the value if any in the row:columnfamily position. Deleting a row would add the special value to each columnfamily for the designated row.\n\nNot pretty. ","from":"developer"},{"body":"After playing in HADOOP-1784, here's some thoughts on this issue.\n\nHADOOP-1784 added deleteAll. It takes a timestamp and a column name. Internally it asks a new getKeys method for a list of all keys that match the passed column and that are older than or equal to the passed timestamp.\n\nTo delete a whole row, we could add a deleteAll that just took a timestamp and a version of getKeys that returned all entries on a row -- not just all entries in a row that match the passed column. This new deleteAll would then work like the current deleteAll adding a delete cell for every key returned by the getKeys call.\n\nDeleting all members of a column family would be a matter of removing all store files under the column family HStore (and any matching entries in memcache).","from":"developer"},{"body":"HADOOP-2339 does the last item mentioned above","from":"developer"},{"body":"Seems like this is pretty important for the API to be complete. Elevating to Major.","from":"developer"},{"body":"Here goes.\n\nThis patch adds deleteAll(Text row) and deleteAll(Text row, long ts). It deletes all the cells for the row.","from":"developer"},{"body":"There was some unnecessary garbage in HStore.java that this patch removes.","from":"developer"},{"body":"Bryan Duxbury - 06/Dec/07 08:58 AM\n> There was some unnecessary garbage in HStore.java that this patch removes. \n\nI see no changes to HStore in this patch. Please explain.","from":"developer"},{"body":"I meant that the original patch (1550.patch) had added some garbage \nto HStore.java which I removed in the second patch, since it wasn't \nused for anything.\n\n\n\n","from":"developer"},{"body":"Tests pass locally, except for TestScanner2, which has shown to fail even without the patch.","from":"developer"},{"body":"Couple of comments on patch B:\n\n+ The public HTable.deleteAll methods needs javadoc (Whats there is not javadoc). \n+ The javadoc in HRegionInterface looks like it was copy/pasted (I think your patch will fail in the javadoc validation up on hudson). Same in HRegion.","from":"developer"},{"body":"This patch cleans up stack's javadoc suggestions.","from":"developer"},{"body":"v3 is missing the unit test that was present in v2. Is that intentional?\n\nLooking at the v2 unit test, it adds values to column families only -- not to qualified column families. Does the deletaAll work if the you add a value to COLUMN_FAMILY[0] + \"some_qualifier\"?","from":"developer"},{"body":"Stack was right. My test case was under-done. After expanding the test cases, it was clear that the patch didn't work.","from":"developer"},{"body":"Phew. This patch includes the missing TestDeleteAll, which is expanded to better test the deleteAll behavior, and it passes now. The local tests suite passed all except for TestScanner2, which has a known issue at the moment.","from":"developer"},{"body":"Only critical things are that teardown should pass the miniHdfs instance to StaticTestEnvironment.shutdownDfs (I know, the other tests need to do this too -- we need to fix this) and you ahve redundant imports (I'll give you $5 dollars if you come by and let me help you do your eclipse setup).\n\nOtherwise, non-blocking minor items:\n\n+ You might add test that adds a pure column family, one w/o qualifier -- just to be sure that works. You might assert that cell values actually made it into hbase before you check to see if they have been deleted.\n+ line lengths are generally <= 80 FYI.\n+ Do you need to call the 'fail' -- just asking? Any harm letting the null fall into the subsequent assertEquals?\n+ Could you do 'Test.getLength() <= 0' instead of getColumn().toString().equals(\"\") and save on a convertion to String? Same here 'origin.getColumn().equals(new Text())'.\n\nThis is interesting:\n\n- if(HLogEdit.isDeleted(readval.get())) {\n+ if(isDeleted(readkey, readval.get(), true, deletes)) {\n\nDId you find a bug in our getFull?\n\nChange this javadoc -- 'Test that the @param target ...' -- to target (drop the @param -- might confuse javadoc tool or at minimum, looks odd in produced javadoc). Same for @param origin.\n\nJust FYI, this:\n\n{code}\n+ if (target.getRow().equals(origin.getRow())) {\n+ // check the timestamp\n+ return target.getTimestamp() <= origin.getTimestamp();\n+ } else {\n+ return false;\n+ }\n{code}\ncould be written as:\n{code}\nreturn (target.getRow().equals(origin.getRow())) ? target.getTimestamp() <= origin.getTimestamp();: false;\n{code}\n\nJust FYI.\n\nOtherwise, good on you. You've (nearly) nailed a tough nugget.\n\n\n","from":"developer"},{"body":"OK, fixed the javadoc issues and lengthy lines. \n\nThe null check in assertCellValuePresent is necessary because if the result is null, attempting to create a string out of it throws an exception instead of a nice assertion failure.\n\nAnd yes, I did in fact find a bug in getFull. It would only check if the cell was deleted in the HStore, not the memcache. This bug was introduced by me some time ago when I \"fixed\" getFull to not return deleted values. Now it's *really* fixed.","from":"developer"},{"body":"Added running StaticTestEnvironment minihdfs shutdown. Its close of fs before shutting down the minidfs seems to be what fixed the occasionaly hudson hang. Also added for my own edification to the unit test a check that pure columnfamily cell entries get removed.\n\n+1 on patch. Ran the unit test and it passed.","from":"developer"},{"body":"Sending to Hudson.","from":"developer"},{"body":"Committed. Resolving. Thanks Bryan.\n\nOnly thing outstanding is deleting all cells from a particular family. Made a new issue HADOOP-2384 for that.","from":"developer"}],"created":"2007-06-29T19:59:55.000+0000","description":"There is no support in hbase currently for deleting a row -- i.e. remove all columns and their versions keyed by a particular row id. Nor is there a means of passing in a row id and column family name having hbase delete all members of the column family (for the designated row).","issue_id":"12372715","key":"HBASE-309","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2007-12-08T00:45:02.000+0000","role":"fixed_distractor","summary":"[hbase] No means of deleting a'row' nor all members of a column family"} {"case_id":"12477785","cluster":"DISTRACTOR-HBASE-3130","comments":[{"body":"Moving out of 0.92.0. Pull it back in if you think different.","created":"2011-06-10T22:45:41.360+0000"},{"body":"We will need this at Salesforce.com. So I am likely going to work on this soon.\n\nWhat is the exact problem?\nIf it would as simple as just recreating ReplicationZookeeper or using RecoverableZookeeper J-D would have probably just done it :)\n","created":"2011-08-31T03:27:14.505+0000"},{"body":"RecoverableZookeeper only recovers from recoverable exceptions, SessionExpired isn't one of those. Basically you have to catch it on the paths that use ZK, mostly hidden behind ReplicationZookeeper. Treat it as a normal exception, retry it every once in a while.","created":"2011-08-31T03:31:49.711+0000"},{"body":"Looks like it got even worse recently, we got a situation where the SessionExpired was treated like if it was the RS's own and it FATAL'ed.","created":"2011-09-09T22:52:27.593+0000"},{"body":"@J-D:\nCan you post snippet of server log containing the FATAL portion ?","created":"2011-09-10T13:55:23.751+0000"},{"body":"Yes that would be useful. Somebody at Salesforce is currently looking at this issue anyway.","created":"2011-09-10T16:42:13.536+0000"},{"body":"Here it goes:\n\n{quote}\n2011-09-09 19:44:28,224 FATAL org.apache.hadoop.hbase.regionserver.HRegionServer: ABORTING region server serverName=sv4r17s40,60020,1313587209632, load=(requests=4292, regions=186, usedHeap=11929, maxHeap=24749): connection to cluster: 5-0x130d4937f890066 connection to cluster: 5-0x130d4937f890066 received expired from ZooKeeper, aborting\norg.apache.zookeeper.KeeperException$SessionExpiredException: KeeperErrorCode = Session expired\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher.connectionEvent(ZooKeeperWatcher.java:343)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher.process(ZooKeeperWatcher.java:261)\n\tat org.apache.zookeeper.ClientCnxn$EventThread.processEvent(ClientCnxn.java:530)\n\tat org.apache.zookeeper.ClientCnxn$EventThread.run(ClientCnxn.java:506)\n\n{quote}\n\nAs you can see it's pretty generic, I could trace it was the peer connection with the \"connection to cluster\". Moreover the fix will take place around ReplicationPeer which contains a ZKW that requires an Abortable which, at the moment, is the RS itself. Instead we should pass our own, or maybe ReplicationSource should implement it.","created":"2011-09-10T17:05:42.751+0000"},{"body":"Hi all,\n\nI am the \"somebody at salesforce\" that is currently looking at this issue. I should have a sketch patch for you guys shortly.\n\nHere is the general gist of what I am doing:\n1. I am now specifically catching the SESSIONEXPIRED KeeperExceptions that are thrown by methods which use a ZookeeperWatcher from ReplicationPeer.\n\n2. I am adding a public method to ReplicationZookeeper that is responsible for retrying/opening a new connection to the peer zookeeper cluster. This method will take a ReplicationPeer (the old rp from the peer we are trying to reconnect to) and an Abortable (which will be the ReplicationSource). It will return a new ReplicationPeer with a fresh connection.\n\n3. I will modify ReplicationSource so that it implements the Abortable interface.\n\n4. Currently I am going back and forth about two retry strategies:\nA. There is a sleepMultiplier and a maxNumberOfRetries. The source will try to reconnect until the maxNumberOfRetries is exceeded, after which it will abort.\n\nB. Same as A, except after exceeding maxNumberOfRetries it will continue retrying indefinitely, but at a very low frequency (i.e every 5 min). With this approach I would implement it in a way such that replication could always be turned off, at which point the retries would stop.\n\nThoughts/comments/suggestions are always appreciated. I am excited to contribute!\n\nThanks,\nChris Trezzo","created":"2011-09-11T02:28:42.216+0000"},{"body":"Hey Chris :)\nI'd say 4.B. until an admin explicitly stops replication or removes the peer.\nI would also probably handle this in ReplicationSource's main loop.\n","created":"2011-09-11T02:47:52.309+0000"},{"body":"This seems like an important bug fix, can we but this into 0.92 even after we branched it?","created":"2011-09-13T18:29:08.336+0000"},{"body":"@Lars Bugs and doc improvements yes. Features no.","created":"2011-09-13T18:34:02.152+0000"},{"body":"@J-D While testing we encountered the problem above (where the RegionServer closes) too.\n\nShould we fix that part in a separate jira, as this one might take a bit longer? The stop gap change is to log the problem and move on, i.e. simply pass an adhoc Abortable a peer's zk watcher.\n","created":"2011-09-14T21:03:12.955+0000"},{"body":"@Chris\n\nSorry for the late answer.\n\nbq. 1. I am now specifically catching the SESSIONEXPIRED KeeperExceptions that are thrown by methods which use a ZookeeperWatcher from ReplicationPeer.\n\nThey are all catch at a super low lovel (ZKW) so I don't think you need that.\n\nbq. 2. I am adding a public method to ReplicationZookeeper that is responsible for retrying/opening a new connection to the peer zookeeper cluster.\n\nThere might be some synchronization problems with ReplicationSource, watch out.\n\nbq. 3. I will modify ReplicationSource so that it implements the Abortable interface.\n\nI'm -0, I'd prefer if all the handling was kept in ReplicationZookeeper. Ideally for ReplicationSource it would just see that it can't reach the peer for some reason and retry.\n\nbq. 4. Currently I am going back and forth about two retry strategies:\n\nDefinitely 4B, that's how the code works at the moment. Wait until the peer comes back, the operator can always turn it off.","created":"2011-09-15T18:33:30.212+0000"},{"body":"@J-D\n\nThanks for the comments. I am attaching a first patch that I worked on with Lars. There are a few things that are slightly different from your comments above, but I figured I would post it anyways just to get the ball rolling. Let me know what you think.\n\nHere is an overview of the changes:\n\n1. Added a new reloadZkWatcher method to ReplicationPeer. This method is responsible for reseting the ReplicationPeer's ZookeeperWatcher when the connection for the old watcher dies.\n\n2. The crux of the change is in ReplicationSource.chooseSinks. We catch the KeeperException there, and if it is a ConnectionLoss or SessionExpired we reset that peer's ZookeeperWatcher using the new method from (1). Retries are handled by the existing loops in the callers of chooseSinks (i.e. connectToPeers and shipEdits).\n\n3. ReplicationPeer now implements Abortable and is passed in as the Abortable for the peer's ZookeeperWatcher. The abort method simply logs the fact that it was called and returns.\n\nI ran TestReplication, TestMasterReplication and TestMultiSlaveReplication tests with no failures. I also manually tested the case where the slave cluster fails, and confirmed that the ReplicationSource does recover and resumes replication once the peer cluster comes back.\n\nThanks,\nChris","created":"2011-09-15T20:50:24.774+0000"},{"body":"Made Chris a contributor, assigning to him.","created":"2011-09-15T21:39:28.304+0000"},{"body":"Some comments on the contribution itself since you're a newcomer:\n\n - The logging should be done like we do it elsewhere in the code, by using the Log class provided by Commons Logging and then using the LogFactory. ReplicationSource is an example.\n - When logging, try to include relevant information because with all different threads logging it can be hard to follow. For example, this could tell which peer we're talking about:\nbq. Log.info(\"Refreshing ZookeeperWatcher\");\n - Lines should be max 80 characters long (the one where you're catching SessionExpiredException for example).\n\nMore info in this chapter of the book: http://hbase.apache.org/book/developer.html\n\nNow on the content, I still think that the management of the connection should be done inside ReplicationZookeeper. In its current form the patch exposes a feature envy between ReplicationSource and ReplicationPeer, might just be better to contain the functionality in between.\n\nThanks for working on this Chris!","created":"2011-09-15T22:32:05.536+0000"},{"body":"@J-D\n\nThanks for the response! Attaching a new patch with your comments in mind.\n\nHere are the changes:\n\n1. Fixed the logging in ReplicationPeer so that it uses the commons logging LogFactory.\n\n2. Fixed the line with 80+ characters.\n\n3. Moved the connection management logic from ReplicationSource down to ReplicationZookeeper.getSlavesAddresses.\n\n4. I kept the ReplicationPeer.reloadZkWatcher method for a couple reasons. (1) In order to move all the ZookeeperWatcher logic out of ReplicationPeer, we would need to have a setZkw method. This potentially allows for null ZookeeperWatchers within a ReplicationPeer (the reloadZkWatcher method seems like a slightly cleaner choice). (2) It is basically doing the exact same thing that the ReplicationPeer constructor is already doing.\n\nLet me know what you think.\n\nThanks,\nChris","created":"2011-09-16T01:43:52.868+0000"},{"body":"+1\n\nI think the patch is nice because it\no Keeps retrying logic in ReplicationSource\no Removes handling KeeperExceptions from ReplicationSource\no Has ReplicationPeer manage its watcher\no Now the zoo keeper failure handling is in ReplicationZookeeper\n\nNit: You could use Collections.emptyList() instead of new ArrayList(0), and then just fall through to the set- and return getRegionServers.\n","created":"2011-09-16T02:21:52.481+0000"},{"body":"Submitting a new patch that addresses the Collections.emptyList() comment from Lars.","created":"2011-09-16T18:50:39.082+0000"},{"body":"I like where this is going.\n\nFurther improvements:\n\n - In RP, the creation of the ZKW should be refactored\n - The logging in abort() could use some more explanation\n - Would it be possible to have a unit test?","created":"2011-09-16T18:57:19.386+0000"},{"body":"@Chris If you need help writing unit test, ask for it.","created":"2011-09-16T18:59:37.401+0000"},{"body":"@Chris \n\nStack was telling me that you were wondering if you need 2 ZK ensembles, I personally don't think so since the session for the master -> slave connection is different from all the others. So you can just expire that one (maybe across all region servers on the master side).","created":"2011-09-16T23:39:28.288+0000"},{"body":"@J-D\n\nThat sounds like a good plan. I'll try writing a test, and if I run into anything weird I'll let you know.","created":"2011-09-17T02:30:39.579+0000"},{"body":"@J-D\n\nNow that I have looked at the test code a bit, I have a question:\n\nMy current understanding is that to kill the master->slave connection, you need to somehow get the session id and session password for the ReplicationPeer's zookeeper session (i.e. you need the ZookeeperWatcher instance). Currently, this is not exposed. Also, this does not seem like something we would want to expose if the only motivation is for testing. Any thoughts?\n\nI could be missing something obvious.\n\nThanks!\nChris","created":"2011-09-21T04:01:33.272+0000"},{"body":"At the very minimum you could do a TestReplicationPeer that tests the session \"recovery\" code. An integration test might be harder since you have to fiddle with the internals, maybe explore the avenue of having a test that resides in the same package (o.a.h.h.r.replication) and expose the methods only there.","created":"2011-09-21T04:58:55.121+0000"},{"body":"Another thing about the latest patch, this line needs to be removed in RP:\n\nbq. * @param zkw zookeeper connection to the peer","created":"2011-09-27T20:29:35.899+0000"},{"body":"Also I'd add that I was able to test the patch (on 0.90) and it really works, proof:\n\nFirst it loses the connection:\n{quote}\n2011-09-27 16:44:54,984 WARN org.apache.hadoop.hbase.replication.ReplicationPeer: connection to cluster: 10.10.30.7:2181:/hbase1-0x132ad0f29d70017 connection to cluster: 10.10.30.7:2181:/hbase1-0x132ad0f29d70017 received expired from ZooKeeper, aborting\norg.apache.zookeeper.KeeperException$SessionExpiredException: KeeperErrorCode = Session expired\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher.connectionEvent(ZooKeeperWatcher.java:343)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher.process(ZooKeeperWatcher.java:261)\n\tat org.apache.zookeeper.ClientCnxn$EventThread.processEvent(ClientCnxn.java:530)\n\tat org.apache.zookeeper.ClientCnxn$EventThread.run(ClientCnxn.java:506)\n2011-09-27 16:44:54,984 INFO org.apache.zookeeper.ClientCnxn: EventThread shut down\n{quote}\n\nThen later when it tries to replicate it tries to talk to ZK again and it works after a reload:\n\n{quote}\n2011-09-27 16:49:03,738 DEBUG org.apache.hadoop.hbase.replication.regionserver.ReplicationSource: Since we are unable to replicate, sleeping 1000 times 10\n2011-09-27 16:49:13,738 WARN org.apache.hadoop.hbase.replication.ReplicationZookeeper: Lost the ZooKeeper connection for peer 1\norg.apache.zookeeper.KeeperException$SessionExpiredException: KeeperErrorCode = Session expired for /hbase1/rs\n\tat org.apache.zookeeper.KeeperException.create(KeeperException.java:118)\n\tat org.apache.zookeeper.KeeperException.create(KeeperException.java:42)\n\tat org.apache.zookeeper.ZooKeeper.getChildren(ZooKeeper.java:1243)\n\tat org.apache.hadoop.hbase.zookeeper.ZKUtil.listChildrenNoWatch(ZKUtil.java:389)\n\tat org.apache.hadoop.hbase.zookeeper.ZKUtil.listChildrenAndGetAsAddresses(ZKUtil.java:355)\n\tat org.apache.hadoop.hbase.replication.ReplicationZookeeper.fetchSlavesAddresses(ReplicationZookeeper.java:268)\n\tat org.apache.hadoop.hbase.replication.ReplicationZookeeper.getSlavesAddresses(ReplicationZookeeper.java:239)\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSource.chooseSinks(ReplicationSource.java:205)\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSource.shipEdits(ReplicationSource.java:588)\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSource.run(ReplicationSource.java:341)\n2011-09-27 16:49:13,772 INFO org.apache.zookeeper.ZooKeeper: Initiating client connection, connectString=10.10.30.7:2181 sessionTimeout=20000 watcher=connection to cluster: 10.10.30.7:2181:/hbase1\n2011-09-27 16:49:13,773 INFO org.apache.zookeeper.ClientCnxn: Opening socket connection to server /10.10.30.7:2181\n2011-09-27 16:49:14,111 INFO org.apache.zookeeper.ClientCnxn: Socket connection established to hbasedev.sfo.stumble.net/10.10.30.7:2181, initiating session\n2011-09-27 16:49:14,140 INFO org.apache.zookeeper.ClientCnxn: Session establishment complete on server hbasedev.sfo.stumble.net/10.10.30.7:2181, sessionid = 0x132ad0f29d70024, negotiated timeout = 20000\n{quote}","created":"2011-09-27T23:54:57.261+0000"},{"body":"You sound like you are surprised :)\n\n\n","created":"2011-09-28T02:12:45.691+0000"},{"body":"More like joyful amusement, it's been broken for so long... we're pushing this in prod really soon.","created":"2011-09-28T02:16:33.456+0000"},{"body":"I finally had some time last night to look at the test code. Here is a new patch that addresses the above comments and adds a new test. TestReplicationPeer verifies the refresh ZooKeeperWatcher functionality in ReplicationPeer.\n\nAs per J-D's comment above, doing more of an integration test at a higher level seems to be quite tricky and may require a large change.\n\nLet me know what you guys think.\n\nThanks!\nChris","created":"2011-09-28T22:42:24.430+0000"},{"body":"I am also glad the patch brings joyful amusement :-)","created":"2011-09-28T22:51:26.041+0000"},{"body":"Here's another iteration on the patch, I did the following:\n\n - Modified the test to not create a full cluster but just ZK, had to add an option in the testing utility to not create an HTable. That also fixed the test.\n - Removed the copyright in RP and fixed some javadoc.\n - Also in RP the constructor now calls reload (that was the refactoring I mentioned earlier).\n\nLet me know if it makes sense to you.","created":"2011-09-29T22:35:25.222+0000"},{"body":"That looks good to me!","created":"2011-09-29T23:01:48.220+0000"},{"body":"+1","created":"2011-09-29T23:03:39.316+0000"},{"body":"Committed to 0.92 and trunk, thanks for the good work Chris!","created":"2011-09-29T23:41:18.212+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T12:41:58.579+0000"}],"conversations":[{"body":"Currently ReplicationSource cannot recover when its zookeeper connection to its remote cluster expires. HLogs are still being tracked, but a cluster restart is required to continue replication (or a rolling restart).","from":"reporter","subject":"[replication] ReplicationSource can't recover from session expired on remote clusters"},{"body":"Moving out of 0.92.0. Pull it back in if you think different.","from":"developer"},{"body":"We will need this at Salesforce.com. So I am likely going to work on this soon.\n\nWhat is the exact problem?\nIf it would as simple as just recreating ReplicationZookeeper or using RecoverableZookeeper J-D would have probably just done it :)\n","from":"developer"},{"body":"RecoverableZookeeper only recovers from recoverable exceptions, SessionExpired isn't one of those. Basically you have to catch it on the paths that use ZK, mostly hidden behind ReplicationZookeeper. Treat it as a normal exception, retry it every once in a while.","from":"developer"},{"body":"Looks like it got even worse recently, we got a situation where the SessionExpired was treated like if it was the RS's own and it FATAL'ed.","from":"developer"},{"body":"@J-D:\nCan you post snippet of server log containing the FATAL portion ?","from":"developer"},{"body":"Yes that would be useful. Somebody at Salesforce is currently looking at this issue anyway.","from":"developer"},{"body":"Here it goes:\n\n{quote}\n2011-09-09 19:44:28,224 FATAL org.apache.hadoop.hbase.regionserver.HRegionServer: ABORTING region server serverName=sv4r17s40,60020,1313587209632, load=(requests=4292, regions=186, usedHeap=11929, maxHeap=24749): connection to cluster: 5-0x130d4937f890066 connection to cluster: 5-0x130d4937f890066 received expired from ZooKeeper, aborting\norg.apache.zookeeper.KeeperException$SessionExpiredException: KeeperErrorCode = Session expired\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher.connectionEvent(ZooKeeperWatcher.java:343)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher.process(ZooKeeperWatcher.java:261)\n\tat org.apache.zookeeper.ClientCnxn$EventThread.processEvent(ClientCnxn.java:530)\n\tat org.apache.zookeeper.ClientCnxn$EventThread.run(ClientCnxn.java:506)\n\n{quote}\n\nAs you can see it's pretty generic, I could trace it was the peer connection with the \"connection to cluster\". Moreover the fix will take place around ReplicationPeer which contains a ZKW that requires an Abortable which, at the moment, is the RS itself. Instead we should pass our own, or maybe ReplicationSource should implement it.","from":"developer"},{"body":"Hi all,\n\nI am the \"somebody at salesforce\" that is currently looking at this issue. I should have a sketch patch for you guys shortly.\n\nHere is the general gist of what I am doing:\n1. I am now specifically catching the SESSIONEXPIRED KeeperExceptions that are thrown by methods which use a ZookeeperWatcher from ReplicationPeer.\n\n2. I am adding a public method to ReplicationZookeeper that is responsible for retrying/opening a new connection to the peer zookeeper cluster. This method will take a ReplicationPeer (the old rp from the peer we are trying to reconnect to) and an Abortable (which will be the ReplicationSource). It will return a new ReplicationPeer with a fresh connection.\n\n3. I will modify ReplicationSource so that it implements the Abortable interface.\n\n4. Currently I am going back and forth about two retry strategies:\nA. There is a sleepMultiplier and a maxNumberOfRetries. The source will try to reconnect until the maxNumberOfRetries is exceeded, after which it will abort.\n\nB. Same as A, except after exceeding maxNumberOfRetries it will continue retrying indefinitely, but at a very low frequency (i.e every 5 min). With this approach I would implement it in a way such that replication could always be turned off, at which point the retries would stop.\n\nThoughts/comments/suggestions are always appreciated. I am excited to contribute!\n\nThanks,\nChris Trezzo","from":"developer"},{"body":"Hey Chris :)\nI'd say 4.B. until an admin explicitly stops replication or removes the peer.\nI would also probably handle this in ReplicationSource's main loop.\n","from":"developer"},{"body":"This seems like an important bug fix, can we but this into 0.92 even after we branched it?","from":"developer"},{"body":"@Lars Bugs and doc improvements yes. Features no.","from":"developer"},{"body":"@J-D While testing we encountered the problem above (where the RegionServer closes) too.\n\nShould we fix that part in a separate jira, as this one might take a bit longer? The stop gap change is to log the problem and move on, i.e. simply pass an adhoc Abortable a peer's zk watcher.\n","from":"developer"},{"body":"@Chris\n\nSorry for the late answer.\n\nbq. 1. I am now specifically catching the SESSIONEXPIRED KeeperExceptions that are thrown by methods which use a ZookeeperWatcher from ReplicationPeer.\n\nThey are all catch at a super low lovel (ZKW) so I don't think you need that.\n\nbq. 2. I am adding a public method to ReplicationZookeeper that is responsible for retrying/opening a new connection to the peer zookeeper cluster.\n\nThere might be some synchronization problems with ReplicationSource, watch out.\n\nbq. 3. I will modify ReplicationSource so that it implements the Abortable interface.\n\nI'm -0, I'd prefer if all the handling was kept in ReplicationZookeeper. Ideally for ReplicationSource it would just see that it can't reach the peer for some reason and retry.\n\nbq. 4. Currently I am going back and forth about two retry strategies:\n\nDefinitely 4B, that's how the code works at the moment. Wait until the peer comes back, the operator can always turn it off.","from":"developer"},{"body":"@J-D\n\nThanks for the comments. I am attaching a first patch that I worked on with Lars. There are a few things that are slightly different from your comments above, but I figured I would post it anyways just to get the ball rolling. Let me know what you think.\n\nHere is an overview of the changes:\n\n1. Added a new reloadZkWatcher method to ReplicationPeer. This method is responsible for reseting the ReplicationPeer's ZookeeperWatcher when the connection for the old watcher dies.\n\n2. The crux of the change is in ReplicationSource.chooseSinks. We catch the KeeperException there, and if it is a ConnectionLoss or SessionExpired we reset that peer's ZookeeperWatcher using the new method from (1). Retries are handled by the existing loops in the callers of chooseSinks (i.e. connectToPeers and shipEdits).\n\n3. ReplicationPeer now implements Abortable and is passed in as the Abortable for the peer's ZookeeperWatcher. The abort method simply logs the fact that it was called and returns.\n\nI ran TestReplication, TestMasterReplication and TestMultiSlaveReplication tests with no failures. I also manually tested the case where the slave cluster fails, and confirmed that the ReplicationSource does recover and resumes replication once the peer cluster comes back.\n\nThanks,\nChris","from":"developer"},{"body":"Made Chris a contributor, assigning to him.","from":"developer"},{"body":"Some comments on the contribution itself since you're a newcomer:\n\n - The logging should be done like we do it elsewhere in the code, by using the Log class provided by Commons Logging and then using the LogFactory. ReplicationSource is an example.\n - When logging, try to include relevant information because with all different threads logging it can be hard to follow. For example, this could tell which peer we're talking about:\nbq. Log.info(\"Refreshing ZookeeperWatcher\");\n - Lines should be max 80 characters long (the one where you're catching SessionExpiredException for example).\n\nMore info in this chapter of the book: http://hbase.apache.org/book/developer.html\n\nNow on the content, I still think that the management of the connection should be done inside ReplicationZookeeper. In its current form the patch exposes a feature envy between ReplicationSource and ReplicationPeer, might just be better to contain the functionality in between.\n\nThanks for working on this Chris!","from":"developer"},{"body":"@J-D\n\nThanks for the response! Attaching a new patch with your comments in mind.\n\nHere are the changes:\n\n1. Fixed the logging in ReplicationPeer so that it uses the commons logging LogFactory.\n\n2. Fixed the line with 80+ characters.\n\n3. Moved the connection management logic from ReplicationSource down to ReplicationZookeeper.getSlavesAddresses.\n\n4. I kept the ReplicationPeer.reloadZkWatcher method for a couple reasons. (1) In order to move all the ZookeeperWatcher logic out of ReplicationPeer, we would need to have a setZkw method. This potentially allows for null ZookeeperWatchers within a ReplicationPeer (the reloadZkWatcher method seems like a slightly cleaner choice). (2) It is basically doing the exact same thing that the ReplicationPeer constructor is already doing.\n\nLet me know what you think.\n\nThanks,\nChris","from":"developer"},{"body":"+1\n\nI think the patch is nice because it\no Keeps retrying logic in ReplicationSource\no Removes handling KeeperExceptions from ReplicationSource\no Has ReplicationPeer manage its watcher\no Now the zoo keeper failure handling is in ReplicationZookeeper\n\nNit: You could use Collections.emptyList() instead of new ArrayList(0), and then just fall through to the set- and return getRegionServers.\n","from":"developer"},{"body":"Submitting a new patch that addresses the Collections.emptyList() comment from Lars.","from":"developer"},{"body":"I like where this is going.\n\nFurther improvements:\n\n - In RP, the creation of the ZKW should be refactored\n - The logging in abort() could use some more explanation\n - Would it be possible to have a unit test?","from":"developer"},{"body":"@Chris If you need help writing unit test, ask for it.","from":"developer"},{"body":"@Chris \n\nStack was telling me that you were wondering if you need 2 ZK ensembles, I personally don't think so since the session for the master -> slave connection is different from all the others. So you can just expire that one (maybe across all region servers on the master side).","from":"developer"},{"body":"@J-D\n\nThat sounds like a good plan. I'll try writing a test, and if I run into anything weird I'll let you know.","from":"developer"},{"body":"@J-D\n\nNow that I have looked at the test code a bit, I have a question:\n\nMy current understanding is that to kill the master->slave connection, you need to somehow get the session id and session password for the ReplicationPeer's zookeeper session (i.e. you need the ZookeeperWatcher instance). Currently, this is not exposed. Also, this does not seem like something we would want to expose if the only motivation is for testing. Any thoughts?\n\nI could be missing something obvious.\n\nThanks!\nChris","from":"developer"},{"body":"At the very minimum you could do a TestReplicationPeer that tests the session \"recovery\" code. An integration test might be harder since you have to fiddle with the internals, maybe explore the avenue of having a test that resides in the same package (o.a.h.h.r.replication) and expose the methods only there.","from":"developer"},{"body":"Another thing about the latest patch, this line needs to be removed in RP:\n\nbq. * @param zkw zookeeper connection to the peer","from":"developer"},{"body":"Also I'd add that I was able to test the patch (on 0.90) and it really works, proof:\n\nFirst it loses the connection:\n{quote}\n2011-09-27 16:44:54,984 WARN org.apache.hadoop.hbase.replication.ReplicationPeer: connection to cluster: 10.10.30.7:2181:/hbase1-0x132ad0f29d70017 connection to cluster: 10.10.30.7:2181:/hbase1-0x132ad0f29d70017 received expired from ZooKeeper, aborting\norg.apache.zookeeper.KeeperException$SessionExpiredException: KeeperErrorCode = Session expired\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher.connectionEvent(ZooKeeperWatcher.java:343)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher.process(ZooKeeperWatcher.java:261)\n\tat org.apache.zookeeper.ClientCnxn$EventThread.processEvent(ClientCnxn.java:530)\n\tat org.apache.zookeeper.ClientCnxn$EventThread.run(ClientCnxn.java:506)\n2011-09-27 16:44:54,984 INFO org.apache.zookeeper.ClientCnxn: EventThread shut down\n{quote}\n\nThen later when it tries to replicate it tries to talk to ZK again and it works after a reload:\n\n{quote}\n2011-09-27 16:49:03,738 DEBUG org.apache.hadoop.hbase.replication.regionserver.ReplicationSource: Since we are unable to replicate, sleeping 1000 times 10\n2011-09-27 16:49:13,738 WARN org.apache.hadoop.hbase.replication.ReplicationZookeeper: Lost the ZooKeeper connection for peer 1\norg.apache.zookeeper.KeeperException$SessionExpiredException: KeeperErrorCode = Session expired for /hbase1/rs\n\tat org.apache.zookeeper.KeeperException.create(KeeperException.java:118)\n\tat org.apache.zookeeper.KeeperException.create(KeeperException.java:42)\n\tat org.apache.zookeeper.ZooKeeper.getChildren(ZooKeeper.java:1243)\n\tat org.apache.hadoop.hbase.zookeeper.ZKUtil.listChildrenNoWatch(ZKUtil.java:389)\n\tat org.apache.hadoop.hbase.zookeeper.ZKUtil.listChildrenAndGetAsAddresses(ZKUtil.java:355)\n\tat org.apache.hadoop.hbase.replication.ReplicationZookeeper.fetchSlavesAddresses(ReplicationZookeeper.java:268)\n\tat org.apache.hadoop.hbase.replication.ReplicationZookeeper.getSlavesAddresses(ReplicationZookeeper.java:239)\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSource.chooseSinks(ReplicationSource.java:205)\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSource.shipEdits(ReplicationSource.java:588)\n\tat org.apache.hadoop.hbase.replication.regionserver.ReplicationSource.run(ReplicationSource.java:341)\n2011-09-27 16:49:13,772 INFO org.apache.zookeeper.ZooKeeper: Initiating client connection, connectString=10.10.30.7:2181 sessionTimeout=20000 watcher=connection to cluster: 10.10.30.7:2181:/hbase1\n2011-09-27 16:49:13,773 INFO org.apache.zookeeper.ClientCnxn: Opening socket connection to server /10.10.30.7:2181\n2011-09-27 16:49:14,111 INFO org.apache.zookeeper.ClientCnxn: Socket connection established to hbasedev.sfo.stumble.net/10.10.30.7:2181, initiating session\n2011-09-27 16:49:14,140 INFO org.apache.zookeeper.ClientCnxn: Session establishment complete on server hbasedev.sfo.stumble.net/10.10.30.7:2181, sessionid = 0x132ad0f29d70024, negotiated timeout = 20000\n{quote}","from":"developer"},{"body":"You sound like you are surprised :)\n\n\n","from":"developer"},{"body":"More like joyful amusement, it's been broken for so long... we're pushing this in prod really soon.","from":"developer"},{"body":"I finally had some time last night to look at the test code. Here is a new patch that addresses the above comments and adds a new test. TestReplicationPeer verifies the refresh ZooKeeperWatcher functionality in ReplicationPeer.\n\nAs per J-D's comment above, doing more of an integration test at a higher level seems to be quite tricky and may require a large change.\n\nLet me know what you guys think.\n\nThanks!\nChris","from":"developer"},{"body":"I am also glad the patch brings joyful amusement :-)","from":"developer"},{"body":"Here's another iteration on the patch, I did the following:\n\n - Modified the test to not create a full cluster but just ZK, had to add an option in the testing utility to not create an HTable. That also fixed the test.\n - Removed the copyright in RP and fixed some javadoc.\n - Also in RP the constructor now calls reload (that was the refactoring I mentioned earlier).\n\nLet me know if it makes sense to you.","from":"developer"},{"body":"That looks good to me!","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed to 0.92 and trunk, thanks for the good work Chris!","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2010-10-19T21:38:01.000+0000","description":"Currently ReplicationSource cannot recover when its zookeeper connection to its remote cluster expires. HLogs are still being tracked, but a cluster restart is required to continue replication (or a rolling restart).","issue_id":"12477785","key":"HBASE-3130","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2011-09-29T23:41:18.000+0000","role":"fixed_distractor","summary":"[replication] ReplicationSource can't recover from session expired on remote clusters"} {"case_id":"12478686","cluster":"DISTRACTOR-HBASE-3170","comments":[{"body":"We are hitting this too, this is a really unexpected behaviour. Why getting empty key should return data of the first row in table? Reproduced in CDH3u4 (0.90.6):\n\n{code}\nhbase(main):005:0> create 'emptykey', {NAME=>'data', VERSION=>1}\n0 row(s) in 0.2070 seconds\n\nhbase(main):011:0> get 'emptykey', '' \nCOLUMN CELL \n0 row(s) in 0.0120 seconds\n\nhbase(main):006:0> put 'emptykey', 'a', 'data:a', '1234'\n0 row(s) in 0.1980 seconds\n\nhbase(main):007:0> put 'emptykey', 'b', 'data:b', '5678'\n0 row(s) in 0.0070 seconds\n\nhbase(main):008:0> scan 'emptykey' \nROW COLUMN+CELL \n a column=data:a, timestamp=1350989443394, value=1234 \n b column=data:b, timestamp=1350989450499, value=5678 \n2 row(s) in 0.0660 seconds\n\nhbase(main):009:0> get 'emptykey', ''\nCOLUMN CELL \n data:a timestamp=1350989443394, value=1234 \n1 row(s) in 0.0120 seconds\n{code}\n\nIt works the same way also using thrift.\n\nWe can even see, that empty key is supported in fact.\n{code}\nhbase(main):012:0> put 'emptykey', '', 'data:c', '90' \n0 row(s) in 0.0130 seconds\n\nhbase(main):013:0> get 'emptykey', '' \nCOLUMN CELL \n data:c timestamp=1350989869682, value=90 \n1 row(s) in 0.0120 seconds\n\nhbase(main):018:0> scan 'emptykey' \nROW COLUMN+CELL \n column=data:c, timestamp=1350989869682, value=90 \n a column=data:a, timestamp=1350989933922, value=1234 \n b column=data:b, timestamp=1350989937820, value=5678 \n3 row(s) in 0.0180 seconds\n{code}","created":"2012-10-23T11:00:13.438+0000"},{"body":"Bad bug. Pulling into 0.96 so we fix it soon.","created":"2012-10-23T21:44:43.583+0000"},{"body":"Attached is a patch. One change is to keep track of the fact that we are doing a 'get' as opposed to a 'scan'. When this was done, it broke the case of a legitimate 'get' with an empty rowkey. So there is a change to handle that case - basically, in the RegionScannerImpl.nextInternal I do special checks for empty start & end keys (in the new method isStopRowConsideringEmptyRows).. \n\nThanks to Enis for some useful discussions..","created":"2013-01-19T02:33:33.762+0000"},{"body":"Let's see what hudson says.","created":"2013-01-19T02:33:44.811+0000"},{"body":"{code}\n+public class TestGetEmptyRowKey {\n{code}\nThe new test would be a medium test, right ? Please add annotation.\n{code}\n+ //don't stop\n+ return false;\n{code}\nYou can save the comment such as the one above if you add javadoc for isStopRowConsideringEmptyRows().","created":"2013-01-19T02:54:45.396+0000"},{"body":"It turns out that we need to keep the original behavior of Scan.isGetScan()\n\nPatch v2 introduced two new fields for RegionScannerImpl and made the logic in isStopRowConsideringEmptyRows() symmetrical.\n\nThe following tests passed:\n\n 1430 mt -Dtest=TestScanWithBloomError,TestStoreFile,TestGetEmptyRowKey\n 1433 mt -Dtest=TestHRegion,TestColumnSeeking,TestMultiColumnScanner","created":"2013-01-19T06:41:58.265+0000"},{"body":"I realized that the boolean field, RegionScannerImpl.scanCreatedFromGet, is redundant with RegionScannerImpl.stopRowFromScan.","created":"2013-01-19T14:36:37.194+0000"},{"body":"I ran TestConstraint twice locally and it passed.","created":"2013-01-19T16:09:21.198+0000"},{"body":"bq. It turns out that we need to keep the original behavior of Scan.isGetScan()\n\nWhy is that? Seems odd having isGetScan and isScanCreatedFromGet. I'd think that a method named isGetScan would return true if this was a Scan made to service a Get?\n\nIt seems like isGetScan is trying to look at scanner specs to see if a Get scan rather than just at whether or not the Scan constructor that takes a Get was used.\n\nLooking at this, if you pass a Get a null or empty row, should we just throw an exception immediately and not even start the Scan going? Let a Get with an empty row be illegal?","created":"2013-01-19T22:31:45.094+0000"},{"body":"bq. Seems odd having isGetScan and isScanCreatedFromGet.\nI agree. isGetScan() might have started with what we normally expect. Currently isGetScan() is used to serve other use cases.\n\nbq. should we just throw an exception immediately\nThat would simplify this issue a lot.\n\nLet'see what [~devaraj] says about the above proposal.","created":"2013-01-19T23:11:21.712+0000"},{"body":"Okay this patch is a slight rework on my previous patch. It fixes the problem at hand I think and makes the distinction between scan and get clearer...\n\nOn the point about disabling Get with empty row, I was also having the same opinion until I started digging into why Get with empty row key is broken at all. The patch does try to distinguish between scan and get and in the process fixes the issue reported (and would prevent future breakages as there is a testcase to protect it). I think this should be considered.\n\nThoughts?","created":"2013-01-20T18:14:09.432+0000"},{"body":"The following test failure seems to be related:\n{code}\nFailed tests: testMultipleTimestampRanges[0](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=NONE, compr=NONE, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n testMultipleTimestampRanges[1](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=ROW, compr=NONE, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n testMultipleTimestampRanges[2](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=ROWCOL, compr=NONE, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n testMultipleTimestampRanges[3](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=NONE, compr=GZ, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n testMultipleTimestampRanges[4](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=ROW, compr=GZ, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n testMultipleTimestampRanges[5](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=ROWCOL, compr=GZ, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n{code}","created":"2013-01-20T19:05:35.151+0000"},{"body":"Ok it seems like some tests construct the Scan object using the default constructor, and then access the setters of Scan to set the fields. The new field I added doesn't get set in those cases. So in this patch, I retain the check in isGetScan to check for those other fields (as in the original code).","created":"2013-01-20T19:23:14.476+0000"},{"body":"bq. On the point about disabling Get with empty row, I was also having the same opinion until I started digging into why Get with empty row key is broken at all.\n\nCan you say more? Would it be hard just throwing IllegalArgument if Get is passed a null or empty row?\n\nv3 patch seems to keep isScanCreatedFromGet beside a method called isGetScan. I'd think this'll confuse (and yeah, will think that not all Gets use the Get constructor.","created":"2013-01-20T22:53:00.225+0000"},{"body":"Well I just wanted to fix the problem reported by Benoit if possible :-) \n\nPlease have a look at the 3170-4.patch. That's the last iteration from me and passes all tests.","created":"2013-01-21T00:51:05.732+0000"},{"body":"v5 factors out a bit of common code into a method. Patch looks good to me. We could probably just throw an exception if you try and pass a Get a null row too... put the query out of its misery earlier rather than later. Was wondering about the test. We spin up a mini cluster instance to do the null key test. Should this test just be added to another test that has already put up a cluster?\n\nGood stuff lads.","created":"2013-01-21T17:10:58.793+0000"},{"body":"bq. We could probably just throw an exception if you try and pass a Get a null row too.. Was wondering about the test.\nOk [~stack], will review the patch from that point of view..","created":"2013-01-21T17:51:12.610+0000"},{"body":"How about adding the new test to:\n{code}\npublic class TestFromClientSide3 {\n{code}","created":"2013-01-21T18:28:46.108+0000"},{"body":"This puts in the test in TestFromClientSide3. On the handling of null row key, the RPC handler will throw a NPE immediately since the ProtoBufUtil tries to convert the PB request to a regular request and that will throw NPE. I think handling that should be outside the scope of this jira since there could be other such nulls in requests. [~stack] what do you think?","created":"2013-01-21T19:19:54.166+0000"},{"body":"Committed to trunk. Thanks lads.","created":"2013-01-21T20:12:51.442+0000"},{"body":"The patch for 0.94.","created":"2013-01-22T23:34:45.083+0000"},{"body":"Marking closed.","created":"2013-09-23T18:30:22.345+0000"}],"conversations":[{"body":"I'm no longer sure about the expected behavior when using an empty row key (e.g. a 0-byte long byte array). I assumed that this was a legitimate row key, just like having an empty column qualifier is allowed. But it seems that the RegionServer considers the empty row key to be whatever the first row key is.\n{code}\nVersion: 0.89.20100830, r0da2890b242584a8a5648d83532742ca7243346b, Sat Sep 18 15:30:09 PDT 2010\n\nhbase(main):001:0> scan 'tsdb-uid', {LIMIT => 1}\nROW COLUMN+CELL \n \\x00 column=id:metrics, timestamp=1288375187699, value=foo \n \\x00 column=id:tagk, timestamp=1287522021046, value=bar \n \\x00 column=id:tagv, timestamp=1288111387685, value=qux \n1 row(s) in 0.4610 seconds\n\nhbase(main):002:0> get 'tsdb-uid', ''\nCOLUMN CELL \n id:metrics timestamp=1288375187699, value=foo \n id:tagk timestamp=1287522021046, value=bar \n id:tagv timestamp=1288111387685, value=qux \n3 row(s) in 0.0910 seconds\n\nhbase(main):003:0> get 'tsdb-uid', \"\\000\"\nCOLUMN CELL \n id:metrics timestamp=1288375187699, value=foo \n id:tagk timestamp=1287522021046, value=bar \n id:tagv timestamp=1288111387685, value=qux \n3 row(s) in 0.0550 seconds\n{code}\n\nThis isn't a parsing problem with the command-line of the shell. I can reproduce this behavior both with plain Java code and with my asynchbase client.\n\nSince I don't actually have a row with an empty row key, I expected that the first {{get}} would return nothing.","from":"reporter","subject":"RegionServer confused about empty row keys"},{"body":"We are hitting this too, this is a really unexpected behaviour. Why getting empty key should return data of the first row in table? Reproduced in CDH3u4 (0.90.6):\n\n{code}\nhbase(main):005:0> create 'emptykey', {NAME=>'data', VERSION=>1}\n0 row(s) in 0.2070 seconds\n\nhbase(main):011:0> get 'emptykey', '' \nCOLUMN CELL \n0 row(s) in 0.0120 seconds\n\nhbase(main):006:0> put 'emptykey', 'a', 'data:a', '1234'\n0 row(s) in 0.1980 seconds\n\nhbase(main):007:0> put 'emptykey', 'b', 'data:b', '5678'\n0 row(s) in 0.0070 seconds\n\nhbase(main):008:0> scan 'emptykey' \nROW COLUMN+CELL \n a column=data:a, timestamp=1350989443394, value=1234 \n b column=data:b, timestamp=1350989450499, value=5678 \n2 row(s) in 0.0660 seconds\n\nhbase(main):009:0> get 'emptykey', ''\nCOLUMN CELL \n data:a timestamp=1350989443394, value=1234 \n1 row(s) in 0.0120 seconds\n{code}\n\nIt works the same way also using thrift.\n\nWe can even see, that empty key is supported in fact.\n{code}\nhbase(main):012:0> put 'emptykey', '', 'data:c', '90' \n0 row(s) in 0.0130 seconds\n\nhbase(main):013:0> get 'emptykey', '' \nCOLUMN CELL \n data:c timestamp=1350989869682, value=90 \n1 row(s) in 0.0120 seconds\n\nhbase(main):018:0> scan 'emptykey' \nROW COLUMN+CELL \n column=data:c, timestamp=1350989869682, value=90 \n a column=data:a, timestamp=1350989933922, value=1234 \n b column=data:b, timestamp=1350989937820, value=5678 \n3 row(s) in 0.0180 seconds\n{code}","from":"developer"},{"body":"Bad bug. Pulling into 0.96 so we fix it soon.","from":"developer"},{"body":"Attached is a patch. One change is to keep track of the fact that we are doing a 'get' as opposed to a 'scan'. When this was done, it broke the case of a legitimate 'get' with an empty rowkey. So there is a change to handle that case - basically, in the RegionScannerImpl.nextInternal I do special checks for empty start & end keys (in the new method isStopRowConsideringEmptyRows).. \n\nThanks to Enis for some useful discussions..","from":"developer"},{"body":"Let's see what hudson says.","from":"developer"},{"body":"{code}\n+public class TestGetEmptyRowKey {\n{code}\nThe new test would be a medium test, right ? Please add annotation.\n{code}\n+ //don't stop\n+ return false;\n{code}\nYou can save the comment such as the one above if you add javadoc for isStopRowConsideringEmptyRows().","from":"developer"},{"body":"It turns out that we need to keep the original behavior of Scan.isGetScan()\n\nPatch v2 introduced two new fields for RegionScannerImpl and made the logic in isStopRowConsideringEmptyRows() symmetrical.\n\nThe following tests passed:\n\n 1430 mt -Dtest=TestScanWithBloomError,TestStoreFile,TestGetEmptyRowKey\n 1433 mt -Dtest=TestHRegion,TestColumnSeeking,TestMultiColumnScanner","from":"developer"},{"body":"I realized that the boolean field, RegionScannerImpl.scanCreatedFromGet, is redundant with RegionScannerImpl.stopRowFromScan.","from":"developer"},{"body":"I ran TestConstraint twice locally and it passed.","from":"developer"},{"body":"bq. It turns out that we need to keep the original behavior of Scan.isGetScan()\n\nWhy is that? Seems odd having isGetScan and isScanCreatedFromGet. I'd think that a method named isGetScan would return true if this was a Scan made to service a Get?\n\nIt seems like isGetScan is trying to look at scanner specs to see if a Get scan rather than just at whether or not the Scan constructor that takes a Get was used.\n\nLooking at this, if you pass a Get a null or empty row, should we just throw an exception immediately and not even start the Scan going? Let a Get with an empty row be illegal?","from":"developer"},{"body":"bq. Seems odd having isGetScan and isScanCreatedFromGet.\nI agree. isGetScan() might have started with what we normally expect. Currently isGetScan() is used to serve other use cases.\n\nbq. should we just throw an exception immediately\nThat would simplify this issue a lot.\n\nLet'see what [~devaraj] says about the above proposal.","from":"developer"},{"body":"Okay this patch is a slight rework on my previous patch. It fixes the problem at hand I think and makes the distinction between scan and get clearer...\n\nOn the point about disabling Get with empty row, I was also having the same opinion until I started digging into why Get with empty row key is broken at all. The patch does try to distinguish between scan and get and in the process fixes the issue reported (and would prevent future breakages as there is a testcase to protect it). I think this should be considered.\n\nThoughts?","from":"developer"},{"body":"The following test failure seems to be related:\n{code}\nFailed tests: testMultipleTimestampRanges[0](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=NONE, compr=NONE, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n testMultipleTimestampRanges[1](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=ROW, compr=NONE, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n testMultipleTimestampRanges[2](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=ROWCOL, compr=NONE, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n testMultipleTimestampRanges[3](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=NONE, compr=GZ, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n testMultipleTimestampRanges[4](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=ROW, compr=GZ, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n testMultipleTimestampRanges[5](org.apache.hadoop.hbase.regionserver.TestSeekOptimizations): Expected and actual KV arrays differ at position 0: row1/myCF:qual0/2999/Put/vlen=0/ts=0 (length 3) vs. (length 0). Bloom=ROWCOL, compr=GZ, Scan: all columns, row=1, maxVersions=1, lazySeek=false\n{code}","from":"developer"},{"body":"Ok it seems like some tests construct the Scan object using the default constructor, and then access the setters of Scan to set the fields. The new field I added doesn't get set in those cases. So in this patch, I retain the check in isGetScan to check for those other fields (as in the original code).","from":"developer"},{"body":"bq. On the point about disabling Get with empty row, I was also having the same opinion until I started digging into why Get with empty row key is broken at all.\n\nCan you say more? Would it be hard just throwing IllegalArgument if Get is passed a null or empty row?\n\nv3 patch seems to keep isScanCreatedFromGet beside a method called isGetScan. I'd think this'll confuse (and yeah, will think that not all Gets use the Get constructor.","from":"developer"},{"body":"Well I just wanted to fix the problem reported by Benoit if possible :-) \n\nPlease have a look at the 3170-4.patch. That's the last iteration from me and passes all tests.","from":"developer"},{"body":"v5 factors out a bit of common code into a method. Patch looks good to me. We could probably just throw an exception if you try and pass a Get a null row too... put the query out of its misery earlier rather than later. Was wondering about the test. We spin up a mini cluster instance to do the null key test. Should this test just be added to another test that has already put up a cluster?\n\nGood stuff lads.","from":"developer"},{"body":"bq. We could probably just throw an exception if you try and pass a Get a null row too.. Was wondering about the test.\nOk [~stack], will review the patch from that point of view..","from":"developer"},{"body":"How about adding the new test to:\n{code}\npublic class TestFromClientSide3 {\n{code}","from":"developer"},{"body":"This puts in the test in TestFromClientSide3. On the handling of null row key, the RPC handler will throw a NPE immediately since the ProtoBufUtil tries to convert the PB request to a regular request and that will throw NPE. I think handling that should be outside the scope of this jira since there could be other such nulls in requests. [~stack] what do you think?","from":"developer"},{"body":"Committed to trunk. Thanks lads.","from":"developer"},{"body":"The patch for 0.94.","from":"developer"},{"body":"Marking closed.","from":"developer"}],"created":"2010-10-29T18:23:39.000+0000","description":"I'm no longer sure about the expected behavior when using an empty row key (e.g. a 0-byte long byte array). I assumed that this was a legitimate row key, just like having an empty column qualifier is allowed. But it seems that the RegionServer considers the empty row key to be whatever the first row key is.\n{code}\nVersion: 0.89.20100830, r0da2890b242584a8a5648d83532742ca7243346b, Sat Sep 18 15:30:09 PDT 2010\n\nhbase(main):001:0> scan 'tsdb-uid', {LIMIT => 1}\nROW COLUMN+CELL \n \\x00 column=id:metrics, timestamp=1288375187699, value=foo \n \\x00 column=id:tagk, timestamp=1287522021046, value=bar \n \\x00 column=id:tagv, timestamp=1288111387685, value=qux \n1 row(s) in 0.4610 seconds\n\nhbase(main):002:0> get 'tsdb-uid', ''\nCOLUMN CELL \n id:metrics timestamp=1288375187699, value=foo \n id:tagk timestamp=1287522021046, value=bar \n id:tagv timestamp=1288111387685, value=qux \n3 row(s) in 0.0910 seconds\n\nhbase(main):003:0> get 'tsdb-uid', \"\\000\"\nCOLUMN CELL \n id:metrics timestamp=1288375187699, value=foo \n id:tagk timestamp=1287522021046, value=bar \n id:tagv timestamp=1288111387685, value=qux \n3 row(s) in 0.0550 seconds\n{code}\n\nThis isn't a parsing problem with the command-line of the shell. I can reproduce this behavior both with plain Java code and with my asynchbase client.\n\nSince I don't actually have a row with an empty row key, I expected that the first {{get}} would return nothing.","issue_id":"12478686","key":"HBASE-3170","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2013-01-21T20:12:51.000+0000","role":"fixed_distractor","summary":"RegionServer confused about empty row keys"} {"case_id":"12495604","cluster":"DISTRACTOR-HBASE-3444","comments":[{"body":"Also results in java.lang.StringIndexOutOfBoundsException in Bytes.toBytesBinary if the byte[] ends in a '\\'\n\n{code}\nbyte[] bytes = new byte[]{'a','b','c','\\'};\nBytes.toBytesBinary(Bytes.toStringBinary(bytes));\n{code}","created":"2012-04-03T09:30:01.916+0000"},{"body":"Do you have a patch for us Simon? Thanks.","created":"2012-04-03T15:06:07.385+0000"},{"body":"I was looking into a way to fix this so that I can make it reversible. The main problem as I see is that the '\\' character is both a printable character in Bytes.toStringBinary and an indicator for hex characters in Bytes.toBytesBinary. Either Bytes.toStringBinary should not consider '\\' as printable OR Bytes.toBytesBinary should handle lone '\\' cases where it doesn't prefix a hex number.\n\nBelow are modifications to two methods that handle either case. Whichever one you choose, you don't have to choose the other.\n\n# *Modify Bytes.toStringBinary to not consider '\\' as a printable character*\nRemoves \\\\ from the if statement\n{code}\n public static String toStringBinary(final byte [] b, int off, int len) {\n StringBuilder result = new StringBuilder();\n try {\n String first = new String(b, off, len, \"ISO-8859-1\");\n for (int i = 0; i < first.length() ; ++i ) {\n int ch = first.charAt(i) & 0xFF;\n if ( (ch >= '0' && ch <= '9')\n || (ch >= 'A' && ch <= 'Z')\n || (ch >= 'a' && ch <= 'z')\n || \" `~!@#$%^&*()-_=+[]{}|;:'\\\",.<>/?\".indexOf(ch) >= 0 ) { // Change made here to remove '\\\\'\n result.append(first.charAt(i));\n } else {\n result.append(String.format(\"\\\\x%02X\", ch));\n }\n }\n } catch (UnsupportedEncodingException e) {\n System.err.println(\"ISO-8859-1 not supported?\");\n }\n return result.toString();\n }\n{code}\n\n# *Modify Bytes.toBytesBinary to consider standalone '\\'*\nThe problem is that the last '\\' is causing out of bounds issues. Just check to see if there is more to the array.\n{code}\n public static byte [] toBytesBinary(String in) {\n // this may be bigger than we need, but lets be safe.\n byte [] b = new byte[in.length()];\n int size = 0;\n for (int i = 0; i < in.length(); ++i) {\n char ch = in.charAt(i);\n if (ch == '\\\\') {\n // begin hex escape:\n char next = i+1 < in.length() ? in.charAt(i+1) : ch; // Change made here to check for array out of bounds\n if (next != 'x') {\n // invalid escape sequence, ignore this one.\n b[size++] = (byte)ch;\n continue;\n }\n // ok, take next 2 hex digits.\n char hd1 = in.charAt(i+2);\n char hd2 = in.charAt(i+3);\n\n // they need to be A-F0-9:\n if (!isHexDigit(hd1) ||\n !isHexDigit(hd2)) {\n // bogus escape code, ignore:\n continue;\n }\n // turn hex ASCII digit -> number\n byte d = (byte) ((toBinaryFromHex((byte)hd1) << 4) + toBinaryFromHex((byte)hd2));\n\n b[size++] = d;\n i += 3; // skip 3\n } else {\n b[size++] = (byte) ch;\n }\n }\n // resize:\n byte [] b2 = new byte[size];\n System.arraycopy(b, 0, b2, 0, size);\n return b2;\n }\n{code}\n\n# *Test case for both*\n{code}\n public void testToStringBinary_toBytesBinary_Reversable() throws Exception {\n String bytes = Bytes.toStringBinary(Bytes.toBytes(2.17));\n assertEquals(2.17, Bytes.toDouble(Bytes.toBytesBinary(bytes)), 0); \n }\n{code}","created":"2012-10-10T16:29:04.427+0000"},{"body":"Nice work Ken. Hopefully this will get picked up soon.","created":"2012-10-10T16:49:00.764+0000"},{"body":"I would choose \"Modify Bytes.toBytesBinary to consider standalone '\\'\"","created":"2012-10-10T17:12:45.752+0000"},{"body":"I agree that we should change toBytesBinary. Mind making a patch [~dallmkp]? Thanks for looking at this.","created":"2012-10-11T04:49:13.001+0000"},{"body":"Patch based on the code Ken posted.","created":"2012-10-11T18:07:52.530+0000"},{"body":"This patch looks the same as what we have in trunk:\n\n{noformat}\n if (ch == '\\\\' && in.length() > i+1 && in.charAt(i+1) == 'x') \n{noformat}\n\nDid I miss anything?","created":"2012-10-12T17:50:49.509+0000"},{"body":"@Jimmy:\nThe patch modifies the above line of code:\n{code}\n- if (ch == '\\\\' && in.length() > i+1 && in.charAt(i+1) == 'x') {\n{code}\nCan you take a second look ?","created":"2012-10-12T17:53:39.519+0000"},{"body":"To me, the patch moved the second part of the condition inside the block, so essentially, it is the same, right?","created":"2012-10-12T19:01:21.896+0000"},{"body":"No the patch is not the same as current code. See this boundary check:\n{code}\n+ char next = i+1 < in.length() ? in.charAt(i+1) : ch; // check for array out of bounds\n{code}\nWithout the above, in.charAt(i+1) would produce ArrayIndexOutOfBoundsException\n\nThanks","created":"2012-10-12T19:07:12.195+0000"},{"body":"What is the trunk issue [~jxiang]? Is this a dup then?","created":"2012-10-12T19:07:13.999+0000"},{"body":"@Stack, A duplicate of HBASE-6518, or the other way around?\n\n@Ted, jvm will check in.length() > i+1, at first, only if it is true, then it will check in.charAt(i+1)\nso there is no ArrayIndexOutOfBoundsException issue actually. Without the patch, the new unit test should be ok.\n","created":"2012-10-12T19:19:30.293+0000"},{"body":"Looks like HBASE-6518 has fixed this issue in trunk.\n\nThanks Jimmy.","created":"2012-10-12T19:29:03.968+0000"},{"body":"Here is test only from the patch. When I run it on trunk it passes. Let me commit the test as part of this issue.","created":"2012-10-12T19:47:41.101+0000"},{"body":"Committed the test only. It proves that HBASE-6518 works. Thanks for the patch Ken.","created":"2012-10-12T19:50:38.016+0000"},{"body":"Marking closed.","created":"2013-09-23T18:31:12.974+0000"}],"conversations":[{"body":"Bytes.toStringBinary() doesn't escape \\.\n\nOtherwise the transformation isn't reversible\n\nbyte[] a = {'\\', 'x' , '0', '0'}\n\nBytes.toBytesBinary(Bytes.toStringBinary(a)) won't be equal to a\n","from":"reporter","subject":"Test to prove Bytes.toBytesBinary and Bytes.toStringBinary() is reversible"},{"body":"Also results in java.lang.StringIndexOutOfBoundsException in Bytes.toBytesBinary if the byte[] ends in a '\\'\n\n{code}\nbyte[] bytes = new byte[]{'a','b','c','\\'};\nBytes.toBytesBinary(Bytes.toStringBinary(bytes));\n{code}","from":"developer"},{"body":"Do you have a patch for us Simon? Thanks.","from":"developer"},{"body":"I was looking into a way to fix this so that I can make it reversible. The main problem as I see is that the '\\' character is both a printable character in Bytes.toStringBinary and an indicator for hex characters in Bytes.toBytesBinary. Either Bytes.toStringBinary should not consider '\\' as printable OR Bytes.toBytesBinary should handle lone '\\' cases where it doesn't prefix a hex number.\n\nBelow are modifications to two methods that handle either case. Whichever one you choose, you don't have to choose the other.\n\n# *Modify Bytes.toStringBinary to not consider '\\' as a printable character*\nRemoves \\\\ from the if statement\n{code}\n public static String toStringBinary(final byte [] b, int off, int len) {\n StringBuilder result = new StringBuilder();\n try {\n String first = new String(b, off, len, \"ISO-8859-1\");\n for (int i = 0; i < first.length() ; ++i ) {\n int ch = first.charAt(i) & 0xFF;\n if ( (ch >= '0' && ch <= '9')\n || (ch >= 'A' && ch <= 'Z')\n || (ch >= 'a' && ch <= 'z')\n || \" `~!@#$%^&*()-_=+[]{}|;:'\\\",.<>/?\".indexOf(ch) >= 0 ) { // Change made here to remove '\\\\'\n result.append(first.charAt(i));\n } else {\n result.append(String.format(\"\\\\x%02X\", ch));\n }\n }\n } catch (UnsupportedEncodingException e) {\n System.err.println(\"ISO-8859-1 not supported?\");\n }\n return result.toString();\n }\n{code}\n\n# *Modify Bytes.toBytesBinary to consider standalone '\\'*\nThe problem is that the last '\\' is causing out of bounds issues. Just check to see if there is more to the array.\n{code}\n public static byte [] toBytesBinary(String in) {\n // this may be bigger than we need, but lets be safe.\n byte [] b = new byte[in.length()];\n int size = 0;\n for (int i = 0; i < in.length(); ++i) {\n char ch = in.charAt(i);\n if (ch == '\\\\') {\n // begin hex escape:\n char next = i+1 < in.length() ? in.charAt(i+1) : ch; // Change made here to check for array out of bounds\n if (next != 'x') {\n // invalid escape sequence, ignore this one.\n b[size++] = (byte)ch;\n continue;\n }\n // ok, take next 2 hex digits.\n char hd1 = in.charAt(i+2);\n char hd2 = in.charAt(i+3);\n\n // they need to be A-F0-9:\n if (!isHexDigit(hd1) ||\n !isHexDigit(hd2)) {\n // bogus escape code, ignore:\n continue;\n }\n // turn hex ASCII digit -> number\n byte d = (byte) ((toBinaryFromHex((byte)hd1) << 4) + toBinaryFromHex((byte)hd2));\n\n b[size++] = d;\n i += 3; // skip 3\n } else {\n b[size++] = (byte) ch;\n }\n }\n // resize:\n byte [] b2 = new byte[size];\n System.arraycopy(b, 0, b2, 0, size);\n return b2;\n }\n{code}\n\n# *Test case for both*\n{code}\n public void testToStringBinary_toBytesBinary_Reversable() throws Exception {\n String bytes = Bytes.toStringBinary(Bytes.toBytes(2.17));\n assertEquals(2.17, Bytes.toDouble(Bytes.toBytesBinary(bytes)), 0); \n }\n{code}","from":"developer"},{"body":"Nice work Ken. Hopefully this will get picked up soon.","from":"developer"},{"body":"I would choose \"Modify Bytes.toBytesBinary to consider standalone '\\'\"","from":"developer"},{"body":"I agree that we should change toBytesBinary. Mind making a patch [~dallmkp]? Thanks for looking at this.","from":"developer"},{"body":"Patch based on the code Ken posted.","from":"developer"},{"body":"This patch looks the same as what we have in trunk:\n\n{noformat}\n if (ch == '\\\\' && in.length() > i+1 && in.charAt(i+1) == 'x') \n{noformat}\n\nDid I miss anything?","from":"developer"},{"body":"@Jimmy:\nThe patch modifies the above line of code:\n{code}\n- if (ch == '\\\\' && in.length() > i+1 && in.charAt(i+1) == 'x') {\n{code}\nCan you take a second look ?","from":"developer"},{"body":"To me, the patch moved the second part of the condition inside the block, so essentially, it is the same, right?","from":"developer"},{"body":"No the patch is not the same as current code. See this boundary check:\n{code}\n+ char next = i+1 < in.length() ? in.charAt(i+1) : ch; // check for array out of bounds\n{code}\nWithout the above, in.charAt(i+1) would produce ArrayIndexOutOfBoundsException\n\nThanks","from":"developer"},{"body":"What is the trunk issue [~jxiang]? Is this a dup then?","from":"developer"},{"body":"@Stack, A duplicate of HBASE-6518, or the other way around?\n\n@Ted, jvm will check in.length() > i+1, at first, only if it is true, then it will check in.charAt(i+1)\nso there is no ArrayIndexOutOfBoundsException issue actually. Without the patch, the new unit test should be ok.\n","from":"developer"},{"body":"Looks like HBASE-6518 has fixed this issue in trunk.\n\nThanks Jimmy.","from":"developer"},{"body":"Here is test only from the patch. When I run it on trunk it passes. Let me commit the test as part of this issue.","from":"developer"},{"body":"Committed the test only. It proves that HBASE-6518 works. Thanks for the patch Ken.","from":"developer"},{"body":"Marking closed.","from":"developer"}],"created":"2011-01-14T17:14:53.000+0000","description":"Bytes.toStringBinary() doesn't escape \\.\n\nOtherwise the transformation isn't reversible\n\nbyte[] a = {'\\', 'x' , '0', '0'}\n\nBytes.toBytesBinary(Bytes.toStringBinary(a)) won't be equal to a\n","issue_id":"12495604","key":"HBASE-3444","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-10-12T19:50:38.000+0000","role":"fixed_distractor","summary":"Test to prove Bytes.toBytesBinary and Bytes.toStringBinary() is reversible"} {"case_id":"12497941","cluster":"DISTRACTOR-HBASE-3513","comments":[{"body":"this patch is a little WIP, it's based on 0.90 with this commit cherry picked in:\n\nhttps://github.com/stumbleupon/hbase/commit/67a5b46d069f0e7660d2de998c7b311103708fb4\n","created":"2011-02-08T01:43:29.411+0000"},{"body":"That patch looks fine. I don't see a pom edit though included in your patch.","created":"2011-02-08T04:56:46.311+0000"},{"body":"Sorry, I meant, the patch looks great. Its cool getting ICV and parallelGet into our thrift interface.","created":"2011-02-08T04:57:21.432+0000"},{"body":"This is a duplicate of HBASE-3117.\n\nAlso Thrift 0.5 isn't available in any official repository. The issue is still unresolved (THRIFT-363).","created":"2011-02-11T14:29:47.479+0000"},{"body":"And now there is 0.6.0 which we should use I presume?","created":"2011-02-13T12:57:47.838+0000"},{"body":"thrift 0.5.0 is in the apache maven official repo... there was no\n0.6.0 when i went looking.\n\n","created":"2011-02-13T20:44:57.568+0000"},{"body":"Ryan, I'm pretty sure Thrift is *not* in any repository. As I said the issue is still unresolved and they are looking for help getting it into the repository.\n\nhttps://repository.apache.org/index.html#nexus-search;quick~thrift","created":"2011-02-14T13:19:37.539+0000"},{"body":"Try this search:\n\nhttps://repository.apache.org/index.html#nexus-search;quick~libthrift\n","created":"2011-02-14T23:43:03.079+0000"},{"body":"I prefer 0.6.0 rather than 0.5.0, because THRIFT-923 supports asynchronous client in C++.\n\n* https://issues.apache.org/jira/browse/THRIFT-923\n* https://github.com/apache/thrift/blob/trunk/CHANGES\n\nFYI","created":"2011-03-02T20:19:56.955+0000"},{"body":"a patch against current trunk.","created":"2011-03-03T02:07:35.135+0000"},{"body":"Kazuki, 0.6.0 doesnt exist yet (in maven) so we cant depend on it.\n\nThe on the wire serialization should be the same between 0.5.0 and 0.6.0, so you can use whatever version of C++ client you wish. This is about the _hbase serverside_ thrift version.","created":"2011-03-03T02:14:54.463+0000"},{"body":"Hi, ryan. Thanks for the comment. \n\n> hbase serverside thrift version.\n\nThat's right. I'll test the wire compatibility in my environment.\n\nBig thanks for your patch anyway!","created":"2011-03-03T11:19:11.225+0000"},{"body":"+1 on patch (if tests pass). This breaks old thrift clients? Needs to go into INCOMPATIBLE CHANGES section. Add a release note on resolve?","created":"2011-03-05T00:20:38.273+0000"},{"body":"it is wire compatible with older clients, tested with an older SU client. There is no actual incompatibility!","created":"2011-03-05T00:23:27.760+0000"},{"body":"Oh, I never saw your answer. Sorry Ryan I never searched for _libthrift_, thanks for pointing it out Stack. As this is a custom release by the Hadoop folks are we sure this is not in any way customized by them?\n\nI haven't looked at the Maven stuff in a while but by being in a custom group we might get a different version by accident as a transitive dependency if any other project decides to do the same as Hadoop.","created":"2011-03-10T09:11:48.563+0000"},{"body":"I didn't notice the thrift 0.5.0 a hadoop built thing.\n\nThanks for the warning @LarsF.","created":"2011-03-11T05:09:40.936+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T12:42:12.135+0000"}],"conversations":[{"body":"We should upgrade our thrift to 0.5.0, it is the latest and greatest and is in apache maven repo.\n\nDoing some testing with a thrift 0.5.0 server, and an older pre-release php client shows the two are on-wire compatible.\n\nGiven that the upgrade is entirely on the server side, and has no wire-impact this should be a relatively low-impact change.","from":"reporter","subject":"upgrade thrift to 0.5.0 and use mvn version"},{"body":"this patch is a little WIP, it's based on 0.90 with this commit cherry picked in:\n\nhttps://github.com/stumbleupon/hbase/commit/67a5b46d069f0e7660d2de998c7b311103708fb4\n","from":"developer"},{"body":"That patch looks fine. I don't see a pom edit though included in your patch.","from":"developer"},{"body":"Sorry, I meant, the patch looks great. Its cool getting ICV and parallelGet into our thrift interface.","from":"developer"},{"body":"This is a duplicate of HBASE-3117.\n\nAlso Thrift 0.5 isn't available in any official repository. The issue is still unresolved (THRIFT-363).","from":"developer"},{"body":"And now there is 0.6.0 which we should use I presume?","from":"developer"},{"body":"thrift 0.5.0 is in the apache maven official repo... there was no\n0.6.0 when i went looking.\n\n","from":"developer"},{"body":"Ryan, I'm pretty sure Thrift is *not* in any repository. As I said the issue is still unresolved and they are looking for help getting it into the repository.\n\nhttps://repository.apache.org/index.html#nexus-search;quick~thrift","from":"developer"},{"body":"Try this search:\n\nhttps://repository.apache.org/index.html#nexus-search;quick~libthrift\n","from":"developer"},{"body":"I prefer 0.6.0 rather than 0.5.0, because THRIFT-923 supports asynchronous client in C++.\n\n* https://issues.apache.org/jira/browse/THRIFT-923\n* https://github.com/apache/thrift/blob/trunk/CHANGES\n\nFYI","from":"developer"},{"body":"a patch against current trunk.","from":"developer"},{"body":"Kazuki, 0.6.0 doesnt exist yet (in maven) so we cant depend on it.\n\nThe on the wire serialization should be the same between 0.5.0 and 0.6.0, so you can use whatever version of C++ client you wish. This is about the _hbase serverside_ thrift version.","from":"developer"},{"body":"Hi, ryan. Thanks for the comment. \n\n> hbase serverside thrift version.\n\nThat's right. I'll test the wire compatibility in my environment.\n\nBig thanks for your patch anyway!","from":"developer"},{"body":"+1 on patch (if tests pass). This breaks old thrift clients? Needs to go into INCOMPATIBLE CHANGES section. Add a release note on resolve?","from":"developer"},{"body":"it is wire compatible with older clients, tested with an older SU client. There is no actual incompatibility!","from":"developer"},{"body":"Oh, I never saw your answer. Sorry Ryan I never searched for _libthrift_, thanks for pointing it out Stack. As this is a custom release by the Hadoop folks are we sure this is not in any way customized by them?\n\nI haven't looked at the Maven stuff in a while but by being in a custom group we might get a different version by accident as a transitive dependency if any other project decides to do the same as Hadoop.","from":"developer"},{"body":"I didn't notice the thrift 0.5.0 a hadoop built thing.\n\nThanks for the warning @LarsF.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2011-02-08T01:38:44.000+0000","description":"We should upgrade our thrift to 0.5.0, it is the latest and greatest and is in apache maven repo.\n\nDoing some testing with a thrift 0.5.0 server, and an older pre-release php client shows the two are on-wire compatible.\n\nGiven that the upgrade is entirely on the server side, and has no wire-impact this should be a relatively low-impact change.","issue_id":"12497941","key":"HBASE-3513","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2011-03-06T03:10:39.000+0000","role":"fixed_distractor","summary":"upgrade thrift to 0.5.0 and use mvn version"} {"case_id":"12501056","cluster":"DISTRACTOR-HBASE-3617","comments":[{"body":"We do all this stuff in the balance:\n\n{code}\n try {\n // TODO: We should consider making this look more like it does for the\n // region open where we catch all throwables and never abort\n if (serverManager.sendRegionClose(server, state.getRegion())) {\n LOG.debug(\"Sent CLOSE to \" + server + \" for region \" +\n region.getRegionNameAsString());\n return;\n }\n // This never happens. Currently regionserver close always return true.\n LOG.debug(\"Server \" + server + \" region CLOSE RPC returned false for \" +\n region.getEncodedName());\n } catch (NotServingRegionException nsre) {\n LOG.info(\"Server \" + server + \" returned \" + nsre + \" for \" +\n region.getEncodedName());\n // Presume that master has stale data. Presume remote side just split.\n // Presume that the split message when it comes in will fix up the master's\n // in memory cluster state.\n return;\n } catch (ConnectException e) {\n LOG.info(\"Failed connect to \" + server + \", message=\" + e.getMessage() +\n \", region=\" + region.getEncodedName());\n // Presume that regionserver just failed and we haven't got expired\n // server from zk yet. Let expired server deal with clean up.\n } catch (java.net.SocketTimeoutException e) {\n LOG.info(\"Server \" + server + \" returned \" + e.getMessage() + \" for \" +\n region.getEncodedName());\n // Presume retry or server will expire.\n } catch (EOFException e) {\n LOG.info(\"Server \" + server + \" returned \" + e.getMessage() + \" for \" +\n region.getEncodedName());\n // Presume retry or server will expire.\n } catch (RemoteException re) {\n IOException ioe = re.unwrapRemoteException();\n if (ioe instanceof NotServingRegionException) {\n // Failed to close, so pass through and reassign\n LOG.debug(\"Server \" + server + \" returned \" + ioe + \" for \" +\n region.getEncodedName());\n } else if (ioe instanceof EOFException) {\n // Failed to close, so pass through and reassign\n LOG.debug(\"Server \" + server + \" returned \" + ioe + \" for \" +\n region.getEncodedName());\n } else {\n this.master.abort(\"Remote unexpected exception\", ioe);\n }\n } catch (Throwable t) {\n // For now call abort if unexpected exception -- radical, but will get\n // fellas attention. St.Ack 20101012\n this.master.abort(\"Remote unexpected exception\", t);\n }\n }\n{code}","created":"2011-03-10T19:29:01.481+0000"},{"body":"We should catch IOException.","created":"2011-03-10T19:47:01.470+0000"},{"body":"J-D just saw a IOException where message is 'Connection Reset'... an exception we should be catching.\n\nLets do what Ted suggests, catch all IOEs.","created":"2011-03-11T00:26:16.770+0000"},{"body":"Catch all IOExceptions in unassign()","created":"2011-03-12T06:20:12.484+0000"},{"body":"Catch IOException in unassign()","created":"2011-03-12T08:06:35.081+0000"},{"body":"Your patch has other change pollution Ted but thats OK... just watch out for it next time.\n\nLet me try out this change. I think its the right thing to do. We might have to be even more extreme. See how the assign method catches Throwable, not just IOEs. We need to be a little careful here. This is pretty big change, especially on a point release.","created":"2011-03-12T19:56:51.009+0000"},{"body":"Remove unrelated changes in other files. Pardon me.","created":"2011-03-12T21:31:12.051+0000"},{"body":"Second attempt after discussing with J-D","created":"2011-03-22T01:49:02.735+0000"},{"body":"Unwrap RemoteException if necessary","created":"2011-03-22T01:58:47.167+0000"},{"body":"Here is what I applied; its same as the wrapper around assign (this adds same wrapper around unassign).","created":"2011-03-22T22:22:08.064+0000"},{"body":"Applied trunk and branch. Thanks for the patch Ted.","created":"2011-03-22T22:22:30.546+0000"},{"body":"Bringing into 0.90.4. This patch was not applied to the branch as it says above (TRUNK CHANGES.txt has it in the 0.90.2 section).","created":"2011-05-12T18:01:25.696+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T12:43:12.986+0000"}],"conversations":[{"body":"Via Tatsuya up on the list:\n\n{code}\n2011-03-10 07:48:39,192 FATAL org.apache.hadoop.hbase.master.HMaster:\nRemote unexpected exception\njava.net.NoRouteToHostException: No route to host\n at sun.nio.ch.SocketChannelImpl.checkConnect(Native Method)\n at\nsun.nio.ch.SocketChannelImpl.finishConnect(SocketChannelImpl.java:567)\n at\norg.apache.hadoop.net.SocketIOWithTimeout.connect(SocketIOWithTimeout.java:\n206)\n at org.apache.hadoop.net.NetUtils.connect(NetUtils.java:408)\n at org.apache.hadoop.hbase.ipc.HBaseClient\n$Connection.setupIOstreams(HBaseClient.java:328)\n at\norg.apache.hadoop.hbase.ipc.HBaseClient.getConnection(HBaseClient.java:\n883)\n at\norg.apache.hadoop.hbase.ipc.HBaseClient.call(HBaseClient.java:750)\n at org.apache.hadoop.hbase.ipc.HBaseRPC\n$Invoker.invoke(HBaseRPC.java:257)\n at $Proxy6.closeRegion(Unknown Source)\n at\norg.apache.hadoop.hbase.master.ServerManager.sendRegionClose(ServerManager.java:\n589)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.unassign(AssignmentManager.java:\n1093)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.unassign(AssignmentManager.java:\n1040)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.balance(AssignmentManager.java:\n1831)\n at org.apache.hadoop.hbase.master.HMaster.balance(HMaster.java:\n692)\n at org.apache.hadoop.hbase.master.HMaster$1.chore(HMaster.java:\n583)\n at org.apache.hadoop.hbase.Chore.run(Chore.java:66)\n2011-03-10 07:48:39,192 INFO org.apache.hadoop.hbase.master.HMaster:\nAborting\n2011-03-10 07:48:39,192 INFO org.apache.hadoop.hbase.master.HMaster:\nbalance hri=SpecialObject_Speed_Test,,\n1299710751983.f0e5544339870a510c338b3029979d3e.,\nsrc=ap13.secur2,60020,1299710609447,\ndest=ap12.secur2,60020,1299710609148\n2011-03-10 07:48:39,192 DEBUG\norg.apache.hadoop.hbase.master.AssignmentManager: Starting\nunassignment of region SpecialObject_Speed_Test,,\n1299710751983.f0e5544339870a510c338b3029979d3e. (offlining)\n2011-03-10 07:48:39,852 DEBUG org.apache.hadoop.hbase.master.HMaster:\nStopping service threads\n2011-03-10 07:48:39,852 INFO org.apache.hadoop.ipc.HBaseServer:\nStopping server on 60000\n2011-03-10 07:48:39,852 FATAL org.apache.hadoop.hbase.master.HMaster:\nRemote unexpected exception\njava.io.InterruptedIOException: Interruped while waiting for IO on\nchannel java.nio.channels.SocketChannel[connection-pending remote=/\n10.X.X.18:60020]. 19340 millis timeout left.\n at org.apache.hadoop.net.SocketIOWithTimeout\n$SelectorPool.select(SocketIOWithTimeout.java:349)\n at\norg.apache.hadoop.net.SocketIOWithTimeout.connect(SocketIOWithTimeout.java:\n203)\n at org.apache.hadoop.net.NetUtils.connect(NetUtils.java:408)\n at org.apache.hadoop.hbase.ipc.HBaseClient\n$Connection.setupIOstreams(HBaseClient.java:328)\n at\norg.apache.hadoop.hbase.ipc.HBaseClient.getConnection(HBaseClient.java:\n883)\n at\norg.apache.hadoop.hbase.ipc.HBaseClient.call(HBaseClient.java:750)\n at org.apache.hadoop.hbase.ipc.HBaseRPC\n$Invoker.invoke(HBaseRPC.java:257)\n at $Proxy6.closeRegion(Unknown Source)\n at\norg.apache.hadoop.hbase.master.ServerManager.sendRegionClose(ServerManager.java:\n589)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.unassign(AssignmentManager.java:\n1093)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.unassign(AssignmentManager.java:\n1040)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.balance(AssignmentManager.java:\n1831)\n at org.apache.hadoop.hbase.master.HMaster.balance(HMaster.java:\n692)\n at org.apache.hadoop.hbase.master.HMaster$1.chore(HMaster.java:\n583)\n at org.apache.hadoop.hbase.Chore.run(Chore.java:66)\n2011-03-10 07:48:39,852 INFO org.apache.hadoop.hbase.master.HMaster:\nAborting\n{code}","from":"reporter","subject":"NoRouteToHostException during balancing will cause Master abort"},{"body":"We do all this stuff in the balance:\n\n{code}\n try {\n // TODO: We should consider making this look more like it does for the\n // region open where we catch all throwables and never abort\n if (serverManager.sendRegionClose(server, state.getRegion())) {\n LOG.debug(\"Sent CLOSE to \" + server + \" for region \" +\n region.getRegionNameAsString());\n return;\n }\n // This never happens. Currently regionserver close always return true.\n LOG.debug(\"Server \" + server + \" region CLOSE RPC returned false for \" +\n region.getEncodedName());\n } catch (NotServingRegionException nsre) {\n LOG.info(\"Server \" + server + \" returned \" + nsre + \" for \" +\n region.getEncodedName());\n // Presume that master has stale data. Presume remote side just split.\n // Presume that the split message when it comes in will fix up the master's\n // in memory cluster state.\n return;\n } catch (ConnectException e) {\n LOG.info(\"Failed connect to \" + server + \", message=\" + e.getMessage() +\n \", region=\" + region.getEncodedName());\n // Presume that regionserver just failed and we haven't got expired\n // server from zk yet. Let expired server deal with clean up.\n } catch (java.net.SocketTimeoutException e) {\n LOG.info(\"Server \" + server + \" returned \" + e.getMessage() + \" for \" +\n region.getEncodedName());\n // Presume retry or server will expire.\n } catch (EOFException e) {\n LOG.info(\"Server \" + server + \" returned \" + e.getMessage() + \" for \" +\n region.getEncodedName());\n // Presume retry or server will expire.\n } catch (RemoteException re) {\n IOException ioe = re.unwrapRemoteException();\n if (ioe instanceof NotServingRegionException) {\n // Failed to close, so pass through and reassign\n LOG.debug(\"Server \" + server + \" returned \" + ioe + \" for \" +\n region.getEncodedName());\n } else if (ioe instanceof EOFException) {\n // Failed to close, so pass through and reassign\n LOG.debug(\"Server \" + server + \" returned \" + ioe + \" for \" +\n region.getEncodedName());\n } else {\n this.master.abort(\"Remote unexpected exception\", ioe);\n }\n } catch (Throwable t) {\n // For now call abort if unexpected exception -- radical, but will get\n // fellas attention. St.Ack 20101012\n this.master.abort(\"Remote unexpected exception\", t);\n }\n }\n{code}","from":"developer"},{"body":"We should catch IOException.","from":"developer"},{"body":"J-D just saw a IOException where message is 'Connection Reset'... an exception we should be catching.\n\nLets do what Ted suggests, catch all IOEs.","from":"developer"},{"body":"Catch all IOExceptions in unassign()","from":"developer"},{"body":"Catch IOException in unassign()","from":"developer"},{"body":"Your patch has other change pollution Ted but thats OK... just watch out for it next time.\n\nLet me try out this change. I think its the right thing to do. We might have to be even more extreme. See how the assign method catches Throwable, not just IOEs. We need to be a little careful here. This is pretty big change, especially on a point release.","from":"developer"},{"body":"Remove unrelated changes in other files. Pardon me.","from":"developer"},{"body":"Second attempt after discussing with J-D","from":"developer"},{"body":"Unwrap RemoteException if necessary","from":"developer"},{"body":"Here is what I applied; its same as the wrapper around assign (this adds same wrapper around unassign).","from":"developer"},{"body":"Applied trunk and branch. Thanks for the patch Ted.","from":"developer"},{"body":"Bringing into 0.90.4. This patch was not applied to the branch as it says above (TRUNK CHANGES.txt has it in the 0.90.2 section).","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2011-03-10T19:28:13.000+0000","description":"Via Tatsuya up on the list:\n\n{code}\n2011-03-10 07:48:39,192 FATAL org.apache.hadoop.hbase.master.HMaster:\nRemote unexpected exception\njava.net.NoRouteToHostException: No route to host\n at sun.nio.ch.SocketChannelImpl.checkConnect(Native Method)\n at\nsun.nio.ch.SocketChannelImpl.finishConnect(SocketChannelImpl.java:567)\n at\norg.apache.hadoop.net.SocketIOWithTimeout.connect(SocketIOWithTimeout.java:\n206)\n at org.apache.hadoop.net.NetUtils.connect(NetUtils.java:408)\n at org.apache.hadoop.hbase.ipc.HBaseClient\n$Connection.setupIOstreams(HBaseClient.java:328)\n at\norg.apache.hadoop.hbase.ipc.HBaseClient.getConnection(HBaseClient.java:\n883)\n at\norg.apache.hadoop.hbase.ipc.HBaseClient.call(HBaseClient.java:750)\n at org.apache.hadoop.hbase.ipc.HBaseRPC\n$Invoker.invoke(HBaseRPC.java:257)\n at $Proxy6.closeRegion(Unknown Source)\n at\norg.apache.hadoop.hbase.master.ServerManager.sendRegionClose(ServerManager.java:\n589)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.unassign(AssignmentManager.java:\n1093)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.unassign(AssignmentManager.java:\n1040)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.balance(AssignmentManager.java:\n1831)\n at org.apache.hadoop.hbase.master.HMaster.balance(HMaster.java:\n692)\n at org.apache.hadoop.hbase.master.HMaster$1.chore(HMaster.java:\n583)\n at org.apache.hadoop.hbase.Chore.run(Chore.java:66)\n2011-03-10 07:48:39,192 INFO org.apache.hadoop.hbase.master.HMaster:\nAborting\n2011-03-10 07:48:39,192 INFO org.apache.hadoop.hbase.master.HMaster:\nbalance hri=SpecialObject_Speed_Test,,\n1299710751983.f0e5544339870a510c338b3029979d3e.,\nsrc=ap13.secur2,60020,1299710609447,\ndest=ap12.secur2,60020,1299710609148\n2011-03-10 07:48:39,192 DEBUG\norg.apache.hadoop.hbase.master.AssignmentManager: Starting\nunassignment of region SpecialObject_Speed_Test,,\n1299710751983.f0e5544339870a510c338b3029979d3e. (offlining)\n2011-03-10 07:48:39,852 DEBUG org.apache.hadoop.hbase.master.HMaster:\nStopping service threads\n2011-03-10 07:48:39,852 INFO org.apache.hadoop.ipc.HBaseServer:\nStopping server on 60000\n2011-03-10 07:48:39,852 FATAL org.apache.hadoop.hbase.master.HMaster:\nRemote unexpected exception\njava.io.InterruptedIOException: Interruped while waiting for IO on\nchannel java.nio.channels.SocketChannel[connection-pending remote=/\n10.X.X.18:60020]. 19340 millis timeout left.\n at org.apache.hadoop.net.SocketIOWithTimeout\n$SelectorPool.select(SocketIOWithTimeout.java:349)\n at\norg.apache.hadoop.net.SocketIOWithTimeout.connect(SocketIOWithTimeout.java:\n203)\n at org.apache.hadoop.net.NetUtils.connect(NetUtils.java:408)\n at org.apache.hadoop.hbase.ipc.HBaseClient\n$Connection.setupIOstreams(HBaseClient.java:328)\n at\norg.apache.hadoop.hbase.ipc.HBaseClient.getConnection(HBaseClient.java:\n883)\n at\norg.apache.hadoop.hbase.ipc.HBaseClient.call(HBaseClient.java:750)\n at org.apache.hadoop.hbase.ipc.HBaseRPC\n$Invoker.invoke(HBaseRPC.java:257)\n at $Proxy6.closeRegion(Unknown Source)\n at\norg.apache.hadoop.hbase.master.ServerManager.sendRegionClose(ServerManager.java:\n589)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.unassign(AssignmentManager.java:\n1093)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.unassign(AssignmentManager.java:\n1040)\n at\norg.apache.hadoop.hbase.master.AssignmentManager.balance(AssignmentManager.java:\n1831)\n at org.apache.hadoop.hbase.master.HMaster.balance(HMaster.java:\n692)\n at org.apache.hadoop.hbase.master.HMaster$1.chore(HMaster.java:\n583)\n at org.apache.hadoop.hbase.Chore.run(Chore.java:66)\n2011-03-10 07:48:39,852 INFO org.apache.hadoop.hbase.master.HMaster:\nAborting\n{code}","issue_id":"12501056","key":"HBASE-3617","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2011-03-22T22:22:30.000+0000","role":"fixed_distractor","summary":"NoRouteToHostException during balancing will cause Master abort"} {"case_id":"12501811","cluster":"DISTRACTOR-HBASE-3671","comments":[{"body":"Small fix. Please review Jon/JD","created":"2011-03-18T17:44:22.866+0000"},{"body":"Minor comment:\n{code}\n+ LOG.warn(\"Skipping the onining of \" + regionInfo.getRegionNameAsString() +\n+ \" because regions is NOT in RIT -- presuming this is because it SPLIT\");\n{code}\nshould be:\n{code}\n+ LOG.warn(\"Skipping the onlining of \" + regionInfo.getRegionNameAsString() +\n+ \" because region is NOT in RIT -- presuming this is because it SPLIT\");\n{code}\n","created":"2011-03-18T18:39:07.143+0000"},{"body":"+1 after the typo Ted reported and if it passes all the tests.","created":"2011-03-18T18:44:20.573+0000"},{"body":"+1 too, thanks for the quick turnaround guys!","created":"2011-03-18T18:45:07.485+0000"},{"body":"Thanks for reviews Ted and J-D. Will commit soon w/ Ted suggestion.","created":"2011-03-18T23:50:41.072+0000"},{"body":"Thanks reviews B, Ted, and J-D.","created":"2011-03-19T20:22:00.097+0000"},{"body":"Looks like there is inconsistency between javadoc and log message.\nFor isRegionInTransition():\n{code}\n * @return Returns null if passed region is not in transition else the current\n{code}\nBut the change in OpenedRegionHandler says:\n{code}\n if (this.assignmentManager.isRegionInTransition(regionInfo) == null) {\n this.assignmentManager.regionOnline(regionInfo, serverInfo);\n } else {\n LOG.warn(\"Skipping the onlining of \" + regionInfo.getRegionNameAsString() +\n \" because regions is NOT in RIT -- presuming this is because it SPLIT\");\n }\n{code}\n\nBTW TestMultipleTimestamps starts to fail after this change.","created":"2011-03-20T00:59:18.848+0000"},{"body":"@Ted You are right. Thanks for spotting this. Let me fix.","created":"2011-03-20T01:50:23.390+0000"}],"conversations":[{"body":"This issue is about adding a workaround to 0.90 until we get proper fix in 0.92 (HBASE-3559).\n\nHere is the sequence of events:\n\n1. We start to process OPENED region event.\n2. We receive a SPLIT of this region report.\n3. SPLIT processing offline the region and onlines daughters.\n4. Metascanner runs and clears out the region from .META. deleting it\n5. The OPENED handler runs. Marks the region online in Master memory.\n6. Balancer runs. Trys to balance a region that has been deleted.\n\nLoops for ever.\n\nHere is excerpt from logs. It happened during startup, lots going on. Could happen on regionserver crash I suppose, maybe, but we're susceptible during cluster start:\n\n{code}\n# We assign the region\n2011-03-16 15:18:29,053 DEBUG org.apache.hadoop.hbase.zookeeper.ZKAssign: master:60000-0x22e286f0b9c98f1 Async create of unassigned node for 3516b74d0c9d4458c2f2f715249e3f78 with OFFLINE state\n...\n2011-03-16 15:18:32,298 DEBUG org.apache.hadoop.hbase.master.AssignmentManager$CreateUnassignedAsyncCallback: rs=tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. state=OFFLINE, ts=1300313909053, server=sv4borg39,60020,1300313564807\n...\n2011-03-16 15:18:32,732 DEBUG org.apache.hadoop.hbase.master.AssignmentManager$ExistsUnassignedAsyncCallback: rs=tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. state=OFFLINE, ts=1300313909053\n...\n2011-03-16 15:23:02,114 DEBUG org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher: master:60000-0x22e286f0b9c98f1 Received ZooKeeper Event, type=NodeDataChanged, state=SyncConnected, path=/prodjobs/unassigned/3516b74d0c9d4458c2f2f715249e3f78\n...\n2011-03-16 15:23:02,183 DEBUG org.apache.hadoop.hbase.zookeeper.ZKUtil: master:60000-0x22e286f0b9c98f1 Retrieved 127 byte(s) of data from znode /prodjobs/unassigned/3516b74d0c9d4458c2f2f715249e3f78 and set watcher; region=tsdb,^@^D2McZ@^@^@^A^@^@G^@^@^L^@^@f^@^@^U^@^@�^@^@(^@^C^G,1299401073466.3516b74d0c9d4458c2f2f715249e3f78., server=sv4borg39,60020,1300313564807, state=RS_ZK_REGION_OPENED\n2011-03-16 15:23:02,183 DEBUG org.apache.hadoop.hbase.master.AssignmentManager: Handling transition=RS_ZK_REGION_OPENED, server=sv4borg39,60020,1300313564807, region=3516b74d0c9d4458c2f2f715249e3f78\n\n# At this point we've queued an Excecutor to run to process the OPENED event. Now in comes the SPLIT.\n2011-03-16 15:23:18,199 INFO org.apache.hadoop.hbase.master.ServerManager: Received REGION_SPLIT: tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78.: Daughters; tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1300314189812.74c51400bb8dfa127fadfd11a04d72f2., tsdb,\\x00\\x042MmD\\x88\\x00\\x00\\x01\\x00\\x00S\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x029\\x00\\x00(\\x00\\x03\\x03,1300314189812.87b061739a11d0f9d02acfb92ef961a2. from sv4borg39,60020,1300313564807\n2011-03-16 15:23:18,870 WARN org.apache.hadoop.hbase.master.AssignmentManager: Split report has RIT node (shouldnt have one): REGION => {NAME => 'tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78.', STARTKEY => '\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07', ENDKEY => '\\x00\\x043L\\xE7\\xF50\\x00\\x00\\x01\\x00\\x00I\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x0E\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x02u', ENCODED => 3516b74d0c9d4458c2f2f715249e3f78, TABLE => {{NAME => 'tsdb', FAMILIES => [{NAME => 't', BLOOMFILTER => 'NONE', REPLICATION_SCOPE => '0', VERSIONS => '3', COMPRESSION => 'LZO', TTL => '2147483647', BLOCKSIZE => '65536', IN_MEMORY => 'false', BLOCKCACHE => 'true'}]}} node: region=tsdb,^@^D2McZ@^@^@^A^@^@G^@^@^L^@^@f^@^@^U^@^@�^@^@(^@^C^G,1299401073466.3516b74d0c9d4458c2f2f715249e3f78., server=sv4borg39,60020,1300313564807, state=RS_ZK_REGION_OPENED\n\n# Now metascanner runs and actually removes the parent region, deleting it all\n\n2011-03-16 15:28:34,352 INFO org.apache.hadoop.hbase.catalog.MetaEditor: Deleted daughter reference tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1300314189812.74c51400bb8dfa127fadfd11a04d72f2., qualifier=splitA, from parent tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78.\n2011-03-16 15:28:34,356 INFO org.apache.hadoop.hbase.catalog.MetaEditor: Deleted daughter reference tsdb,\\x00\\x042MmD\\x88\\x00\\x00\\x01\\x00\\x00S\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x029\\x00\\x00(\\x00\\x03\\x03,1300314189812.87b061739a11d0f9d02acfb92ef961a2., qualifier=splitB, from parent tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78.\n2011-03-16 15:28:34,356 DEBUG org.apache.hadoop.hbase.master.CatalogJanitor: Deleting region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. because daughter splits no longer hold references\n2011-03-16 15:28:34,356 DEBUG org.apache.hadoop.hbase.master.CatalogJanitor: Deleting region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. because daughter splits no longer hold references\n\n2011-03-16 15:28:34,444 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: DELETING region hdfs://sv4borg29:9000/hbase/tsdb/3516b74d0c9d4458c2f2f715249e3f78\n\n2011-03-16 15:28:34,542 INFO org.apache.hadoop.hbase.catalog.MetaEditor: Deleted region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. from META\n\n# Now the OPENED executor runs, a good while after the above\n#\n\n2011-03-16 15:30:26,679 DEBUG org.apache.hadoop.hbase.master.handler.OpenedRegionHandler: Handling OPENED event for 3516b74d0c9d4458c2f2f715249e3f78; deleting unassigned node\n2011-03-16 15:30:26,679 DEBUG org.apache.hadoop.hbase.zookeeper.ZKAssign: master:60000-0x22e286f0b9c98f1 Deleting existing unassigned node for 3516b74d0c9d4458c2f2f715249e3f78 that is in expected state RS_ZK_REGION_OPENED\n\n\n2011-03-16 15:30:26,725 DEBUG org.apache.hadoop.hbase.zookeeper.ZKUtil: master:60000-0x22e286f0b9c98f1 Retrieved 127 byte(s) of data from znode /prodjobs/unassigned/3516b74d0c9d4458c2f2f715249e3f78; data=region=tsdb,^@^D2McZ@^@^@^A^@^@G^@^@^L^@^@f^@^@^U^@^@�^@^@(^@^C^G,1299401073466.3516b74d0c9d4458c2f2f715249e3f78., server=sv4borg39,60020,1300313564807, state=RS_ZK_REGION_OPENED\n\n\n2011-03-16 15:30:26,875 DEBUG org.apache.hadoop.hbase.zookeeper.ZKAssign: master:60000-0x22e286f0b9c98f1 Successfully deleted unassigned node for region 3516b74d0c9d4458c2f2f715249e3f78 in expected state RS_ZK_REGION_OPENED\n\n\n2011-03-16 15:30:27,051 DEBUG org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher: master:60000-0x22e286f0b9c98f1 Received ZooKeeper Event, type=NodeDeleted, state=SyncConnected, path=/prodjobs/unassigned/3516b74d0c9d4458c2f2f715249e3f78\n\n2011-03-16 15:30:27,051 DEBUG org.apache.hadoop.hbase.master.handler.OpenedRegionHandler: Opened region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. on sv4borg39,60020,1300313564807\n\n# Now we have a region online in Master's memory but its not out in .META. nor in the FS.\n# The balancer runs\n\n2011-03-16 23:18:41,716 INFO org.apache.hadoop.hbase.master.HMaster: balance hri=tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78., src=sv4borg39,60020,1300313564807, dest=sv4borg33,60020,1300342574666\n\n2011-03-16 23:18:41,716 DEBUG org.apache.hadoop.hbase.master.AssignmentManager: Starting unassignment of region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. (offlining)\n\n2011-03-16 23:18:41,718 DEBUG org.apache.hadoop.hbase.master.AssignmentManager: Server serverName=sv4borg39,60020,1300313564807, load=(requests=2, regions=504, usedHeap=929, maxHeap=6973) returned org.apache.hadoop.hbase.NotServingRegionException: org.apache.hadoop.hbase.NotServingRegionException: Received close for tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. but we are not serving it for 3516b74d0c9d4458c2f2f715249e3f78\n\n2011-03-16 23:20:34,436 INFO org.apache.hadoop.hbase.master.AssignmentManager: Regions in transition timed out: tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. state=PENDING_CLOSE, ts=1300342802734\n\n2011-03-16 23:20:34,437 INFO org.apache.hadoop.hbase.master.AssignmentManager: Region has been PENDING_CLOSE for too long, running forced unassign again on region=tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78.\n\n2011-03-16 23:20:34,437 DEBUG org.apache.hadoop.hbase.zookeeper.ZKUtil: master:60000-0x22e286f0b9c98f1 Set watcher on existing znode /prodjobs/unassigned/3516b74d0c9d4458c2f2f715249e3f78\n\n2011-03-16 23:20:34,437 DEBUG org.apache.hadoop.hbase.master.AssignmentManager: Starting unassignment of region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. (offlining)\n2011-03-16 23:20:34,438 DEBUG org.apache.hadoop.hbase.master.AssignmentManager: Attempting to unassign region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. which is already pending close but forcing an additional close\n\n{code}\n\nAd infinitum","from":"reporter","subject":"Split report before we finish parent region open; workaround till 0.92; Race between split and OPENED processing"},{"body":"Small fix. Please review Jon/JD","from":"developer"},{"body":"Minor comment:\n{code}\n+ LOG.warn(\"Skipping the onining of \" + regionInfo.getRegionNameAsString() +\n+ \" because regions is NOT in RIT -- presuming this is because it SPLIT\");\n{code}\nshould be:\n{code}\n+ LOG.warn(\"Skipping the onlining of \" + regionInfo.getRegionNameAsString() +\n+ \" because region is NOT in RIT -- presuming this is because it SPLIT\");\n{code}\n","from":"developer"},{"body":"+1 after the typo Ted reported and if it passes all the tests.","from":"developer"},{"body":"+1 too, thanks for the quick turnaround guys!","from":"developer"},{"body":"Thanks for reviews Ted and J-D. Will commit soon w/ Ted suggestion.","from":"developer"},{"body":"Thanks reviews B, Ted, and J-D.","from":"developer"},{"body":"Looks like there is inconsistency between javadoc and log message.\nFor isRegionInTransition():\n{code}\n * @return Returns null if passed region is not in transition else the current\n{code}\nBut the change in OpenedRegionHandler says:\n{code}\n if (this.assignmentManager.isRegionInTransition(regionInfo) == null) {\n this.assignmentManager.regionOnline(regionInfo, serverInfo);\n } else {\n LOG.warn(\"Skipping the onlining of \" + regionInfo.getRegionNameAsString() +\n \" because regions is NOT in RIT -- presuming this is because it SPLIT\");\n }\n{code}\n\nBTW TestMultipleTimestamps starts to fail after this change.","from":"developer"},{"body":"@Ted You are right. Thanks for spotting this. Let me fix.","from":"developer"}],"created":"2011-03-18T17:42:10.000+0000","description":"This issue is about adding a workaround to 0.90 until we get proper fix in 0.92 (HBASE-3559).\n\nHere is the sequence of events:\n\n1. We start to process OPENED region event.\n2. We receive a SPLIT of this region report.\n3. SPLIT processing offline the region and onlines daughters.\n4. Metascanner runs and clears out the region from .META. deleting it\n5. The OPENED handler runs. Marks the region online in Master memory.\n6. Balancer runs. Trys to balance a region that has been deleted.\n\nLoops for ever.\n\nHere is excerpt from logs. It happened during startup, lots going on. Could happen on regionserver crash I suppose, maybe, but we're susceptible during cluster start:\n\n{code}\n# We assign the region\n2011-03-16 15:18:29,053 DEBUG org.apache.hadoop.hbase.zookeeper.ZKAssign: master:60000-0x22e286f0b9c98f1 Async create of unassigned node for 3516b74d0c9d4458c2f2f715249e3f78 with OFFLINE state\n...\n2011-03-16 15:18:32,298 DEBUG org.apache.hadoop.hbase.master.AssignmentManager$CreateUnassignedAsyncCallback: rs=tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. state=OFFLINE, ts=1300313909053, server=sv4borg39,60020,1300313564807\n...\n2011-03-16 15:18:32,732 DEBUG org.apache.hadoop.hbase.master.AssignmentManager$ExistsUnassignedAsyncCallback: rs=tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. state=OFFLINE, ts=1300313909053\n...\n2011-03-16 15:23:02,114 DEBUG org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher: master:60000-0x22e286f0b9c98f1 Received ZooKeeper Event, type=NodeDataChanged, state=SyncConnected, path=/prodjobs/unassigned/3516b74d0c9d4458c2f2f715249e3f78\n...\n2011-03-16 15:23:02,183 DEBUG org.apache.hadoop.hbase.zookeeper.ZKUtil: master:60000-0x22e286f0b9c98f1 Retrieved 127 byte(s) of data from znode /prodjobs/unassigned/3516b74d0c9d4458c2f2f715249e3f78 and set watcher; region=tsdb,^@^D2McZ@^@^@^A^@^@G^@^@^L^@^@f^@^@^U^@^@�^@^@(^@^C^G,1299401073466.3516b74d0c9d4458c2f2f715249e3f78., server=sv4borg39,60020,1300313564807, state=RS_ZK_REGION_OPENED\n2011-03-16 15:23:02,183 DEBUG org.apache.hadoop.hbase.master.AssignmentManager: Handling transition=RS_ZK_REGION_OPENED, server=sv4borg39,60020,1300313564807, region=3516b74d0c9d4458c2f2f715249e3f78\n\n# At this point we've queued an Excecutor to run to process the OPENED event. Now in comes the SPLIT.\n2011-03-16 15:23:18,199 INFO org.apache.hadoop.hbase.master.ServerManager: Received REGION_SPLIT: tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78.: Daughters; tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1300314189812.74c51400bb8dfa127fadfd11a04d72f2., tsdb,\\x00\\x042MmD\\x88\\x00\\x00\\x01\\x00\\x00S\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x029\\x00\\x00(\\x00\\x03\\x03,1300314189812.87b061739a11d0f9d02acfb92ef961a2. from sv4borg39,60020,1300313564807\n2011-03-16 15:23:18,870 WARN org.apache.hadoop.hbase.master.AssignmentManager: Split report has RIT node (shouldnt have one): REGION => {NAME => 'tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78.', STARTKEY => '\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07', ENDKEY => '\\x00\\x043L\\xE7\\xF50\\x00\\x00\\x01\\x00\\x00I\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x0E\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x02u', ENCODED => 3516b74d0c9d4458c2f2f715249e3f78, TABLE => {{NAME => 'tsdb', FAMILIES => [{NAME => 't', BLOOMFILTER => 'NONE', REPLICATION_SCOPE => '0', VERSIONS => '3', COMPRESSION => 'LZO', TTL => '2147483647', BLOCKSIZE => '65536', IN_MEMORY => 'false', BLOCKCACHE => 'true'}]}} node: region=tsdb,^@^D2McZ@^@^@^A^@^@G^@^@^L^@^@f^@^@^U^@^@�^@^@(^@^C^G,1299401073466.3516b74d0c9d4458c2f2f715249e3f78., server=sv4borg39,60020,1300313564807, state=RS_ZK_REGION_OPENED\n\n# Now metascanner runs and actually removes the parent region, deleting it all\n\n2011-03-16 15:28:34,352 INFO org.apache.hadoop.hbase.catalog.MetaEditor: Deleted daughter reference tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1300314189812.74c51400bb8dfa127fadfd11a04d72f2., qualifier=splitA, from parent tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78.\n2011-03-16 15:28:34,356 INFO org.apache.hadoop.hbase.catalog.MetaEditor: Deleted daughter reference tsdb,\\x00\\x042MmD\\x88\\x00\\x00\\x01\\x00\\x00S\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x029\\x00\\x00(\\x00\\x03\\x03,1300314189812.87b061739a11d0f9d02acfb92ef961a2., qualifier=splitB, from parent tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78.\n2011-03-16 15:28:34,356 DEBUG org.apache.hadoop.hbase.master.CatalogJanitor: Deleting region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. because daughter splits no longer hold references\n2011-03-16 15:28:34,356 DEBUG org.apache.hadoop.hbase.master.CatalogJanitor: Deleting region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. because daughter splits no longer hold references\n\n2011-03-16 15:28:34,444 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: DELETING region hdfs://sv4borg29:9000/hbase/tsdb/3516b74d0c9d4458c2f2f715249e3f78\n\n2011-03-16 15:28:34,542 INFO org.apache.hadoop.hbase.catalog.MetaEditor: Deleted region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. from META\n\n# Now the OPENED executor runs, a good while after the above\n#\n\n2011-03-16 15:30:26,679 DEBUG org.apache.hadoop.hbase.master.handler.OpenedRegionHandler: Handling OPENED event for 3516b74d0c9d4458c2f2f715249e3f78; deleting unassigned node\n2011-03-16 15:30:26,679 DEBUG org.apache.hadoop.hbase.zookeeper.ZKAssign: master:60000-0x22e286f0b9c98f1 Deleting existing unassigned node for 3516b74d0c9d4458c2f2f715249e3f78 that is in expected state RS_ZK_REGION_OPENED\n\n\n2011-03-16 15:30:26,725 DEBUG org.apache.hadoop.hbase.zookeeper.ZKUtil: master:60000-0x22e286f0b9c98f1 Retrieved 127 byte(s) of data from znode /prodjobs/unassigned/3516b74d0c9d4458c2f2f715249e3f78; data=region=tsdb,^@^D2McZ@^@^@^A^@^@G^@^@^L^@^@f^@^@^U^@^@�^@^@(^@^C^G,1299401073466.3516b74d0c9d4458c2f2f715249e3f78., server=sv4borg39,60020,1300313564807, state=RS_ZK_REGION_OPENED\n\n\n2011-03-16 15:30:26,875 DEBUG org.apache.hadoop.hbase.zookeeper.ZKAssign: master:60000-0x22e286f0b9c98f1 Successfully deleted unassigned node for region 3516b74d0c9d4458c2f2f715249e3f78 in expected state RS_ZK_REGION_OPENED\n\n\n2011-03-16 15:30:27,051 DEBUG org.apache.hadoop.hbase.zookeeper.ZooKeeperWatcher: master:60000-0x22e286f0b9c98f1 Received ZooKeeper Event, type=NodeDeleted, state=SyncConnected, path=/prodjobs/unassigned/3516b74d0c9d4458c2f2f715249e3f78\n\n2011-03-16 15:30:27,051 DEBUG org.apache.hadoop.hbase.master.handler.OpenedRegionHandler: Opened region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. on sv4borg39,60020,1300313564807\n\n# Now we have a region online in Master's memory but its not out in .META. nor in the FS.\n# The balancer runs\n\n2011-03-16 23:18:41,716 INFO org.apache.hadoop.hbase.master.HMaster: balance hri=tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78., src=sv4borg39,60020,1300313564807, dest=sv4borg33,60020,1300342574666\n\n2011-03-16 23:18:41,716 DEBUG org.apache.hadoop.hbase.master.AssignmentManager: Starting unassignment of region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. (offlining)\n\n2011-03-16 23:18:41,718 DEBUG org.apache.hadoop.hbase.master.AssignmentManager: Server serverName=sv4borg39,60020,1300313564807, load=(requests=2, regions=504, usedHeap=929, maxHeap=6973) returned org.apache.hadoop.hbase.NotServingRegionException: org.apache.hadoop.hbase.NotServingRegionException: Received close for tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. but we are not serving it for 3516b74d0c9d4458c2f2f715249e3f78\n\n2011-03-16 23:20:34,436 INFO org.apache.hadoop.hbase.master.AssignmentManager: Regions in transition timed out: tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. state=PENDING_CLOSE, ts=1300342802734\n\n2011-03-16 23:20:34,437 INFO org.apache.hadoop.hbase.master.AssignmentManager: Region has been PENDING_CLOSE for too long, running forced unassign again on region=tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78.\n\n2011-03-16 23:20:34,437 DEBUG org.apache.hadoop.hbase.zookeeper.ZKUtil: master:60000-0x22e286f0b9c98f1 Set watcher on existing znode /prodjobs/unassigned/3516b74d0c9d4458c2f2f715249e3f78\n\n2011-03-16 23:20:34,437 DEBUG org.apache.hadoop.hbase.master.AssignmentManager: Starting unassignment of region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. (offlining)\n2011-03-16 23:20:34,438 DEBUG org.apache.hadoop.hbase.master.AssignmentManager: Attempting to unassign region tsdb,\\x00\\x042McZ@\\x00\\x00\\x01\\x00\\x00G\\x00\\x00\\x0C\\x00\\x00f\\x00\\x00\\x15\\x00\\x00\\xA9\\x00\\x00(\\x00\\x03\\x07,1299401073466.3516b74d0c9d4458c2f2f715249e3f78. which is already pending close but forcing an additional close\n\n{code}\n\nAd infinitum","issue_id":"12501811","key":"HBASE-3671","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2011-03-19T20:22:00.000+0000","role":"fixed_distractor","summary":"Split report before we finish parent region open; workaround till 0.92; Race between split and OPENED processing"} {"case_id":"12501846","cluster":"DISTRACTOR-HBASE-3674","comments":[{"body":"Here is a first cut at a patch.","created":"2011-03-18T23:57:52.229+0000"},{"body":"Submitting patch, needs a test though.","created":"2011-03-19T00:00:59.737+0000"},{"body":"+1","created":"2011-03-22T01:59:08.516+0000"},{"body":"Committed branch and trunk. Thanks for the review Nicolas","created":"2011-03-22T02:08:38.553+0000"},{"body":"This change got overwritten when HBASE-1364 was integrated.\n\nThe change has to be added in HLogSplitter in the method getNextLogLine\n\nstatic private Entry getNextLogLine(Reader in, Path path, boolean skipErrors)\n throws CorruptedLogFileException, IOException {\n try {\n return in.next();\n } catch (EOFException eof) {\n // truncated files are expected if a RS crashes (see HBASE-2643)\n LOG.info(\"EOF from hlog \" + path + \". continuing\");\n return null;\n } catch (IOException e) {\n // If the IOE resulted from bad file format,\n // then this problem is idempotent and retrying won't help\n if (e.getCause() instanceof ParseException) {\n LOG.warn(\"ParseException from hlog \" + path + \". continuing\");\n return null;\n }\n\n\nIt might also be necessary to add this change to getReader(...) method\n\n protected Reader getReader(FileSystem fs, FileStatus file, Configuration conf,\n boolean skipErrors)\n throws IOException, CorruptedLogFileException {\n Path path = file.getPath();\n long length = file.getLen();\n Reader in;\n\n\n // Check for possibly empty file. With appends, currently Hadoop reports a\n // zero length even if the file has been sync'd. Revisit if HDFS-376 or\n // HDFS-878 is committed.\n if (length <= 0) {\n LOG.warn(\"File \" + path + \" might be still open, length is 0\");\n }\n\n try {\n recoverFileLease(fs, path, conf);\n try {\n in = getReader(fs, path, conf);\n } catch (EOFException e) {\n if (length <= 0) {\n // TODO should we ignore an empty, not-last log file if skip.errors\n // is false? Either way, the caller should decide what to do. E.g.\n // ignore if this is the last log in sequence.\n // TODO is this scenario still possible if the log has been\n // recovered (i.e. closed)\n LOG.warn(\"Could not open \" + path + \" for reading. File is empty\", e);\n return null;\n\n\n","created":"2011-04-21T19:09:43.790+0000"},{"body":"The patch sets the hbase.hlog.split.skip.errors to true by default. I am wondering why the CheckSumException was not ignored as originally proposed?\n\nThis patch is there in the trunk. In the serialized log splitting case hbase.hlog.split.skip.errors is set to true. But in the distributed log splitting case hbase.hlog.split.skip.errors is set to false by default.","created":"2011-04-21T22:41:00.612+0000"},{"body":"I actually missed when this went in. Having the default be to skip errors seems really really dangerous.... especially as a change in a release branch.","created":"2011-04-22T00:04:28.289+0000"},{"body":"bq. I actually missed when this went in. Having the default be to skip errors seems really really dangerous.... especially as a change in a release branch.\n\nDefault was true but we would keep going anyways; it was incorrectly implemented. I opened HBASE-3675 at the time. Suggest that we continue the discussion around how dangerous changing the config is/was over there. In here I'll just make sure the distributed splitter has same flags as the splitter it replaces.\n\nThanks Todd for helping keeping us honest.","created":"2011-04-25T23:34:56.059+0000"},{"body":"bq. The patch sets the hbase.hlog.split.skip.errors to true by default. I am wondering why the CheckSumException was not ignored as originally proposed?\n\nPrakash, there are two patches. The first adds ignoring checksumexception.","created":"2011-04-26T03:49:25.861+0000"},{"body":"How about this Prakash? This adds back the changes and sets default to true for the fail flag.","created":"2011-04-26T03:54:48.129+0000"},{"body":"+1","created":"2011-04-26T05:01:45.357+0000"},{"body":"Committed to TRUNK. Thanks for review Prakash.","created":"2011-04-26T18:32:52.603+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T12:42:56.135+0000"}],"conversations":[{"body":"In short, a ChecksumException will fail log processing for a server so we skip out w/o archiving logs. On restart, we'll then reprocess the logs -- hit the checksumexception anew, usually -- and so on.\n\nHere is the splitLog method (edited):\n\n{code}\n private List splitLog(final FileStatus[] logfiles) throws IOException {\n ....\n outputSink.startWriterThreads(entryBuffers);\n \n try {\n int i = 0;\n for (FileStatus log : logfiles) {\n Path logPath = log.getPath();\n long logLength = log.getLen();\n splitSize += logLength;\n LOG.debug(\"Splitting hlog \" + (i++ + 1) + \" of \" + logfiles.length\n + \": \" + logPath + \", length=\" + logLength);\n try {\n recoverFileLease(fs, logPath, conf);\n parseHLog(log, entryBuffers, fs, conf);\n processedLogs.add(logPath);\n } catch (EOFException eof) {\n // truncated files are expected if a RS crashes (see HBASE-2643)\n LOG.info(\"EOF from hlog \" + logPath + \". Continuing\");\n processedLogs.add(logPath);\n } catch (FileNotFoundException fnfe) {\n // A file may be missing if the region server was able to archive it\n // before shutting down. This means the edits were persisted already\n LOG.info(\"A log was missing \" + logPath +\n \", probably because it was moved by the\" +\n \" now dead region server. Continuing\");\n processedLogs.add(logPath);\n } catch (IOException e) {\n // If the IOE resulted from bad file format,\n // then this problem is idempotent and retrying won't help\n if (e.getCause() instanceof ParseException ||\n e.getCause() instanceof ChecksumException) {\n LOG.warn(\"ParseException from hlog \" + logPath + \". continuing\");\n processedLogs.add(logPath);\n } else {\n if (skipErrors) {\n LOG.info(\"Got while parsing hlog \" + logPath +\n \". Marking as corrupted\", e);\n corruptedLogs.add(logPath);\n } else {\n throw e;\n }\n }\n }\n }\n if (fs.listStatus(srcDir).length > processedLogs.size()\n + corruptedLogs.size()) {\n throw new OrphanHLogAfterSplitException(\n \"Discovered orphan hlog after split. Maybe the \"\n + \"HRegionServer was not dead when we started\");\n }\n archiveLogs(srcDir, corruptedLogs, processedLogs, oldLogDir, fs, conf); \n } finally {\n splits = outputSink.finishWritingAndClose();\n }\n return splits;\n }\n{code}\n\nNotice how we'll only archive logs only if we successfully split all logs. We won't archive 31 of 35 files if we happen to get a checksum exception on file 32.\n\nI think we should treat a ChecksumException the same as a ParseException; a retry will not fix it if HDFS could not get around the ChecksumException (seems like in our case all replicas were corrupt).\n\nHere is a play-by-play from the logs:\n\n{code}\n813572 2011-03-18 20:31:44,687 DEBUG org.apache.hadoop.hbase.regionserver.wal.HLogSplitter: Splitting hlog 34 of 35: hdfs://sv2borg170:9000/hbase/.logs/sv2borg182,60020,1300384550664/sv2borg182%3A60020.1300461329481, length=150 65662813573 2011-03-18 20:31:44,687 INFO org.apache.hadoop.hbase.util.FSUtils: Recovering file hdfs://sv2borg170:9000/hbase/.logs/sv2borg182,60020,1300384550664/sv2borg182%3A60020.1300461329481\n....\n813617 2011-03-18 20:31:46,238 INFO org.apache.hadoop.fs.FSInputChecker: Found checksum error: b[0, 512]=000000cd000000502037383661376439656265643938636463343433386132343631323633303239371d6170695f6163636573735f746f6b656e5f7374 6174735f6275636b65740000000d9fa4d5dc0000012ec9c7cbaf00ffffffff000000010000006d0000005d00000008002337626262663764626431616561366234616130656334383436653732333132643a32390764656661756c746170695f616e64726f69645f6c6f67676564 696e5f73686172655f70656e64696e675f696e69740000012ec956b02804000000000000000100000000ffffffff4e128eca0eb078d0652b0abac467fd09000000cd000000502034663166613763666165333930666332653138346233393931303132623366331d6170695f6163 636573735f746f6b656e5f73746174735f6275636b65740000000d9fa4d5dd0000012ec9c7cbaf00ffffffff000000010000006d0000005d00000008002366303734323966643036323862636530336238333938356239316237386633353a32390764656661756c746170695f61 6e64726f69645f6c6f67676564696e5f73686172655f70656e64696e675f696e69740000012ec9569f1804000000000000000100000000000000d30000004e2066663763393964303633343339666531666461633761616632613964643631331b6170695f6163636573735f746f 6b656e5f73746174735f68\n813618 org.apache.hadoop.fs.ChecksumException: Checksum error: /blk_7781725413191608261:of:/hbase/.logs/sv2borg182,60020,1300384550664/sv2borg182%3A60020.1300461329481 at 15064576\n813619 at org.apache.hadoop.fs.FSInputChecker.verifySum(FSInputChecker.java:277)\n813620 at org.apache.hadoop.fs.FSInputChecker.readChecksumChunk(FSInputChecker.java:241)\n813621 at org.apache.hadoop.fs.FSInputChecker.fill(FSInputChecker.java:176)\n813622 at org.apache.hadoop.fs.FSInputChecker.read1(FSInputChecker.java:193)\n813623 at org.apache.hadoop.fs.FSInputChecker.read(FSInputChecker.java:158)\n813624 at org.apache.hadoop.hdfs.DFSClient$BlockReader.read(DFSClient.java:1175)\n813625 at org.apache.hadoop.hdfs.DFSClient$DFSInputStream.readBuffer(DFSClient.java:1807)\n813626 at org.apache.hadoop.hdfs.DFSClient$DFSInputStream.read(DFSClient.java:1859)\n813627 at java.io.DataInputStream.read(DataInputStream.java:132)\n813628 at java.io.DataInputStream.readFully(DataInputStream.java:178)\n813629 at org.apache.hadoop.io.DataOutputBuffer$Buffer.write(DataOutputBuffer.java:63)\n813630 at org.apache.hadoop.io.DataOutputBuffer.write(DataOutputBuffer.java:101)\n813631 at org.apache.hadoop.io.SequenceFile$Reader.next(SequenceFile.java:1937)\n813632 at org.apache.hadoop.io.SequenceFile$Reader.next(SequenceFile.java:1837)\n813633 at org.apache.hadoop.io.SequenceFile$Reader.next(SequenceFile.java:1883)\n813634 at org.apache.hadoop.hbase.regionserver.wal.SequenceFileLogReader.next(SequenceFileLogReader.java:198)\n813635 at org.apache.hadoop.hbase.regionserver.wal.SequenceFileLogReader.next(SequenceFileLogReader.java:172)\n813636 at org.apache.hadoop.hbase.regionserver.wal.HLogSplitter.parseHLog(HLogSplitter.java:429)\n813637 at org.apache.hadoop.hbase.regionserver.wal.HLogSplitter.splitLog(HLogSplitter.java:262)\n813638 at org.apache.hadoop.hbase.regionserver.wal.HLogSplitter.splitLog(HLogSplitter.java:188)\n813639 at org.apache.hadoop.hbase.master.MasterFileSystem.splitLog(MasterFileSystem.java:197)\n813640 at org.apache.hadoop.hbase.master.MasterFileSystem.splitLogAfterStartup(MasterFileSystem.java:181)\n813641 at org.apache.hadoop.hbase.master.HMaster.finishInitialization(HMaster.java:384)\n813642 at org.apache.hadoop.hbase.master.HMaster.run(HMaster.java:283)\n813643 2011-03-18 20:31:46,239 WARN org.apache.hadoop.hdfs.DFSClient: Found Checksum error for blk_7781725413191608261_14589573 from 10.20.20.182:50010 at 15064576\n813644 2011-03-18 20:31:46,240 INFO org.apache.hadoop.hdfs.DFSClient: Could not obtain block blk_7781725413191608261_14589573 from any node: java.io.IOException: No live nodes contain current block. Will get new block locations from namenode and retry...\n813645 2011-03-18 20:31:49,243 DEBUG org.apache.hadoop.hbase.regionserver.wal.HLogSplitter: Pushed=80624 entries from hdfs://sv2borg170:9000/hbase/.logs/sv2borg182,60020,1300384550664/sv2borg182%3A60020.1300461329481\n....\n{code}\n\nSee code above. On exception we'll dump edits read so far from this block, close out all writers tying off recovered.edits so far written. We'll skip archiving these files because we only archive if all files are processed; we won't archive files 30 of 35 if we failed splitting on file 31.\n\nI think checksumexception should be treated same as a ParseException\n\n \n\n","from":"reporter","subject":"Treat ChecksumException as we would a ParseException splitting logs; else we replay split on every restart"},{"body":"Here is a first cut at a patch.","from":"developer"},{"body":"Submitting patch, needs a test though.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed branch and trunk. Thanks for the review Nicolas","from":"developer"},{"body":"This change got overwritten when HBASE-1364 was integrated.\n\nThe change has to be added in HLogSplitter in the method getNextLogLine\n\nstatic private Entry getNextLogLine(Reader in, Path path, boolean skipErrors)\n throws CorruptedLogFileException, IOException {\n try {\n return in.next();\n } catch (EOFException eof) {\n // truncated files are expected if a RS crashes (see HBASE-2643)\n LOG.info(\"EOF from hlog \" + path + \". continuing\");\n return null;\n } catch (IOException e) {\n // If the IOE resulted from bad file format,\n // then this problem is idempotent and retrying won't help\n if (e.getCause() instanceof ParseException) {\n LOG.warn(\"ParseException from hlog \" + path + \". continuing\");\n return null;\n }\n\n\nIt might also be necessary to add this change to getReader(...) method\n\n protected Reader getReader(FileSystem fs, FileStatus file, Configuration conf,\n boolean skipErrors)\n throws IOException, CorruptedLogFileException {\n Path path = file.getPath();\n long length = file.getLen();\n Reader in;\n\n\n // Check for possibly empty file. With appends, currently Hadoop reports a\n // zero length even if the file has been sync'd. Revisit if HDFS-376 or\n // HDFS-878 is committed.\n if (length <= 0) {\n LOG.warn(\"File \" + path + \" might be still open, length is 0\");\n }\n\n try {\n recoverFileLease(fs, path, conf);\n try {\n in = getReader(fs, path, conf);\n } catch (EOFException e) {\n if (length <= 0) {\n // TODO should we ignore an empty, not-last log file if skip.errors\n // is false? Either way, the caller should decide what to do. E.g.\n // ignore if this is the last log in sequence.\n // TODO is this scenario still possible if the log has been\n // recovered (i.e. closed)\n LOG.warn(\"Could not open \" + path + \" for reading. File is empty\", e);\n return null;\n\n\n","from":"developer"},{"body":"The patch sets the hbase.hlog.split.skip.errors to true by default. I am wondering why the CheckSumException was not ignored as originally proposed?\n\nThis patch is there in the trunk. In the serialized log splitting case hbase.hlog.split.skip.errors is set to true. But in the distributed log splitting case hbase.hlog.split.skip.errors is set to false by default.","from":"developer"},{"body":"I actually missed when this went in. Having the default be to skip errors seems really really dangerous.... especially as a change in a release branch.","from":"developer"},{"body":"bq. I actually missed when this went in. Having the default be to skip errors seems really really dangerous.... especially as a change in a release branch.\n\nDefault was true but we would keep going anyways; it was incorrectly implemented. I opened HBASE-3675 at the time. Suggest that we continue the discussion around how dangerous changing the config is/was over there. In here I'll just make sure the distributed splitter has same flags as the splitter it replaces.\n\nThanks Todd for helping keeping us honest.","from":"developer"},{"body":"bq. The patch sets the hbase.hlog.split.skip.errors to true by default. I am wondering why the CheckSumException was not ignored as originally proposed?\n\nPrakash, there are two patches. The first adds ignoring checksumexception.","from":"developer"},{"body":"How about this Prakash? This adds back the changes and sets default to true for the fail flag.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Committed to TRUNK. Thanks for review Prakash.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2011-03-18T23:49:33.000+0000","description":"In short, a ChecksumException will fail log processing for a server so we skip out w/o archiving logs. On restart, we'll then reprocess the logs -- hit the checksumexception anew, usually -- and so on.\n\nHere is the splitLog method (edited):\n\n{code}\n private List splitLog(final FileStatus[] logfiles) throws IOException {\n ....\n outputSink.startWriterThreads(entryBuffers);\n \n try {\n int i = 0;\n for (FileStatus log : logfiles) {\n Path logPath = log.getPath();\n long logLength = log.getLen();\n splitSize += logLength;\n LOG.debug(\"Splitting hlog \" + (i++ + 1) + \" of \" + logfiles.length\n + \": \" + logPath + \", length=\" + logLength);\n try {\n recoverFileLease(fs, logPath, conf);\n parseHLog(log, entryBuffers, fs, conf);\n processedLogs.add(logPath);\n } catch (EOFException eof) {\n // truncated files are expected if a RS crashes (see HBASE-2643)\n LOG.info(\"EOF from hlog \" + logPath + \". Continuing\");\n processedLogs.add(logPath);\n } catch (FileNotFoundException fnfe) {\n // A file may be missing if the region server was able to archive it\n // before shutting down. This means the edits were persisted already\n LOG.info(\"A log was missing \" + logPath +\n \", probably because it was moved by the\" +\n \" now dead region server. Continuing\");\n processedLogs.add(logPath);\n } catch (IOException e) {\n // If the IOE resulted from bad file format,\n // then this problem is idempotent and retrying won't help\n if (e.getCause() instanceof ParseException ||\n e.getCause() instanceof ChecksumException) {\n LOG.warn(\"ParseException from hlog \" + logPath + \". continuing\");\n processedLogs.add(logPath);\n } else {\n if (skipErrors) {\n LOG.info(\"Got while parsing hlog \" + logPath +\n \". Marking as corrupted\", e);\n corruptedLogs.add(logPath);\n } else {\n throw e;\n }\n }\n }\n }\n if (fs.listStatus(srcDir).length > processedLogs.size()\n + corruptedLogs.size()) {\n throw new OrphanHLogAfterSplitException(\n \"Discovered orphan hlog after split. Maybe the \"\n + \"HRegionServer was not dead when we started\");\n }\n archiveLogs(srcDir, corruptedLogs, processedLogs, oldLogDir, fs, conf); \n } finally {\n splits = outputSink.finishWritingAndClose();\n }\n return splits;\n }\n{code}\n\nNotice how we'll only archive logs only if we successfully split all logs. We won't archive 31 of 35 files if we happen to get a checksum exception on file 32.\n\nI think we should treat a ChecksumException the same as a ParseException; a retry will not fix it if HDFS could not get around the ChecksumException (seems like in our case all replicas were corrupt).\n\nHere is a play-by-play from the logs:\n\n{code}\n813572 2011-03-18 20:31:44,687 DEBUG org.apache.hadoop.hbase.regionserver.wal.HLogSplitter: Splitting hlog 34 of 35: hdfs://sv2borg170:9000/hbase/.logs/sv2borg182,60020,1300384550664/sv2borg182%3A60020.1300461329481, length=150 65662813573 2011-03-18 20:31:44,687 INFO org.apache.hadoop.hbase.util.FSUtils: Recovering file hdfs://sv2borg170:9000/hbase/.logs/sv2borg182,60020,1300384550664/sv2borg182%3A60020.1300461329481\n....\n813617 2011-03-18 20:31:46,238 INFO org.apache.hadoop.fs.FSInputChecker: Found checksum error: b[0, 512]=000000cd000000502037383661376439656265643938636463343433386132343631323633303239371d6170695f6163636573735f746f6b656e5f7374 6174735f6275636b65740000000d9fa4d5dc0000012ec9c7cbaf00ffffffff000000010000006d0000005d00000008002337626262663764626431616561366234616130656334383436653732333132643a32390764656661756c746170695f616e64726f69645f6c6f67676564 696e5f73686172655f70656e64696e675f696e69740000012ec956b02804000000000000000100000000ffffffff4e128eca0eb078d0652b0abac467fd09000000cd000000502034663166613763666165333930666332653138346233393931303132623366331d6170695f6163 636573735f746f6b656e5f73746174735f6275636b65740000000d9fa4d5dd0000012ec9c7cbaf00ffffffff000000010000006d0000005d00000008002366303734323966643036323862636530336238333938356239316237386633353a32390764656661756c746170695f61 6e64726f69645f6c6f67676564696e5f73686172655f70656e64696e675f696e69740000012ec9569f1804000000000000000100000000000000d30000004e2066663763393964303633343339666531666461633761616632613964643631331b6170695f6163636573735f746f 6b656e5f73746174735f68\n813618 org.apache.hadoop.fs.ChecksumException: Checksum error: /blk_7781725413191608261:of:/hbase/.logs/sv2borg182,60020,1300384550664/sv2borg182%3A60020.1300461329481 at 15064576\n813619 at org.apache.hadoop.fs.FSInputChecker.verifySum(FSInputChecker.java:277)\n813620 at org.apache.hadoop.fs.FSInputChecker.readChecksumChunk(FSInputChecker.java:241)\n813621 at org.apache.hadoop.fs.FSInputChecker.fill(FSInputChecker.java:176)\n813622 at org.apache.hadoop.fs.FSInputChecker.read1(FSInputChecker.java:193)\n813623 at org.apache.hadoop.fs.FSInputChecker.read(FSInputChecker.java:158)\n813624 at org.apache.hadoop.hdfs.DFSClient$BlockReader.read(DFSClient.java:1175)\n813625 at org.apache.hadoop.hdfs.DFSClient$DFSInputStream.readBuffer(DFSClient.java:1807)\n813626 at org.apache.hadoop.hdfs.DFSClient$DFSInputStream.read(DFSClient.java:1859)\n813627 at java.io.DataInputStream.read(DataInputStream.java:132)\n813628 at java.io.DataInputStream.readFully(DataInputStream.java:178)\n813629 at org.apache.hadoop.io.DataOutputBuffer$Buffer.write(DataOutputBuffer.java:63)\n813630 at org.apache.hadoop.io.DataOutputBuffer.write(DataOutputBuffer.java:101)\n813631 at org.apache.hadoop.io.SequenceFile$Reader.next(SequenceFile.java:1937)\n813632 at org.apache.hadoop.io.SequenceFile$Reader.next(SequenceFile.java:1837)\n813633 at org.apache.hadoop.io.SequenceFile$Reader.next(SequenceFile.java:1883)\n813634 at org.apache.hadoop.hbase.regionserver.wal.SequenceFileLogReader.next(SequenceFileLogReader.java:198)\n813635 at org.apache.hadoop.hbase.regionserver.wal.SequenceFileLogReader.next(SequenceFileLogReader.java:172)\n813636 at org.apache.hadoop.hbase.regionserver.wal.HLogSplitter.parseHLog(HLogSplitter.java:429)\n813637 at org.apache.hadoop.hbase.regionserver.wal.HLogSplitter.splitLog(HLogSplitter.java:262)\n813638 at org.apache.hadoop.hbase.regionserver.wal.HLogSplitter.splitLog(HLogSplitter.java:188)\n813639 at org.apache.hadoop.hbase.master.MasterFileSystem.splitLog(MasterFileSystem.java:197)\n813640 at org.apache.hadoop.hbase.master.MasterFileSystem.splitLogAfterStartup(MasterFileSystem.java:181)\n813641 at org.apache.hadoop.hbase.master.HMaster.finishInitialization(HMaster.java:384)\n813642 at org.apache.hadoop.hbase.master.HMaster.run(HMaster.java:283)\n813643 2011-03-18 20:31:46,239 WARN org.apache.hadoop.hdfs.DFSClient: Found Checksum error for blk_7781725413191608261_14589573 from 10.20.20.182:50010 at 15064576\n813644 2011-03-18 20:31:46,240 INFO org.apache.hadoop.hdfs.DFSClient: Could not obtain block blk_7781725413191608261_14589573 from any node: java.io.IOException: No live nodes contain current block. Will get new block locations from namenode and retry...\n813645 2011-03-18 20:31:49,243 DEBUG org.apache.hadoop.hbase.regionserver.wal.HLogSplitter: Pushed=80624 entries from hdfs://sv2borg170:9000/hbase/.logs/sv2borg182,60020,1300384550664/sv2borg182%3A60020.1300461329481\n....\n{code}\n\nSee code above. On exception we'll dump edits read so far from this block, close out all writers tying off recovered.edits so far written. We'll skip archiving these files because we only archive if all files are processed; we won't archive files 30 of 35 if we failed splitting on file 31.\n\nI think checksumexception should be treated same as a ParseException\n\n \n\n","issue_id":"12501846","key":"HBASE-3674","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2011-03-22T02:08:38.000+0000","role":"fixed_distractor","summary":"Treat ChecksumException as we would a ParseException splitting logs; else we replay split on every restart"} {"case_id":"12512427","cluster":"DISTRACTOR-HBASE-4052","comments":[{"body":"As per my analysis the problem is that,\n\nWhen we do a remove all the online regions are closed.\nIn the Enable table flow \n{noformat}\n private List regionsToAssign(final List regionsInMeta)\n throws IOException {\n final List onlineRegions =\n this.assignmentManager.getRegionsOfTable(tableName);\n regionsInMeta.removeAll(onlineRegions);\n return regionsInMeta;\n }\n{noformat}\nWe remove the regions if it is already online.\n\nBut as per the bug, enable is called after switching,\n\nSo while the standby master becomes active, in the rebuildUserRegion api,\nwe add all the regions from the Meta and we dont check if it is already disabled.\n\nSo when the flow comes to enable (regionsToAssign()) we consider the region to be onlined already. \nFinally when we try to scan the table, we get NotServingRegionException.\n\nCorrect me if am wrong in my analysis. \n","created":"2011-07-01T05:01:11.725+0000"},{"body":"What you say seems plausible Ramkrishna. I traced your reasoning above and yes, because rebuildUserRegion adds all regions without regard to whether table is enabled/disabled, when it comes to run the enable, when it asks what regions to enable, they all look as though they are already online because of what getRegionsOfTable returns (all regions that make up the table pulled from this.regions over in AM).\n\nGood stuff.\n\nDo you have a patch Ramkrishna?","created":"2011-07-01T23:18:15.361+0000"},{"body":"Thanks Stack for your comments.\nI am working on the patch. \nI have 2 scenarios to consider\n1. The AM may get killed when the state is in DISABLING but still the regions are not closed\n2. The AM may get killed when the state is in DISABLING but the regions are closed.\n\nSo can we check both for DISABLING and DISABLED state.\nI will provide my patch ASAP.\n\nThanks","created":"2011-07-02T17:15:24.074+0000"},{"body":"I have some doubts and like to get some suggestion before proceeding.\n \nFollowing scenarios needs to be considered.\n Scenario 1:\n===========\nAll the regions are disabled and the state in zookeeper is DISABLED.\nScenario:2\n==========\n \nThe regions are offlined but the AM went down when the zookeeper state was DISABLING.\n \nScenario:3\n=========\n \nThe regions are not yet offlined(or only few regions are offlined) and the AM went down when the zookeeper state was DISABLING.\n \nNow when we do a switch of the master or on restart scenario of master,\nhow can we decide which regions were offlined and which are not.\n \nThough we can get the state of the table as either DISABLED or DISABLING, region wise i am not able to infer in what state the region is.\n \nSo what brings me to get this info is \n \nThe soln should be like we need to check for the state of the table while populating the regions map in master startup.\n \nChecking only for DISABLED state:\n==========================\nCheck for disabled state and those regions that are not in the DISABLED state add it into the regions map in master startup.\n \nIf i check only for the DISABLED state and if the table is in DISABLING state and \n \nafter master retry (or switch) if i try to enable then we will not be able to scan the table because while enabling none of the regions will be enabled\nas the regions in META table and the regions that i have populated in the regions map are same.\nSo I will be getting the same issue as in the description of the defect.\n \nChecking for DISABLED and DISABLING state:\n===================================\nif i check the state of the zookeeper for DISABLED and DISABLING and while restart of master(switch) only those regions which are not in DISABLED or DISABLING state is populated.\nWhen i again try to enable the region if the region was not offlined as part of disable flow(Scenario:3), the waitUntilDone in BulkAssigner is not aware that the region was \nalready onlined and keeps on waiting as the waitUntilDone() sees for the number of regions to become online from the regions map and the actual count it gets from the meta table.\nThis makes enable to go in a loop.\n \nAm i clear with the problem? so is it like before enabling any table do we need to check the state of the table and if it is DISABLING make all those regions to go to\noffline mode.","created":"2011-07-05T15:20:41.035+0000"},{"body":"I think one source of confusion in the above description is that DISABLED and DISABLING states are TableState.\nThere is no concept of disabling/disabled region.\n\nI would opt for Checking for DISABLED and DISABLING TableState.\n\nI think the BulkAssigner in this case refers to EnableTableHandler.BulkEnabler which depends on assignmentManager.getRegionsOfTable() for the number of online regions.\n\nYou need to write a new method, e.g. assignmentManager.getOnlineRegionsOfTable(), and call it in place of assignmentManager.getRegionsOfTable() in EnableTableHandler.regionsToAssign()","created":"2011-07-05T18:24:57.865+0000"},{"body":"Hi,\nSorry for using wrong terminologies. I was aware that the Tables are only disabled and regions are only onlined or offlined.\nI would like to attach the scenarios that needs to be addressed as part of this bug\n\n{noformat}\nIf a table T1 has three regions R1, R2 and R3 \nWe issue a disable command for T1\nScenario 1:\n===========\nAll the regions R1, R2 and R3 are offlined and the state in zookeeper for the table is DISABLED and then the Active master went down,\nR1- Offlined\nR2- Offlined\t\t\tT1-DISABLED\nR3-Offlined\n\tThis is straight forward scenario and the handling is also simple.\nScenario:2\n==========\nAll the regions R1, R2 and R3 are offlined but the Active Master went down when the zookeeper state for the table was in DISABLING.\nR1- Offlined\nR2- Offlined\t\t\tT1-DISABLING\nR3-Offlined\nScenario:3\n========\nThe regions R1, R2 and R3 are not yet offlined and the Active Master went down when the zookeeper state was DISABLING.\nR1- Online\nR2- Online\t\tT1-DISABLING\nR3-Online\nScenario:4\n========\nThe regions R1, R2 are offlined and R3 are not yet offlined and the Active Master went down when the zookeeper state was DISABLING.\nR1- offlined\nR2- offlined\t\tT1-DISABLING\nR3-Online\n\n{noformat}\n\n","created":"2011-07-06T10:59:43.613+0000"},{"body":"The above looks good to me Ramkrishna.","created":"2011-07-07T23:18:59.838+0000"},{"body":"This is the test for the scenario 1\nwhere the table state is DISABLED","created":"2011-07-08T07:19:45.715+0000"},{"body":"The .bmp files has the sequence of changes that needs to be done.\n\nThe other 3 scenarios is difficult to reproduce through testcase.\n\nPlease let me know if the solution is fine.\nProvide your comments and any scenarios needs to be verified.\nI am working on the patch and testing it. Will upload the patch sooner.","created":"2011-07-08T07:28:41.571+0000"},{"body":"{code}\n+ disabledTableRegions.add(disablingTableName);\n{code}\nLooks like disabledTableRegions should be renamed disablingTables.","created":"2011-07-08T15:28:18.056+0000"},{"body":"Yes that can be renamed. I will change it. \nAny scenarios to be verified Ted?","created":"2011-07-08T18:28:11.355+0000"},{"body":"Running your unit test, I got:\n{code}\ntestForCheckingIfEnableAndDisableWorksFineAfterSwitch(org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable) Time elapsed: 24.594 sec <<< ERROR!\njava.lang.reflect.UndeclaredThrowableException\n at $Proxy12.isMasterRunning(Unknown Source)\n at org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.getMaster(HConnectionManager.java:553)\n at org.apache.hadoop.hbase.client.HBaseAdmin.(HBaseAdmin.java:101)\n at org.apache.hadoop.hbase.HBaseTestingUtility.getHBaseAdmin(HBaseTestingUtility.java:1162)\n at org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable.testForCheckingIfEnableAndDisableWorksFineAfterSwitch(TestMasterRestartAfterDisablingTable.java:105)\n{code}\nI think we should wait for backup master to come up before doing:\n{code}\n log(\"Enabling table\\n\");\n TEST_UTIL.getHBaseAdmin().enableTable(table);\n{code}\n","created":"2011-07-08T20:39:30.630+0000"},{"body":"Patch for TRUNK.\nramkrishna's patch seems to be based on an old version of 0.90 which wouldn't apply cleanly on 0.90 branch\n\nAlso the indentation was off on several lines.\nPlease use the following for new files added:\n{code}\n+ * Copyright 2011 The Apache Software Foundation\n{code} ","created":"2011-07-08T21:09:15.137+0000"},{"body":"This test failure is new:\n{code}\ntestRPCException(org.apache.hadoop.hbase.master.TestHMasterRPCException) Time elapsed: 1.543 sec <<< FAILURE!\njava.lang.AssertionError: Unexpected throwable: java.net.SocketTimeoutException: Call to us01-ciqps1-grid06.carrieriq.com/10.202.50.106:19386 failed on socket timeout exception: java.net.SocketTimeoutException: 100 millis timeout while waiting for channel to be ready for read. ch : java.nio.channels.SocketChannel[connected local=/10.202.50.106:48181 remote=us01-ciqps1-grid06.carrieriq.com/10.202.50.106:19386]\n at org.junit.Assert.fail(Assert.java:91)\n at org.apache.hadoop.hbase.master.TestHMasterRPCException.testRPCException(TestHMasterRPCException.java:59)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n{code}","created":"2011-07-09T00:08:15.986+0000"},{"body":"Hi ted\nI had taken the 0.90 branch code\nand applied the patch on that.\nSo do I need to apply patch only\non the trunk? \nAlso the test case ran cleanly in\nmy setup anyway will check it\nSorry for the mistakes\nI will rectify them .","created":"2011-07-09T02:30:43.662+0000"},{"body":"When I tried to apply your patch on 0.90 branch:\n{noformat}\ntyu-mbp:90hbase tyu$ patch -p0 -i HBASE-4052.patch \n(Stripping trailing CRs from patch.)\npatching file src/main/java/org/apache/hadoop/hbase/master/AssignmentManager.java\nHunk #1 succeeded at 59 with fuzz 1 (offset -1 lines).\nHunk #2 FAILED at 1152.\nHunk #3 FAILED at 1445.\nHunk #4 FAILED at 1460.\nHunk #5 succeeded at 1570 (offset 86 lines).\n3 out of 5 hunks FAILED -- saving rejects to file src/main/java/org/apache/hadoop/hbase/master/AssignmentManager.java.rej\n(Stripping trailing CRs from patch.)\npatching file src/test/java/org/apache/hadoop/hbase/master/TestMasterRestartAfterDisablingTable.java\n{noformat}\nPlease come up with clean patch for 0.90 branch. There have been 39 bug fixes after release of 0.90.3\n\nSince this is a bug, the fix will go to both 0.90 and TRUNK.","created":"2011-07-09T02:46:37.735+0000"},{"body":"The patch HBASE-4052-1-0.90_1.patch is for 0.90 branch.\nThe patch HBASE-4052-1-trunk_1.patch is for trunk branch.\nThe patch HBASE-4052-1-_TestCode.patch is the test code.","created":"2011-07-11T13:40:14.559+0000"},{"body":"I still got the following error (in TRUNK):\n{noformat}\ntestForCheckingIfEnableAndDisableWorksFineAfterSwitch(org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable) Time elapsed: 23.153 sec <<< ERROR!\njava.lang.reflect.UndeclaredThrowableException\n at $Proxy12.isMasterRunning(Unknown Source)\n at org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.getMaster(HConnectionManager.java:553)\n at org.apache.hadoop.hbase.client.HBaseAdmin.(HBaseAdmin.java:101)\n at org.apache.hadoop.hbase.HBaseTestingUtility.getHBaseAdmin(HBaseTestingUtility.java:1162)\n at org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable.testForCheckingIfEnableAndDisableWorksFineAfterSwitch(TestMasterRestartAfterDisablingTable.java:100)\n...\nCaused by: java.io.IOException: Connection reset by peer\n at sun.nio.ch.FileDispatcher.read0(Native Method)\n at sun.nio.ch.SocketDispatcher.read(SocketDispatcher.java:21)\n at sun.nio.ch.IOUtil.readIntoNativeBuffer(IOUtil.java:202)\n at sun.nio.ch.IOUtil.read(IOUtil.java:175)\n{noformat}\n","created":"2011-07-11T16:15:40.392+0000"},{"body":"Or the test hangs:\n{noformat}\n\"main\" prio=5 tid=103000800 nid=0x100601000 in Object.wait() [1005fe000]\n java.lang.Thread.State: WAITING (on object monitor)\n at java.lang.Object.wait(Native Method)\n - waiting on <7a46c9828> (a org.apache.hadoop.hbase.ipc.HBaseClient$Call)\n at java.lang.Object.wait(Object.java:485)\n at org.apache.hadoop.hbase.ipc.HBaseClient.call(HBaseClient.java:835)\n - locked <7a46c9828> (a org.apache.hadoop.hbase.ipc.HBaseClient$Call)\n at org.apache.hadoop.hbase.ipc.WritableRpcEngine$Invoker.invoke(WritableRpcEngine.java:142)\n at $Proxy12.isMasterRunning(Unknown Source)\n at org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.getMaster(HConnectionManager.java:553)\n at org.apache.hadoop.hbase.client.HBaseAdmin.(HBaseAdmin.java:101)\n at org.apache.hadoop.hbase.HBaseTestingUtility.getHBaseAdmin(HBaseTestingUtility.java:1162)\n at org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable.testForCheckingIfEnableAndDisableWorksFineAfterSwitch(TestMasterRestartAfterDisablingTable.java:100)\n{noformat}","created":"2011-07-11T16:31:19.758+0000"},{"body":"Hi Ted,\nI think the test case should work fine in the 0.90 branch.\nIn the trunk when we do a switch the RS is not able to connect to the new Master and we get \n2011-07-12 11:01:22,602 INFO org.apache.hadoop.hbase.master.ServerManager: Waiting on regionserver(s) to checkin\n2011-07-12 11:01:25,602 INFO org.apache.hadoop.hbase.master.ServerManager: Waiting on regionserver(s) to checkin\n2011-07-12 11:01:28,602 INFO org.apache.hadoop.hbase.master.ServerManager: Waiting on regionserver(s) to checkin\n\nSo when this test case is executing also we get the same problem.\n\nI will try to find why this behaviour is found in trunk and not in 0.90 branch. Can you please tell me if the\ntest case is passing in 0.90 branch. Because i have verified locally by running all the test cases in 0.90 branch and the testcases were passing.\n{noformat}\nRunning org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 26.148 sec\nRunning org.apache.hadoop.hbase.regionserver.TestFSErrorsExposed\nTests run: 3, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 45.383 sec\nRunning org.apache.hadoop.hbase.client.replication.TestReplicationAdmin\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 1.044 sec\nRunning org.apache.hadoop.hbase.regionserver.TestScanDeleteTracker\nTests run: 6, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 0.234 sec\nRunning org.apache.hadoop.hbase.client.TestMetaScanner\n\n{noformat}\n","created":"2011-07-12T05:32:25.778+0000"},{"body":"TestMasterRestartAfterDisablingTable passed for 0.90 branch.\n\nI think in TRUNK, the failure may be related to hung TestRestartCluster. See https://builds.apache.org/view/G-L/view/HBase/job/HBase-TRUNK/lastCompletedBuild/artifact/trunk/target/surefire-reports/org.apache.hadoop.hbase.master.TestRestartCluster-output.txt","created":"2011-07-12T06:10:14.075+0000"},{"body":"Thanks Ted.\n\nI will try looking into the issue in trunk why the restart or switch is not working.\nIs there any other issue in the patch or solution? or can it be committed ?","created":"2011-07-12T06:27:16.966+0000"},{"body":"Patch for 0.90 looks good.\nI think patch for TRUNK should be applied around same time as patch for 0.90 is applied.\n\nLet's spend some more time on TRUNK although the cause for region server checkin problem may be somewhere else.","created":"2011-07-12T06:35:20.871+0000"},{"body":"Hi Ted, \nI got the reason why we get the following error while executing the test case \n{noformat}\ntestForCheckingIfEnableAndDisableWorksFineAfterSwitch(org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable) Time elapsed: 23.153 sec <<< ERROR!\njava.lang.reflect.UndeclaredThrowableException\n at $Proxy12.isMasterRunning(Unknown Source)\n at org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.getMaster(HConnectionManager.java:553)\n at org.apache.hadoop.hbase.client.HBaseAdmin.(HBaseAdmin.java:101)\n at org.apache.hadoop.hbase.HBaseTestingUtility.getHBaseAdmin(HBaseTestingUtility.java:1162)\n at org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTab\n{noformat}\n\nI think this is also a bug in HBaseAdmin.getConnection().\n\nThe reason is \nIn HConnectionManager the connection object is cached based on the HConnectionKey.\nThe equals method checks the value of the CONNECTION PROPERTIES.\n\nSuppose if we do a restart/switch of the master and again try to do an enable table operation then in the test code we will create a new HBaseAdmin object.\nBut the connection that the Admin creates to the Master is taken from the cache though it is a new connection.\n\nHere none of the values in the CONNECTION PROPERTIES is changed so we get the same connection object when the previous master was active and hence though the master has been restarted we get the old active master address and hence an exception is thrown.\nWorkaround:\n==========\nSo in order to pass the test case we change the value of one of the CONNECTION PROPERTIES so that the cached connection object is not returned.\n\nI reverted the HBASE-4003 and this test case passed with the above change.\nNot sure of the reason why RS doesnot checkin.\n","created":"2011-07-12T12:08:34.459+0000"},{"body":"The following test consistently failed on Linux with patch for 0.90:\n{noformat}\nqueueFailover(org.apache.hadoop.hbase.replication.TestReplication) Time elapsed: 85.997 sec <<< FAILURE!\njava.lang.AssertionError: Waited too much time for queueFailover replication\n at org.junit.Assert.fail(Assert.java:91)\n at org.apache.hadoop.hbase.replication.TestReplication.queueFailover(TestReplication.java:572)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n{noformat}","created":"2011-07-12T19:18:19.110+0000"},{"body":"Ted, will look into the failure of org.apache.hadoop.hbase.replication.TestReplication.\nBut my local test bed shows success for the same testcase..\n{noformat}\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 0.031 sec\nRunning org.apache.hadoop.hbase.rest.TestStatusResource\nTests run: 2, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 23.785 sec\nRunning org.apache.hadoop.hbase.executor.TestExecutorService\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 2.324 sec\nRunning org.apache.hadoop.hbase.client.TestFromClientSide\nTests run: 42, Failures: 0, Errors: 0, Skipped: 3, Time elapsed: 279.871 sec\nRunning org.apache.hadoop.hbase.replication.TestReplication\nTests run: 7, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 213.828 sec\n{noformat}\n\nAnyways will look into it if this patch is the root cause.","created":"2011-07-13T04:07:28.584+0000"},{"body":"Patch looks good to me Ramkrishna. I tried to run your test but was getting NPE on trunk unrelated seemingly to your patch; I think this breaks tests 'HEAD is now at 46924a1... HBASE-3904 Addendum that fixes number of retries (Ita Pai)'","created":"2011-07-13T06:04:45.839+0000"},{"body":"I applied an addendum for HBASE-3904 v6 to TRUNK. NPE is fixed.\nHowever, HBASE-4087 is required for the new unit test to pass.","created":"2011-07-13T09:01:32.127+0000"},{"body":"Integrated to branch and TRUNK.\n\nThanks for the review Stack.\n\n@ramkrishna:\nIn the future, please concatenate test file patch to main patch.","created":"2011-07-16T02:00:41.420+0000"},{"body":"Thanks for the review Ted and Stack.\n\n","created":"2011-07-18T04:00:09.096+0000"},{"body":"@Ted, one small correction in the patch that was applied \n\nActually the patch catches the NotServingRegionException as RemoteException and then\nchecks for the instanceof NotServingRegionException .\n{noformat}\ncatch (Throwable t) {\n if (t instanceof RemoteException) {\n t = ((RemoteException)t).unwrapRemoteException();\n }\n{noformat}\nBut the latest AssignmentManager.java in trunk shows that it was applied in \n{noformat}\ncatch (NotServingRegionException nsre) {\n LOG.info(\"Server \" + server + \" returned \" + nsre + \" for \" +\n region.getEncodedName());\n // Presume that master has stale data. Presume remote side just split.\n // Presume that the split message when it comes in will fix up the master's\n // in memory cluster state.\n if (checkIfRegionBelongsToDisabling(region)) {\n{noformat}\nSo the scenario of partial disabling is failing to recover back to DISABLED state.\nCan you please check and reapply the patch. \n","created":"2011-07-18T07:45:02.650+0000"},{"body":"Handling of NotServingRegionException has been moved to the last catch block.","created":"2011-07-18T11:46:09.595+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T11:52:26.262+0000"}],"conversations":[{"body":"Following is the scenario:\n\nStart RS and Active and standby masters\nCreate table and insert data.\nDisable the table.\nStop the active master and switch to the standby master.\nNow enable the table.\nDo a scan on the enabled table.\nNotServingRegionException is Thrown.\n\nBut the same works well when we dont switch the master.\n","from":"reporter","subject":"Enabling a table after master switch does not allow table scan, throwing NotServingRegionException"},{"body":"As per my analysis the problem is that,\n\nWhen we do a remove all the online regions are closed.\nIn the Enable table flow \n{noformat}\n private List regionsToAssign(final List regionsInMeta)\n throws IOException {\n final List onlineRegions =\n this.assignmentManager.getRegionsOfTable(tableName);\n regionsInMeta.removeAll(onlineRegions);\n return regionsInMeta;\n }\n{noformat}\nWe remove the regions if it is already online.\n\nBut as per the bug, enable is called after switching,\n\nSo while the standby master becomes active, in the rebuildUserRegion api,\nwe add all the regions from the Meta and we dont check if it is already disabled.\n\nSo when the flow comes to enable (regionsToAssign()) we consider the region to be onlined already. \nFinally when we try to scan the table, we get NotServingRegionException.\n\nCorrect me if am wrong in my analysis. \n","from":"developer"},{"body":"What you say seems plausible Ramkrishna. I traced your reasoning above and yes, because rebuildUserRegion adds all regions without regard to whether table is enabled/disabled, when it comes to run the enable, when it asks what regions to enable, they all look as though they are already online because of what getRegionsOfTable returns (all regions that make up the table pulled from this.regions over in AM).\n\nGood stuff.\n\nDo you have a patch Ramkrishna?","from":"developer"},{"body":"Thanks Stack for your comments.\nI am working on the patch. \nI have 2 scenarios to consider\n1. The AM may get killed when the state is in DISABLING but still the regions are not closed\n2. The AM may get killed when the state is in DISABLING but the regions are closed.\n\nSo can we check both for DISABLING and DISABLED state.\nI will provide my patch ASAP.\n\nThanks","from":"developer"},{"body":"I have some doubts and like to get some suggestion before proceeding.\n \nFollowing scenarios needs to be considered.\n Scenario 1:\n===========\nAll the regions are disabled and the state in zookeeper is DISABLED.\nScenario:2\n==========\n \nThe regions are offlined but the AM went down when the zookeeper state was DISABLING.\n \nScenario:3\n=========\n \nThe regions are not yet offlined(or only few regions are offlined) and the AM went down when the zookeeper state was DISABLING.\n \nNow when we do a switch of the master or on restart scenario of master,\nhow can we decide which regions were offlined and which are not.\n \nThough we can get the state of the table as either DISABLED or DISABLING, region wise i am not able to infer in what state the region is.\n \nSo what brings me to get this info is \n \nThe soln should be like we need to check for the state of the table while populating the regions map in master startup.\n \nChecking only for DISABLED state:\n==========================\nCheck for disabled state and those regions that are not in the DISABLED state add it into the regions map in master startup.\n \nIf i check only for the DISABLED state and if the table is in DISABLING state and \n \nafter master retry (or switch) if i try to enable then we will not be able to scan the table because while enabling none of the regions will be enabled\nas the regions in META table and the regions that i have populated in the regions map are same.\nSo I will be getting the same issue as in the description of the defect.\n \nChecking for DISABLED and DISABLING state:\n===================================\nif i check the state of the zookeeper for DISABLED and DISABLING and while restart of master(switch) only those regions which are not in DISABLED or DISABLING state is populated.\nWhen i again try to enable the region if the region was not offlined as part of disable flow(Scenario:3), the waitUntilDone in BulkAssigner is not aware that the region was \nalready onlined and keeps on waiting as the waitUntilDone() sees for the number of regions to become online from the regions map and the actual count it gets from the meta table.\nThis makes enable to go in a loop.\n \nAm i clear with the problem? so is it like before enabling any table do we need to check the state of the table and if it is DISABLING make all those regions to go to\noffline mode.","from":"developer"},{"body":"I think one source of confusion in the above description is that DISABLED and DISABLING states are TableState.\nThere is no concept of disabling/disabled region.\n\nI would opt for Checking for DISABLED and DISABLING TableState.\n\nI think the BulkAssigner in this case refers to EnableTableHandler.BulkEnabler which depends on assignmentManager.getRegionsOfTable() for the number of online regions.\n\nYou need to write a new method, e.g. assignmentManager.getOnlineRegionsOfTable(), and call it in place of assignmentManager.getRegionsOfTable() in EnableTableHandler.regionsToAssign()","from":"developer"},{"body":"Hi,\nSorry for using wrong terminologies. I was aware that the Tables are only disabled and regions are only onlined or offlined.\nI would like to attach the scenarios that needs to be addressed as part of this bug\n\n{noformat}\nIf a table T1 has three regions R1, R2 and R3 \nWe issue a disable command for T1\nScenario 1:\n===========\nAll the regions R1, R2 and R3 are offlined and the state in zookeeper for the table is DISABLED and then the Active master went down,\nR1- Offlined\nR2- Offlined\t\t\tT1-DISABLED\nR3-Offlined\n\tThis is straight forward scenario and the handling is also simple.\nScenario:2\n==========\nAll the regions R1, R2 and R3 are offlined but the Active Master went down when the zookeeper state for the table was in DISABLING.\nR1- Offlined\nR2- Offlined\t\t\tT1-DISABLING\nR3-Offlined\nScenario:3\n========\nThe regions R1, R2 and R3 are not yet offlined and the Active Master went down when the zookeeper state was DISABLING.\nR1- Online\nR2- Online\t\tT1-DISABLING\nR3-Online\nScenario:4\n========\nThe regions R1, R2 are offlined and R3 are not yet offlined and the Active Master went down when the zookeeper state was DISABLING.\nR1- offlined\nR2- offlined\t\tT1-DISABLING\nR3-Online\n\n{noformat}\n\n","from":"developer"},{"body":"The above looks good to me Ramkrishna.","from":"developer"},{"body":"This is the test for the scenario 1\nwhere the table state is DISABLED","from":"developer"},{"body":"The .bmp files has the sequence of changes that needs to be done.\n\nThe other 3 scenarios is difficult to reproduce through testcase.\n\nPlease let me know if the solution is fine.\nProvide your comments and any scenarios needs to be verified.\nI am working on the patch and testing it. Will upload the patch sooner.","from":"developer"},{"body":"{code}\n+ disabledTableRegions.add(disablingTableName);\n{code}\nLooks like disabledTableRegions should be renamed disablingTables.","from":"developer"},{"body":"Yes that can be renamed. I will change it. \nAny scenarios to be verified Ted?","from":"developer"},{"body":"Running your unit test, I got:\n{code}\ntestForCheckingIfEnableAndDisableWorksFineAfterSwitch(org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable) Time elapsed: 24.594 sec <<< ERROR!\njava.lang.reflect.UndeclaredThrowableException\n at $Proxy12.isMasterRunning(Unknown Source)\n at org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.getMaster(HConnectionManager.java:553)\n at org.apache.hadoop.hbase.client.HBaseAdmin.(HBaseAdmin.java:101)\n at org.apache.hadoop.hbase.HBaseTestingUtility.getHBaseAdmin(HBaseTestingUtility.java:1162)\n at org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable.testForCheckingIfEnableAndDisableWorksFineAfterSwitch(TestMasterRestartAfterDisablingTable.java:105)\n{code}\nI think we should wait for backup master to come up before doing:\n{code}\n log(\"Enabling table\\n\");\n TEST_UTIL.getHBaseAdmin().enableTable(table);\n{code}\n","from":"developer"},{"body":"Patch for TRUNK.\nramkrishna's patch seems to be based on an old version of 0.90 which wouldn't apply cleanly on 0.90 branch\n\nAlso the indentation was off on several lines.\nPlease use the following for new files added:\n{code}\n+ * Copyright 2011 The Apache Software Foundation\n{code} ","from":"developer"},{"body":"This test failure is new:\n{code}\ntestRPCException(org.apache.hadoop.hbase.master.TestHMasterRPCException) Time elapsed: 1.543 sec <<< FAILURE!\njava.lang.AssertionError: Unexpected throwable: java.net.SocketTimeoutException: Call to us01-ciqps1-grid06.carrieriq.com/10.202.50.106:19386 failed on socket timeout exception: java.net.SocketTimeoutException: 100 millis timeout while waiting for channel to be ready for read. ch : java.nio.channels.SocketChannel[connected local=/10.202.50.106:48181 remote=us01-ciqps1-grid06.carrieriq.com/10.202.50.106:19386]\n at org.junit.Assert.fail(Assert.java:91)\n at org.apache.hadoop.hbase.master.TestHMasterRPCException.testRPCException(TestHMasterRPCException.java:59)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n{code}","from":"developer"},{"body":"Hi ted\nI had taken the 0.90 branch code\nand applied the patch on that.\nSo do I need to apply patch only\non the trunk? \nAlso the test case ran cleanly in\nmy setup anyway will check it\nSorry for the mistakes\nI will rectify them .","from":"developer"},{"body":"When I tried to apply your patch on 0.90 branch:\n{noformat}\ntyu-mbp:90hbase tyu$ patch -p0 -i HBASE-4052.patch \n(Stripping trailing CRs from patch.)\npatching file src/main/java/org/apache/hadoop/hbase/master/AssignmentManager.java\nHunk #1 succeeded at 59 with fuzz 1 (offset -1 lines).\nHunk #2 FAILED at 1152.\nHunk #3 FAILED at 1445.\nHunk #4 FAILED at 1460.\nHunk #5 succeeded at 1570 (offset 86 lines).\n3 out of 5 hunks FAILED -- saving rejects to file src/main/java/org/apache/hadoop/hbase/master/AssignmentManager.java.rej\n(Stripping trailing CRs from patch.)\npatching file src/test/java/org/apache/hadoop/hbase/master/TestMasterRestartAfterDisablingTable.java\n{noformat}\nPlease come up with clean patch for 0.90 branch. There have been 39 bug fixes after release of 0.90.3\n\nSince this is a bug, the fix will go to both 0.90 and TRUNK.","from":"developer"},{"body":"The patch HBASE-4052-1-0.90_1.patch is for 0.90 branch.\nThe patch HBASE-4052-1-trunk_1.patch is for trunk branch.\nThe patch HBASE-4052-1-_TestCode.patch is the test code.","from":"developer"},{"body":"I still got the following error (in TRUNK):\n{noformat}\ntestForCheckingIfEnableAndDisableWorksFineAfterSwitch(org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable) Time elapsed: 23.153 sec <<< ERROR!\njava.lang.reflect.UndeclaredThrowableException\n at $Proxy12.isMasterRunning(Unknown Source)\n at org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.getMaster(HConnectionManager.java:553)\n at org.apache.hadoop.hbase.client.HBaseAdmin.(HBaseAdmin.java:101)\n at org.apache.hadoop.hbase.HBaseTestingUtility.getHBaseAdmin(HBaseTestingUtility.java:1162)\n at org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable.testForCheckingIfEnableAndDisableWorksFineAfterSwitch(TestMasterRestartAfterDisablingTable.java:100)\n...\nCaused by: java.io.IOException: Connection reset by peer\n at sun.nio.ch.FileDispatcher.read0(Native Method)\n at sun.nio.ch.SocketDispatcher.read(SocketDispatcher.java:21)\n at sun.nio.ch.IOUtil.readIntoNativeBuffer(IOUtil.java:202)\n at sun.nio.ch.IOUtil.read(IOUtil.java:175)\n{noformat}\n","from":"developer"},{"body":"Or the test hangs:\n{noformat}\n\"main\" prio=5 tid=103000800 nid=0x100601000 in Object.wait() [1005fe000]\n java.lang.Thread.State: WAITING (on object monitor)\n at java.lang.Object.wait(Native Method)\n - waiting on <7a46c9828> (a org.apache.hadoop.hbase.ipc.HBaseClient$Call)\n at java.lang.Object.wait(Object.java:485)\n at org.apache.hadoop.hbase.ipc.HBaseClient.call(HBaseClient.java:835)\n - locked <7a46c9828> (a org.apache.hadoop.hbase.ipc.HBaseClient$Call)\n at org.apache.hadoop.hbase.ipc.WritableRpcEngine$Invoker.invoke(WritableRpcEngine.java:142)\n at $Proxy12.isMasterRunning(Unknown Source)\n at org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.getMaster(HConnectionManager.java:553)\n at org.apache.hadoop.hbase.client.HBaseAdmin.(HBaseAdmin.java:101)\n at org.apache.hadoop.hbase.HBaseTestingUtility.getHBaseAdmin(HBaseTestingUtility.java:1162)\n at org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable.testForCheckingIfEnableAndDisableWorksFineAfterSwitch(TestMasterRestartAfterDisablingTable.java:100)\n{noformat}","from":"developer"},{"body":"Hi Ted,\nI think the test case should work fine in the 0.90 branch.\nIn the trunk when we do a switch the RS is not able to connect to the new Master and we get \n2011-07-12 11:01:22,602 INFO org.apache.hadoop.hbase.master.ServerManager: Waiting on regionserver(s) to checkin\n2011-07-12 11:01:25,602 INFO org.apache.hadoop.hbase.master.ServerManager: Waiting on regionserver(s) to checkin\n2011-07-12 11:01:28,602 INFO org.apache.hadoop.hbase.master.ServerManager: Waiting on regionserver(s) to checkin\n\nSo when this test case is executing also we get the same problem.\n\nI will try to find why this behaviour is found in trunk and not in 0.90 branch. Can you please tell me if the\ntest case is passing in 0.90 branch. Because i have verified locally by running all the test cases in 0.90 branch and the testcases were passing.\n{noformat}\nRunning org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 26.148 sec\nRunning org.apache.hadoop.hbase.regionserver.TestFSErrorsExposed\nTests run: 3, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 45.383 sec\nRunning org.apache.hadoop.hbase.client.replication.TestReplicationAdmin\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 1.044 sec\nRunning org.apache.hadoop.hbase.regionserver.TestScanDeleteTracker\nTests run: 6, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 0.234 sec\nRunning org.apache.hadoop.hbase.client.TestMetaScanner\n\n{noformat}\n","from":"developer"},{"body":"TestMasterRestartAfterDisablingTable passed for 0.90 branch.\n\nI think in TRUNK, the failure may be related to hung TestRestartCluster. See https://builds.apache.org/view/G-L/view/HBase/job/HBase-TRUNK/lastCompletedBuild/artifact/trunk/target/surefire-reports/org.apache.hadoop.hbase.master.TestRestartCluster-output.txt","from":"developer"},{"body":"Thanks Ted.\n\nI will try looking into the issue in trunk why the restart or switch is not working.\nIs there any other issue in the patch or solution? or can it be committed ?","from":"developer"},{"body":"Patch for 0.90 looks good.\nI think patch for TRUNK should be applied around same time as patch for 0.90 is applied.\n\nLet's spend some more time on TRUNK although the cause for region server checkin problem may be somewhere else.","from":"developer"},{"body":"Hi Ted, \nI got the reason why we get the following error while executing the test case \n{noformat}\ntestForCheckingIfEnableAndDisableWorksFineAfterSwitch(org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTable) Time elapsed: 23.153 sec <<< ERROR!\njava.lang.reflect.UndeclaredThrowableException\n at $Proxy12.isMasterRunning(Unknown Source)\n at org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.getMaster(HConnectionManager.java:553)\n at org.apache.hadoop.hbase.client.HBaseAdmin.(HBaseAdmin.java:101)\n at org.apache.hadoop.hbase.HBaseTestingUtility.getHBaseAdmin(HBaseTestingUtility.java:1162)\n at org.apache.hadoop.hbase.master.TestMasterRestartAfterDisablingTab\n{noformat}\n\nI think this is also a bug in HBaseAdmin.getConnection().\n\nThe reason is \nIn HConnectionManager the connection object is cached based on the HConnectionKey.\nThe equals method checks the value of the CONNECTION PROPERTIES.\n\nSuppose if we do a restart/switch of the master and again try to do an enable table operation then in the test code we will create a new HBaseAdmin object.\nBut the connection that the Admin creates to the Master is taken from the cache though it is a new connection.\n\nHere none of the values in the CONNECTION PROPERTIES is changed so we get the same connection object when the previous master was active and hence though the master has been restarted we get the old active master address and hence an exception is thrown.\nWorkaround:\n==========\nSo in order to pass the test case we change the value of one of the CONNECTION PROPERTIES so that the cached connection object is not returned.\n\nI reverted the HBASE-4003 and this test case passed with the above change.\nNot sure of the reason why RS doesnot checkin.\n","from":"developer"},{"body":"The following test consistently failed on Linux with patch for 0.90:\n{noformat}\nqueueFailover(org.apache.hadoop.hbase.replication.TestReplication) Time elapsed: 85.997 sec <<< FAILURE!\njava.lang.AssertionError: Waited too much time for queueFailover replication\n at org.junit.Assert.fail(Assert.java:91)\n at org.apache.hadoop.hbase.replication.TestReplication.queueFailover(TestReplication.java:572)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n{noformat}","from":"developer"},{"body":"Ted, will look into the failure of org.apache.hadoop.hbase.replication.TestReplication.\nBut my local test bed shows success for the same testcase..\n{noformat}\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 0.031 sec\nRunning org.apache.hadoop.hbase.rest.TestStatusResource\nTests run: 2, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 23.785 sec\nRunning org.apache.hadoop.hbase.executor.TestExecutorService\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 2.324 sec\nRunning org.apache.hadoop.hbase.client.TestFromClientSide\nTests run: 42, Failures: 0, Errors: 0, Skipped: 3, Time elapsed: 279.871 sec\nRunning org.apache.hadoop.hbase.replication.TestReplication\nTests run: 7, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 213.828 sec\n{noformat}\n\nAnyways will look into it if this patch is the root cause.","from":"developer"},{"body":"Patch looks good to me Ramkrishna. I tried to run your test but was getting NPE on trunk unrelated seemingly to your patch; I think this breaks tests 'HEAD is now at 46924a1... HBASE-3904 Addendum that fixes number of retries (Ita Pai)'","from":"developer"},{"body":"I applied an addendum for HBASE-3904 v6 to TRUNK. NPE is fixed.\nHowever, HBASE-4087 is required for the new unit test to pass.","from":"developer"},{"body":"Integrated to branch and TRUNK.\n\nThanks for the review Stack.\n\n@ramkrishna:\nIn the future, please concatenate test file patch to main patch.","from":"developer"},{"body":"Thanks for the review Ted and Stack.\n\n","from":"developer"},{"body":"@Ted, one small correction in the patch that was applied \n\nActually the patch catches the NotServingRegionException as RemoteException and then\nchecks for the instanceof NotServingRegionException .\n{noformat}\ncatch (Throwable t) {\n if (t instanceof RemoteException) {\n t = ((RemoteException)t).unwrapRemoteException();\n }\n{noformat}\nBut the latest AssignmentManager.java in trunk shows that it was applied in \n{noformat}\ncatch (NotServingRegionException nsre) {\n LOG.info(\"Server \" + server + \" returned \" + nsre + \" for \" +\n region.getEncodedName());\n // Presume that master has stale data. Presume remote side just split.\n // Presume that the split message when it comes in will fix up the master's\n // in memory cluster state.\n if (checkIfRegionBelongsToDisabling(region)) {\n{noformat}\nSo the scenario of partial disabling is failing to recover back to DISABLED state.\nCan you please check and reapply the patch. \n","from":"developer"},{"body":"Handling of NotServingRegionException has been moved to the last catch block.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2011-07-01T04:47:30.000+0000","description":"Following is the scenario:\n\nStart RS and Active and standby masters\nCreate table and insert data.\nDisable the table.\nStop the active master and switch to the standby master.\nNow enable the table.\nDo a scan on the enabled table.\nNotServingRegionException is Thrown.\n\nBut the same works well when we dont switch the master.\n","issue_id":"12512427","key":"HBASE-4052","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2011-07-16T02:40:20.000+0000","role":"fixed_distractor","summary":"Enabling a table after master switch does not allow table scan, throwing NotServingRegionException"} {"case_id":"12513170","cluster":"DISTRACTOR-HBASE-4077","comments":[{"body":"Why not check against lid being null to prevent NPE ?","created":"2011-07-07T18:59:22.447+0000"},{"body":"I followed the same pattern used in the other update functions in HRegion, which include:\n\ncheckAndMutate\nput\nincrement\nincrementColumnValue\n\nI originally did a check against null, but after looking through the code, decided consistency with the other functions was something I liked better. I'm in no way married to this, so I am game to change it to a null check if that is what folks want. Let me know.","created":"2011-07-07T20:38:53.760+0000"},{"body":"My thinking was to make WrongRegionException prominent in region server log.","created":"2011-07-07T21:02:08.682+0000"},{"body":"WrongRegionException actually is and will still be logged as an ERROR level log with the patch I committed. I only posted the NPE exception in the description of this ticket, but the WrongRegionException is also in the logs as well.","created":"2011-07-07T21:06:49.681+0000"},{"body":"Looking at the code again, delete() was the only method that didn't follow the pattern.\nIf WrongRegionException was thrown from getLock(), the inner finally block would be skipped, along with it the NPE.\n\nThanks for the patch Adam.","created":"2011-07-07T21:26:38.818+0000"},{"body":"No problem Ted. My pleasure!","created":"2011-07-07T21:30:49.106+0000"},{"body":"@Ted You going to apply? (+1 from me).","created":"2011-07-07T21:33:49.282+0000"},{"body":"Integrated to branch and TRUNK.\nI shortened indentation to 2 spaces.\n\nThanks for the re view Stack.","created":"2011-07-07T21:42:06.765+0000"},{"body":"This is an interesting test failure:\nhttps://builds.apache.org/view/G-L/view/HBase/job/hbase-0.90/lastCompletedBuild/testReport/org.apache.hadoop.hbase/TestZooKeeper/testClientSessionExpired/\n\nI refreshed 0.90 branch on a Linux box and TestZooKeeper passes standalone.","created":"2011-07-08T03:45:47.219+0000"},{"body":"https://builds.apache.org/view/G-L/view/HBase/job/hbase-0.90/226/console was successful.","created":"2011-07-08T04:15:54.681+0000"},{"body":"Yes. Its wacky. The testing didn't get off the ground. Chalk it up to jenkins randomness.","created":"2011-07-08T05:30:56.056+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T11:55:30.765+0000"}],"conversations":[{"body":"In the HRegion.delete function, If getLock throws a WrongRegionException, no lock id is ever returned, yet in the finally block, it tries to release the row lock using that lock id (which is null). This causes an NPE in the finally clause, and the closeRegionOperation() to never execute, keeping a read lock open forever.\n\nERROR org.apache.hadoop.hbase.regionserver.HRegionServer: \njava.lang.NullPointerException \nat org.apache.hadoop.hbase.util.Bytes.compareTo(Bytes.java:840) \nat org.apache.hadoop.hbase.util.Bytes$ByteArrayComparator.compare(Bytes.java:108) \nat org.apache.hadoop.hbase.util.Bytes$ByteArrayComparator.compare(Bytes.java:100) \nat java.util.TreeMap.getEntryUsingComparator(TreeMap.java:351) \nat java.util.TreeMap.getEntry(TreeMap.java:322) \nat java.util.TreeMap.remove(TreeMap.java:580) \nat java.util.TreeSet.remove(TreeSet.java:259) \nat org.apache.hadoop.hbase.regionserver.HRegion.releaseRowLock(HRegion.java:2145) \nat org.apache.hadoop.hbase.regionserver.HRegion.delete(HRegion.java:1174) \nat org.apache.hadoop.hbase.regionserver.HRegionServer.delete(HRegionServer.java:1914) \nat sun.reflect.GeneratedMethodAccessor22.invoke(Unknown Source) \nat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25) \nat java.lang.reflect.Method.invoke(Method.java:597) \nat org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:570) \nat org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:1039)\n\nWhen the region later attempts to close, the write lock can never be acquired, and the region remains in transition forever.","from":"reporter","subject":"Deadlock if WrongRegionException is thrown from getLock in HRegion.delete"},{"body":"Why not check against lid being null to prevent NPE ?","from":"developer"},{"body":"I followed the same pattern used in the other update functions in HRegion, which include:\n\ncheckAndMutate\nput\nincrement\nincrementColumnValue\n\nI originally did a check against null, but after looking through the code, decided consistency with the other functions was something I liked better. I'm in no way married to this, so I am game to change it to a null check if that is what folks want. Let me know.","from":"developer"},{"body":"My thinking was to make WrongRegionException prominent in region server log.","from":"developer"},{"body":"WrongRegionException actually is and will still be logged as an ERROR level log with the patch I committed. I only posted the NPE exception in the description of this ticket, but the WrongRegionException is also in the logs as well.","from":"developer"},{"body":"Looking at the code again, delete() was the only method that didn't follow the pattern.\nIf WrongRegionException was thrown from getLock(), the inner finally block would be skipped, along with it the NPE.\n\nThanks for the patch Adam.","from":"developer"},{"body":"No problem Ted. My pleasure!","from":"developer"},{"body":"@Ted You going to apply? (+1 from me).","from":"developer"},{"body":"Integrated to branch and TRUNK.\nI shortened indentation to 2 spaces.\n\nThanks for the re view Stack.","from":"developer"},{"body":"This is an interesting test failure:\nhttps://builds.apache.org/view/G-L/view/HBase/job/hbase-0.90/lastCompletedBuild/testReport/org.apache.hadoop.hbase/TestZooKeeper/testClientSessionExpired/\n\nI refreshed 0.90 branch on a Linux box and TestZooKeeper passes standalone.","from":"developer"},{"body":"https://builds.apache.org/view/G-L/view/HBase/job/hbase-0.90/226/console was successful.","from":"developer"},{"body":"Yes. Its wacky. The testing didn't get off the ground. Chalk it up to jenkins randomness.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2011-07-07T18:23:21.000+0000","description":"In the HRegion.delete function, If getLock throws a WrongRegionException, no lock id is ever returned, yet in the finally block, it tries to release the row lock using that lock id (which is null). This causes an NPE in the finally clause, and the closeRegionOperation() to never execute, keeping a read lock open forever.\n\nERROR org.apache.hadoop.hbase.regionserver.HRegionServer: \njava.lang.NullPointerException \nat org.apache.hadoop.hbase.util.Bytes.compareTo(Bytes.java:840) \nat org.apache.hadoop.hbase.util.Bytes$ByteArrayComparator.compare(Bytes.java:108) \nat org.apache.hadoop.hbase.util.Bytes$ByteArrayComparator.compare(Bytes.java:100) \nat java.util.TreeMap.getEntryUsingComparator(TreeMap.java:351) \nat java.util.TreeMap.getEntry(TreeMap.java:322) \nat java.util.TreeMap.remove(TreeMap.java:580) \nat java.util.TreeSet.remove(TreeSet.java:259) \nat org.apache.hadoop.hbase.regionserver.HRegion.releaseRowLock(HRegion.java:2145) \nat org.apache.hadoop.hbase.regionserver.HRegion.delete(HRegion.java:1174) \nat org.apache.hadoop.hbase.regionserver.HRegionServer.delete(HRegionServer.java:1914) \nat sun.reflect.GeneratedMethodAccessor22.invoke(Unknown Source) \nat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25) \nat java.lang.reflect.Method.invoke(Method.java:597) \nat org.apache.hadoop.hbase.ipc.HBaseRPC$Server.call(HBaseRPC.java:570) \nat org.apache.hadoop.hbase.ipc.HBaseServer$Handler.run(HBaseServer.java:1039)\n\nWhen the region later attempts to close, the write lock can never be acquired, and the region remains in transition forever.","issue_id":"12513170","key":"HBASE-4077","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2011-07-07T21:43:36.000+0000","role":"fixed_distractor","summary":"Deadlock if WrongRegionException is thrown from getLock in HRegion.delete"} {"case_id":"12520535","cluster":"DISTRACTOR-HBASE-4277","comments":[{"body":"+1 on patch.","created":"2011-08-29T14:11:25.209+0000"},{"body":"Patch looks good to me (Thanks for doing this RAM). Marking critical on 0.90.5. Before committing, I'd like to check that the addition of new methods to Interface do not break rolling restart.","created":"2011-08-29T18:48:01.125+0000"},{"body":"For rolling restart verification the following steps were taken\n-> Start 2 RS with 0.90.x version one with patch other without patch.\n-> Have a client with the patched version\n-> Call the new api added in this defect for a region in the patched version\n This works fine.\n-> Call a region in the unpatch version\n We get NoSuchMethod found exception.\n\n","created":"2011-09-14T10:33:26.864+0000"},{"body":"This is a 0.90.5 issue, not for 0.92.0. Moving it out.","created":"2011-09-17T00:32:11.160+0000"},{"body":"+1 on commit to 0.90. I brought up a patched server and it was able to take loads fine. I then talked to it with an UNPATCHED client doing moves and gets. Seems fine.","created":"2011-10-26T05:42:45.169+0000"},{"body":"I will see if the patch still applies if not will prepare an updated on and then commit it.\nThanks for your review Stack","created":"2011-10-31T05:42:41.508+0000"},{"body":"The patch applies as is.. If it is ok I can go ahead and commit it. Pls share your comments.","created":"2011-10-31T13:17:22.389+0000"},{"body":"+1 on commit","created":"2011-10-31T16:23:10.245+0000"},{"body":"Tests passes...\n{code}\n testClockSkewDetection(org.apache.hadoop.hbase.master.TestClockSkewDetection): hostname can't be null\n testScanner(org.apache.hadoop.hbase.regionserver.TestScanner): hostname can't be null\n{code}\nThis failure is not due to the patch for this JIRA.\n\n","created":"2011-11-01T06:07:07.984+0000"},{"body":"++1 on commit (smile)","created":"2011-11-01T16:22:49.736+0000"},{"body":"Integrated to 0.90.5. Thanks for your reviews Stack and Ted.","created":"2011-11-02T04:54:02.062+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T11:55:55.188+0000"}],"conversations":[{"body":"As suggested by Stack in HBASE-4217 creating a new issue to provide a patch for 0.90.x version.\n\n\nWe had some sort of an outage this morning due to a few racks losing power, and some regions were left in the following state:\n\nERROR: Region UNKNOWN_REGION on sv4r17s9:60020, key=e32bbe1f48c9b3633c557dc0291b90a3, not on HDFS or in META but deployed on sv4r17s9:60020\n\nThat region was deleted by the master but the region server never got the memo. Right now there's no way to force close it because HRS.closeRegion requires an HRI and the only way to create one is to get it from .META. which in our case doesn't contain a row for that region. Basically we have to wait until that server is dead to get rid of the region and make hbck happy.\n\nThe required change is to have closeRegion accept an encoded name in both HBA (when the RS address is provided) and HRS since it's able to find it anyways from it's list of live regions.\nbq.If a 0.90 version, we maybe should do that in another issue.","from":"reporter","subject":"HRS.closeRegion should be able to close regions with only the encoded name"},{"body":"+1 on patch.","from":"developer"},{"body":"Patch looks good to me (Thanks for doing this RAM). Marking critical on 0.90.5. Before committing, I'd like to check that the addition of new methods to Interface do not break rolling restart.","from":"developer"},{"body":"For rolling restart verification the following steps were taken\n-> Start 2 RS with 0.90.x version one with patch other without patch.\n-> Have a client with the patched version\n-> Call the new api added in this defect for a region in the patched version\n This works fine.\n-> Call a region in the unpatch version\n We get NoSuchMethod found exception.\n\n","from":"developer"},{"body":"This is a 0.90.5 issue, not for 0.92.0. Moving it out.","from":"developer"},{"body":"+1 on commit to 0.90. I brought up a patched server and it was able to take loads fine. I then talked to it with an UNPATCHED client doing moves and gets. Seems fine.","from":"developer"},{"body":"I will see if the patch still applies if not will prepare an updated on and then commit it.\nThanks for your review Stack","from":"developer"},{"body":"The patch applies as is.. If it is ok I can go ahead and commit it. Pls share your comments.","from":"developer"},{"body":"+1 on commit","from":"developer"},{"body":"Tests passes...\n{code}\n testClockSkewDetection(org.apache.hadoop.hbase.master.TestClockSkewDetection): hostname can't be null\n testScanner(org.apache.hadoop.hbase.regionserver.TestScanner): hostname can't be null\n{code}\nThis failure is not due to the patch for this JIRA.\n\n","from":"developer"},{"body":"++1 on commit (smile)","from":"developer"},{"body":"Integrated to 0.90.5. Thanks for your reviews Stack and Ted.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2011-08-29T08:18:08.000+0000","description":"As suggested by Stack in HBASE-4217 creating a new issue to provide a patch for 0.90.x version.\n\n\nWe had some sort of an outage this morning due to a few racks losing power, and some regions were left in the following state:\n\nERROR: Region UNKNOWN_REGION on sv4r17s9:60020, key=e32bbe1f48c9b3633c557dc0291b90a3, not on HDFS or in META but deployed on sv4r17s9:60020\n\nThat region was deleted by the master but the region server never got the memo. Right now there's no way to force close it because HRS.closeRegion requires an HRI and the only way to create one is to get it from .META. which in our case doesn't contain a row for that region. Basically we have to wait until that server is dead to get rid of the region and make hbck happy.\n\nThe required change is to have closeRegion accept an encoded name in both HBA (when the RS address is provided) and HRS since it's able to find it anyways from it's list of live regions.\nbq.If a 0.90 version, we maybe should do that in another issue.","issue_id":"12520535","key":"HBASE-4277","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2011-11-02T04:54:02.000+0000","role":"fixed_distractor","summary":"HRS.closeRegion should be able to close regions with only the encoded name"} {"case_id":"12524808","cluster":"DISTRACTOR-HBASE-4496","comments":[{"body":"I'll come up with a patch.","created":"2011-09-27T16:58:15.004+0000"},{"body":"In a simple test this shaves about 20% of scanning time (with caching switched off), in addition to avoiding LRU and GC churn.","created":"2011-09-27T18:22:13.261+0000"},{"body":"Good find Lars.","created":"2011-09-27T18:36:56.485+0000"},{"body":"This fixes the problem for me.\n\nTestHFileBlock, TestCacheOnWrite, TestHFileBlockIndex, and TestHFileWriterV2 still pass.\n\nIt's not pretty, though. Lots of places where cacheBlock now needs to be passed around.\nI erred on the side of caution, where is wasn't clear whether the block should be cached of not, I kept it being cached (as before).\n\nSomebody with more knowledge about HFileReaderV2 should have a look. If there's something nicer to do, please let me know.\nIt is not clear to be how to write a test for this.","created":"2011-09-27T18:52:06.756+0000"},{"body":"I can roll this in to my changes for CacheConfig. Don't have the JIRA off hand (but the point of it is to make it so we dont' have to change a bunch of constructors all the time).\n\nI'm hoping to have a diff out tonight or tomorrow morning.","created":"2011-09-27T23:51:52.123+0000"},{"body":"That sounds like a plan Jon. Let me know if you'd like me to have a look at the combined patch.","created":"2011-09-28T03:29:17.020+0000"},{"body":"So this exact issue actually triggered why I was having a hard time getting TestCacheOnWrite to pass. The test was previously relying on some broken/inconsistent behavior in which it passes a single instance of a reader with a null block cache but that was removed with the latest CacheConfig stuff.\n\nMy latest patch for HBASE-4422 actually just changes the always true to always false :) I'm going to talk to Mikhail tomorrow (Friday) about the issue here and see if he has any thoughts.","created":"2011-09-30T01:08:36.024+0000"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/\n-----------------------------------------------------------\n\nReview request for hbase, Jonathan Gray and Lars Hofhansl.\n\n\nSummary\n-------\n\nThis fixes a couple of long-existing code issues in HFile v2:\n- Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\n- Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\n- Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\n- Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\n\n\nThis addresses bug HBASE-4496.\n https://issues.apache.org/jira/browse/HBASE-4496\n\n\nDiffs\n-----\n\n src/test/java/org/apache/hadoop/hbase/io/hfile/TestHFileBlockIndex.java 4dc1367 \n src/test/java/org/apache/hadoop/hbase/io/hfile/TestCacheOnWrite.java 5e98375 \n src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java b429819 \n src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlockIndex.java 953896e \n src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV1.java 13d5e70 \n src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java 1cf7767 \n src/main/java/org/apache/hadoop/hbase/io/hfile/HFile.java eec566e \n\nDiff: https://reviews.apache.org/r/2136/diff\n\n\nTesting\n-------\n\nThis is in production in Facebook's hbase-89 branch. \n\nStill testing this open-source patch -- please don't commit yet.\n\n\nThanks,\n\nMikhail\n\n","created":"2011-09-30T20:42:45.669+0000"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/#review2232\n-----------------------------------------------------------\n\nShip it!\n\n\nI love it. This is what I should have done with HBASE-4496 if I had had more knowledge about the reader code.\nI'll do some more manual testing with your patch applied.\nThis will create extra merging work for HBASE-4422 and HBASE-4344\n\n\nsrc/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlockIndex.java\n\n\n This is good.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java\n\n\n I like this. HFileReaderV2 implementing HFileBlock.BasicReader was strange.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java\n\n\n No more casting, awesome.\n \n *Very* minor nit, but why not do reader.getDataBlockIndexReader().seekToDataBlock(...) as you do below?\n\n\n- Lars\n\n\nOn 2011-09-30 20:41:01, Mikhail Bautin wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/2136/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2011-09-30 20:41:01)\nbq. \nbq. \nbq. Review request for hbase, Jonathan Gray and Lars Hofhansl.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. This fixes a couple of long-existing code issues in HFile v2:\nbq. - Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\nbq. - Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\nbq. - Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\nbq. - Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\nbq. \nbq. \nbq. This addresses bug HBASE-4496.\nbq. https://issues.apache.org/jira/browse/HBASE-4496\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestHFileBlockIndex.java 4dc1367 \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestCacheOnWrite.java 5e98375 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java b429819 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlockIndex.java 953896e \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV1.java 13d5e70 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java 1cf7767 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFile.java eec566e \nbq. \nbq. Diff: https://reviews.apache.org/r/2136/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. This is in production in Facebook's hbase-89 branch. \nbq. \nbq. Still testing this open-source patch -- please don't commit yet.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Mikhail\nbq. \nbq.\n\n","created":"2011-09-30T21:05:45.723+0000"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/#review2234\n-----------------------------------------------------------\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java\n\n\n This is mainly to keep lines <80 characters and for readability.\n\n\n- Mikhail\n\n\nOn 2011-09-30 20:41:01, Mikhail Bautin wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/2136/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2011-09-30 20:41:01)\nbq. \nbq. \nbq. Review request for hbase, Jonathan Gray and Lars Hofhansl.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. This fixes a couple of long-existing code issues in HFile v2:\nbq. - Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\nbq. - Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\nbq. - Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\nbq. - Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\nbq. \nbq. \nbq. This addresses bug HBASE-4496.\nbq. https://issues.apache.org/jira/browse/HBASE-4496\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestHFileBlockIndex.java 4dc1367 \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestCacheOnWrite.java 5e98375 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java b429819 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlockIndex.java 953896e \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV1.java 13d5e70 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java 1cf7767 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFile.java eec566e \nbq. \nbq. Diff: https://reviews.apache.org/r/2136/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. This is in production in Facebook's hbase-89 branch. \nbq. \nbq. Still testing this open-source patch -- please don't commit yet.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Mikhail\nbq. \nbq.\n\n","created":"2011-09-30T21:09:45.711+0000"},{"body":"Tried Mikhail's patch with my test scenario. The scan's cacheBlock is now correctly honored.\n","created":"2011-10-02T20:29:12.812+0000"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/#review2256\n-----------------------------------------------------------\n\nShip it!\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java\n\n\n an is not needed\n\n\n- Ted\n\n\nOn 2011-09-30 20:41:01, Mikhail Bautin wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/2136/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2011-09-30 20:41:01)\nbq. \nbq. \nbq. Review request for hbase, Jonathan Gray and Lars Hofhansl.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. This fixes a couple of long-existing code issues in HFile v2:\nbq. - Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\nbq. - Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\nbq. - Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\nbq. - Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\nbq. \nbq. \nbq. This addresses bug HBASE-4496.\nbq. https://issues.apache.org/jira/browse/HBASE-4496\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestHFileBlockIndex.java 4dc1367 \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestCacheOnWrite.java 5e98375 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java b429819 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlockIndex.java 953896e \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV1.java 13d5e70 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java 1cf7767 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFile.java eec566e \nbq. \nbq. Diff: https://reviews.apache.org/r/2136/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. This is in production in Facebook's hbase-89 branch. \nbq. \nbq. Still testing this open-source patch -- please don't commit yet.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Mikhail\nbq. \nbq.\n\n","created":"2011-10-02T20:43:34.168+0000"},{"body":"Integrated to 0.92 and TRUNK.\n\nThanks for the patch Mikhail.\n\nThanks for the review Lars.","created":"2011-10-02T21:30:53.591+0000"},{"body":"Mikhail did the work, he deserves the credit :)","created":"2011-10-02T21:34:41.930+0000"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/\n-----------------------------------------------------------\n\n(Updated 2011-10-02 21:49:56.235130)\n\n\nReview request for hbase, Jonathan Gray and Lars Hofhansl.\n\n\nChanges\n-------\n\nAddressing Ted's comment and fixing TestBlocksRead, since we are now reading fewer blocks in a couple of cases. Re-running test suite just in case.\n\n\nSummary\n-------\n\nThis fixes a couple of long-existing code issues in HFile v2:\n- Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\n- Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\n- Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\n- Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\n\n\nThis addresses bug HBASE-4496.\n https://issues.apache.org/jira/browse/HBASE-4496\n\n\nDiffs (updated)\n-----\n\n src/test/java/org/apache/hadoop/hbase/regionserver/TestBlocksRead.java 019fddd \n\nDiff: https://reviews.apache.org/r/2136/diff\n\n\nTesting\n-------\n\nThis is in production in Facebook's hbase-89 branch. \n\nStill testing this open-source patch -- please don't commit yet.\n\n\nThanks,\n\nMikhail\n\n","created":"2011-10-02T21:50:34.157+0000"},{"body":"\n\nbq. On 2011-10-02 20:43:29, Ted Yu wrote:\nbq. > src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java, line 924\nbq. > \nbq. >\nbq. > an is not needed\n\nSorry -- did not realize that this was already committed. The result was that in the second version of this patch the only changes left were in TestBlocksRead. I will re-run the whole test suite again.\n\n\n- Mikhail\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/#review2256\n-----------------------------------------------------------\n\n\nOn 2011-10-02 21:49:56, Mikhail Bautin wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/2136/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2011-10-02 21:49:56)\nbq. \nbq. \nbq. Review request for hbase, Jonathan Gray and Lars Hofhansl.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. This fixes a couple of long-existing code issues in HFile v2:\nbq. - Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\nbq. - Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\nbq. - Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\nbq. - Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\nbq. \nbq. \nbq. This addresses bug HBASE-4496.\nbq. https://issues.apache.org/jira/browse/HBASE-4496\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/test/java/org/apache/hadoop/hbase/regionserver/TestBlocksRead.java 019fddd \nbq. \nbq. Diff: https://reviews.apache.org/r/2136/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. This is in production in Facebook's hbase-89 branch. \nbq. \nbq. Still testing this open-source patch -- please don't commit yet.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Mikhail\nbq. \nbq.\n\n","created":"2011-10-02T21:54:34.158+0000"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/\n-----------------------------------------------------------\n\n(Updated 2011-10-02 21:58:43.274175)\n\n\nReview request for hbase, Jonathan Gray and Lars Hofhansl.\n\n\nChanges\n-------\n\nFixing comment to be consistent with the expected number of blocks in code.\n\n\nSummary\n-------\n\nThis fixes a couple of long-existing code issues in HFile v2:\n- Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\n- Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\n- Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\n- Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\n\n\nThis addresses bug HBASE-4496.\n https://issues.apache.org/jira/browse/HBASE-4496\n\n\nDiffs (updated)\n-----\n\n src/test/java/org/apache/hadoop/hbase/regionserver/TestBlocksRead.java 019fddd \n\nDiff: https://reviews.apache.org/r/2136/diff\n\n\nTesting\n-------\n\nThis is in production in Facebook's hbase-89 branch. \n\nStill testing this open-source patch -- please don't commit yet.\n\n\nThanks,\n\nMikhail\n\n","created":"2011-10-02T22:00:34.160+0000"},{"body":"Fixes for TestBlocksRead.java","created":"2011-10-02T22:19:56.593+0000"},{"body":"Addendum integrated to 0.92 and TRUNK.","created":"2011-10-02T22:22:48.907+0000"}],"conversations":[{"body":"While testing the LRU cache during the scanning I noticed quite some churn in the cache even when Scan.cacheBlocks is set to false. After debugging this, I found that HFile V2 always caches blocks in the LRU cache regardless of the cacheBlocks setting.\n\nHere's a trace (from Eclipse) showing the problem:\n\nHFileReaderV2.readBlock(long, int, boolean, boolean, boolean) line: 279\t\nHFileReaderV2.readBlockData(long, long, int, boolean) line: 219\t\nHFileBlockIndex$BlockIndexReader.seekToDataBlock(byte[], int, int, HFileBlock) line: 191\t\nHFileReaderV2$ScannerV2.seekTo(byte[], int, int, boolean) line: 502\t\nHFileReaderV2$ScannerV2.reseekTo(byte[], int, int) line: 539\t\nStoreFileScanner.reseekAtOrAfter(HFileScanner, KeyValue) line: 151\t\nStoreFileScanner.reseek(KeyValue) line: 110\t\nKeyValueHeap.reseek(KeyValue) line: 255\t\nStoreScanner.reseek(KeyValue) line: 409\t\nStoreScanner.next(List, int) line: 304\t\nKeyValueHeap.next(List, int) line: 114\t\nKeyValueHeap.next(List) line: 143\t\nHRegion$RegionScannerImpl.nextRow(byte[]) line: 2774\t\nHRegion$RegionScannerImpl.nextInternal(int) line: 2722\t\nHRegion$RegionScannerImpl.next(List, int) line: 2682\t\nHRegion$RegionScannerImpl.next(List) line: 2699\t\nHRegionServer.next(long, int) line: 2092\t\n\nEvery scanner.next causes a reseek, which eventually causes a call to HFileBlockIndex$BlockIndexReader.seekToDataBlock(...) at which point the cacheBlocks information is lost. HFileReaderV2.readBlockData calls HFileReaderV2.readBlock with cacheBlocks set unconditionally to true.\n\nThe fix is not immediately clear, unless we want to pass cacheBlocks to HFileBlockIndex$BlockIndexReader.seekToDataBlock and then on to HFileBlock.BasicReader.readBlockData and all its implementers, which is ugly as readBlockData should not care about caching.\n\nAvoiding caching during scans is somewhat important for us.\n","from":"reporter","subject":"HFile V2 does not honor setCacheBlocks when scanning."},{"body":"I'll come up with a patch.","from":"developer"},{"body":"In a simple test this shaves about 20% of scanning time (with caching switched off), in addition to avoiding LRU and GC churn.","from":"developer"},{"body":"Good find Lars.","from":"developer"},{"body":"This fixes the problem for me.\n\nTestHFileBlock, TestCacheOnWrite, TestHFileBlockIndex, and TestHFileWriterV2 still pass.\n\nIt's not pretty, though. Lots of places where cacheBlock now needs to be passed around.\nI erred on the side of caution, where is wasn't clear whether the block should be cached of not, I kept it being cached (as before).\n\nSomebody with more knowledge about HFileReaderV2 should have a look. If there's something nicer to do, please let me know.\nIt is not clear to be how to write a test for this.","from":"developer"},{"body":"I can roll this in to my changes for CacheConfig. Don't have the JIRA off hand (but the point of it is to make it so we dont' have to change a bunch of constructors all the time).\n\nI'm hoping to have a diff out tonight or tomorrow morning.","from":"developer"},{"body":"That sounds like a plan Jon. Let me know if you'd like me to have a look at the combined patch.","from":"developer"},{"body":"So this exact issue actually triggered why I was having a hard time getting TestCacheOnWrite to pass. The test was previously relying on some broken/inconsistent behavior in which it passes a single instance of a reader with a null block cache but that was removed with the latest CacheConfig stuff.\n\nMy latest patch for HBASE-4422 actually just changes the always true to always false :) I'm going to talk to Mikhail tomorrow (Friday) about the issue here and see if he has any thoughts.","from":"developer"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/\n-----------------------------------------------------------\n\nReview request for hbase, Jonathan Gray and Lars Hofhansl.\n\n\nSummary\n-------\n\nThis fixes a couple of long-existing code issues in HFile v2:\n- Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\n- Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\n- Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\n- Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\n\n\nThis addresses bug HBASE-4496.\n https://issues.apache.org/jira/browse/HBASE-4496\n\n\nDiffs\n-----\n\n src/test/java/org/apache/hadoop/hbase/io/hfile/TestHFileBlockIndex.java 4dc1367 \n src/test/java/org/apache/hadoop/hbase/io/hfile/TestCacheOnWrite.java 5e98375 \n src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java b429819 \n src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlockIndex.java 953896e \n src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV1.java 13d5e70 \n src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java 1cf7767 \n src/main/java/org/apache/hadoop/hbase/io/hfile/HFile.java eec566e \n\nDiff: https://reviews.apache.org/r/2136/diff\n\n\nTesting\n-------\n\nThis is in production in Facebook's hbase-89 branch. \n\nStill testing this open-source patch -- please don't commit yet.\n\n\nThanks,\n\nMikhail\n\n","from":"developer"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/#review2232\n-----------------------------------------------------------\n\nShip it!\n\n\nI love it. This is what I should have done with HBASE-4496 if I had had more knowledge about the reader code.\nI'll do some more manual testing with your patch applied.\nThis will create extra merging work for HBASE-4422 and HBASE-4344\n\n\nsrc/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlockIndex.java\n\n\n This is good.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java\n\n\n I like this. HFileReaderV2 implementing HFileBlock.BasicReader was strange.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java\n\n\n No more casting, awesome.\n \n *Very* minor nit, but why not do reader.getDataBlockIndexReader().seekToDataBlock(...) as you do below?\n\n\n- Lars\n\n\nOn 2011-09-30 20:41:01, Mikhail Bautin wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/2136/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2011-09-30 20:41:01)\nbq. \nbq. \nbq. Review request for hbase, Jonathan Gray and Lars Hofhansl.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. This fixes a couple of long-existing code issues in HFile v2:\nbq. - Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\nbq. - Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\nbq. - Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\nbq. - Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\nbq. \nbq. \nbq. This addresses bug HBASE-4496.\nbq. https://issues.apache.org/jira/browse/HBASE-4496\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestHFileBlockIndex.java 4dc1367 \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestCacheOnWrite.java 5e98375 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java b429819 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlockIndex.java 953896e \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV1.java 13d5e70 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java 1cf7767 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFile.java eec566e \nbq. \nbq. Diff: https://reviews.apache.org/r/2136/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. This is in production in Facebook's hbase-89 branch. \nbq. \nbq. Still testing this open-source patch -- please don't commit yet.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Mikhail\nbq. \nbq.\n\n","from":"developer"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/#review2234\n-----------------------------------------------------------\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java\n\n\n This is mainly to keep lines <80 characters and for readability.\n\n\n- Mikhail\n\n\nOn 2011-09-30 20:41:01, Mikhail Bautin wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/2136/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2011-09-30 20:41:01)\nbq. \nbq. \nbq. Review request for hbase, Jonathan Gray and Lars Hofhansl.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. This fixes a couple of long-existing code issues in HFile v2:\nbq. - Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\nbq. - Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\nbq. - Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\nbq. - Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\nbq. \nbq. \nbq. This addresses bug HBASE-4496.\nbq. https://issues.apache.org/jira/browse/HBASE-4496\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestHFileBlockIndex.java 4dc1367 \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestCacheOnWrite.java 5e98375 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java b429819 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlockIndex.java 953896e \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV1.java 13d5e70 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java 1cf7767 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFile.java eec566e \nbq. \nbq. Diff: https://reviews.apache.org/r/2136/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. This is in production in Facebook's hbase-89 branch. \nbq. \nbq. Still testing this open-source patch -- please don't commit yet.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Mikhail\nbq. \nbq.\n\n","from":"developer"},{"body":"Tried Mikhail's patch with my test scenario. The scan's cacheBlock is now correctly honored.\n","from":"developer"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/#review2256\n-----------------------------------------------------------\n\nShip it!\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java\n\n\n an is not needed\n\n\n- Ted\n\n\nOn 2011-09-30 20:41:01, Mikhail Bautin wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/2136/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2011-09-30 20:41:01)\nbq. \nbq. \nbq. Review request for hbase, Jonathan Gray and Lars Hofhansl.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. This fixes a couple of long-existing code issues in HFile v2:\nbq. - Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\nbq. - Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\nbq. - Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\nbq. - Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\nbq. \nbq. \nbq. This addresses bug HBASE-4496.\nbq. https://issues.apache.org/jira/browse/HBASE-4496\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestHFileBlockIndex.java 4dc1367 \nbq. src/test/java/org/apache/hadoop/hbase/io/hfile/TestCacheOnWrite.java 5e98375 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java b429819 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlockIndex.java 953896e \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV1.java 13d5e70 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFileReaderV2.java 1cf7767 \nbq. src/main/java/org/apache/hadoop/hbase/io/hfile/HFile.java eec566e \nbq. \nbq. Diff: https://reviews.apache.org/r/2136/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. This is in production in Facebook's hbase-89 branch. \nbq. \nbq. Still testing this open-source patch -- please don't commit yet.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Mikhail\nbq. \nbq.\n\n","from":"developer"},{"body":"Integrated to 0.92 and TRUNK.\n\nThanks for the patch Mikhail.\n\nThanks for the review Lars.","from":"developer"},{"body":"Mikhail did the work, he deserves the credit :)","from":"developer"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/\n-----------------------------------------------------------\n\n(Updated 2011-10-02 21:49:56.235130)\n\n\nReview request for hbase, Jonathan Gray and Lars Hofhansl.\n\n\nChanges\n-------\n\nAddressing Ted's comment and fixing TestBlocksRead, since we are now reading fewer blocks in a couple of cases. Re-running test suite just in case.\n\n\nSummary\n-------\n\nThis fixes a couple of long-existing code issues in HFile v2:\n- Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\n- Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\n- Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\n- Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\n\n\nThis addresses bug HBASE-4496.\n https://issues.apache.org/jira/browse/HBASE-4496\n\n\nDiffs (updated)\n-----\n\n src/test/java/org/apache/hadoop/hbase/regionserver/TestBlocksRead.java 019fddd \n\nDiff: https://reviews.apache.org/r/2136/diff\n\n\nTesting\n-------\n\nThis is in production in Facebook's hbase-89 branch. \n\nStill testing this open-source patch -- please don't commit yet.\n\n\nThanks,\n\nMikhail\n\n","from":"developer"},{"body":"\n\nbq. On 2011-10-02 20:43:29, Ted Yu wrote:\nbq. > src/main/java/org/apache/hadoop/hbase/io/hfile/HFileBlock.java, line 924\nbq. > \nbq. >\nbq. > an is not needed\n\nSorry -- did not realize that this was already committed. The result was that in the second version of this patch the only changes left were in TestBlocksRead. I will re-run the whole test suite again.\n\n\n- Mikhail\n\n\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/#review2256\n-----------------------------------------------------------\n\n\nOn 2011-10-02 21:49:56, Mikhail Bautin wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/2136/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2011-10-02 21:49:56)\nbq. \nbq. \nbq. Review request for hbase, Jonathan Gray and Lars Hofhansl.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. This fixes a couple of long-existing code issues in HFile v2:\nbq. - Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\nbq. - Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\nbq. - Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\nbq. - Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\nbq. \nbq. \nbq. This addresses bug HBASE-4496.\nbq. https://issues.apache.org/jira/browse/HBASE-4496\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/test/java/org/apache/hadoop/hbase/regionserver/TestBlocksRead.java 019fddd \nbq. \nbq. Diff: https://reviews.apache.org/r/2136/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. This is in production in Facebook's hbase-89 branch. \nbq. \nbq. Still testing this open-source patch -- please don't commit yet.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Mikhail\nbq. \nbq.\n\n","from":"developer"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/2136/\n-----------------------------------------------------------\n\n(Updated 2011-10-02 21:58:43.274175)\n\n\nReview request for hbase, Jonathan Gray and Lars Hofhansl.\n\n\nChanges\n-------\n\nFixing comment to be consistent with the expected number of blocks in code.\n\n\nSummary\n-------\n\nThis fixes a couple of long-existing code issues in HFile v2:\n- Making seekBefore cache the previous block it has to read when the scanner happens to be at the first key of a block (this was a performance regression introduced in HFile v2).\n- Fixing the accounting of the number of blocks read for the one-level index case in HFileBlockIndex.seekToDataBlock if the current block is the same as the requested block.\n- Getting rid of HFileBlock.BasicReader, which was used both by FSReaderV2 and HFileReaderV2, but the former did not cache blocks (a source of confusion).\n- Adding a new interface HFile.CachingBlockReader instead, which is implemented by HFile readers and passed to HFileBlockIndex.\n\n\nThis addresses bug HBASE-4496.\n https://issues.apache.org/jira/browse/HBASE-4496\n\n\nDiffs (updated)\n-----\n\n src/test/java/org/apache/hadoop/hbase/regionserver/TestBlocksRead.java 019fddd \n\nDiff: https://reviews.apache.org/r/2136/diff\n\n\nTesting\n-------\n\nThis is in production in Facebook's hbase-89 branch. \n\nStill testing this open-source patch -- please don't commit yet.\n\n\nThanks,\n\nMikhail\n\n","from":"developer"},{"body":"Fixes for TestBlocksRead.java","from":"developer"},{"body":"Addendum integrated to 0.92 and TRUNK.","from":"developer"}],"created":"2011-09-27T07:22:27.000+0000","description":"While testing the LRU cache during the scanning I noticed quite some churn in the cache even when Scan.cacheBlocks is set to false. After debugging this, I found that HFile V2 always caches blocks in the LRU cache regardless of the cacheBlocks setting.\n\nHere's a trace (from Eclipse) showing the problem:\n\nHFileReaderV2.readBlock(long, int, boolean, boolean, boolean) line: 279\t\nHFileReaderV2.readBlockData(long, long, int, boolean) line: 219\t\nHFileBlockIndex$BlockIndexReader.seekToDataBlock(byte[], int, int, HFileBlock) line: 191\t\nHFileReaderV2$ScannerV2.seekTo(byte[], int, int, boolean) line: 502\t\nHFileReaderV2$ScannerV2.reseekTo(byte[], int, int) line: 539\t\nStoreFileScanner.reseekAtOrAfter(HFileScanner, KeyValue) line: 151\t\nStoreFileScanner.reseek(KeyValue) line: 110\t\nKeyValueHeap.reseek(KeyValue) line: 255\t\nStoreScanner.reseek(KeyValue) line: 409\t\nStoreScanner.next(List, int) line: 304\t\nKeyValueHeap.next(List, int) line: 114\t\nKeyValueHeap.next(List) line: 143\t\nHRegion$RegionScannerImpl.nextRow(byte[]) line: 2774\t\nHRegion$RegionScannerImpl.nextInternal(int) line: 2722\t\nHRegion$RegionScannerImpl.next(List, int) line: 2682\t\nHRegion$RegionScannerImpl.next(List) line: 2699\t\nHRegionServer.next(long, int) line: 2092\t\n\nEvery scanner.next causes a reseek, which eventually causes a call to HFileBlockIndex$BlockIndexReader.seekToDataBlock(...) at which point the cacheBlocks information is lost. HFileReaderV2.readBlockData calls HFileReaderV2.readBlock with cacheBlocks set unconditionally to true.\n\nThe fix is not immediately clear, unless we want to pass cacheBlocks to HFileBlockIndex$BlockIndexReader.seekToDataBlock and then on to HFileBlock.BasicReader.readBlockData and all its implementers, which is ugly as readBlockData should not care about caching.\n\nAvoiding caching during scans is somewhat important for us.\n","issue_id":"12524808","key":"HBASE-4496","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2011-10-02T21:30:59.000+0000","role":"fixed_distractor","summary":"HFile V2 does not honor setCacheBlocks when scanning."} {"case_id":"12526464","cluster":"DISTRACTOR-HBASE-4565","comments":[{"body":"@Suraj Do you have a patch for us?","created":"2011-10-10T16:50:56.382+0000"},{"body":"Working on one ... assigning this to myself.","created":"2011-10-13T16:55:34.335+0000"},{"body":"Update: I haven't been able to get it to work in a cross-platform way as yet. The best I've been able to do is to exec cygpath -u ${project.build.directory} to get a POSIX directory ... however, this would involve replacing the ${project.build.directory} with a new variable for cygwin.\n\nThe workaround for now is to manually update the path to /cygdrive/c/path/to/project/build/directory wherever it execs out to shell.\nI'll continue to look into other options. ","created":"2011-10-14T08:39:04.978+0000"},{"body":"+1. Not having to dual boot on my desktop would be awesome.","created":"2011-10-26T01:03:13.104+0000"},{"body":"Patch to support maven build on cygwin. Same approach as being done in hadoop-0.23 pom.xml to support cygwin.","created":"2011-11-11T18:24:58.638+0000"},{"body":"Same patch against hbase-0.92 ","created":"2011-11-11T19:05:48.190+0000"},{"body":"Trunk patch - just ran dos2unix to remove Windows line endings ... no other change.","created":"2011-11-11T19:11:15.446+0000"},{"body":"The tarball in the previous patches had an extra directory. Attaching v3 versions for both trunk and 0.92","created":"2011-11-13T06:41:09.723+0000"},{"body":"Really helpful. Thanks a lot. \n\nAlso patch seems to be failing on trunk. I am not sure it its already out of sync with trunk.\n\n$ patch -p0 < pom.patch\n(Stripping trailing CRs from patch.)\npatching file pom.xml\nHunk #1 FAILED at 696.\npatch unexpectedly ends in middle of line\nHunk #2 succeeded at 1328 with fuzz 1 (offset 90 lines).\n1 out of 2 hunks FAILED -- saving rejects to file pom.xml.rej\n\nI manually patched it and works beautifully. ","created":"2011-12-27T15:25:14.293+0000"},{"body":"Ah - looks like it is out of sync with 0.92 and trunk now with the hbase-security pom updates. I'll rebase and put out a new patch version.","created":"2012-01-04T15:25:44.663+0000"},{"body":"@Suraj: Are you still working on this. I think we want this in 0.94.","created":"2012-03-21T17:23:06.901+0000"},{"body":"Moving out of 0.94, pull back if you feel differently.","created":"2012-03-22T15:38:06.077+0000"},{"body":"Is it ok if I rebase this patch to trunk?\nI need it to build in my windows env.","created":"2012-03-28T10:09:00.191+0000"},{"body":"@Laxman - Yes, please go ahead.\n\n@Lars - Ok, no problem.","created":"2012-03-28T14:35:08.082+0000"},{"body":"It does not look like the rebase to trunk was posted? [~svarma] Want to rebase?","created":"2012-09-22T22:04:11.164+0000"},{"body":"This is no longer an issue on trunk, it appears. The build script modularization changes have completely done away with the copynativelibs.sh which caused the original issue. I am able to build from trunk successfully via cygwin now.","created":"2012-09-26T05:02:11.974+0000"},{"body":"[~svarma] So we should apply the patch to 0.92 and 0.94? The v3 patch still works on windows? Thanks for checking trunk.","created":"2012-09-26T05:38:50.870+0000"},{"body":"Yes, this is an issue in 0.92 and 0.94. I don't think the v3 patch still works as the security profile changes were made after that. I'll rebase and upload new patches for 0.92 and 0.94.","created":"2012-09-28T18:55:13.199+0000"},{"body":"Rebased patch to both 0.92 and 0.94. Please find attached v4 versions of the patch.\nTested on a Windows 7 / cygwin environment and build ran successfully for both 0.92 head and 0.94 head.","created":"2012-09-29T06:55:29.492+0000"},{"body":"To clarify ... the v4 patches are for 0.92 and 0.94 branches. (trunk doesn't need any patch).","created":"2012-09-29T07:11:36.988+0000"},{"body":"Committed to 0.94 and to 0.92. I checked patch doesn't break linux packag'ing. Thanks for the patch Suraj.","created":"2012-09-29T21:06:36.216+0000"}],"conversations":[{"body":"This is broken in both 0.92 as well as trunk pom.xml\n\nHere's a sample maven log snippet from trunk (from Mayuresh on user mailing list)\n\n[INFO] [antrun:run {execution: package}]\n[INFO] Executing tasks\n\nmain:\n [mkdir] Created dir: D:\\workspace\\mkshirsa\\hbase-trunk\\target\\hbase-0.93-SNAPSHOT\\hbase-0.93-SNAPSHOT\\lib\\native\\${build.platform}\n [exec] ls: cannot access D:workspacemkshirsahbase-trunktarget/nativelib: No such file or directory\n [exec] tar (child): Cannot connect to D: resolve failed\n[INFO] ------------------------------------------------------------------------\n[ERROR] BUILD ERROR\n[INFO] ------------------------------------------------------------------------\n[INFO] An Ant BuildException has occured: exec returned: 3328\n\nThere are two issues: \n1) The ant run task below doesn't resolve the windows file separator returned by the project.build.directory - this causes the above resolve failed.\n\n\n if [ `ls ${project.build.directory}/nativelib | wc -l` -ne 0]; then\n\n\n2) The tar argument value below also has a similar issue in that the path arg doesn't resolve right.\n\n\n \n \n \n\n\nIn both cases, the fix would probably be to use a cross-platform way to handle the directory locations. \n","from":"reporter","subject":"Maven HBase build broken on cygwin with copynativelib.sh call."},{"body":"@Suraj Do you have a patch for us?","from":"developer"},{"body":"Working on one ... assigning this to myself.","from":"developer"},{"body":"Update: I haven't been able to get it to work in a cross-platform way as yet. The best I've been able to do is to exec cygpath -u ${project.build.directory} to get a POSIX directory ... however, this would involve replacing the ${project.build.directory} with a new variable for cygwin.\n\nThe workaround for now is to manually update the path to /cygdrive/c/path/to/project/build/directory wherever it execs out to shell.\nI'll continue to look into other options. ","from":"developer"},{"body":"+1. Not having to dual boot on my desktop would be awesome.","from":"developer"},{"body":"Patch to support maven build on cygwin. Same approach as being done in hadoop-0.23 pom.xml to support cygwin.","from":"developer"},{"body":"Same patch against hbase-0.92 ","from":"developer"},{"body":"Trunk patch - just ran dos2unix to remove Windows line endings ... no other change.","from":"developer"},{"body":"The tarball in the previous patches had an extra directory. Attaching v3 versions for both trunk and 0.92","from":"developer"},{"body":"Really helpful. Thanks a lot. \n\nAlso patch seems to be failing on trunk. I am not sure it its already out of sync with trunk.\n\n$ patch -p0 < pom.patch\n(Stripping trailing CRs from patch.)\npatching file pom.xml\nHunk #1 FAILED at 696.\npatch unexpectedly ends in middle of line\nHunk #2 succeeded at 1328 with fuzz 1 (offset 90 lines).\n1 out of 2 hunks FAILED -- saving rejects to file pom.xml.rej\n\nI manually patched it and works beautifully. ","from":"developer"},{"body":"Ah - looks like it is out of sync with 0.92 and trunk now with the hbase-security pom updates. I'll rebase and put out a new patch version.","from":"developer"},{"body":"@Suraj: Are you still working on this. I think we want this in 0.94.","from":"developer"},{"body":"Moving out of 0.94, pull back if you feel differently.","from":"developer"},{"body":"Is it ok if I rebase this patch to trunk?\nI need it to build in my windows env.","from":"developer"},{"body":"@Laxman - Yes, please go ahead.\n\n@Lars - Ok, no problem.","from":"developer"},{"body":"It does not look like the rebase to trunk was posted? [~svarma] Want to rebase?","from":"developer"},{"body":"This is no longer an issue on trunk, it appears. The build script modularization changes have completely done away with the copynativelibs.sh which caused the original issue. I am able to build from trunk successfully via cygwin now.","from":"developer"},{"body":"[~svarma] So we should apply the patch to 0.92 and 0.94? The v3 patch still works on windows? Thanks for checking trunk.","from":"developer"},{"body":"Yes, this is an issue in 0.92 and 0.94. I don't think the v3 patch still works as the security profile changes were made after that. I'll rebase and upload new patches for 0.92 and 0.94.","from":"developer"},{"body":"Rebased patch to both 0.92 and 0.94. Please find attached v4 versions of the patch.\nTested on a Windows 7 / cygwin environment and build ran successfully for both 0.92 head and 0.94 head.","from":"developer"},{"body":"To clarify ... the v4 patches are for 0.92 and 0.94 branches. (trunk doesn't need any patch).","from":"developer"},{"body":"Committed to 0.94 and to 0.92. I checked patch doesn't break linux packag'ing. Thanks for the patch Suraj.","from":"developer"}],"created":"2011-10-10T16:23:35.000+0000","description":"This is broken in both 0.92 as well as trunk pom.xml\n\nHere's a sample maven log snippet from trunk (from Mayuresh on user mailing list)\n\n[INFO] [antrun:run {execution: package}]\n[INFO] Executing tasks\n\nmain:\n [mkdir] Created dir: D:\\workspace\\mkshirsa\\hbase-trunk\\target\\hbase-0.93-SNAPSHOT\\hbase-0.93-SNAPSHOT\\lib\\native\\${build.platform}\n [exec] ls: cannot access D:workspacemkshirsahbase-trunktarget/nativelib: No such file or directory\n [exec] tar (child): Cannot connect to D: resolve failed\n[INFO] ------------------------------------------------------------------------\n[ERROR] BUILD ERROR\n[INFO] ------------------------------------------------------------------------\n[INFO] An Ant BuildException has occured: exec returned: 3328\n\nThere are two issues: \n1) The ant run task below doesn't resolve the windows file separator returned by the project.build.directory - this causes the above resolve failed.\n\n\n if [ `ls ${project.build.directory}/nativelib | wc -l` -ne 0]; then\n\n\n2) The tar argument value below also has a similar issue in that the path arg doesn't resolve right.\n\n\n \n \n \n\n\nIn both cases, the fix would probably be to use a cross-platform way to handle the directory locations. \n","issue_id":"12526464","key":"HBASE-4565","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-09-29T21:06:36.000+0000","role":"fixed_distractor","summary":"Maven HBase build broken on cygwin with copynativelib.sh call."} {"case_id":"12528693","cluster":"DISTRACTOR-HBASE-4671","comments":[{"body":"(Changing 127.0.1.1 to 127.0.0.1 in the hosts file that is.)","created":"2011-10-25T16:01:34.039+0000"},{"body":"It's actual for Ubuntu 12.04 too.","created":"2012-06-23T14:59:00.170+0000"},{"body":"From the hbase reference guide: http://hbase.apache.org/book.html#os\n\n{noformat}\n2.2.3. Loopback IP\n\nHBase expects the loopback IP address to be 127.0.0.1. Ubuntu and some other distributions, for example, will default to 127.0.1.1 and this will cause problems for you.\n\n/etc/hosts should look something like this:\n\n 127.0.0.1 localhost\n 127.0.0.1 ubuntu.ubuntu-domain ubuntu\n{noformat}\n","created":"2012-06-23T15:18:08.839+0000"},{"body":"nkeywal, thank you for reference and quick response! \nYou made me curious :) So, I've looked into sources and found that some services start on 'localhost' and some on 'ubuntu' (accordingly to your example) domain name. \nI'll try to look into sources more deeply and create a patch if it's possible. ","created":"2012-06-23T16:50:42.639+0000"},{"body":"Marking as fixed. Looks like a case of RTFM - unless we actually need to change some hardcoded values?","created":"2012-08-28T17:46:06.913+0000"}],"conversations":[{"body":"When /etc/hosts contains following lines (and this is not uncommon) it will cause HBaseTestingUtility to malfunction.\n127.0.0.1\tlocalhost\n127.0.1.1\tmyMachineName\n\nSymptoms:\n2011-10-25 17:38:30,875 WARN master.AssignmentManager - Failed assignment of -ROOT-,,0.70236052 to serverName=localhost,34462,1319557102914, load=(requests=0, regions=0, usedHeap=46, maxHeap=865), trying to assign elsewhere instead; retry=0\norg.apache.hadoop.hbase.client.RetriesExhaustedException: Failed setting up proxy interface org.apache.hadoop.hbase.ipc.HRegionInterface to /127.0.0.1:34462 after attempts=1\n\nbecause\n\n2011-10-25 17:38:28,371 INFO regionserver.HRegionServer - Serving as localhost,34462,1319557102914, RPC listening on /127.0.1.1:34462, sessionid=0x1333bbb7a180002\n\ncaused by /127.0.0.1:34462 vs /127.0.1.1:34462\n\nWorkaround:\nChanging 127.0.1.1 to 127.0.0.1 works.\n\nPermanent solution:\nDunno, my understanding of inner workings is not sufficient enough. Although it seems like it has something to do with changing the machine name from myMachineName to localhost during the test:\n2011-10-25 17:38:28,056 INFO regionserver.HRegionServer - Master passed us address to use. Was=myMachineName:34462, Now=localhost:34462","from":"reporter","subject":"HBaseTestingUtility unable to connect to regionserver because of 127.0.0.1 / 127.0.1.1 discrepancy"},{"body":"(Changing 127.0.1.1 to 127.0.0.1 in the hosts file that is.)","from":"developer"},{"body":"It's actual for Ubuntu 12.04 too.","from":"developer"},{"body":"From the hbase reference guide: http://hbase.apache.org/book.html#os\n\n{noformat}\n2.2.3. Loopback IP\n\nHBase expects the loopback IP address to be 127.0.0.1. Ubuntu and some other distributions, for example, will default to 127.0.1.1 and this will cause problems for you.\n\n/etc/hosts should look something like this:\n\n 127.0.0.1 localhost\n 127.0.0.1 ubuntu.ubuntu-domain ubuntu\n{noformat}\n","from":"developer"},{"body":"nkeywal, thank you for reference and quick response! \nYou made me curious :) So, I've looked into sources and found that some services start on 'localhost' and some on 'ubuntu' (accordingly to your example) domain name. \nI'll try to look into sources more deeply and create a patch if it's possible. ","from":"developer"},{"body":"Marking as fixed. Looks like a case of RTFM - unless we actually need to change some hardcoded values?","from":"developer"}],"created":"2011-10-25T15:42:49.000+0000","description":"When /etc/hosts contains following lines (and this is not uncommon) it will cause HBaseTestingUtility to malfunction.\n127.0.0.1\tlocalhost\n127.0.1.1\tmyMachineName\n\nSymptoms:\n2011-10-25 17:38:30,875 WARN master.AssignmentManager - Failed assignment of -ROOT-,,0.70236052 to serverName=localhost,34462,1319557102914, load=(requests=0, regions=0, usedHeap=46, maxHeap=865), trying to assign elsewhere instead; retry=0\norg.apache.hadoop.hbase.client.RetriesExhaustedException: Failed setting up proxy interface org.apache.hadoop.hbase.ipc.HRegionInterface to /127.0.0.1:34462 after attempts=1\n\nbecause\n\n2011-10-25 17:38:28,371 INFO regionserver.HRegionServer - Serving as localhost,34462,1319557102914, RPC listening on /127.0.1.1:34462, sessionid=0x1333bbb7a180002\n\ncaused by /127.0.0.1:34462 vs /127.0.1.1:34462\n\nWorkaround:\nChanging 127.0.1.1 to 127.0.0.1 works.\n\nPermanent solution:\nDunno, my understanding of inner workings is not sufficient enough. Although it seems like it has something to do with changing the machine name from myMachineName to localhost during the test:\n2011-10-25 17:38:28,056 INFO regionserver.HRegionServer - Master passed us address to use. Was=myMachineName:34462, Now=localhost:34462","issue_id":"12528693","key":"HBASE-4671","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-08-28T17:46:06.000+0000","role":"fixed_distractor","summary":"HBaseTestingUtility unable to connect to regionserver because of 127.0.0.1 / 127.0.1.1 discrepancy"} {"case_id":"12534821","cluster":"DISTRACTOR-HBASE-5010","comments":[{"body":"mbautin requested code review of \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Modifying scanner selection in StoreScanner to take TTL into account, so that we don't scan StoreFiles that only contain expired keys.\n\n This diff is for 89-fb, but there will be a very similar diff for trunk.\n\nTEST PLAN\n Unit tests (existing ones and a new one).\n Deploy to cluster and scan an existing table that only contains expired keys. Verify that next() calls don't time out. Previously, the scanner would try to go through the whole table in an attempt to find a non-expired key.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestSelectScannersUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n\nMANAGE HERALD DIFFERENTIAL RULES\n https://reviews.facebook.net/herald/view/differential/\n\nWHY DID I GET THIS EMAIL?\n https://reviews.facebook.net/herald/transcript/1911/\n\nTip: use the X-Herald-Rules header to filter Herald messages in your client.\n","created":"2011-12-18T01:19:30.618+0000"},{"body":"Kannan has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n Mikhail: Nice work, and unit test. Compaction related comment inline.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:133 pre-existing issue: this constructor is no longer just for major compactions. Minor compactions also use this.\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:154 If we do a similar check in this constructor, we would get the same optimization for compactions also.\n\n As we talked about offline, if the entire file has expired data, then we can avoid adding the scanner to the KeyValueHeap below. So for CFs which have routinely expiring data due to TTL, compactions would have to read a lot less data too or could essentially turn into feather-weight ops which just delete unnecessary/old files.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-22T02:27:30.314+0000"},{"body":"tedyu has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n Nice work.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java:80 oldestUnexpiredTS is missing after @param\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:149 This change makes the line over 80 chars.\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:169 Line too long.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-22T19:03:30.674+0000"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Addressing Kannan's and Ted's comments.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestSelectScannersUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n","created":"2011-12-23T02:59:30.540+0000"},{"body":"lhofhansl has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n lgtm\n See minor comment inline.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java:832 Rather than making this public, should locate the TestClass in same package and this \"package private\". That would reduce the change of anybody accidentally using it.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-23T20:42:30.205+0000"},{"body":"mbautin requested code review of \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, tedyu, JIRA\n\n This is the trunk version of D909. The main difference is that there is a minVersions CF setting in trunk, and when minVersions is not zero, we can't exclude StoreFiles based on TTL, because we might have to retrieve KVs with expired timestamps to comply with the minVersions requirement.\n\nTEST PLAN\n Unit tests (including a new one).\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ExplicitColumnTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanWildcardColumnTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/Store.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompaction.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestExplicitColumnTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMinVersions.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestScanWildcardColumnTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestSelectScannersUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreScanner.java\n\nMANAGE HERALD DIFFERENTIAL RULES\n https://reviews.facebook.net/herald/view/differential/\n\nWHY DID I GET THIS EMAIL?\n https://reviews.facebook.net/herald/transcript/2127/\n\nTip: use the X-Herald-Rules header to filter Herald messages in your client.\n","created":"2011-12-25T23:36:30.531+0000"},{"body":"tedyu has commented on the revision \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\n\n Looks good.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java:1207 Should we mention oldestUnexpiredTS here ?\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n","created":"2011-12-25T23:48:30.773+0000"},{"body":"MinVersions is my doing. I'll take a look as soon as I get a chance. \n\n\n","created":"2011-12-25T23:54:30.641+0000"},{"body":"lhofhansl has commented on the revision \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\n\n Looks pretty good. See minor comments inline.\n minVersions handling looks correct to me.\n\n I think we should call out somewhere that this optimization cannot be used with MIN_VERSIONS enabled, along with some guess for the typical performance penalty.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java:737 As said in my 0.90 review, this is better package private with the test class in the same package.\n src/main/java/org/apache/hadoop/hbase/regionserver/ExplicitColumnTracker.java:74 Javadoc for oldestUnexpiredTS is missing.\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java:134 Javadoc for oldestUnexpiredTS\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n","created":"2011-12-26T04:14:30.463+0000"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, tedyu, JIRA\n\n Addressing Ted's comment and getting rid of unused ttl parameters in a couple of places.\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ExplicitColumnTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanWildcardColumnTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/Store.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompaction.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestExplicitColumnTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMinVersions.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestScanWildcardColumnTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestSelectScannersUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreScanner.java\n","created":"2011-12-26T11:58:30.898+0000"},{"body":"Moved the new test to src/test/java/org/apache/hadoop/hbase/io/hfile/TestScannerSelectionUsingTTL.java\n\nAdded javadoc for oldestUnexpiredTS\n\nAdded test category for TestScannerSelectionUsingTTL","created":"2011-12-26T16:03:17.270+0000"},{"body":"Fix compilation error in TestScannerSelectionUsingTTL","created":"2011-12-26T16:20:11.014+0000"},{"body":"Test failures reported by Hadoop QA were due to NumberFormatException\n\nPatch integrated to TRUNK.\n\nThanks for the patch Mikhail.\n\nThanks for the review Lars.\n\nI tried applying 5010.patch to 0.92 but got some rejected changes.","created":"2011-12-26T19:37:45.495+0000"},{"body":"@Ted: thanks for taking care of the patch! A couple of questions:\n(1) Do you know why these spurious NumberFormatExceptions occur? Has this happened before?\n(2) Do you think we need this optimization in 0.92?\n\n@Lars: thanks for the review!\n","created":"2011-12-26T21:47:01.292+0000"},{"body":"We have a bunch of other optimizations in 0.94 only. I think this should be 0.94 only as well.","created":"2011-12-26T22:06:32.698+0000"},{"body":"The exception was Due to mapreduce-3583","created":"2011-12-26T22:09:42.947+0000"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\n\n Lars: replying to your comments inline. I started addressing them but Ted addressed some of them and committed the patch first :) Not sure if anything was done about emphasizing the incompatibility of this fix with MIN_VERSIONS, though. At least, this fix does not make things worse if MIN_VERSIONS is nonzero. A performance penalty could only be incurred when there are entire HFiles consisting of expired entries and MIN_VERSIONS is nonzero.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java:737 Done.\n src/main/java/org/apache/hadoop/hbase/regionserver/ExplicitColumnTracker.java:74 Done.\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java:134 Done.\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n","created":"2011-12-26T22:12:30.401+0000"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Addressing Lars's comment.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/io/hfile/TestScannerSelectionUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n","created":"2011-12-26T22:20:30.364+0000"},{"body":"mbautin has abandoned the revision \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\n\n This was committed into trunk by @tedyu, closing the revision.\n\n @tedyu: could you please include a line of the form \"Differential Revision: D1017\" in commit messages in the future? That way the revision could be automatically marked as \"committed\" in Phabricator.\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n","created":"2011-12-26T23:00:30.513+0000"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n See some replies to comments inline.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:133 Removed the line about major compactions.\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:149 Fixed.\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java:80 Fixed.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-27T03:20:31.231+0000"},{"body":"Kannan has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n The compaction code path doesn't yet get the benefit of this optimization. Spoke with Mikhail offline, and he'll update the diff to handle this case.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-29T00:25:31.158+0000"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Addressing Kannan's comment about doing the same optimization during compactions. Adding a compaction test to the unit test, and verifying that we don't read expired files using per-CF metrics.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/Store.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/metrics/SchemaMetrics.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/io/hfile/TestScannerSelectionUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n","created":"2011-12-29T01:57:30.341+0000"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:219 Assuming that a default scan will go through all KVs. Please let me know if this is incorrect.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-29T01:59:30.165+0000"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:154 This is addressed in the most recent version of the diff.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-29T01:59:31.459+0000"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Actually, this is where I addressed Kannan's comment about compactions:\n\n https://reviews.facebook.net/D909?vs=2679&id=3015&whitespace=ignore-all\n\n (see the new line 154 in StoreScanner.java:\n\n scanners = selectScannersFrom(scanners);\n\n )\n\n I have spent quite a bit of time making sure that the unit test is testing the optimization during compactions. I am now invoking compactions in two different ways, for a total of 6 parameterized instances of the test, but it is still really quick. All of our compaction codepaths go through StoreScanner constructors, so we should have the optimization in all of them.\n\n All unit tests pass.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/Store.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/metrics/SchemaMetrics.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/io/hfile/TestScannerSelectionUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompaction.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n","created":"2011-12-29T03:18:31.552+0000"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:155 This is why we filter out expired StoreFiles on compactions.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-29T03:22:31.487+0000"},{"body":"Kannan has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:209 scan & columns arguments no longer used?\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:224 and we are using this.scan & this.columns here?\n\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-29T03:52:30.654+0000"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:209 Yes, that's correct.\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:224 Yes, this is the way it was in the original fix (89-fb version).\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-29T19:19:30.079+0000"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Addressing the issue that Kannan pointed out: getScanner does not use its arguments. It turned out that getScanner was called with this.scan and this.columns in all callsites, so I have removed those arguments.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/Store.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/metrics/SchemaMetrics.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/io/hfile/TestScannerSelectionUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompaction.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n","created":"2011-12-29T19:54:30.683+0000"},{"body":"Kannan has accepted the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n awesome optimization and nice work!\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","created":"2011-12-29T19:58:30.268+0000"},{"body":"Can we close this? Its been applied to trunk. Its not in fb-0.89?","created":"2012-01-13T23:33:04.095+0000"},{"body":"Actually... Does this work together with HBASE-4721 correctly?\nHBASE-4721 allows keeping delete markers longer than the TTL set for the store.\n","created":"2012-01-14T04:51:00.305+0000"},{"body":"nspiegelberg has committed the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nCOMMIT\n https://reviews.facebook.net/rHBASEEIGHTNINEFBBRANCH1232732\n","created":"2012-01-18T03:13:40.462+0000"},{"body":"@Stack: there a small item pending here. The 89-fb patch contains more testing for the compaction case, which has to be ported to trunk.\n\n@Lars: I will work with Prakash to ensure this does not break HBASE-4721. I believe HBASE-4721 is only used for a replication approach that is still in development.\n\nI think we'll keep this JIRA open until the two above items are resolved, unless anyone objects.\n","created":"2012-01-18T23:46:59.703+0000"},{"body":"A follow-up fix was submitted as part of HBASE-5274 to bring the trunk fix for this issue to parity with the 89-fb fix. Resolving.","created":"2012-01-27T02:07:07.037+0000"},{"body":"This change is doesn't break HBASE-4721.\n\nHBASE-4721 introduced another parameter called hbase.hstore.time.to.purge.deletes to keep deletes even after major compactions. But hbase.hstore.time.to.purge.deletes doesn't override the TTL of the store.\n\nPasting the comment from code which hopefully makes it clear that this diff works with HBASE-4721\n\n // By default, when hbase.hstore.time.to.purge.deletes is 0ms, a delete\n // marker is always removed during a major compaction. If set to non-zero\n // value then major compaction will try to keep a delete marker around for\n // the given number of milliseconds. We want to keep the delete markers\n // around a bit longer because old puts might appear out-of-order. For\n // example, during log replication between two clusters.\n //\n // If the delete marker has lived longer than its column-family's TTL then\n // the delete marker will be removed even if time.to.purge.deletes has not\n // passed. This is because all the Puts that this delete marker can influence\n // would have also expired. (Removing of delete markers on col family TTL will\n // not happen if min-versions is set to non-zero)\n //","created":"2012-01-27T04:53:38.015+0000"},{"body":"Thanks for clarifying Prakash.","created":"2012-01-27T04:56:02.322+0000"},{"body":"Why HBASE-5510 updates coming in this JIRA?\n\nRegards\nRam\n\n","created":"2012-03-06T09:04:56.568+0000"},{"body":"@Ram: I don't see any mentions of HBASE-5510 in this JIRA, except for your comment. What updates are you referring to?","created":"2012-03-06T18:11:11.036+0000"},{"body":"@Mikhail\nThe commit related updates that usually comes up once a commit is done was appearing in this JIRA. But it was for HBASE-5510. Ted removed them as it was not related to this JIRA. May be that deleted part you are not able to view now. :)\nSorry if the above comment had confused you.","created":"2012-03-06T18:13:27.603+0000"},{"body":"Actually here is the reason for those confusing updates. Ted seems to have specified HBASE-5010 instead of HBASE-5510 in the commit message.\n\ncommit 5d773d9fa176cb056b993fdff8a2853f75315ec8\nAuthor: tedyu \nDate: Mon Mar 5 10:41:03 2012\n\n HBASE-5010 Pass region info in LoadBalancer.randomAssignment(List servers) (Anoop Sam \n \n git-svn-id: http://svn.apache.org/repos/asf/hbase/trunk@1297155 13f79535-47bb-0310-9956-ffa450edef68\n\n","created":"2012-03-06T18:17:28.325+0000"}],"conversations":[{"body":"In ScanWildcardColumnTracker we have\n\n{code:java}\n \n this.oldestStamp = EnvironmentEdgeManager.currentTimeMillis() - ttl;\n\n ...\n\n private boolean isExpired(long timestamp) {\n return timestamp < oldestStamp;\n }\n{code}\n\nbut this time range filtering does not participate in HFile selection. In one real case this caused next() calls to time out because all KVs in a table got expired, but next() had to iterate over the whole table to find that out. We should be able to filter out those HFiles right away. I think a reasonable approach is to add a \"default timerange filter\" to every scan for a CF with a finite TTL and utilize existing filtering in StoreFile.Reader.passesTimerangeFilter.\n","from":"reporter","subject":"Filter HFiles based on TTL"},{"body":"mbautin requested code review of \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Modifying scanner selection in StoreScanner to take TTL into account, so that we don't scan StoreFiles that only contain expired keys.\n\n This diff is for 89-fb, but there will be a very similar diff for trunk.\n\nTEST PLAN\n Unit tests (existing ones and a new one).\n Deploy to cluster and scan an existing table that only contains expired keys. Verify that next() calls don't time out. Previously, the scanner would try to go through the whole table in an attempt to find a non-expired key.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestSelectScannersUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n\nMANAGE HERALD DIFFERENTIAL RULES\n https://reviews.facebook.net/herald/view/differential/\n\nWHY DID I GET THIS EMAIL?\n https://reviews.facebook.net/herald/transcript/1911/\n\nTip: use the X-Herald-Rules header to filter Herald messages in your client.\n","from":"developer"},{"body":"Kannan has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n Mikhail: Nice work, and unit test. Compaction related comment inline.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:133 pre-existing issue: this constructor is no longer just for major compactions. Minor compactions also use this.\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:154 If we do a similar check in this constructor, we would get the same optimization for compactions also.\n\n As we talked about offline, if the entire file has expired data, then we can avoid adding the scanner to the KeyValueHeap below. So for CFs which have routinely expiring data due to TTL, compactions would have to read a lot less data too or could essentially turn into feather-weight ops which just delete unnecessary/old files.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"tedyu has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n Nice work.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java:80 oldestUnexpiredTS is missing after @param\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:149 This change makes the line over 80 chars.\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:169 Line too long.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Addressing Kannan's and Ted's comments.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestSelectScannersUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n","from":"developer"},{"body":"lhofhansl has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n lgtm\n See minor comment inline.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java:832 Rather than making this public, should locate the TestClass in same package and this \"package private\". That would reduce the change of anybody accidentally using it.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"mbautin requested code review of \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, tedyu, JIRA\n\n This is the trunk version of D909. The main difference is that there is a minVersions CF setting in trunk, and when minVersions is not zero, we can't exclude StoreFiles based on TTL, because we might have to retrieve KVs with expired timestamps to comply with the minVersions requirement.\n\nTEST PLAN\n Unit tests (including a new one).\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ExplicitColumnTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanWildcardColumnTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/Store.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompaction.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestExplicitColumnTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMinVersions.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestScanWildcardColumnTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestSelectScannersUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreScanner.java\n\nMANAGE HERALD DIFFERENTIAL RULES\n https://reviews.facebook.net/herald/view/differential/\n\nWHY DID I GET THIS EMAIL?\n https://reviews.facebook.net/herald/transcript/2127/\n\nTip: use the X-Herald-Rules header to filter Herald messages in your client.\n","from":"developer"},{"body":"tedyu has commented on the revision \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\n\n Looks good.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java:1207 Should we mention oldestUnexpiredTS here ?\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n","from":"developer"},{"body":"MinVersions is my doing. I'll take a look as soon as I get a chance. \n\n\n","from":"developer"},{"body":"lhofhansl has commented on the revision \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\n\n Looks pretty good. See minor comments inline.\n minVersions handling looks correct to me.\n\n I think we should call out somewhere that this optimization cannot be used with MIN_VERSIONS enabled, along with some guess for the typical performance penalty.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java:737 As said in my 0.90 review, this is better package private with the test class in the same package.\n src/main/java/org/apache/hadoop/hbase/regionserver/ExplicitColumnTracker.java:74 Javadoc for oldestUnexpiredTS is missing.\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java:134 Javadoc for oldestUnexpiredTS\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n","from":"developer"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, tedyu, JIRA\n\n Addressing Ted's comment and getting rid of unused ttl parameters in a couple of places.\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ExplicitColumnTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanWildcardColumnTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/Store.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompaction.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestExplicitColumnTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMinVersions.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestScanWildcardColumnTracker.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestSelectScannersUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreScanner.java\n","from":"developer"},{"body":"Moved the new test to src/test/java/org/apache/hadoop/hbase/io/hfile/TestScannerSelectionUsingTTL.java\n\nAdded javadoc for oldestUnexpiredTS\n\nAdded test category for TestScannerSelectionUsingTTL","from":"developer"},{"body":"Fix compilation error in TestScannerSelectionUsingTTL","from":"developer"},{"body":"Test failures reported by Hadoop QA were due to NumberFormatException\n\nPatch integrated to TRUNK.\n\nThanks for the patch Mikhail.\n\nThanks for the review Lars.\n\nI tried applying 5010.patch to 0.92 but got some rejected changes.","from":"developer"},{"body":"@Ted: thanks for taking care of the patch! A couple of questions:\n(1) Do you know why these spurious NumberFormatExceptions occur? Has this happened before?\n(2) Do you think we need this optimization in 0.92?\n\n@Lars: thanks for the review!\n","from":"developer"},{"body":"We have a bunch of other optimizations in 0.94 only. I think this should be 0.94 only as well.","from":"developer"},{"body":"The exception was Due to mapreduce-3583","from":"developer"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\n\n Lars: replying to your comments inline. I started addressing them but Ted addressed some of them and committed the patch first :) Not sure if anything was done about emphasizing the incompatibility of this fix with MIN_VERSIONS, though. At least, this fix does not make things worse if MIN_VERSIONS is nonzero. A performance penalty could only be incurred when there are entire HFiles consisting of expired entries and MIN_VERSIONS is nonzero.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java:737 Done.\n src/main/java/org/apache/hadoop/hbase/regionserver/ExplicitColumnTracker.java:74 Done.\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java:134 Done.\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n","from":"developer"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Addressing Lars's comment.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/io/hfile/TestScannerSelectionUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n","from":"developer"},{"body":"mbautin has abandoned the revision \"[jira] [HBASE-5010] Filter HFiles based on TTL\".\n\n This was committed into trunk by @tedyu, closing the revision.\n\n @tedyu: could you please include a line of the form \"Differential Revision: D1017\" in commit messages in the future? That way the revision could be automatically marked as \"committed\" in Phabricator.\n\nREVISION DETAIL\n https://reviews.facebook.net/D1017\n","from":"developer"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n See some replies to comments inline.\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:133 Removed the line about major compactions.\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:149 Fixed.\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java:80 Fixed.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"Kannan has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n The compaction code path doesn't yet get the benefit of this optimization. Spoke with Mikhail offline, and he'll update the diff to handle this case.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Addressing Kannan's comment about doing the same optimization during compactions. Adding a compaction test to the unit test, and verifying that we don't read expired files using per-CF metrics.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/Store.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/metrics/SchemaMetrics.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/io/hfile/TestScannerSelectionUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n","from":"developer"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:219 Assuming that a default scan will go through all KVs. Please let me know if this is incorrect.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:154 This is addressed in the most recent version of the diff.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Actually, this is where I addressed Kannan's comment about compactions:\n\n https://reviews.facebook.net/D909?vs=2679&id=3015&whitespace=ignore-all\n\n (see the new line 154 in StoreScanner.java:\n\n scanners = selectScannersFrom(scanners);\n\n )\n\n I have spent quite a bit of time making sure that the unit test is testing the optimization during compactions. I am now invoking compactions in two different ways, for a total of 6 parameterized instances of the test, but it is still really quick. All of our compaction codepaths go through StoreScanner constructors, so we should have the optimization in all of them.\n\n All unit tests pass.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/Store.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/metrics/SchemaMetrics.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/io/hfile/TestScannerSelectionUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompaction.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n","from":"developer"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:155 This is why we filter out expired StoreFiles on compactions.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"Kannan has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:209 scan & columns arguments no longer used?\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:224 and we are using this.scan & this.columns here?\n\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"mbautin has commented on the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nINLINE COMMENTS\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:209 Yes, that's correct.\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java:224 Yes, this is the way it was in the original fix (89-fb version).\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"mbautin updated the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\nReviewers: Kannan, Liyin, JIRA\n\n Addressing the issue that Kannan pointed out: getScanner does not use its arguments. It turned out that getScanner was called with this.scan and this.columns in all callsites, so I have removed those arguments.\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nAFFECTED FILES\n src/main/java/org/apache/hadoop/hbase/io/hfile/LruBlockCache.java\n src/main/java/org/apache/hadoop/hbase/regionserver/KeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/MemStore.java\n src/main/java/org/apache/hadoop/hbase/regionserver/NonLazyKeyValueScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/ScanQueryMatcher.java\n src/main/java/org/apache/hadoop/hbase/regionserver/Store.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFile.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreFileScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/StoreScanner.java\n src/main/java/org/apache/hadoop/hbase/regionserver/TimeRangeTracker.java\n src/main/java/org/apache/hadoop/hbase/regionserver/metrics/SchemaMetrics.java\n src/main/java/org/apache/hadoop/hbase/util/Threads.java\n src/test/java/org/apache/hadoop/hbase/io/hfile/TestScannerSelectionUsingTTL.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompaction.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestCompoundBloomFilter.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestMemStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestQueryMatcher.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStore.java\n src/test/java/org/apache/hadoop/hbase/regionserver/TestStoreFile.java\n","from":"developer"},{"body":"Kannan has accepted the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\n awesome optimization and nice work!\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n","from":"developer"},{"body":"Can we close this? Its been applied to trunk. Its not in fb-0.89?","from":"developer"},{"body":"Actually... Does this work together with HBASE-4721 correctly?\nHBASE-4721 allows keeping delete markers longer than the TTL set for the store.\n","from":"developer"},{"body":"nspiegelberg has committed the revision \"[jira] [HBASE-5010] [89-fb] Filter HFiles based on TTL\".\n\nREVISION DETAIL\n https://reviews.facebook.net/D909\n\nCOMMIT\n https://reviews.facebook.net/rHBASEEIGHTNINEFBBRANCH1232732\n","from":"developer"},{"body":"@Stack: there a small item pending here. The 89-fb patch contains more testing for the compaction case, which has to be ported to trunk.\n\n@Lars: I will work with Prakash to ensure this does not break HBASE-4721. I believe HBASE-4721 is only used for a replication approach that is still in development.\n\nI think we'll keep this JIRA open until the two above items are resolved, unless anyone objects.\n","from":"developer"},{"body":"A follow-up fix was submitted as part of HBASE-5274 to bring the trunk fix for this issue to parity with the 89-fb fix. Resolving.","from":"developer"},{"body":"This change is doesn't break HBASE-4721.\n\nHBASE-4721 introduced another parameter called hbase.hstore.time.to.purge.deletes to keep deletes even after major compactions. But hbase.hstore.time.to.purge.deletes doesn't override the TTL of the store.\n\nPasting the comment from code which hopefully makes it clear that this diff works with HBASE-4721\n\n // By default, when hbase.hstore.time.to.purge.deletes is 0ms, a delete\n // marker is always removed during a major compaction. If set to non-zero\n // value then major compaction will try to keep a delete marker around for\n // the given number of milliseconds. We want to keep the delete markers\n // around a bit longer because old puts might appear out-of-order. For\n // example, during log replication between two clusters.\n //\n // If the delete marker has lived longer than its column-family's TTL then\n // the delete marker will be removed even if time.to.purge.deletes has not\n // passed. This is because all the Puts that this delete marker can influence\n // would have also expired. (Removing of delete markers on col family TTL will\n // not happen if min-versions is set to non-zero)\n //","from":"developer"},{"body":"Thanks for clarifying Prakash.","from":"developer"},{"body":"Why HBASE-5510 updates coming in this JIRA?\n\nRegards\nRam\n\n","from":"developer"},{"body":"@Ram: I don't see any mentions of HBASE-5510 in this JIRA, except for your comment. What updates are you referring to?","from":"developer"},{"body":"@Mikhail\nThe commit related updates that usually comes up once a commit is done was appearing in this JIRA. But it was for HBASE-5510. Ted removed them as it was not related to this JIRA. May be that deleted part you are not able to view now. :)\nSorry if the above comment had confused you.","from":"developer"},{"body":"Actually here is the reason for those confusing updates. Ted seems to have specified HBASE-5010 instead of HBASE-5510 in the commit message.\n\ncommit 5d773d9fa176cb056b993fdff8a2853f75315ec8\nAuthor: tedyu \nDate: Mon Mar 5 10:41:03 2012\n\n HBASE-5010 Pass region info in LoadBalancer.randomAssignment(List servers) (Anoop Sam \n \n git-svn-id: http://svn.apache.org/repos/asf/hbase/trunk@1297155 13f79535-47bb-0310-9956-ffa450edef68\n\n","from":"developer"}],"created":"2011-12-12T19:09:06.000+0000","description":"In ScanWildcardColumnTracker we have\n\n{code:java}\n \n this.oldestStamp = EnvironmentEdgeManager.currentTimeMillis() - ttl;\n\n ...\n\n private boolean isExpired(long timestamp) {\n return timestamp < oldestStamp;\n }\n{code}\n\nbut this time range filtering does not participate in HFile selection. In one real case this caused next() calls to time out because all KVs in a table got expired, but next() had to iterate over the whole table to find that out. We should be able to filter out those HFiles right away. I think a reasonable approach is to add a \"default timerange filter\" to every scan for a CF with a finite TTL and utilize existing filtering in StoreFile.Reader.passesTimerangeFilter.\n","issue_id":"12534821","key":"HBASE-5010","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-01-27T02:07:06.000+0000","role":"fixed_distractor","summary":"Filter HFiles based on TTL"} {"case_id":"12537717","cluster":"DISTRACTOR-HBASE-5153","comments":[{"body":"Plz share your comments, thank you.","created":"2012-01-09T08:54:16.988+0000"},{"body":"Your idea is correct.\nFew comments\n{code}\nCopyright 2010 The Apache Software Foundation\n{code}\nThis can be removed.\nAlso add javadoc for the new class added.","created":"2012-01-09T09:52:49.695+0000"},{"body":"Thank you, Rama. I update the patch:)","created":"2012-01-09T10:07:51.873+0000"},{"body":"Can you add a unit test ?\nTesting this in cluster is also desirable. ","created":"2012-01-09T12:11:08.894+0000"},{"body":"Thanks, Ted. I'll add the unit test code immediately.","created":"2012-01-09T12:33:37.743+0000"},{"body":"{code}\n+ if (this.closed)\n+ throw new ConnectionClosedException(toString() + \" closed\");\n{code}\nPlease enclose the throw statement in curly braces.\n\nPlease also prepare patch for TRUNK.","created":"2012-01-09T15:23:29.955+0000"},{"body":"Only do that in flushCommits maybe not enough. I'll go though the code and give a more considerate approach. and also will give a patch for TRUNK.","created":"2012-01-10T14:22:53.417+0000"},{"body":"@Jieshan:\nCan you prepare a patch for trunk ?","created":"2012-01-11T14:22:33.637+0000"},{"body":"Patch looks good. I like your addition of a specific Exception for closed state.\n\nDoes this have to be public Jieshan?\n\n{code}\ngetRegionServerWithRetries\n{code}\n\nSame for processBatch and getRegionLocation.\n\nIf public should be in HTableInterface but they seem implementation methods rather than something that should be part of public interface.\n\nA style nit -- i.e. not important but if you are going to redo the patch you miight want to address it -- is that you do this in handleConnectionClosedException....\n\n{code}\n+ if (ioe instanceof ConnectionClosedException) {\n{code}\n\nand the whole method is dealing with the case where above is true. I'd suggest that you might do:\n\n{code}\nif (!(ioe instanceof ConnectionClosedException)) return;\n{code}\n\n... then you save a whole indent and its clear that the method is all about dealing with ConnectionClosedException.\n\nIs it right including this in HTable?\n\n{code}\ngetPauseTime\n{code}\n\nIn trunk that is in a new ConnectionUtils class. Maybe you have to do it for 0.90?\n\nI'm wondering if the class ConnectionClosedException needs to be public also? Its only used in this package, right?\n\n\n","created":"2012-01-11T17:27:12.795+0000"},{"body":"@Ted, the patch for TRUNK seems very different, and i still need some time to check it. hope i can provide today:)\n\n@Stack, I think ConnectionUtils is reasonable. I can add it:). I will update the patch.\n\nThank you all.","created":"2012-01-12T03:54:05.450+0000"},{"body":"Jieshan's patch for TRUNK.","created":"2012-01-12T17:53:27.622+0000"},{"body":"For the trunk patch:\n{code}\n+ * @param actions\n+ * The collection of actions.\n+ * @param tableName\n+ * Name of the hbase table\n{code}\nThe explanation should be on the same line as parameter name.\n{code}\n+ * Check the connection. If it has been closed, we should recreate the\n+ * Connection. See HBASE-5153.\n+ * \n+ * @param ioe\n+ * The IOE throwed by the HConnection methods.\n{code}\nConnection on second line above should start with lower case 'c'.\nthrowed should be thrown.\n{code}\n+ throw new RetriesExhaustedException(\n+ \"Retry to recreate connection in HTable, but failed.\");\n{code}\nPlease include the value of tries in the exception message.\n\nGood work.","created":"2012-01-12T18:13:11.924+0000"},{"body":"I updated the patches. ","created":"2012-01-13T03:48:16.453+0000"},{"body":"+1 on latest patches.","created":"2012-01-13T06:17:20.735+0000"},{"body":"The trunk and 0.92 patch won't work with HBASE-4805.\nI would rather see a way to be able to reuse the connection rather than getting a new one as this defeats the whole idea of HBASE-4805.\n","created":"2012-01-13T06:23:00.792+0000"},{"body":"Is the only time when this happens a loss of the ZK connection?\nAll that a new HConnectionImplementation is doing is a call to setupZookeeperTrackers.\nHConnectionImplementation.abort already does this. Why would that succeed in the new Connection? Is it because we retry?\n\nMaybe HConnectionImplementation.resetZooKeeperTrackers should have the retry logic.\n\n-1 on current patches.\n","created":"2012-01-13T06:34:39.856+0000"},{"body":"Maybe retry in resetZooKeeperTrackers is better. We try our best to re-use the original connection. The worst case is even if we retried max times and still fail, then abort, but i think we're responsible for letting the user level know this. ","created":"2012-01-13T08:38:32.221+0000"},{"body":"@Lars:\nHbase-4805 is not in 0.90\nSo in spirit, this Jira should be able to go into 0.90\nWhat do you think ?","created":"2012-01-13T13:00:33.128+0000"},{"body":"Thanks, Ted.\n\nWithout HBASE-3065, 0.90 doesn't handle the ConnectionLossException correctly. Consider the below case:\n1. Somewhere trigger a HConnection#abort. \n2. Suppose the check of \"if (t instanceof KeeperException.SessionExpiredException)\" is true. Then called the resetZooKeeperTrackers().\n3. A ConnectionLossException occur during ZookeeperNodeTracker#start. then trigger a new HConnection#abort. At this scenario, the previous abort may print a log of \n\"Reconnected successfully. This disconnect could have been caused by a network partition or a long-running GC pause......\"\n4. The new abort carry a Throwable with a type which is not KeeperException.SessionExpiredException. so this time abort directly.\n\nIt seems a recursion here.\n \nEither re-use the old connection by resetZooKeeperTrackers, or re-create the connection, the ZookeeperWatcher will be a new one. So I still think the patch for 0.90 is reasonable.\n\nTrunk patch will be made big changes.\n\nSo any other good suggestions? Thanks.\n","created":"2012-01-13T14:20:39.424+0000"},{"body":"I think we can solve HBASE-5084 first by not bundling the HConnection and ExecutorService.\nAfter that, this JIRA's patch for TRUNK can be modified to take effect when HConnection is managed internally.","created":"2012-01-13T14:58:16.679+0000"},{"body":"@Ted: This could go into 0.90. Downside then is that 0.90 would behave quite differently.\nI feel ambivalent about HBASE-5084 (seems we're over engineering this, but if you think we need that, then let's fix that first and then revisit this).\n\n@Jieshan & Ted: Maybe have some retry logic resetZooKeeperTrackers *and* give HConnection an extra \"reset\" method? That way app code can redo the reset if needed.\n","created":"2012-01-13T16:46:57.471+0000"},{"body":"@Lars\nOk Lars. Tomorrow may be we can upload a new patch with a retry logic in resetZooKeeperTrackers. I will discuss with Jieshan tomorrow. Thanks Lars.","created":"2012-01-13T17:53:45.940+0000"},{"body":"Thanks Ram (and Jieshan and Ted).\nIn my previous comments I did not mean to be abrasive. Thanks for finding this problem and working on a patch for it (I completely missed this when I worked on HBASE-4805).\n","created":"2012-01-13T18:00:57.222+0000"},{"body":"No problem Lars. Its good to get the best suggestions and ideas. :)","created":"2012-01-13T18:04:41.696+0000"},{"body":"This is the new patch for 90 according to Lars' suggestion. Tell me if anything still need to change.\n\nThank you Ted, Lars & Ram.","created":"2012-01-14T11:50:46.978+0000"},{"body":"For patch v5, setupZookeeperTrackers():\n{code}\n- this.rootRegionTracker.start();\n+ return this.rootRegionTracker.start(allowAbort)\n+ && this.masterAddressTracker.start(allowAbort);\n{code}\nIf rootRegionTracker.start() returns false, masterAddressTracker.start() would be skipped.\nSince setupZookeeperTrackers() is called by HConnectionImplementation ctor, I suggest making setupZookeeperTrackers() reentrant (checking masterAddressTracker and rootRegionTracker against null) and calling it in HConnectionImplementation ctor.","created":"2012-01-14T15:35:06.220+0000"},{"body":"{code}\n+ * Close the original connection and create a new one.\n+ * @throws ZooKeeperConnectionException if unable to connect to zookeeper\n+ */\n+ public void resetConnection() throws ZooKeeperConnectionException;\n{code}\nThe first line should read 'Closes the original connection and creates a new one.'\nSince resetZooKeeperTrackersWithRetries() is called underneath, I think naming this new method resetZooKeeperTrackersWithRetries() seems better.\n{code}\n+/**\n+ * Thrown when HConnection has been closed.\n+ */\n+public class InvalidConnectionException extends IOException {\n{code}\nGoing over the patch, the new exception is indeed used to indicate closed connection. I think naming it ClosedConnectionException is better.","created":"2012-01-14T15:47:05.296+0000"},{"body":"Thanks Jieshan. This patch is great.\n+1 after Ted's comments are fixed.\n\nAre you planning to make a 0.92 and trunk patch as well?\n(After Stacks changes in trunk that might be not completely trivial).","created":"2012-01-14T17:32:57.723+0000"},{"body":"Oh, and should probably change the title of this jira.","created":"2012-01-14T17:36:31.831+0000"},{"body":"Changes:\n1. Add a method \"unregisterListener\" in ZookeeperWatcher. This will be called in ZookeeperNodeTracker's stop method.\n2. Stop the trackers if start failed during the retries.\n\nThank you , Ted.\n\n@Lars, I will make the patches for trunk and 92, patch for 90 seems more urgent:)..Thank you. ","created":"2012-01-16T05:56:24.164+0000"},{"body":"Minor comment:\nIn setupZookeeperTrackers(), after calling stop(), the masterAddressTracker / rootRegionTracker should be set to null.\n\nPlease run patch v6 over test suite.\n\nThanks","created":"2012-01-16T06:07:30.255+0000"},{"body":"The test is still running. Before I get the results, I want your comments:).\nThank you.","created":"2012-01-16T06:11:05.365+0000"},{"body":"I will upload the new patch after the tests finish(Including this minor change), Thanks, Ted.","created":"2012-01-16T06:42:28.262+0000"},{"body":"Ran the tests again, got the same results: 5 tests failed due to the hostName problem. Please find the results from the attachment \"TestResults-hbase5153.out\".","created":"2012-01-16T15:57:30.649+0000"},{"body":"@Jieshan:\nPlease find a machine which has access to internet to run the test suite.\nmaven needs to download artifacts.\n\nI ran the 5 tests on MacBook and they passed.\n{code}\n 839 mt -Dtest=TestClockSkewDetection\n 840 mt -Dtest=TestScanner\n 841 mt -Dtest=TestCatalogTrackerOnCluster\n 842 mt -Dtest=TestCatalogTracker\n{code}\n+1 on latest patch.","created":"2012-01-16T16:37:56.957+0000"},{"body":"+1 on latest patch","created":"2012-01-16T22:51:44.978+0000"},{"body":"Patch for TRUNK.","created":"2012-01-17T00:41:13.130+0000"},{"body":"+1 on trunk patch. 0.92 is similar?","created":"2012-01-17T00:50:01.765+0000"},{"body":"Integrated to 0.90 and TRUNK.\n\nWaiting for 0.92 security build to pass before integrating to 0.92\n\nThanks for the patch Jieshan.\n\nThanks for the review Lars.","created":"2012-01-17T04:07:39.963+0000"},{"body":"Patch for 0.92 branch","created":"2012-01-17T04:33:26.913+0000"},{"body":"Integrated to 0.92 branch.","created":"2012-01-17T04:43:21.620+0000"},{"body":"Thank you, Ted.","created":"2012-01-17T04:47:54.558+0000"},{"body":"Thanks a lot Ted. Nice of you .:)","created":"2012-01-17T04:52:51.666+0000"},{"body":"When I ran TestMergeTool on TRUNK, I got:\n{code}\n\"main\" prio=5 tid=102801000 nid=0x100601000 waiting on condition [1005fb000]\n java.lang.Thread.State: TIMED_WAITING (sleeping)\n\tat java.lang.Thread.sleep(Native Method)\n\tat java.lang.Thread.sleep(Thread.java:302)\n\tat java.util.concurrent.TimeUnit.sleep(TimeUnit.java:328)\n\tat org.apache.hadoop.hbase.util.RetryCounter.sleepUntilNextRetry(RetryCounter.java:55)\n\tat org.apache.hadoop.hbase.zookeeper.RecoverableZooKeeper.exists(RecoverableZooKeeper.java:171)\n\tat org.apache.hadoop.hbase.zookeeper.ZKUtil.watchAndCheckExists(ZKUtil.java:230)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:80)\n\t- locked <784a83930> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <7854b9898> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <7854b98e0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <7854b9928> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78f991ad0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790b9f540> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790b9f560> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790b9f580> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790b9f5a0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78f98ae20> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78f98ae40> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78f98ae60> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78e8b6ac8> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78e8b6ae8> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78e8b6b08> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e640> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e660> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e600> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e620> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e5c0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e5e0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e580> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e5a0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e540> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e560> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e500> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.(HConnectionManager.java:577)\n\tat org.apache.hadoop.hbase.client.HConnectionManager.getConnection(HConnectionManager.java:184)\n\t- locked <790c46258> (a org.apache.hadoop.hbase.client.HConnectionManager$1)\n\tat org.apache.hadoop.hbase.client.HBaseAdmin.(HBaseAdmin.java:98)\n\tat org.apache.hadoop.hbase.client.HBaseAdmin.checkHBaseAvailable(HBaseAdmin.java:1570)\n\tat org.apache.hadoop.hbase.util.Merge.run(Merge.java:94)\n\tat org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:65)\n\tat org.apache.hadoop.hbase.util.TestMergeTool.mergeAndVerify(TestMergeTool.java:190)\n\tat org.apache.hadoop.hbase.util.TestMergeTool.testMergeTool(TestMergeTool.java:268)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat junit.framework.TestCase.runTest(TestCase.java:168)\n\tat junit.framework.TestCase.runBare(TestCase.java:134)\n\tat junit.framework.TestResult$1.protect(TestResult.java:110)\n\tat junit.framework.TestResult.runProtected(TestResult.java:128)\n\tat junit.framework.TestResult.run(TestResult.java:113)\n\tat junit.framework.TestCase.run(TestCase.java:124)\n\tat junit.framework.TestSuite.runTest(TestSuite.java:243)\n\tat junit.framework.TestSuite.run(TestSuite.java:238)\n\tat org.junit.internal.runners.JUnit38ClassRunner.run(JUnit38ClassRunner.java:83)\n\tat org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n\tat org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n\tat org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n\tat org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n\tat org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\n{code}","created":"2012-01-17T16:13:08.726+0000"},{"body":"Reverted from 0.92 and TRUNK due to failed Jenkins builds.","created":"2012-01-17T17:17:26.922+0000"},{"body":"Was that repeatable? Might be a problem with the test (relying on the fact that an HConnection would just give up after one try).","created":"2012-01-17T18:09:21.292+0000"},{"body":"TestMergeTool hung in several builds.\nThat was why I ran it on MacBook.","created":"2012-01-17T18:14:33.650+0000"},{"body":"FYI: [^HBASE-5153-V6-90-minorchange.patch] also fixes HBASE-5289 (linked)","created":"2012-01-26T21:03:28.184+0000"},{"body":"Also it occurred to me that another nice change would be to be able specify the retry count for resetZooKeeperTrackersWithRetries different from the other operations. \nThe thinking is this:\nWhile the ZK is not reachable the HConnection (and any other HConnection) is essentially not usable. In some settings it might be good to have the connection just sit there, and retry until the connection is bad. Maybe for another jira.\n\nWhere are we with this generally?\nIs it just TestMergeTool hanging? If so I'll have a look at it today.","created":"2012-01-26T21:21:00.996+0000"},{"body":"There were two failed tests:\nhttps://builds.apache.org/job/HBase-0.92-security/81/\n\nIf you can resolve the hanging TestMergeTool, that would be great.\n\nI am on-call this week, FYI","created":"2012-01-26T21:30:48.270+0000"},{"body":"So is this change in 0.90 now? I'm confused. Should revert it from there too, I guess.\nI will see what's up with TestMergeTool in trunk now.","created":"2012-01-27T01:36:40.121+0000"},{"body":"So here's the problem. This is hanging while validating that HBase is not running via HBaseAdmin.checkHBaseAvailable, which just attempts to create a new HBaseAdmin after it sets hbase.client.retries.number to 1. However HConnectionImpl caches hbase.client.retries.number in numRetries, and hence if ZK is not running resetZooKeeperTrackersWithRetries will retry for a while.\nThe simplest fix would be for resetZooKeeperTrackersWithRetries to ignore he cached setting and to retrieve the value again from the setting. While I am at it, I'll also add another option to a different number of retries here.","created":"2012-01-27T01:57:52.504+0000"},{"body":"Thanks for tracking down the issue, Lars.\nIf you can upload the latest 5153-trunk.txt to reviewboard first followed by your new patch, that would help us know your changes easily.","created":"2012-01-27T02:39:10.014+0000"},{"body":"Sure... There's a bit more to this too. resetZooKeeperTrackersWithRetries on its last try calls setupZookeeperTrackers with allow aborts, which will call resetZooKeeperTrackersWithRetries again. Leading to an endless loop. Need to think about how to refactor this.","created":"2012-01-27T03:04:51.021+0000"},{"body":"It should not leading to an endless loop. Unless, each retry will get a ZookeeperLossException. If this exception happened for long time, Zookeeper must has some problem. so when create a new Zookeeper instance, it already thrown a Exception. So it won't be an endless loop:\n{noformat}\nif ((t instanceof KeeperException.SessionExpiredException)\n || (t instanceof KeeperException.ConnectionLossException)) {\n try {\n LOG.info(\"This client just lost it's session with ZooKeeper, trying\" +\n \" to reconnect.\");\n resetZooKeeperTrackersWithRetries();\n LOG.info(\"Reconnected successfully. This disconnect could have been\" +\n \" caused by a network partition or a long-running GC pause,\" +\n \" either way it's recommended that you verify your environment.\");\n return;\n } catch (ZooKeeperConnectionException e) {\n LOG.error(\"Could not reconnect to ZooKeeper after session\" +\n \" expiration, aborting\");\n t = e;\n }\n }\n if (t != null) LOG.fatal(msg, t);\n else LOG.fatal(msg);\n HConnectionManager.deleteStaleConnection(this);\n{noformat}","created":"2012-01-27T03:40:21.390+0000"},{"body":"As discussed with Ted. Trunk and 92 already including a retry logic in RecoverableZooKeeper. So that makes the retry logic in resetZooKeeperTrackersWithRetries less important.","created":"2012-01-27T03:45:16.200+0000"},{"body":"See the stack trace I pasted here:\nhttps://issues.apache.org/jira/browse/HBASE-5153?focusedCommentId=13187774&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-13187774\n\nbq. If this exception happened for long time, Zookeeper must has some problem.\nWe should be prepared when the above indeed happens. I am sure this scenario is possible.\n\nSee also this part of the code:\n{code}\n+ } catch (ZooKeeperConnectionException zkce) {\n+ if (isLastTime) {\n+ throw zkce;\n+ }\n+ }\n{code}\n","created":"2012-01-27T03:49:10.263+0000"},{"body":"Thanks, Ted...I will take a look at that stack trace.\n \nIf a ZooKeeperConnectionException thrown by the below code:\n{noformat}\n try {\n if (setupZookeeperTrackers(isLastTime)) {\n break;\n }\n } catch (ZooKeeperConnectionException zkce) {\n if (isLastTime) {\n throw zkce;\n }\n }\n{noformat}\n\nIf will be catched in abort method, then calling LOG.fatal(msg, t);\n\nNo problem here. Don't know whether I get you correctly:(.","created":"2012-01-27T03:59:30.308+0000"},{"body":"I have a new patch... Testing it now.\nIt is mostly what you have Jieshan, but it does not need all the changes to the ZookeeperNodeTracker and subclasses.\nI'll attach it soon... Then please let me know what you think.","created":"2012-01-27T04:14:23.126+0000"},{"body":"There was also some weird stuff in HBaseAdmin.checkHBaseAvailable. It did set the client retry to 1, but left the the retry count on in RecoverableZookeeper, which leads to long, unnecessary waits.","created":"2012-01-27T04:16:00.709+0000"},{"body":"Please let me know what you think of 5153-trunk-v2.txt... Thanks","created":"2012-01-27T04:20:40.245+0000"},{"body":"@Jieshan: The endless loop happens when ZK is actually down.","created":"2012-01-27T04:22:31.361+0000"},{"body":"bq. As discussed with Ted. Trunk and 92 already including a retry logic in RecoverableZooKeeper. So that makes the retry logic in resetZooKeeperTrackersWithRetries less important.\n\nAre you guys saying we do not need this in 0.92+?","created":"2012-01-27T04:26:47.574+0000"},{"body":"Thanks Lars. It's nice of you:)","created":"2012-01-27T04:26:58.500+0000"},{"body":"\"The endless loop happens when ZK is actually down.\"\nIf ZK is actually down, the below code will throw a Exception:\n this.zooKeeper = getZooKeeperWatcher();\nThen catched by the below code:\n{noformat}\n try {\n LOG.info(\"This client just lost it's session with ZooKeeper, trying\" +\n \" to reconnect.\");\n resetZooKeeperTrackersWithRetries();\n LOG.info(\"Reconnected successfully. This disconnect could have been\" +\n \" caused by a network partition or a long-running GC pause,\" +\n \" either way it's recommended that you verify your environment.\");\n return;\n } catch (ZooKeeperConnectionException e) {\n LOG.error(\"Could not reconnect to ZooKeeper after session\" +\n \" expiration, aborting\");\n t = e;\n }\n if (t != null) LOG.fatal(msg, t);\n else LOG.fatal(msg);\n HConnectionManager.deleteStaleConnection(this);\n{noformat}\n\nIt should not be a endless loop. Does that make sense?\n","created":"2012-01-27T04:30:27.180+0000"},{"body":"bq. Are you guys saying we do not need this in 0.92+?\n\nWe need this, but the patch for 0.92+ should take notice of this:)","created":"2012-01-27T04:35:05.313+0000"},{"body":"{code}\n+ if (isResettingZKTrackers) {\n+ return;\n+ }\n{code}\nI was thinking about something similar to the above.\n\nIn resetZooKeeperTrackersWithRetries(), shall we call zooKeeper.close() before resetting zooKeeper?\n{code}\n this.zooKeeper.close();\n this.zooKeeper = null;\n{code}\n\nThanks for the help.\n\nI think we need an addendum for 0.90","created":"2012-01-27T04:46:07.911+0000"},{"body":"@Lars:\nAdding a \"isResettingZKTrackers\" sounds good to me. One doubt: Is it necessary to add the keyword of \"volatile\"?","created":"2012-01-27T04:47:10.183+0000"},{"body":"@Jieshan: Hmm... I do see the endless loop in the debugger. It happens from HBaseAdmin.checkHBaseAvailable when HBase is actually down. We get into resetZooKeeperTrackersWithRetries, which calls setupZookeeperTrackers, which causes abort to be called, which calls resetZooKeeperTrackersWithRetries.\nI am not sure getZooKeeperWatcher() would throw if ZK is not available.\n\n@Ted: calling this.zooKeeper.close() (if it is not null) first seems prudent.\n\n@Jieshan: Good catch, yes it should be volatile.\n","created":"2012-01-27T04:52:57.287+0000"},{"body":"@Lars:\nDuring the retry, if we get any exceptions, Zookeeper and the Trackers also need to close.\n{noformat}\n+ try {\n+ setupZookeeperTrackers();\n+ break;\n+ } catch (ZooKeeperConnectionException zkce) {\n+ if (tries >= this.numRetries) {\n+ throw zkce;\n+ }\n+ }\n{noformat}","created":"2012-01-27T06:48:37.437+0000"},{"body":"@Lars:\nOne more doubt basing on the previous comment, regading on the below code:\n{noformat}\n+ try {\n+ setupZookeeperTrackers();\n+ break;\n+ } catch (ZooKeeperConnectionException zkce) {\n+ if (tries >= this.numRetries) {\n+ throw zkce;\n+ }\n+ }\n{noformat}\nif ZookeeperNodeTracker#start get an exception, then catches and calls abort(Doesn't throw it). The above retries will be break under this situation. So this retries takes no effects.Correct me if am wrong. ","created":"2012-01-27T08:35:31.645+0000"},{"body":"This is the addendum for 90.","created":"2012-01-27T10:22:11.329+0000"},{"body":"+1 on addendum for 0.90","created":"2012-01-27T14:43:18.160+0000"},{"body":"@Ted\nThanks for the review.\n@Lars\nCan you review the 0.90 patch and provide your comments.\nIf it is ok i can integrate to 0.90 and take another RC. :)\n","created":"2012-01-27T14:47:01.186+0000"},{"body":"+1 on addendum. Technically we should change ZooKeeperNodeTracker.start back to be void, but that is not necessary.\nI'll update the trunk patch soon.","created":"2012-01-27T16:43:36.746+0000"},{"body":"@Jieshan... I see your point. Need to think about that.","created":"2012-01-27T16:45:41.808+0000"},{"body":"Actually then the addendum is not right either. If ZooKeeperNodeTracker.start gets an exception it'll call abort (and hence not return false)","created":"2012-01-27T16:57:12.619+0000"},{"body":"I take this back... The 0.90 addendum is good. If abort is called from the reset method it'll return right away, and in ZooKeeperNodeTracker.start fall through to return false.\n\nThis entire thing is too complicated. Let's hold off on the 0.92 and trunk patches, but leave it in 0.90 (although I am not a big fan of diverging code bases either).","created":"2012-01-27T19:04:09.310+0000"},{"body":"Dropping Fix versions as Lars suggested.","created":"2012-01-27T19:37:08.337+0000"},{"body":"Here's a minimal patch for trunk.\n\nIf the HConnection is not usable for any reason, the using application can attempt to reset it.\nI kept CloseConnectionException and the idea that ZooKeeperWatcher can unregister a listener.\n\nIf ZK is really not available that is probably an event that the application has to deal with anyway.\n\nLong term HConnection should *not* hold persistent connections to ZK anyway, but only establish a connection when needed.","created":"2012-01-27T22:43:17.466+0000"},{"body":"Noticed one flaw in the 0.90 addendum patch. Should be\n{code}\n} catch (ZooKeeperConnectionException zkce) {\n if (tries+1 >= this.numRetries) {\n throw zkce;\n }\n}\n{code}\n(note the +1 in there)","created":"2012-01-27T22:44:54.313+0000"},{"body":"For resetConnection() in the new patch, I see:\n{code}\n+ this.closed = false;\n{code}\nIn HConnectionImplementation.close(), we have:\n{code}\n if (master != null) {\n if (stopProxy) {\n HBaseRPC.stopProxy(master);\n }\n master = null;\n{code}\nThis seems to imply that once a connection is closed, we cannot simply change this.closed to false in resetConnection().","created":"2012-01-27T23:08:44.225+0000"},{"body":"That is just guard against closed = true in HConnectionImplementation.abort().\nabort probably not set closed to true, but then a bunch of other places would need to be changed too.","created":"2012-01-27T23:37:41.341+0000"},{"body":"@Lars\nAddressing your comments.\nBased on your feedback will commit the patch if it is ok.","created":"2012-01-28T05:52:52.425+0000"},{"body":"@Ram: Sure. I now feel the whole logic is too complicated, but we'll not fix that any time soon. So +1 on addendum.","created":"2012-01-28T06:27:44.834+0000"},{"body":"@Lars\nI understand your point. But as this JIRA is now committed in 0.90 so this addendum is needed.\nI think we can ensure such things doesn't happen for other JIRAs.\nGood on you Lars.","created":"2012-01-28T06:43:52.437+0000"},{"body":"This got committed to 0.90.6.\nSorry for not resolving it at that time.","created":"2012-03-19T04:24:55.978+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T11:53:58.270+0000"}],"conversations":[{"body":"HBASE-4893 is related to this issue. In that issue, we know, if multi-threads share a same connection, once this connection got abort in one thread, the other threads will got a \"HConnectionManager$HConnectionImplementation@18fb1f7 closed\" exception.\n\nIt solve the problem of \"stale connection can't removed\". But the orignal HTable instance cann't be continue to use. The connection in HTable should be recreated.\n\nActually, there's two aproach to solve this:\n1. In user code, once catch an IOE, close connection and re-create HTable instance. We can use this as a workaround.\n2. In HBase Client side, catch this exception, and re-create connection.\n","from":"reporter","subject":"Add retry logic in HConnectionImplementation#resetZooKeeperTrackers"},{"body":"Plz share your comments, thank you.","from":"developer"},{"body":"Your idea is correct.\nFew comments\n{code}\nCopyright 2010 The Apache Software Foundation\n{code}\nThis can be removed.\nAlso add javadoc for the new class added.","from":"developer"},{"body":"Thank you, Rama. I update the patch:)","from":"developer"},{"body":"Can you add a unit test ?\nTesting this in cluster is also desirable. ","from":"developer"},{"body":"Thanks, Ted. I'll add the unit test code immediately.","from":"developer"},{"body":"{code}\n+ if (this.closed)\n+ throw new ConnectionClosedException(toString() + \" closed\");\n{code}\nPlease enclose the throw statement in curly braces.\n\nPlease also prepare patch for TRUNK.","from":"developer"},{"body":"Only do that in flushCommits maybe not enough. I'll go though the code and give a more considerate approach. and also will give a patch for TRUNK.","from":"developer"},{"body":"@Jieshan:\nCan you prepare a patch for trunk ?","from":"developer"},{"body":"Patch looks good. I like your addition of a specific Exception for closed state.\n\nDoes this have to be public Jieshan?\n\n{code}\ngetRegionServerWithRetries\n{code}\n\nSame for processBatch and getRegionLocation.\n\nIf public should be in HTableInterface but they seem implementation methods rather than something that should be part of public interface.\n\nA style nit -- i.e. not important but if you are going to redo the patch you miight want to address it -- is that you do this in handleConnectionClosedException....\n\n{code}\n+ if (ioe instanceof ConnectionClosedException) {\n{code}\n\nand the whole method is dealing with the case where above is true. I'd suggest that you might do:\n\n{code}\nif (!(ioe instanceof ConnectionClosedException)) return;\n{code}\n\n... then you save a whole indent and its clear that the method is all about dealing with ConnectionClosedException.\n\nIs it right including this in HTable?\n\n{code}\ngetPauseTime\n{code}\n\nIn trunk that is in a new ConnectionUtils class. Maybe you have to do it for 0.90?\n\nI'm wondering if the class ConnectionClosedException needs to be public also? Its only used in this package, right?\n\n\n","from":"developer"},{"body":"@Ted, the patch for TRUNK seems very different, and i still need some time to check it. hope i can provide today:)\n\n@Stack, I think ConnectionUtils is reasonable. I can add it:). I will update the patch.\n\nThank you all.","from":"developer"},{"body":"Jieshan's patch for TRUNK.","from":"developer"},{"body":"For the trunk patch:\n{code}\n+ * @param actions\n+ * The collection of actions.\n+ * @param tableName\n+ * Name of the hbase table\n{code}\nThe explanation should be on the same line as parameter name.\n{code}\n+ * Check the connection. If it has been closed, we should recreate the\n+ * Connection. See HBASE-5153.\n+ * \n+ * @param ioe\n+ * The IOE throwed by the HConnection methods.\n{code}\nConnection on second line above should start with lower case 'c'.\nthrowed should be thrown.\n{code}\n+ throw new RetriesExhaustedException(\n+ \"Retry to recreate connection in HTable, but failed.\");\n{code}\nPlease include the value of tries in the exception message.\n\nGood work.","from":"developer"},{"body":"I updated the patches. ","from":"developer"},{"body":"+1 on latest patches.","from":"developer"},{"body":"The trunk and 0.92 patch won't work with HBASE-4805.\nI would rather see a way to be able to reuse the connection rather than getting a new one as this defeats the whole idea of HBASE-4805.\n","from":"developer"},{"body":"Is the only time when this happens a loss of the ZK connection?\nAll that a new HConnectionImplementation is doing is a call to setupZookeeperTrackers.\nHConnectionImplementation.abort already does this. Why would that succeed in the new Connection? Is it because we retry?\n\nMaybe HConnectionImplementation.resetZooKeeperTrackers should have the retry logic.\n\n-1 on current patches.\n","from":"developer"},{"body":"Maybe retry in resetZooKeeperTrackers is better. We try our best to re-use the original connection. The worst case is even if we retried max times and still fail, then abort, but i think we're responsible for letting the user level know this. ","from":"developer"},{"body":"@Lars:\nHbase-4805 is not in 0.90\nSo in spirit, this Jira should be able to go into 0.90\nWhat do you think ?","from":"developer"},{"body":"Thanks, Ted.\n\nWithout HBASE-3065, 0.90 doesn't handle the ConnectionLossException correctly. Consider the below case:\n1. Somewhere trigger a HConnection#abort. \n2. Suppose the check of \"if (t instanceof KeeperException.SessionExpiredException)\" is true. Then called the resetZooKeeperTrackers().\n3. A ConnectionLossException occur during ZookeeperNodeTracker#start. then trigger a new HConnection#abort. At this scenario, the previous abort may print a log of \n\"Reconnected successfully. This disconnect could have been caused by a network partition or a long-running GC pause......\"\n4. The new abort carry a Throwable with a type which is not KeeperException.SessionExpiredException. so this time abort directly.\n\nIt seems a recursion here.\n \nEither re-use the old connection by resetZooKeeperTrackers, or re-create the connection, the ZookeeperWatcher will be a new one. So I still think the patch for 0.90 is reasonable.\n\nTrunk patch will be made big changes.\n\nSo any other good suggestions? Thanks.\n","from":"developer"},{"body":"I think we can solve HBASE-5084 first by not bundling the HConnection and ExecutorService.\nAfter that, this JIRA's patch for TRUNK can be modified to take effect when HConnection is managed internally.","from":"developer"},{"body":"@Ted: This could go into 0.90. Downside then is that 0.90 would behave quite differently.\nI feel ambivalent about HBASE-5084 (seems we're over engineering this, but if you think we need that, then let's fix that first and then revisit this).\n\n@Jieshan & Ted: Maybe have some retry logic resetZooKeeperTrackers *and* give HConnection an extra \"reset\" method? That way app code can redo the reset if needed.\n","from":"developer"},{"body":"@Lars\nOk Lars. Tomorrow may be we can upload a new patch with a retry logic in resetZooKeeperTrackers. I will discuss with Jieshan tomorrow. Thanks Lars.","from":"developer"},{"body":"Thanks Ram (and Jieshan and Ted).\nIn my previous comments I did not mean to be abrasive. Thanks for finding this problem and working on a patch for it (I completely missed this when I worked on HBASE-4805).\n","from":"developer"},{"body":"No problem Lars. Its good to get the best suggestions and ideas. :)","from":"developer"},{"body":"This is the new patch for 90 according to Lars' suggestion. Tell me if anything still need to change.\n\nThank you Ted, Lars & Ram.","from":"developer"},{"body":"For patch v5, setupZookeeperTrackers():\n{code}\n- this.rootRegionTracker.start();\n+ return this.rootRegionTracker.start(allowAbort)\n+ && this.masterAddressTracker.start(allowAbort);\n{code}\nIf rootRegionTracker.start() returns false, masterAddressTracker.start() would be skipped.\nSince setupZookeeperTrackers() is called by HConnectionImplementation ctor, I suggest making setupZookeeperTrackers() reentrant (checking masterAddressTracker and rootRegionTracker against null) and calling it in HConnectionImplementation ctor.","from":"developer"},{"body":"{code}\n+ * Close the original connection and create a new one.\n+ * @throws ZooKeeperConnectionException if unable to connect to zookeeper\n+ */\n+ public void resetConnection() throws ZooKeeperConnectionException;\n{code}\nThe first line should read 'Closes the original connection and creates a new one.'\nSince resetZooKeeperTrackersWithRetries() is called underneath, I think naming this new method resetZooKeeperTrackersWithRetries() seems better.\n{code}\n+/**\n+ * Thrown when HConnection has been closed.\n+ */\n+public class InvalidConnectionException extends IOException {\n{code}\nGoing over the patch, the new exception is indeed used to indicate closed connection. I think naming it ClosedConnectionException is better.","from":"developer"},{"body":"Thanks Jieshan. This patch is great.\n+1 after Ted's comments are fixed.\n\nAre you planning to make a 0.92 and trunk patch as well?\n(After Stacks changes in trunk that might be not completely trivial).","from":"developer"},{"body":"Oh, and should probably change the title of this jira.","from":"developer"},{"body":"Changes:\n1. Add a method \"unregisterListener\" in ZookeeperWatcher. This will be called in ZookeeperNodeTracker's stop method.\n2. Stop the trackers if start failed during the retries.\n\nThank you , Ted.\n\n@Lars, I will make the patches for trunk and 92, patch for 90 seems more urgent:)..Thank you. ","from":"developer"},{"body":"Minor comment:\nIn setupZookeeperTrackers(), after calling stop(), the masterAddressTracker / rootRegionTracker should be set to null.\n\nPlease run patch v6 over test suite.\n\nThanks","from":"developer"},{"body":"The test is still running. Before I get the results, I want your comments:).\nThank you.","from":"developer"},{"body":"I will upload the new patch after the tests finish(Including this minor change), Thanks, Ted.","from":"developer"},{"body":"Ran the tests again, got the same results: 5 tests failed due to the hostName problem. Please find the results from the attachment \"TestResults-hbase5153.out\".","from":"developer"},{"body":"@Jieshan:\nPlease find a machine which has access to internet to run the test suite.\nmaven needs to download artifacts.\n\nI ran the 5 tests on MacBook and they passed.\n{code}\n 839 mt -Dtest=TestClockSkewDetection\n 840 mt -Dtest=TestScanner\n 841 mt -Dtest=TestCatalogTrackerOnCluster\n 842 mt -Dtest=TestCatalogTracker\n{code}\n+1 on latest patch.","from":"developer"},{"body":"+1 on latest patch","from":"developer"},{"body":"Patch for TRUNK.","from":"developer"},{"body":"+1 on trunk patch. 0.92 is similar?","from":"developer"},{"body":"Integrated to 0.90 and TRUNK.\n\nWaiting for 0.92 security build to pass before integrating to 0.92\n\nThanks for the patch Jieshan.\n\nThanks for the review Lars.","from":"developer"},{"body":"Patch for 0.92 branch","from":"developer"},{"body":"Integrated to 0.92 branch.","from":"developer"},{"body":"Thank you, Ted.","from":"developer"},{"body":"Thanks a lot Ted. Nice of you .:)","from":"developer"},{"body":"When I ran TestMergeTool on TRUNK, I got:\n{code}\n\"main\" prio=5 tid=102801000 nid=0x100601000 waiting on condition [1005fb000]\n java.lang.Thread.State: TIMED_WAITING (sleeping)\n\tat java.lang.Thread.sleep(Native Method)\n\tat java.lang.Thread.sleep(Thread.java:302)\n\tat java.util.concurrent.TimeUnit.sleep(TimeUnit.java:328)\n\tat org.apache.hadoop.hbase.util.RetryCounter.sleepUntilNextRetry(RetryCounter.java:55)\n\tat org.apache.hadoop.hbase.zookeeper.RecoverableZooKeeper.exists(RecoverableZooKeeper.java:171)\n\tat org.apache.hadoop.hbase.zookeeper.ZKUtil.watchAndCheckExists(ZKUtil.java:230)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:80)\n\t- locked <784a83930> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <7854b9898> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <7854b98e0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <7854b9928> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78f991ad0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790b9f540> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790b9f560> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790b9f580> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790b9f5a0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78f98ae20> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78f98ae40> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78f98ae60> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78e8b6ac8> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78e8b6ae8> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <78e8b6b08> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e640> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e660> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e600> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e620> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e5c0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e5e0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e580> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e5a0> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e540> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e560> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.resetZooKeeperTrackersWithRetries(HConnectionManager.java:625)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.abort(HConnectionManager.java:1711)\n\tat org.apache.hadoop.hbase.zookeeper.ZooKeeperNodeTracker.start(ZooKeeperNodeTracker.java:93)\n\t- locked <790d0e500> (a org.apache.hadoop.hbase.MasterAddressTracker)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.setupZookeeperTrackers(HConnectionManager.java:590)\n\t- locked <790b70180> (a org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation.(HConnectionManager.java:577)\n\tat org.apache.hadoop.hbase.client.HConnectionManager.getConnection(HConnectionManager.java:184)\n\t- locked <790c46258> (a org.apache.hadoop.hbase.client.HConnectionManager$1)\n\tat org.apache.hadoop.hbase.client.HBaseAdmin.(HBaseAdmin.java:98)\n\tat org.apache.hadoop.hbase.client.HBaseAdmin.checkHBaseAvailable(HBaseAdmin.java:1570)\n\tat org.apache.hadoop.hbase.util.Merge.run(Merge.java:94)\n\tat org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:65)\n\tat org.apache.hadoop.hbase.util.TestMergeTool.mergeAndVerify(TestMergeTool.java:190)\n\tat org.apache.hadoop.hbase.util.TestMergeTool.testMergeTool(TestMergeTool.java:268)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat junit.framework.TestCase.runTest(TestCase.java:168)\n\tat junit.framework.TestCase.runBare(TestCase.java:134)\n\tat junit.framework.TestResult$1.protect(TestResult.java:110)\n\tat junit.framework.TestResult.runProtected(TestResult.java:128)\n\tat junit.framework.TestResult.run(TestResult.java:113)\n\tat junit.framework.TestCase.run(TestCase.java:124)\n\tat junit.framework.TestSuite.runTest(TestSuite.java:243)\n\tat junit.framework.TestSuite.run(TestSuite.java:238)\n\tat org.junit.internal.runners.JUnit38ClassRunner.run(JUnit38ClassRunner.java:83)\n\tat org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n\tat org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n\tat org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n\tat org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n\tat org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\n{code}","from":"developer"},{"body":"Reverted from 0.92 and TRUNK due to failed Jenkins builds.","from":"developer"},{"body":"Was that repeatable? Might be a problem with the test (relying on the fact that an HConnection would just give up after one try).","from":"developer"},{"body":"TestMergeTool hung in several builds.\nThat was why I ran it on MacBook.","from":"developer"},{"body":"FYI: [^HBASE-5153-V6-90-minorchange.patch] also fixes HBASE-5289 (linked)","from":"developer"},{"body":"Also it occurred to me that another nice change would be to be able specify the retry count for resetZooKeeperTrackersWithRetries different from the other operations. \nThe thinking is this:\nWhile the ZK is not reachable the HConnection (and any other HConnection) is essentially not usable. In some settings it might be good to have the connection just sit there, and retry until the connection is bad. Maybe for another jira.\n\nWhere are we with this generally?\nIs it just TestMergeTool hanging? If so I'll have a look at it today.","from":"developer"},{"body":"There were two failed tests:\nhttps://builds.apache.org/job/HBase-0.92-security/81/\n\nIf you can resolve the hanging TestMergeTool, that would be great.\n\nI am on-call this week, FYI","from":"developer"},{"body":"So is this change in 0.90 now? I'm confused. Should revert it from there too, I guess.\nI will see what's up with TestMergeTool in trunk now.","from":"developer"},{"body":"So here's the problem. This is hanging while validating that HBase is not running via HBaseAdmin.checkHBaseAvailable, which just attempts to create a new HBaseAdmin after it sets hbase.client.retries.number to 1. However HConnectionImpl caches hbase.client.retries.number in numRetries, and hence if ZK is not running resetZooKeeperTrackersWithRetries will retry for a while.\nThe simplest fix would be for resetZooKeeperTrackersWithRetries to ignore he cached setting and to retrieve the value again from the setting. While I am at it, I'll also add another option to a different number of retries here.","from":"developer"},{"body":"Thanks for tracking down the issue, Lars.\nIf you can upload the latest 5153-trunk.txt to reviewboard first followed by your new patch, that would help us know your changes easily.","from":"developer"},{"body":"Sure... There's a bit more to this too. resetZooKeeperTrackersWithRetries on its last try calls setupZookeeperTrackers with allow aborts, which will call resetZooKeeperTrackersWithRetries again. Leading to an endless loop. Need to think about how to refactor this.","from":"developer"},{"body":"It should not leading to an endless loop. Unless, each retry will get a ZookeeperLossException. If this exception happened for long time, Zookeeper must has some problem. so when create a new Zookeeper instance, it already thrown a Exception. So it won't be an endless loop:\n{noformat}\nif ((t instanceof KeeperException.SessionExpiredException)\n || (t instanceof KeeperException.ConnectionLossException)) {\n try {\n LOG.info(\"This client just lost it's session with ZooKeeper, trying\" +\n \" to reconnect.\");\n resetZooKeeperTrackersWithRetries();\n LOG.info(\"Reconnected successfully. This disconnect could have been\" +\n \" caused by a network partition or a long-running GC pause,\" +\n \" either way it's recommended that you verify your environment.\");\n return;\n } catch (ZooKeeperConnectionException e) {\n LOG.error(\"Could not reconnect to ZooKeeper after session\" +\n \" expiration, aborting\");\n t = e;\n }\n }\n if (t != null) LOG.fatal(msg, t);\n else LOG.fatal(msg);\n HConnectionManager.deleteStaleConnection(this);\n{noformat}","from":"developer"},{"body":"As discussed with Ted. Trunk and 92 already including a retry logic in RecoverableZooKeeper. So that makes the retry logic in resetZooKeeperTrackersWithRetries less important.","from":"developer"},{"body":"See the stack trace I pasted here:\nhttps://issues.apache.org/jira/browse/HBASE-5153?focusedCommentId=13187774&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-13187774\n\nbq. If this exception happened for long time, Zookeeper must has some problem.\nWe should be prepared when the above indeed happens. I am sure this scenario is possible.\n\nSee also this part of the code:\n{code}\n+ } catch (ZooKeeperConnectionException zkce) {\n+ if (isLastTime) {\n+ throw zkce;\n+ }\n+ }\n{code}\n","from":"developer"},{"body":"Thanks, Ted...I will take a look at that stack trace.\n \nIf a ZooKeeperConnectionException thrown by the below code:\n{noformat}\n try {\n if (setupZookeeperTrackers(isLastTime)) {\n break;\n }\n } catch (ZooKeeperConnectionException zkce) {\n if (isLastTime) {\n throw zkce;\n }\n }\n{noformat}\n\nIf will be catched in abort method, then calling LOG.fatal(msg, t);\n\nNo problem here. Don't know whether I get you correctly:(.","from":"developer"},{"body":"I have a new patch... Testing it now.\nIt is mostly what you have Jieshan, but it does not need all the changes to the ZookeeperNodeTracker and subclasses.\nI'll attach it soon... Then please let me know what you think.","from":"developer"},{"body":"There was also some weird stuff in HBaseAdmin.checkHBaseAvailable. It did set the client retry to 1, but left the the retry count on in RecoverableZookeeper, which leads to long, unnecessary waits.","from":"developer"},{"body":"Please let me know what you think of 5153-trunk-v2.txt... Thanks","from":"developer"},{"body":"@Jieshan: The endless loop happens when ZK is actually down.","from":"developer"},{"body":"bq. As discussed with Ted. Trunk and 92 already including a retry logic in RecoverableZooKeeper. So that makes the retry logic in resetZooKeeperTrackersWithRetries less important.\n\nAre you guys saying we do not need this in 0.92+?","from":"developer"},{"body":"Thanks Lars. It's nice of you:)","from":"developer"},{"body":"\"The endless loop happens when ZK is actually down.\"\nIf ZK is actually down, the below code will throw a Exception:\n this.zooKeeper = getZooKeeperWatcher();\nThen catched by the below code:\n{noformat}\n try {\n LOG.info(\"This client just lost it's session with ZooKeeper, trying\" +\n \" to reconnect.\");\n resetZooKeeperTrackersWithRetries();\n LOG.info(\"Reconnected successfully. This disconnect could have been\" +\n \" caused by a network partition or a long-running GC pause,\" +\n \" either way it's recommended that you verify your environment.\");\n return;\n } catch (ZooKeeperConnectionException e) {\n LOG.error(\"Could not reconnect to ZooKeeper after session\" +\n \" expiration, aborting\");\n t = e;\n }\n if (t != null) LOG.fatal(msg, t);\n else LOG.fatal(msg);\n HConnectionManager.deleteStaleConnection(this);\n{noformat}\n\nIt should not be a endless loop. Does that make sense?\n","from":"developer"},{"body":"bq. Are you guys saying we do not need this in 0.92+?\n\nWe need this, but the patch for 0.92+ should take notice of this:)","from":"developer"},{"body":"{code}\n+ if (isResettingZKTrackers) {\n+ return;\n+ }\n{code}\nI was thinking about something similar to the above.\n\nIn resetZooKeeperTrackersWithRetries(), shall we call zooKeeper.close() before resetting zooKeeper?\n{code}\n this.zooKeeper.close();\n this.zooKeeper = null;\n{code}\n\nThanks for the help.\n\nI think we need an addendum for 0.90","from":"developer"},{"body":"@Lars:\nAdding a \"isResettingZKTrackers\" sounds good to me. One doubt: Is it necessary to add the keyword of \"volatile\"?","from":"developer"},{"body":"@Jieshan: Hmm... I do see the endless loop in the debugger. It happens from HBaseAdmin.checkHBaseAvailable when HBase is actually down. We get into resetZooKeeperTrackersWithRetries, which calls setupZookeeperTrackers, which causes abort to be called, which calls resetZooKeeperTrackersWithRetries.\nI am not sure getZooKeeperWatcher() would throw if ZK is not available.\n\n@Ted: calling this.zooKeeper.close() (if it is not null) first seems prudent.\n\n@Jieshan: Good catch, yes it should be volatile.\n","from":"developer"},{"body":"@Lars:\nDuring the retry, if we get any exceptions, Zookeeper and the Trackers also need to close.\n{noformat}\n+ try {\n+ setupZookeeperTrackers();\n+ break;\n+ } catch (ZooKeeperConnectionException zkce) {\n+ if (tries >= this.numRetries) {\n+ throw zkce;\n+ }\n+ }\n{noformat}","from":"developer"},{"body":"@Lars:\nOne more doubt basing on the previous comment, regading on the below code:\n{noformat}\n+ try {\n+ setupZookeeperTrackers();\n+ break;\n+ } catch (ZooKeeperConnectionException zkce) {\n+ if (tries >= this.numRetries) {\n+ throw zkce;\n+ }\n+ }\n{noformat}\nif ZookeeperNodeTracker#start get an exception, then catches and calls abort(Doesn't throw it). The above retries will be break under this situation. So this retries takes no effects.Correct me if am wrong. ","from":"developer"},{"body":"This is the addendum for 90.","from":"developer"},{"body":"+1 on addendum for 0.90","from":"developer"},{"body":"@Ted\nThanks for the review.\n@Lars\nCan you review the 0.90 patch and provide your comments.\nIf it is ok i can integrate to 0.90 and take another RC. :)\n","from":"developer"},{"body":"+1 on addendum. Technically we should change ZooKeeperNodeTracker.start back to be void, but that is not necessary.\nI'll update the trunk patch soon.","from":"developer"},{"body":"@Jieshan... I see your point. Need to think about that.","from":"developer"},{"body":"Actually then the addendum is not right either. If ZooKeeperNodeTracker.start gets an exception it'll call abort (and hence not return false)","from":"developer"},{"body":"I take this back... The 0.90 addendum is good. If abort is called from the reset method it'll return right away, and in ZooKeeperNodeTracker.start fall through to return false.\n\nThis entire thing is too complicated. Let's hold off on the 0.92 and trunk patches, but leave it in 0.90 (although I am not a big fan of diverging code bases either).","from":"developer"},{"body":"Dropping Fix versions as Lars suggested.","from":"developer"},{"body":"Here's a minimal patch for trunk.\n\nIf the HConnection is not usable for any reason, the using application can attempt to reset it.\nI kept CloseConnectionException and the idea that ZooKeeperWatcher can unregister a listener.\n\nIf ZK is really not available that is probably an event that the application has to deal with anyway.\n\nLong term HConnection should *not* hold persistent connections to ZK anyway, but only establish a connection when needed.","from":"developer"},{"body":"Noticed one flaw in the 0.90 addendum patch. Should be\n{code}\n} catch (ZooKeeperConnectionException zkce) {\n if (tries+1 >= this.numRetries) {\n throw zkce;\n }\n}\n{code}\n(note the +1 in there)","from":"developer"},{"body":"For resetConnection() in the new patch, I see:\n{code}\n+ this.closed = false;\n{code}\nIn HConnectionImplementation.close(), we have:\n{code}\n if (master != null) {\n if (stopProxy) {\n HBaseRPC.stopProxy(master);\n }\n master = null;\n{code}\nThis seems to imply that once a connection is closed, we cannot simply change this.closed to false in resetConnection().","from":"developer"},{"body":"That is just guard against closed = true in HConnectionImplementation.abort().\nabort probably not set closed to true, but then a bunch of other places would need to be changed too.","from":"developer"},{"body":"@Lars\nAddressing your comments.\nBased on your feedback will commit the patch if it is ok.","from":"developer"},{"body":"@Ram: Sure. I now feel the whole logic is too complicated, but we'll not fix that any time soon. So +1 on addendum.","from":"developer"},{"body":"@Lars\nI understand your point. But as this JIRA is now committed in 0.90 so this addendum is needed.\nI think we can ensure such things doesn't happen for other JIRAs.\nGood on you Lars.","from":"developer"},{"body":"This got committed to 0.90.6.\nSorry for not resolving it at that time.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2012-01-09T08:52:15.000+0000","description":"HBASE-4893 is related to this issue. In that issue, we know, if multi-threads share a same connection, once this connection got abort in one thread, the other threads will got a \"HConnectionManager$HConnectionImplementation@18fb1f7 closed\" exception.\n\nIt solve the problem of \"stale connection can't removed\". But the orignal HTable instance cann't be continue to use. The connection in HTable should be recreated.\n\nActually, there's two aproach to solve this:\n1. In user code, once catch an IOE, close connection and re-create HTable instance. We can use this as a workaround.\n2. In HBase Client side, catch this exception, and re-create connection.\n","issue_id":"12537717","key":"HBASE-5153","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-03-19T04:24:55.000+0000","role":"fixed_distractor","summary":"Add retry logic in HConnectionImplementation#resetZooKeeperTrackers"} {"case_id":"12549929","cluster":"DISTRACTOR-HBASE-5741","comments":[{"body":"Yes, the javadoc says it half correct when it says that \"the table must exist in HBase if you are not using the option importtsv.bulk.output\".\nThe behavior, in absence of the above property is to do inserts into the provided table; so it throws an exception if the table doesn't exist.\n\nWhen we use this option, importtsv tries to configure job by making an attempt to read all the regions of this table and then use the start keys of these regions as partition-markers for the TotalOrderPartitioner class. This usage eliminates the first start-row (which is always a EMPTY_BYTE_ARRAY), so even if we create a new table, it will be useless as for using it for configuring a job. \n\nFor the case with LoadIncrementalHFile, the destination of the directory containing the hFiles is assumed to be in a specific format:\n//\n //\n\n --where we are storing the hfiles for a specific column family in a separate sub-directory.\n\nIf we do create a table based on the parameters given to the importtsv command, it will not be useful for the case of bulkload usecase as in the importtsv job, we dump hfiles based on rows; so all coulmn famlies for a specific row lands in one hfile. \n\nIt would be great to know if we actually use this workflow: create HFiles from importtsv job, and then use bulkload to insert those HFiles in HBase table.\n\nI think we should change the javadoc.\n\nPlease let me know if you have any questions; hbase-mapreduce use cases are exciting.","created":"2012-04-08T00:59:15.031+0000"},{"body":"btw, by \"making an attempt to read all the regions of this table and then use the start keys of these regions\", I mean get the start rows for all the existing regions of the table and then use them for job configuration.","created":"2012-04-08T01:01:59.427+0000"},{"body":"Thanks for the feedback. We I have customers who are using the workflow you mentioned. This is why I discovered the confusing javadoc reference.","created":"2012-04-09T16:59:03.340+0000"},{"body":"You are right. I am on it now. Thanks.","created":"2012-04-10T16:56:22.303+0000"},{"body":"A patch for checking/adding non-existant table in hbase in case we are using bulkoutput option. Added a test in the TestImportTsv. Please comment.","created":"2012-04-11T04:11:47.860+0000"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/4700/\n-----------------------------------------------------------\n\nReview request for hbase.\n\n\nSummary\n-------\n\nThere is a bulk output option in the importtsv workload. It outputs the HFiles in a user defined directory. The current code assumes that a table with its name equal to the given output directory exists, and throws an exception otherwise. Here is a patch for creating a table in case it doesn't exist.\n\n\nThis addresses bug HBase-5741.\n https://issues.apache.org/jira/browse/HBase-5741\n\n\nDiffs\n-----\n\n src/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java ab22fc4 \n src/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java ac30a62 \n\nDiff: https://reviews.apache.org/r/4700/diff\n\n\nTesting\n-------\n\nAdded a new test for bulkoutput; All importtsv tests pass.\n\n\nThanks,\n\nHimanshu\n\n","created":"2012-04-11T17:17:16.101+0000"},{"body":"Uploaded the patch on rb; \nhttps://reviews.apache.org/r/4700/","created":"2012-04-11T17:18:29.954+0000"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/4700/#review6857\n-----------------------------------------------------------\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n I think the new method should be called doesTableExist() or tableExists()\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n Insert space between for and (.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n Insert space between if and (.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n Insert one leading space before cfSet\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n Can this method be package private ?\n Since HBaseAdmin is created inside, a better name would be createHBaseAdmin().\n\n\n\nsrc/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java\n\n\n Insert space between if and (.\n\n\n\nsrc/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java\n\n\n Insert space between else and {.\n\n\n- Ted\n\n\nOn 2012-04-11 17:16:59, Himanshu Vashishtha wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/4700/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2012-04-11 17:16:59)\nbq. \nbq. \nbq. Review request for hbase.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. There is a bulk output option in the importtsv workload. It outputs the HFiles in a user defined directory. The current code assumes that a table with its name equal to the given output directory exists, and throws an exception otherwise. Here is a patch for creating a table in case it doesn't exist.\nbq. \nbq. \nbq. This addresses bug HBase-5741.\nbq. https://issues.apache.org/jira/browse/HBase-5741\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java ab22fc4 \nbq. src/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java ac30a62 \nbq. \nbq. Diff: https://reviews.apache.org/r/4700/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. Added a new test for bulkoutput; All importtsv tests pass.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Himanshu\nbq. \nbq.\n\n","created":"2012-04-11T17:45:16.721+0000"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/4700/\n-----------------------------------------------------------\n\n(Updated 2012-04-11 19:01:39.660787)\n\n\nReview request for hbase.\n\n\nChanges\n-------\n\nChanges done as per Ted.\n\n\nSummary\n-------\n\nThere is a bulk output option in the importtsv workload. It outputs the HFiles in a user defined directory. The current code assumes that a table with its name equal to the given output directory exists, and throws an exception otherwise. Here is a patch for creating a table in case it doesn't exist.\n\n\nThis addresses bug HBase-5741.\n https://issues.apache.org/jira/browse/HBase-5741\n\n\nDiffs (updated)\n-----\n\n src/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java ab22fc4 \n src/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java ac30a62 \n\nDiff: https://reviews.apache.org/r/4700/diff\n\n\nTesting\n-------\n\nAdded a new test for bulkoutput; All importtsv tests pass.\n\n\nThanks,\n\nHimanshu\n\n","created":"2012-04-11T19:03:14.124+0000"},{"body":"Patch version 2 after rb.","created":"2012-04-11T19:25:19.303+0000"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/4700/#review6860\n-----------------------------------------------------------\n\nShip it!\n\n\nPlease attach new patch to JIRA after addressing minor comments.\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n The parentheses around hbaseAdmin.tableExists() are not needed.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n Move this to line 263.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n 'bothered about' -> 'concerned with'\n\n\n- Ted\n\n\nOn 2012-04-11 19:01:39, Himanshu Vashishtha wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/4700/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2012-04-11 19:01:39)\nbq. \nbq. \nbq. Review request for hbase.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. There is a bulk output option in the importtsv workload. It outputs the HFiles in a user defined directory. The current code assumes that a table with its name equal to the given output directory exists, and throws an exception otherwise. Here is a patch for creating a table in case it doesn't exist.\nbq. \nbq. \nbq. This addresses bug HBase-5741.\nbq. https://issues.apache.org/jira/browse/HBase-5741\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java ab22fc4 \nbq. src/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java ac30a62 \nbq. \nbq. Diff: https://reviews.apache.org/r/4700/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. Added a new test for bulkoutput; All importtsv tests pass.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Himanshu\nbq. \nbq.\n\n","created":"2012-04-11T20:27:13.764+0000"},{"body":"Patch v3 addresses latest review comments.","created":"2012-04-11T22:23:09.862+0000"},{"body":"Patch for 0.94 where TestImportTsv passed.","created":"2012-04-11T22:53:28.141+0000"},{"body":"I would integrate patches to trunk and 0.94 tomorrow morning if there is no objection.","created":"2012-04-11T23:22:16.722+0000"},{"body":"Sorry for chiming in late here, but do we actually want ImportTsv to create the table for us?\nPersonally I'd rather have it fail, which forces me to think about how I want create the table (pre split, compression, etc).\n","created":"2012-04-11T23:55:38.450+0000"},{"body":"How about adding an option to ImportTsv, default to false, allowing user to auto-create the table ?","created":"2012-04-12T00:03:05.732+0000"},{"body":"Import does not do that either. At least we should be consistent between the various importing tools.\nIt seems overkill to burden these tools with that.\n\nThat said, I am not opposed, just questioning the usefulness.","created":"2012-04-12T00:06:09.204+0000"},{"body":"As Clint mentioned, there are some workloads that use ImportTsv to create HFiles, which later on are used in the bulk import method : my motivation for the patch.\n\nWhat you suggest Lars, failing the importtsv, or adding a similar thing in the Import tool too? I think later makes it consistent.","created":"2012-04-12T00:27:08.172+0000"},{"body":"Even HFileOutputFormat uses the table in HBase to find out about splits. We will rarely bulk import into an empty table that is not pre-split...? (I'm assuming here, you folks at Cloudera see many more use cases that I do)\n\nIf you believe strongly that we should add this, let's do it :)\nWe can do Import in a separate patch, or not do that as we see fit.\n","created":"2012-04-12T00:33:30.017+0000"},{"body":"Moving to 0.94.1, since we're still discussing (again, sorry about the late chime in)","created":"2012-04-12T16:12:01.244+0000"},{"body":"I am waiting for Clint to give me some numbers of such use cases to make the cut; (Apparently, he is on vacation these days and will be back this Monday)\nThanks.","created":"2012-04-13T18:41:33.573+0000"},{"body":"Lars, et al, thank you for your questions and comments. I may be missing something, so please correct me if I'm wrong, but here's how I see the situation:\n\n1) importtsv will auto-split the input data in region-server-sized Hfiles, based on hbase.hregion.max.filesize (because we use the HFileOutputFormat and it checks this parameter from the client configs).\n\n2) completebulkload then distributes the resulting HFiles to the region servers, thereby populating your cluster with all the regions needed to host your data. Alleviating all the memstore flushes, compactions, splitting, and other overhead associated with running billions of individual put()s to incrementally load your data. This is why we recommend the bulk load process to initially populate HBase tables.\n\n3) pre-splitting your table (based on rowkeys, as is the only mechanism) is useless in this scenario, because it will not take into account the size of data associated with each key. The customer would literally have to walk through their data and determine how many Gigs of data are associated with each rowkey in order to preemptively split their tables in a bulk load scenario...and even then, HBase only takes those pre-splits as \"suggestions\" and if the data ends up needing to be split differently, that will happen automatically. So what purpose does pre-splitting serve in a bulk load? Think \"initial load\" of massive amounts of data.\n\n4) Lars had a good point about compression, yes, typically a customer *should* set up their table with compression first, but they typically don't know this and they can do it after the fact without consequence.\n\nIn conclusion, we provide this handy command-line tool to help our HBase customers (most of whom are complete newbies and don't know what they're doing) get their HBase tables up and running, yet we give them conflicting information about how to use the tool. If we need to tell them to pre-create the table, then we should tell them so...clearly. However, I think it is just simpler and more intuitive if the tool did that automatically, as the javadocs indicate it will. Since region splits are determined internally by the HFileOutputFormat anyway (if they use the bulkload option), what's the harm? This enables people to get their data loaded into HBase with the least amount of education, ramp-up, java development, data mining, etc., etc.","created":"2012-04-17T22:15:24.587+0000"},{"body":"Thanks Clint. No harm, I just think it will hide complexity that should not be hidden. I also don't think it's too much to ask a user to create the table first (but we should definitely fix the Javadoc in that case).\n\nThat all said, I am +-0 on this. It's a simple change, and v3 looks good.","created":"2012-04-17T23:59:43.531+0000"},{"body":"Integrated to trunk.\n\nThanks for the patch Himanshu.\n\nWill wait for 0.94.0 to come out before integrating to 0.94","created":"2012-04-18T00:33:26.087+0000"},{"body":"Integrated to 0.94 as well.","created":"2012-04-19T16:22:42.321+0000"},{"body":"This was committed a while back... Marking it accordingly.","created":"2012-05-17T16:44:57.107+0000"}],"conversations":[{"body":"The usage statement for the \"importtsv\" command to hbase claims this:\n\n\"Note: if you do not use this option, then the target table must already exist in HBase\" (in reference to the \"importtsv.bulk.output\" command-line option)\n\nThe truth is, the table must exist no matter what, importtsv cannot and will not create it for you.\n\nThis is the case because the createSubmittableJob method of ImportTsv does not even attempt to check if the table exists already, much less create it:\n\n(From org.apache.hadoop.hbase.mapreduce.ImportTsv.java)\n\n305 HTable table = new HTable(conf, tableName);\n\nThe HTable method signature in use there assumes the table exists and runs a meta scan on it:\n\n(From org.apache.hadoop.hbase.client.HTable.java)\n\n142 * Creates an object to access a HBase table.\n...\n151 public HTable(Configuration conf, final String tableName)\n\nWhat we should do inside of createSubmittableJob is something similar to what the \"completebulkloads\" command would do:\n\n(Taken from org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles.java)\n\n690 boolean tableExists = this.doesTableExist(tableName);\n691 if (!tableExists) this.createTable(tableName,dirPath);\n\nCurrently the docs are misleading, the table in fact must exist prior to running importtsv. We should check if it exists rather than assume it's already there and throw the below exception:\n\n12/03/14 17:15:42 WARN client.HConnectionManager$HConnectionImplementation: Encountered problems when prefetch META table: \norg.apache.hadoop.hbase.TableNotFoundException: Cannot find row in .META. for table: myTable2, row=myTable2,,99999999999999\n\tat org.apache.hadoop.hbase.client.MetaScanner.metaScan(MetaScanner.java:150)\n...\n","from":"reporter","subject":"ImportTsv does not check for table existence "},{"body":"Yes, the javadoc says it half correct when it says that \"the table must exist in HBase if you are not using the option importtsv.bulk.output\".\nThe behavior, in absence of the above property is to do inserts into the provided table; so it throws an exception if the table doesn't exist.\n\nWhen we use this option, importtsv tries to configure job by making an attempt to read all the regions of this table and then use the start keys of these regions as partition-markers for the TotalOrderPartitioner class. This usage eliminates the first start-row (which is always a EMPTY_BYTE_ARRAY), so even if we create a new table, it will be useless as for using it for configuring a job. \n\nFor the case with LoadIncrementalHFile, the destination of the directory containing the hFiles is assumed to be in a specific format:\n//\n //\n\n --where we are storing the hfiles for a specific column family in a separate sub-directory.\n\nIf we do create a table based on the parameters given to the importtsv command, it will not be useful for the case of bulkload usecase as in the importtsv job, we dump hfiles based on rows; so all coulmn famlies for a specific row lands in one hfile. \n\nIt would be great to know if we actually use this workflow: create HFiles from importtsv job, and then use bulkload to insert those HFiles in HBase table.\n\nI think we should change the javadoc.\n\nPlease let me know if you have any questions; hbase-mapreduce use cases are exciting.","from":"developer"},{"body":"btw, by \"making an attempt to read all the regions of this table and then use the start keys of these regions\", I mean get the start rows for all the existing regions of the table and then use them for job configuration.","from":"developer"},{"body":"Thanks for the feedback. We I have customers who are using the workflow you mentioned. This is why I discovered the confusing javadoc reference.","from":"developer"},{"body":"You are right. I am on it now. Thanks.","from":"developer"},{"body":"A patch for checking/adding non-existant table in hbase in case we are using bulkoutput option. Added a test in the TestImportTsv. Please comment.","from":"developer"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/4700/\n-----------------------------------------------------------\n\nReview request for hbase.\n\n\nSummary\n-------\n\nThere is a bulk output option in the importtsv workload. It outputs the HFiles in a user defined directory. The current code assumes that a table with its name equal to the given output directory exists, and throws an exception otherwise. Here is a patch for creating a table in case it doesn't exist.\n\n\nThis addresses bug HBase-5741.\n https://issues.apache.org/jira/browse/HBase-5741\n\n\nDiffs\n-----\n\n src/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java ab22fc4 \n src/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java ac30a62 \n\nDiff: https://reviews.apache.org/r/4700/diff\n\n\nTesting\n-------\n\nAdded a new test for bulkoutput; All importtsv tests pass.\n\n\nThanks,\n\nHimanshu\n\n","from":"developer"},{"body":"Uploaded the patch on rb; \nhttps://reviews.apache.org/r/4700/","from":"developer"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/4700/#review6857\n-----------------------------------------------------------\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n I think the new method should be called doesTableExist() or tableExists()\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n Insert space between for and (.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n Insert space between if and (.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n Insert one leading space before cfSet\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n Can this method be package private ?\n Since HBaseAdmin is created inside, a better name would be createHBaseAdmin().\n\n\n\nsrc/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java\n\n\n Insert space between if and (.\n\n\n\nsrc/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java\n\n\n Insert space between else and {.\n\n\n- Ted\n\n\nOn 2012-04-11 17:16:59, Himanshu Vashishtha wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/4700/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2012-04-11 17:16:59)\nbq. \nbq. \nbq. Review request for hbase.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. There is a bulk output option in the importtsv workload. It outputs the HFiles in a user defined directory. The current code assumes that a table with its name equal to the given output directory exists, and throws an exception otherwise. Here is a patch for creating a table in case it doesn't exist.\nbq. \nbq. \nbq. This addresses bug HBase-5741.\nbq. https://issues.apache.org/jira/browse/HBase-5741\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java ab22fc4 \nbq. src/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java ac30a62 \nbq. \nbq. Diff: https://reviews.apache.org/r/4700/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. Added a new test for bulkoutput; All importtsv tests pass.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Himanshu\nbq. \nbq.\n\n","from":"developer"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/4700/\n-----------------------------------------------------------\n\n(Updated 2012-04-11 19:01:39.660787)\n\n\nReview request for hbase.\n\n\nChanges\n-------\n\nChanges done as per Ted.\n\n\nSummary\n-------\n\nThere is a bulk output option in the importtsv workload. It outputs the HFiles in a user defined directory. The current code assumes that a table with its name equal to the given output directory exists, and throws an exception otherwise. Here is a patch for creating a table in case it doesn't exist.\n\n\nThis addresses bug HBase-5741.\n https://issues.apache.org/jira/browse/HBase-5741\n\n\nDiffs (updated)\n-----\n\n src/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java ab22fc4 \n src/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java ac30a62 \n\nDiff: https://reviews.apache.org/r/4700/diff\n\n\nTesting\n-------\n\nAdded a new test for bulkoutput; All importtsv tests pass.\n\n\nThanks,\n\nHimanshu\n\n","from":"developer"},{"body":"Patch version 2 after rb.","from":"developer"},{"body":"\n-----------------------------------------------------------\nThis is an automatically generated e-mail. To reply, visit:\nhttps://reviews.apache.org/r/4700/#review6860\n-----------------------------------------------------------\n\nShip it!\n\n\nPlease attach new patch to JIRA after addressing minor comments.\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n The parentheses around hbaseAdmin.tableExists() are not needed.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n Move this to line 263.\n\n\n\nsrc/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java\n\n\n 'bothered about' -> 'concerned with'\n\n\n- Ted\n\n\nOn 2012-04-11 19:01:39, Himanshu Vashishtha wrote:\nbq. \nbq. -----------------------------------------------------------\nbq. This is an automatically generated e-mail. To reply, visit:\nbq. https://reviews.apache.org/r/4700/\nbq. -----------------------------------------------------------\nbq. \nbq. (Updated 2012-04-11 19:01:39)\nbq. \nbq. \nbq. Review request for hbase.\nbq. \nbq. \nbq. Summary\nbq. -------\nbq. \nbq. There is a bulk output option in the importtsv workload. It outputs the HFiles in a user defined directory. The current code assumes that a table with its name equal to the given output directory exists, and throws an exception otherwise. Here is a patch for creating a table in case it doesn't exist.\nbq. \nbq. \nbq. This addresses bug HBase-5741.\nbq. https://issues.apache.org/jira/browse/HBase-5741\nbq. \nbq. \nbq. Diffs\nbq. -----\nbq. \nbq. src/main/java/org/apache/hadoop/hbase/mapreduce/ImportTsv.java ab22fc4 \nbq. src/test/java/org/apache/hadoop/hbase/mapreduce/TestImportTsv.java ac30a62 \nbq. \nbq. Diff: https://reviews.apache.org/r/4700/diff\nbq. \nbq. \nbq. Testing\nbq. -------\nbq. \nbq. Added a new test for bulkoutput; All importtsv tests pass.\nbq. \nbq. \nbq. Thanks,\nbq. \nbq. Himanshu\nbq. \nbq.\n\n","from":"developer"},{"body":"Patch v3 addresses latest review comments.","from":"developer"},{"body":"Patch for 0.94 where TestImportTsv passed.","from":"developer"},{"body":"I would integrate patches to trunk and 0.94 tomorrow morning if there is no objection.","from":"developer"},{"body":"Sorry for chiming in late here, but do we actually want ImportTsv to create the table for us?\nPersonally I'd rather have it fail, which forces me to think about how I want create the table (pre split, compression, etc).\n","from":"developer"},{"body":"How about adding an option to ImportTsv, default to false, allowing user to auto-create the table ?","from":"developer"},{"body":"Import does not do that either. At least we should be consistent between the various importing tools.\nIt seems overkill to burden these tools with that.\n\nThat said, I am not opposed, just questioning the usefulness.","from":"developer"},{"body":"As Clint mentioned, there are some workloads that use ImportTsv to create HFiles, which later on are used in the bulk import method : my motivation for the patch.\n\nWhat you suggest Lars, failing the importtsv, or adding a similar thing in the Import tool too? I think later makes it consistent.","from":"developer"},{"body":"Even HFileOutputFormat uses the table in HBase to find out about splits. We will rarely bulk import into an empty table that is not pre-split...? (I'm assuming here, you folks at Cloudera see many more use cases that I do)\n\nIf you believe strongly that we should add this, let's do it :)\nWe can do Import in a separate patch, or not do that as we see fit.\n","from":"developer"},{"body":"Moving to 0.94.1, since we're still discussing (again, sorry about the late chime in)","from":"developer"},{"body":"I am waiting for Clint to give me some numbers of such use cases to make the cut; (Apparently, he is on vacation these days and will be back this Monday)\nThanks.","from":"developer"},{"body":"Lars, et al, thank you for your questions and comments. I may be missing something, so please correct me if I'm wrong, but here's how I see the situation:\n\n1) importtsv will auto-split the input data in region-server-sized Hfiles, based on hbase.hregion.max.filesize (because we use the HFileOutputFormat and it checks this parameter from the client configs).\n\n2) completebulkload then distributes the resulting HFiles to the region servers, thereby populating your cluster with all the regions needed to host your data. Alleviating all the memstore flushes, compactions, splitting, and other overhead associated with running billions of individual put()s to incrementally load your data. This is why we recommend the bulk load process to initially populate HBase tables.\n\n3) pre-splitting your table (based on rowkeys, as is the only mechanism) is useless in this scenario, because it will not take into account the size of data associated with each key. The customer would literally have to walk through their data and determine how many Gigs of data are associated with each rowkey in order to preemptively split their tables in a bulk load scenario...and even then, HBase only takes those pre-splits as \"suggestions\" and if the data ends up needing to be split differently, that will happen automatically. So what purpose does pre-splitting serve in a bulk load? Think \"initial load\" of massive amounts of data.\n\n4) Lars had a good point about compression, yes, typically a customer *should* set up their table with compression first, but they typically don't know this and they can do it after the fact without consequence.\n\nIn conclusion, we provide this handy command-line tool to help our HBase customers (most of whom are complete newbies and don't know what they're doing) get their HBase tables up and running, yet we give them conflicting information about how to use the tool. If we need to tell them to pre-create the table, then we should tell them so...clearly. However, I think it is just simpler and more intuitive if the tool did that automatically, as the javadocs indicate it will. Since region splits are determined internally by the HFileOutputFormat anyway (if they use the bulkload option), what's the harm? This enables people to get their data loaded into HBase with the least amount of education, ramp-up, java development, data mining, etc., etc.","from":"developer"},{"body":"Thanks Clint. No harm, I just think it will hide complexity that should not be hidden. I also don't think it's too much to ask a user to create the table first (but we should definitely fix the Javadoc in that case).\n\nThat all said, I am +-0 on this. It's a simple change, and v3 looks good.","from":"developer"},{"body":"Integrated to trunk.\n\nThanks for the patch Himanshu.\n\nWill wait for 0.94.0 to come out before integrating to 0.94","from":"developer"},{"body":"Integrated to 0.94 as well.","from":"developer"},{"body":"This was committed a while back... Marking it accordingly.","from":"developer"}],"created":"2012-04-06T20:11:46.000+0000","description":"The usage statement for the \"importtsv\" command to hbase claims this:\n\n\"Note: if you do not use this option, then the target table must already exist in HBase\" (in reference to the \"importtsv.bulk.output\" command-line option)\n\nThe truth is, the table must exist no matter what, importtsv cannot and will not create it for you.\n\nThis is the case because the createSubmittableJob method of ImportTsv does not even attempt to check if the table exists already, much less create it:\n\n(From org.apache.hadoop.hbase.mapreduce.ImportTsv.java)\n\n305 HTable table = new HTable(conf, tableName);\n\nThe HTable method signature in use there assumes the table exists and runs a meta scan on it:\n\n(From org.apache.hadoop.hbase.client.HTable.java)\n\n142 * Creates an object to access a HBase table.\n...\n151 public HTable(Configuration conf, final String tableName)\n\nWhat we should do inside of createSubmittableJob is something similar to what the \"completebulkloads\" command would do:\n\n(Taken from org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles.java)\n\n690 boolean tableExists = this.doesTableExist(tableName);\n691 if (!tableExists) this.createTable(tableName,dirPath);\n\nCurrently the docs are misleading, the table in fact must exist prior to running importtsv. We should check if it exists rather than assume it's already there and throw the below exception:\n\n12/03/14 17:15:42 WARN client.HConnectionManager$HConnectionImplementation: Encountered problems when prefetch META table: \norg.apache.hadoop.hbase.TableNotFoundException: Cannot find row in .META. for table: myTable2, row=myTable2,,99999999999999\n\tat org.apache.hadoop.hbase.client.MetaScanner.metaScan(MetaScanner.java:150)\n...\n","issue_id":"12549929","key":"HBASE-5741","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-05-17T16:44:56.000+0000","role":"fixed_distractor","summary":"ImportTsv does not check for table existence "} {"case_id":"12554413","cluster":"DISTRACTOR-HBASE-5966","comments":[{"body":"-Dhadoop.profile=24 gives the same result, the profiles seem to be identical (?).","created":"2012-05-09T01:07:08.865+0000"},{"body":"Andy, I assume you are seeing more than one test fail with similar errors?","created":"2012-05-09T04:02:30.237+0000"},{"body":"Anything that uses the minicluster fails. I only checked the output of a few but all had the same type and order of exceptions. \n","created":"2012-05-09T04:09:04.134+0000"},{"body":"Btw, for my education, -Dhadoop.profile just flips some code in hbase and doesn't mean we are getting different versions of hadoop-mapreduce, correct?","created":"2012-05-09T04:09:20.493+0000"},{"body":"It's a maven activation that chooses the 0.23+ dependency set. You'll see it in the root POM. \n","created":"2012-05-09T04:16:53.999+0000"},{"body":"I am also facing the same issue. Testcases using Map-Red minicluster are failing with the same kind of exception when run with 0.23 version","created":"2012-05-09T04:41:57.269+0000"},{"body":"This is the exception that i got while running testcases. Some other testcases are also failing with the same error.\n{code:xml} \n-------------------------------------------------------------------------------\nTest set: org.apache.hadoop.hbase.mapreduce.TestTableInputFormatScan\n-------------------------------------------------------------------------------\nTests run: 1, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 177.132 sec <<< FAILURE!\norg.apache.hadoop.hbase.mapreduce.TestTableInputFormatScan Time elapsed: 0 sec <<< ERROR!\norg.apache.hadoop.yarn.YarnException: Failed to Start org.apache.hadoop.mapred.MiniMRCluster\n\tat org.apache.hadoop.yarn.service.CompositeService.start(CompositeService.java:78)\n\tat org.apache.hadoop.mapred.MiniMRClientClusterFactory.create(MiniMRClientClusterFactory.java:67)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:180)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:170)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:162)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:154)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:147)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:140)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:133)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:128)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniMapReduceCluster(HBaseTestingUtility.java:1269)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniMapReduceCluster(HBaseTestingUtility.java:1256)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableInputFormatScan.setUpBeforeClass(TestTableInputFormatScan.java:83)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:45)\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15)\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:42)\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:27)\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:30)\n\tat org.junit.runners.ParentRunner.run(ParentRunner.java:300)\n\tat org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n\tat org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n\tat org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n\tat org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n\tat org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\nCaused by: org.apache.hadoop.yarn.YarnException: java.io.IOException: ResourceManager failed to start. Final state is INITED\n\tat org.apache.hadoop.yarn.server.MiniYARNCluster$ResourceManagerWrapper.start(MiniYARNCluster.java:152)\n\tat org.apache.hadoop.yarn.service.CompositeService.start(CompositeService.java:68)\n\t... 34 more\nCaused by: java.io.IOException: ResourceManager failed to start. Final state is INITED\n\tat org.apache.hadoop.yarn.server.MiniYARNCluster$ResourceManagerWrapper.start(MiniYARNCluster.java:146)\n\t... 35 more\n{code}","created":"2012-05-09T05:11:17.494+0000"},{"body":"I think this is coming after the changes of MAPREDUCE-3867.\n\nI have attached the patch to fix this.\n\n@Ashutosh: I don't think your failure is related to this because it impacts only the client to RM communication. As per your exception stack, RM is failing to start. Can you check the detailed log and environment for more details?\n","created":"2012-05-09T09:17:57.991+0000"},{"body":"I have updated the patch to fix the above failures.","created":"2012-05-09T10:22:24.036+0000"},{"body":"Testing with this patch ported to 0.94, via\n\n{noformat}\nmvn -PlocalTests -Psecurity \\\n -Dhadoop.profile=23 -Dhadoop.version=2.0.0-SNAPSHOT \\\n clean test \\\n -Dtest=org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n{noformat}\n\nI see this failure:\n\n{noformat}\nTests run: 1, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 45.473 sec <<< FAILURE!\ntestMultiRegionTable(org.apache.hadoop.hbase.mapreduce.TestTableMapReduce) Time elapsed: 11.876 sec <<< ERROR!\njava.lang.NullPointerException\n\tat org.apache.hadoop.net.DNS.reverseDns(DNS.java:93)\n\tat org.apache.hadoop.hbase.mapreduce.TableInputFormatBase.reverseDNS(TableInputFormatBase.java:200)\n\tat org.apache.hadoop.hbase.mapreduce.TableInputFormatBase.getSplits(TableInputFormatBase.java:165)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.writeNewSplits(JobSubmitter.java:451)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.writeSplits(JobSubmitter.java:468)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.submitJobInternal(JobSubmitter.java:360)\n\tat org.apache.hadoop.mapreduce.Job$11.run(Job.java:1226)\n\tat org.apache.hadoop.mapreduce.Job$11.run(Job.java:1223)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:416)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1232)\n\tat org.apache.hadoop.mapreduce.Job.submit(Job.java:1223)\n\tat org.apache.hadoop.mapreduce.Job.waitForCompletion(Job.java:1244)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.runTestOnTable(TestTableMapReduce.java:151)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.testMultiRegionTable(TestTableMapReduce.java:129)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:45)\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15)\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:42)\n\tat org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:20)\n\tat org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:47)\n\tat org.junit.rules.RunRules.evaluate(RunRules.java:18)\n\tat org.junit.runners.ParentRunner.runLeaf(ParentRunner.java:263)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:68)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:47)\n\tat org.junit.runners.ParentRunner$3.run(ParentRunner.java:231)\n\tat org.junit.runners.ParentRunner$1.schedule(ParentRunner.java:60)\n\tat org.junit.runners.ParentRunner.runChildren(ParentRunner.java:229)\n\tat org.junit.runners.ParentRunner.access$000(ParentRunner.java:50)\n\tat org.junit.runners.ParentRunner$2.evaluate(ParentRunner.java:222)\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:28)\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:30)\n\tat org.junit.runners.ParentRunner.run(ParentRunner.java:300)\n\tat org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n\tat org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n\tat org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n\tat org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n\tat org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\n{noformat}\n\nand probably there are more but this is the only one I've tried so far.\n\nIs it just me?","created":"2012-05-09T16:47:41.512+0000"},{"body":"I saw lots of failures too. I am testing with patch for 5963 and 5966 now.","created":"2012-05-09T16:50:57.230+0000"},{"body":"Above failure may have been due to locally applied patches for other ongoing issues with 2.0.0-alpha, so I ran it again with a clean 0.94 branch checkout and see this:\n\n{noformat}\n-------------------------------------------------------------------------------\nTest set: org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n-------------------------------------------------------------------------------\nTests run: 1, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 205.635 sec <<< FAILURE!\norg.apache.hadoop.hbase.mapreduce.TestTableMapReduce Time elapsed: 0 sec <<< ERROR!\njava.io.IOException: Shutting down\n\tat org.apache.hadoop.hbase.MiniHBaseCluster.init(MiniHBaseCluster.java:203)\n\tat org.apache.hadoop.hbase.MiniHBaseCluster.(MiniHBaseCluster.java:76)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniHBaseCluster(HBaseTestingUtility.java:632)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniCluster(HBaseTestingUtility.java:606)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniCluster(HBaseTestingUtility.java:554)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniCluster(HBaseTestingUtility.java:523)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.beforeClass(TestTableMapReduce.java:68)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:45)\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15)\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:42)\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:27)\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:30)\n\tat org.junit.runners.ParentRunner.run(ParentRunner.java:300)\n\tat org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n\tat org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n\tat org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n\tat org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n\tat org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\nCaused by: java.lang.RuntimeException: Master not initialized after 200 seconds\n\tat org.apache.hadoop.hbase.util.JVMClusterUtil.startup(JVMClusterUtil.java:206)\n\tat org.apache.hadoop.hbase.LocalHBaseCluster.startup(LocalHBaseCluster.java:422)\n\tat org.apache.hadoop.hbase.MiniHBaseCluster.init(MiniHBaseCluster.java:196)\n\t... 28 more\n{noformat}\n\nSame result on a clean checkout of TRUNK.","created":"2012-05-09T17:03:00.563+0000"},{"body":"Combined with patches from HBASE-5964 and HBASE-5963, I saw this:\n{code}\ntestMultiRegionTable(org.apache.hadoop.hbase.mapreduce.TestTableMapReduce) Time elapsed: 47.489 sec <<< FAILURE!\njava.lang.AssertionError\n at org.junit.Assert.fail(Assert.java:92)\n at org.junit.Assert.assertTrue(Assert.java:43)\n at org.junit.Assert.assertTrue(Assert.java:54)\n at org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.runTestOnTable(TestTableMapReduce.java:151)\n at org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.testMultiRegionTable(TestTableMapReduce.java:129)\n{code}\nIn test output, I found:\n{code}\n2012-05-09 10:05:11,141 ERROR [main] mapreduce.TableInputFormatBase(171): Cannot resolve the host name for /10.249.196.102 because of javax.naming.NameNotFoundException: DNS name not found [response code 3]; remaining name '102.196.249.10.in-addr.arpa'\n{code}","created":"2012-05-09T17:41:08.786+0000"},{"body":"I forgot to apply HBASE-5964 to the fresh checkout. After doing that, I see the same issue that Ted reports.","created":"2012-05-09T18:32:40.552+0000"},{"body":"The patch is based on the patch Andrew posted.\n\nTogether with the patch for HBASE-5975, TestTableMapReduce is green for me.","created":"2012-05-10T04:10:24.614+0000"},{"body":"These tests are green on my box.","created":"2012-05-10T22:22:34.206+0000"},{"body":"Can we push this patch to trunk? Based on our testing, it does fix the MR test failures with HADOOP 2.0.0.","created":"2012-05-11T17:09:58.887+0000"},{"body":"What am I doing wrong? I tried it three times here on my laptop:\n\n{code}\n-------------------------------------------------------------------------------\nTest set: org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n-------------------------------------------------------------------------------\nTests run: 1, Failures: 1, Errors: 0, Skipped: 0, Time elapsed: 64.363 sec <<< FAILURE!\ntestMultiRegionTable(org.apache.hadoop.hbase.mapreduce.TestTableMapReduce) Time elapsed: 30.455 sec <<< FAILURE!\njava.lang.AssertionError\n at org.junit.Assert.fail(Assert.java:92)\n at org.junit.Assert.assertTrue(Assert.java:43)\n at org.junit.Assert.assertTrue(Assert.java:54)\n at org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.runTestOnTable(TestTableMapReduce.java:151)\n at org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.testMultiRegionTable(TestTableMapReduce.java:129)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:45)\n at org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15)\n at org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:42)\n at org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:20)\n at org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:47)\n at org.junit.rules.RunRules.evaluate(RunRules.java:18)\n at org.junit.runners.ParentRunner.runLeaf(ParentRunner.java:263)\n at org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:68)\n at org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:47)\n at org.junit.runners.ParentRunner$3.run(ParentRunner.java:231)\n at org.junit.runners.ParentRunner$1.schedule(ParentRunner.java:60)\n at org.junit.runners.ParentRunner.runChildren(ParentRunner.java:229)\n at org.junit.runners.ParentRunner.access$000(ParentRunner.java:50)\n at org.junit.runners.ParentRunner$2.evaluate(ParentRunner.java:222)\n at org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:28)\n at org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:30)\n at org.junit.runners.ParentRunner.run(ParentRunner.java:300)\n at org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n at org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n at org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n at org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n at org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n at org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n at org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\n{code}\n\nI did this:\n\n{code}\n$ MAVEN_OPTS=\"-Xmx2g\" ~/bin/mvn/bin/mvn -PlocalTests -Psecurity -Dhadoop.profile=23 -Dhadoop.version=2.0.0-SNAPSHOT clean test -Dtest=org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n{code}","created":"2012-05-11T21:32:48.074+0000"},{"body":"Did you set JAVA_HOME?","created":"2012-05-12T01:11:08.422+0000"},{"body":"It passed when I set JAVA_HOME. Thanks for the patch Jimmy.","created":"2012-05-12T05:48:33.518+0000"},{"body":"Discussed with Jimmy. Let's have this in 0.94.1","created":"2012-07-19T21:58:55.789+0000"},{"body":"Attached patch for 0.94. Ran TestTableMapReduce against both 1.0 and 2.0 hadoop profiles, both passed:\n\n\nmvn test -PlocalTests -Dtest=org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n\n-------------------------------------------------------\n T E S T S\n-------------------------------------------------------\nRunning org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 188.087 sec\n\nResults :\n\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0\n\nmvn test -PlocalTests -Dhadoop.profile=2.0 -Dtest=org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n\n-------------------------------------------------------\n T E S T S\n-------------------------------------------------------\nRunning org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 167.49 sec\n\nResults :\n\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0\n","created":"2012-07-19T22:59:06.685+0000"},{"body":"looks good to me, will commit to 0.94 tonight if no objection.","created":"2012-07-19T23:11:22.104+0000"},{"body":"+1","created":"2012-07-19T23:31:31.239+0000"},{"body":"Integrated to 0.94. Thank Greg for the patch, Lars for the review.","created":"2012-07-19T23:44:05.663+0000"}],"conversations":[{"body":"Some fairly recent change in Hadoop 2.0.0-alpha has broken our MapReduce test rigging. Below is a representative error, can be easily reproduced with:\n\n{noformat}\nmvn -PlocalTests -Psecurity \\\n -Dhadoop.profile=23 -Dhadoop.version=2.0.0-SNAPSHOT \\\n clean test \\\n -Dtest=org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n{noformat}\n\nAnd the result:\n\n{noformat}\n-------------------------------------------------------\n T E S T S\n-------------------------------------------------------\nRunning org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\nTests run: 1, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 54.292 sec <<< FAILURE!\n\n-------------------------------------------------------------------------------\nTest set: org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n-------------------------------------------------------------------------------\nTests run: 1, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 54.292 sec <<< FAILURE!\ntestMultiRegionTable(org.apache.hadoop.hbase.mapreduce.TestTableMapReduce) Time elapsed: 21.935 sec <<< ERROR!\njava.lang.reflect.UndeclaredThrowableException\n\tat org.apache.hadoop.yarn.exceptions.impl.pb.YarnRemoteExceptionPBImpl.unwrapAndThrowException(YarnRemoteExceptionPBImpl.java:135)\n\tat org.apache.hadoop.yarn.api.impl.pb.client.ClientRMProtocolPBClientImpl.getNewApplication(ClientRMProtocolPBClientImpl.java:134)\n\tat org.apache.hadoop.mapred.ResourceMgrDelegate.getNewJobID(ResourceMgrDelegate.java:183)\n\tat org.apache.hadoop.mapred.YARNRunner.getNewJobID(YARNRunner.java:216)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.submitJobInternal(JobSubmitter.java:339)\n\tat org.apache.hadoop.mapreduce.Job$11.run(Job.java:1226)\n\tat org.apache.hadoop.mapreduce.Job$11.run(Job.java:1223)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:416)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1232)\n\tat org.apache.hadoop.mapreduce.Job.submit(Job.java:1223)\n\tat org.apache.hadoop.mapreduce.Job.waitForCompletion(Job.java:1244)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.runTestOnTable(TestTableMapReduce.java:151)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.testMultiRegionTable(TestTableMapReduce.java:129)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:45)\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15)\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:42)\n\tat org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:20)\n\tat org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:47)\n\tat org.junit.rules.RunRules.evaluate(RunRules.java:18)\n\tat org.junit.runners.ParentRunner.runLeaf(ParentRunner.java:263)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:68)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:47)\n\tat org.junit.runners.ParentRunner$3.run(ParentRunner.java:231)\n\tat org.junit.runners.ParentRunner$1.schedule(ParentRunner.java:60)\n\tat org.junit.runners.ParentRunner.runChildren(ParentRunner.java:229)\n\tat org.junit.runners.ParentRunner.access$000(ParentRunner.java:50)\n\tat org.junit.runners.ParentRunner$2.evaluate(ParentRunner.java:222)\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:28)\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:30)\n\tat org.junit.runners.ParentRunner.run(ParentRunner.java:300)\n\tat org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n\tat org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n\tat org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n\tat org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n\tat org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\nCaused by: com.google.protobuf.ServiceException: java.net.ConnectException: Call From acer.localdomain/192.168.122.1 to 0.0.0.0:8032 failed on connection exception: java.net.ConnectException: Connection refused; For more details see: http://wiki.apache.org/hadoop/ConnectionRefused\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:188)\n\tat $Proxy89.getNewApplication(Unknown Source)\n\tat org.apache.hadoop.yarn.api.impl.pb.client.ClientRMProtocolPBClientImpl.getNewApplication(ClientRMProtocolPBClientImpl.java:132)\n\t... 45 more\nCaused by: java.net.ConnectException: Call From acer.localdomain/192.168.122.1 to 0.0.0.0:8032 failed on connection exception: java.net.ConnectException: Connection refused; For more details see: http://wiki.apache.org/hadoop/ConnectionRefused\n\tat org.apache.hadoop.net.NetUtils.wrapException(NetUtils.java:725)\n\tat org.apache.hadoop.ipc.Client.call(Client.java:1160)\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:185)\n\t... 47 more\nCaused by: java.net.ConnectException: Connection refused\n\tat sun.nio.ch.SocketChannelImpl.checkConnect(Native Method)\n\tat sun.nio.ch.SocketChannelImpl.finishConnect(SocketChannelImpl.java:592)\n\tat org.apache.hadoop.net.SocketIOWithTimeout.connect(SocketIOWithTimeout.java:206)\n\tat org.apache.hadoop.net.NetUtils.connect(NetUtils.java:522)\n\tat org.apache.hadoop.net.NetUtils.connect(NetUtils.java:487)\n\tat org.apache.hadoop.ipc.Client$Connection.setupConnection(Client.java:469)\n\tat org.apache.hadoop.ipc.Client$Connection.setupIOstreams(Client.java:563)\n\tat org.apache.hadoop.ipc.Client$Connection.access$2000(Client.java:212)\n\tat org.apache.hadoop.ipc.Client.getConnection(Client.java:1266)\n\tat org.apache.hadoop.ipc.Client.call(Client.java:1136)\n\t... 48 more\n{noformat}\n","from":"reporter","subject":"MapReduce based tests broken on Hadoop 2.0.0-alpha"},{"body":"-Dhadoop.profile=24 gives the same result, the profiles seem to be identical (?).","from":"developer"},{"body":"Andy, I assume you are seeing more than one test fail with similar errors?","from":"developer"},{"body":"Anything that uses the minicluster fails. I only checked the output of a few but all had the same type and order of exceptions. \n","from":"developer"},{"body":"Btw, for my education, -Dhadoop.profile just flips some code in hbase and doesn't mean we are getting different versions of hadoop-mapreduce, correct?","from":"developer"},{"body":"It's a maven activation that chooses the 0.23+ dependency set. You'll see it in the root POM. \n","from":"developer"},{"body":"I am also facing the same issue. Testcases using Map-Red minicluster are failing with the same kind of exception when run with 0.23 version","from":"developer"},{"body":"This is the exception that i got while running testcases. Some other testcases are also failing with the same error.\n{code:xml} \n-------------------------------------------------------------------------------\nTest set: org.apache.hadoop.hbase.mapreduce.TestTableInputFormatScan\n-------------------------------------------------------------------------------\nTests run: 1, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 177.132 sec <<< FAILURE!\norg.apache.hadoop.hbase.mapreduce.TestTableInputFormatScan Time elapsed: 0 sec <<< ERROR!\norg.apache.hadoop.yarn.YarnException: Failed to Start org.apache.hadoop.mapred.MiniMRCluster\n\tat org.apache.hadoop.yarn.service.CompositeService.start(CompositeService.java:78)\n\tat org.apache.hadoop.mapred.MiniMRClientClusterFactory.create(MiniMRClientClusterFactory.java:67)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:180)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:170)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:162)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:154)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:147)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:140)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:133)\n\tat org.apache.hadoop.mapred.MiniMRCluster.(MiniMRCluster.java:128)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniMapReduceCluster(HBaseTestingUtility.java:1269)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniMapReduceCluster(HBaseTestingUtility.java:1256)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableInputFormatScan.setUpBeforeClass(TestTableInputFormatScan.java:83)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:45)\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15)\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:42)\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:27)\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:30)\n\tat org.junit.runners.ParentRunner.run(ParentRunner.java:300)\n\tat org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\tat java.lang.reflect.Method.invoke(Method.java:597)\n\tat org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n\tat org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n\tat org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n\tat org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n\tat org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\nCaused by: org.apache.hadoop.yarn.YarnException: java.io.IOException: ResourceManager failed to start. Final state is INITED\n\tat org.apache.hadoop.yarn.server.MiniYARNCluster$ResourceManagerWrapper.start(MiniYARNCluster.java:152)\n\tat org.apache.hadoop.yarn.service.CompositeService.start(CompositeService.java:68)\n\t... 34 more\nCaused by: java.io.IOException: ResourceManager failed to start. Final state is INITED\n\tat org.apache.hadoop.yarn.server.MiniYARNCluster$ResourceManagerWrapper.start(MiniYARNCluster.java:146)\n\t... 35 more\n{code}","from":"developer"},{"body":"I think this is coming after the changes of MAPREDUCE-3867.\n\nI have attached the patch to fix this.\n\n@Ashutosh: I don't think your failure is related to this because it impacts only the client to RM communication. As per your exception stack, RM is failing to start. Can you check the detailed log and environment for more details?\n","from":"developer"},{"body":"I have updated the patch to fix the above failures.","from":"developer"},{"body":"Testing with this patch ported to 0.94, via\n\n{noformat}\nmvn -PlocalTests -Psecurity \\\n -Dhadoop.profile=23 -Dhadoop.version=2.0.0-SNAPSHOT \\\n clean test \\\n -Dtest=org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n{noformat}\n\nI see this failure:\n\n{noformat}\nTests run: 1, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 45.473 sec <<< FAILURE!\ntestMultiRegionTable(org.apache.hadoop.hbase.mapreduce.TestTableMapReduce) Time elapsed: 11.876 sec <<< ERROR!\njava.lang.NullPointerException\n\tat org.apache.hadoop.net.DNS.reverseDns(DNS.java:93)\n\tat org.apache.hadoop.hbase.mapreduce.TableInputFormatBase.reverseDNS(TableInputFormatBase.java:200)\n\tat org.apache.hadoop.hbase.mapreduce.TableInputFormatBase.getSplits(TableInputFormatBase.java:165)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.writeNewSplits(JobSubmitter.java:451)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.writeSplits(JobSubmitter.java:468)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.submitJobInternal(JobSubmitter.java:360)\n\tat org.apache.hadoop.mapreduce.Job$11.run(Job.java:1226)\n\tat org.apache.hadoop.mapreduce.Job$11.run(Job.java:1223)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:416)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1232)\n\tat org.apache.hadoop.mapreduce.Job.submit(Job.java:1223)\n\tat org.apache.hadoop.mapreduce.Job.waitForCompletion(Job.java:1244)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.runTestOnTable(TestTableMapReduce.java:151)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.testMultiRegionTable(TestTableMapReduce.java:129)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:45)\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15)\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:42)\n\tat org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:20)\n\tat org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:47)\n\tat org.junit.rules.RunRules.evaluate(RunRules.java:18)\n\tat org.junit.runners.ParentRunner.runLeaf(ParentRunner.java:263)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:68)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:47)\n\tat org.junit.runners.ParentRunner$3.run(ParentRunner.java:231)\n\tat org.junit.runners.ParentRunner$1.schedule(ParentRunner.java:60)\n\tat org.junit.runners.ParentRunner.runChildren(ParentRunner.java:229)\n\tat org.junit.runners.ParentRunner.access$000(ParentRunner.java:50)\n\tat org.junit.runners.ParentRunner$2.evaluate(ParentRunner.java:222)\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:28)\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:30)\n\tat org.junit.runners.ParentRunner.run(ParentRunner.java:300)\n\tat org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n\tat org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n\tat org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n\tat org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n\tat org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\n{noformat}\n\nand probably there are more but this is the only one I've tried so far.\n\nIs it just me?","from":"developer"},{"body":"I saw lots of failures too. I am testing with patch for 5963 and 5966 now.","from":"developer"},{"body":"Above failure may have been due to locally applied patches for other ongoing issues with 2.0.0-alpha, so I ran it again with a clean 0.94 branch checkout and see this:\n\n{noformat}\n-------------------------------------------------------------------------------\nTest set: org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n-------------------------------------------------------------------------------\nTests run: 1, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 205.635 sec <<< FAILURE!\norg.apache.hadoop.hbase.mapreduce.TestTableMapReduce Time elapsed: 0 sec <<< ERROR!\njava.io.IOException: Shutting down\n\tat org.apache.hadoop.hbase.MiniHBaseCluster.init(MiniHBaseCluster.java:203)\n\tat org.apache.hadoop.hbase.MiniHBaseCluster.(MiniHBaseCluster.java:76)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniHBaseCluster(HBaseTestingUtility.java:632)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniCluster(HBaseTestingUtility.java:606)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniCluster(HBaseTestingUtility.java:554)\n\tat org.apache.hadoop.hbase.HBaseTestingUtility.startMiniCluster(HBaseTestingUtility.java:523)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.beforeClass(TestTableMapReduce.java:68)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:45)\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15)\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:42)\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:27)\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:30)\n\tat org.junit.runners.ParentRunner.run(ParentRunner.java:300)\n\tat org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n\tat org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n\tat org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n\tat org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n\tat org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\nCaused by: java.lang.RuntimeException: Master not initialized after 200 seconds\n\tat org.apache.hadoop.hbase.util.JVMClusterUtil.startup(JVMClusterUtil.java:206)\n\tat org.apache.hadoop.hbase.LocalHBaseCluster.startup(LocalHBaseCluster.java:422)\n\tat org.apache.hadoop.hbase.MiniHBaseCluster.init(MiniHBaseCluster.java:196)\n\t... 28 more\n{noformat}\n\nSame result on a clean checkout of TRUNK.","from":"developer"},{"body":"Combined with patches from HBASE-5964 and HBASE-5963, I saw this:\n{code}\ntestMultiRegionTable(org.apache.hadoop.hbase.mapreduce.TestTableMapReduce) Time elapsed: 47.489 sec <<< FAILURE!\njava.lang.AssertionError\n at org.junit.Assert.fail(Assert.java:92)\n at org.junit.Assert.assertTrue(Assert.java:43)\n at org.junit.Assert.assertTrue(Assert.java:54)\n at org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.runTestOnTable(TestTableMapReduce.java:151)\n at org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.testMultiRegionTable(TestTableMapReduce.java:129)\n{code}\nIn test output, I found:\n{code}\n2012-05-09 10:05:11,141 ERROR [main] mapreduce.TableInputFormatBase(171): Cannot resolve the host name for /10.249.196.102 because of javax.naming.NameNotFoundException: DNS name not found [response code 3]; remaining name '102.196.249.10.in-addr.arpa'\n{code}","from":"developer"},{"body":"I forgot to apply HBASE-5964 to the fresh checkout. After doing that, I see the same issue that Ted reports.","from":"developer"},{"body":"The patch is based on the patch Andrew posted.\n\nTogether with the patch for HBASE-5975, TestTableMapReduce is green for me.","from":"developer"},{"body":"These tests are green on my box.","from":"developer"},{"body":"Can we push this patch to trunk? Based on our testing, it does fix the MR test failures with HADOOP 2.0.0.","from":"developer"},{"body":"What am I doing wrong? I tried it three times here on my laptop:\n\n{code}\n-------------------------------------------------------------------------------\nTest set: org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n-------------------------------------------------------------------------------\nTests run: 1, Failures: 1, Errors: 0, Skipped: 0, Time elapsed: 64.363 sec <<< FAILURE!\ntestMultiRegionTable(org.apache.hadoop.hbase.mapreduce.TestTableMapReduce) Time elapsed: 30.455 sec <<< FAILURE!\njava.lang.AssertionError\n at org.junit.Assert.fail(Assert.java:92)\n at org.junit.Assert.assertTrue(Assert.java:43)\n at org.junit.Assert.assertTrue(Assert.java:54)\n at org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.runTestOnTable(TestTableMapReduce.java:151)\n at org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.testMultiRegionTable(TestTableMapReduce.java:129)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:45)\n at org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15)\n at org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:42)\n at org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:20)\n at org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:47)\n at org.junit.rules.RunRules.evaluate(RunRules.java:18)\n at org.junit.runners.ParentRunner.runLeaf(ParentRunner.java:263)\n at org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:68)\n at org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:47)\n at org.junit.runners.ParentRunner$3.run(ParentRunner.java:231)\n at org.junit.runners.ParentRunner$1.schedule(ParentRunner.java:60)\n at org.junit.runners.ParentRunner.runChildren(ParentRunner.java:229)\n at org.junit.runners.ParentRunner.access$000(ParentRunner.java:50)\n at org.junit.runners.ParentRunner$2.evaluate(ParentRunner.java:222)\n at org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:28)\n at org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:30)\n at org.junit.runners.ParentRunner.run(ParentRunner.java:300)\n at org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n at org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n at org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n at java.lang.reflect.Method.invoke(Method.java:597)\n at org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n at org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n at org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n at org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n at org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\n{code}\n\nI did this:\n\n{code}\n$ MAVEN_OPTS=\"-Xmx2g\" ~/bin/mvn/bin/mvn -PlocalTests -Psecurity -Dhadoop.profile=23 -Dhadoop.version=2.0.0-SNAPSHOT clean test -Dtest=org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n{code}","from":"developer"},{"body":"Did you set JAVA_HOME?","from":"developer"},{"body":"It passed when I set JAVA_HOME. Thanks for the patch Jimmy.","from":"developer"},{"body":"Discussed with Jimmy. Let's have this in 0.94.1","from":"developer"},{"body":"Attached patch for 0.94. Ran TestTableMapReduce against both 1.0 and 2.0 hadoop profiles, both passed:\n\n\nmvn test -PlocalTests -Dtest=org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n\n-------------------------------------------------------\n T E S T S\n-------------------------------------------------------\nRunning org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 188.087 sec\n\nResults :\n\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0\n\nmvn test -PlocalTests -Dhadoop.profile=2.0 -Dtest=org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n\n-------------------------------------------------------\n T E S T S\n-------------------------------------------------------\nRunning org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 167.49 sec\n\nResults :\n\nTests run: 1, Failures: 0, Errors: 0, Skipped: 0\n","from":"developer"},{"body":"looks good to me, will commit to 0.94 tonight if no objection.","from":"developer"},{"body":"+1","from":"developer"},{"body":"Integrated to 0.94. Thank Greg for the patch, Lars for the review.","from":"developer"}],"created":"2012-05-09T01:04:01.000+0000","description":"Some fairly recent change in Hadoop 2.0.0-alpha has broken our MapReduce test rigging. Below is a representative error, can be easily reproduced with:\n\n{noformat}\nmvn -PlocalTests -Psecurity \\\n -Dhadoop.profile=23 -Dhadoop.version=2.0.0-SNAPSHOT \\\n clean test \\\n -Dtest=org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n{noformat}\n\nAnd the result:\n\n{noformat}\n-------------------------------------------------------\n T E S T S\n-------------------------------------------------------\nRunning org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\nTests run: 1, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 54.292 sec <<< FAILURE!\n\n-------------------------------------------------------------------------------\nTest set: org.apache.hadoop.hbase.mapreduce.TestTableMapReduce\n-------------------------------------------------------------------------------\nTests run: 1, Failures: 0, Errors: 1, Skipped: 0, Time elapsed: 54.292 sec <<< FAILURE!\ntestMultiRegionTable(org.apache.hadoop.hbase.mapreduce.TestTableMapReduce) Time elapsed: 21.935 sec <<< ERROR!\njava.lang.reflect.UndeclaredThrowableException\n\tat org.apache.hadoop.yarn.exceptions.impl.pb.YarnRemoteExceptionPBImpl.unwrapAndThrowException(YarnRemoteExceptionPBImpl.java:135)\n\tat org.apache.hadoop.yarn.api.impl.pb.client.ClientRMProtocolPBClientImpl.getNewApplication(ClientRMProtocolPBClientImpl.java:134)\n\tat org.apache.hadoop.mapred.ResourceMgrDelegate.getNewJobID(ResourceMgrDelegate.java:183)\n\tat org.apache.hadoop.mapred.YARNRunner.getNewJobID(YARNRunner.java:216)\n\tat org.apache.hadoop.mapreduce.JobSubmitter.submitJobInternal(JobSubmitter.java:339)\n\tat org.apache.hadoop.mapreduce.Job$11.run(Job.java:1226)\n\tat org.apache.hadoop.mapreduce.Job$11.run(Job.java:1223)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:416)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1232)\n\tat org.apache.hadoop.mapreduce.Job.submit(Job.java:1223)\n\tat org.apache.hadoop.mapreduce.Job.waitForCompletion(Job.java:1244)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.runTestOnTable(TestTableMapReduce.java:151)\n\tat org.apache.hadoop.hbase.mapreduce.TestTableMapReduce.testMultiRegionTable(TestTableMapReduce.java:129)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:45)\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:15)\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:42)\n\tat org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:20)\n\tat org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:47)\n\tat org.junit.rules.RunRules.evaluate(RunRules.java:18)\n\tat org.junit.runners.ParentRunner.runLeaf(ParentRunner.java:263)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:68)\n\tat org.junit.runners.BlockJUnit4ClassRunner.runChild(BlockJUnit4ClassRunner.java:47)\n\tat org.junit.runners.ParentRunner$3.run(ParentRunner.java:231)\n\tat org.junit.runners.ParentRunner$1.schedule(ParentRunner.java:60)\n\tat org.junit.runners.ParentRunner.runChildren(ParentRunner.java:229)\n\tat org.junit.runners.ParentRunner.access$000(ParentRunner.java:50)\n\tat org.junit.runners.ParentRunner$2.evaluate(ParentRunner.java:222)\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:28)\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:30)\n\tat org.junit.runners.ParentRunner.run(ParentRunner.java:300)\n\tat org.apache.maven.surefire.junit4.JUnit4TestSet.execute(JUnit4TestSet.java:53)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.executeTestSet(JUnit4Provider.java:123)\n\tat org.apache.maven.surefire.junit4.JUnit4Provider.invoke(JUnit4Provider.java:104)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.maven.surefire.util.ReflectionUtils.invokeMethodWithArray(ReflectionUtils.java:164)\n\tat org.apache.maven.surefire.booter.ProviderFactory$ProviderProxy.invoke(ProviderFactory.java:110)\n\tat org.apache.maven.surefire.booter.SurefireStarter.invokeProvider(SurefireStarter.java:175)\n\tat org.apache.maven.surefire.booter.SurefireStarter.runSuitesInProcessWhenForked(SurefireStarter.java:81)\n\tat org.apache.maven.surefire.booter.ForkedBooter.main(ForkedBooter.java:68)\nCaused by: com.google.protobuf.ServiceException: java.net.ConnectException: Call From acer.localdomain/192.168.122.1 to 0.0.0.0:8032 failed on connection exception: java.net.ConnectException: Connection refused; For more details see: http://wiki.apache.org/hadoop/ConnectionRefused\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:188)\n\tat $Proxy89.getNewApplication(Unknown Source)\n\tat org.apache.hadoop.yarn.api.impl.pb.client.ClientRMProtocolPBClientImpl.getNewApplication(ClientRMProtocolPBClientImpl.java:132)\n\t... 45 more\nCaused by: java.net.ConnectException: Call From acer.localdomain/192.168.122.1 to 0.0.0.0:8032 failed on connection exception: java.net.ConnectException: Connection refused; For more details see: http://wiki.apache.org/hadoop/ConnectionRefused\n\tat org.apache.hadoop.net.NetUtils.wrapException(NetUtils.java:725)\n\tat org.apache.hadoop.ipc.Client.call(Client.java:1160)\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:185)\n\t... 47 more\nCaused by: java.net.ConnectException: Connection refused\n\tat sun.nio.ch.SocketChannelImpl.checkConnect(Native Method)\n\tat sun.nio.ch.SocketChannelImpl.finishConnect(SocketChannelImpl.java:592)\n\tat org.apache.hadoop.net.SocketIOWithTimeout.connect(SocketIOWithTimeout.java:206)\n\tat org.apache.hadoop.net.NetUtils.connect(NetUtils.java:522)\n\tat org.apache.hadoop.net.NetUtils.connect(NetUtils.java:487)\n\tat org.apache.hadoop.ipc.Client$Connection.setupConnection(Client.java:469)\n\tat org.apache.hadoop.ipc.Client$Connection.setupIOstreams(Client.java:563)\n\tat org.apache.hadoop.ipc.Client$Connection.access$2000(Client.java:212)\n\tat org.apache.hadoop.ipc.Client.getConnection(Client.java:1266)\n\tat org.apache.hadoop.ipc.Client.call(Client.java:1136)\n\t... 48 more\n{noformat}\n","issue_id":"12554413","key":"HBASE-5966","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-07-19T23:44:05.000+0000","role":"fixed_distractor","summary":"MapReduce based tests broken on Hadoop 2.0.0-alpha"} {"case_id":"12395373","cluster":"DISTRACTOR-HBASE-615","comments":[{"body":"There's two ways this could go. Either we can fix it by widening the margin by which you have to be \"over\" the load average before we start rebalancing region assignments, or we have to redo this code to be more sensitive to the load as it is being computed on the assignment side of things. \n\nI'll do the first no matter what as a test because it will be very simple. However, if the problem is that the definitions of \"overloaded\" from the perspective of load reported by the regionserver (this is some combo of regions and requests right now, yes?) and the math done by the rebalancing code are inherently different, then we'll need to either make the existing load balancing on assignment dumber or the rebalancing smarter. In the long run, we're definitely going to want to do the latter, but it requires us to start tracking requests and anything else that goes into the balancing computation at the region level, as well as actually reporting that information when the workers check in with their short list of reassignable regions. That way, when we're deciding how many regions to unassign, we can make informed decisions, rather than just trying to equalize averages.\n\n","created":"2008-05-22T14:11:19.900+0000"},{"body":"Here's a patch for the easiest of the ideas. Jim, can you apply this and try it out with your 51-region table?","created":"2008-05-22T14:12:52.699+0000"},{"body":"Well now it oscillates moving three regions around. There are 40 regions including root and meta, 4 region servers, but the master refuses to give more regions to the region server hosting the meta region. It has 4 regions and the rest have 11, 12 or 13","created":"2008-05-22T22:35:50.465+0000"},{"body":"Please correct me if I am wrong or overlook something.\n\nDuring startup, META will be requested more than other regions. Therefore,\nthe RegionServer that serves META will be considered more \"loaded\" than\nothers. So, we tends not to assign more regions to that one. However,\nour rebalance algo currently considers only # of loaded regions as the \"load\"\nfor region servers. That's the cause of oscillation at startup.\n\nI'm thinking of the possibility that during startup, we just use assign\nevenly to all region servers. Once this is stabilized, we start to consider\n# of requests as part of the server load. Moreover, the \"# of requests\" here\nshould be calculated from a period of time, otherwise, we may moving regions\njust because some spikes.\n\njust my 2 cents.","created":"2008-06-11T16:43:37.026+0000"},{"body":"> Rong-En Fan - 11/Jun/08 09:43 AM\n> Please correct me if I am wrong or overlook something.\n> \n> During startup, META will be requested more than other regions. Therefore,\n> the RegionServer that serves META will be considered more \"loaded\" than\n> others. So, we tends not to assign more regions to that one. However,\n> our rebalance algo currently considers only # of loaded regions as the \"load\"\n> for region servers. That's the cause of oscillation at startup.\n> \n> I'm thinking of the possibility that during startup, we just use assign\n> evenly to all region servers. Once this is stabilized, we start to consider\n>\n> 1. of requests as part of the server load. Moreover, the \"# of requests\" here\n> should be calculated from a period of time, otherwise, we may moving regions\n> just because some spikes.\n\nYou are absolutely correct. During startup, the server hosting the meta region gets all the requests,\nso its one region gets multiplied by the number of requests giving a \"load\" that is far greater than all\nthe other servers which are getting no requests and consequently their \"load\" == number of regions\nthey are serving.\n\nShould be a fairly easy fix to ignore requests during startup.\n\nBTW, I verified this by changing HServerLoad.getLoad() to just return the number of regions. The\ncluster had all regions on-line within a couple of minutes and they were balanced. When the\nnumber of requests was factored in during startup, the cluster did not achieve a steady state\nfor 1/2 hour (after which I gave up).\n\njust my 2 cents.\n","created":"2008-06-18T20:26:08.495+0000"},{"body":"Should we just disable factoring in requests in a region server's load for the moment? It might lead to a worse distribution of regions, but it might also make no difference, if on average all regions are equally busy, which would be the case if you're running big time map/reduce jobs. If you have hot rows/regions, then this could hurt you, but it's hard to say by how much.","created":"2008-06-18T22:02:54.313+0000"},{"body":"Would that mean no balancing? Or just that balancing would be every server has ~ same number of regions? If latter, lets do that for now. It would still be massive improvement over what we have in 0.1 branch. We can make a new issue for making balancing code consider loading.","created":"2008-06-18T22:29:48.219+0000"},{"body":"I believe Bryan is saying the latter way.","created":"2008-06-19T02:51:59.844+0000"},{"body":"Here's a patch that removes the number of requests from the load considerations. It passes unit tests (except TestMetaUtils, which times out, but appears unrelated). \n\nJim, can you test this on your cluster and see if it reduces oscillation?","created":"2008-06-19T15:07:04.916+0000"},{"body":"I actually have the same patch applied locally, the cluster will start in a balance state\nwithin few minutes.","created":"2008-06-19T15:17:52.766+0000"},{"body":"I had applied a similar change to my cluster and and it worked just fine. Reviewed patch. +1","created":"2008-06-19T16:24:09.417+0000"},{"body":"I just committed this.","created":"2008-06-19T16:30:41.518+0000"}],"conversations":[{"body":"When starting a cluster with four region servers and a large table (49 regions) (+root +meta) = 51 total regions, the region balancer oscillates for a very long time and does not seem to reach a steady state.\n\nAdditionally, for whatever reason, it seems reluctant to assign regions to the first of four region servers, which may be the root cause. In my test, the first server had 10 regions assigned, the second and fourth had 13 regions assigned, and the master would continually assign and deassign 2 regions to the third server, which oscillated between 13 and 15 regions. If it assigned the two fluctuating regions to the first server, it would achieve the best balance possible: 12, 13, 13, 13.\n\nAfter 20 minutes, it had not stopped oscillating. An application trying to work against this cluster would run very slowly as it would be continually re-finding the two regions in flux.\n\nWhen the table was being created, regions were nicely balanced. On restart, however, it just would not settle down.\n\nPerhaps the balancer should set a target number of regions for each server which when the server achieved +/- 1 regions, the rebalancer would not try to change unless the number of regions changed.","from":"reporter","subject":"Region balancer oscillates during cluster startup"},{"body":"There's two ways this could go. Either we can fix it by widening the margin by which you have to be \"over\" the load average before we start rebalancing region assignments, or we have to redo this code to be more sensitive to the load as it is being computed on the assignment side of things. \n\nI'll do the first no matter what as a test because it will be very simple. However, if the problem is that the definitions of \"overloaded\" from the perspective of load reported by the regionserver (this is some combo of regions and requests right now, yes?) and the math done by the rebalancing code are inherently different, then we'll need to either make the existing load balancing on assignment dumber or the rebalancing smarter. In the long run, we're definitely going to want to do the latter, but it requires us to start tracking requests and anything else that goes into the balancing computation at the region level, as well as actually reporting that information when the workers check in with their short list of reassignable regions. That way, when we're deciding how many regions to unassign, we can make informed decisions, rather than just trying to equalize averages.\n\n","from":"developer"},{"body":"Here's a patch for the easiest of the ideas. Jim, can you apply this and try it out with your 51-region table?","from":"developer"},{"body":"Well now it oscillates moving three regions around. There are 40 regions including root and meta, 4 region servers, but the master refuses to give more regions to the region server hosting the meta region. It has 4 regions and the rest have 11, 12 or 13","from":"developer"},{"body":"Please correct me if I am wrong or overlook something.\n\nDuring startup, META will be requested more than other regions. Therefore,\nthe RegionServer that serves META will be considered more \"loaded\" than\nothers. So, we tends not to assign more regions to that one. However,\nour rebalance algo currently considers only # of loaded regions as the \"load\"\nfor region servers. That's the cause of oscillation at startup.\n\nI'm thinking of the possibility that during startup, we just use assign\nevenly to all region servers. Once this is stabilized, we start to consider\n# of requests as part of the server load. Moreover, the \"# of requests\" here\nshould be calculated from a period of time, otherwise, we may moving regions\njust because some spikes.\n\njust my 2 cents.","from":"developer"},{"body":"> Rong-En Fan - 11/Jun/08 09:43 AM\n> Please correct me if I am wrong or overlook something.\n> \n> During startup, META will be requested more than other regions. Therefore,\n> the RegionServer that serves META will be considered more \"loaded\" than\n> others. So, we tends not to assign more regions to that one. However,\n> our rebalance algo currently considers only # of loaded regions as the \"load\"\n> for region servers. That's the cause of oscillation at startup.\n> \n> I'm thinking of the possibility that during startup, we just use assign\n> evenly to all region servers. Once this is stabilized, we start to consider\n>\n> 1. of requests as part of the server load. Moreover, the \"# of requests\" here\n> should be calculated from a period of time, otherwise, we may moving regions\n> just because some spikes.\n\nYou are absolutely correct. During startup, the server hosting the meta region gets all the requests,\nso its one region gets multiplied by the number of requests giving a \"load\" that is far greater than all\nthe other servers which are getting no requests and consequently their \"load\" == number of regions\nthey are serving.\n\nShould be a fairly easy fix to ignore requests during startup.\n\nBTW, I verified this by changing HServerLoad.getLoad() to just return the number of regions. The\ncluster had all regions on-line within a couple of minutes and they were balanced. When the\nnumber of requests was factored in during startup, the cluster did not achieve a steady state\nfor 1/2 hour (after which I gave up).\n\njust my 2 cents.\n","from":"developer"},{"body":"Should we just disable factoring in requests in a region server's load for the moment? It might lead to a worse distribution of regions, but it might also make no difference, if on average all regions are equally busy, which would be the case if you're running big time map/reduce jobs. If you have hot rows/regions, then this could hurt you, but it's hard to say by how much.","from":"developer"},{"body":"Would that mean no balancing? Or just that balancing would be every server has ~ same number of regions? If latter, lets do that for now. It would still be massive improvement over what we have in 0.1 branch. We can make a new issue for making balancing code consider loading.","from":"developer"},{"body":"I believe Bryan is saying the latter way.","from":"developer"},{"body":"Here's a patch that removes the number of requests from the load considerations. It passes unit tests (except TestMetaUtils, which times out, but appears unrelated). \n\nJim, can you test this on your cluster and see if it reduces oscillation?","from":"developer"},{"body":"I actually have the same patch applied locally, the cluster will start in a balance state\nwithin few minutes.","from":"developer"},{"body":"I had applied a similar change to my cluster and and it worked just fine. Reviewed patch. +1","from":"developer"},{"body":"I just committed this.","from":"developer"}],"created":"2008-05-06T01:33:13.000+0000","description":"When starting a cluster with four region servers and a large table (49 regions) (+root +meta) = 51 total regions, the region balancer oscillates for a very long time and does not seem to reach a steady state.\n\nAdditionally, for whatever reason, it seems reluctant to assign regions to the first of four region servers, which may be the root cause. In my test, the first server had 10 regions assigned, the second and fourth had 13 regions assigned, and the master would continually assign and deassign 2 regions to the third server, which oscillated between 13 and 15 regions. If it assigned the two fluctuating regions to the first server, it would achieve the best balance possible: 12, 13, 13, 13.\n\nAfter 20 minutes, it had not stopped oscillating. An application trying to work against this cluster would run very slowly as it would be continually re-finding the two regions in flux.\n\nWhen the table was being created, regions were nicely balanced. On restart, however, it just would not settle down.\n\nPerhaps the balancer should set a target number of regions for each server which when the server achieved +/- 1 regions, the rebalancer would not try to change unless the number of regions changed.","issue_id":"12395373","key":"HBASE-615","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2008-06-19T16:30:41.000+0000","role":"fixed_distractor","summary":"Region balancer oscillates during cluster startup"} {"case_id":"12560849","cluster":"DISTRACTOR-HBASE-6220","comments":[{"body":"The \"_avg_time\" is hardcoded in org.apache.hadoop.metrics.util.MetricsTimeVaryingRate in hadoop-core. I could duplicate the functionality in order to change the String, or try to get a change pushed upstream. Both solutions seem like overkill. I also noticed that the underlying metrics package is deprecated and the metrics2 package seems to be the way forward. As I am a \"noob\" does anyone have thoughts on what the proper approach might be?\n\nThanks,\n-Paul","created":"2012-06-26T01:31:44.373+0000"},{"body":"Creating another class is probably the correct way. See how MetricsHistogram does it. Though there is a push to re-vamp the metrics usage in hbase so things could change with that.","created":"2012-06-26T01:49:02.990+0000"},{"body":"I moved over the metrics that used PersistentMetricsTimeVaryingRate to MetricsHistogram because it seemed like a better fit. \n\nNo tests are included, but I am looking for feedback for correct approach, and what areas to test. I could not find any tests that exercised RegionServerMetrics or MasterServerMetrics, so I did not have much to crib off of.","created":"2012-06-26T03:05:54.909+0000"},{"body":"Sorry, I forgot to add that this obviously changes the metrics that are logged, and I am not sure if that is OK. Obviously this is something to consider. \n\nThanks,\n\n-Paul","created":"2012-06-26T03:06:51.003+0000"},{"body":"From the mailing list I assume this is getting swallowed up by the push to use metrics2 in https://issues.apache.org/jira/browse/HBASE-4050. Let me know what action I should take as I am a new contributor.\n\nThanks,\n\n-Paul","created":"2012-07-06T15:46:25.011+0000"},{"body":"looks good to me. Have you verified that the changed metrics then get out to jmx?\n","created":"2012-07-06T16:36:53.658+0000"},{"body":"Patch integrated to trunk.\n\nThanks for the patch, Paul.\n\nThanks for the review, Elliot.","created":"2012-07-06T17:21:01.378+0000"},{"body":"Did Elliott's question of Paul get answered? i.e. if the metrics come out in jmx (and in /jmx servlet?) Good stuff.","created":"2012-07-06T21:25:32.231+0000"},{"body":"@stack I don't think it did. :)\n\nI'll take a look soon, hopefully this weekend, as I'm not completely sure.","created":"2012-07-06T22:30:54.381+0000"},{"body":"@Paul Easy test is browse to /jmx on regionserver. Look for your new metrics messings. Thanks for the contrib.","created":"2012-07-07T10:42:32.222+0000"},{"body":"Apologies for not getting to this earlier.\n\nI noticed that the flush metrics were not changed in the patch. Was there a reason for that?","created":"2012-07-10T13:33:30.856+0000"},{"body":"Sorry I've been away this week with intermittent internet connectivity. I'll try to reply to all of this soon.","created":"2012-07-11T22:27:00.400+0000"},{"body":"Hey all, back from the offline world, replies below:\n\n@stack Yep, everything seems to show up in /jmx just fine.\n\n@David Wang No reason that it wasn't included, I've actually attached another patch to switch those over the Metrics Histogram as well if desired we can include it.\n\nThanks all, and sorry for the slow responses.","created":"2012-07-18T00:32:28.250+0000"},{"body":"I made HBASE-6419 to apply Pauls' amendment.","created":"2012-07-18T14:48:57.710+0000"},{"body":"Should I resolve this or someone else? Thanks.","created":"2012-07-18T19:15:21.322+0000"}],"conversations":[{"body":"PersistentMetricsTimeVaryingRate gets used for metrics that are not time-based, leading to confusing names such as \"avg_time\" for compaction size, etc. You hav to read the code in order to understand that this is actually referring to bytes, not seconds.","from":"reporter","subject":"PersistentMetricsTimeVaryingRate gets used for non-time-based metrics"},{"body":"The \"_avg_time\" is hardcoded in org.apache.hadoop.metrics.util.MetricsTimeVaryingRate in hadoop-core. I could duplicate the functionality in order to change the String, or try to get a change pushed upstream. Both solutions seem like overkill. I also noticed that the underlying metrics package is deprecated and the metrics2 package seems to be the way forward. As I am a \"noob\" does anyone have thoughts on what the proper approach might be?\n\nThanks,\n-Paul","from":"developer"},{"body":"Creating another class is probably the correct way. See how MetricsHistogram does it. Though there is a push to re-vamp the metrics usage in hbase so things could change with that.","from":"developer"},{"body":"I moved over the metrics that used PersistentMetricsTimeVaryingRate to MetricsHistogram because it seemed like a better fit. \n\nNo tests are included, but I am looking for feedback for correct approach, and what areas to test. I could not find any tests that exercised RegionServerMetrics or MasterServerMetrics, so I did not have much to crib off of.","from":"developer"},{"body":"Sorry, I forgot to add that this obviously changes the metrics that are logged, and I am not sure if that is OK. Obviously this is something to consider. \n\nThanks,\n\n-Paul","from":"developer"},{"body":"From the mailing list I assume this is getting swallowed up by the push to use metrics2 in https://issues.apache.org/jira/browse/HBASE-4050. Let me know what action I should take as I am a new contributor.\n\nThanks,\n\n-Paul","from":"developer"},{"body":"looks good to me. Have you verified that the changed metrics then get out to jmx?\n","from":"developer"},{"body":"Patch integrated to trunk.\n\nThanks for the patch, Paul.\n\nThanks for the review, Elliot.","from":"developer"},{"body":"Did Elliott's question of Paul get answered? i.e. if the metrics come out in jmx (and in /jmx servlet?) Good stuff.","from":"developer"},{"body":"@stack I don't think it did. :)\n\nI'll take a look soon, hopefully this weekend, as I'm not completely sure.","from":"developer"},{"body":"@Paul Easy test is browse to /jmx on regionserver. Look for your new metrics messings. Thanks for the contrib.","from":"developer"},{"body":"Apologies for not getting to this earlier.\n\nI noticed that the flush metrics were not changed in the patch. Was there a reason for that?","from":"developer"},{"body":"Sorry I've been away this week with intermittent internet connectivity. I'll try to reply to all of this soon.","from":"developer"},{"body":"Hey all, back from the offline world, replies below:\n\n@stack Yep, everything seems to show up in /jmx just fine.\n\n@David Wang No reason that it wasn't included, I've actually attached another patch to switch those over the Metrics Histogram as well if desired we can include it.\n\nThanks all, and sorry for the slow responses.","from":"developer"},{"body":"I made HBASE-6419 to apply Pauls' amendment.","from":"developer"},{"body":"Should I resolve this or someone else? Thanks.","from":"developer"}],"created":"2012-06-15T23:59:26.000+0000","description":"PersistentMetricsTimeVaryingRate gets used for metrics that are not time-based, leading to confusing names such as \"avg_time\" for compaction size, etc. You hav to read the code in order to understand that this is actually referring to bytes, not seconds.","issue_id":"12560849","key":"HBASE-6220","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-07-18T19:19:04.000+0000","role":"fixed_distractor","summary":"PersistentMetricsTimeVaryingRate gets used for non-time-based metrics"} {"case_id":"12595151","cluster":"DISTRACTOR-HBASE-6239","comments":[{"body":"Easy fix in ReplicationSink, just set the ts that comes with each KV, then I'm adding test that's also good for 0.94 and onward (and it passes).","created":"2012-06-19T19:16:19.535+0000"},{"body":"I would argue that this bug is not \"minor\", because we're talking about data being corrupted by HBase.","created":"2012-06-22T16:25:48.435+0000"},{"body":"+1 on patch.\n\nPity this is so ugly:\n\n{code}\n+ HLog.Entry[] entries = new HLog.Entry[3];\n+ long now = System.currentTimeMillis();\n+ for(int i = 0; i < 3; i++) {\n+ entries[i] = createEntry(TABLE_NAME1, 1, i, KeyValue.Type.Put, now+i);\n+ }\n+ // Kinda ugly, trying to merge all the entries into one\n+ entries[0].getEdit().add(entries[1].getEdit().getKeyValues().get(0));\n+ entries[0].getEdit().add(entries[2].getEdit().getKeyValues().get(0));\n+ HLog.Entry[] entry = new HLog.Entry[1];\n+ entry[0] = entries[0];\n+ SINK.replicateEntries(entry);\n{code}\n\n...but its not a blocker. Commit. Backport to 0.94?","created":"2012-06-28T18:12:51.798+0000"},{"body":"Adding as a potential for 0.90.7 release.","created":"2012-06-29T17:23:18.095+0000"},{"body":"Bumping to 0.90.8 -- I'm personally not concerned about replication in 0.90.7","created":"2012-07-12T00:25:18.699+0000"},{"body":"This means HBase replication will still corrupt timestamps in 0.90.7, which in many cases makes replication useless. Are you sure?","created":"2012-07-12T06:54:44.252+0000"},{"body":"If this is not a blocker for 0.92's where replication is considered robust, I'm not sure it should be a blocker on 0.90 where it is more experimental. That said, if a patch lands for 0.90 I'll gladly accept it -- I'm just not going to block a 0.90.7 because of it.","created":"2012-07-12T07:18:23.993+0000"},{"body":"Integrated to 0.92 branch.\n\nThanks for the patch, J-D.\n\nThanks for the review, Stack.","created":"2012-08-31T05:04:02.458+0000"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","created":"2015-11-20T11:53:12.655+0000"}],"conversations":[{"body":"ReplicationSink assumes that all the KVs for the same row inside a WALEdit will have the same timestamp, which is not necessarily the case.\n\nThis only affects 0.90 and 0.92 since HBASE-5203 fixes it in 0.94","from":"reporter","subject":"[replication] ReplicationSink uses the ts of the first KV for the other KVs in the same row"},{"body":"Easy fix in ReplicationSink, just set the ts that comes with each KV, then I'm adding test that's also good for 0.94 and onward (and it passes).","from":"developer"},{"body":"I would argue that this bug is not \"minor\", because we're talking about data being corrupted by HBase.","from":"developer"},{"body":"+1 on patch.\n\nPity this is so ugly:\n\n{code}\n+ HLog.Entry[] entries = new HLog.Entry[3];\n+ long now = System.currentTimeMillis();\n+ for(int i = 0; i < 3; i++) {\n+ entries[i] = createEntry(TABLE_NAME1, 1, i, KeyValue.Type.Put, now+i);\n+ }\n+ // Kinda ugly, trying to merge all the entries into one\n+ entries[0].getEdit().add(entries[1].getEdit().getKeyValues().get(0));\n+ entries[0].getEdit().add(entries[2].getEdit().getKeyValues().get(0));\n+ HLog.Entry[] entry = new HLog.Entry[1];\n+ entry[0] = entries[0];\n+ SINK.replicateEntries(entry);\n{code}\n\n...but its not a blocker. Commit. Backport to 0.94?","from":"developer"},{"body":"Adding as a potential for 0.90.7 release.","from":"developer"},{"body":"Bumping to 0.90.8 -- I'm personally not concerned about replication in 0.90.7","from":"developer"},{"body":"This means HBase replication will still corrupt timestamps in 0.90.7, which in many cases makes replication useless. Are you sure?","from":"developer"},{"body":"If this is not a blocker for 0.92's where replication is considered robust, I'm not sure it should be a blocker on 0.90 where it is more experimental. That said, if a patch lands for 0.90 I'll gladly accept it -- I'm just not going to block a 0.90.7 because of it.","from":"developer"},{"body":"Integrated to 0.92 branch.\n\nThanks for the patch, J-D.\n\nThanks for the review, Stack.","from":"developer"},{"body":"This issue was closed as part of a bulk closing operation on 2015-11-20. All issues that have been resolved and where all fixVersions have been released have been closed (following discussions on the mailing list).","from":"developer"}],"created":"2012-06-19T19:13:28.000+0000","description":"ReplicationSink assumes that all the KVs for the same row inside a WALEdit will have the same timestamp, which is not necessarily the case.\n\nThis only affects 0.90 and 0.92 since HBASE-5203 fixes it in 0.94","issue_id":"12595151","key":"HBASE-6239","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-10-05T21:52:48.000+0000","role":"fixed_distractor","summary":"[replication] ReplicationSink uses the ts of the first KV for the other KVs in the same row"} {"case_id":"12598225","cluster":"DISTRACTOR-HBASE-6364","comments":[{"body":"Also wanted to capture NKeywal's comments on this from the mailing list: \n\nNKeywal said: \nWhat you're describing -the 35 minutes recovery time- seems to match the code. And it's a bug (still there on trunk). Could you please create a jira for it? If you have the logs it even better.\n\nLowering the ipc.socket.timeout seems to be an acceptable partial workaround. Setting it to 10s seems ok to me. Lower than this... I don't know.\n","created":"2012-07-10T17:12:34.600+0000"},{"body":"That's a possible cause:\n\nin HBaseClient#getConnection()\n{noformat}\n do {\n synchronized (connections) {\n connection = connections.get(remoteId);\n if (connection == null) {\n connection = new Connection(remoteId);\n connections.put(remoteId, connection);\n }\n }\n } while (!connection.addCall(call));\n \n connection.setupIOstreams();\n return connection;\n }\n{noformat}\nConnection#addCall and Connection#setupIOstreams are synchronized. if #setupIOstreams fails, it marks the connection as dead, remove it from the connections list, and throws an exception. In #addCall it returns false if the connection is marked as dead. So\ncase 1 -> Sometimes, we may add the call to a connection that will be marked as dead:\n Thread 1: create the connection, add it to the connections list, call addCall\n Thread 2: get the connection, add it to the calls list \n Thread 1: get into setupIOstreams, fails, and mark the connection as dead, throws an exception, done\n Thread 2: get into setupIOstreams, see that the connection is dead, done. The call has been added to a dead connection\n\ncase 2 -> If we have a lot of threads on a dying connection, we will have:\n Thread 1: goes until setupIOstreams\n All other threads: get the connection from the list, wait on the synchronized addCall \n Thread 1: exit from setupIOstreams with an exception after 20 seconds (socket timeout)\n All other threads: call addCall, as the connection is dead, reloop\n One of these threads will create a connection\n One of these thread will win the race on addCall\n One of them will win the race on setupIOstreams\n Most of them should be waiting on addCall so reloop\n So back to the case 1 or 2.\n\n\nWe would have the same behavior on pure hadoop client (ipc.Client), as the implementation is similar, at least on 1.0.3.\n\n\nSuraj, does this match your analysis? How many region servers and regions did you have during your test? What's the client doing?","created":"2012-07-19T16:07:37.217+0000"},{"body":"stack trace snippet from test.","created":"2012-07-28T06:56:32.712+0000"},{"body":"nkeywal, my test cluster had 5 regionservers with an average of 35 regions per server. The client is doing Get calls from all threads - each thread creates a separate HTable but with the same HBaseConfiguration (so, sharing the same underlying HConnection.)\n\nCase 2 that you mention above is what I think I saw. Several threads blocked on #addCall with a lock owned by the setupIOStreams call (on NetUtils.connect()) which was timing out after 20s. \n\nI finally managed to locate the stack traces from the test ... I'm attaching a snippet from that with some BLOCKED threads and the setupIOStreams call holding the lock. \n\nI believe the same behaviour would occur on the hadoop client as well ... but perhaps because of multiple datanodes, the problem is not as acute? In the HBaseClient case, all threads repeatedly reaches out for META to rediscover the relocated regions. \n(Attached the thread dump snippet)\n","created":"2012-07-28T06:56:47.468+0000"},{"body":"Thanks Suraj. So it seems that we have the root cause. You're right, on ipc.Client it's less likely to occur. Only the NN could have the issue I think, and as the hot failover is not yet deployed in production, it has not been an issue so far. I will ping Todd on this to see what he thinks.\n\nThe only strange thing is that when tested on the 0.96 + minicluster, I have the issue but many threads seem to be able to 'escape', i.e. manage to do the 'addCall' between two blocking calls to setupIOstreams. But it could be a pure multithreading hazard. And I haven't been able to find a better cause than this.\n\nI will propose a patch, likely end of next week. Suraj, If you have some time available to validate it on your env, it would be great.","created":"2012-07-28T09:40:43.519+0000"},{"body":"v1. There are some other ways of doing this, like adding a list of dead servers and a timeout, but it does not pass the unit tests, with numerous failures. I haven't sorted out if there is a common root cause...","created":"2012-08-05T16:56:48.892+0000"},{"body":"likely to be unrelated, worked twice locally. Let's retry.","created":"2012-08-05T18:13:06.069+0000"},{"body":"https://builds.apache.org/job/PreCommit-HBASE-Build/2514/console got aborted.","created":"2012-08-05T22:26:42.051+0000"},{"body":"Patch from N.","created":"2012-08-05T22:27:09.585+0000"},{"body":"Is this alleviated (at least somewhat) by HBASE-6326?","created":"2012-08-06T04:36:02.665+0000"},{"body":"Looking at the issue, no it isn't. NM me. :)","created":"2012-08-06T04:41:37.502+0000"},{"body":"Reanalyzing the fix, there is an issue with the v1: we could have a call added to a dying connection, and this call won't get cleaned up. This is was not possible previously. Will write a v2.","created":"2012-08-06T07:49:13.241+0000"},{"body":"v2, fixes the problem mentioned above, works locally.","created":"2012-08-06T09:54:49.563+0000"},{"body":"v3 just changes a comment, not the code itself.","created":"2012-08-06T10:54:09.999+0000"},{"body":"errors are unrelated imho. Retrying to see if the third executions says something different.","created":"2012-08-06T13:30:35.370+0000"},{"body":"Good one lads.\n\nFix formatting before commit N. Make it same as surrounding code... Add spacings around brackets -- the 'else' -- and the '+' in String concatenations.\n\nI think I understand the notifying that is going on on the end of the addCall method. They line up w/ waits on Call and waits on the calls data member?\n\nWould it be hard making a test of this bit of code?\n\nWhat speed up around recovery are you seeing N? Should we change the default timeout too as Suraj does above?","created":"2012-08-06T16:24:34.931+0000"},{"body":"In addCall():\n{code}\n+ calls.put(call.id, call);\n+ notify();\n{code}\nDo we need a 'synchronized (call)' block for the above notification ?","created":"2012-08-06T16:31:18.461+0000"},{"body":"bq. Fix formatting before commit N.\nOk.\n\nBefore committing, I would be interested by a feedback from Suraj. There are just a few lines of code, so rebasing won't be complicated if he needs some time to test it.\n\nbq. I think I understand the notifying that is going on on the end of the addCall method. They line up w/ waits on Call and waits on the calls data member?\nIt's mainly playing with the synchronized: Connection#addCall & Connection#setupIOstreams are both synchronized, so, on an exception during setupIOstreams, either:\n- you were waiting just before setupIOstreams, and in this case you've been clean up during setupIOstreams exception management\n- you were waiting before the addCall, and in this case you won't be added to the calls list.\n- in both case when you enter yourself in setupIOstreams you are filtered by the test on shouldCloseConnection \n\nbq. Would it be hard making a test of this bit of code?\nIt's a difficult question, because there are both the behavior of this jira and both the generic behavior to be tested.\n1) Just for this jira, when I tested it I added a sleep to simulate a connection timeout. I will provide soon a (small) set of utility functions to better simulated this, with real timeouts. This type of test (more in the category of regression tests than unit tests) could be added to the integration tests may be. I had various issues during the tests, it was more difficult than expected.\n2) Testing the HBaseClient itself would be useful, but the interesting path is the multithreaded one.\n\nbq. What speed up around recovery are you seeing N? Should we change the default timeout too as Suraj does above?\n\nFor the speed up, it's arbitrary, as it depends on the number of RS. on my tests, 20% of the calls were serialized. I.e. with an operation on 20 rs, the fix makes it 3 times faster. But it seems that Suraj had a much worse serialization, on a bigger cluster, so for him we could expect much better results, likely 20 times faster or better.\nAnother point is that in this fix we don't keep a list of dead rs, so we cut the connection attempts only of they are happening while another is taking place. So if he could try it would be great.\n\nFor the default timeout, I think we can cut down the connect timeout. But I think it's safer to make it to 5 seconds, so this fix remains important. I will work on this on another jira.\n\n\nbq. Do we need a 'synchronized (call)' block for the above notification ?\nIt's a notification on the connection itself, and addCall is synchronized, so it's ok. Then 'calls' is a 'ConcurrentSkipListMap' so we can access it concurrently.","created":"2012-08-06T17:51:17.758+0000"},{"body":"bq. Before committing, I would be interested by a feedback from Suraj. There are just a few lines of code, so rebasing won't be complicated if he needs some time to test it.\n\nSounds good.\n\nbq. For the default timeout, I think we can cut down the connect timeout. But I think it's safer to make it to 5 seconds, so this fix remains important. I will work on this on another jira.\n\nSounds good too.\n\nOn test, even if its utility to simulate so you can prove your fix, that'd be great.\n\nI like the numbers you are quoting above.\n\n","created":"2012-08-06T20:26:20.520+0000"},{"body":"Not sure I wrapped my head around the issue completely. But from the discussion here and looking at the patch it looks right.\nThis should be in 0.94 as well.","created":"2012-08-06T22:22:10.053+0000"},{"body":"New version, with the results of the test implementation on the mini cluster:\n- I haven't been able to simulate a connect timeout. So I need to instrument the code a way or another to add a sleep if we want the push the tests in the test suite. I can do that by subclassing HBaseClient for example. It won't be very clean. Is there another solution?\n- When I implemented Suraj scenario, I reproduced exactly his issue (while previously, in my tests the time lost was only a fraction of the number of threads). And my fix was not working on this specific scenario.\n- So I finally added a dead servers list. The expiry time is configurable (hbase.ipc.client.recheckServersTimeout). It's not totally without impact: if someone was actually relying on the time spent on retries (for example if a server comes down & up again), he will be disappointed... I set a small default value for this reason (2 seconds, ie. less than most retries wait time). Increasing it should be done only if there are good reasons for this.\n","created":"2012-08-08T15:40:06.725+0000"},{"body":"For the tests, I tried 2 options:\n- Implementing a specific SocketFactory that would return configurable sockets. Too fragile & too complicated\n- Adding a hook to get a specific HBaseClient implementation, that would add the sleep when necessary. That's in the v6.\n\nOn the long term, I think the best place to add test hooks is NetUtils, but this class is made of static methods and is in the hadoop.net package, not HBase.\n\n\nI'm not a big fan of adding these new tests to the main build, but it can now be done.","created":"2012-08-09T12:32:01.867+0000"},{"body":"We already have a class named DeadServers here: ./hbase-server/src/main/java/org/apache/hadoop/hbase/master/DeadServer.java Maybe name this something else?\n\nDo you think the deadservers will hold reference to the original backing Map? If so, if we keep adding, are we adding to backing Map or to the tailMap? I'm talking about here:\n\n+ deadServers = deadServers.tailMap(now);\n\nI'm afraid we will be retaining reference to backing Map and it will grow without bound? Is that possible?\n\nSo, if default timeout for items in list is two seconds and socket timeout is 20 seconds, will items timeout in the dead list before we ever make use of it?\n\nYou should fix the formatting adding spaces around '+' etc., so it has formatting like the rest of the code in this class... or is that how its done elsewhere in this class? If so, fine.\n\nYour addition where we can override HBaseClient looks generally useful.\n\nUse EnvironmentEdge instead of System.currentMillis...\n\n+ final long start = System.currentTimeMillis();\n\nTest looks good.\n\nIs the HBaseRecoveryTestingUtility generally useful do you think? Maybe we don't add test but add this? (This class needs class comment explaining what goodies it has).\n\nGood stuff.\n","created":"2012-08-09T14:50:51.075+0000"},{"body":"bq. We already have a class named DeadServers here: ./hbase-server/src/main/java/org/apache/hadoop/hbase/master/DeadServer.java Maybe name this something else?\nYou're right. Ok.\n\nbq. I'm afraid we will be retaining reference to backing Map and it will grow without bound? Is that possible?\nIt's even probable :-). I need to fix this.\n\nbq. So, if default timeout for items in list is two seconds and socket timeout is 20 seconds, will items timeout in the dead list before we ever make use of it?\n\nNo, because it's added after the timeout. So:\nt0: connect starts, will wait 20s\nt <19: all threads get in the queue because of the synchronized\nt20: timeout; added to the dead list, synchronized lock freed\nt21: all threads gets in and gets out because of the dead server list\nt22: if a new thread comes in with the wrong server it will wait again\n\nbq. Use EnvironmentEdge instead of System.currentMillis...\nEven in an unit test? I didn't know. Ok.\n\nbq. You should fix the formatting adding spaces around '+' etc., so it has formatting like the rest of the code in this class... or is that how its done elsewhere in this class? If so, fine.\nI will recheck. There are some '+' like this in HBaseClient already :-). But I will add spaces to mines.\n\n\nbq. Is the HBaseRecoveryTestingUtility generally useful do you think? Maybe we don't add test but add this? (This class needs class comment explaining what goodies it has).\nYes, I think so. It helps to write more readable failure tests. It still has rough edges, and may be bugs. Also, some functions should be in HBaseTestingUtility. It will be ready soon, but it's better to wait a little before committing...\n\nThanks for the feedback.","created":"2012-08-09T15:18:08.718+0000"},{"body":"{code}\n+ deadServers.put(expiry, address.toString());\n{code}\nDo we need a MultiMap as the backing store for deadServers (possibly two servers having the same expiration time) ?\n\nThe hbase.ipc.client.recheckServersTimeout is for dead servers. Release note is needed so that people know what to look for.\n","created":"2012-08-09T18:46:19.256+0000"},{"body":"bq. Do we need a MultiMap as the backing store for deadServers (possibly two servers having the same expiration time) ?\nYou're right. I will need to fix this as well.\n\nbq. The hbase.ipc.client.recheckServersTimeout is for dead servers. Release note is needed so that people know what to look for.\nWill do. But the default should be good for most people imho. I will put a warning in the release notes (there could be an issue if a node disappears and comes back if the setting is too high: the client will have to wait for the end of the recheck before seeing the node coming back).","created":"2012-08-09T18:58:32.609+0000"},{"body":"@N:\nThanks for the quick response. Appreciate it.\nnit: please add javadoc for remoteId:\n{code}\n+ * Creates a connection. Can be overridden by a subclass for testing.\n+ */\n+ protected Connection createConnection(ConnectionId remoteId) throws IOException {\n{code}\n{code}\n+ IOException e = new DeadServerIOException(\n+ \"This server is is the dead server list: \"+server);\n{code}\n'is is' -> 'is in'\nDeadServerIOException extends IOException, I think 'IO' doesn't have to appear in the name of exception.\n\nWould be nice if the next patch is put up on review board.","created":"2012-08-09T19:09:57.631+0000"},{"body":"Thanks nkeywal - this looks great. I will try to rebase it back to my version and see if I can arrange a re-test. However, please don't wait for me for your commit as it may take a few days before I can arrange to run this in my environment (perhaps end of next week ... at best). \n\n","created":"2012-08-10T00:27:58.785+0000"},{"body":"@Suraj\nOk, we will see how it takes to finish the review process as well. But I've reproduced your scenario I think.\n\n@All\nv8 with all comments taken into account\nTestHBaseClient to be in the final commit\nTest_HBASE_6364 / HBaseRecoveryTestingUtility not for the final commit\n\nI'm going to create a rb for this as well.","created":"2012-08-10T11:56:12.170+0000"},{"body":"https://reviews.apache.org/r/6521/","created":"2012-08-10T12:03:23.646+0000"},{"body":"v10 with all comments taken into account. Unfortunately, TestReplication failed 3 times in a row. I need to look at this. There is no clear reason, this last patch is very similar to the previous one. Will look at that later...","created":"2012-08-15T10:24:08.824+0000"},{"body":"Ok, tried 3 more times, they're all successful. So I think it's ok. If there is nothing scary in the hadoop-qa-build, I will commit v10 within 2 or 3 days.","created":"2012-08-15T13:30:50.444+0000"},{"body":"@nkeywal Where is v10? I see v9 attached here. Is it what is up in RB?","created":"2012-08-17T23:24:54.199+0000"},{"body":"@stack. Sorry I missed your message. Yes, the last one is v9, and it's the one in RB as well.","created":"2012-08-18T12:13:19.932+0000"},{"body":"I added some comment up on RB but this is close to committable IMO","created":"2012-08-18T22:03:38.346+0000"},{"body":"Version that will be committed if the local tests (in progress) are ok.","created":"2012-08-21T10:46:11.923+0000"},{"body":"Committed revision 1375473.\n\nAfter another fishy and not reproduced error on org.apache.hadoop.hbase.TestMultiVersions","created":"2012-08-21T11:15:14.443+0000"},{"body":"@Nicholas: What's your feeling here... Should this go into 0.94? The patch seems contained to me.","created":"2012-08-21T17:00:19.785+0000"},{"body":"@lars\nI think it's reasonably safe (well I wrote it :-) ). On the other hand to have the issue it solves you need to be a little bit unlucky as well. As the consequences are quite bad when you're unlucky I think it's better to do it. And if there is an issue you can deactivate it (by setting hbase.ipc.client.failed.servers.expiry to zero), or lower the expiry time to something as 100ms. If it breaks something I will fix it for both versions.","created":"2012-08-21T17:17:22.604+0000"},{"body":"+1 on backport so we get benefit of the MTTR work sooner.\n\n@Nkeywal Should we lower the default ipc.socket.timeout too? I don't see that in your patch?","created":"2012-08-21T18:34:52.801+0000"},{"body":"bq. +1 on backport so we get benefit of the MTTR work sooner.\nI can do it if you like.\n\nbq. @Nkeywal Should we lower the default ipc.socket.timeout too? I don't see that in your patch?\nDoable, there will be no side effect, it's used only in HBaseClient. How do you want it:\n- safe: 10 seconds\n- reasonably aggressive: 5 seconds\n- instructive: 1 second\n\n20 seconds is a common timeout value (used in the OS if I remember well). Not subject to GC effects. I don't know for large clusters under large failure conditions (with switches). Especially, in this case, it will be all the clients connecting to a single server (the one with meta), so likely going through a single network link at the end. I would vote 10s; but 5s is doable. 1s could work, who knows?\n","created":"2012-08-21T19:25:13.930+0000"},{"body":"20 is the current default? If so, we can leave it. Should we add a note to perf section of refguide on lowering connection timeout?\n\nThat'd be grand if you'd commit to 0.94.","created":"2012-08-21T21:48:47.937+0000"},{"body":"For connect timeout, yes it's 20s. We could add a note as well, I will create a jira for this. And commit in 0.94.","created":"2012-08-21T21:53:36.429+0000"},{"body":"6364.94.v2.nolargetest.patch contains the patch for 0.94. My own test depends on a class that does not exist in 0.94; so I didn't test it on 0.95 \n\nUnit tests ok, except testClientPoolRoundRobin(org.apache.hadoop.hbase.client.TestFromClientSide): The number of versions of '[B@4c9cde9a:[B@4eda77c1 did not match 4 expected:<4> but was:<3>\n\nfailed once, second try ok. Committed.","created":"2012-08-22T12:35:00.438+0000"},{"body":"Addendum looks good to me.","created":"2012-08-22T16:35:21.125+0000"},{"body":"No failure on the unit tests small & medium with the security profile.\nCommitted revision 1376136.","created":"2012-08-22T16:49:44.123+0000"},{"body":"+1 on addendum","created":"2012-08-22T18:07:50.238+0000"},{"body":"This time it's the usual guilty so I think we're ok.","created":"2012-08-22T18:08:20.623+0000"},{"body":"@Nicholas: Thank you very much for backporting the patch!","created":"2012-08-23T03:23:15.257+0000"},{"body":"This was committed, marking fixed.","created":"2012-08-26T01:17:13.956+0000"},{"body":"Fix up after bulk move overwrote some 0.94.2 fix versions w/ 0.95.0 (Noticed by Lars Hofhansl)","created":"2013-04-07T04:38:04.425+0000"},{"body":"[~nkeywal] What is supposed to happen when the FailedServerException is thrown?\n\nThe below is hard to read -- it is out of our loadtesttool used alot in hbase-it.\n\nA loadtesttool thread is failing with a FailedServerException... \"This server is in the failed servers list'. Reading the above, I thought we were supposed to pause and retry but if you notice, we are doing connection setup using withoutRetries... because we are inside a Process.\n\nAny input would help. Sorry for being short. Have to run (I'm trying to figure failing hbase-it tests up on ec2 and on internal jenkins). Thanks.\n\n{code}\n2013-05-27 17:08:53,193 ERROR [HBaseWriterThread_3] util.MultiThreadedWriter(191): Failed to insert: 48349; region information: cached: region=IntegrationTestDataIngestWithChaosMonkey,8cccccc4,1369699538187.6b2be2c004e633f8eed7de8aff6f7cfd., hostname=a1007.halxg.cloudera.com,53752,1369699532487, seqNum=214684; cache is up to date; errors: Error from [a1007.halxg.cloudera.com:53752] for [979402d0d20fb0f8ded281a8b8687ab9-48349]java.io.IOException: Call to a1007.halxg.cloudera.com/10.20.184.107:53752 failed on local exception: org.apache.hadoop.hbase.ipc.RpcClient$FailedServerException: This server is in the failed servers list: a1007.halxg.cloudera.com/10.20.184.107:53752\n\tat org.apache.hadoop.hbase.ipc.RpcClient.wrapException(RpcClient.java:1368)er$HConnectionImplementation$Process$1.call(HConnectionManager.java:2463)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\nCaused by: org.apache.hadoop.hbase.ipc.RpcClient$FailedServerException: This server is in the failed servers list: a1007.halxg.cloudera.com/10.20.184.107:53752\n\tat org.apache.hadoop.hbase.ipc.RpcClient$Connection.setupIOstreams(RpcClient.java:798)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.getConnection(RpcClient.java:1422)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.call(RpcClient.java:1314)\n\t... 15 more\n\n2013-05-27 17:08:53,193 ERROR [HBaseWriterThread_0] util.MultiThreadedWriter(191): Failed to insert: 48366; region information: cached: region=IntegrationTestDataIngestWithChaosMonkey,8cccccc4,1369699538187.6b2be2c004e633f8eed7de8aff6f7cfd., hostname=a1007.halxg.cloudera.com,53752,1369699532487, seqNum=214684; cache is up to date; errors: Error from [a1007.halxg.cloudera.com:53752] for [92873a55c54f98db38508ba065852cc5-48366]java.io.IOException: Call to a1007.halxg.cloudera.com/10.20.184.107:53752 failed on local exception: org.apache.hadoop.hbase.ipc.RpcClient$FailedServerException: This server is in the failed servers list: a1007.halxg.cloudera.com/10.20.184.107:53752\n\tat org.apache.hadoop.hbase.ipc.RpcClient.wrapException(RpcClient.java:1368)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.call(RpcClient.java:1340)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.callBlockingMethod(RpcClient.java:1540)\n\tat org.apache.hadoop.hbase.ipc.RpcClient$BlockingRpcChannelImplementation.callBlockingMethod(RpcClient.java:1597)\n\tat org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$BlockingStub.multi(ClientProtos.java:21403)\n\tat org.apache.hadoop.hbase.client.MultiServerCallable.call(MultiServerCallable.java:102)\n\tat org.apache.hadoop.hbase.client.MultiServerCallable.call(MultiServerCallable.java:43)\n\tat org.apache.hadoop.hbase.client.ServerCallable.withoutRetries(ServerCallable.java:250)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation$7.call(HConnectionManager.java:1993)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation$7.call(HConnectionManager.java:1988)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation$Process$1.call(HConnectionManager.java:2473)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation$Process$1.call(HConnectionManager.java:2463)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\nCaused by: org.apache.hadoop.hbase.ipc.RpcClient$FailedServerException: This server is in the failed servers list: a1007.halxg.cloudera.com/10.20.184.107:53752\n\tat org.apache.hadoop.hbase.ipc.RpcClient$Connection.setupIOstreams(RpcClient.java:798)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.getConnection(RpcClient.java:1422)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.call(RpcClient.java:1314)\n\t... 15 more\n\n\n{code}","created":"2013-05-30T05:01:06.024+0000"},{"body":"@nkeywal nevermind. The issue I am seeing is not particular to this issue. Please ignore.","created":"2013-05-30T09:33:42.481+0000"}],"conversations":[{"body":"When a server host with a Region Server holding the .META. table is powered down on a live cluster, while the HBase cluster itself detects and reassigns the .META. table, connected HBase Client's take an excessively long time to detect this and re-discover the reassigned .META. \n\nWorkaround: Decrease the ipc.socket.timeout on HBase Client side to a low value (default is 20s leading to 35 minute recovery time; we were able to get acceptable results with 100ms getting a 3 minute recovery) \n\nThis was found during some hardware failure testing scenarios. \n\nTest Case:\n1) Apply load via client app on HBase cluster for several minutes\n2) Power down the region server holding the .META. server (i.e. power off ... and keep it off)\n3) Measure how long it takes for cluster to reassign META table and for client threads to re-lookup and re-orient to the lesser cluster (minus the RS and DN on that host).\n\nObservation:\n1) Client threads spike up to maxThreads size ... and take over 35 mins to recover (i.e. for the thread count to go back to normal) - no client calls are serviced - they just back up on a synchronized method (see #2 below)\n\n2) All the client app threads queue up behind the oahh.ipc.HBaseClient#setupIOStreams method http://tinyurl.com/7js53dj\n\nAfter taking several thread dumps we found that the thread within this synchronized method was blocked on NetUtils.connect(this.socket, remoteId.getAddress(), getSocketTimeout(conf));\n\nThe client thread that gets the synchronized lock would try to connect to the dead RS (till socket times out after 20s), retries, and then the next thread gets in and so forth in a serial manner.\n\nWorkaround:\n-------------------\nDefault ipc.socket.timeout is set to 20s. We dropped this to a low number (1000 ms, 100 ms, etc) on the client side hbase-site.xml. With this setting, the client threads recovered in a couple of minutes by failing fast and re-discovering the .META. table on a reassigned RS.\n\nAssumption: This ipc.socket.timeout is only ever used during the initial \"HConnection\" setup via the NetUtils.connect and should only ever be used when connectivity to a region server is lost and needs to be re-established. i.e it does not affect the normal \"RPC\" actiivity as this is just the connect timeout.\nDuring RS GC periods, any _new_ clients trying to connect will fail and will require .META. table re-lookups.\n\nThis above timeout workaround is only for the HBase client side.","from":"reporter","subject":"Powering down the server host holding the .META. table causes HBase Client to take excessively long to recover and connect to reassigned .META. table"},{"body":"Also wanted to capture NKeywal's comments on this from the mailing list: \n\nNKeywal said: \nWhat you're describing -the 35 minutes recovery time- seems to match the code. And it's a bug (still there on trunk). Could you please create a jira for it? If you have the logs it even better.\n\nLowering the ipc.socket.timeout seems to be an acceptable partial workaround. Setting it to 10s seems ok to me. Lower than this... I don't know.\n","from":"developer"},{"body":"That's a possible cause:\n\nin HBaseClient#getConnection()\n{noformat}\n do {\n synchronized (connections) {\n connection = connections.get(remoteId);\n if (connection == null) {\n connection = new Connection(remoteId);\n connections.put(remoteId, connection);\n }\n }\n } while (!connection.addCall(call));\n \n connection.setupIOstreams();\n return connection;\n }\n{noformat}\nConnection#addCall and Connection#setupIOstreams are synchronized. if #setupIOstreams fails, it marks the connection as dead, remove it from the connections list, and throws an exception. In #addCall it returns false if the connection is marked as dead. So\ncase 1 -> Sometimes, we may add the call to a connection that will be marked as dead:\n Thread 1: create the connection, add it to the connections list, call addCall\n Thread 2: get the connection, add it to the calls list \n Thread 1: get into setupIOstreams, fails, and mark the connection as dead, throws an exception, done\n Thread 2: get into setupIOstreams, see that the connection is dead, done. The call has been added to a dead connection\n\ncase 2 -> If we have a lot of threads on a dying connection, we will have:\n Thread 1: goes until setupIOstreams\n All other threads: get the connection from the list, wait on the synchronized addCall \n Thread 1: exit from setupIOstreams with an exception after 20 seconds (socket timeout)\n All other threads: call addCall, as the connection is dead, reloop\n One of these threads will create a connection\n One of these thread will win the race on addCall\n One of them will win the race on setupIOstreams\n Most of them should be waiting on addCall so reloop\n So back to the case 1 or 2.\n\n\nWe would have the same behavior on pure hadoop client (ipc.Client), as the implementation is similar, at least on 1.0.3.\n\n\nSuraj, does this match your analysis? How many region servers and regions did you have during your test? What's the client doing?","from":"developer"},{"body":"stack trace snippet from test.","from":"developer"},{"body":"nkeywal, my test cluster had 5 regionservers with an average of 35 regions per server. The client is doing Get calls from all threads - each thread creates a separate HTable but with the same HBaseConfiguration (so, sharing the same underlying HConnection.)\n\nCase 2 that you mention above is what I think I saw. Several threads blocked on #addCall with a lock owned by the setupIOStreams call (on NetUtils.connect()) which was timing out after 20s. \n\nI finally managed to locate the stack traces from the test ... I'm attaching a snippet from that with some BLOCKED threads and the setupIOStreams call holding the lock. \n\nI believe the same behaviour would occur on the hadoop client as well ... but perhaps because of multiple datanodes, the problem is not as acute? In the HBaseClient case, all threads repeatedly reaches out for META to rediscover the relocated regions. \n(Attached the thread dump snippet)\n","from":"developer"},{"body":"Thanks Suraj. So it seems that we have the root cause. You're right, on ipc.Client it's less likely to occur. Only the NN could have the issue I think, and as the hot failover is not yet deployed in production, it has not been an issue so far. I will ping Todd on this to see what he thinks.\n\nThe only strange thing is that when tested on the 0.96 + minicluster, I have the issue but many threads seem to be able to 'escape', i.e. manage to do the 'addCall' between two blocking calls to setupIOstreams. But it could be a pure multithreading hazard. And I haven't been able to find a better cause than this.\n\nI will propose a patch, likely end of next week. Suraj, If you have some time available to validate it on your env, it would be great.","from":"developer"},{"body":"v1. There are some other ways of doing this, like adding a list of dead servers and a timeout, but it does not pass the unit tests, with numerous failures. I haven't sorted out if there is a common root cause...","from":"developer"},{"body":"likely to be unrelated, worked twice locally. Let's retry.","from":"developer"},{"body":"https://builds.apache.org/job/PreCommit-HBASE-Build/2514/console got aborted.","from":"developer"},{"body":"Patch from N.","from":"developer"},{"body":"Is this alleviated (at least somewhat) by HBASE-6326?","from":"developer"},{"body":"Looking at the issue, no it isn't. NM me. :)","from":"developer"},{"body":"Reanalyzing the fix, there is an issue with the v1: we could have a call added to a dying connection, and this call won't get cleaned up. This is was not possible previously. Will write a v2.","from":"developer"},{"body":"v2, fixes the problem mentioned above, works locally.","from":"developer"},{"body":"v3 just changes a comment, not the code itself.","from":"developer"},{"body":"errors are unrelated imho. Retrying to see if the third executions says something different.","from":"developer"},{"body":"Good one lads.\n\nFix formatting before commit N. Make it same as surrounding code... Add spacings around brackets -- the 'else' -- and the '+' in String concatenations.\n\nI think I understand the notifying that is going on on the end of the addCall method. They line up w/ waits on Call and waits on the calls data member?\n\nWould it be hard making a test of this bit of code?\n\nWhat speed up around recovery are you seeing N? Should we change the default timeout too as Suraj does above?","from":"developer"},{"body":"In addCall():\n{code}\n+ calls.put(call.id, call);\n+ notify();\n{code}\nDo we need a 'synchronized (call)' block for the above notification ?","from":"developer"},{"body":"bq. Fix formatting before commit N.\nOk.\n\nBefore committing, I would be interested by a feedback from Suraj. There are just a few lines of code, so rebasing won't be complicated if he needs some time to test it.\n\nbq. I think I understand the notifying that is going on on the end of the addCall method. They line up w/ waits on Call and waits on the calls data member?\nIt's mainly playing with the synchronized: Connection#addCall & Connection#setupIOstreams are both synchronized, so, on an exception during setupIOstreams, either:\n- you were waiting just before setupIOstreams, and in this case you've been clean up during setupIOstreams exception management\n- you were waiting before the addCall, and in this case you won't be added to the calls list.\n- in both case when you enter yourself in setupIOstreams you are filtered by the test on shouldCloseConnection \n\nbq. Would it be hard making a test of this bit of code?\nIt's a difficult question, because there are both the behavior of this jira and both the generic behavior to be tested.\n1) Just for this jira, when I tested it I added a sleep to simulate a connection timeout. I will provide soon a (small) set of utility functions to better simulated this, with real timeouts. This type of test (more in the category of regression tests than unit tests) could be added to the integration tests may be. I had various issues during the tests, it was more difficult than expected.\n2) Testing the HBaseClient itself would be useful, but the interesting path is the multithreaded one.\n\nbq. What speed up around recovery are you seeing N? Should we change the default timeout too as Suraj does above?\n\nFor the speed up, it's arbitrary, as it depends on the number of RS. on my tests, 20% of the calls were serialized. I.e. with an operation on 20 rs, the fix makes it 3 times faster. But it seems that Suraj had a much worse serialization, on a bigger cluster, so for him we could expect much better results, likely 20 times faster or better.\nAnother point is that in this fix we don't keep a list of dead rs, so we cut the connection attempts only of they are happening while another is taking place. So if he could try it would be great.\n\nFor the default timeout, I think we can cut down the connect timeout. But I think it's safer to make it to 5 seconds, so this fix remains important. I will work on this on another jira.\n\n\nbq. Do we need a 'synchronized (call)' block for the above notification ?\nIt's a notification on the connection itself, and addCall is synchronized, so it's ok. Then 'calls' is a 'ConcurrentSkipListMap' so we can access it concurrently.","from":"developer"},{"body":"bq. Before committing, I would be interested by a feedback from Suraj. There are just a few lines of code, so rebasing won't be complicated if he needs some time to test it.\n\nSounds good.\n\nbq. For the default timeout, I think we can cut down the connect timeout. But I think it's safer to make it to 5 seconds, so this fix remains important. I will work on this on another jira.\n\nSounds good too.\n\nOn test, even if its utility to simulate so you can prove your fix, that'd be great.\n\nI like the numbers you are quoting above.\n\n","from":"developer"},{"body":"Not sure I wrapped my head around the issue completely. But from the discussion here and looking at the patch it looks right.\nThis should be in 0.94 as well.","from":"developer"},{"body":"New version, with the results of the test implementation on the mini cluster:\n- I haven't been able to simulate a connect timeout. So I need to instrument the code a way or another to add a sleep if we want the push the tests in the test suite. I can do that by subclassing HBaseClient for example. It won't be very clean. Is there another solution?\n- When I implemented Suraj scenario, I reproduced exactly his issue (while previously, in my tests the time lost was only a fraction of the number of threads). And my fix was not working on this specific scenario.\n- So I finally added a dead servers list. The expiry time is configurable (hbase.ipc.client.recheckServersTimeout). It's not totally without impact: if someone was actually relying on the time spent on retries (for example if a server comes down & up again), he will be disappointed... I set a small default value for this reason (2 seconds, ie. less than most retries wait time). Increasing it should be done only if there are good reasons for this.\n","from":"developer"},{"body":"For the tests, I tried 2 options:\n- Implementing a specific SocketFactory that would return configurable sockets. Too fragile & too complicated\n- Adding a hook to get a specific HBaseClient implementation, that would add the sleep when necessary. That's in the v6.\n\nOn the long term, I think the best place to add test hooks is NetUtils, but this class is made of static methods and is in the hadoop.net package, not HBase.\n\n\nI'm not a big fan of adding these new tests to the main build, but it can now be done.","from":"developer"},{"body":"We already have a class named DeadServers here: ./hbase-server/src/main/java/org/apache/hadoop/hbase/master/DeadServer.java Maybe name this something else?\n\nDo you think the deadservers will hold reference to the original backing Map? If so, if we keep adding, are we adding to backing Map or to the tailMap? I'm talking about here:\n\n+ deadServers = deadServers.tailMap(now);\n\nI'm afraid we will be retaining reference to backing Map and it will grow without bound? Is that possible?\n\nSo, if default timeout for items in list is two seconds and socket timeout is 20 seconds, will items timeout in the dead list before we ever make use of it?\n\nYou should fix the formatting adding spaces around '+' etc., so it has formatting like the rest of the code in this class... or is that how its done elsewhere in this class? If so, fine.\n\nYour addition where we can override HBaseClient looks generally useful.\n\nUse EnvironmentEdge instead of System.currentMillis...\n\n+ final long start = System.currentTimeMillis();\n\nTest looks good.\n\nIs the HBaseRecoveryTestingUtility generally useful do you think? Maybe we don't add test but add this? (This class needs class comment explaining what goodies it has).\n\nGood stuff.\n","from":"developer"},{"body":"bq. We already have a class named DeadServers here: ./hbase-server/src/main/java/org/apache/hadoop/hbase/master/DeadServer.java Maybe name this something else?\nYou're right. Ok.\n\nbq. I'm afraid we will be retaining reference to backing Map and it will grow without bound? Is that possible?\nIt's even probable :-). I need to fix this.\n\nbq. So, if default timeout for items in list is two seconds and socket timeout is 20 seconds, will items timeout in the dead list before we ever make use of it?\n\nNo, because it's added after the timeout. So:\nt0: connect starts, will wait 20s\nt <19: all threads get in the queue because of the synchronized\nt20: timeout; added to the dead list, synchronized lock freed\nt21: all threads gets in and gets out because of the dead server list\nt22: if a new thread comes in with the wrong server it will wait again\n\nbq. Use EnvironmentEdge instead of System.currentMillis...\nEven in an unit test? I didn't know. Ok.\n\nbq. You should fix the formatting adding spaces around '+' etc., so it has formatting like the rest of the code in this class... or is that how its done elsewhere in this class? If so, fine.\nI will recheck. There are some '+' like this in HBaseClient already :-). But I will add spaces to mines.\n\n\nbq. Is the HBaseRecoveryTestingUtility generally useful do you think? Maybe we don't add test but add this? (This class needs class comment explaining what goodies it has).\nYes, I think so. It helps to write more readable failure tests. It still has rough edges, and may be bugs. Also, some functions should be in HBaseTestingUtility. It will be ready soon, but it's better to wait a little before committing...\n\nThanks for the feedback.","from":"developer"},{"body":"{code}\n+ deadServers.put(expiry, address.toString());\n{code}\nDo we need a MultiMap as the backing store for deadServers (possibly two servers having the same expiration time) ?\n\nThe hbase.ipc.client.recheckServersTimeout is for dead servers. Release note is needed so that people know what to look for.\n","from":"developer"},{"body":"bq. Do we need a MultiMap as the backing store for deadServers (possibly two servers having the same expiration time) ?\nYou're right. I will need to fix this as well.\n\nbq. The hbase.ipc.client.recheckServersTimeout is for dead servers. Release note is needed so that people know what to look for.\nWill do. But the default should be good for most people imho. I will put a warning in the release notes (there could be an issue if a node disappears and comes back if the setting is too high: the client will have to wait for the end of the recheck before seeing the node coming back).","from":"developer"},{"body":"@N:\nThanks for the quick response. Appreciate it.\nnit: please add javadoc for remoteId:\n{code}\n+ * Creates a connection. Can be overridden by a subclass for testing.\n+ */\n+ protected Connection createConnection(ConnectionId remoteId) throws IOException {\n{code}\n{code}\n+ IOException e = new DeadServerIOException(\n+ \"This server is is the dead server list: \"+server);\n{code}\n'is is' -> 'is in'\nDeadServerIOException extends IOException, I think 'IO' doesn't have to appear in the name of exception.\n\nWould be nice if the next patch is put up on review board.","from":"developer"},{"body":"Thanks nkeywal - this looks great. I will try to rebase it back to my version and see if I can arrange a re-test. However, please don't wait for me for your commit as it may take a few days before I can arrange to run this in my environment (perhaps end of next week ... at best). \n\n","from":"developer"},{"body":"@Suraj\nOk, we will see how it takes to finish the review process as well. But I've reproduced your scenario I think.\n\n@All\nv8 with all comments taken into account\nTestHBaseClient to be in the final commit\nTest_HBASE_6364 / HBaseRecoveryTestingUtility not for the final commit\n\nI'm going to create a rb for this as well.","from":"developer"},{"body":"https://reviews.apache.org/r/6521/","from":"developer"},{"body":"v10 with all comments taken into account. Unfortunately, TestReplication failed 3 times in a row. I need to look at this. There is no clear reason, this last patch is very similar to the previous one. Will look at that later...","from":"developer"},{"body":"Ok, tried 3 more times, they're all successful. So I think it's ok. If there is nothing scary in the hadoop-qa-build, I will commit v10 within 2 or 3 days.","from":"developer"},{"body":"@nkeywal Where is v10? I see v9 attached here. Is it what is up in RB?","from":"developer"},{"body":"@stack. Sorry I missed your message. Yes, the last one is v9, and it's the one in RB as well.","from":"developer"},{"body":"I added some comment up on RB but this is close to committable IMO","from":"developer"},{"body":"Version that will be committed if the local tests (in progress) are ok.","from":"developer"},{"body":"Committed revision 1375473.\n\nAfter another fishy and not reproduced error on org.apache.hadoop.hbase.TestMultiVersions","from":"developer"},{"body":"@Nicholas: What's your feeling here... Should this go into 0.94? The patch seems contained to me.","from":"developer"},{"body":"@lars\nI think it's reasonably safe (well I wrote it :-) ). On the other hand to have the issue it solves you need to be a little bit unlucky as well. As the consequences are quite bad when you're unlucky I think it's better to do it. And if there is an issue you can deactivate it (by setting hbase.ipc.client.failed.servers.expiry to zero), or lower the expiry time to something as 100ms. If it breaks something I will fix it for both versions.","from":"developer"},{"body":"+1 on backport so we get benefit of the MTTR work sooner.\n\n@Nkeywal Should we lower the default ipc.socket.timeout too? I don't see that in your patch?","from":"developer"},{"body":"bq. +1 on backport so we get benefit of the MTTR work sooner.\nI can do it if you like.\n\nbq. @Nkeywal Should we lower the default ipc.socket.timeout too? I don't see that in your patch?\nDoable, there will be no side effect, it's used only in HBaseClient. How do you want it:\n- safe: 10 seconds\n- reasonably aggressive: 5 seconds\n- instructive: 1 second\n\n20 seconds is a common timeout value (used in the OS if I remember well). Not subject to GC effects. I don't know for large clusters under large failure conditions (with switches). Especially, in this case, it will be all the clients connecting to a single server (the one with meta), so likely going through a single network link at the end. I would vote 10s; but 5s is doable. 1s could work, who knows?\n","from":"developer"},{"body":"20 is the current default? If so, we can leave it. Should we add a note to perf section of refguide on lowering connection timeout?\n\nThat'd be grand if you'd commit to 0.94.","from":"developer"},{"body":"For connect timeout, yes it's 20s. We could add a note as well, I will create a jira for this. And commit in 0.94.","from":"developer"},{"body":"6364.94.v2.nolargetest.patch contains the patch for 0.94. My own test depends on a class that does not exist in 0.94; so I didn't test it on 0.95 \n\nUnit tests ok, except testClientPoolRoundRobin(org.apache.hadoop.hbase.client.TestFromClientSide): The number of versions of '[B@4c9cde9a:[B@4eda77c1 did not match 4 expected:<4> but was:<3>\n\nfailed once, second try ok. Committed.","from":"developer"},{"body":"Addendum looks good to me.","from":"developer"},{"body":"No failure on the unit tests small & medium with the security profile.\nCommitted revision 1376136.","from":"developer"},{"body":"+1 on addendum","from":"developer"},{"body":"This time it's the usual guilty so I think we're ok.","from":"developer"},{"body":"@Nicholas: Thank you very much for backporting the patch!","from":"developer"},{"body":"This was committed, marking fixed.","from":"developer"},{"body":"Fix up after bulk move overwrote some 0.94.2 fix versions w/ 0.95.0 (Noticed by Lars Hofhansl)","from":"developer"},{"body":"[~nkeywal] What is supposed to happen when the FailedServerException is thrown?\n\nThe below is hard to read -- it is out of our loadtesttool used alot in hbase-it.\n\nA loadtesttool thread is failing with a FailedServerException... \"This server is in the failed servers list'. Reading the above, I thought we were supposed to pause and retry but if you notice, we are doing connection setup using withoutRetries... because we are inside a Process.\n\nAny input would help. Sorry for being short. Have to run (I'm trying to figure failing hbase-it tests up on ec2 and on internal jenkins). Thanks.\n\n{code}\n2013-05-27 17:08:53,193 ERROR [HBaseWriterThread_3] util.MultiThreadedWriter(191): Failed to insert: 48349; region information: cached: region=IntegrationTestDataIngestWithChaosMonkey,8cccccc4,1369699538187.6b2be2c004e633f8eed7de8aff6f7cfd., hostname=a1007.halxg.cloudera.com,53752,1369699532487, seqNum=214684; cache is up to date; errors: Error from [a1007.halxg.cloudera.com:53752] for [979402d0d20fb0f8ded281a8b8687ab9-48349]java.io.IOException: Call to a1007.halxg.cloudera.com/10.20.184.107:53752 failed on local exception: org.apache.hadoop.hbase.ipc.RpcClient$FailedServerException: This server is in the failed servers list: a1007.halxg.cloudera.com/10.20.184.107:53752\n\tat org.apache.hadoop.hbase.ipc.RpcClient.wrapException(RpcClient.java:1368)er$HConnectionImplementation$Process$1.call(HConnectionManager.java:2463)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\nCaused by: org.apache.hadoop.hbase.ipc.RpcClient$FailedServerException: This server is in the failed servers list: a1007.halxg.cloudera.com/10.20.184.107:53752\n\tat org.apache.hadoop.hbase.ipc.RpcClient$Connection.setupIOstreams(RpcClient.java:798)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.getConnection(RpcClient.java:1422)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.call(RpcClient.java:1314)\n\t... 15 more\n\n2013-05-27 17:08:53,193 ERROR [HBaseWriterThread_0] util.MultiThreadedWriter(191): Failed to insert: 48366; region information: cached: region=IntegrationTestDataIngestWithChaosMonkey,8cccccc4,1369699538187.6b2be2c004e633f8eed7de8aff6f7cfd., hostname=a1007.halxg.cloudera.com,53752,1369699532487, seqNum=214684; cache is up to date; errors: Error from [a1007.halxg.cloudera.com:53752] for [92873a55c54f98db38508ba065852cc5-48366]java.io.IOException: Call to a1007.halxg.cloudera.com/10.20.184.107:53752 failed on local exception: org.apache.hadoop.hbase.ipc.RpcClient$FailedServerException: This server is in the failed servers list: a1007.halxg.cloudera.com/10.20.184.107:53752\n\tat org.apache.hadoop.hbase.ipc.RpcClient.wrapException(RpcClient.java:1368)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.call(RpcClient.java:1340)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.callBlockingMethod(RpcClient.java:1540)\n\tat org.apache.hadoop.hbase.ipc.RpcClient$BlockingRpcChannelImplementation.callBlockingMethod(RpcClient.java:1597)\n\tat org.apache.hadoop.hbase.protobuf.generated.ClientProtos$ClientService$BlockingStub.multi(ClientProtos.java:21403)\n\tat org.apache.hadoop.hbase.client.MultiServerCallable.call(MultiServerCallable.java:102)\n\tat org.apache.hadoop.hbase.client.MultiServerCallable.call(MultiServerCallable.java:43)\n\tat org.apache.hadoop.hbase.client.ServerCallable.withoutRetries(ServerCallable.java:250)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation$7.call(HConnectionManager.java:1993)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation$7.call(HConnectionManager.java:1988)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation$Process$1.call(HConnectionManager.java:2473)\n\tat org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation$Process$1.call(HConnectionManager.java:2463)\n\tat java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:303)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:138)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n\tat java.lang.Thread.run(Thread.java:662)\nCaused by: org.apache.hadoop.hbase.ipc.RpcClient$FailedServerException: This server is in the failed servers list: a1007.halxg.cloudera.com/10.20.184.107:53752\n\tat org.apache.hadoop.hbase.ipc.RpcClient$Connection.setupIOstreams(RpcClient.java:798)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.getConnection(RpcClient.java:1422)\n\tat org.apache.hadoop.hbase.ipc.RpcClient.call(RpcClient.java:1314)\n\t... 15 more\n\n\n{code}","from":"developer"},{"body":"@nkeywal nevermind. The issue I am seeing is not particular to this issue. Please ignore.","from":"developer"}],"created":"2012-07-10T17:06:49.000+0000","description":"When a server host with a Region Server holding the .META. table is powered down on a live cluster, while the HBase cluster itself detects and reassigns the .META. table, connected HBase Client's take an excessively long time to detect this and re-discover the reassigned .META. \n\nWorkaround: Decrease the ipc.socket.timeout on HBase Client side to a low value (default is 20s leading to 35 minute recovery time; we were able to get acceptable results with 100ms getting a 3 minute recovery) \n\nThis was found during some hardware failure testing scenarios. \n\nTest Case:\n1) Apply load via client app on HBase cluster for several minutes\n2) Power down the region server holding the .META. server (i.e. power off ... and keep it off)\n3) Measure how long it takes for cluster to reassign META table and for client threads to re-lookup and re-orient to the lesser cluster (minus the RS and DN on that host).\n\nObservation:\n1) Client threads spike up to maxThreads size ... and take over 35 mins to recover (i.e. for the thread count to go back to normal) - no client calls are serviced - they just back up on a synchronized method (see #2 below)\n\n2) All the client app threads queue up behind the oahh.ipc.HBaseClient#setupIOStreams method http://tinyurl.com/7js53dj\n\nAfter taking several thread dumps we found that the thread within this synchronized method was blocked on NetUtils.connect(this.socket, remoteId.getAddress(), getSocketTimeout(conf));\n\nThe client thread that gets the synchronized lock would try to connect to the dead RS (till socket times out after 20s), retries, and then the next thread gets in and so forth in a serial manner.\n\nWorkaround:\n-------------------\nDefault ipc.socket.timeout is set to 20s. We dropped this to a low number (1000 ms, 100 ms, etc) on the client side hbase-site.xml. With this setting, the client threads recovered in a couple of minutes by failing fast and re-discovering the .META. table on a reassigned RS.\n\nAssumption: This ipc.socket.timeout is only ever used during the initial \"HConnection\" setup via the NetUtils.connect and should only ever be used when connectivity to a region server is lost and needs to be re-established. i.e it does not affect the normal \"RPC\" actiivity as this is just the connect timeout.\nDuring RS GC periods, any _new_ clients trying to connect will fail and will require .META. table re-lookups.\n\nThis above timeout workaround is only for the HBase client side.","issue_id":"12598225","key":"HBASE-6364","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-08-26T01:17:13.000+0000","role":"fixed_distractor","summary":"Powering down the server host holding the .META. table causes HBase Client to take excessively long to recover and connect to reassigned .META. table"} {"case_id":"12604602","cluster":"DISTRACTOR-HBASE-6642","comments":[{"body":"This problem is because below code in disable_all.rb\n{code}\n regex = /^#{regex}$/ unless regex.is_a?(Regexp)\n list = admin.list.grep(regex)\n count = list.size\n\n{code}\nin case of * regex will become ^*$ which is valid in Ruby and admin.list.grep(regex) gives all available tables.\nBut same expression(^*$) in Java returns empty list.\n{code}\n failed = admin.disable_all(regex)\n puts \"#{count - failed.size} tables successfully disabled\"\n puts \"#{failed.size} tables not disabled due to an exception: #{failed.join ','}\" unless failed.size == 0\n{code}\nEven the number of tables to disabled is zero, output is \"#{count - failed.size} tables successfully disabled\".\nThis is applicable for enable_all and drop_all also.","created":"2012-08-27T12:02:20.656+0000"},{"body":"Small change in above comment:\nEven the number of disabled tables is zero, We are priting \"#{count} tables successfully disabled\".\nHere count is total number of available tables.\n","created":"2012-08-27T12:29:46.976+0000"},{"body":"Patch for trunk. If its ok I will upload patch for 94,92.","created":"2012-11-09T08:35:51.088+0000"},{"body":"Pls check the same with other expressions? If things are fine then its ok..","created":"2012-11-09T08:58:22.419+0000"},{"body":"@Ram\ntested with some basic regular expressions like t.*,.*,[t].*,'t\\d\\d\\d','table' and some more also. Its working fine.","created":"2012-11-09T09:34:16.559+0000"},{"body":"Should we also fix the misleading messaging? It should not print a message indicating successful execution when in fact it was not.","created":"2012-11-11T21:06:43.822+0000"},{"body":"Pushing to 0.94.4","created":"2012-11-12T18:26:57.449+0000"},{"body":"@Lars,\nbq. It should not print a message indicating successful execution when in fact it was not.\nIt wont happen. Messaging wise its clear only.","created":"2012-11-14T06:10:55.938+0000"},{"body":"Looking at this again. Are we changing the meaning?\n^{regex}$ requires the table name needs to match in its entirety.\nJust {regex} (depending on how it is called) could match when only a subset of the table name matches.\n","created":"2012-12-19T07:32:34.746+0000"},{"body":"@Lars\nbq.Looking at this again. Are we changing the meaning?\nwith patch can be yes because ruby grep and java patterns may match table names differently.\n\nIf it is not correct you can close this issue as invalid.\nThanks.","created":"2013-01-03T17:50:21.481+0000"},{"body":"I am removing this from 0.92 and 0.94.","created":"2013-02-26T05:02:22.551+0000"},{"body":"This looks like it could be useful but also dangerous. Would like others to ask for this feature (and try it out) before committing. Moving out of 0.95 for now.","created":"2013-04-04T00:17:02.634+0000"},{"body":"Faced the same \"issue\" recently. Might be good to fix that.\n\nNow, since Java and Ruby are not sending back the same result for a regex, I think we should the same API for the 2 filtering. To make sure it applies the same way on the list. Doing that will allow us to keep the current feature, but will fix it? Just a suggestion. Not sure if it's doable...","created":"2013-11-14T16:46:46.989+0000"},{"body":"patch v1 is on the same line as the previous one.\nremoves all the uses of the ruby regex and instead uses the java Pattern, which we use in the HBaseAdmin methods.","created":"2014-02-13T15:51:32.896+0000"},{"body":"Patch lgtm [~mbertozzi].\nIts better to commit to latest versions atleast.\nThanks.\n","created":"2014-02-18T14:12:31.159+0000"},{"body":"+1","created":"2014-02-20T00:04:19.518+0000"},{"body":"committed to trunk\n\n[~lhofhansl] [~stack] [~apurtell] your decision on backporting it to 94, 96, and 98. You may consider this incompatible change if people relies on globs for \"list\" commands, but the other side is the problem with enable_all, delete_all operations which are operating on something that is not the set requested by the user.","created":"2014-02-20T10:12:50.060+0000"},{"body":"+1 for 0.98. Please update this JIRA with a release note. \n\nI was in this part of the code when working on HBASE-9182. I would not claim our shell stuff in this area is the best it can be.","created":"2014-02-20T19:02:16.872+0000"},{"body":"Closing this issue after 0.99.0 release. ","created":"2015-02-21T23:32:23.605+0000"}],"conversations":[{"body":"created few tables. then performing disable_all operation in shell prompt.\nbut it is not performing operation successfully.\n{noformat}\nhbase(main):043:0> disable_all '*'\ntable12\nzk0113\nzk0114\n\nDisable the above 3 tables (y/n)?\ny/\n3 tables successfully disabled\n\njust it is showing the message but operation is not success.\n\nbut the following way only performing successfully\n\n\nhbase(main):043:0> disable_all '*.*'\ntable12\nzk0113\nzk0114\n\nDisable the above 3 tables (y/n)?\ny\n3 tables successfully disabled\n{noformat}\n","from":"reporter","subject":"enable_all,disable_all,drop_all can call \"list\" command with regex directly."},{"body":"This problem is because below code in disable_all.rb\n{code}\n regex = /^#{regex}$/ unless regex.is_a?(Regexp)\n list = admin.list.grep(regex)\n count = list.size\n\n{code}\nin case of * regex will become ^*$ which is valid in Ruby and admin.list.grep(regex) gives all available tables.\nBut same expression(^*$) in Java returns empty list.\n{code}\n failed = admin.disable_all(regex)\n puts \"#{count - failed.size} tables successfully disabled\"\n puts \"#{failed.size} tables not disabled due to an exception: #{failed.join ','}\" unless failed.size == 0\n{code}\nEven the number of tables to disabled is zero, output is \"#{count - failed.size} tables successfully disabled\".\nThis is applicable for enable_all and drop_all also.","from":"developer"},{"body":"Small change in above comment:\nEven the number of disabled tables is zero, We are priting \"#{count} tables successfully disabled\".\nHere count is total number of available tables.\n","from":"developer"},{"body":"Patch for trunk. If its ok I will upload patch for 94,92.","from":"developer"},{"body":"Pls check the same with other expressions? If things are fine then its ok..","from":"developer"},{"body":"@Ram\ntested with some basic regular expressions like t.*,.*,[t].*,'t\\d\\d\\d','table' and some more also. Its working fine.","from":"developer"},{"body":"Should we also fix the misleading messaging? It should not print a message indicating successful execution when in fact it was not.","from":"developer"},{"body":"Pushing to 0.94.4","from":"developer"},{"body":"@Lars,\nbq. It should not print a message indicating successful execution when in fact it was not.\nIt wont happen. Messaging wise its clear only.","from":"developer"},{"body":"Looking at this again. Are we changing the meaning?\n^{regex}$ requires the table name needs to match in its entirety.\nJust {regex} (depending on how it is called) could match when only a subset of the table name matches.\n","from":"developer"},{"body":"@Lars\nbq.Looking at this again. Are we changing the meaning?\nwith patch can be yes because ruby grep and java patterns may match table names differently.\n\nIf it is not correct you can close this issue as invalid.\nThanks.","from":"developer"},{"body":"I am removing this from 0.92 and 0.94.","from":"developer"},{"body":"This looks like it could be useful but also dangerous. Would like others to ask for this feature (and try it out) before committing. Moving out of 0.95 for now.","from":"developer"},{"body":"Faced the same \"issue\" recently. Might be good to fix that.\n\nNow, since Java and Ruby are not sending back the same result for a regex, I think we should the same API for the 2 filtering. To make sure it applies the same way on the list. Doing that will allow us to keep the current feature, but will fix it? Just a suggestion. Not sure if it's doable...","from":"developer"},{"body":"patch v1 is on the same line as the previous one.\nremoves all the uses of the ruby regex and instead uses the java Pattern, which we use in the HBaseAdmin methods.","from":"developer"},{"body":"Patch lgtm [~mbertozzi].\nIts better to commit to latest versions atleast.\nThanks.\n","from":"developer"},{"body":"+1","from":"developer"},{"body":"committed to trunk\n\n[~lhofhansl] [~stack] [~apurtell] your decision on backporting it to 94, 96, and 98. You may consider this incompatible change if people relies on globs for \"list\" commands, but the other side is the problem with enable_all, delete_all operations which are operating on something that is not the set requested by the user.","from":"developer"},{"body":"+1 for 0.98. Please update this JIRA with a release note. \n\nI was in this part of the code when working on HBASE-9182. I would not claim our shell stuff in this area is the best it can be.","from":"developer"},{"body":"Closing this issue after 0.99.0 release. ","from":"developer"}],"created":"2012-08-23T12:40:12.000+0000","description":"created few tables. then performing disable_all operation in shell prompt.\nbut it is not performing operation successfully.\n{noformat}\nhbase(main):043:0> disable_all '*'\ntable12\nzk0113\nzk0114\n\nDisable the above 3 tables (y/n)?\ny/\n3 tables successfully disabled\n\njust it is showing the message but operation is not success.\n\nbut the following way only performing successfully\n\n\nhbase(main):043:0> disable_all '*.*'\ntable12\nzk0113\nzk0114\n\nDisable the above 3 tables (y/n)?\ny\n3 tables successfully disabled\n{noformat}\n","issue_id":"12604602","key":"HBASE-6642","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2014-02-20T10:12:50.000+0000","role":"fixed_distractor","summary":"enable_all,disable_all,drop_all can call \"list\" command with regex directly."} {"case_id":"12398270","cluster":"DISTRACTOR-HBASE-686","comments":[{"body":"a testcase for this issue, i add this to TestHMemcache.java for reproduce and validating my patch.\n /**\n * Test memcache scanner scanning cached rows, see HBASE-686\n * @throws IOException\n */\n public void testScanner_686() throws IOException\n {\n addRows(this.hmemcache);\n long timestamp = System.currentTimeMillis();\n Text[] cols = new Text[COLUMNS_COUNT * ROW_COUNT];\n for (int i = 0; i < ROW_COUNT; i++)\n {\n for (int ii = 0; ii < COLUMNS_COUNT; ii++)\n {\n cols[(ii + (i * COLUMNS_COUNT))] = getColumnName(i, ii);\n }\n }\n //starting from each row, validate results should contain the starting row\n for (int startRowId = 0; startRowId < ROW_COUNT; startRowId++)\n {\n HInternalScannerInterface scanner =\n this.hmemcache.getScanner(timestamp, cols\n , new Text(getRowName(startRowId)));\n HStoreKey key = new HStoreKey();\n TreeMap results = new TreeMap();\n for (int i = 0; scanner.next(key, results); i++)\n {\n int rowId = startRowId + i;\n assertTrue(\"Row name\",\n key.toString().startsWith(getRowName(rowId).toString()));\n assertEquals(\"Count of columns\", COLUMNS_COUNT,\n results.size());\n TreeMap row = new TreeMap();\n for (Map.Entry e : results.entrySet())\n {\n row.put(e.getKey(), e.getValue());\n }\n isExpectedRow(rowId, row);\n // Clear out set. Otherwise row results accumulate.\n results.clear();\n }\n }\n }\n","created":"2008-06-14T08:17:16.722+0000"},{"body":"my patch, due to my local source changed for HBASE-684, the line number is incorrect i think.\n\nIndex: src/java/org/apache/hadoop/hbase/HStore.java\n===================================================================\n--- src/java/org/apache/hadoop/hbase/HStore.java\tSat Jun 14 15:46:26 CST 2008\n+++ src/java/org/apache/hadoop/hbase/HStore.java\tSat Jun 14 15:46:26 CST 2008\n@@ -631,8 +631,7 @@\n if (results.size() > 0) {\n results.clear();\n }\n- while (results.size() <= 0 &&\n- (this.currentRow = getNextRow(this.currentRow)) != null) {\n+ while (results.size() <= 0 && this.currentRow != null) {\n if (deletes.size() > 0) {\n deletes.clear();\n }\n@@ -661,6 +660,7 @@\n }\n results.put(column, c);\n }\n+ this.currentRow = getNextRow(this.currentRow);\n }\n return results.size() > 0;\n }\n","created":"2008-06-14T08:19:09.608+0000"},{"body":"LN,\n\nCould you please post your testcase and patch as attachments to this Jira. Jira has mangled your input to the extent that it is difficult to make a working test program or patch out of either of your posts.\n\nThanks.\n","created":"2008-06-16T21:09:25.779+0000"},{"body":" My test case :\n\n you have to create table 'aaa' with a column family 'xxx'\n\n\n\t\tHTable htable = new HTable(\"aaa\");\n\t\t BatchUpdate batch = new BatchUpdate(\"123456\".getBytes());\n\t\t batch.put(\"xxx\", \"some value\".getBytes());\n\t\t htable.commit(batch);\n\t\t\n\t\t byte[][] cols = new byte[][]{ \"xxx\".getBytes() };\t\t \t\t \n\t\t RowResult row = htable.getRow(\"123456\".getBytes(), cols, new Date().getTime());\n\t\t System.out.println(\"there is a row = \" + row.get(\"xxx\".getBytes()).toString());\n\t\t\n\t\t cols = new byte[][]{ \"xxx:\".getBytes() };\t\t \n\t\t Scanner scanner = htable.getScanner(cols);\n\t\t \n\t\t if (scanner.iterator().hasNext()) {\n\t\t\t System.out.println(\" GOOD !!!! - scanner returns some rows \");\n\t\t } else {\n\t\t\t System.out.println(\" WRONG !!!! - scanner didn't found any rows\");\n\t\t }\t\t \n\n","created":"2008-06-16T21:23:45.855+0000"},{"body":"This patch is listed as affecting hbase-0.1.2, but the APIs you are using only exist in trunk (0.2.0) is this issue for 0.1.2 or for 0.2.0?","created":"2008-06-16T21:34:13.875+0000"},{"body":"From the comments in the issue, this is the best test case I could come up with. It passes tests on hbase-0.1 branch","created":"2008-06-17T01:09:39.846+0000"},{"body":"This is a updated version of the test to work with trunk. Again, it passes with no issues.","created":"2008-06-17T01:14:49.157+0000"},{"body":"This issue cannot be replicated on either the 0.1 branch or trunk. If there are explicit test cases attached which can reproduce the issue, please re-open with both the versions affected and please attach a test case that demonstrates that the test case fails for that release.","created":"2008-06-17T01:22:17.581+0000"},{"body":"here is a modified memcache test cases(0.1.2 release) and my patch, i have add a new case 'testScanner_686'. i also have a small ant task definition to run this case seprately if u need.\n","created":"2008-06-17T05:33:26.596+0000"},{"body":"i can't understand why this issue resolved as CR in 9 hours(before i wake up:-). i HAVE post my testcase when opening this issue, with clearly piont out it is part of TestHMemcache.java, which is the 0.1.2 test source about memcache. \n\nand, the mistake in the source should be wrong iteration control, as i commented in the issue description.\n\njunit result here:\nTestsuite: org.apache.hadoop.hbase.TestHMemcache\nTests run: 7, Failures: 1, Errors: 0, Time elapsed: 0.22 sec\n\nTestcase: testSnapshotting took 0.16 sec\nTestcase: testGetFull took 0.01 sec\nTestcase: testGetNextRow took 0.01 sec\nTestcase: testGetClosest took 0 sec\nTestcase: testScanner_686 took 0.01 sec\n\tFAILED\nRow name\njunit.framework.AssertionFailedError: Row name\n\tat org.apache.hadoop.hbase.TestHMemcache.testScanner_686(TestHMemcache.java:195)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\nTestcase: testScanner took 0.01 sec\nTestcase: testGetRowKeyAtOrBefore took 0.02 sec\n","created":"2008-06-17T05:56:21.828+0000"},{"body":"it's test case for ver 0.2.0 dev from trunk. - sources inside the jar package\n\n this is class with main method which do : \n 1) creates table - if it doesn't exist\n 2) create BatchUpdate with a \"rowKey\" and commit into a table\n 3) make - getRow with \"rowKey\" - and shows on console data in this row\n 4) make - getScanner on whole table and shows all rows - and counts how many rows scanner returns\n\nit returns :\n\ntable : testTable exists. \n data inserted \n-------------------------------------\n getRow : \n isEmpty = false\n row key = 0000\n col val = testvalue\n-------------------------------------\n getScanner : \nscanner find : 0 rows in this table.\n done. \n\n what does it mean ?\ngetRow can gets data by rowKey - this data exists in table, but the scanner doesn't see this row - and return 0 rows.\n\n Antoni","created":"2008-06-17T12:37:37.006+0000"},{"body":"Mea culpa on closing this issue too early.\n\nThe patch and test case provided do indeed demonstrate and fix the problem. All other regression tests passed as well with the patch applied. Committed to 0.1 branch. Thanks for the contribution LN!","created":"2008-06-17T17:14:23.025+0000"},{"body":"It also affects 0.2 from trunk as test case provided by Antonii (in scanner-test-for.ver.0.2.0.dev.jar) shows","created":"2008-06-17T18:33:08.870+0000"},{"body":"Committed to both 0.1 branch and trunk. ","created":"2008-06-17T20:52:04.704+0000"}],"conversations":[{"body":"HTable.obtainScanner methods should return the start row if it exists, although HTable's javadoc didn't clearly desc. but i found the result of htable scanners sometimes contain the start row, sometimes not.\n\nafter more testing and code review, i found it should be a bug in HStore.Memcache.MemcacheScanner. in the constructor it set this.currentRow = firstRow, but when doing next(), there's a this.currentRow = getNextRow(this.currentRow) before fetch result.\n","from":"reporter","subject":"MemcacheScanner didn't return the first row(if it exists), cause HScannerInterface's output incorrect"},{"body":"a testcase for this issue, i add this to TestHMemcache.java for reproduce and validating my patch.\n /**\n * Test memcache scanner scanning cached rows, see HBASE-686\n * @throws IOException\n */\n public void testScanner_686() throws IOException\n {\n addRows(this.hmemcache);\n long timestamp = System.currentTimeMillis();\n Text[] cols = new Text[COLUMNS_COUNT * ROW_COUNT];\n for (int i = 0; i < ROW_COUNT; i++)\n {\n for (int ii = 0; ii < COLUMNS_COUNT; ii++)\n {\n cols[(ii + (i * COLUMNS_COUNT))] = getColumnName(i, ii);\n }\n }\n //starting from each row, validate results should contain the starting row\n for (int startRowId = 0; startRowId < ROW_COUNT; startRowId++)\n {\n HInternalScannerInterface scanner =\n this.hmemcache.getScanner(timestamp, cols\n , new Text(getRowName(startRowId)));\n HStoreKey key = new HStoreKey();\n TreeMap results = new TreeMap();\n for (int i = 0; scanner.next(key, results); i++)\n {\n int rowId = startRowId + i;\n assertTrue(\"Row name\",\n key.toString().startsWith(getRowName(rowId).toString()));\n assertEquals(\"Count of columns\", COLUMNS_COUNT,\n results.size());\n TreeMap row = new TreeMap();\n for (Map.Entry e : results.entrySet())\n {\n row.put(e.getKey(), e.getValue());\n }\n isExpectedRow(rowId, row);\n // Clear out set. Otherwise row results accumulate.\n results.clear();\n }\n }\n }\n","from":"developer"},{"body":"my patch, due to my local source changed for HBASE-684, the line number is incorrect i think.\n\nIndex: src/java/org/apache/hadoop/hbase/HStore.java\n===================================================================\n--- src/java/org/apache/hadoop/hbase/HStore.java\tSat Jun 14 15:46:26 CST 2008\n+++ src/java/org/apache/hadoop/hbase/HStore.java\tSat Jun 14 15:46:26 CST 2008\n@@ -631,8 +631,7 @@\n if (results.size() > 0) {\n results.clear();\n }\n- while (results.size() <= 0 &&\n- (this.currentRow = getNextRow(this.currentRow)) != null) {\n+ while (results.size() <= 0 && this.currentRow != null) {\n if (deletes.size() > 0) {\n deletes.clear();\n }\n@@ -661,6 +660,7 @@\n }\n results.put(column, c);\n }\n+ this.currentRow = getNextRow(this.currentRow);\n }\n return results.size() > 0;\n }\n","from":"developer"},{"body":"LN,\n\nCould you please post your testcase and patch as attachments to this Jira. Jira has mangled your input to the extent that it is difficult to make a working test program or patch out of either of your posts.\n\nThanks.\n","from":"developer"},{"body":" My test case :\n\n you have to create table 'aaa' with a column family 'xxx'\n\n\n\t\tHTable htable = new HTable(\"aaa\");\n\t\t BatchUpdate batch = new BatchUpdate(\"123456\".getBytes());\n\t\t batch.put(\"xxx\", \"some value\".getBytes());\n\t\t htable.commit(batch);\n\t\t\n\t\t byte[][] cols = new byte[][]{ \"xxx\".getBytes() };\t\t \t\t \n\t\t RowResult row = htable.getRow(\"123456\".getBytes(), cols, new Date().getTime());\n\t\t System.out.println(\"there is a row = \" + row.get(\"xxx\".getBytes()).toString());\n\t\t\n\t\t cols = new byte[][]{ \"xxx:\".getBytes() };\t\t \n\t\t Scanner scanner = htable.getScanner(cols);\n\t\t \n\t\t if (scanner.iterator().hasNext()) {\n\t\t\t System.out.println(\" GOOD !!!! - scanner returns some rows \");\n\t\t } else {\n\t\t\t System.out.println(\" WRONG !!!! - scanner didn't found any rows\");\n\t\t }\t\t \n\n","from":"developer"},{"body":"This patch is listed as affecting hbase-0.1.2, but the APIs you are using only exist in trunk (0.2.0) is this issue for 0.1.2 or for 0.2.0?","from":"developer"},{"body":"From the comments in the issue, this is the best test case I could come up with. It passes tests on hbase-0.1 branch","from":"developer"},{"body":"This is a updated version of the test to work with trunk. Again, it passes with no issues.","from":"developer"},{"body":"This issue cannot be replicated on either the 0.1 branch or trunk. If there are explicit test cases attached which can reproduce the issue, please re-open with both the versions affected and please attach a test case that demonstrates that the test case fails for that release.","from":"developer"},{"body":"here is a modified memcache test cases(0.1.2 release) and my patch, i have add a new case 'testScanner_686'. i also have a small ant task definition to run this case seprately if u need.\n","from":"developer"},{"body":"i can't understand why this issue resolved as CR in 9 hours(before i wake up:-). i HAVE post my testcase when opening this issue, with clearly piont out it is part of TestHMemcache.java, which is the 0.1.2 test source about memcache. \n\nand, the mistake in the source should be wrong iteration control, as i commented in the issue description.\n\njunit result here:\nTestsuite: org.apache.hadoop.hbase.TestHMemcache\nTests run: 7, Failures: 1, Errors: 0, Time elapsed: 0.22 sec\n\nTestcase: testSnapshotting took 0.16 sec\nTestcase: testGetFull took 0.01 sec\nTestcase: testGetNextRow took 0.01 sec\nTestcase: testGetClosest took 0 sec\nTestcase: testScanner_686 took 0.01 sec\n\tFAILED\nRow name\njunit.framework.AssertionFailedError: Row name\n\tat org.apache.hadoop.hbase.TestHMemcache.testScanner_686(TestHMemcache.java:195)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)\n\nTestcase: testScanner took 0.01 sec\nTestcase: testGetRowKeyAtOrBefore took 0.02 sec\n","from":"developer"},{"body":"it's test case for ver 0.2.0 dev from trunk. - sources inside the jar package\n\n this is class with main method which do : \n 1) creates table - if it doesn't exist\n 2) create BatchUpdate with a \"rowKey\" and commit into a table\n 3) make - getRow with \"rowKey\" - and shows on console data in this row\n 4) make - getScanner on whole table and shows all rows - and counts how many rows scanner returns\n\nit returns :\n\ntable : testTable exists. \n data inserted \n-------------------------------------\n getRow : \n isEmpty = false\n row key = 0000\n col val = testvalue\n-------------------------------------\n getScanner : \nscanner find : 0 rows in this table.\n done. \n\n what does it mean ?\ngetRow can gets data by rowKey - this data exists in table, but the scanner doesn't see this row - and return 0 rows.\n\n Antoni","from":"developer"},{"body":"Mea culpa on closing this issue too early.\n\nThe patch and test case provided do indeed demonstrate and fix the problem. All other regression tests passed as well with the patch applied. Committed to 0.1 branch. Thanks for the contribution LN!","from":"developer"},{"body":"It also affects 0.2 from trunk as test case provided by Antonii (in scanner-test-for.ver.0.2.0.dev.jar) shows","from":"developer"},{"body":"Committed to both 0.1 branch and trunk. ","from":"developer"}],"created":"2008-06-14T08:12:48.000+0000","description":"HTable.obtainScanner methods should return the start row if it exists, although HTable's javadoc didn't clearly desc. but i found the result of htable scanners sometimes contain the start row, sometimes not.\n\nafter more testing and code review, i found it should be a bug in HStore.Memcache.MemcacheScanner. in the constructor it set this.currentRow = firstRow, but when doing next(), there's a this.currentRow = getNextRow(this.currentRow) before fetch result.\n","issue_id":"12398270","key":"HBASE-686","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2008-06-17T20:52:04.000+0000","role":"fixed_distractor","summary":"MemcacheScanner didn't return the first row(if it exists), cause HScannerInterface's output incorrect"} {"case_id":"12609852","cluster":"DISTRACTOR-HBASE-6916","comments":[{"body":"Attaching a patch that simply replaces the LOG.info statements with exceptions that are thrown. It also fixes the logic in closeRegion.\n\nNow I see:\n\n{noformat}\nhbase(main):001:0> close_region 'thisisaninvalidregion'\n2012-10-01 18:27:51.595 java[29064:1203]\n\nERROR: org.apache.hadoop.hbase.UnknownRegionException: thisisaninvalidregion\n{noformat}","created":"2012-10-02T01:37:56.607+0000"},{"body":"+1 on commit. Put in 0.94 to make Lars happy.","created":"2012-10-02T03:26:19.531+0000"},{"body":"Attaching the trunk patch.","created":"2012-10-03T23:27:52.735+0000"},{"body":"+1","created":"2012-10-03T23:43:07.700+0000"},{"body":"The TestShell failure is legit, it's always been failing but silently:\n\n{noformat}\nWed Oct 3 23:45:04 UTC 2012......Finished in 72.002 seconds. 1) Error:\ntest_close_should_work_without_region_server_name(Hbase::AdminMethodsTest):\nNativeException: org.apache.hadoop.hbase.UnknownRegionException: hbase_create_table_test_table,,0\n org/apache/hadoop/hbase/client/HBaseAdmin.java:1115:in `closeRegion'\n{noformat}\n\nThat region name doesn't exist, it's bogus.\n\nThe other ones seem unrelated.","created":"2012-10-04T00:50:38.694+0000"},{"body":"In this new patch I removed the close_region test since it's broken and I can't think of a nice way to get a full region name in there.","created":"2012-10-04T01:02:53.788+0000"},{"body":"+1 lgtm","created":"2012-10-04T21:15:24.732+0000"},{"body":"Committed to .92, .94, and trunk. Thanks for the reviews guys.","created":"2012-10-04T21:23:17.522+0000"},{"body":"bq. +1 on commit. Put in 0.94 to make Lars happy.\n\nThat made me happy ;)","created":"2012-10-04T22:45:49.914+0000"},{"body":"Fix up after bulk move overwrote some 0.94.2 fix versions w/ 0.95.0 (Noticed by Lars Hofhansl)","created":"2013-04-07T05:07:26.898+0000"}],"conversations":[{"body":"There is a weird interaction between the shell and HBA. When you try to close a region that doesn't exist, it doesn't throw any error:\n\n{noformat}\nhbase(main):029:0> close_region 'thisisaninvalidregion'\n0 row(s) in 0.0580 seconds\n{noformat}\n\nNormally one should get UnknownRegionException. Starting the shell with \"-d\" I see what a non-shell user would see along with a ton of logging from ZK (skipped here):\n\n{noformat}\nINFO client.HBaseAdmin: No server in .META. for thisisaninvalidregion; pair=null\n{noformat}\n\nBut again this is not the right message, it should have shown\n\n{noformat}\nINFO client.HBaseAdmin: No server in .META. for thisisaninvalidregion; pair=null\n{noformat}\n\nAnd this is because that part of the code treats both UnknownRegionException and NoServerForRegionException like if it was the same thing.\n\nThere is also some ugliness in flush, compact, and split but it normally doesn't show since the code treats everything like it's a table and sends a TableNotFoundException.\n\nThis jira is about making sure that the exceptions are correctly coming out.","from":"reporter","subject":"HBA logs at info level errors that won't show in the shell"},{"body":"Attaching a patch that simply replaces the LOG.info statements with exceptions that are thrown. It also fixes the logic in closeRegion.\n\nNow I see:\n\n{noformat}\nhbase(main):001:0> close_region 'thisisaninvalidregion'\n2012-10-01 18:27:51.595 java[29064:1203]\n\nERROR: org.apache.hadoop.hbase.UnknownRegionException: thisisaninvalidregion\n{noformat}","from":"developer"},{"body":"+1 on commit. Put in 0.94 to make Lars happy.","from":"developer"},{"body":"Attaching the trunk patch.","from":"developer"},{"body":"+1","from":"developer"},{"body":"The TestShell failure is legit, it's always been failing but silently:\n\n{noformat}\nWed Oct 3 23:45:04 UTC 2012......Finished in 72.002 seconds. 1) Error:\ntest_close_should_work_without_region_server_name(Hbase::AdminMethodsTest):\nNativeException: org.apache.hadoop.hbase.UnknownRegionException: hbase_create_table_test_table,,0\n org/apache/hadoop/hbase/client/HBaseAdmin.java:1115:in `closeRegion'\n{noformat}\n\nThat region name doesn't exist, it's bogus.\n\nThe other ones seem unrelated.","from":"developer"},{"body":"In this new patch I removed the close_region test since it's broken and I can't think of a nice way to get a full region name in there.","from":"developer"},{"body":"+1 lgtm","from":"developer"},{"body":"Committed to .92, .94, and trunk. Thanks for the reviews guys.","from":"developer"},{"body":"bq. +1 on commit. Put in 0.94 to make Lars happy.\n\nThat made me happy ;)","from":"developer"},{"body":"Fix up after bulk move overwrote some 0.94.2 fix versions w/ 0.95.0 (Noticed by Lars Hofhansl)","from":"developer"}],"created":"2012-10-02T01:35:31.000+0000","description":"There is a weird interaction between the shell and HBA. When you try to close a region that doesn't exist, it doesn't throw any error:\n\n{noformat}\nhbase(main):029:0> close_region 'thisisaninvalidregion'\n0 row(s) in 0.0580 seconds\n{noformat}\n\nNormally one should get UnknownRegionException. Starting the shell with \"-d\" I see what a non-shell user would see along with a ton of logging from ZK (skipped here):\n\n{noformat}\nINFO client.HBaseAdmin: No server in .META. for thisisaninvalidregion; pair=null\n{noformat}\n\nBut again this is not the right message, it should have shown\n\n{noformat}\nINFO client.HBaseAdmin: No server in .META. for thisisaninvalidregion; pair=null\n{noformat}\n\nAnd this is because that part of the code treats both UnknownRegionException and NoServerForRegionException like if it was the same thing.\n\nThere is also some ugliness in flush, compact, and split but it normally doesn't show since the code treats everything like it's a table and sends a TableNotFoundException.\n\nThis jira is about making sure that the exceptions are correctly coming out.","issue_id":"12609852","key":"HBASE-6916","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2012-10-04T21:23:17.000+0000","role":"fixed_distractor","summary":"HBA logs at info level errors that won't show in the shell"} {"case_id":"12400852","cluster":"DISTRACTOR-HBASE-766","comments":[{"body":"Move the getting of file length to AFTER test of file existance.","created":"2008-07-23T18:03:51.110+0000"},{"body":"I committed this after running tests up on cluster to confirm doesn't break anything.","created":"2008-07-23T22:10:48.122+0000"},{"body":"Jon Gray can reliably reproduce. Looks like we are trying to split regions that still hold references.","created":"2008-07-29T07:59:22.502+0000"},{"body":"We dropped check for references; we need to always compact if references.","created":"2008-07-29T17:54:16.247+0000"},{"body":"stack, please confirm if i need to apply both patches or just hs.patch.\n\nHave actually tested with both patches, neither seems to change behavior at all.\n\nAgain, happens during the second split of the problem table:\n\n\n2008-07-29 15:11:51,677 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Completed compaction of 426199566/content store size is 378.0m\n2008-07-29 15:11:51,684 INFO org.apache.hadoop.hbase.regionserver.HRegion: compaction completed on region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384 in 7sec\n2008-07-29 15:11:51,686 INFO org.apache.hadoop.hbase.regionserver.HRegion: Starting split of region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:51,689 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Compactions and cache flushes disabled for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:51,689 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Updates and scanners disabled for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:51,689 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: No more active scanners for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:51,689 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: No more row locks outstanding on region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:51,689 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Started memcache flush for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384. Current region memcache size 1.9m\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Added /hbase/items/426199566/content/mapfiles/7100309375208596278 with 10296 entries, sequence id 2537536, data size 1.9m, file size 2.1m\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Finished memcache flush for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384 in 92ms, sequence id=2537536, compaction requested=true\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/uclusters\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/covisit\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/content\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/clusters\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/readby\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/cfrecs\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/blooms\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/ranked\n2008-07-29 15:11:51,782 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/savedby\n2008-07-29 15:11:51,782 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/sentby\n2008-07-29 15:11:51,782 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/receivedby\n2008-07-29 15:11:51,782 INFO org.apache.hadoop.hbase.regionserver.HRegion: closed items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:52,033 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Opening region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369511688/509041795\n2008-07-29 15:11:52,049 INFO org.apache.hadoop.hbase.regionserver.HStore: HSTORE_LOGINFOFILE 509041795/savedby/4080673342751735986-426199566/8532274243649733439/bottom does not contain a sequence number - ignoring\n2008-07-29 15:11:52,052 ERROR org.apache.hadoop.hbase.regionserver.CompactSplitThread: Compaction failed for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\njava.io.FileNotFoundException: File does not exist: hdfs://mb0:9000/hbase/items/426199566/savedby/mapfiles/8532274243649733439/data\n at org.apache.hadoop.dfs.DistributedFileSystem.getFileStatus(DistributedFileSystem.java:369)\n at org.apache.hadoop.hbase.regionserver.HStoreFile.length(HStoreFile.java:444)\n at org.apache.hadoop.hbase.regionserver.HStore.loadHStoreFiles(HStore.java:437)\n at org.apache.hadoop.hbase.regionserver.HStore.(HStore.java:219)\n at org.apache.hadoop.hbase.regionserver.HRegion.instantiateHStore(HRegion.java:1618)\n at org.apache.hadoop.hbase.regionserver.HRegion.(HRegion.java:466)\n at org.apache.hadoop.hbase.regionserver.HRegion.(HRegion.java:405)\n at org.apache.hadoop.hbase.regionserver.HRegion.splitRegion(HRegion.java:800)\n at org.apache.hadoop.hbase.regionserver.CompactSplitThread.split(CompactSplitThread.java:133)\n at org.apache.hadoop.hbase.regionserver.CompactSplitThread.run(CompactSplitThread.java:86)\n","created":"2008-07-29T22:17:27.664+0000"},{"body":"Patch 766.patch has already been committed. See comment above at 'stack - 23/Jul/08 03:10 PM'. That patch was about something else. hs.patch is latest (JIRA has date if you hover over it). \n\nI took a look up on your cluster. Looking at hb0 or in mb0, src doesn't show my patch in place. Where'd you build the jar?","created":"2008-07-30T04:01:34.540+0000"},{"body":"Stack,\n\nI used the jar you built, confirmed across the cluster it was the one being used, and ran the test... I don't see any additional logging. Check out the region logs on hb5 (where items table was located).\n\nHere is the error again:\n\n2008-07-30 17:41:58,970 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Started memcache flush for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950. Current region memcache size 64.0m\n2008-07-30 17:42:00,424 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Added /hbase/items/1211238171/content/mapfiles/793072914512989834 with 323136 entries, sequence id 2527714, data size 64.0m, file size 68.8m\n2008-07-30 17:42:00,424 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Finished memcache flush for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950 in 1454ms, sequence id=2527714, compaction requested=true\n2008-07-30 17:42:00,424 DEBUG org.apache.hadoop.hbase.regionserver.CompactSplitThread: Compaction requested for region: items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:00,424 INFO org.apache.hadoop.hbase.regionserver.HRegion: starting compaction on region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:00,428 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Compaction size 396611811, skipped 1, 252131417\n2008-07-30 17:42:00,454 DEBUG org.apache.hadoop.hbase.regionserver.HStore: started compaction of 2 files into /hbase/items/compaction.dir/1211238171/content/mapfiles/4240641699209031496\n2008-07-30 17:42:08,256 DEBUG org.apache.hadoop.hbase.regionserver.HStore: moving /hbase/items/compaction.dir/1211238171/content/mapfiles/4240641699209031496 to /hbase/items/1211238171/content/mapfiles/8894612533024691980\n2008-07-30 17:42:08,293 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Completed compaction of 1211238171/content store size is 377.8m\n2008-07-30 17:42:08,299 INFO org.apache.hadoop.hbase.regionserver.HRegion: compaction completed on region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950 in 7sec\n2008-07-30 17:42:08,301 INFO org.apache.hadoop.hbase.regionserver.HRegion: Starting split of region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,304 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Compactions and cache flushes disabled for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,304 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Updates and scanners disabled for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,304 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: No more active scanners for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,305 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: No more row locks outstanding on region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,305 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Started memcache flush for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950. Current region memcache size 1.9m\n2008-07-30 17:42:08,406 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Added /hbase/items/1211238171/content/mapfiles/2539623802632323222 with 10032 entries, sequence id 2537747, data size 1.9m, file size 2.0m\n2008-07-30 17:42:08,406 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Finished memcache flush for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950 in 101ms, sequence id=2537747, compaction requested=true\n2008-07-30 17:42:08,406 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/uclusters\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/covisit\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/content\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/clusters\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/readby\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/cfrecs\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/blooms\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/ranked\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/savedby\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/sentby\n2008-07-30 17:42:08,408 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/receivedby\n2008-07-30 17:42:08,408 INFO org.apache.hadoop.hbase.regionserver.HRegion: closed items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,638 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Opening region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464928303/2132462761\n2008-07-30 17:42:08,650 INFO org.apache.hadoop.hbase.regionserver.HStore: HSTORE_LOGINFOFILE 2132462761/readby/8482419619225016670-1211238171/6573863319629365608/bottom does not contain a sequence number - ignoring\n2008-07-30 17:42:08,653 ERROR org.apache.hadoop.hbase.regionserver.CompactSplitThread: Compaction failed for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\njava.io.FileNotFoundException: File does not exist: hdfs://mb0:9000/hbase/items/1211238171/readby/mapfiles/6573863319629365608/data\n at org.apache.hadoop.dfs.DistributedFileSystem.getFileStatus(DistributedFileSystem.java:369)\n at org.apache.hadoop.hbase.regionserver.HStoreFile.length(HStoreFile.java:444)\n at org.apache.hadoop.hbase.regionserver.HStore.loadHStoreFiles(HStore.java:437)\n at org.apache.hadoop.hbase.regionserver.HStore.(HStore.java:219)\n at org.apache.hadoop.hbase.regionserver.HRegion.instantiateHStore(HRegion.java:1618)\n at org.apache.hadoop.hbase.regionserver.HRegion.(HRegion.java:466)\n at org.apache.hadoop.hbase.regionserver.HRegion.(HRegion.java:405)\n at org.apache.hadoop.hbase.regionserver.HRegion.splitRegion(HRegion.java:800)\n at org.apache.hadoop.hbase.regionserver.CompactSplitThread.split(CompactSplitThread.java:133)\n at org.apache.hadoop.hbase.regionserver.CompactSplitThread.run(CompactSplitThread.java:86)\n","created":"2008-07-31T00:44:45.515+0000"},{"body":"Reopening. FNFE out on jgray's cluser. ","created":"2008-07-31T22:25:20.937+0000"},{"body":"+1, exact same import that recreated the error every time now succeeds","created":"2008-07-31T22:56:09.115+0000"},{"body":"Resolving. JGray says it works.\n\nLooks like I committed this fix already as part of another patch:\n\n------------------------------------------------------------------------\nr680910 | stack | 2008-07-29 21:37:06 -0700 (Tue, 29 Jul 2008) | 1 line\n\nHBASE-783 For single row, single family retrieval, getRow() works half as fast as getScanner().next()\n\n","created":"2008-07-31T23:01:57.936+0000"},{"body":"Re-confirmed that this is fixed on trunk","created":"2008-08-01T01:30:12.628+0000"}],"conversations":[{"body":"From Renaud Delbru up on the list:\n{code}\nHi,\n\nafter our issues (\"Replay of HLog required\", in a precious thread) with HBase, it seems that HBase has corrupted regions.\nWe have, on the three region servers, errors stating that HBase cannot open certain regions because some map files on hdfs are missing (see the log attached).\n\nDo you have any ideas how to fix this ?\n\nThanks.\n\njava.io.FileNotFoundException: File does not exist: hdfs://hadoop1.sindice.net:54310/hbase/page-repository/1105668475/field/mapfiles/5122893264992435570/data\n at org.apache.hadoop.dfs.DistributedFileSystem.getFileStatus(DistributedFileSystem.java:369)\n at org.apache.hadoop.hbase.regionserver.HStoreFile.length(HStoreFile.java:464)\n at org.apache.hadoop.hbase.regionserver.HStore.loadHStoreFiles(HStore.java:409)\n at org.apache.hadoop.hbase.regionserver.HStore.(HStore.java:236)\n at org.apache.hadoop.hbase.regionserver.HRegion.instantiateHStore(HRegion.java:1575)\n at org.apache.hadoop.hbase.regionserver.HRegion.(HRegion.java:451)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.instantiateRegion(HRegionServer.java:901)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.openRegion(HRegionServer.java:876)\n at org.apache.hadoop.hbase.regionserver.HRegionServer$Worker.run(HRegionServer.java:816)\n at java.lang.Thread.run(Thread.java:619)\n\n{code}\n\nI'd thought we'd added handling of this kind of event post HDFS crashings but looking in code, it seems that TRUNK may not have been fixed up properly.","from":"reporter","subject":"FileNotFoundException trying to load HStoreFile 'data'"},{"body":"Move the getting of file length to AFTER test of file existance.","from":"developer"},{"body":"I committed this after running tests up on cluster to confirm doesn't break anything.","from":"developer"},{"body":"Jon Gray can reliably reproduce. Looks like we are trying to split regions that still hold references.","from":"developer"},{"body":"We dropped check for references; we need to always compact if references.","from":"developer"},{"body":"stack, please confirm if i need to apply both patches or just hs.patch.\n\nHave actually tested with both patches, neither seems to change behavior at all.\n\nAgain, happens during the second split of the problem table:\n\n\n2008-07-29 15:11:51,677 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Completed compaction of 426199566/content store size is 378.0m\n2008-07-29 15:11:51,684 INFO org.apache.hadoop.hbase.regionserver.HRegion: compaction completed on region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384 in 7sec\n2008-07-29 15:11:51,686 INFO org.apache.hadoop.hbase.regionserver.HRegion: Starting split of region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:51,689 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Compactions and cache flushes disabled for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:51,689 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Updates and scanners disabled for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:51,689 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: No more active scanners for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:51,689 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: No more row locks outstanding on region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:51,689 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Started memcache flush for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384. Current region memcache size 1.9m\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Added /hbase/items/426199566/content/mapfiles/7100309375208596278 with 10296 entries, sequence id 2537536, data size 1.9m, file size 2.1m\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Finished memcache flush for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384 in 92ms, sequence id=2537536, compaction requested=true\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/uclusters\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/covisit\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/content\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/clusters\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/readby\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/cfrecs\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/blooms\n2008-07-29 15:11:51,781 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/ranked\n2008-07-29 15:11:51,782 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/savedby\n2008-07-29 15:11:51,782 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/sentby\n2008-07-29 15:11:51,782 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 426199566/receivedby\n2008-07-29 15:11:51,782 INFO org.apache.hadoop.hbase.regionserver.HRegion: closed items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\n2008-07-29 15:11:52,033 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Opening region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369511688/509041795\n2008-07-29 15:11:52,049 INFO org.apache.hadoop.hbase.regionserver.HStore: HSTORE_LOGINFOFILE 509041795/savedby/4080673342751735986-426199566/8532274243649733439/bottom does not contain a sequence number - ignoring\n2008-07-29 15:11:52,052 ERROR org.apache.hadoop.hbase.regionserver.CompactSplitThread: Compaction failed for region items,0050dab2-a057-43d9-8e99-0a41a3e6bdbb,1217369332384\njava.io.FileNotFoundException: File does not exist: hdfs://mb0:9000/hbase/items/426199566/savedby/mapfiles/8532274243649733439/data\n at org.apache.hadoop.dfs.DistributedFileSystem.getFileStatus(DistributedFileSystem.java:369)\n at org.apache.hadoop.hbase.regionserver.HStoreFile.length(HStoreFile.java:444)\n at org.apache.hadoop.hbase.regionserver.HStore.loadHStoreFiles(HStore.java:437)\n at org.apache.hadoop.hbase.regionserver.HStore.(HStore.java:219)\n at org.apache.hadoop.hbase.regionserver.HRegion.instantiateHStore(HRegion.java:1618)\n at org.apache.hadoop.hbase.regionserver.HRegion.(HRegion.java:466)\n at org.apache.hadoop.hbase.regionserver.HRegion.(HRegion.java:405)\n at org.apache.hadoop.hbase.regionserver.HRegion.splitRegion(HRegion.java:800)\n at org.apache.hadoop.hbase.regionserver.CompactSplitThread.split(CompactSplitThread.java:133)\n at org.apache.hadoop.hbase.regionserver.CompactSplitThread.run(CompactSplitThread.java:86)\n","from":"developer"},{"body":"Patch 766.patch has already been committed. See comment above at 'stack - 23/Jul/08 03:10 PM'. That patch was about something else. hs.patch is latest (JIRA has date if you hover over it). \n\nI took a look up on your cluster. Looking at hb0 or in mb0, src doesn't show my patch in place. Where'd you build the jar?","from":"developer"},{"body":"Stack,\n\nI used the jar you built, confirmed across the cluster it was the one being used, and ran the test... I don't see any additional logging. Check out the region logs on hb5 (where items table was located).\n\nHere is the error again:\n\n2008-07-30 17:41:58,970 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Started memcache flush for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950. Current region memcache size 64.0m\n2008-07-30 17:42:00,424 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Added /hbase/items/1211238171/content/mapfiles/793072914512989834 with 323136 entries, sequence id 2527714, data size 64.0m, file size 68.8m\n2008-07-30 17:42:00,424 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Finished memcache flush for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950 in 1454ms, sequence id=2527714, compaction requested=true\n2008-07-30 17:42:00,424 DEBUG org.apache.hadoop.hbase.regionserver.CompactSplitThread: Compaction requested for region: items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:00,424 INFO org.apache.hadoop.hbase.regionserver.HRegion: starting compaction on region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:00,428 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Compaction size 396611811, skipped 1, 252131417\n2008-07-30 17:42:00,454 DEBUG org.apache.hadoop.hbase.regionserver.HStore: started compaction of 2 files into /hbase/items/compaction.dir/1211238171/content/mapfiles/4240641699209031496\n2008-07-30 17:42:08,256 DEBUG org.apache.hadoop.hbase.regionserver.HStore: moving /hbase/items/compaction.dir/1211238171/content/mapfiles/4240641699209031496 to /hbase/items/1211238171/content/mapfiles/8894612533024691980\n2008-07-30 17:42:08,293 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Completed compaction of 1211238171/content store size is 377.8m\n2008-07-30 17:42:08,299 INFO org.apache.hadoop.hbase.regionserver.HRegion: compaction completed on region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950 in 7sec\n2008-07-30 17:42:08,301 INFO org.apache.hadoop.hbase.regionserver.HRegion: Starting split of region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,304 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Compactions and cache flushes disabled for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,304 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Updates and scanners disabled for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,304 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: No more active scanners for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,305 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: No more row locks outstanding on region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,305 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Started memcache flush for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950. Current region memcache size 1.9m\n2008-07-30 17:42:08,406 DEBUG org.apache.hadoop.hbase.regionserver.HStore: Added /hbase/items/1211238171/content/mapfiles/2539623802632323222 with 10032 entries, sequence id 2537747, data size 1.9m, file size 2.0m\n2008-07-30 17:42:08,406 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Finished memcache flush for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950 in 101ms, sequence id=2537747, compaction requested=true\n2008-07-30 17:42:08,406 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/uclusters\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/covisit\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/content\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/clusters\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/readby\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/cfrecs\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/blooms\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/ranked\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/savedby\n2008-07-30 17:42:08,407 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/sentby\n2008-07-30 17:42:08,408 DEBUG org.apache.hadoop.hbase.regionserver.HStore: closed 1211238171/receivedby\n2008-07-30 17:42:08,408 INFO org.apache.hadoop.hbase.regionserver.HRegion: closed items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\n2008-07-30 17:42:08,638 DEBUG org.apache.hadoop.hbase.regionserver.HRegion: Opening region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464928303/2132462761\n2008-07-30 17:42:08,650 INFO org.apache.hadoop.hbase.regionserver.HStore: HSTORE_LOGINFOFILE 2132462761/readby/8482419619225016670-1211238171/6573863319629365608/bottom does not contain a sequence number - ignoring\n2008-07-30 17:42:08,653 ERROR org.apache.hadoop.hbase.regionserver.CompactSplitThread: Compaction failed for region items,0050a173-36b8-45a4-a991-2bab7f5c3cd6,1217464581950\njava.io.FileNotFoundException: File does not exist: hdfs://mb0:9000/hbase/items/1211238171/readby/mapfiles/6573863319629365608/data\n at org.apache.hadoop.dfs.DistributedFileSystem.getFileStatus(DistributedFileSystem.java:369)\n at org.apache.hadoop.hbase.regionserver.HStoreFile.length(HStoreFile.java:444)\n at org.apache.hadoop.hbase.regionserver.HStore.loadHStoreFiles(HStore.java:437)\n at org.apache.hadoop.hbase.regionserver.HStore.(HStore.java:219)\n at org.apache.hadoop.hbase.regionserver.HRegion.instantiateHStore(HRegion.java:1618)\n at org.apache.hadoop.hbase.regionserver.HRegion.(HRegion.java:466)\n at org.apache.hadoop.hbase.regionserver.HRegion.(HRegion.java:405)\n at org.apache.hadoop.hbase.regionserver.HRegion.splitRegion(HRegion.java:800)\n at org.apache.hadoop.hbase.regionserver.CompactSplitThread.split(CompactSplitThread.java:133)\n at org.apache.hadoop.hbase.regionserver.CompactSplitThread.run(CompactSplitThread.java:86)\n","from":"developer"},{"body":"Reopening. FNFE out on jgray's cluser. ","from":"developer"},{"body":"+1, exact same import that recreated the error every time now succeeds","from":"developer"},{"body":"Resolving. JGray says it works.\n\nLooks like I committed this fix already as part of another patch:\n\n------------------------------------------------------------------------\nr680910 | stack | 2008-07-29 21:37:06 -0700 (Tue, 29 Jul 2008) | 1 line\n\nHBASE-783 For single row, single family retrieval, getRow() works half as fast as getScanner().next()\n\n","from":"developer"},{"body":"Re-confirmed that this is fixed on trunk","from":"developer"}],"created":"2008-07-23T18:01:57.000+0000","description":"From Renaud Delbru up on the list:\n{code}\nHi,\n\nafter our issues (\"Replay of HLog required\", in a precious thread) with HBase, it seems that HBase has corrupted regions.\nWe have, on the three region servers, errors stating that HBase cannot open certain regions because some map files on hdfs are missing (see the log attached).\n\nDo you have any ideas how to fix this ?\n\nThanks.\n\njava.io.FileNotFoundException: File does not exist: hdfs://hadoop1.sindice.net:54310/hbase/page-repository/1105668475/field/mapfiles/5122893264992435570/data\n at org.apache.hadoop.dfs.DistributedFileSystem.getFileStatus(DistributedFileSystem.java:369)\n at org.apache.hadoop.hbase.regionserver.HStoreFile.length(HStoreFile.java:464)\n at org.apache.hadoop.hbase.regionserver.HStore.loadHStoreFiles(HStore.java:409)\n at org.apache.hadoop.hbase.regionserver.HStore.(HStore.java:236)\n at org.apache.hadoop.hbase.regionserver.HRegion.instantiateHStore(HRegion.java:1575)\n at org.apache.hadoop.hbase.regionserver.HRegion.(HRegion.java:451)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.instantiateRegion(HRegionServer.java:901)\n at org.apache.hadoop.hbase.regionserver.HRegionServer.openRegion(HRegionServer.java:876)\n at org.apache.hadoop.hbase.regionserver.HRegionServer$Worker.run(HRegionServer.java:816)\n at java.lang.Thread.run(Thread.java:619)\n\n{code}\n\nI'd thought we'd added handling of this kind of event post HDFS crashings but looking in code, it seems that TRUNK may not have been fixed up properly.","issue_id":"12400852","key":"HBASE-766","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2008-07-31T23:01:57.000+0000","role":"fixed_distractor","summary":"FileNotFoundException trying to load HStoreFile 'data'"} {"case_id":"12646931","cluster":"DISTRACTOR-HBASE-8521","comments":[{"body":"@Jonathan:\nAre you able to write a unit test exhibiting this behavior ?\n\nThanks","created":"2013-05-10T02:00:11.489+0000"},{"body":"@Ted:\nSeemingly, I can't. It seems the mini cluster that HBaseTestingUtility creates has a different behavior than an actual cluster. What I will do, however, is upload a diff of the test that I wrote to show the behavior that I tried to create, and then I'll also attach the files necessary to recreate on a real cluster.","created":"2013-05-10T04:02:47.984+0000"},{"body":"@Ted he opened this because of me. It's a pretty tough thing to repro.\n\nBulk load files (on trunk) are sorted by [see StoreFile#Comparator]:\n\nSeqID\nSize (reversed)\nTimeStamp\nfilename\n\n\nSo bulkoad files that are larger than the bulkloaded files that come before them can overwrite key values with the same timestamp. We should adopt what 0.89-fb does. Assign seqid to bulk loaded files at bulk load time, by renaming the files.\n\nThis should mean that bulk load files are treated in much the same way as flushed/compacted files.","created":"2013-05-10T04:10:34.181+0000"},{"body":"There are two directories in this tarball: familyDir1, and familyDir2. Each contains a single HFile, and each of them has one cell of data in them.\n\nThe table was created as:\ncreate 'test', {NAME => 'myfam', VERSIONS => 100000, TTL => 1000000000}\n\nIn familyDir1, the HFile's cell contains the value \"oldVal\" for myfam:myqual.\n\nIn familyDir2, the HFile's cell contains the value \"newVal\" for myfam:myqual.","created":"2013-05-10T04:13:50.580+0000"},{"body":"Basically, my process for reproducing is this:\n\n{noformat}\nhadoop fs -put familyDir1 familyDir1\nhadoop fs -put familyDir2 familyDir2\n\ncreate 'test', {NAME => 'myfam', VERSIONS => 100000, TTL => 1000000000}\n\nhbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles hdfs://localhost:8020/user/natty/familyDir1 test\nhbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles hdfs://localhost:8020/user/natty/familyDir2 test\n{noformat}\n\nThe result after this set of operations is:\n\n{noformat}\n1.9.2p320 :001 > scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n1 row(s) in 0.6260 seconds\n{noformat}\n\nIf I only load familyDir2, the output is this:\n\n{noformat}\n1.9.2p320 :001 > scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=newVal \n1 row(s) in 0.5930 seconds\n{noformat}","created":"2013-05-10T04:19:28.383+0000"},{"body":"Here's a diff of the test I wrote. You'll have to forgive the quick-and-dirtiness of it, but it gets the point across. This test passes, despite the real cluster failing the same functional test.\n\nAlso, with full disclosure, this diff was made against 0.92.1 included in CDH4.1.4, so if it doesn't nicely line up with Apache source, I'll go and diff it against the real source code.","created":"2013-05-10T04:23:28.681+0000"},{"body":"bq. We should adopt what 0.89-fb does. Assign seqid to bulk loaded files at bulk load time, by renaming the files.\n+1\n\n@Jonathan:\nThanks for outlining the procedure for reproducing the issue.","created":"2013-05-10T04:26:04.090+0000"},{"body":"According to HBASE-6630, 0.89-fb's fix was ported to 0.95. Any chance it is reasonable to backport to 0.92 or 0.94?","created":"2013-05-14T18:06:25.081+0000"},{"body":"[~lhofhansl]:\nWhat's your opinion on the backport ?","created":"2013-05-14T18:11:18.481+0000"},{"body":"The patch in HBASE-6630 is easy enough to understand (and small too when disregarding the protobuf boilerplate).","created":"2013-05-15T03:20:33.948+0000"},{"body":"+0 from me (seems like a rare use case to me).\n\nIt seems like Jeff and Jonathan need that, so let's get it in.\n","created":"2013-05-15T05:29:00.487+0000"},{"body":"Looks like the following new method needs to be added to HRegionInterface (note the 3rd parameter):\n{code}\n public boolean bulkLoadHFiles(List> familyPaths, byte[] regionName, assignSeqIds)\n throws IOException;\n{code}\nThis means that if some region servers use 0.94.7 jars, the call in LoadIncrementalHFiles utilizing the new method would fail.\n\nFallback to existing, two argument, method can be used in above scenario.\n\nI want to other people's opinion on the compatibility issue.","created":"2013-05-15T16:30:22.748+0000"},{"body":"FYI, I have backported HBASE-6630 to 0.94 and used also some of 0.89fb diff too.\n\nI will run some tests locally with the data attached to this defect and try to reproduce what [~natty] get.","created":"2013-05-16T11:34:20.731+0000"},{"body":"So. First try attached.\n\nPassed all the tests successfuly. (3 tests failed because of compilation issues not related to this patch. Servlets, etc.)\n{code}\nhbase(main):001:0> create 'test', {NAME => 'myfam', VERSIONS => 100000, TTL => 1000000000}\n{code}\n\n{code}\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir1 test\n{code}\n\n{code}\nhbase(main):001:0> scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n1 row(s) in 0.4190 seconds\n{code}\n\n{code}\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir2 test\n{code}\n\n{code}\nhbase(main):001:0> scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=newVal \n1 row(s) in 0.4330 seconds\n{code}\n\nSo as said above, there is come incompatibility between he version.\n\nOne thing that I can change is to not pass the parameter if it's false so new clients can still call old servers.\n\nOpen to comments.","created":"2013-05-16T12:16:46.632+0000"},{"body":"Small update.\n\nI have defaulted the assignSeqIds to false in LoadIncrementalHFiles so if no additionnal parameter is used, it will behave the same way as before.\n\nAlso, if assignSeqIds is false, I'm using the previous existing method signatures in the serverCallable to make sure new client -> old server is still compatible.\n\nIf someone specificaly turn assignSeqIds to true, then they are most probably aware of this new feature and will most probably also know that they need to upgrade the server side too.\n\nWhat I can add is a message in the console when this is turned to true to informe that this is working only with a server >= 0.94.8.\n\nAgain, here are the tests results of this change:\n{code}\nTests in error: \n testBasic(org.apache.hadoop.hbase.regionserver.TestRSStatusServlet): Unresolved compilation problems: (..)\n testWithRegions(org.apache.hadoop.hbase.regionserver.TestRSStatusServlet): Unresolved compilation problems: (..)\n testGetEmpty(org.apache.hadoop.hbase.avro.TestAvroUtil): Unresolved compilation problems: (..)\n\nTests run: 682, Failures: 0, Errors: 3, Skipped: 0\n{code}\nI have compilation errors in my repository for thos 2 classes, not related to the changes I made.\n\n{code}\nbin/hbase -Dhbase.mapreduce.bulkload.assign.sequenceNumbers=true org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir1 test\n\nhbase(main):001:0> scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n1 row(s) in 0.3860 seconds\n\n\nbin/hbase -Dhbase.mapreduce.bulkload.assign.sequenceNumbers=true org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir2 test\n\n\nhbase(main):001:0> scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=newVal \n1 row(s) in 0.3800 seconds\n{code}\n\n\nAnd without the parameter...\n\n{code}\necho \"create 'test', {NAME => 'myfam', VERSIONS => 100000, TTL => 1000000000}\" | bin/hbase shell\n\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir1 test\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir2 test\n\necho \"scan 'test'\" | bin/hbase shell\nHBase Shell; enter 'help' for list of supported commands.\nType \"exit\" to leave the HBase Shell\nVersion 0.94.8-SNAPSHOT, r, Thu May 16 08:43:22 EDT 2013\n\nscan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n{code}\n","created":"2013-05-16T12:48:42.673+0000"},{"body":"I just tried with a new client including this fix to bulkload to an existing cluster running 0.94.7 using the usecase Jonathan had provide above.\n\n{code}\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles hdfs://node3:9000/user/hbase/familyDir1 test\n\n echo \"scan 'test'\" | bin/hbase shell\nHBase Shell; enter 'help' for list of supported commands.\nType \"exit\" to leave the HBase Shell\nVersion 0.94.8-SNAPSHOT, r, Thu May 16 08:43:22 EDT 2013\n\nscan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n1 row(s) in 0.7900 seconds\n\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles hdfs://node3:9000/user/hbase/familyDir2 test\n\necho \"scan 'test'\" | bin/hbase shell\nHBase Shell; enter 'help' for list of supported commands.\nType \"exit\" to leave the HBase Shell\nVersion 0.94.8-SNAPSHOT, r, Thu May 16 08:43:22 EDT 2013\n\nscan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n1 row(s) in 0.6810 seconds\n{code}\n\nThings are consistent and working as before, not any error is displayed.\n\nFull log:\n{code}\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles hdfs://node3:9000/user/hbase/familyDir2 test\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:zookeeper.version=3.4.5-1392090, built on 09/30/2012 17:52 GMT\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:host.name=cloudera.distparser.com\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.version=1.6.0_45\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.vendor=Sun Microsystems Inc.\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.home=/usr/local/jdk1.6.0_45/jre\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.class.path=/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT/bin/../conf:/usr/local/jdk1.6.0_45//lib/tools.jar:/home/jmspaggiari/.m2/repository/asm/asm/3.1/asm-3.1.jar:/home/jmspaggiari/.m2/repository/com/github/stephenc/high-scale-lib/high-scale-lib/1.1.1/high-scale-lib-1.1.1.jar:/home/jmspaggiari/.m2/repository/com/google/code/findbugs/jsr305/1.3.9/jsr305-1.3.9.jar:/home/jmspaggiari/.m2/repository/com/google/guava/guava/11.0.2/guava-11.0.2.jar:/home/jmspaggiari/.m2/repository/com/google/protobuf/protobuf-java/2.4.0a/protobuf-java-2.4.0a.jar:/home/jmspaggiari/.m2/repository/com/sun/jersey/jersey-core/1.8/jersey-core-1.8.jar:/home/jmspaggiari/.m2/repository/com/sun/jersey/jersey-json/1.8/jersey-json-1.8.jar:/home/jmspaggiari/.m2/repository/com/sun/jersey/jersey-server/1.8/jersey-server-1.8.jar:/home/jmspaggiari/.m2/repository/com/sun/xml/bind/jaxb-impl/2.2.3-1/jaxb-impl-2.2.3-1.jar:/home/jmspaggiari/.m2/repository/com/yammer/metrics/metrics-core/2.1.2/metrics-core-2.1.2.jar:/home/jmspaggiari/.m2/repository/commons-beanutils/commons-beanutils/1.7.0/commons-beanutils-1.7.0.jar:/home/jmspaggiari/.m2/repository/commons-beanutils/commons-beanutils-core/1.8.0/commons-beanutils-core-1.8.0.jar:/home/jmspaggiari/.m2/repository/commons-cli/commons-cli/1.2/commons-cli-1.2.jar:/home/jmspaggiari/.m2/repository/commons-codec/commons-codec/1.4/commons-codec-1.4.jar:/home/jmspaggiari/.m2/repository/commons-collections/commons-collections/3.2.1/commons-collections-3.2.1.jar:/home/jmspaggiari/.m2/repository/commons-configuration/commons-configuration/1.6/commons-configuration-1.6.jar:/home/jmspaggiari/.m2/repository/commons-digester/commons-digester/1.8/commons-digester-1.8.jar:/home/jmspaggiari/.m2/repository/commons-el/commons-el/1.0/commons-el-1.0.jar:/home/jmspaggiari/.m2/repository/commons-httpclient/commons-httpclient/3.1/commons-httpclient-3.1.jar:/home/jmspaggiari/.m2/repository/commons-io/commons-io/2.1/commons-io-2.1.jar:/home/jmspaggiari/.m2/repository/commons-lang/commons-lang/2.5/commons-lang-2.5.jar:/home/jmspaggiari/.m2/repository/commons-logging/commons-logging/1.1.1/commons-logging-1.1.1.jar:/home/jmspaggiari/.m2/repository/commons-net/commons-net/1.4.1/commons-net-1.4.1.jar:/home/jmspaggiari/.m2/repository/javax/activation/activation/1.1/activation-1.1.jar:/home/jmspaggiari/.m2/repository/javax/xml/bind/jaxb-api/2.1/jaxb-api-2.1.jar:/home/jmspaggiari/.m2/repository/junit/junit/4.10-HBASE-1/junit-4.10-HBASE-1.jar:/home/jmspaggiari/.m2/repository/log4j/log4j/1.2.16/log4j-1.2.16.jar:/home/jmspaggiari/.m2/repository/org/apache/avro/avro/1.5.3/avro-1.5.3.jar:/home/jmspaggiari/.m2/repository/org/apache/avro/avro-ipc/1.5.3/avro-ipc-1.5.3.jar:/home/jmspaggiari/.m2/repository/org/apache/commons/commons-math/2.1/commons-math-2.1.jar:/home/jmspaggiari/.m2/repository/org/apache/ftpserver/ftplet-api/1.0.0/ftplet-api-1.0.0.jar:/home/jmspaggiari/.m2/repository/org/apache/ftpserver/ftpserver-core/1.0.0/ftpserver-core-1.0.0.jar:/home/jmspaggiari/.m2/repository/org/apache/ftpserver/ftpserver-deprecated/1.0.0-M2/ftpserver-deprecated-1.0.0-M2.jar:/home/jmspaggiari/.m2/repository/org/apache/hadoop/hadoop-core/1.0.4/hadoop-core-1.0.4.jar:/home/jmspaggiari/.m2/repository/org/apache/hadoop/hadoop-test/1.0.4/hadoop-test-1.0.4.jar:/home/jmspaggiari/.m2/repository/org/apache/httpcomponents/httpclient/4.1.2/httpclient-4.1.2.jar:/home/jmspaggiari/.m2/repository/org/apache/httpcomponents/httpcore/4.1.3/httpcore-4.1.3.jar:/home/jmspaggiari/.m2/repository/org/apache/mina/mina-core/2.0.0-M5/mina-core-2.0.0-M5.jar:/home/jmspaggiari/.m2/repository/org/apache/thrift/libthrift/0.8.0/libthrift-0.8.0.jar:/home/jmspaggiari/.m2/repository/org/apache/velocity/velocity/1.7/velocity-1.7.jar:/home/jmspaggiari/.m2/repository/org/apache/zookeeper/zookeeper/3.4.5/zookeeper-3.4.5.jar:/home/jmspaggiari/.m2/repository/org/codehaus/jackson/jackson-core-asl/1.8.8/jackson-core-asl-1.8.8.jar:/home/jmspaggiari/.m2/repository/org/codehaus/jackson/jackson-jaxrs/1.8.8/jackson-jaxrs-1.8.8.jar:/home/jmspaggiari/.m2/repository/org/codehaus/jackson/jackson-mapper-asl/1.8.8/jackson-mapper-asl-1.8.8.jar:/home/jmspaggiari/.m2/repository/org/codehaus/jackson/jackson-xc/1.8.8/jackson-xc-1.8.8.jar:/home/jmspaggiari/.m2/repository/org/codehaus/jettison/jettison/1.1/jettison-1.1.jar:/home/jmspaggiari/.m2/repository/org/eclipse/jdt/core/3.1.1/core-3.1.1.jar:/home/jmspaggiari/.m2/repository/org/jamon/jamon-runtime/2.3.1/jamon-runtime-2.3.1.jar:/home/jmspaggiari/.m2/repository/org/jboss/netty/netty/3.2.4.Final/netty-3.2.4.Final.jar:/home/jmspaggiari/.m2/repository/org/jruby/jruby-complete/1.6.5/jruby-complete-1.6.5.jar:/home/jmspaggiari/.m2/repository/org/mockito/mockito-all/1.8.5/mockito-all-1.8.5.jar:/home/jmspaggiari/.m2/repository/org/mortbay/jetty/jetty/6.1.26/jetty-6.1.26.jar:/home/jmspaggiari/.m2/repository/org/mortbay/jetty/jetty-util/6.1.26/jetty-util-6.1.26.jar:/home/jmspaggiari/.m2/repository/org/mortbay/jetty/jsp-2.1/6.1.14/jsp-2.1-6.1.14.jar:/home/jmspaggiari/.m2/repository/org/mortbay/jetty/jsp-api-2.1/6.1.14/jsp-api-2.1-6.1.14.jar:/home/jmspaggiari/.m2/repository/org/mortbay/jetty/servlet-api-2.5/6.1.14/servlet-api-2.5-6.1.14.jar:/home/jmspaggiari/.m2/repository/org/slf4j/slf4j-api/1.4.3/slf4j-api-1.4.3.jar:/home/jmspaggiari/.m2/repository/org/slf4j/slf4j-log4j12/1.4.3/slf4j-log4j12-1.4.3.jar:/home/jmspaggiari/.m2/repository/org/xerial/snappy/snappy-java/1.0.3.2/snappy-java-1.0.3.2.jar:/home/jmspaggiari/.m2/repository/stax/stax-api/1.0.1/stax-api-1.0.1.jar:/home/jmspaggiari/.m2/repository/tomcat/jasper-compiler/5.5.23/jasper-compiler-5.5.23.jar:/home/jmspaggiari/.m2/repository/tomcat/jasper-runtime/5.5.23/jasper-runtime-5.5.23.jar:/home/jmspaggiari/.m2/repository/xmlenc/xmlenc/0.52/xmlenc-0.52.jar:/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT/bin/../target/classes:/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT/bin/../target/test-classes:/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT/bin/../target:/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT/bin/../lib/*.jar:\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.library.path=/usr/local/jdk1.6.0_45/jre/lib/amd64/server:/usr/local/jdk1.6.0_45/jre/lib/amd64:/usr/local/jdk1.6.0_45/jre/../lib/amd64:/usr/java/packages/lib/amd64:/usr/lib64:/lib64:/lib:/usr/lib\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.io.tmpdir=/tmp\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.compiler=\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:os.name=Linux\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:os.arch=amd64\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:os.version=3.2.0-4-amd64\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:user.name=jmspaggiari\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:user.home=/home/jmspaggiari\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:user.dir=/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Initiating client connection, connectString=latitude:2181,cube:2181,node3:2181 sessionTimeout=180000 watcher=hconnection\n13/05/16 14:38:15 INFO zookeeper.RecoverableZooKeeper: The identifier of this process is 16902@cloudera\n13/05/16 14:38:15 INFO zookeeper.ClientCnxn: Opening socket connection to server latitude/192.168.23.4:2181. Will not attempt to authenticate using SASL (Unable to locate a login configuration)\n13/05/16 14:38:15 INFO zookeeper.ClientCnxn: Socket connection established to latitude/192.168.23.4:2181, initiating session\n13/05/16 14:38:15 INFO zookeeper.ClientCnxn: Session establishment complete on server latitude/192.168.23.4:2181, sessionid = 0x23eadc5ddeb008f, negotiated timeout = 40000\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Initiating client connection, connectString=latitude:2181,cube:2181,node3:2181 sessionTimeout=180000 watcher=catalogtracker-on-org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation@79ee2c2c\n13/05/16 14:38:15 INFO zookeeper.RecoverableZooKeeper: The identifier of this process is 16902@cloudera\n13/05/16 14:38:15 INFO zookeeper.ClientCnxn: Opening socket connection to server latitude/192.168.23.4:2181. Will not attempt to authenticate using SASL (Unable to locate a login configuration)\n13/05/16 14:38:15 DEBUG catalog.CatalogTracker: Starting catalog tracker org.apache.hadoop.hbase.catalog.CatalogTracker@77ff92f5\n13/05/16 14:38:15 INFO zookeeper.ClientCnxn: Socket connection established to latitude/192.168.23.4:2181, initiating session\n13/05/16 14:38:16 INFO zookeeper.ClientCnxn: Session establishment complete on server latitude/192.168.23.4:2181, sessionid = 0x23eadc5ddeb0090, negotiated timeout = 40000\n13/05/16 14:38:16 DEBUG client.HConnectionManager$HConnectionImplementation: Looked up root region location, connection=org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation@79ee2c2c; serverName=node5,60020,1368669066647\n13/05/16 14:38:16 DEBUG client.HConnectionManager$HConnectionImplementation: Cached location for .META.,,1.1028785192 is node6:60020\n13/05/16 14:38:16 DEBUG client.ClientScanner: Creating scanner over .META. starting at key 'test,,'\n13/05/16 14:38:16 DEBUG client.ClientScanner: Advancing internal scanner to startKey at 'test,,'\n13/05/16 14:38:16 DEBUG catalog.CatalogTracker: Stopping catalog tracker org.apache.hadoop.hbase.catalog.CatalogTracker@77ff92f5\n13/05/16 14:38:16 INFO zookeeper.ZooKeeper: Session: 0x23eadc5ddeb0090 closed\n13/05/16 14:38:16 INFO zookeeper.ClientCnxn: EventThread shut down\n13/05/16 14:38:16 DEBUG client.MetaScanner: Scanning .META. starting at row=test,,00000000000000 for max=10 rows using org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation@79ee2c2c\n13/05/16 14:38:16 DEBUG client.HConnectionManager$HConnectionImplementation: Cached location for test,,1368729286037.32038a4d760aa9303643e2985dcd29a5. is node2:60020\n13/05/16 14:38:16 DEBUG client.MetaScanner: Scanning .META. starting at row= for max=2147483647 rows using org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation@79ee2c2c\n13/05/16 14:38:17 DEBUG client.MetaScanner: Scanning .META. starting at row=test,,00000000000000 for max=2147483647 rows using org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation@79ee2c2c\n13/05/16 14:38:17 INFO hfile.CacheConfig: Allocating LruBlockCache with maximum size 246.9m\n13/05/16 14:38:17 INFO util.ChecksumType: Checksum can use java.util.zip.CRC32\n13/05/16 14:38:17 INFO mapreduce.LoadIncrementalHFiles: Trying to load hfile=hdfs://node3:9000/user/hbase/familyDir2/myfam/hfile_1 first=aaaa last=aaaa\n13/05/16 14:38:18 DEBUG mapreduce.LoadIncrementalHFiles: Going to connect to server region=test,,1368729286037.32038a4d760aa9303643e2985dcd29a5., hostname=node2, port=60020 for row \n{code}","created":"2013-05-16T18:44:09.126+0000"},{"body":"Any chance to have someone looking at this one?","created":"2013-06-12T21:15:06.148+0000"},{"body":"patch looks reasonable esp. given it's a backport.\nIt has long lines.\nThese ifs:\n{code}\n+ if (assignSeqIds) {\n+ success = secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken, assignSeqIds);\n+ } else {\n+ success = secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken);\n+ }\n{code}\nseem excessive, just the first call is enough, right?","created":"2013-06-18T22:21:21.803+0000"},{"body":"secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken) calls secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken, false) so if assignSeqIds is false, secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken) and secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken, assignSeqIds) are going to be the same call, therefore the if is not required.\n\nSame for \n{code}\n if (assignSeqIds) {\n success = secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken, assignSeqIds);\n } else {\n success = secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken);\n }\n{code}\n\nI will change that and re-base the patch since it's not working anymore with current 0.94 branch...\n\nThanks for looking at it.","created":"2013-06-19T14:17:40.407+0000"},{"body":"Rebased patch attached. Some modifications required because minSeqId got removed by HBASE-7210. Also applied Ted's comment. I'm not testing with the data attached to this ticket.","created":"2013-08-15T12:12:52.041+0000"},{"body":"This jira will need a title that accurately describes what's being done here.\n\nOn the patch:\n\n - There are 11 \"== Durability.USE_DEFAULT\", can we just have a method somewhere (like in Mutation) that does it and is named \"isDefaultDurability\"?\n - LoadIncrementalHFiles.assignSeqIds should be final.\n - Since bulk loaded files can have sequence ids, we should print it out. StoreFile.toStringDetailed is a candidate for that change, there might be more\n - What's up with the commented out code in HRegion?\n - Is it passing all the unit tests? A trunk version would here to get some Hadoop QA love.\n\nI tested the patch to see if Jonathan's original use case is covered and it looks like it is. I also did some mixed workloads of normal Puts and bulk loaded files.","created":"2013-08-15T21:55:15.563+0000"},{"body":"Hi [~jdcryans], thanks for looking at it.\n\n{quote}\n There are 11 \"== Durability.USE_DEFAULT\", can we just have a method somewhere (like in Mutation) that does it and is named \"isDefaultDurability\"?\n{quote}\nThis is because setWriteToWAL is now deprecated. If I set something locally to do the same thing as what setWriteToWAL was doing, then it's not very clean since we are just bypassing the deprecation warning. Other option is to keep setWriteToWAL calls and add @suppressWarning to avoid them and keep setWriteToWAL? \n\n{quote}\nLoadIncrementalHFiles.assignSeqIds should be final.\n{quote}\nI don't think we can. If I put it final, how are you going to modify it on the constructor? (assignSeqIds = conf.getBoolean(ASSIGN_SEQ_IDS, true))\n\n{quote}\nSince bulk loaded files can have sequence ids, we should print it out. StoreFile.toStringDetailed is a candidate for that change, there might be more\n{quote}\nStoreFile.toStringDetailed already display the sequenceId under certain conditions. I will remove the condition to make sure we always have this information.\n\n{quote}\nWhat's up with the commented out code in HRegion?\n{quote}\nWow, this should not be there. At all! Removed!\n\n{quote}\nIs it passing all the unit tests? A trunk version would here to get some Hadoop QA love.\n{quote}\nYes, it's passing unit tests. I will post the results before EOD. Also, there is no trunk version because it's a backport of HBASE-6630 which is already doing the same thing in trunk.\n\n\nSo, waiting for your recommendation regarding Durability.USE_DEFAULT tests and I will post an updated version.","created":"2013-08-16T18:53:50.083+0000"},{"body":"Update version attached.\n1) final added. Was not working on trunk, but working fine on 0.94.\n2) Removed the Durability modifications since they are not related to this patch. Was only deprecated calls cleanup.\n\nShould be nicer now.","created":"2013-08-20T19:45:57.455+0000"},{"body":"Hum, seems that I pushed the wrong file...","created":"2013-08-24T01:19:15.323+0000"},{"body":"Removed commented code.","created":"2013-08-24T01:24:48.882+0000"},{"body":"[~lhofhansl] any chance to look at that? Want it into 0.94? This patch is already running on a production cluster.","created":"2013-09-13T21:26:26.843+0000"},{"body":"Still +0 on this.\n\nAny strong sponsors? If so, let's get it in.","created":"2013-10-02T00:00:35.443+0000"},{"body":"I'll jump in late to the discussion to add a personal story. \n\nWe're very much relying on this patch since we bulk load everything and our use-case depends on this. We're running with this patch at the moment and we're hoping not to lose it when upgrading. We put automated testing of our feature that relies on this and we spent a lot of time looking at the internals (fun times with the Hfile tool) to check that all was as expected. While the test doesn't provide a full guarantee that the behavior is as expected as opposed to still being a combination of non-deterministic behavior and luck, this runs every day and each day that the test doesn't fail increases our confidence. \n\nI'm very much +1. If there's one patch we ever needed from hbase, this is it. ","created":"2013-10-02T00:24:33.607+0000"},{"body":"Thanks for chiming in Alexandre... I'll count that a \"strong sponsor\" as per above.\nIf there are no objections, I will commit this to 0.94.13 tomorrow.\n","created":"2013-10-02T05:33:06.946+0000"},{"body":"Thanks Lars, I appreciate it. I really do. ","created":"2013-10-02T05:37:54.453+0000"},{"body":"Cool, that's all good news too! Thanks guys. One less on my radar.","created":"2013-10-02T12:59:57.468+0000"},{"body":"[~lhofhansl] do you have time to cut a patch for 0.96 before the next RC? (cc [~stack])","created":"2013-10-02T17:45:38.587+0000"},{"body":"^^^ or [~jmspaggi] :)","created":"2013-10-02T17:46:14.685+0000"},{"body":"It's already there in HBASE-6630... Only 0.94 was missing it. So we should be good.","created":"2013-10-02T18:17:32.794+0000"},{"body":"ACK. Thanks JM.","created":"2013-10-02T18:27:06.130+0000"},{"body":"Committed to the 0.94 branch. Thanks for the patch.","created":"2013-10-02T19:06:41.333+0000"},{"body":"TestLoadIncrementalHFilesSplitRecovery.testBulkLoadPhaseFailure failed... Looks suspicious.","created":"2013-10-02T22:08:46.158+0000"},{"body":"I agree. Let me run that locally few multiple to validate...","created":"2013-10-02T22:11:32.622+0000"},{"body":"Also, looking at v4 now. It adds\n{code}\nprivate void verifyAssignedSequenceNumber(String testName, byte hfileRanges, boolean nonZero)\n{code}\nBut it is never actually called from anywhere.\n","created":"2013-10-02T22:16:05.594+0000"},{"body":"Definitely passes without the patch and fail with it.","created":"2013-10-02T22:18:34.122+0000"},{"body":"Failed also on 0.94.10+HBASE-8521 :( Something might have changed since the last time I merged it. Please revert, I will look at the issue.","created":"2013-10-02T22:21:37.262+0000"},{"body":"verifyAssignedSequenceNumber has been removed from trunk because it's never called. I took it from 0.89. I will remove it. Looking at the test failure now.","created":"2013-10-02T22:27:45.271+0000"},{"body":"Ok. Looking deeper, I'm not sure that it's really an issue.\n\nThe test which failed:\n{code}\n /**\n * Test that shows that exception thrown from the RS side will result in an\n * exception on the LIHFile client.\n */\n{code}\n\nRS is not throwing an exception. But nothing related to this patch. But still, occurs only with this patch activated. So there might be some side effect. I will start to roll it back locally line by line to figure what the issue is...","created":"2013-10-02T22:38:47.874+0000"},{"body":"Found it.\n\nDon't revert. It's related to the test, not to the code. Patch and detailed to come in a minute.","created":"2013-10-02T23:47:00.263+0000"},{"body":"I knew I could wait a few minutes :)","created":"2013-10-02T23:53:26.371+0000"},{"body":"So here is the culprit.\n\nThere is now an extra parameter to the bulkLoadHFiles call where we need to specify the assignSeqNum boolean. However, the mocked connection was still configured to fail on the method call without this parameter (which is not the one being called).\n\nThis addendum is to update the mocked connection configuration.\n\nPassed the test locally. Sorry for this.","created":"2013-10-03T00:00:02.063+0000"},{"body":"Committed. Thanks JM!","created":"2013-10-03T00:10:17.770+0000"}],"conversations":[{"body":"Let's say you have a pre-built HFile that contains a cell:\n\n('rowkey1', 'family1', 'qual1', 1234L, 'value1')\n\nWe bulk load this first HFile. Now, let's create a second HFile that contains a cell that overwrites the first:\n\n('rowkey1', 'family1', 'qual1', 1234L, 'value2')\n\nThat gets bulk loaded into the table, but the value that HBase bubbles up is still 'value1'.\n\nIt seems that there's no way to overwrite a cell for a particular timestamp without an explicit put operation. This seems to be the case even after minor and major compactions happen.\n\nMy guess is that this is pretty closely related to the sequence number work being done on the compaction algorithm via HBASE-7842, but I'm not sure if one of would fix the other.","from":"reporter","subject":"Cells cannot be overwritten with bulk loaded HFiles"},{"body":"@Jonathan:\nAre you able to write a unit test exhibiting this behavior ?\n\nThanks","from":"developer"},{"body":"@Ted:\nSeemingly, I can't. It seems the mini cluster that HBaseTestingUtility creates has a different behavior than an actual cluster. What I will do, however, is upload a diff of the test that I wrote to show the behavior that I tried to create, and then I'll also attach the files necessary to recreate on a real cluster.","from":"developer"},{"body":"@Ted he opened this because of me. It's a pretty tough thing to repro.\n\nBulk load files (on trunk) are sorted by [see StoreFile#Comparator]:\n\nSeqID\nSize (reversed)\nTimeStamp\nfilename\n\n\nSo bulkoad files that are larger than the bulkloaded files that come before them can overwrite key values with the same timestamp. We should adopt what 0.89-fb does. Assign seqid to bulk loaded files at bulk load time, by renaming the files.\n\nThis should mean that bulk load files are treated in much the same way as flushed/compacted files.","from":"developer"},{"body":"There are two directories in this tarball: familyDir1, and familyDir2. Each contains a single HFile, and each of them has one cell of data in them.\n\nThe table was created as:\ncreate 'test', {NAME => 'myfam', VERSIONS => 100000, TTL => 1000000000}\n\nIn familyDir1, the HFile's cell contains the value \"oldVal\" for myfam:myqual.\n\nIn familyDir2, the HFile's cell contains the value \"newVal\" for myfam:myqual.","from":"developer"},{"body":"Basically, my process for reproducing is this:\n\n{noformat}\nhadoop fs -put familyDir1 familyDir1\nhadoop fs -put familyDir2 familyDir2\n\ncreate 'test', {NAME => 'myfam', VERSIONS => 100000, TTL => 1000000000}\n\nhbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles hdfs://localhost:8020/user/natty/familyDir1 test\nhbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles hdfs://localhost:8020/user/natty/familyDir2 test\n{noformat}\n\nThe result after this set of operations is:\n\n{noformat}\n1.9.2p320 :001 > scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n1 row(s) in 0.6260 seconds\n{noformat}\n\nIf I only load familyDir2, the output is this:\n\n{noformat}\n1.9.2p320 :001 > scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=newVal \n1 row(s) in 0.5930 seconds\n{noformat}","from":"developer"},{"body":"Here's a diff of the test I wrote. You'll have to forgive the quick-and-dirtiness of it, but it gets the point across. This test passes, despite the real cluster failing the same functional test.\n\nAlso, with full disclosure, this diff was made against 0.92.1 included in CDH4.1.4, so if it doesn't nicely line up with Apache source, I'll go and diff it against the real source code.","from":"developer"},{"body":"bq. We should adopt what 0.89-fb does. Assign seqid to bulk loaded files at bulk load time, by renaming the files.\n+1\n\n@Jonathan:\nThanks for outlining the procedure for reproducing the issue.","from":"developer"},{"body":"According to HBASE-6630, 0.89-fb's fix was ported to 0.95. Any chance it is reasonable to backport to 0.92 or 0.94?","from":"developer"},{"body":"[~lhofhansl]:\nWhat's your opinion on the backport ?","from":"developer"},{"body":"The patch in HBASE-6630 is easy enough to understand (and small too when disregarding the protobuf boilerplate).","from":"developer"},{"body":"+0 from me (seems like a rare use case to me).\n\nIt seems like Jeff and Jonathan need that, so let's get it in.\n","from":"developer"},{"body":"Looks like the following new method needs to be added to HRegionInterface (note the 3rd parameter):\n{code}\n public boolean bulkLoadHFiles(List> familyPaths, byte[] regionName, assignSeqIds)\n throws IOException;\n{code}\nThis means that if some region servers use 0.94.7 jars, the call in LoadIncrementalHFiles utilizing the new method would fail.\n\nFallback to existing, two argument, method can be used in above scenario.\n\nI want to other people's opinion on the compatibility issue.","from":"developer"},{"body":"FYI, I have backported HBASE-6630 to 0.94 and used also some of 0.89fb diff too.\n\nI will run some tests locally with the data attached to this defect and try to reproduce what [~natty] get.","from":"developer"},{"body":"So. First try attached.\n\nPassed all the tests successfuly. (3 tests failed because of compilation issues not related to this patch. Servlets, etc.)\n{code}\nhbase(main):001:0> create 'test', {NAME => 'myfam', VERSIONS => 100000, TTL => 1000000000}\n{code}\n\n{code}\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir1 test\n{code}\n\n{code}\nhbase(main):001:0> scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n1 row(s) in 0.4190 seconds\n{code}\n\n{code}\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir2 test\n{code}\n\n{code}\nhbase(main):001:0> scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=newVal \n1 row(s) in 0.4330 seconds\n{code}\n\nSo as said above, there is come incompatibility between he version.\n\nOne thing that I can change is to not pass the parameter if it's false so new clients can still call old servers.\n\nOpen to comments.","from":"developer"},{"body":"Small update.\n\nI have defaulted the assignSeqIds to false in LoadIncrementalHFiles so if no additionnal parameter is used, it will behave the same way as before.\n\nAlso, if assignSeqIds is false, I'm using the previous existing method signatures in the serverCallable to make sure new client -> old server is still compatible.\n\nIf someone specificaly turn assignSeqIds to true, then they are most probably aware of this new feature and will most probably also know that they need to upgrade the server side too.\n\nWhat I can add is a message in the console when this is turned to true to informe that this is working only with a server >= 0.94.8.\n\nAgain, here are the tests results of this change:\n{code}\nTests in error: \n testBasic(org.apache.hadoop.hbase.regionserver.TestRSStatusServlet): Unresolved compilation problems: (..)\n testWithRegions(org.apache.hadoop.hbase.regionserver.TestRSStatusServlet): Unresolved compilation problems: (..)\n testGetEmpty(org.apache.hadoop.hbase.avro.TestAvroUtil): Unresolved compilation problems: (..)\n\nTests run: 682, Failures: 0, Errors: 3, Skipped: 0\n{code}\nI have compilation errors in my repository for thos 2 classes, not related to the changes I made.\n\n{code}\nbin/hbase -Dhbase.mapreduce.bulkload.assign.sequenceNumbers=true org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir1 test\n\nhbase(main):001:0> scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n1 row(s) in 0.3860 seconds\n\n\nbin/hbase -Dhbase.mapreduce.bulkload.assign.sequenceNumbers=true org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir2 test\n\n\nhbase(main):001:0> scan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=newVal \n1 row(s) in 0.3800 seconds\n{code}\n\n\nAnd without the parameter...\n\n{code}\necho \"create 'test', {NAME => 'myfam', VERSIONS => 100000, TTL => 1000000000}\" | bin/hbase shell\n\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir1 test\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles familyDir2 test\n\necho \"scan 'test'\" | bin/hbase shell\nHBase Shell; enter 'help' for list of supported commands.\nType \"exit\" to leave the HBase Shell\nVersion 0.94.8-SNAPSHOT, r, Thu May 16 08:43:22 EDT 2013\n\nscan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n{code}\n","from":"developer"},{"body":"I just tried with a new client including this fix to bulkload to an existing cluster running 0.94.7 using the usecase Jonathan had provide above.\n\n{code}\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles hdfs://node3:9000/user/hbase/familyDir1 test\n\n echo \"scan 'test'\" | bin/hbase shell\nHBase Shell; enter 'help' for list of supported commands.\nType \"exit\" to leave the HBase Shell\nVersion 0.94.8-SNAPSHOT, r, Thu May 16 08:43:22 EDT 2013\n\nscan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n1 row(s) in 0.7900 seconds\n\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles hdfs://node3:9000/user/hbase/familyDir2 test\n\necho \"scan 'test'\" | bin/hbase shell\nHBase Shell; enter 'help' for list of supported commands.\nType \"exit\" to leave the HBase Shell\nVersion 0.94.8-SNAPSHOT, r, Thu May 16 08:43:22 EDT 2013\n\nscan 'test'\nROW COLUMN+CELL \n aaaa column=myfam:myqual, timestamp=1368157470713, value=oldVal \n1 row(s) in 0.6810 seconds\n{code}\n\nThings are consistent and working as before, not any error is displayed.\n\nFull log:\n{code}\nbin/hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles hdfs://node3:9000/user/hbase/familyDir2 test\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:zookeeper.version=3.4.5-1392090, built on 09/30/2012 17:52 GMT\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:host.name=cloudera.distparser.com\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.version=1.6.0_45\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.vendor=Sun Microsystems Inc.\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.home=/usr/local/jdk1.6.0_45/jre\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.class.path=/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT/bin/../conf:/usr/local/jdk1.6.0_45//lib/tools.jar:/home/jmspaggiari/.m2/repository/asm/asm/3.1/asm-3.1.jar:/home/jmspaggiari/.m2/repository/com/github/stephenc/high-scale-lib/high-scale-lib/1.1.1/high-scale-lib-1.1.1.jar:/home/jmspaggiari/.m2/repository/com/google/code/findbugs/jsr305/1.3.9/jsr305-1.3.9.jar:/home/jmspaggiari/.m2/repository/com/google/guava/guava/11.0.2/guava-11.0.2.jar:/home/jmspaggiari/.m2/repository/com/google/protobuf/protobuf-java/2.4.0a/protobuf-java-2.4.0a.jar:/home/jmspaggiari/.m2/repository/com/sun/jersey/jersey-core/1.8/jersey-core-1.8.jar:/home/jmspaggiari/.m2/repository/com/sun/jersey/jersey-json/1.8/jersey-json-1.8.jar:/home/jmspaggiari/.m2/repository/com/sun/jersey/jersey-server/1.8/jersey-server-1.8.jar:/home/jmspaggiari/.m2/repository/com/sun/xml/bind/jaxb-impl/2.2.3-1/jaxb-impl-2.2.3-1.jar:/home/jmspaggiari/.m2/repository/com/yammer/metrics/metrics-core/2.1.2/metrics-core-2.1.2.jar:/home/jmspaggiari/.m2/repository/commons-beanutils/commons-beanutils/1.7.0/commons-beanutils-1.7.0.jar:/home/jmspaggiari/.m2/repository/commons-beanutils/commons-beanutils-core/1.8.0/commons-beanutils-core-1.8.0.jar:/home/jmspaggiari/.m2/repository/commons-cli/commons-cli/1.2/commons-cli-1.2.jar:/home/jmspaggiari/.m2/repository/commons-codec/commons-codec/1.4/commons-codec-1.4.jar:/home/jmspaggiari/.m2/repository/commons-collections/commons-collections/3.2.1/commons-collections-3.2.1.jar:/home/jmspaggiari/.m2/repository/commons-configuration/commons-configuration/1.6/commons-configuration-1.6.jar:/home/jmspaggiari/.m2/repository/commons-digester/commons-digester/1.8/commons-digester-1.8.jar:/home/jmspaggiari/.m2/repository/commons-el/commons-el/1.0/commons-el-1.0.jar:/home/jmspaggiari/.m2/repository/commons-httpclient/commons-httpclient/3.1/commons-httpclient-3.1.jar:/home/jmspaggiari/.m2/repository/commons-io/commons-io/2.1/commons-io-2.1.jar:/home/jmspaggiari/.m2/repository/commons-lang/commons-lang/2.5/commons-lang-2.5.jar:/home/jmspaggiari/.m2/repository/commons-logging/commons-logging/1.1.1/commons-logging-1.1.1.jar:/home/jmspaggiari/.m2/repository/commons-net/commons-net/1.4.1/commons-net-1.4.1.jar:/home/jmspaggiari/.m2/repository/javax/activation/activation/1.1/activation-1.1.jar:/home/jmspaggiari/.m2/repository/javax/xml/bind/jaxb-api/2.1/jaxb-api-2.1.jar:/home/jmspaggiari/.m2/repository/junit/junit/4.10-HBASE-1/junit-4.10-HBASE-1.jar:/home/jmspaggiari/.m2/repository/log4j/log4j/1.2.16/log4j-1.2.16.jar:/home/jmspaggiari/.m2/repository/org/apache/avro/avro/1.5.3/avro-1.5.3.jar:/home/jmspaggiari/.m2/repository/org/apache/avro/avro-ipc/1.5.3/avro-ipc-1.5.3.jar:/home/jmspaggiari/.m2/repository/org/apache/commons/commons-math/2.1/commons-math-2.1.jar:/home/jmspaggiari/.m2/repository/org/apache/ftpserver/ftplet-api/1.0.0/ftplet-api-1.0.0.jar:/home/jmspaggiari/.m2/repository/org/apache/ftpserver/ftpserver-core/1.0.0/ftpserver-core-1.0.0.jar:/home/jmspaggiari/.m2/repository/org/apache/ftpserver/ftpserver-deprecated/1.0.0-M2/ftpserver-deprecated-1.0.0-M2.jar:/home/jmspaggiari/.m2/repository/org/apache/hadoop/hadoop-core/1.0.4/hadoop-core-1.0.4.jar:/home/jmspaggiari/.m2/repository/org/apache/hadoop/hadoop-test/1.0.4/hadoop-test-1.0.4.jar:/home/jmspaggiari/.m2/repository/org/apache/httpcomponents/httpclient/4.1.2/httpclient-4.1.2.jar:/home/jmspaggiari/.m2/repository/org/apache/httpcomponents/httpcore/4.1.3/httpcore-4.1.3.jar:/home/jmspaggiari/.m2/repository/org/apache/mina/mina-core/2.0.0-M5/mina-core-2.0.0-M5.jar:/home/jmspaggiari/.m2/repository/org/apache/thrift/libthrift/0.8.0/libthrift-0.8.0.jar:/home/jmspaggiari/.m2/repository/org/apache/velocity/velocity/1.7/velocity-1.7.jar:/home/jmspaggiari/.m2/repository/org/apache/zookeeper/zookeeper/3.4.5/zookeeper-3.4.5.jar:/home/jmspaggiari/.m2/repository/org/codehaus/jackson/jackson-core-asl/1.8.8/jackson-core-asl-1.8.8.jar:/home/jmspaggiari/.m2/repository/org/codehaus/jackson/jackson-jaxrs/1.8.8/jackson-jaxrs-1.8.8.jar:/home/jmspaggiari/.m2/repository/org/codehaus/jackson/jackson-mapper-asl/1.8.8/jackson-mapper-asl-1.8.8.jar:/home/jmspaggiari/.m2/repository/org/codehaus/jackson/jackson-xc/1.8.8/jackson-xc-1.8.8.jar:/home/jmspaggiari/.m2/repository/org/codehaus/jettison/jettison/1.1/jettison-1.1.jar:/home/jmspaggiari/.m2/repository/org/eclipse/jdt/core/3.1.1/core-3.1.1.jar:/home/jmspaggiari/.m2/repository/org/jamon/jamon-runtime/2.3.1/jamon-runtime-2.3.1.jar:/home/jmspaggiari/.m2/repository/org/jboss/netty/netty/3.2.4.Final/netty-3.2.4.Final.jar:/home/jmspaggiari/.m2/repository/org/jruby/jruby-complete/1.6.5/jruby-complete-1.6.5.jar:/home/jmspaggiari/.m2/repository/org/mockito/mockito-all/1.8.5/mockito-all-1.8.5.jar:/home/jmspaggiari/.m2/repository/org/mortbay/jetty/jetty/6.1.26/jetty-6.1.26.jar:/home/jmspaggiari/.m2/repository/org/mortbay/jetty/jetty-util/6.1.26/jetty-util-6.1.26.jar:/home/jmspaggiari/.m2/repository/org/mortbay/jetty/jsp-2.1/6.1.14/jsp-2.1-6.1.14.jar:/home/jmspaggiari/.m2/repository/org/mortbay/jetty/jsp-api-2.1/6.1.14/jsp-api-2.1-6.1.14.jar:/home/jmspaggiari/.m2/repository/org/mortbay/jetty/servlet-api-2.5/6.1.14/servlet-api-2.5-6.1.14.jar:/home/jmspaggiari/.m2/repository/org/slf4j/slf4j-api/1.4.3/slf4j-api-1.4.3.jar:/home/jmspaggiari/.m2/repository/org/slf4j/slf4j-log4j12/1.4.3/slf4j-log4j12-1.4.3.jar:/home/jmspaggiari/.m2/repository/org/xerial/snappy/snappy-java/1.0.3.2/snappy-java-1.0.3.2.jar:/home/jmspaggiari/.m2/repository/stax/stax-api/1.0.1/stax-api-1.0.1.jar:/home/jmspaggiari/.m2/repository/tomcat/jasper-compiler/5.5.23/jasper-compiler-5.5.23.jar:/home/jmspaggiari/.m2/repository/tomcat/jasper-runtime/5.5.23/jasper-runtime-5.5.23.jar:/home/jmspaggiari/.m2/repository/xmlenc/xmlenc/0.52/xmlenc-0.52.jar:/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT/bin/../target/classes:/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT/bin/../target/test-classes:/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT/bin/../target:/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT/bin/../lib/*.jar:\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.library.path=/usr/local/jdk1.6.0_45/jre/lib/amd64/server:/usr/local/jdk1.6.0_45/jre/lib/amd64:/usr/local/jdk1.6.0_45/jre/../lib/amd64:/usr/java/packages/lib/amd64:/usr/lib64:/lib64:/lib:/usr/lib\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.io.tmpdir=/tmp\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:java.compiler=\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:os.name=Linux\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:os.arch=amd64\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:os.version=3.2.0-4-amd64\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:user.name=jmspaggiari\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:user.home=/home/jmspaggiari\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Client environment:user.dir=/home/jmspaggiari/workspace/hbase-0.94.8-SNAPSHOT\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Initiating client connection, connectString=latitude:2181,cube:2181,node3:2181 sessionTimeout=180000 watcher=hconnection\n13/05/16 14:38:15 INFO zookeeper.RecoverableZooKeeper: The identifier of this process is 16902@cloudera\n13/05/16 14:38:15 INFO zookeeper.ClientCnxn: Opening socket connection to server latitude/192.168.23.4:2181. Will not attempt to authenticate using SASL (Unable to locate a login configuration)\n13/05/16 14:38:15 INFO zookeeper.ClientCnxn: Socket connection established to latitude/192.168.23.4:2181, initiating session\n13/05/16 14:38:15 INFO zookeeper.ClientCnxn: Session establishment complete on server latitude/192.168.23.4:2181, sessionid = 0x23eadc5ddeb008f, negotiated timeout = 40000\n13/05/16 14:38:15 INFO zookeeper.ZooKeeper: Initiating client connection, connectString=latitude:2181,cube:2181,node3:2181 sessionTimeout=180000 watcher=catalogtracker-on-org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation@79ee2c2c\n13/05/16 14:38:15 INFO zookeeper.RecoverableZooKeeper: The identifier of this process is 16902@cloudera\n13/05/16 14:38:15 INFO zookeeper.ClientCnxn: Opening socket connection to server latitude/192.168.23.4:2181. Will not attempt to authenticate using SASL (Unable to locate a login configuration)\n13/05/16 14:38:15 DEBUG catalog.CatalogTracker: Starting catalog tracker org.apache.hadoop.hbase.catalog.CatalogTracker@77ff92f5\n13/05/16 14:38:15 INFO zookeeper.ClientCnxn: Socket connection established to latitude/192.168.23.4:2181, initiating session\n13/05/16 14:38:16 INFO zookeeper.ClientCnxn: Session establishment complete on server latitude/192.168.23.4:2181, sessionid = 0x23eadc5ddeb0090, negotiated timeout = 40000\n13/05/16 14:38:16 DEBUG client.HConnectionManager$HConnectionImplementation: Looked up root region location, connection=org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation@79ee2c2c; serverName=node5,60020,1368669066647\n13/05/16 14:38:16 DEBUG client.HConnectionManager$HConnectionImplementation: Cached location for .META.,,1.1028785192 is node6:60020\n13/05/16 14:38:16 DEBUG client.ClientScanner: Creating scanner over .META. starting at key 'test,,'\n13/05/16 14:38:16 DEBUG client.ClientScanner: Advancing internal scanner to startKey at 'test,,'\n13/05/16 14:38:16 DEBUG catalog.CatalogTracker: Stopping catalog tracker org.apache.hadoop.hbase.catalog.CatalogTracker@77ff92f5\n13/05/16 14:38:16 INFO zookeeper.ZooKeeper: Session: 0x23eadc5ddeb0090 closed\n13/05/16 14:38:16 INFO zookeeper.ClientCnxn: EventThread shut down\n13/05/16 14:38:16 DEBUG client.MetaScanner: Scanning .META. starting at row=test,,00000000000000 for max=10 rows using org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation@79ee2c2c\n13/05/16 14:38:16 DEBUG client.HConnectionManager$HConnectionImplementation: Cached location for test,,1368729286037.32038a4d760aa9303643e2985dcd29a5. is node2:60020\n13/05/16 14:38:16 DEBUG client.MetaScanner: Scanning .META. starting at row= for max=2147483647 rows using org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation@79ee2c2c\n13/05/16 14:38:17 DEBUG client.MetaScanner: Scanning .META. starting at row=test,,00000000000000 for max=2147483647 rows using org.apache.hadoop.hbase.client.HConnectionManager$HConnectionImplementation@79ee2c2c\n13/05/16 14:38:17 INFO hfile.CacheConfig: Allocating LruBlockCache with maximum size 246.9m\n13/05/16 14:38:17 INFO util.ChecksumType: Checksum can use java.util.zip.CRC32\n13/05/16 14:38:17 INFO mapreduce.LoadIncrementalHFiles: Trying to load hfile=hdfs://node3:9000/user/hbase/familyDir2/myfam/hfile_1 first=aaaa last=aaaa\n13/05/16 14:38:18 DEBUG mapreduce.LoadIncrementalHFiles: Going to connect to server region=test,,1368729286037.32038a4d760aa9303643e2985dcd29a5., hostname=node2, port=60020 for row \n{code}","from":"developer"},{"body":"Any chance to have someone looking at this one?","from":"developer"},{"body":"patch looks reasonable esp. given it's a backport.\nIt has long lines.\nThese ifs:\n{code}\n+ if (assignSeqIds) {\n+ success = secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken, assignSeqIds);\n+ } else {\n+ success = secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken);\n+ }\n{code}\nseem excessive, just the first call is enough, right?","from":"developer"},{"body":"secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken) calls secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken, false) so if assignSeqIds is false, secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken) and secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken, assignSeqIds) are going to be the same call, therefore the if is not required.\n\nSame for \n{code}\n if (assignSeqIds) {\n success = secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken, assignSeqIds);\n } else {\n success = secureClient.bulkLoadHFiles(famPaths, userToken, bulkToken);\n }\n{code}\n\nI will change that and re-base the patch since it's not working anymore with current 0.94 branch...\n\nThanks for looking at it.","from":"developer"},{"body":"Rebased patch attached. Some modifications required because minSeqId got removed by HBASE-7210. Also applied Ted's comment. I'm not testing with the data attached to this ticket.","from":"developer"},{"body":"This jira will need a title that accurately describes what's being done here.\n\nOn the patch:\n\n - There are 11 \"== Durability.USE_DEFAULT\", can we just have a method somewhere (like in Mutation) that does it and is named \"isDefaultDurability\"?\n - LoadIncrementalHFiles.assignSeqIds should be final.\n - Since bulk loaded files can have sequence ids, we should print it out. StoreFile.toStringDetailed is a candidate for that change, there might be more\n - What's up with the commented out code in HRegion?\n - Is it passing all the unit tests? A trunk version would here to get some Hadoop QA love.\n\nI tested the patch to see if Jonathan's original use case is covered and it looks like it is. I also did some mixed workloads of normal Puts and bulk loaded files.","from":"developer"},{"body":"Hi [~jdcryans], thanks for looking at it.\n\n{quote}\n There are 11 \"== Durability.USE_DEFAULT\", can we just have a method somewhere (like in Mutation) that does it and is named \"isDefaultDurability\"?\n{quote}\nThis is because setWriteToWAL is now deprecated. If I set something locally to do the same thing as what setWriteToWAL was doing, then it's not very clean since we are just bypassing the deprecation warning. Other option is to keep setWriteToWAL calls and add @suppressWarning to avoid them and keep setWriteToWAL? \n\n{quote}\nLoadIncrementalHFiles.assignSeqIds should be final.\n{quote}\nI don't think we can. If I put it final, how are you going to modify it on the constructor? (assignSeqIds = conf.getBoolean(ASSIGN_SEQ_IDS, true))\n\n{quote}\nSince bulk loaded files can have sequence ids, we should print it out. StoreFile.toStringDetailed is a candidate for that change, there might be more\n{quote}\nStoreFile.toStringDetailed already display the sequenceId under certain conditions. I will remove the condition to make sure we always have this information.\n\n{quote}\nWhat's up with the commented out code in HRegion?\n{quote}\nWow, this should not be there. At all! Removed!\n\n{quote}\nIs it passing all the unit tests? A trunk version would here to get some Hadoop QA love.\n{quote}\nYes, it's passing unit tests. I will post the results before EOD. Also, there is no trunk version because it's a backport of HBASE-6630 which is already doing the same thing in trunk.\n\n\nSo, waiting for your recommendation regarding Durability.USE_DEFAULT tests and I will post an updated version.","from":"developer"},{"body":"Update version attached.\n1) final added. Was not working on trunk, but working fine on 0.94.\n2) Removed the Durability modifications since they are not related to this patch. Was only deprecated calls cleanup.\n\nShould be nicer now.","from":"developer"},{"body":"Hum, seems that I pushed the wrong file...","from":"developer"},{"body":"Removed commented code.","from":"developer"},{"body":"[~lhofhansl] any chance to look at that? Want it into 0.94? This patch is already running on a production cluster.","from":"developer"},{"body":"Still +0 on this.\n\nAny strong sponsors? If so, let's get it in.","from":"developer"},{"body":"I'll jump in late to the discussion to add a personal story. \n\nWe're very much relying on this patch since we bulk load everything and our use-case depends on this. We're running with this patch at the moment and we're hoping not to lose it when upgrading. We put automated testing of our feature that relies on this and we spent a lot of time looking at the internals (fun times with the Hfile tool) to check that all was as expected. While the test doesn't provide a full guarantee that the behavior is as expected as opposed to still being a combination of non-deterministic behavior and luck, this runs every day and each day that the test doesn't fail increases our confidence. \n\nI'm very much +1. If there's one patch we ever needed from hbase, this is it. ","from":"developer"},{"body":"Thanks for chiming in Alexandre... I'll count that a \"strong sponsor\" as per above.\nIf there are no objections, I will commit this to 0.94.13 tomorrow.\n","from":"developer"},{"body":"Thanks Lars, I appreciate it. I really do. ","from":"developer"},{"body":"Cool, that's all good news too! Thanks guys. One less on my radar.","from":"developer"},{"body":"[~lhofhansl] do you have time to cut a patch for 0.96 before the next RC? (cc [~stack])","from":"developer"},{"body":"^^^ or [~jmspaggi] :)","from":"developer"},{"body":"It's already there in HBASE-6630... Only 0.94 was missing it. So we should be good.","from":"developer"},{"body":"ACK. Thanks JM.","from":"developer"},{"body":"Committed to the 0.94 branch. Thanks for the patch.","from":"developer"},{"body":"TestLoadIncrementalHFilesSplitRecovery.testBulkLoadPhaseFailure failed... Looks suspicious.","from":"developer"},{"body":"I agree. Let me run that locally few multiple to validate...","from":"developer"},{"body":"Also, looking at v4 now. It adds\n{code}\nprivate void verifyAssignedSequenceNumber(String testName, byte hfileRanges, boolean nonZero)\n{code}\nBut it is never actually called from anywhere.\n","from":"developer"},{"body":"Definitely passes without the patch and fail with it.","from":"developer"},{"body":"Failed also on 0.94.10+HBASE-8521 :( Something might have changed since the last time I merged it. Please revert, I will look at the issue.","from":"developer"},{"body":"verifyAssignedSequenceNumber has been removed from trunk because it's never called. I took it from 0.89. I will remove it. Looking at the test failure now.","from":"developer"},{"body":"Ok. Looking deeper, I'm not sure that it's really an issue.\n\nThe test which failed:\n{code}\n /**\n * Test that shows that exception thrown from the RS side will result in an\n * exception on the LIHFile client.\n */\n{code}\n\nRS is not throwing an exception. But nothing related to this patch. But still, occurs only with this patch activated. So there might be some side effect. I will start to roll it back locally line by line to figure what the issue is...","from":"developer"},{"body":"Found it.\n\nDon't revert. It's related to the test, not to the code. Patch and detailed to come in a minute.","from":"developer"},{"body":"I knew I could wait a few minutes :)","from":"developer"},{"body":"So here is the culprit.\n\nThere is now an extra parameter to the bulkLoadHFiles call where we need to specify the assignSeqNum boolean. However, the mocked connection was still configured to fail on the method call without this parameter (which is not the one being called).\n\nThis addendum is to update the mocked connection configuration.\n\nPassed the test locally. Sorry for this.","from":"developer"},{"body":"Committed. Thanks JM!","from":"developer"}],"created":"2013-05-10T00:46:04.000+0000","description":"Let's say you have a pre-built HFile that contains a cell:\n\n('rowkey1', 'family1', 'qual1', 1234L, 'value1')\n\nWe bulk load this first HFile. Now, let's create a second HFile that contains a cell that overwrites the first:\n\n('rowkey1', 'family1', 'qual1', 1234L, 'value2')\n\nThat gets bulk loaded into the table, but the value that HBase bubbles up is still 'value1'.\n\nIt seems that there's no way to overwrite a cell for a particular timestamp without an explicit put operation. This seems to be the case even after minor and major compactions happen.\n\nMy guess is that this is pretty closely related to the sequence number work being done on the compaction algorithm via HBASE-7842, but I'm not sure if one of would fix the other.","issue_id":"12646931","key":"HBASE-8521","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2013-10-02T19:06:41.000+0000","role":"fixed_distractor","summary":"Cells cannot be overwritten with bulk loaded HFiles"} {"case_id":"12403648","cluster":"DISTRACTOR-HBASE-867","comments":[{"body":"It looks like we should be breaking out of the main while loop in the next method after we pass out a column match -- because store keys are sorted -- but then we fall on the next while loop which just nexts until we hit the next row.\n\nNormally this is fine, if only a few columns in a row, but in Daniel's case its taking forever to move to the next row.\n\nAlso, we won't split a region if only one row so its looking like his store files are large, 1G.","created":"2008-09-04T00:11:06.428+0000"},{"body":"To be clear, if thousands of columns plus -- i.e. a canonical usage -- hbase does not work. Here is some of the problem code form StoreFileScanner#next:\n\n{code}\n...\n while ((keys[i] != null)\n && (Bytes.compareTo(keys[i].getRow(), viableRow.getRow()) == 0)) {\n\n // If we are doing a wild card match or there are multiple matchers\n // per column, we need to scan all the older versions of this row\n // to pick up the rest of the family members\n if(!isWildcardScanner()\n && !isMultipleMatchScanner()\n && (keys[i].getTimestamp() != viableRow.getTimestamp())) {\n break;\n }\n\n if (columnMatch(i)) { \n // We only want the first result for any specific family member\n if(!results.containsKey(keys[i].getColumn())) {\n results.put(keys[i].getColumn(), \n new Cell(vals[i], keys[i].getTimestamp()));\n insertedItem = true;\n }\n } else {\n // Content is sorted. If column no longer matches, break.\n break;\n }\n\n if (!getNext(i)) {\n closeSubScanner(i);\n }\n }\n\n // Advance the current scanner beyond the chosen row, to\n // a valid timestamp, so we're ready next time.\n while ((keys[i] != null) &&\n ((Bytes.compareTo(keys[i].getRow(), viableRow.getRow()) <= 0)\n || (keys[i].getTimestamp() > this.timestamp)\n || (! columnMatch(i)))) {\n getNext(i);\n }\n..\n{code}\n\nThe whiles find next row by getting cells until the row does not match. If many columns per row, then that can take for ever (as its doing in Daniel's case). Need to have a file format that has an index that says where next row is. An option would say whether to get to next row by nexting or instead asking index.\n","created":"2008-09-04T16:28:49.070+0000"},{"body":"Marking this a critical issue. Only fix that I see is new mapfile type that keeps an index of each row offset.","created":"2008-09-04T18:19:05.459+0000"},{"body":"Its critical because this is canonical usage-pattern.","created":"2008-09-04T18:19:29.433+0000"},{"body":"Going to take a look at doing this for 0.19.0 since its embarrassing we don't do the canonical use case.","created":"2008-10-29T17:49:50.700+0000"},{"body":"Moving to 0.20.0. Need access to the MapFile index to do this fix. Currently its private.","created":"2008-11-17T17:39:08.252+0000"},{"body":"From #IRC, here is another case we need to be smarter about:\n\n{code}\n01:13 < BenM> keys[i] = new HStoreKey(HConstants.EMPTY_BYTE_ARRAY, this.store.getHRegionInfo());\n01:13 < BenM> if (firstRow != null && firstRow.length != 0) {\n01:13 < BenM> if (findFirstRow(i, firstRow)) {\n01:13 < BenM> continue;\n01:13 < BenM> }\n01:13 < BenM> }\n01:13 < BenM> while (getNext(i)) {\n01:13 < BenM> if (columnMatch(i)) {\n01:13 < BenM> break;\n01:13 < BenM> }\n01:13 < BenM> }\n{code}\n\nIts setting up scanners after store files have been changed.\n\nIf lots of entries for rows we don't care about, then these iterations will take a long time. Need to be smarter about the seek.","created":"2009-02-20T01:24:31.867+0000"},{"body":"I confirmed that this code is what was causing crashes for me. What happened is that I had a MR job that would launch multiple scanners on a region that made updates to the same column family as they were scanning on (but not the same column). As a result, there were lots of processes that had to grep through all of the irrelevent inserts many times as flushes occurred.\n\nI think that this case could be fixed in 0.19.0, and furthermore I think the fix might actually clean up the code a lot:\n\n(10:58:18 PM) BenM: yeah\n(10:58:21 PM) BenM: was just doing that\n(10:58:30 PM) BenM: IMHO, this is a somewhat easier issue to fix\n(10:58:38 PM) BenM: i think it could be done in a way that cleans up the code\n(10:58:50 PM) BenM: right now, the code just scans through each of the map files\n(10:59:02 PM) BenM: without regard to the relative key positions\n(10:59:12 PM) BenM: i think it could use a priority queue so that it only works on the relevent files\n(11:01:22 PM) St^Ack_: BenM: please expand, I don't follow exactly\n(11:01:50 PM) BenM: lets say we have two map files\n(11:02:09 PM) BenM: one with 1/foo:bar 2/foo:bar 3/foo:bar\n(11:02:17 PM) BenM: (row/family:col)\n(11:02:31 PM) BenM: and the other with 1000/blah:blah 1001/blah:blah\n(11:02:39 PM) BenM: the curent logic is\n(11:02:44 PM) BenM: for each map file:\n(11:02:56 PM) BenM: find the first potential row in this file\n(11:03:08 PM) BenM: look at min(all potential rows)\n(11:03:34 PM) BenM: the algorith should be:\n(11:03:43 PM) BenM: q = new PriorityQueue()\n(11:04:05 PM) BenM: for each map file: insert the HStoreKey of the first key in the file\n(11:04:17 PM) BenM: while(k = q.pop()) {\n(11:04:37 PM) BenM: if (k is intersting) break;\n(11:04:37 PM) BenM: advance k\n(11:04:37 PM) BenM: q.push(k)\n(11:04:38 PM) BenM: }\n(11:05:00 PM) BenM: that way, we don't try to find a matching key in the larger rows","created":"2009-02-20T04:15:20.332+0000"},{"body":"Would like to get this solved as part of 1249 issues.","created":"2009-04-28T17:08:58.043+0000"},{"body":"The idea Ben describes above is part of HBASE-1249. This issue will be resolved as part of 1249.","created":"2009-05-08T03:31:31.659+0000"},{"body":"Issue is probably resolved, but I'm going to continue testing. It's not possible to create JUnit tests because the psuedo-distr. cluster falls over under load this high.\n\nFirst simple test (gets though, not scans... scans next):\n\n{noformat}\nGenerated 1 Puts in 3068ms\n 1 rows, 8 bytes/row key\n 1000000 columns/row, 8 bytes/column key\n 8 bytes/column value\nInserted Put in 23295ms\nGet0 (1000000 KVs) completed in 4031ms\nGet1 (1000000 KVs) completed in 2300ms\nGet2 (1000000 KVs) completed in 2831ms\nGet3 (1000000 KVs) completed in 1707ms\nGet4 (1000000 KVs) completed in 2588ms\nGet5 (1000000 KVs) completed in 2671ms\nGet6 (1000000 KVs) completed in 2442ms\nGet7 (1000000 KVs) completed in 2560ms\nGet8 (1000000 KVs) completed in 2462ms\nGet9 (1000000 KVs) completed in 2822ms\n{noformat}","created":"2009-06-06T21:48:20.948+0000"},{"body":"Whats above saying JGray? That you added single row with a million columns? And then you did 10 gets and they took 2.5ms on average to complete? Which column were you getting?\n\nScans would be good to get numbers for. That was reason for original filing.\n\nGood stuff.","created":"2009-06-08T02:08:14.427+0000"},{"body":"I am doing tests for this issue on a 5+1 node cluster, each node is 2core/2gb and hosting two HDFS and two HBase instances (0.19 cluster still up but it's idle).\n\nUsing a newer version of the HBench tool I posted in HBASE-1501, I'm able to run a number of different tests with high numbers of columns.\n\nMy test is inserting 10 rows, each with 2M columns. I do it in 200 rounds, each round I insert 10k columns in each of the 10 rows.\n\nQualifiers are incremented binary longs (1 -> 2M), so 8 bytes. Values are randomized binary data of fixed length. By varying the size of the value (have tried between 8 and 32 bytes per value), I can get different behavior. \n\nWith not much memory to give the RS, I run into OOME problems when serializing the Result. I'm going to rerun tests at higher value sizes and get some clean logs to look at, making sure I have block caching disabled so it doesn't hog heap.\n\nHowever, with 8 byte values I'm able to import without a problem (causes several splits, in the end we have 5 regions for the 10 rows). In addition to the import test, I'm also scanning these 10 rows in two ways. A full scan (all in family) as well as a skip scan (i'm asking for two specific columns, qualifier=1 and qualifier=1888888, so beginning and end of each row).\n\n{noformat}\nInserted 10 rows each with 2000000 total columns in 344566ms (34456.6ms/row)\n\nSkip Scanner open\nRow [row0] Scanned, Contains 2 Columns (10155 ms)\nRow [row1] Scanned, Contains 2 Columns (9978 ms)\nRow [row2] Scanned, Contains 2 Columns (10675 ms)\nRow [row3] Scanned, Contains 2 Columns (9608 ms)\nRow [row4] Scanned, Contains 2 Columns (11703 ms)\nRow [row5] Scanned, Contains 2 Columns (12103 ms)\nRow [row6] Scanned, Contains 2 Columns (6828 ms)\nRow [row7] Scanned, Contains 2 Columns (6603 ms)\nRow [row8] Scanned, Contains 2 Columns (6331 ms)\nRow [row9] Scanned, Contains 2 Columns (6553 ms)\nScanned 10 rows in 90551ms (9055.1ms/row)\n\nFull Scanner open\nRow [row0] Scanned, Contains 2000000 Columns (14374 ms)\nRow [row1] Scanned, Contains 2000000 Columns (14879 ms)\nRow [row2] Scanned, Contains 2000000 Columns (14053 ms)\nRow [row3] Scanned, Contains 2000000 Columns (14263 ms)\nRow [row4] Scanned, Contains 2000000 Columns (8811 ms)\nRow [row5] Scanned, Contains 2000000 Columns (10327 ms)\nRow [row6] Scanned, Contains 2000000 Columns (9757 ms)\nRow [row7] Scanned, Contains 2000000 Columns (9343 ms)\nRow [row8] Scanned, Contains 2000000 Columns (9526 ms)\nRow [row9] Scanned, Contains 2000000 Columns (10004 ms)\nScanned 10 rows in 115342ms (11534.2ms/row)\n{noformat}\n\nRepeated runs improve performance, and ordering of the two types of scans makes a difference. Block cache is off so we're seeing the effect of the linux file cache.","created":"2009-06-17T19:59:35.484+0000"},{"body":"Given that I can do the above on a smallish cluster without a problem (the only problems come from OOME when serializing these big rows if I grow value size), I am resolving this issue as fixed by HBASE-1304.\n\nTo be able to scale further as far as columns in a row, we need intra-row scanning. I have filed HBASE-1537 currently targeted at 0.21 for now.\n\nThere are still improvements to be made when not returning all columns in a row with millions of columns. That optimization is addressed in HBASE-1517 and slated for 0.20.1.","created":"2009-06-17T21:46:35.429+0000"},{"body":"I'm +1 on resolving this issue because of 1304 work (in spite of what the issue title says) and doing further improvements out in other issues. Let me add note that this issue was resolved to CHANGES.txt","created":"2009-06-17T21:59:45.404+0000"}],"conversations":[{"body":"Our Daniel has uploaded a table that has a column family with millions of columns in it. He can get items from the table promptly specifying row and column. Scanning is another matter. Thread dumping I see we're stuck in the scanner constructor nexting through cells.","from":"reporter","subject":"If millions of columns in a column family, hbase scanner won't come up"},{"body":"It looks like we should be breaking out of the main while loop in the next method after we pass out a column match -- because store keys are sorted -- but then we fall on the next while loop which just nexts until we hit the next row.\n\nNormally this is fine, if only a few columns in a row, but in Daniel's case its taking forever to move to the next row.\n\nAlso, we won't split a region if only one row so its looking like his store files are large, 1G.","from":"developer"},{"body":"To be clear, if thousands of columns plus -- i.e. a canonical usage -- hbase does not work. Here is some of the problem code form StoreFileScanner#next:\n\n{code}\n...\n while ((keys[i] != null)\n && (Bytes.compareTo(keys[i].getRow(), viableRow.getRow()) == 0)) {\n\n // If we are doing a wild card match or there are multiple matchers\n // per column, we need to scan all the older versions of this row\n // to pick up the rest of the family members\n if(!isWildcardScanner()\n && !isMultipleMatchScanner()\n && (keys[i].getTimestamp() != viableRow.getTimestamp())) {\n break;\n }\n\n if (columnMatch(i)) { \n // We only want the first result for any specific family member\n if(!results.containsKey(keys[i].getColumn())) {\n results.put(keys[i].getColumn(), \n new Cell(vals[i], keys[i].getTimestamp()));\n insertedItem = true;\n }\n } else {\n // Content is sorted. If column no longer matches, break.\n break;\n }\n\n if (!getNext(i)) {\n closeSubScanner(i);\n }\n }\n\n // Advance the current scanner beyond the chosen row, to\n // a valid timestamp, so we're ready next time.\n while ((keys[i] != null) &&\n ((Bytes.compareTo(keys[i].getRow(), viableRow.getRow()) <= 0)\n || (keys[i].getTimestamp() > this.timestamp)\n || (! columnMatch(i)))) {\n getNext(i);\n }\n..\n{code}\n\nThe whiles find next row by getting cells until the row does not match. If many columns per row, then that can take for ever (as its doing in Daniel's case). Need to have a file format that has an index that says where next row is. An option would say whether to get to next row by nexting or instead asking index.\n","from":"developer"},{"body":"Marking this a critical issue. Only fix that I see is new mapfile type that keeps an index of each row offset.","from":"developer"},{"body":"Its critical because this is canonical usage-pattern.","from":"developer"},{"body":"Going to take a look at doing this for 0.19.0 since its embarrassing we don't do the canonical use case.","from":"developer"},{"body":"Moving to 0.20.0. Need access to the MapFile index to do this fix. Currently its private.","from":"developer"},{"body":"From #IRC, here is another case we need to be smarter about:\n\n{code}\n01:13 < BenM> keys[i] = new HStoreKey(HConstants.EMPTY_BYTE_ARRAY, this.store.getHRegionInfo());\n01:13 < BenM> if (firstRow != null && firstRow.length != 0) {\n01:13 < BenM> if (findFirstRow(i, firstRow)) {\n01:13 < BenM> continue;\n01:13 < BenM> }\n01:13 < BenM> }\n01:13 < BenM> while (getNext(i)) {\n01:13 < BenM> if (columnMatch(i)) {\n01:13 < BenM> break;\n01:13 < BenM> }\n01:13 < BenM> }\n{code}\n\nIts setting up scanners after store files have been changed.\n\nIf lots of entries for rows we don't care about, then these iterations will take a long time. Need to be smarter about the seek.","from":"developer"},{"body":"I confirmed that this code is what was causing crashes for me. What happened is that I had a MR job that would launch multiple scanners on a region that made updates to the same column family as they were scanning on (but not the same column). As a result, there were lots of processes that had to grep through all of the irrelevent inserts many times as flushes occurred.\n\nI think that this case could be fixed in 0.19.0, and furthermore I think the fix might actually clean up the code a lot:\n\n(10:58:18 PM) BenM: yeah\n(10:58:21 PM) BenM: was just doing that\n(10:58:30 PM) BenM: IMHO, this is a somewhat easier issue to fix\n(10:58:38 PM) BenM: i think it could be done in a way that cleans up the code\n(10:58:50 PM) BenM: right now, the code just scans through each of the map files\n(10:59:02 PM) BenM: without regard to the relative key positions\n(10:59:12 PM) BenM: i think it could use a priority queue so that it only works on the relevent files\n(11:01:22 PM) St^Ack_: BenM: please expand, I don't follow exactly\n(11:01:50 PM) BenM: lets say we have two map files\n(11:02:09 PM) BenM: one with 1/foo:bar 2/foo:bar 3/foo:bar\n(11:02:17 PM) BenM: (row/family:col)\n(11:02:31 PM) BenM: and the other with 1000/blah:blah 1001/blah:blah\n(11:02:39 PM) BenM: the curent logic is\n(11:02:44 PM) BenM: for each map file:\n(11:02:56 PM) BenM: find the first potential row in this file\n(11:03:08 PM) BenM: look at min(all potential rows)\n(11:03:34 PM) BenM: the algorith should be:\n(11:03:43 PM) BenM: q = new PriorityQueue()\n(11:04:05 PM) BenM: for each map file: insert the HStoreKey of the first key in the file\n(11:04:17 PM) BenM: while(k = q.pop()) {\n(11:04:37 PM) BenM: if (k is intersting) break;\n(11:04:37 PM) BenM: advance k\n(11:04:37 PM) BenM: q.push(k)\n(11:04:38 PM) BenM: }\n(11:05:00 PM) BenM: that way, we don't try to find a matching key in the larger rows","from":"developer"},{"body":"Would like to get this solved as part of 1249 issues.","from":"developer"},{"body":"The idea Ben describes above is part of HBASE-1249. This issue will be resolved as part of 1249.","from":"developer"},{"body":"Issue is probably resolved, but I'm going to continue testing. It's not possible to create JUnit tests because the psuedo-distr. cluster falls over under load this high.\n\nFirst simple test (gets though, not scans... scans next):\n\n{noformat}\nGenerated 1 Puts in 3068ms\n 1 rows, 8 bytes/row key\n 1000000 columns/row, 8 bytes/column key\n 8 bytes/column value\nInserted Put in 23295ms\nGet0 (1000000 KVs) completed in 4031ms\nGet1 (1000000 KVs) completed in 2300ms\nGet2 (1000000 KVs) completed in 2831ms\nGet3 (1000000 KVs) completed in 1707ms\nGet4 (1000000 KVs) completed in 2588ms\nGet5 (1000000 KVs) completed in 2671ms\nGet6 (1000000 KVs) completed in 2442ms\nGet7 (1000000 KVs) completed in 2560ms\nGet8 (1000000 KVs) completed in 2462ms\nGet9 (1000000 KVs) completed in 2822ms\n{noformat}","from":"developer"},{"body":"Whats above saying JGray? That you added single row with a million columns? And then you did 10 gets and they took 2.5ms on average to complete? Which column were you getting?\n\nScans would be good to get numbers for. That was reason for original filing.\n\nGood stuff.","from":"developer"},{"body":"I am doing tests for this issue on a 5+1 node cluster, each node is 2core/2gb and hosting two HDFS and two HBase instances (0.19 cluster still up but it's idle).\n\nUsing a newer version of the HBench tool I posted in HBASE-1501, I'm able to run a number of different tests with high numbers of columns.\n\nMy test is inserting 10 rows, each with 2M columns. I do it in 200 rounds, each round I insert 10k columns in each of the 10 rows.\n\nQualifiers are incremented binary longs (1 -> 2M), so 8 bytes. Values are randomized binary data of fixed length. By varying the size of the value (have tried between 8 and 32 bytes per value), I can get different behavior. \n\nWith not much memory to give the RS, I run into OOME problems when serializing the Result. I'm going to rerun tests at higher value sizes and get some clean logs to look at, making sure I have block caching disabled so it doesn't hog heap.\n\nHowever, with 8 byte values I'm able to import without a problem (causes several splits, in the end we have 5 regions for the 10 rows). In addition to the import test, I'm also scanning these 10 rows in two ways. A full scan (all in family) as well as a skip scan (i'm asking for two specific columns, qualifier=1 and qualifier=1888888, so beginning and end of each row).\n\n{noformat}\nInserted 10 rows each with 2000000 total columns in 344566ms (34456.6ms/row)\n\nSkip Scanner open\nRow [row0] Scanned, Contains 2 Columns (10155 ms)\nRow [row1] Scanned, Contains 2 Columns (9978 ms)\nRow [row2] Scanned, Contains 2 Columns (10675 ms)\nRow [row3] Scanned, Contains 2 Columns (9608 ms)\nRow [row4] Scanned, Contains 2 Columns (11703 ms)\nRow [row5] Scanned, Contains 2 Columns (12103 ms)\nRow [row6] Scanned, Contains 2 Columns (6828 ms)\nRow [row7] Scanned, Contains 2 Columns (6603 ms)\nRow [row8] Scanned, Contains 2 Columns (6331 ms)\nRow [row9] Scanned, Contains 2 Columns (6553 ms)\nScanned 10 rows in 90551ms (9055.1ms/row)\n\nFull Scanner open\nRow [row0] Scanned, Contains 2000000 Columns (14374 ms)\nRow [row1] Scanned, Contains 2000000 Columns (14879 ms)\nRow [row2] Scanned, Contains 2000000 Columns (14053 ms)\nRow [row3] Scanned, Contains 2000000 Columns (14263 ms)\nRow [row4] Scanned, Contains 2000000 Columns (8811 ms)\nRow [row5] Scanned, Contains 2000000 Columns (10327 ms)\nRow [row6] Scanned, Contains 2000000 Columns (9757 ms)\nRow [row7] Scanned, Contains 2000000 Columns (9343 ms)\nRow [row8] Scanned, Contains 2000000 Columns (9526 ms)\nRow [row9] Scanned, Contains 2000000 Columns (10004 ms)\nScanned 10 rows in 115342ms (11534.2ms/row)\n{noformat}\n\nRepeated runs improve performance, and ordering of the two types of scans makes a difference. Block cache is off so we're seeing the effect of the linux file cache.","from":"developer"},{"body":"Given that I can do the above on a smallish cluster without a problem (the only problems come from OOME when serializing these big rows if I grow value size), I am resolving this issue as fixed by HBASE-1304.\n\nTo be able to scale further as far as columns in a row, we need intra-row scanning. I have filed HBASE-1537 currently targeted at 0.21 for now.\n\nThere are still improvements to be made when not returning all columns in a row with millions of columns. That optimization is addressed in HBASE-1517 and slated for 0.20.1.","from":"developer"},{"body":"I'm +1 on resolving this issue because of 1304 work (in spite of what the issue title says) and doing further improvements out in other issues. Let me add note that this issue was resolved to CHANGES.txt","from":"developer"}],"created":"2008-09-04T00:06:09.000+0000","description":"Our Daniel has uploaded a table that has a column family with millions of columns in it. He can get items from the table promptly specifying row and column. Scanning is another matter. Thread dumping I see we're stuck in the scanner constructor nexting through cells.","issue_id":"12403648","key":"HBASE-867","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2009-06-17T21:46:35.000+0000","role":"fixed_distractor","summary":"If millions of columns in a column family, hbase scanner won't come up"} {"case_id":"12677625","cluster":"DISTRACTOR-HBASE-9895","comments":[{"body":"No good way to dynamically determine an input file format in 0.94 so introducing a system property such as following in order for Import to load a file using 0.94 deserializer.\n\n{code}\n./bin/hbase -Dhbase.input.version=0.94 org.apache.hadoop.hbase.mapreduce.Import \n{code}","created":"2013-11-07T03:10:11.484+0000"},{"body":"Re-attach because the QA run errors seems not related to this patch.","created":"2013-11-08T01:59:11.914+0000"},{"body":"This looks like magic to me; where is setConf(Configuration) called on the Serialization instance? When will getConf() be non-null?","created":"2013-11-08T18:53:13.119+0000"},{"body":"class ResultSerialization is extended from Configured. Therefore, when mapreduce initializes those class and configuration will be passed to the new instance automatically(magically). ","created":"2013-11-08T19:04:50.340+0000"},{"body":"Alright then. Any chance of adding some kind of test? Maybe a data blob that matches the old format? Looks good otherwise.","created":"2013-11-08T19:27:44.238+0000"},{"body":"I added a test case. Since there is no way to put the binary 0.94 exported file, the new test case testImport94Table will fail in the QA run. \n\nThe new test case run fine in my local env with the binary file is put hbase-server/src/test/resources/org/apache/hadoop/hbase/mapreduce/exportedTableIn94Format.","created":"2013-11-09T02:21:42.735+0000"},{"body":"The sample file is small. Can we add it as a src/test/resource ?","created":"2013-11-09T03:59:01.804+0000"},{"body":"{quote}\nThe sample file is small. Can we add it as a src/test/resource ?\n{quote}\nThe binary data file is added into src/test/resources/org/apache/hadoop/hbase/mapreduce/exportedTableIn94Format. As I mentioned above, patch file can't contain binary file so I have to attach the binary file separately and QA run will fail. When commit the change, I'll commit the binary under src/test/resources/org/apache/hadoop/hbase/mapreduce/exportedTableIn94Format along with the text patch file. Therefore, there won't be any issue. ","created":"2013-11-09T04:32:56.359+0000"},{"body":"{code}\n+ public static final String INPUT_FORMT_VER = \"hbase.input.version\";\n{code}\nThere is no comment above the constant definition.\nFORMT is not a word. Did you intend to say FORMAT or FROM ?","created":"2013-11-09T06:12:57.540+0000"},{"body":"That's a typo and should be INPUT_FORMAT_VER. I'll fix it and add comment when I check in the patch. Thanks for the good catch!","created":"2013-11-10T06:28:11.722+0000"},{"body":"This new test is exactly what I had in mind, perfect. Can you confirm that it's cleaning up the temp filesystem after the test ends, make sure it's not abandoning any cruft in the run directory.\n\nWas removal of {{testMetaExport}} intentional?","created":"2013-11-11T18:02:29.993+0000"},{"body":"Thanks [~ndimiduk] for the good catch. The missed test case is added back. The temporary file is cleaned after each test.\n\nThe v3 patch contains [~tedyu@apache.org] and [~ndimiduk] feedbacks. Thanks.","created":"2013-11-11T19:07:12.120+0000"},{"body":"Excellent. +1","created":"2013-11-11T19:23:18.263+0000"},{"body":"+1 for trunk and 0.96. Please add nice release note since it is my guess a few folks will be looking to see if this facility is present in 0.96 going forward.","created":"2013-11-11T20:53:22.760+0000"},{"body":"Nice patch +1. \nPlease change the param name \"hbase.input.version\", to be hbase.import.version. I guess this can be used with -Dhbase.import.version override when launching the job. ","created":"2013-11-11T21:11:54.545+0000"},{"body":"Thanks for all the reviews! [~enis] I'll rename config setting name to \"hbase.import.version\" when I commit the change. \n[~saint.ack@gmail.com] Yeah, I'll add a release note on this. Thanks.","created":"2013-11-11T23:32:07.262+0000"},{"body":"Thanks for all the reviews & feedbacks! I've integrated the patch into 0.96 and trunk branch.","created":"2013-11-12T02:58:40.425+0000"},{"body":"Released in 0.96.1. Issue closed.","created":"2013-12-16T18:46:53.082+0000"},{"body":"[~jeffreyz] This option should be documented in the usage (run with no args) and in the refguide here http://hbase.apache.org/book.html#import . Mind doing a follow on and adding documentation (cut and past of new usage would be great)? See importTsv for a similar example.","created":"2013-12-17T16:28:51.590+0000"},{"body":"ok, let me add some notes there.","created":"2013-12-17T18:53:37.577+0000"}],"conversations":[{"body":"Basically we PBed org.apache.hadoop.hbase.client.Result so a 0.96 cluster cannot import 0.94 exported files. This issue is annoying because a user can't import his old archive files after upgrade or archives from others who are using 0.94.\n\nThe ideal way is to catch deserialization error and then fall back to 0.94 format for importing.","from":"reporter","subject":"0.96 Import utility can't import an exported file from 0.94"},{"body":"No good way to dynamically determine an input file format in 0.94 so introducing a system property such as following in order for Import to load a file using 0.94 deserializer.\n\n{code}\n./bin/hbase -Dhbase.input.version=0.94 org.apache.hadoop.hbase.mapreduce.Import
\n{code}","from":"developer"},{"body":"Re-attach because the QA run errors seems not related to this patch.","from":"developer"},{"body":"This looks like magic to me; where is setConf(Configuration) called on the Serialization instance? When will getConf() be non-null?","from":"developer"},{"body":"class ResultSerialization is extended from Configured. Therefore, when mapreduce initializes those class and configuration will be passed to the new instance automatically(magically). ","from":"developer"},{"body":"Alright then. Any chance of adding some kind of test? Maybe a data blob that matches the old format? Looks good otherwise.","from":"developer"},{"body":"I added a test case. Since there is no way to put the binary 0.94 exported file, the new test case testImport94Table will fail in the QA run. \n\nThe new test case run fine in my local env with the binary file is put hbase-server/src/test/resources/org/apache/hadoop/hbase/mapreduce/exportedTableIn94Format.","from":"developer"},{"body":"The sample file is small. Can we add it as a src/test/resource ?","from":"developer"},{"body":"{quote}\nThe sample file is small. Can we add it as a src/test/resource ?\n{quote}\nThe binary data file is added into src/test/resources/org/apache/hadoop/hbase/mapreduce/exportedTableIn94Format. As I mentioned above, patch file can't contain binary file so I have to attach the binary file separately and QA run will fail. When commit the change, I'll commit the binary under src/test/resources/org/apache/hadoop/hbase/mapreduce/exportedTableIn94Format along with the text patch file. Therefore, there won't be any issue. ","from":"developer"},{"body":"{code}\n+ public static final String INPUT_FORMT_VER = \"hbase.input.version\";\n{code}\nThere is no comment above the constant definition.\nFORMT is not a word. Did you intend to say FORMAT or FROM ?","from":"developer"},{"body":"That's a typo and should be INPUT_FORMAT_VER. I'll fix it and add comment when I check in the patch. Thanks for the good catch!","from":"developer"},{"body":"This new test is exactly what I had in mind, perfect. Can you confirm that it's cleaning up the temp filesystem after the test ends, make sure it's not abandoning any cruft in the run directory.\n\nWas removal of {{testMetaExport}} intentional?","from":"developer"},{"body":"Thanks [~ndimiduk] for the good catch. The missed test case is added back. The temporary file is cleaned after each test.\n\nThe v3 patch contains [~tedyu@apache.org] and [~ndimiduk] feedbacks. Thanks.","from":"developer"},{"body":"Excellent. +1","from":"developer"},{"body":"+1 for trunk and 0.96. Please add nice release note since it is my guess a few folks will be looking to see if this facility is present in 0.96 going forward.","from":"developer"},{"body":"Nice patch +1. \nPlease change the param name \"hbase.input.version\", to be hbase.import.version. I guess this can be used with -Dhbase.import.version override when launching the job. ","from":"developer"},{"body":"Thanks for all the reviews! [~enis] I'll rename config setting name to \"hbase.import.version\" when I commit the change. \n[~saint.ack@gmail.com] Yeah, I'll add a release note on this. Thanks.","from":"developer"},{"body":"Thanks for all the reviews & feedbacks! I've integrated the patch into 0.96 and trunk branch.","from":"developer"},{"body":"Released in 0.96.1. Issue closed.","from":"developer"},{"body":"[~jeffreyz] This option should be documented in the usage (run with no args) and in the refguide here http://hbase.apache.org/book.html#import . Mind doing a follow on and adding documentation (cut and past of new usage would be great)? See importTsv for a similar example.","from":"developer"},{"body":"ok, let me add some notes there.","from":"developer"}],"created":"2013-11-05T19:06:05.000+0000","description":"Basically we PBed org.apache.hadoop.hbase.client.Result so a 0.96 cluster cannot import 0.94 exported files. This issue is annoying because a user can't import his old archive files after upgrade or archives from others who are using 0.94.\n\nThe ideal way is to catch deserialization error and then fall back to 0.94 format for importing.","issue_id":"12677625","key":"HBASE-9895","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2013-11-12T02:58:40.000+0000","role":"fixed_distractor","summary":"0.96 Import utility can't import an exported file from 0.94"} {"case_id":"12859478","cluster":"DISTRACTOR-SPARK-10309","comments":[{"body":"Why is this targeted for 1.6 ? We are finding this issue with basic Spark SQL executions in our applications. Is the expectation that Tungsten sort will be disabled in an upcoming checkin ?\n\nJob aborted due to stage failure: Task 1 in stage 25.0 failed 4 times, most recent failure: Lost task 1.3 in stage 25.0 (TID 3962, 39.6.64.17): java.io.IOException: Unable to acquire 16777216 bytes of memory\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:368)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.(UnsafeExternalSorter.java:138)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.create(UnsafeExternalSorter.java:106)\n\tat org.apache.spark.sql.execution.UnsafeExternalRowSorter.(UnsafeExternalRowSorter.java:68)\n\tat org.apache.spark.sql.execution.TungstenSort.org$apache$spark$sql$execution$TungstenSort$$preparePartition$1(sort.scala:146)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n\tat org.apache.spark.rdd.MapPartitionsWithPreparationRDD.prepare(MapPartitionsWithPreparationRDD.scala:50)\n\tat org.apache.spark.rdd.ZippedPartitionsBaseRDD$$anonfun$tryPrepareParents$1.applyOrElse(ZippedPartitionsRDD.scala:83)\n\tat org.apache.spark.rdd.ZippedPartitionsBaseRDD$$anonfun$tryPrepareParents$1.applyOrElse(ZippedPartitionsRDD.scala:82)\n\tat scala.runtime.AbstractPartialFunction.apply(AbstractPartialFunction.scala:33)\n\tat scala.collection.TraversableLike$$anonfun$collect$1.apply(TraversableLike.scala:278)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n","created":"2015-09-05T07:34:45.978+0000"},{"body":"[~nadenf] In my case, the job finally finished (after retry), so this seems to be a blocker for me.\n\nCould you provide more information about you job?","created":"2015-09-08T18:57:35.506+0000"},{"body":"We are using Spark Job Server to submit the job.\n\nEach job consists of:\n\n1) Execute an SQL statement against HDFS.\n2) Write the results into HDFS.\n3) Writes the result into MongoDB using just a normal Java adapter.\n\nWe do many of these in parallel. We have 6 node cluster with 30GB allocated to Spark (-xmx30g) and 60GB free. We constantly get these failures.","created":"2015-09-09T00:08:12.605+0000"},{"body":"[~nadenf] Could you post the physical plan here? That could help us to understand the root cause.","created":"2015-09-09T00:41:28.164+0000"},{"body":"This also could be related to https://issues.apache.org/jira/browse/SPARK-10341?filter=-2, could you test with 1.5-RC3?","created":"2015-09-09T00:46:10.560+0000"},{"body":"Still working on the physical plan but we have been testing with the latest branch-1.5.0 releases which included this fix. It doesn't help.","created":"2015-09-09T06:43:30.391+0000"},{"body":"[~nadenf] Thanks for letting us know, just realized that your stacktrace already including that fix.\n\nMaybe there are multiple join/aggregation/sort in your query? You can show the physical plan by `df.explain()` ","created":"2015-09-09T16:53:24.286+0000"},{"body":"Is there currently any workaround this issue?\n\nI am also facing it with the last 1.5.1:\n{code:title=Error|borderStyle=solid}\nCaused by: java.io.IOException: Unable to acquire 33554432 bytes of memory at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:351) at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.(UnsafeExternalSorter.java:138) at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.create(UnsafeExternalSorter.java:106) at org.apache.spark.sql.execution.UnsafeKVExternalSorter.(UnsafeKVExternalSorter.java:74) at org.apache.spark.sql.execution.UnsafeKVExternalSorter.(UnsafeKVExternalSorter.java:56) at org.apache.spark.sql.execution.datasources.DynamicPartitionWriterContainer.writeRows(WriterContainer.scala:339) ... 8 more\n{code}","created":"2015-09-25T07:30:43.940+0000"},{"body":"Has been difficult to get a clean stacktrace/explain trace because we are executing lots of SQL commands in parallel and we don't know which one is failing. We are absolutely doing lots of joins/aggregation/sorts. \n\nI have tried increase shuffle.memoryFraction to 0.8 but that didn't help.\n\nThis is still an issue with the latest Spark 1.5.2 branch.\n\n","created":"2015-10-01T22:02:26.456+0000"},{"body":"Same issue, I got the following stacktrace:\n{noformat}\n15/10/20 18:20:43 INFO UnsafeExternalSorter: Thread 64 spilling sort data of 64.0 KB to disk (0 time so far)\n15/10/20 18:20:43 ERROR Executor: Exception in task 11.3 in stage 2.0 (TID 514)\njava.io.IOException: Unable to acquire 67108864 bytes of memory\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:351)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPageIfNecessary(UnsafeExternalSorter.java:332)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.insertKVRecord(UnsafeExternalSorter.java:461)\n\tat org.apache.spark.sql.execution.UnsafeKVExternalSorter.insertKV(UnsafeKVExternalSorter.java:139)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.switchToSortBasedAggregation(TungstenAggregationIterator.scala:489)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.processInputs(TungstenAggregationIterator.scala:379)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.start(TungstenAggregationIterator.scala:622)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1.org$apache$spark$sql$execution$aggregate$TungstenAggregate$$anonfun$$executePartition$1(TungstenAggregate.scala:110)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:119)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:119)\n\tat org.apache.spark.rdd.MapPartitionsWithPreparationRDD.compute(MapPartitionsWithPreparationRDD.scala:64)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:88)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{noformat}","created":"2015-10-20T18:24:42.943+0000"},{"body":"I guess same issue here also.\n{code} \n15/10/26 15:11:33 INFO UnsafeExternalSorter: Thread 4524 spilling sort data of 64.0 KB to disk (0 time so far)\n15/10/26 15:11:33 INFO Executor: Executor is trying to kill task 135.0 in stage 394.0 (TID 11069)\n15/10/26 15:11:33 INFO UnsafeExternalSorter: Thread 4607 spilling sort data of 64.0 KB to disk (0 time so far)\n15/10/26 15:11:33 ERROR Executor: Managed memory leak detected; size = 67108864 bytes, TID = 11149\n15/10/26 15:11:33 ERROR Executor: Exception in task 92.3 in stage 394.0 (TID 11149)\njava.io.IOException: Unable to acquire 67108864 bytes of memory\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:351)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.(UnsafeExternalSorter.java:138)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.create(UnsafeExternalSorter.java:106)\n\tat org.apache.spark.sql.execution.UnsafeExternalRowSorter.(UnsafeExternalRowSorter.java:68)\n\tat org.apache.spark.sql.execution.TungstenSort.org$apache$spark$sql$execution$TungstenSort$$preparePartition$1(sort.scala:146)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n\tat org.apache.spark.rdd.MapPartitionsWithPreparationRDD.prepare(MapPartitionsWithPreparationRDD.scala:50)\n\tat org.apache.spark.rdd.ZippedPartitionsBaseRDD$$anonfun$tryPrepareParents$1.applyOrElse(ZippedPartitionsRDD.scala:83)\n\tat org.apache.spark.rdd.ZippedPartitionsBaseRDD$$anonfun$tryPrepareParents$1.applyOrElse(ZippedPartitionsRDD.scala:82)\n\tat scala.runtime.AbstractPartialFunction.apply(AbstractPartialFunction.scala:33)\n\tat scala.collection.TraversableLike$$anonfun$collect$1.apply(TraversableLike.scala:278)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n\tat scala.collection.TraversableLike$class.collect(TraversableLike.scala:278)\n\tat scala.collection.AbstractTraversable.collect(Traversable.scala:105)\n\tat org.apache.spark.rdd.ZippedPartitionsBaseRDD.tryPrepareParents(ZippedPartitionsRDD.scala:82)\n\tat org.apache.spark.rdd.ZippedPartitionsRDD2.compute(ZippedPartitionsRDD.scala:97)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsWithPreparationRDD.compute(MapPartitionsWithPreparationRDD.scala:63)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:88)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{code} ","created":"2015-10-26T15:26:24.226+0000"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9241","created":"2015-10-27T19:41:07.450+0000"},{"body":"Is there any work around for this issue. We migrated from 1.1 to 1.5 and our jobs heavily depends on join. I have been trying to get rid of this exception but no luck.\n\nIf someone can at least point where in the code might be the issue? I tried doing few joins on DF instead of SQL context but that also didn't help. Sometime job succeeds (like 5%).","created":"2015-11-05T00:56:27.243+0000"},{"body":"Abhishek, try disabling tungsten?\nie. sqlContext.setConf(\"spark.sql.tungsten.enabled\", \"false\")\n","created":"2015-11-06T19:17:18.815+0000"},{"body":"Did anybody find a solution for this? I also get a lot of these errors (well into my job running, arghhh). ","created":"2015-11-09T14:37:04.768+0000"},{"body":"it's working for me","created":"2015-11-09T16:28:49.707+0000"},{"body":"The current 1.6 branch also looks good.","created":"2015-11-09T16:30:56.627+0000"},{"body":"On Spark 1.5.2, We also face this issue when \"braodcast\" join being used in the DateFrame.\n\nWhy this fix is not merged into Spark 1.5.x release? On our case, the job fails eventually, so I have to disable tungsten by \"spark.sql.tungsten.enabled=false\"","created":"2016-03-20T14:43:44.959+0000"},{"body":"[~java8964] This patch is huge, also depends on other changes, it's not easy to backport to 1.5.x. Why not upgrade to 1.6?","created":"2016-03-21T05:22:52.808+0000"}],"conversations":[{"body":"*=== Update ===*\n\nThis is caused by a mismatch between `Runtime.getRuntime.availableProcessors()` and the number of active tasks in `ShuffleMemoryManager`. A quick reproduction is the following:\n\n{code}\n// My machine only has 8 cores\n$ bin/spark-shell --master local[32]\nscala> val df = sc.parallelize(Seq((1, 1), (2, 2))).toDF(\"a\", \"b\")\nscala> df.as(\"x\").join(df.as(\"y\"), $\"x.a\" === $\"y.a\").count()\n\nCaused by: java.io.IOException: Unable to acquire 2097152 bytes of memory\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:351)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.(UnsafeExternalSorter.java:138)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.create(UnsafeExternalSorter.java:106)\n\tat org.apache.spark.sql.execution.UnsafeExternalRowSorter.(UnsafeExternalRowSorter.java:68)\n\tat org.apache.spark.sql.execution.TungstenSort.org$apache$spark$sql$execution$TungstenSort$$preparePartition$1(sort.scala:120)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$2.apply(sort.scala:143)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$2.apply(sort.scala:143)\n\tat org.apache.spark.rdd.MapPartitionsWithPreparationRDD.prepare(MapPartitionsWithPreparationRDD.scala:50)\n{code}\n\n\n*=== Original ===*\n\nWhile running Q53 of TPCDS (scale = 1500) on 24 nodes cluster (12G memory on executor):\n\n{code}\njava.io.IOException: Unable to acquire 33554432 bytes of memory\n at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:368)\n at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.(UnsafeExternalSorter.java:138)\n at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.create(UnsafeExternalSorter.java:106)\n at org.apache.spark.sql.execution.UnsafeExternalRowSorter.(UnsafeExternalRowSorter.java:68)\n at org.apache.spark.sql.execution.TungstenSort.org$apache$spark$sql$execution$TungstenSort$$preparePartition$1(sort.scala:146)\n at org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n at org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n at org.apache.spark.rdd.MapPartitionsWithPreparationRDD.compute(MapPartitionsWithPreparationRDD.scala:45)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n at org.apache.spark.rdd.ZippedPartitionsRDD2.compute(ZippedPartitionsRDD.scala:88)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:88)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\nThe task could finished after retry.","from":"reporter","subject":"Some tasks failed with Unable to acquire memory"},{"body":"Why is this targeted for 1.6 ? We are finding this issue with basic Spark SQL executions in our applications. Is the expectation that Tungsten sort will be disabled in an upcoming checkin ?\n\nJob aborted due to stage failure: Task 1 in stage 25.0 failed 4 times, most recent failure: Lost task 1.3 in stage 25.0 (TID 3962, 39.6.64.17): java.io.IOException: Unable to acquire 16777216 bytes of memory\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:368)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.(UnsafeExternalSorter.java:138)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.create(UnsafeExternalSorter.java:106)\n\tat org.apache.spark.sql.execution.UnsafeExternalRowSorter.(UnsafeExternalRowSorter.java:68)\n\tat org.apache.spark.sql.execution.TungstenSort.org$apache$spark$sql$execution$TungstenSort$$preparePartition$1(sort.scala:146)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n\tat org.apache.spark.rdd.MapPartitionsWithPreparationRDD.prepare(MapPartitionsWithPreparationRDD.scala:50)\n\tat org.apache.spark.rdd.ZippedPartitionsBaseRDD$$anonfun$tryPrepareParents$1.applyOrElse(ZippedPartitionsRDD.scala:83)\n\tat org.apache.spark.rdd.ZippedPartitionsBaseRDD$$anonfun$tryPrepareParents$1.applyOrElse(ZippedPartitionsRDD.scala:82)\n\tat scala.runtime.AbstractPartialFunction.apply(AbstractPartialFunction.scala:33)\n\tat scala.collection.TraversableLike$$anonfun$collect$1.apply(TraversableLike.scala:278)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n","from":"developer"},{"body":"[~nadenf] In my case, the job finally finished (after retry), so this seems to be a blocker for me.\n\nCould you provide more information about you job?","from":"developer"},{"body":"We are using Spark Job Server to submit the job.\n\nEach job consists of:\n\n1) Execute an SQL statement against HDFS.\n2) Write the results into HDFS.\n3) Writes the result into MongoDB using just a normal Java adapter.\n\nWe do many of these in parallel. We have 6 node cluster with 30GB allocated to Spark (-xmx30g) and 60GB free. We constantly get these failures.","from":"developer"},{"body":"[~nadenf] Could you post the physical plan here? That could help us to understand the root cause.","from":"developer"},{"body":"This also could be related to https://issues.apache.org/jira/browse/SPARK-10341?filter=-2, could you test with 1.5-RC3?","from":"developer"},{"body":"Still working on the physical plan but we have been testing with the latest branch-1.5.0 releases which included this fix. It doesn't help.","from":"developer"},{"body":"[~nadenf] Thanks for letting us know, just realized that your stacktrace already including that fix.\n\nMaybe there are multiple join/aggregation/sort in your query? You can show the physical plan by `df.explain()` ","from":"developer"},{"body":"Is there currently any workaround this issue?\n\nI am also facing it with the last 1.5.1:\n{code:title=Error|borderStyle=solid}\nCaused by: java.io.IOException: Unable to acquire 33554432 bytes of memory at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:351) at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.(UnsafeExternalSorter.java:138) at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.create(UnsafeExternalSorter.java:106) at org.apache.spark.sql.execution.UnsafeKVExternalSorter.(UnsafeKVExternalSorter.java:74) at org.apache.spark.sql.execution.UnsafeKVExternalSorter.(UnsafeKVExternalSorter.java:56) at org.apache.spark.sql.execution.datasources.DynamicPartitionWriterContainer.writeRows(WriterContainer.scala:339) ... 8 more\n{code}","from":"developer"},{"body":"Has been difficult to get a clean stacktrace/explain trace because we are executing lots of SQL commands in parallel and we don't know which one is failing. We are absolutely doing lots of joins/aggregation/sorts. \n\nI have tried increase shuffle.memoryFraction to 0.8 but that didn't help.\n\nThis is still an issue with the latest Spark 1.5.2 branch.\n\n","from":"developer"},{"body":"Same issue, I got the following stacktrace:\n{noformat}\n15/10/20 18:20:43 INFO UnsafeExternalSorter: Thread 64 spilling sort data of 64.0 KB to disk (0 time so far)\n15/10/20 18:20:43 ERROR Executor: Exception in task 11.3 in stage 2.0 (TID 514)\njava.io.IOException: Unable to acquire 67108864 bytes of memory\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:351)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPageIfNecessary(UnsafeExternalSorter.java:332)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.insertKVRecord(UnsafeExternalSorter.java:461)\n\tat org.apache.spark.sql.execution.UnsafeKVExternalSorter.insertKV(UnsafeKVExternalSorter.java:139)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.switchToSortBasedAggregation(TungstenAggregationIterator.scala:489)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.processInputs(TungstenAggregationIterator.scala:379)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.start(TungstenAggregationIterator.scala:622)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1.org$apache$spark$sql$execution$aggregate$TungstenAggregate$$anonfun$$executePartition$1(TungstenAggregate.scala:110)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:119)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:119)\n\tat org.apache.spark.rdd.MapPartitionsWithPreparationRDD.compute(MapPartitionsWithPreparationRDD.scala:64)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:88)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{noformat}","from":"developer"},{"body":"I guess same issue here also.\n{code} \n15/10/26 15:11:33 INFO UnsafeExternalSorter: Thread 4524 spilling sort data of 64.0 KB to disk (0 time so far)\n15/10/26 15:11:33 INFO Executor: Executor is trying to kill task 135.0 in stage 394.0 (TID 11069)\n15/10/26 15:11:33 INFO UnsafeExternalSorter: Thread 4607 spilling sort data of 64.0 KB to disk (0 time so far)\n15/10/26 15:11:33 ERROR Executor: Managed memory leak detected; size = 67108864 bytes, TID = 11149\n15/10/26 15:11:33 ERROR Executor: Exception in task 92.3 in stage 394.0 (TID 11149)\njava.io.IOException: Unable to acquire 67108864 bytes of memory\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:351)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.(UnsafeExternalSorter.java:138)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.create(UnsafeExternalSorter.java:106)\n\tat org.apache.spark.sql.execution.UnsafeExternalRowSorter.(UnsafeExternalRowSorter.java:68)\n\tat org.apache.spark.sql.execution.TungstenSort.org$apache$spark$sql$execution$TungstenSort$$preparePartition$1(sort.scala:146)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n\tat org.apache.spark.rdd.MapPartitionsWithPreparationRDD.prepare(MapPartitionsWithPreparationRDD.scala:50)\n\tat org.apache.spark.rdd.ZippedPartitionsBaseRDD$$anonfun$tryPrepareParents$1.applyOrElse(ZippedPartitionsRDD.scala:83)\n\tat org.apache.spark.rdd.ZippedPartitionsBaseRDD$$anonfun$tryPrepareParents$1.applyOrElse(ZippedPartitionsRDD.scala:82)\n\tat scala.runtime.AbstractPartialFunction.apply(AbstractPartialFunction.scala:33)\n\tat scala.collection.TraversableLike$$anonfun$collect$1.apply(TraversableLike.scala:278)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n\tat scala.collection.TraversableLike$class.collect(TraversableLike.scala:278)\n\tat scala.collection.AbstractTraversable.collect(Traversable.scala:105)\n\tat org.apache.spark.rdd.ZippedPartitionsBaseRDD.tryPrepareParents(ZippedPartitionsRDD.scala:82)\n\tat org.apache.spark.rdd.ZippedPartitionsRDD2.compute(ZippedPartitionsRDD.scala:97)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsWithPreparationRDD.compute(MapPartitionsWithPreparationRDD.scala:63)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:88)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{code} ","from":"developer"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9241","from":"developer"},{"body":"Is there any work around for this issue. We migrated from 1.1 to 1.5 and our jobs heavily depends on join. I have been trying to get rid of this exception but no luck.\n\nIf someone can at least point where in the code might be the issue? I tried doing few joins on DF instead of SQL context but that also didn't help. Sometime job succeeds (like 5%).","from":"developer"},{"body":"Abhishek, try disabling tungsten?\nie. sqlContext.setConf(\"spark.sql.tungsten.enabled\", \"false\")\n","from":"developer"},{"body":"Did anybody find a solution for this? I also get a lot of these errors (well into my job running, arghhh). ","from":"developer"},{"body":"it's working for me","from":"developer"},{"body":"The current 1.6 branch also looks good.","from":"developer"},{"body":"On Spark 1.5.2, We also face this issue when \"braodcast\" join being used in the DateFrame.\n\nWhy this fix is not merged into Spark 1.5.x release? On our case, the job fails eventually, so I have to disable tungsten by \"spark.sql.tungsten.enabled=false\"","from":"developer"},{"body":"[~java8964] This patch is huge, also depends on other changes, it's not easy to backport to 1.5.x. Why not upgrade to 1.6?","from":"developer"}],"created":"2015-08-26T23:46:02.000+0000","description":"*=== Update ===*\n\nThis is caused by a mismatch between `Runtime.getRuntime.availableProcessors()` and the number of active tasks in `ShuffleMemoryManager`. A quick reproduction is the following:\n\n{code}\n// My machine only has 8 cores\n$ bin/spark-shell --master local[32]\nscala> val df = sc.parallelize(Seq((1, 1), (2, 2))).toDF(\"a\", \"b\")\nscala> df.as(\"x\").join(df.as(\"y\"), $\"x.a\" === $\"y.a\").count()\n\nCaused by: java.io.IOException: Unable to acquire 2097152 bytes of memory\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:351)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.(UnsafeExternalSorter.java:138)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.create(UnsafeExternalSorter.java:106)\n\tat org.apache.spark.sql.execution.UnsafeExternalRowSorter.(UnsafeExternalRowSorter.java:68)\n\tat org.apache.spark.sql.execution.TungstenSort.org$apache$spark$sql$execution$TungstenSort$$preparePartition$1(sort.scala:120)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$2.apply(sort.scala:143)\n\tat org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$2.apply(sort.scala:143)\n\tat org.apache.spark.rdd.MapPartitionsWithPreparationRDD.prepare(MapPartitionsWithPreparationRDD.scala:50)\n{code}\n\n\n*=== Original ===*\n\nWhile running Q53 of TPCDS (scale = 1500) on 24 nodes cluster (12G memory on executor):\n\n{code}\njava.io.IOException: Unable to acquire 33554432 bytes of memory\n at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.acquireNewPage(UnsafeExternalSorter.java:368)\n at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.(UnsafeExternalSorter.java:138)\n at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.create(UnsafeExternalSorter.java:106)\n at org.apache.spark.sql.execution.UnsafeExternalRowSorter.(UnsafeExternalRowSorter.java:68)\n at org.apache.spark.sql.execution.TungstenSort.org$apache$spark$sql$execution$TungstenSort$$preparePartition$1(sort.scala:146)\n at org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n at org.apache.spark.sql.execution.TungstenSort$$anonfun$doExecute$3.apply(sort.scala:169)\n at org.apache.spark.rdd.MapPartitionsWithPreparationRDD.compute(MapPartitionsWithPreparationRDD.scala:45)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n at org.apache.spark.rdd.ZippedPartitionsRDD2.compute(ZippedPartitionsRDD.scala:88)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:88)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\nThe task could finished after retry.","issue_id":"12859478","key":"SPARK-10309","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-10-30T06:41:01.000+0000","role":"fixed_distractor","summary":"Some tasks failed with Unable to acquire memory"} {"case_id":"12704673","cluster":"DISTRACTOR-SPARK-1065","comments":[{"body":"Here's some code that reproduces it.\n\n{code}\ntheconf = SparkConf().set(\"spark.executor.memory\", \"5g\").setAppName(\"broadcastfail\").setMaster(cluster_url)\nsc = SparkContext(conf=theconf)\n\nbroadcast_vals = []\nfor i in range(5):\n datas = [[float(i) for i in range(200)] for i in range(100000)]\n val = sc.broadcast(datas).value\n broadcast_vals.append(val)\n\nsc.parallelize([i for i in range(80)]).map(lambda x: sum([len(val) for val in broadcast_vals])).collect()\n{code}\n\nIt generates a few arrays of floats, each of which should take about 160 MB. The executors never end up using much memory, but the driver uses an enormous amount. Both the python and java processes ramp up to multiple GB until I start seeing a bunch of \"OutOfMemoryError: java heap space\".\n\nWith a single 160MB array, the job completes fine, but the driver still uses about 9 GB.\n","created":"2014-02-07T18:01:27.322+0000"},{"body":"I am facing the same issue in my project, where I use PySpark. As a proof of that the big objects I have could easily fit into nodes' memory, I am going to use dummy solution of saving my big objects into HDFS and load them on Python nodes.\n\nDoes anybody have an idea how to fix the issue in a better way? I don't have enough either Scala nor Java knowledge to fix this in Spark core. However, I feel like broadcast variables could be reimplemented on Python side though it seems a bit dangerous idea because we don't want to have separate implementations of one thing in both languages. That will also save memory, because while we use broadcasts through Scala we have 1 copy in JVM, 1 pickled copy in Python and 1 constructed object copy in Python.","created":"2014-08-11T22:06:09.695+0000"},{"body":"I have finished my experiment of using HDFS as a temp storage for my big objects. It showed that my mappers do not leak memory and work pretty stable. However, it takes long time to load these objects on each node.\n\nIs there straightforward way to detect memory leaks in Spark and PySpark?","created":"2014-08-12T17:18:09.031+0000"},{"body":"The broadcast was not used correctly in the above code, it should be used like this:\n\n{code}\nbroadcast_vals = []\nfor i in range(5):\n datas = [[float(i) for i in range(200)] for i in range(100000)]\n val = sc.broadcast(datas)\n broadcast_vals.append(val)\n\nsc.parallelize([i for i in range(80)]).map(lambda x: sum([len(val.value) for val in broadcast_vals])).collect()\n{code}\n\nThe reference of object in Python driver in not necessary in most cases, we will make it optional (no reference by default), then it can reduce the memory used in Python driver.","created":"2014-08-13T00:30:24.856+0000"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/1912","created":"2014-08-13T00:33:10.240+0000"},{"body":"[~davies] Will your PR take into account this fix: [SPARK-2521] Broadcast RDD object (instead of sending it along with every task) https://github.com/apache/spark/commit/7b8cd175254d42c8e82f0aa8eb4b7f3508d8fde2 ?\n\"The patch uses broadcast to send RDD objects and the closures to executors\"","created":"2014-08-13T00:45:39.268+0000"},{"body":"[~davies] I have not noticed that there was that mistake in the example, but I have not used that code. I run into the issue in my own code, where I use broadcasts correctly.\n\nI'm building your branch now and will try it right away. Thank you for your fix!","created":"2014-08-13T00:54:33.783+0000"},{"body":"[~frol], I think broadcast the RDD object is already done by that PR.\n\nBut the serialized closure will still be sent to JVM by py4j. After using broadcast for large datasets, the serialized closure should not be too huge, so I guess it will not be a big issue.","created":"2014-08-13T00:57:59.782+0000"},{"body":"After this patch, the above test can run successfully with about 700M memory in Python driver, 5xxMB memory in JVM driver, and 3G memory in python worker.\n\nIt may triggle another problem when run it with Mesos or YARN, because Spark does not reserve memory for Python worker, this may be fixed in 1.2 release.","created":"2014-08-13T01:02:25.188+0000"},{"body":"[~davies] I understand that if you use broadcast explicitly the closure won't be huge, but the point of that PR was also \"1. Users won't need to decide what to broadcast anymore, unless they would want to use a large object multiple times in different operations\".","created":"2014-08-13T01:04:02.776+0000"},{"body":"[~davies] I use YARN setup so I will see how it goes.","created":"2014-08-13T01:05:41.088+0000"},{"body":"[~davies] I have compiled and run your broadcast branch against my cluster on YARN. It does not leak memory any more! And it is at least 25% faster than my dummy implementation on top of HDFS. Implementation with broadcasts takes 4.5 minutes to finish a task, where my implementation took 6 minutes. More heavy tests are still working.","created":"2014-08-13T01:43:11.289+0000"},{"body":"Heavy tasks completed in 18 minutes each instead of 22 minutes, which is 20% speed up. That is nice!\nI don't see any problems on my YARN cluster. Java nodes eat up to 1.5GB RAM (which is my JVM limit) each and Python daemons eat around 650MB each. Though those numbers are still a bit weird, it is obvious to me that workers don't leak now.\n\nThank you a lot!","created":"2014-08-13T02:37:13.564+0000"},{"body":"Cool, thanks for the tests. If we can compress the data, it will be faster. I will do it in another separate PR.\n\nThe closure is serialized by cloudpickle, so it will be much slower if you do not use broadcast explicitly. We can show an warning if the serialized closure is too big.","created":"2014-08-13T04:57:18.711+0000"}],"conversations":[{"body":"PySpark's driver components may run out of memory when broadcasting large variables (say 1 gigabyte).\n\nBecause PySpark's broadcast is implemented on top of Java Spark's broadcast by broadcasting a pickled Python as a byte array, we may be retaining multiple copies of the large object: a pickled copy in the JVM and a deserialized copy in the Python driver.\n\nThe problem could also be due to memory requirements during pickling.\n\nPySpark is also affected by broadcast variables not being garbage collected. Adding an unpersist() method to broadcast variables may fix this: https://github.com/apache/incubator-spark/pull/543.\n\nAs a first step to fixing this, we should write a failing test to reproduce the error.\n\nThis was discovered by [~sandy]: [\"trouble with broadcast variables on pyspark\"|http://apache-spark-user-list.1001560.n3.nabble.com/trouble-with-broadcast-variables-on-pyspark-tp1301.html].","from":"reporter","subject":"PySpark runs out of memory with large broadcast variables"},{"body":"Here's some code that reproduces it.\n\n{code}\ntheconf = SparkConf().set(\"spark.executor.memory\", \"5g\").setAppName(\"broadcastfail\").setMaster(cluster_url)\nsc = SparkContext(conf=theconf)\n\nbroadcast_vals = []\nfor i in range(5):\n datas = [[float(i) for i in range(200)] for i in range(100000)]\n val = sc.broadcast(datas).value\n broadcast_vals.append(val)\n\nsc.parallelize([i for i in range(80)]).map(lambda x: sum([len(val) for val in broadcast_vals])).collect()\n{code}\n\nIt generates a few arrays of floats, each of which should take about 160 MB. The executors never end up using much memory, but the driver uses an enormous amount. Both the python and java processes ramp up to multiple GB until I start seeing a bunch of \"OutOfMemoryError: java heap space\".\n\nWith a single 160MB array, the job completes fine, but the driver still uses about 9 GB.\n","from":"developer"},{"body":"I am facing the same issue in my project, where I use PySpark. As a proof of that the big objects I have could easily fit into nodes' memory, I am going to use dummy solution of saving my big objects into HDFS and load them on Python nodes.\n\nDoes anybody have an idea how to fix the issue in a better way? I don't have enough either Scala nor Java knowledge to fix this in Spark core. However, I feel like broadcast variables could be reimplemented on Python side though it seems a bit dangerous idea because we don't want to have separate implementations of one thing in both languages. That will also save memory, because while we use broadcasts through Scala we have 1 copy in JVM, 1 pickled copy in Python and 1 constructed object copy in Python.","from":"developer"},{"body":"I have finished my experiment of using HDFS as a temp storage for my big objects. It showed that my mappers do not leak memory and work pretty stable. However, it takes long time to load these objects on each node.\n\nIs there straightforward way to detect memory leaks in Spark and PySpark?","from":"developer"},{"body":"The broadcast was not used correctly in the above code, it should be used like this:\n\n{code}\nbroadcast_vals = []\nfor i in range(5):\n datas = [[float(i) for i in range(200)] for i in range(100000)]\n val = sc.broadcast(datas)\n broadcast_vals.append(val)\n\nsc.parallelize([i for i in range(80)]).map(lambda x: sum([len(val.value) for val in broadcast_vals])).collect()\n{code}\n\nThe reference of object in Python driver in not necessary in most cases, we will make it optional (no reference by default), then it can reduce the memory used in Python driver.","from":"developer"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/1912","from":"developer"},{"body":"[~davies] Will your PR take into account this fix: [SPARK-2521] Broadcast RDD object (instead of sending it along with every task) https://github.com/apache/spark/commit/7b8cd175254d42c8e82f0aa8eb4b7f3508d8fde2 ?\n\"The patch uses broadcast to send RDD objects and the closures to executors\"","from":"developer"},{"body":"[~davies] I have not noticed that there was that mistake in the example, but I have not used that code. I run into the issue in my own code, where I use broadcasts correctly.\n\nI'm building your branch now and will try it right away. Thank you for your fix!","from":"developer"},{"body":"[~frol], I think broadcast the RDD object is already done by that PR.\n\nBut the serialized closure will still be sent to JVM by py4j. After using broadcast for large datasets, the serialized closure should not be too huge, so I guess it will not be a big issue.","from":"developer"},{"body":"After this patch, the above test can run successfully with about 700M memory in Python driver, 5xxMB memory in JVM driver, and 3G memory in python worker.\n\nIt may triggle another problem when run it with Mesos or YARN, because Spark does not reserve memory for Python worker, this may be fixed in 1.2 release.","from":"developer"},{"body":"[~davies] I understand that if you use broadcast explicitly the closure won't be huge, but the point of that PR was also \"1. Users won't need to decide what to broadcast anymore, unless they would want to use a large object multiple times in different operations\".","from":"developer"},{"body":"[~davies] I use YARN setup so I will see how it goes.","from":"developer"},{"body":"[~davies] I have compiled and run your broadcast branch against my cluster on YARN. It does not leak memory any more! And it is at least 25% faster than my dummy implementation on top of HDFS. Implementation with broadcasts takes 4.5 minutes to finish a task, where my implementation took 6 minutes. More heavy tests are still working.","from":"developer"},{"body":"Heavy tasks completed in 18 minutes each instead of 22 minutes, which is 20% speed up. That is nice!\nI don't see any problems on my YARN cluster. Java nodes eat up to 1.5GB RAM (which is my JVM limit) each and Python daemons eat around 650MB each. Though those numbers are still a bit weird, it is obvious to me that workers don't leak now.\n\nThank you a lot!","from":"developer"},{"body":"Cool, thanks for the tests. If we can compress the data, it will be faster. I will do it in another separate PR.\n\nThe closure is serialized by cloudpickle, so it will be much slower if you do not use broadcast explicitly. We can show an warning if the serialized closure is too big.","from":"developer"}],"created":"2014-02-07T11:41:54.000+0000","description":"PySpark's driver components may run out of memory when broadcasting large variables (say 1 gigabyte).\n\nBecause PySpark's broadcast is implemented on top of Java Spark's broadcast by broadcasting a pickled Python as a byte array, we may be retaining multiple copies of the large object: a pickled copy in the JVM and a deserialized copy in the Python driver.\n\nThe problem could also be due to memory requirements during pickling.\n\nPySpark is also affected by broadcast variables not being garbage collected. Adding an unpersist() method to broadcast variables may fix this: https://github.com/apache/incubator-spark/pull/543.\n\nAs a first step to fixing this, we should write a failing test to reproduce the error.\n\nThis was discovered by [~sandy]: [\"trouble with broadcast variables on pyspark\"|http://apache-spark-user-list.1001560.n3.nabble.com/trouble-with-broadcast-variables-on-pyspark-tp1301.html].","issue_id":"12704673","key":"SPARK-1065","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2014-08-17T00:00:30.000+0000","role":"fixed_distractor","summary":"PySpark runs out of memory with large broadcast variables"} {"case_id":"12704733","cluster":"DISTRACTOR-SPARK-1112","comments":[{"body":"Thanks for reporting this! Does Spark hang, or does the worker throw an exception? If the former, would you mind uploading the Spark worker log, and if the latter, can you add the stack trace?","created":"2014-02-20T15:47:07.435+0000"},{"body":"No Exception, and not \"hanging\" in the bad way the executors can sometime hang : if I kill the driver, the workers receive the shutdown signal and exit cleanly.\n\nHere are the logs :\n\nDRIVER :\n\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:13 as 2083 bytes in 0 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:14 as TID 2294 on executor 1: t4.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:14 as 2083 bytes in 0 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:15 as TID 2295 on executor 4: t3.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:15 as 2083 bytes in 1 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:16 as TID 2296 on executor 0: t0.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:16 as 2083 bytes in 2 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:17 as TID 2297 on executor 3: t1.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:17 as 2083 bytes in 1 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:18 as TID 2298 on executor 2: t5.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:18 as 2083 bytes in 1 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:19 as TID 2299 on executor 5: t6.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:19 as 2083 bytes in 1 ms\n\nEXECUTOR :\n14/02/19 15:21:53 INFO Executor: Serialized size of result for 2287 is 17229427\n14/02/19 15:21:53 INFO Executor: Sending result for 2287 directly to driver\n14/02/19 15:21:53 INFO Executor: Serialized size of result for 2299 is 17229262\n14/02/19 15:21:53 INFO Executor: Sending result for 2299 directly to driver\n14/02/19 15:21:53 INFO Executor: Finished task ID 2299\n14/02/19 15:21:53 INFO Executor: Finished task ID 2287\n14/02/19 15:21:53 INFO Executor: Serialized size of result for 2281 is 17229426\n14/02/19 15:21:53 INFO Executor: Sending result for 2281 directly to driver\n14/02/19 15:21:53 INFO Executor: Finished task ID 2281\n\nThere is a timezone difference between driver & executor\n\n","created":"2014-02-20T23:03:43.133+0000"},{"body":"I have a similar issue, I'm on spark-0.9.0 compiled with cdh-4.2.1\n\nFor me, serialized tasks over 10 MB do not reach executors. I've tried this with spark.akka.frameSize set to 160 and 10. The workaround suggested (setting spark.akka.frameSize to 10) does not work for me.\n\nI've confirmed that even if serialized tasks are just under 10MB, the executors do get them and the task is completed.\n\nSpark hangs. There are no exceptions or unusual ERROR/WARN/DEBUG logs in the driver, master, executor or worker daemon logs. The executors just don't seem to have received the tasks. The application UI shows the task status as running, but never progresses.\n\nHere are the last few lines from my driver and one executor:\n\nDRIVER:\n14/02/21 06:37:00 INFO TaskSetManager: Finished TID 797 in 53897 ms on spark-slave01 (progress: 78/80)\n14/02/21 06:37:00 INFO DAGScheduler: Completed ResultTask(9, 58)\n14/02/21 06:37:08 INFO TaskSetManager: Finished TID 768 in 75767 ms on spark-slave02 (progress: 79/80)\n14/02/21 06:37:08 INFO TaskSchedulerImpl: Remove TaskSet 9.0 from pool \n14/02/21 06:37:08 INFO DAGScheduler: Completed ResultTask(9, 69)\n14/02/21 06:37:08 INFO DAGScheduler: Stage 9 (reduceByKeyLocally at SKMeans.scala:174) finished in 99.048 s\n14/02/21 06:37:08 INFO SparkContext: Job finished: reduceByKeyLocally at SKMeans.scala:174, took 99.359019444 s\n14/02/21 06:37:09 INFO SparkContext: Starting job: reduceByKeyLocally at SKMeans.scala:174\n14/02/21 06:37:09 INFO DAGScheduler: Got job 7 (reduceByKeyLocally at SKMeans.scala:174) with 80 output partitions (allowLocal=false)\n14/02/21 06:37:09 INFO DAGScheduler: Final stage: Stage 10 (reduceByKeyLocally at SKMeans.scala:174)\n14/02/21 06:37:09 INFO DAGScheduler: Parents of final stage: List()\n14/02/21 06:37:09 INFO DAGScheduler: Missing parents: List()\n14/02/21 06:37:09 INFO DAGScheduler: Submitting Stage 10 (MapPartitionsRDD[28] at reduceByKeyLocally at SKMeans.scala:174), which has no missing parents\n14/02/21 06:37:10 INFO DAGScheduler: Submitting 80 missing tasks from Stage 10 (MapPartitionsRDD[28] at reduceByKeyLocally at SKMeans.scala:174)\n14/02/21 06:37:10 INFO TaskSchedulerImpl: Adding task set 10.0 with 80 tasks\n14/02/21 06:37:10 INFO TaskSetManager: Starting task 10.0:0 as TID 800 on executor 2: spark-slave01 (PROCESS_LOCAL)\n14/02/21 06:37:10 INFO TaskSetManager: Serialized task 10.0:0 as 10700743 bytes in 19 ms\n...\n\nEXECUTOR: spark-slave01\n14/02/21 06:36:08 DEBUG Executor: Task 798's epoch is 3\n14/02/21 06:36:08 DEBUG CacheManager: Looking for partition rdd_4_60\n14/02/21 06:36:08 DEBUG BlockManager: Getting local block rdd_4_60\n14/02/21 06:36:08 DEBUG BlockManager: Level for block rdd_4_60 is StorageLevel(false, true, true, 1)\n14/02/21 06:36:08 DEBUG BlockManager: Getting block rdd_4_60 from memory\n14/02/21 06:36:08 INFO BlockManager: Found block rdd_4_60 locally\n14/02/21 06:36:21 INFO Executor: Serialized size of result for 765 is 1224222\n14/02/21 06:36:21 INFO Executor: Sending result for 765 directly to driver\n14/02/21 06:36:21 INFO Executor: Finished task ID 765\n14/02/21 06:36:34 INFO Executor: Serialized size of result for 790 is 1262463\n14/02/21 06:36:34 INFO Executor: Sending result for 790 directly to driver\n14/02/21 06:36:34 INFO Executor: Finished task ID 790\n14/02/21 06:36:37 INFO Executor: Serialized size of result for 784 is 1394816\n14/02/21 06:36:37 INFO Executor: Sending result for 784 directly to driver\n14/02/21 06:36:37 INFO Executor: Finished task ID 784\n14/02/21 06:36:38 INFO Executor: Serialized size of result for 787 is 1409571\n14/02/21 06:36:38 INFO Executor: Sending result for 787 directly to driver\n14/02/21 06:36:38 INFO Executor: Finished task ID 787\n14/02/21 06:36:41 INFO Executor: Serialized size of result for 798 is 1270321\n14/02/21 06:36:41 INFO Executor: Sending result for 798 directly to driver\n14/02/21 06:36:41 INFO Executor: Finished task ID 798\n14/02/21 06:36:50 INFO Executor: Serialized size of result for 792 is 1175064\n14/02/21 06:36:50 INFO Executor: Sending result for 792 directly to driver\n14/02/21 06:36:50 INFO Executor: Finished task ID 792\n14/02/21 06:36:52 INFO Executor: Serialized size of result for 794 is 1485354\n14/02/21 06:36:52 INFO Executor: Sending result for 794 directly to driver\n14/02/21 06:36:52 INFO Executor: Finished task ID 794\n14/02/21 06:37:00 INFO Executor: Serialized size of result for 797 is 1615486\n14/02/21 06:37:00 INFO Executor: Sending result for 797 directly to driver\n14/02/21 06:37:00 INFO Executor: Finished task ID 797\n\nRoshan","created":"2014-02-20T23:14:31.051+0000"},{"body":"Roshan, the issue you're seeing is different -- Guillaume's issue is when task results are too large to be sent using Akka (in which case Spark should use a different code path to send task results to the executor); your issue is when the task itself is too large, in which case Spark (in theory!) gives up and throw an error. We should fix both problems, but would you mind opening a separate issue?","created":"2014-02-20T23:54:06.696+0000"},{"body":"Guillaume, just to clarify, the logs you pasted above are for when you set the maximum frame size to 16MiB? I'm asking because the task results seem to be just slightly larger than 16MiB, which isn't the failure case you mentioned in your description.","created":"2014-02-20T23:56:32.547+0000"},{"body":"Sorry, I should have specified it. The frameSize was set to 512 for those logs. I've also tried with 16, and it works when results are over 16MB","created":"2014-02-21T00:04:27.711+0000"},{"body":"Cool thanks for clarifying! Looking into this...","created":"2014-02-21T00:06:49.687+0000"},{"body":"When I set BOTH the property on driver with \n\nSystem.setProperty(\"spark.akka.frameSize\", 128) AND I pass the env parameter to SparkContext with SPARK_JAVA_OPTS = \"-Dspark.akka.frameSize=128\" \n\nThen it seems to works.\n\nSo maybe the problem comes from properties not being passed correctly to workers when executors are instanciated ?\n\nAlso, I'm using packaged binary distribution for CDH4 on a standalone cluster","created":"2014-02-21T01:04:08.692+0000"},{"body":"Hi,\n\nI just realized my driver wasn't picking up spark.akka.frameSize value, because of a problem in the way I was passing it in. However, my executor's were picking this value correctly from their conf/spark-env.sh files.\n\nNow, both sides, the driver and executors print the correct value for frameSize with spark.akka.logAkkaConfig=true.\n\nI also noticed that simply starting the driver with java -Dspark.akka.frameSize=200 does not propagate this automatically to the executors. Not that this is an issue. I guess I was just confused about the configuration.\n\nGuillaume, seems like you have the reverse situation as mine, ie. your drivers are correctly configured with the right frameSize, but the executors are still using the 10MB default?\n\nTo conclude, after ensuring that the driver is correctly configured with the right frameSize, so far, serialized tasks larger than 10MB are being received by the executors and run successfully.\n\nRoshan","created":"2014-02-21T01:17:49.349+0000"},{"body":"You're right Roshan.I was expecting the akka properties set before SparkContext creation to be propagated to the Executors (and based on what the Spark code does, it should be the case). \n\nI think it should be enforced for the whole akka stuff (timeouts and so on), as well as for the rest of spark properties","created":"2014-02-21T05:35:36.780+0000"},{"body":"Hey There,\n\nI spent some time playing with this and couldn't reproduce the issue. The driver should capture the options and pass them to executors. I just tested this with a local cluster and was able to verify the akka frame size is passed to executors even if it's not set in spark-env.sh where the executor launches.\n\n[~roshan] - what happens if you remove the setting from spark-env.sh on the executors and only set it at the driver. Does that work correctly?","created":"2014-02-21T23:33:12.526+0000"},{"body":"In my code I already had some options passed in SPARK_JAVA_OPTS in the env of the SparkContext, but nothing about the akka frameSize. Could it be related ?\n\nSince it was working correctly in 0.8.1, maybe it's related to the new SparkConf ?","created":"2014-02-22T00:16:23.984+0000"},{"body":"[~guillaumepitel] If you aren't setting akka.frameSize in SPARK_JAVA_OPTS then where are you setting it?\n\nWhat I was saying is that if you do System.setProperty(spark.akka.frameSize, XX) before you create the SparkContext it should collect this and send it to the executors correctly. One thing is if you set it after you create the SparkContext it won't work... are you doing this by any chance?","created":"2014-02-22T00:23:20.471+0000"},{"body":"No, I create the SparkContext after setting the properties (and I've nothing on my nodes for configuring the frameSize, it's a per-process configuration). \n\nSo before my workaround, I was just setting the property (and the environment in the UI was showing the right value) and passing a SPARK_JAVA_OPTS to the env of the SparkContext with \n{code}\nSystem.setProperty(\"spark.serializer\", \"org.apache.spark.serializer.KryoSerializer\")\nSystem.setProperty(\"spark.kryo.registrator\", registrator)\nSystem.setProperty(\"spark.kryo.referenceTracking\", \"false\")\nSystem.setProperty(\"spark.kryoserializer.buffer.mb\", bufferSize.toString)\nSystem.setProperty(\"spark.locality.wait\", \"10000\")\nSystem.setProperty(\"spark.hadoop.mapreduce.output.fileoutputformat.compress\", \"true\")\nSystem.setProperty(\"spark.hadoop.mapreduce.output.fileoutputformat.compress.codec\", codec)\nSystem.setProperty(\"spark.hadoop.mapreduce.output.fileoutputformat.compress.type\", \"BLOCK\")\nSystem.setProperty(\"spark.akka.frameSize\", akkaFrameSize.toString)\n\nval sb = new StringBuilder()\nsb.append(\"-Dspark.storage.memoryFraction=\" + sparkMemoryFraction())\nsb.append(\" -Dspark.worker.timeout=\" + sparkWorkerTimeout())\nsb.append(\" -Dspark.akka.askTimeout=\" + sparkAkkaAskTimeout())\nsb.append(\" -Dspark.akka.timeout=\" + sparkAkkaTimeout())\nsb.append(\" -Dspark.shuffle.consolidateFiles=\" + sparkShuffleConsolidateFiles())\n\nvar env = new HashMap[String, String]()\nenv += \"SPARK_JAVA_OPTS\" -> sb.toString()\n\nval sc = new SparkContext(sparkMaster(), appName, sparkHome(), jars(), env)\n{code}\n\nThat caused a problem (the akka frame size seemed to be passed to the executor, but only after the creation of the actorSystem, because it was taking the right code path, but the akka system didn't seem to be properly configured).\n\nNow if I add this to my code :\n\n{code}\nsb.append(\" -Dspark.akka.frameSize=\" + akkaFrameSize)\n{code}\n\nIt works\n","created":"2014-02-22T00:39:31.475+0000"},{"body":"@Patrick Wendell -\n\nCase 1.\nDRIVER:\nI start the driver with java -Dspark.akka.frameSize=200 -D... -cp .....\n\nEXECUTOR:\nspark-env.sh on workers with frameSize specified: \nexport SPARK_JAVA_OPTS='-Dspark.local.dir=/data/spark/tmp -Dspark.akka.logAkkaConfig=true -Dspark.akka.frameSize=160 -Dspark.akka.timeout=100 -Dspark.akka.askTimeout=30 -Dspark.akka.logLifecycleEvents=true -Dspark.worker.timeout=200'\n\nExecutor akka config log reports frameSize=160\n\nCase 2. \nDRIVER:\nI start the driver with java -Dspark.akka.frameSize=200 -D... -cp ....\n\nEXECUTOR:\nspark-env.sh on workers with frameSize NOT specified: \nexport SPARK_JAVA_OPTS='-Dspark.local.dir=/data/spark/tmp -Dspark.akka.logAkkaConfig=true -Dspark.akka.timeout=100 -Dspark.akka.askTimeout=30 -Dspark.akka.logLifecycleEvents=true -Dspark.worker.timeout=200'\n\nExecutor akka config log on workers has no frameSize and executor process(by ps waux) has no frameSize. Here, I suppose it picks up the default 10MB.\n\nCase 3. \nDRIVER:\nThis time I export SPARK_JAVA_OPTS=\"-Dspark.akka.frameSize=220\" on the dirver host, before running the driver with java -Dspark.akka.frameSize=200 -D... -cp ....\nI'm not reading SPARK_JAVA_OPTS in the driver code.\nDriver's Akka config log, says frameSize=200.\n\nEXECUTOR1 on HOST1:\nNo frameSize specified in spark-env\nExecutor akka config log says frameSize=220.\n\nEXECUTOR2 on HOST2:\nframeSize=160 in spark-env\nExecutor akka config log says frameSize=220, \n\nps waux shows the process was started with this cmd:\n/usr/java/default/bin/java -cp sparkJar.jar -Dspark.local.dir=/data/spark/tmp -Dspark.akka.logAkkaConfig=true -Dspark.akka.frameSize=160 -Dspark.akka.timeout=100 -Dspark.akka.askTimeout=30 -Dspark.akka.logLifecycleEvents=true -Dspark.worker.timeout=200 -Dspark.akka.frameSize=220 -Xms3072M -Xmx3072M org.apache.spark.executor.CoarseGrainedExecutorBackend .....\n\nSo the SPARK_JAVA_OPTS from both the executor and the driver are appended, but since java overrides the first frameSize with the second, the executor runs with frameSize=220\n\nCase 4.\nDRIVER:\nThis time I export SPARK_JAVA_OPTS=\"-Dspark.akka.frameSize=220\" on the driver host, but don't specify frameSize in the launch command.\nAgain, I don't SPARK_JAVA_OPTS in the driver code.\nNo frameSize in driver's akka config log. I expect driver picks up the default 10MB frameSize.\n\nExecutors are the same as in case3.\n\nWhat I'm seeing is that, you can specify the frameSize for a driver in the launch command as java -Dspark.akka.frameSize property, and for the executors in their respective spark-env.sh. If you specify SPARK_JAVA_OPTS on the driver side, then this will override the value for the executors, but not for the driver. I don't use spark-class.sh to launch my driver, because somewhere I read that its meant for spark internal classes and examples. I also currently build my SparkContext directly, without using SparkConf.\n\nThis not an issue for me any longer. That being said, it would have saved me loads of time, if the logs had provided some indication that sending serialized tasks to the executors had failed because they were larger than 10MB. Also, the documentation could be a bit clearer about what is set where and which property overrides or is overridden.\n\nThanks.","created":"2014-02-22T03:58:52.640+0000"},{"body":"Hi all, \n\nI'm very new to Spark and doing some tests, I've experienced similar issue.\n(tested with Spark Shell, 0.9.1, r3.8xlarge instance on EC2 - 32 core / 244GiB MEM)\n\nI was trying to broadcast 700MB of data and Spark hangs when I run collect() method for the data. \n\nHere's the strange things :\n1) when I tried \n{code}val userInfo = sc.textFile(\"file:///spark/logs/user_sign_up2.csv\").map{line => val split = line.split(\",\"); (split(1), split)}\nval userInfoMap = userInfo.collectAsMap\n{code}\nit runs well.\n2) when I tried \n{code}val userInfo = sc.textFile(\"file:///spark/logs/user_sign_up2.csv\").map{line => val split = line.split(\",\"); (split(1), split(5))} \nval userInfoMap = userInfo.collectAsMap\n{code}\nSpark hangs.\n3) when I slightly control the data size using sample() method or cutting the data file, it runs well. \n\nOur team investigated logs from master and worker then we found worker finished all tasks but master couldn't retrieve the result from a task the result size larger than 10MB\n\nWe tried to apply the workaround setting spark.akka.frameSize to 9, it works like a charm.\n\nI guess it might hard to reproduce the issue, please contact me if there's need of testing or getting logs. \n\nThanks!","created":"2014-05-29T01:59:32.583+0000"},{"body":"I'm curious, why did you want to make the frameSize this big -- are the tasks themselves also big or just the results? There might be other buffers in Akka that can't be made bigger than this. It's possible that this changed in a newer Akka version (because larger frame sizes used to work before).","created":"2014-05-29T02:17:56.950+0000"},{"body":"[~matei]\nI've found the default of spark.akka.frameSize is 10 from the config document, http://spark.apache.org/docs/0.9.1/configuration.html\njust tried to slightly larger and smaller (11 and 9) values.\n\nI did collect() method on the userInfo and it might contains large data. (edited the first comment.)\n\n","created":"2014-05-29T02:49:45.128+0000"},{"body":"[~matei] Do you know which akka version we should use to be able to use big frame size. ","created":"2014-06-13T23:00:20.620+0000"},{"body":"To follow up this thread, I have done some experiments when the frameSize is around 10MB .\n\n1) spark.akka.frameSize = 10\nIf one of the partition size is very close to 10MB, say 9.97MB, the execution blocks without any exception or warning. Worker finished the task to send the serialized result, and then throw exception saying hadoop IPC client connection stops (changing the logging to debug level). However, the master never receives the results and the program just hangs.\nBut if sizes for all the partitions less than some number btw 9.96MB amd 9.97MB, the program works fine.\n2) spark.akka.frameSize = 9\nwhen the partition size is just a little bit smaller than 9MB, it fails as well.\n\nThis bug behavior is not exactly what spark-1112 is about, could you please guide me how to open a separate bug when the serialization size is very close to 10MB. \n\nThanks a lot","created":"2014-06-16T20:09:57.797+0000"},{"body":"I have filed a bug https://issues.apache.org/jira/browse/SPARK-2156","created":"2014-06-19T00:10:33.592+0000"},{"body":"We were able to reproduce this - thanks for reporting it.","created":"2014-06-19T01:38:18.592+0000"},{"body":"Awesome, looking forward to the fix. At least better error or exception message would be helpful. ","created":"2014-06-19T02:06:31.905+0000"},{"body":"PR: https://github.com/apache/spark/pull/1124","created":"2014-06-19T02:35:30.196+0000"},{"body":"This is fixed in the 1.0 branch via:\nhttps://github.com/apache/spark/pull/1172","created":"2014-06-23T02:49:36.023+0000"},{"body":"Fixed in 1.1.0 via:\nhttps://github.com/apache/spark/pull/1132","created":"2014-06-25T02:06:52.586+0000"},{"body":"Can a clear workaround be specified for this bug please? For those unable to upgrade to run on 1.0.1 or 1.1.0 in production, general instructions on the workaround are required. This is a huge blocker for current production deployments (even on 1.0.0) otherwise. For instance, running a saveAsTextFile() on an RDD (~400MB) causes execution to freeze with the last log statements seen on the driver being:\n\n14/06/25 16:38:55 INFO spark.SparkContext: Starting job: saveAsTextFile at Test.java:99\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Got job 6 (saveAsTextFile at Test.java:99) with 2 output partitions (allowLocal=false)\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Final stage: Stage 6(saveAsTextFile at Test.java:99)\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Parents of final stage: List()\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Missing parents: List()\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Submitting Stage 6 (MappedRDD[558] at saveAsTextFile at Test.java:99), which has no missing parents\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Submitting 2 missing tasks from Stage 6 (MappedRDD[558] at saveAsTextFile at Test.java:99)\n14/06/25 16:38:55 INFO scheduler.TaskSchedulerImpl: Adding task set 6.0 with 2 tasks\n14/06/25 16:38:55 INFO scheduler.TaskSetManager: Starting task 6.0:0 as TID 5 on executor 1: somehost.corp (PROCESS_LOCAL)\n14/06/25 16:38:55 INFO scheduler.TaskSetManager: Serialized task 6.0:0 as 351777 bytes in 36 ms\n14/06/25 16:38:55 INFO scheduler.TaskSetManager: Starting task 6.0:1 as TID 6 on executor 0: someotherhost.corp (PROCESS_LOCAL)\n14/06/25 16:38:55 INFO scheduler.TaskSetManager: Serialized task 6.0:1 as 186453 bytes in 16 ms\n\nThe test setup for reproducing this issue has two slaves (each with 24G) running spark standalone. The driver runs with Xmx 4G.\n\nThanks.","created":"2014-06-25T16:48:06.251+0000"},{"body":"[~reachbach] If you are running on standalone mode, it might work if you go on every node in your cluster and add the following to spark-env.sh:\n\n{code}\nexport SPARK_JAVA_OPTS=\"-Dspark.akka.frameSize=XXX\"\n{code}\n\nHowever, this work around will only work if every job in your cluster is using the same frame size (XXX).\n\nThe main recommendation is to upgrade to 1.0.1. We are very conservative about what we merge into maintenance branches, so we recommend users upgrade immediately once we release them.","created":"2014-06-25T16:57:54.687+0000"},{"body":"This is not resolved yet because it needs to be back ported into 0.9","created":"2014-06-25T16:59:17.584+0000"},{"body":"PR for branch-0.9: https://github.com/apache/spark/pull/1455","created":"2014-07-17T04:01:52.734+0000"},{"body":"Issue resolved by pull request 1455\n[https://github.com/apache/spark/pull/1455]","created":"2014-07-17T04:31:18.367+0000"},{"body":"Does anyone test in version0.9.2,I found it also failed , while v1.0.1 & v1.1.0 is ok. ","created":"2014-07-24T14:42:11.531+0000"},{"body":"Looks like the \"Fix Versions\" accidentally got overwritten during a backport / cherry-pick, so I've restored them based on the issue history.","created":"2014-12-01T08:29:53.087+0000"}],"conversations":[{"body":"When I set the spark.akka.frameSize to something over 10, the messages sent from the executors to the driver completely block the execution if the message is bigger than 10MiB and smaller than the frameSize (if it's above the frameSize, it's ok)\n\nWorkaround is to set the spark.akka.frameSize to 10. In this case, since 0.8.1, the blockManager deal with the data to be sent. It seems slower than akka direct message though.\n\nThe configuration seems to be correctly read (see actorSystemConfig.txt), so I don't see where the 10MiB could come from ","from":"reporter","subject":"When spark.akka.frameSize > 10, task results bigger than 10MiB block execution"},{"body":"Thanks for reporting this! Does Spark hang, or does the worker throw an exception? If the former, would you mind uploading the Spark worker log, and if the latter, can you add the stack trace?","from":"developer"},{"body":"No Exception, and not \"hanging\" in the bad way the executors can sometime hang : if I kill the driver, the workers receive the shutdown signal and exit cleanly.\n\nHere are the logs :\n\nDRIVER :\n\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:13 as 2083 bytes in 0 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:14 as TID 2294 on executor 1: t4.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:14 as 2083 bytes in 0 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:15 as TID 2295 on executor 4: t3.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:15 as 2083 bytes in 1 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:16 as TID 2296 on executor 0: t0.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:16 as 2083 bytes in 2 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:17 as TID 2297 on executor 3: t1.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:17 as 2083 bytes in 1 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:18 as TID 2298 on executor 2: t5.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:18 as 2083 bytes in 1 ms\n14/02/19 16:21:49 INFO TaskSetManager: Starting task 12.0:19 as TID 2299 on executor 5: t6.exensa.loc (PROCESS_LOCAL)\n14/02/19 16:21:49 INFO TaskSetManager: Serialized task 12.0:19 as 2083 bytes in 1 ms\n\nEXECUTOR :\n14/02/19 15:21:53 INFO Executor: Serialized size of result for 2287 is 17229427\n14/02/19 15:21:53 INFO Executor: Sending result for 2287 directly to driver\n14/02/19 15:21:53 INFO Executor: Serialized size of result for 2299 is 17229262\n14/02/19 15:21:53 INFO Executor: Sending result for 2299 directly to driver\n14/02/19 15:21:53 INFO Executor: Finished task ID 2299\n14/02/19 15:21:53 INFO Executor: Finished task ID 2287\n14/02/19 15:21:53 INFO Executor: Serialized size of result for 2281 is 17229426\n14/02/19 15:21:53 INFO Executor: Sending result for 2281 directly to driver\n14/02/19 15:21:53 INFO Executor: Finished task ID 2281\n\nThere is a timezone difference between driver & executor\n\n","from":"developer"},{"body":"I have a similar issue, I'm on spark-0.9.0 compiled with cdh-4.2.1\n\nFor me, serialized tasks over 10 MB do not reach executors. I've tried this with spark.akka.frameSize set to 160 and 10. The workaround suggested (setting spark.akka.frameSize to 10) does not work for me.\n\nI've confirmed that even if serialized tasks are just under 10MB, the executors do get them and the task is completed.\n\nSpark hangs. There are no exceptions or unusual ERROR/WARN/DEBUG logs in the driver, master, executor or worker daemon logs. The executors just don't seem to have received the tasks. The application UI shows the task status as running, but never progresses.\n\nHere are the last few lines from my driver and one executor:\n\nDRIVER:\n14/02/21 06:37:00 INFO TaskSetManager: Finished TID 797 in 53897 ms on spark-slave01 (progress: 78/80)\n14/02/21 06:37:00 INFO DAGScheduler: Completed ResultTask(9, 58)\n14/02/21 06:37:08 INFO TaskSetManager: Finished TID 768 in 75767 ms on spark-slave02 (progress: 79/80)\n14/02/21 06:37:08 INFO TaskSchedulerImpl: Remove TaskSet 9.0 from pool \n14/02/21 06:37:08 INFO DAGScheduler: Completed ResultTask(9, 69)\n14/02/21 06:37:08 INFO DAGScheduler: Stage 9 (reduceByKeyLocally at SKMeans.scala:174) finished in 99.048 s\n14/02/21 06:37:08 INFO SparkContext: Job finished: reduceByKeyLocally at SKMeans.scala:174, took 99.359019444 s\n14/02/21 06:37:09 INFO SparkContext: Starting job: reduceByKeyLocally at SKMeans.scala:174\n14/02/21 06:37:09 INFO DAGScheduler: Got job 7 (reduceByKeyLocally at SKMeans.scala:174) with 80 output partitions (allowLocal=false)\n14/02/21 06:37:09 INFO DAGScheduler: Final stage: Stage 10 (reduceByKeyLocally at SKMeans.scala:174)\n14/02/21 06:37:09 INFO DAGScheduler: Parents of final stage: List()\n14/02/21 06:37:09 INFO DAGScheduler: Missing parents: List()\n14/02/21 06:37:09 INFO DAGScheduler: Submitting Stage 10 (MapPartitionsRDD[28] at reduceByKeyLocally at SKMeans.scala:174), which has no missing parents\n14/02/21 06:37:10 INFO DAGScheduler: Submitting 80 missing tasks from Stage 10 (MapPartitionsRDD[28] at reduceByKeyLocally at SKMeans.scala:174)\n14/02/21 06:37:10 INFO TaskSchedulerImpl: Adding task set 10.0 with 80 tasks\n14/02/21 06:37:10 INFO TaskSetManager: Starting task 10.0:0 as TID 800 on executor 2: spark-slave01 (PROCESS_LOCAL)\n14/02/21 06:37:10 INFO TaskSetManager: Serialized task 10.0:0 as 10700743 bytes in 19 ms\n...\n\nEXECUTOR: spark-slave01\n14/02/21 06:36:08 DEBUG Executor: Task 798's epoch is 3\n14/02/21 06:36:08 DEBUG CacheManager: Looking for partition rdd_4_60\n14/02/21 06:36:08 DEBUG BlockManager: Getting local block rdd_4_60\n14/02/21 06:36:08 DEBUG BlockManager: Level for block rdd_4_60 is StorageLevel(false, true, true, 1)\n14/02/21 06:36:08 DEBUG BlockManager: Getting block rdd_4_60 from memory\n14/02/21 06:36:08 INFO BlockManager: Found block rdd_4_60 locally\n14/02/21 06:36:21 INFO Executor: Serialized size of result for 765 is 1224222\n14/02/21 06:36:21 INFO Executor: Sending result for 765 directly to driver\n14/02/21 06:36:21 INFO Executor: Finished task ID 765\n14/02/21 06:36:34 INFO Executor: Serialized size of result for 790 is 1262463\n14/02/21 06:36:34 INFO Executor: Sending result for 790 directly to driver\n14/02/21 06:36:34 INFO Executor: Finished task ID 790\n14/02/21 06:36:37 INFO Executor: Serialized size of result for 784 is 1394816\n14/02/21 06:36:37 INFO Executor: Sending result for 784 directly to driver\n14/02/21 06:36:37 INFO Executor: Finished task ID 784\n14/02/21 06:36:38 INFO Executor: Serialized size of result for 787 is 1409571\n14/02/21 06:36:38 INFO Executor: Sending result for 787 directly to driver\n14/02/21 06:36:38 INFO Executor: Finished task ID 787\n14/02/21 06:36:41 INFO Executor: Serialized size of result for 798 is 1270321\n14/02/21 06:36:41 INFO Executor: Sending result for 798 directly to driver\n14/02/21 06:36:41 INFO Executor: Finished task ID 798\n14/02/21 06:36:50 INFO Executor: Serialized size of result for 792 is 1175064\n14/02/21 06:36:50 INFO Executor: Sending result for 792 directly to driver\n14/02/21 06:36:50 INFO Executor: Finished task ID 792\n14/02/21 06:36:52 INFO Executor: Serialized size of result for 794 is 1485354\n14/02/21 06:36:52 INFO Executor: Sending result for 794 directly to driver\n14/02/21 06:36:52 INFO Executor: Finished task ID 794\n14/02/21 06:37:00 INFO Executor: Serialized size of result for 797 is 1615486\n14/02/21 06:37:00 INFO Executor: Sending result for 797 directly to driver\n14/02/21 06:37:00 INFO Executor: Finished task ID 797\n\nRoshan","from":"developer"},{"body":"Roshan, the issue you're seeing is different -- Guillaume's issue is when task results are too large to be sent using Akka (in which case Spark should use a different code path to send task results to the executor); your issue is when the task itself is too large, in which case Spark (in theory!) gives up and throw an error. We should fix both problems, but would you mind opening a separate issue?","from":"developer"},{"body":"Guillaume, just to clarify, the logs you pasted above are for when you set the maximum frame size to 16MiB? I'm asking because the task results seem to be just slightly larger than 16MiB, which isn't the failure case you mentioned in your description.","from":"developer"},{"body":"Sorry, I should have specified it. The frameSize was set to 512 for those logs. I've also tried with 16, and it works when results are over 16MB","from":"developer"},{"body":"Cool thanks for clarifying! Looking into this...","from":"developer"},{"body":"When I set BOTH the property on driver with \n\nSystem.setProperty(\"spark.akka.frameSize\", 128) AND I pass the env parameter to SparkContext with SPARK_JAVA_OPTS = \"-Dspark.akka.frameSize=128\" \n\nThen it seems to works.\n\nSo maybe the problem comes from properties not being passed correctly to workers when executors are instanciated ?\n\nAlso, I'm using packaged binary distribution for CDH4 on a standalone cluster","from":"developer"},{"body":"Hi,\n\nI just realized my driver wasn't picking up spark.akka.frameSize value, because of a problem in the way I was passing it in. However, my executor's were picking this value correctly from their conf/spark-env.sh files.\n\nNow, both sides, the driver and executors print the correct value for frameSize with spark.akka.logAkkaConfig=true.\n\nI also noticed that simply starting the driver with java -Dspark.akka.frameSize=200 does not propagate this automatically to the executors. Not that this is an issue. I guess I was just confused about the configuration.\n\nGuillaume, seems like you have the reverse situation as mine, ie. your drivers are correctly configured with the right frameSize, but the executors are still using the 10MB default?\n\nTo conclude, after ensuring that the driver is correctly configured with the right frameSize, so far, serialized tasks larger than 10MB are being received by the executors and run successfully.\n\nRoshan","from":"developer"},{"body":"You're right Roshan.I was expecting the akka properties set before SparkContext creation to be propagated to the Executors (and based on what the Spark code does, it should be the case). \n\nI think it should be enforced for the whole akka stuff (timeouts and so on), as well as for the rest of spark properties","from":"developer"},{"body":"Hey There,\n\nI spent some time playing with this and couldn't reproduce the issue. The driver should capture the options and pass them to executors. I just tested this with a local cluster and was able to verify the akka frame size is passed to executors even if it's not set in spark-env.sh where the executor launches.\n\n[~roshan] - what happens if you remove the setting from spark-env.sh on the executors and only set it at the driver. Does that work correctly?","from":"developer"},{"body":"In my code I already had some options passed in SPARK_JAVA_OPTS in the env of the SparkContext, but nothing about the akka frameSize. Could it be related ?\n\nSince it was working correctly in 0.8.1, maybe it's related to the new SparkConf ?","from":"developer"},{"body":"[~guillaumepitel] If you aren't setting akka.frameSize in SPARK_JAVA_OPTS then where are you setting it?\n\nWhat I was saying is that if you do System.setProperty(spark.akka.frameSize, XX) before you create the SparkContext it should collect this and send it to the executors correctly. One thing is if you set it after you create the SparkContext it won't work... are you doing this by any chance?","from":"developer"},{"body":"No, I create the SparkContext after setting the properties (and I've nothing on my nodes for configuring the frameSize, it's a per-process configuration). \n\nSo before my workaround, I was just setting the property (and the environment in the UI was showing the right value) and passing a SPARK_JAVA_OPTS to the env of the SparkContext with \n{code}\nSystem.setProperty(\"spark.serializer\", \"org.apache.spark.serializer.KryoSerializer\")\nSystem.setProperty(\"spark.kryo.registrator\", registrator)\nSystem.setProperty(\"spark.kryo.referenceTracking\", \"false\")\nSystem.setProperty(\"spark.kryoserializer.buffer.mb\", bufferSize.toString)\nSystem.setProperty(\"spark.locality.wait\", \"10000\")\nSystem.setProperty(\"spark.hadoop.mapreduce.output.fileoutputformat.compress\", \"true\")\nSystem.setProperty(\"spark.hadoop.mapreduce.output.fileoutputformat.compress.codec\", codec)\nSystem.setProperty(\"spark.hadoop.mapreduce.output.fileoutputformat.compress.type\", \"BLOCK\")\nSystem.setProperty(\"spark.akka.frameSize\", akkaFrameSize.toString)\n\nval sb = new StringBuilder()\nsb.append(\"-Dspark.storage.memoryFraction=\" + sparkMemoryFraction())\nsb.append(\" -Dspark.worker.timeout=\" + sparkWorkerTimeout())\nsb.append(\" -Dspark.akka.askTimeout=\" + sparkAkkaAskTimeout())\nsb.append(\" -Dspark.akka.timeout=\" + sparkAkkaTimeout())\nsb.append(\" -Dspark.shuffle.consolidateFiles=\" + sparkShuffleConsolidateFiles())\n\nvar env = new HashMap[String, String]()\nenv += \"SPARK_JAVA_OPTS\" -> sb.toString()\n\nval sc = new SparkContext(sparkMaster(), appName, sparkHome(), jars(), env)\n{code}\n\nThat caused a problem (the akka frame size seemed to be passed to the executor, but only after the creation of the actorSystem, because it was taking the right code path, but the akka system didn't seem to be properly configured).\n\nNow if I add this to my code :\n\n{code}\nsb.append(\" -Dspark.akka.frameSize=\" + akkaFrameSize)\n{code}\n\nIt works\n","from":"developer"},{"body":"@Patrick Wendell -\n\nCase 1.\nDRIVER:\nI start the driver with java -Dspark.akka.frameSize=200 -D... -cp .....\n\nEXECUTOR:\nspark-env.sh on workers with frameSize specified: \nexport SPARK_JAVA_OPTS='-Dspark.local.dir=/data/spark/tmp -Dspark.akka.logAkkaConfig=true -Dspark.akka.frameSize=160 -Dspark.akka.timeout=100 -Dspark.akka.askTimeout=30 -Dspark.akka.logLifecycleEvents=true -Dspark.worker.timeout=200'\n\nExecutor akka config log reports frameSize=160\n\nCase 2. \nDRIVER:\nI start the driver with java -Dspark.akka.frameSize=200 -D... -cp ....\n\nEXECUTOR:\nspark-env.sh on workers with frameSize NOT specified: \nexport SPARK_JAVA_OPTS='-Dspark.local.dir=/data/spark/tmp -Dspark.akka.logAkkaConfig=true -Dspark.akka.timeout=100 -Dspark.akka.askTimeout=30 -Dspark.akka.logLifecycleEvents=true -Dspark.worker.timeout=200'\n\nExecutor akka config log on workers has no frameSize and executor process(by ps waux) has no frameSize. Here, I suppose it picks up the default 10MB.\n\nCase 3. \nDRIVER:\nThis time I export SPARK_JAVA_OPTS=\"-Dspark.akka.frameSize=220\" on the dirver host, before running the driver with java -Dspark.akka.frameSize=200 -D... -cp ....\nI'm not reading SPARK_JAVA_OPTS in the driver code.\nDriver's Akka config log, says frameSize=200.\n\nEXECUTOR1 on HOST1:\nNo frameSize specified in spark-env\nExecutor akka config log says frameSize=220.\n\nEXECUTOR2 on HOST2:\nframeSize=160 in spark-env\nExecutor akka config log says frameSize=220, \n\nps waux shows the process was started with this cmd:\n/usr/java/default/bin/java -cp sparkJar.jar -Dspark.local.dir=/data/spark/tmp -Dspark.akka.logAkkaConfig=true -Dspark.akka.frameSize=160 -Dspark.akka.timeout=100 -Dspark.akka.askTimeout=30 -Dspark.akka.logLifecycleEvents=true -Dspark.worker.timeout=200 -Dspark.akka.frameSize=220 -Xms3072M -Xmx3072M org.apache.spark.executor.CoarseGrainedExecutorBackend .....\n\nSo the SPARK_JAVA_OPTS from both the executor and the driver are appended, but since java overrides the first frameSize with the second, the executor runs with frameSize=220\n\nCase 4.\nDRIVER:\nThis time I export SPARK_JAVA_OPTS=\"-Dspark.akka.frameSize=220\" on the driver host, but don't specify frameSize in the launch command.\nAgain, I don't SPARK_JAVA_OPTS in the driver code.\nNo frameSize in driver's akka config log. I expect driver picks up the default 10MB frameSize.\n\nExecutors are the same as in case3.\n\nWhat I'm seeing is that, you can specify the frameSize for a driver in the launch command as java -Dspark.akka.frameSize property, and for the executors in their respective spark-env.sh. If you specify SPARK_JAVA_OPTS on the driver side, then this will override the value for the executors, but not for the driver. I don't use spark-class.sh to launch my driver, because somewhere I read that its meant for spark internal classes and examples. I also currently build my SparkContext directly, without using SparkConf.\n\nThis not an issue for me any longer. That being said, it would have saved me loads of time, if the logs had provided some indication that sending serialized tasks to the executors had failed because they were larger than 10MB. Also, the documentation could be a bit clearer about what is set where and which property overrides or is overridden.\n\nThanks.","from":"developer"},{"body":"Hi all, \n\nI'm very new to Spark and doing some tests, I've experienced similar issue.\n(tested with Spark Shell, 0.9.1, r3.8xlarge instance on EC2 - 32 core / 244GiB MEM)\n\nI was trying to broadcast 700MB of data and Spark hangs when I run collect() method for the data. \n\nHere's the strange things :\n1) when I tried \n{code}val userInfo = sc.textFile(\"file:///spark/logs/user_sign_up2.csv\").map{line => val split = line.split(\",\"); (split(1), split)}\nval userInfoMap = userInfo.collectAsMap\n{code}\nit runs well.\n2) when I tried \n{code}val userInfo = sc.textFile(\"file:///spark/logs/user_sign_up2.csv\").map{line => val split = line.split(\",\"); (split(1), split(5))} \nval userInfoMap = userInfo.collectAsMap\n{code}\nSpark hangs.\n3) when I slightly control the data size using sample() method or cutting the data file, it runs well. \n\nOur team investigated logs from master and worker then we found worker finished all tasks but master couldn't retrieve the result from a task the result size larger than 10MB\n\nWe tried to apply the workaround setting spark.akka.frameSize to 9, it works like a charm.\n\nI guess it might hard to reproduce the issue, please contact me if there's need of testing or getting logs. \n\nThanks!","from":"developer"},{"body":"I'm curious, why did you want to make the frameSize this big -- are the tasks themselves also big or just the results? There might be other buffers in Akka that can't be made bigger than this. It's possible that this changed in a newer Akka version (because larger frame sizes used to work before).","from":"developer"},{"body":"[~matei]\nI've found the default of spark.akka.frameSize is 10 from the config document, http://spark.apache.org/docs/0.9.1/configuration.html\njust tried to slightly larger and smaller (11 and 9) values.\n\nI did collect() method on the userInfo and it might contains large data. (edited the first comment.)\n\n","from":"developer"},{"body":"[~matei] Do you know which akka version we should use to be able to use big frame size. ","from":"developer"},{"body":"To follow up this thread, I have done some experiments when the frameSize is around 10MB .\n\n1) spark.akka.frameSize = 10\nIf one of the partition size is very close to 10MB, say 9.97MB, the execution blocks without any exception or warning. Worker finished the task to send the serialized result, and then throw exception saying hadoop IPC client connection stops (changing the logging to debug level). However, the master never receives the results and the program just hangs.\nBut if sizes for all the partitions less than some number btw 9.96MB amd 9.97MB, the program works fine.\n2) spark.akka.frameSize = 9\nwhen the partition size is just a little bit smaller than 9MB, it fails as well.\n\nThis bug behavior is not exactly what spark-1112 is about, could you please guide me how to open a separate bug when the serialization size is very close to 10MB. \n\nThanks a lot","from":"developer"},{"body":"I have filed a bug https://issues.apache.org/jira/browse/SPARK-2156","from":"developer"},{"body":"We were able to reproduce this - thanks for reporting it.","from":"developer"},{"body":"Awesome, looking forward to the fix. At least better error or exception message would be helpful. ","from":"developer"},{"body":"PR: https://github.com/apache/spark/pull/1124","from":"developer"},{"body":"This is fixed in the 1.0 branch via:\nhttps://github.com/apache/spark/pull/1172","from":"developer"},{"body":"Fixed in 1.1.0 via:\nhttps://github.com/apache/spark/pull/1132","from":"developer"},{"body":"Can a clear workaround be specified for this bug please? For those unable to upgrade to run on 1.0.1 or 1.1.0 in production, general instructions on the workaround are required. This is a huge blocker for current production deployments (even on 1.0.0) otherwise. For instance, running a saveAsTextFile() on an RDD (~400MB) causes execution to freeze with the last log statements seen on the driver being:\n\n14/06/25 16:38:55 INFO spark.SparkContext: Starting job: saveAsTextFile at Test.java:99\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Got job 6 (saveAsTextFile at Test.java:99) with 2 output partitions (allowLocal=false)\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Final stage: Stage 6(saveAsTextFile at Test.java:99)\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Parents of final stage: List()\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Missing parents: List()\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Submitting Stage 6 (MappedRDD[558] at saveAsTextFile at Test.java:99), which has no missing parents\n14/06/25 16:38:55 INFO scheduler.DAGScheduler: Submitting 2 missing tasks from Stage 6 (MappedRDD[558] at saveAsTextFile at Test.java:99)\n14/06/25 16:38:55 INFO scheduler.TaskSchedulerImpl: Adding task set 6.0 with 2 tasks\n14/06/25 16:38:55 INFO scheduler.TaskSetManager: Starting task 6.0:0 as TID 5 on executor 1: somehost.corp (PROCESS_LOCAL)\n14/06/25 16:38:55 INFO scheduler.TaskSetManager: Serialized task 6.0:0 as 351777 bytes in 36 ms\n14/06/25 16:38:55 INFO scheduler.TaskSetManager: Starting task 6.0:1 as TID 6 on executor 0: someotherhost.corp (PROCESS_LOCAL)\n14/06/25 16:38:55 INFO scheduler.TaskSetManager: Serialized task 6.0:1 as 186453 bytes in 16 ms\n\nThe test setup for reproducing this issue has two slaves (each with 24G) running spark standalone. The driver runs with Xmx 4G.\n\nThanks.","from":"developer"},{"body":"[~reachbach] If you are running on standalone mode, it might work if you go on every node in your cluster and add the following to spark-env.sh:\n\n{code}\nexport SPARK_JAVA_OPTS=\"-Dspark.akka.frameSize=XXX\"\n{code}\n\nHowever, this work around will only work if every job in your cluster is using the same frame size (XXX).\n\nThe main recommendation is to upgrade to 1.0.1. We are very conservative about what we merge into maintenance branches, so we recommend users upgrade immediately once we release them.","from":"developer"},{"body":"This is not resolved yet because it needs to be back ported into 0.9","from":"developer"},{"body":"PR for branch-0.9: https://github.com/apache/spark/pull/1455","from":"developer"},{"body":"Issue resolved by pull request 1455\n[https://github.com/apache/spark/pull/1455]","from":"developer"},{"body":"Does anyone test in version0.9.2,I found it also failed , while v1.0.1 & v1.1.0 is ok. ","from":"developer"},{"body":"Looks like the \"Fix Versions\" accidentally got overwritten during a backport / cherry-pick, so I've restored them based on the issue history.","from":"developer"}],"created":"2014-02-20T08:34:26.000+0000","description":"When I set the spark.akka.frameSize to something over 10, the messages sent from the executors to the driver completely block the execution if the message is bigger than 10MiB and smaller than the frameSize (if it's above the frameSize, it's ok)\n\nWorkaround is to set the spark.akka.frameSize to 10. In this case, since 0.8.1, the blockManager deal with the data to be sent. It seems slower than akka direct message though.\n\nThe configuration seems to be correctly read (see actorSystemConfig.txt), so I don't see where the 10MiB could come from ","issue_id":"12704733","key":"SPARK-1112","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2014-07-17T04:31:18.000+0000","role":"fixed_distractor","summary":"When spark.akka.frameSize > 10, task results bigger than 10MiB block execution"} {"case_id":"12906067","cluster":"DISTRACTOR-SPARK-11191","comments":[{"body":"I will add that the exact same thing happens when you don't use {{TEMPORARY}} i.e.:\n\n{code}\nCREATE FUNCTION testUDF AS 'com.foo.class.UDF';\n{code}","created":"2015-10-19T21:01:35.825+0000"},{"body":"This should be caused by builtin FunctionRegistry in 1.5.1\nhttps://github.com/apache/spark/blob/v1.5.1/sql/hive/src/main/scala/org/apache/spark/sql/hive/HiveContext.scala#L413\n\nThe FunctionRegistry in 1.4.1\nhttps://github.com/apache/spark/blob/v1.4.1/sql/hive/src/main/scala/org/apache/spark/sql/hive/HiveContext.scala#L377","created":"2015-11-01T12:14:55.208+0000"},{"body":"The behaviour of \"ADD JAR\" has not been changed from 1.4 to 1.5, and \"CREATE FUNCTION\" is a native command that we will run it using hive client. So I think the change of built-in FunctionRegistry is not the reason of this bug, I'll look into it.","created":"2015-11-02T06:49:11.902+0000"},{"body":"Ok, it looks like `FunctionRegistry.getFunctionInfo` does not contain permanent and temporary functions information.\n\nThe SELECT clause lookups functions in the order:\n1. underlying.lookupFunction (built-in FunctionRegistry)\n2. FunctionRegistry.getFunctionInfo","created":"2015-11-02T08:40:13.524+0000"},{"body":"One of the problem here is SPARK-11595. However, after fixing SPARK-11595, {{CREATE TEMPORARY FUNCTION}} still doesn't work properly. Still investigating.","created":"2015-11-09T13:24:25.313+0000"},{"body":"This should work in master and 1.6.","created":"2015-11-09T18:08:58.353+0000"},{"body":"Checked 10 hours ago. Pulled from master. The same error.","created":"2015-11-12T07:12:46.549+0000"},{"body":"I have wrote a workaround patch for this issue, but the patch doesn't work for embedded metastore.\n\nhttps://github.com/phstudy/spark/commit/b2e618863b733f9fe5dfc69fdb7bd9ab5379df07","created":"2015-11-12T07:48:09.014+0000"},{"body":"This issue consists of two bugs. One of them is the ADD JAR issue (SPARK-11595), which has just been fixed. The other one is that, {{HiveFunctionRegistry}} should use execution Hive client to lookup temporary functions. I'm fixing this for 1.5 and 1.6.","created":"2015-11-12T13:19:10.889+0000"},{"body":"I think you also need to lookup permanent user-defined functions in metadata Hive for local/remote metastore service. This is a common use case for hiveserver.","created":"2015-11-12T15:01:27.006+0000"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9664","created":"2015-11-12T15:47:04.460+0000"},{"body":"Spark SQL hasn't supported persisted functions yet.","created":"2015-11-12T16:43:51.093+0000"},{"body":"I think this is a Spark ThriftServer issue, not spark SQL.\n\nCurrently, by default Spark ThriftServer will create an embedded metastore in execution path, and you can also setup local/remote metastore service. For this case, Spark ThriftServer should support accessing permanent user-defined functions in the metastore.\n\n\nAnd the Spark SQL document also says \"In addition to the basic SQLContext, you can also create a HiveContext, which provides a superset of the functionality provided by the basic SQLContext. Additional features include the ability to write queries using the more complete HiveQL parser, access to Hive UDFs, and the ability to read data from Hive tables\"\n\nhttp://spark.apache.org/docs/latest/sql-programming-guide.html#starting-point-sqlcontext\n","created":"2015-11-12T17:21:51.313+0000"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9671","created":"2015-11-12T18:51:03.768+0000"},{"body":"Issue resolved by pull request 9664\n[https://github.com/apache/spark/pull/9664]","created":"2015-11-12T20:18:32.614+0000"},{"body":"Sorry that I wasn't clear enough in my previous reply. So in Spark, the Thrift server delegates most of the functionalities to Spark SQL, and Hive function lookup is also one of the case. Currently, we haven't implemented Hive persisted function lookup in Spark SQL, so the Thrift server doesn't support it either.","created":"2015-11-13T04:49:54.044+0000"},{"body":"I have implemented Hive persisted function lookup in this patch: https://github.com/phstudy/spark/commit/b2e618863b733f9fe5dfc69fdb7bd9ab5379df07.\n\nBut I found it does not work for embedded metastore, embedded metastore only allows single connection. Do you have any ideas?","created":"2015-11-13T05:20:12.753+0000"},{"body":"What error message/exception stacktrace did you get when working with embedded metastore?","created":"2015-11-13T05:26:09.776+0000"},{"body":"This is my exception stacktrace.\n\n{code}\n13:34:49,728 ERROR Schema:125 - Failed initialising database.\nUnable to open a test connection to the given database. JDBC url = jdbc:derby:;databaseName=metastore_db;create=true, username = APP. Terminating connection pool (set lazyInit to true if you expect to start your database after your app). Original Exception: ------\njava.sql.SQLException: Failed to start database 'metastore_db' with class loader org.apache.spark.sql.hive.client.IsolatedClientLoader$$anon$1@4c8d45cf, see the next exception for details.\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory40.getSQLException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.Util.newEmbedSQLException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.Util.seeNextException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.EmbedConnection.bootDatabase(Unknown Source)\n\tat org.apache.derby.impl.jdbc.EmbedConnection.(Unknown Source)\n\tat org.apache.derby.impl.jdbc.EmbedConnection40.(Unknown Source)\n\tat org.apache.derby.jdbc.Driver40.getNewEmbedConnection(Unknown Source)\n\tat org.apache.derby.jdbc.InternalDriver.connect(Unknown Source)\n\tat org.apache.derby.jdbc.Driver20.connect(Unknown Source)\n\tat org.apache.derby.jdbc.AutoloadedDriver.connect(Unknown Source)\n\tat java.sql.DriverManager.getConnection(DriverManager.java:664)\n\tat java.sql.DriverManager.getConnection(DriverManager.java:208)\n\tat com.jolbox.bonecp.BoneCP.obtainRawInternalConnection(BoneCP.java:361)\n\tat com.jolbox.bonecp.BoneCP.(BoneCP.java:416)\n\tat com.jolbox.bonecp.BoneCPDataSource.getConnection(BoneCPDataSource.java:120)\n\tat org.datanucleus.store.rdbms.ConnectionFactoryImpl$ManagedConnectionImpl.getConnection(ConnectionFactoryImpl.java:501)\n\tat org.datanucleus.store.rdbms.RDBMSStoreManager.(RDBMSStoreManager.java:298)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:422)\n\tat org.datanucleus.plugin.NonManagedPluginRegistry.createExecutableExtension(NonManagedPluginRegistry.java:631)\n\tat org.datanucleus.plugin.PluginManager.createExecutableExtension(PluginManager.java:301)\n\tat org.datanucleus.NucleusContext.createStoreManagerForProperties(NucleusContext.java:1187)\n\tat org.datanucleus.NucleusContext.initialise(NucleusContext.java:356)\n\tat org.datanucleus.api.jdo.JDOPersistenceManagerFactory.freezeConfiguration(JDOPersistenceManagerFactory.java:775)\n\tat org.datanucleus.api.jdo.JDOPersistenceManagerFactory.createPersistenceManagerFactory(JDOPersistenceManagerFactory.java:333)\n\tat org.datanucleus.api.jdo.JDOPersistenceManagerFactory.getPersistenceManagerFactory(JDOPersistenceManagerFactory.java:202)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:497)\n\tat javax.jdo.JDOHelper$16.run(JDOHelper.java:1965)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.jdo.JDOHelper.invoke(JDOHelper.java:1960)\n\tat javax.jdo.JDOHelper.invokeGetPersistenceManagerFactoryOnImplementation(JDOHelper.java:1166)\n\tat javax.jdo.JDOHelper.getPersistenceManagerFactory(JDOHelper.java:808)\n\tat javax.jdo.JDOHelper.getPersistenceManagerFactory(JDOHelper.java:701)\n\tat org.apache.hadoop.hive.metastore.ObjectStore.getPMF(ObjectStore.java:365)\n\tat org.apache.hadoop.hive.metastore.ObjectStore.getPersistenceManager(ObjectStore.java:394)\n\tat org.apache.hadoop.hive.metastore.ObjectStore.initialize(ObjectStore.java:291)\n\tat org.apache.hadoop.hive.metastore.ObjectStore.setConf(ObjectStore.java:258)\n\tat org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:73)\n\tat org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:133)\n\tat org.apache.hadoop.hive.metastore.RawStoreProxy.(RawStoreProxy.java:57)\n\tat org.apache.hadoop.hive.metastore.RawStoreProxy.getProxy(RawStoreProxy.java:66)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStore$HMSHandler.newRawStore(HiveMetaStore.java:593)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStore$HMSHandler.getMS(HiveMetaStore.java:571)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStore$HMSHandler.createDefaultDB(HiveMetaStore.java:620)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStore$HMSHandler.init(HiveMetaStore.java:461)\n\tat org.apache.hadoop.hive.metastore.RetryingHMSHandler.(RetryingHMSHandler.java:66)\n\tat org.apache.hadoop.hive.metastore.RetryingHMSHandler.getProxy(RetryingHMSHandler.java:72)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStore.newRetryingHMSHandler(HiveMetaStore.java:5762)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStoreClient.(HiveMetaStoreClient.java:199)\n\tat org.apache.hadoop.hive.ql.metadata.SessionHiveMetaStoreClient.(SessionHiveMetaStoreClient.java:74)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:422)\n\tat org.apache.hadoop.hive.metastore.MetaStoreUtils.newInstance(MetaStoreUtils.java:1521)\n\tat org.apache.hadoop.hive.metastore.RetryingMetaStoreClient.(RetryingMetaStoreClient.java:86)\n\tat org.apache.hadoop.hive.metastore.RetryingMetaStoreClient.getProxy(RetryingMetaStoreClient.java:132)\n\tat org.apache.hadoop.hive.metastore.RetryingMetaStoreClient.getProxy(RetryingMetaStoreClient.java:104)\n\tat org.apache.hadoop.hive.ql.metadata.Hive.createMetaStoreClient(Hive.java:3005)\n\tat org.apache.hadoop.hive.ql.metadata.Hive.getMSC(Hive.java:3024)\n\tat org.apache.hadoop.hive.ql.metadata.Hive.getAllDatabases(Hive.java:1234)\n\tat org.apache.hadoop.hive.ql.metadata.Hive.reloadFunctions(Hive.java:174)\n\tat org.apache.hadoop.hive.ql.metadata.Hive.(Hive.java:166)\n\tat org.apache.hadoop.hive.ql.session.SessionState.start(SessionState.java:503)\n\tat org.apache.spark.sql.hive.client.ClientWrapper.(ClientWrapper.scala:173)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:422)\n\tat org.apache.spark.sql.hive.client.IsolatedClientLoader.liftedTree1$1(IsolatedClientLoader.scala:183)\n\tat org.apache.spark.sql.hive.client.IsolatedClientLoader.(IsolatedClientLoader.scala:179)\n\tat org.apache.spark.sql.hive.HiveContext.metadataHive$lzycompute(HiveContext.scala:226)\n\tat org.apache.spark.sql.hive.HiveContext.metadataHive(HiveContext.scala:185)\n\tat org.apache.spark.sql.hive.HiveContext.functionRegistry$lzycompute(HiveContext.scala:413)\n\tat org.apache.spark.sql.hive.HiveContext.functionRegistry(HiveContext.scala:412)\n\tat org.apache.spark.sql.UDFRegistration.(UDFRegistration.scala:40)\n\tat org.apache.spark.sql.SQLContext.(SQLContext.scala:296)\n\tat org.apache.spark.sql.hive.HiveContext.(HiveContext.scala:72)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLEnv$.init(SparkSQLEnv.scala:58)\n\tat org.apache.spark.sql.hive.thriftserver.HiveThriftServer2$.main(HiveThriftServer2.scala:77)\n\tat org.apache.spark.sql.hive.thriftserver.HiveThriftServer2.main(HiveThriftServer2.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:497)\n\tat org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:672)\n\tat org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:180)\n\tat org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:205)\n\tat org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:120)\n\tat org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\nCaused by: java.sql.SQLException: Failed to start database 'metastore_db' with class loader org.apache.spark.sql.hive.client.IsolatedClientLoader$$anon$1@4c8d45cf, see the next exception for details.\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory.getSQLException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory40.wrapArgsForTransportAcrossDRDA(Unknown Source)\n\t... 95 more\nCaused by: java.sql.SQLException: Another instance of Derby may have already booted the database /usr/local/Cellar/apache-spark/1.5.1/libexec/sbin/metastore_db.\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory.getSQLException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory40.wrapArgsForTransportAcrossDRDA(Unknown Source)\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory40.getSQLException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.Util.generateCsSQLException(Unknown Source)\n\t... 92 more\nCaused by: ERROR XSDB6: Another instance of Derby may have already booted the database /usr/local/Cellar/apache-spark/1.5.1/libexec/sbin/metastore_db.\n\tat org.apache.derby.iapi.error.StandardException.newException(Unknown Source)\n\tat org.apache.derby.impl.store.raw.data.BaseDataFileFactory.privGetJBMSLockOnDB(Unknown Source)\n\tat org.apache.derby.impl.store.raw.data.BaseDataFileFactory.run(Unknown Source)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat org.apache.derby.impl.store.raw.data.BaseDataFileFactory.getJBMSLockOnDB(Unknown Source)\n\tat org.apache.derby.impl.store.raw.data.BaseDataFileFactory.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.TopService.bootModule(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.startModule(Unknown Source)\n\tat org.apache.derby.iapi.services.monitor.Monitor.bootServiceModule(Unknown Source)\n\tat org.apache.derby.impl.store.raw.RawStore.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.TopService.bootModule(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.startModule(Unknown Source)\n\tat org.apache.derby.iapi.services.monitor.Monitor.bootServiceModule(Unknown Source)\n\tat org.apache.derby.impl.store.access.RAMAccessManager.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.TopService.bootModule(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.startModule(Unknown Source)\n\tat org.apache.derby.iapi.services.monitor.Monitor.bootServiceModule(Unknown Source)\n\tat org.apache.derby.impl.db.BasicDatabase.bootStore(Unknown Source)\n\tat org.apache.derby.impl.db.BasicDatabase.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.TopService.bootModule(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.bootService(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.startProviderService(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.findProviderAndStartService(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.startPersistentService(Unknown Source)\n\tat org.apache.derby.iapi.services.monitor.Monitor.startPersistentService(Unknown Source)\n\t... 92 more\n------\n...\n{code}","created":"2015-11-13T05:38:04.736+0000"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9737","created":"2015-11-16T13:47:05.358+0000"}],"conversations":[{"body":"Since upgrading to spark 1.5 we've been unable to create and use UDF's when we run in thrift server mode.\n\nOur setup:\nWe start the thrift-server running against yarn in client mode, (we've also built our own spark from github branch-1.5 with the following args: {{-Pyarn -Phive -Phive-thrifeserver}}\n\nIf i run the following after connecting via JDBC (in this case via beeline):\n\n{{add jar 'hdfs://path/to/jar\"}}\n(this command succeeds with no errors)\n\n{{CREATE TEMPORARY FUNCTION testUDF AS 'com.foo.class.UDF';}}\n(this command succeeds with no errors)\n\n{{select testUDF(col1) from table1;}}\n\nI get the following error in the logs:\n\n{code}\norg.apache.spark.sql.AnalysisException: undefined function testUDF; line 1 pos 8\n at org.apache.spark.sql.hive.HiveFunctionRegistry$$anonfun$lookupFunction$2$$anonfun$1.apply(hiveUDFs.scala:58)\n at org.apache.spark.sql.hive.HiveFunctionRegistry$$anonfun$lookupFunction$2$$anonfun$1.apply(hiveUDFs.scala:58)\n at scala.Option.getOrElse(Option.scala:120)\n at org.apache.spark.sql.hive.HiveFunctionRegistry$$anonfun$lookupFunction$2.apply(hiveUDFs.scala:57)\n at org.apache.spark.sql.hive.HiveFunctionRegistry$$anonfun$lookupFunction$2.apply(hiveUDFs.scala:53)\n at scala.util.Try.getOrElse(Try.scala:77)\n at org.apache.spark.sql.hive.HiveFunctionRegistry.lookupFunction(hiveUDFs.scala:53)\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveFunctions$$anonfun$apply$10$$anonfun$applyOrElse$5$$anonfun$applyOrElse$24.apply(Analyzer.scala:506)\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveFunctions$$anonfun$apply$10$$anonfun$applyOrElse$5$$anonfun$applyOrElse$24.apply(Analyzer.scala:506)\n at org.apache.spark.sql.catalyst.analysis.package$.withPosition(package.scala:48)\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveFunctions$$anonfun$apply$10$$anonfun$applyOrElse$5.applyOrElse(Analyzer.scala:505)\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveFunctions$$anonfun$apply$10$$anonfun$applyOrElse$5.applyOrElse(Analyzer.scala:502)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:227)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:227)\n at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:51)\n at org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:226)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:232)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:232)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$4.apply(TreeNode.scala:249)\n{code}\n\n\n(cutting the bulk for ease of report, more than happy to send the full output)\n\n{code}\n15/10/12 14:34:37 ERROR SparkExecuteStatementOperation: Error running hive query:\norg.apache.hive.service.cli.HiveSQLException: org.apache.spark.sql.AnalysisException: undefined function testUDF; line 1 pos 100\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.runInternal(SparkExecuteStatementOperation.scala:259)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation$$anon$1$$anon$2.run(SparkExecuteStatementOperation.scala:171)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1628)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation$$anon$1.run(SparkExecuteStatementOperation.scala:182)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\n\nWhen I ran the same against 1.4 it worked.\n\nI've also changed the {{spark.sql.hive.metastore.version}} version to be 0.13 (similar to what it was in 1.4) and 0.14 but I still get the same errors.\n\nAlso, in 1.5, when you run it against the {{spark-sql}} shell, it works.","from":"reporter","subject":"[1.5] Can't create UDF's using hive thrift service"},{"body":"I will add that the exact same thing happens when you don't use {{TEMPORARY}} i.e.:\n\n{code}\nCREATE FUNCTION testUDF AS 'com.foo.class.UDF';\n{code}","from":"developer"},{"body":"This should be caused by builtin FunctionRegistry in 1.5.1\nhttps://github.com/apache/spark/blob/v1.5.1/sql/hive/src/main/scala/org/apache/spark/sql/hive/HiveContext.scala#L413\n\nThe FunctionRegistry in 1.4.1\nhttps://github.com/apache/spark/blob/v1.4.1/sql/hive/src/main/scala/org/apache/spark/sql/hive/HiveContext.scala#L377","from":"developer"},{"body":"The behaviour of \"ADD JAR\" has not been changed from 1.4 to 1.5, and \"CREATE FUNCTION\" is a native command that we will run it using hive client. So I think the change of built-in FunctionRegistry is not the reason of this bug, I'll look into it.","from":"developer"},{"body":"Ok, it looks like `FunctionRegistry.getFunctionInfo` does not contain permanent and temporary functions information.\n\nThe SELECT clause lookups functions in the order:\n1. underlying.lookupFunction (built-in FunctionRegistry)\n2. FunctionRegistry.getFunctionInfo","from":"developer"},{"body":"One of the problem here is SPARK-11595. However, after fixing SPARK-11595, {{CREATE TEMPORARY FUNCTION}} still doesn't work properly. Still investigating.","from":"developer"},{"body":"This should work in master and 1.6.","from":"developer"},{"body":"Checked 10 hours ago. Pulled from master. The same error.","from":"developer"},{"body":"I have wrote a workaround patch for this issue, but the patch doesn't work for embedded metastore.\n\nhttps://github.com/phstudy/spark/commit/b2e618863b733f9fe5dfc69fdb7bd9ab5379df07","from":"developer"},{"body":"This issue consists of two bugs. One of them is the ADD JAR issue (SPARK-11595), which has just been fixed. The other one is that, {{HiveFunctionRegistry}} should use execution Hive client to lookup temporary functions. I'm fixing this for 1.5 and 1.6.","from":"developer"},{"body":"I think you also need to lookup permanent user-defined functions in metadata Hive for local/remote metastore service. This is a common use case for hiveserver.","from":"developer"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9664","from":"developer"},{"body":"Spark SQL hasn't supported persisted functions yet.","from":"developer"},{"body":"I think this is a Spark ThriftServer issue, not spark SQL.\n\nCurrently, by default Spark ThriftServer will create an embedded metastore in execution path, and you can also setup local/remote metastore service. For this case, Spark ThriftServer should support accessing permanent user-defined functions in the metastore.\n\n\nAnd the Spark SQL document also says \"In addition to the basic SQLContext, you can also create a HiveContext, which provides a superset of the functionality provided by the basic SQLContext. Additional features include the ability to write queries using the more complete HiveQL parser, access to Hive UDFs, and the ability to read data from Hive tables\"\n\nhttp://spark.apache.org/docs/latest/sql-programming-guide.html#starting-point-sqlcontext\n","from":"developer"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9671","from":"developer"},{"body":"Issue resolved by pull request 9664\n[https://github.com/apache/spark/pull/9664]","from":"developer"},{"body":"Sorry that I wasn't clear enough in my previous reply. So in Spark, the Thrift server delegates most of the functionalities to Spark SQL, and Hive function lookup is also one of the case. Currently, we haven't implemented Hive persisted function lookup in Spark SQL, so the Thrift server doesn't support it either.","from":"developer"},{"body":"I have implemented Hive persisted function lookup in this patch: https://github.com/phstudy/spark/commit/b2e618863b733f9fe5dfc69fdb7bd9ab5379df07.\n\nBut I found it does not work for embedded metastore, embedded metastore only allows single connection. Do you have any ideas?","from":"developer"},{"body":"What error message/exception stacktrace did you get when working with embedded metastore?","from":"developer"},{"body":"This is my exception stacktrace.\n\n{code}\n13:34:49,728 ERROR Schema:125 - Failed initialising database.\nUnable to open a test connection to the given database. JDBC url = jdbc:derby:;databaseName=metastore_db;create=true, username = APP. Terminating connection pool (set lazyInit to true if you expect to start your database after your app). Original Exception: ------\njava.sql.SQLException: Failed to start database 'metastore_db' with class loader org.apache.spark.sql.hive.client.IsolatedClientLoader$$anon$1@4c8d45cf, see the next exception for details.\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory40.getSQLException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.Util.newEmbedSQLException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.Util.seeNextException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.EmbedConnection.bootDatabase(Unknown Source)\n\tat org.apache.derby.impl.jdbc.EmbedConnection.(Unknown Source)\n\tat org.apache.derby.impl.jdbc.EmbedConnection40.(Unknown Source)\n\tat org.apache.derby.jdbc.Driver40.getNewEmbedConnection(Unknown Source)\n\tat org.apache.derby.jdbc.InternalDriver.connect(Unknown Source)\n\tat org.apache.derby.jdbc.Driver20.connect(Unknown Source)\n\tat org.apache.derby.jdbc.AutoloadedDriver.connect(Unknown Source)\n\tat java.sql.DriverManager.getConnection(DriverManager.java:664)\n\tat java.sql.DriverManager.getConnection(DriverManager.java:208)\n\tat com.jolbox.bonecp.BoneCP.obtainRawInternalConnection(BoneCP.java:361)\n\tat com.jolbox.bonecp.BoneCP.(BoneCP.java:416)\n\tat com.jolbox.bonecp.BoneCPDataSource.getConnection(BoneCPDataSource.java:120)\n\tat org.datanucleus.store.rdbms.ConnectionFactoryImpl$ManagedConnectionImpl.getConnection(ConnectionFactoryImpl.java:501)\n\tat org.datanucleus.store.rdbms.RDBMSStoreManager.(RDBMSStoreManager.java:298)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:422)\n\tat org.datanucleus.plugin.NonManagedPluginRegistry.createExecutableExtension(NonManagedPluginRegistry.java:631)\n\tat org.datanucleus.plugin.PluginManager.createExecutableExtension(PluginManager.java:301)\n\tat org.datanucleus.NucleusContext.createStoreManagerForProperties(NucleusContext.java:1187)\n\tat org.datanucleus.NucleusContext.initialise(NucleusContext.java:356)\n\tat org.datanucleus.api.jdo.JDOPersistenceManagerFactory.freezeConfiguration(JDOPersistenceManagerFactory.java:775)\n\tat org.datanucleus.api.jdo.JDOPersistenceManagerFactory.createPersistenceManagerFactory(JDOPersistenceManagerFactory.java:333)\n\tat org.datanucleus.api.jdo.JDOPersistenceManagerFactory.getPersistenceManagerFactory(JDOPersistenceManagerFactory.java:202)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:497)\n\tat javax.jdo.JDOHelper$16.run(JDOHelper.java:1965)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.jdo.JDOHelper.invoke(JDOHelper.java:1960)\n\tat javax.jdo.JDOHelper.invokeGetPersistenceManagerFactoryOnImplementation(JDOHelper.java:1166)\n\tat javax.jdo.JDOHelper.getPersistenceManagerFactory(JDOHelper.java:808)\n\tat javax.jdo.JDOHelper.getPersistenceManagerFactory(JDOHelper.java:701)\n\tat org.apache.hadoop.hive.metastore.ObjectStore.getPMF(ObjectStore.java:365)\n\tat org.apache.hadoop.hive.metastore.ObjectStore.getPersistenceManager(ObjectStore.java:394)\n\tat org.apache.hadoop.hive.metastore.ObjectStore.initialize(ObjectStore.java:291)\n\tat org.apache.hadoop.hive.metastore.ObjectStore.setConf(ObjectStore.java:258)\n\tat org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:73)\n\tat org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:133)\n\tat org.apache.hadoop.hive.metastore.RawStoreProxy.(RawStoreProxy.java:57)\n\tat org.apache.hadoop.hive.metastore.RawStoreProxy.getProxy(RawStoreProxy.java:66)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStore$HMSHandler.newRawStore(HiveMetaStore.java:593)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStore$HMSHandler.getMS(HiveMetaStore.java:571)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStore$HMSHandler.createDefaultDB(HiveMetaStore.java:620)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStore$HMSHandler.init(HiveMetaStore.java:461)\n\tat org.apache.hadoop.hive.metastore.RetryingHMSHandler.(RetryingHMSHandler.java:66)\n\tat org.apache.hadoop.hive.metastore.RetryingHMSHandler.getProxy(RetryingHMSHandler.java:72)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStore.newRetryingHMSHandler(HiveMetaStore.java:5762)\n\tat org.apache.hadoop.hive.metastore.HiveMetaStoreClient.(HiveMetaStoreClient.java:199)\n\tat org.apache.hadoop.hive.ql.metadata.SessionHiveMetaStoreClient.(SessionHiveMetaStoreClient.java:74)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:422)\n\tat org.apache.hadoop.hive.metastore.MetaStoreUtils.newInstance(MetaStoreUtils.java:1521)\n\tat org.apache.hadoop.hive.metastore.RetryingMetaStoreClient.(RetryingMetaStoreClient.java:86)\n\tat org.apache.hadoop.hive.metastore.RetryingMetaStoreClient.getProxy(RetryingMetaStoreClient.java:132)\n\tat org.apache.hadoop.hive.metastore.RetryingMetaStoreClient.getProxy(RetryingMetaStoreClient.java:104)\n\tat org.apache.hadoop.hive.ql.metadata.Hive.createMetaStoreClient(Hive.java:3005)\n\tat org.apache.hadoop.hive.ql.metadata.Hive.getMSC(Hive.java:3024)\n\tat org.apache.hadoop.hive.ql.metadata.Hive.getAllDatabases(Hive.java:1234)\n\tat org.apache.hadoop.hive.ql.metadata.Hive.reloadFunctions(Hive.java:174)\n\tat org.apache.hadoop.hive.ql.metadata.Hive.(Hive.java:166)\n\tat org.apache.hadoop.hive.ql.session.SessionState.start(SessionState.java:503)\n\tat org.apache.spark.sql.hive.client.ClientWrapper.(ClientWrapper.scala:173)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:422)\n\tat org.apache.spark.sql.hive.client.IsolatedClientLoader.liftedTree1$1(IsolatedClientLoader.scala:183)\n\tat org.apache.spark.sql.hive.client.IsolatedClientLoader.(IsolatedClientLoader.scala:179)\n\tat org.apache.spark.sql.hive.HiveContext.metadataHive$lzycompute(HiveContext.scala:226)\n\tat org.apache.spark.sql.hive.HiveContext.metadataHive(HiveContext.scala:185)\n\tat org.apache.spark.sql.hive.HiveContext.functionRegistry$lzycompute(HiveContext.scala:413)\n\tat org.apache.spark.sql.hive.HiveContext.functionRegistry(HiveContext.scala:412)\n\tat org.apache.spark.sql.UDFRegistration.(UDFRegistration.scala:40)\n\tat org.apache.spark.sql.SQLContext.(SQLContext.scala:296)\n\tat org.apache.spark.sql.hive.HiveContext.(HiveContext.scala:72)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLEnv$.init(SparkSQLEnv.scala:58)\n\tat org.apache.spark.sql.hive.thriftserver.HiveThriftServer2$.main(HiveThriftServer2.scala:77)\n\tat org.apache.spark.sql.hive.thriftserver.HiveThriftServer2.main(HiveThriftServer2.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:497)\n\tat org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:672)\n\tat org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:180)\n\tat org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:205)\n\tat org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:120)\n\tat org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\nCaused by: java.sql.SQLException: Failed to start database 'metastore_db' with class loader org.apache.spark.sql.hive.client.IsolatedClientLoader$$anon$1@4c8d45cf, see the next exception for details.\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory.getSQLException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory40.wrapArgsForTransportAcrossDRDA(Unknown Source)\n\t... 95 more\nCaused by: java.sql.SQLException: Another instance of Derby may have already booted the database /usr/local/Cellar/apache-spark/1.5.1/libexec/sbin/metastore_db.\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory.getSQLException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory40.wrapArgsForTransportAcrossDRDA(Unknown Source)\n\tat org.apache.derby.impl.jdbc.SQLExceptionFactory40.getSQLException(Unknown Source)\n\tat org.apache.derby.impl.jdbc.Util.generateCsSQLException(Unknown Source)\n\t... 92 more\nCaused by: ERROR XSDB6: Another instance of Derby may have already booted the database /usr/local/Cellar/apache-spark/1.5.1/libexec/sbin/metastore_db.\n\tat org.apache.derby.iapi.error.StandardException.newException(Unknown Source)\n\tat org.apache.derby.impl.store.raw.data.BaseDataFileFactory.privGetJBMSLockOnDB(Unknown Source)\n\tat org.apache.derby.impl.store.raw.data.BaseDataFileFactory.run(Unknown Source)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat org.apache.derby.impl.store.raw.data.BaseDataFileFactory.getJBMSLockOnDB(Unknown Source)\n\tat org.apache.derby.impl.store.raw.data.BaseDataFileFactory.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.TopService.bootModule(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.startModule(Unknown Source)\n\tat org.apache.derby.iapi.services.monitor.Monitor.bootServiceModule(Unknown Source)\n\tat org.apache.derby.impl.store.raw.RawStore.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.TopService.bootModule(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.startModule(Unknown Source)\n\tat org.apache.derby.iapi.services.monitor.Monitor.bootServiceModule(Unknown Source)\n\tat org.apache.derby.impl.store.access.RAMAccessManager.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.TopService.bootModule(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.startModule(Unknown Source)\n\tat org.apache.derby.iapi.services.monitor.Monitor.bootServiceModule(Unknown Source)\n\tat org.apache.derby.impl.db.BasicDatabase.bootStore(Unknown Source)\n\tat org.apache.derby.impl.db.BasicDatabase.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.boot(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.TopService.bootModule(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.bootService(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.startProviderService(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.findProviderAndStartService(Unknown Source)\n\tat org.apache.derby.impl.services.monitor.BaseMonitor.startPersistentService(Unknown Source)\n\tat org.apache.derby.iapi.services.monitor.Monitor.startPersistentService(Unknown Source)\n\t... 92 more\n------\n...\n{code}","from":"developer"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9737","from":"developer"}],"created":"2015-10-19T20:40:47.000+0000","description":"Since upgrading to spark 1.5 we've been unable to create and use UDF's when we run in thrift server mode.\n\nOur setup:\nWe start the thrift-server running against yarn in client mode, (we've also built our own spark from github branch-1.5 with the following args: {{-Pyarn -Phive -Phive-thrifeserver}}\n\nIf i run the following after connecting via JDBC (in this case via beeline):\n\n{{add jar 'hdfs://path/to/jar\"}}\n(this command succeeds with no errors)\n\n{{CREATE TEMPORARY FUNCTION testUDF AS 'com.foo.class.UDF';}}\n(this command succeeds with no errors)\n\n{{select testUDF(col1) from table1;}}\n\nI get the following error in the logs:\n\n{code}\norg.apache.spark.sql.AnalysisException: undefined function testUDF; line 1 pos 8\n at org.apache.spark.sql.hive.HiveFunctionRegistry$$anonfun$lookupFunction$2$$anonfun$1.apply(hiveUDFs.scala:58)\n at org.apache.spark.sql.hive.HiveFunctionRegistry$$anonfun$lookupFunction$2$$anonfun$1.apply(hiveUDFs.scala:58)\n at scala.Option.getOrElse(Option.scala:120)\n at org.apache.spark.sql.hive.HiveFunctionRegistry$$anonfun$lookupFunction$2.apply(hiveUDFs.scala:57)\n at org.apache.spark.sql.hive.HiveFunctionRegistry$$anonfun$lookupFunction$2.apply(hiveUDFs.scala:53)\n at scala.util.Try.getOrElse(Try.scala:77)\n at org.apache.spark.sql.hive.HiveFunctionRegistry.lookupFunction(hiveUDFs.scala:53)\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveFunctions$$anonfun$apply$10$$anonfun$applyOrElse$5$$anonfun$applyOrElse$24.apply(Analyzer.scala:506)\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveFunctions$$anonfun$apply$10$$anonfun$applyOrElse$5$$anonfun$applyOrElse$24.apply(Analyzer.scala:506)\n at org.apache.spark.sql.catalyst.analysis.package$.withPosition(package.scala:48)\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveFunctions$$anonfun$apply$10$$anonfun$applyOrElse$5.applyOrElse(Analyzer.scala:505)\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveFunctions$$anonfun$apply$10$$anonfun$applyOrElse$5.applyOrElse(Analyzer.scala:502)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:227)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:227)\n at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:51)\n at org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:226)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:232)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:232)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$4.apply(TreeNode.scala:249)\n{code}\n\n\n(cutting the bulk for ease of report, more than happy to send the full output)\n\n{code}\n15/10/12 14:34:37 ERROR SparkExecuteStatementOperation: Error running hive query:\norg.apache.hive.service.cli.HiveSQLException: org.apache.spark.sql.AnalysisException: undefined function testUDF; line 1 pos 100\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.runInternal(SparkExecuteStatementOperation.scala:259)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation$$anon$1$$anon$2.run(SparkExecuteStatementOperation.scala:171)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1628)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation$$anon$1.run(SparkExecuteStatementOperation.scala:182)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\n\nWhen I ran the same against 1.4 it worked.\n\nI've also changed the {{spark.sql.hive.metastore.version}} version to be 0.13 (similar to what it was in 1.4) and 0.14 but I still get the same errors.\n\nAlso, in 1.5, when you run it against the {{spark-sql}} shell, it works.","issue_id":"12906067","key":"SPARK-11191","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-11-12T20:18:32.000+0000","role":"fixed_distractor","summary":"[1.5] Can't create UDF's using hive thrift service"} {"case_id":"12907596","cluster":"DISTRACTOR-SPARK-11293","comments":[{"body":"User 'JoshRosen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9260","created":"2015-10-24T01:45:03.165+0000"},{"body":"This was fixed for Spark 1.6.0 via a different patch (memory manager consolidation), but I'm going to try to pull my earlier patch into 1.5.x.","created":"2015-11-03T00:55:49.503+0000"},{"body":"User 'JoshRosen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9427","created":"2015-11-03T00:57:02.372+0000"},{"body":"The memory manager was rewritten there? Could it have introduced a memory leak in a different place or of a different kind? Is there a regression test to verify?","created":"2015-11-17T14:38:13.996+0000"},{"body":"Since the back-port for 1.5 was cancelled, I think this is Fixed.","created":"2015-12-15T20:46:32.428+0000"},{"body":"I have a somewhat contrived example that still leaks in 1.6.0. I started {{spark-shell --master 'local-cluster[2,2,1024]'}} and ran:\n\n{code}\nsc.parallelize(0 to 10000000, 2).map(x => x % 10000 -> x).groupByKey.asInstanceOf[org.apache.spark.rdd.ShuffledRDD[Int, Int, Iterable[Int]]].setKeyOrdering(implicitly[Ordering[Int]]).mapPartitions { it => it.take(1) }.collect\n{code}\n\nI've added extra logging around task memory acquisition so I would be able to see what is not released. These are the logs:\n\n{code}\n16/01/05 17:02:45 INFO Executor: Running task 0.0 in stage 13.0 (TID 24)\n16/01/05 17:02:45 INFO MapOutputTrackerWorker: Updating epoch to 7 and clearing cache\n16/01/05 17:02:45 INFO TorrentBroadcast: Started reading broadcast variable 13\n16/01/05 17:02:45 INFO MemoryStore: Block broadcast_13_piece0 stored as bytes in memory (estimated size 2.3 KB, free 7.6 KB)\n16/01/05 17:02:45 INFO TorrentBroadcast: Reading broadcast variable 13 took 6 ms\n16/01/05 17:02:45 INFO MemoryStore: Block broadcast_13 stored as values in memory (estimated size 4.5 KB, free 12.1 KB)\n16/01/05 17:02:45 INFO MapOutputTrackerWorker: Don't have map outputs for shuffle 6, fetching them\n16/01/05 17:02:45 INFO MapOutputTrackerWorker: Doing the fetch; tracker endpoint = NettyRpcEndpointRef(spark://MapOutputTracker@192.168.0.32:55147)\n16/01/05 17:02:45 INFO MapOutputTrackerWorker: Got the output locations\n16/01/05 17:02:45 INFO ShuffleBlockFetcherIterator: Getting 2 non-empty blocks out of 2 blocks\n16/01/05 17:02:45 INFO ShuffleBlockFetcherIterator: Started 1 remote fetches in 1 ms\n16/01/05 17:02:45 ERROR TaskMemoryManager: Task 24 acquire 5.0 MB for null\n16/01/05 17:02:45 ERROR TaskMemoryManager: Stack trace:\njava.lang.Exception: here\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:187)\n\tat org.apache.spark.util.collection.Spillable$class.maybeSpill(Spillable.scala:82)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.maybeSpill(ExternalAppendOnlyMap.scala:55)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.insertAll(ExternalAppendOnlyMap.scala:158)\n\tat org.apache.spark.Aggregator.combineValuesByKey(Aggregator.scala:45)\n\tat org.apache.spark.shuffle.BlockStoreShuffleReader.read(BlockStoreShuffleReader.scala:89)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:98)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/01/05 17:02:47 ERROR TaskMemoryManager: Task 24 acquire 15.0 MB for null\n16/01/05 17:02:47 ERROR TaskMemoryManager: Stack trace:\njava.lang.Exception: here\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:187)\n\tat org.apache.spark.util.collection.Spillable$class.maybeSpill(Spillable.scala:82)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.maybeSpill(ExternalAppendOnlyMap.scala:55)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.insertAll(ExternalAppendOnlyMap.scala:158)\n\tat org.apache.spark.Aggregator.combineValuesByKey(Aggregator.scala:45)\n\tat org.apache.spark.shuffle.BlockStoreShuffleReader.read(BlockStoreShuffleReader.scala:89)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:98)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/01/05 17:02:49 ERROR TaskMemoryManager: Task 24 acquire 5.0 MB for null\n16/01/05 17:02:49 ERROR TaskMemoryManager: Stack trace:\njava.lang.Exception: here\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:187)\n\tat org.apache.spark.util.collection.Spillable$class.maybeSpill(Spillable.scala:82)\n\tat org.apache.spark.util.collection.ExternalSorter.maybeSpill(ExternalSorter.scala:89)\n\tat org.apache.spark.util.collection.ExternalSorter.maybeSpillCollection(ExternalSorter.scala:220)\n\tat org.apache.spark.util.collection.ExternalSorter.insertAll(ExternalSorter.scala:201)\n\tat org.apache.spark.shuffle.BlockStoreShuffleReader.read(BlockStoreShuffleReader.scala:103)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:98)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/01/05 17:02:49 ERROR TaskMemoryManager: Task 24 acquire 10.5 MB for null\n16/01/05 17:02:49 ERROR TaskMemoryManager: Stack trace:\njava.lang.Exception: here\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:187)\n\tat org.apache.spark.util.collection.Spillable$class.maybeSpill(Spillable.scala:82)\n\tat org.apache.spark.util.collection.ExternalSorter.maybeSpill(ExternalSorter.scala:89)\n\tat org.apache.spark.util.collection.ExternalSorter.maybeSpillCollection(ExternalSorter.scala:220)\n\tat org.apache.spark.util.collection.ExternalSorter.insertAll(ExternalSorter.scala:201)\n\tat org.apache.spark.shuffle.BlockStoreShuffleReader.read(BlockStoreShuffleReader.scala:103)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:98)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/01/05 17:02:49 ERROR TaskMemoryManager: Task 24 release 20.0 MB from null\n16/01/05 17:02:49 ERROR TaskMemoryManager: Stack trace:\njava.lang.Exception: here\n\tat org.apache.spark.memory.TaskMemoryManager.releaseExecutionMemory(TaskMemoryManager.java:197)\n\tat org.apache.spark.util.collection.Spillable$class.releaseMemory(Spillable.scala:111)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.releaseMemory(ExternalAppendOnlyMap.scala:55)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.org$apache$spark$util$collection$ExternalAppendOnlyMap$$freeCurrentMap(ExternalAppendOnlyMap.scala:259)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap$$anonfun$iterator$1.apply$mcV$sp(ExternalAppendOnlyMap.scala:251)\n\tat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\n\tat org.apache.spark.util.collection.ExternalSorter.insertAll(ExternalSorter.scala:197)\n\tat org.apache.spark.shuffle.BlockStoreShuffleReader.read(BlockStoreShuffleReader.scala:103)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:98)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/01/05 17:02:49 ERROR Executor: Managed memory leak detected; size = 16259594 bytes, TID = 24\n{code}\n\nThe issue is that {{ExternalSorter.stop()}} is only called by {{CompletionIterator}} if the iterator is iterated through to the end. But here we only take the first element.\n\nIn practice this happens to us in a {{zipPartitions}} call where we do not iterate both iterators to the end. (It's a kind of join.)\n\nIs it illegal to not iterate an RDD iterator to the end? I think it's not. {{RDD.take}} stops short as well. This issue can probably be reproduced with {{RDD.take}} too. (I tried and failed.)","created":"2016-01-05T16:18:29.598+0000"},{"body":"Sorry, my example was overly complicated. This one triggers the same leak.\n\n{code}\nsc.parallelize(0 to 10000000, 2).map(x => x % 10000 -> x).groupByKey.mapPartitions { it => it.take(1) }.collect\n{code}","created":"2016-01-05T16:32:29.925+0000"},{"body":"so should be reopened or not? is there still a memory leak? is there a new memory leak instead of the old one?","created":"2016-01-13T15:51:12.554+0000"},{"body":"> so should be reopened or not? is there still a memory leak? is there a new memory leak instead of the old one?\n\nI think it should be reopened, because the remaining leak is an edge case of the original problem that was not covered by Josh's fix. I'll reopen it and we'll see what he thinks!","created":"2016-01-14T10:51:40.561+0000"},{"body":"so add 1.6.0 as affected version...","created":"2016-01-14T11:35:07.787+0000"},{"body":"> so add 1.6.0 as affected version...\n\nDone.","created":"2016-01-14T11:36:30.542+0000"},{"body":"Not iterating to the end has a bunch of issues IIRC - including what you mention above. For example, m'mapped buffers are not released, etc.\nUnfortunately, I dont think there is a general clean solution for it. Would be good to see what alternatives exist to resolve this.","created":"2016-02-02T20:12:50.585+0000"},{"body":"Whyle trying to execute Analytics Triangle Count with http://snap.stanford.edu/data/com-Friendster.html (with LiveJournal works ok) on about 24 nodes (spark standalone 1.52, around 32g memory each node, 128 partitions, parallelism cores*nodes*8) I get some errors that are maybe related:\n16/02/18 08:52:11 ERROR Executor: Managed memory leak detected; size = 67108864 bytes, TID = 507\n16/02/18 08:52:11 INFO MapOutputTrackerWorker: Updating epoch to 8 and clearing cache\n16/02/18 08:52:11 INFO TorrentBroadcast: Started reading broadcast variable 6\n16/02/18 08:52:11 ERROR Executor: Exception in task 123.0 in stage 3.0 (TID 507)\njava.lang.OutOfMemoryError: Java heap space\n at scala.reflect.ManifestFactory$$anon$10.newArray(Manifest.scala:122)\n at scala.reflect.ManifestFactory$$anon$10.newArray(Manifest.scala:120)\n at org.apache.spark.util.collection.OpenHashSet.rehash(OpenHashSet.scala:231)\n at org.apache.spark.util.collection.OpenHashSet.rehashIfNeeded(OpenHashSet.scala:166)\n at org.apache.spark.util.collection.OpenHashSet.rehashIfNeeded$mcJ$sp(OpenHashSet.scala:164)\n at org.apache.spark.graphx.util.collection.GraphXPrimitiveKeyOpenHashMap$mcJI$sp.changeValue$mcJI$sp(GraphXPrimitiveKeyOpenHashMap.scala:107)\n at org.apache.spark.graphx.impl.EdgePartitionBuilder.toEdgePartition(EdgePartitionBuilder.scala:58)\n at org.apache.spark.graphx.impl.GraphImpl$$anonfun$4.apply(GraphImpl.scala:115)\n at org.apache.spark.graphx.impl.GraphImpl$$anonfun$4.apply(GraphImpl.scala:109)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$18.apply(RDD.scala:727)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$18.apply(RDD.scala:727)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:300)\n at org.apache.spark.CacheManager.getOrCompute(CacheManager.scala:69)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:262)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:300)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:88)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)","created":"2016-02-18T09:53:01.126+0000"},{"body":"Saw a similar issue loading the Friendster graph into Tinkerpop and running the BulkVertexLoader with the Shuffled Vertex RDD persisted to disk. ","created":"2016-02-26T00:26:26.472+0000"},{"body":"[~joshrosen] any plans to fix this? I believe we are hitting this issue with {{1.6.0}}\n\n{code}\n16/03/17 21:38:38 INFO memory.TaskMemoryManager: Acquired by org.apache.spark.shuffle.sort.ShuffleExternalSorter@6ecba8f1: 32.0 KB\n16/03/17 21:38:38 INFO memory.TaskMemoryManager: 1528015093 bytes of memory were used by task 103134 but are not associated with specific consumers\n16/03/17 21:38:38 INFO memory.TaskMemoryManager: 1528047861 bytes of memory are used for execution and 80608434 bytes of memory are used for storage\n16/03/17 21:38:38 ERROR executor.Executor: Managed memory leak detected; size = 1528015093 bytes, TID = 103134\n16/03/17 21:38:38 ERROR executor.Executor: Exception in task 448.0 in stage 273.0 (TID 103134)\njava.lang.OutOfMemoryError: Unable to acquire 128 bytes of memory, got 0\n\tat org.apache.spark.memory.MemoryConsumer.allocatePage(MemoryConsumer.java:120)\n\tat org.apache.spark.shuffle.sort.ShuffleExternalSorter.acquireNewPageIfNecessary(ShuffleExternalSorter.java:354)\n\tat org.apache.spark.shuffle.sort.ShuffleExternalSorter.insertRecord(ShuffleExternalSorter.java:375)\n\tat org.apache.spark.shuffle.sort.UnsafeShuffleWriter.insertRecordIntoSorter(UnsafeShuffleWriter.java:237)\n\tat org.apache.spark.shuffle.sort.UnsafeShuffleWriter.write(UnsafeShuffleWriter.java:164)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n16/03/17 21:38:38 ERROR util.SparkUncaughtExceptionHandler: Uncaught exception in thread Thread[Executor task launch worker-0,5,main]\njava.lang.OutOfMemoryError: Unable to acquire 128 bytes of memory, got 0\n\tat org.apache.spark.memory.MemoryConsumer.allocatePage(MemoryConsumer.java:120)\n\tat org.apache.spark.shuffle.sort.ShuffleExternalSorter.acquireNewPageIfNecessary(ShuffleExternalSorter.java:354)\n\tat org.apache.spark.shuffle.sort.ShuffleExternalSorter.insertRecord(ShuffleExternalSorter.java:375)\n\tat org.apache.spark.shuffle.sort.UnsafeShuffleWriter.insertRecordIntoSorter(UnsafeShuffleWriter.java:237)\n\tat org.apache.spark.shuffle.sort.UnsafeShuffleWriter.write(UnsafeShuffleWriter.java:164)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n16/03/17 21:38:38 INFO storage.DiskBlockManager: Shutdown hook called\n16/03/17 21:38:38 INFO util.ShutdownHookManager: Shutdown hook called\n{code}","created":"2016-03-17T22:43:37.723+0000"},{"body":"It looks we are getting same errors on 1.6.0 too - any plans to address it?","created":"2016-03-21T11:36:24.385+0000"},{"body":"seem to hit the issue with Spark 1.6.1 not sure if this is relative to this..if yes, it can be fixed in Spark 1.6.2 ?\n{code}\n16/03/22 23:10:26 INFO memory.TaskMemoryManager: Allocate page number 16 (67108864 bytes)\n16/03/22 23:10:26 INFO sort.UnsafeExternalSorter: Thread 221 spilling sort data of 1472.0 MB to disk (0 time so far)\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: Allocate page number 1 (1060044737 bytes)\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: Memory used in task 9302\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: Acquired by org.apache.spark.shuffle.sort.ShuffleExternalSorter@8bac554: 32.0 KB\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: Acquired by org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@7a117b4f: 512.0 MB\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: 0 bytes of memory were used by task 9302 but are not associated with specific consumers\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: 14909439433 bytes of memory are used for execution and 1376877 bytes of memory are used for storage\n16/03/22 23:11:26 WARN memory.TaskMemoryManager: leak 32.0 KB memory from org.apache.spark.shuffle.sort.ShuffleExternalSorter@8bac554\n16/03/22 23:11:26 ERROR executor.Executor: Managed memory leak detected; size = 32768 bytes, TID = 9302\n16/03/22 23:11:26 ERROR executor.Executor: Exception in task 192.0 in stage 153.0 (TID 9302)\njava.lang.OutOfMemoryError: Unable to acquire 1073741824 bytes of memory, got 1060044737\n at org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:91)\n at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.growPointerArrayIfNecessary(UnsafeExternalSorter.java:295)\n at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.insertRecord(UnsafeExternalSorter.java:330)\n at org.apache.spark.sql.execution.UnsafeExternalRowSorter.insertRow(UnsafeExternalRowSorter.java:91)\n at org.apache.spark.sql.execution.UnsafeExternalRowSorter.sort(UnsafeExternalRowSorter.java:168)\n at org.apache.spark.sql.execution.Sort$$anonfun$1.apply(Sort.scala:90)\n at org.apache.spark.sql.execution.Sort$$anonfun$1.apply(Sort.scala:64)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.rdd.ZippedPartitionsRDD2.compute(ZippedPartitionsRDD.scala:88)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{code}","created":"2016-03-22T15:55:52.002+0000"},{"body":"I don't have such verbose logging as above but can confirm that the test case given produces memory leaks in Spark 1.6.1\n{code}\nscala> sc.parallelize(0 to 10000000, 2).map(x => x % 10000 -> x).groupByKey.mapPartitions { it => it.take(1) }.collect\n16/03/23 10:26:11 INFO SparkContext: Starting job: collect at :28\n16/03/23 10:26:11 INFO DAGScheduler: Registering RDD 1 (map at :28)\n16/03/23 10:26:11 INFO DAGScheduler: Got job 0 (collect at :28) with 2 output partitions\n16/03/23 10:26:11 INFO DAGScheduler: Final stage: ResultStage 1 (collect at :28)\n16/03/23 10:26:11 INFO DAGScheduler: Parents of final stage: List(ShuffleMapStage 0)\n16/03/23 10:26:11 INFO DAGScheduler: Missing parents: List(ShuffleMapStage 0)\n16/03/23 10:26:11 INFO DAGScheduler: Submitting ShuffleMapStage 0 (MapPartitionsRDD[1] at map at :28), which has no missing parents\n16/03/23 10:26:11 INFO MemoryStore: Block broadcast_0 stored as values in memory (estimated size 3.5 KB, free 3.5 KB)\n16/03/23 10:26:11 INFO MemoryStore: Block broadcast_0_piece0 stored as bytes in memory (estimated size 1955.0 B, free 5.4 KB)\n16/03/23 10:26:11 INFO BlockManagerInfo: Added broadcast_0_piece0 in memory on localhost:51671 (size: 1955.0 B, free: 511.5 MB)\n16/03/23 10:26:11 INFO SparkContext: Created broadcast 0 from broadcast at DAGScheduler.scala:1006\n16/03/23 10:26:11 INFO DAGScheduler: Submitting 2 missing tasks from ShuffleMapStage 0 (MapPartitionsRDD[1] at map at :28)\n16/03/23 10:26:11 INFO TaskSchedulerImpl: Adding task set 0.0 with 2 tasks\n16/03/23 10:26:11 INFO TaskSetManager: Starting task 0.0 in stage 0.0 (TID 0, localhost, partition 0,PROCESS_LOCAL, 2067 bytes)\n16/03/23 10:26:11 INFO TaskSetManager: Starting task 1.0 in stage 0.0 (TID 1, localhost, partition 1,PROCESS_LOCAL, 2124 bytes)\n16/03/23 10:26:11 INFO Executor: Running task 0.0 in stage 0.0 (TID 0)\n16/03/23 10:26:11 INFO Executor: Running task 1.0 in stage 0.0 (TID 1)\n16/03/23 10:26:21 INFO Executor: Finished task 1.0 in stage 0.0 (TID 1). 1159 bytes result sent to driver\n16/03/23 10:26:21 INFO TaskSetManager: Finished task 1.0 in stage 0.0 (TID 1) in 9152 ms on localhost (1/2)\n16/03/23 10:26:21 INFO Executor: Finished task 0.0 in stage 0.0 (TID 0). 1159 bytes result sent to driver\n16/03/23 10:26:21 INFO TaskSetManager: Finished task 0.0 in stage 0.0 (TID 0) in 9292 ms on localhost (2/2)\n16/03/23 10:26:21 INFO DAGScheduler: ShuffleMapStage 0 (map at :28) finished in 9.303 s\n16/03/23 10:26:21 INFO TaskSchedulerImpl: Removed TaskSet 0.0, whose tasks have all completed, from pool \n16/03/23 10:26:21 INFO DAGScheduler: looking for newly runnable stages\n16/03/23 10:26:21 INFO DAGScheduler: running: Set()\n16/03/23 10:26:21 INFO DAGScheduler: waiting: Set(ResultStage 1)\n16/03/23 10:26:21 INFO DAGScheduler: failed: Set()\n16/03/23 10:26:21 INFO DAGScheduler: Submitting ResultStage 1 (MapPartitionsRDD[3] at mapPartitions at :28), which has no missing parents\n16/03/23 10:26:21 INFO MemoryStore: Block broadcast_1 stored as values in memory (estimated size 4.5 KB, free 9.8 KB)\n16/03/23 10:26:21 INFO MemoryStore: Block broadcast_1_piece0 stored as bytes in memory (estimated size 2.3 KB, free 12.1 KB)\n16/03/23 10:26:21 INFO BlockManagerInfo: Added broadcast_1_piece0 in memory on localhost:51671 (size: 2.3 KB, free: 511.5 MB)\n16/03/23 10:26:21 INFO SparkContext: Created broadcast 1 from broadcast at DAGScheduler.scala:1006\n16/03/23 10:26:21 INFO DAGScheduler: Submitting 2 missing tasks from ResultStage 1 (MapPartitionsRDD[3] at mapPartitions at :28)\n16/03/23 10:26:21 INFO TaskSchedulerImpl: Adding task set 1.0 with 2 tasks\n16/03/23 10:26:21 INFO TaskSetManager: Starting task 0.0 in stage 1.0 (TID 2, localhost, partition 0,NODE_LOCAL, 1894 bytes)\n16/03/23 10:26:21 INFO TaskSetManager: Starting task 1.0 in stage 1.0 (TID 3, localhost, partition 1,NODE_LOCAL, 1894 bytes)\n16/03/23 10:26:21 INFO Executor: Running task 1.0 in stage 1.0 (TID 3)\n16/03/23 10:26:21 INFO Executor: Running task 0.0 in stage 1.0 (TID 2)\n16/03/23 10:26:21 INFO ShuffleBlockFetcherIterator: Getting 2 non-empty blocks out of 2 blocks\n16/03/23 10:26:21 INFO ShuffleBlockFetcherIterator: Getting 2 non-empty blocks out of 2 blocks\n16/03/23 10:26:21 INFO ShuffleBlockFetcherIterator: Started 0 remote fetches in 4 ms\n16/03/23 10:26:21 INFO ShuffleBlockFetcherIterator: Started 0 remote fetches in 5 ms\n16/03/23 10:26:30 ERROR Executor: Managed memory leak detected; size = 36707226 bytes, TID = 3\n16/03/23 10:26:30 ERROR Executor: Managed memory leak detected; size = 20935176 bytes, TID = 2\n...{code}\n\nCan we add 1.6.1 to the affected versions?","created":"2016-03-23T10:37:02.265+0000"},{"body":"I suspect that most of these cases are caused by a CompletionIterator being limited via {{take()}}, causing the onComplete method to not be called. Given the limitations of our current iterator interface, I don't know that there's a straightforward way to deal with this.\n\nDoes anyone have an example of a case where this message appears in a situation that cannot be explained by {{take()}} / {{limit()}}?\n\nI'm thinking that maybe we should lower the severity of this message to WARN.","created":"2016-03-23T17:00:38.789+0000"},{"body":"I've seen a few people misled by the error msg, so I'd like to try to downgrade that. I've created a separate ticket [SPARK-141868 | https://issues.apache.org/jira/browse/SPARK-14168] just for changing the msg, in case there is some fix in store here.","created":"2016-03-25T22:18:00.664+0000"},{"body":"For users that are hitting *real* OOMs from similar error msgs (not just ERROR messages that are false-positives), they may be interested in SPARK-14560, which addresses one related issue from Spillables. Note that this isn't really a \"leak\", but can easily lead to OOMs in a shuffle-to-shuffle stage caused by Spillable collections.","created":"2016-04-12T14:23:03.988+0000"},{"body":"I was using Apache spark 1.6 in EMR with spark streaming in yarn and saw memory leaks in one of the containers. Here are the logs\n\n{code}\n16/04/14 13:49:10 INFO executor.CoarseGrainedExecutorBackend: Got assigned task 2942916\n16/04/14 13:49:10 INFO executor.Executor: Running task 22.0 in stage 35684.0 (TID 2942915)\n16/04/14 13:49:10 INFO executor.Executor: Running task 23.0 in stage 35684.0 (TID 2942916)\n16/04/14 13:49:10 INFO storage.ShuffleBlockFetcherIterator: Getting 94 non-empty blocks out of 94 blocks\n16/04/14 13:49:10 INFO storage.ShuffleBlockFetcherIterator: Getting 94 non-empty blocks out of 94 blocks\n16/04/14 13:49:10 INFO storage.ShuffleBlockFetcherIterator: Started 2 remote fetches in 1 ms\n16/04/14 13:49:10 INFO storage.ShuffleBlockFetcherIterator: Started 2 remote fetches in 1 ms\n16/04/14 13:49:10 INFO storage.MemoryStore: Block input-3-1460583424327 stored as values in memory (estimated size 244.7 KB, free 19.3 MB)\n16/04/14 13:49:10 INFO receiver.BlockGenerator: Pushed block input-3-1460641750200\n16/04/14 13:49:10 INFO storage.MemoryStore: 1 blocks selected for dropping\n16/04/14 13:49:10 INFO storage.BlockManager: Dropping block input-1-1460615659379 from memory\n16/04/14 13:49:10 INFO storage.MemoryStore: 1 blocks selected for dropping\n16/04/14 13:49:10 INFO storage.BlockManager: Dropping block input-1-1460615659380 from memory\n16/04/14 13:49:10 INFO memory.TaskMemoryManager: Memory used in task 2942915\n16/04/14 13:49:10 INFO memory.TaskMemoryManager: Acquired by org.apache.spark.unsafe.map.BytesToBytesMap@34158d5f: 32.3 MB\n16/04/14 13:49:10 INFO memory.TaskMemoryManager: 0 bytes of memory were used by task 2942915 but are not associated with specific consumers\n16/04/14 13:49:10 INFO memory.TaskMemoryManager: 101247172 bytes of memory are used for execution and 3603881260 bytes of memory are used for storage\n16/04/14 13:49:10 WARN memory.TaskMemoryManager: leak 32.3 MB memory from org.apache.spark.unsafe.map.BytesToBytesMap@34158d5f\n16/04/14 13:49:10 ERROR executor.Executor: Managed memory leak detected; size = 33816576 bytes, TID = 2942915\n16/04/14 13:49:10 ERROR executor.Executor: Exception in task 22.0 in stage 35684.0 (TID 2942915)\njava.lang.OutOfMemoryError: Unable to acquire 262144 bytes of memory, got 220032\n\tat org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:91)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.allocate(BytesToBytesMap.java:735)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:197)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:212)\n\tat org.apache.spark.sql.execution.UnsafeFixedWidthAggregationMap.(UnsafeFixedWidthAggregationMap.java:103)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.(TungstenAggregationIterator.scala:483)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:95)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:86)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:710)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:710)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/04/14 13:49:10 INFO executor.Executor: Finished task 23.0 in stage 35684.0 (TID 2942916). 1921 bytes result sent to driver\n16/04/14 13:49:10 INFO executor.CoarseGrainedExecutorBackend: Got assigned task 2942927\n16/04/14 13:49:10 INFO executor.Executor: Running task 34.0 in stage 35684.0 (TID 2942927)\n16/04/14 13:49:10 ERROR util.SparkUncaughtExceptionHandler: Uncaught exception in thread Thread[Executor task launch worker-2,5,main]\njava.lang.OutOfMemoryError: Unable to acquire 262144 bytes of memory, got 220032\n\tat org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:91)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.allocate(BytesToBytesMap.java:735)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:197)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:212)\n\tat org.apache.spark.sql.execution.UnsafeFixedWidthAggregationMap.(UnsafeFixedWidthAggregationMap.java:103)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.(TungstenAggregationIterator.scala:483)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:95)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:86)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:710)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:710)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}","created":"2016-04-20T11:32:12.646+0000"},{"body":"Also biting us in 1.6.1 - we have to repartition our dataset into many thousands of partitions to avoid the following stack:\n\n{code}\n2016-05-08 16:05:11,941 ERROR org.apache.spark.executor.Executor: Managed memory leak detected; size = 5748783225 bytes, TID = 1283\n2016-05-08 16:05:11,948 ERROR org.apache.spark.executor.Executor: Exception in task 116.4 in stage 1.0 (TID 1283)\njava.lang.OutOfMemoryError: Unable to acquire 2383 bytes of memory, got 0\n at org.apache.spark.memory.MemoryConsumer.allocatePage(MemoryConsumer.java:120) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.shuffle.sort.ShuffleExternalSorter.acquireNewPageIfNecessary(ShuffleExternalSorter.java:346) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.shuffle.sort.ShuffleExternalSorter.insertRecord(ShuffleExternalSorter.java:367) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.shuffle.sort.UnsafeShuffleWriter.insertRecordIntoSorter(UnsafeShuffleWriter.java:237) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.shuffle.sort.UnsafeShuffleWriter.write(UnsafeShuffleWriter.java:164) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.scheduler.Task.run(Task.scala:89) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [na:1.8.0_91]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_91]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_91]{code}","created":"2016-05-09T17:56:41.315+0000"},{"body":"User 'lianhuiwang' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13027","created":"2016-05-10T14:13:07.094+0000"},{"body":"I think your patch is not related to this bug, right?","created":"2016-05-31T21:36:23.653+0000"},{"body":"This issue was fixed in 1.6.0 as part of the patch for SPARK-10984 (my original PR for that issue was subsumed by the PR for SPARK-10984). We will not backport this patch to 1.5.x or 1.4.x because we do not plan to have further non-security-fix releases for those branches.","created":"2016-05-31T21:59:01.215+0000"},{"body":"Not sure if this is related, but I am running on spark 2.0.2 through spark job-server and see tons of messages like this:\n{code}\n[2016-12-20 11:49:28,662] WARN he.spark.executor.Executor [] [akka://JobServer/user/context-supervisor/sql-context] - Managed memory leak detected; size = 5762976 bytes, TID = 42621\n[2016-12-20 11:49:28,662] WARN k.memory.TaskMemoryManager [] [akka://JobServer/user/context-supervisor/sql-context] - leak 5.5 MB memory from org.apache.spark.util.collection.ExternalSorter@35f81493\n[2016-12-20 11:49:28,662] WARN he.spark.executor.Executor [] [akka://JobServer/user/context-supervisor/sql-context] - Managed memory leak detected; size = 5762976 bytes, TID = 42622\n[2016-12-20 11:49:28,702] WARN k.memory.TaskMemoryManager [] [akka://JobServer/user/context-supervisor/sql-context] - leak 5.5 MB memory from org.apache.spark.util.collection.ExternalSorter@16da7c1a\n[2016-12-20 11:49:28,702] WARN he.spark.executor.Executor [] [akka://JobServer/user/context-supervisor/sql-context] - Managed memory leak detected; size = 5762976 bytes, TID = 42623\n[2016-12-20 11:49:28,702] WARN k.memory.TaskMemoryManager [] [akka://JobServer/user/context-supervisor/sql-context] - leak 5.5 MB memory from org.apache.spark.util.collection.ExternalSorter@151060cf\n[2016-12-20 11:49:28,702] WARN he.spark.executor.Executor [] [akka://JobServer/user/context-supervisor/sql-context] - Managed memory leak detected; size = 5762976 bytes, TID = 42624\n[Stage 5700:=========================> (44 + 4) / 92][2016-12-20 11:49:35,479] WARN k.memory.TaskMemoryManager [] [akka://JobServer/user/context-supervisor/sql-context] - l\n{code}\nAre managed memory leaks ever expected behavior? Or do they always indicate a memory leak problem? I don't really see the memory going up much in jVisualVM.","created":"2016-12-20T19:54:06.228+0000"},{"body":"I hit the sam problem when I run TPCH test on spark1.6.0.\r\n\r\nmy dataset scale is SF=1000,\r\n\r\nEnvironment as follows:\r\n\r\n1Master 3 Worker,\r\n\r\nonHeapMemory=10g,\r\n\r\noffHeapMemory=20g,\r\n\r\n24threads/Worker\r\n\r\nthe query3 and query17 detected memory leak.Some  logs are as follows: \r\n\r\n9/04/09 21:57:59 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2685\r\n 41 19/04/09 21:58:16 WARN TaskMemoryManager: leak 512.0 MB memory from org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@385b3b3f\r\n 42 19/04/09 21:58:16 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2683\r\n 43 19/04/09 21:58:16 WARN TaskMemoryManager: leak 512.0 MB memory from org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@35be7b55\r\n 44 19/04/09 21:58:16 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2703\r\n 45 19/04/09 21:58:20 WARN TaskMemoryManager: leak 512.0 MB memory from org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@50f93582\r\n 46 19/04/09 21:58:20 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2709\r\n 47 19/04/09 21:58:21 WARN TaskMemoryManager: leak 512.0 MB memory from org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@28e3ec7a\r\n 48 19/04/09 21:58:21 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2723\r\n 49 19/04/09 21:59:50 WARN TaskMemoryManager: leak 512.0 MB memory from org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@5b2f5dbc\r\n 50 19/04/09 21:59:50 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2687\r\n 51 19/04/09 22:00:50 WARN TransportChannelHandler: Exception in connection from hw083/172.18.11.83:42989","created":"2019-04-09T14:40:50.002+0000"}],"conversations":[{"body":"I discovered multiple leaks of shuffle memory while working on my memory manager consolidation patch, which added the ability to do strict memory leak detection for the bookkeeping that used to be performed by the ShuffleMemoryManager. This uncovered a handful of places where tasks can acquire execution/shuffle memory but never release it, starving themselves of memory.\n\nProblems that I found:\n\n* {{ExternalSorter.stop()}} should release the sorter's shuffle/execution memory.\n* BlockStoreShuffleReader should call {{ExternalSorter.stop()}} using a {{CompletionIterator}}.\n* {{ExternalAppendOnlyMap}} exposes no equivalent of {{stop()}} for freeing its resources.","from":"reporter","subject":"ExternalSorter and ExternalAppendOnlyMap should free shuffle memory in their stop() methods"},{"body":"User 'JoshRosen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9260","from":"developer"},{"body":"This was fixed for Spark 1.6.0 via a different patch (memory manager consolidation), but I'm going to try to pull my earlier patch into 1.5.x.","from":"developer"},{"body":"User 'JoshRosen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9427","from":"developer"},{"body":"The memory manager was rewritten there? Could it have introduced a memory leak in a different place or of a different kind? Is there a regression test to verify?","from":"developer"},{"body":"Since the back-port for 1.5 was cancelled, I think this is Fixed.","from":"developer"},{"body":"I have a somewhat contrived example that still leaks in 1.6.0. I started {{spark-shell --master 'local-cluster[2,2,1024]'}} and ran:\n\n{code}\nsc.parallelize(0 to 10000000, 2).map(x => x % 10000 -> x).groupByKey.asInstanceOf[org.apache.spark.rdd.ShuffledRDD[Int, Int, Iterable[Int]]].setKeyOrdering(implicitly[Ordering[Int]]).mapPartitions { it => it.take(1) }.collect\n{code}\n\nI've added extra logging around task memory acquisition so I would be able to see what is not released. These are the logs:\n\n{code}\n16/01/05 17:02:45 INFO Executor: Running task 0.0 in stage 13.0 (TID 24)\n16/01/05 17:02:45 INFO MapOutputTrackerWorker: Updating epoch to 7 and clearing cache\n16/01/05 17:02:45 INFO TorrentBroadcast: Started reading broadcast variable 13\n16/01/05 17:02:45 INFO MemoryStore: Block broadcast_13_piece0 stored as bytes in memory (estimated size 2.3 KB, free 7.6 KB)\n16/01/05 17:02:45 INFO TorrentBroadcast: Reading broadcast variable 13 took 6 ms\n16/01/05 17:02:45 INFO MemoryStore: Block broadcast_13 stored as values in memory (estimated size 4.5 KB, free 12.1 KB)\n16/01/05 17:02:45 INFO MapOutputTrackerWorker: Don't have map outputs for shuffle 6, fetching them\n16/01/05 17:02:45 INFO MapOutputTrackerWorker: Doing the fetch; tracker endpoint = NettyRpcEndpointRef(spark://MapOutputTracker@192.168.0.32:55147)\n16/01/05 17:02:45 INFO MapOutputTrackerWorker: Got the output locations\n16/01/05 17:02:45 INFO ShuffleBlockFetcherIterator: Getting 2 non-empty blocks out of 2 blocks\n16/01/05 17:02:45 INFO ShuffleBlockFetcherIterator: Started 1 remote fetches in 1 ms\n16/01/05 17:02:45 ERROR TaskMemoryManager: Task 24 acquire 5.0 MB for null\n16/01/05 17:02:45 ERROR TaskMemoryManager: Stack trace:\njava.lang.Exception: here\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:187)\n\tat org.apache.spark.util.collection.Spillable$class.maybeSpill(Spillable.scala:82)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.maybeSpill(ExternalAppendOnlyMap.scala:55)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.insertAll(ExternalAppendOnlyMap.scala:158)\n\tat org.apache.spark.Aggregator.combineValuesByKey(Aggregator.scala:45)\n\tat org.apache.spark.shuffle.BlockStoreShuffleReader.read(BlockStoreShuffleReader.scala:89)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:98)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/01/05 17:02:47 ERROR TaskMemoryManager: Task 24 acquire 15.0 MB for null\n16/01/05 17:02:47 ERROR TaskMemoryManager: Stack trace:\njava.lang.Exception: here\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:187)\n\tat org.apache.spark.util.collection.Spillable$class.maybeSpill(Spillable.scala:82)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.maybeSpill(ExternalAppendOnlyMap.scala:55)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.insertAll(ExternalAppendOnlyMap.scala:158)\n\tat org.apache.spark.Aggregator.combineValuesByKey(Aggregator.scala:45)\n\tat org.apache.spark.shuffle.BlockStoreShuffleReader.read(BlockStoreShuffleReader.scala:89)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:98)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/01/05 17:02:49 ERROR TaskMemoryManager: Task 24 acquire 5.0 MB for null\n16/01/05 17:02:49 ERROR TaskMemoryManager: Stack trace:\njava.lang.Exception: here\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:187)\n\tat org.apache.spark.util.collection.Spillable$class.maybeSpill(Spillable.scala:82)\n\tat org.apache.spark.util.collection.ExternalSorter.maybeSpill(ExternalSorter.scala:89)\n\tat org.apache.spark.util.collection.ExternalSorter.maybeSpillCollection(ExternalSorter.scala:220)\n\tat org.apache.spark.util.collection.ExternalSorter.insertAll(ExternalSorter.scala:201)\n\tat org.apache.spark.shuffle.BlockStoreShuffleReader.read(BlockStoreShuffleReader.scala:103)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:98)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/01/05 17:02:49 ERROR TaskMemoryManager: Task 24 acquire 10.5 MB for null\n16/01/05 17:02:49 ERROR TaskMemoryManager: Stack trace:\njava.lang.Exception: here\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:187)\n\tat org.apache.spark.util.collection.Spillable$class.maybeSpill(Spillable.scala:82)\n\tat org.apache.spark.util.collection.ExternalSorter.maybeSpill(ExternalSorter.scala:89)\n\tat org.apache.spark.util.collection.ExternalSorter.maybeSpillCollection(ExternalSorter.scala:220)\n\tat org.apache.spark.util.collection.ExternalSorter.insertAll(ExternalSorter.scala:201)\n\tat org.apache.spark.shuffle.BlockStoreShuffleReader.read(BlockStoreShuffleReader.scala:103)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:98)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/01/05 17:02:49 ERROR TaskMemoryManager: Task 24 release 20.0 MB from null\n16/01/05 17:02:49 ERROR TaskMemoryManager: Stack trace:\njava.lang.Exception: here\n\tat org.apache.spark.memory.TaskMemoryManager.releaseExecutionMemory(TaskMemoryManager.java:197)\n\tat org.apache.spark.util.collection.Spillable$class.releaseMemory(Spillable.scala:111)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.releaseMemory(ExternalAppendOnlyMap.scala:55)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.org$apache$spark$util$collection$ExternalAppendOnlyMap$$freeCurrentMap(ExternalAppendOnlyMap.scala:259)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap$$anonfun$iterator$1.apply$mcV$sp(ExternalAppendOnlyMap.scala:251)\n\tat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\n\tat org.apache.spark.util.collection.ExternalSorter.insertAll(ExternalSorter.scala:197)\n\tat org.apache.spark.shuffle.BlockStoreShuffleReader.read(BlockStoreShuffleReader.scala:103)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:98)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/01/05 17:02:49 ERROR Executor: Managed memory leak detected; size = 16259594 bytes, TID = 24\n{code}\n\nThe issue is that {{ExternalSorter.stop()}} is only called by {{CompletionIterator}} if the iterator is iterated through to the end. But here we only take the first element.\n\nIn practice this happens to us in a {{zipPartitions}} call where we do not iterate both iterators to the end. (It's a kind of join.)\n\nIs it illegal to not iterate an RDD iterator to the end? I think it's not. {{RDD.take}} stops short as well. This issue can probably be reproduced with {{RDD.take}} too. (I tried and failed.)","from":"developer"},{"body":"Sorry, my example was overly complicated. This one triggers the same leak.\n\n{code}\nsc.parallelize(0 to 10000000, 2).map(x => x % 10000 -> x).groupByKey.mapPartitions { it => it.take(1) }.collect\n{code}","from":"developer"},{"body":"so should be reopened or not? is there still a memory leak? is there a new memory leak instead of the old one?","from":"developer"},{"body":"> so should be reopened or not? is there still a memory leak? is there a new memory leak instead of the old one?\n\nI think it should be reopened, because the remaining leak is an edge case of the original problem that was not covered by Josh's fix. I'll reopen it and we'll see what he thinks!","from":"developer"},{"body":"so add 1.6.0 as affected version...","from":"developer"},{"body":"> so add 1.6.0 as affected version...\n\nDone.","from":"developer"},{"body":"Not iterating to the end has a bunch of issues IIRC - including what you mention above. For example, m'mapped buffers are not released, etc.\nUnfortunately, I dont think there is a general clean solution for it. Would be good to see what alternatives exist to resolve this.","from":"developer"},{"body":"Whyle trying to execute Analytics Triangle Count with http://snap.stanford.edu/data/com-Friendster.html (with LiveJournal works ok) on about 24 nodes (spark standalone 1.52, around 32g memory each node, 128 partitions, parallelism cores*nodes*8) I get some errors that are maybe related:\n16/02/18 08:52:11 ERROR Executor: Managed memory leak detected; size = 67108864 bytes, TID = 507\n16/02/18 08:52:11 INFO MapOutputTrackerWorker: Updating epoch to 8 and clearing cache\n16/02/18 08:52:11 INFO TorrentBroadcast: Started reading broadcast variable 6\n16/02/18 08:52:11 ERROR Executor: Exception in task 123.0 in stage 3.0 (TID 507)\njava.lang.OutOfMemoryError: Java heap space\n at scala.reflect.ManifestFactory$$anon$10.newArray(Manifest.scala:122)\n at scala.reflect.ManifestFactory$$anon$10.newArray(Manifest.scala:120)\n at org.apache.spark.util.collection.OpenHashSet.rehash(OpenHashSet.scala:231)\n at org.apache.spark.util.collection.OpenHashSet.rehashIfNeeded(OpenHashSet.scala:166)\n at org.apache.spark.util.collection.OpenHashSet.rehashIfNeeded$mcJ$sp(OpenHashSet.scala:164)\n at org.apache.spark.graphx.util.collection.GraphXPrimitiveKeyOpenHashMap$mcJI$sp.changeValue$mcJI$sp(GraphXPrimitiveKeyOpenHashMap.scala:107)\n at org.apache.spark.graphx.impl.EdgePartitionBuilder.toEdgePartition(EdgePartitionBuilder.scala:58)\n at org.apache.spark.graphx.impl.GraphImpl$$anonfun$4.apply(GraphImpl.scala:115)\n at org.apache.spark.graphx.impl.GraphImpl$$anonfun$4.apply(GraphImpl.scala:109)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$18.apply(RDD.scala:727)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$18.apply(RDD.scala:727)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:300)\n at org.apache.spark.CacheManager.getOrCompute(CacheManager.scala:69)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:262)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:300)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:88)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)","from":"developer"},{"body":"Saw a similar issue loading the Friendster graph into Tinkerpop and running the BulkVertexLoader with the Shuffled Vertex RDD persisted to disk. ","from":"developer"},{"body":"[~joshrosen] any plans to fix this? I believe we are hitting this issue with {{1.6.0}}\n\n{code}\n16/03/17 21:38:38 INFO memory.TaskMemoryManager: Acquired by org.apache.spark.shuffle.sort.ShuffleExternalSorter@6ecba8f1: 32.0 KB\n16/03/17 21:38:38 INFO memory.TaskMemoryManager: 1528015093 bytes of memory were used by task 103134 but are not associated with specific consumers\n16/03/17 21:38:38 INFO memory.TaskMemoryManager: 1528047861 bytes of memory are used for execution and 80608434 bytes of memory are used for storage\n16/03/17 21:38:38 ERROR executor.Executor: Managed memory leak detected; size = 1528015093 bytes, TID = 103134\n16/03/17 21:38:38 ERROR executor.Executor: Exception in task 448.0 in stage 273.0 (TID 103134)\njava.lang.OutOfMemoryError: Unable to acquire 128 bytes of memory, got 0\n\tat org.apache.spark.memory.MemoryConsumer.allocatePage(MemoryConsumer.java:120)\n\tat org.apache.spark.shuffle.sort.ShuffleExternalSorter.acquireNewPageIfNecessary(ShuffleExternalSorter.java:354)\n\tat org.apache.spark.shuffle.sort.ShuffleExternalSorter.insertRecord(ShuffleExternalSorter.java:375)\n\tat org.apache.spark.shuffle.sort.UnsafeShuffleWriter.insertRecordIntoSorter(UnsafeShuffleWriter.java:237)\n\tat org.apache.spark.shuffle.sort.UnsafeShuffleWriter.write(UnsafeShuffleWriter.java:164)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n16/03/17 21:38:38 ERROR util.SparkUncaughtExceptionHandler: Uncaught exception in thread Thread[Executor task launch worker-0,5,main]\njava.lang.OutOfMemoryError: Unable to acquire 128 bytes of memory, got 0\n\tat org.apache.spark.memory.MemoryConsumer.allocatePage(MemoryConsumer.java:120)\n\tat org.apache.spark.shuffle.sort.ShuffleExternalSorter.acquireNewPageIfNecessary(ShuffleExternalSorter.java:354)\n\tat org.apache.spark.shuffle.sort.ShuffleExternalSorter.insertRecord(ShuffleExternalSorter.java:375)\n\tat org.apache.spark.shuffle.sort.UnsafeShuffleWriter.insertRecordIntoSorter(UnsafeShuffleWriter.java:237)\n\tat org.apache.spark.shuffle.sort.UnsafeShuffleWriter.write(UnsafeShuffleWriter.java:164)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n16/03/17 21:38:38 INFO storage.DiskBlockManager: Shutdown hook called\n16/03/17 21:38:38 INFO util.ShutdownHookManager: Shutdown hook called\n{code}","from":"developer"},{"body":"It looks we are getting same errors on 1.6.0 too - any plans to address it?","from":"developer"},{"body":"seem to hit the issue with Spark 1.6.1 not sure if this is relative to this..if yes, it can be fixed in Spark 1.6.2 ?\n{code}\n16/03/22 23:10:26 INFO memory.TaskMemoryManager: Allocate page number 16 (67108864 bytes)\n16/03/22 23:10:26 INFO sort.UnsafeExternalSorter: Thread 221 spilling sort data of 1472.0 MB to disk (0 time so far)\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: Allocate page number 1 (1060044737 bytes)\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: Memory used in task 9302\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: Acquired by org.apache.spark.shuffle.sort.ShuffleExternalSorter@8bac554: 32.0 KB\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: Acquired by org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@7a117b4f: 512.0 MB\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: 0 bytes of memory were used by task 9302 but are not associated with specific consumers\n16/03/22 23:11:26 INFO memory.TaskMemoryManager: 14909439433 bytes of memory are used for execution and 1376877 bytes of memory are used for storage\n16/03/22 23:11:26 WARN memory.TaskMemoryManager: leak 32.0 KB memory from org.apache.spark.shuffle.sort.ShuffleExternalSorter@8bac554\n16/03/22 23:11:26 ERROR executor.Executor: Managed memory leak detected; size = 32768 bytes, TID = 9302\n16/03/22 23:11:26 ERROR executor.Executor: Exception in task 192.0 in stage 153.0 (TID 9302)\njava.lang.OutOfMemoryError: Unable to acquire 1073741824 bytes of memory, got 1060044737\n at org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:91)\n at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.growPointerArrayIfNecessary(UnsafeExternalSorter.java:295)\n at org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.insertRecord(UnsafeExternalSorter.java:330)\n at org.apache.spark.sql.execution.UnsafeExternalRowSorter.insertRow(UnsafeExternalRowSorter.java:91)\n at org.apache.spark.sql.execution.UnsafeExternalRowSorter.sort(UnsafeExternalRowSorter.java:168)\n at org.apache.spark.sql.execution.Sort$$anonfun$1.apply(Sort.scala:90)\n at org.apache.spark.sql.execution.Sort$$anonfun$1.apply(Sort.scala:64)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.rdd.ZippedPartitionsRDD2.compute(ZippedPartitionsRDD.scala:88)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{code}","from":"developer"},{"body":"I don't have such verbose logging as above but can confirm that the test case given produces memory leaks in Spark 1.6.1\n{code}\nscala> sc.parallelize(0 to 10000000, 2).map(x => x % 10000 -> x).groupByKey.mapPartitions { it => it.take(1) }.collect\n16/03/23 10:26:11 INFO SparkContext: Starting job: collect at :28\n16/03/23 10:26:11 INFO DAGScheduler: Registering RDD 1 (map at :28)\n16/03/23 10:26:11 INFO DAGScheduler: Got job 0 (collect at :28) with 2 output partitions\n16/03/23 10:26:11 INFO DAGScheduler: Final stage: ResultStage 1 (collect at :28)\n16/03/23 10:26:11 INFO DAGScheduler: Parents of final stage: List(ShuffleMapStage 0)\n16/03/23 10:26:11 INFO DAGScheduler: Missing parents: List(ShuffleMapStage 0)\n16/03/23 10:26:11 INFO DAGScheduler: Submitting ShuffleMapStage 0 (MapPartitionsRDD[1] at map at :28), which has no missing parents\n16/03/23 10:26:11 INFO MemoryStore: Block broadcast_0 stored as values in memory (estimated size 3.5 KB, free 3.5 KB)\n16/03/23 10:26:11 INFO MemoryStore: Block broadcast_0_piece0 stored as bytes in memory (estimated size 1955.0 B, free 5.4 KB)\n16/03/23 10:26:11 INFO BlockManagerInfo: Added broadcast_0_piece0 in memory on localhost:51671 (size: 1955.0 B, free: 511.5 MB)\n16/03/23 10:26:11 INFO SparkContext: Created broadcast 0 from broadcast at DAGScheduler.scala:1006\n16/03/23 10:26:11 INFO DAGScheduler: Submitting 2 missing tasks from ShuffleMapStage 0 (MapPartitionsRDD[1] at map at :28)\n16/03/23 10:26:11 INFO TaskSchedulerImpl: Adding task set 0.0 with 2 tasks\n16/03/23 10:26:11 INFO TaskSetManager: Starting task 0.0 in stage 0.0 (TID 0, localhost, partition 0,PROCESS_LOCAL, 2067 bytes)\n16/03/23 10:26:11 INFO TaskSetManager: Starting task 1.0 in stage 0.0 (TID 1, localhost, partition 1,PROCESS_LOCAL, 2124 bytes)\n16/03/23 10:26:11 INFO Executor: Running task 0.0 in stage 0.0 (TID 0)\n16/03/23 10:26:11 INFO Executor: Running task 1.0 in stage 0.0 (TID 1)\n16/03/23 10:26:21 INFO Executor: Finished task 1.0 in stage 0.0 (TID 1). 1159 bytes result sent to driver\n16/03/23 10:26:21 INFO TaskSetManager: Finished task 1.0 in stage 0.0 (TID 1) in 9152 ms on localhost (1/2)\n16/03/23 10:26:21 INFO Executor: Finished task 0.0 in stage 0.0 (TID 0). 1159 bytes result sent to driver\n16/03/23 10:26:21 INFO TaskSetManager: Finished task 0.0 in stage 0.0 (TID 0) in 9292 ms on localhost (2/2)\n16/03/23 10:26:21 INFO DAGScheduler: ShuffleMapStage 0 (map at :28) finished in 9.303 s\n16/03/23 10:26:21 INFO TaskSchedulerImpl: Removed TaskSet 0.0, whose tasks have all completed, from pool \n16/03/23 10:26:21 INFO DAGScheduler: looking for newly runnable stages\n16/03/23 10:26:21 INFO DAGScheduler: running: Set()\n16/03/23 10:26:21 INFO DAGScheduler: waiting: Set(ResultStage 1)\n16/03/23 10:26:21 INFO DAGScheduler: failed: Set()\n16/03/23 10:26:21 INFO DAGScheduler: Submitting ResultStage 1 (MapPartitionsRDD[3] at mapPartitions at :28), which has no missing parents\n16/03/23 10:26:21 INFO MemoryStore: Block broadcast_1 stored as values in memory (estimated size 4.5 KB, free 9.8 KB)\n16/03/23 10:26:21 INFO MemoryStore: Block broadcast_1_piece0 stored as bytes in memory (estimated size 2.3 KB, free 12.1 KB)\n16/03/23 10:26:21 INFO BlockManagerInfo: Added broadcast_1_piece0 in memory on localhost:51671 (size: 2.3 KB, free: 511.5 MB)\n16/03/23 10:26:21 INFO SparkContext: Created broadcast 1 from broadcast at DAGScheduler.scala:1006\n16/03/23 10:26:21 INFO DAGScheduler: Submitting 2 missing tasks from ResultStage 1 (MapPartitionsRDD[3] at mapPartitions at :28)\n16/03/23 10:26:21 INFO TaskSchedulerImpl: Adding task set 1.0 with 2 tasks\n16/03/23 10:26:21 INFO TaskSetManager: Starting task 0.0 in stage 1.0 (TID 2, localhost, partition 0,NODE_LOCAL, 1894 bytes)\n16/03/23 10:26:21 INFO TaskSetManager: Starting task 1.0 in stage 1.0 (TID 3, localhost, partition 1,NODE_LOCAL, 1894 bytes)\n16/03/23 10:26:21 INFO Executor: Running task 1.0 in stage 1.0 (TID 3)\n16/03/23 10:26:21 INFO Executor: Running task 0.0 in stage 1.0 (TID 2)\n16/03/23 10:26:21 INFO ShuffleBlockFetcherIterator: Getting 2 non-empty blocks out of 2 blocks\n16/03/23 10:26:21 INFO ShuffleBlockFetcherIterator: Getting 2 non-empty blocks out of 2 blocks\n16/03/23 10:26:21 INFO ShuffleBlockFetcherIterator: Started 0 remote fetches in 4 ms\n16/03/23 10:26:21 INFO ShuffleBlockFetcherIterator: Started 0 remote fetches in 5 ms\n16/03/23 10:26:30 ERROR Executor: Managed memory leak detected; size = 36707226 bytes, TID = 3\n16/03/23 10:26:30 ERROR Executor: Managed memory leak detected; size = 20935176 bytes, TID = 2\n...{code}\n\nCan we add 1.6.1 to the affected versions?","from":"developer"},{"body":"I suspect that most of these cases are caused by a CompletionIterator being limited via {{take()}}, causing the onComplete method to not be called. Given the limitations of our current iterator interface, I don't know that there's a straightforward way to deal with this.\n\nDoes anyone have an example of a case where this message appears in a situation that cannot be explained by {{take()}} / {{limit()}}?\n\nI'm thinking that maybe we should lower the severity of this message to WARN.","from":"developer"},{"body":"I've seen a few people misled by the error msg, so I'd like to try to downgrade that. I've created a separate ticket [SPARK-141868 | https://issues.apache.org/jira/browse/SPARK-14168] just for changing the msg, in case there is some fix in store here.","from":"developer"},{"body":"For users that are hitting *real* OOMs from similar error msgs (not just ERROR messages that are false-positives), they may be interested in SPARK-14560, which addresses one related issue from Spillables. Note that this isn't really a \"leak\", but can easily lead to OOMs in a shuffle-to-shuffle stage caused by Spillable collections.","from":"developer"},{"body":"I was using Apache spark 1.6 in EMR with spark streaming in yarn and saw memory leaks in one of the containers. Here are the logs\n\n{code}\n16/04/14 13:49:10 INFO executor.CoarseGrainedExecutorBackend: Got assigned task 2942916\n16/04/14 13:49:10 INFO executor.Executor: Running task 22.0 in stage 35684.0 (TID 2942915)\n16/04/14 13:49:10 INFO executor.Executor: Running task 23.0 in stage 35684.0 (TID 2942916)\n16/04/14 13:49:10 INFO storage.ShuffleBlockFetcherIterator: Getting 94 non-empty blocks out of 94 blocks\n16/04/14 13:49:10 INFO storage.ShuffleBlockFetcherIterator: Getting 94 non-empty blocks out of 94 blocks\n16/04/14 13:49:10 INFO storage.ShuffleBlockFetcherIterator: Started 2 remote fetches in 1 ms\n16/04/14 13:49:10 INFO storage.ShuffleBlockFetcherIterator: Started 2 remote fetches in 1 ms\n16/04/14 13:49:10 INFO storage.MemoryStore: Block input-3-1460583424327 stored as values in memory (estimated size 244.7 KB, free 19.3 MB)\n16/04/14 13:49:10 INFO receiver.BlockGenerator: Pushed block input-3-1460641750200\n16/04/14 13:49:10 INFO storage.MemoryStore: 1 blocks selected for dropping\n16/04/14 13:49:10 INFO storage.BlockManager: Dropping block input-1-1460615659379 from memory\n16/04/14 13:49:10 INFO storage.MemoryStore: 1 blocks selected for dropping\n16/04/14 13:49:10 INFO storage.BlockManager: Dropping block input-1-1460615659380 from memory\n16/04/14 13:49:10 INFO memory.TaskMemoryManager: Memory used in task 2942915\n16/04/14 13:49:10 INFO memory.TaskMemoryManager: Acquired by org.apache.spark.unsafe.map.BytesToBytesMap@34158d5f: 32.3 MB\n16/04/14 13:49:10 INFO memory.TaskMemoryManager: 0 bytes of memory were used by task 2942915 but are not associated with specific consumers\n16/04/14 13:49:10 INFO memory.TaskMemoryManager: 101247172 bytes of memory are used for execution and 3603881260 bytes of memory are used for storage\n16/04/14 13:49:10 WARN memory.TaskMemoryManager: leak 32.3 MB memory from org.apache.spark.unsafe.map.BytesToBytesMap@34158d5f\n16/04/14 13:49:10 ERROR executor.Executor: Managed memory leak detected; size = 33816576 bytes, TID = 2942915\n16/04/14 13:49:10 ERROR executor.Executor: Exception in task 22.0 in stage 35684.0 (TID 2942915)\njava.lang.OutOfMemoryError: Unable to acquire 262144 bytes of memory, got 220032\n\tat org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:91)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.allocate(BytesToBytesMap.java:735)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:197)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:212)\n\tat org.apache.spark.sql.execution.UnsafeFixedWidthAggregationMap.(UnsafeFixedWidthAggregationMap.java:103)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.(TungstenAggregationIterator.scala:483)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:95)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:86)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:710)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:710)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n16/04/14 13:49:10 INFO executor.Executor: Finished task 23.0 in stage 35684.0 (TID 2942916). 1921 bytes result sent to driver\n16/04/14 13:49:10 INFO executor.CoarseGrainedExecutorBackend: Got assigned task 2942927\n16/04/14 13:49:10 INFO executor.Executor: Running task 34.0 in stage 35684.0 (TID 2942927)\n16/04/14 13:49:10 ERROR util.SparkUncaughtExceptionHandler: Uncaught exception in thread Thread[Executor task launch worker-2,5,main]\njava.lang.OutOfMemoryError: Unable to acquire 262144 bytes of memory, got 220032\n\tat org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:91)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.allocate(BytesToBytesMap.java:735)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:197)\n\tat org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:212)\n\tat org.apache.spark.sql.execution.UnsafeFixedWidthAggregationMap.(UnsafeFixedWidthAggregationMap.java:103)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.(TungstenAggregationIterator.scala:483)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:95)\n\tat org.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:86)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:710)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:710)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}","from":"developer"},{"body":"Also biting us in 1.6.1 - we have to repartition our dataset into many thousands of partitions to avoid the following stack:\n\n{code}\n2016-05-08 16:05:11,941 ERROR org.apache.spark.executor.Executor: Managed memory leak detected; size = 5748783225 bytes, TID = 1283\n2016-05-08 16:05:11,948 ERROR org.apache.spark.executor.Executor: Exception in task 116.4 in stage 1.0 (TID 1283)\njava.lang.OutOfMemoryError: Unable to acquire 2383 bytes of memory, got 0\n at org.apache.spark.memory.MemoryConsumer.allocatePage(MemoryConsumer.java:120) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.shuffle.sort.ShuffleExternalSorter.acquireNewPageIfNecessary(ShuffleExternalSorter.java:346) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.shuffle.sort.ShuffleExternalSorter.insertRecord(ShuffleExternalSorter.java:367) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.shuffle.sort.UnsafeShuffleWriter.insertRecordIntoSorter(UnsafeShuffleWriter.java:237) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.shuffle.sort.UnsafeShuffleWriter.write(UnsafeShuffleWriter.java:164) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.scheduler.Task.run(Task.scala:89) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214) ~[d56f3336b4a0fcc71fe8beb90052dbafd0e88a749bdb4bbb15d37894cf443364-spark-core_2.11-1.6.1.jar:1.6.1]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) [na:1.8.0_91]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) [na:1.8.0_91]\n at java.lang.Thread.run(Thread.java:745) [na:1.8.0_91]{code}","from":"developer"},{"body":"User 'lianhuiwang' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13027","from":"developer"},{"body":"I think your patch is not related to this bug, right?","from":"developer"},{"body":"This issue was fixed in 1.6.0 as part of the patch for SPARK-10984 (my original PR for that issue was subsumed by the PR for SPARK-10984). We will not backport this patch to 1.5.x or 1.4.x because we do not plan to have further non-security-fix releases for those branches.","from":"developer"},{"body":"Not sure if this is related, but I am running on spark 2.0.2 through spark job-server and see tons of messages like this:\n{code}\n[2016-12-20 11:49:28,662] WARN he.spark.executor.Executor [] [akka://JobServer/user/context-supervisor/sql-context] - Managed memory leak detected; size = 5762976 bytes, TID = 42621\n[2016-12-20 11:49:28,662] WARN k.memory.TaskMemoryManager [] [akka://JobServer/user/context-supervisor/sql-context] - leak 5.5 MB memory from org.apache.spark.util.collection.ExternalSorter@35f81493\n[2016-12-20 11:49:28,662] WARN he.spark.executor.Executor [] [akka://JobServer/user/context-supervisor/sql-context] - Managed memory leak detected; size = 5762976 bytes, TID = 42622\n[2016-12-20 11:49:28,702] WARN k.memory.TaskMemoryManager [] [akka://JobServer/user/context-supervisor/sql-context] - leak 5.5 MB memory from org.apache.spark.util.collection.ExternalSorter@16da7c1a\n[2016-12-20 11:49:28,702] WARN he.spark.executor.Executor [] [akka://JobServer/user/context-supervisor/sql-context] - Managed memory leak detected; size = 5762976 bytes, TID = 42623\n[2016-12-20 11:49:28,702] WARN k.memory.TaskMemoryManager [] [akka://JobServer/user/context-supervisor/sql-context] - leak 5.5 MB memory from org.apache.spark.util.collection.ExternalSorter@151060cf\n[2016-12-20 11:49:28,702] WARN he.spark.executor.Executor [] [akka://JobServer/user/context-supervisor/sql-context] - Managed memory leak detected; size = 5762976 bytes, TID = 42624\n[Stage 5700:=========================> (44 + 4) / 92][2016-12-20 11:49:35,479] WARN k.memory.TaskMemoryManager [] [akka://JobServer/user/context-supervisor/sql-context] - l\n{code}\nAre managed memory leaks ever expected behavior? Or do they always indicate a memory leak problem? I don't really see the memory going up much in jVisualVM.","from":"developer"},{"body":"I hit the sam problem when I run TPCH test on spark1.6.0.\r\n\r\nmy dataset scale is SF=1000,\r\n\r\nEnvironment as follows:\r\n\r\n1Master 3 Worker,\r\n\r\nonHeapMemory=10g,\r\n\r\noffHeapMemory=20g,\r\n\r\n24threads/Worker\r\n\r\nthe query3 and query17 detected memory leak.Some  logs are as follows: \r\n\r\n9/04/09 21:57:59 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2685\r\n 41 19/04/09 21:58:16 WARN TaskMemoryManager: leak 512.0 MB memory from org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@385b3b3f\r\n 42 19/04/09 21:58:16 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2683\r\n 43 19/04/09 21:58:16 WARN TaskMemoryManager: leak 512.0 MB memory from org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@35be7b55\r\n 44 19/04/09 21:58:16 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2703\r\n 45 19/04/09 21:58:20 WARN TaskMemoryManager: leak 512.0 MB memory from org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@50f93582\r\n 46 19/04/09 21:58:20 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2709\r\n 47 19/04/09 21:58:21 WARN TaskMemoryManager: leak 512.0 MB memory from org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@28e3ec7a\r\n 48 19/04/09 21:58:21 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2723\r\n 49 19/04/09 21:59:50 WARN TaskMemoryManager: leak 512.0 MB memory from org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter@5b2f5dbc\r\n 50 19/04/09 21:59:50 ERROR Executor: Managed memory leak detected; size = 536870912 bytes, TID = 2687\r\n 51 19/04/09 22:00:50 WARN TransportChannelHandler: Exception in connection from hw083/172.18.11.83:42989","from":"developer"}],"created":"2015-10-24T00:29:05.000+0000","description":"I discovered multiple leaks of shuffle memory while working on my memory manager consolidation patch, which added the ability to do strict memory leak detection for the bookkeeping that used to be performed by the ShuffleMemoryManager. This uncovered a handful of places where tasks can acquire execution/shuffle memory but never release it, starving themselves of memory.\n\nProblems that I found:\n\n* {{ExternalSorter.stop()}} should release the sorter's shuffle/execution memory.\n* BlockStoreShuffleReader should call {{ExternalSorter.stop()}} using a {{CompletionIterator}}.\n* {{ExternalAppendOnlyMap}} exposes no equivalent of {{stop()}} for freeing its resources.","issue_id":"12907596","key":"SPARK-11293","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-05-31T21:59:01.000+0000","role":"fixed_distractor","summary":"ExternalSorter and ExternalAppendOnlyMap should free shuffle memory in their stop() methods"} {"case_id":"12921495","cluster":"DISTRACTOR-SPARK-12312","comments":[{"body":"So I just ran into this issue now trying to write to SQL Server. I have got to agree this is an important issue talking to SQL server - it is almost never allowed to use simple username/password authentication due to the security implications.\r\n\r\nCould there at least be a note in the docs that this is not possible? Say in third paragraph here: [https://spark.apache.org/docs/2.3.2/sql-programming-guide.html#jdbc-to-other-databases] where it talks about using username/password as connection properties? I've spent a very, very long time trying to figure out why this wasn't possible. The way the executors behave is rather odd if you try it, so it wasn't obvious why it didn't work (at least to me).\r\n\r\n ","created":"2018-10-18T02:32:37.975+0000"},{"body":"I agree! Can we please get this implemented as soon as possible?  This prevents us from being compliant with our internal security policies.  ","created":"2018-12-06T02:12:29.279+0000"},{"body":"The PR is discontinued for long time so I've picked this up and working on it.","created":"2020-01-16T11:41:22.188+0000"},{"body":"Re documentation - it really would be helpful and avoid time wasting if the documentation wrre updated to indicate that it is not possible to connect to a kerberisied JDBC at the moment. \r\n\r\nThis could be in the docs Foster mentions above and possibly also in \"Troubleshooting\"\r\n\r\nCan we fix the documentation first please as that seems low hanging fruit.","created":"2020-01-28T14:22:24.681+0000"},{"body":"I originally got around this problem using the approach sketched in this repo: \r\n\r\n[https://github.com/nabacg/krb5sqljdb]\r\n\r\nHope it might help someone.. it saved my project from missing an important deadline when we discovered this issue. Unfortunately as you say it's not very well documented and you often tend to discover this problem in later stage of the project, like when moving from DEV to UAT/PROD.. \r\n\r\n ","created":"2020-01-28T14:47:04.969+0000"},{"body":"[~nabacg] thanks, the approach is more or less clear but I'm creating and automated docker test which makes the PR hard. With manual test it's already working.","created":"2020-01-28T16:00:53.094+0000"},{"body":"My suggestion was for people who need a working solution now and can't wait till there is a new Spark release out and/or can't easily upgrade their cluster installations (which happens in corporate multi-tenant situations, I've been there.. ).  \r\n\r\nMy approach avoids those problems by patching the JDBC driver, instead of Spark. It's not a long term solution, but perhaps will save someone's skin. I certainly worked for me and one of my clients. ","created":"2020-01-28T16:19:51.362+0000"},{"body":"[~nabacg] thanks for sharin, pretty sure there are peoples who keen on using it.\r\nMy intention is to provide something for long term, since you have quite an experience in this area happy to hear you opinion when I've created the PR (and of course if you see gaps which can be filled just share).","created":"2020-01-29T11:34:04.100+0000"},{"body":"Sure, will do [~gsomogyi]","created":"2020-01-29T13:15:05.837+0000"},{"body":"@nabacg funnily enough we came up with almost the same solution but for oracle which so has this issue.\n\n\nBtw oracle is hideous as it tramples all over the static jaas config, wiping it out, which btw msoft also used to do but has fixed recently.\nAt least msoft open sourced the jdbc driver which surely encourages healthy debate and fixes. Come on oracle !!!!","created":"2020-01-29T21:29:49.505+0000"},{"body":"Yes we followed the same approach as referred to by nabacg too, wrapping the ms jdbc driver as explained at [https://datamountaineer.com/2016/01/15/spark-jdbc-sql-server-kerberos/].\r\n\r\nTends to be a bit brittle as ms updates the driver","created":"2020-01-29T22:45:44.674+0000"},{"body":"[~quarkonium] As an initial step the PR what I'm working on will introduce a docker test for postgres to see if later DB versions break something. It could be done for all the other DBs to clean this area up...","created":"2020-01-30T10:11:36.709+0000"},{"body":"After deep analysis it has turned out each database type needs custom implementation so creating subtasks to handle them.","created":"2020-02-19T12:37:58.945+0000"},{"body":"Soon the MS SQL and Oracle server is coming which is not possible to dockerize (fully featuring active directory setup is needed which I can't dockerize).\r\nI'm going to test it manually on cluster but if somebody wants to contribute by double checking just ping me.\r\n","created":"2020-04-22T12:20:19.751+0000"},{"body":"How does the current solution handle Kerberos renewals? Any design doc about the whole support?","created":"2020-04-26T23:03:17.326+0000"},{"body":"Not sure what you mean renewal. Renewal of what?\r\n * When keytab is invalid because the password has changed then the same must happen just like all other use-cases (re-distribute the keytab file which will be picked up properly when new connection is initiated)\r\n * When TGT is the question the solution re-obtains TGT automatically each and every time when new connection is created\r\n\r\nOther thing which can be renewed I'm not aware of.\r\n\r\n ","created":"2020-04-27T08:01:52.225+0000"},{"body":"I am trying to find out the potential conflicts if we implement a JDBC connection pool in JDBC batch or streaming sources/sinks.","created":"2020-04-29T08:00:42.581+0000"},{"body":"I think the current approach can be used to implement pooling. Of course couple of problems must be solved.\r\nWithout super deep consideration I would say task retry would be a good point to re-create connection just like it was in Kafka case.\r\nPooling really worth a design doc because each and every connector makes authentication differently...\r\n","created":"2020-04-29T09:44:33.754+0000"},{"body":"are we not solving this for hive database ?","created":"2020-05-20T15:15:47.599+0000"},{"body":"Not yet planned but we can take a look at this when the rest is done.","created":"2020-05-25T08:36:01.210+0000"},{"body":"Team,\r\n\r\nNeed your help on my issue and i followed entire thread. (Note : for me it is not possible to upgrade my version to 3.1 for now) \r\n\r\ni am using spark 2.4 and getting the same issue while connecting with Kerberos Oracle data base..\r\n\r\n1) When i run from local machine with keytab cache Spark job is running fine.\r\n\r\n2) When i run this  kubernetes platform , it throws with an exception as kerberos principle account authentication failed .\r\n\r\nam using ojdbc8 jar to connect Oracle data base. \r\n\r\n \r\n\r\nThis issue comes like sporadic some times  it is connected and some times it is not connected.  Am not why this inconsistency with 2.4 version.  Please can you suggest is there any work around on this with 2.4 version ? \r\n\r\n \r\n\r\n ","created":"2020-08-12T13:18:22.087+0000"},{"body":"Was this oracle?\n\nWe experienced a problem when running out jobs on yarn but not local.\nThe problem stems from the fact that spark (also flink) run the application\ncode in a context where a Kerberos principal has been setup to talk to\nHadoop.\nThen when oracle jdbc comes along then it has a bug in my opinion in that\ninstead of creating a new principal for the database connection it instead\nreuses the ambient principal used for Hadoop.\nIf you turn on Kerberos/ojdbc trace you can see the Hadoop principal being\nused against the database.\n\nI believe Microsoft also used to suffer from this but was fixed.\n\nWorse to come however. If you use the oracle jdbc Kerberos then you might\nnotice that once oracle has done it's connection then the JAAS config in\nyour app is trashed. Another bug in oracle imho. Again I believe msoft used\nto do the same.\n\nOne solution is to use a single principal for Hadoop and oracle. If you\ncan't do that then you may need to create your own oracle driver wrapper\nthat compensates for the ambient Hadoop principal allowing oracle to\nproceed with the intended id. If you do the latter then you will also\nlikely end up remediating the JAAS corruption issue.\n\nI expect many folk never encounter these issues unless they are in\ncorporate environments where different principal for each resource is a\ncommon pattern.\n\nCheers John\n\n\n","created":"2020-08-12T13:46:00.119+0000"},{"body":"[~Arun Tupuri] and [~johnlon] - I am looking for the same work around to connect to kerberosed Oracle through Spark Scala jdbc job. I want to test this approach in local/client and cluster mode as well. Please can i get any reference for this approach.\r\n\r\n[~Arun Tupuri] - Can i get any reference code to test in local mode to connect to kerberosed oracle using spark.\r\n\r\nThanks,\r\n\r\nRavi.","created":"2020-08-15T14:57:28.538+0000"},{"body":"I think we fixed this in oracle along the lines of this post...\n\nhttps://stackoverflow.com/questions/52662249/javakerberos-authentication-to-sql-server-on-spark-framework\n\nSee the references to datamountaineer and also the git repo mentioned gives\ngood ideas...\nhttps://github.com/nabacg/krb5sqljdb\n\nYou should be able to get it working for oracle as I have. We followed\npretty much the same approach as ...\n\nhttps://github.com/nabacg/krb5sqljdb/blob/master/src/main/scala/net/cabworks/jdbc/Krb5SqlServer.scala\n\n.... but we used java and oracle rather than scala and mssql\n\nIt was fairly straightforward.\n\n\nOn Sat, 15 Aug 2020, 3:58 pm Ravi Tummalapenta (Jira), \n\n","created":"2020-08-15T15:17:00.124+0000"},{"body":"[~johnlon] the above references use keytab and principal to establish kerberos authentication to Oracle. IN my scenario, I do not have control over keytab file, as there will be a sidecar which takes care of setting up the kerb5cache file to my pod. So the executor has to use only the krb5cache file. Is this scenario also works?","created":"2020-09-23T21:15:38.358+0000"},{"body":"Yes - the driver wrapper I wrote accepted both a KT or a ticket cache.,\n\nSee the reference I gave.\n\nIf I recall correctly we provided the option of selecting either approach\nof ...\nUserGroupInformation.loginUserFromKeytabAndReturnUGI(principal, keytabFile)\n\nOR\nUserGroupInformation.getUGIFromTicketCache(principal, cache)\n\nJL\n\nOn Wed, 23 Sep 2020 at 22:17, Prakash Rajendran (Jira) \n\n","created":"2020-09-23T22:54:00.115+0000"},{"body":"Tried spark 3.1.0.  Still can not connect to oracle using kerberos.\r\n\r\nBelow is the configuration. This works when I use only driver, and fails when I start using executor\r\n\r\nconnectionProperties.setProperty(\"jdbc.auth.principal\",EnvConfig().getString(\"gosPrincipal\"))\r\nconnectionProperties.setProperty(\"oracle.net.kerberos5_cc_name\", EnvConfig().getString(\"gosKrb5cache\"))\r\nconnectionProperties.setProperty(\"oracle.net.kerberos5_mutual_authentication\", \"true\")\r\nconnectionProperties.setProperty(\"oracle.net.authentication_services\",\"KERBEROS5\")\r\nconnectionProperties.setProperty(\"driver\", \"oracle.jdbc.driver.OracleDriver\")\r\n\r\n \r\n\r\nError message : : java.sql.SQLRecoverableException: IO Error: The service in process is not supported. Unable to find valid kerberos principal for authentication\r\n\r\n ","created":"2021-06-19T18:09:33.305+0000"}],"conversations":[{"body":"When loading DataFrames from JDBC datasource with Kerberos authentication, remote executors (yarn-client/cluster etc. modes) fail to establish a connection due to lack of Kerberos ticket or ability to generate it. \n\nThis is a real issue when trying to ingest data from kerberized data sources (SQL Server, Oracle) in enterprise environment where exposing simple authentication access is not an option due to IT policy issues.","from":"reporter","subject":"Support JDBC Kerberos w/ keytab"},{"body":"So I just ran into this issue now trying to write to SQL Server. I have got to agree this is an important issue talking to SQL server - it is almost never allowed to use simple username/password authentication due to the security implications.\r\n\r\nCould there at least be a note in the docs that this is not possible? Say in third paragraph here: [https://spark.apache.org/docs/2.3.2/sql-programming-guide.html#jdbc-to-other-databases] where it talks about using username/password as connection properties? I've spent a very, very long time trying to figure out why this wasn't possible. The way the executors behave is rather odd if you try it, so it wasn't obvious why it didn't work (at least to me).\r\n\r\n ","from":"developer"},{"body":"I agree! Can we please get this implemented as soon as possible?  This prevents us from being compliant with our internal security policies.  ","from":"developer"},{"body":"The PR is discontinued for long time so I've picked this up and working on it.","from":"developer"},{"body":"Re documentation - it really would be helpful and avoid time wasting if the documentation wrre updated to indicate that it is not possible to connect to a kerberisied JDBC at the moment. \r\n\r\nThis could be in the docs Foster mentions above and possibly also in \"Troubleshooting\"\r\n\r\nCan we fix the documentation first please as that seems low hanging fruit.","from":"developer"},{"body":"I originally got around this problem using the approach sketched in this repo: \r\n\r\n[https://github.com/nabacg/krb5sqljdb]\r\n\r\nHope it might help someone.. it saved my project from missing an important deadline when we discovered this issue. Unfortunately as you say it's not very well documented and you often tend to discover this problem in later stage of the project, like when moving from DEV to UAT/PROD.. \r\n\r\n ","from":"developer"},{"body":"[~nabacg] thanks, the approach is more or less clear but I'm creating and automated docker test which makes the PR hard. With manual test it's already working.","from":"developer"},{"body":"My suggestion was for people who need a working solution now and can't wait till there is a new Spark release out and/or can't easily upgrade their cluster installations (which happens in corporate multi-tenant situations, I've been there.. ).  \r\n\r\nMy approach avoids those problems by patching the JDBC driver, instead of Spark. It's not a long term solution, but perhaps will save someone's skin. I certainly worked for me and one of my clients. ","from":"developer"},{"body":"[~nabacg] thanks for sharin, pretty sure there are peoples who keen on using it.\r\nMy intention is to provide something for long term, since you have quite an experience in this area happy to hear you opinion when I've created the PR (and of course if you see gaps which can be filled just share).","from":"developer"},{"body":"Sure, will do [~gsomogyi]","from":"developer"},{"body":"@nabacg funnily enough we came up with almost the same solution but for oracle which so has this issue.\n\n\nBtw oracle is hideous as it tramples all over the static jaas config, wiping it out, which btw msoft also used to do but has fixed recently.\nAt least msoft open sourced the jdbc driver which surely encourages healthy debate and fixes. Come on oracle !!!!","from":"developer"},{"body":"Yes we followed the same approach as referred to by nabacg too, wrapping the ms jdbc driver as explained at [https://datamountaineer.com/2016/01/15/spark-jdbc-sql-server-kerberos/].\r\n\r\nTends to be a bit brittle as ms updates the driver","from":"developer"},{"body":"[~quarkonium] As an initial step the PR what I'm working on will introduce a docker test for postgres to see if later DB versions break something. It could be done for all the other DBs to clean this area up...","from":"developer"},{"body":"After deep analysis it has turned out each database type needs custom implementation so creating subtasks to handle them.","from":"developer"},{"body":"Soon the MS SQL and Oracle server is coming which is not possible to dockerize (fully featuring active directory setup is needed which I can't dockerize).\r\nI'm going to test it manually on cluster but if somebody wants to contribute by double checking just ping me.\r\n","from":"developer"},{"body":"How does the current solution handle Kerberos renewals? Any design doc about the whole support?","from":"developer"},{"body":"Not sure what you mean renewal. Renewal of what?\r\n * When keytab is invalid because the password has changed then the same must happen just like all other use-cases (re-distribute the keytab file which will be picked up properly when new connection is initiated)\r\n * When TGT is the question the solution re-obtains TGT automatically each and every time when new connection is created\r\n\r\nOther thing which can be renewed I'm not aware of.\r\n\r\n ","from":"developer"},{"body":"I am trying to find out the potential conflicts if we implement a JDBC connection pool in JDBC batch or streaming sources/sinks.","from":"developer"},{"body":"I think the current approach can be used to implement pooling. Of course couple of problems must be solved.\r\nWithout super deep consideration I would say task retry would be a good point to re-create connection just like it was in Kafka case.\r\nPooling really worth a design doc because each and every connector makes authentication differently...\r\n","from":"developer"},{"body":"are we not solving this for hive database ?","from":"developer"},{"body":"Not yet planned but we can take a look at this when the rest is done.","from":"developer"},{"body":"Team,\r\n\r\nNeed your help on my issue and i followed entire thread. (Note : for me it is not possible to upgrade my version to 3.1 for now) \r\n\r\ni am using spark 2.4 and getting the same issue while connecting with Kerberos Oracle data base..\r\n\r\n1) When i run from local machine with keytab cache Spark job is running fine.\r\n\r\n2) When i run this  kubernetes platform , it throws with an exception as kerberos principle account authentication failed .\r\n\r\nam using ojdbc8 jar to connect Oracle data base. \r\n\r\n \r\n\r\nThis issue comes like sporadic some times  it is connected and some times it is not connected.  Am not why this inconsistency with 2.4 version.  Please can you suggest is there any work around on this with 2.4 version ? \r\n\r\n \r\n\r\n ","from":"developer"},{"body":"Was this oracle?\n\nWe experienced a problem when running out jobs on yarn but not local.\nThe problem stems from the fact that spark (also flink) run the application\ncode in a context where a Kerberos principal has been setup to talk to\nHadoop.\nThen when oracle jdbc comes along then it has a bug in my opinion in that\ninstead of creating a new principal for the database connection it instead\nreuses the ambient principal used for Hadoop.\nIf you turn on Kerberos/ojdbc trace you can see the Hadoop principal being\nused against the database.\n\nI believe Microsoft also used to suffer from this but was fixed.\n\nWorse to come however. If you use the oracle jdbc Kerberos then you might\nnotice that once oracle has done it's connection then the JAAS config in\nyour app is trashed. Another bug in oracle imho. Again I believe msoft used\nto do the same.\n\nOne solution is to use a single principal for Hadoop and oracle. If you\ncan't do that then you may need to create your own oracle driver wrapper\nthat compensates for the ambient Hadoop principal allowing oracle to\nproceed with the intended id. If you do the latter then you will also\nlikely end up remediating the JAAS corruption issue.\n\nI expect many folk never encounter these issues unless they are in\ncorporate environments where different principal for each resource is a\ncommon pattern.\n\nCheers John\n\n\n","from":"developer"},{"body":"[~Arun Tupuri] and [~johnlon] - I am looking for the same work around to connect to kerberosed Oracle through Spark Scala jdbc job. I want to test this approach in local/client and cluster mode as well. Please can i get any reference for this approach.\r\n\r\n[~Arun Tupuri] - Can i get any reference code to test in local mode to connect to kerberosed oracle using spark.\r\n\r\nThanks,\r\n\r\nRavi.","from":"developer"},{"body":"I think we fixed this in oracle along the lines of this post...\n\nhttps://stackoverflow.com/questions/52662249/javakerberos-authentication-to-sql-server-on-spark-framework\n\nSee the references to datamountaineer and also the git repo mentioned gives\ngood ideas...\nhttps://github.com/nabacg/krb5sqljdb\n\nYou should be able to get it working for oracle as I have. We followed\npretty much the same approach as ...\n\nhttps://github.com/nabacg/krb5sqljdb/blob/master/src/main/scala/net/cabworks/jdbc/Krb5SqlServer.scala\n\n.... but we used java and oracle rather than scala and mssql\n\nIt was fairly straightforward.\n\n\nOn Sat, 15 Aug 2020, 3:58 pm Ravi Tummalapenta (Jira), \n\n","from":"developer"},{"body":"[~johnlon] the above references use keytab and principal to establish kerberos authentication to Oracle. IN my scenario, I do not have control over keytab file, as there will be a sidecar which takes care of setting up the kerb5cache file to my pod. So the executor has to use only the krb5cache file. Is this scenario also works?","from":"developer"},{"body":"Yes - the driver wrapper I wrote accepted both a KT or a ticket cache.,\n\nSee the reference I gave.\n\nIf I recall correctly we provided the option of selecting either approach\nof ...\nUserGroupInformation.loginUserFromKeytabAndReturnUGI(principal, keytabFile)\n\nOR\nUserGroupInformation.getUGIFromTicketCache(principal, cache)\n\nJL\n\nOn Wed, 23 Sep 2020 at 22:17, Prakash Rajendran (Jira) \n\n","from":"developer"},{"body":"Tried spark 3.1.0.  Still can not connect to oracle using kerberos.\r\n\r\nBelow is the configuration. This works when I use only driver, and fails when I start using executor\r\n\r\nconnectionProperties.setProperty(\"jdbc.auth.principal\",EnvConfig().getString(\"gosPrincipal\"))\r\nconnectionProperties.setProperty(\"oracle.net.kerberos5_cc_name\", EnvConfig().getString(\"gosKrb5cache\"))\r\nconnectionProperties.setProperty(\"oracle.net.kerberos5_mutual_authentication\", \"true\")\r\nconnectionProperties.setProperty(\"oracle.net.authentication_services\",\"KERBEROS5\")\r\nconnectionProperties.setProperty(\"driver\", \"oracle.jdbc.driver.OracleDriver\")\r\n\r\n \r\n\r\nError message : : java.sql.SQLRecoverableException: IO Error: The service in process is not supported. Unable to find valid kerberos principal for authentication\r\n\r\n ","from":"developer"}],"created":"2015-12-13T16:56:00.000+0000","description":"When loading DataFrames from JDBC datasource with Kerberos authentication, remote executors (yarn-client/cluster etc. modes) fail to establish a connection due to lack of Kerberos ticket or ability to generate it. \n\nThis is a real issue when trying to ingest data from kerberized data sources (SQL Server, Oracle) in enterprise environment where exposing simple authentication access is not an option due to IT policy issues.","issue_id":"12921495","key":"SPARK-12312","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2021-02-10T07:34:28.000+0000","role":"fixed_distractor","summary":"Support JDBC Kerberos w/ keytab"} {"case_id":"12928758","cluster":"DISTRACTOR-SPARK-12717","comments":[{"body":"I'm having the same problem when trying to optimize a production environment.\nIs there any update on this issue? Is it being picked up?\n\nEDIT: It's still an issue in pyspark 2.0.2\n\nAlexey","created":"2016-12-02T21:30:10.055+0000"},{"body":"It still happens in the current master.","created":"2017-01-12T02:02:07.026+0000"},{"body":"I met the same problem with pyspark 2.1.0. Any progress?","created":"2017-04-06T05:57:22.666+0000"},{"body":"Same problem with Python3.\n\nCC: [~davies]","created":"2017-04-14T12:08:28.829+0000"},{"body":"User 'vundela' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17694","created":"2017-04-20T01:47:03.424+0000"},{"body":"Please find the attached log with fix for the following command \nspark2-submit --master local[20] bug_spark.py --parallelism 1000 >& run.log","created":"2017-04-20T18:15:22.051+0000"},{"body":"User 'vundela' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17722","created":"2017-04-21T18:46:05.163+0000"},{"body":"Any progress with this error ?\nI have patched from this thread merged and it's working fine.\nMaybe we can merge this to master? (and other branches)","created":"2017-07-19T08:05:45.124+0000"},{"body":"User 'BryanCutler' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18695","created":"2017-07-20T21:53:02.762+0000"},{"body":"Issue resolved by pull request 18695\nhttps://github.com/apache/spark/pull/18695\n\nbut it looks I can't set {{Assignee}}. Could anyone help me set this and resolve this?","created":"2017-08-01T22:34:17.551+0000"},{"body":"[~hyukjin.kwon] I added you to the Committers group in JIRA, maybe that lets you. But I also just assigned this one.","created":"2017-08-01T22:37:42.255+0000"},{"body":"I see. Thank you. Yes, I am seeing now.","created":"2017-08-01T22:38:33.760+0000"},{"body":"Thanks [~hyukjin.kwon]! What are your thoughts on backporting this for 2.2? ","created":"2017-08-01T23:55:15.683+0000"},{"body":"Would you mind if I ask open a backport? Just want to check if it passes the test for sure.","created":"2017-08-02T00:04:40.093+0000"},{"body":"Sure, I'll open a PR for 2.2 and ping you.","created":"2017-08-02T16:59:04.072+0000"},{"body":"This looks closed, but I still had this same issue again:\npython 3.6.1\nspark 2.2.0","created":"2017-09-25T17:43:53.628+0000"},{"body":"Hi [~avloss], the fix will be in Spark 2.1.2 which will be released soon, Spark 2.2.1 which is still pending, and Spark 2.3 when that is released.","created":"2017-09-25T17:57:14.344+0000"},{"body":"[~bryanc] I use pyspark 2.2.0, got the same error. which pyspark version is ok ? thankyou!","created":"2018-01-15T08:47:25.562+0000"},{"body":"Hi [~codlife], you can use Spark 2.2.1 which was released in December or the upcoming release of 2.3.0, both have this fix.","created":"2018-01-15T17:42:22.887+0000"},{"body":"User 'BryanCutler' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18825","created":"2018-11-19T05:48:37.067+0000"},{"body":"User 'BryanCutler' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18825","created":"2018-11-19T05:49:48.945+0000"}],"conversations":[{"body":"The following multi-threaded program that uses broadcast variables consistently throws exceptions like: *Exception(\"Broadcast variable '18' not loaded!\",)* --- even when run with \"--master local[10]\".\n\n{code:title=bug_spark.py|borderStyle=solid}\ntry: \n import pyspark \nexcept: \n pass \nfrom optparse import OptionParser \n \ndef my_option_parser(): \n op = OptionParser() \n op.add_option(\"--parallelism\", dest=\"parallelism\", type=\"int\", default=20) \n return op \n \ndef do_process(x, w): \n return x * w.value \n \ndef func(name, rdd, conf): \n new_rdd = rdd.map(lambda x : do_process(x, conf)) \n total = new_rdd.reduce(lambda x, y : x + y) \n count = rdd.count() \n print name, 1.0 * total / count \n \nif __name__ == \"__main__\": \n import threading \n op = my_option_parser() \n options, args = op.parse_args() \n sc = pyspark.SparkContext(appName=\"Buggy\") \n data_rdd = sc.parallelize(range(0,1000), 1) \n confs = [ sc.broadcast(i) for i in xrange(options.parallelism) ] \n threads = [ threading.Thread(target=func, args=[\"thread_\" + str(i), data_rdd, confs[i]]) for i in xrange(options.parallelism) ] \n for t in threads: \n t.start() \n for t in threads: \n t.join() \n{code}\n\nAbridged run output:\n\n{code:title=abridge_run.txt|borderStyle=solid}\n% spark-submit --master local[10] bug_spark.py --parallelism 20\n[snip]\n16/01/08 17:10:20 ERROR Executor: Exception in task 0.0 in stage 9.0 (TID 9)\norg.apache.spark.api.python.PythonException: Traceback (most recent call last):\n File \"/Network/Servers/mother.adverplex.com/Volumes/homeland/Users/walker/.spark/spark-1.6.0-bin-hadoop2.6/python/lib/pyspark.zip/pyspark/worker.py\", line 98, in main\n command = pickleSer._read_with_length(infile)\n File \"/Network/Servers/mother.adverplex.com/Volumes/homeland/Users/walker/.spark/spark-1.6.0-bin-hadoop2.6/python/lib/pyspark.zip/pyspark/serializers.py\", line 164, in _read_with_length\n return self.loads(obj)\n File \"/Network/Servers/mother.adverplex.com/Volumes/homeland/Users/walker/.spark/spark-1.6.0-bin-hadoop2.6/python/lib/pyspark.zip/pyspark/serializers.py\", line 422, in loads\n return pickle.loads(obj)\n File \"/Network/Servers/mother.adverplex.com/Volumes/homeland/Users/walker/.spark/spark-1.6.0-bin-hadoop2.6/python/lib/pyspark.zip/pyspark/broadcast.py\", line 39, in _from_id\n raise Exception(\"Broadcast variable '%s' not loaded!\" % bid)\nException: (Exception(\"Broadcast variable '6' not loaded!\",), , (6L,))\n\n\tat org.apache.spark.api.python.PythonRunner$$anon$1.read(PythonRDD.scala:166)\n\tat org.apache.spark.api.python.PythonRunner$$anon$1.(PythonRDD.scala:207)\n\tat org.apache.spark.api.python.PythonRunner.compute(PythonRDD.scala:125)\n\tat org.apache.spark.api.python.PythonRDD.compute(PythonRDD.scala:70)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n[snip]\n{code}","from":"reporter","subject":"pyspark broadcast fails when using multiple threads"},{"body":"I'm having the same problem when trying to optimize a production environment.\nIs there any update on this issue? Is it being picked up?\n\nEDIT: It's still an issue in pyspark 2.0.2\n\nAlexey","from":"developer"},{"body":"It still happens in the current master.","from":"developer"},{"body":"I met the same problem with pyspark 2.1.0. Any progress?","from":"developer"},{"body":"Same problem with Python3.\n\nCC: [~davies]","from":"developer"},{"body":"User 'vundela' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17694","from":"developer"},{"body":"Please find the attached log with fix for the following command \nspark2-submit --master local[20] bug_spark.py --parallelism 1000 >& run.log","from":"developer"},{"body":"User 'vundela' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17722","from":"developer"},{"body":"Any progress with this error ?\nI have patched from this thread merged and it's working fine.\nMaybe we can merge this to master? (and other branches)","from":"developer"},{"body":"User 'BryanCutler' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18695","from":"developer"},{"body":"Issue resolved by pull request 18695\nhttps://github.com/apache/spark/pull/18695\n\nbut it looks I can't set {{Assignee}}. Could anyone help me set this and resolve this?","from":"developer"},{"body":"[~hyukjin.kwon] I added you to the Committers group in JIRA, maybe that lets you. But I also just assigned this one.","from":"developer"},{"body":"I see. Thank you. Yes, I am seeing now.","from":"developer"},{"body":"Thanks [~hyukjin.kwon]! What are your thoughts on backporting this for 2.2? ","from":"developer"},{"body":"Would you mind if I ask open a backport? Just want to check if it passes the test for sure.","from":"developer"},{"body":"Sure, I'll open a PR for 2.2 and ping you.","from":"developer"},{"body":"This looks closed, but I still had this same issue again:\npython 3.6.1\nspark 2.2.0","from":"developer"},{"body":"Hi [~avloss], the fix will be in Spark 2.1.2 which will be released soon, Spark 2.2.1 which is still pending, and Spark 2.3 when that is released.","from":"developer"},{"body":"[~bryanc] I use pyspark 2.2.0, got the same error. which pyspark version is ok ? thankyou!","from":"developer"},{"body":"Hi [~codlife], you can use Spark 2.2.1 which was released in December or the upcoming release of 2.3.0, both have this fix.","from":"developer"},{"body":"User 'BryanCutler' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18825","from":"developer"},{"body":"User 'BryanCutler' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18825","from":"developer"}],"created":"2016-01-08T22:18:19.000+0000","description":"The following multi-threaded program that uses broadcast variables consistently throws exceptions like: *Exception(\"Broadcast variable '18' not loaded!\",)* --- even when run with \"--master local[10]\".\n\n{code:title=bug_spark.py|borderStyle=solid}\ntry: \n import pyspark \nexcept: \n pass \nfrom optparse import OptionParser \n \ndef my_option_parser(): \n op = OptionParser() \n op.add_option(\"--parallelism\", dest=\"parallelism\", type=\"int\", default=20) \n return op \n \ndef do_process(x, w): \n return x * w.value \n \ndef func(name, rdd, conf): \n new_rdd = rdd.map(lambda x : do_process(x, conf)) \n total = new_rdd.reduce(lambda x, y : x + y) \n count = rdd.count() \n print name, 1.0 * total / count \n \nif __name__ == \"__main__\": \n import threading \n op = my_option_parser() \n options, args = op.parse_args() \n sc = pyspark.SparkContext(appName=\"Buggy\") \n data_rdd = sc.parallelize(range(0,1000), 1) \n confs = [ sc.broadcast(i) for i in xrange(options.parallelism) ] \n threads = [ threading.Thread(target=func, args=[\"thread_\" + str(i), data_rdd, confs[i]]) for i in xrange(options.parallelism) ] \n for t in threads: \n t.start() \n for t in threads: \n t.join() \n{code}\n\nAbridged run output:\n\n{code:title=abridge_run.txt|borderStyle=solid}\n% spark-submit --master local[10] bug_spark.py --parallelism 20\n[snip]\n16/01/08 17:10:20 ERROR Executor: Exception in task 0.0 in stage 9.0 (TID 9)\norg.apache.spark.api.python.PythonException: Traceback (most recent call last):\n File \"/Network/Servers/mother.adverplex.com/Volumes/homeland/Users/walker/.spark/spark-1.6.0-bin-hadoop2.6/python/lib/pyspark.zip/pyspark/worker.py\", line 98, in main\n command = pickleSer._read_with_length(infile)\n File \"/Network/Servers/mother.adverplex.com/Volumes/homeland/Users/walker/.spark/spark-1.6.0-bin-hadoop2.6/python/lib/pyspark.zip/pyspark/serializers.py\", line 164, in _read_with_length\n return self.loads(obj)\n File \"/Network/Servers/mother.adverplex.com/Volumes/homeland/Users/walker/.spark/spark-1.6.0-bin-hadoop2.6/python/lib/pyspark.zip/pyspark/serializers.py\", line 422, in loads\n return pickle.loads(obj)\n File \"/Network/Servers/mother.adverplex.com/Volumes/homeland/Users/walker/.spark/spark-1.6.0-bin-hadoop2.6/python/lib/pyspark.zip/pyspark/broadcast.py\", line 39, in _from_id\n raise Exception(\"Broadcast variable '%s' not loaded!\" % bid)\nException: (Exception(\"Broadcast variable '6' not loaded!\",), , (6L,))\n\n\tat org.apache.spark.api.python.PythonRunner$$anon$1.read(PythonRDD.scala:166)\n\tat org.apache.spark.api.python.PythonRunner$$anon$1.(PythonRDD.scala:207)\n\tat org.apache.spark.api.python.PythonRunner.compute(PythonRDD.scala:125)\n\tat org.apache.spark.api.python.PythonRDD.compute(PythonRDD.scala:70)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n[snip]\n{code}","issue_id":"12928758","key":"SPARK-12717","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-08-01T22:40:29.000+0000","role":"fixed_distractor","summary":"pyspark broadcast fails when using multiple threads"} {"case_id":"12929509","cluster":"DISTRACTOR-SPARK-12777","comments":[{"body":"Is this in the shell? if so, there is a different more fundamental problem with case classes in the shell, already reported in JIRA.","created":"2016-01-12T15:31:17.119+0000"},{"body":"Nope, this is just in normal code. Although I have encountered the bug with case classes in the shell as well.","created":"2016-01-12T22:00:34.056+0000"},{"body":"I get the same error in SparkShell, however everything works in a plain application (as shown in the listing)\n{code}\nimport org.apache.spark._\nimport org.apache.spark.sql._\n\ncase class Test(v: (Int, Int))\n\nobject Main {\n\n val conf = new SparkConf().setMaster(\"local\").setAppName(\"testbench\")\n val sc = new SparkContext(conf)\n val sqlContext = new SQLContext(sc)\n\n def main(args: Array[String]): Unit = {\n import sqlContext.implicits._\n\n val rdd = sc.parallelize(\n Seq(\n Test((1,2)),\n Test((3,4))))\n val ds = sqlContext.createDataset(rdd).toDS\n ds.show\n }\n}\n{code}","created":"2016-01-14T00:04:43.906+0000"},{"body":"Concerning the problem with type aliases, I can reproduce them both inside a SparkShell and inside a standalone program. See issue SPARK-12816 and related PR.","created":"2016-01-14T01:40:22.972+0000"},{"body":"[~janstenpickle] Could you please recheck and change the status of the issue.","created":"2016-05-24T17:41:19.406+0000"},{"body":"This works in 2.1:\n\nhttps://databricks-prod-cloudfront.cloud.databricks.com/public/4027ec902e239c93eaaa8714f173bcfc/1023043053387187/408017793305293/2840265927289860/latest.html","created":"2016-12-15T22:47:07.784+0000"},{"body":"The issue reproduces for me in standalone application with Spark 2.2.0 when the type is explicitly specified.\n\n{code:scala}\nimport org.apache.spark.rdd.RDD\nimport org.apache.spark.sql.SparkSession\nimport org.apache.spark.{SparkConf, SparkContext}\n\nobject Main extends App {\n\n val sc = new SparkContext(\"local\", \"hello\", new SparkConf())\n val ss = SparkSession.builder().getOrCreate()\n\n import ss.implicits._\n\n val seq = Seq((1, (2, 3)))\n\n type A = (Int, Int)\n\n val rdd: RDD[(Int, A)] = sc.parallelize(seq)\n //val rdd: RDD[(Int, (Int, Int))] = sc.parallelize(seq) // works fine\n //val rdd = sc.parallelize(seq) // works fine\n\n ss.createDataset(rdd).show()\n\n sc.stop()\n}\n{code}\n\nStacktrace (Spark 2.2.0):\n{code:java}\nscala.ScalaReflectionException: type T1 is not a class\n at scala.reflect.api.Symbols$SymbolApi$class.asClass(Symbols.scala:275)\n at scala.reflect.internal.Symbols$SymbolContextApiImpl.asClass(Symbols.scala:84)\n at org.apache.spark.sql.catalyst.ScalaReflection$.getClassFromType(ScalaReflection.scala:682)\n at org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$dataTypeFor(ScalaReflection.scala:84)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$10.apply(ScalaReflection.scala:614)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$10.apply(ScalaReflection.scala:607)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\n at scala.collection.immutable.List.flatMap(List.scala:344)\n at org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$serializerFor(ScalaReflection.scala:607)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$10.apply(ScalaReflection.scala:619)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$10.apply(ScalaReflection.scala:607)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\n at scala.collection.immutable.List.flatMap(List.scala:344)\n at org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$serializerFor(ScalaReflection.scala:607)\n at org.apache.spark.sql.catalyst.ScalaReflection$.serializerFor(ScalaReflection.scala:438)\n at org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$.apply(ExpressionEncoder.scala:71)\n at org.apache.spark.sql.Encoders$.product(Encoders.scala:275)\n at org.apache.spark.sql.LowPrioritySQLImplicits$class.newProductEncoder(SQLImplicits.scala:233)\n at org.apache.spark.sql.SQLImplicits.newProductEncoder(SQLImplicits.scala:33)\n at Main$.delayedEndpoint$Main$1(Main.scala:20)\n at Main$delayedInit$body.apply(Main.scala:5)\n at scala.Function0$class.apply$mcV$sp(Function0.scala:34)\n at scala.runtime.AbstractFunction0.apply$mcV$sp(AbstractFunction0.scala:12)\n at scala.App$$anonfun$main$1.apply(App.scala:76)\n at scala.App$$anonfun$main$1.apply(App.scala:76)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.generic.TraversableForwarder$class.foreach(TraversableForwarder.scala:35)\n at scala.App$class.main(App.scala:76)\n at Main$.main(Main.scala:5)\n at Main.main(Main.scala)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n{code}\n\nStacktrace (Spark 2.1.0):\n{code:java}\njava.util.NoSuchElementException: head of empty list\n at scala.collection.immutable.Nil$.head(List.scala:420)\n at scala.collection.immutable.Nil$.head(List.scala:417)\n at scala.reflect.internal.tpe.TypeMaps$SubstMap.subst(TypeMaps.scala:709)\n at scala.reflect.internal.tpe.TypeMaps$SubstMap.substFor$1(TypeMaps.scala:717)\n at scala.reflect.internal.tpe.TypeMaps$SubstMap.apply(TypeMaps.scala:732)\n at scala.reflect.internal.Types$Type.subst(Types.scala:705)\n at scala.reflect.internal.Types$TypeApiImpl.substituteTypes(Types.scala:240)\n at scala.reflect.internal.Types$TypeApiImpl.substituteTypes(Types.scala:218)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$getConstructorParameters$1.apply(ScalaReflection.scala:818)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$getConstructorParameters$1.apply(ScalaReflection.scala:817)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.immutable.List.map(List.scala:285)\n at org.apache.spark.sql.catalyst.ScalaReflection$class.getConstructorParameters(ScalaReflection.scala:817)\n at org.apache.spark.sql.catalyst.ScalaReflection$.getConstructorParameters(ScalaReflection.scala:39)\n at org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$serializerFor(ScalaReflection.scala:586)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$9.apply(ScalaReflection.scala:596)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$9.apply(ScalaReflection.scala:587)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\n at scala.collection.immutable.List.flatMap(List.scala:344)\n at org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$serializerFor(ScalaReflection.scala:587)\n at org.apache.spark.sql.catalyst.ScalaReflection$.serializerFor(ScalaReflection.scala:425)\n at org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$.apply(ExpressionEncoder.scala:71)\n at org.apache.spark.sql.Encoders$.product(Encoders.scala:275)\n at org.apache.spark.sql.SQLImplicits.newProductEncoder(SQLImplicits.scala:49)\n at Main$.delayedEndpoint$Main$1(Main.scala:20)\n at Main$delayedInit$body.apply(Main.scala:5)\n at scala.Function0$class.apply$mcV$sp(Function0.scala:34)\n at scala.runtime.AbstractFunction0.apply$mcV$sp(AbstractFunction0.scala:12)\n at scala.App$$anonfun$main$1.apply(App.scala:76)\n at scala.App$$anonfun$main$1.apply(App.scala:76)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.generic.TraversableForwarder$class.foreach(TraversableForwarder.scala:35)\n at scala.App$class.main(App.scala:76)\n at Main$.main(Main.scala:5)\n at Main.main(Main.scala)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n{code}","created":"2017-07-17T10:40:48.712+0000"}],"conversations":[{"body":"Datasets can't seem to handle scala tuples as fields of case classes in datasets.\n\n{code}\nSeq((1,2), (3,4)).toDS().show() //works\n{code}\n\nWhen including a tuple as a field, the code fails:\n{code}\ncase class Test(v: (Int, Int))\n\nSeq(Test((1,2)), Test((3,4)).toDS().show //fails\n{code}\n\n{code}\n UnresolvedException: : Invalid call to dataType on unresolved object, tree: 'name (unresolved.scala:59)\n org.apache.spark.sql.catalyst.analysis.UnresolvedAttribute.dataType(unresolved.scala:59)\n org.apache.spark.sql.catalyst.expressions.GetStructField.org$apache$spark$sql$catalyst$expressions$GetStructField$$field$lzycompute(complexTypeExtractors.scala:107)\n org.apache.spark.sql.catalyst.expressions.GetStructField.org$apache$spark$sql$catalyst$expressions$GetStructField$$field(complexTypeExtractors.scala:107)\n org.apache.spark.sql.catalyst.expressions.GetStructField$$anonfun$toString$1.apply(complexTypeExtractors.scala:111)\n org.apache.spark.sql.catalyst.expressions.GetStructField$$anonfun$toString$1.apply(complexTypeExtractors.scala:111)\n org.apache.spark.sql.catalyst.expressions.GetStructField.toString(complexTypeExtractors.scala:111)\n org.apache.spark.sql.catalyst.expressions.Expression.toString(Expression.scala:217)\n org.apache.spark.sql.catalyst.expressions.Expression.toString(Expression.scala:217)\n org.apache.spark.sql.catalyst.expressions.If.toString(conditionalExpressions.scala:76)\n org.apache.spark.sql.catalyst.expressions.Expression.toString(Expression.scala:217)\n org.apache.spark.sql.catalyst.expressions.Alias.toString(namedExpressions.scala:155)\n org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$argString$1.apply(TreeNode.scala:385)\n org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$argString$1.apply(TreeNode.scala:381)\n org.apache.spark.sql.catalyst.trees.TreeNode.argString(TreeNode.scala:388)\n org.apache.spark.sql.catalyst.trees.TreeNode.simpleString(TreeNode.scala:391)\n org.apache.spark.sql.catalyst.plans.QueryPlan.simpleString(QueryPlan.scala:172)\n org.apache.spark.sql.catalyst.trees.TreeNode.generateTreeString(TreeNode.scala:441)\n org.apache.spark.sql.catalyst.trees.TreeNode.treeString(TreeNode.scala:396)\n org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$5.apply(RuleExecutor.scala:118)\n org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$5.apply(RuleExecutor.scala:119)\n org.apache.spark.Logging$class.logDebug(Logging.scala:62)\n org.apache.spark.sql.catalyst.rules.RuleExecutor.logDebug(RuleExecutor.scala:44)\n org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:115)\n org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:72)\n org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:72)\n org.apache.spark.sql.catalyst.encoders.ExpressionEncoder.resolve(ExpressionEncoder.scala:253)\n org.apache.spark.sql.Dataset.(Dataset.scala:78)\n org.apache.spark.sql.Dataset.(Dataset.scala:89)\n org.apache.spark.sql.SQLContext.createDataset(SQLContext.scala:507)\n org.apache.spark.sql.SQLImplicits.localSeqToDatasetHolder(SQLImplicits.scala:80)\n{code}\n\n\nWhen providing a type alias, the code fails in a different way:\n{code}\ntype TwoInt = (Int, Int)\n\ncase class Test(v: TwoInt)\n\nSeq(Test((1,2)), Test((3,4)).toDS().show //fails\n{code}\n\n{code}\n NoSuchElementException: : head of empty list (ScalaReflection.scala:504)\n org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor$1.apply(ScalaReflection.scala:504)\n org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor$1.apply(ScalaReflection.scala:502)\n org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor(ScalaReflection.scala:502)\n org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor$1.apply(ScalaReflection.scala:509)\n org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor$1.apply(ScalaReflection.scala:502)\n org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor(ScalaReflection.scala:502)\n org.apache.spark.sql.catalyst.ScalaReflection$.extractorsFor(ScalaReflection.scala:394)\n org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$.apply(ExpressionEncoder.scala:54)\n org.apache.spark.sql.SQLImplicits.newProductEncoder(SQLImplicits.scala:41)\n com.intenthq.pipeline.actions.ActionsJobIntegrationSpec.enrich(ActionsJobIntegrationSpec.scala:63)\n com.intenthq.pipeline.actions.ActionsJobIntegrationSpec$$anonfun$is$2.apply(ActionsJobIntegrationSpec.scala:45)\n com.intenthq.pipeline.actions.ActionsJobIntegrationSpec$$anonfun$is$2.apply(ActionsJobIntegrationSpec.scala:45)\n{code}","from":"reporter","subject":"Dataset fields can't be Scala tuples"},{"body":"Is this in the shell? if so, there is a different more fundamental problem with case classes in the shell, already reported in JIRA.","from":"developer"},{"body":"Nope, this is just in normal code. Although I have encountered the bug with case classes in the shell as well.","from":"developer"},{"body":"I get the same error in SparkShell, however everything works in a plain application (as shown in the listing)\n{code}\nimport org.apache.spark._\nimport org.apache.spark.sql._\n\ncase class Test(v: (Int, Int))\n\nobject Main {\n\n val conf = new SparkConf().setMaster(\"local\").setAppName(\"testbench\")\n val sc = new SparkContext(conf)\n val sqlContext = new SQLContext(sc)\n\n def main(args: Array[String]): Unit = {\n import sqlContext.implicits._\n\n val rdd = sc.parallelize(\n Seq(\n Test((1,2)),\n Test((3,4))))\n val ds = sqlContext.createDataset(rdd).toDS\n ds.show\n }\n}\n{code}","from":"developer"},{"body":"Concerning the problem with type aliases, I can reproduce them both inside a SparkShell and inside a standalone program. See issue SPARK-12816 and related PR.","from":"developer"},{"body":"[~janstenpickle] Could you please recheck and change the status of the issue.","from":"developer"},{"body":"This works in 2.1:\n\nhttps://databricks-prod-cloudfront.cloud.databricks.com/public/4027ec902e239c93eaaa8714f173bcfc/1023043053387187/408017793305293/2840265927289860/latest.html","from":"developer"},{"body":"The issue reproduces for me in standalone application with Spark 2.2.0 when the type is explicitly specified.\n\n{code:scala}\nimport org.apache.spark.rdd.RDD\nimport org.apache.spark.sql.SparkSession\nimport org.apache.spark.{SparkConf, SparkContext}\n\nobject Main extends App {\n\n val sc = new SparkContext(\"local\", \"hello\", new SparkConf())\n val ss = SparkSession.builder().getOrCreate()\n\n import ss.implicits._\n\n val seq = Seq((1, (2, 3)))\n\n type A = (Int, Int)\n\n val rdd: RDD[(Int, A)] = sc.parallelize(seq)\n //val rdd: RDD[(Int, (Int, Int))] = sc.parallelize(seq) // works fine\n //val rdd = sc.parallelize(seq) // works fine\n\n ss.createDataset(rdd).show()\n\n sc.stop()\n}\n{code}\n\nStacktrace (Spark 2.2.0):\n{code:java}\nscala.ScalaReflectionException: type T1 is not a class\n at scala.reflect.api.Symbols$SymbolApi$class.asClass(Symbols.scala:275)\n at scala.reflect.internal.Symbols$SymbolContextApiImpl.asClass(Symbols.scala:84)\n at org.apache.spark.sql.catalyst.ScalaReflection$.getClassFromType(ScalaReflection.scala:682)\n at org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$dataTypeFor(ScalaReflection.scala:84)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$10.apply(ScalaReflection.scala:614)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$10.apply(ScalaReflection.scala:607)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\n at scala.collection.immutable.List.flatMap(List.scala:344)\n at org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$serializerFor(ScalaReflection.scala:607)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$10.apply(ScalaReflection.scala:619)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$10.apply(ScalaReflection.scala:607)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\n at scala.collection.immutable.List.flatMap(List.scala:344)\n at org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$serializerFor(ScalaReflection.scala:607)\n at org.apache.spark.sql.catalyst.ScalaReflection$.serializerFor(ScalaReflection.scala:438)\n at org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$.apply(ExpressionEncoder.scala:71)\n at org.apache.spark.sql.Encoders$.product(Encoders.scala:275)\n at org.apache.spark.sql.LowPrioritySQLImplicits$class.newProductEncoder(SQLImplicits.scala:233)\n at org.apache.spark.sql.SQLImplicits.newProductEncoder(SQLImplicits.scala:33)\n at Main$.delayedEndpoint$Main$1(Main.scala:20)\n at Main$delayedInit$body.apply(Main.scala:5)\n at scala.Function0$class.apply$mcV$sp(Function0.scala:34)\n at scala.runtime.AbstractFunction0.apply$mcV$sp(AbstractFunction0.scala:12)\n at scala.App$$anonfun$main$1.apply(App.scala:76)\n at scala.App$$anonfun$main$1.apply(App.scala:76)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.generic.TraversableForwarder$class.foreach(TraversableForwarder.scala:35)\n at scala.App$class.main(App.scala:76)\n at Main$.main(Main.scala:5)\n at Main.main(Main.scala)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n{code}\n\nStacktrace (Spark 2.1.0):\n{code:java}\njava.util.NoSuchElementException: head of empty list\n at scala.collection.immutable.Nil$.head(List.scala:420)\n at scala.collection.immutable.Nil$.head(List.scala:417)\n at scala.reflect.internal.tpe.TypeMaps$SubstMap.subst(TypeMaps.scala:709)\n at scala.reflect.internal.tpe.TypeMaps$SubstMap.substFor$1(TypeMaps.scala:717)\n at scala.reflect.internal.tpe.TypeMaps$SubstMap.apply(TypeMaps.scala:732)\n at scala.reflect.internal.Types$Type.subst(Types.scala:705)\n at scala.reflect.internal.Types$TypeApiImpl.substituteTypes(Types.scala:240)\n at scala.reflect.internal.Types$TypeApiImpl.substituteTypes(Types.scala:218)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$getConstructorParameters$1.apply(ScalaReflection.scala:818)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$getConstructorParameters$1.apply(ScalaReflection.scala:817)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.immutable.List.map(List.scala:285)\n at org.apache.spark.sql.catalyst.ScalaReflection$class.getConstructorParameters(ScalaReflection.scala:817)\n at org.apache.spark.sql.catalyst.ScalaReflection$.getConstructorParameters(ScalaReflection.scala:39)\n at org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$serializerFor(ScalaReflection.scala:586)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$9.apply(ScalaReflection.scala:596)\n at org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$9.apply(ScalaReflection.scala:587)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\n at scala.collection.immutable.List.flatMap(List.scala:344)\n at org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$serializerFor(ScalaReflection.scala:587)\n at org.apache.spark.sql.catalyst.ScalaReflection$.serializerFor(ScalaReflection.scala:425)\n at org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$.apply(ExpressionEncoder.scala:71)\n at org.apache.spark.sql.Encoders$.product(Encoders.scala:275)\n at org.apache.spark.sql.SQLImplicits.newProductEncoder(SQLImplicits.scala:49)\n at Main$.delayedEndpoint$Main$1(Main.scala:20)\n at Main$delayedInit$body.apply(Main.scala:5)\n at scala.Function0$class.apply$mcV$sp(Function0.scala:34)\n at scala.runtime.AbstractFunction0.apply$mcV$sp(AbstractFunction0.scala:12)\n at scala.App$$anonfun$main$1.apply(App.scala:76)\n at scala.App$$anonfun$main$1.apply(App.scala:76)\n at scala.collection.immutable.List.foreach(List.scala:381)\n at scala.collection.generic.TraversableForwarder$class.foreach(TraversableForwarder.scala:35)\n at scala.App$class.main(App.scala:76)\n at Main$.main(Main.scala:5)\n at Main.main(Main.scala)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n{code}","from":"developer"}],"created":"2016-01-12T15:17:24.000+0000","description":"Datasets can't seem to handle scala tuples as fields of case classes in datasets.\n\n{code}\nSeq((1,2), (3,4)).toDS().show() //works\n{code}\n\nWhen including a tuple as a field, the code fails:\n{code}\ncase class Test(v: (Int, Int))\n\nSeq(Test((1,2)), Test((3,4)).toDS().show //fails\n{code}\n\n{code}\n UnresolvedException: : Invalid call to dataType on unresolved object, tree: 'name (unresolved.scala:59)\n org.apache.spark.sql.catalyst.analysis.UnresolvedAttribute.dataType(unresolved.scala:59)\n org.apache.spark.sql.catalyst.expressions.GetStructField.org$apache$spark$sql$catalyst$expressions$GetStructField$$field$lzycompute(complexTypeExtractors.scala:107)\n org.apache.spark.sql.catalyst.expressions.GetStructField.org$apache$spark$sql$catalyst$expressions$GetStructField$$field(complexTypeExtractors.scala:107)\n org.apache.spark.sql.catalyst.expressions.GetStructField$$anonfun$toString$1.apply(complexTypeExtractors.scala:111)\n org.apache.spark.sql.catalyst.expressions.GetStructField$$anonfun$toString$1.apply(complexTypeExtractors.scala:111)\n org.apache.spark.sql.catalyst.expressions.GetStructField.toString(complexTypeExtractors.scala:111)\n org.apache.spark.sql.catalyst.expressions.Expression.toString(Expression.scala:217)\n org.apache.spark.sql.catalyst.expressions.Expression.toString(Expression.scala:217)\n org.apache.spark.sql.catalyst.expressions.If.toString(conditionalExpressions.scala:76)\n org.apache.spark.sql.catalyst.expressions.Expression.toString(Expression.scala:217)\n org.apache.spark.sql.catalyst.expressions.Alias.toString(namedExpressions.scala:155)\n org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$argString$1.apply(TreeNode.scala:385)\n org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$argString$1.apply(TreeNode.scala:381)\n org.apache.spark.sql.catalyst.trees.TreeNode.argString(TreeNode.scala:388)\n org.apache.spark.sql.catalyst.trees.TreeNode.simpleString(TreeNode.scala:391)\n org.apache.spark.sql.catalyst.plans.QueryPlan.simpleString(QueryPlan.scala:172)\n org.apache.spark.sql.catalyst.trees.TreeNode.generateTreeString(TreeNode.scala:441)\n org.apache.spark.sql.catalyst.trees.TreeNode.treeString(TreeNode.scala:396)\n org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$5.apply(RuleExecutor.scala:118)\n org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$5.apply(RuleExecutor.scala:119)\n org.apache.spark.Logging$class.logDebug(Logging.scala:62)\n org.apache.spark.sql.catalyst.rules.RuleExecutor.logDebug(RuleExecutor.scala:44)\n org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:115)\n org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:72)\n org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:72)\n org.apache.spark.sql.catalyst.encoders.ExpressionEncoder.resolve(ExpressionEncoder.scala:253)\n org.apache.spark.sql.Dataset.(Dataset.scala:78)\n org.apache.spark.sql.Dataset.(Dataset.scala:89)\n org.apache.spark.sql.SQLContext.createDataset(SQLContext.scala:507)\n org.apache.spark.sql.SQLImplicits.localSeqToDatasetHolder(SQLImplicits.scala:80)\n{code}\n\n\nWhen providing a type alias, the code fails in a different way:\n{code}\ntype TwoInt = (Int, Int)\n\ncase class Test(v: TwoInt)\n\nSeq(Test((1,2)), Test((3,4)).toDS().show //fails\n{code}\n\n{code}\n NoSuchElementException: : head of empty list (ScalaReflection.scala:504)\n org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor$1.apply(ScalaReflection.scala:504)\n org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor$1.apply(ScalaReflection.scala:502)\n org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor(ScalaReflection.scala:502)\n org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor$1.apply(ScalaReflection.scala:509)\n org.apache.spark.sql.catalyst.ScalaReflection$$anonfun$org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor$1.apply(ScalaReflection.scala:502)\n org.apache.spark.sql.catalyst.ScalaReflection$.org$apache$spark$sql$catalyst$ScalaReflection$$extractorFor(ScalaReflection.scala:502)\n org.apache.spark.sql.catalyst.ScalaReflection$.extractorsFor(ScalaReflection.scala:394)\n org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$.apply(ExpressionEncoder.scala:54)\n org.apache.spark.sql.SQLImplicits.newProductEncoder(SQLImplicits.scala:41)\n com.intenthq.pipeline.actions.ActionsJobIntegrationSpec.enrich(ActionsJobIntegrationSpec.scala:63)\n com.intenthq.pipeline.actions.ActionsJobIntegrationSpec$$anonfun$is$2.apply(ActionsJobIntegrationSpec.scala:45)\n com.intenthq.pipeline.actions.ActionsJobIntegrationSpec$$anonfun$is$2.apply(ActionsJobIntegrationSpec.scala:45)\n{code}","issue_id":"12929509","key":"SPARK-12777","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-12-15T22:47:07.000+0000","role":"fixed_distractor","summary":"Dataset fields can't be Scala tuples"} {"case_id":"12935207","cluster":"DISTRACTOR-SPARK-13087","comments":[{"body":"On latest build, looks like there is no this problem.","created":"2016-02-01T09:37:09.727+0000"},{"body":"I re-built 1.6.1 with commit ddb9633043e82fb2a34c7e0e29b487f635c3c744 this morning and I'm seeing a similar error to above. Similarly we use a custom UDF and the above commit fixes the issue.\n\n{code:sql}\nSELECT\n concat(t_4.firstname,\" \",t_4.lastname) customer_name,\n agg_cust(t_3.customercountestimate1_c2) ctd_customercountestimate1_ok\nFROM\n as.sales t_3\nJOIN\n as.customer t_4\nON\n t_3.key_c1 = t_4.customerkey\nGROUP BY\n concat(t_4.firstname,\" \",t_4.lastname)\n{code}\n\n{code}\norg.apache.spark.sql.catalyst.errors.package$TreeNodeException: Binding attribute, tree: concat(firstname#329, ,lastname#330)#339\n at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:49)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:86)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:85)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:259)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:259)\n at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n at org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:258)\n at org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:249)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$.bindReference(BoundAttribute.scala:85)\n at org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:39)\n at org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:39)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n at scala.collection.immutable.List.foreach(List.scala:318)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:244)\n at scala.collection.AbstractTraversable.map(Traversable.scala:105)\n at org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.bind(GenerateMutableProjection.scala:39)\n at org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.bind(GenerateMutableProjection.scala:33)\n at org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator.generate(CodeGenerator.scala:585)\n at org.apache.spark.sql.execution.SparkPlan.newMutableProjection(SparkPlan.scala:227)\n at org.apache.spark.sql.execution.Exchange.org$apache$spark$sql$execution$Exchange$$getPartitionKeyExtractor$1(Exchange.scala:197)\n at org.apache.spark.sql.execution.Exchange$$anonfun$3.apply(Exchange.scala:209)\n at org.apache.spark.sql.execution.Exchange$$anonfun$3.apply(Exchange.scala:208)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\nCaused by: java.lang.RuntimeException: Couldn't find concat(firstname#329, ,lastname#330)#339 in [firstname#329,lastname#330,customercountestimate1_c2#326]\n at scala.sys.package$.error(package.scala:27)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:92)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:86)\n at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:48)\n ... 34 more\n{code}","created":"2016-02-01T20:34:03.242+0000"},{"body":"Here's a self-contained test case:\n\n{code}\n test(\"group by function\") {\n Seq((1, 2)).toDF(\"a\", \"b\").registerTempTable(\"data\")\n\n checkAnswer(\n sql(\"SELECT floor(a) AS a, collect_set(b) FROM data GROUP BY floor(a) ORDER BY a\"),\n Row(1, 2) :: Nil)\n }\n{code}\n\nLooks like the problem is specific to the fallback path used when we detect a hive aggregate function.","created":"2016-02-02T00:48:39.485+0000"},{"body":"User 'marmbrus' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11011","created":"2016-02-02T01:14:04.232+0000"},{"body":"User 'marmbrus' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11013","created":"2016-02-02T01:26:04.441+0000"},{"body":"https://github.com/apache/spark/pull/11011 has been merged into master and https://github.com/apache/spark/pull/11013 has been merged into branch 1.6.","created":"2016-02-02T08:51:31.637+0000"}],"conversations":[{"body":"This is a regression from 1.5.\n\nAn example of the failure:\n\nWorking with this table...\n{code}\n0: jdbc:hive2://10.1.3.203:10000> DESCRIBE csd_0ae1abc1_a3af_4c63_95b0_9599faca6c3d;\n+-----------------------+------------+----------+--+\n| col_name | data_type | comment |\n+-----------------------+------------+----------+--+\n| c_date | timestamp | NULL |\n| c_count | int | NULL |\n| c_location_fips_code | string | NULL |\n| c_airtemp | float | NULL |\n| c_dewtemp | float | NULL |\n| c_pressure | int | NULL |\n| c_rain | float | NULL |\n| c_snow | float | NULL |\n+-----------------------+------------+----------+--+\n{code}\n...and this query (which isn't necessarily all that sensical or useful, but has been adapted from a similarly failing query that uses a custom UDF where the Spark SQL built-in `day` function has been substituted into this query)...\n{code}\nSELECT day ( c_date ) AS c_date, percentile_approx(c_rain, 0.5) AS c_expr_1256887735 FROM csd_0ae1abc1_a3af_4c63_95b0_9599faca6c3d GROUP BY day ( c_date ) ORDER BY c_date;\n{code}\nSpark 1.5 produces the expected results without error.\n\nIn Spark 1.6, this plan is produced...\n{code}\nExchange rangepartitioning(c_date#63009 ASC,16), None\n+- SortBasedAggregate(key=[dayofmonth(cast(c_date#63011 as date))#63020], functions=[(hiveudaffunction(HiveFunctionWrapper(org.apache.hadoop.hive.ql.udf.generic.GenericUDAFPercentileApprox,org.apache.hadoop.hive.ql.udf.generic.Gene\nricUDAFPercentileApprox@6f211801),c_rain#63017,0.5,false,0,0),mode=Complete,isDistinct=false)], output=[c_date#63009,c_expr_1256887735#63010])\n +- ConvertToSafe\n +- !Sort [dayofmonth(cast(c_date#63011 as date))#63020 ASC], false, 0\n +- !TungstenExchange hashpartitioning(dayofmonth(cast(c_date#63011 as date))#63020,16), None\n +- ConvertToUnsafe\n +- HiveTableScan [c_date#63011,c_rain#63017], MetastoreRelation default, csd_0ae1abc1_a3af_4c63_95b0_9599faca6c3d, None\n{code}\n...which fails with a TreeNodeException and stack traces that include this...\n{code}\nCaused by: ! org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 2842.0 failed 4 times, most recent failure: Lost task 0.3 in stage 2842.0 (TID 15007, ip-10-1-1-59.dev.clearstory.com): org.apache.spark.sql.catalyst.errors.package$TreeNodeException: Binding attribute, tree: dayofmonth(cast(c_date#63011 as date))#63020\n at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:49)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:86)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:85)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:259)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:259)\n at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n at org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:258)\n at org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:249)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$.bindReference(BoundAttribute.scala:85)\n at org.apache.spark.sql.catalyst.expressions.InterpretedMutableProjection$$anonfun$$init$$2.apply(Projection.scala:62)\n at org.apache.spark.sql.catalyst.expressions.InterpretedMutableProjection$$anonfun$$init$$2.apply(Projection.scala:62)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n at scala.collection.immutable.List.foreach(List.scala:318)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:244)\n at scala.collection.AbstractTraversable.map(Traversable.scala:105)\n at org.apache.spark.sql.catalyst.expressions.InterpretedMutableProjection.(Projection.scala:62)\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$newMutableProjection$1.apply(SparkPlan.scala:254)\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$newMutableProjection$1.apply(SparkPlan.scala:254)\n at org.apache.spark.sql.execution.Exchange.org$apache$spark$sql$execution$Exchange$$getPartitionKeyExtractor$1(Exchange.scala:196)\n at org.apache.spark.sql.execution.Exchange$$anonfun$3.apply(Exchange.scala:208)\n at org.apache.spark.sql.execution.Exchange$$anonfun$3.apply(Exchange.scala:207)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\nCaused by: java.lang.RuntimeException: Couldn't find dayofmonth(cast(c_date#63011 as date))#63020 in [c_date#63011,c_rain#63017]\n at scala.sys.package$.error(package.scala:27)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:92)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:86)\n at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:48)\n ... 33 more\n{code}\n\nIt is possible to work around the problem by adding a Project node in case an aggregation is relying on aliases missing in the child plan (https://github.com/mbautin/spark/commit/2e99064b42a6dddf6b94b989c744a1308aacaee2), but it seems there should be a deeper fix that prevents the problem instead of covering for it.\n\n[~yhuai] I think this problem crept in with the changes for SPARK-9830","from":"reporter","subject":"Grouping by a complex expression may lead to incorrect AttributeReferences in aggregations"},{"body":"On latest build, looks like there is no this problem.","from":"developer"},{"body":"I re-built 1.6.1 with commit ddb9633043e82fb2a34c7e0e29b487f635c3c744 this morning and I'm seeing a similar error to above. Similarly we use a custom UDF and the above commit fixes the issue.\n\n{code:sql}\nSELECT\n concat(t_4.firstname,\" \",t_4.lastname) customer_name,\n agg_cust(t_3.customercountestimate1_c2) ctd_customercountestimate1_ok\nFROM\n as.sales t_3\nJOIN\n as.customer t_4\nON\n t_3.key_c1 = t_4.customerkey\nGROUP BY\n concat(t_4.firstname,\" \",t_4.lastname)\n{code}\n\n{code}\norg.apache.spark.sql.catalyst.errors.package$TreeNodeException: Binding attribute, tree: concat(firstname#329, ,lastname#330)#339\n at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:49)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:86)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:85)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:259)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:259)\n at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n at org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:258)\n at org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:249)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$.bindReference(BoundAttribute.scala:85)\n at org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:39)\n at org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:39)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n at scala.collection.immutable.List.foreach(List.scala:318)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:244)\n at scala.collection.AbstractTraversable.map(Traversable.scala:105)\n at org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.bind(GenerateMutableProjection.scala:39)\n at org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.bind(GenerateMutableProjection.scala:33)\n at org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator.generate(CodeGenerator.scala:585)\n at org.apache.spark.sql.execution.SparkPlan.newMutableProjection(SparkPlan.scala:227)\n at org.apache.spark.sql.execution.Exchange.org$apache$spark$sql$execution$Exchange$$getPartitionKeyExtractor$1(Exchange.scala:197)\n at org.apache.spark.sql.execution.Exchange$$anonfun$3.apply(Exchange.scala:209)\n at org.apache.spark.sql.execution.Exchange$$anonfun$3.apply(Exchange.scala:208)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\nCaused by: java.lang.RuntimeException: Couldn't find concat(firstname#329, ,lastname#330)#339 in [firstname#329,lastname#330,customercountestimate1_c2#326]\n at scala.sys.package$.error(package.scala:27)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:92)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:86)\n at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:48)\n ... 34 more\n{code}","from":"developer"},{"body":"Here's a self-contained test case:\n\n{code}\n test(\"group by function\") {\n Seq((1, 2)).toDF(\"a\", \"b\").registerTempTable(\"data\")\n\n checkAnswer(\n sql(\"SELECT floor(a) AS a, collect_set(b) FROM data GROUP BY floor(a) ORDER BY a\"),\n Row(1, 2) :: Nil)\n }\n{code}\n\nLooks like the problem is specific to the fallback path used when we detect a hive aggregate function.","from":"developer"},{"body":"User 'marmbrus' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11011","from":"developer"},{"body":"User 'marmbrus' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11013","from":"developer"},{"body":"https://github.com/apache/spark/pull/11011 has been merged into master and https://github.com/apache/spark/pull/11013 has been merged into branch 1.6.","from":"developer"}],"created":"2016-01-29T18:30:01.000+0000","description":"This is a regression from 1.5.\n\nAn example of the failure:\n\nWorking with this table...\n{code}\n0: jdbc:hive2://10.1.3.203:10000> DESCRIBE csd_0ae1abc1_a3af_4c63_95b0_9599faca6c3d;\n+-----------------------+------------+----------+--+\n| col_name | data_type | comment |\n+-----------------------+------------+----------+--+\n| c_date | timestamp | NULL |\n| c_count | int | NULL |\n| c_location_fips_code | string | NULL |\n| c_airtemp | float | NULL |\n| c_dewtemp | float | NULL |\n| c_pressure | int | NULL |\n| c_rain | float | NULL |\n| c_snow | float | NULL |\n+-----------------------+------------+----------+--+\n{code}\n...and this query (which isn't necessarily all that sensical or useful, but has been adapted from a similarly failing query that uses a custom UDF where the Spark SQL built-in `day` function has been substituted into this query)...\n{code}\nSELECT day ( c_date ) AS c_date, percentile_approx(c_rain, 0.5) AS c_expr_1256887735 FROM csd_0ae1abc1_a3af_4c63_95b0_9599faca6c3d GROUP BY day ( c_date ) ORDER BY c_date;\n{code}\nSpark 1.5 produces the expected results without error.\n\nIn Spark 1.6, this plan is produced...\n{code}\nExchange rangepartitioning(c_date#63009 ASC,16), None\n+- SortBasedAggregate(key=[dayofmonth(cast(c_date#63011 as date))#63020], functions=[(hiveudaffunction(HiveFunctionWrapper(org.apache.hadoop.hive.ql.udf.generic.GenericUDAFPercentileApprox,org.apache.hadoop.hive.ql.udf.generic.Gene\nricUDAFPercentileApprox@6f211801),c_rain#63017,0.5,false,0,0),mode=Complete,isDistinct=false)], output=[c_date#63009,c_expr_1256887735#63010])\n +- ConvertToSafe\n +- !Sort [dayofmonth(cast(c_date#63011 as date))#63020 ASC], false, 0\n +- !TungstenExchange hashpartitioning(dayofmonth(cast(c_date#63011 as date))#63020,16), None\n +- ConvertToUnsafe\n +- HiveTableScan [c_date#63011,c_rain#63017], MetastoreRelation default, csd_0ae1abc1_a3af_4c63_95b0_9599faca6c3d, None\n{code}\n...which fails with a TreeNodeException and stack traces that include this...\n{code}\nCaused by: ! org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 2842.0 failed 4 times, most recent failure: Lost task 0.3 in stage 2842.0 (TID 15007, ip-10-1-1-59.dev.clearstory.com): org.apache.spark.sql.catalyst.errors.package$TreeNodeException: Binding attribute, tree: dayofmonth(cast(c_date#63011 as date))#63020\n at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:49)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:86)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:85)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:259)\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:259)\n at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n at org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:258)\n at org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:249)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$.bindReference(BoundAttribute.scala:85)\n at org.apache.spark.sql.catalyst.expressions.InterpretedMutableProjection$$anonfun$$init$$2.apply(Projection.scala:62)\n at org.apache.spark.sql.catalyst.expressions.InterpretedMutableProjection$$anonfun$$init$$2.apply(Projection.scala:62)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n at scala.collection.immutable.List.foreach(List.scala:318)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:244)\n at scala.collection.AbstractTraversable.map(Traversable.scala:105)\n at org.apache.spark.sql.catalyst.expressions.InterpretedMutableProjection.(Projection.scala:62)\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$newMutableProjection$1.apply(SparkPlan.scala:254)\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$newMutableProjection$1.apply(SparkPlan.scala:254)\n at org.apache.spark.sql.execution.Exchange.org$apache$spark$sql$execution$Exchange$$getPartitionKeyExtractor$1(Exchange.scala:196)\n at org.apache.spark.sql.execution.Exchange$$anonfun$3.apply(Exchange.scala:208)\n at org.apache.spark.sql.execution.Exchange$$anonfun$3.apply(Exchange.scala:207)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$21.apply(RDD.scala:728)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\nCaused by: java.lang.RuntimeException: Couldn't find dayofmonth(cast(c_date#63011 as date))#63020 in [c_date#63011,c_rain#63017]\n at scala.sys.package$.error(package.scala:27)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:92)\n at org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:86)\n at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:48)\n ... 33 more\n{code}\n\nIt is possible to work around the problem by adding a Project node in case an aggregation is relying on aliases missing in the child plan (https://github.com/mbautin/spark/commit/2e99064b42a6dddf6b94b989c744a1308aacaee2), but it seems there should be a deeper fix that prevents the problem instead of covering for it.\n\n[~yhuai] I think this problem crept in with the changes for SPARK-9830","issue_id":"12935207","key":"SPARK-13087","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-02-02T08:52:03.000+0000","role":"fixed_distractor","summary":"Grouping by a complex expression may lead to incorrect AttributeReferences in aggregations"} {"case_id":"12940995","cluster":"DISTRACTOR-SPARK-13431","comments":[{"body":"The good/bad news is that I can reproduce this locally.\n\nIt seems like we are having an ANTLR issue again. We merged 2 grammar changing pull requests in the past two days:\n- [SPARK-13321]: https://github.com/apache/spark/commit/55d6fdf22d1d6379180ac09f364c38982897d9ff\n- [SPARK-13306]: https://github.com/apache/spark/commit/7925071280bfa1570435bde3e93492eaf2167d56\n\nAre you sure only the last one caused this?\n\nWhat is a bit funny is that code too large errors typically pop-up during compilation. Does the shade plugin has its own size limit?","created":"2016-02-22T13:15:29.657+0000"},{"body":"In my experience the 'shade' plugin does a lot more than it says (constant propagation, control-flow optimizations). I didn't find a way to turn them off.\n\n... and I don't know if it's only the last one. The first Jenkins build that failed had only this commit in the list of changes.","created":"2016-02-22T13:28:07.897+0000"},{"body":"Shade doesn't do optimization but it certainly does several transformations. It's possible that something like renaming a package name to a longer one puts some method just over the top.\n\nThe problem isn't shading but huge methods. These should simply be broken down further.","created":"2016-02-22T13:47:59.099+0000"},{"body":"Well, the class file format isn't embedding any names inside the code section, that's what the ConstantPool is for. What the shade plugin does or not is slightly irrelevant, the fact is that the failure happens (only) during shading, and all \"compile\" integration builds in Jenkins are passing without problem. Only the ones that run tests (and run `assembly`) fail.","created":"2016-02-22T14:24:44.916+0000"},{"body":"Hm I think you're right about that. I also don't see why it would pass compilation. Is it possible scalac doesn't enforce the limit but shade does? or is that out of the question since the code would actually fail to load?\n\nThe answer still might be to break up the method.\n\nI did update the shade plugin about 5 days ago but it passed since.\nhttps://github.com/apache/spark/commit/b84404865b90fc07858a8b28c2a7028f96517167","created":"2016-02-22T15:03:13.988+0000"},{"body":"According to AMPLAB's Jenkins, the first failure happened at following build that includes [https://issues.apache.org/jira/browse/SPARK-13306]\nhttps://amplab.cs.berkeley.edu/jenkins/job/spark-master-test-maven-hadoop-2.4/337/","created":"2016-02-22T15:07:05.159+0000"},{"body":"Probably the easiest fix is to break some grammar elements in more files. I *think* the problem is deeper though, and there's some state explosion generated by an ambiguous part of the grammar.\n\nThe Scala compiler would show an error, but the ANTLR generator spits out Java files. Definitely those methods are very close to the limit, I saw javac failing when using the Eclipse compiler, but the Oracle one was slightly better.","created":"2016-02-22T15:20:46.293+0000"},{"body":" The Java Virtual Machine specification limits the size of generated Java byte code for each method in a class to the maximum of 64K bytes. so i think its not something you can avoid at shade or scala compiler level.\nSome highly relevant links: https://groups.google.com/forum/#!topic/scala-internals/f6PwUxc8K7I\nhttp://www.cubrid.org/blog/dev-platform/understanding-jvm-internals/","created":"2016-02-22T15:29:03.234+0000"},{"body":"I succeeded to build Spark by commenting out lines 208 and 209 in ExpressionParser.g\nhttps://github.com/apache/spark/blob/7925071280bfa1570435bde3e93492eaf2167d56/sql/catalyst/src/main/antlr3/org/apache/spark/sql/catalyst/parser/ExpressionParser.g#L208\n\nThese two lines are added by [https://issues.apache.org/jira/browse/SPARK-13306].","created":"2016-02-22T15:59:06.121+0000"},{"body":"Test are currently passing and the if you call {{sparkShell}} from sbt you can work with spark and issue sql commands. So the direct problem is definitely caused by shading. The thing I am actually not getting is why sbt assemblies won't fail.\n\nThat being said, ANTLR spits out huge files which are very near to but not over the 64K limit (It won't build otherwise). I am not too confident that splitting the parser another time is much more than a temporary solution. The current grammar contains a perticulary nasty rule to determine if a keyword is non-reserved (this creates a huge state explosion), we should take a look at that. On the longer run I think that we should do a rewrite of the current parser with ANTLR4 which should simplify parser rules (and thus the size of the resulting parser), and make the construction of logical plans less Rube-Goldbergy/pattern-match-heaven.\n\n","created":"2016-02-22T16:07:06.596+0000"},{"body":"[~kiszk] Are you building with Maven or SBT?","created":"2016-02-22T16:09:30.531+0000"},{"body":"I am using mvn, and executed the following command\n\n{code}\nbuild/mvn -Pyarn -Phadoop-2.4 -DskipTests package\n{code}\n","created":"2016-02-22T16:11:48.304+0000"},{"body":"Sbt assembly is using a different library, I think it's called jarjar. Nothing to do with the maven shade plugin.","created":"2016-02-22T16:14:16.026+0000"},{"body":"The size of a static initializer method of SparkSqlParser_ExpressionParser.java is 59901. It seem to be the largest method among generated Java files. It is still less than 64K bytes.","created":"2016-02-22T17:10:43.120+0000"},{"body":"I meet the same problem with maven, how to resolve this issue?\n\n*build command as follow:*\nbq. mvn clean package -DskipTests -Pscala-2.11 -Phadoop-2.6 -Phive -Phive-thriftserver -Pyarn -Dyarn.version=2.6.0-cdh5.4.7 -Dhadoop.version=2.6.0-cdh5.4.7\n\n*build error message:*\n{quote}\n[INFO] Excluding org.scala-lang:scala-reflect:jar:2.11.7 from the shaded jar.\n[INFO] Excluding org.scala-lang:scala-library:jar:2.11.7 from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-core_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.apache.avro:avro-mapred:jar:hadoop2:1.7.7 from the shaded jar.\n[INFO] Excluding org.apache.avro:avro-ipc:jar:1.7.7 from the shaded jar.\n[INFO] Excluding org.apache.avro:avro:jar:1.7.7 from the shaded jar.\n[INFO] Excluding org.apache.avro:avro-ipc:jar:tests:1.7.7 from the shaded jar.\n[INFO] Excluding org.codehaus.jackson:jackson-core-asl:jar:1.9.13 from the shaded jar.\n[INFO] Excluding org.codehaus.jackson:jackson-mapper-asl:jar:1.9.13 from the shaded jar.\n[INFO] Excluding com.twitter:chill_2.11:jar:0.5.0 from the shaded jar.\n[INFO] Excluding com.esotericsoftware.kryo:kryo:jar:2.21 from the shaded jar.\n[INFO] Excluding com.esotericsoftware.reflectasm:reflectasm:jar:shaded:1.07 from the shaded jar.\n[INFO] Excluding com.esotericsoftware.minlog:minlog:jar:1.2 from the shaded jar.\n[INFO] Excluding org.objenesis:objenesis:jar:1.2 from the shaded jar.\n[INFO] Excluding com.twitter:chill-java:jar:0.5.0 from the shaded jar.\n[INFO] Excluding org.apache.xbean:xbean-asm5-shaded:jar:4.4 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-client:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-common:jar:2.2.0 from the shaded jar.\n[INFO] Excluding commons-cli:commons-cli:jar:1.2 from the shaded jar.\n[INFO] Excluding org.apache.commons:commons-math:jar:2.1 from the shaded jar.\n[INFO] Excluding xmlenc:xmlenc:jar:0.52 from the shaded jar.\n[INFO] Excluding commons-io:commons-io:jar:2.1 from the shaded jar.\n[INFO] Excluding commons-lang:commons-lang:jar:2.6 from the shaded jar.\n[INFO] Excluding commons-configuration:commons-configuration:jar:1.6 from the shaded jar.\n[INFO] Excluding commons-collections:commons-collections:jar:3.2.2 from the shaded jar.\n[INFO] Excluding commons-digester:commons-digester:jar:1.8 from the shaded jar.\n[INFO] Excluding commons-beanutils:commons-beanutils:jar:1.7.0 from the shaded jar.\n[INFO] Excluding commons-beanutils:commons-beanutils-core:jar:1.8.0 from the shaded jar.\n[INFO] Excluding com.google.protobuf:protobuf-java:jar:2.5.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-auth:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.commons:commons-compress:jar:1.4.1 from the shaded jar.\n[INFO] Excluding org.tukaani:xz:jar:1.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-hdfs:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.mortbay.jetty:jetty-util:jar:6.1.26 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-mapreduce-client-app:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-mapreduce-client-common:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-yarn-client:jar:2.2.0 from the shaded jar.\n[INFO] Excluding com.google.inject:guice:jar:3.0 from the shaded jar.\n[INFO] Excluding javax.inject:javax.inject:jar:1 from the shaded jar.\n[INFO] Excluding aopalliance:aopalliance:jar:1.0 from the shaded jar.\n[INFO] Excluding com.sun.jersey.jersey-test-framework:jersey-test-framework-grizzly2:jar:1.9 from the shaded jar.\n[INFO] Excluding com.sun.jersey.jersey-test-framework:jersey-test-framework-core:jar:1.9 from the shaded jar.\n[INFO] Excluding javax.servlet:javax.servlet-api:jar:3.0.1 from the shaded jar.\n[INFO] Excluding com.sun.jersey:jersey-client:jar:1.9 from the shaded jar.\n[INFO] Excluding com.sun.jersey:jersey-grizzly2:jar:1.9 from the shaded jar.\n[INFO] Excluding org.glassfish.grizzly:grizzly-http:jar:2.1.2 from the shaded jar.\n[INFO] Excluding org.glassfish.grizzly:grizzly-framework:jar:2.1.2 from the shaded jar.\n[INFO] Excluding org.glassfish.gmbal:gmbal-api-only:jar:3.0.0-b023 from the shaded jar.\n[INFO] Excluding org.glassfish.external:management-api:jar:3.0.0-b012 from the shaded jar.\n[INFO] Excluding org.glassfish.grizzly:grizzly-http-server:jar:2.1.2 from the shaded jar.\n[INFO] Excluding org.glassfish.grizzly:grizzly-rcm:jar:2.1.2 from the shaded jar.\n[INFO] Excluding org.glassfish.grizzly:grizzly-http-servlet:jar:2.1.2 from the shaded jar.\n[INFO] Excluding org.glassfish:javax.servlet:jar:3.1 from the shaded jar.\n[INFO] Excluding com.sun.jersey:jersey-json:jar:1.9 from the shaded jar.\n[INFO] Excluding org.codehaus.jettison:jettison:jar:1.1 from the shaded jar.\n[INFO] Excluding com.sun.xml.bind:jaxb-impl:jar:2.2.3-1 from the shaded jar.\n[INFO] Excluding javax.xml.bind:jaxb-api:jar:2.2.2 from the shaded jar.\n[INFO] Excluding javax.activation:activation:jar:1.1 from the shaded jar.\n[INFO] Excluding org.codehaus.jackson:jackson-jaxrs:jar:1.9.13 from the shaded jar.\n[INFO] Excluding org.codehaus.jackson:jackson-xc:jar:1.9.13 from the shaded jar.\n[INFO] Excluding com.sun.jersey.contribs:jersey-guice:jar:1.9 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-yarn-server-common:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-mapreduce-client-shuffle:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-yarn-api:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-mapreduce-client-core:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-yarn-common:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-mapreduce-client-jobclient:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-annotations:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-launcher_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-network-common_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-network-shuffle_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.fusesource.leveldbjni:leveldbjni-all:jar:1.8 from the shaded jar.\n[INFO] Excluding com.fasterxml.jackson.core:jackson-annotations:jar:2.5.3 from the shaded jar.\n[INFO] Excluding net.java.dev.jets3t:jets3t:jar:0.7.1 from the shaded jar.\n[INFO] Excluding commons-httpclient:commons-httpclient:jar:3.1 from the shaded jar.\n[INFO] Excluding org.apache.curator:curator-recipes:jar:2.4.0 from the shaded jar.\n[INFO] Excluding org.apache.curator:curator-framework:jar:2.4.0 from the shaded jar.\n[INFO] Excluding org.apache.curator:curator-client:jar:2.4.0 from the shaded jar.\n[INFO] Excluding org.apache.zookeeper:zookeeper:jar:3.4.5 from the shaded jar.\n[INFO] Excluding jline:jline:jar:0.9.94 from the shaded jar.\n[INFO] Excluding org.eclipse.jetty.orbit:javax.servlet:jar:3.0.0.v201112011016 from the shaded jar.\n[INFO] Excluding org.apache.commons:commons-lang3:jar:3.3.2 from the shaded jar.\n[INFO] Excluding org.apache.commons:commons-math3:jar:3.4.1 from the shaded jar.\n[INFO] Excluding com.google.code.findbugs:jsr305:jar:1.3.9 from the shaded jar.\n[INFO] Excluding org.slf4j:slf4j-api:jar:1.7.16 from the shaded jar.\n[INFO] Excluding org.slf4j:jul-to-slf4j:jar:1.7.16 from the shaded jar.\n[INFO] Excluding org.slf4j:jcl-over-slf4j:jar:1.7.16 from the shaded jar.\n[INFO] Excluding log4j:log4j:jar:1.2.17 from the shaded jar.\n[INFO] Excluding org.slf4j:slf4j-log4j12:jar:1.7.16 from the shaded jar.\n[INFO] Excluding com.ning:compress-lzf:jar:1.0.3 from the shaded jar.\n[INFO] Excluding org.xerial.snappy:snappy-java:jar:1.1.2 from the shaded jar.\n[INFO] Excluding net.jpountz.lz4:lz4:jar:1.3.0 from the shaded jar.\n[INFO] Excluding org.roaringbitmap:RoaringBitmap:jar:0.5.11 from the shaded jar.\n[INFO] Excluding commons-net:commons-net:jar:2.2 from the shaded jar.\n[INFO] Excluding org.json4s:json4s-jackson_2.11:jar:3.2.10 from the shaded jar.\n[INFO] Excluding org.json4s:json4s-core_2.11:jar:3.2.10 from the shaded jar.\n[INFO] Excluding org.json4s:json4s-ast_2.11:jar:3.2.10 from the shaded jar.\n[INFO] Excluding org.scala-lang:scalap:jar:2.11.7 from the shaded jar.\n[INFO] Excluding org.scala-lang:scala-compiler:jar:2.11.7 from the shaded jar.\n[INFO] Excluding org.scala-lang.modules:scala-parser-combinators_2.11:jar:1.0.4 from the shaded jar.\n[INFO] Excluding com.sun.jersey:jersey-server:jar:1.9 from the shaded jar.\n[INFO] Excluding asm:asm:jar:3.1 from the shaded jar.\n[INFO] Excluding com.sun.jersey:jersey-core:jar:1.9 from the shaded jar.\n[INFO] Excluding org.apache.mesos:mesos:jar:shaded-protobuf:0.21.1 from the shaded jar.\n[INFO] Excluding io.netty:netty-all:jar:4.0.29.Final from the shaded jar.\n[INFO] Excluding io.netty:netty:jar:3.8.0.Final from the shaded jar.\n[INFO] Excluding com.clearspring.analytics:stream:jar:2.7.0 from the shaded jar.\n[INFO] Excluding io.dropwizard.metrics:metrics-core:jar:3.1.2 from the shaded jar.\n[INFO] Excluding io.dropwizard.metrics:metrics-jvm:jar:3.1.2 from the shaded jar.\n[INFO] Excluding io.dropwizard.metrics:metrics-json:jar:3.1.2 from the shaded jar.\n[INFO] Excluding io.dropwizard.metrics:metrics-graphite:jar:3.1.2 from the shaded jar.\n[INFO] Excluding com.fasterxml.jackson.core:jackson-databind:jar:2.5.3 from the shaded jar.\n[INFO] Excluding com.fasterxml.jackson.core:jackson-core:jar:2.5.3 from the shaded jar.\n[INFO] Excluding com.fasterxml.jackson.module:jackson-module-scala_2.11:jar:2.5.3 from the shaded jar.\n[INFO] Excluding com.thoughtworks.paranamer:paranamer:jar:2.6 from the shaded jar.\n[INFO] Excluding org.apache.ivy:ivy:jar:2.4.0 from the shaded jar.\n[INFO] Excluding oro:oro:jar:2.0.8 from the shaded jar.\n[INFO] Excluding net.razorvine:pyrolite:jar:4.9 from the shaded jar.\n[INFO] Excluding net.sf.py4j:py4j:jar:0.9.1 from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-unsafe_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.codehaus.janino:janino:jar:2.7.8 from the shaded jar.\n[INFO] Excluding org.codehaus.janino:commons-compiler:jar:2.7.8 from the shaded jar.\n[INFO] Excluding org.antlr:antlr-runtime:jar:3.5.2 from the shaded jar.\n[INFO] Excluding commons-codec:commons-codec:jar:1.10 from the shaded jar.\n[INFO] Including org.spark-project.spark:unused:jar:1.0.0 in the shaded jar.\n[INFO] Excluding org.scala-lang.modules:scala-xml_2.11:jar:1.0.2 from the shaded jar.\n[DEBUG] Processing JAR /home/myubuntu/Works/workspace/spark/sql/catalyst/target/spark-catalyst_2.11-2.0.0-SNAPSHOT.jar\n[INFO] ------------------------------------------------------------------------\n[INFO] BUILD FAILURE\n[INFO] ------------------------------------------------------------------------\n[INFO] Total time: 52.733 s (Wall Clock)\n[INFO] Finished at: 2016-02-23T13:26:26+08:00\n[INFO] Final Memory: 38M/744M\n[INFO] ------------------------------------------------------------------------\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-shade-plugin:2.4.3:shade (default) on project spark-catalyst_2.11: Error creating shaded jar: Method code too large! -> [Help 1]\norg.apache.maven.lifecycle.LifecycleExecutionException: Failed to execute goal org.apache.maven.plugins:maven-shade-plugin:2.4.3:shade (default) on project spark-catalyst_2.11: Error creating shaded jar: Method code too large!\n\tat org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:212)\n\tat org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:153)\n\tat org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:145)\n\tat org.apache.maven.lifecycle.internal.LifecycleModuleBuilder.buildProject(LifecycleModuleBuilder.java:116)\n\tat org.apache.maven.lifecycle.internal.builder.multithreaded.MultiThreadedBuilder$1.call(MultiThreadedBuilder.java:185)\n\tat org.apache.maven.lifecycle.internal.builder.multithreaded.MultiThreadedBuilder$1.call(MultiThreadedBuilder.java:181)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.maven.plugin.MojoExecutionException: Error creating shaded jar: Method code too large!\n\tat org.apache.maven.plugins.shade.mojo.ShadeMojo.execute(ShadeMojo.java:540)\n\tat org.apache.maven.plugin.DefaultBuildPluginManager.executeMojo(DefaultBuildPluginManager.java:134)\n\tat org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:207)\n\t... 11 more\nCaused by: java.lang.RuntimeException: Method code too large!\n\tat org.objectweb.asm.MethodWriter.a(Unknown Source)\n\tat org.objectweb.asm.ClassWriter.toByteArray(Unknown Source)\n\tat org.apache.maven.plugins.shade.DefaultShader.addRemappedClass(DefaultShader.java:453)\n\tat org.apache.maven.plugins.shade.DefaultShader.shadeSingleJar(DefaultShader.java:219)\n\tat org.apache.maven.plugins.shade.DefaultShader.shadeJars(DefaultShader.java:179)\n\tat org.apache.maven.plugins.shade.DefaultShader.shade(DefaultShader.java:104)\n\tat org.apache.maven.plugins.shade.mojo.ShadeMojo.execute(ShadeMojo.java:454)\n\t... 13 more\n{quote}","created":"2016-02-23T07:09:53.358+0000"},{"body":"Try build/sbt before the community finalizes the solution. For example, \n{code}\nbuild/sbt -Pyarn -Phadoop-2.6 -Phive -Phive-thriftserver assembly\n{code}","created":"2016-02-23T07:13:18.653+0000"},{"body":"I can't build it with sbt because downloading jar from sbt repository is very slow, even failed.","created":"2016-02-23T07:36:45.128+0000"},{"body":"So there is no other option for making a distro? ","created":"2016-02-23T15:26:10.471+0000"},{"body":"I confirmed that this workaround also enables the make-distribution.sh script to complete properly.","created":"2016-02-23T15:58:17.421+0000"},{"body":"I meet to the same problem, using the repo cloned from https://github.com/apache/spark.git.","created":"2016-02-23T16:12:56.061+0000"},{"body":"I experiencing the same problem when checking out 1.6.0 (git clone git@github.com:apache/spark.git v1.6.0).\n....\n[INFO] Excluding org.apache.ivy:ivy:jar:2.4.0 from the shaded jar.\n[INFO] Excluding oro:oro:jar:2.0.8 from the shaded jar.\n[INFO] Excluding net.razorvine:pyrolite:jar:4.9 from the shaded jar.\n[INFO] Excluding net.sf.py4j:py4j:jar:0.9.1 from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-unsafe_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.codehaus.janino:janino:jar:2.7.8 from the shaded jar.\n[INFO] Excluding org.codehaus.janino:commons-compiler:jar:2.7.8 from the shaded jar.\n[INFO] Excluding org.antlr:antlr-runtime:jar:3.5.2 from the shaded jar.\n[INFO] Excluding commons-codec:commons-codec:jar:1.10 from the shaded jar.\n[INFO] Including org.spark-project.spark:unused:jar:1.0.0 in the shaded jar.\n[INFO] Excluding org.scala-lang.modules:scala-xml_2.11:jar:1.0.2 from the shaded jar.\n[INFO] ------------------------------------------------------------------------\n[INFO] Reactor Summary:\n[INFO] \n[INFO] Spark Project Parent POM ........................... SUCCESS [ 7.063 s]\n[INFO] Spark Project Sketch ............................... SUCCESS [ 7.766 s]\n[INFO] Spark Project Test Tags ............................ SUCCESS [ 3.325 s]\n[INFO] Spark Project Core ................................. SUCCESS [03:13 min]\n[INFO] Spark Project GraphX ............................... SUCCESS [ 23.761 s]\n[INFO] Spark Project ML Library ........................... SUCCESS [01:52 min]\n[INFO] Spark Project Tools ................................ SUCCESS [ 4.876 s]\n[INFO] Spark Project Networking ........................... SUCCESS [ 16.129 s]\n[INFO] Spark Project Shuffle Streaming Service ............ SUCCESS [ 13.367 s]\n[INFO] Spark Project Streaming ............................ SUCCESS [ 53.482 s]\n[INFO] Spark Project Catalyst ............................. FAILURE [02:15 min]\n[INFO] Spark Project SQL .................................. SKIPPED\n[INFO] Spark Project Hive ................................. SKIPPED\n[INFO] Spark Project Docker Integration Tests ............. SKIPPED\n[INFO] Spark Project Unsafe ............................... SKIPPED\n[INFO] Spark Project Assembly ............................. SKIPPED\n[INFO] Spark Project External Twitter ..................... SKIPPED\n[INFO] Spark Project External Flume ....................... SKIPPED\n[INFO] Spark Project External Flume Sink .................. SKIPPED\n[INFO] Spark Project External Flume Assembly .............. SKIPPED\n[INFO] Spark Project External Akka ........................ SKIPPED\n[INFO] Spark Project External MQTT ........................ SKIPPED\n[INFO] Spark Project External MQTT Assembly ............... SKIPPED\n[INFO] Spark Project External ZeroMQ ...................... SKIPPED\n[INFO] Spark Project Examples ............................. SKIPPED\n[INFO] Spark Project REPL ................................. SKIPPED\n[INFO] Spark Project Launcher ............................. SKIPPED\n[INFO] Spark Project External Kafka ....................... SKIPPED\n[INFO] Spark Project External Kafka Assembly .............. SKIPPED\n[INFO] Spark Project YARN ................................. SKIPPED\n[INFO] Spark Project YARN Shuffle Service ................. SKIPPED\n[INFO] ------------------------------------------------------------------------\n[INFO] BUILD FAILURE\n[INFO] ------------------------------------------------------------------------\n[INFO] Total time: 09:34 min\n[INFO] Finished at: 2016-02-23T17:04:28+01:00\n[INFO] Final Memory: 65M/1297M\n[INFO] ------------------------------------------------------------------------\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-shade-plugin:2.4.3:shade (default) on project spark-catalyst_2.10: Error creating shaded jar: Method code too large! -> [Help 1]\n[ERROR] \n[ERROR] To see the full stack trace of the errors, re-run Maven with the -e switch.\n[ERROR] Re-run Maven using the -X switch to enable full debug logging.\n[ERROR] \n[ERROR] For more information about the errors and possible solutions, please read the following articles:\n[ERROR] [Help 1] http://cwiki.apache.org/confluence/display/MAVEN/MojoExecutionException","created":"2016-02-23T16:14:24.106+0000"},{"body":"I think the only way is to revert the offending changes, as mentioned [here|https://issues.apache.org/jira/browse/SPARK-13431?focusedCommentId=15157182&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15157182]","created":"2016-02-23T16:14:49.524+0000"},{"body":"Yes, nobody's suggesting this isn't clearly reproducible. No need to post this.","created":"2016-02-23T16:17:38.281+0000"},{"body":"Sorry for the spam. I only thought I report it as the bug is only filed as affecting v2.0.0 (and I am using v1.6.0)","created":"2016-02-23T16:24:34.681+0000"},{"body":"I identified why this problem occurs only in maven. Shade plugin for maven increases the length of Java bytecode for a method. This increasing happens since shade plugin rewrites Java bytecode to rebuild constant pool.\n\nHere is an output of ``javap -c SparkSqlParser_ExpressionParser.class`` before applying shade plugin. The static initializer ``static{}`` uses ``ldc`` bytecode for accessing constant pool at offset 13, 18, and 23. Each ``ldc`` consume only two bytes. As a result, the bytecode length of this method is *less than 65536*.\n{code}\npublic class org.apache.spark.sql.catalyst.parser.SparkSqlParser_ExpressionParser extends org.antlr.runtime.Parser {\n...\n static {};\n Code:\n 0: bipush 70\n 2: anewarray #1035 // class java/lang/String\n 5: dup\n 6: iconst_0\n 7: ldc_w #1036 // String ...\n 10: aastore\n 11: dup\n 12: iconst_1\n 13: ldc #127 // String\n 15: aastore\n 16: dup\n 17: iconst_2\n 18: ldc #127 // String\n 20: aastore\n 21: dup\n 22: iconst_3\n 23: ldc #127 // String\n 25: aastore\n ...\n 59900: return\n }\n}\n{code}\n\nAfter applying shade plugin, the static initializer ``static{}`` uses ``ldc_w`` bytecode for accessing constant pool at offset 13, 19, and 25. Each ``ldc_w`` consumes three bytes. As a result, the bytecode length of this method is *more than 65535*.\n{code}\n static {};\n Code:\n 0: bipush 70\n 2: anewarray #2965 // class java/lang/String\n 5: dup\n 6: iconst_0\n 7: ldc_w #5240 // String ...\n 10: aastore\n 11: dup\n 12: iconst_1\n 13: ldc_w #2924 // String\n 16: aastore\n 17: dup\n 18: iconst_2\n 19: ldc_w #2924 // String\n 22: aastore\n 23: dup\n 24: iconst_3\n 25: ldc_w #2924 // String\n 28: aastore\n ...\n 65533: lconst_0\n 65534: lastore\n ...\n\n }\n}\n{code}\n\nShading plugin seems to rebuild constant pool based on [this comment|http://svn.apache.org/viewvc/maven/plugins/tags/maven-shade-plugin-2.4.3/src/main/java/org/apache/maven/plugins/shade/DefaultShader.java?view=markup#l417]. To use a lot of constant pool entry due to many definitions of String may increase the entry index of the constant pool. As a result, it leads to replace ``ldc`` with ``ldc_w``. Finally, the length of Java bytecode is increased.\n\nAs a next step, what will we do?\n* Can we avoid this rebuild by an option?\n* Can we create a pull request for shade plugin to avoid this?\n* Can we use another plugin?\n* Can we split ExpressionParser.g into smaller files?\n* Other solutions?\n\n\n","created":"2016-02-23T21:07:51.737+0000"},{"body":"Submitted PR for revert: https://github.com/apache/spark/pull/11329","created":"2016-02-23T21:16:40.820+0000"},{"body":"I'd like to split ExpressionParser.g, or we can't touch it anymore (may break sbt break next time), unless we can switch to ANTR4 soon.","created":"2016-02-23T21:24:06.520+0000"},{"body":"https://github.com/apache/spark/pull/11331","created":"2016-02-23T21:33:45.969+0000"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11331","created":"2016-02-23T21:34:03.211+0000"},{"body":"cool!","created":"2016-02-23T23:39:16.614+0000"},{"body":"Issue resolved by pull request 11331\n[https://github.com/apache/spark/pull/11331]","created":"2016-02-24T05:22:54.006+0000"}],"conversations":[{"body":"Cannot build the project when run the normal build commands:\neg.\n{code}\nbuild/mvn -Phadoop-2.6 -Dhadoop.version=2.6.0 clean package\n./make-distribution.sh --name test --tgz -Phadoop-2.6 \n{code}\n\nIntegration builds are also failing: \n\nhttps://amplab.cs.berkeley.edu/jenkins/job/spark-master-test-maven-hadoop-2.6/229/console\nhttps://ci.typesafe.com/job/mit-docker-test-zk-ref/12/console\n\nIt looks like this is the commit that introduced the issue:\n\nhttps://github.com/apache/spark/commit/7925071280bfa1570435bde3e93492eaf2167d56","from":"reporter","subject":"Maven build fails due to: Method code too large! in Catalyst"},{"body":"The good/bad news is that I can reproduce this locally.\n\nIt seems like we are having an ANTLR issue again. We merged 2 grammar changing pull requests in the past two days:\n- [SPARK-13321]: https://github.com/apache/spark/commit/55d6fdf22d1d6379180ac09f364c38982897d9ff\n- [SPARK-13306]: https://github.com/apache/spark/commit/7925071280bfa1570435bde3e93492eaf2167d56\n\nAre you sure only the last one caused this?\n\nWhat is a bit funny is that code too large errors typically pop-up during compilation. Does the shade plugin has its own size limit?","from":"developer"},{"body":"In my experience the 'shade' plugin does a lot more than it says (constant propagation, control-flow optimizations). I didn't find a way to turn them off.\n\n... and I don't know if it's only the last one. The first Jenkins build that failed had only this commit in the list of changes.","from":"developer"},{"body":"Shade doesn't do optimization but it certainly does several transformations. It's possible that something like renaming a package name to a longer one puts some method just over the top.\n\nThe problem isn't shading but huge methods. These should simply be broken down further.","from":"developer"},{"body":"Well, the class file format isn't embedding any names inside the code section, that's what the ConstantPool is for. What the shade plugin does or not is slightly irrelevant, the fact is that the failure happens (only) during shading, and all \"compile\" integration builds in Jenkins are passing without problem. Only the ones that run tests (and run `assembly`) fail.","from":"developer"},{"body":"Hm I think you're right about that. I also don't see why it would pass compilation. Is it possible scalac doesn't enforce the limit but shade does? or is that out of the question since the code would actually fail to load?\n\nThe answer still might be to break up the method.\n\nI did update the shade plugin about 5 days ago but it passed since.\nhttps://github.com/apache/spark/commit/b84404865b90fc07858a8b28c2a7028f96517167","from":"developer"},{"body":"According to AMPLAB's Jenkins, the first failure happened at following build that includes [https://issues.apache.org/jira/browse/SPARK-13306]\nhttps://amplab.cs.berkeley.edu/jenkins/job/spark-master-test-maven-hadoop-2.4/337/","from":"developer"},{"body":"Probably the easiest fix is to break some grammar elements in more files. I *think* the problem is deeper though, and there's some state explosion generated by an ambiguous part of the grammar.\n\nThe Scala compiler would show an error, but the ANTLR generator spits out Java files. Definitely those methods are very close to the limit, I saw javac failing when using the Eclipse compiler, but the Oracle one was slightly better.","from":"developer"},{"body":" The Java Virtual Machine specification limits the size of generated Java byte code for each method in a class to the maximum of 64K bytes. so i think its not something you can avoid at shade or scala compiler level.\nSome highly relevant links: https://groups.google.com/forum/#!topic/scala-internals/f6PwUxc8K7I\nhttp://www.cubrid.org/blog/dev-platform/understanding-jvm-internals/","from":"developer"},{"body":"I succeeded to build Spark by commenting out lines 208 and 209 in ExpressionParser.g\nhttps://github.com/apache/spark/blob/7925071280bfa1570435bde3e93492eaf2167d56/sql/catalyst/src/main/antlr3/org/apache/spark/sql/catalyst/parser/ExpressionParser.g#L208\n\nThese two lines are added by [https://issues.apache.org/jira/browse/SPARK-13306].","from":"developer"},{"body":"Test are currently passing and the if you call {{sparkShell}} from sbt you can work with spark and issue sql commands. So the direct problem is definitely caused by shading. The thing I am actually not getting is why sbt assemblies won't fail.\n\nThat being said, ANTLR spits out huge files which are very near to but not over the 64K limit (It won't build otherwise). I am not too confident that splitting the parser another time is much more than a temporary solution. The current grammar contains a perticulary nasty rule to determine if a keyword is non-reserved (this creates a huge state explosion), we should take a look at that. On the longer run I think that we should do a rewrite of the current parser with ANTLR4 which should simplify parser rules (and thus the size of the resulting parser), and make the construction of logical plans less Rube-Goldbergy/pattern-match-heaven.\n\n","from":"developer"},{"body":"[~kiszk] Are you building with Maven or SBT?","from":"developer"},{"body":"I am using mvn, and executed the following command\n\n{code}\nbuild/mvn -Pyarn -Phadoop-2.4 -DskipTests package\n{code}\n","from":"developer"},{"body":"Sbt assembly is using a different library, I think it's called jarjar. Nothing to do with the maven shade plugin.","from":"developer"},{"body":"The size of a static initializer method of SparkSqlParser_ExpressionParser.java is 59901. It seem to be the largest method among generated Java files. It is still less than 64K bytes.","from":"developer"},{"body":"I meet the same problem with maven, how to resolve this issue?\n\n*build command as follow:*\nbq. mvn clean package -DskipTests -Pscala-2.11 -Phadoop-2.6 -Phive -Phive-thriftserver -Pyarn -Dyarn.version=2.6.0-cdh5.4.7 -Dhadoop.version=2.6.0-cdh5.4.7\n\n*build error message:*\n{quote}\n[INFO] Excluding org.scala-lang:scala-reflect:jar:2.11.7 from the shaded jar.\n[INFO] Excluding org.scala-lang:scala-library:jar:2.11.7 from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-core_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.apache.avro:avro-mapred:jar:hadoop2:1.7.7 from the shaded jar.\n[INFO] Excluding org.apache.avro:avro-ipc:jar:1.7.7 from the shaded jar.\n[INFO] Excluding org.apache.avro:avro:jar:1.7.7 from the shaded jar.\n[INFO] Excluding org.apache.avro:avro-ipc:jar:tests:1.7.7 from the shaded jar.\n[INFO] Excluding org.codehaus.jackson:jackson-core-asl:jar:1.9.13 from the shaded jar.\n[INFO] Excluding org.codehaus.jackson:jackson-mapper-asl:jar:1.9.13 from the shaded jar.\n[INFO] Excluding com.twitter:chill_2.11:jar:0.5.0 from the shaded jar.\n[INFO] Excluding com.esotericsoftware.kryo:kryo:jar:2.21 from the shaded jar.\n[INFO] Excluding com.esotericsoftware.reflectasm:reflectasm:jar:shaded:1.07 from the shaded jar.\n[INFO] Excluding com.esotericsoftware.minlog:minlog:jar:1.2 from the shaded jar.\n[INFO] Excluding org.objenesis:objenesis:jar:1.2 from the shaded jar.\n[INFO] Excluding com.twitter:chill-java:jar:0.5.0 from the shaded jar.\n[INFO] Excluding org.apache.xbean:xbean-asm5-shaded:jar:4.4 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-client:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-common:jar:2.2.0 from the shaded jar.\n[INFO] Excluding commons-cli:commons-cli:jar:1.2 from the shaded jar.\n[INFO] Excluding org.apache.commons:commons-math:jar:2.1 from the shaded jar.\n[INFO] Excluding xmlenc:xmlenc:jar:0.52 from the shaded jar.\n[INFO] Excluding commons-io:commons-io:jar:2.1 from the shaded jar.\n[INFO] Excluding commons-lang:commons-lang:jar:2.6 from the shaded jar.\n[INFO] Excluding commons-configuration:commons-configuration:jar:1.6 from the shaded jar.\n[INFO] Excluding commons-collections:commons-collections:jar:3.2.2 from the shaded jar.\n[INFO] Excluding commons-digester:commons-digester:jar:1.8 from the shaded jar.\n[INFO] Excluding commons-beanutils:commons-beanutils:jar:1.7.0 from the shaded jar.\n[INFO] Excluding commons-beanutils:commons-beanutils-core:jar:1.8.0 from the shaded jar.\n[INFO] Excluding com.google.protobuf:protobuf-java:jar:2.5.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-auth:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.commons:commons-compress:jar:1.4.1 from the shaded jar.\n[INFO] Excluding org.tukaani:xz:jar:1.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-hdfs:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.mortbay.jetty:jetty-util:jar:6.1.26 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-mapreduce-client-app:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-mapreduce-client-common:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-yarn-client:jar:2.2.0 from the shaded jar.\n[INFO] Excluding com.google.inject:guice:jar:3.0 from the shaded jar.\n[INFO] Excluding javax.inject:javax.inject:jar:1 from the shaded jar.\n[INFO] Excluding aopalliance:aopalliance:jar:1.0 from the shaded jar.\n[INFO] Excluding com.sun.jersey.jersey-test-framework:jersey-test-framework-grizzly2:jar:1.9 from the shaded jar.\n[INFO] Excluding com.sun.jersey.jersey-test-framework:jersey-test-framework-core:jar:1.9 from the shaded jar.\n[INFO] Excluding javax.servlet:javax.servlet-api:jar:3.0.1 from the shaded jar.\n[INFO] Excluding com.sun.jersey:jersey-client:jar:1.9 from the shaded jar.\n[INFO] Excluding com.sun.jersey:jersey-grizzly2:jar:1.9 from the shaded jar.\n[INFO] Excluding org.glassfish.grizzly:grizzly-http:jar:2.1.2 from the shaded jar.\n[INFO] Excluding org.glassfish.grizzly:grizzly-framework:jar:2.1.2 from the shaded jar.\n[INFO] Excluding org.glassfish.gmbal:gmbal-api-only:jar:3.0.0-b023 from the shaded jar.\n[INFO] Excluding org.glassfish.external:management-api:jar:3.0.0-b012 from the shaded jar.\n[INFO] Excluding org.glassfish.grizzly:grizzly-http-server:jar:2.1.2 from the shaded jar.\n[INFO] Excluding org.glassfish.grizzly:grizzly-rcm:jar:2.1.2 from the shaded jar.\n[INFO] Excluding org.glassfish.grizzly:grizzly-http-servlet:jar:2.1.2 from the shaded jar.\n[INFO] Excluding org.glassfish:javax.servlet:jar:3.1 from the shaded jar.\n[INFO] Excluding com.sun.jersey:jersey-json:jar:1.9 from the shaded jar.\n[INFO] Excluding org.codehaus.jettison:jettison:jar:1.1 from the shaded jar.\n[INFO] Excluding com.sun.xml.bind:jaxb-impl:jar:2.2.3-1 from the shaded jar.\n[INFO] Excluding javax.xml.bind:jaxb-api:jar:2.2.2 from the shaded jar.\n[INFO] Excluding javax.activation:activation:jar:1.1 from the shaded jar.\n[INFO] Excluding org.codehaus.jackson:jackson-jaxrs:jar:1.9.13 from the shaded jar.\n[INFO] Excluding org.codehaus.jackson:jackson-xc:jar:1.9.13 from the shaded jar.\n[INFO] Excluding com.sun.jersey.contribs:jersey-guice:jar:1.9 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-yarn-server-common:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-mapreduce-client-shuffle:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-yarn-api:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-mapreduce-client-core:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-yarn-common:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-mapreduce-client-jobclient:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.hadoop:hadoop-annotations:jar:2.2.0 from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-launcher_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-network-common_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-network-shuffle_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.fusesource.leveldbjni:leveldbjni-all:jar:1.8 from the shaded jar.\n[INFO] Excluding com.fasterxml.jackson.core:jackson-annotations:jar:2.5.3 from the shaded jar.\n[INFO] Excluding net.java.dev.jets3t:jets3t:jar:0.7.1 from the shaded jar.\n[INFO] Excluding commons-httpclient:commons-httpclient:jar:3.1 from the shaded jar.\n[INFO] Excluding org.apache.curator:curator-recipes:jar:2.4.0 from the shaded jar.\n[INFO] Excluding org.apache.curator:curator-framework:jar:2.4.0 from the shaded jar.\n[INFO] Excluding org.apache.curator:curator-client:jar:2.4.0 from the shaded jar.\n[INFO] Excluding org.apache.zookeeper:zookeeper:jar:3.4.5 from the shaded jar.\n[INFO] Excluding jline:jline:jar:0.9.94 from the shaded jar.\n[INFO] Excluding org.eclipse.jetty.orbit:javax.servlet:jar:3.0.0.v201112011016 from the shaded jar.\n[INFO] Excluding org.apache.commons:commons-lang3:jar:3.3.2 from the shaded jar.\n[INFO] Excluding org.apache.commons:commons-math3:jar:3.4.1 from the shaded jar.\n[INFO] Excluding com.google.code.findbugs:jsr305:jar:1.3.9 from the shaded jar.\n[INFO] Excluding org.slf4j:slf4j-api:jar:1.7.16 from the shaded jar.\n[INFO] Excluding org.slf4j:jul-to-slf4j:jar:1.7.16 from the shaded jar.\n[INFO] Excluding org.slf4j:jcl-over-slf4j:jar:1.7.16 from the shaded jar.\n[INFO] Excluding log4j:log4j:jar:1.2.17 from the shaded jar.\n[INFO] Excluding org.slf4j:slf4j-log4j12:jar:1.7.16 from the shaded jar.\n[INFO] Excluding com.ning:compress-lzf:jar:1.0.3 from the shaded jar.\n[INFO] Excluding org.xerial.snappy:snappy-java:jar:1.1.2 from the shaded jar.\n[INFO] Excluding net.jpountz.lz4:lz4:jar:1.3.0 from the shaded jar.\n[INFO] Excluding org.roaringbitmap:RoaringBitmap:jar:0.5.11 from the shaded jar.\n[INFO] Excluding commons-net:commons-net:jar:2.2 from the shaded jar.\n[INFO] Excluding org.json4s:json4s-jackson_2.11:jar:3.2.10 from the shaded jar.\n[INFO] Excluding org.json4s:json4s-core_2.11:jar:3.2.10 from the shaded jar.\n[INFO] Excluding org.json4s:json4s-ast_2.11:jar:3.2.10 from the shaded jar.\n[INFO] Excluding org.scala-lang:scalap:jar:2.11.7 from the shaded jar.\n[INFO] Excluding org.scala-lang:scala-compiler:jar:2.11.7 from the shaded jar.\n[INFO] Excluding org.scala-lang.modules:scala-parser-combinators_2.11:jar:1.0.4 from the shaded jar.\n[INFO] Excluding com.sun.jersey:jersey-server:jar:1.9 from the shaded jar.\n[INFO] Excluding asm:asm:jar:3.1 from the shaded jar.\n[INFO] Excluding com.sun.jersey:jersey-core:jar:1.9 from the shaded jar.\n[INFO] Excluding org.apache.mesos:mesos:jar:shaded-protobuf:0.21.1 from the shaded jar.\n[INFO] Excluding io.netty:netty-all:jar:4.0.29.Final from the shaded jar.\n[INFO] Excluding io.netty:netty:jar:3.8.0.Final from the shaded jar.\n[INFO] Excluding com.clearspring.analytics:stream:jar:2.7.0 from the shaded jar.\n[INFO] Excluding io.dropwizard.metrics:metrics-core:jar:3.1.2 from the shaded jar.\n[INFO] Excluding io.dropwizard.metrics:metrics-jvm:jar:3.1.2 from the shaded jar.\n[INFO] Excluding io.dropwizard.metrics:metrics-json:jar:3.1.2 from the shaded jar.\n[INFO] Excluding io.dropwizard.metrics:metrics-graphite:jar:3.1.2 from the shaded jar.\n[INFO] Excluding com.fasterxml.jackson.core:jackson-databind:jar:2.5.3 from the shaded jar.\n[INFO] Excluding com.fasterxml.jackson.core:jackson-core:jar:2.5.3 from the shaded jar.\n[INFO] Excluding com.fasterxml.jackson.module:jackson-module-scala_2.11:jar:2.5.3 from the shaded jar.\n[INFO] Excluding com.thoughtworks.paranamer:paranamer:jar:2.6 from the shaded jar.\n[INFO] Excluding org.apache.ivy:ivy:jar:2.4.0 from the shaded jar.\n[INFO] Excluding oro:oro:jar:2.0.8 from the shaded jar.\n[INFO] Excluding net.razorvine:pyrolite:jar:4.9 from the shaded jar.\n[INFO] Excluding net.sf.py4j:py4j:jar:0.9.1 from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-unsafe_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.codehaus.janino:janino:jar:2.7.8 from the shaded jar.\n[INFO] Excluding org.codehaus.janino:commons-compiler:jar:2.7.8 from the shaded jar.\n[INFO] Excluding org.antlr:antlr-runtime:jar:3.5.2 from the shaded jar.\n[INFO] Excluding commons-codec:commons-codec:jar:1.10 from the shaded jar.\n[INFO] Including org.spark-project.spark:unused:jar:1.0.0 in the shaded jar.\n[INFO] Excluding org.scala-lang.modules:scala-xml_2.11:jar:1.0.2 from the shaded jar.\n[DEBUG] Processing JAR /home/myubuntu/Works/workspace/spark/sql/catalyst/target/spark-catalyst_2.11-2.0.0-SNAPSHOT.jar\n[INFO] ------------------------------------------------------------------------\n[INFO] BUILD FAILURE\n[INFO] ------------------------------------------------------------------------\n[INFO] Total time: 52.733 s (Wall Clock)\n[INFO] Finished at: 2016-02-23T13:26:26+08:00\n[INFO] Final Memory: 38M/744M\n[INFO] ------------------------------------------------------------------------\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-shade-plugin:2.4.3:shade (default) on project spark-catalyst_2.11: Error creating shaded jar: Method code too large! -> [Help 1]\norg.apache.maven.lifecycle.LifecycleExecutionException: Failed to execute goal org.apache.maven.plugins:maven-shade-plugin:2.4.3:shade (default) on project spark-catalyst_2.11: Error creating shaded jar: Method code too large!\n\tat org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:212)\n\tat org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:153)\n\tat org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:145)\n\tat org.apache.maven.lifecycle.internal.LifecycleModuleBuilder.buildProject(LifecycleModuleBuilder.java:116)\n\tat org.apache.maven.lifecycle.internal.builder.multithreaded.MultiThreadedBuilder$1.call(MultiThreadedBuilder.java:185)\n\tat org.apache.maven.lifecycle.internal.builder.multithreaded.MultiThreadedBuilder$1.call(MultiThreadedBuilder.java:181)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.maven.plugin.MojoExecutionException: Error creating shaded jar: Method code too large!\n\tat org.apache.maven.plugins.shade.mojo.ShadeMojo.execute(ShadeMojo.java:540)\n\tat org.apache.maven.plugin.DefaultBuildPluginManager.executeMojo(DefaultBuildPluginManager.java:134)\n\tat org.apache.maven.lifecycle.internal.MojoExecutor.execute(MojoExecutor.java:207)\n\t... 11 more\nCaused by: java.lang.RuntimeException: Method code too large!\n\tat org.objectweb.asm.MethodWriter.a(Unknown Source)\n\tat org.objectweb.asm.ClassWriter.toByteArray(Unknown Source)\n\tat org.apache.maven.plugins.shade.DefaultShader.addRemappedClass(DefaultShader.java:453)\n\tat org.apache.maven.plugins.shade.DefaultShader.shadeSingleJar(DefaultShader.java:219)\n\tat org.apache.maven.plugins.shade.DefaultShader.shadeJars(DefaultShader.java:179)\n\tat org.apache.maven.plugins.shade.DefaultShader.shade(DefaultShader.java:104)\n\tat org.apache.maven.plugins.shade.mojo.ShadeMojo.execute(ShadeMojo.java:454)\n\t... 13 more\n{quote}","from":"developer"},{"body":"Try build/sbt before the community finalizes the solution. For example, \n{code}\nbuild/sbt -Pyarn -Phadoop-2.6 -Phive -Phive-thriftserver assembly\n{code}","from":"developer"},{"body":"I can't build it with sbt because downloading jar from sbt repository is very slow, even failed.","from":"developer"},{"body":"So there is no other option for making a distro? ","from":"developer"},{"body":"I confirmed that this workaround also enables the make-distribution.sh script to complete properly.","from":"developer"},{"body":"I meet to the same problem, using the repo cloned from https://github.com/apache/spark.git.","from":"developer"},{"body":"I experiencing the same problem when checking out 1.6.0 (git clone git@github.com:apache/spark.git v1.6.0).\n....\n[INFO] Excluding org.apache.ivy:ivy:jar:2.4.0 from the shaded jar.\n[INFO] Excluding oro:oro:jar:2.0.8 from the shaded jar.\n[INFO] Excluding net.razorvine:pyrolite:jar:4.9 from the shaded jar.\n[INFO] Excluding net.sf.py4j:py4j:jar:0.9.1 from the shaded jar.\n[INFO] Excluding org.apache.spark:spark-unsafe_2.11:jar:2.0.0-SNAPSHOT from the shaded jar.\n[INFO] Excluding org.codehaus.janino:janino:jar:2.7.8 from the shaded jar.\n[INFO] Excluding org.codehaus.janino:commons-compiler:jar:2.7.8 from the shaded jar.\n[INFO] Excluding org.antlr:antlr-runtime:jar:3.5.2 from the shaded jar.\n[INFO] Excluding commons-codec:commons-codec:jar:1.10 from the shaded jar.\n[INFO] Including org.spark-project.spark:unused:jar:1.0.0 in the shaded jar.\n[INFO] Excluding org.scala-lang.modules:scala-xml_2.11:jar:1.0.2 from the shaded jar.\n[INFO] ------------------------------------------------------------------------\n[INFO] Reactor Summary:\n[INFO] \n[INFO] Spark Project Parent POM ........................... SUCCESS [ 7.063 s]\n[INFO] Spark Project Sketch ............................... SUCCESS [ 7.766 s]\n[INFO] Spark Project Test Tags ............................ SUCCESS [ 3.325 s]\n[INFO] Spark Project Core ................................. SUCCESS [03:13 min]\n[INFO] Spark Project GraphX ............................... SUCCESS [ 23.761 s]\n[INFO] Spark Project ML Library ........................... SUCCESS [01:52 min]\n[INFO] Spark Project Tools ................................ SUCCESS [ 4.876 s]\n[INFO] Spark Project Networking ........................... SUCCESS [ 16.129 s]\n[INFO] Spark Project Shuffle Streaming Service ............ SUCCESS [ 13.367 s]\n[INFO] Spark Project Streaming ............................ SUCCESS [ 53.482 s]\n[INFO] Spark Project Catalyst ............................. FAILURE [02:15 min]\n[INFO] Spark Project SQL .................................. SKIPPED\n[INFO] Spark Project Hive ................................. SKIPPED\n[INFO] Spark Project Docker Integration Tests ............. SKIPPED\n[INFO] Spark Project Unsafe ............................... SKIPPED\n[INFO] Spark Project Assembly ............................. SKIPPED\n[INFO] Spark Project External Twitter ..................... SKIPPED\n[INFO] Spark Project External Flume ....................... SKIPPED\n[INFO] Spark Project External Flume Sink .................. SKIPPED\n[INFO] Spark Project External Flume Assembly .............. SKIPPED\n[INFO] Spark Project External Akka ........................ SKIPPED\n[INFO] Spark Project External MQTT ........................ SKIPPED\n[INFO] Spark Project External MQTT Assembly ............... SKIPPED\n[INFO] Spark Project External ZeroMQ ...................... SKIPPED\n[INFO] Spark Project Examples ............................. SKIPPED\n[INFO] Spark Project REPL ................................. SKIPPED\n[INFO] Spark Project Launcher ............................. SKIPPED\n[INFO] Spark Project External Kafka ....................... SKIPPED\n[INFO] Spark Project External Kafka Assembly .............. SKIPPED\n[INFO] Spark Project YARN ................................. SKIPPED\n[INFO] Spark Project YARN Shuffle Service ................. SKIPPED\n[INFO] ------------------------------------------------------------------------\n[INFO] BUILD FAILURE\n[INFO] ------------------------------------------------------------------------\n[INFO] Total time: 09:34 min\n[INFO] Finished at: 2016-02-23T17:04:28+01:00\n[INFO] Final Memory: 65M/1297M\n[INFO] ------------------------------------------------------------------------\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-shade-plugin:2.4.3:shade (default) on project spark-catalyst_2.10: Error creating shaded jar: Method code too large! -> [Help 1]\n[ERROR] \n[ERROR] To see the full stack trace of the errors, re-run Maven with the -e switch.\n[ERROR] Re-run Maven using the -X switch to enable full debug logging.\n[ERROR] \n[ERROR] For more information about the errors and possible solutions, please read the following articles:\n[ERROR] [Help 1] http://cwiki.apache.org/confluence/display/MAVEN/MojoExecutionException","from":"developer"},{"body":"I think the only way is to revert the offending changes, as mentioned [here|https://issues.apache.org/jira/browse/SPARK-13431?focusedCommentId=15157182&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-15157182]","from":"developer"},{"body":"Yes, nobody's suggesting this isn't clearly reproducible. No need to post this.","from":"developer"},{"body":"Sorry for the spam. I only thought I report it as the bug is only filed as affecting v2.0.0 (and I am using v1.6.0)","from":"developer"},{"body":"I identified why this problem occurs only in maven. Shade plugin for maven increases the length of Java bytecode for a method. This increasing happens since shade plugin rewrites Java bytecode to rebuild constant pool.\n\nHere is an output of ``javap -c SparkSqlParser_ExpressionParser.class`` before applying shade plugin. The static initializer ``static{}`` uses ``ldc`` bytecode for accessing constant pool at offset 13, 18, and 23. Each ``ldc`` consume only two bytes. As a result, the bytecode length of this method is *less than 65536*.\n{code}\npublic class org.apache.spark.sql.catalyst.parser.SparkSqlParser_ExpressionParser extends org.antlr.runtime.Parser {\n...\n static {};\n Code:\n 0: bipush 70\n 2: anewarray #1035 // class java/lang/String\n 5: dup\n 6: iconst_0\n 7: ldc_w #1036 // String ...\n 10: aastore\n 11: dup\n 12: iconst_1\n 13: ldc #127 // String\n 15: aastore\n 16: dup\n 17: iconst_2\n 18: ldc #127 // String\n 20: aastore\n 21: dup\n 22: iconst_3\n 23: ldc #127 // String\n 25: aastore\n ...\n 59900: return\n }\n}\n{code}\n\nAfter applying shade plugin, the static initializer ``static{}`` uses ``ldc_w`` bytecode for accessing constant pool at offset 13, 19, and 25. Each ``ldc_w`` consumes three bytes. As a result, the bytecode length of this method is *more than 65535*.\n{code}\n static {};\n Code:\n 0: bipush 70\n 2: anewarray #2965 // class java/lang/String\n 5: dup\n 6: iconst_0\n 7: ldc_w #5240 // String ...\n 10: aastore\n 11: dup\n 12: iconst_1\n 13: ldc_w #2924 // String\n 16: aastore\n 17: dup\n 18: iconst_2\n 19: ldc_w #2924 // String\n 22: aastore\n 23: dup\n 24: iconst_3\n 25: ldc_w #2924 // String\n 28: aastore\n ...\n 65533: lconst_0\n 65534: lastore\n ...\n\n }\n}\n{code}\n\nShading plugin seems to rebuild constant pool based on [this comment|http://svn.apache.org/viewvc/maven/plugins/tags/maven-shade-plugin-2.4.3/src/main/java/org/apache/maven/plugins/shade/DefaultShader.java?view=markup#l417]. To use a lot of constant pool entry due to many definitions of String may increase the entry index of the constant pool. As a result, it leads to replace ``ldc`` with ``ldc_w``. Finally, the length of Java bytecode is increased.\n\nAs a next step, what will we do?\n* Can we avoid this rebuild by an option?\n* Can we create a pull request for shade plugin to avoid this?\n* Can we use another plugin?\n* Can we split ExpressionParser.g into smaller files?\n* Other solutions?\n\n\n","from":"developer"},{"body":"Submitted PR for revert: https://github.com/apache/spark/pull/11329","from":"developer"},{"body":"I'd like to split ExpressionParser.g, or we can't touch it anymore (may break sbt break next time), unless we can switch to ANTR4 soon.","from":"developer"},{"body":"https://github.com/apache/spark/pull/11331","from":"developer"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11331","from":"developer"},{"body":"cool!","from":"developer"},{"body":"Issue resolved by pull request 11331\n[https://github.com/apache/spark/pull/11331]","from":"developer"}],"created":"2016-02-22T09:49:23.000+0000","description":"Cannot build the project when run the normal build commands:\neg.\n{code}\nbuild/mvn -Phadoop-2.6 -Dhadoop.version=2.6.0 clean package\n./make-distribution.sh --name test --tgz -Phadoop-2.6 \n{code}\n\nIntegration builds are also failing: \n\nhttps://amplab.cs.berkeley.edu/jenkins/job/spark-master-test-maven-hadoop-2.6/229/console\nhttps://ci.typesafe.com/job/mit-docker-test-zk-ref/12/console\n\nIt looks like this is the commit that introduced the issue:\n\nhttps://github.com/apache/spark/commit/7925071280bfa1570435bde3e93492eaf2167d56","issue_id":"12940995","key":"SPARK-13431","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-02-24T05:22:53.000+0000","role":"fixed_distractor","summary":"Maven build fails due to: Method code too large! in Catalyst"} {"case_id":"12941327","cluster":"DISTRACTOR-SPARK-13450","comments":[{"body":"I think we should add a ExternalAppendOnlyArrayBuffer to replace ArrayBuffer.\nI will add in my own branch first.","created":"2016-02-23T08:50:28.219+0000"},{"body":"This isn't a helpful JIRA since you didn't say how you reproduce this at all. ","created":"2016-02-23T09:41:30.098+0000"},{"body":"+1 on reproduceable bugs.\n\nSo you are basically doing a Cartesian Product. Sort-Merge join will cache the right side of the join, it is likely OOM when there are a lot of rows with the same key. You can try to mitigate this problem by actually using a Cartesian Join (will perform horrible), do a broadcast join (if one of the sides fits in memory), or do some sort of a co-group.\n\nIf you want to go ahead and want to improve this, I'd suggest you take a look at spark/sql/core/src/main/scala/org/apache/spark/sql/execution/Window.scala in which we also spill to disk when the number of rows in a partition becomes to large. ","created":"2016-02-23T12:54:27.196+0000"},{"body":"A join has a lot of rows with the same key.","created":"2016-02-24T06:41:05.741+0000"},{"body":"Thanks, I will add in my own branch first.","created":"2016-02-24T11:36:23.155+0000"},{"body":"User 'shenh062326' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11386","created":"2016-02-26T07:39:03.923+0000"},{"body":"I have seen this problem a couple times in prod while trying out jobs over Spark. There have been some discussions in the jira and here are my comments on those:\n\n- How to reproduce ? As [~shenhong] said, even in my case there were keys in the joined relation which were skewed. I was able to grab a heap dump which shows the array buffer grown more than a GB : https://issues.apache.org/jira/secure/attachment/12846382/heap-dump-analysis.png\n- [~hvanhovell] had some suggestions about using Cartesian Join OR co-group. In my case, our users are running Hive SQL queries as-is over Spark. If there are one-off such cases, we could have done that but with automated migration of several jobs, we want the query to just work. OOMs lead to un-reliable behavior and affects users.\n- I looked at `Window.scala` but it only works for unsafe rows. In sort merge join, we may or may not have unsafe rows.\n- There was a PR associated with this jira (https://github.com/apache/spark/pull/11386) but its inactive. I tried to pick it up but it does not apply. Its basically copying ExternalAppendOnlyMap code and introducing a `Buffer` version of it for lists. Instead of taking care of merge conflicts, I was able to implement buffer version by reusing `ExternalAppendOnlyMap` code. Will submit a fresh PR after testing it.\n\n{code}\njava.lang.OutOfMemoryError: Java heap space\nat org.apache.spark.sql.catalyst.expressions.UnsafeRow.copy(UnsafeRow.java:503)\nat org.apache.spark.sql.catalyst.expressions.UnsafeRow.copy(UnsafeRow.java:61)\nat org.apache.spark.sql.execution.joins.SortMergeJoinScanner.bufferMatchingRows(SortMergeJoinExec.scala:756)\nat org.apache.spark.sql.execution.joins.SortMergeJoinScanner.findNextInnerJoinRows(SortMergeJoinExec.scala:660)\nat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$doExecute$1$$anon$1.advanceNext(SortMergeJoinExec.scala:137)\nat org.apache.spark.sql.execution.RowIteratorToScala.hasNext(RowIterator.scala:68)\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\nat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.processInputs(TungstenAggregationIterator.scala:186)\nat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.(TungstenAggregationIterator.scala:355)\nat org.apache.spark.sql.execution.aggregate.HashAggregateExec$$anonfun$doExecute$1$$anonfun$4.apply(HashAggregateExec.scala:103)\nat org.apache.spark.sql.execution.aggregate.HashAggregateExec$$anonfun$doExecute$1$$anonfun$4.apply(HashAggregateExec.scala:94)\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:785)\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:785)\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\nat org.apache.spark.scheduler.Task.run(Task.scala:86)\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\nat java.lang.Thread.run(Thread.java:745)\n{code}\n","created":"2017-01-09T18:46:19.898+0000"},{"body":"ExternalAppendOnlyMap estimate the size of the data saved. In SortMergeJoin, I think we can leverage UnsafeExternalSorter to get more accurate and controllable behavior.","created":"2017-01-09T19:02:31.885+0000"},{"body":"User 'tejasapatil' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16909","created":"2017-02-13T07:14:03.176+0000"}],"conversations":[{"body":" When I run a sql with join, task throw java.lang.OutOfMemoryError and sql failed. I have set spark.executor.memory 4096m.\n SortMergeJoin use a ArrayBuffer[InternalRow] to store bufferedMatches, if the join rows have a lot of same key, it will throw OutOfMemoryError.\n\n{code}\n /** Buffered rows from the buffered side of the join. This is empty if there are no matches. */\n private[this] val bufferedMatches: ArrayBuffer[InternalRow] = new ArrayBuffer[InternalRow]\n{code}\n\n\n Here is the stackTrace:\n{code}\norg.xerial.snappy.SnappyNative.arrayCopy(Native Method)\norg.xerial.snappy.Snappy.arrayCopy(Snappy.java:84)\norg.xerial.snappy.SnappyInputStream.rawRead(SnappyInputStream.java:190)\norg.xerial.snappy.SnappyInputStream.read(SnappyInputStream.java:163)\njava.io.DataInputStream.readFully(DataInputStream.java:195)\njava.io.DataInputStream.readLong(DataInputStream.java:416)\norg.apache.spark.util.collection.unsafe.sort.UnsafeSorterSpillReader.loadNext(UnsafeSorterSpillReader.java:71)\norg.apache.spark.util.collection.unsafe.sort.UnsafeSorterSpillMerger$2.loadNext(UnsafeSorterSpillMerger.java:79)\norg.apache.spark.sql.execution.UnsafeExternalRowSorter$1.next(UnsafeExternalRowSorter.java:136)\norg.apache.spark.sql.execution.UnsafeExternalRowSorter$1.next(UnsafeExternalRowSorter.java:123)\norg.apache.spark.sql.execution.RowIteratorFromScala.advanceNext(RowIterator.scala:84)\norg.apache.spark.sql.execution.joins.SortMergeJoinScanner.advancedBufferedToRowWithNullFreeJoinKey(SortMergeJoin.scala:300)\norg.apache.spark.sql.execution.joins.SortMergeJoinScanner.bufferMatchingRows(SortMergeJoin.scala:329)\norg.apache.spark.sql.execution.joins.SortMergeJoinScanner.findNextInnerJoinRows(SortMergeJoin.scala:229)\norg.apache.spark.sql.execution.joins.SortMergeJoin$$anonfun$doExecute$1$$anon$1.advanceNext(SortMergeJoin.scala:105)\norg.apache.spark.sql.execution.RowIteratorToScala.hasNext(RowIterator.scala:68)\nscala.collection.Iterator$$anon$11.hasNext(Iterator.scala:327)\norg.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:88)\norg.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:86)\norg.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:741)\norg.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:741)\norg.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\norg.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:337)\norg.apache.spark.rdd.RDD.iterator(RDD.scala:301)\norg.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\norg.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:337)\norg.apache.spark.rdd.RDD.iterator(RDD.scala:301)\norg.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\norg.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\norg.apache.spark.scheduler.Task.run(Task.scala:89)\norg.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:215)\njava.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\njava.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\njava.lang.Thread.run(Thread.java:744)\n{code}","from":"reporter","subject":"SortMergeJoin will OOM when join rows have lot of same keys"},{"body":"I think we should add a ExternalAppendOnlyArrayBuffer to replace ArrayBuffer.\nI will add in my own branch first.","from":"developer"},{"body":"This isn't a helpful JIRA since you didn't say how you reproduce this at all. ","from":"developer"},{"body":"+1 on reproduceable bugs.\n\nSo you are basically doing a Cartesian Product. Sort-Merge join will cache the right side of the join, it is likely OOM when there are a lot of rows with the same key. You can try to mitigate this problem by actually using a Cartesian Join (will perform horrible), do a broadcast join (if one of the sides fits in memory), or do some sort of a co-group.\n\nIf you want to go ahead and want to improve this, I'd suggest you take a look at spark/sql/core/src/main/scala/org/apache/spark/sql/execution/Window.scala in which we also spill to disk when the number of rows in a partition becomes to large. ","from":"developer"},{"body":"A join has a lot of rows with the same key.","from":"developer"},{"body":"Thanks, I will add in my own branch first.","from":"developer"},{"body":"User 'shenh062326' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11386","from":"developer"},{"body":"I have seen this problem a couple times in prod while trying out jobs over Spark. There have been some discussions in the jira and here are my comments on those:\n\n- How to reproduce ? As [~shenhong] said, even in my case there were keys in the joined relation which were skewed. I was able to grab a heap dump which shows the array buffer grown more than a GB : https://issues.apache.org/jira/secure/attachment/12846382/heap-dump-analysis.png\n- [~hvanhovell] had some suggestions about using Cartesian Join OR co-group. In my case, our users are running Hive SQL queries as-is over Spark. If there are one-off such cases, we could have done that but with automated migration of several jobs, we want the query to just work. OOMs lead to un-reliable behavior and affects users.\n- I looked at `Window.scala` but it only works for unsafe rows. In sort merge join, we may or may not have unsafe rows.\n- There was a PR associated with this jira (https://github.com/apache/spark/pull/11386) but its inactive. I tried to pick it up but it does not apply. Its basically copying ExternalAppendOnlyMap code and introducing a `Buffer` version of it for lists. Instead of taking care of merge conflicts, I was able to implement buffer version by reusing `ExternalAppendOnlyMap` code. Will submit a fresh PR after testing it.\n\n{code}\njava.lang.OutOfMemoryError: Java heap space\nat org.apache.spark.sql.catalyst.expressions.UnsafeRow.copy(UnsafeRow.java:503)\nat org.apache.spark.sql.catalyst.expressions.UnsafeRow.copy(UnsafeRow.java:61)\nat org.apache.spark.sql.execution.joins.SortMergeJoinScanner.bufferMatchingRows(SortMergeJoinExec.scala:756)\nat org.apache.spark.sql.execution.joins.SortMergeJoinScanner.findNextInnerJoinRows(SortMergeJoinExec.scala:660)\nat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$doExecute$1$$anon$1.advanceNext(SortMergeJoinExec.scala:137)\nat org.apache.spark.sql.execution.RowIteratorToScala.hasNext(RowIterator.scala:68)\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\nat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.processInputs(TungstenAggregationIterator.scala:186)\nat org.apache.spark.sql.execution.aggregate.TungstenAggregationIterator.(TungstenAggregationIterator.scala:355)\nat org.apache.spark.sql.execution.aggregate.HashAggregateExec$$anonfun$doExecute$1$$anonfun$4.apply(HashAggregateExec.scala:103)\nat org.apache.spark.sql.execution.aggregate.HashAggregateExec$$anonfun$doExecute$1$$anonfun$4.apply(HashAggregateExec.scala:94)\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:785)\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:785)\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\nat org.apache.spark.scheduler.Task.run(Task.scala:86)\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\nat java.lang.Thread.run(Thread.java:745)\n{code}\n","from":"developer"},{"body":"ExternalAppendOnlyMap estimate the size of the data saved. In SortMergeJoin, I think we can leverage UnsafeExternalSorter to get more accurate and controllable behavior.","from":"developer"},{"body":"User 'tejasapatil' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16909","from":"developer"}],"created":"2016-02-23T08:44:00.000+0000","description":" When I run a sql with join, task throw java.lang.OutOfMemoryError and sql failed. I have set spark.executor.memory 4096m.\n SortMergeJoin use a ArrayBuffer[InternalRow] to store bufferedMatches, if the join rows have a lot of same key, it will throw OutOfMemoryError.\n\n{code}\n /** Buffered rows from the buffered side of the join. This is empty if there are no matches. */\n private[this] val bufferedMatches: ArrayBuffer[InternalRow] = new ArrayBuffer[InternalRow]\n{code}\n\n\n Here is the stackTrace:\n{code}\norg.xerial.snappy.SnappyNative.arrayCopy(Native Method)\norg.xerial.snappy.Snappy.arrayCopy(Snappy.java:84)\norg.xerial.snappy.SnappyInputStream.rawRead(SnappyInputStream.java:190)\norg.xerial.snappy.SnappyInputStream.read(SnappyInputStream.java:163)\njava.io.DataInputStream.readFully(DataInputStream.java:195)\njava.io.DataInputStream.readLong(DataInputStream.java:416)\norg.apache.spark.util.collection.unsafe.sort.UnsafeSorterSpillReader.loadNext(UnsafeSorterSpillReader.java:71)\norg.apache.spark.util.collection.unsafe.sort.UnsafeSorterSpillMerger$2.loadNext(UnsafeSorterSpillMerger.java:79)\norg.apache.spark.sql.execution.UnsafeExternalRowSorter$1.next(UnsafeExternalRowSorter.java:136)\norg.apache.spark.sql.execution.UnsafeExternalRowSorter$1.next(UnsafeExternalRowSorter.java:123)\norg.apache.spark.sql.execution.RowIteratorFromScala.advanceNext(RowIterator.scala:84)\norg.apache.spark.sql.execution.joins.SortMergeJoinScanner.advancedBufferedToRowWithNullFreeJoinKey(SortMergeJoin.scala:300)\norg.apache.spark.sql.execution.joins.SortMergeJoinScanner.bufferMatchingRows(SortMergeJoin.scala:329)\norg.apache.spark.sql.execution.joins.SortMergeJoinScanner.findNextInnerJoinRows(SortMergeJoin.scala:229)\norg.apache.spark.sql.execution.joins.SortMergeJoin$$anonfun$doExecute$1$$anon$1.advanceNext(SortMergeJoin.scala:105)\norg.apache.spark.sql.execution.RowIteratorToScala.hasNext(RowIterator.scala:68)\nscala.collection.Iterator$$anon$11.hasNext(Iterator.scala:327)\norg.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:88)\norg.apache.spark.sql.execution.aggregate.TungstenAggregate$$anonfun$doExecute$1$$anonfun$2.apply(TungstenAggregate.scala:86)\norg.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:741)\norg.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$20.apply(RDD.scala:741)\norg.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\norg.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:337)\norg.apache.spark.rdd.RDD.iterator(RDD.scala:301)\norg.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\norg.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:337)\norg.apache.spark.rdd.RDD.iterator(RDD.scala:301)\norg.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\norg.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\norg.apache.spark.scheduler.Task.run(Task.scala:89)\norg.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:215)\njava.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\njava.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\njava.lang.Thread.run(Thread.java:744)\n{code}","issue_id":"12941327","key":"SPARK-13450","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-03-15T19:27:32.000+0000","role":"fixed_distractor","summary":"SortMergeJoin will OOM when join rows have lot of same keys"} {"case_id":"12944324","cluster":"DISTRACTOR-SPARK-13478","comments":[{"body":"Actually d'oh, I had secure HBase for the testing anyway, so I just checked and HBase doesn't have the same problem.","created":"2016-02-25T01:53:02.287+0000"},{"body":"For the record, here's the exception you get:\n\n{noformat}\n16/02/24 18:06:48 ERROR transport.TSaslTransport: SASL negotiation failure\njavax.security.sasl.SaslException: GSS initiate failed [Caused by GSSException: No valid credentials provided (Mechanism level: Failed to find any Kerberos tgt)]\n at com.sun.security.sasl.gsskerb.GssKrb5Client.evaluateChallenge(GssKrb5Client.java:212)\n at org.apache.thrift.transport.TSaslClientTransport.handleSaslStartMessage(TSaslClientTransport.java:94)\n at org.apache.thrift.transport.TSaslTransport.open(TSaslTransport.java:271)\n at org.apache.thrift.transport.TSaslClientTransport.open(TSaslClientTransport.java:37)\n[...plus lots of other stuff]\n{noformat}","created":"2016-02-25T02:07:59.175+0000"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11358","created":"2016-02-25T02:16:04.004+0000"},{"body":"Can we get this patch applied to 1.6.x as well?","created":"2017-01-19T20:05:45.373+0000"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16665","created":"2017-01-20T23:43:03.808+0000"},{"body":"[~vanzin] As we know 'spark-submit' command by itself doesn't officially allow the '\\-\\-proxy\\-user' and '\\-\\-principal/\\-ketab' options to be used together. So, how this patch working?","created":"2018-10-15T14:37:25.133+0000"},{"body":"You don't need a keytab to log in to kerberos...","created":"2018-10-15T16:33:21.253+0000"},{"body":"[~vanzin]: If I want to 'spark-submit' with keytab plus principal and still use a proxy user, how will I accomplish this in deploy mode 'cluster'. For, 'client' deploy mode things would probably work if I do a 'kinit' first in the shell and then execute 'spark-submit'. So, will the same method work for 'cluster' deploy mode?","created":"2018-10-19T06:15:25.621+0000"},{"body":"If you have a principal and a keytab you don't need a proxy user. Just use the principal and the keytab.\r\n\r\nIf the principal you're using is not the principal you want the app to run as, then you can't give the app the principal's keytab. Otherwise the other user will have access to it, and that's a security problem.","created":"2018-10-19T16:27:26.159+0000"},{"body":"[~vanzin]: My use case's a bit different. Imagine, I'm kind of like a power user say, 'poweruser'. I can login as 'poweruser' and I have its keytab and principal. And, I have a normal user say, 'user1'. 'user1' doesn't have access to poweruser's keytab or principal. But, being  'poweruser', I'd want to impersonate 'user1' and run 'spark-submit' in a shell. Also, please consider I can't login into user1 shell but poweruser's\r\n\r\nIs this use case valid or you see a potential security concern?","created":"2018-10-20T05:43:03.919+0000"},{"body":"If you run as user1 with the poweruser's keytab, it means the code running as user1 will have access to poweruser's keytab. That may be a security concern depending on your environment.\r\n\r\nBut the REALLY safe way is to deploy the app with user1's keytab. If you can impersonate user1, there's no reason why you wouldn't instead get its keytab, so just do that instead. It's safer since it keeps the poweruser's keytab more protected.\r\n\r\nSo, to repeat this again: always use the keytab of the least privileged user. Or just give up on token renewal and restart the app as needed.","created":"2018-10-20T23:17:17.162+0000"}],"conversations":[{"body":"If you use spark-submit's proxy user support, the code that fetches delegation tokens for the Hive Metastore fails. It seems like the Hive library tries to connect to the Metastore as the proxy user, and it doesn't have a Kerberos TGT for that user, so it fails.\n\nI don't know whether the same issue exists in the HBase code, but I'll make a similar change so that both behave similarly.","from":"reporter","subject":"Fetching delegation tokens for Hive fails when using proxy users"},{"body":"Actually d'oh, I had secure HBase for the testing anyway, so I just checked and HBase doesn't have the same problem.","from":"developer"},{"body":"For the record, here's the exception you get:\n\n{noformat}\n16/02/24 18:06:48 ERROR transport.TSaslTransport: SASL negotiation failure\njavax.security.sasl.SaslException: GSS initiate failed [Caused by GSSException: No valid credentials provided (Mechanism level: Failed to find any Kerberos tgt)]\n at com.sun.security.sasl.gsskerb.GssKrb5Client.evaluateChallenge(GssKrb5Client.java:212)\n at org.apache.thrift.transport.TSaslClientTransport.handleSaslStartMessage(TSaslClientTransport.java:94)\n at org.apache.thrift.transport.TSaslTransport.open(TSaslTransport.java:271)\n at org.apache.thrift.transport.TSaslClientTransport.open(TSaslClientTransport.java:37)\n[...plus lots of other stuff]\n{noformat}","from":"developer"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11358","from":"developer"},{"body":"Can we get this patch applied to 1.6.x as well?","from":"developer"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16665","from":"developer"},{"body":"[~vanzin] As we know 'spark-submit' command by itself doesn't officially allow the '\\-\\-proxy\\-user' and '\\-\\-principal/\\-ketab' options to be used together. So, how this patch working?","from":"developer"},{"body":"You don't need a keytab to log in to kerberos...","from":"developer"},{"body":"[~vanzin]: If I want to 'spark-submit' with keytab plus principal and still use a proxy user, how will I accomplish this in deploy mode 'cluster'. For, 'client' deploy mode things would probably work if I do a 'kinit' first in the shell and then execute 'spark-submit'. So, will the same method work for 'cluster' deploy mode?","from":"developer"},{"body":"If you have a principal and a keytab you don't need a proxy user. Just use the principal and the keytab.\r\n\r\nIf the principal you're using is not the principal you want the app to run as, then you can't give the app the principal's keytab. Otherwise the other user will have access to it, and that's a security problem.","from":"developer"},{"body":"[~vanzin]: My use case's a bit different. Imagine, I'm kind of like a power user say, 'poweruser'. I can login as 'poweruser' and I have its keytab and principal. And, I have a normal user say, 'user1'. 'user1' doesn't have access to poweruser's keytab or principal. But, being  'poweruser', I'd want to impersonate 'user1' and run 'spark-submit' in a shell. Also, please consider I can't login into user1 shell but poweruser's\r\n\r\nIs this use case valid or you see a potential security concern?","from":"developer"},{"body":"If you run as user1 with the poweruser's keytab, it means the code running as user1 will have access to poweruser's keytab. That may be a security concern depending on your environment.\r\n\r\nBut the REALLY safe way is to deploy the app with user1's keytab. If you can impersonate user1, there's no reason why you wouldn't instead get its keytab, so just do that instead. It's safer since it keeps the poweruser's keytab more protected.\r\n\r\nSo, to repeat this again: always use the keytab of the least privileged user. Or just give up on token renewal and restart the app as needed.","from":"developer"}],"created":"2016-02-24T23:29:48.000+0000","description":"If you use spark-submit's proxy user support, the code that fetches delegation tokens for the Hive Metastore fails. It seems like the Hive library tries to connect to the Metastore as the proxy user, and it doesn't have a Kerberos TGT for that user, so it fails.\n\nI don't know whether the same issue exists in the HBase code, but I'll make a similar change so that both behave similarly.","issue_id":"12944324","key":"SPARK-13478","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-01-22T00:23:49.000+0000","role":"fixed_distractor","summary":"Fetching delegation tokens for Hive fails when using proxy users"} {"case_id":"12947600","cluster":"DISTRACTOR-SPARK-13709","comments":[{"body":"Is similar to SPARK-13572","created":"2016-03-21T15:51:04.575+0000"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13865","created":"2016-06-22T23:57:04.593+0000"},{"body":"Issue resolved by pull request 13865\n[https://github.com/apache/spark/pull/13865]","created":"2016-06-24T06:12:52.080+0000"}],"conversations":[{"body":"There is a problem decoding Avro data with SparkSQL when partitioned. The schema and encoded data are valid -- I'm able to decode the data with the avro-tools CLI utility. I'm also able to decode the data with non-partitioned SparkSQL tables, Hive, other tools as well... except partitioned SparkSQL schemas.\n\nFor a simple example, I took the example schema and data found in the Oracle documentation here:\n\n*Schema*\n{code:javascript}\n{\n \"type\": \"record\",\n \"name\": \"MemberInfo\",\n \"namespace\": \"avro\",\n \"fields\": [\n {\"name\": \"name\", \"type\": {\n \"type\": \"record\",\n \"name\": \"FullName\",\n \"fields\": [\n {\"name\": \"first\", \"type\": \"string\"},\n {\"name\": \"last\", \"type\": \"string\"}\n ]\n }},\n {\"name\": \"age\", \"type\": \"int\"},\n {\"name\": \"address\", \"type\": {\n \"type\": \"record\",\n \"name\": \"Address\",\n \"fields\": [\n {\"name\": \"street\", \"type\": \"string\"},\n {\"name\": \"city\", \"type\": \"string\"},\n {\"name\": \"state\", \"type\": \"string\"},\n {\"name\": \"zip\", \"type\": \"int\"}\n ]\n }}\n ]\n} \n{code}\n\n*Data*\n{code:javascript}\n{\n \"name\": {\n \"first\": \"Percival\",\n \"last\": \"Lowell\"\n },\n \"age\": 156,\n \"address\": {\n \"street\": \"Mars Hill Rd\",\n \"city\": \"Flagstaff\",\n \"state\": \"AZ\",\n \"zip\": 86001\n }\n}\n{code}\n\n*Create* (no partitions - works)\nIf I create with no partitions, I'm able to query the data just fine.\n\n{code:sql}\nCREATE EXTERNAL TABLE IF NOT EXISTS foo\nROW FORMAT SERDE 'org.apache.hadoop.hive.serde2.avro.AvroSerDe'\nSTORED AS INPUTFORMAT 'org.apache.hadoop.hive.ql.io.avro.AvroContainerInputFormat'\nOUTPUTFORMAT 'org.apache.hadoop.hive.ql.io.avro.AvroContainerOutputFormat'\nLOCATION '/path/to/data/dir'\nTBLPROPERTIES ('avro.schema.url'='/path/to/schema.avsc');\n{code}\n\n*Create* (partitions -- does NOT work)\nIf I create with no partitions, and then manually add a partition, all of my queries return an error. (I need to manually add partitions because I cannot control the structure of the data directories, so dynamic partitioning is not an option.)\n\n{code:sql}\nCREATE EXTERNAL TABLE IF NOT EXISTS foo\nPARTITIONED BY (ds STRING)\nROW FORMAT SERDE 'org.apache.hadoop.hive.serde2.avro.AvroSerDe'\nSTORED AS INPUTFORMAT 'org.apache.hadoop.hive.ql.io.avro.AvroContainerInputFormat'\nOUTPUTFORMAT 'org.apache.hadoop.hive.ql.io.avro.AvroContainerOutputFormat'\nTBLPROPERTIES ('avro.schema.url'='/path/to/schema.avsc');\n\nALTER TABLE foo ADD PARTITION (ds='1') LOCATION '/path/to/data/dir';\n{code}\n\nThe error:\n\n{code}\nspark-sql> SELECT * FROM foo WHERE ds = '1';\n\nDriver stacktrace:\n at org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1431)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1419)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1418)\n at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n at org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1418)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:799)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:799)\n at scala.Option.foreach(Option.scala:236)\n at org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:799)\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:1640)\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1599)\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1588)\n at org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\n at org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:620)\n at org.apache.spark.SparkContext.runJob(SparkContext.scala:1832)\n at org.apache.spark.SparkContext.runJob(SparkContext.scala:1845)\n at org.apache.spark.SparkContext.runJob(SparkContext.scala:1858)\n at org.apache.spark.SparkContext.runJob(SparkContext.scala:1929)\n at org.apache.spark.rdd.RDD$$anonfun$collect$1.apply(RDD.scala:927)\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:150)\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:111)\n at org.apache.spark.rdd.RDD.withScope(RDD.scala:316)\n at org.apache.spark.rdd.RDD.collect(RDD.scala:926)\n at org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:166)\n at org.apache.spark.sql.execution.SparkPlan.executeCollectPublic(SparkPlan.scala:174)\n at org.apache.spark.sql.hive.HiveContext$QueryExecution.stringResult(HiveContext.scala:635)\n at org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:64)\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:308)\n at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:376)\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$.main(SparkSQLCLIDriver.scala:226)\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.main(SparkSQLCLIDriver.scala)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:731)\n at org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:181)\n at org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:206)\n at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:121)\n at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\nCaused by: org.apache.avro.AvroTypeException: Found avro.FullName, expecting union\n at org.apache.avro.io.ResolvingDecoder.doAction(ResolvingDecoder.java:292)\n at org.apache.avro.io.parsing.Parser.advance(Parser.java:88)\n at org.apache.avro.io.ResolvingDecoder.readIndex(ResolvingDecoder.java:267)\n at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:155)\n at org.apache.avro.generic.GenericDatumReader.readField(GenericDatumReader.java:193)\n at org.apache.avro.generic.GenericDatumReader.readRecord(GenericDatumReader.java:183)\n at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:151)\n at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:142)\n at org.apache.hadoop.hive.serde2.avro.AvroDeserializer$SchemaReEncoder.reencode(AvroDeserializer.java:111)\n at org.apache.hadoop.hive.serde2.avro.AvroDeserializer.deserialize(AvroDeserializer.java:175)\n at org.apache.hadoop.hive.serde2.avro.AvroSerDe.deserialize(AvroSerDe.java:201)\n at org.apache.spark.sql.hive.HadoopTableReader$$anonfun$fillObject$2.apply(TableReader.scala:409)\n at org.apache.spark.sql.hive.HadoopTableReader$$anonfun$fillObject$2.apply(TableReader.scala:408)\n at scala.collection.Iterator$$anon$11.next(Iterator.scala:328)\n at scala.collection.Iterator$$anon$11.next(Iterator.scala:328)\n at scala.collection.Iterator$class.foreach(Iterator.scala:727)\n at scala.collection.AbstractIterator.foreach(Iterator.scala:1157)\n at scala.collection.generic.Growable$class.$plus$plus$eq(Growable.scala:48)\n at scala.collection.mutable.ArrayBuffer.$plus$plus$eq(ArrayBuffer.scala:103)\n at scala.collection.mutable.ArrayBuffer.$plus$plus$eq(ArrayBuffer.scala:47)\n at scala.collection.TraversableOnce$class.to(TraversableOnce.scala:273)\n at scala.collection.AbstractIterator.to(Iterator.scala:1157)\n at scala.collection.TraversableOnce$class.toBuffer(TraversableOnce.scala:265)\n at scala.collection.AbstractIterator.toBuffer(Iterator.scala:1157)\n at scala.collection.TraversableOnce$class.toArray(TraversableOnce.scala:252)\n at scala.collection.AbstractIterator.toArray(Iterator.scala:1157)\n at org.apache.spark.rdd.RDD$$anonfun$collect$1$$anonfun$12.apply(RDD.scala:927)\n at org.apache.spark.rdd.RDD$$anonfun$collect$1$$anonfun$12.apply(RDD.scala:927)\n at org.apache.spark.SparkContext$$anonfun$runJob$5.apply(SparkContext.scala:1858)\n at org.apache.spark.SparkContext$$anonfun$runJob$5.apply(SparkContext.scala:1858)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\n*Additional Info*\nFor what it's worth, I found an issue (DRILL-957) reported in Apache Drill and related fix that look very simliar to this. I'll look that to this issue.\n\nOriginally [posted here|http://stackoverflow.com/questions/35826850/spark-unable-to-decode-avro-when-partitioned] on StackOverflow as a question, but I felt strongly that this is indeed a bug so I created this issue.","from":"reporter","subject":"Spark unable to decode Avro when partitioned"},{"body":"Is similar to SPARK-13572","from":"developer"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13865","from":"developer"},{"body":"Issue resolved by pull request 13865\n[https://github.com/apache/spark/pull/13865]","from":"developer"}],"created":"2016-03-07T02:39:56.000+0000","description":"There is a problem decoding Avro data with SparkSQL when partitioned. The schema and encoded data are valid -- I'm able to decode the data with the avro-tools CLI utility. I'm also able to decode the data with non-partitioned SparkSQL tables, Hive, other tools as well... except partitioned SparkSQL schemas.\n\nFor a simple example, I took the example schema and data found in the Oracle documentation here:\n\n*Schema*\n{code:javascript}\n{\n \"type\": \"record\",\n \"name\": \"MemberInfo\",\n \"namespace\": \"avro\",\n \"fields\": [\n {\"name\": \"name\", \"type\": {\n \"type\": \"record\",\n \"name\": \"FullName\",\n \"fields\": [\n {\"name\": \"first\", \"type\": \"string\"},\n {\"name\": \"last\", \"type\": \"string\"}\n ]\n }},\n {\"name\": \"age\", \"type\": \"int\"},\n {\"name\": \"address\", \"type\": {\n \"type\": \"record\",\n \"name\": \"Address\",\n \"fields\": [\n {\"name\": \"street\", \"type\": \"string\"},\n {\"name\": \"city\", \"type\": \"string\"},\n {\"name\": \"state\", \"type\": \"string\"},\n {\"name\": \"zip\", \"type\": \"int\"}\n ]\n }}\n ]\n} \n{code}\n\n*Data*\n{code:javascript}\n{\n \"name\": {\n \"first\": \"Percival\",\n \"last\": \"Lowell\"\n },\n \"age\": 156,\n \"address\": {\n \"street\": \"Mars Hill Rd\",\n \"city\": \"Flagstaff\",\n \"state\": \"AZ\",\n \"zip\": 86001\n }\n}\n{code}\n\n*Create* (no partitions - works)\nIf I create with no partitions, I'm able to query the data just fine.\n\n{code:sql}\nCREATE EXTERNAL TABLE IF NOT EXISTS foo\nROW FORMAT SERDE 'org.apache.hadoop.hive.serde2.avro.AvroSerDe'\nSTORED AS INPUTFORMAT 'org.apache.hadoop.hive.ql.io.avro.AvroContainerInputFormat'\nOUTPUTFORMAT 'org.apache.hadoop.hive.ql.io.avro.AvroContainerOutputFormat'\nLOCATION '/path/to/data/dir'\nTBLPROPERTIES ('avro.schema.url'='/path/to/schema.avsc');\n{code}\n\n*Create* (partitions -- does NOT work)\nIf I create with no partitions, and then manually add a partition, all of my queries return an error. (I need to manually add partitions because I cannot control the structure of the data directories, so dynamic partitioning is not an option.)\n\n{code:sql}\nCREATE EXTERNAL TABLE IF NOT EXISTS foo\nPARTITIONED BY (ds STRING)\nROW FORMAT SERDE 'org.apache.hadoop.hive.serde2.avro.AvroSerDe'\nSTORED AS INPUTFORMAT 'org.apache.hadoop.hive.ql.io.avro.AvroContainerInputFormat'\nOUTPUTFORMAT 'org.apache.hadoop.hive.ql.io.avro.AvroContainerOutputFormat'\nTBLPROPERTIES ('avro.schema.url'='/path/to/schema.avsc');\n\nALTER TABLE foo ADD PARTITION (ds='1') LOCATION '/path/to/data/dir';\n{code}\n\nThe error:\n\n{code}\nspark-sql> SELECT * FROM foo WHERE ds = '1';\n\nDriver stacktrace:\n at org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1431)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1419)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1418)\n at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n at org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1418)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:799)\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:799)\n at scala.Option.foreach(Option.scala:236)\n at org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:799)\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:1640)\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1599)\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1588)\n at org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\n at org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:620)\n at org.apache.spark.SparkContext.runJob(SparkContext.scala:1832)\n at org.apache.spark.SparkContext.runJob(SparkContext.scala:1845)\n at org.apache.spark.SparkContext.runJob(SparkContext.scala:1858)\n at org.apache.spark.SparkContext.runJob(SparkContext.scala:1929)\n at org.apache.spark.rdd.RDD$$anonfun$collect$1.apply(RDD.scala:927)\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:150)\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:111)\n at org.apache.spark.rdd.RDD.withScope(RDD.scala:316)\n at org.apache.spark.rdd.RDD.collect(RDD.scala:926)\n at org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:166)\n at org.apache.spark.sql.execution.SparkPlan.executeCollectPublic(SparkPlan.scala:174)\n at org.apache.spark.sql.hive.HiveContext$QueryExecution.stringResult(HiveContext.scala:635)\n at org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:64)\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:308)\n at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:376)\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$.main(SparkSQLCLIDriver.scala:226)\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.main(SparkSQLCLIDriver.scala)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:731)\n at org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:181)\n at org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:206)\n at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:121)\n at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\nCaused by: org.apache.avro.AvroTypeException: Found avro.FullName, expecting union\n at org.apache.avro.io.ResolvingDecoder.doAction(ResolvingDecoder.java:292)\n at org.apache.avro.io.parsing.Parser.advance(Parser.java:88)\n at org.apache.avro.io.ResolvingDecoder.readIndex(ResolvingDecoder.java:267)\n at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:155)\n at org.apache.avro.generic.GenericDatumReader.readField(GenericDatumReader.java:193)\n at org.apache.avro.generic.GenericDatumReader.readRecord(GenericDatumReader.java:183)\n at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:151)\n at org.apache.avro.generic.GenericDatumReader.read(GenericDatumReader.java:142)\n at org.apache.hadoop.hive.serde2.avro.AvroDeserializer$SchemaReEncoder.reencode(AvroDeserializer.java:111)\n at org.apache.hadoop.hive.serde2.avro.AvroDeserializer.deserialize(AvroDeserializer.java:175)\n at org.apache.hadoop.hive.serde2.avro.AvroSerDe.deserialize(AvroSerDe.java:201)\n at org.apache.spark.sql.hive.HadoopTableReader$$anonfun$fillObject$2.apply(TableReader.scala:409)\n at org.apache.spark.sql.hive.HadoopTableReader$$anonfun$fillObject$2.apply(TableReader.scala:408)\n at scala.collection.Iterator$$anon$11.next(Iterator.scala:328)\n at scala.collection.Iterator$$anon$11.next(Iterator.scala:328)\n at scala.collection.Iterator$class.foreach(Iterator.scala:727)\n at scala.collection.AbstractIterator.foreach(Iterator.scala:1157)\n at scala.collection.generic.Growable$class.$plus$plus$eq(Growable.scala:48)\n at scala.collection.mutable.ArrayBuffer.$plus$plus$eq(ArrayBuffer.scala:103)\n at scala.collection.mutable.ArrayBuffer.$plus$plus$eq(ArrayBuffer.scala:47)\n at scala.collection.TraversableOnce$class.to(TraversableOnce.scala:273)\n at scala.collection.AbstractIterator.to(Iterator.scala:1157)\n at scala.collection.TraversableOnce$class.toBuffer(TraversableOnce.scala:265)\n at scala.collection.AbstractIterator.toBuffer(Iterator.scala:1157)\n at scala.collection.TraversableOnce$class.toArray(TraversableOnce.scala:252)\n at scala.collection.AbstractIterator.toArray(Iterator.scala:1157)\n at org.apache.spark.rdd.RDD$$anonfun$collect$1$$anonfun$12.apply(RDD.scala:927)\n at org.apache.spark.rdd.RDD$$anonfun$collect$1$$anonfun$12.apply(RDD.scala:927)\n at org.apache.spark.SparkContext$$anonfun$runJob$5.apply(SparkContext.scala:1858)\n at org.apache.spark.SparkContext$$anonfun$runJob$5.apply(SparkContext.scala:1858)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\n*Additional Info*\nFor what it's worth, I found an issue (DRILL-957) reported in Apache Drill and related fix that look very simliar to this. I'll look that to this issue.\n\nOriginally [posted here|http://stackoverflow.com/questions/35826850/spark-unable-to-decode-avro-when-partitioned] on StackOverflow as a question, but I felt strongly that this is indeed a bug so I created this issue.","issue_id":"12947600","key":"SPARK-13709","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-06-24T06:12:51.000+0000","role":"fixed_distractor","summary":"Spark unable to decode Avro when partitioned"} {"case_id":"12948114","cluster":"DISTRACTOR-SPARK-13747","comments":[{"body":"FYI, I switched to branch-1.6, and ran the same example. It will fail with StackOverflow because it submits too many blocking tasks to ForkJoinPool.\n\nTherefore, I'm not sure if it's worth to fix it. In general, the user should not submit many blocking tasks to ForkJoinPool otherwise StackOverflow will happen.","created":"2016-03-08T19:09:50.929+0000"},{"body":"User 'andrewor14' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11586","created":"2016-03-08T20:39:06.326+0000"},{"body":"Looks like it still exists in 2.0.0? Will try to add a relevant test-case, soon","created":"2016-09-27T12:11:23.754+0000"},{"body":"I encounter this in 2.0.1, is there any workaround like having separate SparkSession will help?","created":"2016-10-14T07:16:11.677+0000"},{"body":"[~chinwei] could you post the stack trace here?","created":"2016-10-16T04:16:45.741+0000"},{"body":"java.lang.IllegalArgumentException: spark.sql.execution.id is already set\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2546) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2192) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2199) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset$$anonfun$count$1.apply(Dataset.scala:2227) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset$$anonfun$count$1.apply(Dataset.scala:2226) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.withCallback(Dataset.scala:2559) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.count(Dataset.scala:2226) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n","created":"2016-10-17T02:33:21.036+0000"},{"body":"Could you post the full stack track, please? It would be helpful to know who calls `count`.","created":"2016-10-17T03:38:30.852+0000"},{"body":"It is running on Akka, with forkjoin dispatcher. There are 2 actors running concurrently doing different Spark job using the same SparkSession. I can't give the full stack, but here is the outline:\n\njava.lang.IllegalArgumentException: spark.sql.execution.id is already set\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2546) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2192) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2199) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset$$anonfun$count$1.apply(Dataset.scala:2227) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset$$anonfun$count$1.apply(Dataset.scala:2226) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.withCallback(Dataset.scala:2559) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.count(Dataset.scala:2226) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n\n<-- Here is the code that call the df.count -->\n\nat akka.actor.ActorCell.receiveMessage(ActorCell.scala:526) [akka-actor_2.11-2.4.8.jar:na]\n at akka.actor.ActorCell.invoke(ActorCell.scala:495) [akka-actor_2.11-2.4.8.jar:na]\n at akka.dispatch.Mailbox.processMailbox(Mailbox.scala:257) [akka-actor_2.11-2.4.8.jar:na]\n at akka.dispatch.Mailbox.run(Mailbox.scala:224) [akka-actor_2.11-2.4.8.jar:na]\n at akka.dispatch.Mailbox.exec(Mailbox.scala:234) [akka-actor_2.11-2.4.8.jar:na]\n at scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260) [scala-library-2.11.8.jar:na]\n at scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339) [scala-library-2.11.8.jar:na]\n at scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979) [scala-library-2.11.8.jar:na]\n at scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107) [scala-library-2.11.8.jar:na]\n\n\n","created":"2016-10-17T03:48:02.555+0000"},{"body":"There are other places need to be fixed.","created":"2016-10-17T20:57:54.316+0000"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15520","created":"2016-10-17T21:03:04.764+0000"},{"body":"[~chinwei] Could you test https://github.com/apache/spark/pull/15520 and see if the error is gone?","created":"2016-10-17T21:03:09.872+0000"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15646","created":"2016-10-26T17:44:06.052+0000"},{"body":"This issue still exists. Reopened it.","created":"2016-12-09T07:23:47.914+0000"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16230","created":"2016-12-09T07:32:08.138+0000"},{"body":"Issue resolved by pull request 16230\n[https://github.com/apache/spark/pull/16230]","created":"2016-12-13T17:54:03.094+0000"},{"body":"We have hit this on rare instances in our production environment when calling tableNames on SQLContext. We are using Spark 2.1.0. Are there any possible workarounds that we might try? What is the ETA of spark 2.2?\n\n{code}\nspark.sql.execution.id is already set\norg.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81)\norg.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2765)\norg.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2370)\norg.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\norg.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\norg.apache.spark.sql.Dataset.withCallback(Dataset.scala:2778)\norg.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2375)\norg.apache.spark.sql.Dataset.collect(Dataset.scala:2351)\norg.apache.spark.sql.SQLContext.tableNames(SQLContext.scala:750)\n{code}","created":"2017-03-22T19:49:31.334+0000"},{"body":"I am also running into this issue *sporadically* when collecting the results of joining two dataframes.\n\nThe code that *sporadically* generates this issue is:\n{code}\nval spark = SparkSession.builder().appName(\"application\").master(\"local[*]\").getOrCreate()\nval itemCountry = spark.read.format(\"csv\")\n .option(\"header\", \"true\")\n .schema(StructType(Array(\n StructField(\"itemId\", IntegerType, false),\n StructField(\"countryId\", IntegerType, false))))\n .csv(\"/item_country.csv\") // This file matches the schema provided\nval itemPerformance = spark.read.format(\"csv\")\n .option(\"header\", \"true\")\n .schema(StructType(Array(\n StructField(\"itemId\", IntegerType, false),\n StructField(\"date\", TimestampType, false),\n StructField(\"performance\", IntegerType, false))))\n .csv(\"/item_performance.csv\") // This file matches the schema provided\n\nitemCountry.join(itemPerformance, itemCountry(\"itemId\") === itemPerformance(\"itemId\"))\n .groupBy(\"countryId\")\n .agg(sum(when(to_date(itemPerformance(\"date\")) > to_date(lit(\"2017-01-01\")), itemPerformance(\"performance\")).otherwise(0)).alias(\"performance\")).show()\n{code}\n\nThe stack trace is:\n{code}\njava.lang.IllegalArgumentException: spark.sql.execution.id is already set\nat org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81)\nat org.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2765)\nat org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2370)\nat org.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\nat org.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\nat org.apache.spark.sql.Dataset.withCallback(Dataset.scala:2778)\nat org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2375)\nat org.apache.spark.sql.Dataset.collect(Dataset.scala:2351)\nat .... [Custom caller functions]\n{code}","created":"2017-04-20T12:17:22.930+0000"},{"body":"[~mousa] could you try the master branch? This issue will be fixed in Spark 2.2.0.","created":"2017-04-20T17:45:41.334+0000"},{"body":"[~zsxwing] I did a similar test with join and have the same error in 2.2.0 (actual query here - https://github.com/dnaumenko/spark-realtime-analytics-sample/blob/master/samples/src/main/scala/com/github/sparksample/httpapp/SimpleServer.scala). \n\nMy test setup is a simple akka-http long-running application and separate Gatling script that spawns multiple requests for join query (https://github.com/dnaumenko/spark-realtime-analytics-sample/blob/master/loadtool/src/main/scala/com/github/sparksample/SimpleServerSimulation.scala is test simulation script).\n\n[~barrybecker4] I've managed to fix the problem by switching akka's default executor to thread-pool. But I guess the root cause is that Spark is relying on ThreadLocal variables and manages them incorrectly.","created":"2017-04-25T14:24:49.800+0000"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17763","created":"2017-04-25T21:07:03.614+0000"},{"body":"[~dnaumenko] Unfortunately, Spark uses ThreadLocal variables a lot but ForkJoinPool doesn't support that very well (It's easy to leak ThreadLocal variables to other tasks). Could you check if https://github.com/apache/spark/pull/17763 can fix your issue?","created":"2017-04-25T21:08:32.780+0000"},{"body":"Thanks to [~dnaumenko], we also managed to avoid running into this issue by switching akka’s default executor to “thread-pool-executor”. We are also using akka a lot in our system that is creating Spark’s jobs.\nWhat I need to make sure of that, is this issue really a Spark issue or a fork-join framework issue? As far as I understood, a thread can be reused by its (or another) process before completely finishing its work to perform another work which would pollute its locals. If this is the case then the solution proposed in the pull request https://github.com/apache/spark/pull/17763 won’t resolve this issue as it simply, if I fully got the proposed solution, prevents timeout exceptions from being thrown when calling \"Await.ready\".","created":"2017-05-08T14:14:58.304+0000"},{"body":"There seems to be some related discussion here\nhttp://apache-spark-developers-list.1001551.n3.nabble.com/IllegalArgumentException-spark-sql-execution-id-is-already-set-td19124.html\nWe use job-server (2.0-preview branch version) which uses akka. I believe spark does not. Maybe that is why we periodically see this issue (not sure). How can I switch akka's default executor to be \"thread-pool-executor\"? Is it a config option somewhere?","created":"2017-05-08T14:31:10.691+0000"},{"body":"You can override akka's default executor for your application by adding the following configuration into an \"application.conf\" file that should be in the root of the class path.\n{code}\nakka {\n actor {\n default-dispatcher {\n # Which kind of ExecutorService to use for this dispatcher\n # Valid options:\n # - \"default-executor\" requires a \"default-executor\" section\n # - \"fork-join-executor\" requires a \"fork-join-executor\" section\n # - \"thread-pool-executor\" requires a \"thread-pool-executor\" section\n # - A FQCN of a class extending ExecutorServiceConfigurator\n executor = \"thread-pool-executor\"\n }\n }\n}\n{code}\n\nFurther information can be found at: [http://doc.akka.io/docs/akka/current/general/configuration.html] and [http://doc.akka.io/docs/akka/current/scala/dispatchers.html]","created":"2017-05-08T14:59:48.073+0000"},{"body":"The above solution did help my program to start some additional threads, but it is still failing in some of them.\n\njava.lang.IllegalArgumentException: spark.sql.execution.id is already set\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81)\n\nAlthough I am using akka-http to start a server, I am just executing spark actions inside traditional Futures.\n\nSpark data calls are called within a DSL Routing 'get' requests.\n\nusing 2.1.0 in scala 2.11","created":"2017-05-08T17:14:09.421+0000"},{"body":"I also tried the \"thread-pool-executor\" workaround suggested above, but adding the suggested json to the top level of spark job-servers local.conf file. I still saw the error. The error is difficult to reproduce reliably, but I did see it once after making the change.","created":"2017-05-08T17:47:01.932+0000"},{"body":"@saif1988, just to clarify, did you add the following?\ndefault-dispatcher {\n executor = \"default-executor\"\n}\nHow do you know for sure that it fixes the problem? Did you have a case that reliably reproduced it? My problem is that it is very rare. I can know if it still happens, but absence of the error does not tell me that its truly fixed.\n","created":"2017-05-08T18:39:04.407+0000"},{"body":"Sorry for the confusion. No, it doesnt work. I am currently trying out with using different execution contexts.\n\nMy issue happens always, it is 100% reproducible.\n\nTo simplify what I am doing:\n\n1. akka-http server is started and REST DSL is setup\n2. Inside a get dsl, I call a Spark dataframe which calls collect action from within a Future\n3. the object containing the future calls await.result, as I need the dataframe to respond a 200 to http\n4. the collect method is passed through as an annonymous function. runtime exception poinst at such annonymous function as the callback starter of my exception\n5. The future call is handled by a thread pool manged by spark pool. Uses FAIR scheduling.\n\nWhen my website starts, 4 collects are called simoultaneously. Only one get call returns 200. The others are internal server errors.","created":"2017-05-08T19:10:36.524+0000"},{"body":"[~mousa] This is is because Spark uses ThreadLocal in a fork-join pool. Let me try to clarify the issue.\n\nA fork-join pool allows to run another pending task in the same thread when a running task is calling Await.ready/result. The magic is when someone calls Await.ready/result, it will first check if there is any pending task submitted to the pool, if so, it will call the pending task instead of waiting. (See scala.concurrent#blocking)\n\nIn Spark, the codes hitting this issue have the following pattern.\n\n{code}\ntry {\n check if a thread local is set\n if so, throw an exception\n else\n set the thread local value\n do some work\n Call Await.ready/result to wait for a result // This doesn't clear the thread local value. \n // If the fork-join pool schedules a pending task here,\n // it will see the thread local value.\n do some work\n} finally {\n clear the thread local value.\n}\n{code}\n\nMy PR is basically just not calling `scala.concurrent#blocking`. It just makes a fork-join pool become a normal thread pool executor.\n","created":"2017-05-08T19:15:49.081+0000"},{"body":"I did fix it now 100% sure.\n\n1. Used thread-pool-executor\n2. Replaced implicit global execution context with a FixedThreadPool executor (of size 16 on my case).\n\nI have no idea whether I am doing things properly, but it works flawlessly now.","created":"2017-05-08T19:18:16.658+0000"},{"body":"Good to hear that your workaround was successful. How did you do step 2? (Replaced implicit global execution context with a FixedThreadPool executor) Is that in the json configuration or in the code?","created":"2017-05-08T19:26:58.586+0000"},{"body":"[~revolucion09] If you are not using ForkJoinPool, I'm 100% sure you won't hit this issue.","created":"2017-05-08T19:27:20.213+0000"},{"body":"[~zsxwing] I may be confused then. Which issue I am hitting if not this one? You are correct though, I am not explicitly using ForkJoinPool anywhere on my code.\n\n\n[ERROR] [05/08/2017 13:25:27.735] [Sake-akka.actor.default-dispatcher-3] [akka.actor.ActorSystemImpl(Sake)] Error during processing of request: 'spark.sql.execution.id is already set'. Completing with 500 Internal Server Error response. To change default exception handling behavior, provide a custom ExceptionHandler.\njava.lang.IllegalArgumentException: spark.sql.execution.id is already set\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81)\n at org.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2765)\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2370)\n at org.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\n at org.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\n at org.apache.spark.sql.Dataset.withCallback(Dataset.scala:2778)\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2375)\n at org.apache.spark.sql.Dataset.collect(Dataset.scala:2351)\n at org.cmortech.sake.publisher.Descriptor$$anonfun$11.apply(Descriptor.scala:202)\n at org.cmortech.sake.publisher.Descriptor$$anonfun$11.apply(Descriptor.scala:202)\n at org.cmortech.sake.water.entities.orphanage.Orphan$$anonfun$start$1.apply(Orphan.scala:56)\n at scala.concurrent.impl.Future$PromiseCompletingRunnable.liftedTree1$1(Future.scala:24)\n at scala.concurrent.impl.Future$PromiseCompletingRunnable.run(Future.scala:24)\n at scala.concurrent.impl.ExecutionContextImpl$AdaptedForkJoinTask.exec(ExecutionContextImpl.scala:121)\n at scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n at scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n at scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n at scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n","created":"2017-05-08T19:32:43.387+0000"},{"body":"[~revolucion09] I don't know who created ForkJoinPool but your stack trace says it's inside ForkJoinPool.","created":"2017-05-08T19:35:52.006+0000"},{"body":"Okey, I see, the thread name is \"Sake-akka.actor.default-dispatcher-3\", so I think it's just Akka default-dispatcher.","created":"2017-05-08T19:39:17.945+0000"},{"body":"[~zsxwing] Thanks. My program is pretty much default. Could it be that because I am running inside akka-http dsl, and just by using ExecutionContext.implicits.Global, it does spawn everything on a ForkJoinPool? \n\nGood thing is that I could resolve it by creating my own Executors.\n\n[~barrybecker4]\n private val pool = Executors.newFixedThreadPool(16)\n implicit private val ec = ExecutionContext.fromExecutorService(pool)","created":"2017-05-08T19:42:31.131+0000"},{"body":"[~revolucion09] The default dispatcher uses ForkJoinPool. See http://doc.akka.io/docs/akka/current/scala/dispatchers.html#Default_dispatcher ","created":"2017-05-08T19:45:18.442+0000"},{"body":"[~zsxwing] I've tried to build a Spark from a branch in pull request. Didn't manage to make a complete build (had some problems with R dependencies), so I've replaced only spark-core.jar and it seems like the issue still occurs. Could you please provide a jar/dist for re-test? \n\nAs a side note, the fixed-thread-pool solution works well for us. We will probably stick with it.","created":"2017-05-09T13:34:02.008+0000"},{"body":"[~dnaumenko]\nNonetheless, if I am not mistaken, there are proofs that fork join pools provide significant performance boost in scalable environments, that is why akka uses them by default. Fixed or Cached pool threads are considered dangerous for production environments.","created":"2017-05-09T14:50:25.788+0000"},{"body":"[~revolucion09] It's a bit an off-topic discussion. My 50 cents to it - from my understanding, fork-join pool in Akka helps to keep all processors busy, so you can archive a high message throughput per second (as long as you don't use blocking operations). But in Spark driver program, it's not a critical, cause the most time it will be waiting for worker nodes anyway. ","created":"2017-05-09T15:17:02.170+0000"},{"body":"I am having the same exception.\n\nI am creating a new data source that processes reading batch files asynchronously into a temp folder and then returns them as a data frame.\n\nWithin the buildScan(): RDD[Row] method I have a loop that saves the results of each batch in a parquet file:\n\n val df = spark.sparkContext.parallelize(batchResult.records, 200).toDF() \n df.write.mode(SaveMode.Overwrite).save(tempFile)}\n\nThen once the temp files are all written, buildScan method returns \nI will load all those temp files in parallel and return the union in an RDD like this:\nsqlContext.read\n .schema(schema)\n .load(files: _*) \n .queryExecution.executedPlan. execute().asInstanceOf[RDD[Row]]\n\nI can see the concurrency issue as I am trying to write the temp files at same time I am trying to construct a return RDD. \nIs there a better way of doing this?\nTo work around, I can save the temp files as regular CSV to work around the issue but I prefer to save these files as Parquet files using Spark API.\nWould upgrading to 2.12 fix this issue in my case?\n\n","created":"2017-05-12T18:26:25.032+0000"}],"conversations":[{"body":"Run the following codes may fail\n{code}\n(1 to 100).par.foreach { _ =>\n println(sc.parallelize(1 to 5).map { i => (i, i) }.toDF(\"a\", \"b\").count())\n}\n\njava.lang.IllegalArgumentException: spark.sql.execution.id is already set \n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:87) \n at org.apache.spark.sql.DataFrame.withNewExecutionId(DataFrame.scala:1904) \n at org.apache.spark.sql.DataFrame.collect(DataFrame.scala:1385) \n{code}\n\nThis is because SparkContext.runJob can be suspended when using a ForkJoinPool (e.g.,scala.concurrent.ExecutionContext.Implicits.global) as it calls Await.ready (introduced by https://github.com/apache/spark/pull/9264).\n\nSo when SparkContext.runJob is suspended, ForkJoinPool will run another task in the same thread, however, the local properties has been polluted.","from":"reporter","subject":"Concurrent execution in SQL doesn't work with Scala ForkJoinPool"},{"body":"FYI, I switched to branch-1.6, and ran the same example. It will fail with StackOverflow because it submits too many blocking tasks to ForkJoinPool.\n\nTherefore, I'm not sure if it's worth to fix it. In general, the user should not submit many blocking tasks to ForkJoinPool otherwise StackOverflow will happen.","from":"developer"},{"body":"User 'andrewor14' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11586","from":"developer"},{"body":"Looks like it still exists in 2.0.0? Will try to add a relevant test-case, soon","from":"developer"},{"body":"I encounter this in 2.0.1, is there any workaround like having separate SparkSession will help?","from":"developer"},{"body":"[~chinwei] could you post the stack trace here?","from":"developer"},{"body":"java.lang.IllegalArgumentException: spark.sql.execution.id is already set\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2546) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2192) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2199) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset$$anonfun$count$1.apply(Dataset.scala:2227) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset$$anonfun$count$1.apply(Dataset.scala:2226) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.withCallback(Dataset.scala:2559) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.count(Dataset.scala:2226) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n","from":"developer"},{"body":"Could you post the full stack track, please? It would be helpful to know who calls `count`.","from":"developer"},{"body":"It is running on Akka, with forkjoin dispatcher. There are 2 actors running concurrently doing different Spark job using the same SparkSession. I can't give the full stack, but here is the outline:\n\njava.lang.IllegalArgumentException: spark.sql.execution.id is already set\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2546) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2192) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2199) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset$$anonfun$count$1.apply(Dataset.scala:2227) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset$$anonfun$count$1.apply(Dataset.scala:2226) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.withCallback(Dataset.scala:2559) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n at org.apache.spark.sql.Dataset.count(Dataset.scala:2226) ~[spark-sql_2.11-2.0.1.jar:2.0.1]\n\n<-- Here is the code that call the df.count -->\n\nat akka.actor.ActorCell.receiveMessage(ActorCell.scala:526) [akka-actor_2.11-2.4.8.jar:na]\n at akka.actor.ActorCell.invoke(ActorCell.scala:495) [akka-actor_2.11-2.4.8.jar:na]\n at akka.dispatch.Mailbox.processMailbox(Mailbox.scala:257) [akka-actor_2.11-2.4.8.jar:na]\n at akka.dispatch.Mailbox.run(Mailbox.scala:224) [akka-actor_2.11-2.4.8.jar:na]\n at akka.dispatch.Mailbox.exec(Mailbox.scala:234) [akka-actor_2.11-2.4.8.jar:na]\n at scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260) [scala-library-2.11.8.jar:na]\n at scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339) [scala-library-2.11.8.jar:na]\n at scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979) [scala-library-2.11.8.jar:na]\n at scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107) [scala-library-2.11.8.jar:na]\n\n\n","from":"developer"},{"body":"There are other places need to be fixed.","from":"developer"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15520","from":"developer"},{"body":"[~chinwei] Could you test https://github.com/apache/spark/pull/15520 and see if the error is gone?","from":"developer"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15646","from":"developer"},{"body":"This issue still exists. Reopened it.","from":"developer"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16230","from":"developer"},{"body":"Issue resolved by pull request 16230\n[https://github.com/apache/spark/pull/16230]","from":"developer"},{"body":"We have hit this on rare instances in our production environment when calling tableNames on SQLContext. We are using Spark 2.1.0. Are there any possible workarounds that we might try? What is the ETA of spark 2.2?\n\n{code}\nspark.sql.execution.id is already set\norg.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81)\norg.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2765)\norg.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2370)\norg.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\norg.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\norg.apache.spark.sql.Dataset.withCallback(Dataset.scala:2778)\norg.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2375)\norg.apache.spark.sql.Dataset.collect(Dataset.scala:2351)\norg.apache.spark.sql.SQLContext.tableNames(SQLContext.scala:750)\n{code}","from":"developer"},{"body":"I am also running into this issue *sporadically* when collecting the results of joining two dataframes.\n\nThe code that *sporadically* generates this issue is:\n{code}\nval spark = SparkSession.builder().appName(\"application\").master(\"local[*]\").getOrCreate()\nval itemCountry = spark.read.format(\"csv\")\n .option(\"header\", \"true\")\n .schema(StructType(Array(\n StructField(\"itemId\", IntegerType, false),\n StructField(\"countryId\", IntegerType, false))))\n .csv(\"/item_country.csv\") // This file matches the schema provided\nval itemPerformance = spark.read.format(\"csv\")\n .option(\"header\", \"true\")\n .schema(StructType(Array(\n StructField(\"itemId\", IntegerType, false),\n StructField(\"date\", TimestampType, false),\n StructField(\"performance\", IntegerType, false))))\n .csv(\"/item_performance.csv\") // This file matches the schema provided\n\nitemCountry.join(itemPerformance, itemCountry(\"itemId\") === itemPerformance(\"itemId\"))\n .groupBy(\"countryId\")\n .agg(sum(when(to_date(itemPerformance(\"date\")) > to_date(lit(\"2017-01-01\")), itemPerformance(\"performance\")).otherwise(0)).alias(\"performance\")).show()\n{code}\n\nThe stack trace is:\n{code}\njava.lang.IllegalArgumentException: spark.sql.execution.id is already set\nat org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81)\nat org.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2765)\nat org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2370)\nat org.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\nat org.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\nat org.apache.spark.sql.Dataset.withCallback(Dataset.scala:2778)\nat org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2375)\nat org.apache.spark.sql.Dataset.collect(Dataset.scala:2351)\nat .... [Custom caller functions]\n{code}","from":"developer"},{"body":"[~mousa] could you try the master branch? This issue will be fixed in Spark 2.2.0.","from":"developer"},{"body":"[~zsxwing] I did a similar test with join and have the same error in 2.2.0 (actual query here - https://github.com/dnaumenko/spark-realtime-analytics-sample/blob/master/samples/src/main/scala/com/github/sparksample/httpapp/SimpleServer.scala). \n\nMy test setup is a simple akka-http long-running application and separate Gatling script that spawns multiple requests for join query (https://github.com/dnaumenko/spark-realtime-analytics-sample/blob/master/loadtool/src/main/scala/com/github/sparksample/SimpleServerSimulation.scala is test simulation script).\n\n[~barrybecker4] I've managed to fix the problem by switching akka's default executor to thread-pool. But I guess the root cause is that Spark is relying on ThreadLocal variables and manages them incorrectly.","from":"developer"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17763","from":"developer"},{"body":"[~dnaumenko] Unfortunately, Spark uses ThreadLocal variables a lot but ForkJoinPool doesn't support that very well (It's easy to leak ThreadLocal variables to other tasks). Could you check if https://github.com/apache/spark/pull/17763 can fix your issue?","from":"developer"},{"body":"Thanks to [~dnaumenko], we also managed to avoid running into this issue by switching akka’s default executor to “thread-pool-executor”. We are also using akka a lot in our system that is creating Spark’s jobs.\nWhat I need to make sure of that, is this issue really a Spark issue or a fork-join framework issue? As far as I understood, a thread can be reused by its (or another) process before completely finishing its work to perform another work which would pollute its locals. If this is the case then the solution proposed in the pull request https://github.com/apache/spark/pull/17763 won’t resolve this issue as it simply, if I fully got the proposed solution, prevents timeout exceptions from being thrown when calling \"Await.ready\".","from":"developer"},{"body":"There seems to be some related discussion here\nhttp://apache-spark-developers-list.1001551.n3.nabble.com/IllegalArgumentException-spark-sql-execution-id-is-already-set-td19124.html\nWe use job-server (2.0-preview branch version) which uses akka. I believe spark does not. Maybe that is why we periodically see this issue (not sure). How can I switch akka's default executor to be \"thread-pool-executor\"? Is it a config option somewhere?","from":"developer"},{"body":"You can override akka's default executor for your application by adding the following configuration into an \"application.conf\" file that should be in the root of the class path.\n{code}\nakka {\n actor {\n default-dispatcher {\n # Which kind of ExecutorService to use for this dispatcher\n # Valid options:\n # - \"default-executor\" requires a \"default-executor\" section\n # - \"fork-join-executor\" requires a \"fork-join-executor\" section\n # - \"thread-pool-executor\" requires a \"thread-pool-executor\" section\n # - A FQCN of a class extending ExecutorServiceConfigurator\n executor = \"thread-pool-executor\"\n }\n }\n}\n{code}\n\nFurther information can be found at: [http://doc.akka.io/docs/akka/current/general/configuration.html] and [http://doc.akka.io/docs/akka/current/scala/dispatchers.html]","from":"developer"},{"body":"The above solution did help my program to start some additional threads, but it is still failing in some of them.\n\njava.lang.IllegalArgumentException: spark.sql.execution.id is already set\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81)\n\nAlthough I am using akka-http to start a server, I am just executing spark actions inside traditional Futures.\n\nSpark data calls are called within a DSL Routing 'get' requests.\n\nusing 2.1.0 in scala 2.11","from":"developer"},{"body":"I also tried the \"thread-pool-executor\" workaround suggested above, but adding the suggested json to the top level of spark job-servers local.conf file. I still saw the error. The error is difficult to reproduce reliably, but I did see it once after making the change.","from":"developer"},{"body":"@saif1988, just to clarify, did you add the following?\ndefault-dispatcher {\n executor = \"default-executor\"\n}\nHow do you know for sure that it fixes the problem? Did you have a case that reliably reproduced it? My problem is that it is very rare. I can know if it still happens, but absence of the error does not tell me that its truly fixed.\n","from":"developer"},{"body":"Sorry for the confusion. No, it doesnt work. I am currently trying out with using different execution contexts.\n\nMy issue happens always, it is 100% reproducible.\n\nTo simplify what I am doing:\n\n1. akka-http server is started and REST DSL is setup\n2. Inside a get dsl, I call a Spark dataframe which calls collect action from within a Future\n3. the object containing the future calls await.result, as I need the dataframe to respond a 200 to http\n4. the collect method is passed through as an annonymous function. runtime exception poinst at such annonymous function as the callback starter of my exception\n5. The future call is handled by a thread pool manged by spark pool. Uses FAIR scheduling.\n\nWhen my website starts, 4 collects are called simoultaneously. Only one get call returns 200. The others are internal server errors.","from":"developer"},{"body":"[~mousa] This is is because Spark uses ThreadLocal in a fork-join pool. Let me try to clarify the issue.\n\nA fork-join pool allows to run another pending task in the same thread when a running task is calling Await.ready/result. The magic is when someone calls Await.ready/result, it will first check if there is any pending task submitted to the pool, if so, it will call the pending task instead of waiting. (See scala.concurrent#blocking)\n\nIn Spark, the codes hitting this issue have the following pattern.\n\n{code}\ntry {\n check if a thread local is set\n if so, throw an exception\n else\n set the thread local value\n do some work\n Call Await.ready/result to wait for a result // This doesn't clear the thread local value. \n // If the fork-join pool schedules a pending task here,\n // it will see the thread local value.\n do some work\n} finally {\n clear the thread local value.\n}\n{code}\n\nMy PR is basically just not calling `scala.concurrent#blocking`. It just makes a fork-join pool become a normal thread pool executor.\n","from":"developer"},{"body":"I did fix it now 100% sure.\n\n1. Used thread-pool-executor\n2. Replaced implicit global execution context with a FixedThreadPool executor (of size 16 on my case).\n\nI have no idea whether I am doing things properly, but it works flawlessly now.","from":"developer"},{"body":"Good to hear that your workaround was successful. How did you do step 2? (Replaced implicit global execution context with a FixedThreadPool executor) Is that in the json configuration or in the code?","from":"developer"},{"body":"[~revolucion09] If you are not using ForkJoinPool, I'm 100% sure you won't hit this issue.","from":"developer"},{"body":"[~zsxwing] I may be confused then. Which issue I am hitting if not this one? You are correct though, I am not explicitly using ForkJoinPool anywhere on my code.\n\n\n[ERROR] [05/08/2017 13:25:27.735] [Sake-akka.actor.default-dispatcher-3] [akka.actor.ActorSystemImpl(Sake)] Error during processing of request: 'spark.sql.execution.id is already set'. Completing with 500 Internal Server Error response. To change default exception handling behavior, provide a custom ExceptionHandler.\njava.lang.IllegalArgumentException: spark.sql.execution.id is already set\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:81)\n at org.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2765)\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2370)\n at org.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\n at org.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$collect$1.apply(Dataset.scala:2375)\n at org.apache.spark.sql.Dataset.withCallback(Dataset.scala:2778)\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2375)\n at org.apache.spark.sql.Dataset.collect(Dataset.scala:2351)\n at org.cmortech.sake.publisher.Descriptor$$anonfun$11.apply(Descriptor.scala:202)\n at org.cmortech.sake.publisher.Descriptor$$anonfun$11.apply(Descriptor.scala:202)\n at org.cmortech.sake.water.entities.orphanage.Orphan$$anonfun$start$1.apply(Orphan.scala:56)\n at scala.concurrent.impl.Future$PromiseCompletingRunnable.liftedTree1$1(Future.scala:24)\n at scala.concurrent.impl.Future$PromiseCompletingRunnable.run(Future.scala:24)\n at scala.concurrent.impl.ExecutionContextImpl$AdaptedForkJoinTask.exec(ExecutionContextImpl.scala:121)\n at scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n at scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n at scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n at scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n","from":"developer"},{"body":"[~revolucion09] I don't know who created ForkJoinPool but your stack trace says it's inside ForkJoinPool.","from":"developer"},{"body":"Okey, I see, the thread name is \"Sake-akka.actor.default-dispatcher-3\", so I think it's just Akka default-dispatcher.","from":"developer"},{"body":"[~zsxwing] Thanks. My program is pretty much default. Could it be that because I am running inside akka-http dsl, and just by using ExecutionContext.implicits.Global, it does spawn everything on a ForkJoinPool? \n\nGood thing is that I could resolve it by creating my own Executors.\n\n[~barrybecker4]\n private val pool = Executors.newFixedThreadPool(16)\n implicit private val ec = ExecutionContext.fromExecutorService(pool)","from":"developer"},{"body":"[~revolucion09] The default dispatcher uses ForkJoinPool. See http://doc.akka.io/docs/akka/current/scala/dispatchers.html#Default_dispatcher ","from":"developer"},{"body":"[~zsxwing] I've tried to build a Spark from a branch in pull request. Didn't manage to make a complete build (had some problems with R dependencies), so I've replaced only spark-core.jar and it seems like the issue still occurs. Could you please provide a jar/dist for re-test? \n\nAs a side note, the fixed-thread-pool solution works well for us. We will probably stick with it.","from":"developer"},{"body":"[~dnaumenko]\nNonetheless, if I am not mistaken, there are proofs that fork join pools provide significant performance boost in scalable environments, that is why akka uses them by default. Fixed or Cached pool threads are considered dangerous for production environments.","from":"developer"},{"body":"[~revolucion09] It's a bit an off-topic discussion. My 50 cents to it - from my understanding, fork-join pool in Akka helps to keep all processors busy, so you can archive a high message throughput per second (as long as you don't use blocking operations). But in Spark driver program, it's not a critical, cause the most time it will be waiting for worker nodes anyway. ","from":"developer"},{"body":"I am having the same exception.\n\nI am creating a new data source that processes reading batch files asynchronously into a temp folder and then returns them as a data frame.\n\nWithin the buildScan(): RDD[Row] method I have a loop that saves the results of each batch in a parquet file:\n\n val df = spark.sparkContext.parallelize(batchResult.records, 200).toDF() \n df.write.mode(SaveMode.Overwrite).save(tempFile)}\n\nThen once the temp files are all written, buildScan method returns \nI will load all those temp files in parallel and return the union in an RDD like this:\nsqlContext.read\n .schema(schema)\n .load(files: _*) \n .queryExecution.executedPlan. execute().asInstanceOf[RDD[Row]]\n\nI can see the concurrency issue as I am trying to write the temp files at same time I am trying to construct a return RDD. \nIs there a better way of doing this?\nTo work around, I can save the temp files as regular CSV to work around the issue but I prefer to save these files as Parquet files using Spark API.\nWould upgrading to 2.12 fix this issue in my case?\n\n","from":"developer"}],"created":"2016-03-08T19:07:09.000+0000","description":"Run the following codes may fail\n{code}\n(1 to 100).par.foreach { _ =>\n println(sc.parallelize(1 to 5).map { i => (i, i) }.toDF(\"a\", \"b\").count())\n}\n\njava.lang.IllegalArgumentException: spark.sql.execution.id is already set \n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:87) \n at org.apache.spark.sql.DataFrame.withNewExecutionId(DataFrame.scala:1904) \n at org.apache.spark.sql.DataFrame.collect(DataFrame.scala:1385) \n{code}\n\nThis is because SparkContext.runJob can be suspended when using a ForkJoinPool (e.g.,scala.concurrent.ExecutionContext.Implicits.global) as it calls Await.ready (introduced by https://github.com/apache/spark/pull/9264).\n\nSo when SparkContext.runJob is suspended, ForkJoinPool will run another task in the same thread, however, the local properties has been polluted.","issue_id":"12948114","key":"SPARK-13747","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-05-18T00:22:15.000+0000","role":"fixed_distractor","summary":"Concurrent execution in SQL doesn't work with Scala ForkJoinPool"} {"case_id":"12706259","cluster":"DISTRACTOR-SPARK-1394","comments":[{"body":"This seems to be related to the way the handle_sigchld method in daemon.py works.\nIn order to kill the zombie processes the worker calls os.waitpid on SIGCHLD. however. since using Popen also tries to do that eventually, you get a closed handle.\n\nSince platform.py is a native library, I would guess we should find a solution in pyspark (i.e. change the way handle_sigchld works, or maybe limit the processes it waits on)","created":"2014-04-04T07:51:17.548+0000"},{"body":"I have the same issue with Spark 0.9.1. Is there any workaround?","created":"2014-04-29T06:23:42.302+0000"},{"body":"If you have an __init__.py you are sure to go through, you can add the following code to your file:\n\n{code}\n# PySpark adds a SIGCHLD signal handler, but that breaks other packages, so we remove it\ntry:\n import signal\n signal.signal(signal.SIGCHLD, signal.SIG_DFL)\nexcept: pass\n{code}\n\nIt's a work around, it would be better to have a \"smart\" signal handler that only handles the processes that are direct descendants of the daemon. I might try to get something like that out. ","created":"2014-04-29T06:32:21.278+0000"},{"body":"[~idanzalz] unfortunately, it had helped to avoid only one exception, so I commented signal binding in PySpark and these crashes went away. I hope it will be fixed somehow in next Spark release.","created":"2014-04-30T00:31:14.580+0000"},{"body":"We also bumping into the same issue. My I know, how and where can we comment the signal binding in pyspark?","created":"2014-05-08T16:57:23.208+0000"},{"body":"[~sgottipa] spark/python/pyspark/daemon.py:75 #signal.signal(SIGCHLD, handle_sigchld)","created":"2014-05-09T02:53:03.834+0000"},{"body":"For what it's worth the platform library is unfortunately called by a number of numerical libraries associated to machine learning. In particular\n\nTheano calls on platform\nNumExpr calls on platform\nPandas calls on NumExpr\n\nThese libraries are very popular in the numerics / machine learning space.","created":"2014-05-20T21:54:45.922+0000"},{"body":"[~mrocklin] What? What is your reply for?","created":"2014-05-20T22:06:32.571+0000"},{"body":"Just trying to lend weight to this issue by adding context. \nLet me know if this is an inappropriate use of this forum.","created":"2014-05-20T22:09:17.545+0000"},{"body":"i'm taking a look at this","created":"2014-06-27T18:06:18.959+0000"},{"body":"fyi, this certainly looks like the waitpid(0,...) cleanup handler is cleaning up more than it should be\n\nalso, fyi, if you comment it out you'll start accumulating defunct worker processes, which is not good","created":"2014-06-27T19:15:35.951+0000"},{"body":"the python daemon does two levels of forking. the manager forks workers and each worker forks sub-workers, who then handle connections.\n\nthe master has a sigchld handler to cleanup workers. the workers also have sigchld handlers to clean up sub-workers.\n\nthe problem here is the worker's handler is also installed on the sub-worker, which interferes with calls like platform.system() (but not os.uname, btw!)\n\ni assert that the sub-workers, who do not intentionally fork, should not be responsible for cleaning up unexpected sub-processes, and as such should not have a sigchld handler.\n\nif it's desired to have tighter control of process cleanup, namespaces should be used and it should be the manager's responsibility.\n\npull request - https://github.com/apache/spark/pull/1247","created":"2014-06-27T20:14:01.038+0000"},{"body":"this should be resolved, @aarondav","created":"2014-06-29T09:34:22.603+0000"}],"conversations":[{"body":"A simple program that calls system.platform() on the worker fails most of the time (it works some times but very rarely).\nThis is critical since many libraries call that method (e.g. boto).\n\nHere is the trace of the attempt to call that method:\n\n\n\n$ /usr/local/spark/bin/pyspark\nPython 2.7.3 (default, Feb 27 2014, 20:00:17)\n[GCC 4.6.3] on linux2\nType \"help\", \"copyright\", \"credits\" or \"license\" for more information.\n14/04/02 18:18:37 INFO Utils: Using Spark's default log4j profile: org/apache/spark/log4j-defaults.properties\n14/04/02 18:18:37 WARN Utils: Your hostname, qlika-dev resolves to a loopback address: 127.0.1.1; using 10.33.102.46 instead (on interface eth1)\n14/04/02 18:18:37 WARN Utils: Set SPARK_LOCAL_IP if you need to bind to another address\n14/04/02 18:18:38 INFO Slf4jLogger: Slf4jLogger started\n14/04/02 18:18:38 INFO Remoting: Starting remoting\n14/04/02 18:18:39 INFO Remoting: Remoting started; listening on addresses :[akka.tcp://spark@10.33.102.46:36640]\n14/04/02 18:18:39 INFO Remoting: Remoting now listens on addresses: [akka.tcp://spark@10.33.102.46:36640]\n14/04/02 18:18:39 INFO SparkEnv: Registering BlockManagerMaster\n14/04/02 18:18:39 INFO DiskBlockManager: Created local directory at /tmp/spark-local-20140402181839-919f\n14/04/02 18:18:39 INFO MemoryStore: MemoryStore started with capacity 294.6 MB.\n14/04/02 18:18:39 INFO ConnectionManager: Bound socket to port 43357 with id = ConnectionManagerId(10.33.102.46,43357)\n14/04/02 18:18:39 INFO BlockManagerMaster: Trying to register BlockManager\n14/04/02 18:18:39 INFO BlockManagerMasterActor$BlockManagerInfo: Registering block manager 10.33.102.46:43357 with 294.6 MB RAM\n14/04/02 18:18:39 INFO BlockManagerMaster: Registered BlockManager\n14/04/02 18:18:39 INFO HttpServer: Starting HTTP Server\n14/04/02 18:18:39 INFO HttpBroadcast: Broadcast server started at http://10.33.102.46:51803\n14/04/02 18:18:39 INFO SparkEnv: Registering MapOutputTracker\n14/04/02 18:18:39 INFO HttpFileServer: HTTP File server directory is /tmp/spark-9b38acb0-7b01-4463-b0a6-602bfed05a2b\n14/04/02 18:18:39 INFO HttpServer: Starting HTTP Server\n14/04/02 18:18:40 INFO SparkUI: Started Spark Web UI at http://10.33.102.46:4040\n14/04/02 18:18:40 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /__ / .__/\\_,_/_/ /_/\\_\\ version 0.9.0\n /_/\n\nUsing Python version 2.7.3 (default, Feb 27 2014 20:00:17)\nSpark context available as sc.\n>>> import platform\n>>> sc.parallelize([1]).map(lambda x : platform.system()).collect()\n14/04/02 18:19:17 INFO SparkContext: Starting job: collect at :1\n14/04/02 18:19:17 INFO DAGScheduler: Got job 0 (collect at :1) with 1 output partitions (allowLocal=false)\n14/04/02 18:19:17 INFO DAGScheduler: Final stage: Stage 0 (collect at :1)\n14/04/02 18:19:17 INFO DAGScheduler: Parents of final stage: List()\n14/04/02 18:19:17 INFO DAGScheduler: Missing parents: List()\n14/04/02 18:19:17 INFO DAGScheduler: Submitting Stage 0 (PythonRDD[1] at collect at :1), which has no missing parents\n14/04/02 18:19:17 INFO DAGScheduler: Submitting 1 missing tasks from Stage 0 (PythonRDD[1] at collect at :1)\n14/04/02 18:19:17 INFO TaskSchedulerImpl: Adding task set 0.0 with 1 tasks\n14/04/02 18:19:17 INFO TaskSetManager: Starting task 0.0:0 as TID 0 on executor localhost: localhost (PROCESS_LOCAL)\n14/04/02 18:19:17 INFO TaskSetManager: Serialized task 0.0:0 as 2152 bytes in 12 ms\n14/04/02 18:19:17 INFO Executor: Running task ID 0\nPySpark worker failed with exception:\nTraceback (most recent call last):\n File \"/usr/local/spark/python/pyspark/worker.py\", line 77, in main\n serializer.dump_stream(func(split_index, iterator), outfile)\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 182, in dump_stream\n self.serializer.dump_stream(self._batched(iterator), stream)\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 117, in dump_stream\n for obj in iterator:\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 171, in _batched\n for item in iterator:\n File \"\", line 1, in \n File \"/usr/lib/python2.7/platform.py\", line 1306, in system\n return uname()[0]\n File \"/usr/lib/python2.7/platform.py\", line 1273, in uname\n processor = _syscmd_uname('-p','')\n File \"/usr/lib/python2.7/platform.py\", line 1030, in _syscmd_uname\n rc = f.close()\nIOError: [Errno 10] No child processes\n\n14/04/02 18:19:17 ERROR Executor: Exception in task ID 0\norg.apache.spark.api.python.PythonException: Traceback (most recent call last):\n File \"/usr/local/spark/python/pyspark/worker.py\", line 77, in main\n serializer.dump_stream(func(split_index, iterator), outfile)\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 182, in dump_stream\n self.serializer.dump_stream(self._batched(iterator), stream)\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 117, in dump_stream\n for obj in iterator:\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 171, in _batched\n for item in iterator:\n File \"\", line 1, in \n File \"/usr/lib/python2.7/platform.py\", line 1306, in system\n return uname()[0]\n File \"/usr/lib/python2.7/platform.py\", line 1273, in uname\n processor = _syscmd_uname('-p','')\n File \"/usr/lib/python2.7/platform.py\", line 1030, in _syscmd_uname\n rc = f.close()\nIOError: [Errno 10] No child processes\n\n at org.apache.spark.api.python.PythonRDD$$anon$1.read(PythonRDD.scala:131)\n at org.apache.spark.api.python.PythonRDD$$anon$1.(PythonRDD.scala:153)\n at org.apache.spark.api.python.PythonRDD.compute(PythonRDD.scala:96)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:241)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:232)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:109)\n at org.apache.spark.scheduler.Task.run(Task.scala:53)\n at org.apache.spark.executor.Executor$TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:213)\n at org.apache.spark.deploy.SparkHadoopUtil.runAsUser(SparkHadoopUtil.scala:49)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:178)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:744)\n","from":"reporter","subject":"calling system.platform on worker raises IOError"},{"body":"This seems to be related to the way the handle_sigchld method in daemon.py works.\nIn order to kill the zombie processes the worker calls os.waitpid on SIGCHLD. however. since using Popen also tries to do that eventually, you get a closed handle.\n\nSince platform.py is a native library, I would guess we should find a solution in pyspark (i.e. change the way handle_sigchld works, or maybe limit the processes it waits on)","from":"developer"},{"body":"I have the same issue with Spark 0.9.1. Is there any workaround?","from":"developer"},{"body":"If you have an __init__.py you are sure to go through, you can add the following code to your file:\n\n{code}\n# PySpark adds a SIGCHLD signal handler, but that breaks other packages, so we remove it\ntry:\n import signal\n signal.signal(signal.SIGCHLD, signal.SIG_DFL)\nexcept: pass\n{code}\n\nIt's a work around, it would be better to have a \"smart\" signal handler that only handles the processes that are direct descendants of the daemon. I might try to get something like that out. ","from":"developer"},{"body":"[~idanzalz] unfortunately, it had helped to avoid only one exception, so I commented signal binding in PySpark and these crashes went away. I hope it will be fixed somehow in next Spark release.","from":"developer"},{"body":"We also bumping into the same issue. My I know, how and where can we comment the signal binding in pyspark?","from":"developer"},{"body":"[~sgottipa] spark/python/pyspark/daemon.py:75 #signal.signal(SIGCHLD, handle_sigchld)","from":"developer"},{"body":"For what it's worth the platform library is unfortunately called by a number of numerical libraries associated to machine learning. In particular\n\nTheano calls on platform\nNumExpr calls on platform\nPandas calls on NumExpr\n\nThese libraries are very popular in the numerics / machine learning space.","from":"developer"},{"body":"[~mrocklin] What? What is your reply for?","from":"developer"},{"body":"Just trying to lend weight to this issue by adding context. \nLet me know if this is an inappropriate use of this forum.","from":"developer"},{"body":"i'm taking a look at this","from":"developer"},{"body":"fyi, this certainly looks like the waitpid(0,...) cleanup handler is cleaning up more than it should be\n\nalso, fyi, if you comment it out you'll start accumulating defunct worker processes, which is not good","from":"developer"},{"body":"the python daemon does two levels of forking. the manager forks workers and each worker forks sub-workers, who then handle connections.\n\nthe master has a sigchld handler to cleanup workers. the workers also have sigchld handlers to clean up sub-workers.\n\nthe problem here is the worker's handler is also installed on the sub-worker, which interferes with calls like platform.system() (but not os.uname, btw!)\n\ni assert that the sub-workers, who do not intentionally fork, should not be responsible for cleaning up unexpected sub-processes, and as such should not have a sigchld handler.\n\nif it's desired to have tighter control of process cleanup, namespaces should be used and it should be the manager's responsibility.\n\npull request - https://github.com/apache/spark/pull/1247","from":"developer"},{"body":"this should be resolved, @aarondav","from":"developer"}],"created":"2014-04-02T18:29:07.000+0000","description":"A simple program that calls system.platform() on the worker fails most of the time (it works some times but very rarely).\nThis is critical since many libraries call that method (e.g. boto).\n\nHere is the trace of the attempt to call that method:\n\n\n\n$ /usr/local/spark/bin/pyspark\nPython 2.7.3 (default, Feb 27 2014, 20:00:17)\n[GCC 4.6.3] on linux2\nType \"help\", \"copyright\", \"credits\" or \"license\" for more information.\n14/04/02 18:18:37 INFO Utils: Using Spark's default log4j profile: org/apache/spark/log4j-defaults.properties\n14/04/02 18:18:37 WARN Utils: Your hostname, qlika-dev resolves to a loopback address: 127.0.1.1; using 10.33.102.46 instead (on interface eth1)\n14/04/02 18:18:37 WARN Utils: Set SPARK_LOCAL_IP if you need to bind to another address\n14/04/02 18:18:38 INFO Slf4jLogger: Slf4jLogger started\n14/04/02 18:18:38 INFO Remoting: Starting remoting\n14/04/02 18:18:39 INFO Remoting: Remoting started; listening on addresses :[akka.tcp://spark@10.33.102.46:36640]\n14/04/02 18:18:39 INFO Remoting: Remoting now listens on addresses: [akka.tcp://spark@10.33.102.46:36640]\n14/04/02 18:18:39 INFO SparkEnv: Registering BlockManagerMaster\n14/04/02 18:18:39 INFO DiskBlockManager: Created local directory at /tmp/spark-local-20140402181839-919f\n14/04/02 18:18:39 INFO MemoryStore: MemoryStore started with capacity 294.6 MB.\n14/04/02 18:18:39 INFO ConnectionManager: Bound socket to port 43357 with id = ConnectionManagerId(10.33.102.46,43357)\n14/04/02 18:18:39 INFO BlockManagerMaster: Trying to register BlockManager\n14/04/02 18:18:39 INFO BlockManagerMasterActor$BlockManagerInfo: Registering block manager 10.33.102.46:43357 with 294.6 MB RAM\n14/04/02 18:18:39 INFO BlockManagerMaster: Registered BlockManager\n14/04/02 18:18:39 INFO HttpServer: Starting HTTP Server\n14/04/02 18:18:39 INFO HttpBroadcast: Broadcast server started at http://10.33.102.46:51803\n14/04/02 18:18:39 INFO SparkEnv: Registering MapOutputTracker\n14/04/02 18:18:39 INFO HttpFileServer: HTTP File server directory is /tmp/spark-9b38acb0-7b01-4463-b0a6-602bfed05a2b\n14/04/02 18:18:39 INFO HttpServer: Starting HTTP Server\n14/04/02 18:18:40 INFO SparkUI: Started Spark Web UI at http://10.33.102.46:4040\n14/04/02 18:18:40 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /__ / .__/\\_,_/_/ /_/\\_\\ version 0.9.0\n /_/\n\nUsing Python version 2.7.3 (default, Feb 27 2014 20:00:17)\nSpark context available as sc.\n>>> import platform\n>>> sc.parallelize([1]).map(lambda x : platform.system()).collect()\n14/04/02 18:19:17 INFO SparkContext: Starting job: collect at :1\n14/04/02 18:19:17 INFO DAGScheduler: Got job 0 (collect at :1) with 1 output partitions (allowLocal=false)\n14/04/02 18:19:17 INFO DAGScheduler: Final stage: Stage 0 (collect at :1)\n14/04/02 18:19:17 INFO DAGScheduler: Parents of final stage: List()\n14/04/02 18:19:17 INFO DAGScheduler: Missing parents: List()\n14/04/02 18:19:17 INFO DAGScheduler: Submitting Stage 0 (PythonRDD[1] at collect at :1), which has no missing parents\n14/04/02 18:19:17 INFO DAGScheduler: Submitting 1 missing tasks from Stage 0 (PythonRDD[1] at collect at :1)\n14/04/02 18:19:17 INFO TaskSchedulerImpl: Adding task set 0.0 with 1 tasks\n14/04/02 18:19:17 INFO TaskSetManager: Starting task 0.0:0 as TID 0 on executor localhost: localhost (PROCESS_LOCAL)\n14/04/02 18:19:17 INFO TaskSetManager: Serialized task 0.0:0 as 2152 bytes in 12 ms\n14/04/02 18:19:17 INFO Executor: Running task ID 0\nPySpark worker failed with exception:\nTraceback (most recent call last):\n File \"/usr/local/spark/python/pyspark/worker.py\", line 77, in main\n serializer.dump_stream(func(split_index, iterator), outfile)\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 182, in dump_stream\n self.serializer.dump_stream(self._batched(iterator), stream)\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 117, in dump_stream\n for obj in iterator:\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 171, in _batched\n for item in iterator:\n File \"\", line 1, in \n File \"/usr/lib/python2.7/platform.py\", line 1306, in system\n return uname()[0]\n File \"/usr/lib/python2.7/platform.py\", line 1273, in uname\n processor = _syscmd_uname('-p','')\n File \"/usr/lib/python2.7/platform.py\", line 1030, in _syscmd_uname\n rc = f.close()\nIOError: [Errno 10] No child processes\n\n14/04/02 18:19:17 ERROR Executor: Exception in task ID 0\norg.apache.spark.api.python.PythonException: Traceback (most recent call last):\n File \"/usr/local/spark/python/pyspark/worker.py\", line 77, in main\n serializer.dump_stream(func(split_index, iterator), outfile)\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 182, in dump_stream\n self.serializer.dump_stream(self._batched(iterator), stream)\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 117, in dump_stream\n for obj in iterator:\n File \"/usr/local/spark/python/pyspark/serializers.py\", line 171, in _batched\n for item in iterator:\n File \"\", line 1, in \n File \"/usr/lib/python2.7/platform.py\", line 1306, in system\n return uname()[0]\n File \"/usr/lib/python2.7/platform.py\", line 1273, in uname\n processor = _syscmd_uname('-p','')\n File \"/usr/lib/python2.7/platform.py\", line 1030, in _syscmd_uname\n rc = f.close()\nIOError: [Errno 10] No child processes\n\n at org.apache.spark.api.python.PythonRDD$$anon$1.read(PythonRDD.scala:131)\n at org.apache.spark.api.python.PythonRDD$$anon$1.(PythonRDD.scala:153)\n at org.apache.spark.api.python.PythonRDD.compute(PythonRDD.scala:96)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:241)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:232)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:109)\n at org.apache.spark.scheduler.Task.run(Task.scala:53)\n at org.apache.spark.executor.Executor$TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:213)\n at org.apache.spark.deploy.SparkHadoopUtil.runAsUser(SparkHadoopUtil.scala:49)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:178)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:744)\n","issue_id":"12706259","key":"SPARK-1394","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2014-07-25T21:36:09.000+0000","role":"fixed_distractor","summary":"calling system.platform on worker raises IOError"} {"case_id":"12954016","cluster":"DISTRACTOR-SPARK-14209","comments":[{"body":"If you have preemption enabled you really should be using the external shuffle service to avoid these issues. There's really not much that can be done otherwise.","created":"2016-03-28T17:52:49.232+0000"},{"body":"I am using an external shuffle service, as I mentioned in the description. I can see ESTABLISHED TCP connections between it and the executors, and the process is active.\n\nIs there some other way I can verify that it is being used as it is configured?","created":"2016-03-28T18:11:11.006+0000"},{"body":"If you're using the shuffle service then executors should not be trying to talk to each other at all.\n\nCan you check whether your app's configuration contains {{spark.shuffle.service.enabled=true}}?","created":"2016-03-28T18:17:11.498+0000"},{"body":"spark.shuffle.service.enabled\tis shown to be true in the Environment tab of my Spark UI.\n\nI see the following ShuffleService entries in my nodemanager logs:\nhttps://gist.github.com/milescrawford/dabe066b3f29ea078db6","created":"2016-03-28T18:18:17.829+0000"},{"body":"I don't see entries related to the offending executor (000143) in those logs; did you look at the logs from the right NM?","created":"2016-03-28T18:25:09.185+0000"},{"body":"Sorry, i was just showing that they were in use - let me find an example that was preempted.","created":"2016-03-28T18:39:32.530+0000"},{"body":"Here is all the nodemanager logs related to an application that died in this fashion:\n\nhttps://www.dropbox.com/s/p73w920u3tqn9i2/nodemanager-logs.gz?dl=0\n\nThe application that died in this example is application_1458935819920_0019\n\n","created":"2016-03-28T18:51:56.820+0000"},{"body":"Ah, my bad, I should have noticed this in the original stack trace. This seems to be related to caching and not shuffle, right? Can you confirm that you're caching RDDs in that app?\n\nIt seems that code should be more resilient to executors going down.","created":"2016-03-28T19:29:26.847+0000"},{"body":"Yes, I am certainly caching. I did not realize that caching does not use the same external store.","created":"2016-03-28T19:30:03.286+0000"},{"body":"This code has changed significantly in 2.0 (due to SPARK-12817); need to see if it suffers from the same issue.","created":"2016-03-28T20:07:22.061+0000"},{"body":"It will be useful to know whether the issue is fixed in 2.0, but is waiting for 2.0 may not be practical for some users.\n\nAny chance of a fix for the 1.6.x branches?","created":"2016-03-29T16:27:33.569+0000"},{"body":"2.0 doesn't currently suffer from this problem because, instead, it suffers from SPARK-14252. :-p\n\nIt shouldn't be hard to provide a fix for 1.6, though.","created":"2016-03-29T21:34:25.079+0000"},{"body":"That's fantastic news! let me know if I can help test a fix.","created":"2016-03-29T21:38:47.755+0000"},{"body":"Are you able to share the logs from your driver?","created":"2016-03-29T23:25:26.697+0000"},{"body":"Here you go, from the same application run: https://www.dropbox.com/s/0m5c5rshydhj4am/driver-log.gz?dl=0","created":"2016-03-29T23:53:54.128+0000"},{"body":"Are you using a custom log configuration? I can't find some log messages that I can see in the code, especially from BlockManagerMasterEndpoint and BlockManagerMaster.\n\nBTW I think the real culprit is executor 51 (container 000191 instead of the one you originally mention). Still isolating all the logs related to that executor to try to figure out what's going on.","created":"2016-03-30T00:45:47.266+0000"},{"body":"I don't think there's anything custom about our logging - we have it set to INFO, and a few specific modules are turned down, but not the ones you mention.","created":"2016-03-30T00:57:51.814+0000"},{"body":"That's weird. I don't know whether the funny logs are a symptom of the problem or just something else unrelated. For example, aside from the missing logs, there's stuff like this:\n\n{noformat}\n2016-03-25 23:28:41,722 ERROR o.a.s.s.cluster.YarnClusterScheduler: Lost an executor 51 (already removed): Pending loss reason.\n{noformat}\n\nWhich looks fine, except that log is generated by TaskSchedulerImpl, not YarnClusterScheduler.\n\nTracing the code it doesn't seem like there's a problem anywhere; when the executor dies the BlockManagerMaster should remove it from its internal state, and although some tasks might fail in that window before the bookkeeping is updated, what your logs show shouldn't really happen unless there's something really catastrophic going on, like the BlockManagerMaster being deadlocked or having a ridiculously long message queue or something.","created":"2016-03-30T16:40:01.983+0000"},{"body":"Anything I can do to help you investigate?","created":"2016-03-30T17:20:53.470+0000"},{"body":"Do you have any other application that failed this way? A different set of logs would hopefully provide better information (or at least tell us whether those logs are a one off weirdness).","created":"2016-03-30T17:25:43.418+0000"},{"body":"Here's a tarball of all the container logs for another application which failed this way https://www.dropbox.com/s/4ug8064p09a5c7l/application_1458935819920_0029.tar.gz?dl=0","created":"2016-03-30T18:46:06.657+0000"},{"body":"Those logs show the same weird issues as the previous one... is there a way you can use the default log configuration from Spark? That would give us much better and non-misleading information.\n\nAlso, it didn't seem like that application failed. I see some fetch failures, but that's to be expected when executors die. What's odd about your original log is that tasks failed multiple (up to 100) times and eventually failed the application, and that doesn't seem to be happening for this last set of logs.","created":"2016-03-31T21:51:12.630+0000"},{"body":"Can you be a bit more specific about the default log configuration from spark?\n\nAll we're doing in terms of logging is placing a logback.xml file into our classpath that sets the console logger to level INFO...","created":"2016-04-01T21:47:21.946+0000"},{"body":"bq. All we're doing in terms of logging is placing a logback.xml file into our classpath\n\nIs your application explicitly using logback? Spark's \"default\" logging framework is log4j (and the default log level is already INFO).","created":"2016-04-01T21:51:05.840+0000"},{"body":"yes, we definitely use logback - if you look at the top of the logs, you can see it starting up and parsing its config file.","created":"2016-04-01T21:52:27.937+0000"},{"body":"So, I don't know what to say. Your log configuration is resulting in corrupt logs that make debugging your issue very hard. So unless you can easily replace your logging configuration with Spark's default, it will be hard to make progress here.","created":"2016-04-01T21:55:39.833+0000"},{"body":"Yes. I'm still committed to helping you - was just looking for any insight you might have into how we might be corrupting the logs. I'll dig into it monday.","created":"2016-04-01T22:02:54.592+0000"},{"body":"I have not forgotten this, worked on getting a cleaner reproduction today. Stand by.","created":"2016-04-07T00:05:57.831+0000"},{"body":"So, I have launched the exact same simple application that does some RDD caching one with our standard tooling, and one with zero extra configuration:\n\nHere's the application:\nhttps://gist.github.com/afc552574cdfc65ec71bad744e81f6fd\n\nHere's Spark's default logging for a run:\nhttps://www.dropbox.com/s/78u1y8ydsi3jpnb/default-driver.log\n\nHere's our logging:\nhttps://www.dropbox.com/s/esekrncpj7vprxe/our-driver.log\n\nDoes this example show the corruption you're talking about?","created":"2016-04-08T17:56:51.181+0000"},{"body":"Just had a very similar error when a host was lost: https://www.dropbox.com/s/och5o15sxr7dxv6/host-lost-failure.gz?dl=0\n\nThe log verbosity is turned to warn here.","created":"2016-04-13T21:15:18.359+0000"},{"body":"Just to let you know I haven't forgotten about this, I just haven't had much time for anything lately.","created":"2016-04-21T22:59:35.540+0000"},{"body":"This issue looks a lot like https://issues.apache.org/jira/browse/SPARK-7703 - but that issue is marked as resolved...","created":"2016-05-08T22:55:20.334+0000"},{"body":"Hi [~milesc], I look at your updates after my last update, but I don't see logs from an actual reproduction of the issue.\n\nThe file at https://www.dropbox.com/s/78u1y8ydsi3jpnb/default-driver.log contains log from a Spark application, but not from one that failed in the way you describe here. Similarly, https://www.dropbox.com/s/och5o15sxr7dxv6/host-lost-failure.gz?dl=0 contains a different issue.\n\nIf you look at my previous messages, the log message that shows the problem with your original logs looks like this:\n\n{noformat}\n2016-03-25 23:28:41,722 ERROR o.a.s.s.cluster.YarnClusterScheduler: Lost an executor 51 (already removed): Pending loss reason.\n{noformat}\n\nThe logger that generates that message is from the {{TaskSchedulerImpl}} class; so if you see the message coming from {{YarnClusterScheduler}}, something is wrong. There's also a lot of missing logs that you can see in Spark's code but not in your logs, which are further hints that something is wrong with your logs.\n\nNone of the latest logs show any of those messages, which tell me that they're from apps that did not hit this issue.","created":"2016-05-11T23:19:33.075+0000"},{"body":"BTW this particular error looks like SPARK-9439, although that bug says it's fixed in 1.6.","created":"2016-05-12T17:28:23.318+0000"},{"body":"I believe that this issue should have been fixed by SPARK-17485 in Spark 2.x.","created":"2016-09-21T21:10:44.900+0000"},{"body":"I have backported SPARK-17485 to Spark 1.6.x (for inclusion in Spark 1.6.3), so I believe that this issue should be fixed and therefore I'm going to resolve it.\n\nI believe that this ticket may actually be discussing multiple issues that are related to fetch failures following executor lost but which have different underlying causes and fixes. I think the original issue reported in this JIRA relates to BlockFetchException, an error which occurs due to failed fetches of NON-shuffle blocks (such as broadcasts or cached RDD blocks).\n\nIf you're a user and are still experiencing Spark application failures due to executor pre-emption then please file a new JIRA ticket and make sure to include the Spark version and the portion of the driver log which contains the job failure log message (since that will show which exception / stack trace ultimately triggered the failure, allowing us to distinguish shuffle block fetch failures vs. other types of fetch failures). ","created":"2016-09-22T18:16:13.354+0000"}],"conversations":[{"body":"We have a fair-sharing cluster set up, including the external shuffle service. When a new job arrives, existing jobs are successfully preempted down to fit.\n\nA spate of these messages arrives:\n\tExecutorLostFailure (executor 48 exited unrelated to the running tasks) Reason: Container container_1458935819920_0019_01_000143 on host: ip-10-12-46-235.us-west-2.compute.internal was preempted.\n\nThis seems fine - the problem is that soon thereafter, our whole application fails because it is unable to fetch blocks from the pre-empted containers:\n\norg.apache.spark.storage.BlockFetchException: Failed to fetch block from 1 locations. Most recent failure cause:\n Caused by: java.io.IOException: Failed to connect to ip-10-12-46-235.us-west-2.compute.internal/10.12.46.235:55681\n Caused by: java.net.ConnectException: Connection refused: ip-10-12-46-235.us-west-2.compute.internal/10.12.46.235:55681\n\nFull stack: https://gist.github.com/milescrawford/33a1c1e61d88cc8c6daf\n\nSpark does not attempt to recreate these blocks - the tasks simply fail over and over until the maxTaskAttempts value is reached.\n\nIt appears to me that there is some fault in the way preempted containers are being handled - shouldn't these blocks be recreated on demand?","from":"reporter","subject":"Application failure during preemption."},{"body":"If you have preemption enabled you really should be using the external shuffle service to avoid these issues. There's really not much that can be done otherwise.","from":"developer"},{"body":"I am using an external shuffle service, as I mentioned in the description. I can see ESTABLISHED TCP connections between it and the executors, and the process is active.\n\nIs there some other way I can verify that it is being used as it is configured?","from":"developer"},{"body":"If you're using the shuffle service then executors should not be trying to talk to each other at all.\n\nCan you check whether your app's configuration contains {{spark.shuffle.service.enabled=true}}?","from":"developer"},{"body":"spark.shuffle.service.enabled\tis shown to be true in the Environment tab of my Spark UI.\n\nI see the following ShuffleService entries in my nodemanager logs:\nhttps://gist.github.com/milescrawford/dabe066b3f29ea078db6","from":"developer"},{"body":"I don't see entries related to the offending executor (000143) in those logs; did you look at the logs from the right NM?","from":"developer"},{"body":"Sorry, i was just showing that they were in use - let me find an example that was preempted.","from":"developer"},{"body":"Here is all the nodemanager logs related to an application that died in this fashion:\n\nhttps://www.dropbox.com/s/p73w920u3tqn9i2/nodemanager-logs.gz?dl=0\n\nThe application that died in this example is application_1458935819920_0019\n\n","from":"developer"},{"body":"Ah, my bad, I should have noticed this in the original stack trace. This seems to be related to caching and not shuffle, right? Can you confirm that you're caching RDDs in that app?\n\nIt seems that code should be more resilient to executors going down.","from":"developer"},{"body":"Yes, I am certainly caching. I did not realize that caching does not use the same external store.","from":"developer"},{"body":"This code has changed significantly in 2.0 (due to SPARK-12817); need to see if it suffers from the same issue.","from":"developer"},{"body":"It will be useful to know whether the issue is fixed in 2.0, but is waiting for 2.0 may not be practical for some users.\n\nAny chance of a fix for the 1.6.x branches?","from":"developer"},{"body":"2.0 doesn't currently suffer from this problem because, instead, it suffers from SPARK-14252. :-p\n\nIt shouldn't be hard to provide a fix for 1.6, though.","from":"developer"},{"body":"That's fantastic news! let me know if I can help test a fix.","from":"developer"},{"body":"Are you able to share the logs from your driver?","from":"developer"},{"body":"Here you go, from the same application run: https://www.dropbox.com/s/0m5c5rshydhj4am/driver-log.gz?dl=0","from":"developer"},{"body":"Are you using a custom log configuration? I can't find some log messages that I can see in the code, especially from BlockManagerMasterEndpoint and BlockManagerMaster.\n\nBTW I think the real culprit is executor 51 (container 000191 instead of the one you originally mention). Still isolating all the logs related to that executor to try to figure out what's going on.","from":"developer"},{"body":"I don't think there's anything custom about our logging - we have it set to INFO, and a few specific modules are turned down, but not the ones you mention.","from":"developer"},{"body":"That's weird. I don't know whether the funny logs are a symptom of the problem or just something else unrelated. For example, aside from the missing logs, there's stuff like this:\n\n{noformat}\n2016-03-25 23:28:41,722 ERROR o.a.s.s.cluster.YarnClusterScheduler: Lost an executor 51 (already removed): Pending loss reason.\n{noformat}\n\nWhich looks fine, except that log is generated by TaskSchedulerImpl, not YarnClusterScheduler.\n\nTracing the code it doesn't seem like there's a problem anywhere; when the executor dies the BlockManagerMaster should remove it from its internal state, and although some tasks might fail in that window before the bookkeeping is updated, what your logs show shouldn't really happen unless there's something really catastrophic going on, like the BlockManagerMaster being deadlocked or having a ridiculously long message queue or something.","from":"developer"},{"body":"Anything I can do to help you investigate?","from":"developer"},{"body":"Do you have any other application that failed this way? A different set of logs would hopefully provide better information (or at least tell us whether those logs are a one off weirdness).","from":"developer"},{"body":"Here's a tarball of all the container logs for another application which failed this way https://www.dropbox.com/s/4ug8064p09a5c7l/application_1458935819920_0029.tar.gz?dl=0","from":"developer"},{"body":"Those logs show the same weird issues as the previous one... is there a way you can use the default log configuration from Spark? That would give us much better and non-misleading information.\n\nAlso, it didn't seem like that application failed. I see some fetch failures, but that's to be expected when executors die. What's odd about your original log is that tasks failed multiple (up to 100) times and eventually failed the application, and that doesn't seem to be happening for this last set of logs.","from":"developer"},{"body":"Can you be a bit more specific about the default log configuration from spark?\n\nAll we're doing in terms of logging is placing a logback.xml file into our classpath that sets the console logger to level INFO...","from":"developer"},{"body":"bq. All we're doing in terms of logging is placing a logback.xml file into our classpath\n\nIs your application explicitly using logback? Spark's \"default\" logging framework is log4j (and the default log level is already INFO).","from":"developer"},{"body":"yes, we definitely use logback - if you look at the top of the logs, you can see it starting up and parsing its config file.","from":"developer"},{"body":"So, I don't know what to say. Your log configuration is resulting in corrupt logs that make debugging your issue very hard. So unless you can easily replace your logging configuration with Spark's default, it will be hard to make progress here.","from":"developer"},{"body":"Yes. I'm still committed to helping you - was just looking for any insight you might have into how we might be corrupting the logs. I'll dig into it monday.","from":"developer"},{"body":"I have not forgotten this, worked on getting a cleaner reproduction today. Stand by.","from":"developer"},{"body":"So, I have launched the exact same simple application that does some RDD caching one with our standard tooling, and one with zero extra configuration:\n\nHere's the application:\nhttps://gist.github.com/afc552574cdfc65ec71bad744e81f6fd\n\nHere's Spark's default logging for a run:\nhttps://www.dropbox.com/s/78u1y8ydsi3jpnb/default-driver.log\n\nHere's our logging:\nhttps://www.dropbox.com/s/esekrncpj7vprxe/our-driver.log\n\nDoes this example show the corruption you're talking about?","from":"developer"},{"body":"Just had a very similar error when a host was lost: https://www.dropbox.com/s/och5o15sxr7dxv6/host-lost-failure.gz?dl=0\n\nThe log verbosity is turned to warn here.","from":"developer"},{"body":"Just to let you know I haven't forgotten about this, I just haven't had much time for anything lately.","from":"developer"},{"body":"This issue looks a lot like https://issues.apache.org/jira/browse/SPARK-7703 - but that issue is marked as resolved...","from":"developer"},{"body":"Hi [~milesc], I look at your updates after my last update, but I don't see logs from an actual reproduction of the issue.\n\nThe file at https://www.dropbox.com/s/78u1y8ydsi3jpnb/default-driver.log contains log from a Spark application, but not from one that failed in the way you describe here. Similarly, https://www.dropbox.com/s/och5o15sxr7dxv6/host-lost-failure.gz?dl=0 contains a different issue.\n\nIf you look at my previous messages, the log message that shows the problem with your original logs looks like this:\n\n{noformat}\n2016-03-25 23:28:41,722 ERROR o.a.s.s.cluster.YarnClusterScheduler: Lost an executor 51 (already removed): Pending loss reason.\n{noformat}\n\nThe logger that generates that message is from the {{TaskSchedulerImpl}} class; so if you see the message coming from {{YarnClusterScheduler}}, something is wrong. There's also a lot of missing logs that you can see in Spark's code but not in your logs, which are further hints that something is wrong with your logs.\n\nNone of the latest logs show any of those messages, which tell me that they're from apps that did not hit this issue.","from":"developer"},{"body":"BTW this particular error looks like SPARK-9439, although that bug says it's fixed in 1.6.","from":"developer"},{"body":"I believe that this issue should have been fixed by SPARK-17485 in Spark 2.x.","from":"developer"},{"body":"I have backported SPARK-17485 to Spark 1.6.x (for inclusion in Spark 1.6.3), so I believe that this issue should be fixed and therefore I'm going to resolve it.\n\nI believe that this ticket may actually be discussing multiple issues that are related to fetch failures following executor lost but which have different underlying causes and fixes. I think the original issue reported in this JIRA relates to BlockFetchException, an error which occurs due to failed fetches of NON-shuffle blocks (such as broadcasts or cached RDD blocks).\n\nIf you're a user and are still experiencing Spark application failures due to executor pre-emption then please file a new JIRA ticket and make sure to include the Spark version and the portion of the driver log which contains the job failure log message (since that will show which exception / stack trace ultimately triggered the failure, allowing us to distinguish shuffle block fetch failures vs. other types of fetch failures). ","from":"developer"}],"created":"2016-03-28T17:48:44.000+0000","description":"We have a fair-sharing cluster set up, including the external shuffle service. When a new job arrives, existing jobs are successfully preempted down to fit.\n\nA spate of these messages arrives:\n\tExecutorLostFailure (executor 48 exited unrelated to the running tasks) Reason: Container container_1458935819920_0019_01_000143 on host: ip-10-12-46-235.us-west-2.compute.internal was preempted.\n\nThis seems fine - the problem is that soon thereafter, our whole application fails because it is unable to fetch blocks from the pre-empted containers:\n\norg.apache.spark.storage.BlockFetchException: Failed to fetch block from 1 locations. Most recent failure cause:\n Caused by: java.io.IOException: Failed to connect to ip-10-12-46-235.us-west-2.compute.internal/10.12.46.235:55681\n Caused by: java.net.ConnectException: Connection refused: ip-10-12-46-235.us-west-2.compute.internal/10.12.46.235:55681\n\nFull stack: https://gist.github.com/milescrawford/33a1c1e61d88cc8c6daf\n\nSpark does not attempt to recreate these blocks - the tasks simply fail over and over until the maxTaskAttempts value is reached.\n\nIt appears to me that there is some fault in the way preempted containers are being handled - shouldn't these blocks be recreated on demand?","issue_id":"12954016","key":"SPARK-14209","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-09-22T18:17:47.000+0000","role":"fixed_distractor","summary":"Application failure during preemption."} {"case_id":"12954173","cluster":"DISTRACTOR-SPARK-14228","comments":[{"body":"We have similar issues, reporting fake ERROR about \"Could not find xxxx or it has been stopped\".\n\n{code}\n31048 16/03/10 04:28:38 WARN Dispatcher: Message RemoteProcessDisconnected(xxxxxxx:7077) dropped.\norg.apache.spark.SparkException: Could not find OutputCommitCoordinator or it has been stopped.\n at org.apache.spark.rpc.netty.Dispatcher.postMessage(Dispatcher.scala:161)\n at org.apache.spark.rpc.netty.Dispatcher.postToAll(Dispatcher.scala:109)\n at org.apache.spark.rpc.netty.NettyRpcHandler.connectionTerminated(NettyRpcEnv.scala:630)\n at org.apache.spark.network.server.TransportRequestHandler.channelUnregistered(TransportRequestHandler.java:94)\n at org.apache.spark.network.server.TransportChannelHandler.channelUnregistered(TransportChannelHandler.java:89)\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelUnregistered(AbstractChannelHandlerContext.java:158)\n at io.netty.channel.AbstractChannelHandlerContext.fireChannelUnregistered(AbstractChannelHandlerContext.java:144)\n at io.netty.channel.ChannelInboundHandlerAdapter.channelUnregistered(ChannelInboundHandlerAdapter.java:53)\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelUnregistered(AbstractChannelHandlerContext.java:158)\n at io.netty.channel.AbstractChannelHandlerContext.fireChannelUnregistered(AbstractChannelHandlerContext.java:144)\n at io.netty.channel.ChannelInboundHandlerAdapter.channelUnregistered(ChannelInboundHandlerAdapter.java:53)\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelUnregistered(AbstractChannelHandlerContext.java:158)\n at io.netty.channel.AbstractChannelHandlerContext.fireChannelUnregistered(AbstractChannelHandlerContext.java:144)\n at io.netty.channel.ChannelInboundHandlerAdapter.channelUnregistered(ChannelInboundHandlerAdapter.java:53)\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelUnregistered(AbstractChannelHandlerContext.java:158)\n at io.netty.channel.AbstractChannelHandlerContext.fireChannelUnregistered(AbstractChannelHandlerContext.java:144)\n at io.netty.channel.DefaultChannelPipeline.fireChannelUnregistered(DefaultChannelPipeline.java:739)\n at io.netty.channel.AbstractChannel$AbstractUnsafe$8.run(AbstractChannel.java:659)\n at io.netty.util.concurrent.SingleThreadEventExecutor.runAllTasks(SingleThreadEventExecutor.java:357)\n at io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:357)\n at io.netty.util.concurrent.SingleThreadEventExecutor$2.run(SingleThreadEventExecutor.java:111)\n at java.lang.Thread.run(Thread.java:785)\n{code}","created":"2016-05-19T10:03:52.282+0000"},{"body":"Hi, can you specify the version you were working with? I have received the same error in 1.6.0","created":"2017-03-28T10:23:24.460+0000"},{"body":"I had that error on {{1.6.3}}, too.","created":"2017-03-29T07:28:35.600+0000"},{"body":"I had similar issue running while spark in multithread spark context code.\n\n17/09/18 21:40:09 ERROR Inbox: Ignoring error\norg.apache.spark.SparkException: Could not find CoarseGrainedScheduler or it has been stopped.\n\tat org.apache.spark.rpc.netty.Dispatcher.postMessage(Dispatcher.scala:161)\n\tat org.apache.spark.rpc.netty.Dispatcher.postOneWayMessage(Dispatcher.scala:131)\n\tat org.apache.spark.rpc.netty.NettyRpcEnv.send(NettyRpcEnv.scala:193)\n\tat org.apache.spark.rpc.netty.NettyRpcEndpointRef.send(NettyRpcEnv.scala:520)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend.reviveOffers(CoarseGrainedSchedulerBackend.scala:375)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl.executorLost(TaskSchedulerImpl.scala:497)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend$DriverEndpoint.disableExecutor(CoarseGrainedSchedulerBackend.scala:301)\n\tat org.apache.spark.scheduler.cluster.YarnSchedulerBackend$YarnDriverEndpoint$$anonfun$onDisconnected$1.apply(YarnSchedulerBackend.scala:195)\n\tat org.apache.spark.scheduler.cluster.YarnSchedulerBackend$YarnDriverEndpoint$$anonfun$onDisconnected$1.apply(YarnSchedulerBackend.scala:194)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.cluster.YarnSchedulerBackend$YarnDriverEndpoint.onDisconnected(YarnSchedulerBackend.scala:194)\n\tat org.apache.spark.rpc.netty.Inbox$$anonfun$process$1.apply$mcV$sp(Inbox.scala:142)\n\tat org.apache.spark.rpc.netty.Inbox.safelyCall(Inbox.scala:204)\n\tat org.apache.spark.rpc.netty.Inbox.process(Inbox.scala:100)\n\tat org.apache.spark.rpc.netty.Dispatcher$MessageLoop.run(Dispatcher.scala:215)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)","created":"2017-09-19T21:28:58.711+0000"},{"body":"I encountered the same issue in 1.6.3\r\n\r\n{code}\r\n17/11/03 22:38:33 WARN YarnSchedulerBackend$YarnSchedulerEndpoint: Executor for container container_e02_1509517131757_0001_01_000003 exited because of a YARN event (e.g., pre-emption) and not because of an error in the running job.\r\n17/11/03 22:38:33 ERROR YarnClientSchedulerBackend: Could not find CoarseGrainedScheduler or it has been stopped.\r\norg.apache.spark.SparkException: Could not find CoarseGrainedScheduler or it has been stopped.\r\n at org.apache.spark.rpc.netty.Dispatcher.postMessage(Dispatcher.scala:163)\r\n at org.apache.spark.rpc.netty.Dispatcher.postLocalMessage(Dispatcher.scala:128)\r\n at org.apache.spark.rpc.netty.NettyRpcEnv.ask(NettyRpcEnv.scala:231)\r\n at org.apache.spark.rpc.netty.NettyRpcEndpointRef.ask(NettyRpcEnv.scala:515)\r\n at org.apache.spark.rpc.RpcEndpointRef.ask(RpcEndpointRef.scala:62)\r\n at org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend.removeExecutor(CoarseGrainedSchedulerBackend.scala:392)\r\n at org.apache.spark.scheduler.cluster.YarnSchedulerBackend$YarnSchedulerEndpoint$$anonfun$receive$1.applyOrElse(YarnSchedulerBackend.scala:259)\r\n at org.apache.spark.rpc.netty.Inbox$$anonfun$process$1.apply$mcV$sp(Inbox.scala:116)\r\n at org.apache.spark.rpc.netty.Inbox.safelyCall(Inbox.scala:204)\r\n at org.apache.spark.rpc.netty.Inbox.process(Inbox.scala:100)\r\n at org.apache.spark.rpc.netty.Dispatcher$MessageLoop.run(Dispatcher.scala:217)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n at java.lang.Thread.run(Thread.java:745)\r\n17/11/03 22:38:33 WARN YarnSchedulerBackend$YarnSchedulerEndpoint: Executor for container container_e02_1509517131757_0001_01_000002 exited because of a YARN event (e.g., pre-emption) and not because of an error in the running job.\r\n17/11/03 22:38:33 ERROR YarnClientSchedulerBackend: Could not find CoarseGrainedScheduler or it has been stopped.\r\norg.apache.spark.SparkException: Could not find CoarseGrainedScheduler or it has been stopped.\r\n at org.apache.spark.rpc.netty.Dispatcher.postMessage(Dispatcher.scala:163)\r\n at org.apache.spark.rpc.netty.Dispatcher.postLocalMessage(Dispatcher.scala:128)\r\n at org.apache.spark.rpc.netty.NettyRpcEnv.ask(NettyRpcEnv.scala:231)\r\n at org.apache.spark.rpc.netty.NettyRpcEndpointRef.ask(NettyRpcEnv.scala:515)\r\n at org.apache.spark.rpc.RpcEndpointRef.ask(RpcEndpointRef.scala:62)\r\n at org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend.removeExecutor(CoarseGrainedSchedulerBackend.scala:392)\r\n at org.apache.spark.scheduler.cluster.YarnSchedulerBackend$YarnSchedulerEndpoint$$anonfun$receive$1.applyOrElse(YarnSchedulerBackend.scala:259)\r\n at org.apache.spark.rpc.netty.Inbox$$anonfun$process$1.apply$mcV$sp(Inbox.scala:116)\r\n at org.apache.spark.rpc.netty.Inbox.safelyCall(Inbox.scala:204)\r\n at org.apache.spark.rpc.netty.Inbox.process(Inbox.scala:100)\r\n at org.apache.spark.rpc.netty.Dispatcher$MessageLoop.run(Dispatcher.scala:217)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n at java.lang.Thread.run(Thread.java:745)\r\n{code}\r\n","created":"2017-11-06T03:40:32.603+0000"},{"body":"User 'devaraj-kavali' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19741","created":"2017-11-14T02:48:04.310+0000"},{"body":"Issue resolved by pull request 19741\n[https://github.com/apache/spark/pull/19741]","created":"2017-12-06T18:39:35.090+0000"},{"body":"Using this patch, this problem is still exists. When the number of executors is big, and YarnSchedulerBackend.stopped=False after YarnSchedulerBackend.stop() is running, some executor is stoped, and YarnSchedulerBackend.onDisconnected() will be called, then the problem is exists","created":"2017-12-11T13:44:46.388+0000"},{"body":"[~KaiXinXIaoLei], Thanks for checking this. Is the issue you are mentioning different from the two instances mentioned in the PR, can you create a JIRA with the exception stacktrace?","created":"2017-12-11T17:18:22.171+0000"},{"body":"User 'KaiXinXiaoLei' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19945","created":"2017-12-11T23:55:04.990+0000"}],"conversations":[{"body":"When I start 1000 executors, and then stop the process. It will call SparkContext.stop to stop all executors. But during this process, the executors has been killed will lost of rpc with driver, and try to reviveOffers, but can't find CoarseGrainedScheduler or it has been stopped.\n{quote}\n16/03/29 01:45:45 ERROR YarnScheduler: Lost executor 610 on 51-196-152-8: remote Rpc client disassociated\n16/03/29 01:45:45 ERROR Inbox: Ignoring error\norg.apache.spark.SparkException: Could not find CoarseGrainedScheduler or it has been stopped.\n\tat org.apache.spark.rpc.netty.Dispatcher.postMessage(Dispatcher.scala:161)\n\tat org.apache.spark.rpc.netty.Dispatcher.postOneWayMessage(Dispatcher.scala:131)\n\tat org.apache.spark.rpc.netty.NettyRpcEnv.send(NettyRpcEnv.scala:173)\n\tat org.apache.spark.rpc.netty.NettyRpcEndpointRef.send(NettyRpcEnv.scala:398)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend.reviveOffers(CoarseGrainedSchedulerBackend.scala:314)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl.executorLost(TaskSchedulerImpl.scala:482)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend$DriverEndpoint.removeExecutor(CoarseGrainedSchedulerBackend.scala:261)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend$DriverEndpoint$$anonfun$onDisconnected$1.apply(CoarseGrainedSchedulerBackend.scala:207)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend$DriverEndpoint$$anonfun$onDisconnected$1.apply(CoarseGrainedSchedulerBackend.scala:207)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend$DriverEndpoint.onDisconnected(CoarseGrainedSchedulerBackend.scala:207)\n\tat org.apache.spark.rpc.netty.Inbox$$anonfun$process$1.apply$mcV$sp(Inbox.scala:144)\n\tat org.apache.spark.rpc.netty.Inbox.safelyCall(Inbox.scala:204)\n\tat org.apache.spark.rpc.netty.Inbox.process(Inbox.scala:102)\n\tat org.apache.spark.rpc.netty.Dispatcher$MessageLoop.run(Dispatcher.scala:215)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{quote}\n","from":"reporter","subject":"Lost executor of RPC disassociated, and occurs exception: Could not find CoarseGrainedScheduler or it has been stopped"},{"body":"We have similar issues, reporting fake ERROR about \"Could not find xxxx or it has been stopped\".\n\n{code}\n31048 16/03/10 04:28:38 WARN Dispatcher: Message RemoteProcessDisconnected(xxxxxxx:7077) dropped.\norg.apache.spark.SparkException: Could not find OutputCommitCoordinator or it has been stopped.\n at org.apache.spark.rpc.netty.Dispatcher.postMessage(Dispatcher.scala:161)\n at org.apache.spark.rpc.netty.Dispatcher.postToAll(Dispatcher.scala:109)\n at org.apache.spark.rpc.netty.NettyRpcHandler.connectionTerminated(NettyRpcEnv.scala:630)\n at org.apache.spark.network.server.TransportRequestHandler.channelUnregistered(TransportRequestHandler.java:94)\n at org.apache.spark.network.server.TransportChannelHandler.channelUnregistered(TransportChannelHandler.java:89)\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelUnregistered(AbstractChannelHandlerContext.java:158)\n at io.netty.channel.AbstractChannelHandlerContext.fireChannelUnregistered(AbstractChannelHandlerContext.java:144)\n at io.netty.channel.ChannelInboundHandlerAdapter.channelUnregistered(ChannelInboundHandlerAdapter.java:53)\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelUnregistered(AbstractChannelHandlerContext.java:158)\n at io.netty.channel.AbstractChannelHandlerContext.fireChannelUnregistered(AbstractChannelHandlerContext.java:144)\n at io.netty.channel.ChannelInboundHandlerAdapter.channelUnregistered(ChannelInboundHandlerAdapter.java:53)\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelUnregistered(AbstractChannelHandlerContext.java:158)\n at io.netty.channel.AbstractChannelHandlerContext.fireChannelUnregistered(AbstractChannelHandlerContext.java:144)\n at io.netty.channel.ChannelInboundHandlerAdapter.channelUnregistered(ChannelInboundHandlerAdapter.java:53)\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelUnregistered(AbstractChannelHandlerContext.java:158)\n at io.netty.channel.AbstractChannelHandlerContext.fireChannelUnregistered(AbstractChannelHandlerContext.java:144)\n at io.netty.channel.DefaultChannelPipeline.fireChannelUnregistered(DefaultChannelPipeline.java:739)\n at io.netty.channel.AbstractChannel$AbstractUnsafe$8.run(AbstractChannel.java:659)\n at io.netty.util.concurrent.SingleThreadEventExecutor.runAllTasks(SingleThreadEventExecutor.java:357)\n at io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:357)\n at io.netty.util.concurrent.SingleThreadEventExecutor$2.run(SingleThreadEventExecutor.java:111)\n at java.lang.Thread.run(Thread.java:785)\n{code}","from":"developer"},{"body":"Hi, can you specify the version you were working with? I have received the same error in 1.6.0","from":"developer"},{"body":"I had that error on {{1.6.3}}, too.","from":"developer"},{"body":"I had similar issue running while spark in multithread spark context code.\n\n17/09/18 21:40:09 ERROR Inbox: Ignoring error\norg.apache.spark.SparkException: Could not find CoarseGrainedScheduler or it has been stopped.\n\tat org.apache.spark.rpc.netty.Dispatcher.postMessage(Dispatcher.scala:161)\n\tat org.apache.spark.rpc.netty.Dispatcher.postOneWayMessage(Dispatcher.scala:131)\n\tat org.apache.spark.rpc.netty.NettyRpcEnv.send(NettyRpcEnv.scala:193)\n\tat org.apache.spark.rpc.netty.NettyRpcEndpointRef.send(NettyRpcEnv.scala:520)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend.reviveOffers(CoarseGrainedSchedulerBackend.scala:375)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl.executorLost(TaskSchedulerImpl.scala:497)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend$DriverEndpoint.disableExecutor(CoarseGrainedSchedulerBackend.scala:301)\n\tat org.apache.spark.scheduler.cluster.YarnSchedulerBackend$YarnDriverEndpoint$$anonfun$onDisconnected$1.apply(YarnSchedulerBackend.scala:195)\n\tat org.apache.spark.scheduler.cluster.YarnSchedulerBackend$YarnDriverEndpoint$$anonfun$onDisconnected$1.apply(YarnSchedulerBackend.scala:194)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.cluster.YarnSchedulerBackend$YarnDriverEndpoint.onDisconnected(YarnSchedulerBackend.scala:194)\n\tat org.apache.spark.rpc.netty.Inbox$$anonfun$process$1.apply$mcV$sp(Inbox.scala:142)\n\tat org.apache.spark.rpc.netty.Inbox.safelyCall(Inbox.scala:204)\n\tat org.apache.spark.rpc.netty.Inbox.process(Inbox.scala:100)\n\tat org.apache.spark.rpc.netty.Dispatcher$MessageLoop.run(Dispatcher.scala:215)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)","from":"developer"},{"body":"I encountered the same issue in 1.6.3\r\n\r\n{code}\r\n17/11/03 22:38:33 WARN YarnSchedulerBackend$YarnSchedulerEndpoint: Executor for container container_e02_1509517131757_0001_01_000003 exited because of a YARN event (e.g., pre-emption) and not because of an error in the running job.\r\n17/11/03 22:38:33 ERROR YarnClientSchedulerBackend: Could not find CoarseGrainedScheduler or it has been stopped.\r\norg.apache.spark.SparkException: Could not find CoarseGrainedScheduler or it has been stopped.\r\n at org.apache.spark.rpc.netty.Dispatcher.postMessage(Dispatcher.scala:163)\r\n at org.apache.spark.rpc.netty.Dispatcher.postLocalMessage(Dispatcher.scala:128)\r\n at org.apache.spark.rpc.netty.NettyRpcEnv.ask(NettyRpcEnv.scala:231)\r\n at org.apache.spark.rpc.netty.NettyRpcEndpointRef.ask(NettyRpcEnv.scala:515)\r\n at org.apache.spark.rpc.RpcEndpointRef.ask(RpcEndpointRef.scala:62)\r\n at org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend.removeExecutor(CoarseGrainedSchedulerBackend.scala:392)\r\n at org.apache.spark.scheduler.cluster.YarnSchedulerBackend$YarnSchedulerEndpoint$$anonfun$receive$1.applyOrElse(YarnSchedulerBackend.scala:259)\r\n at org.apache.spark.rpc.netty.Inbox$$anonfun$process$1.apply$mcV$sp(Inbox.scala:116)\r\n at org.apache.spark.rpc.netty.Inbox.safelyCall(Inbox.scala:204)\r\n at org.apache.spark.rpc.netty.Inbox.process(Inbox.scala:100)\r\n at org.apache.spark.rpc.netty.Dispatcher$MessageLoop.run(Dispatcher.scala:217)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n at java.lang.Thread.run(Thread.java:745)\r\n17/11/03 22:38:33 WARN YarnSchedulerBackend$YarnSchedulerEndpoint: Executor for container container_e02_1509517131757_0001_01_000002 exited because of a YARN event (e.g., pre-emption) and not because of an error in the running job.\r\n17/11/03 22:38:33 ERROR YarnClientSchedulerBackend: Could not find CoarseGrainedScheduler or it has been stopped.\r\norg.apache.spark.SparkException: Could not find CoarseGrainedScheduler or it has been stopped.\r\n at org.apache.spark.rpc.netty.Dispatcher.postMessage(Dispatcher.scala:163)\r\n at org.apache.spark.rpc.netty.Dispatcher.postLocalMessage(Dispatcher.scala:128)\r\n at org.apache.spark.rpc.netty.NettyRpcEnv.ask(NettyRpcEnv.scala:231)\r\n at org.apache.spark.rpc.netty.NettyRpcEndpointRef.ask(NettyRpcEnv.scala:515)\r\n at org.apache.spark.rpc.RpcEndpointRef.ask(RpcEndpointRef.scala:62)\r\n at org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend.removeExecutor(CoarseGrainedSchedulerBackend.scala:392)\r\n at org.apache.spark.scheduler.cluster.YarnSchedulerBackend$YarnSchedulerEndpoint$$anonfun$receive$1.applyOrElse(YarnSchedulerBackend.scala:259)\r\n at org.apache.spark.rpc.netty.Inbox$$anonfun$process$1.apply$mcV$sp(Inbox.scala:116)\r\n at org.apache.spark.rpc.netty.Inbox.safelyCall(Inbox.scala:204)\r\n at org.apache.spark.rpc.netty.Inbox.process(Inbox.scala:100)\r\n at org.apache.spark.rpc.netty.Dispatcher$MessageLoop.run(Dispatcher.scala:217)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n at java.lang.Thread.run(Thread.java:745)\r\n{code}\r\n","from":"developer"},{"body":"User 'devaraj-kavali' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19741","from":"developer"},{"body":"Issue resolved by pull request 19741\n[https://github.com/apache/spark/pull/19741]","from":"developer"},{"body":"Using this patch, this problem is still exists. When the number of executors is big, and YarnSchedulerBackend.stopped=False after YarnSchedulerBackend.stop() is running, some executor is stoped, and YarnSchedulerBackend.onDisconnected() will be called, then the problem is exists","from":"developer"},{"body":"[~KaiXinXIaoLei], Thanks for checking this. Is the issue you are mentioning different from the two instances mentioned in the PR, can you create a JIRA with the exception stacktrace?","from":"developer"},{"body":"User 'KaiXinXiaoLei' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19945","from":"developer"}],"created":"2016-03-29T02:59:18.000+0000","description":"When I start 1000 executors, and then stop the process. It will call SparkContext.stop to stop all executors. But during this process, the executors has been killed will lost of rpc with driver, and try to reviveOffers, but can't find CoarseGrainedScheduler or it has been stopped.\n{quote}\n16/03/29 01:45:45 ERROR YarnScheduler: Lost executor 610 on 51-196-152-8: remote Rpc client disassociated\n16/03/29 01:45:45 ERROR Inbox: Ignoring error\norg.apache.spark.SparkException: Could not find CoarseGrainedScheduler or it has been stopped.\n\tat org.apache.spark.rpc.netty.Dispatcher.postMessage(Dispatcher.scala:161)\n\tat org.apache.spark.rpc.netty.Dispatcher.postOneWayMessage(Dispatcher.scala:131)\n\tat org.apache.spark.rpc.netty.NettyRpcEnv.send(NettyRpcEnv.scala:173)\n\tat org.apache.spark.rpc.netty.NettyRpcEndpointRef.send(NettyRpcEnv.scala:398)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend.reviveOffers(CoarseGrainedSchedulerBackend.scala:314)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl.executorLost(TaskSchedulerImpl.scala:482)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend$DriverEndpoint.removeExecutor(CoarseGrainedSchedulerBackend.scala:261)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend$DriverEndpoint$$anonfun$onDisconnected$1.apply(CoarseGrainedSchedulerBackend.scala:207)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend$DriverEndpoint$$anonfun$onDisconnected$1.apply(CoarseGrainedSchedulerBackend.scala:207)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.cluster.CoarseGrainedSchedulerBackend$DriverEndpoint.onDisconnected(CoarseGrainedSchedulerBackend.scala:207)\n\tat org.apache.spark.rpc.netty.Inbox$$anonfun$process$1.apply$mcV$sp(Inbox.scala:144)\n\tat org.apache.spark.rpc.netty.Inbox.safelyCall(Inbox.scala:204)\n\tat org.apache.spark.rpc.netty.Inbox.process(Inbox.scala:102)\n\tat org.apache.spark.rpc.netty.Dispatcher$MessageLoop.run(Dispatcher.scala:215)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{quote}\n","issue_id":"12954173","key":"SPARK-14228","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-12-06T18:39:35.000+0000","role":"fixed_distractor","summary":"Lost executor of RPC disassociated, and occurs exception: Could not find CoarseGrainedScheduler or it has been stopped"} {"case_id":"12956001","cluster":"DISTRACTOR-SPARK-14393","comments":[{"body":"Seems different `MonotonicallyIncreasingID` instances share the same stage, that is the same partition id. However, I'm not 100% sure this is a wrong behaviour because we can easily break this monotonically-increasing semantics by using repartition/coalesce after select(monotonicallyIncreasingId()).","created":"2016-04-06T01:55:14.019+0000"},{"body":"If this isn't a bug, I guess it needs to be clear in the documentation what the intended behaviour is here - I assumed that it would guarantee a monotonically-increasing in the end result.","created":"2016-04-06T04:59:11.832+0000"},{"body":"Still happens in the current master/2.0\n\n{code}\nscala> spark.range(10).repartition(5).select(monotonicallyIncreasingId()).coalesce(1).show()\nwarning: there was one deprecation warning; re-run with -deprecation for details\n+-----------------------------+\n|monotonically_increasing_id()|\n+-----------------------------+\n| 0|\n| 0|\n| 1|\n| 2|\n| 0|\n| 1|\n| 0|\n| 1|\n| 2|\n| 0|\n+-----------------------------+\n{code}","created":"2016-10-08T14:05:50.383+0000"},{"body":"[~hyukjin.kwon] What do u think the issue of this ticket should be fixed, or documented?","created":"2016-10-10T00:58:20.323+0000"},{"body":"Thanks for pinging me!, in my personal opinion, we should document this first if we can have some concrete reasons that why it is virtually the limitation in the current implementation.\nI haven't look into this but I *guess* it can be documented as I *guess* it seems pretty easy to reproduce but it seems not fixed for several months.\nI *think* (and as you already know) it is fine to submit a PR if we have a good reason to support the PR whether it is about documentation or the actual fix :)","created":"2016-10-10T01:36:42.892+0000"},{"body":"yea, since this ticket is totally inactive for a while, I think it's okay to document this somewhere and close for now.","created":"2016-10-10T01:48:40.065+0000"},{"body":"Am I correct to understand that I should simply expect `monotonicallyIncreasingId()` to _not_ produce monotonically increasing ids, sometimes?\n\nIf this is the case, then how can one use this function safely?","created":"2016-10-10T17:25:30.001+0000"},{"body":"Since coalesce() just after monotonicallyIncreasingId() breaks the semantics, it's okay not to do that.\nFor example, coalesce() before monotonicallyIncreasingId() is okay;\n{code}\nscala> spark.range(10).repartition(5).coalesce(1).select(monotonicallyIncreasingId()).show\nwarning: there was one deprecation warning; re-run with -deprecation for details\n+-----------------------------+\n|monotonically_increasing_id()|\n+-----------------------------+\n| 0|\n| 1|\n| 2|\n| 3|\n| 4|\n| 5|\n| 6|\n| 7|\n| 8|\n| 9|\n+-----------------------------+\n{code}","created":"2016-10-11T03:47:38.311+0000"},{"body":"This is a bigger issue. It would happen with (`monotonically_increasing_id`, `rand`, `randn`, etc) x (`coalesce`, `union`, etc). The root cause is that the partition ID used to initialize the operator is not the partition ID associated with the DataFrame where the column was originally defined, which is expected by users.\n\ncc [~rxin@databricks.com] [~yhuai]","created":"2016-10-18T07:42:27.990+0000"},{"body":"User 'mengxr' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15567","created":"2016-10-20T10:02:06.318+0000"},{"body":"User 'felixcheung' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15747","created":"2016-11-03T05:26:03.891+0000"},{"body":"Appears this also happens in 2.0.2.\n\nThanks for fixing this!\n\nI disagree that monotonically_increasing_id function should ever be allowed to create duplicate values in a table given its documentation is \"The generated ID is guaranteed to be monotonically increasing and unique, but not consecutive.\" \n\nWe were certainly relying on it to produce unique values!","created":"2017-02-26T01:38:25.002+0000"}],"conversations":[{"body":"When utilising monotonicallyIncreasingId with a coalesce, it appears that every partition uses the same offset (0) leading to non-monotonically increasing IDs.\n\nSee examples below\n\n{code}\n>>> sqlContext.range(10).select(monotonicallyIncreasingId()).show()\n+---------------------------+\n|monotonicallyincreasingid()|\n+---------------------------+\n| 25769803776|\n| 51539607552|\n| 77309411328|\n| 103079215104|\n| 128849018880|\n| 163208757248|\n| 188978561024|\n| 214748364800|\n| 240518168576|\n| 266287972352|\n+---------------------------+\n\n>>> sqlContext.range(10).select(monotonicallyIncreasingId()).coalesce(1).show()\n+---------------------------+\n|monotonicallyincreasingid()|\n+---------------------------+\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n+---------------------------+\n\n>>> sqlContext.range(10).repartition(5).select(monotonicallyIncreasingId()).coalesce(1).show()\n+---------------------------+\n|monotonicallyincreasingid()|\n+---------------------------+\n| 0|\n| 1|\n| 0|\n| 0|\n| 1|\n| 2|\n| 3|\n| 0|\n| 1|\n| 2|\n+---------------------------+\n{code}","from":"reporter","subject":"values generated by non-deterministic functions shouldn't change after coalesce or union"},{"body":"Seems different `MonotonicallyIncreasingID` instances share the same stage, that is the same partition id. However, I'm not 100% sure this is a wrong behaviour because we can easily break this monotonically-increasing semantics by using repartition/coalesce after select(monotonicallyIncreasingId()).","from":"developer"},{"body":"If this isn't a bug, I guess it needs to be clear in the documentation what the intended behaviour is here - I assumed that it would guarantee a monotonically-increasing in the end result.","from":"developer"},{"body":"Still happens in the current master/2.0\n\n{code}\nscala> spark.range(10).repartition(5).select(monotonicallyIncreasingId()).coalesce(1).show()\nwarning: there was one deprecation warning; re-run with -deprecation for details\n+-----------------------------+\n|monotonically_increasing_id()|\n+-----------------------------+\n| 0|\n| 0|\n| 1|\n| 2|\n| 0|\n| 1|\n| 0|\n| 1|\n| 2|\n| 0|\n+-----------------------------+\n{code}","from":"developer"},{"body":"[~hyukjin.kwon] What do u think the issue of this ticket should be fixed, or documented?","from":"developer"},{"body":"Thanks for pinging me!, in my personal opinion, we should document this first if we can have some concrete reasons that why it is virtually the limitation in the current implementation.\nI haven't look into this but I *guess* it can be documented as I *guess* it seems pretty easy to reproduce but it seems not fixed for several months.\nI *think* (and as you already know) it is fine to submit a PR if we have a good reason to support the PR whether it is about documentation or the actual fix :)","from":"developer"},{"body":"yea, since this ticket is totally inactive for a while, I think it's okay to document this somewhere and close for now.","from":"developer"},{"body":"Am I correct to understand that I should simply expect `monotonicallyIncreasingId()` to _not_ produce monotonically increasing ids, sometimes?\n\nIf this is the case, then how can one use this function safely?","from":"developer"},{"body":"Since coalesce() just after monotonicallyIncreasingId() breaks the semantics, it's okay not to do that.\nFor example, coalesce() before monotonicallyIncreasingId() is okay;\n{code}\nscala> spark.range(10).repartition(5).coalesce(1).select(monotonicallyIncreasingId()).show\nwarning: there was one deprecation warning; re-run with -deprecation for details\n+-----------------------------+\n|monotonically_increasing_id()|\n+-----------------------------+\n| 0|\n| 1|\n| 2|\n| 3|\n| 4|\n| 5|\n| 6|\n| 7|\n| 8|\n| 9|\n+-----------------------------+\n{code}","from":"developer"},{"body":"This is a bigger issue. It would happen with (`monotonically_increasing_id`, `rand`, `randn`, etc) x (`coalesce`, `union`, etc). The root cause is that the partition ID used to initialize the operator is not the partition ID associated with the DataFrame where the column was originally defined, which is expected by users.\n\ncc [~rxin@databricks.com] [~yhuai]","from":"developer"},{"body":"User 'mengxr' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15567","from":"developer"},{"body":"User 'felixcheung' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15747","from":"developer"},{"body":"Appears this also happens in 2.0.2.\n\nThanks for fixing this!\n\nI disagree that monotonically_increasing_id function should ever be allowed to create duplicate values in a table given its documentation is \"The generated ID is guaranteed to be monotonically increasing and unique, but not consecutive.\" \n\nWe were certainly relying on it to produce unique values!","from":"developer"}],"created":"2016-04-05T01:04:15.000+0000","description":"When utilising monotonicallyIncreasingId with a coalesce, it appears that every partition uses the same offset (0) leading to non-monotonically increasing IDs.\n\nSee examples below\n\n{code}\n>>> sqlContext.range(10).select(monotonicallyIncreasingId()).show()\n+---------------------------+\n|monotonicallyincreasingid()|\n+---------------------------+\n| 25769803776|\n| 51539607552|\n| 77309411328|\n| 103079215104|\n| 128849018880|\n| 163208757248|\n| 188978561024|\n| 214748364800|\n| 240518168576|\n| 266287972352|\n+---------------------------+\n\n>>> sqlContext.range(10).select(monotonicallyIncreasingId()).coalesce(1).show()\n+---------------------------+\n|monotonicallyincreasingid()|\n+---------------------------+\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n| 0|\n+---------------------------+\n\n>>> sqlContext.range(10).repartition(5).select(monotonicallyIncreasingId()).coalesce(1).show()\n+---------------------------+\n|monotonicallyincreasingid()|\n+---------------------------+\n| 0|\n| 1|\n| 0|\n| 0|\n| 1|\n| 2|\n| 3|\n| 0|\n| 1|\n| 2|\n+---------------------------+\n{code}","issue_id":"12956001","key":"SPARK-14393","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-11-02T18:41:43.000+0000","role":"fixed_distractor","summary":"values generated by non-deterministic functions shouldn't change after coalesce or union"} {"case_id":"12985890","cluster":"DISTRACTOR-SPARK-16334","comments":[{"body":"hi [~epahomov], by which tool were your parquet files written, SparkSQL or ? In addition, what's the {{WriterVersion}}, {{PARQUET_1_0 (\"v1\")}} or {{PARQUET_2_0 (\"v2\")}}?","created":"2016-07-01T05:30:06.972+0000"},{"body":"Hi, we discovered problem with the same stacktrace in Spark 2.0. In our case it's thrown during DataFrame.rdd.aggregate call. Moreover it somehow depends on volume of data, because it is not thrown when we change filter criteria accordingly. We used SparkSQL to write these parquet files and didn't explicitly specify WriterVersion option so I believe whatever version is set by default was used.","created":"2016-07-05T23:11:55.379+0000"},{"body":"Can't reproduce it on 1.6.1 too. Also in our case it looks like it fails in VectorizedColumnReader in INT96 case for TimestampType, which in turn calls PlainValuesDictionary.decodeToBinary(int id).","created":"2016-07-06T15:50:12.870+0000"},{"body":"I believe it relates to the following change in Spark 2.0:\n\n\"For example, we have implemented a new vectorized Parquet reader that does decompression and decoding in column batches. When decoding integer columns (on disk), this new reader is roughly 9 times faster than the non-vectorized one\" taken from this link:\n\nhttps://databricks.com/blog/2016/05/23/apache-spark-as-a-compiler-joining-a-billion-rows-per-second-on-a-laptop.html\n\nCorresponding JIRA tickets:\nhttps://issues.apache.org/jira/browse/SPARK-12854\nhttps://issues.apache.org/jira/browse/SPARK-14008\n","created":"2016-07-11T20:43:28.864+0000"},{"body":"Could you try to disable the vectorized parquet reader. You can do this issuing the following SQL statement {{SET spark.sql.parquet.enableVectorizedReader = FALSE}}, or by issuing {{spark.conf.set(\"spark.sql.parquet.enableVectorizedReader\", \"false\")}} in the REPL.","created":"2016-07-11T20:58:08.961+0000"},{"body":"It would also be great if we can reproduce this. Could share an example?","created":"2016-07-11T21:13:19.269+0000"},{"body":"Hi Herman,\n\nThank you for reply!\n\nI wasn't able to reproduce this error anymore with this spark.sql.parquet.enableVectorizedReader setting set to false in Spark 2.0.\n\nAs for steps to reproduce: unfortunately I have reproduced it on partitioned parquet files containing client-sensitive data (we are calling SparkSession programmatically from our app), so I can't provide data here. Let's see what I can do.","created":"2016-07-12T02:19:10.301+0000"},{"body":"While I've not been able to reproduce this bug, looking at the stack trace, I think one of the likely causes of this is that we're resetting the dictionary while reading every page in a row batch. This is generally okay but might cause problems if a batch has a combination of dictionary-encoded and non-dictionary-encoded pages.","created":"2016-07-15T21:22:26.023+0000"},{"body":"User 'sameeragarwal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14225","created":"2016-07-15T21:41:05.733+0000"},{"body":"[~epakhomov] can you try the patch and see if it fixes your problem? https://github.com/apache/spark/pull/14225\n","created":"2016-07-15T21:55:49.625+0000"},{"body":"Sure, would test on Monday.","created":"2016-07-15T22:43:51.198+0000"},{"body":"cc [~vivanov] too","created":"2016-07-16T00:01:35.848+0000"},{"body":"[~epakhomov] [~lester] any luck with trying out this change? Also in order to prevent future regressions, it'd be great if you can share an illustrative (anonymized) sample of data on which you're seeing this issue so that we can make it part of the test harness.","created":"2016-07-19T18:21:22.857+0000"},{"body":"[~sameerag] that particular case on which I experienced the problem - now working for me. It's hard for me to reproduce what exactly made the bug appear before - would not go into such much trouble. \n\nThanks","created":"2016-07-19T19:45:35.278+0000"},{"body":"[~sameerag] I have just built branch-2.0 which should have included your patch, but I am still experiencing this issue. The parquet file I am using was written using Spark 1.6 and has ~450 columns.\n\nIssuing {{spark.conf.set(\"spark.sql.parquet.enableVectorizedReader\", \"false\")}} prevents the issue from occurring, so it's definitely the vectorized parquet reader.\n\nLet me know if I can provide any additional information to help resolve this issue.","created":"2016-07-28T20:11:54.971+0000"},{"body":"[~sameer] We can reproduce the problem as well. {{spark.conf.set(\"spark.sql.parquet.enableVectorizedReader\", \"false\")}} help us, too. We let you know, when we tried the patch.","created":"2016-08-04T11:47:38.606+0000"},{"body":"[~keith.j.kraus], [~sebastianherold] -- would it be possible for you share the subset of data that's causing this bug (privately, if you'd like)? If that's not really an option, can you please share some more information about its format (row group/ page encoding etc.) by running [parquet-dump | https://github.com/apache/parquet-mr/blob/master/parquet-tools/src/main/scripts/parquet-dump] on your data?","created":"2016-08-04T19:01:40.818+0000"},{"body":"[~sameerag] Sharing even a subset of the data would be very difficult as it has confidential information. I've gone ahead and ran parquet-dump on one part of the parquet (seems like the tool requires pointing it to a specific .parquet file rather than the directory of parts). I'd prefer to share this privately if possible, how can I send you the command output?","created":"2016-08-04T21:27:18.170+0000"},{"body":"Thanks Keith, that'll work. You can mail it to me at sameer@databricks.com.","created":"2016-08-04T21:37:32.812+0000"},{"body":"Seems like a lot of people still have a problem even after suggested fix","created":"2016-08-23T02:56:07.842+0000"},{"body":"(just reason for reopen)","created":"2016-08-23T02:56:43.951+0000"},{"body":"[~sameerag] I just upgraded to spark 2.0 from 1.6.2 and I am now experiencing this issue too.\nMy parquet file was generated in Pig (running on Tez 0.8.4) using pig-parquet-bundle 1.8.1.\nIf you have not been able to reproduce this issue yet, I could send you a sample of data.","created":"2016-08-31T18:48:57.917+0000"},{"body":"Thanks [~tradersancho] that'd really help. The bug is most likely in a particular ordering of plain and dictionary encoded pages within a row-group but in absence of a repro, we're not quite able to lay a finger on it.","created":"2016-08-31T18:57:58.205+0000"},{"body":"Sorry for the late response. We noticed the error on some intermediate results during a hackathon and couldn’t reproduce it until now. But quite sure the error occurred when we wanted to open Spark 1.6-generated Parquet files with thousands or rows (don’t ask) in Spark 2.0. The data was very sparse, because it has been generated out of ugly JSON files. Maybe this helps others for reproduction.\n\n\n","created":"2016-09-01T07:45:42.283+0000"},{"body":"User 'sameeragarwal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14941","created":"2016-09-02T20:12:07.489+0000"},{"body":"[~tradersancho], [~keith.j.kraus] - Thank you once again for sharing your datasets!\n\nIt seems that the exception was caused by re-using the same dictionary column vector while reading multiple row groups one after the other. What made this particularly hard to catch was that this issue manifested only for a particular type of distribution of dictionary/plain encoded data while we read/populate the underlying bit packed dictionary data into a (reusable) column batch. While I've manually verified that this patch fixes the issues that [~tradersancho] experienced, it'd be great if others can try out the fix as well: https://github.com/apache/spark/pull/14941.\n\nIn the meantime, we should also look into adding some randomized parquet distribution generators that can proactively catch edge cases like this.","created":"2016-09-02T20:26:01.203+0000"},{"body":"Issue resolved by pull request 14941\n[https://github.com/apache/spark/pull/14941]","created":"2016-09-02T22:20:57.294+0000"},{"body":"User 'sameeragarwal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14944","created":"2016-09-02T23:17:08.353+0000"}],"conversations":[{"body":"Query:\n\n{code}\nselect * from blabla where user_id = 415706251\n{code}\n\nError:\n\n{code}\n16/06/30 14:07:27 WARN scheduler.TaskSetManager: Lost task 11.0 in stage 0.0 (TID 3, hadoop6): java.lang.ArrayIndexOutOfBoundsException: 6934\n at org.apache.parquet.column.values.dictionary.PlainValuesDictionary$PlainBinaryDictionary.decodeToBinary(PlainValuesDictionary.java:119)\n at org.apache.spark.sql.execution.datasources.parquet.VectorizedColumnReader.decodeDictionaryIds(VectorizedColumnReader.java:273)\n at org.apache.spark.sql.execution.datasources.parquet.VectorizedColumnReader.readBatch(VectorizedColumnReader.java:170)\n at org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.nextBatch(VectorizedParquetRecordReader.java:230)\n at org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.nextKeyValue(VectorizedParquetRecordReader.java:137)\n at org.apache.spark.sql.execution.datasources.RecordReaderIterator.hasNext(RecordReaderIterator.scala:36)\n at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:91)\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.scan_nextBatch$(Unknown Source)\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:246)\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:240)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:780)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:780)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)\n at org.apache.spark.scheduler.Task.run(Task.scala:85)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\n\nWork on 1.6.1","from":"reporter","subject":"SQL query on parquet table java.lang.ArrayIndexOutOfBoundsException"},{"body":"hi [~epahomov], by which tool were your parquet files written, SparkSQL or ? In addition, what's the {{WriterVersion}}, {{PARQUET_1_0 (\"v1\")}} or {{PARQUET_2_0 (\"v2\")}}?","from":"developer"},{"body":"Hi, we discovered problem with the same stacktrace in Spark 2.0. In our case it's thrown during DataFrame.rdd.aggregate call. Moreover it somehow depends on volume of data, because it is not thrown when we change filter criteria accordingly. We used SparkSQL to write these parquet files and didn't explicitly specify WriterVersion option so I believe whatever version is set by default was used.","from":"developer"},{"body":"Can't reproduce it on 1.6.1 too. Also in our case it looks like it fails in VectorizedColumnReader in INT96 case for TimestampType, which in turn calls PlainValuesDictionary.decodeToBinary(int id).","from":"developer"},{"body":"I believe it relates to the following change in Spark 2.0:\n\n\"For example, we have implemented a new vectorized Parquet reader that does decompression and decoding in column batches. When decoding integer columns (on disk), this new reader is roughly 9 times faster than the non-vectorized one\" taken from this link:\n\nhttps://databricks.com/blog/2016/05/23/apache-spark-as-a-compiler-joining-a-billion-rows-per-second-on-a-laptop.html\n\nCorresponding JIRA tickets:\nhttps://issues.apache.org/jira/browse/SPARK-12854\nhttps://issues.apache.org/jira/browse/SPARK-14008\n","from":"developer"},{"body":"Could you try to disable the vectorized parquet reader. You can do this issuing the following SQL statement {{SET spark.sql.parquet.enableVectorizedReader = FALSE}}, or by issuing {{spark.conf.set(\"spark.sql.parquet.enableVectorizedReader\", \"false\")}} in the REPL.","from":"developer"},{"body":"It would also be great if we can reproduce this. Could share an example?","from":"developer"},{"body":"Hi Herman,\n\nThank you for reply!\n\nI wasn't able to reproduce this error anymore with this spark.sql.parquet.enableVectorizedReader setting set to false in Spark 2.0.\n\nAs for steps to reproduce: unfortunately I have reproduced it on partitioned parquet files containing client-sensitive data (we are calling SparkSession programmatically from our app), so I can't provide data here. Let's see what I can do.","from":"developer"},{"body":"While I've not been able to reproduce this bug, looking at the stack trace, I think one of the likely causes of this is that we're resetting the dictionary while reading every page in a row batch. This is generally okay but might cause problems if a batch has a combination of dictionary-encoded and non-dictionary-encoded pages.","from":"developer"},{"body":"User 'sameeragarwal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14225","from":"developer"},{"body":"[~epakhomov] can you try the patch and see if it fixes your problem? https://github.com/apache/spark/pull/14225\n","from":"developer"},{"body":"Sure, would test on Monday.","from":"developer"},{"body":"cc [~vivanov] too","from":"developer"},{"body":"[~epakhomov] [~lester] any luck with trying out this change? Also in order to prevent future regressions, it'd be great if you can share an illustrative (anonymized) sample of data on which you're seeing this issue so that we can make it part of the test harness.","from":"developer"},{"body":"[~sameerag] that particular case on which I experienced the problem - now working for me. It's hard for me to reproduce what exactly made the bug appear before - would not go into such much trouble. \n\nThanks","from":"developer"},{"body":"[~sameerag] I have just built branch-2.0 which should have included your patch, but I am still experiencing this issue. The parquet file I am using was written using Spark 1.6 and has ~450 columns.\n\nIssuing {{spark.conf.set(\"spark.sql.parquet.enableVectorizedReader\", \"false\")}} prevents the issue from occurring, so it's definitely the vectorized parquet reader.\n\nLet me know if I can provide any additional information to help resolve this issue.","from":"developer"},{"body":"[~sameer] We can reproduce the problem as well. {{spark.conf.set(\"spark.sql.parquet.enableVectorizedReader\", \"false\")}} help us, too. We let you know, when we tried the patch.","from":"developer"},{"body":"[~keith.j.kraus], [~sebastianherold] -- would it be possible for you share the subset of data that's causing this bug (privately, if you'd like)? If that's not really an option, can you please share some more information about its format (row group/ page encoding etc.) by running [parquet-dump | https://github.com/apache/parquet-mr/blob/master/parquet-tools/src/main/scripts/parquet-dump] on your data?","from":"developer"},{"body":"[~sameerag] Sharing even a subset of the data would be very difficult as it has confidential information. I've gone ahead and ran parquet-dump on one part of the parquet (seems like the tool requires pointing it to a specific .parquet file rather than the directory of parts). I'd prefer to share this privately if possible, how can I send you the command output?","from":"developer"},{"body":"Thanks Keith, that'll work. You can mail it to me at sameer@databricks.com.","from":"developer"},{"body":"Seems like a lot of people still have a problem even after suggested fix","from":"developer"},{"body":"(just reason for reopen)","from":"developer"},{"body":"[~sameerag] I just upgraded to spark 2.0 from 1.6.2 and I am now experiencing this issue too.\nMy parquet file was generated in Pig (running on Tez 0.8.4) using pig-parquet-bundle 1.8.1.\nIf you have not been able to reproduce this issue yet, I could send you a sample of data.","from":"developer"},{"body":"Thanks [~tradersancho] that'd really help. The bug is most likely in a particular ordering of plain and dictionary encoded pages within a row-group but in absence of a repro, we're not quite able to lay a finger on it.","from":"developer"},{"body":"Sorry for the late response. We noticed the error on some intermediate results during a hackathon and couldn’t reproduce it until now. But quite sure the error occurred when we wanted to open Spark 1.6-generated Parquet files with thousands or rows (don’t ask) in Spark 2.0. The data was very sparse, because it has been generated out of ugly JSON files. Maybe this helps others for reproduction.\n\n\n","from":"developer"},{"body":"User 'sameeragarwal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14941","from":"developer"},{"body":"[~tradersancho], [~keith.j.kraus] - Thank you once again for sharing your datasets!\n\nIt seems that the exception was caused by re-using the same dictionary column vector while reading multiple row groups one after the other. What made this particularly hard to catch was that this issue manifested only for a particular type of distribution of dictionary/plain encoded data while we read/populate the underlying bit packed dictionary data into a (reusable) column batch. While I've manually verified that this patch fixes the issues that [~tradersancho] experienced, it'd be great if others can try out the fix as well: https://github.com/apache/spark/pull/14941.\n\nIn the meantime, we should also look into adding some randomized parquet distribution generators that can proactively catch edge cases like this.","from":"developer"},{"body":"Issue resolved by pull request 14941\n[https://github.com/apache/spark/pull/14941]","from":"developer"},{"body":"User 'sameeragarwal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14944","from":"developer"}],"created":"2016-06-30T21:16:31.000+0000","description":"Query:\n\n{code}\nselect * from blabla where user_id = 415706251\n{code}\n\nError:\n\n{code}\n16/06/30 14:07:27 WARN scheduler.TaskSetManager: Lost task 11.0 in stage 0.0 (TID 3, hadoop6): java.lang.ArrayIndexOutOfBoundsException: 6934\n at org.apache.parquet.column.values.dictionary.PlainValuesDictionary$PlainBinaryDictionary.decodeToBinary(PlainValuesDictionary.java:119)\n at org.apache.spark.sql.execution.datasources.parquet.VectorizedColumnReader.decodeDictionaryIds(VectorizedColumnReader.java:273)\n at org.apache.spark.sql.execution.datasources.parquet.VectorizedColumnReader.readBatch(VectorizedColumnReader.java:170)\n at org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.nextBatch(VectorizedParquetRecordReader.java:230)\n at org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.nextKeyValue(VectorizedParquetRecordReader.java:137)\n at org.apache.spark.sql.execution.datasources.RecordReaderIterator.hasNext(RecordReaderIterator.scala:36)\n at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:91)\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.scan_nextBatch$(Unknown Source)\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:246)\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:240)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:780)\n at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:780)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)\n at org.apache.spark.scheduler.Task.run(Task.scala:85)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\n\nWork on 1.6.1","issue_id":"12985890","key":"SPARK-16334","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-09-02T22:20:57.000+0000","role":"fixed_distractor","summary":"SQL query on parquet table java.lang.ArrayIndexOutOfBoundsException"} {"case_id":"12988033","cluster":"DISTRACTOR-SPARK-16460","comments":[{"body":"Hi, [~marcelboldt]. Thanks for reporting this! I will submit a patch shortly.\n\nA scala reproducer (for reviewers):\n{code}\nimport org.apache.spark.sql.SparkSession\nimport org.apache.spark.sql.types._\n\nobject SPARK_16460 extends App {\n\n val sdf = SparkSession.builder().master(\"local\").getOrCreate().read\n .schema(StructType(List(\n StructField(\"id\", IntegerType),\n StructField(\"d\", DateType),\n StructField(\"dtwo\", DateType))))\n .option(\"inferSchema\", false.toString)\n .option(\"delimiter\", \"|\")\n .option(\"dateFormat\", \"yyyy-MM-dd\")\n .option(\"nullValue\", \"\")\n .option(\"mode\", \"PERMISSIVE\")\n .csv(\"test.csv\")\n\n sdf.show(1)\n\n}\n{code}","created":"2016-07-09T13:23:19.474+0000"},{"body":"User 'lw-lin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14118","created":"2016-07-09T14:25:10.614+0000"},{"body":"Hi [~proflin]:Some time has passed, but I see that your pull request hasn't been merged yet. So I just wanted to confirm that I pulled and compiled it and I am using it since then successfully. Great job, and thanks!","created":"2016-09-10T12:25:10.989+0000"},{"body":"[~marcelboldt] Oh cool! Thanks for the feedback!","created":"2016-09-14T02:57:58.529+0000"},{"body":"Issue resolved by pull request 14118\n[https://github.com/apache/spark/pull/14118]","created":"2016-09-18T18:26:40.390+0000"}],"conversations":[{"body":"Trying to read a CSV file to Spark (using SparkR) containing just this data row:\n\n{code}\n 1|1998-01-01||\n{code}\n\nUsing Spark 1.6.2 (Hadoop 2.6) gives me \n\n{code}\n > head(sdf)\n id d dtwo\n 1 1 1998-01-01 NA\n{code}\n\nSpark 2.0 preview (Hadoop 2.7, Rev. 14308) fails with error: \n\n{panel}\n> Error in invokeJava(isStatic = TRUE, className, methodName, ...) : \n org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 0.0 failed 1 times, most recent failure: Lost task 0.0 in stage 0.0 (TID 0, localhost): java.text.ParseException: Unparseable date: \"\"\n\tat java.text.DateFormat.parse(DateFormat.java:357)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$.castTo(CSVInferSchema.scala:289)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:98)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:74)\n\tat org.apache.spark.sql.execution.datasources.csv.DefaultSource$$anonfun$buildReader$1$$anonfun$apply$1.apply(DefaultSource.scala:124)\n\tat org.apache.spark.sql.execution.datasources.csv.DefaultSource$$anonfun$buildReader$1$$anonfun$apply$1.apply(DefaultSource.scala:124)\n\tat scala.collection.Iterator$$anon$12.nextCur(Iterator.scala:434)\n\tat scala.collection.Iterator$$anon$12.hasNext(Itera...\n{panel}\n\nThe problem seems indeed the NULL value here as with a valid date in the third CSV column it works.\n\nR code:\n{code}\n #Sys.setenv(SPARK_HOME = 'c:/spark/spark-1.6.2-bin-hadoop2.6') \n Sys.setenv(SPARK_HOME = 'C:/spark/spark-2.0.0-preview-bin-hadoop2.7')\n .libPaths(c(file.path(Sys.getenv(\"SPARK_HOME\"), \"R\", \"lib\"), .libPaths()))\n library(SparkR)\n \n sc <-\n sparkR.init(\n master = \"local\",\n sparkPackages = \"com.databricks:spark-csv_2.11:1.4.0\"\n )\n sqlContext <- sparkRSQL.init(sc)\n \n \n st <- structType(structField(\"id\", \"integer\"), structField(\"d\", \"date\"), structField(\"dtwo\", \"date\"))\n \n sdf <- read.df(\n sqlContext,\n path = \"d:/date_test.csv\",\n source = \"com.databricks.spark.csv\",\n schema = st,\n inferSchema = \"false\",\n delimiter = \"|\",\n dateFormat = \"yyyy-MM-dd\",\n nullValue = \"\",\n mode = \"PERMISSIVE\"\n )\n \n head(sdf)\n \n sparkR.stop()\n{code}","from":"reporter","subject":"Spark 2.0 CSV ignores NULL value in Date format"},{"body":"Hi, [~marcelboldt]. Thanks for reporting this! I will submit a patch shortly.\n\nA scala reproducer (for reviewers):\n{code}\nimport org.apache.spark.sql.SparkSession\nimport org.apache.spark.sql.types._\n\nobject SPARK_16460 extends App {\n\n val sdf = SparkSession.builder().master(\"local\").getOrCreate().read\n .schema(StructType(List(\n StructField(\"id\", IntegerType),\n StructField(\"d\", DateType),\n StructField(\"dtwo\", DateType))))\n .option(\"inferSchema\", false.toString)\n .option(\"delimiter\", \"|\")\n .option(\"dateFormat\", \"yyyy-MM-dd\")\n .option(\"nullValue\", \"\")\n .option(\"mode\", \"PERMISSIVE\")\n .csv(\"test.csv\")\n\n sdf.show(1)\n\n}\n{code}","from":"developer"},{"body":"User 'lw-lin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14118","from":"developer"},{"body":"Hi [~proflin]:Some time has passed, but I see that your pull request hasn't been merged yet. So I just wanted to confirm that I pulled and compiled it and I am using it since then successfully. Great job, and thanks!","from":"developer"},{"body":"[~marcelboldt] Oh cool! Thanks for the feedback!","from":"developer"},{"body":"Issue resolved by pull request 14118\n[https://github.com/apache/spark/pull/14118]","from":"developer"}],"created":"2016-07-09T12:07:48.000+0000","description":"Trying to read a CSV file to Spark (using SparkR) containing just this data row:\n\n{code}\n 1|1998-01-01||\n{code}\n\nUsing Spark 1.6.2 (Hadoop 2.6) gives me \n\n{code}\n > head(sdf)\n id d dtwo\n 1 1 1998-01-01 NA\n{code}\n\nSpark 2.0 preview (Hadoop 2.7, Rev. 14308) fails with error: \n\n{panel}\n> Error in invokeJava(isStatic = TRUE, className, methodName, ...) : \n org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 0.0 failed 1 times, most recent failure: Lost task 0.0 in stage 0.0 (TID 0, localhost): java.text.ParseException: Unparseable date: \"\"\n\tat java.text.DateFormat.parse(DateFormat.java:357)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$.castTo(CSVInferSchema.scala:289)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:98)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:74)\n\tat org.apache.spark.sql.execution.datasources.csv.DefaultSource$$anonfun$buildReader$1$$anonfun$apply$1.apply(DefaultSource.scala:124)\n\tat org.apache.spark.sql.execution.datasources.csv.DefaultSource$$anonfun$buildReader$1$$anonfun$apply$1.apply(DefaultSource.scala:124)\n\tat scala.collection.Iterator$$anon$12.nextCur(Iterator.scala:434)\n\tat scala.collection.Iterator$$anon$12.hasNext(Itera...\n{panel}\n\nThe problem seems indeed the NULL value here as with a valid date in the third CSV column it works.\n\nR code:\n{code}\n #Sys.setenv(SPARK_HOME = 'c:/spark/spark-1.6.2-bin-hadoop2.6') \n Sys.setenv(SPARK_HOME = 'C:/spark/spark-2.0.0-preview-bin-hadoop2.7')\n .libPaths(c(file.path(Sys.getenv(\"SPARK_HOME\"), \"R\", \"lib\"), .libPaths()))\n library(SparkR)\n \n sc <-\n sparkR.init(\n master = \"local\",\n sparkPackages = \"com.databricks:spark-csv_2.11:1.4.0\"\n )\n sqlContext <- sparkRSQL.init(sc)\n \n \n st <- structType(structField(\"id\", \"integer\"), structField(\"d\", \"date\"), structField(\"dtwo\", \"date\"))\n \n sdf <- read.df(\n sqlContext,\n path = \"d:/date_test.csv\",\n source = \"com.databricks.spark.csv\",\n schema = st,\n inferSchema = \"false\",\n delimiter = \"|\",\n dateFormat = \"yyyy-MM-dd\",\n nullValue = \"\",\n mode = \"PERMISSIVE\"\n )\n \n head(sdf)\n \n sparkR.stop()\n{code}","issue_id":"12988033","key":"SPARK-16460","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-09-18T18:26:39.000+0000","role":"fixed_distractor","summary":"Spark 2.0 CSV ignores NULL value in Date format"} {"case_id":"12990565","cluster":"DISTRACTOR-SPARK-16613","comments":[{"body":"Thank you for amazingly clean reporting. I could easily regenerate that. In 1.6 branch, the problem still exists. In the current master, I experienced `StackOverflowError`.\n{code}\nscala> val fstRdd = sc.parallelize(List(\"hi\", \"hello\", \"how\", \"are\", \"you\"))\nfstRdd: org.apache.spark.rdd.RDD[String] = ParallelCollectionRDD[0] at parallelize at :24\n\nscala> val pipeRdd = fstRdd.pipe(\"/Users/dongjoon/spark/len.sh\")\njava.lang.StackOverflowError\n at com.fasterxml.jackson.databind.deser.impl.PropertyBasedCreator.startBuilding(PropertyBasedCreator.java:130)\n{code}\nIt's not final investigation, but it's worth to look inside.","created":"2016-07-19T06:08:05.364+0000"},{"body":"Hi [~finkel], [~dongjoon], I think this relates to how many partitions we'll get after we {{sc.parallelize(...)}}.\nFor instance, {{sc.parallelize(..., 5)}} produces:\n{code}\n2\n5\n3\n3\n3\n{code}\n{{sc.parallelize(..., 8)}} produces:\n{code}\n0\n2\n0\n5\n3\n0\n3\n3\n{code}\nAnd be careful, {{sc.parallelize(..., 1)}} produces only:\n{code}\n2\n{code}","created":"2016-07-19T07:04:12.660+0000"},{"body":"Oh. Thank you for analysis!","created":"2016-07-19T07:07:15.619+0000"},{"body":"Then, it seems not a bug in 1.6.2.","created":"2016-07-19T07:08:10.054+0000"},{"body":"[~dongjoon], at about the same time we reproduced the {{StackOverflowError}} from {{RDD.pipe()}}, :)\nAnd besides that stack overflow, {{RDD.pipe(String)}} also doesn't work with commands with options like \"wc -l\" in 2.0, so I filed SPARK-16620 aiming to fix both issues.","created":"2016-07-19T07:17:24.484+0000"},{"body":"Yep. I thought so. You're very fast! Great.","created":"2016-07-19T07:26:41.614+0000"},{"body":"Interesting, I think an issue here is that the semantics of RDD.pipe() are unclear. Input and output are RDD[String], and it seems like each element of the output must be stdout of running the command on a single input element.\n\nIn reality, the process is run once per partition, and each input element is sent to the stdin, separated by newlines. The lines of the output are parsed as the output of the partition -- which means that a process which outputs many lines of output results in many elements of the RDD.\n\nRight now, the process is still invoked for an empty partition, and for this script, correctly results in \"0\".\n\nThe docs do say \"pipes elements to an external process\" which sort of implies the current behavior.\n\nI think we should, in any event, clarify docs and probably also modify the behavior to output nothing for an empty partition -- not even the result of the process when presented with no input.\n\nI'm reluctant to change the semantics of the method to run one process per input, even if that strikes me as more logical. This means we're kind of stuck with this problem that it's not necessarily possible to match outputs 1:1 with inputs, but that's just a constraint on the type of command you can use with this I guess.\n\nCC [~andrewor14] if available, but moreso [~tejasp]","created":"2016-07-19T09:11:27.627+0000"},{"body":"User 'srowen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14260","created":"2016-07-19T10:49:04.848+0000"},{"body":"After thinking about this more I don't think we should break the old semantics. We can document it though","created":"2016-07-19T18:14:15.925+0000"},{"body":"[~srowen] , [~rxin] : I feel that invoking the pipe command even for empty partitions is bad. Users should be exposed to Spark partitions. Leaving the existing behavior as it is might mean that if number of partitions change OR partitioning scheme changes, the results generated for the same input data can differ. This might be annoying. One would expect Spark's behavior to be same as running the pipe command in standalone way over terminal in which case there won't be any empty partitions as partition is a Spark internal concept.\n\nHive's ScriptOperator invokes the user binary only when it sees a row : https://github.com/apache/hive/blob/master/ql/src/java/org/apache/hadoop/hive/ql/exec/ScriptOperator.java#L339 \n\nSpark's ScriptTransformation as well follows the same convention:\nhttps://github.com/apache/spark/blob/master/sql/hive/src/main/scala/org/apache/spark/sql/hive/execution/ScriptTransformation.scala#L244","created":"2016-07-20T03:05:38.779+0000"},{"body":"But the problem of result changing depending on partitioning always exist -- it is there because how the underlying commands are handling the inputs; it's not there because Spark calls the command based on the input. \n\nWe can add a flag to control it if we really want that behavior.\n","created":"2016-07-20T05:13:38.603+0000"},{"body":"Yeah it's a tough call. The current behavior is at least consistent: entirely partition-oriented, one process per partition exactly, always. I agree it's not quite what I'd expect, but maybe the first thing we can do now is at least update the docs without changing the behavior.","created":"2016-07-20T08:43:36.593+0000"}],"conversations":[{"body":"Suppose we have such Spark code\n\n{code}\nobject PipeExample {\n def main(args: Array[String]) {\n\n val fstRdd = sc.parallelize(List(\"hi\", \"hello\", \"how\", \"are\", \"you\"))\n val pipeRdd = fstRdd.pipe(\"/Users/finkel/spark-pipe-example/src/main/resources/len.sh\")\n\n pipeRdd.collect.foreach(println)\n }\n}\n{code}\n\nIt uses a bash script to convert a string to its length.\n\n{code}\n#!/bin/sh\nread input\nlen=${#input}\necho $len\n{code}\n\nSo far so good, but when I run the code, it prints incorrect output. For example:\n\n{code}\n0\n2\n0\n5\n3\n0\n3\n3\n{code}\n\nI expect to see\n\n{code}\n2\n5\n3\n3\n3\n{code}\n\nwhich is correct output for the app. I think it's a bug. It's expected to see only positive integers and avoid zeros.\n\nEnvironment:\n\n1. Spark version is 1.6.2\n2. Scala version is 2.11.6\n","from":"reporter","subject":"RDD.pipe returns values for empty partitions"},{"body":"Thank you for amazingly clean reporting. I could easily regenerate that. In 1.6 branch, the problem still exists. In the current master, I experienced `StackOverflowError`.\n{code}\nscala> val fstRdd = sc.parallelize(List(\"hi\", \"hello\", \"how\", \"are\", \"you\"))\nfstRdd: org.apache.spark.rdd.RDD[String] = ParallelCollectionRDD[0] at parallelize at :24\n\nscala> val pipeRdd = fstRdd.pipe(\"/Users/dongjoon/spark/len.sh\")\njava.lang.StackOverflowError\n at com.fasterxml.jackson.databind.deser.impl.PropertyBasedCreator.startBuilding(PropertyBasedCreator.java:130)\n{code}\nIt's not final investigation, but it's worth to look inside.","from":"developer"},{"body":"Hi [~finkel], [~dongjoon], I think this relates to how many partitions we'll get after we {{sc.parallelize(...)}}.\nFor instance, {{sc.parallelize(..., 5)}} produces:\n{code}\n2\n5\n3\n3\n3\n{code}\n{{sc.parallelize(..., 8)}} produces:\n{code}\n0\n2\n0\n5\n3\n0\n3\n3\n{code}\nAnd be careful, {{sc.parallelize(..., 1)}} produces only:\n{code}\n2\n{code}","from":"developer"},{"body":"Oh. Thank you for analysis!","from":"developer"},{"body":"Then, it seems not a bug in 1.6.2.","from":"developer"},{"body":"[~dongjoon], at about the same time we reproduced the {{StackOverflowError}} from {{RDD.pipe()}}, :)\nAnd besides that stack overflow, {{RDD.pipe(String)}} also doesn't work with commands with options like \"wc -l\" in 2.0, so I filed SPARK-16620 aiming to fix both issues.","from":"developer"},{"body":"Yep. I thought so. You're very fast! Great.","from":"developer"},{"body":"Interesting, I think an issue here is that the semantics of RDD.pipe() are unclear. Input and output are RDD[String], and it seems like each element of the output must be stdout of running the command on a single input element.\n\nIn reality, the process is run once per partition, and each input element is sent to the stdin, separated by newlines. The lines of the output are parsed as the output of the partition -- which means that a process which outputs many lines of output results in many elements of the RDD.\n\nRight now, the process is still invoked for an empty partition, and for this script, correctly results in \"0\".\n\nThe docs do say \"pipes elements to an external process\" which sort of implies the current behavior.\n\nI think we should, in any event, clarify docs and probably also modify the behavior to output nothing for an empty partition -- not even the result of the process when presented with no input.\n\nI'm reluctant to change the semantics of the method to run one process per input, even if that strikes me as more logical. This means we're kind of stuck with this problem that it's not necessarily possible to match outputs 1:1 with inputs, but that's just a constraint on the type of command you can use with this I guess.\n\nCC [~andrewor14] if available, but moreso [~tejasp]","from":"developer"},{"body":"User 'srowen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14260","from":"developer"},{"body":"After thinking about this more I don't think we should break the old semantics. We can document it though","from":"developer"},{"body":"[~srowen] , [~rxin] : I feel that invoking the pipe command even for empty partitions is bad. Users should be exposed to Spark partitions. Leaving the existing behavior as it is might mean that if number of partitions change OR partitioning scheme changes, the results generated for the same input data can differ. This might be annoying. One would expect Spark's behavior to be same as running the pipe command in standalone way over terminal in which case there won't be any empty partitions as partition is a Spark internal concept.\n\nHive's ScriptOperator invokes the user binary only when it sees a row : https://github.com/apache/hive/blob/master/ql/src/java/org/apache/hadoop/hive/ql/exec/ScriptOperator.java#L339 \n\nSpark's ScriptTransformation as well follows the same convention:\nhttps://github.com/apache/spark/blob/master/sql/hive/src/main/scala/org/apache/spark/sql/hive/execution/ScriptTransformation.scala#L244","from":"developer"},{"body":"But the problem of result changing depending on partitioning always exist -- it is there because how the underlying commands are handling the inputs; it's not there because Spark calls the command based on the input. \n\nWe can add a flag to control it if we really want that behavior.\n","from":"developer"},{"body":"Yeah it's a tough call. The current behavior is at least consistent: entirely partition-oriented, one process per partition exactly, always. I agree it's not quite what I'd expect, but maybe the first thing we can do now is at least update the docs without changing the behavior.","from":"developer"}],"created":"2016-07-18T21:37:03.000+0000","description":"Suppose we have such Spark code\n\n{code}\nobject PipeExample {\n def main(args: Array[String]) {\n\n val fstRdd = sc.parallelize(List(\"hi\", \"hello\", \"how\", \"are\", \"you\"))\n val pipeRdd = fstRdd.pipe(\"/Users/finkel/spark-pipe-example/src/main/resources/len.sh\")\n\n pipeRdd.collect.foreach(println)\n }\n}\n{code}\n\nIt uses a bash script to convert a string to its length.\n\n{code}\n#!/bin/sh\nread input\nlen=${#input}\necho $len\n{code}\n\nSo far so good, but when I run the code, it prints incorrect output. For example:\n\n{code}\n0\n2\n0\n5\n3\n0\n3\n3\n{code}\n\nI expect to see\n\n{code}\n2\n5\n3\n3\n3\n{code}\n\nwhich is correct output for the app. I think it's a bug. It's expected to see only positive integers and avoid zeros.\n\nEnvironment:\n\n1. Spark version is 1.6.2\n2. Scala version is 2.11.6\n","issue_id":"12990565","key":"SPARK-16613","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-07-20T16:49:00.000+0000","role":"fixed_distractor","summary":"RDD.pipe returns values for empty partitions"} {"case_id":"12992127","cluster":"DISTRACTOR-SPARK-16700","comments":[{"body":"I dug into this a bit more: \n\n{{_verify_type({}, struct_schema)}} was already raising a similar exception in Spark 1.6.2, however schema validation wasn't being enforced at all by {{createDataFrame}} : https://github.com/apache/spark/blob/branch-1.6/python/pyspark/sql/context.py#L418\n\nIn 2.0.0, it seems that it is done over each row of the data:\nhttps://github.com/apache/spark/blob/master/python/pyspark/sql/session.py#L504\n\nI think there are 2 issues that should be fixed here:\n - {{_verify_type({}, struct_schema)}} shouldn't raise, because as far as I can tell dicts behave as expected and have their items correctly mapped as struct fields.\n - There should be a way to go back to 1.6.x-like behaviour and disable schema verification in {{createDataFrame}}. The {{prepare()}} function is being map()'d over all the data coming from Python, which I think will definitely hurt performance for large datasets and complex schemas. Leaving it on by default but adding a flag to disable it would be a good solution. Without this users will probably have to implement their own {{createDataFrame}} function like I did.\n","created":"2016-07-25T16:16:42.878+0000"},{"body":"When using `Row` object, but with multiple struct types, also returns similar error:\n\n{code}\n_struct = [\n SparkTypes.StructField('string_field', SparkTypes.StringType(), True),\n SparkTypes.StructField('long_field', SparkTypes.LongType(), True),\n SparkTypes.StructField('double_field', SparkTypes.DoubleType(), True)\n]\n_rdd = sc.parallelize([Row(string_field='1', long_field=1, double_field=1.1)])\n\n## Both methods do not work:\n# _schema = SparkTypes.StructType()\n# for _s in _struct:\n# _schema.add(_s)\n_schema = SparkTypes.StructType(_struct)\n\n_df = sqlContext.createDataFrame(_rdd, schema=_schema)\n_df.take(1)\n{code}\n\nReturned error:\n\n{code}\nDoubleType can not accept object '1' in type \n{code}","created":"2016-07-30T12:32:50.078+0000"},{"body":"There are two separate problems here:\n\n1) Spark 2.0 enforce data type checking when creating a DataFrame, it's safer but slower. It makes sense to have a flag for that (on by default)\n\n2) Row object is similar to named tuple (not dict), the columns are ordered. When it's created in a way like dict, we have no way to know the order of columns, so they are sorted by name, then it does not match with the schema provided. We should check the schema (order of columns) when create a DataFrame from RDD of Row (we assume they matched)","created":"2016-08-01T20:18:36.025+0000"},{"body":"Sent PR https://github.com/apache/spark/pull/14469 to address these, could you help to test and review them?","created":"2016-08-02T23:51:04.977+0000"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14469","created":"2016-08-02T23:51:06.041+0000"},{"body":"The verifySchema flag works great, and the {{dict}} issue seems to be fixed for me. Thanks a lot!!","created":"2016-08-03T22:43:48.389+0000"},{"body":"Issue resolved by pull request 14469\n[https://github.com/apache/spark/pull/14469]","created":"2016-08-15T19:41:49.333+0000"}],"conversations":[{"body":"Hello,\n\nI found this issue while testing my codebase with 2.0.0-rc5\n\nStructType in Spark 1.6.2 accepts the Python type, which is very handy. 2.0.0-rc5 does not and throws an error.\n\nI don't know if this was intended but I'd advocate for this behaviour to remain the same. MapType is probably wasteful when your key names never change and switching to Python tuples would be cumbersome.\n\nHere is a minimal script to reproduce the issue: \n\n{code}\nfrom pyspark import SparkContext\nfrom pyspark.sql import types as SparkTypes\nfrom pyspark.sql import SQLContext\n\n\nsc = SparkContext()\nsqlc = SQLContext(sc)\n\nstruct_schema = SparkTypes.StructType([\n SparkTypes.StructField(\"id\", SparkTypes.LongType())\n])\n\nrdd = sc.parallelize([{\"id\": 0}, {\"id\": 1}])\n\ndf = sqlc.createDataFrame(rdd, struct_schema)\n\nprint df.collect()\n\n# 1.6.2 prints [Row(id=0), Row(id=1)]\n\n# 2.0.0-rc5 raises TypeError: StructType can not accept object {'id': 0} in type \n\n{code}\n\nThanks!","from":"reporter","subject":"StructType doesn't accept Python dicts anymore"},{"body":"I dug into this a bit more: \n\n{{_verify_type({}, struct_schema)}} was already raising a similar exception in Spark 1.6.2, however schema validation wasn't being enforced at all by {{createDataFrame}} : https://github.com/apache/spark/blob/branch-1.6/python/pyspark/sql/context.py#L418\n\nIn 2.0.0, it seems that it is done over each row of the data:\nhttps://github.com/apache/spark/blob/master/python/pyspark/sql/session.py#L504\n\nI think there are 2 issues that should be fixed here:\n - {{_verify_type({}, struct_schema)}} shouldn't raise, because as far as I can tell dicts behave as expected and have their items correctly mapped as struct fields.\n - There should be a way to go back to 1.6.x-like behaviour and disable schema verification in {{createDataFrame}}. The {{prepare()}} function is being map()'d over all the data coming from Python, which I think will definitely hurt performance for large datasets and complex schemas. Leaving it on by default but adding a flag to disable it would be a good solution. Without this users will probably have to implement their own {{createDataFrame}} function like I did.\n","from":"developer"},{"body":"When using `Row` object, but with multiple struct types, also returns similar error:\n\n{code}\n_struct = [\n SparkTypes.StructField('string_field', SparkTypes.StringType(), True),\n SparkTypes.StructField('long_field', SparkTypes.LongType(), True),\n SparkTypes.StructField('double_field', SparkTypes.DoubleType(), True)\n]\n_rdd = sc.parallelize([Row(string_field='1', long_field=1, double_field=1.1)])\n\n## Both methods do not work:\n# _schema = SparkTypes.StructType()\n# for _s in _struct:\n# _schema.add(_s)\n_schema = SparkTypes.StructType(_struct)\n\n_df = sqlContext.createDataFrame(_rdd, schema=_schema)\n_df.take(1)\n{code}\n\nReturned error:\n\n{code}\nDoubleType can not accept object '1' in type \n{code}","from":"developer"},{"body":"There are two separate problems here:\n\n1) Spark 2.0 enforce data type checking when creating a DataFrame, it's safer but slower. It makes sense to have a flag for that (on by default)\n\n2) Row object is similar to named tuple (not dict), the columns are ordered. When it's created in a way like dict, we have no way to know the order of columns, so they are sorted by name, then it does not match with the schema provided. We should check the schema (order of columns) when create a DataFrame from RDD of Row (we assume they matched)","from":"developer"},{"body":"Sent PR https://github.com/apache/spark/pull/14469 to address these, could you help to test and review them?","from":"developer"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14469","from":"developer"},{"body":"The verifySchema flag works great, and the {{dict}} issue seems to be fixed for me. Thanks a lot!!","from":"developer"},{"body":"Issue resolved by pull request 14469\n[https://github.com/apache/spark/pull/14469]","from":"developer"}],"created":"2016-07-25T01:44:52.000+0000","description":"Hello,\n\nI found this issue while testing my codebase with 2.0.0-rc5\n\nStructType in Spark 1.6.2 accepts the Python type, which is very handy. 2.0.0-rc5 does not and throws an error.\n\nI don't know if this was intended but I'd advocate for this behaviour to remain the same. MapType is probably wasteful when your key names never change and switching to Python tuples would be cumbersome.\n\nHere is a minimal script to reproduce the issue: \n\n{code}\nfrom pyspark import SparkContext\nfrom pyspark.sql import types as SparkTypes\nfrom pyspark.sql import SQLContext\n\n\nsc = SparkContext()\nsqlc = SQLContext(sc)\n\nstruct_schema = SparkTypes.StructType([\n SparkTypes.StructField(\"id\", SparkTypes.LongType())\n])\n\nrdd = sc.parallelize([{\"id\": 0}, {\"id\": 1}])\n\ndf = sqlc.createDataFrame(rdd, struct_schema)\n\nprint df.collect()\n\n# 1.6.2 prints [Row(id=0), Row(id=1)]\n\n# 2.0.0-rc5 raises TypeError: StructType can not accept object {'id': 0} in type \n\n{code}\n\nThanks!","issue_id":"12992127","key":"SPARK-16700","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-08-15T19:41:48.000+0000","role":"fixed_distractor","summary":"StructType doesn't accept Python dicts anymore"} {"case_id":"12993828","cluster":"DISTRACTOR-SPARK-16826","comments":[{"body":"The contention is in the JDK class java.net.URL itself, and it does look like an unfortunate bottleneck from ages ago. There's a static, globally synchronized Hashtable here, which explains why multiple JVM (executors) alleviates the problem. We can't fix that directly. See also http://www.supermind.org/blog/580/java-net-url-synchronization-bottleneck\n\nWe also probably have to rely on the behavior of this class to parse a URL. I don't see a good alternative here? You can avoid the lookup if you pass in the URLStreamHandler manually which would mean reimplementing some of the same logic in URL. That then means the overhead of looking up a SecurityManager though.\n\nWhat about just avoiding a load of calls to parseURL at once, is that at all reasonable? like rearranging the pipeline to filter or remove duplicates earlier upstream?","created":"2016-08-01T14:24:53.151+0000"},{"body":"[~srowen] thanks for the pointers! \n\nI'm parsing every hyperlink found in Common Crawl, so there are billions of unique ones, no way around it.\n\nWouldn't it be possible to switch to another implementation with an API similar to java.net.URL? As I understand it we never need the URLStreamHandler in the first place anyway?\n\nI'm not a Java expert but what about {{java.net.URI}} or {{org.apache.catalina.util.URL}} for instance?\n","created":"2016-08-02T01:14:26.178+0000"},{"body":"URI.toURL just follows the same code path. Does URI itself parse all the same fields? Didn't think so because URIs are a superset of URLs.\n\nDefinitely open to suggestions. Anything that can parse the same fields respectably is OK.","created":"2016-08-02T02:04:27.263+0000"},{"body":"Sorry I can't be more helpful on the Java side... But I think there must be some high-quality URL parsing code somewhere in the Apache foundation already :-)","created":"2016-08-02T02:21:19.525+0000"},{"body":"[~srowen] what about this? \nhttps://github.com/sylvinus/spark/commit/98119a08368b1cd1faf3f25a32910ad6717c5c02\n\nThe tests seem to pass and I don't think it uses the problematic code paths in java.net.URL (except for getFile but that may be could be fixed easily)","created":"2016-08-02T02:41:26.395+0000"},{"body":"User 'sylvinus' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14488","created":"2016-08-04T00:09:07.613+0000"},{"body":"Issue resolved by pull request 14488\n[https://github.com/apache/spark/pull/14488]","created":"2016-08-05T19:56:09.053+0000"}],"conversations":[{"body":"Hello!\n\nI'm using {{c4.8xlarge}} instances on EC2 with 36 cores and doing lots of {{parse_url(url, \"host\")}} in Spark SQL.\n\nUnfortunately it seems that there is an internal thread-safe cache in there, and the instances end up being 90% idle.\n\nWhen I view the thread dump for my executors, most of the executor threads are \"BLOCKED\", in that state:\n{code}\njava.util.Hashtable.get(Hashtable.java:362)\njava.net.URL.getURLStreamHandler(URL.java:1135)\njava.net.URL.(URL.java:599)\njava.net.URL.(URL.java:490)\njava.net.URL.(URL.java:439)\norg.apache.spark.sql.catalyst.expressions.ParseUrl.getUrl(stringExpressions.scala:731)\norg.apache.spark.sql.catalyst.expressions.ParseUrl.parseUrlWithoutKey(stringExpressions.scala:772)\norg.apache.spark.sql.catalyst.expressions.ParseUrl.eval(stringExpressions.scala:785)\norg.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificPredicate.eval(Unknown Source)\norg.apache.spark.sql.catalyst.expressions.codegen.GeneratePredicate$$anonfun$create$2.apply(GeneratePredicate.scala:69)\norg.apache.spark.sql.catalyst.expressions.codegen.GeneratePredicate$$anonfun$create$2.apply(GeneratePredicate.scala:69)\norg.apache.spark.sql.execution.FilterExec$$anonfun$17$$anonfun$apply$2.apply(basicPhysicalOperators.scala:203)\norg.apache.spark.sql.execution.FilterExec$$anonfun$17$$anonfun$apply$2.apply(basicPhysicalOperators.scala:202)\nscala.collection.Iterator$$anon$13.hasNext(Iterator.scala:463)\norg.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\norg.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\norg.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\nscala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\norg.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java:147)\norg.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\norg.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\norg.apache.spark.scheduler.Task.run(Task.scala:85)\norg.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\njava.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\njava.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\njava.lang.Thread.run(Thread.java:745)\n{code}\n\nHowever, when I switch from 1 executor with 36 cores to 9 executors with 4 cores, throughput is almost 10x higher and the CPUs are back at ~100% use.\n\nThanks!","from":"reporter","subject":"java.util.Hashtable limits the throughput of PARSE_URL()"},{"body":"The contention is in the JDK class java.net.URL itself, and it does look like an unfortunate bottleneck from ages ago. There's a static, globally synchronized Hashtable here, which explains why multiple JVM (executors) alleviates the problem. We can't fix that directly. See also http://www.supermind.org/blog/580/java-net-url-synchronization-bottleneck\n\nWe also probably have to rely on the behavior of this class to parse a URL. I don't see a good alternative here? You can avoid the lookup if you pass in the URLStreamHandler manually which would mean reimplementing some of the same logic in URL. That then means the overhead of looking up a SecurityManager though.\n\nWhat about just avoiding a load of calls to parseURL at once, is that at all reasonable? like rearranging the pipeline to filter or remove duplicates earlier upstream?","from":"developer"},{"body":"[~srowen] thanks for the pointers! \n\nI'm parsing every hyperlink found in Common Crawl, so there are billions of unique ones, no way around it.\n\nWouldn't it be possible to switch to another implementation with an API similar to java.net.URL? As I understand it we never need the URLStreamHandler in the first place anyway?\n\nI'm not a Java expert but what about {{java.net.URI}} or {{org.apache.catalina.util.URL}} for instance?\n","from":"developer"},{"body":"URI.toURL just follows the same code path. Does URI itself parse all the same fields? Didn't think so because URIs are a superset of URLs.\n\nDefinitely open to suggestions. Anything that can parse the same fields respectably is OK.","from":"developer"},{"body":"Sorry I can't be more helpful on the Java side... But I think there must be some high-quality URL parsing code somewhere in the Apache foundation already :-)","from":"developer"},{"body":"[~srowen] what about this? \nhttps://github.com/sylvinus/spark/commit/98119a08368b1cd1faf3f25a32910ad6717c5c02\n\nThe tests seem to pass and I don't think it uses the problematic code paths in java.net.URL (except for getFile but that may be could be fixed easily)","from":"developer"},{"body":"User 'sylvinus' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14488","from":"developer"},{"body":"Issue resolved by pull request 14488\n[https://github.com/apache/spark/pull/14488]","from":"developer"}],"created":"2016-07-31T21:31:05.000+0000","description":"Hello!\n\nI'm using {{c4.8xlarge}} instances on EC2 with 36 cores and doing lots of {{parse_url(url, \"host\")}} in Spark SQL.\n\nUnfortunately it seems that there is an internal thread-safe cache in there, and the instances end up being 90% idle.\n\nWhen I view the thread dump for my executors, most of the executor threads are \"BLOCKED\", in that state:\n{code}\njava.util.Hashtable.get(Hashtable.java:362)\njava.net.URL.getURLStreamHandler(URL.java:1135)\njava.net.URL.(URL.java:599)\njava.net.URL.(URL.java:490)\njava.net.URL.(URL.java:439)\norg.apache.spark.sql.catalyst.expressions.ParseUrl.getUrl(stringExpressions.scala:731)\norg.apache.spark.sql.catalyst.expressions.ParseUrl.parseUrlWithoutKey(stringExpressions.scala:772)\norg.apache.spark.sql.catalyst.expressions.ParseUrl.eval(stringExpressions.scala:785)\norg.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificPredicate.eval(Unknown Source)\norg.apache.spark.sql.catalyst.expressions.codegen.GeneratePredicate$$anonfun$create$2.apply(GeneratePredicate.scala:69)\norg.apache.spark.sql.catalyst.expressions.codegen.GeneratePredicate$$anonfun$create$2.apply(GeneratePredicate.scala:69)\norg.apache.spark.sql.execution.FilterExec$$anonfun$17$$anonfun$apply$2.apply(basicPhysicalOperators.scala:203)\norg.apache.spark.sql.execution.FilterExec$$anonfun$17$$anonfun$apply$2.apply(basicPhysicalOperators.scala:202)\nscala.collection.Iterator$$anon$13.hasNext(Iterator.scala:463)\norg.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\norg.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\norg.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\nscala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\norg.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java:147)\norg.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\norg.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\norg.apache.spark.scheduler.Task.run(Task.scala:85)\norg.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\njava.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\njava.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\njava.lang.Thread.run(Thread.java:745)\n{code}\n\nHowever, when I switch from 1 executor with 36 cores to 9 executors with 4 cores, throughput is almost 10x higher and the CPUs are back at ~100% use.\n\nThanks!","issue_id":"12993828","key":"SPARK-16826","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-08-05T19:56:08.000+0000","role":"fixed_distractor","summary":"java.util.Hashtable limits the throughput of PARSE_URL()"} {"case_id":"12994868","cluster":"DISTRACTOR-SPARK-16896","comments":[{"body":"[~hyukjin.kwon] cc","created":"2016-08-04T11:43:14.910+0000"},{"body":"I agree with appending a number to the deplicated column names. Could you please take a look of this JIRA? [~rxin] and [~falaki]\n","created":"2016-08-05T00:57:39.021+0000"},{"body":"I suggest we generally follow the restrictions of SparkSQL, which does not accept duplicate column names. So I agree with numbering column names if there are duplicates.","created":"2016-08-05T01:08:45.889+0000"},{"body":"I wonder if anyone is tackling this issue. I want to get started contributing to Apache Spark and this sound like quite an interesting intro task. ","created":"2016-08-09T06:17:32.891+0000"},{"body":"I don't mind if you go ahead (I was looking at this problem though).\n\nOne thing I want to say is, we might better match the behaviour to [read.csv|https://stat.ethz.ch/R-manual/R-devel/library/utils/html/read.table.html] in R if possible in this case.\n\nIn addition, we are handling {{nullValue}} in handling the header with making numbers already. I guess we should clarify and write the behaviour in the PR description including the cases in R.\n\nAlso, do not forget to follow https://cwiki.apache.org/confluence/display/SPARK/Contributing+to+Spark for making a contribution.\n\n","created":"2016-08-09T06:42:18.537+0000"},{"body":"Quick confirmation needed : this change needs to happen in the https://github.com/databricks/spark-csv package if i understand correctly . \n\nWill start working on it later today, will come back with some more newbie questions.","created":"2016-08-09T08:43:32.449+0000"},{"body":"No, the code is in Spark now. ","created":"2016-08-09T08:49:53.317+0000"},{"body":"[~srowen] Oh, I am sorry, I coundn't see the comment above. Please ignore the comment I just left. I thought you left the comment for mine.","created":"2016-08-09T09:26:10.346+0000"},{"body":"Cool. So i am assuming that the entry point for the loading csv functionality that needs to be fixed is the one described in : https://github.com/apache/spark/blob/master/sql/core/src/main/scala/org/apache/spark/sql/DataFrameReader.scala .","created":"2016-08-09T09:50:12.503+0000"},{"body":"[~nlauchande] Just FYI, actual codes that need to be corrected will be around [here|https://github.com/apache/spark/blob/cb1b9d34f37a5574de43f61e7036c4b8b81defbf/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/csv/CSVFileFormat.scala#L61-L67].\n\nMaybe we should check the duplication and then give some numbers. I haven't checked the behaviour in R though.\n\nAlso, please make sure for a test. I believe we usually need a test for a patch.\n","created":"2016-08-09T09:56:30.292+0000"},{"body":"Hi [~nlauchande], are you currently working on this?\n","created":"2016-08-18T10:57:20.016+0000"},{"body":"Hi Hyukjin Kwon i did start. But couldn't make a lot of progress within the last couple of days . Feel free to grab it . I can try find another easier and less critical beginner task .","created":"2016-08-18T11:01:33.557+0000"},{"body":"Yup, then I will work on this and submit a PR within few days. Thank you!","created":"2016-08-22T01:44:19.936+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14745","created":"2016-08-22T04:50:07.284+0000"}],"conversations":[{"body":"It would be great if the library allows us to load csv with duplicate column names. I understand that having duplicate columns in the data is odd but sometimes we get data that has duplicate columns. Getting upstream data like that can happen. We may choose to ignore them but currently there is no way to drop those as we are not able to load them at all. Currently as a pre-processing I loaded the data into R, changed the column names and then make a fixed version with which Spark Java API can work.\n\nBut if talk about other options, e.g. R has read.csv which automatically takes care of such situation by appending a number to the column name.\n\nAlso case sensitivity in column names can also cause problems. I mean if we have columns like\n\nColumnName, columnName\n\nI may want to have them as separate. But the option to do this is not documented.","from":"reporter","subject":"Loading csv with duplicate column names"},{"body":"[~hyukjin.kwon] cc","from":"developer"},{"body":"I agree with appending a number to the deplicated column names. Could you please take a look of this JIRA? [~rxin] and [~falaki]\n","from":"developer"},{"body":"I suggest we generally follow the restrictions of SparkSQL, which does not accept duplicate column names. So I agree with numbering column names if there are duplicates.","from":"developer"},{"body":"I wonder if anyone is tackling this issue. I want to get started contributing to Apache Spark and this sound like quite an interesting intro task. ","from":"developer"},{"body":"I don't mind if you go ahead (I was looking at this problem though).\n\nOne thing I want to say is, we might better match the behaviour to [read.csv|https://stat.ethz.ch/R-manual/R-devel/library/utils/html/read.table.html] in R if possible in this case.\n\nIn addition, we are handling {{nullValue}} in handling the header with making numbers already. I guess we should clarify and write the behaviour in the PR description including the cases in R.\n\nAlso, do not forget to follow https://cwiki.apache.org/confluence/display/SPARK/Contributing+to+Spark for making a contribution.\n\n","from":"developer"},{"body":"Quick confirmation needed : this change needs to happen in the https://github.com/databricks/spark-csv package if i understand correctly . \n\nWill start working on it later today, will come back with some more newbie questions.","from":"developer"},{"body":"No, the code is in Spark now. ","from":"developer"},{"body":"[~srowen] Oh, I am sorry, I coundn't see the comment above. Please ignore the comment I just left. I thought you left the comment for mine.","from":"developer"},{"body":"Cool. So i am assuming that the entry point for the loading csv functionality that needs to be fixed is the one described in : https://github.com/apache/spark/blob/master/sql/core/src/main/scala/org/apache/spark/sql/DataFrameReader.scala .","from":"developer"},{"body":"[~nlauchande] Just FYI, actual codes that need to be corrected will be around [here|https://github.com/apache/spark/blob/cb1b9d34f37a5574de43f61e7036c4b8b81defbf/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/csv/CSVFileFormat.scala#L61-L67].\n\nMaybe we should check the duplication and then give some numbers. I haven't checked the behaviour in R though.\n\nAlso, please make sure for a test. I believe we usually need a test for a patch.\n","from":"developer"},{"body":"Hi [~nlauchande], are you currently working on this?\n","from":"developer"},{"body":"Hi Hyukjin Kwon i did start. But couldn't make a lot of progress within the last couple of days . Feel free to grab it . I can try find another easier and less critical beginner task .","from":"developer"},{"body":"Yup, then I will work on this and submit a PR within few days. Thank you!","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14745","from":"developer"}],"created":"2016-08-04T11:42:55.000+0000","description":"It would be great if the library allows us to load csv with duplicate column names. I understand that having duplicate columns in the data is odd but sometimes we get data that has duplicate columns. Getting upstream data like that can happen. We may choose to ignore them but currently there is no way to drop those as we are not able to load them at all. Currently as a pre-processing I loaded the data into R, changed the column names and then make a fixed version with which Spark Java API can work.\n\nBut if talk about other options, e.g. R has read.csv which automatically takes care of such situation by appending a number to the column name.\n\nAlso case sensitivity in column names can also cause problems. I mean if we have columns like\n\nColumnName, columnName\n\nI may want to have them as separate. But the option to do this is not documented.","issue_id":"12994868","key":"SPARK-16896","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-10-11T02:22:14.000+0000","role":"fixed_distractor","summary":"Loading csv with duplicate column names"} {"case_id":"12995731","cluster":"DISTRACTOR-SPARK-16955","comments":[{"body":"[~dongjoon] Will have time to take a look?","created":"2016-08-08T18:53:20.134+0000"},{"body":"Sure! Thank you, [~yhuai]. I'll take a look this.","created":"2016-08-08T18:54:51.425+0000"},{"body":"`ResolveAggregateFunctions` seems to have a bug to drop the ordinals. I'll make a PR after some testing.","created":"2016-08-08T21:07:31.517+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14546","created":"2016-08-08T21:58:05.859+0000"},{"body":"Hi, [~yhuai].\nCould you review the PR?\nThe root cause was `ResolveAggregateFunctions` removed the ordinal sort orders too early.\nAfter improving the `if` condition to check the resolution is completed, the case works well.","created":"2016-08-09T04:10:15.452+0000"},{"body":"User 'clockfly' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14616","created":"2016-08-12T03:06:07.088+0000"},{"body":"this bug is already fixed by https://github.com/apache/spark/pull/14595 by accident.","created":"2016-08-12T13:23:35.979+0000"}],"conversations":[{"body":"The following queries work\n{code}\nselect a from (select 1 as a) tmp order by 1\nselect a, count(*) from (select 1 as a) tmp group by 1\nselect a, count(*) from (select 1 as a) tmp group by 1 order by a\n{code}\n\nHowever, the following query does not\n{code}\nselect a, count(*) from (select 1 as a) tmp group by 1 order by 1\n{code}\n\n{code}\norg.apache.spark.sql.catalyst.analysis.UnresolvedException: Invalid call to Group by position: '1' exceeds the size of the select list '0'. on unresolved object, tree:\nAggregate [1]\n+- SubqueryAlias tmp\n +- Project [1 AS a#82]\n +- OneRowRelation$\n\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$$anonfun$apply$11$$anonfun$34.apply(Analyzer.scala:749)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$$anonfun$apply$11$$anonfun$34.apply(Analyzer.scala:739)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:244)\n\tat scala.collection.AbstractTraversable.map(Traversable.scala:105)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$$anonfun$apply$11.applyOrElse(Analyzer.scala:739)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$$anonfun$apply$11.applyOrElse(Analyzer.scala:715)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolveOperators(LogicalPlan.scala:60)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$.apply(Analyzer.scala:715)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$.apply(Analyzer.scala:714)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:85)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:82)\n\tat scala.collection.LinearSeqOptimized$class.foldLeft(LinearSeqOptimized.scala:111)\n\tat scala.collection.immutable.List.foldLeft(List.scala:84)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:82)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:74)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:74)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveAggregateFunctions$$anonfun$apply$20.applyOrElse(Analyzer.scala:1237)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveAggregateFunctions$$anonfun$apply$20.applyOrElse(Analyzer.scala:1182)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolveOperators(LogicalPlan.scala:60)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveAggregateFunctions$.apply(Analyzer.scala:1182)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveAggregateFunctions$.apply(Analyzer.scala:1181)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:85)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:82)\n\tat scala.collection.LinearSeqOptimized$class.foldLeft(LinearSeqOptimized.scala:111)\n\tat scala.collection.immutable.List.foldLeft(List.scala:84)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:82)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:74)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:74)\n\tat org.apache.spark.sql.execution.QueryExecution.analyzed$lzycompute(QueryExecution.scala:65)\n\tat org.apache.spark.sql.execution.QueryExecution.analyzed(QueryExecution.scala:63)\n\tat org.apache.spark.sql.execution.QueryExecution.assertAnalyzed(QueryExecution.scala:49)\n\tat org.apache.spark.sql.Dataset$.ofRows(Dataset.scala:64)\n\tat org.apache.spark.sql.SparkSession.sql(SparkSession.scala:582)\n\tat org.apache.spark.sql.SQLContext.sql(SQLContext.scala:682)\n{code}","from":"reporter","subject":"Using ordinals in ORDER BY causes an analysis error when the query has a GROUP BY clause using ordinals"},{"body":"[~dongjoon] Will have time to take a look?","from":"developer"},{"body":"Sure! Thank you, [~yhuai]. I'll take a look this.","from":"developer"},{"body":"`ResolveAggregateFunctions` seems to have a bug to drop the ordinals. I'll make a PR after some testing.","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14546","from":"developer"},{"body":"Hi, [~yhuai].\nCould you review the PR?\nThe root cause was `ResolveAggregateFunctions` removed the ordinal sort orders too early.\nAfter improving the `if` condition to check the resolution is completed, the case works well.","from":"developer"},{"body":"User 'clockfly' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14616","from":"developer"},{"body":"this bug is already fixed by https://github.com/apache/spark/pull/14595 by accident.","from":"developer"}],"created":"2016-08-08T18:50:53.000+0000","description":"The following queries work\n{code}\nselect a from (select 1 as a) tmp order by 1\nselect a, count(*) from (select 1 as a) tmp group by 1\nselect a, count(*) from (select 1 as a) tmp group by 1 order by a\n{code}\n\nHowever, the following query does not\n{code}\nselect a, count(*) from (select 1 as a) tmp group by 1 order by 1\n{code}\n\n{code}\norg.apache.spark.sql.catalyst.analysis.UnresolvedException: Invalid call to Group by position: '1' exceeds the size of the select list '0'. on unresolved object, tree:\nAggregate [1]\n+- SubqueryAlias tmp\n +- Project [1 AS a#82]\n +- OneRowRelation$\n\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$$anonfun$apply$11$$anonfun$34.apply(Analyzer.scala:749)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$$anonfun$apply$11$$anonfun$34.apply(Analyzer.scala:739)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:244)\n\tat scala.collection.AbstractTraversable.map(Traversable.scala:105)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$$anonfun$apply$11.applyOrElse(Analyzer.scala:739)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$$anonfun$apply$11.applyOrElse(Analyzer.scala:715)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolveOperators(LogicalPlan.scala:60)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$.apply(Analyzer.scala:715)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveOrdinalInOrderByAndGroupBy$.apply(Analyzer.scala:714)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:85)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:82)\n\tat scala.collection.LinearSeqOptimized$class.foldLeft(LinearSeqOptimized.scala:111)\n\tat scala.collection.immutable.List.foldLeft(List.scala:84)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:82)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:74)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:74)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveAggregateFunctions$$anonfun$apply$20.applyOrElse(Analyzer.scala:1237)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveAggregateFunctions$$anonfun$apply$20.applyOrElse(Analyzer.scala:1182)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolveOperators(LogicalPlan.scala:60)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveAggregateFunctions$.apply(Analyzer.scala:1182)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveAggregateFunctions$.apply(Analyzer.scala:1181)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:85)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:82)\n\tat scala.collection.LinearSeqOptimized$class.foldLeft(LinearSeqOptimized.scala:111)\n\tat scala.collection.immutable.List.foldLeft(List.scala:84)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:82)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:74)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:74)\n\tat org.apache.spark.sql.execution.QueryExecution.analyzed$lzycompute(QueryExecution.scala:65)\n\tat org.apache.spark.sql.execution.QueryExecution.analyzed(QueryExecution.scala:63)\n\tat org.apache.spark.sql.execution.QueryExecution.assertAnalyzed(QueryExecution.scala:49)\n\tat org.apache.spark.sql.Dataset$.ofRows(Dataset.scala:64)\n\tat org.apache.spark.sql.SparkSession.sql(SparkSession.scala:582)\n\tat org.apache.spark.sql.SQLContext.sql(SQLContext.scala:682)\n{code}","issue_id":"12995731","key":"SPARK-16955","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-08-12T13:24:36.000+0000","role":"fixed_distractor","summary":"Using ordinals in ORDER BY causes an analysis error when the query has a GROUP BY clause using ordinals"} {"case_id":"12995920","cluster":"DISTRACTOR-SPARK-16975","comments":[{"body":"cc [~dongjoon] do you have time to look into this?\n","created":"2016-08-10T19:18:12.254+0000"},{"body":"Oh, sure. It's my pleasure. I'll take a look.","created":"2016-08-10T19:20:37.771+0000"},{"body":"Thank you for pinging me.","created":"2016-08-10T19:21:51.089+0000"},{"body":"Hi, [~immerrr].\nI can not reproduce your situation, but could you change `_locality_code` into `locality_code`?\nSpark 2.0 ignores the path names starting with underscore or dot; `_` or `.`","created":"2016-08-10T19:56:25.892+0000"},{"body":"I made a sample case having similar behaviors. I think this is related closed. [~rxin], how do you think about this?\n\n{code}\nspark-1.6.2-bin-hadoop2.6$ ls /tmp/parquet16/\n_SUCCESS _locality_code=1 _locality_code=3 _locality_code=5 _locality_code=7 _locality_code=9\n_locality_code=0 _locality_code=2 _locality_code=4 _locality_code=6 _locality_code=8\n{code}\n\n{code}\nscala> spark.read.parquet(\"/tmp/parquet16\").show\norg.apache.spark.sql.AnalysisException: Unable to infer schema for ParquetFormat at /tmp/parquet16. It must be specified manually;\nscala> spark.read.parquet(\"/tmp/parquet16/_locality_code=0\").show\n+---+\n| id|\n+---+\n| 0|\n+---+\n{code}","created":"2016-08-10T20:11:49.082+0000"},{"body":"Ah, [~rxin]. \nFor this issue, we should add a migration document for 1.6 .\n\nSpark 2.0 itself has the save problem. We should block the illegal column names. May I make a PR for this?\n\n{code}\nscala> spark.range(10).withColumn(\"_locality_code\", $\"id\").write.partitionBy(\"_locality_code\").save(\"/tmp/parquet20\")\n\nscala> spark.read.parquet(\"/tmp/parquet20\")\norg.apache.spark.sql.AnalysisException: Unable to infer schema for ParquetFormat at /tmp/parquet20. It must be specified manually;\n\nscala> spark.range(10).withColumn(\"locality_code\", $\"id\").write.partitionBy(\"locality_code\").save(\"/tmp/parquet201\")\n\nscala> spark.read.parquet(\"/tmp/parquet201\")\nres6: org.apache.spark.sql.DataFrame = [id: bigint, locality_code: int]\n{code}","created":"2016-08-10T20:17:00.742+0000"},{"body":"Let me dig more. I can find more general solution for this for Spark 1.6 / 2.0.","created":"2016-08-10T20:36:18.169+0000"},{"body":"oh, that's unfortunate. coming from python world, underscore seems a natural prefix for \"internal things\".\n\nwhat bugs me, though, is that spark2.0 had no problems reading up to 31 directories starting with underscores and bugged out only when there were 32 of them.\n\nand i'll try the rename, give me a sec..","created":"2016-08-10T20:38:27.609+0000"},{"body":"In the python, could you give the command string to write parquet?\n\nBTW, I also started to work in order to support `_col=xxx` format.","created":"2016-08-10T20:41:14.317+0000"},{"body":"You mean the one I use to write the data? I don't have the exact string at hand, but it was a straightforward conversion from JSON inferring schema on the way, smth like\n{code}\nsqlContext.read.json('/path/to/json-data').write.partitionBy('_locality_code').parquet('/path/to/parquet-data', mode='overwrite')\n{code}\n\n\nOh, another thing was that when I tried reading subdir-by-subdir. When you read a single subdirectory, the _locality_code column is not present (just like in your example), but for some reason it worked ok when reading just one such subdirectory but failed reading that same directory multiple times. There were other columns starting with underscore, though. I didn't use them to partition the data, but maybe they still somehow affected schema inference.","created":"2016-08-10T20:57:18.978+0000"},{"body":"Okay. Wait a second, please. I'll make PR for you.","created":"2016-08-10T20:58:38.866+0000"},{"body":"I mean, this line from the original report\n{code}\nIn [87]: spark.read.parquet(*([subdirs[0]] * 32))\n{code}\n\nmeans pass subdirs[0] 32 times as parameters to spark.read.parquet, i.e. spark.read.parquet(subdirs[0], subdirs[0], ...)","created":"2016-08-10T20:59:39.978+0000"},{"body":"Yep. And, it raised exceptions, right?","created":"2016-08-10T21:01:58.241+0000"},{"body":"Yes, a seemingly similar one.","created":"2016-08-10T21:03:44.392+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14585","created":"2016-08-10T21:11:05.205+0000"},{"body":"I made a PR, [~immerrr]. I tested only `sql` module tests. After passing Jenkins, could you test the PR in your environment if you have some time?\nI think PySpark also will get benefit of this PR.","created":"2016-08-10T21:15:51.634+0000"},{"body":"Great, thank you! That was so fast!\n\nI'll try to look at it tomorrow, but cannot promise anything as my experience in Scala and building it is next to none.","created":"2016-08-10T21:18:44.625+0000"},{"body":"Thank you. See you tomorrow! :)","created":"2016-08-10T21:20:59.947+0000"},{"body":"I have built the code from the PR and it indeed succeeds reading the data.\n\nI have tried doing {{df.count()}} and now I'm swarmed with warnings like this (they are just keep getting printed endlessly in the terminal): \n{code}\n16/08/11 12:18:51 WARN CorruptStatistics: Ignoring statistics because created_by could not be parsed (see PARQUET-251): parquet-mr version 1.6.0\norg.apache.parquet.VersionParser$VersionParseException: Could not parse created_by: parquet-mr version 1.6.0 using format: (.+) version ((.*) )?\\(build ?(.*)\\)\n\tat org.apache.parquet.VersionParser.parse(VersionParser.java:112)\n\tat org.apache.parquet.CorruptStatistics.shouldIgnoreStatistics(CorruptStatistics.java:60)\n\tat org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetStatistics(ParquetMetadataConverter.java:263)\n\tat org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetMetadata(ParquetMetadataConverter.java:567)\n\tat org.apache.parquet.format.converter.ParquetMetadataConverter.readParquetMetadata(ParquetMetadataConverter.java:544)\n\tat org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:431)\n\tat org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:386)\n\tat org.apache.spark.sql.execution.datasources.parquet.SpecificParquetRecordReaderBase.initialize(SpecificParquetRecordReaderBase.java:107)\n\tat org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.initialize(VectorizedParquetRecordReader.java:109)\n\tat org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anonfun$buildReader$1.apply(ParquetFileFormat.scala:369)\n\tat org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anonfun$buildReader$1.apply(ParquetFileFormat.scala:343)\n\tat org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:122)\n\tat org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:97)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.scan_nextBatch$(Unknown Source)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.agg_doAggregateWithoutKey$(Unknown Source)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n\tat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\n\tat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n\tat org.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java:126)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:86)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{code} ","created":"2016-08-11T12:19:34.850+0000"},{"body":"But it works. I have suppressed WARN logs and {{df.count()}} returned the correct value, despite taking 2x time to finish compared to 1.6.2.","created":"2016-08-11T12:57:48.510+0000"},{"body":"The figures were:\n1.6.2: ~6s\n2.0.0: ~12s\n\nInterestingly enough, after restarting the driver, 2.0 run took far less than that:\n1.6.2: ~5.7s\n2.0.0: ~1.4s\n\nMaybe, some sort of caching that survives restarts is used internally.","created":"2016-08-11T13:03:38.089+0000"},{"body":"Great! Thank you for confirming.","created":"2016-08-11T13:33:12.349+0000"},{"body":"Hi, [~rxin].\nCould you review this PR?","created":"2016-08-12T04:51:08.034+0000"},{"body":"Issue resolved by pull request 14585\n[https://github.com/apache/spark/pull/14585]","created":"2016-08-12T07:09:52.540+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14627","created":"2016-08-13T01:19:03.984+0000"}],"conversations":[{"body":"Spark-2.0.0 seems to have some problems reading a parquet dataset generated by 1.6.2. \n\n{code}\nIn [80]: spark.read.parquet('/path/to/data')\n...\nAnalysisException: u'Unable to infer schema for ParquetFormat at /path/to/data. It must be specified manually;'\n{code}\n\nThe dataset is ~150G and partitioned by _locality_code column. None of the partitions are empty. I have narrowed the failing dataset to the first 32 partitions of the data:\n\n{code}\nIn [82]: spark.read.parquet(*subdirs[:32])\n...\nAnalysisException: u'Unable to infer schema for ParquetFormat at /path/to/data/_locality_code=AQ,/path/to/data/_locality_code=AI. It must be specified manually;'\n{code}\n\nInterestingly, it works OK if you remove any of the partitions from the list:\n{code}\nIn [83]: for i in range(32): spark.read.parquet(*(subdirs[:i] + subdirs[i+1:32]))\n{code}\n\nAnother strange thing is that the schemas for the first and the last 31 partitions of the subset are identical:\n{code}\nIn [84]: spark.read.parquet(*subdirs[:31]).schema.fields == spark.read.parquet(*subdirs[1:32]).schema.fields\nOut[84]: True\n{code}\n\nWhich got me interested and I tried this:\n{code}\nIn [87]: spark.read.parquet(*([subdirs[0]] * 32))\n...\nAnalysisException: u'Unable to infer schema for ParquetFormat at /path/to/data/_locality_code=AQ,/path/to/data/_locality_code=AQ. It must be specified manually;'\n\nIn [88]: spark.read.parquet(*([subdirs[15]] * 32))\n...\nAnalysisException: u'Unable to infer schema for ParquetFormat at /path/to/data/_locality_code=AX,/path/to/data/_locality_code=AX. It must be specified manually;'\n\nIn [89]: spark.read.parquet(*([subdirs[31]] * 32))\n...\nAnalysisException: u'Unable to infer schema for ParquetFormat at /path/to/data/_locality_code=BE,/path/to/data/_locality_code=BE. It must be specified manually;'\n{code}\n\nIf I read the first partition, save it in 2.0 and try to read in the same manner, everything is fine:\n{code}\nIn [100]: spark.read.parquet(subdirs[0]).write.parquet('spark-2.0-test')\n16/08/09 11:03:37 WARN ParquetRecordReader: Can not initialize counter due to context is not a instance of TaskInputOutputContext, but is org.apache.hadoop.mapreduce.task.TaskAttemptContextImpl\n\nIn [101]: df = spark.read.parquet(*(['spark-2.0-test'] * 32))\n{code}\n\nI have originally posted it to user mailing list, but with the last discoveries this clearly seems like a bug.","from":"reporter","subject":"Spark-2.0.0 unable to infer schema for parquet data written by Spark-1.6.2"},{"body":"cc [~dongjoon] do you have time to look into this?\n","from":"developer"},{"body":"Oh, sure. It's my pleasure. I'll take a look.","from":"developer"},{"body":"Thank you for pinging me.","from":"developer"},{"body":"Hi, [~immerrr].\nI can not reproduce your situation, but could you change `_locality_code` into `locality_code`?\nSpark 2.0 ignores the path names starting with underscore or dot; `_` or `.`","from":"developer"},{"body":"I made a sample case having similar behaviors. I think this is related closed. [~rxin], how do you think about this?\n\n{code}\nspark-1.6.2-bin-hadoop2.6$ ls /tmp/parquet16/\n_SUCCESS _locality_code=1 _locality_code=3 _locality_code=5 _locality_code=7 _locality_code=9\n_locality_code=0 _locality_code=2 _locality_code=4 _locality_code=6 _locality_code=8\n{code}\n\n{code}\nscala> spark.read.parquet(\"/tmp/parquet16\").show\norg.apache.spark.sql.AnalysisException: Unable to infer schema for ParquetFormat at /tmp/parquet16. It must be specified manually;\nscala> spark.read.parquet(\"/tmp/parquet16/_locality_code=0\").show\n+---+\n| id|\n+---+\n| 0|\n+---+\n{code}","from":"developer"},{"body":"Ah, [~rxin]. \nFor this issue, we should add a migration document for 1.6 .\n\nSpark 2.0 itself has the save problem. We should block the illegal column names. May I make a PR for this?\n\n{code}\nscala> spark.range(10).withColumn(\"_locality_code\", $\"id\").write.partitionBy(\"_locality_code\").save(\"/tmp/parquet20\")\n\nscala> spark.read.parquet(\"/tmp/parquet20\")\norg.apache.spark.sql.AnalysisException: Unable to infer schema for ParquetFormat at /tmp/parquet20. It must be specified manually;\n\nscala> spark.range(10).withColumn(\"locality_code\", $\"id\").write.partitionBy(\"locality_code\").save(\"/tmp/parquet201\")\n\nscala> spark.read.parquet(\"/tmp/parquet201\")\nres6: org.apache.spark.sql.DataFrame = [id: bigint, locality_code: int]\n{code}","from":"developer"},{"body":"Let me dig more. I can find more general solution for this for Spark 1.6 / 2.0.","from":"developer"},{"body":"oh, that's unfortunate. coming from python world, underscore seems a natural prefix for \"internal things\".\n\nwhat bugs me, though, is that spark2.0 had no problems reading up to 31 directories starting with underscores and bugged out only when there were 32 of them.\n\nand i'll try the rename, give me a sec..","from":"developer"},{"body":"In the python, could you give the command string to write parquet?\n\nBTW, I also started to work in order to support `_col=xxx` format.","from":"developer"},{"body":"You mean the one I use to write the data? I don't have the exact string at hand, but it was a straightforward conversion from JSON inferring schema on the way, smth like\n{code}\nsqlContext.read.json('/path/to/json-data').write.partitionBy('_locality_code').parquet('/path/to/parquet-data', mode='overwrite')\n{code}\n\n\nOh, another thing was that when I tried reading subdir-by-subdir. When you read a single subdirectory, the _locality_code column is not present (just like in your example), but for some reason it worked ok when reading just one such subdirectory but failed reading that same directory multiple times. There were other columns starting with underscore, though. I didn't use them to partition the data, but maybe they still somehow affected schema inference.","from":"developer"},{"body":"Okay. Wait a second, please. I'll make PR for you.","from":"developer"},{"body":"I mean, this line from the original report\n{code}\nIn [87]: spark.read.parquet(*([subdirs[0]] * 32))\n{code}\n\nmeans pass subdirs[0] 32 times as parameters to spark.read.parquet, i.e. spark.read.parquet(subdirs[0], subdirs[0], ...)","from":"developer"},{"body":"Yep. And, it raised exceptions, right?","from":"developer"},{"body":"Yes, a seemingly similar one.","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14585","from":"developer"},{"body":"I made a PR, [~immerrr]. I tested only `sql` module tests. After passing Jenkins, could you test the PR in your environment if you have some time?\nI think PySpark also will get benefit of this PR.","from":"developer"},{"body":"Great, thank you! That was so fast!\n\nI'll try to look at it tomorrow, but cannot promise anything as my experience in Scala and building it is next to none.","from":"developer"},{"body":"Thank you. See you tomorrow! :)","from":"developer"},{"body":"I have built the code from the PR and it indeed succeeds reading the data.\n\nI have tried doing {{df.count()}} and now I'm swarmed with warnings like this (they are just keep getting printed endlessly in the terminal): \n{code}\n16/08/11 12:18:51 WARN CorruptStatistics: Ignoring statistics because created_by could not be parsed (see PARQUET-251): parquet-mr version 1.6.0\norg.apache.parquet.VersionParser$VersionParseException: Could not parse created_by: parquet-mr version 1.6.0 using format: (.+) version ((.*) )?\\(build ?(.*)\\)\n\tat org.apache.parquet.VersionParser.parse(VersionParser.java:112)\n\tat org.apache.parquet.CorruptStatistics.shouldIgnoreStatistics(CorruptStatistics.java:60)\n\tat org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetStatistics(ParquetMetadataConverter.java:263)\n\tat org.apache.parquet.format.converter.ParquetMetadataConverter.fromParquetMetadata(ParquetMetadataConverter.java:567)\n\tat org.apache.parquet.format.converter.ParquetMetadataConverter.readParquetMetadata(ParquetMetadataConverter.java:544)\n\tat org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:431)\n\tat org.apache.parquet.hadoop.ParquetFileReader.readFooter(ParquetFileReader.java:386)\n\tat org.apache.spark.sql.execution.datasources.parquet.SpecificParquetRecordReaderBase.initialize(SpecificParquetRecordReaderBase.java:107)\n\tat org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.initialize(VectorizedParquetRecordReader.java:109)\n\tat org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anonfun$buildReader$1.apply(ParquetFileFormat.scala:369)\n\tat org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anonfun$buildReader$1.apply(ParquetFileFormat.scala:343)\n\tat org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:122)\n\tat org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:97)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.scan_nextBatch$(Unknown Source)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.agg_doAggregateWithoutKey$(Unknown Source)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n\tat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\n\tat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n\tat org.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java:126)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:86)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{code} ","from":"developer"},{"body":"But it works. I have suppressed WARN logs and {{df.count()}} returned the correct value, despite taking 2x time to finish compared to 1.6.2.","from":"developer"},{"body":"The figures were:\n1.6.2: ~6s\n2.0.0: ~12s\n\nInterestingly enough, after restarting the driver, 2.0 run took far less than that:\n1.6.2: ~5.7s\n2.0.0: ~1.4s\n\nMaybe, some sort of caching that survives restarts is used internally.","from":"developer"},{"body":"Great! Thank you for confirming.","from":"developer"},{"body":"Hi, [~rxin].\nCould you review this PR?","from":"developer"},{"body":"Issue resolved by pull request 14585\n[https://github.com/apache/spark/pull/14585]","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14627","from":"developer"}],"created":"2016-08-09T11:04:31.000+0000","description":"Spark-2.0.0 seems to have some problems reading a parquet dataset generated by 1.6.2. \n\n{code}\nIn [80]: spark.read.parquet('/path/to/data')\n...\nAnalysisException: u'Unable to infer schema for ParquetFormat at /path/to/data. It must be specified manually;'\n{code}\n\nThe dataset is ~150G and partitioned by _locality_code column. None of the partitions are empty. I have narrowed the failing dataset to the first 32 partitions of the data:\n\n{code}\nIn [82]: spark.read.parquet(*subdirs[:32])\n...\nAnalysisException: u'Unable to infer schema for ParquetFormat at /path/to/data/_locality_code=AQ,/path/to/data/_locality_code=AI. It must be specified manually;'\n{code}\n\nInterestingly, it works OK if you remove any of the partitions from the list:\n{code}\nIn [83]: for i in range(32): spark.read.parquet(*(subdirs[:i] + subdirs[i+1:32]))\n{code}\n\nAnother strange thing is that the schemas for the first and the last 31 partitions of the subset are identical:\n{code}\nIn [84]: spark.read.parquet(*subdirs[:31]).schema.fields == spark.read.parquet(*subdirs[1:32]).schema.fields\nOut[84]: True\n{code}\n\nWhich got me interested and I tried this:\n{code}\nIn [87]: spark.read.parquet(*([subdirs[0]] * 32))\n...\nAnalysisException: u'Unable to infer schema for ParquetFormat at /path/to/data/_locality_code=AQ,/path/to/data/_locality_code=AQ. It must be specified manually;'\n\nIn [88]: spark.read.parquet(*([subdirs[15]] * 32))\n...\nAnalysisException: u'Unable to infer schema for ParquetFormat at /path/to/data/_locality_code=AX,/path/to/data/_locality_code=AX. It must be specified manually;'\n\nIn [89]: spark.read.parquet(*([subdirs[31]] * 32))\n...\nAnalysisException: u'Unable to infer schema for ParquetFormat at /path/to/data/_locality_code=BE,/path/to/data/_locality_code=BE. It must be specified manually;'\n{code}\n\nIf I read the first partition, save it in 2.0 and try to read in the same manner, everything is fine:\n{code}\nIn [100]: spark.read.parquet(subdirs[0]).write.parquet('spark-2.0-test')\n16/08/09 11:03:37 WARN ParquetRecordReader: Can not initialize counter due to context is not a instance of TaskInputOutputContext, but is org.apache.hadoop.mapreduce.task.TaskAttemptContextImpl\n\nIn [101]: df = spark.read.parquet(*(['spark-2.0-test'] * 32))\n{code}\n\nI have originally posted it to user mailing list, but with the last discoveries this clearly seems like a bug.","issue_id":"12995920","key":"SPARK-16975","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-08-12T07:09:52.000+0000","role":"fixed_distractor","summary":"Spark-2.0.0 unable to infer schema for parquet data written by Spark-1.6.2"} {"case_id":"12997755","cluster":"DISTRACTOR-SPARK-17098","comments":[{"body":"Actually, given the error here I think that the problem could be that sometimes {{WindowExpression.foldable == true}} even though {{WindowExpression}} is {{Unevaluable}}:\n\n{code}\ncase class WindowExpression(\n windowFunction: Expression,\n windowSpec: WindowSpecDefinition) extends Expression with Unevaluable {\n\n override def children: Seq[Expression] = windowFunction :: windowSpec :: Nil\n\n override def dataType: DataType = windowFunction.dataType\n override def foldable: Boolean = windowFunction.foldable\n override def nullable: Boolean = windowFunction.nullable\n\n override def toString: String = s\"$windowFunction $windowSpec\"\n override def sql: String = windowFunction.sql + \" OVER \" + windowSpec.sql\n}\n{code}\n\n/cc [~hvanhovell], FYI. ","created":"2016-08-17T00:11:12.496+0000"},{"body":"Hi, [~joshrosen].\nAccording to your error message `cast(0 as bigint)`, it seems a bug in `NullPropagation` optimizer.\nI'll make a PR for this so that you can review that.","created":"2016-08-17T17:00:53.055+0000"},{"body":"Actually, `SELECT COUNT(1 + NULL) OVER ()` also should return 0, too.","created":"2016-08-17T17:21:37.779+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14689","created":"2016-08-17T17:39:05.342+0000"}],"conversations":[{"body":"Running\n\n{code}\nSELECT COUNT(NULL) OVER ()\n{code}\n\nthrows an UnsupportedOperationException during analysis:\n\n{code}\njava.lang.UnsupportedOperationException: Cannot evaluate expression: cast(0 as bigint) windowspecdefinition(ROWS BETWEEN UNBOUNDED PRECEDING AND UNBOUNDED FOLLOWING)\n\tat org.apache.spark.sql.catalyst.expressions.Unevaluable$class.eval(Expression.scala:221)\n\tat org.apache.spark.sql.catalyst.expressions.WindowExpression.eval(windowExpressions.scala:288)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$$anonfun$apply$18$$anonfun$applyOrElse$3.applyOrElse(Optimizer.scala:759)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$$anonfun$apply$18$$anonfun$applyOrElse$3.applyOrElse(Optimizer.scala:752)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:278)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpressionDown$1(QueryPlan.scala:156)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.org$apache$spark$sql$catalyst$plans$QueryPlan$$recursiveTransform$1(QueryPlan.scala:166)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$org$apache$spark$sql$catalyst$plans$QueryPlan$$recursiveTransform$1$1.apply(QueryPlan.scala:170)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n\tat scala.collection.AbstractTraversable.map(Traversable.scala:104)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.org$apache$spark$sql$catalyst$plans$QueryPlan$$recursiveTransform$1(QueryPlan.scala:170)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$4.apply(QueryPlan.scala:175)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpressionsDown(QueryPlan.scala:175)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$$anonfun$apply$18.applyOrElse(Optimizer.scala:752)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$$anonfun$apply$18.applyOrElse(Optimizer.scala:751)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:278)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:268)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$.apply(Optimizer.scala:751)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$.apply(Optimizer.scala:750)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:85)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:82)\n\tat scala.collection.IndexedSeqOptimized$class.foldl(IndexedSeqOptimized.scala:57)\n\tat scala.collection.IndexedSeqOptimized$class.foldLeft(IndexedSeqOptimized.scala:66)\n\tat scala.collection.mutable.WrappedArray.foldLeft(WrappedArray.scala:35)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:82)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:74)\n\tat scala.collection.immutable.List.foreach(List.scala:381)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:74)\n\tat org.apache.spark.sql.execution.QueryExecution.optimizedPlan$lzycompute(QueryExecution.scala:74)\n\tat org.apache.spark.sql.execution.QueryExecution.optimizedPlan(QueryExecution.scala:74)\n\tat org.apache.spark.sql.execution.QueryExecution.sparkPlan$lzycompute(QueryExecution.scala:78)\n\tat org.apache.spark.sql.execution.QueryExecution.sparkPlan(QueryExecution.scala:76)\n\tat org.apache.spark.sql.execution.QueryExecution.executedPlan$lzycompute(QueryExecution.scala:83)\n\tat org.apache.spark.sql.execution.QueryExecution.executedPlan(QueryExecution.scala:83)\n\tat org.apache.spark.sql.Dataset.withTypedCallback(Dataset.scala:2558)\n\tat org.apache.spark.sql.Dataset.head(Dataset.scala:1924)\n\tat org.apache.spark.sql.Dataset.take(Dataset.scala:2139)\n{code}\n\nGiven that \n\n{code}\nSELECT COUNT(0) OVER ()\n{code}\n\nworks fine my hunch is that this is uncovering a bug in the ordering of our optimizer rules or a bug in the constant-folding rule itself. This particular example is probably unimportant by itself but may be an indicator of other problems.","from":"reporter","subject":"\"SELECT COUNT(NULL) OVER ()\" throws UnsupportedOperationException during analysis"},{"body":"Actually, given the error here I think that the problem could be that sometimes {{WindowExpression.foldable == true}} even though {{WindowExpression}} is {{Unevaluable}}:\n\n{code}\ncase class WindowExpression(\n windowFunction: Expression,\n windowSpec: WindowSpecDefinition) extends Expression with Unevaluable {\n\n override def children: Seq[Expression] = windowFunction :: windowSpec :: Nil\n\n override def dataType: DataType = windowFunction.dataType\n override def foldable: Boolean = windowFunction.foldable\n override def nullable: Boolean = windowFunction.nullable\n\n override def toString: String = s\"$windowFunction $windowSpec\"\n override def sql: String = windowFunction.sql + \" OVER \" + windowSpec.sql\n}\n{code}\n\n/cc [~hvanhovell], FYI. ","from":"developer"},{"body":"Hi, [~joshrosen].\nAccording to your error message `cast(0 as bigint)`, it seems a bug in `NullPropagation` optimizer.\nI'll make a PR for this so that you can review that.","from":"developer"},{"body":"Actually, `SELECT COUNT(1 + NULL) OVER ()` also should return 0, too.","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14689","from":"developer"}],"created":"2016-08-17T00:07:27.000+0000","description":"Running\n\n{code}\nSELECT COUNT(NULL) OVER ()\n{code}\n\nthrows an UnsupportedOperationException during analysis:\n\n{code}\njava.lang.UnsupportedOperationException: Cannot evaluate expression: cast(0 as bigint) windowspecdefinition(ROWS BETWEEN UNBOUNDED PRECEDING AND UNBOUNDED FOLLOWING)\n\tat org.apache.spark.sql.catalyst.expressions.Unevaluable$class.eval(Expression.scala:221)\n\tat org.apache.spark.sql.catalyst.expressions.WindowExpression.eval(windowExpressions.scala:288)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$$anonfun$apply$18$$anonfun$applyOrElse$3.applyOrElse(Optimizer.scala:759)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$$anonfun$apply$18$$anonfun$applyOrElse$3.applyOrElse(Optimizer.scala:752)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:278)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpressionDown$1(QueryPlan.scala:156)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.org$apache$spark$sql$catalyst$plans$QueryPlan$$recursiveTransform$1(QueryPlan.scala:166)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$org$apache$spark$sql$catalyst$plans$QueryPlan$$recursiveTransform$1$1.apply(QueryPlan.scala:170)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n\tat scala.collection.AbstractTraversable.map(Traversable.scala:104)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.org$apache$spark$sql$catalyst$plans$QueryPlan$$recursiveTransform$1(QueryPlan.scala:170)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$4.apply(QueryPlan.scala:175)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpressionsDown(QueryPlan.scala:175)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$$anonfun$apply$18.applyOrElse(Optimizer.scala:752)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$$anonfun$apply$18.applyOrElse(Optimizer.scala:751)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:278)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:268)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$.apply(Optimizer.scala:751)\n\tat org.apache.spark.sql.catalyst.optimizer.ConstantFolding$.apply(Optimizer.scala:750)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:85)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:82)\n\tat scala.collection.IndexedSeqOptimized$class.foldl(IndexedSeqOptimized.scala:57)\n\tat scala.collection.IndexedSeqOptimized$class.foldLeft(IndexedSeqOptimized.scala:66)\n\tat scala.collection.mutable.WrappedArray.foldLeft(WrappedArray.scala:35)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:82)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:74)\n\tat scala.collection.immutable.List.foreach(List.scala:381)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:74)\n\tat org.apache.spark.sql.execution.QueryExecution.optimizedPlan$lzycompute(QueryExecution.scala:74)\n\tat org.apache.spark.sql.execution.QueryExecution.optimizedPlan(QueryExecution.scala:74)\n\tat org.apache.spark.sql.execution.QueryExecution.sparkPlan$lzycompute(QueryExecution.scala:78)\n\tat org.apache.spark.sql.execution.QueryExecution.sparkPlan(QueryExecution.scala:76)\n\tat org.apache.spark.sql.execution.QueryExecution.executedPlan$lzycompute(QueryExecution.scala:83)\n\tat org.apache.spark.sql.execution.QueryExecution.executedPlan(QueryExecution.scala:83)\n\tat org.apache.spark.sql.Dataset.withTypedCallback(Dataset.scala:2558)\n\tat org.apache.spark.sql.Dataset.head(Dataset.scala:1924)\n\tat org.apache.spark.sql.Dataset.take(Dataset.scala:2139)\n{code}\n\nGiven that \n\n{code}\nSELECT COUNT(0) OVER ()\n{code}\n\nworks fine my hunch is that this is uncovering a bug in the ordering of our optimizer rules or a bug in the constant-folding rule itself. This particular example is probably unimportant by itself but may be an indicator of other problems.","issue_id":"12997755","key":"SPARK-17098","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-08-21T20:09:31.000+0000","role":"fixed_distractor","summary":"\"SELECT COUNT(NULL) OVER ()\" throws UnsupportedOperationException during analysis"} {"case_id":"12999534","cluster":"DISTRACTOR-SPARK-17210","comments":[{"body":"User 'zjffdu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14784","created":"2016-08-24T07:44:05.115+0000"},{"body":"cc [~rxin] - this is merged to master and branch-2.0. If we are rolling another RC for Spark 2.0.1 will be good to include this fix. Thanks","created":"2016-09-23T18:41:07.229+0000"},{"body":"[~felixcheung] please set the version to 2.0.2 from now on, since rc was already cut. If there is a new RC, I will change all 2.0.2 tickets to 2.0.1 when I cut the next one.\n","created":"2016-09-23T18:45:47.700+0000"},{"body":"Got it, sorry about that, I should have noticed.","created":"2016-09-24T06:12:59.098+0000"}],"conversations":[{"body":"Here's the code to reproduce this issue. \n{code}\nSys.setenv(SPARK_HOME=\"/Users/jzhang/github/spark\")\n.libPaths(c(file.path(Sys.getenv(), \"R\", \"lib\"), .libPaths()))\nlibrary(SparkR)\nsparkR.session(master=\"yarn-client\", sparkConfig = list(spark.executor.instances=\"1\"))\ndf <- as.DataFrame(mtcars)\nhead(df)\n{code}\n\nAnd this is the exception in executor log.\n{noformat}\n16/08/24 15:33:45 INFO BufferedStreamThread: Fatal error: cannot open file '/Users/jzhang/Temp/hadoop_tmp/nm-local-dir/usercache/jzhang/appcache/application_1471846125517_0022/container_1471846125517_0022_01_000002/sparkr/SparkR/worker/daemon.R': No such file or directory\n16/08/24 15:33:55 ERROR Executor: Exception in task 0.0 in stage 3.0 (TID 6)\njava.net.SocketTimeoutException: Accept timed out\n at java.net.PlainSocketImpl.socketAccept(Native Method)\n at java.net.AbstractPlainSocketImpl.accept(AbstractPlainSocketImpl.java:404)\n at java.net.ServerSocket.implAccept(ServerSocket.java:545)\n at java.net.ServerSocket.accept(ServerSocket.java:513)\n at org.apache.spark.api.r.RRunner$.createRWorker(RRunner.scala:367)\n at org.apache.spark.api.r.RRunner.compute(RRunner.scala:69)\n at org.apache.spark.api.r.BaseRRDD.compute(RRDD.scala:49)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)\n at org.apache.spark.scheduler.Task.run(Task.scala:86)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{noformat}","from":"reporter","subject":"sparkr.zip is not distributed to executors when run sparkr in RStudio"},{"body":"User 'zjffdu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14784","from":"developer"},{"body":"cc [~rxin] - this is merged to master and branch-2.0. If we are rolling another RC for Spark 2.0.1 will be good to include this fix. Thanks","from":"developer"},{"body":"[~felixcheung] please set the version to 2.0.2 from now on, since rc was already cut. If there is a new RC, I will change all 2.0.2 tickets to 2.0.1 when I cut the next one.\n","from":"developer"},{"body":"Got it, sorry about that, I should have noticed.","from":"developer"}],"created":"2016-08-24T07:37:40.000+0000","description":"Here's the code to reproduce this issue. \n{code}\nSys.setenv(SPARK_HOME=\"/Users/jzhang/github/spark\")\n.libPaths(c(file.path(Sys.getenv(), \"R\", \"lib\"), .libPaths()))\nlibrary(SparkR)\nsparkR.session(master=\"yarn-client\", sparkConfig = list(spark.executor.instances=\"1\"))\ndf <- as.DataFrame(mtcars)\nhead(df)\n{code}\n\nAnd this is the exception in executor log.\n{noformat}\n16/08/24 15:33:45 INFO BufferedStreamThread: Fatal error: cannot open file '/Users/jzhang/Temp/hadoop_tmp/nm-local-dir/usercache/jzhang/appcache/application_1471846125517_0022/container_1471846125517_0022_01_000002/sparkr/SparkR/worker/daemon.R': No such file or directory\n16/08/24 15:33:55 ERROR Executor: Exception in task 0.0 in stage 3.0 (TID 6)\njava.net.SocketTimeoutException: Accept timed out\n at java.net.PlainSocketImpl.socketAccept(Native Method)\n at java.net.AbstractPlainSocketImpl.accept(AbstractPlainSocketImpl.java:404)\n at java.net.ServerSocket.implAccept(ServerSocket.java:545)\n at java.net.ServerSocket.accept(ServerSocket.java:513)\n at org.apache.spark.api.r.RRunner$.createRWorker(RRunner.scala:367)\n at org.apache.spark.api.r.RRunner.compute(RRunner.scala:69)\n at org.apache.spark.api.r.BaseRRDD.compute(RRDD.scala:49)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)\n at org.apache.spark.scheduler.Task.run(Task.scala:86)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{noformat}","issue_id":"12999534","key":"SPARK-17210","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-09-23T18:39:37.000+0000","role":"fixed_distractor","summary":"sparkr.zip is not distributed to executors when run sparkr in RStudio"} {"case_id":"12999540","cluster":"DISTRACTOR-SPARK-17211","comments":[{"body":"Hi, [~jseppanen].\n\nThank you for the reporting. But, it seems to work correctly like the following in Apache Spark 2.0.0 for me.\n{code}\n>>> data_df.join(func.broadcast(keys_df), 'key_id').show()\n+--------+---+-----+\n| key_id|foo|value|\n+--------+---+-----+\n|54000002| 1| 2|\n|54000000| 2| 0|\n|54000001| 3| 1|\n+--------+---+-----+\n\n>>> data_df.join(keys_df, 'key_id').show()\n+--------+---+-----+\n| key_id|foo|value|\n+--------+---+-----+\n|54000002| 1| 2|\n|54000001| 3| 1|\n|54000000| 2| 0|\n+--------+---+-----+\n\n>>> spark.version\nu'2.0.0'\n{code}\n\nI tested with the following three environments.\n- Local Mode with spark-2.0.0-bin-hadoop2.7\n- [Databricks CE|https://databricks-prod-cloudfront.cloud.databricks.com/public/4027ec902e239c93eaaa8714f173bcfc/6660119172909095/2828921900907106/5162191866050912/latest.html]\n- Yarn Client mode on Hadoop 2.7.2\n\nIs there something for me to reproduce your problem?","created":"2016-08-24T18:06:21.710+0000"},{"body":"[~dongjoon] [~jseppanen] I am also seeing this issue on EMR using release emr-5.0.0, Amazon Hadoop 2.7.2 and Spark 2.0.0\n\nFor a simple join {{df1.join(df2, \"id\")}} the broadcast causes the fields from df2 to either return \"null\" or get assigned an incorrect value. Disabling the broadcast works as expected.\n\nEverything works fine locally through test cases. Could it be something to do with the EMR environment ?","created":"2016-08-24T19:14:30.117+0000"},{"body":"Up to now, this seems to happen in EMR 5.0.0 only since I tested with official Apache Hadoop 2.7.2 and Apache Spark 2.0.0 in yarn-client mode.","created":"2016-08-24T19:26:05.202+0000"},{"body":"Is there any way to meet that situation without EMR? Unfortunately, I don't have EMR environment.","created":"2016-08-24T19:28:36.664+0000"},{"body":"I ran the following in a Databricks environment with Spark 2.0. Works fine.\n\n{code:java}\nimport spark.implicits._\n\nval a1 = Array((123,1),(234,2),(432,5))\nval a2 = Array((\"abc\",1),(\"bcd\",2),(\"dcb\",5))\nval df1 = sc.parallelize(a1).toDF(\"gid\",\"id\")\nval df2 = sc.parallelize(a2).toDF(\"gname\",\"id\")\ndf1.join(df2,\"id\").show() // WORKS\n+---+---+-----+\n| id|gid|gname|\n+---+---+-----+\n| 5|432| dcb|\n| 2|234| bcd|\n| 1|123| abc|\n+---+---+-----+\ndf1.join(broadcast(df2),\"id\").show() // BROADCASTING - DOES NOT WORK on EMR\n+---+---+-----+\n| id|gid|gname|\n+---+---+-----+\n| 1|123| null|\n| 2|234| null|\n| 5|432| null|\n+---+---+-----+\nbroadcast(df1).join(df2,\"id\").show() // BROADCASTING - DOES NOT WORK on EMR\n{code}","created":"2016-08-24T21:28:23.842+0000"},{"body":"@Himanish - How are you running the spark-shell ?\n\nAre you passing any custom parameters like --driver-memory etc to the spark-shell ?\n","created":"2016-08-25T08:37:15.542+0000"},{"body":"Hi [~gurmukhd] \n\nI am using this command on a cluster with r3.2xlarge node instances ( with 61G memory) : \n{code}\nspark-shell --master yarn --deploy-mode client --num-executors 20 --executor-cores 2 --executor-memory 12g --driver-memory 48g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1\n{code}\n\n*One important thing to note*, if I allocate less memory to the driver {{--driver-memory 12g}} , it works as expected and produces the correct results even with broadcasting. Maybe some weird memory issues during broadcasting ?\n\nHi [~dongjoon] , will it be possible for you to check whether you see similar behavior with high driver memory settings on the three environments you tested on ?\n\nThanks ","created":"2016-08-25T13:02:32.230+0000"},{"body":"Interesting, but the use case does not consume such a large memory. If that is the real situation, we had better update the description.","created":"2016-08-25T13:33:26.995+0000"},{"body":"Thanks, I've also filed the issue with Amazon, since it seems possible that it's EMR specific.","created":"2016-08-30T07:04:27.737+0000"},{"body":"I'll keep it open a while but if we can't reproduce in Spark, not sure what we might do here.","created":"2016-08-31T11:08:01.805+0000"},{"body":"Hi\n\nI can see this in Apache Spark 2.0 as well, running with same node configurations as mentioned above.\n\nApache Hadoop 2.72., Spark 2.0:\n-------------------------------------------\n\n[hadoop@sp1 ~]$ hadoop version\nHadoop 2.7.2\nSubversion Unknown -r Unknown\nCompiled by root on 2016-05-16T03:56Z\nCompiled with protoc 2.5.0\nFrom source with checksum d0fda26633fa762bff87ec759ebe689c\nThis command was run using /opt/cluster/hadoop-2.7.2/share/hadoop/common/hadoop-common-2.7.2.jar\n\n[hadoop@sp1 ~]$ spark-shell --version\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /___/ .__/\\_,_/_/ /_/\\_\\ version 2.0.0\n /_/\n\nBranch\nCompiled by user jenkins on 2016-07-19T21:16:09Z\nRevision\n\n[hadoop@sp1 hadoop]$ spark-shell --master yarn --deploy-mode client --num-executors 20 --executor-cores 2 --executor-memory 12g --driver-memory 48g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1\nSetting default log level to \"WARN\".\nTo adjust logging level use sc.setLogLevel(newLevel).\n16/08/31 04:29:48 WARN util.NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable\n16/08/31 04:29:49 WARN yarn.Client: Neither spark.yarn.jars nor spark.yarn.archive is set, falling back to uploading libraries under SPARK_HOME.\n16/08/31 04:30:19 WARN spark.SparkContext: Use an existing SparkContext, some configuration may not take effect.\nSpark context Web UI available at http://10.0.0.227:4040\nSpark context available as 'sc' (master = yarn, app id = application_1472617754154_0001).\nSpark session available as 'spark'.\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /___/ .__/\\_,_/_/ /_/\\_\\ version 2.0.0\n /_/\n\nUsing Scala version 2.11.8 (Java HotSpot(TM) 64-Bit Server VM, Java 1.8.0_101)\nType in expressions to have them evaluated.\nType :help for more information.\n\nscala> val a1 = Array((123,1),(234,2),(432,5))\na1: Array[(Int, Int)] = Array((123,1), (234,2), (432,5))\n\nscala> val a2 = Array((\"abc\",1),(\"bcd\",2),(\"dcb\",5))\na2: Array[(String, Int)] = Array((abc,1), (bcd,2), (dcb,5))\n\nscala> val df1 = sc.parallelize(a1).toDF(\"gid\",\"id\")\ndf1: org.apache.spark.sql.DataFrame = [gid: int, id: int]\n\nscala> val df2 = sc.parallelize(a2).toDF(\"gname\",\"id\")\ndf2: org.apache.spark.sql.DataFrame = [gname: string, id: int]\n\n\nscala> df1.join(df2,\"id\").show()\n+---+---+-----+\n| id|gid|gname|\n+---+---+-----+\n| 1|123| abc|\n| 2|234| bcd|\n| 5|432| dcb|\n+---+---+-----+\n\nscala> val df2 = sc.parallelize(a2).toDF(\"gname\",\"id\")\ndf2: org.apache.spark.sql.DataFrame = [gname: string, id: int]\n\nscala> df1.join(broadcast(df2),\"id\").show()\n+---+---+-----+\n| id|gid|gname|\n+---+---+-----+\n| 1|123| null|\n| 2|234| null|\n| 5|432| null|\n+---+---+-----+\n\nscala> broadcast(df1).join(df2,\"id\").show()\n+---+---+-----+\n| id|gid|gname|\n+---+---+-----+\n| 0| 1| abc|\n| 0| 2| bcd|\n| 0| 5| dcb|\n+---+---+-----+\n\nIf I reduce the driver memory, this works as well. \n\nIt works on Apache spark 1.6 \n\nAs lot of things have changed in Spark 2.0, it needs to be looked upon. It should give error or OOM, instead of returning NULL or ZERO values.\n\n[~himanish] Although, it will be interesting to understand the use case that on a node with 61 GB, executing with driver memory=48GB, leaving just 12 GB for so many other things, when there are other overheads on the system.\n\nOn DataBricks, are you running with same parameters ?\n","created":"2016-08-31T22:17:35.293+0000"},{"body":"[~gurmukhd] You are seeing the errors outside of EMR environment also, right ? On Databricks I ran it through a scala notebook, not sure what the underlying configuration was. \n\nI would want to add one more thing , while running the real job (with lots of data), I noticed even with low driver memory settings and broadcasting enabled , some joins works fine but others messes up the data (either assigns null or field values get mixed up). What are the other things that would use up 12 GB of memory on the node ?","created":"2016-09-01T01:13:46.742+0000"},{"body":"yes, outside of EMR.\n\nOn a node you will have OS, nodemanger, datanode node daemons using memory.\n\nOne other we might need to look at is java 1.8. I have used java 1.8 on Apache Spark 2.0, for tests.","created":"2016-09-01T03:15:27.508+0000"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14927","created":"2016-09-01T21:00:11.002+0000"},{"body":"Thanks Davies.\n\nCan see the issue with the offsets:\n\nIt picks some random values beyond offset and then tries to do a join on rows, which will obviosuly return NULL as those rows are not there:\n\nscala> val bc=sc.broadcast(df1)\n\nscala> bc.value.view.take(10).foreach(println)\n(abc,1)\n(bcd,2)\n(dcb,5)\n\nscala> bc.value.view.toDF\nres13: org.apache.spark.sql.DataFrame = [_1: string, _2: int]\n\nscala> bc.value.view.toDF(\"gid\", \"id\")\nres14: org.apache.spark.sql.DataFrame = [gid: string, id: int]\n\nscala> bc.value.view.toDF(\"gid\", \"id\").show()\n+---+---+\n|gid| id|\n+---+---+\n|abc| 1|\n|bcd| 2|\n|dcb| 5|\n+---+---+\n\nscala> df1.join(bc.value.view.toDF(\"gid\", \"id\"), \"id\").show()\n+---+---+----+\n| id|gid| gid|\n+---+---+----+\n| 1|123|null|\n| 2|234|null|\n| 5|432|null|\n+---+---+----+\n\n\nscala> df1.join(bc.value.view.toDF(\"gidyy\", \"id\"), \"id\").show()\n+---+---+-----+\n| id|gid|gidyy|\n+---+---+-----+\n| 1|123| null|\n| 2|234| null|\n| 5|432| null|\n+---+---+-----+\n\n\nscala> bc.value.view.toDF(\"gid\", \"id\").show()\n+---+---+\n|gid| id|\n+---+---+\n|abc| 1|\n|bcd| 2|\n|dcb| 5|\n+---+---+\n\n\nscala> df1.show()\n+---+---+\n|gid| id|\n+---+---+\n|123| 1|\n|234| 2|\n|432| 5|\n+---+---+\n\n\nscala> df1.join(bc.value.view.toDF(\"gidyy\", \"id\"), \"id\").show()\n+---+---+-----+\n| id|gid|gidyy|\n+---+---+-----+\n| 1|123| null|\n| 2|234| null|\n| 5|432| null|\n+---+---+-----+\n\n\nscala> bc.value.view.toDF(\"gid\", \"id\").join(df1, \"id\").show()\n+-------+----+---+\n| id| gid|gid|\n+-------+----+---+\n|6513249|null|123|\t<----- Look at the \"id\" where is it picking from ? is is fetching some out of address locations, which it should not be\n|6579042|null|234|\n|6447972|null|432|\n+-------+----+---+\n\n\nscala> bc.value.view.toDF(\"gidy\", \"id\").join(df1, \"id\").show()\n+-------+----+---+\n| id|gidy|gid|\n+-------+----+---+\n|6513249|null|123|\n|6579042|null|234|\n|6447972|null|432|\n+-------+----+---+","created":"2016-09-01T22:52:17.328+0000"},{"body":"Hi,\n\nI have been testing this and it seems to be related to memory pointers compression. It happens when the JVMs of the executors use different memory pointers than the driver. The JVM by default enables UseCompressedOops for heaps under 32 GB. So if both the driver and executors memory heaps are over 32 GB, or if both are under 32 GB, there is no problem.\n\nAs a workaround, disabling UseCompressedOops in the smaller heap seems to work. For example:\nspark-shell --driver-memory 30G --executor-memory 40G --driver-java-options \"-XX:-UseCompressedOops\" ...\n\nOr the other case:\nspark-shell --driver-memory 45G --executor-memory 30G --conf \"spark.executor.extraJavaOptions=-XX:-UseCompressedOops\" ...\n\nI also tested JVM 1.7 and 1.8, with the same results. On EMR or oustide, on YARN or standalone, it doesn't matter. Always tested with Spark 2.0.0 using Scala 2.11.8.","created":"2016-09-02T11:09:44.559+0000"},{"body":"Yeah, almost certainly related to compressed Oops. It'd be great to try master if you can, but I don't know if something else would have addressed it. It seems legitimate though.","created":"2016-09-02T14:17:59.052+0000"},{"body":"I've just tried master from git, exactly the same results.","created":"2016-09-02T16:18:34.442+0000"},{"body":"[~migtor] Could you try this patch ? https://github.com/apache/spark/pull/14927","created":"2016-09-02T16:58:29.135+0000"},{"body":"Oh, now I see the point of this issue.","created":"2016-09-02T17:26:54.360+0000"},{"body":"Thanks,\n\nI have tested by disabling UseCompressedOops \"-XX:-UseCompressedOops\" and as stated by [~migtor], it works only for heap <32 GB","created":"2016-09-02T22:34:09.293+0000"},{"body":"Could you try the patch ? https://github.com/apache/spark/pull/14927","created":"2016-09-02T23:01:45.494+0000"},{"body":"Sure, will update soon with my findings.","created":"2016-09-02T23:26:58.804+0000"},{"body":"[~gurmukhd], you say it works only for heap < 32 GB, which one do you mean? Driver memory or executors? There is only need to disable compressed oops for the heap that is under 32 GB when the other isn't. You can test disabling it for both cases to be on the safe side:\n\nspark-shell ... --driver-java-options \"-XX:-UseCompressedOops\" --conf \"spark.executor.extraJavaOptions=-XX:-UseCompressedOops\" ...\n\n[~davies] I just tried the patch, it seems to work! I couldn't break it although probably it needs more thorough testing to cover other cases. But it does seem to fix the issue, thanks!\n\n","created":"2016-09-02T23:55:26.702+0000"},{"body":"[~davies] After applying the patch, tested with various combination of executor and driver memory. The issue seems to have been fixed.\n\nTested for cases as below:\n\n$ spark-shell --master yarn --deploy-mode client --executor-memory 12g --driver-memory 48g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1\n\n$ spark-shell --master yarn --deploy-mode client --executor-memory 48g --driver-memory 12g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1\n\n$ spark-shell --master yarn --deploy-mode client --executor-memory 12g --driver-memory 12g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1\n\n$ spark-shell --master yarn --deploy-mode client --executor-memory 38g --driver-memory 38g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1","created":"2016-09-03T01:08:16.277+0000"},{"body":"Issue resolved by pull request 14927\n[https://github.com/apache/spark/pull/14927]","created":"2016-09-06T17:47:15.373+0000"}],"conversations":[{"body":"Broadcast join produces incorrect columns in join result, see below for an example. The same join but without using broadcast gives the correct columns.\n\nRunning PySpark on YARN on Amazon EMR 5.0.0.\n\n{noformat}\n\nimport pyspark.sql.functions as func\n\nkeys = [\n (54000000, 0),\n (54000001, 1),\n (54000002, 2),\n]\n\nkeys_df = spark.createDataFrame(keys, ['key_id', 'value']).coalesce(1)\nkeys_df.show()\n# +--------+-----+\n# | key_id|value|\n# +--------+-----+\n# |54000000| 0|\n# |54000001| 1|\n# |54000002| 2|\n# +--------+-----+\n\ndata = [\n (54000002, 1),\n (54000000, 2),\n (54000001, 3),\n]\n\ndata_df = spark.createDataFrame(data, ['key_id', 'foo'])\ndata_df.show()\n# +--------+---+ \n# | key_id|foo|\n# +--------+---+\n# |54000002| 1|\n# |54000000| 2|\n# |54000001| 3|\n# +--------+---+\n\n### INCORRECT ###\n\ndata_df.join(func.broadcast(keys_df), 'key_id').show()\n# +--------+---+--------+ \n# | key_id|foo| value|\n# +--------+---+--------+\n# |54000002| 1|54000002|\n# |54000000| 2|54000000|\n# |54000001| 3|54000001|\n# +--------+---+--------+\n\n### CORRECT ###\n\ndata_df.join(keys_df, 'key_id').show()\n# +--------+---+-----+\n# | key_id|foo|value|\n# +--------+---+-----+\n# |54000000| 2| 0|\n# |54000001| 3| 1|\n# |54000002| 1| 2|\n# +--------+---+-----+\n{noformat}\n","from":"reporter","subject":"Broadcast join produces incorrect results when compressed Oops differs between driver, executor"},{"body":"Hi, [~jseppanen].\n\nThank you for the reporting. But, it seems to work correctly like the following in Apache Spark 2.0.0 for me.\n{code}\n>>> data_df.join(func.broadcast(keys_df), 'key_id').show()\n+--------+---+-----+\n| key_id|foo|value|\n+--------+---+-----+\n|54000002| 1| 2|\n|54000000| 2| 0|\n|54000001| 3| 1|\n+--------+---+-----+\n\n>>> data_df.join(keys_df, 'key_id').show()\n+--------+---+-----+\n| key_id|foo|value|\n+--------+---+-----+\n|54000002| 1| 2|\n|54000001| 3| 1|\n|54000000| 2| 0|\n+--------+---+-----+\n\n>>> spark.version\nu'2.0.0'\n{code}\n\nI tested with the following three environments.\n- Local Mode with spark-2.0.0-bin-hadoop2.7\n- [Databricks CE|https://databricks-prod-cloudfront.cloud.databricks.com/public/4027ec902e239c93eaaa8714f173bcfc/6660119172909095/2828921900907106/5162191866050912/latest.html]\n- Yarn Client mode on Hadoop 2.7.2\n\nIs there something for me to reproduce your problem?","from":"developer"},{"body":"[~dongjoon] [~jseppanen] I am also seeing this issue on EMR using release emr-5.0.0, Amazon Hadoop 2.7.2 and Spark 2.0.0\n\nFor a simple join {{df1.join(df2, \"id\")}} the broadcast causes the fields from df2 to either return \"null\" or get assigned an incorrect value. Disabling the broadcast works as expected.\n\nEverything works fine locally through test cases. Could it be something to do with the EMR environment ?","from":"developer"},{"body":"Up to now, this seems to happen in EMR 5.0.0 only since I tested with official Apache Hadoop 2.7.2 and Apache Spark 2.0.0 in yarn-client mode.","from":"developer"},{"body":"Is there any way to meet that situation without EMR? Unfortunately, I don't have EMR environment.","from":"developer"},{"body":"I ran the following in a Databricks environment with Spark 2.0. Works fine.\n\n{code:java}\nimport spark.implicits._\n\nval a1 = Array((123,1),(234,2),(432,5))\nval a2 = Array((\"abc\",1),(\"bcd\",2),(\"dcb\",5))\nval df1 = sc.parallelize(a1).toDF(\"gid\",\"id\")\nval df2 = sc.parallelize(a2).toDF(\"gname\",\"id\")\ndf1.join(df2,\"id\").show() // WORKS\n+---+---+-----+\n| id|gid|gname|\n+---+---+-----+\n| 5|432| dcb|\n| 2|234| bcd|\n| 1|123| abc|\n+---+---+-----+\ndf1.join(broadcast(df2),\"id\").show() // BROADCASTING - DOES NOT WORK on EMR\n+---+---+-----+\n| id|gid|gname|\n+---+---+-----+\n| 1|123| null|\n| 2|234| null|\n| 5|432| null|\n+---+---+-----+\nbroadcast(df1).join(df2,\"id\").show() // BROADCASTING - DOES NOT WORK on EMR\n{code}","from":"developer"},{"body":"@Himanish - How are you running the spark-shell ?\n\nAre you passing any custom parameters like --driver-memory etc to the spark-shell ?\n","from":"developer"},{"body":"Hi [~gurmukhd] \n\nI am using this command on a cluster with r3.2xlarge node instances ( with 61G memory) : \n{code}\nspark-shell --master yarn --deploy-mode client --num-executors 20 --executor-cores 2 --executor-memory 12g --driver-memory 48g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1\n{code}\n\n*One important thing to note*, if I allocate less memory to the driver {{--driver-memory 12g}} , it works as expected and produces the correct results even with broadcasting. Maybe some weird memory issues during broadcasting ?\n\nHi [~dongjoon] , will it be possible for you to check whether you see similar behavior with high driver memory settings on the three environments you tested on ?\n\nThanks ","from":"developer"},{"body":"Interesting, but the use case does not consume such a large memory. If that is the real situation, we had better update the description.","from":"developer"},{"body":"Thanks, I've also filed the issue with Amazon, since it seems possible that it's EMR specific.","from":"developer"},{"body":"I'll keep it open a while but if we can't reproduce in Spark, not sure what we might do here.","from":"developer"},{"body":"Hi\n\nI can see this in Apache Spark 2.0 as well, running with same node configurations as mentioned above.\n\nApache Hadoop 2.72., Spark 2.0:\n-------------------------------------------\n\n[hadoop@sp1 ~]$ hadoop version\nHadoop 2.7.2\nSubversion Unknown -r Unknown\nCompiled by root on 2016-05-16T03:56Z\nCompiled with protoc 2.5.0\nFrom source with checksum d0fda26633fa762bff87ec759ebe689c\nThis command was run using /opt/cluster/hadoop-2.7.2/share/hadoop/common/hadoop-common-2.7.2.jar\n\n[hadoop@sp1 ~]$ spark-shell --version\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /___/ .__/\\_,_/_/ /_/\\_\\ version 2.0.0\n /_/\n\nBranch\nCompiled by user jenkins on 2016-07-19T21:16:09Z\nRevision\n\n[hadoop@sp1 hadoop]$ spark-shell --master yarn --deploy-mode client --num-executors 20 --executor-cores 2 --executor-memory 12g --driver-memory 48g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1\nSetting default log level to \"WARN\".\nTo adjust logging level use sc.setLogLevel(newLevel).\n16/08/31 04:29:48 WARN util.NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable\n16/08/31 04:29:49 WARN yarn.Client: Neither spark.yarn.jars nor spark.yarn.archive is set, falling back to uploading libraries under SPARK_HOME.\n16/08/31 04:30:19 WARN spark.SparkContext: Use an existing SparkContext, some configuration may not take effect.\nSpark context Web UI available at http://10.0.0.227:4040\nSpark context available as 'sc' (master = yarn, app id = application_1472617754154_0001).\nSpark session available as 'spark'.\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /___/ .__/\\_,_/_/ /_/\\_\\ version 2.0.0\n /_/\n\nUsing Scala version 2.11.8 (Java HotSpot(TM) 64-Bit Server VM, Java 1.8.0_101)\nType in expressions to have them evaluated.\nType :help for more information.\n\nscala> val a1 = Array((123,1),(234,2),(432,5))\na1: Array[(Int, Int)] = Array((123,1), (234,2), (432,5))\n\nscala> val a2 = Array((\"abc\",1),(\"bcd\",2),(\"dcb\",5))\na2: Array[(String, Int)] = Array((abc,1), (bcd,2), (dcb,5))\n\nscala> val df1 = sc.parallelize(a1).toDF(\"gid\",\"id\")\ndf1: org.apache.spark.sql.DataFrame = [gid: int, id: int]\n\nscala> val df2 = sc.parallelize(a2).toDF(\"gname\",\"id\")\ndf2: org.apache.spark.sql.DataFrame = [gname: string, id: int]\n\n\nscala> df1.join(df2,\"id\").show()\n+---+---+-----+\n| id|gid|gname|\n+---+---+-----+\n| 1|123| abc|\n| 2|234| bcd|\n| 5|432| dcb|\n+---+---+-----+\n\nscala> val df2 = sc.parallelize(a2).toDF(\"gname\",\"id\")\ndf2: org.apache.spark.sql.DataFrame = [gname: string, id: int]\n\nscala> df1.join(broadcast(df2),\"id\").show()\n+---+---+-----+\n| id|gid|gname|\n+---+---+-----+\n| 1|123| null|\n| 2|234| null|\n| 5|432| null|\n+---+---+-----+\n\nscala> broadcast(df1).join(df2,\"id\").show()\n+---+---+-----+\n| id|gid|gname|\n+---+---+-----+\n| 0| 1| abc|\n| 0| 2| bcd|\n| 0| 5| dcb|\n+---+---+-----+\n\nIf I reduce the driver memory, this works as well. \n\nIt works on Apache spark 1.6 \n\nAs lot of things have changed in Spark 2.0, it needs to be looked upon. It should give error or OOM, instead of returning NULL or ZERO values.\n\n[~himanish] Although, it will be interesting to understand the use case that on a node with 61 GB, executing with driver memory=48GB, leaving just 12 GB for so many other things, when there are other overheads on the system.\n\nOn DataBricks, are you running with same parameters ?\n","from":"developer"},{"body":"[~gurmukhd] You are seeing the errors outside of EMR environment also, right ? On Databricks I ran it through a scala notebook, not sure what the underlying configuration was. \n\nI would want to add one more thing , while running the real job (with lots of data), I noticed even with low driver memory settings and broadcasting enabled , some joins works fine but others messes up the data (either assigns null or field values get mixed up). What are the other things that would use up 12 GB of memory on the node ?","from":"developer"},{"body":"yes, outside of EMR.\n\nOn a node you will have OS, nodemanger, datanode node daemons using memory.\n\nOne other we might need to look at is java 1.8. I have used java 1.8 on Apache Spark 2.0, for tests.","from":"developer"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14927","from":"developer"},{"body":"Thanks Davies.\n\nCan see the issue with the offsets:\n\nIt picks some random values beyond offset and then tries to do a join on rows, which will obviosuly return NULL as those rows are not there:\n\nscala> val bc=sc.broadcast(df1)\n\nscala> bc.value.view.take(10).foreach(println)\n(abc,1)\n(bcd,2)\n(dcb,5)\n\nscala> bc.value.view.toDF\nres13: org.apache.spark.sql.DataFrame = [_1: string, _2: int]\n\nscala> bc.value.view.toDF(\"gid\", \"id\")\nres14: org.apache.spark.sql.DataFrame = [gid: string, id: int]\n\nscala> bc.value.view.toDF(\"gid\", \"id\").show()\n+---+---+\n|gid| id|\n+---+---+\n|abc| 1|\n|bcd| 2|\n|dcb| 5|\n+---+---+\n\nscala> df1.join(bc.value.view.toDF(\"gid\", \"id\"), \"id\").show()\n+---+---+----+\n| id|gid| gid|\n+---+---+----+\n| 1|123|null|\n| 2|234|null|\n| 5|432|null|\n+---+---+----+\n\n\nscala> df1.join(bc.value.view.toDF(\"gidyy\", \"id\"), \"id\").show()\n+---+---+-----+\n| id|gid|gidyy|\n+---+---+-----+\n| 1|123| null|\n| 2|234| null|\n| 5|432| null|\n+---+---+-----+\n\n\nscala> bc.value.view.toDF(\"gid\", \"id\").show()\n+---+---+\n|gid| id|\n+---+---+\n|abc| 1|\n|bcd| 2|\n|dcb| 5|\n+---+---+\n\n\nscala> df1.show()\n+---+---+\n|gid| id|\n+---+---+\n|123| 1|\n|234| 2|\n|432| 5|\n+---+---+\n\n\nscala> df1.join(bc.value.view.toDF(\"gidyy\", \"id\"), \"id\").show()\n+---+---+-----+\n| id|gid|gidyy|\n+---+---+-----+\n| 1|123| null|\n| 2|234| null|\n| 5|432| null|\n+---+---+-----+\n\n\nscala> bc.value.view.toDF(\"gid\", \"id\").join(df1, \"id\").show()\n+-------+----+---+\n| id| gid|gid|\n+-------+----+---+\n|6513249|null|123|\t<----- Look at the \"id\" where is it picking from ? is is fetching some out of address locations, which it should not be\n|6579042|null|234|\n|6447972|null|432|\n+-------+----+---+\n\n\nscala> bc.value.view.toDF(\"gidy\", \"id\").join(df1, \"id\").show()\n+-------+----+---+\n| id|gidy|gid|\n+-------+----+---+\n|6513249|null|123|\n|6579042|null|234|\n|6447972|null|432|\n+-------+----+---+","from":"developer"},{"body":"Hi,\n\nI have been testing this and it seems to be related to memory pointers compression. It happens when the JVMs of the executors use different memory pointers than the driver. The JVM by default enables UseCompressedOops for heaps under 32 GB. So if both the driver and executors memory heaps are over 32 GB, or if both are under 32 GB, there is no problem.\n\nAs a workaround, disabling UseCompressedOops in the smaller heap seems to work. For example:\nspark-shell --driver-memory 30G --executor-memory 40G --driver-java-options \"-XX:-UseCompressedOops\" ...\n\nOr the other case:\nspark-shell --driver-memory 45G --executor-memory 30G --conf \"spark.executor.extraJavaOptions=-XX:-UseCompressedOops\" ...\n\nI also tested JVM 1.7 and 1.8, with the same results. On EMR or oustide, on YARN or standalone, it doesn't matter. Always tested with Spark 2.0.0 using Scala 2.11.8.","from":"developer"},{"body":"Yeah, almost certainly related to compressed Oops. It'd be great to try master if you can, but I don't know if something else would have addressed it. It seems legitimate though.","from":"developer"},{"body":"I've just tried master from git, exactly the same results.","from":"developer"},{"body":"[~migtor] Could you try this patch ? https://github.com/apache/spark/pull/14927","from":"developer"},{"body":"Oh, now I see the point of this issue.","from":"developer"},{"body":"Thanks,\n\nI have tested by disabling UseCompressedOops \"-XX:-UseCompressedOops\" and as stated by [~migtor], it works only for heap <32 GB","from":"developer"},{"body":"Could you try the patch ? https://github.com/apache/spark/pull/14927","from":"developer"},{"body":"Sure, will update soon with my findings.","from":"developer"},{"body":"[~gurmukhd], you say it works only for heap < 32 GB, which one do you mean? Driver memory or executors? There is only need to disable compressed oops for the heap that is under 32 GB when the other isn't. You can test disabling it for both cases to be on the safe side:\n\nspark-shell ... --driver-java-options \"-XX:-UseCompressedOops\" --conf \"spark.executor.extraJavaOptions=-XX:-UseCompressedOops\" ...\n\n[~davies] I just tried the patch, it seems to work! I couldn't break it although probably it needs more thorough testing to cover other cases. But it does seem to fix the issue, thanks!\n\n","from":"developer"},{"body":"[~davies] After applying the patch, tested with various combination of executor and driver memory. The issue seems to have been fixed.\n\nTested for cases as below:\n\n$ spark-shell --master yarn --deploy-mode client --executor-memory 12g --driver-memory 48g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1\n\n$ spark-shell --master yarn --deploy-mode client --executor-memory 48g --driver-memory 12g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1\n\n$ spark-shell --master yarn --deploy-mode client --executor-memory 12g --driver-memory 12g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1\n\n$ spark-shell --master yarn --deploy-mode client --executor-memory 38g --driver-memory 38g --conf spark.yarn.executor.memoryOverhead=4096 --conf spark.sql.shuffle.partitions=1024 --conf spark.yarn.maxAppAttempts=1","from":"developer"},{"body":"Issue resolved by pull request 14927\n[https://github.com/apache/spark/pull/14927]","from":"developer"}],"created":"2016-08-24T07:53:03.000+0000","description":"Broadcast join produces incorrect columns in join result, see below for an example. The same join but without using broadcast gives the correct columns.\n\nRunning PySpark on YARN on Amazon EMR 5.0.0.\n\n{noformat}\n\nimport pyspark.sql.functions as func\n\nkeys = [\n (54000000, 0),\n (54000001, 1),\n (54000002, 2),\n]\n\nkeys_df = spark.createDataFrame(keys, ['key_id', 'value']).coalesce(1)\nkeys_df.show()\n# +--------+-----+\n# | key_id|value|\n# +--------+-----+\n# |54000000| 0|\n# |54000001| 1|\n# |54000002| 2|\n# +--------+-----+\n\ndata = [\n (54000002, 1),\n (54000000, 2),\n (54000001, 3),\n]\n\ndata_df = spark.createDataFrame(data, ['key_id', 'foo'])\ndata_df.show()\n# +--------+---+ \n# | key_id|foo|\n# +--------+---+\n# |54000002| 1|\n# |54000000| 2|\n# |54000001| 3|\n# +--------+---+\n\n### INCORRECT ###\n\ndata_df.join(func.broadcast(keys_df), 'key_id').show()\n# +--------+---+--------+ \n# | key_id|foo| value|\n# +--------+---+--------+\n# |54000002| 1|54000002|\n# |54000000| 2|54000000|\n# |54000001| 3|54000001|\n# +--------+---+--------+\n\n### CORRECT ###\n\ndata_df.join(keys_df, 'key_id').show()\n# +--------+---+-----+\n# | key_id|foo|value|\n# +--------+---+-----+\n# |54000000| 2| 0|\n# |54000001| 3| 1|\n# |54000002| 1| 2|\n# +--------+---+-----+\n{noformat}\n","issue_id":"12999540","key":"SPARK-17211","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-09-06T17:47:14.000+0000","role":"fixed_distractor","summary":"Broadcast join produces incorrect results when compressed Oops differs between driver, executor"} {"case_id":"13000255","cluster":"DISTRACTOR-SPARK-17251","comments":[{"body":"Ok, I have taken a look at this one. We should make {{OuterReference}} a {{NamedExpression}} and then we are good (have most of the code working locally). \n\nIf we fix this, it will fail analysis because we are using a correlated predicate in a {{Project}}. We could make an exception for IN, but I am just wondering if we support such a weird construct at all.","created":"2016-08-26T16:35:41.477+0000"},{"body":"Hi, [~hvanhovell].\nThe suggested solution seems to extend `OuterReference` with trait `NamedExpression` and eventually to raise the following correct exception. May I create a PR for this?\n{code}\norg.apache.spark.sql.AnalysisException: Correlated predicates are not supported outside of WHERE/HAVING clauses\n{code}","created":"2016-11-24T23:06:34.022+0000"},{"body":"Yeah, go ahead. The only thing is that it should not throw an exception, it must return an empty set.","created":"2016-11-24T23:22:34.128+0000"},{"body":"Thank you! I see.","created":"2016-11-24T23:46:14.865+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16012","created":"2016-11-25T11:31:04.739+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16015","created":"2016-11-26T03:34:05.275+0000"},{"body":"The current PR fixes the ClassCastException, and we now issue an AnalysisException. We should fix this in 2.2.","created":"2016-11-26T23:11:11.814+0000"}],"conversations":[{"body":"The following test case produces a ClassCastException in the analyzer:\n\n{code}\nCREATE TABLE t1(a INTEGER);\nINSERT INTO t1 VALUES(1),(2);\nCREATE TABLE t2(b INTEGER);\nINSERT INTO t2 VALUES(1);\n\nSELECT a FROM t1 WHERE a NOT IN (SELECT a FROM t2);\n{code}\n\nHere's the exception:\n\n{code}\njava.lang.ClassCastException: org.apache.spark.sql.catalyst.expressions.OuterReference cannot be cast to org.apache.spark.sql.catalyst.expressions.NamedExpression\n\tat org.apache.spark.sql.catalyst.plans.logical.Project$$anonfun$1.apply(basicLogicalOperators.scala:48)\n\tat scala.collection.LinearSeqOptimized$class.exists(LinearSeqOptimized.scala:80)\n\tat scala.collection.immutable.List.exists(List.scala:84)\n\tat org.apache.spark.sql.catalyst.plans.logical.Project.resolved$lzycompute(basicLogicalOperators.scala:44)\n\tat org.apache.spark.sql.catalyst.plans.logical.Project.resolved(basicLogicalOperators.scala:43)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$.org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveSubquery$$resolveSubQuery(Analyzer.scala:1091)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveSubquery$$resolveSubQueries$1.applyOrElse(Analyzer.scala:1130)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveSubquery$$resolveSubQueries$1.applyOrElse(Analyzer.scala:1116)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:278)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpressionDown$1(QueryPlan.scala:156)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.org$apache$spark$sql$catalyst$plans$QueryPlan$$recursiveTransform$1(QueryPlan.scala:166)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$4.apply(QueryPlan.scala:175)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpressionsDown(QueryPlan.scala:175)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpressions(QueryPlan.scala:144)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$.org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveSubquery$$resolveSubQueries(Analyzer.scala:1116)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$$anonfun$apply$16.applyOrElse(Analyzer.scala:1148)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$$anonfun$apply$16.applyOrElse(Analyzer.scala:1141)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolveOperators(LogicalPlan.scala:60)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$1.apply(LogicalPlan.scala:58)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$1.apply(LogicalPlan.scala:58)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolveOperators(LogicalPlan.scala:58)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$.apply(Analyzer.scala:1141)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$.apply(Analyzer.scala:909)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:85)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:82)\n\tat scala.collection.LinearSeqOptimized$class.foldLeft(LinearSeqOptimized.scala:111)\n\tat scala.collection.immutable.List.foldLeft(List.scala:84)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:82)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:74)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:74)\n\tat org.apache.spark.sql.execution.QueryExecution.analyzed$lzycompute(QueryExecution.scala:65)\n\tat org.apache.spark.sql.execution.QueryExecution.analyzed(QueryExecution.scala:63)\n\tat org.apache.spark.sql.execution.QueryExecution.assertAnalyzed(QueryExecution.scala:49)\n\tat org.apache.spark.sql.Dataset$.ofRows(Dataset.scala:64)\n\tat org.apache.spark.sql.SparkSession.sql(SparkSession.scala:582)\n\tat org.apache.spark.sql.SQLContext.sql(SQLContext.scala:682)\n{code}\n\nThis bug was discovered while trying to run SQLite bug reports through Spark SQL (see https://www.sqlite.org/src/tktview?name=5e3c886796)\n","from":"reporter","subject":"\"ClassCastException: OuterReference cannot be cast to NamedExpression\" for correlated subquery on the RHS of an IN operator"},{"body":"Ok, I have taken a look at this one. We should make {{OuterReference}} a {{NamedExpression}} and then we are good (have most of the code working locally). \n\nIf we fix this, it will fail analysis because we are using a correlated predicate in a {{Project}}. We could make an exception for IN, but I am just wondering if we support such a weird construct at all.","from":"developer"},{"body":"Hi, [~hvanhovell].\nThe suggested solution seems to extend `OuterReference` with trait `NamedExpression` and eventually to raise the following correct exception. May I create a PR for this?\n{code}\norg.apache.spark.sql.AnalysisException: Correlated predicates are not supported outside of WHERE/HAVING clauses\n{code}","from":"developer"},{"body":"Yeah, go ahead. The only thing is that it should not throw an exception, it must return an empty set.","from":"developer"},{"body":"Thank you! I see.","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16012","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16015","from":"developer"},{"body":"The current PR fixes the ClassCastException, and we now issue an AnalysisException. We should fix this in 2.2.","from":"developer"}],"created":"2016-08-26T02:25:52.000+0000","description":"The following test case produces a ClassCastException in the analyzer:\n\n{code}\nCREATE TABLE t1(a INTEGER);\nINSERT INTO t1 VALUES(1),(2);\nCREATE TABLE t2(b INTEGER);\nINSERT INTO t2 VALUES(1);\n\nSELECT a FROM t1 WHERE a NOT IN (SELECT a FROM t2);\n{code}\n\nHere's the exception:\n\n{code}\njava.lang.ClassCastException: org.apache.spark.sql.catalyst.expressions.OuterReference cannot be cast to org.apache.spark.sql.catalyst.expressions.NamedExpression\n\tat org.apache.spark.sql.catalyst.plans.logical.Project$$anonfun$1.apply(basicLogicalOperators.scala:48)\n\tat scala.collection.LinearSeqOptimized$class.exists(LinearSeqOptimized.scala:80)\n\tat scala.collection.immutable.List.exists(List.scala:84)\n\tat org.apache.spark.sql.catalyst.plans.logical.Project.resolved$lzycompute(basicLogicalOperators.scala:44)\n\tat org.apache.spark.sql.catalyst.plans.logical.Project.resolved(basicLogicalOperators.scala:43)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$.org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveSubquery$$resolveSubQuery(Analyzer.scala:1091)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveSubquery$$resolveSubQueries$1.applyOrElse(Analyzer.scala:1130)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveSubquery$$resolveSubQueries$1.applyOrElse(Analyzer.scala:1116)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:279)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:278)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:284)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpressionDown$1(QueryPlan.scala:156)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.org$apache$spark$sql$catalyst$plans$QueryPlan$$recursiveTransform$1(QueryPlan.scala:166)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$4.apply(QueryPlan.scala:175)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpressionsDown(QueryPlan.scala:175)\n\tat org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpressions(QueryPlan.scala:144)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$.org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveSubquery$$resolveSubQueries(Analyzer.scala:1116)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$$anonfun$apply$16.applyOrElse(Analyzer.scala:1148)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$$anonfun$apply$16.applyOrElse(Analyzer.scala:1141)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$resolveOperators$1.apply(LogicalPlan.scala:61)\n\tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:69)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolveOperators(LogicalPlan.scala:60)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$1.apply(LogicalPlan.scala:58)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan$$anonfun$1.apply(LogicalPlan.scala:58)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:321)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:179)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:319)\n\tat org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolveOperators(LogicalPlan.scala:58)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$.apply(Analyzer.scala:1141)\n\tat org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveSubquery$.apply(Analyzer.scala:909)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:85)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:82)\n\tat scala.collection.LinearSeqOptimized$class.foldLeft(LinearSeqOptimized.scala:111)\n\tat scala.collection.immutable.List.foldLeft(List.scala:84)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:82)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:74)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:74)\n\tat org.apache.spark.sql.execution.QueryExecution.analyzed$lzycompute(QueryExecution.scala:65)\n\tat org.apache.spark.sql.execution.QueryExecution.analyzed(QueryExecution.scala:63)\n\tat org.apache.spark.sql.execution.QueryExecution.assertAnalyzed(QueryExecution.scala:49)\n\tat org.apache.spark.sql.Dataset$.ofRows(Dataset.scala:64)\n\tat org.apache.spark.sql.SparkSession.sql(SparkSession.scala:582)\n\tat org.apache.spark.sql.SQLContext.sql(SQLContext.scala:682)\n{code}\n\nThis bug was discovered while trying to run SQLite bug reports through Spark SQL (see https://www.sqlite.org/src/tktview?name=5e3c886796)\n","issue_id":"13000255","key":"SPARK-17251","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-11-26T23:11:10.000+0000","role":"fixed_distractor","summary":"\"ClassCastException: OuterReference cannot be cast to NamedExpression\" for correlated subquery on the RHS of an IN operator"} {"case_id":"13001680","cluster":"DISTRACTOR-SPARK-17337","comments":[{"body":"This problem originated from the Case 1 documented in SPARK-16951. The script illustrates the problem\n\n{noformat}\nSeq(1,2).toDF(\"c1\").createOrReplaceTempView(\"t1\")\nSeq(1).toDF(\"c2\").createOrReplaceTempView(\"t2\")\n\nscala> sql(\"select * from (select t2.c2+1 as c3 from t1 left join t2 on t1.c1=t2.c2) t3 where c3 not in (select c2 from t2)\").show\n+----+\n| c3|\n+----+\n| 2|\n|null|\n+----+\n{noformat}\n\nThe correct answer is 1 row of (2). From the plan below, the incorrect portion of the plan is the LeftAnti, rewritten from the NOT IN subquery, is pushed down below the (T1 LOJ T2) operation. Because LeftAnti predicate is evaluated to unknown (or null) if any argument of a comparison operator in the predicate is null, e.g. NULL = is evaluated to unknown (which is equivalent to false in the context of a predicate), the LeftAnti predicate cannot be pushed down into a LOJ operation.\n\n{noformat}\nscala> sql(\"select * from (select t2.c2+1 as c3 from t1 left join t2 on t1.c1=t2.c2) t3 where c3 not in (select c2 from t2)\").explain(true)\n== Parsed Logical Plan ==\n'Project [*]\n+- 'Filter NOT 'c3 IN (list#124)\n : +- 'SubqueryAlias list#124\n : +- 'Project ['c2]\n : +- 'UnresolvedRelation `t2`\n +- 'SubqueryAlias t3\n +- 'Project [('t2.c2 + 1) AS c3#123]\n +- 'Join LeftOuter, ('t1.c1 = 't2.c2)\n :- 'UnresolvedRelation `t1`\n +- 'UnresolvedRelation `t2`\n\n== Analyzed Logical Plan ==\nc3: int\nProject [c3#123]\n+- Filter NOT predicate-subquery#124 [(c3#123 = c2#77)]\n : +- SubqueryAlias predicate-subquery#124 [(c3#123 = c2#77)]\n : +- Project [c2#77]\n : +- SubqueryAlias t2\n : +- Project [value#75 AS c2#77]\n : +- LocalRelation [value#75]\n +- SubqueryAlias t3\n +- Project [(c2#77 + 1) AS c3#123]\n +- Join LeftOuter, (c1#3 = c2#77)\n :- SubqueryAlias t1\n : +- Project [value#1 AS c1#3]\n : +- LocalRelation [value#1]\n +- SubqueryAlias t2\n +- Project [value#75 AS c2#77]\n +- LocalRelation [value#75]\n\n== Optimized Logical Plan ==\nProject [(c2#77 + 1) AS c3#123]\n+- Join LeftOuter, (c1#3 = c2#77)\n :- Project [value#1 AS c1#3]\n : +- Join LeftAnti, (isnull(((c2#77 + 1) = c2#77)) || ((c2#77 + 1) = c2#77))\n : :- LocalRelation [value#1]\n : +- LocalRelation [c2#77]\n +- LocalRelation [c2#77]\n{noformat}","created":"2016-08-31T14:42:04.223+0000"},{"body":"@hvanhovell has observed that by disabling the rule {{PushPredicateThroughJoin}} will solve this problem. The predicate is pushed down because the code thinks all the columns in the LeftAnti predicate reference to the T2, the right table of (T1 LOJ T2). The fact is some of the C2#77 in the LeftAnti predicate are from the T2 in the NOT IN subquery.\n\nHowever, {{PushPredicateThroughJoin}} is not the root cause of the problem. The problem is the identifier {{[value#75 AS c2#77]}} is used in two places from the two unrelated references of {{SubqueryAlias t2}}. This can be demonstrated by replacing the T2 in the subquery by a different name, SQ in the example below. We will not see the LeftAnti predicate pushed down below the LOJ.\n\n{noformat}\nSeq(1).toDF(\"cx\").createOrReplaceTempView(\"sq\")\n\nscala> sql(\"select * from (select t2.c2+1 as c3 from t1 left join t2 on t1.c1=t2.c2) t3 where c3 not in (select cx from sq)\").explain(true)\n\n...\n\n== Optimized Logical Plan ==\nProject [(c2#77 + 1) AS c3#137]\n+- Join LeftAnti, (isnull(((c2#77 + 1) = cx#133)) || ((c2#77 + 1) = cx#133))\n :- Join LeftOuter, (c1#3 = c2#77)\n : :- LocalRelation [c1#3]\n : +- LocalRelation [c2#77]\n +- LocalRelation [cx#133]\n...\n{noformat}","created":"2016-08-31T14:43:18.168+0000"},{"body":"User 'nsyca' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14899","created":"2016-08-31T15:52:08.221+0000"},{"body":"The same problem surfaced in different symptoms was discussed in SPARK-13801, SPARK-14040, and SPARK-17154. The problem reported here is a specific pattern. We shall find a solution that addresses the root cause. I am considering closing this JIRA as a duplicate.","created":"2016-09-08T16:01:55.066+0000"},{"body":"User 'hvanhovell' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15761","created":"2016-11-03T23:51:06.371+0000"},{"body":"As commented in the PR, the code was nicely done but the bigger problem remains there, tracked by SPARK-17154. I made a few attempts to fix the root cause of the problem surfaced in many symptoms but it has taken me long and not yet as clean as I want. Also I am seeing [~kousuke] keeps refining his work through a series of PRs for that so I hesitate to try competing with his solution.","created":"2016-11-04T00:13:38.546+0000"},{"body":"[~nsyca] You are right to say that this is part of a larger set of problems which we should address and which given the amount of attempts made to fix this proves that it it non-trivial. It is not a problem to have different (competing) PRs for this problem, this usually leads to solutions. So please submit a PR is feel that this is a better solution.\n\nHowever, for this ticket I just want to solve the correctness issue at hand.","created":"2016-11-04T17:56:49.364+0000"},{"body":"Totally agreed on your approach. We should close off any potential incorrect results at the soonest possible.\n\nThanks for the advise on the approach of different/competing PRs. I am relatively new to the community and try to be careful not to break any etiquette of the community.","created":"2016-11-04T18:01:48.353+0000"},{"body":"User 'hvanhovell' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15772","created":"2016-11-04T20:36:05.965+0000"}],"conversations":[{"body":"While investigating SPARK-16951, I found an incorrect results case from a NOT IN subquery. I thought originally it is an edge case. Further investigation found this is a more general problem.","from":"reporter","subject":"Incomplete algorithm for name resolution in Catalyst paser may lead to incorrect result"},{"body":"This problem originated from the Case 1 documented in SPARK-16951. The script illustrates the problem\n\n{noformat}\nSeq(1,2).toDF(\"c1\").createOrReplaceTempView(\"t1\")\nSeq(1).toDF(\"c2\").createOrReplaceTempView(\"t2\")\n\nscala> sql(\"select * from (select t2.c2+1 as c3 from t1 left join t2 on t1.c1=t2.c2) t3 where c3 not in (select c2 from t2)\").show\n+----+\n| c3|\n+----+\n| 2|\n|null|\n+----+\n{noformat}\n\nThe correct answer is 1 row of (2). From the plan below, the incorrect portion of the plan is the LeftAnti, rewritten from the NOT IN subquery, is pushed down below the (T1 LOJ T2) operation. Because LeftAnti predicate is evaluated to unknown (or null) if any argument of a comparison operator in the predicate is null, e.g. NULL = is evaluated to unknown (which is equivalent to false in the context of a predicate), the LeftAnti predicate cannot be pushed down into a LOJ operation.\n\n{noformat}\nscala> sql(\"select * from (select t2.c2+1 as c3 from t1 left join t2 on t1.c1=t2.c2) t3 where c3 not in (select c2 from t2)\").explain(true)\n== Parsed Logical Plan ==\n'Project [*]\n+- 'Filter NOT 'c3 IN (list#124)\n : +- 'SubqueryAlias list#124\n : +- 'Project ['c2]\n : +- 'UnresolvedRelation `t2`\n +- 'SubqueryAlias t3\n +- 'Project [('t2.c2 + 1) AS c3#123]\n +- 'Join LeftOuter, ('t1.c1 = 't2.c2)\n :- 'UnresolvedRelation `t1`\n +- 'UnresolvedRelation `t2`\n\n== Analyzed Logical Plan ==\nc3: int\nProject [c3#123]\n+- Filter NOT predicate-subquery#124 [(c3#123 = c2#77)]\n : +- SubqueryAlias predicate-subquery#124 [(c3#123 = c2#77)]\n : +- Project [c2#77]\n : +- SubqueryAlias t2\n : +- Project [value#75 AS c2#77]\n : +- LocalRelation [value#75]\n +- SubqueryAlias t3\n +- Project [(c2#77 + 1) AS c3#123]\n +- Join LeftOuter, (c1#3 = c2#77)\n :- SubqueryAlias t1\n : +- Project [value#1 AS c1#3]\n : +- LocalRelation [value#1]\n +- SubqueryAlias t2\n +- Project [value#75 AS c2#77]\n +- LocalRelation [value#75]\n\n== Optimized Logical Plan ==\nProject [(c2#77 + 1) AS c3#123]\n+- Join LeftOuter, (c1#3 = c2#77)\n :- Project [value#1 AS c1#3]\n : +- Join LeftAnti, (isnull(((c2#77 + 1) = c2#77)) || ((c2#77 + 1) = c2#77))\n : :- LocalRelation [value#1]\n : +- LocalRelation [c2#77]\n +- LocalRelation [c2#77]\n{noformat}","from":"developer"},{"body":"@hvanhovell has observed that by disabling the rule {{PushPredicateThroughJoin}} will solve this problem. The predicate is pushed down because the code thinks all the columns in the LeftAnti predicate reference to the T2, the right table of (T1 LOJ T2). The fact is some of the C2#77 in the LeftAnti predicate are from the T2 in the NOT IN subquery.\n\nHowever, {{PushPredicateThroughJoin}} is not the root cause of the problem. The problem is the identifier {{[value#75 AS c2#77]}} is used in two places from the two unrelated references of {{SubqueryAlias t2}}. This can be demonstrated by replacing the T2 in the subquery by a different name, SQ in the example below. We will not see the LeftAnti predicate pushed down below the LOJ.\n\n{noformat}\nSeq(1).toDF(\"cx\").createOrReplaceTempView(\"sq\")\n\nscala> sql(\"select * from (select t2.c2+1 as c3 from t1 left join t2 on t1.c1=t2.c2) t3 where c3 not in (select cx from sq)\").explain(true)\n\n...\n\n== Optimized Logical Plan ==\nProject [(c2#77 + 1) AS c3#137]\n+- Join LeftAnti, (isnull(((c2#77 + 1) = cx#133)) || ((c2#77 + 1) = cx#133))\n :- Join LeftOuter, (c1#3 = c2#77)\n : :- LocalRelation [c1#3]\n : +- LocalRelation [c2#77]\n +- LocalRelation [cx#133]\n...\n{noformat}","from":"developer"},{"body":"User 'nsyca' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/14899","from":"developer"},{"body":"The same problem surfaced in different symptoms was discussed in SPARK-13801, SPARK-14040, and SPARK-17154. The problem reported here is a specific pattern. We shall find a solution that addresses the root cause. I am considering closing this JIRA as a duplicate.","from":"developer"},{"body":"User 'hvanhovell' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15761","from":"developer"},{"body":"As commented in the PR, the code was nicely done but the bigger problem remains there, tracked by SPARK-17154. I made a few attempts to fix the root cause of the problem surfaced in many symptoms but it has taken me long and not yet as clean as I want. Also I am seeing [~kousuke] keeps refining his work through a series of PRs for that so I hesitate to try competing with his solution.","from":"developer"},{"body":"[~nsyca] You are right to say that this is part of a larger set of problems which we should address and which given the amount of attempts made to fix this proves that it it non-trivial. It is not a problem to have different (competing) PRs for this problem, this usually leads to solutions. So please submit a PR is feel that this is a better solution.\n\nHowever, for this ticket I just want to solve the correctness issue at hand.","from":"developer"},{"body":"Totally agreed on your approach. We should close off any potential incorrect results at the soonest possible.\n\nThanks for the advise on the approach of different/competing PRs. I am relatively new to the community and try to be careful not to break any etiquette of the community.","from":"developer"},{"body":"User 'hvanhovell' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15772","from":"developer"}],"created":"2016-08-31T14:40:15.000+0000","description":"While investigating SPARK-16951, I found an incorrect results case from a NOT IN subquery. I thought originally it is an edge case. Further investigation found this is a more general problem.","issue_id":"13001680","key":"SPARK-17337","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-11-04T20:36:52.000+0000","role":"fixed_distractor","summary":"Incomplete algorithm for name resolution in Catalyst paser may lead to incorrect result"} {"case_id":"13002813","cluster":"DISTRACTOR-SPARK-17405","comments":[{"body":"[~joshrosen] Thanks for reporting. I haven't been able to reproduce this because of a catalyst bug I have now `Error: org.apache.spark.sql.catalyst.analysis.UnresolvedException: Invalid call to foldable on unresolved object, tree: 'TIMESTAMP(2012-10-19 00:00:00.0) (state=,code=0)`\nI will look more into this.\nHow much memory is configured for this specific test? One thing is that we added memory management through MemoryConsumer for the generated hashmap, so it correctly accounts that part of memory usage and is more likely to throw OOM.","created":"2016-09-06T07:15:02.381+0000"},{"body":"[~qifan], I believe that you may be able to work around the UnresolvedException by checkout out Spark as of your commit (03d77af9ec4ce9a42affd6ab4381ae5bd3c79a5a) rather than using the current master.\n\nI'm running this query through the Spark ThriftServer, started using the script in {{sbin}}, with default settings on my Macbook Pro. Perhaps the default resource requirements are too high for the amount of {{local\\[*]|| task parallelism? In either case, I think we need to fix this so that the out of the box experience works correctly.","created":"2016-09-06T16:15:08.565+0000"},{"body":"On the Spark Dev list, [~jlaskowski] found a simpler example which triggers this issue:\n\n{quote}\n{code}\nscala> val intsMM = 1 to math.pow(10, 3).toInt\nintsMM: scala.collection.immutable.Range.Inclusive = Range(1, 2, 3, 4,\n5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23,\n24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40,\n41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57,\n58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74,\n75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91,\n92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106,\n107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120,\n121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134,\n135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148,\n149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162,\n163, 164, 165, 166, 167, 168, 169, 1...\nscala> val df = intsMM.toDF(\"n\").withColumn(\"m\", 'n % 2)\ndf: org.apache.spark.sql.DataFrame = [n: int, m: int]\n\nscala> df.groupBy('m).agg(sum('n)).show\n...\n16/09/06 22:28:02 ERROR Executor: Exception in task 6.0 in stage 0.0 (TID 6)\njava.lang.OutOfMemoryError: Unable to acquire 262144 bytes of memory, got 0\n...\n{code}\n\nPlease see https://gist.github.com/jaceklaskowski/906d62b830f6c967a7eee5f8eb6e9237\n{quote}","created":"2016-09-06T21:52:46.130+0000"},{"body":"It definitely got better with the build today Sept, 7th. Yesterday, even such a simple query died {{Seq(1).toDF.groupBy('value).count.show}}.","created":"2016-09-07T04:23:30.375+0000"},{"body":"[~joshrosen][~jlaskowski]Thanks for the comments and suggestions. I have run both of your queries on 03d77af9ec4ce9a42affd6ab4381ae5bd3c79a5a and was able to finish both of them without any exceptions. \nI'll do some static code analysis based on the log from [~jlaskowski]","created":"2016-09-07T20:00:22.914+0000"},{"body":"My hunch is that this is affected by the default number of cores in local mode: I think that my MBP uses 16 tasks by default, while I think that the default parallelism is lower in Jenkins (and perhaps on your machine). If you have trouble reproducing this issue then I'd try explicitly running {{local\\[16]}} or {{local\\[32]}} to see if that can reproduce the issue.","created":"2016-09-07T20:24:49.230+0000"},{"body":"[~joshrosen]\nYes likely. The new hashmap asks for 64MB per task, and the default single-node setting uses only hundreds of memory in total.\nWe decided on 64MB due to our single memory page design for simplicity and performance, and that in production it should hold 64MB * cores << memory_capacity.\nMaybe we should increase default memory a bit? Or is it bad in general to have such upfront cost of 64MB?","created":"2016-09-07T20:47:14.684+0000"},{"body":"One quick fix is to set memory capacity in configuration to make sure memory_capacity > x*cores (x being some number > 64MB)","created":"2016-09-07T22:39:09.087+0000"},{"body":"[~joshrosen] Yes, running local[32] will reproduce the exception. ","created":"2016-09-07T22:48:06.338+0000"},{"body":"User 'ericl' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15016","created":"2016-09-08T21:16:09.279+0000"},{"body":"Issue resolved by pull request 15016\n[https://github.com/apache/spark/pull/15016]","created":"2016-09-08T23:48:08.592+0000"}],"conversations":[{"body":"Prior to SPARK-16525 / https://github.com/apache/spark/pull/14176, the following query ran fine via Beeline / Thrift Server and the Spark shell, but after that patch it is consistently OOMING:\n\n{code}\nCREATE TEMPORARY VIEW table_1(double_col_1, boolean_col_2, timestamp_col_3, smallint_col_4, boolean_col_5, int_col_6, timestamp_col_7, varchar0008_col_8, int_col_9, string_col_10) AS (\n SELECT * FROM (VALUES\n (CAST(-147.818640624 AS DOUBLE), CAST(NULL AS BOOLEAN), TIMESTAMP('2012-10-19 00:00:00.0'), CAST(9 AS SMALLINT), false, 77, TIMESTAMP('2014-07-01 00:00:00.0'), '-945', -646, '722'),\n (CAST(594.195125271 AS DOUBLE), false, TIMESTAMP('2016-12-04 00:00:00.0'), CAST(NULL AS SMALLINT), CAST(NULL AS BOOLEAN), CAST(NULL AS INT), TIMESTAMP('1999-12-26 00:00:00.0'), '250', -861, '55'),\n (CAST(-454.171126363 AS DOUBLE), false, TIMESTAMP('2008-12-13 00:00:00.0'), CAST(NULL AS SMALLINT), false, -783, TIMESTAMP('2010-05-28 00:00:00.0'), '211', -959, CAST(NULL AS STRING)),\n (CAST(437.670945524 AS DOUBLE), true, TIMESTAMP('2011-10-16 00:00:00.0'), CAST(952 AS SMALLINT), true, 297, TIMESTAMP('2013-01-13 00:00:00.0'), '262', CAST(NULL AS INT), '936'),\n (CAST(-387.226759334 AS DOUBLE), false, TIMESTAMP('2019-10-03 00:00:00.0'), CAST(-496 AS SMALLINT), CAST(NULL AS BOOLEAN), -925, TIMESTAMP('2028-06-27 00:00:00.0'), '-657', 948, '18'),\n (CAST(-306.138230875 AS DOUBLE), true, TIMESTAMP('1997-10-07 00:00:00.0'), CAST(332 AS SMALLINT), false, 744, TIMESTAMP('1990-09-22 00:00:00.0'), '-345', 566, '-574'),\n (CAST(675.402140308 AS DOUBLE), false, TIMESTAMP('2017-06-26 00:00:00.0'), CAST(972 AS SMALLINT), true, CAST(NULL AS INT), TIMESTAMP('2026-06-10 00:00:00.0'), '518', 683, '-320'),\n (CAST(734.839647174 AS DOUBLE), true, TIMESTAMP('1995-06-01 00:00:00.0'), CAST(-792 AS SMALLINT), CAST(NULL AS BOOLEAN), CAST(NULL AS INT), TIMESTAMP('2021-07-11 00:00:00.0'), '-318', 564, '142')\n ) as t);\n\nCREATE TEMPORARY VIEW table_3(string_col_1, float_col_2, timestamp_col_3, boolean_col_4, timestamp_col_5, decimal3317_col_6) AS (\n SELECT * FROM (VALUES\n ('88', CAST(191.92508 AS FLOAT), TIMESTAMP('1990-10-25 00:00:00.0'), false, TIMESTAMP('1992-11-02 00:00:00.0'), CAST(NULL AS DECIMAL(33,17))),\n ('-419', CAST(-13.477915 AS FLOAT), TIMESTAMP('1996-03-02 00:00:00.0'), true, CAST(NULL AS TIMESTAMP), -653.51000000000000000BD),\n ('970', CAST(-360.432 AS FLOAT), TIMESTAMP('2010-07-29 00:00:00.0'), false, TIMESTAMP('1995-09-01 00:00:00.0'), -936.48000000000000000BD),\n ('807', CAST(814.30756 AS FLOAT), TIMESTAMP('2019-11-06 00:00:00.0'), false, TIMESTAMP('1996-04-25 00:00:00.0'), 335.56000000000000000BD),\n ('-872', CAST(616.50525 AS FLOAT), TIMESTAMP('2011-08-28 00:00:00.0'), false, TIMESTAMP('2003-07-19 00:00:00.0'), -951.18000000000000000BD),\n ('-167', CAST(-875.35675 AS FLOAT), TIMESTAMP('1995-07-14 00:00:00.0'), false, TIMESTAMP('2005-11-29 00:00:00.0'), 224.89000000000000000BD)\n ) as t);\n\nSELECT\nCAST(MIN(t2.smallint_col_4) AS STRING) AS char_col,\nLEAD(MAX((-387) + (727.64)), 90) OVER (PARTITION BY COALESCE(t2.int_col_9, t2.smallint_col_4, t2.int_col_9) ORDER BY COALESCE(t2.int_col_9, t2.smallint_col_4, t2.int_col_9) DESC, CAST(MIN(t2.smallint_col_4) AS STRING)) AS decimal_col,\nCOALESCE(t2.int_col_9, t2.smallint_col_4, t2.int_col_9) AS int_col\nFROM table_3 t1\nINNER JOIN table_1 t2 ON (((t2.timestamp_col_3) = (t1.timestamp_col_5)) AND ((t2.string_col_10) = (t1.string_col_1))) AND ((t2.string_col_10) = (t1.string_col_1))\nWHERE\n(t2.smallint_col_4) IN (t2.int_col_9, t2.int_col_9)\nGROUP BY\nCOALESCE(t2.int_col_9, t2.smallint_col_4, t2.int_col_9);\n{code}\n\nHere's the OOM:\n\n{code}\norg.apache.hive.service.cli.HiveSQLException: org.apache.spark.SparkException: Job aborted due to stage failure: Task 1 in stage 1.0 failed 1 times, most recent failure: Lost task 1.0 in stage 1.0 (TID 9, localhost): java.lang.OutOfMemoryError: Unable to acquire 262144 bytes of memory, got 0\n at org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:100)\n at org.apache.spark.unsafe.map.BytesToBytesMap.allocate(BytesToBytesMap.java:783)\n at org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:204)\n at org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:219)\n at org.apache.spark.sql.execution.UnsafeFixedWidthAggregationMap.(UnsafeFixedWidthAggregationMap.java:104)\n at org.apache.spark.sql.execution.aggregate.HashAggregateExec.createHashMap(HashAggregateExec.scala:305)\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.agg_doAggregateWithKeys$(Unknown Source)\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\n at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n at org.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java:126)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\n at org.apache.spark.scheduler.Task.run(Task.scala:86)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}","from":"reporter","subject":"Simple aggregation query OOMing after SPARK-16525"},{"body":"[~joshrosen] Thanks for reporting. I haven't been able to reproduce this because of a catalyst bug I have now `Error: org.apache.spark.sql.catalyst.analysis.UnresolvedException: Invalid call to foldable on unresolved object, tree: 'TIMESTAMP(2012-10-19 00:00:00.0) (state=,code=0)`\nI will look more into this.\nHow much memory is configured for this specific test? One thing is that we added memory management through MemoryConsumer for the generated hashmap, so it correctly accounts that part of memory usage and is more likely to throw OOM.","from":"developer"},{"body":"[~qifan], I believe that you may be able to work around the UnresolvedException by checkout out Spark as of your commit (03d77af9ec4ce9a42affd6ab4381ae5bd3c79a5a) rather than using the current master.\n\nI'm running this query through the Spark ThriftServer, started using the script in {{sbin}}, with default settings on my Macbook Pro. Perhaps the default resource requirements are too high for the amount of {{local\\[*]|| task parallelism? In either case, I think we need to fix this so that the out of the box experience works correctly.","from":"developer"},{"body":"On the Spark Dev list, [~jlaskowski] found a simpler example which triggers this issue:\n\n{quote}\n{code}\nscala> val intsMM = 1 to math.pow(10, 3).toInt\nintsMM: scala.collection.immutable.Range.Inclusive = Range(1, 2, 3, 4,\n5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23,\n24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40,\n41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57,\n58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74,\n75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91,\n92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106,\n107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120,\n121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134,\n135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148,\n149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162,\n163, 164, 165, 166, 167, 168, 169, 1...\nscala> val df = intsMM.toDF(\"n\").withColumn(\"m\", 'n % 2)\ndf: org.apache.spark.sql.DataFrame = [n: int, m: int]\n\nscala> df.groupBy('m).agg(sum('n)).show\n...\n16/09/06 22:28:02 ERROR Executor: Exception in task 6.0 in stage 0.0 (TID 6)\njava.lang.OutOfMemoryError: Unable to acquire 262144 bytes of memory, got 0\n...\n{code}\n\nPlease see https://gist.github.com/jaceklaskowski/906d62b830f6c967a7eee5f8eb6e9237\n{quote}","from":"developer"},{"body":"It definitely got better with the build today Sept, 7th. Yesterday, even such a simple query died {{Seq(1).toDF.groupBy('value).count.show}}.","from":"developer"},{"body":"[~joshrosen][~jlaskowski]Thanks for the comments and suggestions. I have run both of your queries on 03d77af9ec4ce9a42affd6ab4381ae5bd3c79a5a and was able to finish both of them without any exceptions. \nI'll do some static code analysis based on the log from [~jlaskowski]","from":"developer"},{"body":"My hunch is that this is affected by the default number of cores in local mode: I think that my MBP uses 16 tasks by default, while I think that the default parallelism is lower in Jenkins (and perhaps on your machine). If you have trouble reproducing this issue then I'd try explicitly running {{local\\[16]}} or {{local\\[32]}} to see if that can reproduce the issue.","from":"developer"},{"body":"[~joshrosen]\nYes likely. The new hashmap asks for 64MB per task, and the default single-node setting uses only hundreds of memory in total.\nWe decided on 64MB due to our single memory page design for simplicity and performance, and that in production it should hold 64MB * cores << memory_capacity.\nMaybe we should increase default memory a bit? Or is it bad in general to have such upfront cost of 64MB?","from":"developer"},{"body":"One quick fix is to set memory capacity in configuration to make sure memory_capacity > x*cores (x being some number > 64MB)","from":"developer"},{"body":"[~joshrosen] Yes, running local[32] will reproduce the exception. ","from":"developer"},{"body":"User 'ericl' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15016","from":"developer"},{"body":"Issue resolved by pull request 15016\n[https://github.com/apache/spark/pull/15016]","from":"developer"}],"created":"2016-09-05T23:42:11.000+0000","description":"Prior to SPARK-16525 / https://github.com/apache/spark/pull/14176, the following query ran fine via Beeline / Thrift Server and the Spark shell, but after that patch it is consistently OOMING:\n\n{code}\nCREATE TEMPORARY VIEW table_1(double_col_1, boolean_col_2, timestamp_col_3, smallint_col_4, boolean_col_5, int_col_6, timestamp_col_7, varchar0008_col_8, int_col_9, string_col_10) AS (\n SELECT * FROM (VALUES\n (CAST(-147.818640624 AS DOUBLE), CAST(NULL AS BOOLEAN), TIMESTAMP('2012-10-19 00:00:00.0'), CAST(9 AS SMALLINT), false, 77, TIMESTAMP('2014-07-01 00:00:00.0'), '-945', -646, '722'),\n (CAST(594.195125271 AS DOUBLE), false, TIMESTAMP('2016-12-04 00:00:00.0'), CAST(NULL AS SMALLINT), CAST(NULL AS BOOLEAN), CAST(NULL AS INT), TIMESTAMP('1999-12-26 00:00:00.0'), '250', -861, '55'),\n (CAST(-454.171126363 AS DOUBLE), false, TIMESTAMP('2008-12-13 00:00:00.0'), CAST(NULL AS SMALLINT), false, -783, TIMESTAMP('2010-05-28 00:00:00.0'), '211', -959, CAST(NULL AS STRING)),\n (CAST(437.670945524 AS DOUBLE), true, TIMESTAMP('2011-10-16 00:00:00.0'), CAST(952 AS SMALLINT), true, 297, TIMESTAMP('2013-01-13 00:00:00.0'), '262', CAST(NULL AS INT), '936'),\n (CAST(-387.226759334 AS DOUBLE), false, TIMESTAMP('2019-10-03 00:00:00.0'), CAST(-496 AS SMALLINT), CAST(NULL AS BOOLEAN), -925, TIMESTAMP('2028-06-27 00:00:00.0'), '-657', 948, '18'),\n (CAST(-306.138230875 AS DOUBLE), true, TIMESTAMP('1997-10-07 00:00:00.0'), CAST(332 AS SMALLINT), false, 744, TIMESTAMP('1990-09-22 00:00:00.0'), '-345', 566, '-574'),\n (CAST(675.402140308 AS DOUBLE), false, TIMESTAMP('2017-06-26 00:00:00.0'), CAST(972 AS SMALLINT), true, CAST(NULL AS INT), TIMESTAMP('2026-06-10 00:00:00.0'), '518', 683, '-320'),\n (CAST(734.839647174 AS DOUBLE), true, TIMESTAMP('1995-06-01 00:00:00.0'), CAST(-792 AS SMALLINT), CAST(NULL AS BOOLEAN), CAST(NULL AS INT), TIMESTAMP('2021-07-11 00:00:00.0'), '-318', 564, '142')\n ) as t);\n\nCREATE TEMPORARY VIEW table_3(string_col_1, float_col_2, timestamp_col_3, boolean_col_4, timestamp_col_5, decimal3317_col_6) AS (\n SELECT * FROM (VALUES\n ('88', CAST(191.92508 AS FLOAT), TIMESTAMP('1990-10-25 00:00:00.0'), false, TIMESTAMP('1992-11-02 00:00:00.0'), CAST(NULL AS DECIMAL(33,17))),\n ('-419', CAST(-13.477915 AS FLOAT), TIMESTAMP('1996-03-02 00:00:00.0'), true, CAST(NULL AS TIMESTAMP), -653.51000000000000000BD),\n ('970', CAST(-360.432 AS FLOAT), TIMESTAMP('2010-07-29 00:00:00.0'), false, TIMESTAMP('1995-09-01 00:00:00.0'), -936.48000000000000000BD),\n ('807', CAST(814.30756 AS FLOAT), TIMESTAMP('2019-11-06 00:00:00.0'), false, TIMESTAMP('1996-04-25 00:00:00.0'), 335.56000000000000000BD),\n ('-872', CAST(616.50525 AS FLOAT), TIMESTAMP('2011-08-28 00:00:00.0'), false, TIMESTAMP('2003-07-19 00:00:00.0'), -951.18000000000000000BD),\n ('-167', CAST(-875.35675 AS FLOAT), TIMESTAMP('1995-07-14 00:00:00.0'), false, TIMESTAMP('2005-11-29 00:00:00.0'), 224.89000000000000000BD)\n ) as t);\n\nSELECT\nCAST(MIN(t2.smallint_col_4) AS STRING) AS char_col,\nLEAD(MAX((-387) + (727.64)), 90) OVER (PARTITION BY COALESCE(t2.int_col_9, t2.smallint_col_4, t2.int_col_9) ORDER BY COALESCE(t2.int_col_9, t2.smallint_col_4, t2.int_col_9) DESC, CAST(MIN(t2.smallint_col_4) AS STRING)) AS decimal_col,\nCOALESCE(t2.int_col_9, t2.smallint_col_4, t2.int_col_9) AS int_col\nFROM table_3 t1\nINNER JOIN table_1 t2 ON (((t2.timestamp_col_3) = (t1.timestamp_col_5)) AND ((t2.string_col_10) = (t1.string_col_1))) AND ((t2.string_col_10) = (t1.string_col_1))\nWHERE\n(t2.smallint_col_4) IN (t2.int_col_9, t2.int_col_9)\nGROUP BY\nCOALESCE(t2.int_col_9, t2.smallint_col_4, t2.int_col_9);\n{code}\n\nHere's the OOM:\n\n{code}\norg.apache.hive.service.cli.HiveSQLException: org.apache.spark.SparkException: Job aborted due to stage failure: Task 1 in stage 1.0 failed 1 times, most recent failure: Lost task 1.0 in stage 1.0 (TID 9, localhost): java.lang.OutOfMemoryError: Unable to acquire 262144 bytes of memory, got 0\n at org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:100)\n at org.apache.spark.unsafe.map.BytesToBytesMap.allocate(BytesToBytesMap.java:783)\n at org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:204)\n at org.apache.spark.unsafe.map.BytesToBytesMap.(BytesToBytesMap.java:219)\n at org.apache.spark.sql.execution.UnsafeFixedWidthAggregationMap.(UnsafeFixedWidthAggregationMap.java:104)\n at org.apache.spark.sql.execution.aggregate.HashAggregateExec.createHashMap(HashAggregateExec.scala:305)\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.agg_doAggregateWithKeys$(Unknown Source)\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\n at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n at org.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java:126)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\n at org.apache.spark.scheduler.Task.run(Task.scala:86)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}","issue_id":"13002813","key":"SPARK-17405","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-09-08T23:48:07.000+0000","role":"fixed_distractor","summary":"Simple aggregation query OOMing after SPARK-16525"} {"case_id":"13003767","cluster":"DISTRACTOR-SPARK-17465","comments":[{"body":"User 'saturday-shi' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15022","created":"2016-09-09T08:36:09.994+0000"},{"body":"It seems that the issue has some negative side effects on processing time.\n\nWhen I ran a streaming application using _updateStateByKey_ on Spark 1.6.0, it showed a strange gradual increasing of processing time: the processing time of a 2s-interval-batch would increase from < 0.1s to 2 ~ 3s after 10 ~ 20 days. (similar to [this issue|https://issues.apache.org/jira/browse/SPARK-13288])\n\nBut after I fixed the issue described in this JIRA, the problem seems disappeared. The processing time stably keeping at < 0.1s, and shows no tendency to increase.\n\nI have no idea of why it happens, but the issue - described in this JIRA - seems to have an actual effect on processing time of applications that using _updateStateByKey_ (or checkpoint).\n","created":"2016-09-13T08:43:23.002+0000"},{"body":"Issue resolved by pull request 15022\n[https://github.com/apache/spark/pull/15022]","created":"2016-09-14T20:47:56.042+0000"},{"body":"Resolved.\n\nIn every task, method _currentUnrollMemory_ will be called several times. It will scan all keys of _unrollMemoryMap_ and _pendingUnrollMemoryMap_, so the processing time is proportional to the map size.\nhttps://github.com/apache/spark/blob/v1.6.0/core/src/main/scala/org/apache/spark/storage/MemoryStore.scala#L540-L542\n\nI have checked the processing time of _currentUnrollMemory_. It just equals to the time increased from before.\n\nHope this will help someone who has a similar issue of increasing processing time when upgrade Spark to 1.6.0 :)","created":"2016-10-03T09:16:17.947+0000"},{"body":"[~saturday_s] Thanks for this fix! It really helped us. We were using 1.6.1 and were seeing processing times increase gradually over period of several days. With 1.6.3 this increase does not happen. Thank you!\n","created":"2017-03-15T01:41:05.156+0000"}],"conversations":[{"body":"After updating Spark from 1.5.0 to 1.6.0, I found that it seems to have a memory leak on my Spark streaming application.\n\nHere is the head of the heap histogram of my application, which has been running about 160 hours:\n{code:borderStyle=solid}\n num #instances #bytes class name\n----------------------------------------------\n 1: 28094 71753976 [B\n 2: 1188086 28514064 java.lang.Long\n 3: 1183844 28412256 scala.collection.mutable.DefaultEntry\n 4: 102242 13098768 \n 5: 102242 12421000 \n 6: 8184 9199032 \n 7: 38 8391584 [Lscala.collection.mutable.HashEntry;\n 8: 8184 7514288 \n 9: 6651 4874080 \n 10: 37197 3438040 [C\n 11: 6423 2445640 \n 12: 8773 1044808 java.lang.Class\n 13: 36869 884856 java.lang.String\n 14: 15715 848368 [[I\n 15: 13690 782808 [S\n 16: 18903 604896 java.util.concurrent.ConcurrentHashMap$HashEntry\n 17: 13 426192 [Lscala.concurrent.forkjoin.ForkJoinTask;\n{code}\nIt shows that *scala.collection.mutable.DefaultEntry* and *java.lang.Long* have unexpected big numbers of instances. In fact, the numbers started growing at streaming process began, and keep growing proportional to total number of tasks.\n\n\n\nAfter some further investigation, I found that the problem is caused by some inappropriate memory management in _releaseUnrollMemoryForThisTask_ and _unrollSafely_ method of class [org.apache.spark.storage.MemoryStore|https://github.com/apache/spark/blob/branch-1.6/core/src/main/scala/org/apache/spark/storage/MemoryStore.scala].\n\nIn Spark 1.6.x, a _releaseUnrollMemoryForThisTask_ operation will be processed only with the parameter _memoryToRelease_ > 0:\nhttps://github.com/apache/spark/blob/branch-1.6/core/src/main/scala/org/apache/spark/storage/MemoryStore.scala#L530-L537\nBut in fact, if a task successfully unrolled all its blocks in memory by _unrollSafely_ method, the memory saved in _unrollMemoryMap_ would be set to zero:\nhttps://github.com/apache/spark/blob/branch-1.6/core/src/main/scala/org/apache/spark/storage/MemoryStore.scala#L322\n\nSo the result is, the memory saved in _unrollMemoryMap_ will be released, but the key of that part of memory will never be removed from the hash map. The hash table will keep increasing, while new tasks keep incoming. Although the speed of increase is comparatively slow (about dozens of bytes per task), it is possible that result into OOM after weeks or months.","from":"reporter","subject":"Inappropriate memory management in `org.apache.spark.storage.MemoryStore` may lead to memory leak"},{"body":"User 'saturday-shi' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15022","from":"developer"},{"body":"It seems that the issue has some negative side effects on processing time.\n\nWhen I ran a streaming application using _updateStateByKey_ on Spark 1.6.0, it showed a strange gradual increasing of processing time: the processing time of a 2s-interval-batch would increase from < 0.1s to 2 ~ 3s after 10 ~ 20 days. (similar to [this issue|https://issues.apache.org/jira/browse/SPARK-13288])\n\nBut after I fixed the issue described in this JIRA, the problem seems disappeared. The processing time stably keeping at < 0.1s, and shows no tendency to increase.\n\nI have no idea of why it happens, but the issue - described in this JIRA - seems to have an actual effect on processing time of applications that using _updateStateByKey_ (or checkpoint).\n","from":"developer"},{"body":"Issue resolved by pull request 15022\n[https://github.com/apache/spark/pull/15022]","from":"developer"},{"body":"Resolved.\n\nIn every task, method _currentUnrollMemory_ will be called several times. It will scan all keys of _unrollMemoryMap_ and _pendingUnrollMemoryMap_, so the processing time is proportional to the map size.\nhttps://github.com/apache/spark/blob/v1.6.0/core/src/main/scala/org/apache/spark/storage/MemoryStore.scala#L540-L542\n\nI have checked the processing time of _currentUnrollMemory_. It just equals to the time increased from before.\n\nHope this will help someone who has a similar issue of increasing processing time when upgrade Spark to 1.6.0 :)","from":"developer"},{"body":"[~saturday_s] Thanks for this fix! It really helped us. We were using 1.6.1 and were seeing processing times increase gradually over period of several days. With 1.6.3 this increase does not happen. Thank you!\n","from":"developer"}],"created":"2016-09-09T06:35:19.000+0000","description":"After updating Spark from 1.5.0 to 1.6.0, I found that it seems to have a memory leak on my Spark streaming application.\n\nHere is the head of the heap histogram of my application, which has been running about 160 hours:\n{code:borderStyle=solid}\n num #instances #bytes class name\n----------------------------------------------\n 1: 28094 71753976 [B\n 2: 1188086 28514064 java.lang.Long\n 3: 1183844 28412256 scala.collection.mutable.DefaultEntry\n 4: 102242 13098768 \n 5: 102242 12421000 \n 6: 8184 9199032 \n 7: 38 8391584 [Lscala.collection.mutable.HashEntry;\n 8: 8184 7514288 \n 9: 6651 4874080 \n 10: 37197 3438040 [C\n 11: 6423 2445640 \n 12: 8773 1044808 java.lang.Class\n 13: 36869 884856 java.lang.String\n 14: 15715 848368 [[I\n 15: 13690 782808 [S\n 16: 18903 604896 java.util.concurrent.ConcurrentHashMap$HashEntry\n 17: 13 426192 [Lscala.concurrent.forkjoin.ForkJoinTask;\n{code}\nIt shows that *scala.collection.mutable.DefaultEntry* and *java.lang.Long* have unexpected big numbers of instances. In fact, the numbers started growing at streaming process began, and keep growing proportional to total number of tasks.\n\n\n\nAfter some further investigation, I found that the problem is caused by some inappropriate memory management in _releaseUnrollMemoryForThisTask_ and _unrollSafely_ method of class [org.apache.spark.storage.MemoryStore|https://github.com/apache/spark/blob/branch-1.6/core/src/main/scala/org/apache/spark/storage/MemoryStore.scala].\n\nIn Spark 1.6.x, a _releaseUnrollMemoryForThisTask_ operation will be processed only with the parameter _memoryToRelease_ > 0:\nhttps://github.com/apache/spark/blob/branch-1.6/core/src/main/scala/org/apache/spark/storage/MemoryStore.scala#L530-L537\nBut in fact, if a task successfully unrolled all its blocks in memory by _unrollSafely_ method, the memory saved in _unrollMemoryMap_ would be set to zero:\nhttps://github.com/apache/spark/blob/branch-1.6/core/src/main/scala/org/apache/spark/storage/MemoryStore.scala#L322\n\nSo the result is, the memory saved in _unrollMemoryMap_ will be released, but the key of that part of memory will never be removed from the hash map. The hash table will keep increasing, while new tasks keep incoming. Although the speed of increase is comparatively slow (about dozens of bytes per task), it is possible that result into OOM after weeks or months.","issue_id":"13003767","key":"SPARK-17465","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-09-14T20:47:55.000+0000","role":"fixed_distractor","summary":"Inappropriate memory management in `org.apache.spark.storage.MemoryStore` may lead to memory leak"} {"case_id":"12713136","cluster":"DISTRACTOR-SPARK-1764","comments":[{"body":"I can semi-reliably recreate this by just running this code:\n\n{quote}\nwhile True:\n sc.parallelize(range(100)).map(lambda n: n * 2).collect()\n{quote}\n\nRunning this on Mesos will eventually crash with \n\nPy4JJavaError: An error occurred while calling o1142.collect.\n: org.apache.spark.SparkException: Job 101 cancelled as part of cancellation of all jobs\n\tat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1033)\n\tat org.apache.spark.scheduler.DAGScheduler.handleJobCancellation(DAGScheduler.scala:998)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply$mcVI$sp(DAGScheduler.scala:499)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply(DAGScheduler.scala:499)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply(DAGScheduler.scala:499)\n\tat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\n\tat org.apache.spark.scheduler.DAGScheduler.doCancelAllJobs(DAGScheduler.scala:499)\n\tat org.apache.spark.scheduler.DAGSchedulerActorSupervisor$$anonfun$2.applyOrElse(DAGScheduler.scala:1151)\n\tat org.apache.spark.scheduler.DAGSchedulerActorSupervisor$$anonfun$2.applyOrElse(DAGScheduler.scala:1147)\n\tat akka.actor.SupervisorStrategy.handleFailure(FaultHandling.scala:295)\n\tat akka.actor.dungeon.FaultHandling$class.handleFailure(FaultHandling.scala:253)\n\tat akka.actor.ActorCell.handleFailure(ActorCell.scala:338)\n\tat akka.actor.ActorCell.invokeAll$1(ActorCell.scala:423)\n\tat akka.actor.ActorCell.systemInvoke(ActorCell.scala:447)\n\tat akka.dispatch.Mailbox.processAllSystemMessages(Mailbox.scala:262)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:218)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n\n\nI0508 18:29:03.623627 7868 sched.cpp:730] Stopping framework '20140508-173240-16842879-5050-24645-0032'\n14/05/08 18:29:04 ERROR OneForOneStrategy: EOF reached before Python server acknowledged\norg.apache.spark.SparkException: EOF reached before Python server acknowledged\n\tat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:416)\n\tat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:387)\n\tat org.apache.spark.Accumulable.$plus$plus$eq(Accumulators.scala:71)\n\tat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:279)\n\tat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:277)\n\tat scala.collection.TraversableLike$WithFilter$$anonfun$foreach$1.apply(TraversableLike.scala:772)\n\tat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n\tat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n\tat scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:226)\n\tat scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:39)\n\tat scala.collection.mutable.HashMap.foreach(HashMap.scala:98)\n\tat scala.collection.TraversableLike$WithFilter.foreach(TraversableLike.scala:771)\n\tat org.apache.spark.Accumulators$.add(Accumulators.scala:277)\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskCompletion(DAGScheduler.scala:818)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor$$anonfun$receive$2.applyOrElse(DAGScheduler.scala:1204)\n\tat akka.actor.ActorCell.receiveMessage(ActorCell.scala:498)\n\tat akka.actor.ActorCell.invoke(ActorCell.scala:456)\n\tat akka.dispatch.Mailbox.processMailbox(Mailbox.scala:237)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:219)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)","created":"2014-05-08T18:29:49.682+0000"},{"body":"I did some more digging into this and I have no idea what's the exact issue. The write to the Python server succeeds (which I checked from the Python side) but the Scala side doesn't seem to be able to read the acknowledgement. \n\nI have also confirmed that it isn't an issue with the Python broadcast server dying, as commented out the exception makes it work fine (!) ","created":"2014-05-12T18:33:22.884+0000"},{"body":"We just ran `sc.parallelize(range(100)).map(lambda n: n * 2).collect()` on a Mesos 0.18.1 cluster with the latest spark and it worked. Could you confirm the Spark & Mesos version you are using (if using master please include the sha/commit hash).","created":"2014-05-12T19:07:10.202+0000"},{"body":"Interesting, including SPARK-1806 in our build made it stop failing... I guess this can be considered fixed then","created":"2014-05-12T21:08:08.036+0000"},{"body":"I have seen this issue in a cut of the main branch take on friday 27th June (last commit https://github.com/srowen/spark/commit/18f29b96c7e0948f5f504e522e5aa8a8d1ab163e )\n\nIt does run successfully for a time (~800 requests) but then fails. This has happened on two different clusters, mesos 0.18.0 and 0.18.2, different sites.\n\nOne interesting side effect is as this is running, the number of file handles in use on the slaves goes up to many thousands. This could be the resource that is running out. One of the slaves (different machines) always dies and the job is eventually aborted.\n\nThis happens with a real app that is actually doing stuff with 18GB RDDs, but also with the simple, but brutal while true command below. A similar loop in scala seems to work fine, indefinitely.\n\nI would be interested to hear if anyone can run this command for a significant amount of time because that would point to our environment in some way.\n\nUsing Python version 2.7.6 (default, Jan 17 2014 10:13:17)\nSparkContext available as sc.\n\nIn [1]: while True: \n sc.parallelize(range(100)).map(lambda n: n * 2).collect()\n ...: \n\nit ran for a while then...\n\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 778 in 63 ms on guavcpt-ch2-a28p.sys.comcast.net (progress: 1/8)\n14/06/27 13:59:08 INFO DAGScheduler: Completed ResultTask(97, 2)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 777 in 70 ms on guavcpt-ch2-a32p.sys.comcast.net (progress: 2/8)\n14/06/27 13:59:08 INFO DAGScheduler: Completed ResultTask(97, 1)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 782 in 104 ms on guavcpt-ch2-a28p.sys.comcast.net (progress: 3/8)\n14/06/27 13:59:08 INFO DAGScheduler: Completed ResultTask(97, 6)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 781 in 105 ms on guavcpt-ch2-a32p.sys.comcast.net (progress: 4/8)\n14/06/27 13:59:08 INFO DAGScheduler: Completed ResultTask(97, 5)\n14/06/27 13:59:08 ERROR DAGSchedulerActorSupervisor: eventProcesserActor failed due to the error EOF reached before Python server acknowledged; shutting down SparkContext\n14/06/27 13:59:08 INFO TaskSchedulerImpl: Cancelling stage 97\n14/06/27 13:59:08 INFO DAGScheduler: Could not cancel tasks for stage 97\njava.lang.UnsupportedOperationException\nat org.apache.spark.scheduler.SchedulerBackend$class.killTask(SchedulerBackend.scala:32)\nat org.apache.spark.scheduler.cluster.mesos.MesosSchedulerBackend.killTask(MesosSchedulerBackend.scala:41)\nat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply$mcVJ$sp(TaskSchedulerImpl.scala:185)\nat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply(TaskSchedulerImpl.scala:183)\nat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply(TaskSchedulerImpl.scala:183)\nat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\nat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3.apply(TaskSchedulerImpl.scala:183)\nat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3.apply(TaskSchedulerImpl.scala:176)\nat scala.Option.foreach(Option.scala:236)\nat org.apache.spark.scheduler.TaskSchedulerImpl.cancelTasks(TaskSchedulerImpl.scala:176)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply$mcVI$sp(DAGScheduler.scala:1066)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply(DAGScheduler.scala:1052)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply(DAGScheduler.scala:1052)\nat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\nat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1052)\nat org.apache.spark.scheduler.DAGScheduler.handleJobCancellation(DAGScheduler.scala:1005)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply$mcVI$sp(DAGScheduler.scala:502)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply(DAGScheduler.scala:502)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply(DAGScheduler.scala:502)\nat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\nat org.apache.spark.scheduler.DAGScheduler.doCancelAllJobs(DAGScheduler.scala:502)\nat org.apache.spark.scheduler.DAGSchedulerActorSupervisor$$anonfun$2.applyOrElse(DAGScheduler.scala:1167)\nat org.apache.spark.scheduler.DAGSchedulerActorSupervisor$$anonfun$2.applyOrElse(DAGScheduler.scala:1162)\nat akka.actor.SupervisorStrategy.handleFailure(FaultHandling.scala:295)\nat akka.actor.dungeon.FaultHandling$class.handleFailure(FaultHandling.scala:253)\nat akka.actor.ActorCell.handleFailure(ActorCell.scala:338)\nat akka.actor.ActorCell.invokeAll$1(ActorCell.scala:423)\nat akka.actor.ActorCell.systemInvoke(ActorCell.scala:447)\nat akka.dispatch.Mailbox.processAllSystemMessages(Mailbox.scala:262)\nat akka.dispatch.Mailbox.run(Mailbox.scala:218)\nat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\nat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\nat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\nat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\nat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 776 in 151 ms on guavcpt-ch2-a31p.sys.comcast.net (progress: 5/8)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 780 in 147 ms on guavcpt-ch2-a31p.sys.comcast.net (progress: 6/8)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 779 in 188 ms on guavcpt-ch2-a27p.sys.comcast.net (progress: 7/8)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 783 in 193 ms on guavcpt-ch2-a27p.sys.comcast.net (progress: 8/8)\n14/06/27 13:59:08 INFO TaskSchedulerImpl: Removed TaskSet 97.0, whose tasks have all completed, from pool \n14/06/27 13:59:09 WARN QueuedThreadPool: 5 threads could not be stopped\n14/06/27 13:59:09 INFO SparkUI: Stopped Spark web UI at http://guavcpt-ch2-a26p.sys.comcast.net:4041\n14/06/27 13:59:09 INFO DAGScheduler: Stopping DAGScheduler\nI0627 13:59:09.129179 14267 sched.cpp:730] Stopping framework '20140605-200137-607132588-5050-19211-0073'\n14/06/27 13:59:09 INFO MesosSchedulerBackend: driver.run() returned with code DRIVER_STOPPED\n14/06/27 13:59:10 INFO MapOutputTrackerMasterActor: MapOutputTrackerActor stopped!\n14/06/27 13:59:10 INFO ConnectionManager: Selector thread was interrupted!\n14/06/27 13:59:10 INFO ConnectionManager: ConnectionManager stopped\n14/06/27 13:59:10 INFO MemoryStore: MemoryStore cleared\n14/06/27 13:59:10 INFO BlockManager: BlockManager stopped\n14/06/27 13:59:10 INFO BlockManagerMasterActor: Stopping BlockManagerMaster\n14/06/27 13:59:10 INFO BlockManagerMaster: BlockManagerMaster stopped\n14/06/27 13:59:10 INFO SparkContext: Successfully stopped SparkContext\n14/06/27 13:59:10 ERROR OneForOneStrategy: EOF reached before Python server acknowledged\norg.apache.spark.SparkException: EOF reached before Python server acknowledged\nat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:613)\nat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:584)\nat org.apache.spark.Accumulable.$plus$plus$eq(Accumulators.scala:72)\nat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:280)\nat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:278)\nat scala.collection.TraversableLike$WithFilter$$anonfun$foreach$1.apply(TraversableLike.scala:772)\nat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\nat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\nat scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:226)\nat scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:39)\nat scala.collection.mutable.HashMap.foreach(HashMap.scala:98)\nat scala.collection.TraversableLike$WithFilter.foreach(TraversableLike.scala:771)\nat org.apache.spark.Accumulators$.add(Accumulators.scala:278)\nat org.apache.spark.scheduler.DAGScheduler.handleTaskCompletion(DAGScheduler.scala:825)\nat org.apache.spark.scheduler.DAGSchedulerEventProcessActor$$anonfun$receive$2.applyOrElse(DAGScheduler.scala:1223)\nat akka.actor.ActorCell.receiveMessage(ActorCell.scala:498)\nat akka.actor.ActorCell.invoke(ActorCell.scala:456)\nat akka.dispatch.Mailbox.processMailbox(Mailbox.scala:237)\nat akka.dispatch.Mailbox.run(Mailbox.scala:219)\nat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\nat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\nat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\nat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\nat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n14/06/27 13:59:10 INFO RemoteActorRefProvider$RemotingTerminator: Shutting down remote daemon.\n14/06/27 13:59:10 INFO RemoteActorRefProvider$RemotingTerminator: Remote daemon shut down; proceeding with flushing remote transports.\n14/06/27 13:59:10 INFO Remoting: Remoting shut down\n\n","created":"2014-07-01T12:27:07.276+0000"},{"body":"I'm not sure how this is related to Mesos, is this reproable using YARN or standalone?","created":"2014-07-17T22:47:03.452+0000"},{"body":"Hi;\nNever used yarn. Doesn't happen on standalone.\n\n","created":"2014-07-18T06:07:07.385+0000"},{"body":"This issue should be fixed in SPARK-2282 [1], I had ran the jobs above against mesos-0.19.1 after more than a hour without problems.\n\n[~therealnb] Could you also verify this?\n\n[1] https://github.com/apache/spark/commit/ef4ff00f87a4e8d38866f163f01741c2673e41da","created":"2014-08-25T19:01:42.988+0000"},{"body":"Hi;\nSadly I moved jobs and I don't have a working Spark environment at the moment (I will be doing some Spark work soon :-). I'll pass this on to the guys that are still there and get them to confirm. \nCheers","created":"2014-08-25T19:15:42.612+0000"},{"body":"FYI the workaround described in SPARK-2282 got me past the same \"EOF reached\" issue.\n\ni.e.\n\necho \"1\" > /proc/sys/net/ipv4/tcp_tw_reuse\necho \"1\" > /proc/sys/net/ipv4/tcp_tw_recycle\n\nthen restart your spark shell (or program).\n\nApparently a fix is slated for Spark 1.1\n\nI am running Spark 1.0.2 under Mesos 0.18.2\n","created":"2014-09-11T23:44:02.112+0000"},{"body":"This is fixed by #2282","created":"2014-09-12T00:01:22.500+0000"}],"conversations":[{"body":"I'm getting \"EOF reached before Python server acknowledged\" while using PySpark on Mesos. The error manifests itself in multiple ways. One is:\n\n{noformat}\n14/05/08 18:10:40 ERROR DAGSchedulerActorSupervisor: eventProcesserActor failed due to the error EOF reached before Python server acknowledged; shutting down SparkContext\n{noformat}\n\nAnd the other has a full stacktrace:\n\n{noformat}\n14/05/08 18:03:06 ERROR OneForOneStrategy: EOF reached before Python server acknowledged\norg.apache.spark.SparkException: EOF reached before Python server acknowledged\n\tat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:416)\n\tat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:387)\n\tat org.apache.spark.Accumulable.$plus$plus$eq(Accumulators.scala:71)\n\tat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:279)\n\tat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:277)\n\tat scala.collection.TraversableLike$WithFilter$$anonfun$foreach$1.apply(TraversableLike.scala:772)\n\tat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n\tat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n\tat scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:226)\n\tat scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:39)\n\tat scala.collection.mutable.HashMap.foreach(HashMap.scala:98)\n\tat scala.collection.TraversableLike$WithFilter.foreach(TraversableLike.scala:771)\n\tat org.apache.spark.Accumulators$.add(Accumulators.scala:277)\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskCompletion(DAGScheduler.scala:818)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor$$anonfun$receive$2.applyOrElse(DAGScheduler.scala:1204)\n\tat akka.actor.ActorCell.receiveMessage(ActorCell.scala:498)\n\tat akka.actor.ActorCell.invoke(ActorCell.scala:456)\n\tat akka.dispatch.Mailbox.processMailbox(Mailbox.scala:237)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:219)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n{noformat}\n\nThis error causes the SparkContext to shutdown. I have not been able to reliably reproduce this bug, it seems to happen randomly, but if you run enough tasks on a SparkContext it'll hapen eventually","from":"reporter","subject":"EOF reached before Python server acknowledged"},{"body":"I can semi-reliably recreate this by just running this code:\n\n{quote}\nwhile True:\n sc.parallelize(range(100)).map(lambda n: n * 2).collect()\n{quote}\n\nRunning this on Mesos will eventually crash with \n\nPy4JJavaError: An error occurred while calling o1142.collect.\n: org.apache.spark.SparkException: Job 101 cancelled as part of cancellation of all jobs\n\tat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1033)\n\tat org.apache.spark.scheduler.DAGScheduler.handleJobCancellation(DAGScheduler.scala:998)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply$mcVI$sp(DAGScheduler.scala:499)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply(DAGScheduler.scala:499)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply(DAGScheduler.scala:499)\n\tat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\n\tat org.apache.spark.scheduler.DAGScheduler.doCancelAllJobs(DAGScheduler.scala:499)\n\tat org.apache.spark.scheduler.DAGSchedulerActorSupervisor$$anonfun$2.applyOrElse(DAGScheduler.scala:1151)\n\tat org.apache.spark.scheduler.DAGSchedulerActorSupervisor$$anonfun$2.applyOrElse(DAGScheduler.scala:1147)\n\tat akka.actor.SupervisorStrategy.handleFailure(FaultHandling.scala:295)\n\tat akka.actor.dungeon.FaultHandling$class.handleFailure(FaultHandling.scala:253)\n\tat akka.actor.ActorCell.handleFailure(ActorCell.scala:338)\n\tat akka.actor.ActorCell.invokeAll$1(ActorCell.scala:423)\n\tat akka.actor.ActorCell.systemInvoke(ActorCell.scala:447)\n\tat akka.dispatch.Mailbox.processAllSystemMessages(Mailbox.scala:262)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:218)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n\n\nI0508 18:29:03.623627 7868 sched.cpp:730] Stopping framework '20140508-173240-16842879-5050-24645-0032'\n14/05/08 18:29:04 ERROR OneForOneStrategy: EOF reached before Python server acknowledged\norg.apache.spark.SparkException: EOF reached before Python server acknowledged\n\tat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:416)\n\tat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:387)\n\tat org.apache.spark.Accumulable.$plus$plus$eq(Accumulators.scala:71)\n\tat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:279)\n\tat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:277)\n\tat scala.collection.TraversableLike$WithFilter$$anonfun$foreach$1.apply(TraversableLike.scala:772)\n\tat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n\tat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n\tat scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:226)\n\tat scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:39)\n\tat scala.collection.mutable.HashMap.foreach(HashMap.scala:98)\n\tat scala.collection.TraversableLike$WithFilter.foreach(TraversableLike.scala:771)\n\tat org.apache.spark.Accumulators$.add(Accumulators.scala:277)\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskCompletion(DAGScheduler.scala:818)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor$$anonfun$receive$2.applyOrElse(DAGScheduler.scala:1204)\n\tat akka.actor.ActorCell.receiveMessage(ActorCell.scala:498)\n\tat akka.actor.ActorCell.invoke(ActorCell.scala:456)\n\tat akka.dispatch.Mailbox.processMailbox(Mailbox.scala:237)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:219)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)","from":"developer"},{"body":"I did some more digging into this and I have no idea what's the exact issue. The write to the Python server succeeds (which I checked from the Python side) but the Scala side doesn't seem to be able to read the acknowledgement. \n\nI have also confirmed that it isn't an issue with the Python broadcast server dying, as commented out the exception makes it work fine (!) ","from":"developer"},{"body":"We just ran `sc.parallelize(range(100)).map(lambda n: n * 2).collect()` on a Mesos 0.18.1 cluster with the latest spark and it worked. Could you confirm the Spark & Mesos version you are using (if using master please include the sha/commit hash).","from":"developer"},{"body":"Interesting, including SPARK-1806 in our build made it stop failing... I guess this can be considered fixed then","from":"developer"},{"body":"I have seen this issue in a cut of the main branch take on friday 27th June (last commit https://github.com/srowen/spark/commit/18f29b96c7e0948f5f504e522e5aa8a8d1ab163e )\n\nIt does run successfully for a time (~800 requests) but then fails. This has happened on two different clusters, mesos 0.18.0 and 0.18.2, different sites.\n\nOne interesting side effect is as this is running, the number of file handles in use on the slaves goes up to many thousands. This could be the resource that is running out. One of the slaves (different machines) always dies and the job is eventually aborted.\n\nThis happens with a real app that is actually doing stuff with 18GB RDDs, but also with the simple, but brutal while true command below. A similar loop in scala seems to work fine, indefinitely.\n\nI would be interested to hear if anyone can run this command for a significant amount of time because that would point to our environment in some way.\n\nUsing Python version 2.7.6 (default, Jan 17 2014 10:13:17)\nSparkContext available as sc.\n\nIn [1]: while True: \n sc.parallelize(range(100)).map(lambda n: n * 2).collect()\n ...: \n\nit ran for a while then...\n\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 778 in 63 ms on guavcpt-ch2-a28p.sys.comcast.net (progress: 1/8)\n14/06/27 13:59:08 INFO DAGScheduler: Completed ResultTask(97, 2)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 777 in 70 ms on guavcpt-ch2-a32p.sys.comcast.net (progress: 2/8)\n14/06/27 13:59:08 INFO DAGScheduler: Completed ResultTask(97, 1)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 782 in 104 ms on guavcpt-ch2-a28p.sys.comcast.net (progress: 3/8)\n14/06/27 13:59:08 INFO DAGScheduler: Completed ResultTask(97, 6)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 781 in 105 ms on guavcpt-ch2-a32p.sys.comcast.net (progress: 4/8)\n14/06/27 13:59:08 INFO DAGScheduler: Completed ResultTask(97, 5)\n14/06/27 13:59:08 ERROR DAGSchedulerActorSupervisor: eventProcesserActor failed due to the error EOF reached before Python server acknowledged; shutting down SparkContext\n14/06/27 13:59:08 INFO TaskSchedulerImpl: Cancelling stage 97\n14/06/27 13:59:08 INFO DAGScheduler: Could not cancel tasks for stage 97\njava.lang.UnsupportedOperationException\nat org.apache.spark.scheduler.SchedulerBackend$class.killTask(SchedulerBackend.scala:32)\nat org.apache.spark.scheduler.cluster.mesos.MesosSchedulerBackend.killTask(MesosSchedulerBackend.scala:41)\nat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply$mcVJ$sp(TaskSchedulerImpl.scala:185)\nat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply(TaskSchedulerImpl.scala:183)\nat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply(TaskSchedulerImpl.scala:183)\nat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\nat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3.apply(TaskSchedulerImpl.scala:183)\nat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3.apply(TaskSchedulerImpl.scala:176)\nat scala.Option.foreach(Option.scala:236)\nat org.apache.spark.scheduler.TaskSchedulerImpl.cancelTasks(TaskSchedulerImpl.scala:176)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply$mcVI$sp(DAGScheduler.scala:1066)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply(DAGScheduler.scala:1052)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply(DAGScheduler.scala:1052)\nat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\nat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1052)\nat org.apache.spark.scheduler.DAGScheduler.handleJobCancellation(DAGScheduler.scala:1005)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply$mcVI$sp(DAGScheduler.scala:502)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply(DAGScheduler.scala:502)\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$doCancelAllJobs$1.apply(DAGScheduler.scala:502)\nat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\nat org.apache.spark.scheduler.DAGScheduler.doCancelAllJobs(DAGScheduler.scala:502)\nat org.apache.spark.scheduler.DAGSchedulerActorSupervisor$$anonfun$2.applyOrElse(DAGScheduler.scala:1167)\nat org.apache.spark.scheduler.DAGSchedulerActorSupervisor$$anonfun$2.applyOrElse(DAGScheduler.scala:1162)\nat akka.actor.SupervisorStrategy.handleFailure(FaultHandling.scala:295)\nat akka.actor.dungeon.FaultHandling$class.handleFailure(FaultHandling.scala:253)\nat akka.actor.ActorCell.handleFailure(ActorCell.scala:338)\nat akka.actor.ActorCell.invokeAll$1(ActorCell.scala:423)\nat akka.actor.ActorCell.systemInvoke(ActorCell.scala:447)\nat akka.dispatch.Mailbox.processAllSystemMessages(Mailbox.scala:262)\nat akka.dispatch.Mailbox.run(Mailbox.scala:218)\nat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\nat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\nat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\nat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\nat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 776 in 151 ms on guavcpt-ch2-a31p.sys.comcast.net (progress: 5/8)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 780 in 147 ms on guavcpt-ch2-a31p.sys.comcast.net (progress: 6/8)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 779 in 188 ms on guavcpt-ch2-a27p.sys.comcast.net (progress: 7/8)\n14/06/27 13:59:08 INFO TaskSetManager: Finished TID 783 in 193 ms on guavcpt-ch2-a27p.sys.comcast.net (progress: 8/8)\n14/06/27 13:59:08 INFO TaskSchedulerImpl: Removed TaskSet 97.0, whose tasks have all completed, from pool \n14/06/27 13:59:09 WARN QueuedThreadPool: 5 threads could not be stopped\n14/06/27 13:59:09 INFO SparkUI: Stopped Spark web UI at http://guavcpt-ch2-a26p.sys.comcast.net:4041\n14/06/27 13:59:09 INFO DAGScheduler: Stopping DAGScheduler\nI0627 13:59:09.129179 14267 sched.cpp:730] Stopping framework '20140605-200137-607132588-5050-19211-0073'\n14/06/27 13:59:09 INFO MesosSchedulerBackend: driver.run() returned with code DRIVER_STOPPED\n14/06/27 13:59:10 INFO MapOutputTrackerMasterActor: MapOutputTrackerActor stopped!\n14/06/27 13:59:10 INFO ConnectionManager: Selector thread was interrupted!\n14/06/27 13:59:10 INFO ConnectionManager: ConnectionManager stopped\n14/06/27 13:59:10 INFO MemoryStore: MemoryStore cleared\n14/06/27 13:59:10 INFO BlockManager: BlockManager stopped\n14/06/27 13:59:10 INFO BlockManagerMasterActor: Stopping BlockManagerMaster\n14/06/27 13:59:10 INFO BlockManagerMaster: BlockManagerMaster stopped\n14/06/27 13:59:10 INFO SparkContext: Successfully stopped SparkContext\n14/06/27 13:59:10 ERROR OneForOneStrategy: EOF reached before Python server acknowledged\norg.apache.spark.SparkException: EOF reached before Python server acknowledged\nat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:613)\nat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:584)\nat org.apache.spark.Accumulable.$plus$plus$eq(Accumulators.scala:72)\nat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:280)\nat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:278)\nat scala.collection.TraversableLike$WithFilter$$anonfun$foreach$1.apply(TraversableLike.scala:772)\nat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\nat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\nat scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:226)\nat scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:39)\nat scala.collection.mutable.HashMap.foreach(HashMap.scala:98)\nat scala.collection.TraversableLike$WithFilter.foreach(TraversableLike.scala:771)\nat org.apache.spark.Accumulators$.add(Accumulators.scala:278)\nat org.apache.spark.scheduler.DAGScheduler.handleTaskCompletion(DAGScheduler.scala:825)\nat org.apache.spark.scheduler.DAGSchedulerEventProcessActor$$anonfun$receive$2.applyOrElse(DAGScheduler.scala:1223)\nat akka.actor.ActorCell.receiveMessage(ActorCell.scala:498)\nat akka.actor.ActorCell.invoke(ActorCell.scala:456)\nat akka.dispatch.Mailbox.processMailbox(Mailbox.scala:237)\nat akka.dispatch.Mailbox.run(Mailbox.scala:219)\nat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\nat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\nat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\nat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\nat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n14/06/27 13:59:10 INFO RemoteActorRefProvider$RemotingTerminator: Shutting down remote daemon.\n14/06/27 13:59:10 INFO RemoteActorRefProvider$RemotingTerminator: Remote daemon shut down; proceeding with flushing remote transports.\n14/06/27 13:59:10 INFO Remoting: Remoting shut down\n\n","from":"developer"},{"body":"I'm not sure how this is related to Mesos, is this reproable using YARN or standalone?","from":"developer"},{"body":"Hi;\nNever used yarn. Doesn't happen on standalone.\n\n","from":"developer"},{"body":"This issue should be fixed in SPARK-2282 [1], I had ran the jobs above against mesos-0.19.1 after more than a hour without problems.\n\n[~therealnb] Could you also verify this?\n\n[1] https://github.com/apache/spark/commit/ef4ff00f87a4e8d38866f163f01741c2673e41da","from":"developer"},{"body":"Hi;\nSadly I moved jobs and I don't have a working Spark environment at the moment (I will be doing some Spark work soon :-). I'll pass this on to the guys that are still there and get them to confirm. \nCheers","from":"developer"},{"body":"FYI the workaround described in SPARK-2282 got me past the same \"EOF reached\" issue.\n\ni.e.\n\necho \"1\" > /proc/sys/net/ipv4/tcp_tw_reuse\necho \"1\" > /proc/sys/net/ipv4/tcp_tw_recycle\n\nthen restart your spark shell (or program).\n\nApparently a fix is slated for Spark 1.1\n\nI am running Spark 1.0.2 under Mesos 0.18.2\n","from":"developer"},{"body":"This is fixed by #2282","from":"developer"}],"created":"2014-05-08T17:52:34.000+0000","description":"I'm getting \"EOF reached before Python server acknowledged\" while using PySpark on Mesos. The error manifests itself in multiple ways. One is:\n\n{noformat}\n14/05/08 18:10:40 ERROR DAGSchedulerActorSupervisor: eventProcesserActor failed due to the error EOF reached before Python server acknowledged; shutting down SparkContext\n{noformat}\n\nAnd the other has a full stacktrace:\n\n{noformat}\n14/05/08 18:03:06 ERROR OneForOneStrategy: EOF reached before Python server acknowledged\norg.apache.spark.SparkException: EOF reached before Python server acknowledged\n\tat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:416)\n\tat org.apache.spark.api.python.PythonAccumulatorParam.addInPlace(PythonRDD.scala:387)\n\tat org.apache.spark.Accumulable.$plus$plus$eq(Accumulators.scala:71)\n\tat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:279)\n\tat org.apache.spark.Accumulators$$anonfun$add$2.apply(Accumulators.scala:277)\n\tat scala.collection.TraversableLike$WithFilter$$anonfun$foreach$1.apply(TraversableLike.scala:772)\n\tat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n\tat scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n\tat scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:226)\n\tat scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:39)\n\tat scala.collection.mutable.HashMap.foreach(HashMap.scala:98)\n\tat scala.collection.TraversableLike$WithFilter.foreach(TraversableLike.scala:771)\n\tat org.apache.spark.Accumulators$.add(Accumulators.scala:277)\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskCompletion(DAGScheduler.scala:818)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor$$anonfun$receive$2.applyOrElse(DAGScheduler.scala:1204)\n\tat akka.actor.ActorCell.receiveMessage(ActorCell.scala:498)\n\tat akka.actor.ActorCell.invoke(ActorCell.scala:456)\n\tat akka.dispatch.Mailbox.processMailbox(Mailbox.scala:237)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:219)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n{noformat}\n\nThis error causes the SparkContext to shutdown. I have not been able to reliably reproduce this bug, it seems to happen randomly, but if you run enough tasks on a SparkContext it'll hapen eventually","issue_id":"12713136","key":"SPARK-1764","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2014-09-12T00:01:22.000+0000","role":"fixed_distractor","summary":"EOF reached before Python server acknowledged"} {"case_id":"13007849","cluster":"DISTRACTOR-SPARK-17685","comments":[{"body":"User 'wangyum' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15259","created":"2016-09-27T11:08:05.733+0000"},{"body":"User 'wangyum' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17920","created":"2017-05-09T08:58:03.053+0000"}],"conversations":[{"body":"The following SQL query reproduces this issue:\n\n{code:sql}\nCREATE TABLE tab1(int int, int2 int, str string);\nCREATE TABLE tab2(int int, int2 int, str string);\nINSERT INTO tab1 values(1,1,'str');\nINSERT INTO tab1 values(2,2,'str');\nINSERT INTO tab2 values(1,1,'str');\nINSERT INTO tab2 values(2,3,'str');\n\nSELECT\n count(*)\nFROM\n (\n SELECT t1.int, t2.int2 \n FROM (SELECT * FROM tab1 LIMIT 1310721) t1\n INNER JOIN (SELECT * FROM tab2 LIMIT 1310721) t2 \n ON (t1.int = t2.int AND t1.int2 = t2.int2)\n ) t;\n{code}\n\nException thrown:\n\n{noformat}\njava.lang.IndexOutOfBoundsException: 1\n\tat scala.collection.LinearSeqOptimized$class.apply(LinearSeqOptimized.scala:65)\n\tat scala.collection.immutable.List.apply(List.scala:84)\n\tat org.apache.spark.sql.catalyst.expressions.BoundReference.doGenCode(BoundAttribute.scala:64)\n\tat org.apache.spark.sql.catalyst.expressions.Expression$$anonfun$genCode$2.apply(Expression.scala:104)\n\tat org.apache.spark.sql.catalyst.expressions.Expression$$anonfun$genCode$2.apply(Expression.scala:101)\n\tat scala.Option.getOrElse(Option.scala:121)\n\tat org.apache.spark.sql.catalyst.expressions.Expression.genCode(Expression.scala:101)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$createJoinKey$1.apply(SortMergeJoinExec.scala:334)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$createJoinKey$1.apply(SortMergeJoinExec.scala:334)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.immutable.List.foreach(List.scala:381)\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n\tat scala.collection.immutable.List.map(List.scala:285)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.createJoinKey(SortMergeJoinExec.scala:334)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.genScanner(SortMergeJoinExec.scala:369)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.doProduce(SortMergeJoinExec.scala:512)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.produce(SortMergeJoinExec.scala:35)\n\tat org.apache.spark.sql.execution.ProjectExec.doProduce(basicPhysicalOperators.scala:40)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.ProjectExec.produce(basicPhysicalOperators.scala:30)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduceWithoutKeys(HashAggregateExec.scala:215)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduce(HashAggregateExec.scala:143)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.produce(HashAggregateExec.scala:37)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduceWithoutKeys(HashAggregateExec.scala:215)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduce(HashAggregateExec.scala:143)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.produce(HashAggregateExec.scala:37)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec.doCodeGen(WholeStageCodegenExec.scala:309)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec.doExecute(WholeStageCodegenExec.scala:347)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:114)\n\tat org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:240)\n\tat org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:287)\n\tat org.apache.spark.sql.execution.SparkPlan.executeCollectPublic(SparkPlan.scala:310)\n\tat org.apache.spark.sql.execution.QueryExecution$$anonfun$hiveResultString$3.apply(QueryExecution.scala:129)\n\tat org.apache.spark.sql.execution.QueryExecution$$anonfun$hiveResultString$3.apply(QueryExecution.scala:128)\n\tat org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:57)\n\tat org.apache.spark.sql.execution.QueryExecution.hiveResultString(QueryExecution.scala:128)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:63)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:331)\n\tat org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:376)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$.main(SparkSQLCLIDriver.scala:247)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.main(SparkSQLCLIDriver.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:729)\n\tat org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:185)\n\tat org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:210)\n\tat org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:124)\n\tat org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\njava.lang.IndexOutOfBoundsException: 1\n\tat scala.collection.LinearSeqOptimized$class.apply(LinearSeqOptimized.scala:65)\n\tat scala.collection.immutable.List.apply(List.scala:84)\n\tat org.apache.spark.sql.catalyst.expressions.BoundReference.doGenCode(BoundAttribute.scala:64)\n\tat org.apache.spark.sql.catalyst.expressions.Expression$$anonfun$genCode$2.apply(Expression.scala:104)\n\tat org.apache.spark.sql.catalyst.expressions.Expression$$anonfun$genCode$2.apply(Expression.scala:101)\n\tat scala.Option.getOrElse(Option.scala:121)\n\tat org.apache.spark.sql.catalyst.expressions.Expression.genCode(Expression.scala:101)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$createJoinKey$1.apply(SortMergeJoinExec.scala:334)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$createJoinKey$1.apply(SortMergeJoinExec.scala:334)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.immutable.List.foreach(List.scala:381)\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n\tat scala.collection.immutable.List.map(List.scala:285)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.createJoinKey(SortMergeJoinExec.scala:334)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.genScanner(SortMergeJoinExec.scala:369)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.doProduce(SortMergeJoinExec.scala:512)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.produce(SortMergeJoinExec.scala:35)\n\tat org.apache.spark.sql.execution.ProjectExec.doProduce(basicPhysicalOperators.scala:40)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.ProjectExec.produce(basicPhysicalOperators.scala:30)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduceWithoutKeys(HashAggregateExec.scala:215)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduce(HashAggregateExec.scala:143)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.produce(HashAggregateExec.scala:37)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduceWithoutKeys(HashAggregateExec.scala:215)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduce(HashAggregateExec.scala:143)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.produce(HashAggregateExec.scala:37)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec.doCodeGen(WholeStageCodegenExec.scala:309)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec.doExecute(WholeStageCodegenExec.scala:347)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:114)\n\tat org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:240)\n\tat org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:287)\n\tat org.apache.spark.sql.execution.SparkPlan.executeCollectPublic(SparkPlan.scala:310)\n\tat org.apache.spark.sql.execution.QueryExecution$$anonfun$hiveResultString$3.apply(QueryExecution.scala:129)\n\tat org.apache.spark.sql.execution.QueryExecution$$anonfun$hiveResultString$3.apply(QueryExecution.scala:128)\n\tat org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:57)\n\tat org.apache.spark.sql.execution.QueryExecution.hiveResultString(QueryExecution.scala:128)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:63)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:331)\n\tat org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:376)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$.main(SparkSQLCLIDriver.scala:247)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.main(SparkSQLCLIDriver.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:729)\n\tat org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:185)\n\tat org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:210)\n\tat org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:124)\n\tat org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\n\n{noformat}\n\nThis bug is a regression. Spark 1.6 doesn't have this issue.","from":"reporter","subject":"WholeStageCodegenExec throws IndexOutOfBoundsException"},{"body":"User 'wangyum' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15259","from":"developer"},{"body":"User 'wangyum' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17920","from":"developer"}],"created":"2016-09-27T08:46:33.000+0000","description":"The following SQL query reproduces this issue:\n\n{code:sql}\nCREATE TABLE tab1(int int, int2 int, str string);\nCREATE TABLE tab2(int int, int2 int, str string);\nINSERT INTO tab1 values(1,1,'str');\nINSERT INTO tab1 values(2,2,'str');\nINSERT INTO tab2 values(1,1,'str');\nINSERT INTO tab2 values(2,3,'str');\n\nSELECT\n count(*)\nFROM\n (\n SELECT t1.int, t2.int2 \n FROM (SELECT * FROM tab1 LIMIT 1310721) t1\n INNER JOIN (SELECT * FROM tab2 LIMIT 1310721) t2 \n ON (t1.int = t2.int AND t1.int2 = t2.int2)\n ) t;\n{code}\n\nException thrown:\n\n{noformat}\njava.lang.IndexOutOfBoundsException: 1\n\tat scala.collection.LinearSeqOptimized$class.apply(LinearSeqOptimized.scala:65)\n\tat scala.collection.immutable.List.apply(List.scala:84)\n\tat org.apache.spark.sql.catalyst.expressions.BoundReference.doGenCode(BoundAttribute.scala:64)\n\tat org.apache.spark.sql.catalyst.expressions.Expression$$anonfun$genCode$2.apply(Expression.scala:104)\n\tat org.apache.spark.sql.catalyst.expressions.Expression$$anonfun$genCode$2.apply(Expression.scala:101)\n\tat scala.Option.getOrElse(Option.scala:121)\n\tat org.apache.spark.sql.catalyst.expressions.Expression.genCode(Expression.scala:101)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$createJoinKey$1.apply(SortMergeJoinExec.scala:334)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$createJoinKey$1.apply(SortMergeJoinExec.scala:334)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.immutable.List.foreach(List.scala:381)\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n\tat scala.collection.immutable.List.map(List.scala:285)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.createJoinKey(SortMergeJoinExec.scala:334)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.genScanner(SortMergeJoinExec.scala:369)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.doProduce(SortMergeJoinExec.scala:512)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.produce(SortMergeJoinExec.scala:35)\n\tat org.apache.spark.sql.execution.ProjectExec.doProduce(basicPhysicalOperators.scala:40)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.ProjectExec.produce(basicPhysicalOperators.scala:30)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduceWithoutKeys(HashAggregateExec.scala:215)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduce(HashAggregateExec.scala:143)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.produce(HashAggregateExec.scala:37)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduceWithoutKeys(HashAggregateExec.scala:215)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduce(HashAggregateExec.scala:143)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.produce(HashAggregateExec.scala:37)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec.doCodeGen(WholeStageCodegenExec.scala:309)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec.doExecute(WholeStageCodegenExec.scala:347)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:114)\n\tat org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:240)\n\tat org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:287)\n\tat org.apache.spark.sql.execution.SparkPlan.executeCollectPublic(SparkPlan.scala:310)\n\tat org.apache.spark.sql.execution.QueryExecution$$anonfun$hiveResultString$3.apply(QueryExecution.scala:129)\n\tat org.apache.spark.sql.execution.QueryExecution$$anonfun$hiveResultString$3.apply(QueryExecution.scala:128)\n\tat org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:57)\n\tat org.apache.spark.sql.execution.QueryExecution.hiveResultString(QueryExecution.scala:128)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:63)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:331)\n\tat org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:376)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$.main(SparkSQLCLIDriver.scala:247)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.main(SparkSQLCLIDriver.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:729)\n\tat org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:185)\n\tat org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:210)\n\tat org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:124)\n\tat org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\njava.lang.IndexOutOfBoundsException: 1\n\tat scala.collection.LinearSeqOptimized$class.apply(LinearSeqOptimized.scala:65)\n\tat scala.collection.immutable.List.apply(List.scala:84)\n\tat org.apache.spark.sql.catalyst.expressions.BoundReference.doGenCode(BoundAttribute.scala:64)\n\tat org.apache.spark.sql.catalyst.expressions.Expression$$anonfun$genCode$2.apply(Expression.scala:104)\n\tat org.apache.spark.sql.catalyst.expressions.Expression$$anonfun$genCode$2.apply(Expression.scala:101)\n\tat scala.Option.getOrElse(Option.scala:121)\n\tat org.apache.spark.sql.catalyst.expressions.Expression.genCode(Expression.scala:101)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$createJoinKey$1.apply(SortMergeJoinExec.scala:334)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$createJoinKey$1.apply(SortMergeJoinExec.scala:334)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n\tat scala.collection.immutable.List.foreach(List.scala:381)\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n\tat scala.collection.immutable.List.map(List.scala:285)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.createJoinKey(SortMergeJoinExec.scala:334)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.genScanner(SortMergeJoinExec.scala:369)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.doProduce(SortMergeJoinExec.scala:512)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec.produce(SortMergeJoinExec.scala:35)\n\tat org.apache.spark.sql.execution.ProjectExec.doProduce(basicPhysicalOperators.scala:40)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.ProjectExec.produce(basicPhysicalOperators.scala:30)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduceWithoutKeys(HashAggregateExec.scala:215)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduce(HashAggregateExec.scala:143)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.produce(HashAggregateExec.scala:37)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduceWithoutKeys(HashAggregateExec.scala:215)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.doProduce(HashAggregateExec.scala:143)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:83)\n\tat org.apache.spark.sql.execution.CodegenSupport$$anonfun$produce$1.apply(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.CodegenSupport$class.produce(WholeStageCodegenExec.scala:78)\n\tat org.apache.spark.sql.execution.aggregate.HashAggregateExec.produce(HashAggregateExec.scala:37)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec.doCodeGen(WholeStageCodegenExec.scala:309)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec.doExecute(WholeStageCodegenExec.scala:347)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:114)\n\tat org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:240)\n\tat org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:287)\n\tat org.apache.spark.sql.execution.SparkPlan.executeCollectPublic(SparkPlan.scala:310)\n\tat org.apache.spark.sql.execution.QueryExecution$$anonfun$hiveResultString$3.apply(QueryExecution.scala:129)\n\tat org.apache.spark.sql.execution.QueryExecution$$anonfun$hiveResultString$3.apply(QueryExecution.scala:128)\n\tat org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:57)\n\tat org.apache.spark.sql.execution.QueryExecution.hiveResultString(QueryExecution.scala:128)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:63)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:331)\n\tat org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:376)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$.main(SparkSQLCLIDriver.scala:247)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.main(SparkSQLCLIDriver.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:729)\n\tat org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:185)\n\tat org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:210)\n\tat org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:124)\n\tat org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\n\n{noformat}\n\nThis bug is a regression. Spark 1.6 doesn't have this issue.","issue_id":"13007849","key":"SPARK-17685","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-05-10T02:46:25.000+0000","role":"fixed_distractor","summary":"WholeStageCodegenExec throws IndexOutOfBoundsException"} {"case_id":"13008817","cluster":"DISTRACTOR-SPARK-17742","comments":[{"body":"I dug into the launcher code to see if I can figure out how it is working and see if I could find the bug. But when I reached LauncherServer's ServerConnection's handle method and found that this is socket programming I found it harder to find where the messages are coming from. Still trying to figure out but maybe someone who knows spark code better will find it easier to find the bug.","created":"2016-09-30T07:35:18.038+0000"},{"body":"This code in \"LocalSchedulerBackend\" causes the issue you're seeing:\n\n{code}\n override def stop() {\n stop(SparkAppHandle.State.FINISHED)\n }\n{code}\n\nThat code is run both when you explicitly stop a SparkContext, or when the VM shuts down, via a shutdown hook. You could remove that line (and similar lines in other backends), and change so that if the child process exist with a 0 exit code, then the app is \"successful\". But need to make sure it still works with YARN (both client and cluster mode), since the logic to report app state is different there.","created":"2016-10-03T18:19:08.898+0000"},{"body":"Hi, is there already a fix for this issue? Because it effectively renders the SparkLauncher unusable.. Or does anybody know at least a workaround for it? Any thoughts [~anshbansal], [~vanzin] ? Thanks!","created":"2017-04-04T07:24:20.280+0000"},{"body":"[~daanvdn] We ended up using kafka messages to communicate to the web app that was using the launcher to launch the job whether the job was complete or failed. Dumped Launcher's states as they are broken.","created":"2017-04-04T07:32:11.219+0000"},{"body":"Thanks for the advice [~anshbansal], I'll give that a try. \nA question to the spark developers: is a fix for this bug already on the roadmap for a next release of spark?\n\n","created":"2017-04-04T09:29:24.166+0000"},{"body":"[~daanvdn] did you implement Marcelo's suggestion? this is pretty WYSIWYG -- if you want to push a change, investigate it and test and then open a PR. If you don't see anything here, nobody's working on it.","created":"2017-04-04T09:31:23.630+0000"},{"body":"Hi [~srowen], unfortunately I'm a java dev, so changing stuff in a scala code base is quite a bit out of my comfort zone. In case that changes, I will definitely look into it :-), \n\nCheers","created":"2017-04-04T10:57:25.384+0000"},{"body":"Is anybody working on this? If no, I can help.","created":"2017-07-10T07:10:56.329+0000"},{"body":"https://github.com/apache/spark/pull/18877","created":"2017-08-15T18:26:59.548+0000"},{"body":"Reopening temporarily for a follow up fix:\nhttps://github.com/apache/spark/pull/19012","created":"2017-08-21T21:22:56.338+0000"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19012","created":"2017-09-01T18:02:04.426+0000"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18877","created":"2018-08-30T07:45:05.467+0000"}],"conversations":[{"body":"I tried to launch an application using the below code. This is dummy code to reproduce the problem. I tried exiting spark with status -1, throwing an exception etc. but in no case did the listener give me failed status. But if a spark job returns -1 or throws an exception from the main method it should be considered as a failure. \n\n{code}\npackage com.example;\n\nimport org.apache.spark.launcher.SparkAppHandle;\nimport org.apache.spark.launcher.SparkLauncher;\n\nimport java.io.IOException;\n\npublic class Main2 {\n\n public static void main(String[] args) throws IOException, InterruptedException {\n SparkLauncher launcher = new SparkLauncher()\n .setSparkHome(\"/opt/spark2\")\n .setAppResource(\"/home/aseem/projects/testsparkjob/build/libs/testsparkjob-1.0-SNAPSHOT.jar\")\n .setMainClass(\"com.example.Main\")\n .setMaster(\"local[2]\");\n\n launcher.startApplication(new MyListener());\n\n Thread.sleep(1000 * 60);\n }\n\n}\n\nclass MyListener implements SparkAppHandle.Listener {\n\n @Override\n public void stateChanged(SparkAppHandle handle) {\n\n System.out.println(\"state changed \" + handle.getState());\n }\n\n @Override\n public void infoChanged(SparkAppHandle handle) {\n System.out.println(\"info changed \" + handle.getState());\n }\n}\n{code}\n\nThe spark job is \n{code}\npackage com.example;\n\nimport org.apache.spark.sql.SparkSession;\nimport java.io.IOException;\n\npublic class Main {\n\n public static void main(String[] args) throws IOException {\n SparkSession sparkSession = SparkSession\n .builder()\n .appName(\"\" + System.currentTimeMillis())\n .getOrCreate();\n\n\n try {\n for (int i = 0; i < 15; i++) {\n Thread.sleep(1000);\n System.out.println(\"sleeping 1\");\n }\n } catch (InterruptedException e) {\n e.printStackTrace();\n }\n// sparkSession.stop();\n\n System.exit(-1);\n }\n\n}\n{code}","from":"reporter","subject":"Spark Launcher does not get failed state in Listener "},{"body":"I dug into the launcher code to see if I can figure out how it is working and see if I could find the bug. But when I reached LauncherServer's ServerConnection's handle method and found that this is socket programming I found it harder to find where the messages are coming from. Still trying to figure out but maybe someone who knows spark code better will find it easier to find the bug.","from":"developer"},{"body":"This code in \"LocalSchedulerBackend\" causes the issue you're seeing:\n\n{code}\n override def stop() {\n stop(SparkAppHandle.State.FINISHED)\n }\n{code}\n\nThat code is run both when you explicitly stop a SparkContext, or when the VM shuts down, via a shutdown hook. You could remove that line (and similar lines in other backends), and change so that if the child process exist with a 0 exit code, then the app is \"successful\". But need to make sure it still works with YARN (both client and cluster mode), since the logic to report app state is different there.","from":"developer"},{"body":"Hi, is there already a fix for this issue? Because it effectively renders the SparkLauncher unusable.. Or does anybody know at least a workaround for it? Any thoughts [~anshbansal], [~vanzin] ? Thanks!","from":"developer"},{"body":"[~daanvdn] We ended up using kafka messages to communicate to the web app that was using the launcher to launch the job whether the job was complete or failed. Dumped Launcher's states as they are broken.","from":"developer"},{"body":"Thanks for the advice [~anshbansal], I'll give that a try. \nA question to the spark developers: is a fix for this bug already on the roadmap for a next release of spark?\n\n","from":"developer"},{"body":"[~daanvdn] did you implement Marcelo's suggestion? this is pretty WYSIWYG -- if you want to push a change, investigate it and test and then open a PR. If you don't see anything here, nobody's working on it.","from":"developer"},{"body":"Hi [~srowen], unfortunately I'm a java dev, so changing stuff in a scala code base is quite a bit out of my comfort zone. In case that changes, I will definitely look into it :-), \n\nCheers","from":"developer"},{"body":"Is anybody working on this? If no, I can help.","from":"developer"},{"body":"https://github.com/apache/spark/pull/18877","from":"developer"},{"body":"Reopening temporarily for a follow up fix:\nhttps://github.com/apache/spark/pull/19012","from":"developer"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19012","from":"developer"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18877","from":"developer"}],"created":"2016-09-30T07:25:11.000+0000","description":"I tried to launch an application using the below code. This is dummy code to reproduce the problem. I tried exiting spark with status -1, throwing an exception etc. but in no case did the listener give me failed status. But if a spark job returns -1 or throws an exception from the main method it should be considered as a failure. \n\n{code}\npackage com.example;\n\nimport org.apache.spark.launcher.SparkAppHandle;\nimport org.apache.spark.launcher.SparkLauncher;\n\nimport java.io.IOException;\n\npublic class Main2 {\n\n public static void main(String[] args) throws IOException, InterruptedException {\n SparkLauncher launcher = new SparkLauncher()\n .setSparkHome(\"/opt/spark2\")\n .setAppResource(\"/home/aseem/projects/testsparkjob/build/libs/testsparkjob-1.0-SNAPSHOT.jar\")\n .setMainClass(\"com.example.Main\")\n .setMaster(\"local[2]\");\n\n launcher.startApplication(new MyListener());\n\n Thread.sleep(1000 * 60);\n }\n\n}\n\nclass MyListener implements SparkAppHandle.Listener {\n\n @Override\n public void stateChanged(SparkAppHandle handle) {\n\n System.out.println(\"state changed \" + handle.getState());\n }\n\n @Override\n public void infoChanged(SparkAppHandle handle) {\n System.out.println(\"info changed \" + handle.getState());\n }\n}\n{code}\n\nThe spark job is \n{code}\npackage com.example;\n\nimport org.apache.spark.sql.SparkSession;\nimport java.io.IOException;\n\npublic class Main {\n\n public static void main(String[] args) throws IOException {\n SparkSession sparkSession = SparkSession\n .builder()\n .appName(\"\" + System.currentTimeMillis())\n .getOrCreate();\n\n\n try {\n for (int i = 0; i < 15; i++) {\n Thread.sleep(1000);\n System.out.println(\"sleeping 1\");\n }\n } catch (InterruptedException e) {\n e.printStackTrace();\n }\n// sparkSession.stop();\n\n System.exit(-1);\n }\n\n}\n{code}","issue_id":"13008817","key":"SPARK-17742","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-08-25T17:04:51.000+0000","role":"fixed_distractor","summary":"Spark Launcher does not get failed state in Listener "} {"case_id":"13009008","cluster":"DISTRACTOR-SPARK-17753","comments":[{"body":"Hi, [~kdhuria].\nRight, `CASE WHEN` seems to have a bug on that. I'll investigate this.","created":"2016-10-01T18:36:04.374+0000"},{"body":"Hmm. It's not so each to fix this. Maybe, I'll try later.","created":"2016-10-01T23:00:35.883+0000"},{"body":"We do not support the GTE/LTE/GT/LT/EQ/NEQ operators. Use the symbolic versions instead.\n\nI will open a PR for using complex expressions in a value based case statement.","created":"2016-10-01T23:49:43.243+0000"},{"body":"Hi Herman,\nI have tried with symbolic also, it doesn't work. Even this query fails with same error.\nSELECT alias.p_double as a0, alias.p_text as a1, NULL as a2 FROM hadoop_tbl_all alias WHERE (1 = (CASE ('aaaaabbbbb' = alias.p_text) WHEN TRUE THEN 1 WHEN FALSE THEN 0 ELSE CAST(NULL AS INT) END))\n\nAny boolean condition after case doesn't work.","created":"2016-10-01T23:53:33.498+0000"},{"body":"Yeah, the current grammar doesn't allow us to use a complex expression as the input for a simple case when statement. I'll create a fix (PR) for this. You can quite easily work around this issue, if it is blocking you.","created":"2016-10-02T00:01:56.618+0000"},{"body":"Yeah, Thanks!","created":"2016-10-02T00:07:47.780+0000"},{"body":"User 'hvanhovell' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15322","created":"2016-10-02T01:07:07.190+0000"},{"body":"how have you solved the following error:-\n\n\n at org.apache.spark.sql.catalyst.parser.ParseException.withCommand(ParseDriver.scala:197)\n at org.apache.spark.sql.catalyst.parser.AbstractSqlParser.parse(ParseDriver.scala:99)\n at org.apache.spark.sql.execution.SparkSqlParser.parse(SparkSqlParser.scala:45)\n at org.apache.spark.sql.catalyst.parser.AbstractSqlParser.parsePlan(ParseDriver.scala:53)\n at org.apache.spark.sql.SparkSession.sql(SparkSession.scala:582)\n at org.apache.spark.sql.SQLContext.sql(SQLContext.scala:682)\n ... 48 elided\n","created":"2016-11-09T10:08:49.636+0000"}],"conversations":[{"body":"Simple case in sql throws parser exception in spark 2.0.\nThe following query as well as similar queries fail in spark 2.0 \n{noformat}\nscala> spark.sql(\"SELECT alias.p_double as a0, alias.p_text as a1, NULL as a2 FROM hadoop_tbl_all alias WHERE (1 = (CASE ('aaaaabbbbb' = alias.p_text) OR (8 LTE LENGTH(alias.p_text)) WHEN TRUE THEN 1 WHEN FALSE THEN 0 ELSE CAST(NULL AS INT) END))\")\norg.apache.spark.sql.catalyst.parser.ParseException:\nmismatched input 'FROM' expecting {, 'WHERE', 'GROUP', 'ORDER', 'HAVING', 'LIMIT', 'LATERAL', 'WINDOW', 'UNION', 'EXCEPT', 'INTERSECT', 'SORT', 'CLUSTER', 'DISTRIBUTE'}(line 1, pos 60)\n\n== SQL ==\nSELECT alias.p_double as a0, alias.p_text as a1, NULL as a2 FROM hadoop_tbl_all alias WHERE (1 = (CASE ('aaaaabbbbb' = alias.p_text) OR (8 LTE LENGTH(alias.p_text)) WHEN TRUE THEN 1 WHEN FALSE THEN 0 ELSE CAST(NULL AS INT) END))\n------------------------------------------------------------^^^\n\n at org.apache.spark.sql.catalyst.parser.ParseException.withCommand(ParseDriver.scala:197)\n at org.apache.spark.sql.catalyst.parser.AbstractSqlParser.parse(ParseDriver.scala:99)\n at org.apache.spark.sql.execution.SparkSqlParser.parse(SparkSqlParser.scala:46)\n at org.apache.spark.sql.catalyst.parser.AbstractSqlParser.parsePlan(ParseDriver.scala:53)\n at org.apache.spark.sql.SparkSession.sql(SparkSession.scala:582)\n ... 48 elided\n{noformat}","from":"reporter","subject":"Simple case in spark sql throws ParseException"},{"body":"Hi, [~kdhuria].\nRight, `CASE WHEN` seems to have a bug on that. I'll investigate this.","from":"developer"},{"body":"Hmm. It's not so each to fix this. Maybe, I'll try later.","from":"developer"},{"body":"We do not support the GTE/LTE/GT/LT/EQ/NEQ operators. Use the symbolic versions instead.\n\nI will open a PR for using complex expressions in a value based case statement.","from":"developer"},{"body":"Hi Herman,\nI have tried with symbolic also, it doesn't work. Even this query fails with same error.\nSELECT alias.p_double as a0, alias.p_text as a1, NULL as a2 FROM hadoop_tbl_all alias WHERE (1 = (CASE ('aaaaabbbbb' = alias.p_text) WHEN TRUE THEN 1 WHEN FALSE THEN 0 ELSE CAST(NULL AS INT) END))\n\nAny boolean condition after case doesn't work.","from":"developer"},{"body":"Yeah, the current grammar doesn't allow us to use a complex expression as the input for a simple case when statement. I'll create a fix (PR) for this. You can quite easily work around this issue, if it is blocking you.","from":"developer"},{"body":"Yeah, Thanks!","from":"developer"},{"body":"User 'hvanhovell' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15322","from":"developer"},{"body":"how have you solved the following error:-\n\n\n at org.apache.spark.sql.catalyst.parser.ParseException.withCommand(ParseDriver.scala:197)\n at org.apache.spark.sql.catalyst.parser.AbstractSqlParser.parse(ParseDriver.scala:99)\n at org.apache.spark.sql.execution.SparkSqlParser.parse(SparkSqlParser.scala:45)\n at org.apache.spark.sql.catalyst.parser.AbstractSqlParser.parsePlan(ParseDriver.scala:53)\n at org.apache.spark.sql.SparkSession.sql(SparkSession.scala:582)\n at org.apache.spark.sql.SQLContext.sql(SQLContext.scala:682)\n ... 48 elided\n","from":"developer"}],"created":"2016-09-30T22:28:17.000+0000","description":"Simple case in sql throws parser exception in spark 2.0.\nThe following query as well as similar queries fail in spark 2.0 \n{noformat}\nscala> spark.sql(\"SELECT alias.p_double as a0, alias.p_text as a1, NULL as a2 FROM hadoop_tbl_all alias WHERE (1 = (CASE ('aaaaabbbbb' = alias.p_text) OR (8 LTE LENGTH(alias.p_text)) WHEN TRUE THEN 1 WHEN FALSE THEN 0 ELSE CAST(NULL AS INT) END))\")\norg.apache.spark.sql.catalyst.parser.ParseException:\nmismatched input 'FROM' expecting {, 'WHERE', 'GROUP', 'ORDER', 'HAVING', 'LIMIT', 'LATERAL', 'WINDOW', 'UNION', 'EXCEPT', 'INTERSECT', 'SORT', 'CLUSTER', 'DISTRIBUTE'}(line 1, pos 60)\n\n== SQL ==\nSELECT alias.p_double as a0, alias.p_text as a1, NULL as a2 FROM hadoop_tbl_all alias WHERE (1 = (CASE ('aaaaabbbbb' = alias.p_text) OR (8 LTE LENGTH(alias.p_text)) WHEN TRUE THEN 1 WHEN FALSE THEN 0 ELSE CAST(NULL AS INT) END))\n------------------------------------------------------------^^^\n\n at org.apache.spark.sql.catalyst.parser.ParseException.withCommand(ParseDriver.scala:197)\n at org.apache.spark.sql.catalyst.parser.AbstractSqlParser.parse(ParseDriver.scala:99)\n at org.apache.spark.sql.execution.SparkSqlParser.parse(SparkSqlParser.scala:46)\n at org.apache.spark.sql.catalyst.parser.AbstractSqlParser.parsePlan(ParseDriver.scala:53)\n at org.apache.spark.sql.SparkSession.sql(SparkSession.scala:582)\n ... 48 elided\n{noformat}","issue_id":"13009008","key":"SPARK-17753","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-10-04T02:33:43.000+0000","role":"fixed_distractor","summary":"Simple case in spark sql throws ParseException"} {"case_id":"13012121","cluster":"DISTRACTOR-SPARK-17920","comments":[{"body":"[~norvellj]\nThis is similar to SPARK-19580. I'm also struggling with the same problem in Amazon EMR. Looking for way how to fix it.","created":"2017-02-20T16:44:55.669+0000"},{"body":"I also get the same problem, i implement a hive serde, and in my implemention i need the parameter Configuration to get some static and dynamic settings, i run it in hive well. but when i use spark's HiveContext to insert data to hive table, in serialize it get the configuration as null, so causing NullPointerException!\nI'm using Spark1.5, i found the problem is in class org.apache.spark.sql.hive.execution.InsertIntoHiveTable, the function newSerializer uses serializer.initialize(null, tableDesc.getProperties) to get the serializer. This also affects Spark1.6 and Spark 2.0","created":"2017-02-27T11:25:35.081+0000"},{"body":"User 'vinodkc' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19779","created":"2017-11-18T09:02:03.894+0000"},{"body":"User 'vinodkc' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19795","created":"2017-11-22T13:01:05.238+0000"},{"body":"User 'cloud-fan' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19799","created":"2017-11-23T11:43:05.140+0000"},{"body":"User 'vinodkc' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19809","created":"2017-11-24T06:13:04.046+0000"}],"conversations":[{"body":"When HiveWriterContainer intializes a serde it explicitly passes null for the Configuration:\n\nhttps://github.com/apache/spark/blob/v2.0.0/sql/hive/src/main/scala/org/apache/spark/sql/hive/hiveWriterContainers.scala#L161\n\nWhen attempting to write to a table stored as Avro with avro.schema.url set, this causes a NullPointerException when it tries to get the FileSystem for the URL:\n\nhttps://github.com/apache/hive/blob/release-2.1.0-rc3/serde/src/java/org/apache/hadoop/hive/serde2/avro/AvroSerdeUtils.java#L153\n\nReproduction:\n{noformat}\nspark-sql> create external table avro_in (a string) stored as avro location '/avro-in/' tblproperties ('avro.schema.url'='/avro-schema/avro.avsc');\n\nspark-sql> create external table avro_out (a string) stored as avro location '/avro-out/' tblproperties ('avro.schema.url'='/avro-schema/avro.avsc');\n\nspark-sql> select * from avro_in;\nhello\nTime taken: 1.986 seconds, Fetched 1 row(s)\n\nspark-sql> insert overwrite table avro_out select * from avro_in;\n\n16/10/13 19:34:47 WARN AvroSerDe: Encountered exception determining schema. Returning signal schema to indicate problem\njava.lang.NullPointerException\n\tat org.apache.hadoop.fs.FileSystem.getDefaultUri(FileSystem.java:182)\n\tat org.apache.hadoop.fs.FileSystem.get(FileSystem.java:174)\n\tat org.apache.hadoop.fs.FileSystem.get(FileSystem.java:359)\n\tat org.apache.hadoop.hive.serde2.avro.AvroSerdeUtils.getSchemaFromFS(AvroSerdeUtils.java:131)\n\tat org.apache.hadoop.hive.serde2.avro.AvroSerdeUtils.determineSchemaOrThrowException(AvroSerdeUtils.java:112)\n\tat org.apache.hadoop.hive.serde2.avro.AvroSerDe.determineSchemaOrReturnErrorSchema(AvroSerDe.java:167)\n\tat org.apache.hadoop.hive.serde2.avro.AvroSerDe.initialize(AvroSerDe.java:103)\n\tat org.apache.spark.sql.hive.SparkHiveWriterContainer.newSerializer(hiveWriterContainers.scala:161)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable.sideEffectResult$lzycompute(InsertIntoHiveTable.scala:236)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable.sideEffectResult(InsertIntoHiveTable.scala:142)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable.doExecute(InsertIntoHiveTable.scala:313)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:114)\n\tat org.apache.spark.sql.execution.QueryExecution.toRdd$lzycompute(QueryExecution.scala:86)\n\tat org.apache.spark.sql.execution.QueryExecution.toRdd(QueryExecution.scala:86)\n\tat org.apache.spark.sql.Dataset.(Dataset.scala:186)\n\tat org.apache.spark.sql.Dataset.(Dataset.scala:167)\n\tat org.apache.spark.sql.Dataset$.ofRows(Dataset.scala:65)\n\tat org.apache.spark.sql.SparkSession.sql(SparkSession.scala:582)\n\tat org.apache.spark.sql.SQLContext.sql(SQLContext.scala:682)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:62)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:331)\n\tat org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:376)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$.main(SparkSQLCLIDriver.scala:247)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.main(SparkSQLCLIDriver.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:729)\n\tat org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:185)\n\tat org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:210)\n\tat org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:124)\n\tat org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\n{noformat}\n\nHive fixed a similar issue in FileSinkOperator in https://issues.apache.org/jira/browse/HIVE-9651","from":"reporter","subject":"HiveWriterContainer passes null configuration to serde.initialize, causing NullPointerException in AvroSerde when using avro.schema.url"},{"body":"[~norvellj]\nThis is similar to SPARK-19580. I'm also struggling with the same problem in Amazon EMR. Looking for way how to fix it.","from":"developer"},{"body":"I also get the same problem, i implement a hive serde, and in my implemention i need the parameter Configuration to get some static and dynamic settings, i run it in hive well. but when i use spark's HiveContext to insert data to hive table, in serialize it get the configuration as null, so causing NullPointerException!\nI'm using Spark1.5, i found the problem is in class org.apache.spark.sql.hive.execution.InsertIntoHiveTable, the function newSerializer uses serializer.initialize(null, tableDesc.getProperties) to get the serializer. This also affects Spark1.6 and Spark 2.0","from":"developer"},{"body":"User 'vinodkc' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19779","from":"developer"},{"body":"User 'vinodkc' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19795","from":"developer"},{"body":"User 'cloud-fan' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19799","from":"developer"},{"body":"User 'vinodkc' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19809","from":"developer"}],"created":"2016-10-13T19:43:39.000+0000","description":"When HiveWriterContainer intializes a serde it explicitly passes null for the Configuration:\n\nhttps://github.com/apache/spark/blob/v2.0.0/sql/hive/src/main/scala/org/apache/spark/sql/hive/hiveWriterContainers.scala#L161\n\nWhen attempting to write to a table stored as Avro with avro.schema.url set, this causes a NullPointerException when it tries to get the FileSystem for the URL:\n\nhttps://github.com/apache/hive/blob/release-2.1.0-rc3/serde/src/java/org/apache/hadoop/hive/serde2/avro/AvroSerdeUtils.java#L153\n\nReproduction:\n{noformat}\nspark-sql> create external table avro_in (a string) stored as avro location '/avro-in/' tblproperties ('avro.schema.url'='/avro-schema/avro.avsc');\n\nspark-sql> create external table avro_out (a string) stored as avro location '/avro-out/' tblproperties ('avro.schema.url'='/avro-schema/avro.avsc');\n\nspark-sql> select * from avro_in;\nhello\nTime taken: 1.986 seconds, Fetched 1 row(s)\n\nspark-sql> insert overwrite table avro_out select * from avro_in;\n\n16/10/13 19:34:47 WARN AvroSerDe: Encountered exception determining schema. Returning signal schema to indicate problem\njava.lang.NullPointerException\n\tat org.apache.hadoop.fs.FileSystem.getDefaultUri(FileSystem.java:182)\n\tat org.apache.hadoop.fs.FileSystem.get(FileSystem.java:174)\n\tat org.apache.hadoop.fs.FileSystem.get(FileSystem.java:359)\n\tat org.apache.hadoop.hive.serde2.avro.AvroSerdeUtils.getSchemaFromFS(AvroSerdeUtils.java:131)\n\tat org.apache.hadoop.hive.serde2.avro.AvroSerdeUtils.determineSchemaOrThrowException(AvroSerdeUtils.java:112)\n\tat org.apache.hadoop.hive.serde2.avro.AvroSerDe.determineSchemaOrReturnErrorSchema(AvroSerDe.java:167)\n\tat org.apache.hadoop.hive.serde2.avro.AvroSerDe.initialize(AvroSerDe.java:103)\n\tat org.apache.spark.sql.hive.SparkHiveWriterContainer.newSerializer(hiveWriterContainers.scala:161)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable.sideEffectResult$lzycompute(InsertIntoHiveTable.scala:236)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable.sideEffectResult(InsertIntoHiveTable.scala:142)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable.doExecute(InsertIntoHiveTable.scala:313)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:115)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:136)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n\tat org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:133)\n\tat org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:114)\n\tat org.apache.spark.sql.execution.QueryExecution.toRdd$lzycompute(QueryExecution.scala:86)\n\tat org.apache.spark.sql.execution.QueryExecution.toRdd(QueryExecution.scala:86)\n\tat org.apache.spark.sql.Dataset.(Dataset.scala:186)\n\tat org.apache.spark.sql.Dataset.(Dataset.scala:167)\n\tat org.apache.spark.sql.Dataset$.ofRows(Dataset.scala:65)\n\tat org.apache.spark.sql.SparkSession.sql(SparkSession.scala:582)\n\tat org.apache.spark.sql.SQLContext.sql(SQLContext.scala:682)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:62)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:331)\n\tat org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:376)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$.main(SparkSQLCLIDriver.scala:247)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.main(SparkSQLCLIDriver.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:729)\n\tat org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:185)\n\tat org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:210)\n\tat org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:124)\n\tat org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\n{noformat}\n\nHive fixed a similar issue in FileSinkOperator in https://issues.apache.org/jira/browse/HIVE-9651","issue_id":"13012121","key":"SPARK-17920","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-11-22T17:22:44.000+0000","role":"fixed_distractor","summary":"HiveWriterContainer passes null configuration to serde.initialize, causing NullPointerException in AvroSerde when using avro.schema.url"} {"case_id":"13012970","cluster":"DISTRACTOR-SPARK-17975","comments":[{"body":"This change resolves the issue for me: https://github.com/jvstein/spark/tree/lda-edgerdd","created":"2016-10-17T19:58:36.874+0000"},{"body":"Adding a link to another issue that seems to be related to EdgeRDD partition problems.","created":"2016-10-17T20:04:11.264+0000"},{"body":"Could you send a link to the repro dataset? I could work on this issue but it looks like you have a fix already. For any fixes we need tests to validate them.","created":"2016-12-22T16:32:18.709+0000"},{"body":"Attaching vertical bar delimited documents (one per line).\n\nWith my quick fix, I'm seeing a lot more persisted RDDs on the \"Storage\" tab. I'm either not cleaning something up or there's another issue related to that.","created":"2017-01-04T19:45:30.336+0000"},{"body":"Thank you for sending the dataset, I'm working on this issue.","created":"2017-01-06T06:00:21.639+0000"},{"body":"User 'imatiach-msft' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16494","created":"2017-01-07T01:06:04.246+0000"},{"body":"I was able to reproduce the issue based on your dataset and I've made the suggested fix in the pull request. I added a test case that had a similar issue to your dataset and could reproduce the error. Thank you!","created":"2017-01-07T01:07:23.076+0000"},{"body":"[SPARK-14804] was just fixed. [~jvstein], do you have time to test master with your code to see if the bug you hit is fixed? Thanks!","created":"2017-01-26T01:42:14.570+0000"},{"body":"[~josephkb] I was able to verify that this issue is now fixed after rebasing to master - can you please close the bug? Thank you!","created":"2017-01-26T18:09:04.578+0000"},{"body":"Will do, thanks!","created":"2017-02-09T18:49:07.728+0000"}],"conversations":[{"body":"I'm able to reproduce the error consistently with a 2000 record text file with each record having 1-5 terms and checkpointing enabled. It looks like the problem was introduced with the resolution for SPARK-13355.\n\nThe EdgeRDD class seems to be lying about it's type in a way that causes RDD.mapPartitionsWithIndex method to be unusable when it's referenced as an RDD of Edge elements.\n\n{code}\nval spark = SparkSession.builder.appName(\"lda\").getOrCreate()\nspark.sparkContext.setCheckpointDir(\"hdfs:///tmp/checkpoints\")\nval data: RDD[(Long, Vector)] = // snip\ndata.setName(\"data\").cache()\nval lda = new LDA\nval optimizer = new EMLDAOptimizer\nlda.setOptimizer(optimizer)\n .setK(10)\n .setMaxIterations(400)\n .setAlpha(-1)\n .setBeta(-1)\n .setCheckpointInterval(7)\nval ldaModel = lda.run(data)\n{code}\n\n{noformat}\n16/10/16 23:53:54 WARN TaskSetManager: Lost task 3.0 in stage 348.0 (TID 1225, server2.domain): java.lang.ClassCastException: scala.Tuple2 cannot be cast to org.apache.spark.graphx.Edge\n\tat org.apache.spark.graphx.EdgeRDD$$anonfun$1$$anonfun$apply$1.apply(EdgeRDD.scala:107)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.graphx.EdgeRDD$$anonfun$1.apply(EdgeRDD.scala:107)\n\tat org.apache.spark.graphx.EdgeRDD$$anonfun$1.apply(EdgeRDD.scala:105)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$25.apply(RDD.scala:820)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$25.apply(RDD.scala:820)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD$$anonfun$8.apply(RDD.scala:332)\n\tat org.apache.spark.rdd.RDD$$anonfun$8.apply(RDD.scala:330)\n\tat org.apache.spark.storage.BlockManager$$anonfun$doPutIterator$1.apply(BlockManager.scala:935)\n\tat org.apache.spark.storage.BlockManager$$anonfun$doPutIterator$1.apply(BlockManager.scala:926)\n\tat org.apache.spark.storage.BlockManager.doPut(BlockManager.scala:866)\n\tat org.apache.spark.storage.BlockManager.doPutIterator(BlockManager.scala:926)\n\tat org.apache.spark.storage.BlockManager.getOrElseUpdate(BlockManager.scala:670)\n\tat org.apache.spark.rdd.RDD.getOrCompute(RDD.scala:330)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:281)\n\tat org.apache.spark.graphx.EdgeRDD.compute(EdgeRDD.scala:50)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:86)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:722)\n{noformat}","from":"reporter","subject":"EMLDAOptimizer fails with ClassCastException on YARN"},{"body":"This change resolves the issue for me: https://github.com/jvstein/spark/tree/lda-edgerdd","from":"developer"},{"body":"Adding a link to another issue that seems to be related to EdgeRDD partition problems.","from":"developer"},{"body":"Could you send a link to the repro dataset? I could work on this issue but it looks like you have a fix already. For any fixes we need tests to validate them.","from":"developer"},{"body":"Attaching vertical bar delimited documents (one per line).\n\nWith my quick fix, I'm seeing a lot more persisted RDDs on the \"Storage\" tab. I'm either not cleaning something up or there's another issue related to that.","from":"developer"},{"body":"Thank you for sending the dataset, I'm working on this issue.","from":"developer"},{"body":"User 'imatiach-msft' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16494","from":"developer"},{"body":"I was able to reproduce the issue based on your dataset and I've made the suggested fix in the pull request. I added a test case that had a similar issue to your dataset and could reproduce the error. Thank you!","from":"developer"},{"body":"[SPARK-14804] was just fixed. [~jvstein], do you have time to test master with your code to see if the bug you hit is fixed? Thanks!","from":"developer"},{"body":"[~josephkb] I was able to verify that this issue is now fixed after rebasing to master - can you please close the bug? Thank you!","from":"developer"},{"body":"Will do, thanks!","from":"developer"}],"created":"2016-10-17T19:49:38.000+0000","description":"I'm able to reproduce the error consistently with a 2000 record text file with each record having 1-5 terms and checkpointing enabled. It looks like the problem was introduced with the resolution for SPARK-13355.\n\nThe EdgeRDD class seems to be lying about it's type in a way that causes RDD.mapPartitionsWithIndex method to be unusable when it's referenced as an RDD of Edge elements.\n\n{code}\nval spark = SparkSession.builder.appName(\"lda\").getOrCreate()\nspark.sparkContext.setCheckpointDir(\"hdfs:///tmp/checkpoints\")\nval data: RDD[(Long, Vector)] = // snip\ndata.setName(\"data\").cache()\nval lda = new LDA\nval optimizer = new EMLDAOptimizer\nlda.setOptimizer(optimizer)\n .setK(10)\n .setMaxIterations(400)\n .setAlpha(-1)\n .setBeta(-1)\n .setCheckpointInterval(7)\nval ldaModel = lda.run(data)\n{code}\n\n{noformat}\n16/10/16 23:53:54 WARN TaskSetManager: Lost task 3.0 in stage 348.0 (TID 1225, server2.domain): java.lang.ClassCastException: scala.Tuple2 cannot be cast to org.apache.spark.graphx.Edge\n\tat org.apache.spark.graphx.EdgeRDD$$anonfun$1$$anonfun$apply$1.apply(EdgeRDD.scala:107)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.graphx.EdgeRDD$$anonfun$1.apply(EdgeRDD.scala:107)\n\tat org.apache.spark.graphx.EdgeRDD$$anonfun$1.apply(EdgeRDD.scala:105)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$25.apply(RDD.scala:820)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$25.apply(RDD.scala:820)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD$$anonfun$8.apply(RDD.scala:332)\n\tat org.apache.spark.rdd.RDD$$anonfun$8.apply(RDD.scala:330)\n\tat org.apache.spark.storage.BlockManager$$anonfun$doPutIterator$1.apply(BlockManager.scala:935)\n\tat org.apache.spark.storage.BlockManager$$anonfun$doPutIterator$1.apply(BlockManager.scala:926)\n\tat org.apache.spark.storage.BlockManager.doPut(BlockManager.scala:866)\n\tat org.apache.spark.storage.BlockManager.doPutIterator(BlockManager.scala:926)\n\tat org.apache.spark.storage.BlockManager.getOrElseUpdate(BlockManager.scala:670)\n\tat org.apache.spark.rdd.RDD.getOrCompute(RDD.scala:330)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:281)\n\tat org.apache.spark.graphx.EdgeRDD.compute(EdgeRDD.scala:50)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:86)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:722)\n{noformat}","issue_id":"13012970","key":"SPARK-17975","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-02-09T18:53:28.000+0000","role":"fixed_distractor","summary":"EMLDAOptimizer fails with ClassCastException on YARN"} {"case_id":"13013426","cluster":"DISTRACTOR-SPARK-18004","comments":[{"body":"So there seems to something going on with the date format here. Could you check what query Spark SQL is sending to Oracle?","created":"2016-11-16T14:41:28.015+0000"},{"body":"Date format as per the physical plan logged by Spark Dataframe: PushedFilters: [LessThan(TS,2016-11-17 19:42:01.057)]\n\nConfirmed the same from Oracle query logs as well: WHERE TS < '2016-11-17 19:42:01.057'","created":"2016-11-17T14:40:23.208+0000"},{"body":"which format should be passed to oracle?","created":"2016-11-17T14:50:21.074+0000"},{"body":"The date/timestamp format Oracle expects may vary from instance to instance based on the configuration of parameters NLS_TIMESTAMP_FORMAT and NLS_DATE_FORMAT . So ideal solution would be to use the to_timestamp and to_date Oracle functions. E.g. to_timestamp('12-01-2012 21:24:00', 'dd-mm-yyyy hh24:mi:ss')\n\nhttp://stackoverflow.com/questions/8855378/oracle-sql-timestamps-in-where-clause\nhttps://docs.oracle.com/cd/B19306_01/server.102/b14225/ch4datetime.htm#i1006312","created":"2016-11-18T08:01:32.308+0000"},{"body":"The current spark Jdbc interface (JdbcDialect) cannot handle formats for database-dependent timestamps and just puts `Timestamp#toString` in where clauses (https://github.com/apache/spark/blob/master/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/jdbc/JDBCRDD.scala#L95). One solution is to transform this `Timestamp` into a database-specific timestamp format like https://github.com/apache/spark/compare/master...maropu:SPARK-18004#diff-5a29ce8f760092fb4a9c1f190cc2f61cR96","created":"2016-11-21T09:30:28.132+0000"},{"body":"a silly but doable workaround: cast TimestampType row to LongType and compare the inner representation...\n\n{code:java}\nsrc.filter(col(\"timestamp_col\").cast(LongType).*(1000l) > condition.asInstanceOf[Timestamp].getTime)\n{code}","created":"2017-05-09T08:31:09.113+0000"},{"body":"The right solution here is to make an explicit cast on the string representation of the timestamp value and not rely on implicit casting by the database\nANSI casting should work in nearly every RDBMS out there if the string is ISO 8601 format. Eg.\n{noformat}\nselect timestamp '2016-10-19 12:54:01.934';\n timestamp\n-------------------------\n 2016-10-19 12:54:01.934\n\n{noformat}","created":"2017-05-12T22:51:10.575+0000"},{"body":"Here's the culprit, a private function that converts a scala value to a SQL literal. Note use of {{toString}} to format dates and timestamps. Seems like there should be some hook to a {{JDBCDialect}} that can handle vendor-specific syntax etc.\n\n{code:lang=java}\n /**\n * Converts value to SQL expression.\n */\n private def compileValue(value: Any): Any = value match {\n case stringValue: String => s\"'${escapeSql(stringValue)}'\"\n case timestampValue: Timestamp => \"'\" + timestampValue + \"'\"\n case dateValue: Date => \"'\" + dateValue + \"'\"\n case arrayValue: Array[Any] => arrayValue.map(compileValue).mkString(\", \")\n case _ => value\n }\n{code}","created":"2017-05-15T14:36:57.808+0000"},{"body":"User 'SharpRay' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18404","created":"2017-06-23T12:13:04.426+0000"},{"body":"The PR 18404 is closed. I will resend the PR later.","created":"2017-06-23T14:18:10.272+0000"},{"body":"User 'SharpRay' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18411","created":"2017-06-24T03:11:03.457+0000"},{"body":"User 'SharpRay' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18451","created":"2017-06-28T08:56:03.984+0000"}],"conversations":[{"body":"DataFrame filter Predicate push-down fails for Oracle Timestamp type columns with Exception java.sql.SQLDataException: ORA-01861: literal does not match format string:\n\nJava source code (this code works fine for mysql & mssql databases) :\n{noformat}\n//DataFrame df = create a DataFrame over an Oracle table\ndf = df.filter(df.col(\"TS\").lt(new java.sql.Timestamp(System.currentTimeMillis())));\n\t\tdf.explain();\n\t\tdf.show();\n{noformat}\n\nLog statements with the Exception:\n{noformat}\nSchema: root\n |-- ID: string (nullable = false)\n |-- TS: timestamp (nullable = true)\n |-- DEVICE_ID: string (nullable = true)\n |-- REPLACEMENT: string (nullable = true)\n{noformat}\n\n{noformat}\n== Physical Plan ==\nFilter (TS#1 < 1476861841934000)\n+- Scan JDBCRelation(jdbc:oracle:thin:@10.0.0.111:1521:orcl,ORATABLE,[Lorg.apache.spark.Partition;@78c74647,{user=user, password=pwd, url=jdbc:oracle:thin:@10.0.0.111:1521:orcl, dbtable=ORATABLE, driver=oracle.jdbc.driver.OracleDriver})[ID#0,TS#1,DEVICE_ID#2,REPLACEMENT#3] PushedFilters: [LessThan(TS,2016-10-19 12:54:01.934)]\n2016-10-19 12:54:04,268 ERROR [Executor task launch worker-0] org.apache.spark.executor.Executor\nException in task 0.0 in stage 0.0 (TID 0)\n\njava.sql.SQLDataException: ORA-01861: literal does not match format string\n\n\tat oracle.jdbc.driver.T4CTTIoer.processError(T4CTTIoer.java:461)\n\tat oracle.jdbc.driver.T4CTTIoer.processError(T4CTTIoer.java:402)\n\tat oracle.jdbc.driver.T4C8Oall.processError(T4C8Oall.java:1065)\n\tat oracle.jdbc.driver.T4CTTIfun.receive(T4CTTIfun.java:681)\n\tat oracle.jdbc.driver.T4CTTIfun.doRPC(T4CTTIfun.java:256)\n\tat oracle.jdbc.driver.T4C8Oall.doOALL(T4C8Oall.java:577)\n\tat oracle.jdbc.driver.T4CPreparedStatement.doOall8(T4CPreparedStatement.java:239)\n\tat oracle.jdbc.driver.T4CPreparedStatement.doOall8(T4CPreparedStatement.java:75)\n\tat oracle.jdbc.driver.T4CPreparedStatement.executeForDescribe(T4CPreparedStatement.java:1043)\n\tat oracle.jdbc.driver.OracleStatement.executeMaybeDescribe(OracleStatement.java:1111)\n\tat oracle.jdbc.driver.OracleStatement.doExecuteWithTimeout(OracleStatement.java:1353)\n\tat oracle.jdbc.driver.OraclePreparedStatement.executeInternal(OraclePreparedStatement.java:4485)\n\tat oracle.jdbc.driver.OraclePreparedStatement.executeQuery(OraclePreparedStatement.java:4566)\n\tat oracle.jdbc.driver.OraclePreparedStatementWrapper.executeQuery(OraclePreparedStatementWrapper.java:5251)\n\tat org.apache.spark.sql.execution.datasources.jdbc.JDBCRDD$$anon$1.(JDBCRDD.scala:383)\n\tat org.apache.spark.sql.execution.datasources.jdbc.JDBCRDD.compute(JDBCRDD.scala:359)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{noformat}","from":"reporter","subject":"DataFrame filter Predicate push-down fails for Oracle Timestamp type columns"},{"body":"So there seems to something going on with the date format here. Could you check what query Spark SQL is sending to Oracle?","from":"developer"},{"body":"Date format as per the physical plan logged by Spark Dataframe: PushedFilters: [LessThan(TS,2016-11-17 19:42:01.057)]\n\nConfirmed the same from Oracle query logs as well: WHERE TS < '2016-11-17 19:42:01.057'","from":"developer"},{"body":"which format should be passed to oracle?","from":"developer"},{"body":"The date/timestamp format Oracle expects may vary from instance to instance based on the configuration of parameters NLS_TIMESTAMP_FORMAT and NLS_DATE_FORMAT . So ideal solution would be to use the to_timestamp and to_date Oracle functions. E.g. to_timestamp('12-01-2012 21:24:00', 'dd-mm-yyyy hh24:mi:ss')\n\nhttp://stackoverflow.com/questions/8855378/oracle-sql-timestamps-in-where-clause\nhttps://docs.oracle.com/cd/B19306_01/server.102/b14225/ch4datetime.htm#i1006312","from":"developer"},{"body":"The current spark Jdbc interface (JdbcDialect) cannot handle formats for database-dependent timestamps and just puts `Timestamp#toString` in where clauses (https://github.com/apache/spark/blob/master/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/jdbc/JDBCRDD.scala#L95). One solution is to transform this `Timestamp` into a database-specific timestamp format like https://github.com/apache/spark/compare/master...maropu:SPARK-18004#diff-5a29ce8f760092fb4a9c1f190cc2f61cR96","from":"developer"},{"body":"a silly but doable workaround: cast TimestampType row to LongType and compare the inner representation...\n\n{code:java}\nsrc.filter(col(\"timestamp_col\").cast(LongType).*(1000l) > condition.asInstanceOf[Timestamp].getTime)\n{code}","from":"developer"},{"body":"The right solution here is to make an explicit cast on the string representation of the timestamp value and not rely on implicit casting by the database\nANSI casting should work in nearly every RDBMS out there if the string is ISO 8601 format. Eg.\n{noformat}\nselect timestamp '2016-10-19 12:54:01.934';\n timestamp\n-------------------------\n 2016-10-19 12:54:01.934\n\n{noformat}","from":"developer"},{"body":"Here's the culprit, a private function that converts a scala value to a SQL literal. Note use of {{toString}} to format dates and timestamps. Seems like there should be some hook to a {{JDBCDialect}} that can handle vendor-specific syntax etc.\n\n{code:lang=java}\n /**\n * Converts value to SQL expression.\n */\n private def compileValue(value: Any): Any = value match {\n case stringValue: String => s\"'${escapeSql(stringValue)}'\"\n case timestampValue: Timestamp => \"'\" + timestampValue + \"'\"\n case dateValue: Date => \"'\" + dateValue + \"'\"\n case arrayValue: Array[Any] => arrayValue.map(compileValue).mkString(\", \")\n case _ => value\n }\n{code}","from":"developer"},{"body":"User 'SharpRay' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18404","from":"developer"},{"body":"The PR 18404 is closed. I will resend the PR later.","from":"developer"},{"body":"User 'SharpRay' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18411","from":"developer"},{"body":"User 'SharpRay' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18451","from":"developer"}],"created":"2016-10-19T07:43:08.000+0000","description":"DataFrame filter Predicate push-down fails for Oracle Timestamp type columns with Exception java.sql.SQLDataException: ORA-01861: literal does not match format string:\n\nJava source code (this code works fine for mysql & mssql databases) :\n{noformat}\n//DataFrame df = create a DataFrame over an Oracle table\ndf = df.filter(df.col(\"TS\").lt(new java.sql.Timestamp(System.currentTimeMillis())));\n\t\tdf.explain();\n\t\tdf.show();\n{noformat}\n\nLog statements with the Exception:\n{noformat}\nSchema: root\n |-- ID: string (nullable = false)\n |-- TS: timestamp (nullable = true)\n |-- DEVICE_ID: string (nullable = true)\n |-- REPLACEMENT: string (nullable = true)\n{noformat}\n\n{noformat}\n== Physical Plan ==\nFilter (TS#1 < 1476861841934000)\n+- Scan JDBCRelation(jdbc:oracle:thin:@10.0.0.111:1521:orcl,ORATABLE,[Lorg.apache.spark.Partition;@78c74647,{user=user, password=pwd, url=jdbc:oracle:thin:@10.0.0.111:1521:orcl, dbtable=ORATABLE, driver=oracle.jdbc.driver.OracleDriver})[ID#0,TS#1,DEVICE_ID#2,REPLACEMENT#3] PushedFilters: [LessThan(TS,2016-10-19 12:54:01.934)]\n2016-10-19 12:54:04,268 ERROR [Executor task launch worker-0] org.apache.spark.executor.Executor\nException in task 0.0 in stage 0.0 (TID 0)\n\njava.sql.SQLDataException: ORA-01861: literal does not match format string\n\n\tat oracle.jdbc.driver.T4CTTIoer.processError(T4CTTIoer.java:461)\n\tat oracle.jdbc.driver.T4CTTIoer.processError(T4CTTIoer.java:402)\n\tat oracle.jdbc.driver.T4C8Oall.processError(T4C8Oall.java:1065)\n\tat oracle.jdbc.driver.T4CTTIfun.receive(T4CTTIfun.java:681)\n\tat oracle.jdbc.driver.T4CTTIfun.doRPC(T4CTTIfun.java:256)\n\tat oracle.jdbc.driver.T4C8Oall.doOALL(T4C8Oall.java:577)\n\tat oracle.jdbc.driver.T4CPreparedStatement.doOall8(T4CPreparedStatement.java:239)\n\tat oracle.jdbc.driver.T4CPreparedStatement.doOall8(T4CPreparedStatement.java:75)\n\tat oracle.jdbc.driver.T4CPreparedStatement.executeForDescribe(T4CPreparedStatement.java:1043)\n\tat oracle.jdbc.driver.OracleStatement.executeMaybeDescribe(OracleStatement.java:1111)\n\tat oracle.jdbc.driver.OracleStatement.doExecuteWithTimeout(OracleStatement.java:1353)\n\tat oracle.jdbc.driver.OraclePreparedStatement.executeInternal(OraclePreparedStatement.java:4485)\n\tat oracle.jdbc.driver.OraclePreparedStatement.executeQuery(OraclePreparedStatement.java:4566)\n\tat oracle.jdbc.driver.OraclePreparedStatementWrapper.executeQuery(OraclePreparedStatementWrapper.java:5251)\n\tat org.apache.spark.sql.execution.datasources.jdbc.JDBCRDD$$anon$1.(JDBCRDD.scala:383)\n\tat org.apache.spark.sql.execution.datasources.jdbc.JDBCRDD.compute(JDBCRDD.scala:359)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:89)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{noformat}","issue_id":"13013426","key":"SPARK-18004","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-07-03T00:41:41.000+0000","role":"fixed_distractor","summary":"DataFrame filter Predicate push-down fails for Oracle Timestamp type columns"} {"case_id":"13013541","cluster":"DISTRACTOR-SPARK-18009","comments":[{"body":"Hi, \nI got the same error with spark 2.0.0 but only if enable spark.sql.thriftServer.incrementalCollect=true.\nDo you have this parameter enabled ?\n","created":"2016-10-25T12:19:13.835+0000"},{"body":"[~dkbiswal] Please fix it tonight. Thanks!","created":"2016-10-25T22:41:06.269+0000"},{"body":"Yes!\nIn my case, it's necessary option for integration with BI tools.","created":"2016-10-26T01:43:52.585+0000"},{"body":"[~smilegator][~jerryjung] [~martha.solarte] Thanks. I am testing a fix and should submit a PR for this soon.","created":"2016-10-26T04:08:27.550+0000"},{"body":"User 'dilipbiswal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15642","created":"2016-10-26T10:29:04.509+0000"},{"body":"Will be the fix available for 2.0.0 ?","created":"2016-10-26T11:19:46.537+0000"},{"body":"[~martha.solarte] Not sure\n[~smilegator] Sean, do we back port to 2.0.0 any more ?","created":"2016-10-26T18:42:13.575+0000"},{"body":"We try to fix it in 2.0.2. To get the fix, you have to use the new release. ","created":"2016-10-26T23:41:45.489+0000"},{"body":"Issue resolved by pull request 15642\n[https://github.com/apache/spark/pull/15642]","created":"2016-10-27T05:16:49.523+0000"}],"conversations":[{"body":"After deploy spark thrift server on YARN, then I tried to execute from the beeline following command.\n> show databases;\nI've got this error message. \n{quote}\nbeeline> !connect jdbc:hive2://localhost:10000 a a\nConnecting to jdbc:hive2://localhost:10000\n16/10/19 22:50:18 INFO Utils: Supplied authorities: localhost:10000\n16/10/19 22:50:18 INFO Utils: Resolved authority: localhost:10000\n16/10/19 22:50:18 INFO HiveConnection: Will try to open client transport with JDBC Uri: jdbc:hive2://localhost:10000\nConnected to: Spark SQL (version 2.0.1)\nDriver: Hive JDBC (version 1.2.1.spark2)\nTransaction isolation: TRANSACTION_REPEATABLE_READ\n0: jdbc:hive2://localhost:10000> show databases;\njava.lang.IllegalStateException: Can't overwrite cause with java.lang.ClassCastException: org.apache.spark.sql.catalyst.expressions.GenericInternalRow cannot be cast to org.apache.spark.sql.catalyst.expressions.UnsafeRow\n\tat java.lang.Throwable.initCause(Throwable.java:456)\n\tat org.apache.hive.service.cli.HiveSQLException.toStackTrace(HiveSQLException.java:236)\n\tat org.apache.hive.service.cli.HiveSQLException.toStackTrace(HiveSQLException.java:236)\n\tat org.apache.hive.service.cli.HiveSQLException.toCause(HiveSQLException.java:197)\n\tat org.apache.hive.service.cli.HiveSQLException.(HiveSQLException.java:108)\n\tat org.apache.hive.jdbc.Utils.verifySuccess(Utils.java:256)\n\tat org.apache.hive.jdbc.Utils.verifySuccessWithInfo(Utils.java:242)\n\tat org.apache.hive.jdbc.HiveQueryResultSet.next(HiveQueryResultSet.java:365)\n\tat org.apache.hive.beeline.BufferedRows.(BufferedRows.java:42)\n\tat org.apache.hive.beeline.BeeLine.print(BeeLine.java:1794)\n\tat org.apache.hive.beeline.Commands.execute(Commands.java:860)\n\tat org.apache.hive.beeline.Commands.sql(Commands.java:713)\n\tat org.apache.hive.beeline.BeeLine.dispatch(BeeLine.java:973)\n\tat org.apache.hive.beeline.BeeLine.execute(BeeLine.java:813)\n\tat org.apache.hive.beeline.BeeLine.begin(BeeLine.java:771)\n\tat org.apache.hive.beeline.BeeLine.mainWithInputRedirection(BeeLine.java:484)\n\tat org.apache.hive.beeline.BeeLine.main(BeeLine.java:467)\nCaused by: org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 669.0 failed 4 times, most recent failure: Lost task 0.3 in stage 669.0 (TID 3519, edw-014-22): java.lang.ClassCastException: org.apache.spark.sql.catalyst.expressions.GenericInternalRow cannot be cast to org.apache.spark.sql.catalyst.expressions.UnsafeRow\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:247)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:240)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:86)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n\nDriver stacktrace:\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:526)\n\tat org.apache.hive.service.cli.HiveSQLException.newInstance(HiveSQLException.java:244)\n\tat org.apache.hive.service.cli.HiveSQLException.toStackTrace(HiveSQLException.java:210)\n\t... 15 more\nError: Error retrieving next row (state=,code=0)\n{quote}\n\n\"add jar\" command also same error occurred.\n{quote}\nadd jar ~/udf.jar\njava.lang.IllegalStateException: Can’t overwrite cause with java.lang.ClassCastException: org.apache.spark.sql.catalyst.expressions.GenericInternalRow cannot be cast to org.apache.spark.sql.catalyst.expressions.UnsafeRow\nat java.lang.Throwable.initCause (Throwable.java:456)\n…\n{quote}","from":"reporter","subject":"Spark 2.0.1 SQL Thrift Error"},{"body":"Hi, \nI got the same error with spark 2.0.0 but only if enable spark.sql.thriftServer.incrementalCollect=true.\nDo you have this parameter enabled ?\n","from":"developer"},{"body":"[~dkbiswal] Please fix it tonight. Thanks!","from":"developer"},{"body":"Yes!\nIn my case, it's necessary option for integration with BI tools.","from":"developer"},{"body":"[~smilegator][~jerryjung] [~martha.solarte] Thanks. I am testing a fix and should submit a PR for this soon.","from":"developer"},{"body":"User 'dilipbiswal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15642","from":"developer"},{"body":"Will be the fix available for 2.0.0 ?","from":"developer"},{"body":"[~martha.solarte] Not sure\n[~smilegator] Sean, do we back port to 2.0.0 any more ?","from":"developer"},{"body":"We try to fix it in 2.0.2. To get the fix, you have to use the new release. ","from":"developer"},{"body":"Issue resolved by pull request 15642\n[https://github.com/apache/spark/pull/15642]","from":"developer"}],"created":"2016-10-19T13:57:24.000+0000","description":"After deploy spark thrift server on YARN, then I tried to execute from the beeline following command.\n> show databases;\nI've got this error message. \n{quote}\nbeeline> !connect jdbc:hive2://localhost:10000 a a\nConnecting to jdbc:hive2://localhost:10000\n16/10/19 22:50:18 INFO Utils: Supplied authorities: localhost:10000\n16/10/19 22:50:18 INFO Utils: Resolved authority: localhost:10000\n16/10/19 22:50:18 INFO HiveConnection: Will try to open client transport with JDBC Uri: jdbc:hive2://localhost:10000\nConnected to: Spark SQL (version 2.0.1)\nDriver: Hive JDBC (version 1.2.1.spark2)\nTransaction isolation: TRANSACTION_REPEATABLE_READ\n0: jdbc:hive2://localhost:10000> show databases;\njava.lang.IllegalStateException: Can't overwrite cause with java.lang.ClassCastException: org.apache.spark.sql.catalyst.expressions.GenericInternalRow cannot be cast to org.apache.spark.sql.catalyst.expressions.UnsafeRow\n\tat java.lang.Throwable.initCause(Throwable.java:456)\n\tat org.apache.hive.service.cli.HiveSQLException.toStackTrace(HiveSQLException.java:236)\n\tat org.apache.hive.service.cli.HiveSQLException.toStackTrace(HiveSQLException.java:236)\n\tat org.apache.hive.service.cli.HiveSQLException.toCause(HiveSQLException.java:197)\n\tat org.apache.hive.service.cli.HiveSQLException.(HiveSQLException.java:108)\n\tat org.apache.hive.jdbc.Utils.verifySuccess(Utils.java:256)\n\tat org.apache.hive.jdbc.Utils.verifySuccessWithInfo(Utils.java:242)\n\tat org.apache.hive.jdbc.HiveQueryResultSet.next(HiveQueryResultSet.java:365)\n\tat org.apache.hive.beeline.BufferedRows.(BufferedRows.java:42)\n\tat org.apache.hive.beeline.BeeLine.print(BeeLine.java:1794)\n\tat org.apache.hive.beeline.Commands.execute(Commands.java:860)\n\tat org.apache.hive.beeline.Commands.sql(Commands.java:713)\n\tat org.apache.hive.beeline.BeeLine.dispatch(BeeLine.java:973)\n\tat org.apache.hive.beeline.BeeLine.execute(BeeLine.java:813)\n\tat org.apache.hive.beeline.BeeLine.begin(BeeLine.java:771)\n\tat org.apache.hive.beeline.BeeLine.mainWithInputRedirection(BeeLine.java:484)\n\tat org.apache.hive.beeline.BeeLine.main(BeeLine.java:467)\nCaused by: org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 669.0 failed 4 times, most recent failure: Lost task 0.3 in stage 669.0 (TID 3519, edw-014-22): java.lang.ClassCastException: org.apache.spark.sql.catalyst.expressions.GenericInternalRow cannot be cast to org.apache.spark.sql.catalyst.expressions.UnsafeRow\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:247)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:240)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:86)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n\nDriver stacktrace:\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:526)\n\tat org.apache.hive.service.cli.HiveSQLException.newInstance(HiveSQLException.java:244)\n\tat org.apache.hive.service.cli.HiveSQLException.toStackTrace(HiveSQLException.java:210)\n\t... 15 more\nError: Error retrieving next row (state=,code=0)\n{quote}\n\n\"add jar\" command also same error occurred.\n{quote}\nadd jar ~/udf.jar\njava.lang.IllegalStateException: Can’t overwrite cause with java.lang.ClassCastException: org.apache.spark.sql.catalyst.expressions.GenericInternalRow cannot be cast to org.apache.spark.sql.catalyst.expressions.UnsafeRow\nat java.lang.Throwable.initCause (Throwable.java:456)\n…\n{quote}","issue_id":"13013541","key":"SPARK-18009","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-10-27T05:16:49.000+0000","role":"fixed_distractor","summary":"Spark 2.0.1 SQL Thrift Error"} {"case_id":"13013718","cluster":"DISTRACTOR-SPARK-18020","comments":[{"body":"I'm actually experiencing the exact same problem.","created":"2016-11-09T10:14:19.762+0000"},{"body":"I'm currently looking into this issue.","created":"2016-12-03T10:56:22.953+0000"},{"body":"I tried to checkpoint with SHARD_END by `IRecordProcessorCheckpointer#checkpoint(ExtendedSequenceNumber.SHARD_END.toString())`, but I got the IllegalArgumentException exception with messages like \"Sequence number must be numeric, but was SHARD_END\". Since I'm not sure this is a expected behaviour, I ask this to aws guys here https://forums.aws.amazon.com/thread.jspa?threadID=244218. ","created":"2016-12-04T11:16:58.561+0000"},{"body":"I tried to make a workaround patch to fix this issue (https://github.com/apache/spark/compare/master...maropu:SPARK-18020) and I manually checked this issue resolved.\nBut, I'm not sure this workaround is acceptable, so I continue to look for other better approaches.","created":"2016-12-04T11:31:02.774+0000"},{"body":"User 'maropu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16213","created":"2016-12-08T11:16:05.346+0000"},{"body":"Resolved by https://github.com/apache/spark/pull/16213","created":"2017-01-26T01:42:24.137+0000"}],"conversations":[{"body":"When a kinesis shard is split or combined and the old shard ends, the Amazon Kinesis Client library [calls IRecordProcessor.shutdown|https://github.com/awslabs/amazon-kinesis-client/blob/v1.7.0/src/main/java/com/amazonaws/services/kinesis/clientlibrary/lib/worker/ShutdownTask.java#L100] and expects that {{IRecordProcessor.shutdown}} must checkpoint the sequence number {{ExtendedSequenceNumber.SHARD_END}} before returning. Unfortunately, spark’s [KinesisRecordProcessor|https://github.com/apache/spark/blob/v2.0.1/external/kinesis-asl/src/main/scala/org/apache/spark/streaming/kinesis/KinesisRecordProcessor.scala] sometimes does not checkpoint SHARD_END. This results in an error message, and spark is then blocked indefinitely from processing any items from the child shards.\n\nThis issue has also been raised on StackOverflow: [resharding while spark running on kinesis stream|http://stackoverflow.com/questions/38898691/resharding-while-spark-running-on-kinesis-stream]\n\nException that is logged:\n{code}\n16/10/19 19:37:49 ERROR worker.ShutdownTask: Application exception. \njava.lang.IllegalArgumentException: Application didn't checkpoint at end of shard shardId-000000000030\n at com.amazonaws.services.kinesis.clientlibrary.lib.worker.ShutdownTask.call(ShutdownTask.java:106)\n at com.amazonaws.services.kinesis.clientlibrary.lib.worker.MetricsCollectingTaskDecorator.call(MetricsCollectingTaskDecorator.java:49)\n at com.amazonaws.services.kinesis.clientlibrary.lib.worker.MetricsCollectingTaskDecorator.call(MetricsCollectingTaskDecorator.java:24)\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\nCommand used to split shard:\n\n{code}\naws kinesis --region us-west-1 split-shard --stream-name my-stream --shard-to-split shardId-000000000030 --new-starting-hash-key 5316911983139663491615228241121378303\n{code}\n\nAfter the spark-streaming job has hung, examining the DynamoDB table indicates that the parent shard processor has not reached {{ExtendedSequenceNumber.SHARD_END}} and the child shards are still at {{ExtendedSequenceNumber.TRIM_HORIZON}} waiting for the parent to finish:\n\n{code}\naws kinesis --region us-west-1 describe-stream --stream-name my-stream\n{\n \"StreamDescription\": {\n \"RetentionPeriodHours\": 24, \n \"StreamName\": \"my-stream\", \n \"Shards\": [\n {\n \"ShardId\": \"shardId-000000000030\", \n \"HashKeyRange\": {\n \"EndingHashKey\": \"10633823966279326983230456482242756606\", \n \"StartingHashKey\": \"0\"\n },\n ...\n }, \n {\n \"ShardId\": \"shardId-000000000062\", \n \"HashKeyRange\": {\n \"EndingHashKey\": \"5316911983139663491615228241121378302\", \n \"StartingHashKey\": \"0\"\n }, \n \"ParentShardId\": \"shardId-000000000030\", \n \"SequenceNumberRange\": {\n \"StartingSequenceNumber\": \"49566806087883755242230188435465744452396445937434624994\"\n }\n }, \n {\n \"ShardId\": \"shardId-000000000063\", \n \"HashKeyRange\": {\n \"EndingHashKey\": \"10633823966279326983230456482242756606\", \n \"StartingHashKey\": \"5316911983139663491615228241121378303\"\n }, \n \"ParentShardId\": \"shardId-000000000030\", \n \"SequenceNumberRange\": {\n \"StartingSequenceNumber\": \"49566806087906055987428719058607280170669094298940605426\"\n }\n },\n ...\n ],\n \"StreamStatus\": \"ACTIVE\"\n }\n}\n\naws dynamodb --region us-west-1 scan --table-name my-processor\n{\n \"Items\": [\n {\n \"leaseOwner\": {\n \"S\": \"localhost:fd385c95-5d19-4678-926f-b6d5f5503cbe\"\n }, \n \"leaseCounter\": {\n \"N\": \"49318\"\n }, \n \"ownerSwitchesSinceCheckpoint\": {\n \"N\": \"62\"\n }, \n \"checkpointSubSequenceNumber\": {\n \"N\": \"0\"\n }, \n \"checkpoint\": {\n \"S\": \"49566573572821264975247582655142547856950135436343247330\"\n }, \n \"parentShardId\": {\n \"SS\": [\n \"shardId-000000000014\"\n ]\n }, \n \"leaseKey\": {\n \"S\": \"shardId-000000000030\"\n }\n }, \n {\n \"leaseOwner\": {\n \"S\": \"localhost:ca44dc83-2580-4bf3-903f-e7ccc8a3ab02\"\n }, \n \"leaseCounter\": {\n \"N\": \"25439\"\n }, \n \"ownerSwitchesSinceCheckpoint\": {\n \"N\": \"69\"\n }, \n \"checkpointSubSequenceNumber\": {\n \"N\": \"0\"\n }, \n \"checkpoint\": {\n \"S\": \"TRIM_HORIZON\"\n }, \n \"parentShardId\": {\n \"SS\": [\n \"shardId-000000000030\"\n ]\n }, \n \"leaseKey\": {\n \"S\": \"shardId-000000000062\"\n }\n }, \n {\n \"leaseOwner\": {\n \"S\": \"localhost:94bf603f-780b-4121-87a4-bdf501723f83\"\n }, \n \"leaseCounter\": {\n \"N\": \"25443\"\n }, \n \"ownerSwitchesSinceCheckpoint\": {\n \"N\": \"59\"\n }, \n \"checkpointSubSequenceNumber\": {\n \"N\": \"0\"\n }, \n \"checkpoint\": {\n \"S\": \"TRIM_HORIZON\"\n }, \n \"parentShardId\": {\n \"SS\": [\n \"shardId-000000000030\"\n ]\n }, \n \"leaseKey\": {\n \"S\": \"shardId-000000000063\"\n }\n },\n ...\n ]\n}\n{code}\n\nWorkaround: I manually edited the DynamoDB table to delete the checkpoints for the parent shards. The child shards were then able to begin processing. I’m not sure whether this resulted in a few items being lost though.","from":"reporter","subject":"Kinesis receiver does not snapshot when shard completes"},{"body":"I'm actually experiencing the exact same problem.","from":"developer"},{"body":"I'm currently looking into this issue.","from":"developer"},{"body":"I tried to checkpoint with SHARD_END by `IRecordProcessorCheckpointer#checkpoint(ExtendedSequenceNumber.SHARD_END.toString())`, but I got the IllegalArgumentException exception with messages like \"Sequence number must be numeric, but was SHARD_END\". Since I'm not sure this is a expected behaviour, I ask this to aws guys here https://forums.aws.amazon.com/thread.jspa?threadID=244218. ","from":"developer"},{"body":"I tried to make a workaround patch to fix this issue (https://github.com/apache/spark/compare/master...maropu:SPARK-18020) and I manually checked this issue resolved.\nBut, I'm not sure this workaround is acceptable, so I continue to look for other better approaches.","from":"developer"},{"body":"User 'maropu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16213","from":"developer"},{"body":"Resolved by https://github.com/apache/spark/pull/16213","from":"developer"}],"created":"2016-10-20T00:02:39.000+0000","description":"When a kinesis shard is split or combined and the old shard ends, the Amazon Kinesis Client library [calls IRecordProcessor.shutdown|https://github.com/awslabs/amazon-kinesis-client/blob/v1.7.0/src/main/java/com/amazonaws/services/kinesis/clientlibrary/lib/worker/ShutdownTask.java#L100] and expects that {{IRecordProcessor.shutdown}} must checkpoint the sequence number {{ExtendedSequenceNumber.SHARD_END}} before returning. Unfortunately, spark’s [KinesisRecordProcessor|https://github.com/apache/spark/blob/v2.0.1/external/kinesis-asl/src/main/scala/org/apache/spark/streaming/kinesis/KinesisRecordProcessor.scala] sometimes does not checkpoint SHARD_END. This results in an error message, and spark is then blocked indefinitely from processing any items from the child shards.\n\nThis issue has also been raised on StackOverflow: [resharding while spark running on kinesis stream|http://stackoverflow.com/questions/38898691/resharding-while-spark-running-on-kinesis-stream]\n\nException that is logged:\n{code}\n16/10/19 19:37:49 ERROR worker.ShutdownTask: Application exception. \njava.lang.IllegalArgumentException: Application didn't checkpoint at end of shard shardId-000000000030\n at com.amazonaws.services.kinesis.clientlibrary.lib.worker.ShutdownTask.call(ShutdownTask.java:106)\n at com.amazonaws.services.kinesis.clientlibrary.lib.worker.MetricsCollectingTaskDecorator.call(MetricsCollectingTaskDecorator.java:49)\n at com.amazonaws.services.kinesis.clientlibrary.lib.worker.MetricsCollectingTaskDecorator.call(MetricsCollectingTaskDecorator.java:24)\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}\n\nCommand used to split shard:\n\n{code}\naws kinesis --region us-west-1 split-shard --stream-name my-stream --shard-to-split shardId-000000000030 --new-starting-hash-key 5316911983139663491615228241121378303\n{code}\n\nAfter the spark-streaming job has hung, examining the DynamoDB table indicates that the parent shard processor has not reached {{ExtendedSequenceNumber.SHARD_END}} and the child shards are still at {{ExtendedSequenceNumber.TRIM_HORIZON}} waiting for the parent to finish:\n\n{code}\naws kinesis --region us-west-1 describe-stream --stream-name my-stream\n{\n \"StreamDescription\": {\n \"RetentionPeriodHours\": 24, \n \"StreamName\": \"my-stream\", \n \"Shards\": [\n {\n \"ShardId\": \"shardId-000000000030\", \n \"HashKeyRange\": {\n \"EndingHashKey\": \"10633823966279326983230456482242756606\", \n \"StartingHashKey\": \"0\"\n },\n ...\n }, \n {\n \"ShardId\": \"shardId-000000000062\", \n \"HashKeyRange\": {\n \"EndingHashKey\": \"5316911983139663491615228241121378302\", \n \"StartingHashKey\": \"0\"\n }, \n \"ParentShardId\": \"shardId-000000000030\", \n \"SequenceNumberRange\": {\n \"StartingSequenceNumber\": \"49566806087883755242230188435465744452396445937434624994\"\n }\n }, \n {\n \"ShardId\": \"shardId-000000000063\", \n \"HashKeyRange\": {\n \"EndingHashKey\": \"10633823966279326983230456482242756606\", \n \"StartingHashKey\": \"5316911983139663491615228241121378303\"\n }, \n \"ParentShardId\": \"shardId-000000000030\", \n \"SequenceNumberRange\": {\n \"StartingSequenceNumber\": \"49566806087906055987428719058607280170669094298940605426\"\n }\n },\n ...\n ],\n \"StreamStatus\": \"ACTIVE\"\n }\n}\n\naws dynamodb --region us-west-1 scan --table-name my-processor\n{\n \"Items\": [\n {\n \"leaseOwner\": {\n \"S\": \"localhost:fd385c95-5d19-4678-926f-b6d5f5503cbe\"\n }, \n \"leaseCounter\": {\n \"N\": \"49318\"\n }, \n \"ownerSwitchesSinceCheckpoint\": {\n \"N\": \"62\"\n }, \n \"checkpointSubSequenceNumber\": {\n \"N\": \"0\"\n }, \n \"checkpoint\": {\n \"S\": \"49566573572821264975247582655142547856950135436343247330\"\n }, \n \"parentShardId\": {\n \"SS\": [\n \"shardId-000000000014\"\n ]\n }, \n \"leaseKey\": {\n \"S\": \"shardId-000000000030\"\n }\n }, \n {\n \"leaseOwner\": {\n \"S\": \"localhost:ca44dc83-2580-4bf3-903f-e7ccc8a3ab02\"\n }, \n \"leaseCounter\": {\n \"N\": \"25439\"\n }, \n \"ownerSwitchesSinceCheckpoint\": {\n \"N\": \"69\"\n }, \n \"checkpointSubSequenceNumber\": {\n \"N\": \"0\"\n }, \n \"checkpoint\": {\n \"S\": \"TRIM_HORIZON\"\n }, \n \"parentShardId\": {\n \"SS\": [\n \"shardId-000000000030\"\n ]\n }, \n \"leaseKey\": {\n \"S\": \"shardId-000000000062\"\n }\n }, \n {\n \"leaseOwner\": {\n \"S\": \"localhost:94bf603f-780b-4121-87a4-bdf501723f83\"\n }, \n \"leaseCounter\": {\n \"N\": \"25443\"\n }, \n \"ownerSwitchesSinceCheckpoint\": {\n \"N\": \"59\"\n }, \n \"checkpointSubSequenceNumber\": {\n \"N\": \"0\"\n }, \n \"checkpoint\": {\n \"S\": \"TRIM_HORIZON\"\n }, \n \"parentShardId\": {\n \"SS\": [\n \"shardId-000000000030\"\n ]\n }, \n \"leaseKey\": {\n \"S\": \"shardId-000000000063\"\n }\n },\n ...\n ]\n}\n{code}\n\nWorkaround: I manually edited the DynamoDB table to delete the checkpoints for the parent shards. The child shards were then able to begin processing. I’m not sure whether this resulted in a few items being lost though.","issue_id":"13013718","key":"SPARK-18020","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-01-26T01:42:24.000+0000","role":"fixed_distractor","summary":"Kinesis receiver does not snapshot when shard completes"} {"case_id":"13014750","cluster":"DISTRACTOR-SPARK-18076","comments":[{"body":"User 'srowen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15610","created":"2016-10-24T14:58:05.720+0000"},{"body":"Issue resolved by pull request 15610\n[https://github.com/apache/spark/pull/15610]","created":"2016-11-02T09:39:33.583+0000"}],"conversations":[{"body":"Many parts of the code use {{DateFormat}} and {{NumberFormat}} instances. Although the behavior of these format is mostly determined by things like format strings, the exact behavior can vary according to the platform's default locale. Although the locale defaults to \"en\", it can be set to something else by env variables. And if it does, it can cause the same code to succeed or fail based just on locale:\n\n{code}\nimport java.text._\nimport java.util._\n\ndef parse(s: String, l: Locale) = new SimpleDateFormat(\"yyyyMMMdd\", l).parse(s)\n\nparse(\"1989Dec31\", Locale.US)\nSun Dec 31 00:00:00 GMT 1989\n\nparse(\"1989Dec31\", Locale.UK)\nSun Dec 31 00:00:00 GMT 1989\n\nparse(\"1989Dec31\", Locale.CHINA)\njava.text.ParseException: Unparseable date: \"1989Dec31\"\n at java.text.DateFormat.parse(DateFormat.java:366)\n at .parse(:18)\n ... 32 elided\n\nparse(\"1989Dec31\", Locale.GERMANY)\njava.text.ParseException: Unparseable date: \"1989Dec31\"\n at java.text.DateFormat.parse(DateFormat.java:366)\n at .parse(:18)\n ... 32 elided\n{code}\n\nWhere not otherwise specified, I believe all instances in the code should default to some fixed value, and that should probably be {{Locale.US}}. This matches the JVM's default, and specifies both language (\"en\") and region (\"US\") to remove ambiguity. This most closely matches what the current code behavior would be (unless default locale was changed), because it will currently default to \"en\".\n\nThis affects SQL date/time functions. At the moment, the only SQL function that lets the user specify language/country is \"sentences\", which is consistent with Hive.\n\nIt affects dates passed in the JSON API. \n\nIt affects some strings rendered in the UI, potentially. Although this isn't a correctness issue, there may be an argument for not letting that vary (?)\n\nIt affects a bunch of instances where dates are formatted into strings for things like IDs or file names, which is far less likely to cause a problem, but worth making consistent.\n\nThe other occurrences are in tests.\n\n\nThe downside to this change is also its upside: the behavior doesn't depend on default JVM locale, but, also can't be affected by the default JVM locale. For example, if you wanted to parse some dates in a way that depended on an non-US locale (not just the format string) then it would no longer be possible. There's no means of specifying this, for example, in SQL functions for parsing dates. However, controlling this by globally changing the locale isn't exactly great either.\n\nThe purpose of this change is to make the current default behavior deterministic and fixed. PR coming.\n\nCC [~hyukjin.kwon]","from":"reporter","subject":"Fix default Locale used in DateFormat, NumberFormat to Locale.US"},{"body":"User 'srowen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15610","from":"developer"},{"body":"Issue resolved by pull request 15610\n[https://github.com/apache/spark/pull/15610]","from":"developer"}],"created":"2016-10-24T14:48:44.000+0000","description":"Many parts of the code use {{DateFormat}} and {{NumberFormat}} instances. Although the behavior of these format is mostly determined by things like format strings, the exact behavior can vary according to the platform's default locale. Although the locale defaults to \"en\", it can be set to something else by env variables. And if it does, it can cause the same code to succeed or fail based just on locale:\n\n{code}\nimport java.text._\nimport java.util._\n\ndef parse(s: String, l: Locale) = new SimpleDateFormat(\"yyyyMMMdd\", l).parse(s)\n\nparse(\"1989Dec31\", Locale.US)\nSun Dec 31 00:00:00 GMT 1989\n\nparse(\"1989Dec31\", Locale.UK)\nSun Dec 31 00:00:00 GMT 1989\n\nparse(\"1989Dec31\", Locale.CHINA)\njava.text.ParseException: Unparseable date: \"1989Dec31\"\n at java.text.DateFormat.parse(DateFormat.java:366)\n at .parse(:18)\n ... 32 elided\n\nparse(\"1989Dec31\", Locale.GERMANY)\njava.text.ParseException: Unparseable date: \"1989Dec31\"\n at java.text.DateFormat.parse(DateFormat.java:366)\n at .parse(:18)\n ... 32 elided\n{code}\n\nWhere not otherwise specified, I believe all instances in the code should default to some fixed value, and that should probably be {{Locale.US}}. This matches the JVM's default, and specifies both language (\"en\") and region (\"US\") to remove ambiguity. This most closely matches what the current code behavior would be (unless default locale was changed), because it will currently default to \"en\".\n\nThis affects SQL date/time functions. At the moment, the only SQL function that lets the user specify language/country is \"sentences\", which is consistent with Hive.\n\nIt affects dates passed in the JSON API. \n\nIt affects some strings rendered in the UI, potentially. Although this isn't a correctness issue, there may be an argument for not letting that vary (?)\n\nIt affects a bunch of instances where dates are formatted into strings for things like IDs or file names, which is far less likely to cause a problem, but worth making consistent.\n\nThe other occurrences are in tests.\n\n\nThe downside to this change is also its upside: the behavior doesn't depend on default JVM locale, but, also can't be affected by the default JVM locale. For example, if you wanted to parse some dates in a way that depended on an non-US locale (not just the format string) then it would no longer be possible. There's no means of specifying this, for example, in SQL functions for parsing dates. However, controlling this by globally changing the locale isn't exactly great either.\n\nThe purpose of this change is to make the current default behavior deterministic and fixed. PR coming.\n\nCC [~hyukjin.kwon]","issue_id":"13014750","key":"SPARK-18076","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-11-02T09:39:32.000+0000","role":"fixed_distractor","summary":"Fix default Locale used in DateFormat, NumberFormat to Locale.US"} {"case_id":"13015605","cluster":"DISTRACTOR-SPARK-18125","comments":[{"body":"Same here:\n\nAfter groupByKey( _._1).reduceGroups((a,b) => (a._1, a._2 ++ a._2)).map(_._2):\n\njava.util.concurrent.ExecutionException: java.lang.Exception: failed to compile: org.codehaus.commons.compiler.CompileException: File 'generated.java', Line 31, Column 69: Unknown variable or type \"value4\"\n\nSpark Version 2.0.1","created":"2016-10-27T08:20:24.633+0000"},{"body":"Move to Priority Critical unless a workaround is identified. This is a very basic functionality.","created":"2016-10-27T20:22:48.799+0000"},{"body":"Could one of you provide a reproducible example.","created":"2016-10-27T22:37:17.013+0000"},{"body":"I tried something like this on master and on branch-2.0:\n{noformat}\nval ds = spark.range(10000).select($\"id\" % 100 as \"grp_id\", array($\"id\")).as[(Long, Seq[Long])] \nds.groupByKey(_._1).reduceGroups((a, b) => (a._1, a._2 ++ b._2)).map(_._2)\n{noformat}","created":"2016-10-27T23:06:19.780+0000"},{"body":"Try this in spark-shell:\n\ncase class Route(src: String, dest: String, cost: Int)\ncase class GroupedRoutes(src: String, dest: String, routes: Seq[Route])\n\nval ds = sc.parallelize(Array(\n Route(\"a\", \"b\", 1),\n Route(\"a\", \"b\", 2),\n Route(\"a\", \"c\", 2),\n Route(\"a\", \"d\", 10),\n Route(\"b\", \"a\", 1),\n Route(\"b\", \"a\", 5),\n Route(\"b\", \"c\", 6))\n ).toDF.as[Route]\n\nval grped = ds.map(r => GroupedRoutes(r.src, r.dest, Seq(r)))\n .groupByKey(r => (r.src, r.dest))\n .reduceGroups { (g1: GroupedRoutes, g2: GroupedRoutes) =>\n GroupedRoutes(g1.src, g1.dest, g1.routes ++ g2.routes)\n }.map(_._2)\n\nSame thing works fine in 2.0.0\n\nOn Thu, Oct 27, 2016 at 4:06 PM, Herman van Hovell (JIRA) \n\n\n\n\n-- \nRegards,\nRay\n","created":"2016-10-27T23:42:40.893+0000"},{"body":"I confirmed this code can reproduce on 2.0.1.\nThis problem occurs due to the similar reason in SPARK-18147\n\nTo call {{ctx.splitExpression}} in {{createStruct.doGenCode}} make *a variable* inaccesssible by splitting the original one function into multiple functions.\nSPARK-14793 seems to have introduced {{ctx.splitExpression}} here to fix other issues on April. Since Ray says this code works in 2.0.0, other changes may introduce this issue.","created":"2016-10-28T10:26:06.733+0000"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15693","created":"2016-10-31T09:09:06.025+0000"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15796","created":"2016-11-07T15:21:05.836+0000"}],"conversations":[{"body":"Code logic looks like this:\n{noformat}\n .groupByKey\n .reduceGroups\n .map(_._2)\n{noformat}\nWorks fine with 2.0.0.\n\n2.0.1 error Message: \n{noformat}\nCaused by: java.util.concurrent.ExecutionException: java.lang.Exception: failed to compile: org.codehaus.commons.compiler.CompileException: File 'generated.java', Line 206, Column 123: Unknown variable or type \"value4\"\n/* 001 */ public java.lang.Object generate(Object[] references) {\n/* 002 */ return new SpecificMutableProjection(references);\n/* 003 */ }\n/* 004 */\n/* 005 */ class SpecificMutableProjection extends org.apache.spark.sql.catalyst.expressions.codegen.BaseMutableProjection {\n/* 006 */\n/* 007 */ private Object[] references;\n/* 008 */ private MutableRow mutableRow;\n/* 009 */ private Object[] values;\n/* 010 */ private java.lang.String errMsg;\n/* 011 */ private java.lang.String errMsg1;\n/* 012 */ private boolean MapObjects_loopIsNull1;\n/* 013 */ private io.mistnet.analytics.lib.ConnLog MapObjects_loopValue0;\n/* 014 */ private java.lang.String errMsg2;\n/* 015 */ private Object[] values1;\n/* 016 */ private boolean MapObjects_loopIsNull3;\n/* 017 */ private java.lang.String MapObjects_loopValue2;\n/* 018 */ private boolean isNull_0;\n/* 019 */ private boolean value_0;\n/* 020 */ private boolean isNull_1;\n/* 021 */ private InternalRow value_1;\n/* 022 */\n/* 023 */ private void apply_4(InternalRow i) {\n/* 024 */\n/* 025 */ boolean isNull52 = MapObjects_loopIsNull1;\n/* 026 */ final double value52 = isNull52 ? -1.0 : MapObjects_loopValue0.ts();\n/* 027 */ if (isNull52) {\n/* 028 */ values1[8] = null;\n/* 029 */ } else {\n/* 030 */ values1[8] = value52;\n/* 031 */ }\n/* 032 */ boolean isNull54 = MapObjects_loopIsNull1;\n/* 033 */ final java.lang.String value54 = isNull54 ? null : (java.lang.String) MapObjects_loopValue0.uid();\n/* 034 */ isNull54 = value54 == null;\n/* 035 */ boolean isNull53 = isNull54;\n/* 036 */ final UTF8String value53 = isNull53 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value54);\n/* 037 */ isNull53 = value53 == null;\n/* 038 */ if (isNull53) {\n/* 039 */ values1[9] = null;\n/* 040 */ } else {\n/* 041 */ values1[9] = value53;\n/* 042 */ }\n/* 043 */ boolean isNull56 = MapObjects_loopIsNull1;\n/* 044 */ final java.lang.String value56 = isNull56 ? null : (java.lang.String) MapObjects_loopValue0.src();\n/* 045 */ isNull56 = value56 == null;\n/* 046 */ boolean isNull55 = isNull56;\n/* 047 */ final UTF8String value55 = isNull55 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value56);\n/* 048 */ isNull55 = value55 == null;\n/* 049 */ if (isNull55) {\n/* 050 */ values1[10] = null;\n/* 051 */ } else {\n/* 052 */ values1[10] = value55;\n/* 053 */ }\n/* 054 */ }\n/* 055 */\n/* 056 */\n/* 057 */ private void apply_7(InternalRow i) {\n/* 058 */\n/* 059 */ boolean isNull69 = MapObjects_loopIsNull1;\n/* 060 */ final scala.Option value69 = isNull69 ? null : (scala.Option) MapObjects_loopValue0.orig_bytes();\n/* 061 */ isNull69 = value69 == null;\n/* 062 */\n/* 063 */ final boolean isNull68 = isNull69 || value69.isEmpty();\n/* 064 */ long value68 = isNull68 ?\n/* 065 */ -1L : (Long) value69.get();\n/* 066 */ if (isNull68) {\n/* 067 */ values1[17] = null;\n/* 068 */ } else {\n/* 069 */ values1[17] = value68;\n/* 070 */ }\n/* 071 */ boolean isNull71 = MapObjects_loopIsNull1;\n/* 072 */ final scala.Option value71 = isNull71 ? null : (scala.Option) MapObjects_loopValue0.resp_bytes();\n/* 073 */ isNull71 = value71 == null;\n/* 074 */\n/* 075 */ final boolean isNull70 = isNull71 || value71.isEmpty();\n/* 076 */ long value70 = isNull70 ?\n/* 077 */ -1L : (Long) value71.get();\n/* 078 */ if (isNull70) {\n/* 079 */ values1[18] = null;\n/* 080 */ } else {\n/* 081 */ values1[18] = value70;\n/* 082 */ }\n/* 083 */ boolean isNull74 = MapObjects_loopIsNull1;\n/* 084 */ final scala.Option value74 = isNull74 ? null : (scala.Option) MapObjects_loopValue0.conn_state();\n/* 085 */ isNull74 = value74 == null;\n/* 086 */\n/* 087 */ final boolean isNull73 = isNull74 || value74.isEmpty();\n/* 088 */ java.lang.String value73 = isNull73 ?\n/* 089 */ null : (java.lang.String) value74.get();\n/* 090 */ boolean isNull72 = isNull73;\n/* 091 */ final UTF8String value72 = isNull72 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value73);\n/* 092 */ isNull72 = value72 == null;\n/* 093 */ if (isNull72) {\n/* 094 */ values1[19] = null;\n/* 095 */ } else {\n/* 096 */ values1[19] = value72;\n/* 097 */ }\n/* 098 */ }\n/* 099 */\n/* 100 */\n/* 101 */ private void apply_1(InternalRow i) {\n/* 102 */\n/* 103 */ boolean isNull37 = MapObjects_loopIsNull1;\n/* 104 */ final scala.Option value37 = isNull37 ? null : (scala.Option) MapObjects_loopValue0.sensor_name();\n/* 105 */ isNull37 = value37 == null;\n/* 106 */\n/* 107 */ final boolean isNull36 = isNull37 || value37.isEmpty();\n/* 108 */ java.lang.String value36 = isNull36 ?\n/* 109 */ null : (java.lang.String) value37.get();\n/* 110 */ boolean isNull35 = isNull36;\n/* 111 */ final UTF8String value35 = isNull35 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value36);\n/* 112 */ isNull35 = value35 == null;\n/* 113 */ if (isNull35) {\n/* 114 */ values1[2] = null;\n/* 115 */ } else {\n/* 116 */ values1[2] = value35;\n/* 117 */ }\n/* 118 */ boolean isNull40 = MapObjects_loopIsNull1;\n/* 119 */ final scala.Option value40 = isNull40 ? null : (scala.Option) MapObjects_loopValue0.ioa_uuid();\n/* 120 */ isNull40 = value40 == null;\n/* 121 */\n/* 122 */ final boolean isNull39 = isNull40 || value40.isEmpty();\n/* 123 */ java.lang.String value39 = isNull39 ?\n/* 124 */ null : (java.lang.String) value40.get();\n/* 125 */ boolean isNull38 = isNull39;\n/* 126 */ final UTF8String value38 = isNull38 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value39);\n/* 127 */ isNull38 = value38 == null;\n/* 128 */ if (isNull38) {\n/* 129 */ values1[3] = null;\n/* 130 */ } else {\n/* 131 */ values1[3] = value38;\n/* 132 */ }\n/* 133 */ }\n/* 134 */\n/* 135 */\n/* 136 */ private void apply_12(InternalRow i) {\n/* 137 */\n/* 138 */ boolean isNull98 = MapObjects_loopIsNull1;\n/* 139 */ final scala.Option value98 = isNull98 ? null : (scala.Option) MapObjects_loopValue0.cc();\n/* 140 */ isNull98 = value98 == null;\n/* 141 */\n/* 142 */ final boolean isNull97 = isNull98 || value98.isEmpty();\n/* 143 */ java.lang.String value97 = isNull97 ?\n/* 144 */ null : (java.lang.String) value98.get();\n/* 145 */ boolean isNull96 = isNull97;\n/* 146 */ final UTF8String value96 = isNull96 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value97);\n/* 147 */ isNull96 = value96 == null;\n/* 148 */ if (isNull96) {\n/* 149 */ values1[29] = null;\n/* 150 */ } else {\n/* 151 */ values1[29] = value96;\n/* 152 */ }\n/* 153 */ boolean isNull101 = MapObjects_loopIsNull1;\n/* 154 */ final scala.Option value101 = isNull101 ? null : (scala.Option) MapObjects_loopValue0.location();\n/* 155 */ isNull101 = value101 == null;\n/* 156 */\n/* 157 */ final boolean isNull100 = isNull101 || value101.isEmpty();\n/* 158 */ java.lang.String value100 = isNull100 ?\n/* 159 */ null : (java.lang.String) value101.get();\n/* 160 */ boolean isNull99 = isNull100;\n/* 161 */ final UTF8String value99 = isNull99 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value100);\n/* 162 */ isNull99 = value99 == null;\n/* 163 */ if (isNull99) {\n/* 164 */ values1[30] = null;\n/* 165 */ } else {\n/* 166 */ values1[30] = value99;\n/* 167 */ }\n/* 168 */ }\n/* 169 */\n/* 170 */\n/* 171 */ private void apply_9(InternalRow i) {\n/* 172 */\n/* 173 */ boolean isNull83 = MapObjects_loopIsNull1;\n/* 174 */ final scala.Option value83 = isNull83 ? null : (scala.Option) MapObjects_loopValue0.history();\n/* 175 */ isNull83 = value83 == null;\n/* 176 */\n/* 177 */ final boolean isNull82 = isNull83 || value83.isEmpty();\n/* 178 */ java.lang.String value82 = isNull82 ?\n/* 179 */ null : (java.lang.String) value83.get();\n/* 180 */ boolean isNull81 = isNull82;\n/* 181 */ final UTF8String value81 = isNull81 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value82);\n/* 182 */ isNull81 = value81 == null;\n/* 183 */ if (isNull81) {\n/* 184 */ values1[23] = null;\n/* 185 */ } else {\n/* 186 */ values1[23] = value81;\n/* 187 */ }\n/* 188 */ boolean isNull85 = MapObjects_loopIsNull1;\n/* 189 */ final scala.Option value85 = isNull85 ? null : (scala.Option) MapObjects_loopValue0.orig_pkts();\n/* 190 */ isNull85 = value85 == null;\n/* 191 */\n/* 192 */ final boolean isNull84 = isNull85 || value85.isEmpty();\n/* 193 */ long value84 = isNull84 ?\n/* 194 */ -1L : (Long) value85.get();\n/* 195 */ if (isNull84) {\n/* 196 */ values1[24] = null;\n/* 197 */ } else {\n/* 198 */ values1[24] = value84;\n/* 199 */ }\n/* 200 */ }\n/* 201 */\n/* 202 */\n/* 203 */ private void apply1_1(InternalRow i) {\n/* 204 */\n/* 205 */ boolean isNull25 = false;\n/* 206 */ final io.mistnet.analytics.scan.SrcDstGrouped value25 = isNull25 ? null : (io.mistnet.analytics.scan.SrcDstGrouped) value4._2();\n/* 207 */ isNull25 = value25 == null;\n/* 208 */\n/* 209 */ if (isNull25) {\n/* 210 */ throw new RuntimeException(errMsg2);\n/* 211 */ }\n/* 212 */\n/* 213 */ boolean isNull23 = false;\n/* 214 */ final scala.collection.Seq value23 = isNull23 ? null : (scala.collection.Seq) value25.cs();\n/* 215 */ isNull23 = value23 == null;\n/* 216 */ ArrayData value22 = null;\n/* 217 */\n/* 218 */ if (!isNull23) {\n/* 219 */\n/* 220 */ InternalRow[] convertedArray1 = null;\n/* 221 */ int dataLength1 = value23.size();\n/* 222 */ convertedArray1 = new InternalRow[dataLength1];\n/* 223 */\n/* 224 */ int loopIndex1 = 0;\n/* 225 */ while (loopIndex1 < dataLength1) {\n/* 226 */ MapObjects_loopValue0 = (io.mistnet.analytics.lib.ConnLog) (value23.apply(loopIndex1));\n/* 227 */ MapObjects_loopIsNull1 = MapObjects_loopValue0 == null;\n/* 228 */\n/* 229 */\n/* 230 */ boolean isNull26 = false;\n/* 231 */ InternalRow value26 = null;\n/* 232 */ if (!false && MapObjects_loopIsNull1) {\n/* 233 */\n/* 234 */ final InternalRow value28 = null;\n/* 235 */ isNull26 = true;\n/* 236 */ value26 = value28;\n/* 237 */ } else {\n/* 238 */\n/* 239 */ boolean isNull29 = false;\n/* 240 */ values1 = new Object[31];apply_0(i);\n/* 241 */ apply_1(i);\n/* 242 */ apply_2(i);\n/* 243 */ apply_3(i);\n/* 244 */ apply_4(i);\n/* 245 */ apply_5(i);\n/* 246 */ apply_6(i);\n/* 247 */ apply_7(i);\n/* 248 */ apply_8(i);\n/* 249 */ apply_9(i);\n/* 250 */ apply_10(i);\n/* 251 */ apply_11(i);\n/* 252 */ apply_12(i);\n/* 253 */ final InternalRow value29 = new org.apache.spark.sql.catalyst.expressions.GenericInternalRow(values1);\n/* 254 */ this.values1 = null;\n/* 255 */ isNull26 = isNull29;\n/* 256 */ value26 = value29;\n/* 257 */ }\n/* 258 */ if (isNull26) {\n/* 259 */ convertedArray1[loopIndex1] = null;\n/* 260 */ } else {\n/* 261 */ convertedArray1[loopIndex1] = value26 instanceof UnsafeRow? value26.copy() : value26;\n/* 262 */ }\n/* 263 */\n/* 264 */ loopIndex1 += 1;\n/* 265 */ }\n/* 266 */\n/* 267 */ value22 = new org.apache.spark.sql.catalyst.util.GenericArrayData(convertedArray1);\n/* 268 */ }\n/* 269 */ if (isNull23) {\n/* 270 */ values[2] = null;\n/* 271 */ } else {\n/* 272 */ values[2] = value22;\n/* 273 */ }\n/* 274 */ }\n/* 275 */\n/* 276 */\n/* 277 */ private void apply_3(InternalRow i) {\n/* 278 */\n/* 279 */ boolean isNull49 = MapObjects_loopIsNull1;\n/* 280 */ final scala.Option value49 = isNull49 ? null : (scala.Option) MapObjects_loopValue0.date();\n/* 281 */ isNull49 = value49 == null;\n/* 282 */\n/* 283 */ final boolean isNull48 = isNull49 || value49.isEmpty();\n/* 284 */ java.lang.String value48 = isNull48 ?\n/* 285 */ null : (java.lang.String) value49.get();\n/* 286 */ boolean isNull47 = isNull48;\n/* 287 */ final UTF8String value47 = isNull47 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value48);\n/* 288 */ isNull47 = value47 == null;\n/* 289 */ if (isNull47) {\n/* 290 */ values1[6] = null;\n/* 291 */ } else {\n/* 292 */ values1[6] = value47;\n/* 293 */ }\n/* 294 */ boolean isNull51 = MapObjects_loopIsNull1;\n/* 295 */ final scala.Option value51 = isNull51 ? null : (scala.Option) MapObjects_loopValue0.hour();\n/* 296 */ isNull51 = value51 == null;\n/* 297 */\n/* 298 */ final boolean isNull50 = isNull51 || value51.isEmpty();\n/* 299 */ int value50 = isNull50 ?\n/* 300 */ -1 : (Integer) value51.get();\n/* 301 */ if (isNull50) {\n/* 302 */ values1[7] = null;\n/* 303 */ } else {\n/* 304 */ values1[7] = value50;\n/* 305 */ }\n/* 306 */ }\n/* 307 */\n/* 308 */\n/* 309 */ private void apply_6(InternalRow i) {\n/* 310 */\n/* 311 */ boolean isNull65 = MapObjects_loopIsNull1;\n/* 312 */ final scala.Option value65 = isNull65 ? null : (scala.Option) MapObjects_loopValue0.service();\n/* 313 */ isNull65 = value65 == null;\n/* 314 */\n/* 315 */ final boolean isNull64 = isNull65 || value65.isEmpty();\n/* 316 */ java.lang.String value64 = isNull64 ?\n/* 317 */ null : (java.lang.String) value65.get();\n/* 318 */ boolean isNull63 = isNull64;\n/* 319 */ final UTF8String value63 = isNull63 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value64);\n/* 320 */ isNull63 = value63 == null;\n/* 321 */ if (isNull63) {\n/* 322 */ values1[15] = null;\n/* 323 */ } else {\n/* 324 */ values1[15] = value63;\n/* 325 */ }\n/* 326 */ boolean isNull67 = MapObjects_loopIsNull1;\n/* 327 */ final scala.Option value67 = isNull67 ? null : (scala.Option) MapObjects_loopValue0.duration();\n/* 328 */ isNull67 = value67 == null;\n/* 329 */\n/* 330 */ final boolean isNull66 = isNull67 || value67.isEmpty();\n/* 331 */ double value66 = isNull66 ?\n/* 332 */ -1.0 : (Double) value67.get();\n/* 333 */ if (isNull66) {\n/* 334 */ values1[16] = null;\n/* 335 */ } else {\n/* 336 */ values1[16] = value66;\n/* 337 */ }\n/* 338 */ }\n/* 339 */\n/* 340 */\n/* 341 */ private void apply_0(InternalRow i) {\n/* 342 */\n/* 343 */ boolean isNull32 = MapObjects_loopIsNull1;\n/* 344 */ final scala.Option value32 = isNull32 ? null : (scala.Option) MapObjects_loopValue0.log_type();\n/* 345 */ isNull32 = value32 == null;\n/* 346 */\n/* 347 */ final boolean isNull31 = isNull32 || value32.isEmpty();\n/* 348 */ java.lang.String value31 = isNull31 ?\n/* 349 */ null : (java.lang.String) value32.get();\n/* 350 */ boolean isNull30 = isNull31;\n/* 351 */ final UTF8String value30 = isNull30 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value31);\n/* 352 */ isNull30 = value30 == null;\n/* 353 */ if (isNull30) {\n/* 354 */ values1[0] = null;\n/* 355 */ } else {\n/* 356 */ values1[0] = value30;\n/* 357 */ }\n/* 358 */ boolean isNull34 = MapObjects_loopIsNull1;\n/* 359 */ final scala.Option value34 = isNull34 ? null : (scala.Option) MapObjects_loopValue0.timestamp();\n/* 360 */ isNull34 = value34 == null;\n/* 361 */\n/* 362 */ final boolean isNull33 = isNull34 || value34.isEmpty();\n/* 363 */ long value33 = isNull33 ?\n/* 364 */ -1L : (Long) value34.get();\n/* 365 */ if (isNull33) {\n/* 366 */ values1[1] = null;\n/* 367 */ } else {\n/* 368 */ values1[1] = value33;\n/* 369 */ }\n/* 370 */ }\n/* 371 */\n/* 372 */\n/* 373 */ private void apply_11(InternalRow i) {\n/* 374 */\n/* 375 */ boolean isNull94 = MapObjects_loopIsNull1;\n/* 376 */ final scala.Option value94 = isNull94 ? null : (scala.Option) MapObjects_loopValue0.tunnel_parents();\n/* 377 */ isNull94 = value94 == null;\n/* 378 */\n/* 379 */ final boolean isNull93 = isNull94 || value94.isEmpty();\n/* 380 */ scala.collection.Seq value93 = isNull93 ?\n/* 381 */ null : (scala.collection.Seq) value94.get();\n/* 382 */ ArrayData value92 = null;\n/* 383 */\n/* 384 */ if (!isNull93) {\n/* 385 */\n/* 386 */ UTF8String[] convertedArray = null;\n/* 387 */ int dataLength = value93.size();\n/* 388 */ convertedArray = new UTF8String[dataLength];\n/* 389 */\n/* 390 */ int loopIndex = 0;\n/* 391 */ while (loopIndex < dataLength) {\n/* 392 */ MapObjects_loopValue2 = (java.lang.String) (value93.apply(loopIndex));\n/* 393 */ MapObjects_loopIsNull3 = MapObjects_loopValue2 == null;\n/* 394 */\n/* 395 */\n/* 396 */ boolean isNull95 = MapObjects_loopIsNull3;\n/* 397 */ final UTF8String value95 = isNull95 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(MapObjects_loopValue2);\n/* 398 */ isNull95 = value95 == null;\n/* 399 */ if (isNull95) {\n/* 400 */ convertedArray[loopIndex] = null;\n/* 401 */ } else {\n/* 402 */ convertedArray[loopIndex] = value95;\n/* 403 */ }\n/* 404 */\n/* 405 */ loopIndex += 1;\n/* 406 */ }\n/* 407 */\n/* 408 */ value92 = new org.apache.spark.sql.catalyst.util.GenericArrayData(convertedArray);\n/* 409 */ }\n/* 410 */ if (isNull93) {\n/* 411 */ values1[28] = null;\n/* 412 */ } else {\n/* 413 */ values1[28] = value92;\n/* 414 */ }\n/* 415 */ }\n/* 416 */\n/* 417 */\n/* 418 */ private void apply_8(InternalRow i) {\n/* 419 */\n/* 420 */ boolean isNull76 = MapObjects_loopIsNull1;\n/* 421 */ final scala.Option value76 = isNull76 ? null : (scala.Option) MapObjects_loopValue0.local_orig();\n/* 422 */ isNull76 = value76 == null;\n/* 423 */\n/* 424 */ final boolean isNull75 = isNull76 || value76.isEmpty();\n/* 425 */ boolean value75 = isNull75 ?\n/* 426 */ false : (Boolean) value76.get();\n/* 427 */ if (isNull75) {\n/* 428 */ values1[20] = null;\n/* 429 */ } else {\n/* 430 */ values1[20] = value75;\n/* 431 */ }\n/* 432 */ boolean isNull78 = MapObjects_loopIsNull1;\n/* 433 */ final scala.Option value78 = isNull78 ? null : (scala.Option) MapObjects_loopValue0.local_resp();\n/* 434 */ isNull78 = value78 == null;\n/* 435 */\n/* 436 */ final boolean isNull77 = isNull78 || value78.isEmpty();\n/* 437 */ boolean value77 = isNull77 ?\n/* 438 */ false : (Boolean) value78.get();\n/* 439 */ if (isNull77) {\n/* 440 */ values1[21] = null;\n/* 441 */ } else {\n/* 442 */ values1[21] = value77;\n/* 443 */ }\n/* 444 */ boolean isNull80 = MapObjects_loopIsNull1;\n/* 445 */ final scala.Option value80 = isNull80 ? null : (scala.Option) MapObjects_loopValue0.missed_bytes();\n/* 446 */ isNull80 = value80 == null;\n/* 447 */\n/* 448 */ final boolean isNull79 = isNull80 || value80.isEmpty();\n/* 449 */ long value79 = isNull79 ?\n/* 450 */ -1L : (Long) value80.get();\n/* 451 */ if (isNull79) {\n/* 452 */ values1[22] = null;\n/* 453 */ } else {\n/* 454 */ values1[22] = value79;\n/* 455 */ }\n/* 456 */ }\n/* 457 */\n/* 458 */\n/* 459 */ private void apply1_0(InternalRow i) {\n/* 460 */\n/* 461 */ boolean isNull17 = false;\n/* 462 */ final io.mistnet.analytics.scan.SrcDstGrouped value17 = isNull17 ? null : (io.mistnet.analytics.scan.SrcDstGrouped) value4._2();\n/* 463 */ isNull17 = value17 == null;\n/* 464 */\n/* 465 */ if (isNull17) {\n/* 466 */ throw new RuntimeException(errMsg);\n/* 467 */ }\n/* 468 */\n/* 469 */ boolean isNull15 = false;\n/* 470 */ final java.lang.String value15 = isNull15 ? null : (java.lang.String) value17.src();\n/* 471 */ isNull15 = value15 == null;\n/* 472 */ boolean isNull14 = isNull15;\n/* 473 */ final UTF8String value14 = isNull14 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value15);\n/* 474 */ isNull14 = value14 == null;\n/* 475 */ if (isNull14) {\n/* 476 */ values[0] = null;\n/* 477 */ } else {\n/* 478 */ values[0] = value14;\n/* 479 */ }\n/* 480 */ boolean isNull21 = false;\n/* 481 */ final io.mistnet.analytics.scan.SrcDstGrouped value21 = isNull21 ? null : (io.mistnet.analytics.scan.SrcDstGrouped) value4._2();\n/* 482 */ isNull21 = value21 == null;\n/* 483 */\n/* 484 */ if (isNull21) {\n/* 485 */ throw new RuntimeException(errMsg1);\n/* 486 */ }\n/* 487 */\n/* 488 */ boolean isNull19 = false;\n/* 489 */ final java.lang.String value19 = isNull19 ? null : (java.lang.String) value21.dest();\n/* 490 */ isNull19 = value19 == null;\n/* 491 */ boolean isNull18 = isNull19;\n/* 492 */ final UTF8String value18 = isNull18 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value19);\n/* 493 */ isNull18 = value18 == null;\n/* 494 */ if (isNull18) {\n/* 495 */ values[1] = null;\n/* 496 */ } else {\n/* 497 */ values[1] = value18;\n/* 498 */ }\n/* 499 */ }\n/* 500 */\n/* 501 */\n/* 502 */ private void apply_2(InternalRow i) {\n/* 503 */\n/* 504 */ boolean isNull43 = MapObjects_loopIsNull1;\n/* 505 */ final scala.Option value43 = isNull43 ? null : (scala.Option) MapObjects_loopValue0.user_uuid();\n/* 506 */ isNull43 = value43 == null;\n/* 507 */\n/* 508 */ final boolean isNull42 = isNull43 || value43.isEmpty();\n/* 509 */ java.lang.String value42 = isNull42 ?\n/* 510 */ null : (java.lang.String) value43.get();\n/* 511 */ boolean isNull41 = isNull42;\n/* 512 */ final UTF8String value41 = isNull41 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value42);\n/* 513 */ isNull41 = value41 == null;\n/* 514 */ if (isNull41) {\n/* 515 */ values1[4] = null;\n/* 516 */ } else {\n/* 517 */ values1[4] = value41;\n/* 518 */ }\n/* 519 */ boolean isNull46 = MapObjects_loopIsNull1;\n/* 520 */ final scala.Option value46 = isNull46 ? null : (scala.Option) MapObjects_loopValue0.host_uuid();\n/* 521 */ isNull46 = value46 == null;\n/* 522 */\n/* 523 */ final boolean isNull45 = isNull46 || value46.isEmpty();\n/* 524 */ java.lang.String value45 = isNull45 ?\n/* 525 */ null : (java.lang.String) value46.get();\n/* 526 */ boolean isNull44 = isNull45;\n/* 527 */ final UTF8String value44 = isNull44 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value45);\n/* 528 */ isNull44 = value44 == null;\n/* 529 */ if (isNull44) {\n/* 530 */ values1[5] = null;\n/* 531 */ } else {\n/* 532 */ values1[5] = value44;\n/* 533 */ }\n/* 534 */ }\n/* 535 */\n/* 536 */\n/* 537 */ private void apply_5(InternalRow i) {\n/* 538 */\n/* 539 */ boolean isNull57 = MapObjects_loopIsNull1;\n/* 540 */ final int value57 = isNull57 ? -1 : MapObjects_loopValue0.src_port();\n/* 541 */ if (isNull57) {\n/* 542 */ values1[11] = null;\n/* 543 */ } else {\n/* 544 */ values1[11] = value57;\n/* 545 */ }\n/* 546 */ boolean isNull59 = MapObjects_loopIsNull1;\n/* 547 */ final java.lang.String value59 = isNull59 ? null : (java.lang.String) MapObjects_loopValue0.dest();\n/* 548 */ isNull59 = value59 == null;\n/* 549 */ boolean isNull58 = isNull59;\n/* 550 */ final UTF8String value58 = isNull58 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value59);\n/* 551 */ isNull58 = value58 == null;\n/* 552 */ if (isNull58) {\n/* 553 */ values1[12] = null;\n/* 554 */ } else {\n/* 555 */ values1[12] = value58;\n/* 556 */ }\n/* 557 */ boolean isNull60 = MapObjects_loopIsNull1;\n/* 558 */ final int value60 = isNull60 ? -1 : MapObjects_loopValue0.dest_port();\n/* 559 */ if (isNull60) {\n/* 560 */ values1[13] = null;\n/* 561 */ } else {\n/* 562 */ values1[13] = value60;\n/* 563 */ }\n/* 564 */ boolean isNull62 = MapObjects_loopIsNull1;\n/* 565 */ final java.lang.String value62 = isNull62 ? null : (java.lang.String) MapObjects_loopValue0.proto();\n/* 566 */ isNull62 = value62 == null;\n/* 567 */ boolean isNull61 = isNull62;\n/* 568 */ final UTF8String value61 = isNull61 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value62);\n/* 569 */ isNull61 = value61 == null;\n/* 570 */ if (isNull61) {\n/* 571 */ values1[14] = null;\n/* 572 */ } else {\n/* 573 */ values1[14] = value61;\n/* 574 */ }\n/* 575 */ }\n/* 576 */\n/* 577 */\n/* 578 */ private void apply_10(InternalRow i) {\n/* 579 */\n/* 580 */ boolean isNull87 = MapObjects_loopIsNull1;\n/* 581 */ final scala.Option value87 = isNull87 ? null : (scala.Option) MapObjects_loopValue0.orig_ip_bytes();\n/* 582 */ isNull87 = value87 == null;\n/* 583 */\n/* 584 */ final boolean isNull86 = isNull87 || value87.isEmpty();\n/* 585 */ long value86 = isNull86 ?\n/* 586 */ -1L : (Long) value87.get();\n/* 587 */ if (isNull86) {\n/* 588 */ values1[25] = null;\n/* 589 */ } else {\n/* 590 */ values1[25] = value86;\n/* 591 */ }\n/* 592 */ boolean isNull89 = MapObjects_loopIsNull1;\n/* 593 */ final scala.Option value89 = isNull89 ? null : (scala.Option) MapObjects_loopValue0.resp_pkts();\n/* 594 */ isNull89 = value89 == null;\n/* 595 */\n/* 596 */ final boolean isNull88 = isNull89 || value89.isEmpty();\n/* 597 */ long value88 = isNull88 ?\n/* 598 */ -1L : (Long) value89.get();\n/* 599 */ if (isNull88) {\n/* 600 */ values1[26] = null;\n/* 601 */ } else {\n/* 602 */ values1[26] = value88;\n/* 603 */ }\n/* 604 */ boolean isNull91 = MapObjects_loopIsNull1;\n/* 605 */ final scala.Option value91 = isNull91 ? null : (scala.Option) MapObjects_loopValue0.resp_ip_bytes();\n/* 606 */ isNull91 = value91 == null;\n/* 607 */\n/* 608 */ final boolean isNull90 = isNull91 || value91.isEmpty();\n/* 609 */ long value90 = isNull90 ?\n/* 610 */ -1L : (Long) value91.get();\n/* 611 */ if (isNull90) {\n/* 612 */ values1[27] = null;\n/* 613 */ } else {\n/* 614 */ values1[27] = value90;\n/* 615 */ }\n/* 616 */ }\n/* 617 */\n/* 618 */\n/* 619 */ public SpecificMutableProjection(Object[] references) {\n/* 620 */ this.references = references;\n/* 621 */ mutableRow = new org.apache.spark.sql.catalyst.expressions.GenericMutableRow(2);\n/* 622 */ this.values = null;\n/* 623 */ this.errMsg = (java.lang.String) references[3];\n/* 624 */ this.errMsg1 = (java.lang.String) references[4];\n/* 625 */\n/* 626 */\n/* 627 */ this.errMsg2 = (java.lang.String) references[5];\n/* 628 */ this.values1 = null;\n/* 629 */\n/* 630 */\n/* 631 */ this.isNull_0 = true;\n/* 632 */ this.value_0 = false;\n/* 633 */ this.isNull_1 = true;\n/* 634 */ this.value_1 = null;\n/* 635 */ }\n/* 636 */\n/* 637 */ public org.apache.spark.sql.catalyst.expressions.codegen.BaseMutableProjection target(MutableRow row) {\n/* 638 */ mutableRow = row;\n/* 639 */ return this;\n/* 640 */ }\n/* 641 */\n/* 642 */ /* Provide immutable access to the last projected row. */\n/* 643 */ public InternalRow currentValue() {\n/* 644 */ return (InternalRow) mutableRow;\n/* 645 */ }\n/* 646 */\n/* 647 */ public java.lang.Object apply(java.lang.Object _i) {\n/* 648 */ InternalRow i = (InternalRow) _i;\n/* 649 */\n/* 650 */\n/* 651 */\n/* 652 */ Object obj = ((Expression) references[0]).eval(null);\n/* 653 */ scala.Tuple2 value1 = (scala.Tuple2) obj;\n/* 654 */\n/* 655 */ boolean isNull2 = false;\n/* 656 */ final boolean value2 = isNull2 ? false : (Boolean) value1._1();\n/* 657 */ this.isNull_0 = isNull2;\n/* 658 */ this.value_0 = value2;\n/* 659 */\n/* 660 */\n/* 661 */ Object obj1 = ((Expression) references[1]).eval(null);\n/* 662 */ scala.Tuple2 value4 = (scala.Tuple2) obj1;\n/* 663 */\n/* 664 */ boolean isNull8 = false;\n/* 665 */ final io.mistnet.analytics.scan.SrcDstGrouped value8 = isNull8 ? null : (io.mistnet.analytics.scan.SrcDstGrouped) value4._2();\n/* 666 */ isNull8 = value8 == null;\n/* 667 */ boolean isNull6 = false;\n/* 668 */ boolean value6 = true;\n/* 669 */\n/* 670 */ if (!false && isNull8) {\n/* 671 */ } else {\n/* 672 */\n/* 673 */ Object obj2 = ((Expression) references[2]).eval(null);\n/* 674 */ scala.None$ value10 = (scala.None$) obj2;\n/* 675 */\n/* 676 */ boolean isNull11 = false;\n/* 677 */ final io.mistnet.analytics.scan.SrcDstGrouped value11 = isNull11 ? null : (io.mistnet.analytics.scan.SrcDstGrouped) value4._2();\n/* 678 */ isNull11 = value11 == null;\n/* 679 */ boolean isNull9 = false || isNull11;\n/* 680 */ final boolean value9 = isNull9 ? false : value10.equals(value11);\n/* 681 */ if (!isNull9 && value9) {\n/* 682 */ } else if (!false && !isNull9) {\n/* 683 */ value6 = false;\n/* 684 */ } else {\n/* 685 */ isNull6 = true;\n/* 686 */ }\n/* 687 */ }\n/* 688 */ boolean isNull5 = false;\n/* 689 */ InternalRow value5 = null;\n/* 690 */ if (!isNull6 && value6) {\n/* 691 */\n/* 692 */ final InternalRow value12 = null;\n/* 693 */ isNull5 = true;\n/* 694 */ value5 = value12;\n/* 695 */ } else {\n/* 696 */\n/* 697 */ boolean isNull13 = false;\n/* 698 */ this.values = new Object[3];apply1_0(i);\n/* 699 */ apply1_1(i);\n/* 700 */ final InternalRow value13 = new org.apache.spark.sql.catalyst.expressions.GenericInternalRow(values);\n/* 701 */ this.values = null;\n/* 702 */ isNull5 = isNull13;\n/* 703 */ value5 = value13;\n/* 704 */ }\n/* 705 */ this.isNull_1 = isNull5;\n/* 706 */ this.value_1 = value5;\n/* 707 */\n/* 708 */ // copy all the results into MutableRow\n/* 709 */\n/* 710 */ if (!this.isNull_0) {\n/* 711 */ mutableRow.setBoolean(0, this.value_0);\n/* 712 */ } else {\n/* 713 */ mutableRow.setNullAt(0);\n/* 714 */ }\n/* 715 */\n/* 716 */ if (!this.isNull_1) {\n/* 717 */ mutableRow.update(1, this.value_1);\n/* 718 */ } else {\n/* 719 */ mutableRow.setNullAt(1);\n/* 720 */ }\n/* 721 */\n/* 722 */ return mutableRow;\n/* 723 */ }\n/* 724 */ }\n\n\tat org.spark_project.guava.util.concurrent.AbstractFuture$Sync.getValue(AbstractFuture.java:306)\n\tat org.spark_project.guava.util.concurrent.AbstractFuture$Sync.get(AbstractFuture.java:293)\n\tat org.spark_project.guava.util.concurrent.AbstractFuture.get(AbstractFuture.java:116)\n\tat org.spark_project.guava.util.concurrent.Uninterruptibles.getUninterruptibly(Uninterruptibles.java:135)\n\tat org.spark_project.guava.cache.LocalCache$Segment.getAndRecordStats(LocalCache.java:2410)\n\tat org.spark_project.guava.cache.LocalCache$Segment.loadSync(LocalCache.java:2380)\n\tat org.spark_project.guava.cache.LocalCache$Segment.lockedGetOrLoad(LocalCache.java:2342)\n\tat org.spark_project.guava.cache.LocalCache$Segment.get(LocalCache.java:2257)\n\tat org.spark_project.guava.cache.LocalCache.get(LocalCache.java:4000)\n\tat org.spark_project.guava.cache.LocalCache.getOrLoad(LocalCache.java:4004)\n\tat org.spark_project.guava.cache.LocalCache$LocalLoadingCache.get(LocalCache.java:4874)\n\tat org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator$.compile(CodeGenerator.scala:841)\n\tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.create(GenerateMutableProjection.scala:140)\n\tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.generate(GenerateMutableProjection.scala:44)\n\tat org.apache.spark.sql.execution.SparkPlan.newMutableProjection(SparkPlan.scala:369)\n\tat org.apache.spark.sql.execution.aggregate.SortAggregateExec$$anonfun$doExecute$1$$anonfun$3$$anonfun$4.apply(SortAggregateExec.scala:93)\n\tat org.apache.spark.sql.execution.aggregate.SortAggregateExec$$anonfun$doExecute$1$$anonfun$3$$anonfun$4.apply(SortAggregateExec.scala:92)\n\tat org.apache.spark.sql.execution.aggregate.AggregationIterator.(AggregationIterator.scala:143)\n\tat org.apache.spark.sql.execution.aggregate.SortBasedAggregationIterator.(SortBasedAggregationIterator.scala:39)\n\tat org.apache.spark.sql.execution.aggregate.SortAggregateExec$$anonfun$doExecute$1$$anonfun$3.apply(SortAggregateExec.scala:84)\n\tat org.apache.spark.sql.execution.aggregate.SortAggregateExec$$anonfun$doExecute$1$$anonfun$3.apply(SortAggregateExec.scala:75)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:86)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{noformat}","from":"reporter","subject":"Spark generated code causes CompileException when groupByKey, reduceGroups and map(_._2) are used"},{"body":"Same here:\n\nAfter groupByKey( _._1).reduceGroups((a,b) => (a._1, a._2 ++ a._2)).map(_._2):\n\njava.util.concurrent.ExecutionException: java.lang.Exception: failed to compile: org.codehaus.commons.compiler.CompileException: File 'generated.java', Line 31, Column 69: Unknown variable or type \"value4\"\n\nSpark Version 2.0.1","from":"developer"},{"body":"Move to Priority Critical unless a workaround is identified. This is a very basic functionality.","from":"developer"},{"body":"Could one of you provide a reproducible example.","from":"developer"},{"body":"I tried something like this on master and on branch-2.0:\n{noformat}\nval ds = spark.range(10000).select($\"id\" % 100 as \"grp_id\", array($\"id\")).as[(Long, Seq[Long])] \nds.groupByKey(_._1).reduceGroups((a, b) => (a._1, a._2 ++ b._2)).map(_._2)\n{noformat}","from":"developer"},{"body":"Try this in spark-shell:\n\ncase class Route(src: String, dest: String, cost: Int)\ncase class GroupedRoutes(src: String, dest: String, routes: Seq[Route])\n\nval ds = sc.parallelize(Array(\n Route(\"a\", \"b\", 1),\n Route(\"a\", \"b\", 2),\n Route(\"a\", \"c\", 2),\n Route(\"a\", \"d\", 10),\n Route(\"b\", \"a\", 1),\n Route(\"b\", \"a\", 5),\n Route(\"b\", \"c\", 6))\n ).toDF.as[Route]\n\nval grped = ds.map(r => GroupedRoutes(r.src, r.dest, Seq(r)))\n .groupByKey(r => (r.src, r.dest))\n .reduceGroups { (g1: GroupedRoutes, g2: GroupedRoutes) =>\n GroupedRoutes(g1.src, g1.dest, g1.routes ++ g2.routes)\n }.map(_._2)\n\nSame thing works fine in 2.0.0\n\nOn Thu, Oct 27, 2016 at 4:06 PM, Herman van Hovell (JIRA) \n\n\n\n\n-- \nRegards,\nRay\n","from":"developer"},{"body":"I confirmed this code can reproduce on 2.0.1.\nThis problem occurs due to the similar reason in SPARK-18147\n\nTo call {{ctx.splitExpression}} in {{createStruct.doGenCode}} make *a variable* inaccesssible by splitting the original one function into multiple functions.\nSPARK-14793 seems to have introduced {{ctx.splitExpression}} here to fix other issues on April. Since Ray says this code works in 2.0.0, other changes may introduce this issue.","from":"developer"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15693","from":"developer"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15796","from":"developer"}],"created":"2016-10-26T22:16:03.000+0000","description":"Code logic looks like this:\n{noformat}\n .groupByKey\n .reduceGroups\n .map(_._2)\n{noformat}\nWorks fine with 2.0.0.\n\n2.0.1 error Message: \n{noformat}\nCaused by: java.util.concurrent.ExecutionException: java.lang.Exception: failed to compile: org.codehaus.commons.compiler.CompileException: File 'generated.java', Line 206, Column 123: Unknown variable or type \"value4\"\n/* 001 */ public java.lang.Object generate(Object[] references) {\n/* 002 */ return new SpecificMutableProjection(references);\n/* 003 */ }\n/* 004 */\n/* 005 */ class SpecificMutableProjection extends org.apache.spark.sql.catalyst.expressions.codegen.BaseMutableProjection {\n/* 006 */\n/* 007 */ private Object[] references;\n/* 008 */ private MutableRow mutableRow;\n/* 009 */ private Object[] values;\n/* 010 */ private java.lang.String errMsg;\n/* 011 */ private java.lang.String errMsg1;\n/* 012 */ private boolean MapObjects_loopIsNull1;\n/* 013 */ private io.mistnet.analytics.lib.ConnLog MapObjects_loopValue0;\n/* 014 */ private java.lang.String errMsg2;\n/* 015 */ private Object[] values1;\n/* 016 */ private boolean MapObjects_loopIsNull3;\n/* 017 */ private java.lang.String MapObjects_loopValue2;\n/* 018 */ private boolean isNull_0;\n/* 019 */ private boolean value_0;\n/* 020 */ private boolean isNull_1;\n/* 021 */ private InternalRow value_1;\n/* 022 */\n/* 023 */ private void apply_4(InternalRow i) {\n/* 024 */\n/* 025 */ boolean isNull52 = MapObjects_loopIsNull1;\n/* 026 */ final double value52 = isNull52 ? -1.0 : MapObjects_loopValue0.ts();\n/* 027 */ if (isNull52) {\n/* 028 */ values1[8] = null;\n/* 029 */ } else {\n/* 030 */ values1[8] = value52;\n/* 031 */ }\n/* 032 */ boolean isNull54 = MapObjects_loopIsNull1;\n/* 033 */ final java.lang.String value54 = isNull54 ? null : (java.lang.String) MapObjects_loopValue0.uid();\n/* 034 */ isNull54 = value54 == null;\n/* 035 */ boolean isNull53 = isNull54;\n/* 036 */ final UTF8String value53 = isNull53 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value54);\n/* 037 */ isNull53 = value53 == null;\n/* 038 */ if (isNull53) {\n/* 039 */ values1[9] = null;\n/* 040 */ } else {\n/* 041 */ values1[9] = value53;\n/* 042 */ }\n/* 043 */ boolean isNull56 = MapObjects_loopIsNull1;\n/* 044 */ final java.lang.String value56 = isNull56 ? null : (java.lang.String) MapObjects_loopValue0.src();\n/* 045 */ isNull56 = value56 == null;\n/* 046 */ boolean isNull55 = isNull56;\n/* 047 */ final UTF8String value55 = isNull55 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value56);\n/* 048 */ isNull55 = value55 == null;\n/* 049 */ if (isNull55) {\n/* 050 */ values1[10] = null;\n/* 051 */ } else {\n/* 052 */ values1[10] = value55;\n/* 053 */ }\n/* 054 */ }\n/* 055 */\n/* 056 */\n/* 057 */ private void apply_7(InternalRow i) {\n/* 058 */\n/* 059 */ boolean isNull69 = MapObjects_loopIsNull1;\n/* 060 */ final scala.Option value69 = isNull69 ? null : (scala.Option) MapObjects_loopValue0.orig_bytes();\n/* 061 */ isNull69 = value69 == null;\n/* 062 */\n/* 063 */ final boolean isNull68 = isNull69 || value69.isEmpty();\n/* 064 */ long value68 = isNull68 ?\n/* 065 */ -1L : (Long) value69.get();\n/* 066 */ if (isNull68) {\n/* 067 */ values1[17] = null;\n/* 068 */ } else {\n/* 069 */ values1[17] = value68;\n/* 070 */ }\n/* 071 */ boolean isNull71 = MapObjects_loopIsNull1;\n/* 072 */ final scala.Option value71 = isNull71 ? null : (scala.Option) MapObjects_loopValue0.resp_bytes();\n/* 073 */ isNull71 = value71 == null;\n/* 074 */\n/* 075 */ final boolean isNull70 = isNull71 || value71.isEmpty();\n/* 076 */ long value70 = isNull70 ?\n/* 077 */ -1L : (Long) value71.get();\n/* 078 */ if (isNull70) {\n/* 079 */ values1[18] = null;\n/* 080 */ } else {\n/* 081 */ values1[18] = value70;\n/* 082 */ }\n/* 083 */ boolean isNull74 = MapObjects_loopIsNull1;\n/* 084 */ final scala.Option value74 = isNull74 ? null : (scala.Option) MapObjects_loopValue0.conn_state();\n/* 085 */ isNull74 = value74 == null;\n/* 086 */\n/* 087 */ final boolean isNull73 = isNull74 || value74.isEmpty();\n/* 088 */ java.lang.String value73 = isNull73 ?\n/* 089 */ null : (java.lang.String) value74.get();\n/* 090 */ boolean isNull72 = isNull73;\n/* 091 */ final UTF8String value72 = isNull72 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value73);\n/* 092 */ isNull72 = value72 == null;\n/* 093 */ if (isNull72) {\n/* 094 */ values1[19] = null;\n/* 095 */ } else {\n/* 096 */ values1[19] = value72;\n/* 097 */ }\n/* 098 */ }\n/* 099 */\n/* 100 */\n/* 101 */ private void apply_1(InternalRow i) {\n/* 102 */\n/* 103 */ boolean isNull37 = MapObjects_loopIsNull1;\n/* 104 */ final scala.Option value37 = isNull37 ? null : (scala.Option) MapObjects_loopValue0.sensor_name();\n/* 105 */ isNull37 = value37 == null;\n/* 106 */\n/* 107 */ final boolean isNull36 = isNull37 || value37.isEmpty();\n/* 108 */ java.lang.String value36 = isNull36 ?\n/* 109 */ null : (java.lang.String) value37.get();\n/* 110 */ boolean isNull35 = isNull36;\n/* 111 */ final UTF8String value35 = isNull35 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value36);\n/* 112 */ isNull35 = value35 == null;\n/* 113 */ if (isNull35) {\n/* 114 */ values1[2] = null;\n/* 115 */ } else {\n/* 116 */ values1[2] = value35;\n/* 117 */ }\n/* 118 */ boolean isNull40 = MapObjects_loopIsNull1;\n/* 119 */ final scala.Option value40 = isNull40 ? null : (scala.Option) MapObjects_loopValue0.ioa_uuid();\n/* 120 */ isNull40 = value40 == null;\n/* 121 */\n/* 122 */ final boolean isNull39 = isNull40 || value40.isEmpty();\n/* 123 */ java.lang.String value39 = isNull39 ?\n/* 124 */ null : (java.lang.String) value40.get();\n/* 125 */ boolean isNull38 = isNull39;\n/* 126 */ final UTF8String value38 = isNull38 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value39);\n/* 127 */ isNull38 = value38 == null;\n/* 128 */ if (isNull38) {\n/* 129 */ values1[3] = null;\n/* 130 */ } else {\n/* 131 */ values1[3] = value38;\n/* 132 */ }\n/* 133 */ }\n/* 134 */\n/* 135 */\n/* 136 */ private void apply_12(InternalRow i) {\n/* 137 */\n/* 138 */ boolean isNull98 = MapObjects_loopIsNull1;\n/* 139 */ final scala.Option value98 = isNull98 ? null : (scala.Option) MapObjects_loopValue0.cc();\n/* 140 */ isNull98 = value98 == null;\n/* 141 */\n/* 142 */ final boolean isNull97 = isNull98 || value98.isEmpty();\n/* 143 */ java.lang.String value97 = isNull97 ?\n/* 144 */ null : (java.lang.String) value98.get();\n/* 145 */ boolean isNull96 = isNull97;\n/* 146 */ final UTF8String value96 = isNull96 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value97);\n/* 147 */ isNull96 = value96 == null;\n/* 148 */ if (isNull96) {\n/* 149 */ values1[29] = null;\n/* 150 */ } else {\n/* 151 */ values1[29] = value96;\n/* 152 */ }\n/* 153 */ boolean isNull101 = MapObjects_loopIsNull1;\n/* 154 */ final scala.Option value101 = isNull101 ? null : (scala.Option) MapObjects_loopValue0.location();\n/* 155 */ isNull101 = value101 == null;\n/* 156 */\n/* 157 */ final boolean isNull100 = isNull101 || value101.isEmpty();\n/* 158 */ java.lang.String value100 = isNull100 ?\n/* 159 */ null : (java.lang.String) value101.get();\n/* 160 */ boolean isNull99 = isNull100;\n/* 161 */ final UTF8String value99 = isNull99 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value100);\n/* 162 */ isNull99 = value99 == null;\n/* 163 */ if (isNull99) {\n/* 164 */ values1[30] = null;\n/* 165 */ } else {\n/* 166 */ values1[30] = value99;\n/* 167 */ }\n/* 168 */ }\n/* 169 */\n/* 170 */\n/* 171 */ private void apply_9(InternalRow i) {\n/* 172 */\n/* 173 */ boolean isNull83 = MapObjects_loopIsNull1;\n/* 174 */ final scala.Option value83 = isNull83 ? null : (scala.Option) MapObjects_loopValue0.history();\n/* 175 */ isNull83 = value83 == null;\n/* 176 */\n/* 177 */ final boolean isNull82 = isNull83 || value83.isEmpty();\n/* 178 */ java.lang.String value82 = isNull82 ?\n/* 179 */ null : (java.lang.String) value83.get();\n/* 180 */ boolean isNull81 = isNull82;\n/* 181 */ final UTF8String value81 = isNull81 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value82);\n/* 182 */ isNull81 = value81 == null;\n/* 183 */ if (isNull81) {\n/* 184 */ values1[23] = null;\n/* 185 */ } else {\n/* 186 */ values1[23] = value81;\n/* 187 */ }\n/* 188 */ boolean isNull85 = MapObjects_loopIsNull1;\n/* 189 */ final scala.Option value85 = isNull85 ? null : (scala.Option) MapObjects_loopValue0.orig_pkts();\n/* 190 */ isNull85 = value85 == null;\n/* 191 */\n/* 192 */ final boolean isNull84 = isNull85 || value85.isEmpty();\n/* 193 */ long value84 = isNull84 ?\n/* 194 */ -1L : (Long) value85.get();\n/* 195 */ if (isNull84) {\n/* 196 */ values1[24] = null;\n/* 197 */ } else {\n/* 198 */ values1[24] = value84;\n/* 199 */ }\n/* 200 */ }\n/* 201 */\n/* 202 */\n/* 203 */ private void apply1_1(InternalRow i) {\n/* 204 */\n/* 205 */ boolean isNull25 = false;\n/* 206 */ final io.mistnet.analytics.scan.SrcDstGrouped value25 = isNull25 ? null : (io.mistnet.analytics.scan.SrcDstGrouped) value4._2();\n/* 207 */ isNull25 = value25 == null;\n/* 208 */\n/* 209 */ if (isNull25) {\n/* 210 */ throw new RuntimeException(errMsg2);\n/* 211 */ }\n/* 212 */\n/* 213 */ boolean isNull23 = false;\n/* 214 */ final scala.collection.Seq value23 = isNull23 ? null : (scala.collection.Seq) value25.cs();\n/* 215 */ isNull23 = value23 == null;\n/* 216 */ ArrayData value22 = null;\n/* 217 */\n/* 218 */ if (!isNull23) {\n/* 219 */\n/* 220 */ InternalRow[] convertedArray1 = null;\n/* 221 */ int dataLength1 = value23.size();\n/* 222 */ convertedArray1 = new InternalRow[dataLength1];\n/* 223 */\n/* 224 */ int loopIndex1 = 0;\n/* 225 */ while (loopIndex1 < dataLength1) {\n/* 226 */ MapObjects_loopValue0 = (io.mistnet.analytics.lib.ConnLog) (value23.apply(loopIndex1));\n/* 227 */ MapObjects_loopIsNull1 = MapObjects_loopValue0 == null;\n/* 228 */\n/* 229 */\n/* 230 */ boolean isNull26 = false;\n/* 231 */ InternalRow value26 = null;\n/* 232 */ if (!false && MapObjects_loopIsNull1) {\n/* 233 */\n/* 234 */ final InternalRow value28 = null;\n/* 235 */ isNull26 = true;\n/* 236 */ value26 = value28;\n/* 237 */ } else {\n/* 238 */\n/* 239 */ boolean isNull29 = false;\n/* 240 */ values1 = new Object[31];apply_0(i);\n/* 241 */ apply_1(i);\n/* 242 */ apply_2(i);\n/* 243 */ apply_3(i);\n/* 244 */ apply_4(i);\n/* 245 */ apply_5(i);\n/* 246 */ apply_6(i);\n/* 247 */ apply_7(i);\n/* 248 */ apply_8(i);\n/* 249 */ apply_9(i);\n/* 250 */ apply_10(i);\n/* 251 */ apply_11(i);\n/* 252 */ apply_12(i);\n/* 253 */ final InternalRow value29 = new org.apache.spark.sql.catalyst.expressions.GenericInternalRow(values1);\n/* 254 */ this.values1 = null;\n/* 255 */ isNull26 = isNull29;\n/* 256 */ value26 = value29;\n/* 257 */ }\n/* 258 */ if (isNull26) {\n/* 259 */ convertedArray1[loopIndex1] = null;\n/* 260 */ } else {\n/* 261 */ convertedArray1[loopIndex1] = value26 instanceof UnsafeRow? value26.copy() : value26;\n/* 262 */ }\n/* 263 */\n/* 264 */ loopIndex1 += 1;\n/* 265 */ }\n/* 266 */\n/* 267 */ value22 = new org.apache.spark.sql.catalyst.util.GenericArrayData(convertedArray1);\n/* 268 */ }\n/* 269 */ if (isNull23) {\n/* 270 */ values[2] = null;\n/* 271 */ } else {\n/* 272 */ values[2] = value22;\n/* 273 */ }\n/* 274 */ }\n/* 275 */\n/* 276 */\n/* 277 */ private void apply_3(InternalRow i) {\n/* 278 */\n/* 279 */ boolean isNull49 = MapObjects_loopIsNull1;\n/* 280 */ final scala.Option value49 = isNull49 ? null : (scala.Option) MapObjects_loopValue0.date();\n/* 281 */ isNull49 = value49 == null;\n/* 282 */\n/* 283 */ final boolean isNull48 = isNull49 || value49.isEmpty();\n/* 284 */ java.lang.String value48 = isNull48 ?\n/* 285 */ null : (java.lang.String) value49.get();\n/* 286 */ boolean isNull47 = isNull48;\n/* 287 */ final UTF8String value47 = isNull47 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value48);\n/* 288 */ isNull47 = value47 == null;\n/* 289 */ if (isNull47) {\n/* 290 */ values1[6] = null;\n/* 291 */ } else {\n/* 292 */ values1[6] = value47;\n/* 293 */ }\n/* 294 */ boolean isNull51 = MapObjects_loopIsNull1;\n/* 295 */ final scala.Option value51 = isNull51 ? null : (scala.Option) MapObjects_loopValue0.hour();\n/* 296 */ isNull51 = value51 == null;\n/* 297 */\n/* 298 */ final boolean isNull50 = isNull51 || value51.isEmpty();\n/* 299 */ int value50 = isNull50 ?\n/* 300 */ -1 : (Integer) value51.get();\n/* 301 */ if (isNull50) {\n/* 302 */ values1[7] = null;\n/* 303 */ } else {\n/* 304 */ values1[7] = value50;\n/* 305 */ }\n/* 306 */ }\n/* 307 */\n/* 308 */\n/* 309 */ private void apply_6(InternalRow i) {\n/* 310 */\n/* 311 */ boolean isNull65 = MapObjects_loopIsNull1;\n/* 312 */ final scala.Option value65 = isNull65 ? null : (scala.Option) MapObjects_loopValue0.service();\n/* 313 */ isNull65 = value65 == null;\n/* 314 */\n/* 315 */ final boolean isNull64 = isNull65 || value65.isEmpty();\n/* 316 */ java.lang.String value64 = isNull64 ?\n/* 317 */ null : (java.lang.String) value65.get();\n/* 318 */ boolean isNull63 = isNull64;\n/* 319 */ final UTF8String value63 = isNull63 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value64);\n/* 320 */ isNull63 = value63 == null;\n/* 321 */ if (isNull63) {\n/* 322 */ values1[15] = null;\n/* 323 */ } else {\n/* 324 */ values1[15] = value63;\n/* 325 */ }\n/* 326 */ boolean isNull67 = MapObjects_loopIsNull1;\n/* 327 */ final scala.Option value67 = isNull67 ? null : (scala.Option) MapObjects_loopValue0.duration();\n/* 328 */ isNull67 = value67 == null;\n/* 329 */\n/* 330 */ final boolean isNull66 = isNull67 || value67.isEmpty();\n/* 331 */ double value66 = isNull66 ?\n/* 332 */ -1.0 : (Double) value67.get();\n/* 333 */ if (isNull66) {\n/* 334 */ values1[16] = null;\n/* 335 */ } else {\n/* 336 */ values1[16] = value66;\n/* 337 */ }\n/* 338 */ }\n/* 339 */\n/* 340 */\n/* 341 */ private void apply_0(InternalRow i) {\n/* 342 */\n/* 343 */ boolean isNull32 = MapObjects_loopIsNull1;\n/* 344 */ final scala.Option value32 = isNull32 ? null : (scala.Option) MapObjects_loopValue0.log_type();\n/* 345 */ isNull32 = value32 == null;\n/* 346 */\n/* 347 */ final boolean isNull31 = isNull32 || value32.isEmpty();\n/* 348 */ java.lang.String value31 = isNull31 ?\n/* 349 */ null : (java.lang.String) value32.get();\n/* 350 */ boolean isNull30 = isNull31;\n/* 351 */ final UTF8String value30 = isNull30 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value31);\n/* 352 */ isNull30 = value30 == null;\n/* 353 */ if (isNull30) {\n/* 354 */ values1[0] = null;\n/* 355 */ } else {\n/* 356 */ values1[0] = value30;\n/* 357 */ }\n/* 358 */ boolean isNull34 = MapObjects_loopIsNull1;\n/* 359 */ final scala.Option value34 = isNull34 ? null : (scala.Option) MapObjects_loopValue0.timestamp();\n/* 360 */ isNull34 = value34 == null;\n/* 361 */\n/* 362 */ final boolean isNull33 = isNull34 || value34.isEmpty();\n/* 363 */ long value33 = isNull33 ?\n/* 364 */ -1L : (Long) value34.get();\n/* 365 */ if (isNull33) {\n/* 366 */ values1[1] = null;\n/* 367 */ } else {\n/* 368 */ values1[1] = value33;\n/* 369 */ }\n/* 370 */ }\n/* 371 */\n/* 372 */\n/* 373 */ private void apply_11(InternalRow i) {\n/* 374 */\n/* 375 */ boolean isNull94 = MapObjects_loopIsNull1;\n/* 376 */ final scala.Option value94 = isNull94 ? null : (scala.Option) MapObjects_loopValue0.tunnel_parents();\n/* 377 */ isNull94 = value94 == null;\n/* 378 */\n/* 379 */ final boolean isNull93 = isNull94 || value94.isEmpty();\n/* 380 */ scala.collection.Seq value93 = isNull93 ?\n/* 381 */ null : (scala.collection.Seq) value94.get();\n/* 382 */ ArrayData value92 = null;\n/* 383 */\n/* 384 */ if (!isNull93) {\n/* 385 */\n/* 386 */ UTF8String[] convertedArray = null;\n/* 387 */ int dataLength = value93.size();\n/* 388 */ convertedArray = new UTF8String[dataLength];\n/* 389 */\n/* 390 */ int loopIndex = 0;\n/* 391 */ while (loopIndex < dataLength) {\n/* 392 */ MapObjects_loopValue2 = (java.lang.String) (value93.apply(loopIndex));\n/* 393 */ MapObjects_loopIsNull3 = MapObjects_loopValue2 == null;\n/* 394 */\n/* 395 */\n/* 396 */ boolean isNull95 = MapObjects_loopIsNull3;\n/* 397 */ final UTF8String value95 = isNull95 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(MapObjects_loopValue2);\n/* 398 */ isNull95 = value95 == null;\n/* 399 */ if (isNull95) {\n/* 400 */ convertedArray[loopIndex] = null;\n/* 401 */ } else {\n/* 402 */ convertedArray[loopIndex] = value95;\n/* 403 */ }\n/* 404 */\n/* 405 */ loopIndex += 1;\n/* 406 */ }\n/* 407 */\n/* 408 */ value92 = new org.apache.spark.sql.catalyst.util.GenericArrayData(convertedArray);\n/* 409 */ }\n/* 410 */ if (isNull93) {\n/* 411 */ values1[28] = null;\n/* 412 */ } else {\n/* 413 */ values1[28] = value92;\n/* 414 */ }\n/* 415 */ }\n/* 416 */\n/* 417 */\n/* 418 */ private void apply_8(InternalRow i) {\n/* 419 */\n/* 420 */ boolean isNull76 = MapObjects_loopIsNull1;\n/* 421 */ final scala.Option value76 = isNull76 ? null : (scala.Option) MapObjects_loopValue0.local_orig();\n/* 422 */ isNull76 = value76 == null;\n/* 423 */\n/* 424 */ final boolean isNull75 = isNull76 || value76.isEmpty();\n/* 425 */ boolean value75 = isNull75 ?\n/* 426 */ false : (Boolean) value76.get();\n/* 427 */ if (isNull75) {\n/* 428 */ values1[20] = null;\n/* 429 */ } else {\n/* 430 */ values1[20] = value75;\n/* 431 */ }\n/* 432 */ boolean isNull78 = MapObjects_loopIsNull1;\n/* 433 */ final scala.Option value78 = isNull78 ? null : (scala.Option) MapObjects_loopValue0.local_resp();\n/* 434 */ isNull78 = value78 == null;\n/* 435 */\n/* 436 */ final boolean isNull77 = isNull78 || value78.isEmpty();\n/* 437 */ boolean value77 = isNull77 ?\n/* 438 */ false : (Boolean) value78.get();\n/* 439 */ if (isNull77) {\n/* 440 */ values1[21] = null;\n/* 441 */ } else {\n/* 442 */ values1[21] = value77;\n/* 443 */ }\n/* 444 */ boolean isNull80 = MapObjects_loopIsNull1;\n/* 445 */ final scala.Option value80 = isNull80 ? null : (scala.Option) MapObjects_loopValue0.missed_bytes();\n/* 446 */ isNull80 = value80 == null;\n/* 447 */\n/* 448 */ final boolean isNull79 = isNull80 || value80.isEmpty();\n/* 449 */ long value79 = isNull79 ?\n/* 450 */ -1L : (Long) value80.get();\n/* 451 */ if (isNull79) {\n/* 452 */ values1[22] = null;\n/* 453 */ } else {\n/* 454 */ values1[22] = value79;\n/* 455 */ }\n/* 456 */ }\n/* 457 */\n/* 458 */\n/* 459 */ private void apply1_0(InternalRow i) {\n/* 460 */\n/* 461 */ boolean isNull17 = false;\n/* 462 */ final io.mistnet.analytics.scan.SrcDstGrouped value17 = isNull17 ? null : (io.mistnet.analytics.scan.SrcDstGrouped) value4._2();\n/* 463 */ isNull17 = value17 == null;\n/* 464 */\n/* 465 */ if (isNull17) {\n/* 466 */ throw new RuntimeException(errMsg);\n/* 467 */ }\n/* 468 */\n/* 469 */ boolean isNull15 = false;\n/* 470 */ final java.lang.String value15 = isNull15 ? null : (java.lang.String) value17.src();\n/* 471 */ isNull15 = value15 == null;\n/* 472 */ boolean isNull14 = isNull15;\n/* 473 */ final UTF8String value14 = isNull14 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value15);\n/* 474 */ isNull14 = value14 == null;\n/* 475 */ if (isNull14) {\n/* 476 */ values[0] = null;\n/* 477 */ } else {\n/* 478 */ values[0] = value14;\n/* 479 */ }\n/* 480 */ boolean isNull21 = false;\n/* 481 */ final io.mistnet.analytics.scan.SrcDstGrouped value21 = isNull21 ? null : (io.mistnet.analytics.scan.SrcDstGrouped) value4._2();\n/* 482 */ isNull21 = value21 == null;\n/* 483 */\n/* 484 */ if (isNull21) {\n/* 485 */ throw new RuntimeException(errMsg1);\n/* 486 */ }\n/* 487 */\n/* 488 */ boolean isNull19 = false;\n/* 489 */ final java.lang.String value19 = isNull19 ? null : (java.lang.String) value21.dest();\n/* 490 */ isNull19 = value19 == null;\n/* 491 */ boolean isNull18 = isNull19;\n/* 492 */ final UTF8String value18 = isNull18 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value19);\n/* 493 */ isNull18 = value18 == null;\n/* 494 */ if (isNull18) {\n/* 495 */ values[1] = null;\n/* 496 */ } else {\n/* 497 */ values[1] = value18;\n/* 498 */ }\n/* 499 */ }\n/* 500 */\n/* 501 */\n/* 502 */ private void apply_2(InternalRow i) {\n/* 503 */\n/* 504 */ boolean isNull43 = MapObjects_loopIsNull1;\n/* 505 */ final scala.Option value43 = isNull43 ? null : (scala.Option) MapObjects_loopValue0.user_uuid();\n/* 506 */ isNull43 = value43 == null;\n/* 507 */\n/* 508 */ final boolean isNull42 = isNull43 || value43.isEmpty();\n/* 509 */ java.lang.String value42 = isNull42 ?\n/* 510 */ null : (java.lang.String) value43.get();\n/* 511 */ boolean isNull41 = isNull42;\n/* 512 */ final UTF8String value41 = isNull41 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value42);\n/* 513 */ isNull41 = value41 == null;\n/* 514 */ if (isNull41) {\n/* 515 */ values1[4] = null;\n/* 516 */ } else {\n/* 517 */ values1[4] = value41;\n/* 518 */ }\n/* 519 */ boolean isNull46 = MapObjects_loopIsNull1;\n/* 520 */ final scala.Option value46 = isNull46 ? null : (scala.Option) MapObjects_loopValue0.host_uuid();\n/* 521 */ isNull46 = value46 == null;\n/* 522 */\n/* 523 */ final boolean isNull45 = isNull46 || value46.isEmpty();\n/* 524 */ java.lang.String value45 = isNull45 ?\n/* 525 */ null : (java.lang.String) value46.get();\n/* 526 */ boolean isNull44 = isNull45;\n/* 527 */ final UTF8String value44 = isNull44 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value45);\n/* 528 */ isNull44 = value44 == null;\n/* 529 */ if (isNull44) {\n/* 530 */ values1[5] = null;\n/* 531 */ } else {\n/* 532 */ values1[5] = value44;\n/* 533 */ }\n/* 534 */ }\n/* 535 */\n/* 536 */\n/* 537 */ private void apply_5(InternalRow i) {\n/* 538 */\n/* 539 */ boolean isNull57 = MapObjects_loopIsNull1;\n/* 540 */ final int value57 = isNull57 ? -1 : MapObjects_loopValue0.src_port();\n/* 541 */ if (isNull57) {\n/* 542 */ values1[11] = null;\n/* 543 */ } else {\n/* 544 */ values1[11] = value57;\n/* 545 */ }\n/* 546 */ boolean isNull59 = MapObjects_loopIsNull1;\n/* 547 */ final java.lang.String value59 = isNull59 ? null : (java.lang.String) MapObjects_loopValue0.dest();\n/* 548 */ isNull59 = value59 == null;\n/* 549 */ boolean isNull58 = isNull59;\n/* 550 */ final UTF8String value58 = isNull58 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value59);\n/* 551 */ isNull58 = value58 == null;\n/* 552 */ if (isNull58) {\n/* 553 */ values1[12] = null;\n/* 554 */ } else {\n/* 555 */ values1[12] = value58;\n/* 556 */ }\n/* 557 */ boolean isNull60 = MapObjects_loopIsNull1;\n/* 558 */ final int value60 = isNull60 ? -1 : MapObjects_loopValue0.dest_port();\n/* 559 */ if (isNull60) {\n/* 560 */ values1[13] = null;\n/* 561 */ } else {\n/* 562 */ values1[13] = value60;\n/* 563 */ }\n/* 564 */ boolean isNull62 = MapObjects_loopIsNull1;\n/* 565 */ final java.lang.String value62 = isNull62 ? null : (java.lang.String) MapObjects_loopValue0.proto();\n/* 566 */ isNull62 = value62 == null;\n/* 567 */ boolean isNull61 = isNull62;\n/* 568 */ final UTF8String value61 = isNull61 ? null : org.apache.spark.unsafe.types.UTF8String.fromString(value62);\n/* 569 */ isNull61 = value61 == null;\n/* 570 */ if (isNull61) {\n/* 571 */ values1[14] = null;\n/* 572 */ } else {\n/* 573 */ values1[14] = value61;\n/* 574 */ }\n/* 575 */ }\n/* 576 */\n/* 577 */\n/* 578 */ private void apply_10(InternalRow i) {\n/* 579 */\n/* 580 */ boolean isNull87 = MapObjects_loopIsNull1;\n/* 581 */ final scala.Option value87 = isNull87 ? null : (scala.Option) MapObjects_loopValue0.orig_ip_bytes();\n/* 582 */ isNull87 = value87 == null;\n/* 583 */\n/* 584 */ final boolean isNull86 = isNull87 || value87.isEmpty();\n/* 585 */ long value86 = isNull86 ?\n/* 586 */ -1L : (Long) value87.get();\n/* 587 */ if (isNull86) {\n/* 588 */ values1[25] = null;\n/* 589 */ } else {\n/* 590 */ values1[25] = value86;\n/* 591 */ }\n/* 592 */ boolean isNull89 = MapObjects_loopIsNull1;\n/* 593 */ final scala.Option value89 = isNull89 ? null : (scala.Option) MapObjects_loopValue0.resp_pkts();\n/* 594 */ isNull89 = value89 == null;\n/* 595 */\n/* 596 */ final boolean isNull88 = isNull89 || value89.isEmpty();\n/* 597 */ long value88 = isNull88 ?\n/* 598 */ -1L : (Long) value89.get();\n/* 599 */ if (isNull88) {\n/* 600 */ values1[26] = null;\n/* 601 */ } else {\n/* 602 */ values1[26] = value88;\n/* 603 */ }\n/* 604 */ boolean isNull91 = MapObjects_loopIsNull1;\n/* 605 */ final scala.Option value91 = isNull91 ? null : (scala.Option) MapObjects_loopValue0.resp_ip_bytes();\n/* 606 */ isNull91 = value91 == null;\n/* 607 */\n/* 608 */ final boolean isNull90 = isNull91 || value91.isEmpty();\n/* 609 */ long value90 = isNull90 ?\n/* 610 */ -1L : (Long) value91.get();\n/* 611 */ if (isNull90) {\n/* 612 */ values1[27] = null;\n/* 613 */ } else {\n/* 614 */ values1[27] = value90;\n/* 615 */ }\n/* 616 */ }\n/* 617 */\n/* 618 */\n/* 619 */ public SpecificMutableProjection(Object[] references) {\n/* 620 */ this.references = references;\n/* 621 */ mutableRow = new org.apache.spark.sql.catalyst.expressions.GenericMutableRow(2);\n/* 622 */ this.values = null;\n/* 623 */ this.errMsg = (java.lang.String) references[3];\n/* 624 */ this.errMsg1 = (java.lang.String) references[4];\n/* 625 */\n/* 626 */\n/* 627 */ this.errMsg2 = (java.lang.String) references[5];\n/* 628 */ this.values1 = null;\n/* 629 */\n/* 630 */\n/* 631 */ this.isNull_0 = true;\n/* 632 */ this.value_0 = false;\n/* 633 */ this.isNull_1 = true;\n/* 634 */ this.value_1 = null;\n/* 635 */ }\n/* 636 */\n/* 637 */ public org.apache.spark.sql.catalyst.expressions.codegen.BaseMutableProjection target(MutableRow row) {\n/* 638 */ mutableRow = row;\n/* 639 */ return this;\n/* 640 */ }\n/* 641 */\n/* 642 */ /* Provide immutable access to the last projected row. */\n/* 643 */ public InternalRow currentValue() {\n/* 644 */ return (InternalRow) mutableRow;\n/* 645 */ }\n/* 646 */\n/* 647 */ public java.lang.Object apply(java.lang.Object _i) {\n/* 648 */ InternalRow i = (InternalRow) _i;\n/* 649 */\n/* 650 */\n/* 651 */\n/* 652 */ Object obj = ((Expression) references[0]).eval(null);\n/* 653 */ scala.Tuple2 value1 = (scala.Tuple2) obj;\n/* 654 */\n/* 655 */ boolean isNull2 = false;\n/* 656 */ final boolean value2 = isNull2 ? false : (Boolean) value1._1();\n/* 657 */ this.isNull_0 = isNull2;\n/* 658 */ this.value_0 = value2;\n/* 659 */\n/* 660 */\n/* 661 */ Object obj1 = ((Expression) references[1]).eval(null);\n/* 662 */ scala.Tuple2 value4 = (scala.Tuple2) obj1;\n/* 663 */\n/* 664 */ boolean isNull8 = false;\n/* 665 */ final io.mistnet.analytics.scan.SrcDstGrouped value8 = isNull8 ? null : (io.mistnet.analytics.scan.SrcDstGrouped) value4._2();\n/* 666 */ isNull8 = value8 == null;\n/* 667 */ boolean isNull6 = false;\n/* 668 */ boolean value6 = true;\n/* 669 */\n/* 670 */ if (!false && isNull8) {\n/* 671 */ } else {\n/* 672 */\n/* 673 */ Object obj2 = ((Expression) references[2]).eval(null);\n/* 674 */ scala.None$ value10 = (scala.None$) obj2;\n/* 675 */\n/* 676 */ boolean isNull11 = false;\n/* 677 */ final io.mistnet.analytics.scan.SrcDstGrouped value11 = isNull11 ? null : (io.mistnet.analytics.scan.SrcDstGrouped) value4._2();\n/* 678 */ isNull11 = value11 == null;\n/* 679 */ boolean isNull9 = false || isNull11;\n/* 680 */ final boolean value9 = isNull9 ? false : value10.equals(value11);\n/* 681 */ if (!isNull9 && value9) {\n/* 682 */ } else if (!false && !isNull9) {\n/* 683 */ value6 = false;\n/* 684 */ } else {\n/* 685 */ isNull6 = true;\n/* 686 */ }\n/* 687 */ }\n/* 688 */ boolean isNull5 = false;\n/* 689 */ InternalRow value5 = null;\n/* 690 */ if (!isNull6 && value6) {\n/* 691 */\n/* 692 */ final InternalRow value12 = null;\n/* 693 */ isNull5 = true;\n/* 694 */ value5 = value12;\n/* 695 */ } else {\n/* 696 */\n/* 697 */ boolean isNull13 = false;\n/* 698 */ this.values = new Object[3];apply1_0(i);\n/* 699 */ apply1_1(i);\n/* 700 */ final InternalRow value13 = new org.apache.spark.sql.catalyst.expressions.GenericInternalRow(values);\n/* 701 */ this.values = null;\n/* 702 */ isNull5 = isNull13;\n/* 703 */ value5 = value13;\n/* 704 */ }\n/* 705 */ this.isNull_1 = isNull5;\n/* 706 */ this.value_1 = value5;\n/* 707 */\n/* 708 */ // copy all the results into MutableRow\n/* 709 */\n/* 710 */ if (!this.isNull_0) {\n/* 711 */ mutableRow.setBoolean(0, this.value_0);\n/* 712 */ } else {\n/* 713 */ mutableRow.setNullAt(0);\n/* 714 */ }\n/* 715 */\n/* 716 */ if (!this.isNull_1) {\n/* 717 */ mutableRow.update(1, this.value_1);\n/* 718 */ } else {\n/* 719 */ mutableRow.setNullAt(1);\n/* 720 */ }\n/* 721 */\n/* 722 */ return mutableRow;\n/* 723 */ }\n/* 724 */ }\n\n\tat org.spark_project.guava.util.concurrent.AbstractFuture$Sync.getValue(AbstractFuture.java:306)\n\tat org.spark_project.guava.util.concurrent.AbstractFuture$Sync.get(AbstractFuture.java:293)\n\tat org.spark_project.guava.util.concurrent.AbstractFuture.get(AbstractFuture.java:116)\n\tat org.spark_project.guava.util.concurrent.Uninterruptibles.getUninterruptibly(Uninterruptibles.java:135)\n\tat org.spark_project.guava.cache.LocalCache$Segment.getAndRecordStats(LocalCache.java:2410)\n\tat org.spark_project.guava.cache.LocalCache$Segment.loadSync(LocalCache.java:2380)\n\tat org.spark_project.guava.cache.LocalCache$Segment.lockedGetOrLoad(LocalCache.java:2342)\n\tat org.spark_project.guava.cache.LocalCache$Segment.get(LocalCache.java:2257)\n\tat org.spark_project.guava.cache.LocalCache.get(LocalCache.java:4000)\n\tat org.spark_project.guava.cache.LocalCache.getOrLoad(LocalCache.java:4004)\n\tat org.spark_project.guava.cache.LocalCache$LocalLoadingCache.get(LocalCache.java:4874)\n\tat org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator$.compile(CodeGenerator.scala:841)\n\tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.create(GenerateMutableProjection.scala:140)\n\tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.generate(GenerateMutableProjection.scala:44)\n\tat org.apache.spark.sql.execution.SparkPlan.newMutableProjection(SparkPlan.scala:369)\n\tat org.apache.spark.sql.execution.aggregate.SortAggregateExec$$anonfun$doExecute$1$$anonfun$3$$anonfun$4.apply(SortAggregateExec.scala:93)\n\tat org.apache.spark.sql.execution.aggregate.SortAggregateExec$$anonfun$doExecute$1$$anonfun$3$$anonfun$4.apply(SortAggregateExec.scala:92)\n\tat org.apache.spark.sql.execution.aggregate.AggregationIterator.(AggregationIterator.scala:143)\n\tat org.apache.spark.sql.execution.aggregate.SortBasedAggregationIterator.(SortBasedAggregationIterator.scala:39)\n\tat org.apache.spark.sql.execution.aggregate.SortAggregateExec$$anonfun$doExecute$1$$anonfun$3.apply(SortAggregateExec.scala:84)\n\tat org.apache.spark.sql.execution.aggregate.SortAggregateExec$$anonfun$doExecute$1$$anonfun$3.apply(SortAggregateExec.scala:75)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:79)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:47)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:86)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{noformat}","issue_id":"13015605","key":"SPARK-18125","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-11-07T11:19:51.000+0000","role":"fixed_distractor","summary":"Spark generated code causes CompileException when groupByKey, reduceGroups and map(_._2) are used"} {"case_id":"13018312","cluster":"DISTRACTOR-SPARK-18281","comments":[{"body":"I'm also seeing the same error with both Python 2.7 and Python 3.5 on Spark 2.0.2 and the Git master when using {{rdd.toLocalIterator()}} or {{df.toLocalIterator()}} for a PySpark RDD and DataFrame, respectively.\n\nOn Spark 1.6.x, {{rdd.toLocalIterator()}} worked correctly.\n\nHere's another example using DataFrames:\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\nit = df.toLocalIterator() # should timeout here with an \"java.net.SocketTimeoutException: Accept timed out\" error\nrow = next(it) # throws an \"Exception: could not open socket\" error\n{code}\n\nResult:\n{code}\nERROR PythonRDD: Error while sending iterator\njava.net.SocketTimeoutException: Accept timed out\n\tat java.net.PlainSocketImpl.socketAccept(Native Method)\n\tat java.net.AbstractPlainSocketImpl.accept(AbstractPlainSocketImpl.java:409)\n\tat java.net.ServerSocket.implAccept(ServerSocket.java:545)\n\tat java.net.ServerSocket.accept(ServerSocket.java:513)\n\tat org.apache.spark.api.python.PythonRDD$$anon$2.run(PythonRDD.scala:697)\n{code}\n\n[~davies] I see that [SPARK-14334 | https://issues.apache.org/jira/browse/SPARK-14334] re-engineered and expanded the {{toLocalIterator}} functionality to DataSets/DataFrames for both Scala/Java & Python. Do you have any thoughts on the issue that is arising now?","created":"2016-12-12T23:34:45.760+0000"},{"body":"[~mwdusenb@us.ibm.com] I can reproduce your issue. I already have the fixing too. If you are not working on this, I will submit a PR for it.\n\nBTW, I can't exactly reproduce the issue reported by [~lminer]:\n\n{code}\nfrom pyspark import SparkContext\nsc = SparkContext()\nrdd = sc.parallelize(range(10))\n[x for x in rdd.toLocalIterator()]\n{code}\n\nBut the following one will be failed:\n{code}\nfrom pyspark import SparkContext\nsc = SparkContext()\nrdd = sc.parallelize(range(10))\nit = rdd.toLocalIterator()\nnext(it)\n{code}\n\nThey are caused by the same bug. I'd fix them together.\n","created":"2016-12-13T03:46:14.416+0000"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16263","created":"2016-12-13T05:57:05.525+0000"},{"body":"[~viirya] Thanks for taking on this bug! I tried out PR, and I'm still running into a socket timeout error for the example I gave above:\n\n{code}\nTraceback (most recent call last):\n File \"\", line 1, in \n File \"/home/mwdusenb/spark/python/pyspark/sql/dataframe.py\", line 416, in toLocalIterator\n peek = next(iter)\n File \"/home/mwdusenb/spark/python/pyspark/rdd.py\", line 140, in _load_from_socket\n for item in serializer.load_stream(rf):\n File \"/home/mwdusenb/spark/python/pyspark/serializers.py\", line 144, in load_stream\n yield self._read_with_length(stream)\n File \"/home/mwdusenb/spark/python/pyspark/serializers.py\", line 161, in _read_with_length\n length = read_int(stream)\n File \"/home/mwdusenb/spark/python/pyspark/serializers.py\", line 555, in read_int\n length = stream.read(4)\n File \"/opt/anaconda3/lib/python3.5/socket.py\", line 575, in readinto\n return self._sock.recv_into(b)\nsocket.timeout: timed out\n{code}\n\nInterestingly, it looks like the {{it = df.toLocalIterator()}} line launches a very large number of {{toLocalIterator}} jobs, and then the Python socket times out while those jobs are running.","created":"2016-12-13T19:20:40.560+0000"},{"body":"Here's another interesting finding. The first (original) example fails with the timeout. However, if you create the DataFrame, do something with it, and then create the iterator, it will work.\n\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\nit = df.toLocalIterator() # FAILS HERE\nrow = next(it)\n{code}\n\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\ndf.count()\nit = df.toLocalIterator() # No longer fails\nrow = next(it)\n{code}\n\nThis leads me to believe there may be something wrong with the creation of the DataFrame.","created":"2016-12-13T19:29:02.279+0000"},{"body":"[~mwdusenb@us.ibm.com] Thanks for reporting this again! I can't reproduce this after applying the PR. However, I think the remaining issue is similar to the change in the PR.\n\nIn JVM side, we just get an iterator of the RDD partitioned results. Once the connection is established, we begin to write elements through the socket. However, if the RDD is not materialized before, the materialization time + network cost might exceed the timeout setting before serving the first element to Python.\n\nThat is why when you materialize the RDD by running {{df.count}}, it will not fail.\n\nI'd change the PR accordingly. May you try it again and see if it solves your tests? Thanks.","created":"2016-12-14T03:42:24.936+0000"},{"body":"[~viirya] Thanks for continuing to work on this! With the latest PR update, the original example now runs correctly when running PySpark in local mode ({{./bin/pyspark}}). However, I then tried the same example on Yarn again ({{./bin/pyspark --master yarn --deploy-mode client}}) and ran into the same socket timeout issues.\n\nI looked into it further and found some interesting findings. In local mode execution, the number of partitions for the DataFrame was {{48}}, while in Yarn execution mode, the same DataFrame started at {{665}} partitions. Looking at the UI, as soon as {{it = df.toLocalIterator()}} is called, it will launch a number of {{toLocalIterator}} jobs equal to the number of partitions, so {{48}} jobs in local mode, and {{665}} in Yarn mode. Currently, the {{it = df.toLocalIterator()}} line blocks until all of those jobs finish. So, if the number of partitions is high enough, the socket timeout will still be triggered.\n\nBelow is a reproducible example (I hope!) in local mode ({{./bin/pyspark}}):\n\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\nit = df.toLocalIterator()\nrow = next(it) # this should work\ndf.rdd.getNumPartitions() # returns `48`\n\n# Now let's break it\ndf2 = df.repartition(700) # increase number of partitions\nit2 = df2.toLocalIterator() # THIS FAILS -> `socket.timeout: timed out`\n{code}","created":"2016-12-14T22:30:19.443+0000"},{"body":"For what its worth I can repro on top of the PR with [~mwdusenb@us.ibm.com]'s example. I think we might be going about fixing this in somewhat of an odd way - the current Python behaviour of toLocalIterator is pretty different than that of Scala, we immediately do a foreach on the Scala iterator which is somewhat strange.\n\nMaybe we could change this to behave more like the Scala toLocalIterator and get rid of these timeouts/hacks around the timeouts. What do people think?","created":"2016-12-14T23:31:27.318+0000"},{"body":"Hi [~holdenk], what you meant for \"we immediately do a foreach on the Scala iterator which is somewhat strange.\"?","created":"2016-12-15T02:31:39.634+0000"},{"body":"[~mwdusenb@us.ibm.com] Thanks for this test case! It is useful to me. However I need to increase the partition number to 1000 to reproduce this issue.\n\nThe additional partitions will increase the time to materialize RDD elements and so cause timeout.\n\nI think we can't set a timeout to the socket reading operation like currently doing as the RDD materialization time is unpredictable. I will keep the connection timeout untouched but unset timeout for socket reading. ","created":"2016-12-15T03:20:00.033+0000"},{"body":"[~mwdusenb@us.ibm.com] BTW, I updated the fixing and if you have time to test it again, that would be great. Thank you.","created":"2016-12-15T03:43:30.426+0000"},{"body":"Issue resolved by pull request 16263\n[https://github.com/apache/spark/pull/16263]","created":"2016-12-20T21:13:08.682+0000"},{"body":"Is this bug really resolved? I am using the latest 2.1.0 release and having the same timeout behaviour as [~mwdusenb@us.ibm.com] described. When using the iterator in local mode it works fine but as soon as moving to cluster it will timeout. I also tested with Mike's example and was able to validate it. Can someone point me to a fix or an alternative with similar functionality?\n\nThanks!","created":"2017-02-25T04:13:21.265+0000"},{"body":"Same here: I see the problem with the latest version too (2.1.0)! The problem appears randomly, for me (I haven't built any minimal example, but the problem looks very much like the one reported here: toLocalIterator used, etc.).","created":"2017-03-12T11:31:38.892+0000"},{"body":"Can you provide some info about your environment? Few reproducible examples we used before are:\n\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\nit = df.toLocalIterator()\nrow = next(it) # this should work\ndf.rdd.getNumPartitions() # returns `48`\n\n# Now let's break it\ndf2 = df.repartition(700) # increase number of partitions\nit2 = df2.toLocalIterator() # THIS FAILS -> `socket.timeout: timed out`\n{code}\n\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\nit = df.toLocalIterator() # FAILS HERE\nrow = next(it)\n{code}\n\nCan you run this examples without failure?\n","created":"2017-03-12T11:41:25.089+0000"},{"body":"Or you have other reproducible examples to test?","created":"2017-03-12T11:42:09.007+0000"},{"body":"Thanks Liang-Chi. Now I do have a minimal example: the small example which is marked above as working is not working on my machine:\n{code}\n~/Downloads/spark-2.1.0-bin-hadoop2.7/bin % PYSPARK_DRIVER_PYTHON=ipython2 ./pyspark\n\nPython 2.7.13 (default, Dec 23 2016, 05:05:58)\nType \"copyright\", \"credits\" or \"license\" for more information.\n\nIPython 5.3.0 -- An enhanced Interactive Python.\n? -> Introduction and overview of IPython's features.\n%quickref -> Quick reference.\nhelp -> Python's own help system.\nobject? -> Details about 'object', use 'object??' for extra details.\n17/03/12 12:46:29 WARN SparkContext: Support for Java 7 is deprecated as of Spark 2.0.0\n2017-03-12 12:46:30.538 java[75598:10832148] Unable to load realm info from SCDynamicStore\n17/03/12 12:46:30 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable\n17/03/12 12:46:33 WARN Utils: Service 'SparkUI' could not bind on port 4040. Attempting port 4041.\n17/03/12 12:46:48 WARN ObjectStore: Failed to get database global_temp, returning NoSuchObjectException\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /__ / .__/\\_,_/_/ /_/\\_\\ version 2.1.0\n /_/\n\nUsing Python version 2.7.13 (default, Dec 23 2016 05:05:58)\nSparkSession available as 'spark'.\nIn [1]: df = spark.createDataFrame([[1],[2],[3]])\n ...: it = df.toLocalIterator()\n ...: row = next(it) # this should work\n ...: df.rdd.getNumPartitions() # returns `48`\n ...: \n\n---------------------------------------------------------------------------\ntimeout Traceback (most recent call last)\n in ()\n 1 df = spark.createDataFrame([[1],[2],[3]])\n 2 it = df.toLocalIterator()\n----> 3 row = next(it) # this should work\n 4 df.rdd.getNumPartitions() # returns `48`\n\n/Users/lebigot/Downloads/spark-2.1.0-bin-hadoop2.7/python/pyspark/rdd.pyc in _load_from_socket(port, serializer)\n 138 try:\n 139 rf = sock.makefile(\"rb\", 65536)\n--> 140 for item in serializer.load_stream(rf):\n 141 yield item\n 142 finally:\n\n/Users/lebigot/Downloads/spark-2.1.0-bin-hadoop2.7/python/pyspark/serializers.pyc in load_stream(self, stream)\n 142 while True:\n 143 try:\n--> 144 yield self._read_with_length(stream)\n 145 except EOFError:\n 146 return\n\n/Users/lebigot/Downloads/spark-2.1.0-bin-hadoop2.7/python/pyspark/serializers.pyc in _read_with_length(self, stream)\n 159\n 160 def _read_with_length(self, stream):\n--> 161 length = read_int(stream)\n 162 if length == SpecialLengths.END_OF_DATA_SECTION:\n 163 raise EOFError\n\n/Users/lebigot/Downloads/spark-2.1.0-bin-hadoop2.7/python/pyspark/serializers.pyc in read_int(stream)\n 553\n 554 def read_int(stream):\n--> 555 length = stream.read(4)\n 556 if not length:\n 557 raise EOFError\n\n/opt/local/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/socket.pyc in read(self, size)\n 382 # fragmentation issues on many platforms.\n 383 try:\n--> 384 data = self._sock.recv(left)\n 385 except error, e:\n 386 if e.args[0] == EINTR:\n\ntimeout: timed out\n{code}\nThe number of partitions is 4.\n\nConfiguration:\n- latest macOS Sierra (10.12.3),\n- IPython, etc. provided through MacPorts,\n- no special Spark configuration except for the verbosity level,\n- nothing else running on my machine (MacBook early 2015).","created":"2017-03-12T11:50:51.779+0000"},{"body":"The second example also fails, but differently:\n{code}\n~/Downloads/spark-2.1.0-bin-hadoop2.7/bin % PYSPARK_DRIVER_PYTHON=ipython2 ./pyspark\n\nPython 2.7.13 (default, Dec 23 2016, 05:05:58)\nType \"copyright\", \"credits\" or \"license\" for more information.\n\nIPython 5.3.0 -- An enhanced Interactive Python.\n? -> Introduction and overview of IPython's features.\n%quickref -> Quick reference.\nhelp -> Python's own help system.\nobject? -> Details about 'object', use 'object??' for extra details.\n17/03/12 13:52:32 WARN SparkContext: Support for Java 7 is deprecated as of Spark 2.0.0\n2017-03-12 13:52:32.493 java[79268:10882992] Unable to load realm info from SCDynamicStore\n17/03/12 13:52:32 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable\n17/03/12 13:52:42 WARN ObjectStore: Failed to get database global_temp, returning NoSuchObjectException\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /__ / .__/\\_,_/_/ /_/\\_\\ version 2.1.0\n /_/\n\nUsing Python version 2.7.13 (default, Dec 23 2016 05:05:58)\nSparkSession available as 'spark'.\n(…)\nIn [2]: df = spark.createDataFrame([[1],[2],[3]])\n ...: \nIn [3]: df.rdd.getNumPartitions()\nOut[3]: 4\nIn [4]: df2 = df.repartition(700) # increase number of partitions\n ...: it2 = df2.toLocalIterator() # THIS FAILS -> `socket.timeout: timed out`\n ...: \n ...: \nIn [5]: 17/03/12 13:53:20 ERROR PythonRDD: Error while sending iterator\njava.net.SocketTimeoutException: Accept timed out\n at java.net.PlainSocketImpl.socketAccept(Native Method)\n at java.net.AbstractPlainSocketImpl.accept(AbstractPlainSocketImpl.java)\n at java.net.ServerSocket.implAccept(ServerSocket.java:530)\n at java.net.ServerSocket.accept(ServerSocket.java:498)\n at org.apache.spark.api.python.PythonRDD$$anon$2.run(PythonRDD.scala:69)\n{code}","created":"2017-03-12T12:54:47.018+0000"},{"body":"Besides, can you also provide the error log?","created":"2017-03-12T13:15:16.971+0000"},{"body":"[~lebigot] Thanks for the error log! It is weird because looks like you run the following old code which is fixed by the PR submitted for this issue: https://github.com/apache/spark/pull/16263\n\n{code}\n/Users/lebigot/Downloads/spark-2.1.0-bin-hadoop2.7/python/pyspark/rdd.pyc in _load_from_socket(port, serializer)\n 138 try:\n 139 rf = sock.makefile(\"rb\", 65536)\n--> 140 for item in serializer.load_stream(rf):\n 141 yield item\n 142 finally:\n{code}\n\nI go to check the detailed changes in Spark 2.1.0: https://issues.apache.org/jira/secure/ReleaseNote.jspa?projectId=12315420&version=12335644\n\nI don't find this issue in the JIRA list. So I think this fixing is not included in Spark 2.1.0 release. \n\n","created":"2017-03-12T13:25:57.064+0000"},{"body":"Oh. Btw, you can see the Fix Version/s of this JIRA is 2.0.3, 2.1.1.","created":"2017-03-12T13:27:32.761+0000"},{"body":"Thanks Liang-Chi. I naively thought that if version 2.0.3 was listed in the fixed version it implied that 2.1 had the fix. I'm looking forward to using the fixed version, then (not sure yet how to do this right now without compiling anything, though).","created":"2017-03-12T15:56:47.983+0000"},{"body":"2.1.1 means 2.1.1 is the first 2.1.x version that has the fix, so, not 2.1.0. 2.0.3 comes after 2.1.0 chronologically.","created":"2017-03-12T16:05:40.164+0000"},{"body":"That is right. So you can try 2.1.1 or latest codebase to test it. Please let me know if this issue happens still. Thanks.","created":"2017-03-13T01:54:04.252+0000"}],"conversations":[{"body":"I run the example straight out of the api docs for toLocalIterator and it gives a time out exception:\n\n{code}\nfrom pyspark import SparkContext\nsc = SparkContext()\nrdd = sc.parallelize(range(10))\n[x for x in rdd.toLocalIterator()]\n{code}\n\nconf file:\nspark.driver.maxResultSize 6G\nspark.executor.extraJavaOptions -XX:+UseG1GC -XX:MaxPermSize=1G -XX:+HeapDumpOnOutOfMemoryError\nspark.executor.memory 16G\nspark.executor.uri foo/spark-2.0.1-bin-hadoop2.7.tgz\nspark.hadoop.fs.s3a.impl org.apache.hadoop.fs.s3a.S3AFileSystem\nspark.hadoop.fs.s3a.buffer.dir /raid0/spark\nspark.hadoop.fs.s3n.buffer.dir /raid0/spark\nspark.hadoop.fs.s3a.connection.timeout 500000\nspark.hadoop.fs.s3n.multipart.uploads.enabled true\nspark.hadoop.mapreduce.fileoutputcommitter.algorithm.version 2\nspark.hadoop.parquet.block.size 2147483648\nspark.hadoop.parquet.enable.summary-metadata false\nspark.jars.packages com.databricks:spark-avro_2.11:3.0.1,com.amazonaws:aws-java-sdk-pom:1.10.34\nspark.local.dir /raid0/spark\nspark.mesos.coarse false\nspark.mesos.constraints priority:1\nspark.network.timeout 600\nspark.rpc.message.maxSize 500\nspark.speculation false\nspark.sql.parquet.mergeSchema false\nspark.sql.planner.externalSort true\nspark.submit.deployMode client\nspark.task.cpus 1\n\nException here:\n{code}\n---------------------------------------------------------------------------\ntimeout Traceback (most recent call last)\n in ()\n 2 sc = SparkContext()\n 3 rdd = sc.parallelize(range(10))\n----> 4 [x for x in rdd.toLocalIterator()]\n\n/foo/spark-2.0.1-bin-hadoop2.7/python/pyspark/rdd.pyc in _load_from_socket(port, serializer)\n 140 try:\n 141 rf = sock.makefile(\"rb\", 65536)\n--> 142 for item in serializer.load_stream(rf):\n 143 yield item\n 144 finally:\n\n/foo/spark-2.0.1-bin-hadoop2.7/python/pyspark/serializers.pyc in load_stream(self, stream)\n 137 while True:\n 138 try:\n--> 139 yield self._read_with_length(stream)\n 140 except EOFError:\n 141 return\n\n/foo/spark-2.0.1-bin-hadoop2.7/python/pyspark/serializers.pyc in _read_with_length(self, stream)\n 154 \n 155 def _read_with_length(self, stream):\n--> 156 length = read_int(stream)\n 157 if length == SpecialLengths.END_OF_DATA_SECTION:\n 158 raise EOFError\n\n/foo/spark-2.0.1-bin-hadoop2.7/python/pyspark/serializers.pyc in read_int(stream)\n 541 \n 542 def read_int(stream):\n--> 543 length = stream.read(4)\n 544 if not length:\n 545 raise EOFError\n\n/usr/lib/python2.7/socket.pyc in read(self, size)\n 378 # fragmentation issues on many platforms.\n 379 try:\n--> 380 data = self._sock.recv(left)\n 381 except error, e:\n 382 if e.args[0] == EINTR:\n\ntimeout: timed out\n{code}\n\n","from":"reporter","subject":"toLocalIterator yields time out error on pyspark2"},{"body":"I'm also seeing the same error with both Python 2.7 and Python 3.5 on Spark 2.0.2 and the Git master when using {{rdd.toLocalIterator()}} or {{df.toLocalIterator()}} for a PySpark RDD and DataFrame, respectively.\n\nOn Spark 1.6.x, {{rdd.toLocalIterator()}} worked correctly.\n\nHere's another example using DataFrames:\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\nit = df.toLocalIterator() # should timeout here with an \"java.net.SocketTimeoutException: Accept timed out\" error\nrow = next(it) # throws an \"Exception: could not open socket\" error\n{code}\n\nResult:\n{code}\nERROR PythonRDD: Error while sending iterator\njava.net.SocketTimeoutException: Accept timed out\n\tat java.net.PlainSocketImpl.socketAccept(Native Method)\n\tat java.net.AbstractPlainSocketImpl.accept(AbstractPlainSocketImpl.java:409)\n\tat java.net.ServerSocket.implAccept(ServerSocket.java:545)\n\tat java.net.ServerSocket.accept(ServerSocket.java:513)\n\tat org.apache.spark.api.python.PythonRDD$$anon$2.run(PythonRDD.scala:697)\n{code}\n\n[~davies] I see that [SPARK-14334 | https://issues.apache.org/jira/browse/SPARK-14334] re-engineered and expanded the {{toLocalIterator}} functionality to DataSets/DataFrames for both Scala/Java & Python. Do you have any thoughts on the issue that is arising now?","from":"developer"},{"body":"[~mwdusenb@us.ibm.com] I can reproduce your issue. I already have the fixing too. If you are not working on this, I will submit a PR for it.\n\nBTW, I can't exactly reproduce the issue reported by [~lminer]:\n\n{code}\nfrom pyspark import SparkContext\nsc = SparkContext()\nrdd = sc.parallelize(range(10))\n[x for x in rdd.toLocalIterator()]\n{code}\n\nBut the following one will be failed:\n{code}\nfrom pyspark import SparkContext\nsc = SparkContext()\nrdd = sc.parallelize(range(10))\nit = rdd.toLocalIterator()\nnext(it)\n{code}\n\nThey are caused by the same bug. I'd fix them together.\n","from":"developer"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16263","from":"developer"},{"body":"[~viirya] Thanks for taking on this bug! I tried out PR, and I'm still running into a socket timeout error for the example I gave above:\n\n{code}\nTraceback (most recent call last):\n File \"\", line 1, in \n File \"/home/mwdusenb/spark/python/pyspark/sql/dataframe.py\", line 416, in toLocalIterator\n peek = next(iter)\n File \"/home/mwdusenb/spark/python/pyspark/rdd.py\", line 140, in _load_from_socket\n for item in serializer.load_stream(rf):\n File \"/home/mwdusenb/spark/python/pyspark/serializers.py\", line 144, in load_stream\n yield self._read_with_length(stream)\n File \"/home/mwdusenb/spark/python/pyspark/serializers.py\", line 161, in _read_with_length\n length = read_int(stream)\n File \"/home/mwdusenb/spark/python/pyspark/serializers.py\", line 555, in read_int\n length = stream.read(4)\n File \"/opt/anaconda3/lib/python3.5/socket.py\", line 575, in readinto\n return self._sock.recv_into(b)\nsocket.timeout: timed out\n{code}\n\nInterestingly, it looks like the {{it = df.toLocalIterator()}} line launches a very large number of {{toLocalIterator}} jobs, and then the Python socket times out while those jobs are running.","from":"developer"},{"body":"Here's another interesting finding. The first (original) example fails with the timeout. However, if you create the DataFrame, do something with it, and then create the iterator, it will work.\n\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\nit = df.toLocalIterator() # FAILS HERE\nrow = next(it)\n{code}\n\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\ndf.count()\nit = df.toLocalIterator() # No longer fails\nrow = next(it)\n{code}\n\nThis leads me to believe there may be something wrong with the creation of the DataFrame.","from":"developer"},{"body":"[~mwdusenb@us.ibm.com] Thanks for reporting this again! I can't reproduce this after applying the PR. However, I think the remaining issue is similar to the change in the PR.\n\nIn JVM side, we just get an iterator of the RDD partitioned results. Once the connection is established, we begin to write elements through the socket. However, if the RDD is not materialized before, the materialization time + network cost might exceed the timeout setting before serving the first element to Python.\n\nThat is why when you materialize the RDD by running {{df.count}}, it will not fail.\n\nI'd change the PR accordingly. May you try it again and see if it solves your tests? Thanks.","from":"developer"},{"body":"[~viirya] Thanks for continuing to work on this! With the latest PR update, the original example now runs correctly when running PySpark in local mode ({{./bin/pyspark}}). However, I then tried the same example on Yarn again ({{./bin/pyspark --master yarn --deploy-mode client}}) and ran into the same socket timeout issues.\n\nI looked into it further and found some interesting findings. In local mode execution, the number of partitions for the DataFrame was {{48}}, while in Yarn execution mode, the same DataFrame started at {{665}} partitions. Looking at the UI, as soon as {{it = df.toLocalIterator()}} is called, it will launch a number of {{toLocalIterator}} jobs equal to the number of partitions, so {{48}} jobs in local mode, and {{665}} in Yarn mode. Currently, the {{it = df.toLocalIterator()}} line blocks until all of those jobs finish. So, if the number of partitions is high enough, the socket timeout will still be triggered.\n\nBelow is a reproducible example (I hope!) in local mode ({{./bin/pyspark}}):\n\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\nit = df.toLocalIterator()\nrow = next(it) # this should work\ndf.rdd.getNumPartitions() # returns `48`\n\n# Now let's break it\ndf2 = df.repartition(700) # increase number of partitions\nit2 = df2.toLocalIterator() # THIS FAILS -> `socket.timeout: timed out`\n{code}","from":"developer"},{"body":"For what its worth I can repro on top of the PR with [~mwdusenb@us.ibm.com]'s example. I think we might be going about fixing this in somewhat of an odd way - the current Python behaviour of toLocalIterator is pretty different than that of Scala, we immediately do a foreach on the Scala iterator which is somewhat strange.\n\nMaybe we could change this to behave more like the Scala toLocalIterator and get rid of these timeouts/hacks around the timeouts. What do people think?","from":"developer"},{"body":"Hi [~holdenk], what you meant for \"we immediately do a foreach on the Scala iterator which is somewhat strange.\"?","from":"developer"},{"body":"[~mwdusenb@us.ibm.com] Thanks for this test case! It is useful to me. However I need to increase the partition number to 1000 to reproduce this issue.\n\nThe additional partitions will increase the time to materialize RDD elements and so cause timeout.\n\nI think we can't set a timeout to the socket reading operation like currently doing as the RDD materialization time is unpredictable. I will keep the connection timeout untouched but unset timeout for socket reading. ","from":"developer"},{"body":"[~mwdusenb@us.ibm.com] BTW, I updated the fixing and if you have time to test it again, that would be great. Thank you.","from":"developer"},{"body":"Issue resolved by pull request 16263\n[https://github.com/apache/spark/pull/16263]","from":"developer"},{"body":"Is this bug really resolved? I am using the latest 2.1.0 release and having the same timeout behaviour as [~mwdusenb@us.ibm.com] described. When using the iterator in local mode it works fine but as soon as moving to cluster it will timeout. I also tested with Mike's example and was able to validate it. Can someone point me to a fix or an alternative with similar functionality?\n\nThanks!","from":"developer"},{"body":"Same here: I see the problem with the latest version too (2.1.0)! The problem appears randomly, for me (I haven't built any minimal example, but the problem looks very much like the one reported here: toLocalIterator used, etc.).","from":"developer"},{"body":"Can you provide some info about your environment? Few reproducible examples we used before are:\n\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\nit = df.toLocalIterator()\nrow = next(it) # this should work\ndf.rdd.getNumPartitions() # returns `48`\n\n# Now let's break it\ndf2 = df.repartition(700) # increase number of partitions\nit2 = df2.toLocalIterator() # THIS FAILS -> `socket.timeout: timed out`\n{code}\n\n{code}\ndf = spark.createDataFrame([[1],[2],[3]])\nit = df.toLocalIterator() # FAILS HERE\nrow = next(it)\n{code}\n\nCan you run this examples without failure?\n","from":"developer"},{"body":"Or you have other reproducible examples to test?","from":"developer"},{"body":"Thanks Liang-Chi. Now I do have a minimal example: the small example which is marked above as working is not working on my machine:\n{code}\n~/Downloads/spark-2.1.0-bin-hadoop2.7/bin % PYSPARK_DRIVER_PYTHON=ipython2 ./pyspark\n\nPython 2.7.13 (default, Dec 23 2016, 05:05:58)\nType \"copyright\", \"credits\" or \"license\" for more information.\n\nIPython 5.3.0 -- An enhanced Interactive Python.\n? -> Introduction and overview of IPython's features.\n%quickref -> Quick reference.\nhelp -> Python's own help system.\nobject? -> Details about 'object', use 'object??' for extra details.\n17/03/12 12:46:29 WARN SparkContext: Support for Java 7 is deprecated as of Spark 2.0.0\n2017-03-12 12:46:30.538 java[75598:10832148] Unable to load realm info from SCDynamicStore\n17/03/12 12:46:30 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable\n17/03/12 12:46:33 WARN Utils: Service 'SparkUI' could not bind on port 4040. Attempting port 4041.\n17/03/12 12:46:48 WARN ObjectStore: Failed to get database global_temp, returning NoSuchObjectException\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /__ / .__/\\_,_/_/ /_/\\_\\ version 2.1.0\n /_/\n\nUsing Python version 2.7.13 (default, Dec 23 2016 05:05:58)\nSparkSession available as 'spark'.\nIn [1]: df = spark.createDataFrame([[1],[2],[3]])\n ...: it = df.toLocalIterator()\n ...: row = next(it) # this should work\n ...: df.rdd.getNumPartitions() # returns `48`\n ...: \n\n---------------------------------------------------------------------------\ntimeout Traceback (most recent call last)\n in ()\n 1 df = spark.createDataFrame([[1],[2],[3]])\n 2 it = df.toLocalIterator()\n----> 3 row = next(it) # this should work\n 4 df.rdd.getNumPartitions() # returns `48`\n\n/Users/lebigot/Downloads/spark-2.1.0-bin-hadoop2.7/python/pyspark/rdd.pyc in _load_from_socket(port, serializer)\n 138 try:\n 139 rf = sock.makefile(\"rb\", 65536)\n--> 140 for item in serializer.load_stream(rf):\n 141 yield item\n 142 finally:\n\n/Users/lebigot/Downloads/spark-2.1.0-bin-hadoop2.7/python/pyspark/serializers.pyc in load_stream(self, stream)\n 142 while True:\n 143 try:\n--> 144 yield self._read_with_length(stream)\n 145 except EOFError:\n 146 return\n\n/Users/lebigot/Downloads/spark-2.1.0-bin-hadoop2.7/python/pyspark/serializers.pyc in _read_with_length(self, stream)\n 159\n 160 def _read_with_length(self, stream):\n--> 161 length = read_int(stream)\n 162 if length == SpecialLengths.END_OF_DATA_SECTION:\n 163 raise EOFError\n\n/Users/lebigot/Downloads/spark-2.1.0-bin-hadoop2.7/python/pyspark/serializers.pyc in read_int(stream)\n 553\n 554 def read_int(stream):\n--> 555 length = stream.read(4)\n 556 if not length:\n 557 raise EOFError\n\n/opt/local/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/socket.pyc in read(self, size)\n 382 # fragmentation issues on many platforms.\n 383 try:\n--> 384 data = self._sock.recv(left)\n 385 except error, e:\n 386 if e.args[0] == EINTR:\n\ntimeout: timed out\n{code}\nThe number of partitions is 4.\n\nConfiguration:\n- latest macOS Sierra (10.12.3),\n- IPython, etc. provided through MacPorts,\n- no special Spark configuration except for the verbosity level,\n- nothing else running on my machine (MacBook early 2015).","from":"developer"},{"body":"The second example also fails, but differently:\n{code}\n~/Downloads/spark-2.1.0-bin-hadoop2.7/bin % PYSPARK_DRIVER_PYTHON=ipython2 ./pyspark\n\nPython 2.7.13 (default, Dec 23 2016, 05:05:58)\nType \"copyright\", \"credits\" or \"license\" for more information.\n\nIPython 5.3.0 -- An enhanced Interactive Python.\n? -> Introduction and overview of IPython's features.\n%quickref -> Quick reference.\nhelp -> Python's own help system.\nobject? -> Details about 'object', use 'object??' for extra details.\n17/03/12 13:52:32 WARN SparkContext: Support for Java 7 is deprecated as of Spark 2.0.0\n2017-03-12 13:52:32.493 java[79268:10882992] Unable to load realm info from SCDynamicStore\n17/03/12 13:52:32 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable\n17/03/12 13:52:42 WARN ObjectStore: Failed to get database global_temp, returning NoSuchObjectException\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /__ / .__/\\_,_/_/ /_/\\_\\ version 2.1.0\n /_/\n\nUsing Python version 2.7.13 (default, Dec 23 2016 05:05:58)\nSparkSession available as 'spark'.\n(…)\nIn [2]: df = spark.createDataFrame([[1],[2],[3]])\n ...: \nIn [3]: df.rdd.getNumPartitions()\nOut[3]: 4\nIn [4]: df2 = df.repartition(700) # increase number of partitions\n ...: it2 = df2.toLocalIterator() # THIS FAILS -> `socket.timeout: timed out`\n ...: \n ...: \nIn [5]: 17/03/12 13:53:20 ERROR PythonRDD: Error while sending iterator\njava.net.SocketTimeoutException: Accept timed out\n at java.net.PlainSocketImpl.socketAccept(Native Method)\n at java.net.AbstractPlainSocketImpl.accept(AbstractPlainSocketImpl.java)\n at java.net.ServerSocket.implAccept(ServerSocket.java:530)\n at java.net.ServerSocket.accept(ServerSocket.java:498)\n at org.apache.spark.api.python.PythonRDD$$anon$2.run(PythonRDD.scala:69)\n{code}","from":"developer"},{"body":"Besides, can you also provide the error log?","from":"developer"},{"body":"[~lebigot] Thanks for the error log! It is weird because looks like you run the following old code which is fixed by the PR submitted for this issue: https://github.com/apache/spark/pull/16263\n\n{code}\n/Users/lebigot/Downloads/spark-2.1.0-bin-hadoop2.7/python/pyspark/rdd.pyc in _load_from_socket(port, serializer)\n 138 try:\n 139 rf = sock.makefile(\"rb\", 65536)\n--> 140 for item in serializer.load_stream(rf):\n 141 yield item\n 142 finally:\n{code}\n\nI go to check the detailed changes in Spark 2.1.0: https://issues.apache.org/jira/secure/ReleaseNote.jspa?projectId=12315420&version=12335644\n\nI don't find this issue in the JIRA list. So I think this fixing is not included in Spark 2.1.0 release. \n\n","from":"developer"},{"body":"Oh. Btw, you can see the Fix Version/s of this JIRA is 2.0.3, 2.1.1.","from":"developer"},{"body":"Thanks Liang-Chi. I naively thought that if version 2.0.3 was listed in the fixed version it implied that 2.1 had the fix. I'm looking forward to using the fixed version, then (not sure yet how to do this right now without compiling anything, though).","from":"developer"},{"body":"2.1.1 means 2.1.1 is the first 2.1.x version that has the fix, so, not 2.1.0. 2.0.3 comes after 2.1.0 chronologically.","from":"developer"},{"body":"That is right. So you can try 2.1.1 or latest codebase to test it. Please let me know if this issue happens still. Thanks.","from":"developer"}],"created":"2016-11-04T23:38:31.000+0000","description":"I run the example straight out of the api docs for toLocalIterator and it gives a time out exception:\n\n{code}\nfrom pyspark import SparkContext\nsc = SparkContext()\nrdd = sc.parallelize(range(10))\n[x for x in rdd.toLocalIterator()]\n{code}\n\nconf file:\nspark.driver.maxResultSize 6G\nspark.executor.extraJavaOptions -XX:+UseG1GC -XX:MaxPermSize=1G -XX:+HeapDumpOnOutOfMemoryError\nspark.executor.memory 16G\nspark.executor.uri foo/spark-2.0.1-bin-hadoop2.7.tgz\nspark.hadoop.fs.s3a.impl org.apache.hadoop.fs.s3a.S3AFileSystem\nspark.hadoop.fs.s3a.buffer.dir /raid0/spark\nspark.hadoop.fs.s3n.buffer.dir /raid0/spark\nspark.hadoop.fs.s3a.connection.timeout 500000\nspark.hadoop.fs.s3n.multipart.uploads.enabled true\nspark.hadoop.mapreduce.fileoutputcommitter.algorithm.version 2\nspark.hadoop.parquet.block.size 2147483648\nspark.hadoop.parquet.enable.summary-metadata false\nspark.jars.packages com.databricks:spark-avro_2.11:3.0.1,com.amazonaws:aws-java-sdk-pom:1.10.34\nspark.local.dir /raid0/spark\nspark.mesos.coarse false\nspark.mesos.constraints priority:1\nspark.network.timeout 600\nspark.rpc.message.maxSize 500\nspark.speculation false\nspark.sql.parquet.mergeSchema false\nspark.sql.planner.externalSort true\nspark.submit.deployMode client\nspark.task.cpus 1\n\nException here:\n{code}\n---------------------------------------------------------------------------\ntimeout Traceback (most recent call last)\n in ()\n 2 sc = SparkContext()\n 3 rdd = sc.parallelize(range(10))\n----> 4 [x for x in rdd.toLocalIterator()]\n\n/foo/spark-2.0.1-bin-hadoop2.7/python/pyspark/rdd.pyc in _load_from_socket(port, serializer)\n 140 try:\n 141 rf = sock.makefile(\"rb\", 65536)\n--> 142 for item in serializer.load_stream(rf):\n 143 yield item\n 144 finally:\n\n/foo/spark-2.0.1-bin-hadoop2.7/python/pyspark/serializers.pyc in load_stream(self, stream)\n 137 while True:\n 138 try:\n--> 139 yield self._read_with_length(stream)\n 140 except EOFError:\n 141 return\n\n/foo/spark-2.0.1-bin-hadoop2.7/python/pyspark/serializers.pyc in _read_with_length(self, stream)\n 154 \n 155 def _read_with_length(self, stream):\n--> 156 length = read_int(stream)\n 157 if length == SpecialLengths.END_OF_DATA_SECTION:\n 158 raise EOFError\n\n/foo/spark-2.0.1-bin-hadoop2.7/python/pyspark/serializers.pyc in read_int(stream)\n 541 \n 542 def read_int(stream):\n--> 543 length = stream.read(4)\n 544 if not length:\n 545 raise EOFError\n\n/usr/lib/python2.7/socket.pyc in read(self, size)\n 378 # fragmentation issues on many platforms.\n 379 try:\n--> 380 data = self._sock.recv(left)\n 381 except error, e:\n 382 if e.args[0] == EINTR:\n\ntimeout: timed out\n{code}\n\n","issue_id":"13018312","key":"SPARK-18281","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-12-20T21:13:08.000+0000","role":"fixed_distractor","summary":"toLocalIterator yields time out error on pyspark2"} {"case_id":"13019420","cluster":"DISTRACTOR-SPARK-18371","comments":[{"body":"I worked the math for [PIDRateEstimator|https://github.com/apache/spark/blob/master/streaming/src/main/scala/org/apache/spark/streaming/scheduler/rate/PIDRateEstimator.scala] and I found that when there are long scheduling delays it increases the 'historicalError' by a LOT, which even when multiplied by integral (=0.2 default), results in a large negative in the formula:\nval newRate = (latestRate - proportional * error - integral * historicalError - derivative * dError).max(minRate)\n\nAs a result, minRate (=100 default) becomes the newRate. Now, when it comes in [DirectKafkaInputDStream|https://github.com/apache/spark/blob/master/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L107] , if you have more than 100 partitions in your kafka topics, you're almost guaranteed to get Math.rounded backpressureRate = 0. Which then [here|https://github.com/apache/spark/blob/master/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L114] leads to returning None. That as a result returns [leaderOffsets|https://github.com/apache/spark/blob/master/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L149] for that batch - hence the giant batch with all records from last batch till head of the queue (leaderOffsets).\n\nProposed solution: Not sure, maybe the minRate could be defaulted to Total number of Partitions in all your kafka topics + some constant. Not sure if anyone has any suggestions to changes in PIDRateEstimator itself.\n\nHere's the math I did from the example in Look_at_batch_at_22_14.png : \n\n*Run1* :\n\nlatestRate = -1\nlatestTime = -1\nlatestError = -1\ntime = 1478587297000\nprocessingDelay = 342000\ndelaySinceUpdate = 1478587298\n\nprocessingRate = (10800000/342000) * 1000 = 31578.94\n\nerror = -1 -31579 = -31580\nhistoricalError = 8 * 31579 / 60000 = 4.21\ndError = (-31580 - 4.21)/ 1478587298 = 0.00002136107254\n\nnewRate = (-1 -(1*-31580) - (0.2*4.21) - (0 * 0.00002136107254)).max(100) = (31578.158).max(100) = 31578.158\n\nlatestTime = 1478587297\nlatestError = 0\nlatestRate = 31578.94\nReturns None - which results in picking up maxRatePerPartition\n\n*Run 2* :\ntime = 1478587615000\nprocessingDelay = 5.3 * 60 * 1000 = 318000\nschedulingDelay = 282000\n\ndelaySinceUpdate = (1478587615000 - 1478587297000) = 318\nprocessingRate = 10800000/318000 * 1000 = 33962.2\nerror = 31578 - 33962 = -2384\nhistoricalError = 282000 * 33962 / 60000 = 159621.4\n\ndError = doesnt matter since multiplied by 0\n\nnewRate = (31578 - (1*-2384) - (0.2*159621) - (0 * dError)).max(100) = (2037.72).max(100) = 2037.72\n\nlatestRate = 2037.72\nlatestError = -2384\nlatestTime = 1478587615000\nReturns newRate = 2037.72\n\n*Run 3* :\ntotalLag = 1972830183\nperpartition lag = 6576100.61\nbackpressureRate = 126000 - 129000\n\ntime = 1478587795000\ndelaySinceUpdate = 1478587795000 - 1478587615000 = 180\n\nprocessingRate = 10800000/180000 * 1000 = 60000\nerror = 2037 - 60000 = -57963\nhistoricalError = 540000 * 60000 / 60000 = 540000\ndError = doesntMatter\n\nnewRate = (2037.72 - (1*-57963) -(0.2*540000) - 0).max(100) = (-48000).max(100) = 100\n\nlatestTime = 1478587795000\nlatestRate = 100\nlatestError = -62384\n\nReturns newRate = 100","created":"2016-11-09T01:01:49.907+0000"},{"body":"Thanks for digging into this. The other thing I noticed when working on\n\nhttps://github.com/apache/spark/pull/15132\n\nis that the return value of getLatestRate was cast to Int, which seems wrong and possibly subject to overflow.\n\nIf you have the ability to test that PR (shouldn't require a spark redeploy, since the kafka jar is standalone), may want to test it out.","created":"2016-11-09T02:59:07.072+0000"},{"body":"I'll try to test it out hopefully soon.","created":"2016-11-09T05:48:14.507+0000"},{"body":"I deep dived into it and found a simple solution. The problem is that [maxRateLimitPerPartition|https://github.com/apache/spark/blob/branch-2.0/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L94] returns {{None}} for an unintended case. {{None}} should only be returned if there is no lag as indicated by this [condition|https://github.com/apache/spark/blob/branch-2.0/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L114]. However, this condition is also true if all backpressureRates are [rounded|https://github.com/apache/spark/blob/branch-2.0/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L107] to zero. I propose a solution, where rounding is omitted at all. This has the nice side-effect that backpressure is more fine-grained and not only an integral multiple of the [batchDuration|https://github.com/apache/spark/blob/branch-2.0/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L117] in seconds. I will open a pull request for it soon.","created":"2017-04-26T13:22:12.329+0000"},{"body":"User 'arzt' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17774","created":"2017-04-26T15:08:03.641+0000"},{"body":"Screenshots:\n\n[before|https://issues.apache.org/jira/secure/attachment/12865156/01.png]\n[after|https://issues.apache.org/jira/secure/attachment/12865158/02.png]","created":"2017-04-26T15:13:50.402+0000"},{"body":"[~seb.arzt] Isnt this the same problem will come with kinesis connectors too?. Looks like your fix is only on the kafka part. We are using kinesis with spark streaming and we are seeing the exact same problem. attaching screenshot for reference.  !Screen Shot 2019-09-16 at 12.27.25 PM.png!","created":"2019-09-18T07:28:14.567+0000"},{"body":"[~rkarthikeyan] at a first glace I cannot find back pressure support in the kinesis receiver yet. I think your problem should be investigated independently. I suggest to create a new ticket with instructions to reproduce your findings.","created":"2019-09-29T13:14:41.601+0000"}],"conversations":[{"body":"When the streaming job is configured with backpressureEnabled=true, it generates a GIANT batch of records if the processing time + scheduled delay is (much) larger than batchDuration. This creates a backlog of records like no other and results in batches queueing for hours until it chews through this giant batch.\nExpectation is that it should reduce the number of records per batch in some time to whatever it can really process.\nAttaching some screen shots where it seems that this issue is quite easily reproducible.","from":"reporter","subject":"Spark Streaming backpressure bug - generates a batch with large number of records"},{"body":"I worked the math for [PIDRateEstimator|https://github.com/apache/spark/blob/master/streaming/src/main/scala/org/apache/spark/streaming/scheduler/rate/PIDRateEstimator.scala] and I found that when there are long scheduling delays it increases the 'historicalError' by a LOT, which even when multiplied by integral (=0.2 default), results in a large negative in the formula:\nval newRate = (latestRate - proportional * error - integral * historicalError - derivative * dError).max(minRate)\n\nAs a result, minRate (=100 default) becomes the newRate. Now, when it comes in [DirectKafkaInputDStream|https://github.com/apache/spark/blob/master/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L107] , if you have more than 100 partitions in your kafka topics, you're almost guaranteed to get Math.rounded backpressureRate = 0. Which then [here|https://github.com/apache/spark/blob/master/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L114] leads to returning None. That as a result returns [leaderOffsets|https://github.com/apache/spark/blob/master/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L149] for that batch - hence the giant batch with all records from last batch till head of the queue (leaderOffsets).\n\nProposed solution: Not sure, maybe the minRate could be defaulted to Total number of Partitions in all your kafka topics + some constant. Not sure if anyone has any suggestions to changes in PIDRateEstimator itself.\n\nHere's the math I did from the example in Look_at_batch_at_22_14.png : \n\n*Run1* :\n\nlatestRate = -1\nlatestTime = -1\nlatestError = -1\ntime = 1478587297000\nprocessingDelay = 342000\ndelaySinceUpdate = 1478587298\n\nprocessingRate = (10800000/342000) * 1000 = 31578.94\n\nerror = -1 -31579 = -31580\nhistoricalError = 8 * 31579 / 60000 = 4.21\ndError = (-31580 - 4.21)/ 1478587298 = 0.00002136107254\n\nnewRate = (-1 -(1*-31580) - (0.2*4.21) - (0 * 0.00002136107254)).max(100) = (31578.158).max(100) = 31578.158\n\nlatestTime = 1478587297\nlatestError = 0\nlatestRate = 31578.94\nReturns None - which results in picking up maxRatePerPartition\n\n*Run 2* :\ntime = 1478587615000\nprocessingDelay = 5.3 * 60 * 1000 = 318000\nschedulingDelay = 282000\n\ndelaySinceUpdate = (1478587615000 - 1478587297000) = 318\nprocessingRate = 10800000/318000 * 1000 = 33962.2\nerror = 31578 - 33962 = -2384\nhistoricalError = 282000 * 33962 / 60000 = 159621.4\n\ndError = doesnt matter since multiplied by 0\n\nnewRate = (31578 - (1*-2384) - (0.2*159621) - (0 * dError)).max(100) = (2037.72).max(100) = 2037.72\n\nlatestRate = 2037.72\nlatestError = -2384\nlatestTime = 1478587615000\nReturns newRate = 2037.72\n\n*Run 3* :\ntotalLag = 1972830183\nperpartition lag = 6576100.61\nbackpressureRate = 126000 - 129000\n\ntime = 1478587795000\ndelaySinceUpdate = 1478587795000 - 1478587615000 = 180\n\nprocessingRate = 10800000/180000 * 1000 = 60000\nerror = 2037 - 60000 = -57963\nhistoricalError = 540000 * 60000 / 60000 = 540000\ndError = doesntMatter\n\nnewRate = (2037.72 - (1*-57963) -(0.2*540000) - 0).max(100) = (-48000).max(100) = 100\n\nlatestTime = 1478587795000\nlatestRate = 100\nlatestError = -62384\n\nReturns newRate = 100","from":"developer"},{"body":"Thanks for digging into this. The other thing I noticed when working on\n\nhttps://github.com/apache/spark/pull/15132\n\nis that the return value of getLatestRate was cast to Int, which seems wrong and possibly subject to overflow.\n\nIf you have the ability to test that PR (shouldn't require a spark redeploy, since the kafka jar is standalone), may want to test it out.","from":"developer"},{"body":"I'll try to test it out hopefully soon.","from":"developer"},{"body":"I deep dived into it and found a simple solution. The problem is that [maxRateLimitPerPartition|https://github.com/apache/spark/blob/branch-2.0/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L94] returns {{None}} for an unintended case. {{None}} should only be returned if there is no lag as indicated by this [condition|https://github.com/apache/spark/blob/branch-2.0/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L114]. However, this condition is also true if all backpressureRates are [rounded|https://github.com/apache/spark/blob/branch-2.0/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L107] to zero. I propose a solution, where rounding is omitted at all. This has the nice side-effect that backpressure is more fine-grained and not only an integral multiple of the [batchDuration|https://github.com/apache/spark/blob/branch-2.0/external/kafka-0-8/src/main/scala/org/apache/spark/streaming/kafka/DirectKafkaInputDStream.scala#L117] in seconds. I will open a pull request for it soon.","from":"developer"},{"body":"User 'arzt' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17774","from":"developer"},{"body":"Screenshots:\n\n[before|https://issues.apache.org/jira/secure/attachment/12865156/01.png]\n[after|https://issues.apache.org/jira/secure/attachment/12865158/02.png]","from":"developer"},{"body":"[~seb.arzt] Isnt this the same problem will come with kinesis connectors too?. Looks like your fix is only on the kafka part. We are using kinesis with spark streaming and we are seeing the exact same problem. attaching screenshot for reference.  !Screen Shot 2019-09-16 at 12.27.25 PM.png!","from":"developer"},{"body":"[~rkarthikeyan] at a first glace I cannot find back pressure support in the kinesis receiver yet. I think your problem should be investigated independently. I suggest to create a new ticket with instructions to reproduce your findings.","from":"developer"}],"created":"2016-11-09T00:26:38.000+0000","description":"When the streaming job is configured with backpressureEnabled=true, it generates a GIANT batch of records if the processing time + scheduled delay is (much) larger than batchDuration. This creates a backlog of records like no other and results in batches queueing for hours until it chews through this giant batch.\nExpectation is that it should reduce the number of records per batch in some time to whatever it can really process.\nAttaching some screen shots where it seems that this issue is quite easily reproducible.","issue_id":"13019420","key":"SPARK-18371","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-03-16T17:33:06.000+0000","role":"fixed_distractor","summary":"Spark Streaming backpressure bug - generates a batch with large number of records"} {"case_id":"13019922","cluster":"DISTRACTOR-SPARK-18403","comments":[{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15845","created":"2016-11-10T20:29:04.474+0000"},{"body":"Please make sure we enable it.\n","created":"2016-11-10T21:44:29.516+0000"},{"body":"Here is a minimal test case (add it to {{ObjectHashAggregateSuite}}) that can be used to reproduce this issue steadily:\n{code}\ntest(\"oom\") {\n withSQLConf(\n SQLConf.USE_OBJECT_HASH_AGG.key -> \"true\",\n SQLConf.OBJECT_AGG_SORT_BASED_FALLBACK_THRESHOLD.key -> \"1\"\n ) {\n Seq(Tuple1(Seq.empty[Int]))\n .toDF(\"c0\")\n .groupBy(lit(1))\n .agg(typed_count($\"c0\"), max($\"c0\"))\n .show()\n }\n}\n{code}\nWhat I observed is that the partial aggregation phase produces a malformed {{UnsafeRow}} after applying the {{resultProjection}} [here|https://github.com/apache/spark/blob/07beb5d21c6803e80733149f1560c71cd3cacc86/sql/core/src/main/scala/org/apache/spark/sql/execution/aggregate/AggregationIterator.scala#L254].\n\nWhen printed, the malformed {{UnsafeRow}} is always\n{noformat}\n[0,0,2000000008,2800000008,100000000000000,5a5a5a5a5a5a5a5a]\n{noformat}\nThe {{5a5a5a5a5a5a5a5a}} is interpreted as the length of an {{ArrayData}}. Therefore, the JVM blows up when trying to allocate a huge array to deep copy this {{ArrayData}} at a later phase.\n\n[~sameer] and [~davies], would you mind to have a look at this issue? Thanks!","created":"2016-11-21T20:40:00.999+0000"},{"body":"The 5a5a5a5a5a5a means that the page has been freed clean. This has been added in https://github.com/apache/spark/commit/44c7c62bcfca74c82ffc4e3c53997fff47bfacac\n\nThis means that we are freeing an in-use buffer page.\n\n","created":"2016-11-21T22:21:47.998+0000"},{"body":"Figured it out. It's caused by a false sharing issue inside {{ObjectAggregationIterator}}. In short, after setting an {{UnsafeArrayData}} to an aggregation buffer, which is a safe row, the underlying buffer of the {{UnsafeArrayData}} gets overwritten when iterator steps forward.\n\nHave to say that this issue is pretty hard to debug. The large array allocation blows up the JVM right away and you can't really find the large array in the heap dump since the allocation itself fails. Therefore, all the heap dumps are super small (~70MB) compared to the heap size (3GB for default SBT tests) and you can't find anything useful in the heap dumps.\n\nI'm opening a PR to fix this issue.","created":"2016-11-22T01:53:02.436+0000"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15976","created":"2016-11-22T06:41:05.496+0000"}],"conversations":[{"body":"This test suite fails occasionally on Jenkins due to OOM errors. I've already reproduced it locally but haven't figured out the root cause.\n\nWe should probably disable it temporarily before getting it fixed so that it doesn't break the PR build too often.","from":"reporter","subject":"ObjectHashAggregateSuite is being flaky (occasional OOM errors)"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15845","from":"developer"},{"body":"Please make sure we enable it.\n","from":"developer"},{"body":"Here is a minimal test case (add it to {{ObjectHashAggregateSuite}}) that can be used to reproduce this issue steadily:\n{code}\ntest(\"oom\") {\n withSQLConf(\n SQLConf.USE_OBJECT_HASH_AGG.key -> \"true\",\n SQLConf.OBJECT_AGG_SORT_BASED_FALLBACK_THRESHOLD.key -> \"1\"\n ) {\n Seq(Tuple1(Seq.empty[Int]))\n .toDF(\"c0\")\n .groupBy(lit(1))\n .agg(typed_count($\"c0\"), max($\"c0\"))\n .show()\n }\n}\n{code}\nWhat I observed is that the partial aggregation phase produces a malformed {{UnsafeRow}} after applying the {{resultProjection}} [here|https://github.com/apache/spark/blob/07beb5d21c6803e80733149f1560c71cd3cacc86/sql/core/src/main/scala/org/apache/spark/sql/execution/aggregate/AggregationIterator.scala#L254].\n\nWhen printed, the malformed {{UnsafeRow}} is always\n{noformat}\n[0,0,2000000008,2800000008,100000000000000,5a5a5a5a5a5a5a5a]\n{noformat}\nThe {{5a5a5a5a5a5a5a5a}} is interpreted as the length of an {{ArrayData}}. Therefore, the JVM blows up when trying to allocate a huge array to deep copy this {{ArrayData}} at a later phase.\n\n[~sameer] and [~davies], would you mind to have a look at this issue? Thanks!","from":"developer"},{"body":"The 5a5a5a5a5a5a means that the page has been freed clean. This has been added in https://github.com/apache/spark/commit/44c7c62bcfca74c82ffc4e3c53997fff47bfacac\n\nThis means that we are freeing an in-use buffer page.\n\n","from":"developer"},{"body":"Figured it out. It's caused by a false sharing issue inside {{ObjectAggregationIterator}}. In short, after setting an {{UnsafeArrayData}} to an aggregation buffer, which is a safe row, the underlying buffer of the {{UnsafeArrayData}} gets overwritten when iterator steps forward.\n\nHave to say that this issue is pretty hard to debug. The large array allocation blows up the JVM right away and you can't really find the large array in the heap dump since the allocation itself fails. Therefore, all the heap dumps are super small (~70MB) compared to the heap size (3GB for default SBT tests) and you can't find anything useful in the heap dumps.\n\nI'm opening a PR to fix this issue.","from":"developer"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/15976","from":"developer"}],"created":"2016-11-10T19:34:09.000+0000","description":"This test suite fails occasionally on Jenkins due to OOM errors. I've already reproduced it locally but haven't figured out the root cause.\n\nWe should probably disable it temporarily before getting it fixed so that it doesn't break the PR build too often.","issue_id":"13019922","key":"SPARK-18403","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-11-10T21:44:29.000+0000","role":"fixed_distractor","summary":"ObjectHashAggregateSuite is being flaky (occasional OOM errors)"} {"case_id":"13019949","cluster":"DISTRACTOR-SPARK-18406","comments":[{"body":"Similar issue in doing basic RDD operations (like checking for empty RDD's) within streaming jobs:\n\n17/01/19 16:43:45 WARN BlockManager: Block input-0-1484840671538 replicated to only 0 peer(s) instead of 1 peers\n17/01/19 16:43:46 WARN MetricsHelper: No metrics scope set in thread RecurringTimer - Kinesis Checkpointer - Worker localhost:65fe618d-c0a7-4fea-b710-0f1b5c6498f2, getMetricsScope returning NullMetricsScope.\n17/01/19 16:44:51 ERROR Executor: Exception in task 0.0 in stage 212.0 (TID 4180)\njava.lang.AssertionError: assertion failed\n\tat scala.Predef$.assert(Predef.scala:156)\n\tat org.apache.spark.storage.BlockInfo.checkInvariants(BlockInfoManager.scala:84)\n\tat org.apache.spark.storage.BlockInfo.readerCount_$eq(BlockInfoManager.scala:66)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:362)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:361)\n\tat scala.Option.foreach(Option.scala:257)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:361)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:356)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\n\tat org.apache.spark.storage.BlockInfoManager.releaseAllLocksForTask(BlockInfoManager.scala:356)\n\tat org.apache.spark.storage.BlockManager.releaseAllLocksForTask(BlockManager.scala:646)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:281)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n17/01/19 16:44:51 WARN TaskSetManager: Lost task 0.0 in stage 212.0 (TID 4180, localhost): java.lang.AssertionError: assertion failed\n\tat scala.Predef$.assert(Predef.scala:156)\n\tat org.apache.spark.storage.BlockInfo.checkInvariants(BlockInfoManager.scala:84)\n\tat org.apache.spark.storage.BlockInfo.readerCount_$eq(BlockInfoManager.scala:66)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:362)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:361)\n\tat scala.Option.foreach(Option.scala:257)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:361)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:356)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\n\tat org.apache.spark.storage.BlockInfoManager.releaseAllLocksForTask(BlockInfoManager.scala:356)\n\tat org.apache.spark.storage.BlockManager.releaseAllLocksForTask(BlockManager.scala:646)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:281)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n\n17/01/19 16:44:51 ERROR TaskSetManager: Task 0 in stage 212.0 failed 1 times; aborting job\n17/01/19 16:44:51 ERROR JobScheduler: Error running job streaming job 1484844270000 ms.0\norg.apache.spark.SparkException: An exception was raised by Python:\nTraceback (most recent call last):\n File \"/root/spark/python/lib/pyspark.zip/pyspark/streaming/util.py\", line 65, in call\n r = self.func(t, *rdds)\n File \"/root/spark/python/lib/pyspark.zip/pyspark/streaming/dstream.py\", line 159, in \n func = lambda t, rdd: old_func(rdd)\n File \"/root/streamtest.py\", line 519, in multiplex\n if not flow_rdd.isEmpty():\n File \"/root/spark/python/lib/pyspark.zip/pyspark/rdd.py\", line 1343, in isEmpty\n return self.getNumPartitions() == 0 or len(self.take(1)) == 0\n File \"/root/spark/python/lib/pyspark.zip/pyspark/rdd.py\", line 1310, in take\n res = self.context.runJob(self, takeUpToNumLeft, p)\n File \"/root/spark/python/lib/pyspark.zip/pyspark/context.py\", line 933, in runJob\n port = self._jvm.PythonRDD.runJob(self._jsc.sc(), mappedRDD._jrdd, partitions)\n File \"/root/spark/python/lib/py4j-0.10.3-src.zip/py4j/java_gateway.py\", line 1133, in __call__\n answer, self.gateway_client, self.target_id, self.name)\n File \"/root/spark/python/lib/pyspark.zip/pyspark/sql/utils.py\", line 63, in deco\n return f(*a, **kw)\n File \"/root/spark/python/lib/py4j-0.10.3-src.zip/py4j/protocol.py\", line 319, in get_return_value\n format(target_id, \".\", name), value)\nPy4JJavaError: An error occurred while calling z:org.apache.spark.api.python.PythonRDD.runJob.","created":"2017-01-19T17:24:10.443+0000"},{"body":"The same JIRA is found under https://issues-test.apache.org/jira/browse/SPARK-18406.\nSame issue observed in spark 2.1.0 as well.\n\nThe issue is observed on some simple spark query that compiles into 3 stages. I have some custom RDD being used in this case, and it is registered for persistence. In the compute() method of the custom RDD, I spawn a new thread to compute the data in the background, and then immediately return an abstract iterator (wrapped under a InterruptibleIterator) that gets data from the background computation on demand. The assertion happens when iterator of the parent RDD reaches the end of the data. This issue doesn't always happen when the custom RDD is used in the query, regardless being used once or multiple times.\n\nThe issue is related to the new thread I created which accesses data from the input iterator of parent RDD. The new thread is missing the TSS(thread-specific-storage) for TaskContext. I see BlockInfoManager is using this TSS TaskContext as key to search the storage.\n\nHere is log showing the task ID being unset in this thread: Line 1819: 2017/04/13 15:01:01.674 [Thread-33]: TRACE storage.BlockInfoManager: Task -1024 releasing lock for rdd_25_0\n\nHowever, I have no way to set TSS for my thread now because the method is made protected as below:\nobject TaskContext {\n ...\nprivate[this] val taskContext: ThreadLocal[TaskContext] = new ThreadLocal[TaskContext]\n// Note: protected[spark] instead of private[spark] to prevent the following two from\n // showing up in JavaDoc.\n /** Set the thread local TaskContext. Internal to Spark. */\n protected[spark] def setTaskContext(tc: TaskContext): Unit = taskContext.set(tc)\n\nJust to confirm my theory, I made the TaskContext.setTaskContext public, and called it in the beginning of my thread. The use cases that were failing consistently with assertion on lock-release now run successful in all scenarios I have tried, which include having different number of src/shuffle partitions, number of executors, async vs. sequential execution (for having multiple downstream custom RDDs pulling data from upstream RDD).","created":"2017-04-14T18:51:17.934+0000"},{"body":"I can see how allowing user-level code to call setTaskContext() can fix this issue but it's not ideal because it still places the burden on the end users to call the setTaskContext() method in their code.\n\nInstead, I think a cleaner fix would be to have the CompletionIterator record the task ID when it's instantiated so that the same task ID can be used even if the completion occurs in a different thread (the idea is to reduce our reliance on thread locals: there are reasons why we couldn't completely remove them (API changes), but there are parts of the internals where we can propagate more efficiently).\n\nTo move forward here, my suggestion is that we write a failing regression test based on the description provided by [~yxiao], then experiment on my suggested approach of more explicit threading of task ids into closeable objects when they're first created.\n\nI'm on vacation this week and won't be able to help with this until Monday, April 24th, so someone else will need to help / review if this is urgent.","created":"2017-04-16T04:33:08.235+0000"},{"body":"Thanks Josh for the quick response!\nThis issue is critical to my company's use cases, where for the purpose of performance we have to use custom RDD to take input from multiple parent RDDs, and use existing computation logic (in a black box) in the background to pull the result.\n","created":"2017-04-16T19:16:59.977+0000"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18076","created":"2017-05-23T23:24:04.351+0000"},{"body":"Thanks for the fix. What spark release will have it?\nCan we get a patch on top of spark 2.1.0?","created":"2017-05-23T23:29:36.942+0000"},{"body":"we will backport this to 2.1 and 2.0 later","created":"2017-05-24T07:55:30.726+0000"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18096","created":"2017-05-24T19:06:03.293+0000"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18099","created":"2017-05-24T21:54:03.962+0000"},{"body":"Hello,\nI am experiencing an issue very similar to this. I am currently trying to do a groupByKeyAndWindow() with batch size of 1, window size of 80, and shift size of 1 from data that is being streamed from Kafka (ver 0.10) with Direct Streaming. Every once in a while, I encounter the AssertionError like so:\n\n17/08/03 22:32:19 ERROR org.apache.spark.executor.Executor: Exception in task 0.0 in stage 20936.0 (TID 4409)\njava.lang.AssertionError: assertion failed\n\tat scala.Predef$.assert(Predef.scala:156)\n\tat org.apache.spark.storage.BlockInfo.checkInvariants(BlockInfoManager.scala:84)\n\tat org.apache.spark.storage.BlockInfo.readerCount_$eq(BlockInfoManager.scala:66)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:367)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:366)\n\tat scala.Option.foreach(Option.scala:257)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:366)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:361)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\n\tat org.apache.spark.storage.BlockInfoManager.releaseAllLocksForTask(BlockInfoManager.scala:361)\n\tat org.apache.spark.storage.BlockManager.releaseAllLocksForTask(BlockManager.scala:736)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:342)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:748)\n17/08/03 22:32:19 ERROR org.apache.spark.executor.Executor: Exception in task 0.1 in stage 20936.0 (TID 4410)\njava.lang.AssertionError: assertion failed\n\tat scala.Predef$.assert(Predef.scala:156)\n\tat org.apache.spark.storage.BlockInfo.checkInvariants(BlockInfoManager.scala:84)\n\tat org.apache.spark.storage.BlockInfo.readerCount_$eq(BlockInfoManager.scala:66)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:367)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:366)\n\tat scala.Option.foreach(Option.scala:257)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:366)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:361)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\n\tat org.apache.spark.storage.BlockInfoManager.releaseAllLocksForTask(BlockInfoManager.scala:361)\n\tat org.apache.spark.storage.BlockManager.releaseAllLocksForTask(BlockManager.scala:736)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:342)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:748)\n17/08/03 22:32:19 ERROR org.apache.spark.util.Utils: Uncaught exception in thread stdout writer for /opt/conda/bin/python\njava.lang.AssertionError: assertion failed: Block rdd_30291_0 is not locked for reading\n\tat scala.Predef$.assert(Predef.scala:170)\n\tat org.apache.spark.storage.BlockInfoManager.unlock(BlockInfoManager.scala:299)\n\tat org.apache.spark.storage.BlockManager.releaseLock(BlockManager.scala:720)\n\tat org.apache.spark.storage.BlockManager$$anonfun$1.apply$mcV$sp(BlockManager.scala:516)\n\tat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\n\tat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:37)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:509)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:333)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1954)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala:269)\n17/08/03 22:32:19 ERROR org.apache.spark.util.SparkUncaughtExceptionHandler: Uncaught exception in thread Thread[stdout writer for /opt/conda/bin/python,5,main]\njava.lang.AssertionError: assertion failed: Block rdd_30291_0 is not locked for reading\n\tat scala.Predef$.assert(Predef.scala:170)\n\tat org.apache.spark.storage.BlockInfoManager.unlock(BlockInfoManager.scala:299)\n\tat org.apache.spark.storage.BlockManager.releaseLock(BlockManager.scala:720)\n\tat org.apache.spark.storage.BlockManager$$anonfun$1.apply$mcV$sp(BlockManager.scala:516)\n\tat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\n\tat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:37)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:509)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:333)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1954)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala:269)\n\nwhich also kills the executor. Sometimes a new executor is spawned to pick up where the dead executor left off but sometimes the whole Spark job also crashes due to this error. I'm running version 2.2.0 on Google Dataproc on a single node cluster. ","created":"2017-08-03T23:38:25.731+0000"},{"body":"we still find this issue in Spark 2.2, I check this case change code and find spark version 2.2 contain these change. but our error log:\r\n17/11/03 08:00:10 ERROR Utils: Uncaught exception in thread stdout writer for python\r\njava.lang.AssertionError: assertion failed: Block input-0-1508745006978 is not locked for reading\r\nat scala.Predef$.assert(Predef.scala:170)\r\nat org.apache.spark.storage.BlockInfoManager.unlock(BlockInfoManager.scala:299)\r\nat org.apache.spark.storage.BlockManager.releaseLock(BlockManager.scala:720)\r\nat org.apache.spark.storage.BlockManager$$anonfun$1.apply$mcV$sp(BlockManager.scala:516)\r\nat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\r\nat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\r\nat scala.collection.Iterator$class.foreach(Iterator.scala:893)\r\nat org.apache.spark.util.CompletionIterator.foreach(CompletionIterator.scala:26)\r\nat org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:509)\r\nat org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:333)\r\nat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1954)\r\nat org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala:269)\r\n","created":"2017-11-16T06:51:05.430+0000"},{"body":"This problem still exists in PythonRunner, since python side uses a pre-fetch model to consume the upstream data, and open another thread to serve output data to downstream operators, thus it's possible the Task finishes first and trigger the task cleanup logic, and then the CompletionIterator try to release the write lock it holds on some blocks and found the lock has already been released. I'll submit a PR to bypass the issue later.","created":"2019-05-01T00:12:19.350+0000"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/24542","created":"2019-05-06T23:39:10.737+0000"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/24542","created":"2019-05-06T23:39:26.275+0000"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/24552","created":"2019-05-08T06:13:40.832+0000"},{"body":"User 'rezasafi' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/24670","created":"2019-05-22T02:25:50.050+0000"}],"conversations":[{"body":"The following log comes from a production streaming job where executors periodically die due to uncaught exceptions during block release:\n\n\n{code}\n16/11/07 17:11:06 INFO CoarseGrainedExecutorBackend: Got assigned task 7921\n16/11/07 17:11:06 INFO Executor: Running task 0.0 in stage 2390.0 (TID 7921)\n16/11/07 17:11:06 INFO CoarseGrainedExecutorBackend: Got assigned task 7922\n16/11/07 17:11:06 INFO Executor: Running task 1.0 in stage 2390.0 (TID 7922)\n16/11/07 17:11:06 INFO CoarseGrainedExecutorBackend: Got assigned task 7923\n16/11/07 17:11:06 INFO Executor: Running task 2.0 in stage 2390.0 (TID 7923)\n16/11/07 17:11:06 INFO TorrentBroadcast: Started reading broadcast variable 2721\n16/11/07 17:11:06 INFO CoarseGrainedExecutorBackend: Got assigned task 7924\n16/11/07 17:11:06 INFO Executor: Running task 3.0 in stage 2390.0 (TID 7924)\n16/11/07 17:11:06 INFO MemoryStore: Block broadcast_2721_piece0 stored as bytes in memory (estimated size 5.0 KB, free 4.9 GB)\n16/11/07 17:11:06 INFO TorrentBroadcast: Reading broadcast variable 2721 took 3 ms\n16/11/07 17:11:06 INFO MemoryStore: Block broadcast_2721 stored as values in memory (estimated size 9.4 KB, free 4.9 GB)\n16/11/07 17:11:06 INFO BlockManager: Found block rdd_2741_1 locally\n16/11/07 17:11:06 INFO BlockManager: Found block rdd_2741_3 locally\n16/11/07 17:11:06 INFO BlockManager: Found block rdd_2741_2 locally\n16/11/07 17:11:06 INFO BlockManager: Found block rdd_2741_4 locally\n16/11/07 17:11:06 INFO PythonRunner: Times: total = 2, boot = -566, init = 567, finish = 1\n16/11/07 17:11:06 INFO PythonRunner: Times: total = 7, boot = -540, init = 541, finish = 6\n16/11/07 17:11:06 INFO Executor: Finished task 2.0 in stage 2390.0 (TID 7923). 1429 bytes result sent to driver\n16/11/07 17:11:06 INFO PythonRunner: Times: total = 8, boot = -532, init = 533, finish = 7\n16/11/07 17:11:06 INFO Executor: Finished task 3.0 in stage 2390.0 (TID 7924). 1429 bytes result sent to driver\n16/11/07 17:11:06 ERROR Executor: Exception in task 0.0 in stage 2390.0 (TID 7921)\njava.lang.AssertionError: assertion failed\n\tat scala.Predef$.assert(Predef.scala:165)\n\tat org.apache.spark.storage.BlockInfo.checkInvariants(BlockInfoManager.scala:84)\n\tat org.apache.spark.storage.BlockInfo.readerCount_$eq(BlockInfoManager.scala:66)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:362)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:361)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:361)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:356)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:727)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1157)\n\tat org.apache.spark.storage.BlockInfoManager.releaseAllLocksForTask(BlockInfoManager.scala:356)\n\tat org.apache.spark.storage.BlockManager.releaseAllLocksForTask(BlockManager.scala:646)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:281)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n16/11/07 17:11:06 INFO CoarseGrainedExecutorBackend: Got assigned task 7925\n16/11/07 17:11:06 INFO Executor: Running task 0.1 in stage 2390.0 (TID 7925)\n16/11/07 17:11:06 INFO BlockManager: Found block rdd_2741_1 locally\n16/11/07 17:11:06 INFO PythonRunner: Times: total = 41, boot = -536, init = 576, finish = 1\n16/11/07 17:11:06 INFO Executor: Finished task 1.0 in stage 2390.0 (TID 7922). 1429 bytes result sent to driver\n16/11/07 17:11:06 ERROR Utils: Uncaught exception in thread stdout writer for /databricks/python/bin/python\njava.lang.AssertionError: assertion failed: Block rdd_2741_1 is not locked for reading\n\tat scala.Predef$.assert(Predef.scala:179)\n\tat org.apache.spark.storage.BlockInfoManager.unlock(BlockInfoManager.scala:294)\n\tat org.apache.spark.storage.BlockManager.releaseLock(BlockManager.scala:630)\n\tat org.apache.spark.storage.BlockManager$$anonfun$1.apply$mcV$sp(BlockManager.scala:434)\n\tat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\n\tat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:39)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:727)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:504)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:328)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1882)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala:269)\n16/11/07 17:11:06 ERROR SparkUncaughtExceptionHandler: Uncaught exception in thread Thread[stdout writer for /databricks/python/bin/python,5,main]\njava.lang.AssertionError: assertion failed: Block rdd_2741_1 is not locked for reading\n\tat scala.Predef$.assert(Predef.scala:179)\n\tat org.apache.spark.storage.BlockInfoManager.unlock(BlockInfoManager.scala:294)\n\tat org.apache.spark.storage.BlockManager.releaseLock(BlockManager.scala:630)\n\tat org.apache.spark.storage.BlockManager$$anonfun$1.apply$mcV$sp(BlockManager.scala:434)\n\tat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\n\tat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:39)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:727)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:504)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:328)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1882)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala\n{code}\n\nI think that there's some sort of internal race condition between a task finishing (TID 7921) and automatically releasing locks and between some \"automatically release locks on hitting the end of an iterator\" logic running in a separate thread. The log above came from a production streaming job where executors periodically died with this type of error.","from":"reporter","subject":"Race between end-of-task and completion iterator read lock release"},{"body":"Similar issue in doing basic RDD operations (like checking for empty RDD's) within streaming jobs:\n\n17/01/19 16:43:45 WARN BlockManager: Block input-0-1484840671538 replicated to only 0 peer(s) instead of 1 peers\n17/01/19 16:43:46 WARN MetricsHelper: No metrics scope set in thread RecurringTimer - Kinesis Checkpointer - Worker localhost:65fe618d-c0a7-4fea-b710-0f1b5c6498f2, getMetricsScope returning NullMetricsScope.\n17/01/19 16:44:51 ERROR Executor: Exception in task 0.0 in stage 212.0 (TID 4180)\njava.lang.AssertionError: assertion failed\n\tat scala.Predef$.assert(Predef.scala:156)\n\tat org.apache.spark.storage.BlockInfo.checkInvariants(BlockInfoManager.scala:84)\n\tat org.apache.spark.storage.BlockInfo.readerCount_$eq(BlockInfoManager.scala:66)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:362)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:361)\n\tat scala.Option.foreach(Option.scala:257)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:361)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:356)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\n\tat org.apache.spark.storage.BlockInfoManager.releaseAllLocksForTask(BlockInfoManager.scala:356)\n\tat org.apache.spark.storage.BlockManager.releaseAllLocksForTask(BlockManager.scala:646)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:281)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n17/01/19 16:44:51 WARN TaskSetManager: Lost task 0.0 in stage 212.0 (TID 4180, localhost): java.lang.AssertionError: assertion failed\n\tat scala.Predef$.assert(Predef.scala:156)\n\tat org.apache.spark.storage.BlockInfo.checkInvariants(BlockInfoManager.scala:84)\n\tat org.apache.spark.storage.BlockInfo.readerCount_$eq(BlockInfoManager.scala:66)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:362)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:361)\n\tat scala.Option.foreach(Option.scala:257)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:361)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:356)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\n\tat org.apache.spark.storage.BlockInfoManager.releaseAllLocksForTask(BlockInfoManager.scala:356)\n\tat org.apache.spark.storage.BlockManager.releaseAllLocksForTask(BlockManager.scala:646)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:281)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n\n17/01/19 16:44:51 ERROR TaskSetManager: Task 0 in stage 212.0 failed 1 times; aborting job\n17/01/19 16:44:51 ERROR JobScheduler: Error running job streaming job 1484844270000 ms.0\norg.apache.spark.SparkException: An exception was raised by Python:\nTraceback (most recent call last):\n File \"/root/spark/python/lib/pyspark.zip/pyspark/streaming/util.py\", line 65, in call\n r = self.func(t, *rdds)\n File \"/root/spark/python/lib/pyspark.zip/pyspark/streaming/dstream.py\", line 159, in \n func = lambda t, rdd: old_func(rdd)\n File \"/root/streamtest.py\", line 519, in multiplex\n if not flow_rdd.isEmpty():\n File \"/root/spark/python/lib/pyspark.zip/pyspark/rdd.py\", line 1343, in isEmpty\n return self.getNumPartitions() == 0 or len(self.take(1)) == 0\n File \"/root/spark/python/lib/pyspark.zip/pyspark/rdd.py\", line 1310, in take\n res = self.context.runJob(self, takeUpToNumLeft, p)\n File \"/root/spark/python/lib/pyspark.zip/pyspark/context.py\", line 933, in runJob\n port = self._jvm.PythonRDD.runJob(self._jsc.sc(), mappedRDD._jrdd, partitions)\n File \"/root/spark/python/lib/py4j-0.10.3-src.zip/py4j/java_gateway.py\", line 1133, in __call__\n answer, self.gateway_client, self.target_id, self.name)\n File \"/root/spark/python/lib/pyspark.zip/pyspark/sql/utils.py\", line 63, in deco\n return f(*a, **kw)\n File \"/root/spark/python/lib/py4j-0.10.3-src.zip/py4j/protocol.py\", line 319, in get_return_value\n format(target_id, \".\", name), value)\nPy4JJavaError: An error occurred while calling z:org.apache.spark.api.python.PythonRDD.runJob.","from":"developer"},{"body":"The same JIRA is found under https://issues-test.apache.org/jira/browse/SPARK-18406.\nSame issue observed in spark 2.1.0 as well.\n\nThe issue is observed on some simple spark query that compiles into 3 stages. I have some custom RDD being used in this case, and it is registered for persistence. In the compute() method of the custom RDD, I spawn a new thread to compute the data in the background, and then immediately return an abstract iterator (wrapped under a InterruptibleIterator) that gets data from the background computation on demand. The assertion happens when iterator of the parent RDD reaches the end of the data. This issue doesn't always happen when the custom RDD is used in the query, regardless being used once or multiple times.\n\nThe issue is related to the new thread I created which accesses data from the input iterator of parent RDD. The new thread is missing the TSS(thread-specific-storage) for TaskContext. I see BlockInfoManager is using this TSS TaskContext as key to search the storage.\n\nHere is log showing the task ID being unset in this thread: Line 1819: 2017/04/13 15:01:01.674 [Thread-33]: TRACE storage.BlockInfoManager: Task -1024 releasing lock for rdd_25_0\n\nHowever, I have no way to set TSS for my thread now because the method is made protected as below:\nobject TaskContext {\n ...\nprivate[this] val taskContext: ThreadLocal[TaskContext] = new ThreadLocal[TaskContext]\n// Note: protected[spark] instead of private[spark] to prevent the following two from\n // showing up in JavaDoc.\n /** Set the thread local TaskContext. Internal to Spark. */\n protected[spark] def setTaskContext(tc: TaskContext): Unit = taskContext.set(tc)\n\nJust to confirm my theory, I made the TaskContext.setTaskContext public, and called it in the beginning of my thread. The use cases that were failing consistently with assertion on lock-release now run successful in all scenarios I have tried, which include having different number of src/shuffle partitions, number of executors, async vs. sequential execution (for having multiple downstream custom RDDs pulling data from upstream RDD).","from":"developer"},{"body":"I can see how allowing user-level code to call setTaskContext() can fix this issue but it's not ideal because it still places the burden on the end users to call the setTaskContext() method in their code.\n\nInstead, I think a cleaner fix would be to have the CompletionIterator record the task ID when it's instantiated so that the same task ID can be used even if the completion occurs in a different thread (the idea is to reduce our reliance on thread locals: there are reasons why we couldn't completely remove them (API changes), but there are parts of the internals where we can propagate more efficiently).\n\nTo move forward here, my suggestion is that we write a failing regression test based on the description provided by [~yxiao], then experiment on my suggested approach of more explicit threading of task ids into closeable objects when they're first created.\n\nI'm on vacation this week and won't be able to help with this until Monday, April 24th, so someone else will need to help / review if this is urgent.","from":"developer"},{"body":"Thanks Josh for the quick response!\nThis issue is critical to my company's use cases, where for the purpose of performance we have to use custom RDD to take input from multiple parent RDDs, and use existing computation logic (in a black box) in the background to pull the result.\n","from":"developer"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18076","from":"developer"},{"body":"Thanks for the fix. What spark release will have it?\nCan we get a patch on top of spark 2.1.0?","from":"developer"},{"body":"we will backport this to 2.1 and 2.0 later","from":"developer"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18096","from":"developer"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18099","from":"developer"},{"body":"Hello,\nI am experiencing an issue very similar to this. I am currently trying to do a groupByKeyAndWindow() with batch size of 1, window size of 80, and shift size of 1 from data that is being streamed from Kafka (ver 0.10) with Direct Streaming. Every once in a while, I encounter the AssertionError like so:\n\n17/08/03 22:32:19 ERROR org.apache.spark.executor.Executor: Exception in task 0.0 in stage 20936.0 (TID 4409)\njava.lang.AssertionError: assertion failed\n\tat scala.Predef$.assert(Predef.scala:156)\n\tat org.apache.spark.storage.BlockInfo.checkInvariants(BlockInfoManager.scala:84)\n\tat org.apache.spark.storage.BlockInfo.readerCount_$eq(BlockInfoManager.scala:66)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:367)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:366)\n\tat scala.Option.foreach(Option.scala:257)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:366)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:361)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\n\tat org.apache.spark.storage.BlockInfoManager.releaseAllLocksForTask(BlockInfoManager.scala:361)\n\tat org.apache.spark.storage.BlockManager.releaseAllLocksForTask(BlockManager.scala:736)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:342)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:748)\n17/08/03 22:32:19 ERROR org.apache.spark.executor.Executor: Exception in task 0.1 in stage 20936.0 (TID 4410)\njava.lang.AssertionError: assertion failed\n\tat scala.Predef$.assert(Predef.scala:156)\n\tat org.apache.spark.storage.BlockInfo.checkInvariants(BlockInfoManager.scala:84)\n\tat org.apache.spark.storage.BlockInfo.readerCount_$eq(BlockInfoManager.scala:66)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:367)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:366)\n\tat scala.Option.foreach(Option.scala:257)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:366)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:361)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\n\tat org.apache.spark.storage.BlockInfoManager.releaseAllLocksForTask(BlockInfoManager.scala:361)\n\tat org.apache.spark.storage.BlockManager.releaseAllLocksForTask(BlockManager.scala:736)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:342)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:748)\n17/08/03 22:32:19 ERROR org.apache.spark.util.Utils: Uncaught exception in thread stdout writer for /opt/conda/bin/python\njava.lang.AssertionError: assertion failed: Block rdd_30291_0 is not locked for reading\n\tat scala.Predef$.assert(Predef.scala:170)\n\tat org.apache.spark.storage.BlockInfoManager.unlock(BlockInfoManager.scala:299)\n\tat org.apache.spark.storage.BlockManager.releaseLock(BlockManager.scala:720)\n\tat org.apache.spark.storage.BlockManager$$anonfun$1.apply$mcV$sp(BlockManager.scala:516)\n\tat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\n\tat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:37)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:509)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:333)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1954)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala:269)\n17/08/03 22:32:19 ERROR org.apache.spark.util.SparkUncaughtExceptionHandler: Uncaught exception in thread Thread[stdout writer for /opt/conda/bin/python,5,main]\njava.lang.AssertionError: assertion failed: Block rdd_30291_0 is not locked for reading\n\tat scala.Predef$.assert(Predef.scala:170)\n\tat org.apache.spark.storage.BlockInfoManager.unlock(BlockInfoManager.scala:299)\n\tat org.apache.spark.storage.BlockManager.releaseLock(BlockManager.scala:720)\n\tat org.apache.spark.storage.BlockManager$$anonfun$1.apply$mcV$sp(BlockManager.scala:516)\n\tat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\n\tat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:37)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:893)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:509)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:333)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1954)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala:269)\n\nwhich also kills the executor. Sometimes a new executor is spawned to pick up where the dead executor left off but sometimes the whole Spark job also crashes due to this error. I'm running version 2.2.0 on Google Dataproc on a single node cluster. ","from":"developer"},{"body":"we still find this issue in Spark 2.2, I check this case change code and find spark version 2.2 contain these change. but our error log:\r\n17/11/03 08:00:10 ERROR Utils: Uncaught exception in thread stdout writer for python\r\njava.lang.AssertionError: assertion failed: Block input-0-1508745006978 is not locked for reading\r\nat scala.Predef$.assert(Predef.scala:170)\r\nat org.apache.spark.storage.BlockInfoManager.unlock(BlockInfoManager.scala:299)\r\nat org.apache.spark.storage.BlockManager.releaseLock(BlockManager.scala:720)\r\nat org.apache.spark.storage.BlockManager$$anonfun$1.apply$mcV$sp(BlockManager.scala:516)\r\nat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\r\nat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\r\nat scala.collection.Iterator$class.foreach(Iterator.scala:893)\r\nat org.apache.spark.util.CompletionIterator.foreach(CompletionIterator.scala:26)\r\nat org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:509)\r\nat org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:333)\r\nat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1954)\r\nat org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala:269)\r\n","from":"developer"},{"body":"This problem still exists in PythonRunner, since python side uses a pre-fetch model to consume the upstream data, and open another thread to serve output data to downstream operators, thus it's possible the Task finishes first and trigger the task cleanup logic, and then the CompletionIterator try to release the write lock it holds on some blocks and found the lock has already been released. I'll submit a PR to bypass the issue later.","from":"developer"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/24542","from":"developer"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/24542","from":"developer"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/24552","from":"developer"},{"body":"User 'rezasafi' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/24670","from":"developer"}],"created":"2016-11-10T21:11:00.000+0000","description":"The following log comes from a production streaming job where executors periodically die due to uncaught exceptions during block release:\n\n\n{code}\n16/11/07 17:11:06 INFO CoarseGrainedExecutorBackend: Got assigned task 7921\n16/11/07 17:11:06 INFO Executor: Running task 0.0 in stage 2390.0 (TID 7921)\n16/11/07 17:11:06 INFO CoarseGrainedExecutorBackend: Got assigned task 7922\n16/11/07 17:11:06 INFO Executor: Running task 1.0 in stage 2390.0 (TID 7922)\n16/11/07 17:11:06 INFO CoarseGrainedExecutorBackend: Got assigned task 7923\n16/11/07 17:11:06 INFO Executor: Running task 2.0 in stage 2390.0 (TID 7923)\n16/11/07 17:11:06 INFO TorrentBroadcast: Started reading broadcast variable 2721\n16/11/07 17:11:06 INFO CoarseGrainedExecutorBackend: Got assigned task 7924\n16/11/07 17:11:06 INFO Executor: Running task 3.0 in stage 2390.0 (TID 7924)\n16/11/07 17:11:06 INFO MemoryStore: Block broadcast_2721_piece0 stored as bytes in memory (estimated size 5.0 KB, free 4.9 GB)\n16/11/07 17:11:06 INFO TorrentBroadcast: Reading broadcast variable 2721 took 3 ms\n16/11/07 17:11:06 INFO MemoryStore: Block broadcast_2721 stored as values in memory (estimated size 9.4 KB, free 4.9 GB)\n16/11/07 17:11:06 INFO BlockManager: Found block rdd_2741_1 locally\n16/11/07 17:11:06 INFO BlockManager: Found block rdd_2741_3 locally\n16/11/07 17:11:06 INFO BlockManager: Found block rdd_2741_2 locally\n16/11/07 17:11:06 INFO BlockManager: Found block rdd_2741_4 locally\n16/11/07 17:11:06 INFO PythonRunner: Times: total = 2, boot = -566, init = 567, finish = 1\n16/11/07 17:11:06 INFO PythonRunner: Times: total = 7, boot = -540, init = 541, finish = 6\n16/11/07 17:11:06 INFO Executor: Finished task 2.0 in stage 2390.0 (TID 7923). 1429 bytes result sent to driver\n16/11/07 17:11:06 INFO PythonRunner: Times: total = 8, boot = -532, init = 533, finish = 7\n16/11/07 17:11:06 INFO Executor: Finished task 3.0 in stage 2390.0 (TID 7924). 1429 bytes result sent to driver\n16/11/07 17:11:06 ERROR Executor: Exception in task 0.0 in stage 2390.0 (TID 7921)\njava.lang.AssertionError: assertion failed\n\tat scala.Predef$.assert(Predef.scala:165)\n\tat org.apache.spark.storage.BlockInfo.checkInvariants(BlockInfoManager.scala:84)\n\tat org.apache.spark.storage.BlockInfo.readerCount_$eq(BlockInfoManager.scala:66)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:362)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2$$anonfun$apply$2.apply(BlockInfoManager.scala:361)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:361)\n\tat org.apache.spark.storage.BlockInfoManager$$anonfun$releaseAllLocksForTask$2.apply(BlockInfoManager.scala:356)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:727)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1157)\n\tat org.apache.spark.storage.BlockInfoManager.releaseAllLocksForTask(BlockInfoManager.scala:356)\n\tat org.apache.spark.storage.BlockManager.releaseAllLocksForTask(BlockManager.scala:646)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:281)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n16/11/07 17:11:06 INFO CoarseGrainedExecutorBackend: Got assigned task 7925\n16/11/07 17:11:06 INFO Executor: Running task 0.1 in stage 2390.0 (TID 7925)\n16/11/07 17:11:06 INFO BlockManager: Found block rdd_2741_1 locally\n16/11/07 17:11:06 INFO PythonRunner: Times: total = 41, boot = -536, init = 576, finish = 1\n16/11/07 17:11:06 INFO Executor: Finished task 1.0 in stage 2390.0 (TID 7922). 1429 bytes result sent to driver\n16/11/07 17:11:06 ERROR Utils: Uncaught exception in thread stdout writer for /databricks/python/bin/python\njava.lang.AssertionError: assertion failed: Block rdd_2741_1 is not locked for reading\n\tat scala.Predef$.assert(Predef.scala:179)\n\tat org.apache.spark.storage.BlockInfoManager.unlock(BlockInfoManager.scala:294)\n\tat org.apache.spark.storage.BlockManager.releaseLock(BlockManager.scala:630)\n\tat org.apache.spark.storage.BlockManager$$anonfun$1.apply$mcV$sp(BlockManager.scala:434)\n\tat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\n\tat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:39)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:727)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:504)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:328)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1882)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala:269)\n16/11/07 17:11:06 ERROR SparkUncaughtExceptionHandler: Uncaught exception in thread Thread[stdout writer for /databricks/python/bin/python,5,main]\njava.lang.AssertionError: assertion failed: Block rdd_2741_1 is not locked for reading\n\tat scala.Predef$.assert(Predef.scala:179)\n\tat org.apache.spark.storage.BlockInfoManager.unlock(BlockInfoManager.scala:294)\n\tat org.apache.spark.storage.BlockManager.releaseLock(BlockManager.scala:630)\n\tat org.apache.spark.storage.BlockManager$$anonfun$1.apply$mcV$sp(BlockManager.scala:434)\n\tat org.apache.spark.util.CompletionIterator$$anon$1.completion(CompletionIterator.scala:46)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:35)\n\tat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:39)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:727)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.api.python.PythonRDD$.writeIteratorToStream(PythonRDD.scala:504)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread$$anonfun$run$3.apply(PythonRDD.scala:328)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1882)\n\tat org.apache.spark.api.python.PythonRunner$WriterThread.run(PythonRDD.scala\n{code}\n\nI think that there's some sort of internal race condition between a task finishing (TID 7921) and automatically releasing locks and between some \"automatically release locks on hitting the end of an iterator\" logic running in a separate thread. The log above came from a production streaming job where executors periodically died with this type of error.","issue_id":"13019949","key":"SPARK-18406","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-05-24T07:55:12.000+0000","role":"fixed_distractor","summary":"Race between end-of-task and completion iterator read lock release"} {"case_id":"13022222","cluster":"DISTRACTOR-SPARK-18527","comments":[{"body":"I am interested to work on this.","created":"2016-11-25T20:08:09.882+0000"},{"body":"The implementation of our own percentile function is tracked in SPARK-16282. I don't think that will make 2.1, so it would be great if you can create a PR.","created":"2016-11-26T13:52:03.292+0000"},{"body":"Interesting [~hvanhovell], good to see that \n{code}percentile(a, array){code}\nis aimed to be [covered|https://github.com/apache/spark/pull/14136/files#diff-a15a6f87f9676612c69435953a13ddd3R127] in the own implementation. Indeed that PR seems quite big for a soon release.\n\nMaybe [~dongjoon] could give starting points for this one, related to the very close PR: https://github.com/apache/spark/pull/13930\n\nThanks!","created":"2016-11-28T11:11:05.276+0000"},{"body":"User 'hvanhovell' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16034","created":"2016-11-28T12:06:05.529+0000"},{"body":"This has been fixed by merging the Percentile UDAF into branch-2.1","created":"2016-11-28T22:59:43.955+0000"},{"body":"Thanks [~hvanhovell]. Will be interesting to compare this fix to the native implementation which was also merged into branch-2.1.","created":"2016-11-29T10:38:55.213+0000"},{"body":"We did not merge my PR (which was a hack TBH), but the Percentile implementation. That is the reason why I closed this.","created":"2016-11-29T10:40:31.251+0000"}],"conversations":[{"body":"Same bug as SPARK-16228 but \n{code}_FUNC_(bigint, array) {code}\ninstead of \n{code}_FUNC_(bigint, double){code}\n\nFix of SPARK-16228 only fixes the non-array case that was hit.\n\n{code}\nsql(\"select percentile(value, array(0.5,0.99)) from values 1,2,3 T(value)\")\n{code}\nfails in Spark 2 shell.\n\n\nLonger example\n{code}\ncase class Record(key: Long, value: String)\nval recordsDF = spark.createDataFrame((1 to 100).map(i => Record(i.toLong, s\"val_$i\")))\n\nrecordsDF.createOrReplaceTempView(\"records\")\nsql(\"SELECT percentile(key, Array(0.95, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1)) AS test FROM records\")\norg.apache.spark.sql.AnalysisException: No handler for Hive UDF 'org.apache.hadoop.hive.ql.udf.UDAFPercentile': org.apache.hadoop.hive.ql.exec.NoMatchingMethodException: No matching method for class org.apache.had\noop.hive.ql.udf.UDAFPercentile with (bigint, array). Possible choices: _FUNC_(bigint, array) _FUNC_(bigint, double) ; line 1 pos 7\n at org.apache.hadoop.hive.ql.exec.FunctionRegistry.getMethodInternal(FunctionRegistry.java:1164)\n at org.apache.hadoop.hive.ql.exec.DefaultUDAFEvaluatorResolver.getEvaluatorClass(DefaultUDAFEvaluatorResolver.java:83)\n at org.apache.hadoop.hive.ql.udf.generic.GenericUDAFBridge.getEvaluator(GenericUDAFBridge.java:56)\n at org.apache.hadoop.hive.ql.udf.generic.AbstractGenericUDAFResolver.getEvaluator(AbstractGenericUDAFResolver.java:47){code}","from":"reporter","subject":"UDAFPercentile (bigint, array) needs explicity cast to double"},{"body":"I am interested to work on this.","from":"developer"},{"body":"The implementation of our own percentile function is tracked in SPARK-16282. I don't think that will make 2.1, so it would be great if you can create a PR.","from":"developer"},{"body":"Interesting [~hvanhovell], good to see that \n{code}percentile(a, array){code}\nis aimed to be [covered|https://github.com/apache/spark/pull/14136/files#diff-a15a6f87f9676612c69435953a13ddd3R127] in the own implementation. Indeed that PR seems quite big for a soon release.\n\nMaybe [~dongjoon] could give starting points for this one, related to the very close PR: https://github.com/apache/spark/pull/13930\n\nThanks!","from":"developer"},{"body":"User 'hvanhovell' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16034","from":"developer"},{"body":"This has been fixed by merging the Percentile UDAF into branch-2.1","from":"developer"},{"body":"Thanks [~hvanhovell]. Will be interesting to compare this fix to the native implementation which was also merged into branch-2.1.","from":"developer"},{"body":"We did not merge my PR (which was a hack TBH), but the Percentile implementation. That is the reason why I closed this.","from":"developer"}],"created":"2016-11-21T15:28:41.000+0000","description":"Same bug as SPARK-16228 but \n{code}_FUNC_(bigint, array) {code}\ninstead of \n{code}_FUNC_(bigint, double){code}\n\nFix of SPARK-16228 only fixes the non-array case that was hit.\n\n{code}\nsql(\"select percentile(value, array(0.5,0.99)) from values 1,2,3 T(value)\")\n{code}\nfails in Spark 2 shell.\n\n\nLonger example\n{code}\ncase class Record(key: Long, value: String)\nval recordsDF = spark.createDataFrame((1 to 100).map(i => Record(i.toLong, s\"val_$i\")))\n\nrecordsDF.createOrReplaceTempView(\"records\")\nsql(\"SELECT percentile(key, Array(0.95, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1)) AS test FROM records\")\norg.apache.spark.sql.AnalysisException: No handler for Hive UDF 'org.apache.hadoop.hive.ql.udf.UDAFPercentile': org.apache.hadoop.hive.ql.exec.NoMatchingMethodException: No matching method for class org.apache.had\noop.hive.ql.udf.UDAFPercentile with (bigint, array). Possible choices: _FUNC_(bigint, array) _FUNC_(bigint, double) ; line 1 pos 7\n at org.apache.hadoop.hive.ql.exec.FunctionRegistry.getMethodInternal(FunctionRegistry.java:1164)\n at org.apache.hadoop.hive.ql.exec.DefaultUDAFEvaluatorResolver.getEvaluatorClass(DefaultUDAFEvaluatorResolver.java:83)\n at org.apache.hadoop.hive.ql.udf.generic.GenericUDAFBridge.getEvaluator(GenericUDAFBridge.java:56)\n at org.apache.hadoop.hive.ql.udf.generic.AbstractGenericUDAFResolver.getEvaluator(AbstractGenericUDAFResolver.java:47){code}","issue_id":"13022222","key":"SPARK-18527","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-11-28T22:59:43.000+0000","role":"fixed_distractor","summary":"UDAFPercentile (bigint, array) needs explicity cast to double"} {"case_id":"13022481","cluster":"DISTRACTOR-SPARK-18539","comments":[{"body":"Interesting. Is it valid for invalid `b`?\n{code}\nsc.read\n .schema(StructType(Seq(StructField(\"a\", IntegerType), StructField(\"b\", IntegerType, nullable = true))))\n .load(\"/tmp/test\")\n .createOrReplaceTempView(\"table\")\n{code}","created":"2016-11-26T11:42:57.641+0000"},{"body":"[~dongjoon] I don't know. However, it works fine in previous versions of spark including Spark 2.0.0. I think this is normal to apply schema with nullable fields.","created":"2016-11-27T09:23:02.678+0000"},{"body":"It looks like the predicates are pushed down to Parquet and Parquet is screaming about that.\nI'm not sure the original behavior is desirable.\n\n[~smilegator]. How do you think about the previous behavior?","created":"2016-12-04T22:36:12.156+0000"},{"body":"The parquet filter push-down of Spark 2.x is different from the Spark 1.x. Since Spark 2.0, parquet scan starts using vectorization. Unfortunately, Spark 2.0.0 had a serious bug in filter push-down : https://issues.apache.org/jira/browse/SPARK-15639 \n\nAfter the fix was merged into Spark 2.0.1, you hit the behavior difference. We always respect user-specified schemas. When user-specified schema does not match the actual data schema, the behaviors are not well-defined. You might hit different errors or unexpected results. The behaviors could be different when you choose different formats.\n\nLet me dig it deeper about this specific issues and might provide you a better answer later","created":"2016-12-05T01:48:27.389+0000"},{"body":"Thank you so much!","created":"2016-12-05T03:04:45.752+0000"},{"body":"Thank you for your reply. I think we need to make user-specified schemas (that do not match the actual data schema) work correctly e.g. I have parquet files that are generated every day and have field `a`, in some time I add nullable field `b`, so I don't want to add that field to previous files, but I want my application can read both of them. It's like schema merging.","created":"2016-12-05T05:40:20.278+0000"},{"body":"The use case makes sense. I see now!","created":"2016-12-05T05:46:36.703+0000"},{"body":"Could you turn off `spark.sql.parquet.filterPushdown`? Is that acceptable in your use case?","created":"2016-12-05T05:50:06.346+0000"},{"body":"If I turn off `spark.sql.parquet.filterPushdown` it works correctly. But I would prefer to use filter pushdown optimization in my case.","created":"2016-12-05T06:24:27.030+0000"},{"body":"Basically, we have to know whether a column exists or not before executing/processing a query. Otherwise, we are unable to know whether a filter can be pushed down or not. \n\nTo completely resolve the issue, we need to infer the schema every time when we need to read the external file. Schema inference is very expensive when we need to scan many many (small) files. Thus, it does not make sense for Spark to do it. Right?","created":"2016-12-05T06:27:11.849+0000"},{"body":"Hmm.. How it works when we use schema merging? Does it also very expensive when we need to scan many files, doesn't it?","created":"2016-12-05T06:32:54.887+0000"},{"body":"Yeah. It is very slow when you have many many small parquet files.","created":"2016-12-05T06:35:35.581+0000"},{"body":"FYI, I checked the other formats, CSV and JSON work as expected. No error is issued. Just an empty data set. However, ORC is even worse. It does not tolerate it if users specify a schema with extra non-existent columns, no matter whether it is used in the filter or not.","created":"2016-12-05T06:47:17.822+0000"},{"body":"If we can neglect the performance when we use schema merging, maybe we can neglect the performance when we use user-specified schemas, possibly we should make a new parameter for using this. I think small files are not very big problem, in this case, so the problem is I can't normally use user-specified schemas with filter pushdown optimization. What do you think?","created":"2016-12-05T06:55:23.077+0000"},{"body":"I think this is another reason why we should make user-specified schemas working correctly with parquet and ORC, it would be great have the same behavior for all formats.","created":"2016-12-05T06:58:50.498+0000"},{"body":"The default of `spark.sql.parquet.mergeSchema` is false. To figure out the schema of parquet, the default behavior is based on a much cheaper solution\n\n{noformat}\n // Always tries the summary files first if users don't require a merged schema. In this case,\n // \"_common_metadata\" is more preferable than \"_metadata\" because it doesn't contain row\n // groups information, and could be much smaller for large Parquet files with lots of row\n // groups. If no summary file is available, falls back to some random part-file.\n{noformat}\n\nIt might not always resolve your case.\n\nIf we introduce such a parameter to always infer the schema and compare the inferred schema with user-specified schemas, we can do extra checking and processing on the schema (e.g., type promotion, schema merging between user-specified schema and actual data schema). The scope will be much bigger. This is a design decision. cc [~rxin] [~marmbrus] [~lian cheng] [~cloud_fan] [~ekhliang]\n\n\n\n","created":"2016-12-05T07:32:00.785+0000"},{"body":"Why don't we fix the parquet reader so it can tolerate non-existent columns?\n","created":"2016-12-05T07:36:22.345+0000"},{"body":"The error is from Parquet.\n{noformat}\n16/11/22 17:43:47 ERROR Executor: Exception in task 0.0 in stage 1.0 (TID 1)\njava.lang.IllegalArgumentException: Column [b] was not found in schema!\n\tat org.apache.parquet.Preconditions.checkArgument(Preconditions.java:55)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.getColumnDescriptor(SchemaCompatibilityValidator.java:190)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumn(SchemaCompatibilityValidator.java:178)\n{noformat}\n\nIf users specify the schema, we do not check whether the user-specified schema is right or wrong. We always respect the schema. If the schema contains the non-existent columns, we push down the filters to the Parquet, and then Parquet returns the above error to us.\n\nBelow is the test case you can try.\n\n{noformat}\n Seq(\"parquet\").foreach { format =>\n\n withTempPath { path =>\n Seq((1, \"abc\"), (2, \"hello\")).toDF(\"a\", \"b\").write.format(format).save(path.toString)\n\n // user-specified schema contains nonexistent columns\n val schema = StructType(\n Seq(StructField(\"a\", IntegerType),\n StructField(\"b\", StringType),\n StructField(\"c\", IntegerType)))\n val readDf = spark.read.schema(schema).format(format).load(path.toString)\n\n // Read the table without any filter\n checkAnswer(readDf, Row(1, \"abc\", null) :: Row(2, \"hello\", null) :: Nil)\n // Read the table with a filter on existing columns\n checkAnswer(readDf.filter(\"a < 2\"), Row(1, \"abc\", null) :: Nil)\n\n val e = intercept[SparkException] {\n // Read the table with a filter on nonexistent columns\n readDf.filter(\"c < 2\").show()\n }.getMessage\n assert(e.contains(\"Column [c] was not found in schema\"))\n\n withSQLConf(SQLConf.PARQUET_FILTER_PUSHDOWN_ENABLED.key -> \"false\") {\n checkAnswer(readDf.filter(\"c < 2\"), Nil)\n }\n }\n }\n{noformat}","created":"2016-12-05T07:41:09.787+0000"},{"body":"Haven't looked deeply into this issue, but my hunch is that this is related to https://issues.apache.org/jira/browse/PARQUET-389, which was fixed in parquet-mr 1.9.0, while Spark is still using 1.8 (in 2.1) and 1.7 (in 2.0).","created":"2016-12-05T17:59:59.719+0000"},{"body":"[~lian cheng] [~rxin] We might be able to capture and process the Parquet-issued exception in our reader. [~xwu0226] will submit a very tiny PR to resolve this issue. ","created":"2016-12-05T20:16:29.116+0000"},{"body":"Yes. I have the fix and will submit PR and cc everyone for review.","created":"2016-12-05T20:31:19.382+0000"},{"body":"User 'xwu0226' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16156","created":"2016-12-05T21:52:06.654+0000"},{"body":"As commented on GitHub, there're two issues right now:\n# This bug also affects the normal Parquet reader code path, where {{ParquetRecordReader}} is a 3rd party class closed for modification. Therefore, we can't capture the exception there.\n# [PR #9940|https://github.com/apache/spark/pull/9940] should have already fixed this issue. But somehow it is broken right now.","created":"2016-12-05T23:26:20.992+0000"},{"body":"[~v-gerasimov], [~smilegator], and [~xwu0226], after some investigation, I don't think this is a bug now.\n\nJust tested the master branch using the following test cases:\n{code}\n for {\n useVectorizedReader <- Seq(true, false)\n mergeSchema <- Seq(true, false)\n } {\n test(s\"foo - mergeSchema: $mergeSchema - vectorized: $useVectorizedReader\") {\n withSQLConf(SQLConf.PARQUET_VECTORIZED_READER_ENABLED.key -> useVectorizedReader.toString) {\n withTempPath { dir =>\n val path = dir.getCanonicalPath\n val df = spark.range(1).coalesce(1)\n\n df.selectExpr(\"id AS a\", \"id AS b\").write.parquet(path)\n df.selectExpr(\"id AS a\").write.mode(\"append\").parquet(path)\n\n assertResult(0) {\n spark.read\n .option(\"mergeSchema\", mergeSchema.toString)\n .parquet(path)\n .filter(\"b < 0\")\n .count()\n }\n }\n }\n }\n }\n{code}\nIt turned out that this issue only happens when schema merging is turned off. This also explains why PR #9940 doesn't prevent PARQUET-389: because the trick PR #9940 employs happens during schema merging phase. On the other hand, you can't expect missing columns to be properly read when schema merging is turned off. Therefore, I don't think it's a bug.\n\nThe fix for the snippet mentioned in the ticket description is easy, just add {{.option(\"mergeSchema\", \"true\")}} to enable schema merging.","created":"2016-12-05T23:40:07.999+0000"},{"body":"Please remind me if I missed anything important, otherwise, we can resolve this ticket as \"Not a Problem\".","created":"2016-12-05T23:53:04.890+0000"},{"body":"I think we will hit the issue if we use user-specified schema. Here is what I tried in spark-shell built from master branch:\n{code}\nval df = spark.range(1).coalesce(1)\ndf.selectExpr(\"id AS a\").write.parquet(\"/Users/xinwu/spark-test/data/spark-18539\")\nval schema = StructType(Seq(StructField(\"a\", IntegerType), StructField(\"b\", IntegerType)))\nspark.read.option(\"mergeSchema\", \"true\").schema(schema).parquet(\"/Users/xinwu/spark-test/data/spark-18539\").filter(\"b is null\").count()\n{code}\n\nThe exception is \n{code}\nCaused by: java.lang.IllegalArgumentException: Column [b] was not found in schema!\n at org.apache.parquet.Preconditions.checkArgument(Preconditions.java:55)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.getColumnDescriptor(SchemaCompatibilityValidator.java:181)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumn(SchemaCompatibilityValidator.java:169)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumnFilterPredicate(SchemaCompatibilityValidator.java:151)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:91)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:58)\n at org.apache.parquet.filter2.predicate.Operators$NotEq.accept(Operators.java:194)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:121)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:58)\n at org.apache.parquet.filter2.predicate.Operators$And.accept(Operators.java:308)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validate(SchemaCompatibilityValidator.java:63)\n at org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:59)\n at org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:40)\n at org.apache.parquet.filter2.compat.FilterCompat$FilterPredicateCompat.accept(FilterCompat.java:126)\n at org.apache.parquet.filter2.compat.RowGroupFilter.filterRowGroups(RowGroupFilter.java:46)\n at org.apache.spark.sql.execution.datasources.parquet.SpecificParquetRecordReaderBase.initialize(SpecificParquetRecordReaderBase.java:110)\n at org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.initialize(VectorizedParquetRecordReader.java:109)\n at org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anonfun$buildReader$1.apply(ParquetFileFormat.scala:377)\n{code}\n\nHere I have one parquet file missing column b and query with user-specified schema (a, b). \n\n","created":"2016-12-06T00:45:43.535+0000"},{"body":"Because we respect user-specified schema, we won't infer schema and schema merging won't step in of course.","created":"2016-12-06T01:22:36.185+0000"},{"body":"Actually I am not sure if this is a valid usage. I tend to think it is not. As you specify a non existing column, and the system reports the column was not found in schema. It looks reasonable to me.\n","created":"2016-12-06T01:26:36.949+0000"},{"body":"[~xwu0226], thanks for the new use case!\n\n[~viirya], I do think this is a valid use case as long as all the missing columns are nullable. The only reason that this use case doesn't work right now is PARQUET-389.\n\nI got some vague idea about a possible cleaner fix for this issue. Will post it later.","created":"2016-12-06T01:55:10.196+0000"},{"body":"That's cool.","created":"2016-12-06T02:37:09.299+0000"},{"body":"[~lian cheng], in Parquet's code, looks like a null column still can have its ColumnChunkMetaData. It won't cause problem even before PARQUET-389, because Parquet will check if all values in the chunk are null.\n\nPARQUET-389 resolves the case there is no ColumnChunkMetaData for a column, i.e., the column is missing from the Parquet file.\n\nSo I am not sure is, in a Parquet file, can a nullable column have no ColumnChunkMetaData like you said?\n\nAppreciate if you can clarify it. Thanks.","created":"2016-12-07T03:22:34.075+0000"},{"body":"[~viirya], sorry for the (super) late reply. What I mentioned was a *nullable* column instead of a *null* column. To be more specific, say we have two Parquet files:\n\n- File {{A}} has columns {{}}\n- File {{B}} has columns {{}}, where {{c}} is marked as nullable (or {{optional}} in the term of Parquet)\n\nThen it should be fine to treat these two files as a single dataset with a merged schema {{}} and you should be able to push down predicates involving {{c}}.\n\nBTW, the Parquet community just made a patch release 1.8.2 that includes a fix for PARQUET-389 and we probably will upgrade to 1.8.2 in 2.2.0. Then we'll have a proper fix for this issue and remove the workaround we did while doing schema merging.","created":"2017-01-26T18:34:20.155+0000"},{"body":"[~lian cheng] Yea, I see. The term {{optional}} is more proper and won't cause misunderstanding. If we can upgrade to 1.8.2, that is great we can remove the workaround. The workaround is actually a hacky solution.","created":"2017-01-27T02:26:02.448+0000"},{"body":"SPARK-19409 upgrades parquet-mr to 1.8.2 and fixed this issue.","created":"2017-02-03T19:11:39.692+0000"}],"conversations":[{"body":"{code}\n import org.apache.spark.SparkConf\n import org.apache.spark.sql.SparkSession\n import org.apache.spark.sql.types.DataTypes._\n import org.apache.spark.sql.types.{StructField, StructType}\n\n val sc = SparkSession.builder().config(new SparkConf().setMaster(\"local\")).getOrCreate()\n\n val jsonRDD = sc.sparkContext.parallelize(Seq(\"\"\"{\"a\":1}\"\"\"))\n\n sc.read\n .schema(StructType(Seq(StructField(\"a\", IntegerType))))\n .json(jsonRDD)\n .write\n .parquet(\"/tmp/test\")\n\n sc.read\n .schema(StructType(Seq(StructField(\"a\", IntegerType), StructField(\"b\", IntegerType, nullable = true))))\n .load(\"/tmp/test\")\n .createOrReplaceTempView(\"table\")\n\n sc.sql(\"select b from table where b is not null\").show()\n{code}\nreturns:\n{code}\n16/11/22 17:43:47 ERROR Executor: Exception in task 0.0 in stage 1.0 (TID 1)\njava.lang.IllegalArgumentException: Column [b] was not found in schema!\n\tat org.apache.parquet.Preconditions.checkArgument(Preconditions.java:55)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.getColumnDescriptor(SchemaCompatibilityValidator.java:190)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumn(SchemaCompatibilityValidator.java:178)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumnFilterPredicate(SchemaCompatibilityValidator.java:160)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:100)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:59)\n\tat org.apache.parquet.filter2.predicate.Operators$NotEq.accept(Operators.java:194)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validate(SchemaCompatibilityValidator.java:64)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:59)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:40)\n\tat org.apache.parquet.filter2.compat.FilterCompat$FilterPredicateCompat.accept(FilterCompat.java:126)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.filterRowGroups(RowGroupFilter.java:46)\n\tat org.apache.spark.sql.execution.datasources.parquet.SpecificParquetRecordReaderBase.initialize(SpecificParquetRecordReaderBase.java:110)\n\tat org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.initialize(VectorizedParquetRecordReader.java:109)\n\tat org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anonfun$buildReader$1.apply(ParquetFileFormat.scala:367)\n\tat org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anonfun$buildReader$1.apply(ParquetFileFormat.scala:341)\n\tat org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:116)\n\tat org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:91)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.scan_nextBatch$(Unknown Source)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n\tat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:246)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:240)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:86)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}\nexpected result:\n{code}\n+---+\n| b|\n+---+\n+---+\n{code}\n\nIt works fine in 2.0.0 and 1.6.2. However, if I only select the nonexisting column (without filter) it also works fine.\n\nQuery plan:\n{code}\n== Parsed Logical Plan ==\n'Project ['b]\n+- 'Filter isnotnull('b)\n +- 'UnresolvedRelation `table`\n\n== Analyzed Logical Plan ==\nb: int\nProject [b#8]\n+- Filter isnotnull(b#8)\n +- SubqueryAlias table\n +- Relation[a#7,b#8] parquet\n\n== Optimized Logical Plan ==\nProject [b#8]\n+- Filter isnotnull(b#8)\n +- Relation[a#7,b#8] parquet\n\n== Physical Plan ==\n*Project [b#8]\n+- *Filter isnotnull(b#8)\n +- *BatchedScan parquet [b#8] Format: ParquetFormat, InputPaths: file:/tmp/test, PartitionFilters: [], PushedFilters: [IsNotNull(b)], ReadSchema: struct\n{code}","from":"reporter","subject":"Cannot filter by nonexisting column in parquet file"},{"body":"Interesting. Is it valid for invalid `b`?\n{code}\nsc.read\n .schema(StructType(Seq(StructField(\"a\", IntegerType), StructField(\"b\", IntegerType, nullable = true))))\n .load(\"/tmp/test\")\n .createOrReplaceTempView(\"table\")\n{code}","from":"developer"},{"body":"[~dongjoon] I don't know. However, it works fine in previous versions of spark including Spark 2.0.0. I think this is normal to apply schema with nullable fields.","from":"developer"},{"body":"It looks like the predicates are pushed down to Parquet and Parquet is screaming about that.\nI'm not sure the original behavior is desirable.\n\n[~smilegator]. How do you think about the previous behavior?","from":"developer"},{"body":"The parquet filter push-down of Spark 2.x is different from the Spark 1.x. Since Spark 2.0, parquet scan starts using vectorization. Unfortunately, Spark 2.0.0 had a serious bug in filter push-down : https://issues.apache.org/jira/browse/SPARK-15639 \n\nAfter the fix was merged into Spark 2.0.1, you hit the behavior difference. We always respect user-specified schemas. When user-specified schema does not match the actual data schema, the behaviors are not well-defined. You might hit different errors or unexpected results. The behaviors could be different when you choose different formats.\n\nLet me dig it deeper about this specific issues and might provide you a better answer later","from":"developer"},{"body":"Thank you so much!","from":"developer"},{"body":"Thank you for your reply. I think we need to make user-specified schemas (that do not match the actual data schema) work correctly e.g. I have parquet files that are generated every day and have field `a`, in some time I add nullable field `b`, so I don't want to add that field to previous files, but I want my application can read both of them. It's like schema merging.","from":"developer"},{"body":"The use case makes sense. I see now!","from":"developer"},{"body":"Could you turn off `spark.sql.parquet.filterPushdown`? Is that acceptable in your use case?","from":"developer"},{"body":"If I turn off `spark.sql.parquet.filterPushdown` it works correctly. But I would prefer to use filter pushdown optimization in my case.","from":"developer"},{"body":"Basically, we have to know whether a column exists or not before executing/processing a query. Otherwise, we are unable to know whether a filter can be pushed down or not. \n\nTo completely resolve the issue, we need to infer the schema every time when we need to read the external file. Schema inference is very expensive when we need to scan many many (small) files. Thus, it does not make sense for Spark to do it. Right?","from":"developer"},{"body":"Hmm.. How it works when we use schema merging? Does it also very expensive when we need to scan many files, doesn't it?","from":"developer"},{"body":"Yeah. It is very slow when you have many many small parquet files.","from":"developer"},{"body":"FYI, I checked the other formats, CSV and JSON work as expected. No error is issued. Just an empty data set. However, ORC is even worse. It does not tolerate it if users specify a schema with extra non-existent columns, no matter whether it is used in the filter or not.","from":"developer"},{"body":"If we can neglect the performance when we use schema merging, maybe we can neglect the performance when we use user-specified schemas, possibly we should make a new parameter for using this. I think small files are not very big problem, in this case, so the problem is I can't normally use user-specified schemas with filter pushdown optimization. What do you think?","from":"developer"},{"body":"I think this is another reason why we should make user-specified schemas working correctly with parquet and ORC, it would be great have the same behavior for all formats.","from":"developer"},{"body":"The default of `spark.sql.parquet.mergeSchema` is false. To figure out the schema of parquet, the default behavior is based on a much cheaper solution\n\n{noformat}\n // Always tries the summary files first if users don't require a merged schema. In this case,\n // \"_common_metadata\" is more preferable than \"_metadata\" because it doesn't contain row\n // groups information, and could be much smaller for large Parquet files with lots of row\n // groups. If no summary file is available, falls back to some random part-file.\n{noformat}\n\nIt might not always resolve your case.\n\nIf we introduce such a parameter to always infer the schema and compare the inferred schema with user-specified schemas, we can do extra checking and processing on the schema (e.g., type promotion, schema merging between user-specified schema and actual data schema). The scope will be much bigger. This is a design decision. cc [~rxin] [~marmbrus] [~lian cheng] [~cloud_fan] [~ekhliang]\n\n\n\n","from":"developer"},{"body":"Why don't we fix the parquet reader so it can tolerate non-existent columns?\n","from":"developer"},{"body":"The error is from Parquet.\n{noformat}\n16/11/22 17:43:47 ERROR Executor: Exception in task 0.0 in stage 1.0 (TID 1)\njava.lang.IllegalArgumentException: Column [b] was not found in schema!\n\tat org.apache.parquet.Preconditions.checkArgument(Preconditions.java:55)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.getColumnDescriptor(SchemaCompatibilityValidator.java:190)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumn(SchemaCompatibilityValidator.java:178)\n{noformat}\n\nIf users specify the schema, we do not check whether the user-specified schema is right or wrong. We always respect the schema. If the schema contains the non-existent columns, we push down the filters to the Parquet, and then Parquet returns the above error to us.\n\nBelow is the test case you can try.\n\n{noformat}\n Seq(\"parquet\").foreach { format =>\n\n withTempPath { path =>\n Seq((1, \"abc\"), (2, \"hello\")).toDF(\"a\", \"b\").write.format(format).save(path.toString)\n\n // user-specified schema contains nonexistent columns\n val schema = StructType(\n Seq(StructField(\"a\", IntegerType),\n StructField(\"b\", StringType),\n StructField(\"c\", IntegerType)))\n val readDf = spark.read.schema(schema).format(format).load(path.toString)\n\n // Read the table without any filter\n checkAnswer(readDf, Row(1, \"abc\", null) :: Row(2, \"hello\", null) :: Nil)\n // Read the table with a filter on existing columns\n checkAnswer(readDf.filter(\"a < 2\"), Row(1, \"abc\", null) :: Nil)\n\n val e = intercept[SparkException] {\n // Read the table with a filter on nonexistent columns\n readDf.filter(\"c < 2\").show()\n }.getMessage\n assert(e.contains(\"Column [c] was not found in schema\"))\n\n withSQLConf(SQLConf.PARQUET_FILTER_PUSHDOWN_ENABLED.key -> \"false\") {\n checkAnswer(readDf.filter(\"c < 2\"), Nil)\n }\n }\n }\n{noformat}","from":"developer"},{"body":"Haven't looked deeply into this issue, but my hunch is that this is related to https://issues.apache.org/jira/browse/PARQUET-389, which was fixed in parquet-mr 1.9.0, while Spark is still using 1.8 (in 2.1) and 1.7 (in 2.0).","from":"developer"},{"body":"[~lian cheng] [~rxin] We might be able to capture and process the Parquet-issued exception in our reader. [~xwu0226] will submit a very tiny PR to resolve this issue. ","from":"developer"},{"body":"Yes. I have the fix and will submit PR and cc everyone for review.","from":"developer"},{"body":"User 'xwu0226' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16156","from":"developer"},{"body":"As commented on GitHub, there're two issues right now:\n# This bug also affects the normal Parquet reader code path, where {{ParquetRecordReader}} is a 3rd party class closed for modification. Therefore, we can't capture the exception there.\n# [PR #9940|https://github.com/apache/spark/pull/9940] should have already fixed this issue. But somehow it is broken right now.","from":"developer"},{"body":"[~v-gerasimov], [~smilegator], and [~xwu0226], after some investigation, I don't think this is a bug now.\n\nJust tested the master branch using the following test cases:\n{code}\n for {\n useVectorizedReader <- Seq(true, false)\n mergeSchema <- Seq(true, false)\n } {\n test(s\"foo - mergeSchema: $mergeSchema - vectorized: $useVectorizedReader\") {\n withSQLConf(SQLConf.PARQUET_VECTORIZED_READER_ENABLED.key -> useVectorizedReader.toString) {\n withTempPath { dir =>\n val path = dir.getCanonicalPath\n val df = spark.range(1).coalesce(1)\n\n df.selectExpr(\"id AS a\", \"id AS b\").write.parquet(path)\n df.selectExpr(\"id AS a\").write.mode(\"append\").parquet(path)\n\n assertResult(0) {\n spark.read\n .option(\"mergeSchema\", mergeSchema.toString)\n .parquet(path)\n .filter(\"b < 0\")\n .count()\n }\n }\n }\n }\n }\n{code}\nIt turned out that this issue only happens when schema merging is turned off. This also explains why PR #9940 doesn't prevent PARQUET-389: because the trick PR #9940 employs happens during schema merging phase. On the other hand, you can't expect missing columns to be properly read when schema merging is turned off. Therefore, I don't think it's a bug.\n\nThe fix for the snippet mentioned in the ticket description is easy, just add {{.option(\"mergeSchema\", \"true\")}} to enable schema merging.","from":"developer"},{"body":"Please remind me if I missed anything important, otherwise, we can resolve this ticket as \"Not a Problem\".","from":"developer"},{"body":"I think we will hit the issue if we use user-specified schema. Here is what I tried in spark-shell built from master branch:\n{code}\nval df = spark.range(1).coalesce(1)\ndf.selectExpr(\"id AS a\").write.parquet(\"/Users/xinwu/spark-test/data/spark-18539\")\nval schema = StructType(Seq(StructField(\"a\", IntegerType), StructField(\"b\", IntegerType)))\nspark.read.option(\"mergeSchema\", \"true\").schema(schema).parquet(\"/Users/xinwu/spark-test/data/spark-18539\").filter(\"b is null\").count()\n{code}\n\nThe exception is \n{code}\nCaused by: java.lang.IllegalArgumentException: Column [b] was not found in schema!\n at org.apache.parquet.Preconditions.checkArgument(Preconditions.java:55)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.getColumnDescriptor(SchemaCompatibilityValidator.java:181)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumn(SchemaCompatibilityValidator.java:169)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumnFilterPredicate(SchemaCompatibilityValidator.java:151)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:91)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:58)\n at org.apache.parquet.filter2.predicate.Operators$NotEq.accept(Operators.java:194)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:121)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:58)\n at org.apache.parquet.filter2.predicate.Operators$And.accept(Operators.java:308)\n at org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validate(SchemaCompatibilityValidator.java:63)\n at org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:59)\n at org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:40)\n at org.apache.parquet.filter2.compat.FilterCompat$FilterPredicateCompat.accept(FilterCompat.java:126)\n at org.apache.parquet.filter2.compat.RowGroupFilter.filterRowGroups(RowGroupFilter.java:46)\n at org.apache.spark.sql.execution.datasources.parquet.SpecificParquetRecordReaderBase.initialize(SpecificParquetRecordReaderBase.java:110)\n at org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.initialize(VectorizedParquetRecordReader.java:109)\n at org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anonfun$buildReader$1.apply(ParquetFileFormat.scala:377)\n{code}\n\nHere I have one parquet file missing column b and query with user-specified schema (a, b). \n\n","from":"developer"},{"body":"Because we respect user-specified schema, we won't infer schema and schema merging won't step in of course.","from":"developer"},{"body":"Actually I am not sure if this is a valid usage. I tend to think it is not. As you specify a non existing column, and the system reports the column was not found in schema. It looks reasonable to me.\n","from":"developer"},{"body":"[~xwu0226], thanks for the new use case!\n\n[~viirya], I do think this is a valid use case as long as all the missing columns are nullable. The only reason that this use case doesn't work right now is PARQUET-389.\n\nI got some vague idea about a possible cleaner fix for this issue. Will post it later.","from":"developer"},{"body":"That's cool.","from":"developer"},{"body":"[~lian cheng], in Parquet's code, looks like a null column still can have its ColumnChunkMetaData. It won't cause problem even before PARQUET-389, because Parquet will check if all values in the chunk are null.\n\nPARQUET-389 resolves the case there is no ColumnChunkMetaData for a column, i.e., the column is missing from the Parquet file.\n\nSo I am not sure is, in a Parquet file, can a nullable column have no ColumnChunkMetaData like you said?\n\nAppreciate if you can clarify it. Thanks.","from":"developer"},{"body":"[~viirya], sorry for the (super) late reply. What I mentioned was a *nullable* column instead of a *null* column. To be more specific, say we have two Parquet files:\n\n- File {{A}} has columns {{}}\n- File {{B}} has columns {{}}, where {{c}} is marked as nullable (or {{optional}} in the term of Parquet)\n\nThen it should be fine to treat these two files as a single dataset with a merged schema {{}} and you should be able to push down predicates involving {{c}}.\n\nBTW, the Parquet community just made a patch release 1.8.2 that includes a fix for PARQUET-389 and we probably will upgrade to 1.8.2 in 2.2.0. Then we'll have a proper fix for this issue and remove the workaround we did while doing schema merging.","from":"developer"},{"body":"[~lian cheng] Yea, I see. The term {{optional}} is more proper and won't cause misunderstanding. If we can upgrade to 1.8.2, that is great we can remove the workaround. The workaround is actually a hacky solution.","from":"developer"},{"body":"SPARK-19409 upgrades parquet-mr to 1.8.2 and fixed this issue.","from":"developer"}],"created":"2016-11-22T11:55:53.000+0000","description":"{code}\n import org.apache.spark.SparkConf\n import org.apache.spark.sql.SparkSession\n import org.apache.spark.sql.types.DataTypes._\n import org.apache.spark.sql.types.{StructField, StructType}\n\n val sc = SparkSession.builder().config(new SparkConf().setMaster(\"local\")).getOrCreate()\n\n val jsonRDD = sc.sparkContext.parallelize(Seq(\"\"\"{\"a\":1}\"\"\"))\n\n sc.read\n .schema(StructType(Seq(StructField(\"a\", IntegerType))))\n .json(jsonRDD)\n .write\n .parquet(\"/tmp/test\")\n\n sc.read\n .schema(StructType(Seq(StructField(\"a\", IntegerType), StructField(\"b\", IntegerType, nullable = true))))\n .load(\"/tmp/test\")\n .createOrReplaceTempView(\"table\")\n\n sc.sql(\"select b from table where b is not null\").show()\n{code}\nreturns:\n{code}\n16/11/22 17:43:47 ERROR Executor: Exception in task 0.0 in stage 1.0 (TID 1)\njava.lang.IllegalArgumentException: Column [b] was not found in schema!\n\tat org.apache.parquet.Preconditions.checkArgument(Preconditions.java:55)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.getColumnDescriptor(SchemaCompatibilityValidator.java:190)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumn(SchemaCompatibilityValidator.java:178)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumnFilterPredicate(SchemaCompatibilityValidator.java:160)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:100)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:59)\n\tat org.apache.parquet.filter2.predicate.Operators$NotEq.accept(Operators.java:194)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validate(SchemaCompatibilityValidator.java:64)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:59)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:40)\n\tat org.apache.parquet.filter2.compat.FilterCompat$FilterPredicateCompat.accept(FilterCompat.java:126)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.filterRowGroups(RowGroupFilter.java:46)\n\tat org.apache.spark.sql.execution.datasources.parquet.SpecificParquetRecordReaderBase.initialize(SpecificParquetRecordReaderBase.java:110)\n\tat org.apache.spark.sql.execution.datasources.parquet.VectorizedParquetRecordReader.initialize(VectorizedParquetRecordReader.java:109)\n\tat org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anonfun$buildReader$1.apply(ParquetFileFormat.scala:367)\n\tat org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anonfun$buildReader$1.apply(ParquetFileFormat.scala:341)\n\tat org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.nextIterator(FileScanRDD.scala:116)\n\tat org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:91)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.scan_nextBatch$(Unknown Source)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n\tat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:246)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$4.apply(SparkPlan.scala:240)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$24.apply(RDD.scala:803)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:86)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}\nexpected result:\n{code}\n+---+\n| b|\n+---+\n+---+\n{code}\n\nIt works fine in 2.0.0 and 1.6.2. However, if I only select the nonexisting column (without filter) it also works fine.\n\nQuery plan:\n{code}\n== Parsed Logical Plan ==\n'Project ['b]\n+- 'Filter isnotnull('b)\n +- 'UnresolvedRelation `table`\n\n== Analyzed Logical Plan ==\nb: int\nProject [b#8]\n+- Filter isnotnull(b#8)\n +- SubqueryAlias table\n +- Relation[a#7,b#8] parquet\n\n== Optimized Logical Plan ==\nProject [b#8]\n+- Filter isnotnull(b#8)\n +- Relation[a#7,b#8] parquet\n\n== Physical Plan ==\n*Project [b#8]\n+- *Filter isnotnull(b#8)\n +- *BatchedScan parquet [b#8] Format: ParquetFormat, InputPaths: file:/tmp/test, PartitionFilters: [], PushedFilters: [IsNotNull(b)], ReadSchema: struct\n{code}","issue_id":"13022481","key":"SPARK-18539","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-02-03T19:12:27.000+0000","role":"fixed_distractor","summary":"Cannot filter by nonexisting column in parquet file"} {"case_id":"13028031","cluster":"DISTRACTOR-SPARK-18857","comments":[{"body":"GC logs for 2 spark versions while running the same query","created":"2016-12-14T08:42:11.936+0000"},{"body":"Thank you for reporting, [~vishalagrwal].\nThen, in the Spark side, could you test on Spark 2.0.0 before SPARK-16563?","created":"2016-12-16T11:22:56.593+0000"},{"body":"We are unable to use incremental collect in a spark version before 2.0.2 due the bug spark-18009\n\nWe will have to take 2.0.2 and change this particular class and build from source code.","created":"2016-12-19T05:28:11.669+0000"},{"body":"we have built Spark from 2.0.2 source code by changing SparkExecuteStatementOperation.scala to pre SPARK-16563 version. this version works fine without causing any thrift server issues.","created":"2016-12-26T09:48:34.731+0000"},{"body":"Thank you for testing and sharing that information!","created":"2016-12-30T10:36:29.352+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16440","created":"2016-12-30T12:56:04.036+0000"},{"body":"Hi, [~vishalagrwal].\nI agree with you. This is an important problem.\nAt least, I made a PR as a first attempt. In any ways, I hope this will be resolved soon.","created":"2016-12-30T12:58:16.699+0000"},{"body":"CC [~alicegugu]","created":"2016-12-31T11:51:58.366+0000"},{"body":"Hi [~vishalagrwal].\nCould you test your case with https://github.com/apache/spark/pull/16440 ?\nAlthough I tried to address the iterator issue you mentioned, it's a memory issue.\nSo, I'm not sure the other parts still consume lots of memory in your case.","created":"2017-01-01T17:25:40.769+0000"},{"body":"thanks. we will test it and confirm.","created":"2017-01-02T04:07:22.329+0000"},{"body":"Thanks. its working fine now for our scenario.","created":"2017-01-02T11:35:11.428+0000"},{"body":"Thank you for testing and confirming!","created":"2017-01-02T23:08:35.086+0000"},{"body":"Issue resolved by pull request 16440\n[https://github.com/apache/spark/pull/16440]","created":"2017-01-10T13:28:17.314+0000"},{"body":"Hi, [~srowen].\nThis is a bug existing 2.0.2 and 2.1.X.\nI'll create a backport for this issue.","created":"2017-01-11T04:55:22.908+0000"},{"body":"Or, could you cherry-pick that please?\nWhen I try to cherry-pick for branch-2.0 and branch-2.1, there was no problem.","created":"2017-01-11T04:58:19.213+0000"},{"body":"I think it's probably OK, if it's a significant problem, and we have a targeted, tested fix here.","created":"2017-01-11T10:28:58.084+0000"}],"conversations":[{"body":"We are trying to run a sql query on our spark cluster and extracting around 200 million records through SparkSQL ThriftServer interface. This query works fine for Spark 1.6.3 version, however for spark 2.0.2, thrift server hangs after fetching data from a few partitions (we are using incremental collect mode with 400 partitions). As per documentation max memory taken up by thrift server should be what is required by the biggest data partition. But we observed that Thrift server is not releasing the old partitions memory whenever the GC occurs even though it has moved to next partition data fetches. which is not the case with 1.6.3 version.\n\nOn further investigation we found that SparkExecuteStatementOperation.scala was modified for \"[SPARK-16563][SQL] fix spark sql thrift server FetchResults bug\" and result set iterator was duplicated to keep a reference to the first set.\n\n+ val (itra, itrb) = iter.duplicate\n+ iterHeader = itra\n+ iter = itrb\n\nWe suspect that this is resulting in the memory not being cleared on GC. To confirm this we created an iterator in our test class and fetched the data once without duplicating and second time with creating a duplicate. we could see that in first instance it ran fine and fetched the entire data set while in second instance driver hanged after fetching data from a few partitions.\n\n\n\n","from":"reporter","subject":"SparkSQL ThriftServer hangs while extracting huge data volumes in incremental collect mode"},{"body":"GC logs for 2 spark versions while running the same query","from":"developer"},{"body":"Thank you for reporting, [~vishalagrwal].\nThen, in the Spark side, could you test on Spark 2.0.0 before SPARK-16563?","from":"developer"},{"body":"We are unable to use incremental collect in a spark version before 2.0.2 due the bug spark-18009\n\nWe will have to take 2.0.2 and change this particular class and build from source code.","from":"developer"},{"body":"we have built Spark from 2.0.2 source code by changing SparkExecuteStatementOperation.scala to pre SPARK-16563 version. this version works fine without causing any thrift server issues.","from":"developer"},{"body":"Thank you for testing and sharing that information!","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16440","from":"developer"},{"body":"Hi, [~vishalagrwal].\nI agree with you. This is an important problem.\nAt least, I made a PR as a first attempt. In any ways, I hope this will be resolved soon.","from":"developer"},{"body":"CC [~alicegugu]","from":"developer"},{"body":"Hi [~vishalagrwal].\nCould you test your case with https://github.com/apache/spark/pull/16440 ?\nAlthough I tried to address the iterator issue you mentioned, it's a memory issue.\nSo, I'm not sure the other parts still consume lots of memory in your case.","from":"developer"},{"body":"thanks. we will test it and confirm.","from":"developer"},{"body":"Thanks. its working fine now for our scenario.","from":"developer"},{"body":"Thank you for testing and confirming!","from":"developer"},{"body":"Issue resolved by pull request 16440\n[https://github.com/apache/spark/pull/16440]","from":"developer"},{"body":"Hi, [~srowen].\nThis is a bug existing 2.0.2 and 2.1.X.\nI'll create a backport for this issue.","from":"developer"},{"body":"Or, could you cherry-pick that please?\nWhen I try to cherry-pick for branch-2.0 and branch-2.1, there was no problem.","from":"developer"},{"body":"I think it's probably OK, if it's a significant problem, and we have a targeted, tested fix here.","from":"developer"}],"created":"2016-12-14T08:37:34.000+0000","description":"We are trying to run a sql query on our spark cluster and extracting around 200 million records through SparkSQL ThriftServer interface. This query works fine for Spark 1.6.3 version, however for spark 2.0.2, thrift server hangs after fetching data from a few partitions (we are using incremental collect mode with 400 partitions). As per documentation max memory taken up by thrift server should be what is required by the biggest data partition. But we observed that Thrift server is not releasing the old partitions memory whenever the GC occurs even though it has moved to next partition data fetches. which is not the case with 1.6.3 version.\n\nOn further investigation we found that SparkExecuteStatementOperation.scala was modified for \"[SPARK-16563][SQL] fix spark sql thrift server FetchResults bug\" and result set iterator was duplicated to keep a reference to the first set.\n\n+ val (itra, itrb) = iter.duplicate\n+ iterHeader = itra\n+ iter = itrb\n\nWe suspect that this is resulting in the memory not being cleared on GC. To confirm this we created an iterator in our test class and fetched the data once without duplicating and second time with creating a duplicate. we could see that in first instance it ran fine and fetched the entire data set while in second instance driver hanged after fetching data from a few partitions.\n\n\n\n","issue_id":"13028031","key":"SPARK-18857","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-01-10T13:28:16.000+0000","role":"fixed_distractor","summary":"SparkSQL ThriftServer hangs while extracting huge data volumes in incremental collect mode"} {"case_id":"13030450","cluster":"DISTRACTOR-SPARK-18993","comments":[{"body":"User 'gatorsmile' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16393","created":"2016-12-24T00:17:03.938+0000"},{"body":"User 'srowen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16418","created":"2016-12-27T21:31:04.202+0000"},{"body":"Issue resolved by pull request 16418\n[https://github.com/apache/spark/pull/16418]","created":"2016-12-28T12:18:06.266+0000"}],"conversations":[{"body":"After https://github.com/apache/spark/pull/16311 is merged, I am unable to build it in my IntelliJ. Got the following compilation error:\n\n{noformat}\nError:scalac: error while loading Object, Missing dependency 'object scala in compiler mirror', required by /Library/Java/JavaVirtualMachines/jdk1.8.0_74.jdk/Contents/Home/jre/lib/rt.jar(java/lang/Object.class)\nError:scalac: Error: object scala in compiler mirror not found.\nscala.reflect.internal.MissingRequirementError: object scala in compiler mirror not found.\n\tat scala.reflect.internal.MissingRequirementError$.signal(MissingRequirementError.scala:17)\n\tat scala.reflect.internal.MissingRequirementError$.notFound(MissingRequirementError.scala:18)\n\tat scala.reflect.internal.Mirrors$RootsBase.getModuleOrClass(Mirrors.scala:53)\n\tat scala.reflect.internal.Mirrors$RootsBase.getModuleOrClass(Mirrors.scala:66)\n\tat scala.reflect.internal.Mirrors$RootsBase.getPackage(Mirrors.scala:173)\n\tat scala.reflect.internal.Definitions$DefinitionsClass.ScalaPackage$lzycompute(Definitions.scala:161)\n\tat scala.reflect.internal.Definitions$DefinitionsClass.ScalaPackage(Definitions.scala:161)\n\tat scala.reflect.internal.Definitions$DefinitionsClass.ScalaPackageClass$lzycompute(Definitions.scala:162)\n\tat scala.reflect.internal.Definitions$DefinitionsClass.ScalaPackageClass(Definitions.scala:162)\n\tat scala.reflect.internal.Definitions$DefinitionsClass.init(Definitions.scala:1395)\n\tat scala.tools.nsc.Global$Run.(Global.scala:1215)\n\tat xsbt.CachedCompiler0$$anon$2.(CompilerInterface.scala:105)\n\tat xsbt.CachedCompiler0.run(CompilerInterface.scala:105)\n\tat xsbt.CachedCompiler0.run(CompilerInterface.scala:94)\n\tat xsbt.CompilerInterface.run(CompilerInterface.scala:22)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat sbt.compiler.AnalyzingCompiler.call(AnalyzingCompiler.scala:101)\n\tat sbt.compiler.AnalyzingCompiler.compile(AnalyzingCompiler.scala:47)\n\tat sbt.compiler.AnalyzingCompiler.compile(AnalyzingCompiler.scala:41)\n\tat org.jetbrains.jps.incremental.scala.local.IdeaIncrementalCompiler.compile(IdeaIncrementalCompiler.scala:29)\n\tat org.jetbrains.jps.incremental.scala.local.LocalServer.compile(LocalServer.scala:26)\n\tat org.jetbrains.jps.incremental.scala.remote.Main$.make(Main.scala:67)\n\tat org.jetbrains.jps.incremental.scala.remote.Main$.nailMain(Main.scala:24)\n\tat org.jetbrains.jps.incremental.scala.remote.Main.nailMain(Main.scala)\n\tat sun.reflect.GeneratedMethodAccessor8.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat com.martiansoftware.nailgun.NGSession.run(NGSession.java:319)\n{noformat}","from":"reporter","subject":"Unable to build/compile Spark in IntelliJ due to missing Scala deps in spark-tags"},{"body":"User 'gatorsmile' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16393","from":"developer"},{"body":"User 'srowen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16418","from":"developer"},{"body":"Issue resolved by pull request 16418\n[https://github.com/apache/spark/pull/16418]","from":"developer"}],"created":"2016-12-24T00:12:10.000+0000","description":"After https://github.com/apache/spark/pull/16311 is merged, I am unable to build it in my IntelliJ. Got the following compilation error:\n\n{noformat}\nError:scalac: error while loading Object, Missing dependency 'object scala in compiler mirror', required by /Library/Java/JavaVirtualMachines/jdk1.8.0_74.jdk/Contents/Home/jre/lib/rt.jar(java/lang/Object.class)\nError:scalac: Error: object scala in compiler mirror not found.\nscala.reflect.internal.MissingRequirementError: object scala in compiler mirror not found.\n\tat scala.reflect.internal.MissingRequirementError$.signal(MissingRequirementError.scala:17)\n\tat scala.reflect.internal.MissingRequirementError$.notFound(MissingRequirementError.scala:18)\n\tat scala.reflect.internal.Mirrors$RootsBase.getModuleOrClass(Mirrors.scala:53)\n\tat scala.reflect.internal.Mirrors$RootsBase.getModuleOrClass(Mirrors.scala:66)\n\tat scala.reflect.internal.Mirrors$RootsBase.getPackage(Mirrors.scala:173)\n\tat scala.reflect.internal.Definitions$DefinitionsClass.ScalaPackage$lzycompute(Definitions.scala:161)\n\tat scala.reflect.internal.Definitions$DefinitionsClass.ScalaPackage(Definitions.scala:161)\n\tat scala.reflect.internal.Definitions$DefinitionsClass.ScalaPackageClass$lzycompute(Definitions.scala:162)\n\tat scala.reflect.internal.Definitions$DefinitionsClass.ScalaPackageClass(Definitions.scala:162)\n\tat scala.reflect.internal.Definitions$DefinitionsClass.init(Definitions.scala:1395)\n\tat scala.tools.nsc.Global$Run.(Global.scala:1215)\n\tat xsbt.CachedCompiler0$$anon$2.(CompilerInterface.scala:105)\n\tat xsbt.CachedCompiler0.run(CompilerInterface.scala:105)\n\tat xsbt.CachedCompiler0.run(CompilerInterface.scala:94)\n\tat xsbt.CompilerInterface.run(CompilerInterface.scala:22)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat sbt.compiler.AnalyzingCompiler.call(AnalyzingCompiler.scala:101)\n\tat sbt.compiler.AnalyzingCompiler.compile(AnalyzingCompiler.scala:47)\n\tat sbt.compiler.AnalyzingCompiler.compile(AnalyzingCompiler.scala:41)\n\tat org.jetbrains.jps.incremental.scala.local.IdeaIncrementalCompiler.compile(IdeaIncrementalCompiler.scala:29)\n\tat org.jetbrains.jps.incremental.scala.local.LocalServer.compile(LocalServer.scala:26)\n\tat org.jetbrains.jps.incremental.scala.remote.Main$.make(Main.scala:67)\n\tat org.jetbrains.jps.incremental.scala.remote.Main$.nailMain(Main.scala:24)\n\tat org.jetbrains.jps.incremental.scala.remote.Main.nailMain(Main.scala)\n\tat sun.reflect.GeneratedMethodAccessor8.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat com.martiansoftware.nailgun.NGSession.run(NGSession.java:319)\n{noformat}","issue_id":"13030450","key":"SPARK-18993","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-12-28T12:18:06.000+0000","role":"fixed_distractor","summary":"Unable to build/compile Spark in IntelliJ due to missing Scala deps in spark-tags"} {"case_id":"13030695","cluster":"DISTRACTOR-SPARK-19012","comments":[{"body":"{{\"11111\"}} fails because it is a numeric literal. Either add a letter to the name, or place the name between backticks, e.g.: {{\"`11111`\"}}\n\nThis is not a bug per se, we could however improve the name parsing.","created":"2016-12-27T13:45:20.440+0000"},{"body":"The type is viewName: String so you would expect any string to work. \n\nWhen executing createOrReplaceTempView(tableOrViewName: String) an ParseException of \n{code}\n== SQL ==\n{tableOrViewName}\n{code}\nis thrown. You might not see that these are related, because it goes on about an SQL ParseException and you just defined a tempView\n\nYou might not expect that the tableOrViewName trigger some SQL parsing. A slighter more clearer exception message (viewName not supported) could be more helpful or add the identifier rules set to the documentation.\n\nAs you said prefixing the tableOrViewName with a non-numerical value does the trick (although for me this feels more like a workaround)","created":"2016-12-27T14:51:39.020+0000"},{"body":"Hi, [~hvanhovell] and [~jzijlstra].\nI'll make a PR to raise `AnalysisException` instead.","created":"2016-12-28T23:20:27.398+0000"},{"body":"[~dongjoon] Could make a PR that puts the name in backticks instead? That is a bit more friendly to the end user. Or do you think we will break stuff, if we do?","created":"2016-12-28T23:22:51.662+0000"},{"body":"Oh, you mean always wrap the name with backticks right?","created":"2016-12-28T23:31:17.721+0000"},{"body":"Yeah, maybe a bit more subtle than that (we need to escape backticks in the name).","created":"2016-12-28T23:32:49.216+0000"},{"body":"No problem. However, we need to raise AnalysisException on empty table table still, ``.","created":"2016-12-28T23:33:06.679+0000"},{"body":"+1","created":"2016-12-28T23:33:27.667+0000"},{"body":"BTW, [~hvanhovell]. I found the existing related issue and testcases.\n{code}\ntest(\"SPARK-12982: Add table name validation in temp table registration\") {\n val df = Seq(\"foo\", \"bar\").map(Tuple1.apply).toDF(\"col\")\n // invalid table name test as below\n intercept[AnalysisException](df.createOrReplaceTempView(\"t~\"))\n // valid table name test as below\n df.createOrReplaceTempView(\"table1\")\n // another invalid table name test as below\n intercept[AnalysisException](df.createOrReplaceTempView(\"#$@sum\"))\n // another invalid table name test as below\n intercept[AnalysisException](df.createOrReplaceTempView(\"table!#\"))\n }\n{code}\n\nTo be consistent with this, we should throw AnalysisException on `createOrReplaceTempView(\"11111\")`.\n\nSo, what we want here is to support `createOrReplaceTempView(\"`11111`\")`. Did I understand clearly?","created":"2016-12-29T00:10:22.038+0000"},{"body":"Ur, actually, we already support `createOrReplaceTempView(\"`11111`\")`.","created":"2016-12-29T00:13:18.115+0000"},{"body":"Yeah, you have a point there. I was wondering if we would hit an issue here.\n\nThe question what we want to support:\n* SQL compatibility. This would be one of the more common use cases. In that case it really does not make sense to support an identifier like '11111', because that would fail in SQL.\n* As much flexibility as you want.\n\n[~jzijlstra] could you explain how you are using this?\n[~dongjoon] lets just make the exception better for now.","created":"2016-12-29T00:21:27.080+0000"},{"body":"Thank you for decision. Yep. I'll make the PR like that.","created":"2016-12-29T00:27:08.896+0000"},{"body":"In API docs and many places, `createOrReplaceTempView` was assumed not to throw any Exceptions.\nIt seems we need to discuss on my PR.","created":"2016-12-29T01:13:58.157+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16427","created":"2016-12-29T01:17:04.606+0000"},{"body":"Good to see that its already being discussed. \n\nMSSQL also has some limitation in tableOrViewNames which is described in the documentation. Maybe updating the annotation of the method would also be enough. Having an Exception with a clear reason would definitely already a fix.\n\n[~hvanhovell]\nWe specify our queries inside a configuration not the code. So we have this in our config:\ndataPath = \"hdfs://\"\ndataQuery: \"SELECT column1, column2 FROM \\[TABLE] WHERE 1 = 1\"\n\nSince we have one SparkSession for the application and the tableOrViewName is coupled to that and we don't want to specify an extra config option for the tableOrViewname, I though I'd just use the hashcode from the dataquery as the tableOrViewName. Use that in the createOrReplaceTempView and replace \\[TABLE] inside the query with that. \n\n{code}\nval path = \"hdfs://....{path}\"\nval dataQuery = \"SELECT * FROM [TABLE] LIMIT 1\"\n\nval tableOrViewName = \"_\" + Math.abs(path.hashCode).toString + Math.abs(qry.hashCode).toString\n\nval df = sparkSession.read.orc(path)\ndf.createOrReplaceTempView(tableOrViewName)\n\nval result = sparkSession.sqlContext.sql(qry.replace(\"[TABLE]\", tableOrViewName)).collect\n{code}\n\nLater I want to check If the tableOrViewName has already been created and not call createOrReplaceTempView everytime, but this is just performance improvement.","created":"2016-12-29T07:55:07.806+0000"},{"body":"Yep. I tried to update the annotation but unfortunately it was reverted that now. (You can see that in my PR.)\n\n> Maybe updating the annotation of the method would also be enough. Having an Exception with a clear reason would definitely already a fix.\n\nChanging annotation on `public` API seems to be handled in a different issue with some more discussion because it affects many other codes (e.g. examples).","created":"2016-12-29T20:05:19.474+0000"},{"body":"Ok, you could also start a table name with {{tbl_}} and that would also make the problem go away.","created":"2016-12-29T20:24:38.310+0000"},{"body":"[~hvanhovell] Its already working for me, I was already prefixing the tableOrViewName. \nI thought you needed an example on how a developers mind work in (mis)using other peoples code.\n\nIts nice to see that its been resolved in just 2 days. \n\n","created":"2016-12-30T08:19:43.542+0000"}],"conversations":[{"body":"Using a viewName where the the fist char is a numerical value on dataframe.createOrReplaceTempView(viewName: String) causes:\n\n{code}\nException in thread \"main\" org.apache.spark.sql.catalyst.parser.ParseException: \nmismatched input '1468079114' expecting {'SELECT', 'FROM', 'ADD', 'AS', 'ALL', 'DISTINCT', 'WHERE', 'GROUP', 'BY', 'GROUPING', 'SETS', 'CUBE', 'ROLLUP', 'ORDER', 'HAVING', 'LIMIT', 'AT', 'OR', 'AND', 'IN', NOT, 'NO', 'EXISTS', 'BETWEEN', 'LIKE', RLIKE, 'IS', 'NULL', 'TRUE', 'FALSE', 'NULLS', 'ASC', 'DESC', 'FOR', 'INTERVAL', 'CASE', 'WHEN', 'THEN', 'ELSE', 'END', 'JOIN', 'CROSS', 'OUTER', 'INNER', 'LEFT', 'SEMI', 'RIGHT', 'FULL', 'NATURAL', 'ON', 'LATERAL', 'WINDOW', 'OVER', 'PARTITION', 'RANGE', 'ROWS', 'UNBOUNDED', 'PRECEDING', 'FOLLOWING', 'CURRENT', 'ROW', 'WITH', 'VALUES', 'CREATE', 'TABLE', 'VIEW', 'REPLACE', 'INSERT', 'DELETE', 'INTO', 'DESCRIBE', 'EXPLAIN', 'FORMAT', 'LOGICAL', 'CODEGEN', 'CAST', 'SHOW', 'TABLES', 'COLUMNS', 'COLUMN', 'USE', 'PARTITIONS', 'FUNCTIONS', 'DROP', 'UNION', 'EXCEPT', 'INTERSECT', 'TO', 'TABLESAMPLE', 'STRATIFY', 'ALTER', 'RENAME', 'ARRAY', 'MAP', 'STRUCT', 'COMMENT', 'SET', 'RESET', 'DATA', 'START', 'TRANSACTION', 'COMMIT', 'ROLLBACK', 'MACRO', 'IF', 'DIV', 'PERCENT', 'BUCKET', 'OUT', 'OF', 'SORT', 'CLUSTER', 'DISTRIBUTE', 'OVERWRITE', 'TRANSFORM', 'REDUCE', 'USING', 'SERDE', 'SERDEPROPERTIES', 'RECORDREADER', 'RECORDWRITER', 'DELIMITED', 'FIELDS', 'TERMINATED', 'COLLECTION', 'ITEMS', 'KEYS', 'ESCAPED', 'LINES', 'SEPARATED', 'FUNCTION', 'EXTENDED', 'REFRESH', 'CLEAR', 'CACHE', 'UNCACHE', 'LAZY', 'FORMATTED', TEMPORARY, 'OPTIONS', 'UNSET', 'TBLPROPERTIES', 'DBPROPERTIES', 'BUCKETS', 'SKEWED', 'STORED', 'DIRECTORIES', 'LOCATION', 'EXCHANGE', 'ARCHIVE', 'UNARCHIVE', 'FILEFORMAT', 'TOUCH', 'COMPACT', 'CONCATENATE', 'CHANGE', 'CASCADE', 'RESTRICT', 'CLUSTERED', 'SORTED', 'PURGE', 'INPUTFORMAT', 'OUTPUTFORMAT', DATABASE, DATABASES, 'DFS', 'TRUNCATE', 'ANALYZE', 'COMPUTE', 'LIST', 'STATISTICS', 'PARTITIONED', 'EXTERNAL', 'DEFINED', 'REVOKE', 'GRANT', 'LOCK', 'UNLOCK', 'MSCK', 'REPAIR', 'RECOVER', 'EXPORT', 'IMPORT', 'LOAD', 'ROLE', 'ROLES', 'COMPACTIONS', 'PRINCIPALS', 'TRANSACTIONS', 'INDEX', 'INDEXES', 'LOCKS', 'OPTION', 'ANTI', 'LOCAL', 'INPATH', 'CURRENT_DATE', 'CURRENT_TIMESTAMP', IDENTIFIER, BACKQUOTED_IDENTIFIER}(line 1, pos 0)\n\n== SQL ==\n11111\n{code}\n\n{code}\nval tableOrViewName = \"11111\" //fails\nval tableOrViewName = \"a1111\" //works\nsparkSession.read.orc(path).createOrReplaceTempView(tableOrViewName)\n{code}\n\n","from":"reporter","subject":"CreateOrReplaceTempView throws org.apache.spark.sql.catalyst.parser.ParseException when viewName first char is numerical"},{"body":"{{\"11111\"}} fails because it is a numeric literal. Either add a letter to the name, or place the name between backticks, e.g.: {{\"`11111`\"}}\n\nThis is not a bug per se, we could however improve the name parsing.","from":"developer"},{"body":"The type is viewName: String so you would expect any string to work. \n\nWhen executing createOrReplaceTempView(tableOrViewName: String) an ParseException of \n{code}\n== SQL ==\n{tableOrViewName}\n{code}\nis thrown. You might not see that these are related, because it goes on about an SQL ParseException and you just defined a tempView\n\nYou might not expect that the tableOrViewName trigger some SQL parsing. A slighter more clearer exception message (viewName not supported) could be more helpful or add the identifier rules set to the documentation.\n\nAs you said prefixing the tableOrViewName with a non-numerical value does the trick (although for me this feels more like a workaround)","from":"developer"},{"body":"Hi, [~hvanhovell] and [~jzijlstra].\nI'll make a PR to raise `AnalysisException` instead.","from":"developer"},{"body":"[~dongjoon] Could make a PR that puts the name in backticks instead? That is a bit more friendly to the end user. Or do you think we will break stuff, if we do?","from":"developer"},{"body":"Oh, you mean always wrap the name with backticks right?","from":"developer"},{"body":"Yeah, maybe a bit more subtle than that (we need to escape backticks in the name).","from":"developer"},{"body":"No problem. However, we need to raise AnalysisException on empty table table still, ``.","from":"developer"},{"body":"+1","from":"developer"},{"body":"BTW, [~hvanhovell]. I found the existing related issue and testcases.\n{code}\ntest(\"SPARK-12982: Add table name validation in temp table registration\") {\n val df = Seq(\"foo\", \"bar\").map(Tuple1.apply).toDF(\"col\")\n // invalid table name test as below\n intercept[AnalysisException](df.createOrReplaceTempView(\"t~\"))\n // valid table name test as below\n df.createOrReplaceTempView(\"table1\")\n // another invalid table name test as below\n intercept[AnalysisException](df.createOrReplaceTempView(\"#$@sum\"))\n // another invalid table name test as below\n intercept[AnalysisException](df.createOrReplaceTempView(\"table!#\"))\n }\n{code}\n\nTo be consistent with this, we should throw AnalysisException on `createOrReplaceTempView(\"11111\")`.\n\nSo, what we want here is to support `createOrReplaceTempView(\"`11111`\")`. Did I understand clearly?","from":"developer"},{"body":"Ur, actually, we already support `createOrReplaceTempView(\"`11111`\")`.","from":"developer"},{"body":"Yeah, you have a point there. I was wondering if we would hit an issue here.\n\nThe question what we want to support:\n* SQL compatibility. This would be one of the more common use cases. In that case it really does not make sense to support an identifier like '11111', because that would fail in SQL.\n* As much flexibility as you want.\n\n[~jzijlstra] could you explain how you are using this?\n[~dongjoon] lets just make the exception better for now.","from":"developer"},{"body":"Thank you for decision. Yep. I'll make the PR like that.","from":"developer"},{"body":"In API docs and many places, `createOrReplaceTempView` was assumed not to throw any Exceptions.\nIt seems we need to discuss on my PR.","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16427","from":"developer"},{"body":"Good to see that its already being discussed. \n\nMSSQL also has some limitation in tableOrViewNames which is described in the documentation. Maybe updating the annotation of the method would also be enough. Having an Exception with a clear reason would definitely already a fix.\n\n[~hvanhovell]\nWe specify our queries inside a configuration not the code. So we have this in our config:\ndataPath = \"hdfs://\"\ndataQuery: \"SELECT column1, column2 FROM \\[TABLE] WHERE 1 = 1\"\n\nSince we have one SparkSession for the application and the tableOrViewName is coupled to that and we don't want to specify an extra config option for the tableOrViewname, I though I'd just use the hashcode from the dataquery as the tableOrViewName. Use that in the createOrReplaceTempView and replace \\[TABLE] inside the query with that. \n\n{code}\nval path = \"hdfs://....{path}\"\nval dataQuery = \"SELECT * FROM [TABLE] LIMIT 1\"\n\nval tableOrViewName = \"_\" + Math.abs(path.hashCode).toString + Math.abs(qry.hashCode).toString\n\nval df = sparkSession.read.orc(path)\ndf.createOrReplaceTempView(tableOrViewName)\n\nval result = sparkSession.sqlContext.sql(qry.replace(\"[TABLE]\", tableOrViewName)).collect\n{code}\n\nLater I want to check If the tableOrViewName has already been created and not call createOrReplaceTempView everytime, but this is just performance improvement.","from":"developer"},{"body":"Yep. I tried to update the annotation but unfortunately it was reverted that now. (You can see that in my PR.)\n\n> Maybe updating the annotation of the method would also be enough. Having an Exception with a clear reason would definitely already a fix.\n\nChanging annotation on `public` API seems to be handled in a different issue with some more discussion because it affects many other codes (e.g. examples).","from":"developer"},{"body":"Ok, you could also start a table name with {{tbl_}} and that would also make the problem go away.","from":"developer"},{"body":"[~hvanhovell] Its already working for me, I was already prefixing the tableOrViewName. \nI thought you needed an example on how a developers mind work in (mis)using other peoples code.\n\nIts nice to see that its been resolved in just 2 days. \n\n","from":"developer"}],"created":"2016-12-27T12:58:48.000+0000","description":"Using a viewName where the the fist char is a numerical value on dataframe.createOrReplaceTempView(viewName: String) causes:\n\n{code}\nException in thread \"main\" org.apache.spark.sql.catalyst.parser.ParseException: \nmismatched input '1468079114' expecting {'SELECT', 'FROM', 'ADD', 'AS', 'ALL', 'DISTINCT', 'WHERE', 'GROUP', 'BY', 'GROUPING', 'SETS', 'CUBE', 'ROLLUP', 'ORDER', 'HAVING', 'LIMIT', 'AT', 'OR', 'AND', 'IN', NOT, 'NO', 'EXISTS', 'BETWEEN', 'LIKE', RLIKE, 'IS', 'NULL', 'TRUE', 'FALSE', 'NULLS', 'ASC', 'DESC', 'FOR', 'INTERVAL', 'CASE', 'WHEN', 'THEN', 'ELSE', 'END', 'JOIN', 'CROSS', 'OUTER', 'INNER', 'LEFT', 'SEMI', 'RIGHT', 'FULL', 'NATURAL', 'ON', 'LATERAL', 'WINDOW', 'OVER', 'PARTITION', 'RANGE', 'ROWS', 'UNBOUNDED', 'PRECEDING', 'FOLLOWING', 'CURRENT', 'ROW', 'WITH', 'VALUES', 'CREATE', 'TABLE', 'VIEW', 'REPLACE', 'INSERT', 'DELETE', 'INTO', 'DESCRIBE', 'EXPLAIN', 'FORMAT', 'LOGICAL', 'CODEGEN', 'CAST', 'SHOW', 'TABLES', 'COLUMNS', 'COLUMN', 'USE', 'PARTITIONS', 'FUNCTIONS', 'DROP', 'UNION', 'EXCEPT', 'INTERSECT', 'TO', 'TABLESAMPLE', 'STRATIFY', 'ALTER', 'RENAME', 'ARRAY', 'MAP', 'STRUCT', 'COMMENT', 'SET', 'RESET', 'DATA', 'START', 'TRANSACTION', 'COMMIT', 'ROLLBACK', 'MACRO', 'IF', 'DIV', 'PERCENT', 'BUCKET', 'OUT', 'OF', 'SORT', 'CLUSTER', 'DISTRIBUTE', 'OVERWRITE', 'TRANSFORM', 'REDUCE', 'USING', 'SERDE', 'SERDEPROPERTIES', 'RECORDREADER', 'RECORDWRITER', 'DELIMITED', 'FIELDS', 'TERMINATED', 'COLLECTION', 'ITEMS', 'KEYS', 'ESCAPED', 'LINES', 'SEPARATED', 'FUNCTION', 'EXTENDED', 'REFRESH', 'CLEAR', 'CACHE', 'UNCACHE', 'LAZY', 'FORMATTED', TEMPORARY, 'OPTIONS', 'UNSET', 'TBLPROPERTIES', 'DBPROPERTIES', 'BUCKETS', 'SKEWED', 'STORED', 'DIRECTORIES', 'LOCATION', 'EXCHANGE', 'ARCHIVE', 'UNARCHIVE', 'FILEFORMAT', 'TOUCH', 'COMPACT', 'CONCATENATE', 'CHANGE', 'CASCADE', 'RESTRICT', 'CLUSTERED', 'SORTED', 'PURGE', 'INPUTFORMAT', 'OUTPUTFORMAT', DATABASE, DATABASES, 'DFS', 'TRUNCATE', 'ANALYZE', 'COMPUTE', 'LIST', 'STATISTICS', 'PARTITIONED', 'EXTERNAL', 'DEFINED', 'REVOKE', 'GRANT', 'LOCK', 'UNLOCK', 'MSCK', 'REPAIR', 'RECOVER', 'EXPORT', 'IMPORT', 'LOAD', 'ROLE', 'ROLES', 'COMPACTIONS', 'PRINCIPALS', 'TRANSACTIONS', 'INDEX', 'INDEXES', 'LOCKS', 'OPTION', 'ANTI', 'LOCAL', 'INPATH', 'CURRENT_DATE', 'CURRENT_TIMESTAMP', IDENTIFIER, BACKQUOTED_IDENTIFIER}(line 1, pos 0)\n\n== SQL ==\n11111\n{code}\n\n{code}\nval tableOrViewName = \"11111\" //fails\nval tableOrViewName = \"a1111\" //works\nsparkSession.read.orc(path).createOrReplaceTempView(tableOrViewName)\n{code}\n\n","issue_id":"13030695","key":"SPARK-19012","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-12-29T20:22:48.000+0000","role":"fixed_distractor","summary":"CreateOrReplaceTempView throws org.apache.spark.sql.catalyst.parser.ParseException when viewName first char is numerical"} {"case_id":"13030996","cluster":"DISTRACTOR-SPARK-19019","comments":[{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16429","created":"2016-12-29T02:47:04.201+0000"},{"body":"Issue resolved by pull request 16429\n[https://github.com/apache/spark/pull/16429]","created":"2017-01-17T17:53:30.331+0000"},{"body":"[~davies] Could it be backported to 1.6 and 2.0?","created":"2017-03-18T17:17:08.793+0000"},{"body":"Would also be interested in the answer to Maciej's question (for 2.0) and when is 2.1.1 scheduled to be released? Thank you!","created":"2017-03-19T02:00:40.888+0000"},{"body":"Let me try to make a PR to backport this if this is confirmed.","created":"2017-03-19T04:41:20.363+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17374","created":"2017-03-21T09:29:03.816+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17375","created":"2017-03-21T09:54:02.791+0000"},{"body":"To solve this problem fully, I had to port cloudpickle change too in the PR. Only fixing hijected one described above dose not fully solve this issue. Please refer the discussion in the PR and the change.","created":"2017-05-02T10:49:57.963+0000"},{"body":"Just got this error post fix on spark 2.1:\n\n\n{code:java}\nTraceback (most recent call last):\n File \"/opt/anaconda3/lib/python3.6/runpy.py\", line 183, in _run_module_as_main\n mod_name, mod_spec, code = _get_module_details(mod_name, _Error)\n File \"/opt/anaconda3/lib/python3.6/runpy.py\", line 109, in _get_module_details\n __import__(pkg_name)\n File \"/usr/hdp/current/spark-client/python/pyspark/__init__.py\", line 41, in \n from pyspark.context import SparkContext\n File \"/usr/hdp/current/spark-client/python/pyspark/context.py\", line 33, in \n from pyspark.java_gateway import launch_gateway\n File \"/usr/hdp/current/spark-client/python/pyspark/java_gateway.py\", line 25, in \n import platform\n File \"/opt/anaconda3/lib/python3.6/platform.py\", line 886, in \n \"system node release version machine processor\")\n File \"/usr/hdp/current/spark-client/python/pyspark/serializers.py\", line 381, in namedtuple\n cls = _old_namedtuple(*args, **kwargs)\n TypeError: namedtuple() missing 3 required keyword-only arguments: 'verbose', 'rename', and 'module'\n{code}\n","created":"2017-08-15T14:50:57.518+0000"},{"body":"I think this was backported into Spark 2.1.1. Was your Spark version, 2.1.1+?","created":"2017-08-15T14:55:45.570+0000"},{"body":"Yea. This was just a pythonpath mishab on our end. 2.1.1 is a-okay.","created":"2017-08-16T10:13:22.935+0000"},{"body":"I am facing the same issue in Spark1.6.0? Was this fixed for Spark 1.6.0 version? If not, are there any plans to do so?","created":"2019-02-25T06:41:06.547+0000"},{"body":"According to fix versions, fixed version of 1.6.x version line is 1.6.4. You need to upgrade to 1.6.4, but I believe 1.6 is EOL and no more support on community. You may need to upgrade the version to 2.3.3 (if you feel more safer to have bugfix versions in minor version) or 2.4.0.","created":"2019-02-25T07:26:25.667+0000"}],"conversations":[{"body":"Currently, PySpark does not work with Python 3.6.0.\n\nRunning {{./bin/pyspark}} simply throws the error as below:\n\n{code}\nTraceback (most recent call last):\n File \".../spark/python/pyspark/shell.py\", line 30, in \n import pyspark\n File \".../spark/python/pyspark/__init__.py\", line 46, in \n from pyspark.context import SparkContext\n File \".../spark/python/pyspark/context.py\", line 36, in \n from pyspark.java_gateway import launch_gateway\n File \".../spark/python/pyspark/java_gateway.py\", line 31, in \n from py4j.java_gateway import java_import, JavaGateway, GatewayClient\n File \"\", line 961, in _find_and_load\n File \"\", line 950, in _find_and_load_unlocked\n File \"\", line 646, in _load_unlocked\n File \"\", line 616, in _load_backward_compatible\n File \".../spark/python/lib/py4j-0.10.4-src.zip/py4j/java_gateway.py\", line 18, in \n File \"/usr/local/Cellar/python3/3.6.0/Frameworks/Python.framework/Versions/3.6/lib/python3.6/pydoc.py\", line 62, in \n import pkgutil\n File \"/usr/local/Cellar/python3/3.6.0/Frameworks/Python.framework/Versions/3.6/lib/python3.6/pkgutil.py\", line 22, in \n ModuleInfo = namedtuple('ModuleInfo', 'module_finder name ispkg')\n File \".../spark/python/pyspark/serializers.py\", line 394, in namedtuple\n cls = _old_namedtuple(*args, **kwargs)\nTypeError: namedtuple() missing 3 required keyword-only arguments: 'verbose', 'rename', and 'module'\n{code}\n\nThe problem is in https://github.com/apache/spark/blob/3c68944b229aaaeeaee3efcbae3e3be9a2914855/python/pyspark/serializers.py#L386-L394 as the error says and the cause seems because the arguments of {{namedtuple}} are now completely keyword-only arguments from Python 3.6.0 (See https://bugs.python.org/issue25628).\n\nWe currently copy this function via {{types.FunctionType}} which does not set the default values of keyword-only arguments (meaning {{namedtuple.__kwdefaults__}}) and this seems causing internally missing values in the function (non-bound arguments).\n\n\nThis ends up as below:\n\n{code}\nimport types\nimport collections\n\ndef _copy_func(f):\n return types.FunctionType(f.__code__, f.__globals__, f.__name__,\n f.__defaults__, f.__closure__)\n\n_old_namedtuple = _copy_func(collections.namedtuple)\n\n_old_namedtuple(, \"b\")\n_old_namedtuple(\"a\")\n{code}\n\n\nIf we call as below:\n\n{code}\n>>> _old_namedtuple(\"a\", \"b\")\nTraceback (most recent call last):\n File \"\", line 1, in \nTypeError: namedtuple() missing 3 required keyword-only arguments: 'verbose', 'rename', and 'module'\n{code}\n\n\nIt throws an exception as above becuase {{__kwdefaults__}} for required keyword arguments seem unset in the copied function. So, if we give explicit value for these,\n\n{code}\n>>> _old_namedtuple(\"a\", \"b\", verbose=False, rename=False, module=None)\n\n{code}\n\nIt works fine.\n\nIt seems now we should properly set these into the hijected one.","from":"reporter","subject":"PySpark does not work with Python 3.6.0"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16429","from":"developer"},{"body":"Issue resolved by pull request 16429\n[https://github.com/apache/spark/pull/16429]","from":"developer"},{"body":"[~davies] Could it be backported to 1.6 and 2.0?","from":"developer"},{"body":"Would also be interested in the answer to Maciej's question (for 2.0) and when is 2.1.1 scheduled to be released? Thank you!","from":"developer"},{"body":"Let me try to make a PR to backport this if this is confirmed.","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17374","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17375","from":"developer"},{"body":"To solve this problem fully, I had to port cloudpickle change too in the PR. Only fixing hijected one described above dose not fully solve this issue. Please refer the discussion in the PR and the change.","from":"developer"},{"body":"Just got this error post fix on spark 2.1:\n\n\n{code:java}\nTraceback (most recent call last):\n File \"/opt/anaconda3/lib/python3.6/runpy.py\", line 183, in _run_module_as_main\n mod_name, mod_spec, code = _get_module_details(mod_name, _Error)\n File \"/opt/anaconda3/lib/python3.6/runpy.py\", line 109, in _get_module_details\n __import__(pkg_name)\n File \"/usr/hdp/current/spark-client/python/pyspark/__init__.py\", line 41, in \n from pyspark.context import SparkContext\n File \"/usr/hdp/current/spark-client/python/pyspark/context.py\", line 33, in \n from pyspark.java_gateway import launch_gateway\n File \"/usr/hdp/current/spark-client/python/pyspark/java_gateway.py\", line 25, in \n import platform\n File \"/opt/anaconda3/lib/python3.6/platform.py\", line 886, in \n \"system node release version machine processor\")\n File \"/usr/hdp/current/spark-client/python/pyspark/serializers.py\", line 381, in namedtuple\n cls = _old_namedtuple(*args, **kwargs)\n TypeError: namedtuple() missing 3 required keyword-only arguments: 'verbose', 'rename', and 'module'\n{code}\n","from":"developer"},{"body":"I think this was backported into Spark 2.1.1. Was your Spark version, 2.1.1+?","from":"developer"},{"body":"Yea. This was just a pythonpath mishab on our end. 2.1.1 is a-okay.","from":"developer"},{"body":"I am facing the same issue in Spark1.6.0? Was this fixed for Spark 1.6.0 version? If not, are there any plans to do so?","from":"developer"},{"body":"According to fix versions, fixed version of 1.6.x version line is 1.6.4. You need to upgrade to 1.6.4, but I believe 1.6 is EOL and no more support on community. You may need to upgrade the version to 2.3.3 (if you feel more safer to have bugfix versions in minor version) or 2.4.0.","from":"developer"}],"created":"2016-12-29T02:40:53.000+0000","description":"Currently, PySpark does not work with Python 3.6.0.\n\nRunning {{./bin/pyspark}} simply throws the error as below:\n\n{code}\nTraceback (most recent call last):\n File \".../spark/python/pyspark/shell.py\", line 30, in \n import pyspark\n File \".../spark/python/pyspark/__init__.py\", line 46, in \n from pyspark.context import SparkContext\n File \".../spark/python/pyspark/context.py\", line 36, in \n from pyspark.java_gateway import launch_gateway\n File \".../spark/python/pyspark/java_gateway.py\", line 31, in \n from py4j.java_gateway import java_import, JavaGateway, GatewayClient\n File \"\", line 961, in _find_and_load\n File \"\", line 950, in _find_and_load_unlocked\n File \"\", line 646, in _load_unlocked\n File \"\", line 616, in _load_backward_compatible\n File \".../spark/python/lib/py4j-0.10.4-src.zip/py4j/java_gateway.py\", line 18, in \n File \"/usr/local/Cellar/python3/3.6.0/Frameworks/Python.framework/Versions/3.6/lib/python3.6/pydoc.py\", line 62, in \n import pkgutil\n File \"/usr/local/Cellar/python3/3.6.0/Frameworks/Python.framework/Versions/3.6/lib/python3.6/pkgutil.py\", line 22, in \n ModuleInfo = namedtuple('ModuleInfo', 'module_finder name ispkg')\n File \".../spark/python/pyspark/serializers.py\", line 394, in namedtuple\n cls = _old_namedtuple(*args, **kwargs)\nTypeError: namedtuple() missing 3 required keyword-only arguments: 'verbose', 'rename', and 'module'\n{code}\n\nThe problem is in https://github.com/apache/spark/blob/3c68944b229aaaeeaee3efcbae3e3be9a2914855/python/pyspark/serializers.py#L386-L394 as the error says and the cause seems because the arguments of {{namedtuple}} are now completely keyword-only arguments from Python 3.6.0 (See https://bugs.python.org/issue25628).\n\nWe currently copy this function via {{types.FunctionType}} which does not set the default values of keyword-only arguments (meaning {{namedtuple.__kwdefaults__}}) and this seems causing internally missing values in the function (non-bound arguments).\n\n\nThis ends up as below:\n\n{code}\nimport types\nimport collections\n\ndef _copy_func(f):\n return types.FunctionType(f.__code__, f.__globals__, f.__name__,\n f.__defaults__, f.__closure__)\n\n_old_namedtuple = _copy_func(collections.namedtuple)\n\n_old_namedtuple(, \"b\")\n_old_namedtuple(\"a\")\n{code}\n\n\nIf we call as below:\n\n{code}\n>>> _old_namedtuple(\"a\", \"b\")\nTraceback (most recent call last):\n File \"\", line 1, in \nTypeError: namedtuple() missing 3 required keyword-only arguments: 'verbose', 'rename', and 'module'\n{code}\n\n\nIt throws an exception as above becuase {{__kwdefaults__}} for required keyword arguments seem unset in the copied function. So, if we give explicit value for these,\n\n{code}\n>>> _old_namedtuple(\"a\", \"b\", verbose=False, rename=False, module=None)\n\n{code}\n\nIt works fine.\n\nIt seems now we should properly set these into the hijected one.","issue_id":"13030996","key":"SPARK-19019","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-01-17T17:53:30.000+0000","role":"fixed_distractor","summary":"PySpark does not work with Python 3.6.0"} {"case_id":"13031350","cluster":"DISTRACTOR-SPARK-19038","comments":[{"body":"Also, since the keytab file name in the staging directory does not match the new Spark config setting, the Kerberos ticket is not properly renewed before expiration.","created":"2017-01-02T01:52:03.130+0000"},{"body":"User 'parente' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16482","created":"2017-01-06T05:28:04.930+0000"},{"body":"Please see the comment I made in Github(https://github.com/apache/spark/pull/16482), from my understanding the behavior is expected.","created":"2017-01-06T06:25:59.564+0000"},{"body":"[~jerryshao] PR 16482 is for a different issue SPARK-19105? \n\nI still see this issue after Spark 2.0 upgrade. \nSee details in https://github.com/apache/spark/pull/16482#issuecomment-279563855\nWe didn't have this problem in Spark 1.5 nor 1.6.\n\nTried to workaround by putting keytab in user's home directory in HDFS but this fails with (expected?)\n{noformat}\nException in thread \"main\" org.apache.spark.SparkException: Keytab file: hdfs:///user/svc_odiprd/.kt does not exist\n at org.apache.spark.deploy.SparkSubmit$.prepareSubmitEnvironment(SparkSubmit.scala:555)\n at org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:158)\n at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:124)\n at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\n{noformat}\ncreated SPARK-19588 to create this separate issue for the workaround.","created":"2017-02-14T03:34:25.132+0000"},{"body":"I think the issue you met is the same as this JIRA mentioned, but PR 16482 tries to use a different way to solve this problem, which is not correct for current Spark on YARN. Let me figure out a decent solution to fix this.","created":"2017-02-14T03:44:10.995+0000"},{"body":"Thank you [~jerryshao]","created":"2017-02-14T03:54:34.800+0000"},{"body":"Another possible workaround is to pass principal and keytab to spark-submit through SPARK_SUBMIT_OPTIONS environment variable\n{code}\nenviron[\"SPARK_SUBMIT_OPTIONS\"] = \"--principal %s --keytab %s\" % (kt_principal, kt_location)\n{code}\ninstead of setting in \n{code}\nSparkConf().set(\"spark.yarn.keytab\", kt_location).set(\"spark.yarn.principal\", kt_principal)\n{code}\n\nbut again this is a breaking change / bug in Spark 2.","created":"2017-02-14T04:06:16.198+0000"},{"body":"User 'jerryshao' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16923","created":"2017-02-14T09:13:03.887+0000"}],"conversations":[{"body":"h2. Stack Trace\n\n{noformat}\nPy4JJavaErrorTraceback (most recent call last)\n in ()\n----> 1 sdf = sql.createDataFrame(df)\n\n/opt/spark2/python/pyspark/sql/context.py in createDataFrame(self, data, schema, samplingRatio, verifySchema)\n 307 Py4JJavaError: ...\n 308 \"\"\"\n--> 309 return self.sparkSession.createDataFrame(data, schema, samplingRatio, verifySchema)\n 310 \n 311 @since(1.3)\n\n/opt/spark2/python/pyspark/sql/session.py in createDataFrame(self, data, schema, samplingRatio, verifySchema)\n 524 rdd, schema = self._createFromLocal(map(prepare, data), schema)\n 525 jrdd = self._jvm.SerDeUtil.toJavaArray(rdd._to_java_object_rdd())\n--> 526 jdf = self._jsparkSession.applySchemaToPythonRDD(jrdd.rdd(), schema.json())\n 527 df = DataFrame(jdf, self._wrapped)\n 528 df._schema = schema\n\n/opt/spark2/python/lib/py4j-0.10.3-src.zip/py4j/java_gateway.py in __call__(self, *args)\n 1131 answer = self.gateway_client.send_command(command)\n 1132 return_value = get_return_value(\n-> 1133 answer, self.gateway_client, self.target_id, self.name)\n 1134 \n 1135 for temp_arg in temp_args:\n\n/opt/spark2/python/pyspark/sql/utils.py in deco(*a, **kw)\n 61 def deco(*a, **kw):\n 62 try:\n---> 63 return f(*a, **kw)\n 64 except py4j.protocol.Py4JJavaError as e:\n 65 s = e.java_exception.toString()\n\n/opt/spark2/python/lib/py4j-0.10.3-src.zip/py4j/protocol.py in get_return_value(answer, gateway_client, target_id, name)\n 317 raise Py4JJavaError(\n 318 \"An error occurred while calling {0}{1}{2}.\\n\".\n--> 319 format(target_id, \".\", name), value)\n 320 else:\n 321 raise Py4JError(\n\nPy4JJavaError: An error occurred while calling o47.applySchemaToPythonRDD.\n: org.apache.spark.SparkException: Keytab file: .keytab-f0b9b814-460e-4fa8-8e7d-029186b696c4 specified in spark.yarn.keytab does not exist\n\tat org.apache.spark.sql.hive.client.HiveClientImpl.(HiveClientImpl.scala:113)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:423)\n\tat org.apache.spark.sql.hive.client.IsolatedClientLoader.createClient(IsolatedClientLoader.scala:258)\n\tat org.apache.spark.sql.hive.HiveUtils$.newClientForMetadata(HiveUtils.scala:359)\n\tat org.apache.spark.sql.hive.HiveUtils$.newClientForMetadata(HiveUtils.scala:263)\n\tat org.apache.spark.sql.hive.HiveSharedState.metadataHive$lzycompute(HiveSharedState.scala:39)\n\tat org.apache.spark.sql.hive.HiveSharedState.metadataHive(HiveSharedState.scala:38)\n\tat org.apache.spark.sql.hive.HiveSharedState.externalCatalog$lzycompute(HiveSharedState.scala:46)\n\tat org.apache.spark.sql.hive.HiveSharedState.externalCatalog(HiveSharedState.scala:45)\n\tat org.apache.spark.sql.hive.HiveSessionState.catalog$lzycompute(HiveSessionState.scala:50)\n\tat org.apache.spark.sql.hive.HiveSessionState.catalog(HiveSessionState.scala:48)\n\tat org.apache.spark.sql.hive.HiveSessionState$$anon$1.(HiveSessionState.scala:63)\n\tat org.apache.spark.sql.hive.HiveSessionState.analyzer$lzycompute(HiveSessionState.scala:63)\n\tat org.apache.spark.sql.hive.HiveSessionState.analyzer(HiveSessionState.scala:62)\n\tat org.apache.spark.sql.execution.QueryExecution.assertAnalyzed(QueryExecution.scala:49)\n\tat org.apache.spark.sql.Dataset$.ofRows(Dataset.scala:64)\n\tat org.apache.spark.sql.SparkSession.applySchemaToPythonRDD(SparkSession.scala:666)\n\tat org.apache.spark.sql.SparkSession.applySchemaToPythonRDD(SparkSession.scala:656)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:237)\n\tat py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:357)\n\tat py4j.Gateway.invoke(Gateway.java:280)\n\tat py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:132)\n\tat py4j.commands.CallCommand.execute(CallCommand.java:79)\n\tat py4j.GatewayConnection.run(GatewayConnection.java:214)\n\tat java.lang.Thread.run(Thread.java:745)\n{noformat}\n\nh2. Steps to reproduce\n\n1. Pass valid --principal=user@REALM and --keytab=/home/user/.keytab to spark-submit\n2. Set spark.sql.catalogImplementation = 'hive'\n3. Set deploy mode to yarn-client\n4. Create a SparkSession and try to use session.createDataFrame()\n\nh2. Observations\n\n* The {{setupCredentials}} function in Client.scala sets {{spark.yarn.keytab}} to a UUID suffixed version of the base keytab filename without any path. For example, {{sparkContext.getConf().getAll()}} shows {{spark.yarn.keytab}} as having value {{.keytab-f0b9b814-460e-4fa8-8e7d-029186b696c4}}\n* When listing the contents of the application staging directory on HDFS, no suffixed file exists. Rather, the keytab file appears in the listing with its original name. For instance, {{hdfs dfs -ls hdfs://home/user/.sparkStaging/appication_big_uuid/}} shows an entry {{hdfs://home/user/.sparkStaging/appication_big_uuid/.keytab}}, but not {{hdfs://home/user/.sparkStaging/appication_big_uuid/.keytab-big-uuid}}.\n* The same exception noted above occurs even after I manually put a copy of the keytab with a filename matching the new value of {{spark.yarn.keytab}} onto HDFS in the staging directory.\n\nh2. Expected Behavior\n\nHiveClientImpl should be able to read {{spark.yarn.keytab}} to find the keytab file and initialize itself properly.\n\nh2. References\n\n* SPARK-8619 also noted trouble with the keytab property getting changed after app startup.","from":"reporter","subject":"Can't find keytab file when using Hive catalog"},{"body":"Also, since the keytab file name in the staging directory does not match the new Spark config setting, the Kerberos ticket is not properly renewed before expiration.","from":"developer"},{"body":"User 'parente' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16482","from":"developer"},{"body":"Please see the comment I made in Github(https://github.com/apache/spark/pull/16482), from my understanding the behavior is expected.","from":"developer"},{"body":"[~jerryshao] PR 16482 is for a different issue SPARK-19105? \n\nI still see this issue after Spark 2.0 upgrade. \nSee details in https://github.com/apache/spark/pull/16482#issuecomment-279563855\nWe didn't have this problem in Spark 1.5 nor 1.6.\n\nTried to workaround by putting keytab in user's home directory in HDFS but this fails with (expected?)\n{noformat}\nException in thread \"main\" org.apache.spark.SparkException: Keytab file: hdfs:///user/svc_odiprd/.kt does not exist\n at org.apache.spark.deploy.SparkSubmit$.prepareSubmitEnvironment(SparkSubmit.scala:555)\n at org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:158)\n at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:124)\n at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\n{noformat}\ncreated SPARK-19588 to create this separate issue for the workaround.","from":"developer"},{"body":"I think the issue you met is the same as this JIRA mentioned, but PR 16482 tries to use a different way to solve this problem, which is not correct for current Spark on YARN. Let me figure out a decent solution to fix this.","from":"developer"},{"body":"Thank you [~jerryshao]","from":"developer"},{"body":"Another possible workaround is to pass principal and keytab to spark-submit through SPARK_SUBMIT_OPTIONS environment variable\n{code}\nenviron[\"SPARK_SUBMIT_OPTIONS\"] = \"--principal %s --keytab %s\" % (kt_principal, kt_location)\n{code}\ninstead of setting in \n{code}\nSparkConf().set(\"spark.yarn.keytab\", kt_location).set(\"spark.yarn.principal\", kt_principal)\n{code}\n\nbut again this is a breaking change / bug in Spark 2.","from":"developer"},{"body":"User 'jerryshao' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16923","from":"developer"}],"created":"2016-12-30T23:22:02.000+0000","description":"h2. Stack Trace\n\n{noformat}\nPy4JJavaErrorTraceback (most recent call last)\n in ()\n----> 1 sdf = sql.createDataFrame(df)\n\n/opt/spark2/python/pyspark/sql/context.py in createDataFrame(self, data, schema, samplingRatio, verifySchema)\n 307 Py4JJavaError: ...\n 308 \"\"\"\n--> 309 return self.sparkSession.createDataFrame(data, schema, samplingRatio, verifySchema)\n 310 \n 311 @since(1.3)\n\n/opt/spark2/python/pyspark/sql/session.py in createDataFrame(self, data, schema, samplingRatio, verifySchema)\n 524 rdd, schema = self._createFromLocal(map(prepare, data), schema)\n 525 jrdd = self._jvm.SerDeUtil.toJavaArray(rdd._to_java_object_rdd())\n--> 526 jdf = self._jsparkSession.applySchemaToPythonRDD(jrdd.rdd(), schema.json())\n 527 df = DataFrame(jdf, self._wrapped)\n 528 df._schema = schema\n\n/opt/spark2/python/lib/py4j-0.10.3-src.zip/py4j/java_gateway.py in __call__(self, *args)\n 1131 answer = self.gateway_client.send_command(command)\n 1132 return_value = get_return_value(\n-> 1133 answer, self.gateway_client, self.target_id, self.name)\n 1134 \n 1135 for temp_arg in temp_args:\n\n/opt/spark2/python/pyspark/sql/utils.py in deco(*a, **kw)\n 61 def deco(*a, **kw):\n 62 try:\n---> 63 return f(*a, **kw)\n 64 except py4j.protocol.Py4JJavaError as e:\n 65 s = e.java_exception.toString()\n\n/opt/spark2/python/lib/py4j-0.10.3-src.zip/py4j/protocol.py in get_return_value(answer, gateway_client, target_id, name)\n 317 raise Py4JJavaError(\n 318 \"An error occurred while calling {0}{1}{2}.\\n\".\n--> 319 format(target_id, \".\", name), value)\n 320 else:\n 321 raise Py4JError(\n\nPy4JJavaError: An error occurred while calling o47.applySchemaToPythonRDD.\n: org.apache.spark.SparkException: Keytab file: .keytab-f0b9b814-460e-4fa8-8e7d-029186b696c4 specified in spark.yarn.keytab does not exist\n\tat org.apache.spark.sql.hive.client.HiveClientImpl.(HiveClientImpl.scala:113)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:423)\n\tat org.apache.spark.sql.hive.client.IsolatedClientLoader.createClient(IsolatedClientLoader.scala:258)\n\tat org.apache.spark.sql.hive.HiveUtils$.newClientForMetadata(HiveUtils.scala:359)\n\tat org.apache.spark.sql.hive.HiveUtils$.newClientForMetadata(HiveUtils.scala:263)\n\tat org.apache.spark.sql.hive.HiveSharedState.metadataHive$lzycompute(HiveSharedState.scala:39)\n\tat org.apache.spark.sql.hive.HiveSharedState.metadataHive(HiveSharedState.scala:38)\n\tat org.apache.spark.sql.hive.HiveSharedState.externalCatalog$lzycompute(HiveSharedState.scala:46)\n\tat org.apache.spark.sql.hive.HiveSharedState.externalCatalog(HiveSharedState.scala:45)\n\tat org.apache.spark.sql.hive.HiveSessionState.catalog$lzycompute(HiveSessionState.scala:50)\n\tat org.apache.spark.sql.hive.HiveSessionState.catalog(HiveSessionState.scala:48)\n\tat org.apache.spark.sql.hive.HiveSessionState$$anon$1.(HiveSessionState.scala:63)\n\tat org.apache.spark.sql.hive.HiveSessionState.analyzer$lzycompute(HiveSessionState.scala:63)\n\tat org.apache.spark.sql.hive.HiveSessionState.analyzer(HiveSessionState.scala:62)\n\tat org.apache.spark.sql.execution.QueryExecution.assertAnalyzed(QueryExecution.scala:49)\n\tat org.apache.spark.sql.Dataset$.ofRows(Dataset.scala:64)\n\tat org.apache.spark.sql.SparkSession.applySchemaToPythonRDD(SparkSession.scala:666)\n\tat org.apache.spark.sql.SparkSession.applySchemaToPythonRDD(SparkSession.scala:656)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:237)\n\tat py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:357)\n\tat py4j.Gateway.invoke(Gateway.java:280)\n\tat py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:132)\n\tat py4j.commands.CallCommand.execute(CallCommand.java:79)\n\tat py4j.GatewayConnection.run(GatewayConnection.java:214)\n\tat java.lang.Thread.run(Thread.java:745)\n{noformat}\n\nh2. Steps to reproduce\n\n1. Pass valid --principal=user@REALM and --keytab=/home/user/.keytab to spark-submit\n2. Set spark.sql.catalogImplementation = 'hive'\n3. Set deploy mode to yarn-client\n4. Create a SparkSession and try to use session.createDataFrame()\n\nh2. Observations\n\n* The {{setupCredentials}} function in Client.scala sets {{spark.yarn.keytab}} to a UUID suffixed version of the base keytab filename without any path. For example, {{sparkContext.getConf().getAll()}} shows {{spark.yarn.keytab}} as having value {{.keytab-f0b9b814-460e-4fa8-8e7d-029186b696c4}}\n* When listing the contents of the application staging directory on HDFS, no suffixed file exists. Rather, the keytab file appears in the listing with its original name. For instance, {{hdfs dfs -ls hdfs://home/user/.sparkStaging/appication_big_uuid/}} shows an entry {{hdfs://home/user/.sparkStaging/appication_big_uuid/.keytab}}, but not {{hdfs://home/user/.sparkStaging/appication_big_uuid/.keytab-big-uuid}}.\n* The same exception noted above occurs even after I manually put a copy of the keytab with a filename matching the new value of {{spark.yarn.keytab}} onto HDFS in the staging directory.\n\nh2. Expected Behavior\n\nHiveClientImpl should be able to read {{spark.yarn.keytab}} to find the keytab file and initialize itself properly.\n\nh2. References\n\n* SPARK-8619 also noted trouble with the keytab property getting changed after app startup.","issue_id":"13031350","key":"SPARK-19038","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-02-24T17:33:30.000+0000","role":"fixed_distractor","summary":"Can't find keytab file when using Hive catalog"} {"case_id":"13032650","cluster":"DISTRACTOR-SPARK-19109","comments":[{"body":"It seems this JIRA describes upgrading the version of Hive dependency which is currently 1.2.1 up to my knowledge. \nThe stacktrace seems related with {{OrcFileOperator.getFileReader}} in Spark side which uses {{OrcFile.createReader}} to infer the schema.\nLet me leave a link related with this.","created":"2017-01-23T03:40:47.545+0000"},{"body":"That's one option, but we could also just backport the patch in HIVE-11592 to our Hive fork.","created":"2017-01-23T16:24:27.303+0000"},{"body":"I meet this problem and resolved .I re-complie source code of hive-exec-1.2.1-spark2.jar of spark-2.1.0.\nFirst download sourcecode of hive-exec-1.2.1-spark2.jar and the website is:\nhttps://github.com/JoshRosen\nSecond: download the patch and put into ReaderImpl.java\nhttps://issues.apache.org/jira/secure/attachment/12750949/HIVE-11592.1.patch \nThird:recompile and package the hive-exec-1.2.1-spark2.jar \nLast replace origin jar in spark-2.1.0/jars\n","created":"2017-08-09T08:04:48.863+0000"},{"body":"Hi, [~nseggert] and [~wangchao2017].\nCould you give us a way to reproduce this?","created":"2017-08-18T07:56:20.587+0000"},{"body":"hi, I meet this problem and resolved by recompile source code of hive-exec-1.2.1-spark2.jar of spark-2.1.0/jars\nThe source code website: https://github.com/JoshRosen \nSecond: download the patch and put into ReaderImpl.java\nhttps://issues.apache.org/jira/secure/attachment/12750949/HIVE-11592.1.patch \nthen put this patch into ReaderImpl.java in Intellij IDE.\nthen, you can recompile and package the source code;\nreplace the origin jar in spark/jars\n\n\n\nsydt2011@126.com\n \nFrom: Dongjoon Hyun (JIRA)\nDate: 2017-08-18 15:57\nTo: sydt2011\nSubject: [jira] [Commented] (SPARK-19109) ORC metadata section can sometimes exceed protobuf message size limit\n \n [ https://issues.apache.org/jira/browse/SPARK-19109?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16131877#comment-16131877 ] \n \nDongjoon Hyun commented on SPARK-19109:\n---------------------------------------\n \nHi, [~nseggert] and [~wangchao2017].\nCould you give us a way to reproduce this?\n \n \n \n \n--\nThis message was sent by Atlassian JIRA\n(v6.4.14#64029)\n","created":"2017-08-21T07:11:00.329+0000"},{"body":"[~dongjoon] I don't have the ability to reproduce it on my end anymore. We did the same thing as [~wangchao2017] and re-built Spark against a version of Hive that contains the fix. If I remember correctly, the file in question as a little under a TB and had a few hundred columns.","created":"2017-08-21T14:42:40.855+0000"},{"body":"Thanks, Nic Eggert and sydt. I see.\nI just wanted to resolve this officially in Apache Spark 2.3, but it's difficult to prove it without a valid test case.\nI'll try to find another way. Thank you all again.","created":"2017-08-21T16:49:53.329+0000"},{"body":"Thanks for your reply. I have compiled hive-exec-1.2.1-spark2.jar and it is about 10M ,it cannot be transfored to you because of your email size limit!\nyou can download the comiled jar in https://github.com/sydt2014/spark-hive \n \n\n\n\nsydt2011@126.com\n \nFrom: Dongjoon Hyun (JIRA)\nDate: 2017-08-22 00:50\nTo: sydt2011\nSubject: [jira] [Commented] (SPARK-19109) ORC metadata section can sometimes exceed protobuf message size limit\n \n [ https://issues.apache.org/jira/browse/SPARK-19109?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16135422#comment-16135422 ] \n \nDongjoon Hyun commented on SPARK-19109:\n---------------------------------------\n \nThanks, Nic Eggert and sydt. I see.\nI just wanted to resolve this officially in Apache Spark 2.3, but it's difficult to prove it without a valid test case.\nI'll try to find another way. Thank you all again.\n \n \n \n \n--\nThis message was sent by Atlassian JIRA\n(v6.4.14#64029)\n","created":"2017-08-22T02:13:00.206+0000"},{"body":"Thanks, [~wangchao2017], but I think there are misunderstanding between you and me. :)\nI already knew the Hive patch because [~nseggert] told that here.\n\nThe approach I want to do is to use new Apache ORC 1.4.0 directly. That's the reason why I ask the way of reproduce only here. I don't want to update the Hive fork of Spark. \n\nHowever, I really thank you for your attention and all your help!","created":"2017-08-22T05:11:32.143+0000"},{"body":"HIVE-11592 is fixed in Hive 1.3.0 and ORC 1.4.1 library has the patch. Since SPARK-20682 / SPARK-20728 / SPARK-22279, we are using native ORC implementation based on ORC 1.4.1. This issue is fixed in Apache Spark default configuration.\r\n{code}\r\npublic static final int PROTOBUF_MESSAGE_MAX_LIMIT = 1024 << 20; // 1GB\r\n{code}","created":"2018-01-17T06:27:00.441+0000"},{"body":"thanks for your attention! I am appreciated for your help !\r\n\r\n\r\n\r\n\r\n中国电信集团公司 王超\r\n部门中心:企业信息化事业部IT研发中心\r\n移动电话:18916929162\r\n工作邮箱:wangch@chinatelecom.cn \r\n通讯地址:上海市浦东新区秀沿西路189中国电信B23\r\n邮政编码:201315\r\n\r\n \r\nFrom: Dongjoon Hyun (JIRA)\r\nDate: 2018-01-17 14:29\r\nTo: sydt2011\r\nSubject: [jira] [Resolved] (SPARK-19109) ORC metadata section can sometimes exceed protobuf message size limit\r\n \r\n [ https://issues.apache.org/jira/browse/SPARK-19109?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]\r\n \r\nDongjoon Hyun resolved SPARK-19109.\r\n-----------------------------------\r\n Resolution: Fixed\r\n Fix Version/s: 2.3.0\r\n \r\n \r\n \r\n \r\n--\r\nThis message was sent by Atlassian JIRA\r\n(v7.6.3#76005)\n","created":"2018-01-17T06:46:00.196+0000"}],"conversations":[{"body":"Basically, Spark inherits HIVE-11592 from its Hive dependency. From that issue:\n\nIf there are too many small stripes and with many columns, the overhead for storing metadata (column stats) can exceed the default protobuf message size of 64MB. Reading such files will throw the following exception\n{code}\nException in thread \"main\" com.google.protobuf.InvalidProtocolBufferException: Protocol message was too large. May be malicious. Use CodedInputStream.setSizeLimit() to increase the size limit.\n at com.google.protobuf.InvalidProtocolBufferException.sizeLimitExceeded(InvalidProtocolBufferException.java:110)\n at com.google.protobuf.CodedInputStream.refillBuffer(CodedInputStream.java:755)\n at com.google.protobuf.CodedInputStream.readRawBytes(CodedInputStream.java:811)\n at com.google.protobuf.CodedInputStream.readBytes(CodedInputStream.java:329)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StringStatistics.(OrcProto.java:1331)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StringStatistics.(OrcProto.java:1281)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StringStatistics$1.parsePartialFrom(OrcProto.java:1374)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StringStatistics$1.parsePartialFrom(OrcProto.java:1369)\n at com.google.protobuf.CodedInputStream.readMessage(CodedInputStream.java:309)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$ColumnStatistics.(OrcProto.java:4887)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$ColumnStatistics.(OrcProto.java:4803)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$ColumnStatistics$1.parsePartialFrom(OrcProto.java:4990)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$ColumnStatistics$1.parsePartialFrom(OrcProto.java:4985)\n at com.google.protobuf.CodedInputStream.readMessage(CodedInputStream.java:309)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StripeStatistics.(OrcProto.java:12925)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StripeStatistics.(OrcProto.java:12872)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StripeStatistics$1.parsePartialFrom(OrcProto.java:12961)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StripeStatistics$1.parsePartialFrom(OrcProto.java:12956)\n at com.google.protobuf.CodedInputStream.readMessage(CodedInputStream.java:309)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Metadata.(OrcProto.java:13599)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Metadata.(OrcProto.java:13546)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Metadata$1.parsePartialFrom(OrcProto.java:13635)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Metadata$1.parsePartialFrom(OrcProto.java:13630)\n at com.google.protobuf.AbstractParser.parsePartialFrom(AbstractParser.java:200)\n at com.google.protobuf.AbstractParser.parseFrom(AbstractParser.java:217)\n at com.google.protobuf.AbstractParser.parseFrom(AbstractParser.java:223)\n at com.google.protobuf.AbstractParser.parseFrom(AbstractParser.java:49)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Metadata.parseFrom(OrcProto.java:13746)\n at org.apache.hadoop.hive.ql.io.orc.ReaderImpl$MetaInfoObjExtractor.(ReaderImpl.java:468)\n at org.apache.hadoop.hive.ql.io.orc.ReaderImpl.(ReaderImpl.java:314)\n at org.apache.hadoop.hive.ql.io.orc.OrcFile.createReader(OrcFile.java:228)\n at org.apache.hadoop.hive.ql.io.orc.FileDump.main(FileDump.java:67)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:606)\n at org.apache.hadoop.util.RunJar.run(RunJar.java:221)\n at org.apache.hadoop.util.RunJar.main(RunJar.java:136)\n{code}\n\nThis is fixed in Hive 1.3, so it should be fairly straightforward to pick up the patch.\n\nAs a side note: Spark's management of its Hive fork/dependency seems incredibly arcane to me. Surely there's a better way than publishing to central from developers' personal repos.","from":"reporter","subject":"ORC metadata section can sometimes exceed protobuf message size limit"},{"body":"It seems this JIRA describes upgrading the version of Hive dependency which is currently 1.2.1 up to my knowledge. \nThe stacktrace seems related with {{OrcFileOperator.getFileReader}} in Spark side which uses {{OrcFile.createReader}} to infer the schema.\nLet me leave a link related with this.","from":"developer"},{"body":"That's one option, but we could also just backport the patch in HIVE-11592 to our Hive fork.","from":"developer"},{"body":"I meet this problem and resolved .I re-complie source code of hive-exec-1.2.1-spark2.jar of spark-2.1.0.\nFirst download sourcecode of hive-exec-1.2.1-spark2.jar and the website is:\nhttps://github.com/JoshRosen\nSecond: download the patch and put into ReaderImpl.java\nhttps://issues.apache.org/jira/secure/attachment/12750949/HIVE-11592.1.patch \nThird:recompile and package the hive-exec-1.2.1-spark2.jar \nLast replace origin jar in spark-2.1.0/jars\n","from":"developer"},{"body":"Hi, [~nseggert] and [~wangchao2017].\nCould you give us a way to reproduce this?","from":"developer"},{"body":"hi, I meet this problem and resolved by recompile source code of hive-exec-1.2.1-spark2.jar of spark-2.1.0/jars\nThe source code website: https://github.com/JoshRosen \nSecond: download the patch and put into ReaderImpl.java\nhttps://issues.apache.org/jira/secure/attachment/12750949/HIVE-11592.1.patch \nthen put this patch into ReaderImpl.java in Intellij IDE.\nthen, you can recompile and package the source code;\nreplace the origin jar in spark/jars\n\n\n\nsydt2011@126.com\n \nFrom: Dongjoon Hyun (JIRA)\nDate: 2017-08-18 15:57\nTo: sydt2011\nSubject: [jira] [Commented] (SPARK-19109) ORC metadata section can sometimes exceed protobuf message size limit\n \n [ https://issues.apache.org/jira/browse/SPARK-19109?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16131877#comment-16131877 ] \n \nDongjoon Hyun commented on SPARK-19109:\n---------------------------------------\n \nHi, [~nseggert] and [~wangchao2017].\nCould you give us a way to reproduce this?\n \n \n \n \n--\nThis message was sent by Atlassian JIRA\n(v6.4.14#64029)\n","from":"developer"},{"body":"[~dongjoon] I don't have the ability to reproduce it on my end anymore. We did the same thing as [~wangchao2017] and re-built Spark against a version of Hive that contains the fix. If I remember correctly, the file in question as a little under a TB and had a few hundred columns.","from":"developer"},{"body":"Thanks, Nic Eggert and sydt. I see.\nI just wanted to resolve this officially in Apache Spark 2.3, but it's difficult to prove it without a valid test case.\nI'll try to find another way. Thank you all again.","from":"developer"},{"body":"Thanks for your reply. I have compiled hive-exec-1.2.1-spark2.jar and it is about 10M ,it cannot be transfored to you because of your email size limit!\nyou can download the comiled jar in https://github.com/sydt2014/spark-hive \n \n\n\n\nsydt2011@126.com\n \nFrom: Dongjoon Hyun (JIRA)\nDate: 2017-08-22 00:50\nTo: sydt2011\nSubject: [jira] [Commented] (SPARK-19109) ORC metadata section can sometimes exceed protobuf message size limit\n \n [ https://issues.apache.org/jira/browse/SPARK-19109?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16135422#comment-16135422 ] \n \nDongjoon Hyun commented on SPARK-19109:\n---------------------------------------\n \nThanks, Nic Eggert and sydt. I see.\nI just wanted to resolve this officially in Apache Spark 2.3, but it's difficult to prove it without a valid test case.\nI'll try to find another way. Thank you all again.\n \n \n \n \n--\nThis message was sent by Atlassian JIRA\n(v6.4.14#64029)\n","from":"developer"},{"body":"Thanks, [~wangchao2017], but I think there are misunderstanding between you and me. :)\nI already knew the Hive patch because [~nseggert] told that here.\n\nThe approach I want to do is to use new Apache ORC 1.4.0 directly. That's the reason why I ask the way of reproduce only here. I don't want to update the Hive fork of Spark. \n\nHowever, I really thank you for your attention and all your help!","from":"developer"},{"body":"HIVE-11592 is fixed in Hive 1.3.0 and ORC 1.4.1 library has the patch. Since SPARK-20682 / SPARK-20728 / SPARK-22279, we are using native ORC implementation based on ORC 1.4.1. This issue is fixed in Apache Spark default configuration.\r\n{code}\r\npublic static final int PROTOBUF_MESSAGE_MAX_LIMIT = 1024 << 20; // 1GB\r\n{code}","from":"developer"},{"body":"thanks for your attention! I am appreciated for your help !\r\n\r\n\r\n\r\n\r\n中国电信集团公司 王超\r\n部门中心:企业信息化事业部IT研发中心\r\n移动电话:18916929162\r\n工作邮箱:wangch@chinatelecom.cn \r\n通讯地址:上海市浦东新区秀沿西路189中国电信B23\r\n邮政编码:201315\r\n\r\n \r\nFrom: Dongjoon Hyun (JIRA)\r\nDate: 2018-01-17 14:29\r\nTo: sydt2011\r\nSubject: [jira] [Resolved] (SPARK-19109) ORC metadata section can sometimes exceed protobuf message size limit\r\n \r\n [ https://issues.apache.org/jira/browse/SPARK-19109?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]\r\n \r\nDongjoon Hyun resolved SPARK-19109.\r\n-----------------------------------\r\n Resolution: Fixed\r\n Fix Version/s: 2.3.0\r\n \r\n \r\n \r\n \r\n--\r\nThis message was sent by Atlassian JIRA\r\n(v7.6.3#76005)\n","from":"developer"}],"created":"2017-01-06T20:04:26.000+0000","description":"Basically, Spark inherits HIVE-11592 from its Hive dependency. From that issue:\n\nIf there are too many small stripes and with many columns, the overhead for storing metadata (column stats) can exceed the default protobuf message size of 64MB. Reading such files will throw the following exception\n{code}\nException in thread \"main\" com.google.protobuf.InvalidProtocolBufferException: Protocol message was too large. May be malicious. Use CodedInputStream.setSizeLimit() to increase the size limit.\n at com.google.protobuf.InvalidProtocolBufferException.sizeLimitExceeded(InvalidProtocolBufferException.java:110)\n at com.google.protobuf.CodedInputStream.refillBuffer(CodedInputStream.java:755)\n at com.google.protobuf.CodedInputStream.readRawBytes(CodedInputStream.java:811)\n at com.google.protobuf.CodedInputStream.readBytes(CodedInputStream.java:329)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StringStatistics.(OrcProto.java:1331)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StringStatistics.(OrcProto.java:1281)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StringStatistics$1.parsePartialFrom(OrcProto.java:1374)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StringStatistics$1.parsePartialFrom(OrcProto.java:1369)\n at com.google.protobuf.CodedInputStream.readMessage(CodedInputStream.java:309)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$ColumnStatistics.(OrcProto.java:4887)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$ColumnStatistics.(OrcProto.java:4803)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$ColumnStatistics$1.parsePartialFrom(OrcProto.java:4990)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$ColumnStatistics$1.parsePartialFrom(OrcProto.java:4985)\n at com.google.protobuf.CodedInputStream.readMessage(CodedInputStream.java:309)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StripeStatistics.(OrcProto.java:12925)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StripeStatistics.(OrcProto.java:12872)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StripeStatistics$1.parsePartialFrom(OrcProto.java:12961)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$StripeStatistics$1.parsePartialFrom(OrcProto.java:12956)\n at com.google.protobuf.CodedInputStream.readMessage(CodedInputStream.java:309)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Metadata.(OrcProto.java:13599)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Metadata.(OrcProto.java:13546)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Metadata$1.parsePartialFrom(OrcProto.java:13635)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Metadata$1.parsePartialFrom(OrcProto.java:13630)\n at com.google.protobuf.AbstractParser.parsePartialFrom(AbstractParser.java:200)\n at com.google.protobuf.AbstractParser.parseFrom(AbstractParser.java:217)\n at com.google.protobuf.AbstractParser.parseFrom(AbstractParser.java:223)\n at com.google.protobuf.AbstractParser.parseFrom(AbstractParser.java:49)\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Metadata.parseFrom(OrcProto.java:13746)\n at org.apache.hadoop.hive.ql.io.orc.ReaderImpl$MetaInfoObjExtractor.(ReaderImpl.java:468)\n at org.apache.hadoop.hive.ql.io.orc.ReaderImpl.(ReaderImpl.java:314)\n at org.apache.hadoop.hive.ql.io.orc.OrcFile.createReader(OrcFile.java:228)\n at org.apache.hadoop.hive.ql.io.orc.FileDump.main(FileDump.java:67)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:606)\n at org.apache.hadoop.util.RunJar.run(RunJar.java:221)\n at org.apache.hadoop.util.RunJar.main(RunJar.java:136)\n{code}\n\nThis is fixed in Hive 1.3, so it should be fairly straightforward to pick up the patch.\n\nAs a side note: Spark's management of its Hive fork/dependency seems incredibly arcane to me. Surely there's a better way than publishing to central from developers' personal repos.","issue_id":"13032650","key":"SPARK-19109","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-01-17T06:28:54.000+0000","role":"fixed_distractor","summary":"ORC metadata section can sometimes exceed protobuf message size limit"} {"case_id":"13037317","cluster":"DISTRACTOR-SPARK-19348","comments":[{"body":"Refer the attached pyspark_pipeline_threads.py file containing the complete code to reproduce the issue. It can be run on the pyspark shell or in a jupyter notebook. ","created":"2017-01-24T10:01:33.615+0000"},{"body":"The problem here is with the @keyword_only decorator that is used in the Pipeline constructor (and every other ML class also). It relies on saving the params, stages in this case, to a static class variable. When multiple threads call the wrapped constructor, it becomes a race condition to read from that static variable before it is over-written by another thread. I can put up a potential fix, but it affects all of PySpark ML so it would need to be checked out carefully.\n\nAs a workaround, you could just protect the construction of {{Pipeline}} with a shared lock. Other calls to {{fit}} etc, should be ok since they don't use that keyword_only decorator.","created":"2017-02-02T23:08:10.564+0000"},{"body":"User 'BryanCutler' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16782","created":"2017-02-02T23:27:06.068+0000"},{"body":"To save folks some time, the keyword_only decorator came in to pyspark/ml/util.py with SPARK-4586\ndesign:\nhttps://docs.google.com/document/d/1vL-4f5Xm-7t-kwVSaBylP_ZPrktPZjaOb2dWONtZU2s/edit\ncommit:\nhttps://github.com/apache/spark/commit/cd4a15366244657c4b7936abe5054754534366f2#diff-dd5670d3fb55faba1859e9778e4026e5","created":"2017-02-07T20:07:47.205+0000"},{"body":"Two things happen with this wrapper.\n\nFirst, it appears to me (after confirming with some simplified examples) that the modifications to incoming arguments made in the body of the wrapped function are lost (to take the pipeline.py example, stages=) when they are not updated in the dictionary that is passed in the calls made from inside the wrapped functions. The decorator explicitly states that it saves the 'actual input arguments' but does not clarify why. It seems as if the wrapped code should either update the dictionary or that lines of code that have no lasting effect should be deleted.\n\nSecond, bryanc has pointed out that passing the kwargs to the wrapped function via a static class variable is thread-unsafe within each of the many ml classes that use the decorator. Passing the kwargs as an instance variable as bryanc has proposed seems satisfactory for the second solution, as would using a thread-local class variable. Both require changes to any files using the decorator. Locking in the decorator, if it could be implemented in spite of the nested calls of decorated functions, could confine the changes to the decorator definition. Passing the wrapper's kwargs dictionary as an additional entry in the kwargs dictionary passed to the wrapped function would be threadsafe but change the public API of many functions. It looks possible for a decorator to introspect the wrapped function's parameters in which case the wrapper could pass the dictionary on one of those keywords (the purpose being to leave the public API intact), then inside the wrapped function there would need to be code to detect the wrapper, retrieve the dictionary and restore the coopted variable. (As-is, the wrapped functions already have wrapper-specific code in order to function.)","created":"2017-02-07T22:19:52.902+0000"},{"body":"Per the above, perhaps a fix could address both threadsafety and the orphaned modifications to variables in the wrapped function bodies. For instance, in Pipeline.__init__() and setParams() the fix could remove references to _input_kwargs and instead invoke setParams(stages=stages) within __init__() and _set(stages=stages) within setParams(), respectively, Parallel changes would be needed in all wrapped functions but the resulting code would be functional, readable, and threadsafe.\n\nA fix for Pipeline alone is insufficient, because multiple pipelines could have stages consisting of instances of the same class, e.g. LogisticRegression. All the classes using @keyword_only need to be addressed by whatever fix is decided upon.","created":"2017-02-09T16:22:33.875+0000"},{"body":"User 'BryanCutler' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17193","created":"2017-03-07T19:20:02.537+0000"},{"body":"Issue resolved by pull request 17195\n[https://github.com/apache/spark/pull/17195]","created":"2017-03-08T04:47:19.977+0000"}],"conversations":[{"body":"When pyspark.ml.Pipeline objects are constructed concurrently in separate python threads, it is observed that the stages used to construct a pipeline object get corrupted i.e the stages supplied to a Pipeline object in one thread appear inside a different Pipeline object constructed in a different thread. \n\nThings work fine if construction of pyspark.ml.Pipeline objects is serialized, so this looks like a thread safety problem with pyspark.ml.Pipeline object construction. \n\nConfirmed that the problem exists with Spark 1.6.x as well as 2.x.\n\nWhile the corruption of the Pipeline stages is easily caught, we need to know if performing other pipeline operations, such as pyspark.ml.pipeline.fit( ) are also affected by the underlying cause of this problem. That is, whether other pipeline operations like pyspark.ml.pipeline.fit( ) may be performed in separate threads (on distinct pipeline objects) concurrently without any cross contamination between them.","from":"reporter","subject":"pyspark.ml.Pipeline gets corrupted under multi threaded use"},{"body":"Refer the attached pyspark_pipeline_threads.py file containing the complete code to reproduce the issue. It can be run on the pyspark shell or in a jupyter notebook. ","from":"developer"},{"body":"The problem here is with the @keyword_only decorator that is used in the Pipeline constructor (and every other ML class also). It relies on saving the params, stages in this case, to a static class variable. When multiple threads call the wrapped constructor, it becomes a race condition to read from that static variable before it is over-written by another thread. I can put up a potential fix, but it affects all of PySpark ML so it would need to be checked out carefully.\n\nAs a workaround, you could just protect the construction of {{Pipeline}} with a shared lock. Other calls to {{fit}} etc, should be ok since they don't use that keyword_only decorator.","from":"developer"},{"body":"User 'BryanCutler' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16782","from":"developer"},{"body":"To save folks some time, the keyword_only decorator came in to pyspark/ml/util.py with SPARK-4586\ndesign:\nhttps://docs.google.com/document/d/1vL-4f5Xm-7t-kwVSaBylP_ZPrktPZjaOb2dWONtZU2s/edit\ncommit:\nhttps://github.com/apache/spark/commit/cd4a15366244657c4b7936abe5054754534366f2#diff-dd5670d3fb55faba1859e9778e4026e5","from":"developer"},{"body":"Two things happen with this wrapper.\n\nFirst, it appears to me (after confirming with some simplified examples) that the modifications to incoming arguments made in the body of the wrapped function are lost (to take the pipeline.py example, stages=) when they are not updated in the dictionary that is passed in the calls made from inside the wrapped functions. The decorator explicitly states that it saves the 'actual input arguments' but does not clarify why. It seems as if the wrapped code should either update the dictionary or that lines of code that have no lasting effect should be deleted.\n\nSecond, bryanc has pointed out that passing the kwargs to the wrapped function via a static class variable is thread-unsafe within each of the many ml classes that use the decorator. Passing the kwargs as an instance variable as bryanc has proposed seems satisfactory for the second solution, as would using a thread-local class variable. Both require changes to any files using the decorator. Locking in the decorator, if it could be implemented in spite of the nested calls of decorated functions, could confine the changes to the decorator definition. Passing the wrapper's kwargs dictionary as an additional entry in the kwargs dictionary passed to the wrapped function would be threadsafe but change the public API of many functions. It looks possible for a decorator to introspect the wrapped function's parameters in which case the wrapper could pass the dictionary on one of those keywords (the purpose being to leave the public API intact), then inside the wrapped function there would need to be code to detect the wrapper, retrieve the dictionary and restore the coopted variable. (As-is, the wrapped functions already have wrapper-specific code in order to function.)","from":"developer"},{"body":"Per the above, perhaps a fix could address both threadsafety and the orphaned modifications to variables in the wrapped function bodies. For instance, in Pipeline.__init__() and setParams() the fix could remove references to _input_kwargs and instead invoke setParams(stages=stages) within __init__() and _set(stages=stages) within setParams(), respectively, Parallel changes would be needed in all wrapped functions but the resulting code would be functional, readable, and threadsafe.\n\nA fix for Pipeline alone is insufficient, because multiple pipelines could have stages consisting of instances of the same class, e.g. LogisticRegression. All the classes using @keyword_only need to be addressed by whatever fix is decided upon.","from":"developer"},{"body":"User 'BryanCutler' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17193","from":"developer"},{"body":"Issue resolved by pull request 17195\n[https://github.com/apache/spark/pull/17195]","from":"developer"}],"created":"2017-01-24T09:54:57.000+0000","description":"When pyspark.ml.Pipeline objects are constructed concurrently in separate python threads, it is observed that the stages used to construct a pipeline object get corrupted i.e the stages supplied to a Pipeline object in one thread appear inside a different Pipeline object constructed in a different thread. \n\nThings work fine if construction of pyspark.ml.Pipeline objects is serialized, so this looks like a thread safety problem with pyspark.ml.Pipeline object construction. \n\nConfirmed that the problem exists with Spark 1.6.x as well as 2.x.\n\nWhile the corruption of the Pipeline stages is easily caught, we need to know if performing other pipeline operations, such as pyspark.ml.pipeline.fit( ) are also affected by the underlying cause of this problem. That is, whether other pipeline operations like pyspark.ml.pipeline.fit( ) may be performed in separate threads (on distinct pipeline objects) concurrently without any cross contamination between them.","issue_id":"13037317","key":"SPARK-19348","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-03-08T04:47:19.000+0000","role":"fixed_distractor","summary":"pyspark.ml.Pipeline gets corrupted under multi threaded use"} {"case_id":"13038203","cluster":"DISTRACTOR-SPARK-19372","comments":[{"body":"I was able to reproduce this. I am thinking how to reduce bytecode size per Java method.","created":"2017-02-03T13:09:05.512+0000"},{"body":"User 'kiszk' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17087","created":"2017-02-27T19:40:03.469+0000"},{"body":"I've seen this as well on parquet files.","created":"2017-03-23T19:23:19.880+0000"},{"body":"I have seen this too before.","created":"2017-03-24T00:22:51.405+0000"},{"body":"I implemented the code to take care of it, and am waiting for the review.","created":"2017-03-24T01:03:29.953+0000"},{"body":"Hi, [~kiszk]. I met this failure also.\nIs it possible to backport this to 2.2.0?","created":"2017-05-26T03:27:46.460+0000"},{"body":"+1 for backporting this to 2.2.0!","created":"2017-05-26T03:34:19.020+0000"},{"body":"I see. Let me create a PR for 2.2.0","created":"2017-05-26T04:31:43.007+0000"},{"body":"User 'kiszk' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18119","created":"2017-05-26T07:29:02.961+0000"},{"body":"Thank you so much all!","created":"2017-05-27T05:59:26.762+0000"},{"body":"Hi [~kiszk], the fix does not work in 2.2.0 for \n\nselect * from temp where not( Field1 = '' and Field2 = '' and Field3 = '' and Field4 = '' and Field5 = '' and BLANK_5 = '' and Field7 = '' and Field8 = '' and Field9 = '' and Field10 = '' and Field11 = '' and Field12 = '' and Field13 = '' and Field14 = '' and Field15 = '' and Field16 = '' and Field17 = '' and Field18 = '' and Field19 = '' and Field20 = '' and Field21 = '' and Field22 = '' and Field23 = '' and Field24 = '' and Field25 = '' and Field26 = '' and Field27 = '' and Field28 = '' and Field29 = '' and Field30 = '' and Field31 = '' and Field32 = '' and Field33 = '' and Field34 = '' and Field35 = '' and Field36 = '' and Field37 = '' and Field38 = '' and Field39 = '' and Field40 = '' and Field41 = '' and Field42 = '' and Field43 = '' and Field44 = '' and Field45 = '' and Field46 = '' and Field47 = '' and Field48 = '' and Field49 = '' and Field50 = '' and Field51 = '' and Field52 = '' and Field53 = '' and Field54 = '' and Field55 = '' and Field56 = '' and Field57 = '' and Field58 = '' and Field59 = '' and Field60 = '' and Field61 = '' and Field62 = '' and Field63 = '' and Field64 = '' and Field65 = '' and Field66 = '' and Field67 = '' and Field68 = '' and Field69 = '' and Field70 = '' and Field71 = '' and Field72 = '' and Field73 = '' and Field74 = '' and Field75 = '' and Field76 = '' and Field77 = '' and Field78 = '' and Field79 = '' and Field80 = '' and Field81 = '' and Field82 = '' and Field83 = '' and Field84 = '' and Field85 = '' and Field86 = '' and Field87 = '' and Field88 = '' and Field89 = '' and Field90 = '' and Field91 = '' and Field92 = '' and Field93 = '' and Field94 = '' and Field95 = '' and Field96 = '' and Field97 = '' and Field98 = '' and Field99 = '' and Field100 = '' and Field101 = '' and Field102 = '' and Field103 = '' and Field104 = '' and Field105 = '' and Field106 = '' and Field107 = '' and Field108 = '' and Field109 = '' and Field110 = '' and Field111 = '' and Field112 = '' and Field113 = '' and Field114 = '' and Field115 = '' and Field116 = '' and Field117 = '' and Field118 = '' and Field119 = '' and Field120 = '' and Field121 = '' and Field122 = '' and Field123 = '' and Field124 = '' and Field125 = '' and Field126 = '' and Field127 = '' and Field128 = '' and Field129 = '' and Field130 = '' and Field131 = '' and Field132 = '' and Field133 = '' and Field134 = '' and Field135 = '' and Field136 = '' and Field137 = '' and Field138 = '' and Field139 = '' and Field140 = '' and Field141 = '' and Field142 = '' and Field143 = '' and Field144 = '' and Field145 = '' and Field146 = '' and Field147 = '' and Field148 = '' and Field149 = '' and Field150 = '' and Field151 = '' and Field152 = '' and Field153 = '' and Field154 = '' and Field155 = '' and Field156 = '' and Field157 = '' and Field158 = '' and Field159 = '' and Field160 = '' and Field161 = '' and Field162 = '' and Field163 = '' and Field164 = '' and Field165 = '' and Field166 = '' and Field167 = '' and Field168 = '' and Field169 = '' and Field170 = '' and Field171 = '' and Field172 = '' and Field173 = '' and Field174 = '' and Field175 = '' and Field176 = '' and Field177 = '' and Field178 = '' and Field179 = '' and Field180 = '' and Field181 = '' and Field182 = '' and Field183 = '' and Field184 = '' and Field185 = '' and Field186 = '' and Field187 = '' and Field188 = '' and Field189 = '' and Field190 = '' and Field191 = '' and Field192 = '' and Field193 = '' and Field194 = '' and Field195 = '' and Field196 = '' and Field197 = '' and Field198 = '' and Field199 = '' and Field200 = '' and Field201 = '' and Field202 = '' and Field203 = '' and Field204 = '' and Field205 = '' and Field206 = '' and Field207 = '' and Field208 = '' and Field209 = '' and Field210 = '' and Field211 = '' and Field212 = '' and Field213 = '' and Field214 = '' and Field215 = '' and Field216 = '' and Field217 = '' and Field218 = '' and Field219 = '' and Field220 = '' and Field221 = '' and Field222 = '' and Field223 = '' and Field224 = '' and Field225 = '' and Field226 = '' and Field227 = '' and Field228 = '' and Field229 = '' and Field230 = '' and Field231 = '' and Field232 = '' and Field233 = '' and Field234 = '' and Field235 = '' and Field236 = '' and Field237 = '' and Field238 = '' and Field239 = '' and Field240 = '' and Field241 = '' and Field242 = '' and Field243 = '' and Field244 = '' and Field245 = '' and Field246 = '' and Field247 = '' and Field248 = '' and Field249 = '' and Field250 = '' and Field251 = '' and Field252 = '' and Field253 = '' and Field254 = '' and Field255 = '' and Field256 = '' and Field257 = '' and Field258 = '' and Field259 = '' and Field260 = '' and Field261 = '' and Field262 = '' and Field263 = '' and Field264 = '' and Field265 = '' and Field266 = '' and Field267 = '' and Field268 = '' and Field269 = '' and Field270 = '' and Field271 = '' and Field272 = '' and Field273 = '' and Field274 = '' and Field275 = '' and Field276 = '' and Field277 = '' and Field278 = '' and Field279 = '' and Field280 = '' and Field281 = '' and Field282 = '' and Field283 = '' and Field284 = '' and Field285 = '' and Field286 = '' and Field287 = '' and Field288 = '' and Field289 = '' and Field290 = '' and Field291 = '' and Field292 = '' and Field293 = '' and Field294 = '' and Field295 = '' and Field296 = '' and Field297 = '' and Field298 = '' and Field299 = '' and Field300 = '' and Field301 = '' and Field302 = '' and Field303 = '' and Field304 = '' and Field305 = '' and Field306 = '' and Field307 = '' and Field308 = '' and Field309 = '' and Field310 = '' and Field311 = '' and Field312 = '' and Field313 = '' and Field314 = '' and Field315 = '' and Field316 = '' and Field317 = '' and Field318 = '' and Field319 = '' and Field320 = '' and Field321 = '' and Field322 = '' and Field323 = '' and Field324 = '' and Field325 = '' and Field326 = '' and Field327 = '' and Field328 = '' and Field329 = '' and Field330 = '' and Field331 = '' and Field332 = '' and Field333 = '' and Field334 = '')\n\n\nThe error thrown is \n\n2017-08-11 16:25:48 ERROR Logging$class:91 - Exception in task 0.0 in stage 0.0 (TID 0)\njava.lang.StackOverflowError\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:370)\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:541)\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:541)\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:541)\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:541)\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:541)\n\nAny advice? Thanks","created":"2017-08-11T08:27:39.372+0000"},{"body":"Thank you for letting us know the problem. I investigate this.","created":"2017-08-13T09:30:46.966+0000"},{"body":"[~srinivasanm] I can reproduce this issue by using the master branch. I think that this is another problem.\nCould you please create another JIRA entry to track this issue? I will work for this.\n","created":"2017-08-13T17:05:17.212+0000"},{"body":"[~kiszk] I created a new ticket , https://issues.apache.org/jira/browse/SPARK-21720\nPlease verify the description. Thanks","created":"2017-08-14T02:14:24.587+0000"},{"body":"Hi, [~kiszk] I met this failure also.\nIs it possible to backport this to 2.1.1?\nAppreciate it!","created":"2017-08-14T18:35:04.148+0000"},{"body":"I opened a PR for backporting this at https://github.com/apache/spark/pull/18942, thanks.","created":"2017-08-14T23:34:01.151+0000"},{"body":"User 'poplav' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18942","created":"2017-08-28T16:55:23.557+0000"}],"conversations":[{"body":"For the attached csv file, the code below causes the exception \"org.codehaus.janino.JaninoRuntimeException: Code of method \"(Lorg/apache/spark/sql/catalyst/InternalRow;)Z\" of class \"org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificPredicate\" grows beyond 64 KB\n\nCode:\n{code:borderStyle=solid}\n val conf = new SparkConf().setMaster(\"local[1]\")\n val sqlContext = SparkSession.builder().config(conf).getOrCreate().sqlContext\n\n val dataframe =\n sqlContext\n .read\n .format(\"com.databricks.spark.csv\")\n .load(\"wide400cols.csv\")\n\n val filter = (0 to 399)\n .foldLeft(lit(false))((e, index) => e.or(dataframe.col(dataframe.columns(index)) =!= s\"column${index+1}\"))\n\n val filtered = dataframe.filter(filter)\n filtered.show(100)\n{code}","from":"reporter","subject":"Code generation for Filter predicate including many OR conditions exceeds JVM method size limit "},{"body":"I was able to reproduce this. I am thinking how to reduce bytecode size per Java method.","from":"developer"},{"body":"User 'kiszk' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17087","from":"developer"},{"body":"I've seen this as well on parquet files.","from":"developer"},{"body":"I have seen this too before.","from":"developer"},{"body":"I implemented the code to take care of it, and am waiting for the review.","from":"developer"},{"body":"Hi, [~kiszk]. I met this failure also.\nIs it possible to backport this to 2.2.0?","from":"developer"},{"body":"+1 for backporting this to 2.2.0!","from":"developer"},{"body":"I see. Let me create a PR for 2.2.0","from":"developer"},{"body":"User 'kiszk' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18119","from":"developer"},{"body":"Thank you so much all!","from":"developer"},{"body":"Hi [~kiszk], the fix does not work in 2.2.0 for \n\nselect * from temp where not( Field1 = '' and Field2 = '' and Field3 = '' and Field4 = '' and Field5 = '' and BLANK_5 = '' and Field7 = '' and Field8 = '' and Field9 = '' and Field10 = '' and Field11 = '' and Field12 = '' and Field13 = '' and Field14 = '' and Field15 = '' and Field16 = '' and Field17 = '' and Field18 = '' and Field19 = '' and Field20 = '' and Field21 = '' and Field22 = '' and Field23 = '' and Field24 = '' and Field25 = '' and Field26 = '' and Field27 = '' and Field28 = '' and Field29 = '' and Field30 = '' and Field31 = '' and Field32 = '' and Field33 = '' and Field34 = '' and Field35 = '' and Field36 = '' and Field37 = '' and Field38 = '' and Field39 = '' and Field40 = '' and Field41 = '' and Field42 = '' and Field43 = '' and Field44 = '' and Field45 = '' and Field46 = '' and Field47 = '' and Field48 = '' and Field49 = '' and Field50 = '' and Field51 = '' and Field52 = '' and Field53 = '' and Field54 = '' and Field55 = '' and Field56 = '' and Field57 = '' and Field58 = '' and Field59 = '' and Field60 = '' and Field61 = '' and Field62 = '' and Field63 = '' and Field64 = '' and Field65 = '' and Field66 = '' and Field67 = '' and Field68 = '' and Field69 = '' and Field70 = '' and Field71 = '' and Field72 = '' and Field73 = '' and Field74 = '' and Field75 = '' and Field76 = '' and Field77 = '' and Field78 = '' and Field79 = '' and Field80 = '' and Field81 = '' and Field82 = '' and Field83 = '' and Field84 = '' and Field85 = '' and Field86 = '' and Field87 = '' and Field88 = '' and Field89 = '' and Field90 = '' and Field91 = '' and Field92 = '' and Field93 = '' and Field94 = '' and Field95 = '' and Field96 = '' and Field97 = '' and Field98 = '' and Field99 = '' and Field100 = '' and Field101 = '' and Field102 = '' and Field103 = '' and Field104 = '' and Field105 = '' and Field106 = '' and Field107 = '' and Field108 = '' and Field109 = '' and Field110 = '' and Field111 = '' and Field112 = '' and Field113 = '' and Field114 = '' and Field115 = '' and Field116 = '' and Field117 = '' and Field118 = '' and Field119 = '' and Field120 = '' and Field121 = '' and Field122 = '' and Field123 = '' and Field124 = '' and Field125 = '' and Field126 = '' and Field127 = '' and Field128 = '' and Field129 = '' and Field130 = '' and Field131 = '' and Field132 = '' and Field133 = '' and Field134 = '' and Field135 = '' and Field136 = '' and Field137 = '' and Field138 = '' and Field139 = '' and Field140 = '' and Field141 = '' and Field142 = '' and Field143 = '' and Field144 = '' and Field145 = '' and Field146 = '' and Field147 = '' and Field148 = '' and Field149 = '' and Field150 = '' and Field151 = '' and Field152 = '' and Field153 = '' and Field154 = '' and Field155 = '' and Field156 = '' and Field157 = '' and Field158 = '' and Field159 = '' and Field160 = '' and Field161 = '' and Field162 = '' and Field163 = '' and Field164 = '' and Field165 = '' and Field166 = '' and Field167 = '' and Field168 = '' and Field169 = '' and Field170 = '' and Field171 = '' and Field172 = '' and Field173 = '' and Field174 = '' and Field175 = '' and Field176 = '' and Field177 = '' and Field178 = '' and Field179 = '' and Field180 = '' and Field181 = '' and Field182 = '' and Field183 = '' and Field184 = '' and Field185 = '' and Field186 = '' and Field187 = '' and Field188 = '' and Field189 = '' and Field190 = '' and Field191 = '' and Field192 = '' and Field193 = '' and Field194 = '' and Field195 = '' and Field196 = '' and Field197 = '' and Field198 = '' and Field199 = '' and Field200 = '' and Field201 = '' and Field202 = '' and Field203 = '' and Field204 = '' and Field205 = '' and Field206 = '' and Field207 = '' and Field208 = '' and Field209 = '' and Field210 = '' and Field211 = '' and Field212 = '' and Field213 = '' and Field214 = '' and Field215 = '' and Field216 = '' and Field217 = '' and Field218 = '' and Field219 = '' and Field220 = '' and Field221 = '' and Field222 = '' and Field223 = '' and Field224 = '' and Field225 = '' and Field226 = '' and Field227 = '' and Field228 = '' and Field229 = '' and Field230 = '' and Field231 = '' and Field232 = '' and Field233 = '' and Field234 = '' and Field235 = '' and Field236 = '' and Field237 = '' and Field238 = '' and Field239 = '' and Field240 = '' and Field241 = '' and Field242 = '' and Field243 = '' and Field244 = '' and Field245 = '' and Field246 = '' and Field247 = '' and Field248 = '' and Field249 = '' and Field250 = '' and Field251 = '' and Field252 = '' and Field253 = '' and Field254 = '' and Field255 = '' and Field256 = '' and Field257 = '' and Field258 = '' and Field259 = '' and Field260 = '' and Field261 = '' and Field262 = '' and Field263 = '' and Field264 = '' and Field265 = '' and Field266 = '' and Field267 = '' and Field268 = '' and Field269 = '' and Field270 = '' and Field271 = '' and Field272 = '' and Field273 = '' and Field274 = '' and Field275 = '' and Field276 = '' and Field277 = '' and Field278 = '' and Field279 = '' and Field280 = '' and Field281 = '' and Field282 = '' and Field283 = '' and Field284 = '' and Field285 = '' and Field286 = '' and Field287 = '' and Field288 = '' and Field289 = '' and Field290 = '' and Field291 = '' and Field292 = '' and Field293 = '' and Field294 = '' and Field295 = '' and Field296 = '' and Field297 = '' and Field298 = '' and Field299 = '' and Field300 = '' and Field301 = '' and Field302 = '' and Field303 = '' and Field304 = '' and Field305 = '' and Field306 = '' and Field307 = '' and Field308 = '' and Field309 = '' and Field310 = '' and Field311 = '' and Field312 = '' and Field313 = '' and Field314 = '' and Field315 = '' and Field316 = '' and Field317 = '' and Field318 = '' and Field319 = '' and Field320 = '' and Field321 = '' and Field322 = '' and Field323 = '' and Field324 = '' and Field325 = '' and Field326 = '' and Field327 = '' and Field328 = '' and Field329 = '' and Field330 = '' and Field331 = '' and Field332 = '' and Field333 = '' and Field334 = '')\n\n\nThe error thrown is \n\n2017-08-11 16:25:48 ERROR Logging$class:91 - Exception in task 0.0 in stage 0.0 (TID 0)\njava.lang.StackOverflowError\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:370)\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:541)\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:541)\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:541)\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:541)\n\tat org.codehaus.janino.CodeContext.flowAnalysis(CodeContext.java:541)\n\nAny advice? Thanks","from":"developer"},{"body":"Thank you for letting us know the problem. I investigate this.","from":"developer"},{"body":"[~srinivasanm] I can reproduce this issue by using the master branch. I think that this is another problem.\nCould you please create another JIRA entry to track this issue? I will work for this.\n","from":"developer"},{"body":"[~kiszk] I created a new ticket , https://issues.apache.org/jira/browse/SPARK-21720\nPlease verify the description. Thanks","from":"developer"},{"body":"Hi, [~kiszk] I met this failure also.\nIs it possible to backport this to 2.1.1?\nAppreciate it!","from":"developer"},{"body":"I opened a PR for backporting this at https://github.com/apache/spark/pull/18942, thanks.","from":"developer"},{"body":"User 'poplav' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18942","from":"developer"}],"created":"2017-01-26T17:26:59.000+0000","description":"For the attached csv file, the code below causes the exception \"org.codehaus.janino.JaninoRuntimeException: Code of method \"(Lorg/apache/spark/sql/catalyst/InternalRow;)Z\" of class \"org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificPredicate\" grows beyond 64 KB\n\nCode:\n{code:borderStyle=solid}\n val conf = new SparkConf().setMaster(\"local[1]\")\n val sqlContext = SparkSession.builder().config(conf).getOrCreate().sqlContext\n\n val dataframe =\n sqlContext\n .read\n .format(\"com.databricks.spark.csv\")\n .load(\"wide400cols.csv\")\n\n val filter = (0 to 399)\n .foldLeft(lit(false))((e, index) => e.or(dataframe.col(dataframe.columns(index)) =!= s\"column${index+1}\"))\n\n val filtered = dataframe.filter(filter)\n filtered.show(100)\n{code}","issue_id":"13038203","key":"SPARK-19372","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-05-16T21:47:42.000+0000","role":"fixed_distractor","summary":"Code generation for Filter predicate including many OR conditions exceeds JVM method size limit "} {"case_id":"13040178","cluster":"DISTRACTOR-SPARK-19451","comments":[{"body":"Good catch! I will dig deeply into code and fix it if it is really a bug.","created":"2017-02-06T01:53:43.582+0000"},{"body":"Thanks [~uncleGen] for your answer.\n\nI was tempted to try to fix it myself.. But not sure of possible side effects.\nIf I can help you in any way ( code / debug / tests ), feel free to ask.\n\nP.S. : I think this affects all version of Spark with Window functions, but I have only tested it on spark 1.6.1 and 2.0.2","created":"2017-02-06T08:31:54.058+0000"},{"body":"User 'uncleGen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16818","created":"2017-02-06T09:44:05.130+0000"},{"body":"[~jchamp] I have taken a fast look through the code, and did not find any strong point to set the type of index as Int. In the window function, there is really a underlying integer overflow issue. I made a pull request (https://github.com/apache/spark/pull/16818), and any suggestion is appreciated.","created":"2017-02-06T09:58:13.549+0000"},{"body":"[~jchamp] how may rows are in your partitions? 2 billion? So this is an oversight, but I am not sure we should even try to support more than {{1 << 32 - 1}} values in a partition.","created":"2017-02-06T10:05:42.197+0000"},{"body":"Let's imagine that this window is used on timestamp values in ms : I can ask for a window with a range between [-2160000000L, 0] and only have a few values inside, not necessarily 2160000000L.\n\nI can understand the limitation for the rowBetween() method but the rangeBetween() method is nice for this kind of usage.","created":"2017-02-06T10:19:23.718+0000"},{"body":"Yeah, you are right about that. We should definitely support this.","created":"2017-02-06T10:23:26.583+0000"},{"body":"Glad to see that I'm not the only one convinced by this usage !\n\nThis probably needs to use different data structures for rowBetween() and rangeBetween()","created":"2017-02-06T12:52:53.540+0000"},{"body":"At the end of the day I would like to support arbitrary literals for range frames. See: https://issues.apache.org/jira/browse/SPARK-9221","created":"2017-02-06T13:11:46.966+0000"},{"body":"Any news on this bug / feature request ?\n\nOr any workaround ? May be using stream I can efficiently do what I want ?","created":"2017-07-04T09:22:47.726+0000"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18540","created":"2017-07-05T09:15:05.028+0000"},{"body":"Nice so see that there is such a pull request for this issue !\n\nIf I can help you in some way, feel free to ask.","created":"2017-07-11T09:12:17.530+0000"}],"conversations":[{"body":"Hi there,\n\nthere seems to be a major limitation in spark window functions and rangeBetween method.\n\nIf I have the following code :\n{code:title=Exemple |borderStyle=solid}\n val tw = Window.orderBy(\"date\")\n .partitionBy(\"id\")\n .rangeBetween( from , 0)\n{code}\n\nEverything seems ok, while *from* value is not too large... Even if the rangeBetween() method supports Long parameters.\nBut.... If i set *-2160000000L* value to *from* it does not work !\n\nIt is probably related to this part of code in the between() method, of the WindowSpec class, called by rangeBetween()\n\n{code:title=between() method|borderStyle=solid}\n val boundaryStart = start match {\n case 0 => CurrentRow\n case Long.MinValue => UnboundedPreceding\n case x if x < 0 => ValuePreceding(-start.toInt)\n case x if x > 0 => ValueFollowing(start.toInt)\n }\n{code}\n( look at this *.toInt* )\n\nDoes anybody know it there's a way to solve / patch this behavior ?\n\nAny help will be appreciated\n\nThx","from":"reporter","subject":"rangeBetween method should accept Long value as boundary"},{"body":"Good catch! I will dig deeply into code and fix it if it is really a bug.","from":"developer"},{"body":"Thanks [~uncleGen] for your answer.\n\nI was tempted to try to fix it myself.. But not sure of possible side effects.\nIf I can help you in any way ( code / debug / tests ), feel free to ask.\n\nP.S. : I think this affects all version of Spark with Window functions, but I have only tested it on spark 1.6.1 and 2.0.2","from":"developer"},{"body":"User 'uncleGen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16818","from":"developer"},{"body":"[~jchamp] I have taken a fast look through the code, and did not find any strong point to set the type of index as Int. In the window function, there is really a underlying integer overflow issue. I made a pull request (https://github.com/apache/spark/pull/16818), and any suggestion is appreciated.","from":"developer"},{"body":"[~jchamp] how may rows are in your partitions? 2 billion? So this is an oversight, but I am not sure we should even try to support more than {{1 << 32 - 1}} values in a partition.","from":"developer"},{"body":"Let's imagine that this window is used on timestamp values in ms : I can ask for a window with a range between [-2160000000L, 0] and only have a few values inside, not necessarily 2160000000L.\n\nI can understand the limitation for the rowBetween() method but the rangeBetween() method is nice for this kind of usage.","from":"developer"},{"body":"Yeah, you are right about that. We should definitely support this.","from":"developer"},{"body":"Glad to see that I'm not the only one convinced by this usage !\n\nThis probably needs to use different data structures for rowBetween() and rangeBetween()","from":"developer"},{"body":"At the end of the day I would like to support arbitrary literals for range frames. See: https://issues.apache.org/jira/browse/SPARK-9221","from":"developer"},{"body":"Any news on this bug / feature request ?\n\nOr any workaround ? May be using stream I can efficiently do what I want ?","from":"developer"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18540","from":"developer"},{"body":"Nice so see that there is such a pull request for this issue !\n\nIf I can help you in some way, feel free to ask.","from":"developer"}],"created":"2017-02-03T17:29:33.000+0000","description":"Hi there,\n\nthere seems to be a major limitation in spark window functions and rangeBetween method.\n\nIf I have the following code :\n{code:title=Exemple |borderStyle=solid}\n val tw = Window.orderBy(\"date\")\n .partitionBy(\"id\")\n .rangeBetween( from , 0)\n{code}\n\nEverything seems ok, while *from* value is not too large... Even if the rangeBetween() method supports Long parameters.\nBut.... If i set *-2160000000L* value to *from* it does not work !\n\nIt is probably related to this part of code in the between() method, of the WindowSpec class, called by rangeBetween()\n\n{code:title=between() method|borderStyle=solid}\n val boundaryStart = start match {\n case 0 => CurrentRow\n case Long.MinValue => UnboundedPreceding\n case x if x < 0 => ValuePreceding(-start.toInt)\n case x if x > 0 => ValueFollowing(start.toInt)\n }\n{code}\n( look at this *.toInt* )\n\nDoes anybody know it there's a way to solve / patch this behavior ?\n\nAny help will be appreciated\n\nThx","issue_id":"13040178","key":"SPARK-19451","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-07-29T17:12:28.000+0000","role":"fixed_distractor","summary":"rangeBetween method should accept Long value as boundary"} {"case_id":"13043870","cluster":"DISTRACTOR-SPARK-19644","comments":[{"body":"Snap shot of heap dump after 50 hours","created":"2017-02-17T06:06:33.125+0000"},{"body":"What you have described so far is not a memory leak in Spark. It's normal for the heap to grow unless it has reason to even garbage collect. You're not evidently running out of memory. You're talking about a heap change of 1.1 to 1.3MB, which is trivial (is this a typo?). I'd close this unless you have a clearer case.","created":"2017-02-17T08:38:53.227+0000"},{"body":"you can ignore that change in memory but if you look in the snapshot the number of instances of class scala.collection.immutable.$colon$colon it had increased too high and it keep on increasing over the period of time","created":"2017-02-17T08:43:50.808+0000"},{"body":"There are just linked list objects. Why are they too high? if you have plenty of heap, Java won't bother GCing until it needs to.","created":"2017-02-17T08:50:10.646+0000"},{"body":"No i had just given 2 GB driver memory and if there is not any reference to them the Full GC should clean them but it is not get cleaned thats why i think there is memory leak","created":"2017-02-17T09:03:26.148+0000"},{"body":"And after 40-50 hours full gc is too frequent that all cores of machines are over utilized and batches start to queue up in streaming and I need to restart the streaming","created":"2017-02-17T09:14:52.241+0000"},{"body":"You only show 400MB of lists in your screen dump.\nRunning out of memory doesn't mean a leak. The question is what is holding on to the memory? This isn't a big heap to begin with.\n","created":"2017-02-17T09:19:09.295+0000"},{"body":"Yes that's right running out of memory doesn't mean a leak but gradual increase in heap size and inability of GC to clear the memory is a memory leak. Ideally the number of linked list objects should not be increasing over the period of time and that increase is suggesting that there is a memory leak. ","created":"2017-02-17T09:28:18.933+0000"},{"body":"This is still not necessarily true in general. Your app could be retaining state. This is why this isn't actionable as-is.","created":"2017-02-17T09:29:56.312+0000"},{"body":"We are not using any state or window operation and not using any check pointing so i don't think that app is retaining state.","created":"2017-02-17T09:34:06.529+0000"},{"body":"[~deenbandhu] Could you check the GC root, please? These objects are from Scala reflection. Did you run the job in Spark shell?","created":"2017-02-17T22:39:40.854+0000"},{"body":"Sorry for the delayed response.\n\nNo, I didn't run in spark shell. I ran using spark submit in client deploy mode on a standalone spark cluster.\nI ran eclipse MAT on heap dump and attached the screenshot of dominator tree. I hope this will help you out to find the cause of memory leak. \n\nAlso attached Path to GC root of the object of `scala.reflect.runtime.JavaUniverse` (for smaller heap dump taken at the application start).\n\nWhen I checked the path to GC root for an object of `scala.collection.immutable.$colon$colon` the path contains the same object(`scala.reflect.runtime.JavaUniverse`)","created":"2017-02-20T14:39:44.552+0000"},{"body":"I have analysed the issue more. I performed some of the experiments as follows and analysed the heapdump using jvisualvm at some intervals.\n\n1. Dstream.foreachRdd(rdd => rdd.map(r => someCaseClass(r)).take(10).foreach(println))\n\n2. Dstream.foreachRdd(rdd => rdd.map(r => someCaseClass(r)).toDF.show(10,false))\n\n3. Dstream.foreachRdd(rdd => rdd.map(r => someCaseClass(r)).toDS.show(10,false))\n\nI Observed that the number of instances of scala.collection.immutable.$colon$colon remain constant in 1 scenario but it keeps on increasing in 2 and 3 scenario. So I think there is something leaky in toDS or toDF function this may help you out to find out the issue.","created":"2017-02-21T09:53:46.643+0000"},{"body":"[~deenbandhu] Do you use Scala 2.10 or Scala 2.11?","created":"2017-02-22T22:44:41.549+0000"},{"body":"I am using scala 2.11","created":"2017-02-23T04:12:28.416+0000"},{"body":"Any updates ??","created":"2017-03-17T13:35:13.039+0000"},{"body":"I don't think there is evidence of a memory leak here. It's not even clear it has GCed","created":"2017-03-18T09:50:54.540+0000"},{"body":"It's not even clear it has GCed ?\n\nThe increase in total GC time is a clear indication of GC ","created":"2017-03-19T15:58:37.580+0000"},{"body":"I don't see any GC time here. \nIs this not simple stuff like you have lots of old jobs and stage metrics info in the driver? There is not much memory used here compared to normal operation. Unless you've tried stuff like restricting the number of retained jobs in the UI I don't think there is evidence of a problem. The dumps don't show anything that odd. ","created":"2017-03-19T16:03:38.735+0000"},{"body":"Yes i have tried restricting the number of jobs retained in UI to 200 and moreover the default value is 1000 for number of retained batches and my batch interval is 10s so for 1000 batches it will take somewhere around 10000 sec which is equal to 3-4 hrs but it keeps on accumulating after that. I think there is something else which is creating problem ","created":"2017-03-19T16:07:43.719+0000"},{"body":"Use jcmd to trigger a full GC on the process to see what happens. ","created":"2017-03-19T16:09:39.189+0000"},{"body":"full gc is trigger so many times and its frequency increases with time because of accumulated memory of that big object ","created":"2017-03-19T16:11:28.286+0000"},{"body":"The weird thing is memory retained by Scala runtime universe. I am still not clear if you are saying you run out memory or not. I also don't recall any other reports like this. If you have leads, post them here. ","created":"2017-03-19T16:17:06.421+0000"},{"body":"Hi, [~deenbandhu].\nCould you try this with the latest versions like 2.1.1 or 2.2.0-RC4?","created":"2017-06-19T05:49:53.879+0000"},{"body":"yes i can try but is there any report of such events in that particular version \n","created":"2017-06-19T05:51:38.049+0000"},{"body":"And the spark cassandra connector is also not out for those spark version. which is a dependency for us","created":"2017-06-19T05:55:01.978+0000"},{"body":"I see. Thank you anyway, [~deenbandhu].","created":"2017-06-19T06:02:23.373+0000"},{"body":"did the issue has been fixed? I am using spark 2.1.0 and i also encount the same scene. The driver memory keep increasing. I analysis the dump heap and find the class scala.collection.immutable.$colon$colon keep increasing.","created":"2017-11-01T07:57:20.198+0000"},{"body":"I happened to investigate a similar issue and found out the leak is caused by Scala reflection. Please see my comment in https://github.com/scala/bug/issues/8302\r\n\r\nMy workaround is calling \"scala.reflect.runtime.universe.asInstanceOf[scala.reflect.runtime.JavaUniverse].undoLog.clear()\" manually to clean up these garbage objects. You need to put this line in the same thread that you create Dataset/DataFrame as the leak happens in a thread local object. I think the best place in Spark streaming is foreachRDD.","created":"2017-11-01T19:54:13.849+0000"},{"body":"By the way, you can confirm this issue by checking if the number of \"scala.reflect.internal.tpe.TypeConstraints$TypeConstraint\" is large.","created":"2017-11-01T19:58:53.992+0000"},{"body":"I added more components since it also affects them. The major issue is when creating an encoder object, it leaks several Scala internal objects due to a Scala memory leak issue.","created":"2017-11-01T20:05:27.543+0000"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19687","created":"2017-11-08T00:38:05.344+0000"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19718","created":"2017-11-10T19:49:03.853+0000"}],"conversations":[{"body":"I am using streaming on the production for some aggregation and fetching data from cassandra and saving data back to cassandra. \r\n\r\nI see a gradual increase in old generation heap capacity from 1161216 Bytes to 1397760 Bytes over a period of six hours.\r\n\r\nAfter 50 hours of processing instances of class scala.collection.immutable.$colon$colon incresed to 12,811,793 which is a huge number. \r\n\r\nI think this is a clear case of memory leak\r\n\r\nUpdated: The root cause is when creating an encoder object, it leaks several Scala internal objects due to a Scala memory leak issue: https://github.com/scala/bug/issues/8302","from":"reporter","subject":"Memory leak in Spark Streaming (Encoder/Scala Reflection)"},{"body":"Snap shot of heap dump after 50 hours","from":"developer"},{"body":"What you have described so far is not a memory leak in Spark. It's normal for the heap to grow unless it has reason to even garbage collect. You're not evidently running out of memory. You're talking about a heap change of 1.1 to 1.3MB, which is trivial (is this a typo?). I'd close this unless you have a clearer case.","from":"developer"},{"body":"you can ignore that change in memory but if you look in the snapshot the number of instances of class scala.collection.immutable.$colon$colon it had increased too high and it keep on increasing over the period of time","from":"developer"},{"body":"There are just linked list objects. Why are they too high? if you have plenty of heap, Java won't bother GCing until it needs to.","from":"developer"},{"body":"No i had just given 2 GB driver memory and if there is not any reference to them the Full GC should clean them but it is not get cleaned thats why i think there is memory leak","from":"developer"},{"body":"And after 40-50 hours full gc is too frequent that all cores of machines are over utilized and batches start to queue up in streaming and I need to restart the streaming","from":"developer"},{"body":"You only show 400MB of lists in your screen dump.\nRunning out of memory doesn't mean a leak. The question is what is holding on to the memory? This isn't a big heap to begin with.\n","from":"developer"},{"body":"Yes that's right running out of memory doesn't mean a leak but gradual increase in heap size and inability of GC to clear the memory is a memory leak. Ideally the number of linked list objects should not be increasing over the period of time and that increase is suggesting that there is a memory leak. ","from":"developer"},{"body":"This is still not necessarily true in general. Your app could be retaining state. This is why this isn't actionable as-is.","from":"developer"},{"body":"We are not using any state or window operation and not using any check pointing so i don't think that app is retaining state.","from":"developer"},{"body":"[~deenbandhu] Could you check the GC root, please? These objects are from Scala reflection. Did you run the job in Spark shell?","from":"developer"},{"body":"Sorry for the delayed response.\n\nNo, I didn't run in spark shell. I ran using spark submit in client deploy mode on a standalone spark cluster.\nI ran eclipse MAT on heap dump and attached the screenshot of dominator tree. I hope this will help you out to find the cause of memory leak. \n\nAlso attached Path to GC root of the object of `scala.reflect.runtime.JavaUniverse` (for smaller heap dump taken at the application start).\n\nWhen I checked the path to GC root for an object of `scala.collection.immutable.$colon$colon` the path contains the same object(`scala.reflect.runtime.JavaUniverse`)","from":"developer"},{"body":"I have analysed the issue more. I performed some of the experiments as follows and analysed the heapdump using jvisualvm at some intervals.\n\n1. Dstream.foreachRdd(rdd => rdd.map(r => someCaseClass(r)).take(10).foreach(println))\n\n2. Dstream.foreachRdd(rdd => rdd.map(r => someCaseClass(r)).toDF.show(10,false))\n\n3. Dstream.foreachRdd(rdd => rdd.map(r => someCaseClass(r)).toDS.show(10,false))\n\nI Observed that the number of instances of scala.collection.immutable.$colon$colon remain constant in 1 scenario but it keeps on increasing in 2 and 3 scenario. So I think there is something leaky in toDS or toDF function this may help you out to find out the issue.","from":"developer"},{"body":"[~deenbandhu] Do you use Scala 2.10 or Scala 2.11?","from":"developer"},{"body":"I am using scala 2.11","from":"developer"},{"body":"Any updates ??","from":"developer"},{"body":"I don't think there is evidence of a memory leak here. It's not even clear it has GCed","from":"developer"},{"body":"It's not even clear it has GCed ?\n\nThe increase in total GC time is a clear indication of GC ","from":"developer"},{"body":"I don't see any GC time here. \nIs this not simple stuff like you have lots of old jobs and stage metrics info in the driver? There is not much memory used here compared to normal operation. Unless you've tried stuff like restricting the number of retained jobs in the UI I don't think there is evidence of a problem. The dumps don't show anything that odd. ","from":"developer"},{"body":"Yes i have tried restricting the number of jobs retained in UI to 200 and moreover the default value is 1000 for number of retained batches and my batch interval is 10s so for 1000 batches it will take somewhere around 10000 sec which is equal to 3-4 hrs but it keeps on accumulating after that. I think there is something else which is creating problem ","from":"developer"},{"body":"Use jcmd to trigger a full GC on the process to see what happens. ","from":"developer"},{"body":"full gc is trigger so many times and its frequency increases with time because of accumulated memory of that big object ","from":"developer"},{"body":"The weird thing is memory retained by Scala runtime universe. I am still not clear if you are saying you run out memory or not. I also don't recall any other reports like this. If you have leads, post them here. ","from":"developer"},{"body":"Hi, [~deenbandhu].\nCould you try this with the latest versions like 2.1.1 or 2.2.0-RC4?","from":"developer"},{"body":"yes i can try but is there any report of such events in that particular version \n","from":"developer"},{"body":"And the spark cassandra connector is also not out for those spark version. which is a dependency for us","from":"developer"},{"body":"I see. Thank you anyway, [~deenbandhu].","from":"developer"},{"body":"did the issue has been fixed? I am using spark 2.1.0 and i also encount the same scene. The driver memory keep increasing. I analysis the dump heap and find the class scala.collection.immutable.$colon$colon keep increasing.","from":"developer"},{"body":"I happened to investigate a similar issue and found out the leak is caused by Scala reflection. Please see my comment in https://github.com/scala/bug/issues/8302\r\n\r\nMy workaround is calling \"scala.reflect.runtime.universe.asInstanceOf[scala.reflect.runtime.JavaUniverse].undoLog.clear()\" manually to clean up these garbage objects. You need to put this line in the same thread that you create Dataset/DataFrame as the leak happens in a thread local object. I think the best place in Spark streaming is foreachRDD.","from":"developer"},{"body":"By the way, you can confirm this issue by checking if the number of \"scala.reflect.internal.tpe.TypeConstraints$TypeConstraint\" is large.","from":"developer"},{"body":"I added more components since it also affects them. The major issue is when creating an encoder object, it leaks several Scala internal objects due to a Scala memory leak issue.","from":"developer"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19687","from":"developer"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19718","from":"developer"}],"created":"2017-02-17T06:00:30.000+0000","description":"I am using streaming on the production for some aggregation and fetching data from cassandra and saving data back to cassandra. \r\n\r\nI see a gradual increase in old generation heap capacity from 1161216 Bytes to 1397760 Bytes over a period of six hours.\r\n\r\nAfter 50 hours of processing instances of class scala.collection.immutable.$colon$colon incresed to 12,811,793 which is a huge number. \r\n\r\nI think this is a clear case of memory leak\r\n\r\nUpdated: The root cause is when creating an encoder object, it leaks several Scala internal objects due to a Scala memory leak issue: https://github.com/scala/bug/issues/8302","issue_id":"13043870","key":"SPARK-19644","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-11-10T22:15:27.000+0000","role":"fixed_distractor","summary":"Memory leak in Spark Streaming (Encoder/Scala Reflection)"} {"case_id":"12717622","cluster":"DISTRACTOR-SPARK-1977","comments":[{"body":"I cannot reproduce this error in v1.0.0. There is an example called `MovieLensALS.scala` under `examples/`, which runs fine with kryo enabled. Did you include other dependencies in your application?","created":"2014-06-04T05:54:56.168+0000"},{"body":"Yeah that example worked fine for me in standalone mode but failed in YARN cluster mode with the same error.\nMaybe serialization wasn't needed/triggered in standalone mode?","created":"2014-06-04T07:16:26.087+0000"},{"body":"This is more likely a version conflict in your dependencies. From the Spark WebUI, you can find the system classpath in the environment tab. Please verify that you don't have two different versions of spark, kryo, or any other related library. Classes may hide inside an assembly jar.","created":"2014-06-04T08:07:15.131+0000"},{"body":"We submit 1 spark-assembly and 1 job assembly jar via spark-submit and there are no other obvious scala/spark/kryo jars in the global classpath. I can reproduce the same exception locally with the following snippet, when kryo.register() is commented out.\n\nI just added mutable BitSet to Twitter chill: https://github.com/twitter/chill/pull/185\n\n{code}\nimport com.twitter.chill._\nimport org.apache.spark.serializer.{KryoSerializer, KryoRegistrator}\nimport org.apache.spark.SparkConf\nimport scala.collection.mutable\n\nclass MyRegistrator extends KryoRegistrator {\n override def registerClasses(kryo: Kryo) {\n // kryo.register(classOf[mutable.BitSet])\n }\n}\n\ncase class OutLinkBlock(elementIds: Array[Int], shouldSend: Array[mutable.BitSet])\n\nobject KryoTest {\n def main(args: Array[String]) {\n println(\"hello\")\n val conf = new SparkConf()\n .set(\"spark.serializer\", \"org.apache.spark.serializer.KryoSerializer\")\n .set(\"spark.kryo.registrator\", classOf[MyRegistrator].getName)\n val serializer = new KryoSerializer(conf).newInstance()\n\n val bytes = serializer.serialize(OutLinkBlock(Array(1, 2, 3), Array(mutable.BitSet(2, 4, 6))))\n serializer.deserialize(bytes).asInstanceOf[OutLinkBlock]\n }\n}\n{code}","created":"2014-06-04T08:44:08.767+0000"},{"body":"In our example code, we only register `Rating` and it works. Could you try adding the following:\n\n{code}\nkryo.register(classOf[Rating])\n{code}\n\nI need to reproduce this problem with `ALS.train`.","created":"2014-06-04T20:50:17.223+0000"},{"body":"We are already doing that :)\nOur job works on YARN with \"register(classOf[mutable.BitSet])\". Without it we get the reported exception.","created":"2014-06-04T20:53:04.846+0000"},{"body":"Did you register `Rating`? I think this is necessary.","created":"2014-06-04T21:43:29.027+0000"},{"body":"Yes we did register 'Rating'. And we had to \"register(classOf[mutable.BitSet])\" in addition to make it work.","created":"2014-06-04T21:56:38.496+0000"},{"body":"Hi [~neville], I just run the MovieLens example on my YARN cluster (hadoop-2.0.5-alpha) with kryo enabled and it works. I use the following command:\n\nbin/spark-submit --master yarn-cluster --class org.apache.spark.examples.mllib.MovieLensALS --num-executors ** --driver-memory ** --executor-memory ** --executor-cores 1 spark-examples-1.0.0-hadoop2.0.5-alpha.jar --rank 5 --numIterations 20 --lambda 1.0 --kryo /path/to/sample_movielens_data.txt","created":"2014-06-04T23:25:36.575+0000"},{"body":"Our YARN cluster runs 2.2.0. We built spark-assembly and spark-examples jars with 1.0.0 release source and the bundled make_distribution.sh. And here's my command:\n\n{code}\nspark-submit --master yarn-cluster --class org.apache.spark.examples.mllib.MovieLensALS --num-executors 2 --executor-memory 2g --driver-memory 2g dist/lib/spark-examples-1.0.0-hadoop2.2.0.jar --kryo --implicitPrefs sample_movielens_data.txt\n{code}\n\nHere's a complete list of classpath from the environment tab.\n{code}\n/etc/hadoop/conf\n/usr/lib/hadoop-hdfs/hadoop-hdfs-2.2.0.2.0.6.0-76-tests.jar\n/usr/lib/hadoop-hdfs/hadoop-hdfs-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-hdfs/hadoop-hdfs-nfs-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-hdfs/lib/asm-3.2.jar\n/usr/lib/hadoop-hdfs/lib/commons-cli-1.2.jar\n/usr/lib/hadoop-hdfs/lib/commons-codec-1.4.jar\n/usr/lib/hadoop-hdfs/lib/commons-daemon-1.0.13.jar\n/usr/lib/hadoop-hdfs/lib/commons-el-1.0.jar\n/usr/lib/hadoop-hdfs/lib/commons-io-2.1.jar\n/usr/lib/hadoop-hdfs/lib/commons-lang-2.5.jar\n/usr/lib/hadoop-hdfs/lib/commons-logging-1.1.1.jar\n/usr/lib/hadoop-hdfs/lib/guava-11.0.2.jar\n/usr/lib/hadoop-hdfs/lib/jackson-core-asl-1.8.8.jar\n/usr/lib/hadoop-hdfs/lib/jackson-mapper-asl-1.8.8.jar\n/usr/lib/hadoop-hdfs/lib/jasper-runtime-5.5.23.jar\n/usr/lib/hadoop-hdfs/lib/jersey-core-1.9.jar\n/usr/lib/hadoop-hdfs/lib/jersey-server-1.9.jar\n/usr/lib/hadoop-hdfs/lib/jetty-6.1.26.jar\n/usr/lib/hadoop-hdfs/lib/jetty-util-6.1.26.jar\n/usr/lib/hadoop-hdfs/lib/jsp-api-2.1.jar\n/usr/lib/hadoop-hdfs/lib/jsr305-1.3.9.jar\n/usr/lib/hadoop-hdfs/lib/log4j-1.2.17.jar\n/usr/lib/hadoop-hdfs/lib/netty-3.6.2.Final.jar\n/usr/lib/hadoop-hdfs/lib/protobuf-java-2.5.0.jar\n/usr/lib/hadoop-hdfs/lib/servlet-api-2.5.jar\n/usr/lib/hadoop-hdfs/lib/xmlenc-0.52.jar\n/usr/lib/hadoop-mapreduce/hadoop-archives-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-datajoin-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-distcp-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-extras-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-gridmix-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-app-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-common-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-core-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-hs-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-hs-plugins-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-jobclient-2.2.0.2.0.6.0-76-tests.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-jobclient-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-shuffle-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-examples-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-rumen-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-streaming-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/lib/aopalliance-1.0.jar\n/usr/lib/hadoop-mapreduce/lib/asm-3.2.jar\n/usr/lib/hadoop-mapreduce/lib/avro-1.7.4.jar\n/usr/lib/hadoop-mapreduce/lib/commons-compress-1.4.1.jar\n/usr/lib/hadoop-mapreduce/lib/commons-io-2.1.jar\n/usr/lib/hadoop-mapreduce/lib/guice-3.0.jar\n/usr/lib/hadoop-mapreduce/lib/guice-servlet-3.0.jar\n/usr/lib/hadoop-mapreduce/lib/hamcrest-core-1.1.jar\n/usr/lib/hadoop-mapreduce/lib/jackson-core-asl-1.8.8.jar\n/usr/lib/hadoop-mapreduce/lib/jackson-mapper-asl-1.8.8.jar\n/usr/lib/hadoop-mapreduce/lib/javax.inject-1.jar\n/usr/lib/hadoop-mapreduce/lib/jersey-core-1.9.jar\n/usr/lib/hadoop-mapreduce/lib/jersey-guice-1.9.jar\n/usr/lib/hadoop-mapreduce/lib/jersey-server-1.9.jar\n/usr/lib/hadoop-mapreduce/lib/junit-4.10.jar\n/usr/lib/hadoop-mapreduce/lib/log4j-1.2.17.jar\n/usr/lib/hadoop-mapreduce/lib/netty-3.6.2.Final.jar\n/usr/lib/hadoop-mapreduce/lib/paranamer-2.3.jar\n/usr/lib/hadoop-mapreduce/lib/protobuf-java-2.5.0.jar\n/usr/lib/hadoop-mapreduce/lib/snappy-java-1.0.4.1.jar\n/usr/lib/hadoop-mapreduce/lib/xz-1.0.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-api-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-applications-distributedshell-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-applications-unmanaged-am-launcher-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-client-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-common-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-server-common-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-server-nodemanager-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-server-resourcemanager-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-server-tests-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-server-web-proxy-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-site-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/lib/aopalliance-1.0.jar\n/usr/lib/hadoop-yarn/lib/asm-3.2.jar\n/usr/lib/hadoop-yarn/lib/avro-1.7.4.jar\n/usr/lib/hadoop-yarn/lib/commons-compress-1.4.1.jar\n/usr/lib/hadoop-yarn/lib/commons-io-2.1.jar\n/usr/lib/hadoop-yarn/lib/guice-3.0.jar\n/usr/lib/hadoop-yarn/lib/guice-servlet-3.0.jar\n/usr/lib/hadoop-yarn/lib/hamcrest-core-1.1.jar\n/usr/lib/hadoop-yarn/lib/jackson-core-asl-1.8.8.jar\n/usr/lib/hadoop-yarn/lib/jackson-mapper-asl-1.8.8.jar\n/usr/lib/hadoop-yarn/lib/javax.inject-1.jar\n/usr/lib/hadoop-yarn/lib/jersey-core-1.9.jar\n/usr/lib/hadoop-yarn/lib/jersey-guice-1.9.jar\n/usr/lib/hadoop-yarn/lib/jersey-server-1.9.jar\n/usr/lib/hadoop-yarn/lib/junit-4.10.jar\n/usr/lib/hadoop-yarn/lib/log4j-1.2.17.jar\n/usr/lib/hadoop-yarn/lib/netty-3.6.2.Final.jar\n/usr/lib/hadoop-yarn/lib/paranamer-2.3.jar\n/usr/lib/hadoop-yarn/lib/protobuf-java-2.5.0.jar\n/usr/lib/hadoop-yarn/lib/snappy-java-1.0.4.1.jar\n/usr/lib/hadoop-yarn/lib/xz-1.0.jar\n/usr/lib/hadoop/hadoop-annotations-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop/hadoop-auth-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop/hadoop-common-2.2.0.2.0.6.0-76-tests.jar\n/usr/lib/hadoop/hadoop-common-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop/hadoop-nfs-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop/lib/activation-1.1.jar\n/usr/lib/hadoop/lib/asm-3.2.jar\n/usr/lib/hadoop/lib/avro-1.7.4.jar\n/usr/lib/hadoop/lib/commons-beanutils-1.7.0.jar\n/usr/lib/hadoop/lib/commons-beanutils-core-1.8.0.jar\n/usr/lib/hadoop/lib/commons-cli-1.2.jar\n/usr/lib/hadoop/lib/commons-codec-1.4.jar\n/usr/lib/hadoop/lib/commons-collections-3.2.1.jar\n/usr/lib/hadoop/lib/commons-compress-1.4.1.jar\n/usr/lib/hadoop/lib/commons-configuration-1.6.jar\n/usr/lib/hadoop/lib/commons-digester-1.8.jar\n/usr/lib/hadoop/lib/commons-el-1.0.jar\n/usr/lib/hadoop/lib/commons-httpclient-3.1.jar\n/usr/lib/hadoop/lib/commons-io-2.1.jar\n/usr/lib/hadoop/lib/commons-lang-2.5.jar\n/usr/lib/hadoop/lib/commons-logging-1.1.1.jar\n/usr/lib/hadoop/lib/commons-math-2.1.jar\n/usr/lib/hadoop/lib/commons-net-3.1.jar\n/usr/lib/hadoop/lib/guava-11.0.2.jar\n/usr/lib/hadoop/lib/jackson-core-asl-1.8.8.jar\n/usr/lib/hadoop/lib/jackson-jaxrs-1.8.8.jar\n/usr/lib/hadoop/lib/jackson-mapper-asl-1.8.8.jar\n/usr/lib/hadoop/lib/jackson-xc-1.8.8.jar\n/usr/lib/hadoop/lib/jasper-compiler-5.5.23.jar\n/usr/lib/hadoop/lib/jasper-runtime-5.5.23.jar\n/usr/lib/hadoop/lib/jaxb-api-2.2.2.jar\n/usr/lib/hadoop/lib/jaxb-impl-2.2.3-1.jar\n/usr/lib/hadoop/lib/jersey-core-1.9.jar\n/usr/lib/hadoop/lib/jersey-json-1.9.jar\n/usr/lib/hadoop/lib/jersey-server-1.9.jar\n/usr/lib/hadoop/lib/jets3t-0.6.1.jar\n/usr/lib/hadoop/lib/jettison-1.1.jar\n/usr/lib/hadoop/lib/jetty-6.1.26.jar\n/usr/lib/hadoop/lib/jetty-util-6.1.26.jar\n/usr/lib/hadoop/lib/jsch-0.1.42.jar\n/usr/lib/hadoop/lib/jsp-api-2.1.jar\n/usr/lib/hadoop/lib/jsr305-1.3.9.jar\n/usr/lib/hadoop/lib/junit-4.8.2.jar\n/usr/lib/hadoop/lib/log4j-1.2.17.jar\n/usr/lib/hadoop/lib/mockito-all-1.8.5.jar\n/usr/lib/hadoop/lib/native/*\n/usr/lib/hadoop/lib/netty-3.6.2.Final.jar\n/usr/lib/hadoop/lib/paranamer-2.3.jar\n/usr/lib/hadoop/lib/protobuf-java-2.5.0.jar\n/usr/lib/hadoop/lib/servlet-api-2.5.jar\n/usr/lib/hadoop/lib/slf4j-api-1.7.5.jar\n/usr/lib/hadoop/lib/slf4j-log4j12-1.7.5.jar\n/usr/lib/hadoop/lib/snappy-java-1.0.4.1.jar\n/usr/lib/hadoop/lib/stax-api-1.0.1.jar\n/usr/lib/hadoop/lib/xmlenc-0.52.jar\n/usr/lib/hadoop/lib/xz-1.0.jar\n/usr/lib/hadoop/lib/zookeeper-3.4.5.jar\n{code}","created":"2014-06-05T01:04:01.915+0000"},{"body":"I can reproduce this depending on the size of the dataset:\n\n{noformat}\nspark-submit mllib-movielens-evaluation-assembly-1.0.jar --master spark://mllib1:7077\n--class com.example.MovieLensALS --rank 10 --numIterations 20 --lambda 1.0 --kryo\nhdfs:/movielens/oversampled.dat\n{noformat}\n\nThe exception will not be thrown for small datasets. It will successfully run with MovieLens 100k and 10M. However, when I run it on a 100M dataset, the exception will be thrown.\n\nMy MovieLensALS is mostly the same as the one shipped with Spark. I just added cross-validation. Rating is registered in Kryo just as in the stock example.\n\n{noformat}\n# cat RELEASE \nSpark 1.0.0 built for Hadoop 2.2.0\n{noformat}\n\n","created":"2014-06-06T07:53:15.539+0000"},{"body":"Update: I also reproduce similar error message for a larger data set (~ 3GB).","created":"2014-06-06T16:48:38.009+0000"},{"body":"[~smolav] and [~coderxiang]:\n\nThanks for testing it! Could you post the exact error message you got with stack trace? Based on your description, it should be caused by the default serialization of kryo. It may treat BitSet as a general Java collection, then run into error in ser/de.","created":"2014-06-06T22:28:35.054+0000"},{"body":"Xiangrui Meng, I can't reproduce it at the moment. It takes a quite big dataset to reproduce and I have my machines busy. But I'm pretty sure the stacktrace is exactly the same as the one posted by Neville Li. My bet is that this will be fixed with next Twitter Chill release: https://github.com/twitter/chill/commit/b47512c2c75b94b7c5945985306fa303576bf90d","created":"2014-06-09T09:36:02.620+0000"},{"body":"I think now I understand when it happens. We use storage level MEMORY_AND_DISK for user/product in/out links, which contains BitSet objects. If the dataset is large, these RDDs will be pushed from in memory storage to on disk storage, where the latter requires serialization. So the easiest way to re-produce this error is changing the storage level of inLinks/outLinks to DISK_ONLY and run with kryo.\n\n[~neville] Instead of mapping mutable.BitSet to immutable.BitSet, which introduces overhead, we can register mutable.BitSet in our MovieLensALS example code and wait for the next Chill release. Does it sound good to you?","created":"2014-07-07T18:40:01.202+0000"},{"body":"[~mengxr] sounds good to me.","created":"2014-07-07T18:42:15.919+0000"},{"body":"Do you mind creating a PR registering mutable.BitSet in MovieLensALS.scala and close PR #925? Thanks!","created":"2014-07-07T18:45:33.889+0000"},{"body":"There you go:\nhttps://github.com/apache/spark/pull/1319","created":"2014-07-07T19:10:00.887+0000"},{"body":"Issue resolved by pull request 1319\n[https://github.com/apache/spark/pull/1319]","created":"2014-07-07T22:08:47.089+0000"},{"body":"[~sinisa_lyh]\nSorry to bother you.\nAccording to https://github.com/twitter/chill/pull/185, the twitter.chill have already had the support of mutable BitSet. However, I tried your code, it still doesn't work, if we make kryo as a comment. The task fails in the last line:\n{code}\nserializer.deserialize(bytes).asInstanceOf[OutLinkBlock]\n{code}\nHave you any ideas how it happens? The error information is as follow:\n{code}\n[error] (run-main) com.esotericsoftware.kryo.KryoException: java.lang.IllegalArgumentException: Can not set final scala.collection.mutable.BitSet field OutLinkBlock.elementIds to scala.collection.mutable.HashSet\n[error] Serialization trace:\n[error] elementIds (OutLinkBlock)\ncom.esotericsoftware.kryo.KryoException: java.lang.IllegalArgumentException: Can not set final scala.collection.mutable.BitSet field OutLinkBlock.elementIds to scala.collection.mutable.HashSet\nSerialization trace:\nelementIds (OutLinkBlock)\n\tat com.esotericsoftware.kryo.serializers.FieldSerializer$ObjectField.read(FieldSerializer.java:626)\n\tat com.esotericsoftware.kryo.serializers.FieldSerializer.read(FieldSerializer.java:221)\n\tat com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:729)\n\tat org.apache.spark.serializer.KryoSerializerInstance.deserialize(KryoSerializer.scala:162)\n\tat KroTest$.main(helloworld.scala:25)\n\tat KroTest.main(helloworld.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:622)\nCaused by: java.lang.IllegalArgumentException: Can not set final scala.collection.mutable.BitSet field OutLinkBlock.elementIds to scala.collection.mutable.HashSet\n\tat sun.reflect.UnsafeFieldAccessorImpl.throwSetIllegalArgumentException(UnsafeFieldAccessorImpl.java:164)\n\tat sun.reflect.UnsafeFieldAccessorImpl.throwSetIllegalArgumentException(UnsafeFieldAccessorImpl.java:168)\n\tat sun.reflect.UnsafeQualifiedObjectFieldAccessorImpl.set(UnsafeQualifiedObjectFieldAccessorImpl.java:83)\n\tat java.lang.reflect.Field.set(Field.java:736)\n\tat com.esotericsoftware.kryo.serializers.FieldSerializer$ObjectField.read(FieldSerializer.java:619)\n\tat com.esotericsoftware.kryo.serializers.FieldSerializer.read(FieldSerializer.java:221)\n\tat com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:729)\n\tat org.apache.spark.serializer.KryoSerializerInstance.deserialize(KryoSerializer.scala:162)\n\tat KroTest$.main(helloworld.scala:25)\n\tat KroTest.main(helloworld.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:622)\n\n{code}","created":"2014-10-23T08:26:44.834+0000"},{"body":"In fact, the problem about transformation between HashSet and BitSet in Kyro happens, if we don't register BitSet manually. \n{code}\ncom.esotericsoftware.kryo.KryoException: java.lang.IllegalArgumentException: Can not set final scala.collection.mutable.BitSet field OutLinkBlock.elementIds to scala.collection.mutable.HashSet\n{code}\nThis will also cause the collapse of spark when we use spark HIVE.","created":"2014-10-23T12:37:10.813+0000"},{"body":"Hi all - concise writeup on how to fix this bug here:\nhttp://tbertinmahieux.com/wp/?author=1\n\nAlso related to:\n\nhttp://apache-spark-user-list.1001560.n3.nabble.com/ALS-implicit-error-pyspark-td16595.html\n\nThanks. \n","created":"2014-10-27T16:04:28.969+0000"},{"body":"User 'nevillelyh' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/925","created":"2014-11-14T06:30:08.595+0000"}],"conversations":[{"body":"OutLinkBlock in ALS.scala has an Array[mutable.BitSet] member.\nKryoSerializer uses AllScalaRegistrar from Twitter chill but it doesn't register mutable.BitSet.\n\nRight now we have to register mutable.BitSet manually. A proper fix would be using immutable.BitSet in ALS or register mutable.BitSet in upstream chill.\n\n{code}\nCaused by: org.apache.spark.SparkException: Job aborted due to stage failure: Task 1724.0:9 failed 4 times, most recent failure: Exception failure in TID 68548 on host lon4-hadoopslave-b232.lon4.spotify.net: com.esotericsoftware.kryo.KryoException: java.lang.ArrayStoreException: scala.collection.mutable.HashSet\nSerialization trace:\nshouldSend (org.apache.spark.mllib.recommendation.OutLinkBlock)\n com.esotericsoftware.kryo.serializers.FieldSerializer$ObjectField.read(FieldSerializer.java:626)\n com.esotericsoftware.kryo.serializers.FieldSerializer.read(FieldSerializer.java:221)\n com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:732)\n com.twitter.chill.Tuple2Serializer.read(TupleSerializers.scala:43)\n com.twitter.chill.Tuple2Serializer.read(TupleSerializers.scala:34)\n com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:732)\n org.apache.spark.serializer.KryoDeserializationStream.readObject(KryoSerializer.scala:115)\n org.apache.spark.serializer.DeserializationStream$$anon$1.getNext(Serializer.scala:125)\n org.apache.spark.util.NextIterator.hasNext(NextIterator.scala:71)\n org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:39)\n org.apache.spark.rdd.CoGroupedRDD$$anonfun$compute$4.apply(CoGroupedRDD.scala:155)\n org.apache.spark.rdd.CoGroupedRDD$$anonfun$compute$4.apply(CoGroupedRDD.scala:154)\n scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n org.apache.spark.rdd.CoGroupedRDD.compute(CoGroupedRDD.scala:154)\n org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n org.apache.spark.rdd.RDD.iterator(RDD.scala:229)\n org.apache.spark.rdd.MappedValuesRDD.compute(MappedValuesRDD.scala:31)\n org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n org.apache.spark.rdd.RDD.iterator(RDD.scala:229)\n org.apache.spark.rdd.FlatMappedValuesRDD.compute(FlatMappedValuesRDD.scala:31)\n org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n org.apache.spark.rdd.RDD.iterator(RDD.scala:229)\n org.apache.spark.rdd.FlatMappedRDD.compute(FlatMappedRDD.scala:33)\n org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n org.apache.spark.CacheManager.getOrCompute(CacheManager.scala:77)\n org.apache.spark.rdd.RDD.iterator(RDD.scala:227)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:111)\n org.apache.spark.scheduler.Task.run(Task.scala:51)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:187)\n java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n java.lang.Thread.run(Thread.java:662)\nDriver stacktrace:\n\tat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1033)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1017)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1015)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n\tat org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1015)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:633)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:633)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:633)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor$$anonfun$receive$2.applyOrElse(DAGScheduler.scala:1207)\n\tat akka.actor.ActorCell.receiveMessage(ActorCell.scala:498)\n\tat akka.actor.ActorCell.invoke(ActorCell.scala:456)\n\tat akka.dispatch.Mailbox.processMailbox(Mailbox.scala:237)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:219)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n{code}","from":"reporter","subject":"mutable.BitSet in ALS not serializable with KryoSerializer"},{"body":"I cannot reproduce this error in v1.0.0. There is an example called `MovieLensALS.scala` under `examples/`, which runs fine with kryo enabled. Did you include other dependencies in your application?","from":"developer"},{"body":"Yeah that example worked fine for me in standalone mode but failed in YARN cluster mode with the same error.\nMaybe serialization wasn't needed/triggered in standalone mode?","from":"developer"},{"body":"This is more likely a version conflict in your dependencies. From the Spark WebUI, you can find the system classpath in the environment tab. Please verify that you don't have two different versions of spark, kryo, or any other related library. Classes may hide inside an assembly jar.","from":"developer"},{"body":"We submit 1 spark-assembly and 1 job assembly jar via spark-submit and there are no other obvious scala/spark/kryo jars in the global classpath. I can reproduce the same exception locally with the following snippet, when kryo.register() is commented out.\n\nI just added mutable BitSet to Twitter chill: https://github.com/twitter/chill/pull/185\n\n{code}\nimport com.twitter.chill._\nimport org.apache.spark.serializer.{KryoSerializer, KryoRegistrator}\nimport org.apache.spark.SparkConf\nimport scala.collection.mutable\n\nclass MyRegistrator extends KryoRegistrator {\n override def registerClasses(kryo: Kryo) {\n // kryo.register(classOf[mutable.BitSet])\n }\n}\n\ncase class OutLinkBlock(elementIds: Array[Int], shouldSend: Array[mutable.BitSet])\n\nobject KryoTest {\n def main(args: Array[String]) {\n println(\"hello\")\n val conf = new SparkConf()\n .set(\"spark.serializer\", \"org.apache.spark.serializer.KryoSerializer\")\n .set(\"spark.kryo.registrator\", classOf[MyRegistrator].getName)\n val serializer = new KryoSerializer(conf).newInstance()\n\n val bytes = serializer.serialize(OutLinkBlock(Array(1, 2, 3), Array(mutable.BitSet(2, 4, 6))))\n serializer.deserialize(bytes).asInstanceOf[OutLinkBlock]\n }\n}\n{code}","from":"developer"},{"body":"In our example code, we only register `Rating` and it works. Could you try adding the following:\n\n{code}\nkryo.register(classOf[Rating])\n{code}\n\nI need to reproduce this problem with `ALS.train`.","from":"developer"},{"body":"We are already doing that :)\nOur job works on YARN with \"register(classOf[mutable.BitSet])\". Without it we get the reported exception.","from":"developer"},{"body":"Did you register `Rating`? I think this is necessary.","from":"developer"},{"body":"Yes we did register 'Rating'. And we had to \"register(classOf[mutable.BitSet])\" in addition to make it work.","from":"developer"},{"body":"Hi [~neville], I just run the MovieLens example on my YARN cluster (hadoop-2.0.5-alpha) with kryo enabled and it works. I use the following command:\n\nbin/spark-submit --master yarn-cluster --class org.apache.spark.examples.mllib.MovieLensALS --num-executors ** --driver-memory ** --executor-memory ** --executor-cores 1 spark-examples-1.0.0-hadoop2.0.5-alpha.jar --rank 5 --numIterations 20 --lambda 1.0 --kryo /path/to/sample_movielens_data.txt","from":"developer"},{"body":"Our YARN cluster runs 2.2.0. We built spark-assembly and spark-examples jars with 1.0.0 release source and the bundled make_distribution.sh. And here's my command:\n\n{code}\nspark-submit --master yarn-cluster --class org.apache.spark.examples.mllib.MovieLensALS --num-executors 2 --executor-memory 2g --driver-memory 2g dist/lib/spark-examples-1.0.0-hadoop2.2.0.jar --kryo --implicitPrefs sample_movielens_data.txt\n{code}\n\nHere's a complete list of classpath from the environment tab.\n{code}\n/etc/hadoop/conf\n/usr/lib/hadoop-hdfs/hadoop-hdfs-2.2.0.2.0.6.0-76-tests.jar\n/usr/lib/hadoop-hdfs/hadoop-hdfs-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-hdfs/hadoop-hdfs-nfs-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-hdfs/lib/asm-3.2.jar\n/usr/lib/hadoop-hdfs/lib/commons-cli-1.2.jar\n/usr/lib/hadoop-hdfs/lib/commons-codec-1.4.jar\n/usr/lib/hadoop-hdfs/lib/commons-daemon-1.0.13.jar\n/usr/lib/hadoop-hdfs/lib/commons-el-1.0.jar\n/usr/lib/hadoop-hdfs/lib/commons-io-2.1.jar\n/usr/lib/hadoop-hdfs/lib/commons-lang-2.5.jar\n/usr/lib/hadoop-hdfs/lib/commons-logging-1.1.1.jar\n/usr/lib/hadoop-hdfs/lib/guava-11.0.2.jar\n/usr/lib/hadoop-hdfs/lib/jackson-core-asl-1.8.8.jar\n/usr/lib/hadoop-hdfs/lib/jackson-mapper-asl-1.8.8.jar\n/usr/lib/hadoop-hdfs/lib/jasper-runtime-5.5.23.jar\n/usr/lib/hadoop-hdfs/lib/jersey-core-1.9.jar\n/usr/lib/hadoop-hdfs/lib/jersey-server-1.9.jar\n/usr/lib/hadoop-hdfs/lib/jetty-6.1.26.jar\n/usr/lib/hadoop-hdfs/lib/jetty-util-6.1.26.jar\n/usr/lib/hadoop-hdfs/lib/jsp-api-2.1.jar\n/usr/lib/hadoop-hdfs/lib/jsr305-1.3.9.jar\n/usr/lib/hadoop-hdfs/lib/log4j-1.2.17.jar\n/usr/lib/hadoop-hdfs/lib/netty-3.6.2.Final.jar\n/usr/lib/hadoop-hdfs/lib/protobuf-java-2.5.0.jar\n/usr/lib/hadoop-hdfs/lib/servlet-api-2.5.jar\n/usr/lib/hadoop-hdfs/lib/xmlenc-0.52.jar\n/usr/lib/hadoop-mapreduce/hadoop-archives-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-datajoin-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-distcp-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-extras-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-gridmix-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-app-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-common-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-core-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-hs-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-hs-plugins-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-jobclient-2.2.0.2.0.6.0-76-tests.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-jobclient-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-client-shuffle-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-mapreduce-examples-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-rumen-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/hadoop-streaming-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-mapreduce/lib/aopalliance-1.0.jar\n/usr/lib/hadoop-mapreduce/lib/asm-3.2.jar\n/usr/lib/hadoop-mapreduce/lib/avro-1.7.4.jar\n/usr/lib/hadoop-mapreduce/lib/commons-compress-1.4.1.jar\n/usr/lib/hadoop-mapreduce/lib/commons-io-2.1.jar\n/usr/lib/hadoop-mapreduce/lib/guice-3.0.jar\n/usr/lib/hadoop-mapreduce/lib/guice-servlet-3.0.jar\n/usr/lib/hadoop-mapreduce/lib/hamcrest-core-1.1.jar\n/usr/lib/hadoop-mapreduce/lib/jackson-core-asl-1.8.8.jar\n/usr/lib/hadoop-mapreduce/lib/jackson-mapper-asl-1.8.8.jar\n/usr/lib/hadoop-mapreduce/lib/javax.inject-1.jar\n/usr/lib/hadoop-mapreduce/lib/jersey-core-1.9.jar\n/usr/lib/hadoop-mapreduce/lib/jersey-guice-1.9.jar\n/usr/lib/hadoop-mapreduce/lib/jersey-server-1.9.jar\n/usr/lib/hadoop-mapreduce/lib/junit-4.10.jar\n/usr/lib/hadoop-mapreduce/lib/log4j-1.2.17.jar\n/usr/lib/hadoop-mapreduce/lib/netty-3.6.2.Final.jar\n/usr/lib/hadoop-mapreduce/lib/paranamer-2.3.jar\n/usr/lib/hadoop-mapreduce/lib/protobuf-java-2.5.0.jar\n/usr/lib/hadoop-mapreduce/lib/snappy-java-1.0.4.1.jar\n/usr/lib/hadoop-mapreduce/lib/xz-1.0.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-api-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-applications-distributedshell-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-applications-unmanaged-am-launcher-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-client-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-common-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-server-common-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-server-nodemanager-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-server-resourcemanager-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-server-tests-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-server-web-proxy-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/hadoop-yarn-site-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop-yarn/lib/aopalliance-1.0.jar\n/usr/lib/hadoop-yarn/lib/asm-3.2.jar\n/usr/lib/hadoop-yarn/lib/avro-1.7.4.jar\n/usr/lib/hadoop-yarn/lib/commons-compress-1.4.1.jar\n/usr/lib/hadoop-yarn/lib/commons-io-2.1.jar\n/usr/lib/hadoop-yarn/lib/guice-3.0.jar\n/usr/lib/hadoop-yarn/lib/guice-servlet-3.0.jar\n/usr/lib/hadoop-yarn/lib/hamcrest-core-1.1.jar\n/usr/lib/hadoop-yarn/lib/jackson-core-asl-1.8.8.jar\n/usr/lib/hadoop-yarn/lib/jackson-mapper-asl-1.8.8.jar\n/usr/lib/hadoop-yarn/lib/javax.inject-1.jar\n/usr/lib/hadoop-yarn/lib/jersey-core-1.9.jar\n/usr/lib/hadoop-yarn/lib/jersey-guice-1.9.jar\n/usr/lib/hadoop-yarn/lib/jersey-server-1.9.jar\n/usr/lib/hadoop-yarn/lib/junit-4.10.jar\n/usr/lib/hadoop-yarn/lib/log4j-1.2.17.jar\n/usr/lib/hadoop-yarn/lib/netty-3.6.2.Final.jar\n/usr/lib/hadoop-yarn/lib/paranamer-2.3.jar\n/usr/lib/hadoop-yarn/lib/protobuf-java-2.5.0.jar\n/usr/lib/hadoop-yarn/lib/snappy-java-1.0.4.1.jar\n/usr/lib/hadoop-yarn/lib/xz-1.0.jar\n/usr/lib/hadoop/hadoop-annotations-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop/hadoop-auth-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop/hadoop-common-2.2.0.2.0.6.0-76-tests.jar\n/usr/lib/hadoop/hadoop-common-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop/hadoop-nfs-2.2.0.2.0.6.0-76.jar\n/usr/lib/hadoop/lib/activation-1.1.jar\n/usr/lib/hadoop/lib/asm-3.2.jar\n/usr/lib/hadoop/lib/avro-1.7.4.jar\n/usr/lib/hadoop/lib/commons-beanutils-1.7.0.jar\n/usr/lib/hadoop/lib/commons-beanutils-core-1.8.0.jar\n/usr/lib/hadoop/lib/commons-cli-1.2.jar\n/usr/lib/hadoop/lib/commons-codec-1.4.jar\n/usr/lib/hadoop/lib/commons-collections-3.2.1.jar\n/usr/lib/hadoop/lib/commons-compress-1.4.1.jar\n/usr/lib/hadoop/lib/commons-configuration-1.6.jar\n/usr/lib/hadoop/lib/commons-digester-1.8.jar\n/usr/lib/hadoop/lib/commons-el-1.0.jar\n/usr/lib/hadoop/lib/commons-httpclient-3.1.jar\n/usr/lib/hadoop/lib/commons-io-2.1.jar\n/usr/lib/hadoop/lib/commons-lang-2.5.jar\n/usr/lib/hadoop/lib/commons-logging-1.1.1.jar\n/usr/lib/hadoop/lib/commons-math-2.1.jar\n/usr/lib/hadoop/lib/commons-net-3.1.jar\n/usr/lib/hadoop/lib/guava-11.0.2.jar\n/usr/lib/hadoop/lib/jackson-core-asl-1.8.8.jar\n/usr/lib/hadoop/lib/jackson-jaxrs-1.8.8.jar\n/usr/lib/hadoop/lib/jackson-mapper-asl-1.8.8.jar\n/usr/lib/hadoop/lib/jackson-xc-1.8.8.jar\n/usr/lib/hadoop/lib/jasper-compiler-5.5.23.jar\n/usr/lib/hadoop/lib/jasper-runtime-5.5.23.jar\n/usr/lib/hadoop/lib/jaxb-api-2.2.2.jar\n/usr/lib/hadoop/lib/jaxb-impl-2.2.3-1.jar\n/usr/lib/hadoop/lib/jersey-core-1.9.jar\n/usr/lib/hadoop/lib/jersey-json-1.9.jar\n/usr/lib/hadoop/lib/jersey-server-1.9.jar\n/usr/lib/hadoop/lib/jets3t-0.6.1.jar\n/usr/lib/hadoop/lib/jettison-1.1.jar\n/usr/lib/hadoop/lib/jetty-6.1.26.jar\n/usr/lib/hadoop/lib/jetty-util-6.1.26.jar\n/usr/lib/hadoop/lib/jsch-0.1.42.jar\n/usr/lib/hadoop/lib/jsp-api-2.1.jar\n/usr/lib/hadoop/lib/jsr305-1.3.9.jar\n/usr/lib/hadoop/lib/junit-4.8.2.jar\n/usr/lib/hadoop/lib/log4j-1.2.17.jar\n/usr/lib/hadoop/lib/mockito-all-1.8.5.jar\n/usr/lib/hadoop/lib/native/*\n/usr/lib/hadoop/lib/netty-3.6.2.Final.jar\n/usr/lib/hadoop/lib/paranamer-2.3.jar\n/usr/lib/hadoop/lib/protobuf-java-2.5.0.jar\n/usr/lib/hadoop/lib/servlet-api-2.5.jar\n/usr/lib/hadoop/lib/slf4j-api-1.7.5.jar\n/usr/lib/hadoop/lib/slf4j-log4j12-1.7.5.jar\n/usr/lib/hadoop/lib/snappy-java-1.0.4.1.jar\n/usr/lib/hadoop/lib/stax-api-1.0.1.jar\n/usr/lib/hadoop/lib/xmlenc-0.52.jar\n/usr/lib/hadoop/lib/xz-1.0.jar\n/usr/lib/hadoop/lib/zookeeper-3.4.5.jar\n{code}","from":"developer"},{"body":"I can reproduce this depending on the size of the dataset:\n\n{noformat}\nspark-submit mllib-movielens-evaluation-assembly-1.0.jar --master spark://mllib1:7077\n--class com.example.MovieLensALS --rank 10 --numIterations 20 --lambda 1.0 --kryo\nhdfs:/movielens/oversampled.dat\n{noformat}\n\nThe exception will not be thrown for small datasets. It will successfully run with MovieLens 100k and 10M. However, when I run it on a 100M dataset, the exception will be thrown.\n\nMy MovieLensALS is mostly the same as the one shipped with Spark. I just added cross-validation. Rating is registered in Kryo just as in the stock example.\n\n{noformat}\n# cat RELEASE \nSpark 1.0.0 built for Hadoop 2.2.0\n{noformat}\n\n","from":"developer"},{"body":"Update: I also reproduce similar error message for a larger data set (~ 3GB).","from":"developer"},{"body":"[~smolav] and [~coderxiang]:\n\nThanks for testing it! Could you post the exact error message you got with stack trace? Based on your description, it should be caused by the default serialization of kryo. It may treat BitSet as a general Java collection, then run into error in ser/de.","from":"developer"},{"body":"Xiangrui Meng, I can't reproduce it at the moment. It takes a quite big dataset to reproduce and I have my machines busy. But I'm pretty sure the stacktrace is exactly the same as the one posted by Neville Li. My bet is that this will be fixed with next Twitter Chill release: https://github.com/twitter/chill/commit/b47512c2c75b94b7c5945985306fa303576bf90d","from":"developer"},{"body":"I think now I understand when it happens. We use storage level MEMORY_AND_DISK for user/product in/out links, which contains BitSet objects. If the dataset is large, these RDDs will be pushed from in memory storage to on disk storage, where the latter requires serialization. So the easiest way to re-produce this error is changing the storage level of inLinks/outLinks to DISK_ONLY and run with kryo.\n\n[~neville] Instead of mapping mutable.BitSet to immutable.BitSet, which introduces overhead, we can register mutable.BitSet in our MovieLensALS example code and wait for the next Chill release. Does it sound good to you?","from":"developer"},{"body":"[~mengxr] sounds good to me.","from":"developer"},{"body":"Do you mind creating a PR registering mutable.BitSet in MovieLensALS.scala and close PR #925? Thanks!","from":"developer"},{"body":"There you go:\nhttps://github.com/apache/spark/pull/1319","from":"developer"},{"body":"Issue resolved by pull request 1319\n[https://github.com/apache/spark/pull/1319]","from":"developer"},{"body":"[~sinisa_lyh]\nSorry to bother you.\nAccording to https://github.com/twitter/chill/pull/185, the twitter.chill have already had the support of mutable BitSet. However, I tried your code, it still doesn't work, if we make kryo as a comment. The task fails in the last line:\n{code}\nserializer.deserialize(bytes).asInstanceOf[OutLinkBlock]\n{code}\nHave you any ideas how it happens? The error information is as follow:\n{code}\n[error] (run-main) com.esotericsoftware.kryo.KryoException: java.lang.IllegalArgumentException: Can not set final scala.collection.mutable.BitSet field OutLinkBlock.elementIds to scala.collection.mutable.HashSet\n[error] Serialization trace:\n[error] elementIds (OutLinkBlock)\ncom.esotericsoftware.kryo.KryoException: java.lang.IllegalArgumentException: Can not set final scala.collection.mutable.BitSet field OutLinkBlock.elementIds to scala.collection.mutable.HashSet\nSerialization trace:\nelementIds (OutLinkBlock)\n\tat com.esotericsoftware.kryo.serializers.FieldSerializer$ObjectField.read(FieldSerializer.java:626)\n\tat com.esotericsoftware.kryo.serializers.FieldSerializer.read(FieldSerializer.java:221)\n\tat com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:729)\n\tat org.apache.spark.serializer.KryoSerializerInstance.deserialize(KryoSerializer.scala:162)\n\tat KroTest$.main(helloworld.scala:25)\n\tat KroTest.main(helloworld.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:622)\nCaused by: java.lang.IllegalArgumentException: Can not set final scala.collection.mutable.BitSet field OutLinkBlock.elementIds to scala.collection.mutable.HashSet\n\tat sun.reflect.UnsafeFieldAccessorImpl.throwSetIllegalArgumentException(UnsafeFieldAccessorImpl.java:164)\n\tat sun.reflect.UnsafeFieldAccessorImpl.throwSetIllegalArgumentException(UnsafeFieldAccessorImpl.java:168)\n\tat sun.reflect.UnsafeQualifiedObjectFieldAccessorImpl.set(UnsafeQualifiedObjectFieldAccessorImpl.java:83)\n\tat java.lang.reflect.Field.set(Field.java:736)\n\tat com.esotericsoftware.kryo.serializers.FieldSerializer$ObjectField.read(FieldSerializer.java:619)\n\tat com.esotericsoftware.kryo.serializers.FieldSerializer.read(FieldSerializer.java:221)\n\tat com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:729)\n\tat org.apache.spark.serializer.KryoSerializerInstance.deserialize(KryoSerializer.scala:162)\n\tat KroTest$.main(helloworld.scala:25)\n\tat KroTest.main(helloworld.scala)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:622)\n\n{code}","from":"developer"},{"body":"In fact, the problem about transformation between HashSet and BitSet in Kyro happens, if we don't register BitSet manually. \n{code}\ncom.esotericsoftware.kryo.KryoException: java.lang.IllegalArgumentException: Can not set final scala.collection.mutable.BitSet field OutLinkBlock.elementIds to scala.collection.mutable.HashSet\n{code}\nThis will also cause the collapse of spark when we use spark HIVE.","from":"developer"},{"body":"Hi all - concise writeup on how to fix this bug here:\nhttp://tbertinmahieux.com/wp/?author=1\n\nAlso related to:\n\nhttp://apache-spark-user-list.1001560.n3.nabble.com/ALS-implicit-error-pyspark-td16595.html\n\nThanks. \n","from":"developer"},{"body":"User 'nevillelyh' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/925","from":"developer"}],"created":"2014-05-30T18:58:57.000+0000","description":"OutLinkBlock in ALS.scala has an Array[mutable.BitSet] member.\nKryoSerializer uses AllScalaRegistrar from Twitter chill but it doesn't register mutable.BitSet.\n\nRight now we have to register mutable.BitSet manually. A proper fix would be using immutable.BitSet in ALS or register mutable.BitSet in upstream chill.\n\n{code}\nCaused by: org.apache.spark.SparkException: Job aborted due to stage failure: Task 1724.0:9 failed 4 times, most recent failure: Exception failure in TID 68548 on host lon4-hadoopslave-b232.lon4.spotify.net: com.esotericsoftware.kryo.KryoException: java.lang.ArrayStoreException: scala.collection.mutable.HashSet\nSerialization trace:\nshouldSend (org.apache.spark.mllib.recommendation.OutLinkBlock)\n com.esotericsoftware.kryo.serializers.FieldSerializer$ObjectField.read(FieldSerializer.java:626)\n com.esotericsoftware.kryo.serializers.FieldSerializer.read(FieldSerializer.java:221)\n com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:732)\n com.twitter.chill.Tuple2Serializer.read(TupleSerializers.scala:43)\n com.twitter.chill.Tuple2Serializer.read(TupleSerializers.scala:34)\n com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:732)\n org.apache.spark.serializer.KryoDeserializationStream.readObject(KryoSerializer.scala:115)\n org.apache.spark.serializer.DeserializationStream$$anon$1.getNext(Serializer.scala:125)\n org.apache.spark.util.NextIterator.hasNext(NextIterator.scala:71)\n org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:39)\n org.apache.spark.rdd.CoGroupedRDD$$anonfun$compute$4.apply(CoGroupedRDD.scala:155)\n org.apache.spark.rdd.CoGroupedRDD$$anonfun$compute$4.apply(CoGroupedRDD.scala:154)\n scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n org.apache.spark.rdd.CoGroupedRDD.compute(CoGroupedRDD.scala:154)\n org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n org.apache.spark.rdd.RDD.iterator(RDD.scala:229)\n org.apache.spark.rdd.MappedValuesRDD.compute(MappedValuesRDD.scala:31)\n org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n org.apache.spark.rdd.RDD.iterator(RDD.scala:229)\n org.apache.spark.rdd.FlatMappedValuesRDD.compute(FlatMappedValuesRDD.scala:31)\n org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n org.apache.spark.rdd.RDD.iterator(RDD.scala:229)\n org.apache.spark.rdd.FlatMappedRDD.compute(FlatMappedRDD.scala:33)\n org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n org.apache.spark.CacheManager.getOrCompute(CacheManager.scala:77)\n org.apache.spark.rdd.RDD.iterator(RDD.scala:227)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:111)\n org.apache.spark.scheduler.Task.run(Task.scala:51)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:187)\n java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:886)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:908)\n java.lang.Thread.run(Thread.java:662)\nDriver stacktrace:\n\tat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1033)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1017)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1015)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n\tat org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1015)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:633)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:633)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:633)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor$$anonfun$receive$2.applyOrElse(DAGScheduler.scala:1207)\n\tat akka.actor.ActorCell.receiveMessage(ActorCell.scala:498)\n\tat akka.actor.ActorCell.invoke(ActorCell.scala:456)\n\tat akka.dispatch.Mailbox.processMailbox(Mailbox.scala:237)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:219)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n{code}","issue_id":"12717622","key":"SPARK-1977","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2014-07-07T22:08:47.000+0000","role":"fixed_distractor","summary":"mutable.BitSet in ALS not serializable with KryoSerializer"} {"case_id":"13049391","cluster":"DISTRACTOR-SPARK-19872","comments":[{"body":"This is a regression from spark 2.0.x.","created":"2017-03-08T21:09:11.158+0000"},{"body":"I reverted `rdd.py` and `serializers.py` to the 2.0.2 branch in github and the code above works without an error.\n\nrdd.py link: https://github.com/apache/spark/blob/branch-2.0/python/pyspark/rdd.py\nserializers.py link: https://github.com/apache/spark/blob/branch-2.0/python/pyspark/serializers.py\n\nrdd diff:\n{code}\n--- tmp2.py 2017-03-08 15:17:45.000000000 -0600\n+++ saved2.py 2017-03-08 15:17:59.000000000 -0600\n@@ -52,8 +52,6 @@\n get_used_memory, ExternalSorter, ExternalGroupBy\n from pyspark.traceback_utils import SCCallSiteSync\n\n-from py4j.java_collections import ListConverter, MapConverter\n-\n\n __all__ = [\"RDD\"]\n\n@@ -137,11 +135,12 @@\n break\n if not sock:\n raise Exception(\"could not open socket\")\n- # The RDD materialization time is unpredicable, if we set a timeout for socket reading\n- # operation, it will very possibly fail. See SPARK-18281.\n- sock.settimeout(None)\n- # The socket will be automatically closed when garbage-collected.\n- return serializer.load_stream(sock.makefile(\"rb\", 65536))\n+ try:\n+ rf = sock.makefile(\"rb\", 65536)\n+ for item in serializer.load_stream(rf):\n+ yield item\n+ finally:\n+ sock.close()\n\n\n def ignore_unicode_prefix(f):\n@@ -264,13 +263,44 @@\n\n def isCheckpointed(self):\n \"\"\"\n- Return whether this RDD has been checkpointed or not\n+ Return whether this RDD is checkpointed and materialized, either reliably or locally.\n \"\"\"\n return self._jrdd.rdd().isCheckpointed()\n\n+ def localCheckpoint(self):\n+ \"\"\"\n+ Mark this RDD for local checkpointing using Spark's existing caching layer.\n+\n+ This method is for users who wish to truncate RDD lineages while skipping the expensive\n+ step of replicating the materialized data in a reliable distributed file system. This is\n+ useful for RDDs with long lineages that need to be truncated periodically (e.g. GraphX).\n+\n+ Local checkpointing sacrifices fault-tolerance for performance. In particular, checkpointed\n+ data is written to ephemeral local storage in the executors instead of to a reliable,\n+ fault-tolerant storage. The effect is that if an executor fails during the computation,\n+ the checkpointed data may no longer be accessible, causing an irrecoverable job failure.\n+\n+ This is NOT safe to use with dynamic allocation, which removes executors along\n+ with their cached blocks. If you must use both features, you are advised to set\n+ L{spark.dynamicAllocation.cachedExecutorIdleTimeout} to a high value.\n+\n+ The checkpoint directory set through L{SparkContext.setCheckpointDir()} is not used.\n+ \"\"\"\n+ self._jrdd.rdd().localCheckpoint()\n+\n+ def isLocallyCheckpointed(self):\n+ \"\"\"\n+ Return whether this RDD is marked for local checkpointing.\n+\n+ Exposed for testing.\n+ \"\"\"\n+ return self._jrdd.rdd().isLocallyCheckpointed()\n+\n def getCheckpointFile(self):\n \"\"\"\n Gets the name of the file to which this RDD was checkpointed\n+\n+ Not defined if RDD is checkpointed locally.\n \"\"\"\n checkpointFile = self._jrdd.rdd().getCheckpointFile()\n if checkpointFile.isDefined():\n@@ -387,6 +417,9 @@\n with replacement: expected number of times each element is chosen; fraction must be >= 0\n :param seed: seed for the random number generator\n\n+ .. note:: This is not guaranteed to provide exactly the fraction specified of the total\n+ count of the given :class:`DataFrame`.\n+\n >>> rdd = sc.parallelize(range(100), 4)\n >>> 6 <= rdd.sample(False, 0.1, 81).count() <= 14\n True\n@@ -425,8 +458,8 @@\n \"\"\"\n Return a fixed-size sampled subset of this RDD.\n\n- Note that this method should only be used if the resulting array is expected\n- to be small, as all the data is loaded into the driver's memory.\n+ .. note:: This method should only be used if the resulting array is expected\n+ to be small, as all the data is loaded into the driver's memory.\n\n >>> rdd = sc.parallelize(range(0, 10))\n >>> len(rdd.takeSample(True, 20, 1))\n@@ -537,7 +570,7 @@\n Return the intersection of this RDD and another one. The output will\n not contain any duplicate elements, even if the input RDDs did.\n\n- Note that this method performs a shuffle internally.\n+ .. note:: This method performs a shuffle internally.\n\n >>> rdd1 = sc.parallelize([1, 10, 2, 3, 4, 5])\n >>> rdd2 = sc.parallelize([1, 6, 2, 3, 7, 8])\n@@ -768,8 +801,9 @@\n def collect(self):\n \"\"\"\n Return a list that contains all of the elements in this RDD.\n- Note that this method should only be used if the resulting array is expected\n- to be small, as all the data is loaded into the driver's memory.\n+\n+ .. note:: This method should only be used if the resulting array is expected\n+ to be small, as all the data is loaded into the driver's memory.\n \"\"\"\n with SCCallSiteSync(self.context) as css:\n port = self.ctx._jvm.PythonRDD.collectAndServe(self._jrdd.rdd())\n@@ -1214,12 +1248,12 @@\n\n def top(self, num, key=None):\n \"\"\"\n- Get the top N elements from a RDD.\n+ Get the top N elements from an RDD.\n\n- Note that this method should only be used if the resulting array is expected\n- to be small, as all the data is loaded into the driver's memory.\n+ .. note:: This method should only be used if the resulting array is expected\n+ to be small, as all the data is loaded into the driver's memory.\n\n- Note: It returns the list sorted in descending order.\n+ .. note:: It returns the list sorted in descending order.\n\n >>> sc.parallelize([10, 4, 2, 12, 3]).top(1)\n [12]\n@@ -1238,11 +1272,11 @@\n\n def takeOrdered(self, num, key=None):\n \"\"\"\n- Get the N elements from a RDD ordered in ascending order or as\n+ Get the N elements from an RDD ordered in ascending order or as\n specified by the optional key function.\n\n- Note that this method should only be used if the resulting array is expected\n- to be small, as all the data is loaded into the driver's memory.\n+ .. note:: this method should only be used if the resulting array is expected\n+ to be small, as all the data is loaded into the driver's memory.\n\n >>> sc.parallelize([10, 1, 2, 9, 3, 4, 5, 6, 7]).takeOrdered(6)\n [1, 2, 3, 4, 5, 6]\n@@ -1263,11 +1297,11 @@\n that partition to estimate the number of additional partitions needed\n to satisfy the limit.\n\n- Note that this method should only be used if the resulting array is expected\n- to be small, as all the data is loaded into the driver's memory.\n-\n Translated from the Scala implementation in RDD#take().\n\n+ .. note:: this method should only be used if the resulting array is expected\n+ to be small, as all the data is loaded into the driver's memory.\n+\n >>> sc.parallelize([2, 3, 4, 5, 6]).cache().take(2)\n [2, 3]\n >>> sc.parallelize([2, 3, 4, 5, 6]).take(10)\n@@ -1331,8 +1365,9 @@\n\n def isEmpty(self):\n \"\"\"\n- Returns true if and only if the RDD contains no elements at all. Note that an RDD\n- may be empty even when it has at least 1 partition.\n+ Returns true if and only if the RDD contains no elements at all.\n+\n+ .. note:: an RDD may be empty even when it has at least 1 partition.\n\n >>> sc.parallelize([]).isEmpty()\n True\n@@ -1523,8 +1558,8 @@\n \"\"\"\n Return the key-value pairs in this RDD to the master as a dictionary.\n\n- Note that this method should only be used if the resulting data is expected\n- to be small, as all the data is loaded into the driver's memory.\n+ .. note:: this method should only be used if the resulting data is expected\n+ to be small, as all the data is loaded into the driver's memory.\n\n >>> m = sc.parallelize([(1, 2), (3, 4)]).collectAsMap()\n >>> m[1]\n@@ -1761,8 +1796,7 @@\n set of aggregation functions.\n\n Turns an RDD[(K, V)] into a result of type RDD[(K, C)], for a \"combined\n- type\" C. Note that V and C can be different -- for example, one might\n- group an RDD of type (Int, Int) into an RDD of type (Int, List[Int]).\n+ type\" C.\n\n Users provide three functions:\n\n@@ -1774,6 +1808,9 @@\n\n In addition, users can control the partitioning of the output RDD.\n\n+ .. note:: V and C can be different -- for example, one might group an RDD of type\n+ (Int, Int) into an RDD of type (Int, List[Int]).\n+\n >>> x = sc.parallelize([(\"a\", 1), (\"b\", 1), (\"a\", 1)])\n >>> def add(a, b): return a + str(b)\n >>> sorted(x.combineByKey(str, add, add).collect())\n@@ -1845,9 +1882,9 @@\n Group the values for each key in the RDD into a single sequence.\n Hash-partitions the resulting RDD with numPartitions partitions.\n\n- Note: If you are grouping in order to perform an aggregation (such as a\n- sum or average) over each key, using reduceByKey or aggregateByKey will\n- provide much better performance.\n+ .. note:: If you are grouping in order to perform an aggregation (such as a\n+ sum or average) over each key, using reduceByKey or aggregateByKey will\n+ provide much better performance.\n\n >>> rdd = sc.parallelize([(\"a\", 1), (\"b\", 1), (\"a\", 1)])\n >>> sorted(rdd.groupByKey().mapValues(len).collect())\n@@ -2018,8 +2055,7 @@\n >>> len(rdd.repartition(10).glom().collect())\n 10\n \"\"\"\n- jrdd = self._jrdd.repartition(numPartitions)\n- return RDD(jrdd, self.ctx, self._jrdd_deserializer)\n+ return self.coalesce(numPartitions, shuffle=True)\n\n def coalesce(self, numPartitions, shuffle=False):\n \"\"\"\n@@ -2030,7 +2066,15 @@\n >>> sc.parallelize([1, 2, 3, 4, 5], 3).coalesce(1).glom().collect()\n [[1, 2, 3, 4, 5]]\n \"\"\"\n- jrdd = self._jrdd.coalesce(numPartitions, shuffle)\n+ if shuffle:\n+ # Decrease the batch size in order to distribute evenly the elements across output\n+ # partitions. Otherwise, repartition will possibly produce highly skewed partitions.\n+ batchSize = min(10, self.ctx._batchSize or 1024)\n+ ser = BatchedSerializer(PickleSerializer(), batchSize)\n+ selfCopy = self._reserialize(ser)\n+ jrdd = selfCopy._jrdd.coalesce(numPartitions, shuffle)\n+ else:\n+ jrdd = self._jrdd.coalesce(numPartitions, shuffle)\n return RDD(jrdd, self.ctx, self._jrdd_deserializer)\n\n def zip(self, other):\n@@ -2316,16 +2360,9 @@\n # The broadcast will have same life cycle as created PythonRDD\n broadcast = sc.broadcast(pickled_command)\n pickled_command = ser.dumps(broadcast)\n- # There is a bug in py4j.java_gateway.JavaClass with auto_convert\n- # https://github.com/bartdag/py4j/issues/161\n- # TODO: use auto_convert once py4j fix the bug\n- broadcast_vars = ListConverter().convert(\n- [x._jbroadcast for x in sc._pickled_broadcast_vars],\n- sc._gateway._gateway_client)\n+ broadcast_vars = [x._jbroadcast for x in sc._pickled_broadcast_vars]\n sc._pickled_broadcast_vars.clear()\n- env = MapConverter().convert(sc.environment, sc._gateway._gateway_client)\n- includes = ListConverter().convert(sc._python_includes, sc._gateway._gateway_client)\n- return pickled_command, broadcast_vars, env, includes\n+ return pickled_command, broadcast_vars, sc.environment, sc._python_includes\n\n\n def _wrap_function(sc, func, deserializer, serializer, profiler=None):\n@@ -2433,4 +2470,4 @@\n\n\n if __name__ == \"__main__\":\n- _test()\n\\ No newline at end of file\n+ _test()\n{code}\n\nserializers diff:\n{code}\n--- tmp.py 2017-03-08 15:13:45.000000000 -0600\n+++ /lib/python2.7/site-packages/pyspark/serializers.py 2017-03-08 15:13:03.000000000 -0600\n@@ -61,7 +61,7 @@\n if sys.version < '3':\n import cPickle as pickle\n protocol = 2\n- from itertools import izip as zip\n+ from itertools import izip as zip, imap as map\n else:\n import pickle\n protocol = 3\n@@ -96,7 +96,12 @@\n raise NotImplementedError\n\n def _load_stream_without_unbatching(self, stream):\n- return self.load_stream(stream)\n+ \"\"\"\n+ Return an iterator of deserialized batches (lists) of objects from the input stream.\n+ if the serializer does not operate on batches the default implementation returns an\n+ iterator of single element lists.\n+ \"\"\"\n+ return map(lambda x: [x], self.load_stream(stream))\n\n # Note: our notion of \"equality\" is that output generated by\n # equal serializers can be deserialized using the same serializer.\n@@ -278,50 +283,57 @@\n return \"AutoBatchedSerializer(%s)\" % self.serializer\n\n\n-class CartesianDeserializer(FramedSerializer):\n+class CartesianDeserializer(Serializer):\n\n \"\"\"\n Deserializes the JavaRDD cartesian() of two PythonRDDs.\n+ Due to pyspark batching we cannot simply use the result of the Java RDD cartesian,\n+ we additionally need to do the cartesian within each pair of batches.\n \"\"\"\n\n def __init__(self, key_ser, val_ser):\n- FramedSerializer.__init__(self)\n self.key_ser = key_ser\n self.val_ser = val_ser\n\n- def prepare_keys_values(self, stream):\n- key_stream = self.key_ser._load_stream_without_unbatching(stream)\n- val_stream = self.val_ser._load_stream_without_unbatching(stream)\n- key_is_batched = isinstance(self.key_ser, BatchedSerializer)\n- val_is_batched = isinstance(self.val_ser, BatchedSerializer)\n- for (keys, vals) in zip(key_stream, val_stream):\n- keys = keys if key_is_batched else [keys]\n- vals = vals if val_is_batched else [vals]\n- yield (keys, vals)\n+ def _load_stream_without_unbatching(self, stream):\n+ key_batch_stream = self.key_ser._load_stream_without_unbatching(stream)\n+ val_batch_stream = self.val_ser._load_stream_without_unbatching(stream)\n+ for (key_batch, val_batch) in zip(key_batch_stream, val_batch_stream):\n+ # for correctness with repeated cartesian/zip this must be returned as one batch\n+ yield product(key_batch, val_batch)\n\n def load_stream(self, stream):\n- for (keys, vals) in self.prepare_keys_values(stream):\n- for pair in product(keys, vals):\n- yield pair\n+ return chain.from_iterable(self._load_stream_without_unbatching(stream))\n\n def __repr__(self):\n return \"CartesianDeserializer(%s, %s)\" % \\\n (str(self.key_ser), str(self.val_ser))\n\n\n-class PairDeserializer(CartesianDeserializer):\n+class PairDeserializer(Serializer):\n\n \"\"\"\n Deserializes the JavaRDD zip() of two PythonRDDs.\n+ Due to pyspark batching we cannot simply use the result of the Java RDD zip,\n+ we additionally need to do the zip within each pair of batches.\n \"\"\"\n\n+ def __init__(self, key_ser, val_ser):\n+ self.key_ser = key_ser\n+ self.val_ser = val_ser\n+\n+ def _load_stream_without_unbatching(self, stream):\n+ key_batch_stream = self.key_ser._load_stream_without_unbatching(stream)\n+ val_batch_stream = self.val_ser._load_stream_without_unbatching(stream)\n+ for (key_batch, val_batch) in zip(key_batch_stream, val_batch_stream):\n+ if len(key_batch) != len(val_batch):\n+ raise ValueError(\"Can not deserialize PairRDD with different number of items\"\n+ \" in batches: (%d, %d)\" % (len(key_batch), len(val_batch)))\n+ # for correctness with repeated cartesian/zip this must be returned as one batch\n+ yield zip(key_batch, val_batch)\n+\n def load_stream(self, stream):\n- for (keys, vals) in self.prepare_keys_values(stream):\n- if len(keys) != len(vals):\n- raise ValueError(\"Can not deserialize RDD with different number of items\"\n- \" in pair: (%d, %d)\" % (len(keys), len(vals)))\n- for pair in zip(keys, vals):\n- yield pair\n+ return chain.from_iterable(self._load_stream_without_unbatching(stream))\n\n def __repr__(self):\n return \"PairDeserializer(%s, %s)\" % (str(self.key_ser), str(self.val_ser))\n@@ -378,6 +390,16 @@\n _old_namedtuple = _copy_func(collections.namedtuple)\n\n def namedtuple(*args, **kwargs):\n+ import sys\n+ if sys.version.startswith('3'):\n+ defaults = {\n+ 'verbose': False,\n+ 'rename': False,\n+ 'module': None,\n+ }\n+ for key, value in defaults.items():\n+ if key not in kwargs:\n+ kwargs[key] = value\n cls = _old_namedtuple(*args, **kwargs)\n return _hack_namedtuple(cls)\n\n@@ -559,4 +581,4 @@\n import doctest\n (failure_count, test_count) = doctest.testmod()\n if failure_count:\n- exit(-1)\n\\ No newline at end of file\n+ exit(-1)\n{code}","created":"2017-03-08T21:23:43.262+0000"},{"body":"Using the Spark 2.1.0 serializers.py and the Spark 2.0.2 rdd.py, the code runs.","created":"2017-03-08T21:45:54.361+0000"},{"body":"(Blocker is for committers to determine)","created":"2017-03-09T08:42:57.379+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17282","created":"2017-03-13T19:14:03.382+0000"},{"body":"Wondering if a test will be added to prevent future regressions.","created":"2017-03-15T21:00:42.999+0000"},{"body":"Yup, test was added.","created":"2017-03-15T22:45:23.232+0000"}],"conversations":[{"body":"I'm receiving the following traceback:\n\n{code}\n>>> sc.textFile('test.txt').repartition(10).collect()\nTraceback (most recent call last):\n File \"\", line 1, in \n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/rdd.py\", line 810, in collect\n return list(_load_from_socket(port, self._jrdd_deserializer))\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/rdd.py\", line 140, in _load_from_socket\n for item in serializer.load_stream(rf):\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/serializers.py\", line 539, in load_stream\n yield self.loads(stream)\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/serializers.py\", line 534, in loads\n return s.decode(\"utf-8\") if self.use_unicode else s\n File \"/Users/brianbruggeman/.envs/dg/lib/python2.7/encodings/utf_8.py\", line 16, in decode\n return codecs.utf_8_decode(input, errors, True)\nUnicodeDecodeError: 'utf8' codec can't decode byte 0x80 in position 0: invalid start byte\n{code}\n\nI created a textfile (text.txt) with standard linux newlines:\n{code}\na\nb\n\nd\ne\nf\ng\nh\ni\nj\nk\nl\n\n{code}\n\nI think ran pyspark:\n{code}\n$ pyspark\nPython 2.7.13 (default, Dec 18 2016, 07:03:39)\n[GCC 4.2.1 Compatible Apple LLVM 8.0.0 (clang-800.0.42.1)] on darwin\nType \"help\", \"copyright\", \"credits\" or \"license\" for more information.\nUsing Spark's default log4j profile: org/apache/spark/log4j-defaults.properties\nSetting default log level to \"WARN\".\nTo adjust logging level use sc.setLogLevel(newLevel). For SparkR, use setLogLevel(newLevel).\n17/03/08 13:59:27 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable\n17/03/08 13:59:32 WARN ObjectStore: Failed to get database global_temp, returning NoSuchObjectException\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /__ / .__/\\_,_/_/ /_/\\_\\ version 2.1.0\n /_/\n\nUsing Python version 2.7.13 (default, Dec 18 2016 07:03:39)\nSparkSession available as 'spark'.\n>>> sc.textFile('test.txt').collect()\n[u'a', u'b', u'c', u'd', u'e', u'f', u'g', u'h', u'i', u'j', u'k', u'l']\n>>> sc.textFile('test.txt', use_unicode=False).collect()\n['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j', 'k', 'l']\n>>> sc.textFile('test.txt', use_unicode=False).repartition(10).collect()\n['\\x80\\x02]q\\x01(U\\x01aU\\x01bU\\x01cU\\x01dU\\x01eU\\x01fU\\x01ge.', '\\x80\\x02]q\\x01(U\\x01hU\\x01iU\\x01jU\\x01kU\\x01le.']\n>>> sc.textFile('test.txt').repartition(10).collect()\nTraceback (most recent call last):\n File \"\", line 1, in \n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/rdd.py\", line 810, in collect\n return list(_load_from_socket(port, self._jrdd_deserializer))\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/rdd.py\", line 140, in _load_from_socket\n for item in serializer.load_stream(rf):\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/serializers.py\", line 539, in load_stream\n yield self.loads(stream)\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/serializers.py\", line 534, in loads\n return s.decode(\"utf-8\") if self.use_unicode else s\n File \"/Users/brianbruggeman/.envs/dg/lib/python2.7/encodings/utf_8.py\", line 16, in decode\n return codecs.utf_8_decode(input, errors, True)\nUnicodeDecodeError: 'utf8' codec can't decode byte 0x80 in position 0: invalid start byte\n{code}\n\nThis really looks like a bug in the `serializers.py` code.","from":"reporter","subject":"UnicodeDecodeError in Pyspark on sc.textFile read with repartition"},{"body":"This is a regression from spark 2.0.x.","from":"developer"},{"body":"I reverted `rdd.py` and `serializers.py` to the 2.0.2 branch in github and the code above works without an error.\n\nrdd.py link: https://github.com/apache/spark/blob/branch-2.0/python/pyspark/rdd.py\nserializers.py link: https://github.com/apache/spark/blob/branch-2.0/python/pyspark/serializers.py\n\nrdd diff:\n{code}\n--- tmp2.py 2017-03-08 15:17:45.000000000 -0600\n+++ saved2.py 2017-03-08 15:17:59.000000000 -0600\n@@ -52,8 +52,6 @@\n get_used_memory, ExternalSorter, ExternalGroupBy\n from pyspark.traceback_utils import SCCallSiteSync\n\n-from py4j.java_collections import ListConverter, MapConverter\n-\n\n __all__ = [\"RDD\"]\n\n@@ -137,11 +135,12 @@\n break\n if not sock:\n raise Exception(\"could not open socket\")\n- # The RDD materialization time is unpredicable, if we set a timeout for socket reading\n- # operation, it will very possibly fail. See SPARK-18281.\n- sock.settimeout(None)\n- # The socket will be automatically closed when garbage-collected.\n- return serializer.load_stream(sock.makefile(\"rb\", 65536))\n+ try:\n+ rf = sock.makefile(\"rb\", 65536)\n+ for item in serializer.load_stream(rf):\n+ yield item\n+ finally:\n+ sock.close()\n\n\n def ignore_unicode_prefix(f):\n@@ -264,13 +263,44 @@\n\n def isCheckpointed(self):\n \"\"\"\n- Return whether this RDD has been checkpointed or not\n+ Return whether this RDD is checkpointed and materialized, either reliably or locally.\n \"\"\"\n return self._jrdd.rdd().isCheckpointed()\n\n+ def localCheckpoint(self):\n+ \"\"\"\n+ Mark this RDD for local checkpointing using Spark's existing caching layer.\n+\n+ This method is for users who wish to truncate RDD lineages while skipping the expensive\n+ step of replicating the materialized data in a reliable distributed file system. This is\n+ useful for RDDs with long lineages that need to be truncated periodically (e.g. GraphX).\n+\n+ Local checkpointing sacrifices fault-tolerance for performance. In particular, checkpointed\n+ data is written to ephemeral local storage in the executors instead of to a reliable,\n+ fault-tolerant storage. The effect is that if an executor fails during the computation,\n+ the checkpointed data may no longer be accessible, causing an irrecoverable job failure.\n+\n+ This is NOT safe to use with dynamic allocation, which removes executors along\n+ with their cached blocks. If you must use both features, you are advised to set\n+ L{spark.dynamicAllocation.cachedExecutorIdleTimeout} to a high value.\n+\n+ The checkpoint directory set through L{SparkContext.setCheckpointDir()} is not used.\n+ \"\"\"\n+ self._jrdd.rdd().localCheckpoint()\n+\n+ def isLocallyCheckpointed(self):\n+ \"\"\"\n+ Return whether this RDD is marked for local checkpointing.\n+\n+ Exposed for testing.\n+ \"\"\"\n+ return self._jrdd.rdd().isLocallyCheckpointed()\n+\n def getCheckpointFile(self):\n \"\"\"\n Gets the name of the file to which this RDD was checkpointed\n+\n+ Not defined if RDD is checkpointed locally.\n \"\"\"\n checkpointFile = self._jrdd.rdd().getCheckpointFile()\n if checkpointFile.isDefined():\n@@ -387,6 +417,9 @@\n with replacement: expected number of times each element is chosen; fraction must be >= 0\n :param seed: seed for the random number generator\n\n+ .. note:: This is not guaranteed to provide exactly the fraction specified of the total\n+ count of the given :class:`DataFrame`.\n+\n >>> rdd = sc.parallelize(range(100), 4)\n >>> 6 <= rdd.sample(False, 0.1, 81).count() <= 14\n True\n@@ -425,8 +458,8 @@\n \"\"\"\n Return a fixed-size sampled subset of this RDD.\n\n- Note that this method should only be used if the resulting array is expected\n- to be small, as all the data is loaded into the driver's memory.\n+ .. note:: This method should only be used if the resulting array is expected\n+ to be small, as all the data is loaded into the driver's memory.\n\n >>> rdd = sc.parallelize(range(0, 10))\n >>> len(rdd.takeSample(True, 20, 1))\n@@ -537,7 +570,7 @@\n Return the intersection of this RDD and another one. The output will\n not contain any duplicate elements, even if the input RDDs did.\n\n- Note that this method performs a shuffle internally.\n+ .. note:: This method performs a shuffle internally.\n\n >>> rdd1 = sc.parallelize([1, 10, 2, 3, 4, 5])\n >>> rdd2 = sc.parallelize([1, 6, 2, 3, 7, 8])\n@@ -768,8 +801,9 @@\n def collect(self):\n \"\"\"\n Return a list that contains all of the elements in this RDD.\n- Note that this method should only be used if the resulting array is expected\n- to be small, as all the data is loaded into the driver's memory.\n+\n+ .. note:: This method should only be used if the resulting array is expected\n+ to be small, as all the data is loaded into the driver's memory.\n \"\"\"\n with SCCallSiteSync(self.context) as css:\n port = self.ctx._jvm.PythonRDD.collectAndServe(self._jrdd.rdd())\n@@ -1214,12 +1248,12 @@\n\n def top(self, num, key=None):\n \"\"\"\n- Get the top N elements from a RDD.\n+ Get the top N elements from an RDD.\n\n- Note that this method should only be used if the resulting array is expected\n- to be small, as all the data is loaded into the driver's memory.\n+ .. note:: This method should only be used if the resulting array is expected\n+ to be small, as all the data is loaded into the driver's memory.\n\n- Note: It returns the list sorted in descending order.\n+ .. note:: It returns the list sorted in descending order.\n\n >>> sc.parallelize([10, 4, 2, 12, 3]).top(1)\n [12]\n@@ -1238,11 +1272,11 @@\n\n def takeOrdered(self, num, key=None):\n \"\"\"\n- Get the N elements from a RDD ordered in ascending order or as\n+ Get the N elements from an RDD ordered in ascending order or as\n specified by the optional key function.\n\n- Note that this method should only be used if the resulting array is expected\n- to be small, as all the data is loaded into the driver's memory.\n+ .. note:: this method should only be used if the resulting array is expected\n+ to be small, as all the data is loaded into the driver's memory.\n\n >>> sc.parallelize([10, 1, 2, 9, 3, 4, 5, 6, 7]).takeOrdered(6)\n [1, 2, 3, 4, 5, 6]\n@@ -1263,11 +1297,11 @@\n that partition to estimate the number of additional partitions needed\n to satisfy the limit.\n\n- Note that this method should only be used if the resulting array is expected\n- to be small, as all the data is loaded into the driver's memory.\n-\n Translated from the Scala implementation in RDD#take().\n\n+ .. note:: this method should only be used if the resulting array is expected\n+ to be small, as all the data is loaded into the driver's memory.\n+\n >>> sc.parallelize([2, 3, 4, 5, 6]).cache().take(2)\n [2, 3]\n >>> sc.parallelize([2, 3, 4, 5, 6]).take(10)\n@@ -1331,8 +1365,9 @@\n\n def isEmpty(self):\n \"\"\"\n- Returns true if and only if the RDD contains no elements at all. Note that an RDD\n- may be empty even when it has at least 1 partition.\n+ Returns true if and only if the RDD contains no elements at all.\n+\n+ .. note:: an RDD may be empty even when it has at least 1 partition.\n\n >>> sc.parallelize([]).isEmpty()\n True\n@@ -1523,8 +1558,8 @@\n \"\"\"\n Return the key-value pairs in this RDD to the master as a dictionary.\n\n- Note that this method should only be used if the resulting data is expected\n- to be small, as all the data is loaded into the driver's memory.\n+ .. note:: this method should only be used if the resulting data is expected\n+ to be small, as all the data is loaded into the driver's memory.\n\n >>> m = sc.parallelize([(1, 2), (3, 4)]).collectAsMap()\n >>> m[1]\n@@ -1761,8 +1796,7 @@\n set of aggregation functions.\n\n Turns an RDD[(K, V)] into a result of type RDD[(K, C)], for a \"combined\n- type\" C. Note that V and C can be different -- for example, one might\n- group an RDD of type (Int, Int) into an RDD of type (Int, List[Int]).\n+ type\" C.\n\n Users provide three functions:\n\n@@ -1774,6 +1808,9 @@\n\n In addition, users can control the partitioning of the output RDD.\n\n+ .. note:: V and C can be different -- for example, one might group an RDD of type\n+ (Int, Int) into an RDD of type (Int, List[Int]).\n+\n >>> x = sc.parallelize([(\"a\", 1), (\"b\", 1), (\"a\", 1)])\n >>> def add(a, b): return a + str(b)\n >>> sorted(x.combineByKey(str, add, add).collect())\n@@ -1845,9 +1882,9 @@\n Group the values for each key in the RDD into a single sequence.\n Hash-partitions the resulting RDD with numPartitions partitions.\n\n- Note: If you are grouping in order to perform an aggregation (such as a\n- sum or average) over each key, using reduceByKey or aggregateByKey will\n- provide much better performance.\n+ .. note:: If you are grouping in order to perform an aggregation (such as a\n+ sum or average) over each key, using reduceByKey or aggregateByKey will\n+ provide much better performance.\n\n >>> rdd = sc.parallelize([(\"a\", 1), (\"b\", 1), (\"a\", 1)])\n >>> sorted(rdd.groupByKey().mapValues(len).collect())\n@@ -2018,8 +2055,7 @@\n >>> len(rdd.repartition(10).glom().collect())\n 10\n \"\"\"\n- jrdd = self._jrdd.repartition(numPartitions)\n- return RDD(jrdd, self.ctx, self._jrdd_deserializer)\n+ return self.coalesce(numPartitions, shuffle=True)\n\n def coalesce(self, numPartitions, shuffle=False):\n \"\"\"\n@@ -2030,7 +2066,15 @@\n >>> sc.parallelize([1, 2, 3, 4, 5], 3).coalesce(1).glom().collect()\n [[1, 2, 3, 4, 5]]\n \"\"\"\n- jrdd = self._jrdd.coalesce(numPartitions, shuffle)\n+ if shuffle:\n+ # Decrease the batch size in order to distribute evenly the elements across output\n+ # partitions. Otherwise, repartition will possibly produce highly skewed partitions.\n+ batchSize = min(10, self.ctx._batchSize or 1024)\n+ ser = BatchedSerializer(PickleSerializer(), batchSize)\n+ selfCopy = self._reserialize(ser)\n+ jrdd = selfCopy._jrdd.coalesce(numPartitions, shuffle)\n+ else:\n+ jrdd = self._jrdd.coalesce(numPartitions, shuffle)\n return RDD(jrdd, self.ctx, self._jrdd_deserializer)\n\n def zip(self, other):\n@@ -2316,16 +2360,9 @@\n # The broadcast will have same life cycle as created PythonRDD\n broadcast = sc.broadcast(pickled_command)\n pickled_command = ser.dumps(broadcast)\n- # There is a bug in py4j.java_gateway.JavaClass with auto_convert\n- # https://github.com/bartdag/py4j/issues/161\n- # TODO: use auto_convert once py4j fix the bug\n- broadcast_vars = ListConverter().convert(\n- [x._jbroadcast for x in sc._pickled_broadcast_vars],\n- sc._gateway._gateway_client)\n+ broadcast_vars = [x._jbroadcast for x in sc._pickled_broadcast_vars]\n sc._pickled_broadcast_vars.clear()\n- env = MapConverter().convert(sc.environment, sc._gateway._gateway_client)\n- includes = ListConverter().convert(sc._python_includes, sc._gateway._gateway_client)\n- return pickled_command, broadcast_vars, env, includes\n+ return pickled_command, broadcast_vars, sc.environment, sc._python_includes\n\n\n def _wrap_function(sc, func, deserializer, serializer, profiler=None):\n@@ -2433,4 +2470,4 @@\n\n\n if __name__ == \"__main__\":\n- _test()\n\\ No newline at end of file\n+ _test()\n{code}\n\nserializers diff:\n{code}\n--- tmp.py 2017-03-08 15:13:45.000000000 -0600\n+++ /lib/python2.7/site-packages/pyspark/serializers.py 2017-03-08 15:13:03.000000000 -0600\n@@ -61,7 +61,7 @@\n if sys.version < '3':\n import cPickle as pickle\n protocol = 2\n- from itertools import izip as zip\n+ from itertools import izip as zip, imap as map\n else:\n import pickle\n protocol = 3\n@@ -96,7 +96,12 @@\n raise NotImplementedError\n\n def _load_stream_without_unbatching(self, stream):\n- return self.load_stream(stream)\n+ \"\"\"\n+ Return an iterator of deserialized batches (lists) of objects from the input stream.\n+ if the serializer does not operate on batches the default implementation returns an\n+ iterator of single element lists.\n+ \"\"\"\n+ return map(lambda x: [x], self.load_stream(stream))\n\n # Note: our notion of \"equality\" is that output generated by\n # equal serializers can be deserialized using the same serializer.\n@@ -278,50 +283,57 @@\n return \"AutoBatchedSerializer(%s)\" % self.serializer\n\n\n-class CartesianDeserializer(FramedSerializer):\n+class CartesianDeserializer(Serializer):\n\n \"\"\"\n Deserializes the JavaRDD cartesian() of two PythonRDDs.\n+ Due to pyspark batching we cannot simply use the result of the Java RDD cartesian,\n+ we additionally need to do the cartesian within each pair of batches.\n \"\"\"\n\n def __init__(self, key_ser, val_ser):\n- FramedSerializer.__init__(self)\n self.key_ser = key_ser\n self.val_ser = val_ser\n\n- def prepare_keys_values(self, stream):\n- key_stream = self.key_ser._load_stream_without_unbatching(stream)\n- val_stream = self.val_ser._load_stream_without_unbatching(stream)\n- key_is_batched = isinstance(self.key_ser, BatchedSerializer)\n- val_is_batched = isinstance(self.val_ser, BatchedSerializer)\n- for (keys, vals) in zip(key_stream, val_stream):\n- keys = keys if key_is_batched else [keys]\n- vals = vals if val_is_batched else [vals]\n- yield (keys, vals)\n+ def _load_stream_without_unbatching(self, stream):\n+ key_batch_stream = self.key_ser._load_stream_without_unbatching(stream)\n+ val_batch_stream = self.val_ser._load_stream_without_unbatching(stream)\n+ for (key_batch, val_batch) in zip(key_batch_stream, val_batch_stream):\n+ # for correctness with repeated cartesian/zip this must be returned as one batch\n+ yield product(key_batch, val_batch)\n\n def load_stream(self, stream):\n- for (keys, vals) in self.prepare_keys_values(stream):\n- for pair in product(keys, vals):\n- yield pair\n+ return chain.from_iterable(self._load_stream_without_unbatching(stream))\n\n def __repr__(self):\n return \"CartesianDeserializer(%s, %s)\" % \\\n (str(self.key_ser), str(self.val_ser))\n\n\n-class PairDeserializer(CartesianDeserializer):\n+class PairDeserializer(Serializer):\n\n \"\"\"\n Deserializes the JavaRDD zip() of two PythonRDDs.\n+ Due to pyspark batching we cannot simply use the result of the Java RDD zip,\n+ we additionally need to do the zip within each pair of batches.\n \"\"\"\n\n+ def __init__(self, key_ser, val_ser):\n+ self.key_ser = key_ser\n+ self.val_ser = val_ser\n+\n+ def _load_stream_without_unbatching(self, stream):\n+ key_batch_stream = self.key_ser._load_stream_without_unbatching(stream)\n+ val_batch_stream = self.val_ser._load_stream_without_unbatching(stream)\n+ for (key_batch, val_batch) in zip(key_batch_stream, val_batch_stream):\n+ if len(key_batch) != len(val_batch):\n+ raise ValueError(\"Can not deserialize PairRDD with different number of items\"\n+ \" in batches: (%d, %d)\" % (len(key_batch), len(val_batch)))\n+ # for correctness with repeated cartesian/zip this must be returned as one batch\n+ yield zip(key_batch, val_batch)\n+\n def load_stream(self, stream):\n- for (keys, vals) in self.prepare_keys_values(stream):\n- if len(keys) != len(vals):\n- raise ValueError(\"Can not deserialize RDD with different number of items\"\n- \" in pair: (%d, %d)\" % (len(keys), len(vals)))\n- for pair in zip(keys, vals):\n- yield pair\n+ return chain.from_iterable(self._load_stream_without_unbatching(stream))\n\n def __repr__(self):\n return \"PairDeserializer(%s, %s)\" % (str(self.key_ser), str(self.val_ser))\n@@ -378,6 +390,16 @@\n _old_namedtuple = _copy_func(collections.namedtuple)\n\n def namedtuple(*args, **kwargs):\n+ import sys\n+ if sys.version.startswith('3'):\n+ defaults = {\n+ 'verbose': False,\n+ 'rename': False,\n+ 'module': None,\n+ }\n+ for key, value in defaults.items():\n+ if key not in kwargs:\n+ kwargs[key] = value\n cls = _old_namedtuple(*args, **kwargs)\n return _hack_namedtuple(cls)\n\n@@ -559,4 +581,4 @@\n import doctest\n (failure_count, test_count) = doctest.testmod()\n if failure_count:\n- exit(-1)\n\\ No newline at end of file\n+ exit(-1)\n{code}","from":"developer"},{"body":"Using the Spark 2.1.0 serializers.py and the Spark 2.0.2 rdd.py, the code runs.","from":"developer"},{"body":"(Blocker is for committers to determine)","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17282","from":"developer"},{"body":"Wondering if a test will be added to prevent future regressions.","from":"developer"},{"body":"Yup, test was added.","from":"developer"}],"created":"2017-03-08T20:50:32.000+0000","description":"I'm receiving the following traceback:\n\n{code}\n>>> sc.textFile('test.txt').repartition(10).collect()\nTraceback (most recent call last):\n File \"\", line 1, in \n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/rdd.py\", line 810, in collect\n return list(_load_from_socket(port, self._jrdd_deserializer))\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/rdd.py\", line 140, in _load_from_socket\n for item in serializer.load_stream(rf):\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/serializers.py\", line 539, in load_stream\n yield self.loads(stream)\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/serializers.py\", line 534, in loads\n return s.decode(\"utf-8\") if self.use_unicode else s\n File \"/Users/brianbruggeman/.envs/dg/lib/python2.7/encodings/utf_8.py\", line 16, in decode\n return codecs.utf_8_decode(input, errors, True)\nUnicodeDecodeError: 'utf8' codec can't decode byte 0x80 in position 0: invalid start byte\n{code}\n\nI created a textfile (text.txt) with standard linux newlines:\n{code}\na\nb\n\nd\ne\nf\ng\nh\ni\nj\nk\nl\n\n{code}\n\nI think ran pyspark:\n{code}\n$ pyspark\nPython 2.7.13 (default, Dec 18 2016, 07:03:39)\n[GCC 4.2.1 Compatible Apple LLVM 8.0.0 (clang-800.0.42.1)] on darwin\nType \"help\", \"copyright\", \"credits\" or \"license\" for more information.\nUsing Spark's default log4j profile: org/apache/spark/log4j-defaults.properties\nSetting default log level to \"WARN\".\nTo adjust logging level use sc.setLogLevel(newLevel). For SparkR, use setLogLevel(newLevel).\n17/03/08 13:59:27 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable\n17/03/08 13:59:32 WARN ObjectStore: Failed to get database global_temp, returning NoSuchObjectException\nWelcome to\n ____ __\n / __/__ ___ _____/ /__\n _\\ \\/ _ \\/ _ `/ __/ '_/\n /__ / .__/\\_,_/_/ /_/\\_\\ version 2.1.0\n /_/\n\nUsing Python version 2.7.13 (default, Dec 18 2016 07:03:39)\nSparkSession available as 'spark'.\n>>> sc.textFile('test.txt').collect()\n[u'a', u'b', u'c', u'd', u'e', u'f', u'g', u'h', u'i', u'j', u'k', u'l']\n>>> sc.textFile('test.txt', use_unicode=False).collect()\n['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j', 'k', 'l']\n>>> sc.textFile('test.txt', use_unicode=False).repartition(10).collect()\n['\\x80\\x02]q\\x01(U\\x01aU\\x01bU\\x01cU\\x01dU\\x01eU\\x01fU\\x01ge.', '\\x80\\x02]q\\x01(U\\x01hU\\x01iU\\x01jU\\x01kU\\x01le.']\n>>> sc.textFile('test.txt').repartition(10).collect()\nTraceback (most recent call last):\n File \"\", line 1, in \n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/rdd.py\", line 810, in collect\n return list(_load_from_socket(port, self._jrdd_deserializer))\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/rdd.py\", line 140, in _load_from_socket\n for item in serializer.load_stream(rf):\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/serializers.py\", line 539, in load_stream\n yield self.loads(stream)\n File \"/usr/local/Cellar/apache-spark/2.1.0/libexec/python/pyspark/serializers.py\", line 534, in loads\n return s.decode(\"utf-8\") if self.use_unicode else s\n File \"/Users/brianbruggeman/.envs/dg/lib/python2.7/encodings/utf_8.py\", line 16, in decode\n return codecs.utf_8_decode(input, errors, True)\nUnicodeDecodeError: 'utf8' codec can't decode byte 0x80 in position 0: invalid start byte\n{code}\n\nThis really looks like a bug in the `serializers.py` code.","issue_id":"13049391","key":"SPARK-19872","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-03-15T17:14:32.000+0000","role":"fixed_distractor","summary":"UnicodeDecodeError in Pyspark on sc.textFile read with repartition"} {"case_id":"13049780","cluster":"DISTRACTOR-SPARK-19888","comments":[{"body":"That stacktrace also shows a concurrent modification exception, yes?. See SPARK-19185 for that\n\nSee e.g. SPARK-19680 for background on why offset out of range may occur on executor when it doesn't on driver. Although if you're using reset latest, unless you have really short retention this is kind of surprising.","created":"2017-03-10T21:32:33.277+0000"},{"body":"[~jrmiller] SPARK-19185 resolved on 2.4.0. Can you re-test it please?","created":"2018-12-12T09:05:48.742+0000"},{"body":"Please reopen it if appears again.","created":"2019-02-14T13:58:49.548+0000"},{"body":"Issue solved in SPARK-19185.","created":"2019-02-14T14:00:49.994+0000"}],"conversations":[{"body":"I was told to post this in a Spark ticket from KAFKA-4396:\n\nI've been seeing a curious error with kafka 0.10 (spark 2.11), these may be two separate errors, I'm not sure. What's puzzling is that I'm setting auto.offset.reset to latest and it's still throwing an OffsetOutOfRangeException, behavior that's contrary to the code. Please help! :)\n\n{code}\nval kafkaParams = Map[String, Object](\n \"group.id\" -> consumerGroup,\n \"bootstrap.servers\" -> bootstrapServers,\n \"key.deserializer\" -> classOf[ByteArrayDeserializer],\n \"value.deserializer\" -> classOf[MessageRowDeserializer],\n \"auto.offset.reset\" -> \"latest\",\n \"enable.auto.commit\" -> (false: java.lang.Boolean),\n \"max.poll.records\" -> persisterConfig.maxPollRecords,\n \"request.timeout.ms\" -> persisterConfig.requestTimeoutMs,\n \"session.timeout.ms\" -> persisterConfig.sessionTimeoutMs,\n \"heartbeat.interval.ms\" -> persisterConfig.heartbeatIntervalMs,\n \"connections.max.idle.ms\"-> persisterConfig.connectionsMaxIdleMs\n )\n{code}\n\n{code}\n16/11/09 23:10:17 INFO BlockManagerInfo: Added broadcast_154_piece0 in memory on xyz (size: 146.3 KB, free: 8.4 GB)\n16/11/09 23:10:23 WARN TaskSetManager: Lost task 15.0 in stage 151.0 (TID 38837, xyz): org.apache.kafka.clients.consumer.OffsetOutOfRangeException: Offsets out of range with no configured reset policy for partitions: {topic=231884473}\n at org.apache.kafka.clients.consumer.internals.Fetcher.parseFetchedData(Fetcher.java:588)\n at org.apache.kafka.clients.consumer.internals.Fetcher.fetchedRecords(Fetcher.java:354)\n at org.apache.kafka.clients.consumer.KafkaConsumer.pollOnce(KafkaConsumer.java:1000)\n at org.apache.kafka.clients.consumer.KafkaConsumer.poll(KafkaConsumer.java:938)\n at org.apache.spark.streaming.kafka010.CachedKafkaConsumer.poll(CachedKafkaConsumer.scala:99)\n at org.apache.spark.streaming.kafka010.CachedKafkaConsumer.get(CachedKafkaConsumer.scala:70)\n at org.apache.spark.streaming.kafka010.KafkaRDD$KafkaRDDIterator.next(KafkaRDD.scala:227)\n at org.apache.spark.streaming.kafka010.KafkaRDD$KafkaRDDIterator.next(KafkaRDD.scala:193)\n at scala.collection.Iterator$$anon$12.nextCur(Iterator.scala:434)\n at scala.collection.Iterator$$anon$12.hasNext(Iterator.scala:440)\n at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n at scala.collection.Iterator$$anon$12.hasNext(Iterator.scala:438)\n at org.apache.spark.sql.execution.datasources.DynamicPartitionWriterContainer.writeRows(WriterContainer.scala:397)\n at org.apache.spark.sql.execution.datasources.InsertIntoHadoopFsRelationCommand$$anonfun$run$1$$anonfun$apply$mcV$sp$1.apply(InsertIntoHadoopFsRelationCommand.scala:143)\n at org.apache.spark.sql.execution.datasources.InsertIntoHadoopFsRelationCommand$$anonfun$run$1$$anonfun$apply$mcV$sp$1.apply(InsertIntoHadoopFsRelationCommand.scala:143)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)\n at org.apache.spark.scheduler.Task.run(Task.scala:85)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n\n16/11/09 23:10:29 INFO TaskSetManager: Finished task 10.0 in stage 154.0 (TID 39388) in 12043 ms on xyz (1/16)\n16/11/09 23:10:31 INFO TaskSetManager: Finished task 0.0 in stage 154.0 (TID 39375) in 13444 ms on xyz (2/16)\n16/11/09 23:10:44 WARN TaskSetManager: Lost task 1.0 in stage 151.0 (TID 38843, xyz): java.util.ConcurrentModificationException: KafkaConsumer is not safe for multi-threaded access\n at org.apache.kafka.clients.consumer.KafkaConsumer.acquire(KafkaConsumer.java:1431)\n at org.apache.kafka.clients.consumer.KafkaConsumer.poll(KafkaConsumer.java:929)\n at org.apache.spark.streaming.kafka010.CachedKafkaConsumer.poll(CachedKafkaConsumer.scala:99)\n at org.apache.spark.streaming.kafka010.CachedKafkaConsumer.get(CachedKafkaConsumer.scala:73)\n at org.apache.spark.streaming.kafka010.KafkaRDD$KafkaRDDIterator.next(KafkaRDD.scala:227)\n at org.apache.spark.streaming.kafka010.KafkaRDD$KafkaRDDIterator.next(KafkaRDD.scala:193)\n at scala.collection.Iterator$$anon$12.nextCur(Iterator.scala:434)\n\n{code}","from":"reporter","subject":"Seeing offsets not resetting even when reset policy is configured explicitly"},{"body":"That stacktrace also shows a concurrent modification exception, yes?. See SPARK-19185 for that\n\nSee e.g. SPARK-19680 for background on why offset out of range may occur on executor when it doesn't on driver. Although if you're using reset latest, unless you have really short retention this is kind of surprising.","from":"developer"},{"body":"[~jrmiller] SPARK-19185 resolved on 2.4.0. Can you re-test it please?","from":"developer"},{"body":"Please reopen it if appears again.","from":"developer"},{"body":"Issue solved in SPARK-19185.","from":"developer"}],"created":"2017-03-09T22:33:08.000+0000","description":"I was told to post this in a Spark ticket from KAFKA-4396:\n\nI've been seeing a curious error with kafka 0.10 (spark 2.11), these may be two separate errors, I'm not sure. What's puzzling is that I'm setting auto.offset.reset to latest and it's still throwing an OffsetOutOfRangeException, behavior that's contrary to the code. Please help! :)\n\n{code}\nval kafkaParams = Map[String, Object](\n \"group.id\" -> consumerGroup,\n \"bootstrap.servers\" -> bootstrapServers,\n \"key.deserializer\" -> classOf[ByteArrayDeserializer],\n \"value.deserializer\" -> classOf[MessageRowDeserializer],\n \"auto.offset.reset\" -> \"latest\",\n \"enable.auto.commit\" -> (false: java.lang.Boolean),\n \"max.poll.records\" -> persisterConfig.maxPollRecords,\n \"request.timeout.ms\" -> persisterConfig.requestTimeoutMs,\n \"session.timeout.ms\" -> persisterConfig.sessionTimeoutMs,\n \"heartbeat.interval.ms\" -> persisterConfig.heartbeatIntervalMs,\n \"connections.max.idle.ms\"-> persisterConfig.connectionsMaxIdleMs\n )\n{code}\n\n{code}\n16/11/09 23:10:17 INFO BlockManagerInfo: Added broadcast_154_piece0 in memory on xyz (size: 146.3 KB, free: 8.4 GB)\n16/11/09 23:10:23 WARN TaskSetManager: Lost task 15.0 in stage 151.0 (TID 38837, xyz): org.apache.kafka.clients.consumer.OffsetOutOfRangeException: Offsets out of range with no configured reset policy for partitions: {topic=231884473}\n at org.apache.kafka.clients.consumer.internals.Fetcher.parseFetchedData(Fetcher.java:588)\n at org.apache.kafka.clients.consumer.internals.Fetcher.fetchedRecords(Fetcher.java:354)\n at org.apache.kafka.clients.consumer.KafkaConsumer.pollOnce(KafkaConsumer.java:1000)\n at org.apache.kafka.clients.consumer.KafkaConsumer.poll(KafkaConsumer.java:938)\n at org.apache.spark.streaming.kafka010.CachedKafkaConsumer.poll(CachedKafkaConsumer.scala:99)\n at org.apache.spark.streaming.kafka010.CachedKafkaConsumer.get(CachedKafkaConsumer.scala:70)\n at org.apache.spark.streaming.kafka010.KafkaRDD$KafkaRDDIterator.next(KafkaRDD.scala:227)\n at org.apache.spark.streaming.kafka010.KafkaRDD$KafkaRDDIterator.next(KafkaRDD.scala:193)\n at scala.collection.Iterator$$anon$12.nextCur(Iterator.scala:434)\n at scala.collection.Iterator$$anon$12.hasNext(Iterator.scala:440)\n at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n at scala.collection.Iterator$$anon$12.hasNext(Iterator.scala:438)\n at org.apache.spark.sql.execution.datasources.DynamicPartitionWriterContainer.writeRows(WriterContainer.scala:397)\n at org.apache.spark.sql.execution.datasources.InsertIntoHadoopFsRelationCommand$$anonfun$run$1$$anonfun$apply$mcV$sp$1.apply(InsertIntoHadoopFsRelationCommand.scala:143)\n at org.apache.spark.sql.execution.datasources.InsertIntoHadoopFsRelationCommand$$anonfun$run$1$$anonfun$apply$mcV$sp$1.apply(InsertIntoHadoopFsRelationCommand.scala:143)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)\n at org.apache.spark.scheduler.Task.run(Task.scala:85)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n\n16/11/09 23:10:29 INFO TaskSetManager: Finished task 10.0 in stage 154.0 (TID 39388) in 12043 ms on xyz (1/16)\n16/11/09 23:10:31 INFO TaskSetManager: Finished task 0.0 in stage 154.0 (TID 39375) in 13444 ms on xyz (2/16)\n16/11/09 23:10:44 WARN TaskSetManager: Lost task 1.0 in stage 151.0 (TID 38843, xyz): java.util.ConcurrentModificationException: KafkaConsumer is not safe for multi-threaded access\n at org.apache.kafka.clients.consumer.KafkaConsumer.acquire(KafkaConsumer.java:1431)\n at org.apache.kafka.clients.consumer.KafkaConsumer.poll(KafkaConsumer.java:929)\n at org.apache.spark.streaming.kafka010.CachedKafkaConsumer.poll(CachedKafkaConsumer.scala:99)\n at org.apache.spark.streaming.kafka010.CachedKafkaConsumer.get(CachedKafkaConsumer.scala:73)\n at org.apache.spark.streaming.kafka010.KafkaRDD$KafkaRDDIterator.next(KafkaRDD.scala:227)\n at org.apache.spark.streaming.kafka010.KafkaRDD$KafkaRDDIterator.next(KafkaRDD.scala:193)\n at scala.collection.Iterator$$anon$12.nextCur(Iterator.scala:434)\n\n{code}","issue_id":"13049780","key":"SPARK-19888","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-02-14T14:00:49.000+0000","role":"fixed_distractor","summary":"Seeing offsets not resetting even when reset policy is configured explicitly"} {"case_id":"13059042","cluster":"DISTRACTOR-SPARK-20086","comments":[{"body":"User 'hvanhovell' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17432","created":"2017-03-26T07:58:03.448+0000"},{"body":"Is there a way I can work around this if I'm stuck on Spark 2.1.0 ?\n\nThanks,\nMatthew","created":"2017-07-14T19:31:32.615+0000"},{"body":"will there be a fix for 2.1.0?\n\nRegards,\nJacky","created":"2017-08-02T22:23:40.379+0000"}],"conversations":[{"body":"original post at\n[stackoverflow | http://stackoverflow.com/questions/43007433/pyspark-2-1-0-error-when-working-with-window-function]\n\nI get error when working with pyspark window function. here is some example code:\n\n{code:title=borderStyle=solid}\n import pyspark\n import pyspark.sql.functions as sf\n import pyspark.sql.types as sparktypes\n from pyspark.sql import window\n \n sc = pyspark.SparkContext()\n sqlc = pyspark.SQLContext(sc)\n rdd = sc.parallelize([(1, 2.0), (1, 3.0), (1, 1.), (1, -2.), (1, -1.)])\n df = sqlc.createDataFrame(rdd, [\"x\", \"AmtPaid\"])\n df.show()\n\n{code}\n\ngives:\n\n\n | x|AmtPaid|\n | 1| 2.0|\n | 1| 3.0|\n | 1| 1.0|\n | 1| -2.0|\n | 1| -1.0|\n\n\nnext, compute cumulative sum\n\n{code:title=test.py|borderStyle=solid}\n win_spec_max = (window.Window\n .partitionBy(['x'])\n .rowsBetween(window.Window.unboundedPreceding, 0)))\n df = df.withColumn('AmtPaidCumSum',\n sf.sum(sf.col('AmtPaid')).over(win_spec_max))\n df.show()\n{code}\n\ngives,\n\n\n | x|AmtPaid|AmtPaidCumSum|\n | 1| 2.0| 2.0|\n | 1| 3.0| 5.0|\n | 1| 1.0| 6.0|\n | 1| -2.0| 4.0|\n | 1| -1.0| 3.0|\n\nnext, compute cumulative max,\n\n{code}\n df = df.withColumn('AmtPaidCumSumMax', sf.max(sf.col('AmtPaidCumSum')).over(win_spec_max))\n\n df.show()\n{code}\n\ngives error log\n\n{noformat}\n Py4JJavaError: An error occurred while calling o2609.showString.\n\n\nwith traceback:\n\n\n Py4JJavaErrorTraceback (most recent call last)\n in ()\n ----> 1 df.show()\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/dataframe.pyc in show(self, n, truncate)\n 316 \"\"\"\n 317 if isinstance(truncate, bool) and truncate:\n --> 318 print(self._jdf.showString(n, 20))\n 319 else:\n 320 print(self._jdf.showString(n, int(truncate)))\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/java_gateway.pyc in __call__(self, *args)\n 1131 answer = self.gateway_client.send_command(command)\n 1132 return_value = get_return_value(\n -> 1133 answer, self.gateway_client, self.target_id, self.name)\n 1134 \n 1135 for temp_arg in temp_args:\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/utils.pyc in deco(*a, **kw)\n 61 def deco(*a, **kw):\n 62 try:\n ---> 63 return f(*a, **kw)\n 64 except py4j.protocol.Py4JJavaError as e:\n 65 s = e.java_exception.toString()\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/protocol.pyc in get_return_value(answer, gateway_client, target_id, name)\n 317 raise Py4JJavaError(\n 318 \"An error occurred while calling {0}{1}{2}.\\n\".\n --> 319 format(target_id, \".\", name), value)\n 320 else:\n 321 raise Py4JError(\n{noformat}\n\nbut interestingly enough, if i introduce another change before sencond window operation, say inserting a column then it does not give that error:\n\n{code}\n df = df.withColumn('MaxBound', sf.lit(6.))\n df.show()\n{code}\n\n\n | x|AmtPaid|AmtPaidCumSum|MaxBound|\n | 1| 2.0| 2.0| 6.0|\n | 1| 3.0| 5.0| 6.0|\n | 1| 1.0| 6.0| 6.0|\n | 1| -2.0| 4.0| 6.0|\n | 1| -1.0| 3.0| 6.0|\n\n{code}\n #then apply the second window operations\n df = df.withColumn('AmtPaidCumSumMax', sf.max(sf.col('AmtPaidCumSum')).over(win_spec_max))\n df.show()\n{code}\n\n | x|AmtPaid|AmtPaidCumSum|MaxBound|AmtPaidCumSumMax|\n | 1| 2.0| 2.0| 6.0| 2.0|\n | 1| 3.0| 5.0| 6.0| 5.0|\n | 1| 1.0| 6.0| 6.0| 6.0|\n | 1| -2.0| 4.0| 6.0| 6.0|\n | 1| -1.0| 3.0| 6.0| 6.0|\n\nI do not understand this behaviour\n\nwell, so far so good, but then I try another operation then again get similar error:\n\n{code}\n def _udf_compare_cumsum_sll(x):\n if x['AmtPaidCumSumMax'] >= x['MaxBound']:\n output = 0\n else:\n output = x['AmtPaid']\n return output\n\n udf_compare_cumsum_sll = sf.udf(_udf_compare_cumsum_sll, sparktypes.FloatType())\n df = df.withColumn('AmtPaidAdjusted', udf_compare_cumsum_sll(sf.struct([df[x] for x in df.columns])))\n df.show()\n{code}\n\ngives,\n\n{noformat}\n Py4JJavaErrorTraceback (most recent call last)\n in ()\n ----> 1 df.show()\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/dataframe.pyc in show(self, n, truncate)\n 316 \"\"\"\n 317 if isinstance(truncate, bool) and truncate:\n --> 318 print(self._jdf.showString(n, 20))\n 319 else:\n 320 print(self._jdf.showString(n, int(truncate)))\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/java_gateway.pyc in __call__(self, *args)\n 1131 answer = self.gateway_client.send_command(command)\n 1132 return_value = get_return_value(\n -> 1133 answer, self.gateway_client, self.target_id, self.name)\n 1134 \n 1135 for temp_arg in temp_args:\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/utils.pyc in deco(*a, **kw)\n 61 def deco(*a, **kw):\n 62 try:\n ---> 63 return f(*a, **kw)\n 64 except py4j.protocol.Py4JJavaError as e:\n 65 s = e.java_exception.toString()\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/protocol.pyc in get_return_value(answer, gateway_client, target_id, name)\n 317 raise Py4JJavaError(\n 318 \"An error occurred while calling {0}{1}{2}.\\n\".\n --> 319 format(target_id, \".\", name), value)\n 320 else:\n 321 raise Py4JError(\n\n Py4JJavaError: An error occurred while calling o91.showString.\n : org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 36.0 failed 1 times, most recent failure: Lost task 0.0 in stage 36.0 (TID 645, localhost, executor driver): org.apache.spark.sql.catalyst.errors.package$TreeNodeException: Binding attribute, tree: AmtPaidCumSum#10\n{noformat} \t\n\nI wonder if someone could reproduce this behaviour ...\n\n\nhere is complete log ..\n\n{noformat}\n Py4JJavaErrorTraceback (most recent call last)\n in ()\n ----> 1 df.show()\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/dataframe.pyc in show(self, n, truncate)\n 316 \"\"\"\n 317 if isinstance(truncate, bool) and truncate:\n --> 318 print(self._jdf.showString(n, 20))\n 319 else:\n 320 print(self._jdf.showString(n, int(truncate)))\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/java_gateway.pyc in __call__(self, *args)\n 1131 answer = self.gateway_client.send_command(command)\n 1132 return_value = get_return_value(\n -> 1133 answer, self.gateway_client, self.target_id, self.name)\n 1134\n 1135 for temp_arg in temp_args:\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/utils.pyc in deco(*a, **kw)\n 61 def deco(*a, **kw):\n 62 try:\n ---> 63 return f(*a, **kw)\n 64 except py4j.protocol.Py4JJavaError as e:\n 65 s = e.java_exception.toString()\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/protocol.pyc in get_return_value(answer, gateway_client, target_id, name)\n 317 raise Py4JJavaError(\n 318 \"An error occurred while calling {0}{1}{2}.\\n\".\n --> 319 format(target_id, \".\", name), value)\n 320 else:\n 321 raise Py4JError(\n\n Py4JJavaError: An error occurred while calling o703.showString.\n : org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 119.0 failed 1 times, most recent failure: Lost task 0.0 in stage 119.0 (TID 1817, localhost, executor driver): org.apache.spark.sql.catalyst.errors.package$TreeNodeException: Binding attribute, tree: AmtPaidCumSum#2076\n \tat org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:56)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:88)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:87)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:288)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:288)\n \tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:70)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:287)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5$$anonfun$apply$11.apply(TreeNode.scala:360)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.immutable.List.foreach(List.scala:381)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.immutable.List.map(List.scala:285)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:358)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:188)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:329)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:277)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$.bindReference(BoundAttribute.scala:87)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:38)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:38)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n \tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.AbstractTraversable.map(Traversable.scala:104)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.bind(GenerateMutableProjection.scala:38)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.generate(GenerateMutableProjection.scala:44)\n \tat org.apache.spark.sql.execution.SparkPlan.newMutableProjection(SparkPlan.scala:353)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1$1.apply(WindowExec.scala:203)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1$1.apply(WindowExec.scala:202)\n \tat org.apache.spark.sql.execution.window.AggregateProcessor$.apply(AggregateProcessor.scala:98)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2.org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1(WindowExec.scala:198)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$6.apply(WindowExec.scala:225)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$6.apply(WindowExec.scala:222)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1$$anonfun$16.apply(WindowExec.scala:318)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1$$anonfun$16.apply(WindowExec.scala:318)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n \tat scala.collection.mutable.ArrayOps$ofRef.foreach(ArrayOps.scala:186)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.mutable.ArrayOps$ofRef.map(ArrayOps.scala:186)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1.(WindowExec.scala:318)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14.apply(WindowExec.scala:290)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14.apply(WindowExec.scala:289)\n \tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:796)\n \tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:796)\n \tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n \tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n \tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n \tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n \tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n \tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n \tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\n \tat org.apache.spark.scheduler.Task.run(Task.scala:99)\n \tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:282)\n \tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n \tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n \tat java.lang.Thread.run(Thread.java:745)\n Caused by: java.lang.RuntimeException: Couldn't find AmtPaidCumSum#2076 in [sum#2299,max#2300,x#2066L,AmtPaid#2067]\n \tat scala.sys.package$.error(package.scala:27)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:94)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:88)\n \tat org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:52)\n \t... 62 more\n\n Driver stacktrace:\n \tat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1435)\n \tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1423)\n \tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1422)\n \tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n \tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n \tat org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1422)\n \tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:802)\n \tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:802)\n \tat scala.Option.foreach(Option.scala:257)\n \tat org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:802)\n \tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:1650)\n \tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1605)\n \tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1594)\n \tat org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\n \tat org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:628)\n \tat org.apache.spark.SparkContext.runJob(SparkContext.scala:1918)\n \tat org.apache.spark.SparkContext.runJob(SparkContext.scala:1931)\n \tat org.apache.spark.SparkContext.runJob(SparkContext.scala:1944)\n \tat org.apache.spark.sql.execution.SparkPlan.executeTake(SparkPlan.scala:333)\n \tat org.apache.spark.sql.execution.CollectLimitExec.executeCollect(limit.scala:38)\n \tat org.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$execute$1$1.apply(Dataset.scala:2371)\n \tat org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:57)\n \tat org.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2765)\n \tat org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2370)\n \tat org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2377)\n \tat org.apache.spark.sql.Dataset$$anonfun$head$1.apply(Dataset.scala:2113)\n \tat org.apache.spark.sql.Dataset$$anonfun$head$1.apply(Dataset.scala:2112)\n \tat org.apache.spark.sql.Dataset.withTypedCallback(Dataset.scala:2795)\n \tat org.apache.spark.sql.Dataset.head(Dataset.scala:2112)\n \tat org.apache.spark.sql.Dataset.take(Dataset.scala:2327)\n \tat org.apache.spark.sql.Dataset.showString(Dataset.scala:248)\n \tat sun.reflect.GeneratedMethodAccessor83.invoke(Unknown Source)\n \tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n \tat java.lang.reflect.Method.invoke(Method.java:498)\n \tat py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:244)\n \tat py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:357)\n \tat py4j.Gateway.invoke(Gateway.java:280)\n \tat py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:132)\n \tat py4j.commands.CallCommand.execute(CallCommand.java:79)\n \tat py4j.GatewayConnection.run(GatewayConnection.java:214)\n \tat java.lang.Thread.run(Thread.java:745)\n Caused by: org.apache.spark.sql.catalyst.errors.package$TreeNodeException: Binding attribute, tree: null\n \tat org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:56)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:88)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:87)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:288)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:288)\n \tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:70)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:287)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5$$anonfun$apply$11.apply(TreeNode.scala:360)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.immutable.List.foreach(List.scala:381)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.immutable.List.map(List.scala:285)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:358)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:188)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:329)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:277)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$.bindReference(BoundAttribute.scala:87)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:38)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:38)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n \tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.AbstractTraversable.map(Traversable.scala:104)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.bind(GenerateMutableProjection.scala:38)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.generate(GenerateMutableProjection.scala:44)\n \tat org.apache.spark.sql.execution.SparkPlan.newMutableProjection(SparkPlan.scala:353)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1$1.apply(WindowExec.scala:203)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1$1.apply(WindowExec.scala:202)\n \tat org.apache.spark.sql.execution.window.AggregateProcessor$.apply(AggregateProcessor.scala:98)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2.org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1(WindowExec.scala:198)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$6.apply(WindowExec.scala:225)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$6.apply(WindowExec.scala:222)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1$$anonfun$16.apply(WindowExec.scala:318)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1$$anonfun$16.apply(WindowExec.scala:318)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n \tat scala.collection.mutable.ArrayOps$ofRef.foreach(ArrayOps.scala:186)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.mutable.ArrayOps$ofRef.map(ArrayOps.scala:186)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1.(WindowExec.scala:318)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14.apply(WindowExec.scala:290)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14.apply(WindowExec.scala:289)\n \tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:796)\n \tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:796)\n \tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n \tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n \tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n \tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n \tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n \tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n \tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\n \tat org.apache.spark.scheduler.Task.run(Task.scala:99)\n \tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:282)\n \tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n \tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n \t... 1 more\n Caused by: java.lang.RuntimeException: Couldn't find AmtPaidCumSum#2076 in [sum#2299,max#2300,x#2066L,AmtPaid#2067]\n \tat scala.sys.package$.error(package.scala:27)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:94)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:88)\n \tat org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:52)\n \t... 62 more\n{noformat}","from":"reporter","subject":"issue with pyspark 2.1.0 window function"},{"body":"User 'hvanhovell' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17432","from":"developer"},{"body":"Is there a way I can work around this if I'm stuck on Spark 2.1.0 ?\n\nThanks,\nMatthew","from":"developer"},{"body":"will there be a fix for 2.1.0?\n\nRegards,\nJacky","from":"developer"}],"created":"2017-03-24T21:31:29.000+0000","description":"original post at\n[stackoverflow | http://stackoverflow.com/questions/43007433/pyspark-2-1-0-error-when-working-with-window-function]\n\nI get error when working with pyspark window function. here is some example code:\n\n{code:title=borderStyle=solid}\n import pyspark\n import pyspark.sql.functions as sf\n import pyspark.sql.types as sparktypes\n from pyspark.sql import window\n \n sc = pyspark.SparkContext()\n sqlc = pyspark.SQLContext(sc)\n rdd = sc.parallelize([(1, 2.0), (1, 3.0), (1, 1.), (1, -2.), (1, -1.)])\n df = sqlc.createDataFrame(rdd, [\"x\", \"AmtPaid\"])\n df.show()\n\n{code}\n\ngives:\n\n\n | x|AmtPaid|\n | 1| 2.0|\n | 1| 3.0|\n | 1| 1.0|\n | 1| -2.0|\n | 1| -1.0|\n\n\nnext, compute cumulative sum\n\n{code:title=test.py|borderStyle=solid}\n win_spec_max = (window.Window\n .partitionBy(['x'])\n .rowsBetween(window.Window.unboundedPreceding, 0)))\n df = df.withColumn('AmtPaidCumSum',\n sf.sum(sf.col('AmtPaid')).over(win_spec_max))\n df.show()\n{code}\n\ngives,\n\n\n | x|AmtPaid|AmtPaidCumSum|\n | 1| 2.0| 2.0|\n | 1| 3.0| 5.0|\n | 1| 1.0| 6.0|\n | 1| -2.0| 4.0|\n | 1| -1.0| 3.0|\n\nnext, compute cumulative max,\n\n{code}\n df = df.withColumn('AmtPaidCumSumMax', sf.max(sf.col('AmtPaidCumSum')).over(win_spec_max))\n\n df.show()\n{code}\n\ngives error log\n\n{noformat}\n Py4JJavaError: An error occurred while calling o2609.showString.\n\n\nwith traceback:\n\n\n Py4JJavaErrorTraceback (most recent call last)\n in ()\n ----> 1 df.show()\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/dataframe.pyc in show(self, n, truncate)\n 316 \"\"\"\n 317 if isinstance(truncate, bool) and truncate:\n --> 318 print(self._jdf.showString(n, 20))\n 319 else:\n 320 print(self._jdf.showString(n, int(truncate)))\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/java_gateway.pyc in __call__(self, *args)\n 1131 answer = self.gateway_client.send_command(command)\n 1132 return_value = get_return_value(\n -> 1133 answer, self.gateway_client, self.target_id, self.name)\n 1134 \n 1135 for temp_arg in temp_args:\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/utils.pyc in deco(*a, **kw)\n 61 def deco(*a, **kw):\n 62 try:\n ---> 63 return f(*a, **kw)\n 64 except py4j.protocol.Py4JJavaError as e:\n 65 s = e.java_exception.toString()\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/protocol.pyc in get_return_value(answer, gateway_client, target_id, name)\n 317 raise Py4JJavaError(\n 318 \"An error occurred while calling {0}{1}{2}.\\n\".\n --> 319 format(target_id, \".\", name), value)\n 320 else:\n 321 raise Py4JError(\n{noformat}\n\nbut interestingly enough, if i introduce another change before sencond window operation, say inserting a column then it does not give that error:\n\n{code}\n df = df.withColumn('MaxBound', sf.lit(6.))\n df.show()\n{code}\n\n\n | x|AmtPaid|AmtPaidCumSum|MaxBound|\n | 1| 2.0| 2.0| 6.0|\n | 1| 3.0| 5.0| 6.0|\n | 1| 1.0| 6.0| 6.0|\n | 1| -2.0| 4.0| 6.0|\n | 1| -1.0| 3.0| 6.0|\n\n{code}\n #then apply the second window operations\n df = df.withColumn('AmtPaidCumSumMax', sf.max(sf.col('AmtPaidCumSum')).over(win_spec_max))\n df.show()\n{code}\n\n | x|AmtPaid|AmtPaidCumSum|MaxBound|AmtPaidCumSumMax|\n | 1| 2.0| 2.0| 6.0| 2.0|\n | 1| 3.0| 5.0| 6.0| 5.0|\n | 1| 1.0| 6.0| 6.0| 6.0|\n | 1| -2.0| 4.0| 6.0| 6.0|\n | 1| -1.0| 3.0| 6.0| 6.0|\n\nI do not understand this behaviour\n\nwell, so far so good, but then I try another operation then again get similar error:\n\n{code}\n def _udf_compare_cumsum_sll(x):\n if x['AmtPaidCumSumMax'] >= x['MaxBound']:\n output = 0\n else:\n output = x['AmtPaid']\n return output\n\n udf_compare_cumsum_sll = sf.udf(_udf_compare_cumsum_sll, sparktypes.FloatType())\n df = df.withColumn('AmtPaidAdjusted', udf_compare_cumsum_sll(sf.struct([df[x] for x in df.columns])))\n df.show()\n{code}\n\ngives,\n\n{noformat}\n Py4JJavaErrorTraceback (most recent call last)\n in ()\n ----> 1 df.show()\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/dataframe.pyc in show(self, n, truncate)\n 316 \"\"\"\n 317 if isinstance(truncate, bool) and truncate:\n --> 318 print(self._jdf.showString(n, 20))\n 319 else:\n 320 print(self._jdf.showString(n, int(truncate)))\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/java_gateway.pyc in __call__(self, *args)\n 1131 answer = self.gateway_client.send_command(command)\n 1132 return_value = get_return_value(\n -> 1133 answer, self.gateway_client, self.target_id, self.name)\n 1134 \n 1135 for temp_arg in temp_args:\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/utils.pyc in deco(*a, **kw)\n 61 def deco(*a, **kw):\n 62 try:\n ---> 63 return f(*a, **kw)\n 64 except py4j.protocol.Py4JJavaError as e:\n 65 s = e.java_exception.toString()\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/protocol.pyc in get_return_value(answer, gateway_client, target_id, name)\n 317 raise Py4JJavaError(\n 318 \"An error occurred while calling {0}{1}{2}.\\n\".\n --> 319 format(target_id, \".\", name), value)\n 320 else:\n 321 raise Py4JError(\n\n Py4JJavaError: An error occurred while calling o91.showString.\n : org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 36.0 failed 1 times, most recent failure: Lost task 0.0 in stage 36.0 (TID 645, localhost, executor driver): org.apache.spark.sql.catalyst.errors.package$TreeNodeException: Binding attribute, tree: AmtPaidCumSum#10\n{noformat} \t\n\nI wonder if someone could reproduce this behaviour ...\n\n\nhere is complete log ..\n\n{noformat}\n Py4JJavaErrorTraceback (most recent call last)\n in ()\n ----> 1 df.show()\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/dataframe.pyc in show(self, n, truncate)\n 316 \"\"\"\n 317 if isinstance(truncate, bool) and truncate:\n --> 318 print(self._jdf.showString(n, 20))\n 319 else:\n 320 print(self._jdf.showString(n, int(truncate)))\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/java_gateway.pyc in __call__(self, *args)\n 1131 answer = self.gateway_client.send_command(command)\n 1132 return_value = get_return_value(\n -> 1133 answer, self.gateway_client, self.target_id, self.name)\n 1134\n 1135 for temp_arg in temp_args:\n\n /Users/<>/spark-2.1.0-bin-hadoop2.7/python/pyspark/sql/utils.pyc in deco(*a, **kw)\n 61 def deco(*a, **kw):\n 62 try:\n ---> 63 return f(*a, **kw)\n 64 except py4j.protocol.Py4JJavaError as e:\n 65 s = e.java_exception.toString()\n\n /Users/<>/.virtualenvs/<>/lib/python2.7/site-packages/py4j/protocol.pyc in get_return_value(answer, gateway_client, target_id, name)\n 317 raise Py4JJavaError(\n 318 \"An error occurred while calling {0}{1}{2}.\\n\".\n --> 319 format(target_id, \".\", name), value)\n 320 else:\n 321 raise Py4JError(\n\n Py4JJavaError: An error occurred while calling o703.showString.\n : org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 119.0 failed 1 times, most recent failure: Lost task 0.0 in stage 119.0 (TID 1817, localhost, executor driver): org.apache.spark.sql.catalyst.errors.package$TreeNodeException: Binding attribute, tree: AmtPaidCumSum#2076\n \tat org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:56)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:88)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:87)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:288)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:288)\n \tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:70)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:287)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5$$anonfun$apply$11.apply(TreeNode.scala:360)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.immutable.List.foreach(List.scala:381)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.immutable.List.map(List.scala:285)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:358)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:188)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:329)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:277)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$.bindReference(BoundAttribute.scala:87)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:38)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:38)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n \tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.AbstractTraversable.map(Traversable.scala:104)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.bind(GenerateMutableProjection.scala:38)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.generate(GenerateMutableProjection.scala:44)\n \tat org.apache.spark.sql.execution.SparkPlan.newMutableProjection(SparkPlan.scala:353)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1$1.apply(WindowExec.scala:203)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1$1.apply(WindowExec.scala:202)\n \tat org.apache.spark.sql.execution.window.AggregateProcessor$.apply(AggregateProcessor.scala:98)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2.org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1(WindowExec.scala:198)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$6.apply(WindowExec.scala:225)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$6.apply(WindowExec.scala:222)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1$$anonfun$16.apply(WindowExec.scala:318)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1$$anonfun$16.apply(WindowExec.scala:318)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n \tat scala.collection.mutable.ArrayOps$ofRef.foreach(ArrayOps.scala:186)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.mutable.ArrayOps$ofRef.map(ArrayOps.scala:186)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1.(WindowExec.scala:318)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14.apply(WindowExec.scala:290)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14.apply(WindowExec.scala:289)\n \tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:796)\n \tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:796)\n \tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n \tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n \tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n \tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n \tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n \tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n \tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\n \tat org.apache.spark.scheduler.Task.run(Task.scala:99)\n \tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:282)\n \tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n \tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n \tat java.lang.Thread.run(Thread.java:745)\n Caused by: java.lang.RuntimeException: Couldn't find AmtPaidCumSum#2076 in [sum#2299,max#2300,x#2066L,AmtPaid#2067]\n \tat scala.sys.package$.error(package.scala:27)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:94)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:88)\n \tat org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:52)\n \t... 62 more\n\n Driver stacktrace:\n \tat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1435)\n \tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1423)\n \tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1422)\n \tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n \tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n \tat org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1422)\n \tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:802)\n \tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:802)\n \tat scala.Option.foreach(Option.scala:257)\n \tat org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:802)\n \tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:1650)\n \tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1605)\n \tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1594)\n \tat org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\n \tat org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:628)\n \tat org.apache.spark.SparkContext.runJob(SparkContext.scala:1918)\n \tat org.apache.spark.SparkContext.runJob(SparkContext.scala:1931)\n \tat org.apache.spark.SparkContext.runJob(SparkContext.scala:1944)\n \tat org.apache.spark.sql.execution.SparkPlan.executeTake(SparkPlan.scala:333)\n \tat org.apache.spark.sql.execution.CollectLimitExec.executeCollect(limit.scala:38)\n \tat org.apache.spark.sql.Dataset$$anonfun$org$apache$spark$sql$Dataset$$execute$1$1.apply(Dataset.scala:2371)\n \tat org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:57)\n \tat org.apache.spark.sql.Dataset.withNewExecutionId(Dataset.scala:2765)\n \tat org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$execute$1(Dataset.scala:2370)\n \tat org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collect(Dataset.scala:2377)\n \tat org.apache.spark.sql.Dataset$$anonfun$head$1.apply(Dataset.scala:2113)\n \tat org.apache.spark.sql.Dataset$$anonfun$head$1.apply(Dataset.scala:2112)\n \tat org.apache.spark.sql.Dataset.withTypedCallback(Dataset.scala:2795)\n \tat org.apache.spark.sql.Dataset.head(Dataset.scala:2112)\n \tat org.apache.spark.sql.Dataset.take(Dataset.scala:2327)\n \tat org.apache.spark.sql.Dataset.showString(Dataset.scala:248)\n \tat sun.reflect.GeneratedMethodAccessor83.invoke(Unknown Source)\n \tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n \tat java.lang.reflect.Method.invoke(Method.java:498)\n \tat py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:244)\n \tat py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:357)\n \tat py4j.Gateway.invoke(Gateway.java:280)\n \tat py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:132)\n \tat py4j.commands.CallCommand.execute(CallCommand.java:79)\n \tat py4j.GatewayConnection.run(GatewayConnection.java:214)\n \tat java.lang.Thread.run(Thread.java:745)\n Caused by: org.apache.spark.sql.catalyst.errors.package$TreeNodeException: Binding attribute, tree: null\n \tat org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:56)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:88)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1.applyOrElse(BoundAttribute.scala:87)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:288)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:288)\n \tat org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:70)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:287)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformDown$1.apply(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5$$anonfun$apply$11.apply(TreeNode.scala:360)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.immutable.List.foreach(List.scala:381)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.immutable.List.map(List.scala:285)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$5.apply(TreeNode.scala:358)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:188)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildren(TreeNode.scala:329)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:293)\n \tat org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:277)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$.bindReference(BoundAttribute.scala:87)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:38)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$$anonfun$bind$1.apply(GenerateMutableProjection.scala:38)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n \tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.AbstractTraversable.map(Traversable.scala:104)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.bind(GenerateMutableProjection.scala:38)\n \tat org.apache.spark.sql.catalyst.expressions.codegen.GenerateMutableProjection$.generate(GenerateMutableProjection.scala:44)\n \tat org.apache.spark.sql.execution.SparkPlan.newMutableProjection(SparkPlan.scala:353)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1$1.apply(WindowExec.scala:203)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1$1.apply(WindowExec.scala:202)\n \tat org.apache.spark.sql.execution.window.AggregateProcessor$.apply(AggregateProcessor.scala:98)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2.org$apache$spark$sql$execution$window$WindowExec$$anonfun$$processor$1(WindowExec.scala:198)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$6.apply(WindowExec.scala:225)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$windowFrameExpressionFactoryPairs$2$$anonfun$6.apply(WindowExec.scala:222)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1$$anonfun$16.apply(WindowExec.scala:318)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1$$anonfun$16.apply(WindowExec.scala:318)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n \tat scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n \tat scala.collection.mutable.ArrayOps$ofRef.foreach(ArrayOps.scala:186)\n \tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n \tat scala.collection.mutable.ArrayOps$ofRef.map(ArrayOps.scala:186)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14$$anon$1.(WindowExec.scala:318)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14.apply(WindowExec.scala:290)\n \tat org.apache.spark.sql.execution.window.WindowExec$$anonfun$14.apply(WindowExec.scala:289)\n \tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:796)\n \tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1$$anonfun$apply$23.apply(RDD.scala:796)\n \tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n \tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n \tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n \tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n \tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n \tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n \tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\n \tat org.apache.spark.scheduler.Task.run(Task.scala:99)\n \tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:282)\n \tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n \tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n \t... 1 more\n Caused by: java.lang.RuntimeException: Couldn't find AmtPaidCumSum#2076 in [sum#2299,max#2300,x#2066L,AmtPaid#2067]\n \tat scala.sys.package$.error(package.scala:27)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:94)\n \tat org.apache.spark.sql.catalyst.expressions.BindReferences$$anonfun$bindReference$1$$anonfun$applyOrElse$1.apply(BoundAttribute.scala:88)\n \tat org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:52)\n \t... 62 more\n{noformat}","issue_id":"13059042","key":"SPARK-20086","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-03-26T20:48:57.000+0000","role":"fixed_distractor","summary":"issue with pyspark 2.1.0 window function"} {"case_id":"13061083","cluster":"DISTRACTOR-SPARK-20202","comments":[{"body":"I see wide agreement on that. One question I have is, is including Hive this way merely a really-not-nice-to-have or actually not allowed? I think the question is whether sources are available, right? because releases can't have binary-only parts. I plead ignorance, I have never myself paid much attention to this integration. \n\nIf it's not then this sounds like something has to change for releases beyond 2.1.1 and this can be targeted as a Blocker accordingly.\n\nDoes this depend on refactoring or changes in Hive? IIRC the problem was hive-exec being an uber-jar, but it's been a long time since I read any of that discussion.","created":"2017-04-03T10:29:39.611+0000"},{"body":"As an Apache member, the Spark project can't release binary artifacts that aren't made from its Apache code base. So either, the Spark project needs to use Hive's release artifacts or it needs to formally fork Hive and move the fork into its git repository at Apache and rename it away from org.apache.hive to org.apache.spark. The current path is not allowed.\n\nHive is in the middle of rolling releases and thus this is a good time to make requests. The old uber jar (hive-exec) is already released separately with the classifier \"core.\" It looks like we are using the same protobuf (2.5.0) and kryo (3.0.3) versions.","created":"2017-04-03T11:15:47.021+0000"},{"body":"Agree. I think the logic was that Spark had released its own source/binary version of Hive, and then used that in Spark. I don't think anybody believes that's a good solution in the long term; it was a work-around for hive-exec's packaging IIRC. Once whatever that is is resolved this can go away, but I defer to those who know the issue better on the details.\n\nWhat I'm not clear on is whether the current org.spark-hive situation is streeetching the source/binary policy so far that it breaks, enough that no more releases can happen without it. Best to make it go away ASAP anyway. But I don't know if changes in Hive 2.5 help integration with Hive 1.x. It may require either temporarily blessing the fork, or more jar surgery to un-uberize the hive-exec jar or something.","created":"2017-04-03T11:23:07.715+0000"},{"body":"I should also say here that the Hive community is willing to help. We are in the process of rolling releases so if Spark needs a change, we can work together to get this done.","created":"2017-04-03T11:25:53.996+0000"},{"body":"It is against Apache policy to release binaries that aren't part of your project.","created":"2017-04-03T13:16:23.483+0000"},{"body":"Alrighty, you can leave the status for now, but generally committers set Blocker. I'm not entirely clear this blocks a release, not yet.\n\nYou're absolutely right, but, the hive fork with binaries and source is part of this project. At least, that's the idea. For example, this is notionally voted on and released with each Spark release, but the binary/source of this fork project isn't separately, explicitly, voted on and separately released. I think that should occur for avoidance of doubt, that this is a blessed artifact of the Spark project. Would this answer your process and policy concerns about the release? It's not pretty but I think that's within the law.\n\nOf course, it's no answer in the long term. The goal is to not have to use the fork at all. If Hive packaging changes are already in place to make it unnecessary, great (is that all there is to it, everyone?) I don't know if that presents a solution for earlier versions of Hive. This fork thing may persist in existing branches, but it has to at least be released and used in a proper way. This may need fixes right now.","created":"2017-04-03T13:37:52.818+0000"},{"body":"[~marmbrus] as release manager of the moment, I suggest we actually formally vote on release the org.spark-project.hive artifact, as I'm not clear we ever did formally. That much seems like a must-have. I don't know that it requires re-releasing the artifacts, but at least having a meaningful review of what it is, and agreeing (or not) that it's what the PMC wants to release, would I believe resolve doubts about the legitimacy of that artifact.\n\nThere's still more to do no doubt, to get rid of the fork. This might include seeing if Hive 1.2.x can provide an un-uberized artifact.","created":"2017-04-04T14:23:57.482+0000"},{"body":"# the ugliness need to inset the spark thrift stuff under the hive thrift stuff is obsolete, can be cut entirely.\n# with the shading of kryo not needed, an unshaded hive *may* work. I forget which troublespots there were last time, probably the usual suspects: jackson, guava, etc.\n# Hive 1.2.x refuses to work with Hadoop 3; it considers that an unsupported version. For basic client-side testing, you can build Hadoop 3 with a fake version (e..g {{mvn install -DskipShade -Ddeclared.hadoop.version=2.11}}, but as hadoop version is something which NN/DNs care about, not something that's really going to work in real systems. Presumably later hive versions will address that.\n\nIf hive take over ownership of the spark 1.2.1-spark branch, this could be done first simply by pulling the spark branch into the Hive repo as a branch, defining the artifact naming properly and releasing it. If that is done, before any release of that 1.2.x branch is done, there's a couple of outstanding PRs to pull in (groovy version for security reasons, ... ).. A quick import & re-release would be the fast way to get this out as an asf-approved binary","created":"2017-04-05T13:41:37.852+0000"},{"body":"+1 for a release of the Spark fork from the Hive community. While the reasons for the fork appear to be fixed in the latest version, there's a lot of work to do to get Spark on a newer Hive version. And for patch releases like 2.0.3 and 2.1.1, I don't think updating Hive is an option.\n\nI'm also all for getting master on a real Hive release. A release of alternate Hive binaries was inappropriate. I think that if a third-party organization had done the same, it would be entirely reasonable to treat it as a trademark violation and ask them to stop.\n\nhttp://www.apache.org/foundation/marks/faq/#products","created":"2017-04-05T18:22:34.164+0000"},{"body":"Yes this is really important. The proper way to do this is to publish a proper version of Hive with the right dependency declared (rather than including all the dependencies in a uber jar). Looks like there are broad support to do this. I'm going to create a JIRA ticket on Hive and add a dependency on this.\n\nThis ticket will depend on that.","created":"2017-04-05T21:48:48.406+0000"},{"body":"I've created a ticket on the Hive side to publish 1.2.x: https://issues.apache.org/jira/browse/HIVE-16391\n\nUntil that is resolved, I also wonder if there are other things we should do. For example, vote on the current fork to rectify it?\n\n","created":"2017-04-05T22:05:09.361+0000"},{"body":"One thing I do recall as trouble here was that ivy resolution was different from mvns, and fixing up all the transitives was a troublespot. Patches here need to be tested against SBT and maven —as jenkins only does SBT, the mvn builds will have to be manual. I don't remember which specific dependency was the problem.","created":"2017-04-10T14:05:56.946+0000"},{"body":"Would it possible make sense to untarget this from the maintenance releases (1.6.X, 2.0.X, 2.1.X) and instead focus on the future versions?","created":"2017-04-12T00:54:09.619+0000"},{"body":"There are no currently targeted version, are there?\n","created":"2017-04-12T00:56:02.881+0000"},{"body":"Oh right, sorry I was misreading the intent of Affects Version/s.","created":"2017-04-12T16:51:57.623+0000"},{"body":"How about upgrade Hive directly to 2.3.2. In fact, I've completed the initial work and have been running for a few days.\r\n [https://github.com/apache/spark/pull/20659]","created":"2018-03-05T14:15:47.646+0000"},{"body":"Prefer newer Hive also","created":"2018-06-01T16:46:20.438+0000"},{"body":"What is our plan to to fix this issue, are we going to use new Hive version, or we are still stick to 1.2?\r\n\r\nIf we're still stick to 1.2, [~stevel@apache.org] and I will take this issue and make the ball rolling in Hive community.","created":"2018-06-04T03:44:46.801+0000"},{"body":"I think you could split things into two\r\n\r\n# a modified hive 1.2.1.x for hadoop 3, with a new package name in the maven builds. (joy, profiles!) and some work with the hive team to get this officially published by them. Strength: easy for people to backport into shipping 2.2, 2.3 builds just by changing the POM\r\n\r\n# the bigger move to Hive 2. This will be the best for future, but is bound to have more surprises. There's even the possibility that the hive team might have to make some changes too, which isn't impossible if the timelines line up.","created":"2018-06-04T17:25:04.833+0000"},{"body":"OK, for the 1st, I've already started working on it locally. Looks like it is not a big change, only some POM changes are enough, I will submit a patch to Hive community.","created":"2018-06-05T03:00:24.765+0000"},{"body":"Hi all, what do you guys think about replacing it to Hive 2.3.x in the near future (like Spark 3.0.0) given SPARK-23710, and keeping the fork for now?\r\n\r\nLooks [~q79969786] completed the initial try at SPARK-23710 and now it sounds pretty much feasible as an option now although it sounds there are still some investigations; however, I believe that we can focus on getting through if we have the explicit plan here. I think we are mostly all positive on this option as a final goal anyway but I felt like we need to make sure on this.\r\nIf the above can be set as the goal for this JIRA to get rid of the fork completely, \\*I personally think\\* Hive side also can focus on landing other fixes to the more resent versions without diverting the efforts to maintain an old branch.\r\n\r\nUntil then, I think we could probably consider keeping the fork for now and landing some minor fixes if there're some strong reasons for it. For example, Hadoop 3 support is blocked by one liner fix in the fork. \\*I personally think\\* it is the easiest way to land this fix into the fork. I believe this is pretty reasonable.\r\n\r\nWhat do you guys think about this?\r\n\r\n","created":"2018-06-23T15:52:51.762+0000"},{"body":"[~owen.omalley] and [~rxin], what do you think about the suggestion above? I tried to check all other contexts hard at all my best and ^ was my current conclusion to get through this issue mostly smoothly and easily. ","created":"2018-06-25T03:52:22.751+0000"},{"body":"Would you guys please give some thought on this when you guys are available?","created":"2018-06-29T07:45:21.558+0000"},{"body":"kindly ping [~owen.omalley] and [~rxin]. I would like to make a progress further on this since it's blocked for a while but it's pretty important to make up for this affair. ","created":"2018-07-03T17:19:20.253+0000"},{"body":"Hey [~owen.omalley] and [~rxin], I know I see many sensitive things for example the policy stuff frankly; however, this one needs some input from you guys before proceeding further ...","created":"2018-07-11T02:45:04.924+0000"},{"body":"If you want to try and put together a PR that actually does it, that could work too. But note that it's a lot of work to upgrade execution Hive. Probably 10X more work than Hive publishing the exec jar.\r\n\r\n ","created":"2018-07-11T17:46:56.533+0000"},{"body":"I was thinking we target it for 3.0.0 (otherwise 4.0.0 might make sense ... ). It might be a lot of work indeed but I believe this is what we should do as a final goal which we should do anyway sometime ...\r\n\r\nUntill then, I wanted to propose to keep the fork as a temporary solution until 3.0.0 and we target to upgrade it to 2.3.x in 3.0.0 as the goal and target version ... Branch-2.4 will be cut out soon and I think we would go for 3.0.0 for the next release if I am not mistaken.","created":"2018-07-11T17:59:58.512+0000"},{"body":"Yea you can try and see how difficult it is.\r\n\r\n ","created":"2018-07-11T18:01:02.666+0000"},{"body":"[~rxin], there was an initial try above already though which at least made the regression tests we wrote so far passed. I talked with [~q79969786] before and she's willing to finish this. For this, I need more supports from you and other guys to go this way ..\r\n\r\nI get your point too on the other hand. So, do you think we should rather not explicitly target it since it's pretty difficult and we should better let Hive publish 1.2.x first rather then keeping the fork since it's unclear if we make it in 3.0.0?","created":"2018-07-11T18:16:24.636+0000"},{"body":"I am asking this to set the goal for this JIRA as of the current status and make a progress on this. I left some comments here because to me it looked [~q79969786]'s try is kind of a new fact arrived here to be considered.\r\n\r\nIf publishing Hive is still preferred to get through here for any reason, I will help go with Saisai's patch in HIVE-16391. If keeping the fork and upgrade could be set as a goal for now, I will try to help go with Yumming's way and make a fix to the fork.\r\n\r\nWhich one do you prefer?","created":"2018-07-12T02:18:14.531+0000"},{"body":"How like will there be a hive release? HIVE-16391 is still open?\r\n\r\nStay with hive 1.2 will slowly become a big problem for us within a few months...","created":"2018-07-12T06:04:57.800+0000"},{"body":"I think we are unclear about how we are going to deal with this and it's been left open for a while ..\r\n\r\n[~rxin], do you maybe have some preference in [my comment above|https://issues.apache.org/jira/browse/SPARK-20202?focusedCommentId=16541034&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#comment-16541034]?\r\n\r\n1. Go with Saisai's patch in HIVE-16391\r\n - Publishing Hive 1.2.x could be easier but will give some overhead to Hive side (e.g., maintaining the old branches, for example, backports).\r\n - If I understood correctly, we have less problems (e.g., policy stuff) if we go publishing Hive 1.2.x HIVE-16391\r\n\r\n2. Target the upgrade with [~q79969786]'s fix, and add some fixes to our current fork when there's strong reasons\r\n - It is difficult but [~q79969786] made and completed an initial try about the upgrade. It still need some further investigation (e.g., see [SPARK-23710|https://issues.apache.org/jira/browse/SPARK-23710]) but the try made the regression tests passed at least. She's willing to finish this.\r\n - If we miss the Hive upgrade to 2.3.x in Spark 3.0.0, we should probably target 4.0.0 with upper version of Hive, which I guess make this upgrade even harder.\r\n - Looks we implicitly agree upon this should be the final goal in the long term.\r\n\r\nSee also [~stevel@apache.org]'s [comment above|https://issues.apache.org/jira/browse/SPARK-20202?focusedCommentId=16500560&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#comment-16500560].\r\n\r\nI am re-raising and giving some refreshes here because I personally see:\r\n\r\n- Few facts arrived here since the JIRA was open. So, it looked to me it might be better we consider the possible options again.\r\n- Looks we are quite unclear on this about how we should get through this to me.\r\n- To me, I am sure we need to share and feel in the same way for this JIRA and, it looks I need some more supports from you guys before we go ahead because it'd be a kind of not easily revertible changes. \r\n- Branch-2.4 will be cut out soon and we will go for Spark 3.0.0 if I am not mistaken.\r\n\r\nI know there are many sensitive things going on here; however, please kindly consider and give some inputs. I am sure we all feel that we should resolve this.\r\nLastly, FWIW, I am doing this on my own rather individually if it matters to anyone in any case.","created":"2018-07-15T04:29:21.291+0000"},{"body":"Hi, All.\r\nI set the target version to `3.1.0`. Please join the discussion if you have any concerns.\r\n- https://lists.apache.org/thread.html/eca4e55c717f35f41c029e227fa9be0a7ee2c8a6f378fcce8f9fd4ff@%3Cdev.spark.apache.org%3E","created":"2019-11-19T22:55:46.625+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29936","created":"2020-10-03T22:24:33.342+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29936","created":"2020-10-03T22:25:02.185+0000"},{"body":"Issue resolved by pull request 29936\n[https://github.com/apache/spark/pull/29936]","created":"2020-10-05T22:30:15.876+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29973","created":"2020-10-08T03:00:53.650+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29973","created":"2020-10-08T03:00:55.782+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29973","created":"2020-10-08T03:01:46.194+0000"}],"conversations":[{"body":"Spark can't continue to depend on their fork of Hive and must move to standard Hive versions.","from":"reporter","subject":"Remove references to org.spark-project.hive"},{"body":"I see wide agreement on that. One question I have is, is including Hive this way merely a really-not-nice-to-have or actually not allowed? I think the question is whether sources are available, right? because releases can't have binary-only parts. I plead ignorance, I have never myself paid much attention to this integration. \n\nIf it's not then this sounds like something has to change for releases beyond 2.1.1 and this can be targeted as a Blocker accordingly.\n\nDoes this depend on refactoring or changes in Hive? IIRC the problem was hive-exec being an uber-jar, but it's been a long time since I read any of that discussion.","from":"developer"},{"body":"As an Apache member, the Spark project can't release binary artifacts that aren't made from its Apache code base. So either, the Spark project needs to use Hive's release artifacts or it needs to formally fork Hive and move the fork into its git repository at Apache and rename it away from org.apache.hive to org.apache.spark. The current path is not allowed.\n\nHive is in the middle of rolling releases and thus this is a good time to make requests. The old uber jar (hive-exec) is already released separately with the classifier \"core.\" It looks like we are using the same protobuf (2.5.0) and kryo (3.0.3) versions.","from":"developer"},{"body":"Agree. I think the logic was that Spark had released its own source/binary version of Hive, and then used that in Spark. I don't think anybody believes that's a good solution in the long term; it was a work-around for hive-exec's packaging IIRC. Once whatever that is is resolved this can go away, but I defer to those who know the issue better on the details.\n\nWhat I'm not clear on is whether the current org.spark-hive situation is streeetching the source/binary policy so far that it breaks, enough that no more releases can happen without it. Best to make it go away ASAP anyway. But I don't know if changes in Hive 2.5 help integration with Hive 1.x. It may require either temporarily blessing the fork, or more jar surgery to un-uberize the hive-exec jar or something.","from":"developer"},{"body":"I should also say here that the Hive community is willing to help. We are in the process of rolling releases so if Spark needs a change, we can work together to get this done.","from":"developer"},{"body":"It is against Apache policy to release binaries that aren't part of your project.","from":"developer"},{"body":"Alrighty, you can leave the status for now, but generally committers set Blocker. I'm not entirely clear this blocks a release, not yet.\n\nYou're absolutely right, but, the hive fork with binaries and source is part of this project. At least, that's the idea. For example, this is notionally voted on and released with each Spark release, but the binary/source of this fork project isn't separately, explicitly, voted on and separately released. I think that should occur for avoidance of doubt, that this is a blessed artifact of the Spark project. Would this answer your process and policy concerns about the release? It's not pretty but I think that's within the law.\n\nOf course, it's no answer in the long term. The goal is to not have to use the fork at all. If Hive packaging changes are already in place to make it unnecessary, great (is that all there is to it, everyone?) I don't know if that presents a solution for earlier versions of Hive. This fork thing may persist in existing branches, but it has to at least be released and used in a proper way. This may need fixes right now.","from":"developer"},{"body":"[~marmbrus] as release manager of the moment, I suggest we actually formally vote on release the org.spark-project.hive artifact, as I'm not clear we ever did formally. That much seems like a must-have. I don't know that it requires re-releasing the artifacts, but at least having a meaningful review of what it is, and agreeing (or not) that it's what the PMC wants to release, would I believe resolve doubts about the legitimacy of that artifact.\n\nThere's still more to do no doubt, to get rid of the fork. This might include seeing if Hive 1.2.x can provide an un-uberized artifact.","from":"developer"},{"body":"# the ugliness need to inset the spark thrift stuff under the hive thrift stuff is obsolete, can be cut entirely.\n# with the shading of kryo not needed, an unshaded hive *may* work. I forget which troublespots there were last time, probably the usual suspects: jackson, guava, etc.\n# Hive 1.2.x refuses to work with Hadoop 3; it considers that an unsupported version. For basic client-side testing, you can build Hadoop 3 with a fake version (e..g {{mvn install -DskipShade -Ddeclared.hadoop.version=2.11}}, but as hadoop version is something which NN/DNs care about, not something that's really going to work in real systems. Presumably later hive versions will address that.\n\nIf hive take over ownership of the spark 1.2.1-spark branch, this could be done first simply by pulling the spark branch into the Hive repo as a branch, defining the artifact naming properly and releasing it. If that is done, before any release of that 1.2.x branch is done, there's a couple of outstanding PRs to pull in (groovy version for security reasons, ... ).. A quick import & re-release would be the fast way to get this out as an asf-approved binary","from":"developer"},{"body":"+1 for a release of the Spark fork from the Hive community. While the reasons for the fork appear to be fixed in the latest version, there's a lot of work to do to get Spark on a newer Hive version. And for patch releases like 2.0.3 and 2.1.1, I don't think updating Hive is an option.\n\nI'm also all for getting master on a real Hive release. A release of alternate Hive binaries was inappropriate. I think that if a third-party organization had done the same, it would be entirely reasonable to treat it as a trademark violation and ask them to stop.\n\nhttp://www.apache.org/foundation/marks/faq/#products","from":"developer"},{"body":"Yes this is really important. The proper way to do this is to publish a proper version of Hive with the right dependency declared (rather than including all the dependencies in a uber jar). Looks like there are broad support to do this. I'm going to create a JIRA ticket on Hive and add a dependency on this.\n\nThis ticket will depend on that.","from":"developer"},{"body":"I've created a ticket on the Hive side to publish 1.2.x: https://issues.apache.org/jira/browse/HIVE-16391\n\nUntil that is resolved, I also wonder if there are other things we should do. For example, vote on the current fork to rectify it?\n\n","from":"developer"},{"body":"One thing I do recall as trouble here was that ivy resolution was different from mvns, and fixing up all the transitives was a troublespot. Patches here need to be tested against SBT and maven —as jenkins only does SBT, the mvn builds will have to be manual. I don't remember which specific dependency was the problem.","from":"developer"},{"body":"Would it possible make sense to untarget this from the maintenance releases (1.6.X, 2.0.X, 2.1.X) and instead focus on the future versions?","from":"developer"},{"body":"There are no currently targeted version, are there?\n","from":"developer"},{"body":"Oh right, sorry I was misreading the intent of Affects Version/s.","from":"developer"},{"body":"How about upgrade Hive directly to 2.3.2. In fact, I've completed the initial work and have been running for a few days.\r\n [https://github.com/apache/spark/pull/20659]","from":"developer"},{"body":"Prefer newer Hive also","from":"developer"},{"body":"What is our plan to to fix this issue, are we going to use new Hive version, or we are still stick to 1.2?\r\n\r\nIf we're still stick to 1.2, [~stevel@apache.org] and I will take this issue and make the ball rolling in Hive community.","from":"developer"},{"body":"I think you could split things into two\r\n\r\n# a modified hive 1.2.1.x for hadoop 3, with a new package name in the maven builds. (joy, profiles!) and some work with the hive team to get this officially published by them. Strength: easy for people to backport into shipping 2.2, 2.3 builds just by changing the POM\r\n\r\n# the bigger move to Hive 2. This will be the best for future, but is bound to have more surprises. There's even the possibility that the hive team might have to make some changes too, which isn't impossible if the timelines line up.","from":"developer"},{"body":"OK, for the 1st, I've already started working on it locally. Looks like it is not a big change, only some POM changes are enough, I will submit a patch to Hive community.","from":"developer"},{"body":"Hi all, what do you guys think about replacing it to Hive 2.3.x in the near future (like Spark 3.0.0) given SPARK-23710, and keeping the fork for now?\r\n\r\nLooks [~q79969786] completed the initial try at SPARK-23710 and now it sounds pretty much feasible as an option now although it sounds there are still some investigations; however, I believe that we can focus on getting through if we have the explicit plan here. I think we are mostly all positive on this option as a final goal anyway but I felt like we need to make sure on this.\r\nIf the above can be set as the goal for this JIRA to get rid of the fork completely, \\*I personally think\\* Hive side also can focus on landing other fixes to the more resent versions without diverting the efforts to maintain an old branch.\r\n\r\nUntil then, I think we could probably consider keeping the fork for now and landing some minor fixes if there're some strong reasons for it. For example, Hadoop 3 support is blocked by one liner fix in the fork. \\*I personally think\\* it is the easiest way to land this fix into the fork. I believe this is pretty reasonable.\r\n\r\nWhat do you guys think about this?\r\n\r\n","from":"developer"},{"body":"[~owen.omalley] and [~rxin], what do you think about the suggestion above? I tried to check all other contexts hard at all my best and ^ was my current conclusion to get through this issue mostly smoothly and easily. ","from":"developer"},{"body":"Would you guys please give some thought on this when you guys are available?","from":"developer"},{"body":"kindly ping [~owen.omalley] and [~rxin]. I would like to make a progress further on this since it's blocked for a while but it's pretty important to make up for this affair. ","from":"developer"},{"body":"Hey [~owen.omalley] and [~rxin], I know I see many sensitive things for example the policy stuff frankly; however, this one needs some input from you guys before proceeding further ...","from":"developer"},{"body":"If you want to try and put together a PR that actually does it, that could work too. But note that it's a lot of work to upgrade execution Hive. Probably 10X more work than Hive publishing the exec jar.\r\n\r\n ","from":"developer"},{"body":"I was thinking we target it for 3.0.0 (otherwise 4.0.0 might make sense ... ). It might be a lot of work indeed but I believe this is what we should do as a final goal which we should do anyway sometime ...\r\n\r\nUntill then, I wanted to propose to keep the fork as a temporary solution until 3.0.0 and we target to upgrade it to 2.3.x in 3.0.0 as the goal and target version ... Branch-2.4 will be cut out soon and I think we would go for 3.0.0 for the next release if I am not mistaken.","from":"developer"},{"body":"Yea you can try and see how difficult it is.\r\n\r\n ","from":"developer"},{"body":"[~rxin], there was an initial try above already though which at least made the regression tests we wrote so far passed. I talked with [~q79969786] before and she's willing to finish this. For this, I need more supports from you and other guys to go this way ..\r\n\r\nI get your point too on the other hand. So, do you think we should rather not explicitly target it since it's pretty difficult and we should better let Hive publish 1.2.x first rather then keeping the fork since it's unclear if we make it in 3.0.0?","from":"developer"},{"body":"I am asking this to set the goal for this JIRA as of the current status and make a progress on this. I left some comments here because to me it looked [~q79969786]'s try is kind of a new fact arrived here to be considered.\r\n\r\nIf publishing Hive is still preferred to get through here for any reason, I will help go with Saisai's patch in HIVE-16391. If keeping the fork and upgrade could be set as a goal for now, I will try to help go with Yumming's way and make a fix to the fork.\r\n\r\nWhich one do you prefer?","from":"developer"},{"body":"How like will there be a hive release? HIVE-16391 is still open?\r\n\r\nStay with hive 1.2 will slowly become a big problem for us within a few months...","from":"developer"},{"body":"I think we are unclear about how we are going to deal with this and it's been left open for a while ..\r\n\r\n[~rxin], do you maybe have some preference in [my comment above|https://issues.apache.org/jira/browse/SPARK-20202?focusedCommentId=16541034&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#comment-16541034]?\r\n\r\n1. Go with Saisai's patch in HIVE-16391\r\n - Publishing Hive 1.2.x could be easier but will give some overhead to Hive side (e.g., maintaining the old branches, for example, backports).\r\n - If I understood correctly, we have less problems (e.g., policy stuff) if we go publishing Hive 1.2.x HIVE-16391\r\n\r\n2. Target the upgrade with [~q79969786]'s fix, and add some fixes to our current fork when there's strong reasons\r\n - It is difficult but [~q79969786] made and completed an initial try about the upgrade. It still need some further investigation (e.g., see [SPARK-23710|https://issues.apache.org/jira/browse/SPARK-23710]) but the try made the regression tests passed at least. She's willing to finish this.\r\n - If we miss the Hive upgrade to 2.3.x in Spark 3.0.0, we should probably target 4.0.0 with upper version of Hive, which I guess make this upgrade even harder.\r\n - Looks we implicitly agree upon this should be the final goal in the long term.\r\n\r\nSee also [~stevel@apache.org]'s [comment above|https://issues.apache.org/jira/browse/SPARK-20202?focusedCommentId=16500560&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#comment-16500560].\r\n\r\nI am re-raising and giving some refreshes here because I personally see:\r\n\r\n- Few facts arrived here since the JIRA was open. So, it looked to me it might be better we consider the possible options again.\r\n- Looks we are quite unclear on this about how we should get through this to me.\r\n- To me, I am sure we need to share and feel in the same way for this JIRA and, it looks I need some more supports from you guys before we go ahead because it'd be a kind of not easily revertible changes. \r\n- Branch-2.4 will be cut out soon and we will go for Spark 3.0.0 if I am not mistaken.\r\n\r\nI know there are many sensitive things going on here; however, please kindly consider and give some inputs. I am sure we all feel that we should resolve this.\r\nLastly, FWIW, I am doing this on my own rather individually if it matters to anyone in any case.","from":"developer"},{"body":"Hi, All.\r\nI set the target version to `3.1.0`. Please join the discussion if you have any concerns.\r\n- https://lists.apache.org/thread.html/eca4e55c717f35f41c029e227fa9be0a7ee2c8a6f378fcce8f9fd4ff@%3Cdev.spark.apache.org%3E","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29936","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29936","from":"developer"},{"body":"Issue resolved by pull request 29936\n[https://github.com/apache/spark/pull/29936]","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29973","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29973","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29973","from":"developer"}],"created":"2017-04-03T10:23:25.000+0000","description":"Spark can't continue to depend on their fork of Hive and must move to standard Hive versions.","issue_id":"13061083","key":"SPARK-20202","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2020-10-05T22:30:15.000+0000","role":"fixed_distractor","summary":"Remove references to org.spark-project.hive"} {"case_id":"13062513","cluster":"DISTRACTOR-SPARK-20256","comments":[{"body":"I am working on a fix and creating simulated test cases for this issue. ","created":"2017-04-07T17:59:24.766+0000"},{"body":"Hi, [~xwu0226].\nIs there any progress?","created":"2017-04-11T20:51:51.177+0000"},{"body":"Yes. I am working on it. \nMy proposal is to revert the SPARK-18050 change, then add a try-catch over externalCatalog.createDatabase(...) and log the error of existing default database from Hive into DEBUG log. \n\nI am trying to create a unit-test case to simulate the permission issue, which I have some difficulty. \n\n","created":"2017-04-11T21:16:21.567+0000"},{"body":"Hi, [~xwu0226].\nAre you still preparing a unit test case?","created":"2017-07-01T21:46:33.773+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18501","created":"2017-07-02T00:45:05.708+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18530","created":"2017-07-04T17:42:06.781+0000"}],"conversations":[{"body":"In a cluster setup with production Hive running, when the user wants to run spark-shell using the production Hive metastore, hive-site.xml is copied to SPARK_HOME/conf. So when spark-shell is being started, it tries to check database existence of \"default\" database from Hive metastore. Yet, since this user may not have READ/WRITE access to the configured Hive warehouse directory done by Hive itself, such permission error will prevent spark-shell or any spark application with Hive support enabled from starting at all. \n\nExample error:\n{code}To adjust logging level use sc.setLogLevel(newLevel). For SparkR, use setLogLevel(newLevel).\njava.lang.IllegalArgumentException: Error while instantiating 'org.apache.spark.sql.hive.HiveSessionState':\n at org.apache.spark.sql.SparkSession$.org$apache$spark$sql$SparkSession$$reflect(SparkSession.scala:981)\n at org.apache.spark.sql.SparkSession.sessionState$lzycompute(SparkSession.scala:110)\n at org.apache.spark.sql.SparkSession.sessionState(SparkSession.scala:109)\n at org.apache.spark.sql.SparkSession$Builder$$anonfun$getOrCreate$5.apply(SparkSession.scala:878)\n at org.apache.spark.sql.SparkSession$Builder$$anonfun$getOrCreate$5.apply(SparkSession.scala:878)\n at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:99)\n at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:99)\n at scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:230)\n at scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:40)\n at scala.collection.mutable.HashMap.foreach(HashMap.scala:99)\n at org.apache.spark.sql.SparkSession$Builder.getOrCreate(SparkSession.scala:878)\n at org.apache.spark.repl.Main$.createSparkSession(Main.scala:95)\n ... 47 elided\nCaused by: java.lang.reflect.InvocationTargetException: org.apache.spark.sql.AnalysisException: org.apache.hadoop.hive.ql.metadata.HiveException: MetaException(message:java.security.AccessControlException: Permission denied: user=notebook, access=READ, inode=\"/apps/hive/warehouse\":hive:hadoop:drwxrwx---\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.check(FSPermissionChecker.java:320)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:219)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:190)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1728)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1712)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPathAccess(FSDirectory.java:1686)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkAccess(FSNamesystem.java:8238)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.checkAccess(NameNodeRpcServer.java:1933)\n\tat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.checkAccess(ClientNamenodeProtocolServerSideTranslatorPB.java:1455)\n\tat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n\tat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2049)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2045)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1697)\n\tat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2045)\n);\n at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n at java.lang.reflect.Constructor.newInstance(Constructor.java:423)\n at org.apache.spark.sql.SparkSession$.org$apache$spark$sql$SparkSession$$reflect(SparkSession.scala:978)\n ... 58 more\nCaused by: org.apache.spark.sql.AnalysisException: org.apache.hadoop.hive.ql.metadata.HiveException: MetaException(message:java.security.AccessControlException: Permission denied: user=notebook, access=READ, inode=\"/apps/hive/warehouse\":hive:hadoop:drwxrwx---\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.check(FSPermissionChecker.java:320)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:219)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:190)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1728)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1712)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPathAccess(FSDirectory.java:1686)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkAccess(FSNamesystem.java:8238)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.checkAccess(NameNodeRpcServer.java:1933)\n\tat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.checkAccess(ClientNamenodeProtocolServerSideTranslatorPB.java:1455)\n\tat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n\tat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2049)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2045)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1697)\n\tat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2045)\n);\n at org.apache.spark.sql.hive.HiveExternalCatalog.withClient(HiveExternalCatalog.scala:98)\n at org.apache.spark.sql.hive.HiveExternalCatalog.databaseExists(HiveExternalCatalog.scala:169)\n at org.apache.spark.sql.internal.SharedState.(SharedState.scala:96)\n at org.apache.spark.sql.SparkSession$$anonfun$sharedState$1.apply(SparkSession.scala:101)\n at org.apache.spark.sql.SparkSession$$anonfun$sharedState$1.apply(SparkSession.scala:101)\n at scala.Option.getOrElse(Option.scala:121)\n at org.apache.spark.sql.SparkSession.sharedState$lzycompute(SparkSession.scala:101)\n at org.apache.spark.sql.SparkSession.sharedState(SparkSession.scala:100)\n at org.apache.spark.sql.internal.SessionState.(SessionState.scala:157)\n at org.apache.spark.sql.hive.HiveSessionState.(HiveSessionState.scala:32)\n ... 63 more\nCaused by: org.apache.hadoop.hive.ql.metadata.HiveException: MetaException(message:java.security.AccessControlException: Permission denied: user=notebook, access=READ, inode=\"/apps/hive/warehouse\":hive:hadoop:drwxrwx---\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.check(FSPermissionChecker.java:320)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:219)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:190)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1728)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1712)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPathAccess(FSDirectory.java:1686)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkAccess(FSNamesystem.java:8238)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.checkAccess(NameNodeRpcServer.java:1933)\n\tat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.checkAccess(ClientNamenodeProtocolServerSideTranslatorPB.java:1455)\n\tat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n\tat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2049)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2045)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1697)\n\tat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2045)\n)\n at org.apache.hadoop.hive.ql.metadata.Hive.getDatabase(Hive.java:1305)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getDatabaseOption$1.apply(HiveClientImpl.scala:336)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getDatabaseOption$1.apply(HiveClientImpl.scala:336)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$withHiveState$1.apply(HiveClientImpl.scala:279)\n at org.apache.spark.sql.hive.client.HiveClientImpl.liftedTree1$1(HiveClientImpl.scala:226)\n at org.apache.spark.sql.hive.client.HiveClientImpl.retryLocked(HiveClientImpl.scala:225)\n at org.apache.spark.sql.hive.client.HiveClientImpl.withHiveState(HiveClientImpl.scala:268)\n at org.apache.spark.sql.hive.client.HiveClientImpl.getDatabaseOption(HiveClientImpl.scala:335)\n at org.apache.spark.sql.hive.HiveExternalCatalog$$anonfun$databaseExists$1.apply$mcZ$sp(HiveExternalCatalog.scala:170)\n at org.apache.spark.sql.hive.HiveExternalCatalog$$anonfun$databaseExists$1.apply(HiveExternalCatalog.scala:170)\n at org.apache.spark.sql.hive.HiveExternalCatalog$$anonfun$databaseExists$1.apply(HiveExternalCatalog.scala:170)\n at org.apache.spark.sql.hive.HiveExternalCatalog.withClient(HiveExternalCatalog.scala:95)\n ... 72 more\nCaused by: org.apache.hadoop.hive.metastore.api.MetaException: java.security.AccessControlException: Permission denied: user=notebook, access=READ, inode=\"/apps/hive/warehouse\":hive:hadoop:drwxrwx---\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.check(FSPermissionChecker.java:320)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:219)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:190)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1728)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1712)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPathAccess(FSDirectory.java:1686)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkAccess(FSNamesystem.java:8238)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.checkAccess(NameNodeRpcServer.java:1933)\n\tat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.checkAccess(ClientNamenodeProtocolServerSideTranslatorPB.java:1455)\n\tat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n\tat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2049)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2045)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1697)\n\tat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2045)\n\n at org.apache.hadoop.hive.metastore.api.ThriftHiveMetastore$get_database_result$get_database_resultStandardScheme.read(ThriftHiveMetastore.java:15345)\n at org.apache.hadoop.hive.metastore.api.ThriftHiveMetastore$get_database_result$get_database_resultStandardScheme.read(ThriftHiveMetastore.java:15313)\n at org.apache.hadoop.hive.metastore.api.ThriftHiveMetastore$get_database_result.read(ThriftHiveMetastore.java:15244)\n at org.apache.thrift.TServiceClient.receiveBase(TServiceClient.java:78)\n at org.apache.hadoop.hive.metastore.api.ThriftHiveMetastore$Client.recv_get_database(ThriftHiveMetastore.java:654)\n at org.apache.hadoop.hive.metastore.api.ThriftHiveMetastore$Client.get_database(ThriftHiveMetastore.java:641)\n at org.apache.hadoop.hive.metastore.HiveMetaStoreClient.getDatabase(HiveMetaStoreClient.java:1158)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n at org.apache.hadoop.hive.metastore.RetryingMetaStoreClient.invoke(RetryingMetaStoreClient.java:156)\n at com.sun.proxy.$Proxy19.getDatabase(Unknown Source)\n at org.apache.hadoop.hive.ql.metadata.Hive.getDatabase(Hive.java:1301)\n ... 83 more{code}\n\nThe root cause of this is a regression introduced by SPARK-18050, which tries to avoid annoying error message from Hive complaining about default database already exists. An if-condition was added to check \"default\" database existence before calling createDatabase triggers this permission error. \n\nI think this is a bit of critical because this error prevent spark from being used in Hive support enabled mode. ","from":"reporter","subject":"Fail to start SparkContext/SparkSession with Hive support enabled when user does not have read/write privilege to Hive metastore warehouse dir"},{"body":"I am working on a fix and creating simulated test cases for this issue. ","from":"developer"},{"body":"Hi, [~xwu0226].\nIs there any progress?","from":"developer"},{"body":"Yes. I am working on it. \nMy proposal is to revert the SPARK-18050 change, then add a try-catch over externalCatalog.createDatabase(...) and log the error of existing default database from Hive into DEBUG log. \n\nI am trying to create a unit-test case to simulate the permission issue, which I have some difficulty. \n\n","from":"developer"},{"body":"Hi, [~xwu0226].\nAre you still preparing a unit test case?","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18501","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18530","from":"developer"}],"created":"2017-04-07T17:52:41.000+0000","description":"In a cluster setup with production Hive running, when the user wants to run spark-shell using the production Hive metastore, hive-site.xml is copied to SPARK_HOME/conf. So when spark-shell is being started, it tries to check database existence of \"default\" database from Hive metastore. Yet, since this user may not have READ/WRITE access to the configured Hive warehouse directory done by Hive itself, such permission error will prevent spark-shell or any spark application with Hive support enabled from starting at all. \n\nExample error:\n{code}To adjust logging level use sc.setLogLevel(newLevel). For SparkR, use setLogLevel(newLevel).\njava.lang.IllegalArgumentException: Error while instantiating 'org.apache.spark.sql.hive.HiveSessionState':\n at org.apache.spark.sql.SparkSession$.org$apache$spark$sql$SparkSession$$reflect(SparkSession.scala:981)\n at org.apache.spark.sql.SparkSession.sessionState$lzycompute(SparkSession.scala:110)\n at org.apache.spark.sql.SparkSession.sessionState(SparkSession.scala:109)\n at org.apache.spark.sql.SparkSession$Builder$$anonfun$getOrCreate$5.apply(SparkSession.scala:878)\n at org.apache.spark.sql.SparkSession$Builder$$anonfun$getOrCreate$5.apply(SparkSession.scala:878)\n at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:99)\n at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:99)\n at scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:230)\n at scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:40)\n at scala.collection.mutable.HashMap.foreach(HashMap.scala:99)\n at org.apache.spark.sql.SparkSession$Builder.getOrCreate(SparkSession.scala:878)\n at org.apache.spark.repl.Main$.createSparkSession(Main.scala:95)\n ... 47 elided\nCaused by: java.lang.reflect.InvocationTargetException: org.apache.spark.sql.AnalysisException: org.apache.hadoop.hive.ql.metadata.HiveException: MetaException(message:java.security.AccessControlException: Permission denied: user=notebook, access=READ, inode=\"/apps/hive/warehouse\":hive:hadoop:drwxrwx---\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.check(FSPermissionChecker.java:320)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:219)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:190)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1728)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1712)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPathAccess(FSDirectory.java:1686)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkAccess(FSNamesystem.java:8238)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.checkAccess(NameNodeRpcServer.java:1933)\n\tat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.checkAccess(ClientNamenodeProtocolServerSideTranslatorPB.java:1455)\n\tat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n\tat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2049)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2045)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1697)\n\tat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2045)\n);\n at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n at java.lang.reflect.Constructor.newInstance(Constructor.java:423)\n at org.apache.spark.sql.SparkSession$.org$apache$spark$sql$SparkSession$$reflect(SparkSession.scala:978)\n ... 58 more\nCaused by: org.apache.spark.sql.AnalysisException: org.apache.hadoop.hive.ql.metadata.HiveException: MetaException(message:java.security.AccessControlException: Permission denied: user=notebook, access=READ, inode=\"/apps/hive/warehouse\":hive:hadoop:drwxrwx---\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.check(FSPermissionChecker.java:320)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:219)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:190)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1728)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1712)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPathAccess(FSDirectory.java:1686)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkAccess(FSNamesystem.java:8238)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.checkAccess(NameNodeRpcServer.java:1933)\n\tat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.checkAccess(ClientNamenodeProtocolServerSideTranslatorPB.java:1455)\n\tat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n\tat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2049)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2045)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1697)\n\tat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2045)\n);\n at org.apache.spark.sql.hive.HiveExternalCatalog.withClient(HiveExternalCatalog.scala:98)\n at org.apache.spark.sql.hive.HiveExternalCatalog.databaseExists(HiveExternalCatalog.scala:169)\n at org.apache.spark.sql.internal.SharedState.(SharedState.scala:96)\n at org.apache.spark.sql.SparkSession$$anonfun$sharedState$1.apply(SparkSession.scala:101)\n at org.apache.spark.sql.SparkSession$$anonfun$sharedState$1.apply(SparkSession.scala:101)\n at scala.Option.getOrElse(Option.scala:121)\n at org.apache.spark.sql.SparkSession.sharedState$lzycompute(SparkSession.scala:101)\n at org.apache.spark.sql.SparkSession.sharedState(SparkSession.scala:100)\n at org.apache.spark.sql.internal.SessionState.(SessionState.scala:157)\n at org.apache.spark.sql.hive.HiveSessionState.(HiveSessionState.scala:32)\n ... 63 more\nCaused by: org.apache.hadoop.hive.ql.metadata.HiveException: MetaException(message:java.security.AccessControlException: Permission denied: user=notebook, access=READ, inode=\"/apps/hive/warehouse\":hive:hadoop:drwxrwx---\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.check(FSPermissionChecker.java:320)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:219)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:190)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1728)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1712)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPathAccess(FSDirectory.java:1686)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkAccess(FSNamesystem.java:8238)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.checkAccess(NameNodeRpcServer.java:1933)\n\tat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.checkAccess(ClientNamenodeProtocolServerSideTranslatorPB.java:1455)\n\tat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n\tat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2049)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2045)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1697)\n\tat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2045)\n)\n at org.apache.hadoop.hive.ql.metadata.Hive.getDatabase(Hive.java:1305)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getDatabaseOption$1.apply(HiveClientImpl.scala:336)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getDatabaseOption$1.apply(HiveClientImpl.scala:336)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$withHiveState$1.apply(HiveClientImpl.scala:279)\n at org.apache.spark.sql.hive.client.HiveClientImpl.liftedTree1$1(HiveClientImpl.scala:226)\n at org.apache.spark.sql.hive.client.HiveClientImpl.retryLocked(HiveClientImpl.scala:225)\n at org.apache.spark.sql.hive.client.HiveClientImpl.withHiveState(HiveClientImpl.scala:268)\n at org.apache.spark.sql.hive.client.HiveClientImpl.getDatabaseOption(HiveClientImpl.scala:335)\n at org.apache.spark.sql.hive.HiveExternalCatalog$$anonfun$databaseExists$1.apply$mcZ$sp(HiveExternalCatalog.scala:170)\n at org.apache.spark.sql.hive.HiveExternalCatalog$$anonfun$databaseExists$1.apply(HiveExternalCatalog.scala:170)\n at org.apache.spark.sql.hive.HiveExternalCatalog$$anonfun$databaseExists$1.apply(HiveExternalCatalog.scala:170)\n at org.apache.spark.sql.hive.HiveExternalCatalog.withClient(HiveExternalCatalog.scala:95)\n ... 72 more\nCaused by: org.apache.hadoop.hive.metastore.api.MetaException: java.security.AccessControlException: Permission denied: user=notebook, access=READ, inode=\"/apps/hive/warehouse\":hive:hadoop:drwxrwx---\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.check(FSPermissionChecker.java:320)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:219)\n\tat org.apache.hadoop.hdfs.server.namenode.FSPermissionChecker.checkPermission(FSPermissionChecker.java:190)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1728)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPermission(FSDirectory.java:1712)\n\tat org.apache.hadoop.hdfs.server.namenode.FSDirectory.checkPathAccess(FSDirectory.java:1686)\n\tat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkAccess(FSNamesystem.java:8238)\n\tat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.checkAccess(NameNodeRpcServer.java:1933)\n\tat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.checkAccess(ClientNamenodeProtocolServerSideTranslatorPB.java:1455)\n\tat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n\tat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n\tat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2049)\n\tat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2045)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1697)\n\tat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2045)\n\n at org.apache.hadoop.hive.metastore.api.ThriftHiveMetastore$get_database_result$get_database_resultStandardScheme.read(ThriftHiveMetastore.java:15345)\n at org.apache.hadoop.hive.metastore.api.ThriftHiveMetastore$get_database_result$get_database_resultStandardScheme.read(ThriftHiveMetastore.java:15313)\n at org.apache.hadoop.hive.metastore.api.ThriftHiveMetastore$get_database_result.read(ThriftHiveMetastore.java:15244)\n at org.apache.thrift.TServiceClient.receiveBase(TServiceClient.java:78)\n at org.apache.hadoop.hive.metastore.api.ThriftHiveMetastore$Client.recv_get_database(ThriftHiveMetastore.java:654)\n at org.apache.hadoop.hive.metastore.api.ThriftHiveMetastore$Client.get_database(ThriftHiveMetastore.java:641)\n at org.apache.hadoop.hive.metastore.HiveMetaStoreClient.getDatabase(HiveMetaStoreClient.java:1158)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n at org.apache.hadoop.hive.metastore.RetryingMetaStoreClient.invoke(RetryingMetaStoreClient.java:156)\n at com.sun.proxy.$Proxy19.getDatabase(Unknown Source)\n at org.apache.hadoop.hive.ql.metadata.Hive.getDatabase(Hive.java:1301)\n ... 83 more{code}\n\nThe root cause of this is a regression introduced by SPARK-18050, which tries to avoid annoying error message from Hive complaining about default database already exists. An if-condition was added to check \"default\" database existence before calling createDatabase triggers this permission error. \n\nI think this is a bit of critical because this error prevent spark from being used in Hive support enabled mode. ","issue_id":"13062513","key":"SPARK-20256","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-07-04T16:50:02.000+0000","role":"fixed_distractor","summary":"Fail to start SparkContext/SparkSession with Hive support enabled when user does not have read/write privilege to Hive metastore warehouse dir"} {"case_id":"13064425","cluster":"DISTRACTOR-SPARK-20356","comments":[{"body":"really quite dangerous bug","created":"2017-04-18T11:52:44.825+0000"},{"body":"Here is a reproduction in scala:\n{noformat}\nval df1 = Seq((\"a\", 1), (\"b\", 1), (\"c\", 2)).toDF(\"item\", \"group\")\nval df2 = Seq((\"a\", 1), (\"b\", 2), (\"c\", 3)).toDF(\"item\", \"id\")\nval df3 = df1.join(df2, Seq(\"item\")).select($\"id\", $\"group\".as(\"item\")).distinct()\n\ndf3.unpersist()\nval agg_without_cache = df3.groupBy($\"item\").count()\nagg_without_cache.show()\n\ndf3.cache()\nval agg_with_cache = df3.groupBy($\"item\").count()\nagg_with_cache.show()\n{noformat}","created":"2017-04-18T14:18:45.947+0000"},{"body":"Commit 5ed397baa758c29c54a853d3f8fee0ad44e97c14 doesn't seem to have this issue.","created":"2017-04-18T17:17:22.628+0000"},{"body":"[~viirya] [~hvanhovell] [~cloud_fan] [~smilegator]\nI took a quick look and it seems the issue started happening after this [pr|https://github.com/apache/spark/pull/17175]. We are\nchanging the output partitioning information of InMemoryTableScanExec as part of the fix ( id, item -> item, item) causing a missing shuffle in the operators\nabove InMemoryTableScan. Changing to use the child's output partitioning like before fixes the issue. \n\nI am a little new to this code :-) And this is what i have found so far. Hope this helps.\n","created":"2017-04-18T20:45:15.914+0000"},{"body":"[~dkbiswal] Thanks for pinging me. I will look into this.","created":"2017-04-19T00:28:37.055+0000"},{"body":"[~hvanhovell] I can't reproduce it with your example code. They both output:\n{code}\n+----+-----+\n|item|count|\n+----+-----+\n| 1| 2|\n| 2| 1|\n+----+-----+\n{code}\n\nAm I missing something?","created":"2017-04-19T02:02:13.870+0000"},{"body":"I think I found the reason of the issue. I am working on it.","created":"2017-04-19T02:30:32.489+0000"},{"body":"[~viirya] Did you try from spark-shell or from one of our query suites ? I could reproduce it from spark-shell fine. From our query suites i had to force the number of shuffle partitions to reproduce it.\n{code}\ntest(\"cache defect\") {\n withSQLConf(\"spark.sql.shuffle.partitions\" -> \"200\") {\n val df1 = Seq((\"a\", 1), (\"b\", 1), (\"c\", 2)).toDF(\"item\", \"group\")\n val df2 = Seq((\"a\", 1), (\"b\", 2), (\"c\", 3)).toDF(\"item\", \"id\")\n val df3 = df1.join(df2, Seq(\"item\")).select($\"id\", $\"group\".as(\"item\")).distinct()\n\n df3.explain(true)\n\n df3.unpersist()\n val agg_without_cache = df3.groupBy($\"item\").count()\n agg_without_cache.show()\n\n df3.cache()\n val agg_with_cache = df3.groupBy($\"item\").count()\n agg_with_cache.explain(true)\n agg_with_cache.show()\n }\n }\n{code}","created":"2017-04-19T02:31:29.448+0000"},{"body":"[~dkbiswal] Yeah, right. Thanks. We need to force the partition to reproduce this.","created":"2017-04-19T02:34:16.130+0000"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17679","created":"2017-04-19T03:31:02.751+0000"}],"conversations":[{"body":"I'm experiencing a bug with the head version of spark as of 4/17/2017. After joining to dataframes, renaming a column and invoking distinct, the results of the aggregation is incorrect after caching the dataframe. The following code snippet consistently reproduces the error.\n\nfrom pyspark.sql import SparkSession\nimport pyspark.sql.functions as sf\nimport pandas as pd\n\nspark = SparkSession.builder.master(\"local\").appName(\"Word Count\").getOrCreate()\n\nmapping_sdf = spark.createDataFrame(pd.DataFrame([\n {\"ITEM\": \"a\", \"GROUP\": 1},\n {\"ITEM\": \"b\", \"GROUP\": 1},\n {\"ITEM\": \"c\", \"GROUP\": 2}\n]))\n\nitems_sdf = spark.createDataFrame(pd.DataFrame([\n {\"ITEM\": \"a\", \"ID\": 1},\n {\"ITEM\": \"b\", \"ID\": 2},\n {\"ITEM\": \"c\", \"ID\": 3}\n]))\n\nmapped_sdf = \\\n items_sdf.join(mapping_sdf, on='ITEM').select(\"ID\", sf.col(\"GROUP\").alias('ITEM')).distinct()\n\nprint(mapped_sdf.groupBy(\"ITEM\").count().count()) # Prints 2, correct\nmapped_sdf.cache()\nprint(mapped_sdf.groupBy(\"ITEM\").count().count()) # Prints 3, incorrect\n\nThe next code snippet is almost the same after the first except I don't call distinct on the dataframe. This snippet performs as expected:\n\nmapped_sdf = \\\n items_sdf.join(mapping_sdf, on='ITEM').select(\"ID\", sf.col(\"GROUP\").alias('ITEM'))\n\nprint(mapped_sdf.groupBy(\"ITEM\").count().count()) # Prints 2, correct\nmapped_sdf.cache()\nprint(mapped_sdf.groupBy(\"ITEM\").count().count()) # Prints 2, correct\n\nI don't experience this bug with spark 2.1 or event earlier versions for 2.2","from":"reporter","subject":"Spark sql group by returns incorrect results after join + distinct transformations"},{"body":"really quite dangerous bug","from":"developer"},{"body":"Here is a reproduction in scala:\n{noformat}\nval df1 = Seq((\"a\", 1), (\"b\", 1), (\"c\", 2)).toDF(\"item\", \"group\")\nval df2 = Seq((\"a\", 1), (\"b\", 2), (\"c\", 3)).toDF(\"item\", \"id\")\nval df3 = df1.join(df2, Seq(\"item\")).select($\"id\", $\"group\".as(\"item\")).distinct()\n\ndf3.unpersist()\nval agg_without_cache = df3.groupBy($\"item\").count()\nagg_without_cache.show()\n\ndf3.cache()\nval agg_with_cache = df3.groupBy($\"item\").count()\nagg_with_cache.show()\n{noformat}","from":"developer"},{"body":"Commit 5ed397baa758c29c54a853d3f8fee0ad44e97c14 doesn't seem to have this issue.","from":"developer"},{"body":"[~viirya] [~hvanhovell] [~cloud_fan] [~smilegator]\nI took a quick look and it seems the issue started happening after this [pr|https://github.com/apache/spark/pull/17175]. We are\nchanging the output partitioning information of InMemoryTableScanExec as part of the fix ( id, item -> item, item) causing a missing shuffle in the operators\nabove InMemoryTableScan. Changing to use the child's output partitioning like before fixes the issue. \n\nI am a little new to this code :-) And this is what i have found so far. Hope this helps.\n","from":"developer"},{"body":"[~dkbiswal] Thanks for pinging me. I will look into this.","from":"developer"},{"body":"[~hvanhovell] I can't reproduce it with your example code. They both output:\n{code}\n+----+-----+\n|item|count|\n+----+-----+\n| 1| 2|\n| 2| 1|\n+----+-----+\n{code}\n\nAm I missing something?","from":"developer"},{"body":"I think I found the reason of the issue. I am working on it.","from":"developer"},{"body":"[~viirya] Did you try from spark-shell or from one of our query suites ? I could reproduce it from spark-shell fine. From our query suites i had to force the number of shuffle partitions to reproduce it.\n{code}\ntest(\"cache defect\") {\n withSQLConf(\"spark.sql.shuffle.partitions\" -> \"200\") {\n val df1 = Seq((\"a\", 1), (\"b\", 1), (\"c\", 2)).toDF(\"item\", \"group\")\n val df2 = Seq((\"a\", 1), (\"b\", 2), (\"c\", 3)).toDF(\"item\", \"id\")\n val df3 = df1.join(df2, Seq(\"item\")).select($\"id\", $\"group\".as(\"item\")).distinct()\n\n df3.explain(true)\n\n df3.unpersist()\n val agg_without_cache = df3.groupBy($\"item\").count()\n agg_without_cache.show()\n\n df3.cache()\n val agg_with_cache = df3.groupBy($\"item\").count()\n agg_with_cache.explain(true)\n agg_with_cache.show()\n }\n }\n{code}","from":"developer"},{"body":"[~dkbiswal] Yeah, right. Thanks. We need to force the partition to reproduce this.","from":"developer"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17679","from":"developer"}],"created":"2017-04-17T16:30:37.000+0000","description":"I'm experiencing a bug with the head version of spark as of 4/17/2017. After joining to dataframes, renaming a column and invoking distinct, the results of the aggregation is incorrect after caching the dataframe. The following code snippet consistently reproduces the error.\n\nfrom pyspark.sql import SparkSession\nimport pyspark.sql.functions as sf\nimport pandas as pd\n\nspark = SparkSession.builder.master(\"local\").appName(\"Word Count\").getOrCreate()\n\nmapping_sdf = spark.createDataFrame(pd.DataFrame([\n {\"ITEM\": \"a\", \"GROUP\": 1},\n {\"ITEM\": \"b\", \"GROUP\": 1},\n {\"ITEM\": \"c\", \"GROUP\": 2}\n]))\n\nitems_sdf = spark.createDataFrame(pd.DataFrame([\n {\"ITEM\": \"a\", \"ID\": 1},\n {\"ITEM\": \"b\", \"ID\": 2},\n {\"ITEM\": \"c\", \"ID\": 3}\n]))\n\nmapped_sdf = \\\n items_sdf.join(mapping_sdf, on='ITEM').select(\"ID\", sf.col(\"GROUP\").alias('ITEM')).distinct()\n\nprint(mapped_sdf.groupBy(\"ITEM\").count().count()) # Prints 2, correct\nmapped_sdf.cache()\nprint(mapped_sdf.groupBy(\"ITEM\").count().count()) # Prints 3, incorrect\n\nThe next code snippet is almost the same after the first except I don't call distinct on the dataframe. This snippet performs as expected:\n\nmapped_sdf = \\\n items_sdf.join(mapping_sdf, on='ITEM').select(\"ID\", sf.col(\"GROUP\").alias('ITEM'))\n\nprint(mapped_sdf.groupBy(\"ITEM\").count().count()) # Prints 2, correct\nmapped_sdf.cache()\nprint(mapped_sdf.groupBy(\"ITEM\").count().count()) # Prints 2, correct\n\nI don't experience this bug with spark 2.1 or event earlier versions for 2.2","issue_id":"13064425","key":"SPARK-20356","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-04-19T08:03:04.000+0000","role":"fixed_distractor","summary":"Spark sql group by returns incorrect results after join + distinct transformations"} {"case_id":"13066960","cluster":"DISTRACTOR-SPARK-20466","comments":[{"body":"NPE_log describes the detailed info.","created":"2017-04-26T07:15:55.622+0000"},{"body":"[~kellyzly] How to reproduce it?","created":"2017-06-21T04:45:47.431+0000"},{"body":"[~q79969786]: found the NPE when runing Hive on Spark on TPCx-BB. But i did not remember to run which query to find this exception. But it is better to add a judge( if conf != null ) in [code|https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/rdd/HadoopRDD.scala#L374]","created":"2017-06-21T04:53:26.118+0000"},{"body":"Hm, I think the question is how the JobConf is ever null here. I think adding a null check here would only be a band-aid, or at least, something that would need to be taken care of consistently across many more classes.","created":"2017-06-21T10:57:22.885+0000"},{"body":"I just hit this issue in Hive-on-Spark when running some TPC-DS queries. It seems to be intermittent, re-tries of the task succeed. I have a very similar stack trace:\n\n{code}\njava.lang.NullPointerException\n at org.apache.spark.rdd.HadoopRDD$.addLocalConfiguration(HadoopRDD.scala:364)\n at org.apache.spark.rdd.HadoopRDD$$anon$1.<init>(HadoopRDD.scala:238)\n at org.apache.spark.rdd.HadoopRDD.compute(HadoopRDD.scala:211)\n at org.apache.spark.rdd.HadoopRDD.compute(HadoopRDD.scala:101)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:242)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\n at java.lang.Thread.run(Thread.java:748)\n{code}\n\nThe {{JobConf}} object can be {{null}} if {{HadoopRDD#getJobConf}} returns {{null}}. Looks like there is a race condition in {{#getJobConf}} [here|https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/rdd/HadoopRDD.scala#L160]. The method {{HadoopRDD.containsCachedMetadata}} looks into an internal metadata cache - {{SparkEnv#hadoopJobMetadata}}. This cache uses soft references, so the JVM may reclaim entries from the map whenever there is some GC pressure. In which case, any get request on the key will return a {{null}}. The race condition is that the {{#getJobConf}} method first checks if the cache contains the key, and then retrieves it. In between the {{containsKey}} and {{get}} its possible the the key is GCed by the JVM. This would cause {{#getJobConf}} to return {{null}}.\n\nThe fix should be pretty simple, don't use the {{containsKey(key)}} method on the cache, just run a {{get(key)}} and check if it returns {{null}} or not.\n\nHappy to create a PR if other agrees with my analysis.","created":"2017-09-18T21:32:14.467+0000"},{"body":"[~stakiar]: \nthis exception happened on which query of tpcds? I found in another benchmark test(TPCx-BB)\n{quote}\nThis cache uses soft references, so the JVM may reclaim entries from the map whenever there is some GC pressure. In which case, any get request on the key will return a null. The race condition is that the #getJobConf method first checks if the cache contains the key, and then retrieves. In between the containsKey and get its possible the the key is GCed by the JVM. \n{quote}\nthis exception is because {{HadoopRDD.containsCachedMetadata(jobConfCacheKey)}} returns soft reference and it will return {{null}} when GC happens? If it changes to \n{code}\n else if ( HadoopRDD.getCachedMetadata(jobConfCacheKey) != null) {\n logDebug(\"Re-using cached JobConf\")\n HadoopRDD.getCachedMetadata(jobConfCacheKey).asInstanceOf[JobConf]\n }\n{code}\n HadoopRDD.getCachedMetadata(jobConfCacheKey) will not return null if GC happens?\n\n\n\n","created":"2017-09-19T02:17:24.978+0000"},{"body":"I can't remember exactly which query, I was running a chain of about 10 TPC-DS queries, all in the same HoS session. It was a 1 TB Parquet dataset though.\n\nThe exception is because {{HadoopRDD.containsCachedMetadata}} can return {{true}}, but then a future call to {{HadoopRDD.getCachedMetadata}} on the same key, can return {{null}}. This can happen if the JVM decided to GC some of the entires in the metadata cache (which it can since [soft references|https://docs.oracle.com/javase/7/docs/api/java/lang/ref/SoftReference.html] are used).\n\nWe would have to change it to something like:\n\n{code}\n} else {\n Object conf = HadoopRDD.getCachedMetadata(jobConfCacheKey)\n if (conf != null) {\n logDebug(\"Re-using cached JobConf\")\n HadoopRDD.getCachedMetadata(jobConfCacheKey).asInstanceOf[JobConf]\n }\n}\n{code}\n\nI'm not sure how to write it exactly in Scala, but something like that. Once you create {{Object conf}} and point it to the result of {{HadoopRDD.getCachedMetdata(jobConfCacheKey)}}, you then have a hard reference to the object and it can no longer be GCd.","created":"2017-09-20T04:51:40.672+0000"},{"body":"[~srowen], [~vanzin] does my analysis of this bug make sense? If you agree this is a bug, I can make a PR to fix it.","created":"2017-09-25T18:08:21.598+0000"},{"body":"Explanation makes sense, but your proposed solution is also race-prone (two calls to {{getCachedMetadata}}). You want something like:\n\n{code}\nOption(HadoopRDD.getCachedMetadata(jobConfCacheKey)).getOrElse( /* recover from when there's no cached metadata */\n{code}\n","created":"2017-09-25T18:19:41.986+0000"},{"body":"[~vanzin] thanks for taking a look. Yes, you are right. I'll start working on a PR.","created":"2017-09-25T18:37:06.671+0000"},{"body":"User 'sahilTakiar' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19413","created":"2017-10-02T21:04:03.929+0000"}],"conversations":[{"body":"in spark2.0.2, it throws NPE\n{code}\n 17/04/23 08:19:55 ERROR executor.Executor: Exception in task 439.0 in stage 16.0 (TID 986)$ \njava.lang.NullPointerException$\n^Iat org.apache.spark.rdd.HadoopRDD$.addLocalConfiguration(HadoopRDD.scala:373)$\n^Iat org.apache.spark.rdd.HadoopRDD$$anon$1.(HadoopRDD.scala:243)$\n^Iat org.apache.spark.rdd.HadoopRDD.compute(HadoopRDD.scala:208)$\n^Iat org.apache.spark.rdd.HadoopRDD.compute(HadoopRDD.scala:101)$\n^Iat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)$\n^Iat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)$\n^Iat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)$\n^Iat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)$\n^Iat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)$\n^Iat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)$\n^Iat org.apache.spark.scheduler.Task.run(Task.scala:86)$\n^Iat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)$\n^Iat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)$\n^Iat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)$\n^Iat java.lang.Thread.run(Thread.java:745)$\n{code}\n\nsuggestion to add some code to avoid NPE\n\n{code} \n\n /** Add Hadoop configuration specific to a single partition and attempt. */\n def addLocalConfiguration(jobTrackerId: String, jobId: Int, splitId: Int, attemptId: Int,\n conf: JobConf) {\n val jobID = new JobID(jobTrackerId, jobId)\n val taId = new TaskAttemptID(new TaskID(jobID, TaskType.MAP, splitId), attemptId)\n if ( conf != null){\n conf.set(\"mapred.tip.id\", taId.getTaskID.toString)\n conf.set(\"mapred.task.id\", taId.toString)\n conf.setBoolean(\"mapred.task.is.map\", true)\n conf.setInt(\"mapred.task.partition\", splitId)\n conf.set(\"mapred.job.id\", jobID.toString)\n }\n }\n\n\n{code}","from":"reporter","subject":"HadoopRDD#addLocalConfiguration throws NPE"},{"body":"NPE_log describes the detailed info.","from":"developer"},{"body":"[~kellyzly] How to reproduce it?","from":"developer"},{"body":"[~q79969786]: found the NPE when runing Hive on Spark on TPCx-BB. But i did not remember to run which query to find this exception. But it is better to add a judge( if conf != null ) in [code|https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/rdd/HadoopRDD.scala#L374]","from":"developer"},{"body":"Hm, I think the question is how the JobConf is ever null here. I think adding a null check here would only be a band-aid, or at least, something that would need to be taken care of consistently across many more classes.","from":"developer"},{"body":"I just hit this issue in Hive-on-Spark when running some TPC-DS queries. It seems to be intermittent, re-tries of the task succeed. I have a very similar stack trace:\n\n{code}\njava.lang.NullPointerException\n at org.apache.spark.rdd.HadoopRDD$.addLocalConfiguration(HadoopRDD.scala:364)\n at org.apache.spark.rdd.HadoopRDD$$anon$1.<init>(HadoopRDD.scala:238)\n at org.apache.spark.rdd.HadoopRDD.compute(HadoopRDD.scala:211)\n at org.apache.spark.rdd.HadoopRDD.compute(HadoopRDD.scala:101)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:306)\n at org.apache.spark.rdd.RDD.iterator(RDD.scala:270)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:73)\n at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:242)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\n at java.lang.Thread.run(Thread.java:748)\n{code}\n\nThe {{JobConf}} object can be {{null}} if {{HadoopRDD#getJobConf}} returns {{null}}. Looks like there is a race condition in {{#getJobConf}} [here|https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/rdd/HadoopRDD.scala#L160]. The method {{HadoopRDD.containsCachedMetadata}} looks into an internal metadata cache - {{SparkEnv#hadoopJobMetadata}}. This cache uses soft references, so the JVM may reclaim entries from the map whenever there is some GC pressure. In which case, any get request on the key will return a {{null}}. The race condition is that the {{#getJobConf}} method first checks if the cache contains the key, and then retrieves it. In between the {{containsKey}} and {{get}} its possible the the key is GCed by the JVM. This would cause {{#getJobConf}} to return {{null}}.\n\nThe fix should be pretty simple, don't use the {{containsKey(key)}} method on the cache, just run a {{get(key)}} and check if it returns {{null}} or not.\n\nHappy to create a PR if other agrees with my analysis.","from":"developer"},{"body":"[~stakiar]: \nthis exception happened on which query of tpcds? I found in another benchmark test(TPCx-BB)\n{quote}\nThis cache uses soft references, so the JVM may reclaim entries from the map whenever there is some GC pressure. In which case, any get request on the key will return a null. The race condition is that the #getJobConf method first checks if the cache contains the key, and then retrieves. In between the containsKey and get its possible the the key is GCed by the JVM. \n{quote}\nthis exception is because {{HadoopRDD.containsCachedMetadata(jobConfCacheKey)}} returns soft reference and it will return {{null}} when GC happens? If it changes to \n{code}\n else if ( HadoopRDD.getCachedMetadata(jobConfCacheKey) != null) {\n logDebug(\"Re-using cached JobConf\")\n HadoopRDD.getCachedMetadata(jobConfCacheKey).asInstanceOf[JobConf]\n }\n{code}\n HadoopRDD.getCachedMetadata(jobConfCacheKey) will not return null if GC happens?\n\n\n\n","from":"developer"},{"body":"I can't remember exactly which query, I was running a chain of about 10 TPC-DS queries, all in the same HoS session. It was a 1 TB Parquet dataset though.\n\nThe exception is because {{HadoopRDD.containsCachedMetadata}} can return {{true}}, but then a future call to {{HadoopRDD.getCachedMetadata}} on the same key, can return {{null}}. This can happen if the JVM decided to GC some of the entires in the metadata cache (which it can since [soft references|https://docs.oracle.com/javase/7/docs/api/java/lang/ref/SoftReference.html] are used).\n\nWe would have to change it to something like:\n\n{code}\n} else {\n Object conf = HadoopRDD.getCachedMetadata(jobConfCacheKey)\n if (conf != null) {\n logDebug(\"Re-using cached JobConf\")\n HadoopRDD.getCachedMetadata(jobConfCacheKey).asInstanceOf[JobConf]\n }\n}\n{code}\n\nI'm not sure how to write it exactly in Scala, but something like that. Once you create {{Object conf}} and point it to the result of {{HadoopRDD.getCachedMetdata(jobConfCacheKey)}}, you then have a hard reference to the object and it can no longer be GCd.","from":"developer"},{"body":"[~srowen], [~vanzin] does my analysis of this bug make sense? If you agree this is a bug, I can make a PR to fix it.","from":"developer"},{"body":"Explanation makes sense, but your proposed solution is also race-prone (two calls to {{getCachedMetadata}}). You want something like:\n\n{code}\nOption(HadoopRDD.getCachedMetadata(jobConfCacheKey)).getOrElse( /* recover from when there's no cached metadata */\n{code}\n","from":"developer"},{"body":"[~vanzin] thanks for taking a look. Yes, you are right. I'll start working on a PR.","from":"developer"},{"body":"User 'sahilTakiar' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19413","from":"developer"}],"created":"2017-04-26T07:13:02.000+0000","description":"in spark2.0.2, it throws NPE\n{code}\n 17/04/23 08:19:55 ERROR executor.Executor: Exception in task 439.0 in stage 16.0 (TID 986)$ \njava.lang.NullPointerException$\n^Iat org.apache.spark.rdd.HadoopRDD$.addLocalConfiguration(HadoopRDD.scala:373)$\n^Iat org.apache.spark.rdd.HadoopRDD$$anon$1.(HadoopRDD.scala:243)$\n^Iat org.apache.spark.rdd.HadoopRDD.compute(HadoopRDD.scala:208)$\n^Iat org.apache.spark.rdd.HadoopRDD.compute(HadoopRDD.scala:101)$\n^Iat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)$\n^Iat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)$\n^Iat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)$\n^Iat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:319)$\n^Iat org.apache.spark.rdd.RDD.iterator(RDD.scala:283)$\n^Iat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:70)$\n^Iat org.apache.spark.scheduler.Task.run(Task.scala:86)$\n^Iat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:274)$\n^Iat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)$\n^Iat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)$\n^Iat java.lang.Thread.run(Thread.java:745)$\n{code}\n\nsuggestion to add some code to avoid NPE\n\n{code} \n\n /** Add Hadoop configuration specific to a single partition and attempt. */\n def addLocalConfiguration(jobTrackerId: String, jobId: Int, splitId: Int, attemptId: Int,\n conf: JobConf) {\n val jobID = new JobID(jobTrackerId, jobId)\n val taId = new TaskAttemptID(new TaskID(jobID, TaskType.MAP, splitId), attemptId)\n if ( conf != null){\n conf.set(\"mapred.tip.id\", taId.getTaskID.toString)\n conf.set(\"mapred.task.id\", taId.toString)\n conf.setBoolean(\"mapred.task.is.map\", true)\n conf.setInt(\"mapred.task.partition\", splitId)\n conf.set(\"mapred.job.id\", jobID.toString)\n }\n }\n\n\n{code}","issue_id":"13066960","key":"SPARK-20466","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-10-03T23:57:14.000+0000","role":"fixed_distractor","summary":"HadoopRDD#addLocalConfiguration throws NPE"} {"case_id":"13068377","cluster":"DISTRACTOR-SPARK-20555","comments":[{"body":"User 'gaborfeher' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17830","created":"2017-05-02T10:26:03.753+0000"},{"body":"Hi,\n\nMaybe it was not clear from the title, but this issue is causing data (precision) loss or crashes when reading from Oracle Databases via Spark. I also have a patch!","created":"2017-05-16T13:34:26.420+0000"},{"body":"User 'gatorsmile' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18408","created":"2017-06-23T17:55:05.267+0000"},{"body":"[~gfeher]\r\nMay i know if first issue has been resolved?\r\n\r\n\"1. DECIMAL(1) becomes BooleanType\r\nIn Orcale, a DECIMAL(1) can have values from -9 to 9.\"\r\n\r\nI am using the spark 2.2.0 but still i am getting Boolean \"false\" when source is having NUMBER(1) as 0. I want it as 0 without customschema Could you please advise?","created":"2017-12-07T16:47:05.793+0000"},{"body":"This issues was fixed, so in theory, you should not be getting Boolean type in Spark if the database has a NUMBER(1) column.\r\n\r\nI checked the code and it seems to be that the fix is still there in Spark 2.2.0:\r\nFix: https://github.com/apache/spark/pull/18408/files\r\nSpark 2.2.0 fixed code: https://github.com/apache/spark/blob/v2.2.0/sql/core/src/main/scala/org/apache/spark/sql/jdbc/OracleDialect.scala#L33\r\nSpark 2.2.0 tests: https://github.com/apache/spark/blob/v2.2.0/external/docker-integration-tests/src/test/scala/org/apache/spark/sql/jdbc/OracleIntegrationSuite.scala#L108\r\n\r\nI suggest you the following: double-check that you have a clean Spark 2.2.0 setup, i.e. there are no remnants of earlier Spark versions loaded with your program. If you have confirmed that, then create an as simple as possible step by step guide how to observe your problem, so that another person can do it easily. Perhaps the best would be to demonstrate your problem in a cleanly installed Spark shell.","created":"2017-12-09T00:30:02.115+0000"}],"conversations":[{"body":"When querying an Oracle database, Spark maps some Oracle numeric data types to incorrect Catalyst data types:\n1. DECIMAL(1) becomes BooleanType\nIn Orcale, a DECIMAL(1) can have values from -9 to 9.\nIn Spark now, values larger than 1 become the boolean value true.\n2. DECIMAL(3,2) becomes IntegerType\nIn Oracle, a DECIMAL(2) can have values like 1.23\nIn Spark now, digits after the decimal point are dropped.\n3. DECIMAL(10) becomes IntegerType\nIn Oracle, a DECIMAL(10) can have the value 9999999999 (ten nines), which is more than 2^31\nSpark throws an exception: \"java.sql.SQLException: Numeric Overflow\"\n\nI think the best solution is to always keep Oracle's decimal types. (In theory we could introduce a FloatType in some case of #2, and fix #3 by only introducing IntegerType for DECIMAL(9). But in my opinion, that would end up complicated and error-prone.)\n\nNote: I think the above problems were introduced as part of https://github.com/apache/spark/pull/14377\nThe main purpose of that PR seems to be converting Spark types to correct Oracle types, and that part seems good to me. But it also adds the inverse conversions. As it turns out in the above examples, that is not possible.","from":"reporter","subject":"Incorrect handling of Oracle's decimal types via JDBC"},{"body":"User 'gaborfeher' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17830","from":"developer"},{"body":"Hi,\n\nMaybe it was not clear from the title, but this issue is causing data (precision) loss or crashes when reading from Oracle Databases via Spark. I also have a patch!","from":"developer"},{"body":"User 'gatorsmile' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18408","from":"developer"},{"body":"[~gfeher]\r\nMay i know if first issue has been resolved?\r\n\r\n\"1. DECIMAL(1) becomes BooleanType\r\nIn Orcale, a DECIMAL(1) can have values from -9 to 9.\"\r\n\r\nI am using the spark 2.2.0 but still i am getting Boolean \"false\" when source is having NUMBER(1) as 0. I want it as 0 without customschema Could you please advise?","from":"developer"},{"body":"This issues was fixed, so in theory, you should not be getting Boolean type in Spark if the database has a NUMBER(1) column.\r\n\r\nI checked the code and it seems to be that the fix is still there in Spark 2.2.0:\r\nFix: https://github.com/apache/spark/pull/18408/files\r\nSpark 2.2.0 fixed code: https://github.com/apache/spark/blob/v2.2.0/sql/core/src/main/scala/org/apache/spark/sql/jdbc/OracleDialect.scala#L33\r\nSpark 2.2.0 tests: https://github.com/apache/spark/blob/v2.2.0/external/docker-integration-tests/src/test/scala/org/apache/spark/sql/jdbc/OracleIntegrationSuite.scala#L108\r\n\r\nI suggest you the following: double-check that you have a clean Spark 2.2.0 setup, i.e. there are no remnants of earlier Spark versions loaded with your program. If you have confirmed that, then create an as simple as possible step by step guide how to observe your problem, so that another person can do it easily. Perhaps the best would be to demonstrate your problem in a cleanly installed Spark shell.","from":"developer"}],"created":"2017-05-02T10:19:48.000+0000","description":"When querying an Oracle database, Spark maps some Oracle numeric data types to incorrect Catalyst data types:\n1. DECIMAL(1) becomes BooleanType\nIn Orcale, a DECIMAL(1) can have values from -9 to 9.\nIn Spark now, values larger than 1 become the boolean value true.\n2. DECIMAL(3,2) becomes IntegerType\nIn Oracle, a DECIMAL(2) can have values like 1.23\nIn Spark now, digits after the decimal point are dropped.\n3. DECIMAL(10) becomes IntegerType\nIn Oracle, a DECIMAL(10) can have the value 9999999999 (ten nines), which is more than 2^31\nSpark throws an exception: \"java.sql.SQLException: Numeric Overflow\"\n\nI think the best solution is to always keep Oracle's decimal types. (In theory we could introduce a FloatType in some case of #2, and fix #3 by only introducing IntegerType for DECIMAL(9). But in my opinion, that would end up complicated and error-prone.)\n\nNote: I think the above problems were introduced as part of https://github.com/apache/spark/pull/14377\nThe main purpose of that PR seems to be converting Spark types to correct Oracle types, and that part seems good to me. But it also adds the inverse conversions. As it turns out in the above examples, that is not possible.","issue_id":"13068377","key":"SPARK-20555","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-06-24T05:11:01.000+0000","role":"fixed_distractor","summary":"Incorrect handling of Oracle's decimal types via JDBC"} {"case_id":"13070377","cluster":"DISTRACTOR-SPARK-20680","comments":[{"body":"[~jiangxb] Do you have time to work on this?","created":"2017-05-10T15:56:22.916+0000"},{"body":"[~hvanhovell]Sure, I'll look at this issue.","created":"2017-05-11T09:52:13.605+0000"},{"body":"User 'LantaoJin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17953","created":"2017-05-11T12:04:02.861+0000"},{"body":"Hi [~jiangxb1987] and [~hvanhovell], I fixed with a very simple way. Could you help to review this?\nBelow is the testing after patched.\n{quote}\nspark-sql> describe bad;\nx int NULL\nz null NULL\nTime taken: 0.486 seconds, Fetched 2 row(s)\n{quote}","created":"2017-05-11T12:11:35.279+0000"},{"body":"User 'LantaoJin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28833","created":"2020-06-15T08:51:45.586+0000"},{"body":"User 'LantaoJin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28833","created":"2020-06-15T08:52:18.000+0000"},{"body":"User 'LantaoJin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28935","created":"2020-06-28T02:31:33.503+0000"},{"body":"User 'LantaoJin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28935","created":"2020-06-28T02:31:45.506+0000"},{"body":"Issue resolved by pull request 28833\n[https://github.com/apache/spark/pull/28833]","created":"2020-07-08T01:58:41.162+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29041","created":"2020-07-08T13:56:39.509+0000"},{"body":"User 'ulysses-you' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29244","created":"2020-07-27T12:25:43.855+0000"},{"body":"User 'ulysses-you' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29423","created":"2020-08-13T05:13:55.574+0000"}],"conversations":[{"body":"Create a HIVE view:\n{quote}\nhive> create table bad as select 1 x, null z from dual;\n{quote}\n\nBecause there's no type, Hive gives it the VOID type:\n{quote}\nhive> describe bad;\nOK\nx\tint\t\nz\tvoid\n{quote}\n\nIn Spark2.0.x, the behaviour to read this view is normal:\n{quote}\nspark-sql> describe bad;\nx int NULL\nz void NULL\nTime taken: 4.431 seconds, Fetched 2 row(s)\n{quote}\n\nBut in Spark2.1.x, it failed with SparkException: Cannot recognize hive type string: void\n{quote}\nspark-sql> describe bad;\n17/05/09 03:12:08 INFO execution.SparkSqlParser: Parsing command: describe bad\n17/05/09 03:12:08 INFO parser.CatalystSqlParser: Parsing command: int\n17/05/09 03:12:08 INFO parser.CatalystSqlParser: Parsing command: void\n17/05/09 03:12:08 ERROR thriftserver.SparkSQLDriver: Failed in [describe bad]\norg.apache.spark.SparkException: Cannot recognize hive type string: void\n at org.apache.spark.sql.hive.client.HiveClientImpl.org$apache$spark$sql$hive$client$HiveClientImpl$$fromHiveColumn(HiveClientImpl.scala:789)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getTableOption$1$$anonfun$apply$11$$anonfun$7.apply(HiveClientImpl.scala:365) \n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getTableOption$1$$anonfun$apply$11$$anonfun$7.apply(HiveClientImpl.scala:365) \n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.Iterator$class.foreach(Iterator.scala:893)\n at scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\n at scala.collection.IterableLike$class.foreach(IterableLike.scala:72)\n at scala.collection.AbstractIterable.foreach(Iterable.scala:54)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.AbstractTraversable.map(Traversable.scala:104)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getTableOption$1$$anonfun$apply$11.apply(HiveClientImpl.scala:365)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getTableOption$1$$anonfun$apply$11.apply(HiveClientImpl.scala:361)\n\nCaused by: org.apache.spark.sql.catalyst.parser.ParseException:\nDataType void() is not supported.(line 1, pos 0)\n\n== SQL == \nvoid \n^^^\n\n ... 61 more\norg.apache.spark.SparkException: Cannot recognize hive type string: void\n\n\n{quote}\n","from":"reporter","subject":"Spark-sql do not support for void column datatype of view"},{"body":"[~jiangxb] Do you have time to work on this?","from":"developer"},{"body":"[~hvanhovell]Sure, I'll look at this issue.","from":"developer"},{"body":"User 'LantaoJin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17953","from":"developer"},{"body":"Hi [~jiangxb1987] and [~hvanhovell], I fixed with a very simple way. Could you help to review this?\nBelow is the testing after patched.\n{quote}\nspark-sql> describe bad;\nx int NULL\nz null NULL\nTime taken: 0.486 seconds, Fetched 2 row(s)\n{quote}","from":"developer"},{"body":"User 'LantaoJin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28833","from":"developer"},{"body":"User 'LantaoJin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28833","from":"developer"},{"body":"User 'LantaoJin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28935","from":"developer"},{"body":"User 'LantaoJin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28935","from":"developer"},{"body":"Issue resolved by pull request 28833\n[https://github.com/apache/spark/pull/28833]","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29041","from":"developer"},{"body":"User 'ulysses-you' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29244","from":"developer"},{"body":"User 'ulysses-you' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29423","from":"developer"}],"created":"2017-05-09T10:28:46.000+0000","description":"Create a HIVE view:\n{quote}\nhive> create table bad as select 1 x, null z from dual;\n{quote}\n\nBecause there's no type, Hive gives it the VOID type:\n{quote}\nhive> describe bad;\nOK\nx\tint\t\nz\tvoid\n{quote}\n\nIn Spark2.0.x, the behaviour to read this view is normal:\n{quote}\nspark-sql> describe bad;\nx int NULL\nz void NULL\nTime taken: 4.431 seconds, Fetched 2 row(s)\n{quote}\n\nBut in Spark2.1.x, it failed with SparkException: Cannot recognize hive type string: void\n{quote}\nspark-sql> describe bad;\n17/05/09 03:12:08 INFO execution.SparkSqlParser: Parsing command: describe bad\n17/05/09 03:12:08 INFO parser.CatalystSqlParser: Parsing command: int\n17/05/09 03:12:08 INFO parser.CatalystSqlParser: Parsing command: void\n17/05/09 03:12:08 ERROR thriftserver.SparkSQLDriver: Failed in [describe bad]\norg.apache.spark.SparkException: Cannot recognize hive type string: void\n at org.apache.spark.sql.hive.client.HiveClientImpl.org$apache$spark$sql$hive$client$HiveClientImpl$$fromHiveColumn(HiveClientImpl.scala:789)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getTableOption$1$$anonfun$apply$11$$anonfun$7.apply(HiveClientImpl.scala:365) \n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getTableOption$1$$anonfun$apply$11$$anonfun$7.apply(HiveClientImpl.scala:365) \n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.Iterator$class.foreach(Iterator.scala:893)\n at scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\n at scala.collection.IterableLike$class.foreach(IterableLike.scala:72)\n at scala.collection.AbstractIterable.foreach(Iterable.scala:54)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.AbstractTraversable.map(Traversable.scala:104)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getTableOption$1$$anonfun$apply$11.apply(HiveClientImpl.scala:365)\n at org.apache.spark.sql.hive.client.HiveClientImpl$$anonfun$getTableOption$1$$anonfun$apply$11.apply(HiveClientImpl.scala:361)\n\nCaused by: org.apache.spark.sql.catalyst.parser.ParseException:\nDataType void() is not supported.(line 1, pos 0)\n\n== SQL == \nvoid \n^^^\n\n ... 61 more\norg.apache.spark.SparkException: Cannot recognize hive type string: void\n\n\n{quote}\n","issue_id":"13070377","key":"SPARK-20680","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2020-07-08T01:58:41.000+0000","role":"fixed_distractor","summary":"Spark-sql do not support for void column datatype of view"} {"case_id":"12719084","cluster":"DISTRACTOR-SPARK-2075","comments":[{"body":"Okay I did some more digging. I think the issue is that the anonymous classes used by saveAsTextFile are not guaranteed to be compiled to the same name every time you compile them in Scala. In the Hadoop 1 build these end up being shortened wheras in the Hadoop 2 build they use the longer names. saveAsTextFile seems to, strangely, be the only affected function. I confirmed this by looking at the difference in the hadoop 1 and 2 jars:\n\n{code}\n$ jar tvf spark-1.0.0-bin-hadoop1/lib/spark-assembly-1.0.0-hadoop*.jar |grep \"rdd\\/RDD\\\\$\" | awk '{ print $8;}' | sort > hadoop1\n$ jar tvf spark-1.0.0-bin-hadoop2/lib/spark-assembly-1.0.0-hadoop*.jar |grep \"rdd\\/RDD\\\\$\" | awk '{ print $8;}' | sort > hadoop2\n$ diff hadoop1 hadoop2\n23a24\n> org/apache/spark/rdd/RDD$$anonfun$28$$anonfun$apply$13.class\n27,29d27\n< org/apache/spark/rdd/RDD$$anonfun$30$$anonfun$apply$13.class\n< org/apache/spark/rdd/RDD$$anonfun$30.class\n< org/apache/spark/rdd/RDD$$anonfun$31.class\n90a89,90\n> org/apache/spark/rdd/RDD$$anonfun$saveAsTextFile$1.class\n> org/apache/spark/rdd/RDD$$anonfun$saveAsTextFile$2.class\n{code}\n\nThis strangely only seems to affect the saveAsTextFile function.\n\nI'm still a bit confused though because I didn't think these anonymous classes would show up in the byte code of the user application, so I don't think it should matter (i.e. this is why Scala probably allows this).\n\n{code}\njavap RDD | grep saveAsText\n public void saveAsTextFile(java.lang.String);\n public void saveAsTextFile(java.lang.String, java.lang.Class);\n{code}\n\n[~paulrbrown] could you explain how you are bundling and submitting your application to the Spark cluster?","created":"2014-06-08T22:11:13.315+0000"},{"body":"The job is run by a Java client that connects to the master (using a SparkContext).\n\nBundling is performed by a Maven build with two shade plugin invocations, one to package a \"driver\" uberjar and one to packager a \"worker\" uberjar. The worker flavor is sent to the worker nodes, the driver contains the code to connect to the master and run the job. The Maven build runs against the JAR from Maven Central, and the deployment uses the Spark 1.0.0 hadoop1 download. (The Spark is staged to S3 once and then downloaded onto master/worker nodes and set up during cluster provisioning.)\n\nThe Maven build uses the usual Scala setup with the library as a dependency and the plugin:\n\n{code}\n \n org.scala-lang\n scala-library\n 2.10.3\n \n{code}\n\n{code}\n \n net.alchim31.maven\n scala-maven-plugin\n \n \n \n compile\n testCompile\n \n \n \n \n 2.10.3\n \n -Xms64m\n -Xmx4096m\n \n \n \n{code}\n","created":"2014-06-09T15:04:53.651+0000"},{"body":"As food for thought, [here|http://docs.oracle.com/javase/specs/jvms/se7/html/jvms-4.html#jvms-4.7.6] is the {{InnerClass}} section of the JVM spec. It looks like there have been some changes from 2.10.3 to 2.10.4 (e.g., [SI-6546|https://issues.scala-lang.org/browse/SI-6546]), but I didn't dig in.\n\nI think the thing most likely to work is to ensure that exactly the same bits are used by all of the distributions and posted to Maven Central. (For some discussion on inner class naming stability, there was quite a bit of it on the Java 8 lambda discussion list, e.g., [this message|http://mail.openjdk.java.net/pipermail/lambda-spec-experts/2013-July/000316.html].)","created":"2014-06-09T15:36:29.738+0000"},{"body":"I see - so the issue is that your closures are compiled to call that anonymous class in your driver. Then on the cluster the anonymous class is named differently. I wonder if scala can produce stable naming for anonymous classes, that would be the best solution because otherwise it sort of breaks closures. I looked around the compiler flags but I don't see anything useful.\n\nOne solution here is to see if we can avoid re-compiling spark-core when we make the different builds. The tricky bit though is that if users compile their own downstream version of Spark we'd have to make sure they don't end up with different versions as well.","created":"2014-06-09T17:46:30.586+0000"},{"body":"The Scala compiler produces stable names for anonymous functions. In fact, that's the reason why the name of the enclosing method is part of the name: so that adding or removing an anonymous function in another method does not change the numbering of the others. Names are assigned by using a per-compilation unit counter and a prefix. Looking at the diff, there's quite a different picture in the two cases (anonymous functions vs. anonymous classes). Are you sure the two jars are built from the same sources?\n\nI don't know how the `assembly` jar is produced, but if it's using some sort of whole-program analysis and dead-code elimination, it might erroneously remove them. It might help to look at the inputs to the assembly and see if the class is already missing.\n\nAnother possibility is running `scalac -optimize` in only one of the two builds. However, looking at current sources I can't see why the inliner would remove those closures (the class is not final, and `map` is not final either, so they can't be resolved and inlined).. ","created":"2014-10-15T14:12:51.197+0000"},{"body":"Is there any more on this?\n\nBuilding Spark from the 1.1.0 tar for Hadoop 1.2.1--all is well. Trying to upgrade Mahout to use Spark 1.1.0. The Mahout 1.0-snapshot source builds and build tests pass with spark 1.1.0 as a maven dependency. Running the Mahout build on some bigger data using my dev machine as a standalone single node Spark cluster. So the same code is running as executed the build tests, just in single node cluster mode. Also since I built Spark i assume it is using the artifact from my .m2 maven cache, but not 100% on that. Anyway I get the class not found error below.\n\nI assume the missing function is the anon function passed to the \n\n{code}\n rdd.map(\n {anon function}\n )saveAsTextFile ???? \n{code}\n\nso shouldn't the function be in the Mahout jar (it isn't)? Isn't this function passed in from Mahout so I don't understand why it matters how Spark was built. \n\nSeveral other users are getting this for Spark 1.0.2. If we are doing something wrong in our build process we'd appreciate a pointer.\n\nHere's the error I get:\n\n14/10/20 17:21:36 WARN scheduler.TaskSetManager: Lost task 0.0 in stage 8.0 (TID 16, 192.168.0.2): java.lang.ClassNotFoundException: org.apache.spark.rdd.RDD$$anonfun$saveAsTextFile$1\n java.net.URLClassLoader$1.run(URLClassLoader.java:202)\n java.security.AccessController.doPrivileged(Native Method)\n java.net.URLClassLoader.findClass(URLClassLoader.java:190)\n java.lang.ClassLoader.loadClass(ClassLoader.java:306)\n java.lang.ClassLoader.loadClass(ClassLoader.java:247)\n java.lang.Class.forName0(Native Method)\n java.lang.Class.forName(Class.java:249)\n org.apache.spark.serializer.JavaDeserializationStream$$anon$1.resolveClass(JavaSerializer.scala:59)\n java.io.ObjectInputStream.readNonProxyDesc(ObjectInputStream.java:1591)\n java.io.ObjectInputStream.readClassDesc(ObjectInputStream.java:1496)\n java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:1750)\n java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1329)\n java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:1970)\n java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:1895)\n java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:1777)\n java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1329)\n java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:1970)\n java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:1895)\n java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:1777)\n java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1329)\n java.io.ObjectInputStream.readObject(ObjectInputStream.java:349)\n org.apache.spark.serializer.JavaDeserializationStream.readObject(JavaSerializer.scala:62)\n org.apache.spark.serializer.JavaSerializerInstance.deserialize(JavaSerializer.scala:87)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:57)\n org.apache.spark.scheduler.Task.run(Task.scala:54)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:177)\n java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:895)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:918)\n java.lang.Thread.run(Thread.java:695)\n ","created":"2014-10-21T00:55:47.528+0000"},{"body":"The name of the missing class is {{org.apache.spark.rdd.RDD$$anonfun$saveAsTextFile$1}}, so unless Mahout is putting classes inside Spark packages, that class should be somewhere *in Spark jars*.","created":"2014-10-21T12:44:26.403+0000"},{"body":"Oops, right. But the function name is being constructed at Mahout build time, right? So the rules for constructing the name are different when building Spark and Mahout OR the function is not being put in the Spark jars?\n\nRebuilding yet again so I can't check the jars just yet. This error was from a clean build of Spark 1.1.0 and Mahout in that sequence. I cleaned the .m2 repo and see that it is filled in only when Mahout is build and is filled in with a repo version of Spark 1.1.0\n\nCould the problem be related to the fact that Spark as executed on my standalone cluster is from a local build but as Mahout is built it uses the repo version of Spark? Since I am using hadoop 1.2.1 I suspect the repo version may not be exactly what I build.","created":"2014-10-21T17:30:25.422+0000"},{"body":"trying mvn install instead of the documented mvn package to put Spark in the maven cache so that when building Mahout it will get exactly the same bits as will later be run on the cluster--WAG","created":"2014-10-21T17:45:30.431+0000"},{"body":"OK solved. The WAG worked.\n\nInstead of 'mvn package ...' use 'mvn install ...' so you'll get the exact version of Spark needed in your local maven cache at ~/.m2\n\nThis should be changed in the instructions until some way to reliable point to the right version of Spark in the Maven repos is implemented. In my case I had to build Spark to target Hadoop 1.2.1 but the Maven repo was not built for that.","created":"2014-10-21T19:00:16.440+0000"},{"body":"It looks like in the end this is a problem caused by mismatching Spark versions at the client/server end, which isn't supported.","created":"2014-12-10T06:59:11.349+0000"},{"body":"It's not a version of Spark that's in question; it's incompatibility between what's published to Maven Central and what's offered for download from the Apache site for what is nominally the *same* version.\n\nThat said, I don't object to closing the issue, but I agree with [~pferrel] above that it's a documentation issue. The perfect-world solution would be to publish *multiple* Spark artifacts with Maven classifiers identifying the expected backplane. This issue will continue to afflict people who build Spark applications with Maven Central artifacts and attempt to connect to Spark clusters build against a different backplane.","created":"2014-12-10T09:05:17.679+0000"},{"body":"Hm on re-reading more closely, you are right, this is not attributable to version mismatch per se. I think it's better to leave it open even if I don't know of any action here. Sorry for the noise.\n\nThe same thing appears to occur in 1.1.1. Maybe this goes away, accidentally, when people all use Hadoop 2.x.\n\nI don't think there will be multiple versions deployed to Maven as I think the theory is that the artifacts only exist for the API, and not whatever other libs they happen to link against. This is a leak in that theory though.","created":"2014-12-10T09:33:10.516+0000"},{"body":"If the explanation is correct this needs to be filed against Spark as putting the wrong or not enough artifacts into maven repos. There would need to be a different artifact for every config option that will change internal naming.\n\nI can't understand why lots of people aren't running into this, all it requires is that you link against the repo artifact and run against a user compiled Spark.","created":"2014-12-10T20:45:59.528+0000"},{"body":"I met the same issue. I had a post in the Spark user mailing list but it does not get archived in http://apache-spark-user-list.1001560.n3.nabble.com/, so I have to describe the issue here:\n\nSteps to reproduce:\n1.\tDownload the official pre-built Spark binary 1.1.1 at http://d3kbcqa49mib13.cloudfront.net/spark-1.1.1-bin-hadoop1.tgz \n2.\tLaunch the Spark cluster in pseudo cluster mode\n3.\tA small scala APP which calls RDD.saveAsObjectFile()\nscalaVersion := \"2.10.4\"\n\nlibraryDependencies ++= Seq(\n \"org.apache.spark\" %% \"spark-core\" % \"1.1.1\"\n)\n\nval sc = new SparkContext(args(0), \"test\") //args[0] is the Spark master URI\n val rdd = sc.parallelize(List(1, 2, 3))\n rdd.saveAsObjectFile(\"/tmp/mysaoftmp\")\n sc.stop\n\nthrows an exception as follows:\n[error] (run-main-0) org.apache.spark.SparkException: Job aborted due to stage failure: Task 1 in stage 0.0 failed 4 times, most recent failure: Lost task 1.3 in stage 0.0 (TID 6, ray-desktop.sh.intel.com): java.lang.ClassCastException: scala.Tuple2 cannot be cast to scala.collection.Iterator\n[error] org.apache.spark.rdd.RDD$$anonfun$13.apply(RDD.scala:596)\n[error] org.apache.spark.rdd.RDD$$anonfun$13.apply(RDD.scala:596)\n[error] org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:35)\n[error] org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n[error] org.apache.spark.rdd.RDD.iterator(RDD.scala:229)\n[error] org.apache.spark.rdd.MappedRDD.compute(MappedRDD.scala:31)\n[error] org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n[error] org.apache.spark.rdd.RDD.iterator(RDD.scala:229)\n[error] org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:62)\n[error] org.apache.spark.scheduler.Task.run(Task.scala:54)\n[error] org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:178)\n[error] java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1146)\n[error] java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n[error] java.lang.Thread.run(Thread.java:701)\n\nAfter investigation, I found that this is caused by bytecode incompatibility issue between RDD.class in spark-core_2.10-1.1.1.jar and the pre-built spark assembly respectively.\n\nThis issue also happens with spark 1.1.0.\n","created":"2014-12-18T11:00:25.602+0000"},{"body":"Dig deeply and found weird things:\n\nIf I used `mvn -Dhadoop.version=1.2.1 -DskipTests clean package -pl core -am` to compile, the `saveAsTextFile` will be:\n{noformat}\npublic void saveAsTextFile(java.lang.String);\n Code:\n 0: aload_0\n 1: new #1577; //class org/apache/spark/rdd/RDD$$anonfun$27\n 4: dup\n 5: aload_0\n 6: invokespecial #1578; //Method org/apache/spark/rdd/RDD$$anonfun$27.\"\":(Lorg/apache/spark/rdd/RDD;)V\n 9: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 12: ldc_w #441; //class scala/Tuple2\n 15: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 18: invokevirtual #447; //Method map:(Lscala/Function1;Lscala/reflect/ClassTag;)Lorg/apache/spark/rdd/RDD;\n 21: astore_2\n 22: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 25: ldc_w #1580; //class org/apache/hadoop/io/NullWritable\n 28: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 31: astore_3\n 32: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 35: ldc_w #1582; //class org/apache/hadoop/io/Text\n 38: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 41: astore 4\n 43: getstatic #21; //Field org/apache/spark/rdd/RDD$.MODULE$:Lorg/apache/spark/rdd/RDD$;\n 46: aload_2\n 47: invokevirtual #23; //Method org/apache/spark/rdd/RDD$.rddToPairRDDFunctions$default$4:(Lorg/apache/spark/rdd/RDD;)Lscala/runtime/Null$;\n 50: astore 5\n 52: getstatic #21; //Field org/apache/spark/rdd/RDD$.MODULE$:Lorg/apache/spark/rdd/RDD$;\n 55: aload_2\n 56: aload_3\n 57: aload 4\n 59: aload 5\n 61: pop\n 62: aconst_null\n 63: invokevirtual #47; //Method org/apache/spark/rdd/RDD$.rddToPairRDDFunctions:(Lorg/apache/spark/rdd/RDD;Lscala/reflect/ClassTag;Lscala/reflect/ClassTag;Lscala/math/Ordering;)Lorg/apache/spark/rdd/PairRDDFunctions;\n 66: aload_1\n 67: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 70: ldc_w #1584; //class org/apache/hadoop/mapred/TextOutputFormat\n 73: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 76: invokevirtual #1588; //Method org/apache/spark/rdd/PairRDDFunctions.saveAsHadoopFile:(Ljava/lang/String;Lscala/reflect/ClassTag;)V\n 79: return\n{noformat}\n\nIf I used `mvn -Pyarn -Phadoop-2.2 -Dhadoop.version=2.2.0 -DskipTests clean package -pl core -am` to compile, the `saveAsTextFile` is different:\n{noformat}\npublic void saveAsTextFile(java.lang.String);\n Code:\n 0: getstatic #21; //Field org/apache/spark/rdd/RDD$.MODULE$:Lorg/apache/spark/rdd/RDD$;\n 3: aload_0\n 4: new #1577; //class org/apache/spark/rdd/RDD$$anonfun$saveAsTextFile$1\n 7: dup\n 8: aload_0\n 9: invokespecial #1578; //Method org/apache/spark/rdd/RDD$$anonfun$saveAsTextFile$1.\"\":(Lorg/apache/spark/rdd/RDD;)V\n 12: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 15: ldc_w #441; //class scala/Tuple2\n 18: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 21: invokevirtual #447; //Method map:(Lscala/Function1;Lscala/reflect/ClassTag;)Lorg/apache/spark/rdd/RDD;\n 24: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 27: ldc_w #1580; //class org/apache/hadoop/io/NullWritable\n 30: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 33: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 36: ldc_w #1582; //class org/apache/hadoop/io/Text\n 39: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 42: getstatic #1587; //Field scala/math/Ordering$.MODULE$:Lscala/math/Ordering$;\n 45: getstatic #471; //Field scala/Predef$.MODULE$:Lscala/Predef$;\n 48: invokevirtual #1591; //Method scala/Predef$.conforms:()Lscala/Predef$$less$colon$less;\n 51: invokevirtual #1595; //Method scala/math/Ordering$.ordered:(Lscala/Function1;)Lscala/math/Ordering;\n 54: invokevirtual #47; //Method org/apache/spark/rdd/RDD$.rddToPairRDDFunctions:(Lorg/apache/spark/rdd/RDD;Lscala/reflect/ClassTag;Lscala/reflect/ClassTag;Lscala/math/Ordering;)Lorg/apache/spark/rdd/PairRDDFunctions;\n 57: aload_1\n 58: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 61: ldc_w #1597; //class org/apache/hadoop/mapred/TextOutputFormat\n 64: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 67: invokevirtual #1601; //Method org/apache/spark/rdd/PairRDDFunctions.saveAsHadoopFile:(Ljava/lang/String;Lscala/reflect/ClassTag;)V\n 70: return\n{noformat}\n\nNote: in hadoop 1.2.1, saveAsTextFile use the default `Ordering` value `null`, while in hadoop 2.2.0, saveAsTextFile will use `Ordering.ordered` to create a new `Ordering`.\n","created":"2014-12-18T11:01:00.957+0000"},{"body":"I think we need to address this issue, at least we need to document it.\n\nIn order to guarrantee the binary compatibility, I think we may enforce the release steps to build the assembly as follows for an official release:\n1. Build all module jars except the assembly module\n2. Publish all modules jars to the central maven repository\n3. Then build the assembly module to assemble the module jars into the final assembly by depending on modular jars in the maven repo\n4. May also push the assembly jar to the central maven repository\n","created":"2014-12-18T11:11:04.507+0000"},{"body":"[~zsxwing] I don't expect the byte code to be identical when compiled against different versions of the underlying library, not necessarily. That's not quite the issue, but rather how these differences change the naming of anonymous functions. [~sunrui] I don't think that's quite the issue either. The assembly *is* made from the module JARs. All are published to Maven Central; please search http://search.maven.org/ . ","created":"2014-12-18T11:16:53.496+0000"},{"body":"Owen, if the official assembly is made from the exact published jars, then the binaries should be exactly the same, why we met this issue? You can see that the RDD.class from the maven spark-core 1.1.0 is different from that in the Spark pre-built 1.1.0 assembly (they have different file size, and in-compatible bytecode ).\n\nSo I doubt there is some mistake in the release step of Spark. ","created":"2014-12-18T11:26:50.074+0000"},{"body":"(Sean) But my understanding is that you are comparing to binaries compiled for a potentially different underlying Hadoop distribution. That is the issue at hand here.","created":"2014-12-18T12:12:54.551+0000"},{"body":"So just to make sure I understand things correctly, is it the case that the jar published to maven (spark-core-1.1.1) is built using Hadoop2 dependencies while the Hadoop1 assembly jar that is distributed is built using Hadoop 1 (obviously...) ?\n\n[~srowen] While I see that we officially support submitting jobs using spark-submit, it is surprising to me that other deployment methods would fail this way (from the user's perspective the Spark versions presumably at compile time and run time presumably match up ?). We should at the very least document this, but it would also be good to see if there is a work around.","created":"2014-12-18T19:03:24.090+0000"},{"body":"[~shivaram] No, I don't think that's the case. Certainly not Hadoop 1 vs 2. Here is how the downloadable distros are built:\n\n{code}\nmake_binary_release \"hadoop1\" \"-Phive -Phive-thriftserver -Dhadoop.version=1.0.4\" &\nmake_binary_release \"hadoop1-scala2.11\" \"-Phive -Dscala-2.11\" &\nmake_binary_release \"cdh4\" \"-Phive -Phive-thriftserver -Dhadoop.version=2.0.0-mr1-cdh4.2.0\" &\nmake_binary_release \"hadoop2.3\" \"-Phadoop-2.3 -Phive -Phive-thriftserver -Pyarn\" &\nmake_binary_release \"hadoop2.4\" \"-Phadoop-2.4 -Phive -Phive-thriftserver -Pyarn\" &\nmake_binary_release \"mapr3\" \"-Pmapr3 -Phive -Phive-thriftserver\" &\nmake_binary_release \"mapr4\" \"-Pmapr4 -Pyarn -Phive -Phive-thriftserver\" &\nmake_binary_release \"hadoop2.4-without-hive\" \"-Phadoop-2.4 -Pyarn\" &\n{code}\n\nA default {{mvn release}} would use {{hadoop.version=1.0.4}}. Somebody can correct me if this isn't the case, but I assume that this is what goes to Maven Central.\n\nOf course, this behavior is not desirable and not by design, and worth a mention. As far as I can tell it has only been observed arising in Hadoop 1 vs Hadoop 2-compiled artifacts. Although the intended public API is identical, it's not actually 100% compatible. Ideally you could get away with using any copy of the public API. \n\nIn these particular cases, the safe practice of always harmonizing binaries on client and server is actually necessary, which is no terrible thing I think. Nothing about this requires using {{spark-submit}}, although, that's a way to make sure you're using the same Spark in your app and cluster.","created":"2014-12-18T20:31:21.566+0000"},{"body":"Hmm -- looking at the release steps it looks like the release on maven should be from Hadoop 1.0.4 [~pwendell] or [~andrewor14] might be able to throw more light on this. (BTW I wonder if we can trace the source of this mismatch for the case reported by [~sunrui] where the distribution with Hadoop1 of Spark 1.1.1 doesn't work with the Maven central jar)\n\nI see your high level point that this is not about spark-submit per se, but about having the exact same binary on the server and as a compile-time dependency. Its just unfortunate that having the same Spark version number isn't sufficient. Also is the workaround right now to rebuild Spark from source using `make-distribution`, do `mvn install`, rebuild the application and deploy Spark using the assembly jar ?","created":"2014-12-18T21:57:23.246+0000"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/3740","created":"2014-12-19T02:35:25.908+0000"},{"body":"Please take a look at my PR","created":"2014-12-19T04:21:43.596+0000"},{"body":"Since `mvn release` is built for Hadoop 1.0.4, I don't understand the reason why there is difference in RDD.class bytecode from mvn spark-core and the pre-built binary for Hadoop 1.x, because they are both built for Hadoop 1.x and has the same version 1.1.0.\n\nAccording to [~zsxwing]'s PR, it seems that it's diffcult to guanranttee same bytecode for Hadoop1.x and 2.x. So maybe we need to pubish two versions of a module to mvn, one is for Hadoop 1.x, the other is for Hadoop 2.x, for example, spark-core_2.10-1.1.0-hadoop1.jar and spark-core_2.10-1.1.0-hadoop2.jar?\n","created":"2014-12-19T05:32:40.600+0000"},{"body":"[~sunrui] From digging in to the various reports of this issue, it seemed to me that in each case the Hadoop version did not match. That is, I do not know that it's true that the issue manifests when the Hadoop version matches; that would indeed be strange. I could have missed it; this is a bit hard to follow. But do you see evidence of this?\n\nI don't think publishing two versions fixes anything, really. The PR might get at the heart of the difference here and resolve it for real. It doesn't happen if you match binaries, which is good practice anyway.","created":"2014-12-19T10:05:28.210+0000"},{"body":"[~srowen] I assume that mvn jars were built for Hadoop 1.x as you said that \"A default mvn release would use hadoop.version=1.0.4. Somebody can correct me if this isn't the case, but I assume that this is what goes to Maven Central.\" :) so it is very possible that mvn jars were actually built for Hadoop 2.x, which is the cause for this issue.\n\nYes, I agree that we should using match binaries. The problem is that I thought I was using the matching binaries but actually not. So we can \n1. For exising releases, document the fact that mvn jars are for Hadoop 2.x. If app depends on mvn jars, it should be used with Hadoop 2.x. If the app is intended to work with Hadoop 1.x, one way is to rebuild Spark source against Hadoop 1.x and publish the module jars to local mvn repo.\n2. For futrure release, we may eliminate the Spark core bytecode incompatibility between Hadoop1.x and 2.x just as Shixong's PR is trying to do. ","created":"2014-12-20T13:07:40.601+0000"},{"body":"[~sunrui] You are asking if the Maven Central artifacts are built for Hadoop 2? No, why do you say that? It is not true that you need to use Hadoop 2, or need to build a custom version. What you should do ideally is match the version you compile against the version you deploy against -- good practice, even if in reality it should not be so strict. \n\nBut the best outcome indeed is to allow compiling against any artifact, and having it work against any build of the same public API version, no matter what the 'backend'. That was intended behavior and hopefully the PRs here get at the nature of the problem.","created":"2014-12-21T10:41:31.248+0000"},{"body":"[~srowen] Yes, I totally agree with you matching version between compilation and depolyment. That's my point also. What I am confused is that an app depends on maven spark-core_2.10-1.1.0.jar (suppose it is built for Hadoop 1.x as you said by default is for Hadoop 1.x for mvn release) does not match the spark-core included in the Spark pre-built binary 1.1.0 for hadoop 1.x. I suppose that Spark pre-built binary 1.1.0 for Hadoop 1.x be assembled from MVN jars, so there should be no mismatching, Am I right? If I am wrong, could anyone tell me the reason of the binary mismatch and tell me the best practice for me to solve my problem?\n\n","created":"2014-12-22T02:06:02.026+0000"},{"body":"[~sunrui] What I can see from this JIRA discussion (and [~srowen] please correct me if I am wrong) is that Hadoop 1 vs. Hadoop 2 is one of the causes of incompatibility. It is _not the only_ reason and I don't think we exactly know why the pre-built binary for 1.1.0 is different from the maven version. \n\nI think the best practice advice is to use the exact same jar in the application and in the runtime. Marking Spark a provided dependency in the application build and using spark-submit is one way of achieving this. Or one can publish a local build to maven and use the same local build to start the cluster etc.","created":"2014-12-22T02:30:54.436+0000"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/3758","created":"2014-12-22T06:21:08.170+0000"},{"body":"[~sunrui] Yes, but I am still not clear that anyone has observed the problem with two binaries built for the same version of Hadoop. Most of the situations listed here do not match that description. I might not understand someone's issue report here. In any event, it sounds like an underlying cause has been fixed already anyway.","created":"2014-12-22T10:12:43.277+0000"}],"conversations":[{"body":"Running a job built against the Maven dep for 1.0.0 and the hadoop1 distribution produces:\n\n{code}\njava.lang.ClassNotFoundException:\norg.apache.spark.rdd.RDD$$anonfun$saveAsTextFile$1\n{code}\n\nHere's what's in the Maven dep as of 1.0.0:\n\n{code}\njar tvf ~/.m2/repository/org/apache/spark/spark-core_2.10/1.0.0/spark-core_2.10-1.0.0.jar | grep 'rdd/RDD' | grep 'saveAs'\n 1519 Mon May 26 13:57:58 PDT 2014 org/apache/spark/rdd/RDD$anonfun$saveAsTextFile$1.class\n 1560 Mon May 26 13:57:58 PDT 2014 org/apache/spark/rdd/RDD$anonfun$saveAsTextFile$2.class\n{code}\n\nAnd here's what's in the hadoop1 distribution:\n\n{code}\njar tvf spark-assembly-1.0.0-hadoop1.0.4.jar| grep 'rdd/RDD' | grep 'saveAs'\n{code}\n\nI.e., it's not there. It is in the hadoop2 distribution:\n\n{code}\njar tvf spark-assembly-1.0.0-hadoop2.2.0.jar| grep 'rdd/RDD' | grep 'saveAs'\n 1519 Mon May 26 07:29:54 PDT 2014 org/apache/spark/rdd/RDD$anonfun$saveAsTextFile$1.class\n 1560 Mon May 26 07:29:54 PDT 2014 org/apache/spark/rdd/RDD$anonfun$saveAsTextFile$2.class\n{code}","from":"reporter","subject":"Anonymous classes are missing from Spark distribution"},{"body":"Okay I did some more digging. I think the issue is that the anonymous classes used by saveAsTextFile are not guaranteed to be compiled to the same name every time you compile them in Scala. In the Hadoop 1 build these end up being shortened wheras in the Hadoop 2 build they use the longer names. saveAsTextFile seems to, strangely, be the only affected function. I confirmed this by looking at the difference in the hadoop 1 and 2 jars:\n\n{code}\n$ jar tvf spark-1.0.0-bin-hadoop1/lib/spark-assembly-1.0.0-hadoop*.jar |grep \"rdd\\/RDD\\\\$\" | awk '{ print $8;}' | sort > hadoop1\n$ jar tvf spark-1.0.0-bin-hadoop2/lib/spark-assembly-1.0.0-hadoop*.jar |grep \"rdd\\/RDD\\\\$\" | awk '{ print $8;}' | sort > hadoop2\n$ diff hadoop1 hadoop2\n23a24\n> org/apache/spark/rdd/RDD$$anonfun$28$$anonfun$apply$13.class\n27,29d27\n< org/apache/spark/rdd/RDD$$anonfun$30$$anonfun$apply$13.class\n< org/apache/spark/rdd/RDD$$anonfun$30.class\n< org/apache/spark/rdd/RDD$$anonfun$31.class\n90a89,90\n> org/apache/spark/rdd/RDD$$anonfun$saveAsTextFile$1.class\n> org/apache/spark/rdd/RDD$$anonfun$saveAsTextFile$2.class\n{code}\n\nThis strangely only seems to affect the saveAsTextFile function.\n\nI'm still a bit confused though because I didn't think these anonymous classes would show up in the byte code of the user application, so I don't think it should matter (i.e. this is why Scala probably allows this).\n\n{code}\njavap RDD | grep saveAsText\n public void saveAsTextFile(java.lang.String);\n public void saveAsTextFile(java.lang.String, java.lang.Class);\n{code}\n\n[~paulrbrown] could you explain how you are bundling and submitting your application to the Spark cluster?","from":"developer"},{"body":"The job is run by a Java client that connects to the master (using a SparkContext).\n\nBundling is performed by a Maven build with two shade plugin invocations, one to package a \"driver\" uberjar and one to packager a \"worker\" uberjar. The worker flavor is sent to the worker nodes, the driver contains the code to connect to the master and run the job. The Maven build runs against the JAR from Maven Central, and the deployment uses the Spark 1.0.0 hadoop1 download. (The Spark is staged to S3 once and then downloaded onto master/worker nodes and set up during cluster provisioning.)\n\nThe Maven build uses the usual Scala setup with the library as a dependency and the plugin:\n\n{code}\n \n org.scala-lang\n scala-library\n 2.10.3\n \n{code}\n\n{code}\n \n net.alchim31.maven\n scala-maven-plugin\n \n \n \n compile\n testCompile\n \n \n \n \n 2.10.3\n \n -Xms64m\n -Xmx4096m\n \n \n \n{code}\n","from":"developer"},{"body":"As food for thought, [here|http://docs.oracle.com/javase/specs/jvms/se7/html/jvms-4.html#jvms-4.7.6] is the {{InnerClass}} section of the JVM spec. It looks like there have been some changes from 2.10.3 to 2.10.4 (e.g., [SI-6546|https://issues.scala-lang.org/browse/SI-6546]), but I didn't dig in.\n\nI think the thing most likely to work is to ensure that exactly the same bits are used by all of the distributions and posted to Maven Central. (For some discussion on inner class naming stability, there was quite a bit of it on the Java 8 lambda discussion list, e.g., [this message|http://mail.openjdk.java.net/pipermail/lambda-spec-experts/2013-July/000316.html].)","from":"developer"},{"body":"I see - so the issue is that your closures are compiled to call that anonymous class in your driver. Then on the cluster the anonymous class is named differently. I wonder if scala can produce stable naming for anonymous classes, that would be the best solution because otherwise it sort of breaks closures. I looked around the compiler flags but I don't see anything useful.\n\nOne solution here is to see if we can avoid re-compiling spark-core when we make the different builds. The tricky bit though is that if users compile their own downstream version of Spark we'd have to make sure they don't end up with different versions as well.","from":"developer"},{"body":"The Scala compiler produces stable names for anonymous functions. In fact, that's the reason why the name of the enclosing method is part of the name: so that adding or removing an anonymous function in another method does not change the numbering of the others. Names are assigned by using a per-compilation unit counter and a prefix. Looking at the diff, there's quite a different picture in the two cases (anonymous functions vs. anonymous classes). Are you sure the two jars are built from the same sources?\n\nI don't know how the `assembly` jar is produced, but if it's using some sort of whole-program analysis and dead-code elimination, it might erroneously remove them. It might help to look at the inputs to the assembly and see if the class is already missing.\n\nAnother possibility is running `scalac -optimize` in only one of the two builds. However, looking at current sources I can't see why the inliner would remove those closures (the class is not final, and `map` is not final either, so they can't be resolved and inlined).. ","from":"developer"},{"body":"Is there any more on this?\n\nBuilding Spark from the 1.1.0 tar for Hadoop 1.2.1--all is well. Trying to upgrade Mahout to use Spark 1.1.0. The Mahout 1.0-snapshot source builds and build tests pass with spark 1.1.0 as a maven dependency. Running the Mahout build on some bigger data using my dev machine as a standalone single node Spark cluster. So the same code is running as executed the build tests, just in single node cluster mode. Also since I built Spark i assume it is using the artifact from my .m2 maven cache, but not 100% on that. Anyway I get the class not found error below.\n\nI assume the missing function is the anon function passed to the \n\n{code}\n rdd.map(\n {anon function}\n )saveAsTextFile ???? \n{code}\n\nso shouldn't the function be in the Mahout jar (it isn't)? Isn't this function passed in from Mahout so I don't understand why it matters how Spark was built. \n\nSeveral other users are getting this for Spark 1.0.2. If we are doing something wrong in our build process we'd appreciate a pointer.\n\nHere's the error I get:\n\n14/10/20 17:21:36 WARN scheduler.TaskSetManager: Lost task 0.0 in stage 8.0 (TID 16, 192.168.0.2): java.lang.ClassNotFoundException: org.apache.spark.rdd.RDD$$anonfun$saveAsTextFile$1\n java.net.URLClassLoader$1.run(URLClassLoader.java:202)\n java.security.AccessController.doPrivileged(Native Method)\n java.net.URLClassLoader.findClass(URLClassLoader.java:190)\n java.lang.ClassLoader.loadClass(ClassLoader.java:306)\n java.lang.ClassLoader.loadClass(ClassLoader.java:247)\n java.lang.Class.forName0(Native Method)\n java.lang.Class.forName(Class.java:249)\n org.apache.spark.serializer.JavaDeserializationStream$$anon$1.resolveClass(JavaSerializer.scala:59)\n java.io.ObjectInputStream.readNonProxyDesc(ObjectInputStream.java:1591)\n java.io.ObjectInputStream.readClassDesc(ObjectInputStream.java:1496)\n java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:1750)\n java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1329)\n java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:1970)\n java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:1895)\n java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:1777)\n java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1329)\n java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:1970)\n java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:1895)\n java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:1777)\n java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1329)\n java.io.ObjectInputStream.readObject(ObjectInputStream.java:349)\n org.apache.spark.serializer.JavaDeserializationStream.readObject(JavaSerializer.scala:62)\n org.apache.spark.serializer.JavaSerializerInstance.deserialize(JavaSerializer.scala:87)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:57)\n org.apache.spark.scheduler.Task.run(Task.scala:54)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:177)\n java.util.concurrent.ThreadPoolExecutor$Worker.runTask(ThreadPoolExecutor.java:895)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:918)\n java.lang.Thread.run(Thread.java:695)\n ","from":"developer"},{"body":"The name of the missing class is {{org.apache.spark.rdd.RDD$$anonfun$saveAsTextFile$1}}, so unless Mahout is putting classes inside Spark packages, that class should be somewhere *in Spark jars*.","from":"developer"},{"body":"Oops, right. But the function name is being constructed at Mahout build time, right? So the rules for constructing the name are different when building Spark and Mahout OR the function is not being put in the Spark jars?\n\nRebuilding yet again so I can't check the jars just yet. This error was from a clean build of Spark 1.1.0 and Mahout in that sequence. I cleaned the .m2 repo and see that it is filled in only when Mahout is build and is filled in with a repo version of Spark 1.1.0\n\nCould the problem be related to the fact that Spark as executed on my standalone cluster is from a local build but as Mahout is built it uses the repo version of Spark? Since I am using hadoop 1.2.1 I suspect the repo version may not be exactly what I build.","from":"developer"},{"body":"trying mvn install instead of the documented mvn package to put Spark in the maven cache so that when building Mahout it will get exactly the same bits as will later be run on the cluster--WAG","from":"developer"},{"body":"OK solved. The WAG worked.\n\nInstead of 'mvn package ...' use 'mvn install ...' so you'll get the exact version of Spark needed in your local maven cache at ~/.m2\n\nThis should be changed in the instructions until some way to reliable point to the right version of Spark in the Maven repos is implemented. In my case I had to build Spark to target Hadoop 1.2.1 but the Maven repo was not built for that.","from":"developer"},{"body":"It looks like in the end this is a problem caused by mismatching Spark versions at the client/server end, which isn't supported.","from":"developer"},{"body":"It's not a version of Spark that's in question; it's incompatibility between what's published to Maven Central and what's offered for download from the Apache site for what is nominally the *same* version.\n\nThat said, I don't object to closing the issue, but I agree with [~pferrel] above that it's a documentation issue. The perfect-world solution would be to publish *multiple* Spark artifacts with Maven classifiers identifying the expected backplane. This issue will continue to afflict people who build Spark applications with Maven Central artifacts and attempt to connect to Spark clusters build against a different backplane.","from":"developer"},{"body":"Hm on re-reading more closely, you are right, this is not attributable to version mismatch per se. I think it's better to leave it open even if I don't know of any action here. Sorry for the noise.\n\nThe same thing appears to occur in 1.1.1. Maybe this goes away, accidentally, when people all use Hadoop 2.x.\n\nI don't think there will be multiple versions deployed to Maven as I think the theory is that the artifacts only exist for the API, and not whatever other libs they happen to link against. This is a leak in that theory though.","from":"developer"},{"body":"If the explanation is correct this needs to be filed against Spark as putting the wrong or not enough artifacts into maven repos. There would need to be a different artifact for every config option that will change internal naming.\n\nI can't understand why lots of people aren't running into this, all it requires is that you link against the repo artifact and run against a user compiled Spark.","from":"developer"},{"body":"I met the same issue. I had a post in the Spark user mailing list but it does not get archived in http://apache-spark-user-list.1001560.n3.nabble.com/, so I have to describe the issue here:\n\nSteps to reproduce:\n1.\tDownload the official pre-built Spark binary 1.1.1 at http://d3kbcqa49mib13.cloudfront.net/spark-1.1.1-bin-hadoop1.tgz \n2.\tLaunch the Spark cluster in pseudo cluster mode\n3.\tA small scala APP which calls RDD.saveAsObjectFile()\nscalaVersion := \"2.10.4\"\n\nlibraryDependencies ++= Seq(\n \"org.apache.spark\" %% \"spark-core\" % \"1.1.1\"\n)\n\nval sc = new SparkContext(args(0), \"test\") //args[0] is the Spark master URI\n val rdd = sc.parallelize(List(1, 2, 3))\n rdd.saveAsObjectFile(\"/tmp/mysaoftmp\")\n sc.stop\n\nthrows an exception as follows:\n[error] (run-main-0) org.apache.spark.SparkException: Job aborted due to stage failure: Task 1 in stage 0.0 failed 4 times, most recent failure: Lost task 1.3 in stage 0.0 (TID 6, ray-desktop.sh.intel.com): java.lang.ClassCastException: scala.Tuple2 cannot be cast to scala.collection.Iterator\n[error] org.apache.spark.rdd.RDD$$anonfun$13.apply(RDD.scala:596)\n[error] org.apache.spark.rdd.RDD$$anonfun$13.apply(RDD.scala:596)\n[error] org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:35)\n[error] org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n[error] org.apache.spark.rdd.RDD.iterator(RDD.scala:229)\n[error] org.apache.spark.rdd.MappedRDD.compute(MappedRDD.scala:31)\n[error] org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:262)\n[error] org.apache.spark.rdd.RDD.iterator(RDD.scala:229)\n[error] org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:62)\n[error] org.apache.spark.scheduler.Task.run(Task.scala:54)\n[error] org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:178)\n[error] java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1146)\n[error] java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n[error] java.lang.Thread.run(Thread.java:701)\n\nAfter investigation, I found that this is caused by bytecode incompatibility issue between RDD.class in spark-core_2.10-1.1.1.jar and the pre-built spark assembly respectively.\n\nThis issue also happens with spark 1.1.0.\n","from":"developer"},{"body":"Dig deeply and found weird things:\n\nIf I used `mvn -Dhadoop.version=1.2.1 -DskipTests clean package -pl core -am` to compile, the `saveAsTextFile` will be:\n{noformat}\npublic void saveAsTextFile(java.lang.String);\n Code:\n 0: aload_0\n 1: new #1577; //class org/apache/spark/rdd/RDD$$anonfun$27\n 4: dup\n 5: aload_0\n 6: invokespecial #1578; //Method org/apache/spark/rdd/RDD$$anonfun$27.\"\":(Lorg/apache/spark/rdd/RDD;)V\n 9: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 12: ldc_w #441; //class scala/Tuple2\n 15: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 18: invokevirtual #447; //Method map:(Lscala/Function1;Lscala/reflect/ClassTag;)Lorg/apache/spark/rdd/RDD;\n 21: astore_2\n 22: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 25: ldc_w #1580; //class org/apache/hadoop/io/NullWritable\n 28: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 31: astore_3\n 32: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 35: ldc_w #1582; //class org/apache/hadoop/io/Text\n 38: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 41: astore 4\n 43: getstatic #21; //Field org/apache/spark/rdd/RDD$.MODULE$:Lorg/apache/spark/rdd/RDD$;\n 46: aload_2\n 47: invokevirtual #23; //Method org/apache/spark/rdd/RDD$.rddToPairRDDFunctions$default$4:(Lorg/apache/spark/rdd/RDD;)Lscala/runtime/Null$;\n 50: astore 5\n 52: getstatic #21; //Field org/apache/spark/rdd/RDD$.MODULE$:Lorg/apache/spark/rdd/RDD$;\n 55: aload_2\n 56: aload_3\n 57: aload 4\n 59: aload 5\n 61: pop\n 62: aconst_null\n 63: invokevirtual #47; //Method org/apache/spark/rdd/RDD$.rddToPairRDDFunctions:(Lorg/apache/spark/rdd/RDD;Lscala/reflect/ClassTag;Lscala/reflect/ClassTag;Lscala/math/Ordering;)Lorg/apache/spark/rdd/PairRDDFunctions;\n 66: aload_1\n 67: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 70: ldc_w #1584; //class org/apache/hadoop/mapred/TextOutputFormat\n 73: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 76: invokevirtual #1588; //Method org/apache/spark/rdd/PairRDDFunctions.saveAsHadoopFile:(Ljava/lang/String;Lscala/reflect/ClassTag;)V\n 79: return\n{noformat}\n\nIf I used `mvn -Pyarn -Phadoop-2.2 -Dhadoop.version=2.2.0 -DskipTests clean package -pl core -am` to compile, the `saveAsTextFile` is different:\n{noformat}\npublic void saveAsTextFile(java.lang.String);\n Code:\n 0: getstatic #21; //Field org/apache/spark/rdd/RDD$.MODULE$:Lorg/apache/spark/rdd/RDD$;\n 3: aload_0\n 4: new #1577; //class org/apache/spark/rdd/RDD$$anonfun$saveAsTextFile$1\n 7: dup\n 8: aload_0\n 9: invokespecial #1578; //Method org/apache/spark/rdd/RDD$$anonfun$saveAsTextFile$1.\"\":(Lorg/apache/spark/rdd/RDD;)V\n 12: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 15: ldc_w #441; //class scala/Tuple2\n 18: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 21: invokevirtual #447; //Method map:(Lscala/Function1;Lscala/reflect/ClassTag;)Lorg/apache/spark/rdd/RDD;\n 24: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 27: ldc_w #1580; //class org/apache/hadoop/io/NullWritable\n 30: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 33: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 36: ldc_w #1582; //class org/apache/hadoop/io/Text\n 39: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 42: getstatic #1587; //Field scala/math/Ordering$.MODULE$:Lscala/math/Ordering$;\n 45: getstatic #471; //Field scala/Predef$.MODULE$:Lscala/Predef$;\n 48: invokevirtual #1591; //Method scala/Predef$.conforms:()Lscala/Predef$$less$colon$less;\n 51: invokevirtual #1595; //Method scala/math/Ordering$.ordered:(Lscala/Function1;)Lscala/math/Ordering;\n 54: invokevirtual #47; //Method org/apache/spark/rdd/RDD$.rddToPairRDDFunctions:(Lorg/apache/spark/rdd/RDD;Lscala/reflect/ClassTag;Lscala/reflect/ClassTag;Lscala/math/Ordering;)Lorg/apache/spark/rdd/PairRDDFunctions;\n 57: aload_1\n 58: getstatic #439; //Field scala/reflect/ClassTag$.MODULE$:Lscala/reflect/ClassTag$;\n 61: ldc_w #1597; //class org/apache/hadoop/mapred/TextOutputFormat\n 64: invokevirtual #445; //Method scala/reflect/ClassTag$.apply:(Ljava/lang/Class;)Lscala/reflect/ClassTag;\n 67: invokevirtual #1601; //Method org/apache/spark/rdd/PairRDDFunctions.saveAsHadoopFile:(Ljava/lang/String;Lscala/reflect/ClassTag;)V\n 70: return\n{noformat}\n\nNote: in hadoop 1.2.1, saveAsTextFile use the default `Ordering` value `null`, while in hadoop 2.2.0, saveAsTextFile will use `Ordering.ordered` to create a new `Ordering`.\n","from":"developer"},{"body":"I think we need to address this issue, at least we need to document it.\n\nIn order to guarrantee the binary compatibility, I think we may enforce the release steps to build the assembly as follows for an official release:\n1. Build all module jars except the assembly module\n2. Publish all modules jars to the central maven repository\n3. Then build the assembly module to assemble the module jars into the final assembly by depending on modular jars in the maven repo\n4. May also push the assembly jar to the central maven repository\n","from":"developer"},{"body":"[~zsxwing] I don't expect the byte code to be identical when compiled against different versions of the underlying library, not necessarily. That's not quite the issue, but rather how these differences change the naming of anonymous functions. [~sunrui] I don't think that's quite the issue either. The assembly *is* made from the module JARs. All are published to Maven Central; please search http://search.maven.org/ . ","from":"developer"},{"body":"Owen, if the official assembly is made from the exact published jars, then the binaries should be exactly the same, why we met this issue? You can see that the RDD.class from the maven spark-core 1.1.0 is different from that in the Spark pre-built 1.1.0 assembly (they have different file size, and in-compatible bytecode ).\n\nSo I doubt there is some mistake in the release step of Spark. ","from":"developer"},{"body":"(Sean) But my understanding is that you are comparing to binaries compiled for a potentially different underlying Hadoop distribution. That is the issue at hand here.","from":"developer"},{"body":"So just to make sure I understand things correctly, is it the case that the jar published to maven (spark-core-1.1.1) is built using Hadoop2 dependencies while the Hadoop1 assembly jar that is distributed is built using Hadoop 1 (obviously...) ?\n\n[~srowen] While I see that we officially support submitting jobs using spark-submit, it is surprising to me that other deployment methods would fail this way (from the user's perspective the Spark versions presumably at compile time and run time presumably match up ?). We should at the very least document this, but it would also be good to see if there is a work around.","from":"developer"},{"body":"[~shivaram] No, I don't think that's the case. Certainly not Hadoop 1 vs 2. Here is how the downloadable distros are built:\n\n{code}\nmake_binary_release \"hadoop1\" \"-Phive -Phive-thriftserver -Dhadoop.version=1.0.4\" &\nmake_binary_release \"hadoop1-scala2.11\" \"-Phive -Dscala-2.11\" &\nmake_binary_release \"cdh4\" \"-Phive -Phive-thriftserver -Dhadoop.version=2.0.0-mr1-cdh4.2.0\" &\nmake_binary_release \"hadoop2.3\" \"-Phadoop-2.3 -Phive -Phive-thriftserver -Pyarn\" &\nmake_binary_release \"hadoop2.4\" \"-Phadoop-2.4 -Phive -Phive-thriftserver -Pyarn\" &\nmake_binary_release \"mapr3\" \"-Pmapr3 -Phive -Phive-thriftserver\" &\nmake_binary_release \"mapr4\" \"-Pmapr4 -Pyarn -Phive -Phive-thriftserver\" &\nmake_binary_release \"hadoop2.4-without-hive\" \"-Phadoop-2.4 -Pyarn\" &\n{code}\n\nA default {{mvn release}} would use {{hadoop.version=1.0.4}}. Somebody can correct me if this isn't the case, but I assume that this is what goes to Maven Central.\n\nOf course, this behavior is not desirable and not by design, and worth a mention. As far as I can tell it has only been observed arising in Hadoop 1 vs Hadoop 2-compiled artifacts. Although the intended public API is identical, it's not actually 100% compatible. Ideally you could get away with using any copy of the public API. \n\nIn these particular cases, the safe practice of always harmonizing binaries on client and server is actually necessary, which is no terrible thing I think. Nothing about this requires using {{spark-submit}}, although, that's a way to make sure you're using the same Spark in your app and cluster.","from":"developer"},{"body":"Hmm -- looking at the release steps it looks like the release on maven should be from Hadoop 1.0.4 [~pwendell] or [~andrewor14] might be able to throw more light on this. (BTW I wonder if we can trace the source of this mismatch for the case reported by [~sunrui] where the distribution with Hadoop1 of Spark 1.1.1 doesn't work with the Maven central jar)\n\nI see your high level point that this is not about spark-submit per se, but about having the exact same binary on the server and as a compile-time dependency. Its just unfortunate that having the same Spark version number isn't sufficient. Also is the workaround right now to rebuild Spark from source using `make-distribution`, do `mvn install`, rebuild the application and deploy Spark using the assembly jar ?","from":"developer"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/3740","from":"developer"},{"body":"Please take a look at my PR","from":"developer"},{"body":"Since `mvn release` is built for Hadoop 1.0.4, I don't understand the reason why there is difference in RDD.class bytecode from mvn spark-core and the pre-built binary for Hadoop 1.x, because they are both built for Hadoop 1.x and has the same version 1.1.0.\n\nAccording to [~zsxwing]'s PR, it seems that it's diffcult to guanranttee same bytecode for Hadoop1.x and 2.x. So maybe we need to pubish two versions of a module to mvn, one is for Hadoop 1.x, the other is for Hadoop 2.x, for example, spark-core_2.10-1.1.0-hadoop1.jar and spark-core_2.10-1.1.0-hadoop2.jar?\n","from":"developer"},{"body":"[~sunrui] From digging in to the various reports of this issue, it seemed to me that in each case the Hadoop version did not match. That is, I do not know that it's true that the issue manifests when the Hadoop version matches; that would indeed be strange. I could have missed it; this is a bit hard to follow. But do you see evidence of this?\n\nI don't think publishing two versions fixes anything, really. The PR might get at the heart of the difference here and resolve it for real. It doesn't happen if you match binaries, which is good practice anyway.","from":"developer"},{"body":"[~srowen] I assume that mvn jars were built for Hadoop 1.x as you said that \"A default mvn release would use hadoop.version=1.0.4. Somebody can correct me if this isn't the case, but I assume that this is what goes to Maven Central.\" :) so it is very possible that mvn jars were actually built for Hadoop 2.x, which is the cause for this issue.\n\nYes, I agree that we should using match binaries. The problem is that I thought I was using the matching binaries but actually not. So we can \n1. For exising releases, document the fact that mvn jars are for Hadoop 2.x. If app depends on mvn jars, it should be used with Hadoop 2.x. If the app is intended to work with Hadoop 1.x, one way is to rebuild Spark source against Hadoop 1.x and publish the module jars to local mvn repo.\n2. For futrure release, we may eliminate the Spark core bytecode incompatibility between Hadoop1.x and 2.x just as Shixong's PR is trying to do. ","from":"developer"},{"body":"[~sunrui] You are asking if the Maven Central artifacts are built for Hadoop 2? No, why do you say that? It is not true that you need to use Hadoop 2, or need to build a custom version. What you should do ideally is match the version you compile against the version you deploy against -- good practice, even if in reality it should not be so strict. \n\nBut the best outcome indeed is to allow compiling against any artifact, and having it work against any build of the same public API version, no matter what the 'backend'. That was intended behavior and hopefully the PRs here get at the nature of the problem.","from":"developer"},{"body":"[~srowen] Yes, I totally agree with you matching version between compilation and depolyment. That's my point also. What I am confused is that an app depends on maven spark-core_2.10-1.1.0.jar (suppose it is built for Hadoop 1.x as you said by default is for Hadoop 1.x for mvn release) does not match the spark-core included in the Spark pre-built binary 1.1.0 for hadoop 1.x. I suppose that Spark pre-built binary 1.1.0 for Hadoop 1.x be assembled from MVN jars, so there should be no mismatching, Am I right? If I am wrong, could anyone tell me the reason of the binary mismatch and tell me the best practice for me to solve my problem?\n\n","from":"developer"},{"body":"[~sunrui] What I can see from this JIRA discussion (and [~srowen] please correct me if I am wrong) is that Hadoop 1 vs. Hadoop 2 is one of the causes of incompatibility. It is _not the only_ reason and I don't think we exactly know why the pre-built binary for 1.1.0 is different from the maven version. \n\nI think the best practice advice is to use the exact same jar in the application and in the runtime. Marking Spark a provided dependency in the application build and using spark-submit is one way of achieving this. Or one can publish a local build to maven and use the same local build to start the cluster etc.","from":"developer"},{"body":"User 'zsxwing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/3758","from":"developer"},{"body":"[~sunrui] Yes, but I am still not clear that anyone has observed the problem with two binaries built for the same version of Hadoop. Most of the situations listed here do not match that description. I might not understand someone's issue report here. In any event, it sounds like an underlying cause has been fixed already anyway.","from":"developer"}],"created":"2014-06-08T19:44:59.000+0000","description":"Running a job built against the Maven dep for 1.0.0 and the hadoop1 distribution produces:\n\n{code}\njava.lang.ClassNotFoundException:\norg.apache.spark.rdd.RDD$$anonfun$saveAsTextFile$1\n{code}\n\nHere's what's in the Maven dep as of 1.0.0:\n\n{code}\njar tvf ~/.m2/repository/org/apache/spark/spark-core_2.10/1.0.0/spark-core_2.10-1.0.0.jar | grep 'rdd/RDD' | grep 'saveAs'\n 1519 Mon May 26 13:57:58 PDT 2014 org/apache/spark/rdd/RDD$anonfun$saveAsTextFile$1.class\n 1560 Mon May 26 13:57:58 PDT 2014 org/apache/spark/rdd/RDD$anonfun$saveAsTextFile$2.class\n{code}\n\nAnd here's what's in the hadoop1 distribution:\n\n{code}\njar tvf spark-assembly-1.0.0-hadoop1.0.4.jar| grep 'rdd/RDD' | grep 'saveAs'\n{code}\n\nI.e., it's not there. It is in the hadoop2 distribution:\n\n{code}\njar tvf spark-assembly-1.0.0-hadoop2.2.0.jar| grep 'rdd/RDD' | grep 'saveAs'\n 1519 Mon May 26 07:29:54 PDT 2014 org/apache/spark/rdd/RDD$anonfun$saveAsTextFile$1.class\n 1560 Mon May 26 07:29:54 PDT 2014 org/apache/spark/rdd/RDD$anonfun$saveAsTextFile$2.class\n{code}","issue_id":"12719084","key":"SPARK-2075","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2014-12-22T06:10:28.000+0000","role":"fixed_distractor","summary":"Anonymous classes are missing from Spark distribution"} {"case_id":"13073754","cluster":"DISTRACTOR-SPARK-20832","comments":[{"body":"I'm working on this.","created":"2017-05-30T17:19:43.286+0000"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18362","created":"2017-06-20T13:25:04.204+0000"},{"body":"Issue resolved by pull request 18362\n[https://github.com/apache/spark/pull/18362]","created":"2017-06-22T12:49:04.867+0000"}],"conversations":[{"body":"In SPARK-17370 (a patch authored by [~ekhliang] and reviewed by me), we added logic to the DAGScheduler to mark external shuffle service instances as unavailable upon task failure when the task failure reason was \"SlaveLost\" and this was known to be caused by worker death. If the Spark Master discovered that a worker was dead then it would notify any drivers with executors on those workers to mark those executors as dead. The linked patch simply piggybacked on this logic to have the executor death notification also imply worker death and to have worker-death-caused-executor-death imply shuffle file loss.\n\nHowever, there are modes of external shuffle service loss which this mechanism does not detect, leaving the system prone race conditions. Consider the following:\n\n* Spark standalone is configured to run an external shuffle service embedded in the Worker.\n* Application has shuffle outputs and executors on Worker A.\n* Stage depending on outputs of tasks that ran on Worker A starts.\n* All executors on worker A are removed due to dying with exceptions, scaling-down via the dynamic allocation APIs, but _not_ due to worker death. Worker A is still healthy at this point.\n* At this point the MapOutputTracker still records map output locations on Worker A's shuffle service. This is expected behavior. \n* Worker A dies at an instant where the application has no executors running on it.\n* The Master knows that Worker A died but does not inform the driver (which had no executors on that worker at the time of its death).\n* Some task from the running stage attempts to fetch map outputs from Worker A but these requests time out because Worker A's shuffle service isn't available.\n* Due to other logic in the scheduler, these preventable FetchFailures don't wind up invaliding the now-invalid unavailable map output locations (this is a distinct bug / behavior which I'll discuss in a separate JIRA ticket).\n* This behavior leads to several unsuccessful stage reattempts and ultimately to a job failure.\n\nA simple way to address this would be to have the Master explicitly notify drivers of all Worker deaths, even if those drivers don't currently have executors. The Spark Standalone scheduler backend can receive the explicit WorkerLost message and can bubble up the right calls to the task scheduler and DAGScheduler to invalidate map output locations from the now-dead external shuffle service.\n\nThis relates to SPARK-20115 in the sense that both tickets aim to address issues where the external shuffle service is unavailable. The key difference is the mechanism for detection: SPARK-20115 marks the external shuffle service as unavailable whenever any fetch failure occurs from it, whereas the proposal here relies on more explicit signals. This JIRA ticket's proposal is scoped only to Spark Standalone mode. As a compromise, we might be able to consider \"all of a single shuffle's outputs lost on a single external shuffle service\" following a fetch failure (to be discussed in separate JIRA). ","from":"reporter","subject":"Standalone master should explicitly inform drivers of worker deaths and invalidate external shuffle service outputs"},{"body":"I'm working on this.","from":"developer"},{"body":"User 'jiangxb1987' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18362","from":"developer"},{"body":"Issue resolved by pull request 18362\n[https://github.com/apache/spark/pull/18362]","from":"developer"}],"created":"2017-05-22T03:19:32.000+0000","description":"In SPARK-17370 (a patch authored by [~ekhliang] and reviewed by me), we added logic to the DAGScheduler to mark external shuffle service instances as unavailable upon task failure when the task failure reason was \"SlaveLost\" and this was known to be caused by worker death. If the Spark Master discovered that a worker was dead then it would notify any drivers with executors on those workers to mark those executors as dead. The linked patch simply piggybacked on this logic to have the executor death notification also imply worker death and to have worker-death-caused-executor-death imply shuffle file loss.\n\nHowever, there are modes of external shuffle service loss which this mechanism does not detect, leaving the system prone race conditions. Consider the following:\n\n* Spark standalone is configured to run an external shuffle service embedded in the Worker.\n* Application has shuffle outputs and executors on Worker A.\n* Stage depending on outputs of tasks that ran on Worker A starts.\n* All executors on worker A are removed due to dying with exceptions, scaling-down via the dynamic allocation APIs, but _not_ due to worker death. Worker A is still healthy at this point.\n* At this point the MapOutputTracker still records map output locations on Worker A's shuffle service. This is expected behavior. \n* Worker A dies at an instant where the application has no executors running on it.\n* The Master knows that Worker A died but does not inform the driver (which had no executors on that worker at the time of its death).\n* Some task from the running stage attempts to fetch map outputs from Worker A but these requests time out because Worker A's shuffle service isn't available.\n* Due to other logic in the scheduler, these preventable FetchFailures don't wind up invaliding the now-invalid unavailable map output locations (this is a distinct bug / behavior which I'll discuss in a separate JIRA ticket).\n* This behavior leads to several unsuccessful stage reattempts and ultimately to a job failure.\n\nA simple way to address this would be to have the Master explicitly notify drivers of all Worker deaths, even if those drivers don't currently have executors. The Spark Standalone scheduler backend can receive the explicit WorkerLost message and can bubble up the right calls to the task scheduler and DAGScheduler to invalidate map output locations from the now-dead external shuffle service.\n\nThis relates to SPARK-20115 in the sense that both tickets aim to address issues where the external shuffle service is unavailable. The key difference is the mechanism for detection: SPARK-20115 marks the external shuffle service as unavailable whenever any fetch failure occurs from it, whereas the proposal here relies on more explicit signals. This JIRA ticket's proposal is scoped only to Spark Standalone mode. As a compromise, we might be able to consider \"all of a single shuffle's outputs lost on a single external shuffle service\" following a fetch failure (to be discussed in separate JIRA). ","issue_id":"13073754","key":"SPARK-20832","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-06-22T12:49:04.000+0000","role":"fixed_distractor","summary":"Standalone master should explicitly inform drivers of worker deaths and invalidate external shuffle service outputs"} {"case_id":"13084233","cluster":"DISTRACTOR-SPARK-21287","comments":[{"body":"User 'maver1ck' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18515","created":"2017-07-03T12:51:03.581+0000"},{"body":"Yeah, I'm familiar with this special value.\n\nIs it not equivalent to setting fetch size to 1? or 0? I don't think so, but I don't recall why. Would it work in your situation and therefore be the right way to do this for, likely, most callers?\n\nI can see an argument for removing this assertion, but, the JDBC API says that negative values aren't allowed. This is really nonstandard MySQL behavior, and it'd be good to find a better way to recommend to users.","created":"2017-07-03T12:54:10.049+0000"},{"body":"No. It's not the same like setting 0 or 1.\nEvery other value makes MySQL JDBC driver to store whole ResultSet in memory.\n\nMaybe we can just remove this assertion ?\nI think this is not so super popular configuration option.","created":"2017-07-03T13:04:47.718+0000"},{"body":"It's not supposed to do that right -- you're saying the MySQL driver doesn't quite use the value as intended?\nHm, so any MySQL user would have the whole resultset pulled back into memory every time? that seems like a problem, yeah, but surely MySQL has fixed this?\nIf not, I could see this being a big enough deal that it has to be allowed as a value even if it is wrong.","created":"2017-07-03T13:10:29.820+0000"},{"body":"Quote\n{quote}\nBy default, ResultSets are completely retrieved and stored in memory. In most cases this is the most efficient way to operate and, due to the design of the MySQL network protocol, is easier to implement. If you are working with ResultSets that have a large number of rows or large values and cannot allocate heap space in your JVM for the memory required, you can tell the driver to stream the results back one row at a time.\n{quote}\nhttps://dev.mysql.com/doc/connector-j/5.1/en/connector-j-reference-implementation-notes.html","created":"2017-07-03T13:59:15.157+0000"},{"body":"I know, but this is what fetch size is supposed to control of course. It need not be all or nothing. I am not sure why they didn't just let a fetch size of 1 select this behavior. Does any nonpositive value actually get this behavior? That would be nice.","created":"2017-07-03T14:14:49.900+0000"},{"body":"This value is very specific to MySQL. Since we are supporting different dialects, we could introduce a dialect-specific checking logics.","created":"2017-07-03T16:46:13.904+0000"},{"body":"Yeah, maybe a method to validate or even set the fetch size that is overridden by the MySQL dialect. That could be used where {{stmt.setFetchSize(options.fetchSize)}} is called in JDBCRDD. And then remove the validation in JDBCOptions.","created":"2017-07-03T17:30:22.472+0000"},{"body":"Do we have any following up for this issue? MySQL is a tier-one database. ","created":"2019-01-28T19:57:11.186+0000"},{"body":"[~bestcastor] Feel free to submit a PR ","created":"2019-01-28T19:59:50.326+0000"},{"body":"[~smilegator] Ok. Anything we should follow up to submit the PR? ","created":"2019-01-28T20:01:33.856+0000"},{"body":"[~smilegator]  [~srowen] Just submitted a PR for this : [https://github.com/apache/spark/pull/26230]\r\n\r\nPlease help review.","created":"2019-10-23T15:02:31.078+0000"},{"body":"Issue resolved by pull request 26244\n[https://github.com/apache/spark/pull/26244]","created":"2019-10-24T19:36:24.935+0000"}],"conversations":[{"body":"MySQL JDBC driver gives possibility to not store ResultSet in memory.\nWe can do this by setting fetchSize to Int.MIN_VALUE.\nUnfortunately this configuration isn't correct in Spark.\n{code}\njava.lang.IllegalArgumentException: requirement failed: Invalid value `-2147483648` for parameter `fetchsize`. The minimum value is 0. When the value is 0, the JDBC driver ignores the value and does the estimates.\n\tat scala.Predef$.require(Predef.scala:224)\n\tat org.apache.spark.sql.execution.datasources.jdbc.JDBCOptions.(JDBCOptions.scala:105)\n\tat org.apache.spark.sql.execution.datasources.jdbc.JDBCOptions.(JDBCOptions.scala:34)\n\tat org.apache.spark.sql.execution.datasources.jdbc.JdbcRelationProvider.createRelation(JdbcRelationProvider.scala:32)\n\tat org.apache.spark.sql.execution.datasources.DataSource.resolveRelation(DataSource.scala:330)\n\tat org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:152)\n\tat org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:125)\n\tat org.apache.spark.sql.DataFrameReader.jdbc(DataFrameReader.scala:166)\n\tat org.apache.spark.sql.DataFrameReader.jdbc(DataFrameReader.scala:206)\n\tat sun.reflect.GeneratedMethodAccessor46.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:244)\n\tat py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:357)\n\tat py4j.Gateway.invoke(Gateway.java:280)\n\tat py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:132)\n\tat py4j.commands.CallCommand.execute(CallCommand.java:79)\n\tat py4j.GatewayConnection.run(GatewayConnection.java:214)\n\tat java.lang.Thread.run(Thread.java:748)\n{code}\n\nhttps://dev.mysql.com/doc/connector-j/5.1/en/connector-j-reference-implementation-notes.html","from":"reporter","subject":"Cannot use Int.MIN_VALUE as Spark SQL fetchsize"},{"body":"User 'maver1ck' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18515","from":"developer"},{"body":"Yeah, I'm familiar with this special value.\n\nIs it not equivalent to setting fetch size to 1? or 0? I don't think so, but I don't recall why. Would it work in your situation and therefore be the right way to do this for, likely, most callers?\n\nI can see an argument for removing this assertion, but, the JDBC API says that negative values aren't allowed. This is really nonstandard MySQL behavior, and it'd be good to find a better way to recommend to users.","from":"developer"},{"body":"No. It's not the same like setting 0 or 1.\nEvery other value makes MySQL JDBC driver to store whole ResultSet in memory.\n\nMaybe we can just remove this assertion ?\nI think this is not so super popular configuration option.","from":"developer"},{"body":"It's not supposed to do that right -- you're saying the MySQL driver doesn't quite use the value as intended?\nHm, so any MySQL user would have the whole resultset pulled back into memory every time? that seems like a problem, yeah, but surely MySQL has fixed this?\nIf not, I could see this being a big enough deal that it has to be allowed as a value even if it is wrong.","from":"developer"},{"body":"Quote\n{quote}\nBy default, ResultSets are completely retrieved and stored in memory. In most cases this is the most efficient way to operate and, due to the design of the MySQL network protocol, is easier to implement. If you are working with ResultSets that have a large number of rows or large values and cannot allocate heap space in your JVM for the memory required, you can tell the driver to stream the results back one row at a time.\n{quote}\nhttps://dev.mysql.com/doc/connector-j/5.1/en/connector-j-reference-implementation-notes.html","from":"developer"},{"body":"I know, but this is what fetch size is supposed to control of course. It need not be all or nothing. I am not sure why they didn't just let a fetch size of 1 select this behavior. Does any nonpositive value actually get this behavior? That would be nice.","from":"developer"},{"body":"This value is very specific to MySQL. Since we are supporting different dialects, we could introduce a dialect-specific checking logics.","from":"developer"},{"body":"Yeah, maybe a method to validate or even set the fetch size that is overridden by the MySQL dialect. That could be used where {{stmt.setFetchSize(options.fetchSize)}} is called in JDBCRDD. And then remove the validation in JDBCOptions.","from":"developer"},{"body":"Do we have any following up for this issue? MySQL is a tier-one database. ","from":"developer"},{"body":"[~bestcastor] Feel free to submit a PR ","from":"developer"},{"body":"[~smilegator] Ok. Anything we should follow up to submit the PR? ","from":"developer"},{"body":"[~smilegator]  [~srowen] Just submitted a PR for this : [https://github.com/apache/spark/pull/26230]\r\n\r\nPlease help review.","from":"developer"},{"body":"Issue resolved by pull request 26244\n[https://github.com/apache/spark/pull/26244]","from":"developer"}],"created":"2017-07-03T12:33:16.000+0000","description":"MySQL JDBC driver gives possibility to not store ResultSet in memory.\nWe can do this by setting fetchSize to Int.MIN_VALUE.\nUnfortunately this configuration isn't correct in Spark.\n{code}\njava.lang.IllegalArgumentException: requirement failed: Invalid value `-2147483648` for parameter `fetchsize`. The minimum value is 0. When the value is 0, the JDBC driver ignores the value and does the estimates.\n\tat scala.Predef$.require(Predef.scala:224)\n\tat org.apache.spark.sql.execution.datasources.jdbc.JDBCOptions.(JDBCOptions.scala:105)\n\tat org.apache.spark.sql.execution.datasources.jdbc.JDBCOptions.(JDBCOptions.scala:34)\n\tat org.apache.spark.sql.execution.datasources.jdbc.JdbcRelationProvider.createRelation(JdbcRelationProvider.scala:32)\n\tat org.apache.spark.sql.execution.datasources.DataSource.resolveRelation(DataSource.scala:330)\n\tat org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:152)\n\tat org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:125)\n\tat org.apache.spark.sql.DataFrameReader.jdbc(DataFrameReader.scala:166)\n\tat org.apache.spark.sql.DataFrameReader.jdbc(DataFrameReader.scala:206)\n\tat sun.reflect.GeneratedMethodAccessor46.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:498)\n\tat py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:244)\n\tat py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:357)\n\tat py4j.Gateway.invoke(Gateway.java:280)\n\tat py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:132)\n\tat py4j.commands.CallCommand.execute(CallCommand.java:79)\n\tat py4j.GatewayConnection.run(GatewayConnection.java:214)\n\tat java.lang.Thread.run(Thread.java:748)\n{code}\n\nhttps://dev.mysql.com/doc/connector-j/5.1/en/connector-j-reference-implementation-notes.html","issue_id":"13084233","key":"SPARK-21287","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-10-24T19:36:24.000+0000","role":"fixed_distractor","summary":"Cannot use Int.MIN_VALUE as Spark SQL fetchsize"} {"case_id":"13086291","cluster":"DISTRACTOR-SPARK-21374","comments":[{"body":"Yeah, org.apache.spark.deploy.SparkHadoopUtil.globPath uses a wrong Hadoop configuration. Welcome to submit a PR to fix it.\n\nRight now as a workaround, you can use the following codes to set your keys:\n\n{code}\nval conf = org.apache.spark.deploy.SparkHadoopUtil.get.conf\nconf.set(...)\n{code}","created":"2017-07-12T21:43:58.084+0000"},{"body":"User 'andrey-tpt' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18623","created":"2017-07-13T11:54:03.964+0000"},{"body":"This is possibly a sign that your new configuration isn't having its auth values picked up, or they are incorrect (i.e properties are wrong). Its working for enabled caching as some other codepath has set them up with the right properties, and so when used in the DF, the previous params are picked up.\n\n# if using Hadoop 2.7.x JARs, switch to s3a and use s3a in the URLs & settings. You don't need to set the fs.s3a.impl field either; done for yoiu.\n# if you can upgrade to Hadoop 2.8 binaries, you can use per-bucket configuration; this does exactly what you want: lets you configure different auth details for different buckets, without having to play these games. See [https://hadoop.apache.org/docs/current/hadoop-aws/tools/hadoop-aws/index.html#Configuring_different_S3_buckets] and [https://docs.hortonworks.com/HDPDocuments/HDP2/HDP-2.6.1/bk_cloud-data-access/content/s3-per-bucket-configs.html]\n\n\ngoing to 2.8 binaries (or anything with the feature backported to a 2.7.x variant) should solve your problem without you having to worry about what you are seeing here.","created":"2017-07-13T12:37:48.626+0000"},{"body":"[~stevel@apache.org] \n\nIndeed, while working on PR and debugging the code I see that code works only accidentally because caching is turned on by default.\n\n1. Thanks for the advice. I doubt that it's related to type of the filesystem - I've only mentioned filesystem explicitly to show why \"awsAcceesKeyId\" not \"access.key\" is used with \"s3\" scheme in the example. Sorry for the confusion. \n \n2. Unfortunately, it's not that easy - this example is only simplified version of what happens in my project. We don't have information about which buckets user will try to access in interactive mode so I can not enumerate them all in configuration. ","created":"2017-07-19T17:05:50.643+0000"},{"body":"I understand...the patch shows the issue. Its only working in some codepaths because the (authenticated) S3 FS instance was already created and cached.","created":"2017-08-01T13:35:58.804+0000"}],"conversations":[{"body":"*Motivation:*\nIn my case I want to disable filesystem cache to be able to change S3's access key and secret key on the fly to read from buckets with different permissions. This works perfectly fine for RDDs but doesn't work for DFs.\n\n*Example (works for RDD but fails for DataFrame):*\n\n{code:java}\nimport org.apache.spark.SparkContext\nimport org.apache.spark.SparkConf\nimport org.apache.spark.sql.SparkSession\n\nobject SimpleApp {\n def main(args: Array[String]) {\n\n val awsAccessKeyId = \"something\"\n val awsSecretKey = \"something else\"\n\n val conf = new SparkConf().setAppName(\"Simple Application\").setMaster(\"local[*]\")\n\n val sc = new SparkContext(conf)\n sc.hadoopConfiguration.set(\"fs.s3.awsAccessKeyId\", awsAccessKeyId)\n sc.hadoopConfiguration.set(\"fs.s3.awsSecretAccessKey\", awsSecretKey)\n sc.hadoopConfiguration.setBoolean(\"fs.s3.impl.disable.cache\",true)\n sc.hadoopConfiguration.set(\"fs.s3.impl\",\"org.apache.hadoop.fs.s3native.NativeS3FileSystem\")\n sc.hadoopConfiguration.set(\"fs.s3.buffer.dir\",\"/tmp\")\n\n val spark = SparkSession.builder().config(conf).getOrCreate()\n\n val rddFile = sc.textFile(\"s3://bucket/file.csv\").count // ok\n val rddGlob = sc.textFile(\"s3://bucket/*\").count // ok\n val dfFile = spark.read.format(\"csv\").load(\"s3://bucket/file.csv\").count // ok\n \n val dfGlob = spark.read.format(\"csv\").load(\"s3://bucket/*\").count \n // IllegalArgumentExcepton. AWS Access Key ID and Secret Access Key must be specified as the username or password (respectively)\n // of a s3 URL, or by setting the fs.s3.awsAccessKeyId or fs.s3.awsSecretAccessKey properties (respectively).\n \n sc.stop()\n }\n}\n\n{code}","from":"reporter","subject":"Reading globbed paths from S3 into DF doesn't work if filesystem caching is disabled"},{"body":"Yeah, org.apache.spark.deploy.SparkHadoopUtil.globPath uses a wrong Hadoop configuration. Welcome to submit a PR to fix it.\n\nRight now as a workaround, you can use the following codes to set your keys:\n\n{code}\nval conf = org.apache.spark.deploy.SparkHadoopUtil.get.conf\nconf.set(...)\n{code}","from":"developer"},{"body":"User 'andrey-tpt' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/18623","from":"developer"},{"body":"This is possibly a sign that your new configuration isn't having its auth values picked up, or they are incorrect (i.e properties are wrong). Its working for enabled caching as some other codepath has set them up with the right properties, and so when used in the DF, the previous params are picked up.\n\n# if using Hadoop 2.7.x JARs, switch to s3a and use s3a in the URLs & settings. You don't need to set the fs.s3a.impl field either; done for yoiu.\n# if you can upgrade to Hadoop 2.8 binaries, you can use per-bucket configuration; this does exactly what you want: lets you configure different auth details for different buckets, without having to play these games. See [https://hadoop.apache.org/docs/current/hadoop-aws/tools/hadoop-aws/index.html#Configuring_different_S3_buckets] and [https://docs.hortonworks.com/HDPDocuments/HDP2/HDP-2.6.1/bk_cloud-data-access/content/s3-per-bucket-configs.html]\n\n\ngoing to 2.8 binaries (or anything with the feature backported to a 2.7.x variant) should solve your problem without you having to worry about what you are seeing here.","from":"developer"},{"body":"[~stevel@apache.org] \n\nIndeed, while working on PR and debugging the code I see that code works only accidentally because caching is turned on by default.\n\n1. Thanks for the advice. I doubt that it's related to type of the filesystem - I've only mentioned filesystem explicitly to show why \"awsAcceesKeyId\" not \"access.key\" is used with \"s3\" scheme in the example. Sorry for the confusion. \n \n2. Unfortunately, it's not that easy - this example is only simplified version of what happens in my project. We don't have information about which buckets user will try to access in interactive mode so I can not enumerate them all in configuration. ","from":"developer"},{"body":"I understand...the patch shows the issue. Its only working in some codepaths because the (authenticated) S3 FS instance was already created and cached.","from":"developer"}],"created":"2017-07-11T15:24:23.000+0000","description":"*Motivation:*\nIn my case I want to disable filesystem cache to be able to change S3's access key and secret key on the fly to read from buckets with different permissions. This works perfectly fine for RDDs but doesn't work for DFs.\n\n*Example (works for RDD but fails for DataFrame):*\n\n{code:java}\nimport org.apache.spark.SparkContext\nimport org.apache.spark.SparkConf\nimport org.apache.spark.sql.SparkSession\n\nobject SimpleApp {\n def main(args: Array[String]) {\n\n val awsAccessKeyId = \"something\"\n val awsSecretKey = \"something else\"\n\n val conf = new SparkConf().setAppName(\"Simple Application\").setMaster(\"local[*]\")\n\n val sc = new SparkContext(conf)\n sc.hadoopConfiguration.set(\"fs.s3.awsAccessKeyId\", awsAccessKeyId)\n sc.hadoopConfiguration.set(\"fs.s3.awsSecretAccessKey\", awsSecretKey)\n sc.hadoopConfiguration.setBoolean(\"fs.s3.impl.disable.cache\",true)\n sc.hadoopConfiguration.set(\"fs.s3.impl\",\"org.apache.hadoop.fs.s3native.NativeS3FileSystem\")\n sc.hadoopConfiguration.set(\"fs.s3.buffer.dir\",\"/tmp\")\n\n val spark = SparkSession.builder().config(conf).getOrCreate()\n\n val rddFile = sc.textFile(\"s3://bucket/file.csv\").count // ok\n val rddGlob = sc.textFile(\"s3://bucket/*\").count // ok\n val dfFile = spark.read.format(\"csv\").load(\"s3://bucket/file.csv\").count // ok\n \n val dfGlob = spark.read.format(\"csv\").load(\"s3://bucket/*\").count \n // IllegalArgumentExcepton. AWS Access Key ID and Secret Access Key must be specified as the username or password (respectively)\n // of a s3 URL, or by setting the fs.s3.awsAccessKeyId or fs.s3.awsSecretAccessKey properties (respectively).\n \n sc.stop()\n }\n}\n\n{code}","issue_id":"13086291","key":"SPARK-21374","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-08-05T05:41:10.000+0000","role":"fixed_distractor","summary":"Reading globbed paths from S3 into DF doesn't work if filesystem caching is disabled"} {"case_id":"13090560","cluster":"DISTRACTOR-SPARK-21549","comments":[{"body":"This affects both mapred (\"mapred.output.dir\") and mapreduce (\"mapreduce.output.fileoutputformat.outputdir\") based OutputFormat's which do not set the properties referenced and is an incompatibility introduced in spark 2.2\n\nWorkaround is to explicitly set the property to a dummy value (which is valid and writable by user - say /tmp).\n\n+CC [~WeiqingYang] \n\n","created":"2017-07-28T20:38:38.545+0000"},{"body":"[~mridulm80], [~WeiqingYang] does it make sense to implement the provided workaround with valid and writable directory just within SparkHadoopMapRedceWriter if the needed property is not set, to prevent affecting all the jobs which don't write to hdfs, at least until there is a better solution?","created":"2017-08-21T18:23:19.134+0000"},{"body":"Linking to SPARK-20045, which highlights the commit logic, especially the abort code, needs to be resilient to failures, including that of the invoked {{OutputCommitter.abort()}} calls from raising exceptions. While people implementing committers should be expected to write resilient abort routines, you can't rely on it. Same for calling fs.delete()...it could also fall, so wrapping everything in an exception handler would at least make abort resilient.","created":"2017-09-19T09:24:53.487+0000"},{"body":"# you can't rely on the committers having output and temp dirs. Subclasses of {{FileOutputCommitter}} *must*, though there's no official mechanism for querying that because {{getOutputPath()}} is private.\n# Hadoop 3.0 has added (MAPREDUCE-6956) a new superclass of {{FileOutputCommitter}}, [PathOutputCommitter|[https://github.com/apache/hadoop/blob/trunk/hadoop-mapreduce-project/hadoop-mapreduce-client/hadoop-mapreduce-client-core/src/main/java/org/apache/hadoop/mapreduce/lib/output/PathOutputCommitter.java] which pulls up getWorkingDir method to be more general (so that you can have output committers which tell spark and hive where their intermediate data should go, without them being subclasses of FileOutputCommitter.\n# I'm happy to pull up {{getOutputPath}} to that class, and if people can give me a patch for it *this week* I'll add it for 3.0 beta 1. \n\nRegarding the committer here, you might want to think of moving off/subclassing {{HadoopMapReduceCommitProtocol}}. This is what I've done in [PathOutputCommitProtoco|https://github.com/hortonworks-spark/cloud-integration/blob/master/spark-cloud-integration/src/main/scala/com/hortonworks/spark/cloud/PathOutputCommitProtocol.scala], though I can see it's still relying on the superclass to get that properly. Again, if we can patch the new PathOutputCommitter class for a getOutputPath I use that here. And yes, while that new mapreduce will take a long time to surface in spark core, you can use it independently, from later this year..\n\nIf you are playing with different committers out of spark's own codebase, pick up the the ORC hive tests from [https://github.com/hortonworks-spark/cloud-integration/tree/master/cloud-examples/src/test/scala/org/apache/spark/sql/sources]. These are just some of the spark sql tests reworked slightly so that they'll work with any FileSystem impl. rather than just local file:// paths\n\nPing me direct if you are playing with new committers, & look at MAPREDUCE-6823 to see if that'd be useful to you (& how it could be improved, given its still not in the codebase)\n\n","created":"2017-09-19T09:43:06.953+0000"},{"body":"User 'szhem' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19294","created":"2017-09-20T13:29:04.050+0000"},{"body":"[~mridulm80], [~WeiqingYang], [~stevel@apache.org], \n\nI've implemented the fix in this PR (https://github.com/apache/spark/pull/19294), which sets user's current working directory (which is typically her home directory in case of distributed filesystems) as output directory.\n\nThe patch allows using OutputFormats which write to external systems, databases, etc. by means of RDD API.\nI far as I understand the requirement for output paths to be specified is only necessary to allow files to be committed to an absolute output location, that is not the case for output formats which write data to external systems. \nSo using user's working directory for such situations seems to be ok.\n \n","created":"2017-09-20T13:38:20.920+0000"},{"body":"The {{newTaskTempFileAbsPath()}} method is an interesting spot of code...I'm still trying to work out when it is actually used. Some committers like {{ManifestFileCommitProtocol}} don't support it all. \n\nHowever, if it is used, then your patch is going to cause problems if the dest FS != the default FS, because then the bit of the protocol which takes that list of temp files and renames() them into their destination is going to fail. I think you'd be better off having the committer fail fast when an absolute path is asked for","created":"2017-09-20T16:29:30.241+0000"},{"body":"[~stevel@apache.org] I've updated PR to prevent using FileSystems at all. \nInstead, there is just an additional check whether there are absolute files to rename during commit.","created":"2017-09-20T20:05:52.172+0000"},{"body":"Issue resolved by pull request 19294\n[https://github.com/apache/spark/pull/19294]","created":"2017-10-07T03:46:25.948+0000"},{"body":"User 'mridulm' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19487","created":"2017-10-13T00:29:03.408+0000"},{"body":"User 'mridulm' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19497","created":"2017-10-14T06:24:04.061+0000"}],"conversations":[{"body":"Spark fails to complete job correctly in case of custom OutputFormat implementations.\n\nThere are OutputFormat implementations which do not need to use *mapreduce.output.fileoutputformat.outputdir* standard hadoop property.\n\n[But spark reads this property from the configuration|https://github.com/apache/spark/blob/v2.2.0/core/src/main/scala/org/apache/spark/internal/io/SparkHadoopMapReduceWriter.scala#L79] while setting up an OutputCommitter\n{code:javascript}\nval committer = FileCommitProtocol.instantiate(\n className = classOf[HadoopMapReduceCommitProtocol].getName,\n jobId = stageId.toString,\n outputPath = conf.value.get(\"mapreduce.output.fileoutputformat.outputdir\"),\n isAppend = false).asInstanceOf[HadoopMapReduceCommitProtocol]\ncommitter.setupJob(jobContext)\n{code}\n... and then uses this property later on while [commiting the job|https://github.com/apache/spark/blob/v2.2.0/core/src/main/scala/org/apache/spark/internal/io/HadoopMapReduceCommitProtocol.scala#L132], [aborting the job|https://github.com/apache/spark/blob/v2.2.0/core/src/main/scala/org/apache/spark/internal/io/HadoopMapReduceCommitProtocol.scala#L141], [creating task's temp path|https://github.com/apache/spark/blob/v2.2.0/core/src/main/scala/org/apache/spark/internal/io/HadoopMapReduceCommitProtocol.scala#L95]\n\nIn that cases when the job completes then following exception is thrown\n{code}\nCan not create a Path from a null string\njava.lang.IllegalArgumentException: Can not create a Path from a null string\n at org.apache.hadoop.fs.Path.checkPathArg(Path.java:123)\n at org.apache.hadoop.fs.Path.(Path.java:135)\n at org.apache.hadoop.fs.Path.(Path.java:89)\n at org.apache.spark.internal.io.HadoopMapReduceCommitProtocol.absPathStagingDir(HadoopMapReduceCommitProtocol.scala:58)\n at org.apache.spark.internal.io.HadoopMapReduceCommitProtocol.abortJob(HadoopMapReduceCommitProtocol.scala:141)\n at org.apache.spark.internal.io.SparkHadoopMapReduceWriter$.write(SparkHadoopMapReduceWriter.scala:106)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$saveAsNewAPIHadoopDataset$1.apply$mcV$sp(PairRDDFunctions.scala:1085)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$saveAsNewAPIHadoopDataset$1.apply(PairRDDFunctions.scala:1085)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$saveAsNewAPIHadoopDataset$1.apply(PairRDDFunctions.scala:1085)\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:112)\n at org.apache.spark.rdd.RDD.withScope(RDD.scala:362)\n at org.apache.spark.rdd.PairRDDFunctions.saveAsNewAPIHadoopDataset(PairRDDFunctions.scala:1084)\n ...\n{code}\n\nSo it seems that all the jobs which use OutputFormats which don't write data into HDFS-compatible file systems are broken.","from":"reporter","subject":"Spark fails to complete job correctly in case of OutputFormat which do not write into hdfs"},{"body":"This affects both mapred (\"mapred.output.dir\") and mapreduce (\"mapreduce.output.fileoutputformat.outputdir\") based OutputFormat's which do not set the properties referenced and is an incompatibility introduced in spark 2.2\n\nWorkaround is to explicitly set the property to a dummy value (which is valid and writable by user - say /tmp).\n\n+CC [~WeiqingYang] \n\n","from":"developer"},{"body":"[~mridulm80], [~WeiqingYang] does it make sense to implement the provided workaround with valid and writable directory just within SparkHadoopMapRedceWriter if the needed property is not set, to prevent affecting all the jobs which don't write to hdfs, at least until there is a better solution?","from":"developer"},{"body":"Linking to SPARK-20045, which highlights the commit logic, especially the abort code, needs to be resilient to failures, including that of the invoked {{OutputCommitter.abort()}} calls from raising exceptions. While people implementing committers should be expected to write resilient abort routines, you can't rely on it. Same for calling fs.delete()...it could also fall, so wrapping everything in an exception handler would at least make abort resilient.","from":"developer"},{"body":"# you can't rely on the committers having output and temp dirs. Subclasses of {{FileOutputCommitter}} *must*, though there's no official mechanism for querying that because {{getOutputPath()}} is private.\n# Hadoop 3.0 has added (MAPREDUCE-6956) a new superclass of {{FileOutputCommitter}}, [PathOutputCommitter|[https://github.com/apache/hadoop/blob/trunk/hadoop-mapreduce-project/hadoop-mapreduce-client/hadoop-mapreduce-client-core/src/main/java/org/apache/hadoop/mapreduce/lib/output/PathOutputCommitter.java] which pulls up getWorkingDir method to be more general (so that you can have output committers which tell spark and hive where their intermediate data should go, without them being subclasses of FileOutputCommitter.\n# I'm happy to pull up {{getOutputPath}} to that class, and if people can give me a patch for it *this week* I'll add it for 3.0 beta 1. \n\nRegarding the committer here, you might want to think of moving off/subclassing {{HadoopMapReduceCommitProtocol}}. This is what I've done in [PathOutputCommitProtoco|https://github.com/hortonworks-spark/cloud-integration/blob/master/spark-cloud-integration/src/main/scala/com/hortonworks/spark/cloud/PathOutputCommitProtocol.scala], though I can see it's still relying on the superclass to get that properly. Again, if we can patch the new PathOutputCommitter class for a getOutputPath I use that here. And yes, while that new mapreduce will take a long time to surface in spark core, you can use it independently, from later this year..\n\nIf you are playing with different committers out of spark's own codebase, pick up the the ORC hive tests from [https://github.com/hortonworks-spark/cloud-integration/tree/master/cloud-examples/src/test/scala/org/apache/spark/sql/sources]. These are just some of the spark sql tests reworked slightly so that they'll work with any FileSystem impl. rather than just local file:// paths\n\nPing me direct if you are playing with new committers, & look at MAPREDUCE-6823 to see if that'd be useful to you (& how it could be improved, given its still not in the codebase)\n\n","from":"developer"},{"body":"User 'szhem' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19294","from":"developer"},{"body":"[~mridulm80], [~WeiqingYang], [~stevel@apache.org], \n\nI've implemented the fix in this PR (https://github.com/apache/spark/pull/19294), which sets user's current working directory (which is typically her home directory in case of distributed filesystems) as output directory.\n\nThe patch allows using OutputFormats which write to external systems, databases, etc. by means of RDD API.\nI far as I understand the requirement for output paths to be specified is only necessary to allow files to be committed to an absolute output location, that is not the case for output formats which write data to external systems. \nSo using user's working directory for such situations seems to be ok.\n \n","from":"developer"},{"body":"The {{newTaskTempFileAbsPath()}} method is an interesting spot of code...I'm still trying to work out when it is actually used. Some committers like {{ManifestFileCommitProtocol}} don't support it all. \n\nHowever, if it is used, then your patch is going to cause problems if the dest FS != the default FS, because then the bit of the protocol which takes that list of temp files and renames() them into their destination is going to fail. I think you'd be better off having the committer fail fast when an absolute path is asked for","from":"developer"},{"body":"[~stevel@apache.org] I've updated PR to prevent using FileSystems at all. \nInstead, there is just an additional check whether there are absolute files to rename during commit.","from":"developer"},{"body":"Issue resolved by pull request 19294\n[https://github.com/apache/spark/pull/19294]","from":"developer"},{"body":"User 'mridulm' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19487","from":"developer"},{"body":"User 'mridulm' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19497","from":"developer"}],"created":"2017-07-27T16:45:58.000+0000","description":"Spark fails to complete job correctly in case of custom OutputFormat implementations.\n\nThere are OutputFormat implementations which do not need to use *mapreduce.output.fileoutputformat.outputdir* standard hadoop property.\n\n[But spark reads this property from the configuration|https://github.com/apache/spark/blob/v2.2.0/core/src/main/scala/org/apache/spark/internal/io/SparkHadoopMapReduceWriter.scala#L79] while setting up an OutputCommitter\n{code:javascript}\nval committer = FileCommitProtocol.instantiate(\n className = classOf[HadoopMapReduceCommitProtocol].getName,\n jobId = stageId.toString,\n outputPath = conf.value.get(\"mapreduce.output.fileoutputformat.outputdir\"),\n isAppend = false).asInstanceOf[HadoopMapReduceCommitProtocol]\ncommitter.setupJob(jobContext)\n{code}\n... and then uses this property later on while [commiting the job|https://github.com/apache/spark/blob/v2.2.0/core/src/main/scala/org/apache/spark/internal/io/HadoopMapReduceCommitProtocol.scala#L132], [aborting the job|https://github.com/apache/spark/blob/v2.2.0/core/src/main/scala/org/apache/spark/internal/io/HadoopMapReduceCommitProtocol.scala#L141], [creating task's temp path|https://github.com/apache/spark/blob/v2.2.0/core/src/main/scala/org/apache/spark/internal/io/HadoopMapReduceCommitProtocol.scala#L95]\n\nIn that cases when the job completes then following exception is thrown\n{code}\nCan not create a Path from a null string\njava.lang.IllegalArgumentException: Can not create a Path from a null string\n at org.apache.hadoop.fs.Path.checkPathArg(Path.java:123)\n at org.apache.hadoop.fs.Path.(Path.java:135)\n at org.apache.hadoop.fs.Path.(Path.java:89)\n at org.apache.spark.internal.io.HadoopMapReduceCommitProtocol.absPathStagingDir(HadoopMapReduceCommitProtocol.scala:58)\n at org.apache.spark.internal.io.HadoopMapReduceCommitProtocol.abortJob(HadoopMapReduceCommitProtocol.scala:141)\n at org.apache.spark.internal.io.SparkHadoopMapReduceWriter$.write(SparkHadoopMapReduceWriter.scala:106)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$saveAsNewAPIHadoopDataset$1.apply$mcV$sp(PairRDDFunctions.scala:1085)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$saveAsNewAPIHadoopDataset$1.apply(PairRDDFunctions.scala:1085)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$saveAsNewAPIHadoopDataset$1.apply(PairRDDFunctions.scala:1085)\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:112)\n at org.apache.spark.rdd.RDD.withScope(RDD.scala:362)\n at org.apache.spark.rdd.PairRDDFunctions.saveAsNewAPIHadoopDataset(PairRDDFunctions.scala:1084)\n ...\n{code}\n\nSo it seems that all the jobs which use OutputFormats which don't write data into HDFS-compatible file systems are broken.","issue_id":"13090560","key":"SPARK-21549","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-10-07T03:46:25.000+0000","role":"fixed_distractor","summary":"Spark fails to complete job correctly in case of OutputFormat which do not write into hdfs"} {"case_id":"13090886","cluster":"DISTRACTOR-SPARK-21565","comments":[{"body":"I believe you need to use a window to group by your event time.","created":"2017-07-31T20:47:43.702+0000"},{"body":"I am not sure, does it mean we can not use watermark without windows - kind of on a microbatch level ? \n\nIf that is the case, it should not work with current_timestamp and should have validation not to allow without windows.","created":"2017-07-31T21:14:08.043+0000"},{"body":"No nothing like the limitations of microbatches. The window can be made trivially small if you want only one timestamp per group, for example {{window(eventTime, \"1 microsecond\")}}\n\nAnd yes this should probably be checked in analysis if this is the intended limitation.","created":"2017-07-31T21:36:56.421+0000"},{"body":"Thanks for reporting it. I can reproduce the error in a unit test. Still investigating it.","created":"2017-08-02T18:12:35.674+0000"},{"body":"Resolved by https://github.com/apache/spark/pull/18840","created":"2017-08-07T20:06:05.240+0000"}],"conversations":[{"body":"*Short Description: *\n\nAggregation query fails with eventTime as watermark column while works with newTimeStamp column generated by running SQL with current_timestamp,\n\n*Exception:*\n\n{code}\nCaused by: java.util.NoSuchElementException: None.get\n\tat scala.None$.get(Option.scala:347)\n\tat scala.None$.get(Option.scala:345)\n\tat org.apache.spark.sql.execution.streaming.StateStoreSaveExec$$anonfun$doExecute$3.apply(statefulOperators.scala:204)\n\tat org.apache.spark.sql.execution.streaming.StateStoreSaveExec$$anonfun$doExecute$3.apply(statefulOperators.scala:172)\n\tat org.apache.spark.sql.execution.streaming.state.package$StateStoreOps$$anonfun$1.apply(package.scala:70)\n\tat org.apache.spark.sql.execution.streaming.state.package$StateStoreOps$$anonfun$1.apply(package.scala:65)\n\tat org.apache.spark.sql.execution.streaming.state.StateStoreRDD.compute(StateStoreRDD.scala:64)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n{code}\n\n*Code to replicate:*\n\n{code}\npackage test\n\nimport java.nio.file.{Files, Path, Paths}\nimport java.text.SimpleDateFormat\n\nimport org.apache.spark.sql.types._\nimport org.apache.spark.sql.{SparkSession}\n\nimport scala.collection.JavaConverters._\n\nobject Test1 {\n\n def main(args: Array[String]) {\n\n val sparkSession = SparkSession\n .builder()\n .master(\"local[*]\")\n .appName(\"Spark SQL basic example\")\n .config(\"spark.some.config.option\", \"some-value\")\n .getOrCreate()\n\n val sdf = new SimpleDateFormat(\"yyyy-MM-dd HH:mm:ss\")\n val checkpointPath = \"target/cp1\"\n val newEventsPath = Paths.get(\"target/newEvents/\").toAbsolutePath\n delete(newEventsPath)\n delete(Paths.get(checkpointPath).toAbsolutePath)\n Files.createDirectories(newEventsPath)\n\n\n val dfNewEvents= newEvents(sparkSession)\n dfNewEvents.createOrReplaceTempView(\"dfNewEvents\")\n\n //The below works - Start\n// val dfNewEvents2 = sparkSession.sql(\"select *,current_timestamp as newTimeStamp from dfNewEvents \").withWatermark(\"newTimeStamp\",\"2 seconds\")\n// dfNewEvents2.createOrReplaceTempView(\"dfNewEvents2\")\n// val groupEvents = sparkSession.sql(\"select symbol,newTimeStamp, count(price) as count1 from dfNewEvents2 group by symbol,newTimeStamp\")\n // End\n \n \n //The below doesn't work - Start\n val dfNewEvents2 = sparkSession.sql(\"select * from dfNewEvents \").withWatermark(\"eventTime\",\"2 seconds\")\n dfNewEvents2.createOrReplaceTempView(\"dfNewEvents2\")\n val groupEvents = sparkSession.sql(\"select symbol,eventTime, count(price) as count1 from dfNewEvents2 group by symbol,eventTime\")\n // - End\n \n \n val query1 = groupEvents.writeStream\n .outputMode(\"append\")\n .format(\"console\")\n .option(\"checkpointLocation\", checkpointPath)\n .start(\"./myop\")\n\n val newEventFile1=newEventsPath.resolve(\"eventNew1.json\")\n Files.write(newEventFile1, List(\n \"\"\"{\"symbol\": \"GOOG\",\"price\":100,\"eventTime\":\"2017-07-25T16:00:00.000-04:00\"}\"\"\",\n \"\"\"{\"symbol\": \"GOOG\",\"price\":200,\"eventTime\":\"2017-07-25T16:00:00.000-04:00\"}\"\"\"\n ).toIterable.asJava)\n query1.processAllAvailable()\n\n sparkSession.streams.awaitAnyTermination(10000)\n\n }\n\n private def newEvents(sparkSession: SparkSession) = {\n val newEvents = Paths.get(\"target/newEvents/\").toAbsolutePath\n delete(newEvents)\n Files.createDirectories(newEvents)\n\n val dfNewEvents = sparkSession.readStream.schema(eventsSchema).json(newEvents.toString)//.withWatermark(\"eventTime\",\"2 seconds\")\n dfNewEvents\n }\n\n private val eventsSchema = StructType(List(\n StructField(\"symbol\", StringType, true),\n StructField(\"price\", DoubleType, true),\n StructField(\"eventTime\", TimestampType, false)\n ))\n\n private def delete(dir: Path) = {\n if(Files.exists(dir)) {\n Files.walk(dir).iterator().asScala.toList\n .map(p => p.toFile)\n .sortWith((o1, o2) => o1.compareTo(o2) > 0)\n .foreach(_.delete)\n }\n }\n\n}\n\n\n{code}","from":"reporter","subject":"aggregate query fails with watermark on eventTime but works with watermark on timestamp column generated by current_timestamp"},{"body":"I believe you need to use a window to group by your event time.","from":"developer"},{"body":"I am not sure, does it mean we can not use watermark without windows - kind of on a microbatch level ? \n\nIf that is the case, it should not work with current_timestamp and should have validation not to allow without windows.","from":"developer"},{"body":"No nothing like the limitations of microbatches. The window can be made trivially small if you want only one timestamp per group, for example {{window(eventTime, \"1 microsecond\")}}\n\nAnd yes this should probably be checked in analysis if this is the intended limitation.","from":"developer"},{"body":"Thanks for reporting it. I can reproduce the error in a unit test. Still investigating it.","from":"developer"},{"body":"Resolved by https://github.com/apache/spark/pull/18840","from":"developer"}],"created":"2017-07-28T20:26:44.000+0000","description":"*Short Description: *\n\nAggregation query fails with eventTime as watermark column while works with newTimeStamp column generated by running SQL with current_timestamp,\n\n*Exception:*\n\n{code}\nCaused by: java.util.NoSuchElementException: None.get\n\tat scala.None$.get(Option.scala:347)\n\tat scala.None$.get(Option.scala:345)\n\tat org.apache.spark.sql.execution.streaming.StateStoreSaveExec$$anonfun$doExecute$3.apply(statefulOperators.scala:204)\n\tat org.apache.spark.sql.execution.streaming.StateStoreSaveExec$$anonfun$doExecute$3.apply(statefulOperators.scala:172)\n\tat org.apache.spark.sql.execution.streaming.state.package$StateStoreOps$$anonfun$1.apply(package.scala:70)\n\tat org.apache.spark.sql.execution.streaming.state.package$StateStoreOps$$anonfun$1.apply(package.scala:65)\n\tat org.apache.spark.sql.execution.streaming.state.StateStoreRDD.compute(StateStoreRDD.scala:64)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n{code}\n\n*Code to replicate:*\n\n{code}\npackage test\n\nimport java.nio.file.{Files, Path, Paths}\nimport java.text.SimpleDateFormat\n\nimport org.apache.spark.sql.types._\nimport org.apache.spark.sql.{SparkSession}\n\nimport scala.collection.JavaConverters._\n\nobject Test1 {\n\n def main(args: Array[String]) {\n\n val sparkSession = SparkSession\n .builder()\n .master(\"local[*]\")\n .appName(\"Spark SQL basic example\")\n .config(\"spark.some.config.option\", \"some-value\")\n .getOrCreate()\n\n val sdf = new SimpleDateFormat(\"yyyy-MM-dd HH:mm:ss\")\n val checkpointPath = \"target/cp1\"\n val newEventsPath = Paths.get(\"target/newEvents/\").toAbsolutePath\n delete(newEventsPath)\n delete(Paths.get(checkpointPath).toAbsolutePath)\n Files.createDirectories(newEventsPath)\n\n\n val dfNewEvents= newEvents(sparkSession)\n dfNewEvents.createOrReplaceTempView(\"dfNewEvents\")\n\n //The below works - Start\n// val dfNewEvents2 = sparkSession.sql(\"select *,current_timestamp as newTimeStamp from dfNewEvents \").withWatermark(\"newTimeStamp\",\"2 seconds\")\n// dfNewEvents2.createOrReplaceTempView(\"dfNewEvents2\")\n// val groupEvents = sparkSession.sql(\"select symbol,newTimeStamp, count(price) as count1 from dfNewEvents2 group by symbol,newTimeStamp\")\n // End\n \n \n //The below doesn't work - Start\n val dfNewEvents2 = sparkSession.sql(\"select * from dfNewEvents \").withWatermark(\"eventTime\",\"2 seconds\")\n dfNewEvents2.createOrReplaceTempView(\"dfNewEvents2\")\n val groupEvents = sparkSession.sql(\"select symbol,eventTime, count(price) as count1 from dfNewEvents2 group by symbol,eventTime\")\n // - End\n \n \n val query1 = groupEvents.writeStream\n .outputMode(\"append\")\n .format(\"console\")\n .option(\"checkpointLocation\", checkpointPath)\n .start(\"./myop\")\n\n val newEventFile1=newEventsPath.resolve(\"eventNew1.json\")\n Files.write(newEventFile1, List(\n \"\"\"{\"symbol\": \"GOOG\",\"price\":100,\"eventTime\":\"2017-07-25T16:00:00.000-04:00\"}\"\"\",\n \"\"\"{\"symbol\": \"GOOG\",\"price\":200,\"eventTime\":\"2017-07-25T16:00:00.000-04:00\"}\"\"\"\n ).toIterable.asJava)\n query1.processAllAvailable()\n\n sparkSession.streams.awaitAnyTermination(10000)\n\n }\n\n private def newEvents(sparkSession: SparkSession) = {\n val newEvents = Paths.get(\"target/newEvents/\").toAbsolutePath\n delete(newEvents)\n Files.createDirectories(newEvents)\n\n val dfNewEvents = sparkSession.readStream.schema(eventsSchema).json(newEvents.toString)//.withWatermark(\"eventTime\",\"2 seconds\")\n dfNewEvents\n }\n\n private val eventsSchema = StructType(List(\n StructField(\"symbol\", StringType, true),\n StructField(\"price\", DoubleType, true),\n StructField(\"eventTime\", TimestampType, false)\n ))\n\n private def delete(dir: Path) = {\n if(Files.exists(dir)) {\n Files.walk(dir).iterator().asScala.toList\n .map(p => p.toFile)\n .sortWith((o1, o2) => o1.compareTo(o2) > 0)\n .foreach(_.delete)\n }\n }\n\n}\n\n\n{code}","issue_id":"13090886","key":"SPARK-21565","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-08-07T20:02:33.000+0000","role":"fixed_distractor","summary":"aggregate query fails with watermark on eventTime but works with watermark on timestamp column generated by current_timestamp"} {"case_id":"13092970","cluster":"DISTRACTOR-SPARK-21657","comments":[{"body":"(Not a bug)\nI doubt this is meant to be efficient at the scale you're using it. Is this a real use case?\nWhat change are you proposing?","created":"2017-08-07T18:38:09.189+0000"},{"body":"Absolutely, this is a real use case. \nWe have a lot of production data that rely on that kind of schema for BI reporting. \nOther Hadoop sql engines, including Hive and Impala scale its time to explode nested collections linearly. \nSpark has exponential complexity to explode nested collection.\nThere is definitely a room for improvement, as after ~40k+ records in a nested collection, most time of the job\nis spent in exploding; after ~200k+ records in a nested collection, Spark is not usable.","created":"2017-08-07T18:46:38.077+0000"},{"body":"[~bjornjons] confirms this problem pertains to Spark 2.2 too.\n","created":"2017-08-10T15:31:12.382+0000"},{"body":"Maybe not very related to this issue. But I'm exploring Generate related code to get hint about this issue. I'm curious why we still don't enable codegen of Generate for now. [~hvanhovell] Maybe you know why it is disabled? Thanks.","created":"2017-08-11T08:54:49.041+0000"},{"body":"[~viirya] Probably, this is what you want? https://github.com/apache/spark/commit/b5c5bd98ea5e8dbfebcf86c5459bdf765f5ceb53","created":"2017-08-15T07:11:34.351+0000"},{"body":"Thank you [~maropu] and [~viirya], that commit is for Spark 2.2 so this problem might be worse in 2.2, but I don't think it's a root cause.\nAs we see the same exponential time complexity to explode a nested array in Spark 2.0 and 2.1.","created":"2017-08-15T16:49:28.758+0000"},{"body":"[~maropu] I've noticed that change. There is a hotfix trying to revert that: https://github.com/apache/spark/pull/17425. But in the end the hotfix doesn't revert it back.\n\nActually I've tried to enable codegen for GenerateExec and ran those tests without failure in local. So I'm wondering why we still disable it.","created":"2017-08-16T03:16:00.026+0000"},{"body":"Hi,\r\nWanted to add that we're facing exactly the same issue. 6 hours work for one row that contains 250k array (of struct of 4 strings).\r\nJust wanted to state that if we explode only the array, e.g, in your example:\r\ncached_df = sqlc.sql('select explode(amft) from ' + table_name)\r\n\r\nit finishes in about 3 mins. \r\nit happens in Spark 2.1 and also 2.2, eventhough SPARK-16998 was resolved.","created":"2017-10-26T13:51:44.306+0000"},{"body":"I suspect that something somewhere is doing something that's linear-time that looks like it should be constant-time, like referencing a linked list by index. See https://issues.apache.org/jira/browse/SPARK-22330 for a similar type of thing (though don't think it's the same issue as this one)","created":"2017-10-26T14:06:16.276+0000"},{"body":"Hi,\r\nJust ran a profiler for this code:\r\n{code:java}\r\nval BASE = 100000000\r\nval N = 100000\r\nval df = sc.parallelize(List((\"1234567890\", (BASE to (BASE+N)).map(x => (x.toString, (x+1).toString, (x+2).toString, (x+3).toString)).toList ))).toDF(\"c1\", \"c_arr\")\r\nval df_exploded = df.select(expr(\"c1\"), explode($\"c_arr\").as(\"c2\"))\r\ndf_exploded.write.mode(\"overwrite\").format(\"json\").save(\"/tmp/blah_explode\")\r\n{code}\r\n\r\nand it looks like [~srowen] is right, most of the time is spent in scala.collection.immutable.List.apply()\t (72.1%). inside:\r\norg.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext() (100%)\r\n\r\nI logged the generated code and found the problematic code:\r\n{code:java}\r\n if (serializefromobject_funcResult1 != null) {\r\n serializefromobject_value5 = (scala.collection.immutable.List) serializefromobject_funcResult1;\r\n } else {\r\n serializefromobject_isNull5 = true;\r\n }\r\n.\r\n.\r\n.\r\n while (serializefromobject_loopIndex < serializefromobject_dataLength) {\r\n MapObjects_loopValue0 = (scala.Tuple4) (serializefromobject_value5.apply(serializefromobject_loopIndex));\r\n{code}\r\n\r\nso that causes the quadratic time complexity.\r\nHowever, I'm not sure where is the code that generates this list instead of array for the exploded array.","created":"2017-10-27T12:52:55.497+0000"},{"body":"What if you call toArray in your code, and explode that? if it's just assuming the column type is constant-time to access at an index, then that would work around it. \r\nIdeally it would generate code that traversed the collection if it doesn't support fast RandomAccess, or implements like this otherwise (instance of IndexedSeq or something). But that might help narrow down the issue.","created":"2017-10-27T13:21:12.370+0000"},{"body":"I Switched to toArray instead of toList in the above code and I did get an improvement by factor of 2. but we still remain with the main bottleneck.\r\nnow the diff in the above example between:\r\n{code:java}\r\nval df_exploded = df.select(expr(\"c1\"), explode($\"c_arr\").as(\"c2\"))\r\n{code}\r\nand:\r\n{code:java}\r\nval df_exploded = df.select(explode($\"c_arr\").as(\"c2\"))\r\n{code}\r\nis 128 secs vs. 3 secs.\r\n\r\nAgain I profiled the former and saw that all the time got consumed in:\r\norg.apache.spark.unsafe.Platform.copyMemory()\t97.548096\t23,991 ms (97.5%)\t\r\n\r\nthe obvious diff between the execution plans is that the former has two WholeStageCodeGen plans and the later just one.\r\nI didn't exactly understood the generated code but I would guess that what happens is that in the problematic case the generated explode code is actually multiplying the long array to all the exploded rows and only filters it in the end.\r\nPlease see if you can verify it or think on a workaround for it.\r\n\r\n","created":"2017-10-29T09:14:19.686+0000"},{"body":"Can you paste the plans? this difference might be down to a different cause.\r\n\r\nThe linear-time-access List issue still look worth solving. [~hvanhovell] [~cloud_fan] are either of you familiar with how the explode code is generated? I also couldn't quite figure out what was generating access to a linked list (immutable.List) where a random-access collection looks more appropriate.","created":"2017-10-29T09:43:27.146+0000"},{"body":"Sure,\r\nthe plan for\r\n{code:java}\r\nval df_exploded = df.select(expr(\"c1\"), explode($\"c_arr\").as(\"c2\")).selectExpr(\"c1\" ,\"c2.*\")\r\n{code}\r\nis \r\n{noformat}\r\n== Parsed Logical Plan ==\r\n'Project [unresolvedalias('c1, None), ArrayBuffer(c2).*]\r\n+- Project [c1#6, c2#25]\r\n +- Generate explode(c_arr#7), true, false, [c2#25]\r\n +- Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Analyzed Logical Plan ==\r\nc1: string, _1: string, _2: string, _3: string, _4: string\r\nProject [c1#6, c2#25._1 AS _1#40, c2#25._2 AS _2#41, c2#25._3 AS _3#42, c2#25._4 AS _4#43]\r\n+- Project [c1#6, c2#25]\r\n +- Generate explode(c_arr#7), true, false, [c2#25]\r\n +- Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Optimized Logical Plan ==\r\nProject [c1#6, c2#25._1 AS _1#40, c2#25._2 AS _2#41, c2#25._3 AS _3#42, c2#25._4 AS _4#43]\r\n+- Generate explode(c_arr#7), true, false, [c2#25]\r\n +- Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(input[0, scala.Tuple2, true])._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(input[0, scala.Tuple2, true])._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Physical Plan ==\r\n*Project [c1#6, c2#25._1 AS _1#40, c2#25._2 AS _2#41, c2#25._3 AS _3#42, c2#25._4 AS _4#43]\r\n+- Generate explode(c_arr#7), true, false, [c2#25]\r\n +- *Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- *SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(input[0, scala.Tuple2, true])._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(input[0, scala.Tuple2, true])._2, None) AS _2#4]\r\n +- Scan ExternalRDDScan[obj#2]\r\n{noformat}\r\nand for:\r\n{code:java}\r\nval df_exploded = df.select(explode($\"c_arr\").as(\"c2\")).selectExpr(\"c2.*\")\r\n{code}\r\nis \r\n{noformat}\r\n== Parsed Logical Plan ==\r\n'Project [ArrayBuffer(c2).*]\r\n+- Project [c2#25]\r\n +- Generate explode(c_arr#7), false, false, [c2#25]\r\n +- Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Analyzed Logical Plan ==\r\n_1: string, _2: string, _3: string, _4: string\r\nProject [c2#25._1 AS _1#38, c2#25._2 AS _2#39, c2#25._3 AS _3#40, c2#25._4 AS _4#41]\r\n+- Project [c2#25]\r\n +- Generate explode(c_arr#7), false, false, [c2#25]\r\n +- Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Optimized Logical Plan ==\r\nProject [c2#25._1 AS _1#38, c2#25._2 AS _2#39, c2#25._3 AS _3#40, c2#25._4 AS _4#41]\r\n+- Generate explode(c_arr#7), false, false, [c2#25]\r\n +- Project [_2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(input[0, scala.Tuple2, true])._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(input[0, scala.Tuple2, true])._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Physical Plan ==\r\n*Project [c2#25._1 AS _1#38, c2#25._2 AS _2#39, c2#25._3 AS _3#40, c2#25._4 AS _4#41]\r\n+- Generate explode(c_arr#7), false, false, [c2#25]\r\n +- *Project [_2#4 AS c_arr#7]\r\n +- *SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(input[0, scala.Tuple2, true])._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(input[0, scala.Tuple2, true])._2, None) AS _2#4]\r\n +- Scan ExternalRDDScan[obj#2]\r\n{noformat}","created":"2017-10-29T10:46:45.230+0000"},{"body":"Thanks [~cloud_fan] for the fast look. You're saying that https://issues.apache.org/jira/browse/SPARK-22385 is a superset of this issue?","created":"2017-10-29T12:39:33.365+0000"},{"body":"I'd say they are different issues, and I haven't figured out the reason for this issue yet, and wanna fix that small issue first.","created":"2017-10-29T13:10:36.288+0000"},{"body":"After futher investigating I believe that my assesment is correct, the former case creates a generator with join=true while the later with join=false, as you can see in plans above (I also debugged). this causes the very long array of size 100k to be duplicated 100k times and afterwards get pruned because its column is not in the final projection. \r\nI'm not sure what's the best way to address this issue - ammend the generate operator according to the projection.\r\nin the meanwhile, in our case, I worked around that by manually adding the outer fields into each of structs of the array and then exploded only the array. it's an ugly solution but reduces our query time from 6 hours to about 2 mins.","created":"2017-10-30T04:18:07.323+0000"},{"body":"ok i found the relevant rule:\r\n{code:java|title=Optimizer.scala.java|borderStyle=solid}\r\n // Turn off `join` for Generate if no column from it's child is used\r\n case p @ Project(_, g: Generate)\r\n if g.join && !g.outer && p.references.subsetOf(g.generatedSet) =>\r\n p.copy(child = g.copy(join = false))\r\n{code}\r\nI'm not sure yet why it doesn't work.","created":"2017-10-30T04:36:50.559+0000"},{"body":"After some debugging, I think I understand the tricky part here.\r\nbecause there are outer fields in the query we set join=true for the Generate class, and because the Generator uses the array as child it can't be removed from the Generate output.\r\nI that because omitting the original column is so common it would make sense to add another attribute to the Generate class, like {{omitChild: Boolean}} and let the Optimizer turn it on with appropriate Rule.\r\nwhat do you think?","created":"2017-10-30T06:25:22.262+0000"},{"body":"User 'uzadude' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19683","created":"2017-11-07T11:54:04.440+0000"},{"body":"Hi,\r\nI created a pull request: https://github.com/apache/spark/pull/19683\r\nwould appreciate if you could take a look.","created":"2017-11-07T11:54:51.546+0000"},{"body":"Thank you [~uzadude] - great investigative work.\r\nWould be great if this patch can make it to the 2.3 release.\r\n","created":"2017-11-09T05:46:44.349+0000"},{"body":"Issue resolved by pull request 19683\n[https://github.com/apache/spark/pull/19683]","created":"2017-12-29T13:09:53.004+0000"},{"body":"Thank you everyone involved. It would be the most exciting fix in 2.3 release for us.","created":"2017-12-29T16:39:34.750+0000"}],"conversations":[{"body":"It can take up to half a day to explode a modest-sized nested collection (0.5m).\nOn a recent Xeon processors.\n\nSee attached pyspark script that reproduces this problem.\n\n{code}\ncached_df = sqlc.sql('select individ, hholdid, explode(amft) from ' + table_name).cache()\nprint sqlc.count()\n{code}\n\nThis script generate a number of tables, with the same total number of records across all nested collection (see `scaling` variable in loops). `scaling` variable scales up how many nested elements in each record, but by the same factor scales down number of records in the table. So total number of records stays the same.\n\nTime grows exponentially (notice log-10 vertical axis scale):\n!ExponentialTimeGrowth.PNG!\n\nAt scaling of 50,000 (see attached pyspark script), it took 7 hours to explode the nested collections (\\!) of 8k records.\n\nAfter 1000 elements in nested collection, time grows exponentially.\n","from":"reporter","subject":"Spark has exponential time complexity to explode(array of structs)"},{"body":"(Not a bug)\nI doubt this is meant to be efficient at the scale you're using it. Is this a real use case?\nWhat change are you proposing?","from":"developer"},{"body":"Absolutely, this is a real use case. \nWe have a lot of production data that rely on that kind of schema for BI reporting. \nOther Hadoop sql engines, including Hive and Impala scale its time to explode nested collections linearly. \nSpark has exponential complexity to explode nested collection.\nThere is definitely a room for improvement, as after ~40k+ records in a nested collection, most time of the job\nis spent in exploding; after ~200k+ records in a nested collection, Spark is not usable.","from":"developer"},{"body":"[~bjornjons] confirms this problem pertains to Spark 2.2 too.\n","from":"developer"},{"body":"Maybe not very related to this issue. But I'm exploring Generate related code to get hint about this issue. I'm curious why we still don't enable codegen of Generate for now. [~hvanhovell] Maybe you know why it is disabled? Thanks.","from":"developer"},{"body":"[~viirya] Probably, this is what you want? https://github.com/apache/spark/commit/b5c5bd98ea5e8dbfebcf86c5459bdf765f5ceb53","from":"developer"},{"body":"Thank you [~maropu] and [~viirya], that commit is for Spark 2.2 so this problem might be worse in 2.2, but I don't think it's a root cause.\nAs we see the same exponential time complexity to explode a nested array in Spark 2.0 and 2.1.","from":"developer"},{"body":"[~maropu] I've noticed that change. There is a hotfix trying to revert that: https://github.com/apache/spark/pull/17425. But in the end the hotfix doesn't revert it back.\n\nActually I've tried to enable codegen for GenerateExec and ran those tests without failure in local. So I'm wondering why we still disable it.","from":"developer"},{"body":"Hi,\r\nWanted to add that we're facing exactly the same issue. 6 hours work for one row that contains 250k array (of struct of 4 strings).\r\nJust wanted to state that if we explode only the array, e.g, in your example:\r\ncached_df = sqlc.sql('select explode(amft) from ' + table_name)\r\n\r\nit finishes in about 3 mins. \r\nit happens in Spark 2.1 and also 2.2, eventhough SPARK-16998 was resolved.","from":"developer"},{"body":"I suspect that something somewhere is doing something that's linear-time that looks like it should be constant-time, like referencing a linked list by index. See https://issues.apache.org/jira/browse/SPARK-22330 for a similar type of thing (though don't think it's the same issue as this one)","from":"developer"},{"body":"Hi,\r\nJust ran a profiler for this code:\r\n{code:java}\r\nval BASE = 100000000\r\nval N = 100000\r\nval df = sc.parallelize(List((\"1234567890\", (BASE to (BASE+N)).map(x => (x.toString, (x+1).toString, (x+2).toString, (x+3).toString)).toList ))).toDF(\"c1\", \"c_arr\")\r\nval df_exploded = df.select(expr(\"c1\"), explode($\"c_arr\").as(\"c2\"))\r\ndf_exploded.write.mode(\"overwrite\").format(\"json\").save(\"/tmp/blah_explode\")\r\n{code}\r\n\r\nand it looks like [~srowen] is right, most of the time is spent in scala.collection.immutable.List.apply()\t (72.1%). inside:\r\norg.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext() (100%)\r\n\r\nI logged the generated code and found the problematic code:\r\n{code:java}\r\n if (serializefromobject_funcResult1 != null) {\r\n serializefromobject_value5 = (scala.collection.immutable.List) serializefromobject_funcResult1;\r\n } else {\r\n serializefromobject_isNull5 = true;\r\n }\r\n.\r\n.\r\n.\r\n while (serializefromobject_loopIndex < serializefromobject_dataLength) {\r\n MapObjects_loopValue0 = (scala.Tuple4) (serializefromobject_value5.apply(serializefromobject_loopIndex));\r\n{code}\r\n\r\nso that causes the quadratic time complexity.\r\nHowever, I'm not sure where is the code that generates this list instead of array for the exploded array.","from":"developer"},{"body":"What if you call toArray in your code, and explode that? if it's just assuming the column type is constant-time to access at an index, then that would work around it. \r\nIdeally it would generate code that traversed the collection if it doesn't support fast RandomAccess, or implements like this otherwise (instance of IndexedSeq or something). But that might help narrow down the issue.","from":"developer"},{"body":"I Switched to toArray instead of toList in the above code and I did get an improvement by factor of 2. but we still remain with the main bottleneck.\r\nnow the diff in the above example between:\r\n{code:java}\r\nval df_exploded = df.select(expr(\"c1\"), explode($\"c_arr\").as(\"c2\"))\r\n{code}\r\nand:\r\n{code:java}\r\nval df_exploded = df.select(explode($\"c_arr\").as(\"c2\"))\r\n{code}\r\nis 128 secs vs. 3 secs.\r\n\r\nAgain I profiled the former and saw that all the time got consumed in:\r\norg.apache.spark.unsafe.Platform.copyMemory()\t97.548096\t23,991 ms (97.5%)\t\r\n\r\nthe obvious diff between the execution plans is that the former has two WholeStageCodeGen plans and the later just one.\r\nI didn't exactly understood the generated code but I would guess that what happens is that in the problematic case the generated explode code is actually multiplying the long array to all the exploded rows and only filters it in the end.\r\nPlease see if you can verify it or think on a workaround for it.\r\n\r\n","from":"developer"},{"body":"Can you paste the plans? this difference might be down to a different cause.\r\n\r\nThe linear-time-access List issue still look worth solving. [~hvanhovell] [~cloud_fan] are either of you familiar with how the explode code is generated? I also couldn't quite figure out what was generating access to a linked list (immutable.List) where a random-access collection looks more appropriate.","from":"developer"},{"body":"Sure,\r\nthe plan for\r\n{code:java}\r\nval df_exploded = df.select(expr(\"c1\"), explode($\"c_arr\").as(\"c2\")).selectExpr(\"c1\" ,\"c2.*\")\r\n{code}\r\nis \r\n{noformat}\r\n== Parsed Logical Plan ==\r\n'Project [unresolvedalias('c1, None), ArrayBuffer(c2).*]\r\n+- Project [c1#6, c2#25]\r\n +- Generate explode(c_arr#7), true, false, [c2#25]\r\n +- Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Analyzed Logical Plan ==\r\nc1: string, _1: string, _2: string, _3: string, _4: string\r\nProject [c1#6, c2#25._1 AS _1#40, c2#25._2 AS _2#41, c2#25._3 AS _3#42, c2#25._4 AS _4#43]\r\n+- Project [c1#6, c2#25]\r\n +- Generate explode(c_arr#7), true, false, [c2#25]\r\n +- Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Optimized Logical Plan ==\r\nProject [c1#6, c2#25._1 AS _1#40, c2#25._2 AS _2#41, c2#25._3 AS _3#42, c2#25._4 AS _4#43]\r\n+- Generate explode(c_arr#7), true, false, [c2#25]\r\n +- Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(input[0, scala.Tuple2, true])._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(input[0, scala.Tuple2, true])._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Physical Plan ==\r\n*Project [c1#6, c2#25._1 AS _1#40, c2#25._2 AS _2#41, c2#25._3 AS _3#42, c2#25._4 AS _4#43]\r\n+- Generate explode(c_arr#7), true, false, [c2#25]\r\n +- *Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- *SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(input[0, scala.Tuple2, true])._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(input[0, scala.Tuple2, true])._2, None) AS _2#4]\r\n +- Scan ExternalRDDScan[obj#2]\r\n{noformat}\r\nand for:\r\n{code:java}\r\nval df_exploded = df.select(explode($\"c_arr\").as(\"c2\")).selectExpr(\"c2.*\")\r\n{code}\r\nis \r\n{noformat}\r\n== Parsed Logical Plan ==\r\n'Project [ArrayBuffer(c2).*]\r\n+- Project [c2#25]\r\n +- Generate explode(c_arr#7), false, false, [c2#25]\r\n +- Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Analyzed Logical Plan ==\r\n_1: string, _2: string, _3: string, _4: string\r\nProject [c2#25._1 AS _1#38, c2#25._2 AS _2#39, c2#25._3 AS _3#40, c2#25._4 AS _4#41]\r\n+- Project [c2#25]\r\n +- Generate explode(c_arr#7), false, false, [c2#25]\r\n +- Project [_1#3 AS c1#6, _2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(assertnotnull(input[0, scala.Tuple2, true]))._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Optimized Logical Plan ==\r\nProject [c2#25._1 AS _1#38, c2#25._2 AS _2#39, c2#25._3 AS _3#40, c2#25._4 AS _4#41]\r\n+- Generate explode(c_arr#7), false, false, [c2#25]\r\n +- Project [_2#4 AS c_arr#7]\r\n +- SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(input[0, scala.Tuple2, true])._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(input[0, scala.Tuple2, true])._2, None) AS _2#4]\r\n +- ExternalRDD [obj#2]\r\n\r\n== Physical Plan ==\r\n*Project [c2#25._1 AS _1#38, c2#25._2 AS _2#39, c2#25._3 AS _3#40, c2#25._4 AS _4#41]\r\n+- Generate explode(c_arr#7), false, false, [c2#25]\r\n +- *Project [_2#4 AS c_arr#7]\r\n +- *SerializeFromObject [staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(input[0, scala.Tuple2, true])._1, true) AS _1#3, mapobjects(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), if (isnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))) null else named_struct(_1, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._1, true), _2, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._2, true), _3, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._3, true), _4, staticinvoke(class org.apache.spark.unsafe.types.UTF8String, StringType, fromString, assertnotnull(lambdavariable(MapObjects_loopValue0, MapObjects_loopIsNull0, ObjectType(class scala.Tuple4), true))._4, true)), assertnotnull(input[0, scala.Tuple2, true])._2, None) AS _2#4]\r\n +- Scan ExternalRDDScan[obj#2]\r\n{noformat}","from":"developer"},{"body":"Thanks [~cloud_fan] for the fast look. You're saying that https://issues.apache.org/jira/browse/SPARK-22385 is a superset of this issue?","from":"developer"},{"body":"I'd say they are different issues, and I haven't figured out the reason for this issue yet, and wanna fix that small issue first.","from":"developer"},{"body":"After futher investigating I believe that my assesment is correct, the former case creates a generator with join=true while the later with join=false, as you can see in plans above (I also debugged). this causes the very long array of size 100k to be duplicated 100k times and afterwards get pruned because its column is not in the final projection. \r\nI'm not sure what's the best way to address this issue - ammend the generate operator according to the projection.\r\nin the meanwhile, in our case, I worked around that by manually adding the outer fields into each of structs of the array and then exploded only the array. it's an ugly solution but reduces our query time from 6 hours to about 2 mins.","from":"developer"},{"body":"ok i found the relevant rule:\r\n{code:java|title=Optimizer.scala.java|borderStyle=solid}\r\n // Turn off `join` for Generate if no column from it's child is used\r\n case p @ Project(_, g: Generate)\r\n if g.join && !g.outer && p.references.subsetOf(g.generatedSet) =>\r\n p.copy(child = g.copy(join = false))\r\n{code}\r\nI'm not sure yet why it doesn't work.","from":"developer"},{"body":"After some debugging, I think I understand the tricky part here.\r\nbecause there are outer fields in the query we set join=true for the Generate class, and because the Generator uses the array as child it can't be removed from the Generate output.\r\nI that because omitting the original column is so common it would make sense to add another attribute to the Generate class, like {{omitChild: Boolean}} and let the Optimizer turn it on with appropriate Rule.\r\nwhat do you think?","from":"developer"},{"body":"User 'uzadude' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19683","from":"developer"},{"body":"Hi,\r\nI created a pull request: https://github.com/apache/spark/pull/19683\r\nwould appreciate if you could take a look.","from":"developer"},{"body":"Thank you [~uzadude] - great investigative work.\r\nWould be great if this patch can make it to the 2.3 release.\r\n","from":"developer"},{"body":"Issue resolved by pull request 19683\n[https://github.com/apache/spark/pull/19683]","from":"developer"},{"body":"Thank you everyone involved. It would be the most exciting fix in 2.3 release for us.","from":"developer"}],"created":"2017-08-07T18:21:33.000+0000","description":"It can take up to half a day to explode a modest-sized nested collection (0.5m).\nOn a recent Xeon processors.\n\nSee attached pyspark script that reproduces this problem.\n\n{code}\ncached_df = sqlc.sql('select individ, hholdid, explode(amft) from ' + table_name).cache()\nprint sqlc.count()\n{code}\n\nThis script generate a number of tables, with the same total number of records across all nested collection (see `scaling` variable in loops). `scaling` variable scales up how many nested elements in each record, but by the same factor scales down number of records in the table. So total number of records stays the same.\n\nTime grows exponentially (notice log-10 vertical axis scale):\n!ExponentialTimeGrowth.PNG!\n\nAt scaling of 50,000 (see attached pyspark script), it took 7 hours to explode the nested collections (\\!) of 8k records.\n\nAfter 1000 elements in nested collection, time grows exponentially.\n","issue_id":"13092970","key":"SPARK-21657","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-12-29T13:09:52.000+0000","role":"fixed_distractor","summary":"Spark has exponential time complexity to explode(array of structs)"} {"case_id":"13101894","cluster":"DISTRACTOR-SPARK-21991","comments":[{"body":"User 'nivox' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19217","created":"2017-09-13T09:17:03.656+0000"},{"body":"Thanks for debugging and diagnosing this [~nivox]! I'm seeing the same issue right now on one of my Spark clusters so am interested in getting your fix in to mainline Spark for my users.\r\n\r\nHave you deployed the change from your linked PR in a live setting, and has it fixed the issue for you?","created":"2017-10-24T08:37:57.420+0000"},{"body":"[~aash] I've tested it on a staging environment and it seemed to resolve the issue. However I didn't deploy it to the affected system since I didn't want to run a patched Spark distribution. For the time being I've setup a script to detect when the issue happen and restart the component. \r\nThis is obviously an ugly workaround and I would much prefer to have Spark handle it.","created":"2017-10-24T08:58:13.406+0000"},{"body":"Thanks for the contribution to Spark [~nivox]! I'll be testing this on some clusters of mine and will echo back here if I see any problems. Cheers!","created":"2017-10-25T17:29:16.168+0000"},{"body":"User 'ash211' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19574","created":"2017-10-25T18:04:04.751+0000"}],"conversations":[{"body":"The way the _LauncherServer_ _acceptConnections_ thread schedules client timeouts causes (non-deterministically) the thread to die with the following exception if the machine is under very high load:\n\n{noformat}\nException in thread \"LauncherServer-1\" java.lang.IllegalStateException: Task already scheduled or cancelled\n at java.util.Timer.sched(Timer.java:401)\n at java.util.Timer.schedule(Timer.java:193)\n at org.apache.spark.launcher.LauncherServer.acceptConnections(LauncherServer.java:249)\n at org.apache.spark.launcher.LauncherServer.access$000(LauncherServer.java:80)\n at org.apache.spark.launcher.LauncherServer$1.run(LauncherServer.java:143)\n{noformat}\n\nThe issue is related to the ordering of actions that the _acceptConnections_ thread uses to handle a client connection:\n\n# create timeout action\n# create client thread\n# start client thread\n# schedule timeout action\n\nUnder normal conditions the scheduling of the timeout action happen before the client thread has a chance to start, however if the machine is under very high load the client thread can receive CPU time before the timeout action gets scheduled.\n\nIf this condition happen, the client thread cancel the timeout action (which is not yet been scheduled) and goes on, but as soon as the _acceptConnections_ thread gets the CPU back, it will try to schedule the timeout action (which has already been canceled) thus raising the exception.\n\nChanging the order in which the client thread gets started and the timeout gets scheduled seems to be sufficient to fix this issue.\n\nAs stated above the issue is non-deterministic, I faced the issue multiple times on a single-node machine submitting a high number of short jobs sequentially, but I couldn't easily create a test reproducing the issue. ","from":"reporter","subject":"[LAUNCHER] LauncherServer acceptConnections thread sometime dies if machine has very high load"},{"body":"User 'nivox' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19217","from":"developer"},{"body":"Thanks for debugging and diagnosing this [~nivox]! I'm seeing the same issue right now on one of my Spark clusters so am interested in getting your fix in to mainline Spark for my users.\r\n\r\nHave you deployed the change from your linked PR in a live setting, and has it fixed the issue for you?","from":"developer"},{"body":"[~aash] I've tested it on a staging environment and it seemed to resolve the issue. However I didn't deploy it to the affected system since I didn't want to run a patched Spark distribution. For the time being I've setup a script to detect when the issue happen and restart the component. \r\nThis is obviously an ugly workaround and I would much prefer to have Spark handle it.","from":"developer"},{"body":"Thanks for the contribution to Spark [~nivox]! I'll be testing this on some clusters of mine and will echo back here if I see any problems. Cheers!","from":"developer"},{"body":"User 'ash211' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19574","from":"developer"}],"created":"2017-09-13T09:02:22.000+0000","description":"The way the _LauncherServer_ _acceptConnections_ thread schedules client timeouts causes (non-deterministically) the thread to die with the following exception if the machine is under very high load:\n\n{noformat}\nException in thread \"LauncherServer-1\" java.lang.IllegalStateException: Task already scheduled or cancelled\n at java.util.Timer.sched(Timer.java:401)\n at java.util.Timer.schedule(Timer.java:193)\n at org.apache.spark.launcher.LauncherServer.acceptConnections(LauncherServer.java:249)\n at org.apache.spark.launcher.LauncherServer.access$000(LauncherServer.java:80)\n at org.apache.spark.launcher.LauncherServer$1.run(LauncherServer.java:143)\n{noformat}\n\nThe issue is related to the ordering of actions that the _acceptConnections_ thread uses to handle a client connection:\n\n# create timeout action\n# create client thread\n# start client thread\n# schedule timeout action\n\nUnder normal conditions the scheduling of the timeout action happen before the client thread has a chance to start, however if the machine is under very high load the client thread can receive CPU time before the timeout action gets scheduled.\n\nIf this condition happen, the client thread cancel the timeout action (which is not yet been scheduled) and goes on, but as soon as the _acceptConnections_ thread gets the CPU back, it will try to schedule the timeout action (which has already been canceled) thus raising the exception.\n\nChanging the order in which the client thread gets started and the timeout gets scheduled seems to be sufficient to fix this issue.\n\nAs stated above the issue is non-deterministic, I faced the issue multiple times on a single-node machine submitting a high number of short jobs sequentially, but I couldn't easily create a test reproducing the issue. ","issue_id":"13101894","key":"SPARK-21991","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-10-25T17:12:21.000+0000","role":"fixed_distractor","summary":"[LAUNCHER] LauncherServer acceptConnections thread sometime dies if machine has very high load"} {"case_id":"13102124","cluster":"DISTRACTOR-SPARK-22000","comments":[{"body":"What's the query?","created":"2017-09-14T03:31:00.168+0000"},{"body":"It would be good to generate {{((Long)value13).toString()}} to reduce # of boxing/unboxing.\nAnyway, as @maropu pointed out, could you please put the query? Then, I will create a PR.","created":"2017-09-14T07:26:55.510+0000"},{"body":"That causes a boxing; better still is String.valueOf","created":"2017-09-14T07:34:21.753+0000"},{"body":"Thank you for good suggestion. I will try to use {{String.valueOf}}.","created":"2017-09-14T07:43:06.272+0000"},{"body":"@members, my code generates this error message but i cannot make sample code to reproduce this issue. ","created":"2017-09-15T02:32:10.406+0000"},{"body":"If there is no sample code, it may take a long time to fix this.\nIs it possible to attach all code or to put code to create all of Dataset or DataFrame?","created":"2017-09-18T17:10:09.601+0000"},{"body":"I've got similar issue. My sample code is attached.","created":"2019-02-20T14:33:49.525+0000"},{"body":"Thanks [~xsergey] I can easily reproduce it in master branch and have a fix. Will raise a PR shortly.\r\nCredit to [~srowen] given I'm leveraging String.valueOf(). Thanks!","created":"2019-02-21T09:15:37.957+0000"},{"body":"[~srowen] I am using spark.2.4.1 version getting same error ... for more details please check this , [https://stackoverflow.com/questions/58593215/inserting-into-cassandra-table-from-spark-dataframe-results-in-org-codehaus-comm]     how to fix this ?","created":"2019-10-28T15:25:56.234+0000"},{"body":"As indicated, it's fixed in 3.0, not 2.4.x. The change is substantial so is hard to back-port, but I'd review a change that includes the three linked PRs above in 2.4, if it can be made to work.","created":"2019-10-28T15:38:09.016+0000"}],"conversations":[{"body":"the error message say that toString is not declared on \"value13\" which is \"long\" type in generated code.\ni think value13 should be Long type.\n\n==error message\nCaused by: org.codehaus.commons.compiler.CompileException: File 'generated.java', Line 70, Column 32: failed to compile: org.codehaus.commons.compiler.CompileException: File 'generated.java', Line 70, Column 32: A method named \"toString\" is not declared in any enclosing class nor any supertype, nor through a static import\n\n\n/* 033 */ private void apply1_2(InternalRow i) {\n/* 034 */\n/* 035 */\n/* 036 */ boolean isNull11 = i.isNullAt(1);\n/* 037 */ UTF8String value11 = isNull11 ? null : (i.getUTF8String(1));\n/* 038 */ boolean isNull10 = true;\n/* 039 */ java.lang.String value10 = null;\n/* 040 */ if (!isNull11) {\n/* 041 */\n/* 042 */ isNull10 = false;\n/* 043 */ if (!isNull10) {\n/* 044 */\n/* 045 */ Object funcResult4 = null;\n/* 046 */ funcResult4 = value11.toString();\n/* 047 */\n/* 048 */ if (funcResult4 != null) {\n/* 049 */ value10 = (java.lang.String) funcResult4;\n/* 050 */ } else {\n/* 051 */ isNull10 = true;\n/* 052 */ }\n/* 053 */\n/* 054 */\n/* 055 */ }\n/* 056 */ }\n/* 057 */ javaBean.setApp(value10);\n/* 058 */\n/* 059 */\n/* 060 */ boolean isNull13 = i.isNullAt(12);\n/* 061 */ long value13 = isNull13 ? -1L : (i.getLong(12));\n/* 062 */ boolean isNull12 = true;\n/* 063 */ java.lang.String value12 = null;\n/* 064 */ if (!isNull13) {\n/* 065 */\n/* 066 */ isNull12 = false;\n/* 067 */ if (!isNull12) {\n/* 068 */\n/* 069 */ Object funcResult5 = null;\n/* 070 */ funcResult5 = value13.toString();\n/* 071 */\n/* 072 */ if (funcResult5 != null) {\n/* 073 */ value12 = (java.lang.String) funcResult5;\n/* 074 */ } else {\n/* 075 */ isNull12 = true;\n/* 076 */ }\n/* 077 */\n/* 078 */\n/* 079 */ }\n/* 080 */ }\n/* 081 */ javaBean.setReasonCode(value12);\n/* 082 */\n/* 083 */ }","from":"reporter","subject":"org.codehaus.commons.compiler.CompileException: toString method is not declared"},{"body":"What's the query?","from":"developer"},{"body":"It would be good to generate {{((Long)value13).toString()}} to reduce # of boxing/unboxing.\nAnyway, as @maropu pointed out, could you please put the query? Then, I will create a PR.","from":"developer"},{"body":"That causes a boxing; better still is String.valueOf","from":"developer"},{"body":"Thank you for good suggestion. I will try to use {{String.valueOf}}.","from":"developer"},{"body":"@members, my code generates this error message but i cannot make sample code to reproduce this issue. ","from":"developer"},{"body":"If there is no sample code, it may take a long time to fix this.\nIs it possible to attach all code or to put code to create all of Dataset or DataFrame?","from":"developer"},{"body":"I've got similar issue. My sample code is attached.","from":"developer"},{"body":"Thanks [~xsergey] I can easily reproduce it in master branch and have a fix. Will raise a PR shortly.\r\nCredit to [~srowen] given I'm leveraging String.valueOf(). Thanks!","from":"developer"},{"body":"[~srowen] I am using spark.2.4.1 version getting same error ... for more details please check this , [https://stackoverflow.com/questions/58593215/inserting-into-cassandra-table-from-spark-dataframe-results-in-org-codehaus-comm]     how to fix this ?","from":"developer"},{"body":"As indicated, it's fixed in 3.0, not 2.4.x. The change is substantial so is hard to back-port, but I'd review a change that includes the three linked PRs above in 2.4, if it can be made to work.","from":"developer"}],"created":"2017-09-14T02:29:43.000+0000","description":"the error message say that toString is not declared on \"value13\" which is \"long\" type in generated code.\ni think value13 should be Long type.\n\n==error message\nCaused by: org.codehaus.commons.compiler.CompileException: File 'generated.java', Line 70, Column 32: failed to compile: org.codehaus.commons.compiler.CompileException: File 'generated.java', Line 70, Column 32: A method named \"toString\" is not declared in any enclosing class nor any supertype, nor through a static import\n\n\n/* 033 */ private void apply1_2(InternalRow i) {\n/* 034 */\n/* 035 */\n/* 036 */ boolean isNull11 = i.isNullAt(1);\n/* 037 */ UTF8String value11 = isNull11 ? null : (i.getUTF8String(1));\n/* 038 */ boolean isNull10 = true;\n/* 039 */ java.lang.String value10 = null;\n/* 040 */ if (!isNull11) {\n/* 041 */\n/* 042 */ isNull10 = false;\n/* 043 */ if (!isNull10) {\n/* 044 */\n/* 045 */ Object funcResult4 = null;\n/* 046 */ funcResult4 = value11.toString();\n/* 047 */\n/* 048 */ if (funcResult4 != null) {\n/* 049 */ value10 = (java.lang.String) funcResult4;\n/* 050 */ } else {\n/* 051 */ isNull10 = true;\n/* 052 */ }\n/* 053 */\n/* 054 */\n/* 055 */ }\n/* 056 */ }\n/* 057 */ javaBean.setApp(value10);\n/* 058 */\n/* 059 */\n/* 060 */ boolean isNull13 = i.isNullAt(12);\n/* 061 */ long value13 = isNull13 ? -1L : (i.getLong(12));\n/* 062 */ boolean isNull12 = true;\n/* 063 */ java.lang.String value12 = null;\n/* 064 */ if (!isNull13) {\n/* 065 */\n/* 066 */ isNull12 = false;\n/* 067 */ if (!isNull12) {\n/* 068 */\n/* 069 */ Object funcResult5 = null;\n/* 070 */ funcResult5 = value13.toString();\n/* 071 */\n/* 072 */ if (funcResult5 != null) {\n/* 073 */ value12 = (java.lang.String) funcResult5;\n/* 074 */ } else {\n/* 075 */ isNull12 = true;\n/* 076 */ }\n/* 077 */\n/* 078 */\n/* 079 */ }\n/* 080 */ }\n/* 081 */ javaBean.setReasonCode(value12);\n/* 082 */\n/* 083 */ }","issue_id":"13102124","key":"SPARK-22000","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-02-27T05:53:12.000+0000","role":"fixed_distractor","summary":"org.codehaus.commons.compiler.CompileException: toString method is not declared"} {"case_id":"13104416","cluster":"DISTRACTOR-SPARK-22109","comments":[{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19331","created":"2017-09-23T10:32:03.607+0000"},{"body":"Issue resolved by pull request 19331\nhttps://github.com/apache/spark/pull/19331","created":"2017-09-23T15:12:20.520+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19333","created":"2017-09-23T15:32:03.877+0000"}],"conversations":[{"body":"If you try to read a partitioned json table, spark automatically tries to read figure out if the partition column is a timestamp based on the first value it sees. So if you really partitioned by a string, and the first value happens to look like a timestamp, then you'll run into errors. Even if you specify a schema, the schema is ignored, and spark still tries to infer a timestamp type for the partition column.\n\nThis is particularly weird because schema-inference does *not* work for regular timestamp columns in a flat table. You have to manually specify the schema to get the column interpreted as a timestamp.\n\nThis problem does not appear to be present for other types. Eg., if I partition by a string column, and the first value happens to look like an int, schema inference is still fine.\n\nHere's a small example:\n\n{noformat}\nval df = Seq(\n (1, \"2015-01-01 00:00:00\", Timestamp.valueOf(\"2015-01-01 00:00:00\")),\n (2, \"2014-01-01 00:00:00\", Timestamp.valueOf(\"2014-01-01 00:00:00\")),\n (3, \"blah\", Timestamp.valueOf(\"2016-01-01 00:00:00\"))).toDF(\"i\", \"str\", \"t\")\n\n\ndf.write.partitionBy(\"str\").json(\"partition_by_str\")\ndf.write.partitionBy(\"t\").json(\"partition_by_t\")\ndf.write.json(\"flat\")\n\nval readStr = spark.read.json(\"partition_by_str\")/*\njava.util.NoSuchElementException: None.get\n at scala.None$.get(Option.scala:347)\n at scala.None$.get(Option.scala:345)\n at org.apache.spark.sql.catalyst.expressions.TimeZoneAwareExpression$class.timeZone(datetimeExpressions.scala:46)\n at org.apache.spark.sql.catalyst.expressions.Cast.timeZone$lzycompute(Cast.scala:172)\n at org.apache.spark.sql.catalyst.expressions.Cast.timeZone(Cast.scala:172)\n at org.apache.spark.sql.catalyst.expressions.Cast$$anonfun$castToString$3$$anonfun$apply$16.apply(Cast.scala:208) at org.apache.spark.sql.catalyst.expressions.Cast$$anonfun$castToString$3$$anonfun$apply$16.apply(Cast.scala:208)\n at org.apache.spark.sql.catalyst.expressions.Cast.org$apache$spark$sql$catalyst$expressions$Cast$$buildCast(Cast.scala:201)\n at org.apache.spark.sql.catalyst.expressions.Cast$$anonfun$castToString$3.apply(Cast.scala:207)\n at org.apache.spark.sql.catalyst.expressions.Cast.nullSafeEval(Cast.scala:533)\n at org.apache.spark.sql.catalyst.expressions.UnaryExpression.eval(Expression.scala:327)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$$anonfun$org$apache$spark$sql$execution$datasources$PartitioningUtils$$resolveTypeConflicts$1.apply(PartitioningUtils.scala:485)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$$anonfun$org$apache$spark$sql$execution$datasources$PartitioningUtils$$resolveTypeConflicts$1.apply(PartitioningUtils.scala:484)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.AbstractTraversable.map(Traversable.scala:104)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$.org$apache$spark$sql$execution$datasources$PartitioningUtils$$resolveTypeConflicts(PartitioningUtils.scala:484)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$$anonfun$15.apply(PartitioningUtils.scala:340)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$$anonfun$15.apply(PartitioningUtils.scala:339)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.immutable.Range.foreach(Range.scala:160)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.AbstractTraversable.map(Traversable.scala:104)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$.resolvePartitions(PartitioningUtils.scala:339)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$.parsePartitions(PartitioningUtils.scala:141)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$.parsePartitions(PartitioningUtils.scala:97)\n at org.apache.spark.sql.execution.datasources.PartitioningAwareFileIndex.inferPartitioning(PartitioningAwareFileIndex.scala:153)\n at org.apache.spark.sql.execution.datasources.InMemoryFileIndex.partitionSpec(InMemoryFileIndex.scala:70)\n at org.apache.spark.sql.execution.datasources.PartitioningAwareFileIndex.partitionSchema(PartitioningAwareFileIndex.scala:50)\n at org.apache.spark.sql.execution.datasources.DataSource.getOrInferFileFormatSchema(DataSource.scala:133)\n at org.apache.spark.sql.execution.datasources.DataSource.resolveRelation(DataSource.scala:366)\n at org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:178)\n at org.apache.spark.sql.DataFrameReader.json(DataFrameReader.scala:333)\n at org.apache.spark.sql.DataFrameReader.json(DataFrameReader.scala:279)\n ... 48 elided\n*/\n\nval readStr = spark.read.schema(df.schema).json(\"partition_by_str\")\n/*\nsame exception\n*/\n\nval readT = spark.read.json(\"partition_by_t\") // OK\nval readT = spark.read.schema(df.schema).json(\"partition_by_t\") // OK\n\nval readFlat = spark.read.json(\"flat\") // NO error, by timestamp column is read a String\nval readFlat = spark.read.schema(df.schema).json(\"flat\") // OK\n{noformat}","from":"reporter","subject":"Reading tables partitioned by columns that look like timestamps has inconsistent schema inference"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19331","from":"developer"},{"body":"Issue resolved by pull request 19331\nhttps://github.com/apache/spark/pull/19331","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19333","from":"developer"}],"created":"2017-09-22T21:40:23.000+0000","description":"If you try to read a partitioned json table, spark automatically tries to read figure out if the partition column is a timestamp based on the first value it sees. So if you really partitioned by a string, and the first value happens to look like a timestamp, then you'll run into errors. Even if you specify a schema, the schema is ignored, and spark still tries to infer a timestamp type for the partition column.\n\nThis is particularly weird because schema-inference does *not* work for regular timestamp columns in a flat table. You have to manually specify the schema to get the column interpreted as a timestamp.\n\nThis problem does not appear to be present for other types. Eg., if I partition by a string column, and the first value happens to look like an int, schema inference is still fine.\n\nHere's a small example:\n\n{noformat}\nval df = Seq(\n (1, \"2015-01-01 00:00:00\", Timestamp.valueOf(\"2015-01-01 00:00:00\")),\n (2, \"2014-01-01 00:00:00\", Timestamp.valueOf(\"2014-01-01 00:00:00\")),\n (3, \"blah\", Timestamp.valueOf(\"2016-01-01 00:00:00\"))).toDF(\"i\", \"str\", \"t\")\n\n\ndf.write.partitionBy(\"str\").json(\"partition_by_str\")\ndf.write.partitionBy(\"t\").json(\"partition_by_t\")\ndf.write.json(\"flat\")\n\nval readStr = spark.read.json(\"partition_by_str\")/*\njava.util.NoSuchElementException: None.get\n at scala.None$.get(Option.scala:347)\n at scala.None$.get(Option.scala:345)\n at org.apache.spark.sql.catalyst.expressions.TimeZoneAwareExpression$class.timeZone(datetimeExpressions.scala:46)\n at org.apache.spark.sql.catalyst.expressions.Cast.timeZone$lzycompute(Cast.scala:172)\n at org.apache.spark.sql.catalyst.expressions.Cast.timeZone(Cast.scala:172)\n at org.apache.spark.sql.catalyst.expressions.Cast$$anonfun$castToString$3$$anonfun$apply$16.apply(Cast.scala:208) at org.apache.spark.sql.catalyst.expressions.Cast$$anonfun$castToString$3$$anonfun$apply$16.apply(Cast.scala:208)\n at org.apache.spark.sql.catalyst.expressions.Cast.org$apache$spark$sql$catalyst$expressions$Cast$$buildCast(Cast.scala:201)\n at org.apache.spark.sql.catalyst.expressions.Cast$$anonfun$castToString$3.apply(Cast.scala:207)\n at org.apache.spark.sql.catalyst.expressions.Cast.nullSafeEval(Cast.scala:533)\n at org.apache.spark.sql.catalyst.expressions.UnaryExpression.eval(Expression.scala:327)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$$anonfun$org$apache$spark$sql$execution$datasources$PartitioningUtils$$resolveTypeConflicts$1.apply(PartitioningUtils.scala:485)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$$anonfun$org$apache$spark$sql$execution$datasources$PartitioningUtils$$resolveTypeConflicts$1.apply(PartitioningUtils.scala:484)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.AbstractTraversable.map(Traversable.scala:104)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$.org$apache$spark$sql$execution$datasources$PartitioningUtils$$resolveTypeConflicts(PartitioningUtils.scala:484)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$$anonfun$15.apply(PartitioningUtils.scala:340)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$$anonfun$15.apply(PartitioningUtils.scala:339)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.immutable.Range.foreach(Range.scala:160)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.AbstractTraversable.map(Traversable.scala:104)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$.resolvePartitions(PartitioningUtils.scala:339)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$.parsePartitions(PartitioningUtils.scala:141)\n at org.apache.spark.sql.execution.datasources.PartitioningUtils$.parsePartitions(PartitioningUtils.scala:97)\n at org.apache.spark.sql.execution.datasources.PartitioningAwareFileIndex.inferPartitioning(PartitioningAwareFileIndex.scala:153)\n at org.apache.spark.sql.execution.datasources.InMemoryFileIndex.partitionSpec(InMemoryFileIndex.scala:70)\n at org.apache.spark.sql.execution.datasources.PartitioningAwareFileIndex.partitionSchema(PartitioningAwareFileIndex.scala:50)\n at org.apache.spark.sql.execution.datasources.DataSource.getOrInferFileFormatSchema(DataSource.scala:133)\n at org.apache.spark.sql.execution.datasources.DataSource.resolveRelation(DataSource.scala:366)\n at org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:178)\n at org.apache.spark.sql.DataFrameReader.json(DataFrameReader.scala:333)\n at org.apache.spark.sql.DataFrameReader.json(DataFrameReader.scala:279)\n ... 48 elided\n*/\n\nval readStr = spark.read.schema(df.schema).json(\"partition_by_str\")\n/*\nsame exception\n*/\n\nval readT = spark.read.json(\"partition_by_t\") // OK\nval readT = spark.read.schema(df.schema).json(\"partition_by_t\") // OK\n\nval readFlat = spark.read.json(\"flat\") // NO error, by timestamp column is read a String\nval readFlat = spark.read.schema(df.schema).json(\"flat\") // OK\n{noformat}","issue_id":"13104416","key":"SPARK-22109","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-09-23T15:12:59.000+0000","role":"fixed_distractor","summary":"Reading tables partitioned by columns that look like timestamps has inconsistent schema inference"} {"case_id":"13112560","cluster":"DISTRACTOR-SPARK-22371","comments":[{"body":"Could you please provide an easy way to reproduce the issue?","created":"2017-11-01T14:13:56.547+0000"},{"body":"Hi, Sorry for late reply.\r\n\r\nFrom our analysis this error seems to come when there are jobs running parallel on datasets and union of all those datasets and Full GC triggered at same time which clears the accumulate of children dataset of union.\r\n\r\nI am attaching a small program from which this error comes frequently.\r\n\r\n[^ShuffleIssue.java]\r\n[^Helper.scala]\r\n[^sampledata]\r\n\r\nSteps for generating sample data\r\n{code:title=UNIX COMMAND |borderStyle=solid}\r\nfor((i=1472428800;i<=1472947200;i=i+86400));do mkdir -p output/testdata/eventtime=$i;cp sampledata output/testdata/eventtime=$i/000001_0;cp sampledata output/testdata/eventtime=$i/000001_1;cp sampledata output/testdata/eventtime=$i/000001_2;cp sampledata output/testdata/eventtime=$i/000001_3;done\r\n{code}","created":"2017-12-15T07:13:16.203+0000"},{"body":"Hi, Any update on the above issue ","created":"2018-03-13T07:25:58.101+0000"},{"body":"We've seen this for the first time on 2.3.0. \r\n\r\nThe scenario that [~mayank.agarwal2305] described of jobs running in parallel on datasets and union of all datasets and full gc triggered\" sounds exactly like our scenario. We've been unable to upgrade because of this issue.","created":"2018-03-16T18:01:25.222+0000"},{"body":"We were hit by this issue on 2.3.0 too while running just \"show tables\".\r\n\r\nStack trace:\r\n\r\njava.lang.IllegalStateException: Attempted to access garbage collected accumulator 793\r\nat org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:265)\r\nat org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:261)\r\nat scala.Option.map(Option.scala:146)\r\nat org.apache.spark.util.AccumulatorContext$.get(AccumulatorV2.scala:261)\r\nat org.apache.spark.util.AccumulatorV2$$anonfun$name$1.apply(AccumulatorV2.scala:87)\r\nat org.apache.spark.util.AccumulatorV2$$anonfun$name$1.apply(AccumulatorV2.scala:87)\r\nat scala.Option.orElse(Option.scala:289)\r\nat org.apache.spark.util.AccumulatorV2.name(AccumulatorV2.scala:87)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3$$anonfun$run$1$$anonfun$apply$mcV$sp$1.apply(TaskResultGetter.scala:103)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3$$anonfun$run$1$$anonfun$apply$mcV$sp$1.apply(TaskResultGetter.scala:102)\r\nat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\r\nat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\r\nat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\nat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\r\nat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\r\nat scala.collection.AbstractTraversable.map(Traversable.scala:104)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3$$anonfun$run$1.apply$mcV$sp(TaskResultGetter.scala:102)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3$$anonfun$run$1.apply(TaskResultGetter.scala:63)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3$$anonfun$run$1.apply(TaskResultGetter.scala:63)\r\nat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1988)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3.run(TaskResultGetter.scala:62)\r\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\nat java.lang.Thread.run(Thread.java:748)","created":"2018-03-30T11:18:17.004+0000"},{"body":"Another data point -- I've seen this happen (in 2.3.0) during cleanup after a task failure:\r\n\r\n{code:none}\r\nCaused by: org.apache.spark.SparkException: Job aborted due to stage failure: Exception while getting task result: java.lang.IllegalStateException: Attempted to access garbage collected accumulator 365\r\n at org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1599)\r\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1587)\r\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1586)\r\n at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\n at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\r\n at org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1586)\r\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:831)\r\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:831)\r\n at scala.Option.foreach(Option.scala:257)\r\n at org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:831)\r\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:1820)\r\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1769)\r\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1758)\r\n at org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\r\n at org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:642)\r\n at org.apache.spark.SparkContext.runJob(SparkContext.scala:2027)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$.write(FileFormatWriter.scala:194)\r\n ... 31 more\r\n{code}","created":"2018-04-17T09:53:10.181+0000"},{"body":"Do we really need to throw an exception from AccumulatorContext.get() when an accumulator is garbage collected? There's a period of time when an accumulator has been garbage collected, but hasn't been removed from AccumulatorContext.originals by ContextCleaner. When an update is received for such accumulator it will throw an exception and kill the whole job. This can happen when a stage completed, but there're still running tasks from other attempts, speculation etc. Since AccumulatorContext.get() returns an Option we could just return None in such case. Before SPARK-20940 this method threw IllegalAccessError which is not a NonFatal, was caught at a lower level and didn't cause job failure.","created":"2018-04-20T16:36:07.469+0000"},{"body":"User 'artemrd' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21114","created":"2018-04-20T17:46:05.268+0000"},{"body":"Got the same problem with 2.3 and also the program stalled:\r\n\r\n{{ Uncaught exception in thread heartbeat-receiver-event-loop-thread}}\r\n{{java.lang.IllegalStateException: Attempted to access garbage collected accumulator 8825}}\r\n{{        at org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:265)}}\r\n{{        at org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:261)}}\r\n{{        at scala.Option.map(Option.scala:146)}}\r\n{{        at org.apache.spark.util.AccumulatorContext$.get(AccumulatorV2.scala:261)}}\r\n{{        at org.apache.spark.util.AccumulatorV2$$anonfun$name$1.apply(AccumulatorV2.scala:87)}}\r\n{{        at org.apache.spark.util.AccumulatorV2$$anonfun$name$1.apply(AccumulatorV2.scala:87)}}\r\n{{        at scala.Option.orElse(Option.scala:289)}}\r\n{{        at org.apache.spark.util.AccumulatorV2.name(AccumulatorV2.scala:87)}}\r\n{{        at org.apache.spark.util.AccumulatorV2.toInfo(AccumulatorV2.scala:108)}}","created":"2018-05-07T16:24:40.484+0000"},{"body":"Issue resolved by pull request 21114\n[https://github.com/apache/spark/pull/21114]","created":"2018-05-17T10:50:35.478+0000"},{"body":"Hello,\r\n\r\nis this fix included in version 2.3.2?\r\n\r\nFrom here, I would say yes: [https://github.com/apache/spark/blob/v2.3.2/core/src/main/scala/org/apache/spark/util/AccumulatorV2.scala]\r\n\r\nIn case it is, can you add 2.3.2 in the \"Fix version\" field of this Jira ticket?","created":"2018-10-04T11:46:01.134+0000"},{"body":"if it's fixed in 2.3.1, it goes without saying that it's fixed in 2.3.2 as well.","created":"2018-10-04T12:00:16.415+0000"}],"conversations":[{"body":"Our Spark Jobs are getting stuck on DagScheduler.runJob as dagscheduler thread is stopped because of *Attempted to access garbage collected accumulator 5605982*.\r\n\r\nfrom our investigation it look like accumulator is cleaned by GC first and same accumulator is used for merging the results from executor on task completion event.\r\n\r\n\r\nAs the error java.lang.IllegalAccessError is LinkageError which is treated as FatalError so dag-scheduler loop is finished with below exception.\r\n\r\n---ERROR stack trace --\r\nException in thread \"dag-scheduler-event-loop\" java.lang.IllegalAccessError: Attempted to access garbage collected accumulator 5605982\r\n\tat org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:253)\r\n\tat org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:249)\r\n\tat scala.Option.map(Option.scala:146)\r\n\tat org.apache.spark.util.AccumulatorContext$.get(AccumulatorV2.scala:249)\r\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$updateAccumulators$1.apply(DAGScheduler.scala:1083)\r\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$updateAccumulators$1.apply(DAGScheduler.scala:1080)\r\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\r\n\tat org.apache.spark.scheduler.DAGScheduler.updateAccumulators(DAGScheduler.scala:1080)\r\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskCompletion(DAGScheduler.scala:1183)\r\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:1647)\r\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1605)\r\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1594)\r\nat org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\r\n\r\nI am attaching the thread dump of driver as well ","from":"reporter","subject":"dag-scheduler-event-loop thread stopped with error Attempted to access garbage collected accumulator 5605982"},{"body":"Could you please provide an easy way to reproduce the issue?","from":"developer"},{"body":"Hi, Sorry for late reply.\r\n\r\nFrom our analysis this error seems to come when there are jobs running parallel on datasets and union of all those datasets and Full GC triggered at same time which clears the accumulate of children dataset of union.\r\n\r\nI am attaching a small program from which this error comes frequently.\r\n\r\n[^ShuffleIssue.java]\r\n[^Helper.scala]\r\n[^sampledata]\r\n\r\nSteps for generating sample data\r\n{code:title=UNIX COMMAND |borderStyle=solid}\r\nfor((i=1472428800;i<=1472947200;i=i+86400));do mkdir -p output/testdata/eventtime=$i;cp sampledata output/testdata/eventtime=$i/000001_0;cp sampledata output/testdata/eventtime=$i/000001_1;cp sampledata output/testdata/eventtime=$i/000001_2;cp sampledata output/testdata/eventtime=$i/000001_3;done\r\n{code}","from":"developer"},{"body":"Hi, Any update on the above issue ","from":"developer"},{"body":"We've seen this for the first time on 2.3.0. \r\n\r\nThe scenario that [~mayank.agarwal2305] described of jobs running in parallel on datasets and union of all datasets and full gc triggered\" sounds exactly like our scenario. We've been unable to upgrade because of this issue.","from":"developer"},{"body":"We were hit by this issue on 2.3.0 too while running just \"show tables\".\r\n\r\nStack trace:\r\n\r\njava.lang.IllegalStateException: Attempted to access garbage collected accumulator 793\r\nat org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:265)\r\nat org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:261)\r\nat scala.Option.map(Option.scala:146)\r\nat org.apache.spark.util.AccumulatorContext$.get(AccumulatorV2.scala:261)\r\nat org.apache.spark.util.AccumulatorV2$$anonfun$name$1.apply(AccumulatorV2.scala:87)\r\nat org.apache.spark.util.AccumulatorV2$$anonfun$name$1.apply(AccumulatorV2.scala:87)\r\nat scala.Option.orElse(Option.scala:289)\r\nat org.apache.spark.util.AccumulatorV2.name(AccumulatorV2.scala:87)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3$$anonfun$run$1$$anonfun$apply$mcV$sp$1.apply(TaskResultGetter.scala:103)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3$$anonfun$run$1$$anonfun$apply$mcV$sp$1.apply(TaskResultGetter.scala:102)\r\nat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\r\nat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\r\nat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\nat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\r\nat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\r\nat scala.collection.AbstractTraversable.map(Traversable.scala:104)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3$$anonfun$run$1.apply$mcV$sp(TaskResultGetter.scala:102)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3$$anonfun$run$1.apply(TaskResultGetter.scala:63)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3$$anonfun$run$1.apply(TaskResultGetter.scala:63)\r\nat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1988)\r\nat org.apache.spark.scheduler.TaskResultGetter$$anon$3.run(TaskResultGetter.scala:62)\r\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\nat java.lang.Thread.run(Thread.java:748)","from":"developer"},{"body":"Another data point -- I've seen this happen (in 2.3.0) during cleanup after a task failure:\r\n\r\n{code:none}\r\nCaused by: org.apache.spark.SparkException: Job aborted due to stage failure: Exception while getting task result: java.lang.IllegalStateException: Attempted to access garbage collected accumulator 365\r\n at org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1599)\r\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1587)\r\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1586)\r\n at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\n at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\r\n at org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1586)\r\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:831)\r\n at org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:831)\r\n at scala.Option.foreach(Option.scala:257)\r\n at org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:831)\r\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:1820)\r\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1769)\r\n at org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1758)\r\n at org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\r\n at org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:642)\r\n at org.apache.spark.SparkContext.runJob(SparkContext.scala:2027)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$.write(FileFormatWriter.scala:194)\r\n ... 31 more\r\n{code}","from":"developer"},{"body":"Do we really need to throw an exception from AccumulatorContext.get() when an accumulator is garbage collected? There's a period of time when an accumulator has been garbage collected, but hasn't been removed from AccumulatorContext.originals by ContextCleaner. When an update is received for such accumulator it will throw an exception and kill the whole job. This can happen when a stage completed, but there're still running tasks from other attempts, speculation etc. Since AccumulatorContext.get() returns an Option we could just return None in such case. Before SPARK-20940 this method threw IllegalAccessError which is not a NonFatal, was caught at a lower level and didn't cause job failure.","from":"developer"},{"body":"User 'artemrd' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21114","from":"developer"},{"body":"Got the same problem with 2.3 and also the program stalled:\r\n\r\n{{ Uncaught exception in thread heartbeat-receiver-event-loop-thread}}\r\n{{java.lang.IllegalStateException: Attempted to access garbage collected accumulator 8825}}\r\n{{        at org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:265)}}\r\n{{        at org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:261)}}\r\n{{        at scala.Option.map(Option.scala:146)}}\r\n{{        at org.apache.spark.util.AccumulatorContext$.get(AccumulatorV2.scala:261)}}\r\n{{        at org.apache.spark.util.AccumulatorV2$$anonfun$name$1.apply(AccumulatorV2.scala:87)}}\r\n{{        at org.apache.spark.util.AccumulatorV2$$anonfun$name$1.apply(AccumulatorV2.scala:87)}}\r\n{{        at scala.Option.orElse(Option.scala:289)}}\r\n{{        at org.apache.spark.util.AccumulatorV2.name(AccumulatorV2.scala:87)}}\r\n{{        at org.apache.spark.util.AccumulatorV2.toInfo(AccumulatorV2.scala:108)}}","from":"developer"},{"body":"Issue resolved by pull request 21114\n[https://github.com/apache/spark/pull/21114]","from":"developer"},{"body":"Hello,\r\n\r\nis this fix included in version 2.3.2?\r\n\r\nFrom here, I would say yes: [https://github.com/apache/spark/blob/v2.3.2/core/src/main/scala/org/apache/spark/util/AccumulatorV2.scala]\r\n\r\nIn case it is, can you add 2.3.2 in the \"Fix version\" field of this Jira ticket?","from":"developer"},{"body":"if it's fixed in 2.3.1, it goes without saying that it's fixed in 2.3.2 as well.","from":"developer"}],"created":"2017-10-27T10:35:42.000+0000","description":"Our Spark Jobs are getting stuck on DagScheduler.runJob as dagscheduler thread is stopped because of *Attempted to access garbage collected accumulator 5605982*.\r\n\r\nfrom our investigation it look like accumulator is cleaned by GC first and same accumulator is used for merging the results from executor on task completion event.\r\n\r\n\r\nAs the error java.lang.IllegalAccessError is LinkageError which is treated as FatalError so dag-scheduler loop is finished with below exception.\r\n\r\n---ERROR stack trace --\r\nException in thread \"dag-scheduler-event-loop\" java.lang.IllegalAccessError: Attempted to access garbage collected accumulator 5605982\r\n\tat org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:253)\r\n\tat org.apache.spark.util.AccumulatorContext$$anonfun$get$1.apply(AccumulatorV2.scala:249)\r\n\tat scala.Option.map(Option.scala:146)\r\n\tat org.apache.spark.util.AccumulatorContext$.get(AccumulatorV2.scala:249)\r\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$updateAccumulators$1.apply(DAGScheduler.scala:1083)\r\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$updateAccumulators$1.apply(DAGScheduler.scala:1080)\r\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\r\n\tat org.apache.spark.scheduler.DAGScheduler.updateAccumulators(DAGScheduler.scala:1080)\r\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskCompletion(DAGScheduler.scala:1183)\r\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:1647)\r\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1605)\r\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1594)\r\nat org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\r\n\r\nI am attaching the thread dump of driver as well ","issue_id":"13112560","key":"SPARK-22371","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-05-17T10:50:35.000+0000","role":"fixed_distractor","summary":"dag-scheduler-event-loop thread stopped with error Attempted to access garbage collected accumulator 5605982"} {"case_id":"13112714","cluster":"DISTRACTOR-SPARK-22373","comments":[{"body":"It looks like a Janino problem, so not sure if it can be fixed in Spark, or worked around if it can't be reproduced.","created":"2017-10-28T09:00:53.437+0000"},{"body":"If you can submit a program that can reproduce this, I could investigate which software component can cause an issue.","created":"2017-11-07T04:35:32.036+0000"},{"body":"Also facing this issue.\r\nFrom our experience, it is more likely to happen when handling data with very complicated schema.\r\nThe code to reproduce this issue can be as simple as reading the data followed by immediately writing it out.","created":"2017-11-27T22:40:36.516+0000"},{"body":"[~mshen] Thank you for reporting this issue. We would appreciate it if you can post the code to reproduce.","created":"2017-11-28T01:03:29.112+0000"},{"body":"The code that would throw this NPE looks like the following:\r\n\r\n{noformat}\r\nimport com.databricks.spark.avro._\r\n\r\nval df = spark.read.avro(\"/path/to/data/with/complicated/schema\")\r\ndf.write.mode(\"overwrite\").avro(\"/path/to/output\")\r\n{noformat}\r\n\r\nIt seems to me that the trigger for getting this NPE is the data instead of the code.\r\nNot sure how much this code would help.\r\n\r\nI do have the generated java code that causes Janino to throw this NPE\r\nI converted the [relevant code | https://github.com/apache/spark/blob/master/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/codegen/CodeGenerator.scala#L1215] in org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator into a standalone application, and running that application against the generated java code finished successfully.\r\nHowever, when using multiple threads to run this application, I started seeing this NPE.\r\nIt appears to me that this issue is indeed related to Janino's thread safety.\r\n\r\n[~kiszk], to further investigate, would it be helpful to provide that generated java code?","created":"2017-11-28T01:47:51.365+0000"},{"body":"Tried running the test application using 10 concurrent threads to compile the generated code I have.\r\nWith the current version of Janino 3.0.0 used by Spark, if I run the test application 100 times, I'm always able to see this NPE happening once or twice.\r\nSwitched to latest version of Janino 3.0.7 and tried multiple times, haven't seen this NPE happening yet.\r\nSeems that this might be fixed by bumping up Janino dependency version.","created":"2017-11-28T17:41:22.118+0000"},{"body":"This happens to me regularly enough using 2.1.1.20, Avro, and more than one executor core that I have abandoned use of multiple cores with Avro.\r\n\r\nI've attempted to make a 100% reproducible test case but failed, so I'm reporting these factors here. \r\n\r\n1. set --conf spark.executor.cores=2 (or any higher number)\r\n2. reading in a certain large Avro file\r\n3. spark 2.1.1.20\r\n4. spark.read.avro(fn).cache.count or other action involving writes; just count doesn't do it.\r\n5. The Avro file contains a key of type Map[String->Array[Byte]], though the values can all be empty arrays. The cardinality of they keyspace is high and the number of keys per map is tens to hundreds.\r\n6. Multiple partitions are necessary to trigger the error.\r\n7. Before the stack trace reported above, I see \"ERROR CodeGenerator: failed to compile: java.lang.NullPointerException\" followed by a dump of generated Java code. \r\n\r\nThe nodes that fail have \"INFO CodeGenerator: Code generated in ##.##### ms\" messages and my theory is that the code generator being used here has a thread safety issue. \r\n\r\n","created":"2017-11-28T19:50:16.122+0000"},{"body":"[~leigjklotz],\r\n\r\nI think bumping up Janino version to 3.0.7 definitely helps to resolve this issue.\r\nI have tried multiple times since yesterday.\r\nFor both the standalone application and my Spark application dealing with data that almost always generate this issue, I no longer see the NPE issue after bumping Janino to 3.0.7.\r\nLooking at Janino's release note, I haven't figured out which patch would fix this issue though.","created":"2017-11-29T01:54:46.034+0000"},{"body":"Attach the standalone testing application as well as the generated Java code that triggers this NPE so other people can also verify.\r\nThe application needs to be compiled pulling spark-catalyst, spark-core, and spark-sql as dependencies.\r\nEasiest way is to put it inside spark-catalyst module and compile spark-catalyst itself.\r\n\r\nWith the application compiled, you can launch it taking the dependency JARs as classpath.\r\nAgain, an easy way is to take Spark distribution's jars directory as classpath.\r\n\r\nLaunch the application like the following:\r\n{noformat}\r\nfor i in `seq 1 100`; do\r\n java -cp \"./*\" org.apache.spark.sql.catalyst.expressions.codegen.CodeGenTester 20 /path/to/generated.java;\r\ndone > output 2>&1 &\r\n{noformat}\r\n\r\nThis will run the application 100 times, each attempting to compile the java code using 20 concurrent threads.\r\nUsing Janino 3.0.0, I can always reproduce this issue.\r\nUsing Janino 3.0.7, this issue is gone.","created":"2017-11-29T03:05:56.820+0000"},{"body":"User 'Victsm' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19839","created":"2017-11-29T03:21:04.205+0000"},{"body":"Created PR https://github.com/apache/spark/pull/19839\r\n\r\n[~sowen] [~kiszk],\r\nCould you help to take a look?","created":"2017-11-29T03:21:38.617+0000"},{"body":"[~mshen] Thank you. I've hand-upgraded janino and commons-compiler to 3.0.7, and did no other dependencies. The NPE has not occurred, and I'm running further tests to make sure there are no other ill effects.\r\n","created":"2017-11-30T01:23:26.191+0000"},{"body":"Issue resolved by pull request 19839\n[https://github.com/apache/spark/pull/19839]","created":"2017-12-01T01:25:23.472+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19866","created":"2017-12-02T12:01:04.616+0000"}],"conversations":[{"body":"Very occasional and retry works.\r\n\r\nFull stack:\r\n17/10/27 21:06:15 ERROR Executor: Exception in task 29.0 in stage 12.0 (TID 758)\r\njava.lang.NullPointerException\r\n\tat org.codehaus.janino.IClass.isAssignableFrom(IClass.java:569)\r\n\tat org.codehaus.janino.UnitCompiler.isWideningReferenceConvertible(UnitCompiler.java:10347)\r\n\tat org.codehaus.janino.UnitCompiler.isMethodInvocationConvertible(UnitCompiler.java:8636)\r\n\tat org.codehaus.janino.UnitCompiler.findMostSpecificIInvocable(UnitCompiler.java:8427)\r\n\tat org.codehaus.janino.UnitCompiler.findMostSpecificIInvocable(UnitCompiler.java:8285)\r\n\tat org.codehaus.janino.UnitCompiler.findIMethod(UnitCompiler.java:8169)\r\n\tat org.codehaus.janino.UnitCompiler.findIMethod(UnitCompiler.java:8071)\r\n\tat org.codehaus.janino.UnitCompiler.compileGet2(UnitCompiler.java:4421)\r\n\tat org.codehaus.janino.UnitCompiler.access$7500(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$12.visitMethodInvocation(UnitCompiler.java:3774)\r\n\tat org.codehaus.janino.UnitCompiler$12.visitMethodInvocation(UnitCompiler.java:3762)\r\n\tat org.codehaus.janino.Java$MethodInvocation.accept(Java.java:4328)\r\n\tat org.codehaus.janino.UnitCompiler.compileGet(UnitCompiler.java:3762)\r\n\tat org.codehaus.janino.UnitCompiler.compileGetValue(UnitCompiler.java:4933)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:3180)\r\n\tat org.codehaus.janino.UnitCompiler.access$5000(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$9.visitMethodInvocation(UnitCompiler.java:3151)\r\n\tat org.codehaus.janino.UnitCompiler$9.visitMethodInvocation(UnitCompiler.java:3139)\r\n\tat org.codehaus.janino.Java$MethodInvocation.accept(Java.java:4328)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:3139)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:2112)\r\n\tat org.codehaus.janino.UnitCompiler.access$1700(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$6.visitExpressionStatement(UnitCompiler.java:1377)\r\n\tat org.codehaus.janino.UnitCompiler$6.visitExpressionStatement(UnitCompiler.java:1370)\r\n\tat org.codehaus.janino.Java$ExpressionStatement.accept(Java.java:2558)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:1370)\r\n\tat org.codehaus.janino.UnitCompiler.compileStatements(UnitCompiler.java:1450)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:2811)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:550)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:890)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:894)\r\n\tat org.codehaus.janino.UnitCompiler.access$600(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitMemberClassDeclaration(UnitCompiler.java:377)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitMemberClassDeclaration(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.Java$MemberClassDeclaration.accept(Java.java:1128)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.UnitCompiler.compileDeclaredMemberTypes(UnitCompiler.java:1209)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:564)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:890)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:894)\r\n\tat org.codehaus.janino.UnitCompiler.access$600(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitMemberClassDeclaration(UnitCompiler.java:377)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitMemberClassDeclaration(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.Java$MemberClassDeclaration.accept(Java.java:1128)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.UnitCompiler.compileDeclaredMemberTypes(UnitCompiler.java:1209)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:564)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:420)\r\n\tat org.codehaus.janino.UnitCompiler.access$400(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitPackageMemberClassDeclaration(UnitCompiler.java:374)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitPackageMemberClassDeclaration(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.Java$AbstractPackageMemberClassDeclaration.accept(Java.java:1309)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.UnitCompiler.compileUnit(UnitCompiler.java:345)\r\n\tat org.codehaus.janino.SimpleCompiler.compileToClassLoader(SimpleCompiler.java:396)\r\n\tat org.codehaus.janino.ClassBodyEvaluator.compileToClass(ClassBodyEvaluator.java:311)\r\n\tat org.codehaus.janino.ClassBodyEvaluator.cook(ClassBodyEvaluator.java:229)\r\n\tat org.codehaus.janino.SimpleCompiler.cook(SimpleCompiler.java:196)\r\n\tat org.codehaus.commons.compiler.Cookable.cook(Cookable.java:91)\r\n\tat org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator$.org$apache$spark$sql$catalyst$expressions$codegen$CodeGenerator$$doCompile(CodeGenerator.scala:959)\r\n\tat org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator$$anon$1.load(CodeGenerator.scala:1026)\r\n\tat org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator$$anon$1.load(CodeGenerator.scala:1023)\r\n\tat org.spark_project.guava.cache.LocalCache$LoadingValueReference.loadFuture(LocalCache.java:3599)\r\n\tat org.spark_project.guava.cache.LocalCache$Segment.loadSync(LocalCache.java:2379)\r\n\tat org.spark_project.guava.cache.LocalCache$Segment.lockedGetOrLoad(LocalCache.java:2342)\r\n\tat org.spark_project.guava.cache.LocalCache$Segment.get(LocalCache.java:2257)\r\n\tat org.spark_project.guava.cache.LocalCache.get(LocalCache.java:4000)\r\n\tat org.spark_project.guava.cache.LocalCache.getOrLoad(LocalCache.java:4004)\r\n\tat org.spark_project.guava.cache.LocalCache$LocalLoadingCache.get(LocalCache.java:4874)\r\n\tat org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator$.compile(CodeGenerator.scala:908)\r\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8.apply(WholeStageCodegenExec.scala:372)\r\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8.apply(WholeStageCodegenExec.scala:371)\r\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$26.apply(RDD.scala:844)\r\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$26.apply(RDD.scala:844)\r\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\r\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\r\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\r\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\r\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)\r\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)\r\n\tat org.apache.spark.scheduler.Task.run(Task.scala:99)\r\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:322)\r\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n\tat java.lang.Thread.run(Thread.java:745)\r\n17/10/27 21:06:15 INFO CodeGenerator: Code generated in 8.896831 ms\r\n\r\nIntermittent nature of problem makes me suspect the cache or a thread-related issue.\r\nSome the SQL that appears in the area of the code line reported in Spark UI:\r\n dense_rank() over (partition by itemid, type order by sum(col_a)+(sum(col_b)/1000000000000000.0) desc) as rank, \r\n ...where cast(mytimestampfield as String) >= '$mydate'\r\n \r\n","from":"reporter","subject":"Intermittent NullPointerException in org.codehaus.janino.IClass.isAssignableFrom"},{"body":"It looks like a Janino problem, so not sure if it can be fixed in Spark, or worked around if it can't be reproduced.","from":"developer"},{"body":"If you can submit a program that can reproduce this, I could investigate which software component can cause an issue.","from":"developer"},{"body":"Also facing this issue.\r\nFrom our experience, it is more likely to happen when handling data with very complicated schema.\r\nThe code to reproduce this issue can be as simple as reading the data followed by immediately writing it out.","from":"developer"},{"body":"[~mshen] Thank you for reporting this issue. We would appreciate it if you can post the code to reproduce.","from":"developer"},{"body":"The code that would throw this NPE looks like the following:\r\n\r\n{noformat}\r\nimport com.databricks.spark.avro._\r\n\r\nval df = spark.read.avro(\"/path/to/data/with/complicated/schema\")\r\ndf.write.mode(\"overwrite\").avro(\"/path/to/output\")\r\n{noformat}\r\n\r\nIt seems to me that the trigger for getting this NPE is the data instead of the code.\r\nNot sure how much this code would help.\r\n\r\nI do have the generated java code that causes Janino to throw this NPE\r\nI converted the [relevant code | https://github.com/apache/spark/blob/master/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/codegen/CodeGenerator.scala#L1215] in org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator into a standalone application, and running that application against the generated java code finished successfully.\r\nHowever, when using multiple threads to run this application, I started seeing this NPE.\r\nIt appears to me that this issue is indeed related to Janino's thread safety.\r\n\r\n[~kiszk], to further investigate, would it be helpful to provide that generated java code?","from":"developer"},{"body":"Tried running the test application using 10 concurrent threads to compile the generated code I have.\r\nWith the current version of Janino 3.0.0 used by Spark, if I run the test application 100 times, I'm always able to see this NPE happening once or twice.\r\nSwitched to latest version of Janino 3.0.7 and tried multiple times, haven't seen this NPE happening yet.\r\nSeems that this might be fixed by bumping up Janino dependency version.","from":"developer"},{"body":"This happens to me regularly enough using 2.1.1.20, Avro, and more than one executor core that I have abandoned use of multiple cores with Avro.\r\n\r\nI've attempted to make a 100% reproducible test case but failed, so I'm reporting these factors here. \r\n\r\n1. set --conf spark.executor.cores=2 (or any higher number)\r\n2. reading in a certain large Avro file\r\n3. spark 2.1.1.20\r\n4. spark.read.avro(fn).cache.count or other action involving writes; just count doesn't do it.\r\n5. The Avro file contains a key of type Map[String->Array[Byte]], though the values can all be empty arrays. The cardinality of they keyspace is high and the number of keys per map is tens to hundreds.\r\n6. Multiple partitions are necessary to trigger the error.\r\n7. Before the stack trace reported above, I see \"ERROR CodeGenerator: failed to compile: java.lang.NullPointerException\" followed by a dump of generated Java code. \r\n\r\nThe nodes that fail have \"INFO CodeGenerator: Code generated in ##.##### ms\" messages and my theory is that the code generator being used here has a thread safety issue. \r\n\r\n","from":"developer"},{"body":"[~leigjklotz],\r\n\r\nI think bumping up Janino version to 3.0.7 definitely helps to resolve this issue.\r\nI have tried multiple times since yesterday.\r\nFor both the standalone application and my Spark application dealing with data that almost always generate this issue, I no longer see the NPE issue after bumping Janino to 3.0.7.\r\nLooking at Janino's release note, I haven't figured out which patch would fix this issue though.","from":"developer"},{"body":"Attach the standalone testing application as well as the generated Java code that triggers this NPE so other people can also verify.\r\nThe application needs to be compiled pulling spark-catalyst, spark-core, and spark-sql as dependencies.\r\nEasiest way is to put it inside spark-catalyst module and compile spark-catalyst itself.\r\n\r\nWith the application compiled, you can launch it taking the dependency JARs as classpath.\r\nAgain, an easy way is to take Spark distribution's jars directory as classpath.\r\n\r\nLaunch the application like the following:\r\n{noformat}\r\nfor i in `seq 1 100`; do\r\n java -cp \"./*\" org.apache.spark.sql.catalyst.expressions.codegen.CodeGenTester 20 /path/to/generated.java;\r\ndone > output 2>&1 &\r\n{noformat}\r\n\r\nThis will run the application 100 times, each attempting to compile the java code using 20 concurrent threads.\r\nUsing Janino 3.0.0, I can always reproduce this issue.\r\nUsing Janino 3.0.7, this issue is gone.","from":"developer"},{"body":"User 'Victsm' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19839","from":"developer"},{"body":"Created PR https://github.com/apache/spark/pull/19839\r\n\r\n[~sowen] [~kiszk],\r\nCould you help to take a look?","from":"developer"},{"body":"[~mshen] Thank you. I've hand-upgraded janino and commons-compiler to 3.0.7, and did no other dependencies. The NPE has not occurred, and I'm running further tests to make sure there are no other ill effects.\r\n","from":"developer"},{"body":"Issue resolved by pull request 19839\n[https://github.com/apache/spark/pull/19839]","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19866","from":"developer"}],"created":"2017-10-27T22:17:56.000+0000","description":"Very occasional and retry works.\r\n\r\nFull stack:\r\n17/10/27 21:06:15 ERROR Executor: Exception in task 29.0 in stage 12.0 (TID 758)\r\njava.lang.NullPointerException\r\n\tat org.codehaus.janino.IClass.isAssignableFrom(IClass.java:569)\r\n\tat org.codehaus.janino.UnitCompiler.isWideningReferenceConvertible(UnitCompiler.java:10347)\r\n\tat org.codehaus.janino.UnitCompiler.isMethodInvocationConvertible(UnitCompiler.java:8636)\r\n\tat org.codehaus.janino.UnitCompiler.findMostSpecificIInvocable(UnitCompiler.java:8427)\r\n\tat org.codehaus.janino.UnitCompiler.findMostSpecificIInvocable(UnitCompiler.java:8285)\r\n\tat org.codehaus.janino.UnitCompiler.findIMethod(UnitCompiler.java:8169)\r\n\tat org.codehaus.janino.UnitCompiler.findIMethod(UnitCompiler.java:8071)\r\n\tat org.codehaus.janino.UnitCompiler.compileGet2(UnitCompiler.java:4421)\r\n\tat org.codehaus.janino.UnitCompiler.access$7500(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$12.visitMethodInvocation(UnitCompiler.java:3774)\r\n\tat org.codehaus.janino.UnitCompiler$12.visitMethodInvocation(UnitCompiler.java:3762)\r\n\tat org.codehaus.janino.Java$MethodInvocation.accept(Java.java:4328)\r\n\tat org.codehaus.janino.UnitCompiler.compileGet(UnitCompiler.java:3762)\r\n\tat org.codehaus.janino.UnitCompiler.compileGetValue(UnitCompiler.java:4933)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:3180)\r\n\tat org.codehaus.janino.UnitCompiler.access$5000(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$9.visitMethodInvocation(UnitCompiler.java:3151)\r\n\tat org.codehaus.janino.UnitCompiler$9.visitMethodInvocation(UnitCompiler.java:3139)\r\n\tat org.codehaus.janino.Java$MethodInvocation.accept(Java.java:4328)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:3139)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:2112)\r\n\tat org.codehaus.janino.UnitCompiler.access$1700(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$6.visitExpressionStatement(UnitCompiler.java:1377)\r\n\tat org.codehaus.janino.UnitCompiler$6.visitExpressionStatement(UnitCompiler.java:1370)\r\n\tat org.codehaus.janino.Java$ExpressionStatement.accept(Java.java:2558)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:1370)\r\n\tat org.codehaus.janino.UnitCompiler.compileStatements(UnitCompiler.java:1450)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:2811)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:550)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:890)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:894)\r\n\tat org.codehaus.janino.UnitCompiler.access$600(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitMemberClassDeclaration(UnitCompiler.java:377)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitMemberClassDeclaration(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.Java$MemberClassDeclaration.accept(Java.java:1128)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.UnitCompiler.compileDeclaredMemberTypes(UnitCompiler.java:1209)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:564)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:890)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:894)\r\n\tat org.codehaus.janino.UnitCompiler.access$600(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitMemberClassDeclaration(UnitCompiler.java:377)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitMemberClassDeclaration(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.Java$MemberClassDeclaration.accept(Java.java:1128)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.UnitCompiler.compileDeclaredMemberTypes(UnitCompiler.java:1209)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:564)\r\n\tat org.codehaus.janino.UnitCompiler.compile2(UnitCompiler.java:420)\r\n\tat org.codehaus.janino.UnitCompiler.access$400(UnitCompiler.java:206)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitPackageMemberClassDeclaration(UnitCompiler.java:374)\r\n\tat org.codehaus.janino.UnitCompiler$2.visitPackageMemberClassDeclaration(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.Java$AbstractPackageMemberClassDeclaration.accept(Java.java:1309)\r\n\tat org.codehaus.janino.UnitCompiler.compile(UnitCompiler.java:369)\r\n\tat org.codehaus.janino.UnitCompiler.compileUnit(UnitCompiler.java:345)\r\n\tat org.codehaus.janino.SimpleCompiler.compileToClassLoader(SimpleCompiler.java:396)\r\n\tat org.codehaus.janino.ClassBodyEvaluator.compileToClass(ClassBodyEvaluator.java:311)\r\n\tat org.codehaus.janino.ClassBodyEvaluator.cook(ClassBodyEvaluator.java:229)\r\n\tat org.codehaus.janino.SimpleCompiler.cook(SimpleCompiler.java:196)\r\n\tat org.codehaus.commons.compiler.Cookable.cook(Cookable.java:91)\r\n\tat org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator$.org$apache$spark$sql$catalyst$expressions$codegen$CodeGenerator$$doCompile(CodeGenerator.scala:959)\r\n\tat org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator$$anon$1.load(CodeGenerator.scala:1026)\r\n\tat org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator$$anon$1.load(CodeGenerator.scala:1023)\r\n\tat org.spark_project.guava.cache.LocalCache$LoadingValueReference.loadFuture(LocalCache.java:3599)\r\n\tat org.spark_project.guava.cache.LocalCache$Segment.loadSync(LocalCache.java:2379)\r\n\tat org.spark_project.guava.cache.LocalCache$Segment.lockedGetOrLoad(LocalCache.java:2342)\r\n\tat org.spark_project.guava.cache.LocalCache$Segment.get(LocalCache.java:2257)\r\n\tat org.spark_project.guava.cache.LocalCache.get(LocalCache.java:4000)\r\n\tat org.spark_project.guava.cache.LocalCache.getOrLoad(LocalCache.java:4004)\r\n\tat org.spark_project.guava.cache.LocalCache$LocalLoadingCache.get(LocalCache.java:4874)\r\n\tat org.apache.spark.sql.catalyst.expressions.codegen.CodeGenerator$.compile(CodeGenerator.scala:908)\r\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8.apply(WholeStageCodegenExec.scala:372)\r\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8.apply(WholeStageCodegenExec.scala:371)\r\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$26.apply(RDD.scala:844)\r\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndex$1$$anonfun$apply$26.apply(RDD.scala:844)\r\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\r\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\r\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\r\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\r\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)\r\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)\r\n\tat org.apache.spark.scheduler.Task.run(Task.scala:99)\r\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:322)\r\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n\tat java.lang.Thread.run(Thread.java:745)\r\n17/10/27 21:06:15 INFO CodeGenerator: Code generated in 8.896831 ms\r\n\r\nIntermittent nature of problem makes me suspect the cache or a thread-related issue.\r\nSome the SQL that appears in the area of the code line reported in Spark UI:\r\n dense_rank() over (partition by itemid, type order by sum(col_a)+(sum(col_b)/1000000000000000.0) desc) as rank, \r\n ...where cast(mytimestampfield as String) >= '$mydate'\r\n \r\n","issue_id":"13112714","key":"SPARK-22373","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-12-01T01:25:23.000+0000","role":"fixed_distractor","summary":"Intermittent NullPointerException in org.codehaus.janino.IClass.isAssignableFrom"} {"case_id":"13126584","cluster":"DISTRACTOR-SPARK-22860","comments":[{"body":"User 'tooptoop4' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21514","created":"2018-06-08T23:30:05.074+0000"},{"body":"gentle ping, fix waiting to be committed","created":"2019-02-11T22:48:14.004+0000"},{"body":"Not just in worker log, but in 'ps -ef' process list","created":"2019-02-17T11:47:21.897+0000"},{"body":"If we concern only about logging them into log file (boundary of this issue) we can try to remove them, but if we also concern about showing them into process list, that is a bit different issue.\r\n\r\nIf I'm not mistaken, we'll have to pass them to CoarseGrainedExecutorBackend at any way, because driver cannot pass these values which executor needs them to connect to driver. Adding level of security doesn't help, because we need to pass any security information to CoarseGrainedExecutorBackend to start from.","created":"2019-02-18T00:46:05.143+0000"},{"body":"User 'HeartSaVioR' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/23820","created":"2019-02-18T08:00:27.748+0000"},{"body":"User 'HeartSaVioR' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/23820","created":"2019-02-18T08:00:43.137+0000"},{"body":"I'm proposing to just redact these values from log message first, because complexities of both are quite different. Given we need to deal with cli arguments, I had to create my own logic instead of taking up existing PR. (Existing PR just removed them in arguments which I wonder it works well.) Sorry about that.","created":"2019-02-18T08:04:15.254+0000"},{"body":"[~kabhwan]  spark.ssl.keyStorePassword and  spark.ssl.keyPassword don't need to be passed to  CoarseGrainedExecutorBackend. Only  spark.ssl.trustStorePassword is used","created":"2019-02-22T21:38:00.094+0000"},{"body":"Issue resolved by pull request 23820\n[https://github.com/apache/spark/pull/23820]","created":"2019-02-26T22:50:13.728+0000"},{"body":"[~kabhwan] can this go in 2.4.5?","created":"2019-12-05T23:37:26.910+0000"}],"conversations":[{"body":"The workers log the spark.ssl.keyStorePassword and spark.ssl.trustStorePassword passed by cli to the executor processes. The ExecutorRunner should escape passwords to not appear in the worker's log files in INFO level. In this example, you can see my 'SuperSecretPassword' in a worker log:\r\n\r\n{code}\r\n17/12/08 08:04:12 INFO ExecutorRunner: Launch command: \"/global/myapp/oem/jdk/bin/java\" \"-cp\" \"/global/myapp/application/myapp_software/thing_loader_lib/core-repository-model-zzz-1.2.3-SNAPSHOT.jar\r\n[...]\r\n:/global/myapp/application/spark-2.1.1-bin-hadoop2.7/jars/*\" \"-Xmx16384M\" \"-Dspark.authenticate.enableSaslEncryption=true\" \"-Dspark.ssl.keyStorePassword=SuperSecretPassword\" \"-Dspark.ssl.keyStore=/global/myapp/application/config/ssl/keystore.jks\" \"-Dspark.ssl.trustStore=/global/myapp/application/config/ssl/truststore.jks\" \"-Dspark.ssl.enabled=true\" \"-Dspark.driver.port=39927\" \"-Dspark.ssl.protocol=TLS\" \"-Dspark.ssl.trustStorePassword=SuperSecretPassword\" \"-Dspark.authenticate=true\" \"-Dmyapp_IMPORT_DATE=2017-10-30\" \"-Dmyapp.config.directory=/global/myapp/application/config\" \"-Dsolr.httpclient.builder.factory=com.company.myapp.loader.auth.LoaderConfigSparkSolrBasicAuthConfigurer\" \"-Djavax.net.ssl.trustStore=/global/myapp/application/config/ssl/truststore.jks\" \"-XX:+UseG1GC\" \"-XX:+UseStringDeduplication\" \"-Dthings.loader.export.zzz_files=false\" \"-Dlog4j.configuration=file:/global/myapp/application/config/spark-executor-log4j.properties\" \"-XX:+HeapDumpOnOutOfMemoryError\" \"-XX:+UseStringDeduplication\" \"org.apache.spark.executor.CoarseGrainedExecutorBackend\" \"--driver-url\" \"spark://CoarseGrainedScheduler@192.168.0.1:39927\" \"--executor-id\" \"2\" \"--hostname\" \"192.168.0.1\" \"--cores\" \"4\" \"--app-id\" \"app-20171208080412-0000\" \"--worker-url\" \"spark://Worker@192.168.0.1:59530\"\r\n{code}","from":"reporter","subject":"Spark workers log ssl passwords passed to the executors"},{"body":"User 'tooptoop4' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21514","from":"developer"},{"body":"gentle ping, fix waiting to be committed","from":"developer"},{"body":"Not just in worker log, but in 'ps -ef' process list","from":"developer"},{"body":"If we concern only about logging them into log file (boundary of this issue) we can try to remove them, but if we also concern about showing them into process list, that is a bit different issue.\r\n\r\nIf I'm not mistaken, we'll have to pass them to CoarseGrainedExecutorBackend at any way, because driver cannot pass these values which executor needs them to connect to driver. Adding level of security doesn't help, because we need to pass any security information to CoarseGrainedExecutorBackend to start from.","from":"developer"},{"body":"User 'HeartSaVioR' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/23820","from":"developer"},{"body":"User 'HeartSaVioR' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/23820","from":"developer"},{"body":"I'm proposing to just redact these values from log message first, because complexities of both are quite different. Given we need to deal with cli arguments, I had to create my own logic instead of taking up existing PR. (Existing PR just removed them in arguments which I wonder it works well.) Sorry about that.","from":"developer"},{"body":"[~kabhwan]  spark.ssl.keyStorePassword and  spark.ssl.keyPassword don't need to be passed to  CoarseGrainedExecutorBackend. Only  spark.ssl.trustStorePassword is used","from":"developer"},{"body":"Issue resolved by pull request 23820\n[https://github.com/apache/spark/pull/23820]","from":"developer"},{"body":"[~kabhwan] can this go in 2.4.5?","from":"developer"}],"created":"2017-12-21T15:50:55.000+0000","description":"The workers log the spark.ssl.keyStorePassword and spark.ssl.trustStorePassword passed by cli to the executor processes. The ExecutorRunner should escape passwords to not appear in the worker's log files in INFO level. In this example, you can see my 'SuperSecretPassword' in a worker log:\r\n\r\n{code}\r\n17/12/08 08:04:12 INFO ExecutorRunner: Launch command: \"/global/myapp/oem/jdk/bin/java\" \"-cp\" \"/global/myapp/application/myapp_software/thing_loader_lib/core-repository-model-zzz-1.2.3-SNAPSHOT.jar\r\n[...]\r\n:/global/myapp/application/spark-2.1.1-bin-hadoop2.7/jars/*\" \"-Xmx16384M\" \"-Dspark.authenticate.enableSaslEncryption=true\" \"-Dspark.ssl.keyStorePassword=SuperSecretPassword\" \"-Dspark.ssl.keyStore=/global/myapp/application/config/ssl/keystore.jks\" \"-Dspark.ssl.trustStore=/global/myapp/application/config/ssl/truststore.jks\" \"-Dspark.ssl.enabled=true\" \"-Dspark.driver.port=39927\" \"-Dspark.ssl.protocol=TLS\" \"-Dspark.ssl.trustStorePassword=SuperSecretPassword\" \"-Dspark.authenticate=true\" \"-Dmyapp_IMPORT_DATE=2017-10-30\" \"-Dmyapp.config.directory=/global/myapp/application/config\" \"-Dsolr.httpclient.builder.factory=com.company.myapp.loader.auth.LoaderConfigSparkSolrBasicAuthConfigurer\" \"-Djavax.net.ssl.trustStore=/global/myapp/application/config/ssl/truststore.jks\" \"-XX:+UseG1GC\" \"-XX:+UseStringDeduplication\" \"-Dthings.loader.export.zzz_files=false\" \"-Dlog4j.configuration=file:/global/myapp/application/config/spark-executor-log4j.properties\" \"-XX:+HeapDumpOnOutOfMemoryError\" \"-XX:+UseStringDeduplication\" \"org.apache.spark.executor.CoarseGrainedExecutorBackend\" \"--driver-url\" \"spark://CoarseGrainedScheduler@192.168.0.1:39927\" \"--executor-id\" \"2\" \"--hostname\" \"192.168.0.1\" \"--cores\" \"4\" \"--app-id\" \"app-20171208080412-0000\" \"--worker-url\" \"spark://Worker@192.168.0.1:59530\"\r\n{code}","issue_id":"13126584","key":"SPARK-22860","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-02-26T22:50:13.000+0000","role":"fixed_distractor","summary":"Spark workers log ssl passwords passed to the executors"} {"case_id":"13128464","cluster":"DISTRACTOR-SPARK-22955","comments":[{"body":"[~tdas] should the order of these two stops be reversed? the comment on stopping the event loop suggests it needs to happen 'first', also.","created":"2018-01-04T14:31:59.827+0000"},{"body":"The issue is still exists, submitted PR in order to fix it.","created":"2019-08-20T10:40:02.590+0000"},{"body":"Resolved by https://github.com/apache/spark/pull/25511","created":"2019-09-03T18:02:28.805+0000"}],"conversations":[{"body":"when I stop a spark-streaming application with parameter \r\n spark.streaming.stopGracefullyOnShutdown, I get ERROR as follows:\r\n\r\n{code:java}\r\n2018-01-04 17:31:17,524 ERROR org.apache.spark.deploy.yarn.ApplicationMaster: RECEIVED SIGNAL TERM\r\n2018-01-04 17:31:17,527 INFO org.apache.spark.streaming.StreamingContext: Invoking stop(stopGracefully=true) from shutdown hook\r\n2018-01-04 17:31:17,530 INFO org.apache.spark.streaming.scheduler.ReceiverTracker: ReceiverTracker stopped\r\n2018-01-04 17:31:17,531 INFO org.apache.spark.streaming.scheduler.JobGenerator: Stopping JobGenerator gracefully\r\n2018-01-04 17:31:17,532 INFO org.apache.spark.streaming.scheduler.JobGenerator: Waiting for all received blocks to be consumed for job generation\r\n2018-01-04 17:31:17,533 INFO org.apache.spark.streaming.scheduler.JobGenerator: Waited for all received blocks to be consumed for job generation\r\n2018-01-04 17:31:17,747 INFO org.apache.spark.streaming.scheduler.JobScheduler: Added jobs for time 1515058267000 ms\r\n2018-01-04 17:31:18,302 INFO org.apache.spark.streaming.scheduler.JobScheduler: Added jobs for time 1515058268000 ms\r\n2018-01-04 17:31:18,785 INFO org.apache.spark.streaming.scheduler.JobScheduler: Added jobs for time 1515058269000 ms\r\n2018-01-04 17:31:19,001 INFO org.apache.spark.streaming.util.RecurringTimer: Stopped timer for JobGenerator after time 1515058279000\r\n2018-01-04 17:31:19,200 INFO org.apache.spark.streaming.scheduler.JobScheduler: Added jobs for time 1515058270000 ms\r\n2018-01-04 17:31:19,207 INFO org.apache.spark.streaming.scheduler.JobGenerator: Stopped generation timer\r\n2018-01-04 17:31:19,207 INFO org.apache.spark.streaming.scheduler.JobGenerator: Waiting for jobs to be processed and checkpoints to be written\r\n2018-01-04 17:31:19,210 ERROR org.apache.spark.streaming.scheduler.JobScheduler: Error generating jobs for time 1515058271000 ms\r\njava.lang.IllegalStateException: This consumer has already been closed.\r\n\tat org.apache.kafka.clients.consumer.KafkaConsumer.ensureNotClosed(KafkaConsumer.java:1417)\r\n\tat org.apache.kafka.clients.consumer.KafkaConsumer.acquire(KafkaConsumer.java:1428)\r\n\tat org.apache.kafka.clients.consumer.KafkaConsumer.poll(KafkaConsumer.java:929)\r\n\tat org.apache.spark.streaming.kafka010.DirectKafkaInputDStream.paranoidPoll(DirectKafkaInputDStream.scala:161)\r\n\tat org.apache.spark.streaming.kafka010.DirectKafkaInputDStream.latestOffsets(DirectKafkaInputDStream.scala:180)\r\n\tat org.apache.spark.streaming.kafka010.DirectKafkaInputDStream.compute(DirectKafkaInputDStream.scala:208)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.TransformedDStream$$anonfun$6.apply(TransformedDStream.scala:42)\r\n\tat org.apache.spark.streaming.dstream.TransformedDStream$$anonfun$6.apply(TransformedDStream.scala:42)\r\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\r\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\r\n\tat scala.collection.immutable.List.foreach(List.scala:381)\r\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\r\n\tat scala.collection.immutable.List.map(List.scala:285)\r\n\tat org.apache.spark.streaming.dstream.TransformedDStream.compute(TransformedDStream.scala:42)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.TransformedDStream.createRDDWithLocalProperties(TransformedDStream.scala:65)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.MappedDStream.compute(MappedDStream.scala:36)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.FlatMappedDStream.compute(FlatMappedDStream.scala:36)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.MappedDStream.compute(MappedDStream.scala:36)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.ShuffledDStream.compute(ShuffledDStream.scala:41)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.ForEachDStream.generateJob(ForEachDStream.scala:48)\r\n\tat org.apache.spark.streaming.DStreamGraph$$anonfun$1.apply(DStreamGraph.scala:122)\r\n\tat org.apache.spark.streaming.DStreamGraph$$anonfun$1.apply(DStreamGraph.scala:121)\r\n\tat scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\r\n\tat scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\r\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\r\n\tat scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\r\n\tat scala.collection.AbstractTraversable.flatMap(Traversable.scala:104)\r\n\tat org.apache.spark.streaming.DStreamGraph.generateJobs(DStreamGraph.scala:121)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator$$anonfun$3.apply(JobGenerator.scala:249)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator$$anonfun$3.apply(JobGenerator.scala:247)\r\n\tat scala.util.Try$.apply(Try.scala:192)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator.generateJobs(JobGenerator.scala:247)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator.org$apache$spark$streaming$scheduler$JobGenerator$$processEvent(JobGenerator.scala:183)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator$$anon$1.onReceive(JobGenerator.scala:89)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator$$anon$1.onReceive(JobGenerator.scala:88)\r\n\tat org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\r\n{code}\r\n\r\nIt looks like KafkaConsumer close before JobGenerator stop , then I view the source code, \r\n!https://raw.githubusercontent.com/smdfj/picture/master/spark/JobGenerator.png!\r\n\r\n\r\n\r\nI find graph.stop() will stop Dstream,then close KafkaConsumer,but JobGenerator.eventLoop has not stopped,so error occured.","from":"reporter","subject":"Error generating jobs when Stopping JobGenerator gracefully"},{"body":"[~tdas] should the order of these two stops be reversed? the comment on stopping the event loop suggests it needs to happen 'first', also.","from":"developer"},{"body":"The issue is still exists, submitted PR in order to fix it.","from":"developer"},{"body":"Resolved by https://github.com/apache/spark/pull/25511","from":"developer"}],"created":"2018-01-04T10:19:34.000+0000","description":"when I stop a spark-streaming application with parameter \r\n spark.streaming.stopGracefullyOnShutdown, I get ERROR as follows:\r\n\r\n{code:java}\r\n2018-01-04 17:31:17,524 ERROR org.apache.spark.deploy.yarn.ApplicationMaster: RECEIVED SIGNAL TERM\r\n2018-01-04 17:31:17,527 INFO org.apache.spark.streaming.StreamingContext: Invoking stop(stopGracefully=true) from shutdown hook\r\n2018-01-04 17:31:17,530 INFO org.apache.spark.streaming.scheduler.ReceiverTracker: ReceiverTracker stopped\r\n2018-01-04 17:31:17,531 INFO org.apache.spark.streaming.scheduler.JobGenerator: Stopping JobGenerator gracefully\r\n2018-01-04 17:31:17,532 INFO org.apache.spark.streaming.scheduler.JobGenerator: Waiting for all received blocks to be consumed for job generation\r\n2018-01-04 17:31:17,533 INFO org.apache.spark.streaming.scheduler.JobGenerator: Waited for all received blocks to be consumed for job generation\r\n2018-01-04 17:31:17,747 INFO org.apache.spark.streaming.scheduler.JobScheduler: Added jobs for time 1515058267000 ms\r\n2018-01-04 17:31:18,302 INFO org.apache.spark.streaming.scheduler.JobScheduler: Added jobs for time 1515058268000 ms\r\n2018-01-04 17:31:18,785 INFO org.apache.spark.streaming.scheduler.JobScheduler: Added jobs for time 1515058269000 ms\r\n2018-01-04 17:31:19,001 INFO org.apache.spark.streaming.util.RecurringTimer: Stopped timer for JobGenerator after time 1515058279000\r\n2018-01-04 17:31:19,200 INFO org.apache.spark.streaming.scheduler.JobScheduler: Added jobs for time 1515058270000 ms\r\n2018-01-04 17:31:19,207 INFO org.apache.spark.streaming.scheduler.JobGenerator: Stopped generation timer\r\n2018-01-04 17:31:19,207 INFO org.apache.spark.streaming.scheduler.JobGenerator: Waiting for jobs to be processed and checkpoints to be written\r\n2018-01-04 17:31:19,210 ERROR org.apache.spark.streaming.scheduler.JobScheduler: Error generating jobs for time 1515058271000 ms\r\njava.lang.IllegalStateException: This consumer has already been closed.\r\n\tat org.apache.kafka.clients.consumer.KafkaConsumer.ensureNotClosed(KafkaConsumer.java:1417)\r\n\tat org.apache.kafka.clients.consumer.KafkaConsumer.acquire(KafkaConsumer.java:1428)\r\n\tat org.apache.kafka.clients.consumer.KafkaConsumer.poll(KafkaConsumer.java:929)\r\n\tat org.apache.spark.streaming.kafka010.DirectKafkaInputDStream.paranoidPoll(DirectKafkaInputDStream.scala:161)\r\n\tat org.apache.spark.streaming.kafka010.DirectKafkaInputDStream.latestOffsets(DirectKafkaInputDStream.scala:180)\r\n\tat org.apache.spark.streaming.kafka010.DirectKafkaInputDStream.compute(DirectKafkaInputDStream.scala:208)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.TransformedDStream$$anonfun$6.apply(TransformedDStream.scala:42)\r\n\tat org.apache.spark.streaming.dstream.TransformedDStream$$anonfun$6.apply(TransformedDStream.scala:42)\r\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\r\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\r\n\tat scala.collection.immutable.List.foreach(List.scala:381)\r\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\r\n\tat scala.collection.immutable.List.map(List.scala:285)\r\n\tat org.apache.spark.streaming.dstream.TransformedDStream.compute(TransformedDStream.scala:42)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.TransformedDStream.createRDDWithLocalProperties(TransformedDStream.scala:65)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.MappedDStream.compute(MappedDStream.scala:36)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.FlatMappedDStream.compute(FlatMappedDStream.scala:36)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.MappedDStream.compute(MappedDStream.scala:36)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.ShuffledDStream.compute(ShuffledDStream.scala:41)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1$$anonfun$apply$7.apply(DStream.scala:342)\r\n\tat scala.util.DynamicVariable.withValue(DynamicVariable.scala:58)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1$$anonfun$1.apply(DStream.scala:341)\r\n\tat org.apache.spark.streaming.dstream.DStream.createRDDWithLocalProperties(DStream.scala:416)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:336)\r\n\tat org.apache.spark.streaming.dstream.DStream$$anonfun$getOrCompute$1.apply(DStream.scala:334)\r\n\tat scala.Option.orElse(Option.scala:289)\r\n\tat org.apache.spark.streaming.dstream.DStream.getOrCompute(DStream.scala:331)\r\n\tat org.apache.spark.streaming.dstream.ForEachDStream.generateJob(ForEachDStream.scala:48)\r\n\tat org.apache.spark.streaming.DStreamGraph$$anonfun$1.apply(DStreamGraph.scala:122)\r\n\tat org.apache.spark.streaming.DStreamGraph$$anonfun$1.apply(DStreamGraph.scala:121)\r\n\tat scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\r\n\tat scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\r\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\r\n\tat scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\r\n\tat scala.collection.AbstractTraversable.flatMap(Traversable.scala:104)\r\n\tat org.apache.spark.streaming.DStreamGraph.generateJobs(DStreamGraph.scala:121)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator$$anonfun$3.apply(JobGenerator.scala:249)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator$$anonfun$3.apply(JobGenerator.scala:247)\r\n\tat scala.util.Try$.apply(Try.scala:192)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator.generateJobs(JobGenerator.scala:247)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator.org$apache$spark$streaming$scheduler$JobGenerator$$processEvent(JobGenerator.scala:183)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator$$anon$1.onReceive(JobGenerator.scala:89)\r\n\tat org.apache.spark.streaming.scheduler.JobGenerator$$anon$1.onReceive(JobGenerator.scala:88)\r\n\tat org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\r\n{code}\r\n\r\nIt looks like KafkaConsumer close before JobGenerator stop , then I view the source code, \r\n!https://raw.githubusercontent.com/smdfj/picture/master/spark/JobGenerator.png!\r\n\r\n\r\n\r\nI find graph.stop() will stop Dstream,then close KafkaConsumer,but JobGenerator.eventLoop has not stopped,so error occured.","issue_id":"13128464","key":"SPARK-22955","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-09-03T18:02:28.000+0000","role":"fixed_distractor","summary":"Error generating jobs when Stopping JobGenerator gracefully"} {"case_id":"13129723","cluster":"DISTRACTOR-SPARK-23020","comments":[{"body":"cc [~vanzin]","created":"2018-01-10T01:06:44.685+0000"},{"body":"I think I found the race in the code, now need to figure out how to fix it... :-/","created":"2018-01-10T18:12:07.167+0000"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20223","created":"2018-01-10T20:56:04.681+0000"},{"body":"Issue resolved by pull request 20223\n[https://github.com/apache/spark/pull/20223]","created":"2018-01-16T06:42:19.727+0000"},{"body":"I had to revert this patch as it broke [{{YarnClusterSuite.timeout to get SparkContext in cluster mode triggers failure}}| https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.7/94/testReport/org.apache.spark.deploy.yarn/YarnClusterSuite/timeout_to_get_SparkContext_in_cluster_mode_triggers_failure/history/] \r\n\r\n[{{SparkLauncherSuite.testInProcessLauncher}}|https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.7/90/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/] seems to be still flaky.","created":"2018-01-17T06:24:15.773+0000"},{"body":"User 'sameeragarwal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20291","created":"2018-01-17T08:20:05.827+0000"},{"body":"Bummer. I'll try to take another look later today.","created":"2018-01-17T17:30:07.871+0000"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20297","created":"2018-01-17T19:45:05.754+0000"},{"body":"Issue resolved by pull request 20297\n[https://github.com/apache/spark/pull/20297]","created":"2018-01-22T06:51:10.364+0000"},{"body":"User 'ueshin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20376","created":"2018-01-24T04:15:03.336+0000"},{"body":"FYI The {{SparkLauncherSuite}} test is still failing occasionally (a lot less common though): https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.6/142/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/","created":"2018-01-24T18:17:17.657+0000"},{"body":"Argh. Feel free to disable it in branch-2.3; please leave it on on master so we can get more info while I look at it.","created":"2018-01-24T18:21:31.440+0000"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20388","created":"2018-01-24T22:15:04.385+0000"},{"body":"I'm sorry but the flakiness in the test still refuses to go away: [https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.7/154/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/.|https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.7/154/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/]\r\n\r\n \r\n\r\nPer Marcelo's suggestion, I'm going to (only) disable this test in 2.3. The master builds are failing similarly ([https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-master-test-maven-hadoop-2.7/4426/testReport/junit/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/)] so I hope it'll not hinder any investigation.","created":"2018-01-29T07:05:20.877+0000"},{"body":":-/\r\n\r\nIt's getting harder and harder to reproduce these races locally... this one may take a while.","created":"2018-01-29T17:41:15.191+0000"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20462","created":"2018-01-31T22:09:03.779+0000"},{"body":"Issue resolved by pull request 20462\n[https://github.com/apache/spark/pull/20462]","created":"2018-02-02T04:02:18.031+0000"},{"body":"Things look pretty stable on master, so I'll post a backport for 2.3.1 so we get the fixes in the next maintenance release.\r\nhttps://amplab.cs.berkeley.edu/jenkins/user/vanzin/my-views/view/Spark/job/spark-master-test-maven-hadoop-2.7/4571/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/","created":"2018-03-06T00:27:16.686+0000"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20743","created":"2018-03-06T00:34:04.484+0000"},{"body":"Hi, All.\r\n\r\nThis seems to fail again in branch 2.3. Can we disable this in branch-2.3 for Apache Spark 2.3.1 at least?\r\n- https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.6/lastCompletedBuild/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/","created":"2018-05-04T16:22:14.116+0000"},{"body":"If that still fails somewhere it means there is still a bug somewhere. I don't think disabling the test is the right thing unless it's actually common enough that it's causing problems. Lots of our tests are flaky.\r\n\r\nThat's like the only failure recently, BTW.\r\nhttps://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.6/lastCompletedBuild/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/","created":"2018-05-04T16:54:57.367+0000"}],"conversations":[{"body":"https://amplab.cs.berkeley.edu/jenkins/job/spark-branch-2.3-test-maven-hadoop-2.7/42/testReport/junit/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/","from":"reporter","subject":"Re-enable Flaky Test: org.apache.spark.launcher.SparkLauncherSuite.testInProcessLauncher"},{"body":"cc [~vanzin]","from":"developer"},{"body":"I think I found the race in the code, now need to figure out how to fix it... :-/","from":"developer"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20223","from":"developer"},{"body":"Issue resolved by pull request 20223\n[https://github.com/apache/spark/pull/20223]","from":"developer"},{"body":"I had to revert this patch as it broke [{{YarnClusterSuite.timeout to get SparkContext in cluster mode triggers failure}}| https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.7/94/testReport/org.apache.spark.deploy.yarn/YarnClusterSuite/timeout_to_get_SparkContext_in_cluster_mode_triggers_failure/history/] \r\n\r\n[{{SparkLauncherSuite.testInProcessLauncher}}|https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.7/90/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/] seems to be still flaky.","from":"developer"},{"body":"User 'sameeragarwal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20291","from":"developer"},{"body":"Bummer. I'll try to take another look later today.","from":"developer"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20297","from":"developer"},{"body":"Issue resolved by pull request 20297\n[https://github.com/apache/spark/pull/20297]","from":"developer"},{"body":"User 'ueshin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20376","from":"developer"},{"body":"FYI The {{SparkLauncherSuite}} test is still failing occasionally (a lot less common though): https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.6/142/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/","from":"developer"},{"body":"Argh. Feel free to disable it in branch-2.3; please leave it on on master so we can get more info while I look at it.","from":"developer"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20388","from":"developer"},{"body":"I'm sorry but the flakiness in the test still refuses to go away: [https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.7/154/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/.|https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.7/154/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/]\r\n\r\n \r\n\r\nPer Marcelo's suggestion, I'm going to (only) disable this test in 2.3. The master builds are failing similarly ([https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-master-test-maven-hadoop-2.7/4426/testReport/junit/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/)] so I hope it'll not hinder any investigation.","from":"developer"},{"body":":-/\r\n\r\nIt's getting harder and harder to reproduce these races locally... this one may take a while.","from":"developer"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20462","from":"developer"},{"body":"Issue resolved by pull request 20462\n[https://github.com/apache/spark/pull/20462]","from":"developer"},{"body":"Things look pretty stable on master, so I'll post a backport for 2.3.1 so we get the fixes in the next maintenance release.\r\nhttps://amplab.cs.berkeley.edu/jenkins/user/vanzin/my-views/view/Spark/job/spark-master-test-maven-hadoop-2.7/4571/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/","from":"developer"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20743","from":"developer"},{"body":"Hi, All.\r\n\r\nThis seems to fail again in branch 2.3. Can we disable this in branch-2.3 for Apache Spark 2.3.1 at least?\r\n- https://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.6/lastCompletedBuild/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/","from":"developer"},{"body":"If that still fails somewhere it means there is still a bug somewhere. I don't think disabling the test is the right thing unless it's actually common enough that it's causing problems. Lots of our tests are flaky.\r\n\r\nThat's like the only failure recently, BTW.\r\nhttps://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.6/lastCompletedBuild/testReport/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/","from":"developer"}],"created":"2018-01-10T01:06:20.000+0000","description":"https://amplab.cs.berkeley.edu/jenkins/job/spark-branch-2.3-test-maven-hadoop-2.7/42/testReport/junit/org.apache.spark.launcher/SparkLauncherSuite/testInProcessLauncher/history/","issue_id":"13129723","key":"SPARK-23020","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-02-02T04:02:18.000+0000","role":"fixed_distractor","summary":"Re-enable Flaky Test: org.apache.spark.launcher.SparkLauncherSuite.testInProcessLauncher"} {"case_id":"13133333","cluster":"DISTRACTOR-SPARK-23200","comments":[{"body":"User 'ssaavedra' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20383","created":"2018-01-24T12:29:04.390+0000"},{"body":"Issue resolved by pull request 20383\r\n\r\nhttps://github.com/apache/spark/pull/20383","created":"2018-01-26T07:27:02.469+0000"},{"body":"Looks like the change was reverted. [~ssaavedra], can you propose this change again with the cleanup? We should target it for 2.4.","created":"2018-03-13T18:30:08.283+0000"},{"body":"Is there any followup here? This seems an important fix to run structural streaming with K8s. cc [~tdas] [~zsxwing]","created":"2018-09-10T14:16:06.041+0000"},{"body":"probably need someone to rebuild on the current config names...","created":"2018-09-11T04:11:58.174+0000"},{"body":"User 'ssaavedra' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22392","created":"2018-09-11T08:33:08.495+0000"},{"body":"User 'ssaavedra' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22392","created":"2018-09-11T08:34:15.897+0000"},{"body":"It seems there wasn't much interest in having streaming working on k8s :)\r\n\r\nI don't currently have an available k8s cluster in which to test the PR, but if I get bandwidth to spawn one I'll test this myself in the next days. Otherwise, I'll build Spark and release the Docker images in Docker Hub for someone else with the resources to reproduce this.\r\n\r\nI recommend you to use the twitter example and have checkpointing configured on a s3a:// bucket.\r\n\r\nSteps: launch with spark-submit once, wait for some checkpoint files to spawn on the bucket, then remove the driver pod, and then re-send with spark-submit again. Check in the logs that the driver loaded successfully and was able to revive the workers and that the names are new (no names from the old instances are shown in the logs as missing).","created":"2018-09-11T08:41:22.527+0000"},{"body":"This is important and should have been a bug not an improvement. Checkpointing needs to work as it is used by default in many cases in production. We should have an integration test for it.\r\n\r\n ","created":"2018-09-14T00:44:13.109+0000"},{"body":"[~cloud_fan]I think this should go in 2.4 even though its a bit late, the remaining PR is trivial. A few properties need to be restored from the checkpoint, and of course it needs testing. I can do the testing if we can get it in 2.4 soon. [~foxish] thoughts?","created":"2018-09-14T01:02:01.278+0000"},{"body":"We should definitely merge it to branch 2.4, but I won't block the release since it's not that critical and it's still in progress. After it's merged, feel free to vote -1 on the RC voting email to include this change, if necessary.","created":"2018-09-17T02:50:40.357+0000"},{"body":"Issue resolved by pull request 22392\n[https://github.com/apache/spark/pull/22392]","created":"2018-09-19T05:10:47.182+0000"}],"conversations":[{"body":"Streaming workloads and restarting from checkpoints may need additional changes, i.e. resetting properties -  see https://github.com/apache-spark-on-k8s/spark/pull/516","from":"reporter","subject":"Reset configuration when restarting from checkpoints"},{"body":"User 'ssaavedra' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20383","from":"developer"},{"body":"Issue resolved by pull request 20383\r\n\r\nhttps://github.com/apache/spark/pull/20383","from":"developer"},{"body":"Looks like the change was reverted. [~ssaavedra], can you propose this change again with the cleanup? We should target it for 2.4.","from":"developer"},{"body":"Is there any followup here? This seems an important fix to run structural streaming with K8s. cc [~tdas] [~zsxwing]","from":"developer"},{"body":"probably need someone to rebuild on the current config names...","from":"developer"},{"body":"User 'ssaavedra' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22392","from":"developer"},{"body":"User 'ssaavedra' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22392","from":"developer"},{"body":"It seems there wasn't much interest in having streaming working on k8s :)\r\n\r\nI don't currently have an available k8s cluster in which to test the PR, but if I get bandwidth to spawn one I'll test this myself in the next days. Otherwise, I'll build Spark and release the Docker images in Docker Hub for someone else with the resources to reproduce this.\r\n\r\nI recommend you to use the twitter example and have checkpointing configured on a s3a:// bucket.\r\n\r\nSteps: launch with spark-submit once, wait for some checkpoint files to spawn on the bucket, then remove the driver pod, and then re-send with spark-submit again. Check in the logs that the driver loaded successfully and was able to revive the workers and that the names are new (no names from the old instances are shown in the logs as missing).","from":"developer"},{"body":"This is important and should have been a bug not an improvement. Checkpointing needs to work as it is used by default in many cases in production. We should have an integration test for it.\r\n\r\n ","from":"developer"},{"body":"[~cloud_fan]I think this should go in 2.4 even though its a bit late, the remaining PR is trivial. A few properties need to be restored from the checkpoint, and of course it needs testing. I can do the testing if we can get it in 2.4 soon. [~foxish] thoughts?","from":"developer"},{"body":"We should definitely merge it to branch 2.4, but I won't block the release since it's not that critical and it's still in progress. After it's merged, feel free to vote -1 on the RC voting email to include this change, if necessary.","from":"developer"},{"body":"Issue resolved by pull request 22392\n[https://github.com/apache/spark/pull/22392]","from":"developer"}],"created":"2018-01-24T10:56:29.000+0000","description":"Streaming workloads and restarting from checkpoints may need additional changes, i.e. resetting properties -  see https://github.com/apache-spark-on-k8s/spark/pull/516","issue_id":"13133333","key":"SPARK-23200","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-09-19T05:10:47.000+0000","role":"fixed_distractor","summary":"Reset configuration when restarting from checkpoints"} {"case_id":"13138167","cluster":"DISTRACTOR-SPARK-23408","comments":[{"body":"User 'jose-torres' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20646","created":"2018-02-20T23:49:05.301+0000"},{"body":"User 'tdas' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20650","created":"2018-02-21T11:14:04.934+0000"},{"body":"Issue resolved by pull request 20650\n[https://github.com/apache/spark/pull/20650]","created":"2018-02-23T20:42:16.344+0000"},{"body":"User 'HeartSaVioR' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/23757","created":"2019-02-11T07:33:55.294+0000"}],"conversations":[{"body":"Seen on an unrelated PR.\r\n\r\nhttps://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/87386/testReport/org.apache.spark.sql.streaming/StreamingOuterJoinSuite/left_outer_early_state_exclusion_on_right/\r\n\r\n{noformat}\r\nsbt.ForkMain$ForkError: org.scalatest.exceptions.TestFailedException: \r\nAssert on query failed: Check total state rows = List(4), updated state rows = List(4): Array(1) did not equal List(4) incorrect updates rows\r\norg.scalatest.Assertions$class.newAssertionFailedException(Assertions.scala:528)\r\n\torg.scalatest.FunSuite.newAssertionFailedException(FunSuite.scala:1560)\r\n\torg.scalatest.Assertions$AssertionsHelper.macroAssert(Assertions.scala:501)\r\n\torg.apache.spark.sql.streaming.StateStoreMetricsTest$$anonfun$assertNumStateRows$1.apply(StateStoreMetricsTest.scala:28)\r\n\torg.apache.spark.sql.streaming.StateStoreMetricsTest$$anonfun$assertNumStateRows$1.apply(StateStoreMetricsTest.scala:23)\r\n\torg.apache.spark.sql.streaming.StreamTest$$anonfun$liftedTree1$1$1$$anonfun$apply$14.apply$mcZ$sp(StreamTest.scala:568)\r\n\torg.apache.spark.sql.streaming.StreamTest$class.verify$1(StreamTest.scala:371)\r\n\torg.apache.spark.sql.streaming.StreamTest$$anonfun$liftedTree1$1$1.apply(StreamTest.scala:568)\r\n\torg.apache.spark.sql.streaming.StreamTest$$anonfun$liftedTree1$1$1.apply(StreamTest.scala:432)\r\n\tscala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\n\r\n\r\n== Progress ==\r\n AddData to MemoryStream[value#19652]: 3,4,5\r\n AddData to MemoryStream[value#19662]: 1,2,3\r\n CheckLastBatch: [3,10,6,9]\r\n=> AssertOnQuery(, Check total state rows = List(4), updated state rows = List(4))\r\n AddData to MemoryStream[value#19652]: 20\r\n AddData to MemoryStream[value#19662]: 21\r\n CheckLastBatch: \r\n AddData to MemoryStream[value#19662]: 20\r\n CheckLastBatch: [20,30,40,60],[4,10,8,null],[5,10,10,null]\r\n\r\n== Stream ==\r\nOutput Mode: Append\r\nStream state: {MemoryStream[value#19652]: 0,MemoryStream[value#19662]: 0}\r\nThread state: alive\r\nThread stack trace: java.lang.Thread.sleep(Native Method)\r\norg.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1.apply$mcZ$sp(MicroBatchExecution.scala:152)\r\norg.apache.spark.sql.execution.streaming.ProcessingTimeExecutor.execute(TriggerExecutor.scala:56)\r\norg.apache.spark.sql.execution.streaming.MicroBatchExecution.runActivatedStream(MicroBatchExecution.scala:120)\r\norg.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:279)\r\norg.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:189)\r\n{noformat}\r\n\r\nNo other failures in the history, though.","from":"reporter","subject":"Flaky test: StreamingOuterJoinSuite.left outer early state exclusion on right"},{"body":"User 'jose-torres' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20646","from":"developer"},{"body":"User 'tdas' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20650","from":"developer"},{"body":"Issue resolved by pull request 20650\n[https://github.com/apache/spark/pull/20650]","from":"developer"},{"body":"User 'HeartSaVioR' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/23757","from":"developer"}],"created":"2018-02-13T13:16:12.000+0000","description":"Seen on an unrelated PR.\r\n\r\nhttps://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/87386/testReport/org.apache.spark.sql.streaming/StreamingOuterJoinSuite/left_outer_early_state_exclusion_on_right/\r\n\r\n{noformat}\r\nsbt.ForkMain$ForkError: org.scalatest.exceptions.TestFailedException: \r\nAssert on query failed: Check total state rows = List(4), updated state rows = List(4): Array(1) did not equal List(4) incorrect updates rows\r\norg.scalatest.Assertions$class.newAssertionFailedException(Assertions.scala:528)\r\n\torg.scalatest.FunSuite.newAssertionFailedException(FunSuite.scala:1560)\r\n\torg.scalatest.Assertions$AssertionsHelper.macroAssert(Assertions.scala:501)\r\n\torg.apache.spark.sql.streaming.StateStoreMetricsTest$$anonfun$assertNumStateRows$1.apply(StateStoreMetricsTest.scala:28)\r\n\torg.apache.spark.sql.streaming.StateStoreMetricsTest$$anonfun$assertNumStateRows$1.apply(StateStoreMetricsTest.scala:23)\r\n\torg.apache.spark.sql.streaming.StreamTest$$anonfun$liftedTree1$1$1$$anonfun$apply$14.apply$mcZ$sp(StreamTest.scala:568)\r\n\torg.apache.spark.sql.streaming.StreamTest$class.verify$1(StreamTest.scala:371)\r\n\torg.apache.spark.sql.streaming.StreamTest$$anonfun$liftedTree1$1$1.apply(StreamTest.scala:568)\r\n\torg.apache.spark.sql.streaming.StreamTest$$anonfun$liftedTree1$1$1.apply(StreamTest.scala:432)\r\n\tscala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\n\r\n\r\n== Progress ==\r\n AddData to MemoryStream[value#19652]: 3,4,5\r\n AddData to MemoryStream[value#19662]: 1,2,3\r\n CheckLastBatch: [3,10,6,9]\r\n=> AssertOnQuery(, Check total state rows = List(4), updated state rows = List(4))\r\n AddData to MemoryStream[value#19652]: 20\r\n AddData to MemoryStream[value#19662]: 21\r\n CheckLastBatch: \r\n AddData to MemoryStream[value#19662]: 20\r\n CheckLastBatch: [20,30,40,60],[4,10,8,null],[5,10,10,null]\r\n\r\n== Stream ==\r\nOutput Mode: Append\r\nStream state: {MemoryStream[value#19652]: 0,MemoryStream[value#19662]: 0}\r\nThread state: alive\r\nThread stack trace: java.lang.Thread.sleep(Native Method)\r\norg.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1.apply$mcZ$sp(MicroBatchExecution.scala:152)\r\norg.apache.spark.sql.execution.streaming.ProcessingTimeExecutor.execute(TriggerExecutor.scala:56)\r\norg.apache.spark.sql.execution.streaming.MicroBatchExecution.runActivatedStream(MicroBatchExecution.scala:120)\r\norg.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:279)\r\norg.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:189)\r\n{noformat}\r\n\r\nNo other failures in the history, though.","issue_id":"13138167","key":"SPARK-23408","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-02-23T20:42:16.000+0000","role":"fixed_distractor","summary":"Flaky test: StreamingOuterJoinSuite.left outer early state exclusion on right"} {"case_id":"13138297","cluster":"DISTRACTOR-SPARK-23416","comments":[{"body":"Thank you for filing this!","created":"2018-02-13T19:24:03.881+0000"},{"body":"I see this failing also with this stacktrace:\r\n\r\n\r\n{code:java}\r\nsbt.ForkMain$ForkError: org.apache.spark.sql.streaming.StreamingQueryException: Query memory [id = cca87cf7-0532-41af-b757-0948ec294c0c, runId = c1830af6-1715-4947-bd76-a1a63482280b] terminated with exception: null\r\n\tat org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:295)\r\n\tat org.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:189)\r\nCaused by: sbt.ForkMain$ForkError: java.lang.InterruptedException: null\r\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer.tryAcquireSharedNanos(AbstractQueuedSynchronizer.java:1326)\r\n\tat scala.concurrent.impl.Promise$DefaultPromise.tryAwait(Promise.scala:208)\r\n\tat scala.concurrent.impl.Promise$DefaultPromise.ready(Promise.scala:218)\r\n\tat scala.concurrent.impl.Promise$DefaultPromise.result(Promise.scala:223)\r\n\tat org.apache.spark.util.ThreadUtils$.awaitResult(ThreadUtils.scala:201)\r\n\tat org.apache.spark.rpc.RpcTimeout.awaitResult(RpcTimeout.scala:75)\r\n\tat org.apache.spark.rpc.RpcEndpointRef.askSync(RpcEndpointRef.scala:92)\r\n\tat org.apache.spark.rpc.RpcEndpointRef.askSync(RpcEndpointRef.scala:76)\r\n\tat org.apache.spark.sql.execution.streaming.continuous.ContinuousExecution.runContinuous(ContinuousExecution.scala:271)\r\n\tat org.apache.spark.sql.execution.streaming.continuous.ContinuousExecution.runActivatedStream(ContinuousExecution.scala:89)\r\n\tat org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:279)\r\n\t... 1 more\r\n{code}\r\n\r\nhttps://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/87401/testReport/org.apache.spark.sql.kafka010/KafkaContinuousSourceStressForDontFailOnDataLossSuite/stress_test_for_failOnDataLoss_false/","created":"2018-02-13T19:30:30.333+0000"},{"body":"I think I see the problem.\r\n * StreamExecution.stop() works by interrupting the stream execution thread. This is not safe in general, and can throw any variety of exceptions.\r\n * StreamExecution.isInterruptedByStop() solves this problem by implementing a whitelist of exceptions which indicate the stop() happened.\r\n * The v2 write path adds calls to ThreadUtils.awaitResult(), which weren't in the V1 write path and (if the interrupt happens to fall in them) throw a new exception which isn't accounted for.\r\n\r\nI'm going to write a PR to add another whitelist entry. This whole edifice is a bit fragile, but I don't have a good solution for that.","created":"2018-02-13T19:47:32.884+0000"},{"body":"User 'jose-torres' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20602","created":"2018-02-13T19:53:04.980+0000"},{"body":"User 'cloud-fan' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20605","created":"2018-02-14T02:22:08.974+0000"},{"body":"Issue resolved by pull request 20605\n[https://github.com/apache/spark/pull/20605]","created":"2018-02-15T09:02:34.986+0000"},{"body":"FYI.\r\nhttps://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/90536 (branch-2.3)\r\nhttps://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-sbt-hadoop-2.7/342/\r\nhttps://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.7/376/\r\nhttps://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-sbt-hadoop-2.7/347/","created":"2018-05-12T14:34:43.169+0000"},{"body":"Sorry. I'm reopening this due to the frequent failures","created":"2018-05-19T17:56:21.263+0000"},{"body":"No problem. I've been working on this since last week.","created":"2018-05-19T18:52:06.348+0000"},{"body":"Thank you so much, [~joseph.torres].","created":"2018-05-20T00:15:27.583+0000"},{"body":"User 'jose-torres' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21384","created":"2018-05-21T19:08:04.766+0000"},{"body":"Issue resolved by pull request 21384\n[https://github.com/apache/spark/pull/21384]","created":"2018-05-24T00:22:03.951+0000"},{"body":"Thank you, [~joseph.torres] and [~tdas] .\r\n\r\nActually, recently failures are reported on `branch-2.3`. Can we have this fix on `branch-2.3` please?","created":"2018-05-24T03:13:02.503+0000"},{"body":"Do you know how to drive that? I'm not sure what the process is.","created":"2018-05-24T19:39:35.231+0000"}],"conversations":[{"body":"I suspect this is a race condition latent in the DataSourceV2 write path, or at least the interaction of that write path with StreamTest.\r\n\r\n[https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/87241/testReport/org.apache.spark.sql.kafka010/KafkaSourceStressForDontFailOnDataLossSuite/stress_test_for_failOnDataLoss_false/]\r\nh3. Error Message\r\n\r\norg.apache.spark.sql.streaming.StreamingQueryException: Query [id = 16b2a2b1-acdd-44ec-902f-531169193169, runId = 9567facb-e305-4554-8622-830519002edb] terminated with exception: Writing job aborted.\r\nh3. Stacktrace\r\n\r\nsbt.ForkMain$ForkError: org.apache.spark.sql.streaming.StreamingQueryException: Query [id = 16b2a2b1-acdd-44ec-902f-531169193169, runId = 9567facb-e305-4554-8622-830519002edb] terminated with exception: Writing job aborted. at org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:295) at org.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:189) Caused by: sbt.ForkMain$ForkError: org.apache.spark.SparkException: Writing job aborted. at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec.doExecute(WriteToDataSourceV2.scala:108) at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:131) at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:127) at org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:155) at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151) at org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:152) at org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:127) at org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:247) at org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:294) at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collectFromPlan(Dataset.scala:3272) at org.apache.spark.sql.Dataset$$anonfun$collect$1.apply(Dataset.scala:2722) at org.apache.spark.sql.Dataset$$anonfun$collect$1.apply(Dataset.scala:2722) at org.apache.spark.sql.Dataset$$anonfun$52.apply(Dataset.scala:3253) at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:77) at org.apache.spark.sql.Dataset.withAction(Dataset.scala:3252) at org.apache.spark.sql.Dataset.collect(Dataset.scala:2722) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$3$$anonfun$apply$15.apply(MicroBatchExecution.scala:488) at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:77) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$3.apply(MicroBatchExecution.scala:483) at org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:271) at org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58) at org.apache.spark.sql.execution.streaming.MicroBatchExecution.org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch(MicroBatchExecution.scala:482) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply$mcV$sp(MicroBatchExecution.scala:133) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:121) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:121) at org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:271) at org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1.apply$mcZ$sp(MicroBatchExecution.scala:121) at org.apache.spark.sql.execution.streaming.ProcessingTimeExecutor.execute(TriggerExecutor.scala:56) at org.apache.spark.sql.execution.streaming.MicroBatchExecution.runActivatedStream(MicroBatchExecution.scala:117) at org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:279) ... 1 more Caused by: sbt.ForkMain$ForkError: java.lang.InterruptedException: null at java.util.concurrent.locks.AbstractQueuedSynchronizer.doAcquireSharedInterruptibly(AbstractQueuedSynchronizer.java:998) at java.util.concurrent.locks.AbstractQueuedSynchronizer.acquireSharedInterruptibly(AbstractQueuedSynchronizer.java:1304) at scala.concurrent.impl.Promise$DefaultPromise.tryAwait(Promise.scala:202) at scala.concurrent.impl.Promise$DefaultPromise.ready(Promise.scala:218) at scala.concurrent.impl.Promise$DefaultPromise.ready(Promise.scala:153) at org.apache.spark.util.ThreadUtils$.awaitReady(ThreadUtils.scala:222) at org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:633) at org.apache.spark.SparkContext.runJob(SparkContext.scala:2027) at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec.doExecute(WriteToDataSourceV2.scala:79) ... 31 more","from":"reporter","subject":"Flaky test: KafkaSourceStressForDontFailOnDataLossSuite.stress test for failOnDataLoss=false"},{"body":"Thank you for filing this!","from":"developer"},{"body":"I see this failing also with this stacktrace:\r\n\r\n\r\n{code:java}\r\nsbt.ForkMain$ForkError: org.apache.spark.sql.streaming.StreamingQueryException: Query memory [id = cca87cf7-0532-41af-b757-0948ec294c0c, runId = c1830af6-1715-4947-bd76-a1a63482280b] terminated with exception: null\r\n\tat org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:295)\r\n\tat org.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:189)\r\nCaused by: sbt.ForkMain$ForkError: java.lang.InterruptedException: null\r\n\tat java.util.concurrent.locks.AbstractQueuedSynchronizer.tryAcquireSharedNanos(AbstractQueuedSynchronizer.java:1326)\r\n\tat scala.concurrent.impl.Promise$DefaultPromise.tryAwait(Promise.scala:208)\r\n\tat scala.concurrent.impl.Promise$DefaultPromise.ready(Promise.scala:218)\r\n\tat scala.concurrent.impl.Promise$DefaultPromise.result(Promise.scala:223)\r\n\tat org.apache.spark.util.ThreadUtils$.awaitResult(ThreadUtils.scala:201)\r\n\tat org.apache.spark.rpc.RpcTimeout.awaitResult(RpcTimeout.scala:75)\r\n\tat org.apache.spark.rpc.RpcEndpointRef.askSync(RpcEndpointRef.scala:92)\r\n\tat org.apache.spark.rpc.RpcEndpointRef.askSync(RpcEndpointRef.scala:76)\r\n\tat org.apache.spark.sql.execution.streaming.continuous.ContinuousExecution.runContinuous(ContinuousExecution.scala:271)\r\n\tat org.apache.spark.sql.execution.streaming.continuous.ContinuousExecution.runActivatedStream(ContinuousExecution.scala:89)\r\n\tat org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:279)\r\n\t... 1 more\r\n{code}\r\n\r\nhttps://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/87401/testReport/org.apache.spark.sql.kafka010/KafkaContinuousSourceStressForDontFailOnDataLossSuite/stress_test_for_failOnDataLoss_false/","from":"developer"},{"body":"I think I see the problem.\r\n * StreamExecution.stop() works by interrupting the stream execution thread. This is not safe in general, and can throw any variety of exceptions.\r\n * StreamExecution.isInterruptedByStop() solves this problem by implementing a whitelist of exceptions which indicate the stop() happened.\r\n * The v2 write path adds calls to ThreadUtils.awaitResult(), which weren't in the V1 write path and (if the interrupt happens to fall in them) throw a new exception which isn't accounted for.\r\n\r\nI'm going to write a PR to add another whitelist entry. This whole edifice is a bit fragile, but I don't have a good solution for that.","from":"developer"},{"body":"User 'jose-torres' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20602","from":"developer"},{"body":"User 'cloud-fan' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/20605","from":"developer"},{"body":"Issue resolved by pull request 20605\n[https://github.com/apache/spark/pull/20605]","from":"developer"},{"body":"FYI.\r\nhttps://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/90536 (branch-2.3)\r\nhttps://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-sbt-hadoop-2.7/342/\r\nhttps://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-maven-hadoop-2.7/376/\r\nhttps://amplab.cs.berkeley.edu/jenkins/view/Spark%20QA%20Test%20(Dashboard)/job/spark-branch-2.3-test-sbt-hadoop-2.7/347/","from":"developer"},{"body":"Sorry. I'm reopening this due to the frequent failures","from":"developer"},{"body":"No problem. I've been working on this since last week.","from":"developer"},{"body":"Thank you so much, [~joseph.torres].","from":"developer"},{"body":"User 'jose-torres' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21384","from":"developer"},{"body":"Issue resolved by pull request 21384\n[https://github.com/apache/spark/pull/21384]","from":"developer"},{"body":"Thank you, [~joseph.torres] and [~tdas] .\r\n\r\nActually, recently failures are reported on `branch-2.3`. Can we have this fix on `branch-2.3` please?","from":"developer"},{"body":"Do you know how to drive that? I'm not sure what the process is.","from":"developer"}],"created":"2018-02-13T19:08:42.000+0000","description":"I suspect this is a race condition latent in the DataSourceV2 write path, or at least the interaction of that write path with StreamTest.\r\n\r\n[https://amplab.cs.berkeley.edu/jenkins/job/SparkPullRequestBuilder/87241/testReport/org.apache.spark.sql.kafka010/KafkaSourceStressForDontFailOnDataLossSuite/stress_test_for_failOnDataLoss_false/]\r\nh3. Error Message\r\n\r\norg.apache.spark.sql.streaming.StreamingQueryException: Query [id = 16b2a2b1-acdd-44ec-902f-531169193169, runId = 9567facb-e305-4554-8622-830519002edb] terminated with exception: Writing job aborted.\r\nh3. Stacktrace\r\n\r\nsbt.ForkMain$ForkError: org.apache.spark.sql.streaming.StreamingQueryException: Query [id = 16b2a2b1-acdd-44ec-902f-531169193169, runId = 9567facb-e305-4554-8622-830519002edb] terminated with exception: Writing job aborted. at org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:295) at org.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:189) Caused by: sbt.ForkMain$ForkError: org.apache.spark.SparkException: Writing job aborted. at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec.doExecute(WriteToDataSourceV2.scala:108) at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:131) at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:127) at org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:155) at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151) at org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:152) at org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:127) at org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:247) at org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:294) at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collectFromPlan(Dataset.scala:3272) at org.apache.spark.sql.Dataset$$anonfun$collect$1.apply(Dataset.scala:2722) at org.apache.spark.sql.Dataset$$anonfun$collect$1.apply(Dataset.scala:2722) at org.apache.spark.sql.Dataset$$anonfun$52.apply(Dataset.scala:3253) at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:77) at org.apache.spark.sql.Dataset.withAction(Dataset.scala:3252) at org.apache.spark.sql.Dataset.collect(Dataset.scala:2722) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$3$$anonfun$apply$15.apply(MicroBatchExecution.scala:488) at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:77) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$3.apply(MicroBatchExecution.scala:483) at org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:271) at org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58) at org.apache.spark.sql.execution.streaming.MicroBatchExecution.org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch(MicroBatchExecution.scala:482) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply$mcV$sp(MicroBatchExecution.scala:133) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:121) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:121) at org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:271) at org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58) at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1.apply$mcZ$sp(MicroBatchExecution.scala:121) at org.apache.spark.sql.execution.streaming.ProcessingTimeExecutor.execute(TriggerExecutor.scala:56) at org.apache.spark.sql.execution.streaming.MicroBatchExecution.runActivatedStream(MicroBatchExecution.scala:117) at org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:279) ... 1 more Caused by: sbt.ForkMain$ForkError: java.lang.InterruptedException: null at java.util.concurrent.locks.AbstractQueuedSynchronizer.doAcquireSharedInterruptibly(AbstractQueuedSynchronizer.java:998) at java.util.concurrent.locks.AbstractQueuedSynchronizer.acquireSharedInterruptibly(AbstractQueuedSynchronizer.java:1304) at scala.concurrent.impl.Promise$DefaultPromise.tryAwait(Promise.scala:202) at scala.concurrent.impl.Promise$DefaultPromise.ready(Promise.scala:218) at scala.concurrent.impl.Promise$DefaultPromise.ready(Promise.scala:153) at org.apache.spark.util.ThreadUtils$.awaitReady(ThreadUtils.scala:222) at org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:633) at org.apache.spark.SparkContext.runJob(SparkContext.scala:2027) at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec.doExecute(WriteToDataSourceV2.scala:79) ... 31 more","issue_id":"13138297","key":"SPARK-23416","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-05-24T00:22:03.000+0000","role":"fixed_distractor","summary":"Flaky test: KafkaSourceStressForDontFailOnDataLossSuite.stress test for failOnDataLoss=false"} {"case_id":"13143777","cluster":"DISTRACTOR-SPARK-23636","comments":[{"body":"It seems in KafkaDataConsumer#close :\r\n{code}\r\n def close(): Unit = consumer.close()\r\n{code}\r\nThe code should catch ConcurrentModificationException and try closing the consumer again.","created":"2018-06-29T05:26:59.359+0000"},{"body":"[~mcdeepak] It should be resolved in SPARK-19185 with 2.4.0 (even if it's different API). Can you re-test it please?","created":"2018-12-12T08:49:22.163+0000"},{"body":"Please reopen if problem re-appears.","created":"2019-01-07T12:43:10.989+0000"}],"conversations":[{"body":"h2.  \r\nh2. Summary\r\n\r\n \r\n\r\nWhile using the KafkaUtils.createRDD API - we receive below listed error, specifically when 1 executor connects to 1 kafka topic-partition, but with more than 1 core & fetches an Array(OffsetRanges)\r\n\r\n \r\n\r\n_I've tagged this issue to \"Structured Streaming\" - as I could not find a more appropriate component_ \r\n\r\n \r\n----\r\nh2. Error Faced\r\n{noformat}\r\njava.util.ConcurrentModificationException: KafkaConsumer is not safe for multi-threaded access{noformat}\r\n Stack Trace\r\n{noformat}\r\nCaused by: org.apache.spark.SparkException: Job aborted due to stage failure: Task 5 in stage 1.0 failed 4 times, most recent failure: Lost task 5.3 in stage 1.0 (TID 17, host, executor 16): java.util.ConcurrentModificationException: KafkaConsumer is not safe for multi-threaded access\r\nat org.apache.kafka.clients.consumer.KafkaConsumer.acquire(KafkaConsumer.java:1629)\r\nat org.apache.kafka.clients.consumer.KafkaConsumer.close(KafkaConsumer.java:1528)\r\nat org.apache.kafka.clients.consumer.KafkaConsumer.close(KafkaConsumer.java:1508)\r\nat org.apache.spark.streaming.kafka010.CachedKafkaConsumer.close(CachedKafkaConsumer.scala:59)\r\nat org.apache.spark.streaming.kafka010.CachedKafkaConsumer$.remove(CachedKafkaConsumer.scala:185)\r\nat org.apache.spark.streaming.kafka010.KafkaRDD$KafkaRDDIterator.(KafkaRDD.scala:204)\r\nat org.apache.spark.streaming.kafka010.KafkaRDD.compute(KafkaRDD.scala:181)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323){noformat}\r\n \r\n----\r\nh2. Config Used to simulate the error\r\n\r\nA session with : \r\n * Executors - 1\r\n * Cores - 2 or More\r\n * Kafka Topic - has only 1 partition\r\n * While fetching - More than one Array of Offset Range , Example \r\n\r\n{noformat}\r\nArray(OffsetRange(\"kafka_topic\",0,608954201,608954202),\r\nOffsetRange(\"kafka_topic\",0,608954202,608954203)\r\n){noformat}\r\n \r\n----\r\nh2. Was this approach working before?\r\n\r\n \r\n\r\nThis was working in spark 1.6.2\r\n\r\nHowever, from spark 2.1 onwards - the approach throws exception\r\n\r\n \r\n----\r\nh2. Why are we fetching from kafka as mentioned above.\r\n\r\n \r\n\r\nThis gives us the capability to establish a connection to Kafka Broker for every spark executor's core, thus each core can fetch/process its own set of messages based on the specified (offset ranges).\r\n\r\n \r\n\r\n \r\n----\r\nh2. Sample Code\r\n\r\n \r\n{quote}scala snippet - on versions spark 2.2.0 or 2.1.0\r\n\r\n// Bunch of imports\r\n\r\nimport kafka.serializer.\\{DefaultDecoder, StringDecoder}\r\n import org.apache.avro.generic.GenericRecord\r\n import org.apache.kafka.clients.consumer.ConsumerRecord\r\n import org.apache.kafka.common.serialization._\r\n import org.apache.spark.rdd.RDD\r\n import org.apache.spark.sql.\\{DataFrame, Row, SQLContext}\r\n import org.apache.spark.sql.Row\r\n import org.apache.spark.sql.hive.HiveContext\r\n import org.apache.spark.sql.types.\\{StringType, StructField, StructType}\r\n import org.apache.spark.streaming.kafka010._\r\n import org.apache.spark.streaming.kafka010.KafkaUtils._\r\n{quote}\r\n{quote}// This forces two connections - from a single executor - to topic-partition .\r\n\r\n// And with 2 cores assigned to 1 executor : each core has a task - pulling respective offsets : OffsetRange(\"kafka_topic\",0,1,2) & OffsetRange(\"kafka_topic\",0,2,3)\r\n\r\nval parallelizedRanges = Array(OffsetRange(\"kafka_topic\",0,1,2), // Fetching sample 2 records \r\n OffsetRange(\"kafka_topic\",0,2,3) // Fetching sample 2 records \r\n )\r\n\r\n \r\n\r\n// Initiate kafka properties\r\n\r\nval kafkaParams1: java.util.Map[String, Object] = new java.util.HashMap()\r\n\r\n// kafkaParams1.put(\"key\",\"val\") add all the parameters such as broker, topic.... Not listing every property here.\r\n\r\n \r\n\r\n// Create RDD\r\n\r\nval rDDConsumerRec: RDD[ConsumerRecord[String, String]] =\r\n createRDD[String, String](sparkContext\r\n , kafkaParams1, parallelizedRanges, LocationStrategies.PreferConsistent)\r\n\r\n \r\n\r\n// Map Function\r\n\r\nval data: RDD[Row] = rDDConsumerRec.map \\{ x => Row(x.topic().toString, x.partition().toString, x.offset().toString, x.timestamp().toString, x.value() ) }\r\n\r\n \r\n\r\n// Create a DataFrame\r\n\r\nval df = sqlContext.createDataFrame(data, StructType(\r\n Seq(\r\n StructField(\"topic\", StringType),\r\n StructField(\"partition\", StringType),\r\n StructField(\"offset\", StringType),\r\n StructField(\"timestamp\", StringType),\r\n StructField(\"value\", BinaryType)\r\n )))\r\n\r\n \r\n\r\ndf.show() //  You will see the error reported.\r\n{quote}\r\n \r\n----\r\n \r\nh2. Similar Issue reported earlier, but on a different API\r\n\r\n \r\n\r\nA similar issue reported for DirectStream is \r\n\r\nhttps://issues.apache.org/jira/browse/SPARK-19185\r\n\r\n \r\n\r\n \r\n\r\n \r\n----\r\nh2. What is the impact - if a fix is not available for this problem?\r\n\r\n \r\n\r\n \r\n\r\nWe have a lot of Spark Applications that are running in production, making parallel connections to the 1 topic-partition from each spark-executor: so parallelism is directly proportional to the num-cores in each executor.\r\n\r\nWith spark 2.1 onwards : we are not allowed to make concurrent connections from 1 executor to 1 topic-partition. Only workaround is to start our applications with executor-cores = 1, with dynamic resource allocation enabled.\r\n\r\nWith above configuration - for every offset range we ask kafka - a new executor is spawned to run the fetch task.\r\n\r\nDownside of Workaround -\r\n\r\nAbove approach is not allowing us to leverage more than 1 spark-core per spark-executor.\r\n\r\nAnd asking for an executor - for each offset range - is costly : in terms of scheduling and allocation.\r\n\r\n \r\n\r\n \r\n\r\n ","from":"reporter","subject":"[SPARK 2.2] | Kafka Consumer | KafkaUtils.createRDD throws Exception - java.util.ConcurrentModificationException: KafkaConsumer is not safe for multi-threaded access"},{"body":"It seems in KafkaDataConsumer#close :\r\n{code}\r\n def close(): Unit = consumer.close()\r\n{code}\r\nThe code should catch ConcurrentModificationException and try closing the consumer again.","from":"developer"},{"body":"[~mcdeepak] It should be resolved in SPARK-19185 with 2.4.0 (even if it's different API). Can you re-test it please?","from":"developer"},{"body":"Please reopen if problem re-appears.","from":"developer"}],"created":"2018-03-09T03:12:22.000+0000","description":"h2.  \r\nh2. Summary\r\n\r\n \r\n\r\nWhile using the KafkaUtils.createRDD API - we receive below listed error, specifically when 1 executor connects to 1 kafka topic-partition, but with more than 1 core & fetches an Array(OffsetRanges)\r\n\r\n \r\n\r\n_I've tagged this issue to \"Structured Streaming\" - as I could not find a more appropriate component_ \r\n\r\n \r\n----\r\nh2. Error Faced\r\n{noformat}\r\njava.util.ConcurrentModificationException: KafkaConsumer is not safe for multi-threaded access{noformat}\r\n Stack Trace\r\n{noformat}\r\nCaused by: org.apache.spark.SparkException: Job aborted due to stage failure: Task 5 in stage 1.0 failed 4 times, most recent failure: Lost task 5.3 in stage 1.0 (TID 17, host, executor 16): java.util.ConcurrentModificationException: KafkaConsumer is not safe for multi-threaded access\r\nat org.apache.kafka.clients.consumer.KafkaConsumer.acquire(KafkaConsumer.java:1629)\r\nat org.apache.kafka.clients.consumer.KafkaConsumer.close(KafkaConsumer.java:1528)\r\nat org.apache.kafka.clients.consumer.KafkaConsumer.close(KafkaConsumer.java:1508)\r\nat org.apache.spark.streaming.kafka010.CachedKafkaConsumer.close(CachedKafkaConsumer.scala:59)\r\nat org.apache.spark.streaming.kafka010.CachedKafkaConsumer$.remove(CachedKafkaConsumer.scala:185)\r\nat org.apache.spark.streaming.kafka010.KafkaRDD$KafkaRDDIterator.(KafkaRDD.scala:204)\r\nat org.apache.spark.streaming.kafka010.KafkaRDD.compute(KafkaRDD.scala:181)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323){noformat}\r\n \r\n----\r\nh2. Config Used to simulate the error\r\n\r\nA session with : \r\n * Executors - 1\r\n * Cores - 2 or More\r\n * Kafka Topic - has only 1 partition\r\n * While fetching - More than one Array of Offset Range , Example \r\n\r\n{noformat}\r\nArray(OffsetRange(\"kafka_topic\",0,608954201,608954202),\r\nOffsetRange(\"kafka_topic\",0,608954202,608954203)\r\n){noformat}\r\n \r\n----\r\nh2. Was this approach working before?\r\n\r\n \r\n\r\nThis was working in spark 1.6.2\r\n\r\nHowever, from spark 2.1 onwards - the approach throws exception\r\n\r\n \r\n----\r\nh2. Why are we fetching from kafka as mentioned above.\r\n\r\n \r\n\r\nThis gives us the capability to establish a connection to Kafka Broker for every spark executor's core, thus each core can fetch/process its own set of messages based on the specified (offset ranges).\r\n\r\n \r\n\r\n \r\n----\r\nh2. Sample Code\r\n\r\n \r\n{quote}scala snippet - on versions spark 2.2.0 or 2.1.0\r\n\r\n// Bunch of imports\r\n\r\nimport kafka.serializer.\\{DefaultDecoder, StringDecoder}\r\n import org.apache.avro.generic.GenericRecord\r\n import org.apache.kafka.clients.consumer.ConsumerRecord\r\n import org.apache.kafka.common.serialization._\r\n import org.apache.spark.rdd.RDD\r\n import org.apache.spark.sql.\\{DataFrame, Row, SQLContext}\r\n import org.apache.spark.sql.Row\r\n import org.apache.spark.sql.hive.HiveContext\r\n import org.apache.spark.sql.types.\\{StringType, StructField, StructType}\r\n import org.apache.spark.streaming.kafka010._\r\n import org.apache.spark.streaming.kafka010.KafkaUtils._\r\n{quote}\r\n{quote}// This forces two connections - from a single executor - to topic-partition .\r\n\r\n// And with 2 cores assigned to 1 executor : each core has a task - pulling respective offsets : OffsetRange(\"kafka_topic\",0,1,2) & OffsetRange(\"kafka_topic\",0,2,3)\r\n\r\nval parallelizedRanges = Array(OffsetRange(\"kafka_topic\",0,1,2), // Fetching sample 2 records \r\n OffsetRange(\"kafka_topic\",0,2,3) // Fetching sample 2 records \r\n )\r\n\r\n \r\n\r\n// Initiate kafka properties\r\n\r\nval kafkaParams1: java.util.Map[String, Object] = new java.util.HashMap()\r\n\r\n// kafkaParams1.put(\"key\",\"val\") add all the parameters such as broker, topic.... Not listing every property here.\r\n\r\n \r\n\r\n// Create RDD\r\n\r\nval rDDConsumerRec: RDD[ConsumerRecord[String, String]] =\r\n createRDD[String, String](sparkContext\r\n , kafkaParams1, parallelizedRanges, LocationStrategies.PreferConsistent)\r\n\r\n \r\n\r\n// Map Function\r\n\r\nval data: RDD[Row] = rDDConsumerRec.map \\{ x => Row(x.topic().toString, x.partition().toString, x.offset().toString, x.timestamp().toString, x.value() ) }\r\n\r\n \r\n\r\n// Create a DataFrame\r\n\r\nval df = sqlContext.createDataFrame(data, StructType(\r\n Seq(\r\n StructField(\"topic\", StringType),\r\n StructField(\"partition\", StringType),\r\n StructField(\"offset\", StringType),\r\n StructField(\"timestamp\", StringType),\r\n StructField(\"value\", BinaryType)\r\n )))\r\n\r\n \r\n\r\ndf.show() //  You will see the error reported.\r\n{quote}\r\n \r\n----\r\n \r\nh2. Similar Issue reported earlier, but on a different API\r\n\r\n \r\n\r\nA similar issue reported for DirectStream is \r\n\r\nhttps://issues.apache.org/jira/browse/SPARK-19185\r\n\r\n \r\n\r\n \r\n\r\n \r\n----\r\nh2. What is the impact - if a fix is not available for this problem?\r\n\r\n \r\n\r\n \r\n\r\nWe have a lot of Spark Applications that are running in production, making parallel connections to the 1 topic-partition from each spark-executor: so parallelism is directly proportional to the num-cores in each executor.\r\n\r\nWith spark 2.1 onwards : we are not allowed to make concurrent connections from 1 executor to 1 topic-partition. Only workaround is to start our applications with executor-cores = 1, with dynamic resource allocation enabled.\r\n\r\nWith above configuration - for every offset range we ask kafka - a new executor is spawned to run the fetch task.\r\n\r\nDownside of Workaround -\r\n\r\nAbove approach is not allowing us to leverage more than 1 spark-core per spark-executor.\r\n\r\nAnd asking for an executor - for each offset range - is costly : in terms of scheduling and allocation.\r\n\r\n \r\n\r\n \r\n\r\n ","issue_id":"13143777","key":"SPARK-23636","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-01-07T12:43:10.000+0000","role":"fixed_distractor","summary":"[SPARK 2.2] | Kafka Consumer | KafkaUtils.createRDD throws Exception - java.util.ConcurrentModificationException: KafkaConsumer is not safe for multi-threaded access"} {"case_id":"13149097","cluster":"DISTRACTOR-SPARK-23829","comments":[{"body":"In 2.4 it's fixed as it's using 2.0.0. I think an upgrade will solve this issue (if description about versions is correct).","created":"2018-09-17T14:21:13.810+0000"},{"body":"Confirmed, upgrading to spark 2.4 and kafka 2.0 resolves this issue.","created":"2018-12-04T23:09:10.577+0000"},{"body":"[~collin-scangarella] thanks for your efforts!","created":"2018-12-05T08:30:33.623+0000"},{"body":"{color:#14892c}I am using 2.4.0-chd6.3.0, and got this error again.{color}\r\n\r\nLogical Plan:\r\n TypedFilter , interface org.apache.spark.sql.Row, [StructField(personId,LongType,true), StructField(code,LongType,true), StructField(eventTime,LongType,true), StructField(processTime,LongType,true)], createexternalrow(personId#35L, code#36L, eventTime#33L, processTime#34L, StructField(personId,LongType,true), StructField(code,LongType,true), StructField(eventTime,LongType,true), StructField(processTime,LongType,true))\r\n +- Project [json#30.personId AS personId#35L, json#30.code AS code#36L, unix_timestamp(to_utc_timestamp(cast(json#30.eventTime as timestamp), GMT-8), yyyy-MM-dd HH:mm:ss, Some(Asia/Shanghai)) AS eventTime#33L, unix_timestamp(timestamp#12, yyyy-MM-dd HH:mm:ss, Some(Asia/Shanghai)) AS processTime#34L|#33L, unix_timestamp(timestamp#12, yyyy-MM-dd HH:mm:ss, Some(Asia/Shanghai)) AS processTime#34L]\r\n +- Project [jsontostructs(StructField(code,LongType,true), StructField(eventTime,StringType,true), StructField(personId,LongType,true), cast(value#8 as string), Some(Asia/Shanghai)) AS json#30, timestamp#12|#8 as string), Some(Asia/Shanghai)) AS json#30, timestamp#12]\r\n +- StreamingExecutionRelation KafkaV2[Subscribe[capp-events]], [key#7, value#8, topic#9, partition#10, offset#11L, timestamp#12, timestampType#13|#7, value#8, topic#9, partition#10, offset#11L, timestamp#12, timestampType#13]\r\n\r\nat org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:295)\r\n at org.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:189)\r\n Caused by: org.apache.spark.SparkException: Writing job aborted.\r\n at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec.doExecute(WriteToDataSourceV2Exec.scala:92)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:131)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:155)\r\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\r\n at org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:152)\r\n at org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:247)\r\n at org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:296)\r\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collectFromPlan(Dataset.scala:3383)\r\n at org.apache.spark.sql.Dataset$$anonfun$collect$1.apply(Dataset.scala:2782)","created":"2019-08-16T03:35:54.692+0000"},{"body":"[~leo.zhi] Just seen your comment. From the stack you've provided I'm not able to tell whether it's the same issue or not. Please open a new jira and attach full stacktrace and logs. If you ping me on the jira I can take a look next week.","created":"2019-09-06T19:53:00.248+0000"},{"body":"Hi [~gsomogyi], I am experiencing the same issue leo.zhi mentioned in his comment. Was wondering if there has been any progress on it/ticket number to follow?\r\n\r\nWe are using *2.4.0-cdh6.1.1*\r\n\r\n \r\n\r\nIn our case it starts with:\r\n\r\n \r\n Lost task 0.0 in stage 5965.0 (TID 15060, ..., executor 1): java.util.concurrent.TimeoutException: Cannot fetch record for offset 813529 in 2048 milliseconds\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer.fetchData(KafkaDataConsumer.scala:488)\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer.org$apache$spark$sql$kafka010$InternalKafkaConsumer$$fetchRecord(KafkaDataConsumer.scala:371)\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer$$anonfun$get$1.apply(KafkaDataConsumer.scala:251)\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer$$anonfun$get$1.apply(KafkaDataConsumer.scala:234)\r\n at org.apache.spark.util.UninterruptibleThread.runUninterruptibly(UninterruptibleThread.scala:77)\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer.runUninterruptiblyIfPossible(KafkaDataConsumer.scala:209)\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer.get(KafkaDataConsumer.scala:234)\r\n at org.apache.spark.sql.kafka010.KafkaDataConsumer$class.get(KafkaDataConsumer.scala:64)\r\n at org.apache.spark.sql.kafka010.KafkaDataConsumer$CachedKafkaDataConsumer.get(KafkaDataConsumer.scala:500)\r\n at org.apache.spark.sql.kafka010.KafkaMicroBatchInputPartitionReader.next(KafkaMicroBatchReader.scala:337)\r\n at org.apache.spark.sql.execution.datasources.v2.DataSourceRDD$$anon$1.hasNext(DataSourceRDD.scala:49)\r\n at org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:37)\r\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)\r\n at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\n at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$11$$anon$1.hasNext(WholeStageCodegenExec.scala:624)\r\n at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage2.processNext(Unknown Source)\r\n at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\n at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$11$$anon$1.hasNext(WholeStageCodegenExec.scala:624)\r\n at org.apache.spark.sql.execution.datasources.v2.DataWritingSparkTask$$anonfun$run$3.apply(WriteToDataSourceV2Exec.scala:117)\r\n at org.apache.spark.sql.execution.datasources.v2.DataWritingSparkTask$$anonfun$run$3.apply(WriteToDataSourceV2Exec.scala:116)\r\n at org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1394)\r\n at org.apache.spark.sql.execution.datasources.v2.DataWritingSparkTask$.run(WriteToDataSourceV2Exec.scala:146)\r\n at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec$$anonfun$doExecute$2.apply(WriteToDataSourceV2Exec.scala:67)\r\n at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec$$anonfun$doExecute$2.apply(WriteToDataSourceV2Exec.scala:66)\r\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:90)\r\n at org.apache.spark.scheduler.Task.run(Task.scala:121)\r\n at org.apache.spark.executor.Executor$TaskRunner$$anonfun$11.apply(Executor.scala:407)\r\n at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1360)\r\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:413)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n at java.lang.Thread.run(Thread.java:748)\r\n  \r\n And then:\r\n  \r\n Logical plan:\r\n ...\r\n Caused by: org.apache.spark.SparkException: Writing job aborted.\r\n at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec.doExecute(WriteToDataSourceV2Exec.scala:92)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:131)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:155)\r\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\r\n at org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:152)\r\n at org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:247)\r\n at org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:296)\r\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collectFromPlan(Dataset.scala:3383)\r\n at org.apache.spark.sql.Dataset$$anonfun$collect$1.apply(Dataset.scala:2782)\r\n at org.apache.spark.sql.Dataset$$anonfun$collect$1.apply(Dataset.scala:2782)\r\n at org.apache.spark.sql.Dataset$$anonfun$53.apply(Dataset.scala:3364)\r\n at org.apache.spark.sql.execution.SQLExecution$$anonfun$withNewExecutionId$1.apply(SQLExecution.scala:78)\r\n at org.apache.spark.sql.execution.SQLExecution$.withSQLConfPropagated(SQLExecution.scala:125)\r\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:73)\r\n at org.apache.spark.sql.Dataset.withAction(Dataset.scala:3363)\r\n at org.apache.spark.sql.Dataset.collect(Dataset.scala:2782)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$5$$anonfun$apply$17.apply(MicroBatchExecution.scala:537)\r\n at org.apache.spark.sql.execution.SQLExecution$$anonfun$withNewExecutionId$1.apply(SQLExecution.scala:78)\r\n at org.apache.spark.sql.execution.SQLExecution$.withSQLConfPropagated(SQLExecution.scala:125)\r\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:73)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$5.apply(MicroBatchExecution.scala:532)\r\n at org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:351)\r\n at org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution.org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch(MicroBatchExecution.scala:531)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply$mcV$sp(MicroBatchExecution.scala:198)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:166)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:166)\r\n at org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:351)\r\n at org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1.apply$mcZ$sp(MicroBatchExecution.scala:166)\r\n at org.apache.spark.sql.execution.streaming.ProcessingTimeExecutor.execute(TriggerExecutor.scala:56)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution.runActivatedStream(MicroBatchExecution.scala:160)\r\n at org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:279)\r\n  ","created":"2020-01-28T01:53:25.314+0000"},{"body":"Cannot fetch record means Spark initiated the fetch but was timing out. First I would take a look at what happened on the Kafka side (since Kafka didn't respond in time). If you still think it's a Spark issue I would like to reproduce it with vanilla Spark + create a new jira with logs.","created":"2020-01-28T09:05:19.376+0000"},{"body":"[~gsomogyi]  me new to kafka and spark , what do you mean spark-vanila version here ?","created":"2020-02-04T09:46:11.414+0000"},{"body":"[~BdLearner] I mean upstream Spark.","created":"2020-02-04T12:46:55.765+0000"},{"body":"[~gsomogyi] Isn't the fix backported to spark 2.3?\r\n We are restricted to Spark 2.3 due to HDP and facing this issue often. Tried excluding and implicitly adding kafka-clients 2.1.1 (our kafka version), but it doesn't seem to help.\r\n\r\nIs there a way around on Spark 2.3?","created":"2020-07-07T07:01:35.321+0000"},{"body":"This is not backported to 2.3.\r\nHere is the list of Kafka changes relevant from Spark point of view: https://gist.github.com/gaborgsomogyi/3476c32d69ff2087ed5d7d031653c7a9\r\nKafka client can be hacked around manually but I discourage because it's risky.\r\n\r\nI understand the upgrade constraints but I highly encourage to make further efforts to jump to 2.4 because that's not the only severe issue which 2.3 contains!\r\n","created":"2020-07-07T08:53:49.523+0000"},{"body":"Thanks for the update [~gsomogyi]\r\n\r\nThe investigation doesn't seem to mention the root cause, Do we know why this happens and Is there some steps to reproduce this problem consistently?\r\n\r\nI was able to replicate it consistently with transactional commits. When end offset = N, Spark 2.3 fails with a TimeoutException if (N-1)^th^ offset was a transaction-commit offset and there are no more messages after that.","created":"2020-07-09T08:48:18.911+0000"},{"body":"Since this is race issue there is no consistent way to reproduce. Please reproduce it with our latest release (3.x) and open a Jira with the driver and executor logs.","created":"2020-07-16T10:47:21.887+0000"}],"conversations":[{"body":"In spark 2.3 , it provides a source \"spark-sql-kafka-0-10_2.11\".\r\n\r\n \r\n\r\nWhen I wanted to read from my kafka-0.10.2.1 cluster, it throws out an error \"*java.util.concurrent.TimeoutException: Cannot fetch record xxxx for offset in 12000 milliseconds*\"  frequently , and the job thus failed.\r\n\r\n \r\n\r\nI searched on google & stackoverflow for a while, and found many other people who got this excption too, and nobody gave an answer.\r\n\r\n \r\n\r\nI debuged the source code, found nothing, but I guess it's because the dependency spark-sql-kafka-0-10_2.11 is using.\r\n\r\n \r\n{code:java}\r\n\r\n org.apache.spark\r\n spark-sql-kafka-0-10_2.11\r\n 2.3.0\r\n \r\n \r\n kafka-clients\r\n org.apache.kafka\r\n \r\n \r\n\r\n\r\n org.apache.kafka\r\n kafka-clients\r\n 0.10.2.1\r\n{code}\r\nI excluded it from maven ,and added another version , rerun the code , and now it works.\r\n\r\n \r\n\r\nI guess something is wrong on kafka-clients0.10.0.1 working with kafka0.10.2.1, or more kafka versions. \r\n\r\n \r\n\r\nHope for an explanation.\r\n\r\nHere is the error stack.\r\n{code:java}\r\n[ERROR] 2018-03-30 13:34:11,404 [stream execution thread for [id = 83076cf1-4bf0-4c82-a0b3-23d8432f5964, runId = b3e18aa6-358f-43f6-a077-e34db0822df6]] org.apache.spark.sql.execution.streaming.MicroBatchExecution logError - Query [id = 83076cf1-4bf0-4c82-a0b3-23d8432f5964, runId = b3e18aa6-358f-43f6-a077-e34db0822df6] terminated with error\r\norg.apache.spark.SparkException: Job aborted due to stage failure: Task 6 in stage 0.0 failed 1 times, most recent failure: Lost task 6.0 in stage 0.0 (TID 6, localhost, executor driver): java.util.concurrent.TimeoutException: Cannot fetch record for offset 6481521 in 120000 milliseconds\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.org$apache$spark$sql$kafka010$CachedKafkaConsumer$$fetchData(CachedKafkaConsumer.scala:230)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer$$anonfun$get$1.apply(CachedKafkaConsumer.scala:122)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer$$anonfun$get$1.apply(CachedKafkaConsumer.scala:106)\r\nat org.apache.spark.util.UninterruptibleThread.runUninterruptibly(UninterruptibleThread.scala:77)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.runUninterruptiblyIfPossible(CachedKafkaConsumer.scala:68)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.get(CachedKafkaConsumer.scala:106)\r\nat org.apache.spark.sql.kafka010.KafkaSourceRDD$$anon$1.getNext(KafkaSourceRDD.scala:157)\r\nat org.apache.spark.sql.kafka010.KafkaSourceRDD$$anon$1.getNext(KafkaSourceRDD.scala:148)\r\nat org.apache.spark.util.NextIterator.hasNext(NextIterator.scala:73)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)\r\nat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\nat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$10$$anon$1.hasNext(WholeStageCodegenExec.scala:614)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$12.hasNext(Iterator.scala:440)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage2.processNext(Unknown Source)\r\nat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\nat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$10$$anon$1.hasNext(WholeStageCodegenExec.scala:614)\r\nat org.apache.spark.sql.execution.aggregate.ObjectHashAggregateExec$$anonfun$doExecute$1$$anonfun$2.apply(ObjectHashAggregateExec.scala:107)\r\nat org.apache.spark.sql.execution.aggregate.ObjectHashAggregateExec$$anonfun$doExecute$1$$anonfun$2.apply(ObjectHashAggregateExec.scala:105)\r\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndexInternal$1$$anonfun$apply$24.apply(RDD.scala:818)\r\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndexInternal$1$$anonfun$apply$24.apply(RDD.scala:818)\r\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:324)\r\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:288)\r\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:324)\r\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:288)\r\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)\r\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)\r\nat org.apache.spark.scheduler.Task.run(Task.scala:109)\r\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)\r\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\nat java.lang.Thread.run(Thread.java:748)\r\n\r\nDriver stacktrace:\r\nat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1599)\r\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1587)\r\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1586)\r\nat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\nat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\r\nat org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1586)\r\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:831)\r\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:831)\r\nat scala.Option.foreach(Option.scala:257)\r\nat org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:831)\r\nat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:1820)\r\nat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1769)\r\nat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1758)\r\nat org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\r\nat org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:642)\r\nat org.apache.spark.SparkContext.runJob(SparkContext.scala:2027)\r\nat org.apache.spark.SparkContext.runJob(SparkContext.scala:2048)\r\nat org.apache.spark.SparkContext.runJob(SparkContext.scala:2067)\r\nat org.apache.spark.SparkContext.runJob(SparkContext.scala:2092)\r\nat org.apache.spark.rdd.RDD$$anonfun$foreachPartition$1.apply(RDD.scala:929)\r\nat org.apache.spark.rdd.RDD$$anonfun$foreachPartition$1.apply(RDD.scala:927)\r\nat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\r\nat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:112)\r\nat org.apache.spark.rdd.RDD.withScope(RDD.scala:363)\r\nat org.apache.spark.rdd.RDD.foreachPartition(RDD.scala:927)\r\nat org.apache.spark.sql.execution.streaming.ForeachSink.addBatch(ForeachSink.scala:49)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$3$$anonfun$apply$16.apply(MicroBatchExecution.scala:477)\r\nat org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:77)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$3.apply(MicroBatchExecution.scala:475)\r\nat org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:271)\r\nat org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution.org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch(MicroBatchExecution.scala:474)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply$mcV$sp(MicroBatchExecution.scala:133)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:121)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:121)\r\nat org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:271)\r\nat org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1.apply$mcZ$sp(MicroBatchExecution.scala:121)\r\nat org.apache.spark.sql.execution.streaming.ProcessingTimeExecutor.execute(TriggerExecutor.scala:56)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution.runActivatedStream(MicroBatchExecution.scala:117)\r\nat org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:279)\r\nat org.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:189)\r\nCaused by: java.util.concurrent.TimeoutException: Cannot fetch record for offset 6481521 in 120000 milliseconds\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.org$apache$spark$sql$kafka010$CachedKafkaConsumer$$fetchData(CachedKafkaConsumer.scala:230)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer$$anonfun$get$1.apply(CachedKafkaConsumer.scala:122)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer$$anonfun$get$1.apply(CachedKafkaConsumer.scala:106)\r\nat org.apache.spark.util.UninterruptibleThread.runUninterruptibly(UninterruptibleThread.scala:77)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.runUninterruptiblyIfPossible(CachedKafkaConsumer.scala:68)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.get(CachedKafkaConsumer.scala:106)\r\nat org.apache.spark.sql.kafka010.KafkaSourceRDD$$anon$1.getNext(KafkaSourceRDD.scala:157)\r\nat org.apache.spark.sql.kafka010.KafkaSourceRDD$$anon$1.getNext(KafkaSourceRDD.scala:148)\r\nat org.apache.spark.util.NextIterator.hasNext(NextIterator.scala:73)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)\r\nat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\nat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$10$$anon$1.hasNext(WholeStageCodegenExec.scala:614)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$12.hasNext(Iterator.scala:440)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage2.processNext(Unknown Source)\r\nat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\nat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$10$$anon$1.hasNext(WholeStageCodegenExec.scala:614)\r\nat org.apache.spark.sql.execution.aggregate.ObjectHashAggregateExec$$anonfun$doExecute$1$$anonfun$2.apply(ObjectHashAggregateExec.scala:107)\r\nat org.apache.spark.sql.execution.aggregate.ObjectHashAggregateExec$$anonfun$doExecute$1$$anonfun$2.apply(ObjectHashAggregateExec.scala:105)\r\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndexInternal$1$$anonfun$apply$24.apply(RDD.scala:818)\r\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndexInternal$1$$anonfun$apply$24.apply(RDD.scala:818)\r\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:324)\r\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:288)\r\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:324)\r\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:288)\r\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)\r\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)\r\nat org.apache.spark.scheduler.Task.run(Task.scala:109)\r\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)\r\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\nat java.lang.Thread.run(Thread.java:748)\r\n{code}","from":"reporter","subject":"spark-sql-kafka source in spark 2.3 causes reading stream failure frequently"},{"body":"In 2.4 it's fixed as it's using 2.0.0. I think an upgrade will solve this issue (if description about versions is correct).","from":"developer"},{"body":"Confirmed, upgrading to spark 2.4 and kafka 2.0 resolves this issue.","from":"developer"},{"body":"[~collin-scangarella] thanks for your efforts!","from":"developer"},{"body":"{color:#14892c}I am using 2.4.0-chd6.3.0, and got this error again.{color}\r\n\r\nLogical Plan:\r\n TypedFilter , interface org.apache.spark.sql.Row, [StructField(personId,LongType,true), StructField(code,LongType,true), StructField(eventTime,LongType,true), StructField(processTime,LongType,true)], createexternalrow(personId#35L, code#36L, eventTime#33L, processTime#34L, StructField(personId,LongType,true), StructField(code,LongType,true), StructField(eventTime,LongType,true), StructField(processTime,LongType,true))\r\n +- Project [json#30.personId AS personId#35L, json#30.code AS code#36L, unix_timestamp(to_utc_timestamp(cast(json#30.eventTime as timestamp), GMT-8), yyyy-MM-dd HH:mm:ss, Some(Asia/Shanghai)) AS eventTime#33L, unix_timestamp(timestamp#12, yyyy-MM-dd HH:mm:ss, Some(Asia/Shanghai)) AS processTime#34L|#33L, unix_timestamp(timestamp#12, yyyy-MM-dd HH:mm:ss, Some(Asia/Shanghai)) AS processTime#34L]\r\n +- Project [jsontostructs(StructField(code,LongType,true), StructField(eventTime,StringType,true), StructField(personId,LongType,true), cast(value#8 as string), Some(Asia/Shanghai)) AS json#30, timestamp#12|#8 as string), Some(Asia/Shanghai)) AS json#30, timestamp#12]\r\n +- StreamingExecutionRelation KafkaV2[Subscribe[capp-events]], [key#7, value#8, topic#9, partition#10, offset#11L, timestamp#12, timestampType#13|#7, value#8, topic#9, partition#10, offset#11L, timestamp#12, timestampType#13]\r\n\r\nat org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:295)\r\n at org.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:189)\r\n Caused by: org.apache.spark.SparkException: Writing job aborted.\r\n at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec.doExecute(WriteToDataSourceV2Exec.scala:92)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:131)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:155)\r\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\r\n at org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:152)\r\n at org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:247)\r\n at org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:296)\r\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collectFromPlan(Dataset.scala:3383)\r\n at org.apache.spark.sql.Dataset$$anonfun$collect$1.apply(Dataset.scala:2782)","from":"developer"},{"body":"[~leo.zhi] Just seen your comment. From the stack you've provided I'm not able to tell whether it's the same issue or not. Please open a new jira and attach full stacktrace and logs. If you ping me on the jira I can take a look next week.","from":"developer"},{"body":"Hi [~gsomogyi], I am experiencing the same issue leo.zhi mentioned in his comment. Was wondering if there has been any progress on it/ticket number to follow?\r\n\r\nWe are using *2.4.0-cdh6.1.1*\r\n\r\n \r\n\r\nIn our case it starts with:\r\n\r\n \r\n Lost task 0.0 in stage 5965.0 (TID 15060, ..., executor 1): java.util.concurrent.TimeoutException: Cannot fetch record for offset 813529 in 2048 milliseconds\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer.fetchData(KafkaDataConsumer.scala:488)\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer.org$apache$spark$sql$kafka010$InternalKafkaConsumer$$fetchRecord(KafkaDataConsumer.scala:371)\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer$$anonfun$get$1.apply(KafkaDataConsumer.scala:251)\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer$$anonfun$get$1.apply(KafkaDataConsumer.scala:234)\r\n at org.apache.spark.util.UninterruptibleThread.runUninterruptibly(UninterruptibleThread.scala:77)\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer.runUninterruptiblyIfPossible(KafkaDataConsumer.scala:209)\r\n at org.apache.spark.sql.kafka010.InternalKafkaConsumer.get(KafkaDataConsumer.scala:234)\r\n at org.apache.spark.sql.kafka010.KafkaDataConsumer$class.get(KafkaDataConsumer.scala:64)\r\n at org.apache.spark.sql.kafka010.KafkaDataConsumer$CachedKafkaDataConsumer.get(KafkaDataConsumer.scala:500)\r\n at org.apache.spark.sql.kafka010.KafkaMicroBatchInputPartitionReader.next(KafkaMicroBatchReader.scala:337)\r\n at org.apache.spark.sql.execution.datasources.v2.DataSourceRDD$$anon$1.hasNext(DataSourceRDD.scala:49)\r\n at org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:37)\r\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)\r\n at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\n at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$11$$anon$1.hasNext(WholeStageCodegenExec.scala:624)\r\n at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\n at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage2.processNext(Unknown Source)\r\n at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\n at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$11$$anon$1.hasNext(WholeStageCodegenExec.scala:624)\r\n at org.apache.spark.sql.execution.datasources.v2.DataWritingSparkTask$$anonfun$run$3.apply(WriteToDataSourceV2Exec.scala:117)\r\n at org.apache.spark.sql.execution.datasources.v2.DataWritingSparkTask$$anonfun$run$3.apply(WriteToDataSourceV2Exec.scala:116)\r\n at org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1394)\r\n at org.apache.spark.sql.execution.datasources.v2.DataWritingSparkTask$.run(WriteToDataSourceV2Exec.scala:146)\r\n at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec$$anonfun$doExecute$2.apply(WriteToDataSourceV2Exec.scala:67)\r\n at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec$$anonfun$doExecute$2.apply(WriteToDataSourceV2Exec.scala:66)\r\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:90)\r\n at org.apache.spark.scheduler.Task.run(Task.scala:121)\r\n at org.apache.spark.executor.Executor$TaskRunner$$anonfun$11.apply(Executor.scala:407)\r\n at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1360)\r\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:413)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n at java.lang.Thread.run(Thread.java:748)\r\n  \r\n And then:\r\n  \r\n Logical plan:\r\n ...\r\n Caused by: org.apache.spark.SparkException: Writing job aborted.\r\n at org.apache.spark.sql.execution.datasources.v2.WriteToDataSourceV2Exec.doExecute(WriteToDataSourceV2Exec.scala:92)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:131)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:155)\r\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\r\n at org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:152)\r\n at org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:247)\r\n at org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:296)\r\n at org.apache.spark.sql.Dataset.org$apache$spark$sql$Dataset$$collectFromPlan(Dataset.scala:3383)\r\n at org.apache.spark.sql.Dataset$$anonfun$collect$1.apply(Dataset.scala:2782)\r\n at org.apache.spark.sql.Dataset$$anonfun$collect$1.apply(Dataset.scala:2782)\r\n at org.apache.spark.sql.Dataset$$anonfun$53.apply(Dataset.scala:3364)\r\n at org.apache.spark.sql.execution.SQLExecution$$anonfun$withNewExecutionId$1.apply(SQLExecution.scala:78)\r\n at org.apache.spark.sql.execution.SQLExecution$.withSQLConfPropagated(SQLExecution.scala:125)\r\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:73)\r\n at org.apache.spark.sql.Dataset.withAction(Dataset.scala:3363)\r\n at org.apache.spark.sql.Dataset.collect(Dataset.scala:2782)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$5$$anonfun$apply$17.apply(MicroBatchExecution.scala:537)\r\n at org.apache.spark.sql.execution.SQLExecution$$anonfun$withNewExecutionId$1.apply(SQLExecution.scala:78)\r\n at org.apache.spark.sql.execution.SQLExecution$.withSQLConfPropagated(SQLExecution.scala:125)\r\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:73)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$5.apply(MicroBatchExecution.scala:532)\r\n at org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:351)\r\n at org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution.org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch(MicroBatchExecution.scala:531)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply$mcV$sp(MicroBatchExecution.scala:198)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:166)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:166)\r\n at org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:351)\r\n at org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1.apply$mcZ$sp(MicroBatchExecution.scala:166)\r\n at org.apache.spark.sql.execution.streaming.ProcessingTimeExecutor.execute(TriggerExecutor.scala:56)\r\n at org.apache.spark.sql.execution.streaming.MicroBatchExecution.runActivatedStream(MicroBatchExecution.scala:160)\r\n at org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:279)\r\n  ","from":"developer"},{"body":"Cannot fetch record means Spark initiated the fetch but was timing out. First I would take a look at what happened on the Kafka side (since Kafka didn't respond in time). If you still think it's a Spark issue I would like to reproduce it with vanilla Spark + create a new jira with logs.","from":"developer"},{"body":"[~gsomogyi]  me new to kafka and spark , what do you mean spark-vanila version here ?","from":"developer"},{"body":"[~BdLearner] I mean upstream Spark.","from":"developer"},{"body":"[~gsomogyi] Isn't the fix backported to spark 2.3?\r\n We are restricted to Spark 2.3 due to HDP and facing this issue often. Tried excluding and implicitly adding kafka-clients 2.1.1 (our kafka version), but it doesn't seem to help.\r\n\r\nIs there a way around on Spark 2.3?","from":"developer"},{"body":"This is not backported to 2.3.\r\nHere is the list of Kafka changes relevant from Spark point of view: https://gist.github.com/gaborgsomogyi/3476c32d69ff2087ed5d7d031653c7a9\r\nKafka client can be hacked around manually but I discourage because it's risky.\r\n\r\nI understand the upgrade constraints but I highly encourage to make further efforts to jump to 2.4 because that's not the only severe issue which 2.3 contains!\r\n","from":"developer"},{"body":"Thanks for the update [~gsomogyi]\r\n\r\nThe investigation doesn't seem to mention the root cause, Do we know why this happens and Is there some steps to reproduce this problem consistently?\r\n\r\nI was able to replicate it consistently with transactional commits. When end offset = N, Spark 2.3 fails with a TimeoutException if (N-1)^th^ offset was a transaction-commit offset and there are no more messages after that.","from":"developer"},{"body":"Since this is race issue there is no consistent way to reproduce. Please reproduce it with our latest release (3.x) and open a Jira with the driver and executor logs.","from":"developer"}],"created":"2018-03-30T05:38:59.000+0000","description":"In spark 2.3 , it provides a source \"spark-sql-kafka-0-10_2.11\".\r\n\r\n \r\n\r\nWhen I wanted to read from my kafka-0.10.2.1 cluster, it throws out an error \"*java.util.concurrent.TimeoutException: Cannot fetch record xxxx for offset in 12000 milliseconds*\"  frequently , and the job thus failed.\r\n\r\n \r\n\r\nI searched on google & stackoverflow for a while, and found many other people who got this excption too, and nobody gave an answer.\r\n\r\n \r\n\r\nI debuged the source code, found nothing, but I guess it's because the dependency spark-sql-kafka-0-10_2.11 is using.\r\n\r\n \r\n{code:java}\r\n\r\n org.apache.spark\r\n spark-sql-kafka-0-10_2.11\r\n 2.3.0\r\n \r\n \r\n kafka-clients\r\n org.apache.kafka\r\n \r\n \r\n\r\n\r\n org.apache.kafka\r\n kafka-clients\r\n 0.10.2.1\r\n{code}\r\nI excluded it from maven ,and added another version , rerun the code , and now it works.\r\n\r\n \r\n\r\nI guess something is wrong on kafka-clients0.10.0.1 working with kafka0.10.2.1, or more kafka versions. \r\n\r\n \r\n\r\nHope for an explanation.\r\n\r\nHere is the error stack.\r\n{code:java}\r\n[ERROR] 2018-03-30 13:34:11,404 [stream execution thread for [id = 83076cf1-4bf0-4c82-a0b3-23d8432f5964, runId = b3e18aa6-358f-43f6-a077-e34db0822df6]] org.apache.spark.sql.execution.streaming.MicroBatchExecution logError - Query [id = 83076cf1-4bf0-4c82-a0b3-23d8432f5964, runId = b3e18aa6-358f-43f6-a077-e34db0822df6] terminated with error\r\norg.apache.spark.SparkException: Job aborted due to stage failure: Task 6 in stage 0.0 failed 1 times, most recent failure: Lost task 6.0 in stage 0.0 (TID 6, localhost, executor driver): java.util.concurrent.TimeoutException: Cannot fetch record for offset 6481521 in 120000 milliseconds\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.org$apache$spark$sql$kafka010$CachedKafkaConsumer$$fetchData(CachedKafkaConsumer.scala:230)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer$$anonfun$get$1.apply(CachedKafkaConsumer.scala:122)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer$$anonfun$get$1.apply(CachedKafkaConsumer.scala:106)\r\nat org.apache.spark.util.UninterruptibleThread.runUninterruptibly(UninterruptibleThread.scala:77)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.runUninterruptiblyIfPossible(CachedKafkaConsumer.scala:68)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.get(CachedKafkaConsumer.scala:106)\r\nat org.apache.spark.sql.kafka010.KafkaSourceRDD$$anon$1.getNext(KafkaSourceRDD.scala:157)\r\nat org.apache.spark.sql.kafka010.KafkaSourceRDD$$anon$1.getNext(KafkaSourceRDD.scala:148)\r\nat org.apache.spark.util.NextIterator.hasNext(NextIterator.scala:73)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)\r\nat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\nat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$10$$anon$1.hasNext(WholeStageCodegenExec.scala:614)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$12.hasNext(Iterator.scala:440)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage2.processNext(Unknown Source)\r\nat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\nat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$10$$anon$1.hasNext(WholeStageCodegenExec.scala:614)\r\nat org.apache.spark.sql.execution.aggregate.ObjectHashAggregateExec$$anonfun$doExecute$1$$anonfun$2.apply(ObjectHashAggregateExec.scala:107)\r\nat org.apache.spark.sql.execution.aggregate.ObjectHashAggregateExec$$anonfun$doExecute$1$$anonfun$2.apply(ObjectHashAggregateExec.scala:105)\r\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndexInternal$1$$anonfun$apply$24.apply(RDD.scala:818)\r\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndexInternal$1$$anonfun$apply$24.apply(RDD.scala:818)\r\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:324)\r\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:288)\r\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:324)\r\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:288)\r\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)\r\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)\r\nat org.apache.spark.scheduler.Task.run(Task.scala:109)\r\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)\r\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\nat java.lang.Thread.run(Thread.java:748)\r\n\r\nDriver stacktrace:\r\nat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1599)\r\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1587)\r\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1586)\r\nat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\r\nat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48)\r\nat org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1586)\r\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:831)\r\nat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:831)\r\nat scala.Option.foreach(Option.scala:257)\r\nat org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:831)\r\nat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.doOnReceive(DAGScheduler.scala:1820)\r\nat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1769)\r\nat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1758)\r\nat org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\r\nat org.apache.spark.scheduler.DAGScheduler.runJob(DAGScheduler.scala:642)\r\nat org.apache.spark.SparkContext.runJob(SparkContext.scala:2027)\r\nat org.apache.spark.SparkContext.runJob(SparkContext.scala:2048)\r\nat org.apache.spark.SparkContext.runJob(SparkContext.scala:2067)\r\nat org.apache.spark.SparkContext.runJob(SparkContext.scala:2092)\r\nat org.apache.spark.rdd.RDD$$anonfun$foreachPartition$1.apply(RDD.scala:929)\r\nat org.apache.spark.rdd.RDD$$anonfun$foreachPartition$1.apply(RDD.scala:927)\r\nat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\r\nat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:112)\r\nat org.apache.spark.rdd.RDD.withScope(RDD.scala:363)\r\nat org.apache.spark.rdd.RDD.foreachPartition(RDD.scala:927)\r\nat org.apache.spark.sql.execution.streaming.ForeachSink.addBatch(ForeachSink.scala:49)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$3$$anonfun$apply$16.apply(MicroBatchExecution.scala:477)\r\nat org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:77)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch$3.apply(MicroBatchExecution.scala:475)\r\nat org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:271)\r\nat org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution.org$apache$spark$sql$execution$streaming$MicroBatchExecution$$runBatch(MicroBatchExecution.scala:474)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply$mcV$sp(MicroBatchExecution.scala:133)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:121)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1$$anonfun$apply$mcZ$sp$1.apply(MicroBatchExecution.scala:121)\r\nat org.apache.spark.sql.execution.streaming.ProgressReporter$class.reportTimeTaken(ProgressReporter.scala:271)\r\nat org.apache.spark.sql.execution.streaming.StreamExecution.reportTimeTaken(StreamExecution.scala:58)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution$$anonfun$runActivatedStream$1.apply$mcZ$sp(MicroBatchExecution.scala:121)\r\nat org.apache.spark.sql.execution.streaming.ProcessingTimeExecutor.execute(TriggerExecutor.scala:56)\r\nat org.apache.spark.sql.execution.streaming.MicroBatchExecution.runActivatedStream(MicroBatchExecution.scala:117)\r\nat org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runStream(StreamExecution.scala:279)\r\nat org.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:189)\r\nCaused by: java.util.concurrent.TimeoutException: Cannot fetch record for offset 6481521 in 120000 milliseconds\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.org$apache$spark$sql$kafka010$CachedKafkaConsumer$$fetchData(CachedKafkaConsumer.scala:230)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer$$anonfun$get$1.apply(CachedKafkaConsumer.scala:122)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer$$anonfun$get$1.apply(CachedKafkaConsumer.scala:106)\r\nat org.apache.spark.util.UninterruptibleThread.runUninterruptibly(UninterruptibleThread.scala:77)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.runUninterruptiblyIfPossible(CachedKafkaConsumer.scala:68)\r\nat org.apache.spark.sql.kafka010.CachedKafkaConsumer.get(CachedKafkaConsumer.scala:106)\r\nat org.apache.spark.sql.kafka010.KafkaSourceRDD$$anon$1.getNext(KafkaSourceRDD.scala:157)\r\nat org.apache.spark.sql.kafka010.KafkaSourceRDD$$anon$1.getNext(KafkaSourceRDD.scala:148)\r\nat org.apache.spark.util.NextIterator.hasNext(NextIterator.scala:73)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)\r\nat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\nat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$10$$anon$1.hasNext(WholeStageCodegenExec.scala:614)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat scala.collection.Iterator$$anon$12.hasNext(Iterator.scala:440)\r\nat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:409)\r\nat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage2.processNext(Unknown Source)\r\nat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\nat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$10$$anon$1.hasNext(WholeStageCodegenExec.scala:614)\r\nat org.apache.spark.sql.execution.aggregate.ObjectHashAggregateExec$$anonfun$doExecute$1$$anonfun$2.apply(ObjectHashAggregateExec.scala:107)\r\nat org.apache.spark.sql.execution.aggregate.ObjectHashAggregateExec$$anonfun$doExecute$1$$anonfun$2.apply(ObjectHashAggregateExec.scala:105)\r\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndexInternal$1$$anonfun$apply$24.apply(RDD.scala:818)\r\nat org.apache.spark.rdd.RDD$$anonfun$mapPartitionsWithIndexInternal$1$$anonfun$apply$24.apply(RDD.scala:818)\r\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:324)\r\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:288)\r\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:324)\r\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:288)\r\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)\r\nat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)\r\nat org.apache.spark.scheduler.Task.run(Task.scala:109)\r\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)\r\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\nat java.lang.Thread.run(Thread.java:748)\r\n{code}","issue_id":"13149097","key":"SPARK-23829","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-12-05T08:35:42.000+0000","role":"fixed_distractor","summary":"spark-sql-kafka source in spark 2.3 causes reading stream failure frequently"} {"case_id":"13152273","cluster":"DISTRACTOR-SPARK-23977","comments":[{"body":"User 'steveloughran' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21066","created":"2018-04-13T13:17:05.742+0000"},{"body":"[~stevel@apache.org] is the intention of this ticket to incorporate the Hadoop libraries within Spark itself, ie. no Hadoop dependency?\r\nTrying to understand whether this is a viable solution for Spark on Kubernetes writing to S3","created":"2018-05-07T01:34:48.796+0000"},{"body":"It will need the hadoop-aws module and deoendencies as that is where the core code is. This patch just does the binding to the InsertIntoHadoopFS relation (move to Hadoop MRv2 FileOutputFormat & expect the new superclass, PathOutputCommtter, rather than always a FileOutputcommitter, and for Parquet, something similar with a ParquetOutputCommitter.\r\n\r\n+its only in Hadoop 3.1, though you can backport to branch-2, especially if you are prepared to bump up the minimum java version to 8 in that branch.\r\n\r\nt should work on k8s, given it works standalone. All it needs is an endpoint supporting the multipart upload operation of S3, which includes some non-AWS object stores.\r\n\r\n Look @ the HADOOP-13786 work and the paper [a zero rename committer|https://github.com/steveloughran/zero-rename-committer/releases/download/tag_draft_003/a_zero_rename_committer.pdf]. \r\n\r\nAnd there's some integration tests downstream in https://github.com/hortonworks-spark/cloud-integration . I can help set you up to run those, if you email me directly. Essentially: you need to choose which stores to test against from: s3, openstack, azure, and configure them\r\n\r\nNote that of the two variant committers, \"staging\" and \"magic\", the magic one needs a consistent S3 endpoint, which you only get on AWS S3 with an external services, usually dynamo DB based (S3mper, EMR consisent S3, S3Guard). The staging one needs enough local HDD to buffer the output of all active tasks, but doesn't need that consistency for its own query. You will need a plan for chaining together work though, which is inevitably one of \"consistency layer\" or \"wait long enough between writer and reader that you expect the metadata to be consistent\"\r\n\r\nFinally, if you are using spark to write directly to S3 today, without any consistency layer, then your commit algorithm had better not be mimicing directory rename by list + copy + delete. You need this code for safe as well as performant committing of work to S3.\r\n\r\n","created":"2018-05-07T13:38:24.239+0000"},{"body":"I'm removing the target version, since we are not going to merge it to 2.4","created":"2018-09-10T13:58:46.794+0000"},{"body":"Issue resolved by pull request 24970\n[https://github.com/apache/spark/pull/24970]","created":"2019-08-15T17:16:11.421+0000"},{"body":"[~stevel@apache.org] How to workaround following exception during the execution of \"INSERT OVERWRITE\" in spark.sql (spark 3.1.1 with hadoop 3.2)?\r\n{quote}  if (dynamicPartitionOverwrite) {\r\n\r\n    // until there's explicit extensions to the PathOutputCommitProtocols\r\n\r\n    // to support the spark mechanism, it's left to the individual committer\r\n\r\n    // choice to handle partitioning.\r\n\r\n    throw new IOException(PathOutputCommitProtocol.UNSUPPORTED)\r\n\r\n  }\r\n{quote}","created":"2021-03-25T21:39:14.312+0000"},{"body":"use the partitioned committer and configure it to do the right thing when updating partitions (merge, delete everything already there, fail)","created":"2021-03-26T12:29:58.526+0000"},{"body":"[~stevel@apache.org] Thanks for the info. Below are the related (key, value) we used:\r\n # spark.hadoop.fs.s3a.committer.name — partitioned\r\n # spark.hadoop.mapreduce.outputcommitter.factory.scheme.s3a — org.apache.hadoop.fs.s3a.commit.S3ACommitterFactory\r\n # spark.sql.sources.commitProtocolClass — org.apache.spark.internal.io.cloud.PathOutputCommitProtocol\r\n # spark.sql.parquet.output.committer.class — org.apache.spark.internal.io.cloud.BindingParquetOutputCommitter\r\n\r\n3 & 4 appear to be necessary to ensure S3A committers being used by Spark for parquet outputs, except that \"INSERT OVERWRITE\" is blocked by the dynamicPartitionOverwrite exception. It will be helpful and appreciated if you can patiently elaborate on the proper way to \"use the partitioned committer and configure it to do the right thing ...\" in Spark. For example:\r\n * PathOutputCommitProtocol appears to be needed but its constructor fails with the exception.\r\n * S3A committers honor \"fs.s3a.committer.staging.conflict-mode\" which needs to be \"replace\" for \"INSERT OVERWRITE\" but \"append\" for \"INSERT INTO\". So it is spark.sql query specific. How to make spark.sql automatically set the right value?\r\n * Does above require code change in Spark or there is a configuration-only way?","created":"2021-03-30T21:14:57.662+0000"},{"body":"the spark settings don't make it down from sql; you can do more at the RDD API level.\r\n\r\nThe problem is that the spark partition insert logic all relies on renaming which has the O(data) performance penalty as well as the other commit correctness issues. Yes, something to do the pushdown could be done, or with the multipart APIs of HADOOP-13186 give Spark a standard API to implement a zero rename committer in its own code.\r\n\r\nHowever, focus is on things like Iceberg and Delta lake, which offer more in terms of : \r\n* atomic job commit\r\n* avoid the performance and cost issues of relying on directory tree scan as a way to identify source files.; the performance issues of doing all IO down a single shard of S3 storage.\r\nI do not disagree with the direction of that work; we have to view the S3A committers (and IBM's Stocator + AWS EMR spark committers) as the final attempts to maintain that \"it's just a directory tree\" model into a cloud world where directories don't always exist, and listing them is measurable in hundreds of milliseconds. ","created":"2021-04-03T16:27:56.058+0000"},{"body":"Thank you very much [~stevel@apache.org] for your explanations.\r\n\r\nI am experiencing the same problems as [~danzhi] and your comments helped me a lot.\r\n\r\nSad magic committer does not work with dynamic partition overwrite because it has an amazing performance when writing loads of JSON partitioned data.","created":"2021-11-01T10:33:12.243+0000"},{"body":"[~gumartinm] can I draw your attention to Apache Iceberg?\r\n\r\nmeanwhile\r\nMAPREDUCE-7341 adds a high performance targeting abfs and gcs; all tree scanning is in task commit, which is atomic; job commit aggressively parallelised and optimized for stores whose listStatusIterator calls are incremental with prefetching: we can start processing at the first page of task manifests found in a listing well the second Page is still being retrieved. Also in there: rate limiting, IO Statistics Collection.\r\n\r\nHADOOP-17833 I will pick up some of that work, including incremental loading and rate limiting. And if we can keep reads and writes below the S3 IOPS limits, we should be able to avoid situations where we have to start sleeping and re-trying.\r\n\r\nHADOOP-17981 is my homework this week -emergency work to deal with a rare but current failure in abfs under heavy load.\r\n","created":"2021-11-02T10:06:20.126+0000"},{"body":"Hi all,\r\n\r\n \r\n\r\nI follow the [recommendations|https://spark.apache.org/docs/latest/cloud-integration.html#parquet-io-settings] and getting the following warning:\r\n{code:java}\r\n2022-02-16 15:22:03.292 WARN FlowThread0 ParquetOutputFormat: Setting parquet.enable.summary-metadata is deprecated, please use parquet.summary.metadata.level {code}\r\nWhat would be the recommended value for `parquet.summary.metadata.level ` ? \r\n ","created":"2022-02-16T15:28:11.650+0000"}],"conversations":[{"body":"Hadoop 3.1 adds a mechanism for job-specific and store-specific committers (MAPREDUCE-6823, MAPREDUCE-6956), and one key implementation, S3A committers, HADOOP-13786\r\n\r\nThese committers deliver high-performance output of MR and spark jobs to S3, and offer the key semantics which Spark depends on: no visible output until job commit, a failure of a task at an stage, including partway through task commit, can be handled by executing and committing another task attempt. \r\n\r\nIn contrast, the FileOutputFormat commit algorithms on S3 have issues:\r\n\r\n* Awful performance because files are copied by rename\r\n* FileOutputFormat v1: weak task commit failure recovery semantics as the (v1) expectation: \"directory renames are atomic\" doesn't hold.\r\n* S3 metadata eventual consistency can cause rename to miss files or fail entirely (SPARK-15849)\r\n\r\nNote also that FileOutputFormat \"v2\" commit algorithm doesn't offer any of the commit semantics w.r.t observability of or recovery from task commit failure, on any filesystem.\r\n\r\nThe S3A committers address these by way of uploading all data to the destination through multipart uploads, uploads which are only completed in job commit.\r\n\r\nThe new {{PathOutputCommitter}} factory mechanism allows applications to work with the S3A committers and any other, by adding a plugin mechanism into the MRv2 FileOutputFormat class, where it job config and filesystem configuration options can dynamically choose the output committer.\r\n\r\nSpark can use these with some binding classes to \r\n\r\n# Add a subclass of {{HadoopMapReduceCommitProtocol}} which uses the MRv2 classes and {{PathOutputCommitterFactory}} to create the committers.\r\n# Add a {{BindingParquetOutputCommitter extends ParquetOutputCommitter}}\r\nto wire up Parquet output even when code requires the committer to be a subclass of {{ParquetOutputCommitter}}\r\n\r\nThis patch builds on SPARK-23807 for setting up the dependencies.","from":"reporter","subject":"Add commit protocol binding to Hadoop 3.1 PathOutputCommitter mechanism"},{"body":"User 'steveloughran' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21066","from":"developer"},{"body":"[~stevel@apache.org] is the intention of this ticket to incorporate the Hadoop libraries within Spark itself, ie. no Hadoop dependency?\r\nTrying to understand whether this is a viable solution for Spark on Kubernetes writing to S3","from":"developer"},{"body":"It will need the hadoop-aws module and deoendencies as that is where the core code is. This patch just does the binding to the InsertIntoHadoopFS relation (move to Hadoop MRv2 FileOutputFormat & expect the new superclass, PathOutputCommtter, rather than always a FileOutputcommitter, and for Parquet, something similar with a ParquetOutputCommitter.\r\n\r\n+its only in Hadoop 3.1, though you can backport to branch-2, especially if you are prepared to bump up the minimum java version to 8 in that branch.\r\n\r\nt should work on k8s, given it works standalone. All it needs is an endpoint supporting the multipart upload operation of S3, which includes some non-AWS object stores.\r\n\r\n Look @ the HADOOP-13786 work and the paper [a zero rename committer|https://github.com/steveloughran/zero-rename-committer/releases/download/tag_draft_003/a_zero_rename_committer.pdf]. \r\n\r\nAnd there's some integration tests downstream in https://github.com/hortonworks-spark/cloud-integration . I can help set you up to run those, if you email me directly. Essentially: you need to choose which stores to test against from: s3, openstack, azure, and configure them\r\n\r\nNote that of the two variant committers, \"staging\" and \"magic\", the magic one needs a consistent S3 endpoint, which you only get on AWS S3 with an external services, usually dynamo DB based (S3mper, EMR consisent S3, S3Guard). The staging one needs enough local HDD to buffer the output of all active tasks, but doesn't need that consistency for its own query. You will need a plan for chaining together work though, which is inevitably one of \"consistency layer\" or \"wait long enough between writer and reader that you expect the metadata to be consistent\"\r\n\r\nFinally, if you are using spark to write directly to S3 today, without any consistency layer, then your commit algorithm had better not be mimicing directory rename by list + copy + delete. You need this code for safe as well as performant committing of work to S3.\r\n\r\n","from":"developer"},{"body":"I'm removing the target version, since we are not going to merge it to 2.4","from":"developer"},{"body":"Issue resolved by pull request 24970\n[https://github.com/apache/spark/pull/24970]","from":"developer"},{"body":"[~stevel@apache.org] How to workaround following exception during the execution of \"INSERT OVERWRITE\" in spark.sql (spark 3.1.1 with hadoop 3.2)?\r\n{quote}  if (dynamicPartitionOverwrite) {\r\n\r\n    // until there's explicit extensions to the PathOutputCommitProtocols\r\n\r\n    // to support the spark mechanism, it's left to the individual committer\r\n\r\n    // choice to handle partitioning.\r\n\r\n    throw new IOException(PathOutputCommitProtocol.UNSUPPORTED)\r\n\r\n  }\r\n{quote}","from":"developer"},{"body":"use the partitioned committer and configure it to do the right thing when updating partitions (merge, delete everything already there, fail)","from":"developer"},{"body":"[~stevel@apache.org] Thanks for the info. Below are the related (key, value) we used:\r\n # spark.hadoop.fs.s3a.committer.name — partitioned\r\n # spark.hadoop.mapreduce.outputcommitter.factory.scheme.s3a — org.apache.hadoop.fs.s3a.commit.S3ACommitterFactory\r\n # spark.sql.sources.commitProtocolClass — org.apache.spark.internal.io.cloud.PathOutputCommitProtocol\r\n # spark.sql.parquet.output.committer.class — org.apache.spark.internal.io.cloud.BindingParquetOutputCommitter\r\n\r\n3 & 4 appear to be necessary to ensure S3A committers being used by Spark for parquet outputs, except that \"INSERT OVERWRITE\" is blocked by the dynamicPartitionOverwrite exception. It will be helpful and appreciated if you can patiently elaborate on the proper way to \"use the partitioned committer and configure it to do the right thing ...\" in Spark. For example:\r\n * PathOutputCommitProtocol appears to be needed but its constructor fails with the exception.\r\n * S3A committers honor \"fs.s3a.committer.staging.conflict-mode\" which needs to be \"replace\" for \"INSERT OVERWRITE\" but \"append\" for \"INSERT INTO\". So it is spark.sql query specific. How to make spark.sql automatically set the right value?\r\n * Does above require code change in Spark or there is a configuration-only way?","from":"developer"},{"body":"the spark settings don't make it down from sql; you can do more at the RDD API level.\r\n\r\nThe problem is that the spark partition insert logic all relies on renaming which has the O(data) performance penalty as well as the other commit correctness issues. Yes, something to do the pushdown could be done, or with the multipart APIs of HADOOP-13186 give Spark a standard API to implement a zero rename committer in its own code.\r\n\r\nHowever, focus is on things like Iceberg and Delta lake, which offer more in terms of : \r\n* atomic job commit\r\n* avoid the performance and cost issues of relying on directory tree scan as a way to identify source files.; the performance issues of doing all IO down a single shard of S3 storage.\r\nI do not disagree with the direction of that work; we have to view the S3A committers (and IBM's Stocator + AWS EMR spark committers) as the final attempts to maintain that \"it's just a directory tree\" model into a cloud world where directories don't always exist, and listing them is measurable in hundreds of milliseconds. ","from":"developer"},{"body":"Thank you very much [~stevel@apache.org] for your explanations.\r\n\r\nI am experiencing the same problems as [~danzhi] and your comments helped me a lot.\r\n\r\nSad magic committer does not work with dynamic partition overwrite because it has an amazing performance when writing loads of JSON partitioned data.","from":"developer"},{"body":"[~gumartinm] can I draw your attention to Apache Iceberg?\r\n\r\nmeanwhile\r\nMAPREDUCE-7341 adds a high performance targeting abfs and gcs; all tree scanning is in task commit, which is atomic; job commit aggressively parallelised and optimized for stores whose listStatusIterator calls are incremental with prefetching: we can start processing at the first page of task manifests found in a listing well the second Page is still being retrieved. Also in there: rate limiting, IO Statistics Collection.\r\n\r\nHADOOP-17833 I will pick up some of that work, including incremental loading and rate limiting. And if we can keep reads and writes below the S3 IOPS limits, we should be able to avoid situations where we have to start sleeping and re-trying.\r\n\r\nHADOOP-17981 is my homework this week -emergency work to deal with a rare but current failure in abfs under heavy load.\r\n","from":"developer"},{"body":"Hi all,\r\n\r\n \r\n\r\nI follow the [recommendations|https://spark.apache.org/docs/latest/cloud-integration.html#parquet-io-settings] and getting the following warning:\r\n{code:java}\r\n2022-02-16 15:22:03.292 WARN FlowThread0 ParquetOutputFormat: Setting parquet.enable.summary-metadata is deprecated, please use parquet.summary.metadata.level {code}\r\nWhat would be the recommended value for `parquet.summary.metadata.level ` ? \r\n ","from":"developer"}],"created":"2018-04-13T12:39:58.000+0000","description":"Hadoop 3.1 adds a mechanism for job-specific and store-specific committers (MAPREDUCE-6823, MAPREDUCE-6956), and one key implementation, S3A committers, HADOOP-13786\r\n\r\nThese committers deliver high-performance output of MR and spark jobs to S3, and offer the key semantics which Spark depends on: no visible output until job commit, a failure of a task at an stage, including partway through task commit, can be handled by executing and committing another task attempt. \r\n\r\nIn contrast, the FileOutputFormat commit algorithms on S3 have issues:\r\n\r\n* Awful performance because files are copied by rename\r\n* FileOutputFormat v1: weak task commit failure recovery semantics as the (v1) expectation: \"directory renames are atomic\" doesn't hold.\r\n* S3 metadata eventual consistency can cause rename to miss files or fail entirely (SPARK-15849)\r\n\r\nNote also that FileOutputFormat \"v2\" commit algorithm doesn't offer any of the commit semantics w.r.t observability of or recovery from task commit failure, on any filesystem.\r\n\r\nThe S3A committers address these by way of uploading all data to the destination through multipart uploads, uploads which are only completed in job commit.\r\n\r\nThe new {{PathOutputCommitter}} factory mechanism allows applications to work with the S3A committers and any other, by adding a plugin mechanism into the MRv2 FileOutputFormat class, where it job config and filesystem configuration options can dynamically choose the output committer.\r\n\r\nSpark can use these with some binding classes to \r\n\r\n# Add a subclass of {{HadoopMapReduceCommitProtocol}} which uses the MRv2 classes and {{PathOutputCommitterFactory}} to create the committers.\r\n# Add a {{BindingParquetOutputCommitter extends ParquetOutputCommitter}}\r\nto wire up Parquet output even when code requires the committer to be a subclass of {{ParquetOutputCommitter}}\r\n\r\nThis patch builds on SPARK-23807 for setting up the dependencies.","issue_id":"13152273","key":"SPARK-23977","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-08-15T17:16:11.000+0000","role":"fixed_distractor","summary":"Add commit protocol binding to Hadoop 3.1 PathOutputCommitter mechanism"} {"case_id":"13153439","cluster":"DISTRACTOR-SPARK-24018","comments":[{"body":"Hi, I have got exactly same problem with spark-2.3.1 without hadoop and I have spent nice morning till I have found this issue. It would be fine to fix this issue.","created":"2018-06-29T08:47:53.659+0000"},{"body":"I believe this is limited to spark-shell and was caused by SPARK-18646. Reverting it seems to fix the issue for me.\r\n\r\nI don't know if there is a simple solution that fixes both this and the user classpath issue from that.","created":"2018-07-06T22:26:15.893+0000"},{"body":"I don't think this is only related to the Spark shell: in my case I don't use any user classpath, snappy-1.1.2.6 is just nowhere to be found in Spark's or Hadoop's ClassPath. I first got this issue using spark-submit. The parquet lib version provided by Spark is incompatible with snappy-1.0.4.1 found in Hadoop's classpath.\r\n\r\n[~pclay] did you start spark-shell with any arguments? Maybe snappy is shaded in one of the JARs you have in your classpath?","created":"2018-07-07T18:06:22.956+0000"},{"body":"I believe we are both partially correct in that a fix (with Spark 2.3.0) does require snappy-java-1.1.2, and it was caused by SPARK-18646. The native library loader of Snappy 1.0.4 [uses a self-described hack|https://github.com/xerial/snappy-java/blob/snappy-java-1.0.4/src/main/java/org/xerial/snappy/SnappyLoader.java#L175] to inject the loader onto the root class loader. The hack was [later removed|https://github.com/xerial/snappy-java/commit/06f007a08#diff-a1c8fc77f8] in 1.1.2, which allows the non-inheriting class loader to pick it up.\r\n\r\n \r\n\r\nI believed this only affects spark-shell, because neither pyspark (the REPL and spark-submit) nor\r\n{code:java}\r\n./bin/spark-submit --class org.apache.spark.examples.sql.SQLDataSourceExample examples/jars/spark-examples_2.11-2.3.0.jar{code}\r\nhave this issue. What repro did you have without spark-shell?\r\n\r\n \r\n\r\nI don't believe this is related to Parquet versioning because this also throws:\r\n{code:java}\r\nscala> import org.xerial.snappy.Snappy \r\nimport org.xerial.snappy.Snappy \r\n\r\nscala> sc.parallelize(Seq(\"foo\")).map(Snappy.compress).collect \r\n2018-07-09 13:44:14 ERROR Executor:91 - Exception in task 11.0 in stage 0.0 (TID 11) \r\njava.lang.UnsatisfiedLinkError: org.xerial.snappy.SnappyNative.maxCompressedLength(I)I\r\n...{code}\r\n \r\n\r\nIn answer to your last question I did not pass any arguments to spark-shell. All I did to repro was\r\n{code:java}\r\nexport SPARK_DIST_CLASSPATH=$(~/Downloads/hadoop-2.8.3/bin/hadoop classpath)\r\n~/Downloads/spark-2.3.0-bin-without-hadoop/bin/spark-shell{code}\r\n \r\n\r\n ","created":"2018-07-09T20:49:29.624+0000"},{"body":"Oh, indeed you are right! I was mistaken when I thought that I had the error using spark-submit, I just verified again and it works, thanks for the explanation!","created":"2018-07-10T00:03:11.602+0000"},{"body":"It may be fixed by [SPARK-24927|https://issues.apache.org/jira/browse/SPARK-24927].\r\n","created":"2018-09-22T11:51:30.006+0000"},{"body":"I confirm the fix that appeared in Spark 2.3.2","created":"2018-12-10T16:44:26.332+0000"}],"conversations":[{"body":"On a brand-new installation of Spark 2.3.0 with a user-provided hadoop-2.8.3, Spark fails to read or write dataframes in parquet format with snappy compression.\r\n\r\nThis is due to an incompatibility between the snappy-java version that is required by parquet (parquet is provided in Spark jars but snappy isn't) and the version that is available from hadoop-2.8.3.\r\n\r\n \r\n\r\nSteps to reproduce:\r\n * Download and extract hadoop-2.8.3\r\n * Download and extract spark-2.3.0-without-hadoop\r\n * export JAVA_HOME, HADOOP_HOME, SPARK_HOME, PATH\r\n * Following instructions from [https://spark.apache.org/docs/latest/hadoop-provided.html], set SPARK_DIST_CLASSPATH=$(hadoop classpath) in spark-env.sh\r\n * Start a spark-shell, enter the following:\r\n\r\n \r\n{code:java}\r\nimport spark.implicits._\r\nval df = List(1, 2, 3, 4).toDF\r\ndf.write\r\n .format(\"parquet\")\r\n .option(\"compression\", \"snappy\")\r\n .mode(\"overwrite\")\r\n .save(\"test.parquet\")\r\n{code}\r\n \r\n\r\n \r\n\r\nThis fails with the following:\r\n{noformat}\r\njava.lang.UnsatisfiedLinkError: org.xerial.snappy.SnappyNative.maxCompressedLength(I)I\r\n at org.xerial.snappy.SnappyNative.maxCompressedLength(Native Method)\r\n at org.xerial.snappy.Snappy.maxCompressedLength(Snappy.java:316)\r\n at org.apache.parquet.hadoop.codec.SnappyCompressor.compress(SnappyCompressor.java:67)\r\n at org.apache.hadoop.io.compress.CompressorStream.compress(CompressorStream.java:81)\r\n at org.apache.hadoop.io.compress.CompressorStream.finish(CompressorStream.java:92)\r\n at org.apache.parquet.hadoop.CodecFactory$BytesCompressor.compress(CodecFactory.java:112)\r\n at org.apache.parquet.hadoop.ColumnChunkPageWriteStore$ColumnChunkPageWriter.writePage(ColumnChunkPageWriteStore.java:93)\r\n at org.apache.parquet.column.impl.ColumnWriterV1.writePage(ColumnWriterV1.java:150)\r\n at org.apache.parquet.column.impl.ColumnWriterV1.flush(ColumnWriterV1.java:238)\r\n at org.apache.parquet.column.impl.ColumnWriteStoreV1.flush(ColumnWriteStoreV1.java:121)\r\n at org.apache.parquet.hadoop.InternalParquetRecordWriter.flushRowGroupToStore(InternalParquetRecordWriter.java:167)\r\n at org.apache.parquet.hadoop.InternalParquetRecordWriter.close(InternalParquetRecordWriter.java:109)\r\n at org.apache.parquet.hadoop.ParquetRecordWriter.close(ParquetRecordWriter.java:163)\r\n at org.apache.spark.sql.execution.datasources.parquet.ParquetOutputWriter.close(ParquetOutputWriter.scala:42)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$SingleDirectoryWriteTask.releaseResources(FileFormatWriter.scala:405)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$SingleDirectoryWriteTask.execute(FileFormatWriter.scala:396)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:269)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:267)\r\n at org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1411)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:272)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:197)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:196)\r\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87) at org.apache.spark.scheduler.Task.run(Task.scala:109)\r\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n at java.lang.Thread.run(Thread.java:748){noformat}\r\n \r\n\r\n  Downloading snappy-java-1.1.2.6.jar and placing it in Sparks's jar folder solves the issue.","from":"reporter","subject":"Spark-without-hadoop package fails to create or read parquet files with snappy compression"},{"body":"Hi, I have got exactly same problem with spark-2.3.1 without hadoop and I have spent nice morning till I have found this issue. It would be fine to fix this issue.","from":"developer"},{"body":"I believe this is limited to spark-shell and was caused by SPARK-18646. Reverting it seems to fix the issue for me.\r\n\r\nI don't know if there is a simple solution that fixes both this and the user classpath issue from that.","from":"developer"},{"body":"I don't think this is only related to the Spark shell: in my case I don't use any user classpath, snappy-1.1.2.6 is just nowhere to be found in Spark's or Hadoop's ClassPath. I first got this issue using spark-submit. The parquet lib version provided by Spark is incompatible with snappy-1.0.4.1 found in Hadoop's classpath.\r\n\r\n[~pclay] did you start spark-shell with any arguments? Maybe snappy is shaded in one of the JARs you have in your classpath?","from":"developer"},{"body":"I believe we are both partially correct in that a fix (with Spark 2.3.0) does require snappy-java-1.1.2, and it was caused by SPARK-18646. The native library loader of Snappy 1.0.4 [uses a self-described hack|https://github.com/xerial/snappy-java/blob/snappy-java-1.0.4/src/main/java/org/xerial/snappy/SnappyLoader.java#L175] to inject the loader onto the root class loader. The hack was [later removed|https://github.com/xerial/snappy-java/commit/06f007a08#diff-a1c8fc77f8] in 1.1.2, which allows the non-inheriting class loader to pick it up.\r\n\r\n \r\n\r\nI believed this only affects spark-shell, because neither pyspark (the REPL and spark-submit) nor\r\n{code:java}\r\n./bin/spark-submit --class org.apache.spark.examples.sql.SQLDataSourceExample examples/jars/spark-examples_2.11-2.3.0.jar{code}\r\nhave this issue. What repro did you have without spark-shell?\r\n\r\n \r\n\r\nI don't believe this is related to Parquet versioning because this also throws:\r\n{code:java}\r\nscala> import org.xerial.snappy.Snappy \r\nimport org.xerial.snappy.Snappy \r\n\r\nscala> sc.parallelize(Seq(\"foo\")).map(Snappy.compress).collect \r\n2018-07-09 13:44:14 ERROR Executor:91 - Exception in task 11.0 in stage 0.0 (TID 11) \r\njava.lang.UnsatisfiedLinkError: org.xerial.snappy.SnappyNative.maxCompressedLength(I)I\r\n...{code}\r\n \r\n\r\nIn answer to your last question I did not pass any arguments to spark-shell. All I did to repro was\r\n{code:java}\r\nexport SPARK_DIST_CLASSPATH=$(~/Downloads/hadoop-2.8.3/bin/hadoop classpath)\r\n~/Downloads/spark-2.3.0-bin-without-hadoop/bin/spark-shell{code}\r\n \r\n\r\n ","from":"developer"},{"body":"Oh, indeed you are right! I was mistaken when I thought that I had the error using spark-submit, I just verified again and it works, thanks for the explanation!","from":"developer"},{"body":"It may be fixed by [SPARK-24927|https://issues.apache.org/jira/browse/SPARK-24927].\r\n","from":"developer"},{"body":"I confirm the fix that appeared in Spark 2.3.2","from":"developer"}],"created":"2018-04-18T18:45:53.000+0000","description":"On a brand-new installation of Spark 2.3.0 with a user-provided hadoop-2.8.3, Spark fails to read or write dataframes in parquet format with snappy compression.\r\n\r\nThis is due to an incompatibility between the snappy-java version that is required by parquet (parquet is provided in Spark jars but snappy isn't) and the version that is available from hadoop-2.8.3.\r\n\r\n \r\n\r\nSteps to reproduce:\r\n * Download and extract hadoop-2.8.3\r\n * Download and extract spark-2.3.0-without-hadoop\r\n * export JAVA_HOME, HADOOP_HOME, SPARK_HOME, PATH\r\n * Following instructions from [https://spark.apache.org/docs/latest/hadoop-provided.html], set SPARK_DIST_CLASSPATH=$(hadoop classpath) in spark-env.sh\r\n * Start a spark-shell, enter the following:\r\n\r\n \r\n{code:java}\r\nimport spark.implicits._\r\nval df = List(1, 2, 3, 4).toDF\r\ndf.write\r\n .format(\"parquet\")\r\n .option(\"compression\", \"snappy\")\r\n .mode(\"overwrite\")\r\n .save(\"test.parquet\")\r\n{code}\r\n \r\n\r\n \r\n\r\nThis fails with the following:\r\n{noformat}\r\njava.lang.UnsatisfiedLinkError: org.xerial.snappy.SnappyNative.maxCompressedLength(I)I\r\n at org.xerial.snappy.SnappyNative.maxCompressedLength(Native Method)\r\n at org.xerial.snappy.Snappy.maxCompressedLength(Snappy.java:316)\r\n at org.apache.parquet.hadoop.codec.SnappyCompressor.compress(SnappyCompressor.java:67)\r\n at org.apache.hadoop.io.compress.CompressorStream.compress(CompressorStream.java:81)\r\n at org.apache.hadoop.io.compress.CompressorStream.finish(CompressorStream.java:92)\r\n at org.apache.parquet.hadoop.CodecFactory$BytesCompressor.compress(CodecFactory.java:112)\r\n at org.apache.parquet.hadoop.ColumnChunkPageWriteStore$ColumnChunkPageWriter.writePage(ColumnChunkPageWriteStore.java:93)\r\n at org.apache.parquet.column.impl.ColumnWriterV1.writePage(ColumnWriterV1.java:150)\r\n at org.apache.parquet.column.impl.ColumnWriterV1.flush(ColumnWriterV1.java:238)\r\n at org.apache.parquet.column.impl.ColumnWriteStoreV1.flush(ColumnWriteStoreV1.java:121)\r\n at org.apache.parquet.hadoop.InternalParquetRecordWriter.flushRowGroupToStore(InternalParquetRecordWriter.java:167)\r\n at org.apache.parquet.hadoop.InternalParquetRecordWriter.close(InternalParquetRecordWriter.java:109)\r\n at org.apache.parquet.hadoop.ParquetRecordWriter.close(ParquetRecordWriter.java:163)\r\n at org.apache.spark.sql.execution.datasources.parquet.ParquetOutputWriter.close(ParquetOutputWriter.scala:42)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$SingleDirectoryWriteTask.releaseResources(FileFormatWriter.scala:405)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$SingleDirectoryWriteTask.execute(FileFormatWriter.scala:396)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:269)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:267)\r\n at org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1411)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:272)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:197)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:196)\r\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87) at org.apache.spark.scheduler.Task.run(Task.scala:109)\r\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n at java.lang.Thread.run(Thread.java:748){noformat}\r\n \r\n\r\n  Downloading snappy-java-1.1.2.6.jar and placing it in Sparks's jar folder solves the issue.","issue_id":"13153439","key":"SPARK-24018","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-12-10T16:44:26.000+0000","role":"fixed_distractor","summary":"Spark-without-hadoop package fails to create or read parquet files with snappy compression"} {"case_id":"13161595","cluster":"DISTRACTOR-SPARK-24373","comments":[{"body":"We found after upgrading to Spark 2.3, many of our production systems runs slower. This is founded to be caused by \"df.cache(); df.count()\" no longer eagerly cache df correctly. I think this might be a regression from 2.2. \r\n\r\nI think any one uses  \"df.cache() df.count()\" to cache data eagerly will be affected.","created":"2018-05-23T22:13:18.533+0000"},{"body":"This is a reproduce:\r\n{code:java}\r\nval myUDF = udf((x: Long) => { println(\"xxxx\"); x + 1 })\r\n\r\nval df1 = spark.range(0, 1).toDF(\"s\").select(myUDF($\"s\"))\r\ndf1.cache()\r\ndf1.count()\r\n// No xxxx printed\r\n{code}\r\nIt appears the issue is related to UDF:\r\n{code:java}\r\nval df1 = spark.range(0, 1).toDF(\"s\").select(myUDF($\"s\"))\r\ndf1.cache()\r\ndf1.groupBy().count().explain()\r\n\r\n== Physical Plan ==\r\n*(2) HashAggregate(keys=[], functions=[count(1)])\r\n+- Exchange SinglePartition\r\n +- *(1) HashAggregate(keys=[], functions=[partial_count(1)])\r\n +- *(1) Project\r\n +- *(1) Range (0, 1, step=1, splits=2)\r\n{code}\r\nWithout UDF it uses \"count\" materialize cache:\r\n{code:java}\r\nval df1 = spark.range(0, 1).toDF(\"s\").select($\"s\" + 1)\r\ndf1.cache()\r\ndf1.groupBy().count().explain()\r\n\r\n== Physical Plan ==\r\n*(2) HashAggregate(keys=[], functions=[count(1)])\r\n +- Exchange SinglePartition\r\n +- *(1) HashAggregate(keys=[], functions=[partial_count(1)])\r\n +- *(1) InMemoryTableScan\r\n +- InMemoryRelation [(s + 1)#179L], CachedRDDBuilder(true,10000,StorageLevel(disk, memory, deserialized, 1 replicas),*(1) Project [(id#175L + 1) AS (s + 1)#179L]\r\n +- *(1) Range (0, 1, step=1, splits=2) ,None)\r\n +- *(1) Project [(id#175L + 1) AS (s + 1)#179L]\r\n +- *(1) Range (0, 1, step=1, splits=2)\r\n\r\n{code}\r\n ","created":"2018-05-24T14:05:00.562+0000"},{"body":"We are also facing increased runtime duration for our SQL jobs (after upgrading from 2.2.1 to 2.3.0), but didn't trace it down to the root cause. This issue sounds reasonable to me, as we are also using cache() + count() quite often.","created":"2018-05-24T14:46:09.068+0000"},{"body":"I turned on the log trace of RuleExecutor and found that in my example of df1.count() after cache. \r\n\r\n \r\n{code:java}\r\nscala> df1.groupBy().count().explain(true)\r\n\r\n=== Applying Rule org.apache.spark.sql.catalyst.analysis.Analyzer$HandleNullInputsForUDF ===\r\nAggregate [count(1) AS count#40L] Aggregate [count(1) AS count#40L]\r\n!+- Project [value#2L, if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L] +- Project [value#2L, if (isnull(value#2L)) null else if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L] +- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L] +- ExternalRDD [obj#1L]\r\n{code}\r\n that is node\r\n{code:java}\r\nProject [value#2L, if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]{code}\r\nbecomes \r\n{code:java}\r\nProject [value#2L, if (isnull(value#2L)) null else if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]{code}\r\n \r\n\r\nThis will cause a miss in the CacheManager?\r\n\r\nwhich could be confirmed by later applying ColumnPrunning rule's log trace.  \r\n\r\nMay question is: is that supposed protected by AnalysisBarrier ?","created":"2018-05-24T15:37:26.255+0000"},{"body":"This could be the same as SPARK-23309.","created":"2018-05-24T16:57:16.322+0000"},{"body":"It is not apparently to me that they are the same issue though","created":"2018-05-24T18:23:30.627+0000"},{"body":"I guess we should use `planWithBarrier` in the 'RelationalGroupedDataset' or other similar places. Any suggestion?","created":"2018-05-24T21:17:11.798+0000"},{"body":"[~wbzhao] yes, I do agree with you. That is the problem.","created":"2018-05-25T10:51:19.474+0000"},{"body":"User 'mgaido91' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21432","created":"2018-05-25T14:12:07.291+0000"},{"body":"[~icexelloss] [~aweise] Are you also using the Dataset APIs groupBy(), rollup(), cube(), rollup, pivot() and groupByKey()?","created":"2018-05-25T17:54:51.903+0000"},{"body":"We use groupby() and pivot()","created":"2018-05-25T17:56:08.890+0000"},{"body":"BTW, I plan to continue my work of https://github.com/apache/spark/pull/18717, which will add an eager persist/cache API. ","created":"2018-05-25T17:59:36.623+0000"},{"body":"[~smilegator] I think an eager API is not related to the problem experienced here, though.","created":"2018-05-25T18:24:11.035+0000"},{"body":"[~LI,Xiao] That is a good idea :) Eager caching is useful, many times I see additional count just to cache eagerly","created":"2018-05-25T18:52:50.024+0000"},{"body":"{code}\r\n def count(): Long = withAction(\"count\", groupBy().count().queryExecution) { plan =>\r\n plan.executeCollect().head.getLong(0)\r\n }\r\n{code}\r\n\r\nMany Spark users are using df.count() after df.cache() for achieving eager caching. Since our count() API is using `groupBy()`, the impact becomes much bigger. The count() API will not trigger the data materialization when the plans are different after multiple rounds of plan analysis. ","created":"2018-05-25T20:18:51.784+0000"},{"body":"[~smilegator] yes, you're right, the impact would be definitely lower.","created":"2018-05-25T21:09:46.814+0000"},{"body":"In the above example, each time when we re-analyze the plan that is recreated through the Dataset APIs count(), groupBy(), rollup(), cube(), rollup, pivot() and groupByKey(), the Analyzer rule HandleNullInputsForUDF will add the extra IF expression above the UDF in the previously resolved sub-plan. Note, this is not the only rule that could change the analyzed plans if we re-run the analyzer.\r\n\r\nThis is a regression introduced by [https://github.com/apache/spark/pull/17770]. We replaced the original solution (based on the analyzed flag) by the AnalysisBarrier. However, we did not add the AnalysisBarrier on the APIs of RelationalGroupedDataset and KeyValueGroupedDataset.\r\n\r\nTo fix it, we will changes the plan again. We might face some unknown issues. How about adding a temporary flag in Spark 2.3.1? If anything unexpected happens, our users still can change it back to the Spark 2.3.0 behavior?","created":"2018-05-25T21:50:15.846+0000"},{"body":"[~smilegator] do you mean that add AnalysisBarrier to RelationalGroupedDataset and KeyValueGroupedDataset could lead to new bugs?","created":"2018-05-25T21:54:15.499+0000"},{"body":"[~wbzhao] as I answered on the PR, the fix is complete and includes also {{flatMapGroupsInPandas}}.","created":"2018-05-29T13:53:28.469+0000"},{"body":"[~mgaido] Thanks. I didn't look the comment carefully. ","created":"2018-05-29T14:18:29.156+0000"},{"body":"[~icexelloss] This is still possible since the query plans are changed. I am also fine to do it without a flag. If you apply the fix to your internal fork, I would suggest to add a flag. At least, you can turn it off when anything unexpected happens. ","created":"2018-05-30T21:15:38.172+0000"},{"body":"[~smilegator] Thank you for the suggestion.","created":"2018-05-30T21:19:08.168+0000"}],"conversations":[{"body":"Here is the code to reproduce in local mode\r\n{code:java}\r\nscala> val df = sc.range(1, 2).toDF\r\ndf: org.apache.spark.sql.DataFrame = [value: bigint]\r\n\r\nscala> val myudf = udf({x: Long => println(\"xxxx\"); x + 1})\r\nmyudf: org.apache.spark.sql.expressions.UserDefinedFunction = UserDefinedFunction(,LongType,Some(List(LongType)))\r\n\r\nscala> val df1 = df.withColumn(\"value1\", myudf(col(\"value\")))\r\ndf1: org.apache.spark.sql.DataFrame = [value: bigint, value1: bigint]\r\n\r\nscala> df1.cache\r\nres0: df1.type = [value: bigint, value1: bigint]\r\n\r\nscala> df1.count\r\nres1: Long = 1 \r\n\r\nscala> df1.count\r\nres2: Long = 1\r\n\r\nscala> df1.count\r\nres3: Long = 1\r\n{code}\r\n \r\n\r\nin Spark 2.2, you could see it prints \"xxxx\". \r\n\r\nIn the above example, when you do explain. You could see\r\n{code:java}\r\nscala> df1.explain(true)\r\n== Parsed Logical Plan ==\r\n'Project [value#2L, UDF('value) AS value1#5]\r\n+- AnalysisBarrier\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L]\r\n\r\n== Analyzed Logical Plan ==\r\nvalue: bigint, value1: bigint\r\nProject [value#2L, if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L]\r\n\r\n== Optimized Logical Plan ==\r\nInMemoryRelation [value#2L, value1#5L], true, 10000, StorageLevel(disk, memory, deserialized, 1 replicas)\r\n+- *(1) Project [value#2L, UDF(value#2L) AS value1#5L]\r\n+- *(1) SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- Scan ExternalRDDScan[obj#1L]\r\n\r\n== Physical Plan ==\r\n*(1) InMemoryTableScan [value#2L, value1#5L]\r\n+- InMemoryRelation [value#2L, value1#5L], true, 10000, StorageLevel(disk, memory, deserialized, 1 replicas)\r\n+- *(1) Project [value#2L, UDF(value#2L) AS value1#5L]\r\n+- *(1) SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- Scan ExternalRDDScan[obj#1L]\r\n\r\n{code}\r\nbut the ImMemoryTableScan is mising in the following explain()\r\n{code:java}\r\nscala> df1.groupBy().count().explain(true)\r\n== Parsed Logical Plan ==\r\nAggregate [count(1) AS count#170L]\r\n+- Project [value#2L, if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L]\r\n\r\n== Analyzed Logical Plan ==\r\ncount: bigint\r\nAggregate [count(1) AS count#170L]\r\n+- Project [value#2L, if (isnull(value#2L)) null else if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L]\r\n\r\n== Optimized Logical Plan ==\r\nAggregate [count(1) AS count#170L]\r\n+- Project\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L]\r\n\r\n== Physical Plan ==\r\n*(2) HashAggregate(keys=[], functions=[count(1)], output=[count#170L])\r\n+- Exchange SinglePartition\r\n+- *(1) HashAggregate(keys=[], functions=[partial_count(1)], output=[count#175L])\r\n+- *(1) Project\r\n+- *(1) SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- Scan ExternalRDDScan[obj#1L]\r\n{code}\r\n \r\n\r\n ","from":"reporter","subject":"\"df.cache() df.count()\" no longer eagerly caches data when the analyzed plans are different after re-analyzing the plans"},{"body":"We found after upgrading to Spark 2.3, many of our production systems runs slower. This is founded to be caused by \"df.cache(); df.count()\" no longer eagerly cache df correctly. I think this might be a regression from 2.2. \r\n\r\nI think any one uses  \"df.cache() df.count()\" to cache data eagerly will be affected.","from":"developer"},{"body":"This is a reproduce:\r\n{code:java}\r\nval myUDF = udf((x: Long) => { println(\"xxxx\"); x + 1 })\r\n\r\nval df1 = spark.range(0, 1).toDF(\"s\").select(myUDF($\"s\"))\r\ndf1.cache()\r\ndf1.count()\r\n// No xxxx printed\r\n{code}\r\nIt appears the issue is related to UDF:\r\n{code:java}\r\nval df1 = spark.range(0, 1).toDF(\"s\").select(myUDF($\"s\"))\r\ndf1.cache()\r\ndf1.groupBy().count().explain()\r\n\r\n== Physical Plan ==\r\n*(2) HashAggregate(keys=[], functions=[count(1)])\r\n+- Exchange SinglePartition\r\n +- *(1) HashAggregate(keys=[], functions=[partial_count(1)])\r\n +- *(1) Project\r\n +- *(1) Range (0, 1, step=1, splits=2)\r\n{code}\r\nWithout UDF it uses \"count\" materialize cache:\r\n{code:java}\r\nval df1 = spark.range(0, 1).toDF(\"s\").select($\"s\" + 1)\r\ndf1.cache()\r\ndf1.groupBy().count().explain()\r\n\r\n== Physical Plan ==\r\n*(2) HashAggregate(keys=[], functions=[count(1)])\r\n +- Exchange SinglePartition\r\n +- *(1) HashAggregate(keys=[], functions=[partial_count(1)])\r\n +- *(1) InMemoryTableScan\r\n +- InMemoryRelation [(s + 1)#179L], CachedRDDBuilder(true,10000,StorageLevel(disk, memory, deserialized, 1 replicas),*(1) Project [(id#175L + 1) AS (s + 1)#179L]\r\n +- *(1) Range (0, 1, step=1, splits=2) ,None)\r\n +- *(1) Project [(id#175L + 1) AS (s + 1)#179L]\r\n +- *(1) Range (0, 1, step=1, splits=2)\r\n\r\n{code}\r\n ","from":"developer"},{"body":"We are also facing increased runtime duration for our SQL jobs (after upgrading from 2.2.1 to 2.3.0), but didn't trace it down to the root cause. This issue sounds reasonable to me, as we are also using cache() + count() quite often.","from":"developer"},{"body":"I turned on the log trace of RuleExecutor and found that in my example of df1.count() after cache. \r\n\r\n \r\n{code:java}\r\nscala> df1.groupBy().count().explain(true)\r\n\r\n=== Applying Rule org.apache.spark.sql.catalyst.analysis.Analyzer$HandleNullInputsForUDF ===\r\nAggregate [count(1) AS count#40L] Aggregate [count(1) AS count#40L]\r\n!+- Project [value#2L, if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L] +- Project [value#2L, if (isnull(value#2L)) null else if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L] +- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L] +- ExternalRDD [obj#1L]\r\n{code}\r\n that is node\r\n{code:java}\r\nProject [value#2L, if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]{code}\r\nbecomes \r\n{code:java}\r\nProject [value#2L, if (isnull(value#2L)) null else if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]{code}\r\n \r\n\r\nThis will cause a miss in the CacheManager?\r\n\r\nwhich could be confirmed by later applying ColumnPrunning rule's log trace.  \r\n\r\nMay question is: is that supposed protected by AnalysisBarrier ?","from":"developer"},{"body":"This could be the same as SPARK-23309.","from":"developer"},{"body":"It is not apparently to me that they are the same issue though","from":"developer"},{"body":"I guess we should use `planWithBarrier` in the 'RelationalGroupedDataset' or other similar places. Any suggestion?","from":"developer"},{"body":"[~wbzhao] yes, I do agree with you. That is the problem.","from":"developer"},{"body":"User 'mgaido91' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21432","from":"developer"},{"body":"[~icexelloss] [~aweise] Are you also using the Dataset APIs groupBy(), rollup(), cube(), rollup, pivot() and groupByKey()?","from":"developer"},{"body":"We use groupby() and pivot()","from":"developer"},{"body":"BTW, I plan to continue my work of https://github.com/apache/spark/pull/18717, which will add an eager persist/cache API. ","from":"developer"},{"body":"[~smilegator] I think an eager API is not related to the problem experienced here, though.","from":"developer"},{"body":"[~LI,Xiao] That is a good idea :) Eager caching is useful, many times I see additional count just to cache eagerly","from":"developer"},{"body":"{code}\r\n def count(): Long = withAction(\"count\", groupBy().count().queryExecution) { plan =>\r\n plan.executeCollect().head.getLong(0)\r\n }\r\n{code}\r\n\r\nMany Spark users are using df.count() after df.cache() for achieving eager caching. Since our count() API is using `groupBy()`, the impact becomes much bigger. The count() API will not trigger the data materialization when the plans are different after multiple rounds of plan analysis. ","from":"developer"},{"body":"[~smilegator] yes, you're right, the impact would be definitely lower.","from":"developer"},{"body":"In the above example, each time when we re-analyze the plan that is recreated through the Dataset APIs count(), groupBy(), rollup(), cube(), rollup, pivot() and groupByKey(), the Analyzer rule HandleNullInputsForUDF will add the extra IF expression above the UDF in the previously resolved sub-plan. Note, this is not the only rule that could change the analyzed plans if we re-run the analyzer.\r\n\r\nThis is a regression introduced by [https://github.com/apache/spark/pull/17770]. We replaced the original solution (based on the analyzed flag) by the AnalysisBarrier. However, we did not add the AnalysisBarrier on the APIs of RelationalGroupedDataset and KeyValueGroupedDataset.\r\n\r\nTo fix it, we will changes the plan again. We might face some unknown issues. How about adding a temporary flag in Spark 2.3.1? If anything unexpected happens, our users still can change it back to the Spark 2.3.0 behavior?","from":"developer"},{"body":"[~smilegator] do you mean that add AnalysisBarrier to RelationalGroupedDataset and KeyValueGroupedDataset could lead to new bugs?","from":"developer"},{"body":"[~wbzhao] as I answered on the PR, the fix is complete and includes also {{flatMapGroupsInPandas}}.","from":"developer"},{"body":"[~mgaido] Thanks. I didn't look the comment carefully. ","from":"developer"},{"body":"[~icexelloss] This is still possible since the query plans are changed. I am also fine to do it without a flag. If you apply the fix to your internal fork, I would suggest to add a flag. At least, you can turn it off when anything unexpected happens. ","from":"developer"},{"body":"[~smilegator] Thank you for the suggestion.","from":"developer"}],"created":"2018-05-23T21:44:46.000+0000","description":"Here is the code to reproduce in local mode\r\n{code:java}\r\nscala> val df = sc.range(1, 2).toDF\r\ndf: org.apache.spark.sql.DataFrame = [value: bigint]\r\n\r\nscala> val myudf = udf({x: Long => println(\"xxxx\"); x + 1})\r\nmyudf: org.apache.spark.sql.expressions.UserDefinedFunction = UserDefinedFunction(,LongType,Some(List(LongType)))\r\n\r\nscala> val df1 = df.withColumn(\"value1\", myudf(col(\"value\")))\r\ndf1: org.apache.spark.sql.DataFrame = [value: bigint, value1: bigint]\r\n\r\nscala> df1.cache\r\nres0: df1.type = [value: bigint, value1: bigint]\r\n\r\nscala> df1.count\r\nres1: Long = 1 \r\n\r\nscala> df1.count\r\nres2: Long = 1\r\n\r\nscala> df1.count\r\nres3: Long = 1\r\n{code}\r\n \r\n\r\nin Spark 2.2, you could see it prints \"xxxx\". \r\n\r\nIn the above example, when you do explain. You could see\r\n{code:java}\r\nscala> df1.explain(true)\r\n== Parsed Logical Plan ==\r\n'Project [value#2L, UDF('value) AS value1#5]\r\n+- AnalysisBarrier\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L]\r\n\r\n== Analyzed Logical Plan ==\r\nvalue: bigint, value1: bigint\r\nProject [value#2L, if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L]\r\n\r\n== Optimized Logical Plan ==\r\nInMemoryRelation [value#2L, value1#5L], true, 10000, StorageLevel(disk, memory, deserialized, 1 replicas)\r\n+- *(1) Project [value#2L, UDF(value#2L) AS value1#5L]\r\n+- *(1) SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- Scan ExternalRDDScan[obj#1L]\r\n\r\n== Physical Plan ==\r\n*(1) InMemoryTableScan [value#2L, value1#5L]\r\n+- InMemoryRelation [value#2L, value1#5L], true, 10000, StorageLevel(disk, memory, deserialized, 1 replicas)\r\n+- *(1) Project [value#2L, UDF(value#2L) AS value1#5L]\r\n+- *(1) SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- Scan ExternalRDDScan[obj#1L]\r\n\r\n{code}\r\nbut the ImMemoryTableScan is mising in the following explain()\r\n{code:java}\r\nscala> df1.groupBy().count().explain(true)\r\n== Parsed Logical Plan ==\r\nAggregate [count(1) AS count#170L]\r\n+- Project [value#2L, if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L]\r\n\r\n== Analyzed Logical Plan ==\r\ncount: bigint\r\nAggregate [count(1) AS count#170L]\r\n+- Project [value#2L, if (isnull(value#2L)) null else if (isnull(value#2L)) null else UDF(value#2L) AS value1#5L]\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L]\r\n\r\n== Optimized Logical Plan ==\r\nAggregate [count(1) AS count#170L]\r\n+- Project\r\n+- SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- ExternalRDD [obj#1L]\r\n\r\n== Physical Plan ==\r\n*(2) HashAggregate(keys=[], functions=[count(1)], output=[count#170L])\r\n+- Exchange SinglePartition\r\n+- *(1) HashAggregate(keys=[], functions=[partial_count(1)], output=[count#175L])\r\n+- *(1) Project\r\n+- *(1) SerializeFromObject [input[0, bigint, false] AS value#2L]\r\n+- Scan ExternalRDDScan[obj#1L]\r\n{code}\r\n \r\n\r\n ","issue_id":"13161595","key":"SPARK-24373","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-05-28T05:57:49.000+0000","role":"fixed_distractor","summary":"\"df.cache() df.count()\" no longer eagerly caches data when the analyzed plans are different after re-analyzing the plans"} {"case_id":"13162332","cluster":"DISTRACTOR-SPARK-24401","comments":[{"body":"Probably is worth to mention that my dataset is coming from a oracle DB","created":"2018-05-28T06:40:17.646+0000"},{"body":"I followed the repro steps using the file you attached and I was unable to reproduce on current master. May you please check the repro code and try it on current master branch?","created":"2018-05-28T15:19:03.526+0000"},{"body":"I didn't narrow down and I'm not sure that this is related to this issue though, I found v2.2.0 had an incorrect result (fixed in v2.2.1);\r\nAs marco said, I'd be greated if you try on master;\r\n{code}\r\n// v2.2.0\r\nscala> second_agg.show\r\n+----+-----+----------------+-----+ \r\n| id1| id2| maxf| minf|\r\n+----+-----+----------------+-----+\r\n|1498|88586|0.00238636363635|3E-14|\r\n+----+-----+----------------+-----+\r\n\r\n// v2.2.1\r\nscala> second_agg.show\r\n+----+-----+----------------+----------------+ \r\n| id1| id2| maxf| minf|\r\n+----+-----+----------------+----------------+\r\n|1498|88586|3.95833333330000|2.00000000000000|\r\n+----+-----+----------------+----------------+\r\n\r\n// v2.3.0\r\nscala> second_agg.show\r\n+----+-----+----------------+----------------+\r\n| id1| id2| maxf| minf|\r\n+----+-----+----------------+----------------+\r\n|1498|88586|3.95833333330000|2.00000000000000|\r\n+----+-----+----------------+----------------+\r\n\r\n// master\r\nscala> second_agg.show\r\n+----+-----+----------------+----------------+\r\n| id1| id2| maxf| minf|\r\n+----+-----+----------------+----------------+\r\n|1498|88586|3.95833333330000|2.00000000000000|\r\n+----+-----+----------------+----------------+\r\n{code}","created":"2018-05-29T02:16:20.528+0000"},{"body":"Yes, I just tried my code and it works on spark 2.3.0","created":"2018-05-29T05:48:24.907+0000"},{"body":"Was already Fixed on 2.3.0","created":"2018-05-29T05:49:58.147+0000"}],"conversations":[{"body":"Hi, \r\n\r\nI think I found a really ugly bug in spark when performing aggregations with Decimals\r\n\r\nTo reproduce: \r\n\r\n \r\n{code:java}\r\nval df = spark.read.parquet(\"attached file\")\r\nval first_agg = fact_df.groupBy(\"id1\", \"id2\", \"start_date\").agg(mean(\"projection_factor\").alias(\"projection_factor\"))\r\nfirst_agg.show\r\nval second_agg = first_agg.groupBy(\"id1\",\"id2\").agg(max(\"projection_factor\").alias(\"maxf\"), min(\"projection_factor\").alias(\"minf\"))\r\nsecond_agg.show\r\n{code}\r\nFirst aggregation works fine the second aggregation seems to be summing instead of max value. I tried with spark 2.2.0 and 2.3.0 same problem.\r\n\r\nThe dataset as circa 800 Rows and the projection_factor has values from 0 until 100. the result should not be bigger that 5 but with get 265820543091454.... as result back.\r\n\r\n \r\n\r\n \r\n\r\nAs Code not 100% the same but I think there is really a bug there: \r\n\r\n \r\n{code:java}\r\nBigDecimal [] objects = new BigDecimal[]{\r\n new BigDecimal(3.5714285714D),\r\n new BigDecimal(3.5714285714D),\r\n new BigDecimal(3.5714285714D),\r\n new BigDecimal(3.5714285714D)};\r\nRow dataRow = new GenericRow(objects);\r\nRow dataRow2 = new GenericRow(objects);\r\nStructType structType = new StructType()\r\n .add(\"id1\", DataTypes.createDecimalType(38,10), true)\r\n .add(\"id2\", DataTypes.createDecimalType(38,10), true)\r\n .add(\"id3\", DataTypes.createDecimalType(38,10), true)\r\n .add(\"id4\", DataTypes.createDecimalType(38,10), true);\r\n\r\nfinal Dataset dataFrame = sparkSession.createDataFrame(Arrays.asList(dataRow,dataRow2), structType);\r\nSystem.out.println(dataFrame.schema());\r\ndataFrame.show();\r\nfinal Dataset df1 = dataFrame.groupBy(\"id1\",\"id2\")\r\n .agg( mean(\"id3\").alias(\"projection_factor\"));\r\ndf1.show();\r\nfinal Dataset df2 = df1\r\n .groupBy(\"id1\")\r\n .agg(max(\"projection_factor\"));\r\n\r\ndf2.show();\r\n{code}\r\n \r\n\r\nThe df2 should have:\r\n{code:java}\r\n+------------+----------------------+\r\n| id1|max(projection_factor)|\r\n+------------+----------------------+\r\n|3.5714285714| 3.5714285714|\r\n+------------+----------------------+\r\n\r\n{code}\r\ninstead it returns: \r\n{code:java}\r\n+------------+----------------------+\r\n| id1|max(projection_factor)|\r\n+------------+----------------------+\r\n|3.5714285714| 0.00035714285714|\r\n+------------+----------------------+\r\n{code}\r\n \r\n\r\n ","from":"reporter","subject":"Aggreate on Decimal Types does not work"},{"body":"Probably is worth to mention that my dataset is coming from a oracle DB","from":"developer"},{"body":"I followed the repro steps using the file you attached and I was unable to reproduce on current master. May you please check the repro code and try it on current master branch?","from":"developer"},{"body":"I didn't narrow down and I'm not sure that this is related to this issue though, I found v2.2.0 had an incorrect result (fixed in v2.2.1);\r\nAs marco said, I'd be greated if you try on master;\r\n{code}\r\n// v2.2.0\r\nscala> second_agg.show\r\n+----+-----+----------------+-----+ \r\n| id1| id2| maxf| minf|\r\n+----+-----+----------------+-----+\r\n|1498|88586|0.00238636363635|3E-14|\r\n+----+-----+----------------+-----+\r\n\r\n// v2.2.1\r\nscala> second_agg.show\r\n+----+-----+----------------+----------------+ \r\n| id1| id2| maxf| minf|\r\n+----+-----+----------------+----------------+\r\n|1498|88586|3.95833333330000|2.00000000000000|\r\n+----+-----+----------------+----------------+\r\n\r\n// v2.3.0\r\nscala> second_agg.show\r\n+----+-----+----------------+----------------+\r\n| id1| id2| maxf| minf|\r\n+----+-----+----------------+----------------+\r\n|1498|88586|3.95833333330000|2.00000000000000|\r\n+----+-----+----------------+----------------+\r\n\r\n// master\r\nscala> second_agg.show\r\n+----+-----+----------------+----------------+\r\n| id1| id2| maxf| minf|\r\n+----+-----+----------------+----------------+\r\n|1498|88586|3.95833333330000|2.00000000000000|\r\n+----+-----+----------------+----------------+\r\n{code}","from":"developer"},{"body":"Yes, I just tried my code and it works on spark 2.3.0","from":"developer"},{"body":"Was already Fixed on 2.3.0","from":"developer"}],"created":"2018-05-28T06:33:33.000+0000","description":"Hi, \r\n\r\nI think I found a really ugly bug in spark when performing aggregations with Decimals\r\n\r\nTo reproduce: \r\n\r\n \r\n{code:java}\r\nval df = spark.read.parquet(\"attached file\")\r\nval first_agg = fact_df.groupBy(\"id1\", \"id2\", \"start_date\").agg(mean(\"projection_factor\").alias(\"projection_factor\"))\r\nfirst_agg.show\r\nval second_agg = first_agg.groupBy(\"id1\",\"id2\").agg(max(\"projection_factor\").alias(\"maxf\"), min(\"projection_factor\").alias(\"minf\"))\r\nsecond_agg.show\r\n{code}\r\nFirst aggregation works fine the second aggregation seems to be summing instead of max value. I tried with spark 2.2.0 and 2.3.0 same problem.\r\n\r\nThe dataset as circa 800 Rows and the projection_factor has values from 0 until 100. the result should not be bigger that 5 but with get 265820543091454.... as result back.\r\n\r\n \r\n\r\n \r\n\r\nAs Code not 100% the same but I think there is really a bug there: \r\n\r\n \r\n{code:java}\r\nBigDecimal [] objects = new BigDecimal[]{\r\n new BigDecimal(3.5714285714D),\r\n new BigDecimal(3.5714285714D),\r\n new BigDecimal(3.5714285714D),\r\n new BigDecimal(3.5714285714D)};\r\nRow dataRow = new GenericRow(objects);\r\nRow dataRow2 = new GenericRow(objects);\r\nStructType structType = new StructType()\r\n .add(\"id1\", DataTypes.createDecimalType(38,10), true)\r\n .add(\"id2\", DataTypes.createDecimalType(38,10), true)\r\n .add(\"id3\", DataTypes.createDecimalType(38,10), true)\r\n .add(\"id4\", DataTypes.createDecimalType(38,10), true);\r\n\r\nfinal Dataset dataFrame = sparkSession.createDataFrame(Arrays.asList(dataRow,dataRow2), structType);\r\nSystem.out.println(dataFrame.schema());\r\ndataFrame.show();\r\nfinal Dataset df1 = dataFrame.groupBy(\"id1\",\"id2\")\r\n .agg( mean(\"id3\").alias(\"projection_factor\"));\r\ndf1.show();\r\nfinal Dataset df2 = df1\r\n .groupBy(\"id1\")\r\n .agg(max(\"projection_factor\"));\r\n\r\ndf2.show();\r\n{code}\r\n \r\n\r\nThe df2 should have:\r\n{code:java}\r\n+------------+----------------------+\r\n| id1|max(projection_factor)|\r\n+------------+----------------------+\r\n|3.5714285714| 3.5714285714|\r\n+------------+----------------------+\r\n\r\n{code}\r\ninstead it returns: \r\n{code:java}\r\n+------------+----------------------+\r\n| id1|max(projection_factor)|\r\n+------------+----------------------+\r\n|3.5714285714| 0.00035714285714|\r\n+------------+----------------------+\r\n{code}\r\n \r\n\r\n ","issue_id":"13162332","key":"SPARK-24401","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-05-29T05:49:58.000+0000","role":"fixed_distractor","summary":"Aggreate on Decimal Types does not work"} {"case_id":"13164700","cluster":"DISTRACTOR-SPARK-24486","comments":[{"body":"Thank you for reporting a problem.\r\nCould you please let us know which value is shown for each of three results in `sum(...)`?","created":"2018-06-07T15:55:28.468+0000"},{"body":"Thanks for your comment.\r\n\r\nIn all 3 cases (spark 2.2.1, 2.3.0 and latest version from master) I am using a simple test workload to investigate the issue, that is:\r\nspark.read.parquet(\"file:///tmp/deleteme1\").limit(1).show()\r\nThe output is simply the first row of the test table, that is an int value \"1\" and an array of 30000 int elements.\r\n\r\nI'll be happy to provide additional info on the tests and workload. BTW, it should be straighforward to reproduce this issue in a test environemnt if you can spare the time.\r\n\r\n ","created":"2018-06-07T20:13:07.126+0000"},{"body":"[~lucacanali]  May be caused by SPARK-23023. Cloud you use {{collect()}} to test you case. Below is my benchmark:\r\n\r\ncode:\r\n{code:java}\r\nval benchmark = new Benchmark(\"read parquet\", 1, minNumIters = 10)\r\nbenchmark.addCase(\"show\", 5) { _ =>\r\n spark.read.parquet(\"file:///tmp/deleteme1\").limit(1).show()\r\n}\r\nbenchmark.addCase(\"collect\", 5) { _ =>\r\n spark.read.parquet(\"file:///tmp/deleteme1\").limit(1).collect()\r\n}\r\nbenchmark.run()\r\n{code}\r\n{noformat}\r\nJava HotSpot(TM) 64-Bit Server VM 1.8.0_151-b12 on Mac OS X 10.12.6\r\nIntel(R) Core(TM) i7-7820HQ CPU @ 2.90GHz\r\n\r\nread parquet: Best/Avg Time(ms) Rate(M/s) Per Row(ns) Relative\r\n------------------------------------------------------------------------------------------------\r\nshow 14578 / 15174 0.0 14578205056.0 1.0X\r\ncollect 252 / 271 0.0 251586336.0 57.9X\r\n{noformat}","created":"2018-09-22T15:57:34.056+0000"},{"body":"Thanks [~yumwang] for looking at this. Indeed I confirm that using collect instead of show is faster. In addition, testing on Spark master (March 12 2019) I see that show works fast there too (I have not yet looked at which PR fixed this).","created":"2019-03-12T10:08:10.461+0000"}],"conversations":[{"body":"We have found an issue of slow performance in one of our applications when running on Spark 2.3.0 (the same workload does not have a performance issue on Spark 2.2.1). We suspect a regression in the area of handling columns of ArrayType. I have built a simplified test case showing a manifestation of the issue to help with troubleshooting:\r\n\r\n \r\n\r\n \r\n{code:java}\r\n// prepare test data\r\nval stringListValues=Range(1,30000).mkString(\",\")\r\nsql(s\"select 1 as myid, Array($stringListValues) as myarray from range(20000)\").repartition(1).write.parquet(\"file:///tmp/deleteme1\")\r\n\r\n// run test\r\nspark.read.parquet(\"file:///tmp/deleteme1\").limit(1).show(){code}\r\n\r\nPerformance measurements:\r\n\r\n \r\n\r\nOn a desktop-size test system, the test runs in about 2 sec using Spark 2.2.1 (runtime goes down to subsecond in subsequent runs) and takes close to 20 sec on Spark 2.3.0\r\n\r\n \r\n\r\nAdditional drill-down using Spark task metrics data, show that in Spark 2.2.1 only 2 records are read by this workload, while on Spark 2.3.0 all rows in the file are read, which appears anomalous.\r\n\r\nExample:\r\n{code:java}\r\nbin/spark-shell --master local[*] --driver-memory 2g --packages ch.cern.sparkmeasure:spark-measure_2.11:0.11\r\nval stageMetrics = ch.cern.sparkmeasure.StageMetrics(spark) \r\nstageMetrics.runAndMeasure(spark.read.parquet(\"file:///tmp/deleteme1\").limit(1).show())\r\n{code}\r\n \r\n\r\n \r\n\r\nSelected metrics from Spark 2.3.0 run:\r\n\r\n \r\n{noformat}\r\nelapsedTime => 17849 (18 s)\r\nsum(numTasks) => 11\r\nsum(recordsRead) => 20000\r\nsum(bytesRead) => 1136448171 (1083.0 MB){noformat}\r\n \r\n\r\n \r\n\r\nFrom Spark 2.2.1 run:\r\n\r\n \r\n{noformat}\r\nelapsedTime => 1329 (1 s)\r\nsum(numTasks) => 2\r\nsum(recordsRead) => 2\r\nsum(bytesRead) => 269162610 (256.0 MB)\r\n{noformat}\r\n \r\n\r\nNote: Using Spark built from master (as I write this, June 7th 2018) shows the same behavior as found in Spark 2.3.0\r\n\r\n ","from":"reporter","subject":"Slow performance reading ArrayType columns"},{"body":"Thank you for reporting a problem.\r\nCould you please let us know which value is shown for each of three results in `sum(...)`?","from":"developer"},{"body":"Thanks for your comment.\r\n\r\nIn all 3 cases (spark 2.2.1, 2.3.0 and latest version from master) I am using a simple test workload to investigate the issue, that is:\r\nspark.read.parquet(\"file:///tmp/deleteme1\").limit(1).show()\r\nThe output is simply the first row of the test table, that is an int value \"1\" and an array of 30000 int elements.\r\n\r\nI'll be happy to provide additional info on the tests and workload. BTW, it should be straighforward to reproduce this issue in a test environemnt if you can spare the time.\r\n\r\n ","from":"developer"},{"body":"[~lucacanali]  May be caused by SPARK-23023. Cloud you use {{collect()}} to test you case. Below is my benchmark:\r\n\r\ncode:\r\n{code:java}\r\nval benchmark = new Benchmark(\"read parquet\", 1, minNumIters = 10)\r\nbenchmark.addCase(\"show\", 5) { _ =>\r\n spark.read.parquet(\"file:///tmp/deleteme1\").limit(1).show()\r\n}\r\nbenchmark.addCase(\"collect\", 5) { _ =>\r\n spark.read.parquet(\"file:///tmp/deleteme1\").limit(1).collect()\r\n}\r\nbenchmark.run()\r\n{code}\r\n{noformat}\r\nJava HotSpot(TM) 64-Bit Server VM 1.8.0_151-b12 on Mac OS X 10.12.6\r\nIntel(R) Core(TM) i7-7820HQ CPU @ 2.90GHz\r\n\r\nread parquet: Best/Avg Time(ms) Rate(M/s) Per Row(ns) Relative\r\n------------------------------------------------------------------------------------------------\r\nshow 14578 / 15174 0.0 14578205056.0 1.0X\r\ncollect 252 / 271 0.0 251586336.0 57.9X\r\n{noformat}","from":"developer"},{"body":"Thanks [~yumwang] for looking at this. Indeed I confirm that using collect instead of show is faster. In addition, testing on Spark master (March 12 2019) I see that show works fast there too (I have not yet looked at which PR fixed this).","from":"developer"}],"created":"2018-06-07T14:34:19.000+0000","description":"We have found an issue of slow performance in one of our applications when running on Spark 2.3.0 (the same workload does not have a performance issue on Spark 2.2.1). We suspect a regression in the area of handling columns of ArrayType. I have built a simplified test case showing a manifestation of the issue to help with troubleshooting:\r\n\r\n \r\n\r\n \r\n{code:java}\r\n// prepare test data\r\nval stringListValues=Range(1,30000).mkString(\",\")\r\nsql(s\"select 1 as myid, Array($stringListValues) as myarray from range(20000)\").repartition(1).write.parquet(\"file:///tmp/deleteme1\")\r\n\r\n// run test\r\nspark.read.parquet(\"file:///tmp/deleteme1\").limit(1).show(){code}\r\n\r\nPerformance measurements:\r\n\r\n \r\n\r\nOn a desktop-size test system, the test runs in about 2 sec using Spark 2.2.1 (runtime goes down to subsecond in subsequent runs) and takes close to 20 sec on Spark 2.3.0\r\n\r\n \r\n\r\nAdditional drill-down using Spark task metrics data, show that in Spark 2.2.1 only 2 records are read by this workload, while on Spark 2.3.0 all rows in the file are read, which appears anomalous.\r\n\r\nExample:\r\n{code:java}\r\nbin/spark-shell --master local[*] --driver-memory 2g --packages ch.cern.sparkmeasure:spark-measure_2.11:0.11\r\nval stageMetrics = ch.cern.sparkmeasure.StageMetrics(spark) \r\nstageMetrics.runAndMeasure(spark.read.parquet(\"file:///tmp/deleteme1\").limit(1).show())\r\n{code}\r\n \r\n\r\n \r\n\r\nSelected metrics from Spark 2.3.0 run:\r\n\r\n \r\n{noformat}\r\nelapsedTime => 17849 (18 s)\r\nsum(numTasks) => 11\r\nsum(recordsRead) => 20000\r\nsum(bytesRead) => 1136448171 (1083.0 MB){noformat}\r\n \r\n\r\n \r\n\r\nFrom Spark 2.2.1 run:\r\n\r\n \r\n{noformat}\r\nelapsedTime => 1329 (1 s)\r\nsum(numTasks) => 2\r\nsum(recordsRead) => 2\r\nsum(bytesRead) => 269162610 (256.0 MB)\r\n{noformat}\r\n \r\n\r\nNote: Using Spark built from master (as I write this, June 7th 2018) shows the same behavior as found in Spark 2.3.0\r\n\r\n ","issue_id":"13164700","key":"SPARK-24486","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-03-12T10:08:27.000+0000","role":"fixed_distractor","summary":"Slow performance reading ArrayType columns"} {"case_id":"13165595","cluster":"DISTRACTOR-SPARK-24530","comments":[{"body":"In my case, it works correctly on current master (commit 9786ce6). My environment is Ubuntu 18.04, Python 2.7, and Sphinx 1.7.5.\r\n\r\n!pyspark-ml-doc-utuntu18.04-python2.7-sphinx-1.7.5.png! ","created":"2018-06-13T03:12:35.620+0000"},{"body":"Hi, [~mengxr] .\r\n\r\nI got the following locally. It generated correctly. !image-2018-06-13-15-15-51-025.png!\r\n\r\nHowever, as you pointed out, I found that the following status. Some docs are broken.\r\n\r\n2.1.x\r\n\r\n(O) [https://spark.apache.org/docs/2.1.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(O) [https://spark.apache.org/docs/2.1.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.1.2/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n\r\n2.2.x\r\n(O) [https://spark.apache.org/docs/2.2.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.2.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n\r\n2.3.x\r\n(O) [https://spark.apache.org/docs/2.3.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.3.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]","created":"2018-06-13T22:23:52.800+0000"},{"body":"OMG, I am sorry; I was misunderstanding. The documentation is also broken in my environment.","created":"2018-06-14T08:37:09.790+0000"},{"body":"Probably it's related with Sphinx's version just from my vague rough memory. I just felt like I better mention it here at least ...","created":"2018-06-24T06:39:18.562+0000"},{"body":"[~dongjoon] [~hyukjin.kwon] Could you report your system, Python, and Sphinx version?\r\n\r\nMine is: macOS, python 2.7, sphinx 1.6.3.","created":"2018-06-25T18:30:38.578+0000"},{"body":"macOS, Python 2.7.14, Sphinx 1.4.1 shows:\r\n\r\n{code}\r\nclass pyspark.ml.classification.LogisticRegression(*args, **kwargs)[source]\r\nLogistic regression. This class supports multinomial logistic (softmax) and binomial logistic regression.\r\n{code}\r\n","created":"2018-06-26T08:47:36.624+0000"},{"body":"I have another computer: macOS, Python 2.7.14, Sphinx 1.7.2 shows:\r\n\r\n{code}\r\nclass pyspark.ml.classification.LogisticRegression(*args, **kwargs)[source]\r\nLogistic regression. This class supports multinomial logistic (softmax) and binomial logistic regression.\r\n{code}\r\n\r\nI think we need [~dongjoon]'s input.","created":"2018-06-26T08:50:21.000+0000"},{"body":"[~mengxr] and [~hyukjin.kwon]. My environment is macOS, *python 3*, Sphinx v1.6.3. \r\n{code}\r\n~/s/p/docs:master$ make html\r\nsphinx-build -b html -d _build/doctrees . _build/html\r\nRunning Sphinx v1.6.3\r\nmaking output directory...\r\n...\r\n{code}\r\n\r\nAccording to the above reports, many combinations of Python 2.7 and Sphinx looks broken?","created":"2018-06-26T21:42:01.895+0000"},{"body":"Confirmed that macOS, python 3, and Sphinx v1.6.6 can produce correct doc on my machine. I didn't find any reports on Sphinx github. So if we could make a minimal reproducible example, we should report the issue to Sphinx. On our side, we should update the release procedure doc to use Python 3 to generate docs. We should also update the official docs that are broken (2.1.2, 2.2.1, 2.3.1).\r\n\r\n[~hyukjin.kwon] Do you have time to take this ticket? (feel free to say no if you are busy:)\r\n\r\ncc: [~smilegator]\r\n\r\n ","created":"2018-06-27T03:11:31.920+0000"},{"body":"Will take a look on this weekends. Please go ahead if anyone finds some time till then :-).","created":"2018-06-27T04:01:39.247+0000"},{"body":"[~hyukjin.kwon]  Thanks for helping this!","created":"2018-06-27T04:11:49.710+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21659","created":"2018-06-28T17:11:08.910+0000"},{"body":"[~vanzin] and [~jerryshao] FYI.","created":"2018-06-30T09:16:46.929+0000"},{"body":"[~hyukjin.kwon], Spark 2.1.3 and 2.2.2 are on vote, can you please fix the issue and leave the comments in the related threads.","created":"2018-07-02T03:19:54.974+0000"},{"body":"Yea, will post a email to related threads, and try to deal with it very soon. [~mengxr], mind if I set the priority to Critical since we have a workaround to get through this anyway?","created":"2018-07-02T04:25:12.474+0000"},{"body":"related PRs are also open in https://github.com/apache/spark-website/pulls","created":"2018-07-02T18:06:40.544+0000"},{"body":"Hi [~hyukjin.kwon] what is the current status of this JIRA, do you have an ETA about it?","created":"2018-07-04T01:55:25.419+0000"},{"body":"There's workaround for this. To cut it short, Sphinx for Python 3 is required and installed e.g., {{sudo pip3 install sphinx}} but Sphinx for Python 2 should be removed first before installing Sphinx for Python 3 for sure if Sphinx for Python 2 is already installed.\r\n\r\nCurrently, whether it uses Sphinx for Pythom 3 or not can be manually checked by, {{cd python/docs && make clean html}}, checking if the keywords arguments are shown in the equaivelent link above before making the release documentation for Python API.","created":"2018-07-04T02:22:51.488+0000"},{"body":"[~mengxr], I lowered the priority to {{Critical}} for now since I believe this doesn't block the release although it's critical. Please revert my action if you think differently. I don't mind.","created":"2018-07-04T02:24:28.325+0000"},{"body":"Issue resolved by pull request 21659\n[https://github.com/apache/spark/pull/21659]","created":"2018-07-11T02:11:11.372+0000"},{"body":"Note to myself:\r\n\r\nThis ended up with misconfiguration and mistakes in docstrings within PySpark. We should configure differently but then it causes doc generation to be broken currently.\r\nWe should fix docs and configuration, and check the built documentation closely to fix this cleanly, and then should revert the current changes within {{Makefile}}\r\n\r\nAlso see https://github.com/sphinx-doc/sphinx/issues/5142#issuecomment-414634234","created":"2018-08-21T10:58:15.828+0000"},{"body":"User 'cloud-fan' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22607","created":"2018-10-02T07:14:08.143+0000"},{"body":"User 'cloud-fan' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22607","created":"2018-10-02T07:15:22.478+0000"},{"body":"User 'gatorsmile' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/24503","created":"2019-05-01T03:38:47.308+0000"}],"conversations":[{"body":"I generated python docs from master locally using `make html`. However, the generated html doc doesn't render class docs correctly. I attached the screenshot from Spark 2.3 docs and master docs generated on my local. Not sure if this is because my local setup.\r\n\r\ncc: [~dongjoon] Could you help verify?\r\n\r\n \r\n\r\nThe followings are our released doc status. Some recent docs seems to be broken.\r\n\r\n*2.1.x*\r\n\r\n(O) [https://spark.apache.org/docs/2.1.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(O) [https://spark.apache.org/docs/2.1.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.1.2/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n\r\n*2.2.x*\r\n(O) [https://spark.apache.org/docs/2.2.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.2.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n\r\n*2.3.x*\r\n(O) [https://spark.apache.org/docs/2.3.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.3.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]","from":"reporter","subject":"Sphinx doesn't render autodoc_docstring_signature correctly (with Python 2?) and pyspark.ml docs are broken"},{"body":"In my case, it works correctly on current master (commit 9786ce6). My environment is Ubuntu 18.04, Python 2.7, and Sphinx 1.7.5.\r\n\r\n!pyspark-ml-doc-utuntu18.04-python2.7-sphinx-1.7.5.png! ","from":"developer"},{"body":"Hi, [~mengxr] .\r\n\r\nI got the following locally. It generated correctly. !image-2018-06-13-15-15-51-025.png!\r\n\r\nHowever, as you pointed out, I found that the following status. Some docs are broken.\r\n\r\n2.1.x\r\n\r\n(O) [https://spark.apache.org/docs/2.1.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(O) [https://spark.apache.org/docs/2.1.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.1.2/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n\r\n2.2.x\r\n(O) [https://spark.apache.org/docs/2.2.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.2.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n\r\n2.3.x\r\n(O) [https://spark.apache.org/docs/2.3.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.3.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]","from":"developer"},{"body":"OMG, I am sorry; I was misunderstanding. The documentation is also broken in my environment.","from":"developer"},{"body":"Probably it's related with Sphinx's version just from my vague rough memory. I just felt like I better mention it here at least ...","from":"developer"},{"body":"[~dongjoon] [~hyukjin.kwon] Could you report your system, Python, and Sphinx version?\r\n\r\nMine is: macOS, python 2.7, sphinx 1.6.3.","from":"developer"},{"body":"macOS, Python 2.7.14, Sphinx 1.4.1 shows:\r\n\r\n{code}\r\nclass pyspark.ml.classification.LogisticRegression(*args, **kwargs)[source]\r\nLogistic regression. This class supports multinomial logistic (softmax) and binomial logistic regression.\r\n{code}\r\n","from":"developer"},{"body":"I have another computer: macOS, Python 2.7.14, Sphinx 1.7.2 shows:\r\n\r\n{code}\r\nclass pyspark.ml.classification.LogisticRegression(*args, **kwargs)[source]\r\nLogistic regression. This class supports multinomial logistic (softmax) and binomial logistic regression.\r\n{code}\r\n\r\nI think we need [~dongjoon]'s input.","from":"developer"},{"body":"[~mengxr] and [~hyukjin.kwon]. My environment is macOS, *python 3*, Sphinx v1.6.3. \r\n{code}\r\n~/s/p/docs:master$ make html\r\nsphinx-build -b html -d _build/doctrees . _build/html\r\nRunning Sphinx v1.6.3\r\nmaking output directory...\r\n...\r\n{code}\r\n\r\nAccording to the above reports, many combinations of Python 2.7 and Sphinx looks broken?","from":"developer"},{"body":"Confirmed that macOS, python 3, and Sphinx v1.6.6 can produce correct doc on my machine. I didn't find any reports on Sphinx github. So if we could make a minimal reproducible example, we should report the issue to Sphinx. On our side, we should update the release procedure doc to use Python 3 to generate docs. We should also update the official docs that are broken (2.1.2, 2.2.1, 2.3.1).\r\n\r\n[~hyukjin.kwon] Do you have time to take this ticket? (feel free to say no if you are busy:)\r\n\r\ncc: [~smilegator]\r\n\r\n ","from":"developer"},{"body":"Will take a look on this weekends. Please go ahead if anyone finds some time till then :-).","from":"developer"},{"body":"[~hyukjin.kwon]  Thanks for helping this!","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21659","from":"developer"},{"body":"[~vanzin] and [~jerryshao] FYI.","from":"developer"},{"body":"[~hyukjin.kwon], Spark 2.1.3 and 2.2.2 are on vote, can you please fix the issue and leave the comments in the related threads.","from":"developer"},{"body":"Yea, will post a email to related threads, and try to deal with it very soon. [~mengxr], mind if I set the priority to Critical since we have a workaround to get through this anyway?","from":"developer"},{"body":"related PRs are also open in https://github.com/apache/spark-website/pulls","from":"developer"},{"body":"Hi [~hyukjin.kwon] what is the current status of this JIRA, do you have an ETA about it?","from":"developer"},{"body":"There's workaround for this. To cut it short, Sphinx for Python 3 is required and installed e.g., {{sudo pip3 install sphinx}} but Sphinx for Python 2 should be removed first before installing Sphinx for Python 3 for sure if Sphinx for Python 2 is already installed.\r\n\r\nCurrently, whether it uses Sphinx for Pythom 3 or not can be manually checked by, {{cd python/docs && make clean html}}, checking if the keywords arguments are shown in the equaivelent link above before making the release documentation for Python API.","from":"developer"},{"body":"[~mengxr], I lowered the priority to {{Critical}} for now since I believe this doesn't block the release although it's critical. Please revert my action if you think differently. I don't mind.","from":"developer"},{"body":"Issue resolved by pull request 21659\n[https://github.com/apache/spark/pull/21659]","from":"developer"},{"body":"Note to myself:\r\n\r\nThis ended up with misconfiguration and mistakes in docstrings within PySpark. We should configure differently but then it causes doc generation to be broken currently.\r\nWe should fix docs and configuration, and check the built documentation closely to fix this cleanly, and then should revert the current changes within {{Makefile}}\r\n\r\nAlso see https://github.com/sphinx-doc/sphinx/issues/5142#issuecomment-414634234","from":"developer"},{"body":"User 'cloud-fan' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22607","from":"developer"},{"body":"User 'cloud-fan' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22607","from":"developer"},{"body":"User 'gatorsmile' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/24503","from":"developer"}],"created":"2018-06-12T15:22:54.000+0000","description":"I generated python docs from master locally using `make html`. However, the generated html doc doesn't render class docs correctly. I attached the screenshot from Spark 2.3 docs and master docs generated on my local. Not sure if this is because my local setup.\r\n\r\ncc: [~dongjoon] Could you help verify?\r\n\r\n \r\n\r\nThe followings are our released doc status. Some recent docs seems to be broken.\r\n\r\n*2.1.x*\r\n\r\n(O) [https://spark.apache.org/docs/2.1.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(O) [https://spark.apache.org/docs/2.1.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.1.2/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n\r\n*2.2.x*\r\n(O) [https://spark.apache.org/docs/2.2.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.2.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n\r\n*2.3.x*\r\n(O) [https://spark.apache.org/docs/2.3.0/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]\r\n(X) [https://spark.apache.org/docs/2.3.1/api/python/pyspark.ml.html#pyspark.ml.classification.LogisticRegression]","issue_id":"13165595","key":"SPARK-24530","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-07-11T02:11:11.000+0000","role":"fixed_distractor","summary":"Sphinx doesn't render autodoc_docstring_signature correctly (with Python 2?) and pyspark.ml docs are broken"} {"case_id":"13173948","cluster":"DISTRACTOR-SPARK-24894","comments":[{"body":"[~mcheah]. We need to make sure the truncation leads to a valid hostname.","created":"2018-07-24T18:01:49.883+0000"},{"body":"Why is Fix Version 3.0.0? This looks like a bug fix to me, so should've been in a patch version?","created":"2020-02-19T00:16:11.347+0000"}],"conversations":[{"body":"The truncation for hostname happening here [https://github.com/apache/spark/blob/5ff1b9ba1983d5601add62aef64a3e87d07050eb/resource-managers/kubernetes/core/src/main/scala/org/apache/spark/deploy/k8s/features/BasicExecutorFeatureStep.scala#L77]  is a problematic and can lead to DNS names starting with \"-\". \r\n\r\nOriginally filled here : https://github.com/GoogleCloudPlatform/spark-on-k8s-operator/issues/229\r\n\r\n```\r\n{{2018-07-23 21:21:42 ERROR Utils:91 - Uncaught exception in thread kubernetes-pod-allocator io.fabric8.kubernetes.client.KubernetesClientException: Failure executing: POST at: https://kubernetes.default.svc/api/v1/namespaces/default/pods. Message: Pod \"user-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9\" is invalid: spec.hostname: Invalid value: \"-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9\": a DNS-1123 label must consist of lower case alphanumeric characters or '-', and must start and end with an alphanumeric character (e.g. 'my-name', or '123-abc', regex used for validation is '[a-z0-9]([-a-z0-9]*[a-z0-9])?'). Received status: Status(apiVersion=v1, code=422, details=StatusDetails(causes=[StatusCause(field=spec.hostname, message=Invalid value: \"-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9\": a DNS-1123 label must consist of lower case alphanumeric characters or '-', and must start and end with an alphanumeric character (e.g. 'my-name', or '123-abc', regex used for validation is '[a-z0-9]([-a-z0-9]*[a-z0-9])?'), reason=FieldValueInvalid, additionalProperties={})], group=null, kind=Pod, name=user-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9, retryAfterSeconds=null, uid=null, additionalProperties={}), kind=Status, message=Pod \"user-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9\" is invalid: spec.hostname: Invalid value: \"-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9\": a DNS-1123 label must consist of lower case alphanumeric characters or '-', and must start and end with an alphanumeric character (e.g. 'my-name', or '123-abc', regex used for validation is '[a-z0-9]([-a-z0-9]*[a-z0-9])?'), metadata=ListMeta(resourceVersion=null, selfLink=null, additionalProperties={}), reason=Invalid, status=Failure, additionalProperties={}). at io.fabric8.kubernetes.client.dsl.base.OperationSupport.requestFailure(OperationSupport.java:470) at io.fabric8.kubernetes.client.dsl.base.OperationSupport.assertResponseCode(OperationSupport.java:409) at io.fabric8.kubernetes.client.dsl.base.OperationSupport.handleResponse(OperationSupport.java:379) at io.fabric8.kubernetes.client.dsl.base.OperationSupport.handleResponse(OperationSupport.java:343) at io.fabric8.kubernetes.client.dsl.base.OperationSupport.handleCreate(OperationSupport.java:226) at io.fabric8.kubernetes.client.dsl.base.BaseOperation.handleCreate(BaseOperation.java:769) at io.fabric8.kubernetes.client.dsl.base.BaseOperation.create(BaseOperation.java:356) at org.apache.spark.scheduler.cluster.k8s.KubernetesClusterSchedulerBackend$$anon$1$$anonfun$3$$anonfun$apply$3.apply(KubernetesClusterSchedulerBackend.scala:140) at org.apache.spark.scheduler.cluster.k8s.KubernetesClusterSchedulerBackend$$anon$1$$anonfun$3$$anonfun$apply$3.apply(KubernetesClusterSchedulerBackend.scala:140) at org.apache.spark.util.Utils$.tryLog(Utils.scala:1922) at org.apache.spark.scheduler.cluster.k8s.KubernetesClusterSchedulerBackend$$anon$1$$anonfun$3.apply(KubernetesClusterSchedulerBackend.scala:139) at org.apache.spark.scheduler.cluster.k8s.KubernetesClusterSchedulerBackend$$anon$1$$anonfun$3.apply(KubernetesClusterSchedulerBackend.scala:138) at scala.collection.MapLike$MappedValues$$anonfun$foreach$3.apply(MapLike.scala:245) at scala.collection.MapLike$MappedValues$$anonfun$foreach$3.apply(MapLike.scala:245) at scala.collection.TraversableLike$WithFilter$$anonfun$foreach$1.apply(TraversableLike.scala:733) at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:99) at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:99) at scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:230) at scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:40) at scala.collection.mutable.HashMap.foreach(HashMap.scala:99) at scala.collection.TraversableLike$WithFilter.foreach(TraversableLike.scala:732) at scala.collection.MapLike$MappedValues.foreach(MapLike.scala:245) at scala.collection.TraversableLike$class.map(TraversableLike.scala:234) at scala.collection.AbstractTraversable.map(Traversable.scala:104) at org.apache.spark.scheduler.cluster.k8s.KubernetesClusterSchedulerBackend$$anon$1.run(KubernetesClusterSchedulerBackend.scala:145) at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180) at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294) at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149) at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624) at java.lang.Thread.run(Thread.java:748)}}\r\n```","from":"reporter","subject":"Invalid DNS name due to hostname truncation "},{"body":"[~mcheah]. We need to make sure the truncation leads to a valid hostname.","from":"developer"},{"body":"Why is Fix Version 3.0.0? This looks like a bug fix to me, so should've been in a patch version?","from":"developer"}],"created":"2018-07-23T22:53:59.000+0000","description":"The truncation for hostname happening here [https://github.com/apache/spark/blob/5ff1b9ba1983d5601add62aef64a3e87d07050eb/resource-managers/kubernetes/core/src/main/scala/org/apache/spark/deploy/k8s/features/BasicExecutorFeatureStep.scala#L77]  is a problematic and can lead to DNS names starting with \"-\". \r\n\r\nOriginally filled here : https://github.com/GoogleCloudPlatform/spark-on-k8s-operator/issues/229\r\n\r\n```\r\n{{2018-07-23 21:21:42 ERROR Utils:91 - Uncaught exception in thread kubernetes-pod-allocator io.fabric8.kubernetes.client.KubernetesClientException: Failure executing: POST at: https://kubernetes.default.svc/api/v1/namespaces/default/pods. Message: Pod \"user-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9\" is invalid: spec.hostname: Invalid value: \"-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9\": a DNS-1123 label must consist of lower case alphanumeric characters or '-', and must start and end with an alphanumeric character (e.g. 'my-name', or '123-abc', regex used for validation is '[a-z0-9]([-a-z0-9]*[a-z0-9])?'). Received status: Status(apiVersion=v1, code=422, details=StatusDetails(causes=[StatusCause(field=spec.hostname, message=Invalid value: \"-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9\": a DNS-1123 label must consist of lower case alphanumeric characters or '-', and must start and end with an alphanumeric character (e.g. 'my-name', or '123-abc', regex used for validation is '[a-z0-9]([-a-z0-9]*[a-z0-9])?'), reason=FieldValueInvalid, additionalProperties={})], group=null, kind=Pod, name=user-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9, retryAfterSeconds=null, uid=null, additionalProperties={}), kind=Status, message=Pod \"user-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9\" is invalid: spec.hostname: Invalid value: \"-archetypes-all-weekly-1532380861251850404-1532380862321-exec-9\": a DNS-1123 label must consist of lower case alphanumeric characters or '-', and must start and end with an alphanumeric character (e.g. 'my-name', or '123-abc', regex used for validation is '[a-z0-9]([-a-z0-9]*[a-z0-9])?'), metadata=ListMeta(resourceVersion=null, selfLink=null, additionalProperties={}), reason=Invalid, status=Failure, additionalProperties={}). at io.fabric8.kubernetes.client.dsl.base.OperationSupport.requestFailure(OperationSupport.java:470) at io.fabric8.kubernetes.client.dsl.base.OperationSupport.assertResponseCode(OperationSupport.java:409) at io.fabric8.kubernetes.client.dsl.base.OperationSupport.handleResponse(OperationSupport.java:379) at io.fabric8.kubernetes.client.dsl.base.OperationSupport.handleResponse(OperationSupport.java:343) at io.fabric8.kubernetes.client.dsl.base.OperationSupport.handleCreate(OperationSupport.java:226) at io.fabric8.kubernetes.client.dsl.base.BaseOperation.handleCreate(BaseOperation.java:769) at io.fabric8.kubernetes.client.dsl.base.BaseOperation.create(BaseOperation.java:356) at org.apache.spark.scheduler.cluster.k8s.KubernetesClusterSchedulerBackend$$anon$1$$anonfun$3$$anonfun$apply$3.apply(KubernetesClusterSchedulerBackend.scala:140) at org.apache.spark.scheduler.cluster.k8s.KubernetesClusterSchedulerBackend$$anon$1$$anonfun$3$$anonfun$apply$3.apply(KubernetesClusterSchedulerBackend.scala:140) at org.apache.spark.util.Utils$.tryLog(Utils.scala:1922) at org.apache.spark.scheduler.cluster.k8s.KubernetesClusterSchedulerBackend$$anon$1$$anonfun$3.apply(KubernetesClusterSchedulerBackend.scala:139) at org.apache.spark.scheduler.cluster.k8s.KubernetesClusterSchedulerBackend$$anon$1$$anonfun$3.apply(KubernetesClusterSchedulerBackend.scala:138) at scala.collection.MapLike$MappedValues$$anonfun$foreach$3.apply(MapLike.scala:245) at scala.collection.MapLike$MappedValues$$anonfun$foreach$3.apply(MapLike.scala:245) at scala.collection.TraversableLike$WithFilter$$anonfun$foreach$1.apply(TraversableLike.scala:733) at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:99) at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:99) at scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:230) at scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:40) at scala.collection.mutable.HashMap.foreach(HashMap.scala:99) at scala.collection.TraversableLike$WithFilter.foreach(TraversableLike.scala:732) at scala.collection.MapLike$MappedValues.foreach(MapLike.scala:245) at scala.collection.TraversableLike$class.map(TraversableLike.scala:234) at scala.collection.AbstractTraversable.map(Traversable.scala:104) at org.apache.spark.scheduler.cluster.k8s.KubernetesClusterSchedulerBackend$$anon$1.run(KubernetesClusterSchedulerBackend.scala:145) at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180) at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294) at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149) at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624) at java.lang.Thread.run(Thread.java:748)}}\r\n```","issue_id":"13173948","key":"SPARK-24894","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-02-19T23:20:40.000+0000","role":"fixed_distractor","summary":"Invalid DNS name due to hostname truncation "} {"case_id":"13175244","cluster":"DISTRACTOR-SPARK-24950","comments":[{"body":"one solution, of course, is to pin the java version on the upcoming ubuntu workers to one that passes this test, but things like this make build engineers like me die a little bit inside.","created":"2018-07-27T18:41:36.588+0000"},{"body":"It's pretty clear this is down to differences in how time zones are defined, as they change over time and the JDK incorporates updated versions of the standard definitions in each release. \r\n\r\nIt looks like the difference between _171 and _181 is the difference between 2018c and 2018e in this table: http://www.oracle.com/technetwork/java/javase/tzdata-versions-138805.html\r\n\r\nNothing obviously relevant from Oracle's release notes. But I found this in the notes for 2018d:\r\n\r\n[http://mm.icann.org/pipermail/tz-announce/2018-March/000049.html]\r\n\r\n\"Enderbury and Kiritimati skipped New Year's Eve 1994, not New Year's Day 1995.  (Thanks to Kerry Shetline.)\"\r\n\r\nSo the answer is probably that the test has to be updated to reflect the fix to the timezone definition.\r\n\r\n \r\n\r\nOf course, if the test changes, it also starts failing on older Java 8 versions! probably not worth it.\r\n\r\nI'd suggest we resolve it by commenting this out with a note. There's no evidence this is a problem in Spark itself.","created":"2018-07-27T20:03:51.178+0000"},{"body":"sgtm\r\n\r\ni also dug through the java release notes WRT timezone changes and didn't find anything (which i forgot to mention).  sorry about that!  :)\r\n\r\ni'll start by just commenting out the failing TZ (Pacific/Enderbury) and see if that works.\r\n\r\n ","created":"2018-07-27T20:44:28.493+0000"},{"body":"User 'd80tb7' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21901","created":"2018-07-27T20:46:05.865+0000"},{"body":"Hi,\r\n\r\n \r\n\r\njust to say that I looked at this and came to the same conclusion as Sean.  I've submitted a PR which excludes both New Years Eve and New Years day from the test- which should mean it will work on both old and new jvms.","created":"2018-07-27T20:48:28.377+0000"},{"body":"testing this manually:\r\n\r\nhttps://amplab.cs.berkeley.edu/jenkins/job/spark-master-test-sbt-hadoop-2.6-ubuntu-test/878/","created":"2018-07-27T20:54:41.788+0000"},{"body":"User 'srowen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21903","created":"2018-07-27T22:20:06.162+0000"},{"body":"the spark-master-test-sbt build failed on ubuntu, but the the improtant bit is that the DateTimeUtilsSuite tests passed!\r\n\r\nback to unraveling R...  :\\\r\n\r\nthanks for the quick patch, [~d80tb7]!","created":"2018-07-28T01:43:17.483+0000"},{"body":"Issue resolved by pull request 21901\n[https://github.com/apache/spark/pull/21901]","created":"2018-07-28T15:41:18.142+0000"},{"body":"[~srowen] [~d80tb7] i'm thinking that we actually need to backport this change to previous branches.\r\n\r\n:(\r\n\r\n[https://amplab.cs.berkeley.edu/jenkins/view/RISELab%20Infra/job/spark-branch-2.1-test-maven-hadoop-2.7-ubuntu-testing/2/]\r\n{noformat}\r\n- daysToMillis and millisToDays *** FAILED ***\r\n 9131 did not equal 9130 Round trip of 9130 did not work in tz sun.util.calendar.ZoneInfo[id=\"Pacific/Enderbury\",offset=46800000,dstSavings=0,useDaylight=false,transitions=5,lastRule=null] (DateTimeUtilsSuite.scala:554){noformat}","created":"2018-08-09T22:20:01.203+0000"},{"body":"Roger that, will back-port back to 2.1 as best I can.","created":"2018-08-09T22:23:51.481+0000"},{"body":"word.  let's leave this open until for a bit longer as i continue to test.","created":"2018-08-09T22:26:02.088+0000"},{"body":"booyah!  i love watching the build queue pile up.  :)\r\n\r\nthanks [~srowen]!","created":"2018-08-09T22:36:22.960+0000"},{"body":"do we care about 2.0 and 1.6?","created":"2018-08-09T22:52:48.247+0000"}],"conversations":[{"body":"during my travails to port the spark builds to run on ubuntu 16.04LTS, i have encountered a strange and apparently java version-specific failure on *one* specific unit test.\r\n\r\nthe failure is here:\r\n\r\n[https://amplab.cs.berkeley.edu/jenkins/job/spark-master-test-sbt-hadoop-2.6-ubuntu-test/868/testReport/junit/org.apache.spark.sql.catalyst.util/DateTimeUtilsSuite/daysToMillis_and_millisToDays/]\r\n\r\nthe java version on this worker is:\r\n\r\nsknapp@ubuntu-testing:~$ java -version\r\n java version \"1.8.0_181\"\r\n Java(TM) SE Runtime Environment (build 1.8.0_181-b13)\r\n Java HotSpot(TM) 64-Bit Server VM (build 25.181-b13, mixed mode)\r\n\r\nhowever, when i run this exact build on the other ubuntu workers, it passes.  they systems are set up (for the most part) identically except for the java version:\r\n\r\nsknapp@amp-jenkins-staging-worker-02:~$ java -version\r\n java version \"1.8.0_171\"\r\n Java(TM) SE Runtime Environment (build 1.8.0_171-b11)\r\n Java HotSpot(TM) 64-Bit Server VM (build 25.171-b11, mixed mode)\r\n\r\nthere are some minor kernel and other package differences on these ubuntu workers, but nothing that (in my opinion) would affect this test.  i am willing to help investigate this, however.\r\n\r\nthe test also passes on the centos 6.9 workers, which have the following java version installed:\r\n\r\n[sknapp@amp-jenkins-worker-05 ~]$ java -version\r\njava version \"1.8.0_60\"\r\nJava(TM) SE Runtime Environment (build 1.8.0_60-b27)\r\nJava HotSpot(TM) 64-Bit Server VM (build 25.60-b23, mixed mode)my guess is that either:\r\n\r\nsql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/util/DateTimeUtils.scala\r\n\r\nor\r\n\r\nsql/catalyst/src/test/scala/org/apache/spark/sql/catalyst/util/DateTimeUtilsSuite.scala\r\n\r\nis doing something wrong.  i am not a scala expert by any means, so i'd really like some help in trying to un-block the project to port the builds to ubuntu.","from":"reporter","subject":"scala DateTimeUtilsSuite daysToMillis and millisToDays fails w/java 8 181-b13"},{"body":"one solution, of course, is to pin the java version on the upcoming ubuntu workers to one that passes this test, but things like this make build engineers like me die a little bit inside.","from":"developer"},{"body":"It's pretty clear this is down to differences in how time zones are defined, as they change over time and the JDK incorporates updated versions of the standard definitions in each release. \r\n\r\nIt looks like the difference between _171 and _181 is the difference between 2018c and 2018e in this table: http://www.oracle.com/technetwork/java/javase/tzdata-versions-138805.html\r\n\r\nNothing obviously relevant from Oracle's release notes. But I found this in the notes for 2018d:\r\n\r\n[http://mm.icann.org/pipermail/tz-announce/2018-March/000049.html]\r\n\r\n\"Enderbury and Kiritimati skipped New Year's Eve 1994, not New Year's Day 1995.  (Thanks to Kerry Shetline.)\"\r\n\r\nSo the answer is probably that the test has to be updated to reflect the fix to the timezone definition.\r\n\r\n \r\n\r\nOf course, if the test changes, it also starts failing on older Java 8 versions! probably not worth it.\r\n\r\nI'd suggest we resolve it by commenting this out with a note. There's no evidence this is a problem in Spark itself.","from":"developer"},{"body":"sgtm\r\n\r\ni also dug through the java release notes WRT timezone changes and didn't find anything (which i forgot to mention).  sorry about that!  :)\r\n\r\ni'll start by just commenting out the failing TZ (Pacific/Enderbury) and see if that works.\r\n\r\n ","from":"developer"},{"body":"User 'd80tb7' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21901","from":"developer"},{"body":"Hi,\r\n\r\n \r\n\r\njust to say that I looked at this and came to the same conclusion as Sean.  I've submitted a PR which excludes both New Years Eve and New Years day from the test- which should mean it will work on both old and new jvms.","from":"developer"},{"body":"testing this manually:\r\n\r\nhttps://amplab.cs.berkeley.edu/jenkins/job/spark-master-test-sbt-hadoop-2.6-ubuntu-test/878/","from":"developer"},{"body":"User 'srowen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21903","from":"developer"},{"body":"the spark-master-test-sbt build failed on ubuntu, but the the improtant bit is that the DateTimeUtilsSuite tests passed!\r\n\r\nback to unraveling R...  :\\\r\n\r\nthanks for the quick patch, [~d80tb7]!","from":"developer"},{"body":"Issue resolved by pull request 21901\n[https://github.com/apache/spark/pull/21901]","from":"developer"},{"body":"[~srowen] [~d80tb7] i'm thinking that we actually need to backport this change to previous branches.\r\n\r\n:(\r\n\r\n[https://amplab.cs.berkeley.edu/jenkins/view/RISELab%20Infra/job/spark-branch-2.1-test-maven-hadoop-2.7-ubuntu-testing/2/]\r\n{noformat}\r\n- daysToMillis and millisToDays *** FAILED ***\r\n 9131 did not equal 9130 Round trip of 9130 did not work in tz sun.util.calendar.ZoneInfo[id=\"Pacific/Enderbury\",offset=46800000,dstSavings=0,useDaylight=false,transitions=5,lastRule=null] (DateTimeUtilsSuite.scala:554){noformat}","from":"developer"},{"body":"Roger that, will back-port back to 2.1 as best I can.","from":"developer"},{"body":"word.  let's leave this open until for a bit longer as i continue to test.","from":"developer"},{"body":"booyah!  i love watching the build queue pile up.  :)\r\n\r\nthanks [~srowen]!","from":"developer"},{"body":"do we care about 2.0 and 1.6?","from":"developer"}],"created":"2018-07-27T18:20:36.000+0000","description":"during my travails to port the spark builds to run on ubuntu 16.04LTS, i have encountered a strange and apparently java version-specific failure on *one* specific unit test.\r\n\r\nthe failure is here:\r\n\r\n[https://amplab.cs.berkeley.edu/jenkins/job/spark-master-test-sbt-hadoop-2.6-ubuntu-test/868/testReport/junit/org.apache.spark.sql.catalyst.util/DateTimeUtilsSuite/daysToMillis_and_millisToDays/]\r\n\r\nthe java version on this worker is:\r\n\r\nsknapp@ubuntu-testing:~$ java -version\r\n java version \"1.8.0_181\"\r\n Java(TM) SE Runtime Environment (build 1.8.0_181-b13)\r\n Java HotSpot(TM) 64-Bit Server VM (build 25.181-b13, mixed mode)\r\n\r\nhowever, when i run this exact build on the other ubuntu workers, it passes.  they systems are set up (for the most part) identically except for the java version:\r\n\r\nsknapp@amp-jenkins-staging-worker-02:~$ java -version\r\n java version \"1.8.0_171\"\r\n Java(TM) SE Runtime Environment (build 1.8.0_171-b11)\r\n Java HotSpot(TM) 64-Bit Server VM (build 25.171-b11, mixed mode)\r\n\r\nthere are some minor kernel and other package differences on these ubuntu workers, but nothing that (in my opinion) would affect this test.  i am willing to help investigate this, however.\r\n\r\nthe test also passes on the centos 6.9 workers, which have the following java version installed:\r\n\r\n[sknapp@amp-jenkins-worker-05 ~]$ java -version\r\njava version \"1.8.0_60\"\r\nJava(TM) SE Runtime Environment (build 1.8.0_60-b27)\r\nJava HotSpot(TM) 64-Bit Server VM (build 25.60-b23, mixed mode)my guess is that either:\r\n\r\nsql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/util/DateTimeUtils.scala\r\n\r\nor\r\n\r\nsql/catalyst/src/test/scala/org/apache/spark/sql/catalyst/util/DateTimeUtilsSuite.scala\r\n\r\nis doing something wrong.  i am not a scala expert by any means, so i'd really like some help in trying to un-block the project to port the builds to ubuntu.","issue_id":"13175244","key":"SPARK-24950","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-08-09T22:42:14.000+0000","role":"fixed_distractor","summary":"scala DateTimeUtilsSuite daysToMillis and millisToDays fails w/java 8 181-b13"} {"case_id":"13176532","cluster":"DISTRACTOR-SPARK-25003","comments":[{"body":"User 'RussellSpitzer' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21988","created":"2018-08-03T15:58:05.891+0000"},{"body":"User 'RussellSpitzer' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21989","created":"2018-08-03T16:01:06.935+0000"},{"body":"User 'RussellSpitzer' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21990","created":"2018-08-03T16:08:04.380+0000"},{"body":"[~holden.karau] , Wrote up a PR for each branch target since I'm not sure what version you would think best for the update. Please let me know if you have any feedback or advice on how to get an automatic test in :)","created":"2018-08-03T16:08:14.872+0000"},{"body":"User 'RussellSpitzer' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21988","created":"2018-08-07T13:04:08.353+0000"},{"body":"User 'RussellSpitzer' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21989","created":"2018-08-07T13:04:10.603+0000"},{"body":"Issue resolved by pull request 21990\n[https://github.com/apache/spark/pull/21990]","created":"2018-10-18T04:29:41.131+0000"},{"body":"Hi [~hyukjin.kwon], [~rspitzer],   The extension points functionality is not available from pyspark. Would it be possible to get this fix on the v2.4 branch. Are there any issues with back porting it?  Thank you so much. ","created":"2019-04-01T18:52:16.483+0000"},{"body":"There was no interest in putting in OSS 2.4 and 2.2, but I did do this backport for the Datastax Distribution of Spark 2.4 and I can report it is a relatively simple and straightforward process.","created":"2019-04-01T19:47:48.567+0000"},{"body":"Thanks [~rspitzer].  That is good to know. \r\n\r\n[~hyukjin.kwon],  Do you have any concerns if we open a PR to back port to v2.4?","created":"2019-04-02T17:52:01.409+0000"},{"body":"This session stuff logic is a bit convoluted and many session changes were made .. I wouldn't backport it from 3.0 to 2.x unless it's quite serious one.","created":"2019-04-02T22:16:08.595+0000"},{"body":"`The extension points functionality is not available from pyspark` I believe it is a serious bug.\r\n\r\nI wonder what scenarios would be considered serious and backport it to 2.x ?","created":"2022-06-08T05:55:49.491+0000"}],"conversations":[{"body":"When creating a SparkSession here\r\n\r\n[https://github.com/apache/spark/blob/v2.2.2/python/pyspark/sql/session.py#L216]\r\n{code:python}\r\nif jsparkSession is None:\r\n jsparkSession = self._jvm.SparkSession(self._jsc.sc())\r\nself._jsparkSession = jsparkSession\r\n{code}\r\n\r\nI believe it ends up calling the constructor here\r\nhttps://github.com/apache/spark/blob/v2.2.2/sql/core/src/main/scala/org/apache/spark/sql/SparkSession.scala#L85-L87\r\n{code:scala}\r\n private[sql] def this(sc: SparkContext) {\r\n this(sc, None, None, new SparkSessionExtensions)\r\n }\r\n{code}\r\n\r\nWhich creates a new SparkSessionsExtensions object and does not pick up new extensions that could have been set in the config like the companion getOrCreate does.\r\nhttps://github.com/apache/spark/blob/v2.2.2/sql/core/src/main/scala/org/apache/spark/sql/SparkSession.scala#L928-L944\r\n{code:scala}\r\n//in getOrCreate\r\n // Initialize extensions if the user has defined a configurator class.\r\n val extensionConfOption = sparkContext.conf.get(StaticSQLConf.SPARK_SESSION_EXTENSIONS)\r\n if (extensionConfOption.isDefined) {\r\n val extensionConfClassName = extensionConfOption.get\r\n try {\r\n val extensionConfClass = Utils.classForName(extensionConfClassName)\r\n val extensionConf = extensionConfClass.newInstance()\r\n .asInstanceOf[SparkSessionExtensions => Unit]\r\n extensionConf(extensions)\r\n } catch {\r\n // Ignore the error if we cannot find the class or when the class has the wrong type.\r\n case e @ (_: ClassCastException |\r\n _: ClassNotFoundException |\r\n _: NoClassDefFoundError) =>\r\n logWarning(s\"Cannot use $extensionConfClassName to configure session extensions.\", e)\r\n }\r\n }\r\n{code}\r\n\r\nI think a quick fix would be to use the getOrCreate method from the companion object instead of calling the constructor from the SparkContext. Or we could fix this by ensuring that all constructors attempt to pick up custom extensions if they are set.","from":"reporter","subject":"Pyspark Does not use Spark Sql Extensions"},{"body":"User 'RussellSpitzer' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21988","from":"developer"},{"body":"User 'RussellSpitzer' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21989","from":"developer"},{"body":"User 'RussellSpitzer' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21990","from":"developer"},{"body":"[~holden.karau] , Wrote up a PR for each branch target since I'm not sure what version you would think best for the update. Please let me know if you have any feedback or advice on how to get an automatic test in :)","from":"developer"},{"body":"User 'RussellSpitzer' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21988","from":"developer"},{"body":"User 'RussellSpitzer' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/21989","from":"developer"},{"body":"Issue resolved by pull request 21990\n[https://github.com/apache/spark/pull/21990]","from":"developer"},{"body":"Hi [~hyukjin.kwon], [~rspitzer],   The extension points functionality is not available from pyspark. Would it be possible to get this fix on the v2.4 branch. Are there any issues with back porting it?  Thank you so much. ","from":"developer"},{"body":"There was no interest in putting in OSS 2.4 and 2.2, but I did do this backport for the Datastax Distribution of Spark 2.4 and I can report it is a relatively simple and straightforward process.","from":"developer"},{"body":"Thanks [~rspitzer].  That is good to know. \r\n\r\n[~hyukjin.kwon],  Do you have any concerns if we open a PR to back port to v2.4?","from":"developer"},{"body":"This session stuff logic is a bit convoluted and many session changes were made .. I wouldn't backport it from 3.0 to 2.x unless it's quite serious one.","from":"developer"},{"body":"`The extension points functionality is not available from pyspark` I believe it is a serious bug.\r\n\r\nI wonder what scenarios would be considered serious and backport it to 2.x ?","from":"developer"}],"created":"2018-08-02T18:00:47.000+0000","description":"When creating a SparkSession here\r\n\r\n[https://github.com/apache/spark/blob/v2.2.2/python/pyspark/sql/session.py#L216]\r\n{code:python}\r\nif jsparkSession is None:\r\n jsparkSession = self._jvm.SparkSession(self._jsc.sc())\r\nself._jsparkSession = jsparkSession\r\n{code}\r\n\r\nI believe it ends up calling the constructor here\r\nhttps://github.com/apache/spark/blob/v2.2.2/sql/core/src/main/scala/org/apache/spark/sql/SparkSession.scala#L85-L87\r\n{code:scala}\r\n private[sql] def this(sc: SparkContext) {\r\n this(sc, None, None, new SparkSessionExtensions)\r\n }\r\n{code}\r\n\r\nWhich creates a new SparkSessionsExtensions object and does not pick up new extensions that could have been set in the config like the companion getOrCreate does.\r\nhttps://github.com/apache/spark/blob/v2.2.2/sql/core/src/main/scala/org/apache/spark/sql/SparkSession.scala#L928-L944\r\n{code:scala}\r\n//in getOrCreate\r\n // Initialize extensions if the user has defined a configurator class.\r\n val extensionConfOption = sparkContext.conf.get(StaticSQLConf.SPARK_SESSION_EXTENSIONS)\r\n if (extensionConfOption.isDefined) {\r\n val extensionConfClassName = extensionConfOption.get\r\n try {\r\n val extensionConfClass = Utils.classForName(extensionConfClassName)\r\n val extensionConf = extensionConfClass.newInstance()\r\n .asInstanceOf[SparkSessionExtensions => Unit]\r\n extensionConf(extensions)\r\n } catch {\r\n // Ignore the error if we cannot find the class or when the class has the wrong type.\r\n case e @ (_: ClassCastException |\r\n _: ClassNotFoundException |\r\n _: NoClassDefFoundError) =>\r\n logWarning(s\"Cannot use $extensionConfClassName to configure session extensions.\", e)\r\n }\r\n }\r\n{code}\r\n\r\nI think a quick fix would be to use the getOrCreate method from the companion object instead of calling the constructor from the SparkContext. Or we could fix this by ensuring that all constructors attempt to pick up custom extensions if they are set.","issue_id":"13176532","key":"SPARK-25003","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-10-18T04:29:41.000+0000","role":"fixed_distractor","summary":"Pyspark Does not use Spark Sql Extensions"} {"case_id":"13181915","cluster":"DISTRACTOR-SPARK-25271","comments":[{"body":"cc [~cloud_fan]","created":"2018-08-29T13:10:33.373+0000"},{"body":"I think this is a known issue on Hive and Parquet, some context can be found at https://issues.apache.org/jira/browse/HIVE-11625.\r\n\r\nIt can be reproduced by:\r\n{code:java}\r\nsql(\"create table vp_reader STORED AS PARQUET as select map() as a\")\r\n18/08/30 00:07:15 ERROR DataWritableWriter: Parquet record is malformed: empty fields are illegal, the field should be ommited completely instead\r\nparquet.io.ParquetEncodingException: empty fields are illegal, the field should be ommited completely instead \r\n...\r\n{code}\r\nIf you don't store it as Parquet format, it can work:\r\n{code:java}\r\nsql(\"create table vp_reader STORED AS ORC as select map() as a\")\r\nsql(\"select * from vp_reader\").show\r\n+---+\r\n| a|\r\n+---+\r\n| []|\r\n+---+\r\n{code}","created":"2018-08-30T00:07:54.971+0000"},{"body":" [~viirya], It works fine in Spark 2.2.1 version. Below is the details\r\n{code:java}\r\nc:\\spark-2.2.1-bin-hadoop2.7>bin\\spark-shell\r\n Using Spark's default log4j profile: org/apache/spark/log4j-defaults.properties\r\n Setting default log level to \"WARN\".\r\n To adjust logging level use sc.setLogLevel(newLevel). For SparkR, use setLogLevel(newLevel).\r\n Spark context Web UI available at http://localhost:4040\r\n Spark context available as 'sc' (master = local[*], app id = local-1535611823064).\r\n Spark session available as 'spark'.\r\n Welcome to\r\n ____ __\r\n / _/_ ___ ____/ /_\r\n \\ \\/ _ \\/ _ `/ __/ '/\r\n /__/ ./_,// //_\\ version 2.2.1\r\n /_/\r\nUsing Scala version 2.11.8 (Java HotSpot(TM) 64-Bit Server VM, Java 1.8.0_60)\r\n Type in expressions to have them evaluated.\r\n Type :help for more information.\r\nscala> spark.sql(\"create table vp_reader_temp (projects map) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' COLLECTION ITEMS TERMINATED BY ':' MAP KEYS TERMINATED BY '$'\")\r\n res0: org.apache.spark.sql.DataFrame = []\r\nscala> spark.sql(\"LOAD DATA LOCAL INPATH 'parquetReader' INTO TABLE vp_reader_temp\")\r\n res1: org.apache.spark.sql.DataFrame = []\r\nscala> spark.sql(\"create table vp_reader STORED AS PARQUET as select * from vp_reader_temp\")\r\n res2: org.apache.spark.sql.DataFrame = []\r\nscala> spark.sql(\"select * from vp_reader\").collect\r\n res3: Array[org.apache.spark.sql.Row] = Array([Map(1 -> abc, 2 -> pqr, 3 -> xyz)], [Map()])\r\nscala>\r\n{code}\r\n ","created":"2018-08-30T07:17:42.081+0000"},{"body":"After further analyzing the issue i got following details\r\n\r\nIn  SingleDirectoryWriteTask private class(org.apache.spark.sql.execution.datasources.FileFormatWriter File) , currentWriter is  initialized with different outputWriter in spark-2.2.1 and spar-2.3.1, as shown below. \r\n\r\n\r\n{code:java}\r\nSpark-2.3.1= currentWriter is initilized with \"HiveOutputWriter\"\r\nSpark-2.2.1= currentWriter is initilized with \"ParquetOutputWriter\"\r\n{code}\r\n\r\n\r\nSo ParquetOutputWriter may be handling the null/empty values.","created":"2018-09-04T14:25:30.851+0000"},{"body":"[~cloud_fan] [~sowen]  Will this cause a compatibility problem compare to older version, If user has  null record ,then he is getting an exception with the current version where as the older version of spark(2.2.1)  wont throw any exception.\r\n\r\nI think the Output writers has been updated in the below PR\r\n\r\n[https://github.com/apache/spark/pull/20521]","created":"2018-09-04T14:41:27.110+0000"},{"body":"cc [~hyukjin.kwon]","created":"2018-09-04T14:41:48.736+0000"},{"body":"will take a look later but mind if I ask to elaborate\r\n\r\n{quote}\r\nI think the Output writers has been updated in the below PR\r\n\r\nhttps://github.com/apache/spark/pull/20521\r\n{quote}\r\n\r\n? Sounds rather a corner case but still a regression.","created":"2018-09-06T08:21:18.874+0000"},{"body":"As [~S71955] told, The Behaviour changed from the [https://github.com/apache/spark/pull/20521]\r\n\r\nWhile debugging \"spark.sql(\"create table vp_reader STORED AS PARQUET as select * from vp_reader_temp\")\"    I found following  details.\r\n \r\n{code:java}\r\nIn Spark-2.2.1 It is using the \"InsertIntoTable\"(org.apache.spark.sql.hive.execution.createHiveTableAsSelectCommand.run()) which will use ParquetFileFormat as shown below snaps{code}\r\n\r\n1 Figure: It uses \"InsertIntoTable\" for plan generation\r\n\r\n!image-2018-09-07-09-29-33-370.png! \r\n\r\n2 Figure: It is using the \"ParquetFileFormat\" as fileformat\r\n\r\n!image-2018-09-07-09-29-52-899.png! \r\n\r\n{code:java}\r\nBut in Spark-2.3.1, It is using the \"InsertIntoHiveTable\", (org.apache.spark.sql.hive.execution.createHiveTableAsSelectCommand.run()) which will use HiveFileFormat as shown below snap's{code}\r\n  \r\n3 Figure: It uses \"InsertIntoHiveTable\" for plan generation\r\n!image-2018-09-07-09-32-43-892.png! \r\n\r\n\r\n4 Figure: It is using the \"HiveFileFormat\" as fileformat\r\n!image-2018-09-07-09-33-03-095.png!\r\n\r\ncc [~hyukjin.kwon] Let me know any further clarification\r\n\r\n ","created":"2018-09-07T04:13:40.540+0000"},{"body":"Yeah, looks like after some changes, this kind of queries now uses Hive's record writer. So it inherits the issue in Hive.","created":"2018-09-12T03:37:50.057+0000"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22514","created":"2018-09-21T06:27:12.227+0000"},{"body":"Issue resolved by pull request 22514\n[https://github.com/apache/spark/pull/22514]","created":"2018-12-20T02:50:47.745+0000"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/30017","created":"2020-10-12T16:29:05.803+0000"},{"body":"[~viirya] [~apachespark] Hello, I did not have this problem in Spark2.4.0-CDH6.3.2 version, but this problem was repeated in Spark2.4.3 version, I do not understand why the lower version succeeded and the higher version failed, I would like to ask whether the fix of this bug does not support the 2.4.3 version? The following is my version information and error message:\r\n{code:java}\r\nspark version: 2.4.0-cdh6.3.2\r\nhive version: 2.1.1-cdh.6.3.2\r\nscala> spark.sql(\"create table test STORED AS PARQUET as select map() as a\")\r\nscala> sql(\"select * from test\").show\r\n+---+                                                                           \r\n|  a|\r\n+---+\r\n| []|\r\n+---+\r\n\r\n-----------------------------------------------------------------------------------------------------------------\r\nspark version: 2.4.3\r\nhive version: 3.1.2\r\nscala> spark.sql(\"create table test STORED AS PARQUET as select map() as a\")\r\n\r\nCaused by: org.apache.spark.SparkException: Task failed while writing rows.\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:257)\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:170)\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:169)\r\n  at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:90)\r\n  at org.apache.spark.scheduler.Task.run(Task.scala:121)\r\n  at org.apache.spark.executor.Executor$TaskRunner$$anonfun$10.apply(Executor.scala:408)\r\n  at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1360)\r\n  at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:414)\r\n  at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n  at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n  at java.lang.Thread.run(Thread.java:748)\r\nCaused by: java.lang.RuntimeException: Parquet record is malformed: empty fields are illegal, the field should be ommited completely instead\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.write(DataWritableWriter.java:64)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriteSupport.write(DataWritableWriteSupport.java:59)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriteSupport.write(DataWritableWriteSupport.java:31)\r\n  at parquet.hadoop.InternalParquetRecordWriter.write(InternalParquetRecordWriter.java:121)\r\n  at parquet.hadoop.ParquetRecordWriter.write(ParquetRecordWriter.java:123)\r\n  at parquet.hadoop.ParquetRecordWriter.write(ParquetRecordWriter.java:42)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.ParquetRecordWriterWrapper.write(ParquetRecordWriterWrapper.java:111)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.ParquetRecordWriterWrapper.write(ParquetRecordWriterWrapper.java:124)\r\n  at org.apache.spark.sql.hive.execution.HiveOutputWriter.write(HiveFileFormat.scala:149)\r\n  at org.apache.spark.sql.execution.datasources.SingleDirectoryDataWriter.write(FileFormatDataWriter.scala:137)\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:245)\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:242)\r\n  at org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1394)\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:248)\r\n  ... 10 more\r\nCaused by: parquet.io.ParquetEncodingException: empty fields are illegal, the field should be ommited completely instead\r\n  at parquet.io.MessageColumnIO$MessageColumnIORecordConsumer.endField(MessageColumnIO.java:244)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeMap(DataWritableWriter.java:241)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeValue(DataWritableWriter.java:116)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeGroupFields(DataWritableWriter.java:89)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.write(DataWritableWriter.java:60)\r\n  ... 23 more\r\n{code}\r\n ","created":"2022-02-18T02:58:58.812+0000"},{"body":"Based on this JIRA, we only have this fix since 2.4.8.\r\n\r\nI guess 2.4.0-cdh6.3.2 may backport the fix as this is a distribution maintained by the vendor, I don't know about the detail.","created":"2022-02-18T03:11:52.126+0000"}],"conversations":[{"body":"\r\n{code:java}\r\n 1)cat /data/parquet.dat\r\n\r\n1$abc2$pqr:3$xyz\r\nnull{code}\r\n \r\n\r\n\r\n{code:java}\r\n2)spark.sql(\"create table vp_reader_temp (projects map) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' COLLECTION ITEMS TERMINATED BY ':' MAP KEYS TERMINATED BY '$'\")\r\n{code}\r\n\r\n{code:java}\r\n3)spark.sql(\"\r\nLOAD DATA LOCAL INPATH '/data/parquet.dat' INTO TABLE vp_reader_temp\")\r\n{code}\r\n\r\n{code:java}\r\n4)spark.sql(\"create table vp_reader STORED AS PARQUET as select * from vp_reader_temp\")\r\n{code}\r\n\r\n\r\n*Result :* Throwing exception (Working fine with spark 2.2.1)\r\n\r\n{code:java}\r\njava.lang.RuntimeException: Parquet record is malformed: empty fields are illegal, the field should be ommited completely instead\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.write(DataWritableWriter.java:64)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriteSupport.write(DataWritableWriteSupport.java:59)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriteSupport.write(DataWritableWriteSupport.java:31)\r\n\tat org.apache.parquet.hadoop.InternalParquetRecordWriter.write(InternalParquetRecordWriter.java:123)\r\n\tat org.apache.parquet.hadoop.ParquetRecordWriter.write(ParquetRecordWriter.java:180)\r\n\tat org.apache.parquet.hadoop.ParquetRecordWriter.write(ParquetRecordWriter.java:46)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.ParquetRecordWriterWrapper.write(ParquetRecordWriterWrapper.java:112)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.ParquetRecordWriterWrapper.write(ParquetRecordWriterWrapper.java:125)\r\n\tat org.apache.spark.sql.hive.execution.HiveOutputWriter.write(HiveFileFormat.scala:149)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$SingleDirectoryWriteTask.execute(FileFormatWriter.scala:406)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:283)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:281)\r\n\tat org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1438)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:286)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:211)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:210)\r\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\r\n\tat org.apache.spark.scheduler.Task.run(Task.scala:109)\r\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:349)\r\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n\tat java.lang.Thread.run(Thread.java:745)\r\nCaused by: org.apache.parquet.io.ParquetEncodingException: empty fields are illegal, the field should be ommited completely instead\r\n\tat org.apache.parquet.io.MessageColumnIO$MessageColumnIORecordConsumer.endField(MessageColumnIO.java:320)\r\n\tat org.apache.parquet.io.RecordConsumerLoggingWrapper.endField(RecordConsumerLoggingWrapper.java:165)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeMap(DataWritableWriter.java:241)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeValue(DataWritableWriter.java:116)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeGroupFields(DataWritableWriter.java:89)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.write(DataWritableWriter.java:60)\r\n\t... 21 more\r\n{code}\r\n","from":"reporter","subject":"Creating parquet table with all the column null throws exception"},{"body":"cc [~cloud_fan]","from":"developer"},{"body":"I think this is a known issue on Hive and Parquet, some context can be found at https://issues.apache.org/jira/browse/HIVE-11625.\r\n\r\nIt can be reproduced by:\r\n{code:java}\r\nsql(\"create table vp_reader STORED AS PARQUET as select map() as a\")\r\n18/08/30 00:07:15 ERROR DataWritableWriter: Parquet record is malformed: empty fields are illegal, the field should be ommited completely instead\r\nparquet.io.ParquetEncodingException: empty fields are illegal, the field should be ommited completely instead \r\n...\r\n{code}\r\nIf you don't store it as Parquet format, it can work:\r\n{code:java}\r\nsql(\"create table vp_reader STORED AS ORC as select map() as a\")\r\nsql(\"select * from vp_reader\").show\r\n+---+\r\n| a|\r\n+---+\r\n| []|\r\n+---+\r\n{code}","from":"developer"},{"body":" [~viirya], It works fine in Spark 2.2.1 version. Below is the details\r\n{code:java}\r\nc:\\spark-2.2.1-bin-hadoop2.7>bin\\spark-shell\r\n Using Spark's default log4j profile: org/apache/spark/log4j-defaults.properties\r\n Setting default log level to \"WARN\".\r\n To adjust logging level use sc.setLogLevel(newLevel). For SparkR, use setLogLevel(newLevel).\r\n Spark context Web UI available at http://localhost:4040\r\n Spark context available as 'sc' (master = local[*], app id = local-1535611823064).\r\n Spark session available as 'spark'.\r\n Welcome to\r\n ____ __\r\n / _/_ ___ ____/ /_\r\n \\ \\/ _ \\/ _ `/ __/ '/\r\n /__/ ./_,// //_\\ version 2.2.1\r\n /_/\r\nUsing Scala version 2.11.8 (Java HotSpot(TM) 64-Bit Server VM, Java 1.8.0_60)\r\n Type in expressions to have them evaluated.\r\n Type :help for more information.\r\nscala> spark.sql(\"create table vp_reader_temp (projects map) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' COLLECTION ITEMS TERMINATED BY ':' MAP KEYS TERMINATED BY '$'\")\r\n res0: org.apache.spark.sql.DataFrame = []\r\nscala> spark.sql(\"LOAD DATA LOCAL INPATH 'parquetReader' INTO TABLE vp_reader_temp\")\r\n res1: org.apache.spark.sql.DataFrame = []\r\nscala> spark.sql(\"create table vp_reader STORED AS PARQUET as select * from vp_reader_temp\")\r\n res2: org.apache.spark.sql.DataFrame = []\r\nscala> spark.sql(\"select * from vp_reader\").collect\r\n res3: Array[org.apache.spark.sql.Row] = Array([Map(1 -> abc, 2 -> pqr, 3 -> xyz)], [Map()])\r\nscala>\r\n{code}\r\n ","from":"developer"},{"body":"After further analyzing the issue i got following details\r\n\r\nIn  SingleDirectoryWriteTask private class(org.apache.spark.sql.execution.datasources.FileFormatWriter File) , currentWriter is  initialized with different outputWriter in spark-2.2.1 and spar-2.3.1, as shown below. \r\n\r\n\r\n{code:java}\r\nSpark-2.3.1= currentWriter is initilized with \"HiveOutputWriter\"\r\nSpark-2.2.1= currentWriter is initilized with \"ParquetOutputWriter\"\r\n{code}\r\n\r\n\r\nSo ParquetOutputWriter may be handling the null/empty values.","from":"developer"},{"body":"[~cloud_fan] [~sowen]  Will this cause a compatibility problem compare to older version, If user has  null record ,then he is getting an exception with the current version where as the older version of spark(2.2.1)  wont throw any exception.\r\n\r\nI think the Output writers has been updated in the below PR\r\n\r\n[https://github.com/apache/spark/pull/20521]","from":"developer"},{"body":"cc [~hyukjin.kwon]","from":"developer"},{"body":"will take a look later but mind if I ask to elaborate\r\n\r\n{quote}\r\nI think the Output writers has been updated in the below PR\r\n\r\nhttps://github.com/apache/spark/pull/20521\r\n{quote}\r\n\r\n? Sounds rather a corner case but still a regression.","from":"developer"},{"body":"As [~S71955] told, The Behaviour changed from the [https://github.com/apache/spark/pull/20521]\r\n\r\nWhile debugging \"spark.sql(\"create table vp_reader STORED AS PARQUET as select * from vp_reader_temp\")\"    I found following  details.\r\n \r\n{code:java}\r\nIn Spark-2.2.1 It is using the \"InsertIntoTable\"(org.apache.spark.sql.hive.execution.createHiveTableAsSelectCommand.run()) which will use ParquetFileFormat as shown below snaps{code}\r\n\r\n1 Figure: It uses \"InsertIntoTable\" for plan generation\r\n\r\n!image-2018-09-07-09-29-33-370.png! \r\n\r\n2 Figure: It is using the \"ParquetFileFormat\" as fileformat\r\n\r\n!image-2018-09-07-09-29-52-899.png! \r\n\r\n{code:java}\r\nBut in Spark-2.3.1, It is using the \"InsertIntoHiveTable\", (org.apache.spark.sql.hive.execution.createHiveTableAsSelectCommand.run()) which will use HiveFileFormat as shown below snap's{code}\r\n  \r\n3 Figure: It uses \"InsertIntoHiveTable\" for plan generation\r\n!image-2018-09-07-09-32-43-892.png! \r\n\r\n\r\n4 Figure: It is using the \"HiveFileFormat\" as fileformat\r\n!image-2018-09-07-09-33-03-095.png!\r\n\r\ncc [~hyukjin.kwon] Let me know any further clarification\r\n\r\n ","from":"developer"},{"body":"Yeah, looks like after some changes, this kind of queries now uses Hive's record writer. So it inherits the issue in Hive.","from":"developer"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22514","from":"developer"},{"body":"Issue resolved by pull request 22514\n[https://github.com/apache/spark/pull/22514]","from":"developer"},{"body":"User 'viirya' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/30017","from":"developer"},{"body":"[~viirya] [~apachespark] Hello, I did not have this problem in Spark2.4.0-CDH6.3.2 version, but this problem was repeated in Spark2.4.3 version, I do not understand why the lower version succeeded and the higher version failed, I would like to ask whether the fix of this bug does not support the 2.4.3 version? The following is my version information and error message:\r\n{code:java}\r\nspark version: 2.4.0-cdh6.3.2\r\nhive version: 2.1.1-cdh.6.3.2\r\nscala> spark.sql(\"create table test STORED AS PARQUET as select map() as a\")\r\nscala> sql(\"select * from test\").show\r\n+---+                                                                           \r\n|  a|\r\n+---+\r\n| []|\r\n+---+\r\n\r\n-----------------------------------------------------------------------------------------------------------------\r\nspark version: 2.4.3\r\nhive version: 3.1.2\r\nscala> spark.sql(\"create table test STORED AS PARQUET as select map() as a\")\r\n\r\nCaused by: org.apache.spark.SparkException: Task failed while writing rows.\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:257)\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:170)\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:169)\r\n  at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:90)\r\n  at org.apache.spark.scheduler.Task.run(Task.scala:121)\r\n  at org.apache.spark.executor.Executor$TaskRunner$$anonfun$10.apply(Executor.scala:408)\r\n  at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1360)\r\n  at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:414)\r\n  at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n  at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n  at java.lang.Thread.run(Thread.java:748)\r\nCaused by: java.lang.RuntimeException: Parquet record is malformed: empty fields are illegal, the field should be ommited completely instead\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.write(DataWritableWriter.java:64)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriteSupport.write(DataWritableWriteSupport.java:59)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriteSupport.write(DataWritableWriteSupport.java:31)\r\n  at parquet.hadoop.InternalParquetRecordWriter.write(InternalParquetRecordWriter.java:121)\r\n  at parquet.hadoop.ParquetRecordWriter.write(ParquetRecordWriter.java:123)\r\n  at parquet.hadoop.ParquetRecordWriter.write(ParquetRecordWriter.java:42)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.ParquetRecordWriterWrapper.write(ParquetRecordWriterWrapper.java:111)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.ParquetRecordWriterWrapper.write(ParquetRecordWriterWrapper.java:124)\r\n  at org.apache.spark.sql.hive.execution.HiveOutputWriter.write(HiveFileFormat.scala:149)\r\n  at org.apache.spark.sql.execution.datasources.SingleDirectoryDataWriter.write(FileFormatDataWriter.scala:137)\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:245)\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:242)\r\n  at org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1394)\r\n  at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:248)\r\n  ... 10 more\r\nCaused by: parquet.io.ParquetEncodingException: empty fields are illegal, the field should be ommited completely instead\r\n  at parquet.io.MessageColumnIO$MessageColumnIORecordConsumer.endField(MessageColumnIO.java:244)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeMap(DataWritableWriter.java:241)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeValue(DataWritableWriter.java:116)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeGroupFields(DataWritableWriter.java:89)\r\n  at org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.write(DataWritableWriter.java:60)\r\n  ... 23 more\r\n{code}\r\n ","from":"developer"},{"body":"Based on this JIRA, we only have this fix since 2.4.8.\r\n\r\nI guess 2.4.0-cdh6.3.2 may backport the fix as this is a distribution maintained by the vendor, I don't know about the detail.","from":"developer"}],"created":"2018-08-29T13:08:50.000+0000","description":"\r\n{code:java}\r\n 1)cat /data/parquet.dat\r\n\r\n1$abc2$pqr:3$xyz\r\nnull{code}\r\n \r\n\r\n\r\n{code:java}\r\n2)spark.sql(\"create table vp_reader_temp (projects map) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' COLLECTION ITEMS TERMINATED BY ':' MAP KEYS TERMINATED BY '$'\")\r\n{code}\r\n\r\n{code:java}\r\n3)spark.sql(\"\r\nLOAD DATA LOCAL INPATH '/data/parquet.dat' INTO TABLE vp_reader_temp\")\r\n{code}\r\n\r\n{code:java}\r\n4)spark.sql(\"create table vp_reader STORED AS PARQUET as select * from vp_reader_temp\")\r\n{code}\r\n\r\n\r\n*Result :* Throwing exception (Working fine with spark 2.2.1)\r\n\r\n{code:java}\r\njava.lang.RuntimeException: Parquet record is malformed: empty fields are illegal, the field should be ommited completely instead\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.write(DataWritableWriter.java:64)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriteSupport.write(DataWritableWriteSupport.java:59)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriteSupport.write(DataWritableWriteSupport.java:31)\r\n\tat org.apache.parquet.hadoop.InternalParquetRecordWriter.write(InternalParquetRecordWriter.java:123)\r\n\tat org.apache.parquet.hadoop.ParquetRecordWriter.write(ParquetRecordWriter.java:180)\r\n\tat org.apache.parquet.hadoop.ParquetRecordWriter.write(ParquetRecordWriter.java:46)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.ParquetRecordWriterWrapper.write(ParquetRecordWriterWrapper.java:112)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.ParquetRecordWriterWrapper.write(ParquetRecordWriterWrapper.java:125)\r\n\tat org.apache.spark.sql.hive.execution.HiveOutputWriter.write(HiveFileFormat.scala:149)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$SingleDirectoryWriteTask.execute(FileFormatWriter.scala:406)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:283)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:281)\r\n\tat org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1438)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:286)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:211)\r\n\tat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:210)\r\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\r\n\tat org.apache.spark.scheduler.Task.run(Task.scala:109)\r\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:349)\r\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n\tat java.lang.Thread.run(Thread.java:745)\r\nCaused by: org.apache.parquet.io.ParquetEncodingException: empty fields are illegal, the field should be ommited completely instead\r\n\tat org.apache.parquet.io.MessageColumnIO$MessageColumnIORecordConsumer.endField(MessageColumnIO.java:320)\r\n\tat org.apache.parquet.io.RecordConsumerLoggingWrapper.endField(RecordConsumerLoggingWrapper.java:165)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeMap(DataWritableWriter.java:241)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeValue(DataWritableWriter.java:116)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.writeGroupFields(DataWritableWriter.java:89)\r\n\tat org.apache.hadoop.hive.ql.io.parquet.write.DataWritableWriter.write(DataWritableWriter.java:60)\r\n\t... 21 more\r\n{code}\r\n","issue_id":"13181915","key":"SPARK-25271","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-12-20T02:50:47.000+0000","role":"fixed_distractor","summary":"Creating parquet table with all the column null throws exception"} {"case_id":"13187505","cluster":"DISTRACTOR-SPARK-25538","comments":[{"body":"Please do not use Blocker and Critical when reporting issues as they are reserved for committer. Though, I agree this should be a blocker for 2.4.0 as it is a correctness issue. cc [~cloud_fan]","created":"2018-09-26T07:44:10.399+0000"},{"body":"Hi [~Steven Rand], is it possible to upload the data file? We need to be able to reproduce it in order to debug it, thanks! I'll set it as a blocker once I confirm this is a correctness issue.","created":"2018-09-26T08:06:50.884+0000"},{"body":"Hi [~cloud_fan], I'm still trying to create a scrubbed version of the file. It's proving to be difficult since mutating the original DataFrame often prevents the issue from reproducing (e.g., the example in the description where calling {{withColumnRenamed}} before {{distinct.count}} leads to the correct result being printed).","created":"2018-09-26T08:58:37.613+0000"},{"body":"cc [~kiszk] as well","created":"2018-09-26T13:12:47.907+0000"},{"body":"Hi [~Steven Rand], would it be possible to share the schema of this DataFrame?\r\n","created":"2018-09-26T19:04:56.106+0000"},{"body":"[~kiszk], yes, the schema is:\r\n\r\n \r\n{code}\r\nscala> spark.read.parquet(\"hdfs:///data\").printSchema\r\nroot\r\n |-- col_0: string (nullable = true)\r\n |-- col_1: timestamp (nullable = true)\r\n |-- col_2: string (nullable = true)\r\n |-- col_3: timestamp (nullable = true)\r\n |-- col_4: string (nullable = true)\r\n |-- col_5: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_6: string (nullable = true)\r\n |-- col_7: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_8: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_9: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_10: string (nullable = true)\r\n |-- col_11: timestamp (nullable = true)\r\n |-- col_12: integer (nullable = true)\r\n |-- col_13: boolean (nullable = true)\r\n |-- col_14: decimal(38,18) (nullable = true)\r\n |-- col_15: long (nullable = true)\r\n |-- col_16: string (nullable = true)\r\n |-- col_17: integer (nullable = true)\r\n |-- col_18: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_19: string (nullable = true)\r\n |-- col_20: string (nullable = true)\r\n |-- col_21: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_22: string (nullable = true)\r\n |-- col_23: array (nullable = true)\r\n | |-- element: timestamp (containsNull = true)\r\n |-- col_24: string (nullable = true)\r\n |-- col_25: string (nullable = true)\r\n |-- col_26: string (nullable = true)\r\n |-- col_27: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_28: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_29: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_30: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_31: decimal(38,18) (nullable = true)\r\n |-- col_32: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_33: string (nullable = true)\r\n |-- col_34: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_35: decimal(38,18) (nullable = true)\r\n |-- col_36: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_37: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_38: decimal(38,18) (nullable = true)\r\n |-- col_39: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_40: string (nullable = true)\r\n |-- col_41: string (nullable = true)\r\n |-- col_42: string (nullable = true)\r\n |-- col_43: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_44: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_45: string (nullable = true)\r\n |-- col_46: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_47: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_48: string (nullable = true)\r\n |-- col_49: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_50: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_51: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_52: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_53: string (nullable = true)\r\n |-- col_54: decimal(38,18) (nullable = true)\r\n |-- col_55: decimal(38,18) (nullable = true)\r\n |-- col_56: decimal(38,18) (nullable = true)\r\n |-- col_57: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n{code}","created":"2018-09-27T00:17:46.246+0000"},{"body":"Thank for upload a schema. While I looked at the schema, I am still not sure about the reason of this problem.\r\nI would appreciate it if you could find a good input data that can reproduce a problem.","created":"2018-09-28T09:29:06.538+0000"},{"body":"[~kiszk] that makes sense, I'll try to do so. The issue I've been having so far is that when I run the UDF I've written to change the data (while preserving number of duplicate rows), the resulting DataFrame doesn't reproduce the issue.","created":"2018-09-29T02:26:27.643+0000"},{"body":"[~kiszk] I've uploaded a tarball containing parquet files that reproduce the issue but don't contain any of the values in the original dataset. Specifically, some columns have been dropped, all strings have been changed to \"test_string\", all values in col_50 have been changed to 0.0043, and the values in col_14 have all been mapped from their original values to values between 0.001 and 0.0044.\r\n\r\nThis new DataFrame still reproduces issues similar to those in the description:\r\n{code:java}\r\nscala> df.distinct.count\r\nres3: Long = 64\r\n\r\nscala> df.sort(\"col_0\").distinct.count\r\nres4: Long = 73\r\n\r\nscala> df.withColumnRenamed(\"col_0\", \"new\").distinct.count\r\nres5: Long = 63\r\n{code}\r\nI get those inconsistent/wrong results on {{2.4.0-rc2}} and if I check out commit {{a7c19d9c21d59fd0109a7078c80b33d3da03fafd}}, which is SPARK-23713. If I check out the commit immediately before, which is {{fe2b7a4568d65a62da6e6eb00fff05f248b4332c}}, then all three commands return 63.\r\n\r\ncc [~cloud_fan] – IMO this should block the 2.4.0 release.","created":"2018-09-30T05:56:52.515+0000"},{"body":"Thank you. I will check it tonight in Japan.","created":"2018-10-01T01:52:36.417+0000"},{"body":"I was able to reproduce also using limit instead of sort:\r\n{code}\r\nscala> df.limit(80).distinct.count\r\nres83: Long = 72\r\n\r\nscala> df.distinct.count\r\nres84: Long = 64\r\n\r\nscala> df.limit(20).distinct.count\r\nres88: Long = 20\r\n\r\nscala> df.limit(20).distinct.collect.distinct.length\r\nres89: Int = 17\r\n{code}","created":"2018-10-01T14:41:47.154+0000"},{"body":"This test case does not print {{63}} using master branch.\r\n\r\n{code}\r\n test(\"test2\") {\r\n val df = spark.read.parquet(\"file:///SPARK-25538-repro\")\r\n val c1 = df.distinct.count\r\n val c2 = df.sort(\"col_0\").distinct.count\r\n val c3 = df.withColumnRenamed(\"col_0\", \"new\").distinct.count\r\n val c0 = df.count\r\n print(s\"c1=$c1, c2=$c2, c3=$c3, c0=$c0\\n\")\r\n }\r\n\r\nc1=64, c2=73, c3=64, c0=123\r\n{code}","created":"2018-10-01T17:13:44.486+0000"},{"body":"[~mgaido]'s PR, https://github.com/apache/spark/pull/22602, fixes the decimal issue and looks reasonable.","created":"2018-10-01T18:18:11.300+0000"},{"body":"User 'mgaido91' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22602","created":"2018-10-01T19:45:09.329+0000"},{"body":"User 'mgaido91' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22602","created":"2018-10-01T19:46:03.325+0000"},{"body":"Issue resolved by pull request 22602\n[https://github.com/apache/spark/pull/22602]","created":"2018-10-03T14:29:38.796+0000"},{"body":"Thanks all!","created":"2018-10-03T23:47:15.121+0000"}],"conversations":[{"body":"It appears that {{df.distinct.count}} can return incorrect values after SPARK-23713. It's possible that other operations are affected as well; {{distinct}} just happens to be the one that we noticed. I believe that this issue was introduced by SPARK-23713 because I can't reproduce it until that commit, and I've been able to reproduce it after that commit as well as with {{tags/v2.4.0-rc1}}. \r\n\r\nBelow are example spark-shell sessions to illustrate the problem. Unfortunately the data used in these examples can't be uploaded to this Jira ticket. I'll try to create test data which also reproduces the issue, and will upload that if I'm able to do so.\r\n\r\nExample from Spark 2.3.1, which behaves correctly:\r\n\r\n{code}\r\nscala> val df = spark.read.parquet(\"hdfs:///data\")\r\ndf: org.apache.spark.sql.DataFrame = []\r\n\r\nscala> df.count\r\nres0: Long = 123\r\n\r\nscala> df.distinct.count\r\nres1: Long = 115\r\n{code}\r\n\r\nExample from Spark 2.4.0-rc1, which returns different output:\r\n\r\n{code}\r\nscala> val df = spark.read.parquet(\"hdfs:///data\")\r\ndf: org.apache.spark.sql.DataFrame = []\r\n\r\nscala> df.count\r\nres0: Long = 123\r\n\r\nscala> df.distinct.count\r\nres1: Long = 116\r\n\r\nscala> df.sort(\"col_0\").distinct.count\r\nres2: Long = 123\r\n\r\nscala> df.withColumnRenamed(\"col_0\", \"newName\").distinct.count\r\nres3: Long = 115\r\n{code}","from":"reporter","subject":"incorrect row counts after distinct()"},{"body":"Please do not use Blocker and Critical when reporting issues as they are reserved for committer. Though, I agree this should be a blocker for 2.4.0 as it is a correctness issue. cc [~cloud_fan]","from":"developer"},{"body":"Hi [~Steven Rand], is it possible to upload the data file? We need to be able to reproduce it in order to debug it, thanks! I'll set it as a blocker once I confirm this is a correctness issue.","from":"developer"},{"body":"Hi [~cloud_fan], I'm still trying to create a scrubbed version of the file. It's proving to be difficult since mutating the original DataFrame often prevents the issue from reproducing (e.g., the example in the description where calling {{withColumnRenamed}} before {{distinct.count}} leads to the correct result being printed).","from":"developer"},{"body":"cc [~kiszk] as well","from":"developer"},{"body":"Hi [~Steven Rand], would it be possible to share the schema of this DataFrame?\r\n","from":"developer"},{"body":"[~kiszk], yes, the schema is:\r\n\r\n \r\n{code}\r\nscala> spark.read.parquet(\"hdfs:///data\").printSchema\r\nroot\r\n |-- col_0: string (nullable = true)\r\n |-- col_1: timestamp (nullable = true)\r\n |-- col_2: string (nullable = true)\r\n |-- col_3: timestamp (nullable = true)\r\n |-- col_4: string (nullable = true)\r\n |-- col_5: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_6: string (nullable = true)\r\n |-- col_7: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_8: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_9: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_10: string (nullable = true)\r\n |-- col_11: timestamp (nullable = true)\r\n |-- col_12: integer (nullable = true)\r\n |-- col_13: boolean (nullable = true)\r\n |-- col_14: decimal(38,18) (nullable = true)\r\n |-- col_15: long (nullable = true)\r\n |-- col_16: string (nullable = true)\r\n |-- col_17: integer (nullable = true)\r\n |-- col_18: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_19: string (nullable = true)\r\n |-- col_20: string (nullable = true)\r\n |-- col_21: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_22: string (nullable = true)\r\n |-- col_23: array (nullable = true)\r\n | |-- element: timestamp (containsNull = true)\r\n |-- col_24: string (nullable = true)\r\n |-- col_25: string (nullable = true)\r\n |-- col_26: string (nullable = true)\r\n |-- col_27: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_28: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_29: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_30: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_31: decimal(38,18) (nullable = true)\r\n |-- col_32: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_33: string (nullable = true)\r\n |-- col_34: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_35: decimal(38,18) (nullable = true)\r\n |-- col_36: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_37: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_38: decimal(38,18) (nullable = true)\r\n |-- col_39: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_40: string (nullable = true)\r\n |-- col_41: string (nullable = true)\r\n |-- col_42: string (nullable = true)\r\n |-- col_43: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_44: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_45: string (nullable = true)\r\n |-- col_46: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_47: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_48: string (nullable = true)\r\n |-- col_49: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_50: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_51: array (nullable = true)\r\n | |-- element: string (containsNull = true)\r\n |-- col_52: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n |-- col_53: string (nullable = true)\r\n |-- col_54: decimal(38,18) (nullable = true)\r\n |-- col_55: decimal(38,18) (nullable = true)\r\n |-- col_56: decimal(38,18) (nullable = true)\r\n |-- col_57: array (nullable = true)\r\n | |-- element: decimal(38,18) (containsNull = true)\r\n{code}","from":"developer"},{"body":"Thank for upload a schema. While I looked at the schema, I am still not sure about the reason of this problem.\r\nI would appreciate it if you could find a good input data that can reproduce a problem.","from":"developer"},{"body":"[~kiszk] that makes sense, I'll try to do so. The issue I've been having so far is that when I run the UDF I've written to change the data (while preserving number of duplicate rows), the resulting DataFrame doesn't reproduce the issue.","from":"developer"},{"body":"[~kiszk] I've uploaded a tarball containing parquet files that reproduce the issue but don't contain any of the values in the original dataset. Specifically, some columns have been dropped, all strings have been changed to \"test_string\", all values in col_50 have been changed to 0.0043, and the values in col_14 have all been mapped from their original values to values between 0.001 and 0.0044.\r\n\r\nThis new DataFrame still reproduces issues similar to those in the description:\r\n{code:java}\r\nscala> df.distinct.count\r\nres3: Long = 64\r\n\r\nscala> df.sort(\"col_0\").distinct.count\r\nres4: Long = 73\r\n\r\nscala> df.withColumnRenamed(\"col_0\", \"new\").distinct.count\r\nres5: Long = 63\r\n{code}\r\nI get those inconsistent/wrong results on {{2.4.0-rc2}} and if I check out commit {{a7c19d9c21d59fd0109a7078c80b33d3da03fafd}}, which is SPARK-23713. If I check out the commit immediately before, which is {{fe2b7a4568d65a62da6e6eb00fff05f248b4332c}}, then all three commands return 63.\r\n\r\ncc [~cloud_fan] – IMO this should block the 2.4.0 release.","from":"developer"},{"body":"Thank you. I will check it tonight in Japan.","from":"developer"},{"body":"I was able to reproduce also using limit instead of sort:\r\n{code}\r\nscala> df.limit(80).distinct.count\r\nres83: Long = 72\r\n\r\nscala> df.distinct.count\r\nres84: Long = 64\r\n\r\nscala> df.limit(20).distinct.count\r\nres88: Long = 20\r\n\r\nscala> df.limit(20).distinct.collect.distinct.length\r\nres89: Int = 17\r\n{code}","from":"developer"},{"body":"This test case does not print {{63}} using master branch.\r\n\r\n{code}\r\n test(\"test2\") {\r\n val df = spark.read.parquet(\"file:///SPARK-25538-repro\")\r\n val c1 = df.distinct.count\r\n val c2 = df.sort(\"col_0\").distinct.count\r\n val c3 = df.withColumnRenamed(\"col_0\", \"new\").distinct.count\r\n val c0 = df.count\r\n print(s\"c1=$c1, c2=$c2, c3=$c3, c0=$c0\\n\")\r\n }\r\n\r\nc1=64, c2=73, c3=64, c0=123\r\n{code}","from":"developer"},{"body":"[~mgaido]'s PR, https://github.com/apache/spark/pull/22602, fixes the decimal issue and looks reasonable.","from":"developer"},{"body":"User 'mgaido91' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22602","from":"developer"},{"body":"User 'mgaido91' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22602","from":"developer"},{"body":"Issue resolved by pull request 22602\n[https://github.com/apache/spark/pull/22602]","from":"developer"},{"body":"Thanks all!","from":"developer"}],"created":"2018-09-26T05:53:26.000+0000","description":"It appears that {{df.distinct.count}} can return incorrect values after SPARK-23713. It's possible that other operations are affected as well; {{distinct}} just happens to be the one that we noticed. I believe that this issue was introduced by SPARK-23713 because I can't reproduce it until that commit, and I've been able to reproduce it after that commit as well as with {{tags/v2.4.0-rc1}}. \r\n\r\nBelow are example spark-shell sessions to illustrate the problem. Unfortunately the data used in these examples can't be uploaded to this Jira ticket. I'll try to create test data which also reproduces the issue, and will upload that if I'm able to do so.\r\n\r\nExample from Spark 2.3.1, which behaves correctly:\r\n\r\n{code}\r\nscala> val df = spark.read.parquet(\"hdfs:///data\")\r\ndf: org.apache.spark.sql.DataFrame = []\r\n\r\nscala> df.count\r\nres0: Long = 123\r\n\r\nscala> df.distinct.count\r\nres1: Long = 115\r\n{code}\r\n\r\nExample from Spark 2.4.0-rc1, which returns different output:\r\n\r\n{code}\r\nscala> val df = spark.read.parquet(\"hdfs:///data\")\r\ndf: org.apache.spark.sql.DataFrame = []\r\n\r\nscala> df.count\r\nres0: Long = 123\r\n\r\nscala> df.distinct.count\r\nres1: Long = 116\r\n\r\nscala> df.sort(\"col_0\").distinct.count\r\nres2: Long = 123\r\n\r\nscala> df.withColumnRenamed(\"col_0\", \"newName\").distinct.count\r\nres3: Long = 115\r\n{code}","issue_id":"13187505","key":"SPARK-25538","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-10-03T14:29:38.000+0000","role":"fixed_distractor","summary":"incorrect row counts after distinct()"} {"case_id":"13190500","cluster":"DISTRACTOR-SPARK-25694","comments":[{"body":"Please avoid to set target versions which are usually reserved for committers.","created":"2018-10-12T01:42:09.762+0000"},{"body":"Got it, thanks [~hyukjin.kwon] for the suggestion!","created":"2018-10-12T20:51:03.991+0000"},{"body":"[~boyangwa], do you know of a workaround for this? I'm facing the same issue.","created":"2018-12-14T18:14:19.748+0000"},{"body":"[~howardatwork],\r\n\r\n \r\n\r\nI have hit this problem as well, seems no workaround without change in httpclient.  I ever used scalaj.http, but replaced it httpcomponent\r\n\r\n ","created":"2019-03-14T01:16:25.485+0000"},{"body":"Although SPARK-12868 adds `setURLStreamHandlerFactory` for `ADD JARS` commands, this method can be called at most once in a given Java Virtual Machine. So, there is another issue with this.\r\n- https://docs.oracle.com/en/java/javase/11/docs/api/java.base/java/net/URL.html#setURLStreamHandlerFactory(java.net.URLStreamHandlerFactory)","created":"2019-11-01T19:13:16.806+0000"},{"body":"Issue resolved by pull request 26530\n[https://github.com/apache/spark/pull/26530]","created":"2019-11-18T05:46:21.639+0000"},{"body":"i have a question than why FsUrlStreamHandlerFactory will deal with the http schema url. from the code , when protocal is http,method createURLStreamHandler in FsUrlStreamHandlerFactory will return null. ","created":"2020-03-02T03:20:29.784+0000"}],"conversations":[{"body":"URL.setURLStreamHandlerFactory() in SharedState causes URL.openConnection() returns FsUrlConnection object, which is not compatible with HttpURLConnection. This will cause exception when using some third party http library (e.g. scalaj.http).\r\n\r\nThe following code in Spark 2.3.0 introduced the issue: sql/core/src/main/scala/org/apache/spark/sql/internal/SharedState.scala:\r\n{code}\r\nobject SharedState extends Logging {   ...   \r\n URL.setURLStreamHandlerFactory(new FsUrlStreamHandlerFactory())   ...\r\n}\r\n{code}\r\n\r\nHere is the example exception when using scalaj.http in Spark:\r\n{code}\r\n StackTrace: scala.MatchError: org.apache.hadoop.fs.FsUrlConnection:[http://wwww.example.com|http://wwww.example.com/] (of class org.apache.hadoop.fs.FsUrlConnection)\r\n at scalaj.http.HttpRequest.scalaj$http$HttpRequest$$doConnection(Http.scala:343)\r\n at scalaj.http.HttpRequest.exec(Http.scala:335)\r\n at scalaj.http.HttpRequest.asString(Http.scala:455)\r\n{code}\r\n  \r\nOne option to fix the issue is to return null in URLStreamHandlerFactory.createURLStreamHandler when the protocol is http/https, so it will use the default behavior and be compatible with scalaj.http. Following is the code example:\r\n\r\n{code}\r\nclass SparkUrlStreamHandlerFactory extends URLStreamHandlerFactory with Logging {\r\n\r\n private val fsUrlStreamHandlerFactory = new FsUrlStreamHandlerFactory()\r\n\r\n override def createURLStreamHandler(protocol: String): URLStreamHandler = {\r\n val handler = fsUrlStreamHandlerFactory.createURLStreamHandler(protocol)\r\n if (handler == null) {\r\n return null\r\n }\r\n\r\n if (protocol != null &&\r\n (protocol.equalsIgnoreCase(\"http\")\r\n || protocol.equalsIgnoreCase(\"https\"))) {\r\n // return null to use system default URLStreamHandler\r\n null\r\n } else {\r\n handler\r\n }\r\n }\r\n}\r\n{code}\r\n\r\nI would like to get some discussion here before submitting a pull request.\r\n","from":"reporter","subject":"URL.setURLStreamHandlerFactory causing incompatible HttpURLConnection issue"},{"body":"Please avoid to set target versions which are usually reserved for committers.","from":"developer"},{"body":"Got it, thanks [~hyukjin.kwon] for the suggestion!","from":"developer"},{"body":"[~boyangwa], do you know of a workaround for this? I'm facing the same issue.","from":"developer"},{"body":"[~howardatwork],\r\n\r\n \r\n\r\nI have hit this problem as well, seems no workaround without change in httpclient.  I ever used scalaj.http, but replaced it httpcomponent\r\n\r\n ","from":"developer"},{"body":"Although SPARK-12868 adds `setURLStreamHandlerFactory` for `ADD JARS` commands, this method can be called at most once in a given Java Virtual Machine. So, there is another issue with this.\r\n- https://docs.oracle.com/en/java/javase/11/docs/api/java.base/java/net/URL.html#setURLStreamHandlerFactory(java.net.URLStreamHandlerFactory)","from":"developer"},{"body":"Issue resolved by pull request 26530\n[https://github.com/apache/spark/pull/26530]","from":"developer"},{"body":"i have a question than why FsUrlStreamHandlerFactory will deal with the http schema url. from the code , when protocal is http,method createURLStreamHandler in FsUrlStreamHandlerFactory will return null. ","from":"developer"}],"created":"2018-10-09T22:20:15.000+0000","description":"URL.setURLStreamHandlerFactory() in SharedState causes URL.openConnection() returns FsUrlConnection object, which is not compatible with HttpURLConnection. This will cause exception when using some third party http library (e.g. scalaj.http).\r\n\r\nThe following code in Spark 2.3.0 introduced the issue: sql/core/src/main/scala/org/apache/spark/sql/internal/SharedState.scala:\r\n{code}\r\nobject SharedState extends Logging {   ...   \r\n URL.setURLStreamHandlerFactory(new FsUrlStreamHandlerFactory())   ...\r\n}\r\n{code}\r\n\r\nHere is the example exception when using scalaj.http in Spark:\r\n{code}\r\n StackTrace: scala.MatchError: org.apache.hadoop.fs.FsUrlConnection:[http://wwww.example.com|http://wwww.example.com/] (of class org.apache.hadoop.fs.FsUrlConnection)\r\n at scalaj.http.HttpRequest.scalaj$http$HttpRequest$$doConnection(Http.scala:343)\r\n at scalaj.http.HttpRequest.exec(Http.scala:335)\r\n at scalaj.http.HttpRequest.asString(Http.scala:455)\r\n{code}\r\n  \r\nOne option to fix the issue is to return null in URLStreamHandlerFactory.createURLStreamHandler when the protocol is http/https, so it will use the default behavior and be compatible with scalaj.http. Following is the code example:\r\n\r\n{code}\r\nclass SparkUrlStreamHandlerFactory extends URLStreamHandlerFactory with Logging {\r\n\r\n private val fsUrlStreamHandlerFactory = new FsUrlStreamHandlerFactory()\r\n\r\n override def createURLStreamHandler(protocol: String): URLStreamHandler = {\r\n val handler = fsUrlStreamHandlerFactory.createURLStreamHandler(protocol)\r\n if (handler == null) {\r\n return null\r\n }\r\n\r\n if (protocol != null &&\r\n (protocol.equalsIgnoreCase(\"http\")\r\n || protocol.equalsIgnoreCase(\"https\"))) {\r\n // return null to use system default URLStreamHandler\r\n null\r\n } else {\r\n handler\r\n }\r\n }\r\n}\r\n{code}\r\n\r\nI would like to get some discussion here before submitting a pull request.\r\n","issue_id":"13190500","key":"SPARK-25694","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-11-18T05:46:21.000+0000","role":"fixed_distractor","summary":"URL.setURLStreamHandlerFactory causing incompatible HttpURLConnection issue"} {"case_id":"13193716","cluster":"DISTRACTOR-SPARK-25816","comments":[{"body":"User 'peter-toth' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22817","created":"2018-10-24T19:04:57.125+0000"},{"body":"Here is another reproduce that should be related to this same issue:\r\n\r\nval v0 = sqlContext.read.avro(\"final_allDatatypes_Spark.avro\");\r\nval v00 = v0.toDF(v0.schema.fields.indices.view.map(\"\" + _):_*)\r\nval v001 = v00.select($\"0\".as(\"0\"), $\"1\".as(\"1\"),$\"2\".as(\"2\"),$\"3\".as(\"3\"),$\"4\".as(\"4\"),$\"5\".as(\"5\"),$\"6\".as(\"6\"),$\"7\".as(\"7\"),$\"8\".as(\"8\"))\r\nval v013 = $\"8\"\r\nval v010 = map(v013, v013)\r\n \r\nv001.where(map(v013, v010)(v013)(v013)===\"dummy\")\r\n\r\n \r\n\r\norg.apache.spark.sql.AnalysisException: Reference '8' is ambiguous, could be: 8, 8.;\r\n at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolve(LogicalPlan.scala:213)\r\n at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolveChildren(LogicalPlan.scala:97)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$36.apply(Analyzer.scala:822)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$36.apply(Analyzer.scala:824)\r\n at org.apache.spark.sql.catalyst.analysis.package$.withPosition(package.scala:53)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$.org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve(Analyzer.scala:821)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve$2.apply(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve$2.apply(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$4.apply(TreeNode.scala:306)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:187)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapChildren(TreeNode.scala:304)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$.org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve$2.apply(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve$2.apply(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$4.apply(TreeNode.scala:306)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:187)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapChildren(TreeNode.scala:304)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$.org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$apply$9$$anonfun$applyOrElse$36.apply(Analyzer.scala:891)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$apply$9$$anonfun$applyOrElse$36.apply(Analyzer.scala:891)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$1.apply(QueryPlan.scala:107)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$1.apply(QueryPlan.scala:107)\r\n at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:70)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpression$1(QueryPlan.scala:106)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan.org$apache$spark$sql$catalyst$plans$QueryPlan$$recursiveTransform$1(QueryPlan.scala:118)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$2.apply(QueryPlan.scala:127)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:187)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan.mapExpressions(QueryPlan.scala:127)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$apply$9.applyOrElse(Analyzer.scala:891)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$apply$9.applyOrElse(Analyzer.scala:833)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformUp$1.apply(TreeNode.scala:289)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformUp$1.apply(TreeNode.scala:289)\r\n at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:70)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.transformUp(TreeNode.scala:288)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:286)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:286)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$4.apply(TreeNode.scala:306)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:187)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapChildren(TreeNode.scala:304)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.transformUp(TreeNode.scala:286)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$.apply(Analyzer.scala:833)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$.apply(Analyzer.scala:690)\r\n at org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:87)\r\n at org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:84)\r\n at scala.collection.LinearSeqOptimized$class.foldLeft(LinearSeqOptimized.scala:124)\r\n at scala.collection.immutable.List.foldLeft(List.scala:84)\r\n at org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:84)\r\n at org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:76)\r\n at scala.collection.immutable.List.foreach(List.scala:381)\r\n at org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:76)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer.org$apache$spark$sql$catalyst$analysis$Analyzer$$executeSameContext(Analyzer.scala:124)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer.execute(Analyzer.scala:118)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer.executeAndCheck(Analyzer.scala:103)\r\n at org.apache.spark.sql.execution.QueryExecution.analyzed$lzycompute(QueryExecution.scala:57)\r\n at org.apache.spark.sql.execution.QueryExecution.analyzed(QueryExecution.scala:55)\r\n at org.apache.spark.sql.execution.QueryExecution.assertAnalyzed(QueryExecution.scala:47)\r\n at org.apache.spark.sql.Dataset.(Dataset.scala:172)\r\n at org.apache.spark.sql.Dataset.(Dataset.scala:178)\r\n at org.apache.spark.sql.Dataset$.apply(Dataset.scala:65)\r\n at org.apache.spark.sql.Dataset.withTypedPlan(Dataset.scala:3301)\r\n at org.apache.spark.sql.Dataset.filter(Dataset.scala:1458)\r\n at org.apache.spark.sql.Dataset.where(Dataset.scala:1486)\r\n\r\n ","created":"2018-10-26T20:09:13.600+0000"},{"body":"Thanks [~bzhang], It seems both are regressions from 2.2 to 2.3 for the same reason. My submitted PR fixes them.","created":"2018-10-27T12:10:20.361+0000"},{"body":"Thank you, [~bzhang] and [~petertoth]. I also confirmed that this is a bug in 2.3.x and 2.4.0 RC4 (and master). Thanks to [~petertoth], it looks like we can have the fix this for 2.3.3 and 2.4.0 RC5.\r\n\r\ncc [~cloud_fan]","created":"2018-10-28T19:06:35.514+0000"},{"body":"Thanks [~petertoth] and everyone else for the quick fix.","created":"2018-10-29T18:10:35.499+0000"}],"conversations":[{"body":"When there is a duplicate column name in the current Dataframe and orginal Dataframe where current df is selected from, Spark in 2.3.0 and 2.3.1 does not resolve the column correctly when using it in the expression, hence causing casting issue. The same code is working in Spark 2.2.1\r\n\r\nPlease see below code to reproduce the issue\r\n\r\nimport org.apache.spark._\r\nimport org.apache.spark.rdd._\r\nimport org.apache.spark.storage.StorageLevel._\r\nimport org.apache.spark.sql._\r\nimport org.apache.spark.sql.DataFrame\r\nimport org.apache.spark.sql.types._\r\nimport org.apache.spark.sql.functions._\r\nimport org.apache.spark.sql.catalyst.expressions._\r\nimport org.apache.spark.sql.Column\r\n\r\nval v0 = spark.read.parquet(\"/data/home/bzinfa/bz/source.snappy.parquet\")\r\nval v00 = v0.toDF(v0.schema.fields.indices.view.map(\"\" + _):_*)\r\nval v5 = v00.select($\"13\".as(\"0\"),$\"14\".as(\"1\"),$\"15\".as(\"2\"))\r\nval v5_2 = $\"2\"\r\nv5.where(lit(500).<(v5_2(new Column(new MapKeys(v5_2.expr))(lit(0)))))\r\n\r\n//v00's 3rdcolumn is binary and 16th is map\r\n\r\nError:\r\norg.apache.spark.sql.AnalysisException: cannot resolve 'map_keys(`2`)' due to data type mismatch: argument 1 requires map type, however, '`2`' is of binary type.;\r\n \r\n 'Project [0#1591, 1#1592, 2#1593] +- 'Filter (500 < {color:#FF0000}2#1593{color}[map_keys({color:#FF0000}2#1561{color})[0]]) +- Project [13#1572 AS 0#1591, 14#1573 AS 1#1592, 15#1574 AS 2#1593, 2#1561] +- Project [c_bytes#1527 AS 0#1559, c_union#1528 AS 1#1560, c_fixed#1529 AS 2#1561, c_boolean#1530 AS 3#1562, c_float#1531 AS 4#1563, c_double#1532 AS 5#1564, c_int#1533 AS 6#1565, c_long#1534L AS 7#1566L, c_string#1535 AS 8#1567, c_decimal_18_2#1536 AS 9#1568, c_decimal_28_2#1537 AS 10#1569, c_decimal_38_2#1538 AS 11#1570, c_date#1539 AS 12#1571, simple_struct#1540 AS 13#1572, simple_array#1541 AS 14#1573, simple_map#1542 AS 15#1574] +- Relation[c_bytes#1527,c_union#1528,c_fixed#1529,c_boolean#1530,c_float#1531,c_double#1532,c_int#1533,c_long#1534L,c_string#1535,c_decimal_18_2#1536,c_decimal_28_2#1537,c_decimal_38_2#1538,c_date#1539,simple_struct#1540,simple_array#1541,simple_map#1542] parquet","from":"reporter","subject":"Functions does not resolve Columns correctly"},{"body":"User 'peter-toth' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/22817","from":"developer"},{"body":"Here is another reproduce that should be related to this same issue:\r\n\r\nval v0 = sqlContext.read.avro(\"final_allDatatypes_Spark.avro\");\r\nval v00 = v0.toDF(v0.schema.fields.indices.view.map(\"\" + _):_*)\r\nval v001 = v00.select($\"0\".as(\"0\"), $\"1\".as(\"1\"),$\"2\".as(\"2\"),$\"3\".as(\"3\"),$\"4\".as(\"4\"),$\"5\".as(\"5\"),$\"6\".as(\"6\"),$\"7\".as(\"7\"),$\"8\".as(\"8\"))\r\nval v013 = $\"8\"\r\nval v010 = map(v013, v013)\r\n \r\nv001.where(map(v013, v010)(v013)(v013)===\"dummy\")\r\n\r\n \r\n\r\norg.apache.spark.sql.AnalysisException: Reference '8' is ambiguous, could be: 8, 8.;\r\n at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolve(LogicalPlan.scala:213)\r\n at org.apache.spark.sql.catalyst.plans.logical.LogicalPlan.resolveChildren(LogicalPlan.scala:97)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$36.apply(Analyzer.scala:822)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$36.apply(Analyzer.scala:824)\r\n at org.apache.spark.sql.catalyst.analysis.package$.withPosition(package.scala:53)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$.org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve(Analyzer.scala:821)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve$2.apply(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve$2.apply(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$4.apply(TreeNode.scala:306)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:187)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapChildren(TreeNode.scala:304)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$.org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve$2.apply(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve$2.apply(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$4.apply(TreeNode.scala:306)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:187)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapChildren(TreeNode.scala:304)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$.org$apache$spark$sql$catalyst$analysis$Analyzer$ResolveReferences$$resolve(Analyzer.scala:830)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$apply$9$$anonfun$applyOrElse$36.apply(Analyzer.scala:891)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$apply$9$$anonfun$applyOrElse$36.apply(Analyzer.scala:891)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$1.apply(QueryPlan.scala:107)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$1.apply(QueryPlan.scala:107)\r\n at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:70)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan.transformExpression$1(QueryPlan.scala:106)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan.org$apache$spark$sql$catalyst$plans$QueryPlan$$recursiveTransform$1(QueryPlan.scala:118)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan$$anonfun$2.apply(QueryPlan.scala:127)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:187)\r\n at org.apache.spark.sql.catalyst.plans.QueryPlan.mapExpressions(QueryPlan.scala:127)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$apply$9.applyOrElse(Analyzer.scala:891)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$$anonfun$apply$9.applyOrElse(Analyzer.scala:833)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformUp$1.apply(TreeNode.scala:289)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$transformUp$1.apply(TreeNode.scala:289)\r\n at org.apache.spark.sql.catalyst.trees.CurrentOrigin$.withOrigin(TreeNode.scala:70)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.transformUp(TreeNode.scala:288)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:286)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$3.apply(TreeNode.scala:286)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$4.apply(TreeNode.scala:306)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapProductIterator(TreeNode.scala:187)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.mapChildren(TreeNode.scala:304)\r\n at org.apache.spark.sql.catalyst.trees.TreeNode.transformUp(TreeNode.scala:286)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$.apply(Analyzer.scala:833)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer$ResolveReferences$.apply(Analyzer.scala:690)\r\n at org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:87)\r\n at org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1$$anonfun$apply$1.apply(RuleExecutor.scala:84)\r\n at scala.collection.LinearSeqOptimized$class.foldLeft(LinearSeqOptimized.scala:124)\r\n at scala.collection.immutable.List.foldLeft(List.scala:84)\r\n at org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:84)\r\n at org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$execute$1.apply(RuleExecutor.scala:76)\r\n at scala.collection.immutable.List.foreach(List.scala:381)\r\n at org.apache.spark.sql.catalyst.rules.RuleExecutor.execute(RuleExecutor.scala:76)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer.org$apache$spark$sql$catalyst$analysis$Analyzer$$executeSameContext(Analyzer.scala:124)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer.execute(Analyzer.scala:118)\r\n at org.apache.spark.sql.catalyst.analysis.Analyzer.executeAndCheck(Analyzer.scala:103)\r\n at org.apache.spark.sql.execution.QueryExecution.analyzed$lzycompute(QueryExecution.scala:57)\r\n at org.apache.spark.sql.execution.QueryExecution.analyzed(QueryExecution.scala:55)\r\n at org.apache.spark.sql.execution.QueryExecution.assertAnalyzed(QueryExecution.scala:47)\r\n at org.apache.spark.sql.Dataset.(Dataset.scala:172)\r\n at org.apache.spark.sql.Dataset.(Dataset.scala:178)\r\n at org.apache.spark.sql.Dataset$.apply(Dataset.scala:65)\r\n at org.apache.spark.sql.Dataset.withTypedPlan(Dataset.scala:3301)\r\n at org.apache.spark.sql.Dataset.filter(Dataset.scala:1458)\r\n at org.apache.spark.sql.Dataset.where(Dataset.scala:1486)\r\n\r\n ","from":"developer"},{"body":"Thanks [~bzhang], It seems both are regressions from 2.2 to 2.3 for the same reason. My submitted PR fixes them.","from":"developer"},{"body":"Thank you, [~bzhang] and [~petertoth]. I also confirmed that this is a bug in 2.3.x and 2.4.0 RC4 (and master). Thanks to [~petertoth], it looks like we can have the fix this for 2.3.3 and 2.4.0 RC5.\r\n\r\ncc [~cloud_fan]","from":"developer"},{"body":"Thanks [~petertoth] and everyone else for the quick fix.","from":"developer"}],"created":"2018-10-23T23:41:19.000+0000","description":"When there is a duplicate column name in the current Dataframe and orginal Dataframe where current df is selected from, Spark in 2.3.0 and 2.3.1 does not resolve the column correctly when using it in the expression, hence causing casting issue. The same code is working in Spark 2.2.1\r\n\r\nPlease see below code to reproduce the issue\r\n\r\nimport org.apache.spark._\r\nimport org.apache.spark.rdd._\r\nimport org.apache.spark.storage.StorageLevel._\r\nimport org.apache.spark.sql._\r\nimport org.apache.spark.sql.DataFrame\r\nimport org.apache.spark.sql.types._\r\nimport org.apache.spark.sql.functions._\r\nimport org.apache.spark.sql.catalyst.expressions._\r\nimport org.apache.spark.sql.Column\r\n\r\nval v0 = spark.read.parquet(\"/data/home/bzinfa/bz/source.snappy.parquet\")\r\nval v00 = v0.toDF(v0.schema.fields.indices.view.map(\"\" + _):_*)\r\nval v5 = v00.select($\"13\".as(\"0\"),$\"14\".as(\"1\"),$\"15\".as(\"2\"))\r\nval v5_2 = $\"2\"\r\nv5.where(lit(500).<(v5_2(new Column(new MapKeys(v5_2.expr))(lit(0)))))\r\n\r\n//v00's 3rdcolumn is binary and 16th is map\r\n\r\nError:\r\norg.apache.spark.sql.AnalysisException: cannot resolve 'map_keys(`2`)' due to data type mismatch: argument 1 requires map type, however, '`2`' is of binary type.;\r\n \r\n 'Project [0#1591, 1#1592, 2#1593] +- 'Filter (500 < {color:#FF0000}2#1593{color}[map_keys({color:#FF0000}2#1561{color})[0]]) +- Project [13#1572 AS 0#1591, 14#1573 AS 1#1592, 15#1574 AS 2#1593, 2#1561] +- Project [c_bytes#1527 AS 0#1559, c_union#1528 AS 1#1560, c_fixed#1529 AS 2#1561, c_boolean#1530 AS 3#1562, c_float#1531 AS 4#1563, c_double#1532 AS 5#1564, c_int#1533 AS 6#1565, c_long#1534L AS 7#1566L, c_string#1535 AS 8#1567, c_decimal_18_2#1536 AS 9#1568, c_decimal_28_2#1537 AS 10#1569, c_decimal_38_2#1538 AS 11#1570, c_date#1539 AS 12#1571, simple_struct#1540 AS 13#1572, simple_array#1541 AS 14#1573, simple_map#1542 AS 15#1574] +- Relation[c_bytes#1527,c_union#1528,c_fixed#1529,c_boolean#1530,c_float#1531,c_double#1532,c_int#1533,c_long#1534L,c_string#1535,c_decimal_18_2#1536,c_decimal_28_2#1537,c_decimal_38_2#1538,c_date#1539,simple_struct#1540,simple_array#1541,simple_map#1542] parquet","issue_id":"13193716","key":"SPARK-25816","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-10-29T00:57:05.000+0000","role":"fixed_distractor","summary":"Functions does not resolve Columns correctly"} {"case_id":"13198776","cluster":"DISTRACTOR-SPARK-26084","comments":[{"body":"/cc [~maropu] [~hvanhovell] who worked on the PR that may have caused this problem","created":"2018-11-15T22:59:36.914+0000"},{"body":"[~simeons] since you have already propose a solution, do you mind opening a PR.","created":"2018-11-16T11:37:41.214+0000"},{"body":"[~hvanhovell] done [https://github.com/apache/spark/pull/23075]","created":"2018-11-18T01:44:39.439+0000"},{"body":"User 'ssimeonov' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/23075","created":"2018-11-18T01:45:46.962+0000"},{"body":"User 'ssimeonov' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/23075","created":"2018-11-18T01:45:59.772+0000"}],"conversations":[{"body":"[SPARK-18394|https://issues.apache.org/jira/browse/SPARK-18394] introduced a stable ordering in {{AttributeSet.toSeq}} using expression IDs ([PR-18959|https://github.com/apache/spark/pull/18959/files#diff-75576f0ec7f9d8b5032000245217d233R128]) without noticing that {{AggregateExpression.references}} used {{AttributeSet.toSeq}} as a shortcut ([link|https://github.com/apache/spark/blob/5264164a67df498b73facae207eda12ee133be7d/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/aggregate/interfaces.scala#L132]). The net result is that {{AggregateExpression.references}} fails for unresolved aggregate functions.\r\n\r\n{code:scala}\r\norg.apache.spark.sql.catalyst.expressions.aggregate.AggregateExpression(\r\n org.apache.spark.sql.catalyst.expressions.aggregate.Sum(('x + 'y).expr),\r\n mode = org.apache.spark.sql.catalyst.expressions.aggregate.Complete,\r\n isDistinct = false\r\n).references\r\n{code}\r\n\r\nfails with\r\n\r\n{code:scala}\r\norg.apache.spark.sql.catalyst.analysis.UnresolvedException: Invalid call to exprId on unresolved object, tree: 'y\r\n\tat org.apache.spark.sql.catalyst.analysis.UnresolvedAttribute.exprId(unresolved.scala:104)\r\n\tat org.apache.spark.sql.catalyst.expressions.AttributeSet$$anonfun$toSeq$2.apply(AttributeSet.scala:128)\r\n\tat org.apache.spark.sql.catalyst.expressions.AttributeSet$$anonfun$toSeq$2.apply(AttributeSet.scala:128)\r\n\tat scala.math.Ordering$$anon$5.compare(Ordering.scala:122)\r\n\tat java.util.TimSort.countRunAndMakeAscending(TimSort.java:355)\r\n\tat java.util.TimSort.sort(TimSort.java:220)\r\n\tat java.util.Arrays.sort(Arrays.java:1438)\r\n\tat scala.collection.SeqLike$class.sorted(SeqLike.scala:648)\r\n\tat scala.collection.AbstractSeq.sorted(Seq.scala:41)\r\n\tat scala.collection.SeqLike$class.sortBy(SeqLike.scala:623)\r\n\tat scala.collection.AbstractSeq.sortBy(Seq.scala:41)\r\n\tat org.apache.spark.sql.catalyst.expressions.AttributeSet.toSeq(AttributeSet.scala:128)\r\n\tat org.apache.spark.sql.catalyst.expressions.aggregate.AggregateExpression.references(interfaces.scala:201)\r\n{code}\r\n\r\nThe solution is to avoid calling {{toSeq}} as ordering is not important in {{references}} and simplify (and speed up) the implementation to something like\r\n\r\n{code:scala}\r\nmode match {\r\n case Partial | Complete => aggregateFunction.references\r\n case PartialMerge | Final => AttributeSet(aggregateFunction.aggBufferAttributes)\r\n}\r\n{code}","from":"reporter","subject":"AggregateExpression.references fails on unresolved expression trees"},{"body":"/cc [~maropu] [~hvanhovell] who worked on the PR that may have caused this problem","from":"developer"},{"body":"[~simeons] since you have already propose a solution, do you mind opening a PR.","from":"developer"},{"body":"[~hvanhovell] done [https://github.com/apache/spark/pull/23075]","from":"developer"},{"body":"User 'ssimeonov' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/23075","from":"developer"},{"body":"User 'ssimeonov' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/23075","from":"developer"}],"created":"2018-11-15T22:56:04.000+0000","description":"[SPARK-18394|https://issues.apache.org/jira/browse/SPARK-18394] introduced a stable ordering in {{AttributeSet.toSeq}} using expression IDs ([PR-18959|https://github.com/apache/spark/pull/18959/files#diff-75576f0ec7f9d8b5032000245217d233R128]) without noticing that {{AggregateExpression.references}} used {{AttributeSet.toSeq}} as a shortcut ([link|https://github.com/apache/spark/blob/5264164a67df498b73facae207eda12ee133be7d/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/aggregate/interfaces.scala#L132]). The net result is that {{AggregateExpression.references}} fails for unresolved aggregate functions.\r\n\r\n{code:scala}\r\norg.apache.spark.sql.catalyst.expressions.aggregate.AggregateExpression(\r\n org.apache.spark.sql.catalyst.expressions.aggregate.Sum(('x + 'y).expr),\r\n mode = org.apache.spark.sql.catalyst.expressions.aggregate.Complete,\r\n isDistinct = false\r\n).references\r\n{code}\r\n\r\nfails with\r\n\r\n{code:scala}\r\norg.apache.spark.sql.catalyst.analysis.UnresolvedException: Invalid call to exprId on unresolved object, tree: 'y\r\n\tat org.apache.spark.sql.catalyst.analysis.UnresolvedAttribute.exprId(unresolved.scala:104)\r\n\tat org.apache.spark.sql.catalyst.expressions.AttributeSet$$anonfun$toSeq$2.apply(AttributeSet.scala:128)\r\n\tat org.apache.spark.sql.catalyst.expressions.AttributeSet$$anonfun$toSeq$2.apply(AttributeSet.scala:128)\r\n\tat scala.math.Ordering$$anon$5.compare(Ordering.scala:122)\r\n\tat java.util.TimSort.countRunAndMakeAscending(TimSort.java:355)\r\n\tat java.util.TimSort.sort(TimSort.java:220)\r\n\tat java.util.Arrays.sort(Arrays.java:1438)\r\n\tat scala.collection.SeqLike$class.sorted(SeqLike.scala:648)\r\n\tat scala.collection.AbstractSeq.sorted(Seq.scala:41)\r\n\tat scala.collection.SeqLike$class.sortBy(SeqLike.scala:623)\r\n\tat scala.collection.AbstractSeq.sortBy(Seq.scala:41)\r\n\tat org.apache.spark.sql.catalyst.expressions.AttributeSet.toSeq(AttributeSet.scala:128)\r\n\tat org.apache.spark.sql.catalyst.expressions.aggregate.AggregateExpression.references(interfaces.scala:201)\r\n{code}\r\n\r\nThe solution is to avoid calling {{toSeq}} as ordering is not important in {{references}} and simplify (and speed up) the implementation to something like\r\n\r\n{code:scala}\r\nmode match {\r\n case Partial | Complete => aggregateFunction.references\r\n case PartialMerge | Final => AttributeSet(aggregateFunction.aggBufferAttributes)\r\n}\r\n{code}","issue_id":"13198776","key":"SPARK-26084","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2018-11-20T20:59:56.000+0000","role":"fixed_distractor","summary":"AggregateExpression.references fails on unresolved expression trees"} {"case_id":"13208308","cluster":"DISTRACTOR-SPARK-26572","comments":[{"body":"I checked I could reproduce this below and I set \"correctness\" in the Label.\r\nIt seems HashAggregate hits a bug if it has stateful expressions, e.g., monotonically_increasing_id, rand, ...\r\n{code:java}\r\nscala> val baseTable = Seq((1), (1)).toDF(\"idx\")\r\nscala> val distinctWithId = baseTable.distinct.withColumn(\"id\", functions.monotonically_increasing_id())\r\nscala> baseTable.join(distinctWithId, \"idx\").show\r\n+---+------------+ \r\n|idx| id|\r\n+---+------------+\r\n| 1|369367187456|\r\n| 1|369367187457|\r\n+---+------------+\r\n\r\nscala> sql(\"SET spark.sql.codegen.wholeStage=false\")\r\nscala> baseTable.join(distinctWithId, \"idx\").show\r\n+---+------------+\r\n|idx| id|\r\n+---+------------+\r\n| 1|369367187456|\r\n| 1|369367187456|\r\n+---+------------+\r\n{code}\r\nThis is pretty a corner case, so I didn't set a blocker.","created":"2019-02-04T11:34:16.060+0000"},{"body":"Resolved by https://github.com/apache/spark/pull/23731","created":"2019-02-15T01:05:18.993+0000"}],"conversations":[{"body":"When joining a table with projected monotonically_increasing_id column after calling distinct with another table the operators do not get executed in the right order. \r\n\r\nHere is a minimal example:\r\n{code:java}\r\nimport org.apache.spark.sql.{DataFrame, SparkSession, functions}\r\n\r\nobject JoinBug extends App {\r\n\r\n // Spark session setup\r\n val session = SparkSession.builder().master(\"local[*]\").getOrCreate()\r\n import session.sqlContext.implicits._\r\n session.sparkContext.setLogLevel(\"error\")\r\n\r\n // Bug in Spark: \"monotonically_increasing_id\" is pushed down when it shouldn't be. Push down only happens when the\r\n // DF containing the \"monotonically_increasing_id\" expression is on the left side of the join.\r\n val baseTable = Seq((1), (1)).toDF(\"idx\")\r\n val distinctWithId = baseTable.distinct.withColumn(\"id\", functions.monotonically_increasing_id())\r\n val monotonicallyOnRight: DataFrame = baseTable.join(distinctWithId, \"idx\")\r\n val monotonicallyOnLeft: DataFrame = distinctWithId.join(baseTable, \"idx\")\r\n\r\n monotonicallyOnLeft.show // Wrong\r\n monotonicallyOnRight.show // Ok in Spark 2.2.2 - also wrong in Spark 2.4.0\r\n\r\n}\r\n\r\n{code}\r\nIt produces the following output:\r\n{code:java}\r\nWrong:\r\n+---+------------+\r\n|idx| id |\r\n+---+------------+\r\n| 1|369367187456 |\r\n| 1|369367187457 |\r\n+---+------------+\r\n\r\nRight:\r\n+---+------------+\r\n|idx| id |\r\n+---+------------+\r\n| 1|369367187456 |\r\n| 1|369367187456 |\r\n+---+------------+\r\n{code}\r\nWe assume that the join operator triggers a pushdown of expressions (monotonically_increasing_id in this case) which gets pushed down to be executed before distinct. This produces non-distinct rows with unique id's. However it seems like this behavior only appears if the table with the projected expression is on the left side of the join in Spark 2.2.2 (for version 2.4.0 it fails on both joins).","from":"reporter","subject":"Join on distinct column with monotonically_increasing_id produces wrong output"},{"body":"I checked I could reproduce this below and I set \"correctness\" in the Label.\r\nIt seems HashAggregate hits a bug if it has stateful expressions, e.g., monotonically_increasing_id, rand, ...\r\n{code:java}\r\nscala> val baseTable = Seq((1), (1)).toDF(\"idx\")\r\nscala> val distinctWithId = baseTable.distinct.withColumn(\"id\", functions.monotonically_increasing_id())\r\nscala> baseTable.join(distinctWithId, \"idx\").show\r\n+---+------------+ \r\n|idx| id|\r\n+---+------------+\r\n| 1|369367187456|\r\n| 1|369367187457|\r\n+---+------------+\r\n\r\nscala> sql(\"SET spark.sql.codegen.wholeStage=false\")\r\nscala> baseTable.join(distinctWithId, \"idx\").show\r\n+---+------------+\r\n|idx| id|\r\n+---+------------+\r\n| 1|369367187456|\r\n| 1|369367187456|\r\n+---+------------+\r\n{code}\r\nThis is pretty a corner case, so I didn't set a blocker.","from":"developer"},{"body":"Resolved by https://github.com/apache/spark/pull/23731","from":"developer"}],"created":"2019-01-08T13:16:39.000+0000","description":"When joining a table with projected monotonically_increasing_id column after calling distinct with another table the operators do not get executed in the right order. \r\n\r\nHere is a minimal example:\r\n{code:java}\r\nimport org.apache.spark.sql.{DataFrame, SparkSession, functions}\r\n\r\nobject JoinBug extends App {\r\n\r\n // Spark session setup\r\n val session = SparkSession.builder().master(\"local[*]\").getOrCreate()\r\n import session.sqlContext.implicits._\r\n session.sparkContext.setLogLevel(\"error\")\r\n\r\n // Bug in Spark: \"monotonically_increasing_id\" is pushed down when it shouldn't be. Push down only happens when the\r\n // DF containing the \"monotonically_increasing_id\" expression is on the left side of the join.\r\n val baseTable = Seq((1), (1)).toDF(\"idx\")\r\n val distinctWithId = baseTable.distinct.withColumn(\"id\", functions.monotonically_increasing_id())\r\n val monotonicallyOnRight: DataFrame = baseTable.join(distinctWithId, \"idx\")\r\n val monotonicallyOnLeft: DataFrame = distinctWithId.join(baseTable, \"idx\")\r\n\r\n monotonicallyOnLeft.show // Wrong\r\n monotonicallyOnRight.show // Ok in Spark 2.2.2 - also wrong in Spark 2.4.0\r\n\r\n}\r\n\r\n{code}\r\nIt produces the following output:\r\n{code:java}\r\nWrong:\r\n+---+------------+\r\n|idx| id |\r\n+---+------------+\r\n| 1|369367187456 |\r\n| 1|369367187457 |\r\n+---+------------+\r\n\r\nRight:\r\n+---+------------+\r\n|idx| id |\r\n+---+------------+\r\n| 1|369367187456 |\r\n| 1|369367187456 |\r\n+---+------------+\r\n{code}\r\nWe assume that the join operator triggers a pushdown of expressions (monotonically_increasing_id in this case) which gets pushed down to be executed before distinct. This produces non-distinct rows with unique id's. However it seems like this behavior only appears if the table with the projected expression is on the left side of the join in Spark 2.2.2 (for version 2.4.0 it fails on both joins).","issue_id":"13208308","key":"SPARK-26572","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-02-15T01:05:18.000+0000","role":"fixed_distractor","summary":"Join on distinct column with monotonically_increasing_id produces wrong output"} {"case_id":"13214170","cluster":"DISTRACTOR-SPARK-26836","comments":[{"body":"Thank you for reporting, [~treff7es]. \r\nDid you see this issue before at Spark 2.3.x with Databricks Avro libraries?\r\n\r\ncc [~Gengliang.Wang]","created":"2019-02-08T08:08:49.163+0000"},{"body":"I just tried with Spark 2.3.1 and the same issue. I did not specify any Avro library because it is a Hive Avro table originally. I guess it is using Hive's Avro serde for serialization/deserialization (or maybe I'm not right :)).","created":"2019-02-08T09:21:38.152+0000"},{"body":"In the meantime I checked if using the same hive version in Spark would solve this issue by setting:\r\n\r\n \r\n{code:java}\r\nspark.sql.hive.metastore.version 2.3.3\r\nspark.sql.hive.metastore.jars /etc/hadoop/conf:/usr/lib/hadoop/lib/*:/usr/lib/hadoop/.//*:/usr/lib/hadoop-hdfs/./:/usr/lib/hadoop-hdfs/lib/*:/usr/lib/hadoop-hdfs/.//*:/usr/lib/hadoop-yarn/lib/*:/usr/lib/hadoop-yarn/.//*:/usr/lib/hadoop-mapreduce/lib/*:/usr/lib/hadoop-mapreduce/.//*::/etc/tez/conf:/usr/lib/tez/*:/usr/lib/tez/lib/*:/usr/lib/hadoop-lzo/lib/*:/usr/share/aws/aws-java-sdk/*:/usr/share/aws/emr/emrfs/conf:/usr/share/aws/emr/emrfs/lib/*:/usr/share/aws/emr/emrfs/auxlib/*:/usr/share/aws/emr/ddb/lib/emr-ddb-hadoop.jar:/usr/share/aws/emr/goodies/lib/emr-hadoop-goodies.jar:/usr/share/aws/emr/kinesis/lib/emr-kinesis-hadoop.jar:/usr/share/aws/emr/cloudwatch-sink/lib/*:/usr/share/aws/emr/security/conf:/usr/share/aws/emr/security/lib/*:/usr/lib/hive/lib/* \r\n{code}\r\nThe issue still exists regardless the hive metastore version spark is using.","created":"2019-02-08T14:13:46.232+0000"},{"body":"[~dongjoon]I don't think the issue is related to spark-avro lib.\r\n[~treff7es] Can you try reproducing the issue on Hive directly and see what the behavior is? ","created":"2019-02-11T07:47:42.202+0000"},{"body":"On hive directly the query returns with the correct resultset regardless if _avro.schema.url_ is set for the partitions or not :\r\n{code:java}\r\nhive> select * from spark_test;\r\nOK\r\n6 fishfingers and custard Colin Baker 2019-02-05\r\n3 fishfingers and custard Jon Pertwee 2019-02-05\r\n4 fishfingers and custard Tom Baker 2019-02-05\r\n5 fishfingers and custard Peter Davison 2019-02-05\r\n11 fishfingers and custard Matt Smith 2019-02-05\r\n1 fishfingers and custard William Hartnell 2019-02-05\r\n7 fishfingers and custard Sylvester McCoy 2019-02-05\r\n8 fishfingers and custard Paul McGann 2019-02-05\r\n2 fishfingers and custard Patrick Troughton 2019-02-05\r\n9 fishfingers and custard Christopher Eccleston 2019-02-05\r\n10 fishfingers and custard David Tennant 2019-02-05\r\n21 fishfinger Jim Baker 2019-02-06\r\n24 fishfinger Bean Pertwee 2019-02-06\r\nTime taken: 4.291 seconds, Fetched: 13 row(s){code}","created":"2019-02-11T13:04:36.070+0000"},{"body":"And one more thing.\r\n\r\nI also got this warning which I did not get if I run on the table where the partitions do not contain the avro.schema.url property.\r\n{code:java}\r\n19/02/11 14:39:40 WARN AvroDeserializer: Received different schemas. Have to re-encode: {\"type\":\"record\",\"name\":\"doctors\",\"namespace\":\"testing.hive.avro.serde\",\"fields\":[{\"name\":\"number\",\"type\":\"int\",\"doc\":\"Order of playing the role\"},{\"name\":\"extra_field\",\"type\":\"string\",\"default\":\"fishfingers and custard\",\"doc:\":\"an extra field not in the original file\"},{\"name\":\"first_name\",\"type\":\"string\",\"doc\":\"first name of actor playing role\"},{\"name\":\"last_name\",\"type\":\"string\",\"doc\":\"last name of actor playing role\"}]}\r\nSIZE{-3ac2eea4:168dcc8e145:-8000=org.apache.hadoop.hive.serde2.avro.AvroDeserializer$SchemaReEncoder@1429dfec} ID -3ac2eea4:168dcc8e145:-8000{code}","created":"2019-02-11T13:43:22.434+0000"},{"body":"Hi,\r\n\r\n \r\n\r\nWe encounter the same issue with Spark 2.2.0 when reading avro from a partition where avro.schema.url point to a new schema (forward compatibility) but stored avro files were in older schema.\r\n\r\nQuerying with Hive client is OK but with Spark Sql we get \r\n{code}\r\n19/06/20 10:43:47 WARN avro.AvroDeserializer: Received different schemas. Have to re-encode: {\"type\":\"record\",\"name\":\"KeyValuePair\",\"namespace\":\"org.apache.avro.mapreduce\",\"doc\":\"A key/value pair\",\"fields\":[{\"name\":\"key\",\"type\": .............\r\nSIZE{4207b7ac:16b740e5ed1:-8000=org.apache.hadoop.hive.serde2.avro.AvroDeserializer$SchemaReEncoder@17173314} ID 4207b7ac:16b740e5ed1:-8000\r\n19/06/20 10:43:47 ERROR executor.Executor: Exception in task 6.0 in stage 1.0 (TID 22)\r\njava.lang.RuntimeException: Hive internal error: conversion of double to structnot supported yet.\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters$StructConverter.(ObjectInspectorConverters.java:380)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters.getConverter(ObjectInspectorConverters.java:155)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters$StructConverter.(ObjectInspectorConverters.java:374)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters.getConverter(ObjectInspectorConverters.java:155)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters$ListConverter.convert(ObjectInspectorConverters.java:331)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters$StructConverter.convert(ObjectInspectorConverters.java:396)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters$StructConverter.convert(ObjectInspectorConverters.java:396)\r\n\tat org.apache.spark.sql.hive.HadoopTableReader$$anonfun$fillObject$2.apply(TableReader.scala:430)\r\n\tat org.apache.spark.sql.hive.HadoopTableReader$$anonfun$fillObject$2.apply(TableReader.scala:429)\r\n\tat scala.collection.Iterator$$anon$11.next(Iterator.scala:409)\r\n\tat scala.collection.Iterator$$anon$11.next(Iterator.scala:409)\r\n\tat scala.collection.Iterator$$anon$11.next(Iterator.scala:409)\r\n\tat scala.collection.Iterator$$anon$11.next(Iterator.scala:409)\r\n\tat org.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java:149)\r\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)\r\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)\r\n\tat org.apache.spark.scheduler.Task.run(Task.scala:108)\r\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:338)\r\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n\tat java.lang.Thread.run(Thread.java:745)\r\n{code}","created":"2019-06-20T10:03:02.403+0000"},{"body":"cc [~dbtsai]","created":"2019-06-20T19:58:05.560+0000"},{"body":"cc [~Gengliang.Wang]","created":"2020-02-20T11:01:18.697+0000"},{"body":"I am lowering the priority to Critical as it's at least not a regression and doesn't look blocking Spark 3.0; however, indeed we should treat correctness issues at least Critical+.","created":"2020-02-28T02:59:50.478+0000"},{"body":"I am working on this and a PR can be expected this weekend / next week","created":"2021-01-08T16:14:44.857+0000"},{"body":"User 'attilapiros' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/31133","created":"2021-01-11T14:35:55.644+0000"},{"body":"Issue resolved by pull request 31133\n[https://github.com/apache/spark/pull/31133]","created":"2021-02-05T18:56:58.693+0000"}],"conversations":[{"body":"I have a hive avro table where the avro schema is stored on s3 next to the avro files. \r\n\r\nIn the table definiton the avro.schema.url always points to the latest partition's _schema.avsc file which is always the lates schema. (Avro schemas are backward and forward compatible in a table)\r\n\r\nWhen new data comes in, I always add a new partition where the avro.schema.url properties also set to the _schema.avsc which was used when it was added and of course I always update the table avro.schema.url property to the latest one.\r\n\r\nQuerying this table works fine until the schema evolves in a way that a new optional property is added in the middle. \r\n\r\nWhen this happens then after the spark sql query the columns in the old partition gets mixed up and it shows the wrong data for the columns.\r\n\r\nIf I query the table with Hive then everything is perfectly fine and it gives me back the correct columns for the partitions which were created the old schema and for the new which was created the evolved schema.\r\n\r\n \r\n\r\nHere is how I could reproduce with the [doctors.avro|https://github.com/apache/spark/blob/master/sql/hive/src/test/resources/data/files/doctors.avro] example data in sql test suite.\r\n # I have created two partition folder:\r\n{code:java}\r\n[hadoop@ip-192-168-10-158 hadoop]$ hdfs dfs -ls s3://somelocation/doctors/*/\r\nFound 2 items\r\n-rw-rw-rw- 1 hadoop hadoop 418 2019-02-06 12:48 s3://somelocation/doctors\r\n/dt=2019-02-05/_schema.avsc\r\n-rw-rw-rw- 1 hadoop hadoop 521 2019-02-06 12:13 s3://somelocation/doctors\r\n/dt=2019-02-05/doctors.avro\r\nFound 2 items\r\n-rw-rw-rw- 1 hadoop hadoop 580 2019-02-06 12:49 s3://somelocation/doctors\r\n/dt=2019-02-06/_schema.avsc\r\n-rw-rw-rw- 1 hadoop hadoop 577 2019-02-06 12:13 s3://somelocation/doctors\r\n/dt=2019-02-06/doctors_evolved.avro{code}\r\nHere the first partition had data which was created with the schema before evolving and the second one had the evolved one. (the evolved schema is the same as in your testcase except I moved the extra_field column to the last from the second and I generated two lines of avro data with the evolved schema.\r\n # I have created a hive table with the following command:\r\n\r\n \r\n{code:java}\r\nCREATE EXTERNAL TABLE `default.doctors`\r\n PARTITIONED BY (\r\n `dt` string\r\n )\r\n ROW FORMAT SERDE\r\n 'org.apache.hadoop.hive.serde2.avro.AvroSerDe'\r\n WITH SERDEPROPERTIES (\r\n 'avro.schema.url'='s3://somelocation/doctors/\r\n/dt=2019-02-06/_schema.avsc')\r\n STORED AS INPUTFORMAT\r\n 'org.apache.hadoop.hive.ql.io.avro.AvroContainerInputFormat'\r\n OUTPUTFORMAT\r\n 'org.apache.hadoop.hive.ql.io.avro.AvroContainerOutputFormat'\r\n LOCATION\r\n 's3://somelocation/doctors/'\r\n TBLPROPERTIES (\r\n 'transient_lastDdlTime'='1538130975'){code}\r\n \r\n\r\nHere as you can see the table schema url points to the latest schema\r\n\r\n3. I ran an msck _repair table_ to pick up all the partitions.\r\n\r\nFyi: If I run my select * query from here then everything is fine and no columns switch happening.\r\n\r\n4. Then I changed the first partition's avro.schema.url url to points to the schema which is under the partition folder (non-evolved one -> s3://somelocation/doctors/\r\n/dt=2019-02-05/_schema.avsc)\r\n\r\nThen if you ran a _select * from default.spark_test_ then the columns will be mixed up (on the data below the first name column becomes the extra_field column. I guess because in the latest schema it is the second column):\r\n\r\n \r\n{code:java}\r\nnumber,extra_field,first_name,last_name,dt \r\n6,Colin,Baker,null,2019-02-05 \r\n3,Jon,Pertwee,null,2019-02-05 \r\n4,Tom,Baker,null,2019-02-05 \r\n5,Peter,Davison,null,2019-02-05 \r\n11,Matt,Smith,null,2019-02-05 \r\n1,William,Hartnell,null,2019-02-05 \r\n7,Sylvester,McCoy,null,2019-02-05 \r\n8,Paul,McGann,null,2019-02-05 \r\n2,Patrick,Troughton,null,2019-02-05 \r\n9,Christopher,Eccleston,null,2019-02-05 \r\n10,David,Tennant,null,2019-02-05 \r\n21,fishfinger,Jim,Baker,2019-02-06 \r\n24,fishfinger,Bean,Pertwee,2019-02-06\r\n\r\n{code}\r\nIf I try the same query from Hive and not from spark sql then everything is fine and it never switches the columns.\r\n\r\n ","from":"reporter","subject":"Columns get switched in Spark SQL using Avro backed Hive table if schema evolves"},{"body":"Thank you for reporting, [~treff7es]. \r\nDid you see this issue before at Spark 2.3.x with Databricks Avro libraries?\r\n\r\ncc [~Gengliang.Wang]","from":"developer"},{"body":"I just tried with Spark 2.3.1 and the same issue. I did not specify any Avro library because it is a Hive Avro table originally. I guess it is using Hive's Avro serde for serialization/deserialization (or maybe I'm not right :)).","from":"developer"},{"body":"In the meantime I checked if using the same hive version in Spark would solve this issue by setting:\r\n\r\n \r\n{code:java}\r\nspark.sql.hive.metastore.version 2.3.3\r\nspark.sql.hive.metastore.jars /etc/hadoop/conf:/usr/lib/hadoop/lib/*:/usr/lib/hadoop/.//*:/usr/lib/hadoop-hdfs/./:/usr/lib/hadoop-hdfs/lib/*:/usr/lib/hadoop-hdfs/.//*:/usr/lib/hadoop-yarn/lib/*:/usr/lib/hadoop-yarn/.//*:/usr/lib/hadoop-mapreduce/lib/*:/usr/lib/hadoop-mapreduce/.//*::/etc/tez/conf:/usr/lib/tez/*:/usr/lib/tez/lib/*:/usr/lib/hadoop-lzo/lib/*:/usr/share/aws/aws-java-sdk/*:/usr/share/aws/emr/emrfs/conf:/usr/share/aws/emr/emrfs/lib/*:/usr/share/aws/emr/emrfs/auxlib/*:/usr/share/aws/emr/ddb/lib/emr-ddb-hadoop.jar:/usr/share/aws/emr/goodies/lib/emr-hadoop-goodies.jar:/usr/share/aws/emr/kinesis/lib/emr-kinesis-hadoop.jar:/usr/share/aws/emr/cloudwatch-sink/lib/*:/usr/share/aws/emr/security/conf:/usr/share/aws/emr/security/lib/*:/usr/lib/hive/lib/* \r\n{code}\r\nThe issue still exists regardless the hive metastore version spark is using.","from":"developer"},{"body":"[~dongjoon]I don't think the issue is related to spark-avro lib.\r\n[~treff7es] Can you try reproducing the issue on Hive directly and see what the behavior is? ","from":"developer"},{"body":"On hive directly the query returns with the correct resultset regardless if _avro.schema.url_ is set for the partitions or not :\r\n{code:java}\r\nhive> select * from spark_test;\r\nOK\r\n6 fishfingers and custard Colin Baker 2019-02-05\r\n3 fishfingers and custard Jon Pertwee 2019-02-05\r\n4 fishfingers and custard Tom Baker 2019-02-05\r\n5 fishfingers and custard Peter Davison 2019-02-05\r\n11 fishfingers and custard Matt Smith 2019-02-05\r\n1 fishfingers and custard William Hartnell 2019-02-05\r\n7 fishfingers and custard Sylvester McCoy 2019-02-05\r\n8 fishfingers and custard Paul McGann 2019-02-05\r\n2 fishfingers and custard Patrick Troughton 2019-02-05\r\n9 fishfingers and custard Christopher Eccleston 2019-02-05\r\n10 fishfingers and custard David Tennant 2019-02-05\r\n21 fishfinger Jim Baker 2019-02-06\r\n24 fishfinger Bean Pertwee 2019-02-06\r\nTime taken: 4.291 seconds, Fetched: 13 row(s){code}","from":"developer"},{"body":"And one more thing.\r\n\r\nI also got this warning which I did not get if I run on the table where the partitions do not contain the avro.schema.url property.\r\n{code:java}\r\n19/02/11 14:39:40 WARN AvroDeserializer: Received different schemas. Have to re-encode: {\"type\":\"record\",\"name\":\"doctors\",\"namespace\":\"testing.hive.avro.serde\",\"fields\":[{\"name\":\"number\",\"type\":\"int\",\"doc\":\"Order of playing the role\"},{\"name\":\"extra_field\",\"type\":\"string\",\"default\":\"fishfingers and custard\",\"doc:\":\"an extra field not in the original file\"},{\"name\":\"first_name\",\"type\":\"string\",\"doc\":\"first name of actor playing role\"},{\"name\":\"last_name\",\"type\":\"string\",\"doc\":\"last name of actor playing role\"}]}\r\nSIZE{-3ac2eea4:168dcc8e145:-8000=org.apache.hadoop.hive.serde2.avro.AvroDeserializer$SchemaReEncoder@1429dfec} ID -3ac2eea4:168dcc8e145:-8000{code}","from":"developer"},{"body":"Hi,\r\n\r\n \r\n\r\nWe encounter the same issue with Spark 2.2.0 when reading avro from a partition where avro.schema.url point to a new schema (forward compatibility) but stored avro files were in older schema.\r\n\r\nQuerying with Hive client is OK but with Spark Sql we get \r\n{code}\r\n19/06/20 10:43:47 WARN avro.AvroDeserializer: Received different schemas. Have to re-encode: {\"type\":\"record\",\"name\":\"KeyValuePair\",\"namespace\":\"org.apache.avro.mapreduce\",\"doc\":\"A key/value pair\",\"fields\":[{\"name\":\"key\",\"type\": .............\r\nSIZE{4207b7ac:16b740e5ed1:-8000=org.apache.hadoop.hive.serde2.avro.AvroDeserializer$SchemaReEncoder@17173314} ID 4207b7ac:16b740e5ed1:-8000\r\n19/06/20 10:43:47 ERROR executor.Executor: Exception in task 6.0 in stage 1.0 (TID 22)\r\njava.lang.RuntimeException: Hive internal error: conversion of double to structnot supported yet.\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters$StructConverter.(ObjectInspectorConverters.java:380)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters.getConverter(ObjectInspectorConverters.java:155)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters$StructConverter.(ObjectInspectorConverters.java:374)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters.getConverter(ObjectInspectorConverters.java:155)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters$ListConverter.convert(ObjectInspectorConverters.java:331)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters$StructConverter.convert(ObjectInspectorConverters.java:396)\r\n\tat org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorConverters$StructConverter.convert(ObjectInspectorConverters.java:396)\r\n\tat org.apache.spark.sql.hive.HadoopTableReader$$anonfun$fillObject$2.apply(TableReader.scala:430)\r\n\tat org.apache.spark.sql.hive.HadoopTableReader$$anonfun$fillObject$2.apply(TableReader.scala:429)\r\n\tat scala.collection.Iterator$$anon$11.next(Iterator.scala:409)\r\n\tat scala.collection.Iterator$$anon$11.next(Iterator.scala:409)\r\n\tat scala.collection.Iterator$$anon$11.next(Iterator.scala:409)\r\n\tat scala.collection.Iterator$$anon$11.next(Iterator.scala:409)\r\n\tat org.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java:149)\r\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)\r\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)\r\n\tat org.apache.spark.scheduler.Task.run(Task.scala:108)\r\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:338)\r\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n\tat java.lang.Thread.run(Thread.java:745)\r\n{code}","from":"developer"},{"body":"cc [~dbtsai]","from":"developer"},{"body":"cc [~Gengliang.Wang]","from":"developer"},{"body":"I am lowering the priority to Critical as it's at least not a regression and doesn't look blocking Spark 3.0; however, indeed we should treat correctness issues at least Critical+.","from":"developer"},{"body":"I am working on this and a PR can be expected this weekend / next week","from":"developer"},{"body":"User 'attilapiros' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/31133","from":"developer"},{"body":"Issue resolved by pull request 31133\n[https://github.com/apache/spark/pull/31133]","from":"developer"}],"created":"2019-02-06T14:54:48.000+0000","description":"I have a hive avro table where the avro schema is stored on s3 next to the avro files. \r\n\r\nIn the table definiton the avro.schema.url always points to the latest partition's _schema.avsc file which is always the lates schema. (Avro schemas are backward and forward compatible in a table)\r\n\r\nWhen new data comes in, I always add a new partition where the avro.schema.url properties also set to the _schema.avsc which was used when it was added and of course I always update the table avro.schema.url property to the latest one.\r\n\r\nQuerying this table works fine until the schema evolves in a way that a new optional property is added in the middle. \r\n\r\nWhen this happens then after the spark sql query the columns in the old partition gets mixed up and it shows the wrong data for the columns.\r\n\r\nIf I query the table with Hive then everything is perfectly fine and it gives me back the correct columns for the partitions which were created the old schema and for the new which was created the evolved schema.\r\n\r\n \r\n\r\nHere is how I could reproduce with the [doctors.avro|https://github.com/apache/spark/blob/master/sql/hive/src/test/resources/data/files/doctors.avro] example data in sql test suite.\r\n # I have created two partition folder:\r\n{code:java}\r\n[hadoop@ip-192-168-10-158 hadoop]$ hdfs dfs -ls s3://somelocation/doctors/*/\r\nFound 2 items\r\n-rw-rw-rw- 1 hadoop hadoop 418 2019-02-06 12:48 s3://somelocation/doctors\r\n/dt=2019-02-05/_schema.avsc\r\n-rw-rw-rw- 1 hadoop hadoop 521 2019-02-06 12:13 s3://somelocation/doctors\r\n/dt=2019-02-05/doctors.avro\r\nFound 2 items\r\n-rw-rw-rw- 1 hadoop hadoop 580 2019-02-06 12:49 s3://somelocation/doctors\r\n/dt=2019-02-06/_schema.avsc\r\n-rw-rw-rw- 1 hadoop hadoop 577 2019-02-06 12:13 s3://somelocation/doctors\r\n/dt=2019-02-06/doctors_evolved.avro{code}\r\nHere the first partition had data which was created with the schema before evolving and the second one had the evolved one. (the evolved schema is the same as in your testcase except I moved the extra_field column to the last from the second and I generated two lines of avro data with the evolved schema.\r\n # I have created a hive table with the following command:\r\n\r\n \r\n{code:java}\r\nCREATE EXTERNAL TABLE `default.doctors`\r\n PARTITIONED BY (\r\n `dt` string\r\n )\r\n ROW FORMAT SERDE\r\n 'org.apache.hadoop.hive.serde2.avro.AvroSerDe'\r\n WITH SERDEPROPERTIES (\r\n 'avro.schema.url'='s3://somelocation/doctors/\r\n/dt=2019-02-06/_schema.avsc')\r\n STORED AS INPUTFORMAT\r\n 'org.apache.hadoop.hive.ql.io.avro.AvroContainerInputFormat'\r\n OUTPUTFORMAT\r\n 'org.apache.hadoop.hive.ql.io.avro.AvroContainerOutputFormat'\r\n LOCATION\r\n 's3://somelocation/doctors/'\r\n TBLPROPERTIES (\r\n 'transient_lastDdlTime'='1538130975'){code}\r\n \r\n\r\nHere as you can see the table schema url points to the latest schema\r\n\r\n3. I ran an msck _repair table_ to pick up all the partitions.\r\n\r\nFyi: If I run my select * query from here then everything is fine and no columns switch happening.\r\n\r\n4. Then I changed the first partition's avro.schema.url url to points to the schema which is under the partition folder (non-evolved one -> s3://somelocation/doctors/\r\n/dt=2019-02-05/_schema.avsc)\r\n\r\nThen if you ran a _select * from default.spark_test_ then the columns will be mixed up (on the data below the first name column becomes the extra_field column. I guess because in the latest schema it is the second column):\r\n\r\n \r\n{code:java}\r\nnumber,extra_field,first_name,last_name,dt \r\n6,Colin,Baker,null,2019-02-05 \r\n3,Jon,Pertwee,null,2019-02-05 \r\n4,Tom,Baker,null,2019-02-05 \r\n5,Peter,Davison,null,2019-02-05 \r\n11,Matt,Smith,null,2019-02-05 \r\n1,William,Hartnell,null,2019-02-05 \r\n7,Sylvester,McCoy,null,2019-02-05 \r\n8,Paul,McGann,null,2019-02-05 \r\n2,Patrick,Troughton,null,2019-02-05 \r\n9,Christopher,Eccleston,null,2019-02-05 \r\n10,David,Tennant,null,2019-02-05 \r\n21,fishfinger,Jim,Baker,2019-02-06 \r\n24,fishfinger,Bean,Pertwee,2019-02-06\r\n\r\n{code}\r\nIf I try the same query from Hive and not from spark sql then everything is fine and it never switches the columns.\r\n\r\n ","issue_id":"13214170","key":"SPARK-26836","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2021-02-05T18:56:58.000+0000","role":"fixed_distractor","summary":"Columns get switched in Spark SQL using Avro backed Hive table if schema evolves"} {"case_id":"13222381","cluster":"DISTRACTOR-SPARK-27194","comments":[{"body":"Currently looks like from logs the file name for task 200.0 and 200.1(reattempt) expected file name to be, part-00200-blah-blah.c000.snappy.parquet. (refer org.apache.spark.internal.io.HadoopMapReduceCommitProtocol#getFilename)\r\n\r\nMay be we should have taskId_attemptId in the part file name so that rerun tasks do not conflict with older failed tasks.\r\n\r\ncc [~srowen] [~cloud_fan] [~dongjoon] any thoughts.?","created":"2019-03-19T02:37:51.419+0000"},{"body":"Hi, [~ajithshetty]. Just out of curiosity, according to the `Affected Versions`, do you observe this in 2.3.1 and 2.3.2? Is it possible to use 2.3.3 in your cluster since that's already released?","created":"2019-03-19T19:06:12.708+0000"},{"body":"[~dongjoon] Yes i tried with spark 2.3.3, and the issue persist. Here is the operation i performed\r\n{code:java}\r\nspark.sql.sources.partitionOverwriteMode=DYNAMIC{code}\r\n{code:java}\r\ncreate table t1 (i int, part1 int, part2 int) using parquet partitioned by (part1, part2)\r\ninsert into t1 partition(part1=1, part2=1) select 1\r\ninsert overwrite table t1 partition(part1=1, part2=1) select 2\r\ninsert overwrite table t1 partition(part1=2, part2) select 2, 2 // here the exec is killed and task respawns{code}\r\nand here is the full stacktrace as per 2.3.3\r\n{code:java}\r\n2019-03-20 19:58:06 WARN TaskSetManager:66 - Lost task 0.1 in stage 2.0 (TID 3, QWERTY, executor 2): org.apache.spark.SparkException: Task failed while writing rows.\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:285)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:197)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:196)\r\nat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\r\nat org.apache.spark.scheduler.Task.run(Task.scala:109)\r\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)\r\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\nat java.lang.Thread.run(Thread.java:745)\r\nCaused by: org.apache.hadoop.fs.FileAlreadyExistsException: /user/hive/warehouse/t2/.spark-staging-1f1efbfd-7e20-4e0f-a49c-a7fa3eae4cb1/part1=2/part2=2/part-00000-1f1efbfd-7e20-4e0f-a49c-a7fa3eae4cb1.c000.snappy.parquet for client 127.0.0.1 already exists\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFileInternal(FSNamesystem.java:2578)\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFileInt(FSNamesystem.java:2465)\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFile(FSNamesystem.java:2349)\r\nat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.create(NameNodeRpcServer.java:624)\r\nat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.create(ClientNamenodeProtocolServerSideTranslatorPB.java:398)\r\nat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\r\nat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\r\nat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\r\nat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2217)\r\nat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2213)\r\nat java.security.AccessController.doPrivileged(Native Method)\r\nat javax.security.auth.Subject.doAs(Subject.java:422)\r\nat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1746)\r\nat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2213)\r\n\r\nat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\r\nat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\r\nat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\r\nat java.lang.reflect.Constructor.newInstance(Constructor.java:423)\r\nat org.apache.hadoop.ipc.RemoteException.instantiateException(RemoteException.java:106)\r\nat org.apache.hadoop.ipc.RemoteException.unwrapRemoteException(RemoteException.java:73)\r\nat org.apache.hadoop.hdfs.DFSOutputStream.newStreamForCreate(DFSOutputStream.java:1653)\r\nat org.apache.hadoop.hdfs.DFSClient.create(DFSClient.java:1689)\r\nat org.apache.hadoop.hdfs.DFSClient.create(DFSClient.java:1624)\r\nat org.apache.hadoop.hdfs.DistributedFileSystem$7.doCall(DistributedFileSystem.java:448)\r\nat org.apache.hadoop.hdfs.DistributedFileSystem$7.doCall(DistributedFileSystem.java:444)\r\nat org.apache.hadoop.fs.FileSystemLinkResolver.resolve(FileSystemLinkResolver.java:81)\r\nat org.apache.hadoop.hdfs.DistributedFileSystem.create(DistributedFileSystem.java:459)\r\nat org.apache.hadoop.hdfs.DistributedFileSystem.create(DistributedFileSystem.java:387)\r\nat org.apache.hadoop.fs.FileSystem.create(FileSystem.java:911)\r\nat org.apache.hadoop.fs.FileSystem.create(FileSystem.java:892)\r\nat org.apache.parquet.hadoop.ParquetFileWriter.(ParquetFileWriter.java:236)\r\nat org.apache.parquet.hadoop.ParquetOutputFormat.getRecordWriter(ParquetOutputFormat.java:342)\r\nat org.apache.parquet.hadoop.ParquetOutputFormat.getRecordWriter(ParquetOutputFormat.java:302)\r\nat org.apache.spark.sql.execution.datasources.parquet.ParquetOutputWriter.(ParquetOutputWriter.scala:37)\r\nat org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anon$1.newInstance(ParquetFileFormat.scala:151)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$DynamicPartitionWriteTask.org$apache$spark$sql$execution$datasources$FileFormatWriter$DynamicPartitionWriteTask$$newOutputWriter(FileFormatWriter.scala:511)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$DynamicPartitionWriteTask$$anonfun$execute$5.apply(FileFormatWriter.scala:546)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$DynamicPartitionWriteTask$$anonfun$execute$5.apply(FileFormatWriter.scala:527)\r\nat scala.collection.Iterator$class.foreach(Iterator.scala:893)\r\nat scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$DynamicPartitionWriteTask.execute(FileFormatWriter.scala:527)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:269)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:267)\r\nat org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1415)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:272)\r\n... 8 more\r\nCaused by: org.apache.hadoop.ipc.RemoteException(org.apache.hadoop.fs.FileAlreadyExistsException): /user/hive/warehouse/t2/.spark-staging-1f1efbfd-7e20-4e0f-a49c-a7fa3eae4cb1/part1=2/part2=2/part-00000-1f1efbfd-7e20-4e0f-a49c-a7fa3eae4cb1.c000.snappy.parquet for client 127.0.0.1 already exists\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFileInternal(FSNamesystem.java:2578)\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFileInt(FSNamesystem.java:2465)\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFile(FSNamesystem.java:2349)\r\nat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.create(NameNodeRpcServer.java:624)\r\nat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.create(ClientNamenodeProtocolServerSideTranslatorPB.java:398)\r\nat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\r\nat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\r\nat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\r\nat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2217)\r\nat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2213)\r\nat java.security.AccessController.doPrivileged(Native Method)\r\nat javax.security.auth.Subject.doAs(Subject.java:422)\r\nat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1746)\r\nat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2213)\r\n\r\nat org.apache.hadoop.ipc.Client.call(Client.java:1475)\r\nat org.apache.hadoop.ipc.Client.call(Client.java:1412)\r\nat org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:229)\r\nat com.sun.proxy.$Proxy15.create(Unknown Source)\r\nat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolTranslatorPB.create(ClientNamenodeProtocolTranslatorPB.java:296)\r\nat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\nat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\nat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\nat java.lang.reflect.Method.invoke(Method.java:498)\r\nat org.apache.hadoop.io.retry.RetryInvocationHandler.invokeMethod(RetryInvocationHandler.java:191)\r\nat org.apache.hadoop.io.retry.RetryInvocationHandler.invoke(RetryInvocationHandler.java:102)\r\nat com.sun.proxy.$Proxy16.create(Unknown Source)\r\nat org.apache.hadoop.hdfs.DFSOutputStream.newStreamForCreate(DFSOutputStream.java:1648)\r\n... 32 more\r\n{code}\r\n \r\n\r\n \r\n\r\n ","created":"2019-03-20T14:33:18.157+0000"},{"body":"Hi [~dongjoon] , i have some analysis here : [https://github.com/apache/spark/pull/24142#issuecomment-474866759] Please let me know your views","created":"2019-03-20T14:54:48.607+0000"},{"body":"Thank you for checking that, [~ajithshetty].","created":"2019-03-20T22:55:15.627+0000"},{"body":"[~ajithshetty] is this being worked upon","created":"2019-11-18T15:42:45.046+0000"},{"body":"User 'turboFei' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28989","created":"2020-07-03T03:06:18.090+0000"},{"body":"User 'turboFei' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28989","created":"2020-07-03T03:07:48.596+0000"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29000","created":"2020-07-05T20:02:05.828+0000"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29260","created":"2020-07-27T11:43:25.207+0000"},{"body":"Issue resolved by pull request 29000\n[https://github.com/apache/spark/pull/29000]","created":"2020-11-25T12:50:56.724+0000"}],"conversations":[{"body":"When a container fails for some reason (for example when killed by yarn for exceeding memory limits), the subsequent task attempts for the tasks that were running on that container all fail with a FileAlreadyExistsException. The original task attempt does not seem to successfully call abortTask (or at least its \"best effort\" delete is unsuccessful) and clean up the parquet file it was writing to, so when later task attempts try to write to the same spark-staging directory using the same file name, the job fails.\r\n\r\nHere is what transpires in the logs:\r\n\r\nThe container where task 200.0 is running is killed and the task is lost:\r\n{code}\r\n19/02/20 09:33:25 ERROR cluster.YarnClusterScheduler: Lost executor y on t.y.z.com: Container killed by YARN for exceeding memory limits. 8.1 GB of 8 GB physical memory used. Consider boosting spark.yarn.executor.memoryOverhead.\r\n 19/02/20 09:33:25 WARN scheduler.TaskSetManager: Lost task 200.0 in stage 0.0 (TID xxx, t.y.z.com, executor 93): ExecutorLostFailure (executor 93 exited caused by one of the running tasks) Reason: Container killed by YARN for exceeding memory limits. 8.1 GB of 8 GB physical memory used. Consider boosting spark.yarn.executor.memoryOverhead.\r\n{code}\r\n\r\nThe task is re-attempted on a different executor and fails because the part-00200-blah-blah.c000.snappy.parquet file from the first task attempt already exists:\r\n{code}\r\n19/02/20 09:35:01 WARN scheduler.TaskSetManager: Lost task 200.1 in stage 0.0 (TID 594, tn.y.z.com, executor 70): org.apache.spark.SparkException: Task failed while writing rows.\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:285)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:197)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:196)\r\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\r\n at org.apache.spark.scheduler.Task.run(Task.scala:109)\r\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n at java.lang.Thread.run(Thread.java:745)\r\n Caused by: org.apache.hadoop.fs.FileAlreadyExistsException: /user/hive/warehouse/tmp_supply_feb1/.spark-staging-blah-blah-blah/dt=2019-02-17/part-00200-blah-blah.c000.snappy.parquet for client a.b.c.d already exists\r\n{code}\r\n\r\nThe job fails when the the configured task attempts (spark.task.maxFailures) have failed with the same error:\r\n{code}\r\norg.apache.spark.SparkException: Job aborted due to stage failure: Task 200 in stage 0.0 failed 20 times, most recent failure: Lost task 284.19 in stage 0.0 (TID yyy, tm.y.z.com, executor 16): org.apache.spark.SparkException: Task failed while writing rows.\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:285)\r\n ...\r\n Caused by: org.apache.hadoop.fs.FileAlreadyExistsException: /user/hive/warehouse/tmp_supply_feb1/.spark-staging-blah-blah-blah/dt=2019-02-17/part-00200-blah-blah.c000.snappy.parquet for client i.p.a.d already exists\r\n{code}\r\n\r\nSPARK-26682 wasn't the root cause here, since there wasn't any stage reattempt.\r\n\r\nThis issue seems to happen when spark.sql.sources.partitionOverwriteMode=dynamic. \r\n\r\n ","from":"reporter","subject":"Job failures when task attempts do not clean up spark-staging parquet files"},{"body":"Currently looks like from logs the file name for task 200.0 and 200.1(reattempt) expected file name to be, part-00200-blah-blah.c000.snappy.parquet. (refer org.apache.spark.internal.io.HadoopMapReduceCommitProtocol#getFilename)\r\n\r\nMay be we should have taskId_attemptId in the part file name so that rerun tasks do not conflict with older failed tasks.\r\n\r\ncc [~srowen] [~cloud_fan] [~dongjoon] any thoughts.?","from":"developer"},{"body":"Hi, [~ajithshetty]. Just out of curiosity, according to the `Affected Versions`, do you observe this in 2.3.1 and 2.3.2? Is it possible to use 2.3.3 in your cluster since that's already released?","from":"developer"},{"body":"[~dongjoon] Yes i tried with spark 2.3.3, and the issue persist. Here is the operation i performed\r\n{code:java}\r\nspark.sql.sources.partitionOverwriteMode=DYNAMIC{code}\r\n{code:java}\r\ncreate table t1 (i int, part1 int, part2 int) using parquet partitioned by (part1, part2)\r\ninsert into t1 partition(part1=1, part2=1) select 1\r\ninsert overwrite table t1 partition(part1=1, part2=1) select 2\r\ninsert overwrite table t1 partition(part1=2, part2) select 2, 2 // here the exec is killed and task respawns{code}\r\nand here is the full stacktrace as per 2.3.3\r\n{code:java}\r\n2019-03-20 19:58:06 WARN TaskSetManager:66 - Lost task 0.1 in stage 2.0 (TID 3, QWERTY, executor 2): org.apache.spark.SparkException: Task failed while writing rows.\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:285)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:197)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:196)\r\nat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\r\nat org.apache.spark.scheduler.Task.run(Task.scala:109)\r\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)\r\nat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\nat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\nat java.lang.Thread.run(Thread.java:745)\r\nCaused by: org.apache.hadoop.fs.FileAlreadyExistsException: /user/hive/warehouse/t2/.spark-staging-1f1efbfd-7e20-4e0f-a49c-a7fa3eae4cb1/part1=2/part2=2/part-00000-1f1efbfd-7e20-4e0f-a49c-a7fa3eae4cb1.c000.snappy.parquet for client 127.0.0.1 already exists\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFileInternal(FSNamesystem.java:2578)\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFileInt(FSNamesystem.java:2465)\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFile(FSNamesystem.java:2349)\r\nat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.create(NameNodeRpcServer.java:624)\r\nat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.create(ClientNamenodeProtocolServerSideTranslatorPB.java:398)\r\nat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\r\nat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\r\nat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\r\nat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2217)\r\nat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2213)\r\nat java.security.AccessController.doPrivileged(Native Method)\r\nat javax.security.auth.Subject.doAs(Subject.java:422)\r\nat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1746)\r\nat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2213)\r\n\r\nat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\r\nat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\r\nat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\r\nat java.lang.reflect.Constructor.newInstance(Constructor.java:423)\r\nat org.apache.hadoop.ipc.RemoteException.instantiateException(RemoteException.java:106)\r\nat org.apache.hadoop.ipc.RemoteException.unwrapRemoteException(RemoteException.java:73)\r\nat org.apache.hadoop.hdfs.DFSOutputStream.newStreamForCreate(DFSOutputStream.java:1653)\r\nat org.apache.hadoop.hdfs.DFSClient.create(DFSClient.java:1689)\r\nat org.apache.hadoop.hdfs.DFSClient.create(DFSClient.java:1624)\r\nat org.apache.hadoop.hdfs.DistributedFileSystem$7.doCall(DistributedFileSystem.java:448)\r\nat org.apache.hadoop.hdfs.DistributedFileSystem$7.doCall(DistributedFileSystem.java:444)\r\nat org.apache.hadoop.fs.FileSystemLinkResolver.resolve(FileSystemLinkResolver.java:81)\r\nat org.apache.hadoop.hdfs.DistributedFileSystem.create(DistributedFileSystem.java:459)\r\nat org.apache.hadoop.hdfs.DistributedFileSystem.create(DistributedFileSystem.java:387)\r\nat org.apache.hadoop.fs.FileSystem.create(FileSystem.java:911)\r\nat org.apache.hadoop.fs.FileSystem.create(FileSystem.java:892)\r\nat org.apache.parquet.hadoop.ParquetFileWriter.(ParquetFileWriter.java:236)\r\nat org.apache.parquet.hadoop.ParquetOutputFormat.getRecordWriter(ParquetOutputFormat.java:342)\r\nat org.apache.parquet.hadoop.ParquetOutputFormat.getRecordWriter(ParquetOutputFormat.java:302)\r\nat org.apache.spark.sql.execution.datasources.parquet.ParquetOutputWriter.(ParquetOutputWriter.scala:37)\r\nat org.apache.spark.sql.execution.datasources.parquet.ParquetFileFormat$$anon$1.newInstance(ParquetFileFormat.scala:151)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$DynamicPartitionWriteTask.org$apache$spark$sql$execution$datasources$FileFormatWriter$DynamicPartitionWriteTask$$newOutputWriter(FileFormatWriter.scala:511)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$DynamicPartitionWriteTask$$anonfun$execute$5.apply(FileFormatWriter.scala:546)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$DynamicPartitionWriteTask$$anonfun$execute$5.apply(FileFormatWriter.scala:527)\r\nat scala.collection.Iterator$class.foreach(Iterator.scala:893)\r\nat scala.collection.AbstractIterator.foreach(Iterator.scala:1336)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$DynamicPartitionWriteTask.execute(FileFormatWriter.scala:527)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:269)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:267)\r\nat org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1415)\r\nat org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:272)\r\n... 8 more\r\nCaused by: org.apache.hadoop.ipc.RemoteException(org.apache.hadoop.fs.FileAlreadyExistsException): /user/hive/warehouse/t2/.spark-staging-1f1efbfd-7e20-4e0f-a49c-a7fa3eae4cb1/part1=2/part2=2/part-00000-1f1efbfd-7e20-4e0f-a49c-a7fa3eae4cb1.c000.snappy.parquet for client 127.0.0.1 already exists\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFileInternal(FSNamesystem.java:2578)\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFileInt(FSNamesystem.java:2465)\r\nat org.apache.hadoop.hdfs.server.namenode.FSNamesystem.startFile(FSNamesystem.java:2349)\r\nat org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.create(NameNodeRpcServer.java:624)\r\nat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.create(ClientNamenodeProtocolServerSideTranslatorPB.java:398)\r\nat org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\r\nat org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\r\nat org.apache.hadoop.ipc.RPC$Server.call(RPC.java:982)\r\nat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2217)\r\nat org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2213)\r\nat java.security.AccessController.doPrivileged(Native Method)\r\nat javax.security.auth.Subject.doAs(Subject.java:422)\r\nat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1746)\r\nat org.apache.hadoop.ipc.Server$Handler.run(Server.java:2213)\r\n\r\nat org.apache.hadoop.ipc.Client.call(Client.java:1475)\r\nat org.apache.hadoop.ipc.Client.call(Client.java:1412)\r\nat org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:229)\r\nat com.sun.proxy.$Proxy15.create(Unknown Source)\r\nat org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolTranslatorPB.create(ClientNamenodeProtocolTranslatorPB.java:296)\r\nat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\nat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\nat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\nat java.lang.reflect.Method.invoke(Method.java:498)\r\nat org.apache.hadoop.io.retry.RetryInvocationHandler.invokeMethod(RetryInvocationHandler.java:191)\r\nat org.apache.hadoop.io.retry.RetryInvocationHandler.invoke(RetryInvocationHandler.java:102)\r\nat com.sun.proxy.$Proxy16.create(Unknown Source)\r\nat org.apache.hadoop.hdfs.DFSOutputStream.newStreamForCreate(DFSOutputStream.java:1648)\r\n... 32 more\r\n{code}\r\n \r\n\r\n \r\n\r\n ","from":"developer"},{"body":"Hi [~dongjoon] , i have some analysis here : [https://github.com/apache/spark/pull/24142#issuecomment-474866759] Please let me know your views","from":"developer"},{"body":"Thank you for checking that, [~ajithshetty].","from":"developer"},{"body":"[~ajithshetty] is this being worked upon","from":"developer"},{"body":"User 'turboFei' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28989","from":"developer"},{"body":"User 'turboFei' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28989","from":"developer"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29000","from":"developer"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29260","from":"developer"},{"body":"Issue resolved by pull request 29000\n[https://github.com/apache/spark/pull/29000]","from":"developer"}],"created":"2019-03-18T17:43:38.000+0000","description":"When a container fails for some reason (for example when killed by yarn for exceeding memory limits), the subsequent task attempts for the tasks that were running on that container all fail with a FileAlreadyExistsException. The original task attempt does not seem to successfully call abortTask (or at least its \"best effort\" delete is unsuccessful) and clean up the parquet file it was writing to, so when later task attempts try to write to the same spark-staging directory using the same file name, the job fails.\r\n\r\nHere is what transpires in the logs:\r\n\r\nThe container where task 200.0 is running is killed and the task is lost:\r\n{code}\r\n19/02/20 09:33:25 ERROR cluster.YarnClusterScheduler: Lost executor y on t.y.z.com: Container killed by YARN for exceeding memory limits. 8.1 GB of 8 GB physical memory used. Consider boosting spark.yarn.executor.memoryOverhead.\r\n 19/02/20 09:33:25 WARN scheduler.TaskSetManager: Lost task 200.0 in stage 0.0 (TID xxx, t.y.z.com, executor 93): ExecutorLostFailure (executor 93 exited caused by one of the running tasks) Reason: Container killed by YARN for exceeding memory limits. 8.1 GB of 8 GB physical memory used. Consider boosting spark.yarn.executor.memoryOverhead.\r\n{code}\r\n\r\nThe task is re-attempted on a different executor and fails because the part-00200-blah-blah.c000.snappy.parquet file from the first task attempt already exists:\r\n{code}\r\n19/02/20 09:35:01 WARN scheduler.TaskSetManager: Lost task 200.1 in stage 0.0 (TID 594, tn.y.z.com, executor 70): org.apache.spark.SparkException: Task failed while writing rows.\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:285)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:197)\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:196)\r\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\r\n at org.apache.spark.scheduler.Task.run(Task.scala:109)\r\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\r\n at java.lang.Thread.run(Thread.java:745)\r\n Caused by: org.apache.hadoop.fs.FileAlreadyExistsException: /user/hive/warehouse/tmp_supply_feb1/.spark-staging-blah-blah-blah/dt=2019-02-17/part-00200-blah-blah.c000.snappy.parquet for client a.b.c.d already exists\r\n{code}\r\n\r\nThe job fails when the the configured task attempts (spark.task.maxFailures) have failed with the same error:\r\n{code}\r\norg.apache.spark.SparkException: Job aborted due to stage failure: Task 200 in stage 0.0 failed 20 times, most recent failure: Lost task 284.19 in stage 0.0 (TID yyy, tm.y.z.com, executor 16): org.apache.spark.SparkException: Task failed while writing rows.\r\n at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:285)\r\n ...\r\n Caused by: org.apache.hadoop.fs.FileAlreadyExistsException: /user/hive/warehouse/tmp_supply_feb1/.spark-staging-blah-blah-blah/dt=2019-02-17/part-00200-blah-blah.c000.snappy.parquet for client i.p.a.d already exists\r\n{code}\r\n\r\nSPARK-26682 wasn't the root cause here, since there wasn't any stage reattempt.\r\n\r\nThis issue seems to happen when spark.sql.sources.partitionOverwriteMode=dynamic. \r\n\r\n ","issue_id":"13222381","key":"SPARK-27194","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2020-11-25T12:50:56.000+0000","role":"fixed_distractor","summary":"Job failures when task attempts do not clean up spark-staging parquet files"} {"case_id":"13238597","cluster":"DISTRACTOR-SPARK-27991","comments":[{"body":"I've tried to come up with a standalone reproduction of this issue, but so far I've been unable to find one that triggers this error. I've tried creating jobs which run 10000+ mappers shuffling tiny blocks to a single reducer, resulting in thousands of requests in flight, but this has failed to trigger the error posted above.\r\n\r\nHowever, I _did_ manage to get a more complete backtrace from a different internal workload:\r\n{code:java}\r\nCaused by: io.netty.util.internal.OutOfDirectMemoryError: failed to allocate 16777216 byte(s) of direct memory (used: 7918845952, max: 7923040256)\r\n\tat io.netty.util.internal.PlatformDependent.incrementMemoryCounter(PlatformDependent.java:640)\r\n\tat io.netty.util.internal.PlatformDependent.allocateDirectNoCleaner(PlatformDependent.java:594)\r\n\tat io.netty.buffer.PoolArena$DirectArena.allocateDirect(PoolArena.java:764)\r\n\tat io.netty.buffer.PoolArena$DirectArena.newChunk(PoolArena.java:740)\r\n\tat io.netty.buffer.PoolArena.allocateNormal(PoolArena.java:244)\r\n\tat io.netty.buffer.PoolArena.allocate(PoolArena.java:226)\r\n\tat io.netty.buffer.PoolArena.allocate(PoolArena.java:146)\r\n\tat io.netty.buffer.PooledByteBufAllocator.newDirectBuffer(PooledByteBufAllocator.java:324)\r\n\tat io.netty.buffer.AbstractByteBufAllocator.directBuffer(AbstractByteBufAllocator.java:185)\r\n\tat io.netty.buffer.AbstractByteBufAllocator.directBuffer(AbstractByteBufAllocator.java:176)\r\n\tat io.netty.buffer.AbstractByteBufAllocator.ioBuffer(AbstractByteBufAllocator.java:137)\r\n\tat io.netty.channel.DefaultMaxMessagesRecvByteBufAllocator$MaxMessageHandle.allocate(DefaultMaxMessagesRecvByteBufAllocator.java:80)\r\n\tat io.netty.channel.nio.AbstractNioByteChannel$NioByteUnsafe.read(AbstractNioByteChannel.java:122)\r\n\tat io.netty.channel.nio.NioEventLoop.processSelectedKey(NioEventLoop.java:645)\r\n\tat io.netty.channel.nio.NioEventLoop.processSelectedKeysOptimized(NioEventLoop.java:580)\r\n\tat io.netty.channel.nio.NioEventLoop.processSelectedKeys(NioEventLoop.java:497)\r\n\tat io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:459)\r\n\tat io.netty.util.concurrent.SingleThreadEventExecutor$5.run(SingleThreadEventExecutor.java:858)\r\n\tat io.netty.util.concurrent.DefaultThreadFactory$DefaultRunnableDecorator.run(DefaultThreadFactory.java:138)\r\n\t... 1 more{code}\r\nSomething that jumps out to me is the {{DefaultMaxMessagesRecvByteBufAllocator}} (and the {{AdaptiveRecvByteBufAllocator}} in SPARK-24989): maybe there's something about these failing workloads which is leading to significant space wasting in receive buffers, causing tiny blocks to experience huge bloat in space requirements?","created":"2019-07-11T00:27:31.832+0000"},{"body":"User 'Ngone51' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/32287","created":"2021-04-22T05:50:26.576+0000"},{"body":"User 'Ngone51' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/32287","created":"2021-04-22T05:50:42.295+0000"},{"body":"Issue resolved by pull request 32287\n[https://github.com/apache/spark/pull/32287]","created":"2021-05-20T04:27:12.620+0000"}],"conversations":[{"body":"ShuffleBlockFetcherIterator has logic to limit the number of simultaneous block fetches. By default, this logic tries to keep the number of outstanding block fetches [beneath a data size limit|https://github.com/apache/spark/blob/v2.4.3/core/src/main/scala/org/apache/spark/storage/ShuffleBlockFetcherIterator.scala#L274] ({{maxBytesInFlight}}). However, this limiting does not take fixed overheads into account: even though a remote block might be, say, 4KB, there are certain fixed-size internal overheads due to Netty buffer sizes which may cause the actual space requirements to be larger.\r\n\r\nAs a result, if a map stage produces a huge number of extremely tiny blocks then we may see errors like\r\n{code:java}\r\norg.apache.spark.shuffle.FetchFailedException: failed to allocate 16777216 byte(s) of direct memory (used: 39325794304, max: 39325794304)\r\nat org.apache.spark.storage.ShuffleBlockFetcherIterator.throwFetchFailedException(ShuffleBlockFetcherIterator.scala:554)\r\nat org.apache.spark.storage.ShuffleBlockFetcherIterator.next(ShuffleBlockFetcherIterator.scala:485)\r\n[...]\r\nCaused by: io.netty.util.internal.OutOfDirectMemoryError: failed to allocate 16777216 byte(s) of direct memory (used: 39325794304, max: 39325794304)\r\nat io.netty.util.internal.PlatformDependent.incrementMemoryCounter(PlatformDependent.java:640)\r\nat io.netty.util.internal.PlatformDependent.allocateDirectNoCleaner(PlatformDependent.java:594)\r\nat io.netty.buffer.PoolArena$DirectArena.allocateDirect(PoolArena.java:764)\r\nat io.netty.buffer.PoolArena$DirectArena.newChunk(PoolArena.java:740)\r\nat io.netty.buffer.PoolArena.allocateNormal(PoolArena.java:244)\r\nat io.netty.buffer.PoolArena.allocate(PoolArena.java:226)\r\nat io.netty.buffer.PoolArena.allocate(PoolArena.java:146)\r\nat io.netty.buffer.PooledByteBufAllocator.newDirectBuffer(PooledByteBufAllocator.java:324)\r\n[...]{code}\r\nSPARK-24989 is another report of this problem (but with a different proposed fix).\r\n\r\nThis problem can currently be mitigated by setting {{spark.reducer.maxReqsInFlight}} to some some non-IntMax value (SPARK-6166), but this additional manual configuration step is cumbersome.\r\n\r\nInstead, I think that Spark should take these fixed overheads into account in the {{maxBytesInFlight}} calculation: instead of using blocks' actual sizes, use {{Math.min(blockSize, minimumNettyBufferSize)}}. There might be some tricky details involved to make this work on all configurations (e.g. to use a different minimum when direct buffers are disabled, etc.), but I think the core idea behind the fix is pretty simple.\r\n\r\nThis will improve Spark's stability and removes configuration / tuning burden from end users.","from":"reporter","subject":"ShuffleBlockFetcherIterator should take Netty constant-factor overheads into account when limiting number of simultaneous block fetches"},{"body":"I've tried to come up with a standalone reproduction of this issue, but so far I've been unable to find one that triggers this error. I've tried creating jobs which run 10000+ mappers shuffling tiny blocks to a single reducer, resulting in thousands of requests in flight, but this has failed to trigger the error posted above.\r\n\r\nHowever, I _did_ manage to get a more complete backtrace from a different internal workload:\r\n{code:java}\r\nCaused by: io.netty.util.internal.OutOfDirectMemoryError: failed to allocate 16777216 byte(s) of direct memory (used: 7918845952, max: 7923040256)\r\n\tat io.netty.util.internal.PlatformDependent.incrementMemoryCounter(PlatformDependent.java:640)\r\n\tat io.netty.util.internal.PlatformDependent.allocateDirectNoCleaner(PlatformDependent.java:594)\r\n\tat io.netty.buffer.PoolArena$DirectArena.allocateDirect(PoolArena.java:764)\r\n\tat io.netty.buffer.PoolArena$DirectArena.newChunk(PoolArena.java:740)\r\n\tat io.netty.buffer.PoolArena.allocateNormal(PoolArena.java:244)\r\n\tat io.netty.buffer.PoolArena.allocate(PoolArena.java:226)\r\n\tat io.netty.buffer.PoolArena.allocate(PoolArena.java:146)\r\n\tat io.netty.buffer.PooledByteBufAllocator.newDirectBuffer(PooledByteBufAllocator.java:324)\r\n\tat io.netty.buffer.AbstractByteBufAllocator.directBuffer(AbstractByteBufAllocator.java:185)\r\n\tat io.netty.buffer.AbstractByteBufAllocator.directBuffer(AbstractByteBufAllocator.java:176)\r\n\tat io.netty.buffer.AbstractByteBufAllocator.ioBuffer(AbstractByteBufAllocator.java:137)\r\n\tat io.netty.channel.DefaultMaxMessagesRecvByteBufAllocator$MaxMessageHandle.allocate(DefaultMaxMessagesRecvByteBufAllocator.java:80)\r\n\tat io.netty.channel.nio.AbstractNioByteChannel$NioByteUnsafe.read(AbstractNioByteChannel.java:122)\r\n\tat io.netty.channel.nio.NioEventLoop.processSelectedKey(NioEventLoop.java:645)\r\n\tat io.netty.channel.nio.NioEventLoop.processSelectedKeysOptimized(NioEventLoop.java:580)\r\n\tat io.netty.channel.nio.NioEventLoop.processSelectedKeys(NioEventLoop.java:497)\r\n\tat io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:459)\r\n\tat io.netty.util.concurrent.SingleThreadEventExecutor$5.run(SingleThreadEventExecutor.java:858)\r\n\tat io.netty.util.concurrent.DefaultThreadFactory$DefaultRunnableDecorator.run(DefaultThreadFactory.java:138)\r\n\t... 1 more{code}\r\nSomething that jumps out to me is the {{DefaultMaxMessagesRecvByteBufAllocator}} (and the {{AdaptiveRecvByteBufAllocator}} in SPARK-24989): maybe there's something about these failing workloads which is leading to significant space wasting in receive buffers, causing tiny blocks to experience huge bloat in space requirements?","from":"developer"},{"body":"User 'Ngone51' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/32287","from":"developer"},{"body":"User 'Ngone51' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/32287","from":"developer"},{"body":"Issue resolved by pull request 32287\n[https://github.com/apache/spark/pull/32287]","from":"developer"}],"created":"2019-06-10T17:24:25.000+0000","description":"ShuffleBlockFetcherIterator has logic to limit the number of simultaneous block fetches. By default, this logic tries to keep the number of outstanding block fetches [beneath a data size limit|https://github.com/apache/spark/blob/v2.4.3/core/src/main/scala/org/apache/spark/storage/ShuffleBlockFetcherIterator.scala#L274] ({{maxBytesInFlight}}). However, this limiting does not take fixed overheads into account: even though a remote block might be, say, 4KB, there are certain fixed-size internal overheads due to Netty buffer sizes which may cause the actual space requirements to be larger.\r\n\r\nAs a result, if a map stage produces a huge number of extremely tiny blocks then we may see errors like\r\n{code:java}\r\norg.apache.spark.shuffle.FetchFailedException: failed to allocate 16777216 byte(s) of direct memory (used: 39325794304, max: 39325794304)\r\nat org.apache.spark.storage.ShuffleBlockFetcherIterator.throwFetchFailedException(ShuffleBlockFetcherIterator.scala:554)\r\nat org.apache.spark.storage.ShuffleBlockFetcherIterator.next(ShuffleBlockFetcherIterator.scala:485)\r\n[...]\r\nCaused by: io.netty.util.internal.OutOfDirectMemoryError: failed to allocate 16777216 byte(s) of direct memory (used: 39325794304, max: 39325794304)\r\nat io.netty.util.internal.PlatformDependent.incrementMemoryCounter(PlatformDependent.java:640)\r\nat io.netty.util.internal.PlatformDependent.allocateDirectNoCleaner(PlatformDependent.java:594)\r\nat io.netty.buffer.PoolArena$DirectArena.allocateDirect(PoolArena.java:764)\r\nat io.netty.buffer.PoolArena$DirectArena.newChunk(PoolArena.java:740)\r\nat io.netty.buffer.PoolArena.allocateNormal(PoolArena.java:244)\r\nat io.netty.buffer.PoolArena.allocate(PoolArena.java:226)\r\nat io.netty.buffer.PoolArena.allocate(PoolArena.java:146)\r\nat io.netty.buffer.PooledByteBufAllocator.newDirectBuffer(PooledByteBufAllocator.java:324)\r\n[...]{code}\r\nSPARK-24989 is another report of this problem (but with a different proposed fix).\r\n\r\nThis problem can currently be mitigated by setting {{spark.reducer.maxReqsInFlight}} to some some non-IntMax value (SPARK-6166), but this additional manual configuration step is cumbersome.\r\n\r\nInstead, I think that Spark should take these fixed overheads into account in the {{maxBytesInFlight}} calculation: instead of using blocks' actual sizes, use {{Math.min(blockSize, minimumNettyBufferSize)}}. There might be some tricky details involved to make this work on all configurations (e.g. to use a different minimum when direct buffers are disabled, etc.), but I think the core idea behind the fix is pretty simple.\r\n\r\nThis will improve Spark's stability and removes configuration / tuning burden from end users.","issue_id":"13238597","key":"SPARK-27991","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2021-05-20T04:27:12.000+0000","role":"fixed_distractor","summary":"ShuffleBlockFetcherIterator should take Netty constant-factor overheads into account when limiting number of simultaneous block fetches"} {"case_id":"13239002","cluster":"DISTRACTOR-SPARK-28025","comments":[{"body":"Same problem. A different part of the code.","created":"2019-06-12T10:40:10.650+0000"},{"body":"I reproduced  the issue in a  spark-shell session:\r\n{code:java}\r\n____ __\r\n/ __/__ ___ _____/ /__\r\n_\\ \\/ _ \\/ _ `/ __/ '_/\r\n/___/ .__/\\_,_/_/ /_/\\_\\ version 2.4.3\r\n/_/\r\n\r\nscala> import org.apache.spark.sql.execution.streaming._\r\nimport org.apache.spark.sql.execution.streaming._\r\n\r\nscala> val hadoopConf = spark.sparkContext.hadoopConfiguration\r\n\r\nscala> import org.apache.spark.sql.internal.SQLConf\r\nimport org.apache.spark.sql.internal.SQLConf\r\n\r\nscala> SQLConf.STREAMING_CHECKPOINT_FILE_MANAGER_CLASS.parent.key\r\nres1: String = spark.sql.streaming.checkpointFileManagerClass\r\n\r\nscala> hadoopConf.getSQLConf.STREAMING_CHECKPOINT_FILE_MANAGER_CLASS.parent.key)\r\nres2: String = null\r\n\r\n// mount point for the shared PVC: /storage\r\nscala> val glusterCpfm = new org.apache.hadoop.fs.Path(\"/storage/crc-store\")\r\nglusterCpfm: org.apache.hadoop.fs.Path = /storage/crc-store\r\n\r\nscala> val glusterfm = CheckpointFileManager.create(glusterCpfm, hadoopConf)\r\nglusterfm: org.apache.spark.sql.execution.streaming.CheckpointFileManager = org.apache.spark.sql.execution.streaming.FileContextBasedCheckpointFileManager@28d00f54\r\n\r\nscala> glusterfm.isLocal\r\nres17: Boolean = true\r\n\r\nscala> glusterfm.mkdirs(glusterCpfm)\r\n\r\nscala> val atomicFile = glusterfm.createAtomic(new org.apache.hadoop.fs.Path(\"/storage/crc-store/file.log\"), false)\r\natomicFile: org.apache.spark.sql.execution.streaming.CheckpointFileManager.CancellableFSDataOutputStream = org.apache.spark.sql.execution.streaming.CheckpointFileManager$RenameBasedFSDataOutputStream@1c6e065\r\n\r\nscala> atomicFile.writeChars(\"Hello, World\")\r\n\r\nscala> atomicFile.close\r\n\r\n/**\r\n* Inspect the file system\r\n*\r\n* $ cat file.log\r\n* Hello, World\r\n* $ ls -al\r\n* total 5\r\n* drwxr-sr-x. 2 jboss 2000 85 Jun 12 09:44 .\r\n* drwxrwsr-x. 8 root 2000 4096 Jun 12 09:42 ..\r\n* -rw-r--r--. 1 jboss 2000 12 Jun 12 09:44 ..file.log.c6f90863-77d2-494e-b1cc-0d0ed1344f74.tmp.crc\r\n* -rw-r--r--. 1 jboss 2000 24 Jun 12 09:44 file.log\r\n**/\r\n\r\n// Delete the file -- simulate the operation done by the HDFSBackedStateStoreProvider#cleanup\r\n\r\nscala> glusterfm.delete(new org.apache.hadoop.fs.Path(\"/storage/crc-store/file.log\"))\r\n\r\n/**\r\n* Inspect the file system -> .crc file left behind\r\n* $ ls -al\r\n* total 9\r\n* drwxr-sr-x. 2 jboss 2000 4096 Jun 12 09:46 .\r\n* drwxrwsr-x. 8 root 2000 4096 Jun 12 09:42 ..\r\n* -rw-r--r--. 1 jboss 2000 12 Jun 12 09:44 ..file.log.c6f90863-77d2-494e-b1cc-0d0ed1344f74.tmp.crc\r\n**/\r\n{code}","created":"2019-06-12T10:50:08.379+0000"},{"body":"[~gmaas]\r\n\r\nNice finding. Would you like to submit a PR for this? Thanks!\r\n\r\n(If you would want to defer to someone, please ping me so that I could take this up.)","created":"2019-06-12T12:36:19.534+0000"},{"body":"looking at the previous patch, you don't need to call exists() Before the delete as delete is required to be a no-op if the source isnt' there. Saves the cost of a HEAD if you are using an object store as a destination.","created":"2019-06-12T15:45:46.443+0000"},{"body":"There is a workaround (avoid creating crc files if you dont want, in certain envs it is the default [https://cloud.google.com/blog/products/storage-data-transfer/new-file-checksum-feature-lets-you-validate-data-transfers-between-hdfs-and-cloud-storage]) by setting `--conf spark.hadoop.spark.sql.streaming.checkpointFileManagerClass=org.apache.spark.sql.execution.streaming.FileSystemBasedCheckpointFileManager` when using local fs\r\n\r\nand modifying the FileSystemBasedCheckpointFileManager to run `fs.setWriteChecksum(false)` after fs is created.\r\n\r\nReason is the FileContextBasedCheckpointFileManager will use ChecksumFS ([https://github.com/apache/hadoop/blob/73746c5da76d5e39df131534a1ec35dfc5d2529b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/ChecksumFs.java]) under the hoods which will ignore \"\r\n\r\nCreateOpts.checksumParam(ChecksumOpt.createDisabled())\" passed. These settings will only avoid creating checksums for the checksums themshelves and only if the underlying fs supports it. However, it will create the checksum file in any case.\r\n\r\nFileSystemBasedCheckpointFileManager uses [https://github.com/apache/hadoop/blob/73746c5da76d5e39df131534a1ec35dfc5d2529b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/ChecksumFileSystem.java] which allows to avoid checksums when creating files if you set that flag.\r\n\r\nNote that the crc is created when the tmp file is created not during rename or mv.\r\n\r\nI will create a PR shortly. ","created":"2019-06-12T18:17:12.873+0000"},{"body":"Personally I would respect the reporter and encourage to submit a patch by theirselves so that we could have broader contributors, but I assume you two are colleague (same employer) and discussed to decide who to submit a PR.","created":"2019-06-12T20:47:55.298+0000"},{"body":"[~kabhwan] yes we discussed this internally when trying to solve the issue, otherwise I would not intervene. I fully respect people colleagues or not ;)","created":"2019-06-12T23:27:38.194+0000"},{"body":"I just found out that the following config options would suffice to avoid creating crcs for the given case:\r\n\r\n--conf spark.hadoop.spark.sql.streaming.checkpointFileManagerClass=org.apache.spark.sql.execution.streaming.FileSystemBasedCheckpointFileManager --conf spark.hadoop.fs.file.impl=org.apache.hadoop.fs.RawLocalFileSystem\r\n\r\nSo no need to make a PR to disable crc emission if user wants to for this case. I could make one to cover all cases to be able to enable/disable crcs if needed. However, the filesystem hierarchy in hadoop is a bit inconsistent. So the LocalFileSystem will use a a FilterFileSystem that has the flag setWriteChecksum but the DistributedFileSystem does not have it and it is controlled by property [https://github.com/apache/hadoop/blob/533138718cc05b78e0afe583d7a9bd30e8a48fdc/hadoop-hdfs-project/hadoop-hdfs/src/test/java/org/apache/hadoop/hdfs/client/impl/TestBlockReaderLocal.java#L620]\r\n\r\n`dfs.checksum.type`\r\n\r\nBtw there is a nice summary on the topic in chapter 5 of the book: hadoop the definitive guide. \r\n\r\nLooking at the old PR, I am not sure it would work with the LocalFs as it seems that the crc file will be renamed by default as the underlying system supports checksums by default.  ","created":"2019-06-13T10:26:05.035+0000"},{"body":"I'm also affected by performance issue caused by .crc files leak in checkpoint directory. [~skonto] thank you for the workaround, it works.\r\n\r\nWould it be possible to implement cleaning of crc files, when one needs the checksum ?","created":"2019-07-25T15:21:21.229+0000"},{"body":"I still think it should rename crc file as well (so that it can be removed altogether if we purge batches) or just remove the crc file when it renames tmp file. If we want to stick with workaround, at least the workaround should be added to the guide doc. It doesn't look like being resolved.","created":"2019-08-18T08:01:25.610+0000"},{"body":"@[~dongjoon] [~zsxwing] this needs to be re-opened. When using the workaround we recently hit this issue:\r\n\r\n[https://github.com/broadinstitute/gatk/issues/1389]\r\n\r\nwhich can be fixed easily with a derived class like in this PR: [https://github.com/broadinstitute/gatk/pull/1421/files]\r\n\r\nbut this is a bit of inconvenient. \r\n\r\nHowever, I believe as well that this should be fixed in Spark (less surprises) otherwise we need to document it as [~kabhwan] said above.","created":"2019-08-30T11:52:49.403+0000"},{"body":"[~skonto]\r\n\r\nPlease take a look at my PR as my PR didn't follow your workaround. We identified which Hadoop issue we are facing, and took a workaround as deleting crc file manually.","created":"2019-08-30T12:48:28.713+0000"},{"body":"Has anyone considered enhancing org.apache.hadoop.fs.ChecksumFileSystem to say \"if \"file.bytes-per-checksum\" == 0 then checksums are disabled?\r\n\r\nCurrently it fails if bytes per CRC <= 0, but you could make the 0 value a switch to say \"none\".","created":"2019-08-30T12:53:26.349+0000"},{"body":"[~kabhwan] cool I have a look.","created":"2019-08-30T14:06:51.225+0000"},{"body":"[~skonto], this one: [https://github.com/apache/spark/pull/25488]","created":"2019-08-30T14:09:19.566+0000"},{"body":"Thanks I will have a look :)","created":"2019-08-30T14:16:13.752+0000"},{"body":"FYI, I just submitted a patch for HADOOP-16255. Hope we can get rid of workaround sooner.","created":"2019-08-31T04:29:59.901+0000"},{"body":"[~stevel@apache.org] I think adding \"file.bytes-per-checksum = 0\" possibility would be a workaround (at least here). The whole point why one choose ChecksumFileSystem is to have checksum.","created":"2019-09-02T08:53:28.643+0000"},{"body":"[~stevel@apache.org]\r\n\r\nSorry to ping you on the old issue. According to the line we create a temp file, I guess we intend to \"disable\" creating CRC file, so actually it should never be an issue if thing works as intended, whereas it isn't.\r\n\r\n[https://github.com/apache/spark/blob/5effa8ea261ba59214afedc2853d1b248b330ca6/sql/core/src/main/scala/org/apache/spark/sql/execution/streaming/CheckpointFileManager.scala#L306-L311]\r\n{code:java}\r\n// fc = FileContext.getFileContext(...)\r\n\r\n fc.create(\r\n path, EnumSet.of(CREATE, OVERWRITE), CreateOpts.checksumParam(ChecksumOpt.createDisabled())){code}\r\nIs this something we should fix in Hadoop side, or the option is incorrect and we should apply other option?","created":"2020-10-08T08:12:33.065+0000"}],"conversations":[{"body":"The HDFSBackedStateStoreProvider when using the default CheckpointFileManager is leaving '.crc' files behind. There's a .crc file created for each `atomicFile` operation of the CheckpointFileManager.\r\n\r\nOver time, the number of files becomes very large. It makes the state store file system constantly increase in size and, in our case, deteriorates the file system performance.\r\n\r\nHere's a sample of one of our spark storage volumes after 2 days of execution (4 stateful streaming jobs, each on a different sub-dir):\r\n # \r\n{noformat}\r\nTotal files in PVC (used for checkpoints and state store)\r\n$find . | wc -l\r\n431796\r\n\r\n# .crc files\r\n$find . -name \"*.crc\" | wc -l\r\n418053{noformat}\r\n\r\nWith each .crc file taking one storage block, the used storage runs into the GBs of data.\r\n\r\nThese jobs are running on Kubernetes. Our shared storage provider, GlusterFS, shows serious performance deterioration with this large number of files:\r\n{noformat}\r\nDEBUG HDFSBackedStateStoreProvider: fetchFiles() took 29164ms{noformat}\r\n ","from":"reporter","subject":"HDFSBackedStateStoreProvider should not leak .crc files "},{"body":"Same problem. A different part of the code.","from":"developer"},{"body":"I reproduced  the issue in a  spark-shell session:\r\n{code:java}\r\n____ __\r\n/ __/__ ___ _____/ /__\r\n_\\ \\/ _ \\/ _ `/ __/ '_/\r\n/___/ .__/\\_,_/_/ /_/\\_\\ version 2.4.3\r\n/_/\r\n\r\nscala> import org.apache.spark.sql.execution.streaming._\r\nimport org.apache.spark.sql.execution.streaming._\r\n\r\nscala> val hadoopConf = spark.sparkContext.hadoopConfiguration\r\n\r\nscala> import org.apache.spark.sql.internal.SQLConf\r\nimport org.apache.spark.sql.internal.SQLConf\r\n\r\nscala> SQLConf.STREAMING_CHECKPOINT_FILE_MANAGER_CLASS.parent.key\r\nres1: String = spark.sql.streaming.checkpointFileManagerClass\r\n\r\nscala> hadoopConf.getSQLConf.STREAMING_CHECKPOINT_FILE_MANAGER_CLASS.parent.key)\r\nres2: String = null\r\n\r\n// mount point for the shared PVC: /storage\r\nscala> val glusterCpfm = new org.apache.hadoop.fs.Path(\"/storage/crc-store\")\r\nglusterCpfm: org.apache.hadoop.fs.Path = /storage/crc-store\r\n\r\nscala> val glusterfm = CheckpointFileManager.create(glusterCpfm, hadoopConf)\r\nglusterfm: org.apache.spark.sql.execution.streaming.CheckpointFileManager = org.apache.spark.sql.execution.streaming.FileContextBasedCheckpointFileManager@28d00f54\r\n\r\nscala> glusterfm.isLocal\r\nres17: Boolean = true\r\n\r\nscala> glusterfm.mkdirs(glusterCpfm)\r\n\r\nscala> val atomicFile = glusterfm.createAtomic(new org.apache.hadoop.fs.Path(\"/storage/crc-store/file.log\"), false)\r\natomicFile: org.apache.spark.sql.execution.streaming.CheckpointFileManager.CancellableFSDataOutputStream = org.apache.spark.sql.execution.streaming.CheckpointFileManager$RenameBasedFSDataOutputStream@1c6e065\r\n\r\nscala> atomicFile.writeChars(\"Hello, World\")\r\n\r\nscala> atomicFile.close\r\n\r\n/**\r\n* Inspect the file system\r\n*\r\n* $ cat file.log\r\n* Hello, World\r\n* $ ls -al\r\n* total 5\r\n* drwxr-sr-x. 2 jboss 2000 85 Jun 12 09:44 .\r\n* drwxrwsr-x. 8 root 2000 4096 Jun 12 09:42 ..\r\n* -rw-r--r--. 1 jboss 2000 12 Jun 12 09:44 ..file.log.c6f90863-77d2-494e-b1cc-0d0ed1344f74.tmp.crc\r\n* -rw-r--r--. 1 jboss 2000 24 Jun 12 09:44 file.log\r\n**/\r\n\r\n// Delete the file -- simulate the operation done by the HDFSBackedStateStoreProvider#cleanup\r\n\r\nscala> glusterfm.delete(new org.apache.hadoop.fs.Path(\"/storage/crc-store/file.log\"))\r\n\r\n/**\r\n* Inspect the file system -> .crc file left behind\r\n* $ ls -al\r\n* total 9\r\n* drwxr-sr-x. 2 jboss 2000 4096 Jun 12 09:46 .\r\n* drwxrwsr-x. 8 root 2000 4096 Jun 12 09:42 ..\r\n* -rw-r--r--. 1 jboss 2000 12 Jun 12 09:44 ..file.log.c6f90863-77d2-494e-b1cc-0d0ed1344f74.tmp.crc\r\n**/\r\n{code}","from":"developer"},{"body":"[~gmaas]\r\n\r\nNice finding. Would you like to submit a PR for this? Thanks!\r\n\r\n(If you would want to defer to someone, please ping me so that I could take this up.)","from":"developer"},{"body":"looking at the previous patch, you don't need to call exists() Before the delete as delete is required to be a no-op if the source isnt' there. Saves the cost of a HEAD if you are using an object store as a destination.","from":"developer"},{"body":"There is a workaround (avoid creating crc files if you dont want, in certain envs it is the default [https://cloud.google.com/blog/products/storage-data-transfer/new-file-checksum-feature-lets-you-validate-data-transfers-between-hdfs-and-cloud-storage]) by setting `--conf spark.hadoop.spark.sql.streaming.checkpointFileManagerClass=org.apache.spark.sql.execution.streaming.FileSystemBasedCheckpointFileManager` when using local fs\r\n\r\nand modifying the FileSystemBasedCheckpointFileManager to run `fs.setWriteChecksum(false)` after fs is created.\r\n\r\nReason is the FileContextBasedCheckpointFileManager will use ChecksumFS ([https://github.com/apache/hadoop/blob/73746c5da76d5e39df131534a1ec35dfc5d2529b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/ChecksumFs.java]) under the hoods which will ignore \"\r\n\r\nCreateOpts.checksumParam(ChecksumOpt.createDisabled())\" passed. These settings will only avoid creating checksums for the checksums themshelves and only if the underlying fs supports it. However, it will create the checksum file in any case.\r\n\r\nFileSystemBasedCheckpointFileManager uses [https://github.com/apache/hadoop/blob/73746c5da76d5e39df131534a1ec35dfc5d2529b/hadoop-common-project/hadoop-common/src/main/java/org/apache/hadoop/fs/ChecksumFileSystem.java] which allows to avoid checksums when creating files if you set that flag.\r\n\r\nNote that the crc is created when the tmp file is created not during rename or mv.\r\n\r\nI will create a PR shortly. ","from":"developer"},{"body":"Personally I would respect the reporter and encourage to submit a patch by theirselves so that we could have broader contributors, but I assume you two are colleague (same employer) and discussed to decide who to submit a PR.","from":"developer"},{"body":"[~kabhwan] yes we discussed this internally when trying to solve the issue, otherwise I would not intervene. I fully respect people colleagues or not ;)","from":"developer"},{"body":"I just found out that the following config options would suffice to avoid creating crcs for the given case:\r\n\r\n--conf spark.hadoop.spark.sql.streaming.checkpointFileManagerClass=org.apache.spark.sql.execution.streaming.FileSystemBasedCheckpointFileManager --conf spark.hadoop.fs.file.impl=org.apache.hadoop.fs.RawLocalFileSystem\r\n\r\nSo no need to make a PR to disable crc emission if user wants to for this case. I could make one to cover all cases to be able to enable/disable crcs if needed. However, the filesystem hierarchy in hadoop is a bit inconsistent. So the LocalFileSystem will use a a FilterFileSystem that has the flag setWriteChecksum but the DistributedFileSystem does not have it and it is controlled by property [https://github.com/apache/hadoop/blob/533138718cc05b78e0afe583d7a9bd30e8a48fdc/hadoop-hdfs-project/hadoop-hdfs/src/test/java/org/apache/hadoop/hdfs/client/impl/TestBlockReaderLocal.java#L620]\r\n\r\n`dfs.checksum.type`\r\n\r\nBtw there is a nice summary on the topic in chapter 5 of the book: hadoop the definitive guide. \r\n\r\nLooking at the old PR, I am not sure it would work with the LocalFs as it seems that the crc file will be renamed by default as the underlying system supports checksums by default.  ","from":"developer"},{"body":"I'm also affected by performance issue caused by .crc files leak in checkpoint directory. [~skonto] thank you for the workaround, it works.\r\n\r\nWould it be possible to implement cleaning of crc files, when one needs the checksum ?","from":"developer"},{"body":"I still think it should rename crc file as well (so that it can be removed altogether if we purge batches) or just remove the crc file when it renames tmp file. If we want to stick with workaround, at least the workaround should be added to the guide doc. It doesn't look like being resolved.","from":"developer"},{"body":"@[~dongjoon] [~zsxwing] this needs to be re-opened. When using the workaround we recently hit this issue:\r\n\r\n[https://github.com/broadinstitute/gatk/issues/1389]\r\n\r\nwhich can be fixed easily with a derived class like in this PR: [https://github.com/broadinstitute/gatk/pull/1421/files]\r\n\r\nbut this is a bit of inconvenient. \r\n\r\nHowever, I believe as well that this should be fixed in Spark (less surprises) otherwise we need to document it as [~kabhwan] said above.","from":"developer"},{"body":"[~skonto]\r\n\r\nPlease take a look at my PR as my PR didn't follow your workaround. We identified which Hadoop issue we are facing, and took a workaround as deleting crc file manually.","from":"developer"},{"body":"Has anyone considered enhancing org.apache.hadoop.fs.ChecksumFileSystem to say \"if \"file.bytes-per-checksum\" == 0 then checksums are disabled?\r\n\r\nCurrently it fails if bytes per CRC <= 0, but you could make the 0 value a switch to say \"none\".","from":"developer"},{"body":"[~kabhwan] cool I have a look.","from":"developer"},{"body":"[~skonto], this one: [https://github.com/apache/spark/pull/25488]","from":"developer"},{"body":"Thanks I will have a look :)","from":"developer"},{"body":"FYI, I just submitted a patch for HADOOP-16255. Hope we can get rid of workaround sooner.","from":"developer"},{"body":"[~stevel@apache.org] I think adding \"file.bytes-per-checksum = 0\" possibility would be a workaround (at least here). The whole point why one choose ChecksumFileSystem is to have checksum.","from":"developer"},{"body":"[~stevel@apache.org]\r\n\r\nSorry to ping you on the old issue. According to the line we create a temp file, I guess we intend to \"disable\" creating CRC file, so actually it should never be an issue if thing works as intended, whereas it isn't.\r\n\r\n[https://github.com/apache/spark/blob/5effa8ea261ba59214afedc2853d1b248b330ca6/sql/core/src/main/scala/org/apache/spark/sql/execution/streaming/CheckpointFileManager.scala#L306-L311]\r\n{code:java}\r\n// fc = FileContext.getFileContext(...)\r\n\r\n fc.create(\r\n path, EnumSet.of(CREATE, OVERWRITE), CreateOpts.checksumParam(ChecksumOpt.createDisabled())){code}\r\nIs this something we should fix in Hadoop side, or the option is incorrect and we should apply other option?","from":"developer"}],"created":"2019-06-12T10:39:08.000+0000","description":"The HDFSBackedStateStoreProvider when using the default CheckpointFileManager is leaving '.crc' files behind. There's a .crc file created for each `atomicFile` operation of the CheckpointFileManager.\r\n\r\nOver time, the number of files becomes very large. It makes the state store file system constantly increase in size and, in our case, deteriorates the file system performance.\r\n\r\nHere's a sample of one of our spark storage volumes after 2 days of execution (4 stateful streaming jobs, each on a different sub-dir):\r\n # \r\n{noformat}\r\nTotal files in PVC (used for checkpoints and state store)\r\n$find . | wc -l\r\n431796\r\n\r\n# .crc files\r\n$find . -name \"*.crc\" | wc -l\r\n418053{noformat}\r\n\r\nWith each .crc file taking one storage block, the used storage runs into the GBs of data.\r\n\r\nThese jobs are running on Kubernetes. Our shared storage provider, GlusterFS, shows serious performance deterioration with this large number of files:\r\n{noformat}\r\nDEBUG HDFSBackedStateStoreProvider: fetchFiles() took 29164ms{noformat}\r\n ","issue_id":"13239002","key":"SPARK-28025","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2019-08-23T06:12:22.000+0000","role":"fixed_distractor","summary":"HDFSBackedStateStoreProvider should not leak .crc files "} {"case_id":"12732446","cluster":"DISTRACTOR-SPARK-2892","comments":[{"body":"I see the same with 1.0.2 streaming, with or without stopGracefully = true\n\nssc.stop(stopSparkContext = false, stopGracefully = true)\n\nERROR 08:26:21,139 Deregistered receiver for stream 0: Stopped by driver\n WARN 08:26:21,211 Stopped executor without error\n WARN 08:26:21,213 All of the receivers have not deregistered, Map(0 -> ReceiverInfo(0,ActorReceiver-0,null,false,host,Stopped by driver,))\n","created":"2014-09-04T13:08:15.364+0000"},{"body":"I wonder if the ERROR should be a WARN or INFO since it occurs as a result of ReceiverSupervisorImpl receiving a StopReceiver, and \" Deregistered receiver for stream\" seems like the expected behavior.\n\n\nDEBUG 13:00:22,418 Stopping JobScheduler\n INFO 13:00:22,441 Received stop signal\n INFO 13:00:22,441 Sent stop signal to all 1 receivers\n INFO 13:00:22,442 Stopping receiver with message: Stopped by driver: \n INFO 13:00:22,442 Called receiver onStop\n INFO 13:00:22,443 Deregistering receiver 0\nERROR 13:00:22,445 Deregistered receiver for stream 0: Stopped by driver\n INFO 13:00:22,445 Stopped receiver 0","created":"2014-09-04T16:58:08.896+0000"},{"body":"I intended it to be ERROR to catch such issues where receivers dont stop properly. If a program is supposed to shutdown after stopping the streaming context, then this probably not much of a problem as everything of Spark is torn down anyways. But a SparkContext is going to be reused, then this is indeed a problem. ","created":"2014-09-05T17:54:27.530+0000"},{"body":"Any update on this? Will it get fixed for 1.0.3 or 1.1.0","created":"2014-09-18T15:28:45.516+0000"},{"body":"Update?","created":"2014-11-25T16:58:42.788+0000"},{"body":"It looks like this one and the issue mentioned in SPARK-4802 (ReceiverInfo removal at ReceiverTracker upon deregistering receiver) are related. I believe the following warning message is the result of receiverInfo not being removed at ReceiverTracker by the ReceiverTrackerActor when the corresponding receiver is deregistered.\n\n\"WARN ReceiverTracker: All of the receivers have not deregistered, Map(0 -> ReceiverInfo(0,SocketReceiver-0,null,false,localhost,Stopped by driver,))\"\n\nFrom what I can see so far, closing the streaming context stops the receiver only in \"local\" mode.\n\nIn \"cluster\" mode, using the Spark standalone cluster I noticed that when the ReceiverTracker at the driver sends the \"StopReceiver\" message as a result of streaming context close, it couldn't reach to the ReceiverSupervisorImpl's actor that is running at the executor node. At the same time, the ReceiverSupervisorImpl at the executor could send the messages such as RegisterReceiver, AddBlock back to the ReceiverTrackerActor at the driver.\n\nIt would be great if someone could explain what might be going on from ReceiverTracker -> ReceiverSupervisorImpl actor at executor when sending the stop signal in the distributed mode case.\n\nThanks!","created":"2014-12-10T00:54:07.615+0000"},{"body":"This does indeed look like the same issue. Since SPARK-4802 has an open PR I think it should continue there.","created":"2014-12-10T06:23:01.167+0000"},{"body":"[~srowen] SPARK-4802 is only related to the receiverInfo not being removed. This issue is actually much more critical, given that Receivers do not seem to stop other than in local mode. Please reopen.\n","created":"2014-12-10T16:48:58.163+0000"},{"body":"To add more info:\n\nWhen the ReceiverTracker sends the \"StopReceiver\" message to the receiver actor at the executor, it's ReceiverLauncher thread always times out and I notice the corresponding job is cancelled only because of stopping the DAGScheduler. This throws the exception[1] while at the executor side the worker node throws this info[2]\n\nException[1]:\nINFO sparkDriver-akka.actor.default-dispatcher-14 cluster.SparkDeploySchedulerBackend - Asking each executor to shut down\n15:06:53,783 1.1.0.SNAP INFO Thread-40 scheduler.DAGScheduler - Job 1 failed: start at SparkDriver.java:109, took 72.739141 s\nException in thread \"Thread-40\" org.apache.spark.SparkException: Job cancelled because SparkContext was shut down\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$cleanUpAfterSchedulerStop$1.apply(DAGScheduler.scala:702)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$cleanUpAfterSchedulerStop$1.apply(DAGScheduler.scala:701)\n\tat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\n\tat org.apache.spark.scheduler.DAGScheduler.cleanUpAfterSchedulerStop(DAGScheduler.scala:701)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor.postStop(DAGScheduler.scala:1428)\n\tat akka.actor.Actor$class.aroundPostStop(Actor.scala:475)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor.aroundPostStop(DAGScheduler.scala:1375)\n\tat akka.actor.dungeon.FaultHandling$class.akka$actor$dungeon$FaultHandling$$finishTerminate(FaultHandling.scala:210)\n\tat akka.actor.dungeon.FaultHandling$class.terminate(FaultHandling.scala:172)\n\tat akka.actor.ActorCell.terminate(ActorCell.scala:369)\n\tat akka.actor.ActorCell.invokeAll$1(ActorCell.scala:462)\n\tat akka.actor.ActorCell.systemInvoke(ActorCell.scala:478)\n\tat akka.dispatch.Mailbox.processAllSystemMessages(Mailbox.scala:263)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:219)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:393)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n\ninfo[2]\nINFO LocalActorRef: Message [akka.remote.transport.AssociationHandle$Disassociated] from Actor[akka://sparkWorker/deadLetters] to Actor[akka://sparkWorker/system/transports/akkaprotocolmanager.tcp0/akkaProtocol-tcp%3A%2F%2FsparkWorker%40127.0.0.1%3A52219-2#1619424601] was not delivered. [1] dead letters encountered. This logging can be turned off or adjusted with configuration settings 'akka.log-dead-letters' and 'akka.log-dead-letters-during-shutdown'.\n14/12/10 15:06:53 ERROR EndpointWriter: AssociationError [akka.tcp://sparkWorker@localhost:51262] <- [akka.tcp://sparkExecutor@localhost:52217]: Error [Shut down address: akka.tcp://sparkExecutor@localhost:52217] [\nakka.remote.ShutDownAssociation: Shut down address: akka.tcp://sparkExecutor@localhost:52217\nCaused by: akka.remote.transport.Transport$InvalidAssociationException: The remote system terminated the association because it is shutting down.\n\nPlease note that I am running everything on localhost and the same thing works fine on \"local\" mode and the above issue only arises on \"cluster\" mode. I tried changing the hostname to 127.0.0.1 but noticed the same.\n\nAny clues on what might be going on here would help a lot.\nThanks!","created":"2014-12-10T23:25:02.802+0000"},{"body":"OK. The issues may have a common cause but that can be deferred until the other JIRA is resolved. If it turns out to resolve this, great.","created":"2014-12-11T16:59:59.346+0000"},{"body":"[~ilayaperumalg] Now that SPARK-4802 has been solved could you check whether this issue has been resolved?\n\n","created":"2014-12-24T21:48:42.229+0000"},{"body":"[~ilayaperumalg] Ping, any thoughts?","created":"2014-12-31T01:56:28.989+0000"},{"body":"I think that I've found the reason that this works in local mode and not on a real cluster; see SPARK-5035 / https://github.com/apache/spark/pull/3857","created":"2014-12-31T09:24:18.654+0000"},{"body":"[~joshrosen] This indeed fixes the issue on stopping the receiver. thanks!\n\nP.S: I deleted my prior comment from the JIRA as the error I noticed was due to some setup issue on my cluster environment.","created":"2014-12-31T12:07:55.184+0000"},{"body":"I believe this is fixed by SPARK-5035, so I'm closing this.","created":"2015-02-09T20:46:10.346+0000"}],"conversations":[{"body":"Running NetworkWordCount with\n{quote} \nssc.start(); Thread.sleep(10000); ssc.stop(stopSparkContext = false); Thread.sleep(60000)\n{quote}\n\ngives the following error\n\n{quote}\n14/08/06 18:37:13 INFO TaskSetManager: Finished task 0.0 in stage 0.0 (TID 0) in 10047 ms on localhost (1/1)\n14/08/06 18:37:13 INFO DAGScheduler: Stage 0 (runJob at ReceiverTracker.scala:275) finished in 10.056 s\n14/08/06 18:37:13 INFO TaskSchedulerImpl: Removed TaskSet 0.0, whose tasks have all completed, from pool\n14/08/06 18:37:13 INFO SparkContext: Job finished: runJob at ReceiverTracker.scala:275, took 10.179263 s\n14/08/06 18:37:13 INFO ReceiverTracker: All of the receivers have been terminated\n14/08/06 18:37:13 WARN ReceiverTracker: All of the receivers have not deregistered, Map(0 -> ReceiverInfo(0,SocketReceiver-0,null,false,localhost,Stopped by driver,))\n14/08/06 18:37:13 INFO ReceiverTracker: ReceiverTracker stopped\n14/08/06 18:37:13 INFO JobGenerator: Stopping JobGenerator immediately\n14/08/06 18:37:13 INFO RecurringTimer: Stopped timer for JobGenerator after time 1407375433000\n14/08/06 18:37:13 INFO JobGenerator: Stopped JobGenerator\n14/08/06 18:37:13 INFO JobScheduler: Stopped JobScheduler\n14/08/06 18:37:13 INFO StreamingContext: StreamingContext stopped successfully\n14/08/06 18:37:43 INFO SocketReceiver: Stopped receiving\n14/08/06 18:37:43 INFO SocketReceiver: Closed socket to localhost:9999\n{quote}","from":"reporter","subject":"Socket Receiver does not stop when streaming context is stopped"},{"body":"I see the same with 1.0.2 streaming, with or without stopGracefully = true\n\nssc.stop(stopSparkContext = false, stopGracefully = true)\n\nERROR 08:26:21,139 Deregistered receiver for stream 0: Stopped by driver\n WARN 08:26:21,211 Stopped executor without error\n WARN 08:26:21,213 All of the receivers have not deregistered, Map(0 -> ReceiverInfo(0,ActorReceiver-0,null,false,host,Stopped by driver,))\n","from":"developer"},{"body":"I wonder if the ERROR should be a WARN or INFO since it occurs as a result of ReceiverSupervisorImpl receiving a StopReceiver, and \" Deregistered receiver for stream\" seems like the expected behavior.\n\n\nDEBUG 13:00:22,418 Stopping JobScheduler\n INFO 13:00:22,441 Received stop signal\n INFO 13:00:22,441 Sent stop signal to all 1 receivers\n INFO 13:00:22,442 Stopping receiver with message: Stopped by driver: \n INFO 13:00:22,442 Called receiver onStop\n INFO 13:00:22,443 Deregistering receiver 0\nERROR 13:00:22,445 Deregistered receiver for stream 0: Stopped by driver\n INFO 13:00:22,445 Stopped receiver 0","from":"developer"},{"body":"I intended it to be ERROR to catch such issues where receivers dont stop properly. If a program is supposed to shutdown after stopping the streaming context, then this probably not much of a problem as everything of Spark is torn down anyways. But a SparkContext is going to be reused, then this is indeed a problem. ","from":"developer"},{"body":"Any update on this? Will it get fixed for 1.0.3 or 1.1.0","from":"developer"},{"body":"Update?","from":"developer"},{"body":"It looks like this one and the issue mentioned in SPARK-4802 (ReceiverInfo removal at ReceiverTracker upon deregistering receiver) are related. I believe the following warning message is the result of receiverInfo not being removed at ReceiverTracker by the ReceiverTrackerActor when the corresponding receiver is deregistered.\n\n\"WARN ReceiverTracker: All of the receivers have not deregistered, Map(0 -> ReceiverInfo(0,SocketReceiver-0,null,false,localhost,Stopped by driver,))\"\n\nFrom what I can see so far, closing the streaming context stops the receiver only in \"local\" mode.\n\nIn \"cluster\" mode, using the Spark standalone cluster I noticed that when the ReceiverTracker at the driver sends the \"StopReceiver\" message as a result of streaming context close, it couldn't reach to the ReceiverSupervisorImpl's actor that is running at the executor node. At the same time, the ReceiverSupervisorImpl at the executor could send the messages such as RegisterReceiver, AddBlock back to the ReceiverTrackerActor at the driver.\n\nIt would be great if someone could explain what might be going on from ReceiverTracker -> ReceiverSupervisorImpl actor at executor when sending the stop signal in the distributed mode case.\n\nThanks!","from":"developer"},{"body":"This does indeed look like the same issue. Since SPARK-4802 has an open PR I think it should continue there.","from":"developer"},{"body":"[~srowen] SPARK-4802 is only related to the receiverInfo not being removed. This issue is actually much more critical, given that Receivers do not seem to stop other than in local mode. Please reopen.\n","from":"developer"},{"body":"To add more info:\n\nWhen the ReceiverTracker sends the \"StopReceiver\" message to the receiver actor at the executor, it's ReceiverLauncher thread always times out and I notice the corresponding job is cancelled only because of stopping the DAGScheduler. This throws the exception[1] while at the executor side the worker node throws this info[2]\n\nException[1]:\nINFO sparkDriver-akka.actor.default-dispatcher-14 cluster.SparkDeploySchedulerBackend - Asking each executor to shut down\n15:06:53,783 1.1.0.SNAP INFO Thread-40 scheduler.DAGScheduler - Job 1 failed: start at SparkDriver.java:109, took 72.739141 s\nException in thread \"Thread-40\" org.apache.spark.SparkException: Job cancelled because SparkContext was shut down\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$cleanUpAfterSchedulerStop$1.apply(DAGScheduler.scala:702)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$cleanUpAfterSchedulerStop$1.apply(DAGScheduler.scala:701)\n\tat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\n\tat org.apache.spark.scheduler.DAGScheduler.cleanUpAfterSchedulerStop(DAGScheduler.scala:701)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor.postStop(DAGScheduler.scala:1428)\n\tat akka.actor.Actor$class.aroundPostStop(Actor.scala:475)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor.aroundPostStop(DAGScheduler.scala:1375)\n\tat akka.actor.dungeon.FaultHandling$class.akka$actor$dungeon$FaultHandling$$finishTerminate(FaultHandling.scala:210)\n\tat akka.actor.dungeon.FaultHandling$class.terminate(FaultHandling.scala:172)\n\tat akka.actor.ActorCell.terminate(ActorCell.scala:369)\n\tat akka.actor.ActorCell.invokeAll$1(ActorCell.scala:462)\n\tat akka.actor.ActorCell.systemInvoke(ActorCell.scala:478)\n\tat akka.dispatch.Mailbox.processAllSystemMessages(Mailbox.scala:263)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:219)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:393)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n\ninfo[2]\nINFO LocalActorRef: Message [akka.remote.transport.AssociationHandle$Disassociated] from Actor[akka://sparkWorker/deadLetters] to Actor[akka://sparkWorker/system/transports/akkaprotocolmanager.tcp0/akkaProtocol-tcp%3A%2F%2FsparkWorker%40127.0.0.1%3A52219-2#1619424601] was not delivered. [1] dead letters encountered. This logging can be turned off or adjusted with configuration settings 'akka.log-dead-letters' and 'akka.log-dead-letters-during-shutdown'.\n14/12/10 15:06:53 ERROR EndpointWriter: AssociationError [akka.tcp://sparkWorker@localhost:51262] <- [akka.tcp://sparkExecutor@localhost:52217]: Error [Shut down address: akka.tcp://sparkExecutor@localhost:52217] [\nakka.remote.ShutDownAssociation: Shut down address: akka.tcp://sparkExecutor@localhost:52217\nCaused by: akka.remote.transport.Transport$InvalidAssociationException: The remote system terminated the association because it is shutting down.\n\nPlease note that I am running everything on localhost and the same thing works fine on \"local\" mode and the above issue only arises on \"cluster\" mode. I tried changing the hostname to 127.0.0.1 but noticed the same.\n\nAny clues on what might be going on here would help a lot.\nThanks!","from":"developer"},{"body":"OK. The issues may have a common cause but that can be deferred until the other JIRA is resolved. If it turns out to resolve this, great.","from":"developer"},{"body":"[~ilayaperumalg] Now that SPARK-4802 has been solved could you check whether this issue has been resolved?\n\n","from":"developer"},{"body":"[~ilayaperumalg] Ping, any thoughts?","from":"developer"},{"body":"I think that I've found the reason that this works in local mode and not on a real cluster; see SPARK-5035 / https://github.com/apache/spark/pull/3857","from":"developer"},{"body":"[~joshrosen] This indeed fixes the issue on stopping the receiver. thanks!\n\nP.S: I deleted my prior comment from the JIRA as the error I noticed was due to some setup issue on my cluster environment.","from":"developer"},{"body":"I believe this is fixed by SPARK-5035, so I'm closing this.","from":"developer"}],"created":"2014-08-07T01:44:09.000+0000","description":"Running NetworkWordCount with\n{quote} \nssc.start(); Thread.sleep(10000); ssc.stop(stopSparkContext = false); Thread.sleep(60000)\n{quote}\n\ngives the following error\n\n{quote}\n14/08/06 18:37:13 INFO TaskSetManager: Finished task 0.0 in stage 0.0 (TID 0) in 10047 ms on localhost (1/1)\n14/08/06 18:37:13 INFO DAGScheduler: Stage 0 (runJob at ReceiverTracker.scala:275) finished in 10.056 s\n14/08/06 18:37:13 INFO TaskSchedulerImpl: Removed TaskSet 0.0, whose tasks have all completed, from pool\n14/08/06 18:37:13 INFO SparkContext: Job finished: runJob at ReceiverTracker.scala:275, took 10.179263 s\n14/08/06 18:37:13 INFO ReceiverTracker: All of the receivers have been terminated\n14/08/06 18:37:13 WARN ReceiverTracker: All of the receivers have not deregistered, Map(0 -> ReceiverInfo(0,SocketReceiver-0,null,false,localhost,Stopped by driver,))\n14/08/06 18:37:13 INFO ReceiverTracker: ReceiverTracker stopped\n14/08/06 18:37:13 INFO JobGenerator: Stopping JobGenerator immediately\n14/08/06 18:37:13 INFO RecurringTimer: Stopped timer for JobGenerator after time 1407375433000\n14/08/06 18:37:13 INFO JobGenerator: Stopped JobGenerator\n14/08/06 18:37:13 INFO JobScheduler: Stopped JobScheduler\n14/08/06 18:37:13 INFO StreamingContext: StreamingContext stopped successfully\n14/08/06 18:37:43 INFO SocketReceiver: Stopped receiving\n14/08/06 18:37:43 INFO SocketReceiver: Closed socket to localhost:9999\n{quote}","issue_id":"12732446","key":"SPARK-2892","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-02-09T20:46:10.000+0000","role":"fixed_distractor","summary":"Socket Receiver does not stop when streaming context is stopped"} {"case_id":"13259642","cluster":"DISTRACTOR-SPARK-29302","comments":[{"body":"FileCommitProtocol will assign a filename based on task id, isn't? Should not a speculation task have a different task id?","created":"2019-09-30T23:28:18.361+0000"},{"body":"Can you share a reproducer? or this sounds more like a question instead of a bug.","created":"2019-09-30T23:29:26.393+0000"},{"body":"You can add the code below into FileFormatWriterSuite.\r\n{code:java}\r\n test(\"SPARK-29302: for dynamic partition overwrite, a task will concurrent write a same file\" +\r\n \" with its relative speculation task\") {\r\n withTempDir { f =>\r\n val jobId = SparkHadoopWriterUtils.createJobID(new Date(), 1)\r\n val taskId = new TaskID(jobId, TaskType.MAP, 1)\r\n val taskAttemptId0 = new TaskAttemptID(taskId, 0)\r\n val taskAttemptId1 = new TaskAttemptID(taskId, 1)\r\n\r\n val taskAttemptContext0: TaskAttemptContext = {\r\n // Set up the configuration object\r\n val hadoopConf = new Configuration();\r\n hadoopConf.set(\"mapreduce.job.id\", jobId.toString)\r\n hadoopConf.set(\"mapreduce.task.id\", taskAttemptId0.getTaskID.toString)\r\n hadoopConf.set(\"mapreduce.task.attempt.id\", taskAttemptId0.toString)\r\n hadoopConf.setBoolean(\"mapreduce.task.ismap\", true)\r\n hadoopConf.setInt(\"mapreduce.task.partition\", 0)\r\n\r\n new TaskAttemptContextImpl(hadoopConf, taskAttemptId0)\r\n }\r\n\r\n val taskAttemptContext1: TaskAttemptContext = {\r\n // Set up the configuration object\r\n val hadoopConf = new Configuration();\r\n hadoopConf.set(\"mapreduce.job.id\", jobId.toString)\r\n hadoopConf.set(\"mapreduce.task.id\", taskAttemptId1.getTaskID.toString)\r\n hadoopConf.set(\"mapreduce.task.attempt.id\", taskAttemptId1.toString)\r\n hadoopConf.setBoolean(\"mapreduce.task.ismap\", true)\r\n hadoopConf.setInt(\"mapreduce.task.partition\", 0)\r\n\r\n new TaskAttemptContextImpl(hadoopConf, taskAttemptId1)\r\n }\r\n\r\n val committer = new HadoopMapReduceCommitProtocol(jobId.toString, f.getAbsolutePath)\r\n val tf0 = committer.newTaskTempFile(taskAttemptContext0, Some(f.getAbsolutePath), \"ext\")\r\n val tf1 = committer.newTaskTempFile(taskAttemptContext1, Some(f.getAbsolutePath), \"ext\")\r\n assert(tf0 == tf1)\r\n }\r\n{code}\r\n","created":"2019-10-01T02:14:33.302+0000"},{"body":"For dynamic partition overwrite, when executing a task, a determinable path would be specified.\r\n\r\nIn the reproduce suite above, I create two task attempt context with same task id and different attempt id.\r\nAnd specify a output dir for newTaskTempFile method.","created":"2019-10-01T02:24:25.748+0000"},{"body":"Will a speculation task have the same JobID as previous task like above test? \r\n","created":"2019-10-01T02:52:31.377+0000"},{"body":"Yes, they are in a same stage, so they have a same jobId.","created":"2019-10-01T02:59:57.227+0000"},{"body":"No, I meant this:\r\n{code:java}\r\nval jobId = SparkHadoopWriterUtils.createJobID(new Date(), 1)\r\n{code}\r\n\r\nFor each writing job, in setupJob, HadoopMapReduceCommitProtocol will create a JobID. This JobID will be used getFilename. A speculation task will have a different writing job, the JobID should be different?","created":"2019-10-01T04:08:43.600+0000"},{"body":"If this is not an issue, we should close it.","created":"2019-10-04T03:50:29.301+0000"},{"body":"Sorry for the late reply, I was on my National Day holiday for the past eight days. \r\nI just made a simple jobId in the UT above.\r\nIn fact, it was created by a jobIdInstant.\r\nAnd for the tasks of a same job, they are same.\r\nSo, I think this is still an issue.\r\n !screenshot-1.png! \r\n !screenshot-2.png! \r\n\r\n","created":"2019-10-09T08:29:56.403+0000"},{"body":"cc [~cloud_fan]","created":"2019-10-09T08:43:40.413+0000"},{"body":"Hi all, I have also encountered this issue and have a fix still under test.\r\n\r\nWhen we use the dynamic overwrite, we actually use same staging directory for speculative tasks. Meanwhile, FileOutputComitter actually did nothing when commit task.","created":"2019-10-11T06:49:24.667+0000"},{"body":"This bug can be easily reproduced by inserting into a partitioned DataSource table (would not reproduced if using a Hive table) with spark.sql.sources.partitionOverwriteMode=dynamic and spark.speculation=true.","created":"2019-10-11T06:56:01.246+0000"},{"body":"[~dangdangdang]\r\n\r\nHi, I have thought a simple solution.\r\n\r\nWe just need make the file name of a task be unique.\r\nAnd the OutputCommitCoordinator would decide which task file can be committed.\r\n\r\nBut I don't have an appropriate unit test.","created":"2019-10-11T07:46:48.048+0000"},{"body":"Hi [~hzfeiwang], thanks for your contribution.(y)\r\n\r\nI just took a look at your PR, the solution is quite different from how I solved this issue, so I created a PR as an alternative solution which use FileOutputCommitter to avoid writing collision.\r\n\r\nMeanwhile, I think current tests in org.apache.spark.sql.sources.InsertSuite is enough for this case.","created":"2019-10-11T10:40:55.878+0000"},{"body":"i believe we are seeing this issue. it shows up in particular when pre-emption is turned on and we are using dynamic partition overwrite. pre-emption kills tasks, they get restarted, and then they fail again because the output directory already exists (so task throws FileAlreadyExistsException). as a result entire job fails.\r\n\r\nso i dont think this is just a speculative execution issue. this is a general issue with dynamic partition overwrite not being able to recover from task failure.","created":"2020-03-26T19:16:24.009+0000"},{"body":"I agree with [~feiwang], it looks like newTaskTempFile is not robust to speculative execution and task failures when dynamicPartitionOverwrite is enabled IMO.\r\nThis will need to be fixed - it is currently using the same path irrespective of which attempt it is.","created":"2020-04-17T23:02:05.334+0000"},{"body":"User 'turboFei' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28989","created":"2020-07-03T03:07:38.203+0000"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29000","created":"2020-07-05T20:02:46.336+0000"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29000","created":"2020-07-05T20:03:14.246+0000"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29260","created":"2020-07-27T11:45:09.924+0000"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29260","created":"2020-07-27T11:46:00.088+0000"},{"body":"Issue resolved by pull request 29000\n[https://github.com/apache/spark/pull/29000]","created":"2020-11-25T12:52:31.713+0000"}],"conversations":[{"body":"Now, for a dynamic partition overwrite operation, the filename of a task output is determinable.\r\n\r\nSo, if speculation is enabled, would a task conflict with its relative speculation task?\r\n\r\nWould the two tasks concurrent write a same file?\r\n","from":"reporter","subject":"dynamic partition overwrite with speculation enabled"},{"body":"FileCommitProtocol will assign a filename based on task id, isn't? Should not a speculation task have a different task id?","from":"developer"},{"body":"Can you share a reproducer? or this sounds more like a question instead of a bug.","from":"developer"},{"body":"You can add the code below into FileFormatWriterSuite.\r\n{code:java}\r\n test(\"SPARK-29302: for dynamic partition overwrite, a task will concurrent write a same file\" +\r\n \" with its relative speculation task\") {\r\n withTempDir { f =>\r\n val jobId = SparkHadoopWriterUtils.createJobID(new Date(), 1)\r\n val taskId = new TaskID(jobId, TaskType.MAP, 1)\r\n val taskAttemptId0 = new TaskAttemptID(taskId, 0)\r\n val taskAttemptId1 = new TaskAttemptID(taskId, 1)\r\n\r\n val taskAttemptContext0: TaskAttemptContext = {\r\n // Set up the configuration object\r\n val hadoopConf = new Configuration();\r\n hadoopConf.set(\"mapreduce.job.id\", jobId.toString)\r\n hadoopConf.set(\"mapreduce.task.id\", taskAttemptId0.getTaskID.toString)\r\n hadoopConf.set(\"mapreduce.task.attempt.id\", taskAttemptId0.toString)\r\n hadoopConf.setBoolean(\"mapreduce.task.ismap\", true)\r\n hadoopConf.setInt(\"mapreduce.task.partition\", 0)\r\n\r\n new TaskAttemptContextImpl(hadoopConf, taskAttemptId0)\r\n }\r\n\r\n val taskAttemptContext1: TaskAttemptContext = {\r\n // Set up the configuration object\r\n val hadoopConf = new Configuration();\r\n hadoopConf.set(\"mapreduce.job.id\", jobId.toString)\r\n hadoopConf.set(\"mapreduce.task.id\", taskAttemptId1.getTaskID.toString)\r\n hadoopConf.set(\"mapreduce.task.attempt.id\", taskAttemptId1.toString)\r\n hadoopConf.setBoolean(\"mapreduce.task.ismap\", true)\r\n hadoopConf.setInt(\"mapreduce.task.partition\", 0)\r\n\r\n new TaskAttemptContextImpl(hadoopConf, taskAttemptId1)\r\n }\r\n\r\n val committer = new HadoopMapReduceCommitProtocol(jobId.toString, f.getAbsolutePath)\r\n val tf0 = committer.newTaskTempFile(taskAttemptContext0, Some(f.getAbsolutePath), \"ext\")\r\n val tf1 = committer.newTaskTempFile(taskAttemptContext1, Some(f.getAbsolutePath), \"ext\")\r\n assert(tf0 == tf1)\r\n }\r\n{code}\r\n","from":"developer"},{"body":"For dynamic partition overwrite, when executing a task, a determinable path would be specified.\r\n\r\nIn the reproduce suite above, I create two task attempt context with same task id and different attempt id.\r\nAnd specify a output dir for newTaskTempFile method.","from":"developer"},{"body":"Will a speculation task have the same JobID as previous task like above test? \r\n","from":"developer"},{"body":"Yes, they are in a same stage, so they have a same jobId.","from":"developer"},{"body":"No, I meant this:\r\n{code:java}\r\nval jobId = SparkHadoopWriterUtils.createJobID(new Date(), 1)\r\n{code}\r\n\r\nFor each writing job, in setupJob, HadoopMapReduceCommitProtocol will create a JobID. This JobID will be used getFilename. A speculation task will have a different writing job, the JobID should be different?","from":"developer"},{"body":"If this is not an issue, we should close it.","from":"developer"},{"body":"Sorry for the late reply, I was on my National Day holiday for the past eight days. \r\nI just made a simple jobId in the UT above.\r\nIn fact, it was created by a jobIdInstant.\r\nAnd for the tasks of a same job, they are same.\r\nSo, I think this is still an issue.\r\n !screenshot-1.png! \r\n !screenshot-2.png! \r\n\r\n","from":"developer"},{"body":"cc [~cloud_fan]","from":"developer"},{"body":"Hi all, I have also encountered this issue and have a fix still under test.\r\n\r\nWhen we use the dynamic overwrite, we actually use same staging directory for speculative tasks. Meanwhile, FileOutputComitter actually did nothing when commit task.","from":"developer"},{"body":"This bug can be easily reproduced by inserting into a partitioned DataSource table (would not reproduced if using a Hive table) with spark.sql.sources.partitionOverwriteMode=dynamic and spark.speculation=true.","from":"developer"},{"body":"[~dangdangdang]\r\n\r\nHi, I have thought a simple solution.\r\n\r\nWe just need make the file name of a task be unique.\r\nAnd the OutputCommitCoordinator would decide which task file can be committed.\r\n\r\nBut I don't have an appropriate unit test.","from":"developer"},{"body":"Hi [~hzfeiwang], thanks for your contribution.(y)\r\n\r\nI just took a look at your PR, the solution is quite different from how I solved this issue, so I created a PR as an alternative solution which use FileOutputCommitter to avoid writing collision.\r\n\r\nMeanwhile, I think current tests in org.apache.spark.sql.sources.InsertSuite is enough for this case.","from":"developer"},{"body":"i believe we are seeing this issue. it shows up in particular when pre-emption is turned on and we are using dynamic partition overwrite. pre-emption kills tasks, they get restarted, and then they fail again because the output directory already exists (so task throws FileAlreadyExistsException). as a result entire job fails.\r\n\r\nso i dont think this is just a speculative execution issue. this is a general issue with dynamic partition overwrite not being able to recover from task failure.","from":"developer"},{"body":"I agree with [~feiwang], it looks like newTaskTempFile is not robust to speculative execution and task failures when dynamicPartitionOverwrite is enabled IMO.\r\nThis will need to be fixed - it is currently using the same path irrespective of which attempt it is.","from":"developer"},{"body":"User 'turboFei' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28989","from":"developer"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29000","from":"developer"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29000","from":"developer"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29260","from":"developer"},{"body":"User 'WinkerDu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/29260","from":"developer"},{"body":"Issue resolved by pull request 29000\n[https://github.com/apache/spark/pull/29000]","from":"developer"}],"created":"2019-09-30T11:12:45.000+0000","description":"Now, for a dynamic partition overwrite operation, the filename of a task output is determinable.\r\n\r\nSo, if speculation is enabled, would a task conflict with its relative speculation task?\r\n\r\nWould the two tasks concurrent write a same file?\r\n","issue_id":"13259642","key":"SPARK-29302","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2020-11-25T12:52:31.000+0000","role":"fixed_distractor","summary":"dynamic partition overwrite with speculation enabled"} {"case_id":"13262803","cluster":"DISTRACTOR-SPARK-29497","comments":[{"body":"This is happening to us too.  We haven't figured out how to get around it in our own code, other than passing raw functions around.\r\n\r\nWe're using Gradle instead of Maven, so I'm not sure if those variables even exist for us.","created":"2020-12-10T20:09:09.472+0000"},{"body":"I've also been able to reproduce this on Spark 3.1.1 using Scala 2.12.10 on the vanilla Spark-Shell\r\n\r\nSeems like a persistent issue","created":"2021-08-10T16:09:19.924+0000"},{"body":"some more verifying and additional info.\r\ni was able to reproduce this on both spark-3.2.1 (scala 2.12) and spark-3.2.1 (scala 2.13)) - hadoop3.2 on shell as well as via java code.\r\n\r\nlocal scala and java version:\r\n{code:java}\r\n❯ scala -version\r\nScala code runner version 2.12.14 -- Copyright 2002-2021, LAMP/EPFL and Lightbend, Inc.\r\n❯ java --version\r\nopenjdk 11.0.12 2021-07-20{code}\r\nIn my case, I am using map ( same with mapParttion also)\r\n\r\nExample:  simple Function with dummy reducer:\r\n{code:java}\r\nJavaRDD rdd =\r\ndata.toJavaRDD().map(\r\nnew Function() {\r\n@Override\r\npublic Object call(Row v1) throws Exception\r\n{ return v1 !=null; }\r\n}\r\n);\r\nObject result = rdd.reduce(LocalClass::reduceDummy);\r\npublic static Object reduceDummy(Object a, Object b)\r\n{ return null; }\r\n \r\n{code}\r\n \r\n\r\nError:\r\n{code:java}\r\nCaused by: java.lang.ClassCastException: cannot assign instance of java.lang.invoke.SerializedLambda to field org.apache.spark.rdd.MapPartitionsRDD.f of type scala.Function3 in instance of org.apache.spark.rdd.MapPartitionsRDD\r\n{code}\r\nStacktrace:\r\n{code:java}\r\nCaused by: java.lang.ClassCastException: cannot assign instance of java.lang.invoke.SerializedLambda to field org.apache.spark.rdd.MapPartitionsRDD.f of type scala.Function3 in instance of org.apache.spark.rdd.MapPartitionsRDD\r\nat java.base/java.io.ObjectStreamClass$FieldReflector.setObjFieldValues(ObjectStreamClass.java:2205)\r\nat java.base/java.io.ObjectStreamClass$FieldReflector.checkObjectFieldValueTypes(ObjectStreamClass.java:2168)\r\nat java.base/java.io.ObjectStreamClass.checkObjFieldValueTypes(ObjectStreamClass.java:1422)\r\nCaused by: java.lang.ClassCastException: cannot assign instance of java.lang.invoke.SerializedLambda to field org.apache.spark.rdd.MapPartitionsRDD.f of type scala.Function3 in instance of org.apache.spark.rdd.MapPartitionsRDD\r\nat java.base/java.io.ObjectInputStream.defaultCheckFieldValues(ObjectInputStream.java:2480)\r\nat java.base/java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2387)\r\nat java.base/java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2196)\r\nat java.base/java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1679)\r\nat java.base/java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2464)\r\nat java.base/java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2358)\r\nat java.base/java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2196)\r\nat java.base/java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1679)\r\nat java.base/java.io.ObjectInputStream.readObject(ObjectInputStream.java:493)\r\nat java.base/java.io.ObjectInputStream.readObject(ObjectInputStream.java:451)\r\nat org.apache.spark.serializer.JavaDeserializationStream.readObject(JavaSerializer.scala:76)\r\nat org.apache.spark.serializer.JavaSerializerInstance.deserialize(JavaSerializer.scala:115)\r\nat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:83)\r\nat org.apache.spark.scheduler.Task.run(Task.scala:131)\r\nat org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$3(Executor.scala:506)\r\nat org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1462)\r\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:509)\r\nat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)\r\nat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)\r\nat java.base/java.lang.Thread.run(Thread.java:829)\r\n{code}\r\n \r\n\r\n*It works fine via spark test framework, but not with spark standalone:* com.holdenkarau:spark-testing-base_2.12","created":"2022-02-25T17:59:36.854+0000"},{"body":"We have seen the same problem with Spark Connect. As far as I can tell this is either a Scala or Java bug. As soon as you invoke a lambda from a lambda this will fail. A simple non-spark reproduction:\r\n\r\n{code:scala}\r\ncase class Command1() extends Serializable {\r\n val direct: Int => Int = (i: Int) => i + 1\r\n}\r\n\r\ncase class Command2(prev: Command1) extends Serializable {\r\n val indirect: Int => Int = (i: Int) => prev.direct(i)\r\n}\r\n\r\n...\r\nval command2 = Command2(Command1())\r\nSparkSerDeUtils.deserialize[Int => Int](SparkSerDeUtils.serialize(command2.indirect))\r\n{code}\r\n\r\nI am trying to find out if this is a Scala or a Java bug. If you decode this using https://github.com/NickstaDB/SerializationDumper the dump looks reasonable.\r\n","created":"2023-07-30T19:06:58.593+0000"},{"body":"In Java, This looks like by design.\r\n\r\nI was able to make it serializable and resolve the problem.\r\n\r\nFor ex: If you are using `UnaryOperator` you can tell to be serializable via following syntax.\r\n\r\n \r\n{code:java}\r\nUnaryOperator rowTransform = (UnaryOperator & Serializable) row -> { \r\n ....\r\n}{code}\r\nI found this at: [https://stackoverflow.com/questions/22807912/how-to-serialize-a-lambda]\r\n\r\nThere is also good explanation on serialization logic here: [https://stackoverflow.com/questions/28186607/java-lang-classcastexception-using-lambda-expressions-in-spark-job-on-remote-ser]\r\n\r\nHope this helps.\r\n\r\n \r\n\r\n ","created":"2023-07-30T19:17:35.190+0000"},{"body":"I have also tried the following code (inspired by the what was reported in this ticket), and that failed in a similar fashion:\r\n{code:java}\r\ncase class MultipleLambdas() extends Serializable {\r\n val direct: Int => Int = (i: Int) => i + 1\r\n val indirect: Int => Int = (i: Int) => direct(i)\r\n}\r\n...\r\nval ml = MultipleLambdas()\r\nSparkSerDeUtils.deserialize[Int => Int](SparkSerDeUtils.serialize(ml.indirect))\r\n\r\n{code}","created":"2023-07-30T19:21:22.761+0000"},{"body":"[~jaysen] Thanks for the reaction.\r\n\r\nThe scala functions *are* serializable, otherwise UDFs in general wouldn't work. The second SO article you linked definitely helped me to understand why we got weird exceptions when we were deserializing UDFs without having all the required CP entries. I ran a debugger on the serverside, and this does not seem to be the case here (you should have CNFE exceptions in the HandleTable), the 'indirect' lambda does not seem to be properly initialized. It is probably an ordering issue.","created":"2023-07-30T19:35:04.213+0000"},{"body":"This all seems to boil down to the fact that you cannot deserialize an object graph if this contains a self reference and that self-reference is a proxy. In this case readResolve is never called (because its inputs are not fully deserialized), and you get a ClassCastException because you are trying to assign a proxy to the real object. In this particular case SerializedLambda is a serialization proxy. See [this|https://docs.oracle.com/en/java/javase/17/docs/specs/serialization/input.html#the-readresolve-method] for more information.\r\n\r\nAll of the examples mentioned here contain such a self reference, in the form of lamdba -> parent -> lambda. An even simpler example would be the following:\r\n\r\n{code:scala}\r\ncase class SelfRef(start: Int) extends Serializable {\r\n val method: Int => Int = (i: Int) => i + start\r\n}\r\n...\r\nSparkSerDeUtils.deserialize[Int => Int](SparkSerDeUtils.serialize(SelfRef(43).method)) // KABOOM\r\n{code}\r\n\r\nAs far as mitigations go there is not much we can do besides improving error handling (hard because the observed exception also masks dependency problems), and writing down which patterns to avoid.","created":"2023-07-31T02:31:55.264+0000"},{"body":"I have added a check for this to Spark Connect. If someone is brave enough they can do the same thing for other UDFs.","created":"2023-08-01T18:55:18.535+0000"},{"body":"Resolving as Fixed: this change was merged to master. Implementing commit: f54b40202178 [SPARK-29497][CONNECT] Throw error when UDF is not deserializable. (Triaged as a stranded ticket — the fix landed but the JIRA was left unresolved. Verified present on current master and not reverted.)","created":"2026-07-09T01:55:50.594+0000"}],"conversations":[{"body":"Note this is for scala 2.12:\r\n\r\nThere seems to be an issue in spark with serializing a udf that is created from a function assigned to a class member that references another function assigned to a class member. This is similar to https://issues.apache.org/jira/browse/SPARK-25047 but it looks like the resolution has an issue with this case. After trimming it down to the base issue I came up with the following to reproduce:\r\n\r\n \r\n\r\n \r\n{code:java}\r\nobject TestLambdaShell extends Serializable {\r\n val hello: String => String = s => s\"hello $s!\" \r\n val lambdaTest: String => String = hello( _ ) \r\n def functionTest: String => String = hello( _ )\r\n}\r\n\r\nval hello = udf( TestLambdaShell.hello )\r\nval functionTest = udf( TestLambdaShell.functionTest )\r\nval lambdaTest = udf( TestLambdaShell.lambdaTest )\r\n\r\nsc.parallelize(Seq(\"world\"),1).toDF(\"test\").select(hello($\"test\")).show(1)\r\nsc.parallelize(Seq(\"world\"),1).toDF(\"test\").select(functionTest($\"test\")).show(1)\r\nsc.parallelize(Seq(\"world\"),1).toDF(\"test\").select(lambdaTest($\"test\")).show(1)\r\n{code}\r\n \r\n\r\nAll of which works except the last line which results in an exception on the executors:\r\n\r\n \r\n{code:java}\r\nCaused by: java.lang.ClassCastException: cannot assign instance of java.lang.invoke.SerializedLambda to field $$$82b5b23cea489b2712a1db46c77e458$$$$w$TestLambdaShell$.lambdaTest of type scala.Function1 in instance of $$$82b5b23cea489b2712a1db46c77e458$$$$w$TestLambdaShell$\r\n at java.io.ObjectStreamClass$FieldReflector.setObjFieldValues(ObjectStreamClass.java:2133)\r\n at java.io.ObjectStreamClass.setObjFieldValues(ObjectStreamClass.java:1305)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2251)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readArray(ObjectInputStream.java:1933)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1529)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readArray(ObjectInputStream.java:1933)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1529)\r\n at java.io.ObjectInputStream.readArray(ObjectInputStream.java:1933)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1529)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readArray(ObjectInputStream.java:1933)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1529)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readObject(ObjectInputStream.java:422)\r\n at scala.collection.immutable.List$SerializationProxy.readObject(List.scala:488)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at java.io.ObjectStreamClass.invokeReadObject(ObjectStreamClass.java:1058)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2136)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readObject(ObjectInputStream.java:422)\r\n at scala.collection.immutable.List$SerializationProxy.readObject(List.scala:488)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at java.io.ObjectStreamClass.invokeReadObject(ObjectStreamClass.java:1058)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2136)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readObject(ObjectInputStream.java:422)\r\n at org.apache.spark.serializer.JavaDeserializationStream.readObject(JavaSerializer.scala:75)\r\n at org.apache.spark.serializer.JavaSerializerInstance.deserialize(JavaSerializer.scala:114)\r\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:83)\r\n at org.apache.spark.scheduler.Task.run(Task.scala:121)\r\n at org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$3(Executor.scala:411)\r\n at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1360)\r\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:414)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n at java.lang.Thread.run(Thread.java:748)\r\n{code}\r\n \r\n\r\nIn spark 2.2.x I used a class that had something like this that worked fine, now that we've upgraded to 2.12 we ran into a few serialization issues in places, most of which were solved by extending serializable but this case was not fixed by that.\r\n\r\n \r\n\r\nAlso this happens regardless of whether it's done in the shell or in a jar.\r\n\r\n \r\n\r\n \r\n\r\n \r\n\r\nSo after much more debugging, this turns out to be some weird mix of scala 2.12.0 and scala 2.12.8. Spark is compiled on 2.12.8 and so is our own code but I noticed that the maven compiled class did not match the compiled class using 2.12.8 scalac directly. After a lot of digging we realized that scala-compiler actually indirectly depends on scala library 2.12.0 and only when the spark dependency is added does it start using it for some reason. Without the spark dependency and just direct scala 2.12.8 dependencies, the code builds fine and compiles correctly as 2.12.8. \r\n\r\n \r\n\r\nWe were able to fix this using:\r\n\r\n \r\n{code:java}\r\ntrue\r\n2.12\r\n2.12.8\r\n{code}\r\n \r\n\r\nAnd this resolves our issue for our own jars that we create and link to spark. However, my original test case still seems to reproduce in the spark shell and for us also in apache zeppelin so it seems almost like somehow they are also compiling it on 2.12.0 but I'm not quite sure how. In the spark pom.xml it seems to have the fail on multiple versions and compiles fine so I'm not quite sure how this is happening but at least its more isolated now. I'm also wondering if anything else could be affected by this.","from":"reporter","subject":"Cannot assign instance of java.lang.invoke.SerializedLambda to field"},{"body":"This is happening to us too.  We haven't figured out how to get around it in our own code, other than passing raw functions around.\r\n\r\nWe're using Gradle instead of Maven, so I'm not sure if those variables even exist for us.","from":"developer"},{"body":"I've also been able to reproduce this on Spark 3.1.1 using Scala 2.12.10 on the vanilla Spark-Shell\r\n\r\nSeems like a persistent issue","from":"developer"},{"body":"some more verifying and additional info.\r\ni was able to reproduce this on both spark-3.2.1 (scala 2.12) and spark-3.2.1 (scala 2.13)) - hadoop3.2 on shell as well as via java code.\r\n\r\nlocal scala and java version:\r\n{code:java}\r\n❯ scala -version\r\nScala code runner version 2.12.14 -- Copyright 2002-2021, LAMP/EPFL and Lightbend, Inc.\r\n❯ java --version\r\nopenjdk 11.0.12 2021-07-20{code}\r\nIn my case, I am using map ( same with mapParttion also)\r\n\r\nExample:  simple Function with dummy reducer:\r\n{code:java}\r\nJavaRDD rdd =\r\ndata.toJavaRDD().map(\r\nnew Function() {\r\n@Override\r\npublic Object call(Row v1) throws Exception\r\n{ return v1 !=null; }\r\n}\r\n);\r\nObject result = rdd.reduce(LocalClass::reduceDummy);\r\npublic static Object reduceDummy(Object a, Object b)\r\n{ return null; }\r\n \r\n{code}\r\n \r\n\r\nError:\r\n{code:java}\r\nCaused by: java.lang.ClassCastException: cannot assign instance of java.lang.invoke.SerializedLambda to field org.apache.spark.rdd.MapPartitionsRDD.f of type scala.Function3 in instance of org.apache.spark.rdd.MapPartitionsRDD\r\n{code}\r\nStacktrace:\r\n{code:java}\r\nCaused by: java.lang.ClassCastException: cannot assign instance of java.lang.invoke.SerializedLambda to field org.apache.spark.rdd.MapPartitionsRDD.f of type scala.Function3 in instance of org.apache.spark.rdd.MapPartitionsRDD\r\nat java.base/java.io.ObjectStreamClass$FieldReflector.setObjFieldValues(ObjectStreamClass.java:2205)\r\nat java.base/java.io.ObjectStreamClass$FieldReflector.checkObjectFieldValueTypes(ObjectStreamClass.java:2168)\r\nat java.base/java.io.ObjectStreamClass.checkObjFieldValueTypes(ObjectStreamClass.java:1422)\r\nCaused by: java.lang.ClassCastException: cannot assign instance of java.lang.invoke.SerializedLambda to field org.apache.spark.rdd.MapPartitionsRDD.f of type scala.Function3 in instance of org.apache.spark.rdd.MapPartitionsRDD\r\nat java.base/java.io.ObjectInputStream.defaultCheckFieldValues(ObjectInputStream.java:2480)\r\nat java.base/java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2387)\r\nat java.base/java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2196)\r\nat java.base/java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1679)\r\nat java.base/java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2464)\r\nat java.base/java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2358)\r\nat java.base/java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2196)\r\nat java.base/java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1679)\r\nat java.base/java.io.ObjectInputStream.readObject(ObjectInputStream.java:493)\r\nat java.base/java.io.ObjectInputStream.readObject(ObjectInputStream.java:451)\r\nat org.apache.spark.serializer.JavaDeserializationStream.readObject(JavaSerializer.scala:76)\r\nat org.apache.spark.serializer.JavaSerializerInstance.deserialize(JavaSerializer.scala:115)\r\nat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:83)\r\nat org.apache.spark.scheduler.Task.run(Task.scala:131)\r\nat org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$3(Executor.scala:506)\r\nat org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1462)\r\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:509)\r\nat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1128)\r\nat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:628)\r\nat java.base/java.lang.Thread.run(Thread.java:829)\r\n{code}\r\n \r\n\r\n*It works fine via spark test framework, but not with spark standalone:* com.holdenkarau:spark-testing-base_2.12","from":"developer"},{"body":"We have seen the same problem with Spark Connect. As far as I can tell this is either a Scala or Java bug. As soon as you invoke a lambda from a lambda this will fail. A simple non-spark reproduction:\r\n\r\n{code:scala}\r\ncase class Command1() extends Serializable {\r\n val direct: Int => Int = (i: Int) => i + 1\r\n}\r\n\r\ncase class Command2(prev: Command1) extends Serializable {\r\n val indirect: Int => Int = (i: Int) => prev.direct(i)\r\n}\r\n\r\n...\r\nval command2 = Command2(Command1())\r\nSparkSerDeUtils.deserialize[Int => Int](SparkSerDeUtils.serialize(command2.indirect))\r\n{code}\r\n\r\nI am trying to find out if this is a Scala or a Java bug. If you decode this using https://github.com/NickstaDB/SerializationDumper the dump looks reasonable.\r\n","from":"developer"},{"body":"In Java, This looks like by design.\r\n\r\nI was able to make it serializable and resolve the problem.\r\n\r\nFor ex: If you are using `UnaryOperator` you can tell to be serializable via following syntax.\r\n\r\n \r\n{code:java}\r\nUnaryOperator rowTransform = (UnaryOperator & Serializable) row -> { \r\n ....\r\n}{code}\r\nI found this at: [https://stackoverflow.com/questions/22807912/how-to-serialize-a-lambda]\r\n\r\nThere is also good explanation on serialization logic here: [https://stackoverflow.com/questions/28186607/java-lang-classcastexception-using-lambda-expressions-in-spark-job-on-remote-ser]\r\n\r\nHope this helps.\r\n\r\n \r\n\r\n ","from":"developer"},{"body":"I have also tried the following code (inspired by the what was reported in this ticket), and that failed in a similar fashion:\r\n{code:java}\r\ncase class MultipleLambdas() extends Serializable {\r\n val direct: Int => Int = (i: Int) => i + 1\r\n val indirect: Int => Int = (i: Int) => direct(i)\r\n}\r\n...\r\nval ml = MultipleLambdas()\r\nSparkSerDeUtils.deserialize[Int => Int](SparkSerDeUtils.serialize(ml.indirect))\r\n\r\n{code}","from":"developer"},{"body":"[~jaysen] Thanks for the reaction.\r\n\r\nThe scala functions *are* serializable, otherwise UDFs in general wouldn't work. The second SO article you linked definitely helped me to understand why we got weird exceptions when we were deserializing UDFs without having all the required CP entries. I ran a debugger on the serverside, and this does not seem to be the case here (you should have CNFE exceptions in the HandleTable), the 'indirect' lambda does not seem to be properly initialized. It is probably an ordering issue.","from":"developer"},{"body":"This all seems to boil down to the fact that you cannot deserialize an object graph if this contains a self reference and that self-reference is a proxy. In this case readResolve is never called (because its inputs are not fully deserialized), and you get a ClassCastException because you are trying to assign a proxy to the real object. In this particular case SerializedLambda is a serialization proxy. See [this|https://docs.oracle.com/en/java/javase/17/docs/specs/serialization/input.html#the-readresolve-method] for more information.\r\n\r\nAll of the examples mentioned here contain such a self reference, in the form of lamdba -> parent -> lambda. An even simpler example would be the following:\r\n\r\n{code:scala}\r\ncase class SelfRef(start: Int) extends Serializable {\r\n val method: Int => Int = (i: Int) => i + start\r\n}\r\n...\r\nSparkSerDeUtils.deserialize[Int => Int](SparkSerDeUtils.serialize(SelfRef(43).method)) // KABOOM\r\n{code}\r\n\r\nAs far as mitigations go there is not much we can do besides improving error handling (hard because the observed exception also masks dependency problems), and writing down which patterns to avoid.","from":"developer"},{"body":"I have added a check for this to Spark Connect. If someone is brave enough they can do the same thing for other UDFs.","from":"developer"},{"body":"Resolving as Fixed: this change was merged to master. Implementing commit: f54b40202178 [SPARK-29497][CONNECT] Throw error when UDF is not deserializable. (Triaged as a stranded ticket — the fix landed but the JIRA was left unresolved. Verified present on current master and not reverted.)","from":"developer"}],"created":"2019-10-17T09:08:20.000+0000","description":"Note this is for scala 2.12:\r\n\r\nThere seems to be an issue in spark with serializing a udf that is created from a function assigned to a class member that references another function assigned to a class member. This is similar to https://issues.apache.org/jira/browse/SPARK-25047 but it looks like the resolution has an issue with this case. After trimming it down to the base issue I came up with the following to reproduce:\r\n\r\n \r\n\r\n \r\n{code:java}\r\nobject TestLambdaShell extends Serializable {\r\n val hello: String => String = s => s\"hello $s!\" \r\n val lambdaTest: String => String = hello( _ ) \r\n def functionTest: String => String = hello( _ )\r\n}\r\n\r\nval hello = udf( TestLambdaShell.hello )\r\nval functionTest = udf( TestLambdaShell.functionTest )\r\nval lambdaTest = udf( TestLambdaShell.lambdaTest )\r\n\r\nsc.parallelize(Seq(\"world\"),1).toDF(\"test\").select(hello($\"test\")).show(1)\r\nsc.parallelize(Seq(\"world\"),1).toDF(\"test\").select(functionTest($\"test\")).show(1)\r\nsc.parallelize(Seq(\"world\"),1).toDF(\"test\").select(lambdaTest($\"test\")).show(1)\r\n{code}\r\n \r\n\r\nAll of which works except the last line which results in an exception on the executors:\r\n\r\n \r\n{code:java}\r\nCaused by: java.lang.ClassCastException: cannot assign instance of java.lang.invoke.SerializedLambda to field $$$82b5b23cea489b2712a1db46c77e458$$$$w$TestLambdaShell$.lambdaTest of type scala.Function1 in instance of $$$82b5b23cea489b2712a1db46c77e458$$$$w$TestLambdaShell$\r\n at java.io.ObjectStreamClass$FieldReflector.setObjFieldValues(ObjectStreamClass.java:2133)\r\n at java.io.ObjectStreamClass.setObjFieldValues(ObjectStreamClass.java:1305)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2251)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readArray(ObjectInputStream.java:1933)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1529)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readArray(ObjectInputStream.java:1933)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1529)\r\n at java.io.ObjectInputStream.readArray(ObjectInputStream.java:1933)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1529)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readArray(ObjectInputStream.java:1933)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1529)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readObject(ObjectInputStream.java:422)\r\n at scala.collection.immutable.List$SerializationProxy.readObject(List.scala:488)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at java.io.ObjectStreamClass.invokeReadObject(ObjectStreamClass.java:1058)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2136)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readObject(ObjectInputStream.java:422)\r\n at scala.collection.immutable.List$SerializationProxy.readObject(List.scala:488)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at java.io.ObjectStreamClass.invokeReadObject(ObjectStreamClass.java:1058)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2136)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:2245)\r\n at java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:2169)\r\n at java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:2027)\r\n at java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1535)\r\n at java.io.ObjectInputStream.readObject(ObjectInputStream.java:422)\r\n at org.apache.spark.serializer.JavaDeserializationStream.readObject(JavaSerializer.scala:75)\r\n at org.apache.spark.serializer.JavaSerializerInstance.deserialize(JavaSerializer.scala:114)\r\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:83)\r\n at org.apache.spark.scheduler.Task.run(Task.scala:121)\r\n at org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$3(Executor.scala:411)\r\n at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1360)\r\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:414)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n at java.lang.Thread.run(Thread.java:748)\r\n{code}\r\n \r\n\r\nIn spark 2.2.x I used a class that had something like this that worked fine, now that we've upgraded to 2.12 we ran into a few serialization issues in places, most of which were solved by extending serializable but this case was not fixed by that.\r\n\r\n \r\n\r\nAlso this happens regardless of whether it's done in the shell or in a jar.\r\n\r\n \r\n\r\n \r\n\r\n \r\n\r\nSo after much more debugging, this turns out to be some weird mix of scala 2.12.0 and scala 2.12.8. Spark is compiled on 2.12.8 and so is our own code but I noticed that the maven compiled class did not match the compiled class using 2.12.8 scalac directly. After a lot of digging we realized that scala-compiler actually indirectly depends on scala library 2.12.0 and only when the spark dependency is added does it start using it for some reason. Without the spark dependency and just direct scala 2.12.8 dependencies, the code builds fine and compiles correctly as 2.12.8. \r\n\r\n \r\n\r\nWe were able to fix this using:\r\n\r\n \r\n{code:java}\r\ntrue\r\n2.12\r\n2.12.8\r\n{code}\r\n \r\n\r\nAnd this resolves our issue for our own jars that we create and link to spark. However, my original test case still seems to reproduce in the spark shell and for us also in apache zeppelin so it seems almost like somehow they are also compiling it on 2.12.0 but I'm not quite sure how. In the spark pom.xml it seems to have the fail on multiple versions and compiles fine so I'm not quite sure how this is happening but at least its more isolated now. I'm also wondering if anything else could be affected by this.","issue_id":"13262803","key":"SPARK-29497","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2026-07-09T01:55:50.000+0000","role":"fixed_distractor","summary":"Cannot assign instance of java.lang.invoke.SerializedLambda to field"} {"case_id":"13265491","cluster":"DISTRACTOR-SPARK-29683","comments":[{"body":"We're experiencing this as well during HA YARN failover.","created":"2019-11-05T03:12:09.612+0000"},{"body":"User 'cnZach' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28606","created":"2020-06-01T14:47:03.724+0000"},{"body":"I agree with Genmao, these logics that added by SPARK-16630 is too strong condition to make application fail.\r\n\r\n \r\n{code:java}\r\ndef isAllNodeBlacklisted: Boolean = currentBlacklistedYarnNodes.size >= numClusterNodes\r\nval allBlacklistedNodes = excludeNodes ++ schedulerBlacklist ++ allocatorBlacklist.keySet\r\n{code}\r\n \r\n\r\nI think the above logic would work only for partial failure or intermittent issue that numClusterNodes isn't changed, in case of a scheduler or allocator failure become permanent failure that ResourceManager has to remove from its pool which numClusterNodes is changed, the above logic could fail.\r\n\r\ne.g) let's say a cluster has 2 NodeManagers(numClusterNodes = 2), and one NodeManager(N1) has the some issues that cause scheduling failures which ends up increasing schedulerBlacklist.size to 1, and later N1 can't recover from ResourceManager's perspective due to a hardware failure or decommissioned by operator or any other ways, in this case numClusterNodes becomes 1 which makes isAllNodeBlacklisted true, even if there is still 1 NodeManager available and \"spark.yarn.blacklist.executor.launch.blacklisting.enabled\" set to false\r\n\r\nParticularly in cloud environment, resizing of cluster happens all the times, for long-running spark application with many resize operations of cluster, schedulerBlacklist.size could keep increasing while numClusterNodes keep fluctuated, in addition even if currentBlacklistedYarnNodes.size >= numClusterNodes is true case, there could be new nodes would be added quickly. \r\n\r\nI found [e70df2cea46f71461d8d401a420e946f999862c1|https://github.com/apache/spark/commit/e70df2cea46f71461d8d401a420e946f999862c1] was added to handle the case of numClusterNode = 0.\r\n\r\nHowever for other cases as mentioned in this JIRA, I think just removing the following part from [ApplicationMaster|https://github.com/apache/spark/blob/branch-2.4/resource-managers/yarn/src/main/scala/org/apache/spark/deploy/yarn/ApplicationMaster.scala#L535-L538] would more make sense, because isAllNodeBlackListed doesn't necessarily mean running application needs to fail.\r\n\r\n \r\n{code:java}\r\n} else if (allocator.isAllNodeBlacklisted) {\r\n finish(FinalApplicationStatus.FAILED,\r\n ApplicationMaster.EXIT_MAX_EXECUTOR_FAILURES,\r\n \"Due to executor failures all available nodes are blacklisted\")\r\n{code}\r\n \r\n\r\nOr at least the above condition should apply as optional with \"spark.yarn.blacklist.executor.launch.blacklisting.enabled\" or some new configuration because SPARK-16630 added as optional but the above logic impact regardless of any configuration.\r\n\r\nI wonder other's opinion about this.\r\n\r\n ","created":"2020-10-30T22:09:30.402+0000"},{"body":"Facing the same issue with 3.0.1 with streaming:\r\n\r\n{noformat}\r\nApplication Report :\r\nApplication-Id : application_123934893489289_0001\r\nApplication-Name : yyyyy\r\nApplication-Type : SPARK\r\nUser : livy\r\nQueue : default\r\nApplication Priority : 0\r\nStart-Time : 1622842590806\r\nFinish-Time : 1623111223883\r\nProgress : 100%\r\nState : FINISHED\r\nFinal-State : FAILED\r\nTracking-URL : ip-xx.ec2.internal:18080/history/application_1123934893489289_0001/3\r\nRPC Port : 36535\r\nAM Host : ip-10-160-98-55.ec2.internal\r\nAggregate Resource Allocation : 41388024201 MB-seconds, 2854390 vcore-seconds\r\nAggregate Resource Preempted : 0 MB-seconds, 0 vcore-seconds\r\nLog Aggregation Status : TIME_OUT\r\nDiagnostics : Due to executor failures all available nodes are blacklisted\r\nUnmanaged Application : false\r\nApplication Node Label Expression : \r\nAM container Node Label Expression : \r\nTimeoutType : LIFETIME\tExpiryTime : UNLIMITED\tRemainingTime : -1seconds\r\n{noformat}\r\n\r\nIs there a workaround or any updates on this?\r\n","created":"2021-06-14T16:22:57.675+0000"},{"body":"User 'sungpeo' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/35089","created":"2022-01-03T10:48:19.114+0000"},{"body":"User 'sungpeo' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/35089","created":"2022-01-03T10:49:42.296+0000"},{"body":"[~srowen] I think we can close this as this commit solved the issue:\r\nhttps://github.com/apache/spark/commit/e70df2cea46f71461d8d401a420e946f999862c1\r\n\r\nWhat do you think?","created":"2022-11-01T00:23:35.221+0000"}],"conversations":[{"body":"My streaming job will fail *due to executor failures all available nodes are blacklisted*. This exception is thrown only when all node is blacklisted:\r\n{code:java}\r\ndef isAllNodeBlacklisted: Boolean = currentBlacklistedYarnNodes.size >= numClusterNodes\r\n\r\nval allBlacklistedNodes = excludeNodes ++ schedulerBlacklist ++ allocatorBlacklist.keySet\r\n{code}\r\nAfter diving into the code, I found some critical conditions not be handled properly:\r\n - unchecked `excludeNodes`: it comes from user config. If not set properly, it may lead to \"currentBlacklistedYarnNodes.size >= numClusterNodes\". For example, we may set some nodes not in Yarn cluster.\r\n{code:java}\r\nexcludeNodes = (invalid1, invalid2, invalid3)\r\nclusterNodes = (valid1, valid2)\r\n{code}\r\n\r\n - `numClusterNodes` may equals 0: When HA Yarn failover, it will take some time for all NodeManagers to register ResourceManager again. In this case, `numClusterNode` may equals 0 or some other number, and Spark driver failed.\r\n - too strong condition check: Spark driver will fail as long as \"currentBlacklistedYarnNodes.size >= numClusterNodes\". This condition should not indicate a unrecovered fatal. For example, there are some NodeManagers restarting. So we can give some waiting time before job failed.","from":"reporter","subject":"Job failed due to executor failures all available nodes are blacklisted"},{"body":"We're experiencing this as well during HA YARN failover.","from":"developer"},{"body":"User 'cnZach' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/28606","from":"developer"},{"body":"I agree with Genmao, these logics that added by SPARK-16630 is too strong condition to make application fail.\r\n\r\n \r\n{code:java}\r\ndef isAllNodeBlacklisted: Boolean = currentBlacklistedYarnNodes.size >= numClusterNodes\r\nval allBlacklistedNodes = excludeNodes ++ schedulerBlacklist ++ allocatorBlacklist.keySet\r\n{code}\r\n \r\n\r\nI think the above logic would work only for partial failure or intermittent issue that numClusterNodes isn't changed, in case of a scheduler or allocator failure become permanent failure that ResourceManager has to remove from its pool which numClusterNodes is changed, the above logic could fail.\r\n\r\ne.g) let's say a cluster has 2 NodeManagers(numClusterNodes = 2), and one NodeManager(N1) has the some issues that cause scheduling failures which ends up increasing schedulerBlacklist.size to 1, and later N1 can't recover from ResourceManager's perspective due to a hardware failure or decommissioned by operator or any other ways, in this case numClusterNodes becomes 1 which makes isAllNodeBlacklisted true, even if there is still 1 NodeManager available and \"spark.yarn.blacklist.executor.launch.blacklisting.enabled\" set to false\r\n\r\nParticularly in cloud environment, resizing of cluster happens all the times, for long-running spark application with many resize operations of cluster, schedulerBlacklist.size could keep increasing while numClusterNodes keep fluctuated, in addition even if currentBlacklistedYarnNodes.size >= numClusterNodes is true case, there could be new nodes would be added quickly. \r\n\r\nI found [e70df2cea46f71461d8d401a420e946f999862c1|https://github.com/apache/spark/commit/e70df2cea46f71461d8d401a420e946f999862c1] was added to handle the case of numClusterNode = 0.\r\n\r\nHowever for other cases as mentioned in this JIRA, I think just removing the following part from [ApplicationMaster|https://github.com/apache/spark/blob/branch-2.4/resource-managers/yarn/src/main/scala/org/apache/spark/deploy/yarn/ApplicationMaster.scala#L535-L538] would more make sense, because isAllNodeBlackListed doesn't necessarily mean running application needs to fail.\r\n\r\n \r\n{code:java}\r\n} else if (allocator.isAllNodeBlacklisted) {\r\n finish(FinalApplicationStatus.FAILED,\r\n ApplicationMaster.EXIT_MAX_EXECUTOR_FAILURES,\r\n \"Due to executor failures all available nodes are blacklisted\")\r\n{code}\r\n \r\n\r\nOr at least the above condition should apply as optional with \"spark.yarn.blacklist.executor.launch.blacklisting.enabled\" or some new configuration because SPARK-16630 added as optional but the above logic impact regardless of any configuration.\r\n\r\nI wonder other's opinion about this.\r\n\r\n ","from":"developer"},{"body":"Facing the same issue with 3.0.1 with streaming:\r\n\r\n{noformat}\r\nApplication Report :\r\nApplication-Id : application_123934893489289_0001\r\nApplication-Name : yyyyy\r\nApplication-Type : SPARK\r\nUser : livy\r\nQueue : default\r\nApplication Priority : 0\r\nStart-Time : 1622842590806\r\nFinish-Time : 1623111223883\r\nProgress : 100%\r\nState : FINISHED\r\nFinal-State : FAILED\r\nTracking-URL : ip-xx.ec2.internal:18080/history/application_1123934893489289_0001/3\r\nRPC Port : 36535\r\nAM Host : ip-10-160-98-55.ec2.internal\r\nAggregate Resource Allocation : 41388024201 MB-seconds, 2854390 vcore-seconds\r\nAggregate Resource Preempted : 0 MB-seconds, 0 vcore-seconds\r\nLog Aggregation Status : TIME_OUT\r\nDiagnostics : Due to executor failures all available nodes are blacklisted\r\nUnmanaged Application : false\r\nApplication Node Label Expression : \r\nAM container Node Label Expression : \r\nTimeoutType : LIFETIME\tExpiryTime : UNLIMITED\tRemainingTime : -1seconds\r\n{noformat}\r\n\r\nIs there a workaround or any updates on this?\r\n","from":"developer"},{"body":"User 'sungpeo' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/35089","from":"developer"},{"body":"User 'sungpeo' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/35089","from":"developer"},{"body":"[~srowen] I think we can close this as this commit solved the issue:\r\nhttps://github.com/apache/spark/commit/e70df2cea46f71461d8d401a420e946f999862c1\r\n\r\nWhat do you think?","from":"developer"}],"created":"2019-10-31T09:35:34.000+0000","description":"My streaming job will fail *due to executor failures all available nodes are blacklisted*. This exception is thrown only when all node is blacklisted:\r\n{code:java}\r\ndef isAllNodeBlacklisted: Boolean = currentBlacklistedYarnNodes.size >= numClusterNodes\r\n\r\nval allBlacklistedNodes = excludeNodes ++ schedulerBlacklist ++ allocatorBlacklist.keySet\r\n{code}\r\nAfter diving into the code, I found some critical conditions not be handled properly:\r\n - unchecked `excludeNodes`: it comes from user config. If not set properly, it may lead to \"currentBlacklistedYarnNodes.size >= numClusterNodes\". For example, we may set some nodes not in Yarn cluster.\r\n{code:java}\r\nexcludeNodes = (invalid1, invalid2, invalid3)\r\nclusterNodes = (valid1, valid2)\r\n{code}\r\n\r\n - `numClusterNodes` may equals 0: When HA Yarn failover, it will take some time for all NodeManagers to register ResourceManager again. In this case, `numClusterNode` may equals 0 or some other number, and Spark driver failed.\r\n - too strong condition check: Spark driver will fail as long as \"currentBlacklistedYarnNodes.size >= numClusterNodes\". This condition should not indicate a unrecovered fatal. For example, there are some NodeManagers restarting. So we can give some waiting time before job failed.","issue_id":"13265491","key":"SPARK-29683","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2022-11-01T03:08:45.000+0000","role":"fixed_distractor","summary":"Job failed due to executor failures all available nodes are blacklisted"} {"case_id":"13266553","cluster":"DISTRACTOR-SPARK-29773","comments":[{"body":"2.3.x is EOL releases. Can you try out in higher versions?","created":"2019-11-06T12:09:23.610+0000"},{"body":"We will try on 2.4.4","created":"2019-11-06T12:30:42.523+0000"},{"body":"ping [~aermakov], have you tried this out?","created":"2019-11-11T17:13:41.693+0000"},{"body":"This issue has been resolved for Spark 2.4.4","created":"2020-06-22T15:20:43.531+0000"}],"conversations":[{"body":"Unable to process empty ORC files in Hive Table using Spark SQL. It seems that a problem with class org.apache.hadoop.hive.ql.io.orc.OrcInputFormat.getSplits()\r\n\r\nStack trace:\r\n{code:java}\r\n19/10/30 22:29:54 ERROR SparkSQLDriver: Failed in [select distinct _tech_load_dt from dl_raw.tpaccsieee_ut_data_address]\r\norg.apache.spark.sql.catalyst.errors.package$TreeNodeException: execute, tree:\r\nExchange hashpartitioning(_tech_load_dt#1374, 200)\r\n+- *(1) HashAggregate(keys=[_tech_load_dt#1374], functions=[], output=[_tech_load_dt#1374])\r\n +- HiveTableScan [_tech_load_dt#1374], HiveTableRelation `dl_raw`.`tpaccsieee_ut_data_address`, org.apache.hadoop.hive.ql.io.orc.OrcSerde, [address#1307, address_9zp#1308, address_adm#1309, address_md#1310, adress_doc#1311, building#1312, change_date_addr_el#1313, change_date_okato#1314, change_date_окато#1315, city#1316, city_id#1317, cnv_cont_id#1318, code_intercity#1319, code_kladr#1320, code_plan1#1321, date_act#1322, date_change#1323, date_prz_incorrect_code_kladr#1324, date_record#1325, district#1326, district_id#1327, etaj#1328, e_plan#1329, fax#1330, ... 44 more fields] at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:56)\r\n at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec.doExecute(ShuffleExchangeExec.scala:119)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:131)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:155)\r\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\r\n at org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:152)\r\n at org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.InputAdapter.inputRDDs(WholeStageCodegenExec.scala:371)\r\n at org.apache.spark.sql.execution.aggregate.HashAggregateExec.inputRDDs(HashAggregateExec.scala:150)\r\n at org.apache.spark.sql.execution.WholeStageCodegenExec.doExecute(WholeStageCodegenExec.scala:605)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:131)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:155)\r\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\r\n at org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:152)\r\n at org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:247)\r\n at org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:294)\r\n at org.apache.spark.sql.execution.SparkPlan.executeCollectPublic(SparkPlan.scala:324)\r\n at org.apache.spark.sql.execution.QueryExecution.hiveResultString(QueryExecution.scala:122)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLDriver$$anonfun$run$1.apply(SparkSQLDriver.scala:64)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLDriver$$anonfun$run$1.apply(SparkSQLDriver.scala:64)\r\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:77)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:63)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:364)\r\n at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:376)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$.main(SparkSQLCLIDriver.scala:272)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.main(SparkSQLCLIDriver.scala)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at org.apache.spark.deploy.JavaMainApplication.start(SparkApplication.scala:52)\r\n at org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:894)\r\n at org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:198)\r\n at org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:228)\r\n at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:137)\r\n at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\r\nCaused by: java.lang.RuntimeException: serious problem\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat.generateSplitsInfo(OrcInputFormat.java:1021)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat.getSplits(OrcInputFormat.java:1048)\r\n at org.apache.spark.rdd.HadoopRDD.getPartitions(HadoopRDD.scala:200)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.rdd.MapPartitionsRDD.getPartitions(MapPartitionsRDD.scala:35)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.rdd.MapPartitionsRDD.getPartitions(MapPartitionsRDD.scala:35)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.rdd.MapPartitionsRDD.getPartitions(MapPartitionsRDD.scala:35)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.rdd.MapPartitionsRDD.getPartitions(MapPartitionsRDD.scala:35)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.rdd.MapPartitionsRDD.getPartitions(MapPartitionsRDD.scala:35)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.ShuffleDependency.(Dependency.scala:91)\r\n at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec$.prepareShuffleDependency(ShuffleExchangeExec.scala:318)\r\n at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec.prepareShuffleDependency(ShuffleExchangeExec.scala:91)\r\n at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec$$anonfun$doExecute$1.apply(ShuffleExchangeExec.scala:128)\r\n at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec$$anonfun$doExecute$1.apply(ShuffleExchangeExec.scala:119)\r\n at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:52)\r\n ... 38 more\r\nCaused by: java.util.concurrent.ExecutionException: java.lang.IndexOutOfBoundsException: Index: 0\r\n at java.util.concurrent.FutureTask.report(FutureTask.java:122)\r\n at java.util.concurrent.FutureTask.get(FutureTask.java:192)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat.generateSplitsInfo(OrcInputFormat.java:1016)\r\n ... 75 more\r\nCaused by: java.lang.IndexOutOfBoundsException: Index: 0\r\n at java.util.Collections$EmptyList.get(Collections.java:4454)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Type.getSubtypes(OrcProto.java:12240)\r\n at org.apache.hadoop.hive.ql.io.orc.ReaderImpl.getColumnIndicesFromNames(ReaderImpl.java:651)\r\n at org.apache.hadoop.hive.ql.io.orc.ReaderImpl.getRawDataSizeOfColumns(ReaderImpl.java:634)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat$SplitGenerator.populateAndCacheStripeDetails(OrcInputFormat.java:927)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat$SplitGenerator.call(OrcInputFormat.java:836)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat$SplitGenerator.call(OrcInputFormat.java:702)\r\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n at java.lang.Thread.run(Thread.java:748)holder\r\n{code}","from":"reporter","subject":"Unable to process empty ORC files in Hive Table using Spark SQL"},{"body":"2.3.x is EOL releases. Can you try out in higher versions?","from":"developer"},{"body":"We will try on 2.4.4","from":"developer"},{"body":"ping [~aermakov], have you tried this out?","from":"developer"},{"body":"This issue has been resolved for Spark 2.4.4","from":"developer"}],"created":"2019-11-06T11:02:34.000+0000","description":"Unable to process empty ORC files in Hive Table using Spark SQL. It seems that a problem with class org.apache.hadoop.hive.ql.io.orc.OrcInputFormat.getSplits()\r\n\r\nStack trace:\r\n{code:java}\r\n19/10/30 22:29:54 ERROR SparkSQLDriver: Failed in [select distinct _tech_load_dt from dl_raw.tpaccsieee_ut_data_address]\r\norg.apache.spark.sql.catalyst.errors.package$TreeNodeException: execute, tree:\r\nExchange hashpartitioning(_tech_load_dt#1374, 200)\r\n+- *(1) HashAggregate(keys=[_tech_load_dt#1374], functions=[], output=[_tech_load_dt#1374])\r\n +- HiveTableScan [_tech_load_dt#1374], HiveTableRelation `dl_raw`.`tpaccsieee_ut_data_address`, org.apache.hadoop.hive.ql.io.orc.OrcSerde, [address#1307, address_9zp#1308, address_adm#1309, address_md#1310, adress_doc#1311, building#1312, change_date_addr_el#1313, change_date_okato#1314, change_date_окато#1315, city#1316, city_id#1317, cnv_cont_id#1318, code_intercity#1319, code_kladr#1320, code_plan1#1321, date_act#1322, date_change#1323, date_prz_incorrect_code_kladr#1324, date_record#1325, district#1326, district_id#1327, etaj#1328, e_plan#1329, fax#1330, ... 44 more fields] at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:56)\r\n at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec.doExecute(ShuffleExchangeExec.scala:119)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:131)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:155)\r\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\r\n at org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:152)\r\n at org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.InputAdapter.inputRDDs(WholeStageCodegenExec.scala:371)\r\n at org.apache.spark.sql.execution.aggregate.HashAggregateExec.inputRDDs(HashAggregateExec.scala:150)\r\n at org.apache.spark.sql.execution.WholeStageCodegenExec.doExecute(WholeStageCodegenExec.scala:605)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:131)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$1.apply(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan$$anonfun$executeQuery$1.apply(SparkPlan.scala:155)\r\n at org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:151)\r\n at org.apache.spark.sql.execution.SparkPlan.executeQuery(SparkPlan.scala:152)\r\n at org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:127)\r\n at org.apache.spark.sql.execution.SparkPlan.getByteArrayRdd(SparkPlan.scala:247)\r\n at org.apache.spark.sql.execution.SparkPlan.executeCollect(SparkPlan.scala:294)\r\n at org.apache.spark.sql.execution.SparkPlan.executeCollectPublic(SparkPlan.scala:324)\r\n at org.apache.spark.sql.execution.QueryExecution.hiveResultString(QueryExecution.scala:122)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLDriver$$anonfun$run$1.apply(SparkSQLDriver.scala:64)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLDriver$$anonfun$run$1.apply(SparkSQLDriver.scala:64)\r\n at org.apache.spark.sql.execution.SQLExecution$.withNewExecutionId(SQLExecution.scala:77)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLDriver.run(SparkSQLDriver.scala:63)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.processCmd(SparkSQLCLIDriver.scala:364)\r\n at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:376)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$.main(SparkSQLCLIDriver.scala:272)\r\n at org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.main(SparkSQLCLIDriver.scala)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n at java.lang.reflect.Method.invoke(Method.java:498)\r\n at org.apache.spark.deploy.JavaMainApplication.start(SparkApplication.scala:52)\r\n at org.apache.spark.deploy.SparkSubmit$.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:894)\r\n at org.apache.spark.deploy.SparkSubmit$.doRunMain$1(SparkSubmit.scala:198)\r\n at org.apache.spark.deploy.SparkSubmit$.submit(SparkSubmit.scala:228)\r\n at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:137)\r\n at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\r\nCaused by: java.lang.RuntimeException: serious problem\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat.generateSplitsInfo(OrcInputFormat.java:1021)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat.getSplits(OrcInputFormat.java:1048)\r\n at org.apache.spark.rdd.HadoopRDD.getPartitions(HadoopRDD.scala:200)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.rdd.MapPartitionsRDD.getPartitions(MapPartitionsRDD.scala:35)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.rdd.MapPartitionsRDD.getPartitions(MapPartitionsRDD.scala:35)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.rdd.MapPartitionsRDD.getPartitions(MapPartitionsRDD.scala:35)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.rdd.MapPartitionsRDD.getPartitions(MapPartitionsRDD.scala:35)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.rdd.MapPartitionsRDD.getPartitions(MapPartitionsRDD.scala:35)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:253)\r\n at org.apache.spark.rdd.RDD$$anonfun$partitions$2.apply(RDD.scala:251)\r\n at scala.Option.getOrElse(Option.scala:121)\r\n at org.apache.spark.rdd.RDD.partitions(RDD.scala:251)\r\n at org.apache.spark.ShuffleDependency.(Dependency.scala:91)\r\n at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec$.prepareShuffleDependency(ShuffleExchangeExec.scala:318)\r\n at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec.prepareShuffleDependency(ShuffleExchangeExec.scala:91)\r\n at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec$$anonfun$doExecute$1.apply(ShuffleExchangeExec.scala:128)\r\n at org.apache.spark.sql.execution.exchange.ShuffleExchangeExec$$anonfun$doExecute$1.apply(ShuffleExchangeExec.scala:119)\r\n at org.apache.spark.sql.catalyst.errors.package$.attachTree(package.scala:52)\r\n ... 38 more\r\nCaused by: java.util.concurrent.ExecutionException: java.lang.IndexOutOfBoundsException: Index: 0\r\n at java.util.concurrent.FutureTask.report(FutureTask.java:122)\r\n at java.util.concurrent.FutureTask.get(FutureTask.java:192)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat.generateSplitsInfo(OrcInputFormat.java:1016)\r\n ... 75 more\r\nCaused by: java.lang.IndexOutOfBoundsException: Index: 0\r\n at java.util.Collections$EmptyList.get(Collections.java:4454)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcProto$Type.getSubtypes(OrcProto.java:12240)\r\n at org.apache.hadoop.hive.ql.io.orc.ReaderImpl.getColumnIndicesFromNames(ReaderImpl.java:651)\r\n at org.apache.hadoop.hive.ql.io.orc.ReaderImpl.getRawDataSizeOfColumns(ReaderImpl.java:634)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat$SplitGenerator.populateAndCacheStripeDetails(OrcInputFormat.java:927)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat$SplitGenerator.call(OrcInputFormat.java:836)\r\n at org.apache.hadoop.hive.ql.io.orc.OrcInputFormat$SplitGenerator.call(OrcInputFormat.java:702)\r\n at java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)\r\n at java.lang.Thread.run(Thread.java:748)holder\r\n{code}","issue_id":"13266553","key":"SPARK-29773","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2020-06-22T15:21:25.000+0000","role":"fixed_distractor","summary":"Unable to process empty ORC files in Hive Table using Spark SQL"} {"case_id":"12733707","cluster":"DISTRACTOR-SPARK-3005","comments":[{"body":"Some additional driver logs during the spark driver hang:\n{code}\nTRACE [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,908 Logging.scala (line 66) Checking for newly runnable parent stages\n INFO [Result resolver thread-1] 2014-08-13 15:58:15,908 Logging.scala (line 58) Removed TaskSet 1.0, whose tasks have all completed, from pool \nTRACE [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,909 Logging.scala (line 66) running: Set(Stage 1, Stage 2)\nTRACE [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,909 Logging.scala (line 66) waiting: Set(Stage 0)\nTRACE [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,909 Logging.scala (line 66) failed: Set()\nDEBUG [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,909 Logging.scala (line 62) submitStage(Stage 0)\nDEBUG [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,910 Logging.scala (line 62) missing: List(Stage 1, Stage 2)\nDEBUG [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,910 Logging.scala (line 62) submitStage(Stage 1)\nDEBUG [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,910 Logging.scala (line 62) submitStage(Stage 2)\nTRACE [spark-akka.actor.default-dispatcher-3] 2014-08-13 15:58:56,643 Logging.scala (line 66) Checking for hosts with no recent heart beats in BlockManagerMaster.\nTRACE [spark-akka.actor.default-dispatcher-6] 2014-08-13 15:59:56,653 Logging.scala (line 66) Checking for hosts with no recent heart beats in BlockManagerMaster.\nTRACE [spark-akka.actor.default-dispatcher-2] 2014-08-13 16:00:56,652 Logging.scala (line 66) Checking for hosts with no recent heart beats in BlockManagerMaster.\n{code}","created":"2014-08-13T08:48:03.300+0000"},{"body":"a quick fix for fine grained killTask\n","created":"2014-08-14T03:59:55.919+0000"},{"body":"Could adding an empty killTask method to MesosSchedulerBackend fix this problem?\n{code:title=MesosSchedulerBackend.scala}\noverride def killTask(taskId: Long, executorId: String, interruptThread: Boolean) {}\n{code}\nThis works for my tests.","created":"2014-08-14T04:20:23.463+0000"},{"body":"User 'xuzhongxing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/1940","created":"2014-08-14T08:07:50.159+0000"},{"body":"Hey There - the message you are seeing does not itself indicate a failure. It is just saying that Mesos does not support cancellation (see SPARK-1749). The exact reason for the hanging is unclear.","created":"2014-08-24T17:05:34.426+0000"},{"body":"I was using spark to process data from Cassandra. And when Cassandra is under heavy load, the executors on the slaves throws timeout exception, and the tasks fail. Then the driver need to cancel the job.\n\nIn coarse-grained mode, it works fine. The coarse-grained backend has the killTask() method.\n\nBut in fine-grained mode, the backend didn't override the killTask() method. So it throws exception.\n\nI could tune Cassandra to avoid read timeout. But I just want spark to exit when the job fails. Currently it just throws that operation-not-supported exception and hangs there. This is unacceptable behaviour.","created":"2014-08-25T02:02:25.570+0000"},{"body":"[SPARK-1749] didn't fix this problem. It just catches the UnsupportedOperationException and logs it. Then it sets ableToCancelStages = false. This is exactly the reason that causes the hang. Because the code only does cleanup when ableToCancelStages = true.\n{code}\n+ if (ableToCancelStages) {\n+ job.listener.jobFailed(error)\n+ cleanupStateForJobAndIndependentStages(job, resultStage)\n+ listenerBus.post(SparkListenerJobEnd(job.jobId, JobFailed(error)))\n{code}\nThe fact is that in the mesos fine-grained case, it is unnecessary to killTask(). So throwing UnsupportedOperationException and set ableToCancelStages = false is wrong behaviour for this case. We just need to do nothing in killTask() and let the driver do the rest of the cleanup.\n\nThe problem here is in the MesosSchedulerBackend. The MesosSchedulerBackend does not need to kill tasks, and should not throw UnsupportedOperationException. The tasks themselves already died and exited.\n","created":"2014-08-25T02:16:06.549+0000"},{"body":"I think it might be safe to just define the semantics of cancelTask to be \"best effort\" and have the default implementation be empty. From what I can tell the DAGScheduler is already resilient to a task finishing after the stage has been cleaned up (for other reasons I think it needs be resilient to this). So maybe we should just remove this entire ableToCancelStages logic in here and make it simpler. /cc [~kayousterhout] and [~markhamstra] who worked on this code.\n\nHey [~xuzhongxing] what do you mean that the \"tasks themselves already died and exited\"? The code here is designed to cancel outstanding tasks that are still running. For instance, I have a job that has 500 tasks running. Then there is a failure of one task multiple time so I need to fail the stage, but there are still many tasks running. I think those tasks need to be killed still. Otherwise you have zombie tasks running. I think in mesos fine-grained mode we just won't be able to support this feature... but we should make it so it doesn't hang.\n","created":"2014-09-02T05:36:06.652+0000"},{"body":"By \"tasks themselves already died and exited\", I mean that even if we do nothing in killTasks(), there won't be any zombie tasks left on the slaves. This is what I get from testing the Mesos fine-grained mode. If I'm wrong, please correct me. But the logic here is incomplete or inconsistent, and needs to be fixed.","created":"2014-09-02T09:56:07.905+0000"},{"body":"Resolved in https://github.com/apache/spark/pull/2453","created":"2014-10-08T01:32:44.052+0000"}],"conversations":[{"body":"I am using Spark, Mesos, spark-cassandra-connector to do some work on a cassandra cluster.\n\nDuring the job running, I killed the Cassandra daemon to simulate some failure cases. This results in task failures.\n\nIf I run the job in Mesos coarse-grained mode, the spark driver program throws an exception and shutdown cleanly.\n\nBut when I run the job in Mesos fine-grained mode, the spark driver program hangs.\n\nThe spark log is: \n\n{code}\n INFO [spark-akka.actor.default-dispatcher-4] 2014-08-13 15:58:15,794 Logging.scala (line 58) Cancelling stage 1\n INFO [spark-akka.actor.default-dispatcher-4] 2014-08-13 15:58:15,797 Logging.scala (line 79) Could not cancel tasks for stage 1\njava.lang.UnsupportedOperationException\n\tat org.apache.spark.scheduler.SchedulerBackend$class.killTask(SchedulerBackend.scala:32)\n\tat org.apache.spark.scheduler.cluster.mesos.MesosSchedulerBackend.killTask(MesosSchedulerBackend.scala:41)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply$mcVJ$sp(TaskSchedulerImpl.scala:185)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply(TaskSchedulerImpl.scala:183)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply(TaskSchedulerImpl.scala:183)\n\tat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3.apply(TaskSchedulerImpl.scala:183)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3.apply(TaskSchedulerImpl.scala:176)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl.cancelTasks(TaskSchedulerImpl.scala:176)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply$mcVI$sp(DAGScheduler.scala:1075)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply(DAGScheduler.scala:1061)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply(DAGScheduler.scala:1061)\n\tat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\n\tat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1061)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1033)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1031)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n\tat org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1031)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:635)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:635)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:635)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor$$anonfun$receive$2.applyOrElse(DAGScheduler.scala:1234)\n\tat akka.actor.ActorCell.receiveMessage(ActorCell.scala:498)\n\tat akka.actor.ActorCell.invoke(ActorCell.scala:456)\n\tat akka.dispatch.Mailbox.processMailbox(Mailbox.scala:237)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:219)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n{code}","from":"reporter","subject":"Spark with Mesos fine-grained mode throws UnsupportedOperationException in MesosSchedulerBackend.killTask()"},{"body":"Some additional driver logs during the spark driver hang:\n{code}\nTRACE [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,908 Logging.scala (line 66) Checking for newly runnable parent stages\n INFO [Result resolver thread-1] 2014-08-13 15:58:15,908 Logging.scala (line 58) Removed TaskSet 1.0, whose tasks have all completed, from pool \nTRACE [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,909 Logging.scala (line 66) running: Set(Stage 1, Stage 2)\nTRACE [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,909 Logging.scala (line 66) waiting: Set(Stage 0)\nTRACE [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,909 Logging.scala (line 66) failed: Set()\nDEBUG [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,909 Logging.scala (line 62) submitStage(Stage 0)\nDEBUG [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,910 Logging.scala (line 62) missing: List(Stage 1, Stage 2)\nDEBUG [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,910 Logging.scala (line 62) submitStage(Stage 1)\nDEBUG [spark-akka.actor.default-dispatcher-2] 2014-08-13 15:58:15,910 Logging.scala (line 62) submitStage(Stage 2)\nTRACE [spark-akka.actor.default-dispatcher-3] 2014-08-13 15:58:56,643 Logging.scala (line 66) Checking for hosts with no recent heart beats in BlockManagerMaster.\nTRACE [spark-akka.actor.default-dispatcher-6] 2014-08-13 15:59:56,653 Logging.scala (line 66) Checking for hosts with no recent heart beats in BlockManagerMaster.\nTRACE [spark-akka.actor.default-dispatcher-2] 2014-08-13 16:00:56,652 Logging.scala (line 66) Checking for hosts with no recent heart beats in BlockManagerMaster.\n{code}","from":"developer"},{"body":"a quick fix for fine grained killTask\n","from":"developer"},{"body":"Could adding an empty killTask method to MesosSchedulerBackend fix this problem?\n{code:title=MesosSchedulerBackend.scala}\noverride def killTask(taskId: Long, executorId: String, interruptThread: Boolean) {}\n{code}\nThis works for my tests.","from":"developer"},{"body":"User 'xuzhongxing' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/1940","from":"developer"},{"body":"Hey There - the message you are seeing does not itself indicate a failure. It is just saying that Mesos does not support cancellation (see SPARK-1749). The exact reason for the hanging is unclear.","from":"developer"},{"body":"I was using spark to process data from Cassandra. And when Cassandra is under heavy load, the executors on the slaves throws timeout exception, and the tasks fail. Then the driver need to cancel the job.\n\nIn coarse-grained mode, it works fine. The coarse-grained backend has the killTask() method.\n\nBut in fine-grained mode, the backend didn't override the killTask() method. So it throws exception.\n\nI could tune Cassandra to avoid read timeout. But I just want spark to exit when the job fails. Currently it just throws that operation-not-supported exception and hangs there. This is unacceptable behaviour.","from":"developer"},{"body":"[SPARK-1749] didn't fix this problem. It just catches the UnsupportedOperationException and logs it. Then it sets ableToCancelStages = false. This is exactly the reason that causes the hang. Because the code only does cleanup when ableToCancelStages = true.\n{code}\n+ if (ableToCancelStages) {\n+ job.listener.jobFailed(error)\n+ cleanupStateForJobAndIndependentStages(job, resultStage)\n+ listenerBus.post(SparkListenerJobEnd(job.jobId, JobFailed(error)))\n{code}\nThe fact is that in the mesos fine-grained case, it is unnecessary to killTask(). So throwing UnsupportedOperationException and set ableToCancelStages = false is wrong behaviour for this case. We just need to do nothing in killTask() and let the driver do the rest of the cleanup.\n\nThe problem here is in the MesosSchedulerBackend. The MesosSchedulerBackend does not need to kill tasks, and should not throw UnsupportedOperationException. The tasks themselves already died and exited.\n","from":"developer"},{"body":"I think it might be safe to just define the semantics of cancelTask to be \"best effort\" and have the default implementation be empty. From what I can tell the DAGScheduler is already resilient to a task finishing after the stage has been cleaned up (for other reasons I think it needs be resilient to this). So maybe we should just remove this entire ableToCancelStages logic in here and make it simpler. /cc [~kayousterhout] and [~markhamstra] who worked on this code.\n\nHey [~xuzhongxing] what do you mean that the \"tasks themselves already died and exited\"? The code here is designed to cancel outstanding tasks that are still running. For instance, I have a job that has 500 tasks running. Then there is a failure of one task multiple time so I need to fail the stage, but there are still many tasks running. I think those tasks need to be killed still. Otherwise you have zombie tasks running. I think in mesos fine-grained mode we just won't be able to support this feature... but we should make it so it doesn't hang.\n","from":"developer"},{"body":"By \"tasks themselves already died and exited\", I mean that even if we do nothing in killTasks(), there won't be any zombie tasks left on the slaves. This is what I get from testing the Mesos fine-grained mode. If I'm wrong, please correct me. But the logic here is incomplete or inconsistent, and needs to be fixed.","from":"developer"},{"body":"Resolved in https://github.com/apache/spark/pull/2453","from":"developer"}],"created":"2014-08-13T08:18:06.000+0000","description":"I am using Spark, Mesos, spark-cassandra-connector to do some work on a cassandra cluster.\n\nDuring the job running, I killed the Cassandra daemon to simulate some failure cases. This results in task failures.\n\nIf I run the job in Mesos coarse-grained mode, the spark driver program throws an exception and shutdown cleanly.\n\nBut when I run the job in Mesos fine-grained mode, the spark driver program hangs.\n\nThe spark log is: \n\n{code}\n INFO [spark-akka.actor.default-dispatcher-4] 2014-08-13 15:58:15,794 Logging.scala (line 58) Cancelling stage 1\n INFO [spark-akka.actor.default-dispatcher-4] 2014-08-13 15:58:15,797 Logging.scala (line 79) Could not cancel tasks for stage 1\njava.lang.UnsupportedOperationException\n\tat org.apache.spark.scheduler.SchedulerBackend$class.killTask(SchedulerBackend.scala:32)\n\tat org.apache.spark.scheduler.cluster.mesos.MesosSchedulerBackend.killTask(MesosSchedulerBackend.scala:41)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply$mcVJ$sp(TaskSchedulerImpl.scala:185)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply(TaskSchedulerImpl.scala:183)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3$$anonfun$apply$1.apply(TaskSchedulerImpl.scala:183)\n\tat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3.apply(TaskSchedulerImpl.scala:183)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl$$anonfun$cancelTasks$3.apply(TaskSchedulerImpl.scala:176)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.TaskSchedulerImpl.cancelTasks(TaskSchedulerImpl.scala:176)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply$mcVI$sp(DAGScheduler.scala:1075)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply(DAGScheduler.scala:1061)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages$1.apply(DAGScheduler.scala:1061)\n\tat scala.collection.mutable.HashSet.foreach(HashSet.scala:79)\n\tat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1061)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1033)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1031)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n\tat org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1031)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:635)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:635)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:635)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessActor$$anonfun$receive$2.applyOrElse(DAGScheduler.scala:1234)\n\tat akka.actor.ActorCell.receiveMessage(ActorCell.scala:498)\n\tat akka.actor.ActorCell.invoke(ActorCell.scala:456)\n\tat akka.dispatch.Mailbox.processMailbox(Mailbox.scala:237)\n\tat akka.dispatch.Mailbox.run(Mailbox.scala:219)\n\tat akka.dispatch.ForkJoinExecutorConfigurator$AkkaForkJoinTask.exec(AbstractDispatcher.scala:386)\n\tat scala.concurrent.forkjoin.ForkJoinTask.doExec(ForkJoinTask.java:260)\n\tat scala.concurrent.forkjoin.ForkJoinPool$WorkQueue.runTask(ForkJoinPool.java:1339)\n\tat scala.concurrent.forkjoin.ForkJoinPool.runWorker(ForkJoinPool.java:1979)\n\tat scala.concurrent.forkjoin.ForkJoinWorkerThread.run(ForkJoinWorkerThread.java:107)\n{code}","issue_id":"12733707","key":"SPARK-3005","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2014-10-08T01:32:04.000+0000","role":"fixed_distractor","summary":"Spark with Mesos fine-grained mode throws UnsupportedOperationException in MesosSchedulerBackend.killTask()"} {"case_id":"12735492","cluster":"DISTRACTOR-SPARK-3151","comments":[{"body":"User 'squito' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4857","created":"2015-03-02T21:44:47.078+0000"},{"body":"I just hit this with spark 2.1 when processing a disk persisted RDD\nwhile the root cause for this is probably data skew, this seems like a severe limitation on spark.\nthis specific failure is a bit surprising as spark already knows the block size on disk at this point (it maps the entire file from offset 0 to file size), so it should easily be possible to split this into several blocks, after all the code is using ChunkedByteBuffer.\na better approach would be lazily loading the blocks as deserialization progresses , and an even better solution (proposed by [~matei] in [comment|https://issues.apache.org/jira/browse/SPARK-1476?focusedCommentId=13967947&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-13967947] for #1476) would be splitting these block in higher levels of spark (a mapper that produces multiple blocks, cached RDD partition that consists of several blocks...)\n\nI can see the attached PR is closed, is there an expected/in-progress fix for this?\n\n\n2017-07-05 07:04:51,368 UTC\tWARN \ttask-result-getter-0\torg.apache.spark.scheduler.TaskSetManager\t Lost task 131.0 in stage 14.0 (TID 2228, 172.20.1.137, executor 1): java.lang.IllegalArgumentException: Size exceeds Integer.MAX_VALUE\n\tat sun.nio.ch.FileChannelImpl.map(FileChannelImpl.java:869)\n\tat org.apache.spark.storage.DiskStore$$anonfun$getBytes$2.apply(DiskStore.scala:103)\n\tat org.apache.spark.storage.DiskStore$$anonfun$getBytes$2.apply(DiskStore.scala:91)\n\tat org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1303)\n\tat org.apache.spark.storage.DiskStore.getBytes(DiskStore.scala:105)\n\tat org.apache.spark.storage.BlockManager.getLocalValues(BlockManager.scala:465)\n\tat org.apache.spark.storage.BlockManager.getOrElseUpdate(BlockManager.scala:701)\n\tat org.apache.spark.rdd.RDD.getOrCompute(RDD.scala:334)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:285)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.rdd.UnionRDD.compute(UnionRDD.scala:105)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:99)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:282)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:748)","created":"2017-07-05T08:06:49.749+0000"},{"body":"created PR 18855: https://github.com/apache/spark/pull/18855","created":"2017-08-08T13:49:40.923+0000"},{"body":"Issue resolved by pull request 18855\n[https://github.com/apache/spark/pull/18855]","created":"2017-08-17T01:22:48.900+0000"}],"conversations":[{"body":"[DiskStore|https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/storage/DiskStore.scala] attempts to memory map the block file in {{def getBytes}}. If the file is larger than 2GB (Integer.MAX_VALUE) as specified by [FileChannel.map|http://docs.oracle.com/javase/7/docs/api/java/nio/channels/FileChannel.html#map%28java.nio.channels.FileChannel.MapMode,%20long,%20long%29], then the memory map fails.\n\n{code}\nSome(channel.map(MapMode.READ_ONLY, segment.offset, segment.length)) # line 104\n{code}","from":"reporter","subject":"DiskStore attempts to map any size BlockId without checking MappedByteBuffer limit"},{"body":"User 'squito' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4857","from":"developer"},{"body":"I just hit this with spark 2.1 when processing a disk persisted RDD\nwhile the root cause for this is probably data skew, this seems like a severe limitation on spark.\nthis specific failure is a bit surprising as spark already knows the block size on disk at this point (it maps the entire file from offset 0 to file size), so it should easily be possible to split this into several blocks, after all the code is using ChunkedByteBuffer.\na better approach would be lazily loading the blocks as deserialization progresses , and an even better solution (proposed by [~matei] in [comment|https://issues.apache.org/jira/browse/SPARK-1476?focusedCommentId=13967947&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-13967947] for #1476) would be splitting these block in higher levels of spark (a mapper that produces multiple blocks, cached RDD partition that consists of several blocks...)\n\nI can see the attached PR is closed, is there an expected/in-progress fix for this?\n\n\n2017-07-05 07:04:51,368 UTC\tWARN \ttask-result-getter-0\torg.apache.spark.scheduler.TaskSetManager\t Lost task 131.0 in stage 14.0 (TID 2228, 172.20.1.137, executor 1): java.lang.IllegalArgumentException: Size exceeds Integer.MAX_VALUE\n\tat sun.nio.ch.FileChannelImpl.map(FileChannelImpl.java:869)\n\tat org.apache.spark.storage.DiskStore$$anonfun$getBytes$2.apply(DiskStore.scala:103)\n\tat org.apache.spark.storage.DiskStore$$anonfun$getBytes$2.apply(DiskStore.scala:91)\n\tat org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1303)\n\tat org.apache.spark.storage.DiskStore.getBytes(DiskStore.scala:105)\n\tat org.apache.spark.storage.BlockManager.getLocalValues(BlockManager.scala:465)\n\tat org.apache.spark.storage.BlockManager.getOrElseUpdate(BlockManager.scala:701)\n\tat org.apache.spark.rdd.RDD.getOrCompute(RDD.scala:334)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:285)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.rdd.UnionRDD.compute(UnionRDD.scala:105)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:323)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:287)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:99)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:282)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:748)","from":"developer"},{"body":"created PR 18855: https://github.com/apache/spark/pull/18855","from":"developer"},{"body":"Issue resolved by pull request 18855\n[https://github.com/apache/spark/pull/18855]","from":"developer"}],"created":"2014-08-20T18:41:17.000+0000","description":"[DiskStore|https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/storage/DiskStore.scala] attempts to memory map the block file in {{def getBytes}}. If the file is larger than 2GB (Integer.MAX_VALUE) as specified by [FileChannel.map|http://docs.oracle.com/javase/7/docs/api/java/nio/channels/FileChannel.html#map%28java.nio.channels.FileChannel.MapMode,%20long,%20long%29], then the memory map fails.\n\n{code}\nSome(channel.map(MapMode.READ_ONLY, segment.offset, segment.length)) # line 104\n{code}","issue_id":"12735492","key":"SPARK-3151","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-08-17T01:22:48.000+0000","role":"fixed_distractor","summary":"DiskStore attempts to map any size BlockId without checking MappedByteBuffer limit"} {"case_id":"13366410","cluster":"DISTRACTOR-SPARK-34805","comments":[{"body":"I believe this is not a PySpark-specific issue. We have a unit test in [transmogif.ai|https://transmogrif.ai/] where we are specifying [column metadata manually|https://github.com/salesforce/TransmogrifAI/blob/90a0f298f14506a27c84a71de414d53a30cf687f/core/src/test/scala/com/salesforce/op/stages/impl/preparators/SanityCheckerTest.scala#L137] and check whether the metadata is properly passed on to a model that consumes this column. The column metadata is properly given to the column using {{.as(columnName, metadata)}}, but is immediately lost once the select is executed. I've traced the issue to the changes in {{ExpressionEncoder}}:\r\n * In Spark 2.4, [it takes it a schema argument|https://github.com/apache/spark/blob/e89526d2401b3a04719721c923a6f630e555e286/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/encoders/ExpressionEncoder.scala#L222] through which the column metadata is passed along\r\n * In Spark 3.0, [it no longer takes|https://github.com/apache/spark/blob/39889df32a7a916d826e255fda6fc62e2a3d7971/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/encoders/ExpressionEncoder.scala#L232] this schema parameter and it seems like the column metadata is lost as a result\r\n\r\nI can't tell if this was intentional or not, but it renders the metadata argument of the {{.as}} [method|https://github.com/apache/spark/blob/39889df32a7a916d826e255fda6fc62e2a3d7971/sql/core/src/main/scala/org/apache/spark/sql/Column.scala#L1133] in {{Column}} mostly useless.","created":"2021-04-30T16:23:40.446+0000"},{"body":"The problem happens in Scala as well. I attached a scala file [^nested_columns_metadata.scala] to demonstrate the issue. I tried it in the spark-shell of versions 2.4.7, 3.1.2 and 3.2.0, always with the same result. This behavior is a bug, because the documentation for {{StructField}} clearly says that the \"metadata should be preserved during transformation if the content of the column is not modified, e.g, in selection\"","created":"2022-01-19T11:51:51.025+0000"},{"body":"User 'kevinwallimann' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/35270","created":"2022-01-21T10:47:09.575+0000"},{"body":"User 'kevinwallimann' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/35270","created":"2022-01-21T10:47:38.050+0000"},{"body":"Issue resolved by pull request 35270\n[https://github.com/apache/spark/pull/35270]","created":"2022-03-21T16:38:49.577+0000"}],"conversations":[{"body":"For a DataFrame schema with nested StructTypes, where metadata is set for fields in the schema, that metadata is lost when a DataFrame selects nested fields.  For example, suppose\r\n{code:java}\r\ndf.schema.fields[0].dataType.fields[0].metadata\r\n{code}\r\nreturns a non-empty dictionary, then\r\n{code:java}\r\ndf.select('Field0.SubField0').schema.fields[0].metadata{code}\r\nreturns an empty dictionary, where \"Field0\" is the name of the first field in the DataFrame and \"SubField0\" is the name of the first nested field under \"Field0\".\r\n\r\n ","from":"reporter","subject":"PySpark loses metadata in DataFrame fields when selecting nested columns"},{"body":"I believe this is not a PySpark-specific issue. We have a unit test in [transmogif.ai|https://transmogrif.ai/] where we are specifying [column metadata manually|https://github.com/salesforce/TransmogrifAI/blob/90a0f298f14506a27c84a71de414d53a30cf687f/core/src/test/scala/com/salesforce/op/stages/impl/preparators/SanityCheckerTest.scala#L137] and check whether the metadata is properly passed on to a model that consumes this column. The column metadata is properly given to the column using {{.as(columnName, metadata)}}, but is immediately lost once the select is executed. I've traced the issue to the changes in {{ExpressionEncoder}}:\r\n * In Spark 2.4, [it takes it a schema argument|https://github.com/apache/spark/blob/e89526d2401b3a04719721c923a6f630e555e286/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/encoders/ExpressionEncoder.scala#L222] through which the column metadata is passed along\r\n * In Spark 3.0, [it no longer takes|https://github.com/apache/spark/blob/39889df32a7a916d826e255fda6fc62e2a3d7971/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/encoders/ExpressionEncoder.scala#L232] this schema parameter and it seems like the column metadata is lost as a result\r\n\r\nI can't tell if this was intentional or not, but it renders the metadata argument of the {{.as}} [method|https://github.com/apache/spark/blob/39889df32a7a916d826e255fda6fc62e2a3d7971/sql/core/src/main/scala/org/apache/spark/sql/Column.scala#L1133] in {{Column}} mostly useless.","from":"developer"},{"body":"The problem happens in Scala as well. I attached a scala file [^nested_columns_metadata.scala] to demonstrate the issue. I tried it in the spark-shell of versions 2.4.7, 3.1.2 and 3.2.0, always with the same result. This behavior is a bug, because the documentation for {{StructField}} clearly says that the \"metadata should be preserved during transformation if the content of the column is not modified, e.g, in selection\"","from":"developer"},{"body":"User 'kevinwallimann' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/35270","from":"developer"},{"body":"User 'kevinwallimann' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/35270","from":"developer"},{"body":"Issue resolved by pull request 35270\n[https://github.com/apache/spark/pull/35270]","from":"developer"}],"created":"2021-03-19T17:42:48.000+0000","description":"For a DataFrame schema with nested StructTypes, where metadata is set for fields in the schema, that metadata is lost when a DataFrame selects nested fields.  For example, suppose\r\n{code:java}\r\ndf.schema.fields[0].dataType.fields[0].metadata\r\n{code}\r\nreturns a non-empty dictionary, then\r\n{code:java}\r\ndf.select('Field0.SubField0').schema.fields[0].metadata{code}\r\nreturns an empty dictionary, where \"Field0\" is the name of the first field in the DataFrame and \"SubField0\" is the name of the first nested field under \"Field0\".\r\n\r\n ","issue_id":"13366410","key":"SPARK-34805","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2022-03-21T16:38:49.000+0000","role":"fixed_distractor","summary":"PySpark loses metadata in DataFrame fields when selecting nested columns"} {"case_id":"13395114","cluster":"DISTRACTOR-SPARK-36509","comments":[{"body":"User 'sarutak' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/33818","created":"2021-08-24T06:37:00.446+0000"},{"body":"Issue resolved in https://github.com/apache/spark/pull/33818","created":"2021-08-28T09:07:33.151+0000"}],"conversations":[{"body":"This is reproducible with an application that uses less cores than what are available on the workers:\r\n\r\nE.g. with 1 application with 1 executor, when the worker with the executor is killed, the application will not get another executor assigned even if there are enough resources in the cluster. This seems to be a regression, caused by [https://github.com/apache/spark/commit/51de86baed0776304c6184f2c04b6303ef48df90#diff-ca694acef669f50f9b45ca0d32ab6f5a516270bb26b33c4abb704e2dc00a1a03] .\r\n\r\nThat causes an assertion error on the master because it get's an executorStateChange from 'RUNNING' to 'RUNNING' instead of 'FAILED':\r\n{noformat}\r\n2021-08-13 14:04:12,554 [dispatcher-event-loop-2] INFO : I have been elected leader! New state: ALIVE\r\n2021-08-13 14:04:12,554 [dispatcher-event-loop-2] INFO : I have been elected leader! New state: ALIVE\r\n2021-08-13 14:04:56,489 [dispatcher-event-loop-10] INFO : Registering worker 172.27.64.1:58636 with 12 cores, 30.7 GiB RAM\r\n2021-08-13 14:04:59,949 [dispatcher-event-loop-6] INFO : Registering worker 172.27.64.1:58694 with 12 cores, 30.7 GiB RAM\r\n2021-08-13 14:05:20,212 [dispatcher-event-loop-2] INFO : Registering app query-frontend-null-172.27.64.1\r\n2021-08-13 14:05:20,212 [dispatcher-event-loop-2] INFO : Registered app query-frontend-null-172.27.64.1 with ID app-20210813140520-0000\r\n2021-08-13 14:05:20,228 [dispatcher-event-loop-2] INFO : Launching executor app-20210813140520-0000/0 on worker worker-20210813140459-172.27.64.1-58694\r\n2021-08-13 14:05:37,991 [dispatcher-event-loop-9] ERROR: Ignoring errorjava.lang.AssertionError: assertion failed: executor 0 state transfer from RUNNING to RUNNING is illegal at scala.Predef$.assert(Predef.scala:223) at org.apache.spark.deploy.master.Master$$anonfun$receive$1.applyOrElse(Master.scala:323) at org.apache.spark.rpc.netty.Inbox.$anonfun$process$1(Inbox.scala:115)\r\n {noformat}\r\n ","from":"reporter","subject":"Executors don't get rescheduled in standalone mode when worker dies"},{"body":"User 'sarutak' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/33818","from":"developer"},{"body":"Issue resolved in https://github.com/apache/spark/pull/33818","from":"developer"}],"created":"2021-08-13T12:36:43.000+0000","description":"This is reproducible with an application that uses less cores than what are available on the workers:\r\n\r\nE.g. with 1 application with 1 executor, when the worker with the executor is killed, the application will not get another executor assigned even if there are enough resources in the cluster. This seems to be a regression, caused by [https://github.com/apache/spark/commit/51de86baed0776304c6184f2c04b6303ef48df90#diff-ca694acef669f50f9b45ca0d32ab6f5a516270bb26b33c4abb704e2dc00a1a03] .\r\n\r\nThat causes an assertion error on the master because it get's an executorStateChange from 'RUNNING' to 'RUNNING' instead of 'FAILED':\r\n{noformat}\r\n2021-08-13 14:04:12,554 [dispatcher-event-loop-2] INFO : I have been elected leader! New state: ALIVE\r\n2021-08-13 14:04:12,554 [dispatcher-event-loop-2] INFO : I have been elected leader! New state: ALIVE\r\n2021-08-13 14:04:56,489 [dispatcher-event-loop-10] INFO : Registering worker 172.27.64.1:58636 with 12 cores, 30.7 GiB RAM\r\n2021-08-13 14:04:59,949 [dispatcher-event-loop-6] INFO : Registering worker 172.27.64.1:58694 with 12 cores, 30.7 GiB RAM\r\n2021-08-13 14:05:20,212 [dispatcher-event-loop-2] INFO : Registering app query-frontend-null-172.27.64.1\r\n2021-08-13 14:05:20,212 [dispatcher-event-loop-2] INFO : Registered app query-frontend-null-172.27.64.1 with ID app-20210813140520-0000\r\n2021-08-13 14:05:20,228 [dispatcher-event-loop-2] INFO : Launching executor app-20210813140520-0000/0 on worker worker-20210813140459-172.27.64.1-58694\r\n2021-08-13 14:05:37,991 [dispatcher-event-loop-9] ERROR: Ignoring errorjava.lang.AssertionError: assertion failed: executor 0 state transfer from RUNNING to RUNNING is illegal at scala.Predef$.assert(Predef.scala:223) at org.apache.spark.deploy.master.Master$$anonfun$receive$1.applyOrElse(Master.scala:323) at org.apache.spark.rpc.netty.Inbox.$anonfun$process$1(Inbox.scala:115)\r\n {noformat}\r\n ","issue_id":"13395114","key":"SPARK-36509","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2021-08-28T09:07:33.000+0000","role":"fixed_distractor","summary":"Executors don't get rescheduled in standalone mode when worker dies"} {"case_id":"12748414","cluster":"DISTRACTOR-SPARK-3958","comments":[{"body":"Digging into this stacktrace in more detail:\n\nSnappy-java's {{SnappyOutputStream}} writes its own 8-byte header at the beginning of the serialized output. This header consists of an 8-byte magic value followed by two 4-byte version numbers. This header is distinct from Snappy's own 6-byte magic number / header.\n\n{{org.xerial.snappy.SnappyInputStream.readHeader}} is implemented like this (in Snappy-Java 1.1.1.3):\n\n{code}\nprotected void readHeader() throws IOException {\n byte[] header = new byte[SnappyCodec.headerSize()];\n int readBytes = 0;\n while (readBytes < header.length) {\n int ret = in.read(header, readBytes, header.length - readBytes);\n if (ret == -1)\n break;\n readBytes += ret;\n }\n\n // Quick test of the header \n if (readBytes < header.length || header[0] != SnappyCodec.MAGIC_HEADER[0]) {\n // do the default uncompression\n readFully(header, readBytes);\n return;\n }\n\n SnappyCodec codec = SnappyCodec.readHeader(new ByteArrayInputStream(header));\n if (codec.isValidMagicHeader()) {\n // The input data is compressed by SnappyOutputStream\n if (codec.version < SnappyCodec.MINIMUM_COMPATIBLE_VERSION) {\n throw new IOException(String.format(\n \"compressed with imcompatible codec version %d. At least version %d is required\",\n codec.version, SnappyCodec.MINIMUM_COMPATIBLE_VERSION));\n }\n }\n else {\n // (probably) compressed by Snappy.compress(byte[])\n readFully(header, readBytes);\n return;\n }\n }\n{code}\n\nIt starts by attempting to read the 8-byte header. The first {{while}} loop exits when we've either read 8 bytes of header data or if the input stream was closed before it could read a complete header. The following code checks whether the header is unexpectedly short or whether it doesn't match the snappy-java magic header. In our case, we end up taking this branch and calling {{readFully(header, readBytes)}} in order to perform the default Snappy decompression This is the wrong branch to take (since our data was compressed with a SnappyOutputStream), leading to the PARSING_ERROR.\n\nBased on this, I think that the input data to the SnappyInputStream is somehow being corrupted. It's not obvious whether this corruption is causing the input data to be too short or whether the start of the stream has the wrong contents. I'll keep digging and look into adding some size-checking assertions throughout our code.","created":"2014-10-15T21:11:26.005+0000"},{"body":"Removing 1.1.0 as an affected version for now, since the stacktrace that I posted here was from a recent build of master (1.2). If anyone can reproduce this in branch-1.1 or Spark 1.1.0, please let me know.","created":"2014-10-15T22:08:35.954+0000"},{"body":"I think that I can safely rule out problems in TorrentBroadcast's (de-)blockification code: I used ScalaCheck to write some tests to ensure that blockifyObject and unblockifyObject are inverses, plus a similar test for ByteArrayChunkOutputStream: https://github.com/JoshRosen/spark/commit/413be7f6a8d4eb14c69c7db87e2564ed4d776c42?diff=unified","created":"2014-10-15T23:49:20.346+0000"},{"body":"Hi Josh, have you tried other compression like LZO to narrow down the problem?","created":"2014-10-16T01:16:12.868+0000"},{"body":"Hi [~jerryshao],\n\nI don't have a reliable reproduction for this issue yet, so I haven't tried switching compression schemes or broadcast implementations. I'm working with the user who provided this stack trace to see if we can get more logs to provide additional context.","created":"2014-10-16T01:26:45.905+0000"},{"body":"[~davies] ran across this exception while testing a pull request that modifies TorrentBroadcast: https://github.com/apache/spark/pull/2681#issuecomment-59120483\n\nThat PR's reproductioncould be a valuable debugging clue for this issue.","created":"2014-10-17T02:29:24.669+0000"},{"body":"User 'JoshRosen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/2844","created":"2014-10-19T07:08:29.515+0000"},{"body":"Adding 1.1.0 as an affected version, since a user has observed this in 1.1.0, too; see SPARK-4133.","created":"2014-10-29T17:50:08.595+0000"},{"body":"Observed the same exception ( Spark 1.1.0 ). After changing broadcast factory to the HttpBroadcastFactory ( as suggested in the SPARK-4133 ), exception looks like:\njava.io.FileNotFoundException: http://10.8.0.22:44907/broadcast_0\n\nFull logs attached to issue: spark_ex.logs \n\nSeems that something wrong with the handling of the \"broadcast_0\" in the BlockManager:\nI think that something wrong with the handing of the \"broadcast_0\" in the BlockManager:\n14/11/04 17:20:39 INFO MemoryStore: Block broadcast_0 stored as values in memory (estimated size 1216.0 B, free 983.1 MB)\n14/11/04 17:20:39 DEBUG BlockManager: Put block broadcast_0 locally took 84 ms\n14/11/04 17:20:39 DEBUG BlockManager: Putting block broadcast_0 without replication took 86 ms\n14/11/04 17:20:39 DEBUG BlockManager: Getting local block broadcast_0\n14/11/04 17:20:39 DEBUG BlockManager: Level for block broadcast_0 is StorageLevel(true, true, false, true, 1)\n14/11/04 17:20:39 DEBUG BlockManager: Getting block broadcast_0 from memory\n14/11/04 17:20:39 INFO BlockManager: Found block broadcast_0 locally\n14/11/04 17:20:57 WARN BlockManager: Block broadcast_0 already exists on this machine; not re-adding it\n14/11/04 17:20:57 DEBUG BlockManager: Getting local block broadcast_0\n14/11/04 17:20:57 DEBUG BlockManager: Block broadcast_0 not registered locally\n14/11/04 17:20:57 DEBUG BlockManager: Getting remote block broadcast_0\n14/11/04 17:20:57 DEBUG BlockManagerMasterActor: [actor] received message GetLocations(broadcast_0) from Actor[akka://sparkDriver/temp/$l]\n14/11/04 17:20:57 DEBUG BlockManager: Block broadcast_0 not found\n14/11/04 17:20:57 DEBUG BlockManagerMasterActor: [actor] handled message (0.117676 ms) GetLocations(broadcast_0) from Actor[akka://sparkDriver/temp/$l]\n14/11/04 17:20:57 INFO HttpBroadcast: Started reading broadcast variable 0\n14/11/04 17:20:57 DEBUG HttpBroadcast: broadcast read server: http://10.8.0.22:44907 id: broadcast-0\n14/11/04 17:20:57 DEBUG HttpBroadcast: broadcast not using security\n14/11/04 17:20:57 DEBUG RecurringTimer: Callback for BlockGenerator called at time 1415114457200\n14/11/04 17:20:57 ERROR Executor: Exception in task 0.0 in stage 0.0 (TID 0)\njava.io.FileNotFoundException: http://10.8.0.22:44907/broadcast_0\n","created":"2014-11-04T16:10:09.158+0000"},{"body":"At this point I'm not aware of people still hitting this set of issues in newer releases, so per discussion with [~joshrosen], I'd like to close this. Please comment on this JIRA if you are having some variant of this issue in a newer version of Spark, and we'll continue to investigate.","created":"2015-01-21T22:28:49.223+0000"},{"body":"[~pwendell] [~joshrosen] We just hit this bug in one of our production jobs using Spark Streaming 1.4.1. Each task spawned by the streaming job fails down the road.\nThis jobs has been working fine for months, so I'm not clear on whether we can narrow down the conditions to reproduce it.\n\nHere's the exception:\n\n{code}\n[Stage 16049:(0 + 0) / 24][Stage 16056:(0 + 0) / 24][Stage 16058:(0 + 0) / 24]Exception in thread \"main\" org.apache.spark.SparkException: Job aborted due to stage failure: Task 23 in stage 17478.0 failed 6 times, most recent failure: Lost task 23.5 in stage 17478.0 (TID 172352, dnode-6.hdfs.private): java.io.IOException: PARSING_ERROR(2)\n\tat org.xerial.snappy.SnappyNative.throw_error(SnappyNative.java:84)\n\tat org.xerial.snappy.SnappyNative.uncompressedLength(Native Method)\n\tat org.xerial.snappy.Snappy.uncompressedLength(Snappy.java:594)\n\tat org.xerial.snappy.SnappyInputStream.hasNextChunk(SnappyInputStream.java:358)\n\tat org.xerial.snappy.SnappyInputStream.read(SnappyInputStream.java:387)\n\tat java.io.ObjectInputStream$PeekInputStream.peek(ObjectInputStream.java:2296)\n\tat java.io.ObjectInputStream$BlockDataInputStream.peek(ObjectInputStream.java:2589)\n\tat java.io.ObjectInputStream$BlockDataInputStream.peekByte(ObjectInputStream.java:2599)\n\tat java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1319)\n\tat java.io.ObjectInputStream.readObject(ObjectInputStream.java:371)\n\tat org.apache.spark.serializer.JavaDeserializationStream.readObject(JavaSerializer.scala:69)\n\tat org.apache.spark.serializer.DeserializationStream.readKey(Serializer.scala:169)\n\tat org.apache.spark.serializer.DeserializationStream$$anon$2.getNext(Serializer.scala:200)\n\tat org.apache.spark.serializer.DeserializationStream$$anon$2.getNext(Serializer.scala:197)\n\tat org.apache.spark.util.NextIterator.hasNext(NextIterator.scala:71)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:32)\n\tat scala.collection.Iterator$$anon$13.hasNext(Iterator.scala:371)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:32)\n\tat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:39)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.insertAll(ExternalAppendOnlyMap.scala:127)\n\tat org.apache.spark.Aggregator.combineCombinersByKey(Aggregator.scala:91)\n\tat org.apache.spark.shuffle.hash.HashShuffleReader.read(HashShuffleReader.scala:44)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:90)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:277)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:244)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:35)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:277)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:244)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:70)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:70)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n\nDriver stacktrace:\n\tat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1273)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1264)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1263)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n\tat org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1263)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:730)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:730)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:730)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1457)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1418)\n\tat org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\n{code}\n\n(Murphy's law: Bugs reappear the moment you close them)","created":"2016-01-25T08:46:01.785+0000"}],"conversations":[{"body":"TorrentBroadcast deserialization sometimes fails with decompression errors, which are most likely caused by stream-corruption exceptions. For example, this can manifest itself as a Snappy PARSING_ERROR when deserializing a broadcasted task:\n\n{code}\n14/10/14 17:20:55.016 DEBUG BlockManager: Getting local block broadcast_8\n14/10/14 17:20:55.016 DEBUG BlockManager: Block broadcast_8 not registered locally\n14/10/14 17:20:55.016 INFO TorrentBroadcast: Started reading broadcast variable 8\n14/10/14 17:20:55.017 INFO TorrentBroadcast: Reading broadcast variable 8 took 5.3433E-5 s\n14/10/14 17:20:55.017 ERROR Executor: Exception in task 2.0 in stage 8.0 (TID 18)\njava.io.IOException: PARSING_ERROR(2)\n\tat org.xerial.snappy.SnappyNative.throw_error(SnappyNative.java:84)\n\tat org.xerial.snappy.SnappyNative.uncompressedLength(Native Method)\n\tat org.xerial.snappy.Snappy.uncompressedLength(Snappy.java:594)\n\tat org.xerial.snappy.SnappyInputStream.readFully(SnappyInputStream.java:125)\n\tat org.xerial.snappy.SnappyInputStream.readHeader(SnappyInputStream.java:88)\n\tat org.xerial.snappy.SnappyInputStream.(SnappyInputStream.java:58)\n\tat org.apache.spark.io.SnappyCompressionCodec.compressedInputStream(CompressionCodec.scala:128)\n\tat org.apache.spark.broadcast.TorrentBroadcast$.unBlockifyObject(TorrentBroadcast.scala:216)\n\tat org.apache.spark.broadcast.TorrentBroadcast.readObject(TorrentBroadcast.scala:170)\n\tat sun.reflect.GeneratedMethodAccessor92.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:606)\n\tat java.io.ObjectStreamClass.invokeReadObject(ObjectStreamClass.java:1017)\n\tat java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:1893)\n\tat java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:1798)\n\tat java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1350)\n\tat java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:1990)\n\tat java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:1915)\n\tat java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:1798)\n\tat java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1350)\n\tat java.io.ObjectInputStream.readObject(ObjectInputStream.java:370)\n\tat org.apache.spark.serializer.JavaDeserializationStream.readObject(JavaSerializer.scala:62)\n\tat org.apache.spark.serializer.JavaSerializerInstance.deserialize(JavaSerializer.scala:87)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:164)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}\n\nSPARK-3630 is an umbrella ticket for investigating all causes of these Kryo and Snappy deserialization errors. This ticket is for a more narrowly-focused exploration of the TorrentBroadcast version of these errors, since the similar errors that we've seen in sort-based shuffle seem to be explained by a different cause (see SPARK-3948).","from":"reporter","subject":"Possible stream-corruption issues in TorrentBroadcast"},{"body":"Digging into this stacktrace in more detail:\n\nSnappy-java's {{SnappyOutputStream}} writes its own 8-byte header at the beginning of the serialized output. This header consists of an 8-byte magic value followed by two 4-byte version numbers. This header is distinct from Snappy's own 6-byte magic number / header.\n\n{{org.xerial.snappy.SnappyInputStream.readHeader}} is implemented like this (in Snappy-Java 1.1.1.3):\n\n{code}\nprotected void readHeader() throws IOException {\n byte[] header = new byte[SnappyCodec.headerSize()];\n int readBytes = 0;\n while (readBytes < header.length) {\n int ret = in.read(header, readBytes, header.length - readBytes);\n if (ret == -1)\n break;\n readBytes += ret;\n }\n\n // Quick test of the header \n if (readBytes < header.length || header[0] != SnappyCodec.MAGIC_HEADER[0]) {\n // do the default uncompression\n readFully(header, readBytes);\n return;\n }\n\n SnappyCodec codec = SnappyCodec.readHeader(new ByteArrayInputStream(header));\n if (codec.isValidMagicHeader()) {\n // The input data is compressed by SnappyOutputStream\n if (codec.version < SnappyCodec.MINIMUM_COMPATIBLE_VERSION) {\n throw new IOException(String.format(\n \"compressed with imcompatible codec version %d. At least version %d is required\",\n codec.version, SnappyCodec.MINIMUM_COMPATIBLE_VERSION));\n }\n }\n else {\n // (probably) compressed by Snappy.compress(byte[])\n readFully(header, readBytes);\n return;\n }\n }\n{code}\n\nIt starts by attempting to read the 8-byte header. The first {{while}} loop exits when we've either read 8 bytes of header data or if the input stream was closed before it could read a complete header. The following code checks whether the header is unexpectedly short or whether it doesn't match the snappy-java magic header. In our case, we end up taking this branch and calling {{readFully(header, readBytes)}} in order to perform the default Snappy decompression This is the wrong branch to take (since our data was compressed with a SnappyOutputStream), leading to the PARSING_ERROR.\n\nBased on this, I think that the input data to the SnappyInputStream is somehow being corrupted. It's not obvious whether this corruption is causing the input data to be too short or whether the start of the stream has the wrong contents. I'll keep digging and look into adding some size-checking assertions throughout our code.","from":"developer"},{"body":"Removing 1.1.0 as an affected version for now, since the stacktrace that I posted here was from a recent build of master (1.2). If anyone can reproduce this in branch-1.1 or Spark 1.1.0, please let me know.","from":"developer"},{"body":"I think that I can safely rule out problems in TorrentBroadcast's (de-)blockification code: I used ScalaCheck to write some tests to ensure that blockifyObject and unblockifyObject are inverses, plus a similar test for ByteArrayChunkOutputStream: https://github.com/JoshRosen/spark/commit/413be7f6a8d4eb14c69c7db87e2564ed4d776c42?diff=unified","from":"developer"},{"body":"Hi Josh, have you tried other compression like LZO to narrow down the problem?","from":"developer"},{"body":"Hi [~jerryshao],\n\nI don't have a reliable reproduction for this issue yet, so I haven't tried switching compression schemes or broadcast implementations. I'm working with the user who provided this stack trace to see if we can get more logs to provide additional context.","from":"developer"},{"body":"[~davies] ran across this exception while testing a pull request that modifies TorrentBroadcast: https://github.com/apache/spark/pull/2681#issuecomment-59120483\n\nThat PR's reproductioncould be a valuable debugging clue for this issue.","from":"developer"},{"body":"User 'JoshRosen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/2844","from":"developer"},{"body":"Adding 1.1.0 as an affected version, since a user has observed this in 1.1.0, too; see SPARK-4133.","from":"developer"},{"body":"Observed the same exception ( Spark 1.1.0 ). After changing broadcast factory to the HttpBroadcastFactory ( as suggested in the SPARK-4133 ), exception looks like:\njava.io.FileNotFoundException: http://10.8.0.22:44907/broadcast_0\n\nFull logs attached to issue: spark_ex.logs \n\nSeems that something wrong with the handling of the \"broadcast_0\" in the BlockManager:\nI think that something wrong with the handing of the \"broadcast_0\" in the BlockManager:\n14/11/04 17:20:39 INFO MemoryStore: Block broadcast_0 stored as values in memory (estimated size 1216.0 B, free 983.1 MB)\n14/11/04 17:20:39 DEBUG BlockManager: Put block broadcast_0 locally took 84 ms\n14/11/04 17:20:39 DEBUG BlockManager: Putting block broadcast_0 without replication took 86 ms\n14/11/04 17:20:39 DEBUG BlockManager: Getting local block broadcast_0\n14/11/04 17:20:39 DEBUG BlockManager: Level for block broadcast_0 is StorageLevel(true, true, false, true, 1)\n14/11/04 17:20:39 DEBUG BlockManager: Getting block broadcast_0 from memory\n14/11/04 17:20:39 INFO BlockManager: Found block broadcast_0 locally\n14/11/04 17:20:57 WARN BlockManager: Block broadcast_0 already exists on this machine; not re-adding it\n14/11/04 17:20:57 DEBUG BlockManager: Getting local block broadcast_0\n14/11/04 17:20:57 DEBUG BlockManager: Block broadcast_0 not registered locally\n14/11/04 17:20:57 DEBUG BlockManager: Getting remote block broadcast_0\n14/11/04 17:20:57 DEBUG BlockManagerMasterActor: [actor] received message GetLocations(broadcast_0) from Actor[akka://sparkDriver/temp/$l]\n14/11/04 17:20:57 DEBUG BlockManager: Block broadcast_0 not found\n14/11/04 17:20:57 DEBUG BlockManagerMasterActor: [actor] handled message (0.117676 ms) GetLocations(broadcast_0) from Actor[akka://sparkDriver/temp/$l]\n14/11/04 17:20:57 INFO HttpBroadcast: Started reading broadcast variable 0\n14/11/04 17:20:57 DEBUG HttpBroadcast: broadcast read server: http://10.8.0.22:44907 id: broadcast-0\n14/11/04 17:20:57 DEBUG HttpBroadcast: broadcast not using security\n14/11/04 17:20:57 DEBUG RecurringTimer: Callback for BlockGenerator called at time 1415114457200\n14/11/04 17:20:57 ERROR Executor: Exception in task 0.0 in stage 0.0 (TID 0)\njava.io.FileNotFoundException: http://10.8.0.22:44907/broadcast_0\n","from":"developer"},{"body":"At this point I'm not aware of people still hitting this set of issues in newer releases, so per discussion with [~joshrosen], I'd like to close this. Please comment on this JIRA if you are having some variant of this issue in a newer version of Spark, and we'll continue to investigate.","from":"developer"},{"body":"[~pwendell] [~joshrosen] We just hit this bug in one of our production jobs using Spark Streaming 1.4.1. Each task spawned by the streaming job fails down the road.\nThis jobs has been working fine for months, so I'm not clear on whether we can narrow down the conditions to reproduce it.\n\nHere's the exception:\n\n{code}\n[Stage 16049:(0 + 0) / 24][Stage 16056:(0 + 0) / 24][Stage 16058:(0 + 0) / 24]Exception in thread \"main\" org.apache.spark.SparkException: Job aborted due to stage failure: Task 23 in stage 17478.0 failed 6 times, most recent failure: Lost task 23.5 in stage 17478.0 (TID 172352, dnode-6.hdfs.private): java.io.IOException: PARSING_ERROR(2)\n\tat org.xerial.snappy.SnappyNative.throw_error(SnappyNative.java:84)\n\tat org.xerial.snappy.SnappyNative.uncompressedLength(Native Method)\n\tat org.xerial.snappy.Snappy.uncompressedLength(Snappy.java:594)\n\tat org.xerial.snappy.SnappyInputStream.hasNextChunk(SnappyInputStream.java:358)\n\tat org.xerial.snappy.SnappyInputStream.read(SnappyInputStream.java:387)\n\tat java.io.ObjectInputStream$PeekInputStream.peek(ObjectInputStream.java:2296)\n\tat java.io.ObjectInputStream$BlockDataInputStream.peek(ObjectInputStream.java:2589)\n\tat java.io.ObjectInputStream$BlockDataInputStream.peekByte(ObjectInputStream.java:2599)\n\tat java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1319)\n\tat java.io.ObjectInputStream.readObject(ObjectInputStream.java:371)\n\tat org.apache.spark.serializer.JavaDeserializationStream.readObject(JavaSerializer.scala:69)\n\tat org.apache.spark.serializer.DeserializationStream.readKey(Serializer.scala:169)\n\tat org.apache.spark.serializer.DeserializationStream$$anon$2.getNext(Serializer.scala:200)\n\tat org.apache.spark.serializer.DeserializationStream$$anon$2.getNext(Serializer.scala:197)\n\tat org.apache.spark.util.NextIterator.hasNext(NextIterator.scala:71)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:32)\n\tat scala.collection.Iterator$$anon$13.hasNext(Iterator.scala:371)\n\tat org.apache.spark.util.CompletionIterator.hasNext(CompletionIterator.scala:32)\n\tat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:39)\n\tat org.apache.spark.util.collection.ExternalAppendOnlyMap.insertAll(ExternalAppendOnlyMap.scala:127)\n\tat org.apache.spark.Aggregator.combineCombinersByKey(Aggregator.scala:91)\n\tat org.apache.spark.shuffle.hash.HashShuffleReader.read(HashShuffleReader.scala:44)\n\tat org.apache.spark.rdd.ShuffledRDD.compute(ShuffledRDD.scala:90)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:277)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:244)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:35)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:277)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:244)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:70)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:41)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:70)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n\nDriver stacktrace:\n\tat org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$failJobAndIndependentStages(DAGScheduler.scala:1273)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1264)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:1263)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n\tat org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:1263)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:730)\n\tat org.apache.spark.scheduler.DAGScheduler$$anonfun$handleTaskSetFailed$1.apply(DAGScheduler.scala:730)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.DAGScheduler.handleTaskSetFailed(DAGScheduler.scala:730)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1457)\n\tat org.apache.spark.scheduler.DAGSchedulerEventProcessLoop.onReceive(DAGScheduler.scala:1418)\n\tat org.apache.spark.util.EventLoop$$anon$1.run(EventLoop.scala:48)\n{code}\n\n(Murphy's law: Bugs reappear the moment you close them)","from":"developer"}],"created":"2014-10-15T20:54:21.000+0000","description":"TorrentBroadcast deserialization sometimes fails with decompression errors, which are most likely caused by stream-corruption exceptions. For example, this can manifest itself as a Snappy PARSING_ERROR when deserializing a broadcasted task:\n\n{code}\n14/10/14 17:20:55.016 DEBUG BlockManager: Getting local block broadcast_8\n14/10/14 17:20:55.016 DEBUG BlockManager: Block broadcast_8 not registered locally\n14/10/14 17:20:55.016 INFO TorrentBroadcast: Started reading broadcast variable 8\n14/10/14 17:20:55.017 INFO TorrentBroadcast: Reading broadcast variable 8 took 5.3433E-5 s\n14/10/14 17:20:55.017 ERROR Executor: Exception in task 2.0 in stage 8.0 (TID 18)\njava.io.IOException: PARSING_ERROR(2)\n\tat org.xerial.snappy.SnappyNative.throw_error(SnappyNative.java:84)\n\tat org.xerial.snappy.SnappyNative.uncompressedLength(Native Method)\n\tat org.xerial.snappy.Snappy.uncompressedLength(Snappy.java:594)\n\tat org.xerial.snappy.SnappyInputStream.readFully(SnappyInputStream.java:125)\n\tat org.xerial.snappy.SnappyInputStream.readHeader(SnappyInputStream.java:88)\n\tat org.xerial.snappy.SnappyInputStream.(SnappyInputStream.java:58)\n\tat org.apache.spark.io.SnappyCompressionCodec.compressedInputStream(CompressionCodec.scala:128)\n\tat org.apache.spark.broadcast.TorrentBroadcast$.unBlockifyObject(TorrentBroadcast.scala:216)\n\tat org.apache.spark.broadcast.TorrentBroadcast.readObject(TorrentBroadcast.scala:170)\n\tat sun.reflect.GeneratedMethodAccessor92.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:606)\n\tat java.io.ObjectStreamClass.invokeReadObject(ObjectStreamClass.java:1017)\n\tat java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:1893)\n\tat java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:1798)\n\tat java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1350)\n\tat java.io.ObjectInputStream.defaultReadFields(ObjectInputStream.java:1990)\n\tat java.io.ObjectInputStream.readSerialData(ObjectInputStream.java:1915)\n\tat java.io.ObjectInputStream.readOrdinaryObject(ObjectInputStream.java:1798)\n\tat java.io.ObjectInputStream.readObject0(ObjectInputStream.java:1350)\n\tat java.io.ObjectInputStream.readObject(ObjectInputStream.java:370)\n\tat org.apache.spark.serializer.JavaDeserializationStream.readObject(JavaSerializer.scala:62)\n\tat org.apache.spark.serializer.JavaSerializerInstance.deserialize(JavaSerializer.scala:87)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:164)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}\n\nSPARK-3630 is an umbrella ticket for investigating all causes of these Kryo and Snappy deserialization errors. This ticket is for a more narrowly-focused exploration of the TorrentBroadcast version of these errors, since the similar errors that we've seen in sort-based shuffle seem to be explained by a different cause (see SPARK-3948).","issue_id":"12748414","key":"SPARK-3958","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-01-21T22:28:49.000+0000","role":"fixed_distractor","summary":"Possible stream-corruption issues in TorrentBroadcast"} {"case_id":"12748544","cluster":"DISTRACTOR-SPARK-3967","comments":[{"body":"Ensure that the temporary file which the jar file is fetched in is located in the same directory than the target jar file\n","created":"2014-10-16T09:23:16.164+0000"},{"body":"After investigating, it turns out that the problem is when the executor fetches a jar file: the jar is downloaded in a temporary file, always in /d1/yarn/local/nm-local-dir (first directory of yarn.nodemanager.local-dirs), and then moved in one of the directories of yarn.nodemanager.local-dirs:\n--> if it is the same than the temporary file (i.e. /d1/yarn/local/nm-local-dir), then the application continues normally\n--> if it is another one (i.e. /d2/yarn/local/nm-local-dir, /d3/yarn/local/nm-local-dir,...), it fails with the following error:\n14/10/10 14:33:51 ERROR executor.Executor: Exception in task 0.0 in stage 1.0 (TID 0)\njava.io.FileNotFoundException: ./logReader-1.0.10.jar (Permission denied)\n at java.io.FileOutputStream.open(Native Method)\n at java.io.FileOutputStream.(FileOutputStream.java:221)\n at com.google.common.io.Files$FileByteSink.openStream(Files.java:223)\n at com.google.common.io.Files$FileByteSink.openStream(Files.java:211)\n at com.google.common.io.ByteSource.copyTo(ByteSource.java:203)\n at com.google.common.io.Files.copy(Files.java:436)\n at com.google.common.io.Files.move(Files.java:651)\n at org.apache.spark.util.Utils$.fetchFile(Utils.scala:440)\n at org.apache.spark.executor.Executor$$anonfun$org$apache$spark$executor$Executor$$updateDependencies$6.apply(Executor.scala:325)\n at org.apache.spark.executor.Executor$$anonfun$org$apache$spark$executor$Executor$$updateDependencies$6.apply(Executor.scala:323)\n at scala.collection.TraversableLike$WithFilter$$anonfun$foreach$1.apply(TraversableLike.scala:772)\n at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n at scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:226)\n at scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:39)\n at scala.collection.mutable.HashMap.foreach(HashMap.scala:98)\n at scala.collection.TraversableLike$WithFilter.foreach(TraversableLike.scala:771)\n at org.apache.spark.executor.Executor.org$apache$spark$executor$Executor$$updateDependencies(Executor.scala:323)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:158)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n\nI have no idea why the move fails when the source and target files are not on the same partition (it is no more atomic, but it should succeed anyway), for the moment I have worked around the problem with the attached patch (i.e. I ensure that the temp file and the moved file are always on the same partition).","created":"2014-10-16T09:24:04.199+0000"},{"body":"I've been debugging this issue as well and I think I've found an issue in {{org.apache.spark.util.Utils}} that is contributing to / causing the problem:\n\n{{Files.move}} on [line 390|https://github.com/apache/spark/blob/v1.1.0/core/src/main/scala/org/apache/spark/util/Utils.scala#L390] is called even if {{targetFile}} exists and {{tempFile}} and {{targetFile}} are equal.\n\nThe check on [line 379|https://github.com/apache/spark/blob/v1.1.0/core/src/main/scala/org/apache/spark/util/Utils.scala#L379] seems to imply the desire to skip a redundant overwrite if the file is already there and has the contents that it should have.\n\nGating the {{Files.move}} call on a further {{if (!targetFile.exists)}} fixes the issue for me; attached is a patch of the change.\n\nIn practice all of my executors that hit this code path are finding every dependency JAR to already exist and be exactly equal to what they need it to be, meaning they were all needlessly overwriting all of their dependency JARs, and now are all basically no-op-ing in {{Utils.fetchFile}}; I've not determined who/what is putting the JARs there, why the issue only crops up in {{yarn-cluster}} mode (or {{--master yarn --deploy-mode cluster}}), etc., but it seems like either way this patch is probably desirable.\n","created":"2014-10-17T22:28:16.146+0000"},{"body":"Don't redundantly copy executor dependency files in {{Utils.fetchFile}}.","created":"2014-10-17T22:29:36.791+0000"},{"body":"You guys should make PRs for these. I am also not sure if it's so necessary to download the file into a temp directory and move it... it may cause a copy instead of rename, and in fact does here, and so is not like the file appears in the target dir atomically anyway. I'm not sure the code here cleans up the partially downloaded file in case of error and that could leave a broken file in the target dir instead of just a temp dir.\n\nThe change to not copy the file when identical looks sound; I bet you can avoid checking if it exists twice.","created":"2014-10-18T00:32:45.246+0000"},{"body":"User 'ryan-williams' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/2848","created":"2014-10-19T23:11:43.268+0000"},{"body":"User 'preaudc' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/2855","created":"2014-10-20T10:14:34.179+0000"},{"body":"Hi Ryan,\nThanks for your help. You should probably add the same test on the {{Files.move}} on [line 437|https://github.com/apache/spark/blob/v1.1.0/core/src/main/scala/org/apache/spark/util/Utils.scala#L437]","created":"2014-10-20T10:22:00.911+0000"},{"body":"Cool, I'll add it there as well [~preaudc], and take further discussion to [the PR|https://github.com/apache/spark/pull/2848] unless you prefer having it here","created":"2014-10-20T15:40:04.580+0000"},{"body":"That's fine, thanks!","created":"2014-10-20T15:59:54.136+0000"},{"body":"This issue also hapens to me.but i want to know why this issue don't happens to the *yarn-client* mode","created":"2014-10-21T10:09:10.578+0000"},{"body":"I think you this pull request also can be referenced:\nhttps://github.com/apache/spark/pull/1616","created":"2014-10-21T11:34:42.266+0000"},{"body":"To give a quick update on this: I've merged [~preaudc]'s patch into {{master}} and {{branch-1.1}} (and will backport it into {{branch-1.2}} after the 1.2.0 release passes).\n\nI'd also like to include [~rdub]'s patch, too, since it's a pretty good refactoring of this code. That patch is getting a lot closer to merging; I'm going to try to loop back soon to provide more review feedback.","created":"2014-12-08T19:58:39.989+0000"},{"body":"I've merged [~rdub]'s patch (SPARK-4896) into {{master}}, {{branch-1.1}}, and {{branch-1.2}} and have backported the other patch to {{branch-1.2}}.\n\nIt would be great if folks could confirm whether these fixes have resolved this issue, or whether there's still more work to be done.","created":"2014-12-19T23:58:50.017+0000"}],"conversations":[{"body":"Spark applications fail from time to time in yarn-cluster mode (but not in yarn-client mode) when yarn.nodemanager.local-dirs (Hadoop YARN config) is set to a comma-separated list of directories which are located on different disks/partitions.\n\nSteps to reproduce:\n1. Set yarn.nodemanager.local-dirs (in yarn-site.xml) to a list of directories located on different partitions (the more you set, the more likely it will be to reproduce the bug):\n(...)\n\n yarn.nodemanager.local-dirs\n file:/d1/yarn/local/nm-local-dir,file:/d2/yarn/local/nm-local-dir,file:/d3/yarn/local/nm-local-dir,file:/d4/yarn/local/nm-local-dir,file:/d5/yarn/local/nm-local-dir,file:/d6/yarn/local/nm-local-dir,file:/d7/yarn/local/nm-local-dir\n\n(...)\n2. Launch (several times) an application in yarn-cluster mode, it will fail (apparently randomly) from time to time","from":"reporter","subject":"Spark applications fail in yarn-cluster mode when the directories configured in yarn.nodemanager.local-dirs are located on different disks/partitions"},{"body":"Ensure that the temporary file which the jar file is fetched in is located in the same directory than the target jar file\n","from":"developer"},{"body":"After investigating, it turns out that the problem is when the executor fetches a jar file: the jar is downloaded in a temporary file, always in /d1/yarn/local/nm-local-dir (first directory of yarn.nodemanager.local-dirs), and then moved in one of the directories of yarn.nodemanager.local-dirs:\n--> if it is the same than the temporary file (i.e. /d1/yarn/local/nm-local-dir), then the application continues normally\n--> if it is another one (i.e. /d2/yarn/local/nm-local-dir, /d3/yarn/local/nm-local-dir,...), it fails with the following error:\n14/10/10 14:33:51 ERROR executor.Executor: Exception in task 0.0 in stage 1.0 (TID 0)\njava.io.FileNotFoundException: ./logReader-1.0.10.jar (Permission denied)\n at java.io.FileOutputStream.open(Native Method)\n at java.io.FileOutputStream.(FileOutputStream.java:221)\n at com.google.common.io.Files$FileByteSink.openStream(Files.java:223)\n at com.google.common.io.Files$FileByteSink.openStream(Files.java:211)\n at com.google.common.io.ByteSource.copyTo(ByteSource.java:203)\n at com.google.common.io.Files.copy(Files.java:436)\n at com.google.common.io.Files.move(Files.java:651)\n at org.apache.spark.util.Utils$.fetchFile(Utils.scala:440)\n at org.apache.spark.executor.Executor$$anonfun$org$apache$spark$executor$Executor$$updateDependencies$6.apply(Executor.scala:325)\n at org.apache.spark.executor.Executor$$anonfun$org$apache$spark$executor$Executor$$updateDependencies$6.apply(Executor.scala:323)\n at scala.collection.TraversableLike$WithFilter$$anonfun$foreach$1.apply(TraversableLike.scala:772)\n at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n at scala.collection.mutable.HashMap$$anonfun$foreach$1.apply(HashMap.scala:98)\n at scala.collection.mutable.HashTable$class.foreachEntry(HashTable.scala:226)\n at scala.collection.mutable.HashMap.foreachEntry(HashMap.scala:39)\n at scala.collection.mutable.HashMap.foreach(HashMap.scala:98)\n at scala.collection.TraversableLike$WithFilter.foreach(TraversableLike.scala:771)\n at org.apache.spark.executor.Executor.org$apache$spark$executor$Executor$$updateDependencies(Executor.scala:323)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:158)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n\nI have no idea why the move fails when the source and target files are not on the same partition (it is no more atomic, but it should succeed anyway), for the moment I have worked around the problem with the attached patch (i.e. I ensure that the temp file and the moved file are always on the same partition).","from":"developer"},{"body":"I've been debugging this issue as well and I think I've found an issue in {{org.apache.spark.util.Utils}} that is contributing to / causing the problem:\n\n{{Files.move}} on [line 390|https://github.com/apache/spark/blob/v1.1.0/core/src/main/scala/org/apache/spark/util/Utils.scala#L390] is called even if {{targetFile}} exists and {{tempFile}} and {{targetFile}} are equal.\n\nThe check on [line 379|https://github.com/apache/spark/blob/v1.1.0/core/src/main/scala/org/apache/spark/util/Utils.scala#L379] seems to imply the desire to skip a redundant overwrite if the file is already there and has the contents that it should have.\n\nGating the {{Files.move}} call on a further {{if (!targetFile.exists)}} fixes the issue for me; attached is a patch of the change.\n\nIn practice all of my executors that hit this code path are finding every dependency JAR to already exist and be exactly equal to what they need it to be, meaning they were all needlessly overwriting all of their dependency JARs, and now are all basically no-op-ing in {{Utils.fetchFile}}; I've not determined who/what is putting the JARs there, why the issue only crops up in {{yarn-cluster}} mode (or {{--master yarn --deploy-mode cluster}}), etc., but it seems like either way this patch is probably desirable.\n","from":"developer"},{"body":"Don't redundantly copy executor dependency files in {{Utils.fetchFile}}.","from":"developer"},{"body":"You guys should make PRs for these. I am also not sure if it's so necessary to download the file into a temp directory and move it... it may cause a copy instead of rename, and in fact does here, and so is not like the file appears in the target dir atomically anyway. I'm not sure the code here cleans up the partially downloaded file in case of error and that could leave a broken file in the target dir instead of just a temp dir.\n\nThe change to not copy the file when identical looks sound; I bet you can avoid checking if it exists twice.","from":"developer"},{"body":"User 'ryan-williams' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/2848","from":"developer"},{"body":"User 'preaudc' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/2855","from":"developer"},{"body":"Hi Ryan,\nThanks for your help. You should probably add the same test on the {{Files.move}} on [line 437|https://github.com/apache/spark/blob/v1.1.0/core/src/main/scala/org/apache/spark/util/Utils.scala#L437]","from":"developer"},{"body":"Cool, I'll add it there as well [~preaudc], and take further discussion to [the PR|https://github.com/apache/spark/pull/2848] unless you prefer having it here","from":"developer"},{"body":"That's fine, thanks!","from":"developer"},{"body":"This issue also hapens to me.but i want to know why this issue don't happens to the *yarn-client* mode","from":"developer"},{"body":"I think you this pull request also can be referenced:\nhttps://github.com/apache/spark/pull/1616","from":"developer"},{"body":"To give a quick update on this: I've merged [~preaudc]'s patch into {{master}} and {{branch-1.1}} (and will backport it into {{branch-1.2}} after the 1.2.0 release passes).\n\nI'd also like to include [~rdub]'s patch, too, since it's a pretty good refactoring of this code. That patch is getting a lot closer to merging; I'm going to try to loop back soon to provide more review feedback.","from":"developer"},{"body":"I've merged [~rdub]'s patch (SPARK-4896) into {{master}}, {{branch-1.1}}, and {{branch-1.2}} and have backported the other patch to {{branch-1.2}}.\n\nIt would be great if folks could confirm whether these fixes have resolved this issue, or whether there's still more work to be done.","from":"developer"}],"created":"2014-10-16T09:16:25.000+0000","description":"Spark applications fail from time to time in yarn-cluster mode (but not in yarn-client mode) when yarn.nodemanager.local-dirs (Hadoop YARN config) is set to a comma-separated list of directories which are located on different disks/partitions.\n\nSteps to reproduce:\n1. Set yarn.nodemanager.local-dirs (in yarn-site.xml) to a list of directories located on different partitions (the more you set, the more likely it will be to reproduce the bug):\n(...)\n\n yarn.nodemanager.local-dirs\n file:/d1/yarn/local/nm-local-dir,file:/d2/yarn/local/nm-local-dir,file:/d3/yarn/local/nm-local-dir,file:/d4/yarn/local/nm-local-dir,file:/d5/yarn/local/nm-local-dir,file:/d6/yarn/local/nm-local-dir,file:/d7/yarn/local/nm-local-dir\n\n(...)\n2. Launch (several times) an application in yarn-cluster mode, it will fail (apparently randomly) from time to time","issue_id":"12748544","key":"SPARK-3967","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-05-15T13:35:50.000+0000","role":"fixed_distractor","summary":"Spark applications fail in yarn-cluster mode when the directories configured in yarn.nodemanager.local-dirs are located on different disks/partitions"} {"case_id":"13470364","cluster":"DISTRACTOR-SPARK-39696","comments":[{"body":"User 'LuciferYang' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/37206","created":"2022-07-16T18:29:41.500+0000"},{"body":"User 'LuciferYang' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/37206","created":"2022-07-16T18:29:50.406+0000"},{"body":"Can you give more details on how this can be reproduced [~smcmullan] ?\r\nThanks","created":"2022-07-26T02:28:24.424+0000"},{"body":"It happens randomly. Run a job and all good and then again and the exception may, or may not, appear.","created":"2022-09-10T07:59:45.701+0000"},{"body":"same for me. Is it possible to extend any timeout?","created":"2022-09-16T05:08:51.607+0000"},{"body":"Created a test case that consistently reproduces the issue for when using Scala 2.13: [https://github.com/apache/spark/pull/40663]","created":"2023-04-04T14:51:04.434+0000"},{"body":"This is resolved via https://github.com/apache/spark/pull/40663","created":"2023-04-07T04:02:51.636+0000"}],"conversations":[{"body":"{noformat}\r\n2022-06-21 18:17:49.289Z ERROR [executor-heartbeater] org.apache.spark.util.Utils - Uncaught exception in thread executor-heartbeater\r\njava.util.ConcurrentModificationException: mutation occurred during iteration\r\n at scala.collection.mutable.MutationTracker$.checkMutations(MutationTracker.scala:43) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.mutable.CheckedIndexedSeqView$CheckedIterator.hasNext(CheckedIndexedSeqView.scala:47) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOnceOps.copyToArray(IterableOnce.scala:873) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOnceOps.copyToArray$(IterableOnce.scala:869) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.AbstractIterator.copyToArray(Iterator.scala:1293) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOnceOps.copyToArray(IterableOnce.scala:852) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOnceOps.copyToArray$(IterableOnce.scala:852) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.AbstractIterator.copyToArray(Iterator.scala:1293) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.immutable.VectorStatics$.append1IfSpace(Vector.scala:1959) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.immutable.Vector1.appendedAll0(Vector.scala:425) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.immutable.Vector.appendedAll(Vector.scala:203) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.immutable.Vector.appendedAll(Vector.scala:113) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.SeqOps.concat(Seq.scala:187) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.SeqOps.concat$(Seq.scala:187) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.AbstractSeq.concat(Seq.scala:1161) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOps.$plus$plus(Iterable.scala:726) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOps.$plus$plus$(Iterable.scala:726) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.AbstractIterable.$plus$plus(Iterable.scala:926) ~[scala-library-2.13.8.jar:?]\r\n at org.apache.spark.executor.TaskMetrics.accumulators(TaskMetrics.scala:261) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at org.apache.spark.executor.Executor.$anonfun$reportHeartBeat$1(Executor.scala:1042) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at scala.collection.IterableOnceOps.foreach(IterableOnce.scala:563) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOnceOps.foreach$(IterableOnce.scala:561) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.AbstractIterable.foreach(Iterable.scala:926) ~[scala-library-2.13.8.jar:?]\r\n at org.apache.spark.executor.Executor.reportHeartBeat(Executor.scala:1036) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at org.apache.spark.executor.Executor.$anonfun$heartbeater$1(Executor.scala:238) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at scala.runtime.java8.JFunction0$mcV$sp.apply(JFunction0$mcV$sp.scala:18) ~[scala-library-2.13.8.jar:?]\r\n at org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:2066) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at org.apache.spark.Heartbeater$$anon$1.run(Heartbeater.scala:46) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) ~[?:?]\r\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:305) ~[?:?]\r\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:305) ~[?:?]\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) ~[?:?]\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) ~[?:?]\r\n at java.lang.Thread.run(Thread.java:833) ~[?:?] {noformat}","from":"reporter","subject":"Uncaught exception in thread executor-heartbeater java.util.ConcurrentModificationException: mutation occurred during iteration"},{"body":"User 'LuciferYang' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/37206","from":"developer"},{"body":"User 'LuciferYang' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/37206","from":"developer"},{"body":"Can you give more details on how this can be reproduced [~smcmullan] ?\r\nThanks","from":"developer"},{"body":"It happens randomly. Run a job and all good and then again and the exception may, or may not, appear.","from":"developer"},{"body":"same for me. Is it possible to extend any timeout?","from":"developer"},{"body":"Created a test case that consistently reproduces the issue for when using Scala 2.13: [https://github.com/apache/spark/pull/40663]","from":"developer"},{"body":"This is resolved via https://github.com/apache/spark/pull/40663","from":"developer"}],"created":"2022-07-06T14:07:08.000+0000","description":"{noformat}\r\n2022-06-21 18:17:49.289Z ERROR [executor-heartbeater] org.apache.spark.util.Utils - Uncaught exception in thread executor-heartbeater\r\njava.util.ConcurrentModificationException: mutation occurred during iteration\r\n at scala.collection.mutable.MutationTracker$.checkMutations(MutationTracker.scala:43) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.mutable.CheckedIndexedSeqView$CheckedIterator.hasNext(CheckedIndexedSeqView.scala:47) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOnceOps.copyToArray(IterableOnce.scala:873) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOnceOps.copyToArray$(IterableOnce.scala:869) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.AbstractIterator.copyToArray(Iterator.scala:1293) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOnceOps.copyToArray(IterableOnce.scala:852) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOnceOps.copyToArray$(IterableOnce.scala:852) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.AbstractIterator.copyToArray(Iterator.scala:1293) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.immutable.VectorStatics$.append1IfSpace(Vector.scala:1959) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.immutable.Vector1.appendedAll0(Vector.scala:425) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.immutable.Vector.appendedAll(Vector.scala:203) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.immutable.Vector.appendedAll(Vector.scala:113) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.SeqOps.concat(Seq.scala:187) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.SeqOps.concat$(Seq.scala:187) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.AbstractSeq.concat(Seq.scala:1161) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOps.$plus$plus(Iterable.scala:726) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOps.$plus$plus$(Iterable.scala:726) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.AbstractIterable.$plus$plus(Iterable.scala:926) ~[scala-library-2.13.8.jar:?]\r\n at org.apache.spark.executor.TaskMetrics.accumulators(TaskMetrics.scala:261) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at org.apache.spark.executor.Executor.$anonfun$reportHeartBeat$1(Executor.scala:1042) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at scala.collection.IterableOnceOps.foreach(IterableOnce.scala:563) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.IterableOnceOps.foreach$(IterableOnce.scala:561) ~[scala-library-2.13.8.jar:?]\r\n at scala.collection.AbstractIterable.foreach(Iterable.scala:926) ~[scala-library-2.13.8.jar:?]\r\n at org.apache.spark.executor.Executor.reportHeartBeat(Executor.scala:1036) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at org.apache.spark.executor.Executor.$anonfun$heartbeater$1(Executor.scala:238) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at scala.runtime.java8.JFunction0$mcV$sp.apply(JFunction0$mcV$sp.scala:18) ~[scala-library-2.13.8.jar:?]\r\n at org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:2066) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at org.apache.spark.Heartbeater$$anon$1.run(Heartbeater.scala:46) ~[spark-core_2.13-3.3.0.jar:3.3.0]\r\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) ~[?:?]\r\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:305) ~[?:?]\r\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:305) ~[?:?]\r\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) ~[?:?]\r\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) ~[?:?]\r\n at java.lang.Thread.run(Thread.java:833) ~[?:?] {noformat}","issue_id":"13470364","key":"SPARK-39696","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2023-04-07T04:02:51.000+0000","role":"fixed_distractor","summary":"Uncaught exception in thread executor-heartbeater java.util.ConcurrentModificationException: mutation occurred during iteration"} {"case_id":"13551072","cluster":"DISTRACTOR-SPARK-45201","comments":[{"body":"I also tried building and then running Spark with Java17 instead of Java8 with no changes.","created":"2023-09-18T12:44:08.219+0000"},{"body":"I also added a Dockerfile which builds Spark 3.5.0 using the eclipse-temurin:11-jdk-focal image (Ubuntu based), I still get the same error. \r\n\r\nMy container command is as follows:\r\n{code:java}\r\ncommand:\r\n - /bin/bash\r\n - '-c' \r\nargs:\r\n - >-\r\n /opt/entrypoint.sh\r\n /opt/spark/sbin/start-connect-server.sh\r\n --properties-file /opt/spark/spark-properties.conf{code}\r\nThis whole configuration works without issues with Spark 3.4.1. \r\n\r\n ","created":"2023-09-18T16:35:31.775+0000"},{"body":"Out of completeness, this is what my spark-properties.conf looks like (i've censored some info):\r\n{code:java}\r\nspark.connect.grpc.binding.port 15002\r\nspark.driver.host 10.0.0.246\r\nspark.dynamicAllocation.enabled true\r\nspark.dynamicAllocation.minExecutors 0\r\nspark.dynamicAllocation.maxExecutors 5\r\nspark.dynamicAllocation.executorAllocationRatio 0.25\r\nspark.dynamicAllocation.schedulerBacklogTimeout 1s\r\nspark.dynamicAllocation.sustainedSchedulerBacklogTimeout 120s\r\nspark.dynamicAllocation.executorIdleTimeout     300s\r\nspark.dynamicAllocation.cachedExecutorIdleTimeout 1800s\r\nspark.jars.ivy /tmp/.ivy\r\nspark.kubernetes.allocation.driver.readinessTimeout 60s\r\nspark.kubernetes.executor.container.image 0123456789.dkr.ecr.some-aws-region.amazonaws.com/spark-connect-3.5.0:v1.0.18\r\nspark.kubernetes.container.image.pullPolicy Always\r\nspark.kubernetes.driver.pod.name spark-connect-0\r\nspark.kubernetes.executor.podTemplateFile /opt/spark/executor-pod-template.yaml\r\nspark.kubernetes.executor.request.cores 12000m\r\nspark.driver.cores 1\r\nspark.executor.cores 64\r\nspark.kubernetes.namespace spark-connect\r\nspark.master k8s://https://kubernetes.default.svc.cluster.local:443\r\nspark.ui.port 4040\r\nspark.executor.memory 40000m\r\nspark.executor.memoryOverhead 8000m\r\nspark.driver.memory 10240m\r\nspark.driver.memoryOverhead 2048m\r\nspark.executor.extraJavaOptions -XX:+ExitOnOutOfMemoryError -XX:+UseCompressedOops -XX:+UseG1GC\r\nspark.driver.extraJavaOptions -XX:+ExitOnOutOfMemoryError -XX:+UseCompressedOops -XX:+UseG1GC\r\nspark.sql.parquet.datetimeRebaseModeInWrite CORRECTED\r\nspark.driver.extraClassPath /opt/spark/jars/*\r\nspark.driver.extraLibraryPath /opt/hadoop/lib/native\r\nspark.executor.extraClassPath /opt/spark/jars/*\r\nspark.executor.extraLibraryPath /opt/hadoop/lib/native\r\nspark.sql.hive.metastore.jars builtin\r\nspark.hadoop.aws.region some-aws-region\r\nspark.sql.catalogImplementation hive\r\nspark.sql.execution.arrow.pyspark.enabled true\r\nspark.sql.execution.arrow.pyspark.fallback.enabled true\r\nspark.eventLog.enabled false\r\nspark.sql.extensions io.delta.sql.DeltaSparkSessionExtension\r\nspark.sql.catalog.spark_catalog org.apache.spark.sql.delta.catalog.DeltaCatalog\r\nspark.delta.logStore.s3a.impl io.delta.storage.S3DynamoDBLogStore\r\nspark.io.delta.storage.S3DynamoDBLogStore.ddb.tableName dynamodb-table-delta-table-lock\r\nspark.io.delta.storage.S3DynamoDBLogStore.ddb.region some-aws-region\r\nspark.databricks.delta.replaceWhere.constraintCheck.enabled false\r\nspark.databricks.delta.replaceWhere.dataColumns.enabled true\r\nspark.databricks.delta.schema.autoMerge.enabled false\r\nspark.databricks.delta.merge.repartitionBeforeWrite.enabled true\r\nspark.databricks.delta.optimize.repartition.enabled true\r\nspark.databricks.delta.checkpoint.partSize 10\r\nspark.databricks.delta.properties.defaults.dataSkippingNumIndexedCols -1\r\nspark.kubernetes.authenticate.driver.serviceAccountName spark-connect\r\nspark.kubernetes.authenticate.executor.serviceAccountName spark-connect\r\nspark.kubernetes.executor.annotation.eks.amazonaws.com/role-arn arn:aws:iam::0123456789:role/spark-connect-irsa\r\nspark.kubernetes.authenticate.submission.caCertFile /var/run/secrets/kubernetes.io/serviceaccount/ca.crt\r\nspark.kubernetes.authenticate.submission.oauthTokenFile /var/run/secrets/kubernetes.io/serviceaccount/token\r\nspark.hadoop.fs.s3a.aws.credentials.provider com.amazonaws.auth.WebIdentityTokenCredentialsProvider\r\nspark.hadoop.fs.s3a.impl org.apache.hadoop.fs.s3a.S3AFileSystem\r\nspark.hadoop.fs.s3.impl org.apache.hadoop.fs.s3a.S3AFileSystem\r\nspark.hadoop.fs.s3a.fast.upload true\r\nspark.hadoop.fs.s3a.experimental.input.fadvise random\r\nspark.hive.imetastoreclient.factory.class com.amazonaws.glue.catalog.metastore.AWSGlueDataCatalogHiveClientFactory\r\nspark.aws.glue.cache.table.enable true\r\nspark.aws.glue.cache.table.size 1000\r\nspark.aws.glue.cache.table.ttl-mins 30\r\nspark.aws.glue.cache.db.enable true\r\nspark.aws.glue.cache.db.size 1000\r\nspark.aws.glue.cache.db.ttl-mins 30\r\nspark.serializer org.apache.spark.serializer.KryoSerializer\r\nspark.hadoop.mapreduce.input.fileinputformat.list-status.num-threads 64\r\nspark.sql.sources.parallelPartitionDiscovery.threshold 256\r\nspark.connect.grpc.maxInboundMessageSize 1073741824\r\nspark.connect.grpc.arrow.maxBatchSize 64m\r\nspark.sql.catalog.spark_catalog.defaultDatabase experimental {code}","created":"2023-09-18T16:46:24.692+0000"},{"body":"After spending hours analyzing the project pom files, I discovered two things.\r\n\r\nFirst, the shade plugin is relocating the guava/failureaccess package twice in the connect jars (once by the module shade plugin, once by the base project plugin). I created a simple patch to prevent the relocation of failureacces by the base plugin. I am adding the patch file [^spark-3.5.0.patch] to this Jira issue, I do not have time to create a pull request, you can apply the patch by navigating inside the source folder and running:\r\n{code:java}\r\npatch -p1 (LocalCache.java:3511)\r\n    at org.sparkproject.guava.cache.LocalCache$LoadingValueReference.(LocalCache.java:3515)\r\n    at org.sparkproject.guava.cache.LocalCache$Segment.lockedGetOrLoad(LocalCache.java:2168)\r\n    at org.sparkproject.guava.cache.LocalCache$Segment.get(LocalCache.java:2079)\r\n    at org.sparkproject.guava.cache.LocalCache.get(LocalCache.java:4011)\r\n    at org.sparkproject.guava.cache.LocalCache.getOrLoad(LocalCache.java:4034)\r\n    at org.sparkproject.guava.cache.LocalCache$LocalLoadingCache.get(LocalCache.java:5010)\r\n    at org.apache.spark.storage.BlockManagerId$.getCachedBlockManagerId(BlockManagerId.scala:146)\r\n    at org.apache.spark.storage.BlockManagerId$.apply(BlockManagerId.scala:127)\r\n    at org.apache.spark.storage.BlockManager.initialize(BlockManager.scala:536)\r\n    at org.apache.spark.SparkContext.(SparkContext.scala:625)\r\n    at org.apache.spark.SparkContext$.getOrCreate(SparkContext.scala:2888)\r\n    at org.apache.spark.sql.SparkSession$Builder.$anonfun$getOrCreate$2(SparkSession.scala:1099)\r\n    at scala.Option.getOrElse(Option.scala:189)\r\n    at org.apache.spark.sql.SparkSession$Builder.getOrCreate(SparkSession.scala:1093)\r\n    at org.apache.spark.sql.connect.service.SparkConnectServer$.main(SparkConnectServer.scala:34)\r\n    at org.apache.spark.sql.connect.service.SparkConnectServer.main(SparkConnectServer.scala)\r\n    at java.base/jdk.internal.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n    at java.base/jdk.internal.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:77)\r\n    at java.base/jdk.internal.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n    at java.base/java.lang.reflect.Method.invoke(Method.java:568)\r\n    at org.apache.spark.deploy.JavaMainApplication.start(SparkApplication.scala:52)\r\n    at org.apache.spark.deploy.SparkSubmit.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:1029)\r\n    at org.apache.spark.deploy.SparkSubmit.doRunMain$1(SparkSubmit.scala:194)\r\n    at org.apache.spark.deploy.SparkSubmit.submit(SparkSubmit.scala:217)\r\n    at org.apache.spark.deploy.SparkSubmit.doSubmit(SparkSubmit.scala:91)\r\n    at org.apache.spark.deploy.SparkSubmit$$anon$2.doSubmit(SparkSubmit.scala:1120)\r\n    at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:1129)\r\n    at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\r\nCaused by: java.lang.ClassNotFoundException: org.sparkproject.guava.util.concurrent.internal.InternalFutureFailureAccess\r\n    at java.base/jdk.internal.loader.BuiltinClassLoader.loadClass(BuiltinClassLoader.java:641)\r\n    at java.base/jdk.internal.loader.ClassLoaders$AppClassLoader.loadClass(ClassLoaders.java:188)\r\n    at java.base/java.lang.ClassLoader.loadClass(ClassLoader.java:525)\r\n    ... 56 more{code}\r\nMy build command is as follows:\r\n{code:java}\r\nexport MAVEN_OPTS=\"-Xss256m -Xmx8g -XX:ReservedCodeCacheSize=2g\"\r\n./dev/make-distribution.sh --name spark --pip -Pscala-2.12 -Pconnect -Pkubernetes -Phive -Phive-thriftserver -Phadoop-3 -Dhadoop.version=\"3.3.4\" -Dhive.version=\"2.3.9\" -Dhive23.version=\"2.3.9\" -Dhive.version.short=\"2.3\"{code}\r\nI am building Spark using Debian Bookworm, with Java 8 8u382 and Maven 3.8.8.\r\nI get the same error even when I omitt the -Pconnect profile from Maven, and simply add the \"org.apache.spark:spark-connect_2.12:3.5.0\" jar with the appropriate spark config. On the other hand, if I download the pre-built spark package, and add the spark-connect jar, this error does not appear. \r\nWhat could I be possibly missing in my build environment? I have omitted the  yarn, mesos and sparkr profiles (which are used in the distributed built) on purpose, but I do not see how these affect spark connect.\r\n\r\nAny help will be appreciated!\r\n\r\n ","from":"reporter","subject":"NoClassDefFoundError: InternalFutureFailureAccess when compiling Spark 3.5.0"},{"body":"I also tried building and then running Spark with Java17 instead of Java8 with no changes.","from":"developer"},{"body":"I also added a Dockerfile which builds Spark 3.5.0 using the eclipse-temurin:11-jdk-focal image (Ubuntu based), I still get the same error. \r\n\r\nMy container command is as follows:\r\n{code:java}\r\ncommand:\r\n - /bin/bash\r\n - '-c' \r\nargs:\r\n - >-\r\n /opt/entrypoint.sh\r\n /opt/spark/sbin/start-connect-server.sh\r\n --properties-file /opt/spark/spark-properties.conf{code}\r\nThis whole configuration works without issues with Spark 3.4.1. \r\n\r\n ","from":"developer"},{"body":"Out of completeness, this is what my spark-properties.conf looks like (i've censored some info):\r\n{code:java}\r\nspark.connect.grpc.binding.port 15002\r\nspark.driver.host 10.0.0.246\r\nspark.dynamicAllocation.enabled true\r\nspark.dynamicAllocation.minExecutors 0\r\nspark.dynamicAllocation.maxExecutors 5\r\nspark.dynamicAllocation.executorAllocationRatio 0.25\r\nspark.dynamicAllocation.schedulerBacklogTimeout 1s\r\nspark.dynamicAllocation.sustainedSchedulerBacklogTimeout 120s\r\nspark.dynamicAllocation.executorIdleTimeout     300s\r\nspark.dynamicAllocation.cachedExecutorIdleTimeout 1800s\r\nspark.jars.ivy /tmp/.ivy\r\nspark.kubernetes.allocation.driver.readinessTimeout 60s\r\nspark.kubernetes.executor.container.image 0123456789.dkr.ecr.some-aws-region.amazonaws.com/spark-connect-3.5.0:v1.0.18\r\nspark.kubernetes.container.image.pullPolicy Always\r\nspark.kubernetes.driver.pod.name spark-connect-0\r\nspark.kubernetes.executor.podTemplateFile /opt/spark/executor-pod-template.yaml\r\nspark.kubernetes.executor.request.cores 12000m\r\nspark.driver.cores 1\r\nspark.executor.cores 64\r\nspark.kubernetes.namespace spark-connect\r\nspark.master k8s://https://kubernetes.default.svc.cluster.local:443\r\nspark.ui.port 4040\r\nspark.executor.memory 40000m\r\nspark.executor.memoryOverhead 8000m\r\nspark.driver.memory 10240m\r\nspark.driver.memoryOverhead 2048m\r\nspark.executor.extraJavaOptions -XX:+ExitOnOutOfMemoryError -XX:+UseCompressedOops -XX:+UseG1GC\r\nspark.driver.extraJavaOptions -XX:+ExitOnOutOfMemoryError -XX:+UseCompressedOops -XX:+UseG1GC\r\nspark.sql.parquet.datetimeRebaseModeInWrite CORRECTED\r\nspark.driver.extraClassPath /opt/spark/jars/*\r\nspark.driver.extraLibraryPath /opt/hadoop/lib/native\r\nspark.executor.extraClassPath /opt/spark/jars/*\r\nspark.executor.extraLibraryPath /opt/hadoop/lib/native\r\nspark.sql.hive.metastore.jars builtin\r\nspark.hadoop.aws.region some-aws-region\r\nspark.sql.catalogImplementation hive\r\nspark.sql.execution.arrow.pyspark.enabled true\r\nspark.sql.execution.arrow.pyspark.fallback.enabled true\r\nspark.eventLog.enabled false\r\nspark.sql.extensions io.delta.sql.DeltaSparkSessionExtension\r\nspark.sql.catalog.spark_catalog org.apache.spark.sql.delta.catalog.DeltaCatalog\r\nspark.delta.logStore.s3a.impl io.delta.storage.S3DynamoDBLogStore\r\nspark.io.delta.storage.S3DynamoDBLogStore.ddb.tableName dynamodb-table-delta-table-lock\r\nspark.io.delta.storage.S3DynamoDBLogStore.ddb.region some-aws-region\r\nspark.databricks.delta.replaceWhere.constraintCheck.enabled false\r\nspark.databricks.delta.replaceWhere.dataColumns.enabled true\r\nspark.databricks.delta.schema.autoMerge.enabled false\r\nspark.databricks.delta.merge.repartitionBeforeWrite.enabled true\r\nspark.databricks.delta.optimize.repartition.enabled true\r\nspark.databricks.delta.checkpoint.partSize 10\r\nspark.databricks.delta.properties.defaults.dataSkippingNumIndexedCols -1\r\nspark.kubernetes.authenticate.driver.serviceAccountName spark-connect\r\nspark.kubernetes.authenticate.executor.serviceAccountName spark-connect\r\nspark.kubernetes.executor.annotation.eks.amazonaws.com/role-arn arn:aws:iam::0123456789:role/spark-connect-irsa\r\nspark.kubernetes.authenticate.submission.caCertFile /var/run/secrets/kubernetes.io/serviceaccount/ca.crt\r\nspark.kubernetes.authenticate.submission.oauthTokenFile /var/run/secrets/kubernetes.io/serviceaccount/token\r\nspark.hadoop.fs.s3a.aws.credentials.provider com.amazonaws.auth.WebIdentityTokenCredentialsProvider\r\nspark.hadoop.fs.s3a.impl org.apache.hadoop.fs.s3a.S3AFileSystem\r\nspark.hadoop.fs.s3.impl org.apache.hadoop.fs.s3a.S3AFileSystem\r\nspark.hadoop.fs.s3a.fast.upload true\r\nspark.hadoop.fs.s3a.experimental.input.fadvise random\r\nspark.hive.imetastoreclient.factory.class com.amazonaws.glue.catalog.metastore.AWSGlueDataCatalogHiveClientFactory\r\nspark.aws.glue.cache.table.enable true\r\nspark.aws.glue.cache.table.size 1000\r\nspark.aws.glue.cache.table.ttl-mins 30\r\nspark.aws.glue.cache.db.enable true\r\nspark.aws.glue.cache.db.size 1000\r\nspark.aws.glue.cache.db.ttl-mins 30\r\nspark.serializer org.apache.spark.serializer.KryoSerializer\r\nspark.hadoop.mapreduce.input.fileinputformat.list-status.num-threads 64\r\nspark.sql.sources.parallelPartitionDiscovery.threshold 256\r\nspark.connect.grpc.maxInboundMessageSize 1073741824\r\nspark.connect.grpc.arrow.maxBatchSize 64m\r\nspark.sql.catalog.spark_catalog.defaultDatabase experimental {code}","from":"developer"},{"body":"After spending hours analyzing the project pom files, I discovered two things.\r\n\r\nFirst, the shade plugin is relocating the guava/failureaccess package twice in the connect jars (once by the module shade plugin, once by the base project plugin). I created a simple patch to prevent the relocation of failureacces by the base plugin. I am adding the patch file [^spark-3.5.0.patch] to this Jira issue, I do not have time to create a pull request, you can apply the patch by navigating inside the source folder and running:\r\n{code:java}\r\npatch -p1 (LocalCache.java:3511)\r\n    at org.sparkproject.guava.cache.LocalCache$LoadingValueReference.(LocalCache.java:3515)\r\n    at org.sparkproject.guava.cache.LocalCache$Segment.lockedGetOrLoad(LocalCache.java:2168)\r\n    at org.sparkproject.guava.cache.LocalCache$Segment.get(LocalCache.java:2079)\r\n    at org.sparkproject.guava.cache.LocalCache.get(LocalCache.java:4011)\r\n    at org.sparkproject.guava.cache.LocalCache.getOrLoad(LocalCache.java:4034)\r\n    at org.sparkproject.guava.cache.LocalCache$LocalLoadingCache.get(LocalCache.java:5010)\r\n    at org.apache.spark.storage.BlockManagerId$.getCachedBlockManagerId(BlockManagerId.scala:146)\r\n    at org.apache.spark.storage.BlockManagerId$.apply(BlockManagerId.scala:127)\r\n    at org.apache.spark.storage.BlockManager.initialize(BlockManager.scala:536)\r\n    at org.apache.spark.SparkContext.(SparkContext.scala:625)\r\n    at org.apache.spark.SparkContext$.getOrCreate(SparkContext.scala:2888)\r\n    at org.apache.spark.sql.SparkSession$Builder.$anonfun$getOrCreate$2(SparkSession.scala:1099)\r\n    at scala.Option.getOrElse(Option.scala:189)\r\n    at org.apache.spark.sql.SparkSession$Builder.getOrCreate(SparkSession.scala:1093)\r\n    at org.apache.spark.sql.connect.service.SparkConnectServer$.main(SparkConnectServer.scala:34)\r\n    at org.apache.spark.sql.connect.service.SparkConnectServer.main(SparkConnectServer.scala)\r\n    at java.base/jdk.internal.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n    at java.base/jdk.internal.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:77)\r\n    at java.base/jdk.internal.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n    at java.base/java.lang.reflect.Method.invoke(Method.java:568)\r\n    at org.apache.spark.deploy.JavaMainApplication.start(SparkApplication.scala:52)\r\n    at org.apache.spark.deploy.SparkSubmit.org$apache$spark$deploy$SparkSubmit$$runMain(SparkSubmit.scala:1029)\r\n    at org.apache.spark.deploy.SparkSubmit.doRunMain$1(SparkSubmit.scala:194)\r\n    at org.apache.spark.deploy.SparkSubmit.submit(SparkSubmit.scala:217)\r\n    at org.apache.spark.deploy.SparkSubmit.doSubmit(SparkSubmit.scala:91)\r\n    at org.apache.spark.deploy.SparkSubmit$$anon$2.doSubmit(SparkSubmit.scala:1120)\r\n    at org.apache.spark.deploy.SparkSubmit$.main(SparkSubmit.scala:1129)\r\n    at org.apache.spark.deploy.SparkSubmit.main(SparkSubmit.scala)\r\nCaused by: java.lang.ClassNotFoundException: org.sparkproject.guava.util.concurrent.internal.InternalFutureFailureAccess\r\n    at java.base/jdk.internal.loader.BuiltinClassLoader.loadClass(BuiltinClassLoader.java:641)\r\n    at java.base/jdk.internal.loader.ClassLoaders$AppClassLoader.loadClass(ClassLoaders.java:188)\r\n    at java.base/java.lang.ClassLoader.loadClass(ClassLoader.java:525)\r\n    ... 56 more{code}\r\nMy build command is as follows:\r\n{code:java}\r\nexport MAVEN_OPTS=\"-Xss256m -Xmx8g -XX:ReservedCodeCacheSize=2g\"\r\n./dev/make-distribution.sh --name spark --pip -Pscala-2.12 -Pconnect -Pkubernetes -Phive -Phive-thriftserver -Phadoop-3 -Dhadoop.version=\"3.3.4\" -Dhive.version=\"2.3.9\" -Dhive23.version=\"2.3.9\" -Dhive.version.short=\"2.3\"{code}\r\nI am building Spark using Debian Bookworm, with Java 8 8u382 and Maven 3.8.8.\r\nI get the same error even when I omitt the -Pconnect profile from Maven, and simply add the \"org.apache.spark:spark-connect_2.12:3.5.0\" jar with the appropriate spark config. On the other hand, if I download the pre-built spark package, and add the spark-connect jar, this error does not appear. \r\nWhat could I be possibly missing in my build environment? I have omitted the  yarn, mesos and sparkr profiles (which are used in the distributed built) on purpose, but I do not see how these affect spark connect.\r\n\r\nAny help will be appreciated!\r\n\r\n ","issue_id":"13551072","key":"SPARK-45201","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2024-07-24T01:55:01.000+0000","role":"fixed_distractor","summary":"NoClassDefFoundError: InternalFutureFailureAccess when compiling Spark 3.5.0"} {"case_id":"12762466","cluster":"DISTRACTOR-SPARK-4879","comments":[{"body":"Hey Josh, \n\nI have been playing around with your repro above and I think I can consistently trigger the bad behavior by just tweaking the value of {{spark.speculation.multiplier}} and {{spark.speculation.quantile}}.\n\nI set the {{multiplier}} to be 1 and the {{quantile}} to 0.01 so that only 1% of tasks have to finish before any task that takes longer than those 1% of tasks should speculate. \nAs expected, I see a lot of tasks getting speculated. \nAfter running the repro about 5 times, I have seen 2 errors (stack traces at the bottom and the full run from the REPL is attached with this comment). \n\nOne thing I do notice is that the part-00000 associated with Stage 1 was always where I expected it to be in HDFS, and all lines were present (checked using a {{wc -l}})\n\n\n{code}\nscala> 15/01/07 13:44:26 WARN scheduler.TaskSetManager: Lost task 0.1 in stage 0.0 (TID 119, ): java.io.IOException: The temporary job-output directory hdfs://:8020/test6/_temporary doesn't exist!\n org.apache.hadoop.mapred.FileOutputCommitter.getWorkPath(FileOutputCommitter.java:250)\n org.apache.hadoop.mapred.FileOutputFormat.getTaskOutputPath(FileOutputFormat.java:240)\n org.apache.hadoop.mapred.TextOutputFormat.getRecordWriter(TextOutputFormat.java:116)\n org.apache.spark.SparkHadoopWriter.open(SparkHadoopWriter.scala:89)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:980)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:974)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:62)\n org.apache.spark.scheduler.Task.run(Task.scala:54)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:178)\n java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n java.lang.Thread.run(Thread.java:745)\n{code}\n\n{code}\n15/01/07 15:17:39 WARN scheduler.TaskSetManager: Lost task 0.1 in stage 0.0 (TID 120, ): org.apache.hadoop.ipc.RemoteException: No lease on /test7/_temporary/_attempt_201501071517_0000_m_000000_120/part-00000: File does not exist. Holder DFSClient_NONMAPREDUCE_-469253416_73 does not have any open files.\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkLease(FSNamesystem.java:2609)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.analyzeFileState(FSNamesystem.java:2426)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getAdditionalBlock(FSNamesystem.java:2339)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.addBlock(NameNodeRpcServer.java:501)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.addBlock(ClientNamenodeProtocolServerSideTranslatorPB.java:299)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java:44954)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:453)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:1002)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:1752)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:1748)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:415)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1438)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:1746)\n\n org.apache.hadoop.ipc.Client.call(Client.java:1238)\n org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:202)\n com.sun.proxy.$Proxy9.addBlock(Unknown Source)\n sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n java.lang.reflect.Method.invoke(Method.java:606)\n org.apache.hadoop.io.retry.RetryInvocationHandler.invokeMethod(RetryInvocationHandler.java:164)\n org.apache.hadoop.io.retry.RetryInvocationHandler.invoke(RetryInvocationHandler.java:83)\n com.sun.proxy.$Proxy9.addBlock(Unknown Source)\n org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolTranslatorPB.addBlock(ClientNamenodeProtocolTranslatorPB.java:291)\n org.apache.hadoop.hdfs.DFSOutputStream$DataStreamer.locateFollowingBlock(DFSOutputStream.java:1177)\n org.apache.hadoop.hdfs.DFSOutputStream$DataStreamer.nextBlockOutputStream(DFSOutputStream.java:1030)\n org.apache.hadoop.hdfs.DFSOutputStream$DataStreamer.run(DFSOutputStream.java:488)\n{code}\n","created":"2015-01-07T23:47:31.634+0000"},{"body":"Hey Zach,\n\nFrom my last round of attempts (maybe a week ago), I was able to reproduce some sort of \"No lease on ....\" error but not the missing output partitions symptom. I'd still like to try to find a reproduction for the missing files but this could be tricky if it involves a very quick race condition; I don't think that's the case, though, since it's pretty unlikely that the speculated and original task are finishing at the exact same time.","created":"2015-01-09T19:57:44.097+0000"},{"body":"Hey Josh,\n\nI was able to reproduce the missing file using the speculation settings in my previous comment:\n\n{code}\nscala> 15/01/09 18:33:28 WARN scheduler.TaskSetManager: Lost task 42.1 in stage 0.0 (TID 113, ): java.io.IOException: Failed to save output of task: attempt_201501091833_0000_m_000042_113\n org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:160)\n org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:172)\n org.apache.hadoop.mapred.FileOutputCommitter.commitTask(FileOutputCommitter.java:132)\n org.apache.spark.SparkHadoopWriter.commit(SparkHadoopWriter.scala:109)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:991)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:974)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:62)\n org.apache.spark.scheduler.Task.run(Task.scala:54)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:178)\n java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n java.lang.Thread.run(Thread.java:745)\n15/01/09 18:33:47 WARN scheduler.TaskSetManager: Lost task 0.0 in stage 0.0 (TID 0, ): java.io.IOException: The temporary job-output directory hdfs://:8020/test2/_temporary doesn't exist!\n org.apache.hadoop.mapred.FileOutputCommitter.getWorkPath(FileOutputCommitter.java:250)\n org.apache.hadoop.mapred.FileOutputFormat.getTaskOutputPath(FileOutputFormat.java:240)\n org.apache.hadoop.mapred.TextOutputFormat.getRecordWriter(TextOutputFormat.java:116)\n org.apache.spark.SparkHadoopWriter.open(SparkHadoopWriter.scala:89)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:980)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:974)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:62)\n org.apache.spark.scheduler.Task.run(Task.scala:54)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:178)\n java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n java.lang.Thread.run(Thread.java:745)\n{code}\n\nNotice here that there are only 99 part files and part-00042 is missing (as seen in the stacktrace above)\n{code}\n $ hadoop fs -ls /test2 | grep part | wc -l\n99\n $ hadoop fs -ls /test2 | grep part-0004\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00040\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00041\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00043\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00044\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00045\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00046\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00047\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00048\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00049\n{code}\n\n","created":"2015-01-10T02:50:17.087+0000"},{"body":"For clarity, here is the scala code I used in the REPL:\n{code}\nscala> val numTasks = 100\nnumTasks: Int = 100\n\nscala> sc.parallelize(1 to numTasks, numTasks).mapPartitionsWithContext { case (ctx, iter) =>\n | if (ctx.partitionId == 0) { // If this is the one task that should run really slow\n | if (ctx.attemptId == 0) { // If this is the first attempt, run slow\n | Thread.sleep(20 * 1000)\n | }\n | }\n | iter\n | }.map(x => (x, x)).saveAsTextFile(\"/test2\")\n{code}","created":"2015-01-10T02:51:02.229+0000"},{"body":"I think that part of the reproduction issues that I had might have been due to {{attemptId}} returning a unique task attempt ID rather than the attempt number, meaning that only the _first_ run of that test in the REPL would be capable of uncovering the bug.\n\nSee https://github.com/apache/spark/pull/3849 / SPARK-4014 for more context. I'm going to try to merge that patch today, which will let me write a reliable regression test.","created":"2015-01-12T19:31:15.536+0000"},{"body":"I'm not sure that SparkHadoopWriter's use of FileOutputCommitter properly obeys the OutputCommitter contracts in Hadoop. According to the [OutputCommitter Javadoc|https://hadoop.apache.org/docs/current/api/org/apache/hadoop/mapreduce/OutputCommitter.html]\n\n{quote}\nThe methods in this class can be called from several different processes and from several different contexts. It is important to know which process and which context each is called from. Each method should be marked accordingly in its documentation. It is also important to note that not all methods are guaranteed to be called once and only once. If a method is not guaranteed to have this property the output committer needs to handle this appropriately. Also note it will only be in rare situations where they may be called multiple times for the same task.\n{quote}\n\nBased on the documentation, `needsTaskCommit` \" is called from each individual task's process that will output to HDFS, and it is called just for that task.\", so it seems like it should be safe to call this from SparkHadoopWriter.\n\nHowever, maybe we're misusing the `commitTask` method:\n\n{quote}\nIf needsTaskCommit(TaskAttemptContext) returns true and this task is the task that the AM determines finished first, this method is called to commit an individual task's output. This is to mark that tasks output as complete, as commitJob(JobContext) will also be called later on if the entire job finished successfully. This is called from a task's process. This may be called multiple times for the same task, but different task attempts. It should be very rare for this to be called multiple times and requires odd networking failures to make this happen. In the future the Hadoop framework may eliminate this race. \n{quote}\n\nI think that we're missing the \"this task is the task that the AM determines finished first\" part of the equation here. If `needsTaskCommit` is false, then we definitely shouldn't commit (e.g. if it's an original task that lost to a speculated copy), but if it's true then I don't think it's safe to commit; we need some central authority to pick a winner.\n\nLet's see how Hadoop does things, working backwards from actual calls of `commitTask` to see whether they're guarded by some coordination through the AM. It looks like `OutputCommitter` is part of the `mapred` API, so I'll only look at classes in that package:\n\nIn `Task.java`, `committer.commitTask` is only performed after checking `canCommit` through `TaskUmbilicalProtocol`: https://github.com/apache/hadoop/blob/a655973e781caf662b360c96e0fa3f5a873cf676/hadoop-mapreduce-project/hadoop-mapreduce-client/hadoop-mapreduce-client-core/src/main/java/org/apache/hadoop/mapred/Task.java#L1185. According to the Javadocs for TaskAttemptListenerImpl.canCommit (the actual concrete implementation of this method):\n\n{code}\n /**\n * Child checking whether it can commit.\n * \n *
\n * Commit is a two-phased protocol. First the attempt informs the\n * ApplicationMaster that it is\n * {@link #commitPending(TaskAttemptID, TaskStatus)}. Then it repeatedly polls\n * the ApplicationMaster whether it {@link #canCommit(TaskAttemptID)} This is\n * a legacy from the centralized commit protocol handling by the JobTracker.\n */\n @Override\n public boolean canCommit(TaskAttemptID taskAttemptID) throws IOException {\n{code}\n\nThis ends up delegating to `Task.canCommit()`:\n\n{code}\n /**\n * Can the output of the taskAttempt be committed. Note that once the task\n * gives a go for a commit, further canCommit requests from any other attempts\n * should return false.\n * \n * @param taskAttemptID\n * @return whether the attempt's output can be committed or not.\n */\n boolean canCommit(TaskAttemptId taskAttemptID);\n{code}\n\nThere's a bunch of tricky logic that involves communication with the AM (see AttemptCommitPendingTransition and the other transitions in TaskImpl), but it looks like the gist is that the \"winner\" is picked by the AM through some central coordination process. \n\nSo, it looks like the right fix is to implement these same state transitions ourselves. It would be nice if there was a clean way to do this that could be easily backported to maintenance branches. ","created":"2015-01-15T22:29:16.953+0000"},{"body":"User 'JoshRosen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4066","created":"2015-01-16T02:18:09.423+0000"},{"body":"User 'mccheah' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4155","created":"2015-01-22T02:02:17.023+0000"},{"body":"This issue is _really_ hard to reproduce, but I managed to trigger the original bug as part of the testing for my patch. Here's what I ran:\n\n{code}\n~/spark-1.3.0-SNAPSHOT-bin-1.0.4/bin/spark-shell --conf spark.speculation.multiplier=1 --conf spark.speculation.quantile=0.01 --conf spark.speculation=true --conf spark.hadoop.outputCommitCoordination.enabled=false\n{code}\n\n{code}\nval numTasks = 100\nval numTrials = 100\nval outputPath = \"/output-committer-bug-\"\nval sleepDuration = 1000\n\nfor (trial <- 0 to (numTrials - 1)) {\n val outputLocation = outputPath + trial\n sc.parallelize(1 to numTasks, numTasks).mapPartitionsWithContext { case (ctx, iter) =>\n if (ctx.partitionId % 5 == 0) {\n if (ctx.attemptNumber == 0) { // If this is the first attempt, run slow\n Thread.sleep(sleepDuration)\n }\n }\n iter\n }.map(identity).saveAsTextFile(outputLocation)\n Thread.sleep(sleepDuration * 2)\n println(\"TESTING OUTPUT OF TRIAL \" + trial)\n val savedData = sc.textFile(outputLocation).map(_.toInt).collect()\n if (savedData.toSet != (1 to numTasks).toSet) {\n println(\"MISSING: \" + ((1 to numTasks).toSet -- savedData.toSet))\n assert(false)\n }\n println(\"-\" * 80)\n}\n{code}\n\nIt took 22 runs until I actually observed missing output partitions (several of the earlier runs threw spurious exceptions and didn't have missing outputs):\n\n{code}\n[...]\n15/02/10 22:17:21 INFO scheduler.DAGScheduler: Job 66 finished: saveAsTextFile at :39, took 2.479592 s\n15/02/10 22:17:21 WARN scheduler.TaskSetManager: Lost task 75.0 in stage 66.0 (TID 6861, ip-172-31-1-124.us-west-2.compute.internal): java.io.IOException: Failed to save output of task: attempt_201502102217_0066_m_000075_6861\n at org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:160)\n at org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:172)\n at org.apache.hadoop.mapred.FileOutputCommitter.commitTask(FileOutputCommitter.java:132)\n at org.apache.spark.SparkHadoopWriter.performCommit$1(SparkHadoopWriter.scala:113)\n at org.apache.spark.SparkHadoopWriter.commit(SparkHadoopWriter.scala:150)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:1082)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:1059)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:61)\n at org.apache.spark.scheduler.Task.run(Task.scala:64)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:197)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n\n15/02/10 22:17:21 WARN scheduler.TaskSetManager: Lost task 80.0 in stage 66.0 (TID 6866, ip-172-31-11-151.us-west-2.compute.internal): java.io.IOException: The temporary job-output directory hdfs://ec2-54-213-142-80.us-west-2.compute.amazonaws.com:9000/output-committer-bug-22/_temporary doesn't exist!\n at org.apache.hadoop.mapred.FileOutputCommitter.getWorkPath(FileOutputCommitter.java:250)\n at org.apache.hadoop.mapred.FileOutputFormat.getTaskOutputPath(FileOutputFormat.java:244)\n at org.apache.hadoop.mapred.TextOutputFormat.getRecordWriter(TextOutputFormat.java:116)\n at org.apache.spark.SparkHadoopWriter.open(SparkHadoopWriter.scala:91)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:1068)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:1059)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:61)\n at org.apache.spark.scheduler.Task.run(Task.scala:64)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:197)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n\n15/02/10 22:17:21 INFO scheduler.TaskSetManager: Lost task 85.0 in stage 66.0 (TID 6871) on executor ip-172-31-1-124.us-west-2.compute.internal: java.io.IOException (The temporary job-output directory hdfs://ec2-54-213-142-80.us-west-2.compute.amazonaws.com:9000/output-committer-bug-22/_temporary doesn't exist!) [duplicate 1]\n15/02/10 22:17:21 INFO scheduler.TaskSetManager: Lost task 90.0 in stage 66.0 (TID 6876) on executor ip-172-31-1-124.us-west-2.compute.internal: java.io.IOException (The temporary job-output directory hdfs://ec2-54-213-142-80.us-west-2.compute.amazonaws.com:9000/output-committer-bug-22/_temporary doesn't exist!) [duplicate 2]\n15/02/10 22:17:21 INFO scheduler.TaskSetManager: Lost task 95.0 in stage 66.0 (TID 6881) on executor ip-172-31-11-151.us-west-2.compute.internal: java.io.IOException (The temporary job-output directory hdfs://ec2-54-213-142-80.us-west-2.compute.amazonaws.com:9000/output-committer-bug-22/_temporary doesn't exist!) [duplicate 3]\n15/02/10 22:17:21 INFO scheduler.TaskSchedulerImpl: Removed TaskSet 66.0, whose tasks have all completed, from pool\nTESTING OUTPUT OF TRIAL 22\n[...]\n15/02/10 22:17:23 INFO scheduler.TaskSchedulerImpl: Removed TaskSet 67.0, whose tasks have all completed, from pool\n15/02/10 22:17:23 INFO scheduler.DAGScheduler: Stage 67 (collect at :42) finished in 0.158 s\n15/02/10 22:17:23 INFO scheduler.DAGScheduler: Job 67 finished: collect at :42, took 0.165580 s\nMISSING: Set(76)\njava.lang.AssertionError: assertion failed\n[...]\n{code}\n\nAnd I confirmed that it's missing in HDFS:\n\n{code}\n~/ephemeral-hdfs/bin/hadoop fs -ls /output-committer-bug-22 | grep part | wc -l\nWarning: $HADOOP_HOME is deprecated.\n\n99\n{code}\n\nTo test my patch, I'm going to remove the flag that disables it and run this for a huge number of trials to ensure that we don't hit the missing output bug.","created":"2015-02-10T22:21:45.073+0000"},{"body":"This is really great work [~joshrosen]! I really appreciate the effort you're putting into getting this one figured out since these kind of non-deterministic bugs are the most painful for both users and devs to figure out.","created":"2015-02-10T22:27:05.677+0000"},{"body":"Thanks for picking this up [~joshrosen], I see that the hotly-debated patch has finally merged in!\n\nI'm a bit backed up but there will be others from my side that will be testing this. We'll comment on the bug once we have verified the issue is fixed on our side.","created":"2015-02-11T09:59:10.352+0000"},{"body":"Could this happen very very rarely when not using speculative execution?\nOnce in a long while, I have a situation where the OutputCommitter says it wrote the file successfully, but the output file doesn't appear there.","created":"2015-02-17T15:14:07.546+0000"},{"body":"[~romi-totango] what filesystem are you writing to? HDFS / S3 / other ? If this was HDFS then I think you may have a bug, but if it was S3 then there are loose eventual consistency guarantees that might be surprising you.","created":"2015-02-17T16:40:15.758+0000"},{"body":"[~aash] I'm writing to S3. I'm aware of eventual consistency issues, but is it possible that 99.999% files are written and available immediately, and then one file doesn't appear for very long? (couldn't afford to wait until it was available so I ran the process again and it wrote the file immediately).\nStill isn't it somehow possibly related to this bug or similar, that Spark's output writing mechanism thinks it wrote the file, but actually it did not?","created":"2015-02-18T08:28:54.118+0000"},{"body":"That sounds exactly like what I'd expect to happen in S3. The eventual consistency bug is that S3 reports something is written, Spark then believes it has written the file, but when you go to access it the file isn't there yet (at least in the us standard region. If you use another region, there is read-after-write consistency).\n\nIn short, it doesn't sound like you're encountering this particular issue with speculative execution (SPARK-4879), just normal S3 consistency trickiness.","created":"2015-02-18T14:52:31.019+0000"},{"body":"\nWith 1.3 RC, we are still seeing this issue.","created":"2015-03-05T10:08:15.822+0000"},{"body":"Just to add - we are seeing this on hdfs.\nOut of the 75k output files to be written, about 53k did not get moved.\n\nBut on waiting for a really long time (about 20 minutes) all files eventually got moved.\nIs this something to do with FileOutputCommitter itself or side-effect of the patch committed ?\n\nWe have speculative execution turned on; even though the saveAsHadoopFile taskset completes (as per logs), the call itself does not return until entire 'move' is done - just that it took more time to commit the output than to generate it in first place !","created":"2015-03-05T10:13:25.107+0000"},{"body":"[~mridulm80], when you say that \"the call itself does not return until entire 'move' is done\", do you mean that the {{saveAsHadoopFile()}} call blocks in {{saveAsHadoopDataset()}}'s final {{ writer.commitJob()}} call? This ticket addresses a bug where output partitions are missing _after_ control returns from the {{saveAsHadoop*()}} call.\n\nIn the issue addressed by this ticket, partitions were deleted after the job and save action had completed and output remained missing. It sounds like you might be describing a different issue where jobs take a long time to commit. For your workload, is the behavior that you are observing a performance regression in Spark 1.3? If you'd like to run 1.3 without the effects of my patch, just set {{spark.hadoop.outputCommitCoordination.enabled=false}} in your SparkConf to bypass the new output commit coordination logic.","created":"2015-03-05T10:52:16.521+0000"},{"body":"[~joshrosen] The former - the call blocks on saveAsHadoopFile (the very next line was sc.stop() and then exit).\n\nWe cannot use 'spark.hadoop.outputCommitCoordination.enabled=false' since we need to have speculative execution on (turning it off drastically affects job runtime) and disabling commit coordination causes other issues (as reported in this jira).\n\nI was not sure if \n- the dramatically increased to commit was due to the PR from this jira.\n- is orthogonal to this jira and an issue with FileOutputCommitter (as was seen which searching for cause online).\n\nGiven the central coordination blocking on master, I wanted to eliminate causes from our end before digging into FIleOutputCommitter and/or Namenode code.","created":"2015-03-05T16:03:31.388+0000"},{"body":"Can you perhaps jstack or profile the driver and the executors to see if they're blocked in OutputCommitCoordinator logic?","created":"2015-03-05T18:14:25.673+0000"},{"body":"Hi, is there any updates for this issue?","created":"2015-06-04T09:18:47.072+0000"},{"body":"I'm experiencing this issue. Sometimes rdd with 4 partitions is written with 3 parts and _SUCCESS marker is there.","created":"2015-06-15T16:28:01.681+0000"},{"body":"I have a fairly reliable reproduction with Spark 1.4.0 and HDFS. I'm running on 10 EC2 m3.2xlarge instances using the ephemeral HDFS. If {{spark.speculation}} is true, I get hit by this 50% of the time or more. It's a fairly complex workload, not something you can test in a {{spark-shell}}. What I saw was that I saved a 400-partition RDD with {{saveAsNewAPIHadoopFile}} (which returned without error) and when I tried to read it back, the files for partitions 323 and 324 were missing. (In the case that I took a closer look at.) I don't have the logs at hand now, but it's like you describe I think ({{Failed to save output of task}}). I can add them later if it would be useful.\n\nI turned off {{spark.speculation}} and haven't seen the issue since.\n\nIs there anything I could do to help debug this issue?","created":"2015-07-09T14:26:32.356+0000"},{"body":"I wonder if this issue is serious enough to note in the documentation. What do you think about adding a big fat warning for speculative execution until it is fixed? \"Enabling speculative execution may lead to missing output files\"? Or perhaps add a verification pass that checks if all the outputs are present and raises an exception if not.\n\nSilently dropping output files is a horrible bug. We've been debugging a somewhat mythological data corruption issue for about a month, and now we realize that this issue (SPARK-4879) is a very plausible explanation. We have never been able to reproduce it, but we have a log file, and it shows a speculative task for a {{saveAsNewAPIHadoopFile}} stage.","created":"2015-07-10T12:24:52.908+0000"},{"body":"[~darabos], do you think that this issue might have been resolved in an earlier Spark version but inadvertently broken in the upgrade to 1.4.0? If you have an easy reproduction, it might be helpful to see whether the problem occurs on 1.3.1.","created":"2015-07-12T19:14:32.853+0000"},{"body":"Good idea! I'll try with 1.3.1 next week.","created":"2015-07-12T20:40:00.583+0000"},{"body":"I've managed to reproduce on Spark 1.3.1 too (pre-built for Hadoop 1). I ran on EC2 with the {{spark-ec2}} script and used the ephemeral HDFS. I used a 5-machine cluster and repeatedly ran a complex test suite for about 30 minutes until the error was triggered. Here are the relevant logs:\n\n{noformat}\nI2015-07-15 13:45:19,954 TaskSetManager:[task-result-getter-2] Finished task 198.0 in stage 320.0 (TID 13290) in 568 ms on ip-10-153-188-224.ec2.internal (195/200)\nI2015-07-15 13:45:21,174 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Marking task 197 in stage 320.0 (on ip-10-231-214-6.ec2.internal) as speculatable because it ran more than 1240 ms\nI2015-07-15 13:45:21,174 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Marking task 176 in stage 320.0 (on ip-10-231-214-6.ec2.internal) as speculatable because it ran more than 1240 ms\nI2015-07-15 13:45:21,174 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Marking task 194 in stage 320.0 (on ip-10-231-214-6.ec2.internal) as speculatable because it ran more than 1240 ms\nI2015-07-15 13:45:21,174 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Marking task 196 in stage 320.0 (on ip-10-231-214-6.ec2.internal) as speculatable because it ran more than 1240 ms\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Marking task 195 in stage 320.0 (on ip-10-231-214-6.ec2.internal) as speculatable because it ran more than 1240 ms\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Starting task 197.1 in stage 320.0 (TID 13292, ip-10-153-188-224.ec2.internal, PROCESS_LOCAL, 1612 bytes)\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Starting task 176.1 in stage 320.0 (TID 13293, ip-10-63-27-248.ec2.internal, PROCESS_LOCAL, 1612 bytes)\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Starting task 194.1 in stage 320.0 (TID 13294, ip-10-154-1-239.ec2.internal, PROCESS_LOCAL, 1612 bytes)\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Starting task 195.1 in stage 320.0 (TID 13295, ip-10-228-67-34.ec2.internal, PROCESS_LOCAL, 1612 bytes)\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Starting task 196.1 in stage 320.0 (TID 13296, ip-10-153-188-224.ec2.internal, PROCESS_LOCAL, 1612 bytes)\nI2015-07-15 13:45:21,402 TaskSetManager:[task-result-getter-3] Finished task 176.0 in stage 320.0 (TID 13268) in 2223 ms on ip-10-231-214-6.ec2.internal (196/200)\nI2015-07-15 13:45:21,445 TaskSetManager:[task-result-getter-0] Finished task 195.0 in stage 320.0 (TID 13287) in 2113 ms on ip-10-231-214-6.ec2.internal (197/200)\nI2015-07-15 13:45:21,461 TaskSetManager:[task-result-getter-1] Finished task 196.1 in stage 320.0 (TID 13296) in 285 ms on ip-10-153-188-224.ec2.internal (198/200)\nI2015-07-15 13:45:21,464 TaskSetManager:[task-result-getter-2] Finished task 194.1 in stage 320.0 (TID 13294) in 287 ms on ip-10-154-1-239.ec2.internal (199/200)\nI2015-07-15 13:45:21,465 TaskSetManager:[task-result-getter-3] Ignoring task-finished event for 176.1 in stage 320.0 because task 176 has already completed successfully\nI2015-07-15 13:45:21,468 TaskSetManager:[task-result-getter-0] Finished task 197.1 in stage 320.0 (TID 13292) in 292 ms on ip-10-153-188-224.ec2.internal (200/200)\nI2015-07-15 13:45:21,468 DAGScheduler:[dag-scheduler-event-loop] Stage 320 (saveAsNewAPIHadoopFile at HadoopFile.scala:208) finished in 4.802 s\nI2015-07-15 13:45:21,468 DAGScheduler:[DataManager-5] Job 46 finished: saveAsNewAPIHadoopFile at HadoopFile.scala:208, took 4.836626 s\nW2015-07-15 13:45:21,478 TaskSetManager:[task-result-getter-1] Lost task 195.1 in stage 320.0 (TID 13295, ip-10-228-67-34.ec2.internal): java.io.IOException: Failed to save output of task: attempt_201507151345_0628_r_000195_1\n at org.apache.hadoop.mapreduce.lib.output.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:203)\n at org.apache.hadoop.mapreduce.lib.output.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:214)\n at org.apache.hadoop.mapreduce.lib.output.FileOutputCommitter.commitTask(FileOutputCommitter.java:167)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$12.apply(PairRDDFunctions.scala:1009)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$12.apply(PairRDDFunctions.scala:979)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:61)\n at org.apache.spark.scheduler.Task.run(Task.scala:64)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:203)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{noformat}\n\nI checked the output directory and indeed {{part-r-00195}} is missing.\n\nOur speculative execution configuration was:\n\n{noformat}\nspark.speculation true\nspark.speculation.interval 1000\nspark.speculation.quantile 0.90\nspark.speculation.multiplier 2\n{noformat}\n\nWe've been running production code with this setup for almost a year. Until now we hadn't checked for missing files. Such a painful bug.","created":"2015-07-15T14:01:03.165+0000"},{"body":"I think that this is going to be a complex issue to investigate / fix, so I'd like to break down that task into a series of smaller subtasks. Here's my initial proposal (each of these could be done as a separate JIRA / patch / TODO so that we make incremental progress):\n\n- Add automated post-save-action checks to assert that the expected number of output files are present. This will make debugging much easier. There are some corner-cases to address related to S3 and _temporary / _SUCCESS files, but I'm sure we can come up with a solution that addresses them. The error message raised by this assert could have a pointer to this issue and could suggest disabling speculation as a workaround.\n- Write a standalone program which is capable of reproducing the most recently reported set of symptoms for this bug. The symptoms that originally prompted this ticket were cases where a speculative / duplicated task would continue running after the original job had completed, then fail to commit and delete the destination file. It's possible that this issue has resurfaced / wasn't properly fixed, but it's also possible that there is a different race which wasn't addressed by the original patch.\n -- One subtask of this may be logging improvements: we might not be able to come up with a reasonably minimal reproduction until we have more granular logs to help us better pinpoint the race.\n -- If this is non-deterministic, then we can rig our program to try to expose the race, then have it re-try itself hundreds of times in a loop and stop once a failure occurs.\n -- Given the inherent non-determinism, external dependencies, and slow nature of this type of test, I don't think that this can be an automated test in Jenkins. However, I think that the code for this test should live in the Spark codebase. We should create a pre-release manual QA checklist and make sure that it includes a task to manually run this test on EC2.\n- If we gain a better understanding of the races / problems that lead to this bug, we may be able to mock / stub / simulate / interpose in a way that lets us write a unit test to reproduce this bug. If that's possible, then we should do it in order to have a fast-to-run test that can run regularly.\n- Fix the underlying bug.","created":"2015-07-15T16:17:10.026+0000"},{"body":"Thanks, Josh! I wonder when the post-save-action check you suggest would run. Based on the above log snippet I think it's quite likely that at the point when the stage finishes the files are all there. I suspect it's the speculative task which fails after the stage has finished that deletes the file. It may be hard to check at the right point in time.\n\nI'll try to find a smaller reproduction. First it would be great if I could reproduce on my local machine instead of starting EC2 clusters. Then I just need to dig out the key operations from our -entangled mess of a- _highly sophisticated_ codebase.\n\nOne more thing that occurs to me is that perhaps there should never be a line of code that deletes an output file. I haven't had a chance to dig into the code yet, and I'm sure there is a reason for it, but perhaps the same goal could be accomplished without deleting output files. What do you think?","created":"2015-07-15T17:17:55.694+0000"},{"body":"That's a good point about the post-save-action check: code internal to Spark can only perform the check right before we return control back to the user so it's possible that the check will say that everything's fine even if the files are later deleted. This check still might have some value, but it's certainly prone to false-negatives and wouldn't be able to catch all manifestations of this bug.\n\nIn the original issue that prompted this patch, the deletion of the output file occurred inside of Hadoop code: if the renaming of the temporary file to the final output file fails, Hadoop's FileOutputCommitter will attempt to delete the final output file before re-attempting the rename (this is explained in a bit more detail, including a code walkthrough, in the description of this JIRA ticket).","created":"2015-07-17T15:22:50.478+0000"},{"body":"there is blog post http://tech.grammarly.com/blog/posts/Petabyte-Scale-Text-Processing-with-Spark.html with a link to gist by Aaron Davidson: https://gist.github.com/aarondav/c513916e72101bbe14ec which mentions in comments that when using speculation it should be DirectOutputCommitter\n\nmight be somebody will find it useful (imho, it should be in documentation)\n","created":"2015-08-23T09:54:39.343+0000"},{"body":"I'm clearing \"backport-needed\" since it's virtually certain that there will be no more 1.2.x or earlier releases, and so the fix that was committed won't go back further at this point.\n\nIs it something to leave open pending the ongoing conversation here? sounds like there may be more to the fix? ","created":"2015-09-12T13:02:00.559+0000"},{"body":"What about splitting the issue into HDFS commits (interesting its mostly EC2 reports), which is probably fixed, and eventually consistent object stores (S3 on the Apache Hadoop releases (but not amazon's own), which need a different committer (no rename), better checks for file presence (direct stat() is more reliable than a directory listing), and some dedicated test suite which could be targeted straight at s3 —yet still runnable remotely by someone (not jenkins) from their own desktop & build servers. That's essentially what we do in core hadoop to qualify the object stores' base API compatibility.","created":"2015-09-13T12:01:18.232+0000"}],"conversations":[{"body":"When speculative execution is enabled ({{spark.speculation=true}}), jobs that save output files may report that they have completed successfully even though some output partitions written by speculative tasks may be missing.\n\nh3. Reproduction\n\nThis symptom was reported to me by a Spark user and I've been doing my own investigation to try to come up with an in-house reproduction.\n\nI'm still working on a reliable local reproduction for this issue, which is a little tricky because Spark won't schedule speculated tasks on the same host as the original task, so you need an actual (or containerized) multi-host cluster to test speculation. Here's a simple reproduction of some of the symptoms on EC2, which can be run in {{spark-shell}} with {{--conf spark.speculation=true}}:\n\n{code}\n // Rig a job such that all but one of the tasks complete instantly\n // and one task runs for 20 seconds on its first attempt and instantly\n // on its second attempt:\n val numTasks = 100\n sc.parallelize(1 to numTasks, numTasks).repartition(2).mapPartitionsWithContext { case (ctx, iter) =>\n if (ctx.partitionId == 0) { // If this is the one task that should run really slow\n if (ctx.attemptId == 0) { // If this is the first attempt, run slow\n Thread.sleep(20 * 1000)\n }\n }\n iter\n }.map(x => (x, x)).saveAsTextFile(\"/test4\")\n{code}\n\nWhen I run this, I end up with a job that completes quickly (due to speculation) but reports failures from the speculated task:\n\n{code}\n[...]\n14/12/11 01:41:13 INFO scheduler.TaskSetManager: Finished task 37.1 in stage 3.0 (TID 411) in 131 ms on ip-172-31-8-164.us-west-2.compute.internal (100/100)\n14/12/11 01:41:13 INFO scheduler.DAGScheduler: Stage 3 (saveAsTextFile at :22) finished in 0.856 s\n14/12/11 01:41:13 INFO spark.SparkContext: Job finished: saveAsTextFile at :22, took 0.885438374 s\n14/12/11 01:41:13 INFO scheduler.TaskSetManager: Ignoring task-finished event for 70.1 in stage 3.0 because task 70 has already completed successfully\n\nscala> 14/12/11 01:41:13 WARN scheduler.TaskSetManager: Lost task 49.1 in stage 3.0 (TID 413, ip-172-31-8-164.us-west-2.compute.internal): java.io.IOException: Failed to save output of task: attempt_201412110141_0003_m_000049_413\n org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:160)\n org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:172)\n org.apache.hadoop.mapred.FileOutputCommitter.commitTask(FileOutputCommitter.java:132)\n org.apache.spark.SparkHadoopWriter.commit(SparkHadoopWriter.scala:109)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:991)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:974)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:62)\n org.apache.spark.scheduler.Task.run(Task.scala:54)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:177)\n java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n java.lang.Thread.run(Thread.java:745)\n{code}\n\nOne interesting thing to note about this stack trace: if we look at {{FileOutputCommitter.java:160}} ([link|http://grepcode.com/file/repository.cloudera.com/content/repositories/releases/org.apache.hadoop/hadoop-core/2.5.0-mr1-cdh5.2.0/org/apache/hadoop/mapred/FileOutputCommitter.java#160]), this point in the execution seems to correspond to a case where a task completes, attempts to commit its output, fails for some reason, then deletes the destination file, tries again, and fails:\n\n{code}\n if (fs.isFile(taskOutput)) {\n152 Path finalOutputPath = getFinalPath(jobOutputDir, taskOutput, \n153 getTempTaskOutputPath(context));\n154 if (!fs.rename(taskOutput, finalOutputPath)) {\n155 if (!fs.delete(finalOutputPath, true)) {\n156 throw new IOException(\"Failed to delete earlier output of task: \" + \n157 attemptId);\n158 }\n159 if (!fs.rename(taskOutput, finalOutputPath)) {\n160 throw new IOException(\"Failed to save output of task: \" + \n161 \t\t attemptId);\n162 }\n163 }\n{code}\n\nThis could explain why the output file is missing: the second copy of the task keeps running after the job completes and deletes the output written by the other task after failing to commit its own copy of the output.\n\nThere are still a few open questions about how exactly we get into this scenario:\n\n*Why is the second copy of the task allowed to commit its output after the other task / the job has successfully completed?*\n\nTo check whether a task's temporary output should be committed, SparkHadoopWriter calls {{FileOutputCommitter.needsTaskCommit()}}, which returns {{true}} if the tasks's temporary output exists ([link|http://grepcode.com/file/repository.cloudera.com/content/repositories/releases/org.apache.hadoop/hadoop-core/2.5.0-mr1-cdh5.2.0/org/apache/hadoop/mapred/FileOutputCommitter.java#206]). Tihs does not seem to check whether the destination already exists. This means that {{needsTaskCommit}} can return {{true}} for speculative tasks.\n\n*Why does the rename fail?*\n\nI think that what's happening is that the temporary task output files are being deleted once the job has completed, which is causing the {{rename}} to fail because {{FileOutputCommitter.commitTask}} doesn't seem to guard against missing output files.\n\nI'm not sure about this, though, since the stack trace seems to imply that the temporary output file existed. Maybe the filesystem methods are returning stale metadata? Maybe there's a race? I think a race condition seems pretty unlikely, since the time-scale at which it would have to happen doesn't sync up with the scale of the timestamps that I saw in the user report.\n\nh3. Possible Fixes:\n\nThe root problem here might be that speculative copies of tasks are somehow allowed to commit their output. We might be able to fix this by centralizing the \"should this task commit its output\" decision at the driver.\n\n(I have more concrete suggestions of how to do this; to be posted soon)","from":"reporter","subject":"Missing output partitions after job completes with speculative execution"},{"body":"Hey Josh, \n\nI have been playing around with your repro above and I think I can consistently trigger the bad behavior by just tweaking the value of {{spark.speculation.multiplier}} and {{spark.speculation.quantile}}.\n\nI set the {{multiplier}} to be 1 and the {{quantile}} to 0.01 so that only 1% of tasks have to finish before any task that takes longer than those 1% of tasks should speculate. \nAs expected, I see a lot of tasks getting speculated. \nAfter running the repro about 5 times, I have seen 2 errors (stack traces at the bottom and the full run from the REPL is attached with this comment). \n\nOne thing I do notice is that the part-00000 associated with Stage 1 was always where I expected it to be in HDFS, and all lines were present (checked using a {{wc -l}})\n\n\n{code}\nscala> 15/01/07 13:44:26 WARN scheduler.TaskSetManager: Lost task 0.1 in stage 0.0 (TID 119, ): java.io.IOException: The temporary job-output directory hdfs://:8020/test6/_temporary doesn't exist!\n org.apache.hadoop.mapred.FileOutputCommitter.getWorkPath(FileOutputCommitter.java:250)\n org.apache.hadoop.mapred.FileOutputFormat.getTaskOutputPath(FileOutputFormat.java:240)\n org.apache.hadoop.mapred.TextOutputFormat.getRecordWriter(TextOutputFormat.java:116)\n org.apache.spark.SparkHadoopWriter.open(SparkHadoopWriter.scala:89)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:980)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:974)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:62)\n org.apache.spark.scheduler.Task.run(Task.scala:54)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:178)\n java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n java.lang.Thread.run(Thread.java:745)\n{code}\n\n{code}\n15/01/07 15:17:39 WARN scheduler.TaskSetManager: Lost task 0.1 in stage 0.0 (TID 120, ): org.apache.hadoop.ipc.RemoteException: No lease on /test7/_temporary/_attempt_201501071517_0000_m_000000_120/part-00000: File does not exist. Holder DFSClient_NONMAPREDUCE_-469253416_73 does not have any open files.\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.checkLease(FSNamesystem.java:2609)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.analyzeFileState(FSNamesystem.java:2426)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getAdditionalBlock(FSNamesystem.java:2339)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.addBlock(NameNodeRpcServer.java:501)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.addBlock(ClientNamenodeProtocolServerSideTranslatorPB.java:299)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java:44954)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:453)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:1002)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:1752)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:1748)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:415)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1438)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:1746)\n\n org.apache.hadoop.ipc.Client.call(Client.java:1238)\n org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:202)\n com.sun.proxy.$Proxy9.addBlock(Unknown Source)\n sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n java.lang.reflect.Method.invoke(Method.java:606)\n org.apache.hadoop.io.retry.RetryInvocationHandler.invokeMethod(RetryInvocationHandler.java:164)\n org.apache.hadoop.io.retry.RetryInvocationHandler.invoke(RetryInvocationHandler.java:83)\n com.sun.proxy.$Proxy9.addBlock(Unknown Source)\n org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolTranslatorPB.addBlock(ClientNamenodeProtocolTranslatorPB.java:291)\n org.apache.hadoop.hdfs.DFSOutputStream$DataStreamer.locateFollowingBlock(DFSOutputStream.java:1177)\n org.apache.hadoop.hdfs.DFSOutputStream$DataStreamer.nextBlockOutputStream(DFSOutputStream.java:1030)\n org.apache.hadoop.hdfs.DFSOutputStream$DataStreamer.run(DFSOutputStream.java:488)\n{code}\n","from":"developer"},{"body":"Hey Zach,\n\nFrom my last round of attempts (maybe a week ago), I was able to reproduce some sort of \"No lease on ....\" error but not the missing output partitions symptom. I'd still like to try to find a reproduction for the missing files but this could be tricky if it involves a very quick race condition; I don't think that's the case, though, since it's pretty unlikely that the speculated and original task are finishing at the exact same time.","from":"developer"},{"body":"Hey Josh,\n\nI was able to reproduce the missing file using the speculation settings in my previous comment:\n\n{code}\nscala> 15/01/09 18:33:28 WARN scheduler.TaskSetManager: Lost task 42.1 in stage 0.0 (TID 113, ): java.io.IOException: Failed to save output of task: attempt_201501091833_0000_m_000042_113\n org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:160)\n org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:172)\n org.apache.hadoop.mapred.FileOutputCommitter.commitTask(FileOutputCommitter.java:132)\n org.apache.spark.SparkHadoopWriter.commit(SparkHadoopWriter.scala:109)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:991)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:974)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:62)\n org.apache.spark.scheduler.Task.run(Task.scala:54)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:178)\n java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n java.lang.Thread.run(Thread.java:745)\n15/01/09 18:33:47 WARN scheduler.TaskSetManager: Lost task 0.0 in stage 0.0 (TID 0, ): java.io.IOException: The temporary job-output directory hdfs://:8020/test2/_temporary doesn't exist!\n org.apache.hadoop.mapred.FileOutputCommitter.getWorkPath(FileOutputCommitter.java:250)\n org.apache.hadoop.mapred.FileOutputFormat.getTaskOutputPath(FileOutputFormat.java:240)\n org.apache.hadoop.mapred.TextOutputFormat.getRecordWriter(TextOutputFormat.java:116)\n org.apache.spark.SparkHadoopWriter.open(SparkHadoopWriter.scala:89)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:980)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:974)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:62)\n org.apache.spark.scheduler.Task.run(Task.scala:54)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:178)\n java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n java.lang.Thread.run(Thread.java:745)\n{code}\n\nNotice here that there are only 99 part files and part-00042 is missing (as seen in the stacktrace above)\n{code}\n $ hadoop fs -ls /test2 | grep part | wc -l\n99\n $ hadoop fs -ls /test2 | grep part-0004\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00040\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00041\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00043\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00044\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00045\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00046\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00047\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00048\n-rw-r--r-- 3 supergroup 8 2015-01-09 18:33 /test2/part-00049\n{code}\n\n","from":"developer"},{"body":"For clarity, here is the scala code I used in the REPL:\n{code}\nscala> val numTasks = 100\nnumTasks: Int = 100\n\nscala> sc.parallelize(1 to numTasks, numTasks).mapPartitionsWithContext { case (ctx, iter) =>\n | if (ctx.partitionId == 0) { // If this is the one task that should run really slow\n | if (ctx.attemptId == 0) { // If this is the first attempt, run slow\n | Thread.sleep(20 * 1000)\n | }\n | }\n | iter\n | }.map(x => (x, x)).saveAsTextFile(\"/test2\")\n{code}","from":"developer"},{"body":"I think that part of the reproduction issues that I had might have been due to {{attemptId}} returning a unique task attempt ID rather than the attempt number, meaning that only the _first_ run of that test in the REPL would be capable of uncovering the bug.\n\nSee https://github.com/apache/spark/pull/3849 / SPARK-4014 for more context. I'm going to try to merge that patch today, which will let me write a reliable regression test.","from":"developer"},{"body":"I'm not sure that SparkHadoopWriter's use of FileOutputCommitter properly obeys the OutputCommitter contracts in Hadoop. According to the [OutputCommitter Javadoc|https://hadoop.apache.org/docs/current/api/org/apache/hadoop/mapreduce/OutputCommitter.html]\n\n{quote}\nThe methods in this class can be called from several different processes and from several different contexts. It is important to know which process and which context each is called from. Each method should be marked accordingly in its documentation. It is also important to note that not all methods are guaranteed to be called once and only once. If a method is not guaranteed to have this property the output committer needs to handle this appropriately. Also note it will only be in rare situations where they may be called multiple times for the same task.\n{quote}\n\nBased on the documentation, `needsTaskCommit` \" is called from each individual task's process that will output to HDFS, and it is called just for that task.\", so it seems like it should be safe to call this from SparkHadoopWriter.\n\nHowever, maybe we're misusing the `commitTask` method:\n\n{quote}\nIf needsTaskCommit(TaskAttemptContext) returns true and this task is the task that the AM determines finished first, this method is called to commit an individual task's output. This is to mark that tasks output as complete, as commitJob(JobContext) will also be called later on if the entire job finished successfully. This is called from a task's process. This may be called multiple times for the same task, but different task attempts. It should be very rare for this to be called multiple times and requires odd networking failures to make this happen. In the future the Hadoop framework may eliminate this race. \n{quote}\n\nI think that we're missing the \"this task is the task that the AM determines finished first\" part of the equation here. If `needsTaskCommit` is false, then we definitely shouldn't commit (e.g. if it's an original task that lost to a speculated copy), but if it's true then I don't think it's safe to commit; we need some central authority to pick a winner.\n\nLet's see how Hadoop does things, working backwards from actual calls of `commitTask` to see whether they're guarded by some coordination through the AM. It looks like `OutputCommitter` is part of the `mapred` API, so I'll only look at classes in that package:\n\nIn `Task.java`, `committer.commitTask` is only performed after checking `canCommit` through `TaskUmbilicalProtocol`: https://github.com/apache/hadoop/blob/a655973e781caf662b360c96e0fa3f5a873cf676/hadoop-mapreduce-project/hadoop-mapreduce-client/hadoop-mapreduce-client-core/src/main/java/org/apache/hadoop/mapred/Task.java#L1185. According to the Javadocs for TaskAttemptListenerImpl.canCommit (the actual concrete implementation of this method):\n\n{code}\n /**\n * Child checking whether it can commit.\n * \n *
\n * Commit is a two-phased protocol. First the attempt informs the\n * ApplicationMaster that it is\n * {@link #commitPending(TaskAttemptID, TaskStatus)}. Then it repeatedly polls\n * the ApplicationMaster whether it {@link #canCommit(TaskAttemptID)} This is\n * a legacy from the centralized commit protocol handling by the JobTracker.\n */\n @Override\n public boolean canCommit(TaskAttemptID taskAttemptID) throws IOException {\n{code}\n\nThis ends up delegating to `Task.canCommit()`:\n\n{code}\n /**\n * Can the output of the taskAttempt be committed. Note that once the task\n * gives a go for a commit, further canCommit requests from any other attempts\n * should return false.\n * \n * @param taskAttemptID\n * @return whether the attempt's output can be committed or not.\n */\n boolean canCommit(TaskAttemptId taskAttemptID);\n{code}\n\nThere's a bunch of tricky logic that involves communication with the AM (see AttemptCommitPendingTransition and the other transitions in TaskImpl), but it looks like the gist is that the \"winner\" is picked by the AM through some central coordination process. \n\nSo, it looks like the right fix is to implement these same state transitions ourselves. It would be nice if there was a clean way to do this that could be easily backported to maintenance branches. ","from":"developer"},{"body":"User 'JoshRosen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4066","from":"developer"},{"body":"User 'mccheah' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4155","from":"developer"},{"body":"This issue is _really_ hard to reproduce, but I managed to trigger the original bug as part of the testing for my patch. Here's what I ran:\n\n{code}\n~/spark-1.3.0-SNAPSHOT-bin-1.0.4/bin/spark-shell --conf spark.speculation.multiplier=1 --conf spark.speculation.quantile=0.01 --conf spark.speculation=true --conf spark.hadoop.outputCommitCoordination.enabled=false\n{code}\n\n{code}\nval numTasks = 100\nval numTrials = 100\nval outputPath = \"/output-committer-bug-\"\nval sleepDuration = 1000\n\nfor (trial <- 0 to (numTrials - 1)) {\n val outputLocation = outputPath + trial\n sc.parallelize(1 to numTasks, numTasks).mapPartitionsWithContext { case (ctx, iter) =>\n if (ctx.partitionId % 5 == 0) {\n if (ctx.attemptNumber == 0) { // If this is the first attempt, run slow\n Thread.sleep(sleepDuration)\n }\n }\n iter\n }.map(identity).saveAsTextFile(outputLocation)\n Thread.sleep(sleepDuration * 2)\n println(\"TESTING OUTPUT OF TRIAL \" + trial)\n val savedData = sc.textFile(outputLocation).map(_.toInt).collect()\n if (savedData.toSet != (1 to numTasks).toSet) {\n println(\"MISSING: \" + ((1 to numTasks).toSet -- savedData.toSet))\n assert(false)\n }\n println(\"-\" * 80)\n}\n{code}\n\nIt took 22 runs until I actually observed missing output partitions (several of the earlier runs threw spurious exceptions and didn't have missing outputs):\n\n{code}\n[...]\n15/02/10 22:17:21 INFO scheduler.DAGScheduler: Job 66 finished: saveAsTextFile at :39, took 2.479592 s\n15/02/10 22:17:21 WARN scheduler.TaskSetManager: Lost task 75.0 in stage 66.0 (TID 6861, ip-172-31-1-124.us-west-2.compute.internal): java.io.IOException: Failed to save output of task: attempt_201502102217_0066_m_000075_6861\n at org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:160)\n at org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:172)\n at org.apache.hadoop.mapred.FileOutputCommitter.commitTask(FileOutputCommitter.java:132)\n at org.apache.spark.SparkHadoopWriter.performCommit$1(SparkHadoopWriter.scala:113)\n at org.apache.spark.SparkHadoopWriter.commit(SparkHadoopWriter.scala:150)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:1082)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:1059)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:61)\n at org.apache.spark.scheduler.Task.run(Task.scala:64)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:197)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n\n15/02/10 22:17:21 WARN scheduler.TaskSetManager: Lost task 80.0 in stage 66.0 (TID 6866, ip-172-31-11-151.us-west-2.compute.internal): java.io.IOException: The temporary job-output directory hdfs://ec2-54-213-142-80.us-west-2.compute.amazonaws.com:9000/output-committer-bug-22/_temporary doesn't exist!\n at org.apache.hadoop.mapred.FileOutputCommitter.getWorkPath(FileOutputCommitter.java:250)\n at org.apache.hadoop.mapred.FileOutputFormat.getTaskOutputPath(FileOutputFormat.java:244)\n at org.apache.hadoop.mapred.TextOutputFormat.getRecordWriter(TextOutputFormat.java:116)\n at org.apache.spark.SparkHadoopWriter.open(SparkHadoopWriter.scala:91)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:1068)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:1059)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:61)\n at org.apache.spark.scheduler.Task.run(Task.scala:64)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:197)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n\n15/02/10 22:17:21 INFO scheduler.TaskSetManager: Lost task 85.0 in stage 66.0 (TID 6871) on executor ip-172-31-1-124.us-west-2.compute.internal: java.io.IOException (The temporary job-output directory hdfs://ec2-54-213-142-80.us-west-2.compute.amazonaws.com:9000/output-committer-bug-22/_temporary doesn't exist!) [duplicate 1]\n15/02/10 22:17:21 INFO scheduler.TaskSetManager: Lost task 90.0 in stage 66.0 (TID 6876) on executor ip-172-31-1-124.us-west-2.compute.internal: java.io.IOException (The temporary job-output directory hdfs://ec2-54-213-142-80.us-west-2.compute.amazonaws.com:9000/output-committer-bug-22/_temporary doesn't exist!) [duplicate 2]\n15/02/10 22:17:21 INFO scheduler.TaskSetManager: Lost task 95.0 in stage 66.0 (TID 6881) on executor ip-172-31-11-151.us-west-2.compute.internal: java.io.IOException (The temporary job-output directory hdfs://ec2-54-213-142-80.us-west-2.compute.amazonaws.com:9000/output-committer-bug-22/_temporary doesn't exist!) [duplicate 3]\n15/02/10 22:17:21 INFO scheduler.TaskSchedulerImpl: Removed TaskSet 66.0, whose tasks have all completed, from pool\nTESTING OUTPUT OF TRIAL 22\n[...]\n15/02/10 22:17:23 INFO scheduler.TaskSchedulerImpl: Removed TaskSet 67.0, whose tasks have all completed, from pool\n15/02/10 22:17:23 INFO scheduler.DAGScheduler: Stage 67 (collect at :42) finished in 0.158 s\n15/02/10 22:17:23 INFO scheduler.DAGScheduler: Job 67 finished: collect at :42, took 0.165580 s\nMISSING: Set(76)\njava.lang.AssertionError: assertion failed\n[...]\n{code}\n\nAnd I confirmed that it's missing in HDFS:\n\n{code}\n~/ephemeral-hdfs/bin/hadoop fs -ls /output-committer-bug-22 | grep part | wc -l\nWarning: $HADOOP_HOME is deprecated.\n\n99\n{code}\n\nTo test my patch, I'm going to remove the flag that disables it and run this for a huge number of trials to ensure that we don't hit the missing output bug.","from":"developer"},{"body":"This is really great work [~joshrosen]! I really appreciate the effort you're putting into getting this one figured out since these kind of non-deterministic bugs are the most painful for both users and devs to figure out.","from":"developer"},{"body":"Thanks for picking this up [~joshrosen], I see that the hotly-debated patch has finally merged in!\n\nI'm a bit backed up but there will be others from my side that will be testing this. We'll comment on the bug once we have verified the issue is fixed on our side.","from":"developer"},{"body":"Could this happen very very rarely when not using speculative execution?\nOnce in a long while, I have a situation where the OutputCommitter says it wrote the file successfully, but the output file doesn't appear there.","from":"developer"},{"body":"[~romi-totango] what filesystem are you writing to? HDFS / S3 / other ? If this was HDFS then I think you may have a bug, but if it was S3 then there are loose eventual consistency guarantees that might be surprising you.","from":"developer"},{"body":"[~aash] I'm writing to S3. I'm aware of eventual consistency issues, but is it possible that 99.999% files are written and available immediately, and then one file doesn't appear for very long? (couldn't afford to wait until it was available so I ran the process again and it wrote the file immediately).\nStill isn't it somehow possibly related to this bug or similar, that Spark's output writing mechanism thinks it wrote the file, but actually it did not?","from":"developer"},{"body":"That sounds exactly like what I'd expect to happen in S3. The eventual consistency bug is that S3 reports something is written, Spark then believes it has written the file, but when you go to access it the file isn't there yet (at least in the us standard region. If you use another region, there is read-after-write consistency).\n\nIn short, it doesn't sound like you're encountering this particular issue with speculative execution (SPARK-4879), just normal S3 consistency trickiness.","from":"developer"},{"body":"\nWith 1.3 RC, we are still seeing this issue.","from":"developer"},{"body":"Just to add - we are seeing this on hdfs.\nOut of the 75k output files to be written, about 53k did not get moved.\n\nBut on waiting for a really long time (about 20 minutes) all files eventually got moved.\nIs this something to do with FileOutputCommitter itself or side-effect of the patch committed ?\n\nWe have speculative execution turned on; even though the saveAsHadoopFile taskset completes (as per logs), the call itself does not return until entire 'move' is done - just that it took more time to commit the output than to generate it in first place !","from":"developer"},{"body":"[~mridulm80], when you say that \"the call itself does not return until entire 'move' is done\", do you mean that the {{saveAsHadoopFile()}} call blocks in {{saveAsHadoopDataset()}}'s final {{ writer.commitJob()}} call? This ticket addresses a bug where output partitions are missing _after_ control returns from the {{saveAsHadoop*()}} call.\n\nIn the issue addressed by this ticket, partitions were deleted after the job and save action had completed and output remained missing. It sounds like you might be describing a different issue where jobs take a long time to commit. For your workload, is the behavior that you are observing a performance regression in Spark 1.3? If you'd like to run 1.3 without the effects of my patch, just set {{spark.hadoop.outputCommitCoordination.enabled=false}} in your SparkConf to bypass the new output commit coordination logic.","from":"developer"},{"body":"[~joshrosen] The former - the call blocks on saveAsHadoopFile (the very next line was sc.stop() and then exit).\n\nWe cannot use 'spark.hadoop.outputCommitCoordination.enabled=false' since we need to have speculative execution on (turning it off drastically affects job runtime) and disabling commit coordination causes other issues (as reported in this jira).\n\nI was not sure if \n- the dramatically increased to commit was due to the PR from this jira.\n- is orthogonal to this jira and an issue with FileOutputCommitter (as was seen which searching for cause online).\n\nGiven the central coordination blocking on master, I wanted to eliminate causes from our end before digging into FIleOutputCommitter and/or Namenode code.","from":"developer"},{"body":"Can you perhaps jstack or profile the driver and the executors to see if they're blocked in OutputCommitCoordinator logic?","from":"developer"},{"body":"Hi, is there any updates for this issue?","from":"developer"},{"body":"I'm experiencing this issue. Sometimes rdd with 4 partitions is written with 3 parts and _SUCCESS marker is there.","from":"developer"},{"body":"I have a fairly reliable reproduction with Spark 1.4.0 and HDFS. I'm running on 10 EC2 m3.2xlarge instances using the ephemeral HDFS. If {{spark.speculation}} is true, I get hit by this 50% of the time or more. It's a fairly complex workload, not something you can test in a {{spark-shell}}. What I saw was that I saved a 400-partition RDD with {{saveAsNewAPIHadoopFile}} (which returned without error) and when I tried to read it back, the files for partitions 323 and 324 were missing. (In the case that I took a closer look at.) I don't have the logs at hand now, but it's like you describe I think ({{Failed to save output of task}}). I can add them later if it would be useful.\n\nI turned off {{spark.speculation}} and haven't seen the issue since.\n\nIs there anything I could do to help debug this issue?","from":"developer"},{"body":"I wonder if this issue is serious enough to note in the documentation. What do you think about adding a big fat warning for speculative execution until it is fixed? \"Enabling speculative execution may lead to missing output files\"? Or perhaps add a verification pass that checks if all the outputs are present and raises an exception if not.\n\nSilently dropping output files is a horrible bug. We've been debugging a somewhat mythological data corruption issue for about a month, and now we realize that this issue (SPARK-4879) is a very plausible explanation. We have never been able to reproduce it, but we have a log file, and it shows a speculative task for a {{saveAsNewAPIHadoopFile}} stage.","from":"developer"},{"body":"[~darabos], do you think that this issue might have been resolved in an earlier Spark version but inadvertently broken in the upgrade to 1.4.0? If you have an easy reproduction, it might be helpful to see whether the problem occurs on 1.3.1.","from":"developer"},{"body":"Good idea! I'll try with 1.3.1 next week.","from":"developer"},{"body":"I've managed to reproduce on Spark 1.3.1 too (pre-built for Hadoop 1). I ran on EC2 with the {{spark-ec2}} script and used the ephemeral HDFS. I used a 5-machine cluster and repeatedly ran a complex test suite for about 30 minutes until the error was triggered. Here are the relevant logs:\n\n{noformat}\nI2015-07-15 13:45:19,954 TaskSetManager:[task-result-getter-2] Finished task 198.0 in stage 320.0 (TID 13290) in 568 ms on ip-10-153-188-224.ec2.internal (195/200)\nI2015-07-15 13:45:21,174 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Marking task 197 in stage 320.0 (on ip-10-231-214-6.ec2.internal) as speculatable because it ran more than 1240 ms\nI2015-07-15 13:45:21,174 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Marking task 176 in stage 320.0 (on ip-10-231-214-6.ec2.internal) as speculatable because it ran more than 1240 ms\nI2015-07-15 13:45:21,174 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Marking task 194 in stage 320.0 (on ip-10-231-214-6.ec2.internal) as speculatable because it ran more than 1240 ms\nI2015-07-15 13:45:21,174 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Marking task 196 in stage 320.0 (on ip-10-231-214-6.ec2.internal) as speculatable because it ran more than 1240 ms\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Marking task 195 in stage 320.0 (on ip-10-231-214-6.ec2.internal) as speculatable because it ran more than 1240 ms\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Starting task 197.1 in stage 320.0 (TID 13292, ip-10-153-188-224.ec2.internal, PROCESS_LOCAL, 1612 bytes)\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Starting task 176.1 in stage 320.0 (TID 13293, ip-10-63-27-248.ec2.internal, PROCESS_LOCAL, 1612 bytes)\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Starting task 194.1 in stage 320.0 (TID 13294, ip-10-154-1-239.ec2.internal, PROCESS_LOCAL, 1612 bytes)\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Starting task 195.1 in stage 320.0 (TID 13295, ip-10-228-67-34.ec2.internal, PROCESS_LOCAL, 1612 bytes)\nI2015-07-15 13:45:21,175 TaskSetManager:[sparkDriver-akka.actor.default-dispatcher-2] Starting task 196.1 in stage 320.0 (TID 13296, ip-10-153-188-224.ec2.internal, PROCESS_LOCAL, 1612 bytes)\nI2015-07-15 13:45:21,402 TaskSetManager:[task-result-getter-3] Finished task 176.0 in stage 320.0 (TID 13268) in 2223 ms on ip-10-231-214-6.ec2.internal (196/200)\nI2015-07-15 13:45:21,445 TaskSetManager:[task-result-getter-0] Finished task 195.0 in stage 320.0 (TID 13287) in 2113 ms on ip-10-231-214-6.ec2.internal (197/200)\nI2015-07-15 13:45:21,461 TaskSetManager:[task-result-getter-1] Finished task 196.1 in stage 320.0 (TID 13296) in 285 ms on ip-10-153-188-224.ec2.internal (198/200)\nI2015-07-15 13:45:21,464 TaskSetManager:[task-result-getter-2] Finished task 194.1 in stage 320.0 (TID 13294) in 287 ms on ip-10-154-1-239.ec2.internal (199/200)\nI2015-07-15 13:45:21,465 TaskSetManager:[task-result-getter-3] Ignoring task-finished event for 176.1 in stage 320.0 because task 176 has already completed successfully\nI2015-07-15 13:45:21,468 TaskSetManager:[task-result-getter-0] Finished task 197.1 in stage 320.0 (TID 13292) in 292 ms on ip-10-153-188-224.ec2.internal (200/200)\nI2015-07-15 13:45:21,468 DAGScheduler:[dag-scheduler-event-loop] Stage 320 (saveAsNewAPIHadoopFile at HadoopFile.scala:208) finished in 4.802 s\nI2015-07-15 13:45:21,468 DAGScheduler:[DataManager-5] Job 46 finished: saveAsNewAPIHadoopFile at HadoopFile.scala:208, took 4.836626 s\nW2015-07-15 13:45:21,478 TaskSetManager:[task-result-getter-1] Lost task 195.1 in stage 320.0 (TID 13295, ip-10-228-67-34.ec2.internal): java.io.IOException: Failed to save output of task: attempt_201507151345_0628_r_000195_1\n at org.apache.hadoop.mapreduce.lib.output.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:203)\n at org.apache.hadoop.mapreduce.lib.output.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:214)\n at org.apache.hadoop.mapreduce.lib.output.FileOutputCommitter.commitTask(FileOutputCommitter.java:167)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$12.apply(PairRDDFunctions.scala:1009)\n at org.apache.spark.rdd.PairRDDFunctions$$anonfun$12.apply(PairRDDFunctions.scala:979)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:61)\n at org.apache.spark.scheduler.Task.run(Task.scala:64)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:203)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n at java.lang.Thread.run(Thread.java:745)\n{noformat}\n\nI checked the output directory and indeed {{part-r-00195}} is missing.\n\nOur speculative execution configuration was:\n\n{noformat}\nspark.speculation true\nspark.speculation.interval 1000\nspark.speculation.quantile 0.90\nspark.speculation.multiplier 2\n{noformat}\n\nWe've been running production code with this setup for almost a year. Until now we hadn't checked for missing files. Such a painful bug.","from":"developer"},{"body":"I think that this is going to be a complex issue to investigate / fix, so I'd like to break down that task into a series of smaller subtasks. Here's my initial proposal (each of these could be done as a separate JIRA / patch / TODO so that we make incremental progress):\n\n- Add automated post-save-action checks to assert that the expected number of output files are present. This will make debugging much easier. There are some corner-cases to address related to S3 and _temporary / _SUCCESS files, but I'm sure we can come up with a solution that addresses them. The error message raised by this assert could have a pointer to this issue and could suggest disabling speculation as a workaround.\n- Write a standalone program which is capable of reproducing the most recently reported set of symptoms for this bug. The symptoms that originally prompted this ticket were cases where a speculative / duplicated task would continue running after the original job had completed, then fail to commit and delete the destination file. It's possible that this issue has resurfaced / wasn't properly fixed, but it's also possible that there is a different race which wasn't addressed by the original patch.\n -- One subtask of this may be logging improvements: we might not be able to come up with a reasonably minimal reproduction until we have more granular logs to help us better pinpoint the race.\n -- If this is non-deterministic, then we can rig our program to try to expose the race, then have it re-try itself hundreds of times in a loop and stop once a failure occurs.\n -- Given the inherent non-determinism, external dependencies, and slow nature of this type of test, I don't think that this can be an automated test in Jenkins. However, I think that the code for this test should live in the Spark codebase. We should create a pre-release manual QA checklist and make sure that it includes a task to manually run this test on EC2.\n- If we gain a better understanding of the races / problems that lead to this bug, we may be able to mock / stub / simulate / interpose in a way that lets us write a unit test to reproduce this bug. If that's possible, then we should do it in order to have a fast-to-run test that can run regularly.\n- Fix the underlying bug.","from":"developer"},{"body":"Thanks, Josh! I wonder when the post-save-action check you suggest would run. Based on the above log snippet I think it's quite likely that at the point when the stage finishes the files are all there. I suspect it's the speculative task which fails after the stage has finished that deletes the file. It may be hard to check at the right point in time.\n\nI'll try to find a smaller reproduction. First it would be great if I could reproduce on my local machine instead of starting EC2 clusters. Then I just need to dig out the key operations from our -entangled mess of a- _highly sophisticated_ codebase.\n\nOne more thing that occurs to me is that perhaps there should never be a line of code that deletes an output file. I haven't had a chance to dig into the code yet, and I'm sure there is a reason for it, but perhaps the same goal could be accomplished without deleting output files. What do you think?","from":"developer"},{"body":"That's a good point about the post-save-action check: code internal to Spark can only perform the check right before we return control back to the user so it's possible that the check will say that everything's fine even if the files are later deleted. This check still might have some value, but it's certainly prone to false-negatives and wouldn't be able to catch all manifestations of this bug.\n\nIn the original issue that prompted this patch, the deletion of the output file occurred inside of Hadoop code: if the renaming of the temporary file to the final output file fails, Hadoop's FileOutputCommitter will attempt to delete the final output file before re-attempting the rename (this is explained in a bit more detail, including a code walkthrough, in the description of this JIRA ticket).","from":"developer"},{"body":"there is blog post http://tech.grammarly.com/blog/posts/Petabyte-Scale-Text-Processing-with-Spark.html with a link to gist by Aaron Davidson: https://gist.github.com/aarondav/c513916e72101bbe14ec which mentions in comments that when using speculation it should be DirectOutputCommitter\n\nmight be somebody will find it useful (imho, it should be in documentation)\n","from":"developer"},{"body":"I'm clearing \"backport-needed\" since it's virtually certain that there will be no more 1.2.x or earlier releases, and so the fix that was committed won't go back further at this point.\n\nIs it something to leave open pending the ongoing conversation here? sounds like there may be more to the fix? ","from":"developer"},{"body":"What about splitting the issue into HDFS commits (interesting its mostly EC2 reports), which is probably fixed, and eventually consistent object stores (S3 on the Apache Hadoop releases (but not amazon's own), which need a different committer (no rename), better checks for file presence (direct stat() is more reliable than a directory listing), and some dedicated test suite which could be targeted straight at s3 —yet still runnable remotely by someone (not jenkins) from their own desktop & build servers. That's essentially what we do in core hadoop to qualify the object stores' base API compatibility.","from":"developer"}],"created":"2014-12-18T02:15:31.000+0000","description":"When speculative execution is enabled ({{spark.speculation=true}}), jobs that save output files may report that they have completed successfully even though some output partitions written by speculative tasks may be missing.\n\nh3. Reproduction\n\nThis symptom was reported to me by a Spark user and I've been doing my own investigation to try to come up with an in-house reproduction.\n\nI'm still working on a reliable local reproduction for this issue, which is a little tricky because Spark won't schedule speculated tasks on the same host as the original task, so you need an actual (or containerized) multi-host cluster to test speculation. Here's a simple reproduction of some of the symptoms on EC2, which can be run in {{spark-shell}} with {{--conf spark.speculation=true}}:\n\n{code}\n // Rig a job such that all but one of the tasks complete instantly\n // and one task runs for 20 seconds on its first attempt and instantly\n // on its second attempt:\n val numTasks = 100\n sc.parallelize(1 to numTasks, numTasks).repartition(2).mapPartitionsWithContext { case (ctx, iter) =>\n if (ctx.partitionId == 0) { // If this is the one task that should run really slow\n if (ctx.attemptId == 0) { // If this is the first attempt, run slow\n Thread.sleep(20 * 1000)\n }\n }\n iter\n }.map(x => (x, x)).saveAsTextFile(\"/test4\")\n{code}\n\nWhen I run this, I end up with a job that completes quickly (due to speculation) but reports failures from the speculated task:\n\n{code}\n[...]\n14/12/11 01:41:13 INFO scheduler.TaskSetManager: Finished task 37.1 in stage 3.0 (TID 411) in 131 ms on ip-172-31-8-164.us-west-2.compute.internal (100/100)\n14/12/11 01:41:13 INFO scheduler.DAGScheduler: Stage 3 (saveAsTextFile at :22) finished in 0.856 s\n14/12/11 01:41:13 INFO spark.SparkContext: Job finished: saveAsTextFile at :22, took 0.885438374 s\n14/12/11 01:41:13 INFO scheduler.TaskSetManager: Ignoring task-finished event for 70.1 in stage 3.0 because task 70 has already completed successfully\n\nscala> 14/12/11 01:41:13 WARN scheduler.TaskSetManager: Lost task 49.1 in stage 3.0 (TID 413, ip-172-31-8-164.us-west-2.compute.internal): java.io.IOException: Failed to save output of task: attempt_201412110141_0003_m_000049_413\n org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:160)\n org.apache.hadoop.mapred.FileOutputCommitter.moveTaskOutputs(FileOutputCommitter.java:172)\n org.apache.hadoop.mapred.FileOutputCommitter.commitTask(FileOutputCommitter.java:132)\n org.apache.spark.SparkHadoopWriter.commit(SparkHadoopWriter.scala:109)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:991)\n org.apache.spark.rdd.PairRDDFunctions$$anonfun$13.apply(PairRDDFunctions.scala:974)\n org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:62)\n org.apache.spark.scheduler.Task.run(Task.scala:54)\n org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:177)\n java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n java.lang.Thread.run(Thread.java:745)\n{code}\n\nOne interesting thing to note about this stack trace: if we look at {{FileOutputCommitter.java:160}} ([link|http://grepcode.com/file/repository.cloudera.com/content/repositories/releases/org.apache.hadoop/hadoop-core/2.5.0-mr1-cdh5.2.0/org/apache/hadoop/mapred/FileOutputCommitter.java#160]), this point in the execution seems to correspond to a case where a task completes, attempts to commit its output, fails for some reason, then deletes the destination file, tries again, and fails:\n\n{code}\n if (fs.isFile(taskOutput)) {\n152 Path finalOutputPath = getFinalPath(jobOutputDir, taskOutput, \n153 getTempTaskOutputPath(context));\n154 if (!fs.rename(taskOutput, finalOutputPath)) {\n155 if (!fs.delete(finalOutputPath, true)) {\n156 throw new IOException(\"Failed to delete earlier output of task: \" + \n157 attemptId);\n158 }\n159 if (!fs.rename(taskOutput, finalOutputPath)) {\n160 throw new IOException(\"Failed to save output of task: \" + \n161 \t\t attemptId);\n162 }\n163 }\n{code}\n\nThis could explain why the output file is missing: the second copy of the task keeps running after the job completes and deletes the output written by the other task after failing to commit its own copy of the output.\n\nThere are still a few open questions about how exactly we get into this scenario:\n\n*Why is the second copy of the task allowed to commit its output after the other task / the job has successfully completed?*\n\nTo check whether a task's temporary output should be committed, SparkHadoopWriter calls {{FileOutputCommitter.needsTaskCommit()}}, which returns {{true}} if the tasks's temporary output exists ([link|http://grepcode.com/file/repository.cloudera.com/content/repositories/releases/org.apache.hadoop/hadoop-core/2.5.0-mr1-cdh5.2.0/org/apache/hadoop/mapred/FileOutputCommitter.java#206]). Tihs does not seem to check whether the destination already exists. This means that {{needsTaskCommit}} can return {{true}} for speculative tasks.\n\n*Why does the rename fail?*\n\nI think that what's happening is that the temporary task output files are being deleted once the job has completed, which is causing the {{rename}} to fail because {{FileOutputCommitter.commitTask}} doesn't seem to guard against missing output files.\n\nI'm not sure about this, though, since the stack trace seems to imply that the temporary output file existed. Maybe the filesystem methods are returning stale metadata? Maybe there's a race? I think a race condition seems pretty unlikely, since the time-scale at which it would have to happen doesn't sync up with the scale of the timestamps that I saw in the user report.\n\nh3. Possible Fixes:\n\nThe root problem here might be that speculative copies of tasks are somehow allowed to commit their output. We might be able to fix this by centralizing the \"should this task commit its output\" decision at the driver.\n\n(I have more concrete suggestions of how to do this; to be posted soon)","issue_id":"12762466","key":"SPARK-4879","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-09-26T20:28:24.000+0000","role":"fixed_distractor","summary":"Missing output partitions after job completes with speculative execution"} {"case_id":"12763321","cluster":"DISTRACTOR-SPARK-4923","comments":[{"body":"Hey [~pc175@uowmail.edu.au] - we removed this from Maven because it's not meant as a stable API in Spark. Could you talk about which parts of the repl API you are using and how you are using it?","created":"2014-12-22T23:20:23.682+0000"},{"body":"Hey Patrick,\n\nThe following API has been integrated since 1.0.0, IMHO they are stable enough for daily prototyping, creating case class used to be defective but has been fixed long time ago.\nSparkILoop.getAddedJars()\n$SparkIMain.bind\n$SparkIMain.quietBind\n$SparkIMain.interpret\nend of :)\n\nAt first I assume that further development on it has been moved to databricks cloud. But the JIRA ticket was already there in September. So maybe demand on this API from the community is indeed low enough.\nHowever, I would still suggest keeping it, even promoting it into a Developer's API, this would encourage more projects to integrate in a more flexible way, and save prototyping/QA cost by customizing fixtures of REPL. People will still move to databricks cloud, which has far more features than that. Many influential projects already depends on the routinely published Scala-REPL (e.g. playFW), it would be strange for Spark not doing the same.\nWhat do you think? ","created":"2014-12-22T23:42:03.166+0000"},{"body":"Sorry my project is https://github.com/tribbloid/ISpark","created":"2014-12-22T23:45:07.519+0000"},{"body":"Hey [~pc175@uowmail.edu.au], thanks for filling that in - I didn't even realize we had code in there that was bytecode public. By stable I meant that we are promising it is an unchanging API. This is what we usually think about when we release things. For 1.2.0 I refactored our build and found out that we were publishing a bunch of random internal build components, so I took them all out of the published artifacts (examples, our assembly jar, etc) in SPARK-4923.\n\nAnyways - perhaps we could just annotate these as developer API's and be clear that they might change in the future. If you wanted to do that, and re-enable publishing them, I'd be happy to do that.","created":"2014-12-23T01:02:50.085+0000"},{"body":"thank you so much! First patch uploaded","created":"2014-12-23T01:46:48.671+0000"},{"body":"Hey [~pwendell], for some projects like Zeppelin(http://zeppelin-project.org/) depend on this repl jar to work, it would be very nice to publish this jar publicly, as for 1.1 version Spark actually published it. This changes will make several similar projects fail to update to latest Spark version.","created":"2014-12-25T07:24:27.728+0000"},{"body":"Yes - I can retro-actively publish this one for Spark 1.2.","created":"2014-12-25T14:53:21.866+0000"},{"body":"Thanks [~pwendell], when done, I'll start supporting this release it in the [Spark Notebook|https://github.com/andypetrella/spark-notebook] as well.","created":"2014-12-27T11:09:05.407+0000"},{"body":"Hey, a guy waiting for it was complaining that the version number of spark-repl_2.10 on typesafe repo is actually 1.2.0-SNAPSHOT, not the released 1.2.0\nIs it done on purpose?","created":"2014-12-30T15:53:20.030+0000"},{"body":"The typesafe repo would not be where the artifacts are officially released. I don't see any 1.2.x artifacts for the REPL anywhere, including Maven Central. That's the point of this JIRA. If someone published a snapshot release to a different repo I think you'd best ask them.","created":"2014-12-30T16:04:11.228+0000"},{"body":"I am sorry for any misunderstanding. I did not find the SNAPSHOT in the typesafe repo, but indeed somewhere else on GitHub.","created":"2014-12-30T16:30:16.455+0000"},{"body":"We are updating https://github.com/datastax/spark-cassandra-connector which integrates with the REPL, as does DSE, so this is a blocker for our upgrade to spark 1.2.0 as well.","created":"2014-12-30T23:24:34.445+0000"},{"body":"FYI, this is a blocker for us as well: https://github.com/ibm-et/spark-kernel\n\nSpecific issue: https://github.com/ibm-et/spark-kernel/issues/12\n\nWe are using quite a few different public methods from SparkIMain (such as valueOfTerm to pull out variables from the interpreter), not just interpret and bind. The API markings suggested by [~peng] would not be enough for us, [~pwendell].","created":"2015-01-03T07:33:38.724+0000"},{"body":"Just FYI, same for us here https://github.com/NFLabs/zeppelin/issues/260","created":"2015-01-03T14:00:06.661+0000"},{"body":"You are right, in fact 'Dev's API' simply means method is susceptible to changes without deprecation or notice, which the main 3 markings will be least likely to undergo.\nCould you please edit the patch and add more markings?","created":"2015-01-04T00:03:35.927+0000"},{"body":"Hey All,\n\nSorry this has caused a disruption. As I said in the earlier comment. if anyone on these projects can submit a patch that locks down the visibility in that package and opening up things that are specifically needed, I'm fine to keep publishing it (and do so retro-actively for 1.2). We just need to look closely at what we are exposing because this package currently violates Spark's API policy. Because the Scala repl does not itself offer any kind of API stability, it will be hard for Spark to do same. But I think it's fine to just annotate and expose unstable API's here, provided projects understand the implications of depending on them.\n\n[~senkwich] - since you guys are probably the heaviest user, would you be willing to take a crack at this? Basically start by making everything private and then go and unlock things that you need as Developer API's.\n\n- Patrick","created":"2015-01-12T21:58:22.241+0000"},{"body":"[~pwendell], I can definitely do that. Would you prefer a patch in the same form as the one attached? Or would it be better to create a pull request on Github for this with the changes?","created":"2015-01-12T22:07:20.054+0000"},{"body":"[~senkwich] definitely prefer github.","created":"2015-01-12T22:10:11.533+0000"},{"body":"Okay, I'll do that and update this JIRA once I've submitted the pull request.","created":"2015-01-12T22:12:41.157+0000"},{"body":"User 'rcsenkbeil' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4034","created":"2015-01-14T06:33:17.300+0000"},{"body":"As the nice bot has stated, I created a pull request for this issue. I detailed why I marked each method/field public and provided Scaladocs for each of them to make the exposure of the REPL API a little nicer.\n\nAs stated in the pull request, I only tackled Scala 2.10 for now as the Scala 2.11 did not \"appear\" to be ready, although I could easily be mistaken. I merely glanced at the SparkIMain and noticed that it did not have the class server declaration to ship the compiled class files nor was it - or any of the other classes - in the org.apache.spark.repl package.","created":"2015-01-14T06:36:35.598+0000"},{"body":"I updated the title of this to reflect the work that actually happened in Chip's patch. And SPARK-5289 is tracking publishing of the artifacts.","created":"2015-01-20T21:36:26.863+0000"}],"conversations":[{"body":"Spark-repl installation and deployment has been discontinued (see SPARK-3452). But its in the dependency list of a few projects that extends its initialization process.\nPlease remove the 'skip' setting in spark-repl and make it an 'official' API to encourage more platform to integrate with it.","from":"reporter","subject":"Add Developer API to REPL to allow re-publishing the REPL jar"},{"body":"Hey [~pc175@uowmail.edu.au] - we removed this from Maven because it's not meant as a stable API in Spark. Could you talk about which parts of the repl API you are using and how you are using it?","from":"developer"},{"body":"Hey Patrick,\n\nThe following API has been integrated since 1.0.0, IMHO they are stable enough for daily prototyping, creating case class used to be defective but has been fixed long time ago.\nSparkILoop.getAddedJars()\n$SparkIMain.bind\n$SparkIMain.quietBind\n$SparkIMain.interpret\nend of :)\n\nAt first I assume that further development on it has been moved to databricks cloud. But the JIRA ticket was already there in September. So maybe demand on this API from the community is indeed low enough.\nHowever, I would still suggest keeping it, even promoting it into a Developer's API, this would encourage more projects to integrate in a more flexible way, and save prototyping/QA cost by customizing fixtures of REPL. People will still move to databricks cloud, which has far more features than that. Many influential projects already depends on the routinely published Scala-REPL (e.g. playFW), it would be strange for Spark not doing the same.\nWhat do you think? ","from":"developer"},{"body":"Sorry my project is https://github.com/tribbloid/ISpark","from":"developer"},{"body":"Hey [~pc175@uowmail.edu.au], thanks for filling that in - I didn't even realize we had code in there that was bytecode public. By stable I meant that we are promising it is an unchanging API. This is what we usually think about when we release things. For 1.2.0 I refactored our build and found out that we were publishing a bunch of random internal build components, so I took them all out of the published artifacts (examples, our assembly jar, etc) in SPARK-4923.\n\nAnyways - perhaps we could just annotate these as developer API's and be clear that they might change in the future. If you wanted to do that, and re-enable publishing them, I'd be happy to do that.","from":"developer"},{"body":"thank you so much! First patch uploaded","from":"developer"},{"body":"Hey [~pwendell], for some projects like Zeppelin(http://zeppelin-project.org/) depend on this repl jar to work, it would be very nice to publish this jar publicly, as for 1.1 version Spark actually published it. This changes will make several similar projects fail to update to latest Spark version.","from":"developer"},{"body":"Yes - I can retro-actively publish this one for Spark 1.2.","from":"developer"},{"body":"Thanks [~pwendell], when done, I'll start supporting this release it in the [Spark Notebook|https://github.com/andypetrella/spark-notebook] as well.","from":"developer"},{"body":"Hey, a guy waiting for it was complaining that the version number of spark-repl_2.10 on typesafe repo is actually 1.2.0-SNAPSHOT, not the released 1.2.0\nIs it done on purpose?","from":"developer"},{"body":"The typesafe repo would not be where the artifacts are officially released. I don't see any 1.2.x artifacts for the REPL anywhere, including Maven Central. That's the point of this JIRA. If someone published a snapshot release to a different repo I think you'd best ask them.","from":"developer"},{"body":"I am sorry for any misunderstanding. I did not find the SNAPSHOT in the typesafe repo, but indeed somewhere else on GitHub.","from":"developer"},{"body":"We are updating https://github.com/datastax/spark-cassandra-connector which integrates with the REPL, as does DSE, so this is a blocker for our upgrade to spark 1.2.0 as well.","from":"developer"},{"body":"FYI, this is a blocker for us as well: https://github.com/ibm-et/spark-kernel\n\nSpecific issue: https://github.com/ibm-et/spark-kernel/issues/12\n\nWe are using quite a few different public methods from SparkIMain (such as valueOfTerm to pull out variables from the interpreter), not just interpret and bind. The API markings suggested by [~peng] would not be enough for us, [~pwendell].","from":"developer"},{"body":"Just FYI, same for us here https://github.com/NFLabs/zeppelin/issues/260","from":"developer"},{"body":"You are right, in fact 'Dev's API' simply means method is susceptible to changes without deprecation or notice, which the main 3 markings will be least likely to undergo.\nCould you please edit the patch and add more markings?","from":"developer"},{"body":"Hey All,\n\nSorry this has caused a disruption. As I said in the earlier comment. if anyone on these projects can submit a patch that locks down the visibility in that package and opening up things that are specifically needed, I'm fine to keep publishing it (and do so retro-actively for 1.2). We just need to look closely at what we are exposing because this package currently violates Spark's API policy. Because the Scala repl does not itself offer any kind of API stability, it will be hard for Spark to do same. But I think it's fine to just annotate and expose unstable API's here, provided projects understand the implications of depending on them.\n\n[~senkwich] - since you guys are probably the heaviest user, would you be willing to take a crack at this? Basically start by making everything private and then go and unlock things that you need as Developer API's.\n\n- Patrick","from":"developer"},{"body":"[~pwendell], I can definitely do that. Would you prefer a patch in the same form as the one attached? Or would it be better to create a pull request on Github for this with the changes?","from":"developer"},{"body":"[~senkwich] definitely prefer github.","from":"developer"},{"body":"Okay, I'll do that and update this JIRA once I've submitted the pull request.","from":"developer"},{"body":"User 'rcsenkbeil' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4034","from":"developer"},{"body":"As the nice bot has stated, I created a pull request for this issue. I detailed why I marked each method/field public and provided Scaladocs for each of them to make the exposure of the REPL API a little nicer.\n\nAs stated in the pull request, I only tackled Scala 2.10 for now as the Scala 2.11 did not \"appear\" to be ready, although I could easily be mistaken. I merely glanced at the SparkIMain and noticed that it did not have the class server declaration to ship the compiled class files nor was it - or any of the other classes - in the org.apache.spark.repl package.","from":"developer"},{"body":"I updated the title of this to reflect the work that actually happened in Chip's patch. And SPARK-5289 is tracking publishing of the artifacts.","from":"developer"}],"created":"2014-12-22T22:07:39.000+0000","description":"Spark-repl installation and deployment has been discontinued (see SPARK-3452). But its in the dependency list of a few projects that extends its initialization process.\nPlease remove the 'skip' setting in spark-repl and make it an 'official' API to encourage more platform to integrate with it.","issue_id":"12763321","key":"SPARK-4923","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-01-20T21:36:26.000+0000","role":"fixed_distractor","summary":"Add Developer API to REPL to allow re-publishing the REPL jar"} {"case_id":"12763602","cluster":"DISTRACTOR-SPARK-4943","comments":[{"body":"User 'alexliu68' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/3848","created":"2014-12-30T19:38:45.630+0000"},{"body":"One possible approach here is to change the signature of UnresolvedRelation as follows:\n\n{code}\ncase class UnresolvedRelation(tableIdentifier: Seq[String], alias: Option[String])\n{code}\n\nThis way we can leave parsing and handling of backticks up to the parser and let the catalogs interpret the identifiers in a system dependent way.","created":"2014-12-31T20:25:36.024+0000"},{"body":"The seq can be like \n\n{code}\ntableName, databaseName, clusterName, catalog\n{code}\n\nFirst element is tableName, the next one is databaseName, then clusterName and catalog(option).\n\n\nSQL query looks like\n\n{code}\nSELECT test_table.column1 FROM catalog.cluster.database.table AS test_table \nSELECT test_table.column1 FROM cluster.database.table AS test_table \nSELECT test_table.column1 FROM database.table AS test_table \nSELECT table.column1 FROM table\n{code}\n\nPotentially we should not use AS clause here.\n\nParser only allow maximum four levels in full table name. [catalog].[cluster].[database].[table]","created":"2014-12-31T21:26:54.025+0000"},{"body":"Should we also change the signatures of Catalog methods to use {code}tableIdentifier: Seq[String] {code} instead of {code}db: Option[String], tableName: String{code}?\n\n{code}\n\n def tableExists(db: Option[String], tableName: String): Boolean\n\n def lookupRelation(\n databaseName: Option[String],\n tableName: String,\n alias: Option[String] = None): LogicalPlan\n\n def registerTable(databaseName: Option[String], tableName: String, plan: LogicalPlan): Unit\n\n def unregisterTable(databaseName: Option[String], tableName: String): Unit\n\n def unregisterAllTables(): Unit\n\n protected def processDatabaseAndTableName(\n databaseName: Option[String],\n tableName: String): (Option[String], String)\n{code}","created":"2015-01-02T20:55:07.577+0000"},{"body":"The approach of {code}\ncase class UnresolvedRelation(tableIdentifier: Seq[String], alias: Option[String])\n{code} is a little unclear about what's stored in tableIdentifier by simply reading the code.\n\nAnother approach is storing catalog.cluster.database in databaseName and tableName in tableName and keep case class no change\n{code}\ncase class UnresolvedRelation(databaseName: Option[String], tableName: String, alias: Option[String])\n{code}\nso no API changes.\n\nIf we keep clusterName as a separate parameter, then API changes to \n{code}\ncase class UnresolvedRelation(clusterName: Option[String], databaseName: Option[String], tableName: String, alias: Option[String])\n{code}\n\nCatalog API needs change accordingly\n\n","created":"2015-01-05T20:59:09.209+0000"},{"body":"I wouldn't say that the notion of {{tableIdentifier: Seq\\[String\\]}} is unclear. Instead I would say that it is deliberately unspecified in order to be flexible. Some systems have {{clusters}}, some systems have {{databases}}, some systems have {{schema}}, some have {{tables}}. Thus, this API gives us one interface for communicating between the parser and the underlying datastore that makes no assumption about how that datastore is laid out.\n\nIf we do make this change then yes, I agree that we should also make it in the catalog as well. In general our handling of this has always been a little clunky since there is a whole bunch of code that just ignores the database field.\n\nOne question is: what parts of the table identifier Spark SQL handles and what parts we pass on to the datasource? A goal here should be to be able to connect to and join data from multiple sources. Here is what I would propose as an addition to the current API, which only lets you register individual tables.\n\n - Users can register external catalogs which are responsible for producing {{BaseRelation}}s.\n - Each external catalog has a user specified name that is given when registering.\n - There is a notion of the current catalog, which can be changed with {{USE}}. By default, we pass the all the {{tableIdentifiers}} to this default catalog and its up to it to determine what each part means.\n - Users can also specify fully qualified tables when joining multiple data sources. We detect this case when the first {{tableIdentifier}} matches one of the registered catalogs. In this case we strip of the catalog name and pass the remaining {{tableIdentifiers}} to the specified catalog.\n\nWhat do you think?","created":"2015-01-05T21:44:35.845+0000"},{"body":"Catalog part of table identifier should be handled by Spark SQL which calls the registered catalog Context to connect to the underline datasources. cluster/database/scheme/table should be handled by datasource(Cassandra Spark SQL integration can handle cluster, database and table level join).\n\ne.g.\n{code}\nSELECT test1.a, test1.b, test2.c FROM cassandra.cluster.database.table1 AS test1\n LEFT OUTER JOIN mySql.cluster.database.table2 AS test2 ON test1.a = test2.a\n{code}\n\nso cluster.database.table1 is passed to cassandra catalog datasource, cassandra is handled by Spark SQL to call cassandraContext which then call the underline datasource.\n\ncluster.database.table2 is passed to mySql catalog datasource, mySql is handled by Spark SQL to call the mySqlContext which then call the underline datasource.\n\n\nIf USE command is used, then all tableIdentifiers are passed to datasource. e.g.\n{code}\nUSE cassandra\nSELECT test1.a, test1.b, test2.c FROM cluster1.database.table AS test1\n LEFT OUTER JOIN cluster2.database.table AS test2 ON test1.a = test2.a\n{code}\n\ncluster1.database.table1 and cluster2.database.table are passed to cassandra datasource\n\n","created":"2015-01-05T22:13:47.595+0000"},{"body":"For each catalog, the configuration settings should start with catalog name. e.g.\n{noformat}\nset cassandra.cluster.database.table.ttl = 1000\nset cassandra.database.table.ttl =1000 (default cluster)\nset mysql.cluster.database.table.xxx = 200\n{noformat}\n\nIf there's no catalog in the setting string, use the default catalog.","created":"2015-01-05T22:19:29.584+0000"},{"body":"Thanks for your comments Alex. Are you proposing any changes to what I said?\n\nAnother thing I'm confused about is your comment regarding joins. As of now there is no public API for passing that kind of information down into a datasource.\n\nRegarding the configuration. We will pass the datasource a SQLContext and you can do {{.getConf}} using whatever arbitrary string you want. I don't think Spark SQL needs to have any control here.","created":"2015-01-05T23:30:27.880+0000"},{"body":"No changes to your approach. \n\nRegarding cluster1.database.table1 and cluster2.database.table passing to datasources. they are set as tableIdentifier and tableIdentifier is passed to catalog.lookupRelation method where datasource can use it.","created":"2015-01-05T23:42:39.560+0000"},{"body":"User 'alexliu68' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/3941","created":"2015-01-08T05:13:55.780+0000"},{"body":"Issue resolved by pull request 3941\n[https://github.com/apache/spark/pull/3941]","created":"2015-01-10T21:42:56.551+0000"},{"body":"User 'scwf' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4062","created":"2015-01-15T14:17:46.964+0000"}],"conversations":[{"body":"When integrating Spark 1.2.0 with Cassandra SQL, the following query is broken. It was working for Spark 1.1.0 version. Basically we use the table name having dot to include database name \n\n{code}\n[info] java.lang.RuntimeException: [1.29] failure: ``UNION'' expected but `.' found\n[info] \n\n[info] SELECT test1.a FROM sql_test.test1 AS test1 UNION DISTINCT SELECT test2.a FROM sql_test.test2 AS test2\n[info] ^\n[info] at scala.sys.package$.error(package.scala:27)\n[info] at org.apache.spark.sql.catalyst.AbstractSparkSQLParser.apply(SparkSQLParser.scala:33)\n[info] at org.apache.spark.sql.SQLContext$$anonfun$1.apply(SQLContext.scala:79)\n[info] at org.apache.spark.sql.SQLContext$$anonfun$1.apply(SQLContext.scala:79)\n[info] at org.apache.spark.sql.catalyst.SparkSQLParser$$anonfun$org$apache$spark$sql$catalyst$SparkSQLParser$$others$1.apply(SparkSQLParser.scala:174)\n[info] at org.apache.spark.sql.catalyst.SparkSQLParser$$anonfun$org$apache$spark$sql$catalyst$SparkSQLParser$$others$1.apply(SparkSQLParser.scala:173)\n[info] at scala.util.parsing.combinator.Parsers$Success.map(Parsers.scala:136)\n[info] at scala.util.parsing.combinator.Parsers$Success.map(Parsers.scala:135)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$map$1.apply(Parsers.scala:242)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$map$1.apply(Parsers.scala:242)\n[info] at scala.util.parsing.combinator.Parsers$$anon$3.apply(Parsers.scala:222)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$append$1$$anonfun$apply$2.apply(Parsers.scala:254)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$append$1$$anonfun$apply$2.apply(Parsers.scala:254)\n[info] at scala.util.parsing.combinator.Parsers$Failure.append(Parsers.scala:202)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$append$1.apply(Parsers.scala:254)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$append$1.apply(Parsers.scala:254)\n[info] at scala.util.parsing.combinator.Parsers$$anon$3.apply(Parsers.scala:222)\n[info] at scala.util.parsing.combinator.Parsers$$anon$2$$anonfun$apply$14.apply(Parsers.scala:891)\n[info] at scala.util.parsing.combinator.Parsers$$anon$2$$anonfun$apply$14.apply(Parsers.scala:891)\n[info] at scala.util.DynamicVariable.withValue(DynamicVariable.scala:57)\n[info] at scala.util.parsing.combinator.Parsers$$anon$2.apply(Parsers.scala:890)\n[info] at scala.util.parsing.combinator.PackratParsers$$anon$1.apply(PackratParsers.scala:110)\n[info] at org.apache.spark.sql.catalyst.AbstractSparkSQLParser.apply(SparkSQLParser.scala:31)\n[info] at org.apache.spark.sql.SQLContext$$anonfun$parseSql$1.apply(SQLContext.scala:83)\n[info] at org.apache.spark.sql.SQLContext$$anonfun$parseSql$1.apply(SQLContext.scala:83)\n[info] at scala.Option.getOrElse(Option.scala:120)\n[info] at org.apache.spark.sql.SQLContext.parseSql(SQLContext.scala:83)\n[info] at org.apache.spark.sql.cassandra.CassandraSQLContext.cassandraSql(CassandraSQLContext.scala:53)\n[info] at org.apache.spark.sql.cassandra.CassandraSQLContext.sql(CassandraSQLContext.scala:56)\n[info] at com.datastax.spark.connector.sql.CassandraSQLSpec$$anonfun$20.apply$mcV$sp(CassandraSQLSpec.scala:169)\n[info] at com.datastax.spark.connector.sql.CassandraSQLSpec$$anonfun$20.apply(CassandraSQLSpec.scala:168)\n[info] at com.datastax.spark.connector.sql.CassandraSQLSpec$$anonfun$20.apply(CassandraSQLSpec.scala:168)\n[info] at org.scalatest.Transformer$$anonfun$apply$1.apply$mcV$sp(Transformer.scala:22)\n[info] at org.scalatest.OutcomeOf$class.outcomeOf(OutcomeOf.scala:85)\n[info] at org.scalatest.OutcomeOf$.outcomeOf(OutcomeOf.scala:104)\n[info] at org.scalatest.Transformer.apply(Transformer.scala:22)\n[info] at org.scalatest.Transformer.apply(Transformer.scala:20)\n[info] at org.scalatest.FlatSpecLike$$anon$1.apply(FlatSpecLike.scala:1647)\n[info] at org.scalatest.Suite$class.withFixture(Suite.scala:1122)\n[info] at org.scalatest.FlatSpec.withFixture(FlatSpec.scala:1683)\n[info] at org.scalatest.FlatSpecLike$class.invokeWithFixture$1(FlatSpecLike.scala:1644)\n[info] at org.scalatest.FlatSpecLike$$anonfun$runTest$1.apply(FlatSpecLike.scala:1656)\n[info] at org.scalatest.FlatSpecLike$$anonfun$runTest$1.apply(FlatSpecLike.scala:1656)\n[info] at org.scalatest.SuperEngine.runTestImpl(Engine.scala:306)\n[info] at org.scalatest.FlatSpecLike$class.runTest(FlatSpecLike.scala:1656)\n[info] at org.scalatest.FlatSpec.runTest(FlatSpec.scala:1683)\n[info] at org.scalatest.FlatSpecLike$$anonfun$runTests$1.apply(FlatSpecLike.scala:1714)\n[info] at org.scalatest.FlatSpecLike$$anonfun$runTests$1.apply(FlatSpecLike.scala:1714)\n[info] at org.scalatest.SuperEngine$$anonfun$traverseSubNodes$1$1.apply(Engine.scala:413)\n[info] at org.scalatest.SuperEngine$$anonfun$traverseSubNodes$1$1.apply(Engine.scala:401)\n[info] at scala.collection.immutable.List.foreach(List.scala:318)\n[info] at org.scalatest.SuperEngine.traverseSubNodes$1(Engine.scala:401)\n[info] at org.scalatest.SuperEngine.org$scalatest$SuperEngine$$runTestsInBranch(Engine.scala:396)\n[info] at org.scalatest.SuperEngine.runTestsImpl(Engine.scala:483)\n[info] at org.scalatest.FlatSpecLike$class.runTests(FlatSpecLike.scala:1714)\n[info] at org.scalatest.FlatSpec.runTests(FlatSpec.scala:1683)\n[info] at org.scalatest.Suite$class.run(Suite.scala:1424)\n[info] at org.scalatest.FlatSpec.org$scalatest$FlatSpecLike$$super$run(FlatSpec.scala:1683)\n[info] at org.scalatest.FlatSpecLike$$anonfun$run$1.apply(FlatSpecLike.scala:1760)\n[info] at org.scalatest.FlatSpecLike$$anonfun$run$1.apply(FlatSpecLike.scala:1760)\n[info] at org.scalatest.SuperEngine.runImpl(Engine.scala:545)\n[info] at org.scalatest.FlatSpecLike$class.run(FlatSpecLike.scala:1760)\n[info] at org.scalatest.FlatSpec.run(FlatSpec.scala:1683)\n[info] at org.scalatest.tools.Framework.org$scalatest$tools$Framework$$runSuite(Framework.scala:466)\n[info] at org.scalatest.tools.Framework$ScalaTestTask.execute(Framework.scala:677)\n[info] at sbt.ForkMain$Run$2.call(ForkMain.java:294)\n[info] at sbt.ForkMain$Run$2.call(ForkMain.java:284)\n[info] at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n[info] at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n[info] at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n[info] at java.lang.Thread.run(Thread.java:745)\n[info] - should allow to select rows with union distinct clause *** FAILED *** (46 milliseconds)\n[info] java.lang.RuntimeException: [1.29] failure: ``UNION'' expected but `.' found\n{code}","from":"reporter","subject":"Parsing error for query with table name having dot"},{"body":"User 'alexliu68' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/3848","from":"developer"},{"body":"One possible approach here is to change the signature of UnresolvedRelation as follows:\n\n{code}\ncase class UnresolvedRelation(tableIdentifier: Seq[String], alias: Option[String])\n{code}\n\nThis way we can leave parsing and handling of backticks up to the parser and let the catalogs interpret the identifiers in a system dependent way.","from":"developer"},{"body":"The seq can be like \n\n{code}\ntableName, databaseName, clusterName, catalog\n{code}\n\nFirst element is tableName, the next one is databaseName, then clusterName and catalog(option).\n\n\nSQL query looks like\n\n{code}\nSELECT test_table.column1 FROM catalog.cluster.database.table AS test_table \nSELECT test_table.column1 FROM cluster.database.table AS test_table \nSELECT test_table.column1 FROM database.table AS test_table \nSELECT table.column1 FROM table\n{code}\n\nPotentially we should not use AS clause here.\n\nParser only allow maximum four levels in full table name. [catalog].[cluster].[database].[table]","from":"developer"},{"body":"Should we also change the signatures of Catalog methods to use {code}tableIdentifier: Seq[String] {code} instead of {code}db: Option[String], tableName: String{code}?\n\n{code}\n\n def tableExists(db: Option[String], tableName: String): Boolean\n\n def lookupRelation(\n databaseName: Option[String],\n tableName: String,\n alias: Option[String] = None): LogicalPlan\n\n def registerTable(databaseName: Option[String], tableName: String, plan: LogicalPlan): Unit\n\n def unregisterTable(databaseName: Option[String], tableName: String): Unit\n\n def unregisterAllTables(): Unit\n\n protected def processDatabaseAndTableName(\n databaseName: Option[String],\n tableName: String): (Option[String], String)\n{code}","from":"developer"},{"body":"The approach of {code}\ncase class UnresolvedRelation(tableIdentifier: Seq[String], alias: Option[String])\n{code} is a little unclear about what's stored in tableIdentifier by simply reading the code.\n\nAnother approach is storing catalog.cluster.database in databaseName and tableName in tableName and keep case class no change\n{code}\ncase class UnresolvedRelation(databaseName: Option[String], tableName: String, alias: Option[String])\n{code}\nso no API changes.\n\nIf we keep clusterName as a separate parameter, then API changes to \n{code}\ncase class UnresolvedRelation(clusterName: Option[String], databaseName: Option[String], tableName: String, alias: Option[String])\n{code}\n\nCatalog API needs change accordingly\n\n","from":"developer"},{"body":"I wouldn't say that the notion of {{tableIdentifier: Seq\\[String\\]}} is unclear. Instead I would say that it is deliberately unspecified in order to be flexible. Some systems have {{clusters}}, some systems have {{databases}}, some systems have {{schema}}, some have {{tables}}. Thus, this API gives us one interface for communicating between the parser and the underlying datastore that makes no assumption about how that datastore is laid out.\n\nIf we do make this change then yes, I agree that we should also make it in the catalog as well. In general our handling of this has always been a little clunky since there is a whole bunch of code that just ignores the database field.\n\nOne question is: what parts of the table identifier Spark SQL handles and what parts we pass on to the datasource? A goal here should be to be able to connect to and join data from multiple sources. Here is what I would propose as an addition to the current API, which only lets you register individual tables.\n\n - Users can register external catalogs which are responsible for producing {{BaseRelation}}s.\n - Each external catalog has a user specified name that is given when registering.\n - There is a notion of the current catalog, which can be changed with {{USE}}. By default, we pass the all the {{tableIdentifiers}} to this default catalog and its up to it to determine what each part means.\n - Users can also specify fully qualified tables when joining multiple data sources. We detect this case when the first {{tableIdentifier}} matches one of the registered catalogs. In this case we strip of the catalog name and pass the remaining {{tableIdentifiers}} to the specified catalog.\n\nWhat do you think?","from":"developer"},{"body":"Catalog part of table identifier should be handled by Spark SQL which calls the registered catalog Context to connect to the underline datasources. cluster/database/scheme/table should be handled by datasource(Cassandra Spark SQL integration can handle cluster, database and table level join).\n\ne.g.\n{code}\nSELECT test1.a, test1.b, test2.c FROM cassandra.cluster.database.table1 AS test1\n LEFT OUTER JOIN mySql.cluster.database.table2 AS test2 ON test1.a = test2.a\n{code}\n\nso cluster.database.table1 is passed to cassandra catalog datasource, cassandra is handled by Spark SQL to call cassandraContext which then call the underline datasource.\n\ncluster.database.table2 is passed to mySql catalog datasource, mySql is handled by Spark SQL to call the mySqlContext which then call the underline datasource.\n\n\nIf USE command is used, then all tableIdentifiers are passed to datasource. e.g.\n{code}\nUSE cassandra\nSELECT test1.a, test1.b, test2.c FROM cluster1.database.table AS test1\n LEFT OUTER JOIN cluster2.database.table AS test2 ON test1.a = test2.a\n{code}\n\ncluster1.database.table1 and cluster2.database.table are passed to cassandra datasource\n\n","from":"developer"},{"body":"For each catalog, the configuration settings should start with catalog name. e.g.\n{noformat}\nset cassandra.cluster.database.table.ttl = 1000\nset cassandra.database.table.ttl =1000 (default cluster)\nset mysql.cluster.database.table.xxx = 200\n{noformat}\n\nIf there's no catalog in the setting string, use the default catalog.","from":"developer"},{"body":"Thanks for your comments Alex. Are you proposing any changes to what I said?\n\nAnother thing I'm confused about is your comment regarding joins. As of now there is no public API for passing that kind of information down into a datasource.\n\nRegarding the configuration. We will pass the datasource a SQLContext and you can do {{.getConf}} using whatever arbitrary string you want. I don't think Spark SQL needs to have any control here.","from":"developer"},{"body":"No changes to your approach. \n\nRegarding cluster1.database.table1 and cluster2.database.table passing to datasources. they are set as tableIdentifier and tableIdentifier is passed to catalog.lookupRelation method where datasource can use it.","from":"developer"},{"body":"User 'alexliu68' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/3941","from":"developer"},{"body":"Issue resolved by pull request 3941\n[https://github.com/apache/spark/pull/3941]","from":"developer"},{"body":"User 'scwf' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4062","from":"developer"}],"created":"2014-12-24T01:19:51.000+0000","description":"When integrating Spark 1.2.0 with Cassandra SQL, the following query is broken. It was working for Spark 1.1.0 version. Basically we use the table name having dot to include database name \n\n{code}\n[info] java.lang.RuntimeException: [1.29] failure: ``UNION'' expected but `.' found\n[info] \n\n[info] SELECT test1.a FROM sql_test.test1 AS test1 UNION DISTINCT SELECT test2.a FROM sql_test.test2 AS test2\n[info] ^\n[info] at scala.sys.package$.error(package.scala:27)\n[info] at org.apache.spark.sql.catalyst.AbstractSparkSQLParser.apply(SparkSQLParser.scala:33)\n[info] at org.apache.spark.sql.SQLContext$$anonfun$1.apply(SQLContext.scala:79)\n[info] at org.apache.spark.sql.SQLContext$$anonfun$1.apply(SQLContext.scala:79)\n[info] at org.apache.spark.sql.catalyst.SparkSQLParser$$anonfun$org$apache$spark$sql$catalyst$SparkSQLParser$$others$1.apply(SparkSQLParser.scala:174)\n[info] at org.apache.spark.sql.catalyst.SparkSQLParser$$anonfun$org$apache$spark$sql$catalyst$SparkSQLParser$$others$1.apply(SparkSQLParser.scala:173)\n[info] at scala.util.parsing.combinator.Parsers$Success.map(Parsers.scala:136)\n[info] at scala.util.parsing.combinator.Parsers$Success.map(Parsers.scala:135)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$map$1.apply(Parsers.scala:242)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$map$1.apply(Parsers.scala:242)\n[info] at scala.util.parsing.combinator.Parsers$$anon$3.apply(Parsers.scala:222)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$append$1$$anonfun$apply$2.apply(Parsers.scala:254)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$append$1$$anonfun$apply$2.apply(Parsers.scala:254)\n[info] at scala.util.parsing.combinator.Parsers$Failure.append(Parsers.scala:202)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$append$1.apply(Parsers.scala:254)\n[info] at scala.util.parsing.combinator.Parsers$Parser$$anonfun$append$1.apply(Parsers.scala:254)\n[info] at scala.util.parsing.combinator.Parsers$$anon$3.apply(Parsers.scala:222)\n[info] at scala.util.parsing.combinator.Parsers$$anon$2$$anonfun$apply$14.apply(Parsers.scala:891)\n[info] at scala.util.parsing.combinator.Parsers$$anon$2$$anonfun$apply$14.apply(Parsers.scala:891)\n[info] at scala.util.DynamicVariable.withValue(DynamicVariable.scala:57)\n[info] at scala.util.parsing.combinator.Parsers$$anon$2.apply(Parsers.scala:890)\n[info] at scala.util.parsing.combinator.PackratParsers$$anon$1.apply(PackratParsers.scala:110)\n[info] at org.apache.spark.sql.catalyst.AbstractSparkSQLParser.apply(SparkSQLParser.scala:31)\n[info] at org.apache.spark.sql.SQLContext$$anonfun$parseSql$1.apply(SQLContext.scala:83)\n[info] at org.apache.spark.sql.SQLContext$$anonfun$parseSql$1.apply(SQLContext.scala:83)\n[info] at scala.Option.getOrElse(Option.scala:120)\n[info] at org.apache.spark.sql.SQLContext.parseSql(SQLContext.scala:83)\n[info] at org.apache.spark.sql.cassandra.CassandraSQLContext.cassandraSql(CassandraSQLContext.scala:53)\n[info] at org.apache.spark.sql.cassandra.CassandraSQLContext.sql(CassandraSQLContext.scala:56)\n[info] at com.datastax.spark.connector.sql.CassandraSQLSpec$$anonfun$20.apply$mcV$sp(CassandraSQLSpec.scala:169)\n[info] at com.datastax.spark.connector.sql.CassandraSQLSpec$$anonfun$20.apply(CassandraSQLSpec.scala:168)\n[info] at com.datastax.spark.connector.sql.CassandraSQLSpec$$anonfun$20.apply(CassandraSQLSpec.scala:168)\n[info] at org.scalatest.Transformer$$anonfun$apply$1.apply$mcV$sp(Transformer.scala:22)\n[info] at org.scalatest.OutcomeOf$class.outcomeOf(OutcomeOf.scala:85)\n[info] at org.scalatest.OutcomeOf$.outcomeOf(OutcomeOf.scala:104)\n[info] at org.scalatest.Transformer.apply(Transformer.scala:22)\n[info] at org.scalatest.Transformer.apply(Transformer.scala:20)\n[info] at org.scalatest.FlatSpecLike$$anon$1.apply(FlatSpecLike.scala:1647)\n[info] at org.scalatest.Suite$class.withFixture(Suite.scala:1122)\n[info] at org.scalatest.FlatSpec.withFixture(FlatSpec.scala:1683)\n[info] at org.scalatest.FlatSpecLike$class.invokeWithFixture$1(FlatSpecLike.scala:1644)\n[info] at org.scalatest.FlatSpecLike$$anonfun$runTest$1.apply(FlatSpecLike.scala:1656)\n[info] at org.scalatest.FlatSpecLike$$anonfun$runTest$1.apply(FlatSpecLike.scala:1656)\n[info] at org.scalatest.SuperEngine.runTestImpl(Engine.scala:306)\n[info] at org.scalatest.FlatSpecLike$class.runTest(FlatSpecLike.scala:1656)\n[info] at org.scalatest.FlatSpec.runTest(FlatSpec.scala:1683)\n[info] at org.scalatest.FlatSpecLike$$anonfun$runTests$1.apply(FlatSpecLike.scala:1714)\n[info] at org.scalatest.FlatSpecLike$$anonfun$runTests$1.apply(FlatSpecLike.scala:1714)\n[info] at org.scalatest.SuperEngine$$anonfun$traverseSubNodes$1$1.apply(Engine.scala:413)\n[info] at org.scalatest.SuperEngine$$anonfun$traverseSubNodes$1$1.apply(Engine.scala:401)\n[info] at scala.collection.immutable.List.foreach(List.scala:318)\n[info] at org.scalatest.SuperEngine.traverseSubNodes$1(Engine.scala:401)\n[info] at org.scalatest.SuperEngine.org$scalatest$SuperEngine$$runTestsInBranch(Engine.scala:396)\n[info] at org.scalatest.SuperEngine.runTestsImpl(Engine.scala:483)\n[info] at org.scalatest.FlatSpecLike$class.runTests(FlatSpecLike.scala:1714)\n[info] at org.scalatest.FlatSpec.runTests(FlatSpec.scala:1683)\n[info] at org.scalatest.Suite$class.run(Suite.scala:1424)\n[info] at org.scalatest.FlatSpec.org$scalatest$FlatSpecLike$$super$run(FlatSpec.scala:1683)\n[info] at org.scalatest.FlatSpecLike$$anonfun$run$1.apply(FlatSpecLike.scala:1760)\n[info] at org.scalatest.FlatSpecLike$$anonfun$run$1.apply(FlatSpecLike.scala:1760)\n[info] at org.scalatest.SuperEngine.runImpl(Engine.scala:545)\n[info] at org.scalatest.FlatSpecLike$class.run(FlatSpecLike.scala:1760)\n[info] at org.scalatest.FlatSpec.run(FlatSpec.scala:1683)\n[info] at org.scalatest.tools.Framework.org$scalatest$tools$Framework$$runSuite(Framework.scala:466)\n[info] at org.scalatest.tools.Framework$ScalaTestTask.execute(Framework.scala:677)\n[info] at sbt.ForkMain$Run$2.call(ForkMain.java:294)\n[info] at sbt.ForkMain$Run$2.call(ForkMain.java:284)\n[info] at java.util.concurrent.FutureTask.run(FutureTask.java:262)\n[info] at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n[info] at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n[info] at java.lang.Thread.run(Thread.java:745)\n[info] - should allow to select rows with union distinct clause *** FAILED *** (46 milliseconds)\n[info] java.lang.RuntimeException: [1.29] failure: ``UNION'' expected but `.' found\n{code}","issue_id":"12763602","key":"SPARK-4943","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-01-10T21:42:56.000+0000","role":"fixed_distractor","summary":"Parsing error for query with table name having dot"} {"case_id":"12763987","cluster":"DISTRACTOR-SPARK-4988","comments":[{"body":"in `ScalaReflection.scala`, method `convertToScala` `case (d: Decimal, _: DecimalType) => d.toBigDecimal`.\n\nso, `HiveShim.createDecimal(o.asInstanceOf[Decimal].toBigDecimal.underlying())` in `HiveInspectors.scala` report error\n\n","created":"2014-12-29T07:54:33.658+0000"},{"body":"User 'guowei2' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/3821","created":"2014-12-29T08:03:08.159+0000"},{"body":"This is also happening also for comparison\nselect * from tbl where decfield > 2\n\nOne more issue with decimal types is that when a udf is used against a decimal type the argument to udf comes as scala BigDecimal instead of JavaBigDecimal when using java api.","created":"2015-01-23T11:37:09.735+0000"},{"body":"SQL statements and data","created":"2015-06-26T23:26:19.765+0000"},{"body":"Both the spark-sql statements are working. \nHere is the listing if some one want to test it. The same is attached in a file too.\n\n--------\n1. Data file :\n\nbash-3.2$ cat t1.txt\n1,Barney,10.5\n2,Nancy,7.5\n3,Tony,4.5\n5,Fred,3.5\n6,Alok,12.5\n7,Jan,23.5\n8,Barbara,11.5\n9,Mike,6.4\n10,Deron,3.7\n11,Glenn,9.9\n12,Seth,7.8\n13,Gerome,4.5\n14,Alan,34.5\n15,Rohan,33.7\n16,clifford,3.5\n17,Rosstin,1.5\n\n2. Create table with decimal data type.\n\nCREATE EXTERNAL TABLE user(id INT, name STRING, fico Decimal(4,2)) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' LINES TERMINATED BY '\\n' STORED AS TEXTFILE LOCATION '/Users/sudhakarthota/tmp';\n\n3. Check if all the data is appearing for select statement.\n\npark-sql> select * from user;\n1\tBarney\t10.5\n2\tNancy\t7.5\n3\tTony\t4.5\n5\tFred\t3.5\n6\tAlok\t12.5\n7\tJan\t23.5\n8\tBarbara\t11.5\n9\tMike\t6.4\n10\tDeron\t3.7\n11\tGlenn\t9.9\n12\tSeth\t7.8\n13\tGerome\t4.5\n14\tAlan\t34.5\n15\tRohan\t33.7\n16\tclifford\t3.5\n17\tRosstin\t1.5\nTime taken: 0.062 seconds, Fetched 16 row(s)\n\n4. Create table test1 that was failing:\ncreate table test1 as select * from user order by fico limit 10;\n\nspark-sql> create table test1 as select * from user order by fico limit 10;\nrmr: DEPRECATED: Please use 'rm -r' instead.\nDeleted file:///user/hive/warehouse/test1\nTime taken: 0.223 seconds\nspark-sql> \n\n5. select records.\n\nspark-sql> select * from test1; \n17\tRosstin\t1.5\n16\tclifford\t3.5\n5\tFred\t3.5\n10\tDeron\t3.7\n3\tTony\t4.5\n13\tGerome\t4.5\n9\tMike\t6.4\n2\tNancy\t7.5\n12\tSeth\t7.8\n11\tGlenn\t9.9\nTime taken: 0.055 seconds, Fetched 10 row(s)\nspark-sql> \n\n\n6. Do the second test that was failing\n\nspark-sql> select * from user where fico >2;\n1\tBarney\t10.5\n2\tNancy\t7.5\n3\tTony\t4.5\n5\tFred\t3.5\n6\tAlok\t12.5\n7\tJan\t23.5\n8\tBarbara\t11.5\n9\tMike\t6.4\n10\tDeron\t3.7\n11\tGlenn\t9.9\n12\tSeth\t7.8\n13\tGerome\t4.5\n14\tAlan\t34.5\n15\tRohan\t33.7\n16\tclifford\t3.5\nTime taken: 0.061 seconds, Fetched 15 row(s)\nspark-sql> \n \n------\n\nThanks\nSudhakar Thota","created":"2015-06-26T23:26:27.035+0000"},{"body":"This should be fixed after cleaning up Row and InternalRow stuff.","created":"2015-08-06T23:20:15.014+0000"}],"conversations":[{"body":"A table 'test' with a decimal type col.\ncreate table test1 as select * from test order by a limit 10;\n\norg.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 2.0 failed 1 times, most recent failure: Lost task 0.0 in stage 2.0 (TID 2, localhost): java.lang.ClassCastException: scala.math.BigDecimal cannot be cast to org.apache.spark.sql.catalyst.types.decimal.Decimal\n\tat org.apache.spark.sql.hive.HiveInspectors$$anonfun$wrapperFor$2.apply(HiveInspectors.scala:339)\n\tat org.apache.spark.sql.hive.HiveInspectors$$anonfun$wrapperFor$2.apply(HiveInspectors.scala:339)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable$$anonfun$org$apache$spark$sql$hive$execution$InsertIntoHiveTable$$writeToFile$1$1.apply(InsertIntoHiveTable.scala:111)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable$$anonfun$org$apache$spark$sql$hive$execution$InsertIntoHiveTable$$writeToFile$1$1.apply(InsertIntoHiveTable.scala:108)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:727)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable.org$apache$spark$sql$hive$execution$InsertIntoHiveTable$$writeToFile$1(InsertIntoHiveTable.scala:108)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable$$anonfun$saveAsHiveFile$3.apply(InsertIntoHiveTable.scala:87)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable$$anonfun$saveAsHiveFile$3.apply(InsertIntoHiveTable.scala:87)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:61)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:56)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:195)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:744)","from":"reporter","subject":"\"Create table ..as select ..from..order by .. limit 10\" report error when one col is a Decimal"},{"body":"in `ScalaReflection.scala`, method `convertToScala` `case (d: Decimal, _: DecimalType) => d.toBigDecimal`.\n\nso, `HiveShim.createDecimal(o.asInstanceOf[Decimal].toBigDecimal.underlying())` in `HiveInspectors.scala` report error\n\n","from":"developer"},{"body":"User 'guowei2' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/3821","from":"developer"},{"body":"This is also happening also for comparison\nselect * from tbl where decfield > 2\n\nOne more issue with decimal types is that when a udf is used against a decimal type the argument to udf comes as scala BigDecimal instead of JavaBigDecimal when using java api.","from":"developer"},{"body":"SQL statements and data","from":"developer"},{"body":"Both the spark-sql statements are working. \nHere is the listing if some one want to test it. The same is attached in a file too.\n\n--------\n1. Data file :\n\nbash-3.2$ cat t1.txt\n1,Barney,10.5\n2,Nancy,7.5\n3,Tony,4.5\n5,Fred,3.5\n6,Alok,12.5\n7,Jan,23.5\n8,Barbara,11.5\n9,Mike,6.4\n10,Deron,3.7\n11,Glenn,9.9\n12,Seth,7.8\n13,Gerome,4.5\n14,Alan,34.5\n15,Rohan,33.7\n16,clifford,3.5\n17,Rosstin,1.5\n\n2. Create table with decimal data type.\n\nCREATE EXTERNAL TABLE user(id INT, name STRING, fico Decimal(4,2)) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' LINES TERMINATED BY '\\n' STORED AS TEXTFILE LOCATION '/Users/sudhakarthota/tmp';\n\n3. Check if all the data is appearing for select statement.\n\npark-sql> select * from user;\n1\tBarney\t10.5\n2\tNancy\t7.5\n3\tTony\t4.5\n5\tFred\t3.5\n6\tAlok\t12.5\n7\tJan\t23.5\n8\tBarbara\t11.5\n9\tMike\t6.4\n10\tDeron\t3.7\n11\tGlenn\t9.9\n12\tSeth\t7.8\n13\tGerome\t4.5\n14\tAlan\t34.5\n15\tRohan\t33.7\n16\tclifford\t3.5\n17\tRosstin\t1.5\nTime taken: 0.062 seconds, Fetched 16 row(s)\n\n4. Create table test1 that was failing:\ncreate table test1 as select * from user order by fico limit 10;\n\nspark-sql> create table test1 as select * from user order by fico limit 10;\nrmr: DEPRECATED: Please use 'rm -r' instead.\nDeleted file:///user/hive/warehouse/test1\nTime taken: 0.223 seconds\nspark-sql> \n\n5. select records.\n\nspark-sql> select * from test1; \n17\tRosstin\t1.5\n16\tclifford\t3.5\n5\tFred\t3.5\n10\tDeron\t3.7\n3\tTony\t4.5\n13\tGerome\t4.5\n9\tMike\t6.4\n2\tNancy\t7.5\n12\tSeth\t7.8\n11\tGlenn\t9.9\nTime taken: 0.055 seconds, Fetched 10 row(s)\nspark-sql> \n\n\n6. Do the second test that was failing\n\nspark-sql> select * from user where fico >2;\n1\tBarney\t10.5\n2\tNancy\t7.5\n3\tTony\t4.5\n5\tFred\t3.5\n6\tAlok\t12.5\n7\tJan\t23.5\n8\tBarbara\t11.5\n9\tMike\t6.4\n10\tDeron\t3.7\n11\tGlenn\t9.9\n12\tSeth\t7.8\n13\tGerome\t4.5\n14\tAlan\t34.5\n15\tRohan\t33.7\n16\tclifford\t3.5\nTime taken: 0.061 seconds, Fetched 15 row(s)\nspark-sql> \n \n------\n\nThanks\nSudhakar Thota","from":"developer"},{"body":"This should be fixed after cleaning up Row and InternalRow stuff.","from":"developer"}],"created":"2014-12-29T07:48:34.000+0000","description":"A table 'test' with a decimal type col.\ncreate table test1 as select * from test order by a limit 10;\n\norg.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 2.0 failed 1 times, most recent failure: Lost task 0.0 in stage 2.0 (TID 2, localhost): java.lang.ClassCastException: scala.math.BigDecimal cannot be cast to org.apache.spark.sql.catalyst.types.decimal.Decimal\n\tat org.apache.spark.sql.hive.HiveInspectors$$anonfun$wrapperFor$2.apply(HiveInspectors.scala:339)\n\tat org.apache.spark.sql.hive.HiveInspectors$$anonfun$wrapperFor$2.apply(HiveInspectors.scala:339)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable$$anonfun$org$apache$spark$sql$hive$execution$InsertIntoHiveTable$$writeToFile$1$1.apply(InsertIntoHiveTable.scala:111)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable$$anonfun$org$apache$spark$sql$hive$execution$InsertIntoHiveTable$$writeToFile$1$1.apply(InsertIntoHiveTable.scala:108)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:727)\n\tat org.apache.spark.InterruptibleIterator.foreach(InterruptibleIterator.scala:28)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable.org$apache$spark$sql$hive$execution$InsertIntoHiveTable$$writeToFile$1(InsertIntoHiveTable.scala:108)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable$$anonfun$saveAsHiveFile$3.apply(InsertIntoHiveTable.scala:87)\n\tat org.apache.spark.sql.hive.execution.InsertIntoHiveTable$$anonfun$saveAsHiveFile$3.apply(InsertIntoHiveTable.scala:87)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:61)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:56)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:195)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:744)","issue_id":"12763987","key":"SPARK-4988","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-08-06T23:20:14.000+0000","role":"fixed_distractor","summary":"\"Create table ..as select ..from..order by .. limit 10\" report error when one col is a Decimal"} {"case_id":"12767147","cluster":"DISTRACTOR-SPARK-5220","comments":[{"body":"Hi Max, as I said in the mail, this is an expected behavior of receiver and block generator because of locking mechanism of BlockGenerator. The receiver will block on the locks for adding data into BlockGenerator, and the BlockGenerator is waiting for pushing thread to put data into HDFS and BM. Because of unmatched speed, it is expected from my understanding.","created":"2015-01-14T01:39:11.403+0000"},{"body":"Hi Jerry, my point is that keepPushingBlocks in BlockGenerator should not terminate because of any exception thrown by the receiver. With current implementation, if an exception is thrown, even when it is not stopped, keepPushingBlocks will terminate, so does the block pushing thread. It is not a matter of speed, it finished. No subsequent blocks can be pushed to BM because there is not block pushing thread any more after the exception. This behavior contradict its name.","created":"2015-01-14T13:35:09.879+0000"},{"body":"Aha, got it, seems this is make sense for normal receivers, but for reliable receivers, I'm not sure is it still make sense to push the following blocks if the previous block is failed.","created":"2015-01-14T13:42:12.463+0000"},{"body":"Current situation: \n1. The TimeoutException is thrown by ReliableKafkaReceiver, receiver logs the error and do nothing else;\n2. Block pushing thread in BlockGenerator is terminated due to the exception;\n3. ReliableKafkaReceiver is ACTIVE but rogues afterward because blocksForPushing queue is full and no block pushing thread available to clean the queue.\n\nSo when the exception occurs, what user will see is the streaming application keeps running but sits there doing nothing (We are using some supervisor to restart the failed app. But if the app keeps running and doing nothing, we have to have some curator to monitor the data flow and kill the app in such situation). Can ReliableKafkaReceiver handle TimeoutException? For example retry pushing block when a TimeoutException is thrown? Or restart itself? Or at least fail the streaming application?","created":"2015-01-14T13:54:07.657+0000"},{"body":"Yeah, I totally understand what is your requirement, maybe we should figure out a more elegant way to deal with such situation. Currently there's a solution trying to partially fix the problem you mentioned, you can refer to (https://github.com/apache/spark/pull/3655). But I think it is not a thorough way because the push thread will finally exited and block the whole system.","created":"2015-01-14T14:10:39.009+0000"},{"body":"I believe https://github.com/apache/spark/pull/3655 fixes the problem. The block pushing thread will not exit when an exception happens because the exception is handled inside storeBlockAndCommitOffset. The change stops the receiver after three retries, which will also stop the BlockGenerator inside ReliableKafkaReceiver. What would happen when the receiver is stopped? Does the application fail?","created":"2015-01-14T14:42:29.123+0000"},{"body":"[~superxma] is this resolved then?","created":"2015-05-15T13:08:25.835+0000"}],"conversations":[{"body":"I am running a Spark streaming application with ReliableKafkaReceiver. It uses BlockGenerator to push blocks to BlockManager. However, writing WALs to HDFS may time out that causes keepPushingBlocks in BlockGenerator to terminate.\n\n15/01/12 19:07:06 ERROR receiver.BlockGenerator: Error in block pushing thread\njava.util.concurrent.TimeoutException: Futures timed out after [30 seconds]\n at scala.concurrent.impl.Promise$DefaultPromise.ready(Promise.scala:219)\n at scala.concurrent.impl.Promise$DefaultPromise.result(Promise.scala:223)\n at scala.concurrent.Await$$anonfun$result$1.apply(package.scala:107)\n at scala.concurrent.BlockContext$DefaultBlockContext$.blockOn(BlockContext.scala:53)\n at scala.concurrent.Await$.result(package.scala:107)\n at org.apache.spark.streaming.receiver.WriteAheadLogBasedBlockHandler.storeBlock(ReceivedBlockHandler.scala:176)\n at org.apache.spark.streaming.receiver.ReceiverSupervisorImpl.pushAndReportBlock(ReceiverSupervisorImpl.scala:160)\n at org.apache.spark.streaming.receiver.ReceiverSupervisorImpl.pushArrayBuffer(ReceiverSupervisorImpl.scala:126)\n at org.apache.spark.streaming.receiver.Receiver.store(Receiver.scala:124)\n at org.apache.spark.streaming.kafka.ReliableKafkaReceiver.org$apache$spark$streaming$kafka$ReliableKafkaReceiver$$storeBlockAndCommitOffset(ReliableKafkaReceiver.scala:207)\n at org.apache.spark.streaming.kafka.ReliableKafkaReceiver$GeneratedBlockHandler.onPushBlock(ReliableKafkaReceiver.scala:275)\n at org.apache.spark.streaming.receiver.BlockGenerator.pushBlock(BlockGenerator.scala:181)\n at org.apache.spark.streaming.receiver.BlockGenerator.org$apache$spark$streaming$receiver$BlockGenerator$$keepPushingBlocks(BlockGenerator.scala:154)\n at org.apache.spark.streaming.receiver.BlockGenerator$$anon$1.run(BlockGenerator.scala:86)\n\nThen the block pushing thread is done and no subsequent blocks can be pushed into blockManager. In turn this blocks receiver from receiving new data.\n\nSo when running my app and the TimeoutException happens, the ReliableKafkaReceiver stays in ACTIVE status but doesn't do anything at all. The application rogues.","from":"reporter","subject":"keepPushingBlocks in BlockGenerator terminated when an exception occurs, which causes the block pushing thread to terminate and blocks receiver "},{"body":"Hi Max, as I said in the mail, this is an expected behavior of receiver and block generator because of locking mechanism of BlockGenerator. The receiver will block on the locks for adding data into BlockGenerator, and the BlockGenerator is waiting for pushing thread to put data into HDFS and BM. Because of unmatched speed, it is expected from my understanding.","from":"developer"},{"body":"Hi Jerry, my point is that keepPushingBlocks in BlockGenerator should not terminate because of any exception thrown by the receiver. With current implementation, if an exception is thrown, even when it is not stopped, keepPushingBlocks will terminate, so does the block pushing thread. It is not a matter of speed, it finished. No subsequent blocks can be pushed to BM because there is not block pushing thread any more after the exception. This behavior contradict its name.","from":"developer"},{"body":"Aha, got it, seems this is make sense for normal receivers, but for reliable receivers, I'm not sure is it still make sense to push the following blocks if the previous block is failed.","from":"developer"},{"body":"Current situation: \n1. The TimeoutException is thrown by ReliableKafkaReceiver, receiver logs the error and do nothing else;\n2. Block pushing thread in BlockGenerator is terminated due to the exception;\n3. ReliableKafkaReceiver is ACTIVE but rogues afterward because blocksForPushing queue is full and no block pushing thread available to clean the queue.\n\nSo when the exception occurs, what user will see is the streaming application keeps running but sits there doing nothing (We are using some supervisor to restart the failed app. But if the app keeps running and doing nothing, we have to have some curator to monitor the data flow and kill the app in such situation). Can ReliableKafkaReceiver handle TimeoutException? For example retry pushing block when a TimeoutException is thrown? Or restart itself? Or at least fail the streaming application?","from":"developer"},{"body":"Yeah, I totally understand what is your requirement, maybe we should figure out a more elegant way to deal with such situation. Currently there's a solution trying to partially fix the problem you mentioned, you can refer to (https://github.com/apache/spark/pull/3655). But I think it is not a thorough way because the push thread will finally exited and block the whole system.","from":"developer"},{"body":"I believe https://github.com/apache/spark/pull/3655 fixes the problem. The block pushing thread will not exit when an exception happens because the exception is handled inside storeBlockAndCommitOffset. The change stops the receiver after three retries, which will also stop the BlockGenerator inside ReliableKafkaReceiver. What would happen when the receiver is stopped? Does the application fail?","from":"developer"},{"body":"[~superxma] is this resolved then?","from":"developer"}],"created":"2015-01-13T15:08:33.000+0000","description":"I am running a Spark streaming application with ReliableKafkaReceiver. It uses BlockGenerator to push blocks to BlockManager. However, writing WALs to HDFS may time out that causes keepPushingBlocks in BlockGenerator to terminate.\n\n15/01/12 19:07:06 ERROR receiver.BlockGenerator: Error in block pushing thread\njava.util.concurrent.TimeoutException: Futures timed out after [30 seconds]\n at scala.concurrent.impl.Promise$DefaultPromise.ready(Promise.scala:219)\n at scala.concurrent.impl.Promise$DefaultPromise.result(Promise.scala:223)\n at scala.concurrent.Await$$anonfun$result$1.apply(package.scala:107)\n at scala.concurrent.BlockContext$DefaultBlockContext$.blockOn(BlockContext.scala:53)\n at scala.concurrent.Await$.result(package.scala:107)\n at org.apache.spark.streaming.receiver.WriteAheadLogBasedBlockHandler.storeBlock(ReceivedBlockHandler.scala:176)\n at org.apache.spark.streaming.receiver.ReceiverSupervisorImpl.pushAndReportBlock(ReceiverSupervisorImpl.scala:160)\n at org.apache.spark.streaming.receiver.ReceiverSupervisorImpl.pushArrayBuffer(ReceiverSupervisorImpl.scala:126)\n at org.apache.spark.streaming.receiver.Receiver.store(Receiver.scala:124)\n at org.apache.spark.streaming.kafka.ReliableKafkaReceiver.org$apache$spark$streaming$kafka$ReliableKafkaReceiver$$storeBlockAndCommitOffset(ReliableKafkaReceiver.scala:207)\n at org.apache.spark.streaming.kafka.ReliableKafkaReceiver$GeneratedBlockHandler.onPushBlock(ReliableKafkaReceiver.scala:275)\n at org.apache.spark.streaming.receiver.BlockGenerator.pushBlock(BlockGenerator.scala:181)\n at org.apache.spark.streaming.receiver.BlockGenerator.org$apache$spark$streaming$receiver$BlockGenerator$$keepPushingBlocks(BlockGenerator.scala:154)\n at org.apache.spark.streaming.receiver.BlockGenerator$$anon$1.run(BlockGenerator.scala:86)\n\nThen the block pushing thread is done and no subsequent blocks can be pushed into blockManager. In turn this blocks receiver from receiving new data.\n\nSo when running my app and the TimeoutException happens, the ReliableKafkaReceiver stays in ACTIVE status but doesn't do anything at all. The application rogues.","issue_id":"12767147","key":"SPARK-5220","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-05-20T08:19:51.000+0000","role":"fixed_distractor","summary":"keepPushingBlocks in BlockGenerator terminated when an exception occurs, which causes the block pushing thread to terminate and blocks receiver "} {"case_id":"12769489","cluster":"DISTRACTOR-SPARK-5371","comments":[{"body":"User 'marmbrus' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/5278","created":"2015-03-31T02:30:22.492+0000"},{"body":"Issue resolved by pull request 5278\n[https://github.com/apache/spark/pull/5278]","created":"2015-03-31T18:44:07.176+0000"}],"conversations":[{"body":"This SQL session:\n\n{code}\nDROP TABLE\n test1;\nDROP TABLE\n test2;\nCREATE TABLE\n test1\n (\n c11 INT,\n c12 INT,\n c13 INT,\n c14 INT\n );\nCREATE TABLE\n test2\n (\n c21 INT,\n c22 INT,\n c23 INT,\n c24 INT\n );\nSELECT\n MIN(t3.c_1),\n MIN(t3.c_2),\n MIN(t3.c_3),\n MIN(t3.c_4)\nFROM\n (\n SELECT\n SUM(t1.c11) c_1,\n NULL c_2,\n NULL c_3,\n NULL c_4\n FROM\n test1 t1\n UNION ALL\n SELECT\n NULL c_1,\n SUM(t2.c22) c_2,\n SUM(t2.c23) c_3,\n SUM(t2.c24) c_4\n FROM\n test2 t2 ) t3; \n{code}\n\nProduces this error:\n\n{code}\n15/01/23 00:25:21 INFO thriftserver.SparkExecuteStatementOperation: Running query 'SELECT\n MIN(t3.c_1),\n MIN(t3.c_2),\n MIN(t3.c_3),\n MIN(t3.c_4)\nFROM\n (\n SELECT\n SUM(t1.c11) c_1,\n NULL c_2,\n NULL c_3,\n NULL c_4\n FROM\n test1 t1\n UNION ALL\n SELECT\n NULL c_1,\n SUM(t2.c22) c_2,\n SUM(t2.c23) c_3,\n SUM(t2.c24) c_4\n FROM\n test2 t2 ) t3'\n15/01/23 00:25:21 INFO parse.ParseDriver: Parsing command: SELECT\n MIN(t3.c_1),\n MIN(t3.c_2),\n MIN(t3.c_3),\n MIN(t3.c_4)\nFROM\n (\n SELECT\n SUM(t1.c11) c_1,\n NULL c_2,\n NULL c_3,\n NULL c_4\n FROM\n test1 t1\n UNION ALL\n SELECT\n NULL c_1,\n SUM(t2.c22) c_2,\n SUM(t2.c23) c_3,\n SUM(t2.c24) c_4\n FROM\n test2 t2 ) t3\n15/01/23 00:25:21 INFO parse.ParseDriver: Parse Completed\n15/01/23 00:25:21 ERROR thriftserver.SparkExecuteStatementOperation: Error executing query:\njava.util.NoSuchElementException: key not found: c_2#23488\n\tat scala.collection.MapLike$class.default(MapLike.scala:228)\n\tat org.apache.spark.sql.catalyst.expressions.AttributeMap.default(AttributeMap.scala:29)\n\tat scala.collection.MapLike$class.apply(MapLike.scala:141)\n\tat org.apache.spark.sql.catalyst.expressions.AttributeMap.apply(AttributeMap.scala:29)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$1.applyOrElse(Optimizer.scala:77)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$1.applyOrElse(Optimizer.scala:76)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:144)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:135)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$.pushToRight(Optimizer.scala:76)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$apply$1$$anonfun$applyOrElse$6.apply(Optimizer.scala:98)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$apply$1$$anonfun$applyOrElse$6.apply(Optimizer.scala:98)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:244)\n\tat scala.collection.AbstractTraversable.map(Traversable.scala:105)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$apply$1.applyOrElse(Optimizer.scala:98)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$apply$1.applyOrElse(Optimizer.scala:85)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:144)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$4.apply(TreeNode.scala:162)\n\tat scala.collection.Iterator$$anon$11.next(Iterator.scala:328)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:727)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1157)\n\tat scala.collection.generic.Growable$class.$plus$plus$eq(Growable.scala:48)\n\tat scala.collection.mutable.ArrayBuffer.$plus$plus$eq(ArrayBuffer.scala:103)\n\tat scala.collection.mutable.ArrayBuffer.$plus$plus$eq(ArrayBuffer.scala:47)\n\tat scala.collection.TraversableOnce$class.to(TraversableOnce.scala:273)\n\tat scala.collection.AbstractIterator.to(Iterator.scala:1157)\n\tat scala.collection.TraversableOnce$class.toBuffer(TraversableOnce.scala:265)\n\tat scala.collection.AbstractIterator.toBuffer(Iterator.scala:1157)\n\tat scala.collection.TraversableOnce$class.toArray(TraversableOnce.scala:252)\n\tat scala.collection.AbstractIterator.toArray(Iterator.scala:1157)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildrenDown(TreeNode.scala:191)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:147)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:135)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$.apply(Optimizer.scala:85)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$.apply(Optimizer.scala:59)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$apply$1$$anonfun$apply$2.apply(RuleExecutor.scala:61)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$apply$1$$anonfun$apply$2.apply(RuleExecutor.scala:59)\n\tat scala.collection.IndexedSeqOptimized$class.foldl(IndexedSeqOptimized.scala:51)\n\tat scala.collection.IndexedSeqOptimized$class.foldLeft(IndexedSeqOptimized.scala:60)\n\tat scala.collection.mutable.WrappedArray.foldLeft(WrappedArray.scala:34)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$apply$1.apply(RuleExecutor.scala:59)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$apply$1.apply(RuleExecutor.scala:51)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor.apply(RuleExecutor.scala:51)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.optimizedPlan$lzycompute(SQLContext.scala:462)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.optimizedPlan(SQLContext.scala:462)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.sparkPlan$lzycompute(SQLContext.scala:467)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.sparkPlan(SQLContext.scala:465)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.executedPlan$lzycompute(SQLContext.scala:471)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.executedPlan(SQLContext.scala:471)\n\tat org.apache.spark.sql.SchemaRDD.collect(SchemaRDD.scala:463)\n\tat org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:178)\n\tat org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n\tat org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n\tat sun.reflect.GeneratedMethodAccessor61.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:483)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n\tat org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n\tat com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n\tat org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n\tat org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n\tat org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n\tat org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n\tat org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n\tat org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n\tat org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n\tat org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n15/01/23 00:25:22 WARN thrift.ThriftCLIService: Error executing statement:\norg.apache.hive.service.cli.HiveSQLException: java.util.NoSuchElementException: key not found: c_2#23488\n\tat org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:189)\n\tat org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n\tat org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n\tat sun.reflect.GeneratedMethodAccessor61.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:483)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n\tat org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n\tat com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n\tat org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n\tat org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n\tat org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n\tat org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n\tat org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n\tat org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n\tat org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n\tat org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}\n\nSome similar queries work. For example:\n\n{code}\nSELECT\n MIN(t3.c_1),\n MIN(t3.c_2),\n MIN(t3.c_3),\n MIN(t3.c_4)\nFROM\n (\n SELECT\n SUM(t1.c11) c_1,\n SUM(t1.c12) c_2,\n SUM(t1.c13) c_3,\n SUM(t1.c14) c_4\n FROM\n test1 t1\n UNION ALL\n SELECT\n SUM(t2.c21) c_1,\n SUM(t2.c22) c_2,\n SUM(t2.c23) c_3,\n SUM(t2.c24) c_4\n FROM\n test2 t2 ) t3; \n{code}\n\nWorks fine. Notice the only difference is the {{null}}.","from":"reporter","subject":"Failure to analyze query with UNION ALL and double aggregation"},{"body":"User 'marmbrus' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/5278","from":"developer"},{"body":"Issue resolved by pull request 5278\n[https://github.com/apache/spark/pull/5278]","from":"developer"}],"created":"2015-01-23T00:26:56.000+0000","description":"This SQL session:\n\n{code}\nDROP TABLE\n test1;\nDROP TABLE\n test2;\nCREATE TABLE\n test1\n (\n c11 INT,\n c12 INT,\n c13 INT,\n c14 INT\n );\nCREATE TABLE\n test2\n (\n c21 INT,\n c22 INT,\n c23 INT,\n c24 INT\n );\nSELECT\n MIN(t3.c_1),\n MIN(t3.c_2),\n MIN(t3.c_3),\n MIN(t3.c_4)\nFROM\n (\n SELECT\n SUM(t1.c11) c_1,\n NULL c_2,\n NULL c_3,\n NULL c_4\n FROM\n test1 t1\n UNION ALL\n SELECT\n NULL c_1,\n SUM(t2.c22) c_2,\n SUM(t2.c23) c_3,\n SUM(t2.c24) c_4\n FROM\n test2 t2 ) t3; \n{code}\n\nProduces this error:\n\n{code}\n15/01/23 00:25:21 INFO thriftserver.SparkExecuteStatementOperation: Running query 'SELECT\n MIN(t3.c_1),\n MIN(t3.c_2),\n MIN(t3.c_3),\n MIN(t3.c_4)\nFROM\n (\n SELECT\n SUM(t1.c11) c_1,\n NULL c_2,\n NULL c_3,\n NULL c_4\n FROM\n test1 t1\n UNION ALL\n SELECT\n NULL c_1,\n SUM(t2.c22) c_2,\n SUM(t2.c23) c_3,\n SUM(t2.c24) c_4\n FROM\n test2 t2 ) t3'\n15/01/23 00:25:21 INFO parse.ParseDriver: Parsing command: SELECT\n MIN(t3.c_1),\n MIN(t3.c_2),\n MIN(t3.c_3),\n MIN(t3.c_4)\nFROM\n (\n SELECT\n SUM(t1.c11) c_1,\n NULL c_2,\n NULL c_3,\n NULL c_4\n FROM\n test1 t1\n UNION ALL\n SELECT\n NULL c_1,\n SUM(t2.c22) c_2,\n SUM(t2.c23) c_3,\n SUM(t2.c24) c_4\n FROM\n test2 t2 ) t3\n15/01/23 00:25:21 INFO parse.ParseDriver: Parse Completed\n15/01/23 00:25:21 ERROR thriftserver.SparkExecuteStatementOperation: Error executing query:\njava.util.NoSuchElementException: key not found: c_2#23488\n\tat scala.collection.MapLike$class.default(MapLike.scala:228)\n\tat org.apache.spark.sql.catalyst.expressions.AttributeMap.default(AttributeMap.scala:29)\n\tat scala.collection.MapLike$class.apply(MapLike.scala:141)\n\tat org.apache.spark.sql.catalyst.expressions.AttributeMap.apply(AttributeMap.scala:29)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$1.applyOrElse(Optimizer.scala:77)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$1.applyOrElse(Optimizer.scala:76)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:144)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:135)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$.pushToRight(Optimizer.scala:76)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$apply$1$$anonfun$applyOrElse$6.apply(Optimizer.scala:98)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$apply$1$$anonfun$applyOrElse$6.apply(Optimizer.scala:98)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n\tat scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:244)\n\tat scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)\n\tat scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)\n\tat scala.collection.TraversableLike$class.map(TraversableLike.scala:244)\n\tat scala.collection.AbstractTraversable.map(Traversable.scala:105)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$apply$1.applyOrElse(Optimizer.scala:98)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$$anonfun$apply$1.applyOrElse(Optimizer.scala:85)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:144)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode$$anonfun$4.apply(TreeNode.scala:162)\n\tat scala.collection.Iterator$$anon$11.next(Iterator.scala:328)\n\tat scala.collection.Iterator$class.foreach(Iterator.scala:727)\n\tat scala.collection.AbstractIterator.foreach(Iterator.scala:1157)\n\tat scala.collection.generic.Growable$class.$plus$plus$eq(Growable.scala:48)\n\tat scala.collection.mutable.ArrayBuffer.$plus$plus$eq(ArrayBuffer.scala:103)\n\tat scala.collection.mutable.ArrayBuffer.$plus$plus$eq(ArrayBuffer.scala:47)\n\tat scala.collection.TraversableOnce$class.to(TraversableOnce.scala:273)\n\tat scala.collection.AbstractIterator.to(Iterator.scala:1157)\n\tat scala.collection.TraversableOnce$class.toBuffer(TraversableOnce.scala:265)\n\tat scala.collection.AbstractIterator.toBuffer(Iterator.scala:1157)\n\tat scala.collection.TraversableOnce$class.toArray(TraversableOnce.scala:252)\n\tat scala.collection.AbstractIterator.toArray(Iterator.scala:1157)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformChildrenDown(TreeNode.scala:191)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transformDown(TreeNode.scala:147)\n\tat org.apache.spark.sql.catalyst.trees.TreeNode.transform(TreeNode.scala:135)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$.apply(Optimizer.scala:85)\n\tat org.apache.spark.sql.catalyst.optimizer.UnionPushdown$.apply(Optimizer.scala:59)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$apply$1$$anonfun$apply$2.apply(RuleExecutor.scala:61)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$apply$1$$anonfun$apply$2.apply(RuleExecutor.scala:59)\n\tat scala.collection.IndexedSeqOptimized$class.foldl(IndexedSeqOptimized.scala:51)\n\tat scala.collection.IndexedSeqOptimized$class.foldLeft(IndexedSeqOptimized.scala:60)\n\tat scala.collection.mutable.WrappedArray.foldLeft(WrappedArray.scala:34)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$apply$1.apply(RuleExecutor.scala:59)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor$$anonfun$apply$1.apply(RuleExecutor.scala:51)\n\tat scala.collection.immutable.List.foreach(List.scala:318)\n\tat org.apache.spark.sql.catalyst.rules.RuleExecutor.apply(RuleExecutor.scala:51)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.optimizedPlan$lzycompute(SQLContext.scala:462)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.optimizedPlan(SQLContext.scala:462)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.sparkPlan$lzycompute(SQLContext.scala:467)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.sparkPlan(SQLContext.scala:465)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.executedPlan$lzycompute(SQLContext.scala:471)\n\tat org.apache.spark.sql.SQLContext$QueryExecution.executedPlan(SQLContext.scala:471)\n\tat org.apache.spark.sql.SchemaRDD.collect(SchemaRDD.scala:463)\n\tat org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:178)\n\tat org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n\tat org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n\tat sun.reflect.GeneratedMethodAccessor61.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:483)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n\tat org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n\tat com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n\tat org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n\tat org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n\tat org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n\tat org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n\tat org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n\tat org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n\tat org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n\tat org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n15/01/23 00:25:22 WARN thrift.ThriftCLIService: Error executing statement:\norg.apache.hive.service.cli.HiveSQLException: java.util.NoSuchElementException: key not found: c_2#23488\n\tat org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:189)\n\tat org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n\tat org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n\tat sun.reflect.GeneratedMethodAccessor61.invoke(Unknown Source)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:483)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n\tat java.security.AccessController.doPrivileged(Native Method)\n\tat javax.security.auth.Subject.doAs(Subject.java:422)\n\tat org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n\tat org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n\tat org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n\tat com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n\tat org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n\tat org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n\tat org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n\tat org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n\tat org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n\tat org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n\tat org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n\tat org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}\n\nSome similar queries work. For example:\n\n{code}\nSELECT\n MIN(t3.c_1),\n MIN(t3.c_2),\n MIN(t3.c_3),\n MIN(t3.c_4)\nFROM\n (\n SELECT\n SUM(t1.c11) c_1,\n SUM(t1.c12) c_2,\n SUM(t1.c13) c_3,\n SUM(t1.c14) c_4\n FROM\n test1 t1\n UNION ALL\n SELECT\n SUM(t2.c21) c_1,\n SUM(t2.c22) c_2,\n SUM(t2.c23) c_3,\n SUM(t2.c24) c_4\n FROM\n test2 t2 ) t3; \n{code}\n\nWorks fine. Notice the only difference is the {{null}}.","issue_id":"12769489","key":"SPARK-5371","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-03-31T18:44:07.000+0000","role":"fixed_distractor","summary":"Failure to analyze query with UNION ALL and double aggregation"} {"case_id":"13630339","cluster":"DISTRACTOR-SPARK-53759","comments":[{"body":"I don't know if this is related to Windows 11, somehow? I have the exact same issue on pyspark versions 3.5.2 and 4.0.0, but I know that it used to work for 3.5.2 on my previous machine running Windows 10.","created":"2025-10-08T12:56:38.295+0000"},{"body":"Interesting, we have reproduced it on at least Windows 10 22H2, Windows Server 2022 21H2, and Windows 11 Enterprise 24H2","created":"2025-11-06T00:38:54.318+0000"},{"body":"I can confirm the issue on Windows 11 Home 25H2 and Python versions 3.12, 3.13 and 314, it is easily reproducible.\r\n\r\nWe ran into this when we switched from Python 3.10 where createDataFrame worked without an issue.","created":"2026-01-08T13:22:52.195+0000"},{"body":"[~gurwls223], I see that you are the shepherd of this issue. Given its severity, shouldn't it be prioritized and get an actual assignee?","created":"2026-02-12T13:00:04.350+0000"},{"body":"I was able to reproduce this issue on Windows and investigated further to find the root cause.\r\n\r\nThe problem is in the simple-worker codepath, which Windows always uses because {{os.fork()}} is unavailable. When the worker finishes processing a task, it writes results to a buffered socket ({{BufferedRWPair}} with a 64KB buffer created via {{sock.makefile()}}). The data sits in this buffer until someone calls {{flush()}} — but the simple-worker path never does. It relies on Python's garbage collector to flush the buffer during interpreter shutdown.\r\n\r\nThis worked on Python 3.11 because the buffer happened to be flushed before the socket was torn down. On Python 3.12+, CPython [moved GC to the eval breaker|https://github.com/python/cpython/issues/97922], which changed the order objects are finalized during shutdown. Now the socket closes first, and when the buffer tries to flush, the data is silently lost. The JVM, still waiting to read results, sees {{EOFException}}.\r\n\r\nThe daemon path ({{daemon.py}}) doesn't have this problem because it already has an explicit {{outfile.flush()}} in a {{finally}} block. The simple-worker path was just missing the same pattern.\r\n\r\nThis appears to be resolved by [PR #54458|https://github.com/apache/spark/pull/54458] (SPARK-55665), which unified worker socket handling and added the missing flush. I tested every available pre-release on PyPI and confirmed the fix landed in {{pyspark==4.2.0.dev3}} (dev1 and dev2 are still affected). All current stable releases through 4.1.1 have the bug.\r\n\r\n*Workarounds for anyone on a stable release:*\r\n- Install the pre-release: {{pip install --pre pyspark==4.2.0.dev3}}\r\n- Use Python 3.11 (the bug does not manifest on 3.11)\r\n- On Linux/macOS, use the default daemon mode (only the simple-worker path is affected)\r\n\r\nI put together a reproducer with a test matrix and full root cause analysis here: https://github.com/anblanco/spark53759-reproducer\r\n","created":"2026-04-04T23:58:12.149+0000"},{"body":"> This appears to be resolved by PR #54458 (SPARK-55665)\r\n\r\n \r\nWill this be backported to any of the earlier releases of pyspark? ","created":"2026-04-05T09:37:47.762+0000"},{"body":"[~aablanco] Thank you! I've tested it on 4.2.0.dev3 and it works for me.","created":"2026-04-07T10:02:21.344+0000"},{"body":"Issue resolved by pull request 55201\n[https://github.com/apache/spark/pull/55201]","created":"2026-04-08T20:47:17.872+0000"},{"body":"This landed at branch-4.0 via https://github.com/apache/spark/pull/55224","created":"2026-04-08T21:40:10.207+0000"},{"body":"This landed at branch-3.5 via https://github.com/apache/spark/pull/55225","created":"2026-04-08T21:42:00.455+0000"}],"conversations":[{"body":"Python 3.12+ crashes locally on Windows when using the `createDataFrame` API. All dataframe creation methods in the [Quickstart: Dataframe|https://spark.apache.org/docs/latest/api/python/getting_started/quickstart_df.html#DataFrame-Creation] seem to crash but other operations work as expected.\r\n\r\nReproduction:\r\n{code:java}\r\nimport os\r\nimport sys\r\nfrom pyspark.sql import SparkSession\r\n\r\nos.environ[\"PYSPARK_PYTHON\"] = sys.executable\r\nspark = SparkSession.builder.getOrCreate()\r\n\r\ndf = spark.createDataFrame([(1,), (2,)], [\"myint\"])\r\ndf.show() {code}\r\n \r\n\r\nStack trace. This is with \"spark.python.worker.faulthandler.enabled\" enabled, but the stack trace is the same with it disabled:\r\n{code:java}\r\nTraceback (most recent call last):\r\nFile \".py\", line 10, in \r\n df.show()\r\nFile \"/sql/classic/dataframe.py\", line 285, in show\r\n print(self._show_string(n, truncate, vertical))\r\nFile \"/sql/classic/dataframe.py\", line 303, in _show_string\r\n return self._jdf.showString(n, 20, vertical)\r\nFile \"/java_gateway.py\", line 1362, in __call__\r\n return_value = get_return_value(\r\n answer, self.gateway_client, self.target_id, self.name)\r\nFile \"/errors/exceptions/captured.py\", line 282, in deco\r\n return f(*a, **kw)\r\nFile \"/protocol.py\", line 327, in get_return_value\r\n raise Py4JJavaError(\r\n \"An error occurred while calling {0}{1}{2}.\\n\".\r\n format(target_id, \".\", name), value)\r\npy4j.protocol.Py4JJavaError: An error occurred while calling o48.showString.:\r\norg.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 0.0 failed 1 times, most recent failure: Lost task 0.0 in stage 0.0 (TID 0) (executor driver): org.apache.spark.SparkException: Python worker exited unexpectedly (crashed)\r\nat org.apache.spark.api.python.BasePythonRunner$ReaderIterator$$anonfun$1.applyOrElse(PythonRunner.scala:624)\r\nat org.apache.spark.api.python.BasePythonRunner$ReaderIterator$$anonfun$1.applyOrElse(PythonRunner.scala:599)\r\nat scala.runtime.AbstractPartialFunction.apply(AbstractPartialFunction.scala:35)\r\nat org.apache.spark.api.python.PythonRunner$$anon$3.read(PythonRunner.scala:945)\r\nat org.apache.spark.api.python.PythonRunner$$anon$3.read(PythonRunner.scala:925)\r\nat org.apache.spark.api.python.BasePythonRunner$ReaderIterator.hasNext(PythonRunner.scala:532)\r\nat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:37)\r\nat scala.collection.Iterator$$anon$10.hasNext(Iterator.scala:601)\r\nat scala.collection.Iterator$$anon$9.hasNext(Iterator.scala:583)\r\nat scala.collection.Iterator$$anon$9.hasNext(Iterator.scala:583)\r\nat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)\r\nat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\nat org.apache.spark.sql.execution.WholeStageCodegenEvaluatorFactory$WholeStageCodegenPartitionEvaluator$$anon$1.hasNext(WholeStageCodegenEvaluatorFactory.scala:50)\r\nat org.apache.spark.sql.execution.SparkPlan.$anonfun$getByteArrayRdd$1(SparkPlan.scala:402)\r\nat org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2(RDD.scala:901)\r\nat org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2$adapted(RDD.scala:901)\r\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:52)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:374)\r\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:338)\r\nat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:93)\r\nat org.apache.spark.TaskContext.runTaskWithListeners(TaskContext.scala:171)\r\nat org.apache.spark.scheduler.Task.run(Task.scala:147)\r\nat org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$5(Executor.scala:647)\r\nat org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally(SparkErrorUtils.scala:80)\r\nat org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally$(SparkErrorUtils.scala:77) at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:99)\r\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:650)\r\nat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)\r\nat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)\r\nat java.base/java.lang.Thread.run(Thread.java:840)Caused by: java.io.EOFException\r\nat java.base/java.io.DataInputStream.readInt(DataInputStream.java:386)\r\nat org.apache.spark.api.python.PythonRunner$$anon$3.read(PythonRunner.scala:933)\r\n... 26 more{code}\r\n\r\nNote, I originally posted in the Python 3.13 issue but it seemed better to create a new issue since that one was closed. Please let us know if our team can help debug this further, it seems relatively low level. Thank you!","from":"reporter","subject":"Fix missing flush in simple-worker path"},{"body":"I don't know if this is related to Windows 11, somehow? I have the exact same issue on pyspark versions 3.5.2 and 4.0.0, but I know that it used to work for 3.5.2 on my previous machine running Windows 10.","from":"developer"},{"body":"Interesting, we have reproduced it on at least Windows 10 22H2, Windows Server 2022 21H2, and Windows 11 Enterprise 24H2","from":"developer"},{"body":"I can confirm the issue on Windows 11 Home 25H2 and Python versions 3.12, 3.13 and 314, it is easily reproducible.\r\n\r\nWe ran into this when we switched from Python 3.10 where createDataFrame worked without an issue.","from":"developer"},{"body":"[~gurwls223], I see that you are the shepherd of this issue. Given its severity, shouldn't it be prioritized and get an actual assignee?","from":"developer"},{"body":"I was able to reproduce this issue on Windows and investigated further to find the root cause.\r\n\r\nThe problem is in the simple-worker codepath, which Windows always uses because {{os.fork()}} is unavailable. When the worker finishes processing a task, it writes results to a buffered socket ({{BufferedRWPair}} with a 64KB buffer created via {{sock.makefile()}}). The data sits in this buffer until someone calls {{flush()}} — but the simple-worker path never does. It relies on Python's garbage collector to flush the buffer during interpreter shutdown.\r\n\r\nThis worked on Python 3.11 because the buffer happened to be flushed before the socket was torn down. On Python 3.12+, CPython [moved GC to the eval breaker|https://github.com/python/cpython/issues/97922], which changed the order objects are finalized during shutdown. Now the socket closes first, and when the buffer tries to flush, the data is silently lost. The JVM, still waiting to read results, sees {{EOFException}}.\r\n\r\nThe daemon path ({{daemon.py}}) doesn't have this problem because it already has an explicit {{outfile.flush()}} in a {{finally}} block. The simple-worker path was just missing the same pattern.\r\n\r\nThis appears to be resolved by [PR #54458|https://github.com/apache/spark/pull/54458] (SPARK-55665), which unified worker socket handling and added the missing flush. I tested every available pre-release on PyPI and confirmed the fix landed in {{pyspark==4.2.0.dev3}} (dev1 and dev2 are still affected). All current stable releases through 4.1.1 have the bug.\r\n\r\n*Workarounds for anyone on a stable release:*\r\n- Install the pre-release: {{pip install --pre pyspark==4.2.0.dev3}}\r\n- Use Python 3.11 (the bug does not manifest on 3.11)\r\n- On Linux/macOS, use the default daemon mode (only the simple-worker path is affected)\r\n\r\nI put together a reproducer with a test matrix and full root cause analysis here: https://github.com/anblanco/spark53759-reproducer\r\n","from":"developer"},{"body":"> This appears to be resolved by PR #54458 (SPARK-55665)\r\n\r\n \r\nWill this be backported to any of the earlier releases of pyspark? ","from":"developer"},{"body":"[~aablanco] Thank you! I've tested it on 4.2.0.dev3 and it works for me.","from":"developer"},{"body":"Issue resolved by pull request 55201\n[https://github.com/apache/spark/pull/55201]","from":"developer"},{"body":"This landed at branch-4.0 via https://github.com/apache/spark/pull/55224","from":"developer"},{"body":"This landed at branch-3.5 via https://github.com/apache/spark/pull/55225","from":"developer"}],"created":"2025-09-30T11:28:34.000+0000","description":"Python 3.12+ crashes locally on Windows when using the `createDataFrame` API. All dataframe creation methods in the [Quickstart: Dataframe|https://spark.apache.org/docs/latest/api/python/getting_started/quickstart_df.html#DataFrame-Creation] seem to crash but other operations work as expected.\r\n\r\nReproduction:\r\n{code:java}\r\nimport os\r\nimport sys\r\nfrom pyspark.sql import SparkSession\r\n\r\nos.environ[\"PYSPARK_PYTHON\"] = sys.executable\r\nspark = SparkSession.builder.getOrCreate()\r\n\r\ndf = spark.createDataFrame([(1,), (2,)], [\"myint\"])\r\ndf.show() {code}\r\n \r\n\r\nStack trace. This is with \"spark.python.worker.faulthandler.enabled\" enabled, but the stack trace is the same with it disabled:\r\n{code:java}\r\nTraceback (most recent call last):\r\nFile \".py\", line 10, in \r\n df.show()\r\nFile \"/sql/classic/dataframe.py\", line 285, in show\r\n print(self._show_string(n, truncate, vertical))\r\nFile \"/sql/classic/dataframe.py\", line 303, in _show_string\r\n return self._jdf.showString(n, 20, vertical)\r\nFile \"/java_gateway.py\", line 1362, in __call__\r\n return_value = get_return_value(\r\n answer, self.gateway_client, self.target_id, self.name)\r\nFile \"/errors/exceptions/captured.py\", line 282, in deco\r\n return f(*a, **kw)\r\nFile \"/protocol.py\", line 327, in get_return_value\r\n raise Py4JJavaError(\r\n \"An error occurred while calling {0}{1}{2}.\\n\".\r\n format(target_id, \".\", name), value)\r\npy4j.protocol.Py4JJavaError: An error occurred while calling o48.showString.:\r\norg.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 0.0 failed 1 times, most recent failure: Lost task 0.0 in stage 0.0 (TID 0) (executor driver): org.apache.spark.SparkException: Python worker exited unexpectedly (crashed)\r\nat org.apache.spark.api.python.BasePythonRunner$ReaderIterator$$anonfun$1.applyOrElse(PythonRunner.scala:624)\r\nat org.apache.spark.api.python.BasePythonRunner$ReaderIterator$$anonfun$1.applyOrElse(PythonRunner.scala:599)\r\nat scala.runtime.AbstractPartialFunction.apply(AbstractPartialFunction.scala:35)\r\nat org.apache.spark.api.python.PythonRunner$$anon$3.read(PythonRunner.scala:945)\r\nat org.apache.spark.api.python.PythonRunner$$anon$3.read(PythonRunner.scala:925)\r\nat org.apache.spark.api.python.BasePythonRunner$ReaderIterator.hasNext(PythonRunner.scala:532)\r\nat org.apache.spark.InterruptibleIterator.hasNext(InterruptibleIterator.scala:37)\r\nat scala.collection.Iterator$$anon$10.hasNext(Iterator.scala:601)\r\nat scala.collection.Iterator$$anon$9.hasNext(Iterator.scala:583)\r\nat scala.collection.Iterator$$anon$9.hasNext(Iterator.scala:583)\r\nat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)\r\nat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\r\nat org.apache.spark.sql.execution.WholeStageCodegenEvaluatorFactory$WholeStageCodegenPartitionEvaluator$$anon$1.hasNext(WholeStageCodegenEvaluatorFactory.scala:50)\r\nat org.apache.spark.sql.execution.SparkPlan.$anonfun$getByteArrayRdd$1(SparkPlan.scala:402)\r\nat org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2(RDD.scala:901)\r\nat org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2$adapted(RDD.scala:901)\r\nat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:52)\r\nat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:374)\r\nat org.apache.spark.rdd.RDD.iterator(RDD.scala:338)\r\nat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:93)\r\nat org.apache.spark.TaskContext.runTaskWithListeners(TaskContext.scala:171)\r\nat org.apache.spark.scheduler.Task.run(Task.scala:147)\r\nat org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$5(Executor.scala:647)\r\nat org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally(SparkErrorUtils.scala:80)\r\nat org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally$(SparkErrorUtils.scala:77) at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:99)\r\nat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:650)\r\nat java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)\r\nat java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)\r\nat java.base/java.lang.Thread.run(Thread.java:840)Caused by: java.io.EOFException\r\nat java.base/java.io.DataInputStream.readInt(DataInputStream.java:386)\r\nat org.apache.spark.api.python.PythonRunner$$anon$3.read(PythonRunner.scala:933)\r\n... 26 more{code}\r\n\r\nNote, I originally posted in the Python 3.13 issue but it seemed better to create a new issue since that one was closed. Please let us know if our team can help debug this further, it seems relatively low level. Thank you!","issue_id":"13630339","key":"SPARK-53759","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2026-04-08T20:47:17.000+0000","role":"fixed_distractor","summary":"Fix missing flush in simple-worker path"} {"case_id":"12769777","cluster":"DISTRACTOR-SPARK-5391","comments":[{"body":"Same error occurred when a table is created with json serde in hive table and queried from SparkQL.","created":"2015-02-06T05:13:39.497+0000"},{"body":"Is this still a problem? Is there a reason you aren't using the native JSON support?","created":"2015-09-15T21:39:04.789+0000"},{"body":"Haven't tried native JSON but looks promising, so this ticket is probably lower priority.","created":"2015-09-15T21:47:07.697+0000"},{"body":"I tried this, it worked in master, will close this as resolved.","created":"2015-10-13T18:33:41.786+0000"}],"conversations":[{"body":"- Using Spark built from trunk on this commit: https://github.com/apache/spark/commit/bc20a52b34e826895d0dcc1d783c021ebd456ebd\n- Build for Hive13\n- Using this JSON serde: https://github.com/rcongiu/Hive-JSON-Serde\n\nFirst download jar locally:\n{code}\n$ curl http://www.congiu.net/hive-json-serde/1.3/cdh5/json-serde-1.3-jar-with-dependencies.jar > /tmp/json-serde-1.3-jar-with-dependencies.jar\n{code}\n\nThen add it in SparkSQL session:\n{code}\nadd jar /tmp/json-serde-1.3-jar-with-dependencies.jar\n{code}\n\nFinally create table:\n{code}\ncreate table test_json (c1 boolean) ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe';\n{code}\n\nLogs for add jar:\n{code}\n15/01/23 23:48:33 INFO thriftserver.SparkExecuteStatementOperation: Running query 'add jar /tmp/json-serde-1.3-jar-with-dependencies.jar'\n15/01/23 23:48:34 INFO session.SessionState: No Tez session required at this point. hive.execution.engine=mr.\n15/01/23 23:48:34 INFO SessionState: Added /tmp/json-serde-1.3-jar-with-dependencies.jar to class path\n15/01/23 23:48:34 INFO SessionState: Added resource: /tmp/json-serde-1.3-jar-with-dependencies.jar\n15/01/23 23:48:34 INFO spark.SparkContext: Added JAR /tmp/json-serde-1.3-jar-with-dependencies.jar at http://192.168.99.9:51312/jars/json-serde-1.3-jar-with-dependencies.jar with timestamp 1422056914776\n15/01/23 23:48:34 INFO thriftserver.SparkExecuteStatementOperation: Result Schema: List()\n15/01/23 23:48:34 INFO thriftserver.SparkExecuteStatementOperation: Result Schema: List()\n{code}\n\nLogs (with error) for create table:\n{code}\n15/01/23 23:49:00 INFO thriftserver.SparkExecuteStatementOperation: Running query 'create table test_json (c1 boolean) ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe''\n15/01/23 23:49:00 INFO parse.ParseDriver: Parsing command: create table test_json (c1 boolean) ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe'\n15/01/23 23:49:01 INFO parse.ParseDriver: Parse Completed\n15/01/23 23:49:01 INFO session.SessionState: No Tez session required at this point. hive.execution.engine=mr.\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO ql.Driver: Concurrency mode is disabled, not creating a lock manager\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO parse.ParseDriver: Parsing command: create table test_json (c1 boolean) ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe'\n15/01/23 23:49:01 INFO parse.ParseDriver: Parse Completed\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO parse.SemanticAnalyzer: Starting Semantic Analysis\n15/01/23 23:49:01 INFO parse.SemanticAnalyzer: Creating table test_json position=13\n15/01/23 23:49:01 INFO ql.Driver: Semantic Analysis Completed\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO ql.Driver: Returning Hive schema: Schema(fieldSchemas:null, properties:null)\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO ql.Driver: Starting command: create table test_json (c1 boolean) ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe'\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 WARN security.ShellBasedUnixGroupsMapping: got exception trying to get groups for user anonymous\norg.apache.hadoop.util.Shell$ExitCodeException: id: anonymous: No such user\n\n at org.apache.hadoop.util.Shell.runCommand(Shell.java:505)\n at org.apache.hadoop.util.Shell.run(Shell.java:418)\n at org.apache.hadoop.util.Shell$ShellCommandExecutor.execute(Shell.java:650)\n at org.apache.hadoop.util.Shell.execCommand(Shell.java:739)\n at org.apache.hadoop.util.Shell.execCommand(Shell.java:722)\n at org.apache.hadoop.security.ShellBasedUnixGroupsMapping.getUnixGroups(ShellBasedUnixGroupsMapping.java:83)\n at org.apache.hadoop.security.ShellBasedUnixGroupsMapping.getGroups(ShellBasedUnixGroupsMapping.java:52)\n at org.apache.hadoop.security.JniBasedUnixGroupsMappingWithFallback.getGroups(JniBasedUnixGroupsMappingWithFallback.java:50)\n at org.apache.hadoop.security.Groups.getGroups(Groups.java:139)\n at org.apache.hadoop.security.UserGroupInformation.getGroupNames(UserGroupInformation.java:1409)\n at org.apache.hadoop.hive.ql.security.HadoopDefaultAuthenticator.setConf(HadoopDefaultAuthenticator.java:63)\n at org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:73)\n at org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:133)\n at org.apache.hadoop.hive.ql.metadata.HiveUtils.getAuthenticator(HiveUtils.java:424)\n at org.apache.hadoop.hive.ql.session.SessionState.setupAuth(SessionState.java:377)\n at org.apache.hadoop.hive.ql.session.SessionState.getAuthenticator(SessionState.java:867)\n at org.apache.hadoop.hive.ql.session.SessionState.getUserFromAuthenticator(SessionState.java:589)\n at org.apache.hadoop.hive.ql.metadata.Table.getEmptyTable(Table.java:174)\n at org.apache.hadoop.hive.ql.metadata.Table.(Table.java:116)\n at org.apache.hadoop.hive.ql.metadata.Hive.newTable(Hive.java:2566)\n at org.apache.hadoop.hive.ql.exec.DDLTask.createTable(DDLTask.java:4046)\n at org.apache.hadoop.hive.ql.exec.DDLTask.execute(DDLTask.java:281)\n at org.apache.hadoop.hive.ql.exec.Task.executeTask(Task.java:153)\n at org.apache.hadoop.hive.ql.exec.TaskRunner.runSequential(TaskRunner.java:85)\n at org.apache.hadoop.hive.ql.Driver.launchTask(Driver.java:1503)\n at org.apache.hadoop.hive.ql.Driver.execute(Driver.java:1270)\n at org.apache.hadoop.hive.ql.Driver.runInternal(Driver.java:1088)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:911)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:901)\n at org.apache.spark.sql.hive.HiveContext.runHive(HiveContext.scala:292)\n at org.apache.spark.sql.hive.HiveContext.runSqlHive(HiveContext.scala:264)\n at org.apache.spark.sql.hive.execution.HiveNativeCommand.run(HiveNativeCommand.scala:37)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult$lzycompute(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.execute(commands.scala:61)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd$lzycompute(SQLContext.scala:474)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd(SQLContext.scala:474)\n at org.apache.spark.sql.SchemaRDDLike$class.$init$(SchemaRDDLike.scala:58)\n at org.apache.spark.sql.SchemaRDD.(SchemaRDD.scala:107)\n at org.apache.spark.sql.hive.HiveContext.sql(HiveContext.scala:73)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:160)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n at org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n at org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n at org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n at com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n at org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n at org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n at org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n15/01/23 23:49:01 WARN security.UserGroupInformation: No groups available for user anonymous\n15/01/23 23:49:01 WARN security.ShellBasedUnixGroupsMapping: got exception trying to get groups for user anonymous\norg.apache.hadoop.util.Shell$ExitCodeException: id: anonymous: No such user\n\n at org.apache.hadoop.util.Shell.runCommand(Shell.java:505)\n at org.apache.hadoop.util.Shell.run(Shell.java:418)\n at org.apache.hadoop.util.Shell$ShellCommandExecutor.execute(Shell.java:650)\n at org.apache.hadoop.util.Shell.execCommand(Shell.java:739)\n at org.apache.hadoop.util.Shell.execCommand(Shell.java:722)\n at org.apache.hadoop.security.ShellBasedUnixGroupsMapping.getUnixGroups(ShellBasedUnixGroupsMapping.java:83)\n at org.apache.hadoop.security.ShellBasedUnixGroupsMapping.getGroups(ShellBasedUnixGroupsMapping.java:52)\n at org.apache.hadoop.security.JniBasedUnixGroupsMappingWithFallback.getGroups(JniBasedUnixGroupsMappingWithFallback.java:50)\n at org.apache.hadoop.security.Groups.getGroups(Groups.java:139)\n at org.apache.hadoop.security.UserGroupInformation.getGroupNames(UserGroupInformation.java:1409)\n at org.apache.hadoop.hive.ql.security.HadoopDefaultAuthenticator.setConf(HadoopDefaultAuthenticator.java:64)\n at org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:73)\n at org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:133)\n at org.apache.hadoop.hive.ql.metadata.HiveUtils.getAuthenticator(HiveUtils.java:424)\n at org.apache.hadoop.hive.ql.session.SessionState.setupAuth(SessionState.java:377)\n at org.apache.hadoop.hive.ql.session.SessionState.getAuthenticator(SessionState.java:867)\n at org.apache.hadoop.hive.ql.session.SessionState.getUserFromAuthenticator(SessionState.java:589)\n at org.apache.hadoop.hive.ql.metadata.Table.getEmptyTable(Table.java:174)\n at org.apache.hadoop.hive.ql.metadata.Table.(Table.java:116)\n at org.apache.hadoop.hive.ql.metadata.Hive.newTable(Hive.java:2566)\n at org.apache.hadoop.hive.ql.exec.DDLTask.createTable(DDLTask.java:4046)\n at org.apache.hadoop.hive.ql.exec.DDLTask.execute(DDLTask.java:281)\n at org.apache.hadoop.hive.ql.exec.Task.executeTask(Task.java:153)\n at org.apache.hadoop.hive.ql.exec.TaskRunner.runSequential(TaskRunner.java:85)\n at org.apache.hadoop.hive.ql.Driver.launchTask(Driver.java:1503)\n at org.apache.hadoop.hive.ql.Driver.execute(Driver.java:1270)\n at org.apache.hadoop.hive.ql.Driver.runInternal(Driver.java:1088)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:911)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:901)\n at org.apache.spark.sql.hive.HiveContext.runHive(HiveContext.scala:292)\n at org.apache.spark.sql.hive.HiveContext.runSqlHive(HiveContext.scala:264)\n at org.apache.spark.sql.hive.execution.HiveNativeCommand.run(HiveNativeCommand.scala:37)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult$lzycompute(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.execute(commands.scala:61)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd$lzycompute(SQLContext.scala:474)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd(SQLContext.scala:474)\n at org.apache.spark.sql.SchemaRDDLike$class.$init$(SchemaRDDLike.scala:58)\n at org.apache.spark.sql.SchemaRDD.(SchemaRDD.scala:107)\n at org.apache.spark.sql.hive.HiveContext.sql(HiveContext.scala:73)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:160)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n at org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n at org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n at org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n at com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n at org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n at org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n at org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n15/01/23 23:49:01 WARN security.UserGroupInformation: No groups available for user anonymous\n15/01/23 23:49:01 ERROR exec.DDLTask: org.apache.hadoop.hive.ql.metadata.HiveException: Cannot validate serde: org.openx.data.jsonserde.JsonSerDe\n at org.apache.hadoop.hive.ql.exec.DDLTask.validateSerDe(DDLTask.java:3952)\n at org.apache.hadoop.hive.ql.exec.DDLTask.createTable(DDLTask.java:4084)\n at org.apache.hadoop.hive.ql.exec.DDLTask.execute(DDLTask.java:281)\n at org.apache.hadoop.hive.ql.exec.Task.executeTask(Task.java:153)\n at org.apache.hadoop.hive.ql.exec.TaskRunner.runSequential(TaskRunner.java:85)\n at org.apache.hadoop.hive.ql.Driver.launchTask(Driver.java:1503)\n at org.apache.hadoop.hive.ql.Driver.execute(Driver.java:1270)\n at org.apache.hadoop.hive.ql.Driver.runInternal(Driver.java:1088)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:911)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:901)\n at org.apache.spark.sql.hive.HiveContext.runHive(HiveContext.scala:292)\n at org.apache.spark.sql.hive.HiveContext.runSqlHive(HiveContext.scala:264)\n at org.apache.spark.sql.hive.execution.HiveNativeCommand.run(HiveNativeCommand.scala:37)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult$lzycompute(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.execute(commands.scala:61)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd$lzycompute(SQLContext.scala:474)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd(SQLContext.scala:474)\n at org.apache.spark.sql.SchemaRDDLike$class.$init$(SchemaRDDLike.scala:58)\n at org.apache.spark.sql.SchemaRDD.(SchemaRDD.scala:107)\n at org.apache.spark.sql.hive.HiveContext.sql(HiveContext.scala:73)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:160)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n at org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n at org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n at org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n at com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n at org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n at org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n at org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\nCaused by: java.lang.ClassNotFoundException: Class org.openx.data.jsonserde.JsonSerDe not found\n at org.apache.hadoop.conf.Configuration.getClassByName(Configuration.java:1801)\n at org.apache.hadoop.hive.ql.exec.DDLTask.validateSerDe(DDLTask.java:3946)\n ... 47 more\n\n15/01/23 23:49:01 ERROR ql.Driver: FAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Cannot validate serde: org.openx.data.jsonserde.JsonSerDe\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 ERROR hive.HiveContext:\n======================\nHIVE FAILURE OUTPUT\n======================\nSET spark.sql.codegen=false\nSET spark.sql.parquet.binaryAsString=true\nSET spark.sql.parquet.cacheMetadata=true\nSET spark.sql.hive.version=0.13.1\nSET spark.sql.autoBroadcastJoinThreshold=500000\nSET spark.sql.shuffle.partitions=10\nADD JAR /tmp/json-serde-1.3-jar-with-dependencies.jar\nAdded /tmp/json-serde-1.3-jar-with-dependencies.jar to class path\nAdded resource: /tmp/json-serde-1.3-jar-with-dependencies.jar\nFAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Cannot validate serde: org.openx.data.jsonserde.JsonSerDe\n\n======================\nEND HIVE FAILURE OUTPUT\n======================\n\n15/01/23 23:49:01 ERROR thriftserver.SparkExecuteStatementOperation: Error executing query:\norg.apache.spark.sql.execution.QueryExecutionException: FAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Cannot validate serde: org.openx.data.jsonserde.JsonSerDe\n at org.apache.spark.sql.hive.HiveContext.runHive(HiveContext.scala:296)\n at org.apache.spark.sql.hive.HiveContext.runSqlHive(HiveContext.scala:264)\n at org.apache.spark.sql.hive.execution.HiveNativeCommand.run(HiveNativeCommand.scala:37)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult$lzycompute(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.execute(commands.scala:61)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd$lzycompute(SQLContext.scala:474)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd(SQLContext.scala:474)\n at org.apache.spark.sql.SchemaRDDLike$class.$init$(SchemaRDDLike.scala:58)\n at org.apache.spark.sql.SchemaRDD.(SchemaRDD.scala:107)\n at org.apache.spark.sql.hive.HiveContext.sql(HiveContext.scala:73)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:160)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n at org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n at org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n at org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n at com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n at org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n at org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n at org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n15/01/23 23:49:01 WARN thrift.ThriftCLIService: Error executing statement:\norg.apache.hive.service.cli.HiveSQLException: org.apache.spark.sql.execution.QueryExecutionException: FAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Cannot validate serde: org.openx.data.jsonserde.JsonSerDe\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:189)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n at org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n at org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n at org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n at com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n at org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n at org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n at org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}","from":"reporter","subject":"SparkSQL fails to create tables with custom JSON SerDe"},{"body":"Same error occurred when a table is created with json serde in hive table and queried from SparkQL.","from":"developer"},{"body":"Is this still a problem? Is there a reason you aren't using the native JSON support?","from":"developer"},{"body":"Haven't tried native JSON but looks promising, so this ticket is probably lower priority.","from":"developer"},{"body":"I tried this, it worked in master, will close this as resolved.","from":"developer"}],"created":"2015-01-24T00:10:14.000+0000","description":"- Using Spark built from trunk on this commit: https://github.com/apache/spark/commit/bc20a52b34e826895d0dcc1d783c021ebd456ebd\n- Build for Hive13\n- Using this JSON serde: https://github.com/rcongiu/Hive-JSON-Serde\n\nFirst download jar locally:\n{code}\n$ curl http://www.congiu.net/hive-json-serde/1.3/cdh5/json-serde-1.3-jar-with-dependencies.jar > /tmp/json-serde-1.3-jar-with-dependencies.jar\n{code}\n\nThen add it in SparkSQL session:\n{code}\nadd jar /tmp/json-serde-1.3-jar-with-dependencies.jar\n{code}\n\nFinally create table:\n{code}\ncreate table test_json (c1 boolean) ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe';\n{code}\n\nLogs for add jar:\n{code}\n15/01/23 23:48:33 INFO thriftserver.SparkExecuteStatementOperation: Running query 'add jar /tmp/json-serde-1.3-jar-with-dependencies.jar'\n15/01/23 23:48:34 INFO session.SessionState: No Tez session required at this point. hive.execution.engine=mr.\n15/01/23 23:48:34 INFO SessionState: Added /tmp/json-serde-1.3-jar-with-dependencies.jar to class path\n15/01/23 23:48:34 INFO SessionState: Added resource: /tmp/json-serde-1.3-jar-with-dependencies.jar\n15/01/23 23:48:34 INFO spark.SparkContext: Added JAR /tmp/json-serde-1.3-jar-with-dependencies.jar at http://192.168.99.9:51312/jars/json-serde-1.3-jar-with-dependencies.jar with timestamp 1422056914776\n15/01/23 23:48:34 INFO thriftserver.SparkExecuteStatementOperation: Result Schema: List()\n15/01/23 23:48:34 INFO thriftserver.SparkExecuteStatementOperation: Result Schema: List()\n{code}\n\nLogs (with error) for create table:\n{code}\n15/01/23 23:49:00 INFO thriftserver.SparkExecuteStatementOperation: Running query 'create table test_json (c1 boolean) ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe''\n15/01/23 23:49:00 INFO parse.ParseDriver: Parsing command: create table test_json (c1 boolean) ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe'\n15/01/23 23:49:01 INFO parse.ParseDriver: Parse Completed\n15/01/23 23:49:01 INFO session.SessionState: No Tez session required at this point. hive.execution.engine=mr.\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO ql.Driver: Concurrency mode is disabled, not creating a lock manager\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO parse.ParseDriver: Parsing command: create table test_json (c1 boolean) ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe'\n15/01/23 23:49:01 INFO parse.ParseDriver: Parse Completed\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO parse.SemanticAnalyzer: Starting Semantic Analysis\n15/01/23 23:49:01 INFO parse.SemanticAnalyzer: Creating table test_json position=13\n15/01/23 23:49:01 INFO ql.Driver: Semantic Analysis Completed\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO ql.Driver: Returning Hive schema: Schema(fieldSchemas:null, properties:null)\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO ql.Driver: Starting command: create table test_json (c1 boolean) ROW FORMAT SERDE 'org.openx.data.jsonserde.JsonSerDe'\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 WARN security.ShellBasedUnixGroupsMapping: got exception trying to get groups for user anonymous\norg.apache.hadoop.util.Shell$ExitCodeException: id: anonymous: No such user\n\n at org.apache.hadoop.util.Shell.runCommand(Shell.java:505)\n at org.apache.hadoop.util.Shell.run(Shell.java:418)\n at org.apache.hadoop.util.Shell$ShellCommandExecutor.execute(Shell.java:650)\n at org.apache.hadoop.util.Shell.execCommand(Shell.java:739)\n at org.apache.hadoop.util.Shell.execCommand(Shell.java:722)\n at org.apache.hadoop.security.ShellBasedUnixGroupsMapping.getUnixGroups(ShellBasedUnixGroupsMapping.java:83)\n at org.apache.hadoop.security.ShellBasedUnixGroupsMapping.getGroups(ShellBasedUnixGroupsMapping.java:52)\n at org.apache.hadoop.security.JniBasedUnixGroupsMappingWithFallback.getGroups(JniBasedUnixGroupsMappingWithFallback.java:50)\n at org.apache.hadoop.security.Groups.getGroups(Groups.java:139)\n at org.apache.hadoop.security.UserGroupInformation.getGroupNames(UserGroupInformation.java:1409)\n at org.apache.hadoop.hive.ql.security.HadoopDefaultAuthenticator.setConf(HadoopDefaultAuthenticator.java:63)\n at org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:73)\n at org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:133)\n at org.apache.hadoop.hive.ql.metadata.HiveUtils.getAuthenticator(HiveUtils.java:424)\n at org.apache.hadoop.hive.ql.session.SessionState.setupAuth(SessionState.java:377)\n at org.apache.hadoop.hive.ql.session.SessionState.getAuthenticator(SessionState.java:867)\n at org.apache.hadoop.hive.ql.session.SessionState.getUserFromAuthenticator(SessionState.java:589)\n at org.apache.hadoop.hive.ql.metadata.Table.getEmptyTable(Table.java:174)\n at org.apache.hadoop.hive.ql.metadata.Table.(Table.java:116)\n at org.apache.hadoop.hive.ql.metadata.Hive.newTable(Hive.java:2566)\n at org.apache.hadoop.hive.ql.exec.DDLTask.createTable(DDLTask.java:4046)\n at org.apache.hadoop.hive.ql.exec.DDLTask.execute(DDLTask.java:281)\n at org.apache.hadoop.hive.ql.exec.Task.executeTask(Task.java:153)\n at org.apache.hadoop.hive.ql.exec.TaskRunner.runSequential(TaskRunner.java:85)\n at org.apache.hadoop.hive.ql.Driver.launchTask(Driver.java:1503)\n at org.apache.hadoop.hive.ql.Driver.execute(Driver.java:1270)\n at org.apache.hadoop.hive.ql.Driver.runInternal(Driver.java:1088)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:911)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:901)\n at org.apache.spark.sql.hive.HiveContext.runHive(HiveContext.scala:292)\n at org.apache.spark.sql.hive.HiveContext.runSqlHive(HiveContext.scala:264)\n at org.apache.spark.sql.hive.execution.HiveNativeCommand.run(HiveNativeCommand.scala:37)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult$lzycompute(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.execute(commands.scala:61)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd$lzycompute(SQLContext.scala:474)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd(SQLContext.scala:474)\n at org.apache.spark.sql.SchemaRDDLike$class.$init$(SchemaRDDLike.scala:58)\n at org.apache.spark.sql.SchemaRDD.(SchemaRDD.scala:107)\n at org.apache.spark.sql.hive.HiveContext.sql(HiveContext.scala:73)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:160)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n at org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n at org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n at org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n at com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n at org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n at org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n at org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n15/01/23 23:49:01 WARN security.UserGroupInformation: No groups available for user anonymous\n15/01/23 23:49:01 WARN security.ShellBasedUnixGroupsMapping: got exception trying to get groups for user anonymous\norg.apache.hadoop.util.Shell$ExitCodeException: id: anonymous: No such user\n\n at org.apache.hadoop.util.Shell.runCommand(Shell.java:505)\n at org.apache.hadoop.util.Shell.run(Shell.java:418)\n at org.apache.hadoop.util.Shell$ShellCommandExecutor.execute(Shell.java:650)\n at org.apache.hadoop.util.Shell.execCommand(Shell.java:739)\n at org.apache.hadoop.util.Shell.execCommand(Shell.java:722)\n at org.apache.hadoop.security.ShellBasedUnixGroupsMapping.getUnixGroups(ShellBasedUnixGroupsMapping.java:83)\n at org.apache.hadoop.security.ShellBasedUnixGroupsMapping.getGroups(ShellBasedUnixGroupsMapping.java:52)\n at org.apache.hadoop.security.JniBasedUnixGroupsMappingWithFallback.getGroups(JniBasedUnixGroupsMappingWithFallback.java:50)\n at org.apache.hadoop.security.Groups.getGroups(Groups.java:139)\n at org.apache.hadoop.security.UserGroupInformation.getGroupNames(UserGroupInformation.java:1409)\n at org.apache.hadoop.hive.ql.security.HadoopDefaultAuthenticator.setConf(HadoopDefaultAuthenticator.java:64)\n at org.apache.hadoop.util.ReflectionUtils.setConf(ReflectionUtils.java:73)\n at org.apache.hadoop.util.ReflectionUtils.newInstance(ReflectionUtils.java:133)\n at org.apache.hadoop.hive.ql.metadata.HiveUtils.getAuthenticator(HiveUtils.java:424)\n at org.apache.hadoop.hive.ql.session.SessionState.setupAuth(SessionState.java:377)\n at org.apache.hadoop.hive.ql.session.SessionState.getAuthenticator(SessionState.java:867)\n at org.apache.hadoop.hive.ql.session.SessionState.getUserFromAuthenticator(SessionState.java:589)\n at org.apache.hadoop.hive.ql.metadata.Table.getEmptyTable(Table.java:174)\n at org.apache.hadoop.hive.ql.metadata.Table.(Table.java:116)\n at org.apache.hadoop.hive.ql.metadata.Hive.newTable(Hive.java:2566)\n at org.apache.hadoop.hive.ql.exec.DDLTask.createTable(DDLTask.java:4046)\n at org.apache.hadoop.hive.ql.exec.DDLTask.execute(DDLTask.java:281)\n at org.apache.hadoop.hive.ql.exec.Task.executeTask(Task.java:153)\n at org.apache.hadoop.hive.ql.exec.TaskRunner.runSequential(TaskRunner.java:85)\n at org.apache.hadoop.hive.ql.Driver.launchTask(Driver.java:1503)\n at org.apache.hadoop.hive.ql.Driver.execute(Driver.java:1270)\n at org.apache.hadoop.hive.ql.Driver.runInternal(Driver.java:1088)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:911)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:901)\n at org.apache.spark.sql.hive.HiveContext.runHive(HiveContext.scala:292)\n at org.apache.spark.sql.hive.HiveContext.runSqlHive(HiveContext.scala:264)\n at org.apache.spark.sql.hive.execution.HiveNativeCommand.run(HiveNativeCommand.scala:37)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult$lzycompute(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.execute(commands.scala:61)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd$lzycompute(SQLContext.scala:474)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd(SQLContext.scala:474)\n at org.apache.spark.sql.SchemaRDDLike$class.$init$(SchemaRDDLike.scala:58)\n at org.apache.spark.sql.SchemaRDD.(SchemaRDD.scala:107)\n at org.apache.spark.sql.hive.HiveContext.sql(HiveContext.scala:73)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:160)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n at org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n at org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n at org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n at com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n at org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n at org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n at org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n15/01/23 23:49:01 WARN security.UserGroupInformation: No groups available for user anonymous\n15/01/23 23:49:01 ERROR exec.DDLTask: org.apache.hadoop.hive.ql.metadata.HiveException: Cannot validate serde: org.openx.data.jsonserde.JsonSerDe\n at org.apache.hadoop.hive.ql.exec.DDLTask.validateSerDe(DDLTask.java:3952)\n at org.apache.hadoop.hive.ql.exec.DDLTask.createTable(DDLTask.java:4084)\n at org.apache.hadoop.hive.ql.exec.DDLTask.execute(DDLTask.java:281)\n at org.apache.hadoop.hive.ql.exec.Task.executeTask(Task.java:153)\n at org.apache.hadoop.hive.ql.exec.TaskRunner.runSequential(TaskRunner.java:85)\n at org.apache.hadoop.hive.ql.Driver.launchTask(Driver.java:1503)\n at org.apache.hadoop.hive.ql.Driver.execute(Driver.java:1270)\n at org.apache.hadoop.hive.ql.Driver.runInternal(Driver.java:1088)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:911)\n at org.apache.hadoop.hive.ql.Driver.run(Driver.java:901)\n at org.apache.spark.sql.hive.HiveContext.runHive(HiveContext.scala:292)\n at org.apache.spark.sql.hive.HiveContext.runSqlHive(HiveContext.scala:264)\n at org.apache.spark.sql.hive.execution.HiveNativeCommand.run(HiveNativeCommand.scala:37)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult$lzycompute(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.execute(commands.scala:61)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd$lzycompute(SQLContext.scala:474)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd(SQLContext.scala:474)\n at org.apache.spark.sql.SchemaRDDLike$class.$init$(SchemaRDDLike.scala:58)\n at org.apache.spark.sql.SchemaRDD.(SchemaRDD.scala:107)\n at org.apache.spark.sql.hive.HiveContext.sql(HiveContext.scala:73)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:160)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n at org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n at org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n at org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n at com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n at org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n at org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n at org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\nCaused by: java.lang.ClassNotFoundException: Class org.openx.data.jsonserde.JsonSerDe not found\n at org.apache.hadoop.conf.Configuration.getClassByName(Configuration.java:1801)\n at org.apache.hadoop.hive.ql.exec.DDLTask.validateSerDe(DDLTask.java:3946)\n ... 47 more\n\n15/01/23 23:49:01 ERROR ql.Driver: FAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Cannot validate serde: org.openx.data.jsonserde.JsonSerDe\n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 INFO log.PerfLogger: \n15/01/23 23:49:01 ERROR hive.HiveContext:\n======================\nHIVE FAILURE OUTPUT\n======================\nSET spark.sql.codegen=false\nSET spark.sql.parquet.binaryAsString=true\nSET spark.sql.parquet.cacheMetadata=true\nSET spark.sql.hive.version=0.13.1\nSET spark.sql.autoBroadcastJoinThreshold=500000\nSET spark.sql.shuffle.partitions=10\nADD JAR /tmp/json-serde-1.3-jar-with-dependencies.jar\nAdded /tmp/json-serde-1.3-jar-with-dependencies.jar to class path\nAdded resource: /tmp/json-serde-1.3-jar-with-dependencies.jar\nFAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Cannot validate serde: org.openx.data.jsonserde.JsonSerDe\n\n======================\nEND HIVE FAILURE OUTPUT\n======================\n\n15/01/23 23:49:01 ERROR thriftserver.SparkExecuteStatementOperation: Error executing query:\norg.apache.spark.sql.execution.QueryExecutionException: FAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Cannot validate serde: org.openx.data.jsonserde.JsonSerDe\n at org.apache.spark.sql.hive.HiveContext.runHive(HiveContext.scala:296)\n at org.apache.spark.sql.hive.HiveContext.runSqlHive(HiveContext.scala:264)\n at org.apache.spark.sql.hive.execution.HiveNativeCommand.run(HiveNativeCommand.scala:37)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult$lzycompute(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.sideEffectResult(commands.scala:53)\n at org.apache.spark.sql.execution.ExecutedCommand.execute(commands.scala:61)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd$lzycompute(SQLContext.scala:474)\n at org.apache.spark.sql.SQLContext$QueryExecution.toRdd(SQLContext.scala:474)\n at org.apache.spark.sql.SchemaRDDLike$class.$init$(SchemaRDDLike.scala:58)\n at org.apache.spark.sql.SchemaRDD.(SchemaRDD.scala:107)\n at org.apache.spark.sql.hive.HiveContext.sql(HiveContext.scala:73)\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:160)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n at org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n at org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n at org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n at com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n at org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n at org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n at org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n15/01/23 23:49:01 WARN thrift.ThriftCLIService: Error executing statement:\norg.apache.hive.service.cli.HiveSQLException: org.apache.spark.sql.execution.QueryExecutionException: FAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Cannot validate serde: org.openx.data.jsonserde.JsonSerDe\n at org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation.run(Shim13.scala:189)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatementInternal(HiveSessionImpl.java:231)\n at org.apache.hive.service.cli.session.HiveSessionImpl.executeStatement(HiveSessionImpl.java:212)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:483)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:79)\n at org.apache.hive.service.cli.session.HiveSessionProxy.access$000(HiveSessionProxy.java:37)\n at org.apache.hive.service.cli.session.HiveSessionProxy$1.run(HiveSessionProxy.java:64)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1548)\n at org.apache.hadoop.hive.shims.HadoopShimsSecure.doAs(HadoopShimsSecure.java:493)\n at org.apache.hive.service.cli.session.HiveSessionProxy.invoke(HiveSessionProxy.java:60)\n at com.sun.proxy.$Proxy18.executeStatement(Unknown Source)\n at org.apache.hive.service.cli.CLIService.executeStatement(CLIService.java:220)\n at org.apache.hive.service.cli.thrift.ThriftCLIService.ExecuteStatement(ThriftCLIService.java:344)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1313)\n at org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement.getResult(TCLIService.java:1298)\n at org.apache.thrift.ProcessFunction.process(ProcessFunction.java:39)\n at org.apache.thrift.TBaseProcessor.process(TBaseProcessor.java:39)\n at org.apache.hive.service.auth.TSetIpAddressProcessor.process(TSetIpAddressProcessor.java:55)\n at org.apache.thrift.server.TThreadPoolServer$WorkerProcess.run(TThreadPoolServer.java:206)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\n{code}","issue_id":"12769777","key":"SPARK-5391","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-10-13T18:33:56.000+0000","role":"fixed_distractor","summary":"SparkSQL fails to create tables with custom JSON SerDe"} {"case_id":"12769827","cluster":"DISTRACTOR-SPARK-5395","comments":[{"body":"Having the same issue in standalone deployment mode. A single spark-submitted job is spawning a ton of pyspark.daemon instances and depleting the cluster memory even though the appropriate environment variables have been set.","created":"2015-01-26T17:21:27.729+0000"},{"body":"[~mkman84], do you also see this for both spark.python.worker.reuse false/true? (FWIW, the run I pasted above had reuse disabled.)\n\nAlso, do you happen to have a small job that can be used as a repro ([~davies] was asking for one, so far I only managed to trigger this condition using production data).","created":"2015-01-26T23:36:21.224+0000"},{"body":"[~skrasser], I actually only managed to have this reproduced using production data as well (so far). I'll try to write a simple version tomorrow but it seems that it's a mix of both python worker processes not being killed after it's no longer running (causing build up), as well as the python worker exceeding the allocated memory limit. \n\nI think it *may* be related to a couple of specific actions such as groupByKey/cogroup, though I'll still need to do some tests to be sure what's causing this.\n\nI should also add that we haven't modified the default for the python.worker.reuse variable, so in our case it should be using the default of True.","created":"2015-01-26T23:56:26.822+0000"},{"body":"This may prove to be useful...\n\nI'm watching a presently running spark-submitted job, while watching the pyspark.daemon processes. The framework is permitted to only use 8 cores on each node with the default python worker memory of 512mb per node (not the executor memory which is set to higher than this).\n\nIgnoring the exact RDD actions for a moment, it looks like while it transitions from Stage 1 -> Stage 2, it spawned up 8-10 additional pyspark.daemon processes making the box use more cores than it was even allowed to... A few seconds after that, the other 8 processes entered a sleeping state while still holding onto the physical memory it ate up in Stage 1. As soon as Stage 2 finished, practically all of the pyspark.daemons vanished and freed up the memory usage. I was keeping an eye on 2 random nodes and the exact same thing occurred on both. It was also the only currently executing job at the time so there was really no other interference/contention for resources.\n\nI will try to provide a bit more detail on the exact transformations/actions occurring between the 2 stages, although I know a PartionBy and cogroup are occurring at the very least without inspecting the spark-submitted code directly.","created":"2015-01-27T01:04:21.810+0000"},{"body":"Some additional findings from my side: I've managed to trigger the problem using a simpler job on production data that basically does a reduceByKey followed by a count action. I get >20 workers (2 cores per executor) before any tasks in the first stage (reduceByKey) complete (i.e. different from the stage transition behavior you noticed). However, this doesn't occur if I run over a smaller data set, i.e. fewer production data files.\n\nBefore calling reduceByKey I have a coalesce call. Without that the error does not occur (at least in this smaller script). This at first glance looked potentially spilling related (more data per task), but attempting to force spills by setting the worker memory very low did not help with my attempts to get a repro on test data.","created":"2015-01-27T03:14:33.941+0000"},{"body":"Actually I think I know why this happens... I'm thinking the problem really occurs due to the way auto-persistence of specific actions occur. \n\nReduceByKey, GroupByKey, cogroup, etc, are typically heavy actions that get auto-persisted for the reason that the resulting RDD's will most likely be used for something right after. \n\nThe interesting thing is that this memory is outside of the executor memory for the framework (it's what goes into these pyspark daemons that get spawned up temporarily). The other interesting fact is that let's say we leave the default python worker memory set to 512MB, and you have a framework that uses 8 cores on each executor, it spawns up 8 * 512MB (4GB) of python workers while the stage is running. \n\n[~skrasser] In your case, if you chain a bunch of auto-persisting actions (which I believe coalesce is a part of, since instead of dealing with a shuffle read, it instead builds a potentially large array of partitions on the executor), it will spawn an additional 2 python workers per executor for that separate task, while the previous tasks' python workers are left in a sleeping state, waiting for the results of the subsequent task to complete... \n\nIf that's the case, then it should actually be a bit easier showing how a single framework can nuke a single host by creating a crazy chain of coalescing/reduceByKey/GroupByKey/cogrouping actions (which I'm off to try out now haha)\n\nEDIT: I'm almost positive this is what's causing this to occur now. Unfortunately there is no easy way to prevent a single framework from wiping out all of the memory on a single box if it does a huge amount of shuffle writing, with the combination of auto-persisting and chained RDD actions which depend on previous RDD computations... You COULD break up the chain by forcing an intermediate step to DISK using (saveAsPickleFile/saveAsTextFile perhaps), and then in the next step reading it back in. At least that would force the previous python worker daemons to be cleaned up before potentially spawning new ones...\n\nIdeally there should be an environment variable for the max number of python workers allowed to be spawned per executor, because it looks like that doesn't exist as of yet! ","created":"2015-01-27T15:05:23.666+0000"},{"body":"I've definitely seen this behavior when adding additional reduceByKey operations that could execute in parallel (around 2 workers per such operation), but in my small test script it's a single reduceByKey operation followed by a single coalesce statement. I'm running it right now, and it already spiked up to over 30 Python workers per executor, so there must be something else going on on top of this.","created":"2015-01-27T17:31:45.011+0000"},{"body":"Some new findings: I can trigger the problem now just using the {{coalesce}} call. My job now looks like this:\n{code}sc.newAPIHadoopFile().map().map().coalesce().count(){code}\n\nIn the 64 executor case, this occurs when processing 1TB in 1500 files. If I go down to 2 executors, 200GB in 305 files make the worker count go up to 9 (higher as I add more files). With less data, things appear normal.\n\nThat raises the question about what {{coalesce()}} is doing that causes new workers to spawn.","created":"2015-01-28T01:43:49.838+0000"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4238","created":"2015-01-28T07:26:01.979+0000"},{"body":"Thanks Davies!","created":"2015-01-28T17:15:18.408+0000"},{"body":"Thanks! Hopefully that tackles the same problem I was seeing!","created":"2015-01-28T17:20:08.579+0000"},{"body":"I've committed Davies' patch (https://github.com/apache/spark/pull/4238) to {{master}} for inclusion in Spark 1.3.0 and tagged it for later backport to Spark 1.2.2. (I'll cherry-pick the commit after we close the 1.2.1 vote).","created":"2015-01-30T01:30:22.643+0000"},{"body":"I've merged this into `branch-1.2` (1.2.2), so I'm marking this as fixed.","created":"2015-02-17T04:35:59.220+0000"}],"conversations":[{"body":"During job execution a large number of Python worker accumulates eventually causing YARN to kill containers for being over their memory allocation (in the case below that is about 8G for executors plus 6G for overhead per container). \n\nIn this instance, at the time of killing the container 97 pyspark.daemon processes had accumulated.\n\n{noformat}\n2015-01-23 15:36:53,654 INFO [Reporter] yarn.YarnAllocationHandler (Logging.scala:logInfo(59)) - Container marked as failed: container_1421692415636_0052_01_000030. Exit status: 143. Diagnostics: Container [pid=35211,containerID=container_1421692415636_0052_01_000030] is running beyond physical memory limits. Current usage: 14.9 GB of 14.5 GB physical memory used; 41.3 GB of 72.5 GB virtual memory used. Killing container.\nDump of the process-tree for container_1421692415636_0052_01_000030 :\n|- PID PPID PGRPID SESSID CMD_NAME USER_MODE_TIME(MILLIS) SYSTEM_TIME(MILLIS) VMEM_USAGE(BYTES) RSSMEM_USAGE(PAGES) FULL_CMD_LINE\n|- 54101 36625 36625 35211 (python) 78 1 332730368 16834 python -m pyspark.daemon\n|- 52140 36625 36625 35211 (python) 58 1 332730368 16837 python -m pyspark.daemon\n|- 36625 35228 36625 35211 (python) 65 604 331685888 17694 python -m pyspark.daemon\n\t[...]\n{noformat}\n\nThe configuration used uses 64 containers with 2 cores each.\n\nFull output here: https://gist.github.com/skrasser/e3e2ee8dede5ef6b082c\n\nMailinglist discussion: https://www.mail-archive.com/user@spark.apache.org/msg20102.html","from":"reporter","subject":"Large number of Python workers causing resource depletion"},{"body":"Having the same issue in standalone deployment mode. A single spark-submitted job is spawning a ton of pyspark.daemon instances and depleting the cluster memory even though the appropriate environment variables have been set.","from":"developer"},{"body":"[~mkman84], do you also see this for both spark.python.worker.reuse false/true? (FWIW, the run I pasted above had reuse disabled.)\n\nAlso, do you happen to have a small job that can be used as a repro ([~davies] was asking for one, so far I only managed to trigger this condition using production data).","from":"developer"},{"body":"[~skrasser], I actually only managed to have this reproduced using production data as well (so far). I'll try to write a simple version tomorrow but it seems that it's a mix of both python worker processes not being killed after it's no longer running (causing build up), as well as the python worker exceeding the allocated memory limit. \n\nI think it *may* be related to a couple of specific actions such as groupByKey/cogroup, though I'll still need to do some tests to be sure what's causing this.\n\nI should also add that we haven't modified the default for the python.worker.reuse variable, so in our case it should be using the default of True.","from":"developer"},{"body":"This may prove to be useful...\n\nI'm watching a presently running spark-submitted job, while watching the pyspark.daemon processes. The framework is permitted to only use 8 cores on each node with the default python worker memory of 512mb per node (not the executor memory which is set to higher than this).\n\nIgnoring the exact RDD actions for a moment, it looks like while it transitions from Stage 1 -> Stage 2, it spawned up 8-10 additional pyspark.daemon processes making the box use more cores than it was even allowed to... A few seconds after that, the other 8 processes entered a sleeping state while still holding onto the physical memory it ate up in Stage 1. As soon as Stage 2 finished, practically all of the pyspark.daemons vanished and freed up the memory usage. I was keeping an eye on 2 random nodes and the exact same thing occurred on both. It was also the only currently executing job at the time so there was really no other interference/contention for resources.\n\nI will try to provide a bit more detail on the exact transformations/actions occurring between the 2 stages, although I know a PartionBy and cogroup are occurring at the very least without inspecting the spark-submitted code directly.","from":"developer"},{"body":"Some additional findings from my side: I've managed to trigger the problem using a simpler job on production data that basically does a reduceByKey followed by a count action. I get >20 workers (2 cores per executor) before any tasks in the first stage (reduceByKey) complete (i.e. different from the stage transition behavior you noticed). However, this doesn't occur if I run over a smaller data set, i.e. fewer production data files.\n\nBefore calling reduceByKey I have a coalesce call. Without that the error does not occur (at least in this smaller script). This at first glance looked potentially spilling related (more data per task), but attempting to force spills by setting the worker memory very low did not help with my attempts to get a repro on test data.","from":"developer"},{"body":"Actually I think I know why this happens... I'm thinking the problem really occurs due to the way auto-persistence of specific actions occur. \n\nReduceByKey, GroupByKey, cogroup, etc, are typically heavy actions that get auto-persisted for the reason that the resulting RDD's will most likely be used for something right after. \n\nThe interesting thing is that this memory is outside of the executor memory for the framework (it's what goes into these pyspark daemons that get spawned up temporarily). The other interesting fact is that let's say we leave the default python worker memory set to 512MB, and you have a framework that uses 8 cores on each executor, it spawns up 8 * 512MB (4GB) of python workers while the stage is running. \n\n[~skrasser] In your case, if you chain a bunch of auto-persisting actions (which I believe coalesce is a part of, since instead of dealing with a shuffle read, it instead builds a potentially large array of partitions on the executor), it will spawn an additional 2 python workers per executor for that separate task, while the previous tasks' python workers are left in a sleeping state, waiting for the results of the subsequent task to complete... \n\nIf that's the case, then it should actually be a bit easier showing how a single framework can nuke a single host by creating a crazy chain of coalescing/reduceByKey/GroupByKey/cogrouping actions (which I'm off to try out now haha)\n\nEDIT: I'm almost positive this is what's causing this to occur now. Unfortunately there is no easy way to prevent a single framework from wiping out all of the memory on a single box if it does a huge amount of shuffle writing, with the combination of auto-persisting and chained RDD actions which depend on previous RDD computations... You COULD break up the chain by forcing an intermediate step to DISK using (saveAsPickleFile/saveAsTextFile perhaps), and then in the next step reading it back in. At least that would force the previous python worker daemons to be cleaned up before potentially spawning new ones...\n\nIdeally there should be an environment variable for the max number of python workers allowed to be spawned per executor, because it looks like that doesn't exist as of yet! ","from":"developer"},{"body":"I've definitely seen this behavior when adding additional reduceByKey operations that could execute in parallel (around 2 workers per such operation), but in my small test script it's a single reduceByKey operation followed by a single coalesce statement. I'm running it right now, and it already spiked up to over 30 Python workers per executor, so there must be something else going on on top of this.","from":"developer"},{"body":"Some new findings: I can trigger the problem now just using the {{coalesce}} call. My job now looks like this:\n{code}sc.newAPIHadoopFile().map().map().coalesce().count(){code}\n\nIn the 64 executor case, this occurs when processing 1TB in 1500 files. If I go down to 2 executors, 200GB in 305 files make the worker count go up to 9 (higher as I add more files). With less data, things appear normal.\n\nThat raises the question about what {{coalesce()}} is doing that causes new workers to spawn.","from":"developer"},{"body":"User 'davies' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4238","from":"developer"},{"body":"Thanks Davies!","from":"developer"},{"body":"Thanks! Hopefully that tackles the same problem I was seeing!","from":"developer"},{"body":"I've committed Davies' patch (https://github.com/apache/spark/pull/4238) to {{master}} for inclusion in Spark 1.3.0 and tagged it for later backport to Spark 1.2.2. (I'll cherry-pick the commit after we close the 1.2.1 vote).","from":"developer"},{"body":"I've merged this into `branch-1.2` (1.2.2), so I'm marking this as fixed.","from":"developer"}],"created":"2015-01-24T08:19:40.000+0000","description":"During job execution a large number of Python worker accumulates eventually causing YARN to kill containers for being over their memory allocation (in the case below that is about 8G for executors plus 6G for overhead per container). \n\nIn this instance, at the time of killing the container 97 pyspark.daemon processes had accumulated.\n\n{noformat}\n2015-01-23 15:36:53,654 INFO [Reporter] yarn.YarnAllocationHandler (Logging.scala:logInfo(59)) - Container marked as failed: container_1421692415636_0052_01_000030. Exit status: 143. Diagnostics: Container [pid=35211,containerID=container_1421692415636_0052_01_000030] is running beyond physical memory limits. Current usage: 14.9 GB of 14.5 GB physical memory used; 41.3 GB of 72.5 GB virtual memory used. Killing container.\nDump of the process-tree for container_1421692415636_0052_01_000030 :\n|- PID PPID PGRPID SESSID CMD_NAME USER_MODE_TIME(MILLIS) SYSTEM_TIME(MILLIS) VMEM_USAGE(BYTES) RSSMEM_USAGE(PAGES) FULL_CMD_LINE\n|- 54101 36625 36625 35211 (python) 78 1 332730368 16834 python -m pyspark.daemon\n|- 52140 36625 36625 35211 (python) 58 1 332730368 16837 python -m pyspark.daemon\n|- 36625 35228 36625 35211 (python) 65 604 331685888 17694 python -m pyspark.daemon\n\t[...]\n{noformat}\n\nThe configuration used uses 64 containers with 2 cores each.\n\nFull output here: https://gist.github.com/skrasser/e3e2ee8dede5ef6b082c\n\nMailinglist discussion: https://www.mail-archive.com/user@spark.apache.org/msg20102.html","issue_id":"12769827","key":"SPARK-5395","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-02-17T04:35:59.000+0000","role":"fixed_distractor","summary":"Large number of Python workers causing resource depletion"} {"case_id":"12776752","cluster":"DISTRACTOR-SPARK-5945","comments":[{"body":"I encounter stage retry infinitely when a executor lost because a Spark bug which I already fix in SPARK-5259\n\nFor solve stage retry, I add a retry limit.\n\n{code}\ncase FetchFailed{\n ....\n if (disallowStageRetryForTest) {\n abortStage(failedStage, \"Fetch failure will not retry stage due to testing config\")\n } else if (failedStage.attemptId >= maxStageFailures) {\n abortStage(failedStage, s\"Fetch failure will not retry stage\" +\n \" due to reach to max Failure times: \" + maxStageFailures)\n } else if (failedStages.isEmpty && eventProcessActor != null) {\n // Don't schedule an event to resubmit failed stages if failed isn't empty, because\n .....\n}\n{code}\n\n","created":"2015-02-28T03:25:59.858+0000"},{"body":"Hi Imran - I'd be happy to tackle this. Could you please assign it to me? Thank you. ","created":"2015-03-18T20:59:43.838+0000"},{"body":"Hi [~ilganeli],\n\nsorry for taking a while to respond. I think the main issue here is not so much just implementing the code (as [~SuYan] already has shown the small required patch). The big issue is figuring out what the desired semantics are (see the questions I listed above), which means just getting feedback from all the required people on this one. But if you want to drive that process, that sounds great, it would really be appreciated!","created":"2015-03-23T19:28:32.415+0000"},{"body":"User 'ilganeli' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/5636","created":"2015-04-22T18:15:08.264+0000"},{"body":"Commenting here rather than on the github for archiving purposes!\n\nI took at look at the proposed pull request, and I'd be in favor of a much simpler approach, where for each stage, we track the number of failures (from any cause), and then fail the job once a stage fails 4 times (4, for consistency with the max task failures). If the stage succeeds, we can reset the count to 0, to avoid the potential problem [~imranr] mentioned for stages that are re-used by many jobs (so the counter would be numConsecutiveFailures or something like that). This can just be added to the Stage class, I think. This is consistent with the approach we use for tasks, where if a task has failed 4 times (for any reason), we abort the stage.\n\nI also would advocate against adding a configuration parameter for this. I can't imagine a case where someone would want to keep trying after 4 failures, and I think sometimes configuration parameters for things like this lead people to believe they can fix a problem by changing the configuration variable (just up the max number of failures!!) when really there is some bigger underlying issue they should fix. 4 seems to have worked well for tasks, so I'd just use the same default here (and it's always easy to add a configuration variable later on if lots of people say they need it).","created":"2015-04-22T21:35:43.379+0000"},{"body":"[~kayousterhout] - thanks for the review. If I understand correctly, your suggestion would still address [~imranr]'s second comment since the first stage would always (or mostly succeed), e.g. it wouldn't have N consecutive failures so even if subsequent stages fail, those wouldn't count towards the failure count for this particular stage since it would have been reset when it succeeded. \n\nDo you have any thoughts on the first comment? Specifically, is retrying a stage likely to succeed at all or is it a waste of effort in the first place?\n","created":"2015-04-22T22:08:42.577+0000"},{"body":"When there's a fetch failed exception, what happens is that we mark the corresponding map tasks as failed, and then re-run all of the failed tasks from the previous stage. When that's done, we re-run the stage with the original fetch failed exception -- so in the normal case, the stage with the fetch failed exception should succeed the second time.","created":"2015-04-22T22:11:14.394+0000"},{"body":"good point about moving the design discussion to jira, thanks Kay.\n\nYes, totally agree that we want to retry at least once, its definitely \"normal\" that in a big cluster, a node will go bad from time-to-time, but the good thing is Spark knows how to recover. I also agree that we shouldn't care about the cause of the failure, just the failure count.\n\nI totally see your points about the different configuration parameters, but let me just play devil's advocate. yes, configuration parameters are confusing to users, but thats a reason we should have sensible defaults and most users should never need to touch them. That doesn't mean nobody will want them. Tasks and stages are in some ways very different things -- tasks are meant to be very small and lightweight, so failing a few extra times is no big deal. But stages can be really big -- I would imagine in most cases, you actually might want to fail completely if the stage fails even twice, just because you can waste so much time in stage failure. Then again, there might be the other extreme, with really big clusters and unstable hardware, maybe two failures won't be that big a deal, so some users will want it higher.\n\nI disagree that its easy to add the config later. Yes, its easy to make the code change. But its hard to deploy the change in a production environment. And I can see this as a parameter that devops team needs to play with for their exact system / workload / SLAs etc. -- not sure at all, but I think we just don't know, and so we should leave the door open.\n\nI also can't see any reason why anyone would want infinite retries -- but I'm hesitant (and asked for the change) just b/c of changing from the old behavior. I guess int.maxvalue is close enough if somebody needs it?","created":"2015-04-22T22:35:34.036+0000"},{"body":"I realized there might be a cleaner solution here: I wonder if we should just break the FetchFailedException into two subtypes, one of which is UnrecoverableFetchFailedException. If a shuffle block is too large, there's no point in retrying the stage; we should just fail it. I realized maybe this is what you were alluding to in your earlier comment, when you asked whether it's worth it to retry. It seems like, at the point when we throw the exception, we do know whether it's worth it to retry, and it would be cleaner to just throw an appropriate exception. Thoughts?","created":"2015-04-22T22:55:15.367+0000"},{"body":"I think we want to do something more than just handling the case of blocks that are too large. Putting in a special case to avoid retrying at all in that case would be fine, but to me that is a separate issue, the more important thing to do is to put in a general retry limit. There could be all sorts of other reasons for fetch failures (I just looked into a case where an OOM would kill an executor, which lead to 10 hours of stage retry attempts before somebody manually killed the job).","created":"2015-04-23T04:49:03.748+0000"},{"body":"Good point -- that makes sense.\n\nI'm still in favor of not adding a config parameter: I think we're in agreement that it's hard to imagine a case where someone wants this to be *more* than 4, so making it 4 seems strictly better than the current approach where it is infinite. If people complain or want to configure it to be less, we can always change this in a minor release.","created":"2015-04-23T05:09:19.236+0000"},{"body":"So to recap:\na) Move failure count tracking into Stage\nb) Reset failure count on Stage success, so even if that stage is re-submitted due to failures downstream, we never hit the cap\nc) Remove config parameter. ","created":"2015-04-24T00:40:35.391+0000"},{"body":"[~kayousterhout] can you please clarify -- did you want to just hardcode to 4, or did you want to reuse {{spark.task.maxFailures}} for stage failures as well?","created":"2015-04-29T06:21:47.510+0000"},{"body":"I wanted to hardcode to 4 (totally agree with the sentiment you expressed earlier in this thread, that it doesn't make sense / is very confusing to re-use a config parameter for two different things).","created":"2015-04-29T16:20:14.047+0000"},{"body":"blocked by SPARK-7308 b/c you need to know which attempt is being failed","created":"2015-05-26T22:32:44.487+0000"},{"body":"At the moment we have a ton of these infinite retries. A stage is retried a few dozen times, then its parent goes missing and Spark starts retrying the parent until it also goes missing... We are still debugging the cause of our fetch failures, but I just wanted to mention that if there were a {{spark.stage.maxFailures}} option, we would be setting it to 1 at this point.\n\nThanks for all the work on this bug. Even if it's not fixed yet, it's very informative.","created":"2015-07-02T13:38:59.873+0000"},{"body":"I have retargeted this and downgraded it from Blocker to Critical since it's been there for a while and not a regression.","created":"2015-08-19T18:56:40.724+0000"}],"conversations":[{"body":"While investigating SPARK-5928, I noticed some very strange behavior in the way spark retries stages after a FetchFailedException. It seems that on a FetchFailedException, instead of simply killing the task and retrying, Spark aborts the stage and retries. If it just retried the task, the task might fail 4 times and then trigger the usual job killing mechanism. But by killing the stage instead, the max retry logic is skipped (it looks to me like there is no limit for retries on a stage).\n\nAfter a bit of discussion with Kay Ousterhout, it seems the idea is that if a fetch fails, we assume that the block manager we are fetching from has failed, and that it will succeed if we retry the stage w/out that block manager. In that case, it wouldn't make any sense to retry the task, since its doomed to fail every time, so we might as well kill the whole stage. But this raises two questions:\n\n\n1) Is it really safe to assume that a FetchFailedException means that the BlockManager has failed, and ti will work if we just try another one? SPARK-5928 shows that there are at least some cases where that assumption is wrong. Even if we fix that case, this logic seems brittle to the next case we find. I guess the idea is that this behavior is what gives us the \"R\" in RDD ... but it seems like its not really that robust and maybe should be reconsidered.\n\n2) Should stages only be retried a limited number of times? It would be pretty easy to put in a limited number of retries per stage. Though again, we encounter issues with keeping things resilient. Theoretically one stage could have many retries, but due to failures in different stages further downstream, so we might need to track the cause of each retry as well to still have the desired behavior.\n\nIn general it just seems there is some flakiness in the retry logic. This is the only reproducible example I have at the moment, but I vaguely recall hitting other cases of strange behavior w/ retries when trying to run long pipelines. Eg., if one executor is stuck in a GC during a fetch, the fetch fails, but the executor eventually comes back and the stage gets retried again, but the same GC issues happen the second time around, etc.\n\nCopied from SPARK-5928, here's the example program that can regularly produce a loop of stage failures. Note that it will only fail from a remote fetch, so it can't be run locally -- I ran with {{MASTER=yarn-client spark-shell --num-executors 2 --executor-memory 4000m}}\n\n{code}\n val rdd = sc.parallelize(1 to 1e6.toInt, 1).map{ ignore =>\n val n = 3e3.toInt\n val arr = new Array[Byte](n)\n //need to make sure the array doesn't compress to something small\n scala.util.Random.nextBytes(arr)\n arr\n }\n rdd.map { x => (1, x)}.groupByKey().count()\n{code}","from":"reporter","subject":"Spark should not retry a stage infinitely on a FetchFailedException"},{"body":"I encounter stage retry infinitely when a executor lost because a Spark bug which I already fix in SPARK-5259\n\nFor solve stage retry, I add a retry limit.\n\n{code}\ncase FetchFailed{\n ....\n if (disallowStageRetryForTest) {\n abortStage(failedStage, \"Fetch failure will not retry stage due to testing config\")\n } else if (failedStage.attemptId >= maxStageFailures) {\n abortStage(failedStage, s\"Fetch failure will not retry stage\" +\n \" due to reach to max Failure times: \" + maxStageFailures)\n } else if (failedStages.isEmpty && eventProcessActor != null) {\n // Don't schedule an event to resubmit failed stages if failed isn't empty, because\n .....\n}\n{code}\n\n","from":"developer"},{"body":"Hi Imran - I'd be happy to tackle this. Could you please assign it to me? Thank you. ","from":"developer"},{"body":"Hi [~ilganeli],\n\nsorry for taking a while to respond. I think the main issue here is not so much just implementing the code (as [~SuYan] already has shown the small required patch). The big issue is figuring out what the desired semantics are (see the questions I listed above), which means just getting feedback from all the required people on this one. But if you want to drive that process, that sounds great, it would really be appreciated!","from":"developer"},{"body":"User 'ilganeli' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/5636","from":"developer"},{"body":"Commenting here rather than on the github for archiving purposes!\n\nI took at look at the proposed pull request, and I'd be in favor of a much simpler approach, where for each stage, we track the number of failures (from any cause), and then fail the job once a stage fails 4 times (4, for consistency with the max task failures). If the stage succeeds, we can reset the count to 0, to avoid the potential problem [~imranr] mentioned for stages that are re-used by many jobs (so the counter would be numConsecutiveFailures or something like that). This can just be added to the Stage class, I think. This is consistent with the approach we use for tasks, where if a task has failed 4 times (for any reason), we abort the stage.\n\nI also would advocate against adding a configuration parameter for this. I can't imagine a case where someone would want to keep trying after 4 failures, and I think sometimes configuration parameters for things like this lead people to believe they can fix a problem by changing the configuration variable (just up the max number of failures!!) when really there is some bigger underlying issue they should fix. 4 seems to have worked well for tasks, so I'd just use the same default here (and it's always easy to add a configuration variable later on if lots of people say they need it).","from":"developer"},{"body":"[~kayousterhout] - thanks for the review. If I understand correctly, your suggestion would still address [~imranr]'s second comment since the first stage would always (or mostly succeed), e.g. it wouldn't have N consecutive failures so even if subsequent stages fail, those wouldn't count towards the failure count for this particular stage since it would have been reset when it succeeded. \n\nDo you have any thoughts on the first comment? Specifically, is retrying a stage likely to succeed at all or is it a waste of effort in the first place?\n","from":"developer"},{"body":"When there's a fetch failed exception, what happens is that we mark the corresponding map tasks as failed, and then re-run all of the failed tasks from the previous stage. When that's done, we re-run the stage with the original fetch failed exception -- so in the normal case, the stage with the fetch failed exception should succeed the second time.","from":"developer"},{"body":"good point about moving the design discussion to jira, thanks Kay.\n\nYes, totally agree that we want to retry at least once, its definitely \"normal\" that in a big cluster, a node will go bad from time-to-time, but the good thing is Spark knows how to recover. I also agree that we shouldn't care about the cause of the failure, just the failure count.\n\nI totally see your points about the different configuration parameters, but let me just play devil's advocate. yes, configuration parameters are confusing to users, but thats a reason we should have sensible defaults and most users should never need to touch them. That doesn't mean nobody will want them. Tasks and stages are in some ways very different things -- tasks are meant to be very small and lightweight, so failing a few extra times is no big deal. But stages can be really big -- I would imagine in most cases, you actually might want to fail completely if the stage fails even twice, just because you can waste so much time in stage failure. Then again, there might be the other extreme, with really big clusters and unstable hardware, maybe two failures won't be that big a deal, so some users will want it higher.\n\nI disagree that its easy to add the config later. Yes, its easy to make the code change. But its hard to deploy the change in a production environment. And I can see this as a parameter that devops team needs to play with for their exact system / workload / SLAs etc. -- not sure at all, but I think we just don't know, and so we should leave the door open.\n\nI also can't see any reason why anyone would want infinite retries -- but I'm hesitant (and asked for the change) just b/c of changing from the old behavior. I guess int.maxvalue is close enough if somebody needs it?","from":"developer"},{"body":"I realized there might be a cleaner solution here: I wonder if we should just break the FetchFailedException into two subtypes, one of which is UnrecoverableFetchFailedException. If a shuffle block is too large, there's no point in retrying the stage; we should just fail it. I realized maybe this is what you were alluding to in your earlier comment, when you asked whether it's worth it to retry. It seems like, at the point when we throw the exception, we do know whether it's worth it to retry, and it would be cleaner to just throw an appropriate exception. Thoughts?","from":"developer"},{"body":"I think we want to do something more than just handling the case of blocks that are too large. Putting in a special case to avoid retrying at all in that case would be fine, but to me that is a separate issue, the more important thing to do is to put in a general retry limit. There could be all sorts of other reasons for fetch failures (I just looked into a case where an OOM would kill an executor, which lead to 10 hours of stage retry attempts before somebody manually killed the job).","from":"developer"},{"body":"Good point -- that makes sense.\n\nI'm still in favor of not adding a config parameter: I think we're in agreement that it's hard to imagine a case where someone wants this to be *more* than 4, so making it 4 seems strictly better than the current approach where it is infinite. If people complain or want to configure it to be less, we can always change this in a minor release.","from":"developer"},{"body":"So to recap:\na) Move failure count tracking into Stage\nb) Reset failure count on Stage success, so even if that stage is re-submitted due to failures downstream, we never hit the cap\nc) Remove config parameter. ","from":"developer"},{"body":"[~kayousterhout] can you please clarify -- did you want to just hardcode to 4, or did you want to reuse {{spark.task.maxFailures}} for stage failures as well?","from":"developer"},{"body":"I wanted to hardcode to 4 (totally agree with the sentiment you expressed earlier in this thread, that it doesn't make sense / is very confusing to re-use a config parameter for two different things).","from":"developer"},{"body":"blocked by SPARK-7308 b/c you need to know which attempt is being failed","from":"developer"},{"body":"At the moment we have a ton of these infinite retries. A stage is retried a few dozen times, then its parent goes missing and Spark starts retrying the parent until it also goes missing... We are still debugging the cause of our fetch failures, but I just wanted to mention that if there were a {{spark.stage.maxFailures}} option, we would be setting it to 1 at this point.\n\nThanks for all the work on this bug. Even if it's not fixed yet, it's very informative.","from":"developer"},{"body":"I have retargeted this and downgraded it from Blocker to Critical since it's been there for a while and not a regression.","from":"developer"}],"created":"2015-02-23T01:48:58.000+0000","description":"While investigating SPARK-5928, I noticed some very strange behavior in the way spark retries stages after a FetchFailedException. It seems that on a FetchFailedException, instead of simply killing the task and retrying, Spark aborts the stage and retries. If it just retried the task, the task might fail 4 times and then trigger the usual job killing mechanism. But by killing the stage instead, the max retry logic is skipped (it looks to me like there is no limit for retries on a stage).\n\nAfter a bit of discussion with Kay Ousterhout, it seems the idea is that if a fetch fails, we assume that the block manager we are fetching from has failed, and that it will succeed if we retry the stage w/out that block manager. In that case, it wouldn't make any sense to retry the task, since its doomed to fail every time, so we might as well kill the whole stage. But this raises two questions:\n\n\n1) Is it really safe to assume that a FetchFailedException means that the BlockManager has failed, and ti will work if we just try another one? SPARK-5928 shows that there are at least some cases where that assumption is wrong. Even if we fix that case, this logic seems brittle to the next case we find. I guess the idea is that this behavior is what gives us the \"R\" in RDD ... but it seems like its not really that robust and maybe should be reconsidered.\n\n2) Should stages only be retried a limited number of times? It would be pretty easy to put in a limited number of retries per stage. Though again, we encounter issues with keeping things resilient. Theoretically one stage could have many retries, but due to failures in different stages further downstream, so we might need to track the cause of each retry as well to still have the desired behavior.\n\nIn general it just seems there is some flakiness in the retry logic. This is the only reproducible example I have at the moment, but I vaguely recall hitting other cases of strange behavior w/ retries when trying to run long pipelines. Eg., if one executor is stuck in a GC during a fetch, the fetch fails, but the executor eventually comes back and the stage gets retried again, but the same GC issues happen the second time around, etc.\n\nCopied from SPARK-5928, here's the example program that can regularly produce a loop of stage failures. Note that it will only fail from a remote fetch, so it can't be run locally -- I ran with {{MASTER=yarn-client spark-shell --num-executors 2 --executor-memory 4000m}}\n\n{code}\n val rdd = sc.parallelize(1 to 1e6.toInt, 1).map{ ignore =>\n val n = 3e3.toInt\n val arr = new Array[Byte](n)\n //need to make sure the array doesn't compress to something small\n scala.util.Random.nextBytes(arr)\n arr\n }\n rdd.map { x => (1, x)}.groupByKey().count()\n{code}","issue_id":"12776752","key":"SPARK-5945","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-09-03T05:08:47.000+0000","role":"fixed_distractor","summary":"Spark should not retry a stage infinitely on a FetchFailedException"} {"case_id":"12779285","cluster":"DISTRACTOR-SPARK-6152","comments":[{"body":"Spark is not compiled with Java 8. What problem are you reporting?","created":"2015-03-04T09:43:52.707+0000"},{"body":"I've made the description more clear. The problem is occurs when I create a Scala 2.11 project that imports Spark 1.2.1 and I compile my Scala code to Java 8 using \"scalac -target:jvm-1.8 ...\"","created":"2015-03-04T17:38:24.368+0000"},{"body":"To your deleted comment -- yes indeed it looks like the library explicitly does not work with Java 8 since class file version 52 == Java 8","created":"2015-03-19T00:42:39.847+0000"},{"body":"I just released reflectasm-1.10.1 (which now should support java 8 due to the upgrade to asm 5) to maven central.","created":"2015-03-20T13:01:01.776+0000"},{"body":"Nice one. It looks like {{reflectasm}} comes in via {{chill}}. Do you know if/when {{chill}} might consume the newer version? then we could consume that.","created":"2015-03-20T13:04:52.154+0000"},{"body":"If you want to get the new version via {{chill}}, you will also need to wait for a release of {{kryo}} as {{chill}} gets the dependency via {{kryo}}\n\n-I think the root cause here is we are using a shaded library in a third party dependency. Sounds like a bad practise. We should just depend on an unshaded {{reflectasm}} directly-\n\n{{kyro}} just updated to {{reflectasm-1.10.1}} four hours ago, so this dependency train will take a while to arrive.\n\nEDIT: Actually, that was a silly idea. We do need to upgrade {{chill}} otherwise we will just move the failure into {{kyro}}","created":"2015-03-20T16:52:17.917+0000"},{"body":"I'll try to get out a new kryo version as soon as possible...","created":"2015-03-20T17:10:45.559+0000"},{"body":"Btw, chill guys are still on kryo 2.21, kryo right now is 3.0.0. IIRC there were some compatibility issues in kryo they complained about. Perhaps you already should open an issue for chill java 8 support to see what they think about it.","created":"2015-03-20T17:17:32.368+0000"},{"body":"Done: https://github.com/twitter/chill/issues/223","created":"2015-03-20T20:34:53.119+0000"},{"body":"Btw, we just released kryo 3.0.1: https://github.com/EsotericSoftware/kryo/blob/master/CHANGES.md#2240---300-2014-0-04","created":"2015-03-24T23:26:28.558+0000"},{"body":"It appears as if progress on updating chill to use a version of reflectasm that is compatible with Java 1.8 has stalled: https://github.com/twitter/chill/pull/224\n","created":"2015-06-06T03:02:48.985+0000"},{"body":"Chill and Kryo need to be in sync; there's also the need to be compatible with the version Hive uses, (which has historically been addressed with custom versions of Hive).\n\nIf spark could jump to Kryo 3.x, classpath conflict with hive would go away, provided the wire formats of serialized classes were compatible: hive's spark-client JAR uses kryo 2.2.x to talk to spark.","created":"2015-07-29T20:25:40.640+0000"},{"body":"Interesting [~stevel@apache.org]! What kinds of changes do you think this would require -- mostly verifying that there's backward compatibility with those serialized classes?","created":"2015-08-14T20:40:57.955+0000"},{"body":"what changes? Yes, making sure there are no regressions. Hive had to upgrade to deal with bugs Kryo 2.21; for Spark 1.5 there's a special org.spark-project.hive artifact which downgraded to Kryo 2.21; the subset of Hive that spark-hive and spark-hive-thriftserver all work there. Hive would certainly veto any reverting to 2.21; I don't know what their stance would be to a 3.x upgrade on the 1.2 branch ... reluctant would be the default response, I suspect.\n\nGetting everything to 2.24 is more likely, though if 3.x is needed for Java 8 compatibility it could be argued for","created":"2015-09-15T11:31:50.850+0000"},{"body":"Do we actually need reflectasm itself for the closure cleaner or are we just using it as a convenient way to pull in a shaded ASM artifact? Why not publish our own shaded ASM 5 and use that instead?","created":"2015-11-05T22:47:53.296+0000"},{"body":"It turns out that Apache Geronimo has already published shaded ASM 5 artifacts:\n\nhttp://mvnrepository.com/artifact/org.apache.xbean/xbean-asm5-shaded/4.4\n\nHere's the source that was used to produce that shaded artifact:\n\nhttps://github.com/apache/geronimo-xbean/tree/xbean-4.4/xbean-asm5-shaded\n\nThis corresponds to ASM 5.0.4, which is the latest release:\n\nhttps://github.com/apache/geronimo-xbean/blob/xbean-4.4/pom.xml#L67\n\nI'll investigate changing Spark to use this instead of reflectASM's shaded copy.","created":"2015-11-05T22:59:03.557+0000"},{"body":"User 'JoshRosen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9512","created":"2015-11-05T23:46:05.177+0000"},{"body":"Does anyone have a standalone reproduction of this issue that I can use to test my PR? https://github.com/apache/spark/pull/9512\n \nEDIT: just realized that this issue pertains to _Scala_ classes that were compiled with Java 8. Will add a new test to try that out.","created":"2015-11-10T19:42:07.712+0000"},{"body":"Yep, was able to reproduce trivially by running Spark's existing Scala unit tests with JDK 8. I'm going to add some plumbing to the build in order to let us test this in Jenkins.","created":"2015-11-10T19:57:07.942+0000"},{"body":"Issue resolved by pull request 9512\n[https://github.com/apache/spark/pull/9512]","created":"2015-11-11T19:18:23.462+0000"},{"body":"Nice - when is 1.6.0 due for release?","created":"2015-11-13T06:59:39.040+0000"}],"conversations":[{"body":"Spark uses reflectasm to check Scala closures which fails if the *user defined Scala closures* are compiled to Java 8 class version\n\nThe cause is reflectasm does not support Java 8\nhttps://github.com/EsotericSoftware/reflectasm/issues/35\n\nWorkaround:\nDon't compile Scala classes to Java 8, Scala 2.11 does not support nor require any Java 8 features\n\nStack trace:\n{code}\njava.lang.IllegalArgumentException\n\tat com.esotericsoftware.reflectasm.shaded.org.objectweb.asm.ClassReader.(Unknown Source)\n\tat com.esotericsoftware.reflectasm.shaded.org.objectweb.asm.ClassReader.(Unknown Source)\n\tat com.esotericsoftware.reflectasm.shaded.org.objectweb.asm.ClassReader.(Unknown Source)\n\tat org.apache.spark.util.ClosureCleaner$.org$apache$spark$util$ClosureCleaner$$getClassReader(ClosureCleaner.scala:41)\n\tat org.apache.spark.util.ClosureCleaner$.getInnerClasses(ClosureCleaner.scala:84)\n\tat org.apache.spark.util.ClosureCleaner$.clean(ClosureCleaner.scala:107)\n\tat org.apache.spark.SparkContext.clean(SparkContext.scala:1478)\n\tat org.apache.spark.rdd.RDD.map(RDD.scala:288)\n\tat ...my Scala 2.11 compiled to Java 8 code calling into spark\n{code}","from":"reporter","subject":"Spark does not support Java 8 compiled Scala classes"},{"body":"Spark is not compiled with Java 8. What problem are you reporting?","from":"developer"},{"body":"I've made the description more clear. The problem is occurs when I create a Scala 2.11 project that imports Spark 1.2.1 and I compile my Scala code to Java 8 using \"scalac -target:jvm-1.8 ...\"","from":"developer"},{"body":"To your deleted comment -- yes indeed it looks like the library explicitly does not work with Java 8 since class file version 52 == Java 8","from":"developer"},{"body":"I just released reflectasm-1.10.1 (which now should support java 8 due to the upgrade to asm 5) to maven central.","from":"developer"},{"body":"Nice one. It looks like {{reflectasm}} comes in via {{chill}}. Do you know if/when {{chill}} might consume the newer version? then we could consume that.","from":"developer"},{"body":"If you want to get the new version via {{chill}}, you will also need to wait for a release of {{kryo}} as {{chill}} gets the dependency via {{kryo}}\n\n-I think the root cause here is we are using a shaded library in a third party dependency. Sounds like a bad practise. We should just depend on an unshaded {{reflectasm}} directly-\n\n{{kyro}} just updated to {{reflectasm-1.10.1}} four hours ago, so this dependency train will take a while to arrive.\n\nEDIT: Actually, that was a silly idea. We do need to upgrade {{chill}} otherwise we will just move the failure into {{kyro}}","from":"developer"},{"body":"I'll try to get out a new kryo version as soon as possible...","from":"developer"},{"body":"Btw, chill guys are still on kryo 2.21, kryo right now is 3.0.0. IIRC there were some compatibility issues in kryo they complained about. Perhaps you already should open an issue for chill java 8 support to see what they think about it.","from":"developer"},{"body":"Done: https://github.com/twitter/chill/issues/223","from":"developer"},{"body":"Btw, we just released kryo 3.0.1: https://github.com/EsotericSoftware/kryo/blob/master/CHANGES.md#2240---300-2014-0-04","from":"developer"},{"body":"It appears as if progress on updating chill to use a version of reflectasm that is compatible with Java 1.8 has stalled: https://github.com/twitter/chill/pull/224\n","from":"developer"},{"body":"Chill and Kryo need to be in sync; there's also the need to be compatible with the version Hive uses, (which has historically been addressed with custom versions of Hive).\n\nIf spark could jump to Kryo 3.x, classpath conflict with hive would go away, provided the wire formats of serialized classes were compatible: hive's spark-client JAR uses kryo 2.2.x to talk to spark.","from":"developer"},{"body":"Interesting [~stevel@apache.org]! What kinds of changes do you think this would require -- mostly verifying that there's backward compatibility with those serialized classes?","from":"developer"},{"body":"what changes? Yes, making sure there are no regressions. Hive had to upgrade to deal with bugs Kryo 2.21; for Spark 1.5 there's a special org.spark-project.hive artifact which downgraded to Kryo 2.21; the subset of Hive that spark-hive and spark-hive-thriftserver all work there. Hive would certainly veto any reverting to 2.21; I don't know what their stance would be to a 3.x upgrade on the 1.2 branch ... reluctant would be the default response, I suspect.\n\nGetting everything to 2.24 is more likely, though if 3.x is needed for Java 8 compatibility it could be argued for","from":"developer"},{"body":"Do we actually need reflectasm itself for the closure cleaner or are we just using it as a convenient way to pull in a shaded ASM artifact? Why not publish our own shaded ASM 5 and use that instead?","from":"developer"},{"body":"It turns out that Apache Geronimo has already published shaded ASM 5 artifacts:\n\nhttp://mvnrepository.com/artifact/org.apache.xbean/xbean-asm5-shaded/4.4\n\nHere's the source that was used to produce that shaded artifact:\n\nhttps://github.com/apache/geronimo-xbean/tree/xbean-4.4/xbean-asm5-shaded\n\nThis corresponds to ASM 5.0.4, which is the latest release:\n\nhttps://github.com/apache/geronimo-xbean/blob/xbean-4.4/pom.xml#L67\n\nI'll investigate changing Spark to use this instead of reflectASM's shaded copy.","from":"developer"},{"body":"User 'JoshRosen' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9512","from":"developer"},{"body":"Does anyone have a standalone reproduction of this issue that I can use to test my PR? https://github.com/apache/spark/pull/9512\n \nEDIT: just realized that this issue pertains to _Scala_ classes that were compiled with Java 8. Will add a new test to try that out.","from":"developer"},{"body":"Yep, was able to reproduce trivially by running Spark's existing Scala unit tests with JDK 8. I'm going to add some plumbing to the build in order to let us test this in Jenkins.","from":"developer"},{"body":"Issue resolved by pull request 9512\n[https://github.com/apache/spark/pull/9512]","from":"developer"},{"body":"Nice - when is 1.6.0 due for release?","from":"developer"}],"created":"2015-03-04T03:05:34.000+0000","description":"Spark uses reflectasm to check Scala closures which fails if the *user defined Scala closures* are compiled to Java 8 class version\n\nThe cause is reflectasm does not support Java 8\nhttps://github.com/EsotericSoftware/reflectasm/issues/35\n\nWorkaround:\nDon't compile Scala classes to Java 8, Scala 2.11 does not support nor require any Java 8 features\n\nStack trace:\n{code}\njava.lang.IllegalArgumentException\n\tat com.esotericsoftware.reflectasm.shaded.org.objectweb.asm.ClassReader.(Unknown Source)\n\tat com.esotericsoftware.reflectasm.shaded.org.objectweb.asm.ClassReader.(Unknown Source)\n\tat com.esotericsoftware.reflectasm.shaded.org.objectweb.asm.ClassReader.(Unknown Source)\n\tat org.apache.spark.util.ClosureCleaner$.org$apache$spark$util$ClosureCleaner$$getClassReader(ClosureCleaner.scala:41)\n\tat org.apache.spark.util.ClosureCleaner$.getInnerClasses(ClosureCleaner.scala:84)\n\tat org.apache.spark.util.ClosureCleaner$.clean(ClosureCleaner.scala:107)\n\tat org.apache.spark.SparkContext.clean(SparkContext.scala:1478)\n\tat org.apache.spark.rdd.RDD.map(RDD.scala:288)\n\tat ...my Scala 2.11 compiled to Java 8 code calling into spark\n{code}","issue_id":"12779285","key":"SPARK-6152","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-11-11T19:18:23.000+0000","role":"fixed_distractor","summary":"Spark does not support Java 8 compiled Scala classes"} {"case_id":"12818951","cluster":"DISTRACTOR-SPARK-6743","comments":[{"body":"Any thoughts on this?","created":"2015-05-11T22:07:55.214+0000"},{"body":"Sorry, my first example was not very clear. Here is a more precise one:\n\n{code}\n val sqlc = new SQLContext(sc)\n\n val tab0 = sc.parallelize(Seq(\n Tuple1(\"A1\"),\n Tuple1(\"A2\")\n ))\n sqlc.registerDataFrameAsTable(sqlc.createDataFrame(tab0), \"tab0\")\n sqlc.cacheTable(\"tab0\")\n\n val tab1 = sc.parallelize(Seq(\n Tuple1(\"B1\"),\n Tuple1(\"B2\")\n ))\n sqlc.registerDataFrameAsTable(sqlc.createDataFrame(tab1), \"tab1\")\n sqlc.cacheTable(\"tab1\")\n\n /* Succeeds */\n val result1 = sqlc.sql(\"SELECT tab0._1,tab1._1 FROM tab0, tab1 GROUP BY tab0._1,tab1._1 ORDER BY tab0._1, tab1._1\").collect()\n assertResult(Array(Row(\"A1\", \"B1\"), Row(\"A1\", \"B2\"), Row(\"A2\", \"B1\"), Row(\"A2\", \"B2\")))(result1)\n\n /* Fails. Got: Array([A1], [A2]) */\n val result2 = sqlc.sql(\"SELECT tab1._1 FROM tab0, tab1 GROUP BY tab1._1 ORDER BY tab1._1\").collect()\n assertResult(Array(Row(\"B1\"), Row(\"B2\")))(result2)\n{code}","created":"2015-05-14T09:39:58.697+0000"},{"body":"Note that the bug is not related to GROUP BY, that's just a quick way to produce a Project logical plan with an empty projection list from SQL. Builing upon my previous test case, here are some further instances of the bug using logical plans and DataFrames:\n\n{code}\nimport org.apache.spark.sql.catalyst.dsl.plans._\n import org.apache.spark.sql.catalyst.dsl.expressions._\n \n val plan0 = sqlc.table(\"tab0\").logicalPlan.subquery('tab0)\n val plan1 = sqlc.table(\"tab1\").logicalPlan.subquery('tab1)\n \n /* Succeeds */\n val planA = plan0.select('_1 as \"c0\")\n .join(plan1.select('_1 as \"c1\"))\n .select('c0, 'c1)\n .orderBy('c0.asc, 'c1.asc)\n assertResult(Array(Row(\"A1\", \"B1\"), Row(\"A1\", \"B2\"), Row(\"A2\", \"B1\"), Row(\"A2\", \"B2\")))(DataFrame(sqlc, planA).collect())\n\n /* Fails. Got: Array([A1], [A1], [A2], [A2]) */\n val planB = plan0.select('_1 as \"c0\")\n .join(plan1.select('_1 as \"c1\"))\n .select('c1)\n .orderBy('c1.asc)\n assertResult(Array(Row(\"B1\"), Row(\"B1\"), Row(\"B2\"), Row(\"B2\")))(DataFrame(sqlc, planB).collect())\n\n /* Fails. Got: Array([A1], [A1], [A2], [A2]) */\n val planC = plan0.select()\n .join(plan1.select('_1 as \"c1\"))\n .select('c1)\n .orderBy('c1.asc)\n assertResult(Array(Row(\"B1\"), Row(\"B1\"), Row(\"B2\"), Row(\"B2\")))(DataFrame(sqlc, planC).collect())\n{code}","created":"2015-05-14T10:26:08.860+0000"},{"body":"This problem only happens for cached relations. Here is the root of the problem:\n\n{code}\n/* Fails. Got: Array(Row(\"A1\"), Row(\"A2\") */\nassertResult(Array(Row(), Row()))(\n InMemoryColumnarTableScan(Nil, Nil, sqlc.table(\"tab0\").queryExecution.sparkPlan.asInstanceOf[InMemoryColumnarTableScan].relation)\n .execute().collect()\n)\n{code}\n\nInMemoryColumnarTableScan returns the narrowest column when no attributes are requested:\n\n{code}\n // Find the ordinals and data types of the requested columns. If none are requested, use the\n // narrowest (the field with minimum default element size).\n val (requestedColumnIndices, requestedColumnDataTypes) = if (attributes.isEmpty) {\n val (narrowestOrdinal, narrowestDataType) =\n relation.output.zipWithIndex.map { case (a, ordinal) =>\n ordinal -> a.dataType\n } minBy { case (_, dataType) =>\n ColumnType(dataType).defaultSize\n }\n Seq(narrowestOrdinal) -> Seq(narrowestDataType)\n } else {\n attributes.map { a =>\n relation.output.indexWhere(_.exprId == a.exprId) -> a.dataType\n }.unzip\n }\n{code}\n\nIt seems this is what leads to incorrect results.","created":"2015-05-14T12:31:37.429+0000"},{"body":"User 'marmbrus' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6165","created":"2015-05-15T02:08:03.082+0000"},{"body":"Thanks for narrowing this down. I've opened a PR with a fix.","created":"2015-05-15T02:08:03.173+0000"},{"body":"Issue resolved by pull request 6165\n[https://github.com/apache/spark/pull/6165]","created":"2015-05-22T16:45:45.675+0000"}],"conversations":[{"body":"{code:java}\nval sqlContext = new SQLContext(sc)\nval tab0 = sc.parallelize(Seq(\n (83,0,38),\n (26,0,79),\n (43,81,24)\n ))\n sqlContext.registerDataFrameAsTable(sqlContext.createDataFrame(tab0), \"tab0\")\nsqlContext.cacheTable(\"tab0\") \nval df1 = sqlContext.sql(\"SELECT tab0._2, cor0._2 FROM tab0, tab0 cor0 GROUP BY tab0._2, cor0._2\")\nval result1 = df1.collect()\nval df2 = sqlContext.sql(\"SELECT cor0._2 FROM tab0, tab0 cor0 GROUP BY cor0._2\")\nval result2 = df2.collect()\nval df3 = sqlContext.sql(\"SELECT cor0._2 FROM tab0 cor0 GROUP BY cor0._2\")\nval result3 = df3.collect()\n{code}\n\nGiven the previous code, result2 equals to Row(43), Row(83), Row(26), which is wrong. These results correspond to cor0._1, instead of cor0._2. Correct results would be Row(0), Row(81), which are ok for the third query. The first query also produces valid results, and the only difference is that the left side of the join is not empty.","from":"reporter","subject":"Join with empty projection on one side produces invalid results"},{"body":"Any thoughts on this?","from":"developer"},{"body":"Sorry, my first example was not very clear. Here is a more precise one:\n\n{code}\n val sqlc = new SQLContext(sc)\n\n val tab0 = sc.parallelize(Seq(\n Tuple1(\"A1\"),\n Tuple1(\"A2\")\n ))\n sqlc.registerDataFrameAsTable(sqlc.createDataFrame(tab0), \"tab0\")\n sqlc.cacheTable(\"tab0\")\n\n val tab1 = sc.parallelize(Seq(\n Tuple1(\"B1\"),\n Tuple1(\"B2\")\n ))\n sqlc.registerDataFrameAsTable(sqlc.createDataFrame(tab1), \"tab1\")\n sqlc.cacheTable(\"tab1\")\n\n /* Succeeds */\n val result1 = sqlc.sql(\"SELECT tab0._1,tab1._1 FROM tab0, tab1 GROUP BY tab0._1,tab1._1 ORDER BY tab0._1, tab1._1\").collect()\n assertResult(Array(Row(\"A1\", \"B1\"), Row(\"A1\", \"B2\"), Row(\"A2\", \"B1\"), Row(\"A2\", \"B2\")))(result1)\n\n /* Fails. Got: Array([A1], [A2]) */\n val result2 = sqlc.sql(\"SELECT tab1._1 FROM tab0, tab1 GROUP BY tab1._1 ORDER BY tab1._1\").collect()\n assertResult(Array(Row(\"B1\"), Row(\"B2\")))(result2)\n{code}","from":"developer"},{"body":"Note that the bug is not related to GROUP BY, that's just a quick way to produce a Project logical plan with an empty projection list from SQL. Builing upon my previous test case, here are some further instances of the bug using logical plans and DataFrames:\n\n{code}\nimport org.apache.spark.sql.catalyst.dsl.plans._\n import org.apache.spark.sql.catalyst.dsl.expressions._\n \n val plan0 = sqlc.table(\"tab0\").logicalPlan.subquery('tab0)\n val plan1 = sqlc.table(\"tab1\").logicalPlan.subquery('tab1)\n \n /* Succeeds */\n val planA = plan0.select('_1 as \"c0\")\n .join(plan1.select('_1 as \"c1\"))\n .select('c0, 'c1)\n .orderBy('c0.asc, 'c1.asc)\n assertResult(Array(Row(\"A1\", \"B1\"), Row(\"A1\", \"B2\"), Row(\"A2\", \"B1\"), Row(\"A2\", \"B2\")))(DataFrame(sqlc, planA).collect())\n\n /* Fails. Got: Array([A1], [A1], [A2], [A2]) */\n val planB = plan0.select('_1 as \"c0\")\n .join(plan1.select('_1 as \"c1\"))\n .select('c1)\n .orderBy('c1.asc)\n assertResult(Array(Row(\"B1\"), Row(\"B1\"), Row(\"B2\"), Row(\"B2\")))(DataFrame(sqlc, planB).collect())\n\n /* Fails. Got: Array([A1], [A1], [A2], [A2]) */\n val planC = plan0.select()\n .join(plan1.select('_1 as \"c1\"))\n .select('c1)\n .orderBy('c1.asc)\n assertResult(Array(Row(\"B1\"), Row(\"B1\"), Row(\"B2\"), Row(\"B2\")))(DataFrame(sqlc, planC).collect())\n{code}","from":"developer"},{"body":"This problem only happens for cached relations. Here is the root of the problem:\n\n{code}\n/* Fails. Got: Array(Row(\"A1\"), Row(\"A2\") */\nassertResult(Array(Row(), Row()))(\n InMemoryColumnarTableScan(Nil, Nil, sqlc.table(\"tab0\").queryExecution.sparkPlan.asInstanceOf[InMemoryColumnarTableScan].relation)\n .execute().collect()\n)\n{code}\n\nInMemoryColumnarTableScan returns the narrowest column when no attributes are requested:\n\n{code}\n // Find the ordinals and data types of the requested columns. If none are requested, use the\n // narrowest (the field with minimum default element size).\n val (requestedColumnIndices, requestedColumnDataTypes) = if (attributes.isEmpty) {\n val (narrowestOrdinal, narrowestDataType) =\n relation.output.zipWithIndex.map { case (a, ordinal) =>\n ordinal -> a.dataType\n } minBy { case (_, dataType) =>\n ColumnType(dataType).defaultSize\n }\n Seq(narrowestOrdinal) -> Seq(narrowestDataType)\n } else {\n attributes.map { a =>\n relation.output.indexWhere(_.exprId == a.exprId) -> a.dataType\n }.unzip\n }\n{code}\n\nIt seems this is what leads to incorrect results.","from":"developer"},{"body":"User 'marmbrus' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6165","from":"developer"},{"body":"Thanks for narrowing this down. I've opened a PR with a fix.","from":"developer"},{"body":"Issue resolved by pull request 6165\n[https://github.com/apache/spark/pull/6165]","from":"developer"}],"created":"2015-04-07T14:49:47.000+0000","description":"{code:java}\nval sqlContext = new SQLContext(sc)\nval tab0 = sc.parallelize(Seq(\n (83,0,38),\n (26,0,79),\n (43,81,24)\n ))\n sqlContext.registerDataFrameAsTable(sqlContext.createDataFrame(tab0), \"tab0\")\nsqlContext.cacheTable(\"tab0\") \nval df1 = sqlContext.sql(\"SELECT tab0._2, cor0._2 FROM tab0, tab0 cor0 GROUP BY tab0._2, cor0._2\")\nval result1 = df1.collect()\nval df2 = sqlContext.sql(\"SELECT cor0._2 FROM tab0, tab0 cor0 GROUP BY cor0._2\")\nval result2 = df2.collect()\nval df3 = sqlContext.sql(\"SELECT cor0._2 FROM tab0 cor0 GROUP BY cor0._2\")\nval result3 = df3.collect()\n{code}\n\nGiven the previous code, result2 equals to Row(43), Row(83), Row(26), which is wrong. These results correspond to cor0._1, instead of cor0._2. Correct results would be Row(0), Row(81), which are ok for the third query. The first query also produces valid results, and the only difference is that the left side of the join is not empty.","issue_id":"12818951","key":"SPARK-6743","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-05-22T16:45:45.000+0000","role":"fixed_distractor","summary":"Join with empty projection on one side produces invalid results"} {"case_id":"12819481","cluster":"DISTRACTOR-SPARK-6785","comments":[{"body":"Hi Patrick,\nI would like to work on this issue. Seems like the date conversion is thrown off by the time-zone adjustments and the fact that the interchange type is days instead of millis. I am preparing a pull-request which will also include test cases to cover more date conversion scenarios.\n\n","created":"2015-05-13T20:54:40.261+0000"},{"body":"User 'ckadner' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6236","created":"2015-05-18T19:18:02.353+0000"},{"body":"User 'ckadner' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6242","created":"2015-05-18T20:14:02.269+0000"},{"body":"{panel:borderStyle=dashed|borderColor=#ccc|bgColor=#FFFFCE}\nPull Request +[6242|https://github.com/apache/spark/pull/6242]+\n{panel}\n\\\\\nBefore my fix, the from-and-to Java date conversion of dates before 1970 will only work for {{java.sql.Date}} objects that reflect a date and time exactly at midnight in the System's local time zone. \nOtherwise, if the Date's time is just one millisecond before or after midnight, the result of the above conversion will be offset by one day for Dates before 1970 because of a rounding (truncation) flaw in the function {{DateUtils.millisToDays(Long):Int}}\n\n\\\\\n\n{code}\n scala> val df = new SimpleDateFormat(\"yyyy-MM-dd HH:mm:ss\")\n df: java.text.SimpleDateFormat = yyyy-MM-dd HH:mm:ss\n\n scala> val d1 = new Date(df.parse(\"1969-01-01 00:00:00\").getTime)\n d2: java.sql.Date = 1969-01-01\n\t\n scala> val d2 = new Date(df.parse(\"1969-01-01 00:00:01\").getTime)\n d2: java.sql.Date = 1969-01-01\n\n scala> DateUtils.toJavaDate(DateUtils.fromJavaDate(d1))\n res1: java.sql.Date = 1969-01-01\n\t\n scala> DateUtils.toJavaDate(DateUtils.fromJavaDate(d2))\n res2: java.sql.Date = 1969-01-02\n{code}\n\n\\\\\n\nWhat is the code doing and how to fix it:\n\n\\\\\n\n - A {{java.util.Date}} is represented by milliseconds ({{Long}}) since the Epoch (1970/01/01 0:00:00 GMT) with positive numbers for dates after and negative numbers for dates before 1970\n \n - The function {{DateUtils.fromJavaDate(java.util.Date):Int}} calculates the number of full days passed since 1970/01/01 00:00:00 (local time, not UTC), but by using the data type {{Long}} (as opposed to {{Double}}) when converting milliseconds to days it essentially truncates the fractional part of days passed (disregarding the impact of hours, minutes, seconds)\n \n - The function {{DateUtils.toJavaDate(Int):Date}} converts the given number of days into milliseconds and adds it 1970/01/01 00:00:00 (local time, not UTC)\n\n - _Side note: The time-zone offset from UTC is factored in when converting a Date to days and removed when converting days to Date, so the time-zone shifting is neutralized in the round-trip conversion {{toJavaDate(fromJavaDate(java.util.Date))}}._\n \n - The truncation of partial days is not a problem for dates after 1970 since adding a fraction of a day to any date will not flip the calendar to the next day (since all our Dates start 0:00:00 AM)\n \n - That truncation of partial days however is a problem when subtracting even a second from a {{Date}} with time at 0:00:00 AM which should turn the calender back one day to the previous date\n \n - Ideally the date conversion should be done using milliseconds, but since using days has been established already, the fix is to work with {{Double}} to preserve fractions of days and use {{floor()}} instead of the implicit truncate to round to a full number of days ({{Int}})\n\n\\\\\n\nPseudo-code example, adding or subtracting 1 hour to Date \"1970/01/01 0:00:00\" using milliseconds...\n\n{code}\n\"1970-01-01 0:00:00\" + 1 hr = \"1970-01-01 1:00:00\"\n\"1970-01-01 0:00:00\" - 1 hr = \"1969-12-31 23:00:00\"\n{code}\n\n\\\\\n\nSame example, using full days. One hour is about 0.04 days. Using {{trunc()}} versus {{floor()}} we get ... \n\n{code}\ntrunc(+0.04) = +0 --> \"1970-01-01\" + 0 days = \"1970-01-01\" (correct)\nfloor(+0.04) = +0 --> \"1970-01-01\" + 0 days = \"1970-01-01\" (correct)\n\ntrunc(-0.04) = -0 --> \"1970-01-01\" + -0 days = \"1970-01-01\" (incorrect, bug)\nfloor(-0.04) = -1 --> \"1970-01-01\" + -1 day = \"1969-12-31\" (correct, fix)\n{code}\n\n{code} \ndef trunc(d: Dounble): Int = d.toInt\n{code}","created":"2015-05-18T22:15:38.576+0000"},{"body":"User 'ckadner' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6983","created":"2015-06-24T11:38:06.419+0000"},{"body":"Issue resolved by pull request 6983\n[https://github.com/apache/spark/pull/6983]","created":"2015-06-30T19:23:11.638+0000"}],"conversations":[{"body":"{code}\nscala> val d = new Date(100)\nd: java.sql.Date = 1969-12-31\n\nscala> DateUtils.toJavaDate(DateUtils.fromJavaDate(d))\nres1: java.sql.Date = 1970-01-01\n\n{code}","from":"reporter","subject":"DateUtils can not handle date before 1970/01/01 correctly"},{"body":"Hi Patrick,\nI would like to work on this issue. Seems like the date conversion is thrown off by the time-zone adjustments and the fact that the interchange type is days instead of millis. I am preparing a pull-request which will also include test cases to cover more date conversion scenarios.\n\n","from":"developer"},{"body":"User 'ckadner' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6236","from":"developer"},{"body":"User 'ckadner' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6242","from":"developer"},{"body":"{panel:borderStyle=dashed|borderColor=#ccc|bgColor=#FFFFCE}\nPull Request +[6242|https://github.com/apache/spark/pull/6242]+\n{panel}\n\\\\\nBefore my fix, the from-and-to Java date conversion of dates before 1970 will only work for {{java.sql.Date}} objects that reflect a date and time exactly at midnight in the System's local time zone. \nOtherwise, if the Date's time is just one millisecond before or after midnight, the result of the above conversion will be offset by one day for Dates before 1970 because of a rounding (truncation) flaw in the function {{DateUtils.millisToDays(Long):Int}}\n\n\\\\\n\n{code}\n scala> val df = new SimpleDateFormat(\"yyyy-MM-dd HH:mm:ss\")\n df: java.text.SimpleDateFormat = yyyy-MM-dd HH:mm:ss\n\n scala> val d1 = new Date(df.parse(\"1969-01-01 00:00:00\").getTime)\n d2: java.sql.Date = 1969-01-01\n\t\n scala> val d2 = new Date(df.parse(\"1969-01-01 00:00:01\").getTime)\n d2: java.sql.Date = 1969-01-01\n\n scala> DateUtils.toJavaDate(DateUtils.fromJavaDate(d1))\n res1: java.sql.Date = 1969-01-01\n\t\n scala> DateUtils.toJavaDate(DateUtils.fromJavaDate(d2))\n res2: java.sql.Date = 1969-01-02\n{code}\n\n\\\\\n\nWhat is the code doing and how to fix it:\n\n\\\\\n\n - A {{java.util.Date}} is represented by milliseconds ({{Long}}) since the Epoch (1970/01/01 0:00:00 GMT) with positive numbers for dates after and negative numbers for dates before 1970\n \n - The function {{DateUtils.fromJavaDate(java.util.Date):Int}} calculates the number of full days passed since 1970/01/01 00:00:00 (local time, not UTC), but by using the data type {{Long}} (as opposed to {{Double}}) when converting milliseconds to days it essentially truncates the fractional part of days passed (disregarding the impact of hours, minutes, seconds)\n \n - The function {{DateUtils.toJavaDate(Int):Date}} converts the given number of days into milliseconds and adds it 1970/01/01 00:00:00 (local time, not UTC)\n\n - _Side note: The time-zone offset from UTC is factored in when converting a Date to days and removed when converting days to Date, so the time-zone shifting is neutralized in the round-trip conversion {{toJavaDate(fromJavaDate(java.util.Date))}}._\n \n - The truncation of partial days is not a problem for dates after 1970 since adding a fraction of a day to any date will not flip the calendar to the next day (since all our Dates start 0:00:00 AM)\n \n - That truncation of partial days however is a problem when subtracting even a second from a {{Date}} with time at 0:00:00 AM which should turn the calender back one day to the previous date\n \n - Ideally the date conversion should be done using milliseconds, but since using days has been established already, the fix is to work with {{Double}} to preserve fractions of days and use {{floor()}} instead of the implicit truncate to round to a full number of days ({{Int}})\n\n\\\\\n\nPseudo-code example, adding or subtracting 1 hour to Date \"1970/01/01 0:00:00\" using milliseconds...\n\n{code}\n\"1970-01-01 0:00:00\" + 1 hr = \"1970-01-01 1:00:00\"\n\"1970-01-01 0:00:00\" - 1 hr = \"1969-12-31 23:00:00\"\n{code}\n\n\\\\\n\nSame example, using full days. One hour is about 0.04 days. Using {{trunc()}} versus {{floor()}} we get ... \n\n{code}\ntrunc(+0.04) = +0 --> \"1970-01-01\" + 0 days = \"1970-01-01\" (correct)\nfloor(+0.04) = +0 --> \"1970-01-01\" + 0 days = \"1970-01-01\" (correct)\n\ntrunc(-0.04) = -0 --> \"1970-01-01\" + -0 days = \"1970-01-01\" (incorrect, bug)\nfloor(-0.04) = -1 --> \"1970-01-01\" + -1 day = \"1969-12-31\" (correct, fix)\n{code}\n\n{code} \ndef trunc(d: Dounble): Int = d.toInt\n{code}","from":"developer"},{"body":"User 'ckadner' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6983","from":"developer"},{"body":"Issue resolved by pull request 6983\n[https://github.com/apache/spark/pull/6983]","from":"developer"}],"created":"2015-04-08T22:33:56.000+0000","description":"{code}\nscala> val d = new Date(100)\nd: java.sql.Date = 1969-12-31\n\nscala> DateUtils.toJavaDate(DateUtils.fromJavaDate(d))\nres1: java.sql.Date = 1970-01-01\n\n{code}","issue_id":"12819481","key":"SPARK-6785","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-06-30T19:23:11.000+0000","role":"fixed_distractor","summary":"DateUtils can not handle date before 1970/01/01 correctly"} {"case_id":"12828334","cluster":"DISTRACTOR-SPARK-7483","comments":[{"body":"(Updated) Maybe this is a bug...will look into it.\n","created":"2015-05-08T21:21:23.886+0000"},{"body":"Does it fix anything if you give Kryo more info, such as explicit registration of relevant classes?","created":"2015-05-08T21:25:52.127+0000"},{"body":"hmm, class org.apache.spark.mllib.fpm.FPTree is private and in spark-mllib. Spark is registering classes in spark-core in org.apache.spark.serializer.KryoSerializer#toRegister so it is not straightforward how to do that easily.\n\nAnd I do not think this is the case of registering, looking at spark serialization doc \"Finally, if you don’t register your custom classes, Kryo will still work, but it will have to store the full class name with each object, which is wasteful.\"","created":"2015-05-11T09:30:44.785+0000"},{"body":"I agree; it should work, but I'm not sure why it's failing. I'm not that familiar with Kryo, but I'll ask around. Thanks for reporting this!","created":"2015-05-11T17:56:29.201+0000"},{"body":"I encountered the same bug.\nAdding \nsparkConf.registerKryoClasses(Array(classOf[ArrayBuffer[String]], classOf[ListBuffer[String]]))\nseems to fix the problem.","created":"2015-07-02T02:27:03.108+0000"},{"body":"this solves the problem for me as well.\n\nmaybe those classes should be registered in kryo from spark when user code is using kryo serialization?","created":"2015-07-02T15:13:10.302+0000"},{"body":"This does NOT work.\n\nRegistering the classes stops it from crashing, but produces a bug in the FP-Growth algorithm.\n\nSpecifically, the frequency counts for itemsets are wrong.\n\n:(","created":"2015-07-02T23:24:16.602+0000"},{"body":"hmm that's right - adding it in FPGrowthSuite in spark fails first spec.","created":"2015-07-03T07:12:22.885+0000"},{"body":"I've tried to change the nodes property of Summay class inside object FPTree from ListBuffer to ArrayBuffer. And it seems to fix the exception problem as well as producing the right item counts so far. Suspect the problem with KryoSerializer dealing with ListBuffer class, but haven't looked into details. The code looks like:\n\nprivate[fpm] object FPTree {\n\n /** Representing a node in an FP-Tree. */\n class Node[T](val parent: Node[T]) extends Serializable {\n var item: T = _\n var count: Long = 0L\n val children: mutable.Map[T, Node[T]] = mutable.Map.empty\n\n def isRoot: Boolean = parent == null\n }\n\n /** Summary of a item in an FP-Tree. */\n private class Summary[T] extends Serializable {\n var count: Long = 0L\n val nodes: mutable.ArrayBuffer[Node[T]] = mutable.ArrayBuffer.empty\n }\n}\n\n","created":"2015-07-05T22:14:35.541+0000"},{"body":"kyro not support ListBuffer because ListBuffer don't have any \"zero argument constructor\".\nrefer to : https://github.com/EsotericSoftware/kryo#using-standard-java-serialization\n\n\"By default, if a class has a zero argument constructor then it is invoked via ReflectASM or reflection, otherwise an exception is thrown. \"\n\nIs that the reason?\n\nWhen using kyro , use ArrayBuffer instead of ListBuffer","created":"2015-09-24T11:32:29.930+0000"},{"body":"User 'mark800' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11041","created":"2016-02-03T02:20:03.765+0000"},{"body":"Upgrade Chill to 0.7.2, then everything is fine.\n\nNew version registers more Scala classes, including ListBuffer to support Kryo with FPGrowth.","created":"2016-02-03T02:24:05.815+0000"},{"body":"Issue resolved by pull request 11041\n[https://github.com/apache/spark/pull/11041]","created":"2016-02-27T13:51:22.201+0000"}],"conversations":[{"body":"When using FPGrowth algorithm with KryoSerializer - Spark fails with\n\n{code}\nJob aborted due to stage failure: Task 0 in stage 9.0 failed 1 times, most recent failure: Lost task 0.0 in stage 9.0 (TID 16, localhost): com.esotericsoftware.kryo.KryoException: java.lang.IllegalArgumentException: Can not set final scala.collection.mutable.ListBuffer field org.apache.spark.mllib.fpm.FPTree$Summary.nodes to scala.collection.mutable.ArrayBuffer\nSerialization trace:\nnodes (org.apache.spark.mllib.fpm.FPTree$Summary)\norg$apache$spark$mllib$fpm$FPTree$$summaries (org.apache.spark.mllib.fpm.FPTree)\n{code}\n\nThis can be easily reproduced in spark codebase by setting \n{code}\nconf.set(\"spark.serializer\", \"org.apache.spark.serializer.KryoSerializer\")\n{code} and running FPGrowthSuite.\n\n\n","from":"reporter","subject":"[MLLib] Using Kryo with FPGrowth fails with an exception"},{"body":"(Updated) Maybe this is a bug...will look into it.\n","from":"developer"},{"body":"Does it fix anything if you give Kryo more info, such as explicit registration of relevant classes?","from":"developer"},{"body":"hmm, class org.apache.spark.mllib.fpm.FPTree is private and in spark-mllib. Spark is registering classes in spark-core in org.apache.spark.serializer.KryoSerializer#toRegister so it is not straightforward how to do that easily.\n\nAnd I do not think this is the case of registering, looking at spark serialization doc \"Finally, if you don’t register your custom classes, Kryo will still work, but it will have to store the full class name with each object, which is wasteful.\"","from":"developer"},{"body":"I agree; it should work, but I'm not sure why it's failing. I'm not that familiar with Kryo, but I'll ask around. Thanks for reporting this!","from":"developer"},{"body":"I encountered the same bug.\nAdding \nsparkConf.registerKryoClasses(Array(classOf[ArrayBuffer[String]], classOf[ListBuffer[String]]))\nseems to fix the problem.","from":"developer"},{"body":"this solves the problem for me as well.\n\nmaybe those classes should be registered in kryo from spark when user code is using kryo serialization?","from":"developer"},{"body":"This does NOT work.\n\nRegistering the classes stops it from crashing, but produces a bug in the FP-Growth algorithm.\n\nSpecifically, the frequency counts for itemsets are wrong.\n\n:(","from":"developer"},{"body":"hmm that's right - adding it in FPGrowthSuite in spark fails first spec.","from":"developer"},{"body":"I've tried to change the nodes property of Summay class inside object FPTree from ListBuffer to ArrayBuffer. And it seems to fix the exception problem as well as producing the right item counts so far. Suspect the problem with KryoSerializer dealing with ListBuffer class, but haven't looked into details. The code looks like:\n\nprivate[fpm] object FPTree {\n\n /** Representing a node in an FP-Tree. */\n class Node[T](val parent: Node[T]) extends Serializable {\n var item: T = _\n var count: Long = 0L\n val children: mutable.Map[T, Node[T]] = mutable.Map.empty\n\n def isRoot: Boolean = parent == null\n }\n\n /** Summary of a item in an FP-Tree. */\n private class Summary[T] extends Serializable {\n var count: Long = 0L\n val nodes: mutable.ArrayBuffer[Node[T]] = mutable.ArrayBuffer.empty\n }\n}\n\n","from":"developer"},{"body":"kyro not support ListBuffer because ListBuffer don't have any \"zero argument constructor\".\nrefer to : https://github.com/EsotericSoftware/kryo#using-standard-java-serialization\n\n\"By default, if a class has a zero argument constructor then it is invoked via ReflectASM or reflection, otherwise an exception is thrown. \"\n\nIs that the reason?\n\nWhen using kyro , use ArrayBuffer instead of ListBuffer","from":"developer"},{"body":"User 'mark800' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/11041","from":"developer"},{"body":"Upgrade Chill to 0.7.2, then everything is fine.\n\nNew version registers more Scala classes, including ListBuffer to support Kryo with FPGrowth.","from":"developer"},{"body":"Issue resolved by pull request 11041\n[https://github.com/apache/spark/pull/11041]","from":"developer"}],"created":"2015-05-08T11:53:23.000+0000","description":"When using FPGrowth algorithm with KryoSerializer - Spark fails with\n\n{code}\nJob aborted due to stage failure: Task 0 in stage 9.0 failed 1 times, most recent failure: Lost task 0.0 in stage 9.0 (TID 16, localhost): com.esotericsoftware.kryo.KryoException: java.lang.IllegalArgumentException: Can not set final scala.collection.mutable.ListBuffer field org.apache.spark.mllib.fpm.FPTree$Summary.nodes to scala.collection.mutable.ArrayBuffer\nSerialization trace:\nnodes (org.apache.spark.mllib.fpm.FPTree$Summary)\norg$apache$spark$mllib$fpm$FPTree$$summaries (org.apache.spark.mllib.fpm.FPTree)\n{code}\n\nThis can be easily reproduced in spark codebase by setting \n{code}\nconf.set(\"spark.serializer\", \"org.apache.spark.serializer.KryoSerializer\")\n{code} and running FPGrowthSuite.\n\n\n","issue_id":"12828334","key":"SPARK-7483","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-02-27T13:51:22.000+0000","role":"fixed_distractor","summary":"[MLLib] Using Kryo with FPGrowth fails with an exception"} {"case_id":"12832306","cluster":"DISTRACTOR-SPARK-7837","comments":[{"body":"We have made the parquet reader side robust to files left in _temporary. So, this problem should have a much smaller impact. \n\nI am re-targeting it to 1.5. Will keep an eye on it and investigate the root cause.","created":"2015-05-29T00:25:41.099+0000"},{"body":"Seems https://www.mail-archive.com/user@spark.apache.org/msg30327.html is about the same issue.","created":"2015-06-16T05:12:22.041+0000"},{"body":"Facing the Same Issue at my End, I am processing around 68TB of data and transforming it to Parquet and I am also facing it. And this App is on Production, will really like it be resolved on highest priority in 1.5.0.\n\nThanks Guys","created":"2015-07-15T13:04:20.881+0000"},{"body":"[~mkanchwala] One quick clarification question. Did your spark job fail because of the NPE? Or, your job succeeded but you saw the NPE in some speculative tasks? ","created":"2015-07-15T14:48:43.056+0000"},{"body":"Job Succeeded But shown NPE, I am worried about any data loss between this transformation","created":"2015-07-15T16:46:07.245+0000"},{"body":"Also I notice that I completed my transformation to Parquet but after the job completion it is taking too much time to save the data to s3.\n\nInput : s3:///input\nOutput : s3:///output\n\nI have around 27K files and it is writing the data at a speed of 1k / 20 mins and my job completed in an hour, can you share the details with me why it's so slow for parquet?","created":"2015-07-15T17:10:03.170+0000"},{"body":"[~mkanchwala] There is a bug (https://issues.apache.org/jira/browse/SPARK-8406), which potentially may cause data loss for a large job. Please use Spark 1.4.1 as soon as it is released (I think today or tomorrow) or manually apply the fix (https://github.com/nemccarthy/spark/commit/ba365909b964fe5a5851d88f5f7b7edcd1998142) to your Spark 1.4.0 source code.\n\nRegarding the slowness of saving parquet files in S3, a possible cause is that Parquet's original output committer ({{org.apache.parquet.hadoop.ParquetOutputCommitter}}) will first write data in the temporary dir and then move them to the right place when it commits tasks. This behavior is not necessary in most of the cases for S3 because \"S3 supports multiple writers outputting to the same file, where visibility is guaranteed to be atomic\" (https://gist.github.com/aarondav/c513916e72101bbe14ec). Once you upgrade to Spark 1.4.1, you can set {{spark.sql.parquet.output.committer.class}} to {{org.apache.spark.sql.parquet.DirectParquetOutputCommitter}} in your hadoop conf, which will write output files directly to their final locations. The only case that is not safe to use DirectParquetOutputCommitter is when you append data to an existing table. In this case, Spark 1.4.1 will internally switch back to the original Parquet output committer.","created":"2015-07-15T18:05:09.947+0000"},{"body":"Thanks [~yhuai] for the update. Will be looking forward on Spark 1.4.1 release. Also I've filed a seperate issue to the Parent Bug as SPARK-9072 : Parquet : Writing data to S3 very slowly . Can you plan this and make the neccessary changes for the Same.\n\nThanks\n\n","created":"2015-07-15T18:23:14.879+0000"},{"body":"When speculation is on, I can reproduce it every time I run\n{code}\nsc.parallelize((1 to 100), 20).map { i =>\n if (i == 4 || i == 29) Thread.sleep(10000) else Thread.sleep(100)\n i\n}.map(i => Tuple1(i)).toDF(\"i\").write.mode(\"overwrite\").format(\"parquet\").save(\"/home/yin/outputCommitter\")\n{code}","created":"2015-08-14T20:46:25.685+0000"},{"body":"Just a note to people who want to reproduce this issue:\n\n# You need to start a Spark cluster with at least two workers running on two distinct nodes. Speculation isn't enabled when running in local mode or single node cluster. If you only have a single machine, you'll probably have to resort to VMs\n# Don't forget to set {{spark.speculation}} to {{true}} (it's {{false}} by default)","created":"2015-08-16T17:20:26.851+0000"},{"body":"ah i see the reason of the NPE. We actually called close twice. In DefaultWriterContainer's writeRows, we start to write out rows and at the end we call commitTask. In commitTask, we first call writer.close and then we call super.commitTask(). In writer.close, we triggered ParquetRecordWriter's close, which sets columnStore to null. Then, because the speculative task's commit is rejected (i.e. super.commitTask() is rejected by OutputCommitCoordinator), we cal labortTask, which triggers writer.close again. Inside writer.close we call ParquetRecordWriter's close and then we get NPE because columnStore is already set to null.","created":"2015-08-17T02:54:09.157+0000"},{"body":"Good job!","created":"2015-08-17T05:37:31.317+0000"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/8236","created":"2015-08-17T09:43:03.730+0000"},{"body":"Issue resolved by pull request 8236\n[https://github.com/apache/spark/pull/8236]","created":"2015-08-17T16:59:56.469+0000"},{"body":"[~mkanchwala] Based on our investigate, the NPE was caused by calling {{close}} of a Parquet record writer twice. The first time we call {{close}}, parquet sets {{columnStore}} (an parquet internal variable) to null and the second time we call {{close}}, the NPE is triggered.\n\nThis issue does not cause data loss.","created":"2015-08-17T18:01:45.737+0000"},{"body":"User 'LuciferYang' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/37245","created":"2022-07-21T12:45:22.256+0000"},{"body":"User 'LuciferYang' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/37245","created":"2022-07-21T12:46:17.237+0000"}],"conversations":[{"body":"The query is like {{df.orderBy(...).saveAsTable(...)}}.\n\nWhen there is no partitioning columns and there is a skewed key, I found the following exception in speculative tasks. After these failures, seems we could not call {{SparkHadoopMapRedUtil.commitTask}} correctly.\n\n{code}\njava.lang.NullPointerException\n\tat parquet.hadoop.InternalParquetRecordWriter.flushRowGroupToStore(InternalParquetRecordWriter.java:146)\n\tat parquet.hadoop.InternalParquetRecordWriter.close(InternalParquetRecordWriter.java:112)\n\tat parquet.hadoop.ParquetRecordWriter.close(ParquetRecordWriter.java:73)\n\tat org.apache.spark.sql.parquet.ParquetOutputWriter.close(newParquet.scala:115)\n\tat org.apache.spark.sql.sources.DefaultWriterContainer.abortTask(commands.scala:385)\n\tat org.apache.spark.sql.sources.InsertIntoHadoopFsRelation.org$apache$spark$sql$sources$InsertIntoHadoopFsRelation$$writeRows$1(commands.scala:150)\n\tat org.apache.spark.sql.sources.InsertIntoHadoopFsRelation$$anonfun$insert$1.apply(commands.scala:122)\n\tat org.apache.spark.sql.sources.InsertIntoHadoopFsRelation$$anonfun$insert$1.apply(commands.scala:122)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:63)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:70)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}","from":"reporter","subject":"NPE when save as parquet in speculative tasks"},{"body":"We have made the parquet reader side robust to files left in _temporary. So, this problem should have a much smaller impact. \n\nI am re-targeting it to 1.5. Will keep an eye on it and investigate the root cause.","from":"developer"},{"body":"Seems https://www.mail-archive.com/user@spark.apache.org/msg30327.html is about the same issue.","from":"developer"},{"body":"Facing the Same Issue at my End, I am processing around 68TB of data and transforming it to Parquet and I am also facing it. And this App is on Production, will really like it be resolved on highest priority in 1.5.0.\n\nThanks Guys","from":"developer"},{"body":"[~mkanchwala] One quick clarification question. Did your spark job fail because of the NPE? Or, your job succeeded but you saw the NPE in some speculative tasks? ","from":"developer"},{"body":"Job Succeeded But shown NPE, I am worried about any data loss between this transformation","from":"developer"},{"body":"Also I notice that I completed my transformation to Parquet but after the job completion it is taking too much time to save the data to s3.\n\nInput : s3:///input\nOutput : s3:///output\n\nI have around 27K files and it is writing the data at a speed of 1k / 20 mins and my job completed in an hour, can you share the details with me why it's so slow for parquet?","from":"developer"},{"body":"[~mkanchwala] There is a bug (https://issues.apache.org/jira/browse/SPARK-8406), which potentially may cause data loss for a large job. Please use Spark 1.4.1 as soon as it is released (I think today or tomorrow) or manually apply the fix (https://github.com/nemccarthy/spark/commit/ba365909b964fe5a5851d88f5f7b7edcd1998142) to your Spark 1.4.0 source code.\n\nRegarding the slowness of saving parquet files in S3, a possible cause is that Parquet's original output committer ({{org.apache.parquet.hadoop.ParquetOutputCommitter}}) will first write data in the temporary dir and then move them to the right place when it commits tasks. This behavior is not necessary in most of the cases for S3 because \"S3 supports multiple writers outputting to the same file, where visibility is guaranteed to be atomic\" (https://gist.github.com/aarondav/c513916e72101bbe14ec). Once you upgrade to Spark 1.4.1, you can set {{spark.sql.parquet.output.committer.class}} to {{org.apache.spark.sql.parquet.DirectParquetOutputCommitter}} in your hadoop conf, which will write output files directly to their final locations. The only case that is not safe to use DirectParquetOutputCommitter is when you append data to an existing table. In this case, Spark 1.4.1 will internally switch back to the original Parquet output committer.","from":"developer"},{"body":"Thanks [~yhuai] for the update. Will be looking forward on Spark 1.4.1 release. Also I've filed a seperate issue to the Parent Bug as SPARK-9072 : Parquet : Writing data to S3 very slowly . Can you plan this and make the neccessary changes for the Same.\n\nThanks\n\n","from":"developer"},{"body":"When speculation is on, I can reproduce it every time I run\n{code}\nsc.parallelize((1 to 100), 20).map { i =>\n if (i == 4 || i == 29) Thread.sleep(10000) else Thread.sleep(100)\n i\n}.map(i => Tuple1(i)).toDF(\"i\").write.mode(\"overwrite\").format(\"parquet\").save(\"/home/yin/outputCommitter\")\n{code}","from":"developer"},{"body":"Just a note to people who want to reproduce this issue:\n\n# You need to start a Spark cluster with at least two workers running on two distinct nodes. Speculation isn't enabled when running in local mode or single node cluster. If you only have a single machine, you'll probably have to resort to VMs\n# Don't forget to set {{spark.speculation}} to {{true}} (it's {{false}} by default)","from":"developer"},{"body":"ah i see the reason of the NPE. We actually called close twice. In DefaultWriterContainer's writeRows, we start to write out rows and at the end we call commitTask. In commitTask, we first call writer.close and then we call super.commitTask(). In writer.close, we triggered ParquetRecordWriter's close, which sets columnStore to null. Then, because the speculative task's commit is rejected (i.e. super.commitTask() is rejected by OutputCommitCoordinator), we cal labortTask, which triggers writer.close again. Inside writer.close we call ParquetRecordWriter's close and then we get NPE because columnStore is already set to null.","from":"developer"},{"body":"Good job!","from":"developer"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/8236","from":"developer"},{"body":"Issue resolved by pull request 8236\n[https://github.com/apache/spark/pull/8236]","from":"developer"},{"body":"[~mkanchwala] Based on our investigate, the NPE was caused by calling {{close}} of a Parquet record writer twice. The first time we call {{close}}, parquet sets {{columnStore}} (an parquet internal variable) to null and the second time we call {{close}}, the NPE is triggered.\n\nThis issue does not cause data loss.","from":"developer"},{"body":"User 'LuciferYang' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/37245","from":"developer"},{"body":"User 'LuciferYang' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/37245","from":"developer"}],"created":"2015-05-23T00:15:37.000+0000","description":"The query is like {{df.orderBy(...).saveAsTable(...)}}.\n\nWhen there is no partitioning columns and there is a skewed key, I found the following exception in speculative tasks. After these failures, seems we could not call {{SparkHadoopMapRedUtil.commitTask}} correctly.\n\n{code}\njava.lang.NullPointerException\n\tat parquet.hadoop.InternalParquetRecordWriter.flushRowGroupToStore(InternalParquetRecordWriter.java:146)\n\tat parquet.hadoop.InternalParquetRecordWriter.close(InternalParquetRecordWriter.java:112)\n\tat parquet.hadoop.ParquetRecordWriter.close(ParquetRecordWriter.java:73)\n\tat org.apache.spark.sql.parquet.ParquetOutputWriter.close(newParquet.scala:115)\n\tat org.apache.spark.sql.sources.DefaultWriterContainer.abortTask(commands.scala:385)\n\tat org.apache.spark.sql.sources.InsertIntoHadoopFsRelation.org$apache$spark$sql$sources$InsertIntoHadoopFsRelation$$writeRows$1(commands.scala:150)\n\tat org.apache.spark.sql.sources.InsertIntoHadoopFsRelation$$anonfun$insert$1.apply(commands.scala:122)\n\tat org.apache.spark.sql.sources.InsertIntoHadoopFsRelation$$anonfun$insert$1.apply(commands.scala:122)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:63)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:70)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:213)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:745)\n{code}","issue_id":"12832306","key":"SPARK-7837","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-08-17T16:59:56.000+0000","role":"fixed_distractor","summary":"NPE when save as parquet in speculative tasks"} {"case_id":"12832720","cluster":"DISTRACTOR-SPARK-7869","comments":[{"body":"This is also a problem for {{json}} columns, not just {{jsonb}} ones. \n\nIt would be nice to get the JSON as a String column, instead of an error.\n","created":"2015-08-24T17:15:52.895+0000"},{"body":"User '0x0FFF' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/8948","created":"2015-09-30T10:43:05.108+0000"},{"body":"Still not resolves in spark version 1.6. I am seeing the same issue in spark","created":"2016-06-20T13:53:22.354+0000"},{"body":"Hi Nipun, I am using Spark. Is there a way to insert the Jsonb data into postgres. We have a new project in design phase. We are thinking of using Apache Spark + Postgres DB. But we are facing issues while inserting JSONB data type. \nIs there a support for Postgres-JSONB from spark? Can you please help us ? I have posted this question in the issues but no response. We really need help, can you please let us know if there is a way of inserting?? ","created":"2017-02-22T15:12:46.059+0000"}],"conversations":[{"body":"Most of our tables load into dataframes just fine with postgres. However we have a number of tables leveraging the JSONB datatype. Spark will error and refuse to load this table. While asking for Spark to support JSONB might be a tall order in the short term, it would be great if Spark would at least load the table ignoring the columns it can't load or have it be an option.\n{code}\npdf = sql_context.load(source=\"jdbc\", url=url, dbtable=\"table_of_json\")\n\nPy4JJavaError: An error occurred while calling o41.load.\n: java.sql.SQLException: Unsupported type 1111\n at org.apache.spark.sql.jdbc.JDBCRDD$.getCatalystType(JDBCRDD.scala:78)\n at org.apache.spark.sql.jdbc.JDBCRDD$.resolveTable(JDBCRDD.scala:112)\n at org.apache.spark.sql.jdbc.JDBCRelation.(JDBCRelation.scala:133)\n at org.apache.spark.sql.jdbc.DefaultSource.createRelation(JDBCRelation.scala:121)\n at org.apache.spark.sql.sources.ResolvedDataSource$.apply(ddl.scala:219)\n at org.apache.spark.sql.SQLContext.load(SQLContext.scala:697)\n at org.apache.spark.sql.SQLContext.load(SQLContext.scala:685)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:606)\n at py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:231)\n at py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:379)\n at py4j.Gateway.invoke(Gateway.java:259)\n at py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:133)\n at py4j.commands.CallCommand.execute(CallCommand.java:79)\n at py4j.GatewayConnection.run(GatewayConnection.java:207)\n at java.lang.Thread.run(Thread.java:745)\n{code}","from":"reporter","subject":"Spark Data Frame Fails to Load Postgres Tables with JSONB DataType Columns"},{"body":"This is also a problem for {{json}} columns, not just {{jsonb}} ones. \n\nIt would be nice to get the JSON as a String column, instead of an error.\n","from":"developer"},{"body":"User '0x0FFF' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/8948","from":"developer"},{"body":"Still not resolves in spark version 1.6. I am seeing the same issue in spark","from":"developer"},{"body":"Hi Nipun, I am using Spark. Is there a way to insert the Jsonb data into postgres. We have a new project in design phase. We are thinking of using Apache Spark + Postgres DB. But we are facing issues while inserting JSONB data type. \nIs there a support for Postgres-JSONB from spark? Can you please help us ? I have posted this question in the issues but no response. We really need help, can you please let us know if there is a way of inserting?? ","from":"developer"}],"created":"2015-05-26T13:45:12.000+0000","description":"Most of our tables load into dataframes just fine with postgres. However we have a number of tables leveraging the JSONB datatype. Spark will error and refuse to load this table. While asking for Spark to support JSONB might be a tall order in the short term, it would be great if Spark would at least load the table ignoring the columns it can't load or have it be an option.\n{code}\npdf = sql_context.load(source=\"jdbc\", url=url, dbtable=\"table_of_json\")\n\nPy4JJavaError: An error occurred while calling o41.load.\n: java.sql.SQLException: Unsupported type 1111\n at org.apache.spark.sql.jdbc.JDBCRDD$.getCatalystType(JDBCRDD.scala:78)\n at org.apache.spark.sql.jdbc.JDBCRDD$.resolveTable(JDBCRDD.scala:112)\n at org.apache.spark.sql.jdbc.JDBCRelation.(JDBCRelation.scala:133)\n at org.apache.spark.sql.jdbc.DefaultSource.createRelation(JDBCRelation.scala:121)\n at org.apache.spark.sql.sources.ResolvedDataSource$.apply(ddl.scala:219)\n at org.apache.spark.sql.SQLContext.load(SQLContext.scala:697)\n at org.apache.spark.sql.SQLContext.load(SQLContext.scala:685)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:606)\n at py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:231)\n at py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:379)\n at py4j.Gateway.invoke(Gateway.java:259)\n at py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:133)\n at py4j.commands.CallCommand.execute(CallCommand.java:79)\n at py4j.GatewayConnection.run(GatewayConnection.java:207)\n at java.lang.Thread.run(Thread.java:745)\n{code}","issue_id":"12832720","key":"SPARK-7869","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-10-08T06:12:40.000+0000","role":"fixed_distractor","summary":"Spark Data Frame Fails to Load Postgres Tables with JSONB DataType Columns"} {"case_id":"12838410","cluster":"DISTRACTOR-SPARK-8406","comments":[{"body":"This is hitting us hard. Let me know if there is anything we can do to help on this end with contributing a fix or testing. \n\nFYI heres details from the mailing list. \n\n ---\n\nWhen trying to save a data frame with 569610608 rows. \n\n dfc.write.format(\"parquet\").save(“/data/map_parquet_file\")\n\nWe get random results between runs. Caching the data frame in memory makes no difference. It looks like the write out misses some of the RDD partitions. We have an RDD with 6750 partitions. When we write out we get less files out than the number of partitions. When reading the data back in and running a count, we get smaller number of rows. \n\nI’ve tried counting the rows in all different ways. All return the same result, 560214031 rows, missing about 9.4 million rows (0.15%).\n\n qc.read.parquet(\"/data/map_parquet_file\").count\n qc.read.parquet(\"/data/map_parquet_file\").rdd.count\n qc.read.parquet(\"/data/map_parquet_file\").mapPartitions{itr => var c = 0; itr.foreach(_ => c = c + 1); Seq(c).toIterator }.reduce(_ + _)\n\nLooking on HDFS the files, there are 6643 .parquet files. 107 missing partitions (about 0.15%). \n\nThen writing out the same cached DF again to a new file gives 6717 files on hdfs (about 33 files missing or 0.5%);\n\n dfc.write.parquet(“/data/map_parquet_file_2\")\n\nAnd we get 566670107 rows back (about 3million missing ~0.5%); \n\n qc.read.parquet(\"/data/map_parquet_file_2\").count\n\nWriting the same df out to json writes the expected number (6750) of parquet files and returns the right number of rows 569610608. \n\n dfc.write.format(\"json\").save(\"/data/map_parquet_file_3\")\n qc.read.format(\"json\").load(\"/data/map_parquet_file_3\").count\n\nOne thing to note is that the parquet part files on HDFS are not the normal sequential part numbers like for the json output and parquet output in Spark 1.3.\n\npart-r-06151.gz.parquet part-r-118401.gz.parquet part-r-146249.gz.parquet part-r-196755.gz.parquet part-r-35811.gz.parquet part-r-55628.gz.parquet part-r-73497.gz.parquet part-r-97237.gz.parquet\npart-r-06161.gz.parquet part-r-118406.gz.parquet part-r-146254.gz.parquet part-r-196763.gz.parquet part-r-35826.gz.parquet part-r-55647.gz.parquet part-r-73500.gz.parquet _SUCCESS\n\nWe are using MapR 4.0.2 for hdfs.","created":"2015-06-17T08:43:56.154+0000"},{"body":"An example task execution order which causes overwriting:\n\n# Writing a DataFrame with 4 RDD partitions to an empty directory.\n# Task 1 and task 2 get scheduled, while task 3 and task 4 are queued. Both task 1 and task 2 find current max part number to be 0 (because destination directory is empty).\n# Task 1 finishes, generates {{part-r-00001.gz.parquet}}. Current max part number becomes 1.\n# Task 4 gets scheduled, decides to write to {{part-r-00005.gz.parquet}} (5 = current max part number + task ID), but hasn't start writing the file yet.\n# Task 2 finishes, generates {{part-r-00002.gz.parquet}}. Current max part number becomes 2.\n# Task 3 gets scheduled, also decides to write to {{part-r-00005.gz.parquet}} since task 4 hasn't start writing its output file, and task 3 finds current max part number is still 2.\n# Task 4 finishes writing {{part-r-00005.gz.parquet}}\n# Task 3 finishes writing {{part-r-00005.gz.parquet}}\n# Output of task 4 is overwritten.","created":"2015-06-17T10:23:05.044+0000"},{"body":"It seems to me that ORC is not free of this bug, but instead just more likely to avoid a problem, right?","created":"2015-06-17T18:48:43.000+0000"},{"body":"Yeah, just updated the JIRA description. ORC may hit this issue only when two tasks with the same task ID (which means they are in two concurrent jobs) are writing to the same location within the same millisecond.","created":"2015-06-17T20:16:47.130+0000"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6864","created":"2015-06-17T22:43:03.687+0000"},{"body":"[~nemccarthy], thanks again for the report. [Here|https://github.com/apache/spark/pull/6864#issuecomment-113024897] is a summary for better understanding of this issue.","created":"2015-06-18T16:55:35.030+0000"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6932","created":"2015-06-22T08:04:05.614+0000"},{"body":"Issue resolved by pull request 6864\n[https://github.com/apache/spark/pull/6864]","created":"2015-06-22T17:04:14.331+0000"},{"body":"[~nemccarthy] I have merged the fix to both master and branch 1.4.","created":"2015-06-22T17:06:44.272+0000"}],"conversations":[{"body":"To support appending, the Parquet data source tries to find out the max part number of part-files in the destination directory (the in output file name \"part-r-.gz.parquet\") at the beginning of the write job. In 1.3.0, this step happens on driver side before any files are written. However, in 1.4.0, this is moved to task side. Thus, for tasks scheduled later, they may see wrong max part number generated by newly written files by other finished tasks within the same job. This actually causes a race condition. In most cases, this only causes nonconsecutive IDs in output file names. But when the DataFrame contains thousands of RDD partitions, it's likely that two tasks may choose the same part number, thus one of them gets overwritten by the other.\n\nThe following Spark shell snippet can reproduce nonconsecutive part numbers:\n{code}\nsqlContext.range(0, 128).repartition(16).write.mode(\"overwrite\").parquet(\"foo\")\n{code}\n\"16\" can be replaced with any integer that is greater than the default parallelism on your machine (usually it means core number, on my machine it's 8).\n{noformat}\n-rw-r--r-- 3 lian supergroup 0 2015-06-17 00:06 /user/lian/foo/_SUCCESS\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00001.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00002.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00003.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00004.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00005.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00006.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00007.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00008.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00017.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00018.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00019.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00020.gz.parquet\n-rw-r--r-- 3 lian supergroup 352 2015-06-17 00:06 /user/lian/foo/part-r-00021.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00022.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00023.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00024.gz.parquet\n{noformat}\n\nAnd here is another Spark shell snippet for reproducing overwriting:\n{code}\nsqlContext.range(0, 10000).repartition(500).write.mode(\"overwrite\").parquet(\"foo\")\nsqlContext.read.parquet(\"foo\").count()\n{code}\nExpected answer should be {{10000}}, but you may see a number like {{9960}} due to overwriting. The actual number varies for different runs and different nodes.\n\nNotice that the newly added ORC data source is less likely to hit this issue because it uses task ID and {{System.currentTimeMills()}} to generate the output file name. Thus, the ORC data source may hit this issue only when two tasks with the same task ID (which means they are in two concurrent jobs) are writing to the same location within the same millisecond.","from":"reporter","subject":"Race condition when writing Parquet files"},{"body":"This is hitting us hard. Let me know if there is anything we can do to help on this end with contributing a fix or testing. \n\nFYI heres details from the mailing list. \n\n ---\n\nWhen trying to save a data frame with 569610608 rows. \n\n dfc.write.format(\"parquet\").save(“/data/map_parquet_file\")\n\nWe get random results between runs. Caching the data frame in memory makes no difference. It looks like the write out misses some of the RDD partitions. We have an RDD with 6750 partitions. When we write out we get less files out than the number of partitions. When reading the data back in and running a count, we get smaller number of rows. \n\nI’ve tried counting the rows in all different ways. All return the same result, 560214031 rows, missing about 9.4 million rows (0.15%).\n\n qc.read.parquet(\"/data/map_parquet_file\").count\n qc.read.parquet(\"/data/map_parquet_file\").rdd.count\n qc.read.parquet(\"/data/map_parquet_file\").mapPartitions{itr => var c = 0; itr.foreach(_ => c = c + 1); Seq(c).toIterator }.reduce(_ + _)\n\nLooking on HDFS the files, there are 6643 .parquet files. 107 missing partitions (about 0.15%). \n\nThen writing out the same cached DF again to a new file gives 6717 files on hdfs (about 33 files missing or 0.5%);\n\n dfc.write.parquet(“/data/map_parquet_file_2\")\n\nAnd we get 566670107 rows back (about 3million missing ~0.5%); \n\n qc.read.parquet(\"/data/map_parquet_file_2\").count\n\nWriting the same df out to json writes the expected number (6750) of parquet files and returns the right number of rows 569610608. \n\n dfc.write.format(\"json\").save(\"/data/map_parquet_file_3\")\n qc.read.format(\"json\").load(\"/data/map_parquet_file_3\").count\n\nOne thing to note is that the parquet part files on HDFS are not the normal sequential part numbers like for the json output and parquet output in Spark 1.3.\n\npart-r-06151.gz.parquet part-r-118401.gz.parquet part-r-146249.gz.parquet part-r-196755.gz.parquet part-r-35811.gz.parquet part-r-55628.gz.parquet part-r-73497.gz.parquet part-r-97237.gz.parquet\npart-r-06161.gz.parquet part-r-118406.gz.parquet part-r-146254.gz.parquet part-r-196763.gz.parquet part-r-35826.gz.parquet part-r-55647.gz.parquet part-r-73500.gz.parquet _SUCCESS\n\nWe are using MapR 4.0.2 for hdfs.","from":"developer"},{"body":"An example task execution order which causes overwriting:\n\n# Writing a DataFrame with 4 RDD partitions to an empty directory.\n# Task 1 and task 2 get scheduled, while task 3 and task 4 are queued. Both task 1 and task 2 find current max part number to be 0 (because destination directory is empty).\n# Task 1 finishes, generates {{part-r-00001.gz.parquet}}. Current max part number becomes 1.\n# Task 4 gets scheduled, decides to write to {{part-r-00005.gz.parquet}} (5 = current max part number + task ID), but hasn't start writing the file yet.\n# Task 2 finishes, generates {{part-r-00002.gz.parquet}}. Current max part number becomes 2.\n# Task 3 gets scheduled, also decides to write to {{part-r-00005.gz.parquet}} since task 4 hasn't start writing its output file, and task 3 finds current max part number is still 2.\n# Task 4 finishes writing {{part-r-00005.gz.parquet}}\n# Task 3 finishes writing {{part-r-00005.gz.parquet}}\n# Output of task 4 is overwritten.","from":"developer"},{"body":"It seems to me that ORC is not free of this bug, but instead just more likely to avoid a problem, right?","from":"developer"},{"body":"Yeah, just updated the JIRA description. ORC may hit this issue only when two tasks with the same task ID (which means they are in two concurrent jobs) are writing to the same location within the same millisecond.","from":"developer"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6864","from":"developer"},{"body":"[~nemccarthy], thanks again for the report. [Here|https://github.com/apache/spark/pull/6864#issuecomment-113024897] is a summary for better understanding of this issue.","from":"developer"},{"body":"User 'liancheng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/6932","from":"developer"},{"body":"Issue resolved by pull request 6864\n[https://github.com/apache/spark/pull/6864]","from":"developer"},{"body":"[~nemccarthy] I have merged the fix to both master and branch 1.4.","from":"developer"}],"created":"2015-06-17T08:20:35.000+0000","description":"To support appending, the Parquet data source tries to find out the max part number of part-files in the destination directory (the in output file name \"part-r-.gz.parquet\") at the beginning of the write job. In 1.3.0, this step happens on driver side before any files are written. However, in 1.4.0, this is moved to task side. Thus, for tasks scheduled later, they may see wrong max part number generated by newly written files by other finished tasks within the same job. This actually causes a race condition. In most cases, this only causes nonconsecutive IDs in output file names. But when the DataFrame contains thousands of RDD partitions, it's likely that two tasks may choose the same part number, thus one of them gets overwritten by the other.\n\nThe following Spark shell snippet can reproduce nonconsecutive part numbers:\n{code}\nsqlContext.range(0, 128).repartition(16).write.mode(\"overwrite\").parquet(\"foo\")\n{code}\n\"16\" can be replaced with any integer that is greater than the default parallelism on your machine (usually it means core number, on my machine it's 8).\n{noformat}\n-rw-r--r-- 3 lian supergroup 0 2015-06-17 00:06 /user/lian/foo/_SUCCESS\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00001.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00002.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00003.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00004.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00005.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00006.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00007.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00008.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00017.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00018.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00019.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00020.gz.parquet\n-rw-r--r-- 3 lian supergroup 352 2015-06-17 00:06 /user/lian/foo/part-r-00021.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00022.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00023.gz.parquet\n-rw-r--r-- 3 lian supergroup 353 2015-06-17 00:06 /user/lian/foo/part-r-00024.gz.parquet\n{noformat}\n\nAnd here is another Spark shell snippet for reproducing overwriting:\n{code}\nsqlContext.range(0, 10000).repartition(500).write.mode(\"overwrite\").parquet(\"foo\")\nsqlContext.read.parquet(\"foo\").count()\n{code}\nExpected answer should be {{10000}}, but you may see a number like {{9960}} due to overwriting. The actual number varies for different runs and different nodes.\n\nNotice that the newly added ORC data source is less likely to hit this issue because it uses task ID and {{System.currentTimeMills()}} to generate the output file name. Thus, the ORC data source may hit this issue only when two tasks with the same task ID (which means they are in two concurrent jobs) are writing to the same location within the same millisecond.","issue_id":"12838410","key":"SPARK-8406","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-06-22T17:04:14.000+0000","role":"fixed_distractor","summary":"Race condition when writing Parquet files"} {"case_id":"12845820","cluster":"DISTRACTOR-SPARK-9131","comments":[{"body":"[~lfag] [~luispeguerra] would it be possible to get a reproducible case with a small dataset?","created":"2015-07-21T02:05:55.320+0000"},{"body":"It seems the odd result depends on the size of the dataset. I am trying to reduce the data to get a small dataset, but output is ok with around 650.000 cases. However, looking at the same registers, they become odd when the dataset grows to more than one million (roughly speaking).","created":"2015-07-21T10:17:48.064+0000"},{"body":"I see. Even if we have a relatively large dataset, as long as we can reproduce it, it'd be great to have.\n","created":"2015-07-21T19:38:50.156+0000"},{"body":"Agree, I have prepared a dataset.json and is ready to be uploaded. However, its size is too large (more than 600 mb). How can I upload it for you? ","created":"2015-07-22T08:20:07.395+0000"},{"body":"Actually, the file is reduced by compression to 27 mb, closer to the limit but still over it","created":"2015-07-22T08:22:03.485+0000"},{"body":"Would it be possible to upload it to s3 or github?\n","created":"2015-07-22T08:28:31.512+0000"},{"body":"I hope they work fine. I have split them into several files to reach the limit size","created":"2015-07-22T08:29:00.467+0000"},{"body":"By the way, UDFs code should be changed to return StringType(). I changed the data type to string since It does not matter the data type but using the UDFs","created":"2015-07-22T08:34:39.044+0000"},{"body":"I think this maybe fixed by https://github.com/apache/spark/pull/7131\n\n[~luispeguerra] Could you help to confirm that its' fixed in master or not?","created":"2015-08-03T22:42:10.624+0000"},{"body":"Going to close this since it's most likely fixed.\n\n[~lfag] [~luispeguerra] can you try it on branch-1.5? If it doesn't work, we should reopen this.\n","created":"2015-08-05T05:51:47.529+0000"},{"body":"I just ran into something that sounds similar to this that affects 1.3.0, 1.4.1, and 1.5.0:\n\n- https://issues.apache.org/jira/browse/SPARK-10685","created":"2015-09-18T02:52:15.691+0000"}],"conversations":[{"body":"I am having some troubles when using a custom udf in dataframes with pyspark 1.4.\n\nI have rewritten the udf to simplify the problem and it gets even weirder. The udfs I am using do absolutely nothing, they just receive some value and output the same value with the same format.\n\nI show you my code below:\n{code}\nc= a.join(b, a['ID'] == b['ID_new'], 'inner')\n\nc.filter(c['ID'] == '6000000002698917').show()\n\nudf_A = UserDefinedFunction(lambda x: x, DateType())\nudf_B = UserDefinedFunction(lambda x: x, DateType())\nudf_C = UserDefinedFunction(lambda x: x, DateType())\n\nd = c.select(c['ID'], c['t1'].alias('ta'), udf_A(vinc_muestra['t2']).alias('tb'), udf_B(vinc_muestra['t1']).alias('tc'), udf_C(vinc_muestra['t2']).alias('td'))\n\nd.filter(d['ID'] == '6000000002698917').show()\n{code}\n\nI am showing here the results from the outputs:\n{code}\n+----------------+----------------+----------+----------+\n| ID | ID_new | t1\t | t2 |\n+----------------+----------------+----------+----------+\n|6000000002698917| 6000000002698917| 2012-02-28| 2014-02-28|\n|6000000002698917| 6000000002698917| 2012-02-20| 2013-02-20|\n|6000000002698917| 6000000002698917| 2012-02-28| 2014-02-28|\n|6000000002698917| 6000000002698917| 2012-02-20| 2013-02-20|\n|6000000002698917| 6000000002698917| 2012-02-20| 2013-02-20|\n|6000000002698917| 6000000002698917| 2012-02-28| 2014-02-28|\n|6000000002698917| 6000000002698917| 2012-02-28| 2014-02-28|\n|6000000002698917| 6000000002698917| 2012-02-20| 2013-02-20|\n+----------------+----------------+----------+----------+\n\n+----------------+---------------+---------------+------------+------------+\n| ID |\t ta\t |\t tb\t |\t tc\t | td\t |\n+----------------+---------------+---------------+------------+------------+\n|6000000002698917| 2012-02-28| 2007-03-05| 2003-03-05| 2014-02-28|\n|6000000002698917| 2012-02-20| 2007-02-15| 2002-02-15| 2013-02-20|\n|6000000002698917| 2012-02-28| 2007-03-10| 2005-03-10| 2014-02-28|\n|6000000002698917| 2012-02-20| 2007-03-05| 2003-03-05| 2013-02-20|\n|6000000002698917| 2012-02-20| 2013-08-02| 2013-01-02| 2013-02-20|\n|6000000002698917| 2012-02-28| 2007-02-15| 2002-02-15| 2014-02-28|\n|6000000002698917| 2012-02-28| 2007-02-15| 2002-02-15| 2014-02-28|\n|6000000002698917| 2012-02-20| 2014-01-02| 2013-01-02| 2013-02-20|\n+----------------+---------------+---------------+------------+------------+\n{code}\nThe problem here is that values at columns 'tb', 'tc' and 'td' in dataframe 'd' are completely different from values 't1' and 't2' in dataframe c even when my udfs are doing nothing. It seems like if values were somehow got from other registers (or just invented). Results are different between executions (apparently random).\n\nThanks in advance","from":"reporter","subject":"Python UDFs change data values"},{"body":"[~lfag] [~luispeguerra] would it be possible to get a reproducible case with a small dataset?","from":"developer"},{"body":"It seems the odd result depends on the size of the dataset. I am trying to reduce the data to get a small dataset, but output is ok with around 650.000 cases. However, looking at the same registers, they become odd when the dataset grows to more than one million (roughly speaking).","from":"developer"},{"body":"I see. Even if we have a relatively large dataset, as long as we can reproduce it, it'd be great to have.\n","from":"developer"},{"body":"Agree, I have prepared a dataset.json and is ready to be uploaded. However, its size is too large (more than 600 mb). How can I upload it for you? ","from":"developer"},{"body":"Actually, the file is reduced by compression to 27 mb, closer to the limit but still over it","from":"developer"},{"body":"Would it be possible to upload it to s3 or github?\n","from":"developer"},{"body":"I hope they work fine. I have split them into several files to reach the limit size","from":"developer"},{"body":"By the way, UDFs code should be changed to return StringType(). I changed the data type to string since It does not matter the data type but using the UDFs","from":"developer"},{"body":"I think this maybe fixed by https://github.com/apache/spark/pull/7131\n\n[~luispeguerra] Could you help to confirm that its' fixed in master or not?","from":"developer"},{"body":"Going to close this since it's most likely fixed.\n\n[~lfag] [~luispeguerra] can you try it on branch-1.5? If it doesn't work, we should reopen this.\n","from":"developer"},{"body":"I just ran into something that sounds similar to this that affects 1.3.0, 1.4.1, and 1.5.0:\n\n- https://issues.apache.org/jira/browse/SPARK-10685","from":"developer"}],"created":"2015-07-17T07:41:06.000+0000","description":"I am having some troubles when using a custom udf in dataframes with pyspark 1.4.\n\nI have rewritten the udf to simplify the problem and it gets even weirder. The udfs I am using do absolutely nothing, they just receive some value and output the same value with the same format.\n\nI show you my code below:\n{code}\nc= a.join(b, a['ID'] == b['ID_new'], 'inner')\n\nc.filter(c['ID'] == '6000000002698917').show()\n\nudf_A = UserDefinedFunction(lambda x: x, DateType())\nudf_B = UserDefinedFunction(lambda x: x, DateType())\nudf_C = UserDefinedFunction(lambda x: x, DateType())\n\nd = c.select(c['ID'], c['t1'].alias('ta'), udf_A(vinc_muestra['t2']).alias('tb'), udf_B(vinc_muestra['t1']).alias('tc'), udf_C(vinc_muestra['t2']).alias('td'))\n\nd.filter(d['ID'] == '6000000002698917').show()\n{code}\n\nI am showing here the results from the outputs:\n{code}\n+----------------+----------------+----------+----------+\n| ID | ID_new | t1\t | t2 |\n+----------------+----------------+----------+----------+\n|6000000002698917| 6000000002698917| 2012-02-28| 2014-02-28|\n|6000000002698917| 6000000002698917| 2012-02-20| 2013-02-20|\n|6000000002698917| 6000000002698917| 2012-02-28| 2014-02-28|\n|6000000002698917| 6000000002698917| 2012-02-20| 2013-02-20|\n|6000000002698917| 6000000002698917| 2012-02-20| 2013-02-20|\n|6000000002698917| 6000000002698917| 2012-02-28| 2014-02-28|\n|6000000002698917| 6000000002698917| 2012-02-28| 2014-02-28|\n|6000000002698917| 6000000002698917| 2012-02-20| 2013-02-20|\n+----------------+----------------+----------+----------+\n\n+----------------+---------------+---------------+------------+------------+\n| ID |\t ta\t |\t tb\t |\t tc\t | td\t |\n+----------------+---------------+---------------+------------+------------+\n|6000000002698917| 2012-02-28| 2007-03-05| 2003-03-05| 2014-02-28|\n|6000000002698917| 2012-02-20| 2007-02-15| 2002-02-15| 2013-02-20|\n|6000000002698917| 2012-02-28| 2007-03-10| 2005-03-10| 2014-02-28|\n|6000000002698917| 2012-02-20| 2007-03-05| 2003-03-05| 2013-02-20|\n|6000000002698917| 2012-02-20| 2013-08-02| 2013-01-02| 2013-02-20|\n|6000000002698917| 2012-02-28| 2007-02-15| 2002-02-15| 2014-02-28|\n|6000000002698917| 2012-02-28| 2007-02-15| 2002-02-15| 2014-02-28|\n|6000000002698917| 2012-02-20| 2014-01-02| 2013-01-02| 2013-02-20|\n+----------------+---------------+---------------+------------+------------+\n{code}\nThe problem here is that values at columns 'tb', 'tc' and 'td' in dataframe 'd' are completely different from values 't1' and 't2' in dataframe c even when my udfs are doing nothing. It seems like if values were somehow got from other registers (or just invented). Results are different between executions (apparently random).\n\nThanks in advance","issue_id":"12845820","key":"SPARK-9131","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-08-05T05:51:47.000+0000","role":"fixed_distractor","summary":"Python UDFs change data values"} {"case_id":"12849097","cluster":"DISTRACTOR-SPARK-9342","comments":[{"body":"This should be resolved since Spark 2.0. Please check it. If you still hit the issue, reopen it. Thanks!","created":"2016-10-08T04:37:23.813+0000"}],"conversations":[{"body":"The Spark SQL documentation's section on Hive support claims that views are supported. However, even basic view operations fail with exceptions related to column resolution. \n\nFor example,\n\n{code}\n// The test table has columns category & num\nctx.sql(\"create view view1 as select * from test\")\nctx.table(\"view1\").printSchema\n{code}\n\ngenerates\n\n{code}\norg.apache.spark.sql.AnalysisException: cannot resolve 'test.col' given input columns category, num; line 1 pos 7\n\tat org.apache.spark.sql.catalyst.analysis.package$AnalysisErrorAt.failAnalysis(package.scala:42)\n ...\n{code}\n\nYou can see a standalone reproducible example with full spark-shell output demonstrating the problem at [https://gist.github.com/ssimeonov/57164f9d6b928ba0cfde]\n\nThe problem is that {{ctx.sql(\"create view view1 as select * from test\")}} puts the following in the metastore including {{cols:[FieldSchema(name:col, type:string, comment:null)]}} even though the {{test}} table has {{category}} and {{num}} columns:\n\n{code}\n15/07/26 15:47:28 INFO HiveMetaStore: 0: create_table: Table(tableName:view1, dbName:default, owner:ubuntu, createTime:1437925648, lastAccessTime:0, retention:0, sd:StorageDescriptor(cols:[FieldSchema(name:col, type:string, comment:null)], location:null, inputFormat:org.apache.hadoop.mapred.SequenceFileInputFormat, outputFormat:org.apache.hadoop.hive.ql.io.HiveSequenceFileOutputFormat, compressed:false, numBuckets:-1, serdeInfo:SerDeInfo(name:null, serializationLib:null, parameters:{}), bucketCols:[], sortCols:[], parameters:{}, skewedInfo:SkewedInfo(skewedColNames:[], skewedColValues:[], skewedColValueLocationMaps:{})), partitionKeys:[], parameters:{}, viewOriginalText:select * from test, viewExpandedText:select `test`.`col` from `default`.`test`, tableType:VIRTUAL_VIEW)\n15/07/26 15:47:28 INFO audit: ugi=ubuntu\tip=unknown-ip-addr\tcmd=create_table: Table(tableName:view1, dbName:default, owner:ubuntu, createTime:1437925648, lastAccessTime:0, retention:0, sd:StorageDescriptor(cols:[FieldSchema(name:col, type:string, comment:null)], location:null, inputFormat:org.apache.hadoop.mapred.SequenceFileInputFormat, outputFormat:org.apache.hadoop.hive.ql.io.HiveSequenceFileOutputFormat, compressed:false, numBuckets:-1, serdeInfo:SerDeInfo(name:null, serializationLib:null, parameters:{}), bucketCols:[], sortCols:[], parameters:{}, skewedInfo:SkewedInfo(skewedColNames:[], skewedColValues:[], skewedColValueLocationMaps:{})), partitionKeys:[], parameters:{}, viewOriginalText:select * from test, viewExpandedText:select `test`.`col` from `default`.`test`, tableType:VIRTUAL_VIEW)\n{code}","from":"reporter","subject":"Spark SQL views don't work"},{"body":"This should be resolved since Spark 2.0. Please check it. If you still hit the issue, reopen it. Thanks!","from":"developer"}],"created":"2015-07-25T15:07:01.000+0000","description":"The Spark SQL documentation's section on Hive support claims that views are supported. However, even basic view operations fail with exceptions related to column resolution. \n\nFor example,\n\n{code}\n// The test table has columns category & num\nctx.sql(\"create view view1 as select * from test\")\nctx.table(\"view1\").printSchema\n{code}\n\ngenerates\n\n{code}\norg.apache.spark.sql.AnalysisException: cannot resolve 'test.col' given input columns category, num; line 1 pos 7\n\tat org.apache.spark.sql.catalyst.analysis.package$AnalysisErrorAt.failAnalysis(package.scala:42)\n ...\n{code}\n\nYou can see a standalone reproducible example with full spark-shell output demonstrating the problem at [https://gist.github.com/ssimeonov/57164f9d6b928ba0cfde]\n\nThe problem is that {{ctx.sql(\"create view view1 as select * from test\")}} puts the following in the metastore including {{cols:[FieldSchema(name:col, type:string, comment:null)]}} even though the {{test}} table has {{category}} and {{num}} columns:\n\n{code}\n15/07/26 15:47:28 INFO HiveMetaStore: 0: create_table: Table(tableName:view1, dbName:default, owner:ubuntu, createTime:1437925648, lastAccessTime:0, retention:0, sd:StorageDescriptor(cols:[FieldSchema(name:col, type:string, comment:null)], location:null, inputFormat:org.apache.hadoop.mapred.SequenceFileInputFormat, outputFormat:org.apache.hadoop.hive.ql.io.HiveSequenceFileOutputFormat, compressed:false, numBuckets:-1, serdeInfo:SerDeInfo(name:null, serializationLib:null, parameters:{}), bucketCols:[], sortCols:[], parameters:{}, skewedInfo:SkewedInfo(skewedColNames:[], skewedColValues:[], skewedColValueLocationMaps:{})), partitionKeys:[], parameters:{}, viewOriginalText:select * from test, viewExpandedText:select `test`.`col` from `default`.`test`, tableType:VIRTUAL_VIEW)\n15/07/26 15:47:28 INFO audit: ugi=ubuntu\tip=unknown-ip-addr\tcmd=create_table: Table(tableName:view1, dbName:default, owner:ubuntu, createTime:1437925648, lastAccessTime:0, retention:0, sd:StorageDescriptor(cols:[FieldSchema(name:col, type:string, comment:null)], location:null, inputFormat:org.apache.hadoop.mapred.SequenceFileInputFormat, outputFormat:org.apache.hadoop.hive.ql.io.HiveSequenceFileOutputFormat, compressed:false, numBuckets:-1, serdeInfo:SerDeInfo(name:null, serializationLib:null, parameters:{}), bucketCols:[], sortCols:[], parameters:{}, skewedInfo:SkewedInfo(skewedColNames:[], skewedColValues:[], skewedColValueLocationMaps:{})), partitionKeys:[], parameters:{}, viewOriginalText:select * from test, viewExpandedText:select `test`.`col` from `default`.`test`, tableType:VIRTUAL_VIEW)\n{code}","issue_id":"12849097","key":"SPARK-9342","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-10-08T04:37:39.000+0000","role":"fixed_distractor","summary":"Spark SQL views don't work"} {"case_id":"12849916","cluster":"DISTRACTOR-SPARK-9435","comments":[{"body":"Attaching minimal reproduction code.","created":"2015-07-29T12:27:59.262+0000"},{"body":"From a quick glance, the problem is likely that the {{equals}} function on Java UDFs is working correctly. As a workaround you could probably calculate the udf in a nested select.","created":"2015-09-08T19:03:12.294+0000"},{"body":"Same error for me:\n\n{code:java}\n // Register computeDecade() as a SparkSQL function\n sqlContext.udf().register(\"computeDecade\", (Integer year) -> computeDecade(year), DataTypes.StringType);\n\n final List albums = Arrays.asList(new Album(2000, \"1\"), new Album(2000, \"2\"), new Album(2000, \"3\"));\n\n final JavaRDD rdd = javaSc.parallelize(albums);\n final DataFrame df = sqlContext.createDataFrame(rdd, Album.class);\n df.registerTempTable(\"albums\");\n\n final DataFrame dataFrame = sqlContext.sql(\"SELECT computeDecade(year),count(title) \"+\n \" FROM albums \" +\n \" GROUP BY computeDecade(year)\");\n\n{code}","created":"2015-11-12T09:02:44.451+0000"},{"body":"Work-around: *define the UDF using the Scala API instead*\n\n{code:java}\n public static final class ComputeDecadeFn extends AbstractFunction1 implements Serializable {\n @Override\n public String apply(Integer year) {\n return computeDecade(year);\n }\n }\n sqlContext.udf().register(\"computeDecade\", new ComputeDecadeFn(),\n JavaApiHelper.getTypeTag(String.class),\n JavaApiHelper.getTypeTag(Integer.class));\n{code}\n\n You cannot use the lambda expression because the UDF function should be serializable.\n\n The _JavaApiHelper.getTypeTag()_ method comes from *com.datastax.spark.connector.util.JavaApiHelper* here: https://github.com/datastax/spark-cassandra-connector/blob/master/spark-cassandra-connector/src/main/scala/com/datastax/spark/connector/util/JavaApiHelper.scala#L35","created":"2015-11-12T13:57:03.476+0000"},{"body":"This sill happens in the current master - \n\n{code}\nval df = Seq((1, 10), (2, 11), (3, 12)).toDF(\"x\", \"y\")\nval udf = new UDF1[Int, Int] {\n override def call(i: Int): Int = i + 1\n}\n\nspark.udf.register(\"inc\", udf, IntegerType)\ndf.createOrReplaceTempView(\"tmp\")\nspark.sql(\"SELECT inc(y) FROM tmp GROUP BY inc(y)\").show()\n{code}\n\nI tested both Scala and Java ones. and I believe the above one is simpler Scala one to reproduce the same issue.","created":"2017-01-10T04:24:15.035+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16553","created":"2017-01-11T20:46:04.625+0000"},{"body":"Issue resolved by pull request 16553\n[https://github.com/apache/spark/pull/16553]","created":"2017-01-24T06:21:42.567+0000"}],"conversations":[{"body":"If you define a UDF in Java, for example by implementing the UDF1 interface, then try to use that UDF on a column in both the SELECT and GROUP BY clauses of a query, you'll get an error like this:\n\n{code}\n\n\"SELECT inc(y),COUNT(DISTINCT x) FROM test_table GROUP BY inc(y)\"\n\norg.apache.spark.sql.AnalysisException: expression 'y' is neither present in the group by, nor is it an aggregate function. Add to group by or wrap in first() if you don't care which value you get.\n{code}\n\nWe put together a minimal reproduction in the attached Java file, which makes use of the data in the text file attached.\n\nI'm guessing there's some kind of issue with the equality implementation, so Spark can't tell that those two expressions are the same maybe? If you do the same thing from Scala, it works fine.\n\nNote for context: we ran into this issue while working around SPARK-9338.\n","from":"reporter","subject":"Java UDFs don't work with GROUP BY expressions"},{"body":"Attaching minimal reproduction code.","from":"developer"},{"body":"From a quick glance, the problem is likely that the {{equals}} function on Java UDFs is working correctly. As a workaround you could probably calculate the udf in a nested select.","from":"developer"},{"body":"Same error for me:\n\n{code:java}\n // Register computeDecade() as a SparkSQL function\n sqlContext.udf().register(\"computeDecade\", (Integer year) -> computeDecade(year), DataTypes.StringType);\n\n final List albums = Arrays.asList(new Album(2000, \"1\"), new Album(2000, \"2\"), new Album(2000, \"3\"));\n\n final JavaRDD rdd = javaSc.parallelize(albums);\n final DataFrame df = sqlContext.createDataFrame(rdd, Album.class);\n df.registerTempTable(\"albums\");\n\n final DataFrame dataFrame = sqlContext.sql(\"SELECT computeDecade(year),count(title) \"+\n \" FROM albums \" +\n \" GROUP BY computeDecade(year)\");\n\n{code}","from":"developer"},{"body":"Work-around: *define the UDF using the Scala API instead*\n\n{code:java}\n public static final class ComputeDecadeFn extends AbstractFunction1 implements Serializable {\n @Override\n public String apply(Integer year) {\n return computeDecade(year);\n }\n }\n sqlContext.udf().register(\"computeDecade\", new ComputeDecadeFn(),\n JavaApiHelper.getTypeTag(String.class),\n JavaApiHelper.getTypeTag(Integer.class));\n{code}\n\n You cannot use the lambda expression because the UDF function should be serializable.\n\n The _JavaApiHelper.getTypeTag()_ method comes from *com.datastax.spark.connector.util.JavaApiHelper* here: https://github.com/datastax/spark-cassandra-connector/blob/master/spark-cassandra-connector/src/main/scala/com/datastax/spark/connector/util/JavaApiHelper.scala#L35","from":"developer"},{"body":"This sill happens in the current master - \n\n{code}\nval df = Seq((1, 10), (2, 11), (3, 12)).toDF(\"x\", \"y\")\nval udf = new UDF1[Int, Int] {\n override def call(i: Int): Int = i + 1\n}\n\nspark.udf.register(\"inc\", udf, IntegerType)\ndf.createOrReplaceTempView(\"tmp\")\nspark.sql(\"SELECT inc(y) FROM tmp GROUP BY inc(y)\").show()\n{code}\n\nI tested both Scala and Java ones. and I believe the above one is simpler Scala one to reproduce the same issue.","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16553","from":"developer"},{"body":"Issue resolved by pull request 16553\n[https://github.com/apache/spark/pull/16553]","from":"developer"}],"created":"2015-07-29T12:27:13.000+0000","description":"If you define a UDF in Java, for example by implementing the UDF1 interface, then try to use that UDF on a column in both the SELECT and GROUP BY clauses of a query, you'll get an error like this:\n\n{code}\n\n\"SELECT inc(y),COUNT(DISTINCT x) FROM test_table GROUP BY inc(y)\"\n\norg.apache.spark.sql.AnalysisException: expression 'y' is neither present in the group by, nor is it an aggregate function. Add to group by or wrap in first() if you don't care which value you get.\n{code}\n\nWe put together a minimal reproduction in the attached Java file, which makes use of the data in the text file attached.\n\nI'm guessing there's some kind of issue with the equality implementation, so Spark can't tell that those two expressions are the same maybe? If you do the same thing from Scala, it works fine.\n\nNote for context: we ran into this issue while working around SPARK-9338.\n","issue_id":"12849916","key":"SPARK-9435","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-01-24T06:21:42.000+0000","role":"fixed_distractor","summary":"Java UDFs don't work with GROUP BY expressions"} {"case_id":"12922349","cluster":"JIRA-CASSANDRA-c6b13e0fa922","comments":[{"body":"Just a note on v4 and preferred version: Python driver has supported v4 (and defaulted it) since version 2.6.0 ( released in conjunction with Cassandra 2.2.0 pre-release series in June 2015).","created":"2015-12-16T15:06:10.835+0000"},{"body":"Not sure if we want to prefer v3 for Cassandra 2.2 because then we lose UDF/UDA support in the driver, don't we? For example, the v3 protocol doc doesn't mention how to handle schema change events for FUNCTIONs and AGGREGATEs, while the v4 protocol doc does.","created":"2015-12-16T18:57:36.127+0000"},{"body":"bq. we lose UDF/UDA support in the driver, don't we?\n\nWell, not really. We really only lose the _specific_ schema change events, but changes are still notified (as keyspace update if memory serves) so drivers can still deal with it (albeit in a slightly less efficient manner since they don't know what function has been added/updated and they may have to read more to figure that out, but that is unlikely to have much visible impact to users in practice).\n\nBut let me be clear that I'm not pretending any of the solutions above are great, they're not. And if someone has some idea I haven't think of that is much better, please do share. Short of that, I think \"you shouldn't use the nice-but-not-absolutely-necessary additions of the v4 protocol until 3.0\" is better than \"you cannot upgrade to 3.x\".","created":"2015-12-17T10:32:45.333+0000"},{"body":"After some offline discussions, there was some agreement on going with option 3 above: simply document clearly that the protocol v3 should be used when migrating to 3.X. Unless someone has a strong objection to this or something better to suggest in the next day or so, I'll proceed by adding a clear mention in the {{NEWS}} file (including what you lose by sticking to the protocol v3, which as said above is not a whole lot) and sending a mail to the user list to grab attention on this issue.","created":"2016-01-04T12:23:25.327+0000"},{"body":"As said above, I've now pushed an email to the mailing list about this and added instructions (to use the native protocol v3 during upgrades) to the NEWS file, so closing.","created":"2016-01-07T15:28:36.608+0000"}],"conversations":[{"body":"In CASSANDRA-10254, the paging states generated by 3.0 for the native protocol v4 were made 3.0 specific. This was done because the paging state in pre-3.0 versions contains a serialized cell name, but 3.0 doesn't talk in term of cells internally (at least not the pre-3.0 ones) and so using an old-format cell name when we only have 3.0 nodes is inefficient and inelegant.\n\nUnfortunately that change was made on the assumption than the protocol v4 was 3.0 only but it's not, it ended up being released with 2.2 and that completely slipped my mind. So in practice, you can't properly have a mixed 2.2/3.0 cluster if your driver is using the protocol v4.\n\nAnd unfortunately, I don't think there is an easy way to fix that without breaking something. Concretely, I can see 3 choices:\n# we change 3.0 so that it generates old-format paging states on the v4 protocol. The 2 main downsides are that 1) this breaks 3.0 upgrades if the driver is using the v4 protocol, and at least on the java side the only driver versions that support 3.0 will use v4 by default and 2) we're signing off on having sub-optimal paging state until the protocol v5 ships (probably not too soon).\n# we remove the v4 protocol from 2.2. This means 2.2 will have to use v3 before upgrade at the risk of breaking upgrade. This is also bad, but I'm not sure the driver version using the v4 protocol are quite ready yet (at least the java driver is not GA yet) so if we work with the drivers teams to make sure the v3 protocol gets prefered by default on 2.2 in the GA versions of these driver, this might be somewhat transparent to users.\n# we don't change anything code-wise, but we document clearly that you can't upgrade from 2.2 to 3.0 if your clients use protocol v4 (so we leave upgrade broken if the v4 protocol is used as it is currently). This is not great, but we can work with the drivers teams here again to make sure drivers prefer the v3 version for 2.2 nodes so most people don't notice in practice.\n\nI think I'm leaning towards solution 3). It's not great but at least we break no minor upgrades (neither on 2.2, nor on 3.0) which is probably the most important. We'd basically be just adding a new condition on 2.2->3.0 upgrades. We could additionally make 3.0 node completely refuse v4 connections if they know a 2.2 nodes is in the cluster for extra safety.\n\nPing [~omichallat], [~adutra] and [~aholmber] as you might want to be aware of that ticket.","from":"reporter","subject":"Paging state between 2.2 and 3.0 are incompatible on protocol v4"},{"body":"Just a note on v4 and preferred version: Python driver has supported v4 (and defaulted it) since version 2.6.0 ( released in conjunction with Cassandra 2.2.0 pre-release series in June 2015).","from":"developer"},{"body":"Not sure if we want to prefer v3 for Cassandra 2.2 because then we lose UDF/UDA support in the driver, don't we? For example, the v3 protocol doc doesn't mention how to handle schema change events for FUNCTIONs and AGGREGATEs, while the v4 protocol doc does.","from":"developer"},{"body":"bq. we lose UDF/UDA support in the driver, don't we?\n\nWell, not really. We really only lose the _specific_ schema change events, but changes are still notified (as keyspace update if memory serves) so drivers can still deal with it (albeit in a slightly less efficient manner since they don't know what function has been added/updated and they may have to read more to figure that out, but that is unlikely to have much visible impact to users in practice).\n\nBut let me be clear that I'm not pretending any of the solutions above are great, they're not. And if someone has some idea I haven't think of that is much better, please do share. Short of that, I think \"you shouldn't use the nice-but-not-absolutely-necessary additions of the v4 protocol until 3.0\" is better than \"you cannot upgrade to 3.x\".","from":"developer"},{"body":"After some offline discussions, there was some agreement on going with option 3 above: simply document clearly that the protocol v3 should be used when migrating to 3.X. Unless someone has a strong objection to this or something better to suggest in the next day or so, I'll proceed by adding a clear mention in the {{NEWS}} file (including what you lose by sticking to the protocol v3, which as said above is not a whole lot) and sending a mail to the user list to grab attention on this issue.","from":"developer"},{"body":"As said above, I've now pushed an email to the mailing list about this and added instructions (to use the native protocol v3 during upgrades) to the NEWS file, so closing.","from":"developer"}],"created":"2015-12-16T11:24:48.000+0000","description":"In CASSANDRA-10254, the paging states generated by 3.0 for the native protocol v4 were made 3.0 specific. This was done because the paging state in pre-3.0 versions contains a serialized cell name, but 3.0 doesn't talk in term of cells internally (at least not the pre-3.0 ones) and so using an old-format cell name when we only have 3.0 nodes is inefficient and inelegant.\n\nUnfortunately that change was made on the assumption than the protocol v4 was 3.0 only but it's not, it ended up being released with 2.2 and that completely slipped my mind. So in practice, you can't properly have a mixed 2.2/3.0 cluster if your driver is using the protocol v4.\n\nAnd unfortunately, I don't think there is an easy way to fix that without breaking something. Concretely, I can see 3 choices:\n# we change 3.0 so that it generates old-format paging states on the v4 protocol. The 2 main downsides are that 1) this breaks 3.0 upgrades if the driver is using the v4 protocol, and at least on the java side the only driver versions that support 3.0 will use v4 by default and 2) we're signing off on having sub-optimal paging state until the protocol v5 ships (probably not too soon).\n# we remove the v4 protocol from 2.2. This means 2.2 will have to use v3 before upgrade at the risk of breaking upgrade. This is also bad, but I'm not sure the driver version using the v4 protocol are quite ready yet (at least the java driver is not GA yet) so if we work with the drivers teams to make sure the v3 protocol gets prefered by default on 2.2 in the GA versions of these driver, this might be somewhat transparent to users.\n# we don't change anything code-wise, but we document clearly that you can't upgrade from 2.2 to 3.0 if your clients use protocol v4 (so we leave upgrade broken if the v4 protocol is used as it is currently). This is not great, but we can work with the drivers teams here again to make sure drivers prefer the v3 version for 2.2 nodes so most people don't notice in practice.\n\nI think I'm leaning towards solution 3). It's not great but at least we break no minor upgrades (neither on 2.2, nor on 3.0) which is probably the most important. We'd basically be just adding a new condition on 2.2->3.0 upgrades. We could additionally make 3.0 node completely refuse v4 connections if they know a 2.2 nodes is in the cluster for extra safety.\n\nPing [~omichallat], [~adutra] and [~aholmber] as you might want to be aware of that ticket.","issue_id":"12922349","key":"CASSANDRA-10880","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-01-07T15:28:36.000+0000","role":"gold_target","summary":"Paging state between 2.2 and 3.0 are incompatible on protocol v4"} {"case_id":"12932787","cluster":"JIRA-CASSANDRA-18d3ddb47d49","comments":[{"body":"patch submitted","created":"2016-01-20T20:02:57.942+0000"},{"body":"patch file submitted","created":"2016-01-20T20:04:12.475+0000"},{"body":"Thanks for your patch [~iryndin]. I think it still needs to be improved, here are some suggestions:\n1) There's no need to create a {{JsonStringEncoder}} static instance and call {{#getInstance()}} on such instance, as it's a static method.\n2) An even better design would be to implement the encoding methods (i.e. {{#quoteAsString()}}) as part of the {{Json}} class, to avoid leaking {{JsonStringEncoder}} instances, which are really an implementation detail.\n3) Finally, I would like to see some unit tests showing that {{Json}} methods are now thread safe. ","created":"2016-01-21T11:17:22.142+0000"},{"body":"Is there an existing muti-threaded test suite for Cassandra which could be extended for this fix? This seems to be more of an integration test in that it would require a live Cassandra instance with predictable test data. I am not familiar with the code but I am interested in reviewing it and understanding how you test multi-threaded code in Cassandra.\n\nI am able to replicate the issue every time (just clicking on a webapp page 4-8 times makes my new application fail).\n\nI wrote this code to replicate and confirmed the patch does fix the issue.\n\nimport com.datastax.driver.core.*;\nimport com.google.gson.JsonObject;\nimport com.google.gson.JsonParser;\n\nimport java.util.Iterator;\nimport java.util.concurrent.ExecutorService;\nimport java.util.concurrent.Executors;\nimport java.util.concurrent.TimeUnit;\n\n// Load some data with: ./tools/bin/cassandra-stress write n=1000 cl=one -mode native cql3 -log file=~/temp/create_schema.log\n\npublic class JsonBug {\n\n public static void main(String [] arguments) {\n\n String node = \"127.0.0.1\";\n\n Cluster cluster = Cluster.builder()\n .addContactPoint(node)\n .build();\n\n Metadata metadata = cluster.getMetadata();\n\n System.out.printf(\"Connected to cluster: %s\\n\",\n metadata.getClusterName());\n\n for ( Host host : metadata.getAllHosts() ) {\n System.out.printf(\"Datacenter: %s; Host: %s; Rack: %s\\n\",\n host.getDatacenter(), host.getAddress(), host.getRack());\n }\n\n final Session session = cluster.connect(\"keyspace1\");\n\n final PreparedStatement prepStatement = session.prepare(\"select JSON * from standard1\");\n\n\n ExecutorService executorService = Executors.newFixedThreadPool(250);\n\n for(int i=0; i<100; i++) {\n\n executorService.submit(new Runnable() {\n\n public void run() {\n\n BoundStatement boundStatement = new BoundStatement(prepStatement);\n\n ResultSet rs = session.execute(boundStatement);\n\n Iterator rsIterator = rs.iterator();\n\n\n JsonParser parser = new JsonParser();\n\n while(rsIterator.hasNext()) {\n\n Row row = rsIterator.next();\n\n String jsonString = row.getString(0);\n\n //System.out.println(jsonString);\n\n JsonObject jsonObj = (JsonObject) parser.parse(jsonString);\n\n if(jsonObj.get(\"key\") == null) System.out.println(\"No key for \" + jsonString + \"\\n\");\n if(jsonObj.get(\"\\\"C0\\\"\") == null) System.out.println(\"No C0 for \" + jsonString + \"\\n\");\n if(jsonObj.get(\"\\\"C1\\\"\") == null) System.out.println(\"No C1 for \" + jsonString + \"\\n\");\n if(jsonObj.get(\"\\\"C2\\\"\") == null) System.out.println(\"No C2 for \" + jsonString + \"\\n\");\n if(jsonObj.get(\"\\\"C3\\\"\") == null) System.out.println(\"No C3 for \" + jsonString + \"\\n\");\n if(jsonObj.get(\"\\\"C4\\\"\") == null) System.out.println(\"No C4for \" + jsonString + \"\\n\");\n }\n }\n });\n\n }\n\n executorService.shutdown();\n\n try {\n executorService.awaitTermination(Long.MAX_VALUE, TimeUnit.NANOSECONDS);\n } catch (InterruptedException e) {\n e.printStackTrace();\n }\n\n cluster.close();\n }\n}\n","created":"2016-01-26T18:17:27.826+0000"},{"body":"[~henryman], the {{Json}} class can definitely be tested in isolation, no need for integration tests in this case.","created":"2016-01-26T20:13:41.293+0000"},{"body":"I agree. Thanks. I am new to the code. The patch from Ivan is pointing me to the places to look at in the code.","created":"2016-02-01T16:44:41.015+0000"},{"body":"We kind of need to fix this asap and I'm a little bit lost of who is actively working on that. So unless someone else come up with a patch updated with Sergio's comments and proper testing in the next few days, can you have a look [~thobbs].","created":"2016-02-05T17:04:04.941+0000"},{"body":"[~sbtourist] I've pushed a set of branches with your suggested changes and a unit test, would you mind reviewing?\n\nBranches and pending CI runs:\n|[CASSANDRA-11048|https://github.com/thobbs/cassandra/tree/CASSANDRA-11048]|[testall|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-dtest]|\n|[CASSANDRA-11048-3.0|https://github.com/thobbs/cassandra/tree/CASSANDRA-11048-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-3.0-dtest]|\n|[CASSANDRA-11048-trunk|https://github.com/thobbs/cassandra/tree/CASSANDRA-11048-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-trunk-dtest]|","created":"2016-02-06T00:07:29.381+0000"},{"body":"[~thobbs],\n\nthe fix looks good, but the test doesn't seem correct to me, for the following reasons:\n* The call to {{fail(exc.getMessage())}} inside the {{Runnable}} doesn't actually make the test fail, because it doesn't propagate to the main thread: you should either wait on the futures, or \"capture\" the exception by yourself.\n* The call to {{executor.awaitTermination(1, TimeUnit.MINUTES)}} should be asserted true, so the test will fail if the other threads hang for any reason.\n\nOther than that, on a more general note, the method might have been tested in isolation via an actual unit test that only verifies the quoted string is not messed up, without having to create tables and rows, but I see the rest of the test class is made like that, so I will not object.","created":"2016-02-08T11:35:42.585+0000"},{"body":"[~sbtourist] thank you for the careful review. I've pushed a second commit to address your comments. Since the previous test runs look good and this commit only changes the test behavior, I haven't started a second round of tests.\n\nbq. Other than that, on a more general note, the method might have been tested in isolation via an actual unit test that only verifies the quoted string is not messed up, without having to create tables and rows, but I see the rest of the test class is made like that, so I will not object.\n\nThe test is still quite fast to execute, and I prefer the extra (future) safety of testing the full stack here, so I went with the integration test approach.","created":"2016-02-09T21:19:02.151+0000"},{"body":"Looks good now, but it seems it wasn't merged upwards?","created":"2016-02-11T11:17:54.544+0000"},{"body":"bq. Looks good now, but it seems it wasn't merged upwards?\n\nThere weren't any merge conflicts, so the CI tests were the only reason to create the other branches. Since I didn't feel the CI tests were necessary for the second commit, I didn't merge into the other branches. Sorry for the confusion there.\n\nAnyway, that sounds like a +1, so I've committed this as {{6b1bd1745ab9f04f9e379e6d22764b97693ba2ad}} in 2.2 and merge upward to 3.0 and trunk. Thanks!","created":"2016-02-11T16:34:42.619+0000"}],"conversations":[{"body":"{{org.apache.cassandra.cql3.Json}} uses a shared instance of {{JsonStringEncoder}} which is not thread safe (see 1), while {{JsonStringEncoder#getInstance()}} should be used (see 2).\n\nAs a consequence, concurrent {{select JSON}} queries often produce wrong (sometimes unreadable) results.\n\n1. http://grepcode.com/file/repo1.maven.org/maven2/org.codehaus.jackson/jackson-core-asl/1.9.2/org/codehaus/jackson/io/JsonStringEncoder.java\n2. http://grepcode.com/file/repo1.maven.org/maven2/org.codehaus.jackson/jackson-core-asl/1.9.2/org/codehaus/jackson/io/JsonStringEncoder.java#JsonStringEncoder.getInstance%28%29","from":"reporter","subject":"JSON queries are not thread safe"},{"body":"patch submitted","from":"developer"},{"body":"patch file submitted","from":"developer"},{"body":"Thanks for your patch [~iryndin]. I think it still needs to be improved, here are some suggestions:\n1) There's no need to create a {{JsonStringEncoder}} static instance and call {{#getInstance()}} on such instance, as it's a static method.\n2) An even better design would be to implement the encoding methods (i.e. {{#quoteAsString()}}) as part of the {{Json}} class, to avoid leaking {{JsonStringEncoder}} instances, which are really an implementation detail.\n3) Finally, I would like to see some unit tests showing that {{Json}} methods are now thread safe. ","from":"developer"},{"body":"Is there an existing muti-threaded test suite for Cassandra which could be extended for this fix? This seems to be more of an integration test in that it would require a live Cassandra instance with predictable test data. I am not familiar with the code but I am interested in reviewing it and understanding how you test multi-threaded code in Cassandra.\n\nI am able to replicate the issue every time (just clicking on a webapp page 4-8 times makes my new application fail).\n\nI wrote this code to replicate and confirmed the patch does fix the issue.\n\nimport com.datastax.driver.core.*;\nimport com.google.gson.JsonObject;\nimport com.google.gson.JsonParser;\n\nimport java.util.Iterator;\nimport java.util.concurrent.ExecutorService;\nimport java.util.concurrent.Executors;\nimport java.util.concurrent.TimeUnit;\n\n// Load some data with: ./tools/bin/cassandra-stress write n=1000 cl=one -mode native cql3 -log file=~/temp/create_schema.log\n\npublic class JsonBug {\n\n public static void main(String [] arguments) {\n\n String node = \"127.0.0.1\";\n\n Cluster cluster = Cluster.builder()\n .addContactPoint(node)\n .build();\n\n Metadata metadata = cluster.getMetadata();\n\n System.out.printf(\"Connected to cluster: %s\\n\",\n metadata.getClusterName());\n\n for ( Host host : metadata.getAllHosts() ) {\n System.out.printf(\"Datacenter: %s; Host: %s; Rack: %s\\n\",\n host.getDatacenter(), host.getAddress(), host.getRack());\n }\n\n final Session session = cluster.connect(\"keyspace1\");\n\n final PreparedStatement prepStatement = session.prepare(\"select JSON * from standard1\");\n\n\n ExecutorService executorService = Executors.newFixedThreadPool(250);\n\n for(int i=0; i<100; i++) {\n\n executorService.submit(new Runnable() {\n\n public void run() {\n\n BoundStatement boundStatement = new BoundStatement(prepStatement);\n\n ResultSet rs = session.execute(boundStatement);\n\n Iterator rsIterator = rs.iterator();\n\n\n JsonParser parser = new JsonParser();\n\n while(rsIterator.hasNext()) {\n\n Row row = rsIterator.next();\n\n String jsonString = row.getString(0);\n\n //System.out.println(jsonString);\n\n JsonObject jsonObj = (JsonObject) parser.parse(jsonString);\n\n if(jsonObj.get(\"key\") == null) System.out.println(\"No key for \" + jsonString + \"\\n\");\n if(jsonObj.get(\"\\\"C0\\\"\") == null) System.out.println(\"No C0 for \" + jsonString + \"\\n\");\n if(jsonObj.get(\"\\\"C1\\\"\") == null) System.out.println(\"No C1 for \" + jsonString + \"\\n\");\n if(jsonObj.get(\"\\\"C2\\\"\") == null) System.out.println(\"No C2 for \" + jsonString + \"\\n\");\n if(jsonObj.get(\"\\\"C3\\\"\") == null) System.out.println(\"No C3 for \" + jsonString + \"\\n\");\n if(jsonObj.get(\"\\\"C4\\\"\") == null) System.out.println(\"No C4for \" + jsonString + \"\\n\");\n }\n }\n });\n\n }\n\n executorService.shutdown();\n\n try {\n executorService.awaitTermination(Long.MAX_VALUE, TimeUnit.NANOSECONDS);\n } catch (InterruptedException e) {\n e.printStackTrace();\n }\n\n cluster.close();\n }\n}\n","from":"developer"},{"body":"[~henryman], the {{Json}} class can definitely be tested in isolation, no need for integration tests in this case.","from":"developer"},{"body":"I agree. Thanks. I am new to the code. The patch from Ivan is pointing me to the places to look at in the code.","from":"developer"},{"body":"We kind of need to fix this asap and I'm a little bit lost of who is actively working on that. So unless someone else come up with a patch updated with Sergio's comments and proper testing in the next few days, can you have a look [~thobbs].","from":"developer"},{"body":"[~sbtourist] I've pushed a set of branches with your suggested changes and a unit test, would you mind reviewing?\n\nBranches and pending CI runs:\n|[CASSANDRA-11048|https://github.com/thobbs/cassandra/tree/CASSANDRA-11048]|[testall|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-dtest]|\n|[CASSANDRA-11048-3.0|https://github.com/thobbs/cassandra/tree/CASSANDRA-11048-3.0]|[testall|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-3.0-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-3.0-dtest]|\n|[CASSANDRA-11048-trunk|https://github.com/thobbs/cassandra/tree/CASSANDRA-11048-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/thobbs/job/thobbs-CASSANDRA-11048-trunk-dtest]|","from":"developer"},{"body":"[~thobbs],\n\nthe fix looks good, but the test doesn't seem correct to me, for the following reasons:\n* The call to {{fail(exc.getMessage())}} inside the {{Runnable}} doesn't actually make the test fail, because it doesn't propagate to the main thread: you should either wait on the futures, or \"capture\" the exception by yourself.\n* The call to {{executor.awaitTermination(1, TimeUnit.MINUTES)}} should be asserted true, so the test will fail if the other threads hang for any reason.\n\nOther than that, on a more general note, the method might have been tested in isolation via an actual unit test that only verifies the quoted string is not messed up, without having to create tables and rows, but I see the rest of the test class is made like that, so I will not object.","from":"developer"},{"body":"[~sbtourist] thank you for the careful review. I've pushed a second commit to address your comments. Since the previous test runs look good and this commit only changes the test behavior, I haven't started a second round of tests.\n\nbq. Other than that, on a more general note, the method might have been tested in isolation via an actual unit test that only verifies the quoted string is not messed up, without having to create tables and rows, but I see the rest of the test class is made like that, so I will not object.\n\nThe test is still quite fast to execute, and I prefer the extra (future) safety of testing the full stack here, so I went with the integration test approach.","from":"developer"},{"body":"Looks good now, but it seems it wasn't merged upwards?","from":"developer"},{"body":"bq. Looks good now, but it seems it wasn't merged upwards?\n\nThere weren't any merge conflicts, so the CI tests were the only reason to create the other branches. Since I didn't feel the CI tests were necessary for the second commit, I didn't merge into the other branches. Sorry for the confusion there.\n\nAnyway, that sounds like a +1, so I've committed this as {{6b1bd1745ab9f04f9e379e6d22764b97693ba2ad}} in 2.2 and merge upward to 3.0 and trunk. Thanks!","from":"developer"}],"created":"2016-01-20T17:41:10.000+0000","description":"{{org.apache.cassandra.cql3.Json}} uses a shared instance of {{JsonStringEncoder}} which is not thread safe (see 1), while {{JsonStringEncoder#getInstance()}} should be used (see 2).\n\nAs a consequence, concurrent {{select JSON}} queries often produce wrong (sometimes unreadable) results.\n\n1. http://grepcode.com/file/repo1.maven.org/maven2/org.codehaus.jackson/jackson-core-asl/1.9.2/org/codehaus/jackson/io/JsonStringEncoder.java\n2. http://grepcode.com/file/repo1.maven.org/maven2/org.codehaus.jackson/jackson-core-asl/1.9.2/org/codehaus/jackson/io/JsonStringEncoder.java#JsonStringEncoder.getInstance%28%29","issue_id":"12932787","key":"CASSANDRA-11048","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-02-11T16:34:42.000+0000","role":"gold_target","summary":"JSON queries are not thread safe"} {"case_id":"12956572","cluster":"JIRA-CASSANDRA-88571d61eb49","comments":[{"body":"I bisected this, and it looks like this problem was introduced in [CASSANDRA-9986].","created":"2016-04-06T22:13:21.196+0000"},{"body":"I looked into this because it looked interesting. [CASSANDRA-9986] removed the SliceableUnfilteredRowIterator. In {{SinglePartitionReadCommand.queryMemtablesAndSSTablesInTimestampOrder}} and also {{queryMemtablesAndDiskInternal}}, it switched from using {{ClusteringIndexFilter.filter}} to {{filter.getSlices}} and handing these slices to sstable.iterator. This caused the problem.\n\nBefore, if after reduceFilter we still needed to find a static and we had no clustering, the filter would make no attempt to read farther than the static row. Now, if clustering is empty, we produce no slices in {{getSlices}}, and so when the sstable gets an iterator, it never sets a slice for the reader and instead reads the whole partition. It seems to me that AbstractSSTableIterator doesn't correctly handle the case of empty slices in general and that we can reproduce this with a condition like {{a > 7 AND a < 5}}. I also think that the index was unrelated and only caused a flush to disk.\n\n(There's probably some imprecision here, but that should get someone most of the way there.)","created":"2016-04-07T03:28:57.123+0000"},{"body":"bq. It seems to me that AbstractSSTableIterator doesn't correctly handle the case of empty slices in general\n\nAgreed. Do you want to give a shot at a patch?","created":"2016-04-07T08:03:57.403+0000"},{"body":"Cassandra community is awesome! Should buy you a beer, Joel.","created":"2016-04-07T15:37:54.481+0000"},{"body":"I've uploaded a small patch that adds a NoRowsReader (analogous to the noRowsIterator used in the memtable case) as well as tests covering when slices is empty. This is the most elegant and least invasive solution I could find, but I'm pretty unfamiliar with this part of the codebase, so any alternatives suggestion would be appreciated.\n\nCI looks clean relative to upstream.\n||branch||testall||dtest||\n|[11513-trunk|https://github.com/jkni/cassandra/tree/11513-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-11513-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-11513-trunk-dtest]|\n","created":"2016-04-10T19:36:04.233+0000"},{"body":"Lgtm, +1. Committed, thanks.","created":"2016-04-18T13:22:12.151+0000"}],"conversations":[{"body":"[cqlsh 5.0.1 | Cassandra 3.4 | CQL spec 3.4.0 | Native protocol v4]\n\nRun followings,\n{code}\ndrop table if exists test0;\nCREATE TABLE test0 (\n pk int,\n a int,\n b text,\n s text static,\n PRIMARY KEY (pk, a)\n);\n\ninsert into test0 (pk,a,b,s) values (0,1,'b1','hello b1');\ninsert into test0 (pk,a,b,s) values (0,2,'b2','hello b2');\ninsert into test0 (pk,a,b,s) values (0,3,'b3','hello b3');\ncreate index on test0 (b);\ninsert into test0 (pk,a,b,s) values (0,2,'b2 again','b2 again');\n{code}\n\nNow select one record based on primary key, we got all three records.\n{code}\ncqlsh:ops> select * from test0 where pk=0 and a=2;\n\n pk | a | s | b\n----+---+----------+----------\n 0 | 1 | b2 again | b1\n 0 | 2 | b2 again | b2 again\n 0 | 3 | b2 again | b3\n{code}\n\n{code}\ncqlsh:ops> desc test0;\n\nCREATE TABLE ops.test0 (\n pk int,\n a int,\n b text,\n s text static,\n PRIMARY KEY (pk, a)\n) WITH CLUSTERING ORDER BY (a ASC)\n AND bloom_filter_fp_chance = 0.01\n AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\n AND comment = ''\n AND compaction = {'class': 'org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy', 'max_threshold': '32', 'min_threshold': '4'}\n AND compression = {'chunk_length_in_kb': '64', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}\n AND crc_check_chance = 1.0\n AND dclocal_read_repair_chance = 0.1\n AND default_time_to_live = 0\n AND gc_grace_seconds = 864000\n AND max_index_interval = 2048\n AND memtable_flush_period_in_ms = 0\n AND min_index_interval = 128\n AND read_repair_chance = 0.0\n AND speculative_retry = '99PERCENTILE';\nCREATE INDEX test0_b_idx ON ops.test0 (b);\n{code}","from":"reporter","subject":"Result set is not unique on primary key (cql)"},{"body":"I bisected this, and it looks like this problem was introduced in [CASSANDRA-9986].","from":"developer"},{"body":"I looked into this because it looked interesting. [CASSANDRA-9986] removed the SliceableUnfilteredRowIterator. In {{SinglePartitionReadCommand.queryMemtablesAndSSTablesInTimestampOrder}} and also {{queryMemtablesAndDiskInternal}}, it switched from using {{ClusteringIndexFilter.filter}} to {{filter.getSlices}} and handing these slices to sstable.iterator. This caused the problem.\n\nBefore, if after reduceFilter we still needed to find a static and we had no clustering, the filter would make no attempt to read farther than the static row. Now, if clustering is empty, we produce no slices in {{getSlices}}, and so when the sstable gets an iterator, it never sets a slice for the reader and instead reads the whole partition. It seems to me that AbstractSSTableIterator doesn't correctly handle the case of empty slices in general and that we can reproduce this with a condition like {{a > 7 AND a < 5}}. I also think that the index was unrelated and only caused a flush to disk.\n\n(There's probably some imprecision here, but that should get someone most of the way there.)","from":"developer"},{"body":"bq. It seems to me that AbstractSSTableIterator doesn't correctly handle the case of empty slices in general\n\nAgreed. Do you want to give a shot at a patch?","from":"developer"},{"body":"Cassandra community is awesome! Should buy you a beer, Joel.","from":"developer"},{"body":"I've uploaded a small patch that adds a NoRowsReader (analogous to the noRowsIterator used in the memtable case) as well as tests covering when slices is empty. This is the most elegant and least invasive solution I could find, but I'm pretty unfamiliar with this part of the codebase, so any alternatives suggestion would be appreciated.\n\nCI looks clean relative to upstream.\n||branch||testall||dtest||\n|[11513-trunk|https://github.com/jkni/cassandra/tree/11513-trunk]|[testall|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-11513-trunk-testall]|[dtest|http://cassci.datastax.com/view/Dev/view/jkni/job/jkni-11513-trunk-dtest]|\n","from":"developer"},{"body":"Lgtm, +1. Committed, thanks.","from":"developer"}],"created":"2016-04-06T17:47:45.000+0000","description":"[cqlsh 5.0.1 | Cassandra 3.4 | CQL spec 3.4.0 | Native protocol v4]\n\nRun followings,\n{code}\ndrop table if exists test0;\nCREATE TABLE test0 (\n pk int,\n a int,\n b text,\n s text static,\n PRIMARY KEY (pk, a)\n);\n\ninsert into test0 (pk,a,b,s) values (0,1,'b1','hello b1');\ninsert into test0 (pk,a,b,s) values (0,2,'b2','hello b2');\ninsert into test0 (pk,a,b,s) values (0,3,'b3','hello b3');\ncreate index on test0 (b);\ninsert into test0 (pk,a,b,s) values (0,2,'b2 again','b2 again');\n{code}\n\nNow select one record based on primary key, we got all three records.\n{code}\ncqlsh:ops> select * from test0 where pk=0 and a=2;\n\n pk | a | s | b\n----+---+----------+----------\n 0 | 1 | b2 again | b1\n 0 | 2 | b2 again | b2 again\n 0 | 3 | b2 again | b3\n{code}\n\n{code}\ncqlsh:ops> desc test0;\n\nCREATE TABLE ops.test0 (\n pk int,\n a int,\n b text,\n s text static,\n PRIMARY KEY (pk, a)\n) WITH CLUSTERING ORDER BY (a ASC)\n AND bloom_filter_fp_chance = 0.01\n AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}\n AND comment = ''\n AND compaction = {'class': 'org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy', 'max_threshold': '32', 'min_threshold': '4'}\n AND compression = {'chunk_length_in_kb': '64', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}\n AND crc_check_chance = 1.0\n AND dclocal_read_repair_chance = 0.1\n AND default_time_to_live = 0\n AND gc_grace_seconds = 864000\n AND max_index_interval = 2048\n AND memtable_flush_period_in_ms = 0\n AND min_index_interval = 128\n AND read_repair_chance = 0.0\n AND speculative_retry = '99PERCENTILE';\nCREATE INDEX test0_b_idx ON ops.test0 (b);\n{code}","issue_id":"12956572","key":"CASSANDRA-11513","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-04-18T13:22:12.000+0000","role":"gold_target","summary":"Result set is not unique on primary key (cql)"} {"case_id":"12972277","cluster":"JIRA-CASSANDRA-c72ee123687a","comments":[{"body":"I can't find any ways of assigning it to myself, is it a permissions issue?\n\nI've basically made a really simple patch - it just throws an IOException instead of an AssertionError (but it looks like there are many asserts scattered around the code base, any reason why?) with a more helpful message so that the coordinator doesn't see the other nodes as down forever. I've also written a couple of tests and attempted to use the existing logic in QueryProcessor (That looks like the right place to put it) to sanitize the SELECT to begin with. Feel free to leave behind any comments! ","created":"2016-05-24T04:51:30.352+0000"},{"body":"The code looks good overall (with small style issue in {{overloadBuffer}})), but I'd prefer to throw an {{IllegalArgumentException}} instead of {{IOException}} and would use {{assertInvalidThrow}} in the test to make sure the right class of exception is thrown.","created":"2016-05-26T08:22:37.427+0000"},{"body":"I think the proper fix is the added call to {{validateComposite}}. With that, we shouldn't ever reach the assertion in {{ByteBufferUtil}}, and so I'd personally just left it as an assertion. ","created":"2016-05-26T10:10:23.563+0000"},{"body":"Hey Branimir and Sylvain, thanks for reviewing. I've updated the patch and re-ran the tests. \n\n{{overloadBuffer}}: I re-read the Code style documentation, I've made changes to the braces and moved the comment away so the line isn't overly long, let me know if there are any more format changes required.\n\n{{assertInvalidThrow}}: Didn't realise this method existed, updated the tests. \n\n{{ByteBufferUtil}}: I've updated it to throw an IllegalArgumentException. I wasn't certain if it could be reached via another way (I'm not familiar with the code base, naively I assume no) but I guess the idea was that if it did, at least the node will not see the other node as down forever after that until restart. Let me know if you feel strongly to leave it as assert. \n\nI also did try this on C* 3.5, it doesn't suffer from the same problem as 2.1, which I guess is due to the big revamp of the storage engine :)","created":"2016-05-27T03:44:27.153+0000"},{"body":"bq. but I guess the idea was that if it did, at least the node will not see the other node as down forever after that until restart\n\nBut my point is, are you sure that is the case? I mean, is there some place where we catch {{IllegalArgumentException}} and not {{AssertionError}} so that which exception is thrown makes a difference? If there is, then maybe the right fix is to properly catch all exceptions at that particular place.","created":"2016-05-27T07:40:41.911+0000"},{"body":"That makes sense. There is indeed a place in {{writeConnected}} in {{OutboundTCPConnection}} catching Exception. I've updated the patch to make that catch Throwable since an AssertionError extends Error, and the asserts will just have a nicer message. \n\n","created":"2016-05-27T08:11:45.853+0000"},{"body":"Code LGTM. Uploaded for testing:\n|[2.1 branch|https://github.com/blambov/cassandra/tree/11882-2.1]|[utest|http://cassci.datastax.com/job/blambov-11882-2.1-testall/]|[dtest|http://cassci.datastax.com/job/blambov-11882-2.1-dtest/]|\n\nCould you make a patch for 2.2 as well ({{SelectStatement.java}} has some incompatible changes)? And possibly a test that makes sure this is fine in 3.0+?","created":"2016-05-27T08:35:34.373+0000"},{"body":"Sorry for being slightly annoying, but catching {{Throwable}} in {{writeConnected}} is imo right, but it means we can catch stuffs like {{OutOfMemoryException}} for which we kind of want to let the JVM die. So basically, we should call {{JVMStabilityInspector.inspectThrowable()}} at the beginning of the {{catch}} clause. We probably always should have done that in fact really, so 3.0 should probably get that part (and yes, the storage engine revamp is why you can know have larger than 64k values in clustering). ","created":"2016-05-27T09:59:12.006+0000"},{"body":"Sylvain, it's not annoying, I agree with you on that. I've changed it.\n\nBranimir, I've attached the 2 patches again (updated). I've put in the check for the JVM and moved the Insert tests in the 2.1 patch to a new {{InsertTest}} class because they really aren't Selects. I've also removed the {{overloadBuffer}} method and followed the idea of using just a static {{TOO_BIG}} variable that were present in 2.2 and 3, that way we don't need an extra method.\n\nI did retry this on 3 because the tests I ported over doesn't work on Inserts into CK with values larger than 64k. Using trunk, if I attempt (cqlsh) to run an insert query that with a CK that is larger than 64k, it actually works. I can even select it afterwards and it returns the inserted record well, so I mistakenly thought it worked fine on 3. But if I restart Cassandra, it will never succeed in starting up:\n\n{code}\nCaused by: java.lang.AssertionError: 131082\n\tat org.apache.cassandra.utils.ByteBufferUtil.writeWithShortLength(ByteBufferUtil.java:308) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.metadata.StatsMetadata$StatsMetadataSerializer.serialize(StatsMetadata.java:286) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.metadata.StatsMetadata$StatsMetadataSerializer.serialize(StatsMetadata.java:235) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.metadata.MetadataSerializer.serialize(MetadataSerializer.java:75) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.format.big.BigTableWriter.writeMetadata(BigTableWriter.java:378) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.format.big.BigTableWriter.access$300(BigTableWriter.java:51) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.format.big.BigTableWriter$TransactionalProxy.doPrepare(BigTableWriter.java:342) ~[main/:na]\n\tat org.apache.cassandra.utils.concurrent.Transactional$AbstractTransactional.prepareToCommit(Transactional.java:173) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.format.SSTableWriter.prepareToCommit(SSTableWriter.java:280) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.SimpleSSTableMultiWriter.prepareToCommit(SimpleSSTableMultiWriter.java:101) ~[main/:na]\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1145) ~[main/:na]\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1095) ~[main/:na]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) ~[na:1.8.0_40]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) ~[na:1.8.0_40]\n\tat java.lang.Thread.run(Thread.java:745) ~[na:1.8.0_40]\n{code}\n\nIt can also be reproduced by doing the INSERT with a CK having a value larger than 64k, and then calling {{nodetool flush}}. It looks like because writing {{StatsMetadata}} while flushing Memtables it still calls {{writeWithShortLength}}. I'll test with changing it to {{writeWithLength}}. ","created":"2016-05-30T07:45:13.981+0000"},{"body":"That is not a good idea - even if I could get it to work in this version by serializing and deserializing with more bytes than before, it will break backwards compatibility - old SSTables that have been serialized with {{writeWithShortLength}} will not be able to be deserialized.\n\nGiven that it may just be worthwhile to sanitize the input as before to prevent users from inserting a Clustering Key larger than 64k, also because even if the underlying storage structure allowed for it, it's still a pretty big size to be using for Clustering key. I'm also thinking that if someone did insert a CK at the moment larger than 64k, it will keep working until the Memtable is flushed (either by it getting full or C* restarting and the CommitLogs replaying) at which point AssertionErrors are thrown. If C* restarts this way, it will never succeed in starting up (I had to delete the commitlogs folder). Any thoughts? ","created":"2016-05-31T07:38:16.736+0000"},{"body":"I've just included the patch with the input sanitization for now for 3.X (I based it off trunk). ","created":"2016-05-31T08:05:43.523+0000"},{"body":"Uploaded for testing:\n|[2.1 branch|https://github.com/blambov/cassandra/tree/11882-2.1]|[utest|http://cassci.datastax.com/job/blambov-11882-2.1-testall/]|[dtest|http://cassci.datastax.com/job/blambov-11882-2.1-dtest/]|\n|[2.2 branch|https://github.com/blambov/cassandra/tree/11882-2.2]|[utest|http://cassci.datastax.com/job/blambov-11882-2.2-testall/]|[dtest|http://cassci.datastax.com/job/blambov-11882-2.2-dtest/]|\n|[3.0 branch|https://github.com/blambov/cassandra/tree/11882-3.0]|[utest|http://cassci.datastax.com/job/blambov-11882-3.0-testall/]|[dtest|http://cassci.datastax.com/job/blambov-11882-3.0-dtest/]|\n|[trunk|https://github.com/blambov/cassandra/tree/11882]|[utest|http://cassci.datastax.com/job/blambov-11882-testall/]|[dtest|http://cassci.datastax.com/job/blambov-11882-dtest/]|\n","created":"2016-05-31T08:50:02.294+0000"},{"body":"We don't want to add such limitation in 3.0. The fact that don't have a 64K limit for the clustering columns value is a feature in 3.0, not a bug, and we don't want to artificially add it back. The only thing that 3.0 might need is the change to {{OutboundTcpConnection}}, but as far as I can tell, that's about it.","created":"2016-05-31T09:46:15.459+0000"},{"body":"Sylvain, the 64k+ keys currently _do not_ work in 3.x. They are not always serialized properly. We should find a way to fix the problem and make sure they are properly supported, but not as part of this ticket.\n\nWhile they are not working, it's best to validate and error out early rather than accept the writes and let the node fail badly on flush and commit log recovery, thus I _would_ include the validation in the 3.x patch.","created":"2016-05-31T11:52:37.747+0000"},{"body":"I'd love more details on why it doesn't work on 3.0, and in particular what \"not always\" means. But if it doesn't, it's a bug (by opposition to 2.1/2.2 where it's a genuine limitation of the underlying format) and I'd rather we start by checking if it's not an easy fix before rejecting queries, because I suspect we'll just end up forgetting about it and let it exist longer that it has to be. I'm happy to create a separate ticket for that, to be assigned on that ticket and to look at it ASAP. But I'm not happy shoving it under the rug without at least having looked at it.","created":"2016-06-01T07:53:26.830+0000"},{"body":"From [the comments above|https://issues.apache.org/jira/browse/CASSANDRA-11882?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15306355#comment-15306355] I understand the fix will need a change in {{StatsMetadata}} format, which could require major version change or would otherwise be involved enough to warrant a separate ticket.\n\n[~Lerh Low], could you create a new issue explaining what goes wrong with 64k+ keys in 3.0+ and assign it to Sylvain?\n\nI restarted the tests as something went wrong with half of them yesterday.","created":"2016-06-01T10:10:02.821+0000"},{"body":"bq. From the comments above I understand the fix will need a change in StatsMetadata format, which could require major version change or would otherwise be involved enough to warrant a separate ticket.\n\nHum, somehow missed that comment, sorry. I guess that make sense, but this only limit each clustering column value to be <= 64k, while the patch currently limit the total size of all clustering values to be <= 64k. I guess if we don't forget to create a ticket to fix the sstable metadata (which you can indeed feel free to assign to me), then I guess I'm fine refusing larger than 64k values for clustering columns for now, as long as we don't limit the {{Clustering}} as a whole.","created":"2016-06-01T10:41:11.159+0000"},{"body":"Branimir & Sylvian,\n\nThanks for the reviews so far. I was under the impression that {{Clustering}} itself is already a clustering column, good catch. I've retraced through the code path and I've updated the patch so that it only refuses larger than 64k values for clustering columns (i.e an INSERT with ckey1 = 32k and ckey2 = 32k will work, and memtable flushing still works in that case too :)). \n\nI've created the other issue here: [CASSANDRA-11943|https://issues.apache.org/jira/browse/CASSANDRA-11943]. I don't think I have permissions to assign it to Sylvain, feel free to let me know if there is anything unclear with the setup etc. \n\n\n\n","created":"2016-06-02T01:32:10.259+0000"},{"body":"Updated the branches above with the latest version. Tests are now passing (to the extent that their respective base version are) and I'm happy with the code.\n\nMarking ready to commit.","created":"2016-06-15T09:23:05.618+0000"},{"body":"Committed, thanks.","created":"2016-06-23T09:07:41.639+0000"}],"conversations":[{"body":"Setup:\n\n{code}\nCREATE KEYSPACE Blues WITH REPLICATION = { 'class' : 'SimpleStrategy', 'replication_factor' : 2};\nCREATE TABLE test (a text, b text, PRIMARY KEY ((a), b))\n{code}\n\nThere currently doesn't seem to be an existing check for selecting clustering keys that are larger than 64k. So if we proceed to do the following select:\n\n{code}\nCONSISTENCY ALL;\nSELECT * FROM Blues.test WHERE a = 'foo' AND b = 'something larger than 64k';\n{code}\n\nAn AssertionError is thrown in `ByteBufferUtil` with just a number and an error message detailing 'Coordinator node timed out waiting for replica nodes responses' . Additionally, because an error extends Throwable (it's not a subclass of Exception), it's not caught so the connection between the coordinator node and the other nodes which have the replicas seem to be 'stuck' until it's restarted. Any other subsequent queries, even if it's just SELECT where a = 'foo' and b = 'bar', will always return the Coordinator timing out waiting for replica nodes responses'.","from":"reporter","subject":"Clustering Key with ByteBuffer size > 64k throws Assertion Error"},{"body":"I can't find any ways of assigning it to myself, is it a permissions issue?\n\nI've basically made a really simple patch - it just throws an IOException instead of an AssertionError (but it looks like there are many asserts scattered around the code base, any reason why?) with a more helpful message so that the coordinator doesn't see the other nodes as down forever. I've also written a couple of tests and attempted to use the existing logic in QueryProcessor (That looks like the right place to put it) to sanitize the SELECT to begin with. Feel free to leave behind any comments! ","from":"developer"},{"body":"The code looks good overall (with small style issue in {{overloadBuffer}})), but I'd prefer to throw an {{IllegalArgumentException}} instead of {{IOException}} and would use {{assertInvalidThrow}} in the test to make sure the right class of exception is thrown.","from":"developer"},{"body":"I think the proper fix is the added call to {{validateComposite}}. With that, we shouldn't ever reach the assertion in {{ByteBufferUtil}}, and so I'd personally just left it as an assertion. ","from":"developer"},{"body":"Hey Branimir and Sylvain, thanks for reviewing. I've updated the patch and re-ran the tests. \n\n{{overloadBuffer}}: I re-read the Code style documentation, I've made changes to the braces and moved the comment away so the line isn't overly long, let me know if there are any more format changes required.\n\n{{assertInvalidThrow}}: Didn't realise this method existed, updated the tests. \n\n{{ByteBufferUtil}}: I've updated it to throw an IllegalArgumentException. I wasn't certain if it could be reached via another way (I'm not familiar with the code base, naively I assume no) but I guess the idea was that if it did, at least the node will not see the other node as down forever after that until restart. Let me know if you feel strongly to leave it as assert. \n\nI also did try this on C* 3.5, it doesn't suffer from the same problem as 2.1, which I guess is due to the big revamp of the storage engine :)","from":"developer"},{"body":"bq. but I guess the idea was that if it did, at least the node will not see the other node as down forever after that until restart\n\nBut my point is, are you sure that is the case? I mean, is there some place where we catch {{IllegalArgumentException}} and not {{AssertionError}} so that which exception is thrown makes a difference? If there is, then maybe the right fix is to properly catch all exceptions at that particular place.","from":"developer"},{"body":"That makes sense. There is indeed a place in {{writeConnected}} in {{OutboundTCPConnection}} catching Exception. I've updated the patch to make that catch Throwable since an AssertionError extends Error, and the asserts will just have a nicer message. \n\n","from":"developer"},{"body":"Code LGTM. Uploaded for testing:\n|[2.1 branch|https://github.com/blambov/cassandra/tree/11882-2.1]|[utest|http://cassci.datastax.com/job/blambov-11882-2.1-testall/]|[dtest|http://cassci.datastax.com/job/blambov-11882-2.1-dtest/]|\n\nCould you make a patch for 2.2 as well ({{SelectStatement.java}} has some incompatible changes)? And possibly a test that makes sure this is fine in 3.0+?","from":"developer"},{"body":"Sorry for being slightly annoying, but catching {{Throwable}} in {{writeConnected}} is imo right, but it means we can catch stuffs like {{OutOfMemoryException}} for which we kind of want to let the JVM die. So basically, we should call {{JVMStabilityInspector.inspectThrowable()}} at the beginning of the {{catch}} clause. We probably always should have done that in fact really, so 3.0 should probably get that part (and yes, the storage engine revamp is why you can know have larger than 64k values in clustering). ","from":"developer"},{"body":"Sylvain, it's not annoying, I agree with you on that. I've changed it.\n\nBranimir, I've attached the 2 patches again (updated). I've put in the check for the JVM and moved the Insert tests in the 2.1 patch to a new {{InsertTest}} class because they really aren't Selects. I've also removed the {{overloadBuffer}} method and followed the idea of using just a static {{TOO_BIG}} variable that were present in 2.2 and 3, that way we don't need an extra method.\n\nI did retry this on 3 because the tests I ported over doesn't work on Inserts into CK with values larger than 64k. Using trunk, if I attempt (cqlsh) to run an insert query that with a CK that is larger than 64k, it actually works. I can even select it afterwards and it returns the inserted record well, so I mistakenly thought it worked fine on 3. But if I restart Cassandra, it will never succeed in starting up:\n\n{code}\nCaused by: java.lang.AssertionError: 131082\n\tat org.apache.cassandra.utils.ByteBufferUtil.writeWithShortLength(ByteBufferUtil.java:308) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.metadata.StatsMetadata$StatsMetadataSerializer.serialize(StatsMetadata.java:286) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.metadata.StatsMetadata$StatsMetadataSerializer.serialize(StatsMetadata.java:235) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.metadata.MetadataSerializer.serialize(MetadataSerializer.java:75) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.format.big.BigTableWriter.writeMetadata(BigTableWriter.java:378) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.format.big.BigTableWriter.access$300(BigTableWriter.java:51) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.format.big.BigTableWriter$TransactionalProxy.doPrepare(BigTableWriter.java:342) ~[main/:na]\n\tat org.apache.cassandra.utils.concurrent.Transactional$AbstractTransactional.prepareToCommit(Transactional.java:173) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.format.SSTableWriter.prepareToCommit(SSTableWriter.java:280) ~[main/:na]\n\tat org.apache.cassandra.io.sstable.SimpleSSTableMultiWriter.prepareToCommit(SimpleSSTableMultiWriter.java:101) ~[main/:na]\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.flushMemtable(ColumnFamilyStore.java:1145) ~[main/:na]\n\tat org.apache.cassandra.db.ColumnFamilyStore$Flush.run(ColumnFamilyStore.java:1095) ~[main/:na]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) ~[na:1.8.0_40]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) ~[na:1.8.0_40]\n\tat java.lang.Thread.run(Thread.java:745) ~[na:1.8.0_40]\n{code}\n\nIt can also be reproduced by doing the INSERT with a CK having a value larger than 64k, and then calling {{nodetool flush}}. It looks like because writing {{StatsMetadata}} while flushing Memtables it still calls {{writeWithShortLength}}. I'll test with changing it to {{writeWithLength}}. ","from":"developer"},{"body":"That is not a good idea - even if I could get it to work in this version by serializing and deserializing with more bytes than before, it will break backwards compatibility - old SSTables that have been serialized with {{writeWithShortLength}} will not be able to be deserialized.\n\nGiven that it may just be worthwhile to sanitize the input as before to prevent users from inserting a Clustering Key larger than 64k, also because even if the underlying storage structure allowed for it, it's still a pretty big size to be using for Clustering key. I'm also thinking that if someone did insert a CK at the moment larger than 64k, it will keep working until the Memtable is flushed (either by it getting full or C* restarting and the CommitLogs replaying) at which point AssertionErrors are thrown. If C* restarts this way, it will never succeed in starting up (I had to delete the commitlogs folder). Any thoughts? ","from":"developer"},{"body":"I've just included the patch with the input sanitization for now for 3.X (I based it off trunk). ","from":"developer"},{"body":"Uploaded for testing:\n|[2.1 branch|https://github.com/blambov/cassandra/tree/11882-2.1]|[utest|http://cassci.datastax.com/job/blambov-11882-2.1-testall/]|[dtest|http://cassci.datastax.com/job/blambov-11882-2.1-dtest/]|\n|[2.2 branch|https://github.com/blambov/cassandra/tree/11882-2.2]|[utest|http://cassci.datastax.com/job/blambov-11882-2.2-testall/]|[dtest|http://cassci.datastax.com/job/blambov-11882-2.2-dtest/]|\n|[3.0 branch|https://github.com/blambov/cassandra/tree/11882-3.0]|[utest|http://cassci.datastax.com/job/blambov-11882-3.0-testall/]|[dtest|http://cassci.datastax.com/job/blambov-11882-3.0-dtest/]|\n|[trunk|https://github.com/blambov/cassandra/tree/11882]|[utest|http://cassci.datastax.com/job/blambov-11882-testall/]|[dtest|http://cassci.datastax.com/job/blambov-11882-dtest/]|\n","from":"developer"},{"body":"We don't want to add such limitation in 3.0. The fact that don't have a 64K limit for the clustering columns value is a feature in 3.0, not a bug, and we don't want to artificially add it back. The only thing that 3.0 might need is the change to {{OutboundTcpConnection}}, but as far as I can tell, that's about it.","from":"developer"},{"body":"Sylvain, the 64k+ keys currently _do not_ work in 3.x. They are not always serialized properly. We should find a way to fix the problem and make sure they are properly supported, but not as part of this ticket.\n\nWhile they are not working, it's best to validate and error out early rather than accept the writes and let the node fail badly on flush and commit log recovery, thus I _would_ include the validation in the 3.x patch.","from":"developer"},{"body":"I'd love more details on why it doesn't work on 3.0, and in particular what \"not always\" means. But if it doesn't, it's a bug (by opposition to 2.1/2.2 where it's a genuine limitation of the underlying format) and I'd rather we start by checking if it's not an easy fix before rejecting queries, because I suspect we'll just end up forgetting about it and let it exist longer that it has to be. I'm happy to create a separate ticket for that, to be assigned on that ticket and to look at it ASAP. But I'm not happy shoving it under the rug without at least having looked at it.","from":"developer"},{"body":"From [the comments above|https://issues.apache.org/jira/browse/CASSANDRA-11882?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15306355#comment-15306355] I understand the fix will need a change in {{StatsMetadata}} format, which could require major version change or would otherwise be involved enough to warrant a separate ticket.\n\n[~Lerh Low], could you create a new issue explaining what goes wrong with 64k+ keys in 3.0+ and assign it to Sylvain?\n\nI restarted the tests as something went wrong with half of them yesterday.","from":"developer"},{"body":"bq. From the comments above I understand the fix will need a change in StatsMetadata format, which could require major version change or would otherwise be involved enough to warrant a separate ticket.\n\nHum, somehow missed that comment, sorry. I guess that make sense, but this only limit each clustering column value to be <= 64k, while the patch currently limit the total size of all clustering values to be <= 64k. I guess if we don't forget to create a ticket to fix the sstable metadata (which you can indeed feel free to assign to me), then I guess I'm fine refusing larger than 64k values for clustering columns for now, as long as we don't limit the {{Clustering}} as a whole.","from":"developer"},{"body":"Branimir & Sylvian,\n\nThanks for the reviews so far. I was under the impression that {{Clustering}} itself is already a clustering column, good catch. I've retraced through the code path and I've updated the patch so that it only refuses larger than 64k values for clustering columns (i.e an INSERT with ckey1 = 32k and ckey2 = 32k will work, and memtable flushing still works in that case too :)). \n\nI've created the other issue here: [CASSANDRA-11943|https://issues.apache.org/jira/browse/CASSANDRA-11943]. I don't think I have permissions to assign it to Sylvain, feel free to let me know if there is anything unclear with the setup etc. \n\n\n\n","from":"developer"},{"body":"Updated the branches above with the latest version. Tests are now passing (to the extent that their respective base version are) and I'm happy with the code.\n\nMarking ready to commit.","from":"developer"},{"body":"Committed, thanks.","from":"developer"}],"created":"2016-05-24T04:28:38.000+0000","description":"Setup:\n\n{code}\nCREATE KEYSPACE Blues WITH REPLICATION = { 'class' : 'SimpleStrategy', 'replication_factor' : 2};\nCREATE TABLE test (a text, b text, PRIMARY KEY ((a), b))\n{code}\n\nThere currently doesn't seem to be an existing check for selecting clustering keys that are larger than 64k. So if we proceed to do the following select:\n\n{code}\nCONSISTENCY ALL;\nSELECT * FROM Blues.test WHERE a = 'foo' AND b = 'something larger than 64k';\n{code}\n\nAn AssertionError is thrown in `ByteBufferUtil` with just a number and an error message detailing 'Coordinator node timed out waiting for replica nodes responses' . Additionally, because an error extends Throwable (it's not a subclass of Exception), it's not caught so the connection between the coordinator node and the other nodes which have the replicas seem to be 'stuck' until it's restarted. Any other subsequent queries, even if it's just SELECT where a = 'foo' and b = 'bar', will always return the Coordinator timing out waiting for replica nodes responses'.","issue_id":"12972277","key":"CASSANDRA-11882","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2016-06-23T09:07:41.000+0000","role":"gold_target","summary":"Clustering Key with ByteBuffer size > 64k throws Assertion Error"} {"case_id":"13077216","cluster":"JIRA-CASSANDRA-d001181e0a40","comments":[{"body":"There are some issues in sstabledump\n\n1. as you reported, frozen collections\n\n2. non-frozen UDT\n\n\n{quote}\nException in thread \"main\" java.lang.ClassCastException: org.apache.cassandra.db.marshal.UserType cannot be cast to org.apache.cassandra.db.marshal.CollectionType\n\tat org.apache.cassandra.tools.JsonTransformer.serializeCell(JsonTransformer.java:413)\n\tat org.apache.cassandra.tools.JsonTransformer.serializeColumnData(JsonTransformer.java:396)\n\tat org.apache.cassandra.tools.JsonTransformer.serializeRow(JsonTransformer.java:276)\n\tat org.apache.cassandra.tools.JsonTransformer.serializePartition(JsonTransformer.java:210)\n\tat java.util.stream.ForEachOps$ForEachOp$OfRef.accept(ForEachOps.java:184)\n\tat java.util.stream.ReferencePipeline$2$1.accept(ReferencePipeline.java:175)\n\tat java.util.Iterator.forEachRemaining(Iterator.java:116)\n\tat java.util.Spliterators$IteratorSpliterator.forEachRemaining(Spliterators.java:1801)\n\tat java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:481)\n\tat java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:471)\n\tat java.util.stream.ForEachOps$ForEachOp.evaluateSequential(ForEachOps.java:151)\n\tat java.util.stream.ForEachOps$ForEachOp$OfRef.evaluateSequential(ForEachOps.java:174)\n\tat java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:234)\n\tat java.util.stream.ReferencePipeline.forEach(ReferencePipeline.java:418)\n\tat org.apache.cassandra.tools.JsonTransformer.toJson(JsonTransformer.java:100)\n\tat org.apache.cassandra.tools.SSTableExport.main(SSTableExport.java:236)\n{quote}","created":"2017-07-06T08:19:41.101+0000"},{"body":"I am thinking to change all {{type.getString()}} in {{JsonTransformer}} to {{type.toJSONString}} when [13592|https://issues.apache.org/jira/browse/CASSANDRA-13592] merged.\n\nCurrent {{collectionType.getString()}} only generates entire data as byte string. Imo,{{getString}} is not designed for generate json readable values.","created":"2017-07-07T08:30:48.171+0000"},{"body":"The json transformer uses jackson's JsonGenerator so you do not need things designed to be in json output. Switching to {{toJSONString}} would end up double escaped as the jackson generator will try to remove the json encoding from the JSONString output.\n\nIf I remember correctly pretty much all UDTs were not supported initially because the information to deserialize the UDT requires the system schema tables to be read, which may not be available unless running on the same system with exact configuration and requires client initialization (which is dangerous and we dont want to do). If we have adequate info to deserialize the UDTs from the stats metadata it probably just needs a different check to handle it.","created":"2017-07-07T14:10:55.881+0000"},{"body":"I think we could use {{`writeRawValue(String)`}} to avoid double escape. Now that [13592|https://issues.apache.org/jira/browse/CASSANDRA-13592] is merged, we should have correct json representations for all types.. better than using {{type.toString}} which, imo, serves a different purposes.\n\nbq. all UDTs were not supported initially because the information to deserialize the UDT requires the system schema tables to be read (which is dangerous and we dont want to do)\n\nIf there is dependency to local schema tables, I agree not to support UDT and print raw-bytes instead.\n\nBut AFAIK after 8099, there should not be a dependency on schema tables while deserializing UDT.","created":"2017-07-09T11:23:35.417+0000"},{"body":"Some issues are related to {{ColumnMetadata.cellValueType()}} which currently : {{a}}. if CollectionType, returns value's type; {{b}} otherwise, its own type.\n\nIt doesn't handle properly: {{1}}. frozen collection type, {{2}}. non-frozen udt which requires cellPath to retrieve value's type.. \n\nThere are 3 kind of usage for {{ColumnMetadata.cellValueType()}}:\n\n1. to check if column is counter type, it's safe\n\n2. used in MV to check if cell value (base's non key column in view's primary key) is changed. \n in the existing implementation, it will get {{cellValueType}} of a {{frozen collection}} to decode {{frozen collection bytes}}\n it can easily result in runtime error. eg. base has non-key column {{frozen>}} as view's primary key. so {{tuple type}} is used to decode {{frozen}}\n \n {{non-frozen-udt}} cannot be used as view's primary key. the issue here is only with {{frozen-collection}}.\n \n3. in {{sasi}}. haven't check how it is affected. will check it later","created":"2017-07-09T15:36:58.442+0000"},{"body":"first draft of the patch. if it looks good, I will prepare fixes for 2.2/3.0/3.11 as well\n| [trunk|https://github.com/jasonstack/cassandra/commits/CASSANDRA-13573] | [unit|https://circleci.com/gh/jasonstack/cassandra/130] | [dtest|https://github.com/jasonstack/cassandra-dtest-riptano/commits/CASSANDRA-13573] |\n\nunit test: passed.\ndtest: {{cqlsh_tests.cqlsh_tests.TestCqlsh.test_describe}} & {{bootstrap_test.TestBootstrap.consistent_range_movement_false_with_rf1_should_succeed_test}} both are broken for some time\n\nchanges:\n1. use {{type.toJSONString()}} with {{json.writeRawValue()}} instead of {{type.getString()}} to generate readable content \n2. {{column.cellValueType}} now : {{a}}. if non-frozen collection, return value type, {{b}}. otherwise, return column type.\n","created":"2017-07-10T02:59:24.452+0000"},{"body":"About the impact on SASI, now sasi doesn't support {{complex}} type (aka, type with cellPath, non-frozen collection or udt). \n\nBy fixing {{column.cellValueTytpe}}, the {{columnIndex.isLiteral()}} is now properly returning {{false}} if indexed column is {{frozen-collection}}.","created":"2017-07-11T02:12:02.294+0000"},{"body":"The patch looks good to me, excellent job.\n\nThe assertion at {{ColumnMetadata#cellValueType()}} to check that the precondition mentioned in the comment is satisfied makes sense to me.\n\nThere are some minor nits about code style:\n* There are some missed space-after-comma at {{ViewTest#testFrozenCollectionsWithComplicatedInnerType()}}. Also, it would be nice for the sake of uniformity to align the create table and create view statement as they are in similar tests in the same file.\n* It would be good to add a {{@jira_ticket CASSANDRA-13573}} tag in the dtest docstring.\n* This is just an idea, but I think that the dtest {{CREATE TABLE}} and insertion could be more readable using a cell per line, something like:\n{code}\nsession.execute('CREATE TABLE ks.cf ('\n 'key int PRIMARY KEY,'\n 'list_f frozen>,'\n 'set_f frozen>, '\n 'map_f frozen>,'\n 'tuple_f frozen>, '\n 'user_type_f frozen, '\n 'list_v list,'\n 'set_v set,'\n 'map_v map,'\n 'tuple_v tuple,'\n 'user_type_v simple_type)')\n...\nsession.execute(statement, [1,\n [1, 2, 3], # list_f\n {1, 2, 3}, # set_f\n {1: 1, 2: 2, 3: 3}, # map_f\n (9, 9), # map_f\n FrozenUserType('SG', 100100, {'321', '123'}), # user_type_f\n [1, 2], # list_v\n {1, 2}, # set_v\n {1: 1, 2: 2}, # map_v\n (8, 8), # tuple_v\n NonFrozenUserType('SG', 100100) # user_type_v\n ])\n{code}\nWhat do you think?","created":"2017-08-01T10:02:22.701+0000"},{"body":"Can we make sure to test this with CASSANDRA-13683 ? The sstabledump tool was changed to be tool initiated for last few versions which can lead to UDTs mistakenly accessed from the C* system tables which are not available always. We do not want to add it as a requirement to have to run sstabledump on a C* node.","created":"2017-08-01T13:42:06.688+0000"},{"body":"| [trunk|https://github.com/jasonstack/cassandra/commits/CASSANDRA-13573-trunk] | [unit|https://circleci.com/gh/jasonstack/cassandra/356 ] |irrelevant:\nmaterialized_views_test.TestMatyerializedViews.view_tombstones_test \nbootstrap_test.TestBootstrap.consistent_range_movement_false_with_rf1_should_succeed_test\ncql_tests.cqlsh_tests.TestCqlsh.test_describe\n |\n| [3.11|https://github.com/jasonstack/cassandra/commits/CASSANDRA-13573-3.11] | [unit|https://circleci.com/gh/jasonstack/cassandra/360] | passed |\n| [3.0|https://github.com/jasonstack/cassandra/commits/CASSANDRA-13573-3.0] | [unit|https://circleci.com/gh/jasonstack/cassandra/350] | authe_test.TestAuth.system_auth_ks_is_alterable_test irrelevant | \n| [dtest|https://github.com/jasonstack/cassandra-dtest-riptano/commits/CASSANDRA-13573] |\n\nAddressed comments in dtest and src. \n\nI tested with \"clientInitialization()\" with data/schema folder removed. UDT works. ","created":"2017-08-02T05:13:18.595+0000"},{"body":"The changes look good to me. It seems that the CI tests that are finished are ok; it can be committed if the remaining tests pass.\n\nOne tiny detail that I forgot and can be fixed during commit, the comment \"Test that sstabledump against non-frozen udt\" in the dtest should be \"Test sstabledump against non-frozen udt\", without the \"that\".\n\nThanks!","created":"2017-08-03T10:08:29.055+0000"},{"body":"lgtm +1","created":"2017-08-04T19:22:02.073+0000"},{"body":"Committed to 3.0 as [3960260472fcd4e0243f62cc813992f1365197c6|https://github.com/apache/cassandra/commit/3960260472fcd4e0243f62cc813992f1365197c6] and merged into 3.11 and trunk.\n\nDtest committed to master as [959208749d70e5808aec144e87b73e90d56a7f91|https://github.com/apache/cassandra-dtest/commit/959208749d70e5808aec144e87b73e90d56a7f91]","created":"2017-08-08T14:57:32.009+0000"},{"body":"Thank you both for your review","created":"2017-08-08T15:03:45.996+0000"}],"conversations":[{"body":"Schema and data\"\n{noformat}\nCREATE TABLE ks.cf (\n hash blob,\n report_id timeuuid,\n subject_ids frozen>,\n PRIMARY KEY (hash, report_id)\n) WITH CLUSTERING ORDER BY (report_id DESC);\n\nINSERT INTO ks.cf (hash, report_id, subject_ids) VALUES (0x1213, now(), {1,2,4,5});\n{noformat}\n\nsstabledump output is:\n\n{noformat}\nsstabledump mc-1-big-Data.db \n[\n {\n \"partition\" : {\n \"key\" : [ \"1213\" ],\n \"position\" : 0\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 16,\n \"clustering\" : [ \"ec01eed0-49d9-11e7-b39a-97a96f529c02\" ],\n \"liveness_info\" : { \"tstamp\" : \"2017-06-05T10:29:57.434856Z\" },\n \"cells\" : [\n { \"name\" : \"subject_ids\", \"value\" : \"\" }\n ]\n }\n ]\n }\n]\n{noformat}\n\nWhile the values are really there:\n\n{noformat}\ncqlsh:ks> select * from cf ;\n\n hash | report_id | subject_ids\n--------+--------------------------------------+-------------\n 0x1213 | 02bafff0-49d9-11e7-b39a-97a96f529c02 | {1, 2, 4}\n{noformat}\n","from":"reporter","subject":"ColumnMetadata.cellValueType() doesn't return correct type for non-frozen collection"},{"body":"There are some issues in sstabledump\n\n1. as you reported, frozen collections\n\n2. non-frozen UDT\n\n\n{quote}\nException in thread \"main\" java.lang.ClassCastException: org.apache.cassandra.db.marshal.UserType cannot be cast to org.apache.cassandra.db.marshal.CollectionType\n\tat org.apache.cassandra.tools.JsonTransformer.serializeCell(JsonTransformer.java:413)\n\tat org.apache.cassandra.tools.JsonTransformer.serializeColumnData(JsonTransformer.java:396)\n\tat org.apache.cassandra.tools.JsonTransformer.serializeRow(JsonTransformer.java:276)\n\tat org.apache.cassandra.tools.JsonTransformer.serializePartition(JsonTransformer.java:210)\n\tat java.util.stream.ForEachOps$ForEachOp$OfRef.accept(ForEachOps.java:184)\n\tat java.util.stream.ReferencePipeline$2$1.accept(ReferencePipeline.java:175)\n\tat java.util.Iterator.forEachRemaining(Iterator.java:116)\n\tat java.util.Spliterators$IteratorSpliterator.forEachRemaining(Spliterators.java:1801)\n\tat java.util.stream.AbstractPipeline.copyInto(AbstractPipeline.java:481)\n\tat java.util.stream.AbstractPipeline.wrapAndCopyInto(AbstractPipeline.java:471)\n\tat java.util.stream.ForEachOps$ForEachOp.evaluateSequential(ForEachOps.java:151)\n\tat java.util.stream.ForEachOps$ForEachOp$OfRef.evaluateSequential(ForEachOps.java:174)\n\tat java.util.stream.AbstractPipeline.evaluate(AbstractPipeline.java:234)\n\tat java.util.stream.ReferencePipeline.forEach(ReferencePipeline.java:418)\n\tat org.apache.cassandra.tools.JsonTransformer.toJson(JsonTransformer.java:100)\n\tat org.apache.cassandra.tools.SSTableExport.main(SSTableExport.java:236)\n{quote}","from":"developer"},{"body":"I am thinking to change all {{type.getString()}} in {{JsonTransformer}} to {{type.toJSONString}} when [13592|https://issues.apache.org/jira/browse/CASSANDRA-13592] merged.\n\nCurrent {{collectionType.getString()}} only generates entire data as byte string. Imo,{{getString}} is not designed for generate json readable values.","from":"developer"},{"body":"The json transformer uses jackson's JsonGenerator so you do not need things designed to be in json output. Switching to {{toJSONString}} would end up double escaped as the jackson generator will try to remove the json encoding from the JSONString output.\n\nIf I remember correctly pretty much all UDTs were not supported initially because the information to deserialize the UDT requires the system schema tables to be read, which may not be available unless running on the same system with exact configuration and requires client initialization (which is dangerous and we dont want to do). If we have adequate info to deserialize the UDTs from the stats metadata it probably just needs a different check to handle it.","from":"developer"},{"body":"I think we could use {{`writeRawValue(String)`}} to avoid double escape. Now that [13592|https://issues.apache.org/jira/browse/CASSANDRA-13592] is merged, we should have correct json representations for all types.. better than using {{type.toString}} which, imo, serves a different purposes.\n\nbq. all UDTs were not supported initially because the information to deserialize the UDT requires the system schema tables to be read (which is dangerous and we dont want to do)\n\nIf there is dependency to local schema tables, I agree not to support UDT and print raw-bytes instead.\n\nBut AFAIK after 8099, there should not be a dependency on schema tables while deserializing UDT.","from":"developer"},{"body":"Some issues are related to {{ColumnMetadata.cellValueType()}} which currently : {{a}}. if CollectionType, returns value's type; {{b}} otherwise, its own type.\n\nIt doesn't handle properly: {{1}}. frozen collection type, {{2}}. non-frozen udt which requires cellPath to retrieve value's type.. \n\nThere are 3 kind of usage for {{ColumnMetadata.cellValueType()}}:\n\n1. to check if column is counter type, it's safe\n\n2. used in MV to check if cell value (base's non key column in view's primary key) is changed. \n in the existing implementation, it will get {{cellValueType}} of a {{frozen collection}} to decode {{frozen collection bytes}}\n it can easily result in runtime error. eg. base has non-key column {{frozen>}} as view's primary key. so {{tuple type}} is used to decode {{frozen}}\n \n {{non-frozen-udt}} cannot be used as view's primary key. the issue here is only with {{frozen-collection}}.\n \n3. in {{sasi}}. haven't check how it is affected. will check it later","from":"developer"},{"body":"first draft of the patch. if it looks good, I will prepare fixes for 2.2/3.0/3.11 as well\n| [trunk|https://github.com/jasonstack/cassandra/commits/CASSANDRA-13573] | [unit|https://circleci.com/gh/jasonstack/cassandra/130] | [dtest|https://github.com/jasonstack/cassandra-dtest-riptano/commits/CASSANDRA-13573] |\n\nunit test: passed.\ndtest: {{cqlsh_tests.cqlsh_tests.TestCqlsh.test_describe}} & {{bootstrap_test.TestBootstrap.consistent_range_movement_false_with_rf1_should_succeed_test}} both are broken for some time\n\nchanges:\n1. use {{type.toJSONString()}} with {{json.writeRawValue()}} instead of {{type.getString()}} to generate readable content \n2. {{column.cellValueType}} now : {{a}}. if non-frozen collection, return value type, {{b}}. otherwise, return column type.\n","from":"developer"},{"body":"About the impact on SASI, now sasi doesn't support {{complex}} type (aka, type with cellPath, non-frozen collection or udt). \n\nBy fixing {{column.cellValueTytpe}}, the {{columnIndex.isLiteral()}} is now properly returning {{false}} if indexed column is {{frozen-collection}}.","from":"developer"},{"body":"The patch looks good to me, excellent job.\n\nThe assertion at {{ColumnMetadata#cellValueType()}} to check that the precondition mentioned in the comment is satisfied makes sense to me.\n\nThere are some minor nits about code style:\n* There are some missed space-after-comma at {{ViewTest#testFrozenCollectionsWithComplicatedInnerType()}}. Also, it would be nice for the sake of uniformity to align the create table and create view statement as they are in similar tests in the same file.\n* It would be good to add a {{@jira_ticket CASSANDRA-13573}} tag in the dtest docstring.\n* This is just an idea, but I think that the dtest {{CREATE TABLE}} and insertion could be more readable using a cell per line, something like:\n{code}\nsession.execute('CREATE TABLE ks.cf ('\n 'key int PRIMARY KEY,'\n 'list_f frozen>,'\n 'set_f frozen>, '\n 'map_f frozen>,'\n 'tuple_f frozen>, '\n 'user_type_f frozen, '\n 'list_v list,'\n 'set_v set,'\n 'map_v map,'\n 'tuple_v tuple,'\n 'user_type_v simple_type)')\n...\nsession.execute(statement, [1,\n [1, 2, 3], # list_f\n {1, 2, 3}, # set_f\n {1: 1, 2: 2, 3: 3}, # map_f\n (9, 9), # map_f\n FrozenUserType('SG', 100100, {'321', '123'}), # user_type_f\n [1, 2], # list_v\n {1, 2}, # set_v\n {1: 1, 2: 2}, # map_v\n (8, 8), # tuple_v\n NonFrozenUserType('SG', 100100) # user_type_v\n ])\n{code}\nWhat do you think?","from":"developer"},{"body":"Can we make sure to test this with CASSANDRA-13683 ? The sstabledump tool was changed to be tool initiated for last few versions which can lead to UDTs mistakenly accessed from the C* system tables which are not available always. We do not want to add it as a requirement to have to run sstabledump on a C* node.","from":"developer"},{"body":"| [trunk|https://github.com/jasonstack/cassandra/commits/CASSANDRA-13573-trunk] | [unit|https://circleci.com/gh/jasonstack/cassandra/356 ] |irrelevant:\nmaterialized_views_test.TestMatyerializedViews.view_tombstones_test \nbootstrap_test.TestBootstrap.consistent_range_movement_false_with_rf1_should_succeed_test\ncql_tests.cqlsh_tests.TestCqlsh.test_describe\n |\n| [3.11|https://github.com/jasonstack/cassandra/commits/CASSANDRA-13573-3.11] | [unit|https://circleci.com/gh/jasonstack/cassandra/360] | passed |\n| [3.0|https://github.com/jasonstack/cassandra/commits/CASSANDRA-13573-3.0] | [unit|https://circleci.com/gh/jasonstack/cassandra/350] | authe_test.TestAuth.system_auth_ks_is_alterable_test irrelevant | \n| [dtest|https://github.com/jasonstack/cassandra-dtest-riptano/commits/CASSANDRA-13573] |\n\nAddressed comments in dtest and src. \n\nI tested with \"clientInitialization()\" with data/schema folder removed. UDT works. ","from":"developer"},{"body":"The changes look good to me. It seems that the CI tests that are finished are ok; it can be committed if the remaining tests pass.\n\nOne tiny detail that I forgot and can be fixed during commit, the comment \"Test that sstabledump against non-frozen udt\" in the dtest should be \"Test sstabledump against non-frozen udt\", without the \"that\".\n\nThanks!","from":"developer"},{"body":"lgtm +1","from":"developer"},{"body":"Committed to 3.0 as [3960260472fcd4e0243f62cc813992f1365197c6|https://github.com/apache/cassandra/commit/3960260472fcd4e0243f62cc813992f1365197c6] and merged into 3.11 and trunk.\n\nDtest committed to master as [959208749d70e5808aec144e87b73e90d56a7f91|https://github.com/apache/cassandra-dtest/commit/959208749d70e5808aec144e87b73e90d56a7f91]","from":"developer"},{"body":"Thank you both for your review","from":"developer"}],"created":"2017-06-05T10:38:59.000+0000","description":"Schema and data\"\n{noformat}\nCREATE TABLE ks.cf (\n hash blob,\n report_id timeuuid,\n subject_ids frozen>,\n PRIMARY KEY (hash, report_id)\n) WITH CLUSTERING ORDER BY (report_id DESC);\n\nINSERT INTO ks.cf (hash, report_id, subject_ids) VALUES (0x1213, now(), {1,2,4,5});\n{noformat}\n\nsstabledump output is:\n\n{noformat}\nsstabledump mc-1-big-Data.db \n[\n {\n \"partition\" : {\n \"key\" : [ \"1213\" ],\n \"position\" : 0\n },\n \"rows\" : [\n {\n \"type\" : \"row\",\n \"position\" : 16,\n \"clustering\" : [ \"ec01eed0-49d9-11e7-b39a-97a96f529c02\" ],\n \"liveness_info\" : { \"tstamp\" : \"2017-06-05T10:29:57.434856Z\" },\n \"cells\" : [\n { \"name\" : \"subject_ids\", \"value\" : \"\" }\n ]\n }\n ]\n }\n]\n{noformat}\n\nWhile the values are really there:\n\n{noformat}\ncqlsh:ks> select * from cf ;\n\n hash | report_id | subject_ids\n--------+--------------------------------------+-------------\n 0x1213 | 02bafff0-49d9-11e7-b39a-97a96f529c02 | {1, 2, 4}\n{noformat}\n","issue_id":"13077216","key":"CASSANDRA-13573","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2017-08-08T14:58:09.000+0000","role":"gold_target","summary":"ColumnMetadata.cellValueType() doesn't return correct type for non-frozen collection"} {"case_id":"12551799","cluster":"JIRA-CASSANDRA-63efa8001612","comments":[{"body":"I'll note that an initial idea could be to keep the row header as it is (post CASSANDRA-2319), and during compaction to keep the space for the row size and column count, compact all columns, and seek back to write those two values. However, compression forbids us to do that, so we'll have to really remove those part two. However, we can trade the column count by writing a specific marker to mark the end of a row. As for the data size, we can get it from the index.","created":"2012-04-20T15:04:16.797+0000"},{"body":"This will be a big win, I'm excited.","created":"2012-04-20T15:31:47.191+0000"},{"body":"About a years ago ,I tried to do this ,but find it difficult .","created":"2012-04-21T08:48:35.700+0000"},{"body":"Removing the row-level bloom filter would make this a lot simpler.","created":"2012-11-01T22:20:56.408+0000"},{"body":"bq. Removing the row-level bloom filter would make this a lot simpler.\n\nNote that it really matters at the end of the day, but just to make sure we're on the same page, now that the row-level filters have been promoted to the index file, I don't think removing or keeping will have much impact on this ticket. ","created":"2012-11-30T10:35:14.493+0000"},{"body":"[~slebresne] I think you are right about that, and disambiguating the two tickets (at least, in my brain) should make this a little easier.","created":"2012-11-30T22:21:27.228+0000"},{"body":"Ugh, this is almost straightforward. But doing without column count is kind of a bitch. Fortunately we can use 00/00 as the end-of-row marker since empty column names are not allowed.","created":"2013-03-29T23:15:00.175+0000"},{"body":"Removing the column column and row size and using an end-of-row (EOR) marker seems straightforward, but the devil has been in the details. The main challenge is reading to you read columns until the EOR marker, and you return false on the next hasNext(). The problem is if a higher-level iterator calls hasNext() again on that row, you've already moved the the file pointer past the EOR row (and are at the beginning of the next row). So you look for the EOR marker, which you've already read for that row, but the pointer reads either the key for the next row or EOF - either way, not so good. It took me while to work though that one, and I think I'm done resolving the bugs around removing column count.\n\nRemoving the row size has been largely easier, but I'm working through (hopefully) the last issue on SSTS.KeyScanningIterator. It uses row size to calculate the start of the next row. Without that, if you skip a row based solely upon key, then I need to efficiently advance the file pointer the the head of the next row.\n\nHoping to finish very soon.","created":"2013-03-31T13:08:33.780+0000"},{"body":"Sorry, I got caught up in this over the weekend since the more I dug on CASSANDRA-5344 the more it looked like I was actually solving this first. I've pushed my first draft to http://github.com/jbellis/cassandra/tree/4180.\n\nI see DefsTest fail occasionally but it is not 100% reproducible, so I'm not sure if it was caused by my changes. I also see CFSTest log errors sometimes but that one is definitely not related (CASSANDRA-5410).\n\n(Testing was a bitch at first since sstable errors would just cause the schema loader to break violently; hence the option I added to just inject the schema directly, without going through the migration path. That allowed enough tests to run to track down the problems.)\n\nEverything was fairly straightforward except SSTableScanner. (Sounds like the same thing Jason ran into.) I simplified things by noting that seekTo was only used to initialize the scanner to a certain starting point, so I pulled that into the constructor to make seeking mid-iteration a non-concern. (This also allowed removing SSTableBoundedScanner.) I also merged KeyScanningIterator and FilteringKSI; FKSI already had most of the code needed to compute data size from the index entries, which compaction needed to decide whether to use an eager or lazy approach.\n\nThere were a lot of places that just one-off sstable reading that were easy to miss. This smells fishy to me but it wasn't obvious how to re-organize things to make it unnecessary, so I haven't tried to solve that here.\n\nI also haven't tried to update scrub for the new format.","created":"2013-04-01T14:19:24.365+0000"},{"body":"On the whole, LGTM.\n\nI think Jonathan and I had the same basic notion of how to implement this, and we ran into many of the same problems (schema loading borked hard and fast). Admittedly, the one-off sstable reads were what kept kicking me in the shins (repeatedly), and what kept taking me down wrong paths. I agree about the fishy smell with knowledge of the SSTable format spread around, but minimally that's for a another ticket.\n\nWe can probably remove the {code}output != null{code} check in ColumnIndex.Builder.add() as LCR is no longer passing in null (on the first pass). \n\nJonathan mentioned he didn't try to scrub, but I just tried it out locally, and it failed with an OOM error. I'd look into it more, but brain power running out this late at night. I'm attaching a file with the scrub errors. Will look into it tomorrow.","created":"2013-04-02T07:32:36.980+0000"},{"body":"Broke out scrub into CASSANDRA-5429 so we can review this separately.","created":"2013-04-05T15:41:56.180+0000"},{"body":"Pushed updated code to http://github.com/jbellis/cassandra/tree/4180-3. Besides the rebase, this cleans things up by making ColumnIndex the sole entity responsible for writing the row header (thus removing the flag I'd added to it in the earlier version).","created":"2013-04-12T22:48:49.318+0000"},{"body":"On the whole, lgtm. Tests, especially scrub, are working correctly now. I do appreciate the row header creation being centralized now; that was tricky for me when I first started working on this.\n\nHowever, when I create a simple table under 1.2, then go to start this trunk, it fails on launch with:\n\n{code}\n INFO [main] 2013-04-17 06:49:47,456 CacheService.java (line 165) Scheduling row cache save to each 0 seconds (going to save all keys).\n INFO [SSTableBatchOpen:2] 2013-04-17 06:49:47,710 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_keyspaces/system-schema_keyspaces-ib-1 (261 bytes)\n INFO [SSTableBatchOpen:1] 2013-04-17 06:49:47,710 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_keyspaces/system-schema_keyspaces-ib-2 (165 bytes)\n INFO [SSTableBatchOpen:1] 2013-04-17 06:49:48,121 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columnfamilies/system-schema_columnfamilies-ib-3 (691 bytes)\n INFO [SSTableBatchOpen:2] 2013-04-17 06:49:48,121 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columnfamilies/system-schema_columnfamilies-ib-1 (4540 bytes)\n INFO [SSTableBatchOpen:3] 2013-04-17 06:49:48,122 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columnfamilies/system-schema_columnfamilies-ib-2 (699 bytes)\n INFO [SSTableBatchOpen:1] 2013-04-17 06:49:48,144 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columns/system-schema_columns-ib-2 (209 bytes)\n INFO [SSTableBatchOpen:2] 2013-04-17 06:49:48,145 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columns/system-schema_columns-ib-3 (194 bytes)\n INFO [SSTableBatchOpen:3] 2013-04-17 06:49:48,145 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columns/system-schema_columns-ib-1 (3768 bytes)\nINFO [SSTableBatchOpen:1] 2013-04-17 06:49:48,181 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/local/system-local-ib-5 (441 bytes)\nERROR [main] 2013-04-17 06:49:48,652 CassandraDaemon.java (line 454) Exception encountered during startup\njava.lang.OutOfMemoryError: Java heap space\n\tat org.apache.cassandra.io.util.RandomAccessReader.readBytes(RandomAccessReader.java:339)\n\tat org.apache.cassandra.utils.ByteBufferUtil.read(ByteBufferUtil.java:392)\n\tat org.apache.cassandra.utils.ByteBufferUtil.readWithLength(ByteBufferUtil.java:355)\n\tat org.apache.cassandra.db.ColumnSerializer.deserializeColumnBody(ColumnSerializer.java:124)\n\tat org.apache.cassandra.db.OnDiskAtom$Serializer.deserializeFromSSTable(OnDiskAtom.java:82)\n\tat org.apache.cassandra.db.Column$1.computeNext(Column.java:73)\n\tat org.apache.cassandra.db.Column$1.computeNext(Column.java:62)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n\tat org.apache.cassandra.db.columniterator.SimpleSliceReader.computeNext(SimpleSliceReader.java:92)\n\tat org.apache.cassandra.db.columniterator.SimpleSliceReader.computeNext(SimpleSliceReader.java:36)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n\tat org.apache.cassandra.db.columniterator.SSTableSliceIterator.hasNext(SSTableSliceIterator.java:90)\n\tat org.apache.cassandra.db.columniterator.LazyColumnIterator.computeNext(LazyColumnIterator.java:82)\n\tat org.apache.cassandra.db.columniterator.LazyColumnIterator.computeNext(LazyColumnIterator.java:59)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n\tat org.apache.cassandra.db.filter.QueryFilter$2.getNext(QueryFilter.java:136)\n\tat org.apache.cassandra.db.filter.QueryFilter$2.hasNext(QueryFilter.java:119)\n\tat org.apache.cassandra.utils.MergeIterator$OneToOne.computeNext(MergeIterator.java:199)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n\tat org.apache.cassandra.db.filter.SliceQueryFilter.collectReducedColumns(SliceQueryFilter.java:131)\n\tat org.apache.cassandra.db.filter.QueryFilter.collateColumns(QueryFilter.java:101)\n\tat org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:75)\n\tat org.apache.cassandra.db.RowIteratorFactory$2.getReduced(RowIteratorFactory.java:105)\n\tat org.apache.cassandra.db.RowIteratorFactory$2.getReduced(RowIteratorFactory.java:78)\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:114)\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:97)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n{code}\n\nI'll try to take a look today.\n\n","created":"2013-04-17T13:55:08.733+0000"},{"body":"I can confirm that creating same table under trunk, and relaunching, works correctly, as well as basic select querying. So I think it's something in the upgrade path from 1.2 to trunk.","created":"2013-04-17T14:00:38.239+0000"},{"body":"Just more info: tried my 1.2 to trunk upgrade again, and it's not related to my table at all. I simply started 1.2, let it create the system tables, killed it, and launched trunk - same exception. Will check it out.","created":"2013-04-17T14:04:41.962+0000"},{"body":"Rebased on top of CASSANDRA-5511 with a fix or two: http://github.com/jbellis/cassandra/tree/4180-4\n\nNow passes the \"start 1.2, rebuild and start 4180-4\" test.","created":"2013-04-26T21:01:30.076+0000"},{"body":"Rebased again, to https://github.com/jbellis/cassandra/commits/4180-5. [~krummas], can you review?","created":"2013-05-01T15:59:04.019+0000"},{"body":"Minor rebase to https://github.com/jbellis/cassandra/commits/4180-6","created":"2013-05-06T18:43:27.618+0000"},{"body":"in general, lgtm, few issues;\n\nsstable.getPosition(..) can return null in a couple of cases:\nDuring repair:\n{code}\nERROR [ValidationExecutor:2] 2013-05-07 10:43:24,875 CassandraDaemon.java (line 179) Exception in thread Thread[ValidationExecutor:2,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableScanner.(SSTableScanner.java:109)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1038)\n at org.apache.cassandra.db.compaction.AbstractCompactionStrategy.getScanners(AbstractCompactionStrategy.java:214)\n at org.apache.cassandra.db.compaction.CompactionManager$ValidationCompactionIterable.(CompactionManager.java:744)\n at org.apache.cassandra.db.compaction.CompactionManager.doValidationCompaction(CompactionManager.java:648)\n at org.apache.cassandra.db.compaction.CompactionManager.access$600(CompactionManager.java:64)\n at org.apache.cassandra.db.compaction.CompactionManager$8.call(CompactionManager.java:391)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n at java.util.concurrent.FutureTask.run(FutureTask.java:166)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n at java.lang.Thread.run(Thread.java:722)\n{code}\n\nand scrub is not supposed to work until CASSANDRA-5429 right?:\n{code}\n INFO 09:26:43,968 Retrying from row index; data is 261 bytes starting at 295\n WARN 09:26:43,969 Retry failed too. Skipping to next row (retry's stacktrace follows)\njava.lang.AssertionError: Interval min > max\n at org.apache.cassandra.utils.IntervalTree$IntervalNode.(IntervalTree.java:250)\n at org.apache.cassandra.utils.IntervalTree.(IntervalTree.java:72)\n at org.apache.cassandra.utils.IntervalTree.build(IntervalTree.java:81)\n at org.apache.cassandra.db.DeletionInfo.add(DeletionInfo.java:179)\n at org.apache.cassandra.db.AbstractThreadUnsafeSortedColumns.delete(AbstractThreadUnsafeSortedColumns.java:44)\n at org.apache.cassandra.db.ColumnFamily.addAtom(ColumnFamily.java:142)\n at org.apache.cassandra.db.ColumnFamilySerializer.deserializeColumnsFromSSTable(ColumnFamilySerializer.java:177)\n at org.apache.cassandra.io.sstable.SSTableIdentityIterator.getColumnFamilyWithColumns(SSTableIdentityIterator.java:184)\n at org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:114)\n at org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:98)\n at org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:220)\n at org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:226)\n at org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:204)\n at org.apache.cassandra.db.compaction.CompactionManager.scrubOne(CompactionManager.java:430)\n at org.apache.cassandra.db.compaction.CompactionManager.doScrub(CompactionManager.java:419)\n at org.apache.cassandra.db.compaction.CompactionManager.access$300(CompactionManager.java:64)\n at org.apache.cassandra.db.compaction.CompactionManager$3.perform(CompactionManager.java:234)\n at org.apache.cassandra.db.compaction.CompactionManager$2.call(CompactionManager.java:220)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n at java.util.concurrent.FutureTask.run(FutureTask.java:166)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n at java.lang.Thread.run(Thread.java:722)\n WARN 09:26:43,972 Non-fatal error reading row (stacktrace follows)\njava.io.IOError: java.io.IOException: Impossible row size 9223372034707292160\n at org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:170)\n at org.apache.cassandra.db.compaction.CompactionManager.scrubOne(CompactionManager.java:430)\n at org.apache.cassandra.db.compaction.CompactionManager.doScrub(CompactionManager.java:419)\n at org.apache.cassandra.db.compaction.CompactionManager.access$300(CompactionManager.java:64)\n at org.apache.cassandra.db.compaction.CompactionManager$3.perform(CompactionManager.java:234)\n at org.apache.cassandra.db.compaction.CompactionManager$2.call(CompactionManager.java:220)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n at java.util.concurrent.FutureTask.run(FutureTask.java:166)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n at java.lang.Thread.run(Thread.java:722)\nCaused by: java.io.IOException: Impossible row size 9223372034707292160\n ... 11 more\n{code}\n\nnits;\n* version argument unused in IndexedSliceReader#setToRowStart\n* LazilyCompactedRow comment obsolete\n* compactionType unused param in SSTableWriter#createWriter\n","created":"2013-05-07T09:17:46.855+0000"},{"body":"Pushed fixes to -6 branch. Yes, scrub is expected to be totally broken at this point.","created":"2013-05-07T13:36:07.708+0000"},{"body":"ok, this lgtm!","created":"2013-05-07T18:25:09.099+0000"},{"body":"committed!","created":"2013-05-07T19:22:45.360+0000"}],"conversations":[{"body":"LazilyCompactedRow reads all data twice to compact a row which is obviously inefficient. The main reason we do that is to compute the row header. However, CASSANDRA-2319 have removed the main part of that row header. What remains is the size in bytes and the number of columns, but it should be relatively simple to remove those, which would then remove the need for the two-phase compaction.","from":"reporter","subject":"Single-pass compaction for LCR"},{"body":"I'll note that an initial idea could be to keep the row header as it is (post CASSANDRA-2319), and during compaction to keep the space for the row size and column count, compact all columns, and seek back to write those two values. However, compression forbids us to do that, so we'll have to really remove those part two. However, we can trade the column count by writing a specific marker to mark the end of a row. As for the data size, we can get it from the index.","from":"developer"},{"body":"This will be a big win, I'm excited.","from":"developer"},{"body":"About a years ago ,I tried to do this ,but find it difficult .","from":"developer"},{"body":"Removing the row-level bloom filter would make this a lot simpler.","from":"developer"},{"body":"bq. Removing the row-level bloom filter would make this a lot simpler.\n\nNote that it really matters at the end of the day, but just to make sure we're on the same page, now that the row-level filters have been promoted to the index file, I don't think removing or keeping will have much impact on this ticket. ","from":"developer"},{"body":"[~slebresne] I think you are right about that, and disambiguating the two tickets (at least, in my brain) should make this a little easier.","from":"developer"},{"body":"Ugh, this is almost straightforward. But doing without column count is kind of a bitch. Fortunately we can use 00/00 as the end-of-row marker since empty column names are not allowed.","from":"developer"},{"body":"Removing the column column and row size and using an end-of-row (EOR) marker seems straightforward, but the devil has been in the details. The main challenge is reading to you read columns until the EOR marker, and you return false on the next hasNext(). The problem is if a higher-level iterator calls hasNext() again on that row, you've already moved the the file pointer past the EOR row (and are at the beginning of the next row). So you look for the EOR marker, which you've already read for that row, but the pointer reads either the key for the next row or EOF - either way, not so good. It took me while to work though that one, and I think I'm done resolving the bugs around removing column count.\n\nRemoving the row size has been largely easier, but I'm working through (hopefully) the last issue on SSTS.KeyScanningIterator. It uses row size to calculate the start of the next row. Without that, if you skip a row based solely upon key, then I need to efficiently advance the file pointer the the head of the next row.\n\nHoping to finish very soon.","from":"developer"},{"body":"Sorry, I got caught up in this over the weekend since the more I dug on CASSANDRA-5344 the more it looked like I was actually solving this first. I've pushed my first draft to http://github.com/jbellis/cassandra/tree/4180.\n\nI see DefsTest fail occasionally but it is not 100% reproducible, so I'm not sure if it was caused by my changes. I also see CFSTest log errors sometimes but that one is definitely not related (CASSANDRA-5410).\n\n(Testing was a bitch at first since sstable errors would just cause the schema loader to break violently; hence the option I added to just inject the schema directly, without going through the migration path. That allowed enough tests to run to track down the problems.)\n\nEverything was fairly straightforward except SSTableScanner. (Sounds like the same thing Jason ran into.) I simplified things by noting that seekTo was only used to initialize the scanner to a certain starting point, so I pulled that into the constructor to make seeking mid-iteration a non-concern. (This also allowed removing SSTableBoundedScanner.) I also merged KeyScanningIterator and FilteringKSI; FKSI already had most of the code needed to compute data size from the index entries, which compaction needed to decide whether to use an eager or lazy approach.\n\nThere were a lot of places that just one-off sstable reading that were easy to miss. This smells fishy to me but it wasn't obvious how to re-organize things to make it unnecessary, so I haven't tried to solve that here.\n\nI also haven't tried to update scrub for the new format.","from":"developer"},{"body":"On the whole, LGTM.\n\nI think Jonathan and I had the same basic notion of how to implement this, and we ran into many of the same problems (schema loading borked hard and fast). Admittedly, the one-off sstable reads were what kept kicking me in the shins (repeatedly), and what kept taking me down wrong paths. I agree about the fishy smell with knowledge of the SSTable format spread around, but minimally that's for a another ticket.\n\nWe can probably remove the {code}output != null{code} check in ColumnIndex.Builder.add() as LCR is no longer passing in null (on the first pass). \n\nJonathan mentioned he didn't try to scrub, but I just tried it out locally, and it failed with an OOM error. I'd look into it more, but brain power running out this late at night. I'm attaching a file with the scrub errors. Will look into it tomorrow.","from":"developer"},{"body":"Broke out scrub into CASSANDRA-5429 so we can review this separately.","from":"developer"},{"body":"Pushed updated code to http://github.com/jbellis/cassandra/tree/4180-3. Besides the rebase, this cleans things up by making ColumnIndex the sole entity responsible for writing the row header (thus removing the flag I'd added to it in the earlier version).","from":"developer"},{"body":"On the whole, lgtm. Tests, especially scrub, are working correctly now. I do appreciate the row header creation being centralized now; that was tricky for me when I first started working on this.\n\nHowever, when I create a simple table under 1.2, then go to start this trunk, it fails on launch with:\n\n{code}\n INFO [main] 2013-04-17 06:49:47,456 CacheService.java (line 165) Scheduling row cache save to each 0 seconds (going to save all keys).\n INFO [SSTableBatchOpen:2] 2013-04-17 06:49:47,710 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_keyspaces/system-schema_keyspaces-ib-1 (261 bytes)\n INFO [SSTableBatchOpen:1] 2013-04-17 06:49:47,710 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_keyspaces/system-schema_keyspaces-ib-2 (165 bytes)\n INFO [SSTableBatchOpen:1] 2013-04-17 06:49:48,121 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columnfamilies/system-schema_columnfamilies-ib-3 (691 bytes)\n INFO [SSTableBatchOpen:2] 2013-04-17 06:49:48,121 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columnfamilies/system-schema_columnfamilies-ib-1 (4540 bytes)\n INFO [SSTableBatchOpen:3] 2013-04-17 06:49:48,122 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columnfamilies/system-schema_columnfamilies-ib-2 (699 bytes)\n INFO [SSTableBatchOpen:1] 2013-04-17 06:49:48,144 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columns/system-schema_columns-ib-2 (209 bytes)\n INFO [SSTableBatchOpen:2] 2013-04-17 06:49:48,145 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columns/system-schema_columns-ib-3 (194 bytes)\n INFO [SSTableBatchOpen:3] 2013-04-17 06:49:48,145 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/schema_columns/system-schema_columns-ib-1 (3768 bytes)\nINFO [SSTableBatchOpen:1] 2013-04-17 06:49:48,181 SSTableReader.java (line 168) Opening /var/lib/cassandra/data/system/local/system-local-ib-5 (441 bytes)\nERROR [main] 2013-04-17 06:49:48,652 CassandraDaemon.java (line 454) Exception encountered during startup\njava.lang.OutOfMemoryError: Java heap space\n\tat org.apache.cassandra.io.util.RandomAccessReader.readBytes(RandomAccessReader.java:339)\n\tat org.apache.cassandra.utils.ByteBufferUtil.read(ByteBufferUtil.java:392)\n\tat org.apache.cassandra.utils.ByteBufferUtil.readWithLength(ByteBufferUtil.java:355)\n\tat org.apache.cassandra.db.ColumnSerializer.deserializeColumnBody(ColumnSerializer.java:124)\n\tat org.apache.cassandra.db.OnDiskAtom$Serializer.deserializeFromSSTable(OnDiskAtom.java:82)\n\tat org.apache.cassandra.db.Column$1.computeNext(Column.java:73)\n\tat org.apache.cassandra.db.Column$1.computeNext(Column.java:62)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n\tat org.apache.cassandra.db.columniterator.SimpleSliceReader.computeNext(SimpleSliceReader.java:92)\n\tat org.apache.cassandra.db.columniterator.SimpleSliceReader.computeNext(SimpleSliceReader.java:36)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n\tat org.apache.cassandra.db.columniterator.SSTableSliceIterator.hasNext(SSTableSliceIterator.java:90)\n\tat org.apache.cassandra.db.columniterator.LazyColumnIterator.computeNext(LazyColumnIterator.java:82)\n\tat org.apache.cassandra.db.columniterator.LazyColumnIterator.computeNext(LazyColumnIterator.java:59)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n\tat org.apache.cassandra.db.filter.QueryFilter$2.getNext(QueryFilter.java:136)\n\tat org.apache.cassandra.db.filter.QueryFilter$2.hasNext(QueryFilter.java:119)\n\tat org.apache.cassandra.utils.MergeIterator$OneToOne.computeNext(MergeIterator.java:199)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n\tat org.apache.cassandra.db.filter.SliceQueryFilter.collectReducedColumns(SliceQueryFilter.java:131)\n\tat org.apache.cassandra.db.filter.QueryFilter.collateColumns(QueryFilter.java:101)\n\tat org.apache.cassandra.db.filter.QueryFilter.collateOnDiskAtom(QueryFilter.java:75)\n\tat org.apache.cassandra.db.RowIteratorFactory$2.getReduced(RowIteratorFactory.java:105)\n\tat org.apache.cassandra.db.RowIteratorFactory$2.getReduced(RowIteratorFactory.java:78)\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.consume(MergeIterator.java:114)\n\tat org.apache.cassandra.utils.MergeIterator$ManyToOne.computeNext(MergeIterator.java:97)\n\tat com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)\n\tat com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)\n{code}\n\nI'll try to take a look today.\n\n","from":"developer"},{"body":"I can confirm that creating same table under trunk, and relaunching, works correctly, as well as basic select querying. So I think it's something in the upgrade path from 1.2 to trunk.","from":"developer"},{"body":"Just more info: tried my 1.2 to trunk upgrade again, and it's not related to my table at all. I simply started 1.2, let it create the system tables, killed it, and launched trunk - same exception. Will check it out.","from":"developer"},{"body":"Rebased on top of CASSANDRA-5511 with a fix or two: http://github.com/jbellis/cassandra/tree/4180-4\n\nNow passes the \"start 1.2, rebuild and start 4180-4\" test.","from":"developer"},{"body":"Rebased again, to https://github.com/jbellis/cassandra/commits/4180-5. [~krummas], can you review?","from":"developer"},{"body":"Minor rebase to https://github.com/jbellis/cassandra/commits/4180-6","from":"developer"},{"body":"in general, lgtm, few issues;\n\nsstable.getPosition(..) can return null in a couple of cases:\nDuring repair:\n{code}\nERROR [ValidationExecutor:2] 2013-05-07 10:43:24,875 CassandraDaemon.java (line 179) Exception in thread Thread[ValidationExecutor:2,1,main]\njava.lang.NullPointerException\n at org.apache.cassandra.io.sstable.SSTableScanner.(SSTableScanner.java:109)\n at org.apache.cassandra.io.sstable.SSTableReader.getScanner(SSTableReader.java:1038)\n at org.apache.cassandra.db.compaction.AbstractCompactionStrategy.getScanners(AbstractCompactionStrategy.java:214)\n at org.apache.cassandra.db.compaction.CompactionManager$ValidationCompactionIterable.(CompactionManager.java:744)\n at org.apache.cassandra.db.compaction.CompactionManager.doValidationCompaction(CompactionManager.java:648)\n at org.apache.cassandra.db.compaction.CompactionManager.access$600(CompactionManager.java:64)\n at org.apache.cassandra.db.compaction.CompactionManager$8.call(CompactionManager.java:391)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n at java.util.concurrent.FutureTask.run(FutureTask.java:166)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n at java.lang.Thread.run(Thread.java:722)\n{code}\n\nand scrub is not supposed to work until CASSANDRA-5429 right?:\n{code}\n INFO 09:26:43,968 Retrying from row index; data is 261 bytes starting at 295\n WARN 09:26:43,969 Retry failed too. Skipping to next row (retry's stacktrace follows)\njava.lang.AssertionError: Interval min > max\n at org.apache.cassandra.utils.IntervalTree$IntervalNode.(IntervalTree.java:250)\n at org.apache.cassandra.utils.IntervalTree.(IntervalTree.java:72)\n at org.apache.cassandra.utils.IntervalTree.build(IntervalTree.java:81)\n at org.apache.cassandra.db.DeletionInfo.add(DeletionInfo.java:179)\n at org.apache.cassandra.db.AbstractThreadUnsafeSortedColumns.delete(AbstractThreadUnsafeSortedColumns.java:44)\n at org.apache.cassandra.db.ColumnFamily.addAtom(ColumnFamily.java:142)\n at org.apache.cassandra.db.ColumnFamilySerializer.deserializeColumnsFromSSTable(ColumnFamilySerializer.java:177)\n at org.apache.cassandra.io.sstable.SSTableIdentityIterator.getColumnFamilyWithColumns(SSTableIdentityIterator.java:184)\n at org.apache.cassandra.db.compaction.PrecompactedRow.merge(PrecompactedRow.java:114)\n at org.apache.cassandra.db.compaction.PrecompactedRow.(PrecompactedRow.java:98)\n at org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:220)\n at org.apache.cassandra.db.compaction.CompactionController.getCompactedRow(CompactionController.java:226)\n at org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:204)\n at org.apache.cassandra.db.compaction.CompactionManager.scrubOne(CompactionManager.java:430)\n at org.apache.cassandra.db.compaction.CompactionManager.doScrub(CompactionManager.java:419)\n at org.apache.cassandra.db.compaction.CompactionManager.access$300(CompactionManager.java:64)\n at org.apache.cassandra.db.compaction.CompactionManager$3.perform(CompactionManager.java:234)\n at org.apache.cassandra.db.compaction.CompactionManager$2.call(CompactionManager.java:220)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n at java.util.concurrent.FutureTask.run(FutureTask.java:166)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n at java.lang.Thread.run(Thread.java:722)\n WARN 09:26:43,972 Non-fatal error reading row (stacktrace follows)\njava.io.IOError: java.io.IOException: Impossible row size 9223372034707292160\n at org.apache.cassandra.db.compaction.Scrubber.scrub(Scrubber.java:170)\n at org.apache.cassandra.db.compaction.CompactionManager.scrubOne(CompactionManager.java:430)\n at org.apache.cassandra.db.compaction.CompactionManager.doScrub(CompactionManager.java:419)\n at org.apache.cassandra.db.compaction.CompactionManager.access$300(CompactionManager.java:64)\n at org.apache.cassandra.db.compaction.CompactionManager$3.perform(CompactionManager.java:234)\n at org.apache.cassandra.db.compaction.CompactionManager$2.call(CompactionManager.java:220)\n at java.util.concurrent.FutureTask$Sync.innerRun(FutureTask.java:334)\n at java.util.concurrent.FutureTask.run(FutureTask.java:166)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1110)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:603)\n at java.lang.Thread.run(Thread.java:722)\nCaused by: java.io.IOException: Impossible row size 9223372034707292160\n ... 11 more\n{code}\n\nnits;\n* version argument unused in IndexedSliceReader#setToRowStart\n* LazilyCompactedRow comment obsolete\n* compactionType unused param in SSTableWriter#createWriter\n","from":"developer"},{"body":"Pushed fixes to -6 branch. Yes, scrub is expected to be totally broken at this point.","from":"developer"},{"body":"ok, this lgtm!","from":"developer"},{"body":"committed!","from":"developer"}],"created":"2012-04-20T15:00:36.000+0000","description":"LazilyCompactedRow reads all data twice to compact a row which is obviously inefficient. The main reason we do that is to compute the row header. However, CASSANDRA-2319 have removed the main part of that row header. What remains is the size in bytes and the number of columns, but it should be relatively simple to remove those, which would then remove the need for the two-phase compaction.","issue_id":"12551799","key":"CASSANDRA-4180","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-05-07T19:22:45.000+0000","role":"gold_target","summary":"Single-pass compaction for LCR"} {"case_id":"12630222","cluster":"JIRA-CASSANDRA-c5429e9e9c65","comments":[{"body":"It tells you what the problem is:\n\nCaused by: java.lang.RuntimeException: Unable to read cassandra-rackdc.properties\n","created":"2013-01-31T23:34:37.207+0000"},{"body":"Thanks - I didn't see that needle in the (hay) stack. It looks like as a result of CASSANDRA-5155 the file cassandra-rackdc.properties is now required to be in /etc/cassandra. However, this file is not distributed with the standard Debian Cassandra package. Maybe this should be documented in cassandra.yaml or the file should be distributed as part of the Debian package? Thanks for your help.","created":"2013-01-31T23:51:17.275+0000"},{"body":"Hmm, yes. We should make it optional and/or add to the debian package.","created":"2013-01-31T23:54:11.511+0000"},{"body":"+1, but instead of silently swallowing the exception maybe a WARN would be more appropriate.","created":"2013-02-01T12:42:04.608+0000"},{"body":"I agree with @driftx that we should add the WARN statement. Otherwise, +1","created":"2013-02-01T13:02:36.226+0000"},{"body":"Committed to 1.1, 1.2 and trunk, Thanks Brandon...","created":"2013-02-01T20:38:28.596+0000"}],"conversations":[{"body":"Hello, we're using vanilla Cassandra in an EC2 environment and I just tested upgrading from 1.2.0 to 1.2.1 on two test instances. Cassandra fails to start because it's unable to load the Ec2Snitch. Version 1.2.0 was working OK. I have tried this on uninitialized/empty instances and received the same result. Cassandra successfully starts when switching to SimpleSnitch. Log output is below. We're using the official debian package from apache.org. Please let me know if you need any more details, thanks!\n\noutput.log\n{code}\nERROR 21:25:06,684 Fatal configuration error\norg.apache.cassandra.exceptions.ConfigurationException: Error instantiating snitch class 'org.apache.cassandra.locator.Ec2Snitch'.\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:475)\n\tat org.apache.cassandra.config.DatabaseDescriptor.createEndpointSnitch(DatabaseDescriptor.java:525)\n\tat org.apache.cassandra.config.DatabaseDescriptor.loadYaml(DatabaseDescriptor.java:338)\n\tat org.apache.cassandra.config.DatabaseDescriptor.(DatabaseDescriptor.java:122)\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:151)\n\tat org.apache.cassandra.service.CassandraDaemon.init(CassandraDaemon.java:315)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.commons.daemon.support.DaemonLoader.load(DaemonLoader.java:212)\nCaused by: java.lang.reflect.InvocationTargetException\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:532)\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:457)\n\t... 10 more\nCaused by: java.lang.ExceptionInInitializerError\n\tat org.apache.cassandra.locator.Ec2Snitch.(Ec2Snitch.java:65)\n\t... 15 more\nCaused by: java.lang.RuntimeException: Unable to read cassandra-rackdc.properties\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:39)\n\t... 16 more\nCaused by: java.lang.NullPointerException\n\tat java.util.Properties$LineReader.readLine(Properties.java:435)\n\tat java.util.Properties.load0(Properties.java:354)\n\tat java.util.Properties.load(Properties.java:342)\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:35)\n\t... 16 more\nError instantiating snitch class 'org.apache.cassandra.locator.Ec2Snitch'.\nFatal configuration error; unable to start server. See log for stacktrace.\nService exit with a return value of 1\n{code}\n\nsystem.log\n{code}\nINFO [main] 2013-01-31 21:25:05,016 CassandraDaemon.java (line 123) JVM vendor/version: OpenJDK 64-Bit Server VM/1.6.0_24\n\nERROR [main] 2013-01-31 21:24:52,028 DatabaseDescriptor.java (line 509) Fatal configuration error\norg.apache.cassandra.exceptions.ConfigurationException: Error instantiating snitch class 'org.apache.cassandra.locator.Ec2Snitch'.\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:475)\n\tat org.apache.cassandra.config.DatabaseDescriptor.createEndpointSnitch(DatabaseDescriptor.java:525)\n\tat org.apache.cassandra.config.DatabaseDescriptor.loadYaml(DatabaseDescriptor.java:338)\n\tat org.apache.cassandra.config.DatabaseDescriptor.(DatabaseDescriptor.java:122)\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:151)\n\tat org.apache.cassandra.service.CassandraDaemon.init(CassandraDaemon.java:315)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.commons.daemon.support.DaemonLoader.load(DaemonLoader.java:212)\nCaused by: java.lang.reflect.InvocationTargetException\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:532)\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:457)\n\t... 10 more\nCaused by: java.lang.ExceptionInInitializerError\n\tat org.apache.cassandra.locator.Ec2Snitch.(Ec2Snitch.java:65)\n\t... 15 more\nCaused by: java.lang.RuntimeException: Unable to read cassandra-rackdc.properties\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:39)\n\t... 16 more\nCaused by: java.lang.NullPointerException\n\tat java.util.Properties$LineReader.readLine(Properties.java:435)\n\tat java.util.Properties.load0(Properties.java:354)\n\tat java.util.Properties.load(Properties.java:342)\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:35)\n\t... 16 more\n\n INFO [main] 2013-01-31 21:25:06,656 DatabaseDescriptor.java (line 267) Global memtable threshold is enabled at 407MB\nERROR [main] 2013-01-31 21:25:06,684 DatabaseDescriptor.java (line 509) Fatal configuration error\norg.apache.cassandra.exceptions.ConfigurationException: Error instantiating snitch class 'org.apache.cassandra.locator.Ec2Snitch'.\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:475)\n\tat org.apache.cassandra.config.DatabaseDescriptor.createEndpointSnitch(DatabaseDescriptor.java:525)\n\tat org.apache.cassandra.config.DatabaseDescriptor.loadYaml(DatabaseDescriptor.java:338)\n\tat org.apache.cassandra.config.DatabaseDescriptor.(DatabaseDescriptor.java:122)\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:151)\n\tat org.apache.cassandra.service.CassandraDaemon.init(CassandraDaemon.java:315)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.commons.daemon.support.DaemonLoader.load(DaemonLoader.java:212)\nCaused by: java.lang.reflect.InvocationTargetException\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:532)\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:457)\n\t... 10 more\nCaused by: java.lang.ExceptionInInitializerError\n\tat org.apache.cassandra.locator.Ec2Snitch.(Ec2Snitch.java:65)\n\t... 15 more\nCaused by: java.lang.RuntimeException: Unable to read cassandra-rackdc.properties\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:39)\n\t... 16 more\nCaused by: java.lang.NullPointerException\n\tat java.util.Properties$LineReader.readLine(Properties.java:435)\n\tat java.util.Properties.load0(Properties.java:354)\n\tat java.util.Properties.load(Properties.java:342)\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:35)\n\t... 16 more\n{code}","from":"reporter","subject":"Unable to start when using Ec2Snitch"},{"body":"It tells you what the problem is:\n\nCaused by: java.lang.RuntimeException: Unable to read cassandra-rackdc.properties\n","from":"developer"},{"body":"Thanks - I didn't see that needle in the (hay) stack. It looks like as a result of CASSANDRA-5155 the file cassandra-rackdc.properties is now required to be in /etc/cassandra. However, this file is not distributed with the standard Debian Cassandra package. Maybe this should be documented in cassandra.yaml or the file should be distributed as part of the Debian package? Thanks for your help.","from":"developer"},{"body":"Hmm, yes. We should make it optional and/or add to the debian package.","from":"developer"},{"body":"+1, but instead of silently swallowing the exception maybe a WARN would be more appropriate.","from":"developer"},{"body":"I agree with @driftx that we should add the WARN statement. Otherwise, +1","from":"developer"},{"body":"Committed to 1.1, 1.2 and trunk, Thanks Brandon...","from":"developer"}],"created":"2013-01-31T21:37:28.000+0000","description":"Hello, we're using vanilla Cassandra in an EC2 environment and I just tested upgrading from 1.2.0 to 1.2.1 on two test instances. Cassandra fails to start because it's unable to load the Ec2Snitch. Version 1.2.0 was working OK. I have tried this on uninitialized/empty instances and received the same result. Cassandra successfully starts when switching to SimpleSnitch. Log output is below. We're using the official debian package from apache.org. Please let me know if you need any more details, thanks!\n\noutput.log\n{code}\nERROR 21:25:06,684 Fatal configuration error\norg.apache.cassandra.exceptions.ConfigurationException: Error instantiating snitch class 'org.apache.cassandra.locator.Ec2Snitch'.\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:475)\n\tat org.apache.cassandra.config.DatabaseDescriptor.createEndpointSnitch(DatabaseDescriptor.java:525)\n\tat org.apache.cassandra.config.DatabaseDescriptor.loadYaml(DatabaseDescriptor.java:338)\n\tat org.apache.cassandra.config.DatabaseDescriptor.(DatabaseDescriptor.java:122)\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:151)\n\tat org.apache.cassandra.service.CassandraDaemon.init(CassandraDaemon.java:315)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.commons.daemon.support.DaemonLoader.load(DaemonLoader.java:212)\nCaused by: java.lang.reflect.InvocationTargetException\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:532)\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:457)\n\t... 10 more\nCaused by: java.lang.ExceptionInInitializerError\n\tat org.apache.cassandra.locator.Ec2Snitch.(Ec2Snitch.java:65)\n\t... 15 more\nCaused by: java.lang.RuntimeException: Unable to read cassandra-rackdc.properties\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:39)\n\t... 16 more\nCaused by: java.lang.NullPointerException\n\tat java.util.Properties$LineReader.readLine(Properties.java:435)\n\tat java.util.Properties.load0(Properties.java:354)\n\tat java.util.Properties.load(Properties.java:342)\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:35)\n\t... 16 more\nError instantiating snitch class 'org.apache.cassandra.locator.Ec2Snitch'.\nFatal configuration error; unable to start server. See log for stacktrace.\nService exit with a return value of 1\n{code}\n\nsystem.log\n{code}\nINFO [main] 2013-01-31 21:25:05,016 CassandraDaemon.java (line 123) JVM vendor/version: OpenJDK 64-Bit Server VM/1.6.0_24\n\nERROR [main] 2013-01-31 21:24:52,028 DatabaseDescriptor.java (line 509) Fatal configuration error\norg.apache.cassandra.exceptions.ConfigurationException: Error instantiating snitch class 'org.apache.cassandra.locator.Ec2Snitch'.\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:475)\n\tat org.apache.cassandra.config.DatabaseDescriptor.createEndpointSnitch(DatabaseDescriptor.java:525)\n\tat org.apache.cassandra.config.DatabaseDescriptor.loadYaml(DatabaseDescriptor.java:338)\n\tat org.apache.cassandra.config.DatabaseDescriptor.(DatabaseDescriptor.java:122)\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:151)\n\tat org.apache.cassandra.service.CassandraDaemon.init(CassandraDaemon.java:315)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.commons.daemon.support.DaemonLoader.load(DaemonLoader.java:212)\nCaused by: java.lang.reflect.InvocationTargetException\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:532)\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:457)\n\t... 10 more\nCaused by: java.lang.ExceptionInInitializerError\n\tat org.apache.cassandra.locator.Ec2Snitch.(Ec2Snitch.java:65)\n\t... 15 more\nCaused by: java.lang.RuntimeException: Unable to read cassandra-rackdc.properties\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:39)\n\t... 16 more\nCaused by: java.lang.NullPointerException\n\tat java.util.Properties$LineReader.readLine(Properties.java:435)\n\tat java.util.Properties.load0(Properties.java:354)\n\tat java.util.Properties.load(Properties.java:342)\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:35)\n\t... 16 more\n\n INFO [main] 2013-01-31 21:25:06,656 DatabaseDescriptor.java (line 267) Global memtable threshold is enabled at 407MB\nERROR [main] 2013-01-31 21:25:06,684 DatabaseDescriptor.java (line 509) Fatal configuration error\norg.apache.cassandra.exceptions.ConfigurationException: Error instantiating snitch class 'org.apache.cassandra.locator.Ec2Snitch'.\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:475)\n\tat org.apache.cassandra.config.DatabaseDescriptor.createEndpointSnitch(DatabaseDescriptor.java:525)\n\tat org.apache.cassandra.config.DatabaseDescriptor.loadYaml(DatabaseDescriptor.java:338)\n\tat org.apache.cassandra.config.DatabaseDescriptor.(DatabaseDescriptor.java:122)\n\tat org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:151)\n\tat org.apache.cassandra.service.CassandraDaemon.init(CassandraDaemon.java:315)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:616)\n\tat org.apache.commons.daemon.support.DaemonLoader.load(DaemonLoader.java:212)\nCaused by: java.lang.reflect.InvocationTargetException\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n\tat sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:57)\n\tat sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n\tat java.lang.reflect.Constructor.newInstance(Constructor.java:532)\n\tat org.apache.cassandra.utils.FBUtilities.construct(FBUtilities.java:457)\n\t... 10 more\nCaused by: java.lang.ExceptionInInitializerError\n\tat org.apache.cassandra.locator.Ec2Snitch.(Ec2Snitch.java:65)\n\t... 15 more\nCaused by: java.lang.RuntimeException: Unable to read cassandra-rackdc.properties\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:39)\n\t... 16 more\nCaused by: java.lang.NullPointerException\n\tat java.util.Properties$LineReader.readLine(Properties.java:435)\n\tat java.util.Properties.load0(Properties.java:354)\n\tat java.util.Properties.load(Properties.java:342)\n\tat org.apache.cassandra.locator.SnitchProperties.(SnitchProperties.java:35)\n\t... 16 more\n{code}","issue_id":"12630222","key":"CASSANDRA-5212","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2013-02-01T20:38:28.000+0000","role":"gold_target","summary":"Unable to start when using Ec2Snitch"} {"case_id":"12690431","cluster":"JIRA-CASSANDRA-fa595c03c0ec","comments":[{"body":"There seem to be some issues with indexing on fields that are part of your primary key (see also https://issues.apache.org/jira/browse/CASSANDRA-6470). You may want to try revising your schema and/or query to remove the need to index key columns.\n\nI was also able to reproduce this error in cassandra 2.0.1 and 2.0.4 using cqlsh or the java driver (I tried two versions of the 2.0.0rc, reproduction code below).\n\npublic static void main(String[] args) {\n\tcom.datastax.driver.core.Session session = null;\n\ttry {\n\t\tString address = \"localhost\";\n\t\tsession = new com.datastax.driver.core.Cluster.Builder().addContactPoint(address).build().connect();\n\t\t//create keyspace, table, and indices\n\t\tsession.execute(\"create keyspace if not exists testing with replication = { 'class':'SimpleStrategy', 'replication_factor':1 }\");\n\t\tsession.execute(\"drop table if exists testing.timerangetest\");\n\t\t// testing queries by time range as described in https://stackoverflow.com/questions/4667040/storing-time-ranges-in-cassandra\n\t\t//creating the table using \"i_end\" as a partition key reproduces https://issues.apache.org/jira/browse/CASSANDRA-6612\n\t\tsession.execute(\"create table if not exists testing.timerangetest (id text, end timestamp, i_eq_dummy blob, i_start timestamp, i_end timestamp, primary key ((id, end)))\");\n\t\t//creating the table using \"i_end\" as a cluster key reproduces https://issues.apache.org/jira/browse/CASSANDRA-6470\n\t\t//session.execute(\"create table if not exists testing.timerangetest (id text, i_eq_dummy blob, i_start timestamp, i_end timestamp, primary key (id, i_end))\");\n\n\t\tsession.execute(\"create index if not exists on testing.timerangetest (i_start)\");\n\t\tsession.execute(\"create index if not exists on testing.timerangetest (i_end)\");\n\t\tsession.execute(\"create index if not exists on testing.timerangetest (i_eq_dummy)\");\n\t\t//insert some values\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row1', 5, 0x00, 1, 5)\");\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row2', 13, 0x00, 3, 13)\");\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row3', 17, 0x00, 12, 17)\");\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row4', 22, 0x00, 16, 22)\");\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row5', 24, 0x00, 21, 24)\");\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row6', 23, 0x00, 4, 23)\");\n\t\t//query for everything\n\t\tSystem.out.println(\"all records:\");\n\t\tfor (com.datastax.driver.core.Row r : session.execute(\"select * from testing.timerangetest\").all()) {\n\t\t\tSystem.out.println(\" \" + r.getString(\"id\") + \" @ \" + r.getDate(\"i_start\").getTime() + \"-\" + r.getDate(\"i_end\").getTime());\n\t\t}\n\t\t//query for records that were active between 10 and 20\n\t\tSystem.out.println(\"active records from 10-20:\");\n\t\tfor (com.datastax.driver.core.Row r : session.execute(\"select * from testing.timerangetest where i_eq_dummy=0x00 and i_end >= 10 and i_start <= 20 allow filtering\").all()) {\n\t\t\tSystem.out.println(\" \" + r.getString(\"id\") + \" @ \" + r.getDate(\"i_start\").getTime() + \"-\" + r.getDate(\"i_end\").getTime());\n\t\t}\n\t\t//query for records that were active between 10 and 20 where id='row2'\n\t\tSystem.out.println(\"active records from 10-20 where id='row2' and end=13:\");\n\t\tfor (com.datastax.driver.core.Row r : session.execute(\"select * from testing.timerangetest where id='row2' and end=13 and i_eq_dummy=0x00 and i_end >= 10 and i_start <= 20 allow filtering\").all()) {\n\t\t\tSystem.out.println(\" \" + r.getString(\"id\") + \" @ \" + r.getDate(\"i_start\").getTime() + \"-\" + r.getDate(\"i_end\").getTime());\n\t\t}\n\t} finally {\n\t\tif (session != null) {\n\t\t\tcom.datastax.driver.core.Cluster c = session.getCluster();\n\t\t\tsession.close();\n\t\t\tc.close();\n\t\t}\n\t}\n}","created":"2014-01-31T14:07:38.190+0000"},{"body":"[~enigmacurry] let's see if we can reproduce this in 2.0/2.1 HEAD with cqlsh and the latest java-driver, as above.","created":"2014-07-29T21:27:15.474+0000"},{"body":"[~mshuler] I do not see this error on 2.0 HEAD with cqlsh, but the query instead hangs forever, even with only one row of data inserted/","created":"2014-08-01T18:12:16.917+0000"},{"body":"I do am able to reproduce with the example of the description (which I've pushed as a dtest).\n\nThere is actually two small problem in SecondaryIndexManager. The first one is in the {{hasIndexFor}} that clearly have a bogus logic when there is more than one searcher. It's this error that made the code take the wrong path and trigger the assertion. That said, once fixed, we still have a problem: the {{search}} is unhappy because we have more than one searcher. And indeed, we've never supported more than one searcher because even if we have multiple indexes, we should still have only one searcher for all of them (though we may have other searcher for custom indexes). That \"grouping\" of indexes is done by {{getIndexSearchersForQuery}}. However, that method was grouping using the class name, which didn't result in all internal index sharing the same searcher since we have different concrete implementation depending on whether the index is on a partition key column, clustering key one or regular one. Anyway, attaching a slightly hackish but simple solution that just make sure we group all internal indexes properly in {{getIndexSearchersForQuery}} as we should. We should probably overhaul the {{SecondaryIndexManager}} class at some point because it's pretty confusing imo but that's not for this ticket.","created":"2014-08-05T15:21:54.869+0000"},{"body":"[~beobal] to review","created":"2014-08-05T21:01:38.582+0000"},{"body":"I agree that SIM is in dire need of an overhaul and because of that this is a slightly hackish solution. That aside, the patch looks good to me. Just one thing to note, the bogus logic in SIM.hasIndexFor is already fixed in 2.1 by CASSANDRA-7525","created":"2014-08-07T15:11:26.842+0000"},{"body":"Committed, thanks","created":"2014-08-07T16:37:00.648+0000"}],"conversations":[{"body":"I am trying out Cassandra for the first time and running it locally for simple session management db. [Cassandra-2.0.4, CQL3, datastax driver 2.0.0-rc2]\n\nThe following count query works fine when there is no data in the table:\n{code}\nselect count(*) from session_data where app_name=? and account=? and last_access > ?\n{code}\n\nBut after even a single row is inserted into the table, the query fails with the following error:\n{code}\n java.lang.AssertionError\n\tat org.apache.cassandra.db.filter.ExtendedFilter$WithClauses.getExtraFilter(ExtendedFilter.java:258)\n\tat org.apache.cassandra.db.ColumnFamilyStore.filter(ColumnFamilyStore.java:1719)\n\tat org.apache.cassandra.db.ColumnFamilyStore.getRangeSlice(ColumnFamilyStore.java:1674)\n\tat org.apache.cassandra.db.PagedRangeCommand.executeLocally(PagedRangeCommand.java:111)\n\tat org.apache.cassandra.service.StorageProxy$LocalRangeSliceRunnable.runMayThrow(StorageProxy.java:1418)\n\tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1931)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:744)\n{code}\n\nHere is the schema I am using:\n\n{code}\n CREATE KEYSPACE session WITH replication= {'class': 'SimpleStrategy', 'replication_factor': 1};\n\n CREATE TABLE session_data (\n username text,\n session_id text,\n app_name text,\n account text,\n last_access timestamp,\n created_on timestamp,\n PRIMARY KEY (username, session_id, app_name, account)\n );\n\n create index sessionIndex ON session_data (session_id);\n create index sessionAppName ON session_data (app_name);\n create index lastAccessIndex ON session_data (last_access);\n{code}\n","from":"reporter","subject":"Query failing due to AssertionError"},{"body":"There seem to be some issues with indexing on fields that are part of your primary key (see also https://issues.apache.org/jira/browse/CASSANDRA-6470). You may want to try revising your schema and/or query to remove the need to index key columns.\n\nI was also able to reproduce this error in cassandra 2.0.1 and 2.0.4 using cqlsh or the java driver (I tried two versions of the 2.0.0rc, reproduction code below).\n\npublic static void main(String[] args) {\n\tcom.datastax.driver.core.Session session = null;\n\ttry {\n\t\tString address = \"localhost\";\n\t\tsession = new com.datastax.driver.core.Cluster.Builder().addContactPoint(address).build().connect();\n\t\t//create keyspace, table, and indices\n\t\tsession.execute(\"create keyspace if not exists testing with replication = { 'class':'SimpleStrategy', 'replication_factor':1 }\");\n\t\tsession.execute(\"drop table if exists testing.timerangetest\");\n\t\t// testing queries by time range as described in https://stackoverflow.com/questions/4667040/storing-time-ranges-in-cassandra\n\t\t//creating the table using \"i_end\" as a partition key reproduces https://issues.apache.org/jira/browse/CASSANDRA-6612\n\t\tsession.execute(\"create table if not exists testing.timerangetest (id text, end timestamp, i_eq_dummy blob, i_start timestamp, i_end timestamp, primary key ((id, end)))\");\n\t\t//creating the table using \"i_end\" as a cluster key reproduces https://issues.apache.org/jira/browse/CASSANDRA-6470\n\t\t//session.execute(\"create table if not exists testing.timerangetest (id text, i_eq_dummy blob, i_start timestamp, i_end timestamp, primary key (id, i_end))\");\n\n\t\tsession.execute(\"create index if not exists on testing.timerangetest (i_start)\");\n\t\tsession.execute(\"create index if not exists on testing.timerangetest (i_end)\");\n\t\tsession.execute(\"create index if not exists on testing.timerangetest (i_eq_dummy)\");\n\t\t//insert some values\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row1', 5, 0x00, 1, 5)\");\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row2', 13, 0x00, 3, 13)\");\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row3', 17, 0x00, 12, 17)\");\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row4', 22, 0x00, 16, 22)\");\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row5', 24, 0x00, 21, 24)\");\n\t\tsession.execute(\"insert into testing.timerangetest (id, end, i_eq_dummy, i_start, i_end) values ('row6', 23, 0x00, 4, 23)\");\n\t\t//query for everything\n\t\tSystem.out.println(\"all records:\");\n\t\tfor (com.datastax.driver.core.Row r : session.execute(\"select * from testing.timerangetest\").all()) {\n\t\t\tSystem.out.println(\" \" + r.getString(\"id\") + \" @ \" + r.getDate(\"i_start\").getTime() + \"-\" + r.getDate(\"i_end\").getTime());\n\t\t}\n\t\t//query for records that were active between 10 and 20\n\t\tSystem.out.println(\"active records from 10-20:\");\n\t\tfor (com.datastax.driver.core.Row r : session.execute(\"select * from testing.timerangetest where i_eq_dummy=0x00 and i_end >= 10 and i_start <= 20 allow filtering\").all()) {\n\t\t\tSystem.out.println(\" \" + r.getString(\"id\") + \" @ \" + r.getDate(\"i_start\").getTime() + \"-\" + r.getDate(\"i_end\").getTime());\n\t\t}\n\t\t//query for records that were active between 10 and 20 where id='row2'\n\t\tSystem.out.println(\"active records from 10-20 where id='row2' and end=13:\");\n\t\tfor (com.datastax.driver.core.Row r : session.execute(\"select * from testing.timerangetest where id='row2' and end=13 and i_eq_dummy=0x00 and i_end >= 10 and i_start <= 20 allow filtering\").all()) {\n\t\t\tSystem.out.println(\" \" + r.getString(\"id\") + \" @ \" + r.getDate(\"i_start\").getTime() + \"-\" + r.getDate(\"i_end\").getTime());\n\t\t}\n\t} finally {\n\t\tif (session != null) {\n\t\t\tcom.datastax.driver.core.Cluster c = session.getCluster();\n\t\t\tsession.close();\n\t\t\tc.close();\n\t\t}\n\t}\n}","from":"developer"},{"body":"[~enigmacurry] let's see if we can reproduce this in 2.0/2.1 HEAD with cqlsh and the latest java-driver, as above.","from":"developer"},{"body":"[~mshuler] I do not see this error on 2.0 HEAD with cqlsh, but the query instead hangs forever, even with only one row of data inserted/","from":"developer"},{"body":"I do am able to reproduce with the example of the description (which I've pushed as a dtest).\n\nThere is actually two small problem in SecondaryIndexManager. The first one is in the {{hasIndexFor}} that clearly have a bogus logic when there is more than one searcher. It's this error that made the code take the wrong path and trigger the assertion. That said, once fixed, we still have a problem: the {{search}} is unhappy because we have more than one searcher. And indeed, we've never supported more than one searcher because even if we have multiple indexes, we should still have only one searcher for all of them (though we may have other searcher for custom indexes). That \"grouping\" of indexes is done by {{getIndexSearchersForQuery}}. However, that method was grouping using the class name, which didn't result in all internal index sharing the same searcher since we have different concrete implementation depending on whether the index is on a partition key column, clustering key one or regular one. Anyway, attaching a slightly hackish but simple solution that just make sure we group all internal indexes properly in {{getIndexSearchersForQuery}} as we should. We should probably overhaul the {{SecondaryIndexManager}} class at some point because it's pretty confusing imo but that's not for this ticket.","from":"developer"},{"body":"[~beobal] to review","from":"developer"},{"body":"I agree that SIM is in dire need of an overhaul and because of that this is a slightly hackish solution. That aside, the patch looks good to me. Just one thing to note, the bogus logic in SIM.hasIndexFor is already fixed in 2.1 by CASSANDRA-7525","from":"developer"},{"body":"Committed, thanks","from":"developer"}],"created":"2014-01-22T23:20:18.000+0000","description":"I am trying out Cassandra for the first time and running it locally for simple session management db. [Cassandra-2.0.4, CQL3, datastax driver 2.0.0-rc2]\n\nThe following count query works fine when there is no data in the table:\n{code}\nselect count(*) from session_data where app_name=? and account=? and last_access > ?\n{code}\n\nBut after even a single row is inserted into the table, the query fails with the following error:\n{code}\n java.lang.AssertionError\n\tat org.apache.cassandra.db.filter.ExtendedFilter$WithClauses.getExtraFilter(ExtendedFilter.java:258)\n\tat org.apache.cassandra.db.ColumnFamilyStore.filter(ColumnFamilyStore.java:1719)\n\tat org.apache.cassandra.db.ColumnFamilyStore.getRangeSlice(ColumnFamilyStore.java:1674)\n\tat org.apache.cassandra.db.PagedRangeCommand.executeLocally(PagedRangeCommand.java:111)\n\tat org.apache.cassandra.service.StorageProxy$LocalRangeSliceRunnable.runMayThrow(StorageProxy.java:1418)\n\tat org.apache.cassandra.service.StorageProxy$DroppableRunnable.run(StorageProxy.java:1931)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:744)\n{code}\n\nHere is the schema I am using:\n\n{code}\n CREATE KEYSPACE session WITH replication= {'class': 'SimpleStrategy', 'replication_factor': 1};\n\n CREATE TABLE session_data (\n username text,\n session_id text,\n app_name text,\n account text,\n last_access timestamp,\n created_on timestamp,\n PRIMARY KEY (username, session_id, app_name, account)\n );\n\n create index sessionIndex ON session_data (session_id);\n create index sessionAppName ON session_data (app_name);\n create index lastAccessIndex ON session_data (last_access);\n{code}\n","issue_id":"12690431","key":"CASSANDRA-6612","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-08-07T16:37:00.000+0000","role":"gold_target","summary":"Query failing due to AssertionError"} {"case_id":"12746343","cluster":"JIRA-CASSANDRA-d3ad58e022d0","comments":[{"body":"This should be fixed in 2.1.1. (CASSANDRA-7784)","created":"2014-10-24T15:49:15.341+0000"},{"body":"[~yukim] our production cluster been freshly migrated from 2.0.11 to 2.1.1 yesterday. Everything is running well except for this error which is happening on every cluster node after migration:\n\n{code}\nERROR [CompactionExecutor:1] 2014-11-07 07:54:24,599 CassandraDaemon.java:153 - Exception in thread Thread[CompactionExecutor:1,1,main]\njava.lang.NullPointerException: null\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.1.jar:2.1.1]\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.1.jar:2.1.1]\n at org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:236) ~[apache-cassandra-2.1.1.jar:2.1.1]\n at org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1089) ~[apache-cassandra-2.1.1.jar:2.1.1]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) ~[na:1.7.0_72]\n at java.util.concurrent.FutureTask.run(FutureTask.java:262) ~[na:1.7.0_72]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_72]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_72]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_72]\n{code}\n\nLet me know if I can provide more info.","created":"2014-11-07T08:58:02.238+0000"},{"body":"Okay, thanks for reporting.\n\nI thought CacheService.java:475 indicates that saving key cache for table that does not exist any more, and CASSANDRA-7784 added schema existance check before serializing cache. Thus I marked this as duplicate.\n\nDo you still see the error constantly? Or just after the migration?","created":"2014-11-07T16:32:35.960+0000"},{"body":"[~yukim]\n\nyes after 2 days I can still see the error happening at least once a day on each node.\n\nNow, that the 2.1.1 cluster has been running it's been 2 days and having only this error which is not breaking our application, I am now upgrading the sstables which I purposely delayed to make sure that the cluster was stable. I will report again after sstables on all nodes are migrated to let you know if whether or not this issue is persisting.","created":"2014-11-10T09:00:17.638+0000"},{"body":"I still have the same issue running since a few weeks w/ 2.1.1 and I already upgraded the sstables on that node. There are a lot in my system.log. In 2.0.10 there was not such an issue. I still have most nodes running on 2.0.10. I'm curious for Yuki's report, if upgrading the whole cluster fixes the problem. You can see my system.log at CASSANDRA-8192.","created":"2014-11-11T08:43:22.144+0000"},{"body":"Today I got that failure during nodetool upgradesstables after upgrading from 2.0.10 to 2.1.2 w/ finalizer-patch CASSANDRA-6283\n{noformat}\nERROR [CompactionExecutor:65] 2014-11-19 09:56:02,453 CassandraDaemon.java:153 - Exception in thread Thread[CompactionExecutor:65,1,main]\njava.lang.NullPointerException: null\n\tat org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.2-SNAPSHOT.jar:2.1.2-SNAPSHOT]\n\tat org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.2-SNAPSHOT.jar:2.1.2-SNAPSHOT]\n\tat org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:274) ~[apache-cassandra-2.1.2-SNAPSHOT.jar:2.1.2-SNAPSHOT]\n\tat org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1088) ~[apache-cassandra-2.1.2-SNAPSHOT.jar:2.1.2-SNAPSHOT]\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source) ~[na:1.7.0_55]\n\tat java.util.concurrent.FutureTask.run(Unknown Source) ~[na:1.7.0_55]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [na:1.7.0_55]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [na:1.7.0_55]\n\tat java.lang.Thread.run(Unknown Source) [na:1.7.0_55]\n{noformat}","created":"2014-11-19T11:32:27.761+0000"},{"body":"reopening for investigation","created":"2014-11-19T12:31:59.350+0000"},{"body":"[~mokemokechicken] [~Andie78] error definitely still occurring in a cluster where all nodes are running 2.1.2: error happening against every node of the cluster on regular basis. The same cluster running 2.0.11 did not have that particular issue before migration.","created":"2014-11-24T20:29:27.043+0000"},{"body":"I have exact the same issue in v. 2.1.2 on all nodes in my cluster:\n\n{noformat}\nERROR [CompactionExecutor:1528] 2014-12-19 05:38:47,155 CassandraDaemon.java:153 - Exception in thread Thread[CompactionExecutor:1528,1,main]\njava.lang.NullPointerException: null\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.2.jar:2.1.2]\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.2.jar:2.1.2]\n at org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:274) ~[apache-cassandra-2.1.2.jar:2.1.2]\n at org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1088) ~[apache-cassandra-2.1.2.jar:2.1.2]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) ~[na:1.7.0_72]\n at java.util.concurrent.FutureTask.run(FutureTask.java:262) ~[na:1.7.0_72]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_72]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_72]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_72]\nINFO [CompactionExecutor:1529] 2014-12-19 05:38:47,313 AutoSavingCache.java:302 - Saved CounterCache (6149 items) in 130 ms\n{noformat}","created":"2014-12-19T09:02:38.374+0000"},{"body":"This looks like an issue with serializing data for a dropped table. We should also be looking up the cfMetadata with our UUID, else we can have some weirdness with a serialization crossing a drop/recreate boundary, which would be especially problematic with different schema.","created":"2015-01-13T16:42:46.230+0000"},{"body":"It happened once again during repair session:\n\n{noformat}\nINFO [CompactionExecutor:243] 2015-02-02 06:30:00,845 CompactionTask.java:251 - Compacted 2 sstables to []. 319 bytes to 0 (~0% of original) in 1ms = 0.000000MB/s. 2 total partitions merged to 0. Partition merge counts were {2:1, }\nERROR [CompactionExecutor:244] 2015-02-02 06:39:52,980 CassandraDaemon.java:153 - Exception in thread Thread[CompactionExecutor:244,1,main]\njava.lang.NullPointerException: null\n\tat org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.2.jar:2.1.2]\n\tat org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.2.jar:2.1.2]\n\tat org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:274) ~[apache-cassandra-2.1.2.jar:2.1.2]\n\tat org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1088) ~[apache-cassandra-2.1.2.jar:2.1.2]\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) ~[na:1.7.0_72]\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262) ~[na:1.7.0_72]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_72]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_72]\n\tat java.lang.Thread.run(Thread.java:745) [na:1.7.0_72]\n{noformat}\n\nPlease take a look especially at this weird log message before exception: \"Compacted 2 sstables to []\".","created":"2015-02-02T09:19:24.314+0000"},{"body":"This is the error I get as well, hundreds of times on each node:\n\n~~~\nERROR [CompactionExecutor:2257] 2015-02-11 06:41:37,649 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2257,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2257] 2015-02-11 07:11:11,555 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2257,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2338] 2015-02-11 10:41:38,631 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2338,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2353] 2015-02-11 14:41:38,629 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2353,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 04:57:17,884 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 04:57:29,277 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 05:23:36,772 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2295] 2015-02-12 11:23:56,165 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2295,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2380] 2015-02-12 12:57:29,437 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2380,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2362] 2015-02-12 19:27:15,902 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2362,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2317] 2015-02-12 21:35:06,668 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2317,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2332] 2015-02-13 02:28:00,758 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2332,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2381] 2015-02-13 12:33:33,254 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2381,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2392] 2015-02-13 16:40:06,309 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2392,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2307] 2015-02-13 20:26:42,196 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2307,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2386] 2015-02-14 02:52:56,053 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2386,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2386] 2015-02-14 04:16:04,030 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2386,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:30:30,397 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,166 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,341 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,496 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-15 00:42:54,965 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\n~~~\n\nAny fix around that?\n\nAs far I can see seems that I can't compact my sstables, I don't have a huge table but seems is always stuck between 1 or 2% during compaction.","created":"2015-02-16T19:28:30.907+0000"},{"body":"This is the error I get as well, hundreds of times on each node:\n\n~~~\nERROR [CompactionExecutor:2257] 2015-02-11 06:41:37,649 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2257,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2257] 2015-02-11 07:11:11,555 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2257,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2338] 2015-02-11 10:41:38,631 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2338,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2353] 2015-02-11 14:41:38,629 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2353,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 04:57:17,884 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 04:57:29,277 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 05:23:36,772 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2295] 2015-02-12 11:23:56,165 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2295,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2380] 2015-02-12 12:57:29,437 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2380,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2362] 2015-02-12 19:27:15,902 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2362,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2317] 2015-02-12 21:35:06,668 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2317,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2332] 2015-02-13 02:28:00,758 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2332,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2381] 2015-02-13 12:33:33,254 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2381,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2392] 2015-02-13 16:40:06,309 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2392,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2307] 2015-02-13 20:26:42,196 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2307,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2386] 2015-02-14 02:52:56,053 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2386,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2386] 2015-02-14 04:16:04,030 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2386,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:30:30,397 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,166 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,341 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,496 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-15 00:42:54,965 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\n~~~\n\nAny fix around that?\n\nAs far I can see seems that I can't compact my sstables, I don't have a huge table but seems is always stuck between 1 or 2% during compaction.","created":"2015-02-16T19:28:31.985+0000"},{"body":"I'm getting the same error - https://cpaste.org/pysq5smf0.","created":"2015-02-18T14:15:55.660+0000"},{"body":"We are seeing this issue as well. These logs are from a 2.1.3 node during compaction while joining a mixed 2.1.3/2.1.2 cluster.\n\n{code}\nERROR [CompactionExecutor:30] 2015-02-19 08:24:47,058 CassandraDaemon.java:167 - Exception in thread Thread[CompactionExecutor:30,1,main]\njava.lang.NullPointerException: null\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:274) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1152) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source) ~[na:1.7.0_55]\n at java.util.concurrent.FutureTask.run(Unknown Source) ~[na:1.7.0_55]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [na:1.7.0_55]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [na:1.7.0_55]\n at java.lang.Thread.run(Unknown Source) [na:1.7.0_55]\n{code}\n\nAlso (maybe) worth noting, unlike Rafal's example the sstable merge message immediately before the error does not have any immediate red flags","created":"2015-02-19T16:52:23.270+0000"},{"body":"Does this influence for you guys the compaction progress? My nodes once they trigger this error seems to be stuck at 1/2% without going forward.\n\n!http://fs.daddye.it/Cswq+!","created":"2015-02-19T23:26:41.355+0000"},{"body":"Davide, when we see this error it usually corresponds to compaction stalling, although the specific percentage seems to differ from case to case.","created":"2015-02-19T23:44:33.375+0000"},{"body":"I've increased the logging verbosity and captured the event again, there's only a single log message from the thread in question before the NullPointerException.\n\n{code}\nDEBUG [CompactionExecutor:56] 2015-02-20 02:55:13,257 AutoSavingCache.java:230 \n- Deleting old KeyCache files.\n\nERROR [CompactionExecutor:56] 2015-02-20 02:55:14,314 CassandraDaemon.java:167 \n- Exception in thread Thread[CompactionExecutor:56,1,main]\njava.lang.NullPointerException: null\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.seriali\nze(CacheService.java:475) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.seriali\nze(CacheService.java:463) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavi\nngCache.java:274) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.db.compaction.CompactionManager$11.run(Compacti\nonManager.java:1152) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source) \n~[na:1.7.0_55]\n at java.util.concurrent.FutureTask.run(Unknown Source) ~[na:1.7.0_55]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [n\na:1.7.0_55]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [\nna:1.7.0_55]\n at java.lang.Thread.run(Unknown Source) [na:1.7.0_55]\n{code}\n\nIf there's any log in particular that would be useful for debugging, please let me know. I've stashed away the offending log file.","created":"2015-02-20T03:15:47.400+0000"},{"body":"Attaching a simple patch that deals with the raciness in KeyCacheSerializer , plus looks up the CFMetaData by the id vs ks/cf names, per Benedict's suggestion.","created":"2015-02-20T04:17:36.965+0000"},{"body":"Thanks a lot. When is planned the 2.1.4 to be released? ","created":"2015-02-21T19:46:39.070+0000"},{"body":"+1, although I think this code could do with being refactored, as there's a bit of a poor isolation of concerns - the caller and callee of CacheSerializer methods repeat much of the same work.","created":"2015-03-03T15:19:45.073+0000"},{"body":"Agreed, but hesitant to do that in 2.1.x. I'll open a separate 3.0 ticket for just that.","created":"2015-03-03T19:55:10.352+0000"},{"body":"bq. but hesitant to do that in 2.1.x\n\nAgreed","created":"2015-03-03T19:57:07.168+0000"},{"body":"[~benedict] Committed.\n\nSpeaking of refactoring - I see how it could be improved, but, on second thought, I can't find any particular duplication there (aside from checking for cf id existence, and maybe cache entry existence). Did you have anything particular in mind? If so, can you open the refactoring ticket? Thanks.","created":"2015-03-03T22:11:13.590+0000"},{"body":"Eh, it's not a burning desire to make better, so I'll leave it thanks. More than enough to do. A brief handwavy outline is that the current abstraction adds complexity by not really separating concerns; the duplication is not a problem of the cost of the work but of the cognitive burden (which is essentially why this happened in the first place). But since it's not a significant pain point, it's not worth agonising over either. ","created":"2015-03-04T00:00:11.399+0000"},{"body":"Would it be safe to patch a 2.1.3 cluster with the attatched (8067.txt) patch?","created":"2015-03-06T08:21:47.773+0000"},{"body":"Yes.","created":"2015-03-06T08:24:05.423+0000"},{"body":"Is there a way to fix the issue without the patch? When the 2.1.4 is planned to be released? ","created":"2015-03-12T17:56:13.578+0000"},{"body":"[~DAddYE] In a week or two.","created":"2015-03-13T01:05:10.287+0000"}],"conversations":[{"body":"Hi,\n\nI have this stack trace in the logs of Cassandra server (v2.1)\n\n{code}\nERROR [CompactionExecutor:14] 2014-10-06 23:32:02,098 CassandraDaemon.java:166 - Exception in thread Thread[CompactionExecutor:14,1,main]\njava.lang.NullPointerException: null\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:225) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1061) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source) ~[na:1.7.0]\n at java.util.concurrent.FutureTask$Sync.innerRun(Unknown Source) ~[na:1.7.0]\n at java.util.concurrent.FutureTask.run(Unknown Source) ~[na:1.7.0]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [na:1.7.0]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [na:1.7.0]\n at java.lang.Thread.run(Unknown Source) [na:1.7.0]\n{code}\n\nIt may not be critical because this error occured in the AutoSavingCache. However the line 475 is about the CFMetaData so it may hide bigger issue...\n\n{code}\n 474 CFMetaData cfm = Schema.instance.getCFMetaData(key.desc.ksname, key.desc.cfname);\n 475 cfm.comparator.rowIndexEntrySerializer().serialize(entry, out);\n{code}\n\nRegards,\nEric","from":"reporter","subject":"NullPointerException in KeyCacheSerializer"},{"body":"This should be fixed in 2.1.1. (CASSANDRA-7784)","from":"developer"},{"body":"[~yukim] our production cluster been freshly migrated from 2.0.11 to 2.1.1 yesterday. Everything is running well except for this error which is happening on every cluster node after migration:\n\n{code}\nERROR [CompactionExecutor:1] 2014-11-07 07:54:24,599 CassandraDaemon.java:153 - Exception in thread Thread[CompactionExecutor:1,1,main]\njava.lang.NullPointerException: null\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.1.jar:2.1.1]\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.1.jar:2.1.1]\n at org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:236) ~[apache-cassandra-2.1.1.jar:2.1.1]\n at org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1089) ~[apache-cassandra-2.1.1.jar:2.1.1]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) ~[na:1.7.0_72]\n at java.util.concurrent.FutureTask.run(FutureTask.java:262) ~[na:1.7.0_72]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_72]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_72]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_72]\n{code}\n\nLet me know if I can provide more info.","from":"developer"},{"body":"Okay, thanks for reporting.\n\nI thought CacheService.java:475 indicates that saving key cache for table that does not exist any more, and CASSANDRA-7784 added schema existance check before serializing cache. Thus I marked this as duplicate.\n\nDo you still see the error constantly? Or just after the migration?","from":"developer"},{"body":"[~yukim]\n\nyes after 2 days I can still see the error happening at least once a day on each node.\n\nNow, that the 2.1.1 cluster has been running it's been 2 days and having only this error which is not breaking our application, I am now upgrading the sstables which I purposely delayed to make sure that the cluster was stable. I will report again after sstables on all nodes are migrated to let you know if whether or not this issue is persisting.","from":"developer"},{"body":"I still have the same issue running since a few weeks w/ 2.1.1 and I already upgraded the sstables on that node. There are a lot in my system.log. In 2.0.10 there was not such an issue. I still have most nodes running on 2.0.10. I'm curious for Yuki's report, if upgrading the whole cluster fixes the problem. You can see my system.log at CASSANDRA-8192.","from":"developer"},{"body":"Today I got that failure during nodetool upgradesstables after upgrading from 2.0.10 to 2.1.2 w/ finalizer-patch CASSANDRA-6283\n{noformat}\nERROR [CompactionExecutor:65] 2014-11-19 09:56:02,453 CassandraDaemon.java:153 - Exception in thread Thread[CompactionExecutor:65,1,main]\njava.lang.NullPointerException: null\n\tat org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.2-SNAPSHOT.jar:2.1.2-SNAPSHOT]\n\tat org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.2-SNAPSHOT.jar:2.1.2-SNAPSHOT]\n\tat org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:274) ~[apache-cassandra-2.1.2-SNAPSHOT.jar:2.1.2-SNAPSHOT]\n\tat org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1088) ~[apache-cassandra-2.1.2-SNAPSHOT.jar:2.1.2-SNAPSHOT]\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source) ~[na:1.7.0_55]\n\tat java.util.concurrent.FutureTask.run(Unknown Source) ~[na:1.7.0_55]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [na:1.7.0_55]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [na:1.7.0_55]\n\tat java.lang.Thread.run(Unknown Source) [na:1.7.0_55]\n{noformat}","from":"developer"},{"body":"reopening for investigation","from":"developer"},{"body":"[~mokemokechicken] [~Andie78] error definitely still occurring in a cluster where all nodes are running 2.1.2: error happening against every node of the cluster on regular basis. The same cluster running 2.0.11 did not have that particular issue before migration.","from":"developer"},{"body":"I have exact the same issue in v. 2.1.2 on all nodes in my cluster:\n\n{noformat}\nERROR [CompactionExecutor:1528] 2014-12-19 05:38:47,155 CassandraDaemon.java:153 - Exception in thread Thread[CompactionExecutor:1528,1,main]\njava.lang.NullPointerException: null\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.2.jar:2.1.2]\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.2.jar:2.1.2]\n at org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:274) ~[apache-cassandra-2.1.2.jar:2.1.2]\n at org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1088) ~[apache-cassandra-2.1.2.jar:2.1.2]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) ~[na:1.7.0_72]\n at java.util.concurrent.FutureTask.run(FutureTask.java:262) ~[na:1.7.0_72]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_72]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_72]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_72]\nINFO [CompactionExecutor:1529] 2014-12-19 05:38:47,313 AutoSavingCache.java:302 - Saved CounterCache (6149 items) in 130 ms\n{noformat}","from":"developer"},{"body":"This looks like an issue with serializing data for a dropped table. We should also be looking up the cfMetadata with our UUID, else we can have some weirdness with a serialization crossing a drop/recreate boundary, which would be especially problematic with different schema.","from":"developer"},{"body":"It happened once again during repair session:\n\n{noformat}\nINFO [CompactionExecutor:243] 2015-02-02 06:30:00,845 CompactionTask.java:251 - Compacted 2 sstables to []. 319 bytes to 0 (~0% of original) in 1ms = 0.000000MB/s. 2 total partitions merged to 0. Partition merge counts were {2:1, }\nERROR [CompactionExecutor:244] 2015-02-02 06:39:52,980 CassandraDaemon.java:153 - Exception in thread Thread[CompactionExecutor:244,1,main]\njava.lang.NullPointerException: null\n\tat org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.2.jar:2.1.2]\n\tat org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.2.jar:2.1.2]\n\tat org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:274) ~[apache-cassandra-2.1.2.jar:2.1.2]\n\tat org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1088) ~[apache-cassandra-2.1.2.jar:2.1.2]\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) ~[na:1.7.0_72]\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262) ~[na:1.7.0_72]\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145) ~[na:1.7.0_72]\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615) [na:1.7.0_72]\n\tat java.lang.Thread.run(Thread.java:745) [na:1.7.0_72]\n{noformat}\n\nPlease take a look especially at this weird log message before exception: \"Compacted 2 sstables to []\".","from":"developer"},{"body":"This is the error I get as well, hundreds of times on each node:\n\n~~~\nERROR [CompactionExecutor:2257] 2015-02-11 06:41:37,649 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2257,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2257] 2015-02-11 07:11:11,555 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2257,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2338] 2015-02-11 10:41:38,631 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2338,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2353] 2015-02-11 14:41:38,629 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2353,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 04:57:17,884 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 04:57:29,277 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 05:23:36,772 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2295] 2015-02-12 11:23:56,165 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2295,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2380] 2015-02-12 12:57:29,437 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2380,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2362] 2015-02-12 19:27:15,902 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2362,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2317] 2015-02-12 21:35:06,668 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2317,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2332] 2015-02-13 02:28:00,758 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2332,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2381] 2015-02-13 12:33:33,254 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2381,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2392] 2015-02-13 16:40:06,309 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2392,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2307] 2015-02-13 20:26:42,196 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2307,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2386] 2015-02-14 02:52:56,053 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2386,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2386] 2015-02-14 04:16:04,030 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2386,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:30:30,397 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,166 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,341 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,496 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-15 00:42:54,965 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\n~~~\n\nAny fix around that?\n\nAs far I can see seems that I can't compact my sstables, I don't have a huge table but seems is always stuck between 1 or 2% during compaction.","from":"developer"},{"body":"This is the error I get as well, hundreds of times on each node:\n\n~~~\nERROR [CompactionExecutor:2257] 2015-02-11 06:41:37,649 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2257,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2257] 2015-02-11 07:11:11,555 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2257,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2338] 2015-02-11 10:41:38,631 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2338,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2353] 2015-02-11 14:41:38,629 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2353,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 04:57:17,884 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 04:57:29,277 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2313] 2015-02-12 05:23:36,772 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2313,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2295] 2015-02-12 11:23:56,165 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2295,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2380] 2015-02-12 12:57:29,437 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2380,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2362] 2015-02-12 19:27:15,902 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2362,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2317] 2015-02-12 21:35:06,668 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2317,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2332] 2015-02-13 02:28:00,758 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2332,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2381] 2015-02-13 12:33:33,254 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2381,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2392] 2015-02-13 16:40:06,309 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2392,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2307] 2015-02-13 20:26:42,196 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2307,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2386] 2015-02-14 02:52:56,053 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2386,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2386] 2015-02-14 04:16:04,030 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2386,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:30:30,397 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,166 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,341 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-14 23:35:20,496 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\nERROR [CompactionExecutor:2401] 2015-02-15 00:42:54,965 CassandraDaemon.java [line 153] Exception in thread Thread[CompactionExecutor:2401,1,main]\njava.lang.NullPointerException: null\n~~~\n\nAny fix around that?\n\nAs far I can see seems that I can't compact my sstables, I don't have a huge table but seems is always stuck between 1 or 2% during compaction.","from":"developer"},{"body":"I'm getting the same error - https://cpaste.org/pysq5smf0.","from":"developer"},{"body":"We are seeing this issue as well. These logs are from a 2.1.3 node during compaction while joining a mixed 2.1.3/2.1.2 cluster.\n\n{code}\nERROR [CompactionExecutor:30] 2015-02-19 08:24:47,058 CassandraDaemon.java:167 - Exception in thread Thread[CompactionExecutor:30,1,main]\njava.lang.NullPointerException: null\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:274) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1152) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source) ~[na:1.7.0_55]\n at java.util.concurrent.FutureTask.run(Unknown Source) ~[na:1.7.0_55]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [na:1.7.0_55]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [na:1.7.0_55]\n at java.lang.Thread.run(Unknown Source) [na:1.7.0_55]\n{code}\n\nAlso (maybe) worth noting, unlike Rafal's example the sstable merge message immediately before the error does not have any immediate red flags","from":"developer"},{"body":"Does this influence for you guys the compaction progress? My nodes once they trigger this error seems to be stuck at 1/2% without going forward.\n\n!http://fs.daddye.it/Cswq+!","from":"developer"},{"body":"Davide, when we see this error it usually corresponds to compaction stalling, although the specific percentage seems to differ from case to case.","from":"developer"},{"body":"I've increased the logging verbosity and captured the event again, there's only a single log message from the thread in question before the NullPointerException.\n\n{code}\nDEBUG [CompactionExecutor:56] 2015-02-20 02:55:13,257 AutoSavingCache.java:230 \n- Deleting old KeyCache files.\n\nERROR [CompactionExecutor:56] 2015-02-20 02:55:14,314 CassandraDaemon.java:167 \n- Exception in thread Thread[CompactionExecutor:56,1,main]\njava.lang.NullPointerException: null\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.seriali\nze(CacheService.java:475) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.seriali\nze(CacheService.java:463) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavi\nngCache.java:274) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at org.apache.cassandra.db.compaction.CompactionManager$11.run(Compacti\nonManager.java:1152) ~[apache-cassandra-2.1.3.jar:2.1.3]\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source) \n~[na:1.7.0_55]\n at java.util.concurrent.FutureTask.run(Unknown Source) ~[na:1.7.0_55]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [n\na:1.7.0_55]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [\nna:1.7.0_55]\n at java.lang.Thread.run(Unknown Source) [na:1.7.0_55]\n{code}\n\nIf there's any log in particular that would be useful for debugging, please let me know. I've stashed away the offending log file.","from":"developer"},{"body":"Attaching a simple patch that deals with the raciness in KeyCacheSerializer , plus looks up the CFMetaData by the id vs ks/cf names, per Benedict's suggestion.","from":"developer"},{"body":"Thanks a lot. When is planned the 2.1.4 to be released? ","from":"developer"},{"body":"+1, although I think this code could do with being refactored, as there's a bit of a poor isolation of concerns - the caller and callee of CacheSerializer methods repeat much of the same work.","from":"developer"},{"body":"Agreed, but hesitant to do that in 2.1.x. I'll open a separate 3.0 ticket for just that.","from":"developer"},{"body":"bq. but hesitant to do that in 2.1.x\n\nAgreed","from":"developer"},{"body":"[~benedict] Committed.\n\nSpeaking of refactoring - I see how it could be improved, but, on second thought, I can't find any particular duplication there (aside from checking for cf id existence, and maybe cache entry existence). Did you have anything particular in mind? If so, can you open the refactoring ticket? Thanks.","from":"developer"},{"body":"Eh, it's not a burning desire to make better, so I'll leave it thanks. More than enough to do. A brief handwavy outline is that the current abstraction adds complexity by not really separating concerns; the duplication is not a problem of the cost of the work but of the cognitive burden (which is essentially why this happened in the first place). But since it's not a significant pain point, it's not worth agonising over either. ","from":"developer"},{"body":"Would it be safe to patch a 2.1.3 cluster with the attatched (8067.txt) patch?","from":"developer"},{"body":"Yes.","from":"developer"},{"body":"Is there a way to fix the issue without the patch? When the 2.1.4 is planned to be released? ","from":"developer"},{"body":"[~DAddYE] In a week or two.","from":"developer"}],"created":"2014-10-07T08:02:42.000+0000","description":"Hi,\n\nI have this stack trace in the logs of Cassandra server (v2.1)\n\n{code}\nERROR [CompactionExecutor:14] 2014-10-06 23:32:02,098 CassandraDaemon.java:166 - Exception in thread Thread[CompactionExecutor:14,1,main]\njava.lang.NullPointerException: null\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:475) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at org.apache.cassandra.service.CacheService$KeyCacheSerializer.serialize(CacheService.java:463) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at org.apache.cassandra.cache.AutoSavingCache$Writer.saveCache(AutoSavingCache.java:225) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at org.apache.cassandra.db.compaction.CompactionManager$11.run(CompactionManager.java:1061) ~[apache-cassandra-2.1.0.jar:2.1.0]\n at java.util.concurrent.Executors$RunnableAdapter.call(Unknown Source) ~[na:1.7.0]\n at java.util.concurrent.FutureTask$Sync.innerRun(Unknown Source) ~[na:1.7.0]\n at java.util.concurrent.FutureTask.run(Unknown Source) ~[na:1.7.0]\n at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source) [na:1.7.0]\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source) [na:1.7.0]\n at java.lang.Thread.run(Unknown Source) [na:1.7.0]\n{code}\n\nIt may not be critical because this error occured in the AutoSavingCache. However the line 475 is about the CFMetaData so it may hide bigger issue...\n\n{code}\n 474 CFMetaData cfm = Schema.instance.getCFMetaData(key.desc.ksname, key.desc.cfname);\n 475 cfm.comparator.rowIndexEntrySerializer().serialize(entry, out);\n{code}\n\nRegards,\nEric","issue_id":"12746343","key":"CASSANDRA-8067","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-03-03T22:11:13.000+0000","role":"gold_target","summary":"NullPointerException in KeyCacheSerializer"} {"case_id":"12754248","cluster":"JIRA-CASSANDRA-a80341721d34","comments":[{"body":"The test {{TestCQL.order_by_with_in_test}} is failing for the same reason and began at the same commit.","created":"2014-11-10T20:51:31.416+0000"},{"body":"Despite the commit being in 2.1.x as well, the tests are passing there with no error. This is due to CASSANDRA-4911.","created":"2014-11-10T20:57:36.840+0000"},{"body":"However, in the 2.1 test {{TestCQL.in_order_by_without_selecting_test}}, the following exception has begun appearing at the same time:\n\nhttp://cassci.datastax.com/job/cassandra-2.1_dtest/lastCompletedBuild/testReport/cql_tests/TestCQL/in_order_by_without_selecting_test/history/\n\n{code}ERROR [SharedPool-Worker-3] 2014-11-10 14:58:28,591 Message.java\n:538 - Unexpected exception during request; channel = [id: 0x5aa\ne53fd, /127.0.0.1:63101 => /127.0.0.1:9042]\njava.lang.AssertionError: null\n at org.apache.cassandra.cql3.ResultSet.addRow(ResultSet.java:63) ~[main/:na]\n at org.apache.cassandra.cql3.statements.Selection$ResultSetBuilder.newRow(Selection.java:333) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.processColumnFamily(SelectStatement.java:1227) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.process(SelectStatement.java:1161) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.processResults(SelectStatement.java:290) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:267) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:215) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:64) ~[main/:na]\n at org.apache.cassandra.cql3.QueryProcessor.processStatement(QueryProcessor.java:226) ~[main/:na]\n at org.apache.cassandra.cql3.QueryProcessor.process(QueryProcessor.java:248) ~[main/:na]\n at org.apache.cassandra.transport.messages.QueryMessage.execute(QueryMessage.java:119) ~[main/:na]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:439) [main/:na]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:335) [main/:na]\n at io.netty.channel.SimpleChannelInboundHandler.channelRead(SimpleChannelInboundHandler.java:105) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:333) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext.access$700(AbstractChannelHandlerContext.java:32) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext$8.run(AbstractChannelHandlerContext.java:324) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_67]\n at org.apache.cassandra.concurrent.AbstractTracingAwareExecutorService$FutureTask.run(AbstractTracingAwareExecutorService.java:164) [main/:na]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_67]{code}","created":"2014-11-10T21:00:56.858+0000"},{"body":"In 2.0, the problem was in {{SelectStatement.processOrderingClause()}} where {{ColumnIdentifier.Raw}} instances were being compared with {{ColumnIdentifier}} instances.\n\nIn 2.1 and trunk, the behavior of {{processOrderingClause()}} was changed, so that problem didn't appear. However, there were two additional issues:\n\nFirst, in {{Selection.usingFunctions()}} simply checked {{instanceof ColumnIdentifier}}, which is false for {{ColumnIdentifier.Raw}}, causing all selections to be treated as function/field selections ({{SelectionWithFunctions}}).\n\nSecond, there was an existing bug with {{SelectionWithFunctions}}. In {{Selection.addColumnForOrdering()}}, the internal list of {{columns}} is updated, which works for {{SimpleSelection}}, but is insufficient for {{SelectionWithFunctions}}, which depends on its pre-built list of {{Selectors}}. So, if there were any functions or field selections, {{ORDER BY}} columns would not be properly added to the list of selections. Since the previously mentioned bug resulted in all selections being {{SelectionWithFunctions}}, the normal selection case was also broken.\n\n The 2.1 and trunk patches override {{addColumnForOrdering()}} in {{SelectionWithFunctions}} to update the selectors, and I pushed an [additional dtest|https://github.com/riptano/cassandra-dtest/commit/da86a8eae38179c12b3576ec3d4beaaa0750144c] to cover selecting functions with {{ORDER BY}} clauses.","created":"2014-11-12T23:15:36.140+0000"},{"body":"[~blerer] can you review?","created":"2014-11-12T23:17:13.074+0000"},{"body":"Just a few nits:\n* The patch for 2.0 does not compile as ModificationStatement need to be modified\n* I can understand why you used {{processSelection}} and {{selectionsNeedProcessing}} but as everything else is about functions I find it a bit confusing. \n* Renaming {{SelectionWithFunctions}} to {{SelectionWithProcessing}} only in version 2.1 is not really consistent\n* It will be good if you can add also some unit tests to test the ordering behavior with functions","created":"2014-11-17T18:05:36.247+0000"},{"body":"v2 patches are attached. I also have branches here: [2.0|https://github.com/thobbs/cassandra/tree/CASSANDRA-8286], [2.1|https://github.com/thobbs/cassandra/tree/CASSANDRA-8286-2.1], and [trunk|https://github.com/thobbs/cassandra/tree/CASSANDRA-8286-trunk].\n\nbq. The patch for 2.0 does not compile as ModificationStatement need to be modified\n\nFixed\n\nbq. I can understand why you used processSelection and selectionsNeedProcessing but as everything else is about functions I find it a bit confusing.\n\nIn 2.1 and 3.0, it's also handling field selection for UDTs. Once CASSANDRA-7396 is complete, that will be another non-function case.\n\nbq. Renaming SelectionWithFunctions to SelectionWithProcessing only in version 2.1 is not really consistent\n\nFixed, thanks.\n\nbq. It will be good if you can add also some unit tests to test the ordering behavior with functions\n\nThe latest patches has some additional unit tests to cover ordering and ordering with functions and field selections.","created":"2014-11-18T20:37:23.230+0000"},{"body":"I like the new unit tests.\n+1","created":"2014-11-19T14:05:53.589+0000"},{"body":"Thanks, committed as 1945384fdf1d0bac18d6f75e5f864f1aca5b49db.","created":"2014-11-19T17:14:47.645+0000"}],"conversations":[{"body":"The dtest {{cql_tests.py:TestCQL.order_by_multikey_test}} is now failing in 2.0:\nhttp://cassci.datastax.com/job/cassandra-2.0_dtest/lastCompletedBuild/testReport/cql_tests/TestCQL/order_by_multikey_test/history/\n\nThis failure began at the commit for CASSANDRA-8178.\n\nThe error message reads {code}======================================================================\nERROR: order_by_multikey_test (cql_tests.TestCQL)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/Users/philipthompson/cstar/cassandra-dtest/dtest.py\", line 524, in wrapped\n f(obj)\n File \"/Users/philipthompson/cstar/cassandra-dtest/cql_tests.py\", line 1807, in order_by_multikey_test\n res = cursor.execute(\"SELECT col1 FROM test WHERE my_id in('key1', 'key2', 'key3') ORDER BY col1;\")\n File \"/Library/Python/2.7/site-packages/cassandra/cluster.py\", line 1281, in execute\n result = future.result(timeout)\n File \"/Library/Python/2.7/site-packages/cassandra/cluster.py\", line 2771, in result\n raise self._final_exception\nInvalidRequest: code=2200 [Invalid query] message=\"ORDER BY could not be used on columns missing in select clause.\"{code}\n\nand occurs at the query {{SELECT col1 FROM test WHERE my_id in('key1', 'key2', 'key3') ORDER BY col1;}}","from":"reporter","subject":"Regression in ORDER BY"},{"body":"The test {{TestCQL.order_by_with_in_test}} is failing for the same reason and began at the same commit.","from":"developer"},{"body":"Despite the commit being in 2.1.x as well, the tests are passing there with no error. This is due to CASSANDRA-4911.","from":"developer"},{"body":"However, in the 2.1 test {{TestCQL.in_order_by_without_selecting_test}}, the following exception has begun appearing at the same time:\n\nhttp://cassci.datastax.com/job/cassandra-2.1_dtest/lastCompletedBuild/testReport/cql_tests/TestCQL/in_order_by_without_selecting_test/history/\n\n{code}ERROR [SharedPool-Worker-3] 2014-11-10 14:58:28,591 Message.java\n:538 - Unexpected exception during request; channel = [id: 0x5aa\ne53fd, /127.0.0.1:63101 => /127.0.0.1:9042]\njava.lang.AssertionError: null\n at org.apache.cassandra.cql3.ResultSet.addRow(ResultSet.java:63) ~[main/:na]\n at org.apache.cassandra.cql3.statements.Selection$ResultSetBuilder.newRow(Selection.java:333) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.processColumnFamily(SelectStatement.java:1227) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.process(SelectStatement.java:1161) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.processResults(SelectStatement.java:290) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:267) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:215) ~[main/:na]\n at org.apache.cassandra.cql3.statements.SelectStatement.execute(SelectStatement.java:64) ~[main/:na]\n at org.apache.cassandra.cql3.QueryProcessor.processStatement(QueryProcessor.java:226) ~[main/:na]\n at org.apache.cassandra.cql3.QueryProcessor.process(QueryProcessor.java:248) ~[main/:na]\n at org.apache.cassandra.transport.messages.QueryMessage.execute(QueryMessage.java:119) ~[main/:na]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:439) [main/:na]\n at org.apache.cassandra.transport.Message$Dispatcher.channelRead0(Message.java:335) [main/:na]\n at io.netty.channel.SimpleChannelInboundHandler.channelRead(SimpleChannelInboundHandler.java:105) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:333) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext.access$700(AbstractChannelHandlerContext.java:32) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at io.netty.channel.AbstractChannelHandlerContext$8.run(AbstractChannelHandlerContext.java:324) [netty-all-4.0.23.Final.jar:4.0.23.Final]\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471) [na:1.7.0_67]\n at org.apache.cassandra.concurrent.AbstractTracingAwareExecutorService$FutureTask.run(AbstractTracingAwareExecutorService.java:164) [main/:na]\n at org.apache.cassandra.concurrent.SEPWorker.run(SEPWorker.java:105) [main/:na]\n at java.lang.Thread.run(Thread.java:745) [na:1.7.0_67]{code}","from":"developer"},{"body":"In 2.0, the problem was in {{SelectStatement.processOrderingClause()}} where {{ColumnIdentifier.Raw}} instances were being compared with {{ColumnIdentifier}} instances.\n\nIn 2.1 and trunk, the behavior of {{processOrderingClause()}} was changed, so that problem didn't appear. However, there were two additional issues:\n\nFirst, in {{Selection.usingFunctions()}} simply checked {{instanceof ColumnIdentifier}}, which is false for {{ColumnIdentifier.Raw}}, causing all selections to be treated as function/field selections ({{SelectionWithFunctions}}).\n\nSecond, there was an existing bug with {{SelectionWithFunctions}}. In {{Selection.addColumnForOrdering()}}, the internal list of {{columns}} is updated, which works for {{SimpleSelection}}, but is insufficient for {{SelectionWithFunctions}}, which depends on its pre-built list of {{Selectors}}. So, if there were any functions or field selections, {{ORDER BY}} columns would not be properly added to the list of selections. Since the previously mentioned bug resulted in all selections being {{SelectionWithFunctions}}, the normal selection case was also broken.\n\n The 2.1 and trunk patches override {{addColumnForOrdering()}} in {{SelectionWithFunctions}} to update the selectors, and I pushed an [additional dtest|https://github.com/riptano/cassandra-dtest/commit/da86a8eae38179c12b3576ec3d4beaaa0750144c] to cover selecting functions with {{ORDER BY}} clauses.","from":"developer"},{"body":"[~blerer] can you review?","from":"developer"},{"body":"Just a few nits:\n* The patch for 2.0 does not compile as ModificationStatement need to be modified\n* I can understand why you used {{processSelection}} and {{selectionsNeedProcessing}} but as everything else is about functions I find it a bit confusing. \n* Renaming {{SelectionWithFunctions}} to {{SelectionWithProcessing}} only in version 2.1 is not really consistent\n* It will be good if you can add also some unit tests to test the ordering behavior with functions","from":"developer"},{"body":"v2 patches are attached. I also have branches here: [2.0|https://github.com/thobbs/cassandra/tree/CASSANDRA-8286], [2.1|https://github.com/thobbs/cassandra/tree/CASSANDRA-8286-2.1], and [trunk|https://github.com/thobbs/cassandra/tree/CASSANDRA-8286-trunk].\n\nbq. The patch for 2.0 does not compile as ModificationStatement need to be modified\n\nFixed\n\nbq. I can understand why you used processSelection and selectionsNeedProcessing but as everything else is about functions I find it a bit confusing.\n\nIn 2.1 and 3.0, it's also handling field selection for UDTs. Once CASSANDRA-7396 is complete, that will be another non-function case.\n\nbq. Renaming SelectionWithFunctions to SelectionWithProcessing only in version 2.1 is not really consistent\n\nFixed, thanks.\n\nbq. It will be good if you can add also some unit tests to test the ordering behavior with functions\n\nThe latest patches has some additional unit tests to cover ordering and ordering with functions and field selections.","from":"developer"},{"body":"I like the new unit tests.\n+1","from":"developer"},{"body":"Thanks, committed as 1945384fdf1d0bac18d6f75e5f864f1aca5b49db.","from":"developer"}],"created":"2014-11-10T20:48:26.000+0000","description":"The dtest {{cql_tests.py:TestCQL.order_by_multikey_test}} is now failing in 2.0:\nhttp://cassci.datastax.com/job/cassandra-2.0_dtest/lastCompletedBuild/testReport/cql_tests/TestCQL/order_by_multikey_test/history/\n\nThis failure began at the commit for CASSANDRA-8178.\n\nThe error message reads {code}======================================================================\nERROR: order_by_multikey_test (cql_tests.TestCQL)\n----------------------------------------------------------------------\nTraceback (most recent call last):\n File \"/Users/philipthompson/cstar/cassandra-dtest/dtest.py\", line 524, in wrapped\n f(obj)\n File \"/Users/philipthompson/cstar/cassandra-dtest/cql_tests.py\", line 1807, in order_by_multikey_test\n res = cursor.execute(\"SELECT col1 FROM test WHERE my_id in('key1', 'key2', 'key3') ORDER BY col1;\")\n File \"/Library/Python/2.7/site-packages/cassandra/cluster.py\", line 1281, in execute\n result = future.result(timeout)\n File \"/Library/Python/2.7/site-packages/cassandra/cluster.py\", line 2771, in result\n raise self._final_exception\nInvalidRequest: code=2200 [Invalid query] message=\"ORDER BY could not be used on columns missing in select clause.\"{code}\n\nand occurs at the query {{SELECT col1 FROM test WHERE my_id in('key1', 'key2', 'key3') ORDER BY col1;}}","issue_id":"12754248","key":"CASSANDRA-8286","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2014-11-19T17:14:47.000+0000","role":"gold_target","summary":"Regression in ORDER BY"} {"case_id":"12777272","cluster":"JIRA-CASSANDRA-a3ad97e7b7f2","comments":[{"body":"That's odd, it sounds like the cache code but I don't know of any changes there in 2.1.3. It would probably be useful to have the full heap dump to confirm.","created":"2015-02-25T21:06:08.881+0000"},{"body":"bq. it sounds like the cache code\n\nIt's a regular HashMap, not a CLHM, causing this heap pressure.\n\n{quote}\n\"CompactionExecutor:14\" daemon prio=10 tid=0x00002aaace7a7800 nid=0x7e31 runnable [0x0000000042a79000]\n java.lang.Thread.State: RUNNABLE\n\tat java.util.HashMap.transfer(HashMap.java:605)\n\tat java.util.HashMap.resize(HashMap.java:585)\n\tat java.util.HashMap.addEntry(HashMap.java:883)\n\tat java.util.HashMap.put(HashMap.java:509)\n\tat java.util.HashSet.add(HashSet.java:217)\n\tat org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy.createOverlapChain(SizeTieredCompactionStrategy.java:235)\n\tat org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy.filterColdSSTables(SizeTieredCompactionStrategy.java:202)\n\tat org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy.getNextBackgroundSSTables(SizeTieredCompactionStrategy.java:85)\n\tat org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy.getNextBackgroundTask(SizeTieredCompactionStrategy.java:320)\n\t- locked <0x000000064f254560> (a org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy)\n\tat org.apache.cassandra.db.compaction.WrappingCompactionStrategy.getNextBackgroundTask(WrappingCompactionStrategy.java:72)\n\t- locked <0x000000064f254500> (a org.apache.cassandra.db.compaction.WrappingCompactionStrategy)\n\tat org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:234)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:744)\n{quote}\n\nLooks like a viable candidate, and was introduced recently. I can't see anything fundamentally broken about it, although it does have O(n^2) memory utilisation in the number of sstables. This would require around 8500 sstables to be considered for compaction to create the current heap pressure, which is both: 1) possible for some users (so probably we need to rethink); and 2) also pretty uncommon? so perhaps not the issue here?\n\nCould we get a list of the data files present on disk for this table, your log file, and perhaps a heap dump?\n","created":"2015-02-26T14:36:04.652+0000"},{"body":"Sounds likely\n\nWorkaround would be to set cold_reads_to_omit to 0.0, but we should limit the amount of entries in that map.","created":"2015-02-26T16:41:01.020+0000"},{"body":"One way to do this might be, for each reader:\n\n* find its overlaps;\n** for each overlap, find _its_ set of overlaps, but discard these unless they're both larger than the one we've most recently kept, _and_ the estimated win is sufficient\n* remove all of the (first layer of) overlaps from the readers we visit, and continue with the next remaining reader\n\nThis gives us linear space complexity, but possibly still quadratic time complexity in the worst case, which is trickier to improve.","created":"2015-02-27T01:52:53.950+0000"},{"body":"The heapdump is still being uploaded, I find these Entry objects is indeed generated by SizeTieredCompactionStrategy.createOverlapChain. So do you still need the heapdump?\n\nBesides, in the node which has this issue , there is a table which has 8935 SSTables because this table is rarely being read.","created":"2015-02-27T06:10:05.808+0000"},{"body":"I submit a patch to fix this issue. \nThe idea is very simple: keep only one overlapChain in heap whose estimateCompactionGain < 0.7 and has the max size from any other overlapChains. \n\nIt is still O(N^2) time complexity and generates O(N^2) Entry objects, but useless Entry objects will be recycled because it only keeps at most O(N) Entry objects simultaneously.","created":"2015-02-27T07:11:39.625+0000"},{"body":"I tested my patch in my node on 2.1.3. The patch works to me that it won't exhaust the heap although the gc pressure is high.\n\nFurthermore, if my understanding is not wrong, I think there is an improvement on this O(N^2) algorithm:\n\n\nIf createOverlapChain(a, overlapMap) return \\{a,b,c}, createOverlapChain(b, overlapMap) must also return \\{a,b,c} because \"overlapping\" is a bidirectional relationship. So if we have already get \\{a,b,c} by SStableReader a, we needn't check b and c at all.\n\n\nSo this algorithm is O(N) both in time and space.","created":"2015-02-27T09:57:15.907+0000"},{"body":"I upload a new patch according to my new algorithm. We can compare three jstat outputs with different versions' code in my cluster and see that v1 will not exhaust the heap but still have high gc pressure, while v2's gc pressure is much lower.","created":"2015-02-27T10:19:39.846+0000"},{"body":"bq. If createOverlapChain(a, overlapMap) return {a,b,c}, createOverlapChain(b, overlapMap) must also return {a,b,c} because \"overlapping\" is a bidirectional relationship. So if we have already get {a,b,c} by SStableReader a, we needn't check b and c at all.\n\nThis is unfortunately not true. _a_ may overlap _b_ and _c_, but _b_ may only overlap _a_, and not _c_, for instance. I knocked together a patch on the plane with an average time complexity closer to O(N) but with worst case still O(N^2), and achieved by filtering out sets of readers that are wholly contained in other sets of overlapping readers. This may not be desirable, though, as it may be that by considering a smaller group of readers we see a larger yield from rewriting.\n\nI think the best (medium term) solution to this complexity is to take ownership of the hyperloglog implementation and to build a structure that can yield overlaps efficiently from a multitude of hyperloglogs. If each hyperloglog bucket maintained an ordered map from bucket cardinality to the set of objects with that cardinality, an efficient search could be performed for overlap candidates. Within each bucket we could find the largest overlapping cardinalities, then take the intersection of these across all buckets, leaving the top-ranked partition overlaps, and applying correction for the amount of clustering column overlap between the candidates. We could then post-process to include any files that are almost wholly overlapping with this set of best overlaps. This is just a quick off-the-cuff sketch of a design, and no doubt there are better and more developed approaches along these lines.","created":"2015-02-28T21:10:03.251+0000"},{"body":"That's a lot of complexity for something whose benefit I'm unclear on (I'm talking about the \"ignore cold sstables\" feature).\n\nTo the best of my knowledge, CASSANDRA-8635 has never been observed as a problem in the wild. I'd be in favor of reverting it and documenting \"don't enable cold sstable elision if you do a lot of overwrites.\"","created":"2015-03-01T03:23:30.430+0000"},{"body":"{quote}\nThis is unfortunately not true. a may overlap b and c, but b may only overlap a, and not c, for instance.\n{quote}\nYour case is just like this, right?\n{noformat}\nb -----\na -----\nc ------\n{noformat}\nHowever, in the code of createOverlapChain, all the overlapping SSTables with b will be pushed into the queue(for this case is a). When a is popped, all the overlapping SSTables with a(for this case is c) will also be pushed, until there is no more SSTables that overlap any of SSTables in overlapChain(\\{a,b,c} for this case). Just like the comment on this function says: \"if we have 3 sstables, a, b, c where a overlaps with b, but not c and b overlaps with c, all sstables would be returned.\"\n\nSo in graph theory, this function input \"Map> m\" as an undirected graph and \"SSTableReader s\" as a node of this graph, will return the connected component which contains this node. If this connected component is \\{a,b,c}, input a, b or c will return the same connected component.","created":"2015-03-01T07:10:02.627+0000"},{"body":"Furthermore, if my algorithm is correct, my patch-v2 still need a 2-d loop to get the undirected graph. The space complexity is still O(N^2) if all the SSTables overlap each other and the graph is a complete graph. So I think there is a way to optimize it:\n\nUse a 2-d bitmap matrix with the boolean which means whether coldSSTables[i] overlaps coldSSTables[j] to replace the 2-d collection of SSTableReader. Each HashMap$Entry need 32*8=256 bits but each bitmap boolean needs only 1 bit. This method may need more memory in some case, but need much less in the worst case.\n\nIf my algorithm is not wrong, we can discuss if this optimization is needed or not :)","created":"2015-03-01T07:41:35.796+0000"},{"body":"bq. That's a lot of complexity for something whose benefit I'm unclear on (I'm talking about the \"ignore cold sstables\" feature).\n\nQuite possibly. I've no idea how much use this is. The ability to select overlaps efficiently I assumed would be of wider applicability, but it looks like we don't use it anywhere else. However we might in future? The compaction exercise Marcus runs to intro devs to C* utilises it, so not sure if that is just a toy problem or of wider potential.\n\nbq. Just like the comment on this function says: \"if we have 3 sstables, a, b, c where a overlaps with b, but not c and b overlaps with c, all sstables would be returned.\"\n\nGood point. With that realisation, the code I knocked out on the plane is trivially made guaranteed O(n lg n) (for sorting, O(n) for the remainder of the work), with O(n) space complexity. It is untested, and could be refactored not to use an Iterable (which might be less code complexity), but I upload it [here|https://github.com/belliottsmith/cassandra/tree/8860-alt] for you and Marcus to decide what to do with.","created":"2015-03-01T10:40:49.581+0000"},{"body":"We should probably remove this feature alltogether (cold_reads_to_omit), using DTCS suits these kinds of workloads much better","created":"2015-03-01T12:55:50.755+0000"},{"body":"+1","created":"2015-03-02T15:18:50.853+0000"},{"body":"I had a look at the code and I have some questions:\n\nIn L241, if previ==prevj, it will return an empty list? Should this be readers.subList(previ, prevj+1)?\n\nIn L255, it will find the last sstable that overlaps sstable[i], but sstable[j] may not have the biggest maxColumnNames so there may be another between i and j that overlaps sstable[j+1] but sstable[j] doesn't overlap [j+1]. The chain will be broken. I think we should record the sstable with the biggest maxColumnNames from i to j and compare it with each sstable[j+1].","created":"2015-03-03T04:27:47.509+0000"},{"body":"[~yangzhe1991] yes, both statements are correct, but it's kind of moot since it looks like we'll be removing it :)","created":"2015-03-03T10:17:42.316+0000"},{"body":"patch to remove option attached\n\nkeeping the actual parameter in 2.1 if anyone has automated creating tables etc, but will remove entirely in 3.0","created":"2015-03-03T10:50:56.828+0000"},{"body":"+1, patch looks good","created":"2015-03-03T18:16:41.126+0000"},{"body":"committed, thanks","created":"2015-03-04T08:10:35.061+0000"},{"body":"Note we are going to try a simpler patch on 2.1.3 (because we believe we are seeing this too, though our heaps are simply too big to dump reasonably it would seem). We're just reverting the cold_reads_to_omit default to 0.0, and adding logging.\n\nWe could probably change the schema for all tables to do this explicitly, but I haven't looked at whether this affects LCS L0\n\nThis is just an FYI to see if it helps us - we are seeing memory issue that we did not see in testing with 2.1.3, but may well be related to a proliferation of sstables","created":"2015-03-05T02:51:54.779+0000"},{"body":"{quote}We should probably remove this feature alltogether (cold_reads_to_omit), using DTCS suits these kinds of workloads much better{quote}\nFor the record (not that anyone asked... ;D) I am +1 here; seems like lots of complexity for a questionable potential win, and initial implementation has exposed some serious potential edge cases. DTCS FTW.","created":"2015-03-05T18:35:50.296+0000"}],"conversations":[{"body":"While I upgrading my cluster to 2.1.3, I find some nodes (not all) may have GC issue after the node restarting successfully. Old gen grows very fast and most of the space can not be recycled after setting its status to normal immediately. The qps of both reading and writing are very low and there is no heavy compaction.\n\nJmap result seems strange that there are too many java.util.HashMap$Entry objects in heap, where in my experience the \"[B\" is usually the No1.\n\nIf I downgrade it to 2.1.1, this issue will not appear.\n\nI uploaded conf files and jstack/jmap outputs. I'll upload heap dump if someone need it.","from":"reporter","subject":"Remove cold_reads_to_omit from STCS"},{"body":"That's odd, it sounds like the cache code but I don't know of any changes there in 2.1.3. It would probably be useful to have the full heap dump to confirm.","from":"developer"},{"body":"bq. it sounds like the cache code\n\nIt's a regular HashMap, not a CLHM, causing this heap pressure.\n\n{quote}\n\"CompactionExecutor:14\" daemon prio=10 tid=0x00002aaace7a7800 nid=0x7e31 runnable [0x0000000042a79000]\n java.lang.Thread.State: RUNNABLE\n\tat java.util.HashMap.transfer(HashMap.java:605)\n\tat java.util.HashMap.resize(HashMap.java:585)\n\tat java.util.HashMap.addEntry(HashMap.java:883)\n\tat java.util.HashMap.put(HashMap.java:509)\n\tat java.util.HashSet.add(HashSet.java:217)\n\tat org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy.createOverlapChain(SizeTieredCompactionStrategy.java:235)\n\tat org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy.filterColdSSTables(SizeTieredCompactionStrategy.java:202)\n\tat org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy.getNextBackgroundSSTables(SizeTieredCompactionStrategy.java:85)\n\tat org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy.getNextBackgroundTask(SizeTieredCompactionStrategy.java:320)\n\t- locked <0x000000064f254560> (a org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy)\n\tat org.apache.cassandra.db.compaction.WrappingCompactionStrategy.getNextBackgroundTask(WrappingCompactionStrategy.java:72)\n\t- locked <0x000000064f254500> (a org.apache.cassandra.db.compaction.WrappingCompactionStrategy)\n\tat org.apache.cassandra.db.compaction.CompactionManager$BackgroundCompactionTask.run(CompactionManager.java:234)\n\tat java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:471)\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:262)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)\n\tat java.lang.Thread.run(Thread.java:744)\n{quote}\n\nLooks like a viable candidate, and was introduced recently. I can't see anything fundamentally broken about it, although it does have O(n^2) memory utilisation in the number of sstables. This would require around 8500 sstables to be considered for compaction to create the current heap pressure, which is both: 1) possible for some users (so probably we need to rethink); and 2) also pretty uncommon? so perhaps not the issue here?\n\nCould we get a list of the data files present on disk for this table, your log file, and perhaps a heap dump?\n","from":"developer"},{"body":"Sounds likely\n\nWorkaround would be to set cold_reads_to_omit to 0.0, but we should limit the amount of entries in that map.","from":"developer"},{"body":"One way to do this might be, for each reader:\n\n* find its overlaps;\n** for each overlap, find _its_ set of overlaps, but discard these unless they're both larger than the one we've most recently kept, _and_ the estimated win is sufficient\n* remove all of the (first layer of) overlaps from the readers we visit, and continue with the next remaining reader\n\nThis gives us linear space complexity, but possibly still quadratic time complexity in the worst case, which is trickier to improve.","from":"developer"},{"body":"The heapdump is still being uploaded, I find these Entry objects is indeed generated by SizeTieredCompactionStrategy.createOverlapChain. So do you still need the heapdump?\n\nBesides, in the node which has this issue , there is a table which has 8935 SSTables because this table is rarely being read.","from":"developer"},{"body":"I submit a patch to fix this issue. \nThe idea is very simple: keep only one overlapChain in heap whose estimateCompactionGain < 0.7 and has the max size from any other overlapChains. \n\nIt is still O(N^2) time complexity and generates O(N^2) Entry objects, but useless Entry objects will be recycled because it only keeps at most O(N) Entry objects simultaneously.","from":"developer"},{"body":"I tested my patch in my node on 2.1.3. The patch works to me that it won't exhaust the heap although the gc pressure is high.\n\nFurthermore, if my understanding is not wrong, I think there is an improvement on this O(N^2) algorithm:\n\n\nIf createOverlapChain(a, overlapMap) return \\{a,b,c}, createOverlapChain(b, overlapMap) must also return \\{a,b,c} because \"overlapping\" is a bidirectional relationship. So if we have already get \\{a,b,c} by SStableReader a, we needn't check b and c at all.\n\n\nSo this algorithm is O(N) both in time and space.","from":"developer"},{"body":"I upload a new patch according to my new algorithm. We can compare three jstat outputs with different versions' code in my cluster and see that v1 will not exhaust the heap but still have high gc pressure, while v2's gc pressure is much lower.","from":"developer"},{"body":"bq. If createOverlapChain(a, overlapMap) return {a,b,c}, createOverlapChain(b, overlapMap) must also return {a,b,c} because \"overlapping\" is a bidirectional relationship. So if we have already get {a,b,c} by SStableReader a, we needn't check b and c at all.\n\nThis is unfortunately not true. _a_ may overlap _b_ and _c_, but _b_ may only overlap _a_, and not _c_, for instance. I knocked together a patch on the plane with an average time complexity closer to O(N) but with worst case still O(N^2), and achieved by filtering out sets of readers that are wholly contained in other sets of overlapping readers. This may not be desirable, though, as it may be that by considering a smaller group of readers we see a larger yield from rewriting.\n\nI think the best (medium term) solution to this complexity is to take ownership of the hyperloglog implementation and to build a structure that can yield overlaps efficiently from a multitude of hyperloglogs. If each hyperloglog bucket maintained an ordered map from bucket cardinality to the set of objects with that cardinality, an efficient search could be performed for overlap candidates. Within each bucket we could find the largest overlapping cardinalities, then take the intersection of these across all buckets, leaving the top-ranked partition overlaps, and applying correction for the amount of clustering column overlap between the candidates. We could then post-process to include any files that are almost wholly overlapping with this set of best overlaps. This is just a quick off-the-cuff sketch of a design, and no doubt there are better and more developed approaches along these lines.","from":"developer"},{"body":"That's a lot of complexity for something whose benefit I'm unclear on (I'm talking about the \"ignore cold sstables\" feature).\n\nTo the best of my knowledge, CASSANDRA-8635 has never been observed as a problem in the wild. I'd be in favor of reverting it and documenting \"don't enable cold sstable elision if you do a lot of overwrites.\"","from":"developer"},{"body":"{quote}\nThis is unfortunately not true. a may overlap b and c, but b may only overlap a, and not c, for instance.\n{quote}\nYour case is just like this, right?\n{noformat}\nb -----\na -----\nc ------\n{noformat}\nHowever, in the code of createOverlapChain, all the overlapping SSTables with b will be pushed into the queue(for this case is a). When a is popped, all the overlapping SSTables with a(for this case is c) will also be pushed, until there is no more SSTables that overlap any of SSTables in overlapChain(\\{a,b,c} for this case). Just like the comment on this function says: \"if we have 3 sstables, a, b, c where a overlaps with b, but not c and b overlaps with c, all sstables would be returned.\"\n\nSo in graph theory, this function input \"Map> m\" as an undirected graph and \"SSTableReader s\" as a node of this graph, will return the connected component which contains this node. If this connected component is \\{a,b,c}, input a, b or c will return the same connected component.","from":"developer"},{"body":"Furthermore, if my algorithm is correct, my patch-v2 still need a 2-d loop to get the undirected graph. The space complexity is still O(N^2) if all the SSTables overlap each other and the graph is a complete graph. So I think there is a way to optimize it:\n\nUse a 2-d bitmap matrix with the boolean which means whether coldSSTables[i] overlaps coldSSTables[j] to replace the 2-d collection of SSTableReader. Each HashMap$Entry need 32*8=256 bits but each bitmap boolean needs only 1 bit. This method may need more memory in some case, but need much less in the worst case.\n\nIf my algorithm is not wrong, we can discuss if this optimization is needed or not :)","from":"developer"},{"body":"bq. That's a lot of complexity for something whose benefit I'm unclear on (I'm talking about the \"ignore cold sstables\" feature).\n\nQuite possibly. I've no idea how much use this is. The ability to select overlaps efficiently I assumed would be of wider applicability, but it looks like we don't use it anywhere else. However we might in future? The compaction exercise Marcus runs to intro devs to C* utilises it, so not sure if that is just a toy problem or of wider potential.\n\nbq. Just like the comment on this function says: \"if we have 3 sstables, a, b, c where a overlaps with b, but not c and b overlaps with c, all sstables would be returned.\"\n\nGood point. With that realisation, the code I knocked out on the plane is trivially made guaranteed O(n lg n) (for sorting, O(n) for the remainder of the work), with O(n) space complexity. It is untested, and could be refactored not to use an Iterable (which might be less code complexity), but I upload it [here|https://github.com/belliottsmith/cassandra/tree/8860-alt] for you and Marcus to decide what to do with.","from":"developer"},{"body":"We should probably remove this feature alltogether (cold_reads_to_omit), using DTCS suits these kinds of workloads much better","from":"developer"},{"body":"+1","from":"developer"},{"body":"I had a look at the code and I have some questions:\n\nIn L241, if previ==prevj, it will return an empty list? Should this be readers.subList(previ, prevj+1)?\n\nIn L255, it will find the last sstable that overlaps sstable[i], but sstable[j] may not have the biggest maxColumnNames so there may be another between i and j that overlaps sstable[j+1] but sstable[j] doesn't overlap [j+1]. The chain will be broken. I think we should record the sstable with the biggest maxColumnNames from i to j and compare it with each sstable[j+1].","from":"developer"},{"body":"[~yangzhe1991] yes, both statements are correct, but it's kind of moot since it looks like we'll be removing it :)","from":"developer"},{"body":"patch to remove option attached\n\nkeeping the actual parameter in 2.1 if anyone has automated creating tables etc, but will remove entirely in 3.0","from":"developer"},{"body":"+1, patch looks good","from":"developer"},{"body":"committed, thanks","from":"developer"},{"body":"Note we are going to try a simpler patch on 2.1.3 (because we believe we are seeing this too, though our heaps are simply too big to dump reasonably it would seem). We're just reverting the cold_reads_to_omit default to 0.0, and adding logging.\n\nWe could probably change the schema for all tables to do this explicitly, but I haven't looked at whether this affects LCS L0\n\nThis is just an FYI to see if it helps us - we are seeing memory issue that we did not see in testing with 2.1.3, but may well be related to a proliferation of sstables","from":"developer"},{"body":"{quote}We should probably remove this feature alltogether (cold_reads_to_omit), using DTCS suits these kinds of workloads much better{quote}\nFor the record (not that anyone asked... ;D) I am +1 here; seems like lots of complexity for a questionable potential win, and initial implementation has exposed some serious potential edge cases. DTCS FTW.","from":"developer"}],"created":"2015-02-24T19:47:07.000+0000","description":"While I upgrading my cluster to 2.1.3, I find some nodes (not all) may have GC issue after the node restarting successfully. Old gen grows very fast and most of the space can not be recycled after setting its status to normal immediately. The qps of both reading and writing are very low and there is no heavy compaction.\n\nJmap result seems strange that there are too many java.util.HashMap$Entry objects in heap, where in my experience the \"[B\" is usually the No1.\n\nIf I downgrade it to 2.1.1, this issue will not appear.\n\nI uploaded conf files and jstack/jmap outputs. I'll upload heap dump if someone need it.","issue_id":"12777272","key":"CASSANDRA-8860","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-03-04T08:10:35.000+0000","role":"gold_target","summary":"Remove cold_reads_to_omit from STCS"} {"case_id":"12842222","cluster":"JIRA-CASSANDRA-102a1b80a28c","comments":[{"body":"As per discussion on CASSANDRA-9719, one part of checking that this ticket is finished will be making sure it passes the {{upgrade_internal_auth_test.py:TestAuthUpgrade.upgrade_to_30_test}} dtest.","created":"2015-07-09T18:19:38.060+0000"},{"body":"I'm currently still debugging a couple of failing upgrade tests and waiting on cassci for normal test results, but my branch should be close enough to start the review process.\n\nCassandra branch: https://github.com/thobbs/cassandra/tree/CASSANDRA-9704\nUpgrade dtest branch: https://github.com/thobbs/cassandra-dtest/tree/8099-backwards-compat","created":"2015-07-15T21:40:44.221+0000"},{"body":"I forgot to attach a patch for the required 2.1/2.2 changes as well. Basically, when paging in 2.x, our check to see if a new page contains the same row that the previous page ended on looks for an exact cell name match. This is fine in 2.x because we will return partial rows at the end of the page (just the row marker cell). However, in 3.0, we always return full rows. While we _could_ make some very hacky changes to 3.0 to enable returning a partial row at the end of the page, this seems like the cleanest solution.\n\nThe attached patch (and [branch|https://github.com/thobbs/cassandra/tree/CASSANDRA-9704-2.1-forward-compat]) makes those changes.","created":"2015-07-16T16:02:28.569+0000"},{"body":"There are two remaining failing tests (that I'm aware of) in the new {{upgrade_tests/cql_tests.py}}:\n* {{only_pk_test}}: this requires a schema migration fix in CASSANDRA-6717. This [this comment|https://issues.apache.org/jira/browse/CASSANDRA-6717?focusedCommentId=14626965&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-14626965] for details. I've added a {{@require()}} skip to this test for now.\n* {{noncomposite_static_cf_test}}: For the query {{\"SELECT * FROM users WHERE userid = f47ac10b-58cc-4372-a567-0e02b2c3d479\"}}, the 2.1 node sends a read command that gets translated to this in 3.0: {{Read(ks.users columns=age, firstname, lastname rowFilter= limits= key=550e8400-e29b-41d4-a716-446655440000 filter=names(), nowInSec=1437159837)}}. Normally in 3.0 we would use a slice query for this, but the above query _should_ work. However, there appears to be some sort of internal problem handling the query, resulting in the response containing three cells: firstname, lastname, and a duplicate of lastname. I'm guessing this is caused by some misuse of a Reusable object, but it's damned hard to track down. I suspect that CASSANDRA-9705 would fix this, but I haven't verified this yet (see comments below about merging), so I've added a comment and left the test failing.\n\nOn the topic of CASSANDRA-9705, it looks like this branch is going to require quite a few changes to accommodate that. I suggest waiting until 9705 has been committed before rebasing and committing this branch.","created":"2015-07-17T19:48:19.820+0000"},{"body":"I just ported {{static_columns_paging_test}}, but that's failing in the same manner as in CASSANDRA-9775, so it doesn't appear to be introduced by the new changes. I've left the {{@require('9775')}} decorator on the test for now.\n\nI also remembered that we should really be testing Thrift queries against mixed-version clusters as well, so I'll begin the process of writing those. I'm tempted to start from the pycassa integration tests, as those are the most thorough Thrift tests we have (in my opinion), but that may take some time to convert.","created":"2015-07-17T22:24:27.943+0000"},{"body":"[~thobbs] Can unrequire {{only_pk_test}} now that CASSANDRA-9874 fixed the issue.","created":"2015-07-23T15:48:12.050+0000"},{"body":"bq. I forgot to attach a patch for the required 2.1/2.2 changes as well.\n\nForgot to mention it here, but I had created CASSANDRA-9849 for that and the patch for those changes is now committed.","created":"2015-07-24T07:30:47.506+0000"},{"body":"Alright, that patch looks pretty good overall. Good job [~thobbs], that must not have been all fun.\n\nI did rebase that branch (which was pretty simple), and notice a few minor points:\n* some serializedSize methods weren't exactly mimicking their serialize counterpart.\n* the serialization of legacy ReadResponse was a tad hard to follow. For ranges at least, the number of partition was serialized in {{ReadResponse}} but deserialized in {{UnfilteredPartitionIterators}} for instance. So I simply moved all serialization of read response to {{ReadResponse}}, using the {{UnfilteredPartitionIterators}} serializer only for the current messaging version.\n* {{ReadResponse.LegacyRemoteDataResponse}} was imo more complicated that it had to be. Instead of serializing/deserializing internally to that class, I've just changed so it reads the partition in an {{ArrayBackedPartition}} directly (it's what pre-3.0 node do after all). This simplified the code quite a bit, avoid things like the {{preprocessLegacyResults}}. \n* There was some code duplication between {{PartitionUpdate}} and {{UnfilteredPartitionIterators}} serializer so I've consolidated that (I'm not sure {{UnfilteredPartitionIterators}} wasn't handling row deletion properly in particular).\n\nAnyway, Tyler is currently on vacation and it feels like there is no point of letting that code sit here so I committed with [my rebase/updates|https://github.com/pcmanus/cassandra/commits/9704] (it's not breaking any cassci test as far as I can tell).\n\nNow I ran Tyler's dtest branch locally and while most tests are working, a few are erroring out and I haven't had the time to look more closely. So I've created CASSANDRA-9893 to follow up on those. We should also commit those tests to the dtest branch so cassci run them, but I had a bit of trouble running them locally so I think someone that has access to cassci should look at it. \n","created":"2015-07-24T13:56:14.022+0000"},{"body":"It looks like {{PartitionUpdateSerializer.serialize()}} is not implemented for {{version < MessagingService.VERSION_30}}. The code is [commented out|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/db/partitions/PartitionUpdate.java#L704-L707] and throwing {{UnsupportedOperationException}}, which would mean that 3.0 coordinators can't compatibly forward to 2.2 nodes.","created":"2015-08-03T17:35:57.726+0000"},{"body":"Reopening for now. We do have CASSANDRA-9893, but this feels more appropriate (feel free to re-resolve if you disagree).","created":"2015-08-03T17:39:31.208+0000"},{"body":"It looks like Sylvain accidentally forgot to push his commit to the apache repo. I've rebased on the latest {{cassandra-3.0}} and am working on fixing test failures in [this branch|https://github.com/thobbs/cassandra/tree/CASSANDRA-9704-rebase].","created":"2015-08-04T21:35:33.785+0000"},{"body":"All dtest failures on existing tests appear to not be regressions, so I've committed this as {{8c64cefd19d706003d4b33b333274dbf17c9cb34}} to 3.0 and merged to trunk.\n\nA few of the new upgrade dtests are still failing. I will document those in CASSANDRA-9893 so that somebody can continue the work there.","created":"2015-08-07T22:48:20.924+0000"}],"conversations":[{"body":"The currently committed patch for CASSANDRA-8099 has left backward compatibility on the wire as a TODO. This ticket is to track the actual doing (of which I know [~thobbs] has already done a good chunk).","from":"reporter","subject":"On-wire backward compatibility for 8099"},{"body":"As per discussion on CASSANDRA-9719, one part of checking that this ticket is finished will be making sure it passes the {{upgrade_internal_auth_test.py:TestAuthUpgrade.upgrade_to_30_test}} dtest.","from":"developer"},{"body":"I'm currently still debugging a couple of failing upgrade tests and waiting on cassci for normal test results, but my branch should be close enough to start the review process.\n\nCassandra branch: https://github.com/thobbs/cassandra/tree/CASSANDRA-9704\nUpgrade dtest branch: https://github.com/thobbs/cassandra-dtest/tree/8099-backwards-compat","from":"developer"},{"body":"I forgot to attach a patch for the required 2.1/2.2 changes as well. Basically, when paging in 2.x, our check to see if a new page contains the same row that the previous page ended on looks for an exact cell name match. This is fine in 2.x because we will return partial rows at the end of the page (just the row marker cell). However, in 3.0, we always return full rows. While we _could_ make some very hacky changes to 3.0 to enable returning a partial row at the end of the page, this seems like the cleanest solution.\n\nThe attached patch (and [branch|https://github.com/thobbs/cassandra/tree/CASSANDRA-9704-2.1-forward-compat]) makes those changes.","from":"developer"},{"body":"There are two remaining failing tests (that I'm aware of) in the new {{upgrade_tests/cql_tests.py}}:\n* {{only_pk_test}}: this requires a schema migration fix in CASSANDRA-6717. This [this comment|https://issues.apache.org/jira/browse/CASSANDRA-6717?focusedCommentId=14626965&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-14626965] for details. I've added a {{@require()}} skip to this test for now.\n* {{noncomposite_static_cf_test}}: For the query {{\"SELECT * FROM users WHERE userid = f47ac10b-58cc-4372-a567-0e02b2c3d479\"}}, the 2.1 node sends a read command that gets translated to this in 3.0: {{Read(ks.users columns=age, firstname, lastname rowFilter= limits= key=550e8400-e29b-41d4-a716-446655440000 filter=names(), nowInSec=1437159837)}}. Normally in 3.0 we would use a slice query for this, but the above query _should_ work. However, there appears to be some sort of internal problem handling the query, resulting in the response containing three cells: firstname, lastname, and a duplicate of lastname. I'm guessing this is caused by some misuse of a Reusable object, but it's damned hard to track down. I suspect that CASSANDRA-9705 would fix this, but I haven't verified this yet (see comments below about merging), so I've added a comment and left the test failing.\n\nOn the topic of CASSANDRA-9705, it looks like this branch is going to require quite a few changes to accommodate that. I suggest waiting until 9705 has been committed before rebasing and committing this branch.","from":"developer"},{"body":"I just ported {{static_columns_paging_test}}, but that's failing in the same manner as in CASSANDRA-9775, so it doesn't appear to be introduced by the new changes. I've left the {{@require('9775')}} decorator on the test for now.\n\nI also remembered that we should really be testing Thrift queries against mixed-version clusters as well, so I'll begin the process of writing those. I'm tempted to start from the pycassa integration tests, as those are the most thorough Thrift tests we have (in my opinion), but that may take some time to convert.","from":"developer"},{"body":"[~thobbs] Can unrequire {{only_pk_test}} now that CASSANDRA-9874 fixed the issue.","from":"developer"},{"body":"bq. I forgot to attach a patch for the required 2.1/2.2 changes as well.\n\nForgot to mention it here, but I had created CASSANDRA-9849 for that and the patch for those changes is now committed.","from":"developer"},{"body":"Alright, that patch looks pretty good overall. Good job [~thobbs], that must not have been all fun.\n\nI did rebase that branch (which was pretty simple), and notice a few minor points:\n* some serializedSize methods weren't exactly mimicking their serialize counterpart.\n* the serialization of legacy ReadResponse was a tad hard to follow. For ranges at least, the number of partition was serialized in {{ReadResponse}} but deserialized in {{UnfilteredPartitionIterators}} for instance. So I simply moved all serialization of read response to {{ReadResponse}}, using the {{UnfilteredPartitionIterators}} serializer only for the current messaging version.\n* {{ReadResponse.LegacyRemoteDataResponse}} was imo more complicated that it had to be. Instead of serializing/deserializing internally to that class, I've just changed so it reads the partition in an {{ArrayBackedPartition}} directly (it's what pre-3.0 node do after all). This simplified the code quite a bit, avoid things like the {{preprocessLegacyResults}}. \n* There was some code duplication between {{PartitionUpdate}} and {{UnfilteredPartitionIterators}} serializer so I've consolidated that (I'm not sure {{UnfilteredPartitionIterators}} wasn't handling row deletion properly in particular).\n\nAnyway, Tyler is currently on vacation and it feels like there is no point of letting that code sit here so I committed with [my rebase/updates|https://github.com/pcmanus/cassandra/commits/9704] (it's not breaking any cassci test as far as I can tell).\n\nNow I ran Tyler's dtest branch locally and while most tests are working, a few are erroring out and I haven't had the time to look more closely. So I've created CASSANDRA-9893 to follow up on those. We should also commit those tests to the dtest branch so cassci run them, but I had a bit of trouble running them locally so I think someone that has access to cassci should look at it. \n","from":"developer"},{"body":"It looks like {{PartitionUpdateSerializer.serialize()}} is not implemented for {{version < MessagingService.VERSION_30}}. The code is [commented out|https://github.com/apache/cassandra/blob/trunk/src/java/org/apache/cassandra/db/partitions/PartitionUpdate.java#L704-L707] and throwing {{UnsupportedOperationException}}, which would mean that 3.0 coordinators can't compatibly forward to 2.2 nodes.","from":"developer"},{"body":"Reopening for now. We do have CASSANDRA-9893, but this feels more appropriate (feel free to re-resolve if you disagree).","from":"developer"},{"body":"It looks like Sylvain accidentally forgot to push his commit to the apache repo. I've rebased on the latest {{cassandra-3.0}} and am working on fixing test failures in [this branch|https://github.com/thobbs/cassandra/tree/CASSANDRA-9704-rebase].","from":"developer"},{"body":"All dtest failures on existing tests appear to not be regressions, so I've committed this as {{8c64cefd19d706003d4b33b333274dbf17c9cb34}} to 3.0 and merged to trunk.\n\nA few of the new upgrade dtests are still failing. I will document those in CASSANDRA-9893 so that somebody can continue the work there.","from":"developer"}],"created":"2015-07-02T06:59:35.000+0000","description":"The currently committed patch for CASSANDRA-8099 has left backward compatibility on the wire as a TODO. This ticket is to track the actual doing (of which I know [~thobbs] has already done a good chunk).","issue_id":"12842222","key":"CASSANDRA-9704","metadata":{"source":"Apache Jira"},"project":"CASSANDRA","resolution":"Fixed","resolution_date":"2015-08-07T22:48:20.000+0000","role":"gold_target","summary":"On-wire backward compatibility for 8099"} {"case_id":"12923913","cluster":"JIRA-HADOOP-14bfc47fcbf0","comments":[{"body":"addressing findbugs warnings in 002.","created":"2016-01-12T14:26:59.842+0000"},{"body":"Hi [~iwasakims], the patch lgtm. One nitpick for readability. Callers of handleTimeout can increment {{waiting}} since callers know how long the wait was. e.g.\n{code}\n waiting += soTimeout;\n handleTimeout(e, waiting);\n...\n private void handleTimeout(SocketTimeoutException e, int waiting)\n{code}","created":"2016-02-16T19:05:56.101+0000"},{"body":"I updated the patch based on the suggestion. Thanks, [~arpitagarwal].","created":"2016-02-18T06:25:36.923+0000"},{"body":"Sorry I missed your updated patch [~iwasakims]. +1 lgtm.","created":"2016-03-10T22:49:49.391+0000"},{"body":"Thanks, [~arpitagarwal]. I needed trivial rebasing. Attaching rebased patch.","created":"2016-03-11T00:22:46.581+0000"},{"body":"+1 for the v4 patch. ","created":"2016-03-11T00:29:06.561+0000"},{"body":"Committed to branch-2.6 and above. Thanks. ","created":"2016-03-11T06:49:03.916+0000"},{"body":"this just broke everything. Rolling back across the board.","created":"2016-03-11T16:55:42.372+0000"},{"body":"I apologize for the breakage. Thanks for taking care, [~stevel@apache.org].\n\nNameNode proxy in DFSClient is initialized with the timeout value given by {{Client#getTimeout}} which is -1 by default meaning that timeout is not set. I should have take that into account.","created":"2016-03-12T22:53:39.332+0000"},{"body":"{code}\n final public static int getTimeout(Configuration conf) {\n if (!conf.getBoolean(CommonConfigurationKeys.IPC_CLIENT_PING_KEY,\n CommonConfigurationKeys.IPC_CLIENT_PING_DEFAULT)) {\n return getPingInterval(conf);\n }\n return -1;\n }\n{code}\n\n{{Client#getTimeout}} returns -1 if ipc.client.ping is true, otherwise returns the value of ipc.ping.interval. This seems to reflect the previous behaviour. Current client does not time out even if {{ipc.client.ping = false}} because {{ConnectionId#ConnectionId}} set pingInterval to 0 in that case. This is existing behaviour even before HADOOP-11252.","created":"2016-03-13T00:15:40.423+0000"},{"body":"The value of {{Client#getTimeout}} is used in 2 places in the current hadoop code base.\n\n* DfsClientConf.hdfsTimeout\n* NameNodeProxiesClient#createNonHAProxyWithClientProtocol (HDFS-4646)\n\n{{DfsClientConf.hdfsTimeout}} affects the renewal time of LeaseRenewer. There is no problem to fix {{Client#getTimeout}} to return rgiht timeout value.\n\nThere is a compatibility concern about {{NameNodeProxiesClient#createNonHAProxyWithClientProtocol}} since HDFS-4646 depends on the value of {{Client#getTimeout}} when {{ipc.client.ping = false}}.\n\n{noformat}\nipc.client.ping = false\nipc.ping.interval = 6000 (default)\nipc.client.rpc-timeout.ms = 0(default)\n{noformat}\n\n{{Client#getTimeout}} currently returns 6000 on the configuration above. If we fix to return 0 as timeout value, existing cluster which uses {{ipc.client.ping = false}} expecting that the timeout of namenode proxy is set to ipc.ping.interval lose the timeout setting.","created":"2016-03-13T00:17:08.972+0000"},{"body":"I updated the patch.\n\n* fixed not to set negative soTimeout\n* added test case for negative timeout value is given\n* updated tests based on refactoring by HADOOP-12813\n* added {{Client#getRpcTimeout}} and deprecated {{Client#getTimeout}}\n","created":"2016-03-13T00:35:54.984+0000"},{"body":"attaching 006.\n\n{{Client#getTimeout}} should return the value of ipc.client.rpc-timeout.ms if it is set to > 0.\n","created":"2016-03-13T01:50:40.480+0000"},{"body":"No worries about the breakage, though I think I'd have been happier if 2.6 wasn't picking up so much of these RPC changes, not until they were stable. \n\nCan you add this patch to HDFS and YARN JIRAs so we can verify their miniclusters are happy this time round?","created":"2016-03-13T14:03:18.627+0000"},{"body":"I filed HDFS-9954 to test the patch against HDFS. I will file a YARN task if it works.","created":"2016-03-14T13:03:40.691+0000"},{"body":"It turned out that creating HDFS jira and attaching patch does not invoke HDFS tests. test-patch runs tests based on the contents of the patch. Though test-patch.sh of Yetus provides {{--modulelist}} option to run additional tests, it seems not to be possible to use it via QA build. (Thanks to [~sekikn] for the information.)\n\nI ran tests of HDFS and YARN locally after applying 006 and {{mvn install -DskipTests}}. No test failure except for already reported intermittent ones.","created":"2016-03-16T00:26:16.141+0000"},{"body":"+1 the v6 patch lgtm. Ran a few failing tests locally and they passed although it is unfortunate we cannot get a full unit test run. Thanks.","created":"2016-03-29T22:56:34.224+0000"},{"body":"Thanks again, [~arpitagarwal]. I will wait further comments from other reviewers for a day before committing this.","created":"2016-03-30T14:15:52.059+0000"},{"body":"Committed to branch-2.8 and above. Thanks.","created":"2016-04-05T18:35:57.549+0000"}],"conversations":[{"body":"Currently if the value of ipc.client.rpc-timeout.ms is greater than 0, the timeout overrides the ipc.ping.interval and client will throw exception instead of sending ping when the interval is passed. RPC timeout should work without effectively disabling IPC ping.","from":"reporter","subject":"RPC timeout should not override IPC ping interval"},{"body":"addressing findbugs warnings in 002.","from":"developer"},{"body":"Hi [~iwasakims], the patch lgtm. One nitpick for readability. Callers of handleTimeout can increment {{waiting}} since callers know how long the wait was. e.g.\n{code}\n waiting += soTimeout;\n handleTimeout(e, waiting);\n...\n private void handleTimeout(SocketTimeoutException e, int waiting)\n{code}","from":"developer"},{"body":"I updated the patch based on the suggestion. Thanks, [~arpitagarwal].","from":"developer"},{"body":"Sorry I missed your updated patch [~iwasakims]. +1 lgtm.","from":"developer"},{"body":"Thanks, [~arpitagarwal]. I needed trivial rebasing. Attaching rebased patch.","from":"developer"},{"body":"+1 for the v4 patch. ","from":"developer"},{"body":"Committed to branch-2.6 and above. Thanks. ","from":"developer"},{"body":"this just broke everything. Rolling back across the board.","from":"developer"},{"body":"I apologize for the breakage. Thanks for taking care, [~stevel@apache.org].\n\nNameNode proxy in DFSClient is initialized with the timeout value given by {{Client#getTimeout}} which is -1 by default meaning that timeout is not set. I should have take that into account.","from":"developer"},{"body":"{code}\n final public static int getTimeout(Configuration conf) {\n if (!conf.getBoolean(CommonConfigurationKeys.IPC_CLIENT_PING_KEY,\n CommonConfigurationKeys.IPC_CLIENT_PING_DEFAULT)) {\n return getPingInterval(conf);\n }\n return -1;\n }\n{code}\n\n{{Client#getTimeout}} returns -1 if ipc.client.ping is true, otherwise returns the value of ipc.ping.interval. This seems to reflect the previous behaviour. Current client does not time out even if {{ipc.client.ping = false}} because {{ConnectionId#ConnectionId}} set pingInterval to 0 in that case. This is existing behaviour even before HADOOP-11252.","from":"developer"},{"body":"The value of {{Client#getTimeout}} is used in 2 places in the current hadoop code base.\n\n* DfsClientConf.hdfsTimeout\n* NameNodeProxiesClient#createNonHAProxyWithClientProtocol (HDFS-4646)\n\n{{DfsClientConf.hdfsTimeout}} affects the renewal time of LeaseRenewer. There is no problem to fix {{Client#getTimeout}} to return rgiht timeout value.\n\nThere is a compatibility concern about {{NameNodeProxiesClient#createNonHAProxyWithClientProtocol}} since HDFS-4646 depends on the value of {{Client#getTimeout}} when {{ipc.client.ping = false}}.\n\n{noformat}\nipc.client.ping = false\nipc.ping.interval = 6000 (default)\nipc.client.rpc-timeout.ms = 0(default)\n{noformat}\n\n{{Client#getTimeout}} currently returns 6000 on the configuration above. If we fix to return 0 as timeout value, existing cluster which uses {{ipc.client.ping = false}} expecting that the timeout of namenode proxy is set to ipc.ping.interval lose the timeout setting.","from":"developer"},{"body":"I updated the patch.\n\n* fixed not to set negative soTimeout\n* added test case for negative timeout value is given\n* updated tests based on refactoring by HADOOP-12813\n* added {{Client#getRpcTimeout}} and deprecated {{Client#getTimeout}}\n","from":"developer"},{"body":"attaching 006.\n\n{{Client#getTimeout}} should return the value of ipc.client.rpc-timeout.ms if it is set to > 0.\n","from":"developer"},{"body":"No worries about the breakage, though I think I'd have been happier if 2.6 wasn't picking up so much of these RPC changes, not until they were stable. \n\nCan you add this patch to HDFS and YARN JIRAs so we can verify their miniclusters are happy this time round?","from":"developer"},{"body":"I filed HDFS-9954 to test the patch against HDFS. I will file a YARN task if it works.","from":"developer"},{"body":"It turned out that creating HDFS jira and attaching patch does not invoke HDFS tests. test-patch runs tests based on the contents of the patch. Though test-patch.sh of Yetus provides {{--modulelist}} option to run additional tests, it seems not to be possible to use it via QA build. (Thanks to [~sekikn] for the information.)\n\nI ran tests of HDFS and YARN locally after applying 006 and {{mvn install -DskipTests}}. No test failure except for already reported intermittent ones.","from":"developer"},{"body":"+1 the v6 patch lgtm. Ran a few failing tests locally and they passed although it is unfortunate we cannot get a full unit test run. Thanks.","from":"developer"},{"body":"Thanks again, [~arpitagarwal]. I will wait further comments from other reviewers for a day before committing this.","from":"developer"},{"body":"Committed to branch-2.8 and above. Thanks.","from":"developer"}],"created":"2015-12-23T04:24:37.000+0000","description":"Currently if the value of ipc.client.rpc-timeout.ms is greater than 0, the timeout overrides the ipc.ping.interval and client will throw exception instead of sending ping when the interval is passed. RPC timeout should work without effectively disabling IPC ping.","issue_id":"12923913","key":"HADOOP-12672","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2016-04-05T18:35:57.000+0000","role":"gold_target","summary":"RPC timeout should not override IPC ping interval"} {"case_id":"12969687","cluster":"JIRA-HADOOP-6c3457425f28","comments":[{"body":"The attached v001 patch avoids the unnecessary {{getFileStatus}} call.\n\nThe effect is particularly pronounced when running DistCp with a destination on S3A, where eventual consistency on S3 can cause the {{getFileStatus}} call to fail with {{FileNotFoundException}}. Then, the whole MapReduce task fails, retries, and repeats copying all the data. [~rajesh.balamohan], I know you saw this with some recent large copies to S3A. Would you be interested in trying a test with this patch? So far, I don't have my own repro. Note that this patch is only helpful as long as the DistCp command is not preserving metadata attributes, so don't use the {{-p}} option.\n\nCc [~stevel@apache.org].","created":"2016-05-13T22:56:39.960+0000"},{"body":"Thanks for sharing the patch [~cnauroth]. I tried out the patch with hadoop-2.8 (may 16) and I do not see any more task failures with dist-cp to S3. I see some straggler tasks taking time, but that is completely unrelated to this jira.","created":"2016-05-16T07:00:56.179+0000"},{"body":"You know, I think s3a now has enough instrumentation that the # of times that getFileStatus is called would be measurable. \n\nAt the very least, it'd be good to have a test of DistCp there, to verify that inconsistency problems aren't surfacing. The examples in, say {{TestDistCpViewFs}} , show a start, though I'd expect the new tests to simply throw up IOEs, rather than swallow + fail, the way that class does (and which I have just submitted a patch for, in HADOOP-13148).","created":"2016-05-16T10:11:16.518+0000"},{"body":"Patch v003 adds a new abstract contract test suite for DistCp coverage and concrete test suite subclasses for S3A and WASB. I verified the tests are passing for both hadoop-aws (including running in parallel mode) and hadoop-azure.\n\nI'm going to leave the JIRA issue in Open status instead of Patch Available for now. The v003 patch will potentially hit Jenkins a little hard because of touching multiple modules, so I'd like to get another round of code review feedback first.\n","created":"2016-05-17T23:04:07.053+0000"},{"body":"tested -003 against s3 ireland and azure.\n\n{code}\nRunning org.apache.hadoop.fs.contract.s3a.TestS3AContractDistCp\nTests run: 8, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 223.843 sec - in org.apache.hadoop.fs.contract.s3a.TestS3AContractDistCp\n\n...\nRunning org.apache.hadoop.fs.azure.contract.TestAzureNativeContractDistCp\nTests run: 8, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 65.354 sec - in org.apache.hadoop.fs.azure.contract.TestAzureNativeContractDistCp\n\n{code}\n\nInteresting how much faster azure is. \n\nThe patch, is, as it stands, it's going to add 4 min to a TestS3A* test pattern. Could it be made one of the scaleable tests where it takes a config of option on scale so can be made configurable? There are already some tests which use {{scale.test.operation.count}} to control scale; we could have one on distcp file size, with the large file size being driven by it. Make it something in KB and it could easily be tuned for those of us in a different country from an S3 endpoint.","created":"2016-05-18T19:50:11.511+0000"},{"body":"Interestingly, you're getting a much slower run than me for S3A and a much faster run than me for WASB. I'm in the US Pacific Northwest. My S3 bucket is in US-west-2. My Azure Storage account is in West US.\n\n{code}\nRunning org.apache.hadoop.fs.azure.contract.TestAzureNativeContractDistCp\nTests run: 8, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 140.389 sec - in org.apache.hadoop.fs.azure.contract.TestAzureNativeContractDistCp\n\nRunning org.apache.hadoop.fs.contract.s3a.TestS3AContractDistCp\nTests run: 8, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 143.99 sec - in org.apache.hadoop.fs.contract.s3a.TestS3AContractDistCp\n{code}\n\nbq. Could it be made one of the scaleable tests where it takes a config of option on scale so can be made configurable?\n\nWe definitely could do that, but in my test runs, the large file tests don't show a significantly longer execution time. (See below for my timings.) Are the large file tests a long haul in your environment?\n\nMaybe a more effective change would be to cut down the number of test cases. I could keep just {{deepDirectoryStructureToRemote}}, {{largeFilesToRemote}}, {{deepDirectoryStructureFromRemote}} and {{largeFilesFromRemote}}. If I do that, then my S3A execution time comes down to 90 seconds. I don't think it sacrifices much in terms of coverage.\n\nLet me know your thoughts, and then I'll update the patch.\n\n{code}\n \n \n \n \n \n \n \n \n{code}\n","created":"2016-05-18T23:22:19.685+0000"},{"body":"There's no S3 service in my country, I need to test against a datacentre in a country with a lower tax regime yet still under EU data protection legislation coverage. Ireland; I could benchmark Frankfurt.\n\nIf you think the large files repeat the same coverage as the smaller ones, yes, please unify. Even so, I'd like it to be configurable so that I could set up test runs with smaller datasets —and we have the option of test runs with larger files.\n\nFor those test, it'd be nice if the S3A setup explicitly turned the multipart threshold down (8MB?) and the same for partition sizes, so that it'd test the multipart code path and distcp","created":"2016-05-19T09:38:33.713+0000"},{"body":"BTW, Azure is in ireland too; the performance difference there is clearly not the pipe width and length at my end. Either it is S3 *or* it is how S3A talks to S3","created":"2016-05-19T09:40:38.746+0000"},{"body":"I'm attaching patch v004.\n* Removed redundant single-file tests and small multi-file tests.\n* Introduced {{scale.test.distcp.file.size.kb}} configuration property for tuning test file sizes. The default is 10 MB.\n* Set multi-part configuration properties to 8 MB, so with the default 10 MB file size, the tests will cover multi-part upload.\n\nWith this version of the patch, the S3A test runs in ~55 seconds for me, and the WASB test runs in ~65 seconds. I completed a full parallel-test run against S3 buckets in US-west-2.","created":"2016-05-20T06:50:35.169+0000"},{"body":"+1\n\nlatest patch brings test time down to <60s including all JUnit overhead.\n\nThanks for doing this Chris, especially the tests. They'll be a good bit of regression testing in future","created":"2016-05-20T11:23:47.085+0000"},{"body":"Patch as is breaks 2.8, \n\n{code}\n\n[INFO] ------------------------------------------------------------------------\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.1:testCompile (default-testCompile) on project hadoop-distcp: Compilation failure\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-tools/hadoop-distcp/src/test/java/org/apache/hadoop/tools/contract/AbstractContractDistCpTest.java:[80,25] cannot find symbol\n[ERROR] symbol: method getTestDir()\n[ERROR] location: class org.apache.hadoop.test.GenericTestUtils\n[ERROR] -> [Help 1]\n[ERROR] \n[ERROR] To see the full stack trace of the errors, re-run Maven with the -e switch.\n[ERROR] Re-run Maven using the -X switch to enable full debug logging.\n[ERROR] \n[ERROR] For more information about the errors and possible solutions, please read the following articles:\n[ERROR] [Help 1] http://cwiki.apache.org/confluence/display/MAVEN/MojoFailureException\n[ERROR] \n[ERROR] After correcting the problems, you can resume the build with the command\n{code}\n\nI've reverted the 2.8 patch; left it in branch-2+. Reopening so that a 2.8 specific patch can be added.","created":"2016-05-20T13:04:34.401+0000"},{"body":"Steve, thank you for catching the branch-2.8 problem and reverting. Sorry I didn't catch it myself earlier.\n\nI'm attaching a branch-2.8 patch. {{GenericTestUtils#getTestDir}} was introduced in your HADOOP-12984 patch, targeted to 2.9.0. That's a sizable patch, and I don't want to take on a back-port right now. Instead, this branch-2.8 patch goes back to the pre-HADOOP-12984 strategy of individual tests reading the {{test.build.data}} property directly.","created":"2016-05-20T17:55:23.133+0000"},{"body":"+1\n\napplied the branch-2.8 patch, tested, all is well.","created":"2016-05-21T18:10:58.409+0000"},{"body":"Any chance of creating a patch/applying to 2.7 branch?","created":"2017-06-06T00:11:49.858+0000"},{"body":"we're generally pretty reluctant to put stuff into 2.7.x which isn't a fairly significant bug fix, just because every change is at a risk of breaking something, and we like 2.7.x to be stable. There's enough change in this patch (poms, code, new tests) that it's got the potential to cause trouble\n\nHadoop 2.8.1 is/will be out in the next few days, could you work with that? ","created":"2017-06-06T12:21:29.597+0000"},{"body":"We're using Spark that is pre-built to 2.7 but I can try building Spark against 2.8.1 when it's released to see how it goes.","created":"2017-06-06T17:59:42.838+0000"},{"body":"[~stevel@apache.org] Any idea when 2.8.1 will be released?","created":"2017-07-10T18:28:12.411+0000"},{"body":"a 2.81. did come out, but security related. Go with 2.8.0 for now","created":"2017-07-11T10:28:40.877+0000"}],"conversations":[{"body":"After DistCp copies a file, it calls {{getFileStatus}} to get the {{FileStatus}} from the destination so that it can compare to the source and update metadata if necessary. If the DistCp command was run without the option to preserve metadata attributes, then this additional {{getFileStatus}} call is wasteful.","from":"reporter","subject":"In DistCp, prevent unnecessary getFileStatus call when not preserving metadata."},{"body":"The attached v001 patch avoids the unnecessary {{getFileStatus}} call.\n\nThe effect is particularly pronounced when running DistCp with a destination on S3A, where eventual consistency on S3 can cause the {{getFileStatus}} call to fail with {{FileNotFoundException}}. Then, the whole MapReduce task fails, retries, and repeats copying all the data. [~rajesh.balamohan], I know you saw this with some recent large copies to S3A. Would you be interested in trying a test with this patch? So far, I don't have my own repro. Note that this patch is only helpful as long as the DistCp command is not preserving metadata attributes, so don't use the {{-p}} option.\n\nCc [~stevel@apache.org].","from":"developer"},{"body":"Thanks for sharing the patch [~cnauroth]. I tried out the patch with hadoop-2.8 (may 16) and I do not see any more task failures with dist-cp to S3. I see some straggler tasks taking time, but that is completely unrelated to this jira.","from":"developer"},{"body":"You know, I think s3a now has enough instrumentation that the # of times that getFileStatus is called would be measurable. \n\nAt the very least, it'd be good to have a test of DistCp there, to verify that inconsistency problems aren't surfacing. The examples in, say {{TestDistCpViewFs}} , show a start, though I'd expect the new tests to simply throw up IOEs, rather than swallow + fail, the way that class does (and which I have just submitted a patch for, in HADOOP-13148).","from":"developer"},{"body":"Patch v003 adds a new abstract contract test suite for DistCp coverage and concrete test suite subclasses for S3A and WASB. I verified the tests are passing for both hadoop-aws (including running in parallel mode) and hadoop-azure.\n\nI'm going to leave the JIRA issue in Open status instead of Patch Available for now. The v003 patch will potentially hit Jenkins a little hard because of touching multiple modules, so I'd like to get another round of code review feedback first.\n","from":"developer"},{"body":"tested -003 against s3 ireland and azure.\n\n{code}\nRunning org.apache.hadoop.fs.contract.s3a.TestS3AContractDistCp\nTests run: 8, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 223.843 sec - in org.apache.hadoop.fs.contract.s3a.TestS3AContractDistCp\n\n...\nRunning org.apache.hadoop.fs.azure.contract.TestAzureNativeContractDistCp\nTests run: 8, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 65.354 sec - in org.apache.hadoop.fs.azure.contract.TestAzureNativeContractDistCp\n\n{code}\n\nInteresting how much faster azure is. \n\nThe patch, is, as it stands, it's going to add 4 min to a TestS3A* test pattern. Could it be made one of the scaleable tests where it takes a config of option on scale so can be made configurable? There are already some tests which use {{scale.test.operation.count}} to control scale; we could have one on distcp file size, with the large file size being driven by it. Make it something in KB and it could easily be tuned for those of us in a different country from an S3 endpoint.","from":"developer"},{"body":"Interestingly, you're getting a much slower run than me for S3A and a much faster run than me for WASB. I'm in the US Pacific Northwest. My S3 bucket is in US-west-2. My Azure Storage account is in West US.\n\n{code}\nRunning org.apache.hadoop.fs.azure.contract.TestAzureNativeContractDistCp\nTests run: 8, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 140.389 sec - in org.apache.hadoop.fs.azure.contract.TestAzureNativeContractDistCp\n\nRunning org.apache.hadoop.fs.contract.s3a.TestS3AContractDistCp\nTests run: 8, Failures: 0, Errors: 0, Skipped: 0, Time elapsed: 143.99 sec - in org.apache.hadoop.fs.contract.s3a.TestS3AContractDistCp\n{code}\n\nbq. Could it be made one of the scaleable tests where it takes a config of option on scale so can be made configurable?\n\nWe definitely could do that, but in my test runs, the large file tests don't show a significantly longer execution time. (See below for my timings.) Are the large file tests a long haul in your environment?\n\nMaybe a more effective change would be to cut down the number of test cases. I could keep just {{deepDirectoryStructureToRemote}}, {{largeFilesToRemote}}, {{deepDirectoryStructureFromRemote}} and {{largeFilesFromRemote}}. If I do that, then my S3A execution time comes down to 90 seconds. I don't think it sacrifices much in terms of coverage.\n\nLet me know your thoughts, and then I'll update the patch.\n\n{code}\n \n \n \n \n \n \n \n \n{code}\n","from":"developer"},{"body":"There's no S3 service in my country, I need to test against a datacentre in a country with a lower tax regime yet still under EU data protection legislation coverage. Ireland; I could benchmark Frankfurt.\n\nIf you think the large files repeat the same coverage as the smaller ones, yes, please unify. Even so, I'd like it to be configurable so that I could set up test runs with smaller datasets —and we have the option of test runs with larger files.\n\nFor those test, it'd be nice if the S3A setup explicitly turned the multipart threshold down (8MB?) and the same for partition sizes, so that it'd test the multipart code path and distcp","from":"developer"},{"body":"BTW, Azure is in ireland too; the performance difference there is clearly not the pipe width and length at my end. Either it is S3 *or* it is how S3A talks to S3","from":"developer"},{"body":"I'm attaching patch v004.\n* Removed redundant single-file tests and small multi-file tests.\n* Introduced {{scale.test.distcp.file.size.kb}} configuration property for tuning test file sizes. The default is 10 MB.\n* Set multi-part configuration properties to 8 MB, so with the default 10 MB file size, the tests will cover multi-part upload.\n\nWith this version of the patch, the S3A test runs in ~55 seconds for me, and the WASB test runs in ~65 seconds. I completed a full parallel-test run against S3 buckets in US-west-2.","from":"developer"},{"body":"+1\n\nlatest patch brings test time down to <60s including all JUnit overhead.\n\nThanks for doing this Chris, especially the tests. They'll be a good bit of regression testing in future","from":"developer"},{"body":"Patch as is breaks 2.8, \n\n{code}\n\n[INFO] ------------------------------------------------------------------------\n[ERROR] Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.1:testCompile (default-testCompile) on project hadoop-distcp: Compilation failure\n[ERROR] /Users/stevel/Projects/hadoop-trunk/hadoop-tools/hadoop-distcp/src/test/java/org/apache/hadoop/tools/contract/AbstractContractDistCpTest.java:[80,25] cannot find symbol\n[ERROR] symbol: method getTestDir()\n[ERROR] location: class org.apache.hadoop.test.GenericTestUtils\n[ERROR] -> [Help 1]\n[ERROR] \n[ERROR] To see the full stack trace of the errors, re-run Maven with the -e switch.\n[ERROR] Re-run Maven using the -X switch to enable full debug logging.\n[ERROR] \n[ERROR] For more information about the errors and possible solutions, please read the following articles:\n[ERROR] [Help 1] http://cwiki.apache.org/confluence/display/MAVEN/MojoFailureException\n[ERROR] \n[ERROR] After correcting the problems, you can resume the build with the command\n{code}\n\nI've reverted the 2.8 patch; left it in branch-2+. Reopening so that a 2.8 specific patch can be added.","from":"developer"},{"body":"Steve, thank you for catching the branch-2.8 problem and reverting. Sorry I didn't catch it myself earlier.\n\nI'm attaching a branch-2.8 patch. {{GenericTestUtils#getTestDir}} was introduced in your HADOOP-12984 patch, targeted to 2.9.0. That's a sizable patch, and I don't want to take on a back-port right now. Instead, this branch-2.8 patch goes back to the pre-HADOOP-12984 strategy of individual tests reading the {{test.build.data}} property directly.","from":"developer"},{"body":"+1\n\napplied the branch-2.8 patch, tested, all is well.","from":"developer"},{"body":"Any chance of creating a patch/applying to 2.7 branch?","from":"developer"},{"body":"we're generally pretty reluctant to put stuff into 2.7.x which isn't a fairly significant bug fix, just because every change is at a risk of breaking something, and we like 2.7.x to be stable. There's enough change in this patch (poms, code, new tests) that it's got the potential to cause trouble\n\nHadoop 2.8.1 is/will be out in the next few days, could you work with that? ","from":"developer"},{"body":"We're using Spark that is pre-built to 2.7 but I can try building Spark against 2.8.1 when it's released to see how it goes.","from":"developer"},{"body":"[~stevel@apache.org] Any idea when 2.8.1 will be released?","from":"developer"},{"body":"a 2.81. did come out, but security related. Go with 2.8.0 for now","from":"developer"}],"created":"2016-05-13T22:52:41.000+0000","description":"After DistCp copies a file, it calls {{getFileStatus}} to get the {{FileStatus}} from the destination so that it can compare to the source and update metadata if necessary. If the DistCp command was run without the option to preserve metadata attributes, then this additional {{getFileStatus}} call is wasteful.","issue_id":"12969687","key":"HADOOP-13145","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2016-05-21T18:10:58.000+0000","role":"gold_target","summary":"In DistCp, prevent unnecessary getFileStatus call when not preserving metadata."} {"case_id":"13215528","cluster":"JIRA-HADOOP-0de7dde1714d","comments":[{"body":"thanks, assigned to to me and I'll look next week","created":"2019-02-14T17:09:28.053+0000"},{"body":"yeah, I can replicate this in a test\r\n{code:java}\r\n[ERROR] testReadPastReadahead(org.apache.hadoop.fs.contract.s3a.ITestS3AContractRandomSeek) Time elapsed: 3.061 s <<< ERROR!\r\njava.io.EOFException: End of file reached before reading fully.\r\n\tat org.apache.hadoop.fs.s3a.S3AInputStream.readFully(S3AInputStream.java:707)\r\n\tat org.apache.hadoop.fs.FSDataInputStream.readFully(FSDataInputStream.java:121)\r\n\tat org.apache.hadoop.fs.contract.s3a.ITestS3AContractSeek.testReadPastReadahead(ITestS3AContractSeek.java:79)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n\tat java.lang.reflect.Method.invoke(Method.java:498)\r\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:50)\r\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\r\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:47)\r\n\tat org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\r\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:26)\r\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:27)\r\n\tat org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:55)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:298)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:292)\r\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n\tat java.lang.Thread.run(Thread.java:745)\r\n\r\n[INFO] \r\n\r\n {code}","created":"2019-02-28T16:12:31.564+0000"},{"body":"Hi [~stevel@apache.org], one of my colleagues, Shruti Gumma, has a proposed fix which I'll help him post here today.\r\n\r\nEdit: Shruti is unable to test on S3, so requests you proceed with your fix.","created":"2019-02-28T18:46:48.965+0000"},{"body":"Matt, love to see. it just stuck my own PR up too: lets compare! And if there are extra tests, pull them in.\r\n\r\nRoot cause: using <= over = in the decision making about whether to skip vs close. The situation which triggered the failure was\r\n\r\n* random IO mode (i.e. shorter reads)\r\n* active read\r\n* next read spanned the current active read but went beyond.\r\n\r\nOne like to fix, one for extra debug log, parameterized tests for regression. \r\nThis is going to need backporting to 2.8.+","created":"2019-02-28T19:55:24.338+0000"},{"body":"Yes, I'm thinking at [https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/S3AInputStream.java#L264] we need\r\n\r\n{{&& diff < forwardSeekLimit; // instead of <=}}\r\n\r\nWhat do you think?\r\n\r\nThe big question I have is, some of the text description talks about \"reading past the already active readahead range\", i.e., past {{remainingInCurrentRequest}}, as being a problem, but it seems to me that should be okay; the problem documented so far is *seeking* past {{remainingInCurrentRequest}} (specifically to exactly the end of CurrentRequest, which is incorrectly guarded by the above L264 inequality) and then not doing a stream close, which causes the problem.  Do you know if, say, seeking to a few bytes before the end of CurrentRequest, then reading past it (when the S3 file does indeed have more to read), also causes an EOF, or does the stream machinery handle that case correctly?\r\n\r\nI'm putting together a test platform so I can answer such questions myself, but it will take me a few hours; I haven't worked in s3a before.\r\n\r\n ","created":"2019-03-01T00:43:45.474+0000"},{"body":"Let me add a test for that too","created":"2019-03-04T12:34:59.147+0000"},{"body":"bq. Do you know if, say, seeking to a few bytes before the end of CurrentRequest, then reading past it (when the S3 file does indeed have more to read), also causes an EOF, or does the stream machinery handle that case correctly?\r\n\r\nmatt, this is exactly what is being tested for.\r\nLet me do a variant with a seek to just before EOF and some readbyte and read(byte[]) calls to verify that too.\r\n\r\nAnd yes, a seek for utter completeness","created":"2019-03-04T14:31:58.323+0000"},{"body":"latest PR update adds the tests & makes sure that each parameterized run is switching seek policies","created":"2019-03-04T16:48:45.627+0000"},{"body":"Patch is in trunk; will do a sequence of backports ","created":"2019-03-09T16:30:29.115+0000"},{"body":"now merged into branch-3.0. \r\n\r\nNote in hadoop <= 3.0 normal ==> sequential so no issues there; the parameterized test could be cut down to two. I'll do that for the branch-2 version just to reduce test time, while leaving consistent across 3.x","created":"2019-03-12T11:40:45.645+0000"},{"body":"branch-2 patch in & backported to 2.9 and eventually 2.8; seek tests rerun each time. All happy. Closing as done. \r\n\r\ntrivia: this is be the first patch I get to submit patches for *everywhere*. I am so happy.","created":"2019-03-13T15:58:35.165+0000"},{"body":"Thanks very much [~stevel@apache.org]!","created":"2019-03-13T17:18:30.872+0000"},{"body":"Where do I get started to backport this fix into 3.1.0?","created":"2020-12-22T19:48:52.063+0000"},{"body":"It's in branch-3.1; hadoop-3.1.4 at least.\r\n\r\nIf you want to do it into your own fork of 3.1.0; check out that release then cherrypick the commit which went into branch 3.1( commit 1bace86501ae48c09ff01c2409e6bf6b3cad5408 ) and build a release.","created":"2020-12-23T17:45:59.287+0000"},{"body":"I think the 3.1.0 jars I am getting in my build do not have the fix. Anyways, ended up bumping to 3.1.3 which had the fix.","created":"2020-12-23T18:46:31.086+0000"}],"conversations":[{"body":"When using S3AFileSystem to read Parquet files a specific set of circumstances causes an  EOFException that is not thrown when reading the same file from local disk\r\n\r\nNote this has only been observed under specific circumstances:\r\n  - when the reader is doing a projection (will cause it to do a seek backwards and put the filesystem into random mode)\r\n - when the file is larger than the readahead buffer size\r\n - when the seek behavior of the Parquet reader causes the reader to seek towards the end of the current input stream without reopening, such that the next read on the currently open stream will read past the end of the currently open stream.\r\n\r\nException from Parquet reader is as follows:\r\n{code}\r\nCaused by: java.io.EOFException: Reached the end of stream with 51 bytes left to read\r\n at org.apache.parquet.io.DelegatingSeekableInputStream.readFully(DelegatingSeekableInputStream.java:104)\r\n at org.apache.parquet.io.DelegatingSeekableInputStream.readFullyHeapBuffer(DelegatingSeekableInputStream.java:127)\r\n at org.apache.parquet.io.DelegatingSeekableInputStream.readFully(DelegatingSeekableInputStream.java:91)\r\n at org.apache.parquet.hadoop.ParquetFileReader$ConsecutiveChunkList.readAll(ParquetFileReader.java:1174)\r\n at org.apache.parquet.hadoop.ParquetFileReader.readNextRowGroup(ParquetFileReader.java:805)\r\n at org.apache.parquet.hadoop.InternalParquetRecordReader.checkRead(InternalParquetRecordReader.java:127)\r\n at org.apache.parquet.hadoop.InternalParquetRecordReader.nextKeyValue(InternalParquetRecordReader.java:222)\r\n at org.apache.parquet.hadoop.ParquetRecordReader.nextKeyValue(ParquetRecordReader.java:207)\r\n at org.apache.flink.api.java.hadoop.mapreduce.HadoopInputFormatBase.fetchNext(HadoopInputFormatBase.java:206)\r\n at org.apache.flink.api.java.hadoop.mapreduce.HadoopInputFormatBase.reachedEnd(HadoopInputFormatBase.java:199)\r\n at org.apache.flink.runtime.operators.DataSourceTask.invoke(DataSourceTask.java:190)\r\n at org.apache.flink.runtime.taskmanager.Task.run(Task.java:711)\r\n at java.lang.Thread.run(Thread.java:748)\r\n{code}\r\nThe following example program generate the same root behavior (sans finding a Parquet file that happens to trigger this condition) by purposely reading past the already active readahead range on any file >= 1029 bytes in size.. \r\n\r\n\r\n{code:java}\r\nfinal Configuration conf = new Configuration();\r\nconf.set(\"fs.s3a.readahead.range\", \"1K\");\r\nconf.set(\"fs.s3a.experimental.input.fadvise\", \"random\");\r\n\r\nfinal FileSystem fs = FileSystem.get(path.toUri(), conf);\r\n// forward seek reading across readahead boundary\r\ntry (FSDataInputStream in = fs.open(path)) {\r\n final byte[] temp = new byte[5];\r\n in.readByte();\r\n in.readFully(1023, temp); // <-- works\r\n}\r\n// forward seek reading from end of readahead boundary\r\ntry (FSDataInputStream in = fs.open(path)) {\r\n final byte[] temp = new byte[5];\r\n in.readByte();\r\n in.readFully(1024, temp); // <-- throws EOFException\r\n}\r\n{code}\r\n ","from":"reporter","subject":"Parquet reading S3AFileSystem causes EOF"},{"body":"thanks, assigned to to me and I'll look next week","from":"developer"},{"body":"yeah, I can replicate this in a test\r\n{code:java}\r\n[ERROR] testReadPastReadahead(org.apache.hadoop.fs.contract.s3a.ITestS3AContractRandomSeek) Time elapsed: 3.061 s <<< ERROR!\r\njava.io.EOFException: End of file reached before reading fully.\r\n\tat org.apache.hadoop.fs.s3a.S3AInputStream.readFully(S3AInputStream.java:707)\r\n\tat org.apache.hadoop.fs.FSDataInputStream.readFully(FSDataInputStream.java:121)\r\n\tat org.apache.hadoop.fs.contract.s3a.ITestS3AContractSeek.testReadPastReadahead(ITestS3AContractSeek.java:79)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\r\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\r\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\r\n\tat java.lang.reflect.Method.invoke(Method.java:498)\r\n\tat org.junit.runners.model.FrameworkMethod$1.runReflectiveCall(FrameworkMethod.java:50)\r\n\tat org.junit.internal.runners.model.ReflectiveCallable.run(ReflectiveCallable.java:12)\r\n\tat org.junit.runners.model.FrameworkMethod.invokeExplosively(FrameworkMethod.java:47)\r\n\tat org.junit.internal.runners.statements.InvokeMethod.evaluate(InvokeMethod.java:17)\r\n\tat org.junit.internal.runners.statements.RunBefores.evaluate(RunBefores.java:26)\r\n\tat org.junit.internal.runners.statements.RunAfters.evaluate(RunAfters.java:27)\r\n\tat org.junit.rules.TestWatcher$1.evaluate(TestWatcher.java:55)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:298)\r\n\tat org.junit.internal.runners.statements.FailOnTimeout$CallableStatement.call(FailOnTimeout.java:292)\r\n\tat java.util.concurrent.FutureTask.run(FutureTask.java:266)\r\n\tat java.lang.Thread.run(Thread.java:745)\r\n\r\n[INFO] \r\n\r\n {code}","from":"developer"},{"body":"Hi [~stevel@apache.org], one of my colleagues, Shruti Gumma, has a proposed fix which I'll help him post here today.\r\n\r\nEdit: Shruti is unable to test on S3, so requests you proceed with your fix.","from":"developer"},{"body":"Matt, love to see. it just stuck my own PR up too: lets compare! And if there are extra tests, pull them in.\r\n\r\nRoot cause: using <= over = in the decision making about whether to skip vs close. The situation which triggered the failure was\r\n\r\n* random IO mode (i.e. shorter reads)\r\n* active read\r\n* next read spanned the current active read but went beyond.\r\n\r\nOne like to fix, one for extra debug log, parameterized tests for regression. \r\nThis is going to need backporting to 2.8.+","from":"developer"},{"body":"Yes, I'm thinking at [https://github.com/apache/hadoop/blob/trunk/hadoop-tools/hadoop-aws/src/main/java/org/apache/hadoop/fs/s3a/S3AInputStream.java#L264] we need\r\n\r\n{{&& diff < forwardSeekLimit; // instead of <=}}\r\n\r\nWhat do you think?\r\n\r\nThe big question I have is, some of the text description talks about \"reading past the already active readahead range\", i.e., past {{remainingInCurrentRequest}}, as being a problem, but it seems to me that should be okay; the problem documented so far is *seeking* past {{remainingInCurrentRequest}} (specifically to exactly the end of CurrentRequest, which is incorrectly guarded by the above L264 inequality) and then not doing a stream close, which causes the problem.  Do you know if, say, seeking to a few bytes before the end of CurrentRequest, then reading past it (when the S3 file does indeed have more to read), also causes an EOF, or does the stream machinery handle that case correctly?\r\n\r\nI'm putting together a test platform so I can answer such questions myself, but it will take me a few hours; I haven't worked in s3a before.\r\n\r\n ","from":"developer"},{"body":"Let me add a test for that too","from":"developer"},{"body":"bq. Do you know if, say, seeking to a few bytes before the end of CurrentRequest, then reading past it (when the S3 file does indeed have more to read), also causes an EOF, or does the stream machinery handle that case correctly?\r\n\r\nmatt, this is exactly what is being tested for.\r\nLet me do a variant with a seek to just before EOF and some readbyte and read(byte[]) calls to verify that too.\r\n\r\nAnd yes, a seek for utter completeness","from":"developer"},{"body":"latest PR update adds the tests & makes sure that each parameterized run is switching seek policies","from":"developer"},{"body":"Patch is in trunk; will do a sequence of backports ","from":"developer"},{"body":"now merged into branch-3.0. \r\n\r\nNote in hadoop <= 3.0 normal ==> sequential so no issues there; the parameterized test could be cut down to two. I'll do that for the branch-2 version just to reduce test time, while leaving consistent across 3.x","from":"developer"},{"body":"branch-2 patch in & backported to 2.9 and eventually 2.8; seek tests rerun each time. All happy. Closing as done. \r\n\r\ntrivia: this is be the first patch I get to submit patches for *everywhere*. I am so happy.","from":"developer"},{"body":"Thanks very much [~stevel@apache.org]!","from":"developer"},{"body":"Where do I get started to backport this fix into 3.1.0?","from":"developer"},{"body":"It's in branch-3.1; hadoop-3.1.4 at least.\r\n\r\nIf you want to do it into your own fork of 3.1.0; check out that release then cherrypick the commit which went into branch 3.1( commit 1bace86501ae48c09ff01c2409e6bf6b3cad5408 ) and build a release.","from":"developer"},{"body":"I think the 3.1.0 jars I am getting in my build do not have the fix. Anyways, ended up bumping to 3.1.3 which had the fix.","from":"developer"}],"created":"2019-02-13T15:37:53.000+0000","description":"When using S3AFileSystem to read Parquet files a specific set of circumstances causes an  EOFException that is not thrown when reading the same file from local disk\r\n\r\nNote this has only been observed under specific circumstances:\r\n  - when the reader is doing a projection (will cause it to do a seek backwards and put the filesystem into random mode)\r\n - when the file is larger than the readahead buffer size\r\n - when the seek behavior of the Parquet reader causes the reader to seek towards the end of the current input stream without reopening, such that the next read on the currently open stream will read past the end of the currently open stream.\r\n\r\nException from Parquet reader is as follows:\r\n{code}\r\nCaused by: java.io.EOFException: Reached the end of stream with 51 bytes left to read\r\n at org.apache.parquet.io.DelegatingSeekableInputStream.readFully(DelegatingSeekableInputStream.java:104)\r\n at org.apache.parquet.io.DelegatingSeekableInputStream.readFullyHeapBuffer(DelegatingSeekableInputStream.java:127)\r\n at org.apache.parquet.io.DelegatingSeekableInputStream.readFully(DelegatingSeekableInputStream.java:91)\r\n at org.apache.parquet.hadoop.ParquetFileReader$ConsecutiveChunkList.readAll(ParquetFileReader.java:1174)\r\n at org.apache.parquet.hadoop.ParquetFileReader.readNextRowGroup(ParquetFileReader.java:805)\r\n at org.apache.parquet.hadoop.InternalParquetRecordReader.checkRead(InternalParquetRecordReader.java:127)\r\n at org.apache.parquet.hadoop.InternalParquetRecordReader.nextKeyValue(InternalParquetRecordReader.java:222)\r\n at org.apache.parquet.hadoop.ParquetRecordReader.nextKeyValue(ParquetRecordReader.java:207)\r\n at org.apache.flink.api.java.hadoop.mapreduce.HadoopInputFormatBase.fetchNext(HadoopInputFormatBase.java:206)\r\n at org.apache.flink.api.java.hadoop.mapreduce.HadoopInputFormatBase.reachedEnd(HadoopInputFormatBase.java:199)\r\n at org.apache.flink.runtime.operators.DataSourceTask.invoke(DataSourceTask.java:190)\r\n at org.apache.flink.runtime.taskmanager.Task.run(Task.java:711)\r\n at java.lang.Thread.run(Thread.java:748)\r\n{code}\r\nThe following example program generate the same root behavior (sans finding a Parquet file that happens to trigger this condition) by purposely reading past the already active readahead range on any file >= 1029 bytes in size.. \r\n\r\n\r\n{code:java}\r\nfinal Configuration conf = new Configuration();\r\nconf.set(\"fs.s3a.readahead.range\", \"1K\");\r\nconf.set(\"fs.s3a.experimental.input.fadvise\", \"random\");\r\n\r\nfinal FileSystem fs = FileSystem.get(path.toUri(), conf);\r\n// forward seek reading across readahead boundary\r\ntry (FSDataInputStream in = fs.open(path)) {\r\n final byte[] temp = new byte[5];\r\n in.readByte();\r\n in.readFully(1023, temp); // <-- works\r\n}\r\n// forward seek reading from end of readahead boundary\r\ntry (FSDataInputStream in = fs.open(path)) {\r\n final byte[] temp = new byte[5];\r\n in.readByte();\r\n in.readFully(1024, temp); // <-- throws EOFException\r\n}\r\n{code}\r\n ","issue_id":"13215528","key":"HADOOP-16109","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2019-03-13T15:58:45.000+0000","role":"gold_target","summary":"Parquet reading S3AFileSystem causes EOF"} {"case_id":"13315751","cluster":"JIRA-HADOOP-d3dfc3a59a1f","comments":[{"body":"{noformat}\r\nStep 26/37 : RUN pip2 install configparser==4.0.2 pylint==1.9.2\r\n ---> Running in ee3598ed6a5a\r\nCollecting configparser==4.0.2\r\n Downloading https://files.pythonhosted.org/packages/7a/2a/95ed0501cf5d8709490b1d3a3f9b5cf340da6c433f896bbe9ce08dbe6785/configparser-4.0.2-py2.py3-none-any.whl\r\nCollecting pylint==1.9.2\r\n Downloading https://files.pythonhosted.org/packages/f2/95/0ca03c818ba3cd14f2dd4e95df5b7fa232424b7fc6ea1748d27f293bc007/pylint-1.9.2-py2.py3-none-any.whl (690kB)\r\nCollecting singledispatch; python_version < \"3.4\" (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/c5/10/369f50bcd4621b263927b0a1519987a04383d4a98fb10438042ad410cf88/singledispatch-3.4.0.3-py2.py3-none-any.whl\r\nCollecting isort>=4.2.5 (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/c4/4d/b6286cf463f9cfca698b15524e1198856d68080096f05ca7b3437f0af867/isort-5.0.5.tar.gz (77kB)\r\nCollecting backports.functools-lru-cache; python_version == \"2.7\" (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/da/d1/080d2bb13773803648281a49e3918f65b31b7beebf009887a529357fd44a/backports.functools_lru_cache-1.6.1-py2.py3-none-any.whl\r\nCollecting mccabe (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/87/89/479dc97e18549e21354893e4ee4ef36db1d237534982482c3681ee6e7b57/mccabe-0.6.1-py2.py3-none-any.whl\r\nCollecting astroid<2.0,>=1.6 (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/8b/29/0f7ec6fbf28a158886b7de49aee3a77a8a47a7e24c60e9fd6ec98ee2ec02/astroid-1.6.6-py2.py3-none-any.whl (305kB)\r\nCollecting six (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/ee/ff/48bde5c0f013094d729fe4b0316ba2a24774b3ff1c52d924a8a4cb04078a/six-1.15.0-py2.py3-none-any.whl\r\nCollecting enum34>=1.1.3; python_version < \"3.4\" (from astroid<2.0,>=1.6->pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/6f/2c/a9386903ece2ea85e9807e0e062174dc26fdce8b05f216d00491be29fad5/enum34-1.1.10-py2-none-any.whl\r\nCollecting wrapt (from astroid<2.0,>=1.6->pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/82/f7/e43cefbe88c5fd371f4cf0cf5eb3feccd07515af9fd6cf7dbf1d1793a797/wrapt-1.12.1.tar.gz\r\nCollecting lazy-object-proxy (from astroid<2.0,>=1.6->pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/f3/1f/3e31313f557e0b97bd8f9716f502fa85c5fef181f582f816a4796b8f9ee1/lazy_object_proxy-1.5.0-cp27-cp27mu-manylinux1_x86_64.whl (55kB)\r\nBuilding wheels for collected packages: isort, wrapt\r\n Running setup.py bdist_wheel for isort: started\r\n Running setup.py bdist_wheel for isort: finished with status 'error'\r\n Complete output from command /usr/bin/python -u -c \"import setuptools, tokenize;__file__='/tmp/pip-build-ayvPD4/isort/setup.py';exec(compile(getattr(tokenize, 'open', open)(__file__).read().replace('\\r\\n', '\\n'), __file__, 'exec'))\" bdist_wheel -d /tmp/tmpzmXICSpip-wheel- --python-tag cp27:\r\n /usr/lib/python2.7/distutils/dist.py:267: UserWarning: Unknown distribution option: 'python_requires'\r\n warnings.warn(msg)\r\n running bdist_wheel\r\n running build\r\n running build_py\r\n creating build\r\n creating build/lib.linux-x86_64-2.7\r\n creating build/lib.linux-x86_64-2.7/isort\r\n copying isort/sorting.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/setuptools_commands.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/logo.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/wrap_modes.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/__main__.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/pylama_isort.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/hooks.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/main.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/profiles.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/exceptions.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/_version.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/io.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/format.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/output.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/api.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/utils.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/__init__.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/parse.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/comments.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/wrap.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/place.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/settings.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/sections.py -> build/lib.linux-x86_64-2.7/isort\r\n creating build/lib.linux-x86_64-2.7/isort/_future\r\n copying isort/_future/_dataclasses.py -> build/lib.linux-x86_64-2.7/isort/_future\r\n copying isort/_future/__init__.py -> build/lib.linux-x86_64-2.7/isort/_future\r\n creating build/lib.linux-x86_64-2.7/isort/_vendored\r\n creating build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/decoder.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/encoder.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/ordered.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/tz.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/__init__.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n creating build/lib.linux-x86_64-2.7/isort/deprecated\r\n copying isort/deprecated/finders.py -> build/lib.linux-x86_64-2.7/isort/deprecated\r\n copying isort/deprecated/__init__.py -> build/lib.linux-x86_64-2.7/isort/deprecated\r\n creating build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py35.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py2.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py3.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py36.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py37.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py27.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py38.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/__init__.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/all.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py39.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n error: can't copy 'isort/stdlibs': doesn't exist or not a regular file\r\n \r\n ----------------------------------------\r\n Failed building wheel for isort\r\n Running setup.py clean for isort\r\n Running setup.py bdist_wheel for wrapt: started\r\n Running setup.py bdist_wheel for wrapt: finished with status 'done'\r\n Stored in directory: /root/.cache/pip/wheels/b1/c2/ed/d62208260edbd3fa7156545c00ef966f45f2063d0a84f8208a\r\nSuccessfully built wrapt\r\nFailed to build isort\r\nInstalling collected packages: configparser, six, singledispatch, isort, backports.functools-lru-cache, mccabe, enum34, wrapt, lazy-object-proxy, astroid, pylint\r\n Running setup.py install for isort: started\r\n Running setup.py install for isort: finished with status 'error'\r\n Complete output from command /usr/bin/python -u -c \"import setuptools, tokenize;__file__='/tmp/pip-build-ayvPD4/isort/setup.py';exec(compile(getattr(tokenize, 'open', open)(__file__).read().replace('\\r\\n', '\\n'), __file__, 'exec'))\" install --record /tmp/pip-ulpe3F-record/install-record.txt --single-version-externally-managed --compile:\r\n /usr/lib/python2.7/distutils/dist.py:267: UserWarning: Unknown distribution option: 'python_requires'\r\n warnings.warn(msg)\r\n running install\r\n running build\r\n running build_py\r\n creating build\r\n creating build/lib.linux-x86_64-2.7\r\n creating build/lib.linux-x86_64-2.7/isort\r\n copying isort/sorting.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/setuptools_commands.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/logo.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/wrap_modes.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/__main__.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/pylama_isort.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/hooks.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/main.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/profiles.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/exceptions.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/_version.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/io.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/format.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/output.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/api.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/utils.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/__init__.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/parse.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/comments.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/wrap.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/place.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/settings.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/sections.py -> build/lib.linux-x86_64-2.7/isort\r\n creating build/lib.linux-x86_64-2.7/isort/_future\r\n copying isort/_future/_dataclasses.py -> build/lib.linux-x86_64-2.7/isort/_future\r\n copying isort/_future/__init__.py -> build/lib.linux-x86_64-2.7/isort/_future\r\n creating build/lib.linux-x86_64-2.7/isort/_vendored\r\n creating build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/decoder.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/encoder.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/ordered.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/tz.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/__init__.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n creating build/lib.linux-x86_64-2.7/isort/deprecated\r\n copying isort/deprecated/finders.py -> build/lib.linux-x86_64-2.7/isort/deprecated\r\n copying isort/deprecated/__init__.py -> build/lib.linux-x86_64-2.7/isort/deprecated\r\n creating build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py35.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py2.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py3.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py36.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py37.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py27.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py38.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/__init__.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/all.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py39.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n error: can't copy 'isort/stdlibs': doesn't exist or not a regular file\r\n \r\n ----------------------------------------\r\nCommand \"/usr/bin/python -u -c \"import setuptools, tokenize;__file__='/tmp/pip-build-ayvPD4/isort/setup.py';exec(compile(getattr(tokenize, 'open', open)(__file__).read().replace('\\r\\n', '\\n'), __file__, 'exec'))\" install --record /tmp/pip-ulpe3F-record/install-record.txt --single-version-externally-managed --compile\" failed with error code 1 in /tmp/pip-build-ayvPD4/isort/\r\nYou are using pip version 8.1.1, however version 20.1.1 is available.\r\nYou should consider upgrading via the 'pip install --upgrade pip' command.\r\nThe command '/bin/sh -c pip2 install configparser==4.0.2 pylint==1.9.2' returned a non-zero code: 1\r\n{noformat}\r\n","created":"2020-07-08T23:02:58.495+0000"},{"body":"I got no error on trunk. This seems to be branch specific one.","created":"2020-07-08T23:13:43.053+0000"},{"body":"The cause is that the latest version of isort dropped Python 2 support.","created":"2020-07-09T00:33:20.571+0000"},{"body":"Thanks, [~aajisaka]. I created PR.","created":"2020-07-09T02:05:47.801+0000"},{"body":"Committed this to branch-3.3, branch-3.2, brach-3.1 and branch-2.10.","created":"2020-07-09T03:47:21.398+0000"}],"conversations":[{"body":"{noformat}\r\nThe command '/bin/sh -c pip2 install configparser==4.0.2 pylint==1.9.2' returned a non-zero code: 1\r\n{noformat}\r\n","from":"reporter","subject":"Fix failure of docker image creation due to pip2 install error"},{"body":"{noformat}\r\nStep 26/37 : RUN pip2 install configparser==4.0.2 pylint==1.9.2\r\n ---> Running in ee3598ed6a5a\r\nCollecting configparser==4.0.2\r\n Downloading https://files.pythonhosted.org/packages/7a/2a/95ed0501cf5d8709490b1d3a3f9b5cf340da6c433f896bbe9ce08dbe6785/configparser-4.0.2-py2.py3-none-any.whl\r\nCollecting pylint==1.9.2\r\n Downloading https://files.pythonhosted.org/packages/f2/95/0ca03c818ba3cd14f2dd4e95df5b7fa232424b7fc6ea1748d27f293bc007/pylint-1.9.2-py2.py3-none-any.whl (690kB)\r\nCollecting singledispatch; python_version < \"3.4\" (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/c5/10/369f50bcd4621b263927b0a1519987a04383d4a98fb10438042ad410cf88/singledispatch-3.4.0.3-py2.py3-none-any.whl\r\nCollecting isort>=4.2.5 (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/c4/4d/b6286cf463f9cfca698b15524e1198856d68080096f05ca7b3437f0af867/isort-5.0.5.tar.gz (77kB)\r\nCollecting backports.functools-lru-cache; python_version == \"2.7\" (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/da/d1/080d2bb13773803648281a49e3918f65b31b7beebf009887a529357fd44a/backports.functools_lru_cache-1.6.1-py2.py3-none-any.whl\r\nCollecting mccabe (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/87/89/479dc97e18549e21354893e4ee4ef36db1d237534982482c3681ee6e7b57/mccabe-0.6.1-py2.py3-none-any.whl\r\nCollecting astroid<2.0,>=1.6 (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/8b/29/0f7ec6fbf28a158886b7de49aee3a77a8a47a7e24c60e9fd6ec98ee2ec02/astroid-1.6.6-py2.py3-none-any.whl (305kB)\r\nCollecting six (from pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/ee/ff/48bde5c0f013094d729fe4b0316ba2a24774b3ff1c52d924a8a4cb04078a/six-1.15.0-py2.py3-none-any.whl\r\nCollecting enum34>=1.1.3; python_version < \"3.4\" (from astroid<2.0,>=1.6->pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/6f/2c/a9386903ece2ea85e9807e0e062174dc26fdce8b05f216d00491be29fad5/enum34-1.1.10-py2-none-any.whl\r\nCollecting wrapt (from astroid<2.0,>=1.6->pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/82/f7/e43cefbe88c5fd371f4cf0cf5eb3feccd07515af9fd6cf7dbf1d1793a797/wrapt-1.12.1.tar.gz\r\nCollecting lazy-object-proxy (from astroid<2.0,>=1.6->pylint==1.9.2)\r\n Downloading https://files.pythonhosted.org/packages/f3/1f/3e31313f557e0b97bd8f9716f502fa85c5fef181f582f816a4796b8f9ee1/lazy_object_proxy-1.5.0-cp27-cp27mu-manylinux1_x86_64.whl (55kB)\r\nBuilding wheels for collected packages: isort, wrapt\r\n Running setup.py bdist_wheel for isort: started\r\n Running setup.py bdist_wheel for isort: finished with status 'error'\r\n Complete output from command /usr/bin/python -u -c \"import setuptools, tokenize;__file__='/tmp/pip-build-ayvPD4/isort/setup.py';exec(compile(getattr(tokenize, 'open', open)(__file__).read().replace('\\r\\n', '\\n'), __file__, 'exec'))\" bdist_wheel -d /tmp/tmpzmXICSpip-wheel- --python-tag cp27:\r\n /usr/lib/python2.7/distutils/dist.py:267: UserWarning: Unknown distribution option: 'python_requires'\r\n warnings.warn(msg)\r\n running bdist_wheel\r\n running build\r\n running build_py\r\n creating build\r\n creating build/lib.linux-x86_64-2.7\r\n creating build/lib.linux-x86_64-2.7/isort\r\n copying isort/sorting.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/setuptools_commands.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/logo.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/wrap_modes.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/__main__.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/pylama_isort.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/hooks.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/main.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/profiles.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/exceptions.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/_version.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/io.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/format.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/output.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/api.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/utils.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/__init__.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/parse.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/comments.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/wrap.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/place.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/settings.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/sections.py -> build/lib.linux-x86_64-2.7/isort\r\n creating build/lib.linux-x86_64-2.7/isort/_future\r\n copying isort/_future/_dataclasses.py -> build/lib.linux-x86_64-2.7/isort/_future\r\n copying isort/_future/__init__.py -> build/lib.linux-x86_64-2.7/isort/_future\r\n creating build/lib.linux-x86_64-2.7/isort/_vendored\r\n creating build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/decoder.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/encoder.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/ordered.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/tz.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/__init__.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n creating build/lib.linux-x86_64-2.7/isort/deprecated\r\n copying isort/deprecated/finders.py -> build/lib.linux-x86_64-2.7/isort/deprecated\r\n copying isort/deprecated/__init__.py -> build/lib.linux-x86_64-2.7/isort/deprecated\r\n creating build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py35.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py2.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py3.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py36.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py37.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py27.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py38.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/__init__.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/all.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py39.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n error: can't copy 'isort/stdlibs': doesn't exist or not a regular file\r\n \r\n ----------------------------------------\r\n Failed building wheel for isort\r\n Running setup.py clean for isort\r\n Running setup.py bdist_wheel for wrapt: started\r\n Running setup.py bdist_wheel for wrapt: finished with status 'done'\r\n Stored in directory: /root/.cache/pip/wheels/b1/c2/ed/d62208260edbd3fa7156545c00ef966f45f2063d0a84f8208a\r\nSuccessfully built wrapt\r\nFailed to build isort\r\nInstalling collected packages: configparser, six, singledispatch, isort, backports.functools-lru-cache, mccabe, enum34, wrapt, lazy-object-proxy, astroid, pylint\r\n Running setup.py install for isort: started\r\n Running setup.py install for isort: finished with status 'error'\r\n Complete output from command /usr/bin/python -u -c \"import setuptools, tokenize;__file__='/tmp/pip-build-ayvPD4/isort/setup.py';exec(compile(getattr(tokenize, 'open', open)(__file__).read().replace('\\r\\n', '\\n'), __file__, 'exec'))\" install --record /tmp/pip-ulpe3F-record/install-record.txt --single-version-externally-managed --compile:\r\n /usr/lib/python2.7/distutils/dist.py:267: UserWarning: Unknown distribution option: 'python_requires'\r\n warnings.warn(msg)\r\n running install\r\n running build\r\n running build_py\r\n creating build\r\n creating build/lib.linux-x86_64-2.7\r\n creating build/lib.linux-x86_64-2.7/isort\r\n copying isort/sorting.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/setuptools_commands.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/logo.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/wrap_modes.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/__main__.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/pylama_isort.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/hooks.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/main.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/profiles.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/exceptions.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/_version.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/io.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/format.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/output.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/api.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/utils.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/__init__.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/parse.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/comments.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/wrap.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/place.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/settings.py -> build/lib.linux-x86_64-2.7/isort\r\n copying isort/sections.py -> build/lib.linux-x86_64-2.7/isort\r\n creating build/lib.linux-x86_64-2.7/isort/_future\r\n copying isort/_future/_dataclasses.py -> build/lib.linux-x86_64-2.7/isort/_future\r\n copying isort/_future/__init__.py -> build/lib.linux-x86_64-2.7/isort/_future\r\n creating build/lib.linux-x86_64-2.7/isort/_vendored\r\n creating build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/decoder.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/encoder.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/ordered.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/tz.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n copying isort/_vendored/toml/__init__.py -> build/lib.linux-x86_64-2.7/isort/_vendored/toml\r\n creating build/lib.linux-x86_64-2.7/isort/deprecated\r\n copying isort/deprecated/finders.py -> build/lib.linux-x86_64-2.7/isort/deprecated\r\n copying isort/deprecated/__init__.py -> build/lib.linux-x86_64-2.7/isort/deprecated\r\n creating build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py35.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py2.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py3.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py36.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py37.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py27.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py38.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/__init__.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/all.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n copying isort/stdlibs/py39.py -> build/lib.linux-x86_64-2.7/isort/stdlibs\r\n error: can't copy 'isort/stdlibs': doesn't exist or not a regular file\r\n \r\n ----------------------------------------\r\nCommand \"/usr/bin/python -u -c \"import setuptools, tokenize;__file__='/tmp/pip-build-ayvPD4/isort/setup.py';exec(compile(getattr(tokenize, 'open', open)(__file__).read().replace('\\r\\n', '\\n'), __file__, 'exec'))\" install --record /tmp/pip-ulpe3F-record/install-record.txt --single-version-externally-managed --compile\" failed with error code 1 in /tmp/pip-build-ayvPD4/isort/\r\nYou are using pip version 8.1.1, however version 20.1.1 is available.\r\nYou should consider upgrading via the 'pip install --upgrade pip' command.\r\nThe command '/bin/sh -c pip2 install configparser==4.0.2 pylint==1.9.2' returned a non-zero code: 1\r\n{noformat}\r\n","from":"developer"},{"body":"I got no error on trunk. This seems to be branch specific one.","from":"developer"},{"body":"The cause is that the latest version of isort dropped Python 2 support.","from":"developer"},{"body":"Thanks, [~aajisaka]. I created PR.","from":"developer"},{"body":"Committed this to branch-3.3, branch-3.2, brach-3.1 and branch-2.10.","from":"developer"}],"created":"2020-07-08T22:59:50.000+0000","description":"{noformat}\r\nThe command '/bin/sh -c pip2 install configparser==4.0.2 pylint==1.9.2' returned a non-zero code: 1\r\n{noformat}\r\n","issue_id":"13315751","key":"HADOOP-17120","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2020-07-09T03:47:21.000+0000","role":"gold_target","summary":"Fix failure of docker image creation due to pip2 install error"} {"case_id":"12666292","cluster":"JIRA-HADOOP-00eb2e5f269e","comments":[{"body":"This patch has solved issue for me. Please review.","created":"2013-08-30T10:37:42.169+0000"},{"body":"Any more update required on this issue..?\n\nI guess nobody else building hadoop in windows 32 bit.. ;)","created":"2013-11-12T14:25:36.202+0000"},{"body":"this patch works to me.\nI read there has Platform environment variable https://svn.apache.org/repos/asf/hadoop/common/trunk/BUILDING.txt\nis this patch works in both 32 and 64 bit with Platform environment variable specified?\n\nset Platform=x64 (when building on a 64-bit system)\nset Platform=Win32 (when building on a 32-bit system) ","created":"2014-02-13T09:22:58.988+0000"},{"body":"I hope it should work as both configurations are different based on env variable.\n\nAs building in 64-bit doesn't need this patch, it was not tested with the patch in 64bit machine with Platform=x64","created":"2014-02-13T09:27:31.572+0000"},{"body":"Hi [~vinayrpet]. Sorry to let your patch linger so long here. :-) Let's try to get it in now.\n\nThe change here looks good to me, but the patch needs to be rebased. We also now have LibHDFS building on Windows. That one uses CMake, but the build generator is hard-coded to \"Visual Studio 10 Win64\" in hadoop-hdfs-project/hadoop-hdfs/pom.xml. I think we'll need to look for a way to parameterize that for a 32-bit build. Let me know your thoughts.\n\nThanks!","created":"2015-02-15T20:42:51.834+0000"},{"body":"I have been compiling and using Hadoop on both Winx64 and Win32 from past few months.\nFollowing are some inputs from my side.\n\n1. Win 7 SDK is no longer supported from Microsoft. After installing some updates ( NET 4.5.1 SDK or Windows Update kb2455033) , it stopped working and got following error during compilation {code} fatal error C1083: Cannot open include file: 'ammintrin.h': No such file or directory' {code}\n\nMicrosoft has no solution for this. See the first comment in this link https://connect.microsoft.com/VisualStudio/feedback/details/660584/windows-update-kb2455033-breaks-build-with-missing-ammintrin-h They recommend using VC++ 2010 SP1 or later compilers like 2012 Express or 2013 Express. There is no fix for Win7 SDK, which many people may be using.\n\nSo to compile x64 build, I used VS 2013 Express and Win8.1 SDK. This needs upgrading .sln and .vcxproj files.\nI have documented my compilation steps here http://zutai.blogspot.com/2014/06/build-install-and-run-hadoop-24-240-on.html?showComment=1422091525887#c2264594416650430988\n\n2. For Win32 build using VS2013, I ended up modifying some C++ code to overcome these errors in HDFS\n{code}\n[exec] thread_local_storage.obj : error LNK2001: unresolved external symbol _tls_used [D:\\h\\hadoop-2.6.0-src\\hadoop-hdfs-project\\hadoop-hdfs\\target\\native\\hdfs.vcxproj]\n[exec] thread_local_storage.obj : error LNK2001: unresolved external symbol pTlsCallback [D:\\h\\hadoop-2.6.0-src\\hadoop-hdfs-project\\hadoop-hdfs\\target\\native\\hdfs.vcxproj]\n\nError\t1\terror C2440: 'function' : cannot convert from 'DWORD (__cdecl *)(LPVOID)' to 'LPTHREAD_START_ROUTINE'\tD:\\hadoop-git\\hadoop-hdfs-project\\hadoop-hdfs\\src\\main\\native\\libhdfs\\os\\windows\\thread.c \n{code}\n\nIssue HDFS-7774 has been logged by others also for similar error. I had attached a temp workaround patch there.\n\n3. I also see issue HADOOP-11080 to convert hadoop-common to use cmake\n\nOverall I think, following changes may be needed:\n- Parametrize Win32/X64 and VC++ version also ( I think CMake can autodetect)\n- Use CMake for both hadoop-common and hdfs. Let it generate proper build files based on VC+ version and x64/Win32\n- Files that are modified in HDFS-7774 needs to be properly changed to compile on both x64/Win32\n\n\n\n\n\n\n","created":"2015-02-16T07:29:12.350+0000"},{"body":"[~kiranmr], thank you for your detailed investigation into this. I'd like to recommend proceeding in small steps according to the following plan:\n# Keep the scope of HADOOP-9922 limited to enabling a 32-bit build quickly with the existing build infrastructure. Hopefully that will be a small change like Vinay's patch that can get into the 2.7.0 release quickly. It sounds like HDFS-7774 might be a pre-requisite too, but that looks like an easy change for us to get in.\n# Proceed with the CMake conversion tracked in HADOOP-11080. The scope of this would be to retain existing supported build targets while doing the conversion, not to introduce new functionality.\n# After completion of HADOOP-11080, parameterize the CMake build further to support the additional build targets (i.e. different Visual Studio versions). If it turns out this step is trivial, then it could be consolidated with HADOOP-11080.\n\nLet me know your thoughts. Thanks again!","created":"2015-02-16T16:33:56.099+0000"},{"body":"[~cnauroth], agree with your plan. As listed in first 2 points, getting successful 32-bit build can be the first target. ","created":"2015-02-16T19:45:41.532+0000"},{"body":"Uploading the rebased ( and git generated) patch, though I didnt face any problem in using old patch(using patch -p0).\n\nI no longer use 32 bit windows 7 machine. So could not test the changes again. I have tried to verify in 64bit Win8 machine by setting 'Platform=Win32' it didnt work.\n\nPlease someone with 32bit machine can verify?","created":"2015-02-18T05:59:14.354+0000"},{"body":"bq. We also now have LibHDFS building on Windows. That one uses CMake, but the build generator is hard-coded to \"Visual Studio 10 Win64\" in hadoop-hdfs-project/hadoop-hdfs/pom.xml. I think we'll need to look for a way to parameterize that for a 32-bit build. Let me know your thoughts.\nThanks chris for the point. But I am afraid, I no longer use 32bit Windows machine. Currently I am using 64bit machine. So It would be difficult for me to verify. I didnot made any changes related to this in rebased patch.\n","created":"2015-02-18T06:12:03.713+0000"},{"body":"[~vinayrpet], can I take over this issue? Along with native.vcxproj, 32 bit build also needs to be added for winutils.vcxproj and libwinutils.vcxproj. I am preparing a patch to include all these. I can also verify on both 32-bit and 64-bit windows.","created":"2015-02-18T09:29:57.896+0000"},{"body":"bq. Vinayakumar B, can I take over this issue?\nSure. Assigning to you. Thanks [~kiranmr]","created":"2015-02-19T05:01:59.827+0000"},{"body":"A new patch is attached. It adds 32-bit builds for native.vcxproj, winutils.vcxproj and libwinutils.vcxproj\n\nBoth 32-bit and 64-bit builds have been verified using following:\n1. 32bit: \n- Windows 7 using Win7.1 SDK \n- Windows 8.1 using VS 2013 Express \n\n2. 64bit\n- Windows 8.1 using VS 2013 Express","created":"2015-02-21T12:51:52.191+0000"},{"body":"[~cnauroth] check out this patch. This adds 32-bit builds only for VC++ projects in Hadoop-Common. I am planning to handle HDFS related things in HDFS-7774","created":"2015-02-23T05:24:50.638+0000"},{"body":"[~kiranmr], thank you for sharing a patch. This looks good.\n\nWhen I built for 32-bit, there were 5 additional compilation warnings:\n\n{code}\nservice.c(187): warning C4018: '<' : signed/unsigned mismatch [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\nservice.c(380): warning C4018: '<' : signed/unsigned mismatch [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\nservice.c(430): warning C4018: '<' : signed/unsigned mismatch [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\ntask.c(160): warning C4018: '<' : signed/unsigned mismatch [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\ntask.c(195): warning C4018: '<' : signed/unsigned mismatch [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\n{code}\n\nIt looks like we have some code that was trying to compare an {{int}} to a {{size_t}}, and the difference in data type size on 32-bit triggers these warnings. I suspect you can make this work on both 32-bit and 64-bit by switching the declaration of the relevant variables from {{int}} to {{size_t}}.\n\nI think this patch will be ready to go once that is addressed.","created":"2015-02-25T00:45:46.023+0000"},{"body":"Thanks for the review [~cnauroth], I have attached a new patch addressing these warnings and few more.\n\nFor some variables, i have declared as {{unsigned int}} instead of {{size_t}}, as 64-bit build was complaining assigning {{size_t}} to {{ULONG}}\n\nFollowing warnings in 32-bit build are resolved:\n{code}\nlibwinutils.c(2887): warning C4018: '<' : signed/unsigned mismatch [winutils\\libwinutils.vcxproj]\nlibwinutils.c(2899): warning C4018: '<' : signed/unsigned mismatch [winutils\\libwinutils.vcxproj]\nservice.c(187): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\nservice.c(282): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\nservice.c(380): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\nservice.c(430): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\nservice.c(1117): warning C4020: 'AddNodeManagerAndUserACEsToObject' : too many actual parameters [winutils\\winutils.vcxproj]\ntask.c(160): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\ntask.c(195): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\ntask.c(240): warning C4029: declared formal parameter list different from definition [winutils\\winutils.vcxproj]\ntask.c(339): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\n{code}\n\nFollowing warnings in 64-bit build are resolved:\n{code}\nservice.c(282): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\nservice.c(1117): warning C4020: 'AddNodeManagerAndUserACEsToObject' : too many actual parameters [winutils\\winutils.vcxproj]\ntask.c(240): warning C4029: declared formal parameter list different from definition [winutils\\winutils.vcxproj]\ntask.c(339): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\n{code}\n\nOne warning is changed in 64-bit build:\n{code}\n- task.c(312): warning C4133: 'function' : incompatible types - from 'int *' to 'size_t *' [winutils\\winutils.vcxproj]\n+ task.c(312): warning C4133: 'function' : incompatible types - from 'unsigned int *' to 'size_t *' [winutils\\winutils.vcxproj]\n{code}","created":"2015-02-25T07:08:14.241+0000"},{"body":"Thanks for the new patch, Kiran. Thank you also for the extra cleanup of some of the warnings.\n\nI still see one new warning in the 64-bit compile:\n\n{code}\ntask.c(312): warning C4133: 'function' : incompatible types - from 'unsigned int *' to 'size_t *' [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\n{code}\n\nActually, it's not so much a new warning as a different kind of warning. The existing code does have a warning on this line, but it's a different warning. I'd like to suggest that we drop the change in {{AddNodeManagerAndUserACEsToObject}} for now to avoid this.\n\nI don't want to expand the scope of this jira to a large-scale cleanup of warnings. I just wanted to confirm that: 1) the changes do not introduce new warnings for the 64-bit build and 2) the 32-bit build has no additional problematic warnings. I think once we drop the change in {{AddNodeManagerAndUserACEsToObject}}, we'll meet that goal, and I'll be +1.\n\nWe can file a separate jira for a complete cleanup. (FYI [~rusanu], these warnings appear to be related to the YARN security code.)","created":"2015-02-26T00:12:56.430+0000"},{"body":"[~cnauroth], I have reverted the changes in {{AddNodeManagerAndUserACEsToObject}}. 64-bit build is now back to original warnings in method {{AddNodeManagerAndUserACEsToObject}} \n{code}\ntask.c(312): warning C4133: 'function' : incompatible types - from 'int *' to 'size_t *'\ntask.c(339): warning C4018: '<' : signed/unsigned mismatch\n{code}","created":"2015-02-26T03:54:39.210+0000"},{"body":"+1 for the patch. I committed this to trunk and branch-2. Kiran, thank you for contributing the patch. Thanks also to Vinay for his prior work.","created":"2015-02-26T20:46:51.882+0000"},{"body":"I filed HADOOP-11639 for additional follow-up on the compilation warnings that existed before this patch.","created":"2015-02-26T21:57:07.698+0000"},{"body":"Thanks for the review and committing the patch Chris.","created":"2015-03-02T08:17:13.183+0000"},{"body":"Thanks for the review and committing the patch Chris.","created":"2015-03-02T08:17:20.009+0000"},{"body":"x86 build is failing after YARN-2190","created":"2015-03-09T11:09:17.750+0000"}],"conversations":[{"body":"Building Hadoop in windows 32 bit machine fails as native project is not having Win32 configuration","from":"reporter","subject":"hadoop windows native build will fail in 32 bit machine"},{"body":"This patch has solved issue for me. Please review.","from":"developer"},{"body":"Any more update required on this issue..?\n\nI guess nobody else building hadoop in windows 32 bit.. ;)","from":"developer"},{"body":"this patch works to me.\nI read there has Platform environment variable https://svn.apache.org/repos/asf/hadoop/common/trunk/BUILDING.txt\nis this patch works in both 32 and 64 bit with Platform environment variable specified?\n\nset Platform=x64 (when building on a 64-bit system)\nset Platform=Win32 (when building on a 32-bit system) ","from":"developer"},{"body":"I hope it should work as both configurations are different based on env variable.\n\nAs building in 64-bit doesn't need this patch, it was not tested with the patch in 64bit machine with Platform=x64","from":"developer"},{"body":"Hi [~vinayrpet]. Sorry to let your patch linger so long here. :-) Let's try to get it in now.\n\nThe change here looks good to me, but the patch needs to be rebased. We also now have LibHDFS building on Windows. That one uses CMake, but the build generator is hard-coded to \"Visual Studio 10 Win64\" in hadoop-hdfs-project/hadoop-hdfs/pom.xml. I think we'll need to look for a way to parameterize that for a 32-bit build. Let me know your thoughts.\n\nThanks!","from":"developer"},{"body":"I have been compiling and using Hadoop on both Winx64 and Win32 from past few months.\nFollowing are some inputs from my side.\n\n1. Win 7 SDK is no longer supported from Microsoft. After installing some updates ( NET 4.5.1 SDK or Windows Update kb2455033) , it stopped working and got following error during compilation {code} fatal error C1083: Cannot open include file: 'ammintrin.h': No such file or directory' {code}\n\nMicrosoft has no solution for this. See the first comment in this link https://connect.microsoft.com/VisualStudio/feedback/details/660584/windows-update-kb2455033-breaks-build-with-missing-ammintrin-h They recommend using VC++ 2010 SP1 or later compilers like 2012 Express or 2013 Express. There is no fix for Win7 SDK, which many people may be using.\n\nSo to compile x64 build, I used VS 2013 Express and Win8.1 SDK. This needs upgrading .sln and .vcxproj files.\nI have documented my compilation steps here http://zutai.blogspot.com/2014/06/build-install-and-run-hadoop-24-240-on.html?showComment=1422091525887#c2264594416650430988\n\n2. For Win32 build using VS2013, I ended up modifying some C++ code to overcome these errors in HDFS\n{code}\n[exec] thread_local_storage.obj : error LNK2001: unresolved external symbol _tls_used [D:\\h\\hadoop-2.6.0-src\\hadoop-hdfs-project\\hadoop-hdfs\\target\\native\\hdfs.vcxproj]\n[exec] thread_local_storage.obj : error LNK2001: unresolved external symbol pTlsCallback [D:\\h\\hadoop-2.6.0-src\\hadoop-hdfs-project\\hadoop-hdfs\\target\\native\\hdfs.vcxproj]\n\nError\t1\terror C2440: 'function' : cannot convert from 'DWORD (__cdecl *)(LPVOID)' to 'LPTHREAD_START_ROUTINE'\tD:\\hadoop-git\\hadoop-hdfs-project\\hadoop-hdfs\\src\\main\\native\\libhdfs\\os\\windows\\thread.c \n{code}\n\nIssue HDFS-7774 has been logged by others also for similar error. I had attached a temp workaround patch there.\n\n3. I also see issue HADOOP-11080 to convert hadoop-common to use cmake\n\nOverall I think, following changes may be needed:\n- Parametrize Win32/X64 and VC++ version also ( I think CMake can autodetect)\n- Use CMake for both hadoop-common and hdfs. Let it generate proper build files based on VC+ version and x64/Win32\n- Files that are modified in HDFS-7774 needs to be properly changed to compile on both x64/Win32\n\n\n\n\n\n\n","from":"developer"},{"body":"[~kiranmr], thank you for your detailed investigation into this. I'd like to recommend proceeding in small steps according to the following plan:\n# Keep the scope of HADOOP-9922 limited to enabling a 32-bit build quickly with the existing build infrastructure. Hopefully that will be a small change like Vinay's patch that can get into the 2.7.0 release quickly. It sounds like HDFS-7774 might be a pre-requisite too, but that looks like an easy change for us to get in.\n# Proceed with the CMake conversion tracked in HADOOP-11080. The scope of this would be to retain existing supported build targets while doing the conversion, not to introduce new functionality.\n# After completion of HADOOP-11080, parameterize the CMake build further to support the additional build targets (i.e. different Visual Studio versions). If it turns out this step is trivial, then it could be consolidated with HADOOP-11080.\n\nLet me know your thoughts. Thanks again!","from":"developer"},{"body":"[~cnauroth], agree with your plan. As listed in first 2 points, getting successful 32-bit build can be the first target. ","from":"developer"},{"body":"Uploading the rebased ( and git generated) patch, though I didnt face any problem in using old patch(using patch -p0).\n\nI no longer use 32 bit windows 7 machine. So could not test the changes again. I have tried to verify in 64bit Win8 machine by setting 'Platform=Win32' it didnt work.\n\nPlease someone with 32bit machine can verify?","from":"developer"},{"body":"bq. We also now have LibHDFS building on Windows. That one uses CMake, but the build generator is hard-coded to \"Visual Studio 10 Win64\" in hadoop-hdfs-project/hadoop-hdfs/pom.xml. I think we'll need to look for a way to parameterize that for a 32-bit build. Let me know your thoughts.\nThanks chris for the point. But I am afraid, I no longer use 32bit Windows machine. Currently I am using 64bit machine. So It would be difficult for me to verify. I didnot made any changes related to this in rebased patch.\n","from":"developer"},{"body":"[~vinayrpet], can I take over this issue? Along with native.vcxproj, 32 bit build also needs to be added for winutils.vcxproj and libwinutils.vcxproj. I am preparing a patch to include all these. I can also verify on both 32-bit and 64-bit windows.","from":"developer"},{"body":"bq. Vinayakumar B, can I take over this issue?\nSure. Assigning to you. Thanks [~kiranmr]","from":"developer"},{"body":"A new patch is attached. It adds 32-bit builds for native.vcxproj, winutils.vcxproj and libwinutils.vcxproj\n\nBoth 32-bit and 64-bit builds have been verified using following:\n1. 32bit: \n- Windows 7 using Win7.1 SDK \n- Windows 8.1 using VS 2013 Express \n\n2. 64bit\n- Windows 8.1 using VS 2013 Express","from":"developer"},{"body":"[~cnauroth] check out this patch. This adds 32-bit builds only for VC++ projects in Hadoop-Common. I am planning to handle HDFS related things in HDFS-7774","from":"developer"},{"body":"[~kiranmr], thank you for sharing a patch. This looks good.\n\nWhen I built for 32-bit, there were 5 additional compilation warnings:\n\n{code}\nservice.c(187): warning C4018: '<' : signed/unsigned mismatch [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\nservice.c(380): warning C4018: '<' : signed/unsigned mismatch [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\nservice.c(430): warning C4018: '<' : signed/unsigned mismatch [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\ntask.c(160): warning C4018: '<' : signed/unsigned mismatch [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\ntask.c(195): warning C4018: '<' : signed/unsigned mismatch [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\n{code}\n\nIt looks like we have some code that was trying to compare an {{int}} to a {{size_t}}, and the difference in data type size on 32-bit triggers these warnings. I suspect you can make this work on both 32-bit and 64-bit by switching the declaration of the relevant variables from {{int}} to {{size_t}}.\n\nI think this patch will be ready to go once that is addressed.","from":"developer"},{"body":"Thanks for the review [~cnauroth], I have attached a new patch addressing these warnings and few more.\n\nFor some variables, i have declared as {{unsigned int}} instead of {{size_t}}, as 64-bit build was complaining assigning {{size_t}} to {{ULONG}}\n\nFollowing warnings in 32-bit build are resolved:\n{code}\nlibwinutils.c(2887): warning C4018: '<' : signed/unsigned mismatch [winutils\\libwinutils.vcxproj]\nlibwinutils.c(2899): warning C4018: '<' : signed/unsigned mismatch [winutils\\libwinutils.vcxproj]\nservice.c(187): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\nservice.c(282): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\nservice.c(380): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\nservice.c(430): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\nservice.c(1117): warning C4020: 'AddNodeManagerAndUserACEsToObject' : too many actual parameters [winutils\\winutils.vcxproj]\ntask.c(160): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\ntask.c(195): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\ntask.c(240): warning C4029: declared formal parameter list different from definition [winutils\\winutils.vcxproj]\ntask.c(339): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\n{code}\n\nFollowing warnings in 64-bit build are resolved:\n{code}\nservice.c(282): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\nservice.c(1117): warning C4020: 'AddNodeManagerAndUserACEsToObject' : too many actual parameters [winutils\\winutils.vcxproj]\ntask.c(240): warning C4029: declared formal parameter list different from definition [winutils\\winutils.vcxproj]\ntask.c(339): warning C4018: '<' : signed/unsigned mismatch [winutils\\winutils.vcxproj]\n{code}\n\nOne warning is changed in 64-bit build:\n{code}\n- task.c(312): warning C4133: 'function' : incompatible types - from 'int *' to 'size_t *' [winutils\\winutils.vcxproj]\n+ task.c(312): warning C4133: 'function' : incompatible types - from 'unsigned int *' to 'size_t *' [winutils\\winutils.vcxproj]\n{code}","from":"developer"},{"body":"Thanks for the new patch, Kiran. Thank you also for the extra cleanup of some of the warnings.\n\nI still see one new warning in the 64-bit compile:\n\n{code}\ntask.c(312): warning C4133: 'function' : incompatible types - from 'unsigned int *' to 'size_t *' [C:\\hdc\\hadoop-common-project\\hadoop-common\\src\\main\\winutils\\winutils.vcxproj]\n{code}\n\nActually, it's not so much a new warning as a different kind of warning. The existing code does have a warning on this line, but it's a different warning. I'd like to suggest that we drop the change in {{AddNodeManagerAndUserACEsToObject}} for now to avoid this.\n\nI don't want to expand the scope of this jira to a large-scale cleanup of warnings. I just wanted to confirm that: 1) the changes do not introduce new warnings for the 64-bit build and 2) the 32-bit build has no additional problematic warnings. I think once we drop the change in {{AddNodeManagerAndUserACEsToObject}}, we'll meet that goal, and I'll be +1.\n\nWe can file a separate jira for a complete cleanup. (FYI [~rusanu], these warnings appear to be related to the YARN security code.)","from":"developer"},{"body":"[~cnauroth], I have reverted the changes in {{AddNodeManagerAndUserACEsToObject}}. 64-bit build is now back to original warnings in method {{AddNodeManagerAndUserACEsToObject}} \n{code}\ntask.c(312): warning C4133: 'function' : incompatible types - from 'int *' to 'size_t *'\ntask.c(339): warning C4018: '<' : signed/unsigned mismatch\n{code}","from":"developer"},{"body":"+1 for the patch. I committed this to trunk and branch-2. Kiran, thank you for contributing the patch. Thanks also to Vinay for his prior work.","from":"developer"},{"body":"I filed HADOOP-11639 for additional follow-up on the compilation warnings that existed before this patch.","from":"developer"},{"body":"Thanks for the review and committing the patch Chris.","from":"developer"},{"body":"Thanks for the review and committing the patch Chris.","from":"developer"},{"body":"x86 build is failing after YARN-2190","from":"developer"}],"created":"2013-08-30T10:24:51.000+0000","description":"Building Hadoop in windows 32 bit machine fails as native project is not having Win32 configuration","issue_id":"12666292","key":"HADOOP-9922","metadata":{"source":"Apache Jira"},"project":"HADOOP","resolution":"Fixed","resolution_date":"2015-02-26T20:46:51.000+0000","role":"gold_target","summary":"hadoop windows native build will fail in 32 bit machine"} {"case_id":"12998832","cluster":"JIRA-HBASE-8cbaa6d5ed91","comments":[{"body":"upload patch for branch-1.1","created":"2016-08-22T08:47:30.394+0000"},{"body":"Have you thought of sidelining the corrupt snapshot instead of deleting the '.tmp' dir ?","created":"2016-08-22T08:55:02.136+0000"},{"body":"Sounds reasonable, we just need to delete the relates the snapshot not the entire tmp dir. ","created":"2016-08-22T09:02:20.340+0000"},{"body":"patch for master, we delete relates corrupt snapshot dir under tmp","created":"2016-08-22T09:07:36.748+0000"},{"body":"SnapshotCleaner is a bad place to delete the .tmp dir, since the cleaner is not aware of \"in-progress\" snapshots. you'll end up cancelling files from an in-progress snapshot and having the snapshot request to fail.\n\nwe have a cleanup of the .tmp dir on snapshot failure.. why that does not work?\nhttps://github.com/apache/hbase/blob/master/hbase-server/src/main/java/org/apache/hadoop/hbase/master/snapshot/TakeSnapshotHandler.java#L225","created":"2016-08-22T14:02:47.765+0000"},{"body":"I lost the relates hmaster log when create snapshot (deleted by some program). That's very bad.\n\nBut i doubt that if there is any race condition between cleaner and TakeSnapshotHandler? When TakeSnapshotHandler try to delete relates files after create snapshot failed, the SnapshotCleaner read the relates files at the same time, so delete in TakeSnapshotHandler will failed? \n\nLet me write a testcase to test it.","created":"2016-08-22T15:32:40.455+0000"},{"body":".tmp is used to write in-progress snapshots. so unless we add a flag \"snapshot in progress\" that the cleaner verifies before taking actions, the cleaner should not delete .tmp files.\n\nmy guess is that the snapshot may have timed out. TakeSnapshotHandler had cleaned .tmp but one RS (the one with 8e3179c388e10770eba7d35e30f2777f) was still doing trying to snapshot after timeout. HRegion#addRegionToSnapshot() writes the region-manifest you have in .tmp. and that will never be cleaned since TakeSnapshotHandler is already done.\n\nwe may use something like fs.createNonRecursive() when we create the manifest. in this case if a RS is behind it will not be able to create the manifest after the master deleted .tmp via TakeSnapshotHandler.\notherwise we can add an \"in-progress\" lock to the SnapshotManager. so the cleaner and snapshot manager can avoid destroying each other state.","created":"2016-08-22T15:44:17.815+0000"},{"body":"You guess sounds reasonable. \nLet me try to mock this situation in testcase, meanwhile add \"in-progress\" lock to avoid cleaner clean up in-progress snapshots is good to solve the issue. \n\nThanks [~mbertozzi] for your good suggestions. \n\nWill upload the patch.","created":"2016-08-22T15:54:14.104+0000"},{"body":"I could mock the situation with add sleep in HRegion#addRegionToSnapshot, upload patch v1 as [~mbertozzi] suggestions. \n","created":"2016-08-23T10:20:31.385+0000"},{"body":"v1 looks ok, I was hoping for an in-memory \"lock\" in SnapshotManager instead of a file on disk. but I guess it is more work to pass the SnapshotManager around. \n\n+1 we can always optimize stuff later","created":"2016-08-23T13:45:45.026+0000"},{"body":"{quote}\nI guess it is more work to pass the SnapshotManager around.\n{quote}\nYeah, we have to pass the HMaster into Cleaner, it seems not easy currently due to we use {{FileCleanerDelegate}} to construct the SnapshotCleaner, maybe we can add one param in {{FileCleanerDelegate.getDeletableFiles}}, of course, we open another issue for it, wdyt?","created":"2016-08-24T03:14:21.390+0000"},{"body":"yeah, not trivial. let's just commit this patch and think about optimizations later in another jira.","created":"2016-08-24T03:52:19.707+0000"},{"body":"push to branch-1.1+","created":"2016-08-24T06:08:14.680+0000"},{"body":"see HBASE-16490","created":"2016-08-24T06:19:44.156+0000"}],"conversations":[{"body":"We met the problem on our real production cluster, we need to cleanup some data on hbase, we notice the archive folder is much larger than others, so we delete all snapshots of all tables, but the archive folder still grows bigger and bigger. \n\nAfter check the hmaster log, we notice the exception below:\n{code}\n2016-08-22 15:34:33,089 ERROR [f04,16000,1471240833208_ChoreService_1] snapshot.SnapshotHFileCleaner: Exception while checking if files were valid, keeping them just in case.\norg.apache.hadoop.hbase.snapshot.CorruptedSnapshotException: Couldn't read snapshot info from:hdfs://f04/hbase/.hbase-snapshot/.tmp/frog_stastic_2016-08-17/.snapshotinfo\n at org.apache.hadoop.hbase.snapshot.SnapshotDescriptionUtils.readSnapshotInfo(SnapshotDescriptionUtils.java:295)\n at org.apache.hadoop.hbase.snapshot.SnapshotReferenceUtil.getHFileNames(SnapshotReferenceUtil.java:328)\n at org.apache.hadoop.hbase.master.snapshot.SnapshotHFileCleaner$1.filesUnderSnapshot(SnapshotHFileCleaner.java:85)\n at org.apache.hadoop.hbase.master.snapshot.SnapshotFileCache.getSnapshotsInProgress(SnapshotFileCache.java:303)\n at org.apache.hadoop.hbase.master.snapshot.SnapshotFileCache.getUnreferencedFiles(SnapshotFileCache.java:194)\n at org.apache.hadoop.hbase.master.snapshot.SnapshotHFileCleaner.getDeletableFiles(SnapshotHFileCleaner.java:62)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteFiles(CleanerChore.java:233)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:157)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteDirectory(CleanerChore.java:180)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:149)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteDirectory(CleanerChore.java:180)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:149)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteDirectory(CleanerChore.java:180)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:149)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteDirectory(CleanerChore.java:180)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:149)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteDirectory(CleanerChore.java:180)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:149)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.chore(CleanerChore.java:124)\n at org.apache.hadoop.hbase.ScheduledChore.run(ScheduledChore.java:185)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\nCaused by: java.io.FileNotFoundException: File does not exist: /hbase/.hbase-snapshot/.tmp/frog_stastic_2016-08-17/.snapshotinfo\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:71)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1712)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getBlockLocations(NameNodeRpcServer.java:587)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.getBlockLocations(ClientNamenodeProtocolServerSideTranslatorPB.java:365)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n{code}\n\nIt means when SnapshotHFileCleaner begin to cleanup the archive folder, it reads the the snapshot dir to check if any links to hfiles exist, but when read the file /.hbase-snapshot/.tmp/frog_stastic_2016-08-17/.snapshotinfo, corrupt exception thrown out (not sure why the file not found), and cleanup will be failed.\n\nWhen i check the /.hbase-snapshot/.tmp/frog_stastic_2016-08-17, i can see there is only one file exist /hbase/.hbase-snapshot/.tmp/frog_stastic_2016-08-17/region-manifest.8e3179c388e10770eba7d35e30f2777f, /hbase/.hbase-snapshot/.tmp/frog_stastic_2016-08-17/.snapshotinfo missed. \n\n\nI think we should catch up the exception and delete the file to ensure cleanup will go on.\n\n\n","from":"reporter","subject":"archive folder grows bigger and bigger due to corrupt snapshot under tmp dir"},{"body":"upload patch for branch-1.1","from":"developer"},{"body":"Have you thought of sidelining the corrupt snapshot instead of deleting the '.tmp' dir ?","from":"developer"},{"body":"Sounds reasonable, we just need to delete the relates the snapshot not the entire tmp dir. ","from":"developer"},{"body":"patch for master, we delete relates corrupt snapshot dir under tmp","from":"developer"},{"body":"SnapshotCleaner is a bad place to delete the .tmp dir, since the cleaner is not aware of \"in-progress\" snapshots. you'll end up cancelling files from an in-progress snapshot and having the snapshot request to fail.\n\nwe have a cleanup of the .tmp dir on snapshot failure.. why that does not work?\nhttps://github.com/apache/hbase/blob/master/hbase-server/src/main/java/org/apache/hadoop/hbase/master/snapshot/TakeSnapshotHandler.java#L225","from":"developer"},{"body":"I lost the relates hmaster log when create snapshot (deleted by some program). That's very bad.\n\nBut i doubt that if there is any race condition between cleaner and TakeSnapshotHandler? When TakeSnapshotHandler try to delete relates files after create snapshot failed, the SnapshotCleaner read the relates files at the same time, so delete in TakeSnapshotHandler will failed? \n\nLet me write a testcase to test it.","from":"developer"},{"body":".tmp is used to write in-progress snapshots. so unless we add a flag \"snapshot in progress\" that the cleaner verifies before taking actions, the cleaner should not delete .tmp files.\n\nmy guess is that the snapshot may have timed out. TakeSnapshotHandler had cleaned .tmp but one RS (the one with 8e3179c388e10770eba7d35e30f2777f) was still doing trying to snapshot after timeout. HRegion#addRegionToSnapshot() writes the region-manifest you have in .tmp. and that will never be cleaned since TakeSnapshotHandler is already done.\n\nwe may use something like fs.createNonRecursive() when we create the manifest. in this case if a RS is behind it will not be able to create the manifest after the master deleted .tmp via TakeSnapshotHandler.\notherwise we can add an \"in-progress\" lock to the SnapshotManager. so the cleaner and snapshot manager can avoid destroying each other state.","from":"developer"},{"body":"You guess sounds reasonable. \nLet me try to mock this situation in testcase, meanwhile add \"in-progress\" lock to avoid cleaner clean up in-progress snapshots is good to solve the issue. \n\nThanks [~mbertozzi] for your good suggestions. \n\nWill upload the patch.","from":"developer"},{"body":"I could mock the situation with add sleep in HRegion#addRegionToSnapshot, upload patch v1 as [~mbertozzi] suggestions. \n","from":"developer"},{"body":"v1 looks ok, I was hoping for an in-memory \"lock\" in SnapshotManager instead of a file on disk. but I guess it is more work to pass the SnapshotManager around. \n\n+1 we can always optimize stuff later","from":"developer"},{"body":"{quote}\nI guess it is more work to pass the SnapshotManager around.\n{quote}\nYeah, we have to pass the HMaster into Cleaner, it seems not easy currently due to we use {{FileCleanerDelegate}} to construct the SnapshotCleaner, maybe we can add one param in {{FileCleanerDelegate.getDeletableFiles}}, of course, we open another issue for it, wdyt?","from":"developer"},{"body":"yeah, not trivial. let's just commit this patch and think about optimizations later in another jira.","from":"developer"},{"body":"push to branch-1.1+","from":"developer"},{"body":"see HBASE-16490","from":"developer"}],"created":"2016-08-22T07:56:11.000+0000","description":"We met the problem on our real production cluster, we need to cleanup some data on hbase, we notice the archive folder is much larger than others, so we delete all snapshots of all tables, but the archive folder still grows bigger and bigger. \n\nAfter check the hmaster log, we notice the exception below:\n{code}\n2016-08-22 15:34:33,089 ERROR [f04,16000,1471240833208_ChoreService_1] snapshot.SnapshotHFileCleaner: Exception while checking if files were valid, keeping them just in case.\norg.apache.hadoop.hbase.snapshot.CorruptedSnapshotException: Couldn't read snapshot info from:hdfs://f04/hbase/.hbase-snapshot/.tmp/frog_stastic_2016-08-17/.snapshotinfo\n at org.apache.hadoop.hbase.snapshot.SnapshotDescriptionUtils.readSnapshotInfo(SnapshotDescriptionUtils.java:295)\n at org.apache.hadoop.hbase.snapshot.SnapshotReferenceUtil.getHFileNames(SnapshotReferenceUtil.java:328)\n at org.apache.hadoop.hbase.master.snapshot.SnapshotHFileCleaner$1.filesUnderSnapshot(SnapshotHFileCleaner.java:85)\n at org.apache.hadoop.hbase.master.snapshot.SnapshotFileCache.getSnapshotsInProgress(SnapshotFileCache.java:303)\n at org.apache.hadoop.hbase.master.snapshot.SnapshotFileCache.getUnreferencedFiles(SnapshotFileCache.java:194)\n at org.apache.hadoop.hbase.master.snapshot.SnapshotHFileCleaner.getDeletableFiles(SnapshotHFileCleaner.java:62)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteFiles(CleanerChore.java:233)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:157)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteDirectory(CleanerChore.java:180)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:149)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteDirectory(CleanerChore.java:180)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:149)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteDirectory(CleanerChore.java:180)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:149)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteDirectory(CleanerChore.java:180)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:149)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteDirectory(CleanerChore.java:180)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.checkAndDeleteEntries(CleanerChore.java:149)\n at org.apache.hadoop.hbase.master.cleaner.CleanerChore.chore(CleanerChore.java:124)\n at org.apache.hadoop.hbase.ScheduledChore.run(ScheduledChore.java:185)\n at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511)\n at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180)\n at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\nCaused by: java.io.FileNotFoundException: File does not exist: /hbase/.hbase-snapshot/.tmp/frog_stastic_2016-08-17/.snapshotinfo\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:71)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1712)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getBlockLocations(NameNodeRpcServer.java:587)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.getBlockLocations(ClientNamenodeProtocolServerSideTranslatorPB.java:365)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n{code}\n\nIt means when SnapshotHFileCleaner begin to cleanup the archive folder, it reads the the snapshot dir to check if any links to hfiles exist, but when read the file /.hbase-snapshot/.tmp/frog_stastic_2016-08-17/.snapshotinfo, corrupt exception thrown out (not sure why the file not found), and cleanup will be failed.\n\nWhen i check the /.hbase-snapshot/.tmp/frog_stastic_2016-08-17, i can see there is only one file exist /hbase/.hbase-snapshot/.tmp/frog_stastic_2016-08-17/region-manifest.8e3179c388e10770eba7d35e30f2777f, /hbase/.hbase-snapshot/.tmp/frog_stastic_2016-08-17/.snapshotinfo missed. \n\n\nI think we should catch up the exception and delete the file to ensure cleanup will go on.\n\n\n","issue_id":"12998832","key":"HBASE-16464","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2016-08-24T06:08:35.000+0000","role":"gold_target","summary":"archive folder grows bigger and bigger due to corrupt snapshot under tmp dir"} {"case_id":"13179687","cluster":"JIRA-HBASE-97916ddf28a1","comments":[{"body":"FYI [~apurtell] [~vik.karma] [~abhishek.chouhan]","created":"2018-08-17T21:51:01.567+0000"},{"body":"Might be related to HBASE-20322 CME in StoreScanner causes region server crash. That was the most recent change here. ","created":"2018-08-17T22:08:36.434+0000"},{"body":"Here's a naive patch that treats the symptom but we need to look further I think","created":"2018-08-17T22:20:31.151+0000"},{"body":"Actually that might be the right approach. \r\n\r\nStoreScanner#updateReaders is called from here\r\n{code}\r\n private void notifyChangedReadersObservers(List sfs) throws IOException {\r\n for (ChangedReadersObserver o : this.changedReaderObservers) {\r\n List memStoreScanners;\r\n this.lock.readLock().lock();\r\n try {\r\n memStoreScanners = this.memstore.getScanners(o.getReadPoint());\r\n } finally {\r\n this.lock.readLock().unlock();\r\n }\r\n---> o.updateReaders(sfs, memStoreScanners);\r\n }\r\n }\r\n{code}\r\n\r\nAnd DefaultMemStore#getScanners can return null. \r\n\r\n{code}\r\n public List getScanners(long readPt) {\r\n MemStoreScanner scanner =\r\n new MemStoreScanner(activeSection, snapshotSection, readPt, comparator);\r\n scanner.seek(CellUtil.createCell(HConstants.EMPTY_START_ROW));\r\n if (scanner.peek() == null) {\r\n scanner.close();\r\n ---> return null;\r\n }\r\n return Collections. singletonList(scanner);\r\n }\r\n{code}\r\n","created":"2018-08-17T22:23:42.775+0000"},{"body":"I didn't sync my local branch looks like DefaultMemStore#getScanners was updated recently, Thanks [~apurtell] !","created":"2018-08-17T22:30:04.882+0000"},{"body":"Poke around? Iirc, this update readers is an old font of npes... Just a suggestion.","created":"2018-08-18T00:29:45.275+0000"},{"body":"This is the unintended consequence of application of 2eaa24a1323 (HBASE-17885) after 9ced0c936f4 (HBASE-20322), the former a backport, so it was just something overlooked. The null test added by this patch seems like the right thing to do. I have the branch-1 suite running after this change, so far so good, not expecting problems. ","created":"2018-08-18T00:42:40.066+0000"},{"body":"Ok.","created":"2018-08-18T00:54:32.014+0000"},{"body":"Attaching patch for master that also adds the check for null {{memstoreScanners}}, even if NPE not seen on branch-2+ doesn't hurt to be defensive","created":"2018-08-18T01:25:31.312+0000"},{"body":"I plan to commit these tomorrow. branch-1.3 and up. Let me know if you have any concerns. ","created":"2018-08-18T01:26:14.939+0000"},{"body":"The following is the implementation of CollectionUtils.isEmpty(), which includes a null check.\r\n{code}\r\n public static boolean isEmpty(Collection coll) {\r\n return coll == null || coll.isEmpty();\r\n }\r\n{code}\r\n\r\nSo I don't think we need to add the null checks with CollectionUtils.isEmpty() in the patches.","created":"2018-08-19T16:21:06.006+0000"},{"body":"Also, I think instead of the following \"memStoreScanners != null\", we can use \"!CollectionUtils.isEmpty(memStoreScanners)\" as we don't need to call clearAndClose() when memStoreScanners is empty.\r\n{code}\r\n863\t if (memStoreScanners != null) {\r\n864\t clearAndClose(new ArrayList<>(memStoreScanners));\r\n865\t }\r\n{code}\r\n\r\n","created":"2018-08-19T16:40:32.738+0000"},{"body":"Sure we can do that. And there is another place where I test for a null reference to the collection before an isEmpty check but if that API handles a null passed in as your comment implies (I haven't looked but will) then the preceeding null check can be removed there. Will make these changes. ","created":"2018-08-19T20:18:24.313+0000"},{"body":"CollectionUtils#isEmpty javadoc says \r\n\r\n_Null-safe check if the specified collection is empty._\r\n\r\nCommitting with suggestion. ","created":"2018-08-20T22:46:26.304+0000"},{"body":"With suggested change, patch to branch-2 and up is a no-op","created":"2018-08-20T22:48:27.519+0000"}],"conversations":[{"body":"I see the following NPE in the region server log for a table that is taking heavy writes. \r\nI am not sure how the {{memStoreScanners}} variable gets set to null.\r\n\r\n{code}\r\n2018-08-17 19:59:23,682 DEBUG [MemStoreFlusher.1] regionserver.HRegionFileSystem - Committing store file ...\r\n2018-08-17 19:59:23,684 INFO [MemStoreFlusher.1] regionserver.HStore - Added hdfs://...., entries=919170, sequenceid=275114, filesize=22.6 M\r\n2018-08-17 19:59:23,689 FATAL [MemStoreFlusher.1] regionserver.HRegionServer - ABORTING region server iotperf1dchbase1a-dnds22-2-prd.eng.sfdc.net,60020,1533915690501: Replay of WAL required. Forcing server shutdown\r\norg.apache.hadoop.hbase.DroppedSnapshotException: region: ......\r\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushCacheAndCommit(HRegion.java:2581)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2258)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2220)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:2106)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.flush(HRegion.java:2031)\r\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:508)\r\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:478)\r\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.access$900(MemStoreFlusher.java:76)\r\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher$FlushHandler.run(MemStoreFlusher.java:264)\r\n at java.lang.Thread.run(Thread.java:745)\r\nCaused by: java.lang.NullPointerException\r\n at java.util.ArrayList.(ArrayList.java:177)\r\n at org.apache.hadoop.hbase.regionserver.StoreScanner.updateReaders(StoreScanner.java:827)\r\n at org.apache.hadoop.hbase.regionserver.HStore.notifyChangedReadersObservers(HStore.java:1160)\r\n at org.apache.hadoop.hbase.regionserver.HStore.updateStorefiles(HStore.java:1133)\r\n at org.apache.hadoop.hbase.regionserver.HStore.access$900(HStore.java:120)\r\n at org.apache.hadoop.hbase.regionserver.HStore$StoreFlusherImpl.commit(HStore.java:2487)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushCacheAndCommit(HRegion.java:2536)\r\n ... 9 more\r\n2018-08-17 19:59:23,692 FATAL [MemStoreFlusher.1] regionserver.HRegionServer - RegionServer abort: loaded coprocessors are: [org.apache.hadoop.hbase.security.access.AccessController, org.apache.phoenix.coprocessor.ScanRegionObserver, org.apache.phoenix.coprocessor.UngroupedAggregateRegionObserver, org.apache.phoenix.hbase.index.Indexer, org.apache.phoenix.coprocessor.GroupedAggregateRegionObserver, org.apache.hadoop.hbase.security.token.TokenProvider, org.apache.phoenix.coprocessor.ServerCachingEndpointImpl]\r\n{code}","from":"reporter","subject":"NPE in StoreScanner.updateReaders causes RS to crash "},{"body":"FYI [~apurtell] [~vik.karma] [~abhishek.chouhan]","from":"developer"},{"body":"Might be related to HBASE-20322 CME in StoreScanner causes region server crash. That was the most recent change here. ","from":"developer"},{"body":"Here's a naive patch that treats the symptom but we need to look further I think","from":"developer"},{"body":"Actually that might be the right approach. \r\n\r\nStoreScanner#updateReaders is called from here\r\n{code}\r\n private void notifyChangedReadersObservers(List sfs) throws IOException {\r\n for (ChangedReadersObserver o : this.changedReaderObservers) {\r\n List memStoreScanners;\r\n this.lock.readLock().lock();\r\n try {\r\n memStoreScanners = this.memstore.getScanners(o.getReadPoint());\r\n } finally {\r\n this.lock.readLock().unlock();\r\n }\r\n---> o.updateReaders(sfs, memStoreScanners);\r\n }\r\n }\r\n{code}\r\n\r\nAnd DefaultMemStore#getScanners can return null. \r\n\r\n{code}\r\n public List getScanners(long readPt) {\r\n MemStoreScanner scanner =\r\n new MemStoreScanner(activeSection, snapshotSection, readPt, comparator);\r\n scanner.seek(CellUtil.createCell(HConstants.EMPTY_START_ROW));\r\n if (scanner.peek() == null) {\r\n scanner.close();\r\n ---> return null;\r\n }\r\n return Collections. singletonList(scanner);\r\n }\r\n{code}\r\n","from":"developer"},{"body":"I didn't sync my local branch looks like DefaultMemStore#getScanners was updated recently, Thanks [~apurtell] !","from":"developer"},{"body":"Poke around? Iirc, this update readers is an old font of npes... Just a suggestion.","from":"developer"},{"body":"This is the unintended consequence of application of 2eaa24a1323 (HBASE-17885) after 9ced0c936f4 (HBASE-20322), the former a backport, so it was just something overlooked. The null test added by this patch seems like the right thing to do. I have the branch-1 suite running after this change, so far so good, not expecting problems. ","from":"developer"},{"body":"Ok.","from":"developer"},{"body":"Attaching patch for master that also adds the check for null {{memstoreScanners}}, even if NPE not seen on branch-2+ doesn't hurt to be defensive","from":"developer"},{"body":"I plan to commit these tomorrow. branch-1.3 and up. Let me know if you have any concerns. ","from":"developer"},{"body":"The following is the implementation of CollectionUtils.isEmpty(), which includes a null check.\r\n{code}\r\n public static boolean isEmpty(Collection coll) {\r\n return coll == null || coll.isEmpty();\r\n }\r\n{code}\r\n\r\nSo I don't think we need to add the null checks with CollectionUtils.isEmpty() in the patches.","from":"developer"},{"body":"Also, I think instead of the following \"memStoreScanners != null\", we can use \"!CollectionUtils.isEmpty(memStoreScanners)\" as we don't need to call clearAndClose() when memStoreScanners is empty.\r\n{code}\r\n863\t if (memStoreScanners != null) {\r\n864\t clearAndClose(new ArrayList<>(memStoreScanners));\r\n865\t }\r\n{code}\r\n\r\n","from":"developer"},{"body":"Sure we can do that. And there is another place where I test for a null reference to the collection before an isEmpty check but if that API handles a null passed in as your comment implies (I haven't looked but will) then the preceeding null check can be removed there. Will make these changes. ","from":"developer"},{"body":"CollectionUtils#isEmpty javadoc says \r\n\r\n_Null-safe check if the specified collection is empty._\r\n\r\nCommitting with suggestion. ","from":"developer"},{"body":"With suggested change, patch to branch-2 and up is a no-op","from":"developer"}],"created":"2018-08-17T21:50:22.000+0000","description":"I see the following NPE in the region server log for a table that is taking heavy writes. \r\nI am not sure how the {{memStoreScanners}} variable gets set to null.\r\n\r\n{code}\r\n2018-08-17 19:59:23,682 DEBUG [MemStoreFlusher.1] regionserver.HRegionFileSystem - Committing store file ...\r\n2018-08-17 19:59:23,684 INFO [MemStoreFlusher.1] regionserver.HStore - Added hdfs://...., entries=919170, sequenceid=275114, filesize=22.6 M\r\n2018-08-17 19:59:23,689 FATAL [MemStoreFlusher.1] regionserver.HRegionServer - ABORTING region server iotperf1dchbase1a-dnds22-2-prd.eng.sfdc.net,60020,1533915690501: Replay of WAL required. Forcing server shutdown\r\norg.apache.hadoop.hbase.DroppedSnapshotException: region: ......\r\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushCacheAndCommit(HRegion.java:2581)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2258)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushcache(HRegion.java:2220)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.flushcache(HRegion.java:2106)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.flush(HRegion.java:2031)\r\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:508)\r\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.flushRegion(MemStoreFlusher.java:478)\r\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher.access$900(MemStoreFlusher.java:76)\r\n at org.apache.hadoop.hbase.regionserver.MemStoreFlusher$FlushHandler.run(MemStoreFlusher.java:264)\r\n at java.lang.Thread.run(Thread.java:745)\r\nCaused by: java.lang.NullPointerException\r\n at java.util.ArrayList.(ArrayList.java:177)\r\n at org.apache.hadoop.hbase.regionserver.StoreScanner.updateReaders(StoreScanner.java:827)\r\n at org.apache.hadoop.hbase.regionserver.HStore.notifyChangedReadersObservers(HStore.java:1160)\r\n at org.apache.hadoop.hbase.regionserver.HStore.updateStorefiles(HStore.java:1133)\r\n at org.apache.hadoop.hbase.regionserver.HStore.access$900(HStore.java:120)\r\n at org.apache.hadoop.hbase.regionserver.HStore$StoreFlusherImpl.commit(HStore.java:2487)\r\n at org.apache.hadoop.hbase.regionserver.HRegion.internalFlushCacheAndCommit(HRegion.java:2536)\r\n ... 9 more\r\n2018-08-17 19:59:23,692 FATAL [MemStoreFlusher.1] regionserver.HRegionServer - RegionServer abort: loaded coprocessors are: [org.apache.hadoop.hbase.security.access.AccessController, org.apache.phoenix.coprocessor.ScanRegionObserver, org.apache.phoenix.coprocessor.UngroupedAggregateRegionObserver, org.apache.phoenix.hbase.index.Indexer, org.apache.phoenix.coprocessor.GroupedAggregateRegionObserver, org.apache.hadoop.hbase.security.token.TokenProvider, org.apache.phoenix.coprocessor.ServerCachingEndpointImpl]\r\n{code}","issue_id":"13179687","key":"HBASE-21069","metadata":{"source":"Apache Jira"},"project":"HBASE","resolution":"Fixed","resolution_date":"2018-08-21T16:59:32.000+0000","role":"gold_target","summary":"NPE in StoreScanner.updateReaders causes RS to crash "} {"case_id":"12904858","cluster":"JIRA-SPARK-0490deb2abb9","comments":[{"body":"I tested this case. The problem was, Parquet filters are pushed down regardless of each schema of the splits (or rather files).\n\nWould the predicate pushdown need to be prevented when using mergeSchema option?","created":"2015-10-14T14:08:39.098+0000"},{"body":"Strangely enough, this works perfectly:\n{noformat}\nselect \n col1\nfrom\n `table3`\nwhere\n (CASE WHEN col2 = 2 THEN true ELSE false END) = true;\n{noformat}\n\nAnd returns only the row that contains col2 = 2","created":"2015-10-14T17:23:54.988+0000"},{"body":"In this case, this should be fine because filter is not pushed down to Parquet and data is filtered by Spark filter.\n\nIf you set off spark.sql.parquet.filterPushdown which is true by default, the original case also should work okay. ","created":"2015-10-14T22:58:43.470+0000"},{"body":"Setting the property {{spark.sql.parquet.filterPushdown}} to {{false}} fixed the issue.\n\nKnowing all this, does this indicate a bug in the filter2 implementation of the Parquet library? Maybe this issue should be moved to the Parquet project for someone to look at...\n\n","created":"2015-10-15T12:52:56.576+0000"},{"body":"For me, I think Spark should appropriately set filters for each file, which I think is pretty tricky, or simply prevent filtering for this case. Would anybody give us some feedback please?\n","created":"2015-10-15T13:19:32.141+0000"},{"body":"[~lian cheng] This looks clearly an issue and I made three version of patches. However, I want to be sure of which would be proper before making a PR.\n\n1. Set {{false}} to {{spark.sql.parquet.filterPushdown}} when using {{mergeSchema}}\n2. If {{spark.sql.parquet.filterPushdown}} is {{true}}, retrieve all the schema of every part-files (and also merged one) and check if each can accept the given schema and then, apply the filter only when they all can accept, which I think it's a bit over-implemented.\n3. If {{spark.sql.parquet.filterPushdown}} is {{true}}, retrieve all the schema of every part-files (and also merged one) and apply the filter to each split (rather split) that can accept the filter which (I think it's hacky) ends up different configurations for each task in a job.\n\nWould you please give me some feedbacks?","created":"2015-10-20T01:02:28.769+0000"},{"body":"Quoted from my reply on the user list:\n\nFor 1: This one is pretty simple and safe, I'd like to have this for 1.5.2, or 1.5.3 if we can't make it for 1.5.2.\n\nFor 2: I'd like to have this for Spark master. Actually we only need to calculate the intersection of all file schemata. We can make ParquetRelation.mergeSchemaInParallel return two StructTypes, the first one is the original merged schema, the other is the intersection of all file schemata, which only contains fields that exist in all file schemata. Then we decide which filter to pushed down according to the second StructType.\n\nFor 3: The idea with which I came up at first was similar to this one. Instead of pulling all file schemata to driver side, we can push filter push-down code to executor side. Namely, passing candidate filters to executor side, and compute the Parquet filter predicates according to individual file schema. I haven't looked into this direction in depth, but we can probably put this part into CatalystReadSupport, which is now initialized on executor side. However, correctness of this approach can only be guaranteed by the defensive filtering we do in Spark SQL (i.e. apply all the filters no matter they are pushed down or not), but we are considering to remove it because it imposes unnecessary performance cost. This makes me hesitant to go along this way.\n\nFrom my side, I think this is a bug of Parquet. Parquet was designed to support schema evolution. When scanning a Parquet file, if a column exists in the requested schema but is missing in the file schema, that column is filled with null. This should also hold for pushed-down filter predicates. For example, if filter \"a = 1\" is pushed down but column \"a\" doesn't exist in the Parquet file being scanned, it's safe to assume \"a\" is null in all records and drop all of them. On the contrary, if \"a IS NULL\" is pushed down, all records should be preserved.\n\nFiled PARQUET-389 to track this issue.","created":"2015-10-28T08:30:08.859+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9327","created":"2015-10-28T08:45:03.546+0000"},{"body":"[~rxin] I think this should be a blocker for 1.5.2 because we turned on Parquet filter push-down by default in 1.5, and makes this issue a regression compared to 1.4.","created":"2015-10-29T07:07:20.348+0000"},{"body":"Issue resolved by pull request 9327\n[https://github.com/apache/spark/pull/9327]","created":"2015-10-30T10:23:43.945+0000"}],"conversations":[{"body":"When evolving a schema in parquet files, spark properly expose all columns found in the different parquet files but when trying to query the data, it is not possible to apply a filter on a column that is not present in all files.\n\nTo reproduce:\n*SQL:*\n{noformat}\ncreate table `table1` STORED AS PARQUET LOCATION 'hdfs://:/path/to/table/id=1/' as select 1 as `col1`;\ncreate table `table2` STORED AS PARQUET LOCATION 'hdfs://:/path/to/table/id=2/' as select 1 as `col1`, 2 as `col2`;\ncreate table `table3` USING org.apache.spark.sql.parquet OPTIONS (path \"hdfs://:/path/to/table\");\nselect col1 from `table3` where col2 = 2;\n{noformat}\n\nThe last select will output the following Stack Trace:\n{noformat}\nAn error occurred when executing the SQL command:\nselect col1 from `table3` where col2 = 2\n\n[Simba][HiveJDBCDriver](500051) ERROR processing query/statement. Error Code: 0, SQL state: TStatus(statusCode:ERROR_STATUS, infoMessages:[*org.apache.hive.service.cli.HiveSQLException:org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 7212.0 failed 4 times, most recent failure: Lost task 0.3 in stage 7212.0 (TID 138449, 208.92.52.88): java.lang.IllegalArgumentException: Column [col2] was not found in schema!\n\tat org.apache.parquet.Preconditions.checkArgument(Preconditions.java:55)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.getColumnDescriptor(SchemaCompatibilityValidator.java:190)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumn(SchemaCompatibilityValidator.java:178)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumnFilterPredicate(SchemaCompatibilityValidator.java:160)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:94)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:59)\n\tat org.apache.parquet.filter2.predicate.Operators$Eq.accept(Operators.java:180)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validate(SchemaCompatibilityValidator.java:64)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:59)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:40)\n\tat org.apache.parquet.filter2.compat.FilterCompat$FilterPredicateCompat.accept(FilterCompat.java:126)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.filterRowGroups(RowGroupFilter.java:46)\n\tat org.apache.parquet.hadoop.ParquetRecordReader.initializeInternalReader(ParquetRecordReader.java:160)\n\tat org.apache.parquet.hadoop.ParquetRecordReader.initialize(ParquetRecordReader.java:140)\n\tat org.apache.spark.rdd.SqlNewHadoopRDD$$anon$1.(SqlNewHadoopRDD.scala:155)\n\tat org.apache.spark.rdd.SqlNewHadoopRDD.compute(SqlNewHadoopRDD.scala:120)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.UnionRDD.compute(UnionRDD.scala:87)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:88)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n\nDriver stacktrace::26:25, org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation:runInternal:SparkExecuteStatementOperation.scala:259, org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation:run:SparkExecuteStatementOperation.scala:144, org.apache.hive.service.cli.session.HiveSessionImpl:executeStatementInternal:HiveSessionImpl.java:388, org.apache.hive.service.cli.session.HiveSessionImpl:executeStatement:HiveSessionImpl.java:369, sun.reflect.GeneratedMethodAccessor134:invoke::-1, sun.reflect.DelegatingMethodAccessorImpl:invoke:DelegatingMethodAccessorImpl.java:43, java.lang.reflect.Method:invoke:Method.java:497, org.apache.hive.service.cli.session.HiveSessionProxy:invoke:HiveSessionProxy.java:78, org.apache.hive.service.cli.session.HiveSessionProxy:access$000:HiveSessionProxy.java:36, org.apache.hive.service.cli.session.HiveSessionProxy$1:run:HiveSessionProxy.java:63, java.security.AccessController:doPrivileged:AccessController.java:-2, javax.security.auth.Subject:doAs:Subject.java:422, org.apache.hadoop.security.UserGroupInformation:doAs:UserGroupInformation.java:1628, org.apache.hive.service.cli.session.HiveSessionProxy:invoke:HiveSessionProxy.java:59, com.sun.proxy.$Proxy25:executeStatement::-1, org.apache.hive.service.cli.CLIService:executeStatement:CLIService.java:261, org.apache.hive.service.cli.thrift.ThriftCLIService:ExecuteStatement:ThriftCLIService.java:486, org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement:getResult:TCLIService.java:1313, org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement:getResult:TCLIService.java:1298, org.apache.thrift.ProcessFunction:process:ProcessFunction.java:39, org.apache.thrift.TBaseProcessor:process:TBaseProcessor.java:39, org.apache.hive.service.auth.TSetIpAddressProcessor:process:TSetIpAddressProcessor.java:56, org.apache.thrift.server.TThreadPoolServer$WorkerProcess:run:TThreadPoolServer.java:285, java.util.concurrent.ThreadPoolExecutor:runWorker:ThreadPoolExecutor.java:1142, java.util.concurrent.ThreadPoolExecutor$Worker:run:ThreadPoolExecutor.java:617, java.lang.Thread:run:Thread.java:745], errorCode:0, errorMessage:org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 7212.0 failed 4 times, most recent failure: Lost task 0.3 in stage 7212.0 (TID 138449, 208.92.52.88): java.lang.IllegalArgumentException: Column [col2] was not found in schema!\n\tat org.apache.parquet.Preconditions.checkArgument(Preconditions.java:55)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.getColumnDescriptor(SchemaCompatibilityValidator.java:190)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumn(SchemaCompatibilityValidator.java:178)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumnFilterPredicate(SchemaCompatibilityValidator.java:160)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:94)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:59)\n\tat org.apache.parquet.filter2.predicate.Operators$Eq.accept(Operators.java:180)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validate(SchemaCompatibilityValidator.java:64)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:59)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:40)\n\tat org.apache.parquet.filter2.compat.FilterCompat$FilterPredicateCompat.accept(FilterCompat.java:126)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.filterRowGroups(RowGroupFilter.java:46)\n\tat org.apache.parquet.hadoop.ParquetRecordReader.initializeInternalReader(ParquetRecordReader.java:160)\n\tat org.apache.parquet.hadoop.ParquetRecordReader.initialize(ParquetRecordReader.java:140)\n\tat org.apache.spark.rdd.SqlNewHadoopRDD$$anon$1.(SqlNewHadoopRDD.scala:155)\n\tat org.apache.spark.rdd.SqlNewHadoopRDD.compute(SqlNewHadoopRDD.scala:120)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.UnionRDD.compute(UnionRDD.scala:87)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:88)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n\nDriver stacktrace:), Query: select col1 from `table3` where col2 = 2. [SQL State=HY000, DB Errorcode=500051]\n\nExecution time: 0.44s\n\n1 statement failed.\n{noformat}","from":"reporter","subject":"Parquet filters push-down may cause exception when schema merging is turned on"},{"body":"I tested this case. The problem was, Parquet filters are pushed down regardless of each schema of the splits (or rather files).\n\nWould the predicate pushdown need to be prevented when using mergeSchema option?","from":"developer"},{"body":"Strangely enough, this works perfectly:\n{noformat}\nselect \n col1\nfrom\n `table3`\nwhere\n (CASE WHEN col2 = 2 THEN true ELSE false END) = true;\n{noformat}\n\nAnd returns only the row that contains col2 = 2","from":"developer"},{"body":"In this case, this should be fine because filter is not pushed down to Parquet and data is filtered by Spark filter.\n\nIf you set off spark.sql.parquet.filterPushdown which is true by default, the original case also should work okay. ","from":"developer"},{"body":"Setting the property {{spark.sql.parquet.filterPushdown}} to {{false}} fixed the issue.\n\nKnowing all this, does this indicate a bug in the filter2 implementation of the Parquet library? Maybe this issue should be moved to the Parquet project for someone to look at...\n\n","from":"developer"},{"body":"For me, I think Spark should appropriately set filters for each file, which I think is pretty tricky, or simply prevent filtering for this case. Would anybody give us some feedback please?\n","from":"developer"},{"body":"[~lian cheng] This looks clearly an issue and I made three version of patches. However, I want to be sure of which would be proper before making a PR.\n\n1. Set {{false}} to {{spark.sql.parquet.filterPushdown}} when using {{mergeSchema}}\n2. If {{spark.sql.parquet.filterPushdown}} is {{true}}, retrieve all the schema of every part-files (and also merged one) and check if each can accept the given schema and then, apply the filter only when they all can accept, which I think it's a bit over-implemented.\n3. If {{spark.sql.parquet.filterPushdown}} is {{true}}, retrieve all the schema of every part-files (and also merged one) and apply the filter to each split (rather split) that can accept the filter which (I think it's hacky) ends up different configurations for each task in a job.\n\nWould you please give me some feedbacks?","from":"developer"},{"body":"Quoted from my reply on the user list:\n\nFor 1: This one is pretty simple and safe, I'd like to have this for 1.5.2, or 1.5.3 if we can't make it for 1.5.2.\n\nFor 2: I'd like to have this for Spark master. Actually we only need to calculate the intersection of all file schemata. We can make ParquetRelation.mergeSchemaInParallel return two StructTypes, the first one is the original merged schema, the other is the intersection of all file schemata, which only contains fields that exist in all file schemata. Then we decide which filter to pushed down according to the second StructType.\n\nFor 3: The idea with which I came up at first was similar to this one. Instead of pulling all file schemata to driver side, we can push filter push-down code to executor side. Namely, passing candidate filters to executor side, and compute the Parquet filter predicates according to individual file schema. I haven't looked into this direction in depth, but we can probably put this part into CatalystReadSupport, which is now initialized on executor side. However, correctness of this approach can only be guaranteed by the defensive filtering we do in Spark SQL (i.e. apply all the filters no matter they are pushed down or not), but we are considering to remove it because it imposes unnecessary performance cost. This makes me hesitant to go along this way.\n\nFrom my side, I think this is a bug of Parquet. Parquet was designed to support schema evolution. When scanning a Parquet file, if a column exists in the requested schema but is missing in the file schema, that column is filled with null. This should also hold for pushed-down filter predicates. For example, if filter \"a = 1\" is pushed down but column \"a\" doesn't exist in the Parquet file being scanned, it's safe to assume \"a\" is null in all records and drop all of them. On the contrary, if \"a IS NULL\" is pushed down, all records should be preserved.\n\nFiled PARQUET-389 to track this issue.","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9327","from":"developer"},{"body":"[~rxin] I think this should be a blocker for 1.5.2 because we turned on Parquet filter push-down by default in 1.5, and makes this issue a regression compared to 1.4.","from":"developer"},{"body":"Issue resolved by pull request 9327\n[https://github.com/apache/spark/pull/9327]","from":"developer"}],"created":"2015-10-14T13:11:41.000+0000","description":"When evolving a schema in parquet files, spark properly expose all columns found in the different parquet files but when trying to query the data, it is not possible to apply a filter on a column that is not present in all files.\n\nTo reproduce:\n*SQL:*\n{noformat}\ncreate table `table1` STORED AS PARQUET LOCATION 'hdfs://:/path/to/table/id=1/' as select 1 as `col1`;\ncreate table `table2` STORED AS PARQUET LOCATION 'hdfs://:/path/to/table/id=2/' as select 1 as `col1`, 2 as `col2`;\ncreate table `table3` USING org.apache.spark.sql.parquet OPTIONS (path \"hdfs://:/path/to/table\");\nselect col1 from `table3` where col2 = 2;\n{noformat}\n\nThe last select will output the following Stack Trace:\n{noformat}\nAn error occurred when executing the SQL command:\nselect col1 from `table3` where col2 = 2\n\n[Simba][HiveJDBCDriver](500051) ERROR processing query/statement. Error Code: 0, SQL state: TStatus(statusCode:ERROR_STATUS, infoMessages:[*org.apache.hive.service.cli.HiveSQLException:org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 7212.0 failed 4 times, most recent failure: Lost task 0.3 in stage 7212.0 (TID 138449, 208.92.52.88): java.lang.IllegalArgumentException: Column [col2] was not found in schema!\n\tat org.apache.parquet.Preconditions.checkArgument(Preconditions.java:55)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.getColumnDescriptor(SchemaCompatibilityValidator.java:190)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumn(SchemaCompatibilityValidator.java:178)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumnFilterPredicate(SchemaCompatibilityValidator.java:160)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:94)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:59)\n\tat org.apache.parquet.filter2.predicate.Operators$Eq.accept(Operators.java:180)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validate(SchemaCompatibilityValidator.java:64)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:59)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:40)\n\tat org.apache.parquet.filter2.compat.FilterCompat$FilterPredicateCompat.accept(FilterCompat.java:126)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.filterRowGroups(RowGroupFilter.java:46)\n\tat org.apache.parquet.hadoop.ParquetRecordReader.initializeInternalReader(ParquetRecordReader.java:160)\n\tat org.apache.parquet.hadoop.ParquetRecordReader.initialize(ParquetRecordReader.java:140)\n\tat org.apache.spark.rdd.SqlNewHadoopRDD$$anon$1.(SqlNewHadoopRDD.scala:155)\n\tat org.apache.spark.rdd.SqlNewHadoopRDD.compute(SqlNewHadoopRDD.scala:120)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.UnionRDD.compute(UnionRDD.scala:87)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:88)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n\nDriver stacktrace::26:25, org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation:runInternal:SparkExecuteStatementOperation.scala:259, org.apache.spark.sql.hive.thriftserver.SparkExecuteStatementOperation:run:SparkExecuteStatementOperation.scala:144, org.apache.hive.service.cli.session.HiveSessionImpl:executeStatementInternal:HiveSessionImpl.java:388, org.apache.hive.service.cli.session.HiveSessionImpl:executeStatement:HiveSessionImpl.java:369, sun.reflect.GeneratedMethodAccessor134:invoke::-1, sun.reflect.DelegatingMethodAccessorImpl:invoke:DelegatingMethodAccessorImpl.java:43, java.lang.reflect.Method:invoke:Method.java:497, org.apache.hive.service.cli.session.HiveSessionProxy:invoke:HiveSessionProxy.java:78, org.apache.hive.service.cli.session.HiveSessionProxy:access$000:HiveSessionProxy.java:36, org.apache.hive.service.cli.session.HiveSessionProxy$1:run:HiveSessionProxy.java:63, java.security.AccessController:doPrivileged:AccessController.java:-2, javax.security.auth.Subject:doAs:Subject.java:422, org.apache.hadoop.security.UserGroupInformation:doAs:UserGroupInformation.java:1628, org.apache.hive.service.cli.session.HiveSessionProxy:invoke:HiveSessionProxy.java:59, com.sun.proxy.$Proxy25:executeStatement::-1, org.apache.hive.service.cli.CLIService:executeStatement:CLIService.java:261, org.apache.hive.service.cli.thrift.ThriftCLIService:ExecuteStatement:ThriftCLIService.java:486, org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement:getResult:TCLIService.java:1313, org.apache.hive.service.cli.thrift.TCLIService$Processor$ExecuteStatement:getResult:TCLIService.java:1298, org.apache.thrift.ProcessFunction:process:ProcessFunction.java:39, org.apache.thrift.TBaseProcessor:process:TBaseProcessor.java:39, org.apache.hive.service.auth.TSetIpAddressProcessor:process:TSetIpAddressProcessor.java:56, org.apache.thrift.server.TThreadPoolServer$WorkerProcess:run:TThreadPoolServer.java:285, java.util.concurrent.ThreadPoolExecutor:runWorker:ThreadPoolExecutor.java:1142, java.util.concurrent.ThreadPoolExecutor$Worker:run:ThreadPoolExecutor.java:617, java.lang.Thread:run:Thread.java:745], errorCode:0, errorMessage:org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 7212.0 failed 4 times, most recent failure: Lost task 0.3 in stage 7212.0 (TID 138449, 208.92.52.88): java.lang.IllegalArgumentException: Column [col2] was not found in schema!\n\tat org.apache.parquet.Preconditions.checkArgument(Preconditions.java:55)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.getColumnDescriptor(SchemaCompatibilityValidator.java:190)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumn(SchemaCompatibilityValidator.java:178)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validateColumnFilterPredicate(SchemaCompatibilityValidator.java:160)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:94)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.visit(SchemaCompatibilityValidator.java:59)\n\tat org.apache.parquet.filter2.predicate.Operators$Eq.accept(Operators.java:180)\n\tat org.apache.parquet.filter2.predicate.SchemaCompatibilityValidator.validate(SchemaCompatibilityValidator.java:64)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:59)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.visit(RowGroupFilter.java:40)\n\tat org.apache.parquet.filter2.compat.FilterCompat$FilterPredicateCompat.accept(FilterCompat.java:126)\n\tat org.apache.parquet.filter2.compat.RowGroupFilter.filterRowGroups(RowGroupFilter.java:46)\n\tat org.apache.parquet.hadoop.ParquetRecordReader.initializeInternalReader(ParquetRecordReader.java:160)\n\tat org.apache.parquet.hadoop.ParquetRecordReader.initialize(ParquetRecordReader.java:140)\n\tat org.apache.spark.rdd.SqlNewHadoopRDD$$anon$1.(SqlNewHadoopRDD.scala:155)\n\tat org.apache.spark.rdd.SqlNewHadoopRDD.compute(SqlNewHadoopRDD.scala:120)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.UnionRDD.compute(UnionRDD.scala:87)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:38)\n\tat org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:297)\n\tat org.apache.spark.rdd.RDD.iterator(RDD.scala:264)\n\tat org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:88)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\n\nDriver stacktrace:), Query: select col1 from `table3` where col2 = 2. [SQL State=HY000, DB Errorcode=500051]\n\nExecution time: 0.44s\n\n1 statement failed.\n{noformat}","issue_id":"12904858","key":"SPARK-11103","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-10-30T10:23:43.000+0000","role":"gold_target","summary":"Parquet filters push-down may cause exception when schema merging is turned on"} {"case_id":"12906888","cluster":"JIRA-SPARK-d5e7a51e2bf0","comments":[{"body":"I am able to recreate this issue on 1.6.0 with hive 1.2.1.. I am looking into it. ","created":"2015-10-22T00:55:21.696+0000"},{"body":"Root cause found, fix is being tested. will submit PR shortly.","created":"2015-10-22T23:02:48.521+0000"},{"body":"User 'xwu0226' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9326","created":"2015-10-28T08:06:07.320+0000"},{"body":"Issue resolved by pull request 9326\n[https://github.com/apache/spark/pull/9326]","created":"2015-10-29T14:58:36.717+0000"},{"body":"I picked it in branch 1.5. Will update the fix version later.","created":"2015-10-29T14:59:05.249+0000"},{"body":"OK I add 1.5.3 as the fix version. If we re-cut 1.5.2, I will change the fix version.","created":"2015-10-29T15:09:05.207+0000"},{"body":"[~yhuai] Thank you!","created":"2015-10-29T15:37:51.063+0000"},{"body":"Hi [~yhuai]\n\nI am running into an issue that might be related to this. In short, when doing reads from in memory cached tables, i see that spark is scanning all the columns, as opposed to non-cached parquet tables, when only the required columns are read. For example, i have a table sales_report which is ~800MB on disk stored as parquet, with about 200 columns. When running a query on the non cached table:\n{code}\nselect count(distinct make ) from sales_report; \n{code} \nI see on the spark UI that the job reads about 5MB.\nLooking at the explain of the query, i see that it only reads the necessary column from parquet, as expected. \n{code}\n explain select count(distinct make) from sales_report;\n| == Physical Plan == \n |\n| TungstenAggregate(key=[], functions=[(count(make#7871),mode=Complete,isDistinct=true)], output=[_c0#7870L]) \n| TungstenAggregate(key=[make#7871], functions=[], output=[make#7871]) \n| TungstenExchange SinglePartition \n| TungstenAggregate(key=[make#7871], functions=[], output=[make#7871]) \n| Scan ParquetRelation[hdfs://redacted:8020/apps/hive/warehouse/redacted/sales_report][make#7871] \n+----------------------------------------------------------------------------------------------------------------+\n{code}\n\nHowever when I run \n{code}\ncache table sales_report\n{code}\n\nand do explain on the same query, i see that spark is trying to read all the columns in the table. When running the query, it reads 800MB from the memory. \n\n{code}\nexplain select count(distinct make) from sales_report; \n\n| == Physical Plan == \n| TungstenAggregate(key=[], functions=[(count(make#13337),mode=Complete,isDistinct=true)], output=[_c0#17135L]) \n| TungstenAggregate(key=[make#13337], functions=[], output=[make#13337]) \n| TungstenExchange SinglePartition \n| TungstenAggregate(key=[make#13337], functions=[], output=[make#13337]) \n| InMemoryColumnarTableScan [make#13337], (InMemoryRelation [make#13337,makeref#13338,registration#13339,chassis#13340,derivative#13341,derivativeid#13342,registrationdate#133 |\n{code}\n\nI tried disabling Tungsten, but that didn't change this behavior. \n\n{code}\n set spark.sql.tungsten.enabled=false;\n explain select count(distinct make ) from sales_report; \n| == Physical Plan == \n| Aggregate false, [CombineAndCount(partialSets#19514) AS _c0#18328L] \n| Exchange SinglePartition \n| Aggregate true, [AddToHashSet(make#13337) AS partialSets#19514] \n| InMemoryColumnarTableScan [make#13337], (InMemoryRelation [make#13337,makeref#13338,registration#13339,chassis#13340,derivative#13341,derivativeid#13342,registrationdate#1334 \n{code} \n\nI don't know if this is a bug or expected behavior, but it results in queries over cached and uncached tables having the same run time over large datasets. I am running Spark 1.5.0. \n\nThanks\nGurgen ","created":"2015-11-12T21:20:24.804+0000"},{"body":"When building the cache, we are going to read all of the columns the first time. Subsuquent scans will pass over only the columns required.\n\nNote you should avoid using caching if the table does not fit in memory as the construction of in-memory buffers is pretty expensive.","created":"2015-11-12T21:25:01.128+0000"}],"conversations":[{"body":"Since upgrading to 1.5.1, using the {{CACHE TABLE}} works great for all tables except for parquet tables, likely related to the parquet native reader.\n\nHere are steps for parquet table:\n\n{code}\ncreate table test_parquet stored as parquet as select 1;\nexplain select * from test_parquet;\n{code}\n\nWith output:\n\n{code}\n== Physical Plan ==\nScan ParquetRelation[hdfs://192.168.99.9/user/hive/warehouse/test_parquet][_c0#141]\n{code}\n\nAnd then caching:\n\n{code}\ncache table test_parquet;\nexplain select * from test_parquet;\n{code}\n\nWith output:\n\n{code}\n== Physical Plan ==\nScan ParquetRelation[hdfs://192.168.99.9/user/hive/warehouse/test_parquet][_c0#174]\n{code}\n\nNote it isn't cached. I have included spark log output for the {{cache table}} and {{explain}} statements below.\n\n---\n\nHere's the same for non-parquet table:\n\n{code}\ncache table test_no_parquet;\nexplain select * from test_no_parquet;\n{code}\n\nWith output:\n\n{code}\n== Physical Plan ==\nHiveTableScan [_c0#210], (MetastoreRelation default, test_no_parquet, None)\n{code}\n\nAnd then caching:\n\n{code}\ncache table test_no_parquet;\nexplain select * from test_no_parquet;\n{code}\n\nWith output:\n\n{code}\n== Physical Plan ==\nInMemoryColumnarTableScan [_c0#229], (InMemoryRelation [_c0#229], true, 10000, StorageLevel(true, true, false, true, 1), (HiveTableScan [_c0#211], (MetastoreRelation default, test_no_parquet, None)), Some(test_no_parquet))\n{code}\n\nNot that the table seems to be cached.\n---\n\nNote that if the flag {{spark.sql.hive.convertMetastoreParquet}} is set to {{false}}, parquet tables work the same as non-parquet tables with caching. This is a reasonable workaround for us, but ideally, we would like to benefit from the native reading.\n\n---\n\nSpark logs for {{cache table}} for {{test_parquet}}:\n\n{code}\n15/10/21 21:22:05 INFO thriftserver.SparkExecuteStatementOperation: Running query 'cache table test_parquet' with 20ee2ab9-5242-4783-81cf-46115ed72610\n15/10/21 21:22:05 INFO metastore.HiveMetaStore: 49: get_table : db=default tbl=test_parquet\n15/10/21 21:22:05 INFO HiveMetaStore.audit: ugi=vagrant\tip=unknown-ip-addr\tcmd=get_table : db=default tbl=test_parquet\n15/10/21 21:22:05 INFO metastore.HiveMetaStore: 49: Opening raw store with implemenation class:org.apache.hadoop.hive.metastore.ObjectStore\n15/10/21 21:22:05 INFO metastore.ObjectStore: ObjectStore, initialize called\n15/10/21 21:22:05 INFO DataNucleus.Query: Reading in results for query \"org.datanucleus.store.rdbms.query.SQLQuery@0\" since the connection used is closing\n15/10/21 21:22:05 INFO metastore.MetaStoreDirectSql: Using direct SQL, underlying DB is MYSQL\n15/10/21 21:22:05 INFO metastore.ObjectStore: Initialized ObjectStore\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(215680) called with curMem=4196713, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_59 stored as values in memory (estimated size 210.6 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(20265) called with curMem=4412393, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_59_piece0 stored as bytes in memory (estimated size 19.8 KB, free 128.3 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_59_piece0 in memory on 192.168.99.9:50262 (size: 19.8 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO spark.SparkContext: Created broadcast 59 from run at AccessController.java:-2\n15/10/21 21:22:05 INFO metastore.HiveMetaStore: 49: get_table : db=default tbl=test_parquet\n15/10/21 21:22:05 INFO HiveMetaStore.audit: ugi=vagrant\tip=unknown-ip-addr\tcmd=get_table : db=default tbl=test_parquet\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(215680) called with curMem=4432658, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_60 stored as values in memory (estimated size 210.6 KB, free 128.1 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Removed broadcast_58_piece0 on 192.168.99.9:50262 in memory (size: 19.8 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Removed broadcast_57_piece0 on 192.168.99.9:50262 in memory (size: 21.1 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Removed broadcast_57_piece0 on slave2:46912 in memory (size: 21.1 KB, free: 534.5 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Removed broadcast_57_piece0 on slave0:46599 in memory (size: 21.1 KB, free: 534.3 MB)\n15/10/21 21:22:05 INFO spark.ContextCleaner: Cleaned accumulator 86\n15/10/21 21:22:05 INFO spark.ContextCleaner: Cleaned accumulator 84\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(20265) called with curMem=4327620, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_60_piece0 stored as bytes in memory (estimated size 19.8 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_60_piece0 in memory on 192.168.99.9:50262 (size: 19.8 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO spark.SparkContext: Created broadcast 60 from run at AccessController.java:-2\n15/10/21 21:22:05 INFO spark.SparkContext: Starting job: run at AccessController.java:-2\n15/10/21 21:22:05 INFO parquet.ParquetRelation: Reading Parquet file(s) from hdfs://192.168.99.9/user/hive/warehouse/test_parquet/part-r-00000-7cf64eb9-76ca-47c7-92aa-eb5ba879faae.gz.parquet\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Registering RDD 171 (run at AccessController.java:-2)\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Got job 24 (run at AccessController.java:-2) with 1 output partitions\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Final stage: ResultStage 34(run at AccessController.java:-2)\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Parents of final stage: List(ShuffleMapStage 33)\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Missing parents: List(ShuffleMapStage 33)\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Submitting ShuffleMapStage 33 (MapPartitionsRDD[171] at run at AccessController.java:-2), which has no missing parents\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(9472) called with curMem=4347885, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_61 stored as values in memory (estimated size 9.3 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(4838) called with curMem=4357357, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_61_piece0 stored as bytes in memory (estimated size 4.7 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_61_piece0 in memory on 192.168.99.9:50262 (size: 4.7 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO spark.SparkContext: Created broadcast 61 from broadcast at DAGScheduler.scala:861\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Submitting 1 missing tasks from ShuffleMapStage 33 (MapPartitionsRDD[171] at run at AccessController.java:-2)\n15/10/21 21:22:05 INFO cluster.YarnScheduler: Adding task set 33.0 with 1 tasks\n15/10/21 21:22:05 INFO scheduler.FairSchedulableBuilder: Added task set TaskSet_33 tasks to pool default\n15/10/21 21:22:05 INFO scheduler.TaskSetManager: Starting task 0.0 in stage 33.0 (TID 45, slave2, NODE_LOCAL, 2234 bytes)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_61_piece0 in memory on slave2:46912 (size: 4.7 KB, free: 534.5 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_60_piece0 in memory on slave2:46912 (size: 19.8 KB, free: 534.4 MB)\n15/10/21 21:22:05 INFO scheduler.TaskSetManager: Finished task 0.0 in stage 33.0 (TID 45) in 105 ms on slave2 (1/1)\n15/10/21 21:22:05 INFO cluster.YarnScheduler: Removed TaskSet 33.0, whose tasks have all completed, from pool default\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: ShuffleMapStage 33 (run at AccessController.java:-2) finished in 0.105 s\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: looking for newly runnable stages\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: running: Set()\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: waiting: Set(ResultStage 34)\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: failed: Set()\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Missing parents for ResultStage 34: List()\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: Finished stage: org.apache.spark.scheduler.StageInfo@532f49c8\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Submitting ResultStage 34 (MapPartitionsRDD[174] at run at AccessController.java:-2), which is now runnable\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: task runtime:(count: 1, mean: 105.000000, stdev: 0.000000, max: 105.000000, min: 105.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: shuffle bytes written:(count: 1, mean: 49.000000, stdev: 0.000000, max: 49.000000, min: 49.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t49.0 B\t49.0 B\t49.0 B\t49.0 B\t49.0 B\t49.0 B\t49.0 B\t49.0 B\t49.0 B\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(10440) called with curMem=4362195, maxMem=139009720\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: task result size:(count: 1, mean: 2381.000000, stdev: 0.000000, max: 2381.000000, min: 2381.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_62 stored as values in memory (estimated size 10.2 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: executor (non-fetch) time pct: (count: 1, mean: 68.571429, stdev: 0.000000, max: 68.571429, min: 68.571429)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t69 %\t69 %\t69 %\t69 %\t69 %\t69 %\t69 %\t69 %\t69 %\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: other time pct: (count: 1, mean: 31.428571, stdev: 0.000000, max: 31.428571, min: 31.428571)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t31 %\t31 %\t31 %\t31 %\t31 %\t31 %\t31 %\t31 %\t31 %\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(5358) called with curMem=4372635, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_62_piece0 stored as bytes in memory (estimated size 5.2 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_62_piece0 in memory on 192.168.99.9:50262 (size: 5.2 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO spark.SparkContext: Created broadcast 62 from broadcast at DAGScheduler.scala:861\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Submitting 1 missing tasks from ResultStage 34 (MapPartitionsRDD[174] at run at AccessController.java:-2)\n15/10/21 21:22:05 INFO cluster.YarnScheduler: Adding task set 34.0 with 1 tasks\n15/10/21 21:22:05 INFO scheduler.FairSchedulableBuilder: Added task set TaskSet_34 tasks to pool default\n15/10/21 21:22:05 INFO scheduler.TaskSetManager: Starting task 0.0 in stage 34.0 (TID 46, slave2, PROCESS_LOCAL, 1914 bytes)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_62_piece0 in memory on slave2:46912 (size: 5.2 KB, free: 534.4 MB)\n15/10/21 21:22:05 INFO spark.MapOutputTrackerMasterEndpoint: Asked to send map output locations for shuffle 9 to slave2:43867\n15/10/21 21:22:05 INFO spark.MapOutputTrackerMaster: Size of output statuses for shuffle 9 is 135 bytes\n15/10/21 21:22:05 INFO scheduler.TaskSetManager: Finished task 0.0 in stage 34.0 (TID 46) in 48 ms on slave2 (1/1)\n15/10/21 21:22:05 INFO cluster.YarnScheduler: Removed TaskSet 34.0, whose tasks have all completed, from pool default\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: ResultStage 34 (run at AccessController.java:-2) finished in 0.047 s\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: Finished stage: org.apache.spark.scheduler.StageInfo@37a20848\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: task runtime:(count: 1, mean: 48.000000, stdev: 0.000000, max: 48.000000, min: 48.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: fetch wait time:(count: 1, mean: 0.000000, stdev: 0.000000, max: 0.000000, min: 0.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: remote bytes read:(count: 1, mean: 0.000000, stdev: 0.000000, max: 0.000000, min: 0.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0.0 B\t0.0 B\t0.0 B\t0.0 B\t0.0 B\t0.0 B\t0.0 B\t0.0 B\t0.0 B\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: task result size:(count: 1, mean: 1737.000000, stdev: 0.000000, max: 1737.000000, min: 1737.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: executor (non-fetch) time pct: (count: 1, mean: 29.166667, stdev: 0.000000, max: 29.166667, min: 29.166667)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t29 %\t29 %\t29 %\t29 %\t29 %\t29 %\t29 %\t29 %\t29 %\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: fetch wait time pct: (count: 1, mean: 0.000000, stdev: 0.000000, max: 0.000000, min: 0.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t 0 %\t 0 %\t 0 %\t 0 %\t 0 %\t 0 %\t 0 %\t 0 %\t 0 %\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: other time pct: (count: 1, mean: 70.833333, stdev: 0.000000, max: 70.833333, min: 70.833333)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t71 %\t71 %\t71 %\t71 %\t71 %\t71 %\t71 %\t71 %\t71 %\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Job 24 finished: run at AccessController.java:-2, took 0.175295 s\n{code}\n\nSpark logs for {{explain}} for {{test_parquet}}:\n\n{code}\n15/10/21 21:23:19 INFO thriftserver.SparkExecuteStatementOperation: Running query 'explain select * from test_parquet' with bae9c0bf-57f9-4c80-b745-3f0202469f3f\n15/10/21 21:23:19 INFO parse.ParseDriver: Parsing command: explain select * from test_parquet\n15/10/21 21:23:19 INFO parse.ParseDriver: Parse Completed\n15/10/21 21:23:19 INFO metastore.HiveMetaStore: 50: get_table : db=default tbl=test_parquet\n15/10/21 21:23:19 INFO HiveMetaStore.audit: ugi=vagrant\tip=unknown-ip-addr\tcmd=get_table : db=default tbl=test_parquet\n15/10/21 21:23:19 INFO metastore.HiveMetaStore: 50: Opening raw store with implemenation class:org.apache.hadoop.hive.metastore.ObjectStore\n15/10/21 21:23:19 INFO metastore.ObjectStore: ObjectStore, initialize called\n15/10/21 21:23:19 INFO DataNucleus.Query: Reading in results for query \"org.datanucleus.store.rdbms.query.SQLQuery@0\" since the connection used is closing\n15/10/21 21:23:19 INFO metastore.MetaStoreDirectSql: Using direct SQL, underlying DB is MYSQL\n15/10/21 21:23:19 INFO metastore.ObjectStore: Initialized ObjectStore\n15/10/21 21:23:19 INFO storage.MemoryStore: ensureFreeSpace(215680) called with curMem=4377993, maxMem=139009720\n15/10/21 21:23:19 INFO storage.MemoryStore: Block broadcast_63 stored as values in memory (estimated size 210.6 KB, free 128.2 MB)\n15/10/21 21:23:19 INFO storage.MemoryStore: ensureFreeSpace(20265) called with curMem=4593673, maxMem=139009720\n15/10/21 21:23:19 INFO storage.MemoryStore: Block broadcast_63_piece0 stored as bytes in memory (estimated size 19.8 KB, free 128.2 MB)\n15/10/21 21:23:19 INFO storage.BlockManagerInfo: Added broadcast_63_piece0 in memory on 192.168.99.9:50262 (size: 19.8 KB, free: 132.2 MB)\n15/10/21 21:23:19 INFO spark.SparkContext: Created broadcast 63 from run at AccessController.java:-2\n15/10/21 21:23:19 INFO thriftserver.SparkExecuteStatementOperation: Result Schema: List(plan#262)\n15/10/21 21:23:19 INFO thriftserver.SparkExecuteStatementOperation: Result Schema: List(plan#262)\n15/10/21 21:23:19 INFO thriftserver.SparkExecuteStatementOperation: Result Schema: List(plan#262)\n{code}","from":"reporter","subject":"[1.5] Table cache for Parquet broken in 1.5"},{"body":"I am able to recreate this issue on 1.6.0 with hive 1.2.1.. I am looking into it. ","from":"developer"},{"body":"Root cause found, fix is being tested. will submit PR shortly.","from":"developer"},{"body":"User 'xwu0226' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/9326","from":"developer"},{"body":"Issue resolved by pull request 9326\n[https://github.com/apache/spark/pull/9326]","from":"developer"},{"body":"I picked it in branch 1.5. Will update the fix version later.","from":"developer"},{"body":"OK I add 1.5.3 as the fix version. If we re-cut 1.5.2, I will change the fix version.","from":"developer"},{"body":"[~yhuai] Thank you!","from":"developer"},{"body":"Hi [~yhuai]\n\nI am running into an issue that might be related to this. In short, when doing reads from in memory cached tables, i see that spark is scanning all the columns, as opposed to non-cached parquet tables, when only the required columns are read. For example, i have a table sales_report which is ~800MB on disk stored as parquet, with about 200 columns. When running a query on the non cached table:\n{code}\nselect count(distinct make ) from sales_report; \n{code} \nI see on the spark UI that the job reads about 5MB.\nLooking at the explain of the query, i see that it only reads the necessary column from parquet, as expected. \n{code}\n explain select count(distinct make) from sales_report;\n| == Physical Plan == \n |\n| TungstenAggregate(key=[], functions=[(count(make#7871),mode=Complete,isDistinct=true)], output=[_c0#7870L]) \n| TungstenAggregate(key=[make#7871], functions=[], output=[make#7871]) \n| TungstenExchange SinglePartition \n| TungstenAggregate(key=[make#7871], functions=[], output=[make#7871]) \n| Scan ParquetRelation[hdfs://redacted:8020/apps/hive/warehouse/redacted/sales_report][make#7871] \n+----------------------------------------------------------------------------------------------------------------+\n{code}\n\nHowever when I run \n{code}\ncache table sales_report\n{code}\n\nand do explain on the same query, i see that spark is trying to read all the columns in the table. When running the query, it reads 800MB from the memory. \n\n{code}\nexplain select count(distinct make) from sales_report; \n\n| == Physical Plan == \n| TungstenAggregate(key=[], functions=[(count(make#13337),mode=Complete,isDistinct=true)], output=[_c0#17135L]) \n| TungstenAggregate(key=[make#13337], functions=[], output=[make#13337]) \n| TungstenExchange SinglePartition \n| TungstenAggregate(key=[make#13337], functions=[], output=[make#13337]) \n| InMemoryColumnarTableScan [make#13337], (InMemoryRelation [make#13337,makeref#13338,registration#13339,chassis#13340,derivative#13341,derivativeid#13342,registrationdate#133 |\n{code}\n\nI tried disabling Tungsten, but that didn't change this behavior. \n\n{code}\n set spark.sql.tungsten.enabled=false;\n explain select count(distinct make ) from sales_report; \n| == Physical Plan == \n| Aggregate false, [CombineAndCount(partialSets#19514) AS _c0#18328L] \n| Exchange SinglePartition \n| Aggregate true, [AddToHashSet(make#13337) AS partialSets#19514] \n| InMemoryColumnarTableScan [make#13337], (InMemoryRelation [make#13337,makeref#13338,registration#13339,chassis#13340,derivative#13341,derivativeid#13342,registrationdate#1334 \n{code} \n\nI don't know if this is a bug or expected behavior, but it results in queries over cached and uncached tables having the same run time over large datasets. I am running Spark 1.5.0. \n\nThanks\nGurgen ","from":"developer"},{"body":"When building the cache, we are going to read all of the columns the first time. Subsuquent scans will pass over only the columns required.\n\nNote you should avoid using caching if the table does not fit in memory as the construction of in-memory buffers is pretty expensive.","from":"developer"}],"created":"2015-10-21T21:23:31.000+0000","description":"Since upgrading to 1.5.1, using the {{CACHE TABLE}} works great for all tables except for parquet tables, likely related to the parquet native reader.\n\nHere are steps for parquet table:\n\n{code}\ncreate table test_parquet stored as parquet as select 1;\nexplain select * from test_parquet;\n{code}\n\nWith output:\n\n{code}\n== Physical Plan ==\nScan ParquetRelation[hdfs://192.168.99.9/user/hive/warehouse/test_parquet][_c0#141]\n{code}\n\nAnd then caching:\n\n{code}\ncache table test_parquet;\nexplain select * from test_parquet;\n{code}\n\nWith output:\n\n{code}\n== Physical Plan ==\nScan ParquetRelation[hdfs://192.168.99.9/user/hive/warehouse/test_parquet][_c0#174]\n{code}\n\nNote it isn't cached. I have included spark log output for the {{cache table}} and {{explain}} statements below.\n\n---\n\nHere's the same for non-parquet table:\n\n{code}\ncache table test_no_parquet;\nexplain select * from test_no_parquet;\n{code}\n\nWith output:\n\n{code}\n== Physical Plan ==\nHiveTableScan [_c0#210], (MetastoreRelation default, test_no_parquet, None)\n{code}\n\nAnd then caching:\n\n{code}\ncache table test_no_parquet;\nexplain select * from test_no_parquet;\n{code}\n\nWith output:\n\n{code}\n== Physical Plan ==\nInMemoryColumnarTableScan [_c0#229], (InMemoryRelation [_c0#229], true, 10000, StorageLevel(true, true, false, true, 1), (HiveTableScan [_c0#211], (MetastoreRelation default, test_no_parquet, None)), Some(test_no_parquet))\n{code}\n\nNot that the table seems to be cached.\n---\n\nNote that if the flag {{spark.sql.hive.convertMetastoreParquet}} is set to {{false}}, parquet tables work the same as non-parquet tables with caching. This is a reasonable workaround for us, but ideally, we would like to benefit from the native reading.\n\n---\n\nSpark logs for {{cache table}} for {{test_parquet}}:\n\n{code}\n15/10/21 21:22:05 INFO thriftserver.SparkExecuteStatementOperation: Running query 'cache table test_parquet' with 20ee2ab9-5242-4783-81cf-46115ed72610\n15/10/21 21:22:05 INFO metastore.HiveMetaStore: 49: get_table : db=default tbl=test_parquet\n15/10/21 21:22:05 INFO HiveMetaStore.audit: ugi=vagrant\tip=unknown-ip-addr\tcmd=get_table : db=default tbl=test_parquet\n15/10/21 21:22:05 INFO metastore.HiveMetaStore: 49: Opening raw store with implemenation class:org.apache.hadoop.hive.metastore.ObjectStore\n15/10/21 21:22:05 INFO metastore.ObjectStore: ObjectStore, initialize called\n15/10/21 21:22:05 INFO DataNucleus.Query: Reading in results for query \"org.datanucleus.store.rdbms.query.SQLQuery@0\" since the connection used is closing\n15/10/21 21:22:05 INFO metastore.MetaStoreDirectSql: Using direct SQL, underlying DB is MYSQL\n15/10/21 21:22:05 INFO metastore.ObjectStore: Initialized ObjectStore\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(215680) called with curMem=4196713, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_59 stored as values in memory (estimated size 210.6 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(20265) called with curMem=4412393, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_59_piece0 stored as bytes in memory (estimated size 19.8 KB, free 128.3 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_59_piece0 in memory on 192.168.99.9:50262 (size: 19.8 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO spark.SparkContext: Created broadcast 59 from run at AccessController.java:-2\n15/10/21 21:22:05 INFO metastore.HiveMetaStore: 49: get_table : db=default tbl=test_parquet\n15/10/21 21:22:05 INFO HiveMetaStore.audit: ugi=vagrant\tip=unknown-ip-addr\tcmd=get_table : db=default tbl=test_parquet\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(215680) called with curMem=4432658, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_60 stored as values in memory (estimated size 210.6 KB, free 128.1 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Removed broadcast_58_piece0 on 192.168.99.9:50262 in memory (size: 19.8 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Removed broadcast_57_piece0 on 192.168.99.9:50262 in memory (size: 21.1 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Removed broadcast_57_piece0 on slave2:46912 in memory (size: 21.1 KB, free: 534.5 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Removed broadcast_57_piece0 on slave0:46599 in memory (size: 21.1 KB, free: 534.3 MB)\n15/10/21 21:22:05 INFO spark.ContextCleaner: Cleaned accumulator 86\n15/10/21 21:22:05 INFO spark.ContextCleaner: Cleaned accumulator 84\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(20265) called with curMem=4327620, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_60_piece0 stored as bytes in memory (estimated size 19.8 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_60_piece0 in memory on 192.168.99.9:50262 (size: 19.8 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO spark.SparkContext: Created broadcast 60 from run at AccessController.java:-2\n15/10/21 21:22:05 INFO spark.SparkContext: Starting job: run at AccessController.java:-2\n15/10/21 21:22:05 INFO parquet.ParquetRelation: Reading Parquet file(s) from hdfs://192.168.99.9/user/hive/warehouse/test_parquet/part-r-00000-7cf64eb9-76ca-47c7-92aa-eb5ba879faae.gz.parquet\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Registering RDD 171 (run at AccessController.java:-2)\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Got job 24 (run at AccessController.java:-2) with 1 output partitions\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Final stage: ResultStage 34(run at AccessController.java:-2)\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Parents of final stage: List(ShuffleMapStage 33)\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Missing parents: List(ShuffleMapStage 33)\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Submitting ShuffleMapStage 33 (MapPartitionsRDD[171] at run at AccessController.java:-2), which has no missing parents\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(9472) called with curMem=4347885, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_61 stored as values in memory (estimated size 9.3 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(4838) called with curMem=4357357, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_61_piece0 stored as bytes in memory (estimated size 4.7 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_61_piece0 in memory on 192.168.99.9:50262 (size: 4.7 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO spark.SparkContext: Created broadcast 61 from broadcast at DAGScheduler.scala:861\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Submitting 1 missing tasks from ShuffleMapStage 33 (MapPartitionsRDD[171] at run at AccessController.java:-2)\n15/10/21 21:22:05 INFO cluster.YarnScheduler: Adding task set 33.0 with 1 tasks\n15/10/21 21:22:05 INFO scheduler.FairSchedulableBuilder: Added task set TaskSet_33 tasks to pool default\n15/10/21 21:22:05 INFO scheduler.TaskSetManager: Starting task 0.0 in stage 33.0 (TID 45, slave2, NODE_LOCAL, 2234 bytes)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_61_piece0 in memory on slave2:46912 (size: 4.7 KB, free: 534.5 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_60_piece0 in memory on slave2:46912 (size: 19.8 KB, free: 534.4 MB)\n15/10/21 21:22:05 INFO scheduler.TaskSetManager: Finished task 0.0 in stage 33.0 (TID 45) in 105 ms on slave2 (1/1)\n15/10/21 21:22:05 INFO cluster.YarnScheduler: Removed TaskSet 33.0, whose tasks have all completed, from pool default\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: ShuffleMapStage 33 (run at AccessController.java:-2) finished in 0.105 s\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: looking for newly runnable stages\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: running: Set()\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: waiting: Set(ResultStage 34)\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: failed: Set()\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Missing parents for ResultStage 34: List()\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: Finished stage: org.apache.spark.scheduler.StageInfo@532f49c8\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Submitting ResultStage 34 (MapPartitionsRDD[174] at run at AccessController.java:-2), which is now runnable\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: task runtime:(count: 1, mean: 105.000000, stdev: 0.000000, max: 105.000000, min: 105.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\t105.0 ms\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: shuffle bytes written:(count: 1, mean: 49.000000, stdev: 0.000000, max: 49.000000, min: 49.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t49.0 B\t49.0 B\t49.0 B\t49.0 B\t49.0 B\t49.0 B\t49.0 B\t49.0 B\t49.0 B\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(10440) called with curMem=4362195, maxMem=139009720\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: task result size:(count: 1, mean: 2381.000000, stdev: 0.000000, max: 2381.000000, min: 2381.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\t2.3 KB\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_62 stored as values in memory (estimated size 10.2 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: executor (non-fetch) time pct: (count: 1, mean: 68.571429, stdev: 0.000000, max: 68.571429, min: 68.571429)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t69 %\t69 %\t69 %\t69 %\t69 %\t69 %\t69 %\t69 %\t69 %\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: other time pct: (count: 1, mean: 31.428571, stdev: 0.000000, max: 31.428571, min: 31.428571)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t31 %\t31 %\t31 %\t31 %\t31 %\t31 %\t31 %\t31 %\t31 %\n15/10/21 21:22:05 INFO storage.MemoryStore: ensureFreeSpace(5358) called with curMem=4372635, maxMem=139009720\n15/10/21 21:22:05 INFO storage.MemoryStore: Block broadcast_62_piece0 stored as bytes in memory (estimated size 5.2 KB, free 128.4 MB)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_62_piece0 in memory on 192.168.99.9:50262 (size: 5.2 KB, free: 132.2 MB)\n15/10/21 21:22:05 INFO spark.SparkContext: Created broadcast 62 from broadcast at DAGScheduler.scala:861\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Submitting 1 missing tasks from ResultStage 34 (MapPartitionsRDD[174] at run at AccessController.java:-2)\n15/10/21 21:22:05 INFO cluster.YarnScheduler: Adding task set 34.0 with 1 tasks\n15/10/21 21:22:05 INFO scheduler.FairSchedulableBuilder: Added task set TaskSet_34 tasks to pool default\n15/10/21 21:22:05 INFO scheduler.TaskSetManager: Starting task 0.0 in stage 34.0 (TID 46, slave2, PROCESS_LOCAL, 1914 bytes)\n15/10/21 21:22:05 INFO storage.BlockManagerInfo: Added broadcast_62_piece0 in memory on slave2:46912 (size: 5.2 KB, free: 534.4 MB)\n15/10/21 21:22:05 INFO spark.MapOutputTrackerMasterEndpoint: Asked to send map output locations for shuffle 9 to slave2:43867\n15/10/21 21:22:05 INFO spark.MapOutputTrackerMaster: Size of output statuses for shuffle 9 is 135 bytes\n15/10/21 21:22:05 INFO scheduler.TaskSetManager: Finished task 0.0 in stage 34.0 (TID 46) in 48 ms on slave2 (1/1)\n15/10/21 21:22:05 INFO cluster.YarnScheduler: Removed TaskSet 34.0, whose tasks have all completed, from pool default\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: ResultStage 34 (run at AccessController.java:-2) finished in 0.047 s\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: Finished stage: org.apache.spark.scheduler.StageInfo@37a20848\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: task runtime:(count: 1, mean: 48.000000, stdev: 0.000000, max: 48.000000, min: 48.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\t48.0 ms\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: fetch wait time:(count: 1, mean: 0.000000, stdev: 0.000000, max: 0.000000, min: 0.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\t0.0 ms\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: remote bytes read:(count: 1, mean: 0.000000, stdev: 0.000000, max: 0.000000, min: 0.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0.0 B\t0.0 B\t0.0 B\t0.0 B\t0.0 B\t0.0 B\t0.0 B\t0.0 B\t0.0 B\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: task result size:(count: 1, mean: 1737.000000, stdev: 0.000000, max: 1737.000000, min: 1737.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\t1737.0 B\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: executor (non-fetch) time pct: (count: 1, mean: 29.166667, stdev: 0.000000, max: 29.166667, min: 29.166667)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t29 %\t29 %\t29 %\t29 %\t29 %\t29 %\t29 %\t29 %\t29 %\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: fetch wait time pct: (count: 1, mean: 0.000000, stdev: 0.000000, max: 0.000000, min: 0.000000)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t 0 %\t 0 %\t 0 %\t 0 %\t 0 %\t 0 %\t 0 %\t 0 %\t 0 %\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: other time pct: (count: 1, mean: 70.833333, stdev: 0.000000, max: 70.833333, min: 70.833333)\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t0%\t5%\t10%\t25%\t50%\t75%\t90%\t95%\t100%\n15/10/21 21:22:05 INFO scheduler.StatsReportListener: \t71 %\t71 %\t71 %\t71 %\t71 %\t71 %\t71 %\t71 %\t71 %\n15/10/21 21:22:05 INFO scheduler.DAGScheduler: Job 24 finished: run at AccessController.java:-2, took 0.175295 s\n{code}\n\nSpark logs for {{explain}} for {{test_parquet}}:\n\n{code}\n15/10/21 21:23:19 INFO thriftserver.SparkExecuteStatementOperation: Running query 'explain select * from test_parquet' with bae9c0bf-57f9-4c80-b745-3f0202469f3f\n15/10/21 21:23:19 INFO parse.ParseDriver: Parsing command: explain select * from test_parquet\n15/10/21 21:23:19 INFO parse.ParseDriver: Parse Completed\n15/10/21 21:23:19 INFO metastore.HiveMetaStore: 50: get_table : db=default tbl=test_parquet\n15/10/21 21:23:19 INFO HiveMetaStore.audit: ugi=vagrant\tip=unknown-ip-addr\tcmd=get_table : db=default tbl=test_parquet\n15/10/21 21:23:19 INFO metastore.HiveMetaStore: 50: Opening raw store with implemenation class:org.apache.hadoop.hive.metastore.ObjectStore\n15/10/21 21:23:19 INFO metastore.ObjectStore: ObjectStore, initialize called\n15/10/21 21:23:19 INFO DataNucleus.Query: Reading in results for query \"org.datanucleus.store.rdbms.query.SQLQuery@0\" since the connection used is closing\n15/10/21 21:23:19 INFO metastore.MetaStoreDirectSql: Using direct SQL, underlying DB is MYSQL\n15/10/21 21:23:19 INFO metastore.ObjectStore: Initialized ObjectStore\n15/10/21 21:23:19 INFO storage.MemoryStore: ensureFreeSpace(215680) called with curMem=4377993, maxMem=139009720\n15/10/21 21:23:19 INFO storage.MemoryStore: Block broadcast_63 stored as values in memory (estimated size 210.6 KB, free 128.2 MB)\n15/10/21 21:23:19 INFO storage.MemoryStore: ensureFreeSpace(20265) called with curMem=4593673, maxMem=139009720\n15/10/21 21:23:19 INFO storage.MemoryStore: Block broadcast_63_piece0 stored as bytes in memory (estimated size 19.8 KB, free 128.2 MB)\n15/10/21 21:23:19 INFO storage.BlockManagerInfo: Added broadcast_63_piece0 in memory on 192.168.99.9:50262 (size: 19.8 KB, free: 132.2 MB)\n15/10/21 21:23:19 INFO spark.SparkContext: Created broadcast 63 from run at AccessController.java:-2\n15/10/21 21:23:19 INFO thriftserver.SparkExecuteStatementOperation: Result Schema: List(plan#262)\n15/10/21 21:23:19 INFO thriftserver.SparkExecuteStatementOperation: Result Schema: List(plan#262)\n15/10/21 21:23:19 INFO thriftserver.SparkExecuteStatementOperation: Result Schema: List(plan#262)\n{code}","issue_id":"12906888","key":"SPARK-11246","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-10-29T14:58:36.000+0000","role":"gold_target","summary":"[1.5] Table cache for Parquet broken in 1.5"} {"case_id":"12914940","cluster":"JIRA-SPARK-608f77bb30b3","comments":[{"body":"[~milad.bourhani@gmail.com] I think it is not related to SPARK-9844. My cluster did not show those logs and I can still reproduce the problem.","created":"2015-11-20T18:56:10.637+0000"},{"body":"OK, good that you narrowed the problem down! In any case, let me know if you want me to report more details about my environment.","created":"2015-11-20T19:49:42.436+0000"},{"body":"I've attached the logs on SPARK-3947, just to keep all the attachments there, of course feel free to move/copy them here :)","created":"2015-11-23T12:54:41.603+0000"},{"body":"I tried \n{code}\nval q1 = sql(\"\"\"\nselect store_country,\n store_region,\n gm(amount)\nfrom receipts\nwhere amount > 50\n and store_country = 'italy'\ngroup by store_country, store_region\n\"\"\")\n{code}\n\nSeems the result is good. Looks like using built-in functions and UDAF somehow triggers the problem.","created":"2015-11-24T23:53:36.055+0000"},{"body":"The root cause is that we generate ExprId for ScalaUDAF in Executor, then grouping key and aggregation buffer could have same Id (always bound to grouping key, because it come first), cause wrong result.\n\nBack port https://github.com/apache/spark/pull/9093 into 1.5 branch, so it could be fixed.","created":"2015-12-11T20:42:33.856+0000"},{"body":"Thanks [~davies]! [~milad.bourhani@gmail.com] Can you try our latest branch 1.5 and see it is fixed for your case?","created":"2015-12-11T21:39:42.109+0000"},{"body":"[~davies] btw, which exprId was generated at executor side?","created":"2015-12-11T21:40:50.071+0000"},{"body":"Sure, I'll give it a go next week :) I'll write the results here.","created":"2015-12-11T21:49:38.744+0000"},{"body":"Hi,\n\nI can no longer replicate the bug on branch 1.5 :)\n\nI've tried several times running my example both with local and clustered mode and the result is consistently the same. Specifically, I have:\n- built Spark,\n- then built my project using Spark dependencies with version 1.5.3-SNAPSHOT,\n- run {{./make-distribution.sh}},\n- used the result in the {{dist}} directory to start up the cluster.\n\nThank you!","created":"2015-12-15T08:54:06.815+0000"},{"body":"Thank you for the update!","created":"2015-12-15T18:41:07.098+0000"}],"conversations":[{"body":"I could not reproduce it in 1.6 branch (it can be easily reproduced in 1.5). I think it is an issue in 1.5 branch.\n\nTry the following in spark 1.5 (with a cluster) and you can see the problem.\n\n{code}\nimport java.math.BigDecimal\n\nimport org.apache.spark.sql.expressions.MutableAggregationBuffer\nimport org.apache.spark.sql.expressions.UserDefinedAggregateFunction\nimport org.apache.spark.sql.Row\nimport org.apache.spark.sql.types.{StructType, StructField, DataType, DoubleType, LongType}\n\nclass GeometricMean extends UserDefinedAggregateFunction {\n def inputSchema: StructType =\n StructType(StructField(\"value\", DoubleType) :: Nil)\n\n def bufferSchema: StructType = StructType(\n StructField(\"count\", LongType) ::\n StructField(\"product\", DoubleType) :: Nil\n )\n\n def dataType: DataType = DoubleType\n\n def deterministic: Boolean = true\n\n def initialize(buffer: MutableAggregationBuffer): Unit = {\n buffer(0) = 0L\n buffer(1) = 1.0\n }\n\n def update(buffer: MutableAggregationBuffer,input: Row): Unit = {\n buffer(0) = buffer.getAs[Long](0) + 1\n buffer(1) = buffer.getAs[Double](1) * input.getAs[Double](0)\n }\n\n def merge(buffer1: MutableAggregationBuffer, buffer2: Row): Unit = {\n buffer1(0) = buffer1.getAs[Long](0) + buffer2.getAs[Long](0)\n buffer1(1) = buffer1.getAs[Double](1) * buffer2.getAs[Double](1)\n }\n\n def evaluate(buffer: Row): Any = {\n math.pow(buffer.getDouble(1), 1.0d / buffer.getLong(0))\n }\n}\n\nsqlContext.udf.register(\"gm\", new GeometricMean)\n\nval df = Seq(\n (1, \"italy\", \"emilia\", 42, BigDecimal.valueOf(100, 0), \"john\"),\n (2, \"italy\", \"toscana\", 42, BigDecimal.valueOf(505, 1), \"jim\"),\n (3, \"italy\", \"puglia\", 42, BigDecimal.valueOf(70, 0), \"jenn\"),\n (4, \"italy\", \"emilia\", 42, BigDecimal.valueOf(75 ,0), \"jack\"),\n (5, \"uk\", \"london\", 42, BigDecimal.valueOf(200 ,0), \"carl\"),\n (6, \"italy\", \"emilia\", 42, BigDecimal.valueOf(42, 0), \"john\")).\n toDF(\"receipt_id\", \"store_country\", \"store_region\", \"store_id\", \"amount\", \"seller_name\")\ndf.registerTempTable(\"receipts\")\n \nval q = sql(\"\"\"\nselect store_country,\n store_region,\n avg(amount),\n sum(amount),\n gm(amount)\nfrom receipts\nwhere amount > 50\n and store_country = 'italy'\ngroup by store_country, store_region\n\"\"\")\n\nq.show\n{code}\n","from":"reporter","subject":"UDAF may nondeterministically generate wrong results"},{"body":"[~milad.bourhani@gmail.com] I think it is not related to SPARK-9844. My cluster did not show those logs and I can still reproduce the problem.","from":"developer"},{"body":"OK, good that you narrowed the problem down! In any case, let me know if you want me to report more details about my environment.","from":"developer"},{"body":"I've attached the logs on SPARK-3947, just to keep all the attachments there, of course feel free to move/copy them here :)","from":"developer"},{"body":"I tried \n{code}\nval q1 = sql(\"\"\"\nselect store_country,\n store_region,\n gm(amount)\nfrom receipts\nwhere amount > 50\n and store_country = 'italy'\ngroup by store_country, store_region\n\"\"\")\n{code}\n\nSeems the result is good. Looks like using built-in functions and UDAF somehow triggers the problem.","from":"developer"},{"body":"The root cause is that we generate ExprId for ScalaUDAF in Executor, then grouping key and aggregation buffer could have same Id (always bound to grouping key, because it come first), cause wrong result.\n\nBack port https://github.com/apache/spark/pull/9093 into 1.5 branch, so it could be fixed.","from":"developer"},{"body":"Thanks [~davies]! [~milad.bourhani@gmail.com] Can you try our latest branch 1.5 and see it is fixed for your case?","from":"developer"},{"body":"[~davies] btw, which exprId was generated at executor side?","from":"developer"},{"body":"Sure, I'll give it a go next week :) I'll write the results here.","from":"developer"},{"body":"Hi,\n\nI can no longer replicate the bug on branch 1.5 :)\n\nI've tried several times running my example both with local and clustered mode and the result is consistently the same. Specifically, I have:\n- built Spark,\n- then built my project using Spark dependencies with version 1.5.3-SNAPSHOT,\n- run {{./make-distribution.sh}},\n- used the result in the {{dist}} directory to start up the cluster.\n\nThank you!","from":"developer"},{"body":"Thank you for the update!","from":"developer"}],"created":"2015-11-20T18:48:14.000+0000","description":"I could not reproduce it in 1.6 branch (it can be easily reproduced in 1.5). I think it is an issue in 1.5 branch.\n\nTry the following in spark 1.5 (with a cluster) and you can see the problem.\n\n{code}\nimport java.math.BigDecimal\n\nimport org.apache.spark.sql.expressions.MutableAggregationBuffer\nimport org.apache.spark.sql.expressions.UserDefinedAggregateFunction\nimport org.apache.spark.sql.Row\nimport org.apache.spark.sql.types.{StructType, StructField, DataType, DoubleType, LongType}\n\nclass GeometricMean extends UserDefinedAggregateFunction {\n def inputSchema: StructType =\n StructType(StructField(\"value\", DoubleType) :: Nil)\n\n def bufferSchema: StructType = StructType(\n StructField(\"count\", LongType) ::\n StructField(\"product\", DoubleType) :: Nil\n )\n\n def dataType: DataType = DoubleType\n\n def deterministic: Boolean = true\n\n def initialize(buffer: MutableAggregationBuffer): Unit = {\n buffer(0) = 0L\n buffer(1) = 1.0\n }\n\n def update(buffer: MutableAggregationBuffer,input: Row): Unit = {\n buffer(0) = buffer.getAs[Long](0) + 1\n buffer(1) = buffer.getAs[Double](1) * input.getAs[Double](0)\n }\n\n def merge(buffer1: MutableAggregationBuffer, buffer2: Row): Unit = {\n buffer1(0) = buffer1.getAs[Long](0) + buffer2.getAs[Long](0)\n buffer1(1) = buffer1.getAs[Double](1) * buffer2.getAs[Double](1)\n }\n\n def evaluate(buffer: Row): Any = {\n math.pow(buffer.getDouble(1), 1.0d / buffer.getLong(0))\n }\n}\n\nsqlContext.udf.register(\"gm\", new GeometricMean)\n\nval df = Seq(\n (1, \"italy\", \"emilia\", 42, BigDecimal.valueOf(100, 0), \"john\"),\n (2, \"italy\", \"toscana\", 42, BigDecimal.valueOf(505, 1), \"jim\"),\n (3, \"italy\", \"puglia\", 42, BigDecimal.valueOf(70, 0), \"jenn\"),\n (4, \"italy\", \"emilia\", 42, BigDecimal.valueOf(75 ,0), \"jack\"),\n (5, \"uk\", \"london\", 42, BigDecimal.valueOf(200 ,0), \"carl\"),\n (6, \"italy\", \"emilia\", 42, BigDecimal.valueOf(42, 0), \"john\")).\n toDF(\"receipt_id\", \"store_country\", \"store_region\", \"store_id\", \"amount\", \"seller_name\")\ndf.registerTempTable(\"receipts\")\n \nval q = sql(\"\"\"\nselect store_country,\n store_region,\n avg(amount),\n sum(amount),\n gm(amount)\nfrom receipts\nwhere amount > 50\n and store_country = 'italy'\ngroup by store_country, store_region\n\"\"\")\n\nq.show\n{code}\n","issue_id":"12914940","key":"SPARK-11885","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-12-11T20:44:37.000+0000","role":"gold_target","summary":"UDAF may nondeterministically generate wrong results"} {"case_id":"12960283","cluster":"JIRA-SPARK-ab69a396fd9c","comments":[{"body":"Changing generatedOrdering in LazilyGeneratedOrdering to \n{noformat}\nprivate[this] lazy val generatedOrdering = GenerateOrdering.generate(ordering)\n{noformat} \nsolves the issue and the query runs fine. Thought of checking with committers opinion before posting the PR for this.","created":"2016-04-20T09:55:50.731+0000"},{"body":"User 'rajeshbalamohan' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/12661","created":"2016-04-25T13:15:03.986+0000"},{"body":"User 'bomeng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13141","created":"2016-05-16T23:48:03.585+0000"},{"body":"User 'sameeragarwal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13466","created":"2016-06-02T08:20:04.859+0000"},{"body":"[~rajesh.balamohan] [~bomeng] Instead of relying on null checks or lazy vals, it might be more efficient to just make LazilyGeneratedOrdering KryoSerializable. Let us know your thoughts!","created":"2016-06-02T08:23:40.619+0000"},{"body":"I agree that we should avoid null checks or lazy vals.\nWhen I introduced {{LazilyGeneratedOrdering}}, we decided not to use lazy val because it might cause a performance penalty.\nPlease see https://github.com/apache/spark/pull/10894#discussion_r52691349.","created":"2016-06-02T12:20:26.122+0000"},{"body":" I think this is a good approach. Thanks.","created":"2016-06-02T17:55:50.104+0000"}],"conversations":[{"body":"codebase: spark master\n\nDataSet: TPC-DS\n\nClient: $SPARK_HOME/bin/beeline\n\nExample query to reproduce the issue: \nselect i_item_id from item order by i_item_id limit 10;\n\nExplain plan output\n{noformat}\nexplain select i_item_id from item order by i_item_id limit 10;\n+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+--+\n| plan |\n+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+--+\n| == Physical Plan ==\nTakeOrderedAndProject(limit=10, orderBy=[i_item_id#1229 ASC], output=[i_item_id#1229])\n+- WholeStageCodegen\n : +- Project [i_item_id#1229]\n : +- Scan HadoopFiles[i_item_id#1229] Format: ORC, PushedFilters: [], ReadSchema: struct |\n+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+--+\n{noformat}\n\nException:\n{noformat}\nTaskResultGetter: Exception while getting task result\ncom.esotericsoftware.kryo.KryoException: java.lang.NullPointerException\nSerialization trace:\nunderlying (org.apache.spark.util.BoundedPriorityQueue)\n\tat com.esotericsoftware.kryo.serializers.ObjectField.read(ObjectField.java:144)\n\tat com.esotericsoftware.kryo.serializers.FieldSerializer.read(FieldSerializer.java:551)\n\tat com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:790)\n\tat com.twitter.chill.SomeSerializer.read(SomeSerializer.scala:25)\n\tat com.twitter.chill.SomeSerializer.read(SomeSerializer.scala:19)\n\tat com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:790)\n\tat org.apache.spark.serializer.KryoSerializerInstance.deserialize(KryoSerializer.scala:312)\n\tat org.apache.spark.scheduler.DirectTaskResult.value(TaskResult.scala:87)\n\tat org.apache.spark.scheduler.TaskResultGetter$$anon$2$$anonfun$run$1.apply$mcV$sp(TaskResultGetter.scala:66)\n\tat org.apache.spark.scheduler.TaskResultGetter$$anon$2$$anonfun$run$1.apply(TaskResultGetter.scala:57)\n\tat org.apache.spark.scheduler.TaskResultGetter$$anon$2$$anonfun$run$1.apply(TaskResultGetter.scala:57)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1791)\n\tat org.apache.spark.scheduler.TaskResultGetter$$anon$2.run(TaskResultGetter.scala:56)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: java.lang.NullPointerException\n\tat org.apache.spark.sql.catalyst.expressions.codegen.LazilyGeneratedOrdering.compare(GenerateOrdering.scala:157)\n\tat org.apache.spark.sql.catalyst.expressions.codegen.LazilyGeneratedOrdering.compare(GenerateOrdering.scala:148)\n\tat scala.math.Ordering$$anon$4.compare(Ordering.scala:111)\n\tat java.util.PriorityQueue.siftUpUsingComparator(PriorityQueue.java:669)\n\tat java.util.PriorityQueue.siftUp(PriorityQueue.java:645)\n\tat java.util.PriorityQueue.offer(PriorityQueue.java:344)\n\tat java.util.PriorityQueue.add(PriorityQueue.java:321)\n\tat com.twitter.chill.java.PriorityQueueSerializer.read(PriorityQueueSerializer.java:78)\n\tat com.twitter.chill.java.PriorityQueueSerializer.read(PriorityQueueSerializer.java:31)\n\tat com.esotericsoftware.kryo.Kryo.readObject(Kryo.java:708)\n\tat com.esotericsoftware.kryo.serializers.ObjectField.read(ObjectField.java:125)\n{noformat}","from":"reporter","subject":"LazilyGenerateOrdering throws NullPointerException"},{"body":"Changing generatedOrdering in LazilyGeneratedOrdering to \n{noformat}\nprivate[this] lazy val generatedOrdering = GenerateOrdering.generate(ordering)\n{noformat} \nsolves the issue and the query runs fine. Thought of checking with committers opinion before posting the PR for this.","from":"developer"},{"body":"User 'rajeshbalamohan' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/12661","from":"developer"},{"body":"User 'bomeng' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13141","from":"developer"},{"body":"User 'sameeragarwal' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13466","from":"developer"},{"body":"[~rajesh.balamohan] [~bomeng] Instead of relying on null checks or lazy vals, it might be more efficient to just make LazilyGeneratedOrdering KryoSerializable. Let us know your thoughts!","from":"developer"},{"body":"I agree that we should avoid null checks or lazy vals.\nWhen I introduced {{LazilyGeneratedOrdering}}, we decided not to use lazy val because it might cause a performance penalty.\nPlease see https://github.com/apache/spark/pull/10894#discussion_r52691349.","from":"developer"},{"body":" I think this is a good approach. Thanks.","from":"developer"}],"created":"2016-04-20T09:53:04.000+0000","description":"codebase: spark master\n\nDataSet: TPC-DS\n\nClient: $SPARK_HOME/bin/beeline\n\nExample query to reproduce the issue: \nselect i_item_id from item order by i_item_id limit 10;\n\nExplain plan output\n{noformat}\nexplain select i_item_id from item order by i_item_id limit 10;\n+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+--+\n| plan |\n+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+--+\n| == Physical Plan ==\nTakeOrderedAndProject(limit=10, orderBy=[i_item_id#1229 ASC], output=[i_item_id#1229])\n+- WholeStageCodegen\n : +- Project [i_item_id#1229]\n : +- Scan HadoopFiles[i_item_id#1229] Format: ORC, PushedFilters: [], ReadSchema: struct |\n+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+--+\n{noformat}\n\nException:\n{noformat}\nTaskResultGetter: Exception while getting task result\ncom.esotericsoftware.kryo.KryoException: java.lang.NullPointerException\nSerialization trace:\nunderlying (org.apache.spark.util.BoundedPriorityQueue)\n\tat com.esotericsoftware.kryo.serializers.ObjectField.read(ObjectField.java:144)\n\tat com.esotericsoftware.kryo.serializers.FieldSerializer.read(FieldSerializer.java:551)\n\tat com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:790)\n\tat com.twitter.chill.SomeSerializer.read(SomeSerializer.scala:25)\n\tat com.twitter.chill.SomeSerializer.read(SomeSerializer.scala:19)\n\tat com.esotericsoftware.kryo.Kryo.readClassAndObject(Kryo.java:790)\n\tat org.apache.spark.serializer.KryoSerializerInstance.deserialize(KryoSerializer.scala:312)\n\tat org.apache.spark.scheduler.DirectTaskResult.value(TaskResult.scala:87)\n\tat org.apache.spark.scheduler.TaskResultGetter$$anon$2$$anonfun$run$1.apply$mcV$sp(TaskResultGetter.scala:66)\n\tat org.apache.spark.scheduler.TaskResultGetter$$anon$2$$anonfun$run$1.apply(TaskResultGetter.scala:57)\n\tat org.apache.spark.scheduler.TaskResultGetter$$anon$2$$anonfun$run$1.apply(TaskResultGetter.scala:57)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1791)\n\tat org.apache.spark.scheduler.TaskResultGetter$$anon$2.run(TaskResultGetter.scala:56)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: java.lang.NullPointerException\n\tat org.apache.spark.sql.catalyst.expressions.codegen.LazilyGeneratedOrdering.compare(GenerateOrdering.scala:157)\n\tat org.apache.spark.sql.catalyst.expressions.codegen.LazilyGeneratedOrdering.compare(GenerateOrdering.scala:148)\n\tat scala.math.Ordering$$anon$4.compare(Ordering.scala:111)\n\tat java.util.PriorityQueue.siftUpUsingComparator(PriorityQueue.java:669)\n\tat java.util.PriorityQueue.siftUp(PriorityQueue.java:645)\n\tat java.util.PriorityQueue.offer(PriorityQueue.java:344)\n\tat java.util.PriorityQueue.add(PriorityQueue.java:321)\n\tat com.twitter.chill.java.PriorityQueueSerializer.read(PriorityQueueSerializer.java:78)\n\tat com.twitter.chill.java.PriorityQueueSerializer.read(PriorityQueueSerializer.java:31)\n\tat com.esotericsoftware.kryo.Kryo.readObject(Kryo.java:708)\n\tat com.esotericsoftware.kryo.serializers.ObjectField.read(ObjectField.java:125)\n{noformat}","issue_id":"12960283","key":"SPARK-14752","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-06-02T17:57:45.000+0000","role":"gold_target","summary":"LazilyGenerateOrdering throws NullPointerException"} {"case_id":"12963109","cluster":"JIRA-SPARK-b89aff5afa41","comments":[{"body":"I have tried on master branch, it works fine with the latest code. ","created":"2016-04-27T22:50:58.421+0000"},{"body":"[~bomeng] I have just retested it with the last master commit be317d4a90b3ca906fefeb438f89a09b1c7da5a8 and I am still getting the same error.\n\nhave you tested this with HDFS?\n\n{code:title=spakr-shell}\nscala> ds.write.mode(org.apache.spark.sql.SaveMode.Overwrite).format(\"parquet\").partitionBy(\"id\").save(\"/user/spark/test.parquet\")\nSLF4J: Failed to load class \"org.slf4j.impl.StaticLoggerBinder\". \nSLF4J: Defaulting to no-operation (NOP) logger implementation\nSLF4J: See http://www.slf4j.org/codes.html#StaticLoggerBinder for further details.\njava.io.FileNotFoundException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n...\n...\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invokeMethod(RetryInvocationHandler.java:187)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invoke(RetryInvocationHandler.java:102)\n at com.sun.proxy.$Proxy14.getBlockLocations(Unknown Source)\n at org.apache.hadoop.hdfs.DFSClient.callGetBlockLocations(DFSClient.java:1240)\n ... 78 more\n{code}","created":"2016-04-28T06:46:01.540+0000"},{"body":"[~syepes] I faced the same exception when I try to query partitioned table on HDFS. Using the latest commit on master branch.","created":"2016-05-02T20:13:20.051+0000"},{"body":"Try explicitly specifying the scheme in the URL","created":"2016-05-11T11:15:38.885+0000"},{"body":"Hello [~sowen] using the full URL I still get the same error:\n\n{code}\nscala> spark.read.format(\"parquet\").load(\"hdfs://master:8020/user/spark/test.parquet\").show(1)\njava.io.FileNotFoundException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n{code}\n","created":"2016-05-11T11:47:43.788+0000"},{"body":"Does the data exist? is it readable? some basic debugging info like that would be useful","created":"2016-05-11T11:50:47.851+0000"},{"body":"[~sowen], The partitioned data exists and is readable by Spark, I can read it if I manually specify the partition:\n\n{code}\nscala> spark.read.format(\"parquet\").load(\"hdfs://master:8020/user/spark/test.parquet/id=0\").show\n+-----+ \n| text|\n+-----+\n|hello|\n|world|\n+-----+\n\nscala> spark.read.format(\"parquet\").load(\"hdfs://master:8020/user/spark/test.parquet/id=1\").show\n+-----+ \n| text|\n+-----+\n|hello|\n|there|\n+-----+\n{code}","created":"2016-05-11T12:17:42.204+0000"},{"body":"I have the same issue reading a partitioned parquet table using Spark 2.0.0 (which was saved originally using Spark 1.6.1)","created":"2016-05-11T21:58:51.091+0000"},{"body":"I think this issue was introduced around SPARK-13664, but the thing is that there have been many underlining changes.\nIf you need anymore debugging info let me know. ","created":"2016-05-12T09:19:52.846+0000"},{"body":"As you can see in the description writing is also broken as it throws an exception. It doesn't write the schema properly (https://github.com/apache/spark/blob/bc3760d405cc8c3ffcd957b188afa8b7e3b1f824/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/DataSource.scala#L431) because of the `resolveRelation` eventually hitting the same kind of `FileNotFoundException`.","created":"2016-05-17T08:49:12.701+0000"},{"body":"I can recreate the problem with hdfs location. and I have a patch for it now. I will submit a PR soon. \n\nThe actual results now is following, as expected:\n{code}\nscala> spark.read.format(\"parquet\").load(\"hdfs://bdavm009.svl.ibm.com:8020/user/spark/SPARK-14959_part\").show\n+-----+---+\n| text| id|\n+-----+---+\n|hello| 0|\n|world| 0|\n|hello| 1|\n|there| 1|\n+-----+---+\n\n spark.read.format(\"orc\").load(\"hdfs://bdavm009.svl.ibm.com:8020/user/spark/SPARK-14959_orc\").show\n+-----+---+\n| text| id|\n+-----+---+\n|hello| 0|\n|world| 0|\n|hello| 1|\n|there| 1|\n+-----+---+\n{code}","created":"2016-06-01T23:55:40.846+0000"},{"body":"User 'xwu0226' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13463","created":"2016-06-02T05:10:03.790+0000"},{"body":"Issue resolved by pull request 13463\n[https://github.com/apache/spark/pull/13463]","created":"2016-06-03T05:49:48.975+0000"}],"conversations":[{"body":"Hello,\n\nI have noticed that in the pasts days there is an issue when trying to read partitioned files from HDFS.\n\nI am running on Spark master branch #c544356\n\n\nThe write actually works but the read fails.\n\n{code:title=Issue Reproduction}\ncase class Data(id: Int, text: String)\nval ds = spark.createDataset( Seq(Data(0, \"hello\"), Data(1, \"hello\"), Data(0, \"world\"), Data(1, \"there\")) )\n\nscala> ds.write.mode(org.apache.spark.sql.SaveMode.Overwrite).format(\"parquet\").partitionBy(\"id\").save(\"/user/spark/test.parquet\")\nSLF4J: Failed to load class \"org.slf4j.impl.StaticLoggerBinder\". \nSLF4J: Defaulting to no-operation (NOP) logger implementation\nSLF4J: See http://www.slf4j.org/codes.html#StaticLoggerBinder for further details.\njava.io.FileNotFoundException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1712)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getBlockLocations(NameNodeRpcServer.java:652)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.getBlockLocations(ClientNamenodeProtocolServerSideTranslatorPB.java:365)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:969)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2151)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2147)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1657)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:2145)\n\n at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n at java.lang.reflect.Constructor.newInstance(Constructor.java:423)\n at org.apache.hadoop.ipc.RemoteException.instantiateException(RemoteException.java:106)\n at org.apache.hadoop.ipc.RemoteException.unwrapRemoteException(RemoteException.java:73)\n at org.apache.hadoop.hdfs.DFSClient.callGetBlockLocations(DFSClient.java:1242)\n at org.apache.hadoop.hdfs.DFSClient.getLocatedBlocks(DFSClient.java:1227)\n at org.apache.hadoop.hdfs.DFSClient.getBlockLocations(DFSClient.java:1285)\n at org.apache.hadoop.hdfs.DistributedFileSystem$1.doCall(DistributedFileSystem.java:221)\n at org.apache.hadoop.hdfs.DistributedFileSystem$1.doCall(DistributedFileSystem.java:217)\n at org.apache.hadoop.fs.FileSystemLinkResolver.resolve(FileSystemLinkResolver.java:81)\n at org.apache.hadoop.hdfs.DistributedFileSystem.getFileBlockLocations(DistributedFileSystem.java:228)\n at org.apache.hadoop.hdfs.DistributedFileSystem.getFileBlockLocations(DistributedFileSystem.java:209)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9$$anonfun$apply$4.apply(fileSourceInterfaces.scala:372)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9$$anonfun$apply$4.apply(fileSourceInterfaces.scala:360)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n at scala.collection.mutable.ArrayOps$ofRef.foreach(ArrayOps.scala:186)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.mutable.ArrayOps$ofRef.map(ArrayOps.scala:186)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9.apply(fileSourceInterfaces.scala:360)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9.apply(fileSourceInterfaces.scala:348)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n at scala.collection.mutable.WrappedArray.foreach(WrappedArray.scala:35)\n at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\n at scala.collection.AbstractTraversable.flatMap(Traversable.scala:104)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.listLeafFiles(fileSourceInterfaces.scala:348)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.refresh(fileSourceInterfaces.scala:447)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.(fileSourceInterfaces.scala:291)\n at org.apache.spark.sql.execution.datasources.DataSource.resolveRelation(DataSource.scala:314)\n at org.apache.spark.sql.execution.datasources.DataSource.write(DataSource.scala:431)\n at org.apache.spark.sql.DataFrameWriter.save(DataFrameWriter.scala:246)\n at org.apache.spark.sql.DataFrameWriter.save(DataFrameWriter.scala:229)\n ... 48 elided\nCaused by: org.apache.hadoop.ipc.RemoteException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1712)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getBlockLocations(NameNodeRpcServer.java:652)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.getBlockLocations(ClientNamenodeProtocolServerSideTranslatorPB.java:365)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:969)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2151)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2147)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1657)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:2145)\n\n at org.apache.hadoop.ipc.Client.call(Client.java:1476)\n at org.apache.hadoop.ipc.Client.call(Client.java:1407)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:229)\n at com.sun.proxy.$Proxy13.getBlockLocations(Unknown Source)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolTranslatorPB.getBlockLocations(ClientNamenodeProtocolTranslatorPB.java:255)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invokeMethod(RetryInvocationHandler.java:187)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invoke(RetryInvocationHandler.java:102)\n at com.sun.proxy.$Proxy14.getBlockLocations(Unknown Source)\n at org.apache.hadoop.hdfs.DFSClient.callGetBlockLocations(DFSClient.java:1240)\n ... 78 more\n\n// Reading the specific partitioned data works \nscala> spark.read.format(\"parquet\").load(\"/user/spark/test.parquet/id=0\").show\n+-----+ \n| text|\n+-----+\n|hello|\n|world|\n+-----+\n\n// Reading all the partitions fails\nscala> spark.read.format(\"parquet\").load(\"/user/spark/test.parquet\").show\njava.io.FileNotFoundException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1712)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getBlockLocations(NameNodeRpcServer.java:652)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.getBlockLocations(ClientNamenodeProtocolServerSideTranslatorPB.java:365)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:969)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2151)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2147)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1657)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:2145)\n\n at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n at java.lang.reflect.Constructor.newInstance(Constructor.java:423)\n at org.apache.hadoop.ipc.RemoteException.instantiateException(RemoteException.java:106)\n at org.apache.hadoop.ipc.RemoteException.unwrapRemoteException(RemoteException.java:73)\n at org.apache.hadoop.hdfs.DFSClient.callGetBlockLocations(DFSClient.java:1242)\n at org.apache.hadoop.hdfs.DFSClient.getLocatedBlocks(DFSClient.java:1227)\n at org.apache.hadoop.hdfs.DFSClient.getBlockLocations(DFSClient.java:1285)\n at org.apache.hadoop.hdfs.DistributedFileSystem$1.doCall(DistributedFileSystem.java:221)\n at org.apache.hadoop.hdfs.DistributedFileSystem$1.doCall(DistributedFileSystem.java:217)\n at org.apache.hadoop.fs.FileSystemLinkResolver.resolve(FileSystemLinkResolver.java:81)\n at org.apache.hadoop.hdfs.DistributedFileSystem.getFileBlockLocations(DistributedFileSystem.java:228)\n at org.apache.hadoop.hdfs.DistributedFileSystem.getFileBlockLocations(DistributedFileSystem.java:209)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9$$anonfun$apply$4.apply(fileSourceInterfaces.scala:372)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9$$anonfun$apply$4.apply(fileSourceInterfaces.scala:360)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n at scala.collection.mutable.ArrayOps$ofRef.foreach(ArrayOps.scala:186)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.mutable.ArrayOps$ofRef.map(ArrayOps.scala:186)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9.apply(fileSourceInterfaces.scala:360)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9.apply(fileSourceInterfaces.scala:348)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n at scala.collection.mutable.WrappedArray.foreach(WrappedArray.scala:35)\n at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\n at scala.collection.AbstractTraversable.flatMap(Traversable.scala:104)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.listLeafFiles(fileSourceInterfaces.scala:348)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.refresh(fileSourceInterfaces.scala:447)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.(fileSourceInterfaces.scala:291)\n at org.apache.spark.sql.execution.datasources.DataSource.resolveRelation(DataSource.scala:314)\n at org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:132)\n at org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:142)\n ... 48 elided\nCaused by: org.apache.hadoop.ipc.RemoteException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1712)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getBlockLocations(NameNodeRpcServer.java:652)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.getBlockLocations(ClientNamenodeProtocolServerSideTranslatorPB.java:365)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:969)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2151)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2147)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1657)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:2145)\n\n at org.apache.hadoop.ipc.Client.call(Client.java:1476)\n at org.apache.hadoop.ipc.Client.call(Client.java:1407)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:229)\n at com.sun.proxy.$Proxy13.getBlockLocations(Unknown Source)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolTranslatorPB.getBlockLocations(ClientNamenodeProtocolTranslatorPB.java:255)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invokeMethod(RetryInvocationHandler.java:187)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invoke(RetryInvocationHandler.java:102)\n at com.sun.proxy.$Proxy14.getBlockLocations(Unknown Source)\n at org.apache.hadoop.hdfs.DFSClient.callGetBlockLocations(DFSClient.java:1240)\n ... 77 more\n{code}\n\n\n\n","from":"reporter","subject":"​Problem Reading partitioned ORC or Parquet files"},{"body":"I have tried on master branch, it works fine with the latest code. ","from":"developer"},{"body":"[~bomeng] I have just retested it with the last master commit be317d4a90b3ca906fefeb438f89a09b1c7da5a8 and I am still getting the same error.\n\nhave you tested this with HDFS?\n\n{code:title=spakr-shell}\nscala> ds.write.mode(org.apache.spark.sql.SaveMode.Overwrite).format(\"parquet\").partitionBy(\"id\").save(\"/user/spark/test.parquet\")\nSLF4J: Failed to load class \"org.slf4j.impl.StaticLoggerBinder\". \nSLF4J: Defaulting to no-operation (NOP) logger implementation\nSLF4J: See http://www.slf4j.org/codes.html#StaticLoggerBinder for further details.\njava.io.FileNotFoundException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n...\n...\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invokeMethod(RetryInvocationHandler.java:187)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invoke(RetryInvocationHandler.java:102)\n at com.sun.proxy.$Proxy14.getBlockLocations(Unknown Source)\n at org.apache.hadoop.hdfs.DFSClient.callGetBlockLocations(DFSClient.java:1240)\n ... 78 more\n{code}","from":"developer"},{"body":"[~syepes] I faced the same exception when I try to query partitioned table on HDFS. Using the latest commit on master branch.","from":"developer"},{"body":"Try explicitly specifying the scheme in the URL","from":"developer"},{"body":"Hello [~sowen] using the full URL I still get the same error:\n\n{code}\nscala> spark.read.format(\"parquet\").load(\"hdfs://master:8020/user/spark/test.parquet\").show(1)\njava.io.FileNotFoundException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n{code}\n","from":"developer"},{"body":"Does the data exist? is it readable? some basic debugging info like that would be useful","from":"developer"},{"body":"[~sowen], The partitioned data exists and is readable by Spark, I can read it if I manually specify the partition:\n\n{code}\nscala> spark.read.format(\"parquet\").load(\"hdfs://master:8020/user/spark/test.parquet/id=0\").show\n+-----+ \n| text|\n+-----+\n|hello|\n|world|\n+-----+\n\nscala> spark.read.format(\"parquet\").load(\"hdfs://master:8020/user/spark/test.parquet/id=1\").show\n+-----+ \n| text|\n+-----+\n|hello|\n|there|\n+-----+\n{code}","from":"developer"},{"body":"I have the same issue reading a partitioned parquet table using Spark 2.0.0 (which was saved originally using Spark 1.6.1)","from":"developer"},{"body":"I think this issue was introduced around SPARK-13664, but the thing is that there have been many underlining changes.\nIf you need anymore debugging info let me know. ","from":"developer"},{"body":"As you can see in the description writing is also broken as it throws an exception. It doesn't write the schema properly (https://github.com/apache/spark/blob/bc3760d405cc8c3ffcd957b188afa8b7e3b1f824/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/DataSource.scala#L431) because of the `resolveRelation` eventually hitting the same kind of `FileNotFoundException`.","from":"developer"},{"body":"I can recreate the problem with hdfs location. and I have a patch for it now. I will submit a PR soon. \n\nThe actual results now is following, as expected:\n{code}\nscala> spark.read.format(\"parquet\").load(\"hdfs://bdavm009.svl.ibm.com:8020/user/spark/SPARK-14959_part\").show\n+-----+---+\n| text| id|\n+-----+---+\n|hello| 0|\n|world| 0|\n|hello| 1|\n|there| 1|\n+-----+---+\n\n spark.read.format(\"orc\").load(\"hdfs://bdavm009.svl.ibm.com:8020/user/spark/SPARK-14959_orc\").show\n+-----+---+\n| text| id|\n+-----+---+\n|hello| 0|\n|world| 0|\n|hello| 1|\n|there| 1|\n+-----+---+\n{code}","from":"developer"},{"body":"User 'xwu0226' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13463","from":"developer"},{"body":"Issue resolved by pull request 13463\n[https://github.com/apache/spark/pull/13463]","from":"developer"}],"created":"2016-04-27T15:09:23.000+0000","description":"Hello,\n\nI have noticed that in the pasts days there is an issue when trying to read partitioned files from HDFS.\n\nI am running on Spark master branch #c544356\n\n\nThe write actually works but the read fails.\n\n{code:title=Issue Reproduction}\ncase class Data(id: Int, text: String)\nval ds = spark.createDataset( Seq(Data(0, \"hello\"), Data(1, \"hello\"), Data(0, \"world\"), Data(1, \"there\")) )\n\nscala> ds.write.mode(org.apache.spark.sql.SaveMode.Overwrite).format(\"parquet\").partitionBy(\"id\").save(\"/user/spark/test.parquet\")\nSLF4J: Failed to load class \"org.slf4j.impl.StaticLoggerBinder\". \nSLF4J: Defaulting to no-operation (NOP) logger implementation\nSLF4J: See http://www.slf4j.org/codes.html#StaticLoggerBinder for further details.\njava.io.FileNotFoundException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1712)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getBlockLocations(NameNodeRpcServer.java:652)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.getBlockLocations(ClientNamenodeProtocolServerSideTranslatorPB.java:365)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:969)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2151)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2147)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1657)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:2145)\n\n at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n at java.lang.reflect.Constructor.newInstance(Constructor.java:423)\n at org.apache.hadoop.ipc.RemoteException.instantiateException(RemoteException.java:106)\n at org.apache.hadoop.ipc.RemoteException.unwrapRemoteException(RemoteException.java:73)\n at org.apache.hadoop.hdfs.DFSClient.callGetBlockLocations(DFSClient.java:1242)\n at org.apache.hadoop.hdfs.DFSClient.getLocatedBlocks(DFSClient.java:1227)\n at org.apache.hadoop.hdfs.DFSClient.getBlockLocations(DFSClient.java:1285)\n at org.apache.hadoop.hdfs.DistributedFileSystem$1.doCall(DistributedFileSystem.java:221)\n at org.apache.hadoop.hdfs.DistributedFileSystem$1.doCall(DistributedFileSystem.java:217)\n at org.apache.hadoop.fs.FileSystemLinkResolver.resolve(FileSystemLinkResolver.java:81)\n at org.apache.hadoop.hdfs.DistributedFileSystem.getFileBlockLocations(DistributedFileSystem.java:228)\n at org.apache.hadoop.hdfs.DistributedFileSystem.getFileBlockLocations(DistributedFileSystem.java:209)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9$$anonfun$apply$4.apply(fileSourceInterfaces.scala:372)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9$$anonfun$apply$4.apply(fileSourceInterfaces.scala:360)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n at scala.collection.mutable.ArrayOps$ofRef.foreach(ArrayOps.scala:186)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.mutable.ArrayOps$ofRef.map(ArrayOps.scala:186)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9.apply(fileSourceInterfaces.scala:360)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9.apply(fileSourceInterfaces.scala:348)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n at scala.collection.mutable.WrappedArray.foreach(WrappedArray.scala:35)\n at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\n at scala.collection.AbstractTraversable.flatMap(Traversable.scala:104)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.listLeafFiles(fileSourceInterfaces.scala:348)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.refresh(fileSourceInterfaces.scala:447)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.(fileSourceInterfaces.scala:291)\n at org.apache.spark.sql.execution.datasources.DataSource.resolveRelation(DataSource.scala:314)\n at org.apache.spark.sql.execution.datasources.DataSource.write(DataSource.scala:431)\n at org.apache.spark.sql.DataFrameWriter.save(DataFrameWriter.scala:246)\n at org.apache.spark.sql.DataFrameWriter.save(DataFrameWriter.scala:229)\n ... 48 elided\nCaused by: org.apache.hadoop.ipc.RemoteException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1712)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getBlockLocations(NameNodeRpcServer.java:652)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.getBlockLocations(ClientNamenodeProtocolServerSideTranslatorPB.java:365)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:969)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2151)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2147)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1657)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:2145)\n\n at org.apache.hadoop.ipc.Client.call(Client.java:1476)\n at org.apache.hadoop.ipc.Client.call(Client.java:1407)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:229)\n at com.sun.proxy.$Proxy13.getBlockLocations(Unknown Source)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolTranslatorPB.getBlockLocations(ClientNamenodeProtocolTranslatorPB.java:255)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invokeMethod(RetryInvocationHandler.java:187)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invoke(RetryInvocationHandler.java:102)\n at com.sun.proxy.$Proxy14.getBlockLocations(Unknown Source)\n at org.apache.hadoop.hdfs.DFSClient.callGetBlockLocations(DFSClient.java:1240)\n ... 78 more\n\n// Reading the specific partitioned data works \nscala> spark.read.format(\"parquet\").load(\"/user/spark/test.parquet/id=0\").show\n+-----+ \n| text|\n+-----+\n|hello|\n|world|\n+-----+\n\n// Reading all the partitions fails\nscala> spark.read.format(\"parquet\").load(\"/user/spark/test.parquet\").show\njava.io.FileNotFoundException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1712)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getBlockLocations(NameNodeRpcServer.java:652)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.getBlockLocations(ClientNamenodeProtocolServerSideTranslatorPB.java:365)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:969)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2151)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2147)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1657)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:2145)\n\n at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method)\n at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)\n at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)\n at java.lang.reflect.Constructor.newInstance(Constructor.java:423)\n at org.apache.hadoop.ipc.RemoteException.instantiateException(RemoteException.java:106)\n at org.apache.hadoop.ipc.RemoteException.unwrapRemoteException(RemoteException.java:73)\n at org.apache.hadoop.hdfs.DFSClient.callGetBlockLocations(DFSClient.java:1242)\n at org.apache.hadoop.hdfs.DFSClient.getLocatedBlocks(DFSClient.java:1227)\n at org.apache.hadoop.hdfs.DFSClient.getBlockLocations(DFSClient.java:1285)\n at org.apache.hadoop.hdfs.DistributedFileSystem$1.doCall(DistributedFileSystem.java:221)\n at org.apache.hadoop.hdfs.DistributedFileSystem$1.doCall(DistributedFileSystem.java:217)\n at org.apache.hadoop.fs.FileSystemLinkResolver.resolve(FileSystemLinkResolver.java:81)\n at org.apache.hadoop.hdfs.DistributedFileSystem.getFileBlockLocations(DistributedFileSystem.java:228)\n at org.apache.hadoop.hdfs.DistributedFileSystem.getFileBlockLocations(DistributedFileSystem.java:209)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9$$anonfun$apply$4.apply(fileSourceInterfaces.scala:372)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9$$anonfun$apply$4.apply(fileSourceInterfaces.scala:360)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)\n at scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n at scala.collection.mutable.ArrayOps$ofRef.foreach(ArrayOps.scala:186)\n at scala.collection.TraversableLike$class.map(TraversableLike.scala:234)\n at scala.collection.mutable.ArrayOps$ofRef.map(ArrayOps.scala:186)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9.apply(fileSourceInterfaces.scala:360)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog$$anonfun$9.apply(fileSourceInterfaces.scala:348)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)\n at scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33)\n at scala.collection.mutable.WrappedArray.foreach(WrappedArray.scala:35)\n at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)\n at scala.collection.AbstractTraversable.flatMap(Traversable.scala:104)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.listLeafFiles(fileSourceInterfaces.scala:348)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.refresh(fileSourceInterfaces.scala:447)\n at org.apache.spark.sql.execution.datasources.HDFSFileCatalog.(fileSourceInterfaces.scala:291)\n at org.apache.spark.sql.execution.datasources.DataSource.resolveRelation(DataSource.scala:314)\n at org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:132)\n at org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:142)\n ... 48 elided\nCaused by: org.apache.hadoop.ipc.RemoteException: Path is not a file: /user/spark/test.parquet/id=0\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:75)\n at org.apache.hadoop.hdfs.server.namenode.INodeFile.valueOf(INodeFile.java:61)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocationsInt(FSNamesystem.java:1828)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1799)\n at org.apache.hadoop.hdfs.server.namenode.FSNamesystem.getBlockLocations(FSNamesystem.java:1712)\n at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.getBlockLocations(NameNodeRpcServer.java:652)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolServerSideTranslatorPB.getBlockLocations(ClientNamenodeProtocolServerSideTranslatorPB.java:365)\n at org.apache.hadoop.hdfs.protocol.proto.ClientNamenodeProtocolProtos$ClientNamenodeProtocol$2.callBlockingMethod(ClientNamenodeProtocolProtos.java)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Server$ProtoBufRpcInvoker.call(ProtobufRpcEngine.java:616)\n at org.apache.hadoop.ipc.RPC$Server.call(RPC.java:969)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2151)\n at org.apache.hadoop.ipc.Server$Handler$1.run(Server.java:2147)\n at java.security.AccessController.doPrivileged(Native Method)\n at javax.security.auth.Subject.doAs(Subject.java:422)\n at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1657)\n at org.apache.hadoop.ipc.Server$Handler.run(Server.java:2145)\n\n at org.apache.hadoop.ipc.Client.call(Client.java:1476)\n at org.apache.hadoop.ipc.Client.call(Client.java:1407)\n at org.apache.hadoop.ipc.ProtobufRpcEngine$Invoker.invoke(ProtobufRpcEngine.java:229)\n at com.sun.proxy.$Proxy13.getBlockLocations(Unknown Source)\n at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolTranslatorPB.getBlockLocations(ClientNamenodeProtocolTranslatorPB.java:255)\n at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)\n at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n at java.lang.reflect.Method.invoke(Method.java:498)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invokeMethod(RetryInvocationHandler.java:187)\n at org.apache.hadoop.io.retry.RetryInvocationHandler.invoke(RetryInvocationHandler.java:102)\n at com.sun.proxy.$Proxy14.getBlockLocations(Unknown Source)\n at org.apache.hadoop.hdfs.DFSClient.callGetBlockLocations(DFSClient.java:1240)\n ... 77 more\n{code}\n\n\n\n","issue_id":"12963109","key":"SPARK-14959","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-06-03T05:49:48.000+0000","role":"gold_target","summary":"​Problem Reading partitioned ORC or Parquet files"} {"case_id":"12982431","cluster":"JIRA-SPARK-7bd344e3c947","comments":[{"body":"Use the latest master, I was not able to reproduce, here is my code:\n{quote}\n val a = Seq((\"Alice\", 1)).toDF(\"name\", \"age\").describe()\n val b = Seq((\"Bob\", 2)).toDF(\"name\", \"grade\").describe()\n\n a.show()\n b.show()\n\n a.join(b, Seq(\"summary\")).show()\n{quote}\nAnything I am missing? I am using Scala 2.11, does it only happen to Scala 2.10? ","created":"2016-06-23T21:36:44.817+0000"},{"body":"Hi, [~davies] and [~bomeng].\nIf you don't mind, I'll make a PR soon.\nI resolved the problem and tested on both Scala and Python shell.","created":"2016-06-24T21:27:17.232+0000"},{"body":"Of course, with Scala 2.10.","created":"2016-06-24T21:28:13.595+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13900","created":"2016-06-24T21:40:04.294+0000"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13902","created":"2016-06-25T02:24:04.534+0000"},{"body":"Issue resolved by pull request 13902\n[https://github.com/apache/spark/pull/13902]","created":"2016-06-25T05:31:12.650+0000"}],"conversations":[{"body":"descripbe() of DataFrame use Seq() (it's a Iterator actually) to create another DataFrame, which can not be serialized in Scala 2.10.\n\n{code}\norg.apache.spark.SparkException: Task not serializable\n\tat org.apache.spark.util.ClosureCleaner$.ensureSerializable(ClosureCleaner.scala:304)\n\tat org.apache.spark.util.ClosureCleaner$.org$apache$spark$util$ClosureCleaner$$clean(ClosureCleaner.scala:294)\n\tat org.apache.spark.util.ClosureCleaner$.clean(ClosureCleaner.scala:122)\n\tat org.apache.spark.SparkContext.clean(SparkContext.scala:2060)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1.apply(RDD.scala:707)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1.apply(RDD.scala:706)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:150)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:111)\n\tat org.apache.spark.rdd.RDD.withScope(RDD.scala:316)\n\tat org.apache.spark.rdd.RDD.mapPartitions(RDD.scala:706)\n\tat org.apache.spark.sql.execution.ConvertToUnsafe.doExecute(rowFormatConverters.scala:38)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$5.apply(SparkPlan.scala:132)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$5.apply(SparkPlan.scala:130)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:150)\n\tat org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:130)\n\tat org.apache.spark.sql.execution.joins.BroadcastHashJoin$$anonfun$broadcastFuture$1$$anonfun$apply$1.apply(BroadcastHashJoin.scala:82)\n\tat org.apache.spark.sql.execution.joins.BroadcastHashJoin$$anonfun$broadcastFuture$1$$anonfun$apply$1.apply(BroadcastHashJoin.scala:79)\n\tat org.apache.spark.sql.execution.SQLExecution$.withExecutionId(SQLExecution.scala:100)\n\tat org.apache.spark.sql.execution.joins.BroadcastHashJoin$$anonfun$broadcastFuture$1.apply(BroadcastHashJoin.scala:79)\n\tat org.apache.spark.sql.execution.joins.BroadcastHashJoin$$anonfun$broadcastFuture$1.apply(BroadcastHashJoin.scala:79)\n\tat scala.concurrent.impl.Future$PromiseCompletingRunnable.liftedTree1$1(Future.scala:24)\n\tat scala.concurrent.impl.Future$PromiseCompletingRunnable.run(Future.scala:24)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: java.io.NotSerializableException: scala.collection.Iterator$$anon$11\nSerialization stack:\n\t- object not serializable (class: scala.collection.Iterator$$anon$11, value: empty iterator)\n\t- field (class: scala.collection.Iterator$$anonfun$toStream$1, name: $outer, type: interface scala.collection.Iterator)\n\t- object (class scala.collection.Iterator$$anonfun$toStream$1, )\n\t- field (class: scala.collection.immutable.Stream$Cons, name: tl, type: interface scala.Function0)\n\t- object (class scala.collection.immutable.Stream$Cons, Stream(WrappedArray(1), WrappedArray(2.0), WrappedArray(NaN), WrappedArray(2), WrappedArray(2)))\n\t- field (class: scala.collection.immutable.Stream$$anonfun$zip$1, name: $outer, type: class scala.collection.immutable.Stream)\n\t- object (class scala.collection.immutable.Stream$$anonfun$zip$1, )\n\t- field (class: scala.collection.immutable.Stream$Cons, name: tl, type: interface scala.Function0)\n\t- object (class scala.collection.immutable.Stream$Cons, Stream((WrappedArray(1),(count,)), (WrappedArray(2.0),(mean,)), (WrappedArray(NaN),(stddev,)), (WrappedArray(2),(min,)), (WrappedArray(2),(max,))))\n\t- field (class: scala.collection.immutable.Stream$$anonfun$map$1, name: $outer, type: class scala.collection.immutable.Stream)\n\t- object (class scala.collection.immutable.Stream$$anonfun$map$1, )\n\t- field (class: scala.collection.immutable.Stream$Cons, name: tl, type: interface scala.Function0)\n\t- object (class scala.collection.immutable.Stream$Cons, Stream([count,1], [mean,2.0], [stddev,NaN], [min,2], [max,2]))\n\t- field (class: scala.collection.immutable.Stream$$anonfun$map$1, name: $outer, type: class scala.collection.immutable.Stream)\n\t- object (class scala.collection.immutable.Stream$$anonfun$map$1, )\n\t- field (class: scala.collection.immutable.Stream$Cons, name: tl, type: interface scala.Function0)\n\t- object (class scala.collection.immutable.Stream$Cons, Stream([count,1], [mean,2.0], [stddev,NaN], [min,2], [max,2]))\n\t- field (class: org.apache.spark.sql.execution.LocalTableScan, name: rows, type: interface scala.collection.Seq)\n\t- object (class org.apache.spark.sql.execution.LocalTableScan, LocalTableScan [summary#633,grade#634], [[count,1],[mean,2.0],[stddev,NaN],[min,2],[max,2]]\n)\n\t- field (class: org.apache.spark.sql.execution.ConvertToUnsafe, name: child, type: class org.apache.spark.sql.execution.SparkPlan)\n\t- object (class org.apache.spark.sql.execution.ConvertToUnsafe, ConvertToUnsafe\n+- LocalTableScan [summary#633,grade#634], [[count,1],[mean,2.0],[stddev,NaN],[min,2],[max,2]]\n)\n\t- field (class: org.apache.spark.sql.execution.ConvertToUnsafe$$anonfun$1, name: $outer, type: class org.apache.spark.sql.execution.ConvertToUnsafe)\n\t- object (class org.apache.spark.sql.execution.ConvertToUnsafe$$anonfun$1, )\n\tat org.apache.spark.serializer.SerializationDebugger$.improveException(SerializationDebugger.scala:40)\n\tat org.apache.spark.serializer.JavaSerializationStream.writeObject(JavaSerializer.scala:47)\n\tat org.apache.spark.serializer.JavaSerializerInstance.serialize(JavaSerializer.scala:101)\n\tat org.apache.spark.util.ClosureCleaner$.ensureSerializable(ClosureCleaner.scala:301)\n\t... 24 more\n\n{code}\n\nhttps://databricks-prod-cloudfront.cloud.databricks.com/public/4027ec902e239c93eaaa8714f173bcfc/1953143414968711/3514600112485120/2850462372213371/latest.html","from":"reporter","subject":"Can't join describe() of DataFrame in Scala 2.10"},{"body":"Use the latest master, I was not able to reproduce, here is my code:\n{quote}\n val a = Seq((\"Alice\", 1)).toDF(\"name\", \"age\").describe()\n val b = Seq((\"Bob\", 2)).toDF(\"name\", \"grade\").describe()\n\n a.show()\n b.show()\n\n a.join(b, Seq(\"summary\")).show()\n{quote}\nAnything I am missing? I am using Scala 2.11, does it only happen to Scala 2.10? ","from":"developer"},{"body":"Hi, [~davies] and [~bomeng].\nIf you don't mind, I'll make a PR soon.\nI resolved the problem and tested on both Scala and Python shell.","from":"developer"},{"body":"Of course, with Scala 2.10.","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13900","from":"developer"},{"body":"User 'dongjoon-hyun' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/13902","from":"developer"},{"body":"Issue resolved by pull request 13902\n[https://github.com/apache/spark/pull/13902]","from":"developer"}],"created":"2016-06-23T19:07:32.000+0000","description":"descripbe() of DataFrame use Seq() (it's a Iterator actually) to create another DataFrame, which can not be serialized in Scala 2.10.\n\n{code}\norg.apache.spark.SparkException: Task not serializable\n\tat org.apache.spark.util.ClosureCleaner$.ensureSerializable(ClosureCleaner.scala:304)\n\tat org.apache.spark.util.ClosureCleaner$.org$apache$spark$util$ClosureCleaner$$clean(ClosureCleaner.scala:294)\n\tat org.apache.spark.util.ClosureCleaner$.clean(ClosureCleaner.scala:122)\n\tat org.apache.spark.SparkContext.clean(SparkContext.scala:2060)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1.apply(RDD.scala:707)\n\tat org.apache.spark.rdd.RDD$$anonfun$mapPartitions$1.apply(RDD.scala:706)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:150)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:111)\n\tat org.apache.spark.rdd.RDD.withScope(RDD.scala:316)\n\tat org.apache.spark.rdd.RDD.mapPartitions(RDD.scala:706)\n\tat org.apache.spark.sql.execution.ConvertToUnsafe.doExecute(rowFormatConverters.scala:38)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$5.apply(SparkPlan.scala:132)\n\tat org.apache.spark.sql.execution.SparkPlan$$anonfun$execute$5.apply(SparkPlan.scala:130)\n\tat org.apache.spark.rdd.RDDOperationScope$.withScope(RDDOperationScope.scala:150)\n\tat org.apache.spark.sql.execution.SparkPlan.execute(SparkPlan.scala:130)\n\tat org.apache.spark.sql.execution.joins.BroadcastHashJoin$$anonfun$broadcastFuture$1$$anonfun$apply$1.apply(BroadcastHashJoin.scala:82)\n\tat org.apache.spark.sql.execution.joins.BroadcastHashJoin$$anonfun$broadcastFuture$1$$anonfun$apply$1.apply(BroadcastHashJoin.scala:79)\n\tat org.apache.spark.sql.execution.SQLExecution$.withExecutionId(SQLExecution.scala:100)\n\tat org.apache.spark.sql.execution.joins.BroadcastHashJoin$$anonfun$broadcastFuture$1.apply(BroadcastHashJoin.scala:79)\n\tat org.apache.spark.sql.execution.joins.BroadcastHashJoin$$anonfun$broadcastFuture$1.apply(BroadcastHashJoin.scala:79)\n\tat scala.concurrent.impl.Future$PromiseCompletingRunnable.liftedTree1$1(Future.scala:24)\n\tat scala.concurrent.impl.Future$PromiseCompletingRunnable.run(Future.scala:24)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:745)\nCaused by: java.io.NotSerializableException: scala.collection.Iterator$$anon$11\nSerialization stack:\n\t- object not serializable (class: scala.collection.Iterator$$anon$11, value: empty iterator)\n\t- field (class: scala.collection.Iterator$$anonfun$toStream$1, name: $outer, type: interface scala.collection.Iterator)\n\t- object (class scala.collection.Iterator$$anonfun$toStream$1, )\n\t- field (class: scala.collection.immutable.Stream$Cons, name: tl, type: interface scala.Function0)\n\t- object (class scala.collection.immutable.Stream$Cons, Stream(WrappedArray(1), WrappedArray(2.0), WrappedArray(NaN), WrappedArray(2), WrappedArray(2)))\n\t- field (class: scala.collection.immutable.Stream$$anonfun$zip$1, name: $outer, type: class scala.collection.immutable.Stream)\n\t- object (class scala.collection.immutable.Stream$$anonfun$zip$1, )\n\t- field (class: scala.collection.immutable.Stream$Cons, name: tl, type: interface scala.Function0)\n\t- object (class scala.collection.immutable.Stream$Cons, Stream((WrappedArray(1),(count,)), (WrappedArray(2.0),(mean,)), (WrappedArray(NaN),(stddev,)), (WrappedArray(2),(min,)), (WrappedArray(2),(max,))))\n\t- field (class: scala.collection.immutable.Stream$$anonfun$map$1, name: $outer, type: class scala.collection.immutable.Stream)\n\t- object (class scala.collection.immutable.Stream$$anonfun$map$1, )\n\t- field (class: scala.collection.immutable.Stream$Cons, name: tl, type: interface scala.Function0)\n\t- object (class scala.collection.immutable.Stream$Cons, Stream([count,1], [mean,2.0], [stddev,NaN], [min,2], [max,2]))\n\t- field (class: scala.collection.immutable.Stream$$anonfun$map$1, name: $outer, type: class scala.collection.immutable.Stream)\n\t- object (class scala.collection.immutable.Stream$$anonfun$map$1, )\n\t- field (class: scala.collection.immutable.Stream$Cons, name: tl, type: interface scala.Function0)\n\t- object (class scala.collection.immutable.Stream$Cons, Stream([count,1], [mean,2.0], [stddev,NaN], [min,2], [max,2]))\n\t- field (class: org.apache.spark.sql.execution.LocalTableScan, name: rows, type: interface scala.collection.Seq)\n\t- object (class org.apache.spark.sql.execution.LocalTableScan, LocalTableScan [summary#633,grade#634], [[count,1],[mean,2.0],[stddev,NaN],[min,2],[max,2]]\n)\n\t- field (class: org.apache.spark.sql.execution.ConvertToUnsafe, name: child, type: class org.apache.spark.sql.execution.SparkPlan)\n\t- object (class org.apache.spark.sql.execution.ConvertToUnsafe, ConvertToUnsafe\n+- LocalTableScan [summary#633,grade#634], [[count,1],[mean,2.0],[stddev,NaN],[min,2],[max,2]]\n)\n\t- field (class: org.apache.spark.sql.execution.ConvertToUnsafe$$anonfun$1, name: $outer, type: class org.apache.spark.sql.execution.ConvertToUnsafe)\n\t- object (class org.apache.spark.sql.execution.ConvertToUnsafe$$anonfun$1, )\n\tat org.apache.spark.serializer.SerializationDebugger$.improveException(SerializationDebugger.scala:40)\n\tat org.apache.spark.serializer.JavaSerializationStream.writeObject(JavaSerializer.scala:47)\n\tat org.apache.spark.serializer.JavaSerializerInstance.serialize(JavaSerializer.scala:101)\n\tat org.apache.spark.util.ClosureCleaner$.ensureSerializable(ClosureCleaner.scala:301)\n\t... 24 more\n\n{code}\n\nhttps://databricks-prod-cloudfront.cloud.databricks.com/public/4027ec902e239c93eaaa8714f173bcfc/1953143414968711/3514600112485120/2850462372213371/latest.html","issue_id":"12982431","key":"SPARK-16173","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2016-06-25T05:31:12.000+0000","role":"gold_target","summary":"Can't join describe() of DataFrame in Scala 2.10"} {"case_id":"13025335","cluster":"JIRA-SPARK-dcb981be150c","comments":[{"body":"you mean a query below and you'd like to load the second line only?\n{code}\n>> test.csv <<\n1 0,2014-xx-xx\n2 1,2014-01-01 \n\nscala> import org.apache.spark.sql.types._\nscala> val schema = new StructType().add(\"a\", IntegerType).add(\"b\", DateType)\nscala> spark.read.format(\"csv\").schema(schema).load(\"test.csv\").show\n\n16/12/04 21:21:56 ERROR Executor: Exception in task 0.0 in stage 32.0 (TID 32)\njava.lang.NumberFormatException: For input string: \"xx\"\n at java.lang.NumberFormatException.forInputString(NumberFormatException.java:65)\n at java.lang.Integer.parseInt(Integer.java:580)\n at java.lang.Integer.parseInt(Integer.java:615)\n at java.sql.Date.valueOf(Date.java:134)\n at org.apache.spark.sql.catalyst.util.DateTimeUtils$.stringToTime(DateTimeUtils.scala:137)\n at org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$9.apply$mcI$sp(CSVInferSchema.scala:290)\n at org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$8.apply(CSVInferSchema.scala:290)\n at org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$8.apply(CSVInferSchema.scala:290)\n at scala.util.Try.getOrElse(Try.scala:79)\n at org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$.castTo(CSVInferSchema.scala:287)\n at org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:121)\n at org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:90)\n at org.apache.spark.sql.execution.datasources.csv.CSVFileFormat$$anonfun$buildReader$1$$anonfun$apply$2.apply(CSVFileFormat.scala:173)\n at org.apache.spark.sql.execution.datasources.csv.CSVFileFormat$$anonfun$buildReader$1$$anonfun$apply$2.apply(CSVFileFormat.scala:172)\n{code}","created":"2016-12-04T12:26:26.955+0000"},{"body":"Yes, my understanding was that it should put nullify the value if it fails to parse it in PERMISSIVE mode or drop the whole row (line) in DROPMALFORMED as described in the docs: http://spark.apache.org/docs/latest/api/scala/index.html#org.apache.spark.sql.DataFrameReader, i.e.: \n* mode (default PERMISSIVE): allows a mode for dealing with corrupt records during parsing.\n** PERMISSIVE : sets other fields to null when it meets a corrupted record. When a schema is set by user, it sets null for extra fields.\n** DROPMALFORMED : ignores the whole corrupted records.\n** FAILFAST : throws an exception when it meets corrupted records.","created":"2016-12-04T12:43:09.618+0000"},{"body":"`DROPMALFORMED` works well in this query though, `PERMISSIVE` still throws the exception. But, `PERMISSIVE` mode just fills null in case that the number of fields spark parses is lower than a given schema length (IIUC, in the spark doc., this case is called `corrupted records`). Since the current implementation does not fill null for malformed fields (the format errors above, or something), it seems this is an expected behaviour. cc: [~hyukjin.kwon]","created":"2016-12-04T14:02:51.345+0000"},{"body":"Additionally, in our basic stance, it seems this csv format keeps almost the same behavior with `com.databricks.spark.csv`. I quickly checked that the databricks one throws an exception in this query;\n{code}\nscala> sqlContext.read.format(\"com.databricks.spark.csv\").schema(schema).option(\"mode\", \"PERMISSIVE\").load(\"../test.csv\").show\n16/12/04 23:17:43 ERROR Executor: Exception in task 0.0 in stage 0.0 (TID 0)\njava.lang.NumberFormatException: For input string: \"xx\"\n at java.lang.NumberFormatException.forInputString(NumberFormatException.java:65)\n at java.lang.Integer.parseInt(Integer.java:580)\n at java.lang.Integer.parseInt(Integer.java:615)\n at java.sql.Date.valueOf(Date.java:134)\n at com.databricks.spark.csv.util.TypeCast$.castTo(TypeCast.scala:74)\n at com.databricks.spark.csv.CsvRelation$$anonfun$buildScan$2.apply(CsvRelation.scala:121)\n at com.databricks.spark.csv.CsvRelation$$anonfun$buildScan$2.apply(CsvRelation.scala:108)\n ...\n{code}","created":"2016-12-04T14:30:58.544+0000"},{"body":"Anyway, we can easily fix this like this: https://github.com/apache/spark/compare/master...maropu:SPARK-18699\n\n{code}\nscala> spark.read.format(\"csv\").schema(schema).option(\"mode\", \"PERMISSIVE\").load(\"test.csv\").show\n16/12/04 23:46:55 WARN CSVRelation: Fill NULL in a field because a malformed token detected: 2014-xx-xx\n+---+----------+\n| a| b|\n+---+----------+\n| 0| null|\n| 1|2014-01-01|\n+---+----------+\n{code}","created":"2016-12-04T14:55:13.643+0000"},{"body":"Thank you for cc'ing me. Yup, I noticed this but I decided to just leave as is because I noticed that this is not a regression comparing to the external CSV library as you said below.\n\nI persuaded myself and was thinking that the parse mode is related with parsing itself. JSON has some standard JSON types so IIRC JSON loads the data as null in {{PERMISSIVE}} mode in this case but parsing CSV is not related with types not within plain texts so it throws an exception.\n\nHowever, if this sounds odds in practice, I agree with fixing this although I still am worried of changing the exiting behaviour.\n\nCould I ask what you think [~falaki] please?","created":"2016-12-04T15:34:11.819+0000"},{"body":"While I don't argue that some other packages have similar behaviour, I think the PERMISSIVE mode should be, well, as permissive as possible, since CSVs have very little standards and no types. In ma case I had just one odd value in almost 1 TB set and the job crushed at the very end after about an hour. To go around the issue one needs to manually parse each line, which is not the end of the world, but I wanted to use CSV reader exactly for the confidence of not writing extra code. IMO the mode for error detection should be FAILFAST. Moreover, if I really need to check the data, I read it differently anyway.\nBTW thanks for looking into this.","created":"2016-12-04T22:24:58.410+0000"},{"body":"yea, I'm also working on large csv files now and, certainly, I think this current behavior makes it difficult to find incorrect records in them. Logging with meaningful waning messages (e.g., including the incorrect records and line numbers) helps much to me.","created":"2016-12-09T04:55:27.783+0000"},{"body":"This issue is also seen in cases where a timestamp field (among many) is empty (just \" \"). If I use permissive mode, then any instances where one of those timestamp fields are empty automatically results in the entire row getting dropped. In my case that's over 90% of the rows.","created":"2016-12-14T17:13:23.639+0000"},{"body":"BTW, maybe you could try to set {{nullValue}} to {{\" \"}} or set {{ignoreLeadingWhiteSpace}} and {{ignoreTrailingWhiteSpace}} to {{true}} for now.","created":"2016-12-15T01:37:20.431+0000"},{"body":"Thanks for the reply! Unfortunately neither of those options work in my case.\n","created":"2016-12-15T02:33:58.780+0000"},{"body":"Fixed in PR https://github.com/apache/spark/pull/16319\n\n(not merged into the tree as of now).\n\nEnjoy","created":"2016-12-17T03:05:24.493+0000"},{"body":"User 'kubatyszko' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16319","created":"2016-12-17T03:07:05.822+0000"},{"body":"This fix makes some sensible to me and what do you think? cc: [~hyukjin.kwon]\nIf yes, I try to make a pr based on this (https://github.com/apache/spark/compare/master...maropu:SPARK-18699). If no, I think it's okay to set \"Won't Fix\".","created":"2017-02-13T02:25:58.762+0000"},{"body":"Thanks for cc'ing me. For me, it is reasonable to me too (although I am not supposed to decide what to add into Spark). Maybe, it might be even nicer if we could have some references such as similar libraries e.g., read.csv in R if applicable and/or good explanations for potential use cases.","created":"2017-02-13T10:18:20.461+0000"},{"body":"User 'maropu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16928","created":"2017-02-14T15:29:04.599+0000"},{"body":"Issue resolved by pull request 16928\n[https://github.com/apache/spark/pull/16928]","created":"2017-02-23T20:09:50.553+0000"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17142","created":"2017-03-02T17:48:03.408+0000"}],"conversations":[{"body":"If CSV is read and the schema contains any other type than String, exception is thrown when the string value in CSV is malformed; e.g. if the timestamp does not match the defined one, an exception is thrown:\n{code}\nCaused by: java.lang.IllegalArgumentException\n\tat java.sql.Date.valueOf(Date.java:143)\n\tat org.apache.spark.sql.catalyst.util.DateTimeUtils$.stringToTime(DateTimeUtils.scala:137)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$6.apply$mcJ$sp(CSVInferSchema.scala:272)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$6.apply(CSVInferSchema.scala:272)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$6.apply(CSVInferSchema.scala:272)\n\tat scala.util.Try.getOrElse(Try.scala:79)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$.castTo(CSVInferSchema.scala:269)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:116)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:85)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVFileFormat$$anonfun$buildReader$1$$anonfun$apply$2.apply(CSVFileFormat.scala:128)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVFileFormat$$anonfun$buildReader$1$$anonfun$apply$2.apply(CSVFileFormat.scala:127)\n\tat scala.collection.Iterator$$anon$12.nextCur(Iterator.scala:434)\n\tat scala.collection.Iterator$$anon$12.hasNext(Iterator.scala:440)\n\tat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n\tat org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:91)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n\tat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\n\tat org.apache.spark.sql.execution.datasources.DefaultWriterContainer$$anonfun$writeRows$1.apply$mcV$sp(WriterContainer.scala:253)\n\tat org.apache.spark.sql.execution.datasources.DefaultWriterContainer$$anonfun$writeRows$1.apply(WriterContainer.scala:252)\n\tat org.apache.spark.sql.execution.datasources.DefaultWriterContainer$$anonfun$writeRows$1.apply(WriterContainer.scala:252)\n\tat org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1348)\n\tat org.apache.spark.sql.execution.datasources.DefaultWriterContainer.writeRows(WriterContainer.scala:258)\n\t... 8 more\n{code}\n\nIt behaves similarly with Integer and Long types, from what I've seen.\n\nTo my understanding modes PERMISSIVE and DROPMALFORMED should just null the value or drop the line, but instead they kill the job.","from":"reporter","subject":"Spark CSV parsing types other than String throws exception when malformed"},{"body":"you mean a query below and you'd like to load the second line only?\n{code}\n>> test.csv <<\n1 0,2014-xx-xx\n2 1,2014-01-01 \n\nscala> import org.apache.spark.sql.types._\nscala> val schema = new StructType().add(\"a\", IntegerType).add(\"b\", DateType)\nscala> spark.read.format(\"csv\").schema(schema).load(\"test.csv\").show\n\n16/12/04 21:21:56 ERROR Executor: Exception in task 0.0 in stage 32.0 (TID 32)\njava.lang.NumberFormatException: For input string: \"xx\"\n at java.lang.NumberFormatException.forInputString(NumberFormatException.java:65)\n at java.lang.Integer.parseInt(Integer.java:580)\n at java.lang.Integer.parseInt(Integer.java:615)\n at java.sql.Date.valueOf(Date.java:134)\n at org.apache.spark.sql.catalyst.util.DateTimeUtils$.stringToTime(DateTimeUtils.scala:137)\n at org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$9.apply$mcI$sp(CSVInferSchema.scala:290)\n at org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$8.apply(CSVInferSchema.scala:290)\n at org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$8.apply(CSVInferSchema.scala:290)\n at scala.util.Try.getOrElse(Try.scala:79)\n at org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$.castTo(CSVInferSchema.scala:287)\n at org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:121)\n at org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:90)\n at org.apache.spark.sql.execution.datasources.csv.CSVFileFormat$$anonfun$buildReader$1$$anonfun$apply$2.apply(CSVFileFormat.scala:173)\n at org.apache.spark.sql.execution.datasources.csv.CSVFileFormat$$anonfun$buildReader$1$$anonfun$apply$2.apply(CSVFileFormat.scala:172)\n{code}","from":"developer"},{"body":"Yes, my understanding was that it should put nullify the value if it fails to parse it in PERMISSIVE mode or drop the whole row (line) in DROPMALFORMED as described in the docs: http://spark.apache.org/docs/latest/api/scala/index.html#org.apache.spark.sql.DataFrameReader, i.e.: \n* mode (default PERMISSIVE): allows a mode for dealing with corrupt records during parsing.\n** PERMISSIVE : sets other fields to null when it meets a corrupted record. When a schema is set by user, it sets null for extra fields.\n** DROPMALFORMED : ignores the whole corrupted records.\n** FAILFAST : throws an exception when it meets corrupted records.","from":"developer"},{"body":"`DROPMALFORMED` works well in this query though, `PERMISSIVE` still throws the exception. But, `PERMISSIVE` mode just fills null in case that the number of fields spark parses is lower than a given schema length (IIUC, in the spark doc., this case is called `corrupted records`). Since the current implementation does not fill null for malformed fields (the format errors above, or something), it seems this is an expected behaviour. cc: [~hyukjin.kwon]","from":"developer"},{"body":"Additionally, in our basic stance, it seems this csv format keeps almost the same behavior with `com.databricks.spark.csv`. I quickly checked that the databricks one throws an exception in this query;\n{code}\nscala> sqlContext.read.format(\"com.databricks.spark.csv\").schema(schema).option(\"mode\", \"PERMISSIVE\").load(\"../test.csv\").show\n16/12/04 23:17:43 ERROR Executor: Exception in task 0.0 in stage 0.0 (TID 0)\njava.lang.NumberFormatException: For input string: \"xx\"\n at java.lang.NumberFormatException.forInputString(NumberFormatException.java:65)\n at java.lang.Integer.parseInt(Integer.java:580)\n at java.lang.Integer.parseInt(Integer.java:615)\n at java.sql.Date.valueOf(Date.java:134)\n at com.databricks.spark.csv.util.TypeCast$.castTo(TypeCast.scala:74)\n at com.databricks.spark.csv.CsvRelation$$anonfun$buildScan$2.apply(CsvRelation.scala:121)\n at com.databricks.spark.csv.CsvRelation$$anonfun$buildScan$2.apply(CsvRelation.scala:108)\n ...\n{code}","from":"developer"},{"body":"Anyway, we can easily fix this like this: https://github.com/apache/spark/compare/master...maropu:SPARK-18699\n\n{code}\nscala> spark.read.format(\"csv\").schema(schema).option(\"mode\", \"PERMISSIVE\").load(\"test.csv\").show\n16/12/04 23:46:55 WARN CSVRelation: Fill NULL in a field because a malformed token detected: 2014-xx-xx\n+---+----------+\n| a| b|\n+---+----------+\n| 0| null|\n| 1|2014-01-01|\n+---+----------+\n{code}","from":"developer"},{"body":"Thank you for cc'ing me. Yup, I noticed this but I decided to just leave as is because I noticed that this is not a regression comparing to the external CSV library as you said below.\n\nI persuaded myself and was thinking that the parse mode is related with parsing itself. JSON has some standard JSON types so IIRC JSON loads the data as null in {{PERMISSIVE}} mode in this case but parsing CSV is not related with types not within plain texts so it throws an exception.\n\nHowever, if this sounds odds in practice, I agree with fixing this although I still am worried of changing the exiting behaviour.\n\nCould I ask what you think [~falaki] please?","from":"developer"},{"body":"While I don't argue that some other packages have similar behaviour, I think the PERMISSIVE mode should be, well, as permissive as possible, since CSVs have very little standards and no types. In ma case I had just one odd value in almost 1 TB set and the job crushed at the very end after about an hour. To go around the issue one needs to manually parse each line, which is not the end of the world, but I wanted to use CSV reader exactly for the confidence of not writing extra code. IMO the mode for error detection should be FAILFAST. Moreover, if I really need to check the data, I read it differently anyway.\nBTW thanks for looking into this.","from":"developer"},{"body":"yea, I'm also working on large csv files now and, certainly, I think this current behavior makes it difficult to find incorrect records in them. Logging with meaningful waning messages (e.g., including the incorrect records and line numbers) helps much to me.","from":"developer"},{"body":"This issue is also seen in cases where a timestamp field (among many) is empty (just \" \"). If I use permissive mode, then any instances where one of those timestamp fields are empty automatically results in the entire row getting dropped. In my case that's over 90% of the rows.","from":"developer"},{"body":"BTW, maybe you could try to set {{nullValue}} to {{\" \"}} or set {{ignoreLeadingWhiteSpace}} and {{ignoreTrailingWhiteSpace}} to {{true}} for now.","from":"developer"},{"body":"Thanks for the reply! Unfortunately neither of those options work in my case.\n","from":"developer"},{"body":"Fixed in PR https://github.com/apache/spark/pull/16319\n\n(not merged into the tree as of now).\n\nEnjoy","from":"developer"},{"body":"User 'kubatyszko' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16319","from":"developer"},{"body":"This fix makes some sensible to me and what do you think? cc: [~hyukjin.kwon]\nIf yes, I try to make a pr based on this (https://github.com/apache/spark/compare/master...maropu:SPARK-18699). If no, I think it's okay to set \"Won't Fix\".","from":"developer"},{"body":"Thanks for cc'ing me. For me, it is reasonable to me too (although I am not supposed to decide what to add into Spark). Maybe, it might be even nicer if we could have some references such as similar libraries e.g., read.csv in R if applicable and/or good explanations for potential use cases.","from":"developer"},{"body":"User 'maropu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16928","from":"developer"},{"body":"Issue resolved by pull request 16928\n[https://github.com/apache/spark/pull/16928]","from":"developer"},{"body":"User 'HyukjinKwon' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/17142","from":"developer"}],"created":"2016-12-03T16:33:28.000+0000","description":"If CSV is read and the schema contains any other type than String, exception is thrown when the string value in CSV is malformed; e.g. if the timestamp does not match the defined one, an exception is thrown:\n{code}\nCaused by: java.lang.IllegalArgumentException\n\tat java.sql.Date.valueOf(Date.java:143)\n\tat org.apache.spark.sql.catalyst.util.DateTimeUtils$.stringToTime(DateTimeUtils.scala:137)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$6.apply$mcJ$sp(CSVInferSchema.scala:272)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$6.apply(CSVInferSchema.scala:272)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$$anonfun$castTo$6.apply(CSVInferSchema.scala:272)\n\tat scala.util.Try.getOrElse(Try.scala:79)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVTypeCast$.castTo(CSVInferSchema.scala:269)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:116)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVRelation$$anonfun$csvParser$3.apply(CSVRelation.scala:85)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVFileFormat$$anonfun$buildReader$1$$anonfun$apply$2.apply(CSVFileFormat.scala:128)\n\tat org.apache.spark.sql.execution.datasources.csv.CSVFileFormat$$anonfun$buildReader$1$$anonfun$apply$2.apply(CSVFileFormat.scala:127)\n\tat scala.collection.Iterator$$anon$12.nextCur(Iterator.scala:434)\n\tat scala.collection.Iterator$$anon$12.hasNext(Iterator.scala:440)\n\tat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n\tat org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:91)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n\tat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:370)\n\tat org.apache.spark.sql.execution.datasources.DefaultWriterContainer$$anonfun$writeRows$1.apply$mcV$sp(WriterContainer.scala:253)\n\tat org.apache.spark.sql.execution.datasources.DefaultWriterContainer$$anonfun$writeRows$1.apply(WriterContainer.scala:252)\n\tat org.apache.spark.sql.execution.datasources.DefaultWriterContainer$$anonfun$writeRows$1.apply(WriterContainer.scala:252)\n\tat org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1348)\n\tat org.apache.spark.sql.execution.datasources.DefaultWriterContainer.writeRows(WriterContainer.scala:258)\n\t... 8 more\n{code}\n\nIt behaves similarly with Integer and Long types, from what I've seen.\n\nTo my understanding modes PERMISSIVE and DROPMALFORMED should just null the value or drop the line, but instead they kill the job.","issue_id":"13025335","key":"SPARK-18699","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-02-23T20:09:50.000+0000","role":"gold_target","summary":"Spark CSV parsing types other than String throws exception when malformed"} {"case_id":"13035951","cluster":"JIRA-SPARK-fdca525f0b01","comments":[{"body":"I haven't been successful in creating a test case to reproduce this. In my attempts, I see retries sometimes failing to fetch block *status*, which happens in the initialization of the {{ShuffleBlockFetcherIterator}}. since that is outside user code, it does get handled correctly, and the shuffle data is regenerated. But, the problem is clearly there, and I've seen it effect real users.\n\nopen to ideas for the test case & reproduction.","created":"2017-01-18T18:28:45.408+0000"},{"body":"User 'squito' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16639","created":"2017-01-19T06:32:05.852+0000"},{"body":"This all makes sense, and the PR is a good effort to fix this kind of \"accidental\" swallowing of FetchFailedException. I guess my only real question is if we should allow for the possibility of a FetchFailedException not only being caught, but also the failure being remedied by some means other than the usual handling in the driver. I'm not sure exactly how or why that kind of \"fix the fetch failure before Spark tries to handle it\" would be done; and it would seem that something like that would be prone to subtle errors, so maybe we should just set it in stone that nobody but the driver should try to fix a fetch failure -- which would make the approach of your \"guarantee that the FetchFailedException is seen by the driver\" PR completely correct.","created":"2017-01-19T18:40:40.582+0000"},{"body":"[~markhamstra]\nbq. I guess my only real question is if we should allow for the possibility of a FetchFailedException not only being caught, but also the failure being remedied by some means other than the usual handling in the driver.\n\nI wondered about this too -- in fact, in the end, the pr *does* allow that. You'll notice that if the fetch failure is set in the task context, but the task succeeds, than I log an error but otherwise let the task continue. We only send it back to the driver if there is also a task failure.\n\nI decided to do it that way in case there is some crazy existing code out there which relies on this behavior ... but honestly I'm not sure that is the right decision, maybe it should still send the fetch failure back to the driver and fail the task in that case too.","created":"2017-01-19T19:48:55.943+0000"},{"body":"Ok, I haven't read your PR closely yet, so I missed that.\n\nThis question looks like something that could use more eyes and insights. [~kayousterhout][~matei][~rxin@databricks.com]\n","created":"2017-01-19T20:09:12.819+0000"}],"conversations":[{"body":"The scheduler handles node failures by looking for a special {{FetchFailedException}} thrown by the shuffle block fetcher. This is handled in {{Executor}} and then passed as a special msg back to the driver: https://github.com/apache/spark/blob/278fa1eb305220a85c816c948932d6af8fa619aa/core/src/main/scala/org/apache/spark/executor/Executor.scala#L403\n\nHowever, user code exists in between the shuffle block fetcher and that catch block -- it could intercept the exception, wrap it with something else, and throw a different exception. If that happens, spark treats it as an ordinary task failure, and retries the task, rather than regenerating the missing shuffle data. The task eventually is retried 4 times, its doomed to fail each time, and the job is failed.\n\nYou might think that no user code should do that -- but even sparksql does it:\nhttps://github.com/apache/spark/blob/278fa1eb305220a85c816c948932d6af8fa619aa/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/FileFormatWriter.scala#L214\n\nHere's an example stack trace. This is from Spark 1.6, so the sql code is not the same, but the problem is still there:\n\n{noformat}\n17/01/13 19:18:02 WARN scheduler.TaskSetManager: Lost task 0.0 in stage 1983.0 (TID 304851, xxx): org.apache.spark.SparkException: Task failed while writing rows.\n at org.apache.spark.sql.execution.datasources.DynamicPartitionWriterContainer.writeRows(WriterContainer.scala:414)\n at org.apache.spark.sql.execution.datasources.InsertIntoHadoopFsRelation$$anonfun$run$1$$anonfun$apply$mcV$sp$3.apply(InsertIntoHadoopFsRelation.scala:150)\n at org.apache.spark.sql.execution.datasources.InsertIntoHadoopFsRelation$$anonfun$run$1$$anonfun$apply$mcV$sp$3.apply(InsertIntoHadoopFsRelation.scala:150)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.spark.shuffle.FetchFailedException: Failed to connect to xxx/yyy:zzz\n at org.apache.spark.storage.ShuffleBlockFetcherIterator.throwFetchFailedException(ShuffleBlockFetcherIterator.scala:323)\n...\n17/01/13 19:19:29 ERROR scheduler.TaskSetManager: Task 0 in stage 1983.0 failed 4 times; aborting job\n{noformat}\n\nI think the right fix here is to also set a fetch failure status in the {{TaskContextImpl}}, so the executor can check that instead of just one exception.","from":"reporter","subject":"FetchFailures can be hidden by user (or sql) exception handling"},{"body":"I haven't been successful in creating a test case to reproduce this. In my attempts, I see retries sometimes failing to fetch block *status*, which happens in the initialization of the {{ShuffleBlockFetcherIterator}}. since that is outside user code, it does get handled correctly, and the shuffle data is regenerated. But, the problem is clearly there, and I've seen it effect real users.\n\nopen to ideas for the test case & reproduction.","from":"developer"},{"body":"User 'squito' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/16639","from":"developer"},{"body":"This all makes sense, and the PR is a good effort to fix this kind of \"accidental\" swallowing of FetchFailedException. I guess my only real question is if we should allow for the possibility of a FetchFailedException not only being caught, but also the failure being remedied by some means other than the usual handling in the driver. I'm not sure exactly how or why that kind of \"fix the fetch failure before Spark tries to handle it\" would be done; and it would seem that something like that would be prone to subtle errors, so maybe we should just set it in stone that nobody but the driver should try to fix a fetch failure -- which would make the approach of your \"guarantee that the FetchFailedException is seen by the driver\" PR completely correct.","from":"developer"},{"body":"[~markhamstra]\nbq. I guess my only real question is if we should allow for the possibility of a FetchFailedException not only being caught, but also the failure being remedied by some means other than the usual handling in the driver.\n\nI wondered about this too -- in fact, in the end, the pr *does* allow that. You'll notice that if the fetch failure is set in the task context, but the task succeeds, than I log an error but otherwise let the task continue. We only send it back to the driver if there is also a task failure.\n\nI decided to do it that way in case there is some crazy existing code out there which relies on this behavior ... but honestly I'm not sure that is the right decision, maybe it should still send the fetch failure back to the driver and fail the task in that case too.","from":"developer"},{"body":"Ok, I haven't read your PR closely yet, so I missed that.\n\nThis question looks like something that could use more eyes and insights. [~kayousterhout][~matei][~rxin@databricks.com]\n","from":"developer"}],"created":"2017-01-18T18:22:01.000+0000","description":"The scheduler handles node failures by looking for a special {{FetchFailedException}} thrown by the shuffle block fetcher. This is handled in {{Executor}} and then passed as a special msg back to the driver: https://github.com/apache/spark/blob/278fa1eb305220a85c816c948932d6af8fa619aa/core/src/main/scala/org/apache/spark/executor/Executor.scala#L403\n\nHowever, user code exists in between the shuffle block fetcher and that catch block -- it could intercept the exception, wrap it with something else, and throw a different exception. If that happens, spark treats it as an ordinary task failure, and retries the task, rather than regenerating the missing shuffle data. The task eventually is retried 4 times, its doomed to fail each time, and the job is failed.\n\nYou might think that no user code should do that -- but even sparksql does it:\nhttps://github.com/apache/spark/blob/278fa1eb305220a85c816c948932d6af8fa619aa/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/FileFormatWriter.scala#L214\n\nHere's an example stack trace. This is from Spark 1.6, so the sql code is not the same, but the problem is still there:\n\n{noformat}\n17/01/13 19:18:02 WARN scheduler.TaskSetManager: Lost task 0.0 in stage 1983.0 (TID 304851, xxx): org.apache.spark.SparkException: Task failed while writing rows.\n at org.apache.spark.sql.execution.datasources.DynamicPartitionWriterContainer.writeRows(WriterContainer.scala:414)\n at org.apache.spark.sql.execution.datasources.InsertIntoHadoopFsRelation$$anonfun$run$1$$anonfun$apply$mcV$sp$3.apply(InsertIntoHadoopFsRelation.scala:150)\n at org.apache.spark.sql.execution.datasources.InsertIntoHadoopFsRelation$$anonfun$run$1$$anonfun$apply$mcV$sp$3.apply(InsertIntoHadoopFsRelation.scala:150)\n at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:66)\n at org.apache.spark.scheduler.Task.run(Task.scala:89)\n at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)\n at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n at java.lang.Thread.run(Thread.java:745)\nCaused by: org.apache.spark.shuffle.FetchFailedException: Failed to connect to xxx/yyy:zzz\n at org.apache.spark.storage.ShuffleBlockFetcherIterator.throwFetchFailedException(ShuffleBlockFetcherIterator.scala:323)\n...\n17/01/13 19:19:29 ERROR scheduler.TaskSetManager: Task 0 in stage 1983.0 failed 4 times; aborting job\n{noformat}\n\nI think the right fix here is to also set a fetch failure status in the {{TaskContextImpl}}, so the executor can check that instead of just one exception.","issue_id":"13035951","key":"SPARK-19276","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-03-03T00:48:09.000+0000","role":"gold_target","summary":"FetchFailures can be hidden by user (or sql) exception handling"} {"case_id":"13099575","cluster":"JIRA-SPARK-aeb0851a7b1c","comments":[{"body":"Note that UnsafeExternalSorter.spill appears twice on the stack trace, so it's nested spilling: the first triggered spilling triggers another spilling through UnsafeInMemorySorter.reset.\n\nPossibly it's messing up something by nested-spilling itself twice?\nOr messing something with\n{code:java}\n if (trigger != this) {\n if (readingIterator != null) {\n return readingIterator.spill();\n }\n return 0L; // this should throw exception\n }\n{code}\nin spill()","created":"2017-09-04T09:53:15.376+0000"},{"body":"Thank you for your report. Could you please attach a program that can reproduce this issue?","created":"2017-09-06T20:46:36.898+0000"},{"body":"[~kiszk] unfortunately I don't have a small scale repro. I hit it several times when running sql queries operating on several TB of data on a cluster with ~20 nodes / ~300 cores.\nLooking at the code, I'm quite sure it's caused by UnsafeInMemorySorter.reset:\n{code}\n consumer.freeArray(array);\n array = consumer.allocateArray(initialSize);\n{code}\nwhere allocating this array just after it was freed fails with another OOM, causing a nested spill.\nI think nested spilling is invalid, and thus acquiring memory in a way that can cause spill is invalid on code paths that are already during spilling.","created":"2017-09-06T21:08:50.399+0000"},{"body":"If you cannot provide a repro, could you please run your program with the latest master branch?\nSPARK-21319 may alleviate this issue.","created":"2017-09-08T17:42:14.264+0000"},{"body":"[~juliuszsompolski], I've followed the stack trace you've attached (as much as it's possible with master's code) and I tend to agree with your assumption about UnsafeInMemorySorter.reset.\nit seems that allocateArray did fail on OOM and triggered a nested spill, the thing is that by this point array points to an already freed block, the first time array is actually accessed (TimSort invokes the comparator) it fails on an NPE (assuming asserts are really turned off on your env).\n\ni think the way to solve this is by temporarily setting array (and the relevant pos, capacity, etc) to null/zero/some other value indicating a currently unreadable/empty buffer.\n\n[~kiszk], what do you think?","created":"2017-09-09T19:45:51.103+0000"},{"body":"User 'eyalfa' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19181","created":"2017-09-10T14:46:04.989+0000"},{"body":"opened PR: https://github.com/apache/spark/pull/19181","created":"2017-09-10T14:46:13.677+0000"},{"body":"User 'eyalfa' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19481","created":"2017-10-12T09:36:07.056+0000"},{"body":"Hi, [~hvanhovell] and [~eyalfa].\r\nI added `2.2.1` in fixed versions since it's merged into `branch-2.2` today.","created":"2017-10-12T20:10:26.100+0000"}],"conversations":[{"body":"I see NPE during sorting with the following stacktrace:\n{code}\njava.lang.NullPointerException\n\tat org.apache.spark.memory.TaskMemoryManager.getPage(TaskMemoryManager.java:383)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeInMemorySorter$SortComparator.compare(UnsafeInMemorySorter.java:63)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeInMemorySorter$SortComparator.compare(UnsafeInMemorySorter.java:43)\n\tat org.apache.spark.util.collection.TimSort.countRunAndMakeAscending(TimSort.java:270)\n\tat org.apache.spark.util.collection.TimSort.sort(TimSort.java:142)\n\tat org.apache.spark.util.collection.Sorter.sort(Sorter.scala:37)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeInMemorySorter.getSortedIterator(UnsafeInMemorySorter.java:345)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.spill(UnsafeExternalSorter.java:206)\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:203)\n\tat org.apache.spark.memory.TaskMemoryManager.allocatePage(TaskMemoryManager.java:281)\n\tat org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:90)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeInMemorySorter.reset(UnsafeInMemorySorter.java:173)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.spill(UnsafeExternalSorter.java:221)\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:203)\n\tat org.apache.spark.memory.TaskMemoryManager.allocatePage(TaskMemoryManager.java:281)\n\tat org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:90)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.growPointerArrayIfNecessary(UnsafeExternalSorter.java:349)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.insertRecord(UnsafeExternalSorter.java:400)\n\tat org.apache.spark.sql.execution.UnsafeExternalRowSorter.insertRow(UnsafeExternalRowSorter.java:109)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.sort_addToSorter$(Unknown Source)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n\tat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:395)\n\tat org.apache.spark.sql.execution.RowIteratorFromScala.advanceNext(RowIterator.scala:83)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinScanner.advancedStreamed(SortMergeJoinExec.scala:778)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinScanner.findNextInnerJoinRows(SortMergeJoinExec.scala:685)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$doExecute$1$$anon$2.advanceNext(SortMergeJoinExec.scala:259)\n\tat org.apache.spark.sql.execution.RowIteratorToScala.hasNext(RowIterator.scala:68)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.agg_doAggregateWithKeys$(Unknown Source)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n\tat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:395)\n\tat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n\tat org.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java:125)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:108)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:346)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:748)\n{code}\n","from":"reporter","subject":"NullPointerException in UnsafeExternalSorter.spill()"},{"body":"Note that UnsafeExternalSorter.spill appears twice on the stack trace, so it's nested spilling: the first triggered spilling triggers another spilling through UnsafeInMemorySorter.reset.\n\nPossibly it's messing up something by nested-spilling itself twice?\nOr messing something with\n{code:java}\n if (trigger != this) {\n if (readingIterator != null) {\n return readingIterator.spill();\n }\n return 0L; // this should throw exception\n }\n{code}\nin spill()","from":"developer"},{"body":"Thank you for your report. Could you please attach a program that can reproduce this issue?","from":"developer"},{"body":"[~kiszk] unfortunately I don't have a small scale repro. I hit it several times when running sql queries operating on several TB of data on a cluster with ~20 nodes / ~300 cores.\nLooking at the code, I'm quite sure it's caused by UnsafeInMemorySorter.reset:\n{code}\n consumer.freeArray(array);\n array = consumer.allocateArray(initialSize);\n{code}\nwhere allocating this array just after it was freed fails with another OOM, causing a nested spill.\nI think nested spilling is invalid, and thus acquiring memory in a way that can cause spill is invalid on code paths that are already during spilling.","from":"developer"},{"body":"If you cannot provide a repro, could you please run your program with the latest master branch?\nSPARK-21319 may alleviate this issue.","from":"developer"},{"body":"[~juliuszsompolski], I've followed the stack trace you've attached (as much as it's possible with master's code) and I tend to agree with your assumption about UnsafeInMemorySorter.reset.\nit seems that allocateArray did fail on OOM and triggered a nested spill, the thing is that by this point array points to an already freed block, the first time array is actually accessed (TimSort invokes the comparator) it fails on an NPE (assuming asserts are really turned off on your env).\n\ni think the way to solve this is by temporarily setting array (and the relevant pos, capacity, etc) to null/zero/some other value indicating a currently unreadable/empty buffer.\n\n[~kiszk], what do you think?","from":"developer"},{"body":"User 'eyalfa' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19181","from":"developer"},{"body":"opened PR: https://github.com/apache/spark/pull/19181","from":"developer"},{"body":"User 'eyalfa' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/19481","from":"developer"},{"body":"Hi, [~hvanhovell] and [~eyalfa].\r\nI added `2.2.1` in fixed versions since it's merged into `branch-2.2` today.","from":"developer"}],"created":"2017-09-04T09:48:44.000+0000","description":"I see NPE during sorting with the following stacktrace:\n{code}\njava.lang.NullPointerException\n\tat org.apache.spark.memory.TaskMemoryManager.getPage(TaskMemoryManager.java:383)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeInMemorySorter$SortComparator.compare(UnsafeInMemorySorter.java:63)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeInMemorySorter$SortComparator.compare(UnsafeInMemorySorter.java:43)\n\tat org.apache.spark.util.collection.TimSort.countRunAndMakeAscending(TimSort.java:270)\n\tat org.apache.spark.util.collection.TimSort.sort(TimSort.java:142)\n\tat org.apache.spark.util.collection.Sorter.sort(Sorter.scala:37)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeInMemorySorter.getSortedIterator(UnsafeInMemorySorter.java:345)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.spill(UnsafeExternalSorter.java:206)\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:203)\n\tat org.apache.spark.memory.TaskMemoryManager.allocatePage(TaskMemoryManager.java:281)\n\tat org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:90)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeInMemorySorter.reset(UnsafeInMemorySorter.java:173)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.spill(UnsafeExternalSorter.java:221)\n\tat org.apache.spark.memory.TaskMemoryManager.acquireExecutionMemory(TaskMemoryManager.java:203)\n\tat org.apache.spark.memory.TaskMemoryManager.allocatePage(TaskMemoryManager.java:281)\n\tat org.apache.spark.memory.MemoryConsumer.allocateArray(MemoryConsumer.java:90)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.growPointerArrayIfNecessary(UnsafeExternalSorter.java:349)\n\tat org.apache.spark.util.collection.unsafe.sort.UnsafeExternalSorter.insertRecord(UnsafeExternalSorter.java:400)\n\tat org.apache.spark.sql.execution.UnsafeExternalRowSorter.insertRow(UnsafeExternalRowSorter.java:109)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.sort_addToSorter$(Unknown Source)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n\tat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:395)\n\tat org.apache.spark.sql.execution.RowIteratorFromScala.advanceNext(RowIterator.scala:83)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinScanner.advancedStreamed(SortMergeJoinExec.scala:778)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinScanner.findNextInnerJoinRows(SortMergeJoinExec.scala:685)\n\tat org.apache.spark.sql.execution.joins.SortMergeJoinExec$$anonfun$doExecute$1$$anon$2.advanceNext(SortMergeJoinExec.scala:259)\n\tat org.apache.spark.sql.execution.RowIteratorToScala.hasNext(RowIterator.scala:68)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.agg_doAggregateWithKeys$(Unknown Source)\n\tat org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)\n\tat org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)\n\tat org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:395)\n\tat scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)\n\tat org.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java:125)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)\n\tat org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)\n\tat org.apache.spark.scheduler.Task.run(Task.scala:108)\n\tat org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:346)\n\tat java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)\n\tat java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)\n\tat java.lang.Thread.run(Thread.java:748)\n{code}\n","issue_id":"13099575","key":"SPARK-21907","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2017-10-10T20:58:10.000+0000","role":"gold_target","summary":"NullPointerException in UnsafeExternalSorter.spill()"} {"case_id":"12777649","cluster":"JIRA-SPARK-34e8f81d6695","comments":[{"body":"User 'piaozhexiu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4771","created":"2015-02-25T19:38:56.811+0000"},{"body":"Although a fix is possible for Hadoop 2.2+, it is not clear there is any way to avoid a race with HDFS's shutdown hook before that. It would be moderately painful to solve this with reflection, and probably not worth it. This can be resolved with the approach in the PR above for 2.2+.","created":"2015-03-01T08:57:16.922+0000"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/5560","created":"2015-04-17T18:44:01.405+0000"},{"body":"Issue resolved by pull request 5560\n[https://github.com/apache/spark/pull/5560]","created":"2015-04-22T00:34:06.922+0000"},{"body":"User 'nishkamravi2' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/5672","created":"2015-04-24T01:03:09.117+0000"}],"conversations":[{"body":"This is a regression of SPARK-2261. In branch-1.3 and master, {{EventLoggingListener}} throws \"{{java.io.IOException: Filesystem closed}}\" when ctrl+c or ctrl+d the spark-sql shell.\n\nThe root cause is that DFSClient is already shut down before EventLoggingListener invokes the following HDFS methods, and thus, DFSClient.isClientRunning() check fails-\n{code}\nLine #135: hadoopDataStream.foreach(hadoopFlushMethod.invoke(_))\nLine #187: if (fileSystem.exists(target)) {\n{code}\nThe followings are full stack trace-\n{code}\njava.lang.reflect.InvocationTargetException\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:606)\n\tat org.apache.spark.scheduler.EventLoggingListener$$anonfun$logEvent$3.apply(EventLoggingListener.scala:135)\n\tat org.apache.spark.scheduler.EventLoggingListener$$anonfun$logEvent$3.apply(EventLoggingListener.scala:135)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.EventLoggingListener.logEvent(EventLoggingListener.scala:135)\n\tat org.apache.spark.scheduler.EventLoggingListener.onApplicationEnd(EventLoggingListener.scala:170)\n\tat org.apache.spark.scheduler.SparkListenerBus$class.onPostEvent(SparkListenerBus.scala:54)\n\tat org.apache.spark.scheduler.LiveListenerBus.onPostEvent(LiveListenerBus.scala:31)\n\tat org.apache.spark.scheduler.LiveListenerBus.onPostEvent(LiveListenerBus.scala:31)\n\tat org.apache.spark.util.ListenerBus$class.postToAll(ListenerBus.scala:53)\n\tat org.apache.spark.util.AsynchronousListenerBus.postToAll(AsynchronousListenerBus.scala:36)\n\tat org.apache.spark.util.AsynchronousListenerBus$$anon$1$$anonfun$run$1.apply$mcV$sp(AsynchronousListenerBus.scala:76)\n\tat org.apache.spark.util.AsynchronousListenerBus$$anon$1$$anonfun$run$1.apply(AsynchronousListenerBus.scala:61)\n\tat org.apache.spark.util.AsynchronousListenerBus$$anon$1$$anonfun$run$1.apply(AsynchronousListenerBus.scala:61)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1613)\n\tat org.apache.spark.util.AsynchronousListenerBus$$anon$1.run(AsynchronousListenerBus.scala:60)\nCaused by: java.io.IOException: Filesystem closed\n\tat org.apache.hadoop.hdfs.DFSClient.checkOpen(DFSClient.java:707)\n\tat org.apache.hadoop.hdfs.DFSOutputStream.flushOrSync(DFSOutputStream.java:1843)\n\tat org.apache.hadoop.hdfs.DFSOutputStream.hflush(DFSOutputStream.java:1804)\n\tat org.apache.hadoop.fs.FSDataOutputStream.hflush(FSDataOutputStream.java:127)\n\t... 19 more\n{code}\n{code}\nException in thread \"Thread-3\" java.io.IOException: Filesystem closed\n\tat org.apache.hadoop.hdfs.DFSClient.checkOpen(DFSClient.java:707)\n\tat org.apache.hadoop.hdfs.DFSClient.getFileInfo(DFSClient.java:1760)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem$17.doCall(DistributedFileSystem.java:1124)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem$17.doCall(DistributedFileSystem.java:1120)\n\tat org.apache.hadoop.fs.FileSystemLinkResolver.resolve(FileSystemLinkResolver.java:81)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem.getFileStatus(DistributedFileSystem.java:1120)\n\tat org.apache.hadoop.fs.FileSystem.exists(FileSystem.java:1398)\n\tat org.apache.spark.scheduler.EventLoggingListener.stop(EventLoggingListener.scala:187)\n\tat org.apache.spark.SparkContext$$anonfun$stop$4.apply(SparkContext.scala:1379)\n\tat org.apache.spark.SparkContext$$anonfun$stop$4.apply(SparkContext.scala:1379)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.SparkContext.stop(SparkContext.scala:1379)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLEnv$.stop(SparkSQLEnv.scala:66)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$$anon$1.run(SparkSQLCLIDriver.scala:107)\n{code}\n","from":"reporter","subject":"java.io.IOException: Filesystem is thrown when ctrl+c or ctrl+d spark-sql on YARN "},{"body":"User 'piaozhexiu' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/4771","from":"developer"},{"body":"Although a fix is possible for Hadoop 2.2+, it is not clear there is any way to avoid a race with HDFS's shutdown hook before that. It would be moderately painful to solve this with reflection, and probably not worth it. This can be resolved with the approach in the PR above for 2.2+.","from":"developer"},{"body":"User 'vanzin' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/5560","from":"developer"},{"body":"Issue resolved by pull request 5560\n[https://github.com/apache/spark/pull/5560]","from":"developer"},{"body":"User 'nishkamravi2' has created a pull request for this issue:\nhttps://github.com/apache/spark/pull/5672","from":"developer"}],"created":"2015-02-25T19:35:03.000+0000","description":"This is a regression of SPARK-2261. In branch-1.3 and master, {{EventLoggingListener}} throws \"{{java.io.IOException: Filesystem closed}}\" when ctrl+c or ctrl+d the spark-sql shell.\n\nThe root cause is that DFSClient is already shut down before EventLoggingListener invokes the following HDFS methods, and thus, DFSClient.isClientRunning() check fails-\n{code}\nLine #135: hadoopDataStream.foreach(hadoopFlushMethod.invoke(_))\nLine #187: if (fileSystem.exists(target)) {\n{code}\nThe followings are full stack trace-\n{code}\njava.lang.reflect.InvocationTargetException\n\tat sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)\n\tat sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:57)\n\tat sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)\n\tat java.lang.reflect.Method.invoke(Method.java:606)\n\tat org.apache.spark.scheduler.EventLoggingListener$$anonfun$logEvent$3.apply(EventLoggingListener.scala:135)\n\tat org.apache.spark.scheduler.EventLoggingListener$$anonfun$logEvent$3.apply(EventLoggingListener.scala:135)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.scheduler.EventLoggingListener.logEvent(EventLoggingListener.scala:135)\n\tat org.apache.spark.scheduler.EventLoggingListener.onApplicationEnd(EventLoggingListener.scala:170)\n\tat org.apache.spark.scheduler.SparkListenerBus$class.onPostEvent(SparkListenerBus.scala:54)\n\tat org.apache.spark.scheduler.LiveListenerBus.onPostEvent(LiveListenerBus.scala:31)\n\tat org.apache.spark.scheduler.LiveListenerBus.onPostEvent(LiveListenerBus.scala:31)\n\tat org.apache.spark.util.ListenerBus$class.postToAll(ListenerBus.scala:53)\n\tat org.apache.spark.util.AsynchronousListenerBus.postToAll(AsynchronousListenerBus.scala:36)\n\tat org.apache.spark.util.AsynchronousListenerBus$$anon$1$$anonfun$run$1.apply$mcV$sp(AsynchronousListenerBus.scala:76)\n\tat org.apache.spark.util.AsynchronousListenerBus$$anon$1$$anonfun$run$1.apply(AsynchronousListenerBus.scala:61)\n\tat org.apache.spark.util.AsynchronousListenerBus$$anon$1$$anonfun$run$1.apply(AsynchronousListenerBus.scala:61)\n\tat org.apache.spark.util.Utils$.logUncaughtExceptions(Utils.scala:1613)\n\tat org.apache.spark.util.AsynchronousListenerBus$$anon$1.run(AsynchronousListenerBus.scala:60)\nCaused by: java.io.IOException: Filesystem closed\n\tat org.apache.hadoop.hdfs.DFSClient.checkOpen(DFSClient.java:707)\n\tat org.apache.hadoop.hdfs.DFSOutputStream.flushOrSync(DFSOutputStream.java:1843)\n\tat org.apache.hadoop.hdfs.DFSOutputStream.hflush(DFSOutputStream.java:1804)\n\tat org.apache.hadoop.fs.FSDataOutputStream.hflush(FSDataOutputStream.java:127)\n\t... 19 more\n{code}\n{code}\nException in thread \"Thread-3\" java.io.IOException: Filesystem closed\n\tat org.apache.hadoop.hdfs.DFSClient.checkOpen(DFSClient.java:707)\n\tat org.apache.hadoop.hdfs.DFSClient.getFileInfo(DFSClient.java:1760)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem$17.doCall(DistributedFileSystem.java:1124)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem$17.doCall(DistributedFileSystem.java:1120)\n\tat org.apache.hadoop.fs.FileSystemLinkResolver.resolve(FileSystemLinkResolver.java:81)\n\tat org.apache.hadoop.hdfs.DistributedFileSystem.getFileStatus(DistributedFileSystem.java:1120)\n\tat org.apache.hadoop.fs.FileSystem.exists(FileSystem.java:1398)\n\tat org.apache.spark.scheduler.EventLoggingListener.stop(EventLoggingListener.scala:187)\n\tat org.apache.spark.SparkContext$$anonfun$stop$4.apply(SparkContext.scala:1379)\n\tat org.apache.spark.SparkContext$$anonfun$stop$4.apply(SparkContext.scala:1379)\n\tat scala.Option.foreach(Option.scala:236)\n\tat org.apache.spark.SparkContext.stop(SparkContext.scala:1379)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLEnv$.stop(SparkSQLEnv.scala:66)\n\tat org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver$$anon$1.run(SparkSQLCLIDriver.scala:107)\n{code}\n","issue_id":"12777649","key":"SPARK-6014","metadata":{"source":"Apache Jira"},"project":"SPARK","resolution":"Fixed","resolution_date":"2015-04-22T00:34:06.000+0000","role":"gold_target","summary":"java.io.IOException: Filesystem is thrown when ctrl+c or ctrl+d spark-sql on YARN "} {"case_id":"12845388","cluster":"JIRA-SPARK-3d34cfc41158","comments":[{"body":"It's reasonable to have a field for limit in jdbc dialect. Would that just solve this problem? ","created":"2015-07-15T22:48:36.285+0000"},{"body":"If there were also a way to get tableExists to work correctly (even the 0=1 ugly method is perhaps better than nothing) for those JDBC sources that don't support LIMIT. \n\nAnd, because both current sources that have coded JdbcDialect entries (MySQL and Postgres) do support LIMIT, perhaps the default should be [no support] and MySQL and Postgres should be coded as [does support].\n\nThanks, Bob\n","created":"2015-07-16T00:47:29.416+0000"},{"body":"I'm having second thoughts, if a JdbcDialect entry for LIMIT support is considered, about making [no support] the default, as this may affect existing code that depends on the current implementation. \n\nUsing the same value for every JdbcDialect would effectively make this a no-op however, and, in future, you might consider allowing JdbcDialect settings (keyed by driver pattern, as it is currently) to be specified in a config file (or the existing config file, I couldn't find it in the current settings), so that adding a driver for a data source that uses a different dialect can be accomplished without a custom build. I can file a separate JIRA for that change if you consider it worthwhile.\n\nThanks, Bob","created":"2015-07-16T20:00:10.044+0000"},{"body":"That might make sense -- but in general it is pretty annoying to add config files. Adding a new dialect actually isn't that hard. \n\nShould we just move tableExists into the dialect itself?\n","created":"2015-07-20T00:48:22.683+0000"},{"body":"Could make it part of the JdbcDialect. Or a string in the JdbcDialect that is used by tableExists that returns at most 1 row. With the current configuration (\"SELECT 1 FROM $table LIMIT 1\") being the default. Then you wouldn't have to change how the query works to determine if the table exists. Your call...\n\nAbout the config file... my thinking here is that to add a new dialect (say if I wanted non-standard dialect entries for DB2 JDBC driver) requires a new release or a modified version of the base source. Unless there's a easy way to inject a new JdbcDialect that and without creating a whole new data source that I'm missing. Not all customers want to recompile the source (or run a modified version of the source) to support a new JDBC data source. And would require a database of that type to be available for build testing, as I'm guessing MySQL and Postgres are now. Much less overhead for the build/test team than another config file (or making it part of current a config file). Allowing separation of JDBC-specific driver settings from the base code would accomplish that without custom builds. But you know the target audience better than I do.\n\nCheers, and thanks, \nBob\n \n","created":"2015-07-20T19:30:25.679+0000"},{"body":"Right now you do have to call JdbcDialects.registerDialect, but we can make it available as an option to pass through to the jdbc data source. In that case, it should alleviate your concern about custom build right? All that's needed is to have a JdbcDialect implementation on the classpath.\n\n","created":"2015-07-21T06:18:01.728+0000"},{"body":"Great, I didn't realize that JdbcDialects.registerDialect was a public API, passing it through to the jdbc data source would do it. \n\nCheers, and thanks, Bob","created":"2015-07-21T17:36:02.811+0000"},{"body":"Want to submit a pull request? :)","created":"2015-07-21T17:36:26.639+0000"},{"body":"That was quick. Not sure that I have all the pieces in place for building right now, is it required? ;-) Currently I was just browsing the source code to figure out what would be required to add/use a new fully supported JDBC-based data source (how all the pieces worked) and came across the hardcoded SQL statement. ","created":"2015-07-21T17:49:47.445+0000"},{"body":"@Bob, Reynold\n\nI ran into the same issue when trying to write data frames into an existing table in DB2 database.\nif you are not working on the fix, I would like to give it a try. \n\nThank you for the analysis of the issue , my understanding is fix should do the following to address table exists problem with LIMIT syntax.:\n\n-- Add table Exists method to the JdbcDialect interface, and allow dialects override the method as required for specific databases.\n-- Default implementation of table exists method should use DatabaseMetaData.getTables() to find if table exists. If that particular interface is not implement use the query \"select 1 from $table where 1=0\".\n-- Add table exist method that use LIMIT query to MySQL , and Postgres dialects.\n\n* Enhancing registering of dialect : (I think this may have to be separate Jira to avoid confusion).\n\n@Reynold : I am not understanding your comment on adding option to pass through the jdbc data source. If you can give an example that will be great. \n\nAre you referring to some thing like the following ?\n df.write.option(\"datasource.jdbc.dialects\" \"org.apache.DerbyDialect\").jdbc(\"jdbc:deryby://